Expand diagnosis eval fixtures
This commit is contained in:
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,39 @@
|
||||
## Context
|
||||
|
||||
`diagnosis-eval-harness` already provides fixed case definitions, fixture-mode evaluation, JSON / Markdown report writing, and focused evaluator tests. The current baseline is incomplete because three fixed cases intentionally point to missing fixtures.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Add representative trace fixtures for every fixed diagnosis case.
|
||||
- Save a baseline report that can be reviewed and compared after future Agent changes.
|
||||
- Keep the baseline reproducible in offline mode.
|
||||
- Document how to regenerate the baseline.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not change production Agent runtime behavior.
|
||||
- Do not require live infrastructure or a real LLM.
|
||||
- Do not introduce a new LLM-based grader.
|
||||
- Do not expand the case set beyond the existing five fixed MVP diagnosis cases.
|
||||
|
||||
## Decisions
|
||||
|
||||
- Use checked-in fixture traces instead of live service calls.
|
||||
- Rationale: the goal is a stable regression baseline that can run in CI or interview environments without external dependencies.
|
||||
- Alternative considered: start the application and call the trace API. That is useful later, but it introduces infrastructure noise before the baseline is complete.
|
||||
|
||||
- Save baseline reports under `mvp/eval/reports`.
|
||||
- Rationale: reports are reviewable artifacts, not transient build output, and they show the expected current behavior of the baseline.
|
||||
- Alternative considered: generate reports only in tests. That verifies behavior but does not give an easy artifact to show or diff.
|
||||
|
||||
- Keep fixture outcomes representative rather than forcing every case to pass.
|
||||
- Rationale: a baseline should reflect expected behavior, including low-confidence or degraded cases, as long as the outcome is explicit and stable.
|
||||
- Alternative considered: make every fixture pass. That looks cleaner but hides important degraded-path behavior.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- Fixture data can drift from real runtime traces. Mitigation: keep fixtures shaped like `DiagnosisTraceResponse` and add tests that load every referenced fixture.
|
||||
- A saved baseline report can become stale after intentional rule changes. Mitigation: document regeneration steps and update the report in the same change as rule or fixture updates.
|
||||
- Keyword-based checks are coarse. Mitigation: this change keeps the deterministic harness simple and leaves semantic scoring as a later improvement.
|
||||
@@ -0,0 +1,27 @@
|
||||
## Why
|
||||
|
||||
The diagnosis evaluation harness is implemented, but the baseline is still incomplete because only two of the five fixed cases have trace fixtures. Completing the fixture set and saving a baseline report makes the harness useful as a practical regression signal for interview demos and future Agent changes.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add trace fixtures for the remaining fixed diagnosis cases: Redis timeout, slow response, and JVM memory risk.
|
||||
- Add a reproducible baseline report generated from the full fixture set.
|
||||
- Document how to regenerate and interpret the baseline.
|
||||
- Keep the evaluator deterministic and offline; no live MySQL, Redis, Milvus, or LLM service is required.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `diagnosis-eval-harness`: Extend the existing evaluation harness requirement so the fixed MVP case set has complete fixture coverage and a saved baseline report.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affects `mvp/eval/cases`, `mvp/eval/fixtures`, and eval documentation.
|
||||
- May add baseline output files under `mvp/eval/reports`.
|
||||
- May add or update focused evaluator tests to assert full fixture coverage and report generation.
|
||||
- No production runtime API or database schema changes are expected.
|
||||
+27
@@ -0,0 +1,27 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Evaluation harness SHALL provide complete fixture coverage for fixed cases
|
||||
The system SHALL include an offline trace fixture for every fixed MVP diagnosis evaluation case.
|
||||
|
||||
#### Scenario: Every case resolves to a fixture file
|
||||
- **WHEN** the evaluator loads the fixed case definition file
|
||||
- **THEN** every case's `traceFixture` value SHALL resolve to an existing JSON fixture file
|
||||
|
||||
#### Scenario: Fixture files are loadable as diagnosis traces
|
||||
- **WHEN** each referenced fixture is loaded
|
||||
- **THEN** it SHALL deserialize into the trace response shape used by the evaluator
|
||||
|
||||
### Requirement: Evaluation harness SHALL preserve a reproducible baseline report
|
||||
The system SHALL preserve a generated baseline report for the full fixed fixture set.
|
||||
|
||||
#### Scenario: Baseline report includes all fixed cases
|
||||
- **WHEN** the baseline report is generated from the fixed case file and fixture directory
|
||||
- **THEN** the report SHALL include one result for every fixed case
|
||||
|
||||
#### Scenario: Baseline report is reviewable
|
||||
- **WHEN** the baseline report is written
|
||||
- **THEN** it SHALL be available in JSON and Markdown formats under the eval documentation area
|
||||
|
||||
#### Scenario: Baseline regeneration is documented
|
||||
- **WHEN** a developer changes fixtures or evaluator rules
|
||||
- **THEN** the eval documentation SHALL explain how to regenerate the baseline report
|
||||
@@ -0,0 +1,19 @@
|
||||
## 1. Fixture Coverage
|
||||
|
||||
- [x] 1.1 Add Redis timeout trace fixture referenced by the fixed case file.
|
||||
- [x] 1.2 Add slow response trace fixture referenced by the fixed case file.
|
||||
- [x] 1.3 Add JVM memory risk trace fixture referenced by the fixed case file.
|
||||
- [x] 1.4 Verify every `traceFixture` in `diagnosis-cases.json` resolves to an existing fixture file.
|
||||
|
||||
## 2. Baseline Reports
|
||||
|
||||
- [x] 2.1 Generate a full baseline JSON report for all fixed cases.
|
||||
- [x] 2.2 Generate a full baseline Markdown report for review.
|
||||
- [x] 2.3 Document how to regenerate and interpret the baseline reports.
|
||||
|
||||
## 3. Tests And Validation
|
||||
|
||||
- [x] 3.1 Add or update focused tests for full fixture coverage and baseline report generation.
|
||||
- [x] 3.2 Run evaluator tests.
|
||||
- [x] 3.3 Run compile verification.
|
||||
- [x] 3.4 Run OpenSpec validation.
|
||||
Reference in New Issue
Block a user