Expand diagnosis eval fixtures

This commit is contained in:
aruo
2026-07-05 00:27:57 +08:00
parent ca5c61fabf
commit 4c7c53b024
21 changed files with 622 additions and 11 deletions
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-04
@@ -0,0 +1,39 @@
## Context
`diagnosis-eval-harness` already provides fixed case definitions, fixture-mode evaluation, JSON / Markdown report writing, and focused evaluator tests. The current baseline is incomplete because three fixed cases intentionally point to missing fixtures.
## Goals / Non-Goals
**Goals:**
- Add representative trace fixtures for every fixed diagnosis case.
- Save a baseline report that can be reviewed and compared after future Agent changes.
- Keep the baseline reproducible in offline mode.
- Document how to regenerate the baseline.
**Non-Goals:**
- Do not change production Agent runtime behavior.
- Do not require live infrastructure or a real LLM.
- Do not introduce a new LLM-based grader.
- Do not expand the case set beyond the existing five fixed MVP diagnosis cases.
## Decisions
- Use checked-in fixture traces instead of live service calls.
- Rationale: the goal is a stable regression baseline that can run in CI or interview environments without external dependencies.
- Alternative considered: start the application and call the trace API. That is useful later, but it introduces infrastructure noise before the baseline is complete.
- Save baseline reports under `mvp/eval/reports`.
- Rationale: reports are reviewable artifacts, not transient build output, and they show the expected current behavior of the baseline.
- Alternative considered: generate reports only in tests. That verifies behavior but does not give an easy artifact to show or diff.
- Keep fixture outcomes representative rather than forcing every case to pass.
- Rationale: a baseline should reflect expected behavior, including low-confidence or degraded cases, as long as the outcome is explicit and stable.
- Alternative considered: make every fixture pass. That looks cleaner but hides important degraded-path behavior.
## Risks / Trade-offs
- Fixture data can drift from real runtime traces. Mitigation: keep fixtures shaped like `DiagnosisTraceResponse` and add tests that load every referenced fixture.
- A saved baseline report can become stale after intentional rule changes. Mitigation: document regeneration steps and update the report in the same change as rule or fixture updates.
- Keyword-based checks are coarse. Mitigation: this change keeps the deterministic harness simple and leaves semantic scoring as a later improvement.
@@ -0,0 +1,27 @@
## Why
The diagnosis evaluation harness is implemented, but the baseline is still incomplete because only two of the five fixed cases have trace fixtures. Completing the fixture set and saving a baseline report makes the harness useful as a practical regression signal for interview demos and future Agent changes.
## What Changes
- Add trace fixtures for the remaining fixed diagnosis cases: Redis timeout, slow response, and JVM memory risk.
- Add a reproducible baseline report generated from the full fixture set.
- Document how to regenerate and interpret the baseline.
- Keep the evaluator deterministic and offline; no live MySQL, Redis, Milvus, or LLM service is required.
## Capabilities
### New Capabilities
- None.
### Modified Capabilities
- `diagnosis-eval-harness`: Extend the existing evaluation harness requirement so the fixed MVP case set has complete fixture coverage and a saved baseline report.
## Impact
- Affects `mvp/eval/cases`, `mvp/eval/fixtures`, and eval documentation.
- May add baseline output files under `mvp/eval/reports`.
- May add or update focused evaluator tests to assert full fixture coverage and report generation.
- No production runtime API or database schema changes are expected.
@@ -0,0 +1,27 @@
## ADDED Requirements
### Requirement: Evaluation harness SHALL provide complete fixture coverage for fixed cases
The system SHALL include an offline trace fixture for every fixed MVP diagnosis evaluation case.
#### Scenario: Every case resolves to a fixture file
- **WHEN** the evaluator loads the fixed case definition file
- **THEN** every case's `traceFixture` value SHALL resolve to an existing JSON fixture file
#### Scenario: Fixture files are loadable as diagnosis traces
- **WHEN** each referenced fixture is loaded
- **THEN** it SHALL deserialize into the trace response shape used by the evaluator
### Requirement: Evaluation harness SHALL preserve a reproducible baseline report
The system SHALL preserve a generated baseline report for the full fixed fixture set.
#### Scenario: Baseline report includes all fixed cases
- **WHEN** the baseline report is generated from the fixed case file and fixture directory
- **THEN** the report SHALL include one result for every fixed case
#### Scenario: Baseline report is reviewable
- **WHEN** the baseline report is written
- **THEN** it SHALL be available in JSON and Markdown formats under the eval documentation area
#### Scenario: Baseline regeneration is documented
- **WHEN** a developer changes fixtures or evaluator rules
- **THEN** the eval documentation SHALL explain how to regenerate the baseline report
@@ -0,0 +1,19 @@
## 1. Fixture Coverage
- [x] 1.1 Add Redis timeout trace fixture referenced by the fixed case file.
- [x] 1.2 Add slow response trace fixture referenced by the fixed case file.
- [x] 1.3 Add JVM memory risk trace fixture referenced by the fixed case file.
- [x] 1.4 Verify every `traceFixture` in `diagnosis-cases.json` resolves to an existing fixture file.
## 2. Baseline Reports
- [x] 2.1 Generate a full baseline JSON report for all fixed cases.
- [x] 2.2 Generate a full baseline Markdown report for review.
- [x] 2.3 Document how to regenerate and interpret the baseline reports.
## 3. Tests And Validation
- [x] 3.1 Add or update focused tests for full fixture coverage and baseline report generation.
- [x] 3.2 Run evaluator tests.
- [x] 3.3 Run compile verification.
- [x] 3.4 Run OpenSpec validation.