Expand diagnosis eval fixtures
This commit is contained in:
@@ -0,0 +1,27 @@
|
||||
## Why
|
||||
|
||||
The diagnosis evaluation harness is implemented, but the baseline is still incomplete because only two of the five fixed cases have trace fixtures. Completing the fixture set and saving a baseline report makes the harness useful as a practical regression signal for interview demos and future Agent changes.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add trace fixtures for the remaining fixed diagnosis cases: Redis timeout, slow response, and JVM memory risk.
|
||||
- Add a reproducible baseline report generated from the full fixture set.
|
||||
- Document how to regenerate and interpret the baseline.
|
||||
- Keep the evaluator deterministic and offline; no live MySQL, Redis, Milvus, or LLM service is required.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `diagnosis-eval-harness`: Extend the existing evaluation harness requirement so the fixed MVP case set has complete fixture coverage and a saved baseline report.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affects `mvp/eval/cases`, `mvp/eval/fixtures`, and eval documentation.
|
||||
- May add baseline output files under `mvp/eval/reports`.
|
||||
- May add or update focused evaluator tests to assert full fixture coverage and report generation.
|
||||
- No production runtime API or database schema changes are expected.
|
||||
Reference in New Issue
Block a user