962 B
962 B
Brief: diagnosis-eval-harness
Background
The MVP has a runnable demo and hardened evidence trace semantics, but it still lacks a fixed regression baseline for Agent diagnosis quality. P1-B creates a small evaluation harness that can validate diagnosis traces against fixed cases and produce repeatable reports.
Goals
- Define fixed diagnosis cases for the MVP demo domain.
- Validate trace evidence, verifier verdicts, answer keywords, and degraded-output behavior.
- Produce JSON and Markdown reports for interview and regression use.
- Keep the first version offline by supporting trace fixtures.
Scope
- Evaluation case definitions
- Trace fixture shape
- Rule-based evaluator
- JSON / Markdown report output
- Focused offline tests and docs
Non-Goals
- No LLM-as-judge
- No live end-to-end runtime requirement
- No production API
- No chat or verifier runtime change
Related OpenSpec
openspec/changes/diagnosis-eval-harness/