32 lines
962 B
Markdown
32 lines
962 B
Markdown
# Brief: diagnosis-eval-harness
|
|
|
|
## Background
|
|
|
|
The MVP has a runnable demo and hardened evidence trace semantics, but it still lacks a fixed regression baseline for Agent diagnosis quality. P1-B creates a small evaluation harness that can validate diagnosis traces against fixed cases and produce repeatable reports.
|
|
|
|
## Goals
|
|
|
|
1. Define fixed diagnosis cases for the MVP demo domain.
|
|
2. Validate trace evidence, verifier verdicts, answer keywords, and degraded-output behavior.
|
|
3. Produce JSON and Markdown reports for interview and regression use.
|
|
4. Keep the first version offline by supporting trace fixtures.
|
|
|
|
## Scope
|
|
|
|
- Evaluation case definitions
|
|
- Trace fixture shape
|
|
- Rule-based evaluator
|
|
- JSON / Markdown report output
|
|
- Focused offline tests and docs
|
|
|
|
## Non-Goals
|
|
|
|
- No LLM-as-judge
|
|
- No live end-to-end runtime requirement
|
|
- No production API
|
|
- No chat or verifier runtime change
|
|
|
|
## Related OpenSpec
|
|
|
|
`openspec/changes/diagnosis-eval-harness/`
|