62 lines
1.7 KiB
Markdown
62 lines
1.7 KiB
Markdown
# Diagnosis Eval Harness
|
|
|
|
This folder contains the first fixed-case evaluation set for the MVP diagnosis Agent.
|
|
|
|
## Scope
|
|
|
|
- Case definitions: `cases/diagnosis-cases.json`
|
|
- Offline trace fixtures: `fixtures/*.json`
|
|
- Field definitions: `schema.md`
|
|
- Baseline reports: `reports/baseline-report.json` and `reports/baseline-report.md`
|
|
- Baseline diff sample: `reports/baseline-diff-sample.json` and `reports/baseline-diff-sample.md`
|
|
- Evaluator implementation: `DiagnosisTraceEvaluator`
|
|
- Report writer: `DiagnosisEvalReportWriter`
|
|
|
|
## Current Mode
|
|
|
|
The first version evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.
|
|
|
|
## Verification
|
|
|
|
Run the focused evaluator test:
|
|
|
|
```powershell
|
|
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test
|
|
```
|
|
|
|
The committed baseline report represents the current fixed fixture set:
|
|
|
|
```text
|
|
5 fixed cases
|
|
5 passing fixture evaluations
|
|
2 PASS verdicts
|
|
3 LOW_CONFID verdicts
|
|
```
|
|
|
|
When fixtures or evaluator rules change, regenerate the report from the same case file and fixture directory, then update both JSON and Markdown outputs together.
|
|
|
|
## Interview Story
|
|
|
|
The harness gives the MVP a repeatable baseline:
|
|
|
|
```text
|
|
fixed diagnosis case
|
|
-> saved or runtime trace
|
|
-> rule-based trace validation
|
|
-> JSON / Markdown report
|
|
-> regression signal for prompts, tools, retrieval, and verifier behavior
|
|
```
|
|
|
|
## Baseline Diff
|
|
|
|
Baseline diff compares a current report against `reports/baseline-report.json`.
|
|
|
|
```text
|
|
baseline report
|
|
current report
|
|
-> deterministic diff
|
|
-> regressions, improvements, and changed signals
|
|
```
|
|
|
|
Use it to answer: did a prompt, tool, retrieval, or verifier change make the Agent worse than the fixed baseline?
|