Files
SuperBizAgent-java/mvp/eval/README.md
T
2026-07-05 00:59:53 +08:00

62 lines
1.7 KiB
Markdown

# Diagnosis Eval Harness
This folder contains the first fixed-case evaluation set for the MVP diagnosis Agent.
## Scope
- Case definitions: `cases/diagnosis-cases.json`
- Offline trace fixtures: `fixtures/*.json`
- Field definitions: `schema.md`
- Baseline reports: `reports/baseline-report.json` and `reports/baseline-report.md`
- Baseline diff sample: `reports/baseline-diff-sample.json` and `reports/baseline-diff-sample.md`
- Evaluator implementation: `DiagnosisTraceEvaluator`
- Report writer: `DiagnosisEvalReportWriter`
## Current Mode
The first version evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.
## Verification
Run the focused evaluator test:
```powershell
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test
```
The committed baseline report represents the current fixed fixture set:
```text
5 fixed cases
5 passing fixture evaluations
2 PASS verdicts
3 LOW_CONFID verdicts
```
When fixtures or evaluator rules change, regenerate the report from the same case file and fixture directory, then update both JSON and Markdown outputs together.
## Interview Story
The harness gives the MVP a repeatable baseline:
```text
fixed diagnosis case
-> saved or runtime trace
-> rule-based trace validation
-> JSON / Markdown report
-> regression signal for prompts, tools, retrieval, and verifier behavior
```
## Baseline Diff
Baseline diff compares a current report against `reports/baseline-report.json`.
```text
baseline report
current report
-> deterministic diff
-> regressions, improvements, and changed signals
```
Use it to answer: did a prompt, tool, retrieval, or verifier change make the Agent worse than the fixed baseline?