Add diagnosis eval baseline diff
This commit is contained in:
@@ -8,6 +8,7 @@ This folder contains the first fixed-case evaluation set for the MVP diagnosis A
|
||||
- Offline trace fixtures: `fixtures/*.json`
|
||||
- Field definitions: `schema.md`
|
||||
- Baseline reports: `reports/baseline-report.json` and `reports/baseline-report.md`
|
||||
- Baseline diff sample: `reports/baseline-diff-sample.json` and `reports/baseline-diff-sample.md`
|
||||
- Evaluator implementation: `DiagnosisTraceEvaluator`
|
||||
- Report writer: `DiagnosisEvalReportWriter`
|
||||
|
||||
@@ -45,3 +46,16 @@ fixed diagnosis case
|
||||
-> JSON / Markdown report
|
||||
-> regression signal for prompts, tools, retrieval, and verifier behavior
|
||||
```
|
||||
|
||||
## Baseline Diff
|
||||
|
||||
Baseline diff compares a current report against `reports/baseline-report.json`.
|
||||
|
||||
```text
|
||||
baseline report
|
||||
current report
|
||||
-> deterministic diff
|
||||
-> regressions, improvements, and changed signals
|
||||
```
|
||||
|
||||
Use it to answer: did a prompt, tool, retrieval, or verifier change make the Agent worse than the fixed baseline?
|
||||
|
||||
Reference in New Issue
Block a user