2.3 KiB
2.3 KiB
Context
The eval harness now has a complete five-case fixture baseline and saved JSON / Markdown baseline reports. The missing piece is a deterministic comparison step that explains whether a new report is better, worse, or just different from the baseline.
Goals / Non-Goals
Goals:
- Compare two
DiagnosisEvalReportobjects without requiring external services. - Surface aggregate regressions such as pass-rate drops, verdict distribution shifts, and cost increases.
- Surface per-case regressions such as pass-to-fail changes, missing evidence coverage, verdict changes, keyword coverage loss, and missing cases.
- Write JSON and Markdown diff outputs for review.
Non-Goals:
- Do not run the Agent or regenerate traces.
- Do not introduce LLM-as-judge.
- Do not change evaluator scoring rules.
- Do not block on performance thresholds beyond simple numeric diff signals.
Decisions
-
Decision: Compare report DTOs instead of raw traces.
- Reason:
DiagnosisEvalReportis already the stable structured output of the evaluator and is cheaper to diff than trace internals. - Alternative considered: compare raw trace fixtures. That would expose more detail but duplicate evaluator responsibilities.
- Reason:
-
Decision: Classify each diff item as
REGRESSION,IMPROVEMENT, orCHANGED.- Reason: interview and CI usage both need a quick answer to "did this get worse?" while still preserving neutral changes.
- Alternative considered: only output numeric deltas. That is harder to scan and less actionable.
-
Decision: Keep thresholds explicit and conservative.
- Reason: pass/fail and missing evidence are hard regressions; tool calls and duration are cost signals that should be visible even if not always blocking.
- Alternative considered: fail only on pass-rate drop. That misses cases where quality stays green but cost or confidence behavior changes.
Risks / Trade-offs
- Report comparison can only see fields already captured by
DiagnosisEvalReport. Mitigation: use this as the first regression layer and add richer report fields later if needed. - Duration may fluctuate in live runs. Mitigation: fixture baseline uses stable durations; live-mode thresholds can be added later.
- Verdict distribution changes can be intentional. Mitigation: classify them as
CHANGEDunless they coincide with per-case regressions.