## Context The eval harness now has a complete five-case fixture baseline and saved JSON / Markdown baseline reports. The missing piece is a deterministic comparison step that explains whether a new report is better, worse, or just different from the baseline. ## Goals / Non-Goals **Goals:** - Compare two `DiagnosisEvalReport` objects without requiring external services. - Surface aggregate regressions such as pass-rate drops, verdict distribution shifts, and cost increases. - Surface per-case regressions such as pass-to-fail changes, missing evidence coverage, verdict changes, keyword coverage loss, and missing cases. - Write JSON and Markdown diff outputs for review. **Non-Goals:** - Do not run the Agent or regenerate traces. - Do not introduce LLM-as-judge. - Do not change evaluator scoring rules. - Do not block on performance thresholds beyond simple numeric diff signals. ## Decisions - Decision: Compare report DTOs instead of raw traces. - Reason: `DiagnosisEvalReport` is already the stable structured output of the evaluator and is cheaper to diff than trace internals. - Alternative considered: compare raw trace fixtures. That would expose more detail but duplicate evaluator responsibilities. - Decision: Classify each diff item as `REGRESSION`, `IMPROVEMENT`, or `CHANGED`. - Reason: interview and CI usage both need a quick answer to "did this get worse?" while still preserving neutral changes. - Alternative considered: only output numeric deltas. That is harder to scan and less actionable. - Decision: Keep thresholds explicit and conservative. - Reason: pass/fail and missing evidence are hard regressions; tool calls and duration are cost signals that should be visible even if not always blocking. - Alternative considered: fail only on pass-rate drop. That misses cases where quality stays green but cost or confidence behavior changes. ## Risks / Trade-offs - Report comparison can only see fields already captured by `DiagnosisEvalReport`. Mitigation: use this as the first regression layer and add richer report fields later if needed. - Duration may fluctuate in live runs. Mitigation: fixture baseline uses stable durations; live-mode thresholds can be added later. - Verdict distribution changes can be intentional. Mitigation: classify them as `CHANGED` unless they coincide with per-case regressions.