Files
2026-07-05 00:59:53 +08:00

2.3 KiB

Context

The eval harness now has a complete five-case fixture baseline and saved JSON / Markdown baseline reports. The missing piece is a deterministic comparison step that explains whether a new report is better, worse, or just different from the baseline.

Goals / Non-Goals

Goals:

  • Compare two DiagnosisEvalReport objects without requiring external services.
  • Surface aggregate regressions such as pass-rate drops, verdict distribution shifts, and cost increases.
  • Surface per-case regressions such as pass-to-fail changes, missing evidence coverage, verdict changes, keyword coverage loss, and missing cases.
  • Write JSON and Markdown diff outputs for review.

Non-Goals:

  • Do not run the Agent or regenerate traces.
  • Do not introduce LLM-as-judge.
  • Do not change evaluator scoring rules.
  • Do not block on performance thresholds beyond simple numeric diff signals.

Decisions

  • Decision: Compare report DTOs instead of raw traces.

    • Reason: DiagnosisEvalReport is already the stable structured output of the evaluator and is cheaper to diff than trace internals.
    • Alternative considered: compare raw trace fixtures. That would expose more detail but duplicate evaluator responsibilities.
  • Decision: Classify each diff item as REGRESSION, IMPROVEMENT, or CHANGED.

    • Reason: interview and CI usage both need a quick answer to "did this get worse?" while still preserving neutral changes.
    • Alternative considered: only output numeric deltas. That is harder to scan and less actionable.
  • Decision: Keep thresholds explicit and conservative.

    • Reason: pass/fail and missing evidence are hard regressions; tool calls and duration are cost signals that should be visible even if not always blocking.
    • Alternative considered: fail only on pass-rate drop. That misses cases where quality stays green but cost or confidence behavior changes.

Risks / Trade-offs

  • Report comparison can only see fields already captured by DiagnosisEvalReport. Mitigation: use this as the first regression layer and add richer report fields later if needed.
  • Duration may fluctuate in live runs. Mitigation: fixture baseline uses stable durations; live-mode thresholds can be added later.
  • Verdict distribution changes can be intentional. Mitigation: classify them as CHANGED unless they coincide with per-case regressions.