Files
SuperBizAgent-java/openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures/design.md
T
2026-07-05 00:27:57 +08:00

2.3 KiB

Context

diagnosis-eval-harness already provides fixed case definitions, fixture-mode evaluation, JSON / Markdown report writing, and focused evaluator tests. The current baseline is incomplete because three fixed cases intentionally point to missing fixtures.

Goals / Non-Goals

Goals:

  • Add representative trace fixtures for every fixed diagnosis case.
  • Save a baseline report that can be reviewed and compared after future Agent changes.
  • Keep the baseline reproducible in offline mode.
  • Document how to regenerate the baseline.

Non-Goals:

  • Do not change production Agent runtime behavior.
  • Do not require live infrastructure or a real LLM.
  • Do not introduce a new LLM-based grader.
  • Do not expand the case set beyond the existing five fixed MVP diagnosis cases.

Decisions

  • Use checked-in fixture traces instead of live service calls.

    • Rationale: the goal is a stable regression baseline that can run in CI or interview environments without external dependencies.
    • Alternative considered: start the application and call the trace API. That is useful later, but it introduces infrastructure noise before the baseline is complete.
  • Save baseline reports under mvp/eval/reports.

    • Rationale: reports are reviewable artifacts, not transient build output, and they show the expected current behavior of the baseline.
    • Alternative considered: generate reports only in tests. That verifies behavior but does not give an easy artifact to show or diff.
  • Keep fixture outcomes representative rather than forcing every case to pass.

    • Rationale: a baseline should reflect expected behavior, including low-confidence or degraded cases, as long as the outcome is explicit and stable.
    • Alternative considered: make every fixture pass. That looks cleaner but hides important degraded-path behavior.

Risks / Trade-offs

  • Fixture data can drift from real runtime traces. Mitigation: keep fixtures shaped like DiagnosisTraceResponse and add tests that load every referenced fixture.
  • A saved baseline report can become stale after intentional rule changes. Mitigation: document regeneration steps and update the report in the same change as rule or fixture updates.
  • Keyword-based checks are coarse. Mitigation: this change keeps the deterministic harness simple and leaves semantic scoring as a later improvement.