2.3 KiB
2.3 KiB
Context
diagnosis-eval-harness already provides fixed case definitions, fixture-mode evaluation, JSON / Markdown report writing, and focused evaluator tests. The current baseline is incomplete because three fixed cases intentionally point to missing fixtures.
Goals / Non-Goals
Goals:
- Add representative trace fixtures for every fixed diagnosis case.
- Save a baseline report that can be reviewed and compared after future Agent changes.
- Keep the baseline reproducible in offline mode.
- Document how to regenerate the baseline.
Non-Goals:
- Do not change production Agent runtime behavior.
- Do not require live infrastructure or a real LLM.
- Do not introduce a new LLM-based grader.
- Do not expand the case set beyond the existing five fixed MVP diagnosis cases.
Decisions
-
Use checked-in fixture traces instead of live service calls.
- Rationale: the goal is a stable regression baseline that can run in CI or interview environments without external dependencies.
- Alternative considered: start the application and call the trace API. That is useful later, but it introduces infrastructure noise before the baseline is complete.
-
Save baseline reports under
mvp/eval/reports.- Rationale: reports are reviewable artifacts, not transient build output, and they show the expected current behavior of the baseline.
- Alternative considered: generate reports only in tests. That verifies behavior but does not give an easy artifact to show or diff.
-
Keep fixture outcomes representative rather than forcing every case to pass.
- Rationale: a baseline should reflect expected behavior, including low-confidence or degraded cases, as long as the outcome is explicit and stable.
- Alternative considered: make every fixture pass. That looks cleaner but hides important degraded-path behavior.
Risks / Trade-offs
- Fixture data can drift from real runtime traces. Mitigation: keep fixtures shaped like
DiagnosisTraceResponseand add tests that load every referenced fixture. - A saved baseline report can become stale after intentional rule changes. Mitigation: document regeneration steps and update the report in the same change as rule or fixture updates.
- Keyword-based checks are coarse. Mitigation: this change keeps the deterministic harness simple and leaves semantic scoring as a later improvement.