## Context `diagnosis-eval-harness` already provides fixed case definitions, fixture-mode evaluation, JSON / Markdown report writing, and focused evaluator tests. The current baseline is incomplete because three fixed cases intentionally point to missing fixtures. ## Goals / Non-Goals **Goals:** - Add representative trace fixtures for every fixed diagnosis case. - Save a baseline report that can be reviewed and compared after future Agent changes. - Keep the baseline reproducible in offline mode. - Document how to regenerate the baseline. **Non-Goals:** - Do not change production Agent runtime behavior. - Do not require live infrastructure or a real LLM. - Do not introduce a new LLM-based grader. - Do not expand the case set beyond the existing five fixed MVP diagnosis cases. ## Decisions - Use checked-in fixture traces instead of live service calls. - Rationale: the goal is a stable regression baseline that can run in CI or interview environments without external dependencies. - Alternative considered: start the application and call the trace API. That is useful later, but it introduces infrastructure noise before the baseline is complete. - Save baseline reports under `mvp/eval/reports`. - Rationale: reports are reviewable artifacts, not transient build output, and they show the expected current behavior of the baseline. - Alternative considered: generate reports only in tests. That verifies behavior but does not give an easy artifact to show or diff. - Keep fixture outcomes representative rather than forcing every case to pass. - Rationale: a baseline should reflect expected behavior, including low-confidence or degraded cases, as long as the outcome is explicit and stable. - Alternative considered: make every fixture pass. That looks cleaner but hides important degraded-path behavior. ## Risks / Trade-offs - Fixture data can drift from real runtime traces. Mitigation: keep fixtures shaped like `DiagnosisTraceResponse` and add tests that load every referenced fixture. - A saved baseline report can become stale after intentional rule changes. Mitigation: document regeneration steps and update the report in the same change as rule or fixture updates. - Keyword-based checks are coarse. Mitigation: this change keeps the deterministic harness simple and leaves semantic scoring as a later improvement.