40 lines
2.3 KiB
Markdown
40 lines
2.3 KiB
Markdown
## Context
|
|
|
|
`diagnosis-eval-harness` already provides fixed case definitions, fixture-mode evaluation, JSON / Markdown report writing, and focused evaluator tests. The current baseline is incomplete because three fixed cases intentionally point to missing fixtures.
|
|
|
|
## Goals / Non-Goals
|
|
|
|
**Goals:**
|
|
|
|
- Add representative trace fixtures for every fixed diagnosis case.
|
|
- Save a baseline report that can be reviewed and compared after future Agent changes.
|
|
- Keep the baseline reproducible in offline mode.
|
|
- Document how to regenerate the baseline.
|
|
|
|
**Non-Goals:**
|
|
|
|
- Do not change production Agent runtime behavior.
|
|
- Do not require live infrastructure or a real LLM.
|
|
- Do not introduce a new LLM-based grader.
|
|
- Do not expand the case set beyond the existing five fixed MVP diagnosis cases.
|
|
|
|
## Decisions
|
|
|
|
- Use checked-in fixture traces instead of live service calls.
|
|
- Rationale: the goal is a stable regression baseline that can run in CI or interview environments without external dependencies.
|
|
- Alternative considered: start the application and call the trace API. That is useful later, but it introduces infrastructure noise before the baseline is complete.
|
|
|
|
- Save baseline reports under `mvp/eval/reports`.
|
|
- Rationale: reports are reviewable artifacts, not transient build output, and they show the expected current behavior of the baseline.
|
|
- Alternative considered: generate reports only in tests. That verifies behavior but does not give an easy artifact to show or diff.
|
|
|
|
- Keep fixture outcomes representative rather than forcing every case to pass.
|
|
- Rationale: a baseline should reflect expected behavior, including low-confidence or degraded cases, as long as the outcome is explicit and stable.
|
|
- Alternative considered: make every fixture pass. That looks cleaner but hides important degraded-path behavior.
|
|
|
|
## Risks / Trade-offs
|
|
|
|
- Fixture data can drift from real runtime traces. Mitigation: keep fixtures shaped like `DiagnosisTraceResponse` and add tests that load every referenced fixture.
|
|
- A saved baseline report can become stale after intentional rule changes. Mitigation: document regeneration steps and update the report in the same change as rule or fixture updates.
|
|
- Keyword-based checks are coarse. Mitigation: this change keeps the deterministic harness simple and leaves semantic scoring as a later improvement.
|