## Context The project now has the pieces needed for trace-based evaluation: - `diagnosis_session` stores final answer, status, duration, counts, feedback, and `self_evaluation` - `agent_step` stores ordered agent execution records - `tool_invocation` stores evidence tool calls with normalized evidence semantics - `GET /api/diagnosis/{sessionId}/trace` can aggregate one diagnosis trace for demo review - `evidence-trace-hardening` defined stable `supported`, `no_evidence`, `deduped`, and `failed` semantics P1-B should not add another runtime agent. It should create a repeatable evaluation surface that can be used after changing prompts, retrieval behavior, tools, or verifier logic. ## Goals / Non-Goals **Goals:** - Define fixed MVP diagnosis cases with expected evidence and verdict rules. - Build a deterministic evaluator that can validate a diagnosis trace against a case definition. - Produce JSON and Markdown reports with pass/fail status and key metrics. - Keep the first version usable without a real LLM by allowing fixture trace inputs. - Leave room for a later runtime mode that queries the trace API after a demo run. **Non-Goals:** - No LLM-as-judge in this slice. - No automatic prompt optimization. - No new production API. - No change to chat, verifier, retrieval, upload, or feedback behavior. - No requirement to start MySQL/Redis/Milvus/LLM for the first offline evaluator. ## Decisions | Decision | Choice | Alternative Considered | Rationale | |---|---|---|---| | Evaluation source | Start with fixture / persisted trace JSON input | Always run live `/api/chat` first | Keeps the first harness deterministic and avoids mixing quality checks with external infrastructure availability. | | Judging strategy | Rule-based trace validation | LLM-as-judge | The immediate goal is regression signal for evidence coverage and degraded behavior, not subjective answer scoring. | | Case format | Static JSON/YAML case definitions | Hard-coded Java tests only | Case files are easier to inspect and explain in interviews. | | Report format | JSON plus Markdown | Console-only output | JSON supports automation; Markdown supports quick human review. | | Metrics | Evidence coverage, verdict distribution, tool-call count, duration, answer keyword coverage | Full semantic correctness | These metrics are available from existing trace data and align with the MVP's observable contract. | ## Risks / Trade-offs - [Risk] Rule-based keyword checks can be brittle. -> Mitigation: keep checks focused on required evidence, verdicts, and high-signal root-cause terms rather than exact answer text. - [Risk] Fixture-only evaluation may drift from runtime behavior. -> Mitigation: design the evaluator around the same trace response shape so runtime traces can be fed in later. - [Risk] Metrics may encourage gaming tool counts. -> Mitigation: report tool counts as cost/efficiency signals, not the sole pass/fail criterion. - [Risk] Too many cases can slow iteration. -> Mitigation: start with 5 MVP cases and keep each case small. ## Migration Plan - No deployment migration is required. - The harness is additive and can be run locally as a test or script. - Rollback is deleting the eval case files, runner, and report docs. ## Open Questions - Should runtime trace API polling be included in the first implementation, or left as a follow-up after the fixture validator lands?