# Diagnosis Eval Harness This folder contains the fixed offline regression set for the MVP diagnosis Agent. ## Background The current diagnosis chain is: ```text Planner -> Executor -> Gatekeeper -> Verifier -> Composer -> final answer ``` Stages 1-4 introduced Executor V2 structured output, deterministic Gatekeeper audit, Verifier `claim_checks`, and Composer final-answer rendering. Stage 5 makes those audit fields part of the offline regression harness so future prompt, tool, or chain changes can be checked without relying on a one-off demo. ## Scope - Case definitions: `cases/diagnosis-cases.json` - Offline trace fixtures: `fixtures/*.json` - Field definitions: `schema.md` - Baseline reports: `reports/baseline-report.json` and `reports/baseline-report.md` - Baseline diff sample: `reports/baseline-diff-sample.json` and `reports/baseline-diff-sample.md` - Evaluator implementation: `DiagnosisTraceEvaluator` - Report writer: `DiagnosisEvalReportWriter` ## Current Mode The baseline evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM. The committed baseline currently contains: ```text 12 fixed cases 12 passing fixture evaluations 5 PASS verdicts 6 LOW_CONFID verdicts 1 REJECT verdict ``` The V2 evidence-pipeline matrix covers: - Positive supported evidence for a narrow HighCPU observation. - No-evidence `negative_observation` using `$.no_evidence`. - Gatekeeper failure for a fabricated tool invocation reference. - Unsupported claim filtering before the final answer. - Composer fallback rendering without raw Executor JSON leakage. - Gatekeeper rule set version audit for new matrix fixtures. - Prompt audit version checks for planner, executor, verifier, and composer prompts. - Gatekeeper rule metadata checks for enabled rule id and default severity. ## Verification Run the focused evaluator test: ```powershell mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test ``` Run the broader phase-5 regression set: ```powershell mvn "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test ``` When fixtures or evaluator rules change, regenerate both baseline reports from the same case file and fixture directory, then update JSON and Markdown together. ## Regression Signal The harness is deterministic code, not an LLM judge: ```text fixed diagnosis case -> saved trace fixture -> rule-based trace validation -> JSON / Markdown report -> regression signal for prompts, tools, retrieval, verifier, and composer behavior ``` Stage 5 adds these V2 checks: - `gatekeeper_result.status=fail` cannot coexist with Verifier `PASS`. - Required V2 fixtures must include `gatekeeper_result`, `claim_checks`, and `composer_output`. - `claim_checks` must be structurally auditable. - Composer output must record whether normal parsing or fallback rendering was used. - Gatekeeper rule set version can be asserted per fixture. - Prompt audit version and per-prompt versions can be asserted per fixture. - Gatekeeper rule metadata can be required per fixture. - Final answers must not leak raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`. - Configured unsupported claim keywords must not appear as confirmed final-answer content. ## Baseline Diff Baseline diff compares a current report against `reports/baseline-report.json`. ```text baseline report current report -> deterministic diff -> regressions, improvements, and changed signals ``` Use it to answer: did a prompt, tool, retrieval, verifier, or composer change make the Agent worse than the fixed baseline?