# diagnosis-eval-harness Specification ## Purpose Provide a repeatable offline evaluation harness for MVP diagnosis Agent behavior, so prompt, tool, retrieval, and verifier changes can be checked against fixed trace-based regression cases. ## Requirements ### Requirement: Evaluation harness SHALL define fixed diagnosis cases The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria. #### Scenario: Case definition includes expected evidence - **WHEN** an evaluation case is defined - **THEN** it SHALL include a case id, user question, expected root-cause keywords, required evidence tools, and allowed verifier verdicts #### Scenario: Case definition can express forbidden behavior - **WHEN** a case has known unsafe behavior - **THEN** the case definition SHALL be able to declare forbidden answer keywords or forbidden verdicts ### Requirement: Evaluation harness SHALL validate diagnosis traces The system SHALL validate a diagnosis trace against the corresponding case definition using deterministic rules. #### Scenario: Evidence coverage validation - **WHEN** a trace is evaluated - **THEN** the evaluator SHALL verify that required evidence tools appear in `toolInvocations` or verifier trace summaries #### Scenario: Verifier evaluation validation - **WHEN** a trace is evaluated - **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.verdict` exists - **AND** the verdict SHALL be one of the case's allowed verdicts #### Scenario: Answer keyword validation - **WHEN** a trace is evaluated - **THEN** the evaluator SHALL verify that the final answer includes at least the configured minimum root-cause keyword coverage #### Scenario: Degraded output validation - **WHEN** a trace verdict is `REJECT` - **THEN** the evaluator SHALL verify that the final answer follows the degraded-output contract rather than leaking an unverified executor answer ### Requirement: Evaluation harness SHALL report quality and cost signals The system SHALL produce a report that summarizes pass/fail results and key trace metrics. #### Scenario: JSON report output - **WHEN** an evaluation run completes - **THEN** the evaluator SHALL output a JSON report with per-case result, failed checks, verdict, evidence coverage, tool-call count, and duration #### Scenario: Markdown report output - **WHEN** an evaluation run completes - **THEN** the evaluator SHALL output a Markdown report suitable for review in the repository #### Scenario: Aggregate metrics - **WHEN** multiple cases are evaluated - **THEN** the report SHALL include aggregate pass rate, verdict distribution, average tool-call count, and average duration when available ### Requirement: Evaluation harness SHALL support offline fixture mode The first evaluator version SHALL be runnable without live MySQL, Redis, Milvus, or LLM services by evaluating saved trace fixtures. #### Scenario: Fixture trace evaluation - **WHEN** the evaluator is run against a directory of trace fixture files - **THEN** it SHALL evaluate each trace file against its matching case definition - **AND** it SHALL not require a running application service #### Scenario: Missing fixture is reported clearly - **WHEN** a case has no matching trace fixture - **THEN** the evaluator SHALL mark the case as not run or failed with a clear reason ### Requirement: Evaluation harness SHALL provide complete fixture coverage for fixed cases The system SHALL include an offline trace fixture for every fixed MVP diagnosis evaluation case. #### Scenario: Every case resolves to a fixture file - **WHEN** the evaluator loads the fixed case definition file - **THEN** every case's `traceFixture` value SHALL resolve to an existing JSON fixture file #### Scenario: Fixture files are loadable as diagnosis traces - **WHEN** each referenced fixture is loaded - **THEN** it SHALL deserialize into the trace response shape used by the evaluator ### Requirement: Evaluation harness SHALL preserve a reproducible baseline report The system SHALL preserve a generated baseline report for the full fixed fixture set. #### Scenario: Baseline report includes all fixed cases - **WHEN** the baseline report is generated from the fixed case file and fixture directory - **THEN** the report SHALL include one result for every fixed case #### Scenario: Baseline report is reviewable - **WHEN** the baseline report is written - **THEN** it SHALL be available in JSON and Markdown formats under the eval documentation area #### Scenario: Baseline regeneration is documented - **WHEN** a developer changes fixtures or evaluator rules - **THEN** the eval documentation SHALL explain how to regenerate the baseline report ### Requirement: Evaluation harness SHALL compare reports against a baseline The system SHALL compare a current diagnosis evaluation report against a saved baseline report using deterministic rules. #### Scenario: Aggregate regression detection - **WHEN** the current report has a lower pass rate than the baseline report - **THEN** the diff SHALL record a regression with the old value, new value, and delta #### Scenario: Cost signal detection - **WHEN** average tool-call count or average duration changes between reports - **THEN** the diff SHALL record the baseline value, current value, and delta #### Scenario: Verdict distribution comparison - **WHEN** verdict counts differ between reports - **THEN** the diff SHALL record the verdict distribution changes ### Requirement: Evaluation harness SHALL compare case-level report results The system SHALL compare case results by case id and report actionable per-case changes. #### Scenario: Case pass/fail regression - **WHEN** a case changes from passing in the baseline to failing in the current report - **THEN** the diff SHALL record a regression for that case #### Scenario: Evidence coverage regression - **WHEN** a required evidence tool changes from covered to uncovered for a case - **THEN** the diff SHALL record a regression naming the case and tool #### Scenario: Missing case detection - **WHEN** a baseline case is absent from the current report - **THEN** the diff SHALL record a regression for the missing case #### Scenario: New case detection - **WHEN** a current report contains a case absent from the baseline - **THEN** the diff SHALL record the case as a non-regression change ### Requirement: Evaluation harness SHALL report baseline diff results The system SHALL expose baseline diff output in structured JSON and reviewable Markdown formats. #### Scenario: JSON diff output - **WHEN** a baseline diff is written as JSON - **THEN** it SHALL include aggregate summary fields and detailed diff items #### Scenario: Markdown diff output - **WHEN** a baseline diff is written as Markdown - **THEN** it SHALL include a readable summary and a table of diff items ### Requirement: Evaluation harness SHALL validate Executor V2 audit closure The evaluation harness SHALL be able to validate the V2 audit chain from Executor structured output through Gatekeeper, Verifier claim checks, Composer output, and final user answer. #### Scenario: Required V2 audit fields are present - **WHEN** an evaluation case requires V2 audit closure - **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.gatekeeper_result` exists - **AND** it SHALL verify that `selfEvaluation.verifier_evaluation.claim_checks` exists - **AND** it SHALL verify that `selfEvaluation.verifier_evaluation.composer_output` exists #### Scenario: Gatekeeper failure cannot pass verification - **WHEN** a trace has `gatekeeper_result.status` equal to `fail` - **THEN** the evaluator SHALL fail the case if `selfEvaluation.verifier_evaluation.verdict` is `PASS` #### Scenario: Claim checks are auditable - **WHEN** an evaluation case requires claim checks - **THEN** the evaluator SHALL verify that each claim check includes `claim_id`, `verification`, and `detail` - **AND** each verification SHALL be one of the V2 claim verification values recognized by the Verifier contract #### Scenario: Composer output is auditable - **WHEN** an evaluation case requires Composer output - **THEN** the evaluator SHALL verify that `composer_output` records whether parsed output or fallback rendering was used ### Requirement: Evaluation harness SHALL prevent unsupported claims from leaking into final answers The evaluation harness SHALL detect configured unsafe or unsupported claim text when it appears in the final user-facing answer. #### Scenario: Unsupported final-answer claim is rejected - **WHEN** an evaluation case declares forbidden confirmed-claim keywords - **THEN** the evaluator SHALL fail the case if the final answer contains any of those keywords #### Scenario: Raw Executor JSON is not user-facing - **WHEN** a trace is evaluated under V2 audit closure - **THEN** the evaluator SHALL fail the case if the final answer contains raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id` #### Scenario: Composer fallback still avoids raw JSON leakage - **WHEN** a trace records Composer fallback rendering - **THEN** the evaluator SHALL still enforce final-answer raw JSON leakage checks ### Requirement: Evaluation harness SHALL expose an evidence-pipeline matrix The evaluation harness SHALL include fixed cases that demonstrate the current Chat evidence pipeline across positive evidence, narrow-scope observation, no-evidence negative observation, low-confidence filtering, reject safety, and composer fallback behavior. #### Scenario: Matrix cases are fixture backed - **WHEN** the fixed diagnosis case file is evaluated - **THEN** each matrix case SHALL resolve to an offline trace fixture - **AND** evaluation SHALL not require a live LLM or running application #### Scenario: Matrix cases preserve V2 audit closure - **WHEN** a matrix case requires V2 audit closure - **THEN** its fixture SHALL include `gatekeeper_result` - **AND** it SHALL include `claim_checks` - **AND** it SHALL include `composer_output` ### Requirement: Evaluation harness SHALL validate Gatekeeper rule set version when requested The evaluation harness SHALL be able to assert the Gatekeeper rule set version recorded in a fixture. #### Scenario: Expected rule set version matches - **WHEN** an evaluation case declares `expectedGatekeeperRuleSetVersion` - **AND** the fixture has the same `selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` - **THEN** the rule set version check SHALL pass #### Scenario: Expected rule set version mismatches - **WHEN** an evaluation case declares `expectedGatekeeperRuleSetVersion` - **AND** the fixture has a different or missing rule set version - **THEN** the case SHALL fail with a clear failed check