# diagnosis-eval-harness Specification ## Purpose Provide a repeatable offline evaluation harness for MVP diagnosis Agent behavior, so prompt, tool, retrieval, and verifier changes can be checked against fixed trace-based regression cases. ## Requirements ### Requirement: Evaluation harness SHALL define fixed diagnosis cases The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria. #### Scenario: Case definition includes expected evidence - **WHEN** an evaluation case is defined - **THEN** it SHALL include a case id, user question, expected root-cause keywords, required evidence tools, and allowed verifier verdicts #### Scenario: Case definition can express forbidden behavior - **WHEN** a case has known unsafe behavior - **THEN** the case definition SHALL be able to declare forbidden answer keywords or forbidden verdicts ### Requirement: Evaluation harness SHALL validate diagnosis traces The system SHALL validate a diagnosis trace against the corresponding case definition using deterministic rules. #### Scenario: Evidence coverage validation - **WHEN** a trace is evaluated - **THEN** the evaluator SHALL verify that required evidence tools appear in `toolInvocations` or verifier trace summaries #### Scenario: Verifier evaluation validation - **WHEN** a trace is evaluated - **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.verdict` exists - **AND** the verdict SHALL be one of the case's allowed verdicts #### Scenario: Answer keyword validation - **WHEN** a trace is evaluated - **THEN** the evaluator SHALL verify that the final answer includes at least the configured minimum root-cause keyword coverage #### Scenario: Degraded output validation - **WHEN** a trace verdict is `REJECT` - **THEN** the evaluator SHALL verify that the final answer follows the degraded-output contract rather than leaking an unverified executor answer ### Requirement: Evaluation harness SHALL report quality and cost signals The system SHALL produce a report that summarizes pass/fail results and key trace metrics. #### Scenario: JSON report output - **WHEN** an evaluation run completes - **THEN** the evaluator SHALL output a JSON report with per-case result, failed checks, verdict, evidence coverage, tool-call count, and duration #### Scenario: Markdown report output - **WHEN** an evaluation run completes - **THEN** the evaluator SHALL output a Markdown report suitable for review in the repository #### Scenario: Aggregate metrics - **WHEN** multiple cases are evaluated - **THEN** the report SHALL include aggregate pass rate, verdict distribution, average tool-call count, and average duration when available ### Requirement: Evaluation harness SHALL support offline fixture mode The first evaluator version SHALL be runnable without live MySQL, Redis, Milvus, or LLM services by evaluating saved trace fixtures. #### Scenario: Fixture trace evaluation - **WHEN** the evaluator is run against a directory of trace fixture files - **THEN** it SHALL evaluate each trace file against its matching case definition - **AND** it SHALL not require a running application service #### Scenario: Missing fixture is reported clearly - **WHEN** a case has no matching trace fixture - **THEN** the evaluator SHALL mark the case as not run or failed with a clear reason ### Requirement: Evaluation harness SHALL provide complete fixture coverage for fixed cases The system SHALL include an offline trace fixture for every fixed MVP diagnosis evaluation case. #### Scenario: Every case resolves to a fixture file - **WHEN** the evaluator loads the fixed case definition file - **THEN** every case's `traceFixture` value SHALL resolve to an existing JSON fixture file #### Scenario: Fixture files are loadable as diagnosis traces - **WHEN** each referenced fixture is loaded - **THEN** it SHALL deserialize into the trace response shape used by the evaluator ### Requirement: Evaluation harness SHALL preserve a reproducible baseline report The system SHALL preserve a generated baseline report for the full fixed fixture set. #### Scenario: Baseline report includes all fixed cases - **WHEN** the baseline report is generated from the fixed case file and fixture directory - **THEN** the report SHALL include one result for every fixed case #### Scenario: Baseline report is reviewable - **WHEN** the baseline report is written - **THEN** it SHALL be available in JSON and Markdown formats under the eval documentation area #### Scenario: Baseline regeneration is documented - **WHEN** a developer changes fixtures or evaluator rules - **THEN** the eval documentation SHALL explain how to regenerate the baseline report ### Requirement: Evaluation harness SHALL compare reports against a baseline The system SHALL compare a current diagnosis evaluation report against a saved baseline report using deterministic rules. #### Scenario: Aggregate regression detection - **WHEN** the current report has a lower pass rate than the baseline report - **THEN** the diff SHALL record a regression with the old value, new value, and delta #### Scenario: Cost signal detection - **WHEN** average tool-call count or average duration changes between reports - **THEN** the diff SHALL record the baseline value, current value, and delta #### Scenario: Verdict distribution comparison - **WHEN** verdict counts differ between reports - **THEN** the diff SHALL record the verdict distribution changes ### Requirement: Evaluation harness SHALL compare case-level report results The system SHALL compare case results by case id and report actionable per-case changes. #### Scenario: Case pass/fail regression - **WHEN** a case changes from passing in the baseline to failing in the current report - **THEN** the diff SHALL record a regression for that case #### Scenario: Evidence coverage regression - **WHEN** a required evidence tool changes from covered to uncovered for a case - **THEN** the diff SHALL record a regression naming the case and tool #### Scenario: Missing case detection - **WHEN** a baseline case is absent from the current report - **THEN** the diff SHALL record a regression for the missing case #### Scenario: New case detection - **WHEN** a current report contains a case absent from the baseline - **THEN** the diff SHALL record the case as a non-regression change ### Requirement: Evaluation harness SHALL report baseline diff results The system SHALL expose baseline diff output in structured JSON and reviewable Markdown formats. #### Scenario: JSON diff output - **WHEN** a baseline diff is written as JSON - **THEN** it SHALL include aggregate summary fields and detailed diff items #### Scenario: Markdown diff output - **WHEN** a baseline diff is written as Markdown - **THEN** it SHALL include a readable summary and a table of diff items