9.1 KiB
diagnosis-eval-harness Specification
Purpose
Provide a repeatable offline evaluation harness for MVP diagnosis Agent behavior, so prompt, tool, retrieval, and verifier changes can be checked against fixed trace-based regression cases.
Requirements
Requirement: Evaluation harness SHALL define fixed diagnosis cases
The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria.
Scenario: Case definition includes expected evidence
- WHEN an evaluation case is defined
- THEN it SHALL include a case id, user question, expected root-cause keywords, required evidence tools, and allowed verifier verdicts
Scenario: Case definition can express forbidden behavior
- WHEN a case has known unsafe behavior
- THEN the case definition SHALL be able to declare forbidden answer keywords or forbidden verdicts
Requirement: Evaluation harness SHALL validate diagnosis traces
The system SHALL validate a diagnosis trace against the corresponding case definition using deterministic rules.
Scenario: Evidence coverage validation
- WHEN a trace is evaluated
- THEN the evaluator SHALL verify that required evidence tools appear in
toolInvocationsor verifier trace summaries
Scenario: Verifier evaluation validation
- WHEN a trace is evaluated
- THEN the evaluator SHALL verify that
selfEvaluation.verifier_evaluation.verdictexists - AND the verdict SHALL be one of the case's allowed verdicts
Scenario: Answer keyword validation
- WHEN a trace is evaluated
- THEN the evaluator SHALL verify that the final answer includes at least the configured minimum root-cause keyword coverage
Scenario: Degraded output validation
- WHEN a trace verdict is
REJECT - THEN the evaluator SHALL verify that the final answer follows the degraded-output contract rather than leaking an unverified executor answer
Requirement: Evaluation harness SHALL report quality and cost signals
The system SHALL produce a report that summarizes pass/fail results and key trace metrics.
Scenario: JSON report output
- WHEN an evaluation run completes
- THEN the evaluator SHALL output a JSON report with per-case result, failed checks, verdict, evidence coverage, tool-call count, and duration
Scenario: Markdown report output
- WHEN an evaluation run completes
- THEN the evaluator SHALL output a Markdown report suitable for review in the repository
Scenario: Aggregate metrics
- WHEN multiple cases are evaluated
- THEN the report SHALL include aggregate pass rate, verdict distribution, average tool-call count, and average duration when available
Requirement: Evaluation harness SHALL support offline fixture mode
The first evaluator version SHALL be runnable without live MySQL, Redis, Milvus, or LLM services by evaluating saved trace fixtures.
Scenario: Fixture trace evaluation
- WHEN the evaluator is run against a directory of trace fixture files
- THEN it SHALL evaluate each trace file against its matching case definition
- AND it SHALL not require a running application service
Scenario: Missing fixture is reported clearly
- WHEN a case has no matching trace fixture
- THEN the evaluator SHALL mark the case as not run or failed with a clear reason
Requirement: Evaluation harness SHALL provide complete fixture coverage for fixed cases
The system SHALL include an offline trace fixture for every fixed MVP diagnosis evaluation case.
Scenario: Every case resolves to a fixture file
- WHEN the evaluator loads the fixed case definition file
- THEN every case's
traceFixturevalue SHALL resolve to an existing JSON fixture file
Scenario: Fixture files are loadable as diagnosis traces
- WHEN each referenced fixture is loaded
- THEN it SHALL deserialize into the trace response shape used by the evaluator
Requirement: Evaluation harness SHALL preserve a reproducible baseline report
The system SHALL preserve a generated baseline report for the full fixed fixture set.
Scenario: Baseline report includes all fixed cases
- WHEN the baseline report is generated from the fixed case file and fixture directory
- THEN the report SHALL include one result for every fixed case
Scenario: Baseline report is reviewable
- WHEN the baseline report is written
- THEN it SHALL be available in JSON and Markdown formats under the eval documentation area
Scenario: Baseline regeneration is documented
- WHEN a developer changes fixtures or evaluator rules
- THEN the eval documentation SHALL explain how to regenerate the baseline report
Requirement: Evaluation harness SHALL compare reports against a baseline
The system SHALL compare a current diagnosis evaluation report against a saved baseline report using deterministic rules.
Scenario: Aggregate regression detection
- WHEN the current report has a lower pass rate than the baseline report
- THEN the diff SHALL record a regression with the old value, new value, and delta
Scenario: Cost signal detection
- WHEN average tool-call count or average duration changes between reports
- THEN the diff SHALL record the baseline value, current value, and delta
Scenario: Verdict distribution comparison
- WHEN verdict counts differ between reports
- THEN the diff SHALL record the verdict distribution changes
Requirement: Evaluation harness SHALL compare case-level report results
The system SHALL compare case results by case id and report actionable per-case changes.
Scenario: Case pass/fail regression
- WHEN a case changes from passing in the baseline to failing in the current report
- THEN the diff SHALL record a regression for that case
Scenario: Evidence coverage regression
- WHEN a required evidence tool changes from covered to uncovered for a case
- THEN the diff SHALL record a regression naming the case and tool
Scenario: Missing case detection
- WHEN a baseline case is absent from the current report
- THEN the diff SHALL record a regression for the missing case
Scenario: New case detection
- WHEN a current report contains a case absent from the baseline
- THEN the diff SHALL record the case as a non-regression change
Requirement: Evaluation harness SHALL report baseline diff results
The system SHALL expose baseline diff output in structured JSON and reviewable Markdown formats.
Scenario: JSON diff output
- WHEN a baseline diff is written as JSON
- THEN it SHALL include aggregate summary fields and detailed diff items
Scenario: Markdown diff output
- WHEN a baseline diff is written as Markdown
- THEN it SHALL include a readable summary and a table of diff items
Requirement: Evaluation harness SHALL validate Executor V2 audit closure
The evaluation harness SHALL be able to validate the V2 audit chain from Executor structured output through Gatekeeper, Verifier claim checks, Composer output, and final user answer.
Scenario: Required V2 audit fields are present
- WHEN an evaluation case requires V2 audit closure
- THEN the evaluator SHALL verify that
selfEvaluation.verifier_evaluation.gatekeeper_resultexists - AND it SHALL verify that
selfEvaluation.verifier_evaluation.claim_checksexists - AND it SHALL verify that
selfEvaluation.verifier_evaluation.composer_outputexists
Scenario: Gatekeeper failure cannot pass verification
- WHEN a trace has
gatekeeper_result.statusequal tofail - THEN the evaluator SHALL fail the case if
selfEvaluation.verifier_evaluation.verdictisPASS
Scenario: Claim checks are auditable
- WHEN an evaluation case requires claim checks
- THEN the evaluator SHALL verify that each claim check includes
claim_id,verification, anddetail - AND each verification SHALL be one of the V2 claim verification values recognized by the Verifier contract
Scenario: Composer output is auditable
- WHEN an evaluation case requires Composer output
- THEN the evaluator SHALL verify that
composer_outputrecords whether parsed output or fallback rendering was used
Requirement: Evaluation harness SHALL prevent unsupported claims from leaking into final answers
The evaluation harness SHALL detect configured unsafe or unsupported claim text when it appears in the final user-facing answer.
Scenario: Unsupported final-answer claim is rejected
- WHEN an evaluation case declares forbidden confirmed-claim keywords
- THEN the evaluator SHALL fail the case if the final answer contains any of those keywords
Scenario: Raw Executor JSON is not user-facing
- WHEN a trace is evaluated under V2 audit closure
- THEN the evaluator SHALL fail the case if the final answer contains raw Executor protocol markers such as
executor_evidence_v2,answer_version,evidence_bindings, orclaim_id
Scenario: Composer fallback still avoids raw JSON leakage
- WHEN a trace records Composer fallback rendering
- THEN the evaluator SHALL still enforce final-answer raw JSON leakage checks