6.7 KiB
diagnosis-eval-harness Specification
Purpose
Provide a repeatable offline evaluation harness for MVP diagnosis Agent behavior, so prompt, tool, retrieval, and verifier changes can be checked against fixed trace-based regression cases.
Requirements
Requirement: Evaluation harness SHALL define fixed diagnosis cases
The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria.
Scenario: Case definition includes expected evidence
- WHEN an evaluation case is defined
- THEN it SHALL include a case id, user question, expected root-cause keywords, required evidence tools, and allowed verifier verdicts
Scenario: Case definition can express forbidden behavior
- WHEN a case has known unsafe behavior
- THEN the case definition SHALL be able to declare forbidden answer keywords or forbidden verdicts
Requirement: Evaluation harness SHALL validate diagnosis traces
The system SHALL validate a diagnosis trace against the corresponding case definition using deterministic rules.
Scenario: Evidence coverage validation
- WHEN a trace is evaluated
- THEN the evaluator SHALL verify that required evidence tools appear in
toolInvocationsor verifier trace summaries
Scenario: Verifier evaluation validation
- WHEN a trace is evaluated
- THEN the evaluator SHALL verify that
selfEvaluation.verifier_evaluation.verdictexists - AND the verdict SHALL be one of the case's allowed verdicts
Scenario: Answer keyword validation
- WHEN a trace is evaluated
- THEN the evaluator SHALL verify that the final answer includes at least the configured minimum root-cause keyword coverage
Scenario: Degraded output validation
- WHEN a trace verdict is
REJECT - THEN the evaluator SHALL verify that the final answer follows the degraded-output contract rather than leaking an unverified executor answer
Requirement: Evaluation harness SHALL report quality and cost signals
The system SHALL produce a report that summarizes pass/fail results and key trace metrics.
Scenario: JSON report output
- WHEN an evaluation run completes
- THEN the evaluator SHALL output a JSON report with per-case result, failed checks, verdict, evidence coverage, tool-call count, and duration
Scenario: Markdown report output
- WHEN an evaluation run completes
- THEN the evaluator SHALL output a Markdown report suitable for review in the repository
Scenario: Aggregate metrics
- WHEN multiple cases are evaluated
- THEN the report SHALL include aggregate pass rate, verdict distribution, average tool-call count, and average duration when available
Requirement: Evaluation harness SHALL support offline fixture mode
The first evaluator version SHALL be runnable without live MySQL, Redis, Milvus, or LLM services by evaluating saved trace fixtures.
Scenario: Fixture trace evaluation
- WHEN the evaluator is run against a directory of trace fixture files
- THEN it SHALL evaluate each trace file against its matching case definition
- AND it SHALL not require a running application service
Scenario: Missing fixture is reported clearly
- WHEN a case has no matching trace fixture
- THEN the evaluator SHALL mark the case as not run or failed with a clear reason
Requirement: Evaluation harness SHALL provide complete fixture coverage for fixed cases
The system SHALL include an offline trace fixture for every fixed MVP diagnosis evaluation case.
Scenario: Every case resolves to a fixture file
- WHEN the evaluator loads the fixed case definition file
- THEN every case's
traceFixturevalue SHALL resolve to an existing JSON fixture file
Scenario: Fixture files are loadable as diagnosis traces
- WHEN each referenced fixture is loaded
- THEN it SHALL deserialize into the trace response shape used by the evaluator
Requirement: Evaluation harness SHALL preserve a reproducible baseline report
The system SHALL preserve a generated baseline report for the full fixed fixture set.
Scenario: Baseline report includes all fixed cases
- WHEN the baseline report is generated from the fixed case file and fixture directory
- THEN the report SHALL include one result for every fixed case
Scenario: Baseline report is reviewable
- WHEN the baseline report is written
- THEN it SHALL be available in JSON and Markdown formats under the eval documentation area
Scenario: Baseline regeneration is documented
- WHEN a developer changes fixtures or evaluator rules
- THEN the eval documentation SHALL explain how to regenerate the baseline report
Requirement: Evaluation harness SHALL compare reports against a baseline
The system SHALL compare a current diagnosis evaluation report against a saved baseline report using deterministic rules.
Scenario: Aggregate regression detection
- WHEN the current report has a lower pass rate than the baseline report
- THEN the diff SHALL record a regression with the old value, new value, and delta
Scenario: Cost signal detection
- WHEN average tool-call count or average duration changes between reports
- THEN the diff SHALL record the baseline value, current value, and delta
Scenario: Verdict distribution comparison
- WHEN verdict counts differ between reports
- THEN the diff SHALL record the verdict distribution changes
Requirement: Evaluation harness SHALL compare case-level report results
The system SHALL compare case results by case id and report actionable per-case changes.
Scenario: Case pass/fail regression
- WHEN a case changes from passing in the baseline to failing in the current report
- THEN the diff SHALL record a regression for that case
Scenario: Evidence coverage regression
- WHEN a required evidence tool changes from covered to uncovered for a case
- THEN the diff SHALL record a regression naming the case and tool
Scenario: Missing case detection
- WHEN a baseline case is absent from the current report
- THEN the diff SHALL record a regression for the missing case
Scenario: New case detection
- WHEN a current report contains a case absent from the baseline
- THEN the diff SHALL record the case as a non-regression change
Requirement: Evaluation harness SHALL report baseline diff results
The system SHALL expose baseline diff output in structured JSON and reviewable Markdown formats.
Scenario: JSON diff output
- WHEN a baseline diff is written as JSON
- THEN it SHALL include aggregate summary fields and detailed diff items
Scenario: Markdown diff output
- WHEN a baseline diff is written as Markdown
- THEN it SHALL include a readable summary and a table of diff items