173 lines
9.1 KiB
Markdown
173 lines
9.1 KiB
Markdown
# diagnosis-eval-harness Specification
|
|
|
|
## Purpose
|
|
Provide a repeatable offline evaluation harness for MVP diagnosis Agent behavior, so prompt, tool, retrieval, and verifier changes can be checked against fixed trace-based regression cases.
|
|
|
|
## Requirements
|
|
|
|
### Requirement: Evaluation harness SHALL define fixed diagnosis cases
|
|
The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria.
|
|
|
|
#### Scenario: Case definition includes expected evidence
|
|
- **WHEN** an evaluation case is defined
|
|
- **THEN** it SHALL include a case id, user question, expected root-cause keywords, required evidence tools, and allowed verifier verdicts
|
|
|
|
#### Scenario: Case definition can express forbidden behavior
|
|
- **WHEN** a case has known unsafe behavior
|
|
- **THEN** the case definition SHALL be able to declare forbidden answer keywords or forbidden verdicts
|
|
|
|
### Requirement: Evaluation harness SHALL validate diagnosis traces
|
|
The system SHALL validate a diagnosis trace against the corresponding case definition using deterministic rules.
|
|
|
|
#### Scenario: Evidence coverage validation
|
|
- **WHEN** a trace is evaluated
|
|
- **THEN** the evaluator SHALL verify that required evidence tools appear in `toolInvocations` or verifier trace summaries
|
|
|
|
#### Scenario: Verifier evaluation validation
|
|
- **WHEN** a trace is evaluated
|
|
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.verdict` exists
|
|
- **AND** the verdict SHALL be one of the case's allowed verdicts
|
|
|
|
#### Scenario: Answer keyword validation
|
|
- **WHEN** a trace is evaluated
|
|
- **THEN** the evaluator SHALL verify that the final answer includes at least the configured minimum root-cause keyword coverage
|
|
|
|
#### Scenario: Degraded output validation
|
|
- **WHEN** a trace verdict is `REJECT`
|
|
- **THEN** the evaluator SHALL verify that the final answer follows the degraded-output contract rather than leaking an unverified executor answer
|
|
|
|
### Requirement: Evaluation harness SHALL report quality and cost signals
|
|
The system SHALL produce a report that summarizes pass/fail results and key trace metrics.
|
|
|
|
#### Scenario: JSON report output
|
|
- **WHEN** an evaluation run completes
|
|
- **THEN** the evaluator SHALL output a JSON report with per-case result, failed checks, verdict, evidence coverage, tool-call count, and duration
|
|
|
|
#### Scenario: Markdown report output
|
|
- **WHEN** an evaluation run completes
|
|
- **THEN** the evaluator SHALL output a Markdown report suitable for review in the repository
|
|
|
|
#### Scenario: Aggregate metrics
|
|
- **WHEN** multiple cases are evaluated
|
|
- **THEN** the report SHALL include aggregate pass rate, verdict distribution, average tool-call count, and average duration when available
|
|
|
|
### Requirement: Evaluation harness SHALL support offline fixture mode
|
|
The first evaluator version SHALL be runnable without live MySQL, Redis, Milvus, or LLM services by evaluating saved trace fixtures.
|
|
|
|
#### Scenario: Fixture trace evaluation
|
|
- **WHEN** the evaluator is run against a directory of trace fixture files
|
|
- **THEN** it SHALL evaluate each trace file against its matching case definition
|
|
- **AND** it SHALL not require a running application service
|
|
|
|
#### Scenario: Missing fixture is reported clearly
|
|
- **WHEN** a case has no matching trace fixture
|
|
- **THEN** the evaluator SHALL mark the case as not run or failed with a clear reason
|
|
|
|
### Requirement: Evaluation harness SHALL provide complete fixture coverage for fixed cases
|
|
The system SHALL include an offline trace fixture for every fixed MVP diagnosis evaluation case.
|
|
|
|
#### Scenario: Every case resolves to a fixture file
|
|
- **WHEN** the evaluator loads the fixed case definition file
|
|
- **THEN** every case's `traceFixture` value SHALL resolve to an existing JSON fixture file
|
|
|
|
#### Scenario: Fixture files are loadable as diagnosis traces
|
|
- **WHEN** each referenced fixture is loaded
|
|
- **THEN** it SHALL deserialize into the trace response shape used by the evaluator
|
|
|
|
### Requirement: Evaluation harness SHALL preserve a reproducible baseline report
|
|
The system SHALL preserve a generated baseline report for the full fixed fixture set.
|
|
|
|
#### Scenario: Baseline report includes all fixed cases
|
|
- **WHEN** the baseline report is generated from the fixed case file and fixture directory
|
|
- **THEN** the report SHALL include one result for every fixed case
|
|
|
|
#### Scenario: Baseline report is reviewable
|
|
- **WHEN** the baseline report is written
|
|
- **THEN** it SHALL be available in JSON and Markdown formats under the eval documentation area
|
|
|
|
#### Scenario: Baseline regeneration is documented
|
|
- **WHEN** a developer changes fixtures or evaluator rules
|
|
- **THEN** the eval documentation SHALL explain how to regenerate the baseline report
|
|
|
|
### Requirement: Evaluation harness SHALL compare reports against a baseline
|
|
The system SHALL compare a current diagnosis evaluation report against a saved baseline report using deterministic rules.
|
|
|
|
#### Scenario: Aggregate regression detection
|
|
- **WHEN** the current report has a lower pass rate than the baseline report
|
|
- **THEN** the diff SHALL record a regression with the old value, new value, and delta
|
|
|
|
#### Scenario: Cost signal detection
|
|
- **WHEN** average tool-call count or average duration changes between reports
|
|
- **THEN** the diff SHALL record the baseline value, current value, and delta
|
|
|
|
#### Scenario: Verdict distribution comparison
|
|
- **WHEN** verdict counts differ between reports
|
|
- **THEN** the diff SHALL record the verdict distribution changes
|
|
|
|
### Requirement: Evaluation harness SHALL compare case-level report results
|
|
The system SHALL compare case results by case id and report actionable per-case changes.
|
|
|
|
#### Scenario: Case pass/fail regression
|
|
- **WHEN** a case changes from passing in the baseline to failing in the current report
|
|
- **THEN** the diff SHALL record a regression for that case
|
|
|
|
#### Scenario: Evidence coverage regression
|
|
- **WHEN** a required evidence tool changes from covered to uncovered for a case
|
|
- **THEN** the diff SHALL record a regression naming the case and tool
|
|
|
|
#### Scenario: Missing case detection
|
|
- **WHEN** a baseline case is absent from the current report
|
|
- **THEN** the diff SHALL record a regression for the missing case
|
|
|
|
#### Scenario: New case detection
|
|
- **WHEN** a current report contains a case absent from the baseline
|
|
- **THEN** the diff SHALL record the case as a non-regression change
|
|
|
|
### Requirement: Evaluation harness SHALL report baseline diff results
|
|
The system SHALL expose baseline diff output in structured JSON and reviewable Markdown formats.
|
|
|
|
#### Scenario: JSON diff output
|
|
- **WHEN** a baseline diff is written as JSON
|
|
- **THEN** it SHALL include aggregate summary fields and detailed diff items
|
|
|
|
#### Scenario: Markdown diff output
|
|
- **WHEN** a baseline diff is written as Markdown
|
|
- **THEN** it SHALL include a readable summary and a table of diff items
|
|
|
|
### Requirement: Evaluation harness SHALL validate Executor V2 audit closure
|
|
The evaluation harness SHALL be able to validate the V2 audit chain from Executor structured output through Gatekeeper, Verifier claim checks, Composer output, and final user answer.
|
|
|
|
#### Scenario: Required V2 audit fields are present
|
|
- **WHEN** an evaluation case requires V2 audit closure
|
|
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.gatekeeper_result` exists
|
|
- **AND** it SHALL verify that `selfEvaluation.verifier_evaluation.claim_checks` exists
|
|
- **AND** it SHALL verify that `selfEvaluation.verifier_evaluation.composer_output` exists
|
|
|
|
#### Scenario: Gatekeeper failure cannot pass verification
|
|
- **WHEN** a trace has `gatekeeper_result.status` equal to `fail`
|
|
- **THEN** the evaluator SHALL fail the case if `selfEvaluation.verifier_evaluation.verdict` is `PASS`
|
|
|
|
#### Scenario: Claim checks are auditable
|
|
- **WHEN** an evaluation case requires claim checks
|
|
- **THEN** the evaluator SHALL verify that each claim check includes `claim_id`, `verification`, and `detail`
|
|
- **AND** each verification SHALL be one of the V2 claim verification values recognized by the Verifier contract
|
|
|
|
#### Scenario: Composer output is auditable
|
|
- **WHEN** an evaluation case requires Composer output
|
|
- **THEN** the evaluator SHALL verify that `composer_output` records whether parsed output or fallback rendering was used
|
|
|
|
### Requirement: Evaluation harness SHALL prevent unsupported claims from leaking into final answers
|
|
The evaluation harness SHALL detect configured unsafe or unsupported claim text when it appears in the final user-facing answer.
|
|
|
|
#### Scenario: Unsupported final-answer claim is rejected
|
|
- **WHEN** an evaluation case declares forbidden confirmed-claim keywords
|
|
- **THEN** the evaluator SHALL fail the case if the final answer contains any of those keywords
|
|
|
|
#### Scenario: Raw Executor JSON is not user-facing
|
|
- **WHEN** a trace is evaluated under V2 audit closure
|
|
- **THEN** the evaluator SHALL fail the case if the final answer contains raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`
|
|
|
|
#### Scenario: Composer fallback still avoids raw JSON leakage
|
|
- **WHEN** a trace records Composer fallback rendering
|
|
- **THEN** the evaluator SHALL still enforce final-answer raw JSON leakage checks
|