Files
SuperBizAgent-java/openspec/specs/diagnosis-eval-harness/spec.md
T

173 lines
9.1 KiB
Markdown

# diagnosis-eval-harness Specification
## Purpose
Provide a repeatable offline evaluation harness for MVP diagnosis Agent behavior, so prompt, tool, retrieval, and verifier changes can be checked against fixed trace-based regression cases.
## Requirements
### Requirement: Evaluation harness SHALL define fixed diagnosis cases
The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria.
#### Scenario: Case definition includes expected evidence
- **WHEN** an evaluation case is defined
- **THEN** it SHALL include a case id, user question, expected root-cause keywords, required evidence tools, and allowed verifier verdicts
#### Scenario: Case definition can express forbidden behavior
- **WHEN** a case has known unsafe behavior
- **THEN** the case definition SHALL be able to declare forbidden answer keywords or forbidden verdicts
### Requirement: Evaluation harness SHALL validate diagnosis traces
The system SHALL validate a diagnosis trace against the corresponding case definition using deterministic rules.
#### Scenario: Evidence coverage validation
- **WHEN** a trace is evaluated
- **THEN** the evaluator SHALL verify that required evidence tools appear in `toolInvocations` or verifier trace summaries
#### Scenario: Verifier evaluation validation
- **WHEN** a trace is evaluated
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.verdict` exists
- **AND** the verdict SHALL be one of the case's allowed verdicts
#### Scenario: Answer keyword validation
- **WHEN** a trace is evaluated
- **THEN** the evaluator SHALL verify that the final answer includes at least the configured minimum root-cause keyword coverage
#### Scenario: Degraded output validation
- **WHEN** a trace verdict is `REJECT`
- **THEN** the evaluator SHALL verify that the final answer follows the degraded-output contract rather than leaking an unverified executor answer
### Requirement: Evaluation harness SHALL report quality and cost signals
The system SHALL produce a report that summarizes pass/fail results and key trace metrics.
#### Scenario: JSON report output
- **WHEN** an evaluation run completes
- **THEN** the evaluator SHALL output a JSON report with per-case result, failed checks, verdict, evidence coverage, tool-call count, and duration
#### Scenario: Markdown report output
- **WHEN** an evaluation run completes
- **THEN** the evaluator SHALL output a Markdown report suitable for review in the repository
#### Scenario: Aggregate metrics
- **WHEN** multiple cases are evaluated
- **THEN** the report SHALL include aggregate pass rate, verdict distribution, average tool-call count, and average duration when available
### Requirement: Evaluation harness SHALL support offline fixture mode
The first evaluator version SHALL be runnable without live MySQL, Redis, Milvus, or LLM services by evaluating saved trace fixtures.
#### Scenario: Fixture trace evaluation
- **WHEN** the evaluator is run against a directory of trace fixture files
- **THEN** it SHALL evaluate each trace file against its matching case definition
- **AND** it SHALL not require a running application service
#### Scenario: Missing fixture is reported clearly
- **WHEN** a case has no matching trace fixture
- **THEN** the evaluator SHALL mark the case as not run or failed with a clear reason
### Requirement: Evaluation harness SHALL provide complete fixture coverage for fixed cases
The system SHALL include an offline trace fixture for every fixed MVP diagnosis evaluation case.
#### Scenario: Every case resolves to a fixture file
- **WHEN** the evaluator loads the fixed case definition file
- **THEN** every case's `traceFixture` value SHALL resolve to an existing JSON fixture file
#### Scenario: Fixture files are loadable as diagnosis traces
- **WHEN** each referenced fixture is loaded
- **THEN** it SHALL deserialize into the trace response shape used by the evaluator
### Requirement: Evaluation harness SHALL preserve a reproducible baseline report
The system SHALL preserve a generated baseline report for the full fixed fixture set.
#### Scenario: Baseline report includes all fixed cases
- **WHEN** the baseline report is generated from the fixed case file and fixture directory
- **THEN** the report SHALL include one result for every fixed case
#### Scenario: Baseline report is reviewable
- **WHEN** the baseline report is written
- **THEN** it SHALL be available in JSON and Markdown formats under the eval documentation area
#### Scenario: Baseline regeneration is documented
- **WHEN** a developer changes fixtures or evaluator rules
- **THEN** the eval documentation SHALL explain how to regenerate the baseline report
### Requirement: Evaluation harness SHALL compare reports against a baseline
The system SHALL compare a current diagnosis evaluation report against a saved baseline report using deterministic rules.
#### Scenario: Aggregate regression detection
- **WHEN** the current report has a lower pass rate than the baseline report
- **THEN** the diff SHALL record a regression with the old value, new value, and delta
#### Scenario: Cost signal detection
- **WHEN** average tool-call count or average duration changes between reports
- **THEN** the diff SHALL record the baseline value, current value, and delta
#### Scenario: Verdict distribution comparison
- **WHEN** verdict counts differ between reports
- **THEN** the diff SHALL record the verdict distribution changes
### Requirement: Evaluation harness SHALL compare case-level report results
The system SHALL compare case results by case id and report actionable per-case changes.
#### Scenario: Case pass/fail regression
- **WHEN** a case changes from passing in the baseline to failing in the current report
- **THEN** the diff SHALL record a regression for that case
#### Scenario: Evidence coverage regression
- **WHEN** a required evidence tool changes from covered to uncovered for a case
- **THEN** the diff SHALL record a regression naming the case and tool
#### Scenario: Missing case detection
- **WHEN** a baseline case is absent from the current report
- **THEN** the diff SHALL record a regression for the missing case
#### Scenario: New case detection
- **WHEN** a current report contains a case absent from the baseline
- **THEN** the diff SHALL record the case as a non-regression change
### Requirement: Evaluation harness SHALL report baseline diff results
The system SHALL expose baseline diff output in structured JSON and reviewable Markdown formats.
#### Scenario: JSON diff output
- **WHEN** a baseline diff is written as JSON
- **THEN** it SHALL include aggregate summary fields and detailed diff items
#### Scenario: Markdown diff output
- **WHEN** a baseline diff is written as Markdown
- **THEN** it SHALL include a readable summary and a table of diff items
### Requirement: Evaluation harness SHALL validate Executor V2 audit closure
The evaluation harness SHALL be able to validate the V2 audit chain from Executor structured output through Gatekeeper, Verifier claim checks, Composer output, and final user answer.
#### Scenario: Required V2 audit fields are present
- **WHEN** an evaluation case requires V2 audit closure
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.gatekeeper_result` exists
- **AND** it SHALL verify that `selfEvaluation.verifier_evaluation.claim_checks` exists
- **AND** it SHALL verify that `selfEvaluation.verifier_evaluation.composer_output` exists
#### Scenario: Gatekeeper failure cannot pass verification
- **WHEN** a trace has `gatekeeper_result.status` equal to `fail`
- **THEN** the evaluator SHALL fail the case if `selfEvaluation.verifier_evaluation.verdict` is `PASS`
#### Scenario: Claim checks are auditable
- **WHEN** an evaluation case requires claim checks
- **THEN** the evaluator SHALL verify that each claim check includes `claim_id`, `verification`, and `detail`
- **AND** each verification SHALL be one of the V2 claim verification values recognized by the Verifier contract
#### Scenario: Composer output is auditable
- **WHEN** an evaluation case requires Composer output
- **THEN** the evaluator SHALL verify that `composer_output` records whether parsed output or fallback rendering was used
### Requirement: Evaluation harness SHALL prevent unsupported claims from leaking into final answers
The evaluation harness SHALL detect configured unsafe or unsupported claim text when it appears in the final user-facing answer.
#### Scenario: Unsupported final-answer claim is rejected
- **WHEN** an evaluation case declares forbidden confirmed-claim keywords
- **THEN** the evaluator SHALL fail the case if the final answer contains any of those keywords
#### Scenario: Raw Executor JSON is not user-facing
- **WHEN** a trace is evaluated under V2 audit closure
- **THEN** the evaluator SHALL fail the case if the final answer contains raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`
#### Scenario: Composer fallback still avoids raw JSON leakage
- **WHEN** a trace records Composer fallback rendering
- **THEN** the evaluator SHALL still enforce final-answer raw JSON leakage checks