Files
SuperBizAgent-java/openspec/specs/diagnosis-eval-harness/spec.md
T
2026-07-05 00:27:57 +08:00

4.6 KiB

diagnosis-eval-harness Specification

Purpose

Provide a repeatable offline evaluation harness for MVP diagnosis Agent behavior, so prompt, tool, retrieval, and verifier changes can be checked against fixed trace-based regression cases.

Requirements

Requirement: Evaluation harness SHALL define fixed diagnosis cases

The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria.

Scenario: Case definition includes expected evidence

  • WHEN an evaluation case is defined
  • THEN it SHALL include a case id, user question, expected root-cause keywords, required evidence tools, and allowed verifier verdicts

Scenario: Case definition can express forbidden behavior

  • WHEN a case has known unsafe behavior
  • THEN the case definition SHALL be able to declare forbidden answer keywords or forbidden verdicts

Requirement: Evaluation harness SHALL validate diagnosis traces

The system SHALL validate a diagnosis trace against the corresponding case definition using deterministic rules.

Scenario: Evidence coverage validation

  • WHEN a trace is evaluated
  • THEN the evaluator SHALL verify that required evidence tools appear in toolInvocations or verifier trace summaries

Scenario: Verifier evaluation validation

  • WHEN a trace is evaluated
  • THEN the evaluator SHALL verify that selfEvaluation.verifier_evaluation.verdict exists
  • AND the verdict SHALL be one of the case's allowed verdicts

Scenario: Answer keyword validation

  • WHEN a trace is evaluated
  • THEN the evaluator SHALL verify that the final answer includes at least the configured minimum root-cause keyword coverage

Scenario: Degraded output validation

  • WHEN a trace verdict is REJECT
  • THEN the evaluator SHALL verify that the final answer follows the degraded-output contract rather than leaking an unverified executor answer

Requirement: Evaluation harness SHALL report quality and cost signals

The system SHALL produce a report that summarizes pass/fail results and key trace metrics.

Scenario: JSON report output

  • WHEN an evaluation run completes
  • THEN the evaluator SHALL output a JSON report with per-case result, failed checks, verdict, evidence coverage, tool-call count, and duration

Scenario: Markdown report output

  • WHEN an evaluation run completes
  • THEN the evaluator SHALL output a Markdown report suitable for review in the repository

Scenario: Aggregate metrics

  • WHEN multiple cases are evaluated
  • THEN the report SHALL include aggregate pass rate, verdict distribution, average tool-call count, and average duration when available

Requirement: Evaluation harness SHALL support offline fixture mode

The first evaluator version SHALL be runnable without live MySQL, Redis, Milvus, or LLM services by evaluating saved trace fixtures.

Scenario: Fixture trace evaluation

  • WHEN the evaluator is run against a directory of trace fixture files
  • THEN it SHALL evaluate each trace file against its matching case definition
  • AND it SHALL not require a running application service

Scenario: Missing fixture is reported clearly

  • WHEN a case has no matching trace fixture
  • THEN the evaluator SHALL mark the case as not run or failed with a clear reason

Requirement: Evaluation harness SHALL provide complete fixture coverage for fixed cases

The system SHALL include an offline trace fixture for every fixed MVP diagnosis evaluation case.

Scenario: Every case resolves to a fixture file

  • WHEN the evaluator loads the fixed case definition file
  • THEN every case's traceFixture value SHALL resolve to an existing JSON fixture file

Scenario: Fixture files are loadable as diagnosis traces

  • WHEN each referenced fixture is loaded
  • THEN it SHALL deserialize into the trace response shape used by the evaluator

Requirement: Evaluation harness SHALL preserve a reproducible baseline report

The system SHALL preserve a generated baseline report for the full fixed fixture set.

Scenario: Baseline report includes all fixed cases

  • WHEN the baseline report is generated from the fixed case file and fixture directory
  • THEN the report SHALL include one result for every fixed case

Scenario: Baseline report is reviewable

  • WHEN the baseline report is written
  • THEN it SHALL be available in JSON and Markdown formats under the eval documentation area

Scenario: Baseline regeneration is documented

  • WHEN a developer changes fixtures or evaluator rules
  • THEN the eval documentation SHALL explain how to regenerate the baseline report