Files

12 KiB

diagnosis-eval-harness Specification

Purpose

Provide a repeatable offline evaluation harness for MVP diagnosis Agent behavior, so prompt, tool, retrieval, and verifier changes can be checked against fixed trace-based regression cases.

Requirements

Requirement: Evaluation harness SHALL define fixed diagnosis cases

The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria.

Scenario: Case definition includes expected evidence

  • WHEN an evaluation case is defined
  • THEN it SHALL include a case id, user question, expected root-cause keywords, required evidence tools, and allowed verifier verdicts

Scenario: Case definition can express forbidden behavior

  • WHEN a case has known unsafe behavior
  • THEN the case definition SHALL be able to declare forbidden answer keywords or forbidden verdicts

Requirement: Evaluation harness SHALL validate diagnosis traces

The system SHALL validate a diagnosis trace against the corresponding case definition using deterministic rules.

Scenario: Evidence coverage validation

  • WHEN a trace is evaluated
  • THEN the evaluator SHALL verify that required evidence tools appear in toolInvocations or verifier trace summaries

Scenario: Verifier evaluation validation

  • WHEN a trace is evaluated
  • THEN the evaluator SHALL verify that selfEvaluation.verifier_evaluation.verdict exists
  • AND the verdict SHALL be one of the case's allowed verdicts

Scenario: Answer keyword validation

  • WHEN a trace is evaluated
  • THEN the evaluator SHALL verify that the final answer includes at least the configured minimum root-cause keyword coverage

Scenario: Degraded output validation

  • WHEN a trace verdict is REJECT
  • THEN the evaluator SHALL verify that the final answer follows the degraded-output contract rather than leaking an unverified executor answer

Requirement: Evaluation harness SHALL report quality and cost signals

The system SHALL produce a report that summarizes pass/fail results and key trace metrics.

Scenario: JSON report output

  • WHEN an evaluation run completes
  • THEN the evaluator SHALL output a JSON report with per-case result, failed checks, verdict, evidence coverage, tool-call count, and duration

Scenario: Markdown report output

  • WHEN an evaluation run completes
  • THEN the evaluator SHALL output a Markdown report suitable for review in the repository

Scenario: Aggregate metrics

  • WHEN multiple cases are evaluated
  • THEN the report SHALL include aggregate pass rate, verdict distribution, average tool-call count, and average duration when available

Requirement: Evaluation harness SHALL support offline fixture mode

The first evaluator version SHALL be runnable without live MySQL, Redis, Milvus, or LLM services by evaluating saved trace fixtures.

Scenario: Fixture trace evaluation

  • WHEN the evaluator is run against a directory of trace fixture files
  • THEN it SHALL evaluate each trace file against its matching case definition
  • AND it SHALL not require a running application service

Scenario: Missing fixture is reported clearly

  • WHEN a case has no matching trace fixture
  • THEN the evaluator SHALL mark the case as not run or failed with a clear reason

Requirement: Evaluation harness SHALL provide complete fixture coverage for fixed cases

The system SHALL include an offline trace fixture for every fixed MVP diagnosis evaluation case.

Scenario: Every case resolves to a fixture file

  • WHEN the evaluator loads the fixed case definition file
  • THEN every case's traceFixture value SHALL resolve to an existing JSON fixture file

Scenario: Fixture files are loadable as diagnosis traces

  • WHEN each referenced fixture is loaded
  • THEN it SHALL deserialize into the trace response shape used by the evaluator

Requirement: Evaluation harness SHALL preserve a reproducible baseline report

The system SHALL preserve a generated baseline report for the full fixed fixture set.

Scenario: Baseline report includes all fixed cases

  • WHEN the baseline report is generated from the fixed case file and fixture directory
  • THEN the report SHALL include one result for every fixed case

Scenario: Baseline report is reviewable

  • WHEN the baseline report is written
  • THEN it SHALL be available in JSON and Markdown formats under the eval documentation area

Scenario: Baseline regeneration is documented

  • WHEN a developer changes fixtures or evaluator rules
  • THEN the eval documentation SHALL explain how to regenerate the baseline report

Requirement: Evaluation harness SHALL compare reports against a baseline

The system SHALL compare a current diagnosis evaluation report against a saved baseline report using deterministic rules.

Scenario: Aggregate regression detection

  • WHEN the current report has a lower pass rate than the baseline report
  • THEN the diff SHALL record a regression with the old value, new value, and delta

Scenario: Cost signal detection

  • WHEN average tool-call count or average duration changes between reports
  • THEN the diff SHALL record the baseline value, current value, and delta

Scenario: Verdict distribution comparison

  • WHEN verdict counts differ between reports
  • THEN the diff SHALL record the verdict distribution changes

Requirement: Evaluation harness SHALL compare case-level report results

The system SHALL compare case results by case id and report actionable per-case changes.

Scenario: Case pass/fail regression

  • WHEN a case changes from passing in the baseline to failing in the current report
  • THEN the diff SHALL record a regression for that case

Scenario: Evidence coverage regression

  • WHEN a required evidence tool changes from covered to uncovered for a case
  • THEN the diff SHALL record a regression naming the case and tool

Scenario: Missing case detection

  • WHEN a baseline case is absent from the current report
  • THEN the diff SHALL record a regression for the missing case

Scenario: New case detection

  • WHEN a current report contains a case absent from the baseline
  • THEN the diff SHALL record the case as a non-regression change

Requirement: Evaluation harness SHALL report baseline diff results

The system SHALL expose baseline diff output in structured JSON and reviewable Markdown formats.

Scenario: JSON diff output

  • WHEN a baseline diff is written as JSON
  • THEN it SHALL include aggregate summary fields and detailed diff items

Scenario: Markdown diff output

  • WHEN a baseline diff is written as Markdown
  • THEN it SHALL include a readable summary and a table of diff items

Requirement: Evaluation harness SHALL validate Executor V2 audit closure

The evaluation harness SHALL be able to validate the V2 audit chain from Executor structured output through Gatekeeper, Verifier claim checks, Composer output, and final user answer.

Scenario: Required V2 audit fields are present

  • WHEN an evaluation case requires V2 audit closure
  • THEN the evaluator SHALL verify that selfEvaluation.verifier_evaluation.gatekeeper_result exists
  • AND it SHALL verify that selfEvaluation.verifier_evaluation.claim_checks exists
  • AND it SHALL verify that selfEvaluation.verifier_evaluation.composer_output exists

Scenario: Gatekeeper failure cannot pass verification

  • WHEN a trace has gatekeeper_result.status equal to fail
  • THEN the evaluator SHALL fail the case if selfEvaluation.verifier_evaluation.verdict is PASS

Scenario: Claim checks are auditable

  • WHEN an evaluation case requires claim checks
  • THEN the evaluator SHALL verify that each claim check includes claim_id, verification, and detail
  • AND each verification SHALL be one of the V2 claim verification values recognized by the Verifier contract

Scenario: Composer output is auditable

  • WHEN an evaluation case requires Composer output
  • THEN the evaluator SHALL verify that composer_output records whether parsed output or fallback rendering was used

Requirement: Evaluation harness SHALL prevent unsupported claims from leaking into final answers

The evaluation harness SHALL detect configured unsafe or unsupported claim text when it appears in the final user-facing answer.

Scenario: Unsupported final-answer claim is rejected

  • WHEN an evaluation case declares forbidden confirmed-claim keywords
  • THEN the evaluator SHALL fail the case if the final answer contains any of those keywords

Scenario: Raw Executor JSON is not user-facing

  • WHEN a trace is evaluated under V2 audit closure
  • THEN the evaluator SHALL fail the case if the final answer contains raw Executor protocol markers such as executor_evidence_v2, answer_version, evidence_bindings, or claim_id

Scenario: Composer fallback still avoids raw JSON leakage

  • WHEN a trace records Composer fallback rendering
  • THEN the evaluator SHALL still enforce final-answer raw JSON leakage checks

Requirement: Evaluation harness SHALL expose an evidence-pipeline matrix

The evaluation harness SHALL include fixed cases that demonstrate the current Chat evidence pipeline across positive evidence, narrow-scope observation, no-evidence negative observation, low-confidence filtering, reject safety, and composer fallback behavior.

Scenario: Matrix cases are fixture backed

  • WHEN the fixed diagnosis case file is evaluated
  • THEN each matrix case SHALL resolve to an offline trace fixture
  • AND evaluation SHALL not require a live LLM or running application

Scenario: Matrix cases preserve V2 audit closure

  • WHEN a matrix case requires V2 audit closure
  • THEN its fixture SHALL include gatekeeper_result
  • AND it SHALL include claim_checks
  • AND it SHALL include composer_output

Requirement: Evaluation harness SHALL validate Gatekeeper rule set version when requested

The evaluation harness SHALL be able to assert the Gatekeeper rule set version recorded in a fixture.

Scenario: Expected rule set version matches

  • WHEN an evaluation case declares expectedGatekeeperRuleSetVersion
  • AND the fixture has the same selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version
  • THEN the rule set version check SHALL pass

Scenario: Expected rule set version mismatches

  • WHEN an evaluation case declares expectedGatekeeperRuleSetVersion
  • AND the fixture has a different or missing rule set version
  • THEN the case SHALL fail with a clear failed check

Requirement: Diagnosis eval SHALL validate trace fixtures deterministically

The diagnosis eval harness SHALL evaluate saved trace fixtures without invoking an LLM judge.

Scenario: prompt audit assertions are enforced

  • WHEN an eval case sets requirePromptAudit=true
  • THEN the evaluator SHALL require verifier_evaluation.prompt_audit.version
  • AND when expectedPromptAuditVersion is configured, it SHALL match exactly
  • AND when expectedPromptVersions is configured, each configured prompt name SHALL appear with the expected version

Scenario: Gatekeeper rule metadata assertions are enforced

  • WHEN an eval case sets requireGatekeeperRules=true
  • THEN the evaluator SHALL require verifier_evaluation.gatekeeper_result.rules to be non-empty
  • AND each rule item SHALL include id, enabled, and default_severity

Scenario: expanded baseline remains passing

  • WHEN the committed fixture set is evaluated
  • THEN every case SHALL pass
  • AND baseline JSON and Markdown reports SHALL reflect the expanded case count and verdict distribution