Files
SuperBizAgent-java/mvp/eval

Diagnosis Eval Harness

This folder contains the fixed offline regression set for the MVP diagnosis Agent.

Background

The current complex Chat diagnosis chain is a bounded StateGraph:

Planner -> Executor -> Gatekeeper -> Verified Input -> Verifier -> Composer -> final answer
                                      |                         |
                                      + bounded evidence retry  + safe Fallback

Executor V2 structured output, deterministic Gatekeeper audit, verified-only Verifier input, claim_checks, Composer rendering, and StateGraph routing are covered by deterministic tests so future prompt, tool, or graph changes can be checked without relying on a one-off demo.

Scope

  • Case definitions: cases/diagnosis-cases.json
  • Offline trace fixtures: fixtures/*.json
  • Field definitions: schema.md
  • Baseline reports: reports/baseline-report.json and reports/baseline-report.md
  • Baseline diff sample: reports/baseline-diff-sample.json and reports/baseline-diff-sample.md
  • Evaluator implementation: DiagnosisTraceEvaluator
  • Report writer: DiagnosisEvalReportWriter

Current Mode

The baseline evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.

The committed baseline currently contains:

12 fixed cases
12 passing fixture evaluations
5 PASS verdicts
6 LOW_CONFID verdicts
1 REJECT verdict

The V2 evidence-pipeline matrix covers:

  • Positive supported evidence for a narrow HighCPU observation.
  • No-evidence negative_observation using $.no_evidence.
  • Gatekeeper failure for a fabricated tool invocation reference.
  • Unsupported claim filtering before the final answer.
  • Composer fallback rendering without raw Executor JSON leakage.
  • Gatekeeper rule set version audit for new matrix fixtures.
  • Prompt audit version checks for planner, executor, verifier, and composer prompts.
  • Gatekeeper rule metadata checks for enabled rule id and default severity.

Verification

Run the focused evaluator test:

mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test

Run the authoritative Graph layers plus the fixed evaluator checks:

mvn -q "-Dtest=DiagnosisGraphWorkflowTest,DiagnosisGraphNodeContractTest,ChatServiceGraphIntegrationTest,DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest" test

When fixtures or evaluator rules change, regenerate both baseline reports from the same case file and fixture directory, then update JSON and Markdown together.

Regression Signal

The harness is deterministic code, not an LLM judge:

fixed diagnosis case
-> saved trace fixture
-> rule-based trace validation
-> JSON / Markdown report
-> regression signal for prompts, tools, retrieval, verifier, and composer behavior

Stage 5 adds these V2 checks:

  • gatekeeper_result.status=fail cannot coexist with Verifier PASS.
  • Required V2 fixtures must include gatekeeper_result, claim_checks, and composer_output.
  • claim_checks must be structurally auditable.
  • Composer output must record whether normal parsing or fallback rendering was used.
  • Gatekeeper rule set version can be asserted per fixture.
  • Prompt audit version and per-prompt versions can be asserted per fixture.
  • Gatekeeper rule metadata can be required per fixture.
  • Final answers must not leak raw Executor protocol markers such as executor_evidence_v2, answer_version, evidence_bindings, or claim_id.
  • Configured unsupported claim keywords must not appear as confirmed final-answer content.

Baseline Diff

Baseline diff compares a current report against reports/baseline-report.json.

baseline report
current report
-> deterministic diff
-> regressions, improvements, and changed signals

Use it to answer: did a prompt, tool, retrieval, verifier, or composer change make the Agent worse than the fixed baseline?