Files
SuperBizAgent-java/mvp/eval/README.md
T

104 lines
3.9 KiB
Markdown

# Diagnosis Eval Harness
This folder contains the fixed offline regression set for the MVP diagnosis Agent.
## Background
The current complex Chat diagnosis chain is a bounded StateGraph:
```text
Planner -> Executor -> Gatekeeper -> Verified Input -> Verifier -> Composer -> final answer
| |
+ bounded evidence retry + safe Fallback
```
Executor V2 structured output, deterministic Gatekeeper audit, verified-only Verifier input, `claim_checks`, Composer rendering, and StateGraph routing are covered by deterministic tests so future prompt, tool, or graph changes can be checked without relying on a one-off demo.
## Scope
- Case definitions: `cases/diagnosis-cases.json`
- Offline trace fixtures: `fixtures/*.json`
- Field definitions: `schema.md`
- Baseline reports: `reports/baseline-report.json` and `reports/baseline-report.md`
- Baseline diff sample: `reports/baseline-diff-sample.json` and `reports/baseline-diff-sample.md`
- Evaluator implementation: `DiagnosisTraceEvaluator`
- Report writer: `DiagnosisEvalReportWriter`
## Current Mode
The baseline evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.
The committed baseline currently contains:
```text
12 fixed cases
12 passing fixture evaluations
5 PASS verdicts
6 LOW_CONFID verdicts
1 REJECT verdict
```
The V2 evidence-pipeline matrix covers:
- Positive supported evidence for a narrow HighCPU observation.
- No-evidence `negative_observation` using `$.no_evidence`.
- Gatekeeper failure for a fabricated tool invocation reference.
- Unsupported claim filtering before the final answer.
- Composer fallback rendering without raw Executor JSON leakage.
- Gatekeeper rule set version audit for new matrix fixtures.
- Prompt audit version checks for planner, executor, verifier, and composer prompts.
- Gatekeeper rule metadata checks for enabled rule id and default severity.
## Verification
Run the focused evaluator test:
```powershell
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test
```
Run the authoritative Graph layers plus the fixed evaluator checks:
```powershell
mvn -q "-Dtest=DiagnosisGraphWorkflowTest,DiagnosisGraphNodeContractTest,ChatServiceGraphIntegrationTest,DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest" test
```
When fixtures or evaluator rules change, regenerate both baseline reports from the same case file and fixture directory, then update JSON and Markdown together.
## Regression Signal
The harness is deterministic code, not an LLM judge:
```text
fixed diagnosis case
-> saved trace fixture
-> rule-based trace validation
-> JSON / Markdown report
-> regression signal for prompts, tools, retrieval, verifier, and composer behavior
```
Stage 5 adds these V2 checks:
- `gatekeeper_result.status=fail` cannot coexist with Verifier `PASS`.
- Required V2 fixtures must include `gatekeeper_result`, `claim_checks`, and `composer_output`.
- `claim_checks` must be structurally auditable.
- Composer output must record whether normal parsing or fallback rendering was used.
- Gatekeeper rule set version can be asserted per fixture.
- Prompt audit version and per-prompt versions can be asserted per fixture.
- Gatekeeper rule metadata can be required per fixture.
- Final answers must not leak raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`.
- Configured unsupported claim keywords must not appear as confirmed final-answer content.
## Baseline Diff
Baseline diff compares a current report against `reports/baseline-report.json`.
```text
baseline report
current report
-> deterministic diff
-> regressions, improvements, and changed signals
```
Use it to answer: did a prompt, tool, retrieval, verifier, or composer change make the Agent worse than the fixed baseline?