104 lines
3.9 KiB
Markdown
104 lines
3.9 KiB
Markdown
# Diagnosis Eval Harness
|
|
|
|
This folder contains the fixed offline regression set for the MVP diagnosis Agent.
|
|
|
|
## Background
|
|
|
|
The current complex Chat diagnosis chain is a bounded StateGraph:
|
|
|
|
```text
|
|
Planner -> Executor -> Gatekeeper -> Verified Input -> Verifier -> Composer -> final answer
|
|
| |
|
|
+ bounded evidence retry + safe Fallback
|
|
```
|
|
|
|
Executor V2 structured output, deterministic Gatekeeper audit, verified-only Verifier input, `claim_checks`, Composer rendering, and StateGraph routing are covered by deterministic tests so future prompt, tool, or graph changes can be checked without relying on a one-off demo.
|
|
|
|
## Scope
|
|
|
|
- Case definitions: `cases/diagnosis-cases.json`
|
|
- Offline trace fixtures: `fixtures/*.json`
|
|
- Field definitions: `schema.md`
|
|
- Baseline reports: `reports/baseline-report.json` and `reports/baseline-report.md`
|
|
- Baseline diff sample: `reports/baseline-diff-sample.json` and `reports/baseline-diff-sample.md`
|
|
- Evaluator implementation: `DiagnosisTraceEvaluator`
|
|
- Report writer: `DiagnosisEvalReportWriter`
|
|
|
|
## Current Mode
|
|
|
|
The baseline evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.
|
|
|
|
The committed baseline currently contains:
|
|
|
|
```text
|
|
12 fixed cases
|
|
12 passing fixture evaluations
|
|
5 PASS verdicts
|
|
6 LOW_CONFID verdicts
|
|
1 REJECT verdict
|
|
```
|
|
|
|
The V2 evidence-pipeline matrix covers:
|
|
|
|
- Positive supported evidence for a narrow HighCPU observation.
|
|
- No-evidence `negative_observation` using `$.no_evidence`.
|
|
- Gatekeeper failure for a fabricated tool invocation reference.
|
|
- Unsupported claim filtering before the final answer.
|
|
- Composer fallback rendering without raw Executor JSON leakage.
|
|
- Gatekeeper rule set version audit for new matrix fixtures.
|
|
- Prompt audit version checks for planner, executor, verifier, and composer prompts.
|
|
- Gatekeeper rule metadata checks for enabled rule id and default severity.
|
|
|
|
## Verification
|
|
|
|
Run the focused evaluator test:
|
|
|
|
```powershell
|
|
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test
|
|
```
|
|
|
|
Run the authoritative Graph layers plus the fixed evaluator checks:
|
|
|
|
```powershell
|
|
mvn -q "-Dtest=DiagnosisGraphWorkflowTest,DiagnosisGraphNodeContractTest,ChatServiceGraphIntegrationTest,DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest" test
|
|
```
|
|
|
|
When fixtures or evaluator rules change, regenerate both baseline reports from the same case file and fixture directory, then update JSON and Markdown together.
|
|
|
|
## Regression Signal
|
|
|
|
The harness is deterministic code, not an LLM judge:
|
|
|
|
```text
|
|
fixed diagnosis case
|
|
-> saved trace fixture
|
|
-> rule-based trace validation
|
|
-> JSON / Markdown report
|
|
-> regression signal for prompts, tools, retrieval, verifier, and composer behavior
|
|
```
|
|
|
|
Stage 5 adds these V2 checks:
|
|
|
|
- `gatekeeper_result.status=fail` cannot coexist with Verifier `PASS`.
|
|
- Required V2 fixtures must include `gatekeeper_result`, `claim_checks`, and `composer_output`.
|
|
- `claim_checks` must be structurally auditable.
|
|
- Composer output must record whether normal parsing or fallback rendering was used.
|
|
- Gatekeeper rule set version can be asserted per fixture.
|
|
- Prompt audit version and per-prompt versions can be asserted per fixture.
|
|
- Gatekeeper rule metadata can be required per fixture.
|
|
- Final answers must not leak raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`.
|
|
- Configured unsupported claim keywords must not appear as confirmed final-answer content.
|
|
|
|
## Baseline Diff
|
|
|
|
Baseline diff compares a current report against `reports/baseline-report.json`.
|
|
|
|
```text
|
|
baseline report
|
|
current report
|
|
-> deterministic diff
|
|
-> regressions, improvements, and changed signals
|
|
```
|
|
|
|
Use it to answer: did a prompt, tool, retrieval, verifier, or composer change make the Agent worse than the fixed baseline?
|