feat(eval): add executor audit closure checks

This commit is contained in:
aruo
2026-07-08 10:21:39 +08:00
parent 6015bcbf6f
commit a08672b31e
20 changed files with 1142 additions and 188 deletions
+46 -14
View File
@@ -1,6 +1,16 @@
# Diagnosis Eval Harness
This folder contains the first fixed-case evaluation set for the MVP diagnosis Agent.
This folder contains the fixed offline regression set for the MVP diagnosis Agent.
## Background
The current diagnosis chain is:
```text
Planner -> Executor -> Gatekeeper -> Verifier -> Composer -> final answer
```
Stages 1-4 introduced Executor V2 structured output, deterministic Gatekeeper audit, Verifier `claim_checks`, and Composer final-answer rendering. Stage 5 makes those audit fields part of the offline regression harness so future prompt, tool, or chain changes can be checked without relying on a one-off demo.
## Scope
@@ -14,7 +24,23 @@ This folder contains the first fixed-case evaluation set for the MVP diagnosis A
## Current Mode
The first version evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.
The baseline evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.
The committed baseline currently contains:
```text
8 fixed cases
8 passing fixture evaluations
2 PASS verdicts
5 LOW_CONFID verdicts
1 REJECT verdict
```
The three V2 audit-closure cases cover:
- Gatekeeper failure for a fabricated tool invocation reference.
- Unsupported claim filtering before the final answer.
- Composer fallback rendering without raw Executor JSON leakage.
## Verification
@@ -24,29 +50,35 @@ Run the focused evaluator test:
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test
```
The committed baseline report represents the current fixed fixture set:
Run the broader phase-5 regression set:
```text
5 fixed cases
5 passing fixture evaluations
2 PASS verdicts
3 LOW_CONFID verdicts
```powershell
mvn "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
```
When fixtures or evaluator rules change, regenerate the report from the same case file and fixture directory, then update both JSON and Markdown outputs together.
When fixtures or evaluator rules change, regenerate both baseline reports from the same case file and fixture directory, then update JSON and Markdown together.
## Interview Story
## Regression Signal
The harness gives the MVP a repeatable baseline:
The harness is deterministic code, not an LLM judge:
```text
fixed diagnosis case
-> saved or runtime trace
-> saved trace fixture
-> rule-based trace validation
-> JSON / Markdown report
-> regression signal for prompts, tools, retrieval, and verifier behavior
-> regression signal for prompts, tools, retrieval, verifier, and composer behavior
```
Stage 5 adds these V2 checks:
- `gatekeeper_result.status=fail` cannot coexist with Verifier `PASS`.
- Required V2 fixtures must include `gatekeeper_result`, `claim_checks`, and `composer_output`.
- `claim_checks` must be structurally auditable.
- Composer output must record whether normal parsing or fallback rendering was used.
- Final answers must not leak raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`.
- Configured unsupported claim keywords must not appear as confirmed final-answer content.
## Baseline Diff
Baseline diff compares a current report against `reports/baseline-report.json`.
@@ -58,4 +90,4 @@ current report
-> regressions, improvements, and changed signals
```
Use it to answer: did a prompt, tool, retrieval, or verifier change make the Agent worse than the fixed baseline?
Use it to answer: did a prompt, tool, retrieval, verifier, or composer change make the Agent worse than the fixed baseline?