feat(eval): add executor audit closure checks
This commit is contained in:
+46
-14
@@ -1,6 +1,16 @@
|
||||
# Diagnosis Eval Harness
|
||||
|
||||
This folder contains the first fixed-case evaluation set for the MVP diagnosis Agent.
|
||||
This folder contains the fixed offline regression set for the MVP diagnosis Agent.
|
||||
|
||||
## Background
|
||||
|
||||
The current diagnosis chain is:
|
||||
|
||||
```text
|
||||
Planner -> Executor -> Gatekeeper -> Verifier -> Composer -> final answer
|
||||
```
|
||||
|
||||
Stages 1-4 introduced Executor V2 structured output, deterministic Gatekeeper audit, Verifier `claim_checks`, and Composer final-answer rendering. Stage 5 makes those audit fields part of the offline regression harness so future prompt, tool, or chain changes can be checked without relying on a one-off demo.
|
||||
|
||||
## Scope
|
||||
|
||||
@@ -14,7 +24,23 @@ This folder contains the first fixed-case evaluation set for the MVP diagnosis A
|
||||
|
||||
## Current Mode
|
||||
|
||||
The first version evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.
|
||||
The baseline evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.
|
||||
|
||||
The committed baseline currently contains:
|
||||
|
||||
```text
|
||||
8 fixed cases
|
||||
8 passing fixture evaluations
|
||||
2 PASS verdicts
|
||||
5 LOW_CONFID verdicts
|
||||
1 REJECT verdict
|
||||
```
|
||||
|
||||
The three V2 audit-closure cases cover:
|
||||
|
||||
- Gatekeeper failure for a fabricated tool invocation reference.
|
||||
- Unsupported claim filtering before the final answer.
|
||||
- Composer fallback rendering without raw Executor JSON leakage.
|
||||
|
||||
## Verification
|
||||
|
||||
@@ -24,29 +50,35 @@ Run the focused evaluator test:
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test
|
||||
```
|
||||
|
||||
The committed baseline report represents the current fixed fixture set:
|
||||
Run the broader phase-5 regression set:
|
||||
|
||||
```text
|
||||
5 fixed cases
|
||||
5 passing fixture evaluations
|
||||
2 PASS verdicts
|
||||
3 LOW_CONFID verdicts
|
||||
```powershell
|
||||
mvn "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
|
||||
```
|
||||
|
||||
When fixtures or evaluator rules change, regenerate the report from the same case file and fixture directory, then update both JSON and Markdown outputs together.
|
||||
When fixtures or evaluator rules change, regenerate both baseline reports from the same case file and fixture directory, then update JSON and Markdown together.
|
||||
|
||||
## Interview Story
|
||||
## Regression Signal
|
||||
|
||||
The harness gives the MVP a repeatable baseline:
|
||||
The harness is deterministic code, not an LLM judge:
|
||||
|
||||
```text
|
||||
fixed diagnosis case
|
||||
-> saved or runtime trace
|
||||
-> saved trace fixture
|
||||
-> rule-based trace validation
|
||||
-> JSON / Markdown report
|
||||
-> regression signal for prompts, tools, retrieval, and verifier behavior
|
||||
-> regression signal for prompts, tools, retrieval, verifier, and composer behavior
|
||||
```
|
||||
|
||||
Stage 5 adds these V2 checks:
|
||||
|
||||
- `gatekeeper_result.status=fail` cannot coexist with Verifier `PASS`.
|
||||
- Required V2 fixtures must include `gatekeeper_result`, `claim_checks`, and `composer_output`.
|
||||
- `claim_checks` must be structurally auditable.
|
||||
- Composer output must record whether normal parsing or fallback rendering was used.
|
||||
- Final answers must not leak raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`.
|
||||
- Configured unsupported claim keywords must not appear as confirmed final-answer content.
|
||||
|
||||
## Baseline Diff
|
||||
|
||||
Baseline diff compares a current report against `reports/baseline-report.json`.
|
||||
@@ -58,4 +90,4 @@ current report
|
||||
-> regressions, improvements, and changed signals
|
||||
```
|
||||
|
||||
Use it to answer: did a prompt, tool, retrieval, or verifier change make the Agent worse than the fixed baseline?
|
||||
Use it to answer: did a prompt, tool, retrieval, verifier, or composer change make the Agent worse than the fixed baseline?
|
||||
|
||||
Reference in New Issue
Block a user