feat(eval): add executor audit closure checks
This commit is contained in:
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-08
|
||||
@@ -0,0 +1,49 @@
|
||||
## Context
|
||||
|
||||
Executor Structured Output V2 introduced three audit layers after the original fixed-case eval harness was created:
|
||||
|
||||
- Gatekeeper result in `selfEvaluation.verifier_evaluation.gatekeeper_result`
|
||||
- Verifier claim-level checks in `selfEvaluation.verifier_evaluation.claim_checks`
|
||||
- Composer result in `selfEvaluation.verifier_evaluation.composer_output`
|
||||
|
||||
The existing harness proves that a saved trace has expected answer keywords, evidence tools, verdicts, and basic structured claim bindings. It does not yet prove that the V2 audit chain is internally consistent or that Composer filtered unsupported claims before writing the final user answer.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Make V2 audit fields part of offline regression checks.
|
||||
- Catch physical/protocol regressions: missing Gatekeeper audit, PASS after Gatekeeper fail, missing claim checks, missing Composer audit, raw Executor JSON leakage, and unsupported claims leaking into the final answer.
|
||||
- Add fixture cases that exercise the new checks without starting the application.
|
||||
- Keep the evaluator deterministic and easy to explain to another implementation agent.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not redesign Planner, Executor, Gatekeeper, Verifier, or Composer.
|
||||
- Do not introduce an LLM-based evaluator.
|
||||
- Do not require live MySQL, Redis, Milvus, or application startup for the baseline fixture tests.
|
||||
- Do not add a broad new database schema; trace audit remains read from existing JSON fields.
|
||||
|
||||
## Decisions
|
||||
|
||||
1. Extend eval case expectations instead of hardcoding every V2 rule globally.
|
||||
|
||||
Some legacy or intentionally partial fixtures may not contain all V2 audit fields. Case-level expectations let the fixed baseline explicitly say which trace must prove Gatekeeper, claim checks, Composer, or leakage prevention. The default can remain backward-compatible while new V2 fixtures opt in to stricter checks.
|
||||
|
||||
2. Keep consistency checks lexical and deterministic.
|
||||
|
||||
The evaluator will not determine semantic truth from scratch. It will compare configured keywords against the final answer and audit fields. This is enough to catch the intended regression class: unsupported or contradicted claims being rendered as confirmed final answers.
|
||||
|
||||
3. Treat Gatekeeper failure as a hard regression if paired with `PASS`.
|
||||
|
||||
The production chain already guards this. The eval harness should independently fail any trace where `gatekeeper_result.status=fail` and Verifier verdict remains `PASS`, because that means the audit layer can no longer be trusted.
|
||||
|
||||
4. Record audit signals in eval results.
|
||||
|
||||
Per-case results should expose the observed Gatekeeper status, Composer status, and claim-check count so baseline JSON/Markdown reports remain useful during review.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Keyword-based final-answer checks can miss paraphrases. -> Mitigation: use them only for regression-sensitive fixtures where the unsafe claim keyword is deliberately fixed.
|
||||
- [Risk] Adding too many case fields makes fixtures harder to maintain. -> Mitigation: keep expectation fields small and optional.
|
||||
- [Risk] Baseline report changes may look like a product behavior change. -> Mitigation: document that this phase changes only eval fixtures/rules unless a production bug is discovered and fixed.
|
||||
@@ -0,0 +1,37 @@
|
||||
## Why
|
||||
|
||||
Executor Structured Output V2 已经完成 Executor、Gatekeeper、Verifier、Composer 四段主链路改造,但现有离线评测仍主要检查最终答案关键词、证据工具覆盖和 Verifier verdict。阶段 5 需要把新增的审计字段纳入固定回归门禁,证明 `Planner -> Executor -> Gatekeeper -> Verifier -> Composer -> final answer` 链路不是只在单次 demo 中可用。
|
||||
|
||||
## What Changes
|
||||
|
||||
- Extend the diagnosis eval harness so V2 audit fields are validated as first-class regression checks.
|
||||
- Add fixture coverage for Gatekeeper failure, claim verification filtering, Composer fallback, and raw JSON leakage prevention.
|
||||
- Update baseline reports to reflect the expanded fixed fixture set.
|
||||
- Update eval documentation so another agent can understand the background, stages, data fields, and acceptance commands.
|
||||
- No production protocol change is intended in this phase; production chain behavior should remain unchanged unless tests reveal a bug.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `diagnosis-eval-harness`: add V2 audit-closure requirements for Gatekeeper, Verifier claim checks, Composer output, and final-answer consistency.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected code:
|
||||
- `src/main/java/com/superbiz/agent/eval/*`
|
||||
- `src/test/java/com/superbiz/agent/eval/*`
|
||||
- Affected data:
|
||||
- `mvp/eval/cases/diagnosis-cases.json`
|
||||
- `mvp/eval/fixtures/*.json`
|
||||
- `mvp/eval/reports/baseline-report.*`
|
||||
- `mvp/eval/schema.md`
|
||||
- `mvp/eval/README.md`
|
||||
- Verification:
|
||||
- focused evaluator tests
|
||||
- relevant Executor/Gatekeeper/Verifier/Composer integration tests
|
||||
- OpenSpec validation and archive
|
||||
+38
@@ -0,0 +1,38 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Evaluation harness SHALL validate Executor V2 audit closure
|
||||
The evaluation harness SHALL be able to validate the V2 audit chain from Executor structured output through Gatekeeper, Verifier claim checks, Composer output, and final user answer.
|
||||
|
||||
#### Scenario: Required V2 audit fields are present
|
||||
- **WHEN** an evaluation case requires V2 audit closure
|
||||
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.gatekeeper_result` exists
|
||||
- **AND** it SHALL verify that `selfEvaluation.verifier_evaluation.claim_checks` exists
|
||||
- **AND** it SHALL verify that `selfEvaluation.verifier_evaluation.composer_output` exists
|
||||
|
||||
#### Scenario: Gatekeeper failure cannot pass verification
|
||||
- **WHEN** a trace has `gatekeeper_result.status` equal to `fail`
|
||||
- **THEN** the evaluator SHALL fail the case if `selfEvaluation.verifier_evaluation.verdict` is `PASS`
|
||||
|
||||
#### Scenario: Claim checks are auditable
|
||||
- **WHEN** an evaluation case requires claim checks
|
||||
- **THEN** the evaluator SHALL verify that each claim check includes `claim_id`, `verification`, and `detail`
|
||||
- **AND** each verification SHALL be one of the V2 claim verification values recognized by the Verifier contract
|
||||
|
||||
#### Scenario: Composer output is auditable
|
||||
- **WHEN** an evaluation case requires Composer output
|
||||
- **THEN** the evaluator SHALL verify that `composer_output` records whether parsed output or fallback rendering was used
|
||||
|
||||
### Requirement: Evaluation harness SHALL prevent unsupported claims from leaking into final answers
|
||||
The evaluation harness SHALL detect configured unsafe or unsupported claim text when it appears in the final user-facing answer.
|
||||
|
||||
#### Scenario: Unsupported final-answer claim is rejected
|
||||
- **WHEN** an evaluation case declares forbidden confirmed-claim keywords
|
||||
- **THEN** the evaluator SHALL fail the case if the final answer contains any of those keywords
|
||||
|
||||
#### Scenario: Raw Executor JSON is not user-facing
|
||||
- **WHEN** a trace is evaluated under V2 audit closure
|
||||
- **THEN** the evaluator SHALL fail the case if the final answer contains raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`
|
||||
|
||||
#### Scenario: Composer fallback still avoids raw JSON leakage
|
||||
- **WHEN** a trace records Composer fallback rendering
|
||||
- **THEN** the evaluator SHALL still enforce final-answer raw JSON leakage checks
|
||||
@@ -0,0 +1,23 @@
|
||||
## 1. Eval Contract
|
||||
|
||||
- [x] 1.1 Extend `DiagnosisEvalCase` with optional V2 audit expectations.
|
||||
- [x] 1.2 Extend `DiagnosisEvalResult` and report output with observed audit signals.
|
||||
- [x] 1.3 Add deterministic evaluator checks for Gatekeeper, claim checks, Composer output, and raw JSON leakage.
|
||||
|
||||
## 2. Fixtures and Baseline
|
||||
|
||||
- [x] 2.1 Add fixed eval cases for Gatekeeper failure, unsupported claim filtering, and Composer fallback leakage prevention.
|
||||
- [x] 2.2 Add matching trace fixtures for each new case.
|
||||
- [x] 2.3 Regenerate baseline JSON and Markdown reports.
|
||||
|
||||
## 3. Documentation
|
||||
|
||||
- [x] 3.1 Update eval README with stage 5 scope and verification commands.
|
||||
- [x] 3.2 Rewrite eval schema documentation so V2 audit fields and case expectations are readable.
|
||||
|
||||
## 4. Verification and Archive
|
||||
|
||||
- [x] 4.1 Add or update unit tests for the new V2 audit checks.
|
||||
- [x] 4.2 Run focused evaluator and agent-chain tests.
|
||||
- [x] 4.3 Validate and archive the OpenSpec change.
|
||||
- [x] 4.4 Commit the completed phase 5 changes.
|
||||
@@ -133,3 +133,40 @@ The system SHALL expose baseline diff output in structured JSON and reviewable M
|
||||
#### Scenario: Markdown diff output
|
||||
- **WHEN** a baseline diff is written as Markdown
|
||||
- **THEN** it SHALL include a readable summary and a table of diff items
|
||||
|
||||
### Requirement: Evaluation harness SHALL validate Executor V2 audit closure
|
||||
The evaluation harness SHALL be able to validate the V2 audit chain from Executor structured output through Gatekeeper, Verifier claim checks, Composer output, and final user answer.
|
||||
|
||||
#### Scenario: Required V2 audit fields are present
|
||||
- **WHEN** an evaluation case requires V2 audit closure
|
||||
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.gatekeeper_result` exists
|
||||
- **AND** it SHALL verify that `selfEvaluation.verifier_evaluation.claim_checks` exists
|
||||
- **AND** it SHALL verify that `selfEvaluation.verifier_evaluation.composer_output` exists
|
||||
|
||||
#### Scenario: Gatekeeper failure cannot pass verification
|
||||
- **WHEN** a trace has `gatekeeper_result.status` equal to `fail`
|
||||
- **THEN** the evaluator SHALL fail the case if `selfEvaluation.verifier_evaluation.verdict` is `PASS`
|
||||
|
||||
#### Scenario: Claim checks are auditable
|
||||
- **WHEN** an evaluation case requires claim checks
|
||||
- **THEN** the evaluator SHALL verify that each claim check includes `claim_id`, `verification`, and `detail`
|
||||
- **AND** each verification SHALL be one of the V2 claim verification values recognized by the Verifier contract
|
||||
|
||||
#### Scenario: Composer output is auditable
|
||||
- **WHEN** an evaluation case requires Composer output
|
||||
- **THEN** the evaluator SHALL verify that `composer_output` records whether parsed output or fallback rendering was used
|
||||
|
||||
### Requirement: Evaluation harness SHALL prevent unsupported claims from leaking into final answers
|
||||
The evaluation harness SHALL detect configured unsafe or unsupported claim text when it appears in the final user-facing answer.
|
||||
|
||||
#### Scenario: Unsupported final-answer claim is rejected
|
||||
- **WHEN** an evaluation case declares forbidden confirmed-claim keywords
|
||||
- **THEN** the evaluator SHALL fail the case if the final answer contains any of those keywords
|
||||
|
||||
#### Scenario: Raw Executor JSON is not user-facing
|
||||
- **WHEN** a trace is evaluated under V2 audit closure
|
||||
- **THEN** the evaluator SHALL fail the case if the final answer contains raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`
|
||||
|
||||
#### Scenario: Composer fallback still avoids raw JSON leakage
|
||||
- **WHEN** a trace records Composer fallback rendering
|
||||
- **THEN** the evaluator SHALL still enforce final-answer raw JSON leakage checks
|
||||
|
||||
Reference in New Issue
Block a user