feat(eval): add executor audit closure checks

This commit is contained in:
aruo
2026-07-08 10:21:39 +08:00
parent 6015bcbf6f
commit a08672b31e
20 changed files with 1142 additions and 188 deletions
@@ -133,3 +133,40 @@ The system SHALL expose baseline diff output in structured JSON and reviewable M
#### Scenario: Markdown diff output
- **WHEN** a baseline diff is written as Markdown
- **THEN** it SHALL include a readable summary and a table of diff items
### Requirement: Evaluation harness SHALL validate Executor V2 audit closure
The evaluation harness SHALL be able to validate the V2 audit chain from Executor structured output through Gatekeeper, Verifier claim checks, Composer output, and final user answer.
#### Scenario: Required V2 audit fields are present
- **WHEN** an evaluation case requires V2 audit closure
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.gatekeeper_result` exists
- **AND** it SHALL verify that `selfEvaluation.verifier_evaluation.claim_checks` exists
- **AND** it SHALL verify that `selfEvaluation.verifier_evaluation.composer_output` exists
#### Scenario: Gatekeeper failure cannot pass verification
- **WHEN** a trace has `gatekeeper_result.status` equal to `fail`
- **THEN** the evaluator SHALL fail the case if `selfEvaluation.verifier_evaluation.verdict` is `PASS`
#### Scenario: Claim checks are auditable
- **WHEN** an evaluation case requires claim checks
- **THEN** the evaluator SHALL verify that each claim check includes `claim_id`, `verification`, and `detail`
- **AND** each verification SHALL be one of the V2 claim verification values recognized by the Verifier contract
#### Scenario: Composer output is auditable
- **WHEN** an evaluation case requires Composer output
- **THEN** the evaluator SHALL verify that `composer_output` records whether parsed output or fallback rendering was used
### Requirement: Evaluation harness SHALL prevent unsupported claims from leaking into final answers
The evaluation harness SHALL detect configured unsafe or unsupported claim text when it appears in the final user-facing answer.
#### Scenario: Unsupported final-answer claim is rejected
- **WHEN** an evaluation case declares forbidden confirmed-claim keywords
- **THEN** the evaluator SHALL fail the case if the final answer contains any of those keywords
#### Scenario: Raw Executor JSON is not user-facing
- **WHEN** a trace is evaluated under V2 audit closure
- **THEN** the evaluator SHALL fail the case if the final answer contains raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`
#### Scenario: Composer fallback still avoids raw JSON leakage
- **WHEN** a trace records Composer fallback rendering
- **THEN** the evaluator SHALL still enforce final-answer raw JSON leakage checks