feat(eval): add executor audit closure checks
This commit is contained in:
@@ -133,3 +133,40 @@ The system SHALL expose baseline diff output in structured JSON and reviewable M
|
||||
#### Scenario: Markdown diff output
|
||||
- **WHEN** a baseline diff is written as Markdown
|
||||
- **THEN** it SHALL include a readable summary and a table of diff items
|
||||
|
||||
### Requirement: Evaluation harness SHALL validate Executor V2 audit closure
|
||||
The evaluation harness SHALL be able to validate the V2 audit chain from Executor structured output through Gatekeeper, Verifier claim checks, Composer output, and final user answer.
|
||||
|
||||
#### Scenario: Required V2 audit fields are present
|
||||
- **WHEN** an evaluation case requires V2 audit closure
|
||||
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.gatekeeper_result` exists
|
||||
- **AND** it SHALL verify that `selfEvaluation.verifier_evaluation.claim_checks` exists
|
||||
- **AND** it SHALL verify that `selfEvaluation.verifier_evaluation.composer_output` exists
|
||||
|
||||
#### Scenario: Gatekeeper failure cannot pass verification
|
||||
- **WHEN** a trace has `gatekeeper_result.status` equal to `fail`
|
||||
- **THEN** the evaluator SHALL fail the case if `selfEvaluation.verifier_evaluation.verdict` is `PASS`
|
||||
|
||||
#### Scenario: Claim checks are auditable
|
||||
- **WHEN** an evaluation case requires claim checks
|
||||
- **THEN** the evaluator SHALL verify that each claim check includes `claim_id`, `verification`, and `detail`
|
||||
- **AND** each verification SHALL be one of the V2 claim verification values recognized by the Verifier contract
|
||||
|
||||
#### Scenario: Composer output is auditable
|
||||
- **WHEN** an evaluation case requires Composer output
|
||||
- **THEN** the evaluator SHALL verify that `composer_output` records whether parsed output or fallback rendering was used
|
||||
|
||||
### Requirement: Evaluation harness SHALL prevent unsupported claims from leaking into final answers
|
||||
The evaluation harness SHALL detect configured unsafe or unsupported claim text when it appears in the final user-facing answer.
|
||||
|
||||
#### Scenario: Unsupported final-answer claim is rejected
|
||||
- **WHEN** an evaluation case declares forbidden confirmed-claim keywords
|
||||
- **THEN** the evaluator SHALL fail the case if the final answer contains any of those keywords
|
||||
|
||||
#### Scenario: Raw Executor JSON is not user-facing
|
||||
- **WHEN** a trace is evaluated under V2 audit closure
|
||||
- **THEN** the evaluator SHALL fail the case if the final answer contains raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`
|
||||
|
||||
#### Scenario: Composer fallback still avoids raw JSON leakage
|
||||
- **WHEN** a trace records Composer fallback rendering
|
||||
- **THEN** the evaluator SHALL still enforce final-answer raw JSON leakage checks
|
||||
|
||||
Reference in New Issue
Block a user