feat(eval): add evidence pipeline acceptance closure
This commit is contained in:
+8
-4
@@ -29,18 +29,21 @@ The baseline evaluates saved trace fixtures. It does not start the application a
|
||||
The committed baseline currently contains:
|
||||
|
||||
```text
|
||||
8 fixed cases
|
||||
8 passing fixture evaluations
|
||||
2 PASS verdicts
|
||||
10 fixed cases
|
||||
10 passing fixture evaluations
|
||||
4 PASS verdicts
|
||||
5 LOW_CONFID verdicts
|
||||
1 REJECT verdict
|
||||
```
|
||||
|
||||
The three V2 audit-closure cases cover:
|
||||
The V2 evidence-pipeline matrix covers:
|
||||
|
||||
- Positive supported evidence for a narrow HighCPU observation.
|
||||
- No-evidence `negative_observation` using `$.no_evidence`.
|
||||
- Gatekeeper failure for a fabricated tool invocation reference.
|
||||
- Unsupported claim filtering before the final answer.
|
||||
- Composer fallback rendering without raw Executor JSON leakage.
|
||||
- Gatekeeper rule set version audit for new matrix fixtures.
|
||||
|
||||
## Verification
|
||||
|
||||
@@ -76,6 +79,7 @@ Stage 5 adds these V2 checks:
|
||||
- Required V2 fixtures must include `gatekeeper_result`, `claim_checks`, and `composer_output`.
|
||||
- `claim_checks` must be structurally auditable.
|
||||
- Composer output must record whether normal parsing or fallback rendering was used.
|
||||
- Gatekeeper rule set version can be asserted per fixture.
|
||||
- Final answers must not leak raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`.
|
||||
- Configured unsupported claim keywords must not appear as confirmed final-answer content.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user