feat(demo): add interview quality audit
This commit is contained in:
+8
-4
@@ -29,10 +29,10 @@ The baseline evaluates saved trace fixtures. It does not start the application a
|
||||
The committed baseline currently contains:
|
||||
|
||||
```text
|
||||
10 fixed cases
|
||||
10 passing fixture evaluations
|
||||
4 PASS verdicts
|
||||
5 LOW_CONFID verdicts
|
||||
12 fixed cases
|
||||
12 passing fixture evaluations
|
||||
5 PASS verdicts
|
||||
6 LOW_CONFID verdicts
|
||||
1 REJECT verdict
|
||||
```
|
||||
|
||||
@@ -44,6 +44,8 @@ The V2 evidence-pipeline matrix covers:
|
||||
- Unsupported claim filtering before the final answer.
|
||||
- Composer fallback rendering without raw Executor JSON leakage.
|
||||
- Gatekeeper rule set version audit for new matrix fixtures.
|
||||
- Prompt audit version checks for planner, executor, verifier, and composer prompts.
|
||||
- Gatekeeper rule metadata checks for enabled rule id and default severity.
|
||||
|
||||
## Verification
|
||||
|
||||
@@ -80,6 +82,8 @@ Stage 5 adds these V2 checks:
|
||||
- `claim_checks` must be structurally auditable.
|
||||
- Composer output must record whether normal parsing or fallback rendering was used.
|
||||
- Gatekeeper rule set version can be asserted per fixture.
|
||||
- Prompt audit version and per-prompt versions can be asserted per fixture.
|
||||
- Gatekeeper rule metadata can be required per fixture.
|
||||
- Final answers must not leak raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`.
|
||||
- Configured unsupported claim keywords must not appear as confirmed final-answer content.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user