feat(demo): add interview quality audit

This commit is contained in:
zhuyongxin
2026-07-09 11:18:49 +08:00
parent a6c2d4459c
commit 9c9a0024d4
37 changed files with 1162 additions and 86 deletions
+8 -4
View File
@@ -29,10 +29,10 @@ The baseline evaluates saved trace fixtures. It does not start the application a
The committed baseline currently contains:
```text
10 fixed cases
10 passing fixture evaluations
4 PASS verdicts
5 LOW_CONFID verdicts
12 fixed cases
12 passing fixture evaluations
5 PASS verdicts
6 LOW_CONFID verdicts
1 REJECT verdict
```
@@ -44,6 +44,8 @@ The V2 evidence-pipeline matrix covers:
- Unsupported claim filtering before the final answer.
- Composer fallback rendering without raw Executor JSON leakage.
- Gatekeeper rule set version audit for new matrix fixtures.
- Prompt audit version checks for planner, executor, verifier, and composer prompts.
- Gatekeeper rule metadata checks for enabled rule id and default severity.
## Verification
@@ -80,6 +82,8 @@ Stage 5 adds these V2 checks:
- `claim_checks` must be structurally auditable.
- Composer output must record whether normal parsing or fallback rendering was used.
- Gatekeeper rule set version can be asserted per fixture.
- Prompt audit version and per-prompt versions can be asserted per fixture.
- Gatekeeper rule metadata can be required per fixture.
- Final answers must not leak raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`.
- Configured unsupported claim keywords must not appear as confirmed final-answer content.