docs(openspec): propose interview demo quality audit

This commit is contained in:
zhuyongxin
2026-07-09 10:34:33 +08:00
parent da45fa3fb0
commit a6c2d4459c
11 changed files with 416 additions and 0 deletions
@@ -0,0 +1,117 @@
# Design: Interview Demo Quality Audit
## Overview
The change adds auditability and demo readiness without changing Agent routing.
```text
ChatService prompt resources
-> PromptAuditService
-> verifier_evaluation.prompt_audit
-> Trace API
-> DiagnosisTraceEvaluator
-> baseline report
mvp/demo/scripts/run-interview-demo-check.ps1
-> health/readiness check
-> payment timeout chat
-> trace fetch
-> feedback
-> demo output bundle
```
## Prompt Audit
Add a compact `prompt_audit` object under:
```text
diagnosis_session.self_evaluation.verifier_evaluation.prompt_audit
```
Shape:
```json
{
"version": "chat-prompts-v1",
"prompts": [
{
"name": "chat_planner",
"version": "chat-planner-v1",
"resource": "prompts/chat-planner-prompt.md"
}
]
}
```
Design choices:
- Use explicit local metadata, not full prompt hashes, because the interview goal is explainable version audit rather than cryptographic integrity.
- Keep this metadata in code or a small resource catalog near prompt loading.
- Persist prompt audit with every Chat verifier evaluation, including fallback/degraded paths.
- Do not include full prompt text in trace.
## Eval Expansion
Extend `DiagnosisEvalCase` with optional fields:
- `requirePromptAudit`
- `expectedPromptAuditVersion`
- `expectedPromptVersions`
- `requireGatekeeperRules`
Evaluator behavior:
- If `requirePromptAudit=true`, `verifier_evaluation.prompt_audit.version` must exist.
- If `expectedPromptAuditVersion` is set, it must match.
- If `expectedPromptVersions` is set, each listed prompt name/version pair must exist.
- If `requireGatekeeperRules=true`, `gatekeeper_result.rules` must be a non-empty list and each item must include `id`, `enabled`, and `default_severity`.
Add at least two fixture-backed cases:
- A positive audit closure case that requires prompt audit + Gatekeeper rules.
- A metadata-gap negative case represented as `LOW_CONFID`/safe final answer, used to prove the evaluator catches missing audit metadata when configured.
The baseline must remain fully passing after fixtures are updated.
## Demo Stabilization
Add a PowerShell script:
```text
mvp/demo/scripts/run-interview-demo-check.ps1
```
Responsibilities:
- Accept base URL and session id parameters.
- Check that the service is reachable.
- Run the existing payment-timeout chat request.
- Fetch trace for the same session id.
- Submit useful feedback.
- Write outputs under `mvp/demo/output/`.
- Emit a concise summary with session id, verdict, Gatekeeper rule version, prompt audit version, and output paths.
The script should fail fast with actionable messages when the service is unavailable.
## Documentation
Add/update:
- `mvp/demo/README.md`: mention the preflight script.
- `mvp/demo/ten-minute-interview-demo.md`: use the preflight script as the recommended path.
- `mvp/demo/interview-q-and-a.md`: concise interview answers for Agent engineering tradeoffs.
- `mvp/architecture/harness-quality-gates.md`: record prompt audit as part of the quality gate.
## Verification
Required:
- Targeted unit/eval tests for prompt audit persistence and evaluator checks.
- Regenerated baseline JSON/Markdown reports.
- OpenSpec validation.
E2E:
- If local dependencies are available, run Spring Boot with `mvp-demo` profile and execute the new preflight script.
- If unavailable, record the reason and rely on deterministic unit/eval evidence.