118 lines
3.4 KiB
Markdown
118 lines
3.4 KiB
Markdown
# Design: Interview Demo Quality Audit
|
|
|
|
## Overview
|
|
|
|
The change adds auditability and demo readiness without changing Agent routing.
|
|
|
|
```text
|
|
ChatService prompt resources
|
|
-> PromptAuditService
|
|
-> verifier_evaluation.prompt_audit
|
|
-> Trace API
|
|
-> DiagnosisTraceEvaluator
|
|
-> baseline report
|
|
|
|
mvp/demo/scripts/run-interview-demo-check.ps1
|
|
-> health/readiness check
|
|
-> payment timeout chat
|
|
-> trace fetch
|
|
-> feedback
|
|
-> demo output bundle
|
|
```
|
|
|
|
## Prompt Audit
|
|
|
|
Add a compact `prompt_audit` object under:
|
|
|
|
```text
|
|
diagnosis_session.self_evaluation.verifier_evaluation.prompt_audit
|
|
```
|
|
|
|
Shape:
|
|
|
|
```json
|
|
{
|
|
"version": "chat-prompts-v1",
|
|
"prompts": [
|
|
{
|
|
"name": "chat_planner",
|
|
"version": "chat-planner-v1",
|
|
"resource": "prompts/chat-planner-prompt.md"
|
|
}
|
|
]
|
|
}
|
|
```
|
|
|
|
Design choices:
|
|
|
|
- Use explicit local metadata, not full prompt hashes, because the interview goal is explainable version audit rather than cryptographic integrity.
|
|
- Keep this metadata in code or a small resource catalog near prompt loading.
|
|
- Persist prompt audit with every Chat verifier evaluation, including fallback/degraded paths.
|
|
- Do not include full prompt text in trace.
|
|
|
|
## Eval Expansion
|
|
|
|
Extend `DiagnosisEvalCase` with optional fields:
|
|
|
|
- `requirePromptAudit`
|
|
- `expectedPromptAuditVersion`
|
|
- `expectedPromptVersions`
|
|
- `requireGatekeeperRules`
|
|
|
|
Evaluator behavior:
|
|
|
|
- If `requirePromptAudit=true`, `verifier_evaluation.prompt_audit.version` must exist.
|
|
- If `expectedPromptAuditVersion` is set, it must match.
|
|
- If `expectedPromptVersions` is set, each listed prompt name/version pair must exist.
|
|
- If `requireGatekeeperRules=true`, `gatekeeper_result.rules` must be a non-empty list and each item must include `id`, `enabled`, and `default_severity`.
|
|
|
|
Add at least two fixture-backed cases:
|
|
|
|
- A positive audit closure case that requires prompt audit + Gatekeeper rules.
|
|
- A metadata-gap negative case represented as `LOW_CONFID`/safe final answer, used to prove the evaluator catches missing audit metadata when configured.
|
|
|
|
The baseline must remain fully passing after fixtures are updated.
|
|
|
|
## Demo Stabilization
|
|
|
|
Add a PowerShell script:
|
|
|
|
```text
|
|
mvp/demo/scripts/run-interview-demo-check.ps1
|
|
```
|
|
|
|
Responsibilities:
|
|
|
|
- Accept base URL and session id parameters.
|
|
- Check that the service is reachable.
|
|
- Run the existing payment-timeout chat request.
|
|
- Fetch trace for the same session id.
|
|
- Submit useful feedback.
|
|
- Write outputs under `mvp/demo/output/`.
|
|
- Emit a concise summary with session id, verdict, Gatekeeper rule version, prompt audit version, and output paths.
|
|
|
|
The script should fail fast with actionable messages when the service is unavailable.
|
|
|
|
## Documentation
|
|
|
|
Add/update:
|
|
|
|
- `mvp/demo/README.md`: mention the preflight script.
|
|
- `mvp/demo/ten-minute-interview-demo.md`: use the preflight script as the recommended path.
|
|
- `mvp/demo/interview-q-and-a.md`: concise interview answers for Agent engineering tradeoffs.
|
|
- `mvp/architecture/harness-quality-gates.md`: record prompt audit as part of the quality gate.
|
|
|
|
## Verification
|
|
|
|
Required:
|
|
|
|
- Targeted unit/eval tests for prompt audit persistence and evaluator checks.
|
|
- Regenerated baseline JSON/Markdown reports.
|
|
- OpenSpec validation.
|
|
|
|
E2E:
|
|
|
|
- If local dependencies are available, run Spring Boot with `mvp-demo` profile and execute the new preflight script.
|
|
- If unavailable, record the reason and rely on deterministic unit/eval evidence.
|
|
|