3.4 KiB
3.4 KiB
Design: Interview Demo Quality Audit
Overview
The change adds auditability and demo readiness without changing Agent routing.
ChatService prompt resources
-> PromptAuditService
-> verifier_evaluation.prompt_audit
-> Trace API
-> DiagnosisTraceEvaluator
-> baseline report
mvp/demo/scripts/run-interview-demo-check.ps1
-> health/readiness check
-> payment timeout chat
-> trace fetch
-> feedback
-> demo output bundle
Prompt Audit
Add a compact prompt_audit object under:
diagnosis_session.self_evaluation.verifier_evaluation.prompt_audit
Shape:
{
"version": "chat-prompts-v1",
"prompts": [
{
"name": "chat_planner",
"version": "chat-planner-v1",
"resource": "prompts/chat-planner-prompt.md"
}
]
}
Design choices:
- Use explicit local metadata, not full prompt hashes, because the interview goal is explainable version audit rather than cryptographic integrity.
- Keep this metadata in code or a small resource catalog near prompt loading.
- Persist prompt audit with every Chat verifier evaluation, including fallback/degraded paths.
- Do not include full prompt text in trace.
Eval Expansion
Extend DiagnosisEvalCase with optional fields:
requirePromptAuditexpectedPromptAuditVersionexpectedPromptVersionsrequireGatekeeperRules
Evaluator behavior:
- If
requirePromptAudit=true,verifier_evaluation.prompt_audit.versionmust exist. - If
expectedPromptAuditVersionis set, it must match. - If
expectedPromptVersionsis set, each listed prompt name/version pair must exist. - If
requireGatekeeperRules=true,gatekeeper_result.rulesmust be a non-empty list and each item must includeid,enabled, anddefault_severity.
Add at least two fixture-backed cases:
- A positive audit closure case that requires prompt audit + Gatekeeper rules.
- A metadata-gap negative case represented as
LOW_CONFID/safe final answer, used to prove the evaluator catches missing audit metadata when configured.
The baseline must remain fully passing after fixtures are updated.
Demo Stabilization
Add a PowerShell script:
mvp/demo/scripts/run-interview-demo-check.ps1
Responsibilities:
- Accept base URL and session id parameters.
- Check that the service is reachable.
- Run the existing payment-timeout chat request.
- Fetch trace for the same session id.
- Submit useful feedback.
- Write outputs under
mvp/demo/output/. - Emit a concise summary with session id, verdict, Gatekeeper rule version, prompt audit version, and output paths.
The script should fail fast with actionable messages when the service is unavailable.
Documentation
Add/update:
mvp/demo/README.md: mention the preflight script.mvp/demo/ten-minute-interview-demo.md: use the preflight script as the recommended path.mvp/demo/interview-q-and-a.md: concise interview answers for Agent engineering tradeoffs.mvp/architecture/harness-quality-gates.md: record prompt audit as part of the quality gate.
Verification
Required:
- Targeted unit/eval tests for prompt audit persistence and evaluator checks.
- Regenerated baseline JSON/Markdown reports.
- OpenSpec validation.
E2E:
- If local dependencies are available, run Spring Boot with
mvp-demoprofile and execute the new preflight script. - If unavailable, record the reason and rely on deterministic unit/eval evidence.