Files
SuperBizAgent-java/openspec/changes/archive/2026-07-09-interview-demo-quality-audit/design.md
T

3.4 KiB

Design: Interview Demo Quality Audit

Overview

The change adds auditability and demo readiness without changing Agent routing.

ChatService prompt resources
  -> PromptAuditService
  -> verifier_evaluation.prompt_audit
  -> Trace API
  -> DiagnosisTraceEvaluator
  -> baseline report

mvp/demo/scripts/run-interview-demo-check.ps1
  -> health/readiness check
  -> payment timeout chat
  -> trace fetch
  -> feedback
  -> demo output bundle

Prompt Audit

Add a compact prompt_audit object under:

diagnosis_session.self_evaluation.verifier_evaluation.prompt_audit

Shape:

{
  "version": "chat-prompts-v1",
  "prompts": [
    {
      "name": "chat_planner",
      "version": "chat-planner-v1",
      "resource": "prompts/chat-planner-prompt.md"
    }
  ]
}

Design choices:

  • Use explicit local metadata, not full prompt hashes, because the interview goal is explainable version audit rather than cryptographic integrity.
  • Keep this metadata in code or a small resource catalog near prompt loading.
  • Persist prompt audit with every Chat verifier evaluation, including fallback/degraded paths.
  • Do not include full prompt text in trace.

Eval Expansion

Extend DiagnosisEvalCase with optional fields:

  • requirePromptAudit
  • expectedPromptAuditVersion
  • expectedPromptVersions
  • requireGatekeeperRules

Evaluator behavior:

  • If requirePromptAudit=true, verifier_evaluation.prompt_audit.version must exist.
  • If expectedPromptAuditVersion is set, it must match.
  • If expectedPromptVersions is set, each listed prompt name/version pair must exist.
  • If requireGatekeeperRules=true, gatekeeper_result.rules must be a non-empty list and each item must include id, enabled, and default_severity.

Add at least two fixture-backed cases:

  • A positive audit closure case that requires prompt audit + Gatekeeper rules.
  • A metadata-gap negative case represented as LOW_CONFID/safe final answer, used to prove the evaluator catches missing audit metadata when configured.

The baseline must remain fully passing after fixtures are updated.

Demo Stabilization

Add a PowerShell script:

mvp/demo/scripts/run-interview-demo-check.ps1

Responsibilities:

  • Accept base URL and session id parameters.
  • Check that the service is reachable.
  • Run the existing payment-timeout chat request.
  • Fetch trace for the same session id.
  • Submit useful feedback.
  • Write outputs under mvp/demo/output/.
  • Emit a concise summary with session id, verdict, Gatekeeper rule version, prompt audit version, and output paths.

The script should fail fast with actionable messages when the service is unavailable.

Documentation

Add/update:

  • mvp/demo/README.md: mention the preflight script.
  • mvp/demo/ten-minute-interview-demo.md: use the preflight script as the recommended path.
  • mvp/demo/interview-q-and-a.md: concise interview answers for Agent engineering tradeoffs.
  • mvp/architecture/harness-quality-gates.md: record prompt audit as part of the quality gate.

Verification

Required:

  • Targeted unit/eval tests for prompt audit persistence and evaluator checks.
  • Regenerated baseline JSON/Markdown reports.
  • OpenSpec validation.

E2E:

  • If local dependencies are available, run Spring Boot with mvp-demo profile and execute the new preflight script.
  • If unavailable, record the reason and rely on deterministic unit/eval evidence.