# Change: Interview Demo Quality Audit ## Problem SuperBizAgent MVP is now strong enough to demonstrate traceable Agent engineering, but the interview path still has three gaps: - The live demo has a run script, but no preflight command that checks service readiness and produces a concise interview evidence bundle. - The diagnosis eval baseline covers the V2 evidence pipeline, but it does not yet assert prompt-version audit data and has limited coverage for audit metadata gaps. - Gatekeeper already exposes `rule_set_version`, but prompt versions are not persisted with the verifier evaluation, making prompt changes harder to explain, compare, and roll back in an interview. This change stabilizes the MVP as an interview artifact rather than adding a new diagnosis architecture. ## Proposed Solution Implement a small internal quality/audit increment: 1. Add prompt version audit metadata to Chat verifier evaluation. 2. Extend deterministic diagnosis eval cases/fixtures to assert prompt audit metadata and audit-metadata failures. 3. Add an interview demo preflight script and documentation that can be run before or during a demo to verify service readiness, execute the payment timeout path, fetch trace, and record key audit fields. 4. Add/update MVP documentation for interview Q&A and the new audit/preflight workflow. ## Scope In scope: - Internal `diagnosis_session.self_evaluation.verifier_evaluation` audit JSON. - Diagnosis eval case schema, evaluator checks, fixtures, and baseline reports. - MVP demo scripts/docs. - Architecture/demo documentation for prompt and Gatekeeper version audit. Out of scope: - Public HTTP API changes. - Database schema changes. - New Agent roles, MCP tool server migration, process isolation, or AIOps LLM Verifier. - Replacing existing `Planner -> Executor -> Gatekeeper -> Verifier -> Composer` orchestration. - Guaranteeing live LLM `PASS` for every demo run. Live demo compatibility and deterministic fixture regression are separate acceptance paths. ## Context Constraints From devflow - Evidence Tools produce incident facts and must be recorded in `tool_invocation`. - Chat quality gates are layered: Gatekeeper verifies evidence references, Verifier judges derivability, Composer controls expression. - `diagnosis eval` is deterministic and fixture-backed; no LLM-as-judge. - Demo assets should be runnable, but interview safety should not depend solely on live LLM behavior. - Gatekeeper rule metadata is metadata-only; dynamic rule execution is out of scope. ## Interface Impact Level: L2 internal contract change. Reason: `verifier_evaluation` gains a compact `prompt_audit` object. Existing public API shape remains the same, and the value is exposed only through already-existing trace/self-evaluation JSON. ## Risks - Baseline report churn is expected when adding cases; JSON and Markdown reports must be regenerated together. - Prompt audit must be deterministic and stable enough for eval fixtures; avoid hashing full prompt text with environment-specific content. - Demo preflight must not hardcode secrets and must tolerate local service unavailability with clear failure messages.