3.1 KiB
3.1 KiB
Change: Interview Demo Quality Audit
Problem
SuperBizAgent MVP is now strong enough to demonstrate traceable Agent engineering, but the interview path still has three gaps:
- The live demo has a run script, but no preflight command that checks service readiness and produces a concise interview evidence bundle.
- The diagnosis eval baseline covers the V2 evidence pipeline, but it does not yet assert prompt-version audit data and has limited coverage for audit metadata gaps.
- Gatekeeper already exposes
rule_set_version, but prompt versions are not persisted with the verifier evaluation, making prompt changes harder to explain, compare, and roll back in an interview.
This change stabilizes the MVP as an interview artifact rather than adding a new diagnosis architecture.
Proposed Solution
Implement a small internal quality/audit increment:
- Add prompt version audit metadata to Chat verifier evaluation.
- Extend deterministic diagnosis eval cases/fixtures to assert prompt audit metadata and audit-metadata failures.
- Add an interview demo preflight script and documentation that can be run before or during a demo to verify service readiness, execute the payment timeout path, fetch trace, and record key audit fields.
- Add/update MVP documentation for interview Q&A and the new audit/preflight workflow.
Scope
In scope:
- Internal
diagnosis_session.self_evaluation.verifier_evaluationaudit JSON. - Diagnosis eval case schema, evaluator checks, fixtures, and baseline reports.
- MVP demo scripts/docs.
- Architecture/demo documentation for prompt and Gatekeeper version audit.
Out of scope:
- Public HTTP API changes.
- Database schema changes.
- New Agent roles, MCP tool server migration, process isolation, or AIOps LLM Verifier.
- Replacing existing
Planner -> Executor -> Gatekeeper -> Verifier -> Composerorchestration. - Guaranteeing live LLM
PASSfor every demo run. Live demo compatibility and deterministic fixture regression are separate acceptance paths.
Context Constraints From devflow
- Evidence Tools produce incident facts and must be recorded in
tool_invocation. - Chat quality gates are layered: Gatekeeper verifies evidence references, Verifier judges derivability, Composer controls expression.
diagnosis evalis deterministic and fixture-backed; no LLM-as-judge.- Demo assets should be runnable, but interview safety should not depend solely on live LLM behavior.
- Gatekeeper rule metadata is metadata-only; dynamic rule execution is out of scope.
Interface Impact
Level: L2 internal contract change.
Reason: verifier_evaluation gains a compact prompt_audit object. Existing public API shape remains the same, and the value is exposed only through already-existing trace/self-evaluation JSON.
Risks
- Baseline report churn is expected when adding cases; JSON and Markdown reports must be regenerated together.
- Prompt audit must be deterministic and stable enough for eval fixtures; avoid hashing full prompt text with environment-specific content.
- Demo preflight must not hardcode secrets and must tolerate local service unavailability with clear failure messages.