59 lines
3.1 KiB
Markdown
59 lines
3.1 KiB
Markdown
# Change: Interview Demo Quality Audit
|
|
|
|
## Problem
|
|
|
|
SuperBizAgent MVP is now strong enough to demonstrate traceable Agent engineering, but the interview path still has three gaps:
|
|
|
|
- The live demo has a run script, but no preflight command that checks service readiness and produces a concise interview evidence bundle.
|
|
- The diagnosis eval baseline covers the V2 evidence pipeline, but it does not yet assert prompt-version audit data and has limited coverage for audit metadata gaps.
|
|
- Gatekeeper already exposes `rule_set_version`, but prompt versions are not persisted with the verifier evaluation, making prompt changes harder to explain, compare, and roll back in an interview.
|
|
|
|
This change stabilizes the MVP as an interview artifact rather than adding a new diagnosis architecture.
|
|
|
|
## Proposed Solution
|
|
|
|
Implement a small internal quality/audit increment:
|
|
|
|
1. Add prompt version audit metadata to Chat verifier evaluation.
|
|
2. Extend deterministic diagnosis eval cases/fixtures to assert prompt audit metadata and audit-metadata failures.
|
|
3. Add an interview demo preflight script and documentation that can be run before or during a demo to verify service readiness, execute the payment timeout path, fetch trace, and record key audit fields.
|
|
4. Add/update MVP documentation for interview Q&A and the new audit/preflight workflow.
|
|
|
|
## Scope
|
|
|
|
In scope:
|
|
|
|
- Internal `diagnosis_session.self_evaluation.verifier_evaluation` audit JSON.
|
|
- Diagnosis eval case schema, evaluator checks, fixtures, and baseline reports.
|
|
- MVP demo scripts/docs.
|
|
- Architecture/demo documentation for prompt and Gatekeeper version audit.
|
|
|
|
Out of scope:
|
|
|
|
- Public HTTP API changes.
|
|
- Database schema changes.
|
|
- New Agent roles, MCP tool server migration, process isolation, or AIOps LLM Verifier.
|
|
- Replacing existing `Planner -> Executor -> Gatekeeper -> Verifier -> Composer` orchestration.
|
|
- Guaranteeing live LLM `PASS` for every demo run. Live demo compatibility and deterministic fixture regression are separate acceptance paths.
|
|
|
|
## Context Constraints From devflow
|
|
|
|
- Evidence Tools produce incident facts and must be recorded in `tool_invocation`.
|
|
- Chat quality gates are layered: Gatekeeper verifies evidence references, Verifier judges derivability, Composer controls expression.
|
|
- `diagnosis eval` is deterministic and fixture-backed; no LLM-as-judge.
|
|
- Demo assets should be runnable, but interview safety should not depend solely on live LLM behavior.
|
|
- Gatekeeper rule metadata is metadata-only; dynamic rule execution is out of scope.
|
|
|
|
## Interface Impact
|
|
|
|
Level: L2 internal contract change.
|
|
|
|
Reason: `verifier_evaluation` gains a compact `prompt_audit` object. Existing public API shape remains the same, and the value is exposed only through already-existing trace/self-evaluation JSON.
|
|
|
|
## Risks
|
|
|
|
- Baseline report churn is expected when adding cases; JSON and Markdown reports must be regenerated together.
|
|
- Prompt audit must be deterministic and stable enough for eval fixtures; avoid hashing full prompt text with environment-specific content.
|
|
- Demo preflight must not hardcode secrets and must tolerate local service unavailability with clear failure messages.
|
|
|