docs(openspec): propose interview demo quality audit
This commit is contained in:
@@ -0,0 +1,58 @@
|
||||
# Change: Interview Demo Quality Audit
|
||||
|
||||
## Problem
|
||||
|
||||
SuperBizAgent MVP is now strong enough to demonstrate traceable Agent engineering, but the interview path still has three gaps:
|
||||
|
||||
- The live demo has a run script, but no preflight command that checks service readiness and produces a concise interview evidence bundle.
|
||||
- The diagnosis eval baseline covers the V2 evidence pipeline, but it does not yet assert prompt-version audit data and has limited coverage for audit metadata gaps.
|
||||
- Gatekeeper already exposes `rule_set_version`, but prompt versions are not persisted with the verifier evaluation, making prompt changes harder to explain, compare, and roll back in an interview.
|
||||
|
||||
This change stabilizes the MVP as an interview artifact rather than adding a new diagnosis architecture.
|
||||
|
||||
## Proposed Solution
|
||||
|
||||
Implement a small internal quality/audit increment:
|
||||
|
||||
1. Add prompt version audit metadata to Chat verifier evaluation.
|
||||
2. Extend deterministic diagnosis eval cases/fixtures to assert prompt audit metadata and audit-metadata failures.
|
||||
3. Add an interview demo preflight script and documentation that can be run before or during a demo to verify service readiness, execute the payment timeout path, fetch trace, and record key audit fields.
|
||||
4. Add/update MVP documentation for interview Q&A and the new audit/preflight workflow.
|
||||
|
||||
## Scope
|
||||
|
||||
In scope:
|
||||
|
||||
- Internal `diagnosis_session.self_evaluation.verifier_evaluation` audit JSON.
|
||||
- Diagnosis eval case schema, evaluator checks, fixtures, and baseline reports.
|
||||
- MVP demo scripts/docs.
|
||||
- Architecture/demo documentation for prompt and Gatekeeper version audit.
|
||||
|
||||
Out of scope:
|
||||
|
||||
- Public HTTP API changes.
|
||||
- Database schema changes.
|
||||
- New Agent roles, MCP tool server migration, process isolation, or AIOps LLM Verifier.
|
||||
- Replacing existing `Planner -> Executor -> Gatekeeper -> Verifier -> Composer` orchestration.
|
||||
- Guaranteeing live LLM `PASS` for every demo run. Live demo compatibility and deterministic fixture regression are separate acceptance paths.
|
||||
|
||||
## Context Constraints From devflow
|
||||
|
||||
- Evidence Tools produce incident facts and must be recorded in `tool_invocation`.
|
||||
- Chat quality gates are layered: Gatekeeper verifies evidence references, Verifier judges derivability, Composer controls expression.
|
||||
- `diagnosis eval` is deterministic and fixture-backed; no LLM-as-judge.
|
||||
- Demo assets should be runnable, but interview safety should not depend solely on live LLM behavior.
|
||||
- Gatekeeper rule metadata is metadata-only; dynamic rule execution is out of scope.
|
||||
|
||||
## Interface Impact
|
||||
|
||||
Level: L2 internal contract change.
|
||||
|
||||
Reason: `verifier_evaluation` gains a compact `prompt_audit` object. Existing public API shape remains the same, and the value is exposed only through already-existing trace/self-evaluation JSON.
|
||||
|
||||
## Risks
|
||||
|
||||
- Baseline report churn is expected when adding cases; JSON and Markdown reports must be regenerated together.
|
||||
- Prompt audit must be deterministic and stable enough for eval fixtures; avoid hashing full prompt text with environment-specific content.
|
||||
- Demo preflight must not hardcode secrets and must tolerate local service unavailability with clear failure messages.
|
||||
|
||||
Reference in New Issue
Block a user