# interview-demo-quality-audit Decisions ## Clarify - Entry summary: stabilize the interview demo, expand deterministic eval coverage, and add Prompt/Gatekeeper version audit. - Slug: `interview-demo-quality-audit`. - Devflow scale: `standard`. - Interface impact: L2 internal contract change because `verifier_evaluation` gains `prompt_audit`; no public HTTP API or database schema change. ## Context - `devflow/index.md` used: related entries found for `diagnosis-eval-demo-gatekeeper-closure`, `executor-composer-final-answer`, `verifier-evidence-reference-fidelity`, `mvp-demo-interview-runbook`, and `diagnosis-eval-baseline-diff`. - Relevant glossary: - Evidence Tools produce incident facts and must be recorded in `tool_invocation`. - Verifier should not use skills/runbooks as incident evidence. - Diagnosis Playbook Skill is workflow guidance, not a fact source. - Historical constraints that must enter OpenSpec: - Diagnosis eval is deterministic and fixture-backed; no LLM-as-judge. - Stable demo scenarios are documentation/payloads plus deterministic fixtures; live E2E is a compatibility check, not a guaranteed PASS oracle. - Gatekeeper rule metadata is already metadata-only and should not become dynamic rule execution. - Composer is the final expression layer and must not leak raw Executor JSON. ## Question Pool | ID | Dimension | Mode | Question | Status | |---|---|---|---|---| | Q1 | Terminology | evidence-driven | What should the new audit metadata be called? | Resolved | | Q2 | Boundary | evidence-driven | Does this require public API or schema changes? | Resolved | | Q3 | Acceptance | evidence-driven | Which current assets define deterministic acceptance? | Resolved | | Q4 | Technical | evidence-driven | Where should prompt version metadata live with minimal implementation risk? | Resolved | | Q5 | Scope | user-interview | Should live E2E be mandatory for all scenarios? | Confirmed by objective as conditional | ## Evidence-driven Conclusions - Q1 conclusion: use `prompt_audit` for prompt version metadata and keep existing `gatekeeper_result.rule_set_version`. - Q2 conclusion: keep this as an internal trace/self-evaluation contract change. Do not add endpoints, tables, or new Agent roles. - Q3 conclusion: `DiagnosisTraceEvaluatorTest`, baseline reports, fixed fixtures, and demo scripts define current acceptance style. - Q4 conclusion: add a small Chat prompt audit catalog near `ChatService` prompt loading and persist a compact snapshot with verifier evaluation. - Q5 conclusion: run live E2E with `mvp-demo` profile if dependencies are available; otherwise record the blocker and rely on deterministic eval/unit evidence. ## Specify / Alignment | Check | Status | Notes | |---|---|---| | proposal goals -> proposal | Aligned | Proposal covers demo preflight, eval expansion, prompt audit, and docs. | | proposal scope/constraints -> design | Aligned | Design records no public API/schema changes, prompt audit shape, eval fields, and demo script behavior. | | design decisions -> specs/tasks | Aligned | Specs cover persisted prompt audit, evaluator checks, baseline, and demo script outputs. | | specs observable behavior -> tasks | Aligned | Each requirement has implementation and verification tasks. | ## Audit Input -> processing -> output chain: ```text prompt resource metadata -> ChatService / PromptAudit snapshot -> verifier_evaluation.prompt_audit -> Trace API / eval fixtures -> DiagnosisTraceEvaluator baseline run-interview-demo-check.ps1 -> service readiness -> chat / trace / feedback -> mvp/demo/output summary ``` Architecture risk assessment: 1. The audit shape is intentionally compact and internal; storing full prompt text would create noisy traces and possible sensitive-content risk. 2. Eval should assert versions by explicit metadata, not by prompt content hashes that churn during local prompt edits. 3. Live demo checks may still be LOW_CONFID because LLM output is not deterministic; deterministic fixtures remain the regression source of truth. 4. No devflow/OpenSpec conflict found. ## Commit Gate - `openspec validate interview-demo-quality-audit --strict`: passed. - File completeness: - `proposal.md`: present. - `design.md`: present. - `specs/`: present for `chat-verifier-agent`, `diagnosis-eval-harness`, and `mvp-demo-trace-acceptance`. - `tasks.md`: present. - Consistency: - Proposal goals map to design sections. - Design decisions map to spec requirements and executable tasks. - Task acceptance checks are verifiable. - `.committed` marker created. ## Current Checkpoint - Commit completed. - Apply is authorized by the original objective: "完成后归档提交". ## Pre-apply Research - Capability source: sm-flow built-in apply protocol. `openspec-apply-change` was not invoked directly in this session. - Repository semantic search/LSP note: the requested `codebase-retrieval` and LSP tools were not available in the exposed toolset, so impact analysis used `rg`, direct file reads, OpenSpec/devflow artifacts, and targeted tests. - Reference implementation and reuse: - `ChatService.persistVerifierEvaluation(...)` is the single persistence point for Chat verifier/composer audit data; prompt audit was added there to cover normal, fallback, and degraded Composer paths. - `DiagnosisTraceEvaluator` and `DiagnosisEvalReportWriter` are the deterministic eval extension points; no LLM judge was introduced. - `mvp/demo/scripts/run-payment-timeout-demo.ps1` provided the request/trace/feedback flow reused by the new interview preflight script. - Interface impact remains L2 internal trace contract: `verifier_evaluation.prompt_audit` and eval report fields are added; no public endpoint, table, or request DTO changed. ## Apply Notes - Added compact Chat prompt audit metadata: `chat-prompts-v1`, with planner/executor/verifier/composer prompt versions and resource paths. - Extended diagnosis eval schema, result reporting, baseline fixtures, JSON report, and Markdown report for Prompt audit and Gatekeeper rule metadata. - Added two fixture-backed audit cases: - `prompt-gatekeeper-audit-closure` - `audit-metadata-low-confid` - Added `mvp/demo/scripts/run-interview-demo-check.ps1` to run service readiness, Chat, Trace, feedback, and summary output. - Updated MVP demo/eval/architecture docs to explain `prompt_audit.version`, `gatekeeper_result.rule_set_version`, and deterministic fixture baseline.