6.3 KiB
6.3 KiB
interview-demo-quality-audit Decisions
Clarify
- Entry summary: stabilize the interview demo, expand deterministic eval coverage, and add Prompt/Gatekeeper version audit.
- Slug:
interview-demo-quality-audit. - Devflow scale:
standard. - Interface impact: L2 internal contract change because
verifier_evaluationgainsprompt_audit; no public HTTP API or database schema change.
Context
devflow/index.mdused: related entries found fordiagnosis-eval-demo-gatekeeper-closure,executor-composer-final-answer,verifier-evidence-reference-fidelity,mvp-demo-interview-runbook, anddiagnosis-eval-baseline-diff.- Relevant glossary:
- Evidence Tools produce incident facts and must be recorded in
tool_invocation. - Verifier should not use skills/runbooks as incident evidence.
- Diagnosis Playbook Skill is workflow guidance, not a fact source.
- Evidence Tools produce incident facts and must be recorded in
- Historical constraints that must enter OpenSpec:
- Diagnosis eval is deterministic and fixture-backed; no LLM-as-judge.
- Stable demo scenarios are documentation/payloads plus deterministic fixtures; live E2E is a compatibility check, not a guaranteed PASS oracle.
- Gatekeeper rule metadata is already metadata-only and should not become dynamic rule execution.
- Composer is the final expression layer and must not leak raw Executor JSON.
Question Pool
| ID | Dimension | Mode | Question | Status |
|---|---|---|---|---|
| Q1 | Terminology | evidence-driven | What should the new audit metadata be called? | Resolved |
| Q2 | Boundary | evidence-driven | Does this require public API or schema changes? | Resolved |
| Q3 | Acceptance | evidence-driven | Which current assets define deterministic acceptance? | Resolved |
| Q4 | Technical | evidence-driven | Where should prompt version metadata live with minimal implementation risk? | Resolved |
| Q5 | Scope | user-interview | Should live E2E be mandatory for all scenarios? | Confirmed by objective as conditional |
Evidence-driven Conclusions
- Q1 conclusion: use
prompt_auditfor prompt version metadata and keep existinggatekeeper_result.rule_set_version. - Q2 conclusion: keep this as an internal trace/self-evaluation contract change. Do not add endpoints, tables, or new Agent roles.
- Q3 conclusion:
DiagnosisTraceEvaluatorTest, baseline reports, fixed fixtures, and demo scripts define current acceptance style. - Q4 conclusion: add a small Chat prompt audit catalog near
ChatServiceprompt loading and persist a compact snapshot with verifier evaluation. - Q5 conclusion: run live E2E with
mvp-demoprofile if dependencies are available; otherwise record the blocker and rely on deterministic eval/unit evidence.
Specify / Alignment
| Check | Status | Notes |
|---|---|---|
| proposal goals -> proposal | Aligned | Proposal covers demo preflight, eval expansion, prompt audit, and docs. |
| proposal scope/constraints -> design | Aligned | Design records no public API/schema changes, prompt audit shape, eval fields, and demo script behavior. |
| design decisions -> specs/tasks | Aligned | Specs cover persisted prompt audit, evaluator checks, baseline, and demo script outputs. |
| specs observable behavior -> tasks | Aligned | Each requirement has implementation and verification tasks. |
Audit
Input -> processing -> output chain:
prompt resource metadata
-> ChatService / PromptAudit snapshot
-> verifier_evaluation.prompt_audit
-> Trace API / eval fixtures
-> DiagnosisTraceEvaluator baseline
run-interview-demo-check.ps1
-> service readiness
-> chat / trace / feedback
-> mvp/demo/output summary
Architecture risk assessment:
- The audit shape is intentionally compact and internal; storing full prompt text would create noisy traces and possible sensitive-content risk.
- Eval should assert versions by explicit metadata, not by prompt content hashes that churn during local prompt edits.
- Live demo checks may still be LOW_CONFID because LLM output is not deterministic; deterministic fixtures remain the regression source of truth.
- No devflow/OpenSpec conflict found.
Commit Gate
openspec validate interview-demo-quality-audit --strict: passed.- File completeness:
proposal.md: present.design.md: present.specs/: present forchat-verifier-agent,diagnosis-eval-harness, andmvp-demo-trace-acceptance.tasks.md: present.
- Consistency:
- Proposal goals map to design sections.
- Design decisions map to spec requirements and executable tasks.
- Task acceptance checks are verifiable.
.committedmarker created.
Current Checkpoint
- Commit completed.
- Apply is authorized by the original objective: "完成后归档提交".
Pre-apply Research
- Capability source: sm-flow built-in apply protocol.
openspec-apply-changewas not invoked directly in this session. - Repository semantic search/LSP note: the requested
codebase-retrievaland LSP tools were not available in the exposed toolset, so impact analysis usedrg, direct file reads, OpenSpec/devflow artifacts, and targeted tests. - Reference implementation and reuse:
ChatService.persistVerifierEvaluation(...)is the single persistence point for Chat verifier/composer audit data; prompt audit was added there to cover normal, fallback, and degraded Composer paths.DiagnosisTraceEvaluatorandDiagnosisEvalReportWriterare the deterministic eval extension points; no LLM judge was introduced.mvp/demo/scripts/run-payment-timeout-demo.ps1provided the request/trace/feedback flow reused by the new interview preflight script.
- Interface impact remains L2 internal trace contract:
verifier_evaluation.prompt_auditand eval report fields are added; no public endpoint, table, or request DTO changed.
Apply Notes
- Added compact Chat prompt audit metadata:
chat-prompts-v1, with planner/executor/verifier/composer prompt versions and resource paths. - Extended diagnosis eval schema, result reporting, baseline fixtures, JSON report, and Markdown report for Prompt audit and Gatekeeper rule metadata.
- Added two fixture-backed audit cases:
prompt-gatekeeper-audit-closureaudit-metadata-low-confid
- Added
mvp/demo/scripts/run-interview-demo-check.ps1to run service readiness, Chat, Trace, feedback, and summary output. - Updated MVP demo/eval/architecture docs to explain
prompt_audit.version,gatekeeper_result.rule_set_version, and deterministic fixture baseline.