112 lines
6.3 KiB
Markdown
112 lines
6.3 KiB
Markdown
# interview-demo-quality-audit Decisions
|
|
|
|
## Clarify
|
|
|
|
- Entry summary: stabilize the interview demo, expand deterministic eval coverage, and add Prompt/Gatekeeper version audit.
|
|
- Slug: `interview-demo-quality-audit`.
|
|
- Devflow scale: `standard`.
|
|
- Interface impact: L2 internal contract change because `verifier_evaluation` gains `prompt_audit`; no public HTTP API or database schema change.
|
|
|
|
## Context
|
|
|
|
- `devflow/index.md` used: related entries found for `diagnosis-eval-demo-gatekeeper-closure`, `executor-composer-final-answer`, `verifier-evidence-reference-fidelity`, `mvp-demo-interview-runbook`, and `diagnosis-eval-baseline-diff`.
|
|
- Relevant glossary:
|
|
- Evidence Tools produce incident facts and must be recorded in `tool_invocation`.
|
|
- Verifier should not use skills/runbooks as incident evidence.
|
|
- Diagnosis Playbook Skill is workflow guidance, not a fact source.
|
|
- Historical constraints that must enter OpenSpec:
|
|
- Diagnosis eval is deterministic and fixture-backed; no LLM-as-judge.
|
|
- Stable demo scenarios are documentation/payloads plus deterministic fixtures; live E2E is a compatibility check, not a guaranteed PASS oracle.
|
|
- Gatekeeper rule metadata is already metadata-only and should not become dynamic rule execution.
|
|
- Composer is the final expression layer and must not leak raw Executor JSON.
|
|
|
|
## Question Pool
|
|
|
|
| ID | Dimension | Mode | Question | Status |
|
|
|---|---|---|---|---|
|
|
| Q1 | Terminology | evidence-driven | What should the new audit metadata be called? | Resolved |
|
|
| Q2 | Boundary | evidence-driven | Does this require public API or schema changes? | Resolved |
|
|
| Q3 | Acceptance | evidence-driven | Which current assets define deterministic acceptance? | Resolved |
|
|
| Q4 | Technical | evidence-driven | Where should prompt version metadata live with minimal implementation risk? | Resolved |
|
|
| Q5 | Scope | user-interview | Should live E2E be mandatory for all scenarios? | Confirmed by objective as conditional |
|
|
|
|
## Evidence-driven Conclusions
|
|
|
|
- Q1 conclusion: use `prompt_audit` for prompt version metadata and keep existing `gatekeeper_result.rule_set_version`.
|
|
- Q2 conclusion: keep this as an internal trace/self-evaluation contract change. Do not add endpoints, tables, or new Agent roles.
|
|
- Q3 conclusion: `DiagnosisTraceEvaluatorTest`, baseline reports, fixed fixtures, and demo scripts define current acceptance style.
|
|
- Q4 conclusion: add a small Chat prompt audit catalog near `ChatService` prompt loading and persist a compact snapshot with verifier evaluation.
|
|
- Q5 conclusion: run live E2E with `mvp-demo` profile if dependencies are available; otherwise record the blocker and rely on deterministic eval/unit evidence.
|
|
|
|
## Specify / Alignment
|
|
|
|
| Check | Status | Notes |
|
|
|---|---|---|
|
|
| proposal goals -> proposal | Aligned | Proposal covers demo preflight, eval expansion, prompt audit, and docs. |
|
|
| proposal scope/constraints -> design | Aligned | Design records no public API/schema changes, prompt audit shape, eval fields, and demo script behavior. |
|
|
| design decisions -> specs/tasks | Aligned | Specs cover persisted prompt audit, evaluator checks, baseline, and demo script outputs. |
|
|
| specs observable behavior -> tasks | Aligned | Each requirement has implementation and verification tasks. |
|
|
|
|
## Audit
|
|
|
|
Input -> processing -> output chain:
|
|
|
|
```text
|
|
prompt resource metadata
|
|
-> ChatService / PromptAudit snapshot
|
|
-> verifier_evaluation.prompt_audit
|
|
-> Trace API / eval fixtures
|
|
-> DiagnosisTraceEvaluator baseline
|
|
|
|
run-interview-demo-check.ps1
|
|
-> service readiness
|
|
-> chat / trace / feedback
|
|
-> mvp/demo/output summary
|
|
```
|
|
|
|
Architecture risk assessment:
|
|
|
|
1. The audit shape is intentionally compact and internal; storing full prompt text would create noisy traces and possible sensitive-content risk.
|
|
2. Eval should assert versions by explicit metadata, not by prompt content hashes that churn during local prompt edits.
|
|
3. Live demo checks may still be LOW_CONFID because LLM output is not deterministic; deterministic fixtures remain the regression source of truth.
|
|
4. No devflow/OpenSpec conflict found.
|
|
|
|
## Commit Gate
|
|
|
|
- `openspec validate interview-demo-quality-audit --strict`: passed.
|
|
- File completeness:
|
|
- `proposal.md`: present.
|
|
- `design.md`: present.
|
|
- `specs/`: present for `chat-verifier-agent`, `diagnosis-eval-harness`, and `mvp-demo-trace-acceptance`.
|
|
- `tasks.md`: present.
|
|
- Consistency:
|
|
- Proposal goals map to design sections.
|
|
- Design decisions map to spec requirements and executable tasks.
|
|
- Task acceptance checks are verifiable.
|
|
- `.committed` marker created.
|
|
|
|
## Current Checkpoint
|
|
|
|
- Commit completed.
|
|
- Apply is authorized by the original objective: "完成后归档提交".
|
|
|
|
## Pre-apply Research
|
|
|
|
- Capability source: sm-flow built-in apply protocol. `openspec-apply-change` was not invoked directly in this session.
|
|
- Repository semantic search/LSP note: the requested `codebase-retrieval` and LSP tools were not available in the exposed toolset, so impact analysis used `rg`, direct file reads, OpenSpec/devflow artifacts, and targeted tests.
|
|
- Reference implementation and reuse:
|
|
- `ChatService.persistVerifierEvaluation(...)` is the single persistence point for Chat verifier/composer audit data; prompt audit was added there to cover normal, fallback, and degraded Composer paths.
|
|
- `DiagnosisTraceEvaluator` and `DiagnosisEvalReportWriter` are the deterministic eval extension points; no LLM judge was introduced.
|
|
- `mvp/demo/scripts/run-payment-timeout-demo.ps1` provided the request/trace/feedback flow reused by the new interview preflight script.
|
|
- Interface impact remains L2 internal trace contract: `verifier_evaluation.prompt_audit` and eval report fields are added; no public endpoint, table, or request DTO changed.
|
|
|
|
## Apply Notes
|
|
|
|
- Added compact Chat prompt audit metadata: `chat-prompts-v1`, with planner/executor/verifier/composer prompt versions and resource paths.
|
|
- Extended diagnosis eval schema, result reporting, baseline fixtures, JSON report, and Markdown report for Prompt audit and Gatekeeper rule metadata.
|
|
- Added two fixture-backed audit cases:
|
|
- `prompt-gatekeeper-audit-closure`
|
|
- `audit-metadata-low-confid`
|
|
- Added `mvp/demo/scripts/run-interview-demo-check.ps1` to run service readiness, Chat, Trace, feedback, and summary output.
|
|
- Updated MVP demo/eval/architecture docs to explain `prompt_audit.version`, `gatekeeper_result.rule_set_version`, and deterministic fixture baseline.
|