Files

6.3 KiB

interview-demo-quality-audit Decisions

Clarify

  • Entry summary: stabilize the interview demo, expand deterministic eval coverage, and add Prompt/Gatekeeper version audit.
  • Slug: interview-demo-quality-audit.
  • Devflow scale: standard.
  • Interface impact: L2 internal contract change because verifier_evaluation gains prompt_audit; no public HTTP API or database schema change.

Context

  • devflow/index.md used: related entries found for diagnosis-eval-demo-gatekeeper-closure, executor-composer-final-answer, verifier-evidence-reference-fidelity, mvp-demo-interview-runbook, and diagnosis-eval-baseline-diff.
  • Relevant glossary:
    • Evidence Tools produce incident facts and must be recorded in tool_invocation.
    • Verifier should not use skills/runbooks as incident evidence.
    • Diagnosis Playbook Skill is workflow guidance, not a fact source.
  • Historical constraints that must enter OpenSpec:
    • Diagnosis eval is deterministic and fixture-backed; no LLM-as-judge.
    • Stable demo scenarios are documentation/payloads plus deterministic fixtures; live E2E is a compatibility check, not a guaranteed PASS oracle.
    • Gatekeeper rule metadata is already metadata-only and should not become dynamic rule execution.
    • Composer is the final expression layer and must not leak raw Executor JSON.

Question Pool

ID Dimension Mode Question Status
Q1 Terminology evidence-driven What should the new audit metadata be called? Resolved
Q2 Boundary evidence-driven Does this require public API or schema changes? Resolved
Q3 Acceptance evidence-driven Which current assets define deterministic acceptance? Resolved
Q4 Technical evidence-driven Where should prompt version metadata live with minimal implementation risk? Resolved
Q5 Scope user-interview Should live E2E be mandatory for all scenarios? Confirmed by objective as conditional

Evidence-driven Conclusions

  • Q1 conclusion: use prompt_audit for prompt version metadata and keep existing gatekeeper_result.rule_set_version.
  • Q2 conclusion: keep this as an internal trace/self-evaluation contract change. Do not add endpoints, tables, or new Agent roles.
  • Q3 conclusion: DiagnosisTraceEvaluatorTest, baseline reports, fixed fixtures, and demo scripts define current acceptance style.
  • Q4 conclusion: add a small Chat prompt audit catalog near ChatService prompt loading and persist a compact snapshot with verifier evaluation.
  • Q5 conclusion: run live E2E with mvp-demo profile if dependencies are available; otherwise record the blocker and rely on deterministic eval/unit evidence.

Specify / Alignment

Check Status Notes
proposal goals -> proposal Aligned Proposal covers demo preflight, eval expansion, prompt audit, and docs.
proposal scope/constraints -> design Aligned Design records no public API/schema changes, prompt audit shape, eval fields, and demo script behavior.
design decisions -> specs/tasks Aligned Specs cover persisted prompt audit, evaluator checks, baseline, and demo script outputs.
specs observable behavior -> tasks Aligned Each requirement has implementation and verification tasks.

Audit

Input -> processing -> output chain:

prompt resource metadata
  -> ChatService / PromptAudit snapshot
  -> verifier_evaluation.prompt_audit
  -> Trace API / eval fixtures
  -> DiagnosisTraceEvaluator baseline

run-interview-demo-check.ps1
  -> service readiness
  -> chat / trace / feedback
  -> mvp/demo/output summary

Architecture risk assessment:

  1. The audit shape is intentionally compact and internal; storing full prompt text would create noisy traces and possible sensitive-content risk.
  2. Eval should assert versions by explicit metadata, not by prompt content hashes that churn during local prompt edits.
  3. Live demo checks may still be LOW_CONFID because LLM output is not deterministic; deterministic fixtures remain the regression source of truth.
  4. No devflow/OpenSpec conflict found.

Commit Gate

  • openspec validate interview-demo-quality-audit --strict: passed.
  • File completeness:
    • proposal.md: present.
    • design.md: present.
    • specs/: present for chat-verifier-agent, diagnosis-eval-harness, and mvp-demo-trace-acceptance.
    • tasks.md: present.
  • Consistency:
    • Proposal goals map to design sections.
    • Design decisions map to spec requirements and executable tasks.
    • Task acceptance checks are verifiable.
  • .committed marker created.

Current Checkpoint

  • Commit completed.
  • Apply is authorized by the original objective: "完成后归档提交".

Pre-apply Research

  • Capability source: sm-flow built-in apply protocol. openspec-apply-change was not invoked directly in this session.
  • Repository semantic search/LSP note: the requested codebase-retrieval and LSP tools were not available in the exposed toolset, so impact analysis used rg, direct file reads, OpenSpec/devflow artifacts, and targeted tests.
  • Reference implementation and reuse:
    • ChatService.persistVerifierEvaluation(...) is the single persistence point for Chat verifier/composer audit data; prompt audit was added there to cover normal, fallback, and degraded Composer paths.
    • DiagnosisTraceEvaluator and DiagnosisEvalReportWriter are the deterministic eval extension points; no LLM judge was introduced.
    • mvp/demo/scripts/run-payment-timeout-demo.ps1 provided the request/trace/feedback flow reused by the new interview preflight script.
  • Interface impact remains L2 internal trace contract: verifier_evaluation.prompt_audit and eval report fields are added; no public endpoint, table, or request DTO changed.

Apply Notes

  • Added compact Chat prompt audit metadata: chat-prompts-v1, with planner/executor/verifier/composer prompt versions and resource paths.
  • Extended diagnosis eval schema, result reporting, baseline fixtures, JSON report, and Markdown report for Prompt audit and Gatekeeper rule metadata.
  • Added two fixture-backed audit cases:
    • prompt-gatekeeper-audit-closure
    • audit-metadata-low-confid
  • Added mvp/demo/scripts/run-interview-demo-check.ps1 to run service readiness, Chat, Trace, feedback, and summary output.
  • Updated MVP demo/eval/architecture docs to explain prompt_audit.version, gatekeeper_result.rule_set_version, and deterministic fixture baseline.