Files
SuperBizAgent-java/devflow/projects/2026-07-09-interview-demo-quality-audit/decisions.md
T

4.6 KiB

interview-demo-quality-audit Decisions

Clarify

  • Entry summary: stabilize the interview demo, expand deterministic eval coverage, and add Prompt/Gatekeeper version audit.
  • Slug: interview-demo-quality-audit.
  • Devflow scale: standard.
  • Interface impact: L2 internal contract change because verifier_evaluation gains prompt_audit; no public HTTP API or database schema change.

Context

  • devflow/index.md used: related entries found for diagnosis-eval-demo-gatekeeper-closure, executor-composer-final-answer, verifier-evidence-reference-fidelity, mvp-demo-interview-runbook, and diagnosis-eval-baseline-diff.
  • Relevant glossary:
    • Evidence Tools produce incident facts and must be recorded in tool_invocation.
    • Verifier should not use skills/runbooks as incident evidence.
    • Diagnosis Playbook Skill is workflow guidance, not a fact source.
  • Historical constraints that must enter OpenSpec:
    • Diagnosis eval is deterministic and fixture-backed; no LLM-as-judge.
    • Stable demo scenarios are documentation/payloads plus deterministic fixtures; live E2E is a compatibility check, not a guaranteed PASS oracle.
    • Gatekeeper rule metadata is already metadata-only and should not become dynamic rule execution.
    • Composer is the final expression layer and must not leak raw Executor JSON.

Question Pool

ID Dimension Mode Question Status
Q1 Terminology evidence-driven What should the new audit metadata be called? Resolved
Q2 Boundary evidence-driven Does this require public API or schema changes? Resolved
Q3 Acceptance evidence-driven Which current assets define deterministic acceptance? Resolved
Q4 Technical evidence-driven Where should prompt version metadata live with minimal implementation risk? Resolved
Q5 Scope user-interview Should live E2E be mandatory for all scenarios? Confirmed by objective as conditional

Evidence-driven Conclusions

  • Q1 conclusion: use prompt_audit for prompt version metadata and keep existing gatekeeper_result.rule_set_version.
  • Q2 conclusion: keep this as an internal trace/self-evaluation contract change. Do not add endpoints, tables, or new Agent roles.
  • Q3 conclusion: DiagnosisTraceEvaluatorTest, baseline reports, fixed fixtures, and demo scripts define current acceptance style.
  • Q4 conclusion: add a small Chat prompt audit catalog near ChatService prompt loading and persist a compact snapshot with verifier evaluation.
  • Q5 conclusion: run live E2E with mvp-demo profile if dependencies are available; otherwise record the blocker and rely on deterministic eval/unit evidence.

Specify / Alignment

Check Status Notes
proposal goals -> proposal Aligned Proposal covers demo preflight, eval expansion, prompt audit, and docs.
proposal scope/constraints -> design Aligned Design records no public API/schema changes, prompt audit shape, eval fields, and demo script behavior.
design decisions -> specs/tasks Aligned Specs cover persisted prompt audit, evaluator checks, baseline, and demo script outputs.
specs observable behavior -> tasks Aligned Each requirement has implementation and verification tasks.

Audit

Input -> processing -> output chain:

prompt resource metadata
  -> ChatService / PromptAudit snapshot
  -> verifier_evaluation.prompt_audit
  -> Trace API / eval fixtures
  -> DiagnosisTraceEvaluator baseline

run-interview-demo-check.ps1
  -> service readiness
  -> chat / trace / feedback
  -> mvp/demo/output summary

Architecture risk assessment:

  1. The audit shape is intentionally compact and internal; storing full prompt text would create noisy traces and possible sensitive-content risk.
  2. Eval should assert versions by explicit metadata, not by prompt content hashes that churn during local prompt edits.
  3. Live demo checks may still be LOW_CONFID because LLM output is not deterministic; deterministic fixtures remain the regression source of truth.
  4. No devflow/OpenSpec conflict found.

Commit Gate

  • openspec validate interview-demo-quality-audit --strict: passed.
  • File completeness:
    • proposal.md: present.
    • design.md: present.
    • specs/: present for chat-verifier-agent, diagnosis-eval-harness, and mvp-demo-trace-acceptance.
    • tasks.md: present.
  • Consistency:
    • Proposal goals map to design sections.
    • Design decisions map to spec requirements and executable tasks.
    • Task acceptance checks are verifiable.
  • .committed marker created.

Current Checkpoint

  • Commit completed.
  • Apply is authorized by the original objective: "完成后归档提交".