docs(openspec): propose interview demo quality audit

This commit is contained in:
zhuyongxin
2026-07-09 10:34:33 +08:00
parent da45fa3fb0
commit a6c2d4459c
11 changed files with 416 additions and 0 deletions
@@ -0,0 +1,10 @@
# Acceptance
Acceptance will be filled during Apply/Archive with concrete command output.
Planned checks:
- `openspec validate interview-demo-quality-audit --strict` - passed during commit gate.
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ChatServiceSequentialAgentTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`
- `mvn -q -DskipTests compile`
- `powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1` against `mvp-demo` service if dependencies are available.
@@ -0,0 +1,23 @@
# Interview Demo Quality Audit Brief
## Background
The MVP already demonstrates traceable Agent diagnosis with Planner, Executor, Gatekeeper, Verifier, Composer, evidence tools, trace persistence, and deterministic eval fixtures. The remaining interview-readiness gap is not a new Agent architecture; it is making the demo easier to run and making prompt/rule changes easier to audit.
## Goal
Stabilize the interview demo path, expand fixture-backed evaluation, and persist prompt/Gatekeeper audit metadata so the project can explain and verify Agent behavior during interviews.
## Scope
- Add prompt audit metadata to Chat verifier evaluation.
- Extend deterministic eval cases and baseline reports.
- Add an interview demo preflight/check script.
- Update MVP demo and architecture documentation.
## Non-goals
- No public API or database schema changes.
- No new SubAgent split, MCP migration, process isolation, or AIOps LLM Verifier.
- No guarantee that every live LLM run returns PASS.
@@ -0,0 +1,91 @@
# interview-demo-quality-audit Decisions
## Clarify
- Entry summary: stabilize the interview demo, expand deterministic eval coverage, and add Prompt/Gatekeeper version audit.
- Slug: `interview-demo-quality-audit`.
- Devflow scale: `standard`.
- Interface impact: L2 internal contract change because `verifier_evaluation` gains `prompt_audit`; no public HTTP API or database schema change.
## Context
- `devflow/index.md` used: related entries found for `diagnosis-eval-demo-gatekeeper-closure`, `executor-composer-final-answer`, `verifier-evidence-reference-fidelity`, `mvp-demo-interview-runbook`, and `diagnosis-eval-baseline-diff`.
- Relevant glossary:
- Evidence Tools produce incident facts and must be recorded in `tool_invocation`.
- Verifier should not use skills/runbooks as incident evidence.
- Diagnosis Playbook Skill is workflow guidance, not a fact source.
- Historical constraints that must enter OpenSpec:
- Diagnosis eval is deterministic and fixture-backed; no LLM-as-judge.
- Stable demo scenarios are documentation/payloads plus deterministic fixtures; live E2E is a compatibility check, not a guaranteed PASS oracle.
- Gatekeeper rule metadata is already metadata-only and should not become dynamic rule execution.
- Composer is the final expression layer and must not leak raw Executor JSON.
## Question Pool
| ID | Dimension | Mode | Question | Status |
|---|---|---|---|---|
| Q1 | Terminology | evidence-driven | What should the new audit metadata be called? | Resolved |
| Q2 | Boundary | evidence-driven | Does this require public API or schema changes? | Resolved |
| Q3 | Acceptance | evidence-driven | Which current assets define deterministic acceptance? | Resolved |
| Q4 | Technical | evidence-driven | Where should prompt version metadata live with minimal implementation risk? | Resolved |
| Q5 | Scope | user-interview | Should live E2E be mandatory for all scenarios? | Confirmed by objective as conditional |
## Evidence-driven Conclusions
- Q1 conclusion: use `prompt_audit` for prompt version metadata and keep existing `gatekeeper_result.rule_set_version`.
- Q2 conclusion: keep this as an internal trace/self-evaluation contract change. Do not add endpoints, tables, or new Agent roles.
- Q3 conclusion: `DiagnosisTraceEvaluatorTest`, baseline reports, fixed fixtures, and demo scripts define current acceptance style.
- Q4 conclusion: add a small Chat prompt audit catalog near `ChatService` prompt loading and persist a compact snapshot with verifier evaluation.
- Q5 conclusion: run live E2E with `mvp-demo` profile if dependencies are available; otherwise record the blocker and rely on deterministic eval/unit evidence.
## Specify / Alignment
| Check | Status | Notes |
|---|---|---|
| proposal goals -> proposal | Aligned | Proposal covers demo preflight, eval expansion, prompt audit, and docs. |
| proposal scope/constraints -> design | Aligned | Design records no public API/schema changes, prompt audit shape, eval fields, and demo script behavior. |
| design decisions -> specs/tasks | Aligned | Specs cover persisted prompt audit, evaluator checks, baseline, and demo script outputs. |
| specs observable behavior -> tasks | Aligned | Each requirement has implementation and verification tasks. |
## Audit
Input -> processing -> output chain:
```text
prompt resource metadata
-> ChatService / PromptAudit snapshot
-> verifier_evaluation.prompt_audit
-> Trace API / eval fixtures
-> DiagnosisTraceEvaluator baseline
run-interview-demo-check.ps1
-> service readiness
-> chat / trace / feedback
-> mvp/demo/output summary
```
Architecture risk assessment:
1. The audit shape is intentionally compact and internal; storing full prompt text would create noisy traces and possible sensitive-content risk.
2. Eval should assert versions by explicit metadata, not by prompt content hashes that churn during local prompt edits.
3. Live demo checks may still be LOW_CONFID because LLM output is not deterministic; deterministic fixtures remain the regression source of truth.
4. No devflow/OpenSpec conflict found.
## Commit Gate
- `openspec validate interview-demo-quality-audit --strict`: passed.
- File completeness:
- `proposal.md`: present.
- `design.md`: present.
- `specs/`: present for `chat-verifier-agent`, `diagnosis-eval-harness`, and `mvp-demo-trace-acceptance`.
- `tasks.md`: present.
- Consistency:
- Proposal goals map to design sections.
- Design decisions map to spec requirements and executable tasks.
- Task acceptance checks are verifiable.
- `.committed` marker created.
## Current Checkpoint
- Commit completed.
- Apply is authorized by the original objective: "完成后归档提交".
@@ -0,0 +1,25 @@
# Evidence
## Context Files Read
- `devflow/index.md`
- `devflow/glossary/CONTEXT.md`
- `devflow/projects/2026-07-08-diagnosis-eval-demo-gatekeeper-closure/decisions.md`
- `devflow/projects/2026-07-08-executor-composer-final-answer/decisions.md`
- `mvp/architecture/current-mvp-architecture.md`
- `mvp/architecture/agent-orchestration.md`
- `mvp/architecture/executor-evidence-pipeline-refactor.md`
- `mvp/architecture/harness-quality-gates.md`
- `mvp/demo/README.md`
- `mvp/demo/ten-minute-interview-demo.md`
- `mvp/eval/README.md`
- `mvp/eval/cases/diagnosis-cases.json`
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- `src/main/java/com/superbiz/agent/service/ExecutorGatekeeperService.java`
- `src/main/java/com/superbiz/agent/eval/DiagnosisTraceEvaluator.java`
- `src/main/resources/gatekeeper/gatekeeper-rules.json`
## Tooling Note
The required `codebase-retrieval` and LSP tools were not exposed in this session. Impact analysis used `rg`, direct file reads, existing OpenSpec/devflow artifacts, and targeted tests instead.
@@ -0,0 +1 @@
committed
@@ -0,0 +1,117 @@
# Design: Interview Demo Quality Audit
## Overview
The change adds auditability and demo readiness without changing Agent routing.
```text
ChatService prompt resources
-> PromptAuditService
-> verifier_evaluation.prompt_audit
-> Trace API
-> DiagnosisTraceEvaluator
-> baseline report
mvp/demo/scripts/run-interview-demo-check.ps1
-> health/readiness check
-> payment timeout chat
-> trace fetch
-> feedback
-> demo output bundle
```
## Prompt Audit
Add a compact `prompt_audit` object under:
```text
diagnosis_session.self_evaluation.verifier_evaluation.prompt_audit
```
Shape:
```json
{
"version": "chat-prompts-v1",
"prompts": [
{
"name": "chat_planner",
"version": "chat-planner-v1",
"resource": "prompts/chat-planner-prompt.md"
}
]
}
```
Design choices:
- Use explicit local metadata, not full prompt hashes, because the interview goal is explainable version audit rather than cryptographic integrity.
- Keep this metadata in code or a small resource catalog near prompt loading.
- Persist prompt audit with every Chat verifier evaluation, including fallback/degraded paths.
- Do not include full prompt text in trace.
## Eval Expansion
Extend `DiagnosisEvalCase` with optional fields:
- `requirePromptAudit`
- `expectedPromptAuditVersion`
- `expectedPromptVersions`
- `requireGatekeeperRules`
Evaluator behavior:
- If `requirePromptAudit=true`, `verifier_evaluation.prompt_audit.version` must exist.
- If `expectedPromptAuditVersion` is set, it must match.
- If `expectedPromptVersions` is set, each listed prompt name/version pair must exist.
- If `requireGatekeeperRules=true`, `gatekeeper_result.rules` must be a non-empty list and each item must include `id`, `enabled`, and `default_severity`.
Add at least two fixture-backed cases:
- A positive audit closure case that requires prompt audit + Gatekeeper rules.
- A metadata-gap negative case represented as `LOW_CONFID`/safe final answer, used to prove the evaluator catches missing audit metadata when configured.
The baseline must remain fully passing after fixtures are updated.
## Demo Stabilization
Add a PowerShell script:
```text
mvp/demo/scripts/run-interview-demo-check.ps1
```
Responsibilities:
- Accept base URL and session id parameters.
- Check that the service is reachable.
- Run the existing payment-timeout chat request.
- Fetch trace for the same session id.
- Submit useful feedback.
- Write outputs under `mvp/demo/output/`.
- Emit a concise summary with session id, verdict, Gatekeeper rule version, prompt audit version, and output paths.
The script should fail fast with actionable messages when the service is unavailable.
## Documentation
Add/update:
- `mvp/demo/README.md`: mention the preflight script.
- `mvp/demo/ten-minute-interview-demo.md`: use the preflight script as the recommended path.
- `mvp/demo/interview-q-and-a.md`: concise interview answers for Agent engineering tradeoffs.
- `mvp/architecture/harness-quality-gates.md`: record prompt audit as part of the quality gate.
## Verification
Required:
- Targeted unit/eval tests for prompt audit persistence and evaluator checks.
- Regenerated baseline JSON/Markdown reports.
- OpenSpec validation.
E2E:
- If local dependencies are available, run Spring Boot with `mvp-demo` profile and execute the new preflight script.
- If unavailable, record the reason and rely on deterministic unit/eval evidence.
@@ -0,0 +1,58 @@
# Change: Interview Demo Quality Audit
## Problem
SuperBizAgent MVP is now strong enough to demonstrate traceable Agent engineering, but the interview path still has three gaps:
- The live demo has a run script, but no preflight command that checks service readiness and produces a concise interview evidence bundle.
- The diagnosis eval baseline covers the V2 evidence pipeline, but it does not yet assert prompt-version audit data and has limited coverage for audit metadata gaps.
- Gatekeeper already exposes `rule_set_version`, but prompt versions are not persisted with the verifier evaluation, making prompt changes harder to explain, compare, and roll back in an interview.
This change stabilizes the MVP as an interview artifact rather than adding a new diagnosis architecture.
## Proposed Solution
Implement a small internal quality/audit increment:
1. Add prompt version audit metadata to Chat verifier evaluation.
2. Extend deterministic diagnosis eval cases/fixtures to assert prompt audit metadata and audit-metadata failures.
3. Add an interview demo preflight script and documentation that can be run before or during a demo to verify service readiness, execute the payment timeout path, fetch trace, and record key audit fields.
4. Add/update MVP documentation for interview Q&A and the new audit/preflight workflow.
## Scope
In scope:
- Internal `diagnosis_session.self_evaluation.verifier_evaluation` audit JSON.
- Diagnosis eval case schema, evaluator checks, fixtures, and baseline reports.
- MVP demo scripts/docs.
- Architecture/demo documentation for prompt and Gatekeeper version audit.
Out of scope:
- Public HTTP API changes.
- Database schema changes.
- New Agent roles, MCP tool server migration, process isolation, or AIOps LLM Verifier.
- Replacing existing `Planner -> Executor -> Gatekeeper -> Verifier -> Composer` orchestration.
- Guaranteeing live LLM `PASS` for every demo run. Live demo compatibility and deterministic fixture regression are separate acceptance paths.
## Context Constraints From devflow
- Evidence Tools produce incident facts and must be recorded in `tool_invocation`.
- Chat quality gates are layered: Gatekeeper verifies evidence references, Verifier judges derivability, Composer controls expression.
- `diagnosis eval` is deterministic and fixture-backed; no LLM-as-judge.
- Demo assets should be runnable, but interview safety should not depend solely on live LLM behavior.
- Gatekeeper rule metadata is metadata-only; dynamic rule execution is out of scope.
## Interface Impact
Level: L2 internal contract change.
Reason: `verifier_evaluation` gains a compact `prompt_audit` object. Existing public API shape remains the same, and the value is exposed only through already-existing trace/self-evaluation JSON.
## Risks
- Baseline report churn is expected when adding cases; JSON and Markdown reports must be regenerated together.
- Prompt audit must be deterministic and stable enough for eval fixtures; avoid hashing full prompt text with environment-specific content.
- Demo preflight must not hardcode secrets and must tolerate local service unavailability with clear failure messages.
@@ -0,0 +1,16 @@
## MODIFIED Requirements
### Requirement: Verifier SHALL be observable
The Verifier's verdict SHALL be persisted for observability.
#### Scenario: prompt audit written to verifier evaluation
- **WHEN** the Chat verifier evaluation is persisted
- **THEN** the system SHALL include a `prompt_audit` object under `diagnosis_session.self_evaluation.verifier_evaluation`
- **AND** `prompt_audit.version` SHALL identify the Chat prompt audit catalog version
- **AND** `prompt_audit.prompts` SHALL include the planner, executor, verifier, and composer prompt names and versions
- **AND** full prompt text SHALL NOT be persisted in `prompt_audit`
#### Scenario: prompt audit available on fallback paths
- **WHEN** Chat verifier parsing fails, Composer parsing fails, or Chat produces a degraded answer
- **THEN** the persisted verifier evaluation SHALL still include `prompt_audit`
@@ -0,0 +1,21 @@
## MODIFIED Requirements
### Requirement: Diagnosis eval SHALL validate trace fixtures deterministically
The diagnosis eval harness SHALL evaluate saved trace fixtures without invoking an LLM judge.
#### Scenario: prompt audit assertions are enforced
- **WHEN** an eval case sets `requirePromptAudit=true`
- **THEN** the evaluator SHALL require `verifier_evaluation.prompt_audit.version`
- **AND** when `expectedPromptAuditVersion` is configured, it SHALL match exactly
- **AND** when `expectedPromptVersions` is configured, each configured prompt name SHALL appear with the expected version
#### Scenario: Gatekeeper rule metadata assertions are enforced
- **WHEN** an eval case sets `requireGatekeeperRules=true`
- **THEN** the evaluator SHALL require `verifier_evaluation.gatekeeper_result.rules` to be non-empty
- **AND** each rule item SHALL include `id`, `enabled`, and `default_severity`
#### Scenario: expanded baseline remains passing
- **WHEN** the committed fixture set is evaluated
- **THEN** every case SHALL pass
- **AND** baseline JSON and Markdown reports SHALL reflect the expanded case count and verdict distribution
@@ -0,0 +1,22 @@
## MODIFIED Requirements
### Requirement: MVP demo SHALL be reproducible for interviews
The MVP demo SHALL provide a repeatable way to show a diagnosis answer, trace, verifier evaluation, and feedback.
#### Scenario: interview demo check script records an evidence bundle
- **WHEN** the user runs the interview demo check script against a running `mvp-demo` service
- **THEN** the script SHALL submit a fixed Chat diagnosis request
- **AND** it SHALL fetch the trace for the same session id
- **AND** it SHALL submit useful feedback for that session
- **AND** it SHALL write chat, trace, feedback, and summary outputs under `mvp/demo/output/`
#### Scenario: interview demo check fails with actionable readiness output
- **WHEN** the target service is not reachable
- **THEN** the script SHALL fail before issuing diagnosis requests
- **AND** the failure message SHALL name the base URL and the expected startup profile
#### Scenario: interview documentation explains audit fields
- **WHEN** an interviewer asks how prompt or Gatekeeper changes are audited
- **THEN** the demo documentation SHALL point to `prompt_audit.version` and `gatekeeper_result.rule_set_version`
- **AND** it SHALL explain that deterministic eval fixtures are the regression source of truth
@@ -0,0 +1,32 @@
## 1. OpenSpec And devflow
- [x] 1.1 Create OpenSpec proposal/design/spec/tasks for `interview-demo-quality-audit`.
- [x] 1.2 Record context, question pool, interface impact, audit, and verification plan in devflow decisions.
- [x] 1.3 Pass OpenSpec validation and create `.committed`.
## 2. Prompt/Gatekeeper Version Audit
- [ ] 2.1 Add compact Chat prompt audit metadata for planner, executor, verifier, and composer prompts.
- [ ] 2.2 Persist `prompt_audit` under `verifier_evaluation` for Chat verifier/composer outcomes.
- [ ] 2.3 Add focused tests proving prompt audit appears in persisted verifier evaluation.
- [ ] 2.4 Extend eval checks for prompt audit and Gatekeeper rule metadata.
## 3. Eval Expansion
- [ ] 3.1 Extend diagnosis eval case/result schema for prompt audit fields.
- [ ] 3.2 Add fixture-backed cases for audit closure coverage.
- [ ] 3.3 Regenerate baseline JSON and Markdown reports.
- [ ] 3.4 Update eval docs/schema.
## 4. Interview Demo Stabilization
- [ ] 4.1 Add `run-interview-demo-check.ps1` with service preflight, chat, trace, feedback, and summary output.
- [ ] 4.2 Update demo README and 10-minute script to use the preflight path.
- [ ] 4.3 Add interview Q&A documentation focused on Agent engineering tradeoffs.
## 5. Verification And Archive
- [ ] 5.1 Run targeted tests for ChatService/prompt audit and diagnosis eval.
- [ ] 5.2 Run relevant broader regression tests.
- [ ] 5.3 Run E2E demo check with `mvp-demo` profile if dependencies are available; otherwise record the blocker.
- [ ] 5.4 Archive the OpenSpec change, update devflow artifacts, and commit implementation + archive.