diff --git a/devflow/projects/2026-07-09-interview-demo-quality-audit/acceptance.md b/devflow/projects/2026-07-09-interview-demo-quality-audit/acceptance.md new file mode 100644 index 0000000..1f05c0f --- /dev/null +++ b/devflow/projects/2026-07-09-interview-demo-quality-audit/acceptance.md @@ -0,0 +1,10 @@ +# Acceptance + +Acceptance will be filled during Apply/Archive with concrete command output. + +Planned checks: + +- `openspec validate interview-demo-quality-audit --strict` - passed during commit gate. +- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ChatServiceSequentialAgentTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest" test` +- `mvn -q -DskipTests compile` +- `powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1` against `mvp-demo` service if dependencies are available. diff --git a/devflow/projects/2026-07-09-interview-demo-quality-audit/brief.md b/devflow/projects/2026-07-09-interview-demo-quality-audit/brief.md new file mode 100644 index 0000000..c6300f1 --- /dev/null +++ b/devflow/projects/2026-07-09-interview-demo-quality-audit/brief.md @@ -0,0 +1,23 @@ +# Interview Demo Quality Audit Brief + +## Background + +The MVP already demonstrates traceable Agent diagnosis with Planner, Executor, Gatekeeper, Verifier, Composer, evidence tools, trace persistence, and deterministic eval fixtures. The remaining interview-readiness gap is not a new Agent architecture; it is making the demo easier to run and making prompt/rule changes easier to audit. + +## Goal + +Stabilize the interview demo path, expand fixture-backed evaluation, and persist prompt/Gatekeeper audit metadata so the project can explain and verify Agent behavior during interviews. + +## Scope + +- Add prompt audit metadata to Chat verifier evaluation. +- Extend deterministic eval cases and baseline reports. +- Add an interview demo preflight/check script. +- Update MVP demo and architecture documentation. + +## Non-goals + +- No public API or database schema changes. +- No new SubAgent split, MCP migration, process isolation, or AIOps LLM Verifier. +- No guarantee that every live LLM run returns PASS. + diff --git a/devflow/projects/2026-07-09-interview-demo-quality-audit/decisions.md b/devflow/projects/2026-07-09-interview-demo-quality-audit/decisions.md new file mode 100644 index 0000000..bb2bcc7 --- /dev/null +++ b/devflow/projects/2026-07-09-interview-demo-quality-audit/decisions.md @@ -0,0 +1,91 @@ +# interview-demo-quality-audit Decisions + +## Clarify + +- Entry summary: stabilize the interview demo, expand deterministic eval coverage, and add Prompt/Gatekeeper version audit. +- Slug: `interview-demo-quality-audit`. +- Devflow scale: `standard`. +- Interface impact: L2 internal contract change because `verifier_evaluation` gains `prompt_audit`; no public HTTP API or database schema change. + +## Context + +- `devflow/index.md` used: related entries found for `diagnosis-eval-demo-gatekeeper-closure`, `executor-composer-final-answer`, `verifier-evidence-reference-fidelity`, `mvp-demo-interview-runbook`, and `diagnosis-eval-baseline-diff`. +- Relevant glossary: + - Evidence Tools produce incident facts and must be recorded in `tool_invocation`. + - Verifier should not use skills/runbooks as incident evidence. + - Diagnosis Playbook Skill is workflow guidance, not a fact source. +- Historical constraints that must enter OpenSpec: + - Diagnosis eval is deterministic and fixture-backed; no LLM-as-judge. + - Stable demo scenarios are documentation/payloads plus deterministic fixtures; live E2E is a compatibility check, not a guaranteed PASS oracle. + - Gatekeeper rule metadata is already metadata-only and should not become dynamic rule execution. + - Composer is the final expression layer and must not leak raw Executor JSON. + +## Question Pool + +| ID | Dimension | Mode | Question | Status | +|---|---|---|---|---| +| Q1 | Terminology | evidence-driven | What should the new audit metadata be called? | Resolved | +| Q2 | Boundary | evidence-driven | Does this require public API or schema changes? | Resolved | +| Q3 | Acceptance | evidence-driven | Which current assets define deterministic acceptance? | Resolved | +| Q4 | Technical | evidence-driven | Where should prompt version metadata live with minimal implementation risk? | Resolved | +| Q5 | Scope | user-interview | Should live E2E be mandatory for all scenarios? | Confirmed by objective as conditional | + +## Evidence-driven Conclusions + +- Q1 conclusion: use `prompt_audit` for prompt version metadata and keep existing `gatekeeper_result.rule_set_version`. +- Q2 conclusion: keep this as an internal trace/self-evaluation contract change. Do not add endpoints, tables, or new Agent roles. +- Q3 conclusion: `DiagnosisTraceEvaluatorTest`, baseline reports, fixed fixtures, and demo scripts define current acceptance style. +- Q4 conclusion: add a small Chat prompt audit catalog near `ChatService` prompt loading and persist a compact snapshot with verifier evaluation. +- Q5 conclusion: run live E2E with `mvp-demo` profile if dependencies are available; otherwise record the blocker and rely on deterministic eval/unit evidence. + +## Specify / Alignment + +| Check | Status | Notes | +|---|---|---| +| proposal goals -> proposal | Aligned | Proposal covers demo preflight, eval expansion, prompt audit, and docs. | +| proposal scope/constraints -> design | Aligned | Design records no public API/schema changes, prompt audit shape, eval fields, and demo script behavior. | +| design decisions -> specs/tasks | Aligned | Specs cover persisted prompt audit, evaluator checks, baseline, and demo script outputs. | +| specs observable behavior -> tasks | Aligned | Each requirement has implementation and verification tasks. | + +## Audit + +Input -> processing -> output chain: + +```text +prompt resource metadata + -> ChatService / PromptAudit snapshot + -> verifier_evaluation.prompt_audit + -> Trace API / eval fixtures + -> DiagnosisTraceEvaluator baseline + +run-interview-demo-check.ps1 + -> service readiness + -> chat / trace / feedback + -> mvp/demo/output summary +``` + +Architecture risk assessment: + +1. The audit shape is intentionally compact and internal; storing full prompt text would create noisy traces and possible sensitive-content risk. +2. Eval should assert versions by explicit metadata, not by prompt content hashes that churn during local prompt edits. +3. Live demo checks may still be LOW_CONFID because LLM output is not deterministic; deterministic fixtures remain the regression source of truth. +4. No devflow/OpenSpec conflict found. + +## Commit Gate + +- `openspec validate interview-demo-quality-audit --strict`: passed. +- File completeness: + - `proposal.md`: present. + - `design.md`: present. + - `specs/`: present for `chat-verifier-agent`, `diagnosis-eval-harness`, and `mvp-demo-trace-acceptance`. + - `tasks.md`: present. +- Consistency: + - Proposal goals map to design sections. + - Design decisions map to spec requirements and executable tasks. + - Task acceptance checks are verifiable. +- `.committed` marker created. + +## Current Checkpoint + +- Commit completed. +- Apply is authorized by the original objective: "完成后归档提交". diff --git a/devflow/projects/2026-07-09-interview-demo-quality-audit/evidence.md b/devflow/projects/2026-07-09-interview-demo-quality-audit/evidence.md new file mode 100644 index 0000000..76d6616 --- /dev/null +++ b/devflow/projects/2026-07-09-interview-demo-quality-audit/evidence.md @@ -0,0 +1,25 @@ +# Evidence + +## Context Files Read + +- `devflow/index.md` +- `devflow/glossary/CONTEXT.md` +- `devflow/projects/2026-07-08-diagnosis-eval-demo-gatekeeper-closure/decisions.md` +- `devflow/projects/2026-07-08-executor-composer-final-answer/decisions.md` +- `mvp/architecture/current-mvp-architecture.md` +- `mvp/architecture/agent-orchestration.md` +- `mvp/architecture/executor-evidence-pipeline-refactor.md` +- `mvp/architecture/harness-quality-gates.md` +- `mvp/demo/README.md` +- `mvp/demo/ten-minute-interview-demo.md` +- `mvp/eval/README.md` +- `mvp/eval/cases/diagnosis-cases.json` +- `src/main/java/com/superbiz/agent/service/ChatService.java` +- `src/main/java/com/superbiz/agent/service/ExecutorGatekeeperService.java` +- `src/main/java/com/superbiz/agent/eval/DiagnosisTraceEvaluator.java` +- `src/main/resources/gatekeeper/gatekeeper-rules.json` + +## Tooling Note + +The required `codebase-retrieval` and LSP tools were not exposed in this session. Impact analysis used `rg`, direct file reads, existing OpenSpec/devflow artifacts, and targeted tests instead. + diff --git a/openspec/changes/interview-demo-quality-audit/.committed b/openspec/changes/interview-demo-quality-audit/.committed new file mode 100644 index 0000000..d0fe822 --- /dev/null +++ b/openspec/changes/interview-demo-quality-audit/.committed @@ -0,0 +1 @@ +committed diff --git a/openspec/changes/interview-demo-quality-audit/design.md b/openspec/changes/interview-demo-quality-audit/design.md new file mode 100644 index 0000000..6edc341 --- /dev/null +++ b/openspec/changes/interview-demo-quality-audit/design.md @@ -0,0 +1,117 @@ +# Design: Interview Demo Quality Audit + +## Overview + +The change adds auditability and demo readiness without changing Agent routing. + +```text +ChatService prompt resources + -> PromptAuditService + -> verifier_evaluation.prompt_audit + -> Trace API + -> DiagnosisTraceEvaluator + -> baseline report + +mvp/demo/scripts/run-interview-demo-check.ps1 + -> health/readiness check + -> payment timeout chat + -> trace fetch + -> feedback + -> demo output bundle +``` + +## Prompt Audit + +Add a compact `prompt_audit` object under: + +```text +diagnosis_session.self_evaluation.verifier_evaluation.prompt_audit +``` + +Shape: + +```json +{ + "version": "chat-prompts-v1", + "prompts": [ + { + "name": "chat_planner", + "version": "chat-planner-v1", + "resource": "prompts/chat-planner-prompt.md" + } + ] +} +``` + +Design choices: + +- Use explicit local metadata, not full prompt hashes, because the interview goal is explainable version audit rather than cryptographic integrity. +- Keep this metadata in code or a small resource catalog near prompt loading. +- Persist prompt audit with every Chat verifier evaluation, including fallback/degraded paths. +- Do not include full prompt text in trace. + +## Eval Expansion + +Extend `DiagnosisEvalCase` with optional fields: + +- `requirePromptAudit` +- `expectedPromptAuditVersion` +- `expectedPromptVersions` +- `requireGatekeeperRules` + +Evaluator behavior: + +- If `requirePromptAudit=true`, `verifier_evaluation.prompt_audit.version` must exist. +- If `expectedPromptAuditVersion` is set, it must match. +- If `expectedPromptVersions` is set, each listed prompt name/version pair must exist. +- If `requireGatekeeperRules=true`, `gatekeeper_result.rules` must be a non-empty list and each item must include `id`, `enabled`, and `default_severity`. + +Add at least two fixture-backed cases: + +- A positive audit closure case that requires prompt audit + Gatekeeper rules. +- A metadata-gap negative case represented as `LOW_CONFID`/safe final answer, used to prove the evaluator catches missing audit metadata when configured. + +The baseline must remain fully passing after fixtures are updated. + +## Demo Stabilization + +Add a PowerShell script: + +```text +mvp/demo/scripts/run-interview-demo-check.ps1 +``` + +Responsibilities: + +- Accept base URL and session id parameters. +- Check that the service is reachable. +- Run the existing payment-timeout chat request. +- Fetch trace for the same session id. +- Submit useful feedback. +- Write outputs under `mvp/demo/output/`. +- Emit a concise summary with session id, verdict, Gatekeeper rule version, prompt audit version, and output paths. + +The script should fail fast with actionable messages when the service is unavailable. + +## Documentation + +Add/update: + +- `mvp/demo/README.md`: mention the preflight script. +- `mvp/demo/ten-minute-interview-demo.md`: use the preflight script as the recommended path. +- `mvp/demo/interview-q-and-a.md`: concise interview answers for Agent engineering tradeoffs. +- `mvp/architecture/harness-quality-gates.md`: record prompt audit as part of the quality gate. + +## Verification + +Required: + +- Targeted unit/eval tests for prompt audit persistence and evaluator checks. +- Regenerated baseline JSON/Markdown reports. +- OpenSpec validation. + +E2E: + +- If local dependencies are available, run Spring Boot with `mvp-demo` profile and execute the new preflight script. +- If unavailable, record the reason and rely on deterministic unit/eval evidence. + diff --git a/openspec/changes/interview-demo-quality-audit/proposal.md b/openspec/changes/interview-demo-quality-audit/proposal.md new file mode 100644 index 0000000..f898e73 --- /dev/null +++ b/openspec/changes/interview-demo-quality-audit/proposal.md @@ -0,0 +1,58 @@ +# Change: Interview Demo Quality Audit + +## Problem + +SuperBizAgent MVP is now strong enough to demonstrate traceable Agent engineering, but the interview path still has three gaps: + +- The live demo has a run script, but no preflight command that checks service readiness and produces a concise interview evidence bundle. +- The diagnosis eval baseline covers the V2 evidence pipeline, but it does not yet assert prompt-version audit data and has limited coverage for audit metadata gaps. +- Gatekeeper already exposes `rule_set_version`, but prompt versions are not persisted with the verifier evaluation, making prompt changes harder to explain, compare, and roll back in an interview. + +This change stabilizes the MVP as an interview artifact rather than adding a new diagnosis architecture. + +## Proposed Solution + +Implement a small internal quality/audit increment: + +1. Add prompt version audit metadata to Chat verifier evaluation. +2. Extend deterministic diagnosis eval cases/fixtures to assert prompt audit metadata and audit-metadata failures. +3. Add an interview demo preflight script and documentation that can be run before or during a demo to verify service readiness, execute the payment timeout path, fetch trace, and record key audit fields. +4. Add/update MVP documentation for interview Q&A and the new audit/preflight workflow. + +## Scope + +In scope: + +- Internal `diagnosis_session.self_evaluation.verifier_evaluation` audit JSON. +- Diagnosis eval case schema, evaluator checks, fixtures, and baseline reports. +- MVP demo scripts/docs. +- Architecture/demo documentation for prompt and Gatekeeper version audit. + +Out of scope: + +- Public HTTP API changes. +- Database schema changes. +- New Agent roles, MCP tool server migration, process isolation, or AIOps LLM Verifier. +- Replacing existing `Planner -> Executor -> Gatekeeper -> Verifier -> Composer` orchestration. +- Guaranteeing live LLM `PASS` for every demo run. Live demo compatibility and deterministic fixture regression are separate acceptance paths. + +## Context Constraints From devflow + +- Evidence Tools produce incident facts and must be recorded in `tool_invocation`. +- Chat quality gates are layered: Gatekeeper verifies evidence references, Verifier judges derivability, Composer controls expression. +- `diagnosis eval` is deterministic and fixture-backed; no LLM-as-judge. +- Demo assets should be runnable, but interview safety should not depend solely on live LLM behavior. +- Gatekeeper rule metadata is metadata-only; dynamic rule execution is out of scope. + +## Interface Impact + +Level: L2 internal contract change. + +Reason: `verifier_evaluation` gains a compact `prompt_audit` object. Existing public API shape remains the same, and the value is exposed only through already-existing trace/self-evaluation JSON. + +## Risks + +- Baseline report churn is expected when adding cases; JSON and Markdown reports must be regenerated together. +- Prompt audit must be deterministic and stable enough for eval fixtures; avoid hashing full prompt text with environment-specific content. +- Demo preflight must not hardcode secrets and must tolerate local service unavailability with clear failure messages. + diff --git a/openspec/changes/interview-demo-quality-audit/specs/chat-verifier-agent/spec.md b/openspec/changes/interview-demo-quality-audit/specs/chat-verifier-agent/spec.md new file mode 100644 index 0000000..ee304ee --- /dev/null +++ b/openspec/changes/interview-demo-quality-audit/specs/chat-verifier-agent/spec.md @@ -0,0 +1,16 @@ +## MODIFIED Requirements + +### Requirement: Verifier SHALL be observable +The Verifier's verdict SHALL be persisted for observability. + +#### Scenario: prompt audit written to verifier evaluation +- **WHEN** the Chat verifier evaluation is persisted +- **THEN** the system SHALL include a `prompt_audit` object under `diagnosis_session.self_evaluation.verifier_evaluation` +- **AND** `prompt_audit.version` SHALL identify the Chat prompt audit catalog version +- **AND** `prompt_audit.prompts` SHALL include the planner, executor, verifier, and composer prompt names and versions +- **AND** full prompt text SHALL NOT be persisted in `prompt_audit` + +#### Scenario: prompt audit available on fallback paths +- **WHEN** Chat verifier parsing fails, Composer parsing fails, or Chat produces a degraded answer +- **THEN** the persisted verifier evaluation SHALL still include `prompt_audit` + diff --git a/openspec/changes/interview-demo-quality-audit/specs/diagnosis-eval-harness/spec.md b/openspec/changes/interview-demo-quality-audit/specs/diagnosis-eval-harness/spec.md new file mode 100644 index 0000000..e70a342 --- /dev/null +++ b/openspec/changes/interview-demo-quality-audit/specs/diagnosis-eval-harness/spec.md @@ -0,0 +1,21 @@ +## MODIFIED Requirements + +### Requirement: Diagnosis eval SHALL validate trace fixtures deterministically +The diagnosis eval harness SHALL evaluate saved trace fixtures without invoking an LLM judge. + +#### Scenario: prompt audit assertions are enforced +- **WHEN** an eval case sets `requirePromptAudit=true` +- **THEN** the evaluator SHALL require `verifier_evaluation.prompt_audit.version` +- **AND** when `expectedPromptAuditVersion` is configured, it SHALL match exactly +- **AND** when `expectedPromptVersions` is configured, each configured prompt name SHALL appear with the expected version + +#### Scenario: Gatekeeper rule metadata assertions are enforced +- **WHEN** an eval case sets `requireGatekeeperRules=true` +- **THEN** the evaluator SHALL require `verifier_evaluation.gatekeeper_result.rules` to be non-empty +- **AND** each rule item SHALL include `id`, `enabled`, and `default_severity` + +#### Scenario: expanded baseline remains passing +- **WHEN** the committed fixture set is evaluated +- **THEN** every case SHALL pass +- **AND** baseline JSON and Markdown reports SHALL reflect the expanded case count and verdict distribution + diff --git a/openspec/changes/interview-demo-quality-audit/specs/mvp-demo-trace-acceptance/spec.md b/openspec/changes/interview-demo-quality-audit/specs/mvp-demo-trace-acceptance/spec.md new file mode 100644 index 0000000..7a3df8d --- /dev/null +++ b/openspec/changes/interview-demo-quality-audit/specs/mvp-demo-trace-acceptance/spec.md @@ -0,0 +1,22 @@ +## MODIFIED Requirements + +### Requirement: MVP demo SHALL be reproducible for interviews +The MVP demo SHALL provide a repeatable way to show a diagnosis answer, trace, verifier evaluation, and feedback. + +#### Scenario: interview demo check script records an evidence bundle +- **WHEN** the user runs the interview demo check script against a running `mvp-demo` service +- **THEN** the script SHALL submit a fixed Chat diagnosis request +- **AND** it SHALL fetch the trace for the same session id +- **AND** it SHALL submit useful feedback for that session +- **AND** it SHALL write chat, trace, feedback, and summary outputs under `mvp/demo/output/` + +#### Scenario: interview demo check fails with actionable readiness output +- **WHEN** the target service is not reachable +- **THEN** the script SHALL fail before issuing diagnosis requests +- **AND** the failure message SHALL name the base URL and the expected startup profile + +#### Scenario: interview documentation explains audit fields +- **WHEN** an interviewer asks how prompt or Gatekeeper changes are audited +- **THEN** the demo documentation SHALL point to `prompt_audit.version` and `gatekeeper_result.rule_set_version` +- **AND** it SHALL explain that deterministic eval fixtures are the regression source of truth + diff --git a/openspec/changes/interview-demo-quality-audit/tasks.md b/openspec/changes/interview-demo-quality-audit/tasks.md new file mode 100644 index 0000000..3e32d1c --- /dev/null +++ b/openspec/changes/interview-demo-quality-audit/tasks.md @@ -0,0 +1,32 @@ +## 1. OpenSpec And devflow + +- [x] 1.1 Create OpenSpec proposal/design/spec/tasks for `interview-demo-quality-audit`. +- [x] 1.2 Record context, question pool, interface impact, audit, and verification plan in devflow decisions. +- [x] 1.3 Pass OpenSpec validation and create `.committed`. + +## 2. Prompt/Gatekeeper Version Audit + +- [ ] 2.1 Add compact Chat prompt audit metadata for planner, executor, verifier, and composer prompts. +- [ ] 2.2 Persist `prompt_audit` under `verifier_evaluation` for Chat verifier/composer outcomes. +- [ ] 2.3 Add focused tests proving prompt audit appears in persisted verifier evaluation. +- [ ] 2.4 Extend eval checks for prompt audit and Gatekeeper rule metadata. + +## 3. Eval Expansion + +- [ ] 3.1 Extend diagnosis eval case/result schema for prompt audit fields. +- [ ] 3.2 Add fixture-backed cases for audit closure coverage. +- [ ] 3.3 Regenerate baseline JSON and Markdown reports. +- [ ] 3.4 Update eval docs/schema. + +## 4. Interview Demo Stabilization + +- [ ] 4.1 Add `run-interview-demo-check.ps1` with service preflight, chat, trace, feedback, and summary output. +- [ ] 4.2 Update demo README and 10-minute script to use the preflight path. +- [ ] 4.3 Add interview Q&A documentation focused on Agent engineering tradeoffs. + +## 5. Verification And Archive + +- [ ] 5.1 Run targeted tests for ChatService/prompt audit and diagnosis eval. +- [ ] 5.2 Run relevant broader regression tests. +- [ ] 5.3 Run E2E demo check with `mvp-demo` profile if dependencies are available; otherwise record the blocker. +- [ ] 5.4 Archive the OpenSpec change, update devflow artifacts, and commit implementation + archive.