feat(demo): add interview quality audit

This commit is contained in:
zhuyongxin
2026-07-09 11:18:49 +08:00
parent a6c2d4459c
commit 9c9a0024d4
37 changed files with 1162 additions and 86 deletions
@@ -0,0 +1 @@
committed
@@ -0,0 +1,117 @@
# Design: Interview Demo Quality Audit
## Overview
The change adds auditability and demo readiness without changing Agent routing.
```text
ChatService prompt resources
-> PromptAuditService
-> verifier_evaluation.prompt_audit
-> Trace API
-> DiagnosisTraceEvaluator
-> baseline report
mvp/demo/scripts/run-interview-demo-check.ps1
-> health/readiness check
-> payment timeout chat
-> trace fetch
-> feedback
-> demo output bundle
```
## Prompt Audit
Add a compact `prompt_audit` object under:
```text
diagnosis_session.self_evaluation.verifier_evaluation.prompt_audit
```
Shape:
```json
{
"version": "chat-prompts-v1",
"prompts": [
{
"name": "chat_planner",
"version": "chat-planner-v1",
"resource": "prompts/chat-planner-prompt.md"
}
]
}
```
Design choices:
- Use explicit local metadata, not full prompt hashes, because the interview goal is explainable version audit rather than cryptographic integrity.
- Keep this metadata in code or a small resource catalog near prompt loading.
- Persist prompt audit with every Chat verifier evaluation, including fallback/degraded paths.
- Do not include full prompt text in trace.
## Eval Expansion
Extend `DiagnosisEvalCase` with optional fields:
- `requirePromptAudit`
- `expectedPromptAuditVersion`
- `expectedPromptVersions`
- `requireGatekeeperRules`
Evaluator behavior:
- If `requirePromptAudit=true`, `verifier_evaluation.prompt_audit.version` must exist.
- If `expectedPromptAuditVersion` is set, it must match.
- If `expectedPromptVersions` is set, each listed prompt name/version pair must exist.
- If `requireGatekeeperRules=true`, `gatekeeper_result.rules` must be a non-empty list and each item must include `id`, `enabled`, and `default_severity`.
Add at least two fixture-backed cases:
- A positive audit closure case that requires prompt audit + Gatekeeper rules.
- A metadata-gap negative case represented as `LOW_CONFID`/safe final answer, used to prove the evaluator catches missing audit metadata when configured.
The baseline must remain fully passing after fixtures are updated.
## Demo Stabilization
Add a PowerShell script:
```text
mvp/demo/scripts/run-interview-demo-check.ps1
```
Responsibilities:
- Accept base URL and session id parameters.
- Check that the service is reachable.
- Run the existing payment-timeout chat request.
- Fetch trace for the same session id.
- Submit useful feedback.
- Write outputs under `mvp/demo/output/`.
- Emit a concise summary with session id, verdict, Gatekeeper rule version, prompt audit version, and output paths.
The script should fail fast with actionable messages when the service is unavailable.
## Documentation
Add/update:
- `mvp/demo/README.md`: mention the preflight script.
- `mvp/demo/ten-minute-interview-demo.md`: use the preflight script as the recommended path.
- `mvp/demo/interview-q-and-a.md`: concise interview answers for Agent engineering tradeoffs.
- `mvp/architecture/harness-quality-gates.md`: record prompt audit as part of the quality gate.
## Verification
Required:
- Targeted unit/eval tests for prompt audit persistence and evaluator checks.
- Regenerated baseline JSON/Markdown reports.
- OpenSpec validation.
E2E:
- If local dependencies are available, run Spring Boot with `mvp-demo` profile and execute the new preflight script.
- If unavailable, record the reason and rely on deterministic unit/eval evidence.
@@ -0,0 +1,58 @@
# Change: Interview Demo Quality Audit
## Problem
SuperBizAgent MVP is now strong enough to demonstrate traceable Agent engineering, but the interview path still has three gaps:
- The live demo has a run script, but no preflight command that checks service readiness and produces a concise interview evidence bundle.
- The diagnosis eval baseline covers the V2 evidence pipeline, but it does not yet assert prompt-version audit data and has limited coverage for audit metadata gaps.
- Gatekeeper already exposes `rule_set_version`, but prompt versions are not persisted with the verifier evaluation, making prompt changes harder to explain, compare, and roll back in an interview.
This change stabilizes the MVP as an interview artifact rather than adding a new diagnosis architecture.
## Proposed Solution
Implement a small internal quality/audit increment:
1. Add prompt version audit metadata to Chat verifier evaluation.
2. Extend deterministic diagnosis eval cases/fixtures to assert prompt audit metadata and audit-metadata failures.
3. Add an interview demo preflight script and documentation that can be run before or during a demo to verify service readiness, execute the payment timeout path, fetch trace, and record key audit fields.
4. Add/update MVP documentation for interview Q&A and the new audit/preflight workflow.
## Scope
In scope:
- Internal `diagnosis_session.self_evaluation.verifier_evaluation` audit JSON.
- Diagnosis eval case schema, evaluator checks, fixtures, and baseline reports.
- MVP demo scripts/docs.
- Architecture/demo documentation for prompt and Gatekeeper version audit.
Out of scope:
- Public HTTP API changes.
- Database schema changes.
- New Agent roles, MCP tool server migration, process isolation, or AIOps LLM Verifier.
- Replacing existing `Planner -> Executor -> Gatekeeper -> Verifier -> Composer` orchestration.
- Guaranteeing live LLM `PASS` for every demo run. Live demo compatibility and deterministic fixture regression are separate acceptance paths.
## Context Constraints From devflow
- Evidence Tools produce incident facts and must be recorded in `tool_invocation`.
- Chat quality gates are layered: Gatekeeper verifies evidence references, Verifier judges derivability, Composer controls expression.
- `diagnosis eval` is deterministic and fixture-backed; no LLM-as-judge.
- Demo assets should be runnable, but interview safety should not depend solely on live LLM behavior.
- Gatekeeper rule metadata is metadata-only; dynamic rule execution is out of scope.
## Interface Impact
Level: L2 internal contract change.
Reason: `verifier_evaluation` gains a compact `prompt_audit` object. Existing public API shape remains the same, and the value is exposed only through already-existing trace/self-evaluation JSON.
## Risks
- Baseline report churn is expected when adding cases; JSON and Markdown reports must be regenerated together.
- Prompt audit must be deterministic and stable enough for eval fixtures; avoid hashing full prompt text with environment-specific content.
- Demo preflight must not hardcode secrets and must tolerate local service unavailability with clear failure messages.
@@ -0,0 +1,16 @@
## MODIFIED Requirements
### Requirement: Verifier SHALL be observable
The Verifier's verdict SHALL be persisted for observability.
#### Scenario: prompt audit written to verifier evaluation
- **WHEN** the Chat verifier evaluation is persisted
- **THEN** the system SHALL include a `prompt_audit` object under `diagnosis_session.self_evaluation.verifier_evaluation`
- **AND** `prompt_audit.version` SHALL identify the Chat prompt audit catalog version
- **AND** `prompt_audit.prompts` SHALL include the planner, executor, verifier, and composer prompt names and versions
- **AND** full prompt text SHALL NOT be persisted in `prompt_audit`
#### Scenario: prompt audit available on fallback paths
- **WHEN** Chat verifier parsing fails, Composer parsing fails, or Chat produces a degraded answer
- **THEN** the persisted verifier evaluation SHALL still include `prompt_audit`
@@ -0,0 +1,21 @@
## MODIFIED Requirements
### Requirement: Diagnosis eval SHALL validate trace fixtures deterministically
The diagnosis eval harness SHALL evaluate saved trace fixtures without invoking an LLM judge.
#### Scenario: prompt audit assertions are enforced
- **WHEN** an eval case sets `requirePromptAudit=true`
- **THEN** the evaluator SHALL require `verifier_evaluation.prompt_audit.version`
- **AND** when `expectedPromptAuditVersion` is configured, it SHALL match exactly
- **AND** when `expectedPromptVersions` is configured, each configured prompt name SHALL appear with the expected version
#### Scenario: Gatekeeper rule metadata assertions are enforced
- **WHEN** an eval case sets `requireGatekeeperRules=true`
- **THEN** the evaluator SHALL require `verifier_evaluation.gatekeeper_result.rules` to be non-empty
- **AND** each rule item SHALL include `id`, `enabled`, and `default_severity`
#### Scenario: expanded baseline remains passing
- **WHEN** the committed fixture set is evaluated
- **THEN** every case SHALL pass
- **AND** baseline JSON and Markdown reports SHALL reflect the expanded case count and verdict distribution
@@ -0,0 +1,22 @@
## MODIFIED Requirements
### Requirement: MVP demo SHALL be reproducible for interviews
The MVP demo SHALL provide a repeatable way to show a diagnosis answer, trace, verifier evaluation, and feedback.
#### Scenario: interview demo check script records an evidence bundle
- **WHEN** the user runs the interview demo check script against a running `mvp-demo` service
- **THEN** the script SHALL submit a fixed Chat diagnosis request
- **AND** it SHALL fetch the trace for the same session id
- **AND** it SHALL submit useful feedback for that session
- **AND** it SHALL write chat, trace, feedback, and summary outputs under `mvp/demo/output/`
#### Scenario: interview demo check fails with actionable readiness output
- **WHEN** the target service is not reachable
- **THEN** the script SHALL fail before issuing diagnosis requests
- **AND** the failure message SHALL name the base URL and the expected startup profile
#### Scenario: interview documentation explains audit fields
- **WHEN** an interviewer asks how prompt or Gatekeeper changes are audited
- **THEN** the demo documentation SHALL point to `prompt_audit.version` and `gatekeeper_result.rule_set_version`
- **AND** it SHALL explain that deterministic eval fixtures are the regression source of truth
@@ -0,0 +1,32 @@
## 1. OpenSpec And devflow
- [x] 1.1 Create OpenSpec proposal/design/spec/tasks for `interview-demo-quality-audit`.
- [x] 1.2 Record context, question pool, interface impact, audit, and verification plan in devflow decisions.
- [x] 1.3 Pass OpenSpec validation and create `.committed`.
## 2. Prompt/Gatekeeper Version Audit
- [x] 2.1 Add compact Chat prompt audit metadata for planner, executor, verifier, and composer prompts.
- [x] 2.2 Persist `prompt_audit` under `verifier_evaluation` for Chat verifier/composer outcomes.
- [x] 2.3 Add focused tests proving prompt audit appears in persisted verifier evaluation.
- [x] 2.4 Extend eval checks for prompt audit and Gatekeeper rule metadata.
## 3. Eval Expansion
- [x] 3.1 Extend diagnosis eval case/result schema for prompt audit fields.
- [x] 3.2 Add fixture-backed cases for audit closure coverage.
- [x] 3.3 Regenerate baseline JSON and Markdown reports.
- [x] 3.4 Update eval docs/schema.
## 4. Interview Demo Stabilization
- [x] 4.1 Add `run-interview-demo-check.ps1` with service preflight, chat, trace, feedback, and summary output.
- [x] 4.2 Update demo README and 10-minute script to use the preflight path.
- [x] 4.3 Add interview Q&A documentation focused on Agent engineering tradeoffs.
## 5. Verification And Archive
- [x] 5.1 Run targeted tests for ChatService/prompt audit and diagnosis eval.
- [x] 5.2 Run relevant broader regression tests.
- [x] 5.3 Run E2E demo check with `mvp-demo` profile if dependencies are available; otherwise record the blocker.
- [x] 5.4 Archive the OpenSpec change, update devflow artifacts, and commit implementation + archive.