docs(openspec): propose interview demo quality audit
This commit is contained in:
@@ -0,0 +1 @@
|
||||
committed
|
||||
@@ -0,0 +1,117 @@
|
||||
# Design: Interview Demo Quality Audit
|
||||
|
||||
## Overview
|
||||
|
||||
The change adds auditability and demo readiness without changing Agent routing.
|
||||
|
||||
```text
|
||||
ChatService prompt resources
|
||||
-> PromptAuditService
|
||||
-> verifier_evaluation.prompt_audit
|
||||
-> Trace API
|
||||
-> DiagnosisTraceEvaluator
|
||||
-> baseline report
|
||||
|
||||
mvp/demo/scripts/run-interview-demo-check.ps1
|
||||
-> health/readiness check
|
||||
-> payment timeout chat
|
||||
-> trace fetch
|
||||
-> feedback
|
||||
-> demo output bundle
|
||||
```
|
||||
|
||||
## Prompt Audit
|
||||
|
||||
Add a compact `prompt_audit` object under:
|
||||
|
||||
```text
|
||||
diagnosis_session.self_evaluation.verifier_evaluation.prompt_audit
|
||||
```
|
||||
|
||||
Shape:
|
||||
|
||||
```json
|
||||
{
|
||||
"version": "chat-prompts-v1",
|
||||
"prompts": [
|
||||
{
|
||||
"name": "chat_planner",
|
||||
"version": "chat-planner-v1",
|
||||
"resource": "prompts/chat-planner-prompt.md"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
Design choices:
|
||||
|
||||
- Use explicit local metadata, not full prompt hashes, because the interview goal is explainable version audit rather than cryptographic integrity.
|
||||
- Keep this metadata in code or a small resource catalog near prompt loading.
|
||||
- Persist prompt audit with every Chat verifier evaluation, including fallback/degraded paths.
|
||||
- Do not include full prompt text in trace.
|
||||
|
||||
## Eval Expansion
|
||||
|
||||
Extend `DiagnosisEvalCase` with optional fields:
|
||||
|
||||
- `requirePromptAudit`
|
||||
- `expectedPromptAuditVersion`
|
||||
- `expectedPromptVersions`
|
||||
- `requireGatekeeperRules`
|
||||
|
||||
Evaluator behavior:
|
||||
|
||||
- If `requirePromptAudit=true`, `verifier_evaluation.prompt_audit.version` must exist.
|
||||
- If `expectedPromptAuditVersion` is set, it must match.
|
||||
- If `expectedPromptVersions` is set, each listed prompt name/version pair must exist.
|
||||
- If `requireGatekeeperRules=true`, `gatekeeper_result.rules` must be a non-empty list and each item must include `id`, `enabled`, and `default_severity`.
|
||||
|
||||
Add at least two fixture-backed cases:
|
||||
|
||||
- A positive audit closure case that requires prompt audit + Gatekeeper rules.
|
||||
- A metadata-gap negative case represented as `LOW_CONFID`/safe final answer, used to prove the evaluator catches missing audit metadata when configured.
|
||||
|
||||
The baseline must remain fully passing after fixtures are updated.
|
||||
|
||||
## Demo Stabilization
|
||||
|
||||
Add a PowerShell script:
|
||||
|
||||
```text
|
||||
mvp/demo/scripts/run-interview-demo-check.ps1
|
||||
```
|
||||
|
||||
Responsibilities:
|
||||
|
||||
- Accept base URL and session id parameters.
|
||||
- Check that the service is reachable.
|
||||
- Run the existing payment-timeout chat request.
|
||||
- Fetch trace for the same session id.
|
||||
- Submit useful feedback.
|
||||
- Write outputs under `mvp/demo/output/`.
|
||||
- Emit a concise summary with session id, verdict, Gatekeeper rule version, prompt audit version, and output paths.
|
||||
|
||||
The script should fail fast with actionable messages when the service is unavailable.
|
||||
|
||||
## Documentation
|
||||
|
||||
Add/update:
|
||||
|
||||
- `mvp/demo/README.md`: mention the preflight script.
|
||||
- `mvp/demo/ten-minute-interview-demo.md`: use the preflight script as the recommended path.
|
||||
- `mvp/demo/interview-q-and-a.md`: concise interview answers for Agent engineering tradeoffs.
|
||||
- `mvp/architecture/harness-quality-gates.md`: record prompt audit as part of the quality gate.
|
||||
|
||||
## Verification
|
||||
|
||||
Required:
|
||||
|
||||
- Targeted unit/eval tests for prompt audit persistence and evaluator checks.
|
||||
- Regenerated baseline JSON/Markdown reports.
|
||||
- OpenSpec validation.
|
||||
|
||||
E2E:
|
||||
|
||||
- If local dependencies are available, run Spring Boot with `mvp-demo` profile and execute the new preflight script.
|
||||
- If unavailable, record the reason and rely on deterministic unit/eval evidence.
|
||||
|
||||
@@ -0,0 +1,58 @@
|
||||
# Change: Interview Demo Quality Audit
|
||||
|
||||
## Problem
|
||||
|
||||
SuperBizAgent MVP is now strong enough to demonstrate traceable Agent engineering, but the interview path still has three gaps:
|
||||
|
||||
- The live demo has a run script, but no preflight command that checks service readiness and produces a concise interview evidence bundle.
|
||||
- The diagnosis eval baseline covers the V2 evidence pipeline, but it does not yet assert prompt-version audit data and has limited coverage for audit metadata gaps.
|
||||
- Gatekeeper already exposes `rule_set_version`, but prompt versions are not persisted with the verifier evaluation, making prompt changes harder to explain, compare, and roll back in an interview.
|
||||
|
||||
This change stabilizes the MVP as an interview artifact rather than adding a new diagnosis architecture.
|
||||
|
||||
## Proposed Solution
|
||||
|
||||
Implement a small internal quality/audit increment:
|
||||
|
||||
1. Add prompt version audit metadata to Chat verifier evaluation.
|
||||
2. Extend deterministic diagnosis eval cases/fixtures to assert prompt audit metadata and audit-metadata failures.
|
||||
3. Add an interview demo preflight script and documentation that can be run before or during a demo to verify service readiness, execute the payment timeout path, fetch trace, and record key audit fields.
|
||||
4. Add/update MVP documentation for interview Q&A and the new audit/preflight workflow.
|
||||
|
||||
## Scope
|
||||
|
||||
In scope:
|
||||
|
||||
- Internal `diagnosis_session.self_evaluation.verifier_evaluation` audit JSON.
|
||||
- Diagnosis eval case schema, evaluator checks, fixtures, and baseline reports.
|
||||
- MVP demo scripts/docs.
|
||||
- Architecture/demo documentation for prompt and Gatekeeper version audit.
|
||||
|
||||
Out of scope:
|
||||
|
||||
- Public HTTP API changes.
|
||||
- Database schema changes.
|
||||
- New Agent roles, MCP tool server migration, process isolation, or AIOps LLM Verifier.
|
||||
- Replacing existing `Planner -> Executor -> Gatekeeper -> Verifier -> Composer` orchestration.
|
||||
- Guaranteeing live LLM `PASS` for every demo run. Live demo compatibility and deterministic fixture regression are separate acceptance paths.
|
||||
|
||||
## Context Constraints From devflow
|
||||
|
||||
- Evidence Tools produce incident facts and must be recorded in `tool_invocation`.
|
||||
- Chat quality gates are layered: Gatekeeper verifies evidence references, Verifier judges derivability, Composer controls expression.
|
||||
- `diagnosis eval` is deterministic and fixture-backed; no LLM-as-judge.
|
||||
- Demo assets should be runnable, but interview safety should not depend solely on live LLM behavior.
|
||||
- Gatekeeper rule metadata is metadata-only; dynamic rule execution is out of scope.
|
||||
|
||||
## Interface Impact
|
||||
|
||||
Level: L2 internal contract change.
|
||||
|
||||
Reason: `verifier_evaluation` gains a compact `prompt_audit` object. Existing public API shape remains the same, and the value is exposed only through already-existing trace/self-evaluation JSON.
|
||||
|
||||
## Risks
|
||||
|
||||
- Baseline report churn is expected when adding cases; JSON and Markdown reports must be regenerated together.
|
||||
- Prompt audit must be deterministic and stable enough for eval fixtures; avoid hashing full prompt text with environment-specific content.
|
||||
- Demo preflight must not hardcode secrets and must tolerate local service unavailability with clear failure messages.
|
||||
|
||||
@@ -0,0 +1,16 @@
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Verifier SHALL be observable
|
||||
The Verifier's verdict SHALL be persisted for observability.
|
||||
|
||||
#### Scenario: prompt audit written to verifier evaluation
|
||||
- **WHEN** the Chat verifier evaluation is persisted
|
||||
- **THEN** the system SHALL include a `prompt_audit` object under `diagnosis_session.self_evaluation.verifier_evaluation`
|
||||
- **AND** `prompt_audit.version` SHALL identify the Chat prompt audit catalog version
|
||||
- **AND** `prompt_audit.prompts` SHALL include the planner, executor, verifier, and composer prompt names and versions
|
||||
- **AND** full prompt text SHALL NOT be persisted in `prompt_audit`
|
||||
|
||||
#### Scenario: prompt audit available on fallback paths
|
||||
- **WHEN** Chat verifier parsing fails, Composer parsing fails, or Chat produces a degraded answer
|
||||
- **THEN** the persisted verifier evaluation SHALL still include `prompt_audit`
|
||||
|
||||
@@ -0,0 +1,21 @@
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Diagnosis eval SHALL validate trace fixtures deterministically
|
||||
The diagnosis eval harness SHALL evaluate saved trace fixtures without invoking an LLM judge.
|
||||
|
||||
#### Scenario: prompt audit assertions are enforced
|
||||
- **WHEN** an eval case sets `requirePromptAudit=true`
|
||||
- **THEN** the evaluator SHALL require `verifier_evaluation.prompt_audit.version`
|
||||
- **AND** when `expectedPromptAuditVersion` is configured, it SHALL match exactly
|
||||
- **AND** when `expectedPromptVersions` is configured, each configured prompt name SHALL appear with the expected version
|
||||
|
||||
#### Scenario: Gatekeeper rule metadata assertions are enforced
|
||||
- **WHEN** an eval case sets `requireGatekeeperRules=true`
|
||||
- **THEN** the evaluator SHALL require `verifier_evaluation.gatekeeper_result.rules` to be non-empty
|
||||
- **AND** each rule item SHALL include `id`, `enabled`, and `default_severity`
|
||||
|
||||
#### Scenario: expanded baseline remains passing
|
||||
- **WHEN** the committed fixture set is evaluated
|
||||
- **THEN** every case SHALL pass
|
||||
- **AND** baseline JSON and Markdown reports SHALL reflect the expanded case count and verdict distribution
|
||||
|
||||
@@ -0,0 +1,22 @@
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: MVP demo SHALL be reproducible for interviews
|
||||
The MVP demo SHALL provide a repeatable way to show a diagnosis answer, trace, verifier evaluation, and feedback.
|
||||
|
||||
#### Scenario: interview demo check script records an evidence bundle
|
||||
- **WHEN** the user runs the interview demo check script against a running `mvp-demo` service
|
||||
- **THEN** the script SHALL submit a fixed Chat diagnosis request
|
||||
- **AND** it SHALL fetch the trace for the same session id
|
||||
- **AND** it SHALL submit useful feedback for that session
|
||||
- **AND** it SHALL write chat, trace, feedback, and summary outputs under `mvp/demo/output/`
|
||||
|
||||
#### Scenario: interview demo check fails with actionable readiness output
|
||||
- **WHEN** the target service is not reachable
|
||||
- **THEN** the script SHALL fail before issuing diagnosis requests
|
||||
- **AND** the failure message SHALL name the base URL and the expected startup profile
|
||||
|
||||
#### Scenario: interview documentation explains audit fields
|
||||
- **WHEN** an interviewer asks how prompt or Gatekeeper changes are audited
|
||||
- **THEN** the demo documentation SHALL point to `prompt_audit.version` and `gatekeeper_result.rule_set_version`
|
||||
- **AND** it SHALL explain that deterministic eval fixtures are the regression source of truth
|
||||
|
||||
@@ -0,0 +1,32 @@
|
||||
## 1. OpenSpec And devflow
|
||||
|
||||
- [x] 1.1 Create OpenSpec proposal/design/spec/tasks for `interview-demo-quality-audit`.
|
||||
- [x] 1.2 Record context, question pool, interface impact, audit, and verification plan in devflow decisions.
|
||||
- [x] 1.3 Pass OpenSpec validation and create `.committed`.
|
||||
|
||||
## 2. Prompt/Gatekeeper Version Audit
|
||||
|
||||
- [ ] 2.1 Add compact Chat prompt audit metadata for planner, executor, verifier, and composer prompts.
|
||||
- [ ] 2.2 Persist `prompt_audit` under `verifier_evaluation` for Chat verifier/composer outcomes.
|
||||
- [ ] 2.3 Add focused tests proving prompt audit appears in persisted verifier evaluation.
|
||||
- [ ] 2.4 Extend eval checks for prompt audit and Gatekeeper rule metadata.
|
||||
|
||||
## 3. Eval Expansion
|
||||
|
||||
- [ ] 3.1 Extend diagnosis eval case/result schema for prompt audit fields.
|
||||
- [ ] 3.2 Add fixture-backed cases for audit closure coverage.
|
||||
- [ ] 3.3 Regenerate baseline JSON and Markdown reports.
|
||||
- [ ] 3.4 Update eval docs/schema.
|
||||
|
||||
## 4. Interview Demo Stabilization
|
||||
|
||||
- [ ] 4.1 Add `run-interview-demo-check.ps1` with service preflight, chat, trace, feedback, and summary output.
|
||||
- [ ] 4.2 Update demo README and 10-minute script to use the preflight path.
|
||||
- [ ] 4.3 Add interview Q&A documentation focused on Agent engineering tradeoffs.
|
||||
|
||||
## 5. Verification And Archive
|
||||
|
||||
- [ ] 5.1 Run targeted tests for ChatService/prompt audit and diagnosis eval.
|
||||
- [ ] 5.2 Run relevant broader regression tests.
|
||||
- [ ] 5.3 Run E2E demo check with `mvp-demo` profile if dependencies are available; otherwise record the blocker.
|
||||
- [ ] 5.4 Archive the OpenSpec change, update devflow artifacts, and commit implementation + archive.
|
||||
Reference in New Issue
Block a user