feat(demo): add interview quality audit
This commit is contained in:
@@ -1,10 +1,46 @@
|
||||
# Acceptance
|
||||
|
||||
Acceptance will be filled during Apply/Archive with concrete command output.
|
||||
## Static Verification
|
||||
|
||||
Planned checks:
|
||||
- `openspec validate interview-demo-quality-audit --strict`
|
||||
- Result: passed.
|
||||
- Coverage: OpenSpec proposal/design/spec/tasks consistency.
|
||||
- PowerShell parser/runtime readiness check:
|
||||
- Command: `powershell -NoProfile -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1 -BaseUrl http://127.0.0.1:1 -OutputDir target/demo-check-syntax`
|
||||
- Result: expected failure with actionable readiness message.
|
||||
- Coverage: script parses under Windows PowerShell and fails before issuing diagnosis requests when service is unreachable.
|
||||
|
||||
- `openspec validate interview-demo-quality-audit --strict` - passed during commit gate.
|
||||
## Script Verification
|
||||
|
||||
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
|
||||
- Result: passed.
|
||||
- Coverage: 12/12 fixed eval fixtures, Prompt audit evaluator checks, Gatekeeper rule metadata checks, regenerated baseline reports.
|
||||
- `mvn -q "-Dtest=ChatServiceSequentialAgentTest" test`
|
||||
- Result: passed.
|
||||
- Coverage: Chat verifier evaluation persists `prompt_audit`.
|
||||
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ChatServiceSequentialAgentTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`
|
||||
- Result: passed.
|
||||
- Coverage: broader eval, baseline diff, Chat sequential flow, Gatekeeper, and Verifier input hook regression set.
|
||||
- `mvn -q -DskipTests compile`
|
||||
- `powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1` against `mvp-demo` service if dependencies are available.
|
||||
- Result: passed.
|
||||
- Coverage: main source compilation.
|
||||
|
||||
## E2E Verification
|
||||
|
||||
- Started service with:
|
||||
- `mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo`
|
||||
- Ran:
|
||||
- `powershell -NoProfile -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1 -BaseUrl http://localhost:9900 -SessionId mvp-demo-interview-quality-audit-001`
|
||||
- Result: passed.
|
||||
- Summary:
|
||||
- `chatSuccess=true`
|
||||
- `verdict=LOW_CONFID`
|
||||
- `gatekeeperStatus=fail`
|
||||
- `gatekeeperRuleSetVersion=gatekeeper-rules-v1`
|
||||
- `promptAuditVersion=chat-prompts-v1`
|
||||
- tools included `lookup_knowledge`, `query_logs`, `query_metrics`, and `get_available_log_topics`
|
||||
- Note: live E2E remains a compatibility check, not the deterministic PASS oracle. The fixed fixture baseline is the regression source of truth.
|
||||
|
||||
## Not Verified
|
||||
|
||||
- Browser UI inspection was not required for this change because the scope is backend trace/eval/demo script documentation, not frontend behavior.
|
||||
|
||||
@@ -89,3 +89,23 @@ Architecture risk assessment:
|
||||
|
||||
- Commit completed.
|
||||
- Apply is authorized by the original objective: "完成后归档提交".
|
||||
|
||||
## Pre-apply Research
|
||||
|
||||
- Capability source: sm-flow built-in apply protocol. `openspec-apply-change` was not invoked directly in this session.
|
||||
- Repository semantic search/LSP note: the requested `codebase-retrieval` and LSP tools were not available in the exposed toolset, so impact analysis used `rg`, direct file reads, OpenSpec/devflow artifacts, and targeted tests.
|
||||
- Reference implementation and reuse:
|
||||
- `ChatService.persistVerifierEvaluation(...)` is the single persistence point for Chat verifier/composer audit data; prompt audit was added there to cover normal, fallback, and degraded Composer paths.
|
||||
- `DiagnosisTraceEvaluator` and `DiagnosisEvalReportWriter` are the deterministic eval extension points; no LLM judge was introduced.
|
||||
- `mvp/demo/scripts/run-payment-timeout-demo.ps1` provided the request/trace/feedback flow reused by the new interview preflight script.
|
||||
- Interface impact remains L2 internal trace contract: `verifier_evaluation.prompt_audit` and eval report fields are added; no public endpoint, table, or request DTO changed.
|
||||
|
||||
## Apply Notes
|
||||
|
||||
- Added compact Chat prompt audit metadata: `chat-prompts-v1`, with planner/executor/verifier/composer prompt versions and resource paths.
|
||||
- Extended diagnosis eval schema, result reporting, baseline fixtures, JSON report, and Markdown report for Prompt audit and Gatekeeper rule metadata.
|
||||
- Added two fixture-backed audit cases:
|
||||
- `prompt-gatekeeper-audit-closure`
|
||||
- `audit-metadata-low-confid`
|
||||
- Added `mvp/demo/scripts/run-interview-demo-check.ps1` to run service readiness, Chat, Trace, feedback, and summary output.
|
||||
- Updated MVP demo/eval/architecture docs to explain `prompt_audit.version`, `gatekeeper_result.rule_set_version`, and deterministic fixture baseline.
|
||||
|
||||
@@ -23,3 +23,36 @@
|
||||
|
||||
The required `codebase-retrieval` and LSP tools were not exposed in this session. Impact analysis used `rg`, direct file reads, existing OpenSpec/devflow artifacts, and targeted tests instead.
|
||||
|
||||
## Implementation Evidence
|
||||
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- Adds `prompt_audit` under `verifier_evaluation` through the shared `persistVerifierEvaluation(...)` path.
|
||||
- Uses compact metadata only: audit version, prompt names, prompt versions, and resource paths.
|
||||
- `src/main/java/com/superbiz/agent/eval/DiagnosisTraceEvaluator.java`
|
||||
- Adds deterministic checks for `requirePromptAudit`, `expectedPromptAuditVersion`, `expectedPromptVersions`, and `requireGatekeeperRules`.
|
||||
- `src/main/java/com/superbiz/agent/eval/DiagnosisEvalReportWriter.java`
|
||||
- Adds Prompt Audit and Gatekeeper rule count columns to Markdown reports.
|
||||
- `mvp/eval/cases/diagnosis-cases.json`
|
||||
- Expands fixed baseline to 12 fixture-backed cases.
|
||||
- `mvp/eval/fixtures/prompt-gatekeeper-audit-closure-pass.json`
|
||||
- Positive PASS fixture proving Prompt audit and Gatekeeper rule metadata closure.
|
||||
- `mvp/eval/fixtures/audit-metadata-low-confid.json`
|
||||
- LOW_CONFID fixture proving safe answer behavior while audit metadata remains present.
|
||||
- `mvp/demo/scripts/run-interview-demo-check.ps1`
|
||||
- Adds service readiness, Chat, Trace, feedback, and summary output for interview preflight.
|
||||
|
||||
## Verification Evidence
|
||||
|
||||
- OpenSpec:
|
||||
- `openspec validate interview-demo-quality-audit --strict`: passed before archive.
|
||||
- `openspec validate --specs --strict`: 10 specs passed after merging deltas into main specs.
|
||||
- Unit/eval:
|
||||
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`: passed.
|
||||
- `mvn -q "-Dtest=ChatServiceSequentialAgentTest" test`: passed.
|
||||
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ChatServiceSequentialAgentTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`: passed.
|
||||
- Compile:
|
||||
- `mvn -q -DskipTests compile`: passed.
|
||||
- E2E:
|
||||
- Started `mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo`.
|
||||
- Ran `mvp/demo/scripts/run-interview-demo-check.ps1` against `http://localhost:9900`.
|
||||
- Summary recorded `chatSuccess=true`, `verdict=LOW_CONFID`, `gatekeeperRuleSetVersion=gatekeeper-rules-v1`, and `promptAuditVersion=chat-prompts-v1`.
|
||||
|
||||
Reference in New Issue
Block a user