feat(demo): add interview quality audit

This commit is contained in:
zhuyongxin
2026-07-09 11:18:49 +08:00
parent a6c2d4459c
commit 9c9a0024d4
37 changed files with 1162 additions and 86 deletions
@@ -0,0 +1,32 @@
## 1. OpenSpec And devflow
- [x] 1.1 Create OpenSpec proposal/design/spec/tasks for `interview-demo-quality-audit`.
- [x] 1.2 Record context, question pool, interface impact, audit, and verification plan in devflow decisions.
- [x] 1.3 Pass OpenSpec validation and create `.committed`.
## 2. Prompt/Gatekeeper Version Audit
- [x] 2.1 Add compact Chat prompt audit metadata for planner, executor, verifier, and composer prompts.
- [x] 2.2 Persist `prompt_audit` under `verifier_evaluation` for Chat verifier/composer outcomes.
- [x] 2.3 Add focused tests proving prompt audit appears in persisted verifier evaluation.
- [x] 2.4 Extend eval checks for prompt audit and Gatekeeper rule metadata.
## 3. Eval Expansion
- [x] 3.1 Extend diagnosis eval case/result schema for prompt audit fields.
- [x] 3.2 Add fixture-backed cases for audit closure coverage.
- [x] 3.3 Regenerate baseline JSON and Markdown reports.
- [x] 3.4 Update eval docs/schema.
## 4. Interview Demo Stabilization
- [x] 4.1 Add `run-interview-demo-check.ps1` with service preflight, chat, trace, feedback, and summary output.
- [x] 4.2 Update demo README and 10-minute script to use the preflight path.
- [x] 4.3 Add interview Q&A documentation focused on Agent engineering tradeoffs.
## 5. Verification And Archive
- [x] 5.1 Run targeted tests for ChatService/prompt audit and diagnosis eval.
- [x] 5.2 Run relevant broader regression tests.
- [x] 5.3 Run E2E demo check with `mvp-demo` profile if dependencies are available; otherwise record the blocker.
- [x] 5.4 Archive the OpenSpec change, update devflow artifacts, and commit implementation + archive.
@@ -1,32 +0,0 @@
## 1. OpenSpec And devflow
- [x] 1.1 Create OpenSpec proposal/design/spec/tasks for `interview-demo-quality-audit`.
- [x] 1.2 Record context, question pool, interface impact, audit, and verification plan in devflow decisions.
- [x] 1.3 Pass OpenSpec validation and create `.committed`.
## 2. Prompt/Gatekeeper Version Audit
- [ ] 2.1 Add compact Chat prompt audit metadata for planner, executor, verifier, and composer prompts.
- [ ] 2.2 Persist `prompt_audit` under `verifier_evaluation` for Chat verifier/composer outcomes.
- [ ] 2.3 Add focused tests proving prompt audit appears in persisted verifier evaluation.
- [ ] 2.4 Extend eval checks for prompt audit and Gatekeeper rule metadata.
## 3. Eval Expansion
- [ ] 3.1 Extend diagnosis eval case/result schema for prompt audit fields.
- [ ] 3.2 Add fixture-backed cases for audit closure coverage.
- [ ] 3.3 Regenerate baseline JSON and Markdown reports.
- [ ] 3.4 Update eval docs/schema.
## 4. Interview Demo Stabilization
- [ ] 4.1 Add `run-interview-demo-check.ps1` with service preflight, chat, trace, feedback, and summary output.
- [ ] 4.2 Update demo README and 10-minute script to use the preflight path.
- [ ] 4.3 Add interview Q&A documentation focused on Agent engineering tradeoffs.
## 5. Verification And Archive
- [ ] 5.1 Run targeted tests for ChatService/prompt audit and diagnosis eval.
- [ ] 5.2 Run relevant broader regression tests.
- [ ] 5.3 Run E2E demo check with `mvp-demo` profile if dependencies are available; otherwise record the blocker.
- [ ] 5.4 Archive the OpenSpec change, update devflow artifacts, and commit implementation + archive.
@@ -145,6 +145,17 @@ The Verifier's verdict and downstream final-answer composition SHALL be persiste
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result`
- **AND** existing verifier fields such as `verdict`, `facts_checked`, `executor_output_parse_status`, and `tool_trace_summary` SHALL be preserved
#### Scenario: prompt audit written to verifier evaluation
- **WHEN** the Chat verifier evaluation is persisted
- **THEN** the system SHALL include a `prompt_audit` object under `diagnosis_session.self_evaluation.verifier_evaluation`
- **AND** `prompt_audit.version` SHALL identify the Chat prompt audit catalog version
- **AND** `prompt_audit.prompts` SHALL include the planner, executor, verifier, and composer prompt names and versions
- **AND** full prompt text SHALL NOT be persisted in `prompt_audit`
#### Scenario: prompt audit available on fallback paths
- **WHEN** Chat verifier parsing fails, Composer parsing fails, or Chat produces a degraded answer
- **THEN** the persisted verifier evaluation SHALL still include `prompt_audit`
### Requirement: self_evaluation SHALL be a container object
The `diagnosis_session.self_evaluation` field SHALL store multiple evaluation channels in one JSON object.
@@ -195,3 +195,22 @@ The evaluation harness SHALL be able to assert the Gatekeeper rule set version r
- **WHEN** an evaluation case declares `expectedGatekeeperRuleSetVersion`
- **AND** the fixture has a different or missing rule set version
- **THEN** the case SHALL fail with a clear failed check
### Requirement: Diagnosis eval SHALL validate trace fixtures deterministically
The diagnosis eval harness SHALL evaluate saved trace fixtures without invoking an LLM judge.
#### Scenario: prompt audit assertions are enforced
- **WHEN** an eval case sets `requirePromptAudit=true`
- **THEN** the evaluator SHALL require `verifier_evaluation.prompt_audit.version`
- **AND** when `expectedPromptAuditVersion` is configured, it SHALL match exactly
- **AND** when `expectedPromptVersions` is configured, each configured prompt name SHALL appear with the expected version
#### Scenario: Gatekeeper rule metadata assertions are enforced
- **WHEN** an eval case sets `requireGatekeeperRules=true`
- **THEN** the evaluator SHALL require `verifier_evaluation.gatekeeper_result.rules` to be non-empty
- **AND** each rule item SHALL include `id`, `enabled`, and `default_severity`
#### Scenario: expanded baseline remains passing
- **WHEN** the committed fixture set is evaluated
- **THEN** every case SHALL pass
- **AND** baseline JSON and Markdown reports SHALL reflect the expanded case count and verdict distribution
@@ -56,6 +56,26 @@ The MVP demo SHALL provide scripts and request payloads for running the payment-
- **WHEN** the demo script finishes successfully
- **THEN** it SHALL write chat, trace, and feedback responses under a demo output directory
### Requirement: MVP demo SHALL be reproducible for interviews
The MVP demo SHALL provide a repeatable way to show a diagnosis answer, trace, verifier evaluation, and feedback.
#### Scenario: interview demo check script records an evidence bundle
- **WHEN** the user runs the interview demo check script against a running `mvp-demo` service
- **THEN** the script SHALL submit a fixed Chat diagnosis request
- **AND** it SHALL fetch the trace for the same session id
- **AND** it SHALL submit useful feedback for that session
- **AND** it SHALL write chat, trace, feedback, and summary outputs under `mvp/demo/output/`
#### Scenario: interview demo check fails with actionable readiness output
- **WHEN** the target service is not reachable
- **THEN** the script SHALL fail before issuing diagnosis requests
- **AND** the failure message SHALL name the base URL and the expected startup profile
#### Scenario: interview documentation explains audit fields
- **WHEN** an interviewer asks how prompt or Gatekeeper changes are audited
- **THEN** the demo documentation SHALL point to `prompt_audit.version` and `gatekeeper_result.rule_set_version`
- **AND** it SHALL explain that deterministic eval fixtures are the regression source of truth
### Requirement: MVP demo SHALL provide a trace inspection checklist
The MVP demo SHALL document which trace fields to inspect for evidence, verifier behavior, and session-level auditability.