feat(demo): add interview quality audit
This commit is contained in:
@@ -0,0 +1 @@
|
||||
ready
|
||||
@@ -0,0 +1,32 @@
|
||||
## 1. OpenSpec And devflow
|
||||
|
||||
- [x] 1.1 Create OpenSpec proposal/design/spec/tasks for `interview-demo-quality-audit`.
|
||||
- [x] 1.2 Record context, question pool, interface impact, audit, and verification plan in devflow decisions.
|
||||
- [x] 1.3 Pass OpenSpec validation and create `.committed`.
|
||||
|
||||
## 2. Prompt/Gatekeeper Version Audit
|
||||
|
||||
- [x] 2.1 Add compact Chat prompt audit metadata for planner, executor, verifier, and composer prompts.
|
||||
- [x] 2.2 Persist `prompt_audit` under `verifier_evaluation` for Chat verifier/composer outcomes.
|
||||
- [x] 2.3 Add focused tests proving prompt audit appears in persisted verifier evaluation.
|
||||
- [x] 2.4 Extend eval checks for prompt audit and Gatekeeper rule metadata.
|
||||
|
||||
## 3. Eval Expansion
|
||||
|
||||
- [x] 3.1 Extend diagnosis eval case/result schema for prompt audit fields.
|
||||
- [x] 3.2 Add fixture-backed cases for audit closure coverage.
|
||||
- [x] 3.3 Regenerate baseline JSON and Markdown reports.
|
||||
- [x] 3.4 Update eval docs/schema.
|
||||
|
||||
## 4. Interview Demo Stabilization
|
||||
|
||||
- [x] 4.1 Add `run-interview-demo-check.ps1` with service preflight, chat, trace, feedback, and summary output.
|
||||
- [x] 4.2 Update demo README and 10-minute script to use the preflight path.
|
||||
- [x] 4.3 Add interview Q&A documentation focused on Agent engineering tradeoffs.
|
||||
|
||||
## 5. Verification And Archive
|
||||
|
||||
- [x] 5.1 Run targeted tests for ChatService/prompt audit and diagnosis eval.
|
||||
- [x] 5.2 Run relevant broader regression tests.
|
||||
- [x] 5.3 Run E2E demo check with `mvp-demo` profile if dependencies are available; otherwise record the blocker.
|
||||
- [x] 5.4 Archive the OpenSpec change, update devflow artifacts, and commit implementation + archive.
|
||||
@@ -1,32 +0,0 @@
|
||||
## 1. OpenSpec And devflow
|
||||
|
||||
- [x] 1.1 Create OpenSpec proposal/design/spec/tasks for `interview-demo-quality-audit`.
|
||||
- [x] 1.2 Record context, question pool, interface impact, audit, and verification plan in devflow decisions.
|
||||
- [x] 1.3 Pass OpenSpec validation and create `.committed`.
|
||||
|
||||
## 2. Prompt/Gatekeeper Version Audit
|
||||
|
||||
- [ ] 2.1 Add compact Chat prompt audit metadata for planner, executor, verifier, and composer prompts.
|
||||
- [ ] 2.2 Persist `prompt_audit` under `verifier_evaluation` for Chat verifier/composer outcomes.
|
||||
- [ ] 2.3 Add focused tests proving prompt audit appears in persisted verifier evaluation.
|
||||
- [ ] 2.4 Extend eval checks for prompt audit and Gatekeeper rule metadata.
|
||||
|
||||
## 3. Eval Expansion
|
||||
|
||||
- [ ] 3.1 Extend diagnosis eval case/result schema for prompt audit fields.
|
||||
- [ ] 3.2 Add fixture-backed cases for audit closure coverage.
|
||||
- [ ] 3.3 Regenerate baseline JSON and Markdown reports.
|
||||
- [ ] 3.4 Update eval docs/schema.
|
||||
|
||||
## 4. Interview Demo Stabilization
|
||||
|
||||
- [ ] 4.1 Add `run-interview-demo-check.ps1` with service preflight, chat, trace, feedback, and summary output.
|
||||
- [ ] 4.2 Update demo README and 10-minute script to use the preflight path.
|
||||
- [ ] 4.3 Add interview Q&A documentation focused on Agent engineering tradeoffs.
|
||||
|
||||
## 5. Verification And Archive
|
||||
|
||||
- [ ] 5.1 Run targeted tests for ChatService/prompt audit and diagnosis eval.
|
||||
- [ ] 5.2 Run relevant broader regression tests.
|
||||
- [ ] 5.3 Run E2E demo check with `mvp-demo` profile if dependencies are available; otherwise record the blocker.
|
||||
- [ ] 5.4 Archive the OpenSpec change, update devflow artifacts, and commit implementation + archive.
|
||||
@@ -145,6 +145,17 @@ The Verifier's verdict and downstream final-answer composition SHALL be persiste
|
||||
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result`
|
||||
- **AND** existing verifier fields such as `verdict`, `facts_checked`, `executor_output_parse_status`, and `tool_trace_summary` SHALL be preserved
|
||||
|
||||
#### Scenario: prompt audit written to verifier evaluation
|
||||
- **WHEN** the Chat verifier evaluation is persisted
|
||||
- **THEN** the system SHALL include a `prompt_audit` object under `diagnosis_session.self_evaluation.verifier_evaluation`
|
||||
- **AND** `prompt_audit.version` SHALL identify the Chat prompt audit catalog version
|
||||
- **AND** `prompt_audit.prompts` SHALL include the planner, executor, verifier, and composer prompt names and versions
|
||||
- **AND** full prompt text SHALL NOT be persisted in `prompt_audit`
|
||||
|
||||
#### Scenario: prompt audit available on fallback paths
|
||||
- **WHEN** Chat verifier parsing fails, Composer parsing fails, or Chat produces a degraded answer
|
||||
- **THEN** the persisted verifier evaluation SHALL still include `prompt_audit`
|
||||
|
||||
### Requirement: self_evaluation SHALL be a container object
|
||||
The `diagnosis_session.self_evaluation` field SHALL store multiple evaluation channels in one JSON object.
|
||||
|
||||
|
||||
@@ -195,3 +195,22 @@ The evaluation harness SHALL be able to assert the Gatekeeper rule set version r
|
||||
- **WHEN** an evaluation case declares `expectedGatekeeperRuleSetVersion`
|
||||
- **AND** the fixture has a different or missing rule set version
|
||||
- **THEN** the case SHALL fail with a clear failed check
|
||||
|
||||
### Requirement: Diagnosis eval SHALL validate trace fixtures deterministically
|
||||
The diagnosis eval harness SHALL evaluate saved trace fixtures without invoking an LLM judge.
|
||||
|
||||
#### Scenario: prompt audit assertions are enforced
|
||||
- **WHEN** an eval case sets `requirePromptAudit=true`
|
||||
- **THEN** the evaluator SHALL require `verifier_evaluation.prompt_audit.version`
|
||||
- **AND** when `expectedPromptAuditVersion` is configured, it SHALL match exactly
|
||||
- **AND** when `expectedPromptVersions` is configured, each configured prompt name SHALL appear with the expected version
|
||||
|
||||
#### Scenario: Gatekeeper rule metadata assertions are enforced
|
||||
- **WHEN** an eval case sets `requireGatekeeperRules=true`
|
||||
- **THEN** the evaluator SHALL require `verifier_evaluation.gatekeeper_result.rules` to be non-empty
|
||||
- **AND** each rule item SHALL include `id`, `enabled`, and `default_severity`
|
||||
|
||||
#### Scenario: expanded baseline remains passing
|
||||
- **WHEN** the committed fixture set is evaluated
|
||||
- **THEN** every case SHALL pass
|
||||
- **AND** baseline JSON and Markdown reports SHALL reflect the expanded case count and verdict distribution
|
||||
|
||||
@@ -56,6 +56,26 @@ The MVP demo SHALL provide scripts and request payloads for running the payment-
|
||||
- **WHEN** the demo script finishes successfully
|
||||
- **THEN** it SHALL write chat, trace, and feedback responses under a demo output directory
|
||||
|
||||
### Requirement: MVP demo SHALL be reproducible for interviews
|
||||
The MVP demo SHALL provide a repeatable way to show a diagnosis answer, trace, verifier evaluation, and feedback.
|
||||
|
||||
#### Scenario: interview demo check script records an evidence bundle
|
||||
- **WHEN** the user runs the interview demo check script against a running `mvp-demo` service
|
||||
- **THEN** the script SHALL submit a fixed Chat diagnosis request
|
||||
- **AND** it SHALL fetch the trace for the same session id
|
||||
- **AND** it SHALL submit useful feedback for that session
|
||||
- **AND** it SHALL write chat, trace, feedback, and summary outputs under `mvp/demo/output/`
|
||||
|
||||
#### Scenario: interview demo check fails with actionable readiness output
|
||||
- **WHEN** the target service is not reachable
|
||||
- **THEN** the script SHALL fail before issuing diagnosis requests
|
||||
- **AND** the failure message SHALL name the base URL and the expected startup profile
|
||||
|
||||
#### Scenario: interview documentation explains audit fields
|
||||
- **WHEN** an interviewer asks how prompt or Gatekeeper changes are audited
|
||||
- **THEN** the demo documentation SHALL point to `prompt_audit.version` and `gatekeeper_result.rule_set_version`
|
||||
- **AND** it SHALL explain that deterministic eval fixtures are the regression source of truth
|
||||
|
||||
### Requirement: MVP demo SHALL provide a trace inspection checklist
|
||||
The MVP demo SHALL document which trace fields to inspect for evidence, verifier behavior, and session-level auditability.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user