docs(openspec): propose interview demo quality audit

This commit is contained in:
zhuyongxin
2026-07-09 10:34:33 +08:00
parent da45fa3fb0
commit a6c2d4459c
11 changed files with 416 additions and 0 deletions
@@ -0,0 +1,16 @@
## MODIFIED Requirements
### Requirement: Verifier SHALL be observable
The Verifier's verdict SHALL be persisted for observability.
#### Scenario: prompt audit written to verifier evaluation
- **WHEN** the Chat verifier evaluation is persisted
- **THEN** the system SHALL include a `prompt_audit` object under `diagnosis_session.self_evaluation.verifier_evaluation`
- **AND** `prompt_audit.version` SHALL identify the Chat prompt audit catalog version
- **AND** `prompt_audit.prompts` SHALL include the planner, executor, verifier, and composer prompt names and versions
- **AND** full prompt text SHALL NOT be persisted in `prompt_audit`
#### Scenario: prompt audit available on fallback paths
- **WHEN** Chat verifier parsing fails, Composer parsing fails, or Chat produces a degraded answer
- **THEN** the persisted verifier evaluation SHALL still include `prompt_audit`
@@ -0,0 +1,21 @@
## MODIFIED Requirements
### Requirement: Diagnosis eval SHALL validate trace fixtures deterministically
The diagnosis eval harness SHALL evaluate saved trace fixtures without invoking an LLM judge.
#### Scenario: prompt audit assertions are enforced
- **WHEN** an eval case sets `requirePromptAudit=true`
- **THEN** the evaluator SHALL require `verifier_evaluation.prompt_audit.version`
- **AND** when `expectedPromptAuditVersion` is configured, it SHALL match exactly
- **AND** when `expectedPromptVersions` is configured, each configured prompt name SHALL appear with the expected version
#### Scenario: Gatekeeper rule metadata assertions are enforced
- **WHEN** an eval case sets `requireGatekeeperRules=true`
- **THEN** the evaluator SHALL require `verifier_evaluation.gatekeeper_result.rules` to be non-empty
- **AND** each rule item SHALL include `id`, `enabled`, and `default_severity`
#### Scenario: expanded baseline remains passing
- **WHEN** the committed fixture set is evaluated
- **THEN** every case SHALL pass
- **AND** baseline JSON and Markdown reports SHALL reflect the expanded case count and verdict distribution
@@ -0,0 +1,22 @@
## MODIFIED Requirements
### Requirement: MVP demo SHALL be reproducible for interviews
The MVP demo SHALL provide a repeatable way to show a diagnosis answer, trace, verifier evaluation, and feedback.
#### Scenario: interview demo check script records an evidence bundle
- **WHEN** the user runs the interview demo check script against a running `mvp-demo` service
- **THEN** the script SHALL submit a fixed Chat diagnosis request
- **AND** it SHALL fetch the trace for the same session id
- **AND** it SHALL submit useful feedback for that session
- **AND** it SHALL write chat, trace, feedback, and summary outputs under `mvp/demo/output/`
#### Scenario: interview demo check fails with actionable readiness output
- **WHEN** the target service is not reachable
- **THEN** the script SHALL fail before issuing diagnosis requests
- **AND** the failure message SHALL name the base URL and the expected startup profile
#### Scenario: interview documentation explains audit fields
- **WHEN** an interviewer asks how prompt or Gatekeeper changes are audited
- **THEN** the demo documentation SHALL point to `prompt_audit.version` and `gatekeeper_result.rule_set_version`
- **AND** it SHALL explain that deterministic eval fixtures are the regression source of truth