docs(openspec): propose interview demo quality audit
This commit is contained in:
@@ -0,0 +1,10 @@
|
|||||||
|
# Acceptance
|
||||||
|
|
||||||
|
Acceptance will be filled during Apply/Archive with concrete command output.
|
||||||
|
|
||||||
|
Planned checks:
|
||||||
|
|
||||||
|
- `openspec validate interview-demo-quality-audit --strict` - passed during commit gate.
|
||||||
|
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ChatServiceSequentialAgentTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`
|
||||||
|
- `mvn -q -DskipTests compile`
|
||||||
|
- `powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1` against `mvp-demo` service if dependencies are available.
|
||||||
@@ -0,0 +1,23 @@
|
|||||||
|
# Interview Demo Quality Audit Brief
|
||||||
|
|
||||||
|
## Background
|
||||||
|
|
||||||
|
The MVP already demonstrates traceable Agent diagnosis with Planner, Executor, Gatekeeper, Verifier, Composer, evidence tools, trace persistence, and deterministic eval fixtures. The remaining interview-readiness gap is not a new Agent architecture; it is making the demo easier to run and making prompt/rule changes easier to audit.
|
||||||
|
|
||||||
|
## Goal
|
||||||
|
|
||||||
|
Stabilize the interview demo path, expand fixture-backed evaluation, and persist prompt/Gatekeeper audit metadata so the project can explain and verify Agent behavior during interviews.
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
- Add prompt audit metadata to Chat verifier evaluation.
|
||||||
|
- Extend deterministic eval cases and baseline reports.
|
||||||
|
- Add an interview demo preflight/check script.
|
||||||
|
- Update MVP demo and architecture documentation.
|
||||||
|
|
||||||
|
## Non-goals
|
||||||
|
|
||||||
|
- No public API or database schema changes.
|
||||||
|
- No new SubAgent split, MCP migration, process isolation, or AIOps LLM Verifier.
|
||||||
|
- No guarantee that every live LLM run returns PASS.
|
||||||
|
|
||||||
@@ -0,0 +1,91 @@
|
|||||||
|
# interview-demo-quality-audit Decisions
|
||||||
|
|
||||||
|
## Clarify
|
||||||
|
|
||||||
|
- Entry summary: stabilize the interview demo, expand deterministic eval coverage, and add Prompt/Gatekeeper version audit.
|
||||||
|
- Slug: `interview-demo-quality-audit`.
|
||||||
|
- Devflow scale: `standard`.
|
||||||
|
- Interface impact: L2 internal contract change because `verifier_evaluation` gains `prompt_audit`; no public HTTP API or database schema change.
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
- `devflow/index.md` used: related entries found for `diagnosis-eval-demo-gatekeeper-closure`, `executor-composer-final-answer`, `verifier-evidence-reference-fidelity`, `mvp-demo-interview-runbook`, and `diagnosis-eval-baseline-diff`.
|
||||||
|
- Relevant glossary:
|
||||||
|
- Evidence Tools produce incident facts and must be recorded in `tool_invocation`.
|
||||||
|
- Verifier should not use skills/runbooks as incident evidence.
|
||||||
|
- Diagnosis Playbook Skill is workflow guidance, not a fact source.
|
||||||
|
- Historical constraints that must enter OpenSpec:
|
||||||
|
- Diagnosis eval is deterministic and fixture-backed; no LLM-as-judge.
|
||||||
|
- Stable demo scenarios are documentation/payloads plus deterministic fixtures; live E2E is a compatibility check, not a guaranteed PASS oracle.
|
||||||
|
- Gatekeeper rule metadata is already metadata-only and should not become dynamic rule execution.
|
||||||
|
- Composer is the final expression layer and must not leak raw Executor JSON.
|
||||||
|
|
||||||
|
## Question Pool
|
||||||
|
|
||||||
|
| ID | Dimension | Mode | Question | Status |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| Q1 | Terminology | evidence-driven | What should the new audit metadata be called? | Resolved |
|
||||||
|
| Q2 | Boundary | evidence-driven | Does this require public API or schema changes? | Resolved |
|
||||||
|
| Q3 | Acceptance | evidence-driven | Which current assets define deterministic acceptance? | Resolved |
|
||||||
|
| Q4 | Technical | evidence-driven | Where should prompt version metadata live with minimal implementation risk? | Resolved |
|
||||||
|
| Q5 | Scope | user-interview | Should live E2E be mandatory for all scenarios? | Confirmed by objective as conditional |
|
||||||
|
|
||||||
|
## Evidence-driven Conclusions
|
||||||
|
|
||||||
|
- Q1 conclusion: use `prompt_audit` for prompt version metadata and keep existing `gatekeeper_result.rule_set_version`.
|
||||||
|
- Q2 conclusion: keep this as an internal trace/self-evaluation contract change. Do not add endpoints, tables, or new Agent roles.
|
||||||
|
- Q3 conclusion: `DiagnosisTraceEvaluatorTest`, baseline reports, fixed fixtures, and demo scripts define current acceptance style.
|
||||||
|
- Q4 conclusion: add a small Chat prompt audit catalog near `ChatService` prompt loading and persist a compact snapshot with verifier evaluation.
|
||||||
|
- Q5 conclusion: run live E2E with `mvp-demo` profile if dependencies are available; otherwise record the blocker and rely on deterministic eval/unit evidence.
|
||||||
|
|
||||||
|
## Specify / Alignment
|
||||||
|
|
||||||
|
| Check | Status | Notes |
|
||||||
|
|---|---|---|
|
||||||
|
| proposal goals -> proposal | Aligned | Proposal covers demo preflight, eval expansion, prompt audit, and docs. |
|
||||||
|
| proposal scope/constraints -> design | Aligned | Design records no public API/schema changes, prompt audit shape, eval fields, and demo script behavior. |
|
||||||
|
| design decisions -> specs/tasks | Aligned | Specs cover persisted prompt audit, evaluator checks, baseline, and demo script outputs. |
|
||||||
|
| specs observable behavior -> tasks | Aligned | Each requirement has implementation and verification tasks. |
|
||||||
|
|
||||||
|
## Audit
|
||||||
|
|
||||||
|
Input -> processing -> output chain:
|
||||||
|
|
||||||
|
```text
|
||||||
|
prompt resource metadata
|
||||||
|
-> ChatService / PromptAudit snapshot
|
||||||
|
-> verifier_evaluation.prompt_audit
|
||||||
|
-> Trace API / eval fixtures
|
||||||
|
-> DiagnosisTraceEvaluator baseline
|
||||||
|
|
||||||
|
run-interview-demo-check.ps1
|
||||||
|
-> service readiness
|
||||||
|
-> chat / trace / feedback
|
||||||
|
-> mvp/demo/output summary
|
||||||
|
```
|
||||||
|
|
||||||
|
Architecture risk assessment:
|
||||||
|
|
||||||
|
1. The audit shape is intentionally compact and internal; storing full prompt text would create noisy traces and possible sensitive-content risk.
|
||||||
|
2. Eval should assert versions by explicit metadata, not by prompt content hashes that churn during local prompt edits.
|
||||||
|
3. Live demo checks may still be LOW_CONFID because LLM output is not deterministic; deterministic fixtures remain the regression source of truth.
|
||||||
|
4. No devflow/OpenSpec conflict found.
|
||||||
|
|
||||||
|
## Commit Gate
|
||||||
|
|
||||||
|
- `openspec validate interview-demo-quality-audit --strict`: passed.
|
||||||
|
- File completeness:
|
||||||
|
- `proposal.md`: present.
|
||||||
|
- `design.md`: present.
|
||||||
|
- `specs/`: present for `chat-verifier-agent`, `diagnosis-eval-harness`, and `mvp-demo-trace-acceptance`.
|
||||||
|
- `tasks.md`: present.
|
||||||
|
- Consistency:
|
||||||
|
- Proposal goals map to design sections.
|
||||||
|
- Design decisions map to spec requirements and executable tasks.
|
||||||
|
- Task acceptance checks are verifiable.
|
||||||
|
- `.committed` marker created.
|
||||||
|
|
||||||
|
## Current Checkpoint
|
||||||
|
|
||||||
|
- Commit completed.
|
||||||
|
- Apply is authorized by the original objective: "完成后归档提交".
|
||||||
@@ -0,0 +1,25 @@
|
|||||||
|
# Evidence
|
||||||
|
|
||||||
|
## Context Files Read
|
||||||
|
|
||||||
|
- `devflow/index.md`
|
||||||
|
- `devflow/glossary/CONTEXT.md`
|
||||||
|
- `devflow/projects/2026-07-08-diagnosis-eval-demo-gatekeeper-closure/decisions.md`
|
||||||
|
- `devflow/projects/2026-07-08-executor-composer-final-answer/decisions.md`
|
||||||
|
- `mvp/architecture/current-mvp-architecture.md`
|
||||||
|
- `mvp/architecture/agent-orchestration.md`
|
||||||
|
- `mvp/architecture/executor-evidence-pipeline-refactor.md`
|
||||||
|
- `mvp/architecture/harness-quality-gates.md`
|
||||||
|
- `mvp/demo/README.md`
|
||||||
|
- `mvp/demo/ten-minute-interview-demo.md`
|
||||||
|
- `mvp/eval/README.md`
|
||||||
|
- `mvp/eval/cases/diagnosis-cases.json`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/ExecutorGatekeeperService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/eval/DiagnosisTraceEvaluator.java`
|
||||||
|
- `src/main/resources/gatekeeper/gatekeeper-rules.json`
|
||||||
|
|
||||||
|
## Tooling Note
|
||||||
|
|
||||||
|
The required `codebase-retrieval` and LSP tools were not exposed in this session. Impact analysis used `rg`, direct file reads, existing OpenSpec/devflow artifacts, and targeted tests instead.
|
||||||
|
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
committed
|
||||||
@@ -0,0 +1,117 @@
|
|||||||
|
# Design: Interview Demo Quality Audit
|
||||||
|
|
||||||
|
## Overview
|
||||||
|
|
||||||
|
The change adds auditability and demo readiness without changing Agent routing.
|
||||||
|
|
||||||
|
```text
|
||||||
|
ChatService prompt resources
|
||||||
|
-> PromptAuditService
|
||||||
|
-> verifier_evaluation.prompt_audit
|
||||||
|
-> Trace API
|
||||||
|
-> DiagnosisTraceEvaluator
|
||||||
|
-> baseline report
|
||||||
|
|
||||||
|
mvp/demo/scripts/run-interview-demo-check.ps1
|
||||||
|
-> health/readiness check
|
||||||
|
-> payment timeout chat
|
||||||
|
-> trace fetch
|
||||||
|
-> feedback
|
||||||
|
-> demo output bundle
|
||||||
|
```
|
||||||
|
|
||||||
|
## Prompt Audit
|
||||||
|
|
||||||
|
Add a compact `prompt_audit` object under:
|
||||||
|
|
||||||
|
```text
|
||||||
|
diagnosis_session.self_evaluation.verifier_evaluation.prompt_audit
|
||||||
|
```
|
||||||
|
|
||||||
|
Shape:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"version": "chat-prompts-v1",
|
||||||
|
"prompts": [
|
||||||
|
{
|
||||||
|
"name": "chat_planner",
|
||||||
|
"version": "chat-planner-v1",
|
||||||
|
"resource": "prompts/chat-planner-prompt.md"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Design choices:
|
||||||
|
|
||||||
|
- Use explicit local metadata, not full prompt hashes, because the interview goal is explainable version audit rather than cryptographic integrity.
|
||||||
|
- Keep this metadata in code or a small resource catalog near prompt loading.
|
||||||
|
- Persist prompt audit with every Chat verifier evaluation, including fallback/degraded paths.
|
||||||
|
- Do not include full prompt text in trace.
|
||||||
|
|
||||||
|
## Eval Expansion
|
||||||
|
|
||||||
|
Extend `DiagnosisEvalCase` with optional fields:
|
||||||
|
|
||||||
|
- `requirePromptAudit`
|
||||||
|
- `expectedPromptAuditVersion`
|
||||||
|
- `expectedPromptVersions`
|
||||||
|
- `requireGatekeeperRules`
|
||||||
|
|
||||||
|
Evaluator behavior:
|
||||||
|
|
||||||
|
- If `requirePromptAudit=true`, `verifier_evaluation.prompt_audit.version` must exist.
|
||||||
|
- If `expectedPromptAuditVersion` is set, it must match.
|
||||||
|
- If `expectedPromptVersions` is set, each listed prompt name/version pair must exist.
|
||||||
|
- If `requireGatekeeperRules=true`, `gatekeeper_result.rules` must be a non-empty list and each item must include `id`, `enabled`, and `default_severity`.
|
||||||
|
|
||||||
|
Add at least two fixture-backed cases:
|
||||||
|
|
||||||
|
- A positive audit closure case that requires prompt audit + Gatekeeper rules.
|
||||||
|
- A metadata-gap negative case represented as `LOW_CONFID`/safe final answer, used to prove the evaluator catches missing audit metadata when configured.
|
||||||
|
|
||||||
|
The baseline must remain fully passing after fixtures are updated.
|
||||||
|
|
||||||
|
## Demo Stabilization
|
||||||
|
|
||||||
|
Add a PowerShell script:
|
||||||
|
|
||||||
|
```text
|
||||||
|
mvp/demo/scripts/run-interview-demo-check.ps1
|
||||||
|
```
|
||||||
|
|
||||||
|
Responsibilities:
|
||||||
|
|
||||||
|
- Accept base URL and session id parameters.
|
||||||
|
- Check that the service is reachable.
|
||||||
|
- Run the existing payment-timeout chat request.
|
||||||
|
- Fetch trace for the same session id.
|
||||||
|
- Submit useful feedback.
|
||||||
|
- Write outputs under `mvp/demo/output/`.
|
||||||
|
- Emit a concise summary with session id, verdict, Gatekeeper rule version, prompt audit version, and output paths.
|
||||||
|
|
||||||
|
The script should fail fast with actionable messages when the service is unavailable.
|
||||||
|
|
||||||
|
## Documentation
|
||||||
|
|
||||||
|
Add/update:
|
||||||
|
|
||||||
|
- `mvp/demo/README.md`: mention the preflight script.
|
||||||
|
- `mvp/demo/ten-minute-interview-demo.md`: use the preflight script as the recommended path.
|
||||||
|
- `mvp/demo/interview-q-and-a.md`: concise interview answers for Agent engineering tradeoffs.
|
||||||
|
- `mvp/architecture/harness-quality-gates.md`: record prompt audit as part of the quality gate.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
Required:
|
||||||
|
|
||||||
|
- Targeted unit/eval tests for prompt audit persistence and evaluator checks.
|
||||||
|
- Regenerated baseline JSON/Markdown reports.
|
||||||
|
- OpenSpec validation.
|
||||||
|
|
||||||
|
E2E:
|
||||||
|
|
||||||
|
- If local dependencies are available, run Spring Boot with `mvp-demo` profile and execute the new preflight script.
|
||||||
|
- If unavailable, record the reason and rely on deterministic unit/eval evidence.
|
||||||
|
|
||||||
@@ -0,0 +1,58 @@
|
|||||||
|
# Change: Interview Demo Quality Audit
|
||||||
|
|
||||||
|
## Problem
|
||||||
|
|
||||||
|
SuperBizAgent MVP is now strong enough to demonstrate traceable Agent engineering, but the interview path still has three gaps:
|
||||||
|
|
||||||
|
- The live demo has a run script, but no preflight command that checks service readiness and produces a concise interview evidence bundle.
|
||||||
|
- The diagnosis eval baseline covers the V2 evidence pipeline, but it does not yet assert prompt-version audit data and has limited coverage for audit metadata gaps.
|
||||||
|
- Gatekeeper already exposes `rule_set_version`, but prompt versions are not persisted with the verifier evaluation, making prompt changes harder to explain, compare, and roll back in an interview.
|
||||||
|
|
||||||
|
This change stabilizes the MVP as an interview artifact rather than adding a new diagnosis architecture.
|
||||||
|
|
||||||
|
## Proposed Solution
|
||||||
|
|
||||||
|
Implement a small internal quality/audit increment:
|
||||||
|
|
||||||
|
1. Add prompt version audit metadata to Chat verifier evaluation.
|
||||||
|
2. Extend deterministic diagnosis eval cases/fixtures to assert prompt audit metadata and audit-metadata failures.
|
||||||
|
3. Add an interview demo preflight script and documentation that can be run before or during a demo to verify service readiness, execute the payment timeout path, fetch trace, and record key audit fields.
|
||||||
|
4. Add/update MVP documentation for interview Q&A and the new audit/preflight workflow.
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
In scope:
|
||||||
|
|
||||||
|
- Internal `diagnosis_session.self_evaluation.verifier_evaluation` audit JSON.
|
||||||
|
- Diagnosis eval case schema, evaluator checks, fixtures, and baseline reports.
|
||||||
|
- MVP demo scripts/docs.
|
||||||
|
- Architecture/demo documentation for prompt and Gatekeeper version audit.
|
||||||
|
|
||||||
|
Out of scope:
|
||||||
|
|
||||||
|
- Public HTTP API changes.
|
||||||
|
- Database schema changes.
|
||||||
|
- New Agent roles, MCP tool server migration, process isolation, or AIOps LLM Verifier.
|
||||||
|
- Replacing existing `Planner -> Executor -> Gatekeeper -> Verifier -> Composer` orchestration.
|
||||||
|
- Guaranteeing live LLM `PASS` for every demo run. Live demo compatibility and deterministic fixture regression are separate acceptance paths.
|
||||||
|
|
||||||
|
## Context Constraints From devflow
|
||||||
|
|
||||||
|
- Evidence Tools produce incident facts and must be recorded in `tool_invocation`.
|
||||||
|
- Chat quality gates are layered: Gatekeeper verifies evidence references, Verifier judges derivability, Composer controls expression.
|
||||||
|
- `diagnosis eval` is deterministic and fixture-backed; no LLM-as-judge.
|
||||||
|
- Demo assets should be runnable, but interview safety should not depend solely on live LLM behavior.
|
||||||
|
- Gatekeeper rule metadata is metadata-only; dynamic rule execution is out of scope.
|
||||||
|
|
||||||
|
## Interface Impact
|
||||||
|
|
||||||
|
Level: L2 internal contract change.
|
||||||
|
|
||||||
|
Reason: `verifier_evaluation` gains a compact `prompt_audit` object. Existing public API shape remains the same, and the value is exposed only through already-existing trace/self-evaluation JSON.
|
||||||
|
|
||||||
|
## Risks
|
||||||
|
|
||||||
|
- Baseline report churn is expected when adding cases; JSON and Markdown reports must be regenerated together.
|
||||||
|
- Prompt audit must be deterministic and stable enough for eval fixtures; avoid hashing full prompt text with environment-specific content.
|
||||||
|
- Demo preflight must not hardcode secrets and must tolerate local service unavailability with clear failure messages.
|
||||||
|
|
||||||
@@ -0,0 +1,16 @@
|
|||||||
|
## MODIFIED Requirements
|
||||||
|
|
||||||
|
### Requirement: Verifier SHALL be observable
|
||||||
|
The Verifier's verdict SHALL be persisted for observability.
|
||||||
|
|
||||||
|
#### Scenario: prompt audit written to verifier evaluation
|
||||||
|
- **WHEN** the Chat verifier evaluation is persisted
|
||||||
|
- **THEN** the system SHALL include a `prompt_audit` object under `diagnosis_session.self_evaluation.verifier_evaluation`
|
||||||
|
- **AND** `prompt_audit.version` SHALL identify the Chat prompt audit catalog version
|
||||||
|
- **AND** `prompt_audit.prompts` SHALL include the planner, executor, verifier, and composer prompt names and versions
|
||||||
|
- **AND** full prompt text SHALL NOT be persisted in `prompt_audit`
|
||||||
|
|
||||||
|
#### Scenario: prompt audit available on fallback paths
|
||||||
|
- **WHEN** Chat verifier parsing fails, Composer parsing fails, or Chat produces a degraded answer
|
||||||
|
- **THEN** the persisted verifier evaluation SHALL still include `prompt_audit`
|
||||||
|
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
## MODIFIED Requirements
|
||||||
|
|
||||||
|
### Requirement: Diagnosis eval SHALL validate trace fixtures deterministically
|
||||||
|
The diagnosis eval harness SHALL evaluate saved trace fixtures without invoking an LLM judge.
|
||||||
|
|
||||||
|
#### Scenario: prompt audit assertions are enforced
|
||||||
|
- **WHEN** an eval case sets `requirePromptAudit=true`
|
||||||
|
- **THEN** the evaluator SHALL require `verifier_evaluation.prompt_audit.version`
|
||||||
|
- **AND** when `expectedPromptAuditVersion` is configured, it SHALL match exactly
|
||||||
|
- **AND** when `expectedPromptVersions` is configured, each configured prompt name SHALL appear with the expected version
|
||||||
|
|
||||||
|
#### Scenario: Gatekeeper rule metadata assertions are enforced
|
||||||
|
- **WHEN** an eval case sets `requireGatekeeperRules=true`
|
||||||
|
- **THEN** the evaluator SHALL require `verifier_evaluation.gatekeeper_result.rules` to be non-empty
|
||||||
|
- **AND** each rule item SHALL include `id`, `enabled`, and `default_severity`
|
||||||
|
|
||||||
|
#### Scenario: expanded baseline remains passing
|
||||||
|
- **WHEN** the committed fixture set is evaluated
|
||||||
|
- **THEN** every case SHALL pass
|
||||||
|
- **AND** baseline JSON and Markdown reports SHALL reflect the expanded case count and verdict distribution
|
||||||
|
|
||||||
@@ -0,0 +1,22 @@
|
|||||||
|
## MODIFIED Requirements
|
||||||
|
|
||||||
|
### Requirement: MVP demo SHALL be reproducible for interviews
|
||||||
|
The MVP demo SHALL provide a repeatable way to show a diagnosis answer, trace, verifier evaluation, and feedback.
|
||||||
|
|
||||||
|
#### Scenario: interview demo check script records an evidence bundle
|
||||||
|
- **WHEN** the user runs the interview demo check script against a running `mvp-demo` service
|
||||||
|
- **THEN** the script SHALL submit a fixed Chat diagnosis request
|
||||||
|
- **AND** it SHALL fetch the trace for the same session id
|
||||||
|
- **AND** it SHALL submit useful feedback for that session
|
||||||
|
- **AND** it SHALL write chat, trace, feedback, and summary outputs under `mvp/demo/output/`
|
||||||
|
|
||||||
|
#### Scenario: interview demo check fails with actionable readiness output
|
||||||
|
- **WHEN** the target service is not reachable
|
||||||
|
- **THEN** the script SHALL fail before issuing diagnosis requests
|
||||||
|
- **AND** the failure message SHALL name the base URL and the expected startup profile
|
||||||
|
|
||||||
|
#### Scenario: interview documentation explains audit fields
|
||||||
|
- **WHEN** an interviewer asks how prompt or Gatekeeper changes are audited
|
||||||
|
- **THEN** the demo documentation SHALL point to `prompt_audit.version` and `gatekeeper_result.rule_set_version`
|
||||||
|
- **AND** it SHALL explain that deterministic eval fixtures are the regression source of truth
|
||||||
|
|
||||||
@@ -0,0 +1,32 @@
|
|||||||
|
## 1. OpenSpec And devflow
|
||||||
|
|
||||||
|
- [x] 1.1 Create OpenSpec proposal/design/spec/tasks for `interview-demo-quality-audit`.
|
||||||
|
- [x] 1.2 Record context, question pool, interface impact, audit, and verification plan in devflow decisions.
|
||||||
|
- [x] 1.3 Pass OpenSpec validation and create `.committed`.
|
||||||
|
|
||||||
|
## 2. Prompt/Gatekeeper Version Audit
|
||||||
|
|
||||||
|
- [ ] 2.1 Add compact Chat prompt audit metadata for planner, executor, verifier, and composer prompts.
|
||||||
|
- [ ] 2.2 Persist `prompt_audit` under `verifier_evaluation` for Chat verifier/composer outcomes.
|
||||||
|
- [ ] 2.3 Add focused tests proving prompt audit appears in persisted verifier evaluation.
|
||||||
|
- [ ] 2.4 Extend eval checks for prompt audit and Gatekeeper rule metadata.
|
||||||
|
|
||||||
|
## 3. Eval Expansion
|
||||||
|
|
||||||
|
- [ ] 3.1 Extend diagnosis eval case/result schema for prompt audit fields.
|
||||||
|
- [ ] 3.2 Add fixture-backed cases for audit closure coverage.
|
||||||
|
- [ ] 3.3 Regenerate baseline JSON and Markdown reports.
|
||||||
|
- [ ] 3.4 Update eval docs/schema.
|
||||||
|
|
||||||
|
## 4. Interview Demo Stabilization
|
||||||
|
|
||||||
|
- [ ] 4.1 Add `run-interview-demo-check.ps1` with service preflight, chat, trace, feedback, and summary output.
|
||||||
|
- [ ] 4.2 Update demo README and 10-minute script to use the preflight path.
|
||||||
|
- [ ] 4.3 Add interview Q&A documentation focused on Agent engineering tradeoffs.
|
||||||
|
|
||||||
|
## 5. Verification And Archive
|
||||||
|
|
||||||
|
- [ ] 5.1 Run targeted tests for ChatService/prompt audit and diagnosis eval.
|
||||||
|
- [ ] 5.2 Run relevant broader regression tests.
|
||||||
|
- [ ] 5.3 Run E2E demo check with `mvp-demo` profile if dependencies are available; otherwise record the blocker.
|
||||||
|
- [ ] 5.4 Archive the OpenSpec change, update devflow artifacts, and commit implementation + archive.
|
||||||
Reference in New Issue
Block a user