Files

59 lines
3.4 KiB
Markdown

# Acceptance: chat-verifier-agent
## Classification
standard
## Task Status
| Task | Status | Notes |
| --- | --- | --- |
| Verifier prompt | Done | Strict JSON schema, verdict matrix, fact classifications, and `evidence_refs` are defined. |
| VerifierInputHook | Done | Explicit verifier payload replaces raw conversation history. |
| ChatService integration | Done | Planner, executor, and verifier are called explicitly with max two rounds. |
| Verdict routing | Done | PASS, LOW_CONFID, and REJECT paths are handled in code. |
| Trace summary | Done | Evidence summaries include `trace_ref` and `source_invocation_ids`. |
| self_evaluation merge | Done | `rule_evaluation` and `verifier_evaluation` are preserved independently. |
| Verifier observability | Done | `verifier_evaluation` persists facts, evidence refs, trace summary, rationale, score, and round. |
## Static Verification
- [x] OpenSpec artifacts exist: `proposal.md`, `design.md`, `specs/chat-verifier-agent/spec.md`, `tasks.md`, `.committed`.
- [x] `change.json` exists and has `metadata.status = committed`.
- [x] `.archive-ready` exists.
- [x] devflow archive-prep files exist: `brief.md`, `evidence.md`, `decisions.md`, `acceptance.md`.
- [x] `devflow/index.md` contains `chat-verifier-agent` with status `archived`.
## Script Verification
- [x] `mvn -q -DskipTests compile` passed.
## Runtime Verification
- [x] POST `/api/chat` with a complex question returned successfully.
- [x] Runtime session `9138f064` showed planner, executor, and verifier execution in logs.
- [x] Runtime session `9138f064` wrote `verifier_evaluation.verdict = LOW_CONFID`.
- [x] Runtime session `9138f064` wrote `facts_checked[*].evidence_refs`.
- [x] Runtime session `9138f064` wrote `tool_trace_summary[*].source_invocation_ids`.
- [x] LOW_CONFID final answer included disclaimer and verifier-derived evidence gaps.
## Unverified
| Scenario | Reason | Risk | Follow-up |
| --- | --- | --- | --- |
| PASS runtime path | The exercised complex runtime case produced LOW_CONFID. | Low; PASS routing is simple pass-through after parsed verifier decision. | Add a fixture or deterministic verifier test if this becomes product-critical. |
| REJECT runtime path | No forced contradiction case was run after traceability changes. | Medium; REJECT is the safety-critical degraded path. | Add a targeted test with a fabricated claim and evidence contradiction. |
| Document-path-level evidence mapping | Current implementation records invocation ids and source document labels, not guaranteed canonical document paths for every retrieval mode. | Low for current audit need; medium for future UI drill-down. | Extend retrieval details with canonical document paths in a later change. |
## Remaining Risks
1. Verifier output still depends on model compliance with JSON schema; code falls back to LOW_CONFID on missing or invalid output.
2. `AgentLoggingHook` is shared by several agent paths; current changes preserve compile and runtime behavior but should be watched in AiOps flows.
3. `SupervisorAgent` construction remains as legacy residue in `ChatService`; runtime orchestration is explicit, but a later cleanup should remove unused supervisor construction.
## Archive State
- [x] OpenSpec change is archive-ready.
- [x] OpenSpec change has been moved to `openspec/changes/archive/2026-07-03-chat-verifier-agent/`.
- [x] Main spec exists at `openspec/specs/chat-verifier-agent/spec.md`.