3.4 KiB
3.4 KiB
Acceptance: chat-verifier-agent
Classification
standard
Task Status
| Task | Status | Notes |
|---|---|---|
| Verifier prompt | Done | Strict JSON schema, verdict matrix, fact classifications, and evidence_refs are defined. |
| VerifierInputHook | Done | Explicit verifier payload replaces raw conversation history. |
| ChatService integration | Done | Planner, executor, and verifier are called explicitly with max two rounds. |
| Verdict routing | Done | PASS, LOW_CONFID, and REJECT paths are handled in code. |
| Trace summary | Done | Evidence summaries include trace_ref and source_invocation_ids. |
| self_evaluation merge | Done | rule_evaluation and verifier_evaluation are preserved independently. |
| Verifier observability | Done | verifier_evaluation persists facts, evidence refs, trace summary, rationale, score, and round. |
Static Verification
- OpenSpec artifacts exist:
proposal.md,design.md,specs/chat-verifier-agent/spec.md,tasks.md,.committed. change.jsonexists and hasmetadata.status = committed..archive-readyexists.- devflow archive-prep files exist:
brief.md,evidence.md,decisions.md,acceptance.md. devflow/index.mdcontainschat-verifier-agentwith statusarchived.
Script Verification
mvn -q -DskipTests compilepassed.
Runtime Verification
- POST
/api/chatwith a complex question returned successfully. - Runtime session
9138f064showed planner, executor, and verifier execution in logs. - Runtime session
9138f064wroteverifier_evaluation.verdict = LOW_CONFID. - Runtime session
9138f064wrotefacts_checked[*].evidence_refs. - Runtime session
9138f064wrotetool_trace_summary[*].source_invocation_ids. - LOW_CONFID final answer included disclaimer and verifier-derived evidence gaps.
Unverified
| Scenario | Reason | Risk | Follow-up |
|---|---|---|---|
| PASS runtime path | The exercised complex runtime case produced LOW_CONFID. | Low; PASS routing is simple pass-through after parsed verifier decision. | Add a fixture or deterministic verifier test if this becomes product-critical. |
| REJECT runtime path | No forced contradiction case was run after traceability changes. | Medium; REJECT is the safety-critical degraded path. | Add a targeted test with a fabricated claim and evidence contradiction. |
| Document-path-level evidence mapping | Current implementation records invocation ids and source document labels, not guaranteed canonical document paths for every retrieval mode. | Low for current audit need; medium for future UI drill-down. | Extend retrieval details with canonical document paths in a later change. |
Remaining Risks
- Verifier output still depends on model compliance with JSON schema; code falls back to LOW_CONFID on missing or invalid output.
AgentLoggingHookis shared by several agent paths; current changes preserve compile and runtime behavior but should be watched in AiOps flows.SupervisorAgentconstruction remains as legacy residue inChatService; runtime orchestration is explicit, but a later cleanup should remove unused supervisor construction.
Archive State
- OpenSpec change is archive-ready.
- OpenSpec change has been moved to
openspec/changes/archive/2026-07-03-chat-verifier-agent/. - Main spec exists at
openspec/specs/chat-verifier-agent/spec.md.