Files
SuperBizAgent-java/devflow/projects/2026-07-02-chat-verifier-agent/acceptance.md
T

3.4 KiB

Acceptance: chat-verifier-agent

Classification

standard

Task Status

Task Status Notes
Verifier prompt Done Strict JSON schema, verdict matrix, fact classifications, and evidence_refs are defined.
VerifierInputHook Done Explicit verifier payload replaces raw conversation history.
ChatService integration Done Planner, executor, and verifier are called explicitly with max two rounds.
Verdict routing Done PASS, LOW_CONFID, and REJECT paths are handled in code.
Trace summary Done Evidence summaries include trace_ref and source_invocation_ids.
self_evaluation merge Done rule_evaluation and verifier_evaluation are preserved independently.
Verifier observability Done verifier_evaluation persists facts, evidence refs, trace summary, rationale, score, and round.

Static Verification

  • OpenSpec artifacts exist: proposal.md, design.md, specs/chat-verifier-agent/spec.md, tasks.md, .committed.
  • change.json exists and has metadata.status = committed.
  • .archive-ready exists.
  • devflow archive-prep files exist: brief.md, evidence.md, decisions.md, acceptance.md.
  • devflow/index.md contains chat-verifier-agent with status archived.

Script Verification

  • mvn -q -DskipTests compile passed.

Runtime Verification

  • POST /api/chat with a complex question returned successfully.
  • Runtime session 9138f064 showed planner, executor, and verifier execution in logs.
  • Runtime session 9138f064 wrote verifier_evaluation.verdict = LOW_CONFID.
  • Runtime session 9138f064 wrote facts_checked[*].evidence_refs.
  • Runtime session 9138f064 wrote tool_trace_summary[*].source_invocation_ids.
  • LOW_CONFID final answer included disclaimer and verifier-derived evidence gaps.

Unverified

Scenario Reason Risk Follow-up
PASS runtime path The exercised complex runtime case produced LOW_CONFID. Low; PASS routing is simple pass-through after parsed verifier decision. Add a fixture or deterministic verifier test if this becomes product-critical.
REJECT runtime path No forced contradiction case was run after traceability changes. Medium; REJECT is the safety-critical degraded path. Add a targeted test with a fabricated claim and evidence contradiction.
Document-path-level evidence mapping Current implementation records invocation ids and source document labels, not guaranteed canonical document paths for every retrieval mode. Low for current audit need; medium for future UI drill-down. Extend retrieval details with canonical document paths in a later change.

Remaining Risks

  1. Verifier output still depends on model compliance with JSON schema; code falls back to LOW_CONFID on missing or invalid output.
  2. AgentLoggingHook is shared by several agent paths; current changes preserve compile and runtime behavior but should be watched in AiOps flows.
  3. SupervisorAgent construction remains as legacy residue in ChatService; runtime orchestration is explicit, but a later cleanup should remove unused supervisor construction.

Archive State

  • OpenSpec change is archive-ready.
  • OpenSpec change has been moved to openspec/changes/archive/2026-07-03-chat-verifier-agent/.
  • Main spec exists at openspec/specs/chat-verifier-agent/spec.md.