# Tasks: chat-verifier-agent ## 1. Verifier Prompt - [x] 1.1 Create `src/main/resources/prompts/chat-verifier-prompt.md`. - [x] 1.2 Define fixed fact classifications: `direct_evidence`, `indirect_support`, `no_evidence`, `contradicted`. - [x] 1.3 Define critical fact scope, verdict matrix, and `groundedness_score` mapping. - [x] 1.4 Define strict JSON output schema: `verdict`, `groundedness_score`, `critical_fact_count`, `facts_checked`, `rationale`. - [x] 1.5 Forbid Markdown, code fences, schema-extra fields, and text outside the JSON object. - [x] 1.6 Require `facts_checked[*].evidence_refs` for traceability to tool evidence. ## 2. Verifier Input Hook - [x] 2.1 Add `VerifierInputHook.java` as a `MessagesModelHook` running at `BEFORE_MODEL`. - [x] 2.2 Replace raw verifier history with explicit payload fields: `original_query`, `executor_final_answer`, `tool_trace_summary`, `retry_context`. - [x] 2.3 Persist the current round `tool_trace_summary` in `VerifierContextHolder` for later verifier evaluation storage. ## 3. ChatService Integration - [x] 3.1 Load `chatVerifierPrompt` and add `buildChatVerifierAgent()`. - [x] 3.2 Add configurable `verifier.low-confidence-threshold`. - [x] 3.3 Implement explicit per-round orchestration in `ChatService`: planner call, executor call, verifier call. - [x] 3.4 Keep max two outer rounds and inject `retry_context` only for the second round. - [x] 3.5 Parse verifier JSON directly and fall back to `LOW_CONFID` when verifier output is missing or invalid. - [x] 3.6 Keep `SupervisorAgent` construction as legacy residue only; runtime orchestration no longer depends on prompt-only supervisor sequencing. ## 4. Verdict Routing And User Output - [x] 4.1 Route `PASS` to the executor answer. - [x] 4.2 Route `LOW_CONFID` to a fixed disclaimer plus executor answer. - [x] 4.3 Route `REJECT` to degraded output without passing through the raw unverified answer. - [x] 4.4 Build LOW_CONFID gap lists only from verifier-identified gaps. - [x] 4.5 Build DEGRADED confirmed facts, gaps, and next-step suggestions from verifier facts and trace summary. ## 5. Trace Summary And Observability - [x] 5.1 Add `ToolTraceSummaryService` to build verifier evidence summaries from `tool_invocation`. - [x] 5.2 Include only evidence-bearing tools by default. - [x] 5.3 Compact repeated calls by tool and topic domain. - [x] 5.4 Preserve `source_invocation_ids`, `trace_ref`, query samples, retrieval layers, relevance levels, and source document labels. - [x] 5.5 Parse and persist `facts_checked[*].evidence_refs`. - [x] 5.6 Persist `verifier_evaluation.tool_trace_summary` and `traceability_version`. - [x] 5.7 Store concise verifier summaries in `agent_step.thought` while preserving fuller verifier output in `model_output` / `self_evaluation`. ## 6. self_evaluation Merge Semantics - [x] 6.1 Add `SelfEvaluationMergeService`. - [x] 6.2 Write verifier results under `verifier_evaluation`. - [x] 6.3 Write rule scoring under `rule_evaluation`. - [x] 6.4 Preserve the other channel with read-modify-write semantics. ## 7. Verification - [x] 7.1 Compile verification: `mvn -q -DskipTests compile`. - [x] 7.2 Runtime verification: `/api/chat` complex request reached `planner -> executor -> verifier`. - [x] 7.3 Runtime verification: session `9138f064` persisted `verifier_evaluation.facts_checked[*].evidence_refs`. - [x] 7.4 Runtime verification: session `9138f064` persisted `tool_trace_summary[*].source_invocation_ids`. - [x] 7.5 Runtime verification: LOW_CONFID user output included disclaimer and verifier-derived gaps.