3.5 KiB
3.5 KiB
Tasks: chat-verifier-agent
1. Verifier Prompt
- 1.1 Create
src/main/resources/prompts/chat-verifier-prompt.md. - 1.2 Define fixed fact classifications:
direct_evidence,indirect_support,no_evidence,contradicted. - 1.3 Define critical fact scope, verdict matrix, and
groundedness_scoremapping. - 1.4 Define strict JSON output schema:
verdict,groundedness_score,critical_fact_count,facts_checked,rationale. - 1.5 Forbid Markdown, code fences, schema-extra fields, and text outside the JSON object.
- 1.6 Require
facts_checked[*].evidence_refsfor traceability to tool evidence.
2. Verifier Input Hook
- 2.1 Add
VerifierInputHook.javaas aMessagesModelHookrunning atBEFORE_MODEL. - 2.2 Replace raw verifier history with explicit payload fields:
original_query,executor_final_answer,tool_trace_summary,retry_context. - 2.3 Persist the current round
tool_trace_summaryinVerifierContextHolderfor later verifier evaluation storage.
3. ChatService Integration
- 3.1 Load
chatVerifierPromptand addbuildChatVerifierAgent(). - 3.2 Add configurable
verifier.low-confidence-threshold. - 3.3 Implement explicit per-round orchestration in
ChatService: planner call, executor call, verifier call. - 3.4 Keep max two outer rounds and inject
retry_contextonly for the second round. - 3.5 Parse verifier JSON directly and fall back to
LOW_CONFIDwhen verifier output is missing or invalid. - 3.6 Keep
SupervisorAgentconstruction as legacy residue only; runtime orchestration no longer depends on prompt-only supervisor sequencing.
4. Verdict Routing And User Output
- 4.1 Route
PASSto the executor answer. - 4.2 Route
LOW_CONFIDto a fixed disclaimer plus executor answer. - 4.3 Route
REJECTto degraded output without passing through the raw unverified answer. - 4.4 Build LOW_CONFID gap lists only from verifier-identified gaps.
- 4.5 Build DEGRADED confirmed facts, gaps, and next-step suggestions from verifier facts and trace summary.
5. Trace Summary And Observability
- 5.1 Add
ToolTraceSummaryServiceto build verifier evidence summaries fromtool_invocation. - 5.2 Include only evidence-bearing tools by default.
- 5.3 Compact repeated calls by tool and topic domain.
- 5.4 Preserve
source_invocation_ids,trace_ref, query samples, retrieval layers, relevance levels, and source document labels. - 5.5 Parse and persist
facts_checked[*].evidence_refs. - 5.6 Persist
verifier_evaluation.tool_trace_summaryandtraceability_version. - 5.7 Store concise verifier summaries in
agent_step.thoughtwhile preserving fuller verifier output inmodel_output/self_evaluation.
6. self_evaluation Merge Semantics
- 6.1 Add
SelfEvaluationMergeService. - 6.2 Write verifier results under
verifier_evaluation. - 6.3 Write rule scoring under
rule_evaluation. - 6.4 Preserve the other channel with read-modify-write semantics.
7. Verification
- 7.1 Compile verification:
mvn -q -DskipTests compile. - 7.2 Runtime verification:
/api/chatcomplex request reachedplanner -> executor -> verifier. - 7.3 Runtime verification: session
9138f064persistedverifier_evaluation.facts_checked[*].evidence_refs. - 7.4 Runtime verification: session
9138f064persistedtool_trace_summary[*].source_invocation_ids. - 7.5 Runtime verification: LOW_CONFID user output included disclaimer and verifier-derived gaps.