feat: add chat verifier agent

This commit is contained in:
zhuyongxin
2026-07-03 10:54:33 +08:00
parent 4f5316d473
commit 9050487307
28 changed files with 3200 additions and 209 deletions
@@ -0,0 +1,58 @@
# Tasks: chat-verifier-agent
## 1. Verifier Prompt
- [x] 1.1 Create `src/main/resources/prompts/chat-verifier-prompt.md`.
- [x] 1.2 Define fixed fact classifications: `direct_evidence`, `indirect_support`, `no_evidence`, `contradicted`.
- [x] 1.3 Define critical fact scope, verdict matrix, and `groundedness_score` mapping.
- [x] 1.4 Define strict JSON output schema: `verdict`, `groundedness_score`, `critical_fact_count`, `facts_checked`, `rationale`.
- [x] 1.5 Forbid Markdown, code fences, schema-extra fields, and text outside the JSON object.
- [x] 1.6 Require `facts_checked[*].evidence_refs` for traceability to tool evidence.
## 2. Verifier Input Hook
- [x] 2.1 Add `VerifierInputHook.java` as a `MessagesModelHook` running at `BEFORE_MODEL`.
- [x] 2.2 Replace raw verifier history with explicit payload fields: `original_query`, `executor_final_answer`, `tool_trace_summary`, `retry_context`.
- [x] 2.3 Persist the current round `tool_trace_summary` in `VerifierContextHolder` for later verifier evaluation storage.
## 3. ChatService Integration
- [x] 3.1 Load `chatVerifierPrompt` and add `buildChatVerifierAgent()`.
- [x] 3.2 Add configurable `verifier.low-confidence-threshold`.
- [x] 3.3 Implement explicit per-round orchestration in `ChatService`: planner call, executor call, verifier call.
- [x] 3.4 Keep max two outer rounds and inject `retry_context` only for the second round.
- [x] 3.5 Parse verifier JSON directly and fall back to `LOW_CONFID` when verifier output is missing or invalid.
- [x] 3.6 Keep `SupervisorAgent` construction as legacy residue only; runtime orchestration no longer depends on prompt-only supervisor sequencing.
## 4. Verdict Routing And User Output
- [x] 4.1 Route `PASS` to the executor answer.
- [x] 4.2 Route `LOW_CONFID` to a fixed disclaimer plus executor answer.
- [x] 4.3 Route `REJECT` to degraded output without passing through the raw unverified answer.
- [x] 4.4 Build LOW_CONFID gap lists only from verifier-identified gaps.
- [x] 4.5 Build DEGRADED confirmed facts, gaps, and next-step suggestions from verifier facts and trace summary.
## 5. Trace Summary And Observability
- [x] 5.1 Add `ToolTraceSummaryService` to build verifier evidence summaries from `tool_invocation`.
- [x] 5.2 Include only evidence-bearing tools by default.
- [x] 5.3 Compact repeated calls by tool and topic domain.
- [x] 5.4 Preserve `source_invocation_ids`, `trace_ref`, query samples, retrieval layers, relevance levels, and source document labels.
- [x] 5.5 Parse and persist `facts_checked[*].evidence_refs`.
- [x] 5.6 Persist `verifier_evaluation.tool_trace_summary` and `traceability_version`.
- [x] 5.7 Store concise verifier summaries in `agent_step.thought` while preserving fuller verifier output in `model_output` / `self_evaluation`.
## 6. self_evaluation Merge Semantics
- [x] 6.1 Add `SelfEvaluationMergeService`.
- [x] 6.2 Write verifier results under `verifier_evaluation`.
- [x] 6.3 Write rule scoring under `rule_evaluation`.
- [x] 6.4 Preserve the other channel with read-modify-write semantics.
## 7. Verification
- [x] 7.1 Compile verification: `mvn -q -DskipTests compile`.
- [x] 7.2 Runtime verification: `/api/chat` complex request reached `planner -> executor -> verifier`.
- [x] 7.3 Runtime verification: session `9138f064` persisted `verifier_evaluation.facts_checked[*].evidence_refs`.
- [x] 7.4 Runtime verification: session `9138f064` persisted `tool_trace_summary[*].source_invocation_ids`.
- [x] 7.5 Runtime verification: LOW_CONFID user output included disclaimer and verifier-derived gaps.