feat: add chat verifier agent
This commit is contained in:
@@ -0,0 +1,58 @@
|
||||
# Tasks: chat-verifier-agent
|
||||
|
||||
## 1. Verifier Prompt
|
||||
|
||||
- [x] 1.1 Create `src/main/resources/prompts/chat-verifier-prompt.md`.
|
||||
- [x] 1.2 Define fixed fact classifications: `direct_evidence`, `indirect_support`, `no_evidence`, `contradicted`.
|
||||
- [x] 1.3 Define critical fact scope, verdict matrix, and `groundedness_score` mapping.
|
||||
- [x] 1.4 Define strict JSON output schema: `verdict`, `groundedness_score`, `critical_fact_count`, `facts_checked`, `rationale`.
|
||||
- [x] 1.5 Forbid Markdown, code fences, schema-extra fields, and text outside the JSON object.
|
||||
- [x] 1.6 Require `facts_checked[*].evidence_refs` for traceability to tool evidence.
|
||||
|
||||
## 2. Verifier Input Hook
|
||||
|
||||
- [x] 2.1 Add `VerifierInputHook.java` as a `MessagesModelHook` running at `BEFORE_MODEL`.
|
||||
- [x] 2.2 Replace raw verifier history with explicit payload fields: `original_query`, `executor_final_answer`, `tool_trace_summary`, `retry_context`.
|
||||
- [x] 2.3 Persist the current round `tool_trace_summary` in `VerifierContextHolder` for later verifier evaluation storage.
|
||||
|
||||
## 3. ChatService Integration
|
||||
|
||||
- [x] 3.1 Load `chatVerifierPrompt` and add `buildChatVerifierAgent()`.
|
||||
- [x] 3.2 Add configurable `verifier.low-confidence-threshold`.
|
||||
- [x] 3.3 Implement explicit per-round orchestration in `ChatService`: planner call, executor call, verifier call.
|
||||
- [x] 3.4 Keep max two outer rounds and inject `retry_context` only for the second round.
|
||||
- [x] 3.5 Parse verifier JSON directly and fall back to `LOW_CONFID` when verifier output is missing or invalid.
|
||||
- [x] 3.6 Keep `SupervisorAgent` construction as legacy residue only; runtime orchestration no longer depends on prompt-only supervisor sequencing.
|
||||
|
||||
## 4. Verdict Routing And User Output
|
||||
|
||||
- [x] 4.1 Route `PASS` to the executor answer.
|
||||
- [x] 4.2 Route `LOW_CONFID` to a fixed disclaimer plus executor answer.
|
||||
- [x] 4.3 Route `REJECT` to degraded output without passing through the raw unverified answer.
|
||||
- [x] 4.4 Build LOW_CONFID gap lists only from verifier-identified gaps.
|
||||
- [x] 4.5 Build DEGRADED confirmed facts, gaps, and next-step suggestions from verifier facts and trace summary.
|
||||
|
||||
## 5. Trace Summary And Observability
|
||||
|
||||
- [x] 5.1 Add `ToolTraceSummaryService` to build verifier evidence summaries from `tool_invocation`.
|
||||
- [x] 5.2 Include only evidence-bearing tools by default.
|
||||
- [x] 5.3 Compact repeated calls by tool and topic domain.
|
||||
- [x] 5.4 Preserve `source_invocation_ids`, `trace_ref`, query samples, retrieval layers, relevance levels, and source document labels.
|
||||
- [x] 5.5 Parse and persist `facts_checked[*].evidence_refs`.
|
||||
- [x] 5.6 Persist `verifier_evaluation.tool_trace_summary` and `traceability_version`.
|
||||
- [x] 5.7 Store concise verifier summaries in `agent_step.thought` while preserving fuller verifier output in `model_output` / `self_evaluation`.
|
||||
|
||||
## 6. self_evaluation Merge Semantics
|
||||
|
||||
- [x] 6.1 Add `SelfEvaluationMergeService`.
|
||||
- [x] 6.2 Write verifier results under `verifier_evaluation`.
|
||||
- [x] 6.3 Write rule scoring under `rule_evaluation`.
|
||||
- [x] 6.4 Preserve the other channel with read-modify-write semantics.
|
||||
|
||||
## 7. Verification
|
||||
|
||||
- [x] 7.1 Compile verification: `mvn -q -DskipTests compile`.
|
||||
- [x] 7.2 Runtime verification: `/api/chat` complex request reached `planner -> executor -> verifier`.
|
||||
- [x] 7.3 Runtime verification: session `9138f064` persisted `verifier_evaluation.facts_checked[*].evidence_refs`.
|
||||
- [x] 7.4 Runtime verification: session `9138f064` persisted `tool_trace_summary[*].source_invocation_ids`.
|
||||
- [x] 7.5 Runtime verification: LOW_CONFID user output included disclaimer and verifier-derived gaps.
|
||||
Reference in New Issue
Block a user