Files

3.5 KiB

Tasks: chat-verifier-agent

1. Verifier Prompt

  • 1.1 Create src/main/resources/prompts/chat-verifier-prompt.md.
  • 1.2 Define fixed fact classifications: direct_evidence, indirect_support, no_evidence, contradicted.
  • 1.3 Define critical fact scope, verdict matrix, and groundedness_score mapping.
  • 1.4 Define strict JSON output schema: verdict, groundedness_score, critical_fact_count, facts_checked, rationale.
  • 1.5 Forbid Markdown, code fences, schema-extra fields, and text outside the JSON object.
  • 1.6 Require facts_checked[*].evidence_refs for traceability to tool evidence.

2. Verifier Input Hook

  • 2.1 Add VerifierInputHook.java as a MessagesModelHook running at BEFORE_MODEL.
  • 2.2 Replace raw verifier history with explicit payload fields: original_query, executor_final_answer, tool_trace_summary, retry_context.
  • 2.3 Persist the current round tool_trace_summary in VerifierContextHolder for later verifier evaluation storage.

3. ChatService Integration

  • 3.1 Load chatVerifierPrompt and add buildChatVerifierAgent().
  • 3.2 Add configurable verifier.low-confidence-threshold.
  • 3.3 Implement explicit per-round orchestration in ChatService: planner call, executor call, verifier call.
  • 3.4 Keep max two outer rounds and inject retry_context only for the second round.
  • 3.5 Parse verifier JSON directly and fall back to LOW_CONFID when verifier output is missing or invalid.
  • 3.6 Keep SupervisorAgent construction as legacy residue only; runtime orchestration no longer depends on prompt-only supervisor sequencing.

4. Verdict Routing And User Output

  • 4.1 Route PASS to the executor answer.
  • 4.2 Route LOW_CONFID to a fixed disclaimer plus executor answer.
  • 4.3 Route REJECT to degraded output without passing through the raw unverified answer.
  • 4.4 Build LOW_CONFID gap lists only from verifier-identified gaps.
  • 4.5 Build DEGRADED confirmed facts, gaps, and next-step suggestions from verifier facts and trace summary.

5. Trace Summary And Observability

  • 5.1 Add ToolTraceSummaryService to build verifier evidence summaries from tool_invocation.
  • 5.2 Include only evidence-bearing tools by default.
  • 5.3 Compact repeated calls by tool and topic domain.
  • 5.4 Preserve source_invocation_ids, trace_ref, query samples, retrieval layers, relevance levels, and source document labels.
  • 5.5 Parse and persist facts_checked[*].evidence_refs.
  • 5.6 Persist verifier_evaluation.tool_trace_summary and traceability_version.
  • 5.7 Store concise verifier summaries in agent_step.thought while preserving fuller verifier output in model_output / self_evaluation.

6. self_evaluation Merge Semantics

  • 6.1 Add SelfEvaluationMergeService.
  • 6.2 Write verifier results under verifier_evaluation.
  • 6.3 Write rule scoring under rule_evaluation.
  • 6.4 Preserve the other channel with read-modify-write semantics.

7. Verification

  • 7.1 Compile verification: mvn -q -DskipTests compile.
  • 7.2 Runtime verification: /api/chat complex request reached planner -> executor -> verifier.
  • 7.3 Runtime verification: session 9138f064 persisted verifier_evaluation.facts_checked[*].evidence_refs.
  • 7.4 Runtime verification: session 9138f064 persisted tool_trace_summary[*].source_invocation_ids.
  • 7.5 Runtime verification: LOW_CONFID user output included disclaimer and verifier-derived gaps.