feat: add chat verifier agent

This commit is contained in:
zhuyongxin
2026-07-03 10:54:33 +08:00
parent 4f5316d473
commit 9050487307
28 changed files with 3200 additions and 209 deletions
@@ -0,0 +1,52 @@
# Evidence: chat-verifier-agent
## Code Evidence
### Complex chat path now invokes verifier deterministically
- File: `src/main/java/com/superbiz/agent/service/ChatService.java`
- Evidence: `executeChatComplex` calls planner, executor, then verifier directly through `callAgent(...)`.
- Conclusion: runtime no longer depends on prompt-only Supervisor behavior to call verifier.
### Verifier receives explicit inputs
- File: `src/main/java/com/superbiz/agent/hook/VerifierInputHook.java`
- Evidence: the hook builds a JSON payload with `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
- Conclusion: verifier input is stable and does not depend on guessing the last assistant message from raw history.
### Tool evidence is traceable to persisted invocations
- File: `src/main/java/com/superbiz/agent/service/ToolTraceSummaryService.java`
- Evidence: summaries include `trace_ref`, `source_invocation_ids`, `query_samples`, `retrieval_layers`, `relevance_levels`, and `source_documents`.
- Conclusion: verifier facts can be correlated with actual `tool_invocation` rows.
### Verifier facts preserve evidence references
- File: `src/main/java/com/superbiz/agent/service/ChatService.java`
- Evidence: verifier parsing preserves `facts_checked[*].evidence_refs` and persists `tool_trace_summary` under `verifier_evaluation`.
- Conclusion: `self_evaluation` now contains both verifier judgments and the evidence index used to form them.
### Evaluation channels no longer overwrite each other
- File: `src/main/java/com/superbiz/agent/service/SelfEvaluationMergeService.java`
- Evidence: rule and verifier evaluations are merged into separate keys.
- Conclusion: asynchronous rule scoring preserves verifier output.
### Verifier logging is less noisy
- File: `src/main/java/com/superbiz/agent/hook/AgentLoggingHook.java`
- Evidence: verifier `thought` stores a concise verdict summary, while fuller model output remains available in structured storage.
- Conclusion: `agent_step.thought` is no longer a misleading place for full verifier JSON.
## Runtime Evidence
- Compile verification passed: `mvn -q -DskipTests compile`.
- Runtime session `9138f064` executed `planner -> executor -> verifier`.
- Runtime session `9138f064` persisted `verifier_evaluation.facts_checked[*].evidence_refs`.
- Runtime session `9138f064` persisted `verifier_evaluation.tool_trace_summary[*].source_invocation_ids`.
## Design Evidence
- `LOW_CONFID` returns a fixed disclaimer and verifier-derived gaps.
- `REJECT` returns degraded output and does not pass through the raw Executor answer.
- `retry_context` is derived from verifier-identified missing evidence facts.