2.8 KiB
2.8 KiB
Why
A real Chat diagnosis run for e2e-mysql-skill-rag-20260706-2230 completed, but the runtime evidence exposed gaps that the existing offline tests did not catch: Executor read the selected playbook twice, Verifier marked real log evidence as missing because the summary was too lossy, and lookup_knowledge.retrieval_details did not expose the modular RAG trace keys expected by the architecture.
This change closes the MVP live diagnosis loop so skill use, evidence tools, modular RAG trace details, and Verifier decisions can be trusted from actual diagnosis_session, agent_step, tool_invocation, and log evidence.
What Changes
- Add runtime-observable checks for skill selection and skill reading:
- Planner selects a skill from metadata without
read_skill. - Executor reads the selected skill on demand.
- Repeated
read_skillcalls for the same selected skill are avoided in a single diagnosis run.
- Planner selects a skill from metadata without
- Preserve verifier-useful evidence in
tool_trace_summaryso concrete log and metric facts are not lost behind generic truncated previews. - Ensure
lookup_knowledgepersists modular RAG trace details undertool_invocation.retrieval_detailsusing stable keys for query transform, retrieval trace, context pack summary, rerank trace, fallback reason, and evidence summaries. - Add focused regression coverage for the live-run failure modes without requiring a live LLM.
- No production API shape change is intended.
Capabilities
New Capabilities
None.
Modified Capabilities
diagnosis-playbook-skills: Executor should read the selected playbook once per diagnosis execution path unless a retry round explicitly requires a fresh read.chat-verifier-agent: Verifier-facingtool_trace_summaryshould preserve enough concrete evidence facts from persisted tool rows for direct evidence classification.evidence-trace-hardening: Evidence summaries should not downgrade successful evidence to no-evidence merely because raw outputs were truncated.rag-knowledge-retrieval:lookup_knowledge.retrieval_detailsshould persist modular RAG trace keys for live trace and evaluation consumers.
Impact
- Affected code:
ChatServiceand prompt/hook wiring around Planner, Executor, and Verifier.ToolTraceSummaryServiceevidence summarization.ToolInvocationRecorderand/orLookupKnowledgeToolretrieval details persistence.- Focused tests under
src/test/java.
- Affected data:
- Existing
diagnosis_session,agent_step, andtool_invocationtable shapes remain unchanged. tool_invocation.retrieval_detailsJSON gains or restores stable modular RAG fields.
- Existing
- Affected quality gates:
- Live diagnosis evidence can be assessed from Trace API, MySQL rows, and logs.
- Offline tests cover the observed live-run regressions.