Files

2.8 KiB

Why

A real Chat diagnosis run for e2e-mysql-skill-rag-20260706-2230 completed, but the runtime evidence exposed gaps that the existing offline tests did not catch: Executor read the selected playbook twice, Verifier marked real log evidence as missing because the summary was too lossy, and lookup_knowledge.retrieval_details did not expose the modular RAG trace keys expected by the architecture.

This change closes the MVP live diagnosis loop so skill use, evidence tools, modular RAG trace details, and Verifier decisions can be trusted from actual diagnosis_session, agent_step, tool_invocation, and log evidence.

What Changes

  • Add runtime-observable checks for skill selection and skill reading:
    • Planner selects a skill from metadata without read_skill.
    • Executor reads the selected skill on demand.
    • Repeated read_skill calls for the same selected skill are avoided in a single diagnosis run.
  • Preserve verifier-useful evidence in tool_trace_summary so concrete log and metric facts are not lost behind generic truncated previews.
  • Ensure lookup_knowledge persists modular RAG trace details under tool_invocation.retrieval_details using stable keys for query transform, retrieval trace, context pack summary, rerank trace, fallback reason, and evidence summaries.
  • Add focused regression coverage for the live-run failure modes without requiring a live LLM.
  • No production API shape change is intended.

Capabilities

New Capabilities

None.

Modified Capabilities

  • diagnosis-playbook-skills: Executor should read the selected playbook once per diagnosis execution path unless a retry round explicitly requires a fresh read.
  • chat-verifier-agent: Verifier-facing tool_trace_summary should preserve enough concrete evidence facts from persisted tool rows for direct evidence classification.
  • evidence-trace-hardening: Evidence summaries should not downgrade successful evidence to no-evidence merely because raw outputs were truncated.
  • rag-knowledge-retrieval: lookup_knowledge.retrieval_details should persist modular RAG trace keys for live trace and evaluation consumers.

Impact

  • Affected code:
    • ChatService and prompt/hook wiring around Planner, Executor, and Verifier.
    • ToolTraceSummaryService evidence summarization.
    • ToolInvocationRecorder and/or LookupKnowledgeTool retrieval details persistence.
    • Focused tests under src/test/java.
  • Affected data:
    • Existing diagnosis_session, agent_step, and tool_invocation table shapes remain unchanged.
    • tool_invocation.retrieval_details JSON gains or restores stable modular RAG fields.
  • Affected quality gates:
    • Live diagnosis evidence can be assessed from Trace API, MySQL rows, and logs.
    • Offline tests cover the observed live-run regressions.