44 lines
2.8 KiB
Markdown
44 lines
2.8 KiB
Markdown
## Why
|
|
|
|
A real Chat diagnosis run for `e2e-mysql-skill-rag-20260706-2230` completed, but the runtime evidence exposed gaps that the existing offline tests did not catch: Executor read the selected playbook twice, Verifier marked real log evidence as missing because the summary was too lossy, and `lookup_knowledge.retrieval_details` did not expose the modular RAG trace keys expected by the architecture.
|
|
|
|
This change closes the MVP live diagnosis loop so skill use, evidence tools, modular RAG trace details, and Verifier decisions can be trusted from actual `diagnosis_session`, `agent_step`, `tool_invocation`, and log evidence.
|
|
|
|
## What Changes
|
|
|
|
- Add runtime-observable checks for skill selection and skill reading:
|
|
- Planner selects a skill from metadata without `read_skill`.
|
|
- Executor reads the selected skill on demand.
|
|
- Repeated `read_skill` calls for the same selected skill are avoided in a single diagnosis run.
|
|
- Preserve verifier-useful evidence in `tool_trace_summary` so concrete log and metric facts are not lost behind generic truncated previews.
|
|
- Ensure `lookup_knowledge` persists modular RAG trace details under `tool_invocation.retrieval_details` using stable keys for query transform, retrieval trace, context pack summary, rerank trace, fallback reason, and evidence summaries.
|
|
- Add focused regression coverage for the live-run failure modes without requiring a live LLM.
|
|
- No production API shape change is intended.
|
|
|
|
## Capabilities
|
|
|
|
### New Capabilities
|
|
|
|
None.
|
|
|
|
### Modified Capabilities
|
|
|
|
- `diagnosis-playbook-skills`: Executor should read the selected playbook once per diagnosis execution path unless a retry round explicitly requires a fresh read.
|
|
- `chat-verifier-agent`: Verifier-facing `tool_trace_summary` should preserve enough concrete evidence facts from persisted tool rows for direct evidence classification.
|
|
- `evidence-trace-hardening`: Evidence summaries should not downgrade successful evidence to no-evidence merely because raw outputs were truncated.
|
|
- `rag-knowledge-retrieval`: `lookup_knowledge.retrieval_details` should persist modular RAG trace keys for live trace and evaluation consumers.
|
|
|
|
## Impact
|
|
|
|
- Affected code:
|
|
- `ChatService` and prompt/hook wiring around Planner, Executor, and Verifier.
|
|
- `ToolTraceSummaryService` evidence summarization.
|
|
- `ToolInvocationRecorder` and/or `LookupKnowledgeTool` retrieval details persistence.
|
|
- Focused tests under `src/test/java`.
|
|
- Affected data:
|
|
- Existing `diagnosis_session`, `agent_step`, and `tool_invocation` table shapes remain unchanged.
|
|
- `tool_invocation.retrieval_details` JSON gains or restores stable modular RAG fields.
|
|
- Affected quality gates:
|
|
- Live diagnosis evidence can be assessed from Trace API, MySQL rows, and logs.
|
|
- Offline tests cover the observed live-run regressions.
|