Files
SuperBizAgent-java/openspec/changes/archive/2026-07-06-live-diagnosis-skill-observability/proposal.md
T

44 lines
2.8 KiB
Markdown

## Why
A real Chat diagnosis run for `e2e-mysql-skill-rag-20260706-2230` completed, but the runtime evidence exposed gaps that the existing offline tests did not catch: Executor read the selected playbook twice, Verifier marked real log evidence as missing because the summary was too lossy, and `lookup_knowledge.retrieval_details` did not expose the modular RAG trace keys expected by the architecture.
This change closes the MVP live diagnosis loop so skill use, evidence tools, modular RAG trace details, and Verifier decisions can be trusted from actual `diagnosis_session`, `agent_step`, `tool_invocation`, and log evidence.
## What Changes
- Add runtime-observable checks for skill selection and skill reading:
- Planner selects a skill from metadata without `read_skill`.
- Executor reads the selected skill on demand.
- Repeated `read_skill` calls for the same selected skill are avoided in a single diagnosis run.
- Preserve verifier-useful evidence in `tool_trace_summary` so concrete log and metric facts are not lost behind generic truncated previews.
- Ensure `lookup_knowledge` persists modular RAG trace details under `tool_invocation.retrieval_details` using stable keys for query transform, retrieval trace, context pack summary, rerank trace, fallback reason, and evidence summaries.
- Add focused regression coverage for the live-run failure modes without requiring a live LLM.
- No production API shape change is intended.
## Capabilities
### New Capabilities
None.
### Modified Capabilities
- `diagnosis-playbook-skills`: Executor should read the selected playbook once per diagnosis execution path unless a retry round explicitly requires a fresh read.
- `chat-verifier-agent`: Verifier-facing `tool_trace_summary` should preserve enough concrete evidence facts from persisted tool rows for direct evidence classification.
- `evidence-trace-hardening`: Evidence summaries should not downgrade successful evidence to no-evidence merely because raw outputs were truncated.
- `rag-knowledge-retrieval`: `lookup_knowledge.retrieval_details` should persist modular RAG trace keys for live trace and evaluation consumers.
## Impact
- Affected code:
- `ChatService` and prompt/hook wiring around Planner, Executor, and Verifier.
- `ToolTraceSummaryService` evidence summarization.
- `ToolInvocationRecorder` and/or `LookupKnowledgeTool` retrieval details persistence.
- Focused tests under `src/test/java`.
- Affected data:
- Existing `diagnosis_session`, `agent_step`, and `tool_invocation` table shapes remain unchanged.
- `tool_invocation.retrieval_details` JSON gains or restores stable modular RAG fields.
- Affected quality gates:
- Live diagnosis evidence can be assessed from Trace API, MySQL rows, and logs.
- Offline tests cover the observed live-run regressions.