4.2 KiB
Context
The real Chat run e2e-mysql-skill-rag-20260706-2230 proved the MVP can execute Planner, Executor, evidence tools, Verifier, Trace API, MySQL persistence, and logs end to end. It also exposed runtime gaps not covered by the current offline tests:
- Planner correctly selected
diagnose-mysql-connection-poolfrom metadata and did not callread_skill. - Executor called
read_skilltwice for the same selected skill in one diagnosis path. tool_invocationpersisted concrete log evidence, but the verifier-facing summary was lossy enough that Verifier labeled several existing facts asno_evidence.lookup_knowledgerows did not expose the modular RAG trace keys expected by the architecture inretrieval_details.
The fix should preserve table schemas and public APIs. The evidence contract should improve at JSON-detail and summary levels.
Goals / Non-Goals
Goals:
- Keep Planner metadata-only and Executor on-demand skill reading.
- Avoid duplicate
read_skillcalls for the same selected skill in a single Executor diagnosis path. - Preserve concrete evidence snippets in
tool_trace_summaryfor logs, metrics, and knowledge retrieval without passing full raw outputs. - Persist modular RAG details for
lookup_knowledgeintool_invocation.retrieval_details. - Add offline regression tests that reproduce the observed live-run symptoms without requiring a real LLM.
Non-Goals:
- No new database columns or table migrations.
- No new production API endpoint.
- No change to the public
/api/chator Trace API response shape beyond richer existing JSON/detail fields. - No model-based reranker, new retrieval backend, or extra Agent round policy change.
Decisions
D1. Treat duplicate read_skill as an Executor behavior constraint
The selected playbook should be read once and retained in the active execution context. The implementation can enforce this through prompt constraints and, where feasible, deterministic message/context shaping so the Executor sees that the skill has already been loaded.
Alternative considered: persist read_skill calls in tool_invocation and deduplicate from DB state. Rejected because read_skill is workflow guidance, not evidence; persisting it as evidence would blur the boundary already documented for playbook skills.
D2. Improve summaries at the evidence boundary, not by giving Verifier full raw outputs
ToolTraceSummaryService should extract compact concrete facts from output_preview/details and include them in output_summary. For example, log summaries should retain matched service, level, message, and important metric keys; metrics summaries should retain alert names/services; lookup summaries should retain source titles and modular trace highlights.
Alternative considered: pass full tool outputs to Verifier. Rejected because it increases token cost and reintroduces noisy/raw evidence into model inputs.
D3. Keep modular RAG observability in retrieval_details
lookup_knowledge should write JSON fields for query transform, retrieval trace, context pack summary, rerank trace, fallback reason, and evidence summaries. Existing relational columns remain the coarse query surface; JSON details carry the richer pipeline state.
Alternative considered: add new columns for each RAG trace field. Rejected because the current MVP trace model intentionally keeps schema stable and uses JSON details for retrieval-specific expansion.
D4. Verify with focused tests plus one real rerun
Offline tests should cover duplicate skill-read suppression, evidence summary preservation, and RAG detail persistence. After code changes, rerun the same real diagnosis flow and inspect Trace API, MySQL, and logs.
Risks / Trade-offs
- Summary extraction may still miss domain-specific facts -> keep extraction conservative and test the observed MySQL/log cases first.
- Prompt-only duplicate
read_skillprevention may be probabilistic -> prefer deterministic context markers if available, and verify with a real rerun. - Richer summaries increase verifier tokens -> cap snippets and avoid full raw outputs.
- Existing historical rows will not gain new RAG detail fields -> acceptable because the change affects new invocations only.