Files
SuperBizAgent-java/openspec/changes/archive/2026-07-06-live-diagnosis-skill-observability/design.md
T

59 lines
4.2 KiB
Markdown

## Context
The real Chat run `e2e-mysql-skill-rag-20260706-2230` proved the MVP can execute Planner, Executor, evidence tools, Verifier, Trace API, MySQL persistence, and logs end to end. It also exposed runtime gaps not covered by the current offline tests:
- Planner correctly selected `diagnose-mysql-connection-pool` from metadata and did not call `read_skill`.
- Executor called `read_skill` twice for the same selected skill in one diagnosis path.
- `tool_invocation` persisted concrete log evidence, but the verifier-facing summary was lossy enough that Verifier labeled several existing facts as `no_evidence`.
- `lookup_knowledge` rows did not expose the modular RAG trace keys expected by the architecture in `retrieval_details`.
The fix should preserve table schemas and public APIs. The evidence contract should improve at JSON-detail and summary levels.
## Goals / Non-Goals
**Goals:**
- Keep Planner metadata-only and Executor on-demand skill reading.
- Avoid duplicate `read_skill` calls for the same selected skill in a single Executor diagnosis path.
- Preserve concrete evidence snippets in `tool_trace_summary` for logs, metrics, and knowledge retrieval without passing full raw outputs.
- Persist modular RAG details for `lookup_knowledge` in `tool_invocation.retrieval_details`.
- Add offline regression tests that reproduce the observed live-run symptoms without requiring a real LLM.
**Non-Goals:**
- No new database columns or table migrations.
- No new production API endpoint.
- No change to the public `/api/chat` or Trace API response shape beyond richer existing JSON/detail fields.
- No model-based reranker, new retrieval backend, or extra Agent round policy change.
## Decisions
### D1. Treat duplicate `read_skill` as an Executor behavior constraint
The selected playbook should be read once and retained in the active execution context. The implementation can enforce this through prompt constraints and, where feasible, deterministic message/context shaping so the Executor sees that the skill has already been loaded.
Alternative considered: persist `read_skill` calls in `tool_invocation` and deduplicate from DB state. Rejected because `read_skill` is workflow guidance, not evidence; persisting it as evidence would blur the boundary already documented for playbook skills.
### D2. Improve summaries at the evidence boundary, not by giving Verifier full raw outputs
`ToolTraceSummaryService` should extract compact concrete facts from `output_preview`/details and include them in `output_summary`. For example, log summaries should retain matched service, level, message, and important metric keys; metrics summaries should retain alert names/services; lookup summaries should retain source titles and modular trace highlights.
Alternative considered: pass full tool outputs to Verifier. Rejected because it increases token cost and reintroduces noisy/raw evidence into model inputs.
### D3. Keep modular RAG observability in `retrieval_details`
`lookup_knowledge` should write JSON fields for query transform, retrieval trace, context pack summary, rerank trace, fallback reason, and evidence summaries. Existing relational columns remain the coarse query surface; JSON details carry the richer pipeline state.
Alternative considered: add new columns for each RAG trace field. Rejected because the current MVP trace model intentionally keeps schema stable and uses JSON details for retrieval-specific expansion.
### D4. Verify with focused tests plus one real rerun
Offline tests should cover duplicate skill-read suppression, evidence summary preservation, and RAG detail persistence. After code changes, rerun the same real diagnosis flow and inspect Trace API, MySQL, and logs.
## Risks / Trade-offs
- Summary extraction may still miss domain-specific facts -> keep extraction conservative and test the observed MySQL/log cases first.
- Prompt-only duplicate `read_skill` prevention may be probabilistic -> prefer deterministic context markers if available, and verify with a real rerun.
- Richer summaries increase verifier tokens -> cap snippets and avoid full raw outputs.
- Existing historical rows will not gain new RAG detail fields -> acceptable because the change affects new invocations only.