## 1. Reproduce And Baseline Evidence - [x] 1.1 Run a real `/api/chat` diagnosis with a fixed MySQL/HikariCP session id - [x] 1.2 Inspect Trace API, MySQL rows, and logs for step/tool counts, skill reads, RAG details, and verifier output - [x] 1.3 Record the live-run finding in change evidence after implementation is verified ## 2. Skill Read Boundary - [x] 2.1 Locate Executor prompt/context code that allows repeated `read_skill` calls - [x] 2.2 Add guidance/context so Executor reuses the loaded selected skill within one diagnosis path - [x] 2.3 Add focused test coverage that Planner has no `read_skill` and Executor reads the selected skill once in a single path ## 3. Verifier Evidence Summary - [x] 3.1 Inspect `ToolTraceSummaryService` merge and truncation behavior for log, metric, and knowledge rows - [x] 3.2 Preserve bounded concrete facts from successful `query_logs` and `query_metrics` rows in `tool_trace_summary` - [x] 3.3 Add tests for merged/truncated rows that still contain direct evidence ## 4. Modular RAG Details - [x] 4.1 Inspect `ToolInvocationRecorder` and `LookupKnowledgeTool` detail assembly - [x] 4.2 Persist modular RAG keys in `tool_invocation.retrieval_details` for new `lookup_knowledge` rows - [x] 4.3 Add tests that assert `query_transform`, `retrieval_trace`, `context_pack_summary`, `rerank_trace`, and `evidence_blocks` ## 5. Verification - [x] 5.1 Run focused unit tests for skill, trace summary, and RAG recorder behavior - [x] 5.2 Run OpenSpec strict validation for `live-diagnosis-skill-observability` - [x] 5.3 Rerun the real diagnosis flow and inspect Trace API, MySQL, and logs - [x] 5.4 Decide whether remaining issues require a product/design decision before further code changes