Files

31 lines
1.7 KiB
Markdown

## 1. Reproduce And Baseline Evidence
- [x] 1.1 Run a real `/api/chat` diagnosis with a fixed MySQL/HikariCP session id
- [x] 1.2 Inspect Trace API, MySQL rows, and logs for step/tool counts, skill reads, RAG details, and verifier output
- [x] 1.3 Record the live-run finding in change evidence after implementation is verified
## 2. Skill Read Boundary
- [x] 2.1 Locate Executor prompt/context code that allows repeated `read_skill` calls
- [x] 2.2 Add guidance/context so Executor reuses the loaded selected skill within one diagnosis path
- [x] 2.3 Add focused test coverage that Planner has no `read_skill` and Executor reads the selected skill once in a single path
## 3. Verifier Evidence Summary
- [x] 3.1 Inspect `ToolTraceSummaryService` merge and truncation behavior for log, metric, and knowledge rows
- [x] 3.2 Preserve bounded concrete facts from successful `query_logs` and `query_metrics` rows in `tool_trace_summary`
- [x] 3.3 Add tests for merged/truncated rows that still contain direct evidence
## 4. Modular RAG Details
- [x] 4.1 Inspect `ToolInvocationRecorder` and `LookupKnowledgeTool` detail assembly
- [x] 4.2 Persist modular RAG keys in `tool_invocation.retrieval_details` for new `lookup_knowledge` rows
- [x] 4.3 Add tests that assert `query_transform`, `retrieval_trace`, `context_pack_summary`, `rerank_trace`, and `evidence_blocks`
## 5. Verification
- [x] 5.1 Run focused unit tests for skill, trace summary, and RAG recorder behavior
- [x] 5.2 Run OpenSpec strict validation for `live-diagnosis-skill-observability`
- [x] 5.3 Rerun the real diagnosis flow and inspect Trace API, MySQL, and logs
- [x] 5.4 Decide whether remaining issues require a product/design decision before further code changes