1.7 KiB
1.7 KiB
1. Reproduce And Baseline Evidence
- 1.1 Run a real
/api/chatdiagnosis with a fixed MySQL/HikariCP session id - 1.2 Inspect Trace API, MySQL rows, and logs for step/tool counts, skill reads, RAG details, and verifier output
- 1.3 Record the live-run finding in change evidence after implementation is verified
2. Skill Read Boundary
- 2.1 Locate Executor prompt/context code that allows repeated
read_skillcalls - 2.2 Add guidance/context so Executor reuses the loaded selected skill within one diagnosis path
- 2.3 Add focused test coverage that Planner has no
read_skilland Executor reads the selected skill once in a single path
3. Verifier Evidence Summary
- 3.1 Inspect
ToolTraceSummaryServicemerge and truncation behavior for log, metric, and knowledge rows - 3.2 Preserve bounded concrete facts from successful
query_logsandquery_metricsrows intool_trace_summary - 3.3 Add tests for merged/truncated rows that still contain direct evidence
4. Modular RAG Details
- 4.1 Inspect
ToolInvocationRecorderandLookupKnowledgeTooldetail assembly - 4.2 Persist modular RAG keys in
tool_invocation.retrieval_detailsfor newlookup_knowledgerows - 4.3 Add tests that assert
query_transform,retrieval_trace,context_pack_summary,rerank_trace, andevidence_blocks
5. Verification
- 5.1 Run focused unit tests for skill, trace summary, and RAG recorder behavior
- 5.2 Run OpenSpec strict validation for
live-diagnosis-skill-observability - 5.3 Rerun the real diagnosis flow and inspect Trace API, MySQL, and logs
- 5.4 Decide whether remaining issues require a product/design decision before further code changes