68 lines
2.9 KiB
Markdown
68 lines
2.9 KiB
Markdown
## Evidence
|
|
|
|
### Baseline Live Finding
|
|
|
|
Session `e2e-mysql-skill-rag-20260706-2230` proved the Chat diagnosis path could run end to end but exposed runtime gaps:
|
|
|
|
- Planner selected the MySQL connection pool skill metadata path but Executor read the same selected skill twice.
|
|
- Verifier-facing summaries lost concrete log/metric evidence behind truncated raw previews.
|
|
- New `lookup_knowledge` rows did not consistently expose modular RAG trace keys in `tool_invocation.retrieval_details`.
|
|
- LOW_CONFID output still included raw Executor conclusions, so unsupported claims could reach the user.
|
|
|
|
### Final Live Verification
|
|
|
|
Session: `e2e-mysql-skill-rag-20260706-2355`
|
|
|
|
Trace/API and MySQL evidence:
|
|
|
|
- `diagnosis_session`: `id=73`, `status=SUCCESS`, `agent_flow=CHAT`, `step_count=9`, `tool_call_count=9`.
|
|
- Verifier: `LOW_CONFID`, `groundedness_score=0.5`.
|
|
- `agent_step` distribution:
|
|
- planner: 1 step, 0 tool steps, `selected_skill` present, no `read_skill`.
|
|
- executor: 7 steps, only first executor step contains `read_skill`.
|
|
- verifier: 1 step, no `read_skill`.
|
|
- `tool_invocation` distribution:
|
|
- `lookup_knowledge`: 1
|
|
- `get_available_log_topics`: 1
|
|
- `query_metrics`: 1
|
|
- `query_logs`: 6
|
|
- `lookup_knowledge` row:
|
|
- `retrieval_layer=L0+L1`
|
|
- `l0_match_count=5`
|
|
- `l1_match_count=3`
|
|
- `relevance_level=REFERENCE`
|
|
- `modularRagDetails=true`
|
|
- Verifier `tool_trace_summary` marks the generic mock system event row as `success=false`, `evidence_level=none`, `no_hit_invocation_count=1`.
|
|
- User-facing LOW_CONFID answer includes only direct evidence in `已确认信息`, moves unsupported facts to `当前缺口`, and does not include the raw Executor report.
|
|
|
|
Log evidence:
|
|
|
|
- `logs/codex-spring-run-live.out.log` shows classpath skill loading, Planner output containing `"selected_skill": "diagnose-mysql-connection-pool"`, and exactly one `read_skill` call for that skill in the final run.
|
|
- No `保存 tool_invocation 失败` or MySQL `Data too long` errors appeared for the final run.
|
|
|
|
### Commands
|
|
|
|
```powershell
|
|
openspec validate live-diagnosis-skill-observability --strict
|
|
```
|
|
|
|
Result: passed.
|
|
|
|
```powershell
|
|
mvn -q "-Dtest=ToolTraceSummaryServiceTest,ChatServiceSequentialAgentTest,ToolInvocationRecorderTest" test
|
|
```
|
|
|
|
Result: passed.
|
|
|
|
```powershell
|
|
java -cp "target;$(Get-Content target\e2e-classpath.txt)" E2eDbInspect e2e-mysql-skill-rag-20260706-2355
|
|
```
|
|
|
|
Result: confirmed the final live MySQL row counts, skill read boundary, and modular RAG details described above.
|
|
|
|
### Remaining Known Limits
|
|
|
|
- The Alibaba classpath registry still logs `Loaded 6 skills from classpath: skills`; the project-level `SingleSkillRegistry` exposes only the active skill to Planner/Executor after registry construction.
|
|
- LLM planning remains model-driven; the prompt and metadata hook now require `selected_skill`, and the final live run followed that contract.
|
|
- `read_skill` remains workflow guidance and is not persisted as a `tool_invocation` evidence row.
|