fix(agent): harden live diagnosis skill observability

This commit is contained in:
aruo
2026-07-07 00:20:34 +08:00
parent 64adb998cf
commit b3315ead52
20 changed files with 792 additions and 26 deletions
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-06
@@ -0,0 +1,58 @@
## Context
The real Chat run `e2e-mysql-skill-rag-20260706-2230` proved the MVP can execute Planner, Executor, evidence tools, Verifier, Trace API, MySQL persistence, and logs end to end. It also exposed runtime gaps not covered by the current offline tests:
- Planner correctly selected `diagnose-mysql-connection-pool` from metadata and did not call `read_skill`.
- Executor called `read_skill` twice for the same selected skill in one diagnosis path.
- `tool_invocation` persisted concrete log evidence, but the verifier-facing summary was lossy enough that Verifier labeled several existing facts as `no_evidence`.
- `lookup_knowledge` rows did not expose the modular RAG trace keys expected by the architecture in `retrieval_details`.
The fix should preserve table schemas and public APIs. The evidence contract should improve at JSON-detail and summary levels.
## Goals / Non-Goals
**Goals:**
- Keep Planner metadata-only and Executor on-demand skill reading.
- Avoid duplicate `read_skill` calls for the same selected skill in a single Executor diagnosis path.
- Preserve concrete evidence snippets in `tool_trace_summary` for logs, metrics, and knowledge retrieval without passing full raw outputs.
- Persist modular RAG details for `lookup_knowledge` in `tool_invocation.retrieval_details`.
- Add offline regression tests that reproduce the observed live-run symptoms without requiring a real LLM.
**Non-Goals:**
- No new database columns or table migrations.
- No new production API endpoint.
- No change to the public `/api/chat` or Trace API response shape beyond richer existing JSON/detail fields.
- No model-based reranker, new retrieval backend, or extra Agent round policy change.
## Decisions
### D1. Treat duplicate `read_skill` as an Executor behavior constraint
The selected playbook should be read once and retained in the active execution context. The implementation can enforce this through prompt constraints and, where feasible, deterministic message/context shaping so the Executor sees that the skill has already been loaded.
Alternative considered: persist `read_skill` calls in `tool_invocation` and deduplicate from DB state. Rejected because `read_skill` is workflow guidance, not evidence; persisting it as evidence would blur the boundary already documented for playbook skills.
### D2. Improve summaries at the evidence boundary, not by giving Verifier full raw outputs
`ToolTraceSummaryService` should extract compact concrete facts from `output_preview`/details and include them in `output_summary`. For example, log summaries should retain matched service, level, message, and important metric keys; metrics summaries should retain alert names/services; lookup summaries should retain source titles and modular trace highlights.
Alternative considered: pass full tool outputs to Verifier. Rejected because it increases token cost and reintroduces noisy/raw evidence into model inputs.
### D3. Keep modular RAG observability in `retrieval_details`
`lookup_knowledge` should write JSON fields for query transform, retrieval trace, context pack summary, rerank trace, fallback reason, and evidence summaries. Existing relational columns remain the coarse query surface; JSON details carry the richer pipeline state.
Alternative considered: add new columns for each RAG trace field. Rejected because the current MVP trace model intentionally keeps schema stable and uses JSON details for retrieval-specific expansion.
### D4. Verify with focused tests plus one real rerun
Offline tests should cover duplicate skill-read suppression, evidence summary preservation, and RAG detail persistence. After code changes, rerun the same real diagnosis flow and inspect Trace API, MySQL, and logs.
## Risks / Trade-offs
- Summary extraction may still miss domain-specific facts -> keep extraction conservative and test the observed MySQL/log cases first.
- Prompt-only duplicate `read_skill` prevention may be probabilistic -> prefer deterministic context markers if available, and verify with a real rerun.
- Richer summaries increase verifier tokens -> cap snippets and avoid full raw outputs.
- Existing historical rows will not gain new RAG detail fields -> acceptable because the change affects new invocations only.
@@ -0,0 +1,67 @@
## Evidence
### Baseline Live Finding
Session `e2e-mysql-skill-rag-20260706-2230` proved the Chat diagnosis path could run end to end but exposed runtime gaps:
- Planner selected the MySQL connection pool skill metadata path but Executor read the same selected skill twice.
- Verifier-facing summaries lost concrete log/metric evidence behind truncated raw previews.
- New `lookup_knowledge` rows did not consistently expose modular RAG trace keys in `tool_invocation.retrieval_details`.
- LOW_CONFID output still included raw Executor conclusions, so unsupported claims could reach the user.
### Final Live Verification
Session: `e2e-mysql-skill-rag-20260706-2355`
Trace/API and MySQL evidence:
- `diagnosis_session`: `id=73`, `status=SUCCESS`, `agent_flow=CHAT`, `step_count=9`, `tool_call_count=9`.
- Verifier: `LOW_CONFID`, `groundedness_score=0.5`.
- `agent_step` distribution:
- planner: 1 step, 0 tool steps, `selected_skill` present, no `read_skill`.
- executor: 7 steps, only first executor step contains `read_skill`.
- verifier: 1 step, no `read_skill`.
- `tool_invocation` distribution:
- `lookup_knowledge`: 1
- `get_available_log_topics`: 1
- `query_metrics`: 1
- `query_logs`: 6
- `lookup_knowledge` row:
- `retrieval_layer=L0+L1`
- `l0_match_count=5`
- `l1_match_count=3`
- `relevance_level=REFERENCE`
- `modularRagDetails=true`
- Verifier `tool_trace_summary` marks the generic mock system event row as `success=false`, `evidence_level=none`, `no_hit_invocation_count=1`.
- User-facing LOW_CONFID answer includes only direct evidence in `已确认信息`, moves unsupported facts to `当前缺口`, and does not include the raw Executor report.
Log evidence:
- `logs/codex-spring-run-live.out.log` shows classpath skill loading, Planner output containing `"selected_skill": "diagnose-mysql-connection-pool"`, and exactly one `read_skill` call for that skill in the final run.
- No `保存 tool_invocation 失败` or MySQL `Data too long` errors appeared for the final run.
### Commands
```powershell
openspec validate live-diagnosis-skill-observability --strict
```
Result: passed.
```powershell
mvn -q "-Dtest=ToolTraceSummaryServiceTest,ChatServiceSequentialAgentTest,ToolInvocationRecorderTest" test
```
Result: passed.
```powershell
java -cp "target;$(Get-Content target\e2e-classpath.txt)" E2eDbInspect e2e-mysql-skill-rag-20260706-2355
```
Result: confirmed the final live MySQL row counts, skill read boundary, and modular RAG details described above.
### Remaining Known Limits
- The Alibaba classpath registry still logs `Loaded 6 skills from classpath: skills`; the project-level `SingleSkillRegistry` exposes only the active skill to Planner/Executor after registry construction.
- LLM planning remains model-driven; the prompt and metadata hook now require `selected_skill`, and the final live run followed that contract.
- `read_skill` remains workflow guidance and is not persisted as a `tool_invocation` evidence row.
@@ -0,0 +1,43 @@
## Why
A real Chat diagnosis run for `e2e-mysql-skill-rag-20260706-2230` completed, but the runtime evidence exposed gaps that the existing offline tests did not catch: Executor read the selected playbook twice, Verifier marked real log evidence as missing because the summary was too lossy, and `lookup_knowledge.retrieval_details` did not expose the modular RAG trace keys expected by the architecture.
This change closes the MVP live diagnosis loop so skill use, evidence tools, modular RAG trace details, and Verifier decisions can be trusted from actual `diagnosis_session`, `agent_step`, `tool_invocation`, and log evidence.
## What Changes
- Add runtime-observable checks for skill selection and skill reading:
- Planner selects a skill from metadata without `read_skill`.
- Executor reads the selected skill on demand.
- Repeated `read_skill` calls for the same selected skill are avoided in a single diagnosis run.
- Preserve verifier-useful evidence in `tool_trace_summary` so concrete log and metric facts are not lost behind generic truncated previews.
- Ensure `lookup_knowledge` persists modular RAG trace details under `tool_invocation.retrieval_details` using stable keys for query transform, retrieval trace, context pack summary, rerank trace, fallback reason, and evidence summaries.
- Add focused regression coverage for the live-run failure modes without requiring a live LLM.
- No production API shape change is intended.
## Capabilities
### New Capabilities
None.
### Modified Capabilities
- `diagnosis-playbook-skills`: Executor should read the selected playbook once per diagnosis execution path unless a retry round explicitly requires a fresh read.
- `chat-verifier-agent`: Verifier-facing `tool_trace_summary` should preserve enough concrete evidence facts from persisted tool rows for direct evidence classification.
- `evidence-trace-hardening`: Evidence summaries should not downgrade successful evidence to no-evidence merely because raw outputs were truncated.
- `rag-knowledge-retrieval`: `lookup_knowledge.retrieval_details` should persist modular RAG trace keys for live trace and evaluation consumers.
## Impact
- Affected code:
- `ChatService` and prompt/hook wiring around Planner, Executor, and Verifier.
- `ToolTraceSummaryService` evidence summarization.
- `ToolInvocationRecorder` and/or `LookupKnowledgeTool` retrieval details persistence.
- Focused tests under `src/test/java`.
- Affected data:
- Existing `diagnosis_session`, `agent_step`, and `tool_invocation` table shapes remain unchanged.
- `tool_invocation.retrieval_details` JSON gains or restores stable modular RAG fields.
- Affected quality gates:
- Live diagnosis evidence can be assessed from Trace API, MySQL rows, and logs.
- Offline tests cover the observed live-run regressions.
@@ -0,0 +1,27 @@
## ADDED Requirements
### Requirement: Verifier evidence summaries SHALL preserve concrete supporting facts
The verifier-facing `tool_trace_summary` SHALL preserve compact concrete facts from persisted evidence-tool outputs so direct evidence is not misclassified as missing merely because raw output was truncated.
#### Scenario: Log evidence contains a concrete matching message
- **WHEN** a persisted `query_logs` invocation output contains a concrete log message matching a critical fact
- **THEN** the generated `tool_trace_summary` SHALL include that message or a bounded excerpt of it in `output_summary`
- **AND** Verifier SHALL be able to reference the invocation id as direct evidence
#### Scenario: Metrics evidence contains concrete alert fields
- **WHEN** a persisted `query_metrics` invocation output contains alert names, services, or metric values
- **THEN** the generated `tool_trace_summary` SHALL include the relevant alert names, services, and bounded metric values
- **AND** it SHALL NOT imply unsupported alerts that are absent from the tool output
#### Scenario: Summary remains bounded
- **WHEN** a tool output is large
- **THEN** the generated `tool_trace_summary` SHALL remain bounded
- **AND** it SHALL preserve concrete facts before generic boilerplate or low-value formatting
### Requirement: Verifier low-confidence output SHALL not present unsupported claims as confirmed
When Verifier returns `LOW_CONFID`, user-facing output SHALL clearly separate confirmed facts from evidence gaps and SHALL NOT leave unsupported Executor claims formatted as confirmed findings.
#### Scenario: LOW_CONFID with critical evidence gaps
- **WHEN** Verifier labels critical facts as `no_evidence`
- **THEN** the final user-facing response SHALL identify those gaps from verifier output
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
@@ -0,0 +1,26 @@
## ADDED Requirements
### Requirement: Executor SHALL avoid duplicate playbook reads in one diagnosis path
The system SHALL prevent repeated reads of the same selected diagnosis playbook during a single Executor diagnosis path unless a new retry round explicitly requests a fresh playbook read.
#### Scenario: Executor has already read the selected skill
- **WHEN** the Executor has already called `read_skill` for the selected skill in the current diagnosis path
- **THEN** the Executor SHALL continue using the already loaded playbook instructions
- **AND** it SHALL NOT call `read_skill` again for the same skill in that path
#### Scenario: Retry round may read selected skill again
- **WHEN** ChatService starts a distinct retry round after Verifier requests evidence supplementation
- **THEN** the Executor MAY call `read_skill` again for the selected skill
- **AND** the repeated read SHALL be attributable to the new round rather than the same Executor path
### Requirement: Live traces SHALL expose skill selection boundaries
The system SHALL make the Planner and Executor skill boundary auditable from persisted agent steps and logs.
#### Scenario: Planner selects without reading skill body
- **WHEN** a Planner step selects a diagnosis playbook
- **THEN** the persisted Planner output SHALL include the selected skill name
- **AND** the Planner step SHALL NOT include a `read_skill` tool call
#### Scenario: Executor reads selected skill
- **WHEN** the Executor starts a playbook-backed diagnosis
- **THEN** the persisted Executor step SHALL show `read_skill` for the selected skill before evidence-tool execution
@@ -0,0 +1,25 @@
## ADDED Requirements
### Requirement: Evidence summaries SHALL preserve strongest concrete support across merged invocations
When multiple invocations are merged into one verifier-facing summary entry, the summary SHALL preserve concrete support from the strongest successful invocation.
#### Scenario: Merged log calls include one direct hit
- **WHEN** repeated `query_logs` calls for the same topic are merged
- **AND** at least one call contains a concrete direct evidence hit
- **THEN** the merged summary SHALL include a bounded direct-hit excerpt
- **AND** it SHALL preserve all contributing `source_invocation_ids`
#### Scenario: Truncated preview still contains direct evidence
- **WHEN** a persisted output preview is marked truncated
- **AND** the preview contains a concrete direct evidence hit
- **THEN** the summary SHALL treat the row as successful evidence
- **AND** it SHALL NOT downgrade the row to no-evidence only because the raw output was truncated
### Requirement: Live trace counts SHALL distinguish model steps from evidence tool invocations
The persisted trace SHALL make it possible to audit model step counts separately from evidence tool invocation counts.
#### Scenario: Diagnosis session exposes aggregate counts
- **WHEN** a diagnosis session completes
- **THEN** `diagnosis_session.step_count` SHALL count persisted Agent model steps
- **AND** `diagnosis_session.tool_call_count` SHALL count persisted evidence-tool invocation rows
- **AND** helper workflow calls that are not evidence rows SHALL be auditable from agent steps or logs without inflating `tool_invocation`
@@ -0,0 +1,22 @@
## ADDED Requirements
### Requirement: Lookup knowledge SHALL persist modular RAG trace details
The `lookup_knowledge` tool SHALL persist modular RAG pipeline details in `tool_invocation.retrieval_details` for every new lookup invocation.
#### Scenario: Evidence lookup persists modular detail keys
- **WHEN** `lookup_knowledge` returns usable evidence
- **THEN** `retrieval_details` SHALL include `query_transform`
- **AND** it SHALL include `retrieval_trace`
- **AND** it SHALL include `context_pack_summary`
- **AND** it SHALL include `rerank_trace`
- **AND** it SHALL include `evidence_blocks`
#### Scenario: Fallback lookup persists fallback reason
- **WHEN** `lookup_knowledge` performs an unfiltered retry after filtered retrieval fails or is low quality
- **THEN** `retrieval_details.retrieval_trace` SHALL include the selected attempt
- **AND** `retrieval_details.fallback_reason` SHALL preserve the fallback reason
#### Scenario: No-evidence lookup still preserves trace
- **WHEN** `lookup_knowledge` returns no usable evidence
- **THEN** `retrieval_details` SHALL still include retrieval trace information for attempted retrieval paths
- **AND** it SHALL not include full evidence content as duplicated trace data
@@ -0,0 +1,30 @@
## 1. Reproduce And Baseline Evidence
- [x] 1.1 Run a real `/api/chat` diagnosis with a fixed MySQL/HikariCP session id
- [x] 1.2 Inspect Trace API, MySQL rows, and logs for step/tool counts, skill reads, RAG details, and verifier output
- [x] 1.3 Record the live-run finding in change evidence after implementation is verified
## 2. Skill Read Boundary
- [x] 2.1 Locate Executor prompt/context code that allows repeated `read_skill` calls
- [x] 2.2 Add guidance/context so Executor reuses the loaded selected skill within one diagnosis path
- [x] 2.3 Add focused test coverage that Planner has no `read_skill` and Executor reads the selected skill once in a single path
## 3. Verifier Evidence Summary
- [x] 3.1 Inspect `ToolTraceSummaryService` merge and truncation behavior for log, metric, and knowledge rows
- [x] 3.2 Preserve bounded concrete facts from successful `query_logs` and `query_metrics` rows in `tool_trace_summary`
- [x] 3.3 Add tests for merged/truncated rows that still contain direct evidence
## 4. Modular RAG Details
- [x] 4.1 Inspect `ToolInvocationRecorder` and `LookupKnowledgeTool` detail assembly
- [x] 4.2 Persist modular RAG keys in `tool_invocation.retrieval_details` for new `lookup_knowledge` rows
- [x] 4.3 Add tests that assert `query_transform`, `retrieval_trace`, `context_pack_summary`, `rerank_trace`, and `evidence_blocks`
## 5. Verification
- [x] 5.1 Run focused unit tests for skill, trace summary, and RAG recorder behavior
- [x] 5.2 Run OpenSpec strict validation for `live-diagnosis-skill-observability`
- [x] 5.3 Rerun the real diagnosis flow and inspect Trace API, MySQL, and logs
- [x] 5.4 Decide whether remaining issues require a product/design decision before further code changes
@@ -204,3 +204,29 @@ The verifier integration SHALL continue to work when evidence summaries distingu
- **THEN** those entries SHALL be treated as no-new-evidence
- **AND** they SHALL NOT be interpreted as fresh direct support for the answer
### Requirement: Verifier evidence summaries SHALL preserve concrete supporting facts
The verifier-facing `tool_trace_summary` SHALL preserve compact concrete facts from persisted evidence-tool outputs so direct evidence is not misclassified as missing merely because raw output was truncated.
#### Scenario: Log evidence contains a concrete matching message
- **WHEN** a persisted `query_logs` invocation output contains a concrete log message matching a critical fact
- **THEN** the generated `tool_trace_summary` SHALL include that message or a bounded excerpt of it in `output_summary`
- **AND** Verifier SHALL be able to reference the invocation id as direct evidence
#### Scenario: Metrics evidence contains concrete alert fields
- **WHEN** a persisted `query_metrics` invocation output contains alert names, services, or metric values
- **THEN** the generated `tool_trace_summary` SHALL include the relevant alert names, services, and bounded metric values
- **AND** it SHALL NOT imply unsupported alerts that are absent from the tool output
#### Scenario: Summary remains bounded
- **WHEN** a tool output is large
- **THEN** the generated `tool_trace_summary` SHALL remain bounded
- **AND** it SHALL preserve concrete facts before generic boilerplate or low-value formatting
### Requirement: Verifier low-confidence output SHALL not present unsupported claims as confirmed
When Verifier returns `LOW_CONFID`, user-facing output SHALL clearly separate confirmed facts from evidence gaps and SHALL NOT leave unsupported Executor claims formatted as confirmed findings.
#### Scenario: LOW_CONFID with critical evidence gaps
- **WHEN** Verifier labels critical facts as `no_evidence`
- **THEN** the final user-facing response SHALL identify those gaps from verifier output
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
@@ -3,9 +3,7 @@
## Purpose
Provide versionable diagnosis playbook skills for high-frequency MVP troubleshooting flows, loaded through progressive disclosure so agents can follow scenario-specific evidence workflows without bloating every prompt.
## Requirements
### Requirement: Skill catalog SHALL expose diagnosis playbooks compactly
The system SHALL provide a compact skill catalog containing each playbook skill name and description.
@@ -54,3 +52,28 @@ The system SHALL provide playbooks for the existing fixed diagnosis evaluation s
- **WHEN** the case is payment timeout, MySQL pool exhaustion, Redis timeout, slow response, or JVM memory risk
- **THEN** a matching diagnosis skill SHALL exist
- **AND** the skill SHALL state required evidence tools and low-confidence behavior
### Requirement: Executor SHALL avoid duplicate playbook reads in one diagnosis path
The system SHALL prevent repeated reads of the same selected diagnosis playbook during a single Executor diagnosis path unless a new retry round explicitly requests a fresh playbook read.
#### Scenario: Executor has already read the selected skill
- **WHEN** the Executor has already called `read_skill` for the selected skill in the current diagnosis path
- **THEN** the Executor SHALL continue using the already loaded playbook instructions
- **AND** it SHALL NOT call `read_skill` again for the same skill in that path
#### Scenario: Retry round may read selected skill again
- **WHEN** ChatService starts a distinct retry round after Verifier requests evidence supplementation
- **THEN** the Executor MAY call `read_skill` again for the selected skill
- **AND** the repeated read SHALL be attributable to the new round rather than the same Executor path
### Requirement: Live traces SHALL expose skill selection boundaries
The system SHALL make the Planner and Executor skill boundary auditable from persisted agent steps and logs.
#### Scenario: Planner selects without reading skill body
- **WHEN** a Planner step selects a diagnosis playbook
- **THEN** the persisted Planner output SHALL include the selected skill name
- **AND** the Planner step SHALL NOT include a `read_skill` tool call
#### Scenario: Executor reads selected skill
- **WHEN** the Executor starts a playbook-backed diagnosis
- **THEN** the persisted Executor step SHALL show `read_skill` for the selected skill before evidence-tool execution
@@ -84,3 +84,27 @@ The system SHALL provide offline tests for the hardened evidence contract and de
- **WHEN** the test suite runs the focused `ChatService` fallback tests
- **THEN** it SHALL verify the missing-output, invalid-JSON, `LOW_CONFID`, and `REJECT` degraded paths without requiring a real LLM or database
### Requirement: Evidence summaries SHALL preserve strongest concrete support across merged invocations
When multiple invocations are merged into one verifier-facing summary entry, the summary SHALL preserve concrete support from the strongest successful invocation.
#### Scenario: Merged log calls include one direct hit
- **WHEN** repeated `query_logs` calls for the same topic are merged
- **AND** at least one call contains a concrete direct evidence hit
- **THEN** the merged summary SHALL include a bounded direct-hit excerpt
- **AND** it SHALL preserve all contributing `source_invocation_ids`
#### Scenario: Truncated preview still contains direct evidence
- **WHEN** a persisted output preview is marked truncated
- **AND** the preview contains a concrete direct evidence hit
- **THEN** the summary SHALL treat the row as successful evidence
- **AND** it SHALL NOT downgrade the row to no-evidence only because the raw output was truncated
### Requirement: Live trace counts SHALL distinguish model steps from evidence tool invocations
The persisted trace SHALL make it possible to audit model step counts separately from evidence tool invocation counts.
#### Scenario: Diagnosis session exposes aggregate counts
- **WHEN** a diagnosis session completes
- **THEN** `diagnosis_session.step_count` SHALL count persisted Agent model steps
- **AND** `diagnosis_session.tool_call_count` SHALL count persisted evidence-tool invocation rows
- **AND** helper workflow calls that are not evidence rows SHALL be auditable from agent steps or logs without inflating `tool_invocation`
@@ -227,3 +227,24 @@ The post-retrieval flow SHALL rerank vector candidates using deterministic rule-
- **THEN** the reranker MAY boost the candidate
- **AND** the evidence block SHALL record the hint as a hit reason
- **AND** the system SHALL NOT treat the L0 hint itself as fact evidence
### Requirement: Lookup knowledge SHALL persist modular RAG trace details
The `lookup_knowledge` tool SHALL persist modular RAG pipeline details in `tool_invocation.retrieval_details` for every new lookup invocation.
#### Scenario: Evidence lookup persists modular detail keys
- **WHEN** `lookup_knowledge` returns usable evidence
- **THEN** `retrieval_details` SHALL include `query_transform`
- **AND** it SHALL include `retrieval_trace`
- **AND** it SHALL include `context_pack_summary`
- **AND** it SHALL include `rerank_trace`
- **AND** it SHALL include `evidence_blocks`
#### Scenario: Fallback lookup persists fallback reason
- **WHEN** `lookup_knowledge` performs an unfiltered retry after filtered retrieval fails or is low quality
- **THEN** `retrieval_details.retrieval_trace` SHALL include the selected attempt
- **AND** `retrieval_details.fallback_reason` SHALL preserve the fallback reason
#### Scenario: No-evidence lookup still preserves trace
- **WHEN** `lookup_knowledge` returns no usable evidence
- **THEN** `retrieval_details` SHALL still include retrieval trace information for attempted retrieval paths
- **AND** it SHALL not include full evidence content as duplicated trace data
@@ -544,6 +544,10 @@ public class ChatService {
private ReactAgent buildChatExecutorAgent(ChatModel chatModel, ToolCallback[] toolCallbacks,
List<Map<String, String>> history, String retryContext) {
StringBuilder prompt = new StringBuilder(chatExecutorPrompt);
prompt.append("\n\n--- Skill 读取约束 ---\n")
.append("如果 planner_plan 已给出 selected_skill,本轮 Executor 只允许对该 skill 调用一次 read_skill。")
.append("读取后必须复用已加载的 playbook 指令继续执行证据工具,不要为了检查支持文件、确认流程或生成报告再次读取同一个 skill。")
.append("只有 ChatService 启动新的补证据 retry round 时,才可以重新读取 selected_skill。\n");
if (!history.isEmpty()) {
prompt.append("\n\n--- 对话历史 ---\n");
for (Map<String, String> msg : history) {
@@ -764,14 +768,29 @@ public class ChatService {
private String buildLowConfidenceOutput(String executorAnswer, VerifierDecision decision) {
StringBuilder output = new StringBuilder(LOW_CONFID_DISCLAIMER);
output.append("\n\n").append(executorAnswer == null ? "" : executorAnswer);
List<String> confirmedFacts = extractConfirmedFacts(decision);
output.append("\n\n已确认信息:");
if (confirmedFacts.isEmpty()) {
output.append("\n- 暂无可稳定确认的信息");
} else {
for (String fact : confirmedFacts) {
output.append("\n- ").append(fact);
}
}
List<String> gaps = extractEvidenceGaps(decision);
if (!gaps.isEmpty()) {
output.append("\n\n当前缺口:");
if (!gaps.isEmpty()) {
for (String gap : gaps) {
output.append("\n- ").append(gap);
}
} else {
output.append("\n- 当前缺少足够的直接证据支撑核心结论");
}
output.append("\n\n建议下一步:");
for (String suggestion : buildNextStepSuggestions(decision)) {
output.append("\n- ").append(suggestion);
}
return output.toString();
}
@@ -813,7 +832,7 @@ public class ChatService {
for (Map<String, Object> fact : decision.factsChecked()) {
String verification = String.valueOf(fact.get("verification"));
boolean critical = Boolean.TRUE.equals(fact.get("is_critical"));
if (critical && ("direct_evidence".equals(verification) || "indirect_support".equals(verification))) {
if (critical && "direct_evidence".equals(verification)) {
confirmedFacts.add(String.valueOf(fact.get("fact")));
}
}
@@ -138,21 +138,11 @@ public class ToolInvocationRecorder {
if (record.evidenceBlockCount() != null) {
details.put("evidence_block_count", record.evidenceBlockCount());
}
if (record.evidenceBlocks() != null && !record.evidenceBlocks().isEmpty()) {
details.put("evidence_blocks", record.evidenceBlocks());
}
if (record.queryTransform() != null && !record.queryTransform().isEmpty()) {
details.put("query_transform", record.queryTransform());
}
if (record.retrievalTrace() != null && !record.retrievalTrace().isEmpty()) {
details.put("retrieval_trace", record.retrievalTrace());
}
if (record.contextPack() != null && !record.contextPack().isEmpty()) {
details.put("context_pack_summary", record.contextPack());
}
if (record.rerankTrace() != null && !record.rerankTrace().isEmpty()) {
details.put("rerank_trace", record.rerankTrace());
}
details.put("evidence_blocks", record.evidenceBlocks() == null ? List.of() : record.evidenceBlocks());
details.put("query_transform", record.queryTransform() == null ? Map.of() : record.queryTransform());
details.put("retrieval_trace", record.retrievalTrace() == null ? Map.of() : record.retrievalTrace());
details.put("context_pack_summary", record.contextPack() == null ? Map.of() : record.contextPack());
details.put("rerank_trace", record.rerankTrace() == null ? Map.of() : record.rerankTrace());
if (record.fallbackReason() != null && !record.fallbackReason().isBlank()) {
details.put("fallback_reason", record.fallbackReason());
}
@@ -253,7 +243,7 @@ public class ToolInvocationRecorder {
String dedupReason,
int durationMs) {
RetrievalTrace trace = result != null ? result.getRetrievalTrace() : null;
String layer = trace != null ? trace.getSelectedAttempt() : null;
String layer = recordedRetrievalLayer(query, trace);
String outputPreview = null;
int outputLength = 0;
@@ -462,5 +452,20 @@ public class ToolInvocationRecorder {
}
return scores;
}
private static String recordedRetrievalLayer(KnowledgeQuery query, RetrievalTrace trace) {
if (trace == null || trace.getSelectedAttempt() == null) {
return null;
}
boolean hasL0 = query != null && query.getL0MatchCount() != null && query.getL0MatchCount() > 0;
boolean hasL1 = trace.getSelectedAttempt().contains("VECTOR");
if (hasL0 && hasL1) {
return "L0+L1";
}
if (hasL1) {
return "L1";
}
return hasL0 ? "L0" : null;
}
}
}
@@ -15,6 +15,8 @@ import java.util.List;
import java.util.Locale;
import java.util.Map;
import java.util.Set;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
/**
* Builds a verifier-facing evidence index from persisted tool invocations.
@@ -25,6 +27,7 @@ public class ToolTraceSummaryService {
private static final TypeReference<LinkedHashMap<String, Object>> MAP_TYPE = new TypeReference<>() {};
private static final Set<String> EVIDENCE_TOOLS = Set.of("lookup_knowledge", "query_logs", "query_metrics", "query_order");
private static final Pattern JSON_STRING_FIELD = Pattern.compile("\"%s\"\\s*:\\s*\"((?:\\\\.|[^\"])*)\"");
private final ToolInvocationRepository toolInvocationRepository;
private final ObjectMapper objectMapper = new ObjectMapper();
@@ -128,18 +131,156 @@ public class ToolTraceSummaryService {
if ("lookup_knowledge".equals(invocation.getToolName())) {
String relevance = invocation.getRelevanceLevel() != null ? invocation.getRelevanceLevel() : "UNKNOWN";
String trace = extractLookupTraceSummary(invocation);
String preview = invocation.getOutputPreview() != null && !invocation.getOutputPreview().isBlank()
? truncate(invocation.getOutputPreview(), 160)
? truncate(invocation.getOutputPreview(), 240)
: "no preview";
return "matched domain=" + topicDomain + ", relevance=" + relevance + ", preview=" + preview;
return "matched domain=" + topicDomain + ", relevance=" + relevance + trace + ", preview=" + preview;
}
if ("query_logs".equals(invocation.getToolName())) {
String concrete = extractLogEvidenceSummary(invocation.getOutputPreview());
if (!concrete.isBlank()) {
return concrete;
}
}
if ("query_metrics".equals(invocation.getToolName())) {
String concrete = extractMetricsEvidenceSummary(invocation.getOutputPreview());
if (!concrete.isBlank()) {
return concrete;
}
}
if (invocation.getOutputPreview() != null && !invocation.getOutputPreview().isBlank()) {
return truncate(invocation.getOutputPreview(), 160);
return truncate(invocation.getOutputPreview(), 240);
}
return "evidence retrieved without preview";
}
private String extractLookupTraceSummary(ToolInvocation invocation) {
if (invocation.getRetrievalDetails() == null || invocation.getRetrievalDetails().isBlank()) {
return "";
}
try {
Map<String, Object> details = objectMapper.readValue(invocation.getRetrievalDetails(), MAP_TYPE);
List<String> sources = new ArrayList<>();
Object evidenceBlocks = details.get("evidence_blocks");
if (evidenceBlocks instanceof List<?> blocks) {
for (Object block : blocks) {
if (block instanceof Map<?, ?> blockMap) {
Object title = blockMap.get("title");
Object source = blockMap.get("source");
String label = title != null ? String.valueOf(title) : String.valueOf(source);
if (label != null && !label.isBlank() && !"null".equals(label)) {
sources.add(label);
}
}
if (sources.size() >= 3) {
break;
}
}
}
Object retrievalTrace = details.get("retrieval_trace");
String selectedAttempt = "";
if (retrievalTrace instanceof Map<?, ?> traceMap && traceMap.get("selected_attempt") != null) {
selectedAttempt = ", selected_attempt=" + traceMap.get("selected_attempt");
}
return (sources.isEmpty() ? "" : ", sources=" + truncate(String.join("|", sources), 180)) + selectedAttempt;
} catch (Exception e) {
log.debug("Failed to parse lookup retrieval details", e);
return "";
}
}
private String extractLogEvidenceSummary(String outputPreview) {
if (outputPreview == null || outputPreview.isBlank()) {
return "";
}
List<String> messages = extractJsonStringFields(outputPreview, "message", 3);
List<String> services = extractJsonStringFields(outputPreview, "service", 3);
List<String> levels = extractJsonStringFields(outputPreview, "level", 3);
List<String> timestamps = extractJsonStringFields(outputPreview, "timestamp", 3);
if (messages.isEmpty()) {
return "";
}
List<String> rows = new ArrayList<>();
for (int i = 0; i < messages.size(); i++) {
String prefix = labelAt(timestamps, i) + labelAt(levels, i) + labelAt(services, i);
rows.add((prefix.isBlank() ? "" : prefix + " ") + truncate(messages.get(i), 220));
}
return "log_evidence: " + truncate(String.join(" | ", rows), 520);
}
private boolean hasOnlyGenericMockLogMessages(String outputPreview) {
List<String> messages = extractJsonStringFields(outputPreview, "message", 3);
if (messages.isEmpty()) {
return false;
}
return messages.stream()
.allMatch(message -> message.startsWith("日志消息 #") && message.contains("查询条件:"));
}
private String extractMetricsEvidenceSummary(String outputPreview) {
if (outputPreview == null || outputPreview.isBlank()) {
return "";
}
List<String> alertNames = extractJsonStringFields(outputPreview, "alert_name", 5);
List<String> descriptions = extractJsonStringFields(outputPreview, "description", 5);
List<String> services = extractJsonStringFields(outputPreview, "service", 5);
if (alertNames.isEmpty() && descriptions.isEmpty()) {
return "";
}
List<String> rows = new ArrayList<>();
int count = Math.max(alertNames.size(), descriptions.size());
for (int i = 0; i < Math.min(5, count); i++) {
StringBuilder row = new StringBuilder();
if (i < alertNames.size()) {
row.append(alertNames.get(i));
}
if (i < services.size()) {
if (!row.isEmpty()) {
row.append(" ");
}
row.append("service=").append(services.get(i));
}
if (i < descriptions.size()) {
if (!row.isEmpty()) {
row.append(": ");
}
row.append(descriptions.get(i));
}
rows.add(truncate(row.toString(), 220));
}
return "metric_evidence: " + truncate(String.join(" | ", rows), 520);
}
private List<String> extractJsonStringFields(String text, String field, int limit) {
Pattern pattern = Pattern.compile(String.format(JSON_STRING_FIELD.pattern(), Pattern.quote(field)));
Matcher matcher = pattern.matcher(text);
List<String> values = new ArrayList<>();
while (matcher.find() && values.size() < limit) {
values.add(unescapeJsonString(matcher.group(1)));
}
return values;
}
private String unescapeJsonString(String value) {
return value == null ? "" : value
.replace("\\\"", "\"")
.replace("\\\\", "\\")
.replace("\\n", "\n")
.replace("\\r", "\r")
.replace("\\t", "\t");
}
private String labelAt(List<String> values, int index) {
if (index >= values.size() || values.get(index) == null || values.get(index).isBlank()) {
return "";
}
return "[" + values.get(index) + "]";
}
private String determineEvidenceLevel(ToolInvocation invocation) {
String evidenceStatus = extractEvidenceStatus(invocation);
if (!Boolean.TRUE.equals(invocation.getSuccess())) {
@@ -149,6 +290,10 @@ public class ToolTraceSummaryService {
|| ToolInvocationRecorder.EVIDENCE_STATUS_DEDUPED.equals(evidenceStatus)) {
return "none";
}
if ("query_logs".equals(invocation.getToolName())
&& hasOnlyGenericMockLogMessages(invocation.getOutputPreview())) {
return "none";
}
if ("PRECISE".equals(invocation.getRelevanceLevel()) || "HIGHLY_RELEVANT".equals(invocation.getRelevanceLevel())) {
return "direct";
}
@@ -283,12 +428,44 @@ public class ToolTraceSummaryService {
}
String invocationEvidenceLevel = determineEvidenceLevel(invocation);
if ("none".equals(invocationEvidenceLevel)) {
noHitCount++;
if (outputSummary == null || outputSummary.isBlank()) {
outputSummary = extractOutputSummary(invocation, topicDomain);
}
return;
}
if (!success || evidenceRank(invocationEvidenceLevel) > evidenceRank(evidenceLevel)) {
success = true;
evidenceLevel = invocationEvidenceLevel;
outputSummary = extractOutputSummary(invocation, topicDomain);
} else if (evidenceRank(invocationEvidenceLevel) == evidenceRank(evidenceLevel)) {
String candidateSummary = extractOutputSummary(invocation, topicDomain);
if (isMoreConcrete(candidateSummary, outputSummary)) {
outputSummary = candidateSummary;
}
}
}
private boolean isMoreConcrete(String candidate, String current) {
return concretenessScore(candidate) > concretenessScore(current);
}
private int concretenessScore(String summary) {
if (summary == null || summary.isBlank()) {
return 0;
}
int score = summary.length() > 160 ? 2 : 1;
if (summary.contains("log_evidence") || summary.contains("metric_evidence")) {
score += 5;
}
if (summary.contains("连接池耗尽") || summary.contains("OutOfMemoryError")
|| summary.contains("扫描行数") || summary.contains("HighMemoryUsage")
|| summary.contains("HighCPUUsage")) {
score += 4;
}
return score;
}
int relevanceScore(String answer) {
int score = success ? 10 : 0;
@@ -9,6 +9,8 @@
```json
{
"selected_skill": "匹配的 skill 名称;如果没有匹配则为 null",
"selection_reason": "选择该 skill 的原因;如果没有匹配则说明不使用 skill",
"plan": ["步骤1描述", "步骤2描述", "步骤3描述"],
"reasoning": "规划思路说明"
}
@@ -18,6 +20,9 @@
- 每个步骤应该是一个可以独立执行的任务
- 步骤要具体可操作,不要模糊
- 如果问题需要查知识库,明确在步骤中说明要查什么
- 如果存在 `skill_catalog`,必须先根据 skill 的 name/description 判断是否匹配用户问题
- 如果匹配某个诊断 Skill,必须在顶层 `selected_skill` 填入 skill 名称,并在计划第一步说明 Executor 需要读取该 skill
- Planner 只能选择 skill 元数据,不能调用 `read_skill`,也不能编造 skill 正文内容
## 知识库检索规则
- 制定步骤前,先查看下方 `available_knowledge_domains`(如果存在)
@@ -86,10 +86,61 @@ class ChatServiceSequentialAgentTest {
"sequential-low-confidence-session"
);
assertTrue(result.answer().contains("EXECUTOR_FINAL_ANSWER"));
assertTrue(result.answer().startsWith("以下结论基于当前已获取证据"));
assertFalse(result.answer().contains("EXECUTOR_FINAL_ANSWER"));
assertTrue(result.answer().contains("当前缺口"));
assertEquals(List.of("chat_planner", "chat_executor", "chat_verifier"), chatModel.agentCalls);
}
@Test
void executeChatComplexLowConfidenceConfirmedFactsOnlyUseDirectEvidence() throws Exception {
ChatService chatService = createChatService();
ScriptedChatModel chatModel = new ScriptedChatModel("""
{
"verdict": "LOW_CONFID",
"groundedness_score": 0.37,
"critical_fact_count": 3,
"facts_checked": [
{
"fact": "连接池耗尽 active=50/50",
"is_critical": true,
"verification": "direct_evidence",
"detail": "log evidence",
"evidence_refs": []
},
{
"fact": "临时扩容连接池到 80",
"is_critical": true,
"verification": "indirect_support",
"detail": "suggestion inferred from evidence",
"evidence_refs": []
},
{
"fact": "OOM 导致连接泄漏",
"is_critical": true,
"verification": "no_evidence",
"detail": "missing OOM log",
"evidence_refs": []
}
],
"rationale": "scripted low confidence"
}
""");
ChatService.ChatResult result = chatService.executeChatComplex(
chatModel,
new ToolCallback[0],
"请分析 MySQL 连接池耗尽",
List.of(),
"sequential-low-confid-direct-only-session"
);
assertTrue(result.answer().contains("已确认信息:\n- 连接池耗尽 active=50/50"));
assertFalse(result.answer().contains("临时扩容连接池到 80"));
assertTrue(result.answer().contains("当前缺口:\n- OOM 导致连接泄漏:missing OOM log"));
assertFalse(result.answer().contains("EXECUTOR_FINAL_ANSWER"));
}
@Test
void executeChatComplexFallsBackToLowConfidenceWhenVerifierOutputMissing() throws Exception {
ChatService chatService = createChatService();
@@ -243,6 +294,7 @@ class ChatServiceSequentialAgentTest {
assertTrue(chatModel.executorPromptText.contains("## Skills System"));
assertTrue(chatModel.executorPromptText.contains("diagnose-mysql-connection-pool"));
assertTrue(chatModel.executorPromptText.contains("read_skill"));
assertTrue(chatModel.executorPromptText.contains("只允许对该 skill 调用一次 read_skill"));
assertFalse(chatModel.verifierPromptText.contains("diagnose-mysql-connection-pool"));
assertFalse(chatModel.verifierPromptText.contains("read_skill"));
}
@@ -113,6 +113,10 @@ class ToolInvocationRecorderTest {
assertTrue(saved.getRetrievalDetails().contains("\"evidence_candidate_count\":2"));
assertTrue(saved.getRetrievalDetails().contains("\"evidence_block_count\":1"));
assertTrue(saved.getRetrievalDetails().contains("\"evidence_blocks\""));
assertTrue(saved.getRetrievalDetails().contains("\"query_transform\""));
assertTrue(saved.getRetrievalDetails().contains("\"retrieval_trace\""));
assertTrue(saved.getRetrievalDetails().contains("\"context_pack_summary\""));
assertTrue(saved.getRetrievalDetails().contains("\"rerank_trace\""));
}
@Test
@@ -176,7 +180,9 @@ class ToolInvocationRecorderTest {
assertEquals(1, record.evidenceBlockCount());
assertEquals(1, record.evidenceBlocks().size());
assertTrue(String.valueOf(record.evidenceBlocks().get(0).get("content_preview")).endsWith("..."));
assertEquals("UNFILTERED_VECTOR", record.retrievalLayer());
assertEquals("L1", record.retrievalLayer());
assertEquals(3, record.l1MatchCount());
assertTrue(record.retrievalTrace().containsKey("selected_attempt"));
assertTrue(record.contextPack().containsKey("included_sources"));
}
}
@@ -62,6 +62,7 @@ class ToolTraceSummaryServiceTest {
assertEquals("direct", logsSummary.get("evidence_level"));
assertEquals(2, logsSummary.get("invocation_count"));
assertEquals(1, logsSummary.get("no_hit_invocation_count"));
assertTrue(String.valueOf(logsSummary.get("output_summary")).contains("payment timeout stack trace"));
Map<String, Object> metricsSummary = summaries.stream()
.filter(item -> "query_metrics".equals(item.get("tool_name")))
@@ -72,4 +73,111 @@ class ToolTraceSummaryServiceTest {
assertEquals(1, metricsSummary.get("failed_invocation_count"));
assertTrue(String.valueOf(metricsSummary.get("output_summary")).contains("call failed"));
}
@Test
void buildVerifierTraceSummaryPreservesConcreteFactsFromTruncatedLogAndMetricRows() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
String logPreview = """
{
"success" : true,
"logs" : [ {
"timestamp" : "2026-07-06 22:15:45",
"level" : "ERROR",
"service" : "order-service",
"message" : "数据库连接池耗尽: Cannot acquire connection from pool, active: 50/50, waiting: 23, timeout: 30000ms"
} ]
}
""";
String metricPreview = """
{
"success" : true,
"alerts" : [ {
"alert_name" : "HighCPUUsage",
"service" : "payment-service",
"description" : "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。"
} ]
}
""";
when(repository.findBySessionIdOrderByIdAsc("session-2")).thenReturn(List.of(
ToolInvocation.builder()
.id(10L)
.sessionId("session-2")
.toolName("query_logs")
.inputParams("{\"query\":\"pool\"}")
.outputPreview(logPreview)
.retrievalDetails("{\"retrieved_domains\":[\"application-logs\"],\"evidence_status\":\"supported\"}")
.isTruncated(true)
.success(true)
.build(),
ToolInvocation.builder()
.id(11L)
.sessionId("session-2")
.toolName("query_metrics")
.inputParams("{\"query\":\"active_prometheus_alerts\"}")
.outputPreview(metricPreview)
.retrievalDetails("{\"retrieved_domains\":[\"prometheus_alerts\"],\"evidence_status\":\"supported\"}")
.isTruncated(true)
.success(true)
.build()
));
ToolTraceSummaryService service = new ToolTraceSummaryService(repository);
List<Map<String, Object>> summaries = service.buildVerifierTraceSummary("session-2", "连接池耗尽 HighCPUUsage");
Map<String, Object> logsSummary = summaries.stream()
.filter(item -> "query_logs".equals(item.get("tool_name")))
.findFirst()
.orElseThrow();
assertTrue(String.valueOf(logsSummary.get("output_summary")).contains("连接池耗尽"));
assertTrue(String.valueOf(logsSummary.get("output_summary")).contains("active: 50/50"));
assertEquals(List.of(10L), logsSummary.get("source_invocation_ids"));
Map<String, Object> metricsSummary = summaries.stream()
.filter(item -> "query_metrics".equals(item.get("tool_name")))
.findFirst()
.orElseThrow();
assertTrue(String.valueOf(metricsSummary.get("output_summary")).contains("HighCPUUsage"));
assertTrue(String.valueOf(metricsSummary.get("output_summary")).contains("payment-service"));
}
@Test
void buildVerifierTraceSummaryDoesNotTreatGenericMockLogsAsDirectEvidence() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
String genericLogPreview = """
{
"success" : true,
"logs" : [ {
"timestamp" : "2026-07-06 23:44:41",
"level" : "ERROR",
"service" : "generic-service",
"message" : "日志消息 #0, 查询条件: service:payment-service"
} ]
}
""";
when(repository.findBySessionIdOrderByIdAsc("session-3")).thenReturn(List.of(
ToolInvocation.builder()
.id(20L)
.sessionId("session-3")
.toolName("query_logs")
.inputParams("{\"query\":\"service:payment-service\"}")
.outputPreview(genericLogPreview)
.retrievalDetails("{\"retrieved_domains\":[\"system-metrics\"],\"evidence_status\":\"supported\"}")
.success(true)
.build()
));
ToolTraceSummaryService service = new ToolTraceSummaryService(repository);
List<Map<String, Object>> summaries = service.buildVerifierTraceSummary("session-3", "payment-service timeout");
Map<String, Object> logsSummary = summaries.stream()
.filter(item -> "query_logs".equals(item.get("tool_name")))
.findFirst()
.orElseThrow();
assertEquals(Boolean.FALSE, logsSummary.get("success"));
assertEquals("none", logsSummary.get("evidence_level"));
assertEquals(1, logsSummary.get("no_hit_invocation_count"));
assertTrue(String.valueOf(logsSummary.get("output_summary")).contains("日志消息 #0"));
}
}