fix(agent): harden live diagnosis skill observability

This commit is contained in:
aruo
2026-07-07 00:20:34 +08:00
parent 64adb998cf
commit b3315ead52
20 changed files with 792 additions and 26 deletions
@@ -204,3 +204,29 @@ The verifier integration SHALL continue to work when evidence summaries distingu
- **THEN** those entries SHALL be treated as no-new-evidence
- **AND** they SHALL NOT be interpreted as fresh direct support for the answer
### Requirement: Verifier evidence summaries SHALL preserve concrete supporting facts
The verifier-facing `tool_trace_summary` SHALL preserve compact concrete facts from persisted evidence-tool outputs so direct evidence is not misclassified as missing merely because raw output was truncated.
#### Scenario: Log evidence contains a concrete matching message
- **WHEN** a persisted `query_logs` invocation output contains a concrete log message matching a critical fact
- **THEN** the generated `tool_trace_summary` SHALL include that message or a bounded excerpt of it in `output_summary`
- **AND** Verifier SHALL be able to reference the invocation id as direct evidence
#### Scenario: Metrics evidence contains concrete alert fields
- **WHEN** a persisted `query_metrics` invocation output contains alert names, services, or metric values
- **THEN** the generated `tool_trace_summary` SHALL include the relevant alert names, services, and bounded metric values
- **AND** it SHALL NOT imply unsupported alerts that are absent from the tool output
#### Scenario: Summary remains bounded
- **WHEN** a tool output is large
- **THEN** the generated `tool_trace_summary` SHALL remain bounded
- **AND** it SHALL preserve concrete facts before generic boilerplate or low-value formatting
### Requirement: Verifier low-confidence output SHALL not present unsupported claims as confirmed
When Verifier returns `LOW_CONFID`, user-facing output SHALL clearly separate confirmed facts from evidence gaps and SHALL NOT leave unsupported Executor claims formatted as confirmed findings.
#### Scenario: LOW_CONFID with critical evidence gaps
- **WHEN** Verifier labels critical facts as `no_evidence`
- **THEN** the final user-facing response SHALL identify those gaps from verifier output
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
@@ -3,9 +3,7 @@
## Purpose
Provide versionable diagnosis playbook skills for high-frequency MVP troubleshooting flows, loaded through progressive disclosure so agents can follow scenario-specific evidence workflows without bloating every prompt.
## Requirements
### Requirement: Skill catalog SHALL expose diagnosis playbooks compactly
The system SHALL provide a compact skill catalog containing each playbook skill name and description.
@@ -54,3 +52,28 @@ The system SHALL provide playbooks for the existing fixed diagnosis evaluation s
- **WHEN** the case is payment timeout, MySQL pool exhaustion, Redis timeout, slow response, or JVM memory risk
- **THEN** a matching diagnosis skill SHALL exist
- **AND** the skill SHALL state required evidence tools and low-confidence behavior
### Requirement: Executor SHALL avoid duplicate playbook reads in one diagnosis path
The system SHALL prevent repeated reads of the same selected diagnosis playbook during a single Executor diagnosis path unless a new retry round explicitly requests a fresh playbook read.
#### Scenario: Executor has already read the selected skill
- **WHEN** the Executor has already called `read_skill` for the selected skill in the current diagnosis path
- **THEN** the Executor SHALL continue using the already loaded playbook instructions
- **AND** it SHALL NOT call `read_skill` again for the same skill in that path
#### Scenario: Retry round may read selected skill again
- **WHEN** ChatService starts a distinct retry round after Verifier requests evidence supplementation
- **THEN** the Executor MAY call `read_skill` again for the selected skill
- **AND** the repeated read SHALL be attributable to the new round rather than the same Executor path
### Requirement: Live traces SHALL expose skill selection boundaries
The system SHALL make the Planner and Executor skill boundary auditable from persisted agent steps and logs.
#### Scenario: Planner selects without reading skill body
- **WHEN** a Planner step selects a diagnosis playbook
- **THEN** the persisted Planner output SHALL include the selected skill name
- **AND** the Planner step SHALL NOT include a `read_skill` tool call
#### Scenario: Executor reads selected skill
- **WHEN** the Executor starts a playbook-backed diagnosis
- **THEN** the persisted Executor step SHALL show `read_skill` for the selected skill before evidence-tool execution
@@ -84,3 +84,27 @@ The system SHALL provide offline tests for the hardened evidence contract and de
- **WHEN** the test suite runs the focused `ChatService` fallback tests
- **THEN** it SHALL verify the missing-output, invalid-JSON, `LOW_CONFID`, and `REJECT` degraded paths without requiring a real LLM or database
### Requirement: Evidence summaries SHALL preserve strongest concrete support across merged invocations
When multiple invocations are merged into one verifier-facing summary entry, the summary SHALL preserve concrete support from the strongest successful invocation.
#### Scenario: Merged log calls include one direct hit
- **WHEN** repeated `query_logs` calls for the same topic are merged
- **AND** at least one call contains a concrete direct evidence hit
- **THEN** the merged summary SHALL include a bounded direct-hit excerpt
- **AND** it SHALL preserve all contributing `source_invocation_ids`
#### Scenario: Truncated preview still contains direct evidence
- **WHEN** a persisted output preview is marked truncated
- **AND** the preview contains a concrete direct evidence hit
- **THEN** the summary SHALL treat the row as successful evidence
- **AND** it SHALL NOT downgrade the row to no-evidence only because the raw output was truncated
### Requirement: Live trace counts SHALL distinguish model steps from evidence tool invocations
The persisted trace SHALL make it possible to audit model step counts separately from evidence tool invocation counts.
#### Scenario: Diagnosis session exposes aggregate counts
- **WHEN** a diagnosis session completes
- **THEN** `diagnosis_session.step_count` SHALL count persisted Agent model steps
- **AND** `diagnosis_session.tool_call_count` SHALL count persisted evidence-tool invocation rows
- **AND** helper workflow calls that are not evidence rows SHALL be auditable from agent steps or logs without inflating `tool_invocation`
@@ -227,3 +227,24 @@ The post-retrieval flow SHALL rerank vector candidates using deterministic rule-
- **THEN** the reranker MAY boost the candidate
- **AND** the evidence block SHALL record the hint as a hit reason
- **AND** the system SHALL NOT treat the L0 hint itself as fact evidence
### Requirement: Lookup knowledge SHALL persist modular RAG trace details
The `lookup_knowledge` tool SHALL persist modular RAG pipeline details in `tool_invocation.retrieval_details` for every new lookup invocation.
#### Scenario: Evidence lookup persists modular detail keys
- **WHEN** `lookup_knowledge` returns usable evidence
- **THEN** `retrieval_details` SHALL include `query_transform`
- **AND** it SHALL include `retrieval_trace`
- **AND** it SHALL include `context_pack_summary`
- **AND** it SHALL include `rerank_trace`
- **AND** it SHALL include `evidence_blocks`
#### Scenario: Fallback lookup persists fallback reason
- **WHEN** `lookup_knowledge` performs an unfiltered retry after filtered retrieval fails or is low quality
- **THEN** `retrieval_details.retrieval_trace` SHALL include the selected attempt
- **AND** `retrieval_details.fallback_reason` SHALL preserve the fallback reason
#### Scenario: No-evidence lookup still preserves trace
- **WHEN** `lookup_knowledge` returns no usable evidence
- **THEN** `retrieval_details` SHALL still include retrieval trace information for attempted retrieval paths
- **AND** it SHALL not include full evidence content as duplicated trace data