fix(agent): harden live diagnosis skill observability

This commit is contained in:
aruo
2026-07-07 00:20:34 +08:00
parent 64adb998cf
commit b3315ead52
20 changed files with 792 additions and 26 deletions
@@ -204,3 +204,29 @@ The verifier integration SHALL continue to work when evidence summaries distingu
- **THEN** those entries SHALL be treated as no-new-evidence
- **AND** they SHALL NOT be interpreted as fresh direct support for the answer
### Requirement: Verifier evidence summaries SHALL preserve concrete supporting facts
The verifier-facing `tool_trace_summary` SHALL preserve compact concrete facts from persisted evidence-tool outputs so direct evidence is not misclassified as missing merely because raw output was truncated.
#### Scenario: Log evidence contains a concrete matching message
- **WHEN** a persisted `query_logs` invocation output contains a concrete log message matching a critical fact
- **THEN** the generated `tool_trace_summary` SHALL include that message or a bounded excerpt of it in `output_summary`
- **AND** Verifier SHALL be able to reference the invocation id as direct evidence
#### Scenario: Metrics evidence contains concrete alert fields
- **WHEN** a persisted `query_metrics` invocation output contains alert names, services, or metric values
- **THEN** the generated `tool_trace_summary` SHALL include the relevant alert names, services, and bounded metric values
- **AND** it SHALL NOT imply unsupported alerts that are absent from the tool output
#### Scenario: Summary remains bounded
- **WHEN** a tool output is large
- **THEN** the generated `tool_trace_summary` SHALL remain bounded
- **AND** it SHALL preserve concrete facts before generic boilerplate or low-value formatting
### Requirement: Verifier low-confidence output SHALL not present unsupported claims as confirmed
When Verifier returns `LOW_CONFID`, user-facing output SHALL clearly separate confirmed facts from evidence gaps and SHALL NOT leave unsupported Executor claims formatted as confirmed findings.
#### Scenario: LOW_CONFID with critical evidence gaps
- **WHEN** Verifier labels critical facts as `no_evidence`
- **THEN** the final user-facing response SHALL identify those gaps from verifier output
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions