fix(agent): harden live diagnosis skill observability
This commit is contained in:
@@ -204,3 +204,29 @@ The verifier integration SHALL continue to work when evidence summaries distingu
|
||||
- **THEN** those entries SHALL be treated as no-new-evidence
|
||||
- **AND** they SHALL NOT be interpreted as fresh direct support for the answer
|
||||
|
||||
### Requirement: Verifier evidence summaries SHALL preserve concrete supporting facts
|
||||
The verifier-facing `tool_trace_summary` SHALL preserve compact concrete facts from persisted evidence-tool outputs so direct evidence is not misclassified as missing merely because raw output was truncated.
|
||||
|
||||
#### Scenario: Log evidence contains a concrete matching message
|
||||
- **WHEN** a persisted `query_logs` invocation output contains a concrete log message matching a critical fact
|
||||
- **THEN** the generated `tool_trace_summary` SHALL include that message or a bounded excerpt of it in `output_summary`
|
||||
- **AND** Verifier SHALL be able to reference the invocation id as direct evidence
|
||||
|
||||
#### Scenario: Metrics evidence contains concrete alert fields
|
||||
- **WHEN** a persisted `query_metrics` invocation output contains alert names, services, or metric values
|
||||
- **THEN** the generated `tool_trace_summary` SHALL include the relevant alert names, services, and bounded metric values
|
||||
- **AND** it SHALL NOT imply unsupported alerts that are absent from the tool output
|
||||
|
||||
#### Scenario: Summary remains bounded
|
||||
- **WHEN** a tool output is large
|
||||
- **THEN** the generated `tool_trace_summary` SHALL remain bounded
|
||||
- **AND** it SHALL preserve concrete facts before generic boilerplate or low-value formatting
|
||||
|
||||
### Requirement: Verifier low-confidence output SHALL not present unsupported claims as confirmed
|
||||
When Verifier returns `LOW_CONFID`, user-facing output SHALL clearly separate confirmed facts from evidence gaps and SHALL NOT leave unsupported Executor claims formatted as confirmed findings.
|
||||
|
||||
#### Scenario: LOW_CONFID with critical evidence gaps
|
||||
- **WHEN** Verifier labels critical facts as `no_evidence`
|
||||
- **THEN** the final user-facing response SHALL identify those gaps from verifier output
|
||||
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
|
||||
|
||||
|
||||
Reference in New Issue
Block a user