feat(agent): harden verifier evidence references
This commit is contained in:
@@ -108,3 +108,44 @@ The persisted trace SHALL make it possible to audit model step counts separately
|
||||
- **AND** `diagnosis_session.tool_call_count` SHALL count persisted evidence-tool invocation rows
|
||||
- **AND** helper workflow calls that are not evidence rows SHALL be auditable from agent steps or logs without inflating `tool_invocation`
|
||||
|
||||
### Requirement: Evidence tools SHALL persist minimal evidence refs
|
||||
Evidence-bearing tool invocations SHALL persist claim-addressable evidence references in `tool_invocation.retrieval_details.evidence_refs`.
|
||||
|
||||
#### Scenario: Metrics alerts produce evidence refs
|
||||
- **WHEN** a `query_metrics` invocation returns alert entries
|
||||
- **THEN** the persisted retrieval details SHALL include one `evidence_refs` item per usable alert
|
||||
- **AND** each item SHALL include `raw_path` formatted as `$.alerts[i]`
|
||||
- **AND** each item SHALL include bounded `text` containing concrete alert facts such as alert name, state, service, current value, and duration when available
|
||||
|
||||
#### Scenario: Logs produce evidence refs
|
||||
- **WHEN** a `query_logs` invocation returns log entries
|
||||
- **THEN** the persisted retrieval details SHALL include one `evidence_refs` item per usable log
|
||||
- **AND** each item SHALL include `raw_path` formatted as `$.logs[i]`
|
||||
- **AND** each item SHALL include bounded `text` containing concrete log facts such as timestamp, level, service, and message when available
|
||||
|
||||
#### Scenario: Knowledge lookup produces evidence refs
|
||||
- **WHEN** a `lookup_knowledge` invocation returns evidence blocks
|
||||
- **THEN** the persisted retrieval details SHALL include one `evidence_refs` item per usable evidence block
|
||||
- **AND** each item SHALL include `raw_path` formatted as `$.evidence_blocks[i]`
|
||||
- **AND** each item SHALL include bounded `text` containing concrete block content, title, or source when available
|
||||
|
||||
#### Scenario: Evidence ref extraction does not infer diagnosis
|
||||
- **WHEN** the recorder creates `evidence_refs`
|
||||
- **THEN** it SHALL only copy or format concrete tool output fields
|
||||
- **AND** it SHALL NOT infer root cause, remediation, or diagnosis conclusions
|
||||
|
||||
### Requirement: Log mock no-hit semantics SHALL avoid placeholder evidence
|
||||
The log query mock SHALL distinguish positive mock evidence from no-hit results without using placeholder service logs as evidence.
|
||||
|
||||
#### Scenario: HikariCP positive query returns order-service pool evidence
|
||||
- **WHEN** a `query_logs` request targets `order-service` and HikariCP connection-pool exhaustion terms
|
||||
- **THEN** the tool SHALL return order-service HikariCP-related log entries
|
||||
- **AND** the returned evidence SHALL include concrete terms such as `HikariPool`, `active=50/50`, `waiting`, or `request timed out after 30000ms`
|
||||
- **AND** it SHALL NOT return `generic-service` placeholder logs
|
||||
|
||||
#### Scenario: HikariCP no-hit query returns no evidence
|
||||
- **WHEN** a `query_logs` request targets a service without matching HikariCP mock evidence
|
||||
- **THEN** the tool SHALL return an empty `logs` array
|
||||
- **AND** it SHALL mark the output as `evidence_status=no_evidence`
|
||||
- **AND** it SHALL NOT return `generic-service` placeholder logs
|
||||
|
||||
|
||||
Reference in New Issue
Block a user