feat(agent): harden verifier evidence references

This commit is contained in:
aruo
2026-07-08 16:12:56 +08:00
parent a08672b31e
commit 7b8c75e571
29 changed files with 3426 additions and 94 deletions
+1
View File
@@ -10,6 +10,7 @@
| 2026-07-07 | executor-gatekeeper-hook | Chat质量门禁/证据归因 | Gatekeeper, verifier payload, source_invocation_ids, tool_name match, self_evaluation | openspec/changes/archive/2026-07-07-executor-gatekeeper-hook | archived |
| 2026-07-07 | executor-verifier-claim-checks | Chat质量门禁/证据归因 | Verifier claim_checks, facts_checked compatibility, effective verdict guardrail, malformed output downgrade | openspec/changes/archive/2026-07-07-executor-verifier-claim-checks | archived |
| 2026-07-08 | executor-composer-final-answer | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
| 2026-07-08 | verifier-evidence-reference-fidelity | Chat质量门禁/证据归因 | evidence_refs, raw_path, Gatekeeper severity, verifier evidence excerpt, HikariCP mock, no_evidence | openspec/changes/archive/2026-07-08-verifier-evidence-reference-fidelity | archived |
| 2026-07-06 | rag-eval-pipeline-closure | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
| 2026-07-06 | modular-rag-pipeline | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
@@ -0,0 +1,52 @@
# Acceptance
## Static Verification
- `openspec validate verifier-evidence-reference-fidelity --strict`: passed.
- `openspec validate --specs`: passed.
## Script Verification
- `mvn "-Dtest=ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,QueryLogsToolsTest,ChatServiceSequentialAgentTest" test`
- Passed: 38 tests.
- `mvn "-Dtest=ExecutorGatekeeperServiceTest" test`
- Passed: 9 tests.
- `mvn test`
- Failed on unrelated environment-gated `MilvusConnectionTest.connect`: `MILVUS_TOKEN` environment variable was not set.
- Other executed tests in the run progressed until that single failure; focused tests for this change passed.
## End-to-End Verification
The Java service was restarted with `mvn spring-boot:run`; logs were written under `logs/`.
| Case | Session | Result | Gatekeeper Audit |
|---|---|---|---|
| HikariCP positive | `iss007-hikari-positive-20260708-1553` | PASS; confirmed order-service HikariCP timeout and pool saturation logs | `pass / none`, checked_bindings=2 |
| HikariCP negative | `iss007-hikari-negative-20260708-1555` | LOW_CONFID; no `generic-service`; no false positive for inventory-service | `fail / low_confid` |
| HighMemoryUsage positive | `iss007-memory-positive-20260708-1558` | PASS; confirmed HighMemoryUsage 91%, did not confirm memory leak | `pass / none`, checked_bindings=1 |
| SlowResponse positive | `iss007-slow-positive-20260708-1600` | PASS; confirmed SlowResponse and slow request logs, no DB pool root cause | `pass / none`, checked_bindings=7 |
| Narrow HighCPUUsage | `iss007-narrow-highcpu-20260708-1602` | PASS; only covered payment-service HighCPUUsage | `pass / none`, checked_bindings=1 |
## Database Audit
`scripts/query_mysql.py` was used to verify:
- `diagnosis_session.self_evaluation.verifier_evaluation.verdict`
- `gatekeeper_result.status`
- `gatekeeper_result.severity`
- `gatekeeper_result.checked_bindings`
- no-hit HikariCP query rows persist `evidence_status=no_evidence`
## Remaining Risk
- Negative no-hit claims still have incomplete precise references when Executor uses `$.logs` for empty arrays. Gatekeeper correctly downgrades to `LOW_CONFID`.
- Prompt-only scope control is improved but not a hard contract. A future `scope_contract` may still be needed.
- Full test suite requires `MILVUS_TOKEN` to pass `MilvusConnectionTest`.
## OpenSpec Archive
- `openspec archive verifier-evidence-reference-fidelity --yes`: succeeded.
- Main specs updated:
- `openspec/specs/chat-verifier-agent/spec.md`
- `openspec/specs/evidence-trace-hardening/spec.md`
- Non-blocking warning: proposal did not use OpenSpec's preferred `## Why` / `## What Changes` headers, but archive completed.
@@ -0,0 +1,44 @@
# Verifier Evidence Reference Fidelity
## Background
ISS-007 came from end-to-end diagnosis cases where raw tool output and Executor `evidence_excerpt` contained enough facts, but Verifier still returned `LOW_CONFID` because the verifier-facing summary compressed away key details.
The affected flow is:
```text
chat_planner
-> chat_executor
-> VerifierInputHook / Gatekeeper
-> chat_verifier
-> chat_composer
```
The change hardens the evidence handoff between Executor, Gatekeeper, and Verifier.
## Goal
Make Executor cite concrete tool evidence, make Gatekeeper validate that citation with code, and make Verifier judge whether verified evidence can derive the claim.
## Scope
- Persist `tool_invocation.retrieval_details.evidence_refs`.
- Use `source_invocation_id + raw_path + evidence_excerpt` as the precise evidence binding.
- Add Gatekeeper `severity` and checked binding audit.
- Keep Gatekeeper in the Verifier hook path.
- Keep `tool_trace_summary` as navigation/audit context, not the only evidence source.
- Fix HikariCP mock positive/no-hit behavior.
- Tighten Executor/Verifier prompts for narrow-scope evidence handling.
## Non-goals
- No Planner `scope_contract`.
- No new database table.
- No full JSONPath engine.
- No change to external HTTP API.
- No retry rollback from Gatekeeper to Executor in this phase.
## OpenSpec
- Change: `openspec/changes/verifier-evidence-reference-fidelity`
- Source issue: `mvp/issues/ISS-007-verifier-evidence-summary-fidelity.md`
@@ -0,0 +1,54 @@
# Decisions
## Evidence Reference
Use `source_invocation_id + raw_path + evidence_excerpt` as the precise evidence reference for Executor claim bindings.
Reason:
- Invocation ID alone only identifies a tool call, not the evidence inside it.
- `raw_path` is enough for the first version when paired with `retrieval_details.evidence_refs`.
- `evidence_excerpt` remains the text Verifier reads, but only after Gatekeeper validates it.
## Raw Path
Only support stable locators in the first version:
- `$.alerts[i]`
- `$.logs[i]`
- `$.evidence_blocks[i]`
No full JSONPath engine is introduced.
## Gatekeeper Severity
Gatekeeper output includes:
- `status`
- `severity`
- `checked_bindings`
- `failed_rules`
- `warnings`
- `errors`
Severity meaning:
- `none`: precise references passed.
- `low_confid`: evidence is missing or incomplete, but not fabricated.
- `reject`: fabricated ID, wrong tool, unknown raw path, or mismatched excerpt.
## Verifier Boundary
Verifier uses verified claim-local excerpts as primary derivability evidence. `tool_trace_summary` remains available for navigation and audit, but no longer needs to carry every concrete fact.
## Hook Placement
Gatekeeper remains in the Verifier input hook path. This version does not retry Executor on Gatekeeper failure.
## Planner
Planner is not changed. `scope_contract` remains a later-stage idea. This phase uses prompt constraints to reduce narrow-scope over-expansion.
## Database
No new tables. Evidence refs are stored in `tool_invocation.retrieval_details.evidence_refs`; audit is stored in `diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result`.
@@ -0,0 +1,33 @@
# Evidence
## Existing Context
- Existing `chat-verifier-agent` spec still used `source_invocation_ids` and `tool_trace_summary` as the main verifier evidence context.
- Existing `evidence-trace-hardening` spec already established `tool_invocation.retrieval_details` as the right place for structured tool-specific facts.
- Prior devflow projects established that runbook/skill content is guidance, not incident evidence.
## Code Findings
- `VerifierInputHook` previously backfilled plural `source_invocation_ids` from `tool_trace_summary` by tool name.
- `ExecutorGatekeeperService` previously validated invocation existence and tool name, but not `raw_path` or excerpt authenticity.
- `ToolInvocationRecorder` persisted retrieval details but did not generate claim-addressable `evidence_refs`.
- `QueryLogsTools` could fall back to `generic-service` placeholder logs on no-hit.
## Implementation Evidence
- `ToolInvocationRecorder` now extracts:
- `$.alerts[i]` for `query_metrics`
- `$.logs[i]` for `query_logs`
- `$.evidence_blocks[i]` for `lookup_knowledge`
- `ExecutorGatekeeperService` now validates:
- invocation existence
- tool name
- raw path presence
- `retrieval_details.evidence_refs`
- excerpt similarity/support
- `VerifierInputHook` only auto-fills a singular `source_invocation_id` when exactly one candidate exists and never invents `raw_path`.
- `QueryLogsTools` returns HikariCP mock logs for `order-service` and returns empty no-hit results for unrelated services.
## Residual Finding
The HikariCP negative E2E no longer has generic-service pollution, but the model still issued an extra broad HikariCP query without the service filter and used order-service as context. This is a remaining narrow-scope behavior issue, not a mock evidence pollution issue.
File diff suppressed because it is too large Load Diff
+1
View File
@@ -8,6 +8,7 @@
| ISS-004 | Executor 域级检索水位控制(Phase 2) | 低 | 待规划 | [ISS-004-executor-domain-hard-limit.md](ISS-004-executor-domain-hard-limit.md) |
| ISS-005 | 证据链补齐与降级契约收敛 | 高 | 已归档 | [ISS-005-evidence-trace-hardening.md](ISS-005-evidence-trace-hardening.md) |
| ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) |
| ISS-007 | Verifier 证据摘要保真与工具命中质量问题 | 高 | 已实施 | [ISS-007-verifier-evidence-summary-fidelity.md](ISS-007-verifier-evidence-summary-fidelity.md) |
| executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [executor-evidence-attribution-hallucination.md](executor-evidence-attribution-hallucination.md) |
| executor-self-evidence-loop-design-note | Executor 自证循环与证据摘要链路设计记录 | 高 | 已形成方向 | [executor-self-evidence-loop-design-note.md](executor-self-evidence-loop-design-note.md) |
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) |
@@ -0,0 +1,73 @@
# Decisions
## Discover Context
- Source issue: `mvp/issues/ISS-007-verifier-evidence-summary-fidelity.md`.
- Related existing specs: `chat-verifier-agent`, `evidence-trace-hardening`.
- Related devflow records: `executor-gatekeeper-hook`, `executor-verifier-claim-checks`, `executor-composer-final-answer`, `evidence-trace-hardening`.
- Current repo instruction requested semantic code search and LSP confirmation before code changes; those tools are not exposed in this environment, so implementation will use `rg`, direct code reading, and focused tests as fallback evidence.
## Classification
- sm-flow scale: `standard`.
- Reason: internal contract change across recorder, hook, Gatekeeper, verifier prompt, mock tools, tests, and E2E validation.
## Question Pool
| Question | Mode | Resolution |
|---|---|---|
| Should Planner change or add `scope_contract`? | user-interview, already resolved in issue | No. This change is prompt-first and does not alter Planner. |
| Should Gatekeeper move to Executor hook for retry? | user-interview, already resolved in issue | No. Gatekeeper remains in Verifier hook path for this version. |
| Should new DB tables be added for evidence refs or audit? | user-interview, already resolved in issue | No. Use `tool_invocation.retrieval_details.evidence_refs` and existing `self_evaluation.verifier_evaluation.gatekeeper_result`. |
| Should `tool_trace_summary` remain the primary evidence source? | user-interview, already resolved in issue | No. It becomes navigation/audit context; verified claim-local excerpts become primary evidence for derivability. |
| Should full JSONPath be supported? | user-interview, already resolved in issue | No. Only stable raw paths listed in design are supported. |
## Cross-artifact Alignment
| Check | Status |
|---|---|
| Issue background and target -> proposal | aligned |
| Proposal scope and non-goals -> design | aligned |
| Design protocol and risks -> specs/tasks | aligned |
| Specs observable behavior -> tasks acceptance | aligned |
## Architecture Audit
The change keeps the existing agent orchestration and only hardens the evidence payload between Executor, Gatekeeper, and Verifier. The highest coupling risk is transition compatibility from plural `source_invocation_ids` to singular `source_invocation_id`; the implementation must accept old data but only allow precise new bindings to pass. No new DB table is introduced, reducing migration risk. Verifier prompt and effective verdict guardrails must agree on `gatekeeper_result.severity`, otherwise the model could still produce a PASS that runtime later must downgrade.
## Commit Gate
Passed on 2026-07-08.
- `openspec validate verifier-evidence-reference-fidelity --strict`: passed.
- `openspec validate --specs`: passed.
- File completeness: proposal, design, specs, tasks present.
- Consistency: proposal -> design -> specs -> tasks aligned.
- Apply authorization: user requested continuing implementation and fixing autonomously; proceed to Apply.
## Apply Verification
Focused tests:
- `mvn "-Dtest=ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,QueryLogsToolsTest,ChatServiceSequentialAgentTest" test`: passed, 38 tests.
- `mvn "-Dtest=ExecutorGatekeeperServiceTest" test`: passed, 9 tests after adding the old-invocation-without-evidence-refs case.
OpenSpec:
- `openspec validate verifier-evidence-reference-fidelity --strict`: passed.
- `openspec validate --specs`: passed.
End-to-end sessions:
| Case | Session | Result | Gatekeeper |
|---|---|---|---|
| HikariCP positive | `iss007-hikari-positive-20260708-1553` | `PASS`; answer confirmed order-service HikariCP timeout and pool saturation logs | `pass / none`, checked_bindings=2 |
| HikariCP negative | `iss007-hikari-negative-20260708-1555` | `LOW_CONFID`; no `generic-service` pollution; no positive inventory-service confirmation | `fail / low_confid` due incomplete negative evidence references |
| HighMemoryUsage positive | `iss007-memory-positive-20260708-1558` | `PASS`; answer confirmed HighMemoryUsage 91% without confirming memory leak | `pass / none`, checked_bindings=1 |
| SlowResponse positive | `iss007-slow-positive-20260708-1600` | `PASS`; answer confirmed SlowResponse and slow request logs without DB-pool root cause | `pass / none`, checked_bindings=7 |
| Narrow HighCPUUsage | `iss007-narrow-highcpu-20260708-1602` | `PASS`; answer only covered payment-service HighCPUUsage | `pass / none`, checked_bindings=1 |
Residual observation:
- HikariCP negative no longer returns `generic-service`, and no-hit rows persist `evidence_status=no_evidence`.
- The model still issued an extra broad HikariCP query without the `inventory-service` filter and used order-service as a context claim. This is a remaining narrow-scope behavior issue, not a mock data pollution issue. It is acceptable for this change because the final verdict did not become a false positive for inventory-service.
@@ -0,0 +1,180 @@
# Design
## Data Flow
```text
Evidence tool result
-> ToolInvocationRecorder
persists tool_invocation.retrieval_details.evidence_refs
-> Executor
emits claim-local evidence_bindings
-> VerifierInputHook
preserves structured payload and runs Gatekeeper
-> ExecutorGatekeeperService
validates reference authenticity
-> Verifier
checks derivability from verified excerpts
-> ChatService / Composer
enforces effective verdict and safe final answer
```
## Evidence Reference Contract
`tool_invocation.retrieval_details.evidence_refs` is an array of minimal evidence references:
```json
{
"evidence_refs": [
{
"raw_path": "$.alerts[0]",
"text": "HighMemoryUsage firing, service=order-service, current=91%, duration=15m"
}
]
}
```
Supported `raw_path` formats in this change:
- `$.alerts[i]` for `query_metrics`
- `$.logs[i]` for `query_logs`
- `$.evidence_blocks[i]` for `lookup_knowledge`
Unsupported in this change:
- Deep JSONPath such as `$.alerts[1].description`
- Nested aliases such as `$.retrieval_details.evidence_blocks[0]`
- Filter expressions
- Tool-specific metadata beyond `raw_path` and `text`
## Executor Binding Contract
Executor V2 claim bindings should use:
```json
{
"tool_name": "query_metrics",
"source_invocation_id": 12345,
"raw_path": "$.alerts[0]",
"evidence_excerpt": "HighMemoryUsage firing, service=order-service, current=91%, duration=15m"
}
```
Compatibility:
- Legacy `source_invocation_ids` may still be read for transition.
- New precise validation requires singular `source_invocation_id` plus `raw_path`.
- Missing `raw_path` is `LOW_CONFID`, not `PASS`.
## Gatekeeper Severity
Gatekeeper output includes:
```json
{
"status": "fail",
"severity": "reject",
"checked_bindings": [],
"failed_rules": [],
"warnings": [],
"errors": []
}
```
Severity mapping:
- `status=pass`, `severity=none`: all checked bindings are authentic.
- `status=fail`, `severity=low_confid`: evidence is missing or incomplete but not fabricated.
- `status=fail`, `severity=reject`: Executor cites fabricated or mismatched evidence.
`reject` cases:
- Invocation ID does not exist.
- Invocation belongs to another session.
- Tool name mismatches persisted invocation.
- `raw_path` is not present in `retrieval_details.evidence_refs`.
- `evidence_excerpt` clearly mismatches the system-side `text`.
`low_confid` cases:
- Confirmed claim has empty `evidence_bindings`.
- `source_invocation_id` exists but `raw_path` is missing.
- Old invocation lacks `evidence_refs`.
- `evidence_excerpt` is too short or generic to compare.
- Output is incomplete without concrete fabricated IDs or paths.
## Verifier Behavior
Verifier should treat `gatekeeper_result.severity` as a hard boundary:
- `reject`: effective result cannot be `PASS`; fabricated-reference cases should become `REJECT`.
- `low_confid`: effective result cannot be `PASS`; missing-reference cases should become `LOW_CONFID`.
- `none`: Verifier judges whether claim text is derivable from verified excerpts.
`tool_trace_summary` remains available as navigation and audit context, but claim-local verified excerpts are the primary evidence for derivability.
## Hook Placement
Gatekeeper remains in the Verifier hook path for this change. Retry or rollback into Executor is not implemented in this stage.
Existing ad hoc verifier-hook validation should be replaced by Gatekeeper output. The hook may normalize compatibility fields, but it should not independently decide pass/fail outside Gatekeeper semantics.
## Auto-backfill
`VerifierInputHook` may only auto-fill a missing `source_invocation_id` when exactly one invocation candidate exists for the binding's `tool_name`.
Rules:
- Never bulk-fill multiple IDs.
- Never auto-fill `raw_path`.
- Add a warning when auto-fill occurs.
- Auto-filled binding without `raw_path` must remain `LOW_CONFID`.
## Prompt Constraints
Executor prompt changes are prompt-first, contract-later:
- Executor is an evidence collector and micro-fact extractor.
- Narrow-scope questions should usually produce one claim and at most two claims.
- Do not limit evidence binding count.
- Emit observation or negative observation, not root-cause certainty, for narrow confirmation questions.
- Runbook and skill content cannot become current incident facts.
- Recommended actions, if any, are evidence-collection next steps, not remediation actions.
## HikariCP Mock Behavior
`query_logs` should support positive mock hits for:
- `HikariCP`
- `HikariPool`
- `connection pool`
- `数据库连接池`
- `连接池耗尽`
- `active=50/50`
- `waiting`
- `request timed out after 30000ms`
- `order-service`
Positive output should include order-service HikariCP log records. No-hit should return `logs=[]` and `evidence_status=no_evidence`, not `generic-service` placeholder logs.
## Audit
Gatekeeper result must be persisted under:
```text
diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result
```
The minimum persisted fields are:
- `status`
- `severity`
- `checked_bindings`
- `failed_rules`
- `warnings`
- `errors`
## Interface Impact
Impact level: L2 internal contract change.
The change affects internal Agent payload JSON, persisted audit JSON, prompts, and test fixtures. It does not add public HTTP endpoints or new database tables.
@@ -0,0 +1,62 @@
# verifier-evidence-reference-fidelity
## Problem
Recent end-to-end checks show that some narrow diagnosis questions still become `LOW_CONFID` even when the raw tool output and Executor evidence excerpts contain enough concrete evidence. The failure is caused by evidence being compressed or lost before Verifier reasoning, plus mock log no-hit behavior that can return `generic-service` placeholder logs.
Current weak points:
- Verifier still relies too heavily on `tool_trace_summary.output_summary`.
- Executor evidence bindings identify tool invocations but do not precisely locate evidence inside the invocation output.
- Gatekeeper validates invocation IDs and tool names, but does not yet verify `raw_path` and excerpt fidelity.
- `query_logs` mock data cannot reliably produce a positive HikariCP connection-pool exhaustion path and can pollute no-hit results with placeholder logs.
- Narrow-scope Executor answers can still over-expand into unrelated claims.
## Proposed Change
Introduce a claim-local evidence reference protocol:
```text
tool raw output
-> ToolInvocationRecorder stores retrieval_details.evidence_refs
-> Executor outputs claim + source_invocation_id + raw_path + evidence_excerpt
-> Gatekeeper verifies that the reference is real
-> Verifier judges whether verified evidence can derive the claim
-> Composer only expresses Verifier-allowed material
```
The first implementation keeps the orchestration unchanged. Gatekeeper remains in the Verifier input hook path. Planner `scope_contract` is out of scope for this change.
## Scope
- Add minimal `retrieval_details.evidence_refs` extraction for `query_metrics`, `query_logs`, and `lookup_knowledge`.
- Tighten Executor evidence bindings to prefer singular `source_invocation_id`, stable `raw_path`, and `evidence_excerpt`.
- Extend Gatekeeper to validate `source_invocation_id + raw_path + evidence_excerpt`.
- Add Gatekeeper severity: `none`, `low_confid`, `reject`.
- Make Verifier consume verified `evidence_excerpt` as the primary claim-local evidence.
- Tighten `VerifierInputHook` auto-backfill: only unique invocation candidate, never `raw_path`, and no `PASS` without a precise reference.
- Fix HikariCP mock log matching and no-hit behavior.
- Update Executor and Verifier prompts for narrow-scope and derivability behavior.
- Add focused tests and end-to-end checks for the minimum acceptance matrix.
## Non-goals
- Do not change Planner output.
- Do not implement Planner `scope_contract`.
- Do not add new database tables.
- Do not implement a general JSONPath engine.
- Do not make `tool_trace_summary` the primary evidence source again.
- Do not allow Executor `diagnosis_summary` or `user_facing_answer` to re-enter the V2 contract.
## Context Constraints
- `tool_invocation.retrieval_details` is the preferred place for tool-specific structured details.
- `DiagnosisSession.selfEvaluation.verifier_evaluation.gatekeeper_result` is the existing audit container and must be preserved.
- `read_skill` / runbook guidance is not incident evidence.
- Existing Composer routing must keep raw Executor JSON out of normal user answers.
## Risks
- Existing tests or prompts may still assume plural `source_invocation_ids`.
- Some old invocations will not have `evidence_refs`; those must downgrade to `LOW_CONFID`, not `PASS`.
- Similarity checks must tolerate formatting changes without accepting unrelated text.
@@ -0,0 +1,110 @@
## ADDED Requirements
### Requirement: Executor evidence bindings SHALL support precise evidence references
Executor V2 evidence bindings SHALL support precise evidence references that locate evidence inside a persisted tool invocation.
#### Scenario: Precise evidence binding contains invocation path and excerpt
- **WHEN** Executor binds evidence to a claim
- **THEN** the binding SHOULD include singular `source_invocation_id`
- **AND** the binding SHOULD include `raw_path`
- **AND** the binding SHALL include `tool_name` and `evidence_excerpt`
- **AND** the `raw_path` SHALL be interpreted relative to the referenced tool invocation's `retrieval_details.evidence_refs`
#### Scenario: Legacy plural invocation ids remain compatibility only
- **WHEN** Executor emits legacy `source_invocation_ids`
- **THEN** the system MAY read them for compatibility
- **AND** they SHALL NOT be sufficient for a precise Gatekeeper pass without `raw_path`
### Requirement: Gatekeeper SHALL validate evidence reference fidelity
Gatekeeper SHALL validate that Executor evidence bindings point to real current-session evidence references before Verifier uses them as primary evidence.
#### Scenario: Valid precise binding passes
- **WHEN** a binding's `source_invocation_id` exists in the current session
- **AND** the binding's `tool_name` matches the persisted invocation
- **AND** the binding's `raw_path` exists in `retrieval_details.evidence_refs`
- **AND** the binding's `evidence_excerpt` is supported by the matching evidence ref text
- **THEN** Gatekeeper SHALL return `status=pass`
- **AND** Gatekeeper SHALL return `severity=none`
#### Scenario: Missing raw path is low confidence
- **WHEN** a binding references an existing invocation
- **AND** the binding omits `raw_path`
- **THEN** Gatekeeper SHALL return `status=fail`
- **AND** Gatekeeper SHALL return `severity=low_confid`
- **AND** the effective verifier result SHALL NOT be `PASS`
#### Scenario: Old invocation without evidence refs is low confidence
- **WHEN** a binding references an existing invocation
- **AND** the invocation does not contain `retrieval_details.evidence_refs`
- **THEN** Gatekeeper SHALL return `status=fail`
- **AND** Gatekeeper SHALL return `severity=low_confid`
- **AND** the effective verifier result SHALL NOT be `PASS`
#### Scenario: Unknown raw path is rejected
- **WHEN** a binding references an existing invocation
- **AND** the binding's `raw_path` is absent from that invocation's `retrieval_details.evidence_refs`
- **THEN** Gatekeeper SHALL return `status=fail`
- **AND** Gatekeeper SHALL return `severity=reject`
- **AND** `failed_rules` SHALL include `evidence.raw_path`
#### Scenario: Mismatched excerpt is rejected
- **WHEN** a binding references an existing invocation and raw path
- **AND** the binding's `evidence_excerpt` is not supported by the matching system-side evidence ref text
- **THEN** Gatekeeper SHALL return `status=fail`
- **AND** Gatekeeper SHALL return `severity=reject`
- **AND** `failed_rules` SHALL include `evidence.excerpt_mismatch`
### Requirement: Verifier SHALL use verified claim-local evidence for derivability
Verifier SHALL judge structured claims primarily against Gatekeeper-verified claim-local evidence excerpts.
#### Scenario: Verified excerpt supports direct observation
- **WHEN** `gatekeeper_result.severity=none`
- **AND** a claim's verified evidence excerpts directly contain the claim's concrete facts
- **THEN** Verifier MAY classify that claim as `direct_observation`
#### Scenario: Tool trace summary is navigation context
- **WHEN** `executor_structured_output.claims[].evidence_bindings` are available
- **THEN** Verifier SHALL use `tool_trace_summary` as navigation and audit context
- **AND** it SHALL NOT require `tool_trace_summary.output_summary` to contain every fact already present in verified claim-local evidence
### Requirement: Gatekeeper severity SHALL constrain effective verdict
Runtime effective verdict calculation SHALL treat Gatekeeper severity as a hard upper bound.
#### Scenario: Reject severity prevents PASS
- **WHEN** `gatekeeper_result.severity=reject`
- **AND** the Verifier model returns `verdict=PASS`
- **THEN** ChatService SHALL downgrade the effective verdict
- **AND** the effective verdict SHALL be `REJECT`
#### Scenario: Low confidence severity prevents PASS
- **WHEN** `gatekeeper_result.severity=low_confid`
- **AND** the Verifier model returns `verdict=PASS`
- **THEN** ChatService SHALL downgrade the effective verdict
- **AND** the effective verdict SHALL be `LOW_CONFID`
#### Scenario: Gatekeeper audit includes severity
- **WHEN** verifier evaluation is persisted
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result` SHALL include `status`, `severity`, `checked_bindings`, `failed_rules`, `warnings`, and `errors`
### Requirement: Verifier input hook SHALL only perform narrow compatibility backfill
The verifier input hook SHALL avoid converting broad tool summaries into precise evidence references.
#### Scenario: Unique invocation candidate may be backfilled
- **WHEN** an evidence binding omits `source_invocation_id`
- **AND** exactly one current-session invocation exists for the binding's `tool_name`
- **THEN** the hook MAY backfill `source_invocation_id`
- **AND** it SHALL add a Gatekeeper warning describing the auto-backfill
#### Scenario: Raw path is never backfilled
- **WHEN** an evidence binding omits `raw_path`
- **THEN** the hook SHALL NOT synthesize `raw_path`
- **AND** Gatekeeper SHALL treat the binding as not precise enough to pass
### Requirement: Executor prompt SHALL constrain narrow-scope over-expansion
The Executor prompt SHALL instruct Executor to keep narrow confirmation questions focused on observation-level claims.
#### Scenario: Narrow scope produces minimal observation claims
- **WHEN** the user asks to confirm one specific service, alert, or symptom
- **THEN** Executor SHOULD output the minimum necessary claims, normally one and at most two
- **AND** those claims SHALL be `observation` or `negative_observation` unless current-session evidence proves more
- **AND** Executor SHALL NOT emit unrelated root-cause, remediation, or excluded-topic claims as confirmed facts
@@ -0,0 +1,42 @@
## ADDED Requirements
### Requirement: Evidence tools SHALL persist minimal evidence refs
Evidence-bearing tool invocations SHALL persist claim-addressable evidence references in `tool_invocation.retrieval_details.evidence_refs`.
#### Scenario: Metrics alerts produce evidence refs
- **WHEN** a `query_metrics` invocation returns alert entries
- **THEN** the persisted retrieval details SHALL include one `evidence_refs` item per usable alert
- **AND** each item SHALL include `raw_path` formatted as `$.alerts[i]`
- **AND** each item SHALL include bounded `text` containing concrete alert facts such as alert name, state, service, current value, and duration when available
#### Scenario: Logs produce evidence refs
- **WHEN** a `query_logs` invocation returns log entries
- **THEN** the persisted retrieval details SHALL include one `evidence_refs` item per usable log
- **AND** each item SHALL include `raw_path` formatted as `$.logs[i]`
- **AND** each item SHALL include bounded `text` containing concrete log facts such as timestamp, level, service, and message when available
#### Scenario: Knowledge lookup produces evidence refs
- **WHEN** a `lookup_knowledge` invocation returns evidence blocks
- **THEN** the persisted retrieval details SHALL include one `evidence_refs` item per usable evidence block
- **AND** each item SHALL include `raw_path` formatted as `$.evidence_blocks[i]`
- **AND** each item SHALL include bounded `text` containing concrete block content, title, or source when available
#### Scenario: Evidence ref extraction does not infer diagnosis
- **WHEN** the recorder creates `evidence_refs`
- **THEN** it SHALL only copy or format concrete tool output fields
- **AND** it SHALL NOT infer root cause, remediation, or diagnosis conclusions
### Requirement: Log mock no-hit semantics SHALL avoid placeholder evidence
The log query mock SHALL distinguish positive mock evidence from no-hit results without using placeholder service logs as evidence.
#### Scenario: HikariCP positive query returns order-service pool evidence
- **WHEN** a `query_logs` request targets `order-service` and HikariCP connection-pool exhaustion terms
- **THEN** the tool SHALL return order-service HikariCP-related log entries
- **AND** the returned evidence SHALL include concrete terms such as `HikariPool`, `active=50/50`, `waiting`, or `request timed out after 30000ms`
- **AND** it SHALL NOT return `generic-service` placeholder logs
#### Scenario: HikariCP no-hit query returns no evidence
- **WHEN** a `query_logs` request targets a service without matching HikariCP mock evidence
- **THEN** the tool SHALL return an empty `logs` array
- **AND** it SHALL mark the output as `evidence_status=no_evidence`
- **AND** it SHALL NOT return `generic-service` placeholder logs
@@ -0,0 +1,79 @@
# Tasks
## 1. Evidence reference extraction
- [x] Add `evidence_refs` extraction in `ToolInvocationRecorder` for `query_metrics` alert arrays.
- [x] Add `evidence_refs` extraction in `ToolInvocationRecorder` for `query_logs` log arrays.
- [x] Add `evidence_refs` extraction in `ToolInvocationRecorder` for `lookup_knowledge` evidence blocks.
- [x] Add focused recorder tests for `$.alerts[i]`, `$.logs[i]`, and `$.evidence_blocks[i]`.
Acceptance:
- Persisted `retrieval_details` contains minimal `raw_path` and `text`.
- Extraction does not infer root cause or diagnosis.
## 2. Gatekeeper reference fidelity
- [x] Extend `ExecutorGatekeeperService` to read singular `source_invocation_id`, `raw_path`, and `evidence_excerpt`.
- [x] Keep legacy `source_invocation_ids` compatibility where needed, but require `raw_path` for precise pass.
- [x] Validate invocation existence, session ownership, tool name, `raw_path`, and excerpt similarity.
- [x] Add `severity` and checked binding details to Gatekeeper output.
- [x] Add tests for valid reference, missing `raw_path`, missing `evidence_refs`, unknown `raw_path`, mismatched excerpt, fabricated invocation ID, and tool mismatch.
Acceptance:
- Valid precise references pass.
- Missing precision downgrades to `LOW_CONFID`.
- Fabricated or mismatched references become `REJECT`.
## 3. Verifier input hook and prompts
- [x] Tighten `VerifierInputHook` auto-backfill to only unique single invocation candidates.
- [x] Prevent auto-filled bindings without `raw_path` from passing Gatekeeper.
- [x] Add auto-backfill warnings into Gatekeeper/audit output.
- [x] Update `chat-executor-prompt.md` with evidence-reference and narrow-scope constraints.
- [x] Update `chat-verifier-prompt.md` so verified excerpts are the primary derivability evidence.
- [x] Add focused hook and prompt-sensitive tests where practical.
Acceptance:
- Hook no longer bulk-fills invocation IDs.
- Verifier can see `gatekeeper_result.severity`.
- Executor is instructed to output precise references and avoid over-expansion.
## 4. Effective verdict and audit persistence
- [x] Ensure `severity=reject` prevents effective `PASS` and maps to `REJECT` when appropriate.
- [x] Ensure `severity=low_confid` prevents effective `PASS` and maps to `LOW_CONFID`.
- [x] Ensure `gatekeeper_result` with `severity` is persisted under `verifier_evaluation`.
- [x] Add ChatService or integration tests for effective verdict guardrails.
Acceptance:
- Gatekeeper fail cannot become final PASS.
- Audit JSON contains the minimum Gatekeeper fields.
## 5. HikariCP mock quality
- [x] Add positive HikariCP mock logs for `order-service`.
- [x] Support HikariCP synonym matching.
- [x] Remove `generic-service` placeholder evidence for no-hit cases.
- [x] Add tests for HikariCP positive and negative no-hit behavior.
Acceptance:
- Positive HikariCP query returns order-service logs.
- Negative HikariCP query returns `logs=[]` and `evidence_status=no_evidence`.
## 6. Verification
- [x] Run focused unit tests for recorder, Gatekeeper, hook, ChatService guardrails, and query log mock behavior.
- [x] Run `openspec validate verifier-evidence-reference-fidelity --strict`.
- [x] Run `openspec validate --specs`.
- [x] Start the Java project using `mvn spring-boot:run`.
- [x] Run end-to-end checks for HighMemoryUsage positive, SlowResponse positive, HikariCP positive, HikariCP negative, and narrow forbidden claim.
- [x] Query MySQL audit data with `scripts/query_mysql.py` to confirm persisted `gatekeeper_result`.
Acceptance:
- Minimum E2E matrix passes or any failure is classified as code issue, mock quality issue, retrieval/tool quality issue, or model nondeterminism with evidence.
+111 -1
View File
@@ -1,4 +1,4 @@
# chat-verifier-agent Specification
# chat-verifier-agent Specification
## Purpose
TBD - created by archiving change chat-verifier-agent. Update Purpose after archive.
@@ -119,6 +119,7 @@ The system SHALL use ChatService for explicit single-round `Planner -> Executor
- **THEN** the system SHALL output a degraded result indicating the answer cannot be reliably generated
- **AND** it SHALL NOT pass through the raw Executor answer
- **AND** it SHALL NOT include a root-cause conclusion
### Requirement: Verifier SHALL be observable
The Verifier's verdict and downstream final-answer composition SHALL be persisted for observability.
@@ -421,3 +422,112 @@ The system SHALL use Composer or fixed safe templates for PASS, LOW_CONFID, and
- **AND** it SHALL include only confirmed facts, evidence gaps, and next-step suggestions
- **AND** it SHALL NOT include unverified raw answer content
- **AND** it SHALL NOT include a root-cause conclusion
### Requirement: Executor evidence bindings SHALL support precise evidence references
Executor V2 evidence bindings SHALL support precise evidence references that locate evidence inside a persisted tool invocation.
#### Scenario: Precise evidence binding contains invocation path and excerpt
- **WHEN** Executor binds evidence to a claim
- **THEN** the binding SHOULD include singular `source_invocation_id`
- **AND** the binding SHOULD include `raw_path`
- **AND** the binding SHALL include `tool_name` and `evidence_excerpt`
- **AND** the `raw_path` SHALL be interpreted relative to the referenced tool invocation's `retrieval_details.evidence_refs`
#### Scenario: Legacy plural invocation ids remain compatibility only
- **WHEN** Executor emits legacy `source_invocation_ids`
- **THEN** the system MAY read them for compatibility
- **AND** they SHALL NOT be sufficient for a precise Gatekeeper pass without `raw_path`
### Requirement: Gatekeeper SHALL validate evidence reference fidelity
Gatekeeper SHALL validate that Executor evidence bindings point to real current-session evidence references before Verifier uses them as primary evidence.
#### Scenario: Valid precise binding passes
- **WHEN** a binding's `source_invocation_id` exists in the current session
- **AND** the binding's `tool_name` matches the persisted invocation
- **AND** the binding's `raw_path` exists in `retrieval_details.evidence_refs`
- **AND** the binding's `evidence_excerpt` is supported by the matching evidence ref text
- **THEN** Gatekeeper SHALL return `status=pass`
- **AND** Gatekeeper SHALL return `severity=none`
#### Scenario: Missing raw path is low confidence
- **WHEN** a binding references an existing invocation
- **AND** the binding omits `raw_path`
- **THEN** Gatekeeper SHALL return `status=fail`
- **AND** Gatekeeper SHALL return `severity=low_confid`
- **AND** the effective verifier result SHALL NOT be `PASS`
#### Scenario: Old invocation without evidence refs is low confidence
- **WHEN** a binding references an existing invocation
- **AND** the invocation does not contain `retrieval_details.evidence_refs`
- **THEN** Gatekeeper SHALL return `status=fail`
- **AND** Gatekeeper SHALL return `severity=low_confid`
- **AND** the effective verifier result SHALL NOT be `PASS`
#### Scenario: Unknown raw path is rejected
- **WHEN** a binding references an existing invocation
- **AND** the binding's `raw_path` is absent from that invocation's `retrieval_details.evidence_refs`
- **THEN** Gatekeeper SHALL return `status=fail`
- **AND** Gatekeeper SHALL return `severity=reject`
- **AND** `failed_rules` SHALL include `evidence.raw_path`
#### Scenario: Mismatched excerpt is rejected
- **WHEN** a binding references an existing invocation and raw path
- **AND** the binding's `evidence_excerpt` is not supported by the matching system-side evidence ref text
- **THEN** Gatekeeper SHALL return `status=fail`
- **AND** Gatekeeper SHALL return `severity=reject`
- **AND** `failed_rules` SHALL include `evidence.excerpt_mismatch`
### Requirement: Verifier SHALL use verified claim-local evidence for derivability
Verifier SHALL judge structured claims primarily against Gatekeeper-verified claim-local evidence excerpts.
#### Scenario: Verified excerpt supports direct observation
- **WHEN** `gatekeeper_result.severity=none`
- **AND** a claim's verified evidence excerpts directly contain the claim's concrete facts
- **THEN** Verifier MAY classify that claim as `direct_observation`
#### Scenario: Tool trace summary is navigation context
- **WHEN** `executor_structured_output.claims[].evidence_bindings` are available
- **THEN** Verifier SHALL use `tool_trace_summary` as navigation and audit context
- **AND** it SHALL NOT require `tool_trace_summary.output_summary` to contain every fact already present in verified claim-local evidence
### Requirement: Gatekeeper severity SHALL constrain effective verdict
Runtime effective verdict calculation SHALL treat Gatekeeper severity as a hard upper bound.
#### Scenario: Reject severity prevents PASS
- **WHEN** `gatekeeper_result.severity=reject`
- **AND** the Verifier model returns `verdict=PASS`
- **THEN** ChatService SHALL downgrade the effective verdict
- **AND** the effective verdict SHALL be `REJECT`
#### Scenario: Low confidence severity prevents PASS
- **WHEN** `gatekeeper_result.severity=low_confid`
- **AND** the Verifier model returns `verdict=PASS`
- **THEN** ChatService SHALL downgrade the effective verdict
- **AND** the effective verdict SHALL be `LOW_CONFID`
#### Scenario: Gatekeeper audit includes severity
- **WHEN** verifier evaluation is persisted
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result` SHALL include `status`, `severity`, `checked_bindings`, `failed_rules`, `warnings`, and `errors`
### Requirement: Verifier input hook SHALL only perform narrow compatibility backfill
The verifier input hook SHALL avoid converting broad tool summaries into precise evidence references.
#### Scenario: Unique invocation candidate may be backfilled
- **WHEN** an evidence binding omits `source_invocation_id`
- **AND** exactly one current-session invocation exists for the binding's `tool_name`
- **THEN** the hook MAY backfill `source_invocation_id`
- **AND** it SHALL add a Gatekeeper warning describing the auto-backfill
#### Scenario: Raw path is never backfilled
- **WHEN** an evidence binding omits `raw_path`
- **THEN** the hook SHALL NOT synthesize `raw_path`
- **AND** Gatekeeper SHALL treat the binding as not precise enough to pass
### Requirement: Executor prompt SHALL constrain narrow-scope over-expansion
The Executor prompt SHALL instruct Executor to keep narrow confirmation questions focused on observation-level claims.
#### Scenario: Narrow scope produces minimal observation claims
- **WHEN** the user asks to confirm one specific service, alert, or symptom
- **THEN** Executor SHOULD output the minimum necessary claims, normally one and at most two
- **AND** those claims SHALL be `observation` or `negative_observation` unless current-session evidence proves more
- **AND** Executor SHALL NOT emit unrelated root-cause, remediation, or excluded-topic claims as confirmed facts
@@ -108,3 +108,44 @@ The persisted trace SHALL make it possible to audit model step counts separately
- **AND** `diagnosis_session.tool_call_count` SHALL count persisted evidence-tool invocation rows
- **AND** helper workflow calls that are not evidence rows SHALL be auditable from agent steps or logs without inflating `tool_invocation`
### Requirement: Evidence tools SHALL persist minimal evidence refs
Evidence-bearing tool invocations SHALL persist claim-addressable evidence references in `tool_invocation.retrieval_details.evidence_refs`.
#### Scenario: Metrics alerts produce evidence refs
- **WHEN** a `query_metrics` invocation returns alert entries
- **THEN** the persisted retrieval details SHALL include one `evidence_refs` item per usable alert
- **AND** each item SHALL include `raw_path` formatted as `$.alerts[i]`
- **AND** each item SHALL include bounded `text` containing concrete alert facts such as alert name, state, service, current value, and duration when available
#### Scenario: Logs produce evidence refs
- **WHEN** a `query_logs` invocation returns log entries
- **THEN** the persisted retrieval details SHALL include one `evidence_refs` item per usable log
- **AND** each item SHALL include `raw_path` formatted as `$.logs[i]`
- **AND** each item SHALL include bounded `text` containing concrete log facts such as timestamp, level, service, and message when available
#### Scenario: Knowledge lookup produces evidence refs
- **WHEN** a `lookup_knowledge` invocation returns evidence blocks
- **THEN** the persisted retrieval details SHALL include one `evidence_refs` item per usable evidence block
- **AND** each item SHALL include `raw_path` formatted as `$.evidence_blocks[i]`
- **AND** each item SHALL include bounded `text` containing concrete block content, title, or source when available
#### Scenario: Evidence ref extraction does not infer diagnosis
- **WHEN** the recorder creates `evidence_refs`
- **THEN** it SHALL only copy or format concrete tool output fields
- **AND** it SHALL NOT infer root cause, remediation, or diagnosis conclusions
### Requirement: Log mock no-hit semantics SHALL avoid placeholder evidence
The log query mock SHALL distinguish positive mock evidence from no-hit results without using placeholder service logs as evidence.
#### Scenario: HikariCP positive query returns order-service pool evidence
- **WHEN** a `query_logs` request targets `order-service` and HikariCP connection-pool exhaustion terms
- **THEN** the tool SHALL return order-service HikariCP-related log entries
- **AND** the returned evidence SHALL include concrete terms such as `HikariPool`, `active=50/50`, `waiting`, or `request timed out after 30000ms`
- **AND** it SHALL NOT return `generic-service` placeholder logs
#### Scenario: HikariCP no-hit query returns no evidence
- **WHEN** a `query_logs` request targets a service without matching HikariCP mock evidence
- **THEN** the tool SHALL return an empty `logs` array
- **AND** it SHALL mark the output as `evidence_status=no_evidence`
- **AND** it SHALL NOT return `generic-service` placeholder logs
@@ -289,13 +289,9 @@ public class QueryLogsTools {
logs.addAll(buildSystemEventsLogs(now, normalizedQuery, limit));
break;
default:
logs.addAll(buildGenericLogs(now, normalizedQuery, limit));
return logs;
}
if (logs.isEmpty()) {
logs.addAll(buildGenericLogs(now, normalizedQuery, limit));
}
// 限制返回条数
if (logs.size() > limit) {
@@ -400,6 +396,10 @@ public class QueryLogsTools {
*/
private List<LogEntry> buildApplicationLogs(Instant now, String query, int limit) {
List<LogEntry> logs = new ArrayList<>();
if (isHikariPoolQuery(query) && targetsOrderService(query)) {
logs.addAll(buildHikariPoolLogs(now));
}
// ERROR 级别日志
if (query.contains("error") || query.contains("fatal") || query.contains("500")) {
@@ -514,6 +514,58 @@ public class QueryLogsTools {
return logs;
}
private boolean isHikariPoolQuery(String query) {
return query.contains("hikaricp")
|| query.contains("hikaripool")
|| query.contains("connection pool")
|| query.contains("数据库连接池")
|| query.contains("连接池耗尽")
|| query.contains("active=50/50")
|| query.contains("waiting")
|| query.contains("request timed out after 30000ms");
}
private boolean targetsOrderService(String query) {
if (query.contains("inventory-service") || query.contains("payment-service") || query.contains("user-service")) {
return false;
}
return query.contains("order-service") || isHikariPoolQuery(query);
}
private List<LogEntry> buildHikariPoolLogs(Instant now) {
List<LogEntry> logs = new ArrayList<>();
LogEntry timeout = new LogEntry();
timeout.setTimestamp(FORMATTER.format(now.minus(2, ChronoUnit.MINUTES)));
timeout.setLevel("ERROR");
timeout.setService("order-service");
timeout.setInstance("pod-order-service-5c7d8e9f1-m3n2p");
timeout.setMessage("HikariPool-1 - Connection is not available, request timed out after 30000ms");
timeout.setMetrics(Map.of(
"pool", "HikariPool-1",
"error_type", "ConnectionPoolExhaustedException",
"timeout_ms", "30000"
));
logs.add(timeout);
LogEntry stats = new LogEntry();
stats.setTimestamp(FORMATTER.format(now.minus(1, ChronoUnit.MINUTES)));
stats.setLevel("WARN");
stats.setService("order-service");
stats.setInstance("pod-order-service-5c7d8e9f1-m3n2p");
stats.setMessage("HikariCP pool stats: active=50/50, idle=0, waiting=32");
stats.setMetrics(Map.of(
"pool", "HikariPool-1",
"active", "50",
"max", "50",
"idle", "0",
"waiting", "32"
));
logs.add(stats);
return logs;
}
/**
* 构建数据库慢查询日志(与慢响应告警关联)
@@ -17,9 +17,12 @@ import org.springframework.ai.chat.messages.AssistantMessage;
import org.springframework.ai.chat.messages.Message;
import org.springframework.ai.chat.messages.UserMessage;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.LinkedHashSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
/**
* Replaces verifier history with an explicit structured payload.
@@ -60,14 +63,18 @@ public class VerifierInputHook extends MessagesModelHook {
executorFinalAnswer = extractLastAssistantText(previousMessages);
}
ExecutorOutputParseResult parseResult = parseExecutorOutput(executorFinalAnswer);
VerifierContextHolder.setExecutorStructuredOutput(parseResult.structuredOutput());
VerifierContextHolder.setExecutorOutputParseStatus(parseResult.status());
List<Map<String, Object>> toolTraceSummary =
toolTraceSummaryService.buildVerifierTraceSummary(sessionId, executorFinalAnswer);
VerifierContextHolder.setToolTraceSummary(toolTraceSummary);
ExecutorOutputParseResult parseResult = parseExecutorOutput(executorFinalAnswer);
parseResult = new ExecutorOutputParseResult(
enrichExecutorStructuredOutput(parseResult.structuredOutput(), toolTraceSummary),
parseResult.status()
);
VerifierContextHolder.setExecutorStructuredOutput(parseResult.structuredOutput());
VerifierContextHolder.setExecutorOutputParseStatus(parseResult.status());
Map<String, Object> gatekeeperResult = runGatekeeper(sessionId, parseResult);
VerifierContextHolder.setGatekeeperResult(gatekeeperResult);
@@ -105,6 +112,8 @@ public class VerifierInputHook extends MessagesModelHook {
private Map<String, Object> passGatekeeperResult() {
Map<String, Object> result = new LinkedHashMap<>();
result.put("status", "pass");
result.put("severity", "none");
result.put("checked_bindings", List.of());
result.put("failed_rules", List.of());
result.put("warnings", List.of());
result.put("errors", List.of());
@@ -135,6 +144,129 @@ public class VerifierInputHook extends MessagesModelHook {
}
}
@SuppressWarnings("unchecked")
private Map<String, Object> enrichExecutorStructuredOutput(Map<String, Object> structuredOutput,
List<Map<String, Object>> toolTraceSummary) {
if (structuredOutput == null) {
return null;
}
Map<String, List<Long>> invocationIdsByTool = invocationIdsByTool(toolTraceSummary);
List<Map<String, Object>> warnings = new ArrayList<>();
enrichEvidenceBindingsInSection(structuredOutput.get("claims"), invocationIdsByTool, warnings);
enrichEvidenceBindingsInSection(structuredOutput.get("recommended_actions"), invocationIdsByTool, warnings);
if (!warnings.isEmpty()) {
structuredOutput.put("_gatekeeper_warnings", warnings);
}
return structuredOutput;
}
@SuppressWarnings("unchecked")
private void enrichEvidenceBindingsInSection(Object sectionValue,
Map<String, List<Long>> invocationIdsByTool,
List<Map<String, Object>> warnings) {
if (!(sectionValue instanceof List<?> items)) {
return;
}
for (Object itemValue : items) {
if (!(itemValue instanceof Map<?, ?> item)) {
continue;
}
Object bindingsValue = item.get("evidence_bindings");
if (!(bindingsValue instanceof List<?> bindings)) {
continue;
}
for (Object bindingValue : bindings) {
if (!(bindingValue instanceof Map<?, ?> rawBinding)) {
continue;
}
Map<String, Object> binding = (Map<String, Object>) rawBinding;
String normalizedToolName = normalizeToolName(binding.get("tool_name"));
if (!normalizedToolName.isBlank()) {
binding.put("tool_name", normalizedToolName);
}
if (!hasInvocationId(binding)) {
List<Long> ids = invocationIdsByTool.getOrDefault(normalizedToolName, List.of());
if (ids.size() == 1) {
binding.put("source_invocation_id", ids.get(0));
warnings.add(Map.of(
"rule", "evidence.invocation_auto_backfill",
"message", "source_invocation_id was auto-filled from the unique tool invocation candidate; raw_path remains missing if Executor did not provide it",
"tool_name", normalizedToolName,
"source_invocation_id", ids.get(0)
));
}
}
}
}
}
private Map<String, List<Long>> invocationIdsByTool(List<Map<String, Object>> toolTraceSummary) {
Map<String, Set<Long>> idsByTool = new LinkedHashMap<>();
for (Map<String, Object> summary : toolTraceSummary == null ? List.<Map<String, Object>>of() : toolTraceSummary) {
String toolName = normalizeToolName(summary.get("tool_name"));
if (toolName.isBlank()) {
continue;
}
List<Long> ids = toLongList(summary.get("source_invocation_ids"));
if (ids.isEmpty()) {
continue;
}
idsByTool.computeIfAbsent(toolName, ignored -> new LinkedHashSet<>()).addAll(ids);
}
Map<String, List<Long>> result = new LinkedHashMap<>();
for (Map.Entry<String, Set<Long>> entry : idsByTool.entrySet()) {
result.put(entry.getKey(), new ArrayList<>(entry.getValue()));
}
return result;
}
private boolean hasInvocationId(Map<String, Object> binding) {
if (asLong(binding.get("source_invocation_id")) != null) {
return true;
}
return toLongList(binding.get("source_invocation_ids")).size() == 1;
}
private List<Long> toLongList(Object value) {
if (!(value instanceof List<?> values)) {
return List.of();
}
List<Long> ids = new ArrayList<>();
for (Object item : values) {
Long id = asLong(item);
if (id != null) {
ids.add(id);
}
}
return ids;
}
private Long asLong(Object value) {
if (value instanceof Number number) {
return number.longValue();
}
if (value instanceof String text) {
try {
return Long.parseLong(text);
} catch (NumberFormatException ignored) {
return null;
}
}
return null;
}
private String normalizeToolName(Object value) {
String toolName = value == null ? "" : String.valueOf(value);
return switch (toolName) {
case "lookupKnowledge" -> "lookup_knowledge";
case "queryLogs" -> "query_logs";
case "queryPrometheusAlerts" -> "query_metrics";
case "getAvailableLogTopics" -> "get_available_log_topics";
default -> toolName;
};
}
private String sanitizeJsonPayload(String raw) {
String trimmed = raw.trim();
int fenceStart = trimmed.indexOf("```");
@@ -711,6 +711,13 @@ public class ChatService {
if (gatekeeperResult == null || !"fail".equals(String.valueOf(gatekeeperResult.get("status")))) {
return verdict;
}
String severity = String.valueOf(gatekeeperResult.getOrDefault("severity", ""));
if (ExecutorGatekeeperService.SEVERITY_REJECT.equals(severity)) {
return "REJECT";
}
if (ExecutorGatekeeperService.SEVERITY_LOW_CONFID.equals(severity)) {
return "PASS".equals(verdict) ? "LOW_CONFID" : verdict;
}
if (containsRule(gatekeeperResult.get("failed_rules"), ExecutorGatekeeperService.RULE_INVOCATION_REF)) {
return "REJECT";
}
@@ -885,7 +892,8 @@ public class ChatService {
Optional.ofNullable(VerifierContextHolder.getToolTraceSummary()).orElse(List.of()));
verifierEvaluation.put("gatekeeper_result",
Optional.ofNullable(VerifierContextHolder.getGatekeeperResult())
.orElse(Map.of("status", "pass", "failed_rules", List.of(), "warnings", List.of(), "errors", List.of())));
.orElse(Map.of("status", "pass", "severity", "none", "checked_bindings", List.of(),
"failed_rules", List.of(), "warnings", List.of(), "errors", List.of())));
if (composerOutput != null) {
verifierEvaluation.put("composer_output", composerOutput);
}
@@ -1,5 +1,7 @@
package com.superbiz.agent.service;
import com.fasterxml.jackson.core.type.TypeReference;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.superbiz.agent.domain.entity.ToolInvocation;
import com.superbiz.agent.repository.ToolInvocationRepository;
import org.springframework.stereotype.Service;
@@ -21,12 +23,22 @@ import java.util.stream.Collectors;
public class ExecutorGatekeeperService {
public static final String STATUS_PASS = "pass";
public static final String STATUS_WARN = "warn";
public static final String STATUS_FAIL = "fail";
public static final String SEVERITY_NONE = "none";
public static final String SEVERITY_LOW_CONFID = "low_confid";
public static final String SEVERITY_REJECT = "reject";
public static final String RULE_SCHEMA = "schema.executor_v2";
public static final String RULE_INVOCATION_REF = "evidence.invocation_ref";
public static final String RULE_RAW_PATH = "evidence.raw_path";
public static final String RULE_EXCERPT_MISMATCH = "evidence.excerpt_mismatch";
public static final String RULE_EVIDENCE_MISSING = "evidence.missing";
private static final TypeReference<Map<String, Object>> MAP_TYPE = new TypeReference<>() {
};
private static final double MIN_TOKEN_OVERLAP = 0.5;
private final ToolInvocationRepository toolInvocationRepository;
private final ObjectMapper objectMapper = new ObjectMapper();
public ExecutorGatekeeperService(ToolInvocationRepository toolInvocationRepository) {
this.toolInvocationRepository = toolInvocationRepository;
@@ -39,6 +51,7 @@ public class ExecutorGatekeeperService {
validateSchema(structuredOutput, parseStatus, result);
if (structuredOutput != null) {
validateInvocationRefs(sessionId, structuredOutput, result);
importWarnings(structuredOutput, result);
}
return result.toMap();
}
@@ -49,7 +62,7 @@ public class ExecutorGatekeeperService {
public Map<String, Object> fail(String ruleId, String target, String message) {
GatekeeperResult result = new GatekeeperResult();
result.fail(ruleId, target, message);
result.fail(ruleId, target, message, SEVERITY_REJECT);
return result.toMap();
}
@@ -59,31 +72,35 @@ public class ExecutorGatekeeperService {
String status = parseStatus == null ? "" : String.valueOf(parseStatus.getOrDefault("status", ""));
if (structuredOutput == null) {
if ("valid".equals(status)) {
result.fail(RULE_SCHEMA, "executor_structured_output", "structured output is missing after valid parse");
result.fail(RULE_SCHEMA, "executor_structured_output",
"structured output is missing after valid parse", SEVERITY_LOW_CONFID);
}
return;
}
if (!"executor_evidence_v2".equals(String.valueOf(structuredOutput.get("answer_version")))) {
result.fail(RULE_SCHEMA, "answer_version", "answer_version must be executor_evidence_v2");
result.fail(RULE_SCHEMA, "answer_version", "answer_version must be executor_evidence_v2",
SEVERITY_LOW_CONFID);
}
if (structuredOutput.containsKey("diagnosis_summary")) {
result.fail(RULE_SCHEMA, "diagnosis_summary", "diagnosis_summary is removed from executor_evidence_v2");
result.fail(RULE_SCHEMA, "diagnosis_summary", "diagnosis_summary is removed from executor_evidence_v2",
SEVERITY_REJECT);
}
if (structuredOutput.containsKey("user_facing_answer")) {
result.fail(RULE_SCHEMA, "user_facing_answer", "user_facing_answer is removed from executor_evidence_v2");
result.fail(RULE_SCHEMA, "user_facing_answer", "user_facing_answer is removed from executor_evidence_v2",
SEVERITY_REJECT);
}
Object claimsValue = structuredOutput.get("claims");
if (!(claimsValue instanceof List<?> claims)) {
result.fail(RULE_SCHEMA, "claims", "claims must be an array");
result.fail(RULE_SCHEMA, "claims", "claims must be an array", SEVERITY_LOW_CONFID);
return;
}
for (int i = 0; i < claims.size(); i++) {
String target = "claims[" + i + "]";
Object claimValue = claims.get(i);
if (!(claimValue instanceof Map<?, ?> claim)) {
result.fail(RULE_SCHEMA, target, "claim must be an object");
result.fail(RULE_SCHEMA, target, "claim must be an object", SEVERITY_LOW_CONFID);
continue;
}
requireString(claim, "claim_id", target, result);
@@ -91,11 +108,13 @@ public class ExecutorGatekeeperService {
requireString(claim, "claim_text", target, result);
String supportLevel = stringValue(claim.get("support_level"));
if (!"direct".equals(supportLevel) && !"indirect".equals(supportLevel)) {
result.fail(RULE_SCHEMA, target + ".support_level", "support_level must be direct or indirect");
result.fail(RULE_SCHEMA, target + ".support_level", "support_level must be direct or indirect",
SEVERITY_LOW_CONFID);
}
Object bindings = claim.get("evidence_bindings");
if (!(bindings instanceof List<?> bindingList) || bindingList.isEmpty()) {
result.fail(RULE_SCHEMA, target + ".evidence_bindings", "claims must include non-empty evidence_bindings");
result.fail(RULE_EVIDENCE_MISSING, target + ".evidence_bindings",
"claims must include non-empty evidence_bindings", SEVERITY_LOW_CONFID);
}
}
@@ -106,7 +125,8 @@ public class ExecutorGatekeeperService {
private void validateInvocationRefs(String sessionId, Map<String, Object> structuredOutput, GatekeeperResult result) {
if (sessionId == null || sessionId.isBlank()) {
result.fail(RULE_INVOCATION_REF, "session_id", "session id is required to validate source_invocation_ids");
result.fail(RULE_INVOCATION_REF, "session_id", "session id is required to validate source_invocation_id",
SEVERITY_LOW_CONFID);
return;
}
@@ -132,62 +152,254 @@ public class ExecutorGatekeeperService {
String target = "claims[" + claimIndex + "].evidence_bindings[" + bindingIndex + "]";
Object bindingValue = bindings.get(bindingIndex);
if (!(bindingValue instanceof Map<?, ?> binding)) {
result.fail(RULE_INVOCATION_REF, target, "evidence binding must be an object");
result.fail(RULE_INVOCATION_REF, target, "evidence binding must be an object",
SEVERITY_LOW_CONFID);
continue;
}
validateBindingInvocationIds(binding, validInvocations, target, result);
validateEvidenceBinding(binding, validInvocations, target, claim.get("claim_id"), result);
}
}
Object actionsValue = structuredOutput.get("recommended_actions");
if (!(actionsValue instanceof List<?> actions)) {
return;
}
for (int actionIndex = 0; actionIndex < actions.size(); actionIndex++) {
Object actionValue = actions.get(actionIndex);
if (!(actionValue instanceof Map<?, ?> action)) {
continue;
}
Object bindingsValue = action.get("evidence_bindings");
if (!(bindingsValue instanceof List<?> bindings)) {
continue;
}
for (int bindingIndex = 0; bindingIndex < bindings.size(); bindingIndex++) {
String target = "recommended_actions[" + actionIndex + "].evidence_bindings[" + bindingIndex + "]";
Object bindingValue = bindings.get(bindingIndex);
if (!(bindingValue instanceof Map<?, ?> binding)) {
result.fail(RULE_INVOCATION_REF, target, "evidence binding must be an object",
SEVERITY_LOW_CONFID);
continue;
}
validateEvidenceBinding(binding, validInvocations, target, action.get("action_id"), result);
}
}
}
private void validateBindingInvocationIds(Map<?, ?> binding,
Map<Long, ToolInvocation> validInvocations,
String target,
GatekeeperResult result) {
Object idsValue = binding.get("source_invocation_ids");
if (!(idsValue instanceof List<?> ids) || ids.isEmpty()) {
result.fail(RULE_INVOCATION_REF, target + ".source_invocation_ids",
"source_invocation_ids must be a non-empty array");
private void validateEvidenceBinding(Map<?, ?> binding,
Map<Long, ToolInvocation> validInvocations,
String target,
Object ownerId,
GatekeeperResult result) {
Map<String, Object> checked = new LinkedHashMap<>();
checked.put("claim_id", ownerId == null ? "" : String.valueOf(ownerId));
checked.put("tool_name", stringValue(binding.get("tool_name")));
checked.put("source_invocation_id", binding.get("source_invocation_id"));
checked.put("raw_path", stringValue(binding.get("raw_path")));
Long id = singleInvocationId(binding);
if (id == null) {
checked.put("status", STATUS_FAIL);
checked.put("rule", RULE_INVOCATION_REF);
checked.put("message", "source_invocation_id is required");
result.checked(checked);
result.fail(RULE_INVOCATION_REF, target + ".source_invocation_id",
"source_invocation_id is required", SEVERITY_LOW_CONFID);
return;
}
checked.put("source_invocation_id", id);
String claimedToolName = stringValue(binding.get("tool_name"));
if (claimedToolName.isBlank()) {
result.fail(RULE_INVOCATION_REF, target + ".tool_name", "tool_name is required");
checked.put("status", STATUS_FAIL);
checked.put("rule", RULE_INVOCATION_REF);
checked.put("message", "tool_name is required");
result.checked(checked);
result.fail(RULE_INVOCATION_REF, target + ".tool_name", "tool_name is required", SEVERITY_LOW_CONFID);
return;
}
Set<Long> checkedIds = new HashSet<>();
for (Object idValue : ids) {
Long id = asLong(idValue);
if (id == null) {
result.fail(RULE_INVOCATION_REF, target + ".source_invocation_ids",
"source_invocation_ids must contain numeric ids");
continue;
}
if (!checkedIds.add(id)) {
continue;
}
ToolInvocation invocation = validInvocations.get(id);
if (invocation == null) {
result.fail(RULE_INVOCATION_REF, target, "source_invocation_ids not found in current session: " + id);
continue;
}
if (!claimedToolName.isBlank() && !Objects.equals(claimedToolName, invocation.getToolName())) {
result.fail(RULE_INVOCATION_REF, target + ".tool_name",
"tool_name does not match invocation " + id + ": expected " + invocation.getToolName());
ToolInvocation invocation = validInvocations.get(id);
if (invocation == null) {
checked.put("status", STATUS_FAIL);
checked.put("rule", RULE_INVOCATION_REF);
checked.put("message", "source_invocation_id not found in current session: " + id);
result.checked(checked);
result.fail(RULE_INVOCATION_REF, target,
"source_invocation_id not found in current session: " + id, SEVERITY_REJECT);
return;
}
if (!Objects.equals(claimedToolName, invocation.getToolName())) {
checked.put("status", STATUS_FAIL);
checked.put("rule", RULE_INVOCATION_REF);
checked.put("message", "tool_name does not match invocation " + id + ": expected " + invocation.getToolName());
result.checked(checked);
result.fail(RULE_INVOCATION_REF, target + ".tool_name",
"tool_name does not match invocation " + id + ": expected " + invocation.getToolName(),
SEVERITY_REJECT);
return;
}
String rawPath = stringValue(binding.get("raw_path"));
if (rawPath.isBlank()) {
checked.put("status", STATUS_FAIL);
checked.put("rule", RULE_RAW_PATH);
checked.put("message", "raw_path is required for precise evidence reference");
result.checked(checked);
result.fail(RULE_RAW_PATH, target + ".raw_path",
"raw_path is required for precise evidence reference", SEVERITY_LOW_CONFID);
return;
}
Map<String, String> refs = evidenceRefsByRawPath(invocation.getRetrievalDetails());
if (refs.isEmpty()) {
checked.put("status", STATUS_FAIL);
checked.put("rule", RULE_EVIDENCE_MISSING);
checked.put("message", "invocation has no retrieval_details.evidence_refs");
result.checked(checked);
result.fail(RULE_EVIDENCE_MISSING, target,
"invocation has no retrieval_details.evidence_refs", SEVERITY_LOW_CONFID);
return;
}
String matchedText = refs.get(rawPath);
if (matchedText == null) {
checked.put("status", STATUS_FAIL);
checked.put("rule", RULE_RAW_PATH);
checked.put("message", "raw_path not found in retrieval_details.evidence_refs");
result.checked(checked);
result.fail(RULE_RAW_PATH, target + ".raw_path",
"raw_path not found in retrieval_details.evidence_refs", SEVERITY_REJECT);
return;
}
checked.put("matched_text", matchedText);
String excerpt = stringValue(binding.get("evidence_excerpt"));
if (normalized(excerpt).length() < 8) {
checked.put("status", STATUS_FAIL);
checked.put("rule", RULE_EXCERPT_MISMATCH);
checked.put("message", "evidence_excerpt is too short to compare");
result.checked(checked);
result.fail(RULE_EXCERPT_MISMATCH, target + ".evidence_excerpt",
"evidence_excerpt is too short to compare", SEVERITY_LOW_CONFID);
return;
}
if (!isExcerptSupported(excerpt, matchedText)) {
checked.put("status", STATUS_FAIL);
checked.put("rule", RULE_EXCERPT_MISMATCH);
checked.put("message", "evidence_excerpt is not supported by matched evidence ref text");
result.checked(checked);
result.fail(RULE_EXCERPT_MISMATCH, target + ".evidence_excerpt",
"evidence_excerpt is not supported by matched evidence ref text", SEVERITY_REJECT);
return;
}
checked.put("status", STATUS_PASS);
result.checked(checked);
}
private Long singleInvocationId(Map<?, ?> binding) {
Long singular = asLong(binding.get("source_invocation_id"));
if (singular != null) {
return singular;
}
Object idsValue = binding.get("source_invocation_ids");
if (!(idsValue instanceof List<?> ids) || ids.size() != 1) {
return null;
}
return asLong(ids.get(0));
}
@SuppressWarnings("unchecked")
private void importWarnings(Map<String, Object> structuredOutput, GatekeeperResult result) {
Object warningsValue = structuredOutput.remove("_gatekeeper_warnings");
if (!(warningsValue instanceof List<?> warnings)) {
return;
}
for (Object warning : warnings) {
if (warning instanceof Map<?, ?> map) {
result.warn((Map<String, Object>) map);
} else if (warning != null) {
result.warn(Map.of("message", String.valueOf(warning)));
}
}
}
private Map<String, String> evidenceRefsByRawPath(String retrievalDetails) {
if (retrievalDetails == null || retrievalDetails.isBlank()) {
return Map.of();
}
try {
Map<String, Object> details = objectMapper.readValue(retrievalDetails, MAP_TYPE);
Object refsValue = details.get("evidence_refs");
if (!(refsValue instanceof List<?> refs)) {
return Map.of();
}
Map<String, String> result = new LinkedHashMap<>();
for (Object refValue : refs) {
if (!(refValue instanceof Map<?, ?> ref)) {
continue;
}
String rawPath = stringValue(ref.get("raw_path"));
String text = stringValue(ref.get("text"));
if (!rawPath.isBlank() && !text.isBlank()) {
result.put(rawPath, text);
}
}
return result;
} catch (Exception ignored) {
return Map.of();
}
}
private boolean isExcerptSupported(String excerpt, String matchedText) {
String normalizedExcerpt = normalized(excerpt);
String normalizedMatched = normalized(matchedText);
if (normalizedMatched.contains(normalizedExcerpt) || normalizedExcerpt.contains(normalizedMatched)) {
return true;
}
Set<String> excerptTokens = tokens(normalizedExcerpt);
if (excerptTokens.isEmpty()) {
return false;
}
Set<String> matchedTokens = tokens(normalizedMatched);
int overlap = 0;
for (String token : excerptTokens) {
if (matchedTokens.contains(token)) {
overlap++;
}
}
return (double) overlap / excerptTokens.size() >= MIN_TOKEN_OVERLAP;
}
private String normalized(String value) {
return value == null ? "" : value.toLowerCase()
.replaceAll("[\\p{Punct}\\s,。;:、()【】《》“”‘’]+", " ")
.trim();
}
private Set<String> tokens(String text) {
if (text == null || text.isBlank()) {
return Set.of();
}
Set<String> result = new HashSet<>();
for (String token : text.split("\\s+")) {
if (token.length() >= 2) {
result.add(token);
}
}
return result;
}
private void requireArray(Map<String, Object> output, String field, GatekeeperResult result) {
if (!(output.get(field) instanceof List<?>)) {
result.fail(RULE_SCHEMA, field, field + " must be an array");
result.fail(RULE_SCHEMA, field, field + " must be an array", SEVERITY_LOW_CONFID);
}
}
private void requireString(Map<?, ?> object, String field, String target, GatekeeperResult result) {
if (stringValue(object.get(field)).isBlank()) {
result.fail(RULE_SCHEMA, target + "." + field, field + " is required");
result.fail(RULE_SCHEMA, target + "." + field, field + " is required", SEVERITY_LOW_CONFID);
}
}
@@ -211,23 +423,42 @@ public class ExecutorGatekeeperService {
private static final class GatekeeperResult {
private final List<String> failedRules = new ArrayList<>();
private final List<String> warnings = new ArrayList<>();
private final List<Map<String, Object>> checkedBindings = new ArrayList<>();
private final List<Map<String, Object>> warnings = new ArrayList<>();
private final List<Map<String, Object>> errors = new ArrayList<>();
private String severity = SEVERITY_NONE;
void fail(String ruleId, String target, String message) {
void fail(String ruleId, String target, String message, String failureSeverity) {
if (!failedRules.contains(ruleId)) {
failedRules.add(ruleId);
}
if (SEVERITY_REJECT.equals(failureSeverity)) {
severity = SEVERITY_REJECT;
} else if (!SEVERITY_REJECT.equals(severity)) {
severity = SEVERITY_LOW_CONFID;
}
Map<String, Object> error = new LinkedHashMap<>();
error.put("rule_id", ruleId);
error.put("rule", ruleId);
error.put("target", target);
error.put("message", message);
error.put("severity", failureSeverity);
errors.add(error);
}
void checked(Map<String, Object> checked) {
checkedBindings.add(checked);
}
void warn(Map<String, Object> warning) {
warnings.add(new LinkedHashMap<>(warning));
}
Map<String, Object> toMap() {
Map<String, Object> result = new LinkedHashMap<>();
result.put("status", failedRules.isEmpty() ? (warnings.isEmpty() ? STATUS_PASS : STATUS_WARN) : STATUS_FAIL);
result.put("status", failedRules.isEmpty() ? STATUS_PASS : STATUS_FAIL);
result.put("severity", failedRules.isEmpty() ? SEVERITY_NONE : severity);
result.put("checked_bindings", checkedBindings);
result.put("failed_rules", failedRules);
result.put("warnings", warnings);
result.put("errors", errors);
@@ -1,6 +1,7 @@
package com.superbiz.agent.service;
import com.fasterxml.jackson.core.JsonProcessingException;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.superbiz.agent.domain.entity.ToolInvocation;
import com.superbiz.agent.dto.ContextPack;
@@ -87,6 +88,10 @@ public class ToolInvocationRecorder {
if (extraDetails != null && !extraDetails.isEmpty()) {
details.putAll(extraDetails);
}
List<Map<String, Object>> evidenceRefs = extractEvidenceRefs(toolName, output, details);
if (!evidenceRefs.isEmpty()) {
details.put("evidence_refs", evidenceRefs);
}
ToolInvocation invocation = ToolInvocation.builder()
.toolName(toolName)
@@ -139,6 +144,10 @@ public class ToolInvocationRecorder {
details.put("evidence_block_count", record.evidenceBlockCount());
}
details.put("evidence_blocks", record.evidenceBlocks() == null ? List.of() : record.evidenceBlocks());
List<Map<String, Object>> evidenceRefs = evidenceRefsFromEvidenceBlocks(record.evidenceBlocks());
if (!evidenceRefs.isEmpty()) {
details.put("evidence_refs", evidenceRefs);
}
details.put("query_transform", record.queryTransform() == null ? Map.of() : record.queryTransform());
details.put("retrieval_trace", record.retrievalTrace() == null ? Map.of() : record.retrievalTrace());
details.put("context_pack_summary", record.contextPack() == null ? Map.of() : record.contextPack());
@@ -186,6 +195,132 @@ public class ToolInvocationRecorder {
return success ? EVIDENCE_STATUS_SUPPORTED : EVIDENCE_STATUS_FAILED;
}
private List<Map<String, Object>> extractEvidenceRefs(String toolName, String output, Map<String, Object> details) {
if (output == null || output.isBlank()) {
return List.of();
}
if ("lookup_knowledge".equals(toolName)) {
return evidenceRefsFromEvidenceBlocks(asMapList(details.get("evidence_blocks")));
}
try {
JsonNode root = objectMapper.readTree(output);
if ("query_metrics".equals(toolName)) {
return evidenceRefsFromArray(root.path("alerts"), "$.alerts", this::alertText);
}
if ("query_logs".equals(toolName)) {
return evidenceRefsFromArray(root.path("logs"), "$.logs", this::logText);
}
} catch (Exception e) {
log.debug("extract evidence_refs failed for tool={}", toolName, e);
}
return List.of();
}
private List<Map<String, Object>> evidenceRefsFromArray(JsonNode arrayNode,
String pathPrefix,
java.util.function.Function<JsonNode, String> textExtractor) {
if (!arrayNode.isArray()) {
return List.of();
}
List<Map<String, Object>> refs = new ArrayList<>();
for (int i = 0; i < arrayNode.size(); i++) {
String text = textExtractor.apply(arrayNode.get(i));
if (text == null || text.isBlank()) {
continue;
}
refs.add(Map.of(
"raw_path", pathPrefix + "[" + i + "]",
"text", bounded(text, 500)
));
}
return refs;
}
private String alertText(JsonNode alert) {
List<String> parts = new ArrayList<>();
addPart(parts, textField(alert, "alert_name"));
addPart(parts, textField(alert, "state"));
addPart(parts, textField(alert, "description"));
addPart(parts, "active_at=" + textField(alert, "active_at"));
addPart(parts, "duration=" + textField(alert, "duration"));
return String.join(", ", parts);
}
private String logText(JsonNode log) {
List<String> parts = new ArrayList<>();
addPart(parts, textField(log, "timestamp"));
addPart(parts, textField(log, "level"));
addPart(parts, textField(log, "service"));
addPart(parts, textField(log, "message"));
JsonNode metrics = log.path("metrics");
if (metrics.isObject() && !metrics.isEmpty()) {
addPart(parts, "metrics=" + metrics.toString());
}
return String.join(" ", parts);
}
private List<Map<String, Object>> evidenceRefsFromEvidenceBlocks(List<Map<String, Object>> blocks) {
if (blocks == null || blocks.isEmpty()) {
return List.of();
}
List<Map<String, Object>> refs = new ArrayList<>();
for (int i = 0; i < blocks.size(); i++) {
Map<String, Object> block = blocks.get(i);
String text = firstNonBlank(block.get("content_preview"), block.get("content"),
block.get("title"), block.get("source"));
if (text.isBlank()) {
continue;
}
refs.add(Map.of(
"raw_path", "$.evidence_blocks[" + i + "]",
"text", bounded(text, 500)
));
}
return refs;
}
@SuppressWarnings("unchecked")
private List<Map<String, Object>> asMapList(Object value) {
if (!(value instanceof List<?> list)) {
return List.of();
}
List<Map<String, Object>> result = new ArrayList<>();
for (Object item : list) {
if (item instanceof Map<?, ?> map) {
result.add((Map<String, Object>) map);
}
}
return result;
}
private String textField(JsonNode node, String field) {
JsonNode value = node.path(field);
return value.isMissingNode() || value.isNull() ? "" : value.asText("");
}
private void addPart(List<String> parts, String value) {
if (value != null && !value.isBlank() && !value.endsWith("=")) {
parts.add(value);
}
}
private String firstNonBlank(Object... values) {
for (Object value : values) {
if (value != null && !String.valueOf(value).isBlank()) {
return String.valueOf(value);
}
}
return "";
}
private String bounded(String value, int limit) {
if (value == null) {
return "";
}
return value.length() <= limit ? value : value.substring(0, limit) + "...";
}
private String preview(String output) {
if (output == null) {
return null;
@@ -5,6 +5,7 @@
- 需要外部信息时调用工具,但必须遵守下方的检索约束。
- 严禁凭记忆回答,必须基于本轮工具返回的真实数据。
- 执行完成后,输出严格的证据归因 JSON,供 Verifier 校验。
- 你是证据收集与微观事实提炼器,不是最终答复生成器。
## 规则
- 按顺序执行,不可跳过步骤。
@@ -12,6 +13,8 @@
- runbook、skill、历史案例、知识库中的通用模式只能作为排查指导或建议动作,不能直接写成本次事故的已确认事实。
- 如果检索内容不足以支撑结论,必须显式声明证据不足,严禁补全事故故事。
- 不要使用“通常情况下”“根据经验”“很可能已经发生”等无证据推断词来伪装事实。
- 对窄范围确认问题,只输出与用户问题直接相关的 observation / negative_observation。通常 1 条 claim,最多 2 条 claim;不要限制 evidence_bindings 数量。
- 禁止把根因、修复动作或用户明确排除的服务/主题写成 confirmed claim,除非本轮工具证据直接证明。
## 检索约束
@@ -59,6 +62,7 @@
### recommended_actions
`recommended_actions` 用来放下一步排查或修复动作。
建议可以来自 runbook/skill,但必须说明 reason,不能写成“已确认根因”。
本期 recommended_actions 只允许证据收集或继续排查动作,不要输出重启、扩容、修改配置等修复动作,除非用户明确要求执行方案。
### missing_info
`missing_info` 用来列出无法确认结论所缺少的具体证据。
@@ -80,9 +84,10 @@
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "工具返回中的 evidence block id、trace_ref 或可定位标识",
"tool_name": "lookup_knowledge/query_logs/query_metrics/read_skill 等",
"source_invocation_ids": [],
"source_id": "工具返回中的 evidence block id、trace_ref 或可定位标识,可为空",
"tool_name": "lookup_knowledge/query_logs/query_metrics 等 evidence tool",
"source_invocation_id": null,
"raw_path": "$.alerts[0] / $.logs[0] / $.evidence_blocks[0]",
"evidence_excerpt": "从工具返回中摘取的原话、指标值、日志片段或关键数据"
}
]
@@ -115,5 +120,8 @@
- `claims[*].support_level` 只能是 `direct` 或 `indirect`。
- `claims[*].evidence_bindings` 不能为空。
- `evidence_excerpt` 必须来自工具返回,不允许编造。
- `raw_path` 必须指向工具返回数组中的具体条目:`query_metrics` 使用 `$.alerts[i]`,`query_logs` 使用 `$.logs[i]`,`lookup_knowledge` 使用 `$.evidence_blocks[i]`。
- `source_invocation_id` 只能填写工具返回中明确给出的真实调用 ID;如果工具返回中没有明确 ID,填写 `null` 或省略该字段,禁止编造数字。系统只会在唯一候选工具调用存在时补齐 ID,但不会补齐 `raw_path`。
- 不要再输出 `source_invocation_ids` 作为主要字段;兼容旧字段不作为精确证据引用。
- 如果没有任何可确认事实,`claims` 返回空数组,并在 `missing_info` 说明缺少什么。
- 不要把其它服务、其它历史案例、其它会话的事实迁移为当前会话事实。
@@ -12,7 +12,7 @@
- `executor_final_answer`:Executor 原始输出,仅用于 debug/fallback;当结构化输出有效时,不得从这里抽取额外确认事实
- `executor_structured_output`:如果 Executor 输出了合法证据归因 JSON,这里会提供解析后的对象。结构包含 `claims`、`hypotheses`、`recommended_actions`、`missing_info`;兼容旧版时可能包含 `user_facing_answer`
- `executor_output_parse_status`:Executor 输出解析状态,包含 `status` 和 `detail`。`status` 可能是 `valid` / `missing` / `malformed`
- `tool_trace_summary`:基于真实工具调用整理出的证据索引。每一项都带有:
- `tool_trace_summary`:基于真实工具调用整理出的全局导航和审计索引。它不是唯一证据源;当 claim 有已核验的 `evidence_bindings[].evidence_excerpt` 时,应优先使用 claim-local excerpt 判断可推导性。每一项都带有:
- `trace_ref`
- `tool_name`
- `topic_domain`
@@ -20,7 +20,7 @@
- `input_summary`
- `output_summary`
- `evidence_level`
- `gatekeeper_result`:Executor 结构化输出的确定性校验结果,包含 `status`、`failed_rules`、`warnings`、`errors`
- `gatekeeper_result`:Executor 结构化输出的确定性校验结果,包含 `status`、`severity`、`checked_bindings`、`failed_rules`、`warnings`、`errors`
- `retry_context`:第二轮可选输入;若为空,按首轮处理
## 任务步骤
@@ -29,7 +29,8 @@
如果 `executor_output_parse_status.status="valid"` 且 `executor_structured_output.claims` 存在:
- 优先逐条校验 `executor_structured_output.claims`
- 每个 claim 至少形成一条 `claim_checks`
- 必须检查 claim 的 `evidence_bindings` 是否能对应到 `tool_trace_summary` 中真实存在的 trace、tool 或 source_invocation_ids
- 如果 `gatekeeper_result.severity="none"`,将 claim 的 `evidence_bindings[].evidence_excerpt` 视为已通过代码核验的主证据,判断 `claim_text` 是否能由这些 excerpt 推出
- `tool_trace_summary` 只用于理解工具调用全貌、补充 trace_ref、识别 no_evidence gap,不要求它逐字包含 excerpt 中已经核验过的全部事实
- 不得从 `executor_final_answer` 中抽取不在 claims 里的额外确认事实
如果 structured output 缺失或 malformed:
@@ -58,8 +59,8 @@
- `contradicted`
结构化 claim 的校验规则:
- claim 有真实 evidence binding,且工具摘要直接包含该事实 → `direct_observation`
- claim 有真实 evidence binding,工具摘要没有逐字说明但可以合理推出 → `reasonable_inference`
- claim 有 Gatekeeper 核验通过的 evidence binding,且 `evidence_excerpt` 直接包含该事实 → `direct_observation`
- claim 有 Gatekeeper 核验通过的 evidence binding,excerpt 没有逐字说明但可以合理推出 → `reasonable_inference`
- claim 有部分依据,但写成唯一根因、确认根因或说得过满 → `overstated`
- claim 无法绑定真实 trace、invocation 或 excerpt → `unsupported`
- claim 引入证据外的新服务名、订单号、错误码、指标值、根因 → `external_unknown`
@@ -86,8 +87,10 @@
严格使用以下判定矩阵:
0. 若 `gatekeeper_result.status="fail"`
- 不得输出 `PASS`
- 若 `failed_rules` 包含 `evidence.invocation_ref`,倾向 `REJECT`
- 否则至少输出 `LOW_CONFID`
- 若 `gatekeeper_result.severity="reject"`,输出 `REJECT`
- 若 `gatekeeper_result.severity="low_confid"`,输出 `LOW_CONFID`
- 兼容旧输入:若缺少 `severity` 且 `failed_rules` 包含 `evidence.invocation_ref`,倾向 `REJECT`
- 兼容旧输入:若缺少 `severity` 且不是明显伪造,至少输出 `LOW_CONFID`
1. 若任一关键 claim 为 `contradicted`
- `verdict = "REJECT"`
@@ -0,0 +1,48 @@
package com.superbiz.agent.agent.tool;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.superbiz.agent.service.ToolInvocationRecorder;
import org.junit.jupiter.api.Test;
import org.springframework.test.util.ReflectionTestUtils;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
import static org.mockito.Mockito.mock;
class QueryLogsToolsTest {
private final ObjectMapper objectMapper = new ObjectMapper();
@Test
void queryLogsReturnsHikariPositiveMockForOrderService() throws Exception {
QueryLogsTools tools = new QueryLogsTools(mock(ToolInvocationRecorder.class));
ReflectionTestUtils.setField(tools, "mockEnabled", true);
String output = tools.queryLogs("ap-guangzhou", "application-logs",
"order-service HikariCP connection pool active=50/50 waiting", 10);
JsonNode root = objectMapper.readTree(output);
assertTrue(root.path("success").asBoolean());
assertEquals(2, root.path("logs").size());
assertEquals("order-service", root.path("logs").get(0).path("service").asText());
assertTrue(root.toString().contains("HikariPool-1"));
assertFalse(root.toString().contains("generic-service"));
}
@Test
void queryLogsReturnsEmptyNoHitForOtherServiceHikariQuery() throws Exception {
QueryLogsTools tools = new QueryLogsTools(mock(ToolInvocationRecorder.class));
ReflectionTestUtils.setField(tools, "mockEnabled", true);
String output = tools.queryLogs("ap-guangzhou", "application-logs",
"inventory-service HikariCP connection pool active=50/50 waiting", 10);
JsonNode root = objectMapper.readTree(output);
assertFalse(root.path("success").asBoolean());
assertEquals(0, root.path("logs").size());
assertEquals(0, root.path("total").asInt());
assertFalse(root.toString().contains("generic-service"));
}
}
@@ -95,7 +95,7 @@ class VerifierInputHookTest {
));
ToolInvocationRepository invocationRepository = mock(ToolInvocationRepository.class);
when(invocationRepository.findBySessionIdOrderByIdAsc("structured-v2-session")).thenReturn(List.of(
ToolInvocation.builder().id(101L).sessionId("structured-v2-session").toolName("query_metrics").build()
invocation(101L, "structured-v2-session", "query_metrics", "$.alerts[0]", "active=50 max=50")
));
VerifierInputHook hook = new VerifierInputHook(traceSummaryService,
new ExecutorGatekeeperService(invocationRepository));
@@ -115,7 +115,8 @@ class VerifierInputHookTest {
"source_type": "tool_trace",
"source_id": "trace-1",
"tool_name": "query_metrics",
"source_invocation_ids": [101],
"source_invocation_id": 101,
"raw_path": "$.alerts[0]",
"evidence_excerpt": "active=50 max=50"
}
]
@@ -140,16 +141,96 @@ class VerifierInputHookTest {
assertEquals("连接池 active 达到上限",
payload.path("executor_structured_output").path("claims").get(0).path("claim_text").asText());
assertEquals("pass", payload.path("gatekeeper_result").path("status").asText());
assertEquals("none", payload.path("gatekeeper_result").path("severity").asText());
assertEquals("pass", VerifierContextHolder.getGatekeeperResult().get("status"));
}
@Test
void beforeModelBackfillsOnlyUniqueInvocationIdAndDoesNotPassWithoutRawPath() throws Exception {
ToolTraceSummaryService traceSummaryService = mock(ToolTraceSummaryService.class);
when(traceSummaryService.buildVerifierTraceSummary(anyString(), anyString())).thenReturn(List.of(
Map.of(
"trace_ref", "metrics-1",
"tool_name", "query_metrics",
"source_invocation_ids", List.of(101L)
)
));
ToolInvocationRepository invocationRepository = mock(ToolInvocationRepository.class);
when(invocationRepository.findBySessionIdOrderByIdAsc("backfill-session")).thenReturn(List.of(
invocation(101L, "backfill-session", "query_metrics", "$.alerts[0]",
"CPU 使用率持续超过 80%,当前值为 92%")
));
VerifierInputHook hook = new VerifierInputHook(traceSummaryService,
new ExecutorGatekeeperService(invocationRepository));
String executorOutput = """
{
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "symptom",
"claim_text": "payment-service CPU 使用率超过 92%",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "prometheus-alert-HighCPUUsage",
"tool_name": "queryPrometheusAlerts",
"evidence_excerpt": "CPU 使用率持续超过 80%,当前值为 92%"
}
]
}
],
"hypotheses": [],
"recommended_actions": [
{
"action_text": "restart payment-service",
"reason": "cpu alert is firing",
"evidence_bindings": [
{
"source_type": "tool_trace",
"source_id": "prometheus-alert-HighCPUUsage",
"tool_name": "queryPrometheusAlerts",
"evidence_excerpt": "CPU usage is 92%"
}
]
}
],
"missing_info": []
}
""";
AgentCommand command = hook.beforeModel(
List.of(new AssistantMessage(executorOutput)),
RunnableConfig.builder().addMetadata("sessionId", "backfill-session").build()
);
JsonNode payload = readPayload(command);
JsonNode binding = payload.path("executor_structured_output")
.path("claims").get(0)
.path("evidence_bindings").get(0);
assertEquals("query_metrics", binding.path("tool_name").asText());
assertEquals(101L, binding.path("source_invocation_id").asLong());
JsonNode actionBinding = payload.path("executor_structured_output")
.path("recommended_actions").get(0)
.path("evidence_bindings").get(0);
assertEquals("query_metrics", actionBinding.path("tool_name").asText());
assertEquals(101L, actionBinding.path("source_invocation_id").asLong());
assertFalse(binding.has("raw_path"));
assertEquals("fail", payload.path("gatekeeper_result").path("status").asText());
assertEquals("low_confid", payload.path("gatekeeper_result").path("severity").asText());
assertEquals("evidence.invocation_auto_backfill",
payload.path("gatekeeper_result").path("warnings").get(0).path("rule").asText());
}
@Test
void beforeModelAddsFailingGatekeeperResultForFabricatedInvocationId() throws Exception {
ToolTraceSummaryService traceSummaryService = mock(ToolTraceSummaryService.class);
when(traceSummaryService.buildVerifierTraceSummary(anyString(), anyString())).thenReturn(List.of());
ToolInvocationRepository invocationRepository = mock(ToolInvocationRepository.class);
when(invocationRepository.findBySessionIdOrderByIdAsc("fabricated-invocation-session")).thenReturn(List.of(
ToolInvocation.builder().id(101L).sessionId("fabricated-invocation-session").toolName("query_metrics").build()
invocation(101L, "fabricated-invocation-session", "query_metrics", "$.alerts[0]", "active=50 max=50")
));
VerifierInputHook hook = new VerifierInputHook(traceSummaryService,
new ExecutorGatekeeperService(invocationRepository));
@@ -167,7 +248,8 @@ class VerifierInputHookTest {
{
"source_type": "tool_trace",
"tool_name": "query_metrics",
"source_invocation_ids": [999],
"source_invocation_id": 999,
"raw_path": "$.alerts[0]",
"evidence_excerpt": "active=50 max=50"
}
]
@@ -186,6 +268,7 @@ class VerifierInputHookTest {
JsonNode payload = readPayload(command);
assertEquals("fail", payload.path("gatekeeper_result").path("status").asText());
assertEquals("reject", payload.path("gatekeeper_result").path("severity").asText());
assertEquals("evidence.invocation_ref",
payload.path("gatekeeper_result").path("failed_rules").get(0).asText());
}
@@ -284,4 +367,14 @@ class VerifierInputHookTest {
assertTrue(message instanceof UserMessage);
return objectMapper.readTree(((UserMessage) message).getText());
}
private ToolInvocation invocation(Long id, String sessionId, String toolName, String rawPath, String text) {
return ToolInvocation.builder()
.id(id)
.sessionId(sessionId)
.toolName(toolName)
.retrievalDetails("{\"evidence_refs\":[{\"raw_path\":\"" + rawPath
+ "\",\"text\":\"" + text + "\"}]}")
.build();
}
}
@@ -253,7 +253,8 @@ class ChatServiceSequentialAgentTest {
"source_type": "tool_trace",
"source_id": "trace-1",
"tool_name": "query_metrics",
"source_invocation_ids": [101],
"source_invocation_id": 101,
"raw_path": "$.alerts[0]",
"evidence_excerpt": "active=50 max=50"
}
]
@@ -300,7 +301,8 @@ class ChatServiceSequentialAgentTest {
"source_type": "tool_trace",
"source_id": "trace-1",
"tool_name": "query_metrics",
"source_invocation_ids": [101],
"source_invocation_id": 101,
"raw_path": "$.alerts[0]",
"evidence_excerpt": "active=50 max=50"
}
]
@@ -350,6 +352,7 @@ class ChatServiceSequentialAgentTest {
.id(101L)
.sessionId("sequential-gatekeeper-persist-session")
.toolName("query_metrics")
.retrievalDetails(evidenceRefs("$.alerts[0]", "active=50 max=50"))
.build()));
ScriptedChatModel chatModel = new ScriptedChatModel();
chatModel.executorOutput = """
@@ -366,7 +369,8 @@ class ChatServiceSequentialAgentTest {
"source_type": "tool_trace",
"source_id": "trace-1",
"tool_name": "query_metrics",
"source_invocation_ids": [101],
"source_invocation_id": 101,
"raw_path": "$.alerts[0]",
"evidence_excerpt": "active=50 max=50"
}
]
@@ -393,6 +397,7 @@ class ChatServiceSequentialAgentTest {
@SuppressWarnings("unchecked")
Map<String, Object> gatekeeperResult = (Map<String, Object>) verifierEvaluation.get("gatekeeper_result");
assertEquals("pass", gatekeeperResult.get("status"));
assertEquals("none", gatekeeperResult.get("severity"));
}
@Test
@@ -407,6 +412,7 @@ class ChatServiceSequentialAgentTest {
.id(101L)
.sessionId("sequential-claim-check-session")
.toolName("query_metrics")
.retrievalDetails(evidenceRefs("$.alerts[0]", "active=50 max=50"))
.build()));
ScriptedChatModel chatModel = new ScriptedChatModel("""
{
@@ -466,6 +472,7 @@ class ChatServiceSequentialAgentTest {
.id(101L)
.sessionId("sequential-gatekeeper-fail-session")
.toolName("query_metrics")
.retrievalDetails(evidenceRefs("$.alerts[0]", "active=50 max=50"))
.build()));
ScriptedChatModel chatModel = new ScriptedChatModel("""
{
@@ -493,7 +500,8 @@ class ChatServiceSequentialAgentTest {
{
"source_type": "tool_trace",
"tool_name": "query_metrics",
"source_invocation_ids": [999],
"source_invocation_id": 999,
"raw_path": "$.alerts[0]",
"evidence_excerpt": "active=50 max=50"
}
]
@@ -647,6 +655,7 @@ class ChatServiceSequentialAgentTest {
when(toolInvocationRepository.findBySessionIdOrderByIdAsc(anyString())).thenReturn(List.of(ToolInvocation.builder()
.id(101L)
.toolName("query_metrics")
.retrievalDetails(evidenceRefs("$.alerts[0]", "active=50 max=50"))
.build()));
EvaluationService evaluationService = mock(EvaluationService.class);
@@ -694,7 +703,8 @@ class ChatServiceSequentialAgentTest {
"source_type": "tool_trace",
"source_id": "trace-1",
"tool_name": "query_metrics",
"source_invocation_ids": [101],
"source_invocation_id": 101,
"raw_path": "$.alerts[0]",
"evidence_excerpt": "active=50 max=50"
}
]
@@ -707,6 +717,10 @@ class ChatServiceSequentialAgentTest {
""";
}
private String evidenceRefs(String rawPath, String text) {
return "{\"evidence_refs\":[{\"raw_path\":\"" + rawPath + "\",\"text\":\"" + text + "\"}]}";
}
private static final class ScriptedChatModel implements ChatModel {
private final java.util.ArrayList<String> agentCalls = new java.util.ArrayList<>();
private String promptText = "";
@@ -728,7 +742,8 @@ class ChatServiceSequentialAgentTest {
"source_type": "tool_trace",
"source_id": "trace-1",
"tool_name": "query_metrics",
"source_invocation_ids": [101],
"source_invocation_id": 101,
"raw_path": "$.alerts[0]",
"evidence_excerpt": "active=50 max=50"
}
]
@@ -18,14 +18,18 @@ class ExecutorGatekeeperServiceTest {
void validatePassesForExecutorEvidenceV2WithMatchingInvocation() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
ToolInvocation.builder().id(101L).sessionId("session-1").toolName("query_metrics").build()
invocation(101L, "query_metrics", "$.alerts[0]",
"HighCPUUsage firing, service=payment-service, current=92%, duration=25m")
));
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
Map<String, Object> result = service.validate("session-1", validOutput(101L, "query_metrics"),
Map<String, Object> result = service.validate("session-1",
validOutput(101L, "query_metrics", "$.alerts[0]",
"HighCPUUsage firing, service=payment-service, current=92%"),
Map.of("status", "valid"));
assertEquals("pass", result.get("status"));
assertEquals("none", result.get("severity"));
assertTrue(((List<?>) result.get("failed_rules")).isEmpty());
}
@@ -34,12 +38,13 @@ class ExecutorGatekeeperServiceTest {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of());
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
Map<String, Object> output = validOutput(101L, "query_metrics");
Map<String, Object> output = validOutput(101L, "query_metrics", "$.alerts[0]", "cpu=92");
output.put("user_facing_answer", "旧版最终答案");
Map<String, Object> result = service.validate("session-1", output, Map.of("status", "valid"));
assertEquals("fail", result.get("status"));
assertEquals("reject", result.get("severity"));
assertTrue(((List<?>) result.get("failed_rules")).contains("schema.executor_v2"));
}
@@ -47,14 +52,16 @@ class ExecutorGatekeeperServiceTest {
void validateFailsForFabricatedInvocationId() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
ToolInvocation.builder().id(101L).sessionId("session-1").toolName("query_metrics").build()
invocation(101L, "query_metrics", "$.alerts[0]", "cpu=92")
));
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
Map<String, Object> result = service.validate("session-1", validOutput(999L, "query_metrics"),
Map<String, Object> result = service.validate("session-1",
validOutput(999L, "query_metrics", "$.alerts[0]", "cpu=92"),
Map.of("status", "valid"));
assertEquals("fail", result.get("status"));
assertEquals("reject", result.get("severity"));
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.invocation_ref"));
}
@@ -62,19 +69,137 @@ class ExecutorGatekeeperServiceTest {
void validateFailsForToolNameMismatch() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
ToolInvocation.builder().id(101L).sessionId("session-1").toolName("query_logs").build()
invocation(101L, "query_logs", "$.logs[0]", "cpu=92")
));
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
Map<String, Object> result = service.validate("session-1", validOutput(101L, "query_metrics"),
Map<String, Object> result = service.validate("session-1",
validOutput(101L, "query_metrics", "$.alerts[0]", "cpu=92"),
Map.of("status", "valid"));
assertEquals("fail", result.get("status"));
assertEquals("reject", result.get("severity"));
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.invocation_ref"));
}
@SuppressWarnings("unchecked")
private Map<String, Object> validOutput(Long invocationId, String toolName) {
@Test
void validateDowngradesMissingRawPathToLowConfidence() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
invocation(101L, "query_metrics", "$.alerts[0]", "cpu=92")
));
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
Map<String, Object> output = validOutput(101L, "query_metrics", null, "cpu=92");
Map<String, Object> result = service.validate("session-1", output, Map.of("status", "valid"));
assertEquals("fail", result.get("status"));
assertEquals("low_confid", result.get("severity"));
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.raw_path"));
}
@Test
void validateRejectsUnknownRawPath() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
invocation(101L, "query_metrics", "$.alerts[0]", "cpu=92")
));
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
Map<String, Object> result = service.validate("session-1",
validOutput(101L, "query_metrics", "$.alerts[99]", "cpu=92"),
Map.of("status", "valid"));
assertEquals("fail", result.get("status"));
assertEquals("reject", result.get("severity"));
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.raw_path"));
}
@Test
void validateDowngradesOldInvocationWithoutEvidenceRefsToLowConfidence() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
ToolInvocation.builder().id(101L).sessionId("session-1").toolName("query_metrics")
.retrievalDetails("{\"evidence_status\":\"supported\"}")
.build()
));
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
Map<String, Object> result = service.validate("session-1",
validOutput(101L, "query_metrics", "$.alerts[0]", "cpu=92"),
Map.of("status", "valid"));
assertEquals("fail", result.get("status"));
assertEquals("low_confid", result.get("severity"));
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.missing"));
}
@Test
void validateRejectsMismatchedExcerpt() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
invocation(101L, "query_metrics", "$.alerts[0]", "HighCPUUsage firing service payment-service current 92")
));
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
Map<String, Object> result = service.validate("session-1",
validOutput(101L, "query_metrics", "$.alerts[0]", "HikariCP active=50/50 waiting=32"),
Map.of("status", "valid"));
assertEquals("fail", result.get("status"));
assertEquals("reject", result.get("severity"));
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.excerpt_mismatch"));
}
@Test
void validateFailsForRecommendedActionFabricatedInvocationId() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
invocation(101L, "query_metrics", "$.alerts[0]", "cpu=92")
));
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
Map<String, Object> output = validOutput(101L, "query_metrics", "$.alerts[0]", "cpu=92");
output.put("recommended_actions", List.of(Map.of(
"action_text", "restart service",
"reason", "alert is firing",
"evidence_bindings", List.of(Map.of(
"source_type", "tool_trace",
"source_id", "trace-1",
"tool_name", "query_metrics",
"source_invocation_id", 999L,
"raw_path", "$.alerts[0]",
"evidence_excerpt", "cpu=92"
))
)));
Map<String, Object> result = service.validate("session-1", output, Map.of("status", "valid"));
assertEquals("fail", result.get("status"));
assertEquals("reject", result.get("severity"));
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.invocation_ref"));
}
private ToolInvocation invocation(Long id, String toolName, String rawPath, String text) {
return ToolInvocation.builder()
.id(id)
.sessionId("session-1")
.toolName(toolName)
.retrievalDetails("{\"evidence_refs\":[{\"raw_path\":\"" + rawPath
+ "\",\"text\":\"" + text + "\"}]}")
.build();
}
private Map<String, Object> validOutput(Long invocationId, String toolName, String rawPath, String excerpt) {
Map<String, Object> binding = new java.util.LinkedHashMap<>();
binding.put("source_type", "tool_trace");
binding.put("source_id", "trace-1");
binding.put("tool_name", toolName);
binding.put("source_invocation_id", invocationId);
if (rawPath != null) {
binding.put("raw_path", rawPath);
}
binding.put("evidence_excerpt", excerpt);
return new java.util.LinkedHashMap<>(Map.of(
"answer_version", "executor_evidence_v2",
"claims", List.of(Map.of(
@@ -82,13 +207,7 @@ class ExecutorGatekeeperServiceTest {
"claim_type", "symptom",
"claim_text", "连接池 active 达到上限",
"support_level", "direct",
"evidence_bindings", List.of(Map.of(
"source_type", "tool_trace",
"source_id", "trace-1",
"tool_name", toolName,
"source_invocation_ids", List.of(invocationId),
"evidence_excerpt", "active=50 max=50"
))
"evidence_bindings", List.of(binding)
)),
"hypotheses", List.of(),
"recommended_actions", List.of(),
@@ -1,5 +1,6 @@
package com.superbiz.agent.service;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.superbiz.agent.domain.entity.ToolInvocation;
import com.superbiz.agent.dto.ContextPack;
@@ -23,6 +24,8 @@ import static org.mockito.Mockito.when;
class ToolInvocationRecorderTest {
private final ObjectMapper objectMapper = new ObjectMapper();
@Test
void recordEvidenceToolPreservesNoEvidenceSemantics() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
@@ -56,6 +59,72 @@ class ToolInvocationRecorderTest {
assertTrue(saved.getRetrievalDetails().contains("\"retrieved_domains\":[\"application-logs\"]"));
}
@Test
void recordEvidenceToolExtractsLogEvidenceRefs() throws Exception {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.save(any(ToolInvocation.class))).thenAnswer(invocation -> invocation.getArgument(0));
ToolInvocationRecorder recorder = new ToolInvocationRecorder(repository, new ObjectMapper());
SessionContextHolder.setSessionId("log-ref-session");
try {
recorder.recordEvidenceTool(
"query_logs",
Map.of("query", "HikariCP order-service"),
"""
{"success":true,"logs":[{"timestamp":"2026-07-08 10:00:00","level":"ERROR","service":"order-service","message":"HikariPool-1 - Connection is not available, request timed out after 30000ms","metrics":{"waiting":"32"}}]}
""",
true,
System.currentTimeMillis() - 10,
null,
"application-logs",
ToolInvocationRecorder.EVIDENCE_STATUS_SUPPORTED,
Map.of("log_topic", "application-logs")
);
} finally {
SessionContextHolder.clear();
}
ArgumentCaptor<ToolInvocation> captor = ArgumentCaptor.forClass(ToolInvocation.class);
verify(repository).save(captor.capture());
JsonNode details = objectMapper.readTree(captor.getValue().getRetrievalDetails());
assertEquals("$.logs[0]", details.path("evidence_refs").get(0).path("raw_path").asText());
assertTrue(details.path("evidence_refs").get(0).path("text").asText().contains("HikariPool-1"));
}
@Test
void recordEvidenceToolExtractsMetricEvidenceRefs() throws Exception {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.save(any(ToolInvocation.class))).thenAnswer(invocation -> invocation.getArgument(0));
ToolInvocationRecorder recorder = new ToolInvocationRecorder(repository, new ObjectMapper());
SessionContextHolder.setSessionId("metric-ref-session");
try {
recorder.recordEvidenceTool(
"query_metrics",
Map.of("query", "active_prometheus_alerts"),
"""
{"success":true,"alerts":[{"alert_name":"HighMemoryUsage","state":"firing","description":"服务 order-service 当前值为 91%","active_at":"2026-07-08T10:00:00Z","duration":"15m"}]}
""",
true,
System.currentTimeMillis() - 10,
null,
"prometheus_alerts",
ToolInvocationRecorder.EVIDENCE_STATUS_SUPPORTED,
Map.of("metric_family", "prometheus_alerts")
);
} finally {
SessionContextHolder.clear();
}
ArgumentCaptor<ToolInvocation> captor = ArgumentCaptor.forClass(ToolInvocation.class);
verify(repository).save(captor.capture());
JsonNode details = objectMapper.readTree(captor.getValue().getRetrievalDetails());
assertEquals("$.alerts[0]", details.path("evidence_refs").get(0).path("raw_path").asText());
assertTrue(details.path("evidence_refs").get(0).path("text").asText().contains("HighMemoryUsage"));
}
@Test
void recordLookupKnowledgePreservesRetrievalSpecificFields() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
@@ -113,6 +182,8 @@ class ToolInvocationRecorderTest {
assertTrue(saved.getRetrievalDetails().contains("\"evidence_candidate_count\":2"));
assertTrue(saved.getRetrievalDetails().contains("\"evidence_block_count\":1"));
assertTrue(saved.getRetrievalDetails().contains("\"evidence_blocks\""));
assertTrue(saved.getRetrievalDetails().contains("\"evidence_refs\""));
assertTrue(saved.getRetrievalDetails().contains("\"raw_path\":\"$.evidence_blocks[0]\""));
assertTrue(saved.getRetrievalDetails().contains("\"query_transform\""));
assertTrue(saved.getRetrievalDetails().contains("\"retrieval_trace\""));
assertTrue(saved.getRetrievalDetails().contains("\"context_pack_summary\""));