Compare commits
2
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
7b8c75e571 | ||
|
|
a08672b31e |
@@ -10,6 +10,7 @@
|
|||||||
| 2026-07-07 | executor-gatekeeper-hook | Chat质量门禁/证据归因 | Gatekeeper, verifier payload, source_invocation_ids, tool_name match, self_evaluation | openspec/changes/archive/2026-07-07-executor-gatekeeper-hook | archived |
|
| 2026-07-07 | executor-gatekeeper-hook | Chat质量门禁/证据归因 | Gatekeeper, verifier payload, source_invocation_ids, tool_name match, self_evaluation | openspec/changes/archive/2026-07-07-executor-gatekeeper-hook | archived |
|
||||||
| 2026-07-07 | executor-verifier-claim-checks | Chat质量门禁/证据归因 | Verifier claim_checks, facts_checked compatibility, effective verdict guardrail, malformed output downgrade | openspec/changes/archive/2026-07-07-executor-verifier-claim-checks | archived |
|
| 2026-07-07 | executor-verifier-claim-checks | Chat质量门禁/证据归因 | Verifier claim_checks, facts_checked compatibility, effective verdict guardrail, malformed output downgrade | openspec/changes/archive/2026-07-07-executor-verifier-claim-checks | archived |
|
||||||
| 2026-07-08 | executor-composer-final-answer | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
|
| 2026-07-08 | executor-composer-final-answer | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
|
||||||
|
| 2026-07-08 | verifier-evidence-reference-fidelity | Chat质量门禁/证据归因 | evidence_refs, raw_path, Gatekeeper severity, verifier evidence excerpt, HikariCP mock, no_evidence | openspec/changes/archive/2026-07-08-verifier-evidence-reference-fidelity | archived |
|
||||||
| 2026-07-06 | rag-eval-pipeline-closure | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
|
| 2026-07-06 | rag-eval-pipeline-closure | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
|
||||||
| 2026-07-06 | modular-rag-pipeline | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
|
| 2026-07-06 | modular-rag-pipeline | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
|
||||||
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
|
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
|
||||||
|
|||||||
@@ -0,0 +1,52 @@
|
|||||||
|
# Acceptance
|
||||||
|
|
||||||
|
## Static Verification
|
||||||
|
|
||||||
|
- `openspec validate verifier-evidence-reference-fidelity --strict`: passed.
|
||||||
|
- `openspec validate --specs`: passed.
|
||||||
|
|
||||||
|
## Script Verification
|
||||||
|
|
||||||
|
- `mvn "-Dtest=ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,QueryLogsToolsTest,ChatServiceSequentialAgentTest" test`
|
||||||
|
- Passed: 38 tests.
|
||||||
|
- `mvn "-Dtest=ExecutorGatekeeperServiceTest" test`
|
||||||
|
- Passed: 9 tests.
|
||||||
|
- `mvn test`
|
||||||
|
- Failed on unrelated environment-gated `MilvusConnectionTest.connect`: `MILVUS_TOKEN` environment variable was not set.
|
||||||
|
- Other executed tests in the run progressed until that single failure; focused tests for this change passed.
|
||||||
|
|
||||||
|
## End-to-End Verification
|
||||||
|
|
||||||
|
The Java service was restarted with `mvn spring-boot:run`; logs were written under `logs/`.
|
||||||
|
|
||||||
|
| Case | Session | Result | Gatekeeper Audit |
|
||||||
|
|---|---|---|---|
|
||||||
|
| HikariCP positive | `iss007-hikari-positive-20260708-1553` | PASS; confirmed order-service HikariCP timeout and pool saturation logs | `pass / none`, checked_bindings=2 |
|
||||||
|
| HikariCP negative | `iss007-hikari-negative-20260708-1555` | LOW_CONFID; no `generic-service`; no false positive for inventory-service | `fail / low_confid` |
|
||||||
|
| HighMemoryUsage positive | `iss007-memory-positive-20260708-1558` | PASS; confirmed HighMemoryUsage 91%, did not confirm memory leak | `pass / none`, checked_bindings=1 |
|
||||||
|
| SlowResponse positive | `iss007-slow-positive-20260708-1600` | PASS; confirmed SlowResponse and slow request logs, no DB pool root cause | `pass / none`, checked_bindings=7 |
|
||||||
|
| Narrow HighCPUUsage | `iss007-narrow-highcpu-20260708-1602` | PASS; only covered payment-service HighCPUUsage | `pass / none`, checked_bindings=1 |
|
||||||
|
|
||||||
|
## Database Audit
|
||||||
|
|
||||||
|
`scripts/query_mysql.py` was used to verify:
|
||||||
|
|
||||||
|
- `diagnosis_session.self_evaluation.verifier_evaluation.verdict`
|
||||||
|
- `gatekeeper_result.status`
|
||||||
|
- `gatekeeper_result.severity`
|
||||||
|
- `gatekeeper_result.checked_bindings`
|
||||||
|
- no-hit HikariCP query rows persist `evidence_status=no_evidence`
|
||||||
|
|
||||||
|
## Remaining Risk
|
||||||
|
|
||||||
|
- Negative no-hit claims still have incomplete precise references when Executor uses `$.logs` for empty arrays. Gatekeeper correctly downgrades to `LOW_CONFID`.
|
||||||
|
- Prompt-only scope control is improved but not a hard contract. A future `scope_contract` may still be needed.
|
||||||
|
- Full test suite requires `MILVUS_TOKEN` to pass `MilvusConnectionTest`.
|
||||||
|
|
||||||
|
## OpenSpec Archive
|
||||||
|
|
||||||
|
- `openspec archive verifier-evidence-reference-fidelity --yes`: succeeded.
|
||||||
|
- Main specs updated:
|
||||||
|
- `openspec/specs/chat-verifier-agent/spec.md`
|
||||||
|
- `openspec/specs/evidence-trace-hardening/spec.md`
|
||||||
|
- Non-blocking warning: proposal did not use OpenSpec's preferred `## Why` / `## What Changes` headers, but archive completed.
|
||||||
@@ -0,0 +1,44 @@
|
|||||||
|
# Verifier Evidence Reference Fidelity
|
||||||
|
|
||||||
|
## Background
|
||||||
|
|
||||||
|
ISS-007 came from end-to-end diagnosis cases where raw tool output and Executor `evidence_excerpt` contained enough facts, but Verifier still returned `LOW_CONFID` because the verifier-facing summary compressed away key details.
|
||||||
|
|
||||||
|
The affected flow is:
|
||||||
|
|
||||||
|
```text
|
||||||
|
chat_planner
|
||||||
|
-> chat_executor
|
||||||
|
-> VerifierInputHook / Gatekeeper
|
||||||
|
-> chat_verifier
|
||||||
|
-> chat_composer
|
||||||
|
```
|
||||||
|
|
||||||
|
The change hardens the evidence handoff between Executor, Gatekeeper, and Verifier.
|
||||||
|
|
||||||
|
## Goal
|
||||||
|
|
||||||
|
Make Executor cite concrete tool evidence, make Gatekeeper validate that citation with code, and make Verifier judge whether verified evidence can derive the claim.
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
- Persist `tool_invocation.retrieval_details.evidence_refs`.
|
||||||
|
- Use `source_invocation_id + raw_path + evidence_excerpt` as the precise evidence binding.
|
||||||
|
- Add Gatekeeper `severity` and checked binding audit.
|
||||||
|
- Keep Gatekeeper in the Verifier hook path.
|
||||||
|
- Keep `tool_trace_summary` as navigation/audit context, not the only evidence source.
|
||||||
|
- Fix HikariCP mock positive/no-hit behavior.
|
||||||
|
- Tighten Executor/Verifier prompts for narrow-scope evidence handling.
|
||||||
|
|
||||||
|
## Non-goals
|
||||||
|
|
||||||
|
- No Planner `scope_contract`.
|
||||||
|
- No new database table.
|
||||||
|
- No full JSONPath engine.
|
||||||
|
- No change to external HTTP API.
|
||||||
|
- No retry rollback from Gatekeeper to Executor in this phase.
|
||||||
|
|
||||||
|
## OpenSpec
|
||||||
|
|
||||||
|
- Change: `openspec/changes/verifier-evidence-reference-fidelity`
|
||||||
|
- Source issue: `mvp/issues/ISS-007-verifier-evidence-summary-fidelity.md`
|
||||||
@@ -0,0 +1,54 @@
|
|||||||
|
# Decisions
|
||||||
|
|
||||||
|
## Evidence Reference
|
||||||
|
|
||||||
|
Use `source_invocation_id + raw_path + evidence_excerpt` as the precise evidence reference for Executor claim bindings.
|
||||||
|
|
||||||
|
Reason:
|
||||||
|
|
||||||
|
- Invocation ID alone only identifies a tool call, not the evidence inside it.
|
||||||
|
- `raw_path` is enough for the first version when paired with `retrieval_details.evidence_refs`.
|
||||||
|
- `evidence_excerpt` remains the text Verifier reads, but only after Gatekeeper validates it.
|
||||||
|
|
||||||
|
## Raw Path
|
||||||
|
|
||||||
|
Only support stable locators in the first version:
|
||||||
|
|
||||||
|
- `$.alerts[i]`
|
||||||
|
- `$.logs[i]`
|
||||||
|
- `$.evidence_blocks[i]`
|
||||||
|
|
||||||
|
No full JSONPath engine is introduced.
|
||||||
|
|
||||||
|
## Gatekeeper Severity
|
||||||
|
|
||||||
|
Gatekeeper output includes:
|
||||||
|
|
||||||
|
- `status`
|
||||||
|
- `severity`
|
||||||
|
- `checked_bindings`
|
||||||
|
- `failed_rules`
|
||||||
|
- `warnings`
|
||||||
|
- `errors`
|
||||||
|
|
||||||
|
Severity meaning:
|
||||||
|
|
||||||
|
- `none`: precise references passed.
|
||||||
|
- `low_confid`: evidence is missing or incomplete, but not fabricated.
|
||||||
|
- `reject`: fabricated ID, wrong tool, unknown raw path, or mismatched excerpt.
|
||||||
|
|
||||||
|
## Verifier Boundary
|
||||||
|
|
||||||
|
Verifier uses verified claim-local excerpts as primary derivability evidence. `tool_trace_summary` remains available for navigation and audit, but no longer needs to carry every concrete fact.
|
||||||
|
|
||||||
|
## Hook Placement
|
||||||
|
|
||||||
|
Gatekeeper remains in the Verifier input hook path. This version does not retry Executor on Gatekeeper failure.
|
||||||
|
|
||||||
|
## Planner
|
||||||
|
|
||||||
|
Planner is not changed. `scope_contract` remains a later-stage idea. This phase uses prompt constraints to reduce narrow-scope over-expansion.
|
||||||
|
|
||||||
|
## Database
|
||||||
|
|
||||||
|
No new tables. Evidence refs are stored in `tool_invocation.retrieval_details.evidence_refs`; audit is stored in `diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result`.
|
||||||
@@ -0,0 +1,33 @@
|
|||||||
|
# Evidence
|
||||||
|
|
||||||
|
## Existing Context
|
||||||
|
|
||||||
|
- Existing `chat-verifier-agent` spec still used `source_invocation_ids` and `tool_trace_summary` as the main verifier evidence context.
|
||||||
|
- Existing `evidence-trace-hardening` spec already established `tool_invocation.retrieval_details` as the right place for structured tool-specific facts.
|
||||||
|
- Prior devflow projects established that runbook/skill content is guidance, not incident evidence.
|
||||||
|
|
||||||
|
## Code Findings
|
||||||
|
|
||||||
|
- `VerifierInputHook` previously backfilled plural `source_invocation_ids` from `tool_trace_summary` by tool name.
|
||||||
|
- `ExecutorGatekeeperService` previously validated invocation existence and tool name, but not `raw_path` or excerpt authenticity.
|
||||||
|
- `ToolInvocationRecorder` persisted retrieval details but did not generate claim-addressable `evidence_refs`.
|
||||||
|
- `QueryLogsTools` could fall back to `generic-service` placeholder logs on no-hit.
|
||||||
|
|
||||||
|
## Implementation Evidence
|
||||||
|
|
||||||
|
- `ToolInvocationRecorder` now extracts:
|
||||||
|
- `$.alerts[i]` for `query_metrics`
|
||||||
|
- `$.logs[i]` for `query_logs`
|
||||||
|
- `$.evidence_blocks[i]` for `lookup_knowledge`
|
||||||
|
- `ExecutorGatekeeperService` now validates:
|
||||||
|
- invocation existence
|
||||||
|
- tool name
|
||||||
|
- raw path presence
|
||||||
|
- `retrieval_details.evidence_refs`
|
||||||
|
- excerpt similarity/support
|
||||||
|
- `VerifierInputHook` only auto-fills a singular `source_invocation_id` when exactly one candidate exists and never invents `raw_path`.
|
||||||
|
- `QueryLogsTools` returns HikariCP mock logs for `order-service` and returns empty no-hit results for unrelated services.
|
||||||
|
|
||||||
|
## Residual Finding
|
||||||
|
|
||||||
|
The HikariCP negative E2E no longer has generic-service pollution, but the model still issued an extra broad HikariCP query without the service filter and used order-service as context. This is a remaining narrow-scope behavior issue, not a mock evidence pollution issue.
|
||||||
+46
-14
@@ -1,6 +1,16 @@
|
|||||||
# Diagnosis Eval Harness
|
# Diagnosis Eval Harness
|
||||||
|
|
||||||
This folder contains the first fixed-case evaluation set for the MVP diagnosis Agent.
|
This folder contains the fixed offline regression set for the MVP diagnosis Agent.
|
||||||
|
|
||||||
|
## Background
|
||||||
|
|
||||||
|
The current diagnosis chain is:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Planner -> Executor -> Gatekeeper -> Verifier -> Composer -> final answer
|
||||||
|
```
|
||||||
|
|
||||||
|
Stages 1-4 introduced Executor V2 structured output, deterministic Gatekeeper audit, Verifier `claim_checks`, and Composer final-answer rendering. Stage 5 makes those audit fields part of the offline regression harness so future prompt, tool, or chain changes can be checked without relying on a one-off demo.
|
||||||
|
|
||||||
## Scope
|
## Scope
|
||||||
|
|
||||||
@@ -14,7 +24,23 @@ This folder contains the first fixed-case evaluation set for the MVP diagnosis A
|
|||||||
|
|
||||||
## Current Mode
|
## Current Mode
|
||||||
|
|
||||||
The first version evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.
|
The baseline evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.
|
||||||
|
|
||||||
|
The committed baseline currently contains:
|
||||||
|
|
||||||
|
```text
|
||||||
|
8 fixed cases
|
||||||
|
8 passing fixture evaluations
|
||||||
|
2 PASS verdicts
|
||||||
|
5 LOW_CONFID verdicts
|
||||||
|
1 REJECT verdict
|
||||||
|
```
|
||||||
|
|
||||||
|
The three V2 audit-closure cases cover:
|
||||||
|
|
||||||
|
- Gatekeeper failure for a fabricated tool invocation reference.
|
||||||
|
- Unsupported claim filtering before the final answer.
|
||||||
|
- Composer fallback rendering without raw Executor JSON leakage.
|
||||||
|
|
||||||
## Verification
|
## Verification
|
||||||
|
|
||||||
@@ -24,29 +50,35 @@ Run the focused evaluator test:
|
|||||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test
|
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test
|
||||||
```
|
```
|
||||||
|
|
||||||
The committed baseline report represents the current fixed fixture set:
|
Run the broader phase-5 regression set:
|
||||||
|
|
||||||
```text
|
```powershell
|
||||||
5 fixed cases
|
mvn "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
|
||||||
5 passing fixture evaluations
|
|
||||||
2 PASS verdicts
|
|
||||||
3 LOW_CONFID verdicts
|
|
||||||
```
|
```
|
||||||
|
|
||||||
When fixtures or evaluator rules change, regenerate the report from the same case file and fixture directory, then update both JSON and Markdown outputs together.
|
When fixtures or evaluator rules change, regenerate both baseline reports from the same case file and fixture directory, then update JSON and Markdown together.
|
||||||
|
|
||||||
## Interview Story
|
## Regression Signal
|
||||||
|
|
||||||
The harness gives the MVP a repeatable baseline:
|
The harness is deterministic code, not an LLM judge:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
fixed diagnosis case
|
fixed diagnosis case
|
||||||
-> saved or runtime trace
|
-> saved trace fixture
|
||||||
-> rule-based trace validation
|
-> rule-based trace validation
|
||||||
-> JSON / Markdown report
|
-> JSON / Markdown report
|
||||||
-> regression signal for prompts, tools, retrieval, and verifier behavior
|
-> regression signal for prompts, tools, retrieval, verifier, and composer behavior
|
||||||
```
|
```
|
||||||
|
|
||||||
|
Stage 5 adds these V2 checks:
|
||||||
|
|
||||||
|
- `gatekeeper_result.status=fail` cannot coexist with Verifier `PASS`.
|
||||||
|
- Required V2 fixtures must include `gatekeeper_result`, `claim_checks`, and `composer_output`.
|
||||||
|
- `claim_checks` must be structurally auditable.
|
||||||
|
- Composer output must record whether normal parsing or fallback rendering was used.
|
||||||
|
- Final answers must not leak raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`.
|
||||||
|
- Configured unsupported claim keywords must not appear as confirmed final-answer content.
|
||||||
|
|
||||||
## Baseline Diff
|
## Baseline Diff
|
||||||
|
|
||||||
Baseline diff compares a current report against `reports/baseline-report.json`.
|
Baseline diff compares a current report against `reports/baseline-report.json`.
|
||||||
@@ -58,4 +90,4 @@ current report
|
|||||||
-> regressions, improvements, and changed signals
|
-> regressions, improvements, and changed signals
|
||||||
```
|
```
|
||||||
|
|
||||||
Use it to answer: did a prompt, tool, retrieval, or verifier change make the Agent worse than the fixed baseline?
|
Use it to answer: did a prompt, tool, retrieval, verifier, or composer change make the Agent worse than the fixed baseline?
|
||||||
|
|||||||
@@ -53,5 +53,56 @@
|
|||||||
"requiredEvidenceTools": ["query_metrics", "query_logs"],
|
"requiredEvidenceTools": ["query_metrics", "query_logs"],
|
||||||
"allowedVerdicts": ["LOW_CONFID", "PASS"],
|
"allowedVerdicts": ["LOW_CONFID", "PASS"],
|
||||||
"forbiddenAnswerKeywords": ["可以忽略"]
|
"forbiddenAnswerKeywords": ["可以忽略"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "gatekeeper-fabricated-invocation",
|
||||||
|
"title": "Gatekeeper fabricated invocation",
|
||||||
|
"question": "支付失败是否能确认由日志中的连接池耗尽导致?",
|
||||||
|
"traceFixture": "gatekeeper-fabricated-invocation-reject.json",
|
||||||
|
"expectedRootCauseKeywords": ["证据", "引用", "失败"],
|
||||||
|
"minKeywordMatches": 2,
|
||||||
|
"requiredEvidenceTools": ["query_logs"],
|
||||||
|
"allowedVerdicts": ["REJECT", "LOW_CONFID"],
|
||||||
|
"forbiddenAnswerKeywords": ["已经完全确认"],
|
||||||
|
"requireV2AuditClosure": true,
|
||||||
|
"requireClaimChecks": true,
|
||||||
|
"requireComposerOutput": true,
|
||||||
|
"expectedGatekeeperStatuses": ["fail"],
|
||||||
|
"expectedComposerStatuses": ["valid"],
|
||||||
|
"forbiddenConfirmedClaimKeywords": ["连接池耗尽导致支付失败"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "unsupported-claim-filtering",
|
||||||
|
"title": "Unsupported claim filtering",
|
||||||
|
"question": "订单超时是否可以确认由数据库主库故障导致?",
|
||||||
|
"traceFixture": "unsupported-claim-filtering-low-confid.json",
|
||||||
|
"expectedRootCauseKeywords": ["超时", "证据"],
|
||||||
|
"minKeywordMatches": 2,
|
||||||
|
"requiredEvidenceTools": ["query_logs"],
|
||||||
|
"allowedVerdicts": ["LOW_CONFID"],
|
||||||
|
"forbiddenAnswerKeywords": ["已经确认"],
|
||||||
|
"requireV2AuditClosure": true,
|
||||||
|
"requireClaimChecks": true,
|
||||||
|
"requireComposerOutput": true,
|
||||||
|
"expectedGatekeeperStatuses": ["pass"],
|
||||||
|
"expectedComposerStatuses": ["valid"],
|
||||||
|
"forbiddenConfirmedClaimKeywords": ["主库故障"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "composer-fallback-no-raw-json",
|
||||||
|
"title": "Composer fallback no raw JSON",
|
||||||
|
"question": "库存服务慢响应是否可以直接输出 Executor JSON?",
|
||||||
|
"traceFixture": "composer-fallback-no-raw-json-low-confid.json",
|
||||||
|
"expectedRootCauseKeywords": ["慢响应", "证据"],
|
||||||
|
"minKeywordMatches": 2,
|
||||||
|
"requiredEvidenceTools": ["query_metrics"],
|
||||||
|
"allowedVerdicts": ["LOW_CONFID"],
|
||||||
|
"forbiddenAnswerKeywords": ["executor_evidence_v2", "answer_version", "claim_id"],
|
||||||
|
"requireV2AuditClosure": true,
|
||||||
|
"requireClaimChecks": true,
|
||||||
|
"requireComposerOutput": true,
|
||||||
|
"expectedGatekeeperStatuses": ["pass"],
|
||||||
|
"expectedComposerStatuses": ["composer_malformed"],
|
||||||
|
"forbiddenConfirmedClaimKeywords": ["线程池已经耗尽"]
|
||||||
}
|
}
|
||||||
]
|
]
|
||||||
|
|||||||
@@ -0,0 +1,142 @@
|
|||||||
|
{
|
||||||
|
"session": {
|
||||||
|
"sessionId": "eval-composer-fallback-no-raw-json",
|
||||||
|
"query": "库存服务慢响应是否可以直接输出 Executor JSON?",
|
||||||
|
"status": "SUCCESS",
|
||||||
|
"agentFlow": "CHAT",
|
||||||
|
"totalDurationMs": 47000,
|
||||||
|
"toolCallCount": 1,
|
||||||
|
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:指标显示库存服务出现慢响应。\n\n仍需补充信息:当前没有线程池队列或线程耗尽证据,不能确认线程池方向。\n\n建议动作:补充查询库存服务线程池指标。",
|
||||||
|
"selfEvaluation": {
|
||||||
|
"verifier_evaluation": {
|
||||||
|
"verdict": "LOW_CONFID",
|
||||||
|
"groundedness_score": 0.45,
|
||||||
|
"critical_fact_count": 2,
|
||||||
|
"gatekeeper_result": {
|
||||||
|
"status": "pass",
|
||||||
|
"failed_rules": [],
|
||||||
|
"warnings": [],
|
||||||
|
"errors": []
|
||||||
|
},
|
||||||
|
"executor_structured_output": {
|
||||||
|
"answer_version": "executor_evidence_v2",
|
||||||
|
"claims": [
|
||||||
|
{
|
||||||
|
"claim_id": "claim-slow-response",
|
||||||
|
"claim_type": "symptom",
|
||||||
|
"claim_text": "库存服务出现慢响应",
|
||||||
|
"support_level": "direct",
|
||||||
|
"evidence_bindings": [
|
||||||
|
{
|
||||||
|
"tool_invocation_id": 1,
|
||||||
|
"tool_name": "query_metrics",
|
||||||
|
"evidence_excerpt": "inventory p99 latency increased"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"claim_id": "claim-thread-pool",
|
||||||
|
"claim_type": "root_cause",
|
||||||
|
"claim_text": "线程池已经耗尽",
|
||||||
|
"support_level": "weak",
|
||||||
|
"evidence_bindings": [
|
||||||
|
{
|
||||||
|
"tool_invocation_id": 1,
|
||||||
|
"tool_name": "query_metrics",
|
||||||
|
"evidence_excerpt": "inventory p99 latency increased"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"hypotheses": [],
|
||||||
|
"recommended_actions": [
|
||||||
|
{
|
||||||
|
"action_text": "补充查询库存服务线程池指标",
|
||||||
|
"reason": "当前只有慢响应指标"
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"missing_info": ["线程池队列长度", "活跃线程数"]
|
||||||
|
},
|
||||||
|
"claim_checks": [
|
||||||
|
{
|
||||||
|
"claim_id": "claim-slow-response",
|
||||||
|
"claim_text": "库存服务出现慢响应",
|
||||||
|
"claim_type": "symptom",
|
||||||
|
"verification": "direct_observation",
|
||||||
|
"detail": "指标显示 inventory p99 latency increased",
|
||||||
|
"evidence_refs": [
|
||||||
|
{
|
||||||
|
"tool_invocation_id": 1,
|
||||||
|
"tool_name": "query_metrics"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"claim_id": "claim-thread-pool",
|
||||||
|
"claim_text": "线程池已经耗尽",
|
||||||
|
"claim_type": "root_cause",
|
||||||
|
"verification": "external_unknown",
|
||||||
|
"detail": "没有线程池队列或活跃线程指标,不能确认该结论",
|
||||||
|
"evidence_refs": []
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"facts_checked": [
|
||||||
|
{
|
||||||
|
"fact": "库存服务出现慢响应",
|
||||||
|
"verification": "direct_evidence",
|
||||||
|
"detail": "指标显示 inventory p99 latency increased",
|
||||||
|
"evidence_refs": [
|
||||||
|
{
|
||||||
|
"tool_invocation_id": 1,
|
||||||
|
"tool_name": "query_metrics"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"fact": "线程池已经耗尽",
|
||||||
|
"verification": "external_unknown",
|
||||||
|
"detail": "没有线程池队列或活跃线程指标,不能确认该结论",
|
||||||
|
"evidence_refs": []
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"composer_output": {
|
||||||
|
"status": "composer_malformed",
|
||||||
|
"detail": "used safe fallback rendering",
|
||||||
|
"answer_summary": "指标显示库存服务出现慢响应。",
|
||||||
|
"recommended_actions": [
|
||||||
|
{
|
||||||
|
"action_text": "补充查询库存服务线程池指标",
|
||||||
|
"reason": "当前只有慢响应指标"
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"user_facing_answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:指标显示库存服务出现慢响应。\n\n仍需补充信息:当前没有线程池队列或线程耗尽证据,不能确认线程池方向。\n\n建议动作:补充查询库存服务线程池指标。"
|
||||||
|
},
|
||||||
|
"tool_trace_summary": [
|
||||||
|
{
|
||||||
|
"tool_name": "query_metrics",
|
||||||
|
"success": true,
|
||||||
|
"evidence_level": "direct"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"steps": [],
|
||||||
|
"toolInvocations": [
|
||||||
|
{
|
||||||
|
"id": 1,
|
||||||
|
"sessionId": "eval-composer-fallback-no-raw-json",
|
||||||
|
"toolName": "query_metrics",
|
||||||
|
"outputPreview": "inventory p99 latency increased",
|
||||||
|
"success": true
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"summary": {
|
||||||
|
"persistedStepCount": 3,
|
||||||
|
"returnedStepCount": 3,
|
||||||
|
"persistedToolCallCount": 1,
|
||||||
|
"returnedToolCallCount": 1,
|
||||||
|
"hasVerifierEvaluation": true,
|
||||||
|
"hasFeedback": false
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,110 @@
|
|||||||
|
{
|
||||||
|
"session": {
|
||||||
|
"sessionId": "eval-gatekeeper-fabricated-invocation",
|
||||||
|
"query": "支付失败是否能确认由日志中的连接池耗尽导致?",
|
||||||
|
"status": "SUCCESS",
|
||||||
|
"agentFlow": "CHAT",
|
||||||
|
"totalDurationMs": 39000,
|
||||||
|
"toolCallCount": 1,
|
||||||
|
"answer": "当前无法基于已获取证据生成可靠结论。\n\n证据引用校验失败:Executor 引用了不存在的工具调用记录,因此不能把连接池问题作为确认结论。建议重新收集日志证据后再判断。",
|
||||||
|
"selfEvaluation": {
|
||||||
|
"verifier_evaluation": {
|
||||||
|
"verdict": "REJECT",
|
||||||
|
"groundedness_score": 0.1,
|
||||||
|
"critical_fact_count": 1,
|
||||||
|
"gatekeeper_result": {
|
||||||
|
"status": "fail",
|
||||||
|
"failed_rules": ["evidence.invocation_ref"],
|
||||||
|
"warnings": [],
|
||||||
|
"errors": [
|
||||||
|
{
|
||||||
|
"rule_id": "evidence.invocation_ref",
|
||||||
|
"field": "claims[0].evidence_bindings[0].tool_invocation_id",
|
||||||
|
"message": "tool_invocation_id does not exist"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"executor_structured_output": {
|
||||||
|
"answer_version": "executor_evidence_v2",
|
||||||
|
"claims": [
|
||||||
|
{
|
||||||
|
"claim_id": "claim-1",
|
||||||
|
"claim_type": "root_cause",
|
||||||
|
"claim_text": "连接池耗尽导致支付失败",
|
||||||
|
"support_level": "direct",
|
||||||
|
"evidence_bindings": [
|
||||||
|
{
|
||||||
|
"tool_invocation_id": 99,
|
||||||
|
"tool_name": "query_logs",
|
||||||
|
"evidence_excerpt": "connection pool exhausted"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"hypotheses": [],
|
||||||
|
"recommended_actions": [
|
||||||
|
{
|
||||||
|
"action_text": "重新查询支付服务错误日志",
|
||||||
|
"reason": "当前 Executor 证据引用无法回溯"
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"missing_info": ["需要有效的日志工具调用记录"]
|
||||||
|
},
|
||||||
|
"claim_checks": [
|
||||||
|
{
|
||||||
|
"claim_id": "claim-1",
|
||||||
|
"claim_text": "连接池耗尽导致支付失败",
|
||||||
|
"claim_type": "root_cause",
|
||||||
|
"verification": "unsupported",
|
||||||
|
"detail": "Gatekeeper 已判定证据引用不存在,不能确认该结论",
|
||||||
|
"evidence_refs": []
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"facts_checked": [
|
||||||
|
{
|
||||||
|
"fact": "连接池耗尽导致支付失败",
|
||||||
|
"verification": "unsupported",
|
||||||
|
"detail": "Gatekeeper 已判定证据引用不存在,不能确认该结论",
|
||||||
|
"evidence_refs": []
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"composer_output": {
|
||||||
|
"status": "valid",
|
||||||
|
"answer_summary": "证据引用校验失败,不能确认根因。",
|
||||||
|
"recommended_actions": [
|
||||||
|
{
|
||||||
|
"action_text": "重新查询支付服务错误日志",
|
||||||
|
"reason": "当前 Executor 证据引用无法回溯"
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"user_facing_answer": "当前无法基于已获取证据生成可靠结论。\n\n证据引用校验失败:Executor 引用了不存在的工具调用记录,因此不能把连接池问题作为确认结论。建议重新收集日志证据后再判断。"
|
||||||
|
},
|
||||||
|
"tool_trace_summary": [
|
||||||
|
{
|
||||||
|
"tool_name": "query_logs",
|
||||||
|
"success": true,
|
||||||
|
"evidence_level": "direct"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"steps": [],
|
||||||
|
"toolInvocations": [
|
||||||
|
{
|
||||||
|
"id": 1,
|
||||||
|
"sessionId": "eval-gatekeeper-fabricated-invocation",
|
||||||
|
"toolName": "query_logs",
|
||||||
|
"outputPreview": "payment failed without matching connection pool exhaustion entry",
|
||||||
|
"success": true
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"summary": {
|
||||||
|
"persistedStepCount": 3,
|
||||||
|
"returnedStepCount": 3,
|
||||||
|
"persistedToolCallCount": 1,
|
||||||
|
"returnedToolCallCount": 1,
|
||||||
|
"hasVerifierEvaluation": true,
|
||||||
|
"hasFeedback": false
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,151 @@
|
|||||||
|
{
|
||||||
|
"session": {
|
||||||
|
"sessionId": "eval-unsupported-claim-filtering",
|
||||||
|
"query": "订单超时是否可以确认由数据库主库故障导致?",
|
||||||
|
"status": "SUCCESS",
|
||||||
|
"agentFlow": "CHAT",
|
||||||
|
"totalDurationMs": 44000,
|
||||||
|
"toolCallCount": 1,
|
||||||
|
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:日志显示订单接口出现超时。\n\n仍需补充信息:目前没有数据库故障日志或主库状态证据,不能把该方向写成确认根因。",
|
||||||
|
"selfEvaluation": {
|
||||||
|
"verifier_evaluation": {
|
||||||
|
"verdict": "LOW_CONFID",
|
||||||
|
"groundedness_score": 0.42,
|
||||||
|
"critical_fact_count": 2,
|
||||||
|
"gatekeeper_result": {
|
||||||
|
"status": "pass",
|
||||||
|
"failed_rules": [],
|
||||||
|
"warnings": [],
|
||||||
|
"errors": []
|
||||||
|
},
|
||||||
|
"executor_structured_output": {
|
||||||
|
"answer_version": "executor_evidence_v2",
|
||||||
|
"claims": [
|
||||||
|
{
|
||||||
|
"claim_id": "claim-timeout",
|
||||||
|
"claim_type": "symptom",
|
||||||
|
"claim_text": "订单接口出现超时",
|
||||||
|
"support_level": "direct",
|
||||||
|
"evidence_bindings": [
|
||||||
|
{
|
||||||
|
"tool_invocation_id": 1,
|
||||||
|
"tool_name": "query_logs",
|
||||||
|
"evidence_excerpt": "order api timeout"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"claim_id": "claim-db-primary",
|
||||||
|
"claim_type": "root_cause",
|
||||||
|
"claim_text": "数据库主库故障导致订单超时",
|
||||||
|
"support_level": "weak",
|
||||||
|
"evidence_bindings": [
|
||||||
|
{
|
||||||
|
"tool_invocation_id": 1,
|
||||||
|
"tool_name": "query_logs",
|
||||||
|
"evidence_excerpt": "order api timeout"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"hypotheses": [],
|
||||||
|
"recommended_actions": [
|
||||||
|
{
|
||||||
|
"action_text": "补充查询数据库主库状态和错误日志",
|
||||||
|
"reason": "当前只有订单接口超时日志"
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"missing_info": ["数据库主库状态", "数据库错误日志"]
|
||||||
|
},
|
||||||
|
"claim_checks": [
|
||||||
|
{
|
||||||
|
"claim_id": "claim-timeout",
|
||||||
|
"claim_text": "订单接口出现超时",
|
||||||
|
"claim_type": "symptom",
|
||||||
|
"verification": "direct_observation",
|
||||||
|
"detail": "日志直接记录 order api timeout",
|
||||||
|
"evidence_refs": [
|
||||||
|
{
|
||||||
|
"tool_invocation_id": 1,
|
||||||
|
"tool_name": "query_logs"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"claim_id": "claim-db-primary",
|
||||||
|
"claim_text": "数据库主库故障导致订单超时",
|
||||||
|
"claim_type": "root_cause",
|
||||||
|
"verification": "unsupported",
|
||||||
|
"detail": "日志只能证明订单接口超时,不能证明数据库主库故障",
|
||||||
|
"evidence_refs": [
|
||||||
|
{
|
||||||
|
"tool_invocation_id": 1,
|
||||||
|
"tool_name": "query_logs"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"facts_checked": [
|
||||||
|
{
|
||||||
|
"fact": "订单接口出现超时",
|
||||||
|
"verification": "direct_evidence",
|
||||||
|
"detail": "日志直接记录 order api timeout",
|
||||||
|
"evidence_refs": [
|
||||||
|
{
|
||||||
|
"tool_invocation_id": 1,
|
||||||
|
"tool_name": "query_logs"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"fact": "数据库主库故障导致订单超时",
|
||||||
|
"verification": "unsupported",
|
||||||
|
"detail": "日志只能证明订单接口超时,不能证明数据库主库故障",
|
||||||
|
"evidence_refs": [
|
||||||
|
{
|
||||||
|
"tool_invocation_id": 1,
|
||||||
|
"tool_name": "query_logs"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"composer_output": {
|
||||||
|
"status": "valid",
|
||||||
|
"answer_summary": "日志显示订单接口超时,但数据库方向证据不足。",
|
||||||
|
"recommended_actions": [
|
||||||
|
{
|
||||||
|
"action_text": "补充查询数据库主库状态和错误日志",
|
||||||
|
"reason": "当前只有订单接口超时日志"
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"user_facing_answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:日志显示订单接口出现超时。\n\n仍需补充信息:目前没有数据库故障日志或主库状态证据,不能把该方向写成确认根因。"
|
||||||
|
},
|
||||||
|
"tool_trace_summary": [
|
||||||
|
{
|
||||||
|
"tool_name": "query_logs",
|
||||||
|
"success": true,
|
||||||
|
"evidence_level": "direct"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"steps": [],
|
||||||
|
"toolInvocations": [
|
||||||
|
{
|
||||||
|
"id": 1,
|
||||||
|
"sessionId": "eval-unsupported-claim-filtering",
|
||||||
|
"toolName": "query_logs",
|
||||||
|
"outputPreview": "order api timeout",
|
||||||
|
"success": true
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"summary": {
|
||||||
|
"persistedStepCount": 3,
|
||||||
|
"returnedStepCount": 3,
|
||||||
|
"persistedToolCallCount": 1,
|
||||||
|
"returnedToolCallCount": 1,
|
||||||
|
"hasVerifierEvaluation": true,
|
||||||
|
"hasFeedback": false
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -1,13 +1,14 @@
|
|||||||
{
|
{
|
||||||
"totalCases" : 5,
|
"totalCases" : 8,
|
||||||
"passedCases" : 5,
|
"passedCases" : 8,
|
||||||
"passRate" : 1.0,
|
"passRate" : 1.0,
|
||||||
"verdictDistribution" : {
|
"verdictDistribution" : {
|
||||||
"PASS" : 2,
|
"PASS" : 2,
|
||||||
"LOW_CONFID" : 3
|
"LOW_CONFID" : 5,
|
||||||
|
"REJECT" : 1
|
||||||
},
|
},
|
||||||
"averageToolCallCount" : 2.0,
|
"averageToolCallCount" : 1.625,
|
||||||
"averageDurationMs" : 45800.0,
|
"averageDurationMs" : 44875.0,
|
||||||
"results" : [ {
|
"results" : [ {
|
||||||
"caseId" : "payment-timeout",
|
"caseId" : "payment-timeout",
|
||||||
"title" : "Payment API timeout",
|
"title" : "Payment API timeout",
|
||||||
@@ -21,6 +22,9 @@
|
|||||||
"query_logs" : true,
|
"query_logs" : true,
|
||||||
"query_metrics" : true
|
"query_metrics" : true
|
||||||
},
|
},
|
||||||
|
"gatekeeperStatus" : null,
|
||||||
|
"composerStatus" : null,
|
||||||
|
"claimCheckCount" : null,
|
||||||
"toolCallCount" : 3,
|
"toolCallCount" : 3,
|
||||||
"durationMs" : 42000
|
"durationMs" : 42000
|
||||||
}, {
|
}, {
|
||||||
@@ -35,6 +39,9 @@
|
|||||||
"lookup_knowledge" : true,
|
"lookup_knowledge" : true,
|
||||||
"query_logs" : true
|
"query_logs" : true
|
||||||
},
|
},
|
||||||
|
"gatekeeperStatus" : null,
|
||||||
|
"composerStatus" : null,
|
||||||
|
"claimCheckCount" : null,
|
||||||
"toolCallCount" : 2,
|
"toolCallCount" : 2,
|
||||||
"durationMs" : 51000
|
"durationMs" : 51000
|
||||||
}, {
|
}, {
|
||||||
@@ -48,6 +55,9 @@
|
|||||||
"evidenceCoverage" : {
|
"evidenceCoverage" : {
|
||||||
"query_logs" : true
|
"query_logs" : true
|
||||||
},
|
},
|
||||||
|
"gatekeeperStatus" : null,
|
||||||
|
"composerStatus" : null,
|
||||||
|
"claimCheckCount" : null,
|
||||||
"toolCallCount" : 1,
|
"toolCallCount" : 1,
|
||||||
"durationMs" : 36000
|
"durationMs" : 36000
|
||||||
}, {
|
}, {
|
||||||
@@ -62,6 +72,9 @@
|
|||||||
"query_metrics" : true,
|
"query_metrics" : true,
|
||||||
"query_logs" : true
|
"query_logs" : true
|
||||||
},
|
},
|
||||||
|
"gatekeeperStatus" : null,
|
||||||
|
"composerStatus" : null,
|
||||||
|
"claimCheckCount" : null,
|
||||||
"toolCallCount" : 2,
|
"toolCallCount" : 2,
|
||||||
"durationMs" : 47000
|
"durationMs" : 47000
|
||||||
}, {
|
}, {
|
||||||
@@ -76,7 +89,58 @@
|
|||||||
"query_metrics" : true,
|
"query_metrics" : true,
|
||||||
"query_logs" : true
|
"query_logs" : true
|
||||||
},
|
},
|
||||||
|
"gatekeeperStatus" : null,
|
||||||
|
"composerStatus" : null,
|
||||||
|
"claimCheckCount" : null,
|
||||||
"toolCallCount" : 2,
|
"toolCallCount" : 2,
|
||||||
"durationMs" : 53000
|
"durationMs" : 53000
|
||||||
|
}, {
|
||||||
|
"caseId" : "gatekeeper-fabricated-invocation",
|
||||||
|
"title" : "Gatekeeper fabricated invocation",
|
||||||
|
"passed" : true,
|
||||||
|
"failedChecks" : [ ],
|
||||||
|
"verdict" : "REJECT",
|
||||||
|
"matchedKeywordCount" : 3,
|
||||||
|
"requiredKeywordCount" : 3,
|
||||||
|
"evidenceCoverage" : {
|
||||||
|
"query_logs" : true
|
||||||
|
},
|
||||||
|
"gatekeeperStatus" : "fail",
|
||||||
|
"composerStatus" : "valid",
|
||||||
|
"claimCheckCount" : 1,
|
||||||
|
"toolCallCount" : 1,
|
||||||
|
"durationMs" : 39000
|
||||||
|
}, {
|
||||||
|
"caseId" : "unsupported-claim-filtering",
|
||||||
|
"title" : "Unsupported claim filtering",
|
||||||
|
"passed" : true,
|
||||||
|
"failedChecks" : [ ],
|
||||||
|
"verdict" : "LOW_CONFID",
|
||||||
|
"matchedKeywordCount" : 2,
|
||||||
|
"requiredKeywordCount" : 2,
|
||||||
|
"evidenceCoverage" : {
|
||||||
|
"query_logs" : true
|
||||||
|
},
|
||||||
|
"gatekeeperStatus" : "pass",
|
||||||
|
"composerStatus" : "valid",
|
||||||
|
"claimCheckCount" : 2,
|
||||||
|
"toolCallCount" : 1,
|
||||||
|
"durationMs" : 44000
|
||||||
|
}, {
|
||||||
|
"caseId" : "composer-fallback-no-raw-json",
|
||||||
|
"title" : "Composer fallback no raw JSON",
|
||||||
|
"passed" : true,
|
||||||
|
"failedChecks" : [ ],
|
||||||
|
"verdict" : "LOW_CONFID",
|
||||||
|
"matchedKeywordCount" : 2,
|
||||||
|
"requiredKeywordCount" : 2,
|
||||||
|
"evidenceCoverage" : {
|
||||||
|
"query_metrics" : true
|
||||||
|
},
|
||||||
|
"gatekeeperStatus" : "pass",
|
||||||
|
"composerStatus" : "composer_malformed",
|
||||||
|
"claimCheckCount" : 2,
|
||||||
|
"toolCallCount" : 1,
|
||||||
|
"durationMs" : 47000
|
||||||
} ]
|
} ]
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -1,22 +1,26 @@
|
|||||||
# Diagnosis Eval Report
|
# Diagnosis Eval Report
|
||||||
|
|
||||||
- Total cases: 5
|
- Total cases: 8
|
||||||
- Passed cases: 5
|
- Passed cases: 8
|
||||||
- Pass rate: 100.00%
|
- Pass rate: 100.00%
|
||||||
- Average tool calls: 2.00
|
- Average tool calls: 1.63
|
||||||
- Average duration ms: 45800.00
|
- Average duration ms: 44875.00
|
||||||
|
|
||||||
## Verdict Distribution
|
## Verdict Distribution
|
||||||
|
|
||||||
- PASS: 2
|
- PASS: 2
|
||||||
- LOW_CONFID: 3
|
- LOW_CONFID: 5
|
||||||
|
- REJECT: 1
|
||||||
|
|
||||||
## Cases
|
## Cases
|
||||||
|
|
||||||
| Case | Result | Verdict | Keywords | Tool Calls | Duration ms | Failed Checks |
|
| Case | Result | Verdict | Gatekeeper | Composer | Claim Checks | Keywords | Tool Calls | Duration ms | Failed Checks |
|
||||||
| --- | --- | --- | --- | ---: | ---: | --- |
|
| --- | --- | --- | --- | --- | ---: | --- | ---: | ---: | --- |
|
||||||
| payment-timeout | PASS | PASS | 3/3 | 3 | 42000 | - |
|
| payment-timeout | PASS | PASS | - | - | - | 3/3 | 3 | 42000 | - |
|
||||||
| mysql-pool-exhausted | PASS | LOW_CONFID | 3/3 | 2 | 51000 | - |
|
| mysql-pool-exhausted | PASS | LOW_CONFID | - | - | - | 3/3 | 2 | 51000 | - |
|
||||||
| redis-timeout | PASS | LOW_CONFID | 2/2 | 1 | 36000 | - |
|
| redis-timeout | PASS | LOW_CONFID | - | - | - | 2/2 | 1 | 36000 | - |
|
||||||
| slow-response | PASS | PASS | 2/2 | 2 | 47000 | - |
|
| slow-response | PASS | PASS | - | - | - | 2/2 | 2 | 47000 | - |
|
||||||
| jvm-memory-risk | PASS | LOW_CONFID | 3/3 | 2 | 53000 | - |
|
| jvm-memory-risk | PASS | LOW_CONFID | - | - | - | 3/3 | 2 | 53000 | - |
|
||||||
|
| gatekeeper-fabricated-invocation | PASS | REJECT | fail | valid | 1 | 3/3 | 1 | 39000 | - |
|
||||||
|
| unsupported-claim-filtering | PASS | LOW_CONFID | pass | valid | 2 | 2/2 | 1 | 44000 | - |
|
||||||
|
| composer-fallback-no-raw-json | PASS | LOW_CONFID | pass | composer_malformed | 2 | 2/2 | 1 | 47000 | - |
|
||||||
|
|||||||
+94
-147
@@ -1,203 +1,150 @@
|
|||||||
# Diagnosis Eval Data Schema
|
# Diagnosis Eval Data Schema
|
||||||
|
|
||||||
这份文档记录评测基准里的数据结构。口语化理解就是:
|
这份文档记录 `mvp/eval` 固定评测集的数据结构。评测器读取保存好的 trace fixture,用确定性规则判断这次 Agent 运行是否满足预期。
|
||||||
|
|
||||||
```text
|
```text
|
||||||
用例文件说“我要考什么”
|
case 文件:我要考什么
|
||||||
trace 文件说“Agent 实际做了什么”
|
fixture 文件:Agent 实际做了什么
|
||||||
评测结果说“这次有没有跑偏”
|
评测结果:这条 case 是否通过,哪里失败
|
||||||
汇总报告说“整体稳定性怎么样”
|
baseline report:整套固定集当前认可的结果
|
||||||
```
|
```
|
||||||
|
|
||||||
当前这套评测是代码规则判断,不是再调用一个 LLM 来打分。
|
当前评测不调用 LLM 打分。
|
||||||
|
|
||||||
## 1. 用例定义
|
## 1. Case 定义
|
||||||
|
|
||||||
文件:`mvp/eval/cases/diagnosis-cases.json`
|
文件:`mvp/eval/cases/diagnosis-cases.json`
|
||||||
|
|
||||||
每一条 case 是一个固定考题,告诉评测器“这个问题应该看哪些点、需要哪些证据、哪些结论可以接受”。
|
每条 case 定义一个固定诊断场景。
|
||||||
|
|
||||||
```json
|
```json
|
||||||
{
|
{
|
||||||
"id": "payment-timeout",
|
"id": "unsupported-claim-filtering",
|
||||||
"title": "Payment API timeout",
|
"title": "Unsupported claim filtering",
|
||||||
"question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
|
"question": "订单超时是否可以确认由数据库主库故障导致?",
|
||||||
"traceFixture": "payment-timeout-pass.json",
|
"traceFixture": "unsupported-claim-filtering-low-confid.json",
|
||||||
"expectedRootCauseKeywords": ["支付", "超时", "连接池"],
|
"expectedRootCauseKeywords": ["超时", "证据"],
|
||||||
"minKeywordMatches": 2,
|
"minKeywordMatches": 2,
|
||||||
"requiredEvidenceTools": ["lookup_knowledge", "query_logs", "query_metrics"],
|
"requiredEvidenceTools": ["query_logs"],
|
||||||
"allowedVerdicts": ["PASS", "LOW_CONFID"],
|
"allowedVerdicts": ["LOW_CONFID"],
|
||||||
"forbiddenAnswerKeywords": ["无证据确定"]
|
"forbiddenAnswerKeywords": ["已经确认"],
|
||||||
|
"requireV2AuditClosure": true,
|
||||||
|
"requireClaimChecks": true,
|
||||||
|
"requireComposerOutput": true,
|
||||||
|
"expectedGatekeeperStatuses": ["pass"],
|
||||||
|
"expectedComposerStatuses": ["valid"],
|
||||||
|
"forbiddenConfirmedClaimKeywords": ["主库故障"]
|
||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
字段说明:
|
字段说明:
|
||||||
|
|
||||||
| 字段 | 意思 | 评测器怎么用 |
|
| 字段 | 含义 | 评测器怎么用 |
|
||||||
| --- | --- | --- |
|
| --- | --- | --- |
|
||||||
| `id` | 这条用例的唯一名字 | 出现在报告里,方便定位是哪条 case 挂了 |
|
| `id` | case 唯一标识 | 出现在报告中 |
|
||||||
| `title` | 给人看的标题 | 出现在结果里,方便快速理解场景 |
|
| `title` | 可读标题 | 出现在报告中 |
|
||||||
| `question` | 要问 Agent 的问题 | fixture 模式下不会真的发送给 Agent,但它记录了这条 case 的原始输入 |
|
| `question` | 原始用户问题 | 用于说明场景,fixture 模式不会真实发送给 Agent |
|
||||||
| `traceFixture` | 对应的 trace 文件名 | 评测器会去 `fixtures/` 目录加载这个文件 |
|
| `traceFixture` | 对应 fixture 文件名 | 从 `mvp/eval/fixtures` 加载 |
|
||||||
| `expectedRootCauseKeywords` | 最终回答里希望看到的关键点 | 评测器会在 `session.answer` 里做关键词命中检查 |
|
| `expectedRootCauseKeywords` | 最终答案应覆盖的关键词 | 在 `session.answer` 中做包含判断 |
|
||||||
| `minKeywordMatches` | 至少要命中几个关键词 | 命中数低于这个值,就认为根因覆盖不够 |
|
| `minKeywordMatches` | 最少命中关键词数 | 低于该值则失败 |
|
||||||
| `requiredEvidenceTools` | 这条 case 至少应该用到哪些证据工具 | 评测器会检查 trace 里是否出现这些工具 |
|
| `requiredEvidenceTools` | 必须出现的证据工具 | 从 `toolInvocations` 和 `tool_trace_summary` 中收集 |
|
||||||
| `allowedVerdicts` | Verifier 允许给出的结论 | 比如 `PASS` 或 `LOW_CONFID`,不在列表里就失败 |
|
| `allowedVerdicts` | 允许的 Verifier verdict | verdict 不在列表中则失败 |
|
||||||
| `forbiddenAnswerKeywords` | 回答里不应该出现的危险说法 | 命中这些词,说明回答可能过度自信或不符合降级策略 |
|
| `forbiddenAnswerKeywords` | 最终答案禁止出现的词 | 用于拦截过度自信或危险表达 |
|
||||||
|
| `requireV2AuditClosure` | 是否要求 V2 审计闭环字段 | 要求 `gatekeeper_result`、`claim_checks`、`composer_output` 存在,并检查 raw JSON 泄漏 |
|
||||||
|
| `requireClaimChecks` | 是否要求 `claim_checks` | 要求 claim check 数组存在且非空 |
|
||||||
|
| `requireComposerOutput` | 是否要求 `composer_output` | 要求 Composer 审计存在并带 `status` |
|
||||||
|
| `expectedGatekeeperStatuses` | 允许的 Gatekeeper 状态 | 实际 `gatekeeper_result.status` 不在列表中则失败 |
|
||||||
|
| `expectedComposerStatuses` | 允许的 Composer 状态 | 实际 `composer_output.status` 不在列表中则失败 |
|
||||||
|
| `forbiddenConfirmedClaimKeywords` | 不得进入最终答案的未支持结论关键词 | 用于证明 unsupported/external_unknown claim 被过滤 |
|
||||||
|
|
||||||
## 2. Trace Fixture
|
## 2. Trace Fixture
|
||||||
|
|
||||||
目录:`mvp/eval/fixtures/*.json`
|
目录:`mvp/eval/fixtures/*.json`
|
||||||
|
|
||||||
trace fixture 是一次 Agent 运行后的“留痕快照”。评测器不会关心整个 trace 的所有字段,只读取当前能支撑基准判断的字段。
|
fixture 是一次 Agent 运行后的 trace 快照。评测器只读取当前规则需要的字段。
|
||||||
|
|
||||||
当前会读取这些字段:
|
| Trace 字段 | 含义 | 评测器怎么用 |
|
||||||
|
|
||||||
| Trace 字段 | 意思 | 评测器怎么用 |
|
|
||||||
| --- | --- | --- |
|
| --- | --- | --- |
|
||||||
| `session.answer` | Agent 最终给用户的回答 | 用来检查根因关键词和禁用词 |
|
| `session.answer` | 最终用户答案 | 检查关键词、禁用词、unsupported claim 泄漏、raw JSON 泄漏 |
|
||||||
| `session.totalDurationMs` | 这次运行耗时 | 进入报告,帮助观察性能是否明显变差 |
|
| `session.totalDurationMs` | 运行耗时 | 进入报告 |
|
||||||
| `session.selfEvaluation.verifier_evaluation.verdict` | Verifier 对最终回答的判断 | 必须存在,并且要落在 case 的 `allowedVerdicts` 里 |
|
| `session.selfEvaluation.verifier_evaluation.verdict` | Verifier 判定 | 必须存在并符合 case 的 `allowedVerdicts` |
|
||||||
| `session.selfEvaluation.verifier_evaluation.executor_structured_output.claims[*].evidence_bindings` | Executor 结构化输出里的确认事实证据绑定 | 如果 trace 中存在结构化 Executor 输出,每条 confirmed claim 必须有证据绑定 |
|
| `session.selfEvaluation.verifier_evaluation.gatekeeper_result.status` | Gatekeeper 结果 | V2 case 必须存在;`fail` 不允许搭配 `PASS` |
|
||||||
| `session.selfEvaluation.verifier_evaluation.tool_trace_summary[*].tool_name` | Verifier 总结里看到的工具证据 | 用来补充判断证据工具是否出现 |
|
| `session.selfEvaluation.verifier_evaluation.claim_checks` | Verifier V2 claim 级校验 | V2 case 必须存在;每项需要 `claim_id`、`verification`、`detail` |
|
||||||
| `toolInvocations[*].toolName` | Agent 实际调用过的工具名 | 用来检查 `requiredEvidenceTools` 是否满足 |
|
| `session.selfEvaluation.verifier_evaluation.composer_output.status` | Composer 渲染状态 | V2 case 必须存在;记录 `valid`、`composer_malformed` 等 |
|
||||||
| `toolInvocations[*].success` | 工具调用是否成功 | 当前主要保留在 trace 里,后续可以升级成更严格的成功率检查 |
|
| `session.selfEvaluation.verifier_evaluation.executor_structured_output.claims[*].evidence_bindings` | Executor claim 证据绑定 | 如果结构化输出存在,每条 claim 需要证据绑定 |
|
||||||
|
| `session.selfEvaluation.verifier_evaluation.tool_trace_summary[*].tool_name` | Verifier 看到的工具证据 | 用于补充证据工具覆盖 |
|
||||||
简单说,trace 里最重要的是三类信息:
|
| `toolInvocations[*].toolName` | 实际调用工具名 | 用于检查 `requiredEvidenceTools` |
|
||||||
|
| `toolInvocations[*].success` | 工具调用是否成功 | 当前主要保留在 trace 中 |
|
||||||
```text
|
|
||||||
最终回答:它说了什么
|
|
||||||
工具证据:它查了什么
|
|
||||||
Verifier:它自己有没有承认这个结论可靠
|
|
||||||
```
|
|
||||||
|
|
||||||
## 3. 单条评测结果
|
## 3. 单条评测结果
|
||||||
|
|
||||||
Java 类型:`DiagnosisEvalResult`
|
Java 类型:`DiagnosisEvalResult`
|
||||||
|
|
||||||
这是每条 case 跑完之后的判断结果。
|
| 字段 | 含义 |
|
||||||
|
|
||||||
| 字段 | 意思 |
|
|
||||||
| --- | --- |
|
| --- | --- |
|
||||||
| `caseId` | 对应的 case id |
|
| `caseId` | 对应 case id |
|
||||||
| `title` | case 标题 |
|
| `title` | case 标题 |
|
||||||
| `passed` | 这条 case 是否通过 |
|
| `passed` | 该 case 是否通过 |
|
||||||
| `failedChecks` | 没通过的具体原因,比如缺工具、关键词不够、verdict 不允许 |
|
| `failedChecks` | 失败原因列表 |
|
||||||
| `verdict` | 从 trace 里读出来的 Verifier verdict |
|
| `verdict` | 从 trace 中读到的 Verifier verdict |
|
||||||
| `matchedKeywordCount` | 最终回答命中的关键词数量 |
|
| `matchedKeywordCount` | 最终答案命中的关键词数量 |
|
||||||
| `requiredKeywordCount` | case 定义里一共有多少个关键词 |
|
| `requiredKeywordCount` | case 配置的关键词数量 |
|
||||||
| `evidenceCoverage` | 每个必需工具是否出现,例如 `{ "query_logs": true }` |
|
| `evidenceCoverage` | 每个必需工具是否出现 |
|
||||||
| `toolCallCount` | 本次 trace 里工具调用总数 |
|
| `gatekeeperStatus` | 读到的 `gatekeeper_result.status` |
|
||||||
| `durationMs` | 本次 trace 的耗时 |
|
| `composerStatus` | 读到的 `composer_output.status` |
|
||||||
|
| `claimCheckCount` | `claim_checks` 数量 |
|
||||||
判断通过的口语化规则:
|
| `toolCallCount` | trace 中工具调用总数 |
|
||||||
|
| `durationMs` | trace 总耗时 |
|
||||||
```text
|
|
||||||
回答要说到关键点
|
|
||||||
该查的证据工具要查到
|
|
||||||
Verifier 的结论要在可接受范围内
|
|
||||||
回答不能出现危险的过度自信表达
|
|
||||||
如果是 REJECT,就必须走降级模板
|
|
||||||
```
|
|
||||||
|
|
||||||
## 4. 汇总报告
|
## 4. 汇总报告
|
||||||
|
|
||||||
Java 类型:`DiagnosisEvalReport`
|
Java 类型:`DiagnosisEvalReport`
|
||||||
|
|
||||||
这是整个基准集跑完之后的总结果。
|
| 字段 | 含义 |
|
||||||
|
|
||||||
| 字段 | 意思 |
|
|
||||||
| --- | --- |
|
| --- | --- |
|
||||||
| `totalCases` | 总共评测了多少条 case |
|
| `totalCases` | case 总数 |
|
||||||
| `passedCases` | 通过了多少条 |
|
| `passedCases` | 通过数 |
|
||||||
| `passRate` | 通过率,范围是 `0.0` 到 `1.0` |
|
| `passRate` | 通过率,范围 `0.0` 到 `1.0` |
|
||||||
| `verdictDistribution` | Verifier verdict 的分布,比如有几个 `PASS`、几个 `LOW_CONFID` |
|
| `verdictDistribution` | Verifier verdict 分布 |
|
||||||
| `averageToolCallCount` | 平均每条 case 调用了多少次工具 |
|
| `averageToolCallCount` | 平均工具调用数 |
|
||||||
| `averageDurationMs` | 平均耗时 |
|
| `averageDurationMs` | 平均耗时 |
|
||||||
| `results` | 每条 case 的详细结果列表 |
|
| `results` | 单条 case 结果列表 |
|
||||||
|
|
||||||
## 5. 怎么看这个基准
|
## 5. Stage 5 V2 审计闭环规则
|
||||||
|
|
||||||
这套结构不是为了证明 Agent 永远正确,而是为了在每次改 prompt、工具、检索、Verifier 之后,有一个固定尺子能回答:
|
阶段 5 关注的是“前四段链路是否能被固定评测证明”:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
以前能过的诊断题,现在还过不过?
|
Executor structured output
|
||||||
它是不是少查了某些证据?
|
-> Gatekeeper deterministic audit
|
||||||
它是不是变得更自信但证据不足?
|
-> Verifier claim_checks
|
||||||
它是不是开始输出不该说的话?
|
-> Composer filtered final answer
|
||||||
它是不是明显变慢了?
|
|
||||||
```
|
```
|
||||||
|
|
||||||
所以面试里可以这样讲:
|
新增确定性规则:
|
||||||
|
|
||||||
```text
|
- V2 case 必须有 `gatekeeper_result`、`claim_checks`、`composer_output`。
|
||||||
我没有只看一次 demo 效果,而是把典型诊断场景固化成 case。
|
- `gatekeeper_result.status = fail` 时,Verifier verdict 不能是 `PASS`。
|
||||||
每条 case 都定义预期关键点、必需证据工具和可接受的 verifier 结论。
|
- `claim_checks[*].verification` 只能是 `direct_observation`、`reasonable_inference`、`overstated`、`unsupported`、`external_unknown`、`contradicted`。
|
||||||
Agent 每次运行会留下 trace,评测器用代码规则读取 trace,输出结构化报告。
|
- Composer 输出必须记录 `status`。
|
||||||
这样我改 Agent 的时候,可以用同一套基准判断有没有行为回退。
|
- 最终答案不能泄漏 `executor_evidence_v2`、`answer_version`、`evidence_bindings`、`claim_id`。
|
||||||
```
|
- case 配置的 `forbiddenConfirmedClaimKeywords` 不能出现在最终答案里。
|
||||||
|
|
||||||
## 6. Baseline Diff
|
## 6. Baseline Diff
|
||||||
|
|
||||||
Baseline diff 是拿两份 report 做对比:
|
Baseline diff 比较两份 report:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
baseline report:以前认可的基准结果
|
baseline report:已经认可的基准结果
|
||||||
current report:这次改动后跑出来的新结果
|
current report:当前代码/fixture 跑出的结果
|
||||||
diff report:告诉你哪里变好了、哪里变差了、哪里只是变了
|
diff report:结构化列出退化、改善和普通变化
|
||||||
```
|
```
|
||||||
|
|
||||||
Java 类型:
|
主要退化信号:
|
||||||
|
|
||||||
- `DiagnosisEvalDiffReport`
|
- pass rate 下降。
|
||||||
- `DiagnosisEvalDiffItem`
|
- case 从通过变失败。
|
||||||
|
- 必需证据工具从有变无。
|
||||||
`DiagnosisEvalDiffReport` 字段:
|
- 关键词命中减少。
|
||||||
|
- 工具调用或耗时明显上升。
|
||||||
| 字段 | 意思 |
|
- verdict 分布变化。
|
||||||
| --- | --- |
|
|
||||||
| `baselineTotalCases` | baseline 里有多少条 case |
|
|
||||||
| `currentTotalCases` | current 里有多少条 case |
|
|
||||||
| `baselinePassedCases` | baseline 通过了多少条 |
|
|
||||||
| `currentPassedCases` | current 通过了多少条 |
|
|
||||||
| `baselinePassRate` | baseline 通过率 |
|
|
||||||
| `currentPassRate` | current 通过率 |
|
|
||||||
| `regressionCount` | 退化项数量 |
|
|
||||||
| `improvementCount` | 改善项数量 |
|
|
||||||
| `changedCount` | 普通变化项数量 |
|
|
||||||
| `hasRegression` | 是否存在退化 |
|
|
||||||
| `items` | 具体 diff 明细 |
|
|
||||||
|
|
||||||
`DiagnosisEvalDiffItem` 字段:
|
|
||||||
|
|
||||||
| 字段 | 意思 |
|
|
||||||
| --- | --- |
|
|
||||||
| `type` | `REGRESSION`、`IMPROVEMENT` 或 `CHANGED` |
|
|
||||||
| `scope` | `aggregate` 表示整体指标,`case` 表示单条 case |
|
|
||||||
| `caseId` | 如果是单条 case 变化,这里记录 case id |
|
|
||||||
| `metric` | 哪个指标变了,比如 `passRate` 或 `evidenceCoverage.query_logs` |
|
|
||||||
| `baselineValue` | baseline 里的值 |
|
|
||||||
| `currentValue` | current 里的值 |
|
|
||||||
| `delta` | 数值变化量;非数值变化为空 |
|
|
||||||
| `message` | 给人看的变化说明 |
|
|
||||||
|
|
||||||
口语化判断规则:
|
|
||||||
|
|
||||||
```text
|
|
||||||
pass rate 下降:退化
|
|
||||||
case 从通过变失败:退化
|
|
||||||
证据工具从有变没有:退化
|
|
||||||
关键词命中变少:退化
|
|
||||||
工具调用或耗时升高:成本上升,记为退化信号
|
|
||||||
verdict 分布变化:记录变化,供人工判断是否符合预期
|
|
||||||
```
|
|
||||||
|
|
||||||
面试里可以这样讲:
|
|
||||||
|
|
||||||
```text
|
|
||||||
我把 baseline report 和当前 report 做结构化 diff。
|
|
||||||
它不是再问 LLM,而是用代码比较固定字段。
|
|
||||||
如果某个 case 从 PASS 变 FAIL,或者 query_logs 证据没了,
|
|
||||||
diff 会直接标成 regression。
|
|
||||||
这样 Agent 改动可以用固定基准做回归判断。
|
|
||||||
```
|
|
||||||
|
|||||||
File diff suppressed because it is too large
Load Diff
@@ -8,6 +8,7 @@
|
|||||||
| ISS-004 | Executor 域级检索水位控制(Phase 2) | 低 | 待规划 | [ISS-004-executor-domain-hard-limit.md](ISS-004-executor-domain-hard-limit.md) |
|
| ISS-004 | Executor 域级检索水位控制(Phase 2) | 低 | 待规划 | [ISS-004-executor-domain-hard-limit.md](ISS-004-executor-domain-hard-limit.md) |
|
||||||
| ISS-005 | 证据链补齐与降级契约收敛 | 高 | 已归档 | [ISS-005-evidence-trace-hardening.md](ISS-005-evidence-trace-hardening.md) |
|
| ISS-005 | 证据链补齐与降级契约收敛 | 高 | 已归档 | [ISS-005-evidence-trace-hardening.md](ISS-005-evidence-trace-hardening.md) |
|
||||||
| ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) |
|
| ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) |
|
||||||
|
| ISS-007 | Verifier 证据摘要保真与工具命中质量问题 | 高 | 已实施 | [ISS-007-verifier-evidence-summary-fidelity.md](ISS-007-verifier-evidence-summary-fidelity.md) |
|
||||||
| executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [executor-evidence-attribution-hallucination.md](executor-evidence-attribution-hallucination.md) |
|
| executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [executor-evidence-attribution-hallucination.md](executor-evidence-attribution-hallucination.md) |
|
||||||
| executor-self-evidence-loop-design-note | Executor 自证循环与证据摘要链路设计记录 | 高 | 已形成方向 | [executor-self-evidence-loop-design-note.md](executor-self-evidence-loop-design-note.md) |
|
| executor-self-evidence-loop-design-note | Executor 自证循环与证据摘要链路设计记录 | 高 | 已形成方向 | [executor-self-evidence-loop-design-note.md](executor-self-evidence-loop-design-note.md) |
|
||||||
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) |
|
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) |
|
||||||
|
|||||||
@@ -0,0 +1,2 @@
|
|||||||
|
schema: spec-driven
|
||||||
|
created: 2026-07-08
|
||||||
@@ -0,0 +1,49 @@
|
|||||||
|
## Context
|
||||||
|
|
||||||
|
Executor Structured Output V2 introduced three audit layers after the original fixed-case eval harness was created:
|
||||||
|
|
||||||
|
- Gatekeeper result in `selfEvaluation.verifier_evaluation.gatekeeper_result`
|
||||||
|
- Verifier claim-level checks in `selfEvaluation.verifier_evaluation.claim_checks`
|
||||||
|
- Composer result in `selfEvaluation.verifier_evaluation.composer_output`
|
||||||
|
|
||||||
|
The existing harness proves that a saved trace has expected answer keywords, evidence tools, verdicts, and basic structured claim bindings. It does not yet prove that the V2 audit chain is internally consistent or that Composer filtered unsupported claims before writing the final user answer.
|
||||||
|
|
||||||
|
## Goals / Non-Goals
|
||||||
|
|
||||||
|
**Goals:**
|
||||||
|
|
||||||
|
- Make V2 audit fields part of offline regression checks.
|
||||||
|
- Catch physical/protocol regressions: missing Gatekeeper audit, PASS after Gatekeeper fail, missing claim checks, missing Composer audit, raw Executor JSON leakage, and unsupported claims leaking into the final answer.
|
||||||
|
- Add fixture cases that exercise the new checks without starting the application.
|
||||||
|
- Keep the evaluator deterministic and easy to explain to another implementation agent.
|
||||||
|
|
||||||
|
**Non-Goals:**
|
||||||
|
|
||||||
|
- Do not redesign Planner, Executor, Gatekeeper, Verifier, or Composer.
|
||||||
|
- Do not introduce an LLM-based evaluator.
|
||||||
|
- Do not require live MySQL, Redis, Milvus, or application startup for the baseline fixture tests.
|
||||||
|
- Do not add a broad new database schema; trace audit remains read from existing JSON fields.
|
||||||
|
|
||||||
|
## Decisions
|
||||||
|
|
||||||
|
1. Extend eval case expectations instead of hardcoding every V2 rule globally.
|
||||||
|
|
||||||
|
Some legacy or intentionally partial fixtures may not contain all V2 audit fields. Case-level expectations let the fixed baseline explicitly say which trace must prove Gatekeeper, claim checks, Composer, or leakage prevention. The default can remain backward-compatible while new V2 fixtures opt in to stricter checks.
|
||||||
|
|
||||||
|
2. Keep consistency checks lexical and deterministic.
|
||||||
|
|
||||||
|
The evaluator will not determine semantic truth from scratch. It will compare configured keywords against the final answer and audit fields. This is enough to catch the intended regression class: unsupported or contradicted claims being rendered as confirmed final answers.
|
||||||
|
|
||||||
|
3. Treat Gatekeeper failure as a hard regression if paired with `PASS`.
|
||||||
|
|
||||||
|
The production chain already guards this. The eval harness should independently fail any trace where `gatekeeper_result.status=fail` and Verifier verdict remains `PASS`, because that means the audit layer can no longer be trusted.
|
||||||
|
|
||||||
|
4. Record audit signals in eval results.
|
||||||
|
|
||||||
|
Per-case results should expose the observed Gatekeeper status, Composer status, and claim-check count so baseline JSON/Markdown reports remain useful during review.
|
||||||
|
|
||||||
|
## Risks / Trade-offs
|
||||||
|
|
||||||
|
- [Risk] Keyword-based final-answer checks can miss paraphrases. -> Mitigation: use them only for regression-sensitive fixtures where the unsafe claim keyword is deliberately fixed.
|
||||||
|
- [Risk] Adding too many case fields makes fixtures harder to maintain. -> Mitigation: keep expectation fields small and optional.
|
||||||
|
- [Risk] Baseline report changes may look like a product behavior change. -> Mitigation: document that this phase changes only eval fixtures/rules unless a production bug is discovered and fixed.
|
||||||
@@ -0,0 +1,37 @@
|
|||||||
|
## Why
|
||||||
|
|
||||||
|
Executor Structured Output V2 已经完成 Executor、Gatekeeper、Verifier、Composer 四段主链路改造,但现有离线评测仍主要检查最终答案关键词、证据工具覆盖和 Verifier verdict。阶段 5 需要把新增的审计字段纳入固定回归门禁,证明 `Planner -> Executor -> Gatekeeper -> Verifier -> Composer -> final answer` 链路不是只在单次 demo 中可用。
|
||||||
|
|
||||||
|
## What Changes
|
||||||
|
|
||||||
|
- Extend the diagnosis eval harness so V2 audit fields are validated as first-class regression checks.
|
||||||
|
- Add fixture coverage for Gatekeeper failure, claim verification filtering, Composer fallback, and raw JSON leakage prevention.
|
||||||
|
- Update baseline reports to reflect the expanded fixed fixture set.
|
||||||
|
- Update eval documentation so another agent can understand the background, stages, data fields, and acceptance commands.
|
||||||
|
- No production protocol change is intended in this phase; production chain behavior should remain unchanged unless tests reveal a bug.
|
||||||
|
|
||||||
|
## Capabilities
|
||||||
|
|
||||||
|
### New Capabilities
|
||||||
|
|
||||||
|
- None.
|
||||||
|
|
||||||
|
### Modified Capabilities
|
||||||
|
|
||||||
|
- `diagnosis-eval-harness`: add V2 audit-closure requirements for Gatekeeper, Verifier claim checks, Composer output, and final-answer consistency.
|
||||||
|
|
||||||
|
## Impact
|
||||||
|
|
||||||
|
- Affected code:
|
||||||
|
- `src/main/java/com/superbiz/agent/eval/*`
|
||||||
|
- `src/test/java/com/superbiz/agent/eval/*`
|
||||||
|
- Affected data:
|
||||||
|
- `mvp/eval/cases/diagnosis-cases.json`
|
||||||
|
- `mvp/eval/fixtures/*.json`
|
||||||
|
- `mvp/eval/reports/baseline-report.*`
|
||||||
|
- `mvp/eval/schema.md`
|
||||||
|
- `mvp/eval/README.md`
|
||||||
|
- Verification:
|
||||||
|
- focused evaluator tests
|
||||||
|
- relevant Executor/Gatekeeper/Verifier/Composer integration tests
|
||||||
|
- OpenSpec validation and archive
|
||||||
+38
@@ -0,0 +1,38 @@
|
|||||||
|
## ADDED Requirements
|
||||||
|
|
||||||
|
### Requirement: Evaluation harness SHALL validate Executor V2 audit closure
|
||||||
|
The evaluation harness SHALL be able to validate the V2 audit chain from Executor structured output through Gatekeeper, Verifier claim checks, Composer output, and final user answer.
|
||||||
|
|
||||||
|
#### Scenario: Required V2 audit fields are present
|
||||||
|
- **WHEN** an evaluation case requires V2 audit closure
|
||||||
|
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.gatekeeper_result` exists
|
||||||
|
- **AND** it SHALL verify that `selfEvaluation.verifier_evaluation.claim_checks` exists
|
||||||
|
- **AND** it SHALL verify that `selfEvaluation.verifier_evaluation.composer_output` exists
|
||||||
|
|
||||||
|
#### Scenario: Gatekeeper failure cannot pass verification
|
||||||
|
- **WHEN** a trace has `gatekeeper_result.status` equal to `fail`
|
||||||
|
- **THEN** the evaluator SHALL fail the case if `selfEvaluation.verifier_evaluation.verdict` is `PASS`
|
||||||
|
|
||||||
|
#### Scenario: Claim checks are auditable
|
||||||
|
- **WHEN** an evaluation case requires claim checks
|
||||||
|
- **THEN** the evaluator SHALL verify that each claim check includes `claim_id`, `verification`, and `detail`
|
||||||
|
- **AND** each verification SHALL be one of the V2 claim verification values recognized by the Verifier contract
|
||||||
|
|
||||||
|
#### Scenario: Composer output is auditable
|
||||||
|
- **WHEN** an evaluation case requires Composer output
|
||||||
|
- **THEN** the evaluator SHALL verify that `composer_output` records whether parsed output or fallback rendering was used
|
||||||
|
|
||||||
|
### Requirement: Evaluation harness SHALL prevent unsupported claims from leaking into final answers
|
||||||
|
The evaluation harness SHALL detect configured unsafe or unsupported claim text when it appears in the final user-facing answer.
|
||||||
|
|
||||||
|
#### Scenario: Unsupported final-answer claim is rejected
|
||||||
|
- **WHEN** an evaluation case declares forbidden confirmed-claim keywords
|
||||||
|
- **THEN** the evaluator SHALL fail the case if the final answer contains any of those keywords
|
||||||
|
|
||||||
|
#### Scenario: Raw Executor JSON is not user-facing
|
||||||
|
- **WHEN** a trace is evaluated under V2 audit closure
|
||||||
|
- **THEN** the evaluator SHALL fail the case if the final answer contains raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`
|
||||||
|
|
||||||
|
#### Scenario: Composer fallback still avoids raw JSON leakage
|
||||||
|
- **WHEN** a trace records Composer fallback rendering
|
||||||
|
- **THEN** the evaluator SHALL still enforce final-answer raw JSON leakage checks
|
||||||
@@ -0,0 +1,23 @@
|
|||||||
|
## 1. Eval Contract
|
||||||
|
|
||||||
|
- [x] 1.1 Extend `DiagnosisEvalCase` with optional V2 audit expectations.
|
||||||
|
- [x] 1.2 Extend `DiagnosisEvalResult` and report output with observed audit signals.
|
||||||
|
- [x] 1.3 Add deterministic evaluator checks for Gatekeeper, claim checks, Composer output, and raw JSON leakage.
|
||||||
|
|
||||||
|
## 2. Fixtures and Baseline
|
||||||
|
|
||||||
|
- [x] 2.1 Add fixed eval cases for Gatekeeper failure, unsupported claim filtering, and Composer fallback leakage prevention.
|
||||||
|
- [x] 2.2 Add matching trace fixtures for each new case.
|
||||||
|
- [x] 2.3 Regenerate baseline JSON and Markdown reports.
|
||||||
|
|
||||||
|
## 3. Documentation
|
||||||
|
|
||||||
|
- [x] 3.1 Update eval README with stage 5 scope and verification commands.
|
||||||
|
- [x] 3.2 Rewrite eval schema documentation so V2 audit fields and case expectations are readable.
|
||||||
|
|
||||||
|
## 4. Verification and Archive
|
||||||
|
|
||||||
|
- [x] 4.1 Add or update unit tests for the new V2 audit checks.
|
||||||
|
- [x] 4.2 Run focused evaluator and agent-chain tests.
|
||||||
|
- [x] 4.3 Validate and archive the OpenSpec change.
|
||||||
|
- [x] 4.4 Commit the completed phase 5 changes.
|
||||||
@@ -0,0 +1,73 @@
|
|||||||
|
# Decisions
|
||||||
|
|
||||||
|
## Discover Context
|
||||||
|
|
||||||
|
- Source issue: `mvp/issues/ISS-007-verifier-evidence-summary-fidelity.md`.
|
||||||
|
- Related existing specs: `chat-verifier-agent`, `evidence-trace-hardening`.
|
||||||
|
- Related devflow records: `executor-gatekeeper-hook`, `executor-verifier-claim-checks`, `executor-composer-final-answer`, `evidence-trace-hardening`.
|
||||||
|
- Current repo instruction requested semantic code search and LSP confirmation before code changes; those tools are not exposed in this environment, so implementation will use `rg`, direct code reading, and focused tests as fallback evidence.
|
||||||
|
|
||||||
|
## Classification
|
||||||
|
|
||||||
|
- sm-flow scale: `standard`.
|
||||||
|
- Reason: internal contract change across recorder, hook, Gatekeeper, verifier prompt, mock tools, tests, and E2E validation.
|
||||||
|
|
||||||
|
## Question Pool
|
||||||
|
|
||||||
|
| Question | Mode | Resolution |
|
||||||
|
|---|---|---|
|
||||||
|
| Should Planner change or add `scope_contract`? | user-interview, already resolved in issue | No. This change is prompt-first and does not alter Planner. |
|
||||||
|
| Should Gatekeeper move to Executor hook for retry? | user-interview, already resolved in issue | No. Gatekeeper remains in Verifier hook path for this version. |
|
||||||
|
| Should new DB tables be added for evidence refs or audit? | user-interview, already resolved in issue | No. Use `tool_invocation.retrieval_details.evidence_refs` and existing `self_evaluation.verifier_evaluation.gatekeeper_result`. |
|
||||||
|
| Should `tool_trace_summary` remain the primary evidence source? | user-interview, already resolved in issue | No. It becomes navigation/audit context; verified claim-local excerpts become primary evidence for derivability. |
|
||||||
|
| Should full JSONPath be supported? | user-interview, already resolved in issue | No. Only stable raw paths listed in design are supported. |
|
||||||
|
|
||||||
|
## Cross-artifact Alignment
|
||||||
|
|
||||||
|
| Check | Status |
|
||||||
|
|---|---|
|
||||||
|
| Issue background and target -> proposal | aligned |
|
||||||
|
| Proposal scope and non-goals -> design | aligned |
|
||||||
|
| Design protocol and risks -> specs/tasks | aligned |
|
||||||
|
| Specs observable behavior -> tasks acceptance | aligned |
|
||||||
|
|
||||||
|
## Architecture Audit
|
||||||
|
|
||||||
|
The change keeps the existing agent orchestration and only hardens the evidence payload between Executor, Gatekeeper, and Verifier. The highest coupling risk is transition compatibility from plural `source_invocation_ids` to singular `source_invocation_id`; the implementation must accept old data but only allow precise new bindings to pass. No new DB table is introduced, reducing migration risk. Verifier prompt and effective verdict guardrails must agree on `gatekeeper_result.severity`, otherwise the model could still produce a PASS that runtime later must downgrade.
|
||||||
|
|
||||||
|
## Commit Gate
|
||||||
|
|
||||||
|
Passed on 2026-07-08.
|
||||||
|
|
||||||
|
- `openspec validate verifier-evidence-reference-fidelity --strict`: passed.
|
||||||
|
- `openspec validate --specs`: passed.
|
||||||
|
- File completeness: proposal, design, specs, tasks present.
|
||||||
|
- Consistency: proposal -> design -> specs -> tasks aligned.
|
||||||
|
- Apply authorization: user requested continuing implementation and fixing autonomously; proceed to Apply.
|
||||||
|
|
||||||
|
## Apply Verification
|
||||||
|
|
||||||
|
Focused tests:
|
||||||
|
|
||||||
|
- `mvn "-Dtest=ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,QueryLogsToolsTest,ChatServiceSequentialAgentTest" test`: passed, 38 tests.
|
||||||
|
- `mvn "-Dtest=ExecutorGatekeeperServiceTest" test`: passed, 9 tests after adding the old-invocation-without-evidence-refs case.
|
||||||
|
|
||||||
|
OpenSpec:
|
||||||
|
|
||||||
|
- `openspec validate verifier-evidence-reference-fidelity --strict`: passed.
|
||||||
|
- `openspec validate --specs`: passed.
|
||||||
|
|
||||||
|
End-to-end sessions:
|
||||||
|
|
||||||
|
| Case | Session | Result | Gatekeeper |
|
||||||
|
|---|---|---|---|
|
||||||
|
| HikariCP positive | `iss007-hikari-positive-20260708-1553` | `PASS`; answer confirmed order-service HikariCP timeout and pool saturation logs | `pass / none`, checked_bindings=2 |
|
||||||
|
| HikariCP negative | `iss007-hikari-negative-20260708-1555` | `LOW_CONFID`; no `generic-service` pollution; no positive inventory-service confirmation | `fail / low_confid` due incomplete negative evidence references |
|
||||||
|
| HighMemoryUsage positive | `iss007-memory-positive-20260708-1558` | `PASS`; answer confirmed HighMemoryUsage 91% without confirming memory leak | `pass / none`, checked_bindings=1 |
|
||||||
|
| SlowResponse positive | `iss007-slow-positive-20260708-1600` | `PASS`; answer confirmed SlowResponse and slow request logs without DB-pool root cause | `pass / none`, checked_bindings=7 |
|
||||||
|
| Narrow HighCPUUsage | `iss007-narrow-highcpu-20260708-1602` | `PASS`; answer only covered payment-service HighCPUUsage | `pass / none`, checked_bindings=1 |
|
||||||
|
|
||||||
|
Residual observation:
|
||||||
|
|
||||||
|
- HikariCP negative no longer returns `generic-service`, and no-hit rows persist `evidence_status=no_evidence`.
|
||||||
|
- The model still issued an extra broad HikariCP query without the `inventory-service` filter and used order-service as a context claim. This is a remaining narrow-scope behavior issue, not a mock data pollution issue. It is acceptable for this change because the final verdict did not become a false positive for inventory-service.
|
||||||
@@ -0,0 +1,180 @@
|
|||||||
|
# Design
|
||||||
|
|
||||||
|
## Data Flow
|
||||||
|
|
||||||
|
```text
|
||||||
|
Evidence tool result
|
||||||
|
-> ToolInvocationRecorder
|
||||||
|
persists tool_invocation.retrieval_details.evidence_refs
|
||||||
|
-> Executor
|
||||||
|
emits claim-local evidence_bindings
|
||||||
|
-> VerifierInputHook
|
||||||
|
preserves structured payload and runs Gatekeeper
|
||||||
|
-> ExecutorGatekeeperService
|
||||||
|
validates reference authenticity
|
||||||
|
-> Verifier
|
||||||
|
checks derivability from verified excerpts
|
||||||
|
-> ChatService / Composer
|
||||||
|
enforces effective verdict and safe final answer
|
||||||
|
```
|
||||||
|
|
||||||
|
## Evidence Reference Contract
|
||||||
|
|
||||||
|
`tool_invocation.retrieval_details.evidence_refs` is an array of minimal evidence references:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"evidence_refs": [
|
||||||
|
{
|
||||||
|
"raw_path": "$.alerts[0]",
|
||||||
|
"text": "HighMemoryUsage firing, service=order-service, current=91%, duration=15m"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Supported `raw_path` formats in this change:
|
||||||
|
|
||||||
|
- `$.alerts[i]` for `query_metrics`
|
||||||
|
- `$.logs[i]` for `query_logs`
|
||||||
|
- `$.evidence_blocks[i]` for `lookup_knowledge`
|
||||||
|
|
||||||
|
Unsupported in this change:
|
||||||
|
|
||||||
|
- Deep JSONPath such as `$.alerts[1].description`
|
||||||
|
- Nested aliases such as `$.retrieval_details.evidence_blocks[0]`
|
||||||
|
- Filter expressions
|
||||||
|
- Tool-specific metadata beyond `raw_path` and `text`
|
||||||
|
|
||||||
|
## Executor Binding Contract
|
||||||
|
|
||||||
|
Executor V2 claim bindings should use:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"tool_name": "query_metrics",
|
||||||
|
"source_invocation_id": 12345,
|
||||||
|
"raw_path": "$.alerts[0]",
|
||||||
|
"evidence_excerpt": "HighMemoryUsage firing, service=order-service, current=91%, duration=15m"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Compatibility:
|
||||||
|
|
||||||
|
- Legacy `source_invocation_ids` may still be read for transition.
|
||||||
|
- New precise validation requires singular `source_invocation_id` plus `raw_path`.
|
||||||
|
- Missing `raw_path` is `LOW_CONFID`, not `PASS`.
|
||||||
|
|
||||||
|
## Gatekeeper Severity
|
||||||
|
|
||||||
|
Gatekeeper output includes:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"status": "fail",
|
||||||
|
"severity": "reject",
|
||||||
|
"checked_bindings": [],
|
||||||
|
"failed_rules": [],
|
||||||
|
"warnings": [],
|
||||||
|
"errors": []
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Severity mapping:
|
||||||
|
|
||||||
|
- `status=pass`, `severity=none`: all checked bindings are authentic.
|
||||||
|
- `status=fail`, `severity=low_confid`: evidence is missing or incomplete but not fabricated.
|
||||||
|
- `status=fail`, `severity=reject`: Executor cites fabricated or mismatched evidence.
|
||||||
|
|
||||||
|
`reject` cases:
|
||||||
|
|
||||||
|
- Invocation ID does not exist.
|
||||||
|
- Invocation belongs to another session.
|
||||||
|
- Tool name mismatches persisted invocation.
|
||||||
|
- `raw_path` is not present in `retrieval_details.evidence_refs`.
|
||||||
|
- `evidence_excerpt` clearly mismatches the system-side `text`.
|
||||||
|
|
||||||
|
`low_confid` cases:
|
||||||
|
|
||||||
|
- Confirmed claim has empty `evidence_bindings`.
|
||||||
|
- `source_invocation_id` exists but `raw_path` is missing.
|
||||||
|
- Old invocation lacks `evidence_refs`.
|
||||||
|
- `evidence_excerpt` is too short or generic to compare.
|
||||||
|
- Output is incomplete without concrete fabricated IDs or paths.
|
||||||
|
|
||||||
|
## Verifier Behavior
|
||||||
|
|
||||||
|
Verifier should treat `gatekeeper_result.severity` as a hard boundary:
|
||||||
|
|
||||||
|
- `reject`: effective result cannot be `PASS`; fabricated-reference cases should become `REJECT`.
|
||||||
|
- `low_confid`: effective result cannot be `PASS`; missing-reference cases should become `LOW_CONFID`.
|
||||||
|
- `none`: Verifier judges whether claim text is derivable from verified excerpts.
|
||||||
|
|
||||||
|
`tool_trace_summary` remains available as navigation and audit context, but claim-local verified excerpts are the primary evidence for derivability.
|
||||||
|
|
||||||
|
## Hook Placement
|
||||||
|
|
||||||
|
Gatekeeper remains in the Verifier hook path for this change. Retry or rollback into Executor is not implemented in this stage.
|
||||||
|
|
||||||
|
Existing ad hoc verifier-hook validation should be replaced by Gatekeeper output. The hook may normalize compatibility fields, but it should not independently decide pass/fail outside Gatekeeper semantics.
|
||||||
|
|
||||||
|
## Auto-backfill
|
||||||
|
|
||||||
|
`VerifierInputHook` may only auto-fill a missing `source_invocation_id` when exactly one invocation candidate exists for the binding's `tool_name`.
|
||||||
|
|
||||||
|
Rules:
|
||||||
|
|
||||||
|
- Never bulk-fill multiple IDs.
|
||||||
|
- Never auto-fill `raw_path`.
|
||||||
|
- Add a warning when auto-fill occurs.
|
||||||
|
- Auto-filled binding without `raw_path` must remain `LOW_CONFID`.
|
||||||
|
|
||||||
|
## Prompt Constraints
|
||||||
|
|
||||||
|
Executor prompt changes are prompt-first, contract-later:
|
||||||
|
|
||||||
|
- Executor is an evidence collector and micro-fact extractor.
|
||||||
|
- Narrow-scope questions should usually produce one claim and at most two claims.
|
||||||
|
- Do not limit evidence binding count.
|
||||||
|
- Emit observation or negative observation, not root-cause certainty, for narrow confirmation questions.
|
||||||
|
- Runbook and skill content cannot become current incident facts.
|
||||||
|
- Recommended actions, if any, are evidence-collection next steps, not remediation actions.
|
||||||
|
|
||||||
|
## HikariCP Mock Behavior
|
||||||
|
|
||||||
|
`query_logs` should support positive mock hits for:
|
||||||
|
|
||||||
|
- `HikariCP`
|
||||||
|
- `HikariPool`
|
||||||
|
- `connection pool`
|
||||||
|
- `数据库连接池`
|
||||||
|
- `连接池耗尽`
|
||||||
|
- `active=50/50`
|
||||||
|
- `waiting`
|
||||||
|
- `request timed out after 30000ms`
|
||||||
|
- `order-service`
|
||||||
|
|
||||||
|
Positive output should include order-service HikariCP log records. No-hit should return `logs=[]` and `evidence_status=no_evidence`, not `generic-service` placeholder logs.
|
||||||
|
|
||||||
|
## Audit
|
||||||
|
|
||||||
|
Gatekeeper result must be persisted under:
|
||||||
|
|
||||||
|
```text
|
||||||
|
diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result
|
||||||
|
```
|
||||||
|
|
||||||
|
The minimum persisted fields are:
|
||||||
|
|
||||||
|
- `status`
|
||||||
|
- `severity`
|
||||||
|
- `checked_bindings`
|
||||||
|
- `failed_rules`
|
||||||
|
- `warnings`
|
||||||
|
- `errors`
|
||||||
|
|
||||||
|
## Interface Impact
|
||||||
|
|
||||||
|
Impact level: L2 internal contract change.
|
||||||
|
|
||||||
|
The change affects internal Agent payload JSON, persisted audit JSON, prompts, and test fixtures. It does not add public HTTP endpoints or new database tables.
|
||||||
@@ -0,0 +1,62 @@
|
|||||||
|
# verifier-evidence-reference-fidelity
|
||||||
|
|
||||||
|
## Problem
|
||||||
|
|
||||||
|
Recent end-to-end checks show that some narrow diagnosis questions still become `LOW_CONFID` even when the raw tool output and Executor evidence excerpts contain enough concrete evidence. The failure is caused by evidence being compressed or lost before Verifier reasoning, plus mock log no-hit behavior that can return `generic-service` placeholder logs.
|
||||||
|
|
||||||
|
Current weak points:
|
||||||
|
|
||||||
|
- Verifier still relies too heavily on `tool_trace_summary.output_summary`.
|
||||||
|
- Executor evidence bindings identify tool invocations but do not precisely locate evidence inside the invocation output.
|
||||||
|
- Gatekeeper validates invocation IDs and tool names, but does not yet verify `raw_path` and excerpt fidelity.
|
||||||
|
- `query_logs` mock data cannot reliably produce a positive HikariCP connection-pool exhaustion path and can pollute no-hit results with placeholder logs.
|
||||||
|
- Narrow-scope Executor answers can still over-expand into unrelated claims.
|
||||||
|
|
||||||
|
## Proposed Change
|
||||||
|
|
||||||
|
Introduce a claim-local evidence reference protocol:
|
||||||
|
|
||||||
|
```text
|
||||||
|
tool raw output
|
||||||
|
-> ToolInvocationRecorder stores retrieval_details.evidence_refs
|
||||||
|
-> Executor outputs claim + source_invocation_id + raw_path + evidence_excerpt
|
||||||
|
-> Gatekeeper verifies that the reference is real
|
||||||
|
-> Verifier judges whether verified evidence can derive the claim
|
||||||
|
-> Composer only expresses Verifier-allowed material
|
||||||
|
```
|
||||||
|
|
||||||
|
The first implementation keeps the orchestration unchanged. Gatekeeper remains in the Verifier input hook path. Planner `scope_contract` is out of scope for this change.
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
- Add minimal `retrieval_details.evidence_refs` extraction for `query_metrics`, `query_logs`, and `lookup_knowledge`.
|
||||||
|
- Tighten Executor evidence bindings to prefer singular `source_invocation_id`, stable `raw_path`, and `evidence_excerpt`.
|
||||||
|
- Extend Gatekeeper to validate `source_invocation_id + raw_path + evidence_excerpt`.
|
||||||
|
- Add Gatekeeper severity: `none`, `low_confid`, `reject`.
|
||||||
|
- Make Verifier consume verified `evidence_excerpt` as the primary claim-local evidence.
|
||||||
|
- Tighten `VerifierInputHook` auto-backfill: only unique invocation candidate, never `raw_path`, and no `PASS` without a precise reference.
|
||||||
|
- Fix HikariCP mock log matching and no-hit behavior.
|
||||||
|
- Update Executor and Verifier prompts for narrow-scope and derivability behavior.
|
||||||
|
- Add focused tests and end-to-end checks for the minimum acceptance matrix.
|
||||||
|
|
||||||
|
## Non-goals
|
||||||
|
|
||||||
|
- Do not change Planner output.
|
||||||
|
- Do not implement Planner `scope_contract`.
|
||||||
|
- Do not add new database tables.
|
||||||
|
- Do not implement a general JSONPath engine.
|
||||||
|
- Do not make `tool_trace_summary` the primary evidence source again.
|
||||||
|
- Do not allow Executor `diagnosis_summary` or `user_facing_answer` to re-enter the V2 contract.
|
||||||
|
|
||||||
|
## Context Constraints
|
||||||
|
|
||||||
|
- `tool_invocation.retrieval_details` is the preferred place for tool-specific structured details.
|
||||||
|
- `DiagnosisSession.selfEvaluation.verifier_evaluation.gatekeeper_result` is the existing audit container and must be preserved.
|
||||||
|
- `read_skill` / runbook guidance is not incident evidence.
|
||||||
|
- Existing Composer routing must keep raw Executor JSON out of normal user answers.
|
||||||
|
|
||||||
|
## Risks
|
||||||
|
|
||||||
|
- Existing tests or prompts may still assume plural `source_invocation_ids`.
|
||||||
|
- Some old invocations will not have `evidence_refs`; those must downgrade to `LOW_CONFID`, not `PASS`.
|
||||||
|
- Similarity checks must tolerate formatting changes without accepting unrelated text.
|
||||||
+110
@@ -0,0 +1,110 @@
|
|||||||
|
## ADDED Requirements
|
||||||
|
|
||||||
|
### Requirement: Executor evidence bindings SHALL support precise evidence references
|
||||||
|
Executor V2 evidence bindings SHALL support precise evidence references that locate evidence inside a persisted tool invocation.
|
||||||
|
|
||||||
|
#### Scenario: Precise evidence binding contains invocation path and excerpt
|
||||||
|
- **WHEN** Executor binds evidence to a claim
|
||||||
|
- **THEN** the binding SHOULD include singular `source_invocation_id`
|
||||||
|
- **AND** the binding SHOULD include `raw_path`
|
||||||
|
- **AND** the binding SHALL include `tool_name` and `evidence_excerpt`
|
||||||
|
- **AND** the `raw_path` SHALL be interpreted relative to the referenced tool invocation's `retrieval_details.evidence_refs`
|
||||||
|
|
||||||
|
#### Scenario: Legacy plural invocation ids remain compatibility only
|
||||||
|
- **WHEN** Executor emits legacy `source_invocation_ids`
|
||||||
|
- **THEN** the system MAY read them for compatibility
|
||||||
|
- **AND** they SHALL NOT be sufficient for a precise Gatekeeper pass without `raw_path`
|
||||||
|
|
||||||
|
### Requirement: Gatekeeper SHALL validate evidence reference fidelity
|
||||||
|
Gatekeeper SHALL validate that Executor evidence bindings point to real current-session evidence references before Verifier uses them as primary evidence.
|
||||||
|
|
||||||
|
#### Scenario: Valid precise binding passes
|
||||||
|
- **WHEN** a binding's `source_invocation_id` exists in the current session
|
||||||
|
- **AND** the binding's `tool_name` matches the persisted invocation
|
||||||
|
- **AND** the binding's `raw_path` exists in `retrieval_details.evidence_refs`
|
||||||
|
- **AND** the binding's `evidence_excerpt` is supported by the matching evidence ref text
|
||||||
|
- **THEN** Gatekeeper SHALL return `status=pass`
|
||||||
|
- **AND** Gatekeeper SHALL return `severity=none`
|
||||||
|
|
||||||
|
#### Scenario: Missing raw path is low confidence
|
||||||
|
- **WHEN** a binding references an existing invocation
|
||||||
|
- **AND** the binding omits `raw_path`
|
||||||
|
- **THEN** Gatekeeper SHALL return `status=fail`
|
||||||
|
- **AND** Gatekeeper SHALL return `severity=low_confid`
|
||||||
|
- **AND** the effective verifier result SHALL NOT be `PASS`
|
||||||
|
|
||||||
|
#### Scenario: Old invocation without evidence refs is low confidence
|
||||||
|
- **WHEN** a binding references an existing invocation
|
||||||
|
- **AND** the invocation does not contain `retrieval_details.evidence_refs`
|
||||||
|
- **THEN** Gatekeeper SHALL return `status=fail`
|
||||||
|
- **AND** Gatekeeper SHALL return `severity=low_confid`
|
||||||
|
- **AND** the effective verifier result SHALL NOT be `PASS`
|
||||||
|
|
||||||
|
#### Scenario: Unknown raw path is rejected
|
||||||
|
- **WHEN** a binding references an existing invocation
|
||||||
|
- **AND** the binding's `raw_path` is absent from that invocation's `retrieval_details.evidence_refs`
|
||||||
|
- **THEN** Gatekeeper SHALL return `status=fail`
|
||||||
|
- **AND** Gatekeeper SHALL return `severity=reject`
|
||||||
|
- **AND** `failed_rules` SHALL include `evidence.raw_path`
|
||||||
|
|
||||||
|
#### Scenario: Mismatched excerpt is rejected
|
||||||
|
- **WHEN** a binding references an existing invocation and raw path
|
||||||
|
- **AND** the binding's `evidence_excerpt` is not supported by the matching system-side evidence ref text
|
||||||
|
- **THEN** Gatekeeper SHALL return `status=fail`
|
||||||
|
- **AND** Gatekeeper SHALL return `severity=reject`
|
||||||
|
- **AND** `failed_rules` SHALL include `evidence.excerpt_mismatch`
|
||||||
|
|
||||||
|
### Requirement: Verifier SHALL use verified claim-local evidence for derivability
|
||||||
|
Verifier SHALL judge structured claims primarily against Gatekeeper-verified claim-local evidence excerpts.
|
||||||
|
|
||||||
|
#### Scenario: Verified excerpt supports direct observation
|
||||||
|
- **WHEN** `gatekeeper_result.severity=none`
|
||||||
|
- **AND** a claim's verified evidence excerpts directly contain the claim's concrete facts
|
||||||
|
- **THEN** Verifier MAY classify that claim as `direct_observation`
|
||||||
|
|
||||||
|
#### Scenario: Tool trace summary is navigation context
|
||||||
|
- **WHEN** `executor_structured_output.claims[].evidence_bindings` are available
|
||||||
|
- **THEN** Verifier SHALL use `tool_trace_summary` as navigation and audit context
|
||||||
|
- **AND** it SHALL NOT require `tool_trace_summary.output_summary` to contain every fact already present in verified claim-local evidence
|
||||||
|
|
||||||
|
### Requirement: Gatekeeper severity SHALL constrain effective verdict
|
||||||
|
Runtime effective verdict calculation SHALL treat Gatekeeper severity as a hard upper bound.
|
||||||
|
|
||||||
|
#### Scenario: Reject severity prevents PASS
|
||||||
|
- **WHEN** `gatekeeper_result.severity=reject`
|
||||||
|
- **AND** the Verifier model returns `verdict=PASS`
|
||||||
|
- **THEN** ChatService SHALL downgrade the effective verdict
|
||||||
|
- **AND** the effective verdict SHALL be `REJECT`
|
||||||
|
|
||||||
|
#### Scenario: Low confidence severity prevents PASS
|
||||||
|
- **WHEN** `gatekeeper_result.severity=low_confid`
|
||||||
|
- **AND** the Verifier model returns `verdict=PASS`
|
||||||
|
- **THEN** ChatService SHALL downgrade the effective verdict
|
||||||
|
- **AND** the effective verdict SHALL be `LOW_CONFID`
|
||||||
|
|
||||||
|
#### Scenario: Gatekeeper audit includes severity
|
||||||
|
- **WHEN** verifier evaluation is persisted
|
||||||
|
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result` SHALL include `status`, `severity`, `checked_bindings`, `failed_rules`, `warnings`, and `errors`
|
||||||
|
|
||||||
|
### Requirement: Verifier input hook SHALL only perform narrow compatibility backfill
|
||||||
|
The verifier input hook SHALL avoid converting broad tool summaries into precise evidence references.
|
||||||
|
|
||||||
|
#### Scenario: Unique invocation candidate may be backfilled
|
||||||
|
- **WHEN** an evidence binding omits `source_invocation_id`
|
||||||
|
- **AND** exactly one current-session invocation exists for the binding's `tool_name`
|
||||||
|
- **THEN** the hook MAY backfill `source_invocation_id`
|
||||||
|
- **AND** it SHALL add a Gatekeeper warning describing the auto-backfill
|
||||||
|
|
||||||
|
#### Scenario: Raw path is never backfilled
|
||||||
|
- **WHEN** an evidence binding omits `raw_path`
|
||||||
|
- **THEN** the hook SHALL NOT synthesize `raw_path`
|
||||||
|
- **AND** Gatekeeper SHALL treat the binding as not precise enough to pass
|
||||||
|
|
||||||
|
### Requirement: Executor prompt SHALL constrain narrow-scope over-expansion
|
||||||
|
The Executor prompt SHALL instruct Executor to keep narrow confirmation questions focused on observation-level claims.
|
||||||
|
|
||||||
|
#### Scenario: Narrow scope produces minimal observation claims
|
||||||
|
- **WHEN** the user asks to confirm one specific service, alert, or symptom
|
||||||
|
- **THEN** Executor SHOULD output the minimum necessary claims, normally one and at most two
|
||||||
|
- **AND** those claims SHALL be `observation` or `negative_observation` unless current-session evidence proves more
|
||||||
|
- **AND** Executor SHALL NOT emit unrelated root-cause, remediation, or excluded-topic claims as confirmed facts
|
||||||
+42
@@ -0,0 +1,42 @@
|
|||||||
|
## ADDED Requirements
|
||||||
|
|
||||||
|
### Requirement: Evidence tools SHALL persist minimal evidence refs
|
||||||
|
Evidence-bearing tool invocations SHALL persist claim-addressable evidence references in `tool_invocation.retrieval_details.evidence_refs`.
|
||||||
|
|
||||||
|
#### Scenario: Metrics alerts produce evidence refs
|
||||||
|
- **WHEN** a `query_metrics` invocation returns alert entries
|
||||||
|
- **THEN** the persisted retrieval details SHALL include one `evidence_refs` item per usable alert
|
||||||
|
- **AND** each item SHALL include `raw_path` formatted as `$.alerts[i]`
|
||||||
|
- **AND** each item SHALL include bounded `text` containing concrete alert facts such as alert name, state, service, current value, and duration when available
|
||||||
|
|
||||||
|
#### Scenario: Logs produce evidence refs
|
||||||
|
- **WHEN** a `query_logs` invocation returns log entries
|
||||||
|
- **THEN** the persisted retrieval details SHALL include one `evidence_refs` item per usable log
|
||||||
|
- **AND** each item SHALL include `raw_path` formatted as `$.logs[i]`
|
||||||
|
- **AND** each item SHALL include bounded `text` containing concrete log facts such as timestamp, level, service, and message when available
|
||||||
|
|
||||||
|
#### Scenario: Knowledge lookup produces evidence refs
|
||||||
|
- **WHEN** a `lookup_knowledge` invocation returns evidence blocks
|
||||||
|
- **THEN** the persisted retrieval details SHALL include one `evidence_refs` item per usable evidence block
|
||||||
|
- **AND** each item SHALL include `raw_path` formatted as `$.evidence_blocks[i]`
|
||||||
|
- **AND** each item SHALL include bounded `text` containing concrete block content, title, or source when available
|
||||||
|
|
||||||
|
#### Scenario: Evidence ref extraction does not infer diagnosis
|
||||||
|
- **WHEN** the recorder creates `evidence_refs`
|
||||||
|
- **THEN** it SHALL only copy or format concrete tool output fields
|
||||||
|
- **AND** it SHALL NOT infer root cause, remediation, or diagnosis conclusions
|
||||||
|
|
||||||
|
### Requirement: Log mock no-hit semantics SHALL avoid placeholder evidence
|
||||||
|
The log query mock SHALL distinguish positive mock evidence from no-hit results without using placeholder service logs as evidence.
|
||||||
|
|
||||||
|
#### Scenario: HikariCP positive query returns order-service pool evidence
|
||||||
|
- **WHEN** a `query_logs` request targets `order-service` and HikariCP connection-pool exhaustion terms
|
||||||
|
- **THEN** the tool SHALL return order-service HikariCP-related log entries
|
||||||
|
- **AND** the returned evidence SHALL include concrete terms such as `HikariPool`, `active=50/50`, `waiting`, or `request timed out after 30000ms`
|
||||||
|
- **AND** it SHALL NOT return `generic-service` placeholder logs
|
||||||
|
|
||||||
|
#### Scenario: HikariCP no-hit query returns no evidence
|
||||||
|
- **WHEN** a `query_logs` request targets a service without matching HikariCP mock evidence
|
||||||
|
- **THEN** the tool SHALL return an empty `logs` array
|
||||||
|
- **AND** it SHALL mark the output as `evidence_status=no_evidence`
|
||||||
|
- **AND** it SHALL NOT return `generic-service` placeholder logs
|
||||||
@@ -0,0 +1,79 @@
|
|||||||
|
# Tasks
|
||||||
|
|
||||||
|
## 1. Evidence reference extraction
|
||||||
|
|
||||||
|
- [x] Add `evidence_refs` extraction in `ToolInvocationRecorder` for `query_metrics` alert arrays.
|
||||||
|
- [x] Add `evidence_refs` extraction in `ToolInvocationRecorder` for `query_logs` log arrays.
|
||||||
|
- [x] Add `evidence_refs` extraction in `ToolInvocationRecorder` for `lookup_knowledge` evidence blocks.
|
||||||
|
- [x] Add focused recorder tests for `$.alerts[i]`, `$.logs[i]`, and `$.evidence_blocks[i]`.
|
||||||
|
|
||||||
|
Acceptance:
|
||||||
|
|
||||||
|
- Persisted `retrieval_details` contains minimal `raw_path` and `text`.
|
||||||
|
- Extraction does not infer root cause or diagnosis.
|
||||||
|
|
||||||
|
## 2. Gatekeeper reference fidelity
|
||||||
|
|
||||||
|
- [x] Extend `ExecutorGatekeeperService` to read singular `source_invocation_id`, `raw_path`, and `evidence_excerpt`.
|
||||||
|
- [x] Keep legacy `source_invocation_ids` compatibility where needed, but require `raw_path` for precise pass.
|
||||||
|
- [x] Validate invocation existence, session ownership, tool name, `raw_path`, and excerpt similarity.
|
||||||
|
- [x] Add `severity` and checked binding details to Gatekeeper output.
|
||||||
|
- [x] Add tests for valid reference, missing `raw_path`, missing `evidence_refs`, unknown `raw_path`, mismatched excerpt, fabricated invocation ID, and tool mismatch.
|
||||||
|
|
||||||
|
Acceptance:
|
||||||
|
|
||||||
|
- Valid precise references pass.
|
||||||
|
- Missing precision downgrades to `LOW_CONFID`.
|
||||||
|
- Fabricated or mismatched references become `REJECT`.
|
||||||
|
|
||||||
|
## 3. Verifier input hook and prompts
|
||||||
|
|
||||||
|
- [x] Tighten `VerifierInputHook` auto-backfill to only unique single invocation candidates.
|
||||||
|
- [x] Prevent auto-filled bindings without `raw_path` from passing Gatekeeper.
|
||||||
|
- [x] Add auto-backfill warnings into Gatekeeper/audit output.
|
||||||
|
- [x] Update `chat-executor-prompt.md` with evidence-reference and narrow-scope constraints.
|
||||||
|
- [x] Update `chat-verifier-prompt.md` so verified excerpts are the primary derivability evidence.
|
||||||
|
- [x] Add focused hook and prompt-sensitive tests where practical.
|
||||||
|
|
||||||
|
Acceptance:
|
||||||
|
|
||||||
|
- Hook no longer bulk-fills invocation IDs.
|
||||||
|
- Verifier can see `gatekeeper_result.severity`.
|
||||||
|
- Executor is instructed to output precise references and avoid over-expansion.
|
||||||
|
|
||||||
|
## 4. Effective verdict and audit persistence
|
||||||
|
|
||||||
|
- [x] Ensure `severity=reject` prevents effective `PASS` and maps to `REJECT` when appropriate.
|
||||||
|
- [x] Ensure `severity=low_confid` prevents effective `PASS` and maps to `LOW_CONFID`.
|
||||||
|
- [x] Ensure `gatekeeper_result` with `severity` is persisted under `verifier_evaluation`.
|
||||||
|
- [x] Add ChatService or integration tests for effective verdict guardrails.
|
||||||
|
|
||||||
|
Acceptance:
|
||||||
|
|
||||||
|
- Gatekeeper fail cannot become final PASS.
|
||||||
|
- Audit JSON contains the minimum Gatekeeper fields.
|
||||||
|
|
||||||
|
## 5. HikariCP mock quality
|
||||||
|
|
||||||
|
- [x] Add positive HikariCP mock logs for `order-service`.
|
||||||
|
- [x] Support HikariCP synonym matching.
|
||||||
|
- [x] Remove `generic-service` placeholder evidence for no-hit cases.
|
||||||
|
- [x] Add tests for HikariCP positive and negative no-hit behavior.
|
||||||
|
|
||||||
|
Acceptance:
|
||||||
|
|
||||||
|
- Positive HikariCP query returns order-service logs.
|
||||||
|
- Negative HikariCP query returns `logs=[]` and `evidence_status=no_evidence`.
|
||||||
|
|
||||||
|
## 6. Verification
|
||||||
|
|
||||||
|
- [x] Run focused unit tests for recorder, Gatekeeper, hook, ChatService guardrails, and query log mock behavior.
|
||||||
|
- [x] Run `openspec validate verifier-evidence-reference-fidelity --strict`.
|
||||||
|
- [x] Run `openspec validate --specs`.
|
||||||
|
- [x] Start the Java project using `mvn spring-boot:run`.
|
||||||
|
- [x] Run end-to-end checks for HighMemoryUsage positive, SlowResponse positive, HikariCP positive, HikariCP negative, and narrow forbidden claim.
|
||||||
|
- [x] Query MySQL audit data with `scripts/query_mysql.py` to confirm persisted `gatekeeper_result`.
|
||||||
|
|
||||||
|
Acceptance:
|
||||||
|
|
||||||
|
- Minimum E2E matrix passes or any failure is classified as code issue, mock quality issue, retrieval/tool quality issue, or model nondeterminism with evidence.
|
||||||
@@ -1,4 +1,4 @@
|
|||||||
# chat-verifier-agent Specification
|
# chat-verifier-agent Specification
|
||||||
|
|
||||||
## Purpose
|
## Purpose
|
||||||
TBD - created by archiving change chat-verifier-agent. Update Purpose after archive.
|
TBD - created by archiving change chat-verifier-agent. Update Purpose after archive.
|
||||||
@@ -119,6 +119,7 @@ The system SHALL use ChatService for explicit single-round `Planner -> Executor
|
|||||||
- **THEN** the system SHALL output a degraded result indicating the answer cannot be reliably generated
|
- **THEN** the system SHALL output a degraded result indicating the answer cannot be reliably generated
|
||||||
- **AND** it SHALL NOT pass through the raw Executor answer
|
- **AND** it SHALL NOT pass through the raw Executor answer
|
||||||
- **AND** it SHALL NOT include a root-cause conclusion
|
- **AND** it SHALL NOT include a root-cause conclusion
|
||||||
|
|
||||||
### Requirement: Verifier SHALL be observable
|
### Requirement: Verifier SHALL be observable
|
||||||
The Verifier's verdict and downstream final-answer composition SHALL be persisted for observability.
|
The Verifier's verdict and downstream final-answer composition SHALL be persisted for observability.
|
||||||
|
|
||||||
@@ -421,3 +422,112 @@ The system SHALL use Composer or fixed safe templates for PASS, LOW_CONFID, and
|
|||||||
- **AND** it SHALL include only confirmed facts, evidence gaps, and next-step suggestions
|
- **AND** it SHALL include only confirmed facts, evidence gaps, and next-step suggestions
|
||||||
- **AND** it SHALL NOT include unverified raw answer content
|
- **AND** it SHALL NOT include unverified raw answer content
|
||||||
- **AND** it SHALL NOT include a root-cause conclusion
|
- **AND** it SHALL NOT include a root-cause conclusion
|
||||||
|
|
||||||
|
### Requirement: Executor evidence bindings SHALL support precise evidence references
|
||||||
|
Executor V2 evidence bindings SHALL support precise evidence references that locate evidence inside a persisted tool invocation.
|
||||||
|
|
||||||
|
#### Scenario: Precise evidence binding contains invocation path and excerpt
|
||||||
|
- **WHEN** Executor binds evidence to a claim
|
||||||
|
- **THEN** the binding SHOULD include singular `source_invocation_id`
|
||||||
|
- **AND** the binding SHOULD include `raw_path`
|
||||||
|
- **AND** the binding SHALL include `tool_name` and `evidence_excerpt`
|
||||||
|
- **AND** the `raw_path` SHALL be interpreted relative to the referenced tool invocation's `retrieval_details.evidence_refs`
|
||||||
|
|
||||||
|
#### Scenario: Legacy plural invocation ids remain compatibility only
|
||||||
|
- **WHEN** Executor emits legacy `source_invocation_ids`
|
||||||
|
- **THEN** the system MAY read them for compatibility
|
||||||
|
- **AND** they SHALL NOT be sufficient for a precise Gatekeeper pass without `raw_path`
|
||||||
|
|
||||||
|
### Requirement: Gatekeeper SHALL validate evidence reference fidelity
|
||||||
|
Gatekeeper SHALL validate that Executor evidence bindings point to real current-session evidence references before Verifier uses them as primary evidence.
|
||||||
|
|
||||||
|
#### Scenario: Valid precise binding passes
|
||||||
|
- **WHEN** a binding's `source_invocation_id` exists in the current session
|
||||||
|
- **AND** the binding's `tool_name` matches the persisted invocation
|
||||||
|
- **AND** the binding's `raw_path` exists in `retrieval_details.evidence_refs`
|
||||||
|
- **AND** the binding's `evidence_excerpt` is supported by the matching evidence ref text
|
||||||
|
- **THEN** Gatekeeper SHALL return `status=pass`
|
||||||
|
- **AND** Gatekeeper SHALL return `severity=none`
|
||||||
|
|
||||||
|
#### Scenario: Missing raw path is low confidence
|
||||||
|
- **WHEN** a binding references an existing invocation
|
||||||
|
- **AND** the binding omits `raw_path`
|
||||||
|
- **THEN** Gatekeeper SHALL return `status=fail`
|
||||||
|
- **AND** Gatekeeper SHALL return `severity=low_confid`
|
||||||
|
- **AND** the effective verifier result SHALL NOT be `PASS`
|
||||||
|
|
||||||
|
#### Scenario: Old invocation without evidence refs is low confidence
|
||||||
|
- **WHEN** a binding references an existing invocation
|
||||||
|
- **AND** the invocation does not contain `retrieval_details.evidence_refs`
|
||||||
|
- **THEN** Gatekeeper SHALL return `status=fail`
|
||||||
|
- **AND** Gatekeeper SHALL return `severity=low_confid`
|
||||||
|
- **AND** the effective verifier result SHALL NOT be `PASS`
|
||||||
|
|
||||||
|
#### Scenario: Unknown raw path is rejected
|
||||||
|
- **WHEN** a binding references an existing invocation
|
||||||
|
- **AND** the binding's `raw_path` is absent from that invocation's `retrieval_details.evidence_refs`
|
||||||
|
- **THEN** Gatekeeper SHALL return `status=fail`
|
||||||
|
- **AND** Gatekeeper SHALL return `severity=reject`
|
||||||
|
- **AND** `failed_rules` SHALL include `evidence.raw_path`
|
||||||
|
|
||||||
|
#### Scenario: Mismatched excerpt is rejected
|
||||||
|
- **WHEN** a binding references an existing invocation and raw path
|
||||||
|
- **AND** the binding's `evidence_excerpt` is not supported by the matching system-side evidence ref text
|
||||||
|
- **THEN** Gatekeeper SHALL return `status=fail`
|
||||||
|
- **AND** Gatekeeper SHALL return `severity=reject`
|
||||||
|
- **AND** `failed_rules` SHALL include `evidence.excerpt_mismatch`
|
||||||
|
|
||||||
|
### Requirement: Verifier SHALL use verified claim-local evidence for derivability
|
||||||
|
Verifier SHALL judge structured claims primarily against Gatekeeper-verified claim-local evidence excerpts.
|
||||||
|
|
||||||
|
#### Scenario: Verified excerpt supports direct observation
|
||||||
|
- **WHEN** `gatekeeper_result.severity=none`
|
||||||
|
- **AND** a claim's verified evidence excerpts directly contain the claim's concrete facts
|
||||||
|
- **THEN** Verifier MAY classify that claim as `direct_observation`
|
||||||
|
|
||||||
|
#### Scenario: Tool trace summary is navigation context
|
||||||
|
- **WHEN** `executor_structured_output.claims[].evidence_bindings` are available
|
||||||
|
- **THEN** Verifier SHALL use `tool_trace_summary` as navigation and audit context
|
||||||
|
- **AND** it SHALL NOT require `tool_trace_summary.output_summary` to contain every fact already present in verified claim-local evidence
|
||||||
|
|
||||||
|
### Requirement: Gatekeeper severity SHALL constrain effective verdict
|
||||||
|
Runtime effective verdict calculation SHALL treat Gatekeeper severity as a hard upper bound.
|
||||||
|
|
||||||
|
#### Scenario: Reject severity prevents PASS
|
||||||
|
- **WHEN** `gatekeeper_result.severity=reject`
|
||||||
|
- **AND** the Verifier model returns `verdict=PASS`
|
||||||
|
- **THEN** ChatService SHALL downgrade the effective verdict
|
||||||
|
- **AND** the effective verdict SHALL be `REJECT`
|
||||||
|
|
||||||
|
#### Scenario: Low confidence severity prevents PASS
|
||||||
|
- **WHEN** `gatekeeper_result.severity=low_confid`
|
||||||
|
- **AND** the Verifier model returns `verdict=PASS`
|
||||||
|
- **THEN** ChatService SHALL downgrade the effective verdict
|
||||||
|
- **AND** the effective verdict SHALL be `LOW_CONFID`
|
||||||
|
|
||||||
|
#### Scenario: Gatekeeper audit includes severity
|
||||||
|
- **WHEN** verifier evaluation is persisted
|
||||||
|
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result` SHALL include `status`, `severity`, `checked_bindings`, `failed_rules`, `warnings`, and `errors`
|
||||||
|
|
||||||
|
### Requirement: Verifier input hook SHALL only perform narrow compatibility backfill
|
||||||
|
The verifier input hook SHALL avoid converting broad tool summaries into precise evidence references.
|
||||||
|
|
||||||
|
#### Scenario: Unique invocation candidate may be backfilled
|
||||||
|
- **WHEN** an evidence binding omits `source_invocation_id`
|
||||||
|
- **AND** exactly one current-session invocation exists for the binding's `tool_name`
|
||||||
|
- **THEN** the hook MAY backfill `source_invocation_id`
|
||||||
|
- **AND** it SHALL add a Gatekeeper warning describing the auto-backfill
|
||||||
|
|
||||||
|
#### Scenario: Raw path is never backfilled
|
||||||
|
- **WHEN** an evidence binding omits `raw_path`
|
||||||
|
- **THEN** the hook SHALL NOT synthesize `raw_path`
|
||||||
|
- **AND** Gatekeeper SHALL treat the binding as not precise enough to pass
|
||||||
|
|
||||||
|
### Requirement: Executor prompt SHALL constrain narrow-scope over-expansion
|
||||||
|
The Executor prompt SHALL instruct Executor to keep narrow confirmation questions focused on observation-level claims.
|
||||||
|
|
||||||
|
#### Scenario: Narrow scope produces minimal observation claims
|
||||||
|
- **WHEN** the user asks to confirm one specific service, alert, or symptom
|
||||||
|
- **THEN** Executor SHOULD output the minimum necessary claims, normally one and at most two
|
||||||
|
- **AND** those claims SHALL be `observation` or `negative_observation` unless current-session evidence proves more
|
||||||
|
- **AND** Executor SHALL NOT emit unrelated root-cause, remediation, or excluded-topic claims as confirmed facts
|
||||||
|
|||||||
@@ -133,3 +133,40 @@ The system SHALL expose baseline diff output in structured JSON and reviewable M
|
|||||||
#### Scenario: Markdown diff output
|
#### Scenario: Markdown diff output
|
||||||
- **WHEN** a baseline diff is written as Markdown
|
- **WHEN** a baseline diff is written as Markdown
|
||||||
- **THEN** it SHALL include a readable summary and a table of diff items
|
- **THEN** it SHALL include a readable summary and a table of diff items
|
||||||
|
|
||||||
|
### Requirement: Evaluation harness SHALL validate Executor V2 audit closure
|
||||||
|
The evaluation harness SHALL be able to validate the V2 audit chain from Executor structured output through Gatekeeper, Verifier claim checks, Composer output, and final user answer.
|
||||||
|
|
||||||
|
#### Scenario: Required V2 audit fields are present
|
||||||
|
- **WHEN** an evaluation case requires V2 audit closure
|
||||||
|
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.gatekeeper_result` exists
|
||||||
|
- **AND** it SHALL verify that `selfEvaluation.verifier_evaluation.claim_checks` exists
|
||||||
|
- **AND** it SHALL verify that `selfEvaluation.verifier_evaluation.composer_output` exists
|
||||||
|
|
||||||
|
#### Scenario: Gatekeeper failure cannot pass verification
|
||||||
|
- **WHEN** a trace has `gatekeeper_result.status` equal to `fail`
|
||||||
|
- **THEN** the evaluator SHALL fail the case if `selfEvaluation.verifier_evaluation.verdict` is `PASS`
|
||||||
|
|
||||||
|
#### Scenario: Claim checks are auditable
|
||||||
|
- **WHEN** an evaluation case requires claim checks
|
||||||
|
- **THEN** the evaluator SHALL verify that each claim check includes `claim_id`, `verification`, and `detail`
|
||||||
|
- **AND** each verification SHALL be one of the V2 claim verification values recognized by the Verifier contract
|
||||||
|
|
||||||
|
#### Scenario: Composer output is auditable
|
||||||
|
- **WHEN** an evaluation case requires Composer output
|
||||||
|
- **THEN** the evaluator SHALL verify that `composer_output` records whether parsed output or fallback rendering was used
|
||||||
|
|
||||||
|
### Requirement: Evaluation harness SHALL prevent unsupported claims from leaking into final answers
|
||||||
|
The evaluation harness SHALL detect configured unsafe or unsupported claim text when it appears in the final user-facing answer.
|
||||||
|
|
||||||
|
#### Scenario: Unsupported final-answer claim is rejected
|
||||||
|
- **WHEN** an evaluation case declares forbidden confirmed-claim keywords
|
||||||
|
- **THEN** the evaluator SHALL fail the case if the final answer contains any of those keywords
|
||||||
|
|
||||||
|
#### Scenario: Raw Executor JSON is not user-facing
|
||||||
|
- **WHEN** a trace is evaluated under V2 audit closure
|
||||||
|
- **THEN** the evaluator SHALL fail the case if the final answer contains raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`
|
||||||
|
|
||||||
|
#### Scenario: Composer fallback still avoids raw JSON leakage
|
||||||
|
- **WHEN** a trace records Composer fallback rendering
|
||||||
|
- **THEN** the evaluator SHALL still enforce final-answer raw JSON leakage checks
|
||||||
|
|||||||
@@ -108,3 +108,44 @@ The persisted trace SHALL make it possible to audit model step counts separately
|
|||||||
- **AND** `diagnosis_session.tool_call_count` SHALL count persisted evidence-tool invocation rows
|
- **AND** `diagnosis_session.tool_call_count` SHALL count persisted evidence-tool invocation rows
|
||||||
- **AND** helper workflow calls that are not evidence rows SHALL be auditable from agent steps or logs without inflating `tool_invocation`
|
- **AND** helper workflow calls that are not evidence rows SHALL be auditable from agent steps or logs without inflating `tool_invocation`
|
||||||
|
|
||||||
|
### Requirement: Evidence tools SHALL persist minimal evidence refs
|
||||||
|
Evidence-bearing tool invocations SHALL persist claim-addressable evidence references in `tool_invocation.retrieval_details.evidence_refs`.
|
||||||
|
|
||||||
|
#### Scenario: Metrics alerts produce evidence refs
|
||||||
|
- **WHEN** a `query_metrics` invocation returns alert entries
|
||||||
|
- **THEN** the persisted retrieval details SHALL include one `evidence_refs` item per usable alert
|
||||||
|
- **AND** each item SHALL include `raw_path` formatted as `$.alerts[i]`
|
||||||
|
- **AND** each item SHALL include bounded `text` containing concrete alert facts such as alert name, state, service, current value, and duration when available
|
||||||
|
|
||||||
|
#### Scenario: Logs produce evidence refs
|
||||||
|
- **WHEN** a `query_logs` invocation returns log entries
|
||||||
|
- **THEN** the persisted retrieval details SHALL include one `evidence_refs` item per usable log
|
||||||
|
- **AND** each item SHALL include `raw_path` formatted as `$.logs[i]`
|
||||||
|
- **AND** each item SHALL include bounded `text` containing concrete log facts such as timestamp, level, service, and message when available
|
||||||
|
|
||||||
|
#### Scenario: Knowledge lookup produces evidence refs
|
||||||
|
- **WHEN** a `lookup_knowledge` invocation returns evidence blocks
|
||||||
|
- **THEN** the persisted retrieval details SHALL include one `evidence_refs` item per usable evidence block
|
||||||
|
- **AND** each item SHALL include `raw_path` formatted as `$.evidence_blocks[i]`
|
||||||
|
- **AND** each item SHALL include bounded `text` containing concrete block content, title, or source when available
|
||||||
|
|
||||||
|
#### Scenario: Evidence ref extraction does not infer diagnosis
|
||||||
|
- **WHEN** the recorder creates `evidence_refs`
|
||||||
|
- **THEN** it SHALL only copy or format concrete tool output fields
|
||||||
|
- **AND** it SHALL NOT infer root cause, remediation, or diagnosis conclusions
|
||||||
|
|
||||||
|
### Requirement: Log mock no-hit semantics SHALL avoid placeholder evidence
|
||||||
|
The log query mock SHALL distinguish positive mock evidence from no-hit results without using placeholder service logs as evidence.
|
||||||
|
|
||||||
|
#### Scenario: HikariCP positive query returns order-service pool evidence
|
||||||
|
- **WHEN** a `query_logs` request targets `order-service` and HikariCP connection-pool exhaustion terms
|
||||||
|
- **THEN** the tool SHALL return order-service HikariCP-related log entries
|
||||||
|
- **AND** the returned evidence SHALL include concrete terms such as `HikariPool`, `active=50/50`, `waiting`, or `request timed out after 30000ms`
|
||||||
|
- **AND** it SHALL NOT return `generic-service` placeholder logs
|
||||||
|
|
||||||
|
#### Scenario: HikariCP no-hit query returns no evidence
|
||||||
|
- **WHEN** a `query_logs` request targets a service without matching HikariCP mock evidence
|
||||||
|
- **THEN** the tool SHALL return an empty `logs` array
|
||||||
|
- **AND** it SHALL mark the output as `evidence_status=no_evidence`
|
||||||
|
- **AND** it SHALL NOT return `generic-service` placeholder logs
|
||||||
|
|
||||||
|
|||||||
@@ -289,11 +289,7 @@ public class QueryLogsTools {
|
|||||||
logs.addAll(buildSystemEventsLogs(now, normalizedQuery, limit));
|
logs.addAll(buildSystemEventsLogs(now, normalizedQuery, limit));
|
||||||
break;
|
break;
|
||||||
default:
|
default:
|
||||||
logs.addAll(buildGenericLogs(now, normalizedQuery, limit));
|
return logs;
|
||||||
}
|
|
||||||
|
|
||||||
if (logs.isEmpty()) {
|
|
||||||
logs.addAll(buildGenericLogs(now, normalizedQuery, limit));
|
|
||||||
}
|
}
|
||||||
|
|
||||||
// 限制返回条数
|
// 限制返回条数
|
||||||
@@ -401,6 +397,10 @@ public class QueryLogsTools {
|
|||||||
private List<LogEntry> buildApplicationLogs(Instant now, String query, int limit) {
|
private List<LogEntry> buildApplicationLogs(Instant now, String query, int limit) {
|
||||||
List<LogEntry> logs = new ArrayList<>();
|
List<LogEntry> logs = new ArrayList<>();
|
||||||
|
|
||||||
|
if (isHikariPoolQuery(query) && targetsOrderService(query)) {
|
||||||
|
logs.addAll(buildHikariPoolLogs(now));
|
||||||
|
}
|
||||||
|
|
||||||
// ERROR 级别日志
|
// ERROR 级别日志
|
||||||
if (query.contains("error") || query.contains("fatal") || query.contains("500")) {
|
if (query.contains("error") || query.contains("fatal") || query.contains("500")) {
|
||||||
|
|
||||||
@@ -515,6 +515,58 @@ public class QueryLogsTools {
|
|||||||
return logs;
|
return logs;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
private boolean isHikariPoolQuery(String query) {
|
||||||
|
return query.contains("hikaricp")
|
||||||
|
|| query.contains("hikaripool")
|
||||||
|
|| query.contains("connection pool")
|
||||||
|
|| query.contains("数据库连接池")
|
||||||
|
|| query.contains("连接池耗尽")
|
||||||
|
|| query.contains("active=50/50")
|
||||||
|
|| query.contains("waiting")
|
||||||
|
|| query.contains("request timed out after 30000ms");
|
||||||
|
}
|
||||||
|
|
||||||
|
private boolean targetsOrderService(String query) {
|
||||||
|
if (query.contains("inventory-service") || query.contains("payment-service") || query.contains("user-service")) {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
return query.contains("order-service") || isHikariPoolQuery(query);
|
||||||
|
}
|
||||||
|
|
||||||
|
private List<LogEntry> buildHikariPoolLogs(Instant now) {
|
||||||
|
List<LogEntry> logs = new ArrayList<>();
|
||||||
|
|
||||||
|
LogEntry timeout = new LogEntry();
|
||||||
|
timeout.setTimestamp(FORMATTER.format(now.minus(2, ChronoUnit.MINUTES)));
|
||||||
|
timeout.setLevel("ERROR");
|
||||||
|
timeout.setService("order-service");
|
||||||
|
timeout.setInstance("pod-order-service-5c7d8e9f1-m3n2p");
|
||||||
|
timeout.setMessage("HikariPool-1 - Connection is not available, request timed out after 30000ms");
|
||||||
|
timeout.setMetrics(Map.of(
|
||||||
|
"pool", "HikariPool-1",
|
||||||
|
"error_type", "ConnectionPoolExhaustedException",
|
||||||
|
"timeout_ms", "30000"
|
||||||
|
));
|
||||||
|
logs.add(timeout);
|
||||||
|
|
||||||
|
LogEntry stats = new LogEntry();
|
||||||
|
stats.setTimestamp(FORMATTER.format(now.minus(1, ChronoUnit.MINUTES)));
|
||||||
|
stats.setLevel("WARN");
|
||||||
|
stats.setService("order-service");
|
||||||
|
stats.setInstance("pod-order-service-5c7d8e9f1-m3n2p");
|
||||||
|
stats.setMessage("HikariCP pool stats: active=50/50, idle=0, waiting=32");
|
||||||
|
stats.setMetrics(Map.of(
|
||||||
|
"pool", "HikariPool-1",
|
||||||
|
"active", "50",
|
||||||
|
"max", "50",
|
||||||
|
"idle", "0",
|
||||||
|
"waiting", "32"
|
||||||
|
));
|
||||||
|
logs.add(stats);
|
||||||
|
|
||||||
|
return logs;
|
||||||
|
}
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* 构建数据库慢查询日志(与慢响应告警关联)
|
* 构建数据库慢查询日志(与慢响应告警关联)
|
||||||
*/
|
*/
|
||||||
|
|||||||
@@ -22,4 +22,10 @@ public class DiagnosisEvalCase {
|
|||||||
private List<String> requiredEvidenceTools;
|
private List<String> requiredEvidenceTools;
|
||||||
private List<String> allowedVerdicts;
|
private List<String> allowedVerdicts;
|
||||||
private List<String> forbiddenAnswerKeywords;
|
private List<String> forbiddenAnswerKeywords;
|
||||||
|
private Boolean requireV2AuditClosure;
|
||||||
|
private Boolean requireClaimChecks;
|
||||||
|
private Boolean requireComposerOutput;
|
||||||
|
private List<String> expectedGatekeeperStatuses;
|
||||||
|
private List<String> expectedComposerStatuses;
|
||||||
|
private List<String> forbiddenConfirmedClaimKeywords;
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -46,8 +46,8 @@ public class DiagnosisEvalReportWriter {
|
|||||||
}
|
}
|
||||||
|
|
||||||
builder.append("## Cases\n\n");
|
builder.append("## Cases\n\n");
|
||||||
builder.append("| Case | Result | Verdict | Keywords | Tool Calls | Duration ms | Failed Checks |\n");
|
builder.append("| Case | Result | Verdict | Gatekeeper | Composer | Claim Checks | Keywords | Tool Calls | Duration ms | Failed Checks |\n");
|
||||||
builder.append("| --- | --- | --- | --- | ---: | ---: | --- |\n");
|
builder.append("| --- | --- | --- | --- | --- | ---: | --- | ---: | ---: | --- |\n");
|
||||||
for (DiagnosisEvalResult result : report.getResults()) {
|
for (DiagnosisEvalResult result : report.getResults()) {
|
||||||
builder.append("| ")
|
builder.append("| ")
|
||||||
.append(result.getCaseId())
|
.append(result.getCaseId())
|
||||||
@@ -56,6 +56,12 @@ public class DiagnosisEvalReportWriter {
|
|||||||
.append(" | ")
|
.append(" | ")
|
||||||
.append(valueOrDash(result.getVerdict()))
|
.append(valueOrDash(result.getVerdict()))
|
||||||
.append(" | ")
|
.append(" | ")
|
||||||
|
.append(valueOrDash(result.getGatekeeperStatus()))
|
||||||
|
.append(" | ")
|
||||||
|
.append(valueOrDash(result.getComposerStatus()))
|
||||||
|
.append(" | ")
|
||||||
|
.append(result.getClaimCheckCount() == null ? "-" : result.getClaimCheckCount())
|
||||||
|
.append(" | ")
|
||||||
.append(result.getMatchedKeywordCount()).append("/").append(result.getRequiredKeywordCount())
|
.append(result.getMatchedKeywordCount()).append("/").append(result.getRequiredKeywordCount())
|
||||||
.append(" | ")
|
.append(" | ")
|
||||||
.append(result.getToolCallCount() == null ? "-" : result.getToolCallCount())
|
.append(result.getToolCallCount() == null ? "-" : result.getToolCallCount())
|
||||||
|
|||||||
@@ -22,6 +22,9 @@ public class DiagnosisEvalResult {
|
|||||||
private int matchedKeywordCount;
|
private int matchedKeywordCount;
|
||||||
private int requiredKeywordCount;
|
private int requiredKeywordCount;
|
||||||
private Map<String, Boolean> evidenceCoverage;
|
private Map<String, Boolean> evidenceCoverage;
|
||||||
|
private String gatekeeperStatus;
|
||||||
|
private String composerStatus;
|
||||||
|
private Integer claimCheckCount;
|
||||||
private Integer toolCallCount;
|
private Integer toolCallCount;
|
||||||
private Integer durationMs;
|
private Integer durationMs;
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -20,6 +20,20 @@ public class DiagnosisTraceEvaluator {
|
|||||||
|
|
||||||
private static final TypeReference<List<DiagnosisEvalCase>> CASE_LIST_TYPE = new TypeReference<>() {};
|
private static final TypeReference<List<DiagnosisEvalCase>> CASE_LIST_TYPE = new TypeReference<>() {};
|
||||||
private static final String REJECT_DEGRADED_PREFIX = "当前无法基于已获取证据生成可靠结论";
|
private static final String REJECT_DEGRADED_PREFIX = "当前无法基于已获取证据生成可靠结论";
|
||||||
|
private static final Set<String> RAW_EXECUTOR_MARKERS = Set.of(
|
||||||
|
"executor_evidence_v2",
|
||||||
|
"answer_version",
|
||||||
|
"evidence_bindings",
|
||||||
|
"claim_id"
|
||||||
|
);
|
||||||
|
private static final Set<String> VALID_CLAIM_VERIFICATIONS = Set.of(
|
||||||
|
"direct_observation",
|
||||||
|
"reasonable_inference",
|
||||||
|
"overstated",
|
||||||
|
"unsupported",
|
||||||
|
"external_unknown",
|
||||||
|
"contradicted"
|
||||||
|
);
|
||||||
|
|
||||||
private final ObjectMapper objectMapper;
|
private final ObjectMapper objectMapper;
|
||||||
|
|
||||||
@@ -51,6 +65,9 @@ public class DiagnosisTraceEvaluator {
|
|||||||
.matchedKeywordCount(0)
|
.matchedKeywordCount(0)
|
||||||
.requiredKeywordCount(size(evalCase.getExpectedRootCauseKeywords()))
|
.requiredKeywordCount(size(evalCase.getExpectedRootCauseKeywords()))
|
||||||
.evidenceCoverage(emptyCoverage(evalCase.getRequiredEvidenceTools()))
|
.evidenceCoverage(emptyCoverage(evalCase.getRequiredEvidenceTools()))
|
||||||
|
.gatekeeperStatus(null)
|
||||||
|
.composerStatus(null)
|
||||||
|
.claimCheckCount(null)
|
||||||
.toolCallCount(null)
|
.toolCallCount(null)
|
||||||
.durationMs(null)
|
.durationMs(null)
|
||||||
.build());
|
.build());
|
||||||
@@ -102,6 +119,11 @@ public class DiagnosisTraceEvaluator {
|
|||||||
}
|
}
|
||||||
|
|
||||||
failedChecks.addAll(validateExecutorStructuredOutput(trace));
|
failedChecks.addAll(validateExecutorStructuredOutput(trace));
|
||||||
|
String gatekeeperStatus = extractNestedString(trace, "verifier_evaluation", "gatekeeper_result", "status");
|
||||||
|
String composerStatus = extractNestedString(trace, "verifier_evaluation", "composer_output", "status");
|
||||||
|
Integer claimCheckCount = countList(trace, "verifier_evaluation", "claim_checks");
|
||||||
|
failedChecks.addAll(validateV2AuditClosure(evalCase, trace, normalizedAnswer, verdict,
|
||||||
|
gatekeeperStatus, composerStatus));
|
||||||
|
|
||||||
Integer toolCallCount = trace.getToolInvocations() == null ? 0 : trace.getToolInvocations().size();
|
Integer toolCallCount = trace.getToolInvocations() == null ? 0 : trace.getToolInvocations().size();
|
||||||
Integer durationMs = trace.getSession() == null ? null : trace.getSession().getTotalDurationMs();
|
Integer durationMs = trace.getSession() == null ? null : trace.getSession().getTotalDurationMs();
|
||||||
@@ -115,6 +137,9 @@ public class DiagnosisTraceEvaluator {
|
|||||||
.matchedKeywordCount(matchedKeywordCount)
|
.matchedKeywordCount(matchedKeywordCount)
|
||||||
.requiredKeywordCount(requiredKeywordCount)
|
.requiredKeywordCount(requiredKeywordCount)
|
||||||
.evidenceCoverage(evidenceCoverage)
|
.evidenceCoverage(evidenceCoverage)
|
||||||
|
.gatekeeperStatus(gatekeeperStatus)
|
||||||
|
.composerStatus(composerStatus)
|
||||||
|
.claimCheckCount(claimCheckCount)
|
||||||
.toolCallCount(toolCallCount)
|
.toolCallCount(toolCallCount)
|
||||||
.durationMs(durationMs)
|
.durationMs(durationMs)
|
||||||
.build();
|
.build();
|
||||||
@@ -202,6 +227,91 @@ public class DiagnosisTraceEvaluator {
|
|||||||
return failedChecks;
|
return failedChecks;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
private List<String> validateV2AuditClosure(DiagnosisEvalCase evalCase,
|
||||||
|
DiagnosisTraceResponse trace,
|
||||||
|
String normalizedAnswer,
|
||||||
|
String verdict,
|
||||||
|
String gatekeeperStatus,
|
||||||
|
String composerStatus) {
|
||||||
|
List<String> failedChecks = new ArrayList<>();
|
||||||
|
boolean requireV2AuditClosure = Boolean.TRUE.equals(evalCase.getRequireV2AuditClosure());
|
||||||
|
boolean requireClaimChecks = requireV2AuditClosure || Boolean.TRUE.equals(evalCase.getRequireClaimChecks());
|
||||||
|
boolean requireComposerOutput = requireV2AuditClosure || Boolean.TRUE.equals(evalCase.getRequireComposerOutput());
|
||||||
|
|
||||||
|
Object gatekeeperResult = nestedValue(trace, "verifier_evaluation", "gatekeeper_result");
|
||||||
|
if (requireV2AuditClosure && !(gatekeeperResult instanceof Map<?, ?>)) {
|
||||||
|
failedChecks.add("missing gatekeeper_result");
|
||||||
|
}
|
||||||
|
if ("fail".equals(gatekeeperStatus) && "PASS".equals(verdict)) {
|
||||||
|
failedChecks.add("gatekeeper fail cannot have PASS verdict");
|
||||||
|
}
|
||||||
|
if (!safeList(evalCase.getExpectedGatekeeperStatuses()).isEmpty()
|
||||||
|
&& !safeList(evalCase.getExpectedGatekeeperStatuses()).contains(gatekeeperStatus)) {
|
||||||
|
failedChecks.add("gatekeeper status not expected: " + valueOrMissing(gatekeeperStatus));
|
||||||
|
}
|
||||||
|
|
||||||
|
failedChecks.addAll(validateClaimChecks(trace, requireClaimChecks));
|
||||||
|
|
||||||
|
Object composerOutput = nestedValue(trace, "verifier_evaluation", "composer_output");
|
||||||
|
if (requireComposerOutput && !(composerOutput instanceof Map<?, ?>)) {
|
||||||
|
failedChecks.add("missing composer_output");
|
||||||
|
}
|
||||||
|
if (requireComposerOutput && isBlank(composerStatus)) {
|
||||||
|
failedChecks.add("composer_output missing status");
|
||||||
|
}
|
||||||
|
if (!safeList(evalCase.getExpectedComposerStatuses()).isEmpty()
|
||||||
|
&& !safeList(evalCase.getExpectedComposerStatuses()).contains(composerStatus)) {
|
||||||
|
failedChecks.add("composer status not expected: " + valueOrMissing(composerStatus));
|
||||||
|
}
|
||||||
|
|
||||||
|
for (String forbidden : safeList(evalCase.getForbiddenConfirmedClaimKeywords())) {
|
||||||
|
if (normalizedAnswer.contains(forbidden.toLowerCase(Locale.ROOT))) {
|
||||||
|
failedChecks.add("answer contains forbidden confirmed claim keyword: " + forbidden);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
if (requireV2AuditClosure) {
|
||||||
|
for (String marker : RAW_EXECUTOR_MARKERS) {
|
||||||
|
if (normalizedAnswer.contains(marker.toLowerCase(Locale.ROOT))) {
|
||||||
|
failedChecks.add("answer leaks raw executor marker: " + marker);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return failedChecks;
|
||||||
|
}
|
||||||
|
|
||||||
|
private List<String> validateClaimChecks(DiagnosisTraceResponse trace, boolean required) {
|
||||||
|
Object claimChecks = nestedValue(trace, "verifier_evaluation", "claim_checks");
|
||||||
|
if (!(claimChecks instanceof List<?> claimCheckList)) {
|
||||||
|
return required ? List.of("missing claim_checks") : List.of();
|
||||||
|
}
|
||||||
|
if (required && claimCheckList.isEmpty()) {
|
||||||
|
return List.of("claim_checks is empty");
|
||||||
|
}
|
||||||
|
|
||||||
|
List<String> failedChecks = new ArrayList<>();
|
||||||
|
for (Object item : claimCheckList) {
|
||||||
|
if (!(item instanceof Map<?, ?> claimCheck)) {
|
||||||
|
failedChecks.add("claim_check is not an object");
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
String claimId = stringValue(claimCheck.get("claim_id"));
|
||||||
|
String verification = stringValue(claimCheck.get("verification"));
|
||||||
|
if (isBlank(claimId)) {
|
||||||
|
failedChecks.add("claim_check missing claim_id");
|
||||||
|
}
|
||||||
|
if (isBlank(verification)) {
|
||||||
|
failedChecks.add("claim_check missing verification: " + valueOrMissing(claimId));
|
||||||
|
} else if (!VALID_CLAIM_VERIFICATIONS.contains(verification)) {
|
||||||
|
failedChecks.add("claim_check verification invalid: " + verification);
|
||||||
|
}
|
||||||
|
if (isBlank(stringValue(claimCheck.get("detail")))) {
|
||||||
|
failedChecks.add("claim_check missing detail: " + valueOrMissing(claimId));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return failedChecks;
|
||||||
|
}
|
||||||
|
|
||||||
private Object nestedValue(DiagnosisTraceResponse trace, String firstKey, String secondKey) {
|
private Object nestedValue(DiagnosisTraceResponse trace, String firstKey, String secondKey) {
|
||||||
if (trace.getSession() == null || trace.getSession().getSelfEvaluation() == null) {
|
if (trace.getSession() == null || trace.getSession().getSelfEvaluation() == null) {
|
||||||
return null;
|
return null;
|
||||||
@@ -213,6 +323,20 @@ public class DiagnosisTraceEvaluator {
|
|||||||
return map.get(secondKey);
|
return map.get(secondKey);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
private String extractNestedString(DiagnosisTraceResponse trace, String firstKey, String secondKey, String thirdKey) {
|
||||||
|
Object value = nestedValue(trace, firstKey, secondKey);
|
||||||
|
if (!(value instanceof Map<?, ?> map)) {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
Object nested = map.get(thirdKey);
|
||||||
|
return nested == null ? null : String.valueOf(nested);
|
||||||
|
}
|
||||||
|
|
||||||
|
private Integer countList(DiagnosisTraceResponse trace, String firstKey, String secondKey) {
|
||||||
|
Object value = nestedValue(trace, firstKey, secondKey);
|
||||||
|
return value instanceof List<?> list ? list.size() : null;
|
||||||
|
}
|
||||||
|
|
||||||
private int countMatches(String normalizedAnswer, List<String> keywords) {
|
private int countMatches(String normalizedAnswer, List<String> keywords) {
|
||||||
int count = 0;
|
int count = 0;
|
||||||
for (String keyword : safeList(keywords)) {
|
for (String keyword : safeList(keywords)) {
|
||||||
@@ -242,4 +366,16 @@ public class DiagnosisTraceEvaluator {
|
|||||||
private String nullToEmpty(String value) {
|
private String nullToEmpty(String value) {
|
||||||
return value == null ? "" : value;
|
return value == null ? "" : value;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
private String stringValue(Object value) {
|
||||||
|
return value == null ? null : String.valueOf(value);
|
||||||
|
}
|
||||||
|
|
||||||
|
private boolean isBlank(String value) {
|
||||||
|
return value == null || value.isBlank();
|
||||||
|
}
|
||||||
|
|
||||||
|
private String valueOrMissing(String value) {
|
||||||
|
return isBlank(value) ? "missing" : value;
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -17,9 +17,12 @@ import org.springframework.ai.chat.messages.AssistantMessage;
|
|||||||
import org.springframework.ai.chat.messages.Message;
|
import org.springframework.ai.chat.messages.Message;
|
||||||
import org.springframework.ai.chat.messages.UserMessage;
|
import org.springframework.ai.chat.messages.UserMessage;
|
||||||
|
|
||||||
|
import java.util.ArrayList;
|
||||||
import java.util.LinkedHashMap;
|
import java.util.LinkedHashMap;
|
||||||
|
import java.util.LinkedHashSet;
|
||||||
import java.util.List;
|
import java.util.List;
|
||||||
import java.util.Map;
|
import java.util.Map;
|
||||||
|
import java.util.Set;
|
||||||
|
|
||||||
/**
|
/**
|
||||||
* Replaces verifier history with an explicit structured payload.
|
* Replaces verifier history with an explicit structured payload.
|
||||||
@@ -60,14 +63,18 @@ public class VerifierInputHook extends MessagesModelHook {
|
|||||||
executorFinalAnswer = extractLastAssistantText(previousMessages);
|
executorFinalAnswer = extractLastAssistantText(previousMessages);
|
||||||
}
|
}
|
||||||
|
|
||||||
ExecutorOutputParseResult parseResult = parseExecutorOutput(executorFinalAnswer);
|
|
||||||
VerifierContextHolder.setExecutorStructuredOutput(parseResult.structuredOutput());
|
|
||||||
VerifierContextHolder.setExecutorOutputParseStatus(parseResult.status());
|
|
||||||
|
|
||||||
List<Map<String, Object>> toolTraceSummary =
|
List<Map<String, Object>> toolTraceSummary =
|
||||||
toolTraceSummaryService.buildVerifierTraceSummary(sessionId, executorFinalAnswer);
|
toolTraceSummaryService.buildVerifierTraceSummary(sessionId, executorFinalAnswer);
|
||||||
VerifierContextHolder.setToolTraceSummary(toolTraceSummary);
|
VerifierContextHolder.setToolTraceSummary(toolTraceSummary);
|
||||||
|
|
||||||
|
ExecutorOutputParseResult parseResult = parseExecutorOutput(executorFinalAnswer);
|
||||||
|
parseResult = new ExecutorOutputParseResult(
|
||||||
|
enrichExecutorStructuredOutput(parseResult.structuredOutput(), toolTraceSummary),
|
||||||
|
parseResult.status()
|
||||||
|
);
|
||||||
|
VerifierContextHolder.setExecutorStructuredOutput(parseResult.structuredOutput());
|
||||||
|
VerifierContextHolder.setExecutorOutputParseStatus(parseResult.status());
|
||||||
|
|
||||||
Map<String, Object> gatekeeperResult = runGatekeeper(sessionId, parseResult);
|
Map<String, Object> gatekeeperResult = runGatekeeper(sessionId, parseResult);
|
||||||
VerifierContextHolder.setGatekeeperResult(gatekeeperResult);
|
VerifierContextHolder.setGatekeeperResult(gatekeeperResult);
|
||||||
|
|
||||||
@@ -105,6 +112,8 @@ public class VerifierInputHook extends MessagesModelHook {
|
|||||||
private Map<String, Object> passGatekeeperResult() {
|
private Map<String, Object> passGatekeeperResult() {
|
||||||
Map<String, Object> result = new LinkedHashMap<>();
|
Map<String, Object> result = new LinkedHashMap<>();
|
||||||
result.put("status", "pass");
|
result.put("status", "pass");
|
||||||
|
result.put("severity", "none");
|
||||||
|
result.put("checked_bindings", List.of());
|
||||||
result.put("failed_rules", List.of());
|
result.put("failed_rules", List.of());
|
||||||
result.put("warnings", List.of());
|
result.put("warnings", List.of());
|
||||||
result.put("errors", List.of());
|
result.put("errors", List.of());
|
||||||
@@ -135,6 +144,129 @@ public class VerifierInputHook extends MessagesModelHook {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@SuppressWarnings("unchecked")
|
||||||
|
private Map<String, Object> enrichExecutorStructuredOutput(Map<String, Object> structuredOutput,
|
||||||
|
List<Map<String, Object>> toolTraceSummary) {
|
||||||
|
if (structuredOutput == null) {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
Map<String, List<Long>> invocationIdsByTool = invocationIdsByTool(toolTraceSummary);
|
||||||
|
List<Map<String, Object>> warnings = new ArrayList<>();
|
||||||
|
enrichEvidenceBindingsInSection(structuredOutput.get("claims"), invocationIdsByTool, warnings);
|
||||||
|
enrichEvidenceBindingsInSection(structuredOutput.get("recommended_actions"), invocationIdsByTool, warnings);
|
||||||
|
if (!warnings.isEmpty()) {
|
||||||
|
structuredOutput.put("_gatekeeper_warnings", warnings);
|
||||||
|
}
|
||||||
|
return structuredOutput;
|
||||||
|
}
|
||||||
|
|
||||||
|
@SuppressWarnings("unchecked")
|
||||||
|
private void enrichEvidenceBindingsInSection(Object sectionValue,
|
||||||
|
Map<String, List<Long>> invocationIdsByTool,
|
||||||
|
List<Map<String, Object>> warnings) {
|
||||||
|
if (!(sectionValue instanceof List<?> items)) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
for (Object itemValue : items) {
|
||||||
|
if (!(itemValue instanceof Map<?, ?> item)) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
Object bindingsValue = item.get("evidence_bindings");
|
||||||
|
if (!(bindingsValue instanceof List<?> bindings)) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
for (Object bindingValue : bindings) {
|
||||||
|
if (!(bindingValue instanceof Map<?, ?> rawBinding)) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
Map<String, Object> binding = (Map<String, Object>) rawBinding;
|
||||||
|
String normalizedToolName = normalizeToolName(binding.get("tool_name"));
|
||||||
|
if (!normalizedToolName.isBlank()) {
|
||||||
|
binding.put("tool_name", normalizedToolName);
|
||||||
|
}
|
||||||
|
if (!hasInvocationId(binding)) {
|
||||||
|
List<Long> ids = invocationIdsByTool.getOrDefault(normalizedToolName, List.of());
|
||||||
|
if (ids.size() == 1) {
|
||||||
|
binding.put("source_invocation_id", ids.get(0));
|
||||||
|
warnings.add(Map.of(
|
||||||
|
"rule", "evidence.invocation_auto_backfill",
|
||||||
|
"message", "source_invocation_id was auto-filled from the unique tool invocation candidate; raw_path remains missing if Executor did not provide it",
|
||||||
|
"tool_name", normalizedToolName,
|
||||||
|
"source_invocation_id", ids.get(0)
|
||||||
|
));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private Map<String, List<Long>> invocationIdsByTool(List<Map<String, Object>> toolTraceSummary) {
|
||||||
|
Map<String, Set<Long>> idsByTool = new LinkedHashMap<>();
|
||||||
|
for (Map<String, Object> summary : toolTraceSummary == null ? List.<Map<String, Object>>of() : toolTraceSummary) {
|
||||||
|
String toolName = normalizeToolName(summary.get("tool_name"));
|
||||||
|
if (toolName.isBlank()) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
List<Long> ids = toLongList(summary.get("source_invocation_ids"));
|
||||||
|
if (ids.isEmpty()) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
idsByTool.computeIfAbsent(toolName, ignored -> new LinkedHashSet<>()).addAll(ids);
|
||||||
|
}
|
||||||
|
|
||||||
|
Map<String, List<Long>> result = new LinkedHashMap<>();
|
||||||
|
for (Map.Entry<String, Set<Long>> entry : idsByTool.entrySet()) {
|
||||||
|
result.put(entry.getKey(), new ArrayList<>(entry.getValue()));
|
||||||
|
}
|
||||||
|
return result;
|
||||||
|
}
|
||||||
|
|
||||||
|
private boolean hasInvocationId(Map<String, Object> binding) {
|
||||||
|
if (asLong(binding.get("source_invocation_id")) != null) {
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
return toLongList(binding.get("source_invocation_ids")).size() == 1;
|
||||||
|
}
|
||||||
|
|
||||||
|
private List<Long> toLongList(Object value) {
|
||||||
|
if (!(value instanceof List<?> values)) {
|
||||||
|
return List.of();
|
||||||
|
}
|
||||||
|
List<Long> ids = new ArrayList<>();
|
||||||
|
for (Object item : values) {
|
||||||
|
Long id = asLong(item);
|
||||||
|
if (id != null) {
|
||||||
|
ids.add(id);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return ids;
|
||||||
|
}
|
||||||
|
|
||||||
|
private Long asLong(Object value) {
|
||||||
|
if (value instanceof Number number) {
|
||||||
|
return number.longValue();
|
||||||
|
}
|
||||||
|
if (value instanceof String text) {
|
||||||
|
try {
|
||||||
|
return Long.parseLong(text);
|
||||||
|
} catch (NumberFormatException ignored) {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
|
||||||
|
private String normalizeToolName(Object value) {
|
||||||
|
String toolName = value == null ? "" : String.valueOf(value);
|
||||||
|
return switch (toolName) {
|
||||||
|
case "lookupKnowledge" -> "lookup_knowledge";
|
||||||
|
case "queryLogs" -> "query_logs";
|
||||||
|
case "queryPrometheusAlerts" -> "query_metrics";
|
||||||
|
case "getAvailableLogTopics" -> "get_available_log_topics";
|
||||||
|
default -> toolName;
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
private String sanitizeJsonPayload(String raw) {
|
private String sanitizeJsonPayload(String raw) {
|
||||||
String trimmed = raw.trim();
|
String trimmed = raw.trim();
|
||||||
int fenceStart = trimmed.indexOf("```");
|
int fenceStart = trimmed.indexOf("```");
|
||||||
|
|||||||
@@ -711,6 +711,13 @@ public class ChatService {
|
|||||||
if (gatekeeperResult == null || !"fail".equals(String.valueOf(gatekeeperResult.get("status")))) {
|
if (gatekeeperResult == null || !"fail".equals(String.valueOf(gatekeeperResult.get("status")))) {
|
||||||
return verdict;
|
return verdict;
|
||||||
}
|
}
|
||||||
|
String severity = String.valueOf(gatekeeperResult.getOrDefault("severity", ""));
|
||||||
|
if (ExecutorGatekeeperService.SEVERITY_REJECT.equals(severity)) {
|
||||||
|
return "REJECT";
|
||||||
|
}
|
||||||
|
if (ExecutorGatekeeperService.SEVERITY_LOW_CONFID.equals(severity)) {
|
||||||
|
return "PASS".equals(verdict) ? "LOW_CONFID" : verdict;
|
||||||
|
}
|
||||||
if (containsRule(gatekeeperResult.get("failed_rules"), ExecutorGatekeeperService.RULE_INVOCATION_REF)) {
|
if (containsRule(gatekeeperResult.get("failed_rules"), ExecutorGatekeeperService.RULE_INVOCATION_REF)) {
|
||||||
return "REJECT";
|
return "REJECT";
|
||||||
}
|
}
|
||||||
@@ -885,7 +892,8 @@ public class ChatService {
|
|||||||
Optional.ofNullable(VerifierContextHolder.getToolTraceSummary()).orElse(List.of()));
|
Optional.ofNullable(VerifierContextHolder.getToolTraceSummary()).orElse(List.of()));
|
||||||
verifierEvaluation.put("gatekeeper_result",
|
verifierEvaluation.put("gatekeeper_result",
|
||||||
Optional.ofNullable(VerifierContextHolder.getGatekeeperResult())
|
Optional.ofNullable(VerifierContextHolder.getGatekeeperResult())
|
||||||
.orElse(Map.of("status", "pass", "failed_rules", List.of(), "warnings", List.of(), "errors", List.of())));
|
.orElse(Map.of("status", "pass", "severity", "none", "checked_bindings", List.of(),
|
||||||
|
"failed_rules", List.of(), "warnings", List.of(), "errors", List.of())));
|
||||||
if (composerOutput != null) {
|
if (composerOutput != null) {
|
||||||
verifierEvaluation.put("composer_output", composerOutput);
|
verifierEvaluation.put("composer_output", composerOutput);
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -1,5 +1,7 @@
|
|||||||
package com.superbiz.agent.service;
|
package com.superbiz.agent.service;
|
||||||
|
|
||||||
|
import com.fasterxml.jackson.core.type.TypeReference;
|
||||||
|
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||||
import com.superbiz.agent.domain.entity.ToolInvocation;
|
import com.superbiz.agent.domain.entity.ToolInvocation;
|
||||||
import com.superbiz.agent.repository.ToolInvocationRepository;
|
import com.superbiz.agent.repository.ToolInvocationRepository;
|
||||||
import org.springframework.stereotype.Service;
|
import org.springframework.stereotype.Service;
|
||||||
@@ -21,12 +23,22 @@ import java.util.stream.Collectors;
|
|||||||
public class ExecutorGatekeeperService {
|
public class ExecutorGatekeeperService {
|
||||||
|
|
||||||
public static final String STATUS_PASS = "pass";
|
public static final String STATUS_PASS = "pass";
|
||||||
public static final String STATUS_WARN = "warn";
|
|
||||||
public static final String STATUS_FAIL = "fail";
|
public static final String STATUS_FAIL = "fail";
|
||||||
|
public static final String SEVERITY_NONE = "none";
|
||||||
|
public static final String SEVERITY_LOW_CONFID = "low_confid";
|
||||||
|
public static final String SEVERITY_REJECT = "reject";
|
||||||
public static final String RULE_SCHEMA = "schema.executor_v2";
|
public static final String RULE_SCHEMA = "schema.executor_v2";
|
||||||
public static final String RULE_INVOCATION_REF = "evidence.invocation_ref";
|
public static final String RULE_INVOCATION_REF = "evidence.invocation_ref";
|
||||||
|
public static final String RULE_RAW_PATH = "evidence.raw_path";
|
||||||
|
public static final String RULE_EXCERPT_MISMATCH = "evidence.excerpt_mismatch";
|
||||||
|
public static final String RULE_EVIDENCE_MISSING = "evidence.missing";
|
||||||
|
|
||||||
|
private static final TypeReference<Map<String, Object>> MAP_TYPE = new TypeReference<>() {
|
||||||
|
};
|
||||||
|
private static final double MIN_TOKEN_OVERLAP = 0.5;
|
||||||
|
|
||||||
private final ToolInvocationRepository toolInvocationRepository;
|
private final ToolInvocationRepository toolInvocationRepository;
|
||||||
|
private final ObjectMapper objectMapper = new ObjectMapper();
|
||||||
|
|
||||||
public ExecutorGatekeeperService(ToolInvocationRepository toolInvocationRepository) {
|
public ExecutorGatekeeperService(ToolInvocationRepository toolInvocationRepository) {
|
||||||
this.toolInvocationRepository = toolInvocationRepository;
|
this.toolInvocationRepository = toolInvocationRepository;
|
||||||
@@ -39,6 +51,7 @@ public class ExecutorGatekeeperService {
|
|||||||
validateSchema(structuredOutput, parseStatus, result);
|
validateSchema(structuredOutput, parseStatus, result);
|
||||||
if (structuredOutput != null) {
|
if (structuredOutput != null) {
|
||||||
validateInvocationRefs(sessionId, structuredOutput, result);
|
validateInvocationRefs(sessionId, structuredOutput, result);
|
||||||
|
importWarnings(structuredOutput, result);
|
||||||
}
|
}
|
||||||
return result.toMap();
|
return result.toMap();
|
||||||
}
|
}
|
||||||
@@ -49,7 +62,7 @@ public class ExecutorGatekeeperService {
|
|||||||
|
|
||||||
public Map<String, Object> fail(String ruleId, String target, String message) {
|
public Map<String, Object> fail(String ruleId, String target, String message) {
|
||||||
GatekeeperResult result = new GatekeeperResult();
|
GatekeeperResult result = new GatekeeperResult();
|
||||||
result.fail(ruleId, target, message);
|
result.fail(ruleId, target, message, SEVERITY_REJECT);
|
||||||
return result.toMap();
|
return result.toMap();
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -59,31 +72,35 @@ public class ExecutorGatekeeperService {
|
|||||||
String status = parseStatus == null ? "" : String.valueOf(parseStatus.getOrDefault("status", ""));
|
String status = parseStatus == null ? "" : String.valueOf(parseStatus.getOrDefault("status", ""));
|
||||||
if (structuredOutput == null) {
|
if (structuredOutput == null) {
|
||||||
if ("valid".equals(status)) {
|
if ("valid".equals(status)) {
|
||||||
result.fail(RULE_SCHEMA, "executor_structured_output", "structured output is missing after valid parse");
|
result.fail(RULE_SCHEMA, "executor_structured_output",
|
||||||
|
"structured output is missing after valid parse", SEVERITY_LOW_CONFID);
|
||||||
}
|
}
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
|
||||||
if (!"executor_evidence_v2".equals(String.valueOf(structuredOutput.get("answer_version")))) {
|
if (!"executor_evidence_v2".equals(String.valueOf(structuredOutput.get("answer_version")))) {
|
||||||
result.fail(RULE_SCHEMA, "answer_version", "answer_version must be executor_evidence_v2");
|
result.fail(RULE_SCHEMA, "answer_version", "answer_version must be executor_evidence_v2",
|
||||||
|
SEVERITY_LOW_CONFID);
|
||||||
}
|
}
|
||||||
if (structuredOutput.containsKey("diagnosis_summary")) {
|
if (structuredOutput.containsKey("diagnosis_summary")) {
|
||||||
result.fail(RULE_SCHEMA, "diagnosis_summary", "diagnosis_summary is removed from executor_evidence_v2");
|
result.fail(RULE_SCHEMA, "diagnosis_summary", "diagnosis_summary is removed from executor_evidence_v2",
|
||||||
|
SEVERITY_REJECT);
|
||||||
}
|
}
|
||||||
if (structuredOutput.containsKey("user_facing_answer")) {
|
if (structuredOutput.containsKey("user_facing_answer")) {
|
||||||
result.fail(RULE_SCHEMA, "user_facing_answer", "user_facing_answer is removed from executor_evidence_v2");
|
result.fail(RULE_SCHEMA, "user_facing_answer", "user_facing_answer is removed from executor_evidence_v2",
|
||||||
|
SEVERITY_REJECT);
|
||||||
}
|
}
|
||||||
|
|
||||||
Object claimsValue = structuredOutput.get("claims");
|
Object claimsValue = structuredOutput.get("claims");
|
||||||
if (!(claimsValue instanceof List<?> claims)) {
|
if (!(claimsValue instanceof List<?> claims)) {
|
||||||
result.fail(RULE_SCHEMA, "claims", "claims must be an array");
|
result.fail(RULE_SCHEMA, "claims", "claims must be an array", SEVERITY_LOW_CONFID);
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
for (int i = 0; i < claims.size(); i++) {
|
for (int i = 0; i < claims.size(); i++) {
|
||||||
String target = "claims[" + i + "]";
|
String target = "claims[" + i + "]";
|
||||||
Object claimValue = claims.get(i);
|
Object claimValue = claims.get(i);
|
||||||
if (!(claimValue instanceof Map<?, ?> claim)) {
|
if (!(claimValue instanceof Map<?, ?> claim)) {
|
||||||
result.fail(RULE_SCHEMA, target, "claim must be an object");
|
result.fail(RULE_SCHEMA, target, "claim must be an object", SEVERITY_LOW_CONFID);
|
||||||
continue;
|
continue;
|
||||||
}
|
}
|
||||||
requireString(claim, "claim_id", target, result);
|
requireString(claim, "claim_id", target, result);
|
||||||
@@ -91,11 +108,13 @@ public class ExecutorGatekeeperService {
|
|||||||
requireString(claim, "claim_text", target, result);
|
requireString(claim, "claim_text", target, result);
|
||||||
String supportLevel = stringValue(claim.get("support_level"));
|
String supportLevel = stringValue(claim.get("support_level"));
|
||||||
if (!"direct".equals(supportLevel) && !"indirect".equals(supportLevel)) {
|
if (!"direct".equals(supportLevel) && !"indirect".equals(supportLevel)) {
|
||||||
result.fail(RULE_SCHEMA, target + ".support_level", "support_level must be direct or indirect");
|
result.fail(RULE_SCHEMA, target + ".support_level", "support_level must be direct or indirect",
|
||||||
|
SEVERITY_LOW_CONFID);
|
||||||
}
|
}
|
||||||
Object bindings = claim.get("evidence_bindings");
|
Object bindings = claim.get("evidence_bindings");
|
||||||
if (!(bindings instanceof List<?> bindingList) || bindingList.isEmpty()) {
|
if (!(bindings instanceof List<?> bindingList) || bindingList.isEmpty()) {
|
||||||
result.fail(RULE_SCHEMA, target + ".evidence_bindings", "claims must include non-empty evidence_bindings");
|
result.fail(RULE_EVIDENCE_MISSING, target + ".evidence_bindings",
|
||||||
|
"claims must include non-empty evidence_bindings", SEVERITY_LOW_CONFID);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -106,7 +125,8 @@ public class ExecutorGatekeeperService {
|
|||||||
|
|
||||||
private void validateInvocationRefs(String sessionId, Map<String, Object> structuredOutput, GatekeeperResult result) {
|
private void validateInvocationRefs(String sessionId, Map<String, Object> structuredOutput, GatekeeperResult result) {
|
||||||
if (sessionId == null || sessionId.isBlank()) {
|
if (sessionId == null || sessionId.isBlank()) {
|
||||||
result.fail(RULE_INVOCATION_REF, "session_id", "session id is required to validate source_invocation_ids");
|
result.fail(RULE_INVOCATION_REF, "session_id", "session id is required to validate source_invocation_id",
|
||||||
|
SEVERITY_LOW_CONFID);
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -132,62 +152,254 @@ public class ExecutorGatekeeperService {
|
|||||||
String target = "claims[" + claimIndex + "].evidence_bindings[" + bindingIndex + "]";
|
String target = "claims[" + claimIndex + "].evidence_bindings[" + bindingIndex + "]";
|
||||||
Object bindingValue = bindings.get(bindingIndex);
|
Object bindingValue = bindings.get(bindingIndex);
|
||||||
if (!(bindingValue instanceof Map<?, ?> binding)) {
|
if (!(bindingValue instanceof Map<?, ?> binding)) {
|
||||||
result.fail(RULE_INVOCATION_REF, target, "evidence binding must be an object");
|
result.fail(RULE_INVOCATION_REF, target, "evidence binding must be an object",
|
||||||
|
SEVERITY_LOW_CONFID);
|
||||||
continue;
|
continue;
|
||||||
}
|
}
|
||||||
validateBindingInvocationIds(binding, validInvocations, target, result);
|
validateEvidenceBinding(binding, validInvocations, target, claim.get("claim_id"), result);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
Object actionsValue = structuredOutput.get("recommended_actions");
|
||||||
|
if (!(actionsValue instanceof List<?> actions)) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
for (int actionIndex = 0; actionIndex < actions.size(); actionIndex++) {
|
||||||
|
Object actionValue = actions.get(actionIndex);
|
||||||
|
if (!(actionValue instanceof Map<?, ?> action)) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
Object bindingsValue = action.get("evidence_bindings");
|
||||||
|
if (!(bindingsValue instanceof List<?> bindings)) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
for (int bindingIndex = 0; bindingIndex < bindings.size(); bindingIndex++) {
|
||||||
|
String target = "recommended_actions[" + actionIndex + "].evidence_bindings[" + bindingIndex + "]";
|
||||||
|
Object bindingValue = bindings.get(bindingIndex);
|
||||||
|
if (!(bindingValue instanceof Map<?, ?> binding)) {
|
||||||
|
result.fail(RULE_INVOCATION_REF, target, "evidence binding must be an object",
|
||||||
|
SEVERITY_LOW_CONFID);
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
validateEvidenceBinding(binding, validInvocations, target, action.get("action_id"), result);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
private void validateBindingInvocationIds(Map<?, ?> binding,
|
private void validateEvidenceBinding(Map<?, ?> binding,
|
||||||
Map<Long, ToolInvocation> validInvocations,
|
Map<Long, ToolInvocation> validInvocations,
|
||||||
String target,
|
String target,
|
||||||
|
Object ownerId,
|
||||||
GatekeeperResult result) {
|
GatekeeperResult result) {
|
||||||
Object idsValue = binding.get("source_invocation_ids");
|
Map<String, Object> checked = new LinkedHashMap<>();
|
||||||
if (!(idsValue instanceof List<?> ids) || ids.isEmpty()) {
|
checked.put("claim_id", ownerId == null ? "" : String.valueOf(ownerId));
|
||||||
result.fail(RULE_INVOCATION_REF, target + ".source_invocation_ids",
|
checked.put("tool_name", stringValue(binding.get("tool_name")));
|
||||||
"source_invocation_ids must be a non-empty array");
|
checked.put("source_invocation_id", binding.get("source_invocation_id"));
|
||||||
|
checked.put("raw_path", stringValue(binding.get("raw_path")));
|
||||||
|
|
||||||
|
Long id = singleInvocationId(binding);
|
||||||
|
if (id == null) {
|
||||||
|
checked.put("status", STATUS_FAIL);
|
||||||
|
checked.put("rule", RULE_INVOCATION_REF);
|
||||||
|
checked.put("message", "source_invocation_id is required");
|
||||||
|
result.checked(checked);
|
||||||
|
result.fail(RULE_INVOCATION_REF, target + ".source_invocation_id",
|
||||||
|
"source_invocation_id is required", SEVERITY_LOW_CONFID);
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
checked.put("source_invocation_id", id);
|
||||||
|
|
||||||
String claimedToolName = stringValue(binding.get("tool_name"));
|
String claimedToolName = stringValue(binding.get("tool_name"));
|
||||||
if (claimedToolName.isBlank()) {
|
if (claimedToolName.isBlank()) {
|
||||||
result.fail(RULE_INVOCATION_REF, target + ".tool_name", "tool_name is required");
|
checked.put("status", STATUS_FAIL);
|
||||||
|
checked.put("rule", RULE_INVOCATION_REF);
|
||||||
|
checked.put("message", "tool_name is required");
|
||||||
|
result.checked(checked);
|
||||||
|
result.fail(RULE_INVOCATION_REF, target + ".tool_name", "tool_name is required", SEVERITY_LOW_CONFID);
|
||||||
|
return;
|
||||||
}
|
}
|
||||||
|
|
||||||
Set<Long> checkedIds = new HashSet<>();
|
|
||||||
for (Object idValue : ids) {
|
|
||||||
Long id = asLong(idValue);
|
|
||||||
if (id == null) {
|
|
||||||
result.fail(RULE_INVOCATION_REF, target + ".source_invocation_ids",
|
|
||||||
"source_invocation_ids must contain numeric ids");
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
if (!checkedIds.add(id)) {
|
|
||||||
continue;
|
|
||||||
}
|
|
||||||
ToolInvocation invocation = validInvocations.get(id);
|
ToolInvocation invocation = validInvocations.get(id);
|
||||||
if (invocation == null) {
|
if (invocation == null) {
|
||||||
result.fail(RULE_INVOCATION_REF, target, "source_invocation_ids not found in current session: " + id);
|
checked.put("status", STATUS_FAIL);
|
||||||
|
checked.put("rule", RULE_INVOCATION_REF);
|
||||||
|
checked.put("message", "source_invocation_id not found in current session: " + id);
|
||||||
|
result.checked(checked);
|
||||||
|
result.fail(RULE_INVOCATION_REF, target,
|
||||||
|
"source_invocation_id not found in current session: " + id, SEVERITY_REJECT);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (!Objects.equals(claimedToolName, invocation.getToolName())) {
|
||||||
|
checked.put("status", STATUS_FAIL);
|
||||||
|
checked.put("rule", RULE_INVOCATION_REF);
|
||||||
|
checked.put("message", "tool_name does not match invocation " + id + ": expected " + invocation.getToolName());
|
||||||
|
result.checked(checked);
|
||||||
|
result.fail(RULE_INVOCATION_REF, target + ".tool_name",
|
||||||
|
"tool_name does not match invocation " + id + ": expected " + invocation.getToolName(),
|
||||||
|
SEVERITY_REJECT);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
String rawPath = stringValue(binding.get("raw_path"));
|
||||||
|
if (rawPath.isBlank()) {
|
||||||
|
checked.put("status", STATUS_FAIL);
|
||||||
|
checked.put("rule", RULE_RAW_PATH);
|
||||||
|
checked.put("message", "raw_path is required for precise evidence reference");
|
||||||
|
result.checked(checked);
|
||||||
|
result.fail(RULE_RAW_PATH, target + ".raw_path",
|
||||||
|
"raw_path is required for precise evidence reference", SEVERITY_LOW_CONFID);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
Map<String, String> refs = evidenceRefsByRawPath(invocation.getRetrievalDetails());
|
||||||
|
if (refs.isEmpty()) {
|
||||||
|
checked.put("status", STATUS_FAIL);
|
||||||
|
checked.put("rule", RULE_EVIDENCE_MISSING);
|
||||||
|
checked.put("message", "invocation has no retrieval_details.evidence_refs");
|
||||||
|
result.checked(checked);
|
||||||
|
result.fail(RULE_EVIDENCE_MISSING, target,
|
||||||
|
"invocation has no retrieval_details.evidence_refs", SEVERITY_LOW_CONFID);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
String matchedText = refs.get(rawPath);
|
||||||
|
if (matchedText == null) {
|
||||||
|
checked.put("status", STATUS_FAIL);
|
||||||
|
checked.put("rule", RULE_RAW_PATH);
|
||||||
|
checked.put("message", "raw_path not found in retrieval_details.evidence_refs");
|
||||||
|
result.checked(checked);
|
||||||
|
result.fail(RULE_RAW_PATH, target + ".raw_path",
|
||||||
|
"raw_path not found in retrieval_details.evidence_refs", SEVERITY_REJECT);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
checked.put("matched_text", matchedText);
|
||||||
|
|
||||||
|
String excerpt = stringValue(binding.get("evidence_excerpt"));
|
||||||
|
if (normalized(excerpt).length() < 8) {
|
||||||
|
checked.put("status", STATUS_FAIL);
|
||||||
|
checked.put("rule", RULE_EXCERPT_MISMATCH);
|
||||||
|
checked.put("message", "evidence_excerpt is too short to compare");
|
||||||
|
result.checked(checked);
|
||||||
|
result.fail(RULE_EXCERPT_MISMATCH, target + ".evidence_excerpt",
|
||||||
|
"evidence_excerpt is too short to compare", SEVERITY_LOW_CONFID);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
if (!isExcerptSupported(excerpt, matchedText)) {
|
||||||
|
checked.put("status", STATUS_FAIL);
|
||||||
|
checked.put("rule", RULE_EXCERPT_MISMATCH);
|
||||||
|
checked.put("message", "evidence_excerpt is not supported by matched evidence ref text");
|
||||||
|
result.checked(checked);
|
||||||
|
result.fail(RULE_EXCERPT_MISMATCH, target + ".evidence_excerpt",
|
||||||
|
"evidence_excerpt is not supported by matched evidence ref text", SEVERITY_REJECT);
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
|
||||||
|
checked.put("status", STATUS_PASS);
|
||||||
|
result.checked(checked);
|
||||||
|
}
|
||||||
|
|
||||||
|
private Long singleInvocationId(Map<?, ?> binding) {
|
||||||
|
Long singular = asLong(binding.get("source_invocation_id"));
|
||||||
|
if (singular != null) {
|
||||||
|
return singular;
|
||||||
|
}
|
||||||
|
Object idsValue = binding.get("source_invocation_ids");
|
||||||
|
if (!(idsValue instanceof List<?> ids) || ids.size() != 1) {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
return asLong(ids.get(0));
|
||||||
|
}
|
||||||
|
|
||||||
|
@SuppressWarnings("unchecked")
|
||||||
|
private void importWarnings(Map<String, Object> structuredOutput, GatekeeperResult result) {
|
||||||
|
Object warningsValue = structuredOutput.remove("_gatekeeper_warnings");
|
||||||
|
if (!(warningsValue instanceof List<?> warnings)) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
for (Object warning : warnings) {
|
||||||
|
if (warning instanceof Map<?, ?> map) {
|
||||||
|
result.warn((Map<String, Object>) map);
|
||||||
|
} else if (warning != null) {
|
||||||
|
result.warn(Map.of("message", String.valueOf(warning)));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private Map<String, String> evidenceRefsByRawPath(String retrievalDetails) {
|
||||||
|
if (retrievalDetails == null || retrievalDetails.isBlank()) {
|
||||||
|
return Map.of();
|
||||||
|
}
|
||||||
|
try {
|
||||||
|
Map<String, Object> details = objectMapper.readValue(retrievalDetails, MAP_TYPE);
|
||||||
|
Object refsValue = details.get("evidence_refs");
|
||||||
|
if (!(refsValue instanceof List<?> refs)) {
|
||||||
|
return Map.of();
|
||||||
|
}
|
||||||
|
Map<String, String> result = new LinkedHashMap<>();
|
||||||
|
for (Object refValue : refs) {
|
||||||
|
if (!(refValue instanceof Map<?, ?> ref)) {
|
||||||
continue;
|
continue;
|
||||||
}
|
}
|
||||||
if (!claimedToolName.isBlank() && !Objects.equals(claimedToolName, invocation.getToolName())) {
|
String rawPath = stringValue(ref.get("raw_path"));
|
||||||
result.fail(RULE_INVOCATION_REF, target + ".tool_name",
|
String text = stringValue(ref.get("text"));
|
||||||
"tool_name does not match invocation " + id + ": expected " + invocation.getToolName());
|
if (!rawPath.isBlank() && !text.isBlank()) {
|
||||||
|
result.put(rawPath, text);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
return result;
|
||||||
|
} catch (Exception ignored) {
|
||||||
|
return Map.of();
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private boolean isExcerptSupported(String excerpt, String matchedText) {
|
||||||
|
String normalizedExcerpt = normalized(excerpt);
|
||||||
|
String normalizedMatched = normalized(matchedText);
|
||||||
|
if (normalizedMatched.contains(normalizedExcerpt) || normalizedExcerpt.contains(normalizedMatched)) {
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
Set<String> excerptTokens = tokens(normalizedExcerpt);
|
||||||
|
if (excerptTokens.isEmpty()) {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
Set<String> matchedTokens = tokens(normalizedMatched);
|
||||||
|
int overlap = 0;
|
||||||
|
for (String token : excerptTokens) {
|
||||||
|
if (matchedTokens.contains(token)) {
|
||||||
|
overlap++;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return (double) overlap / excerptTokens.size() >= MIN_TOKEN_OVERLAP;
|
||||||
|
}
|
||||||
|
|
||||||
|
private String normalized(String value) {
|
||||||
|
return value == null ? "" : value.toLowerCase()
|
||||||
|
.replaceAll("[\\p{Punct}\\s,。;:、()【】《》“”‘’]+", " ")
|
||||||
|
.trim();
|
||||||
|
}
|
||||||
|
|
||||||
|
private Set<String> tokens(String text) {
|
||||||
|
if (text == null || text.isBlank()) {
|
||||||
|
return Set.of();
|
||||||
|
}
|
||||||
|
Set<String> result = new HashSet<>();
|
||||||
|
for (String token : text.split("\\s+")) {
|
||||||
|
if (token.length() >= 2) {
|
||||||
|
result.add(token);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return result;
|
||||||
}
|
}
|
||||||
|
|
||||||
private void requireArray(Map<String, Object> output, String field, GatekeeperResult result) {
|
private void requireArray(Map<String, Object> output, String field, GatekeeperResult result) {
|
||||||
if (!(output.get(field) instanceof List<?>)) {
|
if (!(output.get(field) instanceof List<?>)) {
|
||||||
result.fail(RULE_SCHEMA, field, field + " must be an array");
|
result.fail(RULE_SCHEMA, field, field + " must be an array", SEVERITY_LOW_CONFID);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
private void requireString(Map<?, ?> object, String field, String target, GatekeeperResult result) {
|
private void requireString(Map<?, ?> object, String field, String target, GatekeeperResult result) {
|
||||||
if (stringValue(object.get(field)).isBlank()) {
|
if (stringValue(object.get(field)).isBlank()) {
|
||||||
result.fail(RULE_SCHEMA, target + "." + field, field + " is required");
|
result.fail(RULE_SCHEMA, target + "." + field, field + " is required", SEVERITY_LOW_CONFID);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -211,23 +423,42 @@ public class ExecutorGatekeeperService {
|
|||||||
|
|
||||||
private static final class GatekeeperResult {
|
private static final class GatekeeperResult {
|
||||||
private final List<String> failedRules = new ArrayList<>();
|
private final List<String> failedRules = new ArrayList<>();
|
||||||
private final List<String> warnings = new ArrayList<>();
|
private final List<Map<String, Object>> checkedBindings = new ArrayList<>();
|
||||||
|
private final List<Map<String, Object>> warnings = new ArrayList<>();
|
||||||
private final List<Map<String, Object>> errors = new ArrayList<>();
|
private final List<Map<String, Object>> errors = new ArrayList<>();
|
||||||
|
private String severity = SEVERITY_NONE;
|
||||||
|
|
||||||
void fail(String ruleId, String target, String message) {
|
void fail(String ruleId, String target, String message, String failureSeverity) {
|
||||||
if (!failedRules.contains(ruleId)) {
|
if (!failedRules.contains(ruleId)) {
|
||||||
failedRules.add(ruleId);
|
failedRules.add(ruleId);
|
||||||
}
|
}
|
||||||
|
if (SEVERITY_REJECT.equals(failureSeverity)) {
|
||||||
|
severity = SEVERITY_REJECT;
|
||||||
|
} else if (!SEVERITY_REJECT.equals(severity)) {
|
||||||
|
severity = SEVERITY_LOW_CONFID;
|
||||||
|
}
|
||||||
Map<String, Object> error = new LinkedHashMap<>();
|
Map<String, Object> error = new LinkedHashMap<>();
|
||||||
error.put("rule_id", ruleId);
|
error.put("rule_id", ruleId);
|
||||||
|
error.put("rule", ruleId);
|
||||||
error.put("target", target);
|
error.put("target", target);
|
||||||
error.put("message", message);
|
error.put("message", message);
|
||||||
|
error.put("severity", failureSeverity);
|
||||||
errors.add(error);
|
errors.add(error);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
void checked(Map<String, Object> checked) {
|
||||||
|
checkedBindings.add(checked);
|
||||||
|
}
|
||||||
|
|
||||||
|
void warn(Map<String, Object> warning) {
|
||||||
|
warnings.add(new LinkedHashMap<>(warning));
|
||||||
|
}
|
||||||
|
|
||||||
Map<String, Object> toMap() {
|
Map<String, Object> toMap() {
|
||||||
Map<String, Object> result = new LinkedHashMap<>();
|
Map<String, Object> result = new LinkedHashMap<>();
|
||||||
result.put("status", failedRules.isEmpty() ? (warnings.isEmpty() ? STATUS_PASS : STATUS_WARN) : STATUS_FAIL);
|
result.put("status", failedRules.isEmpty() ? STATUS_PASS : STATUS_FAIL);
|
||||||
|
result.put("severity", failedRules.isEmpty() ? SEVERITY_NONE : severity);
|
||||||
|
result.put("checked_bindings", checkedBindings);
|
||||||
result.put("failed_rules", failedRules);
|
result.put("failed_rules", failedRules);
|
||||||
result.put("warnings", warnings);
|
result.put("warnings", warnings);
|
||||||
result.put("errors", errors);
|
result.put("errors", errors);
|
||||||
|
|||||||
@@ -1,6 +1,7 @@
|
|||||||
package com.superbiz.agent.service;
|
package com.superbiz.agent.service;
|
||||||
|
|
||||||
import com.fasterxml.jackson.core.JsonProcessingException;
|
import com.fasterxml.jackson.core.JsonProcessingException;
|
||||||
|
import com.fasterxml.jackson.databind.JsonNode;
|
||||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||||
import com.superbiz.agent.domain.entity.ToolInvocation;
|
import com.superbiz.agent.domain.entity.ToolInvocation;
|
||||||
import com.superbiz.agent.dto.ContextPack;
|
import com.superbiz.agent.dto.ContextPack;
|
||||||
@@ -87,6 +88,10 @@ public class ToolInvocationRecorder {
|
|||||||
if (extraDetails != null && !extraDetails.isEmpty()) {
|
if (extraDetails != null && !extraDetails.isEmpty()) {
|
||||||
details.putAll(extraDetails);
|
details.putAll(extraDetails);
|
||||||
}
|
}
|
||||||
|
List<Map<String, Object>> evidenceRefs = extractEvidenceRefs(toolName, output, details);
|
||||||
|
if (!evidenceRefs.isEmpty()) {
|
||||||
|
details.put("evidence_refs", evidenceRefs);
|
||||||
|
}
|
||||||
|
|
||||||
ToolInvocation invocation = ToolInvocation.builder()
|
ToolInvocation invocation = ToolInvocation.builder()
|
||||||
.toolName(toolName)
|
.toolName(toolName)
|
||||||
@@ -139,6 +144,10 @@ public class ToolInvocationRecorder {
|
|||||||
details.put("evidence_block_count", record.evidenceBlockCount());
|
details.put("evidence_block_count", record.evidenceBlockCount());
|
||||||
}
|
}
|
||||||
details.put("evidence_blocks", record.evidenceBlocks() == null ? List.of() : record.evidenceBlocks());
|
details.put("evidence_blocks", record.evidenceBlocks() == null ? List.of() : record.evidenceBlocks());
|
||||||
|
List<Map<String, Object>> evidenceRefs = evidenceRefsFromEvidenceBlocks(record.evidenceBlocks());
|
||||||
|
if (!evidenceRefs.isEmpty()) {
|
||||||
|
details.put("evidence_refs", evidenceRefs);
|
||||||
|
}
|
||||||
details.put("query_transform", record.queryTransform() == null ? Map.of() : record.queryTransform());
|
details.put("query_transform", record.queryTransform() == null ? Map.of() : record.queryTransform());
|
||||||
details.put("retrieval_trace", record.retrievalTrace() == null ? Map.of() : record.retrievalTrace());
|
details.put("retrieval_trace", record.retrievalTrace() == null ? Map.of() : record.retrievalTrace());
|
||||||
details.put("context_pack_summary", record.contextPack() == null ? Map.of() : record.contextPack());
|
details.put("context_pack_summary", record.contextPack() == null ? Map.of() : record.contextPack());
|
||||||
@@ -186,6 +195,132 @@ public class ToolInvocationRecorder {
|
|||||||
return success ? EVIDENCE_STATUS_SUPPORTED : EVIDENCE_STATUS_FAILED;
|
return success ? EVIDENCE_STATUS_SUPPORTED : EVIDENCE_STATUS_FAILED;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
private List<Map<String, Object>> extractEvidenceRefs(String toolName, String output, Map<String, Object> details) {
|
||||||
|
if (output == null || output.isBlank()) {
|
||||||
|
return List.of();
|
||||||
|
}
|
||||||
|
if ("lookup_knowledge".equals(toolName)) {
|
||||||
|
return evidenceRefsFromEvidenceBlocks(asMapList(details.get("evidence_blocks")));
|
||||||
|
}
|
||||||
|
|
||||||
|
try {
|
||||||
|
JsonNode root = objectMapper.readTree(output);
|
||||||
|
if ("query_metrics".equals(toolName)) {
|
||||||
|
return evidenceRefsFromArray(root.path("alerts"), "$.alerts", this::alertText);
|
||||||
|
}
|
||||||
|
if ("query_logs".equals(toolName)) {
|
||||||
|
return evidenceRefsFromArray(root.path("logs"), "$.logs", this::logText);
|
||||||
|
}
|
||||||
|
} catch (Exception e) {
|
||||||
|
log.debug("extract evidence_refs failed for tool={}", toolName, e);
|
||||||
|
}
|
||||||
|
return List.of();
|
||||||
|
}
|
||||||
|
|
||||||
|
private List<Map<String, Object>> evidenceRefsFromArray(JsonNode arrayNode,
|
||||||
|
String pathPrefix,
|
||||||
|
java.util.function.Function<JsonNode, String> textExtractor) {
|
||||||
|
if (!arrayNode.isArray()) {
|
||||||
|
return List.of();
|
||||||
|
}
|
||||||
|
List<Map<String, Object>> refs = new ArrayList<>();
|
||||||
|
for (int i = 0; i < arrayNode.size(); i++) {
|
||||||
|
String text = textExtractor.apply(arrayNode.get(i));
|
||||||
|
if (text == null || text.isBlank()) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
refs.add(Map.of(
|
||||||
|
"raw_path", pathPrefix + "[" + i + "]",
|
||||||
|
"text", bounded(text, 500)
|
||||||
|
));
|
||||||
|
}
|
||||||
|
return refs;
|
||||||
|
}
|
||||||
|
|
||||||
|
private String alertText(JsonNode alert) {
|
||||||
|
List<String> parts = new ArrayList<>();
|
||||||
|
addPart(parts, textField(alert, "alert_name"));
|
||||||
|
addPart(parts, textField(alert, "state"));
|
||||||
|
addPart(parts, textField(alert, "description"));
|
||||||
|
addPart(parts, "active_at=" + textField(alert, "active_at"));
|
||||||
|
addPart(parts, "duration=" + textField(alert, "duration"));
|
||||||
|
return String.join(", ", parts);
|
||||||
|
}
|
||||||
|
|
||||||
|
private String logText(JsonNode log) {
|
||||||
|
List<String> parts = new ArrayList<>();
|
||||||
|
addPart(parts, textField(log, "timestamp"));
|
||||||
|
addPart(parts, textField(log, "level"));
|
||||||
|
addPart(parts, textField(log, "service"));
|
||||||
|
addPart(parts, textField(log, "message"));
|
||||||
|
JsonNode metrics = log.path("metrics");
|
||||||
|
if (metrics.isObject() && !metrics.isEmpty()) {
|
||||||
|
addPart(parts, "metrics=" + metrics.toString());
|
||||||
|
}
|
||||||
|
return String.join(" ", parts);
|
||||||
|
}
|
||||||
|
|
||||||
|
private List<Map<String, Object>> evidenceRefsFromEvidenceBlocks(List<Map<String, Object>> blocks) {
|
||||||
|
if (blocks == null || blocks.isEmpty()) {
|
||||||
|
return List.of();
|
||||||
|
}
|
||||||
|
List<Map<String, Object>> refs = new ArrayList<>();
|
||||||
|
for (int i = 0; i < blocks.size(); i++) {
|
||||||
|
Map<String, Object> block = blocks.get(i);
|
||||||
|
String text = firstNonBlank(block.get("content_preview"), block.get("content"),
|
||||||
|
block.get("title"), block.get("source"));
|
||||||
|
if (text.isBlank()) {
|
||||||
|
continue;
|
||||||
|
}
|
||||||
|
refs.add(Map.of(
|
||||||
|
"raw_path", "$.evidence_blocks[" + i + "]",
|
||||||
|
"text", bounded(text, 500)
|
||||||
|
));
|
||||||
|
}
|
||||||
|
return refs;
|
||||||
|
}
|
||||||
|
|
||||||
|
@SuppressWarnings("unchecked")
|
||||||
|
private List<Map<String, Object>> asMapList(Object value) {
|
||||||
|
if (!(value instanceof List<?> list)) {
|
||||||
|
return List.of();
|
||||||
|
}
|
||||||
|
List<Map<String, Object>> result = new ArrayList<>();
|
||||||
|
for (Object item : list) {
|
||||||
|
if (item instanceof Map<?, ?> map) {
|
||||||
|
result.add((Map<String, Object>) map);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return result;
|
||||||
|
}
|
||||||
|
|
||||||
|
private String textField(JsonNode node, String field) {
|
||||||
|
JsonNode value = node.path(field);
|
||||||
|
return value.isMissingNode() || value.isNull() ? "" : value.asText("");
|
||||||
|
}
|
||||||
|
|
||||||
|
private void addPart(List<String> parts, String value) {
|
||||||
|
if (value != null && !value.isBlank() && !value.endsWith("=")) {
|
||||||
|
parts.add(value);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private String firstNonBlank(Object... values) {
|
||||||
|
for (Object value : values) {
|
||||||
|
if (value != null && !String.valueOf(value).isBlank()) {
|
||||||
|
return String.valueOf(value);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
return "";
|
||||||
|
}
|
||||||
|
|
||||||
|
private String bounded(String value, int limit) {
|
||||||
|
if (value == null) {
|
||||||
|
return "";
|
||||||
|
}
|
||||||
|
return value.length() <= limit ? value : value.substring(0, limit) + "...";
|
||||||
|
}
|
||||||
|
|
||||||
private String preview(String output) {
|
private String preview(String output) {
|
||||||
if (output == null) {
|
if (output == null) {
|
||||||
return null;
|
return null;
|
||||||
|
|||||||
@@ -5,6 +5,7 @@
|
|||||||
- 需要外部信息时调用工具,但必须遵守下方的检索约束。
|
- 需要外部信息时调用工具,但必须遵守下方的检索约束。
|
||||||
- 严禁凭记忆回答,必须基于本轮工具返回的真实数据。
|
- 严禁凭记忆回答,必须基于本轮工具返回的真实数据。
|
||||||
- 执行完成后,输出严格的证据归因 JSON,供 Verifier 校验。
|
- 执行完成后,输出严格的证据归因 JSON,供 Verifier 校验。
|
||||||
|
- 你是证据收集与微观事实提炼器,不是最终答复生成器。
|
||||||
|
|
||||||
## 规则
|
## 规则
|
||||||
- 按顺序执行,不可跳过步骤。
|
- 按顺序执行,不可跳过步骤。
|
||||||
@@ -12,6 +13,8 @@
|
|||||||
- runbook、skill、历史案例、知识库中的通用模式只能作为排查指导或建议动作,不能直接写成本次事故的已确认事实。
|
- runbook、skill、历史案例、知识库中的通用模式只能作为排查指导或建议动作,不能直接写成本次事故的已确认事实。
|
||||||
- 如果检索内容不足以支撑结论,必须显式声明证据不足,严禁补全事故故事。
|
- 如果检索内容不足以支撑结论,必须显式声明证据不足,严禁补全事故故事。
|
||||||
- 不要使用“通常情况下”“根据经验”“很可能已经发生”等无证据推断词来伪装事实。
|
- 不要使用“通常情况下”“根据经验”“很可能已经发生”等无证据推断词来伪装事实。
|
||||||
|
- 对窄范围确认问题,只输出与用户问题直接相关的 observation / negative_observation。通常 1 条 claim,最多 2 条 claim;不要限制 evidence_bindings 数量。
|
||||||
|
- 禁止把根因、修复动作或用户明确排除的服务/主题写成 confirmed claim,除非本轮工具证据直接证明。
|
||||||
|
|
||||||
## 检索约束
|
## 检索约束
|
||||||
|
|
||||||
@@ -59,6 +62,7 @@
|
|||||||
### recommended_actions
|
### recommended_actions
|
||||||
`recommended_actions` 用来放下一步排查或修复动作。
|
`recommended_actions` 用来放下一步排查或修复动作。
|
||||||
建议可以来自 runbook/skill,但必须说明 reason,不能写成“已确认根因”。
|
建议可以来自 runbook/skill,但必须说明 reason,不能写成“已确认根因”。
|
||||||
|
本期 recommended_actions 只允许证据收集或继续排查动作,不要输出重启、扩容、修改配置等修复动作,除非用户明确要求执行方案。
|
||||||
|
|
||||||
### missing_info
|
### missing_info
|
||||||
`missing_info` 用来列出无法确认结论所缺少的具体证据。
|
`missing_info` 用来列出无法确认结论所缺少的具体证据。
|
||||||
@@ -80,9 +84,10 @@
|
|||||||
"evidence_bindings": [
|
"evidence_bindings": [
|
||||||
{
|
{
|
||||||
"source_type": "tool_trace",
|
"source_type": "tool_trace",
|
||||||
"source_id": "工具返回中的 evidence block id、trace_ref 或可定位标识",
|
"source_id": "工具返回中的 evidence block id、trace_ref 或可定位标识,可为空",
|
||||||
"tool_name": "lookup_knowledge/query_logs/query_metrics/read_skill 等",
|
"tool_name": "lookup_knowledge/query_logs/query_metrics 等 evidence tool",
|
||||||
"source_invocation_ids": [],
|
"source_invocation_id": null,
|
||||||
|
"raw_path": "$.alerts[0] / $.logs[0] / $.evidence_blocks[0]",
|
||||||
"evidence_excerpt": "从工具返回中摘取的原话、指标值、日志片段或关键数据"
|
"evidence_excerpt": "从工具返回中摘取的原话、指标值、日志片段或关键数据"
|
||||||
}
|
}
|
||||||
]
|
]
|
||||||
@@ -115,5 +120,8 @@
|
|||||||
- `claims[*].support_level` 只能是 `direct` 或 `indirect`。
|
- `claims[*].support_level` 只能是 `direct` 或 `indirect`。
|
||||||
- `claims[*].evidence_bindings` 不能为空。
|
- `claims[*].evidence_bindings` 不能为空。
|
||||||
- `evidence_excerpt` 必须来自工具返回,不允许编造。
|
- `evidence_excerpt` 必须来自工具返回,不允许编造。
|
||||||
|
- `raw_path` 必须指向工具返回数组中的具体条目:`query_metrics` 使用 `$.alerts[i]`,`query_logs` 使用 `$.logs[i]`,`lookup_knowledge` 使用 `$.evidence_blocks[i]`。
|
||||||
|
- `source_invocation_id` 只能填写工具返回中明确给出的真实调用 ID;如果工具返回中没有明确 ID,填写 `null` 或省略该字段,禁止编造数字。系统只会在唯一候选工具调用存在时补齐 ID,但不会补齐 `raw_path`。
|
||||||
|
- 不要再输出 `source_invocation_ids` 作为主要字段;兼容旧字段不作为精确证据引用。
|
||||||
- 如果没有任何可确认事实,`claims` 返回空数组,并在 `missing_info` 说明缺少什么。
|
- 如果没有任何可确认事实,`claims` 返回空数组,并在 `missing_info` 说明缺少什么。
|
||||||
- 不要把其它服务、其它历史案例、其它会话的事实迁移为当前会话事实。
|
- 不要把其它服务、其它历史案例、其它会话的事实迁移为当前会话事实。
|
||||||
|
|||||||
@@ -12,7 +12,7 @@
|
|||||||
- `executor_final_answer`:Executor 原始输出,仅用于 debug/fallback;当结构化输出有效时,不得从这里抽取额外确认事实
|
- `executor_final_answer`:Executor 原始输出,仅用于 debug/fallback;当结构化输出有效时,不得从这里抽取额外确认事实
|
||||||
- `executor_structured_output`:如果 Executor 输出了合法证据归因 JSON,这里会提供解析后的对象。结构包含 `claims`、`hypotheses`、`recommended_actions`、`missing_info`;兼容旧版时可能包含 `user_facing_answer`
|
- `executor_structured_output`:如果 Executor 输出了合法证据归因 JSON,这里会提供解析后的对象。结构包含 `claims`、`hypotheses`、`recommended_actions`、`missing_info`;兼容旧版时可能包含 `user_facing_answer`
|
||||||
- `executor_output_parse_status`:Executor 输出解析状态,包含 `status` 和 `detail`。`status` 可能是 `valid` / `missing` / `malformed`
|
- `executor_output_parse_status`:Executor 输出解析状态,包含 `status` 和 `detail`。`status` 可能是 `valid` / `missing` / `malformed`
|
||||||
- `tool_trace_summary`:基于真实工具调用整理出的证据索引。每一项都带有:
|
- `tool_trace_summary`:基于真实工具调用整理出的全局导航和审计索引。它不是唯一证据源;当 claim 有已核验的 `evidence_bindings[].evidence_excerpt` 时,应优先使用 claim-local excerpt 判断可推导性。每一项都带有:
|
||||||
- `trace_ref`
|
- `trace_ref`
|
||||||
- `tool_name`
|
- `tool_name`
|
||||||
- `topic_domain`
|
- `topic_domain`
|
||||||
@@ -20,7 +20,7 @@
|
|||||||
- `input_summary`
|
- `input_summary`
|
||||||
- `output_summary`
|
- `output_summary`
|
||||||
- `evidence_level`
|
- `evidence_level`
|
||||||
- `gatekeeper_result`:Executor 结构化输出的确定性校验结果,包含 `status`、`failed_rules`、`warnings`、`errors`
|
- `gatekeeper_result`:Executor 结构化输出的确定性校验结果,包含 `status`、`severity`、`checked_bindings`、`failed_rules`、`warnings`、`errors`
|
||||||
- `retry_context`:第二轮可选输入;若为空,按首轮处理
|
- `retry_context`:第二轮可选输入;若为空,按首轮处理
|
||||||
|
|
||||||
## 任务步骤
|
## 任务步骤
|
||||||
@@ -29,7 +29,8 @@
|
|||||||
如果 `executor_output_parse_status.status="valid"` 且 `executor_structured_output.claims` 存在:
|
如果 `executor_output_parse_status.status="valid"` 且 `executor_structured_output.claims` 存在:
|
||||||
- 优先逐条校验 `executor_structured_output.claims`
|
- 优先逐条校验 `executor_structured_output.claims`
|
||||||
- 每个 claim 至少形成一条 `claim_checks`
|
- 每个 claim 至少形成一条 `claim_checks`
|
||||||
- 必须检查 claim 的 `evidence_bindings` 是否能对应到 `tool_trace_summary` 中真实存在的 trace、tool 或 source_invocation_ids
|
- 如果 `gatekeeper_result.severity="none"`,将 claim 的 `evidence_bindings[].evidence_excerpt` 视为已通过代码核验的主证据,判断 `claim_text` 是否能由这些 excerpt 推出
|
||||||
|
- `tool_trace_summary` 只用于理解工具调用全貌、补充 trace_ref、识别 no_evidence gap,不要求它逐字包含 excerpt 中已经核验过的全部事实
|
||||||
- 不得从 `executor_final_answer` 中抽取不在 claims 里的额外确认事实
|
- 不得从 `executor_final_answer` 中抽取不在 claims 里的额外确认事实
|
||||||
|
|
||||||
如果 structured output 缺失或 malformed:
|
如果 structured output 缺失或 malformed:
|
||||||
@@ -58,8 +59,8 @@
|
|||||||
- `contradicted`
|
- `contradicted`
|
||||||
|
|
||||||
结构化 claim 的校验规则:
|
结构化 claim 的校验规则:
|
||||||
- claim 有真实 evidence binding,且工具摘要直接包含该事实 → `direct_observation`
|
- claim 有 Gatekeeper 核验通过的 evidence binding,且 `evidence_excerpt` 直接包含该事实 → `direct_observation`
|
||||||
- claim 有真实 evidence binding,工具摘要没有逐字说明但可以合理推出 → `reasonable_inference`
|
- claim 有 Gatekeeper 核验通过的 evidence binding,excerpt 没有逐字说明但可以合理推出 → `reasonable_inference`
|
||||||
- claim 有部分依据,但写成唯一根因、确认根因或说得过满 → `overstated`
|
- claim 有部分依据,但写成唯一根因、确认根因或说得过满 → `overstated`
|
||||||
- claim 无法绑定真实 trace、invocation 或 excerpt → `unsupported`
|
- claim 无法绑定真实 trace、invocation 或 excerpt → `unsupported`
|
||||||
- claim 引入证据外的新服务名、订单号、错误码、指标值、根因 → `external_unknown`
|
- claim 引入证据外的新服务名、订单号、错误码、指标值、根因 → `external_unknown`
|
||||||
@@ -86,8 +87,10 @@
|
|||||||
严格使用以下判定矩阵:
|
严格使用以下判定矩阵:
|
||||||
0. 若 `gatekeeper_result.status="fail"`
|
0. 若 `gatekeeper_result.status="fail"`
|
||||||
- 不得输出 `PASS`
|
- 不得输出 `PASS`
|
||||||
- 若 `failed_rules` 包含 `evidence.invocation_ref`,倾向 `REJECT`
|
- 若 `gatekeeper_result.severity="reject"`,输出 `REJECT`
|
||||||
- 否则至少输出 `LOW_CONFID`
|
- 若 `gatekeeper_result.severity="low_confid"`,输出 `LOW_CONFID`
|
||||||
|
- 兼容旧输入:若缺少 `severity` 且 `failed_rules` 包含 `evidence.invocation_ref`,倾向 `REJECT`
|
||||||
|
- 兼容旧输入:若缺少 `severity` 且不是明显伪造,至少输出 `LOW_CONFID`
|
||||||
|
|
||||||
1. 若任一关键 claim 为 `contradicted`
|
1. 若任一关键 claim 为 `contradicted`
|
||||||
- `verdict = "REJECT"`
|
- `verdict = "REJECT"`
|
||||||
|
|||||||
@@ -0,0 +1,48 @@
|
|||||||
|
package com.superbiz.agent.agent.tool;
|
||||||
|
|
||||||
|
import com.fasterxml.jackson.databind.JsonNode;
|
||||||
|
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||||
|
import com.superbiz.agent.service.ToolInvocationRecorder;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.springframework.test.util.ReflectionTestUtils;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
import static org.mockito.Mockito.mock;
|
||||||
|
|
||||||
|
class QueryLogsToolsTest {
|
||||||
|
|
||||||
|
private final ObjectMapper objectMapper = new ObjectMapper();
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void queryLogsReturnsHikariPositiveMockForOrderService() throws Exception {
|
||||||
|
QueryLogsTools tools = new QueryLogsTools(mock(ToolInvocationRecorder.class));
|
||||||
|
ReflectionTestUtils.setField(tools, "mockEnabled", true);
|
||||||
|
|
||||||
|
String output = tools.queryLogs("ap-guangzhou", "application-logs",
|
||||||
|
"order-service HikariCP connection pool active=50/50 waiting", 10);
|
||||||
|
|
||||||
|
JsonNode root = objectMapper.readTree(output);
|
||||||
|
assertTrue(root.path("success").asBoolean());
|
||||||
|
assertEquals(2, root.path("logs").size());
|
||||||
|
assertEquals("order-service", root.path("logs").get(0).path("service").asText());
|
||||||
|
assertTrue(root.toString().contains("HikariPool-1"));
|
||||||
|
assertFalse(root.toString().contains("generic-service"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void queryLogsReturnsEmptyNoHitForOtherServiceHikariQuery() throws Exception {
|
||||||
|
QueryLogsTools tools = new QueryLogsTools(mock(ToolInvocationRecorder.class));
|
||||||
|
ReflectionTestUtils.setField(tools, "mockEnabled", true);
|
||||||
|
|
||||||
|
String output = tools.queryLogs("ap-guangzhou", "application-logs",
|
||||||
|
"inventory-service HikariCP connection pool active=50/50 waiting", 10);
|
||||||
|
|
||||||
|
JsonNode root = objectMapper.readTree(output);
|
||||||
|
assertFalse(root.path("success").asBoolean());
|
||||||
|
assertEquals(0, root.path("logs").size());
|
||||||
|
assertEquals(0, root.path("total").asInt());
|
||||||
|
assertFalse(root.toString().contains("generic-service"));
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -100,14 +100,14 @@ class DiagnosisEvalBaselineDiffTest {
|
|||||||
}
|
}
|
||||||
|
|
||||||
private void degradeRedisCase(DiagnosisEvalReport report) {
|
private void degradeRedisCase(DiagnosisEvalReport report) {
|
||||||
report.setPassedCases(4);
|
report.setPassedCases(7);
|
||||||
report.setPassRate(0.8);
|
report.setPassRate(0.875);
|
||||||
report.setAverageToolCallCount(3.0);
|
report.setAverageToolCallCount(3.0);
|
||||||
report.setAverageDurationMs(45800.0);
|
report.setAverageDurationMs(44875.0);
|
||||||
report.setVerdictDistribution(new LinkedHashMap<>());
|
report.setVerdictDistribution(new LinkedHashMap<>());
|
||||||
report.getVerdictDistribution().put("PASS", 2L);
|
report.getVerdictDistribution().put("PASS", 2L);
|
||||||
report.getVerdictDistribution().put("LOW_CONFID", 2L);
|
report.getVerdictDistribution().put("LOW_CONFID", 4L);
|
||||||
report.getVerdictDistribution().put("REJECT", 1L);
|
report.getVerdictDistribution().put("REJECT", 2L);
|
||||||
|
|
||||||
DiagnosisEvalResult redis = result(report, "redis-timeout");
|
DiagnosisEvalResult redis = result(report, "redis-timeout");
|
||||||
redis.setPassed(false);
|
redis.setPassed(false);
|
||||||
|
|||||||
@@ -24,11 +24,12 @@ class DiagnosisTraceEvaluatorTest {
|
|||||||
|
|
||||||
DiagnosisEvalReport report = evaluator.evaluate(cases, Path.of("mvp/eval/fixtures"));
|
DiagnosisEvalReport report = evaluator.evaluate(cases, Path.of("mvp/eval/fixtures"));
|
||||||
|
|
||||||
assertEquals(5, report.getTotalCases());
|
assertEquals(8, report.getTotalCases());
|
||||||
assertEquals(5, report.getPassedCases());
|
assertEquals(8, report.getPassedCases());
|
||||||
assertEquals(1.0, report.getPassRate(), 0.001);
|
assertEquals(1.0, report.getPassRate(), 0.001);
|
||||||
assertEquals(2L, report.getVerdictDistribution().get("PASS"));
|
assertEquals(2L, report.getVerdictDistribution().get("PASS"));
|
||||||
assertEquals(3L, report.getVerdictDistribution().get("LOW_CONFID"));
|
assertEquals(5L, report.getVerdictDistribution().get("LOW_CONFID"));
|
||||||
|
assertEquals(1L, report.getVerdictDistribution().get("REJECT"));
|
||||||
|
|
||||||
DiagnosisEvalResult payment = result(report, "payment-timeout");
|
DiagnosisEvalResult payment = result(report, "payment-timeout");
|
||||||
assertTrue(payment.isPassed());
|
assertTrue(payment.isPassed());
|
||||||
@@ -39,6 +40,16 @@ class DiagnosisTraceEvaluatorTest {
|
|||||||
DiagnosisEvalResult redis = result(report, "redis-timeout");
|
DiagnosisEvalResult redis = result(report, "redis-timeout");
|
||||||
assertTrue(redis.isPassed());
|
assertTrue(redis.isPassed());
|
||||||
assertTrue(redis.getEvidenceCoverage().get("query_logs"));
|
assertTrue(redis.getEvidenceCoverage().get("query_logs"));
|
||||||
|
|
||||||
|
DiagnosisEvalResult fabricatedInvocation = result(report, "gatekeeper-fabricated-invocation");
|
||||||
|
assertTrue(fabricatedInvocation.isPassed());
|
||||||
|
assertEquals("fail", fabricatedInvocation.getGatekeeperStatus());
|
||||||
|
assertEquals("valid", fabricatedInvocation.getComposerStatus());
|
||||||
|
assertEquals(1, fabricatedInvocation.getClaimCheckCount());
|
||||||
|
|
||||||
|
DiagnosisEvalResult composerFallback = result(report, "composer-fallback-no-raw-json");
|
||||||
|
assertTrue(composerFallback.isPassed());
|
||||||
|
assertEquals("composer_malformed", composerFallback.getComposerStatus());
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
@@ -108,6 +119,111 @@ class DiagnosisTraceEvaluatorTest {
|
|||||||
"executor confirmed claim missing evidence bindings: claim-unsupported"));
|
"executor confirmed claim missing evidence bindings: claim-unsupported"));
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void evaluateFailsWhenGatekeeperFailStillPassesVerifier() {
|
||||||
|
DiagnosisEvalCase evalCase = DiagnosisEvalCase.builder()
|
||||||
|
.id("gatekeeper-pass-leak")
|
||||||
|
.title("Gatekeeper pass leak")
|
||||||
|
.expectedRootCauseKeywords(List.of())
|
||||||
|
.requiredEvidenceTools(List.of())
|
||||||
|
.allowedVerdicts(List.of("PASS", "LOW_CONFID", "REJECT"))
|
||||||
|
.requireV2AuditClosure(true)
|
||||||
|
.build();
|
||||||
|
DiagnosisTraceResponse trace = DiagnosisTraceResponse.builder()
|
||||||
|
.session(DiagnosisTraceResponse.SessionTrace.builder()
|
||||||
|
.answer("安全回答")
|
||||||
|
.selfEvaluation(java.util.Map.of(
|
||||||
|
"verifier_evaluation", java.util.Map.of(
|
||||||
|
"verdict", "PASS",
|
||||||
|
"gatekeeper_result", java.util.Map.of("status", "fail"),
|
||||||
|
"claim_checks", java.util.List.of(java.util.Map.of(
|
||||||
|
"claim_id", "claim-1",
|
||||||
|
"verification", "unsupported",
|
||||||
|
"detail", "evidence ref invalid"
|
||||||
|
)),
|
||||||
|
"composer_output", java.util.Map.of("status", "valid")
|
||||||
|
)))
|
||||||
|
.build())
|
||||||
|
.toolInvocations(List.of())
|
||||||
|
.build();
|
||||||
|
|
||||||
|
DiagnosisEvalResult result = evaluator.evaluate(evalCase, trace);
|
||||||
|
|
||||||
|
assertFalse(result.isPassed());
|
||||||
|
assertTrue(result.getFailedChecks().contains("gatekeeper fail cannot have PASS verdict"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void evaluateFailsWhenUnsupportedClaimLeaksIntoFinalAnswer() {
|
||||||
|
DiagnosisEvalCase evalCase = DiagnosisEvalCase.builder()
|
||||||
|
.id("unsupported-leak")
|
||||||
|
.title("Unsupported leak")
|
||||||
|
.expectedRootCauseKeywords(List.of())
|
||||||
|
.requiredEvidenceTools(List.of())
|
||||||
|
.allowedVerdicts(List.of("LOW_CONFID"))
|
||||||
|
.requireV2AuditClosure(true)
|
||||||
|
.forbiddenConfirmedClaimKeywords(List.of("主库故障"))
|
||||||
|
.build();
|
||||||
|
DiagnosisTraceResponse trace = DiagnosisTraceResponse.builder()
|
||||||
|
.session(DiagnosisTraceResponse.SessionTrace.builder()
|
||||||
|
.answer("已经确认主库故障。")
|
||||||
|
.selfEvaluation(java.util.Map.of(
|
||||||
|
"verifier_evaluation", java.util.Map.of(
|
||||||
|
"verdict", "LOW_CONFID",
|
||||||
|
"gatekeeper_result", java.util.Map.of("status", "pass"),
|
||||||
|
"claim_checks", java.util.List.of(java.util.Map.of(
|
||||||
|
"claim_id", "claim-1",
|
||||||
|
"verification", "unsupported",
|
||||||
|
"detail", "missing database evidence"
|
||||||
|
)),
|
||||||
|
"composer_output", java.util.Map.of("status", "valid")
|
||||||
|
)))
|
||||||
|
.build())
|
||||||
|
.toolInvocations(List.of())
|
||||||
|
.build();
|
||||||
|
|
||||||
|
DiagnosisEvalResult result = evaluator.evaluate(evalCase, trace);
|
||||||
|
|
||||||
|
assertFalse(result.isPassed());
|
||||||
|
assertTrue(result.getFailedChecks().contains(
|
||||||
|
"answer contains forbidden confirmed claim keyword: 主库故障"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void evaluateFailsWhenFinalAnswerLeaksRawExecutorMarker() {
|
||||||
|
DiagnosisEvalCase evalCase = DiagnosisEvalCase.builder()
|
||||||
|
.id("raw-json-leak")
|
||||||
|
.title("Raw json leak")
|
||||||
|
.expectedRootCauseKeywords(List.of())
|
||||||
|
.requiredEvidenceTools(List.of())
|
||||||
|
.allowedVerdicts(List.of("LOW_CONFID"))
|
||||||
|
.requireV2AuditClosure(true)
|
||||||
|
.build();
|
||||||
|
DiagnosisTraceResponse trace = DiagnosisTraceResponse.builder()
|
||||||
|
.session(DiagnosisTraceResponse.SessionTrace.builder()
|
||||||
|
.answer("answer_version=executor_evidence_v2")
|
||||||
|
.selfEvaluation(java.util.Map.of(
|
||||||
|
"verifier_evaluation", java.util.Map.of(
|
||||||
|
"verdict", "LOW_CONFID",
|
||||||
|
"gatekeeper_result", java.util.Map.of("status", "pass"),
|
||||||
|
"claim_checks", java.util.List.of(java.util.Map.of(
|
||||||
|
"claim_id", "claim-1",
|
||||||
|
"verification", "direct_observation",
|
||||||
|
"detail", "log evidence"
|
||||||
|
)),
|
||||||
|
"composer_output", java.util.Map.of("status", "composer_malformed")
|
||||||
|
)))
|
||||||
|
.build())
|
||||||
|
.toolInvocations(List.of())
|
||||||
|
.build();
|
||||||
|
|
||||||
|
DiagnosisEvalResult result = evaluator.evaluate(evalCase, trace);
|
||||||
|
|
||||||
|
assertFalse(result.isPassed());
|
||||||
|
assertTrue(result.getFailedChecks().contains("answer leaks raw executor marker: executor_evidence_v2"));
|
||||||
|
assertTrue(result.getFailedChecks().contains("answer leaks raw executor marker: answer_version"));
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void reportWriterOutputsJsonAndMarkdown(@TempDir Path tempDir) throws Exception {
|
void reportWriterOutputsJsonAndMarkdown(@TempDir Path tempDir) throws Exception {
|
||||||
DiagnosisEvalReport report = evaluator.evaluate(readCases(), Path.of("mvp/eval/fixtures"));
|
DiagnosisEvalReport report = evaluator.evaluate(readCases(), Path.of("mvp/eval/fixtures"));
|
||||||
|
|||||||
@@ -95,7 +95,7 @@ class VerifierInputHookTest {
|
|||||||
));
|
));
|
||||||
ToolInvocationRepository invocationRepository = mock(ToolInvocationRepository.class);
|
ToolInvocationRepository invocationRepository = mock(ToolInvocationRepository.class);
|
||||||
when(invocationRepository.findBySessionIdOrderByIdAsc("structured-v2-session")).thenReturn(List.of(
|
when(invocationRepository.findBySessionIdOrderByIdAsc("structured-v2-session")).thenReturn(List.of(
|
||||||
ToolInvocation.builder().id(101L).sessionId("structured-v2-session").toolName("query_metrics").build()
|
invocation(101L, "structured-v2-session", "query_metrics", "$.alerts[0]", "active=50 max=50")
|
||||||
));
|
));
|
||||||
VerifierInputHook hook = new VerifierInputHook(traceSummaryService,
|
VerifierInputHook hook = new VerifierInputHook(traceSummaryService,
|
||||||
new ExecutorGatekeeperService(invocationRepository));
|
new ExecutorGatekeeperService(invocationRepository));
|
||||||
@@ -115,7 +115,8 @@ class VerifierInputHookTest {
|
|||||||
"source_type": "tool_trace",
|
"source_type": "tool_trace",
|
||||||
"source_id": "trace-1",
|
"source_id": "trace-1",
|
||||||
"tool_name": "query_metrics",
|
"tool_name": "query_metrics",
|
||||||
"source_invocation_ids": [101],
|
"source_invocation_id": 101,
|
||||||
|
"raw_path": "$.alerts[0]",
|
||||||
"evidence_excerpt": "active=50 max=50"
|
"evidence_excerpt": "active=50 max=50"
|
||||||
}
|
}
|
||||||
]
|
]
|
||||||
@@ -140,16 +141,96 @@ class VerifierInputHookTest {
|
|||||||
assertEquals("连接池 active 达到上限",
|
assertEquals("连接池 active 达到上限",
|
||||||
payload.path("executor_structured_output").path("claims").get(0).path("claim_text").asText());
|
payload.path("executor_structured_output").path("claims").get(0).path("claim_text").asText());
|
||||||
assertEquals("pass", payload.path("gatekeeper_result").path("status").asText());
|
assertEquals("pass", payload.path("gatekeeper_result").path("status").asText());
|
||||||
|
assertEquals("none", payload.path("gatekeeper_result").path("severity").asText());
|
||||||
assertEquals("pass", VerifierContextHolder.getGatekeeperResult().get("status"));
|
assertEquals("pass", VerifierContextHolder.getGatekeeperResult().get("status"));
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void beforeModelBackfillsOnlyUniqueInvocationIdAndDoesNotPassWithoutRawPath() throws Exception {
|
||||||
|
ToolTraceSummaryService traceSummaryService = mock(ToolTraceSummaryService.class);
|
||||||
|
when(traceSummaryService.buildVerifierTraceSummary(anyString(), anyString())).thenReturn(List.of(
|
||||||
|
Map.of(
|
||||||
|
"trace_ref", "metrics-1",
|
||||||
|
"tool_name", "query_metrics",
|
||||||
|
"source_invocation_ids", List.of(101L)
|
||||||
|
)
|
||||||
|
));
|
||||||
|
ToolInvocationRepository invocationRepository = mock(ToolInvocationRepository.class);
|
||||||
|
when(invocationRepository.findBySessionIdOrderByIdAsc("backfill-session")).thenReturn(List.of(
|
||||||
|
invocation(101L, "backfill-session", "query_metrics", "$.alerts[0]",
|
||||||
|
"CPU 使用率持续超过 80%,当前值为 92%")
|
||||||
|
));
|
||||||
|
VerifierInputHook hook = new VerifierInputHook(traceSummaryService,
|
||||||
|
new ExecutorGatekeeperService(invocationRepository));
|
||||||
|
|
||||||
|
String executorOutput = """
|
||||||
|
{
|
||||||
|
"answer_version": "executor_evidence_v2",
|
||||||
|
"claims": [
|
||||||
|
{
|
||||||
|
"claim_id": "claim-1",
|
||||||
|
"claim_type": "symptom",
|
||||||
|
"claim_text": "payment-service CPU 使用率超过 92%",
|
||||||
|
"support_level": "direct",
|
||||||
|
"evidence_bindings": [
|
||||||
|
{
|
||||||
|
"source_type": "tool_trace",
|
||||||
|
"source_id": "prometheus-alert-HighCPUUsage",
|
||||||
|
"tool_name": "queryPrometheusAlerts",
|
||||||
|
"evidence_excerpt": "CPU 使用率持续超过 80%,当前值为 92%"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"hypotheses": [],
|
||||||
|
"recommended_actions": [
|
||||||
|
{
|
||||||
|
"action_text": "restart payment-service",
|
||||||
|
"reason": "cpu alert is firing",
|
||||||
|
"evidence_bindings": [
|
||||||
|
{
|
||||||
|
"source_type": "tool_trace",
|
||||||
|
"source_id": "prometheus-alert-HighCPUUsage",
|
||||||
|
"tool_name": "queryPrometheusAlerts",
|
||||||
|
"evidence_excerpt": "CPU usage is 92%"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"missing_info": []
|
||||||
|
}
|
||||||
|
""";
|
||||||
|
|
||||||
|
AgentCommand command = hook.beforeModel(
|
||||||
|
List.of(new AssistantMessage(executorOutput)),
|
||||||
|
RunnableConfig.builder().addMetadata("sessionId", "backfill-session").build()
|
||||||
|
);
|
||||||
|
|
||||||
|
JsonNode payload = readPayload(command);
|
||||||
|
JsonNode binding = payload.path("executor_structured_output")
|
||||||
|
.path("claims").get(0)
|
||||||
|
.path("evidence_bindings").get(0);
|
||||||
|
assertEquals("query_metrics", binding.path("tool_name").asText());
|
||||||
|
assertEquals(101L, binding.path("source_invocation_id").asLong());
|
||||||
|
JsonNode actionBinding = payload.path("executor_structured_output")
|
||||||
|
.path("recommended_actions").get(0)
|
||||||
|
.path("evidence_bindings").get(0);
|
||||||
|
assertEquals("query_metrics", actionBinding.path("tool_name").asText());
|
||||||
|
assertEquals(101L, actionBinding.path("source_invocation_id").asLong());
|
||||||
|
assertFalse(binding.has("raw_path"));
|
||||||
|
assertEquals("fail", payload.path("gatekeeper_result").path("status").asText());
|
||||||
|
assertEquals("low_confid", payload.path("gatekeeper_result").path("severity").asText());
|
||||||
|
assertEquals("evidence.invocation_auto_backfill",
|
||||||
|
payload.path("gatekeeper_result").path("warnings").get(0).path("rule").asText());
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void beforeModelAddsFailingGatekeeperResultForFabricatedInvocationId() throws Exception {
|
void beforeModelAddsFailingGatekeeperResultForFabricatedInvocationId() throws Exception {
|
||||||
ToolTraceSummaryService traceSummaryService = mock(ToolTraceSummaryService.class);
|
ToolTraceSummaryService traceSummaryService = mock(ToolTraceSummaryService.class);
|
||||||
when(traceSummaryService.buildVerifierTraceSummary(anyString(), anyString())).thenReturn(List.of());
|
when(traceSummaryService.buildVerifierTraceSummary(anyString(), anyString())).thenReturn(List.of());
|
||||||
ToolInvocationRepository invocationRepository = mock(ToolInvocationRepository.class);
|
ToolInvocationRepository invocationRepository = mock(ToolInvocationRepository.class);
|
||||||
when(invocationRepository.findBySessionIdOrderByIdAsc("fabricated-invocation-session")).thenReturn(List.of(
|
when(invocationRepository.findBySessionIdOrderByIdAsc("fabricated-invocation-session")).thenReturn(List.of(
|
||||||
ToolInvocation.builder().id(101L).sessionId("fabricated-invocation-session").toolName("query_metrics").build()
|
invocation(101L, "fabricated-invocation-session", "query_metrics", "$.alerts[0]", "active=50 max=50")
|
||||||
));
|
));
|
||||||
VerifierInputHook hook = new VerifierInputHook(traceSummaryService,
|
VerifierInputHook hook = new VerifierInputHook(traceSummaryService,
|
||||||
new ExecutorGatekeeperService(invocationRepository));
|
new ExecutorGatekeeperService(invocationRepository));
|
||||||
@@ -167,7 +248,8 @@ class VerifierInputHookTest {
|
|||||||
{
|
{
|
||||||
"source_type": "tool_trace",
|
"source_type": "tool_trace",
|
||||||
"tool_name": "query_metrics",
|
"tool_name": "query_metrics",
|
||||||
"source_invocation_ids": [999],
|
"source_invocation_id": 999,
|
||||||
|
"raw_path": "$.alerts[0]",
|
||||||
"evidence_excerpt": "active=50 max=50"
|
"evidence_excerpt": "active=50 max=50"
|
||||||
}
|
}
|
||||||
]
|
]
|
||||||
@@ -186,6 +268,7 @@ class VerifierInputHookTest {
|
|||||||
|
|
||||||
JsonNode payload = readPayload(command);
|
JsonNode payload = readPayload(command);
|
||||||
assertEquals("fail", payload.path("gatekeeper_result").path("status").asText());
|
assertEquals("fail", payload.path("gatekeeper_result").path("status").asText());
|
||||||
|
assertEquals("reject", payload.path("gatekeeper_result").path("severity").asText());
|
||||||
assertEquals("evidence.invocation_ref",
|
assertEquals("evidence.invocation_ref",
|
||||||
payload.path("gatekeeper_result").path("failed_rules").get(0).asText());
|
payload.path("gatekeeper_result").path("failed_rules").get(0).asText());
|
||||||
}
|
}
|
||||||
@@ -284,4 +367,14 @@ class VerifierInputHookTest {
|
|||||||
assertTrue(message instanceof UserMessage);
|
assertTrue(message instanceof UserMessage);
|
||||||
return objectMapper.readTree(((UserMessage) message).getText());
|
return objectMapper.readTree(((UserMessage) message).getText());
|
||||||
}
|
}
|
||||||
|
|
||||||
|
private ToolInvocation invocation(Long id, String sessionId, String toolName, String rawPath, String text) {
|
||||||
|
return ToolInvocation.builder()
|
||||||
|
.id(id)
|
||||||
|
.sessionId(sessionId)
|
||||||
|
.toolName(toolName)
|
||||||
|
.retrievalDetails("{\"evidence_refs\":[{\"raw_path\":\"" + rawPath
|
||||||
|
+ "\",\"text\":\"" + text + "\"}]}")
|
||||||
|
.build();
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -253,7 +253,8 @@ class ChatServiceSequentialAgentTest {
|
|||||||
"source_type": "tool_trace",
|
"source_type": "tool_trace",
|
||||||
"source_id": "trace-1",
|
"source_id": "trace-1",
|
||||||
"tool_name": "query_metrics",
|
"tool_name": "query_metrics",
|
||||||
"source_invocation_ids": [101],
|
"source_invocation_id": 101,
|
||||||
|
"raw_path": "$.alerts[0]",
|
||||||
"evidence_excerpt": "active=50 max=50"
|
"evidence_excerpt": "active=50 max=50"
|
||||||
}
|
}
|
||||||
]
|
]
|
||||||
@@ -300,7 +301,8 @@ class ChatServiceSequentialAgentTest {
|
|||||||
"source_type": "tool_trace",
|
"source_type": "tool_trace",
|
||||||
"source_id": "trace-1",
|
"source_id": "trace-1",
|
||||||
"tool_name": "query_metrics",
|
"tool_name": "query_metrics",
|
||||||
"source_invocation_ids": [101],
|
"source_invocation_id": 101,
|
||||||
|
"raw_path": "$.alerts[0]",
|
||||||
"evidence_excerpt": "active=50 max=50"
|
"evidence_excerpt": "active=50 max=50"
|
||||||
}
|
}
|
||||||
]
|
]
|
||||||
@@ -350,6 +352,7 @@ class ChatServiceSequentialAgentTest {
|
|||||||
.id(101L)
|
.id(101L)
|
||||||
.sessionId("sequential-gatekeeper-persist-session")
|
.sessionId("sequential-gatekeeper-persist-session")
|
||||||
.toolName("query_metrics")
|
.toolName("query_metrics")
|
||||||
|
.retrievalDetails(evidenceRefs("$.alerts[0]", "active=50 max=50"))
|
||||||
.build()));
|
.build()));
|
||||||
ScriptedChatModel chatModel = new ScriptedChatModel();
|
ScriptedChatModel chatModel = new ScriptedChatModel();
|
||||||
chatModel.executorOutput = """
|
chatModel.executorOutput = """
|
||||||
@@ -366,7 +369,8 @@ class ChatServiceSequentialAgentTest {
|
|||||||
"source_type": "tool_trace",
|
"source_type": "tool_trace",
|
||||||
"source_id": "trace-1",
|
"source_id": "trace-1",
|
||||||
"tool_name": "query_metrics",
|
"tool_name": "query_metrics",
|
||||||
"source_invocation_ids": [101],
|
"source_invocation_id": 101,
|
||||||
|
"raw_path": "$.alerts[0]",
|
||||||
"evidence_excerpt": "active=50 max=50"
|
"evidence_excerpt": "active=50 max=50"
|
||||||
}
|
}
|
||||||
]
|
]
|
||||||
@@ -393,6 +397,7 @@ class ChatServiceSequentialAgentTest {
|
|||||||
@SuppressWarnings("unchecked")
|
@SuppressWarnings("unchecked")
|
||||||
Map<String, Object> gatekeeperResult = (Map<String, Object>) verifierEvaluation.get("gatekeeper_result");
|
Map<String, Object> gatekeeperResult = (Map<String, Object>) verifierEvaluation.get("gatekeeper_result");
|
||||||
assertEquals("pass", gatekeeperResult.get("status"));
|
assertEquals("pass", gatekeeperResult.get("status"));
|
||||||
|
assertEquals("none", gatekeeperResult.get("severity"));
|
||||||
}
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
@@ -407,6 +412,7 @@ class ChatServiceSequentialAgentTest {
|
|||||||
.id(101L)
|
.id(101L)
|
||||||
.sessionId("sequential-claim-check-session")
|
.sessionId("sequential-claim-check-session")
|
||||||
.toolName("query_metrics")
|
.toolName("query_metrics")
|
||||||
|
.retrievalDetails(evidenceRefs("$.alerts[0]", "active=50 max=50"))
|
||||||
.build()));
|
.build()));
|
||||||
ScriptedChatModel chatModel = new ScriptedChatModel("""
|
ScriptedChatModel chatModel = new ScriptedChatModel("""
|
||||||
{
|
{
|
||||||
@@ -466,6 +472,7 @@ class ChatServiceSequentialAgentTest {
|
|||||||
.id(101L)
|
.id(101L)
|
||||||
.sessionId("sequential-gatekeeper-fail-session")
|
.sessionId("sequential-gatekeeper-fail-session")
|
||||||
.toolName("query_metrics")
|
.toolName("query_metrics")
|
||||||
|
.retrievalDetails(evidenceRefs("$.alerts[0]", "active=50 max=50"))
|
||||||
.build()));
|
.build()));
|
||||||
ScriptedChatModel chatModel = new ScriptedChatModel("""
|
ScriptedChatModel chatModel = new ScriptedChatModel("""
|
||||||
{
|
{
|
||||||
@@ -493,7 +500,8 @@ class ChatServiceSequentialAgentTest {
|
|||||||
{
|
{
|
||||||
"source_type": "tool_trace",
|
"source_type": "tool_trace",
|
||||||
"tool_name": "query_metrics",
|
"tool_name": "query_metrics",
|
||||||
"source_invocation_ids": [999],
|
"source_invocation_id": 999,
|
||||||
|
"raw_path": "$.alerts[0]",
|
||||||
"evidence_excerpt": "active=50 max=50"
|
"evidence_excerpt": "active=50 max=50"
|
||||||
}
|
}
|
||||||
]
|
]
|
||||||
@@ -647,6 +655,7 @@ class ChatServiceSequentialAgentTest {
|
|||||||
when(toolInvocationRepository.findBySessionIdOrderByIdAsc(anyString())).thenReturn(List.of(ToolInvocation.builder()
|
when(toolInvocationRepository.findBySessionIdOrderByIdAsc(anyString())).thenReturn(List.of(ToolInvocation.builder()
|
||||||
.id(101L)
|
.id(101L)
|
||||||
.toolName("query_metrics")
|
.toolName("query_metrics")
|
||||||
|
.retrievalDetails(evidenceRefs("$.alerts[0]", "active=50 max=50"))
|
||||||
.build()));
|
.build()));
|
||||||
|
|
||||||
EvaluationService evaluationService = mock(EvaluationService.class);
|
EvaluationService evaluationService = mock(EvaluationService.class);
|
||||||
@@ -694,7 +703,8 @@ class ChatServiceSequentialAgentTest {
|
|||||||
"source_type": "tool_trace",
|
"source_type": "tool_trace",
|
||||||
"source_id": "trace-1",
|
"source_id": "trace-1",
|
||||||
"tool_name": "query_metrics",
|
"tool_name": "query_metrics",
|
||||||
"source_invocation_ids": [101],
|
"source_invocation_id": 101,
|
||||||
|
"raw_path": "$.alerts[0]",
|
||||||
"evidence_excerpt": "active=50 max=50"
|
"evidence_excerpt": "active=50 max=50"
|
||||||
}
|
}
|
||||||
]
|
]
|
||||||
@@ -707,6 +717,10 @@ class ChatServiceSequentialAgentTest {
|
|||||||
""";
|
""";
|
||||||
}
|
}
|
||||||
|
|
||||||
|
private String evidenceRefs(String rawPath, String text) {
|
||||||
|
return "{\"evidence_refs\":[{\"raw_path\":\"" + rawPath + "\",\"text\":\"" + text + "\"}]}";
|
||||||
|
}
|
||||||
|
|
||||||
private static final class ScriptedChatModel implements ChatModel {
|
private static final class ScriptedChatModel implements ChatModel {
|
||||||
private final java.util.ArrayList<String> agentCalls = new java.util.ArrayList<>();
|
private final java.util.ArrayList<String> agentCalls = new java.util.ArrayList<>();
|
||||||
private String promptText = "";
|
private String promptText = "";
|
||||||
@@ -728,7 +742,8 @@ class ChatServiceSequentialAgentTest {
|
|||||||
"source_type": "tool_trace",
|
"source_type": "tool_trace",
|
||||||
"source_id": "trace-1",
|
"source_id": "trace-1",
|
||||||
"tool_name": "query_metrics",
|
"tool_name": "query_metrics",
|
||||||
"source_invocation_ids": [101],
|
"source_invocation_id": 101,
|
||||||
|
"raw_path": "$.alerts[0]",
|
||||||
"evidence_excerpt": "active=50 max=50"
|
"evidence_excerpt": "active=50 max=50"
|
||||||
}
|
}
|
||||||
]
|
]
|
||||||
|
|||||||
@@ -18,14 +18,18 @@ class ExecutorGatekeeperServiceTest {
|
|||||||
void validatePassesForExecutorEvidenceV2WithMatchingInvocation() {
|
void validatePassesForExecutorEvidenceV2WithMatchingInvocation() {
|
||||||
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||||
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
|
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
|
||||||
ToolInvocation.builder().id(101L).sessionId("session-1").toolName("query_metrics").build()
|
invocation(101L, "query_metrics", "$.alerts[0]",
|
||||||
|
"HighCPUUsage firing, service=payment-service, current=92%, duration=25m")
|
||||||
));
|
));
|
||||||
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
|
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
|
||||||
|
|
||||||
Map<String, Object> result = service.validate("session-1", validOutput(101L, "query_metrics"),
|
Map<String, Object> result = service.validate("session-1",
|
||||||
|
validOutput(101L, "query_metrics", "$.alerts[0]",
|
||||||
|
"HighCPUUsage firing, service=payment-service, current=92%"),
|
||||||
Map.of("status", "valid"));
|
Map.of("status", "valid"));
|
||||||
|
|
||||||
assertEquals("pass", result.get("status"));
|
assertEquals("pass", result.get("status"));
|
||||||
|
assertEquals("none", result.get("severity"));
|
||||||
assertTrue(((List<?>) result.get("failed_rules")).isEmpty());
|
assertTrue(((List<?>) result.get("failed_rules")).isEmpty());
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -34,12 +38,13 @@ class ExecutorGatekeeperServiceTest {
|
|||||||
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||||
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of());
|
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of());
|
||||||
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
|
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
|
||||||
Map<String, Object> output = validOutput(101L, "query_metrics");
|
Map<String, Object> output = validOutput(101L, "query_metrics", "$.alerts[0]", "cpu=92");
|
||||||
output.put("user_facing_answer", "旧版最终答案");
|
output.put("user_facing_answer", "旧版最终答案");
|
||||||
|
|
||||||
Map<String, Object> result = service.validate("session-1", output, Map.of("status", "valid"));
|
Map<String, Object> result = service.validate("session-1", output, Map.of("status", "valid"));
|
||||||
|
|
||||||
assertEquals("fail", result.get("status"));
|
assertEquals("fail", result.get("status"));
|
||||||
|
assertEquals("reject", result.get("severity"));
|
||||||
assertTrue(((List<?>) result.get("failed_rules")).contains("schema.executor_v2"));
|
assertTrue(((List<?>) result.get("failed_rules")).contains("schema.executor_v2"));
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -47,14 +52,16 @@ class ExecutorGatekeeperServiceTest {
|
|||||||
void validateFailsForFabricatedInvocationId() {
|
void validateFailsForFabricatedInvocationId() {
|
||||||
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||||
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
|
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
|
||||||
ToolInvocation.builder().id(101L).sessionId("session-1").toolName("query_metrics").build()
|
invocation(101L, "query_metrics", "$.alerts[0]", "cpu=92")
|
||||||
));
|
));
|
||||||
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
|
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
|
||||||
|
|
||||||
Map<String, Object> result = service.validate("session-1", validOutput(999L, "query_metrics"),
|
Map<String, Object> result = service.validate("session-1",
|
||||||
|
validOutput(999L, "query_metrics", "$.alerts[0]", "cpu=92"),
|
||||||
Map.of("status", "valid"));
|
Map.of("status", "valid"));
|
||||||
|
|
||||||
assertEquals("fail", result.get("status"));
|
assertEquals("fail", result.get("status"));
|
||||||
|
assertEquals("reject", result.get("severity"));
|
||||||
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.invocation_ref"));
|
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.invocation_ref"));
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -62,19 +69,137 @@ class ExecutorGatekeeperServiceTest {
|
|||||||
void validateFailsForToolNameMismatch() {
|
void validateFailsForToolNameMismatch() {
|
||||||
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||||
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
|
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
|
||||||
ToolInvocation.builder().id(101L).sessionId("session-1").toolName("query_logs").build()
|
invocation(101L, "query_logs", "$.logs[0]", "cpu=92")
|
||||||
));
|
));
|
||||||
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
|
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
|
||||||
|
|
||||||
Map<String, Object> result = service.validate("session-1", validOutput(101L, "query_metrics"),
|
Map<String, Object> result = service.validate("session-1",
|
||||||
|
validOutput(101L, "query_metrics", "$.alerts[0]", "cpu=92"),
|
||||||
Map.of("status", "valid"));
|
Map.of("status", "valid"));
|
||||||
|
|
||||||
assertEquals("fail", result.get("status"));
|
assertEquals("fail", result.get("status"));
|
||||||
|
assertEquals("reject", result.get("severity"));
|
||||||
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.invocation_ref"));
|
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.invocation_ref"));
|
||||||
}
|
}
|
||||||
|
|
||||||
@SuppressWarnings("unchecked")
|
@Test
|
||||||
private Map<String, Object> validOutput(Long invocationId, String toolName) {
|
void validateDowngradesMissingRawPathToLowConfidence() {
|
||||||
|
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||||
|
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
|
||||||
|
invocation(101L, "query_metrics", "$.alerts[0]", "cpu=92")
|
||||||
|
));
|
||||||
|
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
|
||||||
|
Map<String, Object> output = validOutput(101L, "query_metrics", null, "cpu=92");
|
||||||
|
|
||||||
|
Map<String, Object> result = service.validate("session-1", output, Map.of("status", "valid"));
|
||||||
|
|
||||||
|
assertEquals("fail", result.get("status"));
|
||||||
|
assertEquals("low_confid", result.get("severity"));
|
||||||
|
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.raw_path"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void validateRejectsUnknownRawPath() {
|
||||||
|
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||||
|
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
|
||||||
|
invocation(101L, "query_metrics", "$.alerts[0]", "cpu=92")
|
||||||
|
));
|
||||||
|
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
|
||||||
|
|
||||||
|
Map<String, Object> result = service.validate("session-1",
|
||||||
|
validOutput(101L, "query_metrics", "$.alerts[99]", "cpu=92"),
|
||||||
|
Map.of("status", "valid"));
|
||||||
|
|
||||||
|
assertEquals("fail", result.get("status"));
|
||||||
|
assertEquals("reject", result.get("severity"));
|
||||||
|
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.raw_path"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void validateDowngradesOldInvocationWithoutEvidenceRefsToLowConfidence() {
|
||||||
|
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||||
|
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
|
||||||
|
ToolInvocation.builder().id(101L).sessionId("session-1").toolName("query_metrics")
|
||||||
|
.retrievalDetails("{\"evidence_status\":\"supported\"}")
|
||||||
|
.build()
|
||||||
|
));
|
||||||
|
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
|
||||||
|
|
||||||
|
Map<String, Object> result = service.validate("session-1",
|
||||||
|
validOutput(101L, "query_metrics", "$.alerts[0]", "cpu=92"),
|
||||||
|
Map.of("status", "valid"));
|
||||||
|
|
||||||
|
assertEquals("fail", result.get("status"));
|
||||||
|
assertEquals("low_confid", result.get("severity"));
|
||||||
|
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.missing"));
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void validateRejectsMismatchedExcerpt() {
|
||||||
|
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||||
|
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
|
||||||
|
invocation(101L, "query_metrics", "$.alerts[0]", "HighCPUUsage firing service payment-service current 92")
|
||||||
|
));
|
||||||
|
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
|
||||||
|
|
||||||
|
Map<String, Object> result = service.validate("session-1",
|
||||||
|
validOutput(101L, "query_metrics", "$.alerts[0]", "HikariCP active=50/50 waiting=32"),
|
||||||
|
Map.of("status", "valid"));
|
||||||
|
|
||||||
|
assertEquals("fail", result.get("status"));
|
||||||
|
assertEquals("reject", result.get("severity"));
|
||||||
|
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.excerpt_mismatch"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void validateFailsForRecommendedActionFabricatedInvocationId() {
|
||||||
|
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||||
|
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
|
||||||
|
invocation(101L, "query_metrics", "$.alerts[0]", "cpu=92")
|
||||||
|
));
|
||||||
|
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
|
||||||
|
Map<String, Object> output = validOutput(101L, "query_metrics", "$.alerts[0]", "cpu=92");
|
||||||
|
output.put("recommended_actions", List.of(Map.of(
|
||||||
|
"action_text", "restart service",
|
||||||
|
"reason", "alert is firing",
|
||||||
|
"evidence_bindings", List.of(Map.of(
|
||||||
|
"source_type", "tool_trace",
|
||||||
|
"source_id", "trace-1",
|
||||||
|
"tool_name", "query_metrics",
|
||||||
|
"source_invocation_id", 999L,
|
||||||
|
"raw_path", "$.alerts[0]",
|
||||||
|
"evidence_excerpt", "cpu=92"
|
||||||
|
))
|
||||||
|
)));
|
||||||
|
|
||||||
|
Map<String, Object> result = service.validate("session-1", output, Map.of("status", "valid"));
|
||||||
|
|
||||||
|
assertEquals("fail", result.get("status"));
|
||||||
|
assertEquals("reject", result.get("severity"));
|
||||||
|
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.invocation_ref"));
|
||||||
|
}
|
||||||
|
|
||||||
|
private ToolInvocation invocation(Long id, String toolName, String rawPath, String text) {
|
||||||
|
return ToolInvocation.builder()
|
||||||
|
.id(id)
|
||||||
|
.sessionId("session-1")
|
||||||
|
.toolName(toolName)
|
||||||
|
.retrievalDetails("{\"evidence_refs\":[{\"raw_path\":\"" + rawPath
|
||||||
|
+ "\",\"text\":\"" + text + "\"}]}")
|
||||||
|
.build();
|
||||||
|
}
|
||||||
|
|
||||||
|
private Map<String, Object> validOutput(Long invocationId, String toolName, String rawPath, String excerpt) {
|
||||||
|
Map<String, Object> binding = new java.util.LinkedHashMap<>();
|
||||||
|
binding.put("source_type", "tool_trace");
|
||||||
|
binding.put("source_id", "trace-1");
|
||||||
|
binding.put("tool_name", toolName);
|
||||||
|
binding.put("source_invocation_id", invocationId);
|
||||||
|
if (rawPath != null) {
|
||||||
|
binding.put("raw_path", rawPath);
|
||||||
|
}
|
||||||
|
binding.put("evidence_excerpt", excerpt);
|
||||||
return new java.util.LinkedHashMap<>(Map.of(
|
return new java.util.LinkedHashMap<>(Map.of(
|
||||||
"answer_version", "executor_evidence_v2",
|
"answer_version", "executor_evidence_v2",
|
||||||
"claims", List.of(Map.of(
|
"claims", List.of(Map.of(
|
||||||
@@ -82,13 +207,7 @@ class ExecutorGatekeeperServiceTest {
|
|||||||
"claim_type", "symptom",
|
"claim_type", "symptom",
|
||||||
"claim_text", "连接池 active 达到上限",
|
"claim_text", "连接池 active 达到上限",
|
||||||
"support_level", "direct",
|
"support_level", "direct",
|
||||||
"evidence_bindings", List.of(Map.of(
|
"evidence_bindings", List.of(binding)
|
||||||
"source_type", "tool_trace",
|
|
||||||
"source_id", "trace-1",
|
|
||||||
"tool_name", toolName,
|
|
||||||
"source_invocation_ids", List.of(invocationId),
|
|
||||||
"evidence_excerpt", "active=50 max=50"
|
|
||||||
))
|
|
||||||
)),
|
)),
|
||||||
"hypotheses", List.of(),
|
"hypotheses", List.of(),
|
||||||
"recommended_actions", List.of(),
|
"recommended_actions", List.of(),
|
||||||
|
|||||||
@@ -1,5 +1,6 @@
|
|||||||
package com.superbiz.agent.service;
|
package com.superbiz.agent.service;
|
||||||
|
|
||||||
|
import com.fasterxml.jackson.databind.JsonNode;
|
||||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||||
import com.superbiz.agent.domain.entity.ToolInvocation;
|
import com.superbiz.agent.domain.entity.ToolInvocation;
|
||||||
import com.superbiz.agent.dto.ContextPack;
|
import com.superbiz.agent.dto.ContextPack;
|
||||||
@@ -23,6 +24,8 @@ import static org.mockito.Mockito.when;
|
|||||||
|
|
||||||
class ToolInvocationRecorderTest {
|
class ToolInvocationRecorderTest {
|
||||||
|
|
||||||
|
private final ObjectMapper objectMapper = new ObjectMapper();
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void recordEvidenceToolPreservesNoEvidenceSemantics() {
|
void recordEvidenceToolPreservesNoEvidenceSemantics() {
|
||||||
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||||
@@ -56,6 +59,72 @@ class ToolInvocationRecorderTest {
|
|||||||
assertTrue(saved.getRetrievalDetails().contains("\"retrieved_domains\":[\"application-logs\"]"));
|
assertTrue(saved.getRetrievalDetails().contains("\"retrieved_domains\":[\"application-logs\"]"));
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void recordEvidenceToolExtractsLogEvidenceRefs() throws Exception {
|
||||||
|
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||||
|
when(repository.save(any(ToolInvocation.class))).thenAnswer(invocation -> invocation.getArgument(0));
|
||||||
|
ToolInvocationRecorder recorder = new ToolInvocationRecorder(repository, new ObjectMapper());
|
||||||
|
SessionContextHolder.setSessionId("log-ref-session");
|
||||||
|
|
||||||
|
try {
|
||||||
|
recorder.recordEvidenceTool(
|
||||||
|
"query_logs",
|
||||||
|
Map.of("query", "HikariCP order-service"),
|
||||||
|
"""
|
||||||
|
{"success":true,"logs":[{"timestamp":"2026-07-08 10:00:00","level":"ERROR","service":"order-service","message":"HikariPool-1 - Connection is not available, request timed out after 30000ms","metrics":{"waiting":"32"}}]}
|
||||||
|
""",
|
||||||
|
true,
|
||||||
|
System.currentTimeMillis() - 10,
|
||||||
|
null,
|
||||||
|
"application-logs",
|
||||||
|
ToolInvocationRecorder.EVIDENCE_STATUS_SUPPORTED,
|
||||||
|
Map.of("log_topic", "application-logs")
|
||||||
|
);
|
||||||
|
} finally {
|
||||||
|
SessionContextHolder.clear();
|
||||||
|
}
|
||||||
|
|
||||||
|
ArgumentCaptor<ToolInvocation> captor = ArgumentCaptor.forClass(ToolInvocation.class);
|
||||||
|
verify(repository).save(captor.capture());
|
||||||
|
JsonNode details = objectMapper.readTree(captor.getValue().getRetrievalDetails());
|
||||||
|
|
||||||
|
assertEquals("$.logs[0]", details.path("evidence_refs").get(0).path("raw_path").asText());
|
||||||
|
assertTrue(details.path("evidence_refs").get(0).path("text").asText().contains("HikariPool-1"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void recordEvidenceToolExtractsMetricEvidenceRefs() throws Exception {
|
||||||
|
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||||
|
when(repository.save(any(ToolInvocation.class))).thenAnswer(invocation -> invocation.getArgument(0));
|
||||||
|
ToolInvocationRecorder recorder = new ToolInvocationRecorder(repository, new ObjectMapper());
|
||||||
|
SessionContextHolder.setSessionId("metric-ref-session");
|
||||||
|
|
||||||
|
try {
|
||||||
|
recorder.recordEvidenceTool(
|
||||||
|
"query_metrics",
|
||||||
|
Map.of("query", "active_prometheus_alerts"),
|
||||||
|
"""
|
||||||
|
{"success":true,"alerts":[{"alert_name":"HighMemoryUsage","state":"firing","description":"服务 order-service 当前值为 91%","active_at":"2026-07-08T10:00:00Z","duration":"15m"}]}
|
||||||
|
""",
|
||||||
|
true,
|
||||||
|
System.currentTimeMillis() - 10,
|
||||||
|
null,
|
||||||
|
"prometheus_alerts",
|
||||||
|
ToolInvocationRecorder.EVIDENCE_STATUS_SUPPORTED,
|
||||||
|
Map.of("metric_family", "prometheus_alerts")
|
||||||
|
);
|
||||||
|
} finally {
|
||||||
|
SessionContextHolder.clear();
|
||||||
|
}
|
||||||
|
|
||||||
|
ArgumentCaptor<ToolInvocation> captor = ArgumentCaptor.forClass(ToolInvocation.class);
|
||||||
|
verify(repository).save(captor.capture());
|
||||||
|
JsonNode details = objectMapper.readTree(captor.getValue().getRetrievalDetails());
|
||||||
|
|
||||||
|
assertEquals("$.alerts[0]", details.path("evidence_refs").get(0).path("raw_path").asText());
|
||||||
|
assertTrue(details.path("evidence_refs").get(0).path("text").asText().contains("HighMemoryUsage"));
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void recordLookupKnowledgePreservesRetrievalSpecificFields() {
|
void recordLookupKnowledgePreservesRetrievalSpecificFields() {
|
||||||
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
|
||||||
@@ -113,6 +182,8 @@ class ToolInvocationRecorderTest {
|
|||||||
assertTrue(saved.getRetrievalDetails().contains("\"evidence_candidate_count\":2"));
|
assertTrue(saved.getRetrievalDetails().contains("\"evidence_candidate_count\":2"));
|
||||||
assertTrue(saved.getRetrievalDetails().contains("\"evidence_block_count\":1"));
|
assertTrue(saved.getRetrievalDetails().contains("\"evidence_block_count\":1"));
|
||||||
assertTrue(saved.getRetrievalDetails().contains("\"evidence_blocks\""));
|
assertTrue(saved.getRetrievalDetails().contains("\"evidence_blocks\""));
|
||||||
|
assertTrue(saved.getRetrievalDetails().contains("\"evidence_refs\""));
|
||||||
|
assertTrue(saved.getRetrievalDetails().contains("\"raw_path\":\"$.evidence_blocks[0]\""));
|
||||||
assertTrue(saved.getRetrievalDetails().contains("\"query_transform\""));
|
assertTrue(saved.getRetrievalDetails().contains("\"query_transform\""));
|
||||||
assertTrue(saved.getRetrievalDetails().contains("\"retrieval_trace\""));
|
assertTrue(saved.getRetrievalDetails().contains("\"retrieval_trace\""));
|
||||||
assertTrue(saved.getRetrievalDetails().contains("\"context_pack_summary\""));
|
assertTrue(saved.getRetrievalDetails().contains("\"context_pack_summary\""));
|
||||||
|
|||||||
Reference in New Issue
Block a user