Compare commits

...
3 Commits
59 changed files with 3060 additions and 565 deletions
+1
View File
@@ -10,6 +10,7 @@
| 2026-07-07 | executor-gatekeeper-hook | Chat质量门禁/证据归因 | Gatekeeper, verifier payload, source_invocation_ids, tool_name match, self_evaluation | openspec/changes/archive/2026-07-07-executor-gatekeeper-hook | archived | | 2026-07-07 | executor-gatekeeper-hook | Chat质量门禁/证据归因 | Gatekeeper, verifier payload, source_invocation_ids, tool_name match, self_evaluation | openspec/changes/archive/2026-07-07-executor-gatekeeper-hook | archived |
| 2026-07-07 | executor-verifier-claim-checks | Chat质量门禁/证据归因 | Verifier claim_checks, facts_checked compatibility, effective verdict guardrail, malformed output downgrade | openspec/changes/archive/2026-07-07-executor-verifier-claim-checks | archived | | 2026-07-07 | executor-verifier-claim-checks | Chat质量门禁/证据归因 | Verifier claim_checks, facts_checked compatibility, effective verdict guardrail, malformed output downgrade | openspec/changes/archive/2026-07-07-executor-verifier-claim-checks | archived |
| 2026-07-08 | executor-composer-final-answer | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived | | 2026-07-08 | executor-composer-final-answer | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
| 2026-07-08 | diagnosis-eval-demo-gatekeeper-closure | Agent eval/demo/Gatekeeper | diagnosis eval matrix, stable demo scenarios, Gatekeeper rule set version, audit metadata | openspec/changes/archive/2026-07-08-diagnosis-eval-demo-gatekeeper-closure | archived |
| 2026-07-08 | verifier-evidence-reference-fidelity | Chat质量门禁/证据归因 | evidence_refs, raw_path, Gatekeeper severity, verifier evidence excerpt, HikariCP mock, no_evidence | openspec/changes/archive/2026-07-08-verifier-evidence-reference-fidelity | archived | | 2026-07-08 | verifier-evidence-reference-fidelity | Chat质量门禁/证据归因 | evidence_refs, raw_path, Gatekeeper severity, verifier evidence excerpt, HikariCP mock, no_evidence | openspec/changes/archive/2026-07-08-verifier-evidence-reference-fidelity | archived |
| 2026-07-06 | rag-eval-pipeline-closure | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived | | 2026-07-06 | rag-eval-pipeline-closure | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
| 2026-07-06 | modular-rag-pipeline | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived | | 2026-07-06 | modular-rag-pipeline | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
@@ -0,0 +1,48 @@
# diagnosis-eval-demo-gatekeeper-closure Acceptance
## Static / Structure Verification
- `cmd /c openspec validate diagnosis-eval-demo-gatekeeper-closure --strict`
- Result: passed.
- `cmd /c openspec validate --specs`
- Result: passed, 10 specs passed.
## Script Verification
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,VerifierInputHookTest" test`
- Result: 36 tests, 0 failures, 0 errors.
- `mvn "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest,ToolInvocationRecorderTest,QueryLogsToolsTest" test`
- Result: 61 tests, 0 failures, 0 errors.
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`
- Result after E2E startup fix: 23 tests, 0 failures, 0 errors.
## Live E2E Verification
- Start command: `mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"`.
- Demo command: `powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1`.
- Result: chat, trace, and feedback requests completed successfully.
- Output files:
- `mvp/demo/output/chat-response.json`
- `mvp/demo/output/trace-response.json`
- `mvp/demo/output/feedback-response.json`
- Trace observations:
- `hasVerifierEvaluation=true`
- `gatekeeper_result.rule_set_version=gatekeeper-rules-v1`
## Fixed During Verification
- E2E startup initially failed because Spring could not instantiate `ExecutorGatekeeperService`.
- Root cause: two public constructors and no explicit `@Autowired` constructor.
- Fix: annotate the production constructor with `@Autowired`.
## Residual Risk
- The live payment-timeout path can still produce `LOW_CONFID` because model-generated evidence bindings may omit some explicit `source_invocation_id` values.
- This is not a blocker for this change because deterministic matrix behavior is covered by saved fixtures and baseline evaluation.
- Existing Maven warnings remain: duplicate `spring-boot-starter-test` declaration and Lombok `@Builder` default warnings.
## Archive Status
- Devflow archive artifacts created.
- OpenSpec change archived to `openspec/changes/archive/2026-07-08-diagnosis-eval-demo-gatekeeper-closure`.
- Main specs synced by `cmd /c openspec archive diagnosis-eval-demo-gatekeeper-closure --yes`.
@@ -0,0 +1,33 @@
# diagnosis-eval-demo-gatekeeper-closure Brief
## Background
The Chat evidence pipeline already had Executor V2 structured output, deterministic Gatekeeper validation, Verifier claim checks, and Composer final rendering. The missing piece was an interview-ready acceptance story that made the anti-hallucination behavior easy to demonstrate and regress.
## Goal
Close the next three interview-readiness gaps together:
- diagnosis eval fixture matrix
- stable demo data set
- Gatekeeper rule configuration and audit version
## Scope
- Expand `mvp/eval` with matrix-oriented cases, fixtures, and baseline reports.
- Add stable demo request payloads and scenario documentation.
- Add a lightweight local Gatekeeper rule catalog with `rule_set_version` and rule metadata in `gatekeeper_result`.
- Update architecture, demo, and eval docs to describe the current implementation.
## Non-goals
- No new public HTTP endpoint.
- No new database table.
- No Planner `scope_contract`.
- No Gatekeeper retry loop.
- No remote or dynamic rule execution engine.
## OpenSpec
- Change: `openspec/changes/diagnosis-eval-demo-gatekeeper-closure`
- Interface impact: L2 internal contract change.
@@ -0,0 +1,144 @@
# diagnosis-eval-demo-gatekeeper-closure Decisions
## Clarify
- Entry summary: implement the next three interview-readiness items together: diagnosis eval fixture matrix, stable demo data set, and Gatekeeper rule configuration/audit version.
- Slug: `diagnosis-eval-demo-gatekeeper-closure`
- Devflow scale: `standard`
- Interface impact: expected L2 internal contract change because `gatekeeper_result` audit JSON will gain rule metadata/version fields.
## Context
- `devflow/index.md` used: related entries found for diagnosis eval harness, fixture expansion, MVP demo runbook, Gatekeeper hook, and verifier evidence reference fidelity.
- Relevant glossary:
- Evidence Tools produce incident facts and must be recorded in `tool_invocation`.
- Verifier should not use skills/runbooks as incident evidence.
- `tool_invocation.retrieval_details` is the structured evidence/audit home for tool-specific details.
- Historical constraints that must enter OpenSpec:
- Diagnosis eval is offline and deterministic; no LLM-as-judge.
- Demo assets should be runnable, but fixed regression should use saved fixtures.
- Gatekeeper remains in the Verifier hook path.
- No new database table for Gatekeeper audit; use `self_evaluation.verifier_evaluation.gatekeeper_result`.
- `$.no_evidence` is a query no-hit signal, not proof that a problem is impossible.
## Question Pool
| ID | Dimension | Mode | Question | Status |
|---|---|---|---|---|
| Q1 | Terminology | evidence-driven | What names should this change use for the matrix, demo set, and Gatekeeper rule metadata? | Resolved |
| Q2 | Boundary | evidence-driven | Should this change alter public APIs, database schema, Planner output, or retry behavior? | Resolved |
| Q3 | Acceptance | evidence-driven | Which existing tests and baseline assets define the current acceptance style? | Resolved |
| Q4 | Technical | evidence-driven | Where should Gatekeeper rule metadata live with minimal implementation risk? | Pending code research |
| Q5 | Scope | user-interview | Should the stable demo set be documentation/payloads only, or should it include live E2E scripts for all scenarios? | Confirmed |
## Evidence-driven Conclusions
- Q1 conclusion: use `diagnosis eval matrix`, `stable demo scenarios`, and `Gatekeeper rule set version` as terms.
- Q2 conclusion: keep this as an internal contract change. Do not add public endpoints, tables, Planner `scope_contract`, or Gatekeeper retry.
- Q3 conclusion: existing `DiagnosisTraceEvaluatorTest`, `ExecutorGatekeeperServiceTest`, `VerifierInputHookTest`, `ToolInvocationRecorderTest`, and `mvp/eval/reports` define the current acceptance style.
- Q4 conclusion: Gatekeeper metadata should live behind a small rule catalog loaded by `ExecutorGatekeeperService`; the audit output should include a rule set version and enabled rule metadata summary, without adding tables or remote registry.
## User-interview Confirmations
- Q5 confirmed by resumed objective: complete items 1/2/3 with sm-flow, archive, submit, and run end-to-end if necessary.
- Implementation interpretation: stable demo scenarios will be fixed request payloads and runbook docs plus deterministic fixture-backed eval. Live E2E remains necessary only for at least one main path or where unit/fixture evidence is insufficient.
## OpenSpec Backfill
- Created Draft proposal at `openspec/changes/diagnosis-eval-demo-gatekeeper-closure/proposal.md`.
- Context constraints from historical devflow entries were written into the proposal.
- Scope confirmation and Gatekeeper catalog placement were written into the proposal/design.
## Current Checkpoint
- Discover completed.
- No implementation files changed yet.
## Specify / Alignment
### Cross-artifact Alignment
| Check | Status | Notes |
|---|---|---|
| brief/proposal goals -> proposal | Aligned | Proposal covers eval matrix, stable demo scenarios, and Gatekeeper rule catalog/audit version. |
| proposal scope/constraints -> design | Aligned | Design records offline deterministic eval, fixture-backed demo distinction, local rule catalog, and no new table/API. |
| design decisions -> specs/tasks | Aligned | Specs cover eval matrix, rule set version validation, demo scenarios, and Gatekeeper rule metadata; tasks cover matching implementation slices. |
| specs observable behavior -> tasks | Aligned | Each requirement has an executable task and acceptance check. |
### Interface Impact
- Level: L2 internal contract change.
- Reason: `gatekeeper_result` internal audit JSON gains `rule_set_version` and rule metadata summary. Eval case/result fields may gain optional rule set checks. No public HTTP API, database schema, or external DTO contract changes.
## Audit
Input -> processing -> output chain:
```text
mvp/demo request docs + mvp/eval fixtures
-> DiagnosisTraceEvaluator
-> baseline reports
-> interview/demo evidence
Gatekeeper rule catalog
-> ExecutorGatekeeperService
-> VerifierInputHook / ChatService persisted self_evaluation
-> Trace and eval audit
```
Architecture risk assessment:
1. The change is intentionally internal and should not add new public consumers.
2. Gatekeeper catalog must stay metadata-only; dynamic rule execution would be a different, riskier architecture.
3. Fixture-backed demo scenarios should be documented as deterministic regression artifacts, not live LLM guarantees.
4. Baseline report churn is expected and must be committed with case/fixture changes.
5. No devflow/OpenSpec conflict found.
## Commit Gate
- `cmd /c openspec validate diagnosis-eval-demo-gatekeeper-closure --strict`: passed.
- `cmd /c openspec validate --specs`: passed, 10 specs passed.
- File completeness:
- proposal.md: present.
- design.md: present.
- specs: present for `diagnosis-eval-harness`, `mvp-demo-trace-acceptance`, `chat-verifier-agent`.
- tasks.md: present.
- Consistency:
- Proposal concepts have corresponding design sections.
- Design decisions are reflected in specs/tasks.
- Task acceptance checks are verifiable.
## Current Checkpoint
- Commit completed.
- `.committed` marker created.
## Apply Verification
- Focused verification passed:
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,VerifierInputHookTest" test`
- Result: 36 tests, 0 failures, 0 errors.
- Broader relevant regression passed:
- `mvn "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest,ToolInvocationRecorderTest,QueryLogsToolsTest" test`
- Result: 61 tests, 0 failures, 0 errors.
- E2E startup repro found a Spring bean construction issue:
- Command: `mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"`
- Failure: `ExecutorGatekeeperService` had two public constructors and no annotated constructor, so Spring attempted a no-arg constructor and failed with `No default constructor found`.
- Classification: code deviation from OpenSpec implementation intent, not a spec gap.
- Fix: annotate the production constructor with `@Autowired`.
- Post-fix focused regression passed:
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`
- Result: 23 tests, 0 failures, 0 errors.
- Live E2E passed for demo compatibility:
- Start: `mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"`
- Run: `powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1`
- Result: `/api/chat`, `/api/diagnosis/{sessionId}/trace`, and `/api/feedback` completed successfully.
- Trace summary included `hasVerifierEvaluation=true`.
- Persisted Gatekeeper audit included `rule_set_version=gatekeeper-rules-v1`.
- Residual quality note: the live payment-timeout response remained `LOW_CONFID` because some model-produced evidence bindings still lacked explicit `source_invocation_id`; deterministic PASS/LOW_CONFID/REJECT claims are covered by fixture-backed eval.
## Archive Readiness
- OpenSpec tasks 1-4 completed.
- Verification is recorded in devflow acceptance artifacts.
- Remaining known risk: live LLM output is not deterministic and may still produce LOW_CONFID on the payment-timeout path; this is intentionally documented as demo compatibility, not a fixed PASS guarantee.
@@ -0,0 +1,22 @@
# diagnosis-eval-demo-gatekeeper-closure Evidence
## Code And Artifact Evidence
- Gatekeeper rule metadata lives in `src/main/resources/gatekeeper/gatekeeper-rules.json`.
- `ExecutorGatekeeperService` loads the local catalog, uses configured threshold parameters, and emits `rule_set_version` plus enabled rule metadata.
- `VerifierInputHook` and `ChatService` preserve Gatekeeper audit metadata in fallback/default paths.
- `DiagnosisTraceEvaluator` can optionally validate expected Gatekeeper rule set version.
- `mvp/eval/cases/diagnosis-cases.json` now includes narrow-scope and no-evidence matrix cases.
- `mvp/eval/reports/baseline-report.json` and `.md` were regenerated for the expanded fixed matrix.
- `mvp/demo/evidence-pipeline-scenarios.md` documents live vs fixture-backed demo scenarios.
## Decisions
- Keep this phase internal: no public API, no DB schema, no Planner output change.
- Keep Gatekeeper deterministic Java validation; the catalog is metadata/config only.
- Treat live demo as compatibility evidence and fixture-backed eval as deterministic regression evidence.
- Persist audit under the existing `self_evaluation.verifier_evaluation.gatekeeper_result` structure.
## Runtime Finding
The first Maven E2E startup found a real integration issue: `ExecutorGatekeeperService` had multiple public constructors without an annotated constructor, so Spring could not instantiate the service. The fix was to annotate the production constructor with `@Autowired`.
+14 -12
View File
@@ -1,6 +1,6 @@
# MVP 架构文档 # MVP 架构文档
**更新日期**:2026-07-06 **更新日期**:2026-07-08
这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到: 这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到:
@@ -15,7 +15,8 @@
| [current-mvp-architecture.md](current-mvp-architecture.md) | 当前可运行 MVP 的总体架构、链路、持久化和质量门禁 | | [current-mvp-architecture.md](current-mvp-architecture.md) | 当前可运行 MVP 的总体架构、链路、持久化和质量门禁 |
| [interview-one-pager.md](interview-one-pager.md) | 面试一页式架构讲解,包含总图、亮点、取舍和追问回答 | | [interview-one-pager.md](interview-one-pager.md) | 面试一页式架构讲解,包含总图、亮点、取舍和追问回答 |
| [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat SequentialAgent、AIOps SupervisorAgent、工具边界 | | [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat SequentialAgent、AIOps SupervisorAgent、工具边界 |
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Verifier、评测基线组成的质量门禁 | | [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md) | Chat 证据链路当前数据契约,覆盖 Executor V2、Gatekeeper、Verifier、Composer、`evidence_refs` |
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Gatekeeper、Verifier、Composer、评测基线组成的质量门禁 |
| [rag-architecture.md](rag-architecture.md) | RAG/知识检索新架构,覆盖 L0 hint、VectorStore 主路径、SDK fallback、证据追踪 | | [rag-architecture.md](rag-architecture.md) | RAG/知识检索新架构,覆盖 L0 hint、VectorStore 主路径、SDK fallback、证据追踪 |
| [modular-rag-pipeline.md](modular-rag-pipeline.md) | `lookup_knowledge` 模块化 RAG 落地架构,覆盖 pipeline、fallback、evidence-first contract、trace | | [modular-rag-pipeline.md](modular-rag-pipeline.md) | `lookup_knowledge` 模块化 RAG 落地架构,覆盖 pipeline、fallback、evidence-first contract、trace |
| [rag-eval-closure.md](rag-eval-closure.md) | RAG 评测闭环,覆盖 offline baseline、baseline diff、diagnosis eval 和 live acceptance | | [rag-eval-closure.md](rag-eval-closure.md) | RAG 评测闭环,覆盖 offline baseline、baseline diff、diagnosis eval 和 live acceptance |
@@ -28,19 +29,20 @@
## 当前架构一句话 ## 当前架构一句话
SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,执行过程落到 `diagnosis_session`、`agent_step`、`tool_invocation`,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。 SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,Chat 链路由 Gatekeeper 做引用真实性校验、Verifier 做可推导性判断、Composer 生成最终表达;执行过程落到 `diagnosis_session`、`agent_step`、`tool_invocation`,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
## 阅读顺序 ## 阅读顺序
1. 先读 [current-mvp-architecture.md](current-mvp-architecture.md),理解系统边界和主链路。 1. 先读 [current-mvp-architecture.md](current-mvp-architecture.md),理解系统边界和主链路。
2. 面试前读 [interview-one-pager.md](interview-one-pager.md),准备 2-5 分钟讲解。 2. 面试前读 [interview-one-pager.md](interview-one-pager.md),准备 2-5 分钟讲解。
3. 再读 [agent-orchestration.md](agent-orchestration.md),理解当前 Agent 如何协作。 3. 再读 [agent-orchestration.md](agent-orchestration.md),理解当前 Agent 如何协作。
4. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。 4. 接着读 [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md),理解 Chat 证据链路的数据结构和验真边界。
5. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。 5. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。
6. 继续读 [modular-rag-pipeline.md](modular-rag-pipeline.md),看 `lookup_knowledge` 的模块化落地和 evidence-first contract。 6. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。
7. 再读 [rag-eval-closure.md](rag-eval-closure.md),看 RAG baseline 如何形成质量闭环。 7. 继续读 [modular-rag-pipeline.md](modular-rag-pipeline.md),看 `lookup_knowledge` 的模块化落地和 evidence-first contract。
8. 然后读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。 8. 再读 [rag-eval-closure.md](rag-eval-closure.md),看 RAG baseline 如何形成质量闭环。
9. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。 9. 然后读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
10. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。 10. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
11. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。 11. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
12. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。 12. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
13. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
+35 -15
View File
@@ -1,6 +1,6 @@
# Agent 编排架构 # Agent 编排架构
**更新日期**:2026-07-05 **更新日期**:2026-07-08
**状态**:当前可运行架构 **状态**:当前可运行架构
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md` **参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
@@ -8,7 +8,7 @@
旧版 Agent 架构把系统描述为 Supervisor、Planner、SubAgent、Verifier 的团队协作。当前 MVP 保留这个核心思想,但实现更收敛: 旧版 Agent 架构把系统描述为 Supervisor、Planner、SubAgent、Verifier 的团队协作。当前 MVP 保留这个核心思想,但实现更收敛:
- Chat 链路使用固定顺序工作流:`Planner -> Executor -> Verifier`。 - Chat 链路使用固定顺序工作流:`Planner -> Executor -> Gatekeeper -> Verifier -> Composer`。
- AIOps 链路使用 `SupervisorAgent` 调度 `Planner + Executor`,最终由规则评估器做轻量验证。 - AIOps 链路使用 `SupervisorAgent` 调度 `Planner + Executor`,最终由规则评估器做轻量验证。
- 当前没有拆分 ExternalApiSubAgent、InternalErrorSubAgent、DatabaseSubAgent;这些作为后续演进方向保留。 - 当前没有拆分 ExternalApiSubAgent、InternalErrorSubAgent、DatabaseSubAgent;这些作为后续演进方向保留。
- 证据工具不直接散落在各个 Agent 里,而是通过 Spring AI ToolCallback / `@Tool` 统一暴露。 - 证据工具不直接散落在各个 Agent 里,而是通过 Spring AI ToolCallback / `@Tool` 统一暴露。
@@ -23,9 +23,11 @@ flowchart TB
ChatPlanner --> ChatExecutor["chat_executor"] ChatPlanner --> ChatExecutor["chat_executor"]
ChatExecutor --> ChatTools["evidence tools"] ChatExecutor --> ChatTools["evidence tools"]
ChatTools --> ChatExecutor ChatTools --> ChatExecutor
ChatExecutor --> ChatVerifier["chat_verifier"] ChatExecutor --> ChatGatekeeper["ExecutorGatekeeperService"]
ChatGatekeeper --> ChatVerifier["chat_verifier"]
ChatVerifier --> ChatDecision{"PASS / LOW_CONFID / REJECT"} ChatVerifier --> ChatDecision{"PASS / LOW_CONFID / REJECT"}
ChatDecision --> ChatAnswer["final answer"] ChatDecision --> ChatComposer["chat_composer"]
ChatComposer --> ChatAnswer["final answer"]
end end
subgraph AiOps["AIOps diagnosis"] subgraph AiOps["AIOps diagnosis"]
@@ -49,9 +51,11 @@ flowchart TB
ChatService --> Session ChatService --> Session
ChatPlanner --> Step ChatPlanner --> Step
ChatExecutor --> Step ChatExecutor --> Step
ChatGatekeeper --> SelfEval
ChatVerifier --> Step ChatVerifier --> Step
ChatTools --> Invocation ChatTools --> Invocation
ChatDecision --> SelfEval ChatDecision --> SelfEval
ChatComposer --> Step
AiOpsService --> Session AiOpsService --> Session
AiOpsPlanner --> Step AiOpsPlanner --> Step
@@ -68,9 +72,13 @@ Chat 复杂诊断采用 `SequentialAgent`,顺序固定:
chat_planner chat_planner
-> chat_executor -> chat_executor
-> lookup_knowledge / query_logs / query_metrics / date_time -> lookup_knowledge / query_logs / query_metrics / date_time
-> outputs executor_evidence_v2
-> VerifierInputHook / ExecutorGatekeeperService
-> validates source_invocation_id / raw_path / evidence_excerpt
-> chat_verifier -> chat_verifier
-> reads tool_trace_summary -> judges whether verified evidence can derive claims
-> outputs verifier JSON -> chat_composer
-> writes final user-facing answer
``` ```
关键行为: 关键行为:
@@ -78,8 +86,10 @@ chat_planner
| 角色 | 当前职责 | 输出 | | 角色 | 当前职责 | 输出 |
|---|---|---| |---|---|---|
| `chat_planner` | 拆解问题,注入知识域地图和对话历史,给出排查方向 | `planner_plan` | | `chat_planner` | 拆解问题,注入知识域地图和对话历史,给出排查方向 | `planner_plan` |
| `chat_executor` | 按计划调用证据工具,组合工具返回形成诊断答复 | `executor_feedback` | | `chat_executor` | 按计划调用证据工具,抽取带 `source_invocation_id + raw_path + evidence_excerpt` 的微观事实 | `executor_evidence_v2` |
| `chat_verifier` | 只基于已有证据校验 Executor 答案,不做新检索 | `verifier_output` | | `ExecutorGatekeeperService` | 在 Verifier 前做代码级引用验真,拒绝伪造 ID、错配 raw_path、错配 excerpt | `gatekeeper_result` |
| `chat_verifier` | 只判断已验真 evidence excerpt 是否能推出 claim,不做新检索 | `verifier_output` |
| `chat_composer` | 只表达 Verifier 允许输出的 claims、缺口和建议,生成最终用户答复 | `composer_output` |
Chat 链路最多支持两轮验证: Chat 链路最多支持两轮验证:
@@ -90,7 +100,9 @@ sequenceDiagram
participant P as chat_planner participant P as chat_planner
participant E as chat_executor participant E as chat_executor
participant T as tools participant T as tools
participant G as gatekeeper
participant V as chat_verifier participant V as chat_verifier
participant M as chat_composer
participant S as diagnosis_session participant S as diagnosis_session
C->>P: 原始问题 + history + retry_context C->>P: 原始问题 + history + retry_context
@@ -98,14 +110,18 @@ sequenceDiagram
C->>E: planner_plan + 上下文 C->>E: planner_plan + 上下文
E->>T: 调用证据工具 E->>T: 调用证据工具
T-->>E: 证据结果 T-->>E: 证据结果
E-->>C: executor_feedback E-->>C: executor_evidence_v2
C->>V: executor_final_answer + tool_trace_summary C->>G: executor_structured_output + tool_invocation.evidence_refs
G-->>C: gatekeeper_result
C->>V: executor_structured_output + gatekeeper_result + tool_trace_summary
V-->>C: PASS / LOW_CONFID / REJECT V-->>C: PASS / LOW_CONFID / REJECT
C->>S: 写入 verifier_evaluation C->>S: 写入 verifier_evaluation
alt LOW_CONFID 且允许补证据 alt LOW_CONFID 且允许补证据
C->>P: retry_context: 仅补缺失证据 C->>P: retry_context: 仅补缺失证据
else PASS 或 REJECT else PASS 或 REJECT
C-->>S: 保存最终 answer C->>M: allowed_claims + missing_info + recommended_actions
M-->>C: composer_output
C->>S: 保存 Composer 最终 answer
end end
``` ```
@@ -113,7 +129,7 @@ sequenceDiagram
| Verdict | 行为 | | Verdict | 行为 |
|---|---| |---|---|
| `PASS` | 输出 Executor 答案 | | `PASS` | 把 Verifier 允许表达的 claims 交给 Composer 输出 |
| `LOW_CONFID` | 如果分数低于阈值且仍有轮次,构造 `retry_context` 补证据;否则输出低置信提示 | | `LOW_CONFID` | 如果分数低于阈值且仍有轮次,构造 `retry_context` 补证据;否则输出低置信提示 |
| `REJECT` | 输出降级答复,只保留已确认信息和下一步建议 | | `REJECT` | 输出降级答复,只保留已确认信息和下一步建议 |
@@ -177,21 +193,25 @@ flowchart LR
SkillBody --> Executor SkillBody --> Executor
Executor --> EvidenceTools["lookup_knowledge / logs / metrics"] Executor --> EvidenceTools["lookup_knowledge / logs / metrics"]
EvidenceTools --> ToolTrace["tool_invocation evidence"] EvidenceTools --> ToolTrace["tool_invocation evidence"]
Executor --> Verifier["Verifier"] Executor --> Gatekeeper["Gatekeeper"]
Gatekeeper --> Verifier["Verifier"]
ToolTrace --> Verifier ToolTrace --> Verifier
Verifier --> Composer["Composer"]
``` ```
| 角色 | Skill 可见性 | 工具权限 | | 角色 | Skill 可见性 | 工具权限 |
|---|---|---| |---|---|---|
| Planner | 只看 skill name / description,并输出 `selected_skill` | 不暴露 `read_skill` | | Planner | 只看 skill name / description,并输出 `selected_skill` | 不暴露 `read_skill` |
| Executor | 读取 Planner 选中的 skill 正文 | 暴露官方 `read_skill` 和证据工具 | | Executor | 读取 Planner 选中的 skill 正文 | 暴露官方 `read_skill` 和证据工具 |
| Verifier | 不看 skill catalog,也不读 skill 正文 | 只读取 `tool_trace_summary` | | Gatekeeper | 不看 skill catalog,也不读 skill 正文 | 只读取 Executor 输出和 `tool_invocation.retrieval_details.evidence_refs` |
| Verifier | 不看 skill catalog,也不读 skill 正文 | 只读取 Gatekeeper 结果、结构化 claims 和 trace summary |
| Composer | 不看 skill catalog,也不读 skill 正文 | 只读取 Verifier 允许表达的内容 |
## 7. 与旧版设计的差异 ## 7. 与旧版设计的差异
| 旧版设想 | 当前实现 | | 旧版设想 | 当前实现 |
|---|---| |---|---|
| Supervisor + Planner + 多个专科 SubAgent + Verifier | Chat: Planner + Executor + Verifier;AIOps: Supervisor + Planner + Executor | | Supervisor + Planner + 多个专科 SubAgent + Verifier | Chat: Planner + Executor + Gatekeeper + Verifier + Composer;AIOps: Supervisor + Planner + Executor |
| ExternalApiSubAgent / InternalErrorSubAgent / DatabaseSubAgent | 暂未拆分,能力通过通用 Executor + 工具 + Prompt 约束实现 | | ExternalApiSubAgent / InternalErrorSubAgent / DatabaseSubAgent | 暂未拆分,能力通过通用 Executor + 工具 + Prompt 约束实现 |
| 每个 SubAgent 专属工具集 | 当前 Executor 持有统一证据工具集合 | | 每个 SubAgent 专属工具集 | 当前 Executor 持有统一证据工具集合 |
| Verifier 支持 PASS / REVISE / REJECT | 当前 Chat Verifier 输出 PASS / LOW_CONFID / REJECT | | Verifier 支持 PASS / REVISE / REJECT | 当前 Chat Verifier 输出 PASS / LOW_CONFID / REJECT |
+29 -7
View File
@@ -1,6 +1,6 @@
# 当前 MVP 架构 # 当前 MVP 架构
**更新日期**:2026-07-05 **更新日期**:2026-07-08
**状态**:当前可运行架构 **状态**:当前可运行架构
**适用范围**:Demo、面试讲解、后续迭代规划 **适用范围**:Demo、面试讲解、后续迭代规划
@@ -38,7 +38,9 @@ flowchart TB
Supervisor["Supervisor"] Supervisor["Supervisor"]
Planner["Planner"] Planner["Planner"]
Executor["Executor"] Executor["Executor"]
Gatekeeper["Gatekeeper"]
Verifier["Verifier"] Verifier["Verifier"]
Composer["Composer"]
end end
subgraph Tools["Evidence Tools"] subgraph Tools["Evidence Tools"]
@@ -105,7 +107,9 @@ Agent Orchestration
-> Supervisor -> Supervisor
-> Planner -> Planner
-> Executor -> Executor
-> Gatekeeper
-> Verifier -> Verifier
-> Composer
Evidence Tools Evidence Tools
-> lookup_knowledge -> lookup_knowledge
@@ -133,6 +137,7 @@ Persistence
-> Milvus/Zilliz collection -> Milvus/Zilliz collection
Quality Gates Quality Gates
-> executor gatekeeper
-> chat verifier -> chat verifier
-> AIOps rule evaluation -> AIOps rule evaluation
-> diagnosis eval baseline -> diagnosis eval baseline
@@ -150,7 +155,9 @@ sequenceDiagram
participant Planner as Planner Agent participant Planner as Planner Agent
participant Executor as Executor Agent participant Executor as Executor Agent
participant Tool as Evidence Tools participant Tool as Evidence Tools
participant Gatekeeper as Gatekeeper Hook
participant Verifier as Verifier Agent participant Verifier as Verifier Agent
participant Composer as Composer Agent
participant DB as Trace Tables participant DB as Trace Tables
participant Trace as Trace API participant Trace as Trace API
@@ -162,8 +169,12 @@ sequenceDiagram
Executor->>Tool: lookup_knowledge / logs / metrics Executor->>Tool: lookup_knowledge / logs / metrics
Tool->>DB: 写入 tool_invocation Tool->>DB: 写入 tool_invocation
Tool-->>Executor: 返回证据 Tool-->>Executor: 返回证据
Executor->>Verifier: 生成候选诊断并校验 Executor->>Gatekeeper: 输出 executor_evidence_v2
Gatekeeper->>DB: 读取 tool_invocation.evidence_refs 并校验引用
Gatekeeper->>Verifier: 传入已验真的 claims / excerpts
Verifier->>DB: 合并 self_evaluation.verifier_evaluation Verifier->>DB: 合并 self_evaluation.verifier_evaluation
Verifier->>Composer: 传入 allowed_claims / missing_info / actions
Composer->>Chat: 生成最终用户答复
Chat->>DB: 保存 diagnosis_session.answer Chat->>DB: 保存 diagnosis_session.answer
User->>Trace: GET /api/diagnosis/{sessionId}/trace User->>Trace: GET /api/diagnosis/{sessionId}/trace
Trace->>DB: 聚合 session / step / tool Trace->>DB: 聚合 session / step / tool
@@ -180,14 +191,16 @@ POST /api/chat
-> lookup_knowledge -> lookup_knowledge
-> query_logs -> query_logs
-> query_metrics -> query_metrics
-> Verifier 校验最终诊断 -> Gatekeeper 校验 Executor 证据引用真实性
-> Verifier 判断 claim 是否能由已核验证据推出
-> Composer 生成最终用户答复
-> 保存 diagnosis_session -> 保存 diagnosis_session
-> 保存 agent_step -> 保存 agent_step
-> 保存 tool_invocation -> 保存 tool_invocation
-> 合并 self_evaluation.verifier_evaluation -> 合并 self_evaluation.verifier_evaluation
``` ```
Chat 链路的质量门禁是 LLM Verifier。Verifier 输出合并到 `diagnosis_session.self_evaluation.verifier_evaluation`,Trace API 会展示该验证结果。 Chat 链路的质量门禁由三段组成:Gatekeeper 先做代码级引用验真,Verifier 再做 LLM 可推导性判断,Composer 最后控制对用户的表达边界。Gatekeeper、Verifier、Composer 的输出合并到 `diagnosis_session.self_evaluation.verifier_evaluation`,Trace API 会展示该验证结果。
Agent 编排细节见 [agent-orchestration.md](agent-orchestration.md)。 Agent 编排细节见 [agent-orchestration.md](agent-orchestration.md)。
@@ -323,6 +336,7 @@ tool_invocation
-> 工具调用事实 -> 工具调用事实
-> tool_name / input_params / output_preview -> tool_name / input_params / output_preview
-> retrieval_layer / retrieval_details -> retrieval_layer / retrieval_details
-> retrieval_details.evidence_refs
-> relevance_level / dedup_reason -> relevance_level / dedup_reason
-> duration / success -> duration / success
``` ```
@@ -346,12 +360,12 @@ Trace API 聚合:
- 会话状态和最终报告。 - 会话状态和最终报告。
- Agent step 序列。 - Agent step 序列。
- 工具调用和检索细节。 - 工具调用和检索细节。
- Chat verifier 结果。 - Chat Gatekeeper / Verifier / Composer 结果。
- AIOps rule evaluation 结果。 - AIOps rule evaluation 结果。
Trace 是本项目区别于普通问答系统的关键:答案不是孤立文本,而是可以追溯到 Agent 决策、工具调用和证据来源。 Trace 是本项目区别于普通问答系统的关键:答案不是孤立文本,而是可以追溯到 Agent 决策、工具调用和证据来源。
Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gates.md](harness-quality-gates.md),用户反馈与 `self_evaluation` 闭环见 [feedback-architecture.md](feedback-architecture.md)。 Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明见 [harness-quality-gates.md](harness-quality-gates.md),用户反馈与 `self_evaluation` 闭环见 [feedback-architecture.md](feedback-architecture.md)。
## 8. 质量门禁 ## 8. 质量门禁
@@ -359,7 +373,9 @@ Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gate
| 门禁 | 位置 | 作用 | | 门禁 | 位置 | 作用 |
|---|---|---| |---|---|---|
| Chat Verifier | `ChatService` | 校验普通诊断回答质量 | | Executor Gatekeeper | `VerifierInputHook` / `ExecutorGatekeeperService` | 校验 Executor 引用的 invocation、`raw_path`、`evidence_excerpt` 是否真实 |
| Chat Verifier | `ChatService` | 判断已验真证据是否能推出 Executor claims |
| Chat Composer | `ChatService` | 只表达 Verifier 允许输出的内容,避免把 no-evidence 说成已排除 |
| AIOps Rule Evaluation | `AiOpsRuleEvaluationService` | 校验告警诊断是否聚焦 payload 并使用证据 | | AIOps Rule Evaluation | `AiOpsRuleEvaluationService` | 校验告警诊断是否聚焦 payload 并使用证据 |
| Diagnosis Eval Baseline | `mvp/eval/` | 固化诊断 trace 和报告行为 | | Diagnosis Eval Baseline | `mvp/eval/` | 固化诊断 trace 和报告行为 |
| RAG Retrieval Baseline | `eval/rag-retrieval/` | 固化检索召回行为,避免 RAG 重构回退 | | RAG Retrieval Baseline | `eval/rag-retrieval/` | 固化检索召回行为,避免 RAG 重构回退 |
@@ -379,6 +395,11 @@ Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gate
- `title`、`breadcrumb`、`content` 参与 embedding 文本。 - `title`、`breadcrumb`、`content` 参与 embedding 文本。
- `tool_invocation` 记录检索层、relevance level、dedup reason。 - `tool_invocation` 记录检索层、relevance level、dedup reason。
- Chat verifier 和 AIOps rule evaluation 合并进 `self_evaluation`。 - Chat verifier 和 AIOps rule evaluation 合并进 `self_evaluation`。
- Chat Executor 结构化输出 `executor_evidence_v2`,不再直接承担最终用户答复。
- `tool_invocation.retrieval_details.evidence_refs` 支持 `raw_path` 精确引用和 `$.no_evidence` 负向证据。
- Gatekeeper 对 Executor 引用做代码级验真,并在审计中记录 `rule_set_version` 和规则元数据摘要。
- Verifier 只判断可推导性。
- Composer 在 Verifier 之后生成最终用户表达,并限制 negative observation 过度表述。
- RAG offline baseline 和 live acceptance 脚本。 - RAG offline baseline 和 live acceptance 脚本。
暂不作为当前已完成能力声明: 暂不作为当前已完成能力声明:
@@ -406,4 +427,5 @@ Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gate
| Spring AI VectorStore 配置辅助 | `SpringAiVectorStoreSidecarService` | | Spring AI VectorStore 配置辅助 | `SpringAiVectorStoreSidecarService` |
| Trace 聚合 | `DiagnosisTraceService` | | Trace 聚合 | `DiagnosisTraceService` |
| 工具调用记录 | `ToolInvocationRecorder` | | 工具调用记录 | `ToolInvocationRecorder` |
| Executor 引用验真 | `ExecutorGatekeeperService`, `VerifierInputHook` |
| self_evaluation 合并 | `SelfEvaluationMergeService` | | self_evaluation 合并 | `SelfEvaluationMergeService` |
+59 -5
View File
@@ -1,6 +1,6 @@
# 数据模型总览 # 数据模型总览
**更新日期**:2026-07-05 **更新日期**:2026-07-08
**状态**:当前可运行架构 **状态**:当前可运行架构
## 1. 定位 ## 1. 定位
@@ -130,10 +130,35 @@ erDiagram
用途: 用途:
- 给 Trace API 展示证据。 - 给 Trace API 展示证据。
- 给 Verifier 构造 `tool_trace_summary`。 - 给 Gatekeeper 提供 `retrieval_details.evidence_refs` 引用验真源。
- 给 Verifier 构造 `tool_trace_summary` 审计导航。
- 给 `EvaluationService` 计算 evidence score。 - 给 `EvaluationService` 计算 evidence score。
- 给 RAG eval 和人工排查提供检索细节。 - 给 RAG eval 和人工排查提供检索细节。
`retrieval_details.evidence_refs` 是当前 Chat 证据链路的关键字段:
```json
{
"evidence_status": "supported",
"evidence_refs": [
{
"raw_path": "$.logs[0]",
"text": "2026-07-08 23:05:28 ERROR order-service HikariPool-1 - Connection is not available..."
}
]
}
```
字段边界:
| 字段 | 说明 |
|---|---|
| `evidence_status` | 工具证据状态,例如 `supported`、`no_evidence`、`deduped`、`failed` |
| `evidence_refs[].raw_path` | Executor 可引用的稳定路径,例如 `$.logs[0]`、`$.alerts[0]`、`$.evidence_blocks[0]`、`$.no_evidence` |
| `evidence_refs[].text` | 系统抽取的最小证据文本,Gatekeeper 用它核对 `evidence_excerpt` |
`$.no_evidence` 只表示“本次工具查询未检索到匹配证据”,不能被解释为“问题不存在”或“根因已排除”。
## 4. 知识库模型 ## 4. 知识库模型
### api_document ### api_document
@@ -213,9 +238,39 @@ category
边界: 边界:
- `rule_evaluation` 评估证据收集充分度。 - `rule_evaluation` 评估证据收集充分度。
- `verifier_evaluation` 评估 Chat 答案关键事实是否有证据支撑。 - `verifier_evaluation` 评估 Chat 结构化 claims 是否能由已验真证据推出,并保存 Gatekeeper、Verifier、Composer 的审计数据。
- `aiops_rule_evaluation` 评估 AIOps 报告是否聚焦告警并使用证据。 - `aiops_rule_evaluation` 评估 AIOps 报告是否聚焦告警并使用证据。
当前 `verifier_evaluation` 关键结构:
```json
{
"verdict": "PASS",
"groundedness_score": 1.0,
"critical_fact_count": 1,
"claim_checks": [],
"facts_checked": [],
"rationale": "...",
"round": 1,
"traceability_version": "v1",
"executor_output_parse_status": {},
"executor_structured_output": {},
"gatekeeper_result": {},
"composer_output": {},
"tool_trace_summary": []
}
```
必要审计字段:
| 字段 | 说明 |
|---|---|
| `executor_output_parse_status` | Executor 输出是否能解析为 `executor_evidence_v2` |
| `executor_structured_output` | Executor 结构化 claims、hypotheses、recommended_actions、missing_info |
| `gatekeeper_result` | 引用真实性校验结果,包括 rule set version、checked bindings、failed rules、warnings、errors |
| `composer_output` | Composer 最终表达及解析状态 |
| `tool_trace_summary` | Verifier 调用时使用的工具调用导航索引,不是唯一证据源 |
## 7. 数据写入时序 ## 7. 数据写入时序
```mermaid ```mermaid
@@ -253,6 +308,5 @@ sequenceDiagram
1. 增加 run id,支持同 session 多次独立诊断。 1. 增加 run id,支持同 session 多次独立诊断。
2. 强化 `tool_invocation.step_id` 关联。 2. 强化 `tool_invocation.step_id` 关联。
3. 将 evidence block 结构化保存。 3. 将 Gatekeeper 规则配置化时的规则元数据保存为可审计版本。
4. 将 `case_library` 的 rootCause/solution 从完整 answer 中结构化抽取。 4. 将 `case_library` 的 rootCause/solution 从完整 answer 中结构化抽取。
@@ -1,78 +1,44 @@
# Current Chat Agent Data Contracts # Chat Evidence Pipeline Contracts
**状态**:当前实现 **状态**:当前实现
**日期**:2026-07-07 **更新日期**:2026-07-08
**范围**:当前 Chat 复杂诊断链路的数据结构定义 **范围**:Chat 复杂诊断链路中的 Planner、Executor、Gatekeeper、Verifier、Composer 数据契约
当前代码实现是三 Agent 顺序链路: 当前 Chat 复杂诊断链路是:
```text ```text
chat_planner -> chat_executor -> chat_verifier chat_planner
-> chat_executor
-> VerifierInputHook / ExecutorGatekeeperService
-> chat_verifier
-> chat_composer
-> final answer
``` ```
对应 `ChatService.executeChatComplex(...)` 中的 `SequentialAgent`。 设计原则:
- Planner 暂不输出 `scope_contract`。
- Executor 只做证据收集和微观事实提炼,不生成最终用户答案。
- Gatekeeper 在 Verifier 前做代码级引用真实性校验。
- Verifier 判断 claim 是否能由已核验证据推出。
- Composer 只表达 Verifier 允许输出的内容。
--- ---
## 1. Workflow Input ## 1. Planner
由 `ChatService.buildWorkflowInput(...)` 构造,传给 `chat_workflow`。 Planner 当前保持不变,输出 `planner_plan`:
```text
请按固定工作流完成本轮 Planner -> Executor -> Verifier。
--- 用户问题 ---
{question}
--- retry_context ---
{retry_context}
Verifier 完成后由外层代码读取 verifier_output 并决定最终用户输出。
```
| 字段 | 来源 | 定义 |
|---|---|---|
| `question` | 用户输入 | 用户本轮原始问题 |
| `retry_context` | ChatService | 第二轮补证据约束;首轮为空 |
---
## 2. chat_planner
### 2.1 Input
`chat_planner` 的输入来自 workflow input 和 system prompt 追加上下文。
```json ```json
{ {
"question": "用户原始问题", "selected_skill": "diagnose-mysql-connection-pool",
"history": [], "selection_reason": "选择该 skill 的原因",
"available_knowledge_domains": "...", "plan": ["步骤1", "步骤2"],
"skill_catalog": {}, "reasoning": "规划思路"
"retry_context": null
} }
``` ```
| 字段 | 来源 | 定义 | 字段定义:
|---|---|---|
| `question` | workflow input | 用户原始问题 |
| `history` | `ChatService.buildChatPlannerAgent(...)` | 对话历史,拼接到 planner system prompt |
| `available_knowledge_domains` | `KnowledgeDomainService.buildKnowledgeMap()` | 可用知识域地图,拼接到 planner system prompt |
| `skill_catalog` | `PlannerSkillMetadataHook` | Planner 可见的 skill name/description 元数据 |
| `retry_context` | `ChatService` | Verifier 低置信后构造的补证据上下文 |
### 2.2 Output:`planner_plan`
当前 prompt 要求输出 JSON:
```json
{
"selected_skill": "匹配的 skill 名称;如果没有匹配则为 null",
"selection_reason": "选择该 skill 的原因;如果没有匹配则说明不使用 skill",
"plan": ["步骤1描述", "步骤2描述", "步骤3描述"],
"reasoning": "规划思路说明"
}
```
| 字段 | 类型 | 定义 | | 字段 | 类型 | 定义 |
|---|---|---| |---|---|---|
@@ -81,434 +47,397 @@ Verifier 完成后由外层代码读取 verifier_output 并决定最终用户输
| `plan` | array | 给 Executor 的执行步骤 | | `plan` | array | 给 Executor 的执行步骤 |
| `reasoning` | string | 规划思路说明 | | `reasoning` | string | 规划思路说明 |
运行态输出 key: 当前边界:
```text - 不新增 `scope_contract`。
planner_plan - 不要求 Planner 显式列出 forbidden actions。
``` - 窄范围控制先由 Executor Prompt 约束,后续如仍不稳定再引入 Planner contract。
--- ---
## 3. chat_executor ## 2. Executor
### 3.1 Input Executor 输出 `executor_evidence_v2`。它不是最终答复,而是给 Gatekeeper、Verifier、Composer 使用的结构化诊断材料。
`chat_executor` 接收前序 `planner_plan`,并通过 system prompt 获得历史、skill 读取约束、retry 约束和工具权限。 ### 2.1 输出结构
```json ```json
{ {
"planner_plan": {}, "answer_version": "executor_evidence_v2",
"history": [],
"retry_context": null,
"tool_permissions": {
"method_tools": ["dateTimeTools", "lookupKnowledgeTool", "queryMetricsTools", "queryLogsTools"],
"tool_callbacks": []
}
}
```
| 字段 | 来源 | 定义 |
|---|---|---|
| `planner_plan` | `chat_planner` | Planner 输出的计划 |
| `history` | `ChatService.buildChatExecutorAgent(...)` | 对话历史,拼接到 executor system prompt |
| `retry_context` | `ChatService` | 本轮补证据约束 |
| `method_tools` | `ChatService.buildMethodToolsArray()` | Executor 可直接调用的本地工具 |
| `tool_callbacks` | `ToolCallback[]` | 框架发现或外部注入工具 |
| `read_skill` | `SkillsAgentHook` | 当存在 skillRegistry 时,Executor 可读取 Planner 选中的 skill |
### 3.2 Output:`executor_feedback`
当前 `chat-executor-prompt.md` 要求输出一个 JSON 对象,即 `executor_evidence_v1`。
```json
{
"answer_version": "executor_evidence_v1",
"diagnosis_summary": "1-2句话总结,仅包含有证据支撑的事实和证据边界",
"claims": [ "claims": [
{ {
"claim_id": "claim-1", "claim_id": "claim-1",
"claim_type": "root_cause", "claim_type": "observation",
"claim_text": "事实断言或有限结论", "claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
"support_level": "direct", "support_level": "direct",
"evidence_bindings": [ "evidence_bindings": [
{ {
"source_type": "tool_trace", "source_type": "tool_trace",
"source_id": "工具返回中的 evidence block id、trace_ref 或可定位标识", "source_id": "",
"tool_name": "lookup_knowledge/query_logs/query_metrics/read_skill 等", "tool_name": "query_metrics",
"source_invocation_ids": [], "source_invocation_id": 517,
"evidence_excerpt": "从工具返回中摘取的原话、指标值、日志片段或关键数据" "raw_path": "$.alerts[0]",
"evidence_excerpt": "HighCPUUsage, service=payment-service, state=firing, current=92%, duration=25m"
} }
] ]
} }
], ],
"hypotheses": [ "hypotheses": [],
{ "recommended_actions": [],
"hypothesis_text": "未被证实但值得排查的方向", "missing_info": []
"basis": "它基于哪些已知证据或为什么只是推测",
"needed_evidence": ["需要补充的证据"]
}
],
"recommended_actions": [
{
"action_text": "建议动作",
"reason": "为什么建议做这个动作",
"evidence_bindings": []
}
],
"missing_info": [
"导致无法确认完整根因的证据缺口"
],
"user_facing_answer": "面向用户的中文回答。必须与 claims/hypotheses/recommended_actions/missing_info 一致。"
} }
``` ```
| 字段 | 类型 | 定义 | 字段定义:
| 字段 | 类型 | 必填 | 定义 |
|---|---|---:|---|
| `answer_version` | string | 是 | 固定为 `executor_evidence_v2` |
| `claims` | array | 是 | Executor 提出的待验证事实断言 |
| `claims[].claim_id` | string | 是 | claim 标识 |
| `claims[].claim_type` | string | 是 | `observation`、`negative_observation`、`symptom`、`root_cause` 等;窄范围任务只允许前两者 |
| `claims[].claim_text` | string | 是 | 事实断言文本 |
| `claims[].support_level` | string | 是 | `direct` 或 `indirect` |
| `claims[].evidence_bindings` | array | 是 | 支撑该 claim 的证据绑定,不能为空 |
| `evidence_bindings[].source_type` | string | 否 | 当前通常为 `tool_trace` |
| `evidence_bindings[].source_id` | string | 否 | 兼容字段,不作为精确引用主键 |
| `evidence_bindings[].tool_name` | string | 是 | `query_logs`、`query_metrics`、`lookup_knowledge` 等 |
| `evidence_bindings[].source_invocation_id` | number/null | 是 | 来源 `tool_invocation.id`;缺失时 Gatekeeper 只在能唯一匹配时回填 |
| `evidence_bindings[].raw_path` | string | 是 | 工具返回中的稳定定位路径 |
| `evidence_bindings[].evidence_excerpt` | string | 是 | 工具返回中的原文片段或系统抽取的最小证据文本 |
| `hypotheses` | array | 是 | 未证实但值得排查的方向,不是 confirmed fact |
| `recommended_actions` | array | 是 | 下一步动作;本期只允许证据收集或继续排查动作 |
| `missing_info` | array | 是 | 无法确认结论所缺少的证据 |
禁止字段:
- `diagnosis_summary`
- `user_facing_answer`
- `source_invocation_ids` 作为主引用字段
### 2.2 raw_path
当前支持的精确路径:
| 工具 | 正向证据路径 | 负向证据路径 |
|---|---|---| |---|---|---|
| `answer_version` | string | 当前固定为 `executor_evidence_v1` | | `query_metrics` | `$.alerts[i]` | `$.no_evidence` |
| `diagnosis_summary` | string | 有证据边界的简短诊断摘要 | | `query_logs` | `$.logs[i]` | `$.no_evidence` |
| `claims` | array | 已证实或有明确间接支撑的事实断言 | | `lookup_knowledge` | `$.evidence_blocks[i]` | `$.no_evidence` |
| `claims[].claim_id` | string | claim 标识 |
| `claims[].claim_type` | string | claim 类型,例如 `root_cause`、`symptom`、`impact` |
| `claims[].claim_text` | string | 事实断言文本 |
| `claims[].support_level` | string | `direct` 或 `indirect` |
| `claims[].evidence_bindings` | array | 支撑 claim 的证据绑定,不能为空 |
| `evidence_bindings[].source_type` | string | 证据来源类型,例如 `tool_trace` |
| `evidence_bindings[].source_id` | string | evidence block id、trace_ref 或其它定位标识 |
| `evidence_bindings[].tool_name` | string | 来源工具名 |
| `evidence_bindings[].source_invocation_ids` | array | 来源 `tool_invocation.id` |
| `evidence_bindings[].evidence_excerpt` | string | 工具返回中的原话、指标值、日志片段或关键数据 |
| `hypotheses` | array | 未证实但值得排查的方向 |
| `hypotheses[].hypothesis_text` | string | 假设文本 |
| `hypotheses[].basis` | string | 假设依据和未证实原因 |
| `hypotheses[].needed_evidence` | array | 确认该假设还需要的证据 |
| `recommended_actions` | array | 建议动作 |
| `recommended_actions[].action_text` | string | 建议动作文本 |
| `recommended_actions[].reason` | string | 建议原因 |
| `recommended_actions[].evidence_bindings` | array | 建议动作关联证据,可为空 |
| `missing_info` | array | 证据缺口 |
| `user_facing_answer` | string | 候选用户答案,PASS 时由 ChatService 提取输出 |
运行态输出 key: 约束:
```text - `raw_path` 必须指向数组条目或 `$.no_evidence`。
executor_feedback - 禁止字段级子路径,例如 `$.alerts[0].state`、`$.logs[0].message`。
- 同一条工具数组项只能绑定一次;多个字段应合并进同一个 `evidence_excerpt`。
### 2.3 negative_observation
当工具明确返回 no-hit / no-evidence 时,Executor 可以输出 `negative_observation`:
```json
{
"claim_id": "claim-1",
"claim_type": "negative_observation",
"claim_text": "当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。",
"support_level": "direct",
"evidence_bindings": [
{
"tool_name": "query_logs",
"source_invocation_id": 517,
"raw_path": "$.no_evidence",
"evidence_excerpt": "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence"
}
]
}
``` ```
语义边界:
- `$.no_evidence` 只表示“该工具对当前查询返回无匹配证据”。
- 不表示“问题绝对不存在”。
- 不表示“根因被排除”。
- 不表示“系统已经健康”。
- `negative_observation` 的 `evidence_bindings` 只能绑定 `$.no_evidence`,不能混绑其它服务的正向日志。
### 2.4 窄范围任务
窄范围任务指用户只要求确认某个服务、告警、日志、错误、订单或时间窗口。
Executor 必须遵守:
- 只输出 `observation` / `negative_observation`。
- claim 数量通常 1 条,最多 2 条。
- claim 数量限制不限制 `evidence_bindings` 数量。
- 不输出根因、风险、修复建议、经验推断。
- 不把 Runbook / Skill / 知识库通用知识写成当前环境事实。
- 精确查询返回 no-evidence 后,不得放宽关键词、删除服务名或扩大服务范围继续查。
--- ---
## 4. chat_verifier ## 3. Tool Invocation Evidence Refs
### 4.1 Input 工具调用落库到 `tool_invocation`,其中 `retrieval_details.evidence_refs` 是 Gatekeeper 的主校验源。
`VerifierInputHook` 会在 Verifier 调用前替换消息历史,构造显式 JSON payload。 ### 3.1 正向证据
```json
{
"evidence_status": "supported",
"evidence_refs": [
{
"raw_path": "$.logs[0]",
"text": "2026-07-08 23:05:28 ERROR order-service HikariPool-1 - Connection is not available..."
}
]
}
```
### 3.2 负向证据
```json
{
"evidence_status": "no_evidence",
"evidence_refs": [
{
"raw_path": "$.no_evidence",
"text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志"
}
]
}
```
字段定义:
| 字段 | 类型 | 定义 |
|---|---|---|
| `evidence_status` | string | `supported`、`no_evidence`、`deduped`、`failed` |
| `evidence_refs[].raw_path` | string | 证据在工具返回中的稳定定位符 |
| `evidence_refs[].text` | string | 系统抽取的最小证据文本,供 Gatekeeper 和 Verifier 使用 |
---
## 4. Gatekeeper
Gatekeeper 位于 Verifier 前,由 `VerifierInputHook` 触发,负责代码级引用真实性校验。
### 4.1 输入
- `sessionId`
- `executor_structured_output`
- 当前 session 的 `tool_invocation`
### 4.2 输出
```json
{
"status": "pass",
"severity": "none",
"rule_set_version": "gatekeeper-rules-v1",
"rules": [
{
"id": "evidence.raw_path",
"description": "raw_path must exist in retrieval_details.evidence_refs",
"enabled": true,
"default_severity": "reject"
}
],
"checked_bindings": [
{
"claim_id": "claim-1",
"tool_name": "query_logs",
"source_invocation_id": 517,
"raw_path": "$.no_evidence",
"matched_text": "query_logs returned no evidence; ...",
"status": "pass"
}
],
"failed_rules": [],
"warnings": [],
"errors": []
}
```
字段定义:
| 字段 | 类型 | 定义 |
|---|---|---|
| `status` | string | `pass` 或 `fail` |
| `severity` | string | `none`、`low_confid`、`reject` |
| `rule_set_version` | string | 当前加载的 Gatekeeper 规则集版本 |
| `rules` | array | 已启用规则的轻量元数据摘要 |
| `checked_bindings` | array | 每条证据绑定的校验结果 |
| `failed_rules` | array | 失败规则 id |
| `warnings` | array | 自动回填等非阻断信息 |
| `errors` | array | 失败明细 |
校验规则:
- `answer_version` 必须是 `executor_evidence_v2`。
- 不允许 `diagnosis_summary` / `user_facing_answer`。
- 每个 claim 必须有非空 `evidence_bindings`。
- `tool_name` 必须和真实 invocation 对齐。
- `source_invocation_id` 必须存在;缺失时只在 `tool_name + raw_path + evidence_excerpt` 能唯一匹配真实 invocation 时回填。
- `raw_path` 必须存在于 `retrieval_details.evidence_refs`。
- `evidence_excerpt` 必须由 `evidence_refs[].text` 支撑。
- `negative_observation` 只能绑定 `$.no_evidence`。
规则配置:
- 当前规则元数据位于 `src/main/resources/gatekeeper/gatekeeper-rules.json`。
- 规则实现仍是确定性 Java 代码,不执行动态脚本。
- 当前配置只承载规则 id、描述、默认 severity、启用状态和简单参数,例如 excerpt token overlap 阈值。
失败分级:
| 场景 | severity |
|---|---|
| 伪造 invocation id | `reject` |
| tool_name 与 invocation 不匹配 | `reject` |
| raw_path 不存在 | `reject` |
| excerpt 与 matched_text 不匹配 | `reject` |
| negative_observation 绑定正向日志 | `reject` |
| 缺少 raw_path / invocation id 且无法唯一回填 | `low_confid` |
| 旧 invocation 没有 `evidence_refs` | `low_confid` |
---
## 5. Verifier
Verifier 输入由 `VerifierInputHook` 构造:
```json ```json
{ {
"original_query": "用户原始问题", "original_query": "用户原始问题",
"executor_final_answer": "{...executor_feedback raw text...}", "executor_final_answer": "{...executor raw text for debug/fallback only...}",
"executor_structured_output": {}, "executor_structured_output": {
"answer_version": "executor_evidence_v2",
"claims": []
},
"executor_output_parse_status": { "executor_output_parse_status": {
"status": "valid", "status": "valid",
"detail": "parsed executor evidence contract" "detail": "parsed executor evidence contract"
}, },
"tool_trace_summary": [], "tool_trace_summary": [],
"gatekeeper_result": {},
"retry_context": null "retry_context": null
} }
``` ```
| 字段 | 来源 | 定义 | Verifier 职责:
|---|---|---|
| `original_query` | `VerifierContextHolder` | 用户原始问题 |
| `executor_final_answer` | `VerifierContextHolder` 或上一条 AssistantMessage | Executor 原始输出文本 |
| `executor_structured_output` | `VerifierInputHook.parseExecutorOutput(...)` | Executor 输出可解析且包含 `claims` 时的 JSON 对象;否则为 null |
| `executor_output_parse_status.status` | `VerifierInputHook` | `valid` / `missing` / `malformed` |
| `executor_output_parse_status.detail` | `VerifierInputHook` | 解析状态说明 |
| `tool_trace_summary` | `ToolTraceSummaryService.buildVerifierTraceSummary(...)` | 基于真实 `tool_invocation` 构建的证据索引 |
| `retry_context` | `VerifierContextHolder` | 当前补证据上下文 |
### 4.2 `tool_trace_summary` - 不调用工具。
- 不读 skill。
- 不逐字核验 excerpt 真伪;这由 Gatekeeper 完成。
- 只判断 `claim_text` 是否能由已核验的 `evidence_excerpt` 推出。
- 结构化输出有效时,不得从 `executor_final_answer` 抽取额外确认事实。
- 对 `gatekeeper_result.severity=reject` 不得输出 `PASS`。
- 对 `gatekeeper_result.severity=low_confid` 不得输出 `PASS`。
`ToolTraceSummaryService` 聚合 evidence tools: 输出:
```text
lookup_knowledge, query_logs, query_metrics, query_order
```
输出项结构:
```json
{
"trace_ref": "trace-1",
"tool_name": "query_logs",
"success": true,
"input_summary": "query=payment-service timeout",
"output_summary": "log_evidence: ...",
"evidence_level": "direct",
"topic_domain": "general",
"source_invocation_ids": [394],
"invocation_count": 1,
"failed_invocation_count": 0,
"no_hit_invocation_count": 0,
"query_samples": ["payment-service timeout"],
"retrieval_layers": [],
"relevance_levels": [],
"source_documents": []
}
```
| 字段 | 类型 | 定义 |
|---|---|---|
| `trace_ref` | string | Verifier 可引用的证据摘要编号 |
| `tool_name` | string | 聚合后的工具名 |
| `success` | boolean | 是否存在可用证据 |
| `input_summary` | string | 工具输入摘要 |
| `output_summary` | string | 工具输出摘要 |
| `evidence_level` | string | `direct` / `indirect` / `none` |
| `topic_domain` | string | 主题域,优先来自 `retrieval_details.retrieved_domains` |
| `source_invocation_ids` | array | 聚合的 `tool_invocation.id` |
| `invocation_count` | number | 聚合调用次数 |
| `failed_invocation_count` | number | 失败调用次数 |
| `no_hit_invocation_count` | number | 无证据或去重调用次数 |
| `query_samples` | array | 查询样例 |
| `retrieval_layers` | array | 检索层级 |
| `relevance_levels` | array | 相关性等级 |
| `source_documents` | array | 来源文档标签 |
### 4.3 Output:`verifier_output`
当前 `chat-verifier-prompt.md` 要求输出:
```json ```json
{ {
"verdict": "PASS", "verdict": "PASS",
"groundedness_score": 0.8, "groundedness_score": 1.0,
"critical_fact_count": 2, "critical_fact_count": 1,
"facts_checked": [ "claim_checks": [],
{ "facts_checked": [],
"fact": "ERR_TIMEOUT 表示请求超时", "rationale": "..."
"is_critical": true,
"verification": "direct_evidence",
"detail": "知识库文档明确给出该错误码定义",
"evidence_refs": [
{
"trace_ref": "trace-1",
"tool_name": "lookup_knowledge",
"topic_domain": "api",
"source_invocation_ids": [101, 104],
"note": "trace-1 的文档摘要直接给出错误码定义"
}
]
}
],
"rationale": "所有关键事实均有支撑,且至少一条具有直接证据"
} }
``` ```
| 字段 | 类型 | 定义 |
|---|---|---|
| `verdict` | string | `PASS` / `LOW_CONFID` / `REJECT` |
| `groundedness_score` | number | 关键事实证据支撑评分 |
| `critical_fact_count` | number | `facts_checked` 中 `is_critical=true` 的数量 |
| `facts_checked` | array | Verifier 校验过的事实列表 |
| `facts_checked[].fact` | string | 被校验事实 |
| `facts_checked[].is_critical` | boolean | 是否关键事实 |
| `facts_checked[].verification` | string | `direct_evidence` / `indirect_support` / `no_evidence` / `contradicted` |
| `facts_checked[].detail` | string | 校验说明 |
| `facts_checked[].evidence_refs` | array | 证据引用 |
| `evidence_refs[].trace_ref` | string | 引用的 `tool_trace_summary.trace_ref` |
| `evidence_refs[].tool_name` | string | 引用工具 |
| `evidence_refs[].topic_domain` | string | 引用主题域 |
| `evidence_refs[].source_invocation_ids` | array | 引用的 `tool_invocation.id` |
| `evidence_refs[].note` | string | 引用说明 |
| `rationale` | string | verdict 判定理由 |
运行态输出 key:
```text
verifier_output
```
--- ---
## 5. VerifierDecision ## 6. Composer
`ChatService.parseVerifierDecision(...)` 将 `verifier_output` 解析为内部 record: Composer 位于 Verifier 之后,输入是 ChatService 过滤后的允许表达材料。
输入概念:
| 字段 | 定义 |
|---|---|
| `original_query` | 用户原始问题 |
| `verdict` | `PASS` / `LOW_CONFID` / `REJECT` |
| `allowed_claims` | Verifier 允许表达的 claims |
| `allowed_hypotheses` | Verifier 允许表达的假设 |
| `missing_info` | 证据缺口 |
| `recommended_actions` | 允许表达的建议动作 |
| `rationale` | Verifier 判定理由 |
输出:
```json ```json
{ {
"verdict": "LOW_CONFID", "answer_summary": "一句话概括",
"groundednessScore": 0.5, "recommended_actions": [
"criticalFactCount": 2, {
"factsChecked": [], "action_text": "下一步动作",
"rationale": "证据不足", "reason": "原因"
"round": 1 }
}
```
| 字段 | 类型 | 定义 |
|---|---|---|
| `verdict` | string | Verifier verdict |
| `groundednessScore` | number | groundedness score |
| `criticalFactCount` | number | 关键事实数量 |
| `factsChecked` | array | 解析后的 facts_checked |
| `rationale` | string | 判定理由 |
| `round` | number | 当前验证轮次 |
---
## 6. retry_context
当 `LOW_CONFID` 且满足重试条件时,`ChatService.buildRetryContext(...)` 构造:
```json
{
"round": 1,
"missing_evidence_facts": [
"某关键事实:缺少直接证据"
], ],
"instruction": "仅补充以上断言相关证据,不要重复已完成检索" "user_facing_answer": "最终给用户看的中文答案"
} }
``` ```
| 字段 | 类型 | 定义 | 表达边界:
|---|---|---|
| `round` | number | 触发 retry 的轮次 | - Composer 不补事实、不补根因、不调用工具。
| `missing_evidence_facts` | array | 来自 Verifier 的证据缺口 | - 只表达 `allowed_claims`、`allowed_hypotheses`、`missing_info`、`recommended_actions`。
| `instruction` | string | 补证据约束 | - 当 claim 是 `negative_observation` 或证据来自 `$.no_evidence` 时,只能表达“当前查询未检索到 / 本次检索未发现匹配证据”。
- 禁止表达“问题不存在”“已排除该问题”“确认没有”“日志层面已排除”等过度结论。
--- ---
## 7. diagnosis_session.self_evaluation.verifier_evaluation ## 7. Trace Persistence
`ChatService.persistVerifierEvaluation(...)` 将 Verifier 结果合并进 `diagnosis_session.self_evaluation`。 `diagnosis_session.self_evaluation.verifier_evaluation` 持久化:
```json ```json
{ {
"verifier_evaluation": { "verifier_evaluation": {
"verdict": "LOW_CONFID", "verdict": "PASS",
"groundedness_score": 0.5, "groundedness_score": 1.0,
"critical_fact_count": 2, "critical_fact_count": 1,
"claim_checks": [],
"facts_checked": [], "facts_checked": [],
"rationale": "证据不足", "rationale": "...",
"round": 1, "round": 1,
"traceability_version": "v1", "traceability_version": "v1",
"executor_output_parse_status": { "executor_output_parse_status": {},
"status": "valid",
"detail": "parsed executor evidence contract"
},
"executor_structured_output": {}, "executor_structured_output": {},
"gatekeeper_result": {
"rule_set_version": "gatekeeper-rules-v1"
},
"composer_output": {},
"tool_trace_summary": [] "tool_trace_summary": []
} }
} }
``` ```
| 字段 | 类型 | 定义 | Trace API 可用于回放:
|---|---|---|
| `verifier_evaluation.verdict` | string | Verifier verdict | - Executor 输出了哪些 claim。
| `verifier_evaluation.groundedness_score` | number | groundedness score | - 每个 claim 引用了哪些 `source_invocation_id + raw_path + evidence_excerpt`。
| `verifier_evaluation.critical_fact_count` | number | 关键事实数量 | - Gatekeeper 是否通过、是否自动回填。
| `verifier_evaluation.facts_checked` | array | 校验事实列表 | - Verifier 如何判断可推导性。
| `verifier_evaluation.rationale` | string | 判定理由 | - Composer 最终如何表达给用户。
| `verifier_evaluation.round` | number | 验证轮次 |
| `verifier_evaluation.traceability_version` | string | 当前固定为 `v1` |
| `verifier_evaluation.executor_output_parse_status` | object | Executor 输出解析状态 |
| `verifier_evaluation.executor_structured_output` | object/null | 解析后的 Executor 结构化输出 |
| `verifier_evaluation.tool_trace_summary` | array | Verifier 使用的工具证据索引 |
--- ---
## 8. Final Answer Rendering ## 8. 当前已验证样例
ChatService 根据 Verifier verdict 决定最终 `diagnosis_session.answer`。 | 场景 | sessionId | 结果 |
|---|---|---|
| Verdict | 当前行为 | | HighCPUUsage 窄范围正向确认 | `iss008-narrow-highcpu-rerun-20260708-215510` | `PASS`,1 条 `observation`,无越界 claim |
|---|---| | HikariCP negative_observation | `iss009-hikari-negative-latest-20260708-232428` | `PASS`,`raw_path=$.no_evidence`,无过度表达 |
| `PASS` | 优先提取 `executor_feedback.user_facing_answer`;提取失败则使用 executor 原文 |
| `LOW_CONFID` | 输出低置信模板:已确认信息、当前缺口、建议下一步 |
| `REJECT` | 输出降级模板:已确认信息、证据缺口、建议下一步 |
低置信模板使用:
```text
以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。
已确认信息:
- ...
当前缺口:
- ...
建议下一步:
- ...
```
拒绝模板使用:
```text
当前无法基于已获取证据生成可靠结论。
已确认信息:
- ...
证据缺口:
- ...
建议下一步:
- ...
```
--- ---
## 9. Trace Persistence Data ## 9. 仍需记录或后续补强
### 9.1 diagnosis_session 当前架构文档已记录主链路、数据契约和语义边界。后续如果继续实现,建议再补:
| 字段 | 类型 | 定义 | 1. Planner `scope_contract` 的 ADR:只有当 Prompt-first 无法稳定控制越界时再引入。
|---|---|---| 2. 更完整的 Gatekeeper 规则配置化:当前只有本地轻量 metadata/catalog,后续如果做索引层、元数据层、远程规则层,需要单独记录加载顺序、变更审批和回滚策略。
| `session_id` | string | 会话 id | 3. Prompt version 记录:当前 prompt 变更没有版本号,后续如果需要回滚和对比,应记录 prompt version。
| `query` | text | 用户问题 |
| `status` | string | 会话状态 |
| `agent_flow` | string | 当前 Chat 链路为 `CHAT` |
| `total_duration_ms` | number | 总耗时 |
| `total_token_count` | number | 总 token |
| `step_count` | number | agent step 数 |
| `tool_call_count` | number | tool invocation 数 |
| `answer` | longtext | 最终用户答案 |
| `self_evaluation` | json | 包含 verifier_evaluation |
| `feedback` | string | 用户反馈 |
### 9.2 agent_step
| 字段 | 类型 | 定义 |
|---|---|---|
| `session_id` | string | 会话 id |
| `step_index` | number | 步骤序号 |
| `agent_name` | string | `planner` / `executor` / `verifier` |
| `model_input` | text | 模型输入摘要 |
| `model_output` | text | 模型输出摘要 |
| `thought` | text | hook 记录的摘要信息 |
| `has_tool_call` | boolean | 是否包含工具调用 |
| `duration_ms` | number | 模型调用耗时 |
| `token_count` | number | token 数 |
### 9.3 tool_invocation
| 字段 | 类型 | 定义 |
|---|---|---|
| `id` | number | 工具调用 id |
| `session_id` | string | 会话 id |
| `step_id` | number | 对应 agent_step id |
| `tool_name` | string | 工具名 |
| `input_params` | json | 工具输入参数 |
| `output_preview` | text | 工具输出预览 |
| `output_length` | number | 原始输出长度 |
| `retrieval_layer` | string | 检索层 |
| `l0_match_count` | number | L0 命中数 |
| `l1_match_count` | number | L1 命中数 |
| `is_truncated` | boolean | 输出是否截断 |
| `relevance_level` | string | 相关性等级 |
| `dedup_reason` | string | 去重原因 |
| `retrieval_details` | json | 检索细节 |
| `duration_ms` | number | 工具耗时 |
| `success` | boolean | 是否成功 |
| `error_message` | text | 错误信息 |
+43 -12
View File
@@ -1,6 +1,6 @@
# 反馈与自评估架构 # 反馈与自评估架构
**更新日期**:2026-07-05 **更新日期**:2026-07-08
**状态**:当前可运行架构 **状态**:当前可运行架构
**参考历史文档**:`archive/2026-07-05-legacy/confidence-feedback.md` **参考历史文档**:`archive/2026-07-05-legacy/confidence-feedback.md`
@@ -8,7 +8,7 @@
反馈架构包含两条闭环: 反馈架构包含两条闭环:
1. 系统自评估:基于工具调用、Verifier、AIOps 规则检查,写入 `diagnosis_session.self_evaluation`。 1. 系统自评估:基于工具调用、Gatekeeper、Verifier、Composer、AIOps 规则检查,写入 `diagnosis_session.self_evaluation`。
2. 用户反馈:用户标记 `useful` 或 `not_useful`,写入 `diagnosis_session.feedback`,其中 `useful` 会沉淀案例。 2. 用户反馈:用户标记 `useful` 或 `not_useful`,写入 `diagnosis_session.feedback`,其中 `useful` 会沉淀案例。
当前重要边界: 当前重要边界:
@@ -25,9 +25,14 @@ flowchart TD
subgraph SelfEval["Self evaluation"] subgraph SelfEval["Self evaluation"]
Invocation["tool_invocation"] --> RuleEval["EvaluationService: rule_evaluation"] Invocation["tool_invocation"] --> RuleEval["EvaluationService: rule_evaluation"]
Invocation --> EvidenceRefs["evidence_refs"]
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
Invocation --> TraceSummary["ToolTraceSummaryService"] Invocation --> TraceSummary["ToolTraceSummaryService"]
TraceSummary --> Verifier["chat_verifier"] Gatekeeper --> Verifier["chat_verifier"]
TraceSummary --> Verifier
Verifier --> VerifierEval["verifier_evaluation"] Verifier --> VerifierEval["verifier_evaluation"]
Verifier --> Composer["chat_composer"]
Composer --> VerifierEval
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"] Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
AiOpsRule --> AiOpsEval["aiops_rule_evaluation"] AiOpsRule --> AiOpsEval["aiops_rule_evaluation"]
end end
@@ -67,10 +72,15 @@ flowchart TD
"verdict": "PASS", "verdict": "PASS",
"groundedness_score": 0.8, "groundedness_score": 0.8,
"critical_fact_count": 2, "critical_fact_count": 2,
"claim_checks": [],
"facts_checked": [], "facts_checked": [],
"rationale": "...", "rationale": "...",
"round": 1, "round": 1,
"traceability_version": "v1", "traceability_version": "v1",
"executor_output_parse_status": {},
"executor_structured_output": {},
"gatekeeper_result": {},
"composer_output": {},
"tool_trace_summary": [] "tool_trace_summary": []
}, },
"aiops_rule_evaluation": { "aiops_rule_evaluation": {
@@ -117,16 +127,28 @@ flowchart TD
## 5. Chat Verifier 自评估 ## 5. Chat Verifier 自评估
Chat Verifier 校验 Executor 的最终答案是否被证据支撑。 Chat 自评估分三步:
1. Gatekeeper 用代码校验 Executor 的引用是否真实。
2. Verifier 判断已验真的 `evidence_excerpt` 是否能推出 `claim_text`。
3. Composer 只把 Verifier 允许表达的内容写成最终用户答复。
```mermaid ```mermaid
flowchart LR flowchart LR
Answer["executor_final_answer"] --> Verifier["chat_verifier"] ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
Invocation["tool_invocation"] --> Summary["ToolTraceSummaryService"] Invocation["tool_invocation"] --> Summary["ToolTraceSummaryService"]
Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
EvidenceRefs --> Gatekeeper
Gatekeeper --> GateResult["gatekeeper_result"]
Summary --> Evidence["tool_trace_summary"] Summary --> Evidence["tool_trace_summary"]
GateResult --> Verifier["chat_verifier"]
ExecutorOutput --> Verifier
Evidence --> Verifier Evidence --> Verifier
Verifier --> Output["verifier_output JSON"] Verifier --> Output["verifier_output JSON"]
Output --> Composer["chat_composer"]
Composer --> ComposerOutput["composer_output"]
Output --> Merge["SelfEvaluationMergeService.mergeVerifierEvaluation"] Output --> Merge["SelfEvaluationMergeService.mergeVerifierEvaluation"]
ComposerOutput --> Merge
Merge --> Session["diagnosis_session.self_evaluation.verifier_evaluation"] Merge --> Session["diagnosis_session.self_evaluation.verifier_evaluation"]
``` ```
@@ -137,16 +159,25 @@ Verifier 输出:
| `verdict` | `PASS` / `LOW_CONFID` / `REJECT` | | `verdict` | `PASS` / `LOW_CONFID` / `REJECT` |
| `groundedness_score` | 关键事实证据支撑度 | | `groundedness_score` | 关键事实证据支撑度 |
| `critical_fact_count` | 关键事实数量 | | `critical_fact_count` | 关键事实数量 |
| `claim_checks` | 对 Executor 结构化 claims 的逐条可推导性判断 |
| `facts_checked` | 逐条事实校验 | | `facts_checked` | 逐条事实校验 |
| `rationale` | 判定原因 | | `rationale` | 判定原因 |
| `tool_trace_summary` | 本次校验使用的证据索引 | | `executor_structured_output` | Executor 输出的结构化 claims 与证据绑定 |
| `gatekeeper_result` | 引用真实性校验结果 |
| `composer_output` | 最终表达的解析状态和摘要 |
| `tool_trace_summary` | 本次校验使用的工具调用导航索引 |
ChatService 根据 verdict 决定: ChatService 根据 verdict 决定:
- `PASS`:输出 Executor 答案。 - `PASS`:把允许表达的 claims 交给 Composer 输出。
- `LOW_CONFID`:必要时构造 `retry_context` 补证据;否则输出低置信提示。 - `LOW_CONFID`:必要时构造 `retry_context` 补证据;否则输出低置信提示。
- `REJECT`:降级输出,只保留已确认信息。 - `REJECT`:降级输出,只保留已确认信息。
边界:
- `executor_final_answer` 只作为 debug/fallback 上下文;结构化输出有效时,Verifier 不得从中抽取额外确认事实。
- `$.no_evidence` 只能表达“当前查询未检索到匹配证据”,不能表达“已排除/确认没有”。
## 6. AIOps 规则自评估 ## 6. AIOps 规则自评估
AIOps 当前使用 `AiOpsRuleEvaluationService`,结果写入 `aiops_rule_evaluation`。 AIOps 当前使用 `AiOpsRuleEvaluationService`,结果写入 `aiops_rule_evaluation`。
@@ -238,14 +269,14 @@ Trace API 会展示:
近期优先: 近期优先:
1. 将 `rule_evaluation` 与 `verifier_evaluation` 在 Trace API 中结构化展示。 1. 将 `rule_evaluation` 与 `verifier_evaluation` 在 Trace API 中结构化展示。
2. `not_useful` 反馈沉淀 bad case,而不是只写字段。 2. 将 ISS-008 / ISS-009 这类 E2E 通过样例固化进 diagnosis eval fixtures。
3. useful 案例自动提取 faultCategory、errorCode、service、rootCause、solution。 3. `not_useful` 反馈沉淀 bad case,而不是只写字段。
4. AIOps 引入 LLM Verifier。 4. useful 案例自动提取 faultCategory、errorCode、service、rootCause、solution。
5. 把反馈和 eval baseline 打通,形成可回归的质量改进闭环。 5. AIOps 引入 LLM Verifier。
6. 把反馈和 eval baseline 打通,形成可回归的质量改进闭环。
暂不优先: 暂不优先:
- 用用户反馈直接修改 session status。 - 用用户反馈直接修改 session status。
- 仅凭 `evidence_score` 判断答案正确。 - 仅凭 `evidence_score` 判断答案正确。
- 在没有人工审核时自动把 bad case 反向写入 Prompt。 - 在没有人工审核时自动把 bad case 反向写入 Prompt。
+48 -14
View File
@@ -1,6 +1,6 @@
# Harness 与质量门禁架构 # Harness 与质量门禁架构
**更新日期**:2026-07-05 **更新日期**:2026-07-08
**状态**:当前可运行架构 + 后续门禁规划 **状态**:当前可运行架构 + 后续门禁规划
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md` **参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
@@ -21,6 +21,7 @@ Prompt contract
+ Tool boundary + Tool boundary
+ Agent hooks + Agent hooks
+ Trace persistence + Trace persistence
+ Gatekeeper deterministic validation
+ Verifier / rule evaluation + Verifier / rule evaluation
+ Eval baseline + Eval baseline
``` ```
@@ -30,15 +31,19 @@ Prompt contract
```mermaid ```mermaid
flowchart TB flowchart TB
Input["User / AIOps input"] --> Prompt["Prompt contract"] Input["User / AIOps input"] --> Prompt["Prompt contract"]
Prompt --> Agent["Planner / Executor / Verifier"] Prompt --> Agent["Planner / Executor / Verifier / Composer"]
Agent --> Tools["Evidence tools"] Agent --> Tools["Evidence tools"]
Tools --> Invocation["tool_invocation"] Tools --> Invocation["tool_invocation"]
Agent --> StepHook["AgentLoggingHook"] Agent --> StepHook["AgentLoggingHook"]
StepHook --> Step["agent_step"] StepHook --> Step["agent_step"]
Agent --> Session["diagnosis_session"] Agent --> Session["diagnosis_session"]
Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
Agent --> Gatekeeper
Invocation --> TraceSummary["ToolTraceSummaryService"] Invocation --> TraceSummary["ToolTraceSummaryService"]
TraceSummary --> Verifier["chat_verifier"] Gatekeeper --> Verifier["chat_verifier"]
TraceSummary --> Verifier
Verifier --> SelfEval["self_evaluation.verifier_evaluation"] Verifier --> SelfEval["self_evaluation.verifier_evaluation"]
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"] Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
@@ -63,15 +68,18 @@ flowchart TB
| `planner-prompt.md` | AIOps Planner 规划、再规划、输出告警报告 | | `planner-prompt.md` | AIOps Planner 规划、再规划、输出告警报告 |
| `executor-prompt.md` | AIOps Executor 按步骤调用工具 | | `executor-prompt.md` | AIOps Executor 按步骤调用工具 |
| `chat-planner-prompt.md` | Chat 复杂问题规划 | | `chat-planner-prompt.md` | Chat 复杂问题规划 |
| `chat-executor-prompt.md` | Chat 执行工具并形成诊断答复 | | `chat-executor-prompt.md` | Chat 执行工具并输出 `executor_evidence_v2` 微观事实 |
| `chat-verifier-prompt.md` | 校验 Executor 答案是否被工具证据支撑 | | `chat-verifier-prompt.md` | 基于 Gatekeeper 已验真的证据判断 claims 是否可推出 |
| `chat-composer-prompt.md` | 基于 Verifier 允许表达的内容生成最终用户答复 |
Prompt 层当前承担的门禁: Prompt 层当前承担的门禁:
- 禁止凭记忆回答错误码、接口定义、排障步骤。 - 禁止凭记忆回答错误码、接口定义、排障步骤。
- 需要外部信息时必须调用工具。 - 需要外部信息时必须调用工具。
- 工具连续失败或返回空结果时,最终报告必须诚实说明。 - 工具连续失败或返回空结果时,最终报告必须诚实说明。
- Chat Verifier 不允许做新检索,只能校验已有证据。 - Chat Executor 不允许在窄范围问题中扩展根因、风险或修复建议。
- Chat Verifier 不允许做新检索,只能判断已验真证据是否可推出 claims。
- Chat Composer 不允许补事实,尤其不能把 `$.no_evidence` 表达为“已排除/确认没有”。
- AIOps payload 模式必须聚焦输入告警。 - AIOps payload 模式必须聚焦输入告警。
## 4. Trace Hooks ## 4. Trace Hooks
@@ -115,6 +123,7 @@ retrieval_layer
l0_match_count l0_match_count
l1_match_count l1_match_count
retrieval_details retrieval_details
-> evidence_refs
relevance_level relevance_level
dedup_reason dedup_reason
duration_ms duration_ms
@@ -128,23 +137,44 @@ error_message
- 检索结果归一化为 `PRECISE`、`HIGHLY_RELEVANT`、`REFERENCE`。 - 检索结果归一化为 `PRECISE`、`HIGHLY_RELEVANT`、`REFERENCE`。
- 同 session 内重复文档会被 `RetrievedDocTracker` 去重。 - 同 session 内重复文档会被 `RetrievedDocTracker` 去重。
- dedup、no evidence、failed 等状态进入 `retrieval_details.evidence_status`。 - dedup、no evidence、failed 等状态进入 `retrieval_details.evidence_status`。
- `retrieval_details.evidence_refs` 记录可被 Executor 引用的最小证据文本,格式为 `raw_path + text`。
- no-hit / no-evidence 工具结果会生成 `raw_path=$.no_evidence` 的负向证据引用,语义仅限“本次查询未检索到匹配证据”。
## 6. Verifier 门禁 ## 6. Gatekeeper 与 Verifier 门禁
Chat Verifier 的输入不是原始工具日志,而是 `ToolTraceSummaryService` 构造的证据索引。 Chat Verifier 前置一层 Gatekeeper。Gatekeeper 不调用 LLM,只用代码检查 Executor 输出的证据引用是否真实存在。
```mermaid ```mermaid
flowchart LR flowchart LR
Invocation["tool_invocation"] --> Summary["ToolTraceSummaryService"] Invocation["tool_invocation"] --> EvidenceRefs["retrieval_details.evidence_refs"]
ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
EvidenceRefs --> Gatekeeper
Gatekeeper --> GateResult["gatekeeper_result"]
Invocation --> Summary["ToolTraceSummaryService"]
Summary --> EvidenceIndex["tool_trace_summary"] Summary --> EvidenceIndex["tool_trace_summary"]
EvidenceIndex --> Verifier["chat_verifier"] GateResult --> Verifier["chat_verifier"]
ExecutorAnswer["executor_final_answer"] --> Verifier ExecutorOutput --> Verifier
EvidenceIndex --> Verifier
Verifier --> Verdict{"verdict"} Verifier --> Verdict{"verdict"}
Verdict -->|PASS| Pass["输出原答案"] Verdict -->|PASS| Composer["chat_composer"]
Composer --> Pass["输出最终答复"]
Verdict -->|LOW_CONFID| Low["补证据或低置信输出"] Verdict -->|LOW_CONFID| Low["补证据或低置信输出"]
Verdict -->|REJECT| Reject["降级输出"] Verdict -->|REJECT| Reject["降级输出"]
``` ```
Gatekeeper 检查:
| 检查 | 失败语义 |
|---|---|
| `answer_version=executor_evidence_v2` | 非结构化或旧结构输出降为低置信 |
| `source_invocation_id` 真实存在 | 伪造 ID 直接拒绝 |
| `tool_name` 与 invocation 对齐 | 张冠李戴直接拒绝 |
| `raw_path` 存在于 `evidence_refs` | 无中生有直接拒绝 |
| `evidence_excerpt` 由 `evidence_refs[].text` 支撑 | excerpt 编造或错配直接拒绝 |
| `negative_observation` 只能引用 `$.no_evidence` | 用正向日志证明“没查到”直接拒绝 |
Gatekeeper 审计还会记录 `rule_set_version` 和已启用规则元数据摘要。当前规则元数据来自本地 `gatekeeper-rules.json`,规则执行仍是确定性 Java 代码。
Verifier 输出: Verifier 输出:
```json ```json
@@ -152,17 +182,22 @@ Verifier 输出:
"verdict": "PASS|LOW_CONFID|REJECT", "verdict": "PASS|LOW_CONFID|REJECT",
"groundedness_score": 0.8, "groundedness_score": 0.8,
"critical_fact_count": 2, "critical_fact_count": 2,
"claim_checks": [],
"facts_checked": [], "facts_checked": [],
"rationale": "..." "rationale": "..."
} }
``` ```
Verifier 不再逐字核验 excerpt 真伪;这由 Gatekeeper 完成。Verifier 只回答一个问题:`claim_text` 是否能由已经验真的 `evidence_excerpt` 推导出来。
结果写入: 结果写入:
```text ```text
diagnosis_session.self_evaluation.verifier_evaluation diagnosis_session.self_evaluation.verifier_evaluation
``` ```
其中同时持久化 `executor_structured_output`、`gatekeeper_result`、`tool_trace_summary` 和 `composer_output`,用于 Trace 回放。
## 7. AIOps 规则门禁 ## 7. AIOps 规则门禁
AIOps 当前不走 Chat Verifier,而是用 `AiOpsRuleEvaluationService` 做轻量检查。 AIOps 当前不走 Chat Verifier,而是用 `AiOpsRuleEvaluationService` 做轻量检查。
@@ -197,9 +232,8 @@ diagnosis_session.self_evaluation.aiops_rule_evaluation
- 工具参数 schema 校验。 - 工具参数 schema 校验。
- 同一工具调用次数上限。 - 同一工具调用次数上限。
- 工具超时的统一熔断。 - 工具超时的统一熔断。
- 报告中的数值与工具返回值自动对齐校验。 - Gatekeeper 规则远程化或三层分离:索引层、元数据层、规则实现层。
- Prompt 版本记录和回滚。 - Prompt 版本记录和回滚。
- Verifier 对 AIOps 报告的 LLM 级事实校验。 - Verifier 对 AIOps 报告的 LLM 级事实校验。
这些应在评测集扩大后逐步加入,避免一次性把诊断流程卡得过死。 这些应在评测集扩大后逐步加入,避免一次性把诊断流程卡得过死。
+8 -9
View File
@@ -5,7 +5,7 @@
## 1. 一句话 ## 1. 一句话
SuperBizAgent 是一个面向企业故障诊断的可追踪 Agent 系统:它把用户问题或 AIOps 告警转换成 Planner、Executor、Verifier 的诊断链路,所有工具证据、模型步骤、最终答案、自评估和用户反馈都能通过同一个 `sessionId` 回放。 SuperBizAgent 是一个面向企业故障诊断的可追踪 Agent 系统:它把用户问题或 AIOps 告警转换成 Planner、Executor、Gatekeeper、Verifier、Composer 的诊断链路,所有工具证据、模型步骤、最终答案、自评估和用户反馈都能通过同一个 `sessionId` 回放。
## 2. 一张图 ## 2. 一张图
@@ -16,7 +16,7 @@ flowchart TB
API --> Chat["ChatService"] API --> Chat["ChatService"]
API --> AiOps["AiOpsService"] API --> AiOps["AiOpsService"]
Chat --> ChatFlow["Chat: Planner -> Executor -> Verifier"] Chat --> ChatFlow["Chat: Planner -> Executor -> Gatekeeper -> Verifier -> Composer"]
AiOps --> AiOpsFlow["AIOps: Supervisor -> Planner / Executor"] AiOps --> AiOpsFlow["AIOps: Supervisor -> Planner / Executor"]
ChatFlow --> Tools["Evidence Tools"] ChatFlow --> Tools["Evidence Tools"]
@@ -55,14 +55,14 @@ flowchart TB
```text ```text
这个项目不是把问题直接丢给大模型,而是把诊断拆成可审计的执行链路。 这个项目不是把问题直接丢给大模型,而是把诊断拆成可审计的执行链路。
Chat 复杂问题走 Planner -> Executor -> Verifier: Chat 复杂问题走 Planner -> Executor -> Gatekeeper -> Verifier -> Composer:
Planner 负责拆解,Executor 负责调用知识库、日志和指标工具,Verifier 只基于已有工具证据校验最终答案。 Planner 负责拆解,Executor 只负责调用知识库、日志和指标工具并提炼带证据引用的微观事实;Gatekeeper 用代码核对 invocation、raw_path 和 excerpt 是否真实;Verifier 判断这些事实能否由已验真的证据推出;Composer 只把允许表达的结论写成最终答案。
AIOps 告警入口走 Supervisor 调度 Planner/Executor: AIOps 告警入口走 Supervisor 调度 Planner/Executor:
如果请求里有 alert payload,系统会进入 PAYLOAD_TARGETED 模式,报告必须聚焦这个告警,而不是被当前环境中的其他活跃告警带偏。 如果请求里有 alert payload,系统会进入 PAYLOAD_TARGETED 模式,报告必须聚焦这个告警,而不是被当前环境中的其他活跃告警带偏。
所有过程都会落到 diagnosis_session、agent_step、tool_invocation。 所有过程都会落到 diagnosis_session、agent_step、tool_invocation。
所以我可以用一个 sessionId 回放:模型怎么规划、调了哪些工具、工具返回什么、Verifier 怎么判定、用户最后是否反馈有用。 所以我可以用一个 sessionId 回放:模型怎么规划、调了哪些工具、工具返回什么、Gatekeeper 怎么验真、Verifier 怎么判定、Composer 最后怎么表达、用户最后是否反馈有用。
``` ```
## 4. 五个亮点 ## 4. 五个亮点
@@ -72,7 +72,7 @@ AIOps 告警入口走 Supervisor 调度 Planner/Executor:
| 可追踪 Agent | 每次诊断都有 `sessionId`,Trace API 可以回放 session、step、tool | | 可追踪 Agent | 每次诊断都有 `sessionId`,Trace API 可以回放 session、step、tool |
| 显式工具证据链 | `lookup_knowledge`、日志、指标都记录到 `tool_invocation` | | 显式工具证据链 | `lookup_knowledge`、日志、指标都记录到 `tool_invocation` |
| RAG 工程化 | L0 降级为 hint,Spring AI VectorStore 做主检索,SDK fallback 保底 | | RAG 工程化 | L0 降级为 hint,Spring AI VectorStore 做主检索,SDK fallback 保底 |
| 质量门禁 | Chat Verifier 校验 groundedness,AIOps rule evaluation 控制告警聚焦 | | 质量门禁 | Chat Gatekeeper 验引用、Verifier 判可推导、Composer 控表达,AIOps rule evaluation 控制告警聚焦 |
| 反馈闭环 | useful 反馈沉淀 `case_library`,not_useful 保留 bad case 信号 | | 反馈闭环 | useful 反馈沉淀 `case_library`,not_useful 保留 bad case 信号 |
## 5. 三个关键取舍 ## 5. 三个关键取舍
@@ -99,15 +99,14 @@ aiops_rule_evaluation -> AIOps 报告是否聚焦告警并使用证据
| 追问 | 回答方向 | | 追问 | 回答方向 |
|---|---| |---|---|
| 怎么防止幻觉? | Executor 必须用工具;Verifier 只基于 `tool_trace_summary` 校验;LOW_CONFID/REJECT 会降级输出 | | 怎么防止幻觉? | Executor 输出 `executor_evidence_v2`,每个 claim 绑定 `source_invocation_id + raw_path + evidence_excerpt`;Gatekeeper 用 `tool_invocation.retrieval_details.evidence_refs` 核验引用真实性;Verifier 只判断可推导性;Composer 防止把 no-evidence 说成已排除 |
| RAG 质量怎么保证? | offline golden cases + live acceptance + trace inspection 三层验证 | | RAG 质量怎么保证? | offline golden cases + live acceptance + trace inspection 三层验证 |
| 为什么 L0 不直接返回? | L0 子串命中不等于语义相关,当前只做 domain/entity hint 和 metadata filter | | 为什么 L0 不直接返回? | L0 子串命中不等于语义相关,当前只做 domain/entity hint 和 metadata filter |
| AIOps 如何避免跑偏? | payload 模式生成 recommended query,并用 rule evaluation 检查报告聚焦输入告警 | | AIOps 如何避免跑偏? | payload 模式生成 recommended query,并用 rule evaluation 检查报告聚焦输入告警 |
| 下一步怎么演进? | evidence block、邻居 chunk、Playbook、AIOps LLM Verifier、MCP 工具协议化 | | 下一步怎么演进? | 固化 E2E fixture、Prompt version、Gatekeeper 规则配置化、邻居 chunk、AIOps LLM Verifier、MCP 工具协议化 |
## 7. 现场演示入口 ## 7. 现场演示入口
- Demo 脚本:`mvp/demo/ten-minute-interview-demo.md` - Demo 脚本:`mvp/demo/ten-minute-interview-demo.md`
- 故事案例:`interview/story-cases.md` - 故事案例:`interview/story-cases.md`
- 架构细节:`mvp/architecture/README.md` - 架构细节:`mvp/architecture/README.md`
+19 -1
View File
@@ -141,7 +141,8 @@ post-retrieval 层再把检索候选归一为:
- 给 Agent 输出 completeness hint。 - 给 Agent 输出 completeness hint。
- 写入 `tool_invocation.relevance_level`。 - 写入 `tool_invocation.relevance_level`。
- 给 Verifier 构造 `tool_trace_summary`。 - 给 Gatekeeper 提供 `evidence_refs` 引用验真源。
- 给 Verifier 构造 `tool_trace_summary` 审计导航。
- 供 EvaluationService 计算 evidence score。 - 供 EvaluationService 计算 evidence score。
## 6. 文档切片和 metadata ## 6. 文档切片和 metadata
@@ -178,6 +179,7 @@ flowchart LR
LookupResult --> Recorder["ToolInvocationRecorder"] LookupResult --> Recorder["ToolInvocationRecorder"]
Recorder --> Invocation["tool_invocation"] Recorder --> Invocation["tool_invocation"]
Invocation --> Trace["DiagnosisTraceService"] Invocation --> Trace["DiagnosisTraceService"]
Invocation --> Gatekeeper["ExecutorGatekeeperService"]
Invocation --> Summary["ToolTraceSummaryService"] Invocation --> Summary["ToolTraceSummaryService"]
Summary --> Verifier["chat_verifier"] Summary --> Verifier["chat_verifier"]
Invocation --> Eval["EvaluationService / RAG eval"] Invocation --> Eval["EvaluationService / RAG eval"]
@@ -205,9 +207,25 @@ success
- evidence status。 - evidence status。
- dedup reason。 - dedup reason。
- evidence block summaries。 - evidence block summaries。
- evidence refs:`raw_path + text`,用于核对 Executor 的 `evidence_excerpt`。
- context pack summary。 - context pack summary。
- rerank trace。 - rerank trace。
其中 `evidence_refs` 是当前 Chat 证据链路的精确引用源:
```json
{
"evidence_refs": [
{
"raw_path": "$.evidence_blocks[0]",
"text": "最小证据文本"
}
]
}
```
如果检索返回 no evidence,应使用 `raw_path=$.no_evidence` 记录负向证据。它只能说明“本次检索没有匹配证据”,不能作为“问题不存在”的证明。
## 8. 去重与行动记忆 ## 8. 去重与行动记忆
当前 session 级去重由 `RetrievedDocTracker` 负责。 当前 session 级去重由 `RetrievedDocTracker` 负责。
+19 -5
View File
@@ -1,6 +1,6 @@
# 会话与 Trace 生命周期 # 会话与 Trace 生命周期
**更新日期**:2026-07-05 **更新日期**:2026-07-08
**状态**:当前可运行架构 **状态**:当前可运行架构
**参考历史文档**:`archive/2026-07-05-legacy/session-management.md` **参考历史文档**:`archive/2026-07-05-legacy/session-management.md`
@@ -35,6 +35,7 @@ flowchart TD
StepHook --> Step["agent_step"] StepHook --> Step["agent_step"]
Agent --> Tool["Evidence tools"] Agent --> Tool["Evidence tools"]
Tool --> Invocation["tool_invocation"] Tool --> Invocation["tool_invocation"]
Invocation --> Gatekeeper["Gatekeeper evidence validation"]
Agent --> Final{"workflow result"} Agent --> Final{"workflow result"}
Final -->|success| Success["status = SUCCESS, answer saved"] Final -->|success| Success["status = SUCCESS, answer saved"]
@@ -132,6 +133,7 @@ retrieval_layer
l0_match_count l0_match_count
l1_match_count l1_match_count
retrieval_details retrieval_details
-> evidence_refs
relevance_level relevance_level
dedup_reason dedup_reason
duration_ms duration_ms
@@ -139,7 +141,20 @@ success
error_message error_message
``` ```
对 `lookup_knowledge`,`retrieval_details` 会承载 L0/L1、领域、证据状态、去重等检索细节。对非检索工具,检索字段可以为空。 对 `lookup_knowledge`,`retrieval_details` 会承载 L0/L1、领域、证据状态、去重等检索细节。对日志、指标和知识库工具,`retrieval_details.evidence_refs` 会记录 Gatekeeper 可核验的最小证据引用:
```json
{
"evidence_refs": [
{
"raw_path": "$.logs[0]",
"text": "最小证据文本"
}
]
}
```
当工具明确没有返回匹配证据时,可以记录 `raw_path=$.no_evidence`。该路径只表示“本次工具查询未检索到匹配证据”,不表示问题被排除。
## 7. Trace API 聚合 ## 7. Trace API 聚合
@@ -163,7 +178,7 @@ Trace 视图回答的问题:
- 每一步模型输入输出是什么摘要? - 每一步模型输入输出是什么摘要?
- 调用了哪些工具? - 调用了哪些工具?
- 工具返回了什么证据? - 工具返回了什么证据?
- Verifier / AIOps rule 是否通过? - Gatekeeper / Verifier / Composer / AIOps rule 是否通过?
- 用户是否反馈有用? - 用户是否反馈有用?
## 8. Chat 与 AIOps 差异 ## 8. Chat 与 AIOps 差异
@@ -171,7 +186,7 @@ Trace 视图回答的问题:
| 维度 | Chat | AIOps | | 维度 | Chat | AIOps |
|---|---|---| |---|---|---|
| `agent_flow` | `CHAT` | `AI_OPS` | | `agent_flow` | `CHAT` | `AI_OPS` |
| 编排方式 | `SequentialAgent`: Planner -> Executor -> Verifier | `SupervisorAgent`: Planner + Executor | | 编排方式 | `SequentialAgent`: Planner -> Executor -> Gatekeeper -> Verifier -> Composer | `SupervisorAgent`: Planner + Executor |
| 自评估 | `rule_evaluation` + `verifier_evaluation` | `aiops_rule_evaluation` | | 自评估 | `rule_evaluation` + `verifier_evaluation` | `aiops_rule_evaluation` |
| 答案字段 | Chat 最终答复 | 告警分析报告 | | 答案字段 | Chat 最终答复 | 告警分析报告 |
| payload | 用户自然语言 + history | alert payload 或 auto-discovery | | payload | 用户自然语言 + history | alert payload 或 auto-discovery |
@@ -193,4 +208,3 @@ Trace 视图回答的问题:
2. `agent_step` 与 `tool_invocation.step_id` 建立更严格关联。 2. `agent_step` 与 `tool_invocation.step_id` 建立更严格关联。
3. 对多轮同 session 诊断增加 run id,避免复用 session 时历史记录混杂。 3. 对多轮同 session 诊断增加 run id,避免复用 session 时历史记录混杂。
4. 为 Trace 增加导出能力,服务面试演示和回归分析。 4. 为 Trace 增加导出能力,服务面试演示和回归分析。
+14
View File
@@ -6,9 +6,13 @@
- `ten-minute-interview-demo.md`:10 分钟现场演示脚本。 - `ten-minute-interview-demo.md`:10 分钟现场演示脚本。
- `interview-walkthrough.md`:面试讲解话术。 - `interview-walkthrough.md`:面试讲解话术。
- `evidence-pipeline-scenarios.md`:PASS / LOW_CONFID / REJECT / no-evidence 场景矩阵。
- `trace-inspection-checklist.md`:Trace 字段检查清单。 - `trace-inspection-checklist.md`:Trace 字段检查清单。
- `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。 - `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。
- `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。 - `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。
- `requests/narrow-highcpu-chat.json`:窄范围正向观察请求。
- `requests/hikari-no-evidence-chat.json`:no-evidence 负向观察请求。
- `requests/safety-unsupported-claim-chat.json`:安全降级讨论请求。
## 1. 前置条件 ## 1. 前置条件
@@ -167,3 +171,13 @@ AIOps 主线:
-> AIOps rule evaluation -> AIOps rule evaluation
-> Trace API 回放 -> Trace API 回放
``` ```
## 8. Evidence Pipeline 场景矩阵
面试时不要把所有安全场景都压到 live LLM 现场表现上。建议使用:
- `scripts/run-payment-timeout-demo.ps1` 跑主路径。
- `evidence-pipeline-scenarios.md` 讲解 PASS / LOW_CONFID / REJECT / no-evidence 矩阵。
- `mvp/eval/reports/baseline-report.md` 证明固定 fixture 10/10 通过。
这样可以同时展示真实链路和确定性回归能力。
+63
View File
@@ -0,0 +1,63 @@
# Evidence Pipeline Demo Scenarios
这份清单用于面试时说明 Chat 证据链路如何覆盖 `PASS`、`LOW_CONFID`、`REJECT` 和 no-evidence 场景。
重点区别:
- Live demo 证明本地服务、工具、Trace、Feedback 主链路能跑通。
- Fixture-backed eval 证明固定安全场景可以确定性回归,不依赖 LLM 当场随机输出。
## Scenario Matrix
| 场景 | 类型 | 输入/证据 | 期望讲点 |
|---|---|---|---|
| Payment timeout | Live 主路径 | `requests/payment-timeout-chat.json` | 完整 Chat -> Trace -> Feedback 闭环 |
| Narrow HighCPU observation | Live 可尝试 + fixture-backed | `requests/narrow-highcpu-chat.json` / `mvp/eval/fixtures/narrow-highcpu-observation-pass.json` | Executor 只输出观察类 claim,Gatekeeper 验引用,Verifier PASS |
| Hikari no-evidence | Live 可尝试 + fixture-backed | `requests/hikari-no-evidence-chat.json` / `mvp/eval/fixtures/hikari-no-evidence-negative-observation-pass.json` | `$.no_evidence` 只表示本次查询无匹配证据,Composer 不说“已排除” |
| Unsupported claim filtering | Fixture-backed | `requests/safety-unsupported-claim-chat.json` / `mvp/eval/fixtures/unsupported-claim-filtering-low-confid.json` | Verifier 将 unsupported claim 降为 LOW_CONFID,最终答案不确认“主库故障” |
| Fabricated invocation reject | Fixture-backed | `mvp/eval/fixtures/gatekeeper-fabricated-invocation-reject.json` | Gatekeeper 拦截伪造 invocation,最终 REJECT/降级 |
| Composer fallback | Fixture-backed | `mvp/eval/fixtures/composer-fallback-no-raw-json-low-confid.json` | 即使 Composer 输出异常,也不能把 Executor JSON 泄漏给用户 |
## Trace Fields To Inspect
| 能力 | JSON path |
|---|---|
| Executor V2 输出 | `data.session.selfEvaluation.verifier_evaluation.executor_structured_output` |
| Gatekeeper 结果 | `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.status` |
| Gatekeeper 规则版本 | `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` |
| 证据绑定校验 | `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.checked_bindings` |
| Verifier claim checks | `data.session.selfEvaluation.verifier_evaluation.claim_checks` |
| Composer 输出 | `data.session.selfEvaluation.verifier_evaluation.composer_output` |
| 工具证据引用 | `data.toolInvocations[*].retrievalDetails.evidence_refs` |
## How To Present It
```text
我把现场 demo 和固定 eval 分开。
现场 demo 证明系统能跑通真实链路;
fixture-backed eval 证明反幻觉安全场景可以稳定回归。
Gatekeeper 的规则版本也进入 trace,所以后续调整阈值或规则时可以审计。
```
## Optional Live Requests
手动发送某个请求样例:
```powershell
$body = Get-Content -Raw -Encoding UTF8 "mvp/demo/requests/narrow-highcpu-chat.json"
Invoke-RestMethod `
-Method Post `
-Uri "http://localhost:9900/api/chat" `
-ContentType "application/json" `
-Body $body
```
然后查询同一 session:
```powershell
Invoke-RestMethod `
-Method Get `
-Uri "http://localhost:9900/api/diagnosis/mvp-demo-narrow-highcpu-001/trace"
```
注意:除 payment-timeout 主路径外,其它 live 请求是“可尝试”的演示入口;稳定验收以 `mvp/eval` fixture 和 baseline 为准。
+9 -1
View File
@@ -24,6 +24,7 @@
4. 打开 `mvp/demo/output/trace-response.json`。 4. 打开 `mvp/demo/output/trace-response.json`。
5. 指出证据工具和 verifier evaluation。 5. 指出证据工具和 verifier evaluation。
6. 提交 feedback,并展示它挂在同一个 session 上。 6. 提交 feedback,并展示它挂在同一个 session 上。
7. 打开 `evidence-pipeline-scenarios.md`,说明 PASS / LOW_CONFID / REJECT / no-evidence 的固定回归矩阵。
## 3. 命令 ## 3. 命令
@@ -125,6 +126,14 @@ Demo 证明真实链路能跑通,offline eval baseline 证明固定 case 可
这两者分开是有意的:Demo 面向人类审阅,eval 面向自动化信号。 这两者分开是有意的:Demo 面向人类审阅,eval 面向自动化信号。
``` ```
如果被问到怎么防止证据归因幻觉,可以补充:
```text
Executor 的 claim 必须绑定 source_invocation_id、raw_path 和 evidence_excerpt。
Gatekeeper 用代码核验这些引用,并把 rule_set_version 写进 trace。
Verifier 只判断已核验证据能否推出 claim,Composer 只表达允许输出的内容。
```
## 5. 强面试表达 ## 5. 强面试表达
```text ```text
@@ -141,4 +150,3 @@ traceability、evidence persistence、verifier gating、feedback 和 regression
mvp-demo profile mock 了日志和指标,但不是完整生产运行环境。 mvp-demo profile mock 了日志和指标,但不是完整生产运行环境。
密钥清理、默认隔离测试和生产可靠性是后续 hardening 工作。 密钥清理、默认隔离测试和生产可靠性是后续 hardening 工作。
``` ```
@@ -0,0 +1,4 @@
{
"Id": "mvp-demo-hikari-no-evidence-001",
"Question": "确认 inventory-service 当前是否有 HikariCP 连接池耗尽日志。"
}
@@ -0,0 +1,4 @@
{
"Id": "mvp-demo-narrow-highcpu-001",
"Question": "确认 payment-service 当前是否存在 HighCPUUsage 告警。"
}
@@ -0,0 +1,4 @@
{
"Id": "mvp-demo-safety-unsupported-001",
"Question": "订单超时是否可以确认由数据库主库故障导致?请只基于当前证据回答。"
}
+2
View File
@@ -10,6 +10,7 @@
| `data.session.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 | | `data.session.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
| `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace | | `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
| `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 | | `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
| `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | 如果是 Chat V2 链路,是否记录 Gatekeeper 规则版本 | 安全规则可审计、可回归 |
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在同一次诊断上 | | `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在同一次诊断上 |
## 2. Agent 步骤 ## 2. Agent 步骤
@@ -30,6 +31,7 @@
| `data.toolInvocations[*].outputPreview` | 是否有受控长度的证据预览 | 保留证据但不倾倒巨大 payload | | `data.toolInvocations[*].outputPreview` | 是否有受控长度的证据预览 | 保留证据但不倾倒巨大 payload |
| `data.toolInvocations[*].success` | 是否区分成功和失败 | 工具失败对 Verifier 和 reviewer 可见 | | `data.toolInvocations[*].success` | 是否区分成功和失败 | 工具失败对 Verifier 和 reviewer 可见 |
| `data.toolInvocations[*].retrievalDetails` | 是否包含检索 metadata | 检索质量可事后检查 | | `data.toolInvocations[*].retrievalDetails` | 是否包含检索 metadata | 检索质量可事后检查 |
| `data.toolInvocations[*].retrievalDetails.evidence_refs` | 是否包含 `raw_path + text` | Gatekeeper 可以用代码核对 Executor 引用 |
| `data.toolInvocations[*].relevanceLevel` | 是否有相关性等级 | 可解释检索结果强弱 | | `data.toolInvocations[*].relevanceLevel` | 是否有相关性等级 | 可解释检索结果强弱 |
## 4. Summary ## 4. Summary
+8 -4
View File
@@ -29,18 +29,21 @@ The baseline evaluates saved trace fixtures. It does not start the application a
The committed baseline currently contains: The committed baseline currently contains:
```text ```text
8 fixed cases 10 fixed cases
8 passing fixture evaluations 10 passing fixture evaluations
2 PASS verdicts 4 PASS verdicts
5 LOW_CONFID verdicts 5 LOW_CONFID verdicts
1 REJECT verdict 1 REJECT verdict
``` ```
The three V2 audit-closure cases cover: The V2 evidence-pipeline matrix covers:
- Positive supported evidence for a narrow HighCPU observation.
- No-evidence `negative_observation` using `$.no_evidence`.
- Gatekeeper failure for a fabricated tool invocation reference. - Gatekeeper failure for a fabricated tool invocation reference.
- Unsupported claim filtering before the final answer. - Unsupported claim filtering before the final answer.
- Composer fallback rendering without raw Executor JSON leakage. - Composer fallback rendering without raw Executor JSON leakage.
- Gatekeeper rule set version audit for new matrix fixtures.
## Verification ## Verification
@@ -76,6 +79,7 @@ Stage 5 adds these V2 checks:
- Required V2 fixtures must include `gatekeeper_result`, `claim_checks`, and `composer_output`. - Required V2 fixtures must include `gatekeeper_result`, `claim_checks`, and `composer_output`.
- `claim_checks` must be structurally auditable. - `claim_checks` must be structurally auditable.
- Composer output must record whether normal parsing or fallback rendering was used. - Composer output must record whether normal parsing or fallback rendering was used.
- Gatekeeper rule set version can be asserted per fixture.
- Final answers must not leak raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`. - Final answers must not leak raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`.
- Configured unsupported claim keywords must not appear as confirmed final-answer content. - Configured unsupported claim keywords must not appear as confirmed final-answer content.
+36
View File
@@ -1,4 +1,40 @@
[ [
{
"id": "narrow-highcpu-observation",
"title": "Narrow HighCPU observation",
"question": "确认 payment-service 当前是否存在 HighCPUUsage 告警。",
"traceFixture": "narrow-highcpu-observation-pass.json",
"expectedRootCauseKeywords": ["payment-service", "HighCPUUsage", "92%"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["query_metrics"],
"allowedVerdicts": ["PASS"],
"forbiddenAnswerKeywords": ["根因", "修复建议", "通常情况下"],
"requireV2AuditClosure": true,
"requireClaimChecks": true,
"requireComposerOutput": true,
"expectedGatekeeperStatuses": ["pass"],
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
"expectedComposerStatuses": ["valid"],
"forbiddenConfirmedClaimKeywords": ["数据库连接池"]
},
{
"id": "hikari-no-evidence-negative-observation",
"title": "Hikari no-evidence negative observation",
"question": "确认 inventory-service 当前是否有 HikariCP 连接池耗尽日志。",
"traceFixture": "hikari-no-evidence-negative-observation-pass.json",
"expectedRootCauseKeywords": ["未检索到", "HikariCP", "匹配证据"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["query_logs"],
"allowedVerdicts": ["PASS"],
"forbiddenAnswerKeywords": ["已排除", "确认没有", "日志层面已排除"],
"requireV2AuditClosure": true,
"requireClaimChecks": true,
"requireComposerOutput": true,
"expectedGatekeeperStatuses": ["pass"],
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
"expectedComposerStatuses": ["valid"],
"forbiddenConfirmedClaimKeywords": ["已排除 HikariCP"]
},
{ {
"id": "payment-timeout", "id": "payment-timeout",
"title": "Payment API timeout", "title": "Payment API timeout",
@@ -0,0 +1,126 @@
{
"session": {
"sessionId": "eval-hikari-no-evidence-negative-observation",
"query": "确认 inventory-service 当前是否有 HikariCP 连接池耗尽日志。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 21000,
"toolCallCount": 1,
"answer": "本次查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志;这只表示当前查询没有匹配证据,仍不能据此判断系统一定健康。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "PASS",
"groundedness_score": 1.0,
"critical_fact_count": 1,
"gatekeeper_result": {
"status": "pass",
"severity": "none",
"rule_set_version": "gatekeeper-rules-v1",
"rules": [
{
"id": "evidence.raw_path",
"description": "raw_path must exist in retrieval_details.evidence_refs",
"enabled": true,
"default_severity": "reject"
}
],
"checked_bindings": [
{
"claim_id": "claim-1",
"tool_name": "query_logs",
"source_invocation_id": 12,
"raw_path": "$.no_evidence",
"matched_text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志",
"status": "pass"
}
],
"failed_rules": [],
"warnings": [],
"errors": []
},
"executor_structured_output": {
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "negative_observation",
"claim_text": "当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"tool_name": "query_logs",
"source_invocation_id": 12,
"raw_path": "$.no_evidence",
"evidence_excerpt": "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": [
"仅查询了 application-logs 中 inventory-service HikariCP 相关日志"
]
},
"claim_checks": [
{
"claim_id": "claim-1",
"claim_text": "当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。",
"claim_type": "negative_observation",
"verification": "direct_observation",
"detail": "$.no_evidence 只支持本次查询未检索到匹配证据。",
"evidence_refs": [
{
"source_invocation_id": 12,
"raw_path": "$.no_evidence"
}
]
}
],
"facts_checked": [],
"composer_output": {
"status": "valid",
"answer_summary": "本次查询未检索到匹配日志。",
"recommended_actions": [],
"user_facing_answer": "本次查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志;这只表示当前查询没有匹配证据,仍不能据此判断系统一定健康。"
},
"tool_trace_summary": [
{
"tool_name": "query_logs",
"success": true,
"source_invocation_ids": [12],
"evidence_level": "no_evidence"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 12,
"sessionId": "eval-hikari-no-evidence-negative-observation",
"toolName": "query_logs",
"outputPreview": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志",
"retrievalDetails": {
"evidence_status": "no_evidence",
"evidence_refs": [
{
"raw_path": "$.no_evidence",
"text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志"
}
]
},
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 1,
"returnedToolCallCount": 1,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,124 @@
{
"session": {
"sessionId": "eval-narrow-highcpu-observation",
"query": "确认 payment-service 当前是否存在 HighCPUUsage 告警。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 18000,
"toolCallCount": 1,
"answer": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "PASS",
"groundedness_score": 1.0,
"critical_fact_count": 1,
"gatekeeper_result": {
"status": "pass",
"severity": "none",
"rule_set_version": "gatekeeper-rules-v1",
"rules": [
{
"id": "evidence.raw_path",
"description": "raw_path must exist in retrieval_details.evidence_refs",
"enabled": true,
"default_severity": "reject"
}
],
"checked_bindings": [
{
"claim_id": "claim-1",
"tool_name": "query_metrics",
"source_invocation_id": 11,
"raw_path": "$.alerts[0]",
"matched_text": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m",
"status": "pass"
}
],
"failed_rules": [],
"warnings": [],
"errors": []
},
"executor_structured_output": {
"answer_version": "executor_evidence_v2",
"claims": [
{
"claim_id": "claim-1",
"claim_type": "observation",
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
"support_level": "direct",
"evidence_bindings": [
{
"source_type": "tool_trace",
"tool_name": "query_metrics",
"source_invocation_id": 11,
"raw_path": "$.alerts[0]",
"evidence_excerpt": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m"
}
]
}
],
"hypotheses": [],
"recommended_actions": [],
"missing_info": []
},
"claim_checks": [
{
"claim_id": "claim-1",
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
"claim_type": "observation",
"verification": "direct_observation",
"detail": "已核验的指标证据直接包含服务名、告警名和 CPU 当前值。",
"evidence_refs": [
{
"source_invocation_id": 11,
"raw_path": "$.alerts[0]"
}
]
}
],
"facts_checked": [],
"composer_output": {
"status": "valid",
"answer_summary": "payment-service 当前存在 HighCPUUsage 告警。",
"recommended_actions": [],
"user_facing_answer": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。"
},
"tool_trace_summary": [
{
"tool_name": "query_metrics",
"success": true,
"source_invocation_ids": [11],
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 11,
"sessionId": "eval-narrow-highcpu-observation",
"toolName": "query_metrics",
"outputPreview": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m",
"retrievalDetails": {
"evidence_status": "supported",
"evidence_refs": [
{
"raw_path": "$.alerts[0]",
"text": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m"
}
]
},
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 1,
"returnedToolCallCount": 1,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
+47 -5
View File
@@ -1,15 +1,49 @@
{ {
"totalCases" : 8, "totalCases" : 10,
"passedCases" : 8, "passedCases" : 10,
"passRate" : 1.0, "passRate" : 1.0,
"verdictDistribution" : { "verdictDistribution" : {
"PASS" : 2, "PASS" : 4,
"LOW_CONFID" : 5, "LOW_CONFID" : 5,
"REJECT" : 1 "REJECT" : 1
}, },
"averageToolCallCount" : 1.625, "averageToolCallCount" : 1.5,
"averageDurationMs" : 44875.0, "averageDurationMs" : 39800.0,
"results" : [ { "results" : [ {
"caseId" : "narrow-highcpu-observation",
"title" : "Narrow HighCPU observation",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "PASS",
"matchedKeywordCount" : 3,
"requiredKeywordCount" : 3,
"evidenceCoverage" : {
"query_metrics" : true
},
"gatekeeperStatus" : "pass",
"gatekeeperRuleSetVersion" : "gatekeeper-rules-v1",
"composerStatus" : "valid",
"claimCheckCount" : 1,
"toolCallCount" : 1,
"durationMs" : 18000
}, {
"caseId" : "hikari-no-evidence-negative-observation",
"title" : "Hikari no-evidence negative observation",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "PASS",
"matchedKeywordCount" : 3,
"requiredKeywordCount" : 3,
"evidenceCoverage" : {
"query_logs" : true
},
"gatekeeperStatus" : "pass",
"gatekeeperRuleSetVersion" : "gatekeeper-rules-v1",
"composerStatus" : "valid",
"claimCheckCount" : 1,
"toolCallCount" : 1,
"durationMs" : 21000
}, {
"caseId" : "payment-timeout", "caseId" : "payment-timeout",
"title" : "Payment API timeout", "title" : "Payment API timeout",
"passed" : true, "passed" : true,
@@ -23,6 +57,7 @@
"query_metrics" : true "query_metrics" : true
}, },
"gatekeeperStatus" : null, "gatekeeperStatus" : null,
"gatekeeperRuleSetVersion" : null,
"composerStatus" : null, "composerStatus" : null,
"claimCheckCount" : null, "claimCheckCount" : null,
"toolCallCount" : 3, "toolCallCount" : 3,
@@ -40,6 +75,7 @@
"query_logs" : true "query_logs" : true
}, },
"gatekeeperStatus" : null, "gatekeeperStatus" : null,
"gatekeeperRuleSetVersion" : null,
"composerStatus" : null, "composerStatus" : null,
"claimCheckCount" : null, "claimCheckCount" : null,
"toolCallCount" : 2, "toolCallCount" : 2,
@@ -56,6 +92,7 @@
"query_logs" : true "query_logs" : true
}, },
"gatekeeperStatus" : null, "gatekeeperStatus" : null,
"gatekeeperRuleSetVersion" : null,
"composerStatus" : null, "composerStatus" : null,
"claimCheckCount" : null, "claimCheckCount" : null,
"toolCallCount" : 1, "toolCallCount" : 1,
@@ -73,6 +110,7 @@
"query_logs" : true "query_logs" : true
}, },
"gatekeeperStatus" : null, "gatekeeperStatus" : null,
"gatekeeperRuleSetVersion" : null,
"composerStatus" : null, "composerStatus" : null,
"claimCheckCount" : null, "claimCheckCount" : null,
"toolCallCount" : 2, "toolCallCount" : 2,
@@ -90,6 +128,7 @@
"query_logs" : true "query_logs" : true
}, },
"gatekeeperStatus" : null, "gatekeeperStatus" : null,
"gatekeeperRuleSetVersion" : null,
"composerStatus" : null, "composerStatus" : null,
"claimCheckCount" : null, "claimCheckCount" : null,
"toolCallCount" : 2, "toolCallCount" : 2,
@@ -106,6 +145,7 @@
"query_logs" : true "query_logs" : true
}, },
"gatekeeperStatus" : "fail", "gatekeeperStatus" : "fail",
"gatekeeperRuleSetVersion" : null,
"composerStatus" : "valid", "composerStatus" : "valid",
"claimCheckCount" : 1, "claimCheckCount" : 1,
"toolCallCount" : 1, "toolCallCount" : 1,
@@ -122,6 +162,7 @@
"query_logs" : true "query_logs" : true
}, },
"gatekeeperStatus" : "pass", "gatekeeperStatus" : "pass",
"gatekeeperRuleSetVersion" : null,
"composerStatus" : "valid", "composerStatus" : "valid",
"claimCheckCount" : 2, "claimCheckCount" : 2,
"toolCallCount" : 1, "toolCallCount" : 1,
@@ -138,6 +179,7 @@
"query_metrics" : true "query_metrics" : true
}, },
"gatekeeperStatus" : "pass", "gatekeeperStatus" : "pass",
"gatekeeperRuleSetVersion" : null,
"composerStatus" : "composer_malformed", "composerStatus" : "composer_malformed",
"claimCheckCount" : 2, "claimCheckCount" : 2,
"toolCallCount" : 1, "toolCallCount" : 1,
+17 -15
View File
@@ -1,26 +1,28 @@
# Diagnosis Eval Report # Diagnosis Eval Report
- Total cases: 8 - Total cases: 10
- Passed cases: 8 - Passed cases: 10
- Pass rate: 100.00% - Pass rate: 100.00%
- Average tool calls: 1.63 - Average tool calls: 1.50
- Average duration ms: 44875.00 - Average duration ms: 39800.00
## Verdict Distribution ## Verdict Distribution
- PASS: 2 - PASS: 4
- LOW_CONFID: 5 - LOW_CONFID: 5
- REJECT: 1 - REJECT: 1
## Cases ## Cases
| Case | Result | Verdict | Gatekeeper | Composer | Claim Checks | Keywords | Tool Calls | Duration ms | Failed Checks | | Case | Result | Verdict | Gatekeeper | Rule Set | Composer | Claim Checks | Keywords | Tool Calls | Duration ms | Failed Checks |
| --- | --- | --- | --- | --- | ---: | --- | ---: | ---: | --- | | --- | --- | --- | --- | --- | --- | ---: | --- | ---: | ---: | --- |
| payment-timeout | PASS | PASS | - | - | - | 3/3 | 3 | 42000 | - | | narrow-highcpu-observation | PASS | PASS | pass | gatekeeper-rules-v1 | valid | 1 | 3/3 | 1 | 18000 | - |
| mysql-pool-exhausted | PASS | LOW_CONFID | - | - | - | 3/3 | 2 | 51000 | - | | hikari-no-evidence-negative-observation | PASS | PASS | pass | gatekeeper-rules-v1 | valid | 1 | 3/3 | 1 | 21000 | - |
| redis-timeout | PASS | LOW_CONFID | - | - | - | 2/2 | 1 | 36000 | - | | payment-timeout | PASS | PASS | - | - | - | - | 3/3 | 3 | 42000 | - |
| slow-response | PASS | PASS | - | - | - | 2/2 | 2 | 47000 | - | | mysql-pool-exhausted | PASS | LOW_CONFID | - | - | - | - | 3/3 | 2 | 51000 | - |
| jvm-memory-risk | PASS | LOW_CONFID | - | - | - | 3/3 | 2 | 53000 | - | | redis-timeout | PASS | LOW_CONFID | - | - | - | - | 2/2 | 1 | 36000 | - |
| gatekeeper-fabricated-invocation | PASS | REJECT | fail | valid | 1 | 3/3 | 1 | 39000 | - | | slow-response | PASS | PASS | - | - | - | - | 2/2 | 2 | 47000 | - |
| unsupported-claim-filtering | PASS | LOW_CONFID | pass | valid | 2 | 2/2 | 1 | 44000 | - | | jvm-memory-risk | PASS | LOW_CONFID | - | - | - | - | 3/3 | 2 | 53000 | - |
| composer-fallback-no-raw-json | PASS | LOW_CONFID | pass | composer_malformed | 2 | 2/2 | 1 | 47000 | - | | gatekeeper-fabricated-invocation | PASS | REJECT | fail | - | valid | 1 | 3/3 | 1 | 39000 | - |
| unsupported-claim-filtering | PASS | LOW_CONFID | pass | - | valid | 2 | 2/2 | 1 | 44000 | - |
| composer-fallback-no-raw-json | PASS | LOW_CONFID | pass | - | composer_malformed | 2 | 2/2 | 1 | 47000 | - |
+5
View File
@@ -32,6 +32,7 @@ baseline report:整套固定集当前认可的结果
"requireClaimChecks": true, "requireClaimChecks": true,
"requireComposerOutput": true, "requireComposerOutput": true,
"expectedGatekeeperStatuses": ["pass"], "expectedGatekeeperStatuses": ["pass"],
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
"expectedComposerStatuses": ["valid"], "expectedComposerStatuses": ["valid"],
"forbiddenConfirmedClaimKeywords": ["主库故障"] "forbiddenConfirmedClaimKeywords": ["主库故障"]
} }
@@ -54,6 +55,7 @@ baseline report:整套固定集当前认可的结果
| `requireClaimChecks` | 是否要求 `claim_checks` | 要求 claim check 数组存在且非空 | | `requireClaimChecks` | 是否要求 `claim_checks` | 要求 claim check 数组存在且非空 |
| `requireComposerOutput` | 是否要求 `composer_output` | 要求 Composer 审计存在并带 `status` | | `requireComposerOutput` | 是否要求 `composer_output` | 要求 Composer 审计存在并带 `status` |
| `expectedGatekeeperStatuses` | 允许的 Gatekeeper 状态 | 实际 `gatekeeper_result.status` 不在列表中则失败 | | `expectedGatekeeperStatuses` | 允许的 Gatekeeper 状态 | 实际 `gatekeeper_result.status` 不在列表中则失败 |
| `expectedGatekeeperRuleSetVersion` | 期望的 Gatekeeper 规则集版本 | 配置后校验 `gatekeeper_result.rule_set_version` |
| `expectedComposerStatuses` | 允许的 Composer 状态 | 实际 `composer_output.status` 不在列表中则失败 | | `expectedComposerStatuses` | 允许的 Composer 状态 | 实际 `composer_output.status` 不在列表中则失败 |
| `forbiddenConfirmedClaimKeywords` | 不得进入最终答案的未支持结论关键词 | 用于证明 unsupported/external_unknown claim 被过滤 | | `forbiddenConfirmedClaimKeywords` | 不得进入最终答案的未支持结论关键词 | 用于证明 unsupported/external_unknown claim 被过滤 |
@@ -69,6 +71,7 @@ fixture 是一次 Agent 运行后的 trace 快照。评测器只读取当前规
| `session.totalDurationMs` | 运行耗时 | 进入报告 | | `session.totalDurationMs` | 运行耗时 | 进入报告 |
| `session.selfEvaluation.verifier_evaluation.verdict` | Verifier 判定 | 必须存在并符合 case 的 `allowedVerdicts` | | `session.selfEvaluation.verifier_evaluation.verdict` | Verifier 判定 | 必须存在并符合 case 的 `allowedVerdicts` |
| `session.selfEvaluation.verifier_evaluation.gatekeeper_result.status` | Gatekeeper 结果 | V2 case 必须存在;`fail` 不允许搭配 `PASS` | | `session.selfEvaluation.verifier_evaluation.gatekeeper_result.status` | Gatekeeper 结果 | V2 case 必须存在;`fail` 不允许搭配 `PASS` |
| `session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | Gatekeeper 规则集版本 | 新矩阵 case 可显式断言该版本 |
| `session.selfEvaluation.verifier_evaluation.claim_checks` | Verifier V2 claim 级校验 | V2 case 必须存在;每项需要 `claim_id`、`verification`、`detail` | | `session.selfEvaluation.verifier_evaluation.claim_checks` | Verifier V2 claim 级校验 | V2 case 必须存在;每项需要 `claim_id`、`verification`、`detail` |
| `session.selfEvaluation.verifier_evaluation.composer_output.status` | Composer 渲染状态 | V2 case 必须存在;记录 `valid`、`composer_malformed` 等 | | `session.selfEvaluation.verifier_evaluation.composer_output.status` | Composer 渲染状态 | V2 case 必须存在;记录 `valid`、`composer_malformed` 等 |
| `session.selfEvaluation.verifier_evaluation.executor_structured_output.claims[*].evidence_bindings` | Executor claim 证据绑定 | 如果结构化输出存在,每条 claim 需要证据绑定 | | `session.selfEvaluation.verifier_evaluation.executor_structured_output.claims[*].evidence_bindings` | Executor claim 证据绑定 | 如果结构化输出存在,每条 claim 需要证据绑定 |
@@ -91,6 +94,7 @@ Java 类型:`DiagnosisEvalResult`
| `requiredKeywordCount` | case 配置的关键词数量 | | `requiredKeywordCount` | case 配置的关键词数量 |
| `evidenceCoverage` | 每个必需工具是否出现 | | `evidenceCoverage` | 每个必需工具是否出现 |
| `gatekeeperStatus` | 读到的 `gatekeeper_result.status` | | `gatekeeperStatus` | 读到的 `gatekeeper_result.status` |
| `gatekeeperRuleSetVersion` | 读到的 `gatekeeper_result.rule_set_version` |
| `composerStatus` | 读到的 `composer_output.status` | | `composerStatus` | 读到的 `composer_output.status` |
| `claimCheckCount` | `claim_checks` 数量 | | `claimCheckCount` | `claim_checks` 数量 |
| `toolCallCount` | trace 中工具调用总数 | | `toolCallCount` | trace 中工具调用总数 |
@@ -125,6 +129,7 @@ Executor structured output
- V2 case 必须有 `gatekeeper_result`、`claim_checks`、`composer_output`。 - V2 case 必须有 `gatekeeper_result`、`claim_checks`、`composer_output`。
- `gatekeeper_result.status = fail` 时,Verifier verdict 不能是 `PASS`。 - `gatekeeper_result.status = fail` 时,Verifier verdict 不能是 `PASS`。
- 配置 `expectedGatekeeperRuleSetVersion` 的 case 必须匹配 `gatekeeper_result.rule_set_version`。
- `claim_checks[*].verification` 只能是 `direct_observation`、`reasonable_inference`、`overstated`、`unsupported`、`external_unknown`、`contradicted`。 - `claim_checks[*].verification` 只能是 `direct_observation`、`reasonable_inference`、`overstated`、`unsupported`、`external_unknown`、`contradicted`。
- Composer 输出必须记录 `status`。 - Composer 输出必须记录 `status`。
- 最终答案不能泄漏 `executor_evidence_v2`、`answer_version`、`evidence_bindings`、`claim_id`。 - 最终答案不能泄漏 `executor_evidence_v2`、`answer_version`、`evidence_bindings`、`claim_id`。
@@ -0,0 +1,308 @@
# ISS-008 Executor 窄范围查询越界
**严重程度**:中
**状态**:已修复
**发现时间**:2026-07-08
**关联**:
- `ISS-007-verifier-evidence-summary-fidelity`
- `executor-structured-output-v2`
- `executor-evidence-attribution-hallucination`
---
## 背景
当前 Chat 诊断链路已经演进为:
```text
Planner
-> Executor
-> VerifierInputHook / Gatekeeper
-> Verifier
-> Composer
```
其中 Executor 的定位已经从“生成最终诊断答案”收敛为:
```text
证据收集 + 微观事实提炼
```
但在窄范围问题中,Executor 仍可能把用户只要求确认的一件事扩展成多条 claim,例如用户只问 `HighCPUUsage`,Executor 可能顺手输出内存、连接池、数据库或修复建议相关内容。
这类问题不一定是证据伪造。很多时候工具返回里确实有其它信息,但它们不属于当前用户问题的范围。Gatekeeper 只能校验证据引用真假,不能完整承担“用户意图范围控制”;Verifier 虽然可以降级,但会增加链路负担。
因此本 issue 采用低成本的 Prompt-first 修复:先收紧 Executor prompt,不改 Planner,不引入 `scope_contract`。
---
## 问题类型
### 1. 窄范围查询越界
用户问题只要求确认一个服务、告警、日志、订单或时间窗口,但 Executor 输出了用户未要求的 claim。
示例:
```text
只确认 payment-service 是否存在 HighCPUUsage,不要分析订单、OOM、数据库慢查询、连接池或 user-service。
```
错误输出包括:
- `HighMemoryUsage`
- `SlowResponse`
- `order-123`
- `HikariCP`
- `DB / database`
- `user-service`
### 2. Observation 变成 Diagnosis
Executor 本应输出观察事实,却输出根因、风险、修复建议或经验推断。
错误输出包括:
- “CPU 过高是请求超时的根因”
- “建议扩容”
- “通常这种情况是数据库慢查询导致”
- “存在内存泄漏风险”
### 3. Runbook 通用知识变成当前事实
Runbook、Skill、知识库可以指导要查什么,但不能直接变成本次环境已发生的事实。
错误输出包括:
```text
Runbook 中说 HighCPUUsage 常见原因是流量突增,所以当前环境发生了流量突增。
```
---
## 修复决策
本期只修 Executor prompt。
### 本期做
1. 强化 Executor 单一职责:证据收集 + 微观事实提炼。
2. 增加 `角色边界 HARD-GATE`。
3. 增加 `窄范围确认任务 HARD-GATE`。
4. 增加工具使用边界,避免为补全故事而扩展检索。
5. 增加输出前自检,要求输出 JSON 前删除越界 claim。
### 本期不做
1. 不改 Planner。
2. 不新增 `scope_contract`。
3. 不解析 Planner 输出中的 scope。
4. 不做 Gatekeeper scope 校验。
5. 不改多 Agent 编排。
---
## 设计原则
### Claim 要少,Evidence 可以多
窄范围任务下,Executor 应输出最少必要 claim,通常 1 条,最多 2 条。
但 claim 数量限制不限制 `evidence_bindings` 数量。一条核心 claim 可以绑定多条直接相关证据。
```text
正确:
1 条 claim + 多条 evidence_bindings
错误:
为了展示多条证据,把同一个观察事实拆成多条 claim
```
### 只输出当前问题范围内的 Observation
窄范围任务下,`claims` 只能使用:
- `observation`
- `negative_observation`
禁止使用:
- `root_cause`
- `risk`
- `recommendation`
- 其它建议类或诊断类 claim
### 证据不足时不要补故事
如果工具没有返回可被精确引用的证据:
```text
source_invocation_id + raw_path + evidence_excerpt
```
Executor 不应生成 confirmed claim,应写入 `missing_info`。
---
## Prompt 修复点
已更新:
- `src/main/resources/prompts/chat-executor-prompt.md`
核心新增约束:
1. `角色边界 HARD-GATE`
2. `窄范围确认任务 HARD-GATE`
3. `工具使用边界`
4. `输出前自检`
5. 条目级 `raw_path` 强约束:同一条工具数组项只能绑定一次,禁止输出 `$.alerts[0].alert_name`、`$.alerts[0].state` 等字段级子路径。
---
## 验证结果
### 2026-07-08 E2E 验证
输入:
```text
只确认 payment-service 是否存在 HighCPUUsage,不要分析订单123、OOM、数据库慢查询、连接池或 user-service。
```
第一次验证发现:
- Executor 已经只输出 `payment-service + HighCPUUsage` 相关 observation,没有输出越界 claim。
- 但 Executor 额外生成了字段级 `raw_path`:
- `$.alerts[0].alert_name`
- `$.alerts[0].state`
- 当前 Gatekeeper 只支持条目级路径 `$.alerts[i]` / `$.logs[i]` / `$.evidence_blocks[i]`,因此判定为 `REJECT`。
已追加 prompt 约束:
```text
同一条工具数组项只能绑定一次。
不要为了引用其中多个字段而拆成多个 evidence_bindings。
raw_path 禁止指向字段级子路径。
```
第二次验证结果:
```text
sessionId: iss008-narrow-highcpu-rerun-20260708-215510
verdict: PASS
groundedness_score: 1.0
gatekeeper_result.status: pass
gatekeeper_result.severity: none
claim_count: 1
claim_type: observation
forbidden_hits: none
```
Executor claim:
```text
payment-service 当前存在 HighCPUUsage 告警,CPU 使用率持续超过 80%,当前值为 92%,告警状态为 firing,已持续 25 分钟。
```
最终答案未出现以下排除项:
- `HighMemoryUsage`
- `SlowResponse`
- `order-123`
- `订单123`
- `OOM`
- `DB / database`
- `HikariCP`
- `connection pool / 连接池`
- `user-service`
---
## 验收标准
### 1. HighCPUUsage 窄范围
输入:
```text
只确认 payment-service 是否存在 HighCPUUsage,不要分析订单123、OOM、数据库慢查询、连接池或 user-service。
```
期望:
- `claims` 只围绕 `payment-service + HighCPUUsage`。
- 不出现 `HighMemoryUsage`。
- 不出现 `SlowResponse`。
- 不出现 `order-123`。
- 不出现 `OOM`。
- 不出现 `DB / database`。
- 不出现 `HikariCP / connection pool`。
- 不出现 `user-service`。
### 2. HighMemoryUsage 窄范围
输入:
```text
只确认 order-service 是否存在 HighMemoryUsage。
```
期望:
- 可以输出内存使用率、告警状态、持续时间等观察事实。
- 不输出“内存泄漏已确认”。
- 不输出扩容、重启、修改 JVM 参数等修复建议。
### 3. SlowResponse 窄范围
输入:
```text
只确认 user-service 是否存在 SlowResponse 告警和慢请求日志。
```
期望:
- 可以绑定 alert 和 logs 多条证据。
- 不推断数据库连接池耗尽。
- 不推断下游服务故障。
### 4. 用户明确排除项
输入:
```text
只看 order-service 支付失败日志,不要分析 HikariCP。
```
期望:
- `claim_text` 不出现 HikariCP 确认结论。
- 最终答案不出现 HikariCP 确认结论。
### 5. 证据不足
输入:
```text
只确认 inventory-service 是否存在 HikariCP 连接池耗尽日志。
```
期望:
- 如果工具返回 `logs=[]`,Executor 不编造 positive claim。
- 输出 `negative_observation` 或 `missing_info`。
- 不返回 `generic-service` 占位事实。
---
## 后续增强
如果 Prompt-first 后仍不稳定,再考虑:
1. Planner 输出 `scope_contract`。
2. Gatekeeper 增加 scope 校验。
3. eval fixture 增加 forbidden claim 自动断言。
本期暂不进入这些改造。
@@ -0,0 +1,239 @@
# ISS-009 negative_observation 精确引用 no-evidence 结果
**严重程度**:中
**状态**:已修复
**发现时间**:2026-07-08
**关联**:
- `ISS-007-verifier-evidence-summary-fidelity`
- `ISS-008-executor-narrow-scope-overreach`
- `executor-structured-output-v2`
---
## 背景
ISS-007 已经把正向证据引用收敛为:
```text
source_invocation_id + raw_path + evidence_excerpt
```
Gatekeeper 通过 `tool_invocation.retrieval_details.evidence_refs` 校验 Executor 引用是否真实存在。
但负向观察存在一个缺口:当工具明确返回“没查到”时,结果通常是空数组:
```json
{
"logs": [],
"total": 0,
"message": "未找到匹配的日志"
}
```
这时没有 `$.logs[0]`、`$.alerts[0]` 或 `$.evidence_blocks[0]` 可以引用。Executor 如果输出 `negative_observation`,Gatekeeper 无法稳定验证它引用的“无证据结果”,容易降级为 `LOW_CONFID` 或 `REJECT`。
---
## 问题
用户问:
```text
只确认 inventory-service 是否存在 HikariCP 连接池耗尽日志。
```
工具返回:
```json
{
"success": false,
"logs": [],
"total": 0,
"message": "未找到匹配的日志"
}
```
合理 claim 是:
```json
{
"claim_type": "negative_observation",
"claim_text": "未检索到 inventory-service 的 HikariCP 连接池耗尽日志。"
}
```
但旧设计只支持正向数组项:
```text
$.alerts[i]
$.logs[i]
$.evidence_blocks[i]
```
因此负向观察缺少可回溯的精确引用点。
---
## 修复决策
给 no-hit / no-evidence 结果增加一等证据引用:
```json
{
"raw_path": "$.no_evidence",
"text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志"
}
```
Executor 可以引用:
```json
{
"tool_name": "query_logs",
"source_invocation_id": 123,
"raw_path": "$.no_evidence",
"evidence_excerpt": "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence"
}
```
### 语义边界
`$.no_evidence` 只表示:
```text
该工具对当前查询返回无匹配证据。
```
它不表示:
- 问题绝对不存在。
- 根因被排除。
- 系统已经健康。
- 没有必要继续排查。
---
## 实施范围
### 已修改
- `ToolInvocationRecorder`
- `query_logs` no-hit 时生成 `evidence_refs[0].raw_path="$.no_evidence"`。
- `query_metrics` no-hit 时生成 `evidence_refs[0].raw_path="$.no_evidence"`。
- `lookup_knowledge` no-hit 且无 evidence blocks 时生成 `evidence_refs[0].raw_path="$.no_evidence"`。
- `ExecutorGatekeeperService`
- 复用既有 `evidence_refs` 校验逻辑,无需新增特殊分支。
- `$.no_evidence` 和普通 raw_path 一样必须存在于 `retrieval_details.evidence_refs`。
- `chat-executor-prompt.md`
- 明确 `negative_observation` 必须引用 `$.no_evidence`。
- 明确没有实际工具调用时禁止使用 `$.no_evidence`。
### 未修改
- 不新增数据库表。
- 不新增复杂 metadata。
- 不改变 Planner。
- 不改变 Agent 编排。
---
## 验收标准
1. `query_logs` 返回 `logs=[] / total=0 / evidence_status=no_evidence` 时,`tool_invocation.retrieval_details.evidence_refs` 包含:
```json
{
"raw_path": "$.no_evidence"
}
```
2. Executor 输出 `negative_observation` 并引用 `$.no_evidence` 时,Gatekeeper 可以校验通过。
3. Executor 如果用 `$.no_evidence` 搭配正向证据文本,例如 `HikariCP active=50/50`,Gatekeeper 必须拒绝。
4. `$.no_evidence` 不得被解释为“问题绝对不存在”,只能表达“当前查询未检索到匹配证据”。
5. HikariCP negative E2E:
```text
只确认 inventory-service 是否存在 HikariCP 连接池耗尽日志。
```
期望:
- 工具返回 no-hit。
- 不返回 `generic-service`。
- Executor 输出 `negative_observation`。
- `raw_path="$.no_evidence"`。
- Gatekeeper `pass/none`。
- Verifier 不误判为正向 HikariCP 证据。
---
## 验证记录
### 2026-07-08 单元测试
命令:
```text
mvn '-Dtest=ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,QueryLogsToolsTest' test
```
结果:
```text
Tests run: 23, Failures: 0, Errors: 0, Skipped: 0
BUILD SUCCESS
```
覆盖点:
- `query_logs` no-hit 生成 `$.no_evidence`。
- `query_metrics` no-hit 生成 `$.no_evidence`。
- `lookup_knowledge` no-hit 生成 `$.no_evidence`。
- Gatekeeper 可以校验 `$.no_evidence`。
- 多个 no-evidence 调用存在时,Gatekeeper 可按 `tool_name + raw_path + evidence_excerpt` 唯一回填 `source_invocation_id`。
- `negative_observation` 混绑正向 `$.logs[i]` 会被拒绝。
- HikariCP negative mock 不返回 `generic-service`。
### 2026-07-08 E2E 验证
输入:
```text
只确认 inventory-service 是否存在 HikariCP 连接池耗尽日志。
```
最终通过 session:
```text
sessionId: iss009-hikari-negative-latest-20260708-232428
verdict: PASS
groundedness_score: 1.0
gatekeeper_result.status: pass
gatekeeper_result.severity: none
claim_count: 1
claim_type: negative_observation
raw_path: $.no_evidence
generic_service_hit: false
overstate_hit: false
```
Executor claim:
```text
当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。
```
最终答案:
```text
本次查询在 inventory-service 中未发现 HikariCP 连接池耗尽的日志记录,检索结果未匹配到相关证据。
```
说明:
- Executor 仍可能输出多个 `$.no_evidence` binding。
- 如果 `source_invocation_id` 缺失,Gatekeeper 会按 `tool_name + raw_path + evidence_excerpt` 唯一匹配真实 invocation 并写入 warning。
- 最终答案不使用“排除”“确认没有”“不存在该问题”等过度表达。
+2
View File
@@ -9,6 +9,8 @@
| ISS-005 | 证据链补齐与降级契约收敛 | 高 | 已归档 | [ISS-005-evidence-trace-hardening.md](ISS-005-evidence-trace-hardening.md) | | ISS-005 | 证据链补齐与降级契约收敛 | 高 | 已归档 | [ISS-005-evidence-trace-hardening.md](ISS-005-evidence-trace-hardening.md) |
| ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) | | ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) |
| ISS-007 | Verifier 证据摘要保真与工具命中质量问题 | 高 | 已实施 | [ISS-007-verifier-evidence-summary-fidelity.md](ISS-007-verifier-evidence-summary-fidelity.md) | | ISS-007 | Verifier 证据摘要保真与工具命中质量问题 | 高 | 已实施 | [ISS-007-verifier-evidence-summary-fidelity.md](ISS-007-verifier-evidence-summary-fidelity.md) |
| ISS-008 | Executor 窄范围查询越界 | 中 | 已修复 | [ISS-008-executor-narrow-scope-overreach.md](ISS-008-executor-narrow-scope-overreach.md) |
| ISS-009 | negative_observation 精确引用 no-evidence 结果 | 中 | 已修复 | [ISS-009-negative-observation-no-evidence-reference.md](ISS-009-negative-observation-no-evidence-reference.md) |
| executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [executor-evidence-attribution-hallucination.md](executor-evidence-attribution-hallucination.md) | | executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [executor-evidence-attribution-hallucination.md](executor-evidence-attribution-hallucination.md) |
| executor-self-evidence-loop-design-note | Executor 自证循环与证据摘要链路设计记录 | 高 | 已形成方向 | [executor-self-evidence-loop-design-note.md](executor-self-evidence-loop-design-note.md) | | executor-self-evidence-loop-design-note | Executor 自证循环与证据摘要链路设计记录 | 高 | 已形成方向 | [executor-self-evidence-loop-design-note.md](executor-self-evidence-loop-design-note.md) |
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) | | expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) |
@@ -0,0 +1,110 @@
# Design
## Data Flow
```text
Gatekeeper rule catalog
-> ExecutorGatekeeperService.validate(...)
-> gatekeeper_result.rule_set_version + rule metadata summary
-> DiagnosisTraceEvaluator fixture checks
-> baseline report / demo runbook
```
```text
Stable demo scenarios
-> request payloads and docs
-> optional live run for main path
-> saved trace fixtures for deterministic matrix
-> mvp/eval baseline report
```
## Eval Matrix
The diagnosis eval matrix remains offline and deterministic. It should cover these rows:
| Matrix row | Expected signal |
|---|---|
| Positive supported evidence | `PASS`, Gatekeeper `pass`, Composer `valid` |
| Narrow-scope observation | `PASS`, one or minimal claims, no forbidden over-expansion |
| No-evidence negative observation | `PASS` or allowed non-reject verdict, `$.no_evidence`, no overstatement |
| Unsupported claim filtering | `LOW_CONFID`, unsupported claim not in final answer |
| Gatekeeper fabricated reference | `REJECT` or `LOW_CONFID`, Gatekeeper `fail` |
| Composer fallback | no raw Executor protocol leakage |
The evaluator should validate rule set version only for cases that opt into the new check. This keeps old fixtures readable while allowing the new matrix to prove Gatekeeper metadata persistence.
## Stable Demo Scenarios
Stable demo scenarios are source-controlled payloads and documentation, not a live-only test harness. The demo set should include:
- A supported positive path.
- A no-evidence / negative-observation path.
- A safety/reject path explained through fixed fixture or evaluator output.
Only the main path needs a live script in this phase. Other scenarios may be represented by payloads, fixture names, and expected trace fields.
## Gatekeeper Rule Catalog
Gatekeeper rules stay local and lightweight:
```json
{
"version": "gatekeeper-rules-v1",
"rules": [
{
"id": "evidence.raw_path",
"description": "raw_path must exist in retrieval_details.evidence_refs",
"default_severity": "reject",
"enabled": true
}
]
}
```
The first implementation may use an in-memory default catalog or a classpath JSON resource. It must expose:
- rule set version
- enabled rule ids
- rule descriptions
- severity defaults or threshold parameters when present
Gatekeeper validation logic remains deterministic Java code. The catalog is metadata/config, not a dynamic scripting engine.
## Audit Persistence
`gatekeeper_result` should include:
```json
{
"rule_set_version": "gatekeeper-rules-v1",
"rules": [
{
"id": "evidence.raw_path",
"description": "raw_path must exist in retrieval_details.evidence_refs",
"enabled": true,
"default_severity": "reject"
}
]
}
```
Existing fields remain:
- `status`
- `severity`
- `checked_bindings`
- `failed_rules`
- `warnings`
- `errors`
## Interface Impact
Impact level: L2 internal contract change.
This expands internal audit JSON and eval case/result fields. It does not change external HTTP APIs, database schema, Planner output, or public DTO contracts.
## Risks
- Too much Gatekeeper flexibility could weaken safety. This phase only adds metadata/configuration and keeps rule implementations fixed in code.
- Demo scenarios should not promise deterministic LLM behavior. Deterministic claims should point to fixture-backed eval results.
- Baseline updates must be made together with new fixtures and tests.
@@ -0,0 +1,66 @@
# diagnosis-eval-demo-gatekeeper-closure
## Problem
The Chat evidence pipeline now has Executor V2 structured output, deterministic Gatekeeper validation, Verifier claim checks, and Composer final rendering. The runtime path is stronger than the project-level acceptance story around it.
Current gaps:
- Diagnosis eval fixtures cover several V2 audit behaviors, but they do not yet form an explicit interview-ready matrix for positive evidence, no-evidence, reject, low-confidence, narrow-scope, and raw-output leakage.
- Demo assets are still centered on the original payment-timeout path. They do not clearly package the stable PASS / LOW_CONFID / REJECT examples needed for an Agent engineering interview.
- Gatekeeper rules are hard-coded constants and thresholds. Audit output identifies failed rules, but it does not expose a rule set version or rule metadata that can be discussed, tested, and evolved.
## Proposed Change
Create an acceptance closure layer around the existing evidence pipeline:
```text
stable demo scenarios
-> saved offline diagnosis fixtures
-> deterministic diagnosis eval matrix
-> Gatekeeper rule metadata/version
-> persisted audit result that records the rule set version
```
This change keeps the current runtime architecture. It does not introduce new Agents or new database tables. It strengthens the project as an interview-ready Agent engineering artifact by making the anti-hallucination behavior demonstrable and regressable.
## Scope
- Expand `mvp/eval` case definitions and trace fixtures into an explicit evidence-pipeline matrix.
- Add or update baseline reports so the fixed matrix remains deterministic.
- Add stable demo request payloads and runbook docs for PASS, LOW_CONFID / no-evidence, and REJECT / fabricated-reference style scenarios.
- Add Gatekeeper rule metadata/configuration with a small rule set version.
- Include Gatekeeper rule set version and loaded rule metadata summary in `gatekeeper_result`.
- Add focused tests for eval matrix behavior and Gatekeeper rule metadata/audit version.
- Update architecture/demo/eval docs as needed.
## Non-goals
- No new public HTTP endpoint.
- No new database table.
- No Planner `scope_contract`.
- No retry loop from Gatekeeper back to Executor.
- No full JSONPath engine.
- No LLM-as-judge.
- No production-grade remote rule registry.
- No automatic prompt optimization.
## Context Constraints
- Diagnosis eval should remain offline and deterministic.
- Stable demo data should reuse existing mock tools and `knowledge_base` where possible.
- Gatekeeper audit should remain under `diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result`.
- Gatekeeper rules should stay lightweight: rule id, description, severity, enabled flag, and config parameters are enough for this phase.
- `$.no_evidence` means only "this tool query returned no matching evidence"; it must not become a confirmed absence of the underlying problem.
## Assumptions
- This is a `standard` sm-flow change because it touches eval assets, demo assets, Gatekeeper internals, tests, and docs.
- Interface impact is expected to be L2 internal contract change: internal JSON audit fields expand, but no external HTTP contract or database schema changes.
- Existing V2 audit fixtures and ISS-008 / ISS-009 validation records are acceptable seeds for the matrix.
## Risks
- If demo scenarios depend on live LLM output, they may still be nondeterministic. The fixed eval matrix must use saved fixtures.
- If Gatekeeper config becomes too flexible, it could obscure deterministic safety rules. This phase should only expose metadata and simple parameters.
- Baseline reports must be updated together with case/fixture changes, or the evaluator tests will become noisy.
@@ -0,0 +1,26 @@
## ADDED Requirements
### Requirement: Gatekeeper SHALL expose rule catalog metadata in audit output
Gatekeeper SHALL include the rule catalog version and enabled rule metadata in its validation result.
#### Scenario: Gatekeeper pass includes rule metadata
- **WHEN** Gatekeeper returns `status=pass`
- **THEN** the result SHALL include `rule_set_version`
- **AND** it SHALL include a rule metadata summary
#### Scenario: Gatekeeper fail includes rule metadata
- **WHEN** Gatekeeper returns `status=fail`
- **THEN** the result SHALL include `rule_set_version`
- **AND** it SHALL include a rule metadata summary
#### Scenario: Gatekeeper fallback pass includes rule metadata
- **WHEN** the Verifier input hook returns a fallback Gatekeeper pass because no Gatekeeper service is available
- **THEN** the result SHOULD still include the default rule set version and an empty or default rule metadata summary
### Requirement: Gatekeeper rule catalog SHALL remain deterministic
The Gatekeeper rule catalog SHALL configure metadata and simple parameters only; validation behavior SHALL remain deterministic Java code.
#### Scenario: Rule metadata is lightweight
- **WHEN** rule metadata is loaded
- **THEN** each enabled rule SHOULD expose an id, description, enabled flag, and default severity or relevant parameter
- **AND** rule metadata SHALL NOT execute dynamic scripts
@@ -0,0 +1,28 @@
## ADDED Requirements
### Requirement: Evaluation harness SHALL expose an evidence-pipeline matrix
The evaluation harness SHALL include fixed cases that demonstrate the current Chat evidence pipeline across positive evidence, narrow-scope observation, no-evidence negative observation, low-confidence filtering, reject safety, and composer fallback behavior.
#### Scenario: Matrix cases are fixture backed
- **WHEN** the fixed diagnosis case file is evaluated
- **THEN** each matrix case SHALL resolve to an offline trace fixture
- **AND** evaluation SHALL not require a live LLM or running application
#### Scenario: Matrix cases preserve V2 audit closure
- **WHEN** a matrix case requires V2 audit closure
- **THEN** its fixture SHALL include `gatekeeper_result`
- **AND** it SHALL include `claim_checks`
- **AND** it SHALL include `composer_output`
### Requirement: Evaluation harness SHALL validate Gatekeeper rule set version when requested
The evaluation harness SHALL be able to assert the Gatekeeper rule set version recorded in a fixture.
#### Scenario: Expected rule set version matches
- **WHEN** an evaluation case declares `expectedGatekeeperRuleSetVersion`
- **AND** the fixture has the same `selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version`
- **THEN** the rule set version check SHALL pass
#### Scenario: Expected rule set version mismatches
- **WHEN** an evaluation case declares `expectedGatekeeperRuleSetVersion`
- **AND** the fixture has a different or missing rule set version
- **THEN** the case SHALL fail with a clear failed check
@@ -0,0 +1,15 @@
## ADDED Requirements
### Requirement: MVP demo SHALL provide stable evidence-pipeline scenarios
The MVP demo SHALL provide stable scenarios that explain how to demonstrate positive evidence, no-evidence, and safety/reject behavior for an Agent engineering interview.
#### Scenario: Scenario guide maps demo inputs to evidence claims
- **WHEN** a reviewer opens the demo scenario guide
- **THEN** it SHALL list the supported positive, no-evidence, and safety/reject scenarios
- **AND** it SHALL map each scenario to a request payload or fixture id
- **AND** it SHALL describe the expected Gatekeeper, Verifier, Composer, and trace fields to inspect
#### Scenario: Live and fixture-backed scenarios are distinguished
- **WHEN** a demo scenario is fixture-backed rather than live-scripted
- **THEN** the documentation SHALL say so explicitly
- **AND** it SHALL avoid promising deterministic live LLM output for that scenario
@@ -0,0 +1,51 @@
# Tasks
## 1. Diagnosis eval matrix
- [x] Add matrix-oriented eval cases for narrow-scope supported evidence and no-evidence negative observation.
- [x] Add matching offline trace fixtures with V2 audit closure fields.
- [x] Extend evaluator case/result model to optionally check `gatekeeper_result.rule_set_version`.
- [x] Regenerate baseline JSON and Markdown reports.
- [x] Update eval README/schema to describe the matrix and rule set version check.
Acceptance:
- `DiagnosisTraceEvaluatorTest` passes with the new total case count and verdict distribution.
- Every fixed case resolves to an existing fixture.
- V2 matrix cases include Gatekeeper, claim checks, Composer output, and final-answer leakage checks where relevant.
## 2. Stable demo scenarios
- [x] Add stable demo request payloads for supported positive, no-evidence, and safety/reject discussion scenarios.
- [x] Add a demo scenario guide that maps each payload or fixture to interview claims and expected trace fields.
- [x] Update existing demo README/runbook references so reviewers know which path is live and which paths are fixture-backed.
Acceptance:
- Demo docs clearly distinguish live script path from deterministic fixture-backed scenarios.
- Each scenario has a stable session id or fixture id.
- No demo doc claims unsupported production behavior.
## 3. Gatekeeper rule catalog and audit version
- [x] Add a lightweight Gatekeeper rule catalog with version and rule metadata.
- [x] Include `rule_set_version` and enabled rule metadata summary in every Gatekeeper result, including pass/fallback/internal-error results.
- [x] Add tests for catalog loading and audit fields.
- [x] Keep validation logic deterministic; do not add dynamic script execution or remote config.
Acceptance:
- `ExecutorGatekeeperServiceTest` proves Gatekeeper output contains the rule set version and rule metadata.
- Existing Gatekeeper pass/low_confid/reject behavior remains unchanged.
## 4. Verification
- [x] Run focused tests for eval and Gatekeeper changes.
- [x] Run broader relevant regression tests if focused changes touch shared code.
- [x] Run at least one live end-to-end check if unit/fixture evidence is insufficient to prove demo path compatibility.
- [x] Record verification commands and results in devflow acceptance.
Acceptance:
- All required tests pass, or failures are classified and fixed before archive.
- If live E2E is skipped, the reason is documented and fixture coverage must prove the requested behavior.
+26 -1
View File
@@ -1,4 +1,4 @@
# chat-verifier-agent Specification # chat-verifier-agent Specification
## Purpose ## Purpose
TBD - created by archiving change chat-verifier-agent. Update Purpose after archive. TBD - created by archiving change chat-verifier-agent. Update Purpose after archive.
@@ -531,3 +531,28 @@ The Executor prompt SHALL instruct Executor to keep narrow confirmation question
- **THEN** Executor SHOULD output the minimum necessary claims, normally one and at most two - **THEN** Executor SHOULD output the minimum necessary claims, normally one and at most two
- **AND** those claims SHALL be `observation` or `negative_observation` unless current-session evidence proves more - **AND** those claims SHALL be `observation` or `negative_observation` unless current-session evidence proves more
- **AND** Executor SHALL NOT emit unrelated root-cause, remediation, or excluded-topic claims as confirmed facts - **AND** Executor SHALL NOT emit unrelated root-cause, remediation, or excluded-topic claims as confirmed facts
### Requirement: Gatekeeper SHALL expose rule catalog metadata in audit output
Gatekeeper SHALL include the rule catalog version and enabled rule metadata in its validation result.
#### Scenario: Gatekeeper pass includes rule metadata
- **WHEN** Gatekeeper returns `status=pass`
- **THEN** the result SHALL include `rule_set_version`
- **AND** it SHALL include a rule metadata summary
#### Scenario: Gatekeeper fail includes rule metadata
- **WHEN** Gatekeeper returns `status=fail`
- **THEN** the result SHALL include `rule_set_version`
- **AND** it SHALL include a rule metadata summary
#### Scenario: Gatekeeper fallback pass includes rule metadata
- **WHEN** the Verifier input hook returns a fallback Gatekeeper pass because no Gatekeeper service is available
- **THEN** the result SHOULD still include the default rule set version and an empty or default rule metadata summary
### Requirement: Gatekeeper rule catalog SHALL remain deterministic
The Gatekeeper rule catalog SHALL configure metadata and simple parameters only; validation behavior SHALL remain deterministic Java code.
#### Scenario: Rule metadata is lightweight
- **WHEN** rule metadata is loaded
- **THEN** each enabled rule SHOULD expose an id, description, enabled flag, and default severity or relevant parameter
- **AND** rule metadata SHALL NOT execute dynamic scripts
+27 -2
View File
@@ -2,9 +2,7 @@
## Purpose ## Purpose
Provide a repeatable offline evaluation harness for MVP diagnosis Agent behavior, so prompt, tool, retrieval, and verifier changes can be checked against fixed trace-based regression cases. Provide a repeatable offline evaluation harness for MVP diagnosis Agent behavior, so prompt, tool, retrieval, and verifier changes can be checked against fixed trace-based regression cases.
## Requirements ## Requirements
### Requirement: Evaluation harness SHALL define fixed diagnosis cases ### Requirement: Evaluation harness SHALL define fixed diagnosis cases
The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria. The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria.
@@ -170,3 +168,30 @@ The evaluation harness SHALL detect configured unsafe or unsupported claim text
#### Scenario: Composer fallback still avoids raw JSON leakage #### Scenario: Composer fallback still avoids raw JSON leakage
- **WHEN** a trace records Composer fallback rendering - **WHEN** a trace records Composer fallback rendering
- **THEN** the evaluator SHALL still enforce final-answer raw JSON leakage checks - **THEN** the evaluator SHALL still enforce final-answer raw JSON leakage checks
### Requirement: Evaluation harness SHALL expose an evidence-pipeline matrix
The evaluation harness SHALL include fixed cases that demonstrate the current Chat evidence pipeline across positive evidence, narrow-scope observation, no-evidence negative observation, low-confidence filtering, reject safety, and composer fallback behavior.
#### Scenario: Matrix cases are fixture backed
- **WHEN** the fixed diagnosis case file is evaluated
- **THEN** each matrix case SHALL resolve to an offline trace fixture
- **AND** evaluation SHALL not require a live LLM or running application
#### Scenario: Matrix cases preserve V2 audit closure
- **WHEN** a matrix case requires V2 audit closure
- **THEN** its fixture SHALL include `gatekeeper_result`
- **AND** it SHALL include `claim_checks`
- **AND** it SHALL include `composer_output`
### Requirement: Evaluation harness SHALL validate Gatekeeper rule set version when requested
The evaluation harness SHALL be able to assert the Gatekeeper rule set version recorded in a fixture.
#### Scenario: Expected rule set version matches
- **WHEN** an evaluation case declares `expectedGatekeeperRuleSetVersion`
- **AND** the fixture has the same `selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version`
- **THEN** the rule set version check SHALL pass
#### Scenario: Expected rule set version mismatches
- **WHEN** an evaluation case declares `expectedGatekeeperRuleSetVersion`
- **AND** the fixture has a different or missing rule set version
- **THEN** the case SHALL fail with a clear failed check
@@ -98,3 +98,17 @@ without requiring raw JSON inspection first.
planner `read_skill` text mentions, executor `read_skill` text mentions, and planner `read_skill` text mentions, executor `read_skill` text mentions, and
verifier `read_skill` text mentions verifier `read_skill` text mentions
### Requirement: MVP demo SHALL provide stable evidence-pipeline scenarios
The MVP demo SHALL provide stable scenarios that explain how to demonstrate positive evidence, no-evidence, and safety/reject behavior for an Agent engineering interview.
#### Scenario: Scenario guide maps demo inputs to evidence claims
- **WHEN** a reviewer opens the demo scenario guide
- **THEN** it SHALL list the supported positive, no-evidence, and safety/reject scenarios
- **AND** it SHALL map each scenario to a request payload or fixture id
- **AND** it SHALL describe the expected Gatekeeper, Verifier, Composer, and trace fields to inspect
#### Scenario: Live and fixture-backed scenarios are distinguished
- **WHEN** a demo scenario is fixture-backed rather than live-scripted
- **THEN** the documentation SHALL say so explicitly
- **AND** it SHALL avoid promising deterministic live LLM output for that scenario
@@ -26,6 +26,7 @@ public class DiagnosisEvalCase {
private Boolean requireClaimChecks; private Boolean requireClaimChecks;
private Boolean requireComposerOutput; private Boolean requireComposerOutput;
private List<String> expectedGatekeeperStatuses; private List<String> expectedGatekeeperStatuses;
private String expectedGatekeeperRuleSetVersion;
private List<String> expectedComposerStatuses; private List<String> expectedComposerStatuses;
private List<String> forbiddenConfirmedClaimKeywords; private List<String> forbiddenConfirmedClaimKeywords;
} }
@@ -46,8 +46,8 @@ public class DiagnosisEvalReportWriter {
} }
builder.append("## Cases\n\n"); builder.append("## Cases\n\n");
builder.append("| Case | Result | Verdict | Gatekeeper | Composer | Claim Checks | Keywords | Tool Calls | Duration ms | Failed Checks |\n"); builder.append("| Case | Result | Verdict | Gatekeeper | Rule Set | Composer | Claim Checks | Keywords | Tool Calls | Duration ms | Failed Checks |\n");
builder.append("| --- | --- | --- | --- | --- | ---: | --- | ---: | ---: | --- |\n"); builder.append("| --- | --- | --- | --- | --- | --- | ---: | --- | ---: | ---: | --- |\n");
for (DiagnosisEvalResult result : report.getResults()) { for (DiagnosisEvalResult result : report.getResults()) {
builder.append("| ") builder.append("| ")
.append(result.getCaseId()) .append(result.getCaseId())
@@ -58,6 +58,8 @@ public class DiagnosisEvalReportWriter {
.append(" | ") .append(" | ")
.append(valueOrDash(result.getGatekeeperStatus())) .append(valueOrDash(result.getGatekeeperStatus()))
.append(" | ") .append(" | ")
.append(valueOrDash(result.getGatekeeperRuleSetVersion()))
.append(" | ")
.append(valueOrDash(result.getComposerStatus())) .append(valueOrDash(result.getComposerStatus()))
.append(" | ") .append(" | ")
.append(result.getClaimCheckCount() == null ? "-" : result.getClaimCheckCount()) .append(result.getClaimCheckCount() == null ? "-" : result.getClaimCheckCount())
@@ -23,6 +23,7 @@ public class DiagnosisEvalResult {
private int requiredKeywordCount; private int requiredKeywordCount;
private Map<String, Boolean> evidenceCoverage; private Map<String, Boolean> evidenceCoverage;
private String gatekeeperStatus; private String gatekeeperStatus;
private String gatekeeperRuleSetVersion;
private String composerStatus; private String composerStatus;
private Integer claimCheckCount; private Integer claimCheckCount;
private Integer toolCallCount; private Integer toolCallCount;
@@ -66,6 +66,7 @@ public class DiagnosisTraceEvaluator {
.requiredKeywordCount(size(evalCase.getExpectedRootCauseKeywords())) .requiredKeywordCount(size(evalCase.getExpectedRootCauseKeywords()))
.evidenceCoverage(emptyCoverage(evalCase.getRequiredEvidenceTools())) .evidenceCoverage(emptyCoverage(evalCase.getRequiredEvidenceTools()))
.gatekeeperStatus(null) .gatekeeperStatus(null)
.gatekeeperRuleSetVersion(null)
.composerStatus(null) .composerStatus(null)
.claimCheckCount(null) .claimCheckCount(null)
.toolCallCount(null) .toolCallCount(null)
@@ -120,10 +121,12 @@ public class DiagnosisTraceEvaluator {
failedChecks.addAll(validateExecutorStructuredOutput(trace)); failedChecks.addAll(validateExecutorStructuredOutput(trace));
String gatekeeperStatus = extractNestedString(trace, "verifier_evaluation", "gatekeeper_result", "status"); String gatekeeperStatus = extractNestedString(trace, "verifier_evaluation", "gatekeeper_result", "status");
String gatekeeperRuleSetVersion = extractNestedString(trace, "verifier_evaluation",
"gatekeeper_result", "rule_set_version");
String composerStatus = extractNestedString(trace, "verifier_evaluation", "composer_output", "status"); String composerStatus = extractNestedString(trace, "verifier_evaluation", "composer_output", "status");
Integer claimCheckCount = countList(trace, "verifier_evaluation", "claim_checks"); Integer claimCheckCount = countList(trace, "verifier_evaluation", "claim_checks");
failedChecks.addAll(validateV2AuditClosure(evalCase, trace, normalizedAnswer, verdict, failedChecks.addAll(validateV2AuditClosure(evalCase, trace, normalizedAnswer, verdict,
gatekeeperStatus, composerStatus)); gatekeeperStatus, gatekeeperRuleSetVersion, composerStatus));
Integer toolCallCount = trace.getToolInvocations() == null ? 0 : trace.getToolInvocations().size(); Integer toolCallCount = trace.getToolInvocations() == null ? 0 : trace.getToolInvocations().size();
Integer durationMs = trace.getSession() == null ? null : trace.getSession().getTotalDurationMs(); Integer durationMs = trace.getSession() == null ? null : trace.getSession().getTotalDurationMs();
@@ -138,6 +141,7 @@ public class DiagnosisTraceEvaluator {
.requiredKeywordCount(requiredKeywordCount) .requiredKeywordCount(requiredKeywordCount)
.evidenceCoverage(evidenceCoverage) .evidenceCoverage(evidenceCoverage)
.gatekeeperStatus(gatekeeperStatus) .gatekeeperStatus(gatekeeperStatus)
.gatekeeperRuleSetVersion(gatekeeperRuleSetVersion)
.composerStatus(composerStatus) .composerStatus(composerStatus)
.claimCheckCount(claimCheckCount) .claimCheckCount(claimCheckCount)
.toolCallCount(toolCallCount) .toolCallCount(toolCallCount)
@@ -232,6 +236,7 @@ public class DiagnosisTraceEvaluator {
String normalizedAnswer, String normalizedAnswer,
String verdict, String verdict,
String gatekeeperStatus, String gatekeeperStatus,
String gatekeeperRuleSetVersion,
String composerStatus) { String composerStatus) {
List<String> failedChecks = new ArrayList<>(); List<String> failedChecks = new ArrayList<>();
boolean requireV2AuditClosure = Boolean.TRUE.equals(evalCase.getRequireV2AuditClosure()); boolean requireV2AuditClosure = Boolean.TRUE.equals(evalCase.getRequireV2AuditClosure());
@@ -249,6 +254,11 @@ public class DiagnosisTraceEvaluator {
&& !safeList(evalCase.getExpectedGatekeeperStatuses()).contains(gatekeeperStatus)) { && !safeList(evalCase.getExpectedGatekeeperStatuses()).contains(gatekeeperStatus)) {
failedChecks.add("gatekeeper status not expected: " + valueOrMissing(gatekeeperStatus)); failedChecks.add("gatekeeper status not expected: " + valueOrMissing(gatekeeperStatus));
} }
if (!isBlank(evalCase.getExpectedGatekeeperRuleSetVersion())
&& !evalCase.getExpectedGatekeeperRuleSetVersion().equals(gatekeeperRuleSetVersion)) {
failedChecks.add("gatekeeper rule set version not expected: "
+ valueOrMissing(gatekeeperRuleSetVersion));
}
failedChecks.addAll(validateClaimChecks(trace, requireClaimChecks)); failedChecks.addAll(validateClaimChecks(trace, requireClaimChecks));
@@ -9,6 +9,7 @@ import com.fasterxml.jackson.core.type.TypeReference;
import com.fasterxml.jackson.databind.JsonNode; import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper; import com.fasterxml.jackson.databind.ObjectMapper;
import com.superbiz.agent.service.ExecutorGatekeeperService; import com.superbiz.agent.service.ExecutorGatekeeperService;
import com.superbiz.agent.service.GatekeeperRuleCatalog;
import com.superbiz.agent.service.ToolTraceSummaryService; import com.superbiz.agent.service.ToolTraceSummaryService;
import com.superbiz.agent.util.SessionContextHolder; import com.superbiz.agent.util.SessionContextHolder;
import com.superbiz.agent.util.VerifierContextHolder; import com.superbiz.agent.util.VerifierContextHolder;
@@ -110,9 +111,12 @@ public class VerifierInputHook extends MessagesModelHook {
} }
private Map<String, Object> passGatekeeperResult() { private Map<String, Object> passGatekeeperResult() {
GatekeeperRuleCatalog catalog = GatekeeperRuleCatalog.fallback();
Map<String, Object> result = new LinkedHashMap<>(); Map<String, Object> result = new LinkedHashMap<>();
result.put("status", "pass"); result.put("status", "pass");
result.put("severity", "none"); result.put("severity", "none");
result.put("rule_set_version", catalog.version());
result.put("rules", catalog.auditRules());
result.put("checked_bindings", List.of()); result.put("checked_bindings", List.of());
result.put("failed_rules", List.of()); result.put("failed_rules", List.of());
result.put("warnings", List.of()); result.put("warnings", List.of());
@@ -892,8 +892,7 @@ public class ChatService {
Optional.ofNullable(VerifierContextHolder.getToolTraceSummary()).orElse(List.of())); Optional.ofNullable(VerifierContextHolder.getToolTraceSummary()).orElse(List.of()));
verifierEvaluation.put("gatekeeper_result", verifierEvaluation.put("gatekeeper_result",
Optional.ofNullable(VerifierContextHolder.getGatekeeperResult()) Optional.ofNullable(VerifierContextHolder.getGatekeeperResult())
.orElse(Map.of("status", "pass", "severity", "none", "checked_bindings", List.of(), .orElse(defaultGatekeeperPass()));
"failed_rules", List.of(), "warnings", List.of(), "errors", List.of())));
if (composerOutput != null) { if (composerOutput != null) {
verifierEvaluation.put("composer_output", composerOutput); verifierEvaluation.put("composer_output", composerOutput);
} }
@@ -903,6 +902,20 @@ public class ChatService {
diagnosisSessionRepository.save(session); diagnosisSessionRepository.save(session);
} }
private Map<String, Object> defaultGatekeeperPass() {
GatekeeperRuleCatalog catalog = GatekeeperRuleCatalog.fallback();
return Map.of(
"status", "pass",
"severity", "none",
"rule_set_version", catalog.version(),
"rules", catalog.auditRules(),
"checked_bindings", List.of(),
"failed_rules", List.of(),
"warnings", List.of(),
"errors", List.of()
);
}
private ComposerRenderResult composeFinalAnswer(ChatModel chatModel, String originalQuery, private ComposerRenderResult composeFinalAnswer(ChatModel chatModel, String originalQuery,
VerifierDecision decision, RunnableConfig config) { VerifierDecision decision, RunnableConfig config) {
Map<String, Object> composerInput = buildComposerInput(originalQuery, decision); Map<String, Object> composerInput = buildComposerInput(originalQuery, decision);
@@ -4,6 +4,7 @@ import com.fasterxml.jackson.core.type.TypeReference;
import com.fasterxml.jackson.databind.ObjectMapper; import com.fasterxml.jackson.databind.ObjectMapper;
import com.superbiz.agent.domain.entity.ToolInvocation; import com.superbiz.agent.domain.entity.ToolInvocation;
import com.superbiz.agent.repository.ToolInvocationRepository; import com.superbiz.agent.repository.ToolInvocationRepository;
import org.springframework.beans.factory.annotation.Autowired;
import org.springframework.stereotype.Service; import org.springframework.stereotype.Service;
import java.util.ArrayList; import java.util.ArrayList;
@@ -35,19 +36,29 @@ public class ExecutorGatekeeperService {
private static final TypeReference<Map<String, Object>> MAP_TYPE = new TypeReference<>() { private static final TypeReference<Map<String, Object>> MAP_TYPE = new TypeReference<>() {
}; };
private static final double MIN_TOKEN_OVERLAP = 0.5; private static final double DEFAULT_MIN_TOKEN_OVERLAP = 0.5;
private final ToolInvocationRepository toolInvocationRepository; private final ToolInvocationRepository toolInvocationRepository;
private final GatekeeperRuleCatalog ruleCatalog;
private final ObjectMapper objectMapper = new ObjectMapper(); private final ObjectMapper objectMapper = new ObjectMapper();
@Autowired
public ExecutorGatekeeperService(ToolInvocationRepository toolInvocationRepository) { public ExecutorGatekeeperService(ToolInvocationRepository toolInvocationRepository) {
this(toolInvocationRepository, null);
}
public ExecutorGatekeeperService(ToolInvocationRepository toolInvocationRepository,
GatekeeperRuleCatalog ruleCatalog) {
this.toolInvocationRepository = toolInvocationRepository; this.toolInvocationRepository = toolInvocationRepository;
this.ruleCatalog = ruleCatalog == null
? GatekeeperRuleCatalog.loadDefault(objectMapper)
: ruleCatalog;
} }
public Map<String, Object> validate(String sessionId, public Map<String, Object> validate(String sessionId,
Map<String, Object> structuredOutput, Map<String, Object> structuredOutput,
Map<String, Object> parseStatus) { Map<String, Object> parseStatus) {
GatekeeperResult result = new GatekeeperResult(); GatekeeperResult result = new GatekeeperResult(ruleCatalog);
validateSchema(structuredOutput, parseStatus, result); validateSchema(structuredOutput, parseStatus, result);
if (structuredOutput != null) { if (structuredOutput != null) {
validateInvocationRefs(sessionId, structuredOutput, result); validateInvocationRefs(sessionId, structuredOutput, result);
@@ -57,11 +68,11 @@ public class ExecutorGatekeeperService {
} }
public Map<String, Object> pass() { public Map<String, Object> pass() {
return new GatekeeperResult().toMap(); return new GatekeeperResult(ruleCatalog).toMap();
} }
public Map<String, Object> fail(String ruleId, String target, String message) { public Map<String, Object> fail(String ruleId, String target, String message) {
GatekeeperResult result = new GatekeeperResult(); GatekeeperResult result = new GatekeeperResult(ruleCatalog);
result.fail(ruleId, target, message, SEVERITY_REJECT); result.fail(ruleId, target, message, SEVERITY_REJECT);
return result.toMap(); return result.toMap();
} }
@@ -156,7 +167,8 @@ public class ExecutorGatekeeperService {
SEVERITY_LOW_CONFID); SEVERITY_LOW_CONFID);
continue; continue;
} }
validateEvidenceBinding(binding, validInvocations, target, claim.get("claim_id"), result); validateEvidenceBinding(binding, validInvocations, target,
claim.get("claim_id"), claim.get("claim_type"), result);
} }
} }
@@ -181,7 +193,7 @@ public class ExecutorGatekeeperService {
SEVERITY_LOW_CONFID); SEVERITY_LOW_CONFID);
continue; continue;
} }
validateEvidenceBinding(binding, validInvocations, target, action.get("action_id"), result); validateEvidenceBinding(binding, validInvocations, target, action.get("action_id"), null, result);
} }
} }
} }
@@ -190,6 +202,7 @@ public class ExecutorGatekeeperService {
Map<Long, ToolInvocation> validInvocations, Map<Long, ToolInvocation> validInvocations,
String target, String target,
Object ownerId, Object ownerId,
Object ownerType,
GatekeeperResult result) { GatekeeperResult result) {
Map<String, Object> checked = new LinkedHashMap<>(); Map<String, Object> checked = new LinkedHashMap<>();
checked.put("claim_id", ownerId == null ? "" : String.valueOf(ownerId)); checked.put("claim_id", ownerId == null ? "" : String.valueOf(ownerId));
@@ -197,18 +210,6 @@ public class ExecutorGatekeeperService {
checked.put("source_invocation_id", binding.get("source_invocation_id")); checked.put("source_invocation_id", binding.get("source_invocation_id"));
checked.put("raw_path", stringValue(binding.get("raw_path"))); checked.put("raw_path", stringValue(binding.get("raw_path")));
Long id = singleInvocationId(binding);
if (id == null) {
checked.put("status", STATUS_FAIL);
checked.put("rule", RULE_INVOCATION_REF);
checked.put("message", "source_invocation_id is required");
result.checked(checked);
result.fail(RULE_INVOCATION_REF, target + ".source_invocation_id",
"source_invocation_id is required", SEVERITY_LOW_CONFID);
return;
}
checked.put("source_invocation_id", id);
String claimedToolName = stringValue(binding.get("tool_name")); String claimedToolName = stringValue(binding.get("tool_name"));
if (claimedToolName.isBlank()) { if (claimedToolName.isBlank()) {
checked.put("status", STATUS_FAIL); checked.put("status", STATUS_FAIL);
@@ -219,6 +220,46 @@ public class ExecutorGatekeeperService {
return; return;
} }
String rawPath = stringValue(binding.get("raw_path"));
if ("negative_observation".equals(stringValue(ownerType)) && !rawPath.isBlank()
&& !"$.no_evidence".equals(rawPath)) {
checked.put("status", STATUS_FAIL);
checked.put("rule", RULE_RAW_PATH);
checked.put("message", "negative_observation must only bind $.no_evidence references");
result.checked(checked);
result.fail(RULE_RAW_PATH, target + ".raw_path",
"negative_observation must only bind $.no_evidence references", SEVERITY_REJECT);
return;
}
Long id = singleInvocationId(binding);
if (id == null) {
id = uniqueInvocationIdByToolRawPathAndExcerpt(validInvocations, claimedToolName, rawPath,
stringValue(binding.get("evidence_excerpt")));
if (id == null) {
id = uniqueInvocationIdByToolAndRawPath(validInvocations, claimedToolName, rawPath);
}
if (id != null) {
result.warn(Map.of(
"rule", "evidence.invocation_auto_backfill_by_raw_path",
"message", "source_invocation_id was auto-filled from the unique evidence reference candidate",
"tool_name", claimedToolName,
"raw_path", rawPath,
"source_invocation_id", id
));
}
}
if (id == null) {
checked.put("status", STATUS_FAIL);
checked.put("rule", RULE_INVOCATION_REF);
checked.put("message", "source_invocation_id is required");
result.checked(checked);
result.fail(RULE_INVOCATION_REF, target + ".source_invocation_id",
"source_invocation_id is required", SEVERITY_LOW_CONFID);
return;
}
checked.put("source_invocation_id", id);
ToolInvocation invocation = validInvocations.get(id); ToolInvocation invocation = validInvocations.get(id);
if (invocation == null) { if (invocation == null) {
checked.put("status", STATUS_FAIL); checked.put("status", STATUS_FAIL);
@@ -240,7 +281,6 @@ public class ExecutorGatekeeperService {
return; return;
} }
String rawPath = stringValue(binding.get("raw_path"));
if (rawPath.isBlank()) { if (rawPath.isBlank()) {
checked.put("status", STATUS_FAIL); checked.put("status", STATUS_FAIL);
checked.put("rule", RULE_RAW_PATH); checked.put("rule", RULE_RAW_PATH);
@@ -298,6 +338,82 @@ public class ExecutorGatekeeperService {
result.checked(checked); result.checked(checked);
} }
private Long uniqueInvocationIdByToolAndRawPath(Map<Long, ToolInvocation> validInvocations,
String toolName,
String rawPath) {
if (toolName == null || toolName.isBlank() || rawPath == null || rawPath.isBlank()) {
return null;
}
Long matchedId = null;
for (Map.Entry<Long, ToolInvocation> entry : validInvocations.entrySet()) {
ToolInvocation invocation = entry.getValue();
if (!Objects.equals(toolName, invocation.getToolName())) {
continue;
}
if (!evidenceRefsByRawPath(invocation.getRetrievalDetails()).containsKey(rawPath)) {
continue;
}
if (matchedId != null) {
return null;
}
matchedId = entry.getKey();
}
return matchedId;
}
private Long uniqueInvocationIdByToolRawPathAndExcerpt(Map<Long, ToolInvocation> validInvocations,
String toolName,
String rawPath,
String excerpt) {
if (toolName == null || toolName.isBlank()
|| rawPath == null || rawPath.isBlank()
|| excerpt == null || excerpt.isBlank()) {
return null;
}
Long matchedId = null;
for (Map.Entry<Long, ToolInvocation> entry : validInvocations.entrySet()) {
ToolInvocation invocation = entry.getValue();
if (!Objects.equals(toolName, invocation.getToolName())) {
continue;
}
String matchedText = evidenceRefsByRawPath(invocation.getRetrievalDetails()).get(rawPath);
if (matchedText == null || !isBackfillCandidateSupported(rawPath, excerpt, matchedText)) {
continue;
}
if (matchedId != null) {
return null;
}
matchedId = entry.getKey();
}
return matchedId;
}
private boolean isBackfillCandidateSupported(String rawPath, String excerpt, String matchedText) {
if ("$.no_evidence".equals(rawPath)) {
String excerptQuery = semicolonField(excerpt, "query");
String matchedQuery = semicolonField(matchedText, "query");
if (!excerptQuery.isBlank() && !matchedQuery.isBlank()
&& !normalized(excerptQuery).equals(normalized(matchedQuery))) {
return false;
}
}
return isExcerptSupported(excerpt, matchedText);
}
private String semicolonField(String text, String field) {
if (text == null || text.isBlank() || field == null || field.isBlank()) {
return "";
}
String prefix = field + "=";
for (String part : text.split(";")) {
String trimmed = part.trim();
if (trimmed.regionMatches(true, 0, prefix, 0, prefix.length())) {
return trimmed.substring(prefix.length()).trim();
}
}
return "";
}
private Long singleInvocationId(Map<?, ?> binding) { private Long singleInvocationId(Map<?, ?> binding) {
Long singular = asLong(binding.get("source_invocation_id")); Long singular = asLong(binding.get("source_invocation_id"));
if (singular != null) { if (singular != null) {
@@ -369,7 +485,9 @@ public class ExecutorGatekeeperService {
overlap++; overlap++;
} }
} }
return (double) overlap / excerptTokens.size() >= MIN_TOKEN_OVERLAP; double minTokenOverlap = ruleCatalog.doubleParameter(RULE_EXCERPT_MISMATCH,
"min_token_overlap", DEFAULT_MIN_TOKEN_OVERLAP);
return (double) overlap / excerptTokens.size() >= minTokenOverlap;
} }
private String normalized(String value) { private String normalized(String value) {
@@ -422,12 +540,17 @@ public class ExecutorGatekeeperService {
} }
private static final class GatekeeperResult { private static final class GatekeeperResult {
private final GatekeeperRuleCatalog ruleCatalog;
private final List<String> failedRules = new ArrayList<>(); private final List<String> failedRules = new ArrayList<>();
private final List<Map<String, Object>> checkedBindings = new ArrayList<>(); private final List<Map<String, Object>> checkedBindings = new ArrayList<>();
private final List<Map<String, Object>> warnings = new ArrayList<>(); private final List<Map<String, Object>> warnings = new ArrayList<>();
private final List<Map<String, Object>> errors = new ArrayList<>(); private final List<Map<String, Object>> errors = new ArrayList<>();
private String severity = SEVERITY_NONE; private String severity = SEVERITY_NONE;
GatekeeperResult(GatekeeperRuleCatalog ruleCatalog) {
this.ruleCatalog = ruleCatalog == null ? GatekeeperRuleCatalog.fallback() : ruleCatalog;
}
void fail(String ruleId, String target, String message, String failureSeverity) { void fail(String ruleId, String target, String message, String failureSeverity) {
if (!failedRules.contains(ruleId)) { if (!failedRules.contains(ruleId)) {
failedRules.add(ruleId); failedRules.add(ruleId);
@@ -458,6 +581,8 @@ public class ExecutorGatekeeperService {
Map<String, Object> result = new LinkedHashMap<>(); Map<String, Object> result = new LinkedHashMap<>();
result.put("status", failedRules.isEmpty() ? STATUS_PASS : STATUS_FAIL); result.put("status", failedRules.isEmpty() ? STATUS_PASS : STATUS_FAIL);
result.put("severity", failedRules.isEmpty() ? SEVERITY_NONE : severity); result.put("severity", failedRules.isEmpty() ? SEVERITY_NONE : severity);
result.put("rule_set_version", ruleCatalog.version());
result.put("rules", ruleCatalog.auditRules());
result.put("checked_bindings", checkedBindings); result.put("checked_bindings", checkedBindings);
result.put("failed_rules", failedRules); result.put("failed_rules", failedRules);
result.put("warnings", warnings); result.put("warnings", warnings);
@@ -0,0 +1,142 @@
package com.superbiz.agent.service;
import com.fasterxml.jackson.core.type.TypeReference;
import com.fasterxml.jackson.databind.ObjectMapper;
import java.io.InputStream;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Map;
/**
* Lightweight metadata catalog for deterministic Gatekeeper rules.
*/
public class GatekeeperRuleCatalog {
public static final String DEFAULT_RESOURCE = "gatekeeper/gatekeeper-rules.json";
public static final String FALLBACK_VERSION = "gatekeeper-rules-v1";
private static final TypeReference<Map<String, Object>> MAP_TYPE = new TypeReference<>() {
};
private final String version;
private final List<RuleMetadata> rules;
public GatekeeperRuleCatalog(String version, List<RuleMetadata> rules) {
this.version = version == null || version.isBlank() ? FALLBACK_VERSION : version;
this.rules = List.copyOf(rules == null ? List.of() : rules);
}
public static GatekeeperRuleCatalog loadDefault(ObjectMapper objectMapper) {
try (InputStream input = GatekeeperRuleCatalog.class.getClassLoader()
.getResourceAsStream(DEFAULT_RESOURCE)) {
if (input == null) {
return fallback();
}
Map<String, Object> root = objectMapper.readValue(input, MAP_TYPE);
String version = stringValue(root.get("version"));
List<RuleMetadata> rules = new ArrayList<>();
Object rulesValue = root.get("rules");
if (rulesValue instanceof List<?> ruleList) {
for (Object ruleValue : ruleList) {
if (ruleValue instanceof Map<?, ?> ruleMap) {
rules.add(RuleMetadata.from(ruleMap));
}
}
}
return new GatekeeperRuleCatalog(version, rules);
} catch (Exception ignored) {
return fallback();
}
}
public static GatekeeperRuleCatalog fallback() {
return new GatekeeperRuleCatalog(FALLBACK_VERSION, List.of(
new RuleMetadata("schema.executor_v2", "Executor output must match executor_evidence_v2 schema",
true, "low_confid", Map.of()),
new RuleMetadata("evidence.invocation_ref", "source_invocation_id must refer to a real current-session tool invocation",
true, "reject", Map.of()),
new RuleMetadata("evidence.raw_path", "raw_path must exist in retrieval_details.evidence_refs",
true, "reject", Map.of()),
new RuleMetadata("evidence.excerpt_mismatch", "evidence_excerpt must be supported by the matched evidence ref text",
true, "reject", Map.of("min_token_overlap", 0.5)),
new RuleMetadata("evidence.missing", "claims must include usable evidence bindings",
true, "low_confid", Map.of())
));
}
public String version() {
return version;
}
public List<Map<String, Object>> auditRules() {
List<Map<String, Object>> result = new ArrayList<>();
for (RuleMetadata rule : rules) {
if (rule.enabled()) {
result.add(rule.toAuditMap());
}
}
return result;
}
public double doubleParameter(String ruleId, String parameterName, double fallback) {
for (RuleMetadata rule : rules) {
if (!rule.id().equals(ruleId) || !rule.enabled()) {
continue;
}
Object value = rule.parameters().get(parameterName);
if (value instanceof Number number) {
return number.doubleValue();
}
if (value instanceof String text) {
try {
return Double.parseDouble(text);
} catch (NumberFormatException ignored) {
return fallback;
}
}
}
return fallback;
}
private static String stringValue(Object value) {
return value == null ? "" : String.valueOf(value);
}
public record RuleMetadata(String id,
String description,
boolean enabled,
String defaultSeverity,
Map<String, Object> parameters) {
static RuleMetadata from(Map<?, ?> raw) {
String id = stringValue(raw.get("id"));
String description = stringValue(raw.get("description"));
boolean enabled = !(raw.get("enabled") instanceof Boolean value) || value;
String defaultSeverity = stringValue(raw.get("default_severity"));
Map<String, Object> parameters = new LinkedHashMap<>();
Object parametersValue = raw.get("parameters");
if (parametersValue instanceof Map<?, ?> parameterMap) {
for (Map.Entry<?, ?> entry : parameterMap.entrySet()) {
if (entry.getKey() != null) {
parameters.put(String.valueOf(entry.getKey()), entry.getValue());
}
}
}
return new RuleMetadata(id, description, enabled, defaultSeverity, parameters);
}
Map<String, Object> toAuditMap() {
Map<String, Object> result = new LinkedHashMap<>();
result.put("id", id);
result.put("description", description);
result.put("enabled", enabled);
result.put("default_severity", defaultSeverity);
if (!parameters.isEmpty()) {
result.put("parameters", parameters);
}
return result;
}
}
}
@@ -88,7 +88,7 @@ public class ToolInvocationRecorder {
if (extraDetails != null && !extraDetails.isEmpty()) { if (extraDetails != null && !extraDetails.isEmpty()) {
details.putAll(extraDetails); details.putAll(extraDetails);
} }
List<Map<String, Object>> evidenceRefs = extractEvidenceRefs(toolName, output, details); List<Map<String, Object>> evidenceRefs = extractEvidenceRefs(toolName, inputParams, output, details);
if (!evidenceRefs.isEmpty()) { if (!evidenceRefs.isEmpty()) {
details.put("evidence_refs", evidenceRefs); details.put("evidence_refs", evidenceRefs);
} }
@@ -145,6 +145,12 @@ public class ToolInvocationRecorder {
} }
details.put("evidence_blocks", record.evidenceBlocks() == null ? List.of() : record.evidenceBlocks()); details.put("evidence_blocks", record.evidenceBlocks() == null ? List.of() : record.evidenceBlocks());
List<Map<String, Object>> evidenceRefs = evidenceRefsFromEvidenceBlocks(record.evidenceBlocks()); List<Map<String, Object>> evidenceRefs = evidenceRefsFromEvidenceBlocks(record.evidenceBlocks());
if (evidenceRefs.isEmpty()
&& EVIDENCE_STATUS_NO_EVIDENCE.equals(normalizeEvidenceStatus(record.success(), record.evidenceStatus()))) {
Map<String, Object> input = new LinkedHashMap<>();
input.put("query", record.query());
evidenceRefs = List.of(noEvidenceRef("lookup_knowledge", input, null, details));
}
if (!evidenceRefs.isEmpty()) { if (!evidenceRefs.isEmpty()) {
details.put("evidence_refs", evidenceRefs); details.put("evidence_refs", evidenceRefs);
} }
@@ -195,7 +201,10 @@ public class ToolInvocationRecorder {
return success ? EVIDENCE_STATUS_SUPPORTED : EVIDENCE_STATUS_FAILED; return success ? EVIDENCE_STATUS_SUPPORTED : EVIDENCE_STATUS_FAILED;
} }
private List<Map<String, Object>> extractEvidenceRefs(String toolName, String output, Map<String, Object> details) { private List<Map<String, Object>> extractEvidenceRefs(String toolName,
Map<String, Object> inputParams,
String output,
Map<String, Object> details) {
if (output == null || output.isBlank()) { if (output == null || output.isBlank()) {
return List.of(); return List.of();
} }
@@ -205,11 +214,14 @@ public class ToolInvocationRecorder {
try { try {
JsonNode root = objectMapper.readTree(output); JsonNode root = objectMapper.readTree(output);
boolean noEvidence = EVIDENCE_STATUS_NO_EVIDENCE.equals(stringValue(details.get("evidence_status")));
if ("query_metrics".equals(toolName)) { if ("query_metrics".equals(toolName)) {
return evidenceRefsFromArray(root.path("alerts"), "$.alerts", this::alertText); List<Map<String, Object>> refs = evidenceRefsFromArray(root.path("alerts"), "$.alerts", this::alertText);
return refs.isEmpty() && noEvidence ? List.of(noEvidenceRef(toolName, inputParams, root, details)) : refs;
} }
if ("query_logs".equals(toolName)) { if ("query_logs".equals(toolName)) {
return evidenceRefsFromArray(root.path("logs"), "$.logs", this::logText); List<Map<String, Object>> refs = evidenceRefsFromArray(root.path("logs"), "$.logs", this::logText);
return refs.isEmpty() && noEvidence ? List.of(noEvidenceRef(toolName, inputParams, root, details)) : refs;
} }
} catch (Exception e) { } catch (Exception e) {
log.debug("extract evidence_refs failed for tool={}", toolName, e); log.debug("extract evidence_refs failed for tool={}", toolName, e);
@@ -217,6 +229,42 @@ public class ToolInvocationRecorder {
return List.of(); return List.of();
} }
private Map<String, Object> noEvidenceRef(String toolName,
Map<String, Object> inputParams,
JsonNode root,
Map<String, Object> details) {
List<String> parts = new ArrayList<>();
addPart(parts, toolName + " returned no evidence");
addPart(parts, "evidence_status=" + stringValue(details.get("evidence_status")));
String query = firstNonBlank(
root == null ? null : textField(root, "query"),
inputParams == null ? null : inputParams.get("query")
);
if (!query.isBlank()) {
addPart(parts, "query=" + query);
}
String topic = firstNonBlank(
root == null ? null : textField(root, "log_topic"),
inputParams == null ? null : inputParams.get("log_topic"),
details == null ? null : details.get("log_topic"),
details == null ? null : details.get("metric_family")
);
if (!topic.isBlank()) {
addPart(parts, "topic=" + topic);
}
if (root != null && root.has("total")) {
addPart(parts, "total=" + root.path("total").asText());
}
String message = root == null ? "" : textField(root, "message");
if (!message.isBlank()) {
addPart(parts, "message=" + message);
}
return Map.of(
"raw_path", "$.no_evidence",
"text", bounded(String.join("; ", parts), 500)
);
}
private List<Map<String, Object>> evidenceRefsFromArray(JsonNode arrayNode, private List<Map<String, Object>> evidenceRefsFromArray(JsonNode arrayNode,
String pathPrefix, String pathPrefix,
java.util.function.Function<JsonNode, String> textExtractor) { java.util.function.Function<JsonNode, String> textExtractor) {
@@ -314,6 +362,10 @@ public class ToolInvocationRecorder {
return ""; return "";
} }
private String stringValue(Object value) {
return value == null ? "" : String.valueOf(value);
}
private String bounded(String value, int limit) { private String bounded(String value, int limit) {
if (value == null) { if (value == null) {
return ""; return "";
@@ -0,0 +1,38 @@
{
"version": "gatekeeper-rules-v1",
"rules": [
{
"id": "schema.executor_v2",
"description": "Executor output must match executor_evidence_v2 schema",
"enabled": true,
"default_severity": "low_confid"
},
{
"id": "evidence.invocation_ref",
"description": "source_invocation_id must refer to a real current-session tool invocation",
"enabled": true,
"default_severity": "reject"
},
{
"id": "evidence.raw_path",
"description": "raw_path must exist in retrieval_details.evidence_refs",
"enabled": true,
"default_severity": "reject"
},
{
"id": "evidence.excerpt_mismatch",
"description": "evidence_excerpt must be supported by the matched evidence ref text",
"enabled": true,
"default_severity": "reject",
"parameters": {
"min_token_overlap": 0.5
}
},
{
"id": "evidence.missing",
"description": "claims must include usable evidence bindings",
"enabled": true,
"default_severity": "low_confid"
}
]
}
@@ -8,6 +8,8 @@
- 你只能使用输入中的 `allowed_claims`、`allowed_hypotheses`、`missing_info`、`recommended_actions`、`rationale`。 - 你只能使用输入中的 `allowed_claims`、`allowed_hypotheses`、`missing_info`、`recommended_actions`、`rationale`。
- 禁止使用模型经验添加新的服务名、订单号、时间、指标值、错误码、根因或修复理由。 - 禁止使用模型经验添加新的服务名、订单号、时间、指标值、错误码、根因或修复理由。
- 只输出一个合法 JSON 对象,不输出 Markdown,不输出代码块,不输出额外说明。 - 只输出一个合法 JSON 对象,不输出 Markdown,不输出代码块,不输出额外说明。
- 当 `allowed_claims` 中存在 `claim_type=negative_observation`,或证据来自 `$.no_evidence` 时,只能表达“当前查询未检索到 / 本次检索未发现匹配证据”。
- 对 `negative_observation` / `$.no_evidence`,禁止表达“问题不存在”“已排除该问题”“确认没有”“日志层面已排除”等过度结论。
## 输入字段 ## 输入字段
@@ -26,6 +28,7 @@
- 可以表达确认结论。 - 可以表达确认结论。
- 只能使用 `allowed_claims` 和 `recommended_actions`。 - 只能使用 `allowed_claims` 和 `recommended_actions`。
- 只有当 `allowed_claims` 中存在 `claim_type=root_cause` 的 claim 时,才允许表达“根因已确认”。 - 只有当 `allowed_claims` 中存在 `claim_type=root_cause` 的 claim 时,才允许表达“根因已确认”。
- 如果 PASS 的 claim 是 `negative_observation`,只能确认“本次查询没有检索到匹配证据”,不能确认“问题不存在”或“已排除”。
### LOW_CONFID ### LOW_CONFID
@@ -7,15 +7,114 @@
- 执行完成后,输出严格的证据归因 JSON,供 Verifier 校验。 - 执行完成后,输出严格的证据归因 JSON,供 Verifier 校验。
- 你是证据收集与微观事实提炼器,不是最终答复生成器。 - 你是证据收集与微观事实提炼器,不是最终答复生成器。
## 角色边界 HARD-GATE
你只负责证据收集与微观事实提炼,只能输出“当前工具证据可以直接支持的观察事实”。
你不是:
- 根因诊断器。
- 修复方案生成器。
- Runbook 转述器。
- 经验推断器。
- 最终用户答复生成器。
除非本轮工具返回中存在直接证据,否则禁止输出:
- 根因确认。
- 修复建议。
- 扩展排查方向。
- 历史经验。
- 通用知识。
- 与用户问题无关的服务、指标、订单、错误码、组件。
## 规则 ## 规则
- 按顺序执行,不可跳过步骤。 - 按顺序执行,不可跳过步骤。
- 所有事实性结论必须来自本轮 evidence tools 的返回。 - 所有事实性结论必须来自本轮 evidence tools 的返回。
- runbook、skill、历史案例、知识库中的通用模式只能作为排查指导或建议动作,不能直接写成本次事故的已确认事实。 - runbook、skill、历史案例、知识库中的通用模式只能作为排查指导,不能直接写成本次事故的已确认事实。
- 如果检索内容不足以支撑结论,必须显式声明证据不足,严禁补全事故故事。 - 如果检索内容不足以支撑结论,必须显式声明证据不足,严禁补全事故故事。
- 不要使用“通常情况下”“根据经验”“很可能已经发生”等无证据推断词来伪装事实。 - 不要使用“通常情况下”“根据经验”“很可能已经发生”“可能是”“推测”“理论上”等无证据推断词来伪装事实。
- 对窄范围确认问题,只输出与用户问题直接相关的 observation / negative_observation。通常 1 条 claim,最多 2 条 claim;不要限制 evidence_bindings 数量。
- 禁止把根因、修复动作或用户明确排除的服务/主题写成 confirmed claim,除非本轮工具证据直接证明。 - 禁止把根因、修复动作或用户明确排除的服务/主题写成 confirmed claim,除非本轮工具证据直接证明。
## 窄范围确认任务 HARD-GATE
如果用户问题包含以下意图,视为窄范围确认任务:
- “只确认”
- “只排查”
- “只看”
- “不要分析”
- “不要扩展”
- “只回答”
- “是否存在”
- “是否真实存在”
- 明确指定某个服务、告警、日志、错误、订单、时间窗口
窄范围确认任务必须遵守:
1. `claims` 只能输出 `observation` 或 `negative_observation`。
2. claim 数量必须是最少必要数量,通常 1 条,最多 2 条。
3. claim 数量限制不限制 `evidence_bindings` 数量;一条 claim 可以绑定多条直接相关证据。
4. 不得把同一观察事实拆成多条 claim。
5. 只能围绕用户明确要求的目标对象和主题输出 claim。
6. 用户明确排除的对象、服务、告警、订单、数据库、连接池、下游依赖,禁止出现在 claim 中。
7. 禁止输出根因类、风险类或建议类 claim,例如 `root_cause`、`risk`、`recommendation`。
8. 如果证据不足,不要补合理化解释;优先写入 `missing_info`。
9. 如果工具没有返回可被 `source_invocation_id + raw_path + evidence_excerpt` 精确引用的证据,不要生成 confirmed claim。
10. 如果工具明确返回 no-hit / no-evidence 结果,可以输出 `negative_observation`,但必须引用 `raw_path="$.no_evidence"`。
11. 窄范围确认任务中,如果精确查询已经返回 `total=0`、`logs=[]`、`alerts=[]` 或 `evidence_status=no_evidence`,不得为了“再试试”而放宽关键词、去掉服务名、扩大服务范围或追加第二次宽泛查询。
窄范围任务的理想输出是:
- 1 条核心 claim。
- 多条直接相关 `evidence_bindings`。
- 必要的 `missing_info`。
同一条工具数组项只能绑定一次。不要为了引用其中多个字段而拆成多个 `evidence_bindings`。
正确:
```json
{
"raw_path": "$.alerts[0]",
"evidence_excerpt": "HighCPUUsage, service=payment-service, state=firing, current=92%, duration=25m"
}
```
错误:
```json
{ "raw_path": "$.alerts[0].alert_name", "evidence_excerpt": "HighCPUUsage" }
{ "raw_path": "$.alerts[0].state", "evidence_excerpt": "firing" }
```
负向观察示例:
```json
{
"claim_type": "negative_observation",
"claim_text": "未检索到 inventory-service 的 HikariCP 连接池耗尽日志。",
"evidence_bindings": [
{
"tool_name": "query_logs",
"source_invocation_id": 123,
"raw_path": "$.no_evidence",
"evidence_excerpt": "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence"
}
]
}
```
`$.no_evidence` 只表示“该工具对当前查询返回无匹配证据”,不能表示“问题不存在”或“根因被排除”。没有实际调用工具时,禁止使用 `$.no_evidence`。
`negative_observation` 的 `evidence_bindings` 只能绑定 `$.no_evidence`。禁止把其它服务的正向日志或告警绑定到同一个 `negative_observation`,即使这些日志可以说明“不是当前服务”。
输出 `negative_observation` 或基于 `$.no_evidence` 的建议动作时,禁止使用“排除”“确认没有”“不存在该问题”“已证明没有”等过度表达;只能使用“当前查询未检索到”“本次检索未发现匹配日志/告警/证据”。
## 工具使用边界
你只能调用回答当前用户问题所必需的工具。
- 问告警状态:优先使用 `query_metrics`。
- 问日志现象:优先使用 `query_logs`。
- 问知识解释或排查步骤:才使用 `lookup_knowledge`。
- Runbook / Skill / 知识库只能帮助决定“查什么”,不能直接作为“当前环境发生了什么”的证据。
- 如果当前工具结果已经足以回答用户问题,不要继续扩展检索。
- 不要为了补全故事而查询用户没有要求的服务、组件或故障类型。
- 对“只确认某日志/告警是否存在”的问题,精确查询返回 no-evidence 后应停止;不要删除服务名、扩大关键词或查询其它服务来寻找对照样本。
## 检索约束 ## 检索约束
### 1. 判断重复:基于已检索上下文 ### 1. 判断重复:基于已检索上下文
@@ -67,6 +166,19 @@
### missing_info ### missing_info
`missing_info` 用来列出无法确认结论所缺少的具体证据。 `missing_info` 用来列出无法确认结论所缺少的具体证据。
## 输出前自检
在输出 JSON 前,逐项检查:
1. 每条 claim 是否直接回答了用户当前问题?
2. 每条 claim 是否都有真实 `evidence_bindings`?
3. 每个 `evidence_excerpt` 是否来自工具返回原文?
4. 是否出现了用户没有要求的服务、告警、订单、数据库、连接池或下游组件?
5. 是否把 Runbook / Skill / 知识库通用内容写成了当前事实?
6. 是否输出了根因、修复动作、风险判断或经验推断?
只要任一项不通过,删除对应 claim,不要解释。
## 最终输出格式(严格契约) ## 最终输出格式(严格契约)
你必须输出且只能输出一个 JSON 对象,不要输出 Markdown,不要输出代码块,不要输出 JSON 之外的解释文字。 你必须输出且只能输出一个 JSON 对象,不要输出 Markdown,不要输出代码块,不要输出 JSON 之外的解释文字。
@@ -87,7 +199,7 @@
"source_id": "工具返回中的 evidence block id、trace_ref 或可定位标识,可为空", "source_id": "工具返回中的 evidence block id、trace_ref 或可定位标识,可为空",
"tool_name": "lookup_knowledge/query_logs/query_metrics 等 evidence tool", "tool_name": "lookup_knowledge/query_logs/query_metrics 等 evidence tool",
"source_invocation_id": null, "source_invocation_id": null,
"raw_path": "$.alerts[0] / $.logs[0] / $.evidence_blocks[0]", "raw_path": "$.alerts[0] / $.logs[0] / $.evidence_blocks[0] / $.no_evidence",
"evidence_excerpt": "从工具返回中摘取的原话、指标值、日志片段或关键数据" "evidence_excerpt": "从工具返回中摘取的原话、指标值、日志片段或关键数据"
} }
] ]
@@ -121,6 +233,8 @@
- `claims[*].evidence_bindings` 不能为空。 - `claims[*].evidence_bindings` 不能为空。
- `evidence_excerpt` 必须来自工具返回,不允许编造。 - `evidence_excerpt` 必须来自工具返回,不允许编造。
- `raw_path` 必须指向工具返回数组中的具体条目:`query_metrics` 使用 `$.alerts[i]`,`query_logs` 使用 `$.logs[i]`,`lookup_knowledge` 使用 `$.evidence_blocks[i]`。 - `raw_path` 必须指向工具返回数组中的具体条目:`query_metrics` 使用 `$.alerts[i]`,`query_logs` 使用 `$.logs[i]`,`lookup_knowledge` 使用 `$.evidence_blocks[i]`。
- 当且仅当工具明确返回 no-hit / no-evidence 结果时,允许使用 `$.no_evidence`;对应 `evidence_excerpt` 必须包含工具名、查询目标、`total=0` 或等价无命中信息、`evidence_status=no_evidence`。
- `raw_path` 禁止指向字段级子路径,例如 `$.alerts[0].alert_name`、`$.alerts[0].state`、`$.logs[0].message` 都是非法路径。需要引用多个字段时,仍然只使用对应数组条目的 `raw_path`,并把必要字段合并进同一个 `evidence_excerpt`。
- `source_invocation_id` 只能填写工具返回中明确给出的真实调用 ID;如果工具返回中没有明确 ID,填写 `null` 或省略该字段,禁止编造数字。系统只会在唯一候选工具调用存在时补齐 ID,但不会补齐 `raw_path`。 - `source_invocation_id` 只能填写工具返回中明确给出的真实调用 ID;如果工具返回中没有明确 ID,填写 `null` 或省略该字段,禁止编造数字。系统只会在唯一候选工具调用存在时补齐 ID,但不会补齐 `raw_path`。
- 不要再输出 `source_invocation_ids` 作为主要字段;兼容旧字段不作为精确证据引用。 - 不要再输出 `source_invocation_ids` 作为主要字段;兼容旧字段不作为精确证据引用。
- 如果没有任何可确认事实,`claims` 返回空数组,并在 `missing_info` 说明缺少什么。 - 如果没有任何可确认事实,`claims` 返回空数组,并在 `missing_info` 说明缺少什么。
@@ -100,12 +100,12 @@ class DiagnosisEvalBaselineDiffTest {
} }
private void degradeRedisCase(DiagnosisEvalReport report) { private void degradeRedisCase(DiagnosisEvalReport report) {
report.setPassedCases(7); report.setPassedCases(9);
report.setPassRate(0.875); report.setPassRate(0.9);
report.setAverageToolCallCount(3.0); report.setAverageToolCallCount(3.0);
report.setAverageDurationMs(44875.0); report.setAverageDurationMs(39800.0);
report.setVerdictDistribution(new LinkedHashMap<>()); report.setVerdictDistribution(new LinkedHashMap<>());
report.getVerdictDistribution().put("PASS", 2L); report.getVerdictDistribution().put("PASS", 4L);
report.getVerdictDistribution().put("LOW_CONFID", 4L); report.getVerdictDistribution().put("LOW_CONFID", 4L);
report.getVerdictDistribution().put("REJECT", 2L); report.getVerdictDistribution().put("REJECT", 2L);
@@ -24,13 +24,23 @@ class DiagnosisTraceEvaluatorTest {
DiagnosisEvalReport report = evaluator.evaluate(cases, Path.of("mvp/eval/fixtures")); DiagnosisEvalReport report = evaluator.evaluate(cases, Path.of("mvp/eval/fixtures"));
assertEquals(8, report.getTotalCases()); assertEquals(10, report.getTotalCases());
assertEquals(8, report.getPassedCases()); assertEquals(10, report.getPassedCases());
assertEquals(1.0, report.getPassRate(), 0.001); assertEquals(1.0, report.getPassRate(), 0.001);
assertEquals(2L, report.getVerdictDistribution().get("PASS")); assertEquals(4L, report.getVerdictDistribution().get("PASS"));
assertEquals(5L, report.getVerdictDistribution().get("LOW_CONFID")); assertEquals(5L, report.getVerdictDistribution().get("LOW_CONFID"));
assertEquals(1L, report.getVerdictDistribution().get("REJECT")); assertEquals(1L, report.getVerdictDistribution().get("REJECT"));
DiagnosisEvalResult narrowHighCpu = result(report, "narrow-highcpu-observation");
assertTrue(narrowHighCpu.isPassed());
assertEquals("gatekeeper-rules-v1", narrowHighCpu.getGatekeeperRuleSetVersion());
assertEquals("pass", narrowHighCpu.getGatekeeperStatus());
DiagnosisEvalResult hikariNoEvidence = result(report, "hikari-no-evidence-negative-observation");
assertTrue(hikariNoEvidence.isPassed());
assertEquals("gatekeeper-rules-v1", hikariNoEvidence.getGatekeeperRuleSetVersion());
assertEquals("pass", hikariNoEvidence.getGatekeeperStatus());
DiagnosisEvalResult payment = result(report, "payment-timeout"); DiagnosisEvalResult payment = result(report, "payment-timeout");
assertTrue(payment.isPassed()); assertTrue(payment.isPassed());
assertTrue(payment.getEvidenceCoverage().get("lookup_knowledge")); assertTrue(payment.getEvidenceCoverage().get("lookup_knowledge"));
@@ -153,6 +163,39 @@ class DiagnosisTraceEvaluatorTest {
assertTrue(result.getFailedChecks().contains("gatekeeper fail cannot have PASS verdict")); assertTrue(result.getFailedChecks().contains("gatekeeper fail cannot have PASS verdict"));
} }
@Test
void evaluateFailsWhenGatekeeperRuleSetVersionMismatches() {
DiagnosisEvalCase evalCase = DiagnosisEvalCase.builder()
.id("rule-version")
.title("Rule version")
.expectedRootCauseKeywords(List.of())
.requiredEvidenceTools(List.of())
.allowedVerdicts(List.of("PASS"))
.expectedGatekeeperRuleSetVersion("gatekeeper-rules-v1")
.build();
DiagnosisTraceResponse trace = DiagnosisTraceResponse.builder()
.session(DiagnosisTraceResponse.SessionTrace.builder()
.answer("安全回答")
.selfEvaluation(java.util.Map.of(
"verifier_evaluation", java.util.Map.of(
"verdict", "PASS",
"gatekeeper_result", java.util.Map.of(
"status", "pass",
"rule_set_version", "old-rules"
)
)))
.build())
.toolInvocations(List.of())
.build();
DiagnosisEvalResult result = evaluator.evaluate(evalCase, trace);
assertFalse(result.isPassed());
assertTrue(result.getFailedChecks().contains(
"gatekeeper rule set version not expected: old-rules"));
}
@Test @Test
void evaluateFailsWhenUnsupportedClaimLeaksIntoFinalAnswer() { void evaluateFailsWhenUnsupportedClaimLeaksIntoFinalAnswer() {
DiagnosisEvalCase evalCase = DiagnosisEvalCase.builder() DiagnosisEvalCase evalCase = DiagnosisEvalCase.builder()
@@ -8,12 +8,25 @@ import java.util.List;
import java.util.Map; import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals; import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue; import static org.junit.jupiter.api.Assertions.assertTrue;
import static org.mockito.Mockito.mock; import static org.mockito.Mockito.mock;
import static org.mockito.Mockito.when; import static org.mockito.Mockito.when;
class ExecutorGatekeeperServiceTest { class ExecutorGatekeeperServiceTest {
@Test
void ruleCatalogLoadsDefaultMetadata() {
GatekeeperRuleCatalog catalog = GatekeeperRuleCatalog.loadDefault(new com.fasterxml.jackson.databind.ObjectMapper());
assertEquals("gatekeeper-rules-v1", catalog.version());
assertFalse(catalog.auditRules().isEmpty());
assertTrue(catalog.auditRules().stream()
.anyMatch(rule -> "evidence.raw_path".equals(rule.get("id"))));
assertEquals(0.5, catalog.doubleParameter("evidence.excerpt_mismatch",
"min_token_overlap", 0.0), 0.001);
}
@Test @Test
void validatePassesForExecutorEvidenceV2WithMatchingInvocation() { void validatePassesForExecutorEvidenceV2WithMatchingInvocation() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class); ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
@@ -30,9 +43,117 @@ class ExecutorGatekeeperServiceTest {
assertEquals("pass", result.get("status")); assertEquals("pass", result.get("status"));
assertEquals("none", result.get("severity")); assertEquals("none", result.get("severity"));
assertRuleAudit(result);
assertTrue(((List<?>) result.get("failed_rules")).isEmpty()); assertTrue(((List<?>) result.get("failed_rules")).isEmpty());
} }
@Test
void validatePassesForNoEvidenceReference() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
invocation(101L, "query_logs", "$.no_evidence",
"query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志")
));
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
Map<String, Object> result = service.validate("session-1",
validOutput(101L, "query_logs", "$.no_evidence",
"query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence"),
Map.of("status", "valid"));
assertEquals("pass", result.get("status"));
assertEquals("none", result.get("severity"));
assertRuleAudit(result);
assertTrue(((List<?>) result.get("failed_rules")).isEmpty());
}
@Test
void validateBackfillsNoEvidenceInvocationByRawPathWhenToolHasMultipleCalls() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
invocation(101L, "query_logs", "$.no_evidence",
"query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; total=0; message=未找到匹配的日志"),
invocation(102L, "query_logs", "$.logs[0]",
"order-service HikariCP active=50/50 waiting=32")
));
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
Map<String, Object> result = service.validate("session-1",
validOutput(null, "query_logs", "$.no_evidence",
"query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence",
"negative_observation"),
Map.of("status", "valid"));
assertEquals("pass", result.get("status"));
assertEquals("none", result.get("severity"));
assertTrue(((List<?>) result.get("failed_rules")).isEmpty());
assertEquals("evidence.invocation_auto_backfill_by_raw_path",
((Map<?, ?>) ((List<?>) result.get("warnings")).get(0)).get("rule"));
}
@Test
void validateBackfillsNoEvidenceInvocationByExcerptWhenRawPathIsRepeated() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
invocation(101L, "query_logs", "$.no_evidence",
"query_logs returned no evidence; evidence_status=no_evidence; query=service:inventory-service AND HikariCP; total=0; message=未找到匹配的日志"),
invocation(102L, "query_logs", "$.no_evidence",
"query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service; total=0; message=未找到匹配的日志")
));
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
Map<String, Object> result = service.validate("session-1",
validOutput(null, "query_logs", "$.no_evidence",
"query_logs returned no evidence; query=service:inventory-service AND HikariCP; total=0; evidence_status=no_evidence",
"negative_observation"),
Map.of("status", "valid"));
assertEquals("pass", result.get("status"));
assertEquals("none", result.get("severity"));
assertEquals(101L,
((Map<?, ?>) ((List<?>) result.get("checked_bindings")).get(0)).get("source_invocation_id"));
}
@Test
void validateRejectsNoEvidenceReferenceWhenExcerptClaimsPositiveEvidence() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
invocation(101L, "query_logs", "$.no_evidence",
"query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; total=0; message=未找到匹配的日志")
));
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
Map<String, Object> result = service.validate("session-1",
validOutput(101L, "query_logs", "$.no_evidence",
"HikariCP active=50/50 waiting=32"),
Map.of("status", "valid"));
assertEquals("fail", result.get("status"));
assertEquals("reject", result.get("severity"));
assertRuleAudit(result);
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.excerpt_mismatch"));
}
@Test
void validateRejectsPositiveBindingOnNegativeObservation() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.findBySessionIdOrderByIdAsc("session-1")).thenReturn(List.of(
invocation(101L, "query_logs", "$.logs[0]",
"order-service HikariCP active=50/50 waiting=32")
));
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
Map<String, Object> result = service.validate("session-1",
validOutput(101L, "query_logs", "$.logs[0]",
"order-service HikariCP active=50/50 waiting=32",
"negative_observation"),
Map.of("status", "valid"));
assertEquals("fail", result.get("status"));
assertEquals("reject", result.get("severity"));
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.raw_path"));
}
@Test @Test
void validateFailsWhenRemovedFieldsArePresent() { void validateFailsWhenRemovedFieldsArePresent() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class); ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
@@ -62,9 +183,22 @@ class ExecutorGatekeeperServiceTest {
assertEquals("fail", result.get("status")); assertEquals("fail", result.get("status"));
assertEquals("reject", result.get("severity")); assertEquals("reject", result.get("severity"));
assertRuleAudit(result);
assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.invocation_ref")); assertTrue(((List<?>) result.get("failed_rules")).contains("evidence.invocation_ref"));
} }
@Test
void failResultIncludesRuleAuditMetadata() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
ExecutorGatekeeperService service = new ExecutorGatekeeperService(repository);
Map<String, Object> result = service.fail("gatekeeper.internal_error", "gatekeeper", "boom");
assertEquals("fail", result.get("status"));
assertEquals("reject", result.get("severity"));
assertRuleAudit(result);
}
@Test @Test
void validateFailsForToolNameMismatch() { void validateFailsForToolNameMismatch() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class); ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
@@ -191,11 +325,21 @@ class ExecutorGatekeeperServiceTest {
} }
private Map<String, Object> validOutput(Long invocationId, String toolName, String rawPath, String excerpt) { private Map<String, Object> validOutput(Long invocationId, String toolName, String rawPath, String excerpt) {
return validOutput(invocationId, toolName, rawPath, excerpt, "symptom");
}
private Map<String, Object> validOutput(Long invocationId,
String toolName,
String rawPath,
String excerpt,
String claimType) {
Map<String, Object> binding = new java.util.LinkedHashMap<>(); Map<String, Object> binding = new java.util.LinkedHashMap<>();
binding.put("source_type", "tool_trace"); binding.put("source_type", "tool_trace");
binding.put("source_id", "trace-1"); binding.put("source_id", "trace-1");
binding.put("tool_name", toolName); binding.put("tool_name", toolName);
if (invocationId != null) {
binding.put("source_invocation_id", invocationId); binding.put("source_invocation_id", invocationId);
}
if (rawPath != null) { if (rawPath != null) {
binding.put("raw_path", rawPath); binding.put("raw_path", rawPath);
} }
@@ -204,7 +348,7 @@ class ExecutorGatekeeperServiceTest {
"answer_version", "executor_evidence_v2", "answer_version", "executor_evidence_v2",
"claims", List.of(Map.of( "claims", List.of(Map.of(
"claim_id", "claim-1", "claim_id", "claim-1",
"claim_type", "symptom", "claim_type", claimType,
"claim_text", "连接池 active 达到上限", "claim_text", "连接池 active 达到上限",
"support_level", "direct", "support_level", "direct",
"evidence_bindings", List.of(binding) "evidence_bindings", List.of(binding)
@@ -214,4 +358,13 @@ class ExecutorGatekeeperServiceTest {
"missing_info", List.of() "missing_info", List.of()
)); ));
} }
private void assertRuleAudit(Map<String, Object> result) {
assertEquals("gatekeeper-rules-v1", result.get("rule_set_version"));
assertTrue(result.get("rules") instanceof List<?>);
List<?> rules = (List<?>) result.get("rules");
assertFalse(rules.isEmpty());
assertTrue(rules.stream().anyMatch(rule ->
rule instanceof Map<?, ?> map && "evidence.raw_path".equals(map.get("id"))));
}
} }
@@ -27,7 +27,7 @@ class ToolInvocationRecorderTest {
private final ObjectMapper objectMapper = new ObjectMapper(); private final ObjectMapper objectMapper = new ObjectMapper();
@Test @Test
void recordEvidenceToolPreservesNoEvidenceSemantics() { void recordEvidenceToolPreservesNoEvidenceSemantics() throws Exception {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class); ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.save(any(ToolInvocation.class))).thenAnswer(invocation -> invocation.getArgument(0)); when(repository.save(any(ToolInvocation.class))).thenAnswer(invocation -> invocation.getArgument(0));
ToolInvocationRecorder recorder = new ToolInvocationRecorder(repository, new ObjectMapper()); ToolInvocationRecorder recorder = new ToolInvocationRecorder(repository, new ObjectMapper());
@@ -36,8 +36,8 @@ class ToolInvocationRecorderTest {
try { try {
recorder.recordEvidenceTool( recorder.recordEvidenceTool(
"query_logs", "query_logs",
Map.of("query", "timeout"), Map.of("query", "inventory-service HikariCP", "log_topic", "application-logs"),
"{\"success\":false,\"message\":\"未找到匹配的日志\"}", "{\"success\":false,\"query\":\"inventory-service HikariCP\",\"log_topic\":\"application-logs\",\"logs\":[],\"total\":0,\"message\":\"未找到匹配的日志\"}",
true, true,
System.currentTimeMillis() - 10, System.currentTimeMillis() - 10,
null, null,
@@ -57,6 +57,14 @@ class ToolInvocationRecorderTest {
assertEquals(Boolean.TRUE, saved.getSuccess()); assertEquals(Boolean.TRUE, saved.getSuccess());
assertTrue(saved.getRetrievalDetails().contains("\"evidence_status\":\"no_evidence\"")); assertTrue(saved.getRetrievalDetails().contains("\"evidence_status\":\"no_evidence\""));
assertTrue(saved.getRetrievalDetails().contains("\"retrieved_domains\":[\"application-logs\"]")); assertTrue(saved.getRetrievalDetails().contains("\"retrieved_domains\":[\"application-logs\"]"));
JsonNode details = objectMapper.readTree(saved.getRetrievalDetails());
assertEquals("$.no_evidence", details.path("evidence_refs").get(0).path("raw_path").asText());
String text = details.path("evidence_refs").get(0).path("text").asText();
assertTrue(text.contains("query_logs returned no evidence"));
assertTrue(text.contains("inventory-service HikariCP"));
assertTrue(text.contains("total=0"));
assertTrue(text.contains("evidence_status=no_evidence"));
} }
@Test @Test
@@ -125,6 +133,40 @@ class ToolInvocationRecorderTest {
assertTrue(details.path("evidence_refs").get(0).path("text").asText().contains("HighMemoryUsage")); assertTrue(details.path("evidence_refs").get(0).path("text").asText().contains("HighMemoryUsage"));
} }
@Test
void recordEvidenceToolExtractsMetricNoEvidenceRef() throws Exception {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.save(any(ToolInvocation.class))).thenAnswer(invocation -> invocation.getArgument(0));
ToolInvocationRecorder recorder = new ToolInvocationRecorder(repository, new ObjectMapper());
SessionContextHolder.setSessionId("metric-no-evidence-session");
try {
recorder.recordEvidenceTool(
"query_metrics",
Map.of("query", "active_prometheus_alerts"),
"""
{"success":true,"alerts":[],"message":"成功检索到 0 个活动告警"}
""",
true,
System.currentTimeMillis() - 10,
null,
"prometheus_alerts",
ToolInvocationRecorder.EVIDENCE_STATUS_NO_EVIDENCE,
Map.of("metric_family", "prometheus_alerts")
);
} finally {
SessionContextHolder.clear();
}
ArgumentCaptor<ToolInvocation> captor = ArgumentCaptor.forClass(ToolInvocation.class);
verify(repository).save(captor.capture());
JsonNode details = objectMapper.readTree(captor.getValue().getRetrievalDetails());
assertEquals("$.no_evidence", details.path("evidence_refs").get(0).path("raw_path").asText());
assertTrue(details.path("evidence_refs").get(0).path("text").asText().contains("query_metrics returned no evidence"));
assertTrue(details.path("evidence_refs").get(0).path("text").asText().contains("prometheus_alerts"));
}
@Test @Test
void recordLookupKnowledgePreservesRetrievalSpecificFields() { void recordLookupKnowledgePreservesRetrievalSpecificFields() {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class); ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
@@ -190,6 +232,38 @@ class ToolInvocationRecorderTest {
assertTrue(saved.getRetrievalDetails().contains("\"rerank_trace\"")); assertTrue(saved.getRetrievalDetails().contains("\"rerank_trace\""));
} }
@Test
void recordLookupKnowledgeAddsNoEvidenceRefWhenNoBlocks() throws Exception {
ToolInvocationRepository repository = mock(ToolInvocationRepository.class);
when(repository.save(any(ToolInvocation.class))).thenAnswer(invocation -> invocation.getArgument(0));
ToolInvocationRecorder recorder = new ToolInvocationRecorder(repository, new ObjectMapper());
SessionContextHolder.setSessionId("lookup-no-evidence-session");
ToolInvocationRecorder.LookupKnowledgeRecord record = ToolInvocationRecorder.LookupKnowledgeRecord.builder()
.query("inventory-service HikariCP")
.outputPreview("")
.outputLength(0)
.durationMs(12)
.success(true)
.evidenceStatus(ToolInvocationRecorder.EVIDENCE_STATUS_NO_EVIDENCE)
.evidenceBlocks(List.of())
.build();
try {
recorder.recordLookupKnowledge(record);
} finally {
SessionContextHolder.clear();
}
ArgumentCaptor<ToolInvocation> captor = ArgumentCaptor.forClass(ToolInvocation.class);
verify(repository).save(captor.capture());
JsonNode details = objectMapper.readTree(captor.getValue().getRetrievalDetails());
assertEquals("$.no_evidence", details.path("evidence_refs").get(0).path("raw_path").asText());
assertTrue(details.path("evidence_refs").get(0).path("text").asText().contains("lookup_knowledge returned no evidence"));
assertTrue(details.path("evidence_refs").get(0).path("text").asText().contains("inventory-service HikariCP"));
}
@Test @Test
void lookupKnowledgeRecordFromSummarizesEvidenceBlocks() { void lookupKnowledgeRecordFromSummarizesEvidenceBlocks() {
EvidenceBlock block = EvidenceBlock.builder() EvidenceBlock block = EvidenceBlock.builder()