feat(graph): complete stategraph cleanup and acceptance

This commit is contained in:
zhuyongxin
2026-07-20 10:23:27 +08:00
parent 208a231113
commit 190013c901
46 changed files with 1223 additions and 1211 deletions
+4 -4
View File
@@ -1,6 +1,6 @@
# MVP 架构文档
**更新日期**:2026-07-10
**更新日期**:2026-07-17
这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到:
@@ -14,9 +14,9 @@
|---|---|
| [current-mvp-architecture.md](current-mvp-architecture.md) | 当前可运行 MVP 的总体架构、链路、持久化和质量门禁 |
| [interview-one-pager.md](interview-one-pager.md) | 面试一页式架构讲解,包含总图、亮点、取舍和追问回答 |
| [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat SequentialAgent、AIOps SupervisorAgent、工具边界 |
| [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat bounded StateGraph、AIOps SupervisorAgent、工具边界 |
| [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md) | Chat 证据链路当前数据契约,覆盖 Executor V2、Gatekeeper、Verifier、Composer、`evidence_refs` |
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Gatekeeper、Verifier、Composer、评测基线组成的质量门禁 |
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、StateGraph、Trace、Gatekeeper、Verifier、Composer、评测基线组成的质量门禁 |
| [rag-architecture.md](rag-architecture.md) | RAG/知识检索新架构,覆盖 L0 hint、VectorStore 主路径、SDK fallback、证据追踪 |
| [modular-rag-pipeline.md](modular-rag-pipeline.md) | `lookup_knowledge` 模块化 RAG 落地架构,覆盖 pipeline、fallback、evidence-first contract、trace |
| [rag-eval-closure.md](rag-eval-closure.md) | RAG 评测闭环,覆盖 offline baseline、baseline diff、diagnosis eval 和 live acceptance |
@@ -29,7 +29,7 @@
## 当前架构一句话
SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,Chat 链路由 Gatekeeper 做引用真实性校验、Verifier 做可推导性判断、Composer 生成最终表达;多轮会话元数据落到 `chat_session`,每次诊断运行落到 `diagnosis_run`,步骤和工具明细通过 `agent_step.run_id`、`tool_invocation.run_id` 关联,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:复杂 Chat 由有界 StateGraph 显式编排 Planner、Executor、Gatekeeper、Verified Input、Verifier、Composer 与安全 Fallback,Executor 通过工具收集日志、指标和知识库证据;多轮会话元数据落到 `chat_session`,每次诊断运行落到 `diagnosis_run`,步骤和工具明细通过 `run_id` 关联,Graph 路由摘要独立保存为 `orchestration_trace`,最终通过精确 Run Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
## 阅读顺序
+30 -25
View File
@@ -1,6 +1,6 @@
# Agent 编排架构
**更新日期**:2026-07-08
**更新日期**:2026-07-17
**状态**:当前可运行架构
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
@@ -8,7 +8,7 @@
旧版 Agent 架构把系统描述为 Supervisor、Planner、SubAgent、Verifier 的团队协作。当前 MVP 保留这个核心思想,但实现更收敛:
- Chat 链路使用固定顺序工作流:`Planner -> Executor -> Gatekeeper -> Verifier -> Composer`。
- Chat 复杂诊断使用有递归上限的显式 StateGraph;正常路径是 `Planner -> Executor -> Gatekeeper -> Verified Input -> Verifier -> Composer`,条件边负责有限技术重试、一次补证据和安全 Fallback。
- AIOps 链路使用 `SupervisorAgent` 调度 `Planner + Executor`,最终由规则评估器做轻量验证。
- 当前没有拆分 ExternalApiSubAgent、InternalErrorSubAgent、DatabaseSubAgent;这些作为后续演进方向保留。
- 证据工具不直接散落在各个 Agent 里,而是通过 Spring AI ToolCallback / `@Tool` 统一暴露。
@@ -19,15 +19,19 @@
flowchart TB
subgraph Chat["Chat diagnosis"]
ChatIn["POST /api/chat"] --> ChatService["ChatService"]
ChatService --> ChatPlanner["chat_planner"]
ChatService --> ChatGraph["ChatDiagnosisGraphRuntime / StateGraph"]
ChatGraph --> ChatPlanner["Planner Node"]
ChatPlanner --> ChatExecutor["chat_executor"]
ChatExecutor --> ChatTools["evidence tools"]
ChatTools --> ChatExecutor
ChatExecutor --> ChatGatekeeper["ExecutorGatekeeperService"]
ChatGatekeeper --> ChatVerifier["chat_verifier"]
ChatExecutor --> ChatGatekeeper["Gatekeeper Node"]
ChatGatekeeper --> VerifiedInput["Verified Input Node"]
VerifiedInput --> ChatVerifier["Verifier Node"]
ChatVerifier --> ChatDecision{"PASS / LOW_CONFID / REJECT"}
ChatDecision --> ChatComposer["chat_composer"]
ChatDecision --> ChatComposer["Composer Node"]
ChatDecision --> ChatFallback["Fallback Node"]
ChatComposer --> ChatAnswer["final answer"]
ChatFallback --> ChatAnswer
end
subgraph AiOps["AIOps diagnosis"]
@@ -54,6 +58,7 @@ flowchart TB
ChatPlanner --> Step
ChatExecutor --> Step
ChatGatekeeper --> SelfEval
ChatGraph --> Run
ChatVerifier --> Step
ChatTools --> Invocation
ChatDecision --> SelfEval
@@ -69,19 +74,13 @@ flowchart TB
## 3. Chat 编排
Chat 复杂诊断采用 `SequentialAgent`,顺序固定:
Chat 复杂诊断采用 `ChatDiagnosisGraphRuntime` 编译的 bounded StateGraph。它有一条正常路径和显式条件边,不再依赖固定顺序 Agent 或 Verifier Hook:
```text
chat_planner
-> chat_executor
-> lookup_knowledge / query_logs / query_metrics / date_time
-> outputs executor_evidence_v2
-> VerifierInputHook / ExecutorGatekeeperService
-> validates source_invocation_id / raw_path / evidence_excerpt
-> chat_verifier
-> judges whether verified evidence can derive claims
-> chat_composer
-> writes final user-facing answer
START -> PLANNER -> EXECUTOR -> GATEKEEPER -> VERIFIED_INPUT -> VERIFIER -> COMPOSER -> END
| | | | |
+ retry + fallback + fallback + retry + retry/fallback
+ EVIDENCE_RETRY -> PLANNER (最多一次)
```
关键行为:
@@ -90,16 +89,18 @@ chat_planner
|---|---|---|
| `chat_planner` | 拆解问题,注入知识域地图和对话历史,给出排查方向 | `planner_plan` |
| `chat_executor` | 按计划调用证据工具,抽取带 `source_invocation_id + raw_path + evidence_excerpt` 的微观事实 | `executor_evidence_v2` |
| `ExecutorGatekeeperService` | 在 Verifier 前做代码级引用验真,拒绝伪造 ID、错配 raw_path、错配 excerpt | `gatekeeper_result` |
| `chat_verifier` | 只判断已验真 evidence excerpt 是否能推出 claim,不做新检索 | `verifier_output` |
| `GatekeeperNode` / `ExecutorGatekeeperService` | 按当前 `runId` 做代码级引用验真,拒绝伪造 ID、错配 raw_path、错配 excerpt | `gatekeeper_result` |
| `VerifiedInputNode` | 只投影 Gatekeeper 通过的 claims 与 matched evidence,隔离完整工具 Trace | `verified_executor_output`、`verified_evidence` |
| `chat_verifier` | 只判断已验真的 evidence excerpt 是否能推出 claim,不做新检索、不读取完整工具 Trace | `verifier_output` |
| `chat_composer` | 只表达 Verifier 允许输出的 claims、缺口和建议,生成最终用户答复 | `composer_output` |
| `FallbackNode` | 在不可恢复失败或路由上限触发时生成非空安全答复 | `final_answer`、degraded trace |
Chat 链路最多支持两轮验证:
Chat Graph 支持有限技术重试,并只允许一次 evidence retry;所有分支最终进入 Composer 或 Fallback:
```mermaid
sequenceDiagram
autonumber
participant C as ChatService
participant C as ChatService / StateGraph
participant P as chat_planner
participant E as chat_executor
participant T as tools
@@ -114,18 +115,22 @@ sequenceDiagram
E->>T: 调用证据工具
T-->>E: 证据结果
E-->>C: executor_evidence_v2
C->>G: executor_structured_output + tool_invocation.evidence_refs
C->>G: executor_output + run-owned tool_invocation.evidence_refs
G-->>C: gatekeeper_result
C->>V: executor_structured_output + gatekeeper_result + tool_trace_summary
C->>C: VerifiedInputNode projects passed claims/evidence
C->>V: verified_executor_output + verified_evidence + gatekeeper_audit
V-->>C: PASS / LOW_CONFID / REJECT
C->>R: 写入 verifier_evaluation
alt LOW_CONFID 且允许补证据
alt LOW_CONFID 且允许一次补证据
C->>P: retry_context: 仅补缺失证据
else PASS 或 REJECT
else PASS / LOW_CONFID 可输出
C->>M: allowed_claims + missing_info + recommended_actions
M-->>C: composer_output
C->>R: 保存 Composer 最终 answer
else 不可恢复失败
C->>R: Fallback 安全答复
end
C->>R: 保存 orchestration_trace(version/transitions/final_node/termination_reason/degraded/evidence_retry_count)
```
决策语义:
+22 -7
View File
@@ -1,6 +1,6 @@
# 当前 MVP 架构
**更新日期**:2026-07-08
**更新日期**:2026-07-17
**状态**:当前可运行架构
**适用范围**:Demo、面试讲解、后续迭代规划
@@ -36,11 +36,14 @@ flowchart TB
subgraph Agent["Agent Orchestration"]
Supervisor["Supervisor"]
StateGraph["Chat Diagnosis StateGraph"]
Planner["Planner"]
Executor["Executor"]
Gatekeeper["Gatekeeper"]
VerifiedInput["Verified Input"]
Verifier["Verifier"]
Composer["Composer"]
Fallback["Fallback"]
end
subgraph Tools["Evidence Tools"]
@@ -74,7 +77,14 @@ flowchart TB
end
API --> App
ChatService --> Agent
ChatService --> StateGraph
StateGraph --> Planner
StateGraph --> Executor
StateGraph --> Gatekeeper
StateGraph --> VerifiedInput
StateGraph --> Verifier
StateGraph --> Composer
StateGraph --> Fallback
AiOpsService --> Agent
SkillRegistry --> PlannerSkillHook
PlannerSkillHook --> Planner
@@ -374,11 +384,12 @@ Trace API 聚合:
- Agent step 序列。
- 工具调用和检索细节。
- Chat Gatekeeper / Verifier / Composer 结果。
- Chat `run.orchestrationTrace` 路由摘要,独立于 self-evaluation 和步骤/工具明细。
- AIOps rule evaluation 结果。
Trace 是本项目区别于普通问答系统的关键:答案不是孤立文本,而是可以追溯到 Agent 决策、工具调用和证据来源。
Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明见 [harness-quality-gates.md](harness-quality-gates.md),用户反馈与 `self_evaluation` 闭环见 [feedback-architecture.md](feedback-architecture.md)。
Prompt、StateGraph、Gatekeeper、Verifier、Composer 和评测门禁的完整说明见 [harness-quality-gates.md](harness-quality-gates.md),用户反馈与 `self_evaluation` 闭环见 [feedback-architecture.md](feedback-architecture.md)。
## 8. 质量门禁
@@ -386,9 +397,11 @@ Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明
| 门禁 | 位置 | 作用 |
|---|---|---|
| Executor Gatekeeper | `VerifierInputHook` / `ExecutorGatekeeperService` | 校验 Executor 引用的 invocation、`raw_path`、`evidence_excerpt` 是否真实 |
| Chat Verifier | `ChatService` | 判断已验真证据是否能推出 Executor claims |
| Chat Composer | `ChatService` | 只表达 Verifier 允许输出的内容,避免把 no-evidence 说成已排除 |
| Executor Gatekeeper | `GatekeeperNode` / `ExecutorGatekeeperService` | 校验当前 Run 的 invocation、`raw_path`、`evidence_excerpt` 是否真实 |
| Verified Input | `VerifiedInputNode` | 仅投影 Gatekeeper 通过的 claims/evidence,阻断完整工具 Trace 进入 Verifier |
| Chat Verifier | `VerifierNodeAdapter` | 判断已验真证据是否能推出 Executor claims |
| Chat Composer / Fallback | `ComposerNodeAdapter` / `FallbackNode` | 输出受控答复;异常分支也必须安全终止 |
| Graph routing | `DiagnosisGraphWorkflowTest` / `run.orchestrationTrace` | 验证条件边、有限重试、最终节点和终止原因 |
| AIOps Rule Evaluation | `AiOpsRuleEvaluationService` | 校验告警诊断是否聚焦 payload 并使用证据 |
| Diagnosis Eval Baseline | `mvp/eval/` | 固化诊断 trace 和报告行为 |
| RAG Retrieval Baseline | `eval/rag-retrieval/` | 固化检索召回行为,避免 RAG 重构回退 |
@@ -399,6 +412,7 @@ Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明
已经完成:
- Chat 和 AIOps 两条入口链路。
- Chat 复杂诊断已单轨切换到 bounded StateGraph,并持久化 Run-owned `orchestration_trace`。
- 显式 `lookup_knowledge` Agent Tool。
- L0 从最终决策降级为 domain/entity hint。
- `VectorSearchService` 作为稳定检索门面。
@@ -430,6 +444,7 @@ Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明
| 能力 | 代码 |
|---|---|
| Chat 入口与编排 | `ChatController`, `ChatService` |
| Chat StateGraph | `ChatDiagnosisGraphRuntime`, `DiagnosisGraphFactory`, `DiagnosisRealGraphActionsFactory` |
| AIOps 入口与编排 | `ChatController.aiOps`, `AiOpsService` |
| AIOps 规则验证 | `AiOpsRuleEvaluationService` |
| 知识库工具 | `LookupKnowledgeTool` |
@@ -440,5 +455,5 @@ Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明
| Spring AI VectorStore 配置辅助 | `SpringAiVectorStoreSidecarService` |
| Trace 聚合 | `DiagnosisTraceService` |
| 工具调用记录 | `ToolInvocationRecorder` |
| Executor 引用验真 | `ExecutorGatekeeperService`, `VerifierInputHook` |
| Executor 引用验真与投影 | `GatekeeperNode`, `ExecutorGatekeeperService`, `VerifiedInputNode` |
| self_evaluation 合并 | `SelfEvaluationMergeService` |
@@ -1,7 +1,7 @@
# Chat Evidence Pipeline Contracts
**状态**:当前实现
**更新日期**:2026-07-08
**更新日期**:2026-07-17
**范围**:Chat 复杂诊断链路中的 Planner、Executor、Gatekeeper、Verifier、Composer 数据契约
当前 Chat 复杂诊断链路是:
@@ -9,7 +9,8 @@
```text
chat_planner
-> chat_executor
-> VerifierInputHook / ExecutorGatekeeperService
-> GatekeeperNode / ExecutorGatekeeperService
-> VerifiedInputNode
-> chat_verifier
-> chat_composer
-> final answer
@@ -219,13 +220,13 @@ Executor 必须遵守:
## 4. Gatekeeper
Gatekeeper 位于 Verifier 前,由 `VerifierInputHook` 触发,负责代码级引用真实性校验。
Gatekeeper 是 StateGraph 中的显式 Node,调用 `ExecutorGatekeeperService` 对当前 Run 的证据引用做代码级真实性校验。
### 4.1 输入
- `sessionId`
- `sessionId + runId`(来自 `RunnableConfig`,工具查询以 `runId` 为边界)
- `executor_structured_output`
- 当前 session 的 `tool_invocation`
- 当前 run 的 `tool_invocation`
### 4.2 输出
@@ -304,22 +305,20 @@ Gatekeeper 位于 Verifier 前,由 `VerifierInputHook` 触发,负责代码
## 5. Verifier
Verifier 输入由 `VerifierInputHook` 构造:
`VerifiedInputNode` 只保留 Gatekeeper 检查通过的 claim/binding,并为 Verifier 构造最小输入:
```json
{
"original_query": "用户原始问题",
"executor_final_answer": "{...executor raw text for debug/fallback only...}",
"executor_structured_output": {
"diagnosis_context": {
"query": "用户原始问题"
},
"verified_executor_output": {
"answer_version": "executor_evidence_v2",
"claims": []
},
"executor_output_parse_status": {
"status": "valid",
"detail": "parsed executor evidence contract"
},
"tool_trace_summary": [],
"gatekeeper_result": {},
"verified_evidence": [],
"gatekeeper_audit": {},
"verdict_ceiling": "PASS",
"retry_context": null
}
```
@@ -330,7 +329,7 @@ Verifier 职责:
- 不读 skill。
- 不逐字核验 excerpt 真伪;这由 Gatekeeper 完成。
- 只判断 `claim_text` 是否能由已核验的 `evidence_excerpt` 推出。
- 结构化输出有效时,不得从 `executor_final_answer` 抽取额外确认事实。
- 不读取 Executor 原始答复或完整工具 Trace,只读取 verified projection。
- 对 `gatekeeper_result.severity=reject` 不得输出 `PASS`。
- 对 `gatekeeper_result.severity=low_confid` 不得输出 `PASS`。
@@ -410,7 +409,7 @@ Composer 位于 Verifier 之后,输入是 ChatService 过滤后的允许表达
"rule_set_version": "gatekeeper-rules-v1"
},
"composer_output": {},
"tool_trace_summary": []
"verified_evidence": []
}
}
```
@@ -422,6 +421,9 @@ Trace API 可用于回放:
- Gatekeeper 是否通过、是否自动回填。
- Verifier 如何判断可推导性。
- Composer 最终如何表达给用户。
- `run.orchestrationTrace` 如何经过条件边、有限重试并终止。
历史 Run/fixture 的 `verifier_evaluation.tool_trace_summary` 仍可被 Trace UI 或离线评测只读解析,但它是旧链路兼容字段,不是当前 Verifier 输入,也不再由生产链路生成。
---
+11 -12
View File
@@ -1,6 +1,6 @@
# 反馈与自评估架构
**更新日期**:2026-07-10
**更新日期**:2026-07-17
**状态**:当前可运行架构
**参考历史文档**:`archive/2026-07-05-legacy/confidence-feedback.md`
@@ -27,9 +27,8 @@ flowchart TD
Invocation["tool_invocation"] --> RuleEval["EvaluationService: rule_evaluation"]
Invocation --> EvidenceRefs["evidence_refs"]
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
Invocation --> TraceSummary["ToolTraceSummaryService"]
Gatekeeper --> Verifier["chat_verifier"]
TraceSummary --> Verifier
Gatekeeper --> Projection["VerifiedInputNode"]
Projection --> Verifier["chat_verifier"]
Verifier --> VerifierEval["verifier_evaluation"]
Verifier --> Composer["chat_composer"]
Composer --> VerifierEval
@@ -81,7 +80,7 @@ flowchart TD
"executor_structured_output": {},
"gatekeeper_result": {},
"composer_output": {},
"tool_trace_summary": []
"verified_evidence": []
},
"aiops_rule_evaluation": {
"verdict": "...",
@@ -136,14 +135,13 @@ Chat 自评估分三步:
```mermaid
flowchart LR
ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
Invocation["tool_invocation"] --> Summary["ToolTraceSummaryService"]
Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
Invocation["tool_invocation"] --> EvidenceRefs["retrieval_details.evidence_refs"]
EvidenceRefs --> Gatekeeper
Gatekeeper --> GateResult["gatekeeper_result"]
Summary --> Evidence["tool_trace_summary"]
GateResult --> Verifier["chat_verifier"]
ExecutorOutput --> Verifier
Evidence --> Verifier
GateResult --> Projection["VerifiedInputNode"]
ExecutorOutput --> Projection
Projection --> VerifiedOutput["verified_executor_output + verified_evidence"]
VerifiedOutput --> Verifier["chat_verifier"]
Verifier --> Output["verifier_output JSON"]
Output --> Composer["chat_composer"]
Composer --> ComposerOutput["composer_output"]
@@ -163,9 +161,9 @@ Verifier 输出:
| `facts_checked` | 逐条事实校验 |
| `rationale` | 判定原因 |
| `executor_structured_output` | Executor 输出的结构化 claims 与证据绑定 |
| `verified_evidence` | Gatekeeper 通过并投影给 Verifier 的最小 matched evidence |
| `gatekeeper_result` | 引用真实性校验结果 |
| `composer_output` | 最终表达的解析状态和摘要 |
| `tool_trace_summary` | 本次校验使用的工具调用导航索引 |
ChatService 根据 verdict 决定:
@@ -177,6 +175,7 @@ ChatService 根据 verdict 决定:
- `executor_final_answer` 只作为 debug/fallback 上下文;结构化输出有效时,Verifier 不得从中抽取额外确认事实。
- `$.no_evidence` 只能表达“当前查询未检索到匹配证据”,不能表达“已排除/确认没有”。
- `run.orchestrationTrace` 是独立的 StateGraph 路由摘要,不属于 `self_evaluation`;历史 `tool_trace_summary` 仅用于旧 Run/fixture 只读兼容,不是当前 Verifier 输入。
## 6. AIOps 规则自评估
+14 -11
View File
@@ -1,6 +1,6 @@
# Harness 与质量门禁架构
**更新日期**:2026-07-08
**更新日期**:2026-07-17
**状态**:当前可运行架构 + 后续门禁规划
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
@@ -20,6 +20,7 @@ Agent 系统的核心风险不是“没有答案”,而是:
Prompt contract
+ Tool boundary
+ Agent hooks
+ StateGraph routing contract
+ Trace persistence
+ Gatekeeper deterministic validation
+ Verifier / rule evaluation
@@ -31,7 +32,7 @@ Prompt contract
```mermaid
flowchart TB
Input["User / AIOps input"] --> Prompt["Prompt contract"]
Prompt --> Agent["Planner / Executor / Verifier / Composer"]
Prompt --> Agent["Diagnosis StateGraph Nodes"]
Agent --> Tools["Evidence tools"]
Tools --> Invocation["tool_invocation"]
Agent --> StepHook["AgentLoggingHook"]
@@ -41,15 +42,16 @@ flowchart TB
Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
Agent --> Gatekeeper
Invocation --> TraceSummary["ToolTraceSummaryService"]
Gatekeeper --> Verifier["chat_verifier"]
TraceSummary --> Verifier
Gatekeeper --> Projection["VerifiedInputNode"]
Projection --> Verifier["chat_verifier"]
Verifier --> SelfEval["self_evaluation.verifier_evaluation"]
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
AiOpsRule --> AiOpsEval["self_evaluation.aiops_rule_evaluation"]
Run --> TraceAPI["DiagnosisTraceService"]
Agent --> Routing["diagnosis_run.orchestration_trace"]
Routing --> TraceAPI
Step --> TraceAPI
Invocation --> TraceAPI
SelfEval --> TraceAPI
@@ -161,7 +163,7 @@ error_message
## 6. Gatekeeper 与 Verifier 门禁
Chat Verifier 前置一层 Gatekeeper。Gatekeeper 不调用 LLM,只用代码检查 Executor 输出的证据引用是否真实存在。
Chat StateGraph 在 Verifier 前显式执行 Gatekeeper 和 Verified Input。Gatekeeper 不调用 LLM,只用代码检查 Executor 输出的证据引用是否真实存在;Verified Input 只投影通过的 binding。
```mermaid
flowchart LR
@@ -169,11 +171,12 @@ flowchart LR
ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
EvidenceRefs --> Gatekeeper
Gatekeeper --> GateResult["gatekeeper_result"]
Invocation --> Summary["ToolTraceSummaryService"]
Summary --> EvidenceIndex["tool_trace_summary"]
GateResult --> Projection["VerifiedInputNode"]
Projection --> VerifiedClaims["verified_executor_output"]
Projection --> VerifiedEvidence["verified_evidence"]
GateResult --> Verifier["chat_verifier"]
ExecutorOutput --> Verifier
EvidenceIndex --> Verifier
VerifiedClaims --> Verifier
VerifiedEvidence --> Verifier
Verifier --> Verdict{"verdict"}
Verdict -->|PASS| Composer["chat_composer"]
Composer --> Pass["输出最终答复"]
@@ -215,7 +218,7 @@ Verifier 不再逐字核验 excerpt 真伪;这由 Gatekeeper 完成。Verifier
diagnosis_run.self_evaluation.verifier_evaluation
```
其中同时持久化 `executor_structured_output`、`gatekeeper_result`、`tool_trace_summary`、`prompt_audit` 和 `composer_output`,用于 Trace 回放。
其中持久化 verified `executor_structured_output`、`verified_evidence`、`gatekeeper_result`、`prompt_audit` 和 `composer_output`,用于 Trace 回放。Graph 路由另存 `diagnosis_run.orchestration_trace`;历史 `tool_trace_summary` 只作为旧 Run/fixture 的读取兼容字段,不属于当前 Verifier 输入。
## 7. AIOps 规则门禁
+5 -3
View File
@@ -142,7 +142,7 @@ post-retrieval 层再把检索候选归一为:
- 给 Agent 输出 completeness hint。
- 写入 `tool_invocation.relevance_level`。
- 给 Gatekeeper 提供 `evidence_refs` 引用验真源。
- 给 Verifier 构造 `tool_trace_summary` 审计导航。
- 由 Gatekeeper 核验后,经 `VerifiedInputNode` 给 Verifier 构造最小 `verified_evidence` 投影。
- 供 EvaluationService 计算 evidence score。
## 6. 文档切片和 metadata
@@ -180,11 +180,13 @@ flowchart LR
Recorder --> Invocation["tool_invocation"]
Invocation --> Trace["DiagnosisTraceService"]
Invocation --> Gatekeeper["ExecutorGatekeeperService"]
Invocation --> Summary["ToolTraceSummaryService"]
Summary --> Verifier["chat_verifier"]
Gatekeeper --> Projection["VerifiedInputNode"]
Projection --> Verifier["chat_verifier"]
Invocation --> Eval["EvaluationService / RAG eval"]
```
旧 Trace/fixture 中的 `tool_trace_summary` 只保留读取兼容;当前 StateGraph 不再生成它,也不会把完整工具调用摘要输入 Verifier。
`tool_invocation` 中与检索相关的字段:
```text
+12 -5
View File
@@ -1,6 +1,6 @@
# 会话与 Trace 生命周期
**更新日期**:2026-07-10
**更新日期**:2026-07-17
**状态**:当前可运行架构
**参考历史文档**:`archive/2026-07-05-legacy/session-management.md`
@@ -29,7 +29,7 @@ flowchart TD
Session --> Run["create diagnosis_run(runId)"]
Run --> Running["run.status = RUNNING"]
Running --> Agent["Agent workflow"]
Running --> Agent["Chat StateGraph / AIOps workflow"]
Agent --> Context["execution context(sessionId, runId)"]
Context --> StepHook["AgentLoggingHook"]
StepHook --> Step["agent_step(session_id, run_id)"]
@@ -37,7 +37,8 @@ flowchart TD
Tool --> Invocation["tool_invocation(session_id, run_id)"]
Invocation --> Gatekeeper["Gatekeeper evidence validation"]
Agent --> Final{"workflow result"}
Agent --> GraphTrace["Chat: save orchestration_trace"]
GraphTrace --> Final{"workflow result"}
Final -->|success| Success["run.status = SUCCESS, answer saved"]
Final -->|failed| Failed["run.status = FAILED"]
@@ -81,6 +82,7 @@ stateDiagram-v2
| `status` | `diagnosis_run` | 单次运行执行状态 |
| `answer` | `diagnosis_run` | 本次运行最终报告或答复 |
| `self_evaluation` | `diagnosis_run` | 本次运行系统自评估 JSON |
| `orchestration_trace` | `diagnosis_run` | Chat StateGraph 路由摘要;包含 version、transitions、final node、termination reason、degraded 和 evidence retry count |
| `feedback` | `diagnosis_run` | 本次运行用户反馈 |
`feedback` 不修改 `status`。一个执行成功但用户标记 `not_useful` 的 run,仍然应该是 `SUCCESS + feedback=not_useful`。
@@ -115,7 +117,7 @@ ToolInvocationRecorder
-> retrieval_details / evidence_refs
```
Verifier、Gatekeeper 和 EvaluationService 应按 `run_id` 读取工具调用,避免同一 session 的其他 run 参与评分或证据校验。
Gatekeeper 和 EvaluationService 按 `run_id` 读取工具调用,避免同一 session 的其他 run 参与评分或证据校验。Verifier 只读取 `VerifiedInputNode` 生成的 verified projection,不直接读取完整工具调用列表。
## 7. Trace API 聚合
@@ -128,6 +130,7 @@ GET /api/diagnosis/{sessionId}/trace?runId=run-...
```text
diagnosis_run by sessionId + runId
+ run.orchestrationTrace parsed from diagnosis_run.orchestration_trace
+ chat_session metadata when available
+ agent_step where run_id = runId, ordered by the Trace API
+ tool_invocation where run_id = runId order by id
@@ -136,16 +139,20 @@ diagnosis_run by sessionId + runId
当 `runId` 缺失时,Trace API 为兼容旧客户端解析最新 run,并在响应中返回 resolved `runId`。当 `runId` 属于其他 `sessionId` 时,API 必须拒绝,不能泄漏其他会话的 Trace。
`run.orchestrationTrace` 只属于精确 Run 投影,不复制到顶层或 `session`。它解释 Graph 路由;`selfEvaluation` 解释证据/答案质量;`steps` 和 `toolInvocations` 保存详细执行证据,三者职责互不替代。历史 Run 的该字段可以为空。
## 8. Chat 与 AIOps 差异
| 维度 | Chat | AIOps |
|---|---|---|
| `agent_flow` | `CHAT` | `AI_OPS` |
| 编排方式 | `SequentialAgent`: Planner -> Executor -> Gatekeeper -> Verifier -> Composer | `SupervisorAgent`: Planner + Executor |
| 编排方式 | bounded `StateGraph`: Planner / Executor / Gatekeeper / Verified Input / Verifier / Composer / Fallback | `SupervisorAgent`: Planner + Executor |
| 自评估 | `rule_evaluation` + `verifier_evaluation` | `aiops_rule_evaluation` |
| 答案字段 | Chat 最终答复 | 告警分析报告 |
| runId 暴露 | `/api/chat` JSON response | `/api/ai_ops` SSE metadata message |
Chat StateGraph 的权威自动化验收分三层:`DiagnosisGraphWorkflowTest` 验证路由,`DiagnosisGraphNodeContractTest` 验证真实 Node 输入输出,`ChatServiceGraphIntegrationTest` 验证 Run 生命周期、Trace 持久化和对外集成。
## 9. 清理与边界
- Redis 会话历史用于多轮上下文,不是长期审计记录。
+9 -2
View File
@@ -8,7 +8,7 @@
- `interview-walkthrough.md`:面试讲解话术。
- `evidence-pipeline-scenarios.md`:PASS / LOW_CONFID / REJECT / no-evidence 场景矩阵。
- `trace-inspection-checklist.md`:Trace 字段检查清单。
- `scripts/run-interview-demo-check.ps1`:面试预检脚本,包含服务可达性、Chat、Trace、反馈和 summary 输出。
- `scripts/run-interview-demo-check.ps1`:面试预检脚本,绑定 exact runId,强制校验 Run orchestration trace,并输出 Chat、Trace、反馈和 summary。
- `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。
- `interview-q-and-a.md`:面试追问回答,覆盖 Agent 工程取舍、审计和评测。
- `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。
@@ -51,6 +51,8 @@ mvp/demo/output/feedback-response.json
mvp/demo/output/interview-demo-summary.json
```
自动化验收应传入唯一 `-SessionId`,并用 `-OutputDir target/...` 避免覆盖仓库样例。脚本从 Chat 响应取得 exact `runId`,缺少 `data.run.orchestrationTrace` 或 version/final node/termination reason/transitions/degraded/evidence retry count 时会立即失败。summary 额外包含 `orchestrationVersion`、`finalNode`、`terminationReason`、`degraded`、`transitionCount` 和 `evidenceRetryCount`。
手动请求:
```powershell
@@ -100,6 +102,10 @@ Invoke-RestMethod `
- `data.runId` 等于 `$runId`
- `data.session.sessionId` 等于 Chat session id
- `data.run.runId` 等于 `$runId`
- `data.run.orchestrationTrace.version` 非空
- `data.run.orchestrationTrace.final_node` 和 `termination_reason` 非空
- `data.run.orchestrationTrace.transitions` 是本次 Graph 的条件边记录
- `data.run.orchestrationTrace.degraded` 和 `evidence_retry_count` 记录安全降级与补证据次数
- `data.steps` 包含 planner / executor / verifier 等步骤
- `data.toolInvocations` 包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具
- `data.session.selfEvaluation` 包含 verifier 或 rule evaluation
@@ -174,9 +180,10 @@ Chat 主线:
```text
一个 session id + 一个 run id
-> 用户问题
-> 多 Agent 执行
-> bounded StateGraph(Planner / Executor / Gatekeeper / Verified Input / Verifier / Composer / Fallback)
-> 证据工具
-> Verifier / self_evaluation
-> run.orchestrationTrace 路由摘要
-> 最终答案
-> 用户反馈
-> Trace API 回放
+66 -1
View File
@@ -79,6 +79,15 @@ $chatPath = Join-Path $OutputDir "chat-response.json"
$chat | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $chatPath
$runId = $chat.data.runId
if ($chat.data.success -ne $true) {
throw "Chat response was not successful."
}
if ([string]::IsNullOrWhiteSpace([string]$chat.data.answer)) {
throw "Chat response did not include a non-empty answer."
}
if ($chat.data.sessionId -ne $SessionId) {
throw "Chat response sessionId '$($chat.data.sessionId)' did not match requested sessionId '$SessionId'."
}
if (-not $runId) {
throw "Chat response did not include runId; exact trace verification cannot continue."
}
@@ -92,6 +101,51 @@ $trace = Invoke-RestMethod @traceRequest
$tracePath = Join-Path $OutputDir "trace-response.json"
$trace | ConvertTo-Json -Depth 80 | Set-Content -Encoding UTF8 -Path $tracePath
$traceData = Get-TraceData -TraceResponse $trace
if ($null -eq $traceData -or $null -eq $traceData.run) {
throw "Exact Trace response did not include data.run."
}
if ($traceData.runId -ne $runId -or $traceData.run.runId -ne $runId) {
throw "Exact Trace runId did not match Chat runId '$runId'."
}
if ($traceData.run.sessionId -ne $SessionId) {
throw "Exact Trace run did not belong to requested sessionId '$SessionId'."
}
$orchestrationTrace = $traceData.run.orchestrationTrace
if ($null -eq $orchestrationTrace) {
throw "Exact Trace data.run.orchestrationTrace is missing."
}
foreach ($field in @("version", "final_node", "termination_reason")) {
if (-not ($orchestrationTrace.PSObject.Properties.Name -contains $field) -or
[string]::IsNullOrWhiteSpace([string]$orchestrationTrace.$field)) {
throw "Exact Trace data.run.orchestrationTrace.$field is missing."
}
}
foreach ($field in @("transitions", "degraded", "evidence_retry_count")) {
if (-not ($orchestrationTrace.PSObject.Properties.Name -contains $field)) {
throw "Exact Trace data.run.orchestrationTrace.$field is missing."
}
}
if ($null -eq $orchestrationTrace.transitions) {
throw "Exact Trace data.run.orchestrationTrace.transitions must be an array."
}
if ([int]$orchestrationTrace.evidence_retry_count -lt 0) {
throw "Exact Trace data.run.orchestrationTrace.evidence_retry_count must not be negative."
}
if ($traceData.run.status -ne "SUCCESS" -or $traceData.run.agentFlow -ne "CHAT") {
throw "Exact Trace run must be CHAT/SUCCESS."
}
if ([string]::IsNullOrWhiteSpace([string]$traceData.run.answer)) {
throw "Exact Trace run did not include a non-empty answer."
}
if (@($traceData.steps).Count -eq 0 -or @($traceData.toolInvocations).Count -eq 0) {
throw "Exact Trace did not include both Agent steps and tool invocation evidence."
}
if ($null -eq $traceData.run.selfEvaluation) {
throw "Exact Trace run did not include selfEvaluation."
}
$feedbackBody = @{
sessionId = $SessionId
runId = $runId
@@ -105,11 +159,13 @@ $feedbackRequest = @{
Body = $feedbackBody
}
$feedback = Invoke-RestMethod @feedbackRequest
if ($feedback.success -ne $true) {
throw "Feedback request was not successful for runId '$runId'."
}
$feedbackPath = Join-Path $OutputDir "feedback-response.json"
$feedback | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $feedbackPath
$traceData = Get-TraceData -TraceResponse $trace
$selfEvaluation = Get-SelfEvaluation -TraceData $traceData
$verifierEvaluation = $null
if ($null -ne $selfEvaluation) {
@@ -138,6 +194,7 @@ if ($null -ne $promptAudit) {
$promptAuditVersion = $promptAudit.version
}
$toolNames = Get-ToolNames -TraceData $traceData
$transitionCount = @($orchestrationTrace.transitions).Count
$summaryPath = Join-Path $OutputDir "interview-demo-summary.json"
$summary = [ordered]@{
@@ -149,6 +206,12 @@ $summary = [ordered]@{
gatekeeperStatus = $gatekeeperStatus
gatekeeperRuleSetVersion = $gatekeeperRuleSetVersion
promptAuditVersion = $promptAuditVersion
orchestrationVersion = $orchestrationTrace.version
finalNode = $orchestrationTrace.final_node
terminationReason = $orchestrationTrace.termination_reason
degraded = [bool]$orchestrationTrace.degraded
transitionCount = $transitionCount
evidenceRetryCount = [int]$orchestrationTrace.evidence_retry_count
toolNames = $toolNames
paths = [ordered]@{
chat = $chatPath
@@ -165,4 +228,6 @@ Write-Host "Interview demo preflight completed."
Write-Host "Verdict: $($summary.verdict)"
Write-Host "Gatekeeper rules: $($summary.gatekeeperRuleSetVersion)"
Write-Host "Prompt audit: $($summary.promptAuditVersion)"
Write-Host "Graph final node: $($summary.finalNode)"
Write-Host "Graph termination: $($summary.terminationReason)"
Write-Host "Summary: $summaryPath"
+27 -12
View File
@@ -7,16 +7,30 @@
| JSON path | 检查点 | 面试讲点 |
|---|---|---|
| `data.runId` / `data.run.runId` | 是否等于 demo 响应中的 `runId` | `runId` 精确绑定这一次诊断运行 |
| `data.session.sessionId` | 是否等于 `mvp-demo-payment-timeout-001` | `sessionId` 保留多轮上下文,Trace 精确回放依赖 `runId` |
| `data.session.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
| `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
| `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
| `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | 如果是 Chat V2 链路,是否记录 Gatekeeper 规则版本 | 安全规则可审计、可回归 |
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` | 如果是 Chat V2 链路,是否记录 Prompt 审计版本 | Prompt 变更可解释、可回归 |
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.prompts[*].version` | 是否记录 planner / executor / verifier / composer 版本 | 便于定位 Prompt 变更影响 |
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在当前 diagnosis run 上 |
| `data.run.sessionId` | 是否等于本次 Chat 请求的唯一 sessionId | `sessionId` 保留多轮上下文,Trace 精确回放依赖 `runId` |
| `data.run.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
| `data.run.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
| `data.run.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
| `data.run.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | 是否记录 Gatekeeper 规则版本 | 安全规则可审计、可回归 |
| `data.run.selfEvaluation.verifier_evaluation.prompt_audit.version` | 是否记录 Prompt 审计版本 | Prompt 变更可解释、可回归 |
| `data.run.selfEvaluation.verifier_evaluation.prompt_audit.prompts[*].version` | 是否记录 planner / executor / verifier / composer 版本 | 便于定位 Prompt 变更影响 |
| `data.run.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在当前 diagnosis run 上 |
## 2. Agent 步骤
## 2. StateGraph 路由
| JSON path | 检查点 | 面试讲点 |
|---|---|---|
| `data.run.orchestrationTrace.version` | 是否存在当前 trace contract 版本 | 路由摘要可演进、可兼容 |
| `data.run.orchestrationTrace.transitions[*]` | 是否记录实际经过的 Node 和 route | Graph 条件边不是从日志推断 |
| `data.run.orchestrationTrace.final_node` | 最终是 Composer 还是 Fallback | 正常输出与安全降级明确区分 |
| `data.run.orchestrationTrace.termination_reason` | 是否给出终止原因 | 每次 Run 都有可解释终点 |
| `data.run.orchestrationTrace.degraded` | 是否发生安全降级 | fallback 是可审计行为 |
| `data.run.orchestrationTrace.evidence_retry_count` | 是否为 0 或 1 | 补证据循环有硬上限 |
| `interview-demo-summary.json.finalNode` 等摘要字段 | 是否与 exact Trace 一致 | summary 只消费 Run 路由真理源 |
`orchestrationTrace` 负责路由;`selfEvaluation` 负责证据和答案质量;AgentStep/ToolInvocation 负责详细执行与工具证据。三者不能互相替代。
## 3. Agent 步骤
| JSON path | 检查点 | 面试讲点 |
|---|---|---|
@@ -25,7 +39,7 @@
| `data.steps[*].durationMs` | 是否有步骤耗时 | Trace 可用于耗时分析 |
| `data.steps[*].tokenCount` | 如可用,是否记录 token | Trace 可用于模型成本分析 |
## 3. 工具证据
## 4. 工具证据
| JSON path | 检查点 | 面试讲点 |
|---|---|---|
@@ -37,7 +51,7 @@
| `data.toolInvocations[*].retrievalDetails.evidence_refs` | 是否包含 `raw_path + text` | Gatekeeper 可以用代码核对 Executor 引用 |
| `data.toolInvocations[*].relevanceLevel` | 是否有相关性等级 | 可解释检索结果强弱 |
## 4. Summary
## 5. Summary
| JSON path | 检查点 | 面试讲点 |
|---|---|---|
@@ -46,7 +60,7 @@
| `data.summary.hasVerifierEvaluation` | 是否存在 Verifier 结果 | 最终答案经过质量门 |
| `data.summary.hasFeedback` | 提交反馈后是否为 true | 人类反馈闭环完成 |
## 5. 好的结果长什么样
## 6. 好的结果长什么样
```text
同一个 session id + run id
@@ -54,5 +68,6 @@
-> 持久化 agent steps
-> 持久化 evidence tool calls
-> verifier / self-evaluation
-> run.orchestrationTrace routing summary
-> feedback attached to the same run
```
+7 -5
View File
@@ -4,13 +4,15 @@ This folder contains the fixed offline regression set for the MVP diagnosis Agen
## Background
The current diagnosis chain is:
The current complex Chat diagnosis chain is a bounded StateGraph:
```text
Planner -> Executor -> Gatekeeper -> Verifier -> Composer -> final answer
Planner -> Executor -> Gatekeeper -> Verified Input -> Verifier -> Composer -> final answer
| |
+ bounded evidence retry + safe Fallback
```
Stages 1-4 introduced Executor V2 structured output, deterministic Gatekeeper audit, Verifier `claim_checks`, and Composer final-answer rendering. Stage 5 makes those audit fields part of the offline regression harness so future prompt, tool, or chain changes can be checked without relying on a one-off demo.
Executor V2 structured output, deterministic Gatekeeper audit, verified-only Verifier input, `claim_checks`, Composer rendering, and StateGraph routing are covered by deterministic tests so future prompt, tool, or graph changes can be checked without relying on a one-off demo.
## Scope
@@ -55,10 +57,10 @@ Run the focused evaluator test:
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test
```
Run the broader phase-5 regression set:
Run the authoritative Graph layers plus the fixed evaluator checks:
```powershell
mvn "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
mvn -q "-Dtest=DiagnosisGraphWorkflowTest,DiagnosisGraphNodeContractTest,ChatServiceGraphIntegrationTest,DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest" test
```
When fixtures or evaluator rules change, regenerate both baseline reports from the same case file and fixture directory, then update JSON and Markdown together.
+2 -2
View File
@@ -1,6 +1,6 @@
# MVP Issues 索引
**更新日期**:2026-07-16
**更新日期**:2026-07-20
**状态**:按活跃问题、设计笔记、RAG 问题集和已归档问题整理
## 目录约定
@@ -16,7 +16,6 @@
| 名称 | 标题 | 严重程度 | 状态 | 文件 |
|---|---|---|---|---|
| ISS-011 | Chat 诊断 StateGraph 编排改造 | 高 | 待实现 | [active/ISS-011-chat-diagnosis-stategraph-orchestration.md](active/ISS-011-chat-diagnosis-stategraph-orchestration.md) |
| ISS-003 | MVP 设计与实现 Review 收敛 | 高 | 待规划 | [active/ISS-003-mvp-design-implementation-review.md](active/ISS-003-mvp-design-implementation-review.md) |
| ISS-004 | Executor 域级检索水位控制 | 低 | 待规划 | [active/ISS-004-executor-domain-hard-limit.md](active/ISS-004-executor-domain-hard-limit.md) |
| executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [active/executor-evidence-attribution-hallucination.md](active/executor-evidence-attribution-hallucination.md) |
@@ -53,6 +52,7 @@
| 名称 | 标题 | 状态 | 文件 |
|---|---|---|---|
| ISS-011 | Chat 诊断 StateGraph 编排改造 | 已归档 | [archived/ISS-011-chat-diagnosis-stategraph-orchestration.md](archived/ISS-011-chat-diagnosis-stategraph-orchestration.md) |
| ISS-001 | Executor 重复召回同一文档 | 已修复 | [archived/ISS-001-duplicate-retrieval.md](archived/ISS-001-duplicate-retrieval.md) |
| ISS-002 | Executor 无约束重复调用 lookup_knowledge | 已修复 | [archived/ISS-002-executor-unconstrained-lookup.md](archived/ISS-002-executor-unconstrained-lookup.md) |
| ISS-005 | 证据链补齐与降级契约收敛 | 已归档 | [archived/ISS-005-evidence-trace-hardening.md](archived/ISS-005-evidence-trace-hardening.md) |
@@ -1,8 +1,9 @@
# ISS-011 Chat 诊断 StateGraph 编排改造
**状态**:待实现
**状态**:已归档
**严重程度**:高
**发现时间**:2026-07-16
**完成时间**:2026-07-20
**来源**:OnCall / Agent 编排模拟面试、当前 Chat 复杂诊断调用链复核
**预计实施周期**:2–3 个工作日
@@ -871,69 +872,69 @@ Eval baseline
### 编排
- [ ] 每次进入 Planner 阶段时,INVALID_OUTPUT / RETRYABLE_FAILED 最多触发一次技术重试。
- [ ] Planner NON_RETRYABLE_FAILED 或当前阶段第二次技术失败直接进入 Fallback。
- [ ] Planner 技术重试不增加 `evidence_retry_count`,补证据重新进入 Planner 时重置当前阶段的 `planner_retry_count`。
- [ ] Executor FAILED / TOOL_BLOCKED 后不会执行 Gatekeeper 和 Verifier。
- [ ] Executor INVALID_OUTPUT 不重试,不执行 Gatekeeper、Verifier 和模型 Composer。
- [ ] TOOL_BLOCKED 只用于工具层明确阻断且不存在合法 Executor 输出的场景。
- [ ] 工具空结果或工具失败后仍形成合法 Executor 输出时状态为 COMPLETED,并继续 Gatekeeper。
- [ ] Executor 合法 no-evidence 会继续执行 Gatekeeper 和 Verifier。
- [ ] Gatekeeper REJECT 直接进入 Fallback,不执行 Verifier。
- [ ] Gatekeeper LOW_CONFID 且零条已验真 binding 时直接进入 Fallback。
- [ ] Gatekeeper LOW_CONFID 且存在已验真 binding 时,Verifier 只接收通过校验的 binding。
- [ ] Gatekeeper PASS 和可继续的 LOW_CONFID 都经过 Verifier Input Builder。
- [ ] Verifier 只接收通过 binding 对应的 `verified_evidence`,不接收完整 `tool_trace_summary`。
- [ ] 未被 Executor 引用或未通过 Gatekeeper 的工具结果不能进入 Verifier 输入。
- [ ] Gatekeeper LOW_CONFID 路径的 `effective_verdict` 不得升级为 PASS。
- [ ] Gatekeeper 原始 pass/fail + severity 正确标准化为 PASS / LOW_CONFID / REJECT,未知状态安全映射为 REJECT。
- [ ] Verifier 执行状态与诊断 verdict 分离,任何失败状态不得出现在 model/effective verdict 中。
- [ ] Composer 和 Graph 条件边只读取 `effective_verdict`。
- [ ] Verifier INVALID_OUTPUT / RETRYABLE_FAILED 使用相同 verified input 最多技术重试一次,且不重新执行 Gatekeeper、Executor 或工具。
- [ ] Composer INVALID_OUTPUT / RETRYABLE_FAILED 使用相同安全输入最多技术重试一次,且不重新执行 Verifier 或前序节点。
- [ ] Verifier 第二次技术失败或 NON_RETRYABLE_FAILED 的 Fallback 不输出 Executor claim。
- [ ] Composer 第二次技术失败或 NON_RETRYABLE_FAILED 使用确定性安全模板。
- [ ] `verifier_retry_count`、`composer_retry_count` 和 `evidence_retry_count` 互相独立。
- [ ] Verifier LOW_CONFID 最多触发一次 Planner 补证据。
- [ ] LOW_CONFID 补证据循环受一次补查上限和 Graph recursion limit 限制。
- [ ] Gatekeeper verdict ceiling 导致的 LOW_CONFID 不触发补证据。
- [ ] 无法从 `facts_checked` 提取有效 `evidence_gaps` 时不触发补证据。
- [ ] 第二轮 Planner 只输出增量计划,不扩大诊断范围或重复成功查询。
- [ ] 第二轮 Executor 只执行增量查询,但输出完整 `executor_evidence_v2` 快照,而不是仅输出新增片段。
- [ ] 第二轮完整快照包含需要保留的第一轮可信 claims,并由 Gatekeeper 对全部 binding 重新验真。
- [ ] Java 编排层不对两轮 claim 文本进行语义合并。
- [ ] Composer 技术重试耗尽或不可重试失败时使用固定模板结束。
- [x] 每次进入 Planner 阶段时,INVALID_OUTPUT / RETRYABLE_FAILED 最多触发一次技术重试。
- [x] Planner NON_RETRYABLE_FAILED 或当前阶段第二次技术失败直接进入 Fallback。
- [x] Planner 技术重试不增加 `evidence_retry_count`,补证据重新进入 Planner 时重置当前阶段的 `planner_retry_count`。
- [x] Executor FAILED / TOOL_BLOCKED 后不会执行 Gatekeeper 和 Verifier。
- [x] Executor INVALID_OUTPUT 不重试,不执行 Gatekeeper、Verifier 和模型 Composer。
- [x] TOOL_BLOCKED 只用于工具层明确阻断且不存在合法 Executor 输出的场景。
- [x] 工具空结果或工具失败后仍形成合法 Executor 输出时状态为 COMPLETED,并继续 Gatekeeper。
- [x] Executor 合法 no-evidence 会继续执行 Gatekeeper 和 Verifier。
- [x] Gatekeeper REJECT 直接进入 Fallback,不执行 Verifier。
- [x] Gatekeeper LOW_CONFID 且零条已验真 binding 时直接进入 Fallback。
- [x] Gatekeeper LOW_CONFID 且存在已验真 binding 时,Verifier 只接收通过校验的 binding。
- [x] Gatekeeper PASS 和可继续的 LOW_CONFID 都经过 Verifier Input Builder。
- [x] Verifier 只接收通过 binding 对应的 `verified_evidence`,不接收完整 `tool_trace_summary`。
- [x] 未被 Executor 引用或未通过 Gatekeeper 的工具结果不能进入 Verifier 输入。
- [x] Gatekeeper LOW_CONFID 路径的 `effective_verdict` 不得升级为 PASS。
- [x] Gatekeeper 原始 pass/fail + severity 正确标准化为 PASS / LOW_CONFID / REJECT,未知状态安全映射为 REJECT。
- [x] Verifier 执行状态与诊断 verdict 分离,任何失败状态不得出现在 model/effective verdict 中。
- [x] Composer 和 Graph 条件边只读取 `effective_verdict`。
- [x] Verifier INVALID_OUTPUT / RETRYABLE_FAILED 使用相同 verified input 最多技术重试一次,且不重新执行 Gatekeeper、Executor 或工具。
- [x] Composer INVALID_OUTPUT / RETRYABLE_FAILED 使用相同安全输入最多技术重试一次,且不重新执行 Verifier 或前序节点。
- [x] Verifier 第二次技术失败或 NON_RETRYABLE_FAILED 的 Fallback 不输出 Executor claim。
- [x] Composer 第二次技术失败或 NON_RETRYABLE_FAILED 使用确定性安全模板。
- [x] `verifier_retry_count`、`composer_retry_count` 和 `evidence_retry_count` 互相独立。
- [x] Verifier LOW_CONFID 最多触发一次 Planner 补证据。
- [x] LOW_CONFID 补证据循环受一次补查上限和 Graph recursion limit 限制。
- [x] Gatekeeper verdict ceiling 导致的 LOW_CONFID 不触发补证据。
- [x] 无法从 `facts_checked` 提取有效 `evidence_gaps` 时不触发补证据。
- [x] 第二轮 Planner 只输出增量计划,不扩大诊断范围或重复成功查询。
- [x] 第二轮 Executor 只执行增量查询,但输出完整 `executor_evidence_v2` 快照,而不是仅输出新增片段。
- [x] 第二轮完整快照包含需要保留的第一轮可信 claims,并由 Gatekeeper 对全部 binding 重新验真。
- [x] Java 编排层不对两轮 claim 文本进行语义合并。
- [x] Composer 技术重试耗尽或不可重试失败时使用固定模板结束。
### 证据和安全
- [ ] Gatekeeper 规则语义不放宽。
- [ ] Verifier 只消费已验真证据。
- [ ] no-evidence 不得表达为已排除或问题不存在。
- [ ] REJECT 降级不泄漏 Executor 原始答案和未验证根因。
- [ ] Executor INVALID_OUTPUT、Gatekeeper REJECT 和零条可信 binding 的固定 Fallback 不输出任何 Executor claim。
- [ ] 前置验证失败 Fallback 只展示校验状态、工具执行概况、诊断限制和人工复核建议。
- [x] Gatekeeper 规则语义不放宽。
- [x] Verifier 只消费已验真证据。
- [x] no-evidence 不得表达为已排除或问题不存在。
- [x] REJECT 降级不泄漏 Executor 原始答案和未验证根因。
- [x] Executor INVALID_OUTPUT、Gatekeeper REJECT 和零条可信 binding 的固定 Fallback 不输出任何 Executor claim。
- [x] 前置验证失败 Fallback 只展示校验状态、工具执行概况、诊断限制和人工复核建议。
### 数据与审计
- [ ] Graph 使用 runId 作为 threadId。
- [ ] Agent step、tool invocation 和 self_evaluation 仍绑定正确 runId。
- [ ] `orchestration_trace` 只写入当前 diagnosis run,不污染其他 run 或 session 级数据。
- [ ] `orchestration_trace` 不包含 Prompt、模型思考、工具原文和 Graph State 快照。
- [ ] `orchestration_trace.transitions` 由有界 `orchestration_events` 生成,与实际节点执行顺序一致。
- [ ] 可处理异常发生时,已经产生的 orchestration events 能够 best-effort 写入当前 run。
- [ ] Trace 能展示实际节点路径、重试原因和终止原因。
- [ ] 每个新 StateGraph Chat run 的 `run.orchestrationTrace` 非空,且顶层和兼容 `session` 投影不重复该字段。
- [ ] Run 最终状态、答案、耗时、Token 和工具调用数正确回填。
- [ ] 所有成功生成安全响应的终止路径将 Run 标记为 SUCCESS,并通过 verdict 或 `orchestrationTrace.degraded` 表达质量。
- [ ] 只有未处理异常、持久化失败或无法生成安全响应时将 Run 标记为 FAILED。
- [x] Graph 使用 runId 作为 threadId。
- [x] Agent step、tool invocation 和 self_evaluation 仍绑定正确 runId。
- [x] `orchestration_trace` 只写入当前 diagnosis run,不污染其他 run 或 session 级数据。
- [x] `orchestration_trace` 不包含 Prompt、模型思考、工具原文和 Graph State 快照。
- [x] `orchestration_trace.transitions` 由有界 `orchestration_events` 生成,与实际节点执行顺序一致。
- [x] 可处理异常发生时,已经产生的 orchestration events 能够 best-effort 写入当前 run。
- [x] Trace 能展示实际节点路径、重试原因和终止原因。
- [x] 每个新 StateGraph Chat run 的 `run.orchestrationTrace` 非空,且顶层和兼容 `session` 投影不重复该字段。
- [x] Run 最终状态、答案、耗时、Token 和工具调用数正确回填。
- [x] 所有成功生成安全响应的终止路径将 Run 标记为 SUCCESS,并通过 verdict 或 `orchestrationTrace.degraded` 表达质量。
- [x] 只有未处理异常、持久化失败或无法生成安全响应时将 Run 标记为 FAILED。
### 工程质量
- [ ] 新 Graph 测试覆盖所有分支。
- [ ] `ChatServiceSequentialAgentTest` 已由新测试替换。
- [ ] 不保留长期重复的 Sequential 和 Graph 两套实现。
- [ ] 数据库 schema 仅新增 `diagnosis_run.orchestration_trace` nullable JSON 字段。
- [ ] `/api/chat` 和证据协议不变;Trace API 仅在 `run` 对象新增必有的 `orchestrationTrace` 字段。
- [x] 新 Graph 测试覆盖所有分支。
- [x] `ChatServiceSequentialAgentTest` 已由新测试替换。
- [x] 不保留长期重复的 Sequential 和 Graph 两套实现。
- [x] 数据库 schema 仅新增 `diagnosis_run.orchestration_trace` nullable JSON 字段。
- [x] `/api/chat` 和证据协议不变;Trace API 仅在 `run` 对象新增必有的 `orchestrationTrace` 字段。
---