docs(mvp): plan chat diagnosis stategraph refactor
This commit is contained in:
+192
@@ -0,0 +1,192 @@
|
||||
<!doctype html>
|
||||
<html lang="zh-CN">
|
||||
<head>
|
||||
<meta charset="utf-8">
|
||||
<meta name="viewport" content="width=device-width, initial-scale=1">
|
||||
<title>Spring AI Alibaba Graph:诊断编排改造</title>
|
||||
<script src="https://cdn.jsdelivr.net/npm/mermaid/dist/mermaid.min.js"></script>
|
||||
<style>
|
||||
:root { --bg:#f4f7fb; --card:rgba(255,255,255,.86); --text:#172033; --muted:#5c6880; --blue:#2563eb; --line:#dbe3f0; }
|
||||
* { box-sizing:border-box; }
|
||||
body { margin:0; font-family:Inter,"PingFang SC","Microsoft YaHei",sans-serif; background:linear-gradient(135deg,#eef4ff,#f8fafc); color:var(--text); line-height:1.75; }
|
||||
.wrap { max-width:1000px; margin:0 auto; padding:42px 24px 110px; }
|
||||
header,.card { background:var(--card); border:1px solid rgba(255,255,255,.9); box-shadow:0 12px 35px rgba(35,55,90,.09); backdrop-filter:blur(14px); border-radius:20px; padding:28px; margin-bottom:22px; }
|
||||
h1 { margin:0 0 8px; font-size:34px; }
|
||||
h2 { margin-top:0; color:#173c85; }
|
||||
h3 { color:#244c92; }
|
||||
code,pre { font-family:"Cascadia Code",Consolas,monospace; }
|
||||
pre { background:#101827; color:#e5edf9; padding:18px; border-radius:14px; overflow:auto; }
|
||||
.tag { display:inline-block; padding:4px 10px; margin-right:6px; border-radius:999px; background:#e4edff; color:#2453a6; font-size:13px; }
|
||||
.mnemonic-card { background:#fff8cf; border:2px dashed #e6b800; padding:16px; border-radius:14px; }
|
||||
.fission-section { background:#fff1f2; border-left:5px solid #e11d48; padding:18px; border-radius:12px; }
|
||||
.truth { border-left:4px solid #2563eb; background:#eff6ff; padding:14px; border-radius:10px; }
|
||||
details { background:#f8fafc; border:1px solid var(--line); border-radius:12px; padding:12px 15px; margin:9px 0; }
|
||||
summary { cursor:pointer; font-weight:700; }
|
||||
.search { position:fixed; bottom:22px; left:50%; transform:translateX(-50%); width:min(720px,calc(100% - 36px)); background:rgba(15,23,42,.93); padding:12px; border-radius:16px; box-shadow:0 15px 40px rgba(0,0,0,.25); z-index:5; }
|
||||
.search input { width:100%; border:0; outline:0; border-radius:10px; padding:12px 14px; font-size:15px; }
|
||||
.hidden { display:none !important; }
|
||||
ul,ol { padding-left:24px; }
|
||||
</style>
|
||||
</head>
|
||||
<body>
|
||||
<div class="wrap" id="content-area">
|
||||
<header>
|
||||
<h1>Spring AI Alibaba Graph:诊断编排改造</h1>
|
||||
<p>作者:叫我小杨同学的小码酱</p>
|
||||
<span class="tag">StateGraph</span><span class="tag">Agent 编排</span><span class="tag">条件边</span><span class="tag">故障诊断</span>
|
||||
</header>
|
||||
|
||||
<section class="card">
|
||||
<h2>0. 核心摘要</h2>
|
||||
<p><strong>让 ReactAgent 继续负责做事,让 StateGraph 负责下一步去哪里。</strong></p>
|
||||
<p>生活类比:Planner、Executor、Verifier 是医院科室,Graph 是分诊和转诊制度。</p>
|
||||
<p class="truth">官方核心模型是 State、Nodes、Edges。本地依赖 1.1.2.0 已确认支持 StateGraph、条件边、编译配置、中断和 threadId。</p>
|
||||
</section>
|
||||
|
||||
<section class="card">
|
||||
<h2>1. 概念破冰</h2>
|
||||
<div class="mnemonic-card">状态记事实,节点做任务,边管下一步,检查点管恢复。</div>
|
||||
<p>当前 ChatService 已经是半个状态机:SequentialAgent 运行 Planner、Executor、Verifier,外层 Java 再根据 PASS、LOW_CONFID、REJECT 决定 Composer 或重试。Graph 改造的价值,是把分散的控制权显式化。</p>
|
||||
<pre>当前:SequentialAgent + 外层 if/else
|
||||
目标:StateGraph 条件边 + ReactAgent 语义节点</pre>
|
||||
</section>
|
||||
|
||||
<section class="card">
|
||||
<h2>2. 深度解析</h2>
|
||||
<p>Supervisor 适合动态选择专科 Agent;Gatekeeper、Verifier 等强制门禁应由代码边控制。普通工具失败也不需要 interrupt,只有等待人工输入或审批时才暂停。</p>
|
||||
<div class="mermaid">
|
||||
flowchart TD
|
||||
S["START"] --> P["Planner"]
|
||||
P --> E["Executor"]
|
||||
E -- "结构有效" --> G["Gatekeeper"]
|
||||
E -- "阻断或非法" --> F["Fallback"]
|
||||
G -- "允许" --> V["Verifier"]
|
||||
G -- "拒绝" --> F
|
||||
V -- "通过或拒绝" --> C["Composer"]
|
||||
V -- "低置信且有预算" --> R["Retry Guard"]
|
||||
R -- "补证据" --> P
|
||||
R -- "停止" --> C
|
||||
C --> X["END"]
|
||||
F --> X
|
||||
</div>
|
||||
|
||||
<h3>状态设计</h3>
|
||||
<p>保存统一诊断上下文、计划、Executor 结构化输出、Gatekeeper 结果、Verifier verdict、重试轮次和最终结果。不要在 State 里复制所有原始日志或完整思考过程。</p>
|
||||
|
||||
<h3>核心伪代码</h3>
|
||||
<pre>StateGraph graph = new StateGraph("diagnosis_workflow", strategies)
|
||||
.addNode("planner", plannerNode)
|
||||
.addNode("executor", executorNode)
|
||||
.addNode("gatekeeper", gatekeeperNode)
|
||||
.addNode("verifier", verifierNode)
|
||||
.addNode("retry_guard", retryGuardNode)
|
||||
.addNode("composer", composerNode)
|
||||
.addNode("fallback", fallbackNode)
|
||||
.addEdge(START, "planner")
|
||||
.addEdge("planner", "executor")
|
||||
.addConditionalEdges("executor", routeAfterExecutor,
|
||||
Map.of("gatekeeper","gatekeeper",
|
||||
"retry_guard","retry_guard",
|
||||
"fallback","fallback"))
|
||||
.addConditionalEdges("gatekeeper", routeAfterGatekeeper,
|
||||
Map.of("verifier","verifier","fallback","fallback"))
|
||||
.addConditionalEdges("verifier", routeAfterVerifier,
|
||||
Map.of("retry_guard","retry_guard",
|
||||
"composer","composer",
|
||||
"fallback","fallback"))
|
||||
.addConditionalEdges("retry_guard", routeAfterRetry,
|
||||
Map.of("planner","planner","composer","composer"))
|
||||
.addEdge("composer", END)
|
||||
.addEdge("fallback", END);</pre>
|
||||
|
||||
<h3>运行边界</h3>
|
||||
<pre>RunnableConfig config = RunnableConfig.builder()
|
||||
.threadId(runId)
|
||||
.addMetadata("sessionId", sessionId)
|
||||
.addMetadata("runId", runId)
|
||||
.build();</pre>
|
||||
<p>一个 diagnosis_run 使用一个 Graph thread,避免同一 session 下多个 run 共享检查点。MemorySaver 不是跨重启持久化。</p>
|
||||
|
||||
<h3>ReactAgent 的接入</h3>
|
||||
<p>本地 ReactAgent 提供 asNode(boolean, boolean)。当前项目第一阶段更适合用适配节点调用已有 Agent,显式控制输入、outputKey 和解析;状态契约稳定后再评估直接 asNode。</p>
|
||||
</section>
|
||||
|
||||
<section class="card fission-section">
|
||||
<h2>3. 深度裂变</h2>
|
||||
<h3>🔍 搜索内化:改造的是控制权,不是 Agent</h3>
|
||||
<p>Graph 的节点可以是 LLM,也可以是普通 Java 代码;ReactAgent 本身已经是子图。所谓“Multi-Agent 改 Graph”,实际上是把跨 Agent 状态转换交给父 Graph。</p>
|
||||
<p>官方页面示例有 OverAllStaste 拼写错误,实际类型是 OverAllState。网站主分支可能领先于本地依赖,最终必须以项目 JAR 和编译测试为准。</p>
|
||||
</section>
|
||||
|
||||
<section class="card">
|
||||
<h2>4. 实战指南</h2>
|
||||
<ol>
|
||||
<li>先定义节点结果状态,不改 Prompt。</li>
|
||||
<li>把现有 Planner、Executor、Verifier 包装为 Node。</li>
|
||||
<li>Gatekeeper 和固定降级做成 Java Node。</li>
|
||||
<li>迁移现有两轮 LOW_CONFID 控制。</li>
|
||||
<li>加入 Executor 阻断、非法结构和 Gatekeeper REJECT 条件边。</li>
|
||||
<li>保留旧 Sequential 链路作为短期回退。</li>
|
||||
<li>最后再增加 HITL、并行和专科 SubAgent。</li>
|
||||
</ol>
|
||||
<h3>避坑</h3>
|
||||
<ul>
|
||||
<li>不要让 Supervisor 决定是否跳过安全门禁。</li>
|
||||
<li>不要把 no_evidence 当成 Executor 失败。</li>
|
||||
<li>不要让多个 run 共用 sessionId 作为 Graph threadId。</li>
|
||||
<li>不要把 MemorySaver 当成生产持久化。</li>
|
||||
<li>不要未验证 messages 传播就直接大量使用 asNode(true, true)。</li>
|
||||
</ul>
|
||||
</section>
|
||||
|
||||
<section class="card">
|
||||
<h2>5. 温故知新</h2>
|
||||
<h3>FAQ</h3>
|
||||
<details><summary>1. Graph 会替代 ReactAgent 吗?</summary><p>不会,ReactAgent 可以作为子图节点继续使用。</p></details>
|
||||
<details><summary>2. 为什么不用 Supervisor 控制失败?</summary><p>失败跳转是确定性规则,不需要增加一次模型决策。</p></details>
|
||||
<details><summary>3. no_evidence 是否直接 fallback?</summary><p>不一定,合法 no-evidence 引用仍要经过 Gatekeeper 和 Verifier。</p></details>
|
||||
<details><summary>4. Composer 必须是 Agent 吗?</summary><p>正常表达可以使用轻量 Agent,系统失败要保留固定模板。</p></details>
|
||||
<details><summary>5. threadId 用什么?</summary><p>当前数据模型下优先使用 runId,sessionId 作为元数据。</p></details>
|
||||
<details><summary>6. 何时需要 Checkpointer?</summary><p>需要暂停、恢复和检查 Graph 历史状态时。</p></details>
|
||||
<details><summary>7. 能直接使用 agent.asNode 吗?</summary><p>可以,但要验证 outputKey、messages 和父子检查点。</p></details>
|
||||
<details><summary>8. AIOps 要单独 Graph 吗?</summary><p>入口和输出策略独立,诊断核心可以共用。</p></details>
|
||||
|
||||
<h3>自测题</h3>
|
||||
<ol>
|
||||
<li>为什么 Executor 工具阻断不应由 Verifier 决定重试?</li>
|
||||
<li>ReplaceStrategy 和 AppendStrategy 各适合什么状态?</li>
|
||||
<li>为什么合法 no_evidence 仍然需要 Gatekeeper?</li>
|
||||
<li>Supervisor 与条件边的决策权有什么不同?</li>
|
||||
<li>为什么 Graph threadId 更适合使用 runId?</li>
|
||||
<li>什么情况下才应该配置 interruptBefore?</li>
|
||||
</ol>
|
||||
</section>
|
||||
</div>
|
||||
|
||||
<div class="search"><input id="search-input" placeholder="搜索本文内容……"></div>
|
||||
<script>
|
||||
mermaid.initialize({ startOnLoad: true, theme: 'neutral' });
|
||||
window.onload = function() {
|
||||
const input = document.getElementById('search-input');
|
||||
if(!input) return;
|
||||
input.addEventListener('input', (e) => {
|
||||
const term = e.target.value.toLowerCase().trim();
|
||||
const contentArea = document.getElementById('content-area');
|
||||
const blocks = contentArea.querySelectorAll('p, li, blockquote, .fission-section, .mnemonic-card, details, .mermaid');
|
||||
if(term.length === 0) {
|
||||
blocks.forEach(el => el.classList.remove('hidden'));
|
||||
document.querySelectorAll('h1, h2, h3').forEach(el => el.classList.remove('hidden'));
|
||||
return;
|
||||
}
|
||||
blocks.forEach(el => el.classList.add('hidden'));
|
||||
document.querySelectorAll('h1, h2, h3').forEach(el => el.classList.add('hidden'));
|
||||
blocks.forEach(el => {
|
||||
if(el.innerText.toLowerCase().includes(term)) {
|
||||
el.classList.remove('hidden');
|
||||
}
|
||||
});
|
||||
});
|
||||
};
|
||||
</script>
|
||||
</body>
|
||||
</html>
|
||||
+383
@@ -0,0 +1,383 @@
|
||||
# Spring AI Alibaba Graph:诊断编排改造
|
||||
|
||||
作者:叫我小杨同学的小码酱
|
||||
标签:Spring AI Alibaba、StateGraph、Agent 编排、条件边、故障诊断
|
||||
|
||||
## 0. 核心摘要
|
||||
|
||||
一句话:让 ReactAgent 继续负责“做事”,让 StateGraph 负责“下一步去哪里”。
|
||||
|
||||
生活类比:Planner、Executor、Verifier 是医院里的不同科室,Graph 是分诊和转诊制度;不能让某个科室自己决定跳过检验和会诊。
|
||||
|
||||
真理锚点:官方文档将 Graph 概括为 State、Nodes、Edges,核心关系是“节点完成工作,边决定下一步做什么”。当前项目依赖的 `spring-ai-alibaba-graph-core:1.1.2.0` 本地 JAR 已确认提供 `StateGraph.addConditionalEdges(...)`、`CompileConfig.interruptBefore/After(...)` 和 `RunnableConfig.threadId(...)`。
|
||||
|
||||
## 1. 概念破冰
|
||||
|
||||
> 巧记:状态记事实,节点做任务,边管下一步,检查点管恢复。
|
||||
|
||||
当前 `ChatService` 已经像一个半成品状态机:`SequentialAgent` 固定执行 Planner、Executor、Verifier,外层 Java 循环再判断 PASS、LOW_CONFID、REJECT,并决定 Composer 或下一轮。问题不是 Agent 不够多,而是状态转换分散在 `SequentialAgent`、Hook 和外层 `if/else` 中。
|
||||
|
||||
```text
|
||||
当前
|
||||
SequentialAgent: Planner -> Executor -> Verifier
|
||||
|
|
||||
ChatService 外层: PASS / LOW_CONFID / REJECT -> Composer / retry
|
||||
|
||||
目标
|
||||
StateGraph 显式表示所有阶段和条件边
|
||||
ReactAgent 作为图中的语义节点继续复用
|
||||
```
|
||||
|
||||
## 2. 深度解析
|
||||
|
||||
### 2.1 为什么不是直接换成 SupervisorAgent
|
||||
|
||||
SupervisorAgent 适合在多个专科 Agent 之间动态选择,例如 Database、Redis、JVM。当前诊断链路中的 Gatekeeper、Verifier 是不能随意跳过的质量门禁。如果让 LLM Supervisor 决定下一步,关键流程会从代码控制变成模型决策。
|
||||
|
||||
当前真正需要的是确定性路由:
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
S["START"] --> P["Planner"]
|
||||
P --> E["Executor"]
|
||||
E -- "证据结构有效" --> G["Gatekeeper"]
|
||||
E -- "执行阻断或非法输出" --> F["Fallback"]
|
||||
G -- "PASS 或 LOW_CONFID" --> V["Verifier"]
|
||||
G -- "REJECT" --> F
|
||||
V -- "PASS" --> C["Composer"]
|
||||
V -- "LOW_CONFID 且有预算" --> R["Retry Guard"]
|
||||
V -- "REJECT 或无预算" --> C
|
||||
R -- "允许补证据" --> P
|
||||
R -- "停止" --> C
|
||||
C --> X["END"]
|
||||
F --> X
|
||||
```
|
||||
|
||||
### 2.2 State 应保存什么
|
||||
|
||||
状态应该保存跨节点需要共享的原始事实和结构化结果,不保存拼好的 Prompt,也不保存无边界增长的模型思考过程。
|
||||
|
||||
建议的语义状态:
|
||||
|
||||
- `diagnosis_context`:入口适配后的统一诊断上下文。
|
||||
- `planner_plan`:Planner 的结构化计划。
|
||||
- `executor_output`:`executor_evidence_v2`。
|
||||
- `executor_status`:成功、阻断、非法输出或失败。
|
||||
- `gatekeeper_result`:代码验真结果。
|
||||
- `verifier_output`、`verdict`:可推导性结果。
|
||||
- `retry_context`、`round`:有限补证据状态。
|
||||
- `final_answer`:最终表达。
|
||||
- `failure_reason`:确定性的失败原因。
|
||||
|
||||
现有 Trace 已由 `agent_step` 和 `tool_invocation` 持久化,Graph State 不需要复制所有原始日志。
|
||||
|
||||
### 2.3 KeyStrategy 如何选择
|
||||
|
||||
诊断状态大多使用 `ReplaceStrategy`,因为每个阶段产生当前轮的最新结果。只有确实需要累计的轻量事件列表才使用 `AppendStrategy`。
|
||||
|
||||
```java
|
||||
KeyStrategyFactory diagnosisStateStrategies() {
|
||||
return () -> {
|
||||
Map<String, KeyStrategy> strategies = new HashMap<>();
|
||||
strategies.put("diagnosis_context", new ReplaceStrategy());
|
||||
strategies.put("planner_plan", new ReplaceStrategy());
|
||||
strategies.put("executor_output", new ReplaceStrategy());
|
||||
strategies.put("executor_status", new ReplaceStrategy());
|
||||
strategies.put("gatekeeper_result", new ReplaceStrategy());
|
||||
strategies.put("verifier_output", new ReplaceStrategy());
|
||||
strategies.put("verdict", new ReplaceStrategy());
|
||||
strategies.put("retry_context", new ReplaceStrategy());
|
||||
strategies.put("round", new ReplaceStrategy());
|
||||
strategies.put("final_answer", new ReplaceStrategy());
|
||||
strategies.put("failure_reason", new ReplaceStrategy());
|
||||
return strategies;
|
||||
};
|
||||
}
|
||||
```
|
||||
|
||||
### 2.4 ReactAgent 怎样放进 Graph
|
||||
|
||||
本地 `1.1.2.0` 的 `ReactAgent` 提供 `asNode(boolean includeContents, boolean returnReasoningContents)`,可以直接作为子图节点。但当前项目 Planner、Executor、Verifier 的输入组织和输出解析已经有较多定制,第一阶段更推荐使用适配节点显式调用现有 Agent:
|
||||
|
||||
```java
|
||||
var plannerNode = node_async((state, config) -> {
|
||||
DiagnosisContext context = requireContext(state);
|
||||
String prompt = plannerInput(context, state.value("retry_context").orElse(null));
|
||||
|
||||
AssistantMessage response = plannerAgent.call(prompt, childConfig(config, "planner"));
|
||||
PlannerPlan plan = plannerParser.parse(extractText(response));
|
||||
|
||||
return Map.of(
|
||||
"planner_plan", plan,
|
||||
"failure_reason", ""
|
||||
);
|
||||
});
|
||||
```
|
||||
|
||||
适配节点的好处是不会意外把父图全部 `messages` 注入所有 Agent,也能继续复用当前解析器、Hook、Prompt 和 ToolCallback。
|
||||
|
||||
等统一状态契约稳定后,可以评估:
|
||||
|
||||
```java
|
||||
graph.addNode("planner", plannerAgent.asNode(false, false));
|
||||
```
|
||||
|
||||
但需要先验证父子图的 `messages`、outputKey 和 Checkpointer 是否符合预期。
|
||||
|
||||
### 2.5 Executor 节点只报告状态,不决定路由
|
||||
|
||||
```java
|
||||
var executorNode = node_async((state, config) -> {
|
||||
try {
|
||||
PlannerPlan plan = requirePlan(state);
|
||||
DiagnosisContext context = requireContext(state);
|
||||
|
||||
AssistantMessage response = executorAgent.call(
|
||||
executorInput(context, plan),
|
||||
childConfig(config, "executor")
|
||||
);
|
||||
|
||||
ExecutorEvidence output = executorParser.parse(extractText(response));
|
||||
|
||||
if (!output.isStructurallyValid()) {
|
||||
return Map.of(
|
||||
"executor_status", "INVALID_OUTPUT",
|
||||
"failure_reason", "executor_evidence_v2 解析失败"
|
||||
);
|
||||
}
|
||||
|
||||
// no_evidence 仍然是合法结构,需要交给 Gatekeeper 验证真实引用。
|
||||
return Map.of(
|
||||
"executor_status", "COMPLETED",
|
||||
"executor_output", output
|
||||
);
|
||||
}
|
||||
catch (ToolCapabilityBlockedException e) {
|
||||
return Map.of(
|
||||
"executor_status", "TOOL_BLOCKED",
|
||||
"failure_reason", e.getMessage()
|
||||
);
|
||||
}
|
||||
catch (Exception e) {
|
||||
return Map.of(
|
||||
"executor_status", "FAILED",
|
||||
"failure_reason", safeMessage(e)
|
||||
);
|
||||
}
|
||||
});
|
||||
```
|
||||
|
||||
注意:`no_evidence` 不是 Executor 失败。当前项目已经用 `$.no_evidence` 表达“查询成功但无匹配证据”,它仍应进入 Gatekeeper 和 Verifier,防止被过度表达为“问题不存在”。
|
||||
|
||||
### 2.6 Gatekeeper 节点保持纯代码
|
||||
|
||||
```java
|
||||
var gatekeeperNode = node_async((state, config) -> {
|
||||
ExecutorEvidence output = requireExecutorOutput(state);
|
||||
String runId = metadata(config, "runId");
|
||||
|
||||
GatekeeperResult result = executorGatekeeperService.validate(runId, output);
|
||||
|
||||
return Map.of(
|
||||
"gatekeeper_result", result,
|
||||
"gatekeeper_status", result.severity()
|
||||
);
|
||||
});
|
||||
```
|
||||
|
||||
Gatekeeper 不需要改造成 Agent。它负责确定性引用验真,是 Graph 中的普通 Java Node。
|
||||
|
||||
### 2.7 条件边是改造核心
|
||||
|
||||
```java
|
||||
StateGraph graph = new StateGraph("diagnosis_workflow", diagnosisStateStrategies())
|
||||
.addNode("planner", plannerNode)
|
||||
.addNode("executor", executorNode)
|
||||
.addNode("gatekeeper", gatekeeperNode)
|
||||
.addNode("verifier", verifierNode)
|
||||
.addNode("retry_guard", retryGuardNode)
|
||||
.addNode("composer", composerNode)
|
||||
.addNode("fallback", fallbackNode)
|
||||
|
||||
.addEdge(START, "planner")
|
||||
.addEdge("planner", "executor")
|
||||
|
||||
.addConditionalEdges(
|
||||
"executor",
|
||||
edge_async(state -> switch (stringValue(state, "executor_status")) {
|
||||
case "COMPLETED" -> "gatekeeper";
|
||||
case "INVALID_OUTPUT" -> retryAvailable(state) ? "retry_guard" : "fallback";
|
||||
case "TOOL_BLOCKED", "FAILED" -> "fallback";
|
||||
default -> "fallback";
|
||||
}),
|
||||
Map.of(
|
||||
"gatekeeper", "gatekeeper",
|
||||
"retry_guard", "retry_guard",
|
||||
"fallback", "fallback"
|
||||
)
|
||||
)
|
||||
|
||||
.addConditionalEdges(
|
||||
"gatekeeper",
|
||||
edge_async(state -> switch (stringValue(state, "gatekeeper_status")) {
|
||||
case "REJECT" -> "fallback";
|
||||
default -> "verifier";
|
||||
}),
|
||||
Map.of("verifier", "verifier", "fallback", "fallback")
|
||||
)
|
||||
|
||||
.addConditionalEdges(
|
||||
"verifier",
|
||||
edge_async(state -> switch (stringValue(state, "verdict")) {
|
||||
case "LOW_CONFID" -> retryAvailable(state) ? "retry_guard" : "composer";
|
||||
case "PASS", "REJECT" -> "composer";
|
||||
default -> "fallback";
|
||||
}),
|
||||
Map.of(
|
||||
"retry_guard", "retry_guard",
|
||||
"composer", "composer",
|
||||
"fallback", "fallback"
|
||||
)
|
||||
)
|
||||
|
||||
.addConditionalEdges(
|
||||
"retry_guard",
|
||||
edge_async(state -> shouldRetry(state) ? "planner" : "composer"),
|
||||
Map.of("planner", "planner", "composer", "composer")
|
||||
)
|
||||
|
||||
.addEdge("composer", END)
|
||||
.addEdge("fallback", END);
|
||||
```
|
||||
|
||||
### 2.8 编译和执行
|
||||
|
||||
```java
|
||||
SaverConfig saverConfig = SaverConfig.builder()
|
||||
.register(new MemorySaver())
|
||||
.build();
|
||||
|
||||
CompileConfig compileConfig = CompileConfig.builder()
|
||||
.recursionLimit(20)
|
||||
.saverConfig(saverConfig)
|
||||
.build();
|
||||
|
||||
CompiledGraph compiledGraph = graph.compile(compileConfig);
|
||||
|
||||
RunnableConfig runConfig = RunnableConfig.builder()
|
||||
// 一个 diagnosis_run 对应一个 Graph thread,避免同一 session 下多个 run 混状态。
|
||||
.threadId(runId)
|
||||
.addMetadata("sessionId", sessionId)
|
||||
.addMetadata("runId", runId)
|
||||
.build();
|
||||
|
||||
Map<String, Object> initialState = Map.of(
|
||||
"diagnosis_context", diagnosisContext,
|
||||
"round", 1
|
||||
);
|
||||
|
||||
Optional<OverAllState> finalState = compiledGraph.invoke(initialState, runConfig);
|
||||
```
|
||||
|
||||
`MemorySaver` 只适合进程内检查点,不等价于重启后可恢复的持久化。当前项目已有数据库 Trace,可以先把 Graph 用于流程控制;只有真正需要跨进程暂停恢复时,再引入持久 Checkpointer 或显式恢复模型。
|
||||
|
||||
### 2.9 Chat 与 AIOps 怎样共用
|
||||
|
||||
入口适配不同,公共 Graph 接收统一 `DiagnosisContext`:
|
||||
|
||||
```java
|
||||
DiagnosisContext chatContext = chatAdapter.from(question, history);
|
||||
DiagnosisContext aiOpsContext = aiOpsAdapter.from(alertPayload);
|
||||
|
||||
DiagnosisResult chatResult = diagnosisGraph.execute(chatContext, sessionId, runId);
|
||||
DiagnosisResult aiOpsResult = diagnosisGraph.execute(aiOpsContext, sessionId, runId);
|
||||
|
||||
return chatOutputAdapter.render(chatResult);
|
||||
return aiOpsOutputAdapter.render(aiOpsResult);
|
||||
```
|
||||
|
||||
Chat 的普通问答仍走轻量链路;复杂诊断才进入公共 Graph。AIOps 在入口阶段固定主告警范围,Graph 内部继续复用证据收集和验证。
|
||||
|
||||
### 2.10 人工中断怎么放
|
||||
|
||||
普通工具失败不需要 interrupt,条件边即可。只有确实需要等待人工输入或审批时才增加节点:
|
||||
|
||||
```java
|
||||
CompileConfig compileConfig = CompileConfig.builder()
|
||||
.saverConfig(saverConfig)
|
||||
.interruptBefore("human_review")
|
||||
.build();
|
||||
```
|
||||
|
||||
恢复时使用相同 `threadId` 和 checkpoint 信息,并通过 `RunnableConfig.builder(oldConfig).resume()` 或状态更新接口继续。具体恢复协议需要结合当前版本做集成测试,不能只凭文档假设。
|
||||
|
||||
## 3. 深度裂变
|
||||
|
||||
<div class="fission-section">
|
||||
|
||||
### 🔍 搜索内化:真正的改造对象不是 Agent,而是控制权
|
||||
|
||||
官方文档和本地 `1.1.2.0` JAR 均证明 Graph 节点既可以是 LLM,也可以是普通 Java 代码,条件边由状态决定目标节点。`ReactAgent` 本身已经是一个子图,并提供 `asNode(...)` 适配能力。
|
||||
|
||||
因此“从 Multi-Agent 改成 Graph”并不准确。更准确的是:把跨 Agent 的状态转换从高层 Flow 抽出来,交给父 Graph;ReactAgent 继续作为子图存在。
|
||||
|
||||
文档页面示例存在 `OverAllStaste` 拼写错误,实际类名是 `OverAllState`。网站主分支可能领先于本地依赖,因此最终应以项目锁定版本的 JAR 签名和编译测试为准。
|
||||
|
||||
</div>
|
||||
|
||||
## 4. 实战指南
|
||||
|
||||
### 4.1 最小迁移顺序
|
||||
|
||||
1. 定义统一的节点结果状态,不先改 Prompt。
|
||||
2. 把现有 Planner、Executor、Verifier 调用包装为 Graph Node。
|
||||
3. 将 Gatekeeper 和固定降级模板做成普通 Java Node。
|
||||
4. 先迁移当前两轮 LOW_CONFID 循环。
|
||||
5. 为 Executor 阻断、非法结构、Gatekeeper REJECT 增加条件边。
|
||||
6. 保留原 Sequential 链路作为回退,完成行为对比后再删除。
|
||||
7. 最后再考虑 HITL、并行和专科 SubAgent。
|
||||
|
||||
### 4.2 常见反模式
|
||||
|
||||
- 把每个异常都交给 LLM Supervisor 决策。
|
||||
- Graph State 存放所有原始日志和完整思考过程。
|
||||
- 把 `no_evidence` 当作 Executor 执行失败。
|
||||
- 同一 session 的多个 run 共用一个 Graph `threadId`。
|
||||
- 一开始就设计几十个节点和完整 Incident 状态机。
|
||||
- 未验证父子图消息传播就直接大量使用 `ReactAgent.asNode(true, true)`。
|
||||
- 把 `MemorySaver` 当成生产级持久化。
|
||||
|
||||
### 4.3 ROI
|
||||
|
||||
收益:条件分支可见、失败可测试、门禁不可跳过、Trace 更容易与节点对齐。
|
||||
代价:需要维护状态契约、条件边和父子图上下文,并增加 Graph 级测试。
|
||||
判断标准:如果当前只有固定顺序且失败直接结束,SequentialAgent 更简单;当局部重试、降级、HITL 和多入口策略已经出现时,StateGraph 的控制收益开始超过复杂度。
|
||||
|
||||
## 5. 温故知新
|
||||
|
||||
### FAQ
|
||||
|
||||
1. **Graph 会替代 ReactAgent 吗?** 不会,ReactAgent 可以作为 Graph 节点或由适配节点调用。
|
||||
2. **为什么不用 Supervisor 控制失败?** 失败跳转是确定性规则,不应增加一次 LLM 决策。
|
||||
3. **no_evidence 是否直接走 fallback?** 不一定。合法的 no-evidence 引用仍要经过 Gatekeeper 和 Verifier。
|
||||
4. **Composer 是否必须是 Agent?** 正常表达可以是轻量 Agent,系统失败场景应保留固定模板。
|
||||
5. **threadId 用 sessionId 还是 runId?** 当前模型下优先用 runId,避免同一会话多次运行状态串扰。
|
||||
6. **什么时候需要 Checkpointer?** 需要暂停、恢复、查看历史 Graph 状态时;普通 Trace 持久化不自动等于 Graph Checkpoint。
|
||||
7. **能否直接使用 agent.asNode?** 可以,但需要验证输入、outputKey、messages 和父子 Checkpointer 行为。
|
||||
8. **AIOps 是否要单独一张 Graph?** 可以先共用诊断核心,入口范围策略和输出报告保持独立。
|
||||
|
||||
### 自测题
|
||||
|
||||
1. 为什么 Executor 工具阻断不应该由 Verifier 判断是否重试?
|
||||
2. `ReplaceStrategy` 和 `AppendStrategy` 在诊断状态中分别适合什么数据?
|
||||
3. 为什么合法 `no_evidence` 仍然需要 Gatekeeper?
|
||||
4. SupervisorAgent 和 StateGraph 条件边的决策权有什么区别?
|
||||
5. 为什么当前 Graph `threadId` 更适合使用 runId?
|
||||
6. 什么情况下才应该增加 `interruptBefore`?
|
||||
|
||||
### 参考资源
|
||||
|
||||
- https://java2ai.com/docs/frameworks/graph-core/core/core-library
|
||||
- https://java2ai.com/docs/frameworks/graph-core/quick-start
|
||||
- Spring AI Alibaba 本地依赖:`spring-ai-alibaba-graph-core:1.1.2.0`
|
||||
- 当前项目:`ChatService`、`AiOpsService`、`ExecutorGatekeeperService`、`VerifierInputHook`
|
||||
Reference in New Issue
Block a user