Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
4b9cf7c5cc | ||
|
|
190013c901 | ||
|
|
208a231113 | ||
|
|
99e490f227 | ||
|
|
1460dd1e99 | ||
|
|
42ba204532 | ||
|
|
581daffdad | ||
|
|
a36fe72639 |
@@ -103,6 +103,11 @@
|
|||||||
- 使用场景:Trace API、Trace UI、Verifier 审计、评测 fixture 和人工排查。
|
- 使用场景:Trace API、Trace UI、Verifier 审计、评测 fixture 和人工排查。
|
||||||
- 边界:Diagnosis Trace 是聚合视图,不要求单独的 trace 主表;当前 trace 明细由 `agent_step` 和 `tool_invocation` 表承载。
|
- 边界:Diagnosis Trace 是聚合视图,不要求单独的 trace 主表;当前 trace 明细由 `agent_step` 和 `tool_invocation` 表承载。
|
||||||
|
|
||||||
|
### Diagnosis Orchestration Trace
|
||||||
|
- 定义:一次 Diagnosis Run 的紧凑编排审计摘要,记录实际节点路径、条件边原因、技术重试、降级和终止原因。
|
||||||
|
- 使用场景:解释诊断编排为何进入某个节点、为何重试或为何提前终止,并支撑路由验收和人工审计。
|
||||||
|
- 边界:它是 Diagnosis Trace 的编排维度,不是完整事件日志、自评估结果或持久恢复检查点;不保存 Prompt、模型思考、工具原文和完整编排上下文快照。
|
||||||
|
|
||||||
### Flyway
|
### Flyway
|
||||||
- 定义:数据库版本迁移工具,管理 SQL 脚本的版本化执行
|
- 定义:数据库版本迁移工具,管理 SQL 脚本的版本化执行
|
||||||
- 配置:spring.flyway.enabled=true, baseline-on-migrate=true
|
- 配置:spring.flyway.enabled=true, baseline-on-migrate=true
|
||||||
@@ -156,14 +161,14 @@
|
|||||||
- 边界:最终诊断结论必须被 evidence tools 支撑,不能仅由 skill 正文支撑。
|
- 边界:最终诊断结论必须被 evidence tools 支撑,不能仅由 skill 正文支撑。
|
||||||
|
|
||||||
### Verifier Skill Isolation
|
### Verifier Skill Isolation
|
||||||
- 定义:Chat Verifier 与 skill 系统隔离,只校验 Executor 答案和 `tool_trace_summary`。
|
- 定义:Chat Verifier 与 skill 系统隔离,只校验 Gatekeeper 投影后的 `verified_executor_output` 和 `verified_evidence`。
|
||||||
- 使用场景:防止 Verifier 把 playbook 指令当作事实证据;Verifier 只判断已有证据是否支持结论。
|
- 使用场景:防止 Verifier 把 playbook 指令当作事实证据;Verifier 只判断已有证据是否支持结论。
|
||||||
- 边界:Verifier 不接收 `skill_catalog`,不暴露 `read_skill`,不读取 `SKILL.md`。
|
- 边界:Verifier 不接收 `skill_catalog`,不暴露 `read_skill`,不读取 `SKILL.md`、完整 `tool_trace_summary` 或未经验真的 Executor 自由文本。
|
||||||
|
|
||||||
## Diagnosis Playbook Business Rules
|
## Diagnosis Playbook Business Rules
|
||||||
|
|
||||||
- Planner 只看 skill metadata,输出 `selected_skill`、`selection_reason` 和 plan。
|
- Planner 只看 skill metadata,输出 `selected_skill`、`selection_reason` 和 plan。
|
||||||
- Executor 才能调用 `read_skill(selected_skill)`,并且读取 skill 后仍必须调用 evidence tools。
|
- Executor 才能调用 `read_skill(selected_skill)`,并且读取 skill 后仍必须调用 evidence tools。
|
||||||
- Skill 正文不得替代 `lookup_knowledge`、日志、指标或告警数据。
|
- Skill 正文不得替代 `lookup_knowledge`、日志、指标或告警数据。
|
||||||
- Verifier 只基于 `tool_trace_summary` 校验事实,不基于 skill 正文校验事实。
|
- Verifier 只基于 Gatekeeper 通过的 `verified_executor_output` 和 `verified_evidence` 校验事实,不基于 skill 正文、完整工具 Trace 或未验真输出校验事实。
|
||||||
- 当前阶段保留单 active skill 白名单:`diagnose-mysql-connection-pool`。
|
- 当前阶段保留单 active skill 白名单:`diagnose-mysql-connection-pool`。
|
||||||
|
|||||||
@@ -4,6 +4,12 @@
|
|||||||
|
|
||||||
| 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|
| 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|
||||||
|---|---|---|---|---|---|---|
|
|---|---|---|---|---|---|---|
|
||||||
|
| 2026-07-17 | chat-diagnosis-stategraph-cleanup-docs | 清理旧诊断编排闭包,对齐当前文档与 demo contract,并完成 ISS-011 最终 live、日志和数据库验收。 | Chat diagnosis orchestration/cleanup | legacy closure, current docs, orchestration trace, Maven E2E, MySQL ownership | openspec/changes/archive/2026-07-20-chat-diagnosis-stategraph-cleanup-docs | archived |
|
||||||
|
| 2026-07-17 | chat-diagnosis-stategraph-test-suite | 建立 Workflow、Node Contract、Chat Integration 三层权威测试体系并退役旧 Hook implementation tests。 | Chat diagnosis orchestration/testing | workflow test, node contract, Chat integration, coverage matrix, Hook test retirement | openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-test-suite | archived |
|
||||||
|
| 2026-07-17 | chat-diagnosis-stategraph-chatservice-cutover | 将复杂 Chat 单轨切换到 Diagnosis StateGraph,并增加 Run 级 orchestration trace 和 verified-only Verifier 输入。 | Chat diagnosis orchestration/production cutover | ChatService, CompiledGraph stream, runId metadata, orchestration trace, verified-only prompt, V012 | openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-chatservice-cutover | archived |
|
||||||
|
| 2026-07-17 | chat-diagnosis-stategraph-real-nodes | 接入真实 Agent/Java Nodes、显式 Gatekeeper、可信输入投影、关键证据补查与安全 Fallback,暂不切换生产入口。 | Chat diagnosis orchestration/nodes | ReactAgent adapter, Gatekeeper node, verified input, evidence retry, safe fallback, CompiledGraph | openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-real-nodes | archived |
|
||||||
|
| 2026-07-17 | chat-diagnosis-stategraph-routing-skeleton | 实现未接生产入口的 Diagnosis StateGraph 骨架、有限路由和 Fake Node 测试。 | Chat diagnosis orchestration/graph | StateGraph, fake node, conditional edge, retry counter, orchestration events, trace builder | openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-routing-skeleton | archived |
|
||||||
|
| 2026-07-17 | chat-diagnosis-stategraph-design-freeze | 冻结 ISS-011 的 Graph State、条件边、有限重试、安全降级、审计和测试迁移边界。 | Chat diagnosis orchestration/design | StateGraph, runId, Gatekeeper, verified evidence, fallback, orchestration trace, test migration | openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-design-freeze | archived |
|
||||||
| 2026-07-10 | session-run-trace-isolation | 拆分会话态和运行态,引入 runId 隔离 Trace、Feedback、AIOps 和 demo 链路。 | Trace/session/run isolation | chat_session, diagnosis_run, runId, trace exact run, feedback fallback, AIOps SSE metadata, baseline drift | openspec/changes/archive/2026-07-10-session-run-trace-isolation | archived |
|
| 2026-07-10 | session-run-trace-isolation | 拆分会话态和运行态,引入 runId 隔离 Trace、Feedback、AIOps 和 demo 链路。 | Trace/session/run isolation | chat_session, diagnosis_run, runId, trace exact run, feedback fallback, AIOps SSE metadata, baseline drift | openspec/changes/archive/2026-07-10-session-run-trace-isolation | archived |
|
||||||
| 2026-07-09 | interview-demo-quality-audit | 增加面试演示前置质量审计,覆盖 prompt、Gatekeeper 和评测基线。 | Agent eval/demo/Prompt audit | interview demo preflight, prompt_audit, gatekeeper rules, diagnosis baseline, 12 fixtures | openspec/changes/archive/2026-07-09-interview-demo-quality-audit | archived |
|
| 2026-07-09 | interview-demo-quality-audit | 增加面试演示前置质量审计,覆盖 prompt、Gatekeeper 和评测基线。 | Agent eval/demo/Prompt audit | interview demo preflight, prompt_audit, gatekeeper rules, diagnosis baseline, 12 fixtures | openspec/changes/archive/2026-07-09-interview-demo-quality-audit | archived |
|
||||||
| 2026-07-08 | executor-composer-final-answer | 引入 Composer 生成最终回答,只使用 Verifier 允许的结论材料。 | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
|
| 2026-07-08 | executor-composer-final-answer | 引入 Composer 生成最终回答,只使用 Verifier 允许的结论材料。 | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
|
||||||
|
|||||||
+52
@@ -0,0 +1,52 @@
|
|||||||
|
# Chat Diagnosis StateGraph ChatService Cutover 验收
|
||||||
|
|
||||||
|
## 结果
|
||||||
|
|
||||||
|
已接受。OpenSpec tasks 26/26 完成,阶段 3 可归档。
|
||||||
|
|
||||||
|
## 验证
|
||||||
|
|
||||||
|
### 静态验证
|
||||||
|
|
||||||
|
- `git diff --check`:通过。
|
||||||
|
- source/reference isolation:cutover 核心文件中的 `SequentialAgent`、`VerifierContextHolder`、`VerifierInputHook`、`tool_trace_summary` 命中 0;核心 TODO/FIXME/placeholder 命中 0。
|
||||||
|
- schema whitelist:V012 为 1 个 ALTER TABLE、1 个 ADD COLUMN、0 CREATE、0 DROP,目标仅 `diagnosis_run.orchestration_trace JSON NULL`。
|
||||||
|
- `openspec validate chat-diagnosis-stategraph-chatservice-cutover --strict`:通过。
|
||||||
|
- `openspec validate --specs --strict`:14 passed,0 failed。
|
||||||
|
|
||||||
|
### 脚本验证
|
||||||
|
|
||||||
|
- focused Maven regression:Graph runtime/result mapper/real nodes、ChatService cutover、Trace、Controller、Repository、Gatekeeper、Composer、Eval,共 27 suites / 119 tests;0 failures、0 errors、0 skipped。
|
||||||
|
- `mvn -q -DskipTests test-compile`:通过。
|
||||||
|
- `ChatServiceGraphIntegrationTest` 覆盖 SUCCESS、handled Fallback、unhandled failure、blank answer/partial trace 和同 session 多 Run 隔离。
|
||||||
|
|
||||||
|
### 浏览器/人工验证
|
||||||
|
|
||||||
|
- 未运行。阶段 3 的公开协议由 Controller/service tests 覆盖,完整人工/live 验收按用户规则保留到阶段 5。
|
||||||
|
|
||||||
|
### 未验证
|
||||||
|
|
||||||
|
- 未使用 Maven 启动应用做 live E2E。
|
||||||
|
- 未检查 `logs/` 运行日志。
|
||||||
|
- 未执行 `scripts/query_mysql.py` 查询真实数据库。
|
||||||
|
- 原因:用户明确要求只有阶段 5 全部完成后统一执行端到端、日志和数据库验收;阶段 3 只做风险相关自动化验证。
|
||||||
|
- 剩余风险:真实模型/工具调用下的 Prompt 行为、Flyway 在真实 MySQL 的应用结果和最终 Trace 数据须由阶段 5 E2E 证明。
|
||||||
|
|
||||||
|
## 已完成范围
|
||||||
|
|
||||||
|
- 复杂 Chat 唯一生产编排切换到 Diagnosis StateGraph,公开 ChatResult 保持兼容。
|
||||||
|
- Run 生命周期、metrics、Eval、finally cleanup 与 runId 隔离保持;handled Fallback=SUCCESS,未处理/空答案=FAILED。
|
||||||
|
- 新增 Run 级 compact orchestration trace,并只在 Trace run 对象暴露解析结果。
|
||||||
|
- Verifier Prompt/self-evaluation 使用 verified-only 数据,不伪造未发生 verdict/event。
|
||||||
|
- 移除冲突的 Sequential 实现测试并以 Graph/public contract tests 替代。
|
||||||
|
|
||||||
|
## Bug 修复和诊断
|
||||||
|
|
||||||
|
- 替代测试错误引用 `com.superbiz.agent.tool.ToolInvocationRecorder`:通过定义/引用搜索确认类型位于同一 `service` 包,删除错误 import。
|
||||||
|
- no-answer 新测试错误期待 inner cause:分类为测试断言偏差,改为公开 wrapper error,同时保留 FAILED 与真实 partial trace 核心断言。
|
||||||
|
|
||||||
|
## 交接
|
||||||
|
|
||||||
|
- OpenSpec archive:已同步 5 份 delta specs,并归档到 `openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-chatservice-cutover/`。
|
||||||
|
- 下一步:审查精确 Git diff 并完成阶段 3 独立提交,之后才启动阶段 4。
|
||||||
|
- OpenSpec 归档确认:用户已明确授权后续阶段直接实现/归档;归档已完成。
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
# Chat Diagnosis StateGraph ChatService Cutover Brief
|
||||||
|
|
||||||
|
## 背景
|
||||||
|
|
||||||
|
- 用户目标:将 ISS-011 阶段 3 作为独立 sm-flow,正式切换复杂 Chat 生产编排并增加 Run 级 orchestration trace。
|
||||||
|
- 当前问题:真实 Diagnosis Graph Nodes 已存在,但生产入口仍依赖 SequentialAgent、外层重试和 ThreadLocal/Hook 隐式状态。
|
||||||
|
- 关联 OpenSpec:`openspec/changes/chat-diagnosis-stategraph-chatservice-cutover/`
|
||||||
|
- devflow 分档:complex。
|
||||||
|
|
||||||
|
## 范围
|
||||||
|
|
||||||
|
- 本次要做:复杂 Chat 单轨 Graph cutover;复用四类 Agent builder;Graph result/self-evaluation 映射;V012 Run JSON 字段;Trace run-only 投影;verified-only Verifier Prompt;必要回归测试。
|
||||||
|
- 本次不做:不改 `/api/chat` 请求/响应,不做历史 trace 回填,不删除仍供历史代码/测试使用的 Hook/ThreadLocal 类型,不完成阶段 4 测试体系全面收敛,不运行 live E2E/log/DB 验收。
|
||||||
|
- 影响区域:ChatService、Diagnosis Graph runtime/result mapper、DiagnosisRun/Flyway、Trace DTO/service、Verifier Prompt、Graph/Service/Trace tests。
|
||||||
|
|
||||||
|
## OpenSpec 对齐
|
||||||
|
|
||||||
|
- proposal 覆盖状态:已覆盖。
|
||||||
|
- design 覆盖状态:已覆盖。
|
||||||
|
- specs 覆盖状态:已覆盖,5 个 delta capabilities。
|
||||||
|
- tasks 覆盖状态:26/26 已完成。
|
||||||
+207
@@ -0,0 +1,207 @@
|
|||||||
|
# Chat Diagnosis StateGraph ChatService Cutover Decisions
|
||||||
|
|
||||||
|
## Entry Summary
|
||||||
|
|
||||||
|
- 问题:复杂 Chat 仍使用 SequentialAgent + ThreadLocal/Hook 隐式状态机,真实 Graph 尚未成为生产入口,也没有 Run 级 orchestration trace 持久化和 API 投影。
|
||||||
|
- 期望:阶段 3 独立完成生产 cutover、Run/Trace 映射和必要测试,归档并提交后才进入阶段 4。
|
||||||
|
- 分档:complex。
|
||||||
|
- Change:`chat-diagnosis-stategraph-chatservice-cutover`。
|
||||||
|
- 授权:用户已要求后续阶段直接实现,不再逐 checkpoint 等待;阶段门禁、独立 archive/commit 与阶段 5 才 E2E 约束不变。
|
||||||
|
|
||||||
|
## Context Sources
|
||||||
|
|
||||||
|
- `mvp/issues/active/ISS-011-chat-diagnosis-stategraph-orchestration.md` 阶段 3、协议影响、Run 状态和验收章节。
|
||||||
|
- 阶段 0–2 OpenSpec archives、devflow acceptance/decisions 与当前四份 Graph 主 specs。
|
||||||
|
- `devflow/glossary/CONTEXT.md` 中 Chat Session、Diagnosis Run、Diagnosis Trace、Diagnosis Orchestration Trace 的边界。
|
||||||
|
- `ChatService` 生产调用链、四个 Agent builder、Run persistence/self-evaluation/metrics 逻辑。
|
||||||
|
- `DiagnosisRun`、V011 migration、`DiagnosisTraceResponse`、`DiagnosisTraceService` 与相关 tests。
|
||||||
|
- `DiagnosisGraphFactory`、真实 action factory、Node adapters、final trace builder 和本地 CompiledGraph API。
|
||||||
|
|
||||||
|
## Question Pool
|
||||||
|
|
||||||
|
| # | 维度 | 问题 | 模式 | 状态 |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| Q1 | 术语 | orchestration trace 是否属于 self-evaluation 或完整 Trace event log? | evidence-driven | 已解决 |
|
||||||
|
| Q2 | 边界 | 阶段 3 是否改变 `/api/chat`,以及 Trace 字段出现在哪一层? | user-interview(既有冻结决策) | 已确认 |
|
||||||
|
| Q3 | 生命周期 | LOW_CONFID/REJECT/Fallback 是否应使用 FAILED 或新增 DEGRADED? | user-interview(既有冻结决策) | 已确认 |
|
||||||
|
| Q4 | 错误处理 | Graph 异常时如何 best-effort 保存已有编排事实而不伪造 event? | evidence-driven | 已解决 |
|
||||||
|
| Q5 | 技术实现 | 是否复用现有 Agent factory/Hook/ToolCallback 与真实 Node assembly? | evidence-driven | 已解决 |
|
||||||
|
| Q6 | 安全 | Graph Verifier Prompt/self-evaluation 是否可继续使用 raw Executor 和完整 tool trace? | evidence-driven | 已解决 |
|
||||||
|
| Q7 | 验收 | 阶段 3 是否需要新增测试,是否现在运行 live E2E? | user-interview(用户最新规则) | 已确认 |
|
||||||
|
|
||||||
|
## Evidence-driven Findings
|
||||||
|
|
||||||
|
- Q1:glossary 与阶段 0 spec 已定义 orchestration trace 为 Run 级紧凑编排摘要,与 self-evaluation、AgentStep/ToolInvocation 和 checkpoint 分离;无需新增术语或 ADR。
|
||||||
|
- Q4:所有已定义 Agent/Java 失败应由 Graph 路由到 handled Fallback 并返回 final state;只有真实可取得的 final/partial state 才可构造 trace。无法取得 state 的未处理异常标记 Run FAILED,不得生成虚假 transition。
|
||||||
|
- Q5:`DiagnosisRealGraphActionsFactory` 已提供真实 Node assembly;ChatService 现有四个 builder、AgentLoggingHook、Skills hooks、method tools 和 ToolCallbacks 可直接构造 ReactAgent,再通过 `ReactAgentDiagnosisInvoker` 注入,不得复制 Prompt/Agent factory。
|
||||||
|
- Q6:阶段 2 spec 已禁止 Graph Verifier 使用 raw Executor/full tool trace;当前 `chat-verifier-prompt.md` 仍描述旧 Hook payload,是阶段 3 必须同步的明确 gap。self-evaluation 可保存 Graph 的 verified output、Gatekeeper audit 和 Composer audit,但不能重新引入 raw 输入。
|
||||||
|
|
||||||
|
## User-interview Confirmations
|
||||||
|
|
||||||
|
| 问题 | 用户原话/既有确认 | 确认状态 | OpenSpec 回写 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| Q2 外部协议与 Trace 层级 | ISS-011 已冻结 `/api/chat` 不变,Trace 只在 `run.orchestrationTrace` 增解析对象 | 已确认 | proposal |
|
||||||
|
| Q3 Run 生命周期 | ISS-011 已冻结安全 Fallback 为 SUCCESS,不新增 DEGRADED;只有无法生成安全响应的未处理失败为 FAILED | 已确认 | proposal |
|
||||||
|
| Q7 验收节奏 | “端到端只在最后阶段全部完成后才验证;每个阶段如果有必要添加单元测试验收的话,就加” | 已确认 | proposal |
|
||||||
|
|
||||||
|
## Interface Impact
|
||||||
|
|
||||||
|
- 等级:L4。
|
||||||
|
- 原因:复杂 Chat 内部状态机正式切换;DiagnosisRun 增加数据库 JSON 契约;Trace run 对象新增字段;旧固定顺序消费者/测试不再成立。
|
||||||
|
- 保持兼容:`/api/chat` request/response、sessionId/runId、Executor/Verifier/Composer 输出契约不变。
|
||||||
|
- 加法变化:只新增 `diagnosis_run.orchestration_trace` 和 `run.orchestrationTrace`。
|
||||||
|
- 回滚:revert 阶段 3 代码/Prompt/spec,数据库列可保留 nullable;不保留运行时双轨开关。
|
||||||
|
|
||||||
|
## Discover Status
|
||||||
|
|
||||||
|
- `devflow/index.md`:命中阶段 0–2 archives 和 Run/Trace 历史项目。
|
||||||
|
- Glossary:相关术语已存在且无冲突,不需更新。
|
||||||
|
- ADR:阶段 0 已记录难以逆转的状态机/Trace 决策,本阶段没有新的三条件 ADR。
|
||||||
|
- 未解决问题:0。
|
||||||
|
- Draft 产物:当前只创建 proposal + decisions;design/specs/tasks 留到 Commit checkpoint。
|
||||||
|
|
||||||
|
## Grill-with-docs Review
|
||||||
|
|
||||||
|
### Domain model stress test
|
||||||
|
|
||||||
|
- 同一 session 连续两个复杂 Chat run:每次编译/执行使用自己的 runId threadId 和 metadata;orchestration trace 只写各自 DiagnosisRun,session 投影不复制,符合 Session/Run/Trace 领域边界。
|
||||||
|
- Gatekeeper REJECT 或 Planner/Executor 技术失败:Graph 进入确定性 Fallback,返回安全非空答案;Run 为 SUCCESS,trace `degraded=true`,不会把诊断质量塞进 Run status。
|
||||||
|
- Composer 成功但 verdict=LOW_CONFID/REJECT:Run 仍为 SUCCESS;self-evaluation 保存 effective verdict,orchestration trace 保存实际路径,两者职责不混合。
|
||||||
|
- Graph 未处理异常:使用流式 NodeOutput 捕获最后一个真实 state,只有其中已有 events 时才 best-effort 构造 partial trace;没有 event 时不伪造 transition,Run 标记 FAILED。
|
||||||
|
- 历史/AI_OPS run:新增列 nullable,Trace DTO 对旧 run 可返回 null;不增加历史回填或跨 flow 假数据。
|
||||||
|
|
||||||
|
### Design tree conclusions
|
||||||
|
|
||||||
|
- 生产切换采用单轨,不增加 feature flag 或保留 Sequential/Graph 双运行;回滚依靠 Git revert,nullable 列可保留。
|
||||||
|
- ChatService 负责 Run 生命周期和调用一个专用 Graph runtime/orchestrator;复杂 Agent 构造从旧 Sequential 私有流程中搬移或封装复用,不复制第二套 builder。
|
||||||
|
- Graph 执行优先使用可观察的 `CompiledGraph.stream(..., config)` 收集最后真实 state,以支持异常时 best-effort trace;正常终止仍以 END state 的 `final_answer` 为唯一答案来源。
|
||||||
|
- Graph result mapping 形成独立组件:final answer、Verifier evaluation、Composer audit、Gatekeeper audit、trace JSON 均从显式 final state读取,不再访问 VerifierContextHolder。
|
||||||
|
- Verifier Prompt 必须更新为 `diagnosis_context`、`verified_executor_output`、`verified_evidence`、`gatekeeper_audit`、`verdict_ceiling` 和可选 retry context;删除 raw Executor/full tool trace/Hook Gatekeeper 说明。
|
||||||
|
- 阶段 3 必须有 production cutover 与 Trace DTO/Service tests;阶段 4 再做测试体系全面改名、夹具收敛和旧测试删除。
|
||||||
|
|
||||||
|
### Documentation result
|
||||||
|
|
||||||
|
- 术语与 `devflow/glossary/CONTEXT.md` 完全一致,无需修改 glossary。
|
||||||
|
- 状态机、Run status 和 orchestration trace 隔离均来自阶段 0 已归档 ADR/规格,不创建重复 ADR。
|
||||||
|
- proposal 已反映单轨切换、L4 接口影响、流式 partial-state 处理、Prompt 安全边界和阶段 5 E2E 延期。
|
||||||
|
- Grill question pool 全部关闭;无需要再次询问用户的产品取舍。
|
||||||
|
|
||||||
|
## Architecture Audit
|
||||||
|
|
||||||
|
### Module and caller map
|
||||||
|
|
||||||
|
`ChatController` 的 normal/SSE 两个入口都调用 `ChatService.executeChatWithStrategy`,复杂分支进入 `executeChatComplex`;对外仍只消费 `ChatResult(answer, sessionId, runId)`。阶段 3 将内部链路变为 `ChatService Run lifecycle -> complex Chat Graph runtime/Agent assembly -> real Diagnosis Nodes -> Graph result mapper -> DiagnosisRun persistence -> DiagnosisTraceService`。Trace UI/Eval 当前读取兼容 `session.selfEvaluation`,它仍由 Run self-evaluation 投影;新增 orchestration trace 只属于 `run`。AIOps、Feedback、CaseLibrary 和 Run list 继续使用 DiagnosisRun 既有字段,不消费新增 orchestration trace。
|
||||||
|
|
||||||
|
| 模块 | 所有权 | 允许的依赖/影响 |
|
||||||
|
|---|---|---|
|
||||||
|
| ChatController | `/api/chat` normal/SSE 协议 | 继续只依赖 ChatResult;无字段变化 |
|
||||||
|
| ChatService | Chat Session/Diagnosis Run 生命周期、成功/失败保存、metrics、Eval、finally cleanup | 调用一个 Graph runtime/result mapper;不再拥有条件边/重试/Gatekeeper/Composer 路由 |
|
||||||
|
| Complex Chat Graph runtime | 每请求 Agent assembly、initial state、RunnableConfig、CompiledGraph stream | 复用现有 Prompt/tool/skill/logging;不持久化跨 Run 状态 |
|
||||||
|
| Diagnosis Graph | Node status、verified material、events、final answer | 保持阶段 1/2 路由和 counter 所有权 |
|
||||||
|
| Graph result mapper | final/partial state 到安全 evaluation/trace DTO | 不读 ThreadLocal,不查询其他 run,不持久化 raw material |
|
||||||
|
| DiagnosisRun/Flyway | 当前 Run 的 orchestration JSON | 仅一个 nullable JSON 列;历史/AIOps 可为 null |
|
||||||
|
| DiagnosisTraceService/DTO | exact/latest Run 查询和 API 投影 | 只在 RunTrace 加 parsed map;session/top-level/run-list 不重复 |
|
||||||
|
|
||||||
|
### Lifecycle and coupling audit
|
||||||
|
|
||||||
|
ReactAgent 与 CompiledGraph 都按请求构造,因为 Planner/Executor system Prompt 含本次 history/knowledge map;Graph State 和 runtime last-state holder 也是 invocation-scoped,不进入 singleton 可变字段。SessionContextHolder 仍只包围当前请求并在 finally 清理;Graph Node config 以 runId threadId + metadata 绑定 AgentStep、ToolInvocation 和 Gatekeeper。success persistence 在 Eval 之前完成,EvaluationService 再按 runId 合并 rule channel;handled Fallback 与 Composer 使用同一 SUCCESS 路径。Trace JSON 与 self-evaluation 分栏,避免把路线、质量和完整 Agent/tool 明细耦合到一个容器。
|
||||||
|
|
||||||
|
### Consumer impact audit
|
||||||
|
|
||||||
|
- ChatController normal/SSE:ChatResult 协议不变,L4 内部切换不要求调用方迁移。
|
||||||
|
- Trace UI/demo:继续读取兼容 `session.selfEvaluation`;新的 `run.orchestrationTrace` 是加法字段,最终脚本断言留到阶段 5。
|
||||||
|
- DiagnosisTraceEvaluator:现有 fixtures 不变;工具覆盖仍可从 `toolInvocations` 读取,兼容 `executor_structured_output` 保存 verified projection。pre-verification Fallback 质量未来应以 trace degraded 而非伪 verdict 判断,属于阶段 4 测试体系/阶段 5文档收尾。
|
||||||
|
- AIOps/Feedback/CaseLibrary/Run list:实体新增 nullable 字段不改变 builder call sites或查询语义。
|
||||||
|
- 数据库:Hibernate validate 要求 V012 与 entity 同批;回滚代码可忽略保留列。
|
||||||
|
|
||||||
|
### Cross-artifact alignment
|
||||||
|
|
||||||
|
| 对齐链 | 结果 | 证据 |
|
||||||
|
|---|---|---|
|
||||||
|
| brief/proposal 目标、范围、非目标 → proposal | 已对齐 | 单轨 cutover、Run/Trace、Prompt、阶段边界与 E2E 延期均明确 |
|
||||||
|
| proposal 承诺与约束 → design | 已对齐 | 10 项决策覆盖 runtime、state、config、partial state、result、persistence、DB/API/Prompt/tests |
|
||||||
|
| design 架构/接口结论 → specs/tasks | 已对齐 | L4、单轨、Run status、verified-only、Trace 唯一投影和 schema whitelist 均有 requirement/task |
|
||||||
|
| specs 可观察行为 → tasks 可执行切片 | 已对齐 | 6 组 26 个切片覆盖数据、runtime、Prompt、cutover、tests、handoff |
|
||||||
|
|
||||||
|
### Audit result
|
||||||
|
|
||||||
|
未发现与阶段 0–2、Run/Trace ADR 或 glossary 冲突。审计确认不能把 Agent factory 复制进 Node,也不能把 pre-verification Fallback 伪装为 Verifier verdict;两项已在 design/specs/tasks 固定。唯一跨阶段依赖是阶段 4 让测试/Eval 以 orchestration degraded 理解无 Verifier verdict 的安全 Fallback,已记录但不阻塞阶段 3 production correctness。架构风险可接受,cross-artifact gap=0,无未解决接口消费者。
|
||||||
|
|
||||||
|
## Commit Gate
|
||||||
|
|
||||||
|
- schema:spec-driven;proposal/design/specs/tasks 全部 done,applyRequires=`tasks` 已满足。
|
||||||
|
- OpenSpec:当前 change strict validation 通过;14 个主 specs 全部 strict pass。
|
||||||
|
- 规格结构:5 个 delta capabilities、23 条 requirements、72 个 scenarios;tasks 26 个可执行 checkbox。
|
||||||
|
- Cross-artifact:4/4 已对齐,gap=0。
|
||||||
|
- Interface impact:L4;design 已独立记录消费者、加法 DB/Trace 协议、单轨迁移和 Git revert 回滚。
|
||||||
|
- Question pool:所有 evidence-driven 已查证;所有 user-interview 已由 ISS-011/用户原话确认;无未决项。
|
||||||
|
- Preflight:`git diff --check` 通过;Commit checkpoint 尚未修改 Java、SQL、Prompt 或测试。
|
||||||
|
- 结论:Draft OpenSpec 已达到可执行状态,创建 `.committed`,阶段 3 Apply 只能以这些产物为依据。
|
||||||
|
|
||||||
|
## Apply Authorization
|
||||||
|
|
||||||
|
- 用户原话:“直接实现吧,不用找我授权了”。
|
||||||
|
- 本阶段在 Commit gate 后直接进入 Apply;不扩大到阶段 4 全面测试迁移或阶段 5 live E2E。
|
||||||
|
|
||||||
|
## Pre-apply Research
|
||||||
|
|
||||||
|
### Reference implementations read
|
||||||
|
|
||||||
|
- `ChatService.executeChatComplex`、四个 complex Agent builders、Run start/save、metrics、Prompt audit、self-evaluation merge:迁移源实现和外部兼容基线。
|
||||||
|
- `DiagnosisGraphFactory`、`DiagnosisRealGraphActionsFactory`、四个 Agent adapters、Gatekeeper/VerifiedInput/Fallback、`DiagnosisOrchestrationTraceBuilder`:Graph 路由/计数/安全材料唯一真理源。
|
||||||
|
- `DiagnosisRun`、V011 migration、`DiagnosisTraceResponse`、`DiagnosisTraceService`、`DiagnosisTraceServiceTest`:Run JSON 字段和 exact/latest Trace 映射标准。
|
||||||
|
- `ChatController` normal/SSE、`DiagnosisTraceController`:公开协议调用者与响应包装边界。
|
||||||
|
- `SelfEvaluationMergeService`、`EvaluationService`、`DiagnosisTraceEvaluator`:evaluation container、异步 rule merge 和兼容消费者。
|
||||||
|
- 本地 graph-core 1.1.2.0 `javap`:CompiledGraph `stream/invoke/state`、NodeOutput `state()`、RunnableConfig `threadId/metadata` 的真实 API。
|
||||||
|
- `chat-*-prompt.md`:现有 Prompt 组装与 Verifier 旧 Hook payload gap。
|
||||||
|
|
||||||
|
### Technology inventory
|
||||||
|
|
||||||
|
| 类别 | 项目标准 / 本阶段使用 |
|
||||||
|
|---|---|
|
||||||
|
| Graph execution | `CompiledGraph.stream(initial, config)` + NodeOutput.state,当前请求线程阻塞消费 |
|
||||||
|
| Run context | SessionContextHolder + RunnableConfig threadId/metadata;runId 是唯一执行边界 |
|
||||||
|
| Agent assembly | ReactAgent builder、AgentLoggingHook、PlannerSkillMetadataHook/SkillsAgentHook、现有 tools/callbacks |
|
||||||
|
| JSON persistence | Jackson map serialization;实体 String + `@JdbcTypeCode(SqlTypes.JSON)`;Flyway JSON column |
|
||||||
|
| Trace API | Lombok DTO builder + DiagnosisTraceService parsed Map;exact/latest Run 查询 |
|
||||||
|
| Evaluation | SelfEvaluationMergeService container;EvaluationService 按 runId 异步合并 rule channel |
|
||||||
|
| Tests | JUnit 5,通过 public service/runtime/Trace API;只 mock repository/model/tool 外部边界 |
|
||||||
|
| MQ/Consumer | 不涉及 |
|
||||||
|
|
||||||
|
### New infrastructure and reuse
|
||||||
|
|
||||||
|
- 新增一个深接口的 complex Chat Graph runtime/result mapper;复用既有 Graph/Node/Agent builders,不增加新依赖或第二套路由。
|
||||||
|
- 新增 V012、DiagnosisRun 字段和 RunTrace parsed map;不新增表、repository method 或历史 backfill。
|
||||||
|
- TDD tracer bullet 从 Trace DTO/Service 的唯一 Run 投影开始,再进入 runtime final-state mapping,最后切换 ChatService。
|
||||||
|
- 研究未发现 devflow/OpenSpec 冲突,技术清单足以开始实现。
|
||||||
|
|
||||||
|
## Apply Progress
|
||||||
|
|
||||||
|
- TDD Trace slice RED:DiagnosisTraceServiceTest 明确缺少 DiagnosisRun builder 字段和 RunTrace getter。
|
||||||
|
- GREEN:V012、DiagnosisRun JSON 字段、RunTrace parsed map、DiagnosisTraceService mapping 完成;10 个 Trace service tests 通过。
|
||||||
|
- 失败分类:首次 GREEN 运行仅测试夹具用裸 ObjectMapper 无法序列化 LocalDateTime,属于 test harness 偏差;改为项目可用的 `findAndRegisterModules()` 后原始 focused loop 通过,未改业务协议。
|
||||||
|
- Trace API JSON 断言证明顶层/session 无重复字段、run 为解析对象、无 raw 字段;null/invalid JSON fail closed。
|
||||||
|
|
||||||
|
## Apply Completion
|
||||||
|
|
||||||
|
- 复杂 Chat 已单轨切换到 `ChatDiagnosisGraphRuntime`;生产 `ChatService` 不再创建 `SequentialAgent`,不再读取 `VerifierContextHolder`,Graph Verifier 只保留 `AgentLoggingHook`。
|
||||||
|
- runtime 使用 query-only initial state、`threadId=runId` 和 sessionId/runId metadata,流式保留最后真实 state;异常只保存真实 partial state/trace。
|
||||||
|
- `DiagnosisGraphResultMapper` 只从 verified projection 构造兼容 self-evaluation;Verifier 未完成时不伪造 verdict,不持久化 full tool trace/raw Executor。
|
||||||
|
- Verifier Prompt 与 runtime payload 已统一为 verified-only,Prompt audit 更新为 `chat-prompts-v2` / `chat-verifier-v3`。
|
||||||
|
- 旧 `ChatServiceSequentialAgentTest` 因绑定已删除的固定顺序/ThreadLocal/score retry 实现而移除,由 public service/runtime/route tests 替代。
|
||||||
|
- V012 schema 白名单测试确认只新增 `diagnosis_run.orchestration_trace JSON NULL`,无其他 schema object。
|
||||||
|
|
||||||
|
## Verification Summary
|
||||||
|
|
||||||
|
- focused regression:27 suites / 119 tests,0 failures、0 errors、0 skipped。
|
||||||
|
- Maven test compilation:通过。
|
||||||
|
- OpenSpec:当前 change strict pass;主 specs 14/14 strict pass。
|
||||||
|
- 静态门禁:`git diff --check` 通过;cutover 禁用引用 0;核心 TODO/placeholder 0;schema whitelist 为 1 ALTER / 1 ADD COLUMN / 0 CREATE / 0 DROP。
|
||||||
|
- 实现期唯一失败分类:no-answer service test 曾错误期待 inner cause 文本,属于测试断言偏差;按公开 wrapper error + FAILED/partial trace 契约修正后通过,OpenSpec 与生产代码无需变更。
|
||||||
|
- 按阶段门禁未运行 Maven live E2E、未检查 `logs/`、未执行 `scripts/query_mysql.py`;统一保留到阶段 5。
|
||||||
|
|
||||||
|
## Archive Result
|
||||||
|
|
||||||
|
- 5 份 delta specs 已同步:新增 9、修改 13、删除 1 条 requirements。
|
||||||
|
- OpenSpec 已归档到 `openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-chatservice-cutover/`。
|
||||||
|
- CLI 对 proposal 结构给出非阻塞建议(大 change/delta 拆分与 SHALL/scenario 启发式);tasks 26/26、delta specs 和 strict validation 均通过,未形成验收阻塞。
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
# Chat Diagnosis StateGraph ChatService Cutover Evidence
|
||||||
|
|
||||||
|
## 证据
|
||||||
|
|
||||||
|
| 来源 | 证据 | 结论 | 是否已汇报 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `ChatService` + source isolation search | 复杂路径只有一次 Graph runtime 调用,无 SequentialAgent、VerifierContextHolder、VerifierInputHook 或 score retry | 单轨 cutover 完成,简单 Chat 路径未改 | 是 |
|
||||||
|
| `ChatDiagnosisGraphRuntimeTest` / `DiagnosisRealGraphIntegrationTest` | query-only state、runId thread/metadata、Composer/Fallback、empty stream、blank answer、partial trace、Gatekeeper once/twice | Graph identity、路由和真实 partial-state 边界可观察 | 是 |
|
||||||
|
| `DiagnosisGraphResultMapperTest` / `ChatVerifierPromptContractTest` | verified projection、effective verdict 兼容、pre-verification 无伪 verdict、禁用 raw/full trace 字段 | self-evaluation 与 Prompt 安全边界一致 | 是 |
|
||||||
|
| `ChatServiceGraphIntegrationTest` | ChatResult、agent_flow、SUCCESS/Fallback/FAILED、metrics/Eval、partial trace、多 Run 隔离 | 生产生命周期与兼容返回满足阶段 3 规格 | 是 |
|
||||||
|
| `DiagnosisTraceServiceTest` | `run.orchestrationTrace` 解析对象、top/session 无重复、invalid/null fail closed | Trace 字段所有权保持 Run 隔离 | 是 |
|
||||||
|
| `DiagnosisRunSchemaContractTest` | V012 executable SQL 与唯一允许语句精确相等 | schema 仅增加一个 nullable JSON 列 | 是 |
|
||||||
|
| focused Maven suite | 27 suites / 119 tests,0 failures/errors/skipped | Graph、Controller、Repository、Gatekeeper、Composer、Eval 回归通过 | 是 |
|
||||||
|
| OpenSpec/static gates | change strict、14/14 主 specs、test compile、diff check、禁用引用/占位/schema whitelist 全通过 | 规格、编译和源级隔离闭环 | 是 |
|
||||||
|
|
||||||
|
## Evidence-driven 结论
|
||||||
|
|
||||||
|
- orchestration trace 必须独立于 self-evaluation 和详细 Agent/tool Trace;V012 与 RunTrace 的实现保持这一边界。
|
||||||
|
- handled Fallback 表示安全降级答案,Run 仍为 SUCCESS;未处理、空答案或 invariant 失败才是 FAILED。
|
||||||
|
- Verifier/Gatekeeper 只以显式 Graph state 传递可信材料;继续使用 Hook/ThreadLocal 会形成双 Gatekeeper 和不安全输入,因此生产路径已彻底移除该依赖。
|
||||||
|
- 旧 Sequential 测试验证的是已废止内部机制,替换为 public service + runtime + route contract tests 才能保持真实回归价值。
|
||||||
@@ -0,0 +1,49 @@
|
|||||||
|
# Acceptance
|
||||||
|
|
||||||
|
## 静态验证
|
||||||
|
|
||||||
|
- 旧闭包路径不存在,executable legacy refs=0。
|
||||||
|
- current-doc stale architecture refs=0;历史兼容引用有明确限定。
|
||||||
|
- `git diff --check` 通过;无新增 migration/schema 变更;临时调试标记=0。
|
||||||
|
- 当前 change strict 和 16 个 main specs strict 全部通过。
|
||||||
|
|
||||||
|
## 脚本验证
|
||||||
|
|
||||||
|
- `mvn -q -DskipTests test-compile`:通过。
|
||||||
|
- 39-suite authoritative/focused Maven command:157 tests,0 failure/error/skipped。
|
||||||
|
- 归档前 `mvn clean` + test compilation + 43-suite deterministic command:189 tests,0 failure/error/skipped。
|
||||||
|
- `DiagnosisTraceEvaluatorTest` + `DiagnosisEvalBaselineDiffTest`:12/12 baseline,same diff=0。
|
||||||
|
- PowerShell parser + `InterviewDemoScriptContractTest`:通过。
|
||||||
|
- `run-interview-demo-check.ps1 -SessionId iss-011-stage5-20260720015557 -OutputDir target/iss-011-stage5-output-current`:exit 0。
|
||||||
|
- `scripts/query_mysql.py` exact queries:V012、Run JSON、AgentStep/ToolInvocation ownership 全部通过。
|
||||||
|
- `openspec validate --all --strict --no-interactive`:17/17 passed;`git diff --check`、current-doc、schema、debug source、port/temp scope checks 通过。
|
||||||
|
|
||||||
|
## Live E2E
|
||||||
|
|
||||||
|
| 项目 | 结果 |
|
||||||
|
|---|---|
|
||||||
|
| Maven profile | `mvp-demo` |
|
||||||
|
| sessionId | `iss-011-stage5-20260720015557` |
|
||||||
|
| runId | `run-808ac38f-3ad0-4462-a6d0-ed50d8686473` |
|
||||||
|
| Run | `CHAT/SUCCESS` |
|
||||||
|
| answer / metrics | 109 chars / 75964ms / 111802 tokens / 8 steps / 12 tools |
|
||||||
|
| Graph | `stategraph-v1`, `fallback`, `fallback_completed`, degraded=true, 3 transitions, retry=0 |
|
||||||
|
| evaluation / feedback | non-empty / useful |
|
||||||
|
| new ERROR | 0 |
|
||||||
|
| DB ownership | wrong owner=0,wrong-session rows=0 |
|
||||||
|
| process cleanup | owned PIDs stopped,9900 released |
|
||||||
|
|
||||||
|
## 浏览器/人工验证
|
||||||
|
|
||||||
|
- 不适用。本阶段验收入口是 API/PowerShell executable contract,无 UI 改动。
|
||||||
|
|
||||||
|
## 未验证
|
||||||
|
|
||||||
|
- 无 OpenSpec 必需项未验证。
|
||||||
|
|
||||||
|
## 归档状态
|
||||||
|
|
||||||
|
- ISS-011 已归档至 `mvp/issues/archived/ISS-011-chat-diagnosis-stategraph-orchestration.md`。
|
||||||
|
- OpenSpec 归档路径:`openspec/changes/archive/2026-07-20-chat-diagnosis-stategraph-cleanup-docs`。
|
||||||
|
- OpenSpec CLI 已同步主 specs:新增 `chat-diagnosis-stategraph-cleanup-docs`,更新 `mvp-demo-trace-acceptance` 3 项 requirement。
|
||||||
|
- 不 push。
|
||||||
@@ -0,0 +1,32 @@
|
|||||||
|
# Chat Diagnosis StateGraph Cleanup And Final Acceptance
|
||||||
|
|
||||||
|
## 背景
|
||||||
|
|
||||||
|
ISS-011 阶段 0-4 已完成 StateGraph 设计冻结、路由骨架、真实 Nodes、ChatService 单轨切换和三层权威测试。阶段 5 负责删除旧 Hook/ThreadLocal/full-trace service 闭包、对齐当前文档与 demo executable contract,并以唯一 Run 完成最终 Maven、日志和数据库验收。
|
||||||
|
|
||||||
|
## 目标
|
||||||
|
|
||||||
|
- 只保留 bounded StateGraph 复杂 Chat 编排和 verified-only Verifier 输入。
|
||||||
|
- 让 current architecture/eval/demo 文档与 Run-owned orchestration trace 一致。
|
||||||
|
- 先通过确定性回归,再用 Maven `mvp-demo` 证明 exact Chat/Trace/feedback、日志和数据库 ownership。
|
||||||
|
- 所有门禁通过后关闭 ISS-011,并归档阶段 5 OpenSpec。
|
||||||
|
|
||||||
|
## 范围
|
||||||
|
|
||||||
|
- 删除 `VerifierInputHook`、`VerifierContextHolder`、`ToolTraceSummaryService` 及其 focused test。
|
||||||
|
- 更新 current architecture/eval/demo 文档和 interview demo check。
|
||||||
|
- 修复 live 暴露的 Graph event classloader 边界与 nested ReactAgent resume config 问题。
|
||||||
|
- 完成 Graph/Chat/Trace/Eval 回归、Maven live E2E、日志/MySQL 核验、进程清理和 Issue 生命周期收口。
|
||||||
|
|
||||||
|
## 非目标
|
||||||
|
|
||||||
|
- 不修改公开 Chat/feedback API、Executor/Verifier/Composer 业务协议或数据库 schema。
|
||||||
|
- 不重写 archived issues、历史 design notes 和 legacy fixtures。
|
||||||
|
- 不删除旧 Trace/fixture 对 `tool_trace_summary` 的只读兼容。
|
||||||
|
- 不 push,不删除失败尝试的审计数据。
|
||||||
|
|
||||||
|
## 元数据
|
||||||
|
|
||||||
|
- 分档:complex
|
||||||
|
- OpenSpec:`chat-diagnosis-stategraph-cleanup-docs`
|
||||||
|
- 接口影响:L2 内部类型/状态表示修复;外部 API/DTO/schema 不变
|
||||||
@@ -0,0 +1,159 @@
|
|||||||
|
# Chat Diagnosis StateGraph Cleanup, Final Acceptance And Documentation Decisions
|
||||||
|
|
||||||
|
## Entry Summary
|
||||||
|
|
||||||
|
- 问题:ISS-011 运行时已切换且测试体系已收敛,但旧 Hook/ThreadLocal/service死代码、当前架构文档和最终 live 证据尚未闭环。
|
||||||
|
- 期望:阶段 5完成清理、文档、自动化回归、eval、Maven E2E、日志/DB 验收、Issue 归档和独立提交。
|
||||||
|
- 分档:complex;接口影响 L2 内部删除 + 文档/demo 验收增强,外部 API/DB 协议不变。
|
||||||
|
- Change:`chat-diagnosis-stategraph-cleanup-docs`。
|
||||||
|
- 授权:用户已明确要求直接实现;本阶段按此前规则执行唯一最终 E2E。
|
||||||
|
|
||||||
|
## Context Sources
|
||||||
|
|
||||||
|
- ISS-011 阶段 5、测试策略、协议影响、验收标准和冻结决策。
|
||||||
|
- 阶段 0–4 OpenSpec archives、devflow acceptance 与提交 `581daff`、`42ba204`、`1460dd1`、`99e490f`、`208a231`。
|
||||||
|
- 全仓 `VerifierInputHook`/`VerifierContextHolder`/`ToolTraceSummaryService` 定义与引用搜索。
|
||||||
|
- `mvp/architecture/README.md` 列出的 current docs、`mvp/eval/README.md`、`mvp/demo/` scripts/checklist。
|
||||||
|
- `mvp-demo` profile、payment-timeout request、`scripts/query_mysql.py` 和 logs/ 现有布局。
|
||||||
|
|
||||||
|
## Question Pool
|
||||||
|
|
||||||
|
| # | 维度 | 问题 | 模式 | 状态 |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| Q1 | 清理 | 旧 Hook/ThreadLocal/trace summary service 是否还有生产消费者? | evidence-driven | 已解决 |
|
||||||
|
| Q2 | 文档 | 哪些旧引用应更新,哪些历史材料应保留? | evidence-driven | 已解决 |
|
||||||
|
| Q3 | E2E | 最终 live 场景如何绑定唯一 session/run 并证明 Graph 路径? | evidence-driven | 已解决 |
|
||||||
|
| Q4 | 日志/DB | 如何避免用旧日志/latest DB 记录冒充当前证据? | evidence-driven | 已解决 |
|
||||||
|
| Q5 | 验收 | 何时允许启动 Maven、是否需要日志和 DB 查询? | user-interview(用户最新规则) | 已确认 |
|
||||||
|
| Q6 | 关闭 | 何时把 ISS-011 从 active 移到 archived? | evidence-driven | 已解决 |
|
||||||
|
|
||||||
|
## Evidence-driven Findings
|
||||||
|
|
||||||
|
- Q1:旧闭包只有 `VerifierInputHook -> VerifierContextHolder + ToolTraceSummaryService`,以及 `ToolTraceSummaryServiceTest`;ChatService/Graph/Trace/Eval 均无引用,可整体删除。
|
||||||
|
- 实现前规格校正:`ChatVerifierPromptContractTest` 和 `DiagnosisGraphTestSuiteStructureTest` 必须保留旧类型名称的负向字符串断言;这不构成 executable reference。OpenSpec 已收紧为无定义/import/实例化/type-use,允许负向 guard literal。
|
||||||
|
- Q2:current architecture index 仍列 `agent-orchestration.md` 等为当前真理源,因此必须更新;`mvp/issues/design-notes`、archived issues、历史 eval fixtures 保留时间点/兼容语义,不做大规模重写。
|
||||||
|
- Q3:`run-interview-demo-check.ps1` 已用 Chat response runId 查询 exact Trace/feedback,最适合扩展 `run.orchestrationTrace` fail-fast 和 summary,不另建重复脚本。
|
||||||
|
- Q4:E2E 使用唯一 timestamp sessionId;日志记录启动前 byte/time 边界并按 session/run 搜索;DB 所有核心查询带 exact sessionId/runId,另查询错误 ownership count。
|
||||||
|
- Q6:只有实现、回归、eval、live E2E、日志、DB 和 OpenSpec门禁全部通过后,Issue checkbox 才可完成并移动到 archived。
|
||||||
|
|
||||||
|
## User-interview Confirmation
|
||||||
|
|
||||||
|
| 问题 | 用户原话 | 状态 | OpenSpec 回写 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| Q5 最终验收节奏 | “端到端只在最后阶段全部完成后才验证……日志在log文件夹,项目库有查询数据库的py工具” | 已确认 | proposal |
|
||||||
|
|
||||||
|
## Grill-with-docs Result
|
||||||
|
|
||||||
|
- Session/Run/Trace 术语保持不变;新增强调 `orchestration_trace` 是 Run 路由摘要,不属于 self-evaluation 或日志。
|
||||||
|
- StateGraph、Workflow/Node Contract/Chat Integration 属于实现/测试架构术语,不修改业务 glossary。
|
||||||
|
- 当前文档必须使用 explicit Gatekeeper Node、verified-only Verifier 和 bounded evidence retry;历史设计笔记仍可描述当时 Hook 架构。
|
||||||
|
- 删除旧闭包是阶段 0 已冻结单轨迁移的自然收尾,不形成新的难逆转权衡,无需 ADR。
|
||||||
|
|
||||||
|
## Discover Status
|
||||||
|
|
||||||
|
- `devflow/index.md`:命中阶段 0–4 archives。
|
||||||
|
- 接口影响:L2 内部类型删除;外部 API/DTO/DB/Prompt/状态语义无变化。
|
||||||
|
- E2E 入口/脚本/日志/DB 工具已定位;真实执行留到 Apply 最后。
|
||||||
|
- 未解决问题:0。
|
||||||
|
- Draft 产物:proposal + decisions;尚未生成 design/spec/tasks,尚未删除代码或启动应用。
|
||||||
|
|
||||||
|
## Architecture Audit
|
||||||
|
|
||||||
|
### Module and evidence map
|
||||||
|
|
||||||
|
`ChatController -> ChatService -> ChatDiagnosisGraphRuntime -> DiagnosisRealGraphActionsFactory -> explicit Nodes -> DiagnosisGraphResultMapper -> DiagnosisRun/Trace` 是唯一当前 Chat链。旧 `VerifierInputHook -> ToolTraceSummaryService/VerifierContextHolder` 已从主链断开,删除不改变输入/输出或持久化。阶段 5新增的 demo script assertion只消费 exact Trace `run.orchestrationTrace`,DB/log检查是验收消费者,不成为运行时业务依赖。
|
||||||
|
|
||||||
|
| 模块 | 所有权 | 阶段 5动作 |
|
||||||
|
|---|---|---|
|
||||||
|
| Graph/Chat runtime | 路由、Node、Run 生命周期 | 不改行为,仅回归 |
|
||||||
|
| Legacy Hook closure | 旧 Sequential Verifier payload | 整体删除 |
|
||||||
|
| Current architecture docs | 当前实现真理源 | 更新 StateGraph/verified-only/trace |
|
||||||
|
| Historical docs/fixtures | 时间点/兼容记录 | 保留,不冒充当前实现 |
|
||||||
|
| Demo check | live Chat/Trace/feedback executable contract | 增加 exact orchestration trace fail-fast |
|
||||||
|
| logs/MySQL | live运行证据 | 只读本次 session/run |
|
||||||
|
| ISS/OpenSpec/devflow | 生命周期与交接 | 所有门禁通过后归档 |
|
||||||
|
|
||||||
|
### Lifecycle and failure ownership
|
||||||
|
|
||||||
|
- 自动化门禁失败:不启动 live Maven,修复代码/测试/规格后重跑。
|
||||||
|
- live startup失败:应用未 ready,不执行 demo/DB成功声明,先读启动输出和新日志诊断。
|
||||||
|
- Chat/Trace/feedback失败:保留 exact response/run证据,ISS保持 active。
|
||||||
|
- log ERROR:逐条分类;未解释 ERROR阻塞验收。
|
||||||
|
- DB不一致:以 exact run为准,不能用 API成功掩盖 persistence偏差。
|
||||||
|
- finally:无论成功失败都停止本轮进程并确认端口,不扩大到未知已有进程。
|
||||||
|
|
||||||
|
### Consumer and compatibility audit
|
||||||
|
|
||||||
|
- 外部 API/DTO/DB consumer无迁移;demo summary仅加字段。
|
||||||
|
- Trace UI/eval 对历史 `tool_trace_summary` 的读取保留,旧 fixture不批量迁移。
|
||||||
|
- current docs消费者将看到新 StateGraph架构;历史链接仍可追溯 old Hook设计。
|
||||||
|
- Issue move只改变文档位置/index,代码/运行时不依赖该路径。
|
||||||
|
|
||||||
|
### Cross-artifact alignment
|
||||||
|
|
||||||
|
| 上游 → 下游 | 检查内容 | 状态 |
|
||||||
|
|---|---|---|
|
||||||
|
| brief/proposal → proposal | cleanup、current docs、demo、regression、live/log/DB、Issue closure | 已对齐 |
|
||||||
|
| proposal → design | 删除闭包、current/history边界、顺序、identity、日志/DB、cleanup | 已对齐 |
|
||||||
|
| design → specs/tasks | 负向 literal例外、E2E字段、exact evidence、进程清理、Issue gate | 已对齐 |
|
||||||
|
| specs → tasks | 每条 requirement有可执行 cleanup/docs/test/live/log/DB/closure slice | 已对齐 |
|
||||||
|
|
||||||
|
### Audit result
|
||||||
|
|
||||||
|
审计确认阶段 5不需要新运行时抽象或 DB migration;主要风险来自外部 live状态和证据归属,已通过 unique session/run、log boundary、exact DB queries和process ownership缓解。规格误把负向名称 literal 当 executable reference 的 gap 已修正。接口影响 L2,cross-artifact gap=0,无新 ADR。
|
||||||
|
|
||||||
|
## Commit Gate
|
||||||
|
|
||||||
|
- schema:spec-driven;proposal/design/2 delta specs/tasks 全部 done,applyRequires=`tasks` 已满足。
|
||||||
|
- OpenSpec:当前 change strict pass;16 个主 specs strict pass。
|
||||||
|
- Cross-artifact:4/4 已对齐,gap=0;负向 guard literal例外已写入 proposal/design/spec/tasks。
|
||||||
|
- Question pool:5 个 evidence-driven 已解决,1 个 user-interview 已由用户原话确认,无未决项。
|
||||||
|
- Interface impact:L2 internal type removal + demo/docs enhancement;外部协议/DB无变化。
|
||||||
|
- Preflight:`git diff --check` 通过;尚未删除代码、修改 current docs/script或启动应用。
|
||||||
|
- 结论:Draft OpenSpec 达到可执行状态,创建 `.committed` 后进入 Apply。
|
||||||
|
|
||||||
|
## Apply Progress
|
||||||
|
|
||||||
|
### Legacy closure removal
|
||||||
|
|
||||||
|
- 已删除 `VerifierInputHook`、`VerifierContextHolder`、`ToolTraceSummaryService` 和 `ToolTraceSummaryServiceTest`,四个路径均不存在。
|
||||||
|
- `rg` 对 `src/main`、`src/test` 的旧类型扫描仅命中 `ChatVerifierPromptContractTest` 和 `DiagnosisGraphTestSuiteStructureTest` 中的负向守卫字符串;无定义、import、实例化、继承或类型依赖。
|
||||||
|
- 删除后 focused 回归覆盖 Executor parser、Gatekeeper service/node、VerifiedInput、Verifier、Composer、Fallback、Workflow、Node Contract、Chat integration、Trace、result mapper 和结构契约:14 suites / 82 tests,0 failure、0 error、0 skipped。
|
||||||
|
- `mvn -q -DskipTests test-compile` 通过;Graph/shared protocol 真理源保留,Spring 当前链路所需类型可完整编译。
|
||||||
|
|
||||||
|
### Current docs and demo contract
|
||||||
|
|
||||||
|
- architecture index、编排、session/trace、current MVP、evidence pipeline、quality gates、feedback、retrieval 和 eval 文档已切换为 bounded StateGraph、显式 Gatekeeper/Verified Input、verified-only Verifier、有限重试/Fallback 和 Run-owned `orchestration_trace`。
|
||||||
|
- current-doc scan 对 `SequentialAgent`、旧 Hook/ThreadLocal/service 及旧测试类名为 0 命中;`tool_trace_summary` 仅剩 4 处,均明确标注为旧 Run/fixture 只读兼容,不是当前 Verifier 输入。
|
||||||
|
- interview demo check 绑定 Chat 返回的 exact runId,校验 Chat/Trace ownership、Run CHAT/SUCCESS、Agent/tool/self-evaluation、Graph trace 六个字段和 feedback success;summary 新增 orchestration version、final node、termination reason、degraded、transition count 和 evidence retry count。
|
||||||
|
- PowerShell parser 语法检查通过;`InterviewDemoScriptContractTest` 2 tests 通过,覆盖 exact runId URL/response、orchestration fail-fast 和 summary 字段。
|
||||||
|
|
||||||
|
### Final deterministic gates
|
||||||
|
|
||||||
|
- authoritative/focused regression:39 suites / 157 tests,0 failure、0 error、0 skipped;覆盖三层 Graph、全部 Graph Node/router/trace builder、Chat/Trace/Gatekeeper/Composer、Controller、Repository、schema、feedback/tool recorder 和 demo contract。
|
||||||
|
- fixed diagnosis eval:12/12 passed,verdict distribution 为 PASS=5、LOW_CONFID=6、REJECT=1;same-baseline diff 无 regression、0 items。
|
||||||
|
- `mvn -q -DskipTests test-compile` 通过;当前 change strict 通过,16 个主 specs strict 全部通过。
|
||||||
|
- `git diff --check`、legacy executable refs、current-doc stale refs 和 unexpected schema change 检查全部通过。
|
||||||
|
- focused 回归日志中的 Graph ERROR/exception stack trace 来自 `ChatServiceGraphIntegrationTest` 对 FAILED/no-answer/unhandled failure 的显式契约用例,Maven exit 0,不是未解释的 live ERROR。
|
||||||
|
|
||||||
|
### Live failure diagnosis and correction
|
||||||
|
|
||||||
|
- 首次 live identity:sessionId=`iss-011-stage5-20260717140450`,runId=`run-6db680f8-f764-49d1-995f-0e55a4b05a06`。demo contract 在 exact Trace `run.orchestrationTrace=null` 处 fail-fast,未提交 feedback;Run 为 `CHAT/FAILED`,步骤/工具均为 0。
|
||||||
|
- 新日志根因:`DiagnosisOrchestrationTraceBuilder` 收到类名相同但 classloader identity 不同的 `OrchestrationEvent`,`instanceof` 失败并抛出 `orchestration events contain unsupported value`。这是 Spring Boot DevTools live classloader 才暴露的 Graph state 表示缺陷,单元 JVM 未复现。
|
||||||
|
- 冲突分类:代码偏离/运行时兼容 bug,OpenSpec 对 non-empty orchestration trace 和 live Maven startup 的要求正确,不修改验收口径。
|
||||||
|
- RED:新增 portable event map builder 回归,修复前 1 test error;GREEN:Graph state 的 production/test actions 改存 classloader-neutral Map,builder兼容 local record/Map,Node/Workflow assertions 改读 Map。
|
||||||
|
- 修复后先运行 5-suite Graph/Node/Runtime/Chat integration focused gate,再运行完整 39 suites / 157 tests,全部 0 failure/error/skipped;首次 Maven 进程链已按 ownership 停止,9900 已释放。
|
||||||
|
|
||||||
|
- 第二个 live blocker 为外层 Graph `RunnableConfig` 的 resume metadata 被原样传给内层 ReactAgent,触发 `Resume request without a configured checkpoint saver`。回归先证明 nested config 与 outer config 同一且含 `HUMAN_FEEDBACK`,再改为保留 sessionId/runId、剔除 resume/state-update/checkpoint 控制信息的独立配置;5 suites / 50 tests 和随后完整 39 suites / 157 tests 通过,临时 `[DEBUG-ISS011-NODE]` 探针已删除且源码扫描为 0。
|
||||||
|
- 2026-07-17 的一次长请求在工具执行期间遭遇外部 MySQL 瞬时 `Connection is closed`,留下精确 `RUNNING` 失败尝试;仓库查询工具随后证明数据库恢复且 server `wait_timeout=28800`。该失败未被当作验收通过,失败 Run 保留为真实审计记录。
|
||||||
|
|
||||||
|
### Accepted live E2E
|
||||||
|
|
||||||
|
- 启动:`mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"`;启动前 9900 空闲,隐藏进程链为 cmd `25080` -> Maven Java `17860` -> app Java `10732`,readiness 后执行固定 payment-timeout demo。
|
||||||
|
- identity:sessionId=`iss-011-stage5-20260720015557`,runId=`run-808ac38f-3ad0-4462-a6d0-ed50d8686473`;Chat/Trace exact identity 一致,answer 长度 109,feedback request success。
|
||||||
|
- Trace:Run=`CHAT/SUCCESS`,8 AgentSteps、12 ToolInvocations,self-evaluation 非空;`stategraph-v1`,final node=`fallback`,termination=`fallback_completed`,degraded=`true`,3 transitions,evidence retry count=0。
|
||||||
|
- Fallback 原因是 Gatekeeper LOW_CONFID 且本轮模型输出缺少 `source_invocation_id`;这是按冻结契约执行的安全降级,answer 非空且未绕过 Gatekeeper,日志/DB 均可审计。
|
||||||
|
- 最终日志发生跨日 rollover:7 月 20 日 active `application.log`/`chat.log` 全部属于本轮;`application-error.log` 最后写入仍为 7 月 17 日。本轮 `rg " ERROR "` 对 application/chat 为 0,新增 error-file bytes 为 0;session/run、Graph 75964ms、evaluation、exact Trace 和 useful feedback 均有关联日志。
|
||||||
|
- MySQL:V012 `orchestration_trace` 为 nullable JSON;exact Run answer=109、duration=75964、token=111802、steps=8、tools=12、evaluation len=8822、trace len=438、feedback=useful;JSON 路由与 Trace 完全一致。
|
||||||
|
- ownership:AgentStep 8、ToolInvocation 12,各自 distinct session/run=1、wrong owner=0;唯一 session 下 wrong run/step/tool 均为 0。
|
||||||
|
- cleanup:只停止 PID `10732/17860/25080`,最终 9900 已释放,无剩余 owned process。
|
||||||
@@ -0,0 +1,38 @@
|
|||||||
|
# Evidence
|
||||||
|
|
||||||
|
## Source And Dependency Evidence
|
||||||
|
|
||||||
|
- 旧闭包四个文件已删除;`src/main`/`src/test` 旧类型扫描仅剩两个测试中的负向名称守卫,无 definition/import/instantiation/type dependency。
|
||||||
|
- 当前真理源保留 `ExecutorEvidenceParser`、`ExecutorGatekeeperService`、`GatekeeperNode`、`VerifiedInputNode`、`VerifierNodeAdapter` 和 `DiagnosisGraphResultMapper`。
|
||||||
|
- current docs 不再描述 SequentialAgent、Hook Gatekeeper 或 full-trace Verifier;4 处 `tool_trace_summary` 均明确为历史只读兼容。
|
||||||
|
|
||||||
|
## Deterministic Evidence
|
||||||
|
|
||||||
|
- 删除后 focused:14 suites / 82 tests,0 failure/error/skipped。
|
||||||
|
- 最终 authoritative/focused:39 suites / 157 tests,0 failure/error/skipped。
|
||||||
|
- 归档前 `mvn clean` 后重建验证:43 suites / 189 tests,0 failure/error/skipped;额外覆盖 4 个无需外部服务的现存测试类。
|
||||||
|
- diagnosis eval:12/12 passed;PASS=5、LOW_CONFID=6、REJECT=1;same-baseline diff=0。
|
||||||
|
- `mvn -q -DskipTests test-compile`、PowerShell parser、OpenSpec current strict、16 main specs strict、`git diff --check`、legacy/current-doc/schema scans 均通过。
|
||||||
|
|
||||||
|
## Diagnose Evidence
|
||||||
|
|
||||||
|
- DevTools live classloader 使 record `instanceof` 边界失效;portable event map regression 先 RED,Node state 改用 Map 且 builder 兼容 record/Map 后 GREEN。
|
||||||
|
- outer Graph resume metadata 污染 nested ReactAgent;nested config isolation regression 先 RED,保留 session/run metadata并剔除 resume/state-update/checkpoint 控制信息后 GREEN。
|
||||||
|
- 两次修复后均重跑 focused 和完整 deterministic gate;临时调试探针为 0。
|
||||||
|
|
||||||
|
## Accepted Live Evidence
|
||||||
|
|
||||||
|
- sessionId:`iss-011-stage5-20260720015557`
|
||||||
|
- runId:`run-808ac38f-3ad0-4462-a6d0-ed50d8686473`
|
||||||
|
- Maven profile:`mvp-demo`;Chat/Trace/feedback script exit 0。
|
||||||
|
- Run:CHAT/SUCCESS;answer=109 chars;8 steps;12 tools;self-evaluation 非空;feedback=useful。
|
||||||
|
- Graph:stategraph-v1;planner -> executor -> gatekeeper -> fallback;termination=fallback_completed;degraded=true;evidence retries=0。
|
||||||
|
- Logs:本轮 active application/chat 中 ERROR=0,error appender 无新写入;session/run、evaluation、Trace、feedback 可关联。
|
||||||
|
- DB:V012 JSON column 存在;Run fields/JSON 与 API 一致;step/tool wrong owner=0;unique-session wrong rows=0。
|
||||||
|
- Cleanup:owned process chain 已停止,9900 已释放。
|
||||||
|
- Archive preflight:临时 `target/iss-011-stage5*` 目录为 0,OpenSpec strict 17/17,current-doc stale=0,schema diff=0,`git diff --check` 通过。
|
||||||
|
|
||||||
|
## Known Limits
|
||||||
|
|
||||||
|
- 本次真实模型遗漏 `source_invocation_id`,Gatekeeper 按契约降为 LOW_CONFID 并进入安全 Fallback;这是成功且可审计的 degraded Run,不是完整 Composer 正常路径。
|
||||||
|
- 2026-07-17 外部 MySQL 瞬时断连留下一个 RUNNING 失败尝试;它未计入验收且保留审计,不影响 2026-07-20 exact accepted Run。
|
||||||
@@ -0,0 +1,70 @@
|
|||||||
|
# Chat Diagnosis StateGraph Design Freeze Acceptance
|
||||||
|
|
||||||
|
## 结果
|
||||||
|
|
||||||
|
已接受。阶段 0 完成设计冻结,没有修改运行时实现。
|
||||||
|
|
||||||
|
## 验证
|
||||||
|
|
||||||
|
### 静态验证
|
||||||
|
|
||||||
|
- 命令:`git diff --check`
|
||||||
|
- 结果:passed
|
||||||
|
- 备注:仅有现有 LF/CRLF 提示,无 whitespace error。
|
||||||
|
|
||||||
|
- 检查:Git changed/untracked 路径运行时拒绝列表。
|
||||||
|
- 结果:passed,12 个路径,`src/`、Maven、运行配置、脚本、数据库迁移命中 0。
|
||||||
|
- 备注:阶段 0 只包含 Issue、glossary、OpenSpec/devflow 和执行记录。
|
||||||
|
|
||||||
|
- 检查:proposal → design → specs → tasks 四向对齐。
|
||||||
|
- 结果:passed,gap=0。
|
||||||
|
- 备注:状态、路由、Fallback、审计、测试迁移和六阶段边界均闭环。
|
||||||
|
|
||||||
|
### 脚本验证
|
||||||
|
|
||||||
|
- 命令:`openspec validate chat-diagnosis-stategraph-design-freeze --type change --strict --json`
|
||||||
|
- 结果:passed,1/1。
|
||||||
|
- 备注:当前 change 无结构或场景格式问题。
|
||||||
|
|
||||||
|
- 命令:`openspec validate --specs --strict --json`
|
||||||
|
- 结果:passed,11/11。
|
||||||
|
- 备注:归档前主规格基线未回归。
|
||||||
|
|
||||||
|
### 浏览器/人工验证
|
||||||
|
|
||||||
|
- 结果:not run。
|
||||||
|
- 原因:阶段 0 无 UI 或运行行为。
|
||||||
|
|
||||||
|
### 未验证
|
||||||
|
|
||||||
|
- 单元测试:not run。阶段 0 无代码行为,新增或运行单元测试没有新的验收价值。
|
||||||
|
- Maven E2E:not run。用户明确要求仅阶段 5 在全部实现完成后统一执行。
|
||||||
|
- `logs/`:not inspected for stage acceptance;保留到阶段 5。
|
||||||
|
- 数据库:not queried for stage acceptance;保留到阶段 5。
|
||||||
|
- 备注:此前改造前 live baseline 仅为 research,不计入本阶段验收。
|
||||||
|
|
||||||
|
## 已完成范围
|
||||||
|
|
||||||
|
- 冻结最小 Graph State、所有条件边、四类独立计数和终止路径。
|
||||||
|
- 冻结 verified-input、两类 Fallback、Run 状态与 orchestration trace 边界。
|
||||||
|
- 冻结 L4 消费者、兼容、迁移、回滚和测试替换策略。
|
||||||
|
- 创建 ADR-001,并规定阶段 1–5 引用本阶段 archive。
|
||||||
|
- OpenSpec apply tasks 7/7 完成。
|
||||||
|
|
||||||
|
## 已知限制
|
||||||
|
|
||||||
|
- 运行时仍为旧 Sequential 编排,这是阶段 0 的有意状态。
|
||||||
|
- 新设计尚未经过 Graph 编译、Node 单测、Chat 集成或最终 E2E;后续阶段逐项证明。
|
||||||
|
- 主运行时 specs 暂时仍描述旧行为,只新增设计基线 capability。
|
||||||
|
|
||||||
|
## Bug 修复和诊断
|
||||||
|
|
||||||
|
- 更正此前错误流程模型:删除未提交的总 change,改为六个独立 sm-flow/change。
|
||||||
|
- 更正此前错误 E2E 门禁:阶段 0–4 不做 E2E,阶段 5 统一验证。
|
||||||
|
|
||||||
|
## 交接
|
||||||
|
|
||||||
|
- 下一步:阶段 0 Git commit;提交完成后才能创建阶段 1 change。
|
||||||
|
- Delta sync:新增 `chat-diagnosis-stategraph-design-freeze` 主 spec,共 6 个 requirements,无现有 spec 修改/删除。
|
||||||
|
- OpenSpec 归档确认:用户已在目标中明确要求每阶段 archive,并在后续澄清中再次确认,视为已授权。
|
||||||
|
- OpenSpec 归档结果:已同步主 spec,并归档到 `openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-design-freeze/`。
|
||||||
+27
@@ -0,0 +1,27 @@
|
|||||||
|
# ADR-001: StateGraph owns run-scoped diagnosis control
|
||||||
|
|
||||||
|
**状态**:已接受
|
||||||
|
**日期**:2026-07-17
|
||||||
|
|
||||||
|
复杂 Chat 使用 Spring AI Alibaba StateGraph 管理一次 Diagnosis Run 内的顺序、条件边、有限重试、终止和降级;ReactAgent 只执行 Planner、Executor、Verifier、Composer 的语义任务,Gatekeeper、verified-input builder、evidence retry prepare 和固定 Fallback 使用确定性 Java Node。这样可以精确恢复失败位置并形成可审计路径,同时继续复用现有 Agent、工具和证据协议。
|
||||||
|
|
||||||
|
Gatekeeper 从 `VerifierInputHook` 的隐式执行迁为显式且唯一的 Graph Node。Verifier 只能消费 Gatekeeper passed checked bindings 投影出的 `verified_executor_output` 和 `verified_evidence`;前置验证失败的 Fallback 不得输出 Executor claim。
|
||||||
|
|
||||||
|
编排审计只属于当前 `runId`:有界 `orchestration_events` 压缩后写入 `diagnosis_run.orchestration_trace`,Trace API 只在 `run.orchestrationTrace` 暴露解析对象。它不替代 self evaluation、AgentStep、ToolInvocation 或持久 Graph checkpoint,也不新增 trace 明细表。
|
||||||
|
|
||||||
|
## Considered Options
|
||||||
|
|
||||||
|
- 继续在 `ChatService` 外层叠加 `for/if`:拒绝,失败恢复位置、循环上限和路由原因仍然隐式。
|
||||||
|
- 使用 SupervisorAgent:拒绝,当前是固定诊断 Pipeline,不需要动态选择专科 Agent。
|
||||||
|
- 直接将父 Graph State 交给 `ReactAgent.asNode(...)`:首版拒绝,无法证明 messages、outputKey 和私有执行上下文隔离。
|
||||||
|
- 长期保留 Sequential/StateGraph feature flag 双轨:拒绝,会形成两个编排真理源并增加安全规则漂移。
|
||||||
|
|
||||||
|
## Compatibility And Migration
|
||||||
|
|
||||||
|
`/api/chat`、`executor_evidence_v2`、Verifier 和 Composer 输出协议保持不变。Trace API 只增加 run-scoped 字段;数据库只增加 nullable JSON 列,历史 Run 不回填。
|
||||||
|
|
||||||
|
实施必须按六个独立 sm-flow/change 依次完成:设计冻结、路由骨架、真实节点、ChatService/Trace 切换、测试体系、清理与最终验收。阶段 1–5 必须读取阶段 0 archive,偏离本 ADR 时通过当阶段 OpenSpec 显式修正。
|
||||||
|
|
||||||
|
## Rollback
|
||||||
|
|
||||||
|
每阶段使用独立 Git commit,可整体 revert 当前阶段。生产切换后回滚代码时允许保留 nullable `orchestration_trace` 列;不得用配置重新形成长期双轨。
|
||||||
@@ -0,0 +1,20 @@
|
|||||||
|
# Chat Diagnosis StateGraph Design Freeze Brief
|
||||||
|
|
||||||
|
## 背景
|
||||||
|
|
||||||
|
- 用户目标:将 ISS-011 阶段 0–5 分别作为独立 sm-flow,前一阶段 archive 并 Git commit 后才进入下一阶段。
|
||||||
|
- 当前问题:复杂 Chat 的跨 Agent 状态机分散在 ChatService、SequentialAgent、VerifierInputHook 和 ThreadLocal 中;实现前需先冻结统一设计。
|
||||||
|
- 关联 OpenSpec:`openspec/changes/chat-diagnosis-stategraph-design-freeze/`
|
||||||
|
- devflow 分档:complex
|
||||||
|
|
||||||
|
## 范围
|
||||||
|
|
||||||
|
- 本次要做:冻结最小 Graph State、完整路由、有限重试、安全 Fallback、run-scoped 审计、L4 接口影响和测试替换边界。
|
||||||
|
- 本次不做:任何 Java、SQL、Prompt、配置、运行时 spec 或运行行为修改;不运行 Maven E2E。
|
||||||
|
- 影响区域:ISS-011、glossary、OpenSpec 设计基线、阶段 0 ADR 和后续五阶段交接契约。
|
||||||
|
|
||||||
|
## OpenSpec 对齐
|
||||||
|
|
||||||
|
- proposal 覆盖状态:已覆盖阶段 0 目标、范围、非目标、验收和风险。
|
||||||
|
- specs 覆盖状态:新增设计基线 capability,不提前修改运行时 capabilities。
|
||||||
|
- tasks 覆盖状态:7/7 完成,且仅包含文档、ADR 和静态验证。
|
||||||
@@ -0,0 +1,120 @@
|
|||||||
|
# Chat Diagnosis StateGraph Design Freeze Decisions
|
||||||
|
|
||||||
|
## Question Pool
|
||||||
|
|
||||||
|
| # | 维度 | 问题 | 模式 | 状态 |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| Q1 | 术语 | Diagnosis Orchestration Trace 与 Diagnosis Trace、self evaluation、Graph checkpoint 的边界是什么? | evidence-driven | 已解决 |
|
||||||
|
| Q2 | 边界 | ISS-011 是一个总 sm-flow,还是阶段 0–5 各自独立 sm-flow? | user-interview | 已解决 |
|
||||||
|
| Q3 | 验收 | Maven E2E 在每阶段执行还是只在最终阶段执行? | user-interview | 已解决 |
|
||||||
|
| Q4 | 验收 | 每阶段是否强制新增并运行单元测试? | user-interview | 已解决 |
|
||||||
|
| Q5 | 技术 | 锁定的 StateGraph 版本是否提供条件边、config-aware Node/Edge、recursion limit 和 threadId? | evidence-driven | 已解决 |
|
||||||
|
| Q6 | 架构 | Gatekeeper、Verifier 输入与 Composer 的可信材料边界应复用什么现有契约? | evidence-driven | 已解决 |
|
||||||
|
| Q7 | 接口 | 阶段 0 本身和最终设计分别属于什么接口影响等级? | evidence-driven | 已解决 |
|
||||||
|
| Q8 | 归档 | 每阶段是否应在进入下一阶段前独立 archive 和 Git commit? | user-interview | 已解决 |
|
||||||
|
|
||||||
|
## Evidence-driven
|
||||||
|
|
||||||
|
| 结论 | 证据来源 | 是否已汇报用户 |
|
||||||
|
|---|---|---|
|
||||||
|
| Orchestration Trace 是 run 级紧凑编排摘要,不是事件日志、自评估或 Graph checkpoint | `devflow/glossary/CONTEXT.md`、ISS-011 §7.5、session-run-trace-isolation decisions | 已汇报 |
|
||||||
|
| Gatekeeper 从 Hook 迁为显式 Node 是有意替换旧内部架构,不允许双入口 | executor-gatekeeper-hook decisions、ISS-011 §2.3/§7.2/§14 | 已汇报 |
|
||||||
|
| Verifier verified evidence 必须来自 Gatekeeper 通过的 checked bindings;Composer 不得读取 raw Executor/tool output | verifier-evidence-reference-fidelity、executor-composer-final-answer 历史档案、ISS-011 §7.4 | 已汇报 |
|
||||||
|
| 本地 1.1.2.0 API 支持条件边、config-aware Node/Edge、recursion limit、RunnableConfig.threadId 和 ReactAgent.call(input, config) | Maven dependency tree 与本地 JAR `javap` 研究记录 | 已汇报 |
|
||||||
|
| 阶段 0 是文档/规格交付,无运行时接口变化;其冻结的最终目标涉及状态机、DB 和 Trace API,属于 L4 | sm-flow operating-rules、ISS-011 §10/§14 | 已汇报 |
|
||||||
|
| 主 OpenSpec 的外层 groundedness round、Hook Gatekeeper 和完整 tool_trace_summary 输入与冻结设计冲突 | `openspec/specs/chat-verifier-agent/spec.md` 等主规格 | 已汇报 |
|
||||||
|
|
||||||
|
## User-interview
|
||||||
|
|
||||||
|
| 问题原文 | 用户原话 | 确认状态 | OpenSpec 回写 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| ISS-011 的阶段关系如何映射 sm-flow? | “iss-011里每个阶段,都是一个sm-flow,而不是将一整个iss-011打包进一个sm-flow中” | 已确认 | 已回写 proposal 范围与非目标 |
|
||||||
|
| Maven E2E 何时执行? | “端到端只在最后阶段全部完成后才验证” | 已确认 | 已回写 acceptance 与 out of scope |
|
||||||
|
| 阶段单元测试是否强制? | “每个阶段如果有必要添加单元测试验收的话,就加,没有必要的话就跳过单元测试” | 已确认 | 已回写 acceptance |
|
||||||
|
| 每阶段如何进入下一阶段? | “每个阶段需要归档完并提交才能进入下一个阶段” | 已确认 | 已回写阶段门禁 |
|
||||||
|
|
||||||
|
## OpenSpec Input Context
|
||||||
|
|
||||||
|
### devflow index
|
||||||
|
|
||||||
|
已命中并读取:
|
||||||
|
|
||||||
|
- `session-run-trace-isolation`
|
||||||
|
- `executor-gatekeeper-hook`
|
||||||
|
- `verifier-evidence-reference-fidelity`
|
||||||
|
- `executor-evidence-output-contract`
|
||||||
|
- `executor-v2-output-contract`
|
||||||
|
- `executor-verifier-claim-checks`
|
||||||
|
- `executor-composer-final-answer`
|
||||||
|
- `mvp-demo-trace-acceptance`
|
||||||
|
|
||||||
|
### Historical constraints entering OpenSpec
|
||||||
|
|
||||||
|
- `runId` 是运行态和 Trace 的所有权边界。
|
||||||
|
- AgentStep、ToolInvocation 与 self evaluation 的既有 run 绑定必须保留。
|
||||||
|
- checked binding 是 verified evidence 的复用来源。
|
||||||
|
- Composer 的 allowed-material 安全边界不得削弱。
|
||||||
|
- 旧 Gatekeeper-in-Hook 决策在本设计中被显式废止,不能保留并行入口。
|
||||||
|
- 主 OpenSpec 的旧运行时要求只能在对应实现阶段修改,阶段 0 不提前宣称代码已切换。
|
||||||
|
|
||||||
|
## Key Decisions
|
||||||
|
|
||||||
|
### 六个独立交付单元
|
||||||
|
|
||||||
|
阶段 0–5 分别使用独立 slug、OpenSpec change、devflow 项目、Archive 和 Git commit。前一 change 归档并提交后才创建下一 change。
|
||||||
|
|
||||||
|
### 阶段 0 不承载后续实现
|
||||||
|
|
||||||
|
阶段 0 只冻结设计并归档长期上下文。阶段 1–5 的代码、数据库、测试和清理任务不进入本 change 的 tasks。
|
||||||
|
|
||||||
|
### 验收分层
|
||||||
|
|
||||||
|
阶段 0 没有运行时行为变化,不新增或运行单元测试,也不运行 Maven E2E。验收使用 OpenSpec strict validation、结构检查和文档一致性检查。Maven E2E、`logs/` 与数据库只在阶段 5 统一执行。
|
||||||
|
|
||||||
|
### 接口影响
|
||||||
|
|
||||||
|
- 当前 change:L1 文档/设计交付,不改变调用方可观察行为。
|
||||||
|
- 冻结目标:L4,涉及内部状态机语义、Trace API 加字段、数据库契约、迁移和回滚。
|
||||||
|
- 兼容:保持 `/api/chat` 和 Agent 输出协议;Trace API 为加法式 run 字段。
|
||||||
|
- 回滚:后续运行时代码可按阶段 Git revert;nullable DB 列可保留,不引入双轨配置。
|
||||||
|
- 消费者:ChatController/ChatService、Trace DTO/Service/UI/demo、数据库迁移、评测与运维审计。
|
||||||
|
|
||||||
|
## Risks Accepted
|
||||||
|
|
||||||
|
- 阶段 0 archive 后,冻结设计已完成但运行时仍为旧 Sequential;这是有意的阶段性状态。
|
||||||
|
- 六个 changes 的一致性由后续每次 context/commit gate 读取并审计本档案保证。
|
||||||
|
- 不为阶段 0 的纯设计变更新增无行为价值的单元测试。
|
||||||
|
|
||||||
|
## Cross-Artifact 对齐检查
|
||||||
|
|
||||||
|
| 上游 → 下游 | 检查内容 | 状态 |
|
||||||
|
|---|---|---|
|
||||||
|
| brief/prd → proposal | Issue 的阶段 0 目标、范围、非目标和验收已进入 proposal | 已对齐 |
|
||||||
|
| proposal → 设计产物 | 状态、路由、重试、Fallback、审计、测试迁移和六阶段边界均进入 design | 已对齐 |
|
||||||
|
| 设计产物 → specs/tasks | L4 影响、安全边界、终止规则和阶段 0 文档工作均进入 spec 或 tasks | 已对齐 |
|
||||||
|
| specs → tasks | 设计基线、源文档、ADR、严格验证和无运行时变更检查均有可执行任务 | 已对齐 |
|
||||||
|
|
||||||
|
Gap:无。
|
||||||
|
|
||||||
|
## Architecture Audit
|
||||||
|
|
||||||
|
输入链为 `ChatController -> ChatService`,目标处理链为 Run 生命周期 → Graph orchestrator → 白名单 Agent/Java Nodes,输出仍为 `ChatResult + DiagnosisRun + exact RunTrace`。Graph State 只属于当前 run,verified evidence 只来自当前 run 的 passed checked bindings,编排摘要只写当前 Diagnosis Run。旧“Gatekeeper stays in VerifierInputHook”决策被 ISS-011 的显式 Node 有意替代,但旧 Gatekeeper 规则本身继续复用;旧“不新增 trace 主表”和 run ownership 决策保持成立。主要耦合风险是 Hook/Node 双执行、阶段性主 spec 漂移和 Trace 投影重复,design 已分别通过单入口迁移、按阶段修改运行时 specs、唯一 `run.orchestrationTrace` 投影缓解。架构风险已由用户确认的 ISS-011 冻结决策和六阶段交付口径接受,无需返回 grill。
|
||||||
|
|
||||||
|
## Commit Preflight
|
||||||
|
|
||||||
|
- proposal、design、specs、tasks 完整:通过。
|
||||||
|
- 所有 user-interview 已确认:通过。
|
||||||
|
- 未汇报 evidence-driven 结论:无。
|
||||||
|
- L4 接口影响、消费者、兼容、迁移和回滚:已在 design 独立章节记录。
|
||||||
|
- OpenSpec strict validation:change 1/1 通过,主 specs 11/11 通过。
|
||||||
|
- Apply 授权:用户启动目标时要求分阶段完整执行、archive 和 Git commit,后续又确认六个独立 sm-flow;视为当前阶段连续执行授权。
|
||||||
|
|
||||||
|
## Apply Evidence
|
||||||
|
|
||||||
|
- Task 1.1:ISS-011 §3–14 与 committed design 的 State 字段、条件边、四类计数、Fallback、安全审计和测试迁移语义一致;未发现需回写 OpenSpec 的 drift,未修改运行时代码。
|
||||||
|
- Task 1.2:将 glossary 中新 Orchestration Trace 的具体 StateGraph/Graph State 表述收敛为“诊断编排/编排上下文快照”;Verifier 的 verified-output/evidence 名称作为跨节点协议保留。
|
||||||
|
- Task 2.1:创建 ADR-001,记录 StateGraph 控制边界、显式 Gatekeeper、run-scoped trace、替代方案、兼容、迁移和回滚。
|
||||||
|
- Task 2.2:六个独立 change 的唯一边界已写入 design、design-baseline spec、ADR 和根计划;阶段 1–5 必须引用阶段 0 archive。
|
||||||
|
- Task 3.1:最终 strict validation 为 change 1/1、主 specs 11/11;proposal/design/spec/tasks 对状态、Fallback、审计、运行时非目标和测试迁移均形成闭环,cross-artifact gap 为 0。
|
||||||
|
- Task 3.2:Git 枚举 12 个 changed/untracked 路径,运行时拒绝列表命中 0;无 `src/`、Maven、运行配置、脚本或数据库迁移变更。
|
||||||
|
- Task 3.3:单元测试未运行,因为阶段 0 无代码行为;Maven E2E、`logs/` 和数据库核验未运行,因为用户明确要求只在阶段 5 全部实现后统一执行。改造前 research baseline 不计入阶段验收。
|
||||||
@@ -0,0 +1,30 @@
|
|||||||
|
# Chat Diagnosis StateGraph Design Freeze Evidence
|
||||||
|
|
||||||
|
## 证据
|
||||||
|
|
||||||
|
| 来源 | 证据 | 结论 | 是否已汇报 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| ISS-011 §3–14 | 已定义目标流程、26 个 State 字段、有限回边、Fallback、Trace 和测试策略 | 可形成无开放分支的阶段 0 设计基线 | 是 |
|
||||||
|
| `ChatService` / Controller 引用核查 | 生产链为 Controller → ChatService → SequentialAgent,细粒度状态散布在 ChatService | StateGraph 应只接管 Run 内控制,ChatService 保留生命周期 | 是 |
|
||||||
|
| `VerifierInputHook` / `VerifierContextHolder` 引用核查 | Gatekeeper、工具摘要和解析结果通过 Hook/ThreadLocal 隐式传播 | Gatekeeper 与 verified-input 必须迁为显式 Node,不能保留双入口 | 是 |
|
||||||
|
| executor-gatekeeper-hook 历史 decisions | 旧阶段曾决定 Gatekeeper 留在 Hook 且不重试 | 本设计有意替换入口,但继续复用确定性规则 | 是 |
|
||||||
|
| session-run-trace-isolation 历史 decisions | runId 是执行/Trace 所有权边界,AgentStep/ToolInvocation 已承载明细 | orchestration trace 只附着当前 Run,不新增明细主表 | 是 |
|
||||||
|
| Verifier/Composer 历史档案 | checked bindings 和 allowed material 是可信边界 | Builder 不重复验真,Fallback 不泄漏 raw claim | 是 |
|
||||||
|
| 本地 Graph Core 1.1.2.0 JAR | 已验证条件边、config-aware Node/Edge、recursion limit、threadId 与 Agent call API | 阶段 1 可基于锁定签名实现,不采用网上漂移示例 | 是 |
|
||||||
|
| 主 OpenSpec 核查 | 旧 spec 仍描述 Hook Gatekeeper、完整 tool trace 和外层 groundedness round | 阶段 0 不提前改运行时 spec,后续对应实现阶段再修改 | 是 |
|
||||||
|
| 用户原话 | 六个独立 sm-flow;仅阶段 5 E2E;单测按必要性 | 阶段门禁与测试口径已固定 | 是 |
|
||||||
|
|
||||||
|
## Evidence-driven 结论
|
||||||
|
|
||||||
|
- 结论:阶段 0 运行时影响为 L1,但冻结目标为 L4。
|
||||||
|
- 证据:本阶段 changed paths 无 `src/`/SQL/配置;最终目标改变状态机、Trace API 和 DB。
|
||||||
|
- 风险:设计完成不等于运行时完成。
|
||||||
|
- 用户确认:已确认分阶段交付。
|
||||||
|
- 结论:旧 Gatekeeper-in-Hook 决策应被显式 Node 有意替代。
|
||||||
|
- 证据:当前 Hook 引用与 ISS-011 条件边要求冲突。
|
||||||
|
- 风险:迁移期双执行。
|
||||||
|
- 用户确认:ISS-011 冻结决策已确认。
|
||||||
|
- 结论:阶段 0 不需要单元测试或 E2E。
|
||||||
|
- 证据:Git 运行时拒绝列表命中 0,OpenSpec strict 和文档一致性已覆盖本阶段可观察产物。
|
||||||
|
- 风险:运行时正确性仍未证明,将由阶段 1–5 测试和最终 E2E 证明。
|
||||||
|
- 用户确认:已确认。
|
||||||
@@ -0,0 +1,42 @@
|
|||||||
|
# Chat Diagnosis StateGraph Real Nodes 验收
|
||||||
|
|
||||||
|
## 结果
|
||||||
|
|
||||||
|
已接受。OpenSpec tasks 27/27 完成;真实 Nodes 可构造和测试,旧生产 Sequential 路径保持可用,阶段 2 未切换生产入口。
|
||||||
|
|
||||||
|
## 静态验证
|
||||||
|
|
||||||
|
- `openspec validate chat-diagnosis-stategraph-real-nodes --strict`:通过。
|
||||||
|
- `openspec validate --specs --strict`:13 passed,0 failed。
|
||||||
|
- `git diff --check`:通过;只有 LF/CRLF 转换提示,无 whitespace error。
|
||||||
|
- 源码/引用检查:`ChatService` Graph 引用 0;新增源码 forbidden refs 0;protocol 反向依赖 0;DB/Trace/Prompt diff 0。
|
||||||
|
- 装配对齐:无核心 TODO/FIXME/placeholder;Graph factory 继续拥有所有 retry counter 与 Planner reset。
|
||||||
|
|
||||||
|
## 脚本验证
|
||||||
|
|
||||||
|
- 新 protocol/Node/CompiledGraph/Router/Trace focused suite:通过。
|
||||||
|
- `mvn -q "-Dtest=ChatServiceSequentialAgentTest,VerifierInputHookTest,ExecutorGatekeeperServiceTest" test`:通过。
|
||||||
|
- 两组测试合计:100 tests,0 failures,0 errors,0 skipped。
|
||||||
|
- `mvn -q -DskipTests test-compile`:通过。
|
||||||
|
|
||||||
|
## 浏览器/人工验证
|
||||||
|
|
||||||
|
- 未运行。阶段 2 不改变用户入口或 UI,没有独立人工验收价值。
|
||||||
|
|
||||||
|
## 未验证
|
||||||
|
|
||||||
|
- Maven live E2E、`logs/` 日志和 `scripts/query_mysql.py` 数据库核验未运行。
|
||||||
|
- 原因:用户明确要求只在阶段 5 全部实现后统一做最终 E2E;阶段 2 尚未切换生产入口。
|
||||||
|
- 风险:当前验收只证明组件/Graph 契约与旧路径回归,不证明真实生产装配、模型、DB 和 Trace 全链路。
|
||||||
|
|
||||||
|
## 已完成范围
|
||||||
|
|
||||||
|
- 中立共享 protocol 与旧路径行为保持型委托。
|
||||||
|
- 四个 Agent adapters、显式 Gatekeeper、可信投影、关键证据补查、Composer 与两类 Fallback。
|
||||||
|
- 真实 CompiledGraph 装配、critical gap 路由修复和完整 focused test 证据。
|
||||||
|
|
||||||
|
## 交接
|
||||||
|
|
||||||
|
- 下一步:归档本 change、独立提交阶段 2,然后启动阶段 3 `chat-diagnosis-stategraph-chatservice-cutover`。
|
||||||
|
- OpenSpec 归档:已同步主 specs,并归档到 `openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-real-nodes/`。
|
||||||
|
- 归档授权:用户已明确“直接实现吧,不用找我授权了”,授权后续阶段在门禁通过后直接归档和提交。
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
# Chat Diagnosis StateGraph Real Nodes Brief
|
||||||
|
|
||||||
|
## 背景
|
||||||
|
|
||||||
|
- 用户目标:将 ISS-011 阶段 2 作为独立 sm-flow,接入真实 Agent/Java Nodes,并在归档、验收和独立 Git 提交后才进入阶段 3。
|
||||||
|
- 当前问题:阶段 1 只有 Fake Node 路由骨架;Executor/Verifier/Composer 协议逻辑分散在 Hook 与 ChatService,真实 Graph 尚不能安全调用 Agent、Gatekeeper 或投影可信材料。
|
||||||
|
- 关联 OpenSpec:`openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-real-nodes/`
|
||||||
|
- devflow 分档:complex
|
||||||
|
|
||||||
|
## 范围
|
||||||
|
|
||||||
|
- 本次要做:共享无状态 protocol 组件;Planner/Executor/Verifier/Composer adapters;显式 Gatekeeper、Verified Input、Evidence Retry、Fallback Nodes;真实 CompiledGraph 装配和 focused tests。
|
||||||
|
- 本次不做:不切换 ChatService 生产入口,不修改 DB、Trace API、Prompt 契约,不删除旧 Sequential/Hook,不运行 live E2E。
|
||||||
|
- 影响区域:`diagnosis.protocol`、`graph.diagnosis`、`VerifierInputHook`、`ChatService` 共享逻辑委托及对应测试。
|
||||||
|
|
||||||
|
## OpenSpec 对齐
|
||||||
|
|
||||||
|
- proposal 覆盖状态:已覆盖。
|
||||||
|
- design 覆盖状态:已覆盖;接口影响为 L2,生产切换明确延期到阶段 3。
|
||||||
|
- specs 覆盖状态:已覆盖真实 Nodes 安全边界,并修正 critical evidence gap 路由条件。
|
||||||
|
- tasks 覆盖状态:27/27 完成。
|
||||||
@@ -0,0 +1,218 @@
|
|||||||
|
# Chat Diagnosis StateGraph Real Nodes Decisions
|
||||||
|
|
||||||
|
## Question Pool
|
||||||
|
|
||||||
|
| # | 维度 | 问题 | 模式 | 状态 |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| Q1 | 边界 | 阶段 2 是否切换 ChatService/DB/Trace? | user-interview(六阶段已确认) | 已解决 |
|
||||||
|
| Q2 | 复用 | 新 Nodes 如何避免复制 Hook/ChatService 解析与安全渲染? | evidence-driven | 已解决 |
|
||||||
|
| Q3 | Agent | Adapter 如何调用真实 ReactAgent 并保留 Hook/ToolCallback/run config? | evidence-driven | 已解决 |
|
||||||
|
| Q4 | Gatekeeper | 显式 Node 如何保证当前 run、单次调用和 fail-closed 标准化? | evidence-driven | 已解决 |
|
||||||
|
| Q5 | 安全 | passed checked bindings 如何投影为 verified output/evidence? | evidence-driven | 已解决 |
|
||||||
|
| Q6 | 补证据 | 哪些 facts 可触发 evidence retry,completed queries 如何表达? | evidence-driven | 已解决 |
|
||||||
|
| Q7 | Fallback | 前置验证失败与 Composer 后置失败可分别使用哪些材料? | evidence-driven | 已解决 |
|
||||||
|
| Q8 | Prompt | 阶段 2 是否立即修改共享 Verifier Prompt? | evidence-driven | 已解决 |
|
||||||
|
| Q9 | 验收 | 阶段 2 是否需要单元/集成测试与 E2E? | evidence-driven + user rule | 已解决 |
|
||||||
|
|
||||||
|
## Evidence-driven
|
||||||
|
|
||||||
|
| 结论 | 证据来源 | 是否已汇报用户 |
|
||||||
|
|---|---|---|
|
||||||
|
| ReactAgent.call(input, config) 返回 AssistantMessage,现有 AgentLoggingHook 从 config metadata 读取 sessionId/runId | 本地 1.1.2.0 javap、AgentLoggingHook | 已汇报 |
|
||||||
|
| Executor parser 当前在 VerifierInputHook,Verifier/Composer parser 与 renderer 当前在 ChatService | 定向源码阅读和 references | 已汇报 |
|
||||||
|
| checked_bindings 提供 claim_id/tool_name/source_invocation_id/raw_path/matched_text/status | ExecutorGatekeeperService | 已汇报 |
|
||||||
|
| Gatekeeper.validateRun 使用 run-scoped ToolInvocation repository | ExecutorGatekeeperService tests/code | 已汇报 |
|
||||||
|
| Graph path不能注册旧 VerifierInputHook,否则 Gatekeeper 会双执行且输入包含 full tool trace | 阶段 0 ADR、Hook 源码 | 已汇报 |
|
||||||
|
| evidence retry 只允许 critical no_evidence/indirect_support | 阶段 0 design/ISS-011 | 已汇报 |
|
||||||
|
| 阶段 1 router 未检查 is_critical,属于实现偏离 | Router 源码与阶段 0 baseline 对照 | 已汇报 |
|
||||||
|
| 共享 Prompt 暂时仍服务旧 Sequential path,本阶段直接修改会提前破坏生产 payload | ChatService agent builder + VerifierInputHook + prompt | 已汇报 |
|
||||||
|
| Node 输入/安全边界新增且旧解析会重构,单元和 focused regression 必要;E2E 不必要 | 阶段边界和用户规则 | 已汇报 |
|
||||||
|
|
||||||
|
## User-interview
|
||||||
|
|
||||||
|
| 问题原文 | 用户原话 | 确认状态 | OpenSpec 回写 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| 阶段 2 是否独立 sm-flow? | “iss-011里每个阶段,都是一个sm-flow” | 已确认 | 独立 change |
|
||||||
|
| 是否可提前切生产入口? | “每个阶段需要归档完并提交才能进入下一个阶段” | 已确认不可提前阶段 3 | Out of Scope |
|
||||||
|
| 阶段 2 是否执行 E2E? | “端到端只在最后阶段全部完成后才验证” | 已确认不执行 | Acceptance |
|
||||||
|
| 是否添加单元测试? | “如果有必要添加单元测试验收的话,就加” | 已确认规则;本阶段判定必要 | Acceptance |
|
||||||
|
|
||||||
|
## Context And Handoff
|
||||||
|
|
||||||
|
- 阶段 0 archive/commit:design baseline / `581daff`
|
||||||
|
- 阶段 1 archive/commit:routing skeleton / `42ba204`
|
||||||
|
- 当前 change:`chat-diagnosis-stategraph-real-nodes`
|
||||||
|
- 后续 change:`chat-diagnosis-stategraph-chatservice-cutover`,只能在本阶段 archive + commit 后创建。
|
||||||
|
|
||||||
|
## Technical Decisions
|
||||||
|
|
||||||
|
### Shared protocol components
|
||||||
|
|
||||||
|
- 抽取 Executor、Verifier、Composer 解析器,旧 Hook/ChatService 委托新组件。
|
||||||
|
- 抽取 Composer safe input builder / renderer,旧 ChatService 保持相同输出。
|
||||||
|
- JSON sanitization 只存在一份共享实现,不在每个 Node copy。
|
||||||
|
- 抽取过程是行为保持 refactor;现有 focused tests 是回归门禁。
|
||||||
|
|
||||||
|
### Agent adapter boundary
|
||||||
|
|
||||||
|
- `DiagnosisAgentInvoker` 是最小 port:`invoke(String, RunnableConfig) -> String`。
|
||||||
|
- `ReactAgentDiagnosisInvoker` 只包装 `ReactAgent.call(...).getText()`。
|
||||||
|
- Planner/Executor/Verifier/Composer adapters 各自拥有白名单 input projector、parser 和 status event。
|
||||||
|
- tests 使用 fake invoker,另有 ReactAgent wrapper test。
|
||||||
|
|
||||||
|
### Legacy Hook coexistence
|
||||||
|
|
||||||
|
- 旧 Sequential path 在阶段 3 前仍通过 VerifierInputHook 执行 Gatekeeper。
|
||||||
|
- Graph Verifier Agent 不注册该 Hook;显式 Gatekeeper Node 是 Graph 中唯一 validation 入口。
|
||||||
|
- Parser/enricher 可共享,但 Hook 的 legacy payload/prompt 暂不改变。
|
||||||
|
- 阶段 3 切换生产入口并同步 Verifier Prompt;阶段 5 删除旧隐式结构。
|
||||||
|
|
||||||
|
### Gatekeeper normalization
|
||||||
|
|
||||||
|
- raw pass → PASS。
|
||||||
|
- raw fail + severity low_confid → LOW_CONFID。
|
||||||
|
- raw fail + severity reject → REJECT。
|
||||||
|
- 缺失、unknown、异常或不一致 → REJECT。
|
||||||
|
- verified_binding_count 只统计 checked_bindings.status=pass。
|
||||||
|
|
||||||
|
### Verified projection
|
||||||
|
|
||||||
|
- 通过项按 claim_id + source_invocation_id + tool_name + raw_path 与原 binding 精确匹配。
|
||||||
|
- verified_executor_output 只包含 answer_version 与至少一条 passed binding 的 filtered claims;合法零 claim 保持空列表。
|
||||||
|
- verified_evidence 只含 claim_id/source_invocation_id/tool_name/raw_path/matched_text。
|
||||||
|
- hypotheses、失败 binding、未引用 ToolInvocation、raw Executor output 不投影。
|
||||||
|
|
||||||
|
### Evidence retry
|
||||||
|
|
||||||
|
- extractor 要求 `is_critical=true` 且 verification 为 no_evidence/indirect_support。
|
||||||
|
- gap 字段:claim_id(从 `claim-id: text` 提取或空)、fact、verification、reason。
|
||||||
|
- completed queries 由 verified evidence 的 tool/invocation/path 去重生成。
|
||||||
|
- prior verified output/evidence 原样只读进入 retry_context。
|
||||||
|
- 约束固定:max_retry=1、do_not_repeat_successful_queries、only_execute_incremental_queries、preserve_prior_verified_claims。
|
||||||
|
- Executor 输入声明“增量执行、完整输出”,Java 不合并 claims。
|
||||||
|
|
||||||
|
### Interface impact
|
||||||
|
|
||||||
|
- 等级:L2 internal interface;共享 parser 委托保持旧行为。
|
||||||
|
- 新消费者:阶段 3 Graph orchestrator/Agent factory。
|
||||||
|
- 当前外部 API/DB/生产路由:无变化。
|
||||||
|
- 回滚:revert 本阶段提交;旧 Sequential path仍完整。
|
||||||
|
- 有意修正:Router 只允许 critical gap,属于对冻结基线的代码修复。
|
||||||
|
|
||||||
|
## Risks Accepted
|
||||||
|
|
||||||
|
- 真实模型尚未执行;Node contract 通过 fake invoker/real service tests证明,阶段 5 才 E2E。
|
||||||
|
- Prompt 输入说明与 Graph input 的最终同步推迟到阶段 3,避免当前旧生产 Hook 提前不兼容。
|
||||||
|
- 旧 Hook 暂时仍存在,但不进入 Graph action graph;阶段 5 必须删除。
|
||||||
|
|
||||||
|
## Architecture Audit
|
||||||
|
|
||||||
|
### Module and ownership map
|
||||||
|
|
||||||
|
`Diagnosis Context + RunnableConfig` → Agent adapters / deterministic Nodes → owned Graph State fields → `DiagnosisGraphRouter` → next Node or END → `final_answer + orchestration_events`。阶段 2 只提供这条可构造链,阶段 3 才由 `ChatService` 创建 Run、构造 Agent 实例并调用 Graph。
|
||||||
|
|
||||||
|
| 模块 | 数据所有权 | 允许依赖 |
|
||||||
|
|---|---|---|
|
||||||
|
| `diagnosis.protocol` | JSON contract、纯解析结果、安全 input/rendering | Jackson 与纯 DTO;不依赖 Graph/Hook/ChatService/ThreadLocal |
|
||||||
|
| `graph.diagnosis` Agent adapters | 白名单输入、Agent attempt status/event | protocol、注入的 invoker、Graph API |
|
||||||
|
| Gatekeeper Node | raw Gatekeeper result、normalized status/count | ExecutorGatekeeperService、当前 run config |
|
||||||
|
| Verified Input / Retry Prepare | verified projection、critical gaps、retry context | raw result 的只读投影与 protocol DTO |
|
||||||
|
| Router / Factory | 条件边、有限计数、调用次序 | Graph State;不解析 Agent raw output |
|
||||||
|
| legacy Hook / ChatService | 阶段 3 前的生产 Sequential 流程 | 只委托 protocol;不得消费真实 Graph action set |
|
||||||
|
|
||||||
|
### Lifecycle and coupling audit
|
||||||
|
|
||||||
|
Graph State 和 events 都是 invocation-scoped,runId 只从当前 RunnableConfig 获取,Nodes 不持有跨 Run 可变状态。旧路径与 Graph 路径阶段性共享的只有无状态 protocol 组件和 Gatekeeper service,不共享 ThreadLocal 或 Agent output。ReactAgent/invoker 必须构造注入,阶段 2 不复制 ChatService Prompt/Agent factory。唯一有意的阶段性耦合是 legacy consumers 改为委托 protocol,这由旧 focused tests 和完整 revert 保护。阶段 3 前生产入口、DB、Trace、Prompt 均保持隔离。
|
||||||
|
|
||||||
|
### Cross-artifact alignment
|
||||||
|
|
||||||
|
| 对齐链 | 结果 | 证据 |
|
||||||
|
|---|---|---|
|
||||||
|
| brief/proposal 目标、范围、非目标 → proposal | 已对齐 | 独立阶段边界、真实 Nodes、无生产切换均明确 |
|
||||||
|
| proposal 承诺与约束 → design | 已对齐 | invoker、共享组件、Gatekeeper、投影、retry、Fallback、L2 均有决策 |
|
||||||
|
| design 架构/接口结论 → specs/tasks | 已对齐 | 中立 protocol 包、构造注入、fail-closed 和生产隔离均有任务/行为 |
|
||||||
|
| specs 可观察行为 → tasks 可执行切片 | 已对齐 | 每个 requirement 至少由一个实现任务和一个测试/验收任务覆盖 |
|
||||||
|
|
||||||
|
### Audit result
|
||||||
|
|
||||||
|
审计发现共享协议组件包所有权与 Agent 实例装配边界需要显式化,已回写 design/tasks。未发现状态字段、路由计数、run 生命周期或阶段边界的新冲突。接口影响维持 L2,消费者都在本 change 与下一阶段明确范围内。剩余风险是共享抽取的旧行为漂移和 binding 投影泄漏,均有 focused regression 与 mixed-binding tests。架构风险可接受,无未解决问题。
|
||||||
|
|
||||||
|
## Commit Gate
|
||||||
|
|
||||||
|
- proposal/design/specs/tasks:全部存在,OpenSpec status `isComplete=true`。
|
||||||
|
- 当前 change strict validation:通过。
|
||||||
|
- 主 specs strict validation:13 passed,0 failed。
|
||||||
|
- 规格结构:9 requirements、22 scenarios;tasks:27 个 checkbox 切片。
|
||||||
|
- Cross-artifact:4/4 已对齐,gap=0。
|
||||||
|
- 接口影响:L2,已在 design 独立章节记录消费者、兼容和回滚。
|
||||||
|
- Question pool:所有 evidence-driven 已汇报;所有 user-interview 已确认;无未决项。
|
||||||
|
- Preflight scope:`git diff --check` 通过;Commit checkpoint 未修改业务代码。
|
||||||
|
- 结论:Draft OpenSpec 达到可执行状态,创建 `.committed`。
|
||||||
|
|
||||||
|
## Apply Authorization
|
||||||
|
|
||||||
|
- 用户原话:“直接实现吧,不用找我授权了”。
|
||||||
|
- 解释:阶段 2–5 后续 checkpoint 可在前置门禁通过后直接继续,不再因 Apply 或 Archive 授权暂停。
|
||||||
|
- 不扩大范围:每阶段仍须独立 OpenSpec、验收归档、Git commit;阶段 0–4 不做 E2E,阶段 5 才统一执行。
|
||||||
|
|
||||||
|
## Pre-apply Research
|
||||||
|
|
||||||
|
### Reference implementations read
|
||||||
|
|
||||||
|
- `src/main/java/com/superbiz/agent/hook/VerifierInputHook.java`:Executor JSON sanitization、parse status、tool-name normalization、唯一 invocation 回填与 legacy Gatekeeper payload。
|
||||||
|
- `src/main/java/com/superbiz/agent/service/ChatService.java`:Verifier claim/fact parser、Gatekeeper ceiling、Composer allowed-material builder、Composer parser/safe renderer、旧 retry context。
|
||||||
|
- `src/main/java/com/superbiz/agent/service/ExecutorGatekeeperService.java`:`validateRun`、severity normalization source、checked binding 与 matched_text 结构。
|
||||||
|
- `src/main/java/com/superbiz/agent/graph/diagnosis/DiagnosisGraphFactory.java`:config-aware action ports、technical retry/evidence retry counter 所有权和 recursion limit。
|
||||||
|
- `src/main/java/com/superbiz/agent/graph/diagnosis/OrchestrationEvent.java` 与 `src/test/java/com/superbiz/agent/graph/diagnosis/ScriptedDiagnosisGraphActions.java`:每 attempt 单 event 形态。
|
||||||
|
- `src/main/java/com/superbiz/agent/hook/AgentLoggingHook.java`:RunnableConfig metadata 中 sessionId/runId 的审计读取方式。
|
||||||
|
- `src/test/java/com/superbiz/agent/hook/VerifierInputHookTest.java`、`ChatServiceSequentialAgentTest.java`、`DiagnosisGraphRoutingTest.java`:JUnit 5、Mockito 边界替身和 CompiledGraph observable behavior 测试风格。
|
||||||
|
- 本地 `1.1.2.0` JAR `javap`:`ReactAgent.call(String,RunnableConfig)` 与 `RunnableConfig.threadId/metadata` 公共 API。
|
||||||
|
|
||||||
|
### Technology inventory
|
||||||
|
|
||||||
|
| 类别 | 项目标准 / 本阶段使用 |
|
||||||
|
|---|---|
|
||||||
|
| JSON contract | Jackson `ObjectMapper`、`LinkedHashMap` 保持稳定字段顺序;共享 sanitization 只存在一份 |
|
||||||
|
| Graph Node | `AsyncNodeActionWithConfig` 返回 `CompletableFuture<Map<String,Object>>`;state 默认 Replace、events Append |
|
||||||
|
| Agent boundary | 构造注入的 `DiagnosisAgentInvoker`;ReactAgent wrapper 透传同一个 RunnableConfig |
|
||||||
|
| Run scope | `config.metadata("runId")` 为 Gatekeeper 唯一运行边界;缺失即 fail closed |
|
||||||
|
| Error handling | 合法 contract/invalid output/temporary/permanent 分离;unknown exception 不推断为可重试 |
|
||||||
|
| Tests | JUnit 5;只在 ReactAgent、repository 等系统边界使用 fake/mock;Node/CompiledGraph 走公开 action/graph 接口 |
|
||||||
|
| Request/response | 本阶段不改 Controller/DTO/API,不适用 |
|
||||||
|
| MQ/Consumer | 本阶段不涉及,不适用 |
|
||||||
|
| DB/Trace/Prompt | 本阶段禁止修改,阶段 3 处理 |
|
||||||
|
|
||||||
|
### Reuse and new infrastructure
|
||||||
|
|
||||||
|
- 新建中立 `com.superbiz.agent.diagnosis.protocol`:`JsonPayloadSupport`、Executor/Verifier/Composer protocol、safe input/rendering、`EvidenceGapExtractor`;不得依赖 Graph/Hook/ChatService/ThreadLocal。
|
||||||
|
- 新建 `graph.diagnosis` Node 层:invoker wrapper、failure classifier、四个 Agent adapters、Gatekeeper、Verified Input、Retry Prepare、Fallback 和 action assembly。
|
||||||
|
- 不新建 Prompt/Agent factory;阶段 3 通过构造注入已有 ReactAgent 实例。
|
||||||
|
- 不新增 Maven dependency、数据库迁移、配置项或生产 consumer。
|
||||||
|
|
||||||
|
### Pre-apply conclusion
|
||||||
|
|
||||||
|
参考实现、API 签名、异常和测试标准已足以指导实现;未发现 devflow/OpenSpec 冲突。进入 TDD tracer bullet,先锁定共享 Executor parser 的合法 no-evidence 与 malformed 行为。
|
||||||
|
|
||||||
|
## Apply Completion And Verification
|
||||||
|
|
||||||
|
### Assembly alignment
|
||||||
|
|
||||||
|
- 首个共享 protocol 模块与最终真实 Graph action assembly 均逐项对照 design/specs:invoker 显式透传 RunnableConfig,Gatekeeper 单次 fail-closed,Verified Input 精确投影,Verifier/Composer 固定输入重试,critical evidence retry 与两类 Fallback 边界全部落地。
|
||||||
|
- `DiagnosisGraphFactory` 继续独占技术重试计数、evidence retry 计数和 Planner mode/reset;Node 不重复拥有编排计数。
|
||||||
|
- 新增源码中无 `ThreadLocal`、`VerifierInputHook`、`tool_trace_summary`、`raw_executor`、TODO 或 FIXME;protocol 包无 Graph/Hook/ChatService 反向依赖。
|
||||||
|
- `ChatService` 无 Diagnosis Graph/CompiledGraph 引用,生产切换保持在阶段 3;DB migration、Trace DTO/entity/repository 和 prompts 均无 diff。
|
||||||
|
|
||||||
|
### Automated verification
|
||||||
|
|
||||||
|
- 新 protocol/Node/真实 CompiledGraph/Router/Trace focused suite:通过。
|
||||||
|
- 旧 `ChatServiceSequentialAgentTest`、`VerifierInputHookTest`、`ExecutorGatekeeperServiceTest` 回归:通过。
|
||||||
|
- 合计 100 tests,0 failures,0 errors,0 skipped。
|
||||||
|
- `mvn -q -DskipTests test-compile`:通过。
|
||||||
|
- `openspec validate chat-diagnosis-stategraph-real-nodes --strict`:通过。
|
||||||
|
- `openspec validate --specs --strict`:13 passed,0 failed。
|
||||||
|
- `git diff --check`:通过;仅报告 Git 既有 LF/CRLF 转换提示,无 whitespace error。
|
||||||
|
|
||||||
|
### Deferred final verification
|
||||||
|
|
||||||
|
- 阶段 2 按用户确认不运行 Maven live E2E,不启动应用,不检查 `logs/`,不查询数据库。
|
||||||
|
- 上述端到端、日志和 `scripts/query_mysql.py` 数据库核验统一保留到阶段 5 全部实现完成后执行。
|
||||||
@@ -0,0 +1,25 @@
|
|||||||
|
# Chat Diagnosis StateGraph Real Nodes Evidence
|
||||||
|
|
||||||
|
## 证据
|
||||||
|
|
||||||
|
| 来源 | 证据 | 结论 | 是否已汇报 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `VerifierInputHook.java`、`ChatService.java` | Executor/Verifier/Composer 解析与安全渲染原实现 | 抽取到中立 protocol 并让旧路径委托,避免双真理源 | 是 |
|
||||||
|
| `ExecutorGatekeeperService.java` | `validateRun` 与 checked binding/matched_text 契约 | Graph Gatekeeper 必须按当前 runId 单次调用并 fail closed | 是 |
|
||||||
|
| 本地 Graph/ReactAgent 1.1.2.0 API | `ReactAgent.call(String,RunnableConfig)` 与 config-aware Graph action | 最小 invoker 可精确透传输入和当前 Run metadata | 是 |
|
||||||
|
| 阶段 0/1 OpenSpec archives | 路由、计数、状态与安全边界冻结 | 阶段 2 不改变 Graph counter 所有权或生产入口 | 是 |
|
||||||
|
| 新 protocol/Node/CompiledGraph tests | PASS、REJECT、LOW_CONFID、critical retry、固定输入重试和安全 fallback | 真实 Nodes 的可观察路径和材料边界已覆盖 | 是 |
|
||||||
|
| 旧 Sequential/Hook/Gatekeeper tests | 共享抽取后的旧路径回归 | 阶段 2 未破坏当前生产控制流 | 是 |
|
||||||
|
|
||||||
|
## Evidence-driven 结论
|
||||||
|
|
||||||
|
- Graph Verifier 不得注册旧 `VerifierInputHook`,否则会双执行 Gatekeeper 并泄漏完整 tool trace。
|
||||||
|
- Verified Input 必须用 claim/invocation/tool/path 精确关联 passed binding;不能复刻 Gatekeeper 判断或读取未引用工具结果。
|
||||||
|
- Router 与 Retry Prepare 必须共用 critical-gap 提取规则:仅 `is_critical=true` 的 `no_evidence`/`indirect_support`。
|
||||||
|
- 技术重试输入必须字节一致,且不得重跑前序 Node;计数仍由 Graph factory 统一拥有。
|
||||||
|
- 当前 ChatService 无 Graph 引用,DB/Trace/Prompt 无 diff,满足阶段 2 的生产隔离要求。
|
||||||
|
|
||||||
|
## 风险与后续证据
|
||||||
|
|
||||||
|
- 本阶段使用 fake invoker 和 focused tests,不证明真实模型/外部基础设施联通;阶段 5 最终 E2E 统一补证。
|
||||||
|
- Prompt 输入说明、生产装配、Run/Trace 持久化属于阶段 3,不能提前从阶段 2 证据推断已完成。
|
||||||
@@ -0,0 +1,69 @@
|
|||||||
|
# Chat Diagnosis StateGraph Routing Skeleton Acceptance
|
||||||
|
|
||||||
|
## 结果
|
||||||
|
|
||||||
|
已接受。阶段 1 完成未接生产入口的 Graph 骨架和 Fake Node 路由体系。
|
||||||
|
|
||||||
|
## 验证
|
||||||
|
|
||||||
|
### 静态验证
|
||||||
|
|
||||||
|
- 命令:`git diff --check`
|
||||||
|
- 结果:passed。
|
||||||
|
- 检查:production refs、forbidden deps、23 changed paths 白名单。
|
||||||
|
- 结果:passed,outside refs=0,forbidden refs=0,out-of-scope=0。
|
||||||
|
|
||||||
|
### 脚本验证
|
||||||
|
|
||||||
|
- 命令:`mvn -q "-Dtest=DiagnosisGraphRoutingTest,DiagnosisOrchestrationTraceBuilderTest" test`
|
||||||
|
- 结果:passed,35 tests(29 routing + 6 trace),0 failure/error。
|
||||||
|
- 覆盖:正常、三类技术 retry、Executor 不重试、Gatekeeper、evidence retry、Composer、unknown fail-closed、threadId、Append events 和 trace。
|
||||||
|
|
||||||
|
- 命令:`mvn -q "-DskipTests" test`
|
||||||
|
- 结果:passed。
|
||||||
|
- 覆盖:main/test compilation。
|
||||||
|
|
||||||
|
- 命令:`openspec validate chat-diagnosis-stategraph-routing-skeleton --type change --strict --json`
|
||||||
|
- 结果:passed,1/1。
|
||||||
|
|
||||||
|
- 命令:`openspec validate --specs --strict --json`
|
||||||
|
- 结果:passed,12/12(归档前)。
|
||||||
|
|
||||||
|
### 浏览器/人工验证
|
||||||
|
|
||||||
|
- 结果:not run。
|
||||||
|
- 原因:无 UI 或生产入口变化。
|
||||||
|
|
||||||
|
### 未验证
|
||||||
|
|
||||||
|
- Maven E2E:not run,用户要求仅阶段 5 全部实现后统一执行。
|
||||||
|
- `logs/`:not inspected,保留到阶段 5。
|
||||||
|
- 数据库:not queried,保留到阶段 5。
|
||||||
|
- 真实模型/Agent:not invoked,属于阶段 2。
|
||||||
|
|
||||||
|
## 已完成范围
|
||||||
|
|
||||||
|
- 显式 Graph Core direct dependency。
|
||||||
|
- 26 state keys、typed status、topology/route/reason constants。
|
||||||
|
- 八个 config-aware action ports 和真实 CompiledGraph factory。
|
||||||
|
- Planner/Verifier/Composer 技术计数、一次 evidence retry、recursion limit 32。
|
||||||
|
- fail-closed router、events Append 和 trace builder。
|
||||||
|
- 35 个 Fake Node/trace tests。
|
||||||
|
|
||||||
|
## 已知限制
|
||||||
|
|
||||||
|
- Skeleton 没有生产消费者,阶段 3 才切换 ChatService。
|
||||||
|
- Fake Node 只证明控制流,不证明真实 Agent JSON、Gatekeeper 或安全输入映射。
|
||||||
|
- orchestration trace 尚未持久化或通过 API 暴露。
|
||||||
|
|
||||||
|
## Bug 修复和诊断
|
||||||
|
|
||||||
|
- 架构审计发现并修正 Graph Core 传递依赖所有权,改为 BOM 管理的直接依赖。
|
||||||
|
- 自查补齐全部正常节点 threadId、Verifier REJECT 和缺失 effective verdict 场景。
|
||||||
|
|
||||||
|
## 交接
|
||||||
|
|
||||||
|
- 下一步:阶段 1 Git commit;完成后才能创建阶段 2 change。
|
||||||
|
- Delta sync:新增 `chat-diagnosis-stategraph-routing-skeleton` 主 spec,共 6 requirements。
|
||||||
|
- OpenSpec 归档确认:用户已要求每阶段 archive,授权已存在。
|
||||||
|
- OpenSpec 归档结果:已同步主 spec,并归档到 `openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-routing-skeleton/`。
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
# Chat Diagnosis StateGraph Routing Skeleton Brief
|
||||||
|
|
||||||
|
## 背景
|
||||||
|
|
||||||
|
- 用户目标:ISS-011 阶段 1 独立完成 Graph 骨架和 Fake Node 路由测试,archive/commit 后才能进入真实 Node 阶段。
|
||||||
|
- 当前问题:阶段 0 只有设计基线,仓库此前没有可编译 StateGraph 或条件边验证。
|
||||||
|
- 关联 OpenSpec:`openspec/changes/chat-diagnosis-stategraph-routing-skeleton/`
|
||||||
|
- devflow 分档:complex
|
||||||
|
- 前置基线:`openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-design-freeze/`
|
||||||
|
|
||||||
|
## 范围
|
||||||
|
|
||||||
|
- 本次要做:Graph Core 直接依赖、状态/枚举、config-aware action ports、deterministic router、Graph factory、有限计数、events/trace builder、Fake Node tests。
|
||||||
|
- 本次不做:真实 Agent/Gatekeeper、ChatService、Hook、DB、Trace API、旧实现清理、Maven E2E。
|
||||||
|
- 影响区域:`pom.xml`、`com.superbiz.agent.graph.diagnosis`、对应 test 包。
|
||||||
|
|
||||||
|
## OpenSpec 对齐
|
||||||
|
|
||||||
|
- proposal 覆盖状态:已覆盖 skeleton-only 目标、L2 边界、测试与 E2E 非目标。
|
||||||
|
- specs 覆盖状态:6 个 requirements 覆盖编译、Planner、Executor/Gatekeeper、Verifier、Composer、trace。
|
||||||
|
- tasks 覆盖状态:12/12 完成。
|
||||||
@@ -0,0 +1,112 @@
|
|||||||
|
# Chat Diagnosis StateGraph Routing Skeleton Decisions
|
||||||
|
|
||||||
|
## Question Pool
|
||||||
|
|
||||||
|
| # | 维度 | 问题 | 模式 | 状态 |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| Q1 | 术语 | 阶段 1 的 Graph State、event、trace 含义从哪里继承? | evidence-driven | 已解决 |
|
||||||
|
| Q2 | 边界 | 阶段 1 是否接入 ChatService 或真实 Agent/Service? | evidence-driven | 已解决 |
|
||||||
|
| Q3 | 技术 | 锁定 1.1.2.0 是否支持所需 Node/Edge、策略和 recursion limit? | evidence-driven | 已解决 |
|
||||||
|
| Q4 | 技术 | Node action port 是否需要 RunnableConfig? | evidence-driven | 已解决 |
|
||||||
|
| Q5 | 验收 | 阶段 1 是否有必要添加单元测试? | evidence-driven + user rule | 已解决 |
|
||||||
|
| Q6 | 接口 | 新骨架的接口影响等级与消费者是什么? | evidence-driven | 已解决 |
|
||||||
|
| Q7 | 循环 | recursion limit 应取多少,是否替代业务计数? | evidence-driven | 已解决 |
|
||||||
|
| Q8 | 阶段 | 是否在本 change 实现真实 nodes、DB、Trace API 或旧代码清理? | user-interview(已由六阶段口径确认) | 已解决 |
|
||||||
|
|
||||||
|
## Evidence-driven
|
||||||
|
|
||||||
|
| 结论 | 证据来源 | 是否已汇报用户 |
|
||||||
|
|---|---|---|
|
||||||
|
| 阶段 1 必须完整继承阶段 0 baseline | archived design/spec/ADR | 已汇报 |
|
||||||
|
| 本阶段只做 Fake Node 骨架,不接真实模型 | ISS-011 阶段 1/2 边界 | 已汇报 |
|
||||||
|
| 1.1.2.0 支持 AsyncNodeActionWithConfig、AsyncEdgeActionWithConfig、conditional edges、Replace/Append、recursionLimit 和 threadId | 本地 JAR `javap` / `javap -c` | 已汇报 |
|
||||||
|
| AppendStrategy 将 list/collection 追加为有序列表,可用于 node terminal events | 本地 AppendStrategy bytecode | 已汇报 |
|
||||||
|
| 路由是新增行为且分支多,单元测试有必要 | 阶段 1 完成标准与用户“必要则加”规则 | 已汇报 |
|
||||||
|
| 最坏合法路径少于 20 次 Node 执行,limit 32 有安全余量 | 冻结路由矩阵的路径计数 | 已汇报 |
|
||||||
|
| 当前仓库没有 Graph 包或实现 | `rg` / package 目录核查 | 已汇报 |
|
||||||
|
| 生产代码将直接 import Graph Core,必须从传递依赖提升为 BOM 管理的直接依赖 | pom 与 dependency tree / 架构审计 | 已汇报 |
|
||||||
|
|
||||||
|
## User-interview
|
||||||
|
|
||||||
|
| 问题原文 | 用户原话 | 确认状态 | OpenSpec 回写 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| 阶段 1 是否应独立执行 sm-flow? | “iss-011里每个阶段,都是一个sm-flow” | 已确认 | 本 change 独立边界 |
|
||||||
|
| 阶段 1 是否执行 E2E? | “端到端只在最后阶段全部完成后才验证” | 已确认 | Out of Scope / Acceptance |
|
||||||
|
| 阶段 1 是否添加单元测试? | “如果有必要添加单元测试验收的话,就加” | 已确认规则;本阶段判定必要 | Acceptance |
|
||||||
|
| 是否可提前实现阶段 2–5? | “每个阶段需要归档完并提交才能进入下一个阶段” | 已确认不可提前 | Out of Scope |
|
||||||
|
|
||||||
|
## Context And Handoff
|
||||||
|
|
||||||
|
- 前置 archive:`openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-design-freeze/`
|
||||||
|
- 前置 commit:`581daff`
|
||||||
|
- 当前 change:`chat-diagnosis-stategraph-routing-skeleton`
|
||||||
|
- 后续 change:`chat-diagnosis-stategraph-real-nodes`,只能在本阶段 archive + commit 后创建。
|
||||||
|
|
||||||
|
## Technical Decisions
|
||||||
|
|
||||||
|
### Package and API
|
||||||
|
|
||||||
|
- 新包:`com.superbiz.agent.graph.diagnosis`。
|
||||||
|
- `pom.xml` 显式声明 Graph Core,版本继续由现有 BOM 管理。
|
||||||
|
- Graph 工厂接收 config-aware Node action ports,不依赖 Spring Bean 或真实 Agent。
|
||||||
|
- Fake actions 只放 `src/test`。
|
||||||
|
- 状态读取集中在 typed helper/router,避免各 edge 复制字符串解析。
|
||||||
|
|
||||||
|
### Retry ownership
|
||||||
|
|
||||||
|
- Planner/Verifier/Composer Node wrapper 根据进入节点前的上次技术失败状态增加各自 retry count。
|
||||||
|
- 第一次失败时 count=0,允许 self-loop;重入后 count=1,第二次失败直接 Fallback。
|
||||||
|
- Evidence Retry Node 将 run-level evidence count 增加一次并重置 planner retry count。
|
||||||
|
- Edge 只选择 route,不隐式修改状态。
|
||||||
|
|
||||||
|
### Event and trace
|
||||||
|
|
||||||
|
- 每个 Fake/后续真实 Node 返回一个 event list,AppendStrategy 负责累积。
|
||||||
|
- event 包含 node/outcome/reasonCode/attempt,不包含 payload。
|
||||||
|
- trace builder 由相邻 events 生成 transition;final node/reason 来自最后 event。
|
||||||
|
- events 为空时拒绝构造 trace,避免伪造路径。
|
||||||
|
|
||||||
|
### Interface impact
|
||||||
|
|
||||||
|
- 等级:L2 internal interface。
|
||||||
|
- 新消费者:阶段 2 Node adapters、阶段 3 ChatService orchestrator、阶段 4 tests。
|
||||||
|
- 外部 API/DB/运行路径:无变化。
|
||||||
|
- 回滚:revert 本阶段提交即可;因为未接生产入口,没有数据迁移。
|
||||||
|
- 兼容:后续 action adapters 必须实现已冻结 ports,不得传递父 State 全量数据。
|
||||||
|
|
||||||
|
## Risks Accepted
|
||||||
|
- Stage 1 skeleton 会暂时存在但未被生产调用,这是阶段边界要求,不是死代码最终状态。
|
||||||
|
- 真实 Agent 状态映射尚未验证,由阶段 2 独立 change 负责。
|
||||||
|
- Maven E2E 不运行,生产路径完全未改变。
|
||||||
|
|
||||||
|
## Apply Evidence
|
||||||
|
|
||||||
|
- Task 1.1:显式声明 BOM 管理的 Graph Core;新增 26 个 state keys、状态枚举、拓扑/route/reason 常量、默认 Replace + events Append 策略和安全 typed reads。
|
||||||
|
- Task 1.2:新增不可变 event/transition/trace records,构造期拒绝空 routing metadata 和非法计数,map 输出只有冻结字段。
|
||||||
|
- Task 2.1:新增八个 non-null config-aware action ports,Fake/真实 Node 共用同一 Graph 接入面。
|
||||||
|
- Task 2.2:新增纯 router;所有未知/缺失状态 fail closed,LOW_CONFID guard 只读 facts_checked,不构造 retry context。
|
||||||
|
- Task 2.3:Graph Factory 注册八节点和全部条件边;wrapper 独占 retry/evidence control state,compile recursion limit=32。
|
||||||
|
- 首模块 Maven compile:passed(31s),锁定 1.1.2.0 API 假设成立。
|
||||||
|
- Task 3.1:trace builder 从 event 单向派生 transitions/final reason/degraded/evidence count,空或异类 events 显式失败。
|
||||||
|
- Task 4.1:新增严格 FIFO Fake Node fixture,记录 sequence/calls/threadId,每 attempt 只追加一个 terminal event,意外调用立即失败。
|
||||||
|
- Task 4.2:真实 CompiledGraph normal/Planner/Executor/Gatekeeper focused tests 首轮通过(Maven exit 0,约 67s)。
|
||||||
|
- Task 4.3:Verifier/evidence/Composer 路由与独立计数测试通过(Maven exit 0,约 12s)。
|
||||||
|
- Task 4.4:routing + trace focused suite 通过(Maven exit 0,约 28s),覆盖事件顺序、degraded、map 白名单、不可变性和非法输入。
|
||||||
|
- Task 5.1:35 tests(29 routing + 6 trace)全通过;Maven test compilation、change strict 1/1、主 specs 12/12、diff check 均通过。
|
||||||
|
- Task 5.2:现有 production refs=0,新 Graph 对真实 Service/DB/Trace refs=0,23 个 changed paths 全部命中阶段白名单;Maven E2E/log/DB 按用户口径保留到阶段 5。
|
||||||
|
- Review:补齐所有正常节点 threadId 传播、Verifier REJECT→Composer 和 completed-without-verdict fail-closed 用例;增强后 focused suite exit 0。
|
||||||
|
|
||||||
|
## Cross-Artifact 对齐检查
|
||||||
|
|
||||||
|
| 上游 → 下游 | 检查内容 | 状态 |
|
||||||
|
|---|---|---|
|
||||||
|
| brief/prd → proposal | 阶段 1 目标、Fake Node 边界、测试口径和 E2E 非目标 | 已对齐 |
|
||||||
|
| proposal → 设计产物 | direct dependency、状态、ports、router、counter、trace、测试架构 | 已对齐 |
|
||||||
|
| 设计产物 → specs/tasks | 所有可观察路由、终止、安全默认和实现模块 | 已对齐 |
|
||||||
|
| specs → tasks | 编译、全路由、trace、focused tests、生产隔离检查 | 已对齐 |
|
||||||
|
|
||||||
|
Gap:无。
|
||||||
|
|
||||||
|
## Architecture Audit
|
||||||
|
|
||||||
|
输入是 run-scoped 初始 state 与 `RunnableConfig`,处理链是 CompiledGraph → config-aware action ports → deterministic router/counters,输出是最终 state 与纯路由 trace;阶段 1 没有 Controller/Service/DB 消费者。Graph State 由单次 invoke 所有,events 只由 Node append,transitions 只由 builder 派生,避免双写。阶段 2 只实现 action ports,阶段 3 才将 orchestrator 交给 ChatService,因此当前未接生产入口是刻意生命周期边界。审计发现的唯一缺口是 Graph Core 直接依赖所有权,已回写 proposal/design/tasks。与阶段 0 archive 无冲突,L2 风险可由 Fake Node CompiledGraph tests 和 revert 单提交控制。
|
||||||
@@ -0,0 +1,31 @@
|
|||||||
|
# Chat Diagnosis StateGraph Routing Skeleton Evidence
|
||||||
|
|
||||||
|
## 证据
|
||||||
|
|
||||||
|
| 来源 | 证据 | 结论 | 是否已汇报 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| 阶段 0 archive/ADR | 冻结 26 个 state keys、完整路由、四类计数、run audit | 阶段 1 实现未偏离基线 | 是 |
|
||||||
|
| 本地 Graph Core 1.1.2.0 `javap` | config-aware Node/Edge、conditional edge、recursion limit、threadId 存在 | 使用公共锁定 API 可编译 | 是 |
|
||||||
|
| AppendStrategy bytecode | list/collection 按顺序追加 | 每 Node 返回单 event list 可形成实际路径 | 是 |
|
||||||
|
| Maven compile | `mvn -q "-DskipTests" compile` exit 0 | direct dependency 和 Factory API 编译成立 | 是 |
|
||||||
|
| Fake Node CompiledGraph tests | 29 routing tests,0 failure/error | 全条件边、retry、threadId、unknown fail-closed 成立 | 是 |
|
||||||
|
| Trace tests | 6 tests,0 failure/error | transitions、degraded、map 白名单、不可变和非法输入成立 | 是 |
|
||||||
|
| Test compilation | `mvn -q "-DskipTests" test` exit 0 | 全测试源可编译 | 是 |
|
||||||
|
| OpenSpec validation | change 1/1、主 specs 12/12 strict | artifacts 与既有规格无回归 | 是 |
|
||||||
|
| 生产隔离检查 | outside Graph refs=0,Graph 对真实 Service/DB/Trace refs=0 | 阶段 1 未接生产入口 | 是 |
|
||||||
|
| Git 路径白名单 | 23 changed paths,out-of-scope=0 | 无跨阶段文件混入 | 是 |
|
||||||
|
|
||||||
|
## Evidence-driven 结论
|
||||||
|
|
||||||
|
- 结论:直接声明 Graph Core 是正确依赖所有权。
|
||||||
|
- 证据:生产代码直接 import Graph Core;依赖原先仅由 Agent Framework 传递。
|
||||||
|
- 风险:BOM 升级仍需重新跑真实 Graph tests。
|
||||||
|
- 用户确认:不需要,属于构建稳健性。
|
||||||
|
- 结论:recursion limit 32 足够且没有替代业务计数。
|
||||||
|
- 证据:最坏合法路径低于 20;第二次技术失败和第二次 LOW_CONFID tests 均终止。
|
||||||
|
- 风险:未来新增循环必须重算。
|
||||||
|
- 用户确认:不需要,冻结业务上限未变。
|
||||||
|
- 结论:单元测试必要,E2E 不必要。
|
||||||
|
- 证据:阶段新增条件边行为但未接生产入口;35 tests 直接验证 Graph。
|
||||||
|
- 风险:真实 Agent 映射仍留给阶段 2。
|
||||||
|
- 用户确认:符合用户按必要性和最终阶段 E2E 规则。
|
||||||
@@ -0,0 +1,50 @@
|
|||||||
|
# Chat Diagnosis StateGraph Test Suite 验收
|
||||||
|
|
||||||
|
## 结果
|
||||||
|
|
||||||
|
已接受。OpenSpec tasks 21/21 完成,阶段 4为 test-only,无生产行为变更。
|
||||||
|
|
||||||
|
## 验证
|
||||||
|
|
||||||
|
### 静态验证
|
||||||
|
|
||||||
|
- `git diff --check`:通过。
|
||||||
|
- `git diff --name-only HEAD -- src/main`:0 个文件。
|
||||||
|
- test inventory:三个权威类均存在;`ChatServiceSequentialAgentTest`/`VerifierInputHookTest` 均不存在;`ScriptedDiagnosisGraphActions` 定义 1 处。
|
||||||
|
- `openspec validate chat-diagnosis-stategraph-test-suite --strict`:通过。
|
||||||
|
- `openspec validate --specs --strict`:15 passed,0 failed。
|
||||||
|
|
||||||
|
### 脚本验证
|
||||||
|
|
||||||
|
- 三层权威 suite:4 suites / 43 tests,0 failures/errors/skipped。
|
||||||
|
- 完整保留安全回归:31 suites / 126 tests,0 failures/errors/skipped。
|
||||||
|
- `mvn -q -DskipTests test-compile`:通过。
|
||||||
|
|
||||||
|
### 浏览器/人工验证
|
||||||
|
|
||||||
|
- 未运行;阶段 4仅重构自动化测试,且用户要求最终人工/live 验收到阶段 5统一执行。
|
||||||
|
|
||||||
|
### 未验证
|
||||||
|
|
||||||
|
- 未使用 Maven 启动应用做 live E2E。
|
||||||
|
- 未检查 `logs/`。
|
||||||
|
- 未执行 `scripts/query_mysql.py`。
|
||||||
|
- 剩余风险:真实模型、工具、Flyway/MySQL 和最终 Trace 内容仍需阶段 5 E2E/log/DB 证据。
|
||||||
|
|
||||||
|
## 已完成范围
|
||||||
|
|
||||||
|
- 建立 Workflow、Node Contract、Chat Integration 三层权威测试体系。
|
||||||
|
- 补齐 ceiling/second LOW_CONFID、tool-failure legal snapshot、partial-pass REJECT 和 Verifier invalid status 边界。
|
||||||
|
- 删除旧 Hook implementation test,并保留 parser/Gatekeeper/projection/Composer 等安全回归。
|
||||||
|
- 用结构测试和 source gate 防止旧 Sequential/Hook 测试回归。
|
||||||
|
|
||||||
|
## 已知限制
|
||||||
|
|
||||||
|
- 生产 `VerifierInputHook`/`VerifierContextHolder` 类型仍存在但已无生产/权威测试消费者;阶段 5清理。
|
||||||
|
- component tests 与权威层存在有意的分层 overlap,详见 `evidence.md`。
|
||||||
|
|
||||||
|
## 交接
|
||||||
|
|
||||||
|
- OpenSpec archive:已同步 2 份 delta specs,并归档到 `openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-test-suite/`。
|
||||||
|
- 下一步:审查精确 Git diff并完成阶段 4独立提交,之后启动阶段 5。
|
||||||
|
- OpenSpec 归档确认:用户已授权直接执行后续归档;归档已完成。
|
||||||
@@ -0,0 +1,20 @@
|
|||||||
|
# Chat Diagnosis StateGraph Test Suite Brief
|
||||||
|
|
||||||
|
## 背景
|
||||||
|
|
||||||
|
- 用户目标:将 ISS-011 阶段 4作为独立 sm-flow,建立以 Graph 路径和外部行为为中心的新测试体系。
|
||||||
|
- 当前问题:阶段 1–3 覆盖充分但组织分散,旧 Hook payload test 仍会约束已退出生产的隐式状态机。
|
||||||
|
- 关联 OpenSpec:`openspec/changes/chat-diagnosis-stategraph-test-suite/`
|
||||||
|
- devflow 分档:complex;接口影响 L1 test-only。
|
||||||
|
|
||||||
|
## 范围
|
||||||
|
|
||||||
|
- 本次要做:Workflow/Node Contract/Chat Integration 三层权威入口、Issue 路径矩阵补齐、旧 Hook test 退役、保留安全回归。
|
||||||
|
- 本次不做:不改生产源码,不删除生产 Hook/ThreadLocal 类型,不运行 live E2E/log/DB 验收。
|
||||||
|
- 影响区域:Diagnosis Graph tests、ChatService integration test、OpenSpec/devflow。
|
||||||
|
|
||||||
|
## OpenSpec 对齐
|
||||||
|
|
||||||
|
- proposal/design/specs:已覆盖,2 个 delta capabilities。
|
||||||
|
- tasks:21/21 已完成。
|
||||||
|
- 生产行为:无变化,`src/main` diff=0。
|
||||||
@@ -0,0 +1,121 @@
|
|||||||
|
# Chat Diagnosis StateGraph Test Suite Decisions
|
||||||
|
|
||||||
|
## Entry Summary
|
||||||
|
|
||||||
|
- 问题:阶段 1–3 的安全覆盖已充足,但测试命名/组织仍是逐步实现产物,尚未形成 ISS-011 指定的 Workflow、Node Contract、Chat Integration 三层权威体系。
|
||||||
|
- 期望:阶段 4只重构测试结构并补齐矩阵,不改变生产行为;归档并提交后才进入阶段 5。
|
||||||
|
- 分档:complex(路径矩阵广、涉及旧安全测试退役,但接口影响为 L1 test-only)。
|
||||||
|
- Change:`chat-diagnosis-stategraph-test-suite`。
|
||||||
|
- 授权:用户已要求直接实现,阶段门禁与阶段 5 才 live E2E 的约束不变。
|
||||||
|
|
||||||
|
## Context Sources
|
||||||
|
|
||||||
|
- ISS-011 阶段 4、测试策略和验收标准。
|
||||||
|
- `chat-diagnosis-stategraph-design-freeze` 的 test migration requirement。
|
||||||
|
- 阶段 1–3 archive/acceptance 和当前 119-test focused baseline。
|
||||||
|
- `DiagnosisGraphRoutingTest`、各 Node tests、`DiagnosisRealGraphIntegrationTest`、`ChatDiagnosisGraphRuntimeTest`、`ChatServiceGraphIntegrationTest`。
|
||||||
|
- `VerifierInputHookTest`、protocol parser tests、`ExecutorGatekeeperServiceTest` 与 Trace/Controller/Repository/Eval tests。
|
||||||
|
|
||||||
|
## Question Pool
|
||||||
|
|
||||||
|
| # | 维度 | 问题 | 模式 | 状态 |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| Q1 | 术语 | Workflow、Node Contract、Chat Integration 的职责边界是什么? | evidence-driven | 已解决 |
|
||||||
|
| Q2 | 边界 | 是否把所有现有 Graph unit tests 合并成三个巨型类? | evidence-driven | 已解决 |
|
||||||
|
| Q3 | 旧测试 | `VerifierInputHookTest` 在显式 Graph Gatekeeper 后应保留、改写还是删除? | evidence-driven | 已解决 |
|
||||||
|
| Q4 | 验收 | 如何证明阶段 4矩阵完整且没有恢复固定 Sequential 顺序? | evidence-driven | 已解决 |
|
||||||
|
| Q5 | 阶段 | 是否允许为测试可测性改生产代码或执行 live E2E? | user-interview(既有冻结规则) | 已确认 |
|
||||||
|
|
||||||
|
## Evidence-driven Findings
|
||||||
|
|
||||||
|
- Q1:Issue 已明确三类测试;现有 `DiagnosisGraphRoutingTest` 对应 Workflow,分散 Node/Protocol tests 对应 Node Contract,`ChatServiceGraphIntegrationTest` 对应外部生命周期。
|
||||||
|
- Q2:现有细粒度测试失败定位清晰,全部合并会制造大文件;应保留专用 unit tests,同时新增/重命名三层权威入口并共享夹具。
|
||||||
|
- Q3:生产 Graph Verifier 已不注册 Hook,Hook test 仍验证 raw/full-trace/ThreadLocal payload,与 verified-only 生产协议冲突;parser/Gatekeeper/投影行为已有独立 tests,故阶段 4删除 Hook test,生产类型留到阶段 5。
|
||||||
|
- Q4:以 Issue 必需路径清单建立 requirement-to-test matrix;source check 禁止 `SequentialAgent`/`VerifierInputHookTest` 成为新 suite 依赖,并运行保留安全回归。
|
||||||
|
|
||||||
|
## User-interview Confirmation
|
||||||
|
|
||||||
|
| 问题 | 用户原话/既有确认 | 状态 | OpenSpec 回写 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| Q5 阶段边界 | “端到端只在最后阶段全部完成后才验证;每个阶段如果有必要添加单元测试验收的话,就加” | 已确认 | proposal |
|
||||||
|
|
||||||
|
## Grill-with-docs Result
|
||||||
|
|
||||||
|
- 术语不进入业务 glossary:Workflow/Node Contract/Integration 是测试架构术语,不改变 Session、Run、Trace、Gatekeeper 或 evidence gap 领域定义。
|
||||||
|
- 具体场景压力测试:同 session 多 run 属于 Chat Integration;Gatekeeper REJECT/LOW_CONFID 与 retry exhaustion 属于 Workflow;failed binding 过滤与 verified-only payload 属于 Node Contract。
|
||||||
|
- 旧 Hook test 的有效行为已分别迁移到 `ExecutorEvidenceParserTest`、`ExecutorGatekeeperServiceTest`、`GatekeeperNodeTest` 和 `VerifiedInputNodeTest`;删除不会丢失安全真理源。
|
||||||
|
- 没有难以逆转的新架构取舍,不创建 ADR;测试组织可以在保持行为矩阵的前提下继续演进。
|
||||||
|
|
||||||
|
## Discover Status
|
||||||
|
|
||||||
|
- `devflow/index.md`:命中阶段 0–3 archive。
|
||||||
|
- 生产调用链:阶段 3 已冻结且测试阶段默认不修改。
|
||||||
|
- 接口影响:L1 test-only;无 API/DTO/DB/Prompt/运行时消费者变化。
|
||||||
|
- 未解决问题:0。
|
||||||
|
- Draft 产物:proposal + decisions;尚未生成 design/spec/tasks,尚未修改测试代码。
|
||||||
|
|
||||||
|
## Architecture Audit
|
||||||
|
|
||||||
|
### Test ownership map
|
||||||
|
|
||||||
|
`ISS-011 path matrix -> DiagnosisGraphWorkflowTest -> ScriptedDiagnosisGraphActions -> real DiagnosisGraphFactory` 负责控制流;`Node/protocol contracts -> DiagnosisGraphNodeContractTest + focused component tests -> real Node actions/parsers/Gatekeeper` 负责安全投影;`public lifecycle -> ChatServiceGraphIntegrationTest -> Run repositories/Eval/Trace mapping` 负责外部行为。Controller、Repository、Trace、Eval 和 protocol tests 是三层体系的下游安全消费者,不应被重写成 Graph 内部顺序断言。
|
||||||
|
|
||||||
|
### Data and lifecycle ownership
|
||||||
|
|
||||||
|
- Workflow fixtures 只拥有脚本状态、调用计数和 events,不创建 Session/Run。
|
||||||
|
- Node Contract fixtures 只拥有 invoker input/output 和 mock Gatekeeper current-run result,不持久化生产实体。
|
||||||
|
- Chat Integration fixtures 只观察 ChatService public result 和 current Run persistence,不推断内部 Node 次序。
|
||||||
|
- `VerifierInputHookTest` 删除后不产生数据契约缺口:Executor parser、Gatekeeper、passed-binding projection 各自已有单一测试所有者。
|
||||||
|
|
||||||
|
### Coupling risks
|
||||||
|
|
||||||
|
- 最大风险是 Workflow 与 Node Contract 都断言完整路径而重复;设计将 Fake route matrix 与 real-node input/security matrix分开。
|
||||||
|
- `ScriptedDiagnosisGraphActions` 是唯一 Fake topology fixture;不新增第二套 Graph builder。
|
||||||
|
- 测试-only阶段禁止 `src/main` diff,避免为测试便利扩大 production API。
|
||||||
|
- 类重命名使用 Git rename,旧名称仅允许出现在 OpenSpec/devflow迁移说明中。
|
||||||
|
|
||||||
|
### Cross-artifact alignment
|
||||||
|
|
||||||
|
| 上游 → 下游 | 检查内容 | 状态 |
|
||||||
|
|---|---|---|
|
||||||
|
| brief/proposal → proposal | 三层体系、旧 Hook test退役、保留安全回归、阶段 5 E2E 延期 | 已对齐 |
|
||||||
|
| proposal → design | rename-not-copy、职责归属、production diff=0、rollback | 已对齐 |
|
||||||
|
| design → specs/tasks | Workflow/Node/Integration矩阵、Hook test删除、source inventory和验证门禁 | 已对齐 |
|
||||||
|
| specs → tasks | 每类 scenario 均有 rename、补缺、回归或静态验收任务 | 已对齐 |
|
||||||
|
|
||||||
|
### Audit result
|
||||||
|
|
||||||
|
架构审计未发现业务 glossary、阶段 0–3 specs 或生产行为冲突。新 capability 只描述测试验证系统,modified design-freeze requirement 完整保留并细化旧测试替换边界。接口影响保持 L1,cross-artifact gap=0,无需回写生产 spec 或创建 ADR。
|
||||||
|
|
||||||
|
## Commit Gate
|
||||||
|
|
||||||
|
- schema:spec-driven;proposal/design/2 delta specs/tasks 全部 done,applyRequires=`tasks` 已满足。
|
||||||
|
- OpenSpec:当前 change strict pass;15 个主 specs strict pass。
|
||||||
|
- Cross-artifact:4/4 已对齐,gap=0。
|
||||||
|
- Question pool:4 个 evidence-driven 已查证,1 个 user-interview 由用户既有原话确认,无未决项。
|
||||||
|
- Interface impact:L1 test-only;design 有独立影响/回滚章节,默认 `src/main` diff=0。
|
||||||
|
- Preflight:`git diff --check` 通过,当前仅 Draft OpenSpec/devflow 与根目录计划文件,无测试/生产代码修改。
|
||||||
|
- 结论:Draft OpenSpec 达到可执行状态,创建 `.committed` 后进入 Apply。
|
||||||
|
|
||||||
|
## Apply Completion
|
||||||
|
|
||||||
|
- `DiagnosisGraphRoutingTest` 已以 Git rename 演进为 `DiagnosisGraphWorkflowTest`;所有 scripted run 统一断言 events 与真实 sequence 完全一致。
|
||||||
|
- `DiagnosisRealGraphIntegrationTest` 已演进为 `DiagnosisGraphNodeContractTest`;补齐合法工具失败限制、partial-pass REJECT 和 Verifier invalid 无伪 verdict。
|
||||||
|
- Workflow 显式断言 Gatekeeper ceiling 将模型 PASS 限制为 LOW_CONFID 且不触发 evidence retry,以及第二次 LOW_CONFID 不再补证据。
|
||||||
|
- `ChatServiceGraphIntegrationTest` 增加 SessionContextHolder finally cleanup 断言;public Run/Trace/Eval/multi-run 覆盖保持。
|
||||||
|
- `VerifierInputHookTest` 已删除;生产 Hook/ThreadLocal 类型未修改,留待阶段 5清理。
|
||||||
|
- 新增 `DiagnosisGraphTestSuiteStructureTest`,保证三个权威类存在、Sequential/Hook实现测试不存在且权威测试不引用旧实现。
|
||||||
|
|
||||||
|
## Verification Summary
|
||||||
|
|
||||||
|
- 新权威层:4 suites / 43 tests,0 failures/errors/skipped。
|
||||||
|
- 完整保留安全回归:31 suites / 126 tests,0 failures/errors/skipped。
|
||||||
|
- Maven test compilation:通过。
|
||||||
|
- OpenSpec:当前 change strict pass;主 specs 15/15 strict pass。
|
||||||
|
- 静态门禁:`git diff --check` 通过;required authoritative classes=3;legacy tests=0;scripted fixture definitions=1;`src/main` diff=0。
|
||||||
|
- 按阶段门禁未运行 Maven live E2E、未检查 `logs/`、未执行 `scripts/query_mysql.py`;统一保留到阶段 5。
|
||||||
|
|
||||||
|
## Archive Result
|
||||||
|
|
||||||
|
- 2 份 delta specs 已同步:新增 test-suite capability 6 条 requirements,修改 design-freeze test migration requirement 1 条。
|
||||||
|
- OpenSpec 已归档到 `openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-test-suite/`。
|
||||||
@@ -0,0 +1,42 @@
|
|||||||
|
# Chat Diagnosis StateGraph Test Suite Evidence
|
||||||
|
|
||||||
|
## Requirement-to-test Matrix
|
||||||
|
|
||||||
|
| ISS-011 路径/边界 | 权威测试 | 辅助证据 |
|
||||||
|
|---|---|---|
|
||||||
|
| PASS 与精确 event/transition | `DiagnosisGraphWorkflowTest.normalPassPathUsesCompiledGraphAndPreservesEventOrder` | `DiagnosisGraphNodeContractTest.compiledRealNodeGraphCompletesPassPathInExactOrder` |
|
||||||
|
| Planner 一次技术重试/耗尽/非重试失败 | `DiagnosisGraphWorkflowTest.planner*` | Planner adapter focused test |
|
||||||
|
| Executor FAILED/TOOL_BLOCKED/INVALID/no-evidence | `DiagnosisGraphWorkflowTest.executor*` | `DiagnosisGraphNodeContractTest.legalSnapshotAfterToolFailureRemainsCompletedAndReachesGatekeeper` |
|
||||||
|
| Gatekeeper PASS/LOW/REJECT/unknown | `DiagnosisGraphWorkflowTest.gatekeeper*` | Gatekeeper Node/service tests |
|
||||||
|
| partial-pass REJECT 不泄漏 claim | Workflow unsafe route | `DiagnosisGraphNodeContractTest.gatekeeperRejectSkipsVerifierAndUsesPreVerificationFallback` |
|
||||||
|
| verified-only passed-binding projection | Workflow VerifiedInput route | `VerifiedInputNodeTest` + Verifier adapter test |
|
||||||
|
| Verifier 技术重试/耗尽/非重试失败 | `DiagnosisGraphWorkflowTest.verifier*` | `DiagnosisGraphNodeContractTest.verifierInvalidOutputSetsExecutionStatusWithoutFabricatedVerdict` |
|
||||||
|
| critical evidence retry、无 gap、non-critical、ceiling、第二次 LOW | `DiagnosisGraphWorkflowTest.*LowConfidence*` / `secondLowConfidence*` | EvidenceRetryPrepareNodeTest + real full-snapshot integration |
|
||||||
|
| 第二轮完整 snapshot 再验真 | Workflow evidence retry | `DiagnosisGraphNodeContractTest.criticalGapPerformsOneIncrementalRoundAndRevalidatesCompleteSnapshot` |
|
||||||
|
| Verifier REJECT 安全表达 | `DiagnosisGraphWorkflowTest.verifierRejectStillRunsComposerWithSafeMaterial` | ComposerSafeInputBuilderTest |
|
||||||
|
| Composer retry/耗尽/非重试/安全 Fallback | `DiagnosisGraphWorkflowTest.composer*` | ComposerNodeAdapterTest + FallbackNodeTest |
|
||||||
|
| ChatResult/Run/Trace/evaluation/cleanup | `ChatServiceGraphIntegrationTest` | Controller/Trace/Repository/Eval tests |
|
||||||
|
| 同 session 多 run 隔离 | `ChatServiceGraphIntegrationTest.sameSessionCreatesDistinctRunIdsAndKeepsTracePerRun` | Run repository/Trace exact-run tests |
|
||||||
|
| 三层结构与旧实现测试退役 | `DiagnosisGraphTestSuiteStructureTest` | source inventory command |
|
||||||
|
|
||||||
|
## 旧 Hook test 安全映射
|
||||||
|
|
||||||
|
| 旧行为 | 新真理源 |
|
||||||
|
|---|---|
|
||||||
|
| fenced/prefixed/malformed Executor JSON | `ExecutorEvidenceParserTest` |
|
||||||
|
| invocation/raw_path/evidence excerpt真实性 | `ExecutorGatekeeperServiceTest` |
|
||||||
|
| Gatekeeper status/severity/audit | `GatekeeperNodeTest` |
|
||||||
|
| passed binding 精确投影 | `VerifiedInputNodeTest` |
|
||||||
|
| Verifier verified-only payload | `VerifierNodeAdapterTest` + `ChatVerifierPromptContractTest` |
|
||||||
|
|
||||||
|
## 验证证据
|
||||||
|
|
||||||
|
- 31 suites / 126 tests:0 failures、0 errors、0 skipped。
|
||||||
|
- test compilation、当前 change strict、主 specs 15/15、diff check 全通过。
|
||||||
|
- `src/main` diff=0;三个权威类存在;旧 Sequential/Hook tests不存在;Scripted fixture 定义唯一。
|
||||||
|
|
||||||
|
## Intentional overlap and limits
|
||||||
|
|
||||||
|
- Workflow 与 Node Contract 都覆盖 PASS/REJECT,但前者验证 route/event,后者验证真实输入/安全材料;这是分层证据,不是复制 fixture。
|
||||||
|
- 细粒度 component tests 保留以定位失败,不要求全部搬进三个权威类。
|
||||||
|
- 真实模型/工具、日志和数据库行为未在阶段 4验证,统一由阶段 5 live E2E承担。
|
||||||
+192
@@ -0,0 +1,192 @@
|
|||||||
|
<!doctype html>
|
||||||
|
<html lang="zh-CN">
|
||||||
|
<head>
|
||||||
|
<meta charset="utf-8">
|
||||||
|
<meta name="viewport" content="width=device-width, initial-scale=1">
|
||||||
|
<title>Spring AI Alibaba Graph:诊断编排改造</title>
|
||||||
|
<script src="https://cdn.jsdelivr.net/npm/mermaid/dist/mermaid.min.js"></script>
|
||||||
|
<style>
|
||||||
|
:root { --bg:#f4f7fb; --card:rgba(255,255,255,.86); --text:#172033; --muted:#5c6880; --blue:#2563eb; --line:#dbe3f0; }
|
||||||
|
* { box-sizing:border-box; }
|
||||||
|
body { margin:0; font-family:Inter,"PingFang SC","Microsoft YaHei",sans-serif; background:linear-gradient(135deg,#eef4ff,#f8fafc); color:var(--text); line-height:1.75; }
|
||||||
|
.wrap { max-width:1000px; margin:0 auto; padding:42px 24px 110px; }
|
||||||
|
header,.card { background:var(--card); border:1px solid rgba(255,255,255,.9); box-shadow:0 12px 35px rgba(35,55,90,.09); backdrop-filter:blur(14px); border-radius:20px; padding:28px; margin-bottom:22px; }
|
||||||
|
h1 { margin:0 0 8px; font-size:34px; }
|
||||||
|
h2 { margin-top:0; color:#173c85; }
|
||||||
|
h3 { color:#244c92; }
|
||||||
|
code,pre { font-family:"Cascadia Code",Consolas,monospace; }
|
||||||
|
pre { background:#101827; color:#e5edf9; padding:18px; border-radius:14px; overflow:auto; }
|
||||||
|
.tag { display:inline-block; padding:4px 10px; margin-right:6px; border-radius:999px; background:#e4edff; color:#2453a6; font-size:13px; }
|
||||||
|
.mnemonic-card { background:#fff8cf; border:2px dashed #e6b800; padding:16px; border-radius:14px; }
|
||||||
|
.fission-section { background:#fff1f2; border-left:5px solid #e11d48; padding:18px; border-radius:12px; }
|
||||||
|
.truth { border-left:4px solid #2563eb; background:#eff6ff; padding:14px; border-radius:10px; }
|
||||||
|
details { background:#f8fafc; border:1px solid var(--line); border-radius:12px; padding:12px 15px; margin:9px 0; }
|
||||||
|
summary { cursor:pointer; font-weight:700; }
|
||||||
|
.search { position:fixed; bottom:22px; left:50%; transform:translateX(-50%); width:min(720px,calc(100% - 36px)); background:rgba(15,23,42,.93); padding:12px; border-radius:16px; box-shadow:0 15px 40px rgba(0,0,0,.25); z-index:5; }
|
||||||
|
.search input { width:100%; border:0; outline:0; border-radius:10px; padding:12px 14px; font-size:15px; }
|
||||||
|
.hidden { display:none !important; }
|
||||||
|
ul,ol { padding-left:24px; }
|
||||||
|
</style>
|
||||||
|
</head>
|
||||||
|
<body>
|
||||||
|
<div class="wrap" id="content-area">
|
||||||
|
<header>
|
||||||
|
<h1>Spring AI Alibaba Graph:诊断编排改造</h1>
|
||||||
|
<p>作者:叫我小杨同学的小码酱</p>
|
||||||
|
<span class="tag">StateGraph</span><span class="tag">Agent 编排</span><span class="tag">条件边</span><span class="tag">故障诊断</span>
|
||||||
|
</header>
|
||||||
|
|
||||||
|
<section class="card">
|
||||||
|
<h2>0. 核心摘要</h2>
|
||||||
|
<p><strong>让 ReactAgent 继续负责做事,让 StateGraph 负责下一步去哪里。</strong></p>
|
||||||
|
<p>生活类比:Planner、Executor、Verifier 是医院科室,Graph 是分诊和转诊制度。</p>
|
||||||
|
<p class="truth">官方核心模型是 State、Nodes、Edges。本地依赖 1.1.2.0 已确认支持 StateGraph、条件边、编译配置、中断和 threadId。</p>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="card">
|
||||||
|
<h2>1. 概念破冰</h2>
|
||||||
|
<div class="mnemonic-card">状态记事实,节点做任务,边管下一步,检查点管恢复。</div>
|
||||||
|
<p>当前 ChatService 已经是半个状态机:SequentialAgent 运行 Planner、Executor、Verifier,外层 Java 再根据 PASS、LOW_CONFID、REJECT 决定 Composer 或重试。Graph 改造的价值,是把分散的控制权显式化。</p>
|
||||||
|
<pre>当前:SequentialAgent + 外层 if/else
|
||||||
|
目标:StateGraph 条件边 + ReactAgent 语义节点</pre>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="card">
|
||||||
|
<h2>2. 深度解析</h2>
|
||||||
|
<p>Supervisor 适合动态选择专科 Agent;Gatekeeper、Verifier 等强制门禁应由代码边控制。普通工具失败也不需要 interrupt,只有等待人工输入或审批时才暂停。</p>
|
||||||
|
<div class="mermaid">
|
||||||
|
flowchart TD
|
||||||
|
S["START"] --> P["Planner"]
|
||||||
|
P --> E["Executor"]
|
||||||
|
E -- "结构有效" --> G["Gatekeeper"]
|
||||||
|
E -- "阻断或非法" --> F["Fallback"]
|
||||||
|
G -- "允许" --> V["Verifier"]
|
||||||
|
G -- "拒绝" --> F
|
||||||
|
V -- "通过或拒绝" --> C["Composer"]
|
||||||
|
V -- "低置信且有预算" --> R["Retry Guard"]
|
||||||
|
R -- "补证据" --> P
|
||||||
|
R -- "停止" --> C
|
||||||
|
C --> X["END"]
|
||||||
|
F --> X
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<h3>状态设计</h3>
|
||||||
|
<p>保存统一诊断上下文、计划、Executor 结构化输出、Gatekeeper 结果、Verifier verdict、重试轮次和最终结果。不要在 State 里复制所有原始日志或完整思考过程。</p>
|
||||||
|
|
||||||
|
<h3>核心伪代码</h3>
|
||||||
|
<pre>StateGraph graph = new StateGraph("diagnosis_workflow", strategies)
|
||||||
|
.addNode("planner", plannerNode)
|
||||||
|
.addNode("executor", executorNode)
|
||||||
|
.addNode("gatekeeper", gatekeeperNode)
|
||||||
|
.addNode("verifier", verifierNode)
|
||||||
|
.addNode("retry_guard", retryGuardNode)
|
||||||
|
.addNode("composer", composerNode)
|
||||||
|
.addNode("fallback", fallbackNode)
|
||||||
|
.addEdge(START, "planner")
|
||||||
|
.addEdge("planner", "executor")
|
||||||
|
.addConditionalEdges("executor", routeAfterExecutor,
|
||||||
|
Map.of("gatekeeper","gatekeeper",
|
||||||
|
"retry_guard","retry_guard",
|
||||||
|
"fallback","fallback"))
|
||||||
|
.addConditionalEdges("gatekeeper", routeAfterGatekeeper,
|
||||||
|
Map.of("verifier","verifier","fallback","fallback"))
|
||||||
|
.addConditionalEdges("verifier", routeAfterVerifier,
|
||||||
|
Map.of("retry_guard","retry_guard",
|
||||||
|
"composer","composer",
|
||||||
|
"fallback","fallback"))
|
||||||
|
.addConditionalEdges("retry_guard", routeAfterRetry,
|
||||||
|
Map.of("planner","planner","composer","composer"))
|
||||||
|
.addEdge("composer", END)
|
||||||
|
.addEdge("fallback", END);</pre>
|
||||||
|
|
||||||
|
<h3>运行边界</h3>
|
||||||
|
<pre>RunnableConfig config = RunnableConfig.builder()
|
||||||
|
.threadId(runId)
|
||||||
|
.addMetadata("sessionId", sessionId)
|
||||||
|
.addMetadata("runId", runId)
|
||||||
|
.build();</pre>
|
||||||
|
<p>一个 diagnosis_run 使用一个 Graph thread,避免同一 session 下多个 run 共享检查点。MemorySaver 不是跨重启持久化。</p>
|
||||||
|
|
||||||
|
<h3>ReactAgent 的接入</h3>
|
||||||
|
<p>本地 ReactAgent 提供 asNode(boolean, boolean)。当前项目第一阶段更适合用适配节点调用已有 Agent,显式控制输入、outputKey 和解析;状态契约稳定后再评估直接 asNode。</p>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="card fission-section">
|
||||||
|
<h2>3. 深度裂变</h2>
|
||||||
|
<h3>🔍 搜索内化:改造的是控制权,不是 Agent</h3>
|
||||||
|
<p>Graph 的节点可以是 LLM,也可以是普通 Java 代码;ReactAgent 本身已经是子图。所谓“Multi-Agent 改 Graph”,实际上是把跨 Agent 状态转换交给父 Graph。</p>
|
||||||
|
<p>官方页面示例有 OverAllStaste 拼写错误,实际类型是 OverAllState。网站主分支可能领先于本地依赖,最终必须以项目 JAR 和编译测试为准。</p>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="card">
|
||||||
|
<h2>4. 实战指南</h2>
|
||||||
|
<ol>
|
||||||
|
<li>先定义节点结果状态,不改 Prompt。</li>
|
||||||
|
<li>把现有 Planner、Executor、Verifier 包装为 Node。</li>
|
||||||
|
<li>Gatekeeper 和固定降级做成 Java Node。</li>
|
||||||
|
<li>迁移现有两轮 LOW_CONFID 控制。</li>
|
||||||
|
<li>加入 Executor 阻断、非法结构和 Gatekeeper REJECT 条件边。</li>
|
||||||
|
<li>保留旧 Sequential 链路作为短期回退。</li>
|
||||||
|
<li>最后再增加 HITL、并行和专科 SubAgent。</li>
|
||||||
|
</ol>
|
||||||
|
<h3>避坑</h3>
|
||||||
|
<ul>
|
||||||
|
<li>不要让 Supervisor 决定是否跳过安全门禁。</li>
|
||||||
|
<li>不要把 no_evidence 当成 Executor 失败。</li>
|
||||||
|
<li>不要让多个 run 共用 sessionId 作为 Graph threadId。</li>
|
||||||
|
<li>不要把 MemorySaver 当成生产持久化。</li>
|
||||||
|
<li>不要未验证 messages 传播就直接大量使用 asNode(true, true)。</li>
|
||||||
|
</ul>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="card">
|
||||||
|
<h2>5. 温故知新</h2>
|
||||||
|
<h3>FAQ</h3>
|
||||||
|
<details><summary>1. Graph 会替代 ReactAgent 吗?</summary><p>不会,ReactAgent 可以作为子图节点继续使用。</p></details>
|
||||||
|
<details><summary>2. 为什么不用 Supervisor 控制失败?</summary><p>失败跳转是确定性规则,不需要增加一次模型决策。</p></details>
|
||||||
|
<details><summary>3. no_evidence 是否直接 fallback?</summary><p>不一定,合法 no-evidence 引用仍要经过 Gatekeeper 和 Verifier。</p></details>
|
||||||
|
<details><summary>4. Composer 必须是 Agent 吗?</summary><p>正常表达可以使用轻量 Agent,系统失败要保留固定模板。</p></details>
|
||||||
|
<details><summary>5. threadId 用什么?</summary><p>当前数据模型下优先使用 runId,sessionId 作为元数据。</p></details>
|
||||||
|
<details><summary>6. 何时需要 Checkpointer?</summary><p>需要暂停、恢复和检查 Graph 历史状态时。</p></details>
|
||||||
|
<details><summary>7. 能直接使用 agent.asNode 吗?</summary><p>可以,但要验证 outputKey、messages 和父子检查点。</p></details>
|
||||||
|
<details><summary>8. AIOps 要单独 Graph 吗?</summary><p>入口和输出策略独立,诊断核心可以共用。</p></details>
|
||||||
|
|
||||||
|
<h3>自测题</h3>
|
||||||
|
<ol>
|
||||||
|
<li>为什么 Executor 工具阻断不应由 Verifier 决定重试?</li>
|
||||||
|
<li>ReplaceStrategy 和 AppendStrategy 各适合什么状态?</li>
|
||||||
|
<li>为什么合法 no_evidence 仍然需要 Gatekeeper?</li>
|
||||||
|
<li>Supervisor 与条件边的决策权有什么不同?</li>
|
||||||
|
<li>为什么 Graph threadId 更适合使用 runId?</li>
|
||||||
|
<li>什么情况下才应该配置 interruptBefore?</li>
|
||||||
|
</ol>
|
||||||
|
</section>
|
||||||
|
</div>
|
||||||
|
|
||||||
|
<div class="search"><input id="search-input" placeholder="搜索本文内容……"></div>
|
||||||
|
<script>
|
||||||
|
mermaid.initialize({ startOnLoad: true, theme: 'neutral' });
|
||||||
|
window.onload = function() {
|
||||||
|
const input = document.getElementById('search-input');
|
||||||
|
if(!input) return;
|
||||||
|
input.addEventListener('input', (e) => {
|
||||||
|
const term = e.target.value.toLowerCase().trim();
|
||||||
|
const contentArea = document.getElementById('content-area');
|
||||||
|
const blocks = contentArea.querySelectorAll('p, li, blockquote, .fission-section, .mnemonic-card, details, .mermaid');
|
||||||
|
if(term.length === 0) {
|
||||||
|
blocks.forEach(el => el.classList.remove('hidden'));
|
||||||
|
document.querySelectorAll('h1, h2, h3').forEach(el => el.classList.remove('hidden'));
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
blocks.forEach(el => el.classList.add('hidden'));
|
||||||
|
document.querySelectorAll('h1, h2, h3').forEach(el => el.classList.add('hidden'));
|
||||||
|
blocks.forEach(el => {
|
||||||
|
if(el.innerText.toLowerCase().includes(term)) {
|
||||||
|
el.classList.remove('hidden');
|
||||||
|
}
|
||||||
|
});
|
||||||
|
});
|
||||||
|
};
|
||||||
|
</script>
|
||||||
|
</body>
|
||||||
|
</html>
|
||||||
+383
@@ -0,0 +1,383 @@
|
|||||||
|
# Spring AI Alibaba Graph:诊断编排改造
|
||||||
|
|
||||||
|
作者:叫我小杨同学的小码酱
|
||||||
|
标签:Spring AI Alibaba、StateGraph、Agent 编排、条件边、故障诊断
|
||||||
|
|
||||||
|
## 0. 核心摘要
|
||||||
|
|
||||||
|
一句话:让 ReactAgent 继续负责“做事”,让 StateGraph 负责“下一步去哪里”。
|
||||||
|
|
||||||
|
生活类比:Planner、Executor、Verifier 是医院里的不同科室,Graph 是分诊和转诊制度;不能让某个科室自己决定跳过检验和会诊。
|
||||||
|
|
||||||
|
真理锚点:官方文档将 Graph 概括为 State、Nodes、Edges,核心关系是“节点完成工作,边决定下一步做什么”。当前项目依赖的 `spring-ai-alibaba-graph-core:1.1.2.0` 本地 JAR 已确认提供 `StateGraph.addConditionalEdges(...)`、`CompileConfig.interruptBefore/After(...)` 和 `RunnableConfig.threadId(...)`。
|
||||||
|
|
||||||
|
## 1. 概念破冰
|
||||||
|
|
||||||
|
> 巧记:状态记事实,节点做任务,边管下一步,检查点管恢复。
|
||||||
|
|
||||||
|
当前 `ChatService` 已经像一个半成品状态机:`SequentialAgent` 固定执行 Planner、Executor、Verifier,外层 Java 循环再判断 PASS、LOW_CONFID、REJECT,并决定 Composer 或下一轮。问题不是 Agent 不够多,而是状态转换分散在 `SequentialAgent`、Hook 和外层 `if/else` 中。
|
||||||
|
|
||||||
|
```text
|
||||||
|
当前
|
||||||
|
SequentialAgent: Planner -> Executor -> Verifier
|
||||||
|
|
|
||||||
|
ChatService 外层: PASS / LOW_CONFID / REJECT -> Composer / retry
|
||||||
|
|
||||||
|
目标
|
||||||
|
StateGraph 显式表示所有阶段和条件边
|
||||||
|
ReactAgent 作为图中的语义节点继续复用
|
||||||
|
```
|
||||||
|
|
||||||
|
## 2. 深度解析
|
||||||
|
|
||||||
|
### 2.1 为什么不是直接换成 SupervisorAgent
|
||||||
|
|
||||||
|
SupervisorAgent 适合在多个专科 Agent 之间动态选择,例如 Database、Redis、JVM。当前诊断链路中的 Gatekeeper、Verifier 是不能随意跳过的质量门禁。如果让 LLM Supervisor 决定下一步,关键流程会从代码控制变成模型决策。
|
||||||
|
|
||||||
|
当前真正需要的是确定性路由:
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TD
|
||||||
|
S["START"] --> P["Planner"]
|
||||||
|
P --> E["Executor"]
|
||||||
|
E -- "证据结构有效" --> G["Gatekeeper"]
|
||||||
|
E -- "执行阻断或非法输出" --> F["Fallback"]
|
||||||
|
G -- "PASS 或 LOW_CONFID" --> V["Verifier"]
|
||||||
|
G -- "REJECT" --> F
|
||||||
|
V -- "PASS" --> C["Composer"]
|
||||||
|
V -- "LOW_CONFID 且有预算" --> R["Retry Guard"]
|
||||||
|
V -- "REJECT 或无预算" --> C
|
||||||
|
R -- "允许补证据" --> P
|
||||||
|
R -- "停止" --> C
|
||||||
|
C --> X["END"]
|
||||||
|
F --> X
|
||||||
|
```
|
||||||
|
|
||||||
|
### 2.2 State 应保存什么
|
||||||
|
|
||||||
|
状态应该保存跨节点需要共享的原始事实和结构化结果,不保存拼好的 Prompt,也不保存无边界增长的模型思考过程。
|
||||||
|
|
||||||
|
建议的语义状态:
|
||||||
|
|
||||||
|
- `diagnosis_context`:入口适配后的统一诊断上下文。
|
||||||
|
- `planner_plan`:Planner 的结构化计划。
|
||||||
|
- `executor_output`:`executor_evidence_v2`。
|
||||||
|
- `executor_status`:成功、阻断、非法输出或失败。
|
||||||
|
- `gatekeeper_result`:代码验真结果。
|
||||||
|
- `verifier_output`、`verdict`:可推导性结果。
|
||||||
|
- `retry_context`、`round`:有限补证据状态。
|
||||||
|
- `final_answer`:最终表达。
|
||||||
|
- `failure_reason`:确定性的失败原因。
|
||||||
|
|
||||||
|
现有 Trace 已由 `agent_step` 和 `tool_invocation` 持久化,Graph State 不需要复制所有原始日志。
|
||||||
|
|
||||||
|
### 2.3 KeyStrategy 如何选择
|
||||||
|
|
||||||
|
诊断状态大多使用 `ReplaceStrategy`,因为每个阶段产生当前轮的最新结果。只有确实需要累计的轻量事件列表才使用 `AppendStrategy`。
|
||||||
|
|
||||||
|
```java
|
||||||
|
KeyStrategyFactory diagnosisStateStrategies() {
|
||||||
|
return () -> {
|
||||||
|
Map<String, KeyStrategy> strategies = new HashMap<>();
|
||||||
|
strategies.put("diagnosis_context", new ReplaceStrategy());
|
||||||
|
strategies.put("planner_plan", new ReplaceStrategy());
|
||||||
|
strategies.put("executor_output", new ReplaceStrategy());
|
||||||
|
strategies.put("executor_status", new ReplaceStrategy());
|
||||||
|
strategies.put("gatekeeper_result", new ReplaceStrategy());
|
||||||
|
strategies.put("verifier_output", new ReplaceStrategy());
|
||||||
|
strategies.put("verdict", new ReplaceStrategy());
|
||||||
|
strategies.put("retry_context", new ReplaceStrategy());
|
||||||
|
strategies.put("round", new ReplaceStrategy());
|
||||||
|
strategies.put("final_answer", new ReplaceStrategy());
|
||||||
|
strategies.put("failure_reason", new ReplaceStrategy());
|
||||||
|
return strategies;
|
||||||
|
};
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### 2.4 ReactAgent 怎样放进 Graph
|
||||||
|
|
||||||
|
本地 `1.1.2.0` 的 `ReactAgent` 提供 `asNode(boolean includeContents, boolean returnReasoningContents)`,可以直接作为子图节点。但当前项目 Planner、Executor、Verifier 的输入组织和输出解析已经有较多定制,第一阶段更推荐使用适配节点显式调用现有 Agent:
|
||||||
|
|
||||||
|
```java
|
||||||
|
var plannerNode = node_async((state, config) -> {
|
||||||
|
DiagnosisContext context = requireContext(state);
|
||||||
|
String prompt = plannerInput(context, state.value("retry_context").orElse(null));
|
||||||
|
|
||||||
|
AssistantMessage response = plannerAgent.call(prompt, childConfig(config, "planner"));
|
||||||
|
PlannerPlan plan = plannerParser.parse(extractText(response));
|
||||||
|
|
||||||
|
return Map.of(
|
||||||
|
"planner_plan", plan,
|
||||||
|
"failure_reason", ""
|
||||||
|
);
|
||||||
|
});
|
||||||
|
```
|
||||||
|
|
||||||
|
适配节点的好处是不会意外把父图全部 `messages` 注入所有 Agent,也能继续复用当前解析器、Hook、Prompt 和 ToolCallback。
|
||||||
|
|
||||||
|
等统一状态契约稳定后,可以评估:
|
||||||
|
|
||||||
|
```java
|
||||||
|
graph.addNode("planner", plannerAgent.asNode(false, false));
|
||||||
|
```
|
||||||
|
|
||||||
|
但需要先验证父子图的 `messages`、outputKey 和 Checkpointer 是否符合预期。
|
||||||
|
|
||||||
|
### 2.5 Executor 节点只报告状态,不决定路由
|
||||||
|
|
||||||
|
```java
|
||||||
|
var executorNode = node_async((state, config) -> {
|
||||||
|
try {
|
||||||
|
PlannerPlan plan = requirePlan(state);
|
||||||
|
DiagnosisContext context = requireContext(state);
|
||||||
|
|
||||||
|
AssistantMessage response = executorAgent.call(
|
||||||
|
executorInput(context, plan),
|
||||||
|
childConfig(config, "executor")
|
||||||
|
);
|
||||||
|
|
||||||
|
ExecutorEvidence output = executorParser.parse(extractText(response));
|
||||||
|
|
||||||
|
if (!output.isStructurallyValid()) {
|
||||||
|
return Map.of(
|
||||||
|
"executor_status", "INVALID_OUTPUT",
|
||||||
|
"failure_reason", "executor_evidence_v2 解析失败"
|
||||||
|
);
|
||||||
|
}
|
||||||
|
|
||||||
|
// no_evidence 仍然是合法结构,需要交给 Gatekeeper 验证真实引用。
|
||||||
|
return Map.of(
|
||||||
|
"executor_status", "COMPLETED",
|
||||||
|
"executor_output", output
|
||||||
|
);
|
||||||
|
}
|
||||||
|
catch (ToolCapabilityBlockedException e) {
|
||||||
|
return Map.of(
|
||||||
|
"executor_status", "TOOL_BLOCKED",
|
||||||
|
"failure_reason", e.getMessage()
|
||||||
|
);
|
||||||
|
}
|
||||||
|
catch (Exception e) {
|
||||||
|
return Map.of(
|
||||||
|
"executor_status", "FAILED",
|
||||||
|
"failure_reason", safeMessage(e)
|
||||||
|
);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
```
|
||||||
|
|
||||||
|
注意:`no_evidence` 不是 Executor 失败。当前项目已经用 `$.no_evidence` 表达“查询成功但无匹配证据”,它仍应进入 Gatekeeper 和 Verifier,防止被过度表达为“问题不存在”。
|
||||||
|
|
||||||
|
### 2.6 Gatekeeper 节点保持纯代码
|
||||||
|
|
||||||
|
```java
|
||||||
|
var gatekeeperNode = node_async((state, config) -> {
|
||||||
|
ExecutorEvidence output = requireExecutorOutput(state);
|
||||||
|
String runId = metadata(config, "runId");
|
||||||
|
|
||||||
|
GatekeeperResult result = executorGatekeeperService.validate(runId, output);
|
||||||
|
|
||||||
|
return Map.of(
|
||||||
|
"gatekeeper_result", result,
|
||||||
|
"gatekeeper_status", result.severity()
|
||||||
|
);
|
||||||
|
});
|
||||||
|
```
|
||||||
|
|
||||||
|
Gatekeeper 不需要改造成 Agent。它负责确定性引用验真,是 Graph 中的普通 Java Node。
|
||||||
|
|
||||||
|
### 2.7 条件边是改造核心
|
||||||
|
|
||||||
|
```java
|
||||||
|
StateGraph graph = new StateGraph("diagnosis_workflow", diagnosisStateStrategies())
|
||||||
|
.addNode("planner", plannerNode)
|
||||||
|
.addNode("executor", executorNode)
|
||||||
|
.addNode("gatekeeper", gatekeeperNode)
|
||||||
|
.addNode("verifier", verifierNode)
|
||||||
|
.addNode("retry_guard", retryGuardNode)
|
||||||
|
.addNode("composer", composerNode)
|
||||||
|
.addNode("fallback", fallbackNode)
|
||||||
|
|
||||||
|
.addEdge(START, "planner")
|
||||||
|
.addEdge("planner", "executor")
|
||||||
|
|
||||||
|
.addConditionalEdges(
|
||||||
|
"executor",
|
||||||
|
edge_async(state -> switch (stringValue(state, "executor_status")) {
|
||||||
|
case "COMPLETED" -> "gatekeeper";
|
||||||
|
case "INVALID_OUTPUT" -> retryAvailable(state) ? "retry_guard" : "fallback";
|
||||||
|
case "TOOL_BLOCKED", "FAILED" -> "fallback";
|
||||||
|
default -> "fallback";
|
||||||
|
}),
|
||||||
|
Map.of(
|
||||||
|
"gatekeeper", "gatekeeper",
|
||||||
|
"retry_guard", "retry_guard",
|
||||||
|
"fallback", "fallback"
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
.addConditionalEdges(
|
||||||
|
"gatekeeper",
|
||||||
|
edge_async(state -> switch (stringValue(state, "gatekeeper_status")) {
|
||||||
|
case "REJECT" -> "fallback";
|
||||||
|
default -> "verifier";
|
||||||
|
}),
|
||||||
|
Map.of("verifier", "verifier", "fallback", "fallback")
|
||||||
|
)
|
||||||
|
|
||||||
|
.addConditionalEdges(
|
||||||
|
"verifier",
|
||||||
|
edge_async(state -> switch (stringValue(state, "verdict")) {
|
||||||
|
case "LOW_CONFID" -> retryAvailable(state) ? "retry_guard" : "composer";
|
||||||
|
case "PASS", "REJECT" -> "composer";
|
||||||
|
default -> "fallback";
|
||||||
|
}),
|
||||||
|
Map.of(
|
||||||
|
"retry_guard", "retry_guard",
|
||||||
|
"composer", "composer",
|
||||||
|
"fallback", "fallback"
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
.addConditionalEdges(
|
||||||
|
"retry_guard",
|
||||||
|
edge_async(state -> shouldRetry(state) ? "planner" : "composer"),
|
||||||
|
Map.of("planner", "planner", "composer", "composer")
|
||||||
|
)
|
||||||
|
|
||||||
|
.addEdge("composer", END)
|
||||||
|
.addEdge("fallback", END);
|
||||||
|
```
|
||||||
|
|
||||||
|
### 2.8 编译和执行
|
||||||
|
|
||||||
|
```java
|
||||||
|
SaverConfig saverConfig = SaverConfig.builder()
|
||||||
|
.register(new MemorySaver())
|
||||||
|
.build();
|
||||||
|
|
||||||
|
CompileConfig compileConfig = CompileConfig.builder()
|
||||||
|
.recursionLimit(20)
|
||||||
|
.saverConfig(saverConfig)
|
||||||
|
.build();
|
||||||
|
|
||||||
|
CompiledGraph compiledGraph = graph.compile(compileConfig);
|
||||||
|
|
||||||
|
RunnableConfig runConfig = RunnableConfig.builder()
|
||||||
|
// 一个 diagnosis_run 对应一个 Graph thread,避免同一 session 下多个 run 混状态。
|
||||||
|
.threadId(runId)
|
||||||
|
.addMetadata("sessionId", sessionId)
|
||||||
|
.addMetadata("runId", runId)
|
||||||
|
.build();
|
||||||
|
|
||||||
|
Map<String, Object> initialState = Map.of(
|
||||||
|
"diagnosis_context", diagnosisContext,
|
||||||
|
"round", 1
|
||||||
|
);
|
||||||
|
|
||||||
|
Optional<OverAllState> finalState = compiledGraph.invoke(initialState, runConfig);
|
||||||
|
```
|
||||||
|
|
||||||
|
`MemorySaver` 只适合进程内检查点,不等价于重启后可恢复的持久化。当前项目已有数据库 Trace,可以先把 Graph 用于流程控制;只有真正需要跨进程暂停恢复时,再引入持久 Checkpointer 或显式恢复模型。
|
||||||
|
|
||||||
|
### 2.9 Chat 与 AIOps 怎样共用
|
||||||
|
|
||||||
|
入口适配不同,公共 Graph 接收统一 `DiagnosisContext`:
|
||||||
|
|
||||||
|
```java
|
||||||
|
DiagnosisContext chatContext = chatAdapter.from(question, history);
|
||||||
|
DiagnosisContext aiOpsContext = aiOpsAdapter.from(alertPayload);
|
||||||
|
|
||||||
|
DiagnosisResult chatResult = diagnosisGraph.execute(chatContext, sessionId, runId);
|
||||||
|
DiagnosisResult aiOpsResult = diagnosisGraph.execute(aiOpsContext, sessionId, runId);
|
||||||
|
|
||||||
|
return chatOutputAdapter.render(chatResult);
|
||||||
|
return aiOpsOutputAdapter.render(aiOpsResult);
|
||||||
|
```
|
||||||
|
|
||||||
|
Chat 的普通问答仍走轻量链路;复杂诊断才进入公共 Graph。AIOps 在入口阶段固定主告警范围,Graph 内部继续复用证据收集和验证。
|
||||||
|
|
||||||
|
### 2.10 人工中断怎么放
|
||||||
|
|
||||||
|
普通工具失败不需要 interrupt,条件边即可。只有确实需要等待人工输入或审批时才增加节点:
|
||||||
|
|
||||||
|
```java
|
||||||
|
CompileConfig compileConfig = CompileConfig.builder()
|
||||||
|
.saverConfig(saverConfig)
|
||||||
|
.interruptBefore("human_review")
|
||||||
|
.build();
|
||||||
|
```
|
||||||
|
|
||||||
|
恢复时使用相同 `threadId` 和 checkpoint 信息,并通过 `RunnableConfig.builder(oldConfig).resume()` 或状态更新接口继续。具体恢复协议需要结合当前版本做集成测试,不能只凭文档假设。
|
||||||
|
|
||||||
|
## 3. 深度裂变
|
||||||
|
|
||||||
|
<div class="fission-section">
|
||||||
|
|
||||||
|
### 🔍 搜索内化:真正的改造对象不是 Agent,而是控制权
|
||||||
|
|
||||||
|
官方文档和本地 `1.1.2.0` JAR 均证明 Graph 节点既可以是 LLM,也可以是普通 Java 代码,条件边由状态决定目标节点。`ReactAgent` 本身已经是一个子图,并提供 `asNode(...)` 适配能力。
|
||||||
|
|
||||||
|
因此“从 Multi-Agent 改成 Graph”并不准确。更准确的是:把跨 Agent 的状态转换从高层 Flow 抽出来,交给父 Graph;ReactAgent 继续作为子图存在。
|
||||||
|
|
||||||
|
文档页面示例存在 `OverAllStaste` 拼写错误,实际类名是 `OverAllState`。网站主分支可能领先于本地依赖,因此最终应以项目锁定版本的 JAR 签名和编译测试为准。
|
||||||
|
|
||||||
|
</div>
|
||||||
|
|
||||||
|
## 4. 实战指南
|
||||||
|
|
||||||
|
### 4.1 最小迁移顺序
|
||||||
|
|
||||||
|
1. 定义统一的节点结果状态,不先改 Prompt。
|
||||||
|
2. 把现有 Planner、Executor、Verifier 调用包装为 Graph Node。
|
||||||
|
3. 将 Gatekeeper 和固定降级模板做成普通 Java Node。
|
||||||
|
4. 先迁移当前两轮 LOW_CONFID 循环。
|
||||||
|
5. 为 Executor 阻断、非法结构、Gatekeeper REJECT 增加条件边。
|
||||||
|
6. 保留原 Sequential 链路作为回退,完成行为对比后再删除。
|
||||||
|
7. 最后再考虑 HITL、并行和专科 SubAgent。
|
||||||
|
|
||||||
|
### 4.2 常见反模式
|
||||||
|
|
||||||
|
- 把每个异常都交给 LLM Supervisor 决策。
|
||||||
|
- Graph State 存放所有原始日志和完整思考过程。
|
||||||
|
- 把 `no_evidence` 当作 Executor 执行失败。
|
||||||
|
- 同一 session 的多个 run 共用一个 Graph `threadId`。
|
||||||
|
- 一开始就设计几十个节点和完整 Incident 状态机。
|
||||||
|
- 未验证父子图消息传播就直接大量使用 `ReactAgent.asNode(true, true)`。
|
||||||
|
- 把 `MemorySaver` 当成生产级持久化。
|
||||||
|
|
||||||
|
### 4.3 ROI
|
||||||
|
|
||||||
|
收益:条件分支可见、失败可测试、门禁不可跳过、Trace 更容易与节点对齐。
|
||||||
|
代价:需要维护状态契约、条件边和父子图上下文,并增加 Graph 级测试。
|
||||||
|
判断标准:如果当前只有固定顺序且失败直接结束,SequentialAgent 更简单;当局部重试、降级、HITL 和多入口策略已经出现时,StateGraph 的控制收益开始超过复杂度。
|
||||||
|
|
||||||
|
## 5. 温故知新
|
||||||
|
|
||||||
|
### FAQ
|
||||||
|
|
||||||
|
1. **Graph 会替代 ReactAgent 吗?** 不会,ReactAgent 可以作为 Graph 节点或由适配节点调用。
|
||||||
|
2. **为什么不用 Supervisor 控制失败?** 失败跳转是确定性规则,不应增加一次 LLM 决策。
|
||||||
|
3. **no_evidence 是否直接走 fallback?** 不一定。合法的 no-evidence 引用仍要经过 Gatekeeper 和 Verifier。
|
||||||
|
4. **Composer 是否必须是 Agent?** 正常表达可以是轻量 Agent,系统失败场景应保留固定模板。
|
||||||
|
5. **threadId 用 sessionId 还是 runId?** 当前模型下优先用 runId,避免同一会话多次运行状态串扰。
|
||||||
|
6. **什么时候需要 Checkpointer?** 需要暂停、恢复、查看历史 Graph 状态时;普通 Trace 持久化不自动等于 Graph Checkpoint。
|
||||||
|
7. **能否直接使用 agent.asNode?** 可以,但需要验证输入、outputKey、messages 和父子 Checkpointer 行为。
|
||||||
|
8. **AIOps 是否要单独一张 Graph?** 可以先共用诊断核心,入口范围策略和输出报告保持独立。
|
||||||
|
|
||||||
|
### 自测题
|
||||||
|
|
||||||
|
1. 为什么 Executor 工具阻断不应该由 Verifier 判断是否重试?
|
||||||
|
2. `ReplaceStrategy` 和 `AppendStrategy` 在诊断状态中分别适合什么数据?
|
||||||
|
3. 为什么合法 `no_evidence` 仍然需要 Gatekeeper?
|
||||||
|
4. SupervisorAgent 和 StateGraph 条件边的决策权有什么区别?
|
||||||
|
5. 为什么当前 Graph `threadId` 更适合使用 runId?
|
||||||
|
6. 什么情况下才应该增加 `interruptBefore`?
|
||||||
|
|
||||||
|
### 参考资源
|
||||||
|
|
||||||
|
- https://java2ai.com/docs/frameworks/graph-core/core/core-library
|
||||||
|
- https://java2ai.com/docs/frameworks/graph-core/quick-start
|
||||||
|
- Spring AI Alibaba 本地依赖:`spring-ai-alibaba-graph-core:1.1.2.0`
|
||||||
|
- 当前项目:`ChatService`、`AiOpsService`、`ExecutorGatekeeperService`、`VerifierInputHook`
|
||||||
+5
-3
@@ -1,6 +1,6 @@
|
|||||||
# SuperBizAgent MVP 文档
|
# SuperBizAgent MVP 文档
|
||||||
|
|
||||||
**更新日期**:2026-07-10
|
**更新日期**:2026-07-20
|
||||||
|
|
||||||
本目录保存 MVP 阶段的架构、问题、演示、评测和数据表说明。当前材料按“当前入口”和“历史归档”拆开,避免把早期设计稿当成当前实现。
|
本目录保存 MVP 阶段的架构、问题、演示、评测和数据表说明。当前材料按“当前入口”和“历史归档”拆开,避免把早期设计稿当成当前实现。
|
||||||
|
|
||||||
@@ -10,6 +10,7 @@
|
|||||||
|---|---|
|
|---|---|
|
||||||
| [architecture/README.md](architecture/README.md) | 当前 MVP 架构入口 |
|
| [architecture/README.md](architecture/README.md) | 当前 MVP 架构入口 |
|
||||||
| [architecture/current-mvp-architecture.md](architecture/current-mvp-architecture.md) | 当前可运行系统架构 |
|
| [architecture/current-mvp-architecture.md](architecture/current-mvp-architecture.md) | 当前可运行系统架构 |
|
||||||
|
| [architecture/stategraph-runtime-architecture.md](architecture/stategraph-runtime-architecture.md) | 复杂 Chat StateGraph 运行时架构 |
|
||||||
| [architecture/interview-one-pager.md](architecture/interview-one-pager.md) | 面试一页式架构讲解 |
|
| [architecture/interview-one-pager.md](architecture/interview-one-pager.md) | 面试一页式架构讲解 |
|
||||||
| [architecture/agent-orchestration.md](architecture/agent-orchestration.md) | Agent 编排架构 |
|
| [architecture/agent-orchestration.md](architecture/agent-orchestration.md) | Agent 编排架构 |
|
||||||
| [architecture/executor-evidence-pipeline-refactor.md](architecture/executor-evidence-pipeline-refactor.md) | Executor 证据链路改造记录 |
|
| [architecture/executor-evidence-pipeline-refactor.md](architecture/executor-evidence-pipeline-refactor.md) | Executor 证据链路改造记录 |
|
||||||
@@ -39,6 +40,7 @@ mvp/
|
|||||||
architecture/
|
architecture/
|
||||||
README.md
|
README.md
|
||||||
current-mvp-architecture.md
|
current-mvp-architecture.md
|
||||||
|
stategraph-runtime-architecture.md
|
||||||
interview-one-pager.md
|
interview-one-pager.md
|
||||||
agent-orchestration.md
|
agent-orchestration.md
|
||||||
executor-evidence-pipeline-refactor.md
|
executor-evidence-pipeline-refactor.md
|
||||||
@@ -55,7 +57,6 @@ mvp/
|
|||||||
README.md
|
README.md
|
||||||
active/
|
active/
|
||||||
archived/
|
archived/
|
||||||
design-notes/
|
|
||||||
rag/
|
rag/
|
||||||
tables/
|
tables/
|
||||||
README.md
|
README.md
|
||||||
@@ -118,9 +119,10 @@ RAG
|
|||||||
|
|
||||||
## 归档说明
|
## 归档说明
|
||||||
|
|
||||||
历史材料分两类:
|
历史材料分三类:
|
||||||
|
|
||||||
- 旧架构文档:[architecture/archive/2026-07-05-legacy/](architecture/archive/2026-07-05-legacy/)
|
- 旧架构文档:[architecture/archive/2026-07-05-legacy/](architecture/archive/2026-07-05-legacy/)
|
||||||
- 本次文档清理归档:[archive/2026-07-09-doc-cleanup/](archive/2026-07-09-doc-cleanup/)
|
- 本次文档清理归档:[archive/2026-07-09-doc-cleanup/](archive/2026-07-09-doc-cleanup/)
|
||||||
|
- Run-only v2 和 StateGraph 收口后的旧材料:[archive/2026-07-20-doc-cleanup/](archive/2026-07-20-doc-cleanup/)
|
||||||
|
|
||||||
归档文档只用于追溯设计历史。当前实现和后续规划以 `architecture/`、`issues/README.md`、`tables/README.md` 和 OpenSpec/devflow 的最新记录为准。
|
归档文档只用于追溯设计历史。当前实现和后续规划以 `architecture/`、`issues/README.md`、`tables/README.md` 和 OpenSpec/devflow 的最新记录为准。
|
||||||
|
|||||||
+18
-16
@@ -1,6 +1,6 @@
|
|||||||
# MVP 架构文档
|
# MVP 架构文档
|
||||||
|
|
||||||
**更新日期**:2026-07-10
|
**更新日期**:2026-07-20
|
||||||
|
|
||||||
这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到:
|
这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到:
|
||||||
|
|
||||||
@@ -13,10 +13,11 @@
|
|||||||
| 文档 | 用途 |
|
| 文档 | 用途 |
|
||||||
|---|---|
|
|---|---|
|
||||||
| [current-mvp-architecture.md](current-mvp-architecture.md) | 当前可运行 MVP 的总体架构、链路、持久化和质量门禁 |
|
| [current-mvp-architecture.md](current-mvp-architecture.md) | 当前可运行 MVP 的总体架构、链路、持久化和质量门禁 |
|
||||||
|
| [stategraph-runtime-architecture.md](stategraph-runtime-architecture.md) | 最新复杂 Chat StateGraph 运行时权威快照,覆盖 Node 路由、状态边界、Run/Trace 持久化和验收层次 |
|
||||||
| [interview-one-pager.md](interview-one-pager.md) | 面试一页式架构讲解,包含总图、亮点、取舍和追问回答 |
|
| [interview-one-pager.md](interview-one-pager.md) | 面试一页式架构讲解,包含总图、亮点、取舍和追问回答 |
|
||||||
| [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat SequentialAgent、AIOps SupervisorAgent、工具边界 |
|
| [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat bounded StateGraph、AIOps SupervisorAgent、工具边界 |
|
||||||
| [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md) | Chat 证据链路当前数据契约,覆盖 Executor V2、Gatekeeper、Verifier、Composer、`evidence_refs` |
|
| [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md) | Chat 证据链路当前数据契约,覆盖 Executor V2、Gatekeeper、Verifier、Composer、`evidence_refs` |
|
||||||
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Gatekeeper、Verifier、Composer、评测基线组成的质量门禁 |
|
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、StateGraph、Trace、Gatekeeper、Verifier、Composer、评测基线组成的质量门禁 |
|
||||||
| [rag-architecture.md](rag-architecture.md) | RAG/知识检索新架构,覆盖 L0 hint、VectorStore 主路径、SDK fallback、证据追踪 |
|
| [rag-architecture.md](rag-architecture.md) | RAG/知识检索新架构,覆盖 L0 hint、VectorStore 主路径、SDK fallback、证据追踪 |
|
||||||
| [modular-rag-pipeline.md](modular-rag-pipeline.md) | `lookup_knowledge` 模块化 RAG 落地架构,覆盖 pipeline、fallback、evidence-first contract、trace |
|
| [modular-rag-pipeline.md](modular-rag-pipeline.md) | `lookup_knowledge` 模块化 RAG 落地架构,覆盖 pipeline、fallback、evidence-first contract、trace |
|
||||||
| [rag-eval-closure.md](rag-eval-closure.md) | RAG 评测闭环,覆盖 offline baseline、baseline diff、diagnosis eval 和 live acceptance |
|
| [rag-eval-closure.md](rag-eval-closure.md) | RAG 评测闭环,覆盖 offline baseline、baseline diff、diagnosis eval 和 live acceptance |
|
||||||
@@ -29,20 +30,21 @@
|
|||||||
|
|
||||||
## 当前架构一句话
|
## 当前架构一句话
|
||||||
|
|
||||||
SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,Chat 链路由 Gatekeeper 做引用真实性校验、Verifier 做可推导性判断、Composer 生成最终表达;多轮会话元数据落到 `chat_session`,每次诊断运行落到 `diagnosis_run`,步骤和工具明细通过 `agent_step.run_id`、`tool_invocation.run_id` 关联,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
|
SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:复杂 Chat 由有界 StateGraph 显式编排 Planner、Executor、Gatekeeper、Verified Input、Verifier、Composer 与安全 Fallback,Executor 通过工具收集日志、指标和知识库证据;多轮会话元数据落到 `chat_session`,每次诊断运行落到 `diagnosis_run`,步骤和工具明细通过 `run_id` 关联,Graph 路由摘要独立保存为 `orchestration_trace`,最终通过精确 Run Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
|
||||||
|
|
||||||
## 阅读顺序
|
## 阅读顺序
|
||||||
|
|
||||||
1. 先读 [current-mvp-architecture.md](current-mvp-architecture.md),理解系统边界和主链路。
|
1. 先读 [current-mvp-architecture.md](current-mvp-architecture.md),理解系统边界和主链路。
|
||||||
2. 面试前读 [interview-one-pager.md](interview-one-pager.md),准备 2-5 分钟讲解。
|
2. 再读 [stategraph-runtime-architecture.md](stategraph-runtime-architecture.md),理解复杂 Chat 的真实运行时和路由边界。
|
||||||
3. 再读 [agent-orchestration.md](agent-orchestration.md),理解当前 Agent 如何协作。
|
3. 面试前读 [interview-one-pager.md](interview-one-pager.md),准备 2-5 分钟讲解。
|
||||||
4. 接着读 [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md),理解 Chat 证据链路的数据结构和验真边界。
|
4. 再读 [agent-orchestration.md](agent-orchestration.md),理解当前 Agent 如何协作。
|
||||||
5. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。
|
5. 接着读 [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md),理解 Chat 证据链路的数据结构和验真边界。
|
||||||
6. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。
|
6. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。
|
||||||
7. 继续读 [modular-rag-pipeline.md](modular-rag-pipeline.md),看 `lookup_knowledge` 的模块化落地和 evidence-first contract。
|
7. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。
|
||||||
8. 再读 [rag-eval-closure.md](rag-eval-closure.md),看 RAG baseline 如何形成质量闭环。
|
8. 继续读 [modular-rag-pipeline.md](modular-rag-pipeline.md),看 `lookup_knowledge` 的模块化落地和 evidence-first contract。
|
||||||
9. 然后读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
|
9. 再读 [rag-eval-closure.md](rag-eval-closure.md),看 RAG baseline 如何形成质量闭环。
|
||||||
10. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
|
10. 然后读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
|
||||||
11. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
|
11. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
|
||||||
12. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
|
12. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
|
||||||
13. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
|
13. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
|
||||||
|
14. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
# Agent 编排架构
|
# Agent 编排架构
|
||||||
|
|
||||||
**更新日期**:2026-07-08
|
**更新日期**:2026-07-20
|
||||||
**状态**:当前可运行架构
|
**状态**:当前可运行架构
|
||||||
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
||||||
|
|
||||||
@@ -8,7 +8,7 @@
|
|||||||
|
|
||||||
旧版 Agent 架构把系统描述为 Supervisor、Planner、SubAgent、Verifier 的团队协作。当前 MVP 保留这个核心思想,但实现更收敛:
|
旧版 Agent 架构把系统描述为 Supervisor、Planner、SubAgent、Verifier 的团队协作。当前 MVP 保留这个核心思想,但实现更收敛:
|
||||||
|
|
||||||
- Chat 链路使用固定顺序工作流:`Planner -> Executor -> Gatekeeper -> Verifier -> Composer`。
|
- Chat 复杂诊断使用有递归上限的显式 StateGraph;正常路径是 `Planner -> Executor -> Gatekeeper -> Verified Input -> Verifier -> Composer`,条件边负责有限技术重试、一次补证据和安全 Fallback。
|
||||||
- AIOps 链路使用 `SupervisorAgent` 调度 `Planner + Executor`,最终由规则评估器做轻量验证。
|
- AIOps 链路使用 `SupervisorAgent` 调度 `Planner + Executor`,最终由规则评估器做轻量验证。
|
||||||
- 当前没有拆分 ExternalApiSubAgent、InternalErrorSubAgent、DatabaseSubAgent;这些作为后续演进方向保留。
|
- 当前没有拆分 ExternalApiSubAgent、InternalErrorSubAgent、DatabaseSubAgent;这些作为后续演进方向保留。
|
||||||
- 证据工具不直接散落在各个 Agent 里,而是通过 Spring AI ToolCallback / `@Tool` 统一暴露。
|
- 证据工具不直接散落在各个 Agent 里,而是通过 Spring AI ToolCallback / `@Tool` 统一暴露。
|
||||||
@@ -19,15 +19,19 @@
|
|||||||
flowchart TB
|
flowchart TB
|
||||||
subgraph Chat["Chat diagnosis"]
|
subgraph Chat["Chat diagnosis"]
|
||||||
ChatIn["POST /api/chat"] --> ChatService["ChatService"]
|
ChatIn["POST /api/chat"] --> ChatService["ChatService"]
|
||||||
ChatService --> ChatPlanner["chat_planner"]
|
ChatService --> ChatGraph["ChatDiagnosisGraphRuntime / StateGraph"]
|
||||||
|
ChatGraph --> ChatPlanner["Planner Node"]
|
||||||
ChatPlanner --> ChatExecutor["chat_executor"]
|
ChatPlanner --> ChatExecutor["chat_executor"]
|
||||||
ChatExecutor --> ChatTools["evidence tools"]
|
ChatExecutor --> ChatTools["evidence tools"]
|
||||||
ChatTools --> ChatExecutor
|
ChatTools --> ChatExecutor
|
||||||
ChatExecutor --> ChatGatekeeper["ExecutorGatekeeperService"]
|
ChatExecutor --> ChatGatekeeper["Gatekeeper Node"]
|
||||||
ChatGatekeeper --> ChatVerifier["chat_verifier"]
|
ChatGatekeeper --> VerifiedInput["Verified Input Node"]
|
||||||
|
VerifiedInput --> ChatVerifier["Verifier Node"]
|
||||||
ChatVerifier --> ChatDecision{"PASS / LOW_CONFID / REJECT"}
|
ChatVerifier --> ChatDecision{"PASS / LOW_CONFID / REJECT"}
|
||||||
ChatDecision --> ChatComposer["chat_composer"]
|
ChatDecision --> ChatComposer["Composer Node"]
|
||||||
|
ChatDecision --> ChatFallback["Fallback Node"]
|
||||||
ChatComposer --> ChatAnswer["final answer"]
|
ChatComposer --> ChatAnswer["final answer"]
|
||||||
|
ChatFallback --> ChatAnswer
|
||||||
end
|
end
|
||||||
|
|
||||||
subgraph AiOps["AIOps diagnosis"]
|
subgraph AiOps["AIOps diagnosis"]
|
||||||
@@ -54,6 +58,7 @@ flowchart TB
|
|||||||
ChatPlanner --> Step
|
ChatPlanner --> Step
|
||||||
ChatExecutor --> Step
|
ChatExecutor --> Step
|
||||||
ChatGatekeeper --> SelfEval
|
ChatGatekeeper --> SelfEval
|
||||||
|
ChatGraph --> Run
|
||||||
ChatVerifier --> Step
|
ChatVerifier --> Step
|
||||||
ChatTools --> Invocation
|
ChatTools --> Invocation
|
||||||
ChatDecision --> SelfEval
|
ChatDecision --> SelfEval
|
||||||
@@ -69,19 +74,13 @@ flowchart TB
|
|||||||
|
|
||||||
## 3. Chat 编排
|
## 3. Chat 编排
|
||||||
|
|
||||||
Chat 复杂诊断采用 `SequentialAgent`,顺序固定:
|
Chat 复杂诊断采用 `ChatDiagnosisGraphRuntime` 执行、`DiagnosisGraphFactory` 编译的 bounded StateGraph。它有一条正常路径和显式条件边,不再依赖固定顺序 Agent 或隐式前置校验:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
chat_planner
|
START -> PLANNER -> EXECUTOR -> GATEKEEPER -> VERIFIED_INPUT -> VERIFIER -> COMPOSER -> END
|
||||||
-> chat_executor
|
| | | | |
|
||||||
-> lookup_knowledge / query_logs / query_metrics / date_time
|
+ retry + fallback + fallback + retry + retry/fallback
|
||||||
-> outputs executor_evidence_v2
|
+ EVIDENCE_RETRY -> PLANNER (最多一次)
|
||||||
-> VerifierInputHook / ExecutorGatekeeperService
|
|
||||||
-> validates source_invocation_id / raw_path / evidence_excerpt
|
|
||||||
-> chat_verifier
|
|
||||||
-> judges whether verified evidence can derive claims
|
|
||||||
-> chat_composer
|
|
||||||
-> writes final user-facing answer
|
|
||||||
```
|
```
|
||||||
|
|
||||||
关键行为:
|
关键行为:
|
||||||
@@ -90,16 +89,18 @@ chat_planner
|
|||||||
|---|---|---|
|
|---|---|---|
|
||||||
| `chat_planner` | 拆解问题,注入知识域地图和对话历史,给出排查方向 | `planner_plan` |
|
| `chat_planner` | 拆解问题,注入知识域地图和对话历史,给出排查方向 | `planner_plan` |
|
||||||
| `chat_executor` | 按计划调用证据工具,抽取带 `source_invocation_id + raw_path + evidence_excerpt` 的微观事实 | `executor_evidence_v2` |
|
| `chat_executor` | 按计划调用证据工具,抽取带 `source_invocation_id + raw_path + evidence_excerpt` 的微观事实 | `executor_evidence_v2` |
|
||||||
| `ExecutorGatekeeperService` | 在 Verifier 前做代码级引用验真,拒绝伪造 ID、错配 raw_path、错配 excerpt | `gatekeeper_result` |
|
| `GatekeeperNode` / `ExecutorGatekeeperService` | 按当前 `runId` 做代码级引用验真,拒绝伪造 ID、错配 raw_path、错配 excerpt | `gatekeeper_result` |
|
||||||
| `chat_verifier` | 只判断已验真 evidence excerpt 是否能推出 claim,不做新检索 | `verifier_output` |
|
| `VerifiedInputNode` | 只投影 Gatekeeper 通过的 claims 与 matched evidence,隔离完整工具 Trace | `verified_executor_output`、`verified_evidence` |
|
||||||
|
| `chat_verifier` | 只判断已验真的 evidence excerpt 是否能推出 claim,不做新检索、不读取完整工具 Trace | `verifier_output` |
|
||||||
| `chat_composer` | 只表达 Verifier 允许输出的 claims、缺口和建议,生成最终用户答复 | `composer_output` |
|
| `chat_composer` | 只表达 Verifier 允许输出的 claims、缺口和建议,生成最终用户答复 | `composer_output` |
|
||||||
|
| `FallbackNode` | 在不可恢复失败或路由上限触发时生成非空安全答复 | `final_answer`、degraded trace |
|
||||||
|
|
||||||
Chat 链路最多支持两轮验证:
|
Chat Graph 支持有限技术重试,并只允许一次 evidence retry;所有分支最终进入 Composer 或 Fallback:
|
||||||
|
|
||||||
```mermaid
|
```mermaid
|
||||||
sequenceDiagram
|
sequenceDiagram
|
||||||
autonumber
|
autonumber
|
||||||
participant C as ChatService
|
participant C as ChatService / StateGraph
|
||||||
participant P as chat_planner
|
participant P as chat_planner
|
||||||
participant E as chat_executor
|
participant E as chat_executor
|
||||||
participant T as tools
|
participant T as tools
|
||||||
@@ -114,18 +115,22 @@ sequenceDiagram
|
|||||||
E->>T: 调用证据工具
|
E->>T: 调用证据工具
|
||||||
T-->>E: 证据结果
|
T-->>E: 证据结果
|
||||||
E-->>C: executor_evidence_v2
|
E-->>C: executor_evidence_v2
|
||||||
C->>G: executor_structured_output + tool_invocation.evidence_refs
|
C->>G: executor_output + run-owned tool_invocation.evidence_refs
|
||||||
G-->>C: gatekeeper_result
|
G-->>C: gatekeeper_result
|
||||||
C->>V: executor_structured_output + gatekeeper_result + tool_trace_summary
|
C->>C: VerifiedInputNode projects passed claims/evidence
|
||||||
|
C->>V: verified_executor_output + verified_evidence + gatekeeper_audit
|
||||||
V-->>C: PASS / LOW_CONFID / REJECT
|
V-->>C: PASS / LOW_CONFID / REJECT
|
||||||
C->>R: 写入 verifier_evaluation
|
C->>R: 写入 verifier_evaluation
|
||||||
alt LOW_CONFID 且允许补证据
|
alt LOW_CONFID 且允许一次补证据
|
||||||
C->>P: retry_context: 仅补缺失证据
|
C->>P: retry_context: 仅补缺失证据
|
||||||
else PASS 或 REJECT
|
else PASS / LOW_CONFID 可输出
|
||||||
C->>M: allowed_claims + missing_info + recommended_actions
|
C->>M: allowed_claims + missing_info + recommended_actions
|
||||||
M-->>C: composer_output
|
M-->>C: composer_output
|
||||||
C->>R: 保存 Composer 最终 answer
|
C->>R: 保存 Composer 最终 answer
|
||||||
|
else 不可恢复失败
|
||||||
|
C->>R: Fallback 安全答复
|
||||||
end
|
end
|
||||||
|
C->>R: 保存 orchestration_trace(version/transitions/final_node/termination_reason/degraded/evidence_retry_count)
|
||||||
```
|
```
|
||||||
|
|
||||||
决策语义:
|
决策语义:
|
||||||
@@ -134,7 +139,17 @@ sequenceDiagram
|
|||||||
|---|---|
|
|---|---|
|
||||||
| `PASS` | 把 Verifier 允许表达的 claims 交给 Composer 输出 |
|
| `PASS` | 把 Verifier 允许表达的 claims 交给 Composer 输出 |
|
||||||
| `LOW_CONFID` | 如果分数低于阈值且仍有轮次,构造 `retry_context` 补证据;否则输出低置信提示 |
|
| `LOW_CONFID` | 如果分数低于阈值且仍有轮次,构造 `retry_context` 补证据;否则输出低置信提示 |
|
||||||
| `REJECT` | 输出降级答复,只保留已确认信息和下一步建议 |
|
| `REJECT` | Verifier 完成后仍进入 Composer,但 Composer 只能表达允许材料和诊断限制;Gatekeeper REJECT 才直接进入 Fallback |
|
||||||
|
|
||||||
|
### 3.1 运行时边界
|
||||||
|
|
||||||
|
- 外层 Graph 使用 `runId` 作为 `RunnableConfig.threadId`,metadata 只承载 `sessionId/runId` 等业务身份。
|
||||||
|
- `ReactAgentDiagnosisInvoker` 为 nested Agent 创建独立 config,不传播外层 human-feedback、state-update、checkpoint/resume 控制 metadata。
|
||||||
|
- Graph State 默认 replace,只有 `orchestration_events` append;事件使用 portable Map,避免 DevTools classloader 的 record identity 问题。
|
||||||
|
- Planner、Verifier、Composer 各最多一次技术重试;evidence retry 独立计数且最多一次;Graph recursion limit 为 32。
|
||||||
|
- `DiagnosisGraphResultMapper` 负责质量评估,`DiagnosisOrchestrationTraceBuilder` 负责路由摘要,两者分别写入 `self_evaluation` 和 `orchestration_trace`。
|
||||||
|
|
||||||
|
完整运行时拓扑、路由表和失败语义见 [stategraph-runtime-architecture.md](stategraph-runtime-architecture.md)。
|
||||||
|
|
||||||
## 4. AIOps 编排
|
## 4. AIOps 编排
|
||||||
|
|
||||||
@@ -207,7 +222,7 @@ flowchart LR
|
|||||||
| Planner | 只看 skill name / description,并输出 `selected_skill` | 不暴露 `read_skill` |
|
| Planner | 只看 skill name / description,并输出 `selected_skill` | 不暴露 `read_skill` |
|
||||||
| Executor | 读取 Planner 选中的 skill 正文 | 暴露官方 `read_skill` 和证据工具 |
|
| Executor | 读取 Planner 选中的 skill 正文 | 暴露官方 `read_skill` 和证据工具 |
|
||||||
| Gatekeeper | 不看 skill catalog,也不读 skill 正文 | 只读取 Executor 输出和 `tool_invocation.retrieval_details.evidence_refs` |
|
| Gatekeeper | 不看 skill catalog,也不读 skill 正文 | 只读取 Executor 输出和 `tool_invocation.retrieval_details.evidence_refs` |
|
||||||
| Verifier | 不看 skill catalog,也不读 skill 正文 | 只读取 Gatekeeper 结果、结构化 claims 和 trace summary |
|
| Verifier | 不看 skill catalog,也不读 skill 正文 | 只读取 Verified Input 投影和 Gatekeeper audit |
|
||||||
| Composer | 不看 skill catalog,也不读 skill 正文 | 只读取 Verifier 允许表达的内容 |
|
| Composer | 不看 skill catalog,也不读 skill 正文 | 只读取 Verifier 允许表达的内容 |
|
||||||
|
|
||||||
## 7. 与旧版设计的差异
|
## 7. 与旧版设计的差异
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
# 当前 MVP 架构
|
# 当前 MVP 架构
|
||||||
|
|
||||||
**更新日期**:2026-07-08
|
**更新日期**:2026-07-20
|
||||||
**状态**:当前可运行架构
|
**状态**:当前可运行架构
|
||||||
**适用范围**:Demo、面试讲解、后续迭代规划
|
**适用范围**:Demo、面试讲解、后续迭代规划
|
||||||
|
|
||||||
@@ -36,11 +36,14 @@ flowchart TB
|
|||||||
|
|
||||||
subgraph Agent["Agent Orchestration"]
|
subgraph Agent["Agent Orchestration"]
|
||||||
Supervisor["Supervisor"]
|
Supervisor["Supervisor"]
|
||||||
|
StateGraph["Chat Diagnosis StateGraph"]
|
||||||
Planner["Planner"]
|
Planner["Planner"]
|
||||||
Executor["Executor"]
|
Executor["Executor"]
|
||||||
Gatekeeper["Gatekeeper"]
|
Gatekeeper["Gatekeeper"]
|
||||||
|
VerifiedInput["Verified Input"]
|
||||||
Verifier["Verifier"]
|
Verifier["Verifier"]
|
||||||
Composer["Composer"]
|
Composer["Composer"]
|
||||||
|
Fallback["Fallback"]
|
||||||
end
|
end
|
||||||
|
|
||||||
subgraph Tools["Evidence Tools"]
|
subgraph Tools["Evidence Tools"]
|
||||||
@@ -74,7 +77,14 @@ flowchart TB
|
|||||||
end
|
end
|
||||||
|
|
||||||
API --> App
|
API --> App
|
||||||
ChatService --> Agent
|
ChatService --> StateGraph
|
||||||
|
StateGraph --> Planner
|
||||||
|
StateGraph --> Executor
|
||||||
|
StateGraph --> Gatekeeper
|
||||||
|
StateGraph --> VerifiedInput
|
||||||
|
StateGraph --> Verifier
|
||||||
|
StateGraph --> Composer
|
||||||
|
StateGraph --> Fallback
|
||||||
AiOpsService --> Agent
|
AiOpsService --> Agent
|
||||||
SkillRegistry --> PlannerSkillHook
|
SkillRegistry --> PlannerSkillHook
|
||||||
PlannerSkillHook --> Planner
|
PlannerSkillHook --> Planner
|
||||||
@@ -86,8 +96,8 @@ flowchart TB
|
|||||||
RAG --> Store
|
RAG --> Store
|
||||||
Tools --> Invocation
|
Tools --> Invocation
|
||||||
Agent --> Step
|
Agent --> Step
|
||||||
App --> Session
|
App --> ChatSession
|
||||||
TraceService --> Session
|
TraceService --> ChatSession
|
||||||
TraceService --> Step
|
TraceService --> Step
|
||||||
TraceService --> Invocation
|
TraceService --> Invocation
|
||||||
```
|
```
|
||||||
@@ -157,7 +167,8 @@ sequenceDiagram
|
|||||||
participant Planner as Planner Agent
|
participant Planner as Planner Agent
|
||||||
participant Executor as Executor Agent
|
participant Executor as Executor Agent
|
||||||
participant Tool as Evidence Tools
|
participant Tool as Evidence Tools
|
||||||
participant Gatekeeper as Gatekeeper Hook
|
participant Gatekeeper as Gatekeeper Node
|
||||||
|
participant Projection as Verified Input Node
|
||||||
participant Verifier as Verifier Agent
|
participant Verifier as Verifier Agent
|
||||||
participant Composer as Composer Agent
|
participant Composer as Composer Agent
|
||||||
participant DB as Trace Tables
|
participant DB as Trace Tables
|
||||||
@@ -174,11 +185,12 @@ sequenceDiagram
|
|||||||
Tool-->>Executor: 返回证据
|
Tool-->>Executor: 返回证据
|
||||||
Executor->>Gatekeeper: 输出 executor_evidence_v2
|
Executor->>Gatekeeper: 输出 executor_evidence_v2
|
||||||
Gatekeeper->>DB: 读取 tool_invocation.evidence_refs 并校验引用
|
Gatekeeper->>DB: 读取 tool_invocation.evidence_refs 并校验引用
|
||||||
Gatekeeper->>Verifier: 传入已验真的 claims / excerpts
|
Gatekeeper->>Projection: 输出通过验真的 bindings
|
||||||
|
Projection->>Verifier: 只传入 verified claims / evidence
|
||||||
Verifier->>DB: 合并 diagnosis_run.self_evaluation.verifier_evaluation
|
Verifier->>DB: 合并 diagnosis_run.self_evaluation.verifier_evaluation
|
||||||
Verifier->>Composer: 传入 allowed_claims / missing_info / actions
|
Verifier->>Composer: 传入 allowed_claims / missing_info / actions
|
||||||
Composer->>Chat: 生成最终用户答复
|
Composer->>Chat: 生成最终用户答复
|
||||||
Chat->>DB: 保存 diagnosis_run.answer
|
Chat->>DB: 保存 answer / self_evaluation / orchestration_trace
|
||||||
User->>Trace: GET /api/diagnosis/{sessionId}/trace?runId=...
|
User->>Trace: GET /api/diagnosis/{sessionId}/trace?runId=...
|
||||||
Trace->>DB: 聚合 run / step / tool
|
Trace->>DB: 聚合 run / step / tool
|
||||||
Trace-->>User: 返回可回放诊断链路
|
Trace-->>User: 返回可回放诊断链路
|
||||||
@@ -194,19 +206,21 @@ POST /api/chat
|
|||||||
-> lookup_knowledge
|
-> lookup_knowledge
|
||||||
-> query_logs
|
-> query_logs
|
||||||
-> query_metrics
|
-> query_metrics
|
||||||
-> Gatekeeper 校验 Executor 证据引用真实性
|
-> Gatekeeper Node 校验 Executor 证据引用真实性
|
||||||
-> Verifier 判断 claim 是否能由已核验证据推出
|
-> Verified Input Node 只投影通过验真的 claims/evidence
|
||||||
|
-> Verifier 判断 claim 是否能由已验真证据推出
|
||||||
-> Composer 生成最终用户答复
|
-> Composer 生成最终用户答复
|
||||||
-> 保存 chat_session metadata
|
-> 保存 chat_session metadata
|
||||||
-> 保存 diagnosis_run
|
-> 保存 diagnosis_run
|
||||||
-> 保存 agent_step.run_id
|
-> 保存 agent_step.run_id
|
||||||
-> 保存 tool_invocation.run_id
|
-> 保存 tool_invocation.run_id
|
||||||
-> 合并 diagnosis_run.self_evaluation.verifier_evaluation
|
-> 合并 diagnosis_run.self_evaluation.verifier_evaluation
|
||||||
|
-> 保存 diagnosis_run.orchestration_trace
|
||||||
```
|
```
|
||||||
|
|
||||||
Chat 链路的质量门禁由三段组成:Gatekeeper 先做代码级引用验真,Verifier 再做 LLM 可推导性判断,Composer 最后控制对用户的表达边界。Gatekeeper、Verifier、Composer 的输出合并到当前 `diagnosis_run.self_evaluation.verifier_evaluation`,Trace API 会展示该验证结果。
|
Chat 链路的质量门禁由四段组成:Gatekeeper 先做代码级引用验真,Verified Input 再隔离未通过的 binding,Verifier 做 LLM 可推导性判断,Composer 最后控制对用户的表达边界。结构化质量结果合并到 `diagnosis_run.self_evaluation.verifier_evaluation`;Node 路由、重试和降级摘要独立写入 `diagnosis_run.orchestration_trace`。
|
||||||
|
|
||||||
Agent 编排细节见 [agent-orchestration.md](agent-orchestration.md)。
|
最新运行时、路由和状态边界见 [stategraph-runtime-architecture.md](stategraph-runtime-architecture.md),Agent 职责见 [agent-orchestration.md](agent-orchestration.md)。
|
||||||
|
|
||||||
关键代码:
|
关键代码:
|
||||||
|
|
||||||
@@ -333,6 +347,7 @@ diagnosis_run
|
|||||||
-> run_id / session_id
|
-> run_id / session_id
|
||||||
-> query / status / agent_flow / answer
|
-> query / status / agent_flow / answer
|
||||||
-> self_evaluation
|
-> self_evaluation
|
||||||
|
-> orchestration_trace
|
||||||
-> step_count / tool_call_count / duration
|
-> step_count / tool_call_count / duration
|
||||||
|
|
||||||
agent_step
|
agent_step
|
||||||
@@ -355,7 +370,7 @@ tool_invocation
|
|||||||
说明:
|
说明:
|
||||||
|
|
||||||
- 旧的 `diagnosis_record` 已不是当前主模型。
|
- 旧的 `diagnosis_record` 已不是当前主模型。
|
||||||
- `diagnosis_session` 已降级为历史兼容和回滚表,新执行写入 `chat_session + diagnosis_run`。
|
- 当前运行时只使用 `chat_session + diagnosis_run`,不再映射、读取或写入 `diagnosis_session`。
|
||||||
- `api_document` 仍用于文档元数据管理。
|
- `api_document` 仍用于文档元数据管理。
|
||||||
- 文档向量内容存放在 Milvus/Zilliz collection 中。
|
- 文档向量内容存放在 Milvus/Zilliz collection 中。
|
||||||
|
|
||||||
@@ -374,11 +389,12 @@ Trace API 聚合:
|
|||||||
- Agent step 序列。
|
- Agent step 序列。
|
||||||
- 工具调用和检索细节。
|
- 工具调用和检索细节。
|
||||||
- Chat Gatekeeper / Verifier / Composer 结果。
|
- Chat Gatekeeper / Verifier / Composer 结果。
|
||||||
|
- Chat `run.orchestrationTrace` 路由摘要,解析自 `diagnosis_run.orchestration_trace`,独立于 self-evaluation 和步骤/工具明细。
|
||||||
- AIOps rule evaluation 结果。
|
- AIOps rule evaluation 结果。
|
||||||
|
|
||||||
Trace 是本项目区别于普通问答系统的关键:答案不是孤立文本,而是可以追溯到 Agent 决策、工具调用和证据来源。
|
Trace 是本项目区别于普通问答系统的关键:答案不是孤立文本,而是可以追溯到 Agent 决策、工具调用和证据来源。
|
||||||
|
|
||||||
Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明见 [harness-quality-gates.md](harness-quality-gates.md),用户反馈与 `self_evaluation` 闭环见 [feedback-architecture.md](feedback-architecture.md)。
|
Prompt、StateGraph、Gatekeeper、Verifier、Composer 和评测门禁的完整说明见 [harness-quality-gates.md](harness-quality-gates.md),用户反馈与 `self_evaluation` 闭环见 [feedback-architecture.md](feedback-architecture.md)。
|
||||||
|
|
||||||
## 8. 质量门禁
|
## 8. 质量门禁
|
||||||
|
|
||||||
@@ -386,9 +402,11 @@ Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明
|
|||||||
|
|
||||||
| 门禁 | 位置 | 作用 |
|
| 门禁 | 位置 | 作用 |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| Executor Gatekeeper | `VerifierInputHook` / `ExecutorGatekeeperService` | 校验 Executor 引用的 invocation、`raw_path`、`evidence_excerpt` 是否真实 |
|
| Executor Gatekeeper | `GatekeeperNode` / `ExecutorGatekeeperService` | 校验当前 Run 的 invocation、`raw_path`、`evidence_excerpt` 是否真实 |
|
||||||
| Chat Verifier | `ChatService` | 判断已验真证据是否能推出 Executor claims |
|
| Verified Input | `VerifiedInputNode` | 仅投影 Gatekeeper 通过的 claims/evidence,阻断完整工具 Trace 进入 Verifier |
|
||||||
| Chat Composer | `ChatService` | 只表达 Verifier 允许输出的内容,避免把 no-evidence 说成已排除 |
|
| Chat Verifier | `VerifierNodeAdapter` | 判断已验真证据是否能推出 Executor claims |
|
||||||
|
| Chat Composer / Fallback | `ComposerNodeAdapter` / `FallbackNode` | 输出受控答复;异常分支也必须安全终止 |
|
||||||
|
| Graph routing | `DiagnosisGraphWorkflowTest` / `run.orchestrationTrace` | 验证条件边、有限重试、最终节点和终止原因 |
|
||||||
| AIOps Rule Evaluation | `AiOpsRuleEvaluationService` | 校验告警诊断是否聚焦 payload 并使用证据 |
|
| AIOps Rule Evaluation | `AiOpsRuleEvaluationService` | 校验告警诊断是否聚焦 payload 并使用证据 |
|
||||||
| Diagnosis Eval Baseline | `mvp/eval/` | 固化诊断 trace 和报告行为 |
|
| Diagnosis Eval Baseline | `mvp/eval/` | 固化诊断 trace 和报告行为 |
|
||||||
| RAG Retrieval Baseline | `eval/rag-retrieval/` | 固化检索召回行为,避免 RAG 重构回退 |
|
| RAG Retrieval Baseline | `eval/rag-retrieval/` | 固化检索召回行为,避免 RAG 重构回退 |
|
||||||
@@ -399,6 +417,7 @@ Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明
|
|||||||
已经完成:
|
已经完成:
|
||||||
|
|
||||||
- Chat 和 AIOps 两条入口链路。
|
- Chat 和 AIOps 两条入口链路。
|
||||||
|
- Chat 复杂诊断已单轨切换到 bounded StateGraph,并持久化 Run-owned `orchestration_trace`;Graph event 使用 portable Map,nested ReactAgent config 与外层 checkpoint/resume 控制信息隔离。
|
||||||
- 显式 `lookup_knowledge` Agent Tool。
|
- 显式 `lookup_knowledge` Agent Tool。
|
||||||
- L0 从最终决策降级为 domain/entity hint。
|
- L0 从最终决策降级为 domain/entity hint。
|
||||||
- `VectorSearchService` 作为稳定检索门面。
|
- `VectorSearchService` 作为稳定检索门面。
|
||||||
@@ -423,13 +442,16 @@ Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明
|
|||||||
- VectorStore 写入路径全面迁移。
|
- VectorStore 写入路径全面迁移。
|
||||||
- 完整 LLM-based AIOps verifier。
|
- 完整 LLM-based AIOps verifier。
|
||||||
|
|
||||||
后续 Agent 拆分、Skill/Playbook、MCP 工具协议化和进程隔离等方向见 [evolution-roadmap.md](evolution-roadmap.md)。
|
后续 Agent 拆分、Playbook 版本化增强、MCP 工具协议化和进程隔离等方向见 [evolution-roadmap.md](evolution-roadmap.md)。
|
||||||
|
|
||||||
## 10. 关键代码索引
|
## 10. 关键代码索引
|
||||||
|
|
||||||
| 能力 | 代码 |
|
| 能力 | 代码 |
|
||||||
|---|---|
|
|---|---|
|
||||||
| Chat 入口与编排 | `ChatController`, `ChatService` |
|
| Chat 入口与编排 | `ChatController`, `ChatService` |
|
||||||
|
| Chat StateGraph runtime | `ChatDiagnosisGraphRuntime`, `DiagnosisGraphFactory`, `DiagnosisGraphRouter`, `DiagnosisGraphState` |
|
||||||
|
| Chat Node assembly/config isolation | `DiagnosisRealGraphActionsFactory`, `ReactAgentDiagnosisInvoker` |
|
||||||
|
| Graph result/audit mapping | `DiagnosisGraphResultMapper`, `DiagnosisOrchestrationTraceBuilder` |
|
||||||
| AIOps 入口与编排 | `ChatController.aiOps`, `AiOpsService` |
|
| AIOps 入口与编排 | `ChatController.aiOps`, `AiOpsService` |
|
||||||
| AIOps 规则验证 | `AiOpsRuleEvaluationService` |
|
| AIOps 规则验证 | `AiOpsRuleEvaluationService` |
|
||||||
| 知识库工具 | `LookupKnowledgeTool` |
|
| 知识库工具 | `LookupKnowledgeTool` |
|
||||||
@@ -440,5 +462,5 @@ Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明
|
|||||||
| Spring AI VectorStore 配置辅助 | `SpringAiVectorStoreSidecarService` |
|
| Spring AI VectorStore 配置辅助 | `SpringAiVectorStoreSidecarService` |
|
||||||
| Trace 聚合 | `DiagnosisTraceService` |
|
| Trace 聚合 | `DiagnosisTraceService` |
|
||||||
| 工具调用记录 | `ToolInvocationRecorder` |
|
| 工具调用记录 | `ToolInvocationRecorder` |
|
||||||
| Executor 引用验真 | `ExecutorGatekeeperService`, `VerifierInputHook` |
|
| Executor 引用验真与投影 | `GatekeeperNode`, `ExecutorGatekeeperService`, `VerifiedInputNode` |
|
||||||
| self_evaluation 合并 | `SelfEvaluationMergeService` |
|
| self_evaluation 合并 | `SelfEvaluationMergeService` |
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
# 数据模型总览
|
# 数据模型总览
|
||||||
|
|
||||||
**更新日期**:2026-07-10
|
**更新日期**:2026-07-20
|
||||||
**状态**:当前可运行架构
|
**状态**:当前可运行架构
|
||||||
|
|
||||||
## 1. 定位
|
## 1. 定位
|
||||||
@@ -13,7 +13,7 @@
|
|||||||
- 知识库:`api_document`、`knowledge_domain`、Milvus/Zilliz metadata
|
- 知识库:`api_document`、`knowledge_domain`、Milvus/Zilliz metadata
|
||||||
- 反馈沉淀:`case_library`
|
- 反馈沉淀:`case_library`
|
||||||
|
|
||||||
`diagnosis_session` 仍保留为历史兼容和回滚表,不再是新执行写入的主模型。
|
当前运行时只使用 `chat_session + diagnosis_run`;旧 `diagnosis_session` 表不属于当前版本契约。
|
||||||
|
|
||||||
## 2. 总体关系
|
## 2. 总体关系
|
||||||
|
|
||||||
@@ -44,6 +44,7 @@ erDiagram
|
|||||||
varchar agent_flow
|
varchar agent_flow
|
||||||
longtext answer
|
longtext answer
|
||||||
json self_evaluation
|
json self_evaluation
|
||||||
|
json orchestration_trace
|
||||||
varchar feedback
|
varchar feedback
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -111,6 +112,7 @@ erDiagram
|
|||||||
| `agent_flow` | `CHAT` / `AI_OPS` |
|
| `agent_flow` | `CHAT` / `AI_OPS` |
|
||||||
| `answer` | 本次运行最终答复或告警报告 |
|
| `answer` | 本次运行最终答复或告警报告 |
|
||||||
| `self_evaluation` | 本次运行的 rule/verifier/aiops 自评估容器 |
|
| `self_evaluation` | 本次运行的 rule/verifier/aiops 自评估容器 |
|
||||||
|
| `orchestration_trace` | nullable JSON;复杂 Chat 的 StateGraph 路由摘要,非 StateGraph Run 可为空 |
|
||||||
| `feedback` | 本次运行的用户反馈 |
|
| `feedback` | 本次运行的用户反馈 |
|
||||||
|
|
||||||
同一个 `sessionId` 可以有多个 `runId`。Trace、反馈、评测和案例沉淀都应优先使用 `runId`,避免多轮同 session 下的数据混合。
|
同一个 `sessionId` 可以有多个 `runId`。Trace、反馈、评测和案例沉淀都应优先使用 `runId`,避免多轮同 session 下的数据混合。
|
||||||
@@ -139,7 +141,19 @@ erDiagram
|
|||||||
|
|
||||||
`$.no_evidence` 只表示“本次工具查询未检索到匹配证据”,不能被解释为“问题不存在”或“根因已排除”。
|
`$.no_evidence` 只表示“本次工具查询未检索到匹配证据”,不能被解释为“问题不存在”或“根因已排除”。
|
||||||
|
|
||||||
## 5. 反馈沉淀模型
|
## 5. Run 级审计分层
|
||||||
|
|
||||||
|
一次 Run 的可审计信息分为三层,不能互相替代:
|
||||||
|
|
||||||
|
| 层次 | 存储 | 语义 |
|
||||||
|
|---|---|---|
|
||||||
|
| 执行明细 | `agent_step`、`tool_invocation` | 模型步骤、工具输入输出、检索细节和证据事实 |
|
||||||
|
| 质量评估 | `diagnosis_run.self_evaluation` | rule/verifier/aiops 判断、Gatekeeper 审计、Prompt 版本和允许输出材料 |
|
||||||
|
| 编排摘要 | `diagnosis_run.orchestration_trace` | StateGraph transitions、final node、termination reason、degraded、evidence retry count |
|
||||||
|
|
||||||
|
`orchestration_trace` 只通过 exact Run 的 `run.orchestrationTrace` 暴露,Trace 响应不再包含兼容 `session` 投影。Run 状态仍只表达执行生命周期:安全 Fallback 是 `SUCCESS + degraded=true`,只有无法生成安全响应或未处理失败才是 `FAILED`。
|
||||||
|
|
||||||
|
## 6. 反馈沉淀模型
|
||||||
|
|
||||||
`useful` 反馈会触发 `CaseLibraryService.createFromRun`。
|
`useful` 反馈会触发 `CaseLibraryService.createFromRun`。
|
||||||
|
|
||||||
@@ -148,7 +162,7 @@ erDiagram
|
|||||||
| 字段 | 来源 |
|
| 字段 | 来源 |
|
||||||
|---|---|
|
|---|---|
|
||||||
| `case_id` | UUID |
|
| `case_id` | UUID |
|
||||||
| `diagnosis_id` | 新数据为 `diagnosis_run.run_id`;历史数据可能为 `diagnosis_session.session_id` |
|
| `diagnosis_id` | `diagnosis_run.run_id` |
|
||||||
| `source_type` | `AUTO` |
|
| `source_type` | `AUTO` |
|
||||||
| `fault_category` | 当前默认 `GENERAL` |
|
| `fault_category` | 当前默认 `GENERAL` |
|
||||||
| `title` | run query 前 100 字符 |
|
| `title` | run query 前 100 字符 |
|
||||||
@@ -156,7 +170,7 @@ erDiagram
|
|||||||
| `solution` | run answer |
|
| `solution` | run answer |
|
||||||
| `created_by` | `system` |
|
| `created_by` | `system` |
|
||||||
|
|
||||||
## 6. self_evaluation 结构
|
## 7. self_evaluation 结构
|
||||||
|
|
||||||
`diagnosis_run.self_evaluation` 是运行级 JSON 容器:
|
`diagnosis_run.self_evaluation` 是运行级 JSON 容器:
|
||||||
|
|
||||||
@@ -170,16 +184,19 @@ erDiagram
|
|||||||
|
|
||||||
Chat 通常写入 `rule_evaluation` 和 `verifier_evaluation`;AIOps 写入 `aiops_rule_evaluation`。
|
Chat 通常写入 `rule_evaluation` 和 `verifier_evaluation`;AIOps 写入 `aiops_rule_evaluation`。
|
||||||
|
|
||||||
## 7. 当前边界和后续
|
`orchestration_trace` 不放入该 JSON,避免把答案质量和 Graph 路由混成同一审计维度。
|
||||||
|
|
||||||
|
## 8. 当前边界和后续
|
||||||
|
|
||||||
当前边界:
|
当前边界:
|
||||||
|
|
||||||
- `chat_session` 只存会话元数据,不存完整正文历史。
|
- `chat_session` 只存会话元数据,不存完整正文历史。
|
||||||
- `diagnosis_run` 存一次运行的长期审计状态。
|
- `diagnosis_run` 存一次运行的长期审计状态。
|
||||||
- `agent_step.run_id` 和 `tool_invocation.run_id` 是 Trace、Verifier、Eval 的运行边界。
|
- `agent_step.run_id` 和 `tool_invocation.run_id` 是 Trace、Verifier、Eval 的运行边界。
|
||||||
|
- `diagnosis_run.orchestration_trace` 是 nullable Run-owned Graph 摘要;非 StateGraph Run 可以为空。
|
||||||
- 当前实现主要使用逻辑关联,不依赖数据库外键。
|
- 当前实现主要使用逻辑关联,不依赖数据库外键。
|
||||||
- `case_library.diagnosis_id` 是过渡字段,新值按 `run_id` 解释,旧值可能按 `session_id` 解释。
|
- `case_library.diagnosis_id` 是过渡字段,新值按 `run_id` 解释,旧值可能按 `session_id` 解释。
|
||||||
- `diagnosis_session` 只作为历史兼容和回滚表保留。
|
- 当前 Java 运行时不存在 `diagnosis_session` entity/repository 或 fallback。
|
||||||
|
|
||||||
后续可增强:
|
后续可增强:
|
||||||
|
|
||||||
|
|||||||
@@ -1,12 +1,12 @@
|
|||||||
# Agent 架构演进路线
|
# Agent 架构演进路线
|
||||||
|
|
||||||
**更新日期**:2026-07-05
|
**更新日期**:2026-07-20
|
||||||
**状态**:后续演进设计,不代表当前已实现
|
**状态**:后续演进设计,不代表当前已实现
|
||||||
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
||||||
|
|
||||||
## 1. 为什么需要演进路线
|
## 1. 为什么需要演进路线
|
||||||
|
|
||||||
旧版 `agent-architecture.md` 包含很多生产级设想:专科 SubAgent、Skill 体系、进程隔离、回退路由、MCP 工具协议化、进化引擎。它们不应作为当前 MVP 事实写入主架构,但可以作为后续扩展路线。
|
旧版 `agent-architecture.md` 包含很多生产级设想:专科 SubAgent、完整 Skill 治理、进程隔离、跨 Agent 回退、MCP 工具协议化、进化引擎。当前已经落地 bounded StateGraph、安全 Fallback 和基础 Skill/Playbook 接入;本文件只描述它们之上的后续增强。
|
||||||
|
|
||||||
当前原则:
|
当前原则:
|
||||||
|
|
||||||
@@ -18,12 +18,12 @@
|
|||||||
|
|
||||||
```mermaid
|
```mermaid
|
||||||
flowchart TD
|
flowchart TD
|
||||||
MVP["Current MVP: Planner + Executor + Verifier"] --> Split{"Executor 是否过载?"}
|
MVP["Current MVP: bounded StateGraph + evidence gates"] --> Split{"Executor 是否过载?"}
|
||||||
Split -->|是| SubAgents["专科 SubAgent"]
|
Split -->|是| SubAgents["专科 SubAgent"]
|
||||||
Split -->|否| Keep["继续强化通用 Executor"]
|
Split -->|否| Keep["继续强化通用 Executor"]
|
||||||
|
|
||||||
SubAgents --> Skills["Skill / Playbook 体系"]
|
SubAgents --> Skills["Skill / Playbook 版本化治理"]
|
||||||
Skills --> Fallback["回退路由"]
|
Skills --> Fallback["跨 SubAgent 回退路由"]
|
||||||
Fallback --> Isolation["进程或 Pod 隔离"]
|
Fallback --> Isolation["进程或 Pod 隔离"]
|
||||||
|
|
||||||
MVP --> ToolGrowth{"工具数量和来源是否增长?"}
|
MVP --> ToolGrowth{"工具数量和来源是否增长?"}
|
||||||
@@ -59,9 +59,9 @@ flowchart TD
|
|||||||
- 过早拆分会增加 Prompt、评测和 trace 分析成本。
|
- 过早拆分会增加 Prompt、评测和 trace 分析成本。
|
||||||
- 没有足够分类评测前,拆分可能只是移动复杂度。
|
- 没有足够分类评测前,拆分可能只是移动复杂度。
|
||||||
|
|
||||||
## 4. Skill / Playbook 体系
|
## 4. Skill / Playbook 版本化治理
|
||||||
|
|
||||||
旧版设计中的 Skill 可以在当前项目中演进为可版本化的诊断 Playbook。
|
当前已经通过 Planner metadata selection + Executor `read_skill` 接入诊断 Playbook。下一阶段不是重新建设 Skill 入口,而是增加版本、评测、回退和审计治理。
|
||||||
|
|
||||||
```text
|
```text
|
||||||
fault_category
|
fault_category
|
||||||
@@ -73,14 +73,14 @@ fault_category
|
|||||||
-> evaluation checks
|
-> evaluation checks
|
||||||
```
|
```
|
||||||
|
|
||||||
优先落地方向:
|
当前已覆盖的方向:
|
||||||
|
|
||||||
- AIOps 告警处理 Playbook。
|
- AIOps 告警处理 Playbook。
|
||||||
- 支付超时 Playbook。
|
- 支付超时 Playbook。
|
||||||
- MySQL 连接池风险 Playbook。
|
- MySQL 连接池风险 Playbook。
|
||||||
- Redis timeout Playbook。
|
- Redis timeout Playbook。
|
||||||
|
|
||||||
落地前提:
|
后续增强前提:
|
||||||
|
|
||||||
- 每个 Playbook 至少有 3-5 个 eval case。
|
- 每个 Playbook 至少有 3-5 个 eval case。
|
||||||
- Playbook 失败时可以回退到通用 Executor。
|
- Playbook 失败时可以回退到通用 Executor。
|
||||||
|
|||||||
@@ -1,7 +1,7 @@
|
|||||||
# Chat Evidence Pipeline Contracts
|
# Chat Evidence Pipeline Contracts
|
||||||
|
|
||||||
**状态**:当前实现
|
**状态**:当前实现
|
||||||
**更新日期**:2026-07-08
|
**更新日期**:2026-07-17
|
||||||
**范围**:Chat 复杂诊断链路中的 Planner、Executor、Gatekeeper、Verifier、Composer 数据契约
|
**范围**:Chat 复杂诊断链路中的 Planner、Executor、Gatekeeper、Verifier、Composer 数据契约
|
||||||
|
|
||||||
当前 Chat 复杂诊断链路是:
|
当前 Chat 复杂诊断链路是:
|
||||||
@@ -9,7 +9,8 @@
|
|||||||
```text
|
```text
|
||||||
chat_planner
|
chat_planner
|
||||||
-> chat_executor
|
-> chat_executor
|
||||||
-> VerifierInputHook / ExecutorGatekeeperService
|
-> GatekeeperNode / ExecutorGatekeeperService
|
||||||
|
-> VerifiedInputNode
|
||||||
-> chat_verifier
|
-> chat_verifier
|
||||||
-> chat_composer
|
-> chat_composer
|
||||||
-> final answer
|
-> final answer
|
||||||
@@ -219,13 +220,13 @@ Executor 必须遵守:
|
|||||||
|
|
||||||
## 4. Gatekeeper
|
## 4. Gatekeeper
|
||||||
|
|
||||||
Gatekeeper 位于 Verifier 前,由 `VerifierInputHook` 触发,负责代码级引用真实性校验。
|
Gatekeeper 是 StateGraph 中的显式 Node,调用 `ExecutorGatekeeperService` 对当前 Run 的证据引用做代码级真实性校验。
|
||||||
|
|
||||||
### 4.1 输入
|
### 4.1 输入
|
||||||
|
|
||||||
- `sessionId`
|
- `sessionId + runId`(来自 `RunnableConfig`,工具查询以 `runId` 为边界)
|
||||||
- `executor_structured_output`
|
- `executor_structured_output`
|
||||||
- 当前 session 的 `tool_invocation`
|
- 当前 run 的 `tool_invocation`
|
||||||
|
|
||||||
### 4.2 输出
|
### 4.2 输出
|
||||||
|
|
||||||
@@ -304,22 +305,20 @@ Gatekeeper 位于 Verifier 前,由 `VerifierInputHook` 触发,负责代码
|
|||||||
|
|
||||||
## 5. Verifier
|
## 5. Verifier
|
||||||
|
|
||||||
Verifier 输入由 `VerifierInputHook` 构造:
|
`VerifiedInputNode` 只保留 Gatekeeper 检查通过的 claim/binding,并为 Verifier 构造最小输入:
|
||||||
|
|
||||||
```json
|
```json
|
||||||
{
|
{
|
||||||
"original_query": "用户原始问题",
|
"diagnosis_context": {
|
||||||
"executor_final_answer": "{...executor raw text for debug/fallback only...}",
|
"query": "用户原始问题"
|
||||||
"executor_structured_output": {
|
},
|
||||||
|
"verified_executor_output": {
|
||||||
"answer_version": "executor_evidence_v2",
|
"answer_version": "executor_evidence_v2",
|
||||||
"claims": []
|
"claims": []
|
||||||
},
|
},
|
||||||
"executor_output_parse_status": {
|
"verified_evidence": [],
|
||||||
"status": "valid",
|
"gatekeeper_audit": {},
|
||||||
"detail": "parsed executor evidence contract"
|
"verdict_ceiling": "PASS",
|
||||||
},
|
|
||||||
"tool_trace_summary": [],
|
|
||||||
"gatekeeper_result": {},
|
|
||||||
"retry_context": null
|
"retry_context": null
|
||||||
}
|
}
|
||||||
```
|
```
|
||||||
@@ -330,7 +329,7 @@ Verifier 职责:
|
|||||||
- 不读 skill。
|
- 不读 skill。
|
||||||
- 不逐字核验 excerpt 真伪;这由 Gatekeeper 完成。
|
- 不逐字核验 excerpt 真伪;这由 Gatekeeper 完成。
|
||||||
- 只判断 `claim_text` 是否能由已核验的 `evidence_excerpt` 推出。
|
- 只判断 `claim_text` 是否能由已核验的 `evidence_excerpt` 推出。
|
||||||
- 结构化输出有效时,不得从 `executor_final_answer` 抽取额外确认事实。
|
- 不读取 Executor 原始答复或完整工具 Trace,只读取 verified projection。
|
||||||
- 对 `gatekeeper_result.severity=reject` 不得输出 `PASS`。
|
- 对 `gatekeeper_result.severity=reject` 不得输出 `PASS`。
|
||||||
- 对 `gatekeeper_result.severity=low_confid` 不得输出 `PASS`。
|
- 对 `gatekeeper_result.severity=low_confid` 不得输出 `PASS`。
|
||||||
|
|
||||||
@@ -410,7 +409,7 @@ Composer 位于 Verifier 之后,输入是 ChatService 过滤后的允许表达
|
|||||||
"rule_set_version": "gatekeeper-rules-v1"
|
"rule_set_version": "gatekeeper-rules-v1"
|
||||||
},
|
},
|
||||||
"composer_output": {},
|
"composer_output": {},
|
||||||
"tool_trace_summary": []
|
"verified_evidence": []
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
```
|
```
|
||||||
@@ -422,6 +421,9 @@ Trace API 可用于回放:
|
|||||||
- Gatekeeper 是否通过、是否自动回填。
|
- Gatekeeper 是否通过、是否自动回填。
|
||||||
- Verifier 如何判断可推导性。
|
- Verifier 如何判断可推导性。
|
||||||
- Composer 最终如何表达给用户。
|
- Composer 最终如何表达给用户。
|
||||||
|
- `run.orchestrationTrace` 如何经过条件边、有限重试并终止。
|
||||||
|
|
||||||
|
当前版本不生成、读取或展示 `verifier_evaluation.tool_trace_summary`;Verifier 只消费 verified projection。
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
# 反馈与自评估架构
|
# 反馈与自评估架构
|
||||||
|
|
||||||
**更新日期**:2026-07-10
|
**更新日期**:2026-07-17
|
||||||
**状态**:当前可运行架构
|
**状态**:当前可运行架构
|
||||||
**参考历史文档**:`archive/2026-07-05-legacy/confidence-feedback.md`
|
**参考历史文档**:`archive/2026-07-05-legacy/confidence-feedback.md`
|
||||||
|
|
||||||
@@ -27,9 +27,8 @@ flowchart TD
|
|||||||
Invocation["tool_invocation"] --> RuleEval["EvaluationService: rule_evaluation"]
|
Invocation["tool_invocation"] --> RuleEval["EvaluationService: rule_evaluation"]
|
||||||
Invocation --> EvidenceRefs["evidence_refs"]
|
Invocation --> EvidenceRefs["evidence_refs"]
|
||||||
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
|
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
|
||||||
Invocation --> TraceSummary["ToolTraceSummaryService"]
|
Gatekeeper --> Projection["VerifiedInputNode"]
|
||||||
Gatekeeper --> Verifier["chat_verifier"]
|
Projection --> Verifier["chat_verifier"]
|
||||||
TraceSummary --> Verifier
|
|
||||||
Verifier --> VerifierEval["verifier_evaluation"]
|
Verifier --> VerifierEval["verifier_evaluation"]
|
||||||
Verifier --> Composer["chat_composer"]
|
Verifier --> Composer["chat_composer"]
|
||||||
Composer --> VerifierEval
|
Composer --> VerifierEval
|
||||||
@@ -57,7 +56,7 @@ flowchart TD
|
|||||||
|
|
||||||
## 3. self_evaluation JSON
|
## 3. self_evaluation JSON
|
||||||
|
|
||||||
`SelfEvaluationMergeService` 统一维护当前运行的 `diagnosis_run.self_evaluation`。历史兼容数据可能仍存在于 `diagnosis_session.self_evaluation`,但新 Chat/AIOps 执行不再写旧表。
|
`SelfEvaluationMergeService` 只维护当前运行的 `diagnosis_run.self_evaluation`,不再解析或写入旧 `diagnosis_session` 数据。
|
||||||
|
|
||||||
当前结构:
|
当前结构:
|
||||||
|
|
||||||
@@ -81,7 +80,7 @@ flowchart TD
|
|||||||
"executor_structured_output": {},
|
"executor_structured_output": {},
|
||||||
"gatekeeper_result": {},
|
"gatekeeper_result": {},
|
||||||
"composer_output": {},
|
"composer_output": {},
|
||||||
"tool_trace_summary": []
|
"verified_evidence": []
|
||||||
},
|
},
|
||||||
"aiops_rule_evaluation": {
|
"aiops_rule_evaluation": {
|
||||||
"verdict": "...",
|
"verdict": "...",
|
||||||
@@ -90,10 +89,7 @@ flowchart TD
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
兼容逻辑:
|
输入必须使用当前分层 JSON:`rule_evaluation`、`verifier_evaluation`、`aiops_rule_evaluation`。旧扁平 JSON 不再自动包装。
|
||||||
|
|
||||||
- 如果旧 JSON 根节点包含 `evidence_score`,会被包进 `rule_evaluation`。
|
|
||||||
- 如果旧 JSON 根节点包含 `verdict` / `groundedness_score`,会被包进 `verifier_evaluation`。
|
|
||||||
|
|
||||||
## 4. 规则评分
|
## 4. 规则评分
|
||||||
|
|
||||||
@@ -136,14 +132,13 @@ Chat 自评估分三步:
|
|||||||
```mermaid
|
```mermaid
|
||||||
flowchart LR
|
flowchart LR
|
||||||
ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
|
ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
|
||||||
Invocation["tool_invocation"] --> Summary["ToolTraceSummaryService"]
|
Invocation["tool_invocation"] --> EvidenceRefs["retrieval_details.evidence_refs"]
|
||||||
Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
|
|
||||||
EvidenceRefs --> Gatekeeper
|
EvidenceRefs --> Gatekeeper
|
||||||
Gatekeeper --> GateResult["gatekeeper_result"]
|
Gatekeeper --> GateResult["gatekeeper_result"]
|
||||||
Summary --> Evidence["tool_trace_summary"]
|
GateResult --> Projection["VerifiedInputNode"]
|
||||||
GateResult --> Verifier["chat_verifier"]
|
ExecutorOutput --> Projection
|
||||||
ExecutorOutput --> Verifier
|
Projection --> VerifiedOutput["verified_executor_output + verified_evidence"]
|
||||||
Evidence --> Verifier
|
VerifiedOutput --> Verifier["chat_verifier"]
|
||||||
Verifier --> Output["verifier_output JSON"]
|
Verifier --> Output["verifier_output JSON"]
|
||||||
Output --> Composer["chat_composer"]
|
Output --> Composer["chat_composer"]
|
||||||
Composer --> ComposerOutput["composer_output"]
|
Composer --> ComposerOutput["composer_output"]
|
||||||
@@ -163,9 +158,9 @@ Verifier 输出:
|
|||||||
| `facts_checked` | 逐条事实校验 |
|
| `facts_checked` | 逐条事实校验 |
|
||||||
| `rationale` | 判定原因 |
|
| `rationale` | 判定原因 |
|
||||||
| `executor_structured_output` | Executor 输出的结构化 claims 与证据绑定 |
|
| `executor_structured_output` | Executor 输出的结构化 claims 与证据绑定 |
|
||||||
|
| `verified_evidence` | Gatekeeper 通过并投影给 Verifier 的最小 matched evidence |
|
||||||
| `gatekeeper_result` | 引用真实性校验结果 |
|
| `gatekeeper_result` | 引用真实性校验结果 |
|
||||||
| `composer_output` | 最终表达的解析状态和摘要 |
|
| `composer_output` | 最终表达的解析状态和摘要 |
|
||||||
| `tool_trace_summary` | 本次校验使用的工具调用导航索引 |
|
|
||||||
|
|
||||||
ChatService 根据 verdict 决定:
|
ChatService 根据 verdict 决定:
|
||||||
|
|
||||||
@@ -177,6 +172,7 @@ ChatService 根据 verdict 决定:
|
|||||||
|
|
||||||
- `executor_final_answer` 只作为 debug/fallback 上下文;结构化输出有效时,Verifier 不得从中抽取额外确认事实。
|
- `executor_final_answer` 只作为 debug/fallback 上下文;结构化输出有效时,Verifier 不得从中抽取额外确认事实。
|
||||||
- `$.no_evidence` 只能表达“当前查询未检索到匹配证据”,不能表达“已排除/确认没有”。
|
- `$.no_evidence` 只能表达“当前查询未检索到匹配证据”,不能表达“已排除/确认没有”。
|
||||||
|
- `run.orchestrationTrace` 是独立的 StateGraph 路由摘要,不属于 `self_evaluation`;当前版本不生成或读取 `tool_trace_summary`。
|
||||||
|
|
||||||
## 6. AIOps 规则自评估
|
## 6. AIOps 规则自评估
|
||||||
|
|
||||||
@@ -211,7 +207,6 @@ Content-Type: application/json
|
|||||||
"success": true,
|
"success": true,
|
||||||
"message": "反馈已记录",
|
"message": "反馈已记录",
|
||||||
"runId": "run-xxx",
|
"runId": "run-xxx",
|
||||||
"fallbackToLatestRun": false,
|
|
||||||
"caseId": "uuid 或 null"
|
"caseId": "uuid 或 null"
|
||||||
}
|
}
|
||||||
```
|
```
|
||||||
@@ -224,11 +219,7 @@ Content-Type: application/json
|
|||||||
| `not_useful` | 写入 `DiagnosisRun.feedback`,不改变 run status |
|
| `not_useful` | 写入 `DiagnosisRun.feedback`,不改变 run status |
|
||||||
| 其他值 | 返回 HTTP 400 |
|
| 其他值 | 返回 HTTP 400 |
|
||||||
|
|
||||||
兼容行为:
|
新版本要求请求必须携带 `sessionId + runId`。后端验证 `runId` 属于 `sessionId`;缺少 `runId`、Run 不存在或归属错误时直接拒绝,不绑定 latest run,也不回退历史表。
|
||||||
|
|
||||||
- 请求带 `runId` 时,后端验证 `runId` 属于 `sessionId`。
|
|
||||||
- 请求缺少 `runId` 且存在 run-backed 数据时,后端绑定 latest run,并返回 `fallbackToLatestRun=true` 和实际 `runId`。
|
|
||||||
- 仅当没有 `diagnosis_run` 但存在历史 `diagnosis_session` 时,才使用历史 fallback;该路径不声明 latest-run fallback。
|
|
||||||
|
|
||||||
## 8. 案例沉淀
|
## 8. 案例沉淀
|
||||||
|
|
||||||
@@ -239,7 +230,7 @@ Content-Type: application/json
|
|||||||
| CaseLibrary 字段 | 来源 |
|
| CaseLibrary 字段 | 来源 |
|
||||||
|---|---|
|
|---|---|
|
||||||
| `caseId` | UUID |
|
| `caseId` | UUID |
|
||||||
| `diagnosisId` | 新数据为 `DiagnosisRun.runId`;历史数据可能为 `DiagnosisSession.sessionId` |
|
| `diagnosisId` | `DiagnosisRun.runId` |
|
||||||
| `sourceType` | `AUTO` |
|
| `sourceType` | `AUTO` |
|
||||||
| `faultCategory` | 当前固定为 `GENERAL` |
|
| `faultCategory` | 当前固定为 `GENERAL` |
|
||||||
| `title` | `query` 前 100 字符 |
|
| `title` | `query` 前 100 字符 |
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
# Harness 与质量门禁架构
|
# Harness 与质量门禁架构
|
||||||
|
|
||||||
**更新日期**:2026-07-08
|
**更新日期**:2026-07-17
|
||||||
**状态**:当前可运行架构 + 后续门禁规划
|
**状态**:当前可运行架构 + 后续门禁规划
|
||||||
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
||||||
|
|
||||||
@@ -20,6 +20,7 @@ Agent 系统的核心风险不是“没有答案”,而是:
|
|||||||
Prompt contract
|
Prompt contract
|
||||||
+ Tool boundary
|
+ Tool boundary
|
||||||
+ Agent hooks
|
+ Agent hooks
|
||||||
|
+ StateGraph routing contract
|
||||||
+ Trace persistence
|
+ Trace persistence
|
||||||
+ Gatekeeper deterministic validation
|
+ Gatekeeper deterministic validation
|
||||||
+ Verifier / rule evaluation
|
+ Verifier / rule evaluation
|
||||||
@@ -31,7 +32,7 @@ Prompt contract
|
|||||||
```mermaid
|
```mermaid
|
||||||
flowchart TB
|
flowchart TB
|
||||||
Input["User / AIOps input"] --> Prompt["Prompt contract"]
|
Input["User / AIOps input"] --> Prompt["Prompt contract"]
|
||||||
Prompt --> Agent["Planner / Executor / Verifier / Composer"]
|
Prompt --> Agent["Diagnosis StateGraph Nodes"]
|
||||||
Agent --> Tools["Evidence tools"]
|
Agent --> Tools["Evidence tools"]
|
||||||
Tools --> Invocation["tool_invocation"]
|
Tools --> Invocation["tool_invocation"]
|
||||||
Agent --> StepHook["AgentLoggingHook"]
|
Agent --> StepHook["AgentLoggingHook"]
|
||||||
@@ -41,15 +42,16 @@ flowchart TB
|
|||||||
Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
|
Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
|
||||||
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
|
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
|
||||||
Agent --> Gatekeeper
|
Agent --> Gatekeeper
|
||||||
Invocation --> TraceSummary["ToolTraceSummaryService"]
|
Gatekeeper --> Projection["VerifiedInputNode"]
|
||||||
Gatekeeper --> Verifier["chat_verifier"]
|
Projection --> Verifier["chat_verifier"]
|
||||||
TraceSummary --> Verifier
|
|
||||||
Verifier --> SelfEval["self_evaluation.verifier_evaluation"]
|
Verifier --> SelfEval["self_evaluation.verifier_evaluation"]
|
||||||
|
|
||||||
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
|
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
|
||||||
AiOpsRule --> AiOpsEval["self_evaluation.aiops_rule_evaluation"]
|
AiOpsRule --> AiOpsEval["self_evaluation.aiops_rule_evaluation"]
|
||||||
|
|
||||||
Run --> TraceAPI["DiagnosisTraceService"]
|
Run --> TraceAPI["DiagnosisTraceService"]
|
||||||
|
Agent --> Routing["diagnosis_run.orchestration_trace"]
|
||||||
|
Routing --> TraceAPI
|
||||||
Step --> TraceAPI
|
Step --> TraceAPI
|
||||||
Invocation --> TraceAPI
|
Invocation --> TraceAPI
|
||||||
SelfEval --> TraceAPI
|
SelfEval --> TraceAPI
|
||||||
@@ -126,7 +128,7 @@ sequenceDiagram
|
|||||||
- token count。
|
- token count。
|
||||||
- Verifier 的 JSON 输出摘要。
|
- Verifier 的 JSON 输出摘要。
|
||||||
|
|
||||||
新写入必须带 `run_id`;`session_id` 仍保留用于粗粒度排查和历史兼容。
|
新写入必须带 `run_id`;`session_id` 只用于会话归属和粗粒度排查,不能替代 Run 边界。
|
||||||
|
|
||||||
## 5. Tool Invocation 门禁
|
## 5. Tool Invocation 门禁
|
||||||
|
|
||||||
@@ -161,7 +163,7 @@ error_message
|
|||||||
|
|
||||||
## 6. Gatekeeper 与 Verifier 门禁
|
## 6. Gatekeeper 与 Verifier 门禁
|
||||||
|
|
||||||
Chat Verifier 前置一层 Gatekeeper。Gatekeeper 不调用 LLM,只用代码检查 Executor 输出的证据引用是否真实存在。
|
Chat StateGraph 在 Verifier 前显式执行 Gatekeeper 和 Verified Input。Gatekeeper 不调用 LLM,只用代码检查 Executor 输出的证据引用是否真实存在;Verified Input 只投影通过的 binding。
|
||||||
|
|
||||||
```mermaid
|
```mermaid
|
||||||
flowchart LR
|
flowchart LR
|
||||||
@@ -169,11 +171,12 @@ flowchart LR
|
|||||||
ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
|
ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
|
||||||
EvidenceRefs --> Gatekeeper
|
EvidenceRefs --> Gatekeeper
|
||||||
Gatekeeper --> GateResult["gatekeeper_result"]
|
Gatekeeper --> GateResult["gatekeeper_result"]
|
||||||
Invocation --> Summary["ToolTraceSummaryService"]
|
GateResult --> Projection["VerifiedInputNode"]
|
||||||
Summary --> EvidenceIndex["tool_trace_summary"]
|
Projection --> VerifiedClaims["verified_executor_output"]
|
||||||
|
Projection --> VerifiedEvidence["verified_evidence"]
|
||||||
GateResult --> Verifier["chat_verifier"]
|
GateResult --> Verifier["chat_verifier"]
|
||||||
ExecutorOutput --> Verifier
|
VerifiedClaims --> Verifier
|
||||||
EvidenceIndex --> Verifier
|
VerifiedEvidence --> Verifier
|
||||||
Verifier --> Verdict{"verdict"}
|
Verifier --> Verdict{"verdict"}
|
||||||
Verdict -->|PASS| Composer["chat_composer"]
|
Verdict -->|PASS| Composer["chat_composer"]
|
||||||
Composer --> Pass["输出最终答复"]
|
Composer --> Pass["输出最终答复"]
|
||||||
@@ -215,7 +218,7 @@ Verifier 不再逐字核验 excerpt 真伪;这由 Gatekeeper 完成。Verifier
|
|||||||
diagnosis_run.self_evaluation.verifier_evaluation
|
diagnosis_run.self_evaluation.verifier_evaluation
|
||||||
```
|
```
|
||||||
|
|
||||||
其中同时持久化 `executor_structured_output`、`gatekeeper_result`、`tool_trace_summary`、`prompt_audit` 和 `composer_output`,用于 Trace 回放。
|
其中持久化 verified `executor_structured_output`、`verified_evidence`、`gatekeeper_result`、`prompt_audit` 和 `composer_output`,用于 Trace 回放。Graph 路由另存 `diagnosis_run.orchestration_trace`;当前版本不生成或读取 `tool_trace_summary`。
|
||||||
|
|
||||||
## 7. AIOps 规则门禁
|
## 7. AIOps 规则门禁
|
||||||
|
|
||||||
|
|||||||
@@ -1,11 +1,12 @@
|
|||||||
# 面试一页式架构讲解
|
# 面试一页式架构讲解
|
||||||
|
|
||||||
|
**更新日期**:2026-07-20
|
||||||
**用途**:面试现场 2-5 分钟讲清项目
|
**用途**:面试现场 2-5 分钟讲清项目
|
||||||
**适合场景**:开场介绍、架构追问、Demo 前铺垫
|
**适合场景**:开场介绍、架构追问、Demo 前铺垫
|
||||||
|
|
||||||
## 1. 一句话
|
## 1. 一句话
|
||||||
|
|
||||||
SuperBizAgent 是一个面向企业故障诊断的可追踪 Agent 系统:它把用户问题或 AIOps 告警转换成 Planner、Executor、Gatekeeper、Verifier、Composer 的诊断链路,`sessionId` 保留多轮上下文,`runId` 精确绑定一次诊断运行;所有工具证据、模型步骤、最终答案、自评估和用户反馈都能按 `sessionId + runId` 回放。
|
SuperBizAgent 是一个面向企业故障诊断的可追踪 Agent 系统:复杂 Chat 通过 bounded StateGraph 编排 Planner、Executor、Gatekeeper、Verified Input、Verifier、Composer 和安全 Fallback,`sessionId` 保留多轮上下文,`runId` 精确绑定一次诊断运行;工具证据、模型步骤、Graph 路由、自评估、最终答案和用户反馈都能按 `sessionId + runId` 回放。
|
||||||
|
|
||||||
## 2. 一张图
|
## 2. 一张图
|
||||||
|
|
||||||
@@ -16,7 +17,7 @@ flowchart TB
|
|||||||
API --> Chat["ChatService"]
|
API --> Chat["ChatService"]
|
||||||
API --> AiOps["AiOpsService"]
|
API --> AiOps["AiOpsService"]
|
||||||
|
|
||||||
Chat --> ChatFlow["Chat: Planner -> Executor -> Gatekeeper -> Verifier -> Composer"]
|
Chat --> ChatFlow["Chat StateGraph: Planner -> Executor -> Gatekeeper -> Verified Input -> Verifier -> Composer/Fallback"]
|
||||||
AiOps --> AiOpsFlow["AIOps: Supervisor -> Planner / Executor"]
|
AiOps --> AiOpsFlow["AIOps: Supervisor -> Planner / Executor"]
|
||||||
|
|
||||||
ChatFlow --> Tools["Evidence Tools"]
|
ChatFlow --> Tools["Evidence Tools"]
|
||||||
@@ -39,10 +40,14 @@ flowchart TB
|
|||||||
Trace --> Step["agent_step"]
|
Trace --> Step["agent_step"]
|
||||||
Trace --> Invocation["tool_invocation"]
|
Trace --> Invocation["tool_invocation"]
|
||||||
|
|
||||||
Invocation --> Verifier["Verifier / Rule Evaluation"]
|
Invocation --> Gate["Gatekeeper / Rule Evaluation"]
|
||||||
|
Gate --> Projection["Verified Input"]
|
||||||
|
Projection --> Verifier["Verifier"]
|
||||||
Verifier --> SelfEval["self_evaluation"]
|
Verifier --> SelfEval["self_evaluation"]
|
||||||
|
|
||||||
|
Run --> RouteTrace["orchestration_trace"]
|
||||||
Run --> TraceAPI["GET /api/diagnosis/{sessionId}/trace?runId=..."]
|
Run --> TraceAPI["GET /api/diagnosis/{sessionId}/trace?runId=..."]
|
||||||
|
RouteTrace --> TraceAPI
|
||||||
Step --> TraceAPI
|
Step --> TraceAPI
|
||||||
Invocation --> TraceAPI
|
Invocation --> TraceAPI
|
||||||
SelfEval --> TraceAPI
|
SelfEval --> TraceAPI
|
||||||
@@ -56,21 +61,21 @@ flowchart TB
|
|||||||
```text
|
```text
|
||||||
这个项目不是把问题直接丢给大模型,而是把诊断拆成可审计的执行链路。
|
这个项目不是把问题直接丢给大模型,而是把诊断拆成可审计的执行链路。
|
||||||
|
|
||||||
Chat 复杂问题走 Planner -> Executor -> Gatekeeper -> Verifier -> Composer:
|
Chat 复杂问题走 bounded StateGraph:
|
||||||
Planner 负责拆解,Executor 只负责调用知识库、日志和指标工具并提炼带证据引用的微观事实;Gatekeeper 用代码核对 invocation、raw_path 和 excerpt 是否真实;Verifier 判断这些事实能否由已验真的证据推出;Composer 只把允许表达的结论写成最终答案。
|
Planner 负责拆解,Executor 调用知识库、日志和指标工具并提炼带证据引用的微观事实;Gatekeeper 用代码核对 invocation、raw_path 和 excerpt 是否真实;Verified Input 只投影通过的 binding;Verifier 判断这些事实能否由已验真的证据推出;Composer 只把允许表达的结论写成最终答案,不可恢复分支由 Fallback 生成安全答复。
|
||||||
|
|
||||||
AIOps 告警入口走 Supervisor 调度 Planner/Executor:
|
AIOps 告警入口走 Supervisor 调度 Planner/Executor:
|
||||||
如果请求里有 alert payload,系统会进入 PAYLOAD_TARGETED 模式,报告必须聚焦这个告警,而不是被当前环境中的其他活跃告警带偏。
|
如果请求里有 alert payload,系统会进入 PAYLOAD_TARGETED 模式,报告必须聚焦这个告警,而不是被当前环境中的其他活跃告警带偏。
|
||||||
|
|
||||||
会话元数据会落到 chat_session,每次诊断运行会落到 diagnosis_run,步骤和工具明细通过 run_id 关联。
|
会话元数据会落到 chat_session,每次诊断运行会落到 diagnosis_run,步骤和工具明细通过 run_id 关联;证据质量写入 self_evaluation,Graph 路由独立写入 orchestration_trace。
|
||||||
所以我可以用 sessionId + runId 精确回放:模型怎么规划、调了哪些工具、工具返回什么、Gatekeeper 怎么验真、Verifier 怎么判定、Composer 最后怎么表达、用户最后是否反馈有用。
|
所以我可以用 sessionId + runId 精确回放:模型怎么规划、调了哪些工具、Gatekeeper 怎么验真、Graph 为什么重试或降级、Verifier 怎么判定、Composer/Fallback 如何结束、用户最后是否反馈有用。
|
||||||
```
|
```
|
||||||
|
|
||||||
## 4. 五个亮点
|
## 4. 五个亮点
|
||||||
|
|
||||||
| 亮点 | 怎么讲 |
|
| 亮点 | 怎么讲 |
|
||||||
|---|---|
|
|---|---|
|
||||||
| 可追踪 Agent | 每次诊断都有 `runId`,Trace API 可以回放 run、step、tool;同一 `sessionId` 可有多次独立 run |
|
| 可追踪 StateGraph | 每次诊断都有 `runId`,Trace API 可以回放 run、step、tool 和 `run.orchestrationTrace`;同一 `sessionId` 可有多次独立 run |
|
||||||
| 显式工具证据链 | `lookup_knowledge`、日志、指标都记录到 `tool_invocation` |
|
| 显式工具证据链 | `lookup_knowledge`、日志、指标都记录到 `tool_invocation` |
|
||||||
| RAG 工程化 | L0 降级为 hint,Spring AI VectorStore 做主检索,SDK fallback 保底 |
|
| RAG 工程化 | L0 降级为 hint,Spring AI VectorStore 做主检索,SDK fallback 保底 |
|
||||||
| 质量门禁 | Chat Gatekeeper 验引用、Verifier 判可推导、Composer 控表达,AIOps rule evaluation 控制告警聚焦 |
|
| 质量门禁 | Chat Gatekeeper 验引用、Verifier 判可推导、Composer 控表达,AIOps rule evaluation 控制告警聚焦 |
|
||||||
|
|||||||
@@ -142,7 +142,7 @@ post-retrieval 层再把检索候选归一为:
|
|||||||
- 给 Agent 输出 completeness hint。
|
- 给 Agent 输出 completeness hint。
|
||||||
- 写入 `tool_invocation.relevance_level`。
|
- 写入 `tool_invocation.relevance_level`。
|
||||||
- 给 Gatekeeper 提供 `evidence_refs` 引用验真源。
|
- 给 Gatekeeper 提供 `evidence_refs` 引用验真源。
|
||||||
- 给 Verifier 构造 `tool_trace_summary` 审计导航。
|
- 由 Gatekeeper 核验后,经 `VerifiedInputNode` 给 Verifier 构造最小 `verified_evidence` 投影。
|
||||||
- 供 EvaluationService 计算 evidence score。
|
- 供 EvaluationService 计算 evidence score。
|
||||||
|
|
||||||
## 6. 文档切片和 metadata
|
## 6. 文档切片和 metadata
|
||||||
@@ -180,11 +180,13 @@ flowchart LR
|
|||||||
Recorder --> Invocation["tool_invocation"]
|
Recorder --> Invocation["tool_invocation"]
|
||||||
Invocation --> Trace["DiagnosisTraceService"]
|
Invocation --> Trace["DiagnosisTraceService"]
|
||||||
Invocation --> Gatekeeper["ExecutorGatekeeperService"]
|
Invocation --> Gatekeeper["ExecutorGatekeeperService"]
|
||||||
Invocation --> Summary["ToolTraceSummaryService"]
|
Gatekeeper --> Projection["VerifiedInputNode"]
|
||||||
Summary --> Verifier["chat_verifier"]
|
Projection --> Verifier["chat_verifier"]
|
||||||
Invocation --> Eval["EvaluationService / RAG eval"]
|
Invocation --> Eval["EvaluationService / RAG eval"]
|
||||||
```
|
```
|
||||||
|
|
||||||
|
当前 StateGraph 不生成或读取 `tool_trace_summary`,也不会把完整工具调用摘要输入 Verifier。
|
||||||
|
|
||||||
`tool_invocation` 中与检索相关的字段:
|
`tool_invocation` 中与检索相关的字段:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
# 会话与 Trace 生命周期
|
# 会话与 Trace 生命周期
|
||||||
|
|
||||||
**更新日期**:2026-07-10
|
**更新日期**:2026-07-17
|
||||||
**状态**:当前可运行架构
|
**状态**:当前可运行架构
|
||||||
**参考历史文档**:`archive/2026-07-05-legacy/session-management.md`
|
**参考历史文档**:`archive/2026-07-05-legacy/session-management.md`
|
||||||
|
|
||||||
@@ -18,7 +18,7 @@ chat_session(sessionId)
|
|||||||
- `sessionId` 表示多轮会话目录和 Redis 上下文。
|
- `sessionId` 表示多轮会话目录和 Redis 上下文。
|
||||||
- `runId` 表示一次可回放诊断执行。
|
- `runId` 表示一次可回放诊断执行。
|
||||||
- `DiagnosisTraceService` 聚合一个 run 的主记录、步骤和工具调用,形成可回放 Trace。
|
- `DiagnosisTraceService` 聚合一个 run 的主记录、步骤和工具调用,形成可回放 Trace。
|
||||||
- `diagnosis_session` 只保留为历史兼容和回滚表。
|
- 运行时不再映射、读取或写入 `diagnosis_session`;数据库中的旧表不属于当前版本契约。
|
||||||
|
|
||||||
## 2. 生命周期总图
|
## 2. 生命周期总图
|
||||||
|
|
||||||
@@ -29,7 +29,7 @@ flowchart TD
|
|||||||
Session --> Run["create diagnosis_run(runId)"]
|
Session --> Run["create diagnosis_run(runId)"]
|
||||||
Run --> Running["run.status = RUNNING"]
|
Run --> Running["run.status = RUNNING"]
|
||||||
|
|
||||||
Running --> Agent["Agent workflow"]
|
Running --> Agent["Chat StateGraph / AIOps workflow"]
|
||||||
Agent --> Context["execution context(sessionId, runId)"]
|
Agent --> Context["execution context(sessionId, runId)"]
|
||||||
Context --> StepHook["AgentLoggingHook"]
|
Context --> StepHook["AgentLoggingHook"]
|
||||||
StepHook --> Step["agent_step(session_id, run_id)"]
|
StepHook --> Step["agent_step(session_id, run_id)"]
|
||||||
@@ -37,7 +37,8 @@ flowchart TD
|
|||||||
Tool --> Invocation["tool_invocation(session_id, run_id)"]
|
Tool --> Invocation["tool_invocation(session_id, run_id)"]
|
||||||
Invocation --> Gatekeeper["Gatekeeper evidence validation"]
|
Invocation --> Gatekeeper["Gatekeeper evidence validation"]
|
||||||
|
|
||||||
Agent --> Final{"workflow result"}
|
Agent --> GraphTrace["Chat: save orchestration_trace"]
|
||||||
|
GraphTrace --> Final{"workflow result"}
|
||||||
Final -->|success| Success["run.status = SUCCESS, answer saved"]
|
Final -->|success| Success["run.status = SUCCESS, answer saved"]
|
||||||
Final -->|failed| Failed["run.status = FAILED"]
|
Final -->|failed| Failed["run.status = FAILED"]
|
||||||
|
|
||||||
@@ -59,8 +60,7 @@ flowchart TD
|
|||||||
|
|
||||||
- 同一个 `sessionId` 可以贯穿多轮 Chat。
|
- 同一个 `sessionId` 可以贯穿多轮 Chat。
|
||||||
- 每次有效 Chat/AIOps 执行都会创建新的 `runId`。
|
- 每次有效 Chat/AIOps 执行都会创建新的 `runId`。
|
||||||
- Trace 和 Feedback 新客户端应传 `runId`;只传 `sessionId` 时兼容解析 latest run。
|
- Trace 和 Feedback 必须同时传 `sessionId + runId`;后端不解析 latest run,也不回退旧表。
|
||||||
- latest run 排序使用 `diagnosis_run.created_at DESC, id DESC`,不使用 `updated_at`。
|
|
||||||
|
|
||||||
## 4. 运行状态流转
|
## 4. 运行状态流转
|
||||||
|
|
||||||
@@ -81,6 +81,7 @@ stateDiagram-v2
|
|||||||
| `status` | `diagnosis_run` | 单次运行执行状态 |
|
| `status` | `diagnosis_run` | 单次运行执行状态 |
|
||||||
| `answer` | `diagnosis_run` | 本次运行最终报告或答复 |
|
| `answer` | `diagnosis_run` | 本次运行最终报告或答复 |
|
||||||
| `self_evaluation` | `diagnosis_run` | 本次运行系统自评估 JSON |
|
| `self_evaluation` | `diagnosis_run` | 本次运行系统自评估 JSON |
|
||||||
|
| `orchestration_trace` | `diagnosis_run` | Chat StateGraph 路由摘要;包含 version、transitions、final node、termination reason、degraded 和 evidence retry count |
|
||||||
| `feedback` | `diagnosis_run` | 本次运行用户反馈 |
|
| `feedback` | `diagnosis_run` | 本次运行用户反馈 |
|
||||||
|
|
||||||
`feedback` 不修改 `status`。一个执行成功但用户标记 `not_useful` 的 run,仍然应该是 `SUCCESS + feedback=not_useful`。
|
`feedback` 不修改 `status`。一个执行成功但用户标记 `not_useful` 的 run,仍然应该是 `SUCCESS + feedback=not_useful`。
|
||||||
@@ -115,7 +116,7 @@ ToolInvocationRecorder
|
|||||||
-> retrieval_details / evidence_refs
|
-> retrieval_details / evidence_refs
|
||||||
```
|
```
|
||||||
|
|
||||||
Verifier、Gatekeeper 和 EvaluationService 应按 `run_id` 读取工具调用,避免同一 session 的其他 run 参与评分或证据校验。
|
Gatekeeper 和 EvaluationService 按 `run_id` 读取工具调用,避免同一 session 的其他 run 参与评分或证据校验。Verifier 只读取 `VerifiedInputNode` 生成的 verified projection,不直接读取完整工具调用列表。
|
||||||
|
|
||||||
## 7. Trace API 聚合
|
## 7. Trace API 聚合
|
||||||
|
|
||||||
@@ -128,24 +129,29 @@ GET /api/diagnosis/{sessionId}/trace?runId=run-...
|
|||||||
|
|
||||||
```text
|
```text
|
||||||
diagnosis_run by sessionId + runId
|
diagnosis_run by sessionId + runId
|
||||||
|
+ run.orchestrationTrace parsed from diagnosis_run.orchestration_trace
|
||||||
+ chat_session metadata when available
|
+ chat_session metadata when available
|
||||||
+ agent_step where run_id = runId, ordered by the Trace API
|
+ agent_step where run_id = runId, ordered by the Trace API
|
||||||
+ tool_invocation where run_id = runId order by id
|
+ tool_invocation where run_id = runId order by id
|
||||||
-> DiagnosisTraceResponse
|
-> DiagnosisTraceResponse
|
||||||
```
|
```
|
||||||
|
|
||||||
当 `runId` 缺失时,Trace API 为兼容旧客户端解析最新 run,并在响应中返回 resolved `runId`。当 `runId` 属于其他 `sessionId` 时,API 必须拒绝,不能泄漏其他会话的 Trace。
|
当 `runId` 缺失时,Trace API 直接拒绝请求。当 `runId` 属于其他 `sessionId` 时,API 同样拒绝,不能泄漏其他会话的 Trace。
|
||||||
|
|
||||||
|
`run.orchestrationTrace` 只属于精确 Run 投影,响应不再提供兼容 `session` 对象。它解释 Graph 路由;`selfEvaluation` 解释证据/答案质量;`steps` 和 `toolInvocations` 保存详细执行证据,三者职责互不替代。非 StateGraph Run 的该字段可以为空。
|
||||||
|
|
||||||
## 8. Chat 与 AIOps 差异
|
## 8. Chat 与 AIOps 差异
|
||||||
|
|
||||||
| 维度 | Chat | AIOps |
|
| 维度 | Chat | AIOps |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| `agent_flow` | `CHAT` | `AI_OPS` |
|
| `agent_flow` | `CHAT` | `AI_OPS` |
|
||||||
| 编排方式 | `SequentialAgent`: Planner -> Executor -> Gatekeeper -> Verifier -> Composer | `SupervisorAgent`: Planner + Executor |
|
| 编排方式 | bounded `StateGraph`: Planner / Executor / Gatekeeper / Verified Input / Verifier / Composer / Fallback | `SupervisorAgent`: Planner + Executor |
|
||||||
| 自评估 | `rule_evaluation` + `verifier_evaluation` | `aiops_rule_evaluation` |
|
| 自评估 | `rule_evaluation` + `verifier_evaluation` | `aiops_rule_evaluation` |
|
||||||
| 答案字段 | Chat 最终答复 | 告警分析报告 |
|
| 答案字段 | Chat 最终答复 | 告警分析报告 |
|
||||||
| runId 暴露 | `/api/chat` JSON response | `/api/ai_ops` SSE metadata message |
|
| runId 暴露 | `/api/chat` JSON response | `/api/ai_ops` SSE metadata message |
|
||||||
|
|
||||||
|
Chat StateGraph 的权威自动化验收分三层:`DiagnosisGraphWorkflowTest` 验证路由,`DiagnosisGraphNodeContractTest` 验证真实 Node 输入输出,`ChatServiceGraphIntegrationTest` 验证 Run 生命周期、Trace 持久化和对外集成。
|
||||||
|
|
||||||
## 9. 清理与边界
|
## 9. 清理与边界
|
||||||
|
|
||||||
- Redis 会话历史用于多轮上下文,不是长期审计记录。
|
- Redis 会话历史用于多轮上下文,不是长期审计记录。
|
||||||
@@ -157,4 +163,4 @@ diagnosis_run by sessionId + runId
|
|||||||
|
|
||||||
1. Trace API 增加更结构化的 `self_evaluation` 展示。
|
1. Trace API 增加更结构化的 `self_evaluation` 展示。
|
||||||
2. `agent_step` 与 `tool_invocation.step_id` 建立更严格关联。
|
2. `agent_step` 与 `tool_invocation.step_id` 建立更严格关联。
|
||||||
3. 旧 `diagnosis_session` 只读观察期结束后,再评估数据库层面的约束收紧或归档策略。
|
3. 数据库中的旧 `diagnosis_session` 表按独立数据治理任务决定是否物理删除;当前应用不再依赖它。
|
||||||
|
|||||||
@@ -0,0 +1,199 @@
|
|||||||
|
# Chat StateGraph 运行时架构
|
||||||
|
|
||||||
|
**更新日期**:2026-07-20
|
||||||
|
**状态**:当前复杂 Chat 诊断的权威运行时架构
|
||||||
|
**适用范围**:`POST /api/chat` 的复杂诊断路径;简单 Chat 和 AIOps 使用各自链路
|
||||||
|
|
||||||
|
## 1. 架构定位
|
||||||
|
|
||||||
|
复杂 Chat 已单轨切换为 Spring AI Alibaba bounded `StateGraph`。`ChatService` 负责 Run 生命周期和持久化,`ChatDiagnosisGraphRuntime` 负责执行 Graph,`DiagnosisGraphFactory` 负责声明 Node 与条件边,Agent/Java Node 负责各自的语义任务或确定性校验。
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart LR
|
||||||
|
API["POST /api/chat"] --> Chat["ChatService.executeChatComplex"]
|
||||||
|
Chat --> Session["chat_session"]
|
||||||
|
Chat --> Run["diagnosis_run: RUNNING"]
|
||||||
|
Chat --> Actions["DiagnosisRealGraphActionsFactory"]
|
||||||
|
Actions --> Runtime["ChatDiagnosisGraphRuntime"]
|
||||||
|
Runtime --> Factory["DiagnosisGraphFactory"]
|
||||||
|
Factory --> Graph["Compiled StateGraph"]
|
||||||
|
|
||||||
|
Graph --> Planner["PlannerNodeAdapter"]
|
||||||
|
Graph --> Executor["ExecutorNodeAdapter"]
|
||||||
|
Graph --> Gatekeeper["GatekeeperNode"]
|
||||||
|
Graph --> Projection["VerifiedInputNode"]
|
||||||
|
Graph --> Verifier["VerifierNodeAdapter"]
|
||||||
|
Graph --> Retry["EvidenceRetryPrepareNode"]
|
||||||
|
Graph --> Composer["ComposerNodeAdapter"]
|
||||||
|
Graph --> Fallback["FallbackNode"]
|
||||||
|
|
||||||
|
Executor --> Tools["lookup_knowledge / logs / metrics"]
|
||||||
|
Tools --> Invocations["tool_invocation"]
|
||||||
|
Planner --> Steps["agent_step"]
|
||||||
|
Executor --> Steps
|
||||||
|
Verifier --> Steps
|
||||||
|
Composer --> Steps
|
||||||
|
|
||||||
|
Graph --> Mapper["DiagnosisGraphResultMapper"]
|
||||||
|
Graph --> TraceBuilder["DiagnosisOrchestrationTraceBuilder"]
|
||||||
|
Mapper --> SelfEval["diagnosis_run.self_evaluation"]
|
||||||
|
TraceBuilder --> RouteTrace["diagnosis_run.orchestration_trace"]
|
||||||
|
Chat --> RunDone["diagnosis_run: SUCCESS / FAILED"]
|
||||||
|
RunDone --> TraceAPI["exact Run Trace API"]
|
||||||
|
SelfEval --> TraceAPI
|
||||||
|
RouteTrace --> TraceAPI
|
||||||
|
Steps --> TraceAPI
|
||||||
|
Invocations --> TraceAPI
|
||||||
|
```
|
||||||
|
|
||||||
|
## 2. 运行生命周期
|
||||||
|
|
||||||
|
一次复杂 Chat 运行按以下顺序执行:
|
||||||
|
|
||||||
|
1. `ChatService` 解析或创建 `sessionId`,生成唯一 `runId`。
|
||||||
|
2. 确保 `chat_session` 元数据存在,并创建 `diagnosis_run`,初始状态为 `RUNNING`、`agent_flow=CHAT`。
|
||||||
|
3. 构建 Planner、Executor、Verifier、Composer 四个 `ReactAgent`,再由 `DiagnosisRealGraphActionsFactory` 组合 Java Nodes。
|
||||||
|
4. `ChatDiagnosisGraphRuntime` 以 `runId` 作为 Graph `threadId`,将 `sessionId/runId` 放入 `RunnableConfig.metadata`。
|
||||||
|
5. `DiagnosisGraphFactory` 编译 StateGraph 并执行,Graph recursion limit 固定为 32。
|
||||||
|
6. Graph 返回非空 `final_answer` 后,`DiagnosisGraphResultMapper` 生成 verifier evaluation,`DiagnosisOrchestrationTraceBuilder` 压缩路由摘要。
|
||||||
|
7. `ChatService` 保存答案、耗时、步骤数、工具数、自评估和编排摘要,将 Run 标记为 `SUCCESS`。
|
||||||
|
8. 未处理异常会尽力保存 partial state/partial trace,再将 Run 标记为 `FAILED`;能够生成安全 Fallback 的路径仍是 `SUCCESS`,并通过 `degraded=true` 表达质量降级。
|
||||||
|
9. `finally` 清理本轮检索追踪和 session/run ThreadLocal,避免跨 Run 污染。
|
||||||
|
|
||||||
|
## 3. Graph 拓扑
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TD
|
||||||
|
Start([START]) --> Planner[Planner]
|
||||||
|
Planner -->|COMPLETED| Executor[Executor]
|
||||||
|
Planner -->|technical retry once| Planner
|
||||||
|
Planner -->|non-retryable / exhausted| Fallback[Fallback]
|
||||||
|
|
||||||
|
Executor -->|COMPLETED| Gatekeeper[Gatekeeper]
|
||||||
|
Executor -->|INVALID_OUTPUT / TOOL_BLOCKED / FAILED| Fallback
|
||||||
|
|
||||||
|
Gatekeeper -->|PASS| VerifiedInput[Verified Input]
|
||||||
|
Gatekeeper -->|LOW_CONFID + verified bindings| VerifiedInput
|
||||||
|
Gatekeeper -->|REJECT / no verified binding| Fallback
|
||||||
|
|
||||||
|
VerifiedInput --> Verifier[Verifier]
|
||||||
|
Verifier -->|technical retry once| Verifier
|
||||||
|
Verifier -->|LOW_CONFID + critical gap + retry allowed| EvidenceRetry[Evidence Retry]
|
||||||
|
EvidenceRetry -->|EVIDENCE_GAP_ONLY| Planner
|
||||||
|
Verifier -->|completed and no retry| Composer[Composer]
|
||||||
|
Verifier -->|non-retryable / exhausted| Fallback
|
||||||
|
|
||||||
|
Composer -->|COMPLETED| End([END])
|
||||||
|
Composer -->|technical retry once| Composer
|
||||||
|
Composer -->|non-retryable / exhausted| Fallback
|
||||||
|
Fallback --> End
|
||||||
|
```
|
||||||
|
|
||||||
|
路由规则:
|
||||||
|
|
||||||
|
| 节点 | 继续条件 | 重试 | 安全终止 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| Planner | `COMPLETED` 进入 Executor | `INVALID_OUTPUT` / `RETRYABLE_FAILED` 最多一次技术重试 | `NON_RETRYABLE_FAILED` 或重试耗尽进入 Fallback |
|
||||||
|
| Executor | 只有 `COMPLETED` 进入 Gatekeeper | 不做 Graph 技术重试 | 非法输出、工具阻断或执行失败进入 Fallback |
|
||||||
|
| Gatekeeper | `PASS`,或 `LOW_CONFID` 且至少一个 verified binding | 不重试 | `REJECT` 或零 verified binding 进入 Fallback |
|
||||||
|
| Verifier | 完成后由 `effective_verdict` 决定 Composer 或补证据 | 技术失败最多一次;证据补查最多一次 | 不可重试失败或技术重试耗尽进入 Fallback |
|
||||||
|
| Composer | `COMPLETED` 结束 | 技术失败最多一次 | 不可重试失败或重试耗尽进入 Fallback |
|
||||||
|
| Fallback | 生成非空确定性安全答复 | 不重试 | 直接结束并标记 `degraded=true` |
|
||||||
|
|
||||||
|
Evidence retry 只有同时满足以下条件才发生:
|
||||||
|
|
||||||
|
- `evidence_retry_count < 1`。
|
||||||
|
- Gatekeeper 给出的 verifier verdict ceiling 仍允许 `PASS`。
|
||||||
|
- Verifier 输出包含可提取的 critical evidence gap。
|
||||||
|
|
||||||
|
补证据时 Planner 进入 `EVIDENCE_GAP_ONLY`,`planner_retry_count` 重置;Executor 只执行增量查询,但重新输出完整 `executor_evidence_v2` 快照。
|
||||||
|
|
||||||
|
## 4. 状态与执行边界
|
||||||
|
|
||||||
|
### Graph State
|
||||||
|
|
||||||
|
Graph State 只保存跨 Node 的控制信息和结构化结果:
|
||||||
|
|
||||||
|
- `diagnosis_context`、`planner_plan`、`executor_output`。
|
||||||
|
- `gatekeeper_result`、`verified_executor_output`、`verified_evidence`。
|
||||||
|
- `verifier_output`、`composer_output`、`final_answer`。
|
||||||
|
- Planner/Verifier/Composer 技术重试计数和 `evidence_retry_count`。
|
||||||
|
- `orchestration_events` 有界追加事件。
|
||||||
|
|
||||||
|
默认状态键使用 replace strategy,只有 `orchestration_events` 使用 append strategy。事件在 Graph 边界存为 classloader-neutral Map:`node/outcome/reason_code/attempt`,避免 DevTools restart classloader 造成 record 类型身份不一致。
|
||||||
|
|
||||||
|
### RunnableConfig
|
||||||
|
|
||||||
|
外层 Graph config 使用:
|
||||||
|
|
||||||
|
- `threadId = runId`。
|
||||||
|
- metadata 包含 `sessionId` 和 `runId`。
|
||||||
|
|
||||||
|
调用 nested `ReactAgent` 时,`ReactAgentDiagnosisInvoker` 创建独立 config,只保留业务身份与 store,不向子 Agent 传播外层 Graph 的 human-feedback、state-update、checkpoint/resume 控制 metadata,避免父 Graph 恢复语义污染子 Graph。
|
||||||
|
|
||||||
|
## 5. 证据信任边界
|
||||||
|
|
||||||
|
```text
|
||||||
|
Executor output
|
||||||
|
-> source_invocation_id + raw_path + evidence_excerpt
|
||||||
|
-> GatekeeperNode / ExecutorGatekeeperService
|
||||||
|
-> 按当前 runId 读取 tool_invocation
|
||||||
|
-> 验证 invocation ownership、raw_path、excerpt
|
||||||
|
-> VerifiedInputNode
|
||||||
|
-> 只投影通过的 claims/bindings/evidence
|
||||||
|
-> VerifierNodeAdapter
|
||||||
|
-> 只判断已验真证据是否支持 claim
|
||||||
|
-> ComposerNodeAdapter
|
||||||
|
-> 只表达允许输出的结论、限制和建议
|
||||||
|
```
|
||||||
|
|
||||||
|
Verifier 不读取完整工具 Trace,不执行新检索,也不读取 Skill 正文。当前版本不生成或读取 `tool_trace_summary`。
|
||||||
|
|
||||||
|
## 6. Run 级审计模型
|
||||||
|
|
||||||
|
| 审计层 | 存储/API | 回答的问题 |
|
||||||
|
|---|---|---|
|
||||||
|
| 执行明细 | `agent_step`、`tool_invocation` | 模型和工具实际做了什么? |
|
||||||
|
| 证据与答案质量 | `diagnosis_run.self_evaluation` | 引用是否真实、claim 是否可推导、Prompt/Gatekeeper 版本是什么? |
|
||||||
|
| Graph 路由 | `diagnosis_run.orchestration_trace` / `run.orchestrationTrace` | 走了哪些 Node、为何重试或降级、在哪里结束? |
|
||||||
|
|
||||||
|
`orchestration_trace` 当前结构:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"version": "stategraph-v1",
|
||||||
|
"transitions": [
|
||||||
|
{"from": "planner", "to": "executor", "reason_code": "completed", "attempt": 1}
|
||||||
|
],
|
||||||
|
"final_node": "composer",
|
||||||
|
"termination_reason": "composer_completed",
|
||||||
|
"degraded": false,
|
||||||
|
"evidence_retry_count": 0
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
该字段只属于 exact Run 投影,响应不提供兼容 `session` 投影;非 StateGraph Run 可以为空。
|
||||||
|
|
||||||
|
## 7. 验收层次
|
||||||
|
|
||||||
|
| 测试层 | 权威测试 | 覆盖 |
|
||||||
|
|---|---|---|
|
||||||
|
| Workflow | `DiagnosisGraphWorkflowTest` | 全部分支、有限重试、补证据和 Fallback |
|
||||||
|
| Node Contract | `DiagnosisGraphNodeContractTest` | 真实 Node 的输入投影、输出状态和证据边界 |
|
||||||
|
| Runtime | `ChatDiagnosisGraphRuntimeTest` | config、最终状态、partial failure/trace |
|
||||||
|
| Chat Integration | `ChatServiceGraphIntegrationTest` | Run 生命周期、答案、自评估、编排摘要和失败持久化 |
|
||||||
|
| Trace Contract | `DiagnosisTraceServiceTest` | exact run ownership 和 `run.orchestrationTrace` 投影 |
|
||||||
|
| Demo Contract | `InterviewDemoScriptContractTest` | exact runId、Graph 字段和 summary 输出 |
|
||||||
|
|
||||||
|
## 8. 关键代码
|
||||||
|
|
||||||
|
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/graph/diagnosis/ChatDiagnosisGraphRuntime.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/graph/diagnosis/DiagnosisGraphFactory.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/graph/diagnosis/DiagnosisGraphRouter.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/graph/diagnosis/DiagnosisGraphState.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/graph/diagnosis/DiagnosisRealGraphActionsFactory.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/graph/diagnosis/ReactAgentDiagnosisInvoker.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/graph/diagnosis/DiagnosisGraphResultMapper.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/graph/diagnosis/DiagnosisOrchestrationTraceBuilder.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
|
||||||
@@ -0,0 +1,19 @@
|
|||||||
|
# 2026-07-20 MVP 文档清理归档
|
||||||
|
|
||||||
|
本目录保存已被当前 StateGraph、Executor Evidence Pipeline 和 Run-only v2 架构替代的历史材料。归档文件仅用于追溯,不代表当前运行时、接口或数据模型契约。
|
||||||
|
|
||||||
|
## 归档清单
|
||||||
|
|
||||||
|
| 原路径 | 归档文件 | 归档原因 | 当前依据 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `mvp/.backup/database-design-backup-20240622.md` | [database-design-backup-20240622.md](database/database-design-backup-20240622.md) | 早期数据库设计备份,模型和表关系已过时 | [data-model.md](../../architecture/data-model.md) |
|
||||||
|
| `mvp/issues/design-notes/executor-self-evidence-loop-design-note.md` | [executor-self-evidence-loop-design-note.md](issues/design-notes/executor-self-evidence-loop-design-note.md) | 早期问题分析已被可执行证据链契约吸收 | [executor-evidence-pipeline-refactor.md](../../architecture/executor-evidence-pipeline-refactor.md) |
|
||||||
|
| `mvp/issues/design-notes/executor-structured-output-v2.md` | [executor-structured-output-v2.md](issues/design-notes/executor-structured-output-v2.md) | 分阶段实施说明已完成并由当前契约和 OpenSpec 接管 | [executor-evidence-pipeline-refactor.md](../../architecture/executor-evidence-pipeline-refactor.md)、[OpenSpec 主规格](../../../openspec/specs/) |
|
||||||
|
| `mvp/tables/诊断会话表-diagnosis_session.md` | [诊断会话表-diagnosis_session.md](tables/诊断会话表-diagnosis_session.md) | 运行时已不再映射、读取或写入旧会话级诊断表 | [session-trace-lifecycle.md](../../architecture/session-trace-lifecycle.md)、[诊断运行表-diagnosis_run.md](../../tables/诊断运行表-diagnosis_run.md) |
|
||||||
|
|
||||||
|
## 使用约束
|
||||||
|
|
||||||
|
- 当前架构从 `mvp/architecture/README.md` 进入。
|
||||||
|
- 当前问题状态从 `mvp/issues/README.md` 进入。
|
||||||
|
- 当前表模型从 `mvp/tables/README.md` 进入。
|
||||||
|
- 归档中的兼容、回退和阶段状态描述均为历史快照,不得用于推导当前行为。
|
||||||
+2
@@ -1,5 +1,7 @@
|
|||||||
# 数据库设计文档
|
# 数据库设计文档
|
||||||
|
|
||||||
|
> 归档说明:这是 2024 年数据库设计备份,已被当前 `chat_session + diagnosis_run + run-scoped trace detail` 模型替代,仅用于历史追溯。
|
||||||
|
|
||||||
## 一、设计原则
|
## 一、设计原则
|
||||||
|
|
||||||
### 1.1 核心原则
|
### 1.1 核心原则
|
||||||
+2
@@ -1,5 +1,7 @@
|
|||||||
# Executor 自证循环与证据摘要链路设计记录
|
# Executor 自证循环与证据摘要链路设计记录
|
||||||
|
|
||||||
|
> 归档说明:本文是实施前的问题分析。当前实现以 `mvp/architecture/executor-evidence-pipeline-refactor.md` 和 OpenSpec 主规格为准。
|
||||||
|
|
||||||
**状态**:已形成方向,待创建 OpenSpec change
|
**状态**:已形成方向,待创建 OpenSpec change
|
||||||
**严重程度**:高
|
**严重程度**:高
|
||||||
**记录时间**:2026-07-07
|
**记录时间**:2026-07-07
|
||||||
+2
@@ -1,5 +1,7 @@
|
|||||||
# Executor Structured Output V2 可执行设计与实施 Issue
|
# Executor Structured Output V2 可执行设计与实施 Issue
|
||||||
|
|
||||||
|
> 归档说明:本文记录旧的分阶段实施方案,其中的兼容期和待启动状态不再适用。当前实现以 `mvp/architecture/executor-evidence-pipeline-refactor.md` 和 OpenSpec 主规格为准。
|
||||||
|
|
||||||
**状态**:阶段四待启动,前三阶段已归档并提交
|
**状态**:阶段四待启动,前三阶段已归档并提交
|
||||||
**严重程度**:高
|
**严重程度**:高
|
||||||
**创建日期**:2026-07-07
|
**创建日期**:2026-07-07
|
||||||
+4
-2
@@ -1,7 +1,9 @@
|
|||||||
# 诊断会话表:diagnosis_session
|
# 诊断会话表:diagnosis_session
|
||||||
|
|
||||||
**状态**:历史兼容和回滚表
|
> 归档说明:本文记录 Run-only v2 之前的旧表语义。当前 Java 运行时不再映射、读取或写入 `diagnosis_session`;数据库中物理表是否保留由独立数据治理任务决定。
|
||||||
**来源**:`V005__create_session_storage.sql`、`V008__add_answer_to_diagnosis_session.sql`、`DiagnosisSession`
|
|
||||||
|
**状态**:历史快照,不属于当前版本契约
|
||||||
|
**历史来源**:`V005__create_session_storage.sql`、`V008__add_answer_to_diagnosis_session.sql`
|
||||||
|
|
||||||
## 定位
|
## 定位
|
||||||
|
|
||||||
+9
-2
@@ -8,7 +8,7 @@
|
|||||||
- `interview-walkthrough.md`:面试讲解话术。
|
- `interview-walkthrough.md`:面试讲解话术。
|
||||||
- `evidence-pipeline-scenarios.md`:PASS / LOW_CONFID / REJECT / no-evidence 场景矩阵。
|
- `evidence-pipeline-scenarios.md`:PASS / LOW_CONFID / REJECT / no-evidence 场景矩阵。
|
||||||
- `trace-inspection-checklist.md`:Trace 字段检查清单。
|
- `trace-inspection-checklist.md`:Trace 字段检查清单。
|
||||||
- `scripts/run-interview-demo-check.ps1`:面试预检脚本,包含服务可达性、Chat、Trace、反馈和 summary 输出。
|
- `scripts/run-interview-demo-check.ps1`:面试预检脚本,绑定 exact runId,强制校验 Run orchestration trace,并输出 Chat、Trace、反馈和 summary。
|
||||||
- `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。
|
- `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。
|
||||||
- `interview-q-and-a.md`:面试追问回答,覆盖 Agent 工程取舍、审计和评测。
|
- `interview-q-and-a.md`:面试追问回答,覆盖 Agent 工程取舍、审计和评测。
|
||||||
- `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。
|
- `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。
|
||||||
@@ -51,6 +51,8 @@ mvp/demo/output/feedback-response.json
|
|||||||
mvp/demo/output/interview-demo-summary.json
|
mvp/demo/output/interview-demo-summary.json
|
||||||
```
|
```
|
||||||
|
|
||||||
|
自动化验收应传入唯一 `-SessionId`,并用 `-OutputDir target/...` 避免覆盖仓库样例。脚本从 Chat 响应取得 exact `runId`,缺少 `data.run.orchestrationTrace` 或 version/final node/termination reason/transitions/degraded/evidence retry count 时会立即失败。summary 额外包含 `orchestrationVersion`、`finalNode`、`terminationReason`、`degraded`、`transitionCount` 和 `evidenceRetryCount`。
|
||||||
|
|
||||||
手动请求:
|
手动请求:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
@@ -100,6 +102,10 @@ Invoke-RestMethod `
|
|||||||
- `data.runId` 等于 `$runId`
|
- `data.runId` 等于 `$runId`
|
||||||
- `data.session.sessionId` 等于 Chat session id
|
- `data.session.sessionId` 等于 Chat session id
|
||||||
- `data.run.runId` 等于 `$runId`
|
- `data.run.runId` 等于 `$runId`
|
||||||
|
- `data.run.orchestrationTrace.version` 非空
|
||||||
|
- `data.run.orchestrationTrace.final_node` 和 `termination_reason` 非空
|
||||||
|
- `data.run.orchestrationTrace.transitions` 是本次 Graph 的条件边记录
|
||||||
|
- `data.run.orchestrationTrace.degraded` 和 `evidence_retry_count` 记录安全降级与补证据次数
|
||||||
- `data.steps` 包含 planner / executor / verifier 等步骤
|
- `data.steps` 包含 planner / executor / verifier 等步骤
|
||||||
- `data.toolInvocations` 包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具
|
- `data.toolInvocations` 包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具
|
||||||
- `data.session.selfEvaluation` 包含 verifier 或 rule evaluation
|
- `data.session.selfEvaluation` 包含 verifier 或 rule evaluation
|
||||||
@@ -174,9 +180,10 @@ Chat 主线:
|
|||||||
```text
|
```text
|
||||||
一个 session id + 一个 run id
|
一个 session id + 一个 run id
|
||||||
-> 用户问题
|
-> 用户问题
|
||||||
-> 多 Agent 执行
|
-> bounded StateGraph(Planner / Executor / Gatekeeper / Verified Input / Verifier / Composer / Fallback)
|
||||||
-> 证据工具
|
-> 证据工具
|
||||||
-> Verifier / self_evaluation
|
-> Verifier / self_evaluation
|
||||||
|
-> run.orchestrationTrace 路由摘要
|
||||||
-> 最终答案
|
-> 最终答案
|
||||||
-> 用户反馈
|
-> 用户反馈
|
||||||
-> Trace API 回放
|
-> Trace API 回放
|
||||||
|
|||||||
@@ -79,6 +79,15 @@ $chatPath = Join-Path $OutputDir "chat-response.json"
|
|||||||
$chat | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $chatPath
|
$chat | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $chatPath
|
||||||
|
|
||||||
$runId = $chat.data.runId
|
$runId = $chat.data.runId
|
||||||
|
if ($chat.data.success -ne $true) {
|
||||||
|
throw "Chat response was not successful."
|
||||||
|
}
|
||||||
|
if ([string]::IsNullOrWhiteSpace([string]$chat.data.answer)) {
|
||||||
|
throw "Chat response did not include a non-empty answer."
|
||||||
|
}
|
||||||
|
if ($chat.data.sessionId -ne $SessionId) {
|
||||||
|
throw "Chat response sessionId '$($chat.data.sessionId)' did not match requested sessionId '$SessionId'."
|
||||||
|
}
|
||||||
if (-not $runId) {
|
if (-not $runId) {
|
||||||
throw "Chat response did not include runId; exact trace verification cannot continue."
|
throw "Chat response did not include runId; exact trace verification cannot continue."
|
||||||
}
|
}
|
||||||
@@ -92,6 +101,51 @@ $trace = Invoke-RestMethod @traceRequest
|
|||||||
$tracePath = Join-Path $OutputDir "trace-response.json"
|
$tracePath = Join-Path $OutputDir "trace-response.json"
|
||||||
$trace | ConvertTo-Json -Depth 80 | Set-Content -Encoding UTF8 -Path $tracePath
|
$trace | ConvertTo-Json -Depth 80 | Set-Content -Encoding UTF8 -Path $tracePath
|
||||||
|
|
||||||
|
$traceData = Get-TraceData -TraceResponse $trace
|
||||||
|
if ($null -eq $traceData -or $null -eq $traceData.run) {
|
||||||
|
throw "Exact Trace response did not include data.run."
|
||||||
|
}
|
||||||
|
if ($traceData.runId -ne $runId -or $traceData.run.runId -ne $runId) {
|
||||||
|
throw "Exact Trace runId did not match Chat runId '$runId'."
|
||||||
|
}
|
||||||
|
if ($traceData.run.sessionId -ne $SessionId) {
|
||||||
|
throw "Exact Trace run did not belong to requested sessionId '$SessionId'."
|
||||||
|
}
|
||||||
|
|
||||||
|
$orchestrationTrace = $traceData.run.orchestrationTrace
|
||||||
|
if ($null -eq $orchestrationTrace) {
|
||||||
|
throw "Exact Trace data.run.orchestrationTrace is missing."
|
||||||
|
}
|
||||||
|
foreach ($field in @("version", "final_node", "termination_reason")) {
|
||||||
|
if (-not ($orchestrationTrace.PSObject.Properties.Name -contains $field) -or
|
||||||
|
[string]::IsNullOrWhiteSpace([string]$orchestrationTrace.$field)) {
|
||||||
|
throw "Exact Trace data.run.orchestrationTrace.$field is missing."
|
||||||
|
}
|
||||||
|
}
|
||||||
|
foreach ($field in @("transitions", "degraded", "evidence_retry_count")) {
|
||||||
|
if (-not ($orchestrationTrace.PSObject.Properties.Name -contains $field)) {
|
||||||
|
throw "Exact Trace data.run.orchestrationTrace.$field is missing."
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if ($null -eq $orchestrationTrace.transitions) {
|
||||||
|
throw "Exact Trace data.run.orchestrationTrace.transitions must be an array."
|
||||||
|
}
|
||||||
|
if ([int]$orchestrationTrace.evidence_retry_count -lt 0) {
|
||||||
|
throw "Exact Trace data.run.orchestrationTrace.evidence_retry_count must not be negative."
|
||||||
|
}
|
||||||
|
if ($traceData.run.status -ne "SUCCESS" -or $traceData.run.agentFlow -ne "CHAT") {
|
||||||
|
throw "Exact Trace run must be CHAT/SUCCESS."
|
||||||
|
}
|
||||||
|
if ([string]::IsNullOrWhiteSpace([string]$traceData.run.answer)) {
|
||||||
|
throw "Exact Trace run did not include a non-empty answer."
|
||||||
|
}
|
||||||
|
if (@($traceData.steps).Count -eq 0 -or @($traceData.toolInvocations).Count -eq 0) {
|
||||||
|
throw "Exact Trace did not include both Agent steps and tool invocation evidence."
|
||||||
|
}
|
||||||
|
if ($null -eq $traceData.run.selfEvaluation) {
|
||||||
|
throw "Exact Trace run did not include selfEvaluation."
|
||||||
|
}
|
||||||
|
|
||||||
$feedbackBody = @{
|
$feedbackBody = @{
|
||||||
sessionId = $SessionId
|
sessionId = $SessionId
|
||||||
runId = $runId
|
runId = $runId
|
||||||
@@ -105,11 +159,13 @@ $feedbackRequest = @{
|
|||||||
Body = $feedbackBody
|
Body = $feedbackBody
|
||||||
}
|
}
|
||||||
$feedback = Invoke-RestMethod @feedbackRequest
|
$feedback = Invoke-RestMethod @feedbackRequest
|
||||||
|
if ($feedback.success -ne $true) {
|
||||||
|
throw "Feedback request was not successful for runId '$runId'."
|
||||||
|
}
|
||||||
|
|
||||||
$feedbackPath = Join-Path $OutputDir "feedback-response.json"
|
$feedbackPath = Join-Path $OutputDir "feedback-response.json"
|
||||||
$feedback | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $feedbackPath
|
$feedback | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $feedbackPath
|
||||||
|
|
||||||
$traceData = Get-TraceData -TraceResponse $trace
|
|
||||||
$selfEvaluation = Get-SelfEvaluation -TraceData $traceData
|
$selfEvaluation = Get-SelfEvaluation -TraceData $traceData
|
||||||
$verifierEvaluation = $null
|
$verifierEvaluation = $null
|
||||||
if ($null -ne $selfEvaluation) {
|
if ($null -ne $selfEvaluation) {
|
||||||
@@ -138,6 +194,7 @@ if ($null -ne $promptAudit) {
|
|||||||
$promptAuditVersion = $promptAudit.version
|
$promptAuditVersion = $promptAudit.version
|
||||||
}
|
}
|
||||||
$toolNames = Get-ToolNames -TraceData $traceData
|
$toolNames = Get-ToolNames -TraceData $traceData
|
||||||
|
$transitionCount = @($orchestrationTrace.transitions).Count
|
||||||
$summaryPath = Join-Path $OutputDir "interview-demo-summary.json"
|
$summaryPath = Join-Path $OutputDir "interview-demo-summary.json"
|
||||||
|
|
||||||
$summary = [ordered]@{
|
$summary = [ordered]@{
|
||||||
@@ -149,6 +206,12 @@ $summary = [ordered]@{
|
|||||||
gatekeeperStatus = $gatekeeperStatus
|
gatekeeperStatus = $gatekeeperStatus
|
||||||
gatekeeperRuleSetVersion = $gatekeeperRuleSetVersion
|
gatekeeperRuleSetVersion = $gatekeeperRuleSetVersion
|
||||||
promptAuditVersion = $promptAuditVersion
|
promptAuditVersion = $promptAuditVersion
|
||||||
|
orchestrationVersion = $orchestrationTrace.version
|
||||||
|
finalNode = $orchestrationTrace.final_node
|
||||||
|
terminationReason = $orchestrationTrace.termination_reason
|
||||||
|
degraded = [bool]$orchestrationTrace.degraded
|
||||||
|
transitionCount = $transitionCount
|
||||||
|
evidenceRetryCount = [int]$orchestrationTrace.evidence_retry_count
|
||||||
toolNames = $toolNames
|
toolNames = $toolNames
|
||||||
paths = [ordered]@{
|
paths = [ordered]@{
|
||||||
chat = $chatPath
|
chat = $chatPath
|
||||||
@@ -165,4 +228,6 @@ Write-Host "Interview demo preflight completed."
|
|||||||
Write-Host "Verdict: $($summary.verdict)"
|
Write-Host "Verdict: $($summary.verdict)"
|
||||||
Write-Host "Gatekeeper rules: $($summary.gatekeeperRuleSetVersion)"
|
Write-Host "Gatekeeper rules: $($summary.gatekeeperRuleSetVersion)"
|
||||||
Write-Host "Prompt audit: $($summary.promptAuditVersion)"
|
Write-Host "Prompt audit: $($summary.promptAuditVersion)"
|
||||||
|
Write-Host "Graph final node: $($summary.finalNode)"
|
||||||
|
Write-Host "Graph termination: $($summary.terminationReason)"
|
||||||
Write-Host "Summary: $summaryPath"
|
Write-Host "Summary: $summaryPath"
|
||||||
|
|||||||
@@ -7,16 +7,30 @@
|
|||||||
| JSON path | 检查点 | 面试讲点 |
|
| JSON path | 检查点 | 面试讲点 |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| `data.runId` / `data.run.runId` | 是否等于 demo 响应中的 `runId` | `runId` 精确绑定这一次诊断运行 |
|
| `data.runId` / `data.run.runId` | 是否等于 demo 响应中的 `runId` | `runId` 精确绑定这一次诊断运行 |
|
||||||
| `data.session.sessionId` | 是否等于 `mvp-demo-payment-timeout-001` | `sessionId` 保留多轮上下文,Trace 精确回放依赖 `runId` |
|
| `data.run.sessionId` | 是否等于本次 Chat 请求的唯一 sessionId | `sessionId` 保留多轮上下文,Trace 精确回放依赖 `runId` |
|
||||||
| `data.session.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
|
| `data.run.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
|
||||||
| `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
|
| `data.run.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
|
||||||
| `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
|
| `data.run.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
|
||||||
| `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | 如果是 Chat V2 链路,是否记录 Gatekeeper 规则版本 | 安全规则可审计、可回归 |
|
| `data.run.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | 是否记录 Gatekeeper 规则版本 | 安全规则可审计、可回归 |
|
||||||
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` | 如果是 Chat V2 链路,是否记录 Prompt 审计版本 | Prompt 变更可解释、可回归 |
|
| `data.run.selfEvaluation.verifier_evaluation.prompt_audit.version` | 是否记录 Prompt 审计版本 | Prompt 变更可解释、可回归 |
|
||||||
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.prompts[*].version` | 是否记录 planner / executor / verifier / composer 版本 | 便于定位 Prompt 变更影响 |
|
| `data.run.selfEvaluation.verifier_evaluation.prompt_audit.prompts[*].version` | 是否记录 planner / executor / verifier / composer 版本 | 便于定位 Prompt 变更影响 |
|
||||||
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在当前 diagnosis run 上 |
|
| `data.run.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在当前 diagnosis run 上 |
|
||||||
|
|
||||||
## 2. Agent 步骤
|
## 2. StateGraph 路由
|
||||||
|
|
||||||
|
| JSON path | 检查点 | 面试讲点 |
|
||||||
|
|---|---|---|
|
||||||
|
| `data.run.orchestrationTrace.version` | 是否存在当前 trace contract 版本 | 路由摘要可演进、可兼容 |
|
||||||
|
| `data.run.orchestrationTrace.transitions[*]` | 是否记录实际经过的 Node 和 route | Graph 条件边不是从日志推断 |
|
||||||
|
| `data.run.orchestrationTrace.final_node` | 最终是 Composer 还是 Fallback | 正常输出与安全降级明确区分 |
|
||||||
|
| `data.run.orchestrationTrace.termination_reason` | 是否给出终止原因 | 每次 Run 都有可解释终点 |
|
||||||
|
| `data.run.orchestrationTrace.degraded` | 是否发生安全降级 | fallback 是可审计行为 |
|
||||||
|
| `data.run.orchestrationTrace.evidence_retry_count` | 是否为 0 或 1 | 补证据循环有硬上限 |
|
||||||
|
| `interview-demo-summary.json.finalNode` 等摘要字段 | 是否与 exact Trace 一致 | summary 只消费 Run 路由真理源 |
|
||||||
|
|
||||||
|
`orchestrationTrace` 负责路由;`selfEvaluation` 负责证据和答案质量;AgentStep/ToolInvocation 负责详细执行与工具证据。三者不能互相替代。
|
||||||
|
|
||||||
|
## 3. Agent 步骤
|
||||||
|
|
||||||
| JSON path | 检查点 | 面试讲点 |
|
| JSON path | 检查点 | 面试讲点 |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
@@ -25,7 +39,7 @@
|
|||||||
| `data.steps[*].durationMs` | 是否有步骤耗时 | Trace 可用于耗时分析 |
|
| `data.steps[*].durationMs` | 是否有步骤耗时 | Trace 可用于耗时分析 |
|
||||||
| `data.steps[*].tokenCount` | 如可用,是否记录 token | Trace 可用于模型成本分析 |
|
| `data.steps[*].tokenCount` | 如可用,是否记录 token | Trace 可用于模型成本分析 |
|
||||||
|
|
||||||
## 3. 工具证据
|
## 4. 工具证据
|
||||||
|
|
||||||
| JSON path | 检查点 | 面试讲点 |
|
| JSON path | 检查点 | 面试讲点 |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
@@ -37,7 +51,7 @@
|
|||||||
| `data.toolInvocations[*].retrievalDetails.evidence_refs` | 是否包含 `raw_path + text` | Gatekeeper 可以用代码核对 Executor 引用 |
|
| `data.toolInvocations[*].retrievalDetails.evidence_refs` | 是否包含 `raw_path + text` | Gatekeeper 可以用代码核对 Executor 引用 |
|
||||||
| `data.toolInvocations[*].relevanceLevel` | 是否有相关性等级 | 可解释检索结果强弱 |
|
| `data.toolInvocations[*].relevanceLevel` | 是否有相关性等级 | 可解释检索结果强弱 |
|
||||||
|
|
||||||
## 4. Summary
|
## 5. Summary
|
||||||
|
|
||||||
| JSON path | 检查点 | 面试讲点 |
|
| JSON path | 检查点 | 面试讲点 |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
@@ -46,7 +60,7 @@
|
|||||||
| `data.summary.hasVerifierEvaluation` | 是否存在 Verifier 结果 | 最终答案经过质量门 |
|
| `data.summary.hasVerifierEvaluation` | 是否存在 Verifier 结果 | 最终答案经过质量门 |
|
||||||
| `data.summary.hasFeedback` | 提交反馈后是否为 true | 人类反馈闭环完成 |
|
| `data.summary.hasFeedback` | 提交反馈后是否为 true | 人类反馈闭环完成 |
|
||||||
|
|
||||||
## 5. 好的结果长什么样
|
## 6. 好的结果长什么样
|
||||||
|
|
||||||
```text
|
```text
|
||||||
同一个 session id + run id
|
同一个 session id + run id
|
||||||
@@ -54,5 +68,6 @@
|
|||||||
-> 持久化 agent steps
|
-> 持久化 agent steps
|
||||||
-> 持久化 evidence tool calls
|
-> 持久化 evidence tool calls
|
||||||
-> verifier / self-evaluation
|
-> verifier / self-evaluation
|
||||||
|
-> run.orchestrationTrace routing summary
|
||||||
-> feedback attached to the same run
|
-> feedback attached to the same run
|
||||||
```
|
```
|
||||||
|
|||||||
+7
-5
@@ -4,13 +4,15 @@ This folder contains the fixed offline regression set for the MVP diagnosis Agen
|
|||||||
|
|
||||||
## Background
|
## Background
|
||||||
|
|
||||||
The current diagnosis chain is:
|
The current complex Chat diagnosis chain is a bounded StateGraph:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
Planner -> Executor -> Gatekeeper -> Verifier -> Composer -> final answer
|
Planner -> Executor -> Gatekeeper -> Verified Input -> Verifier -> Composer -> final answer
|
||||||
|
| |
|
||||||
|
+ bounded evidence retry + safe Fallback
|
||||||
```
|
```
|
||||||
|
|
||||||
Stages 1-4 introduced Executor V2 structured output, deterministic Gatekeeper audit, Verifier `claim_checks`, and Composer final-answer rendering. Stage 5 makes those audit fields part of the offline regression harness so future prompt, tool, or chain changes can be checked without relying on a one-off demo.
|
Executor V2 structured output, deterministic Gatekeeper audit, verified-only Verifier input, `claim_checks`, Composer rendering, and StateGraph routing are covered by deterministic tests so future prompt, tool, or graph changes can be checked without relying on a one-off demo.
|
||||||
|
|
||||||
## Scope
|
## Scope
|
||||||
|
|
||||||
@@ -55,10 +57,10 @@ Run the focused evaluator test:
|
|||||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test
|
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test
|
||||||
```
|
```
|
||||||
|
|
||||||
Run the broader phase-5 regression set:
|
Run the authoritative Graph layers plus the fixed evaluator checks:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
mvn "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
|
mvn -q "-Dtest=DiagnosisGraphWorkflowTest,DiagnosisGraphNodeContractTest,ChatServiceGraphIntegrationTest,DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest" test
|
||||||
```
|
```
|
||||||
|
|
||||||
When fixtures or evaluator rules change, regenerate both baseline reports from the same case file and fixture directory, then update JSON and Markdown together.
|
When fixtures or evaluator rules change, regenerate both baseline reports from the same case file and fixture directory, then update JSON and Markdown together.
|
||||||
|
|||||||
@@ -1,5 +1,5 @@
|
|||||||
{
|
{
|
||||||
"session": {
|
"run": {
|
||||||
"sessionId": "eval-audit-metadata-low-confid",
|
"sessionId": "eval-audit-metadata-low-confid",
|
||||||
"query": "订单超时是否可以确认由数据库主库故障导致,并检查审计元数据是否完整?",
|
"query": "订单超时是否可以确认由数据库主库故障导致,并检查审计元数据是否完整?",
|
||||||
"status": "SUCCESS",
|
"status": "SUCCESS",
|
||||||
|
|||||||
@@ -1,5 +1,5 @@
|
|||||||
{
|
{
|
||||||
"session": {
|
"run": {
|
||||||
"sessionId": "eval-composer-fallback-no-raw-json",
|
"sessionId": "eval-composer-fallback-no-raw-json",
|
||||||
"query": "库存服务慢响应是否可以直接输出 Executor JSON?",
|
"query": "库存服务慢响应是否可以直接输出 Executor JSON?",
|
||||||
"status": "SUCCESS",
|
"status": "SUCCESS",
|
||||||
|
|||||||
@@ -1,5 +1,5 @@
|
|||||||
{
|
{
|
||||||
"session": {
|
"run": {
|
||||||
"sessionId": "eval-gatekeeper-fabricated-invocation",
|
"sessionId": "eval-gatekeeper-fabricated-invocation",
|
||||||
"query": "支付失败是否能确认由日志中的连接池耗尽导致?",
|
"query": "支付失败是否能确认由日志中的连接池耗尽导致?",
|
||||||
"status": "SUCCESS",
|
"status": "SUCCESS",
|
||||||
|
|||||||
@@ -1,5 +1,5 @@
|
|||||||
{
|
{
|
||||||
"session": {
|
"run": {
|
||||||
"sessionId": "eval-hikari-no-evidence-negative-observation",
|
"sessionId": "eval-hikari-no-evidence-negative-observation",
|
||||||
"query": "确认 inventory-service 当前是否有 HikariCP 连接池耗尽日志。",
|
"query": "确认 inventory-service 当前是否有 HikariCP 连接池耗尽日志。",
|
||||||
"status": "SUCCESS",
|
"status": "SUCCESS",
|
||||||
|
|||||||
@@ -1,5 +1,5 @@
|
|||||||
{
|
{
|
||||||
"session": {
|
"run": {
|
||||||
"sessionId": "eval-jvm-memory-risk",
|
"sessionId": "eval-jvm-memory-risk",
|
||||||
"query": "订单服务内存使用率过高,请判断是否存在 OOM 风险。",
|
"query": "订单服务内存使用率过高,请判断是否存在 OOM 风险。",
|
||||||
"status": "SUCCESS",
|
"status": "SUCCESS",
|
||||||
|
|||||||
@@ -1,5 +1,5 @@
|
|||||||
{
|
{
|
||||||
"session": {
|
"run": {
|
||||||
"sessionId": "eval-mysql-pool",
|
"sessionId": "eval-mysql-pool",
|
||||||
"query": "订单服务大量请求超时,请判断是否和 MySQL 连接池有关。",
|
"query": "订单服务大量请求超时,请判断是否和 MySQL 连接池有关。",
|
||||||
"status": "SUCCESS",
|
"status": "SUCCESS",
|
||||||
|
|||||||
@@ -1,5 +1,5 @@
|
|||||||
{
|
{
|
||||||
"session": {
|
"run": {
|
||||||
"sessionId": "eval-narrow-highcpu-observation",
|
"sessionId": "eval-narrow-highcpu-observation",
|
||||||
"query": "确认 payment-service 当前是否存在 HighCPUUsage 告警。",
|
"query": "确认 payment-service 当前是否存在 HighCPUUsage 告警。",
|
||||||
"status": "SUCCESS",
|
"status": "SUCCESS",
|
||||||
|
|||||||
@@ -1,5 +1,5 @@
|
|||||||
{
|
{
|
||||||
"session": {
|
"run": {
|
||||||
"sessionId": "eval-payment-timeout",
|
"sessionId": "eval-payment-timeout",
|
||||||
"query": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
|
"query": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
|
||||||
"status": "SUCCESS",
|
"status": "SUCCESS",
|
||||||
|
|||||||
@@ -1,5 +1,5 @@
|
|||||||
{
|
{
|
||||||
"session": {
|
"run": {
|
||||||
"sessionId": "eval-prompt-gatekeeper-audit-closure",
|
"sessionId": "eval-prompt-gatekeeper-audit-closure",
|
||||||
"query": "确认 payment-service 当前是否存在 HighCPUUsage 告警,并检查审计元数据是否完整。",
|
"query": "确认 payment-service 当前是否存在 HighCPUUsage 告警,并检查审计元数据是否完整。",
|
||||||
"status": "SUCCESS",
|
"status": "SUCCESS",
|
||||||
|
|||||||
@@ -1,5 +1,5 @@
|
|||||||
{
|
{
|
||||||
"session": {
|
"run": {
|
||||||
"sessionId": "eval-redis-timeout",
|
"sessionId": "eval-redis-timeout",
|
||||||
"query": "支付服务出现 Redis 连接超时,请定位可能原因。",
|
"query": "支付服务出现 Redis 连接超时,请定位可能原因。",
|
||||||
"status": "SUCCESS",
|
"status": "SUCCESS",
|
||||||
|
|||||||
@@ -1,5 +1,5 @@
|
|||||||
{
|
{
|
||||||
"session": {
|
"run": {
|
||||||
"sessionId": "eval-slow-response",
|
"sessionId": "eval-slow-response",
|
||||||
"query": "用户服务 P99 响应时间升高,请结合指标和日志分析。",
|
"query": "用户服务 P99 响应时间升高,请结合指标和日志分析。",
|
||||||
"status": "SUCCESS",
|
"status": "SUCCESS",
|
||||||
|
|||||||
@@ -1,5 +1,5 @@
|
|||||||
{
|
{
|
||||||
"session": {
|
"run": {
|
||||||
"sessionId": "eval-unsupported-claim-filtering",
|
"sessionId": "eval-unsupported-claim-filtering",
|
||||||
"query": "订单超时是否可以确认由数据库主库故障导致?",
|
"query": "订单超时是否可以确认由数据库主库故障导致?",
|
||||||
"status": "SUCCESS",
|
"status": "SUCCESS",
|
||||||
|
|||||||
+5
-10
@@ -1,14 +1,13 @@
|
|||||||
# MVP Issues 索引
|
# MVP Issues 索引
|
||||||
|
|
||||||
**更新日期**:2026-07-10
|
**更新日期**:2026-07-20
|
||||||
**状态**:按活跃问题、设计笔记、RAG 问题集和已归档问题整理
|
**状态**:按活跃问题、RAG 问题集和已归档问题整理
|
||||||
|
|
||||||
## 目录约定
|
## 目录约定
|
||||||
|
|
||||||
| 目录 | 用途 |
|
| 目录 | 用途 |
|
||||||
|---|---|
|
|---|---|
|
||||||
| [active/](active/) | 仍需要规划或实现的问题 |
|
| [active/](active/) | 仍需要规划或实现的问题 |
|
||||||
| [design-notes/](design-notes/) | 已形成方向、用于指导后续实现的设计记录 |
|
|
||||||
| [rag/](rag/) | RAG 子问题集合;多数已合并到 RAG 重构计划 |
|
| [rag/](rag/) | RAG 子问题集合;多数已合并到 RAG 重构计划 |
|
||||||
| [archived/](archived/) | 已修复、已实施或已归档的问题 |
|
| [archived/](archived/) | 已修复、已实施或已归档的问题 |
|
||||||
|
|
||||||
@@ -21,17 +20,12 @@
|
|||||||
| executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [active/executor-evidence-attribution-hallucination.md](active/executor-evidence-attribution-hallucination.md) |
|
| executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [active/executor-evidence-attribution-hallucination.md](active/executor-evidence-attribution-hallucination.md) |
|
||||||
| rag-refactor-plan | RAG 检索重构计划 | 高 | 待规划 | [active/rag-refactor-plan.md](active/rag-refactor-plan.md) |
|
| rag-refactor-plan | RAG 检索重构计划 | 高 | 待规划 | [active/rag-refactor-plan.md](active/rag-refactor-plan.md) |
|
||||||
|
|
||||||
## 设计笔记
|
|
||||||
|
|
||||||
| 名称 | 标题 | 状态 | 文件 |
|
|
||||||
|---|---|---|---|
|
|
||||||
| executor-self-evidence-loop-design-note | Executor 自证循环与证据摘要链路设计记录 | 已形成方向 | [design-notes/executor-self-evidence-loop-design-note.md](design-notes/executor-self-evidence-loop-design-note.md) |
|
|
||||||
| executor-structured-output-v2 | Executor 结构化输出 V2 阶段设计 | 部分已实施,保留为后续改造参考 | [design-notes/executor-structured-output-v2.md](design-notes/executor-structured-output-v2.md) |
|
|
||||||
|
|
||||||
## RAG 问题集
|
## RAG 问题集
|
||||||
|
|
||||||
这些问题已经收敛到 [active/rag-refactor-plan.md](active/rag-refactor-plan.md),单个文件保留用于追溯原始问题和设计背景。
|
这些问题已经收敛到 [active/rag-refactor-plan.md](active/rag-refactor-plan.md),单个文件保留用于追溯原始问题和设计背景。
|
||||||
|
|
||||||
|
已被当前 Executor Evidence Pipeline 和 OpenSpec 替代的早期设计笔记已移至 [2026-07-20 文档清理归档](../archive/2026-07-20-doc-cleanup/README.md)。
|
||||||
|
|
||||||
| 名称 | 标题 | 状态 | 文件 |
|
| 名称 | 标题 | 状态 | 文件 |
|
||||||
|---|---|---|---|
|
|---|---|---|---|
|
||||||
| breadcrumb-embedding-gap | RAG breadcrumb 未参与向量语义 | 已合并到重构计划 | [rag/rag-breadcrumb-embedding-gap.md](rag/rag-breadcrumb-embedding-gap.md) |
|
| breadcrumb-embedding-gap | RAG breadcrumb 未参与向量语义 | 已合并到重构计划 | [rag/rag-breadcrumb-embedding-gap.md](rag/rag-breadcrumb-embedding-gap.md) |
|
||||||
@@ -52,6 +46,7 @@
|
|||||||
|
|
||||||
| 名称 | 标题 | 状态 | 文件 |
|
| 名称 | 标题 | 状态 | 文件 |
|
||||||
|---|---|---|---|
|
|---|---|---|---|
|
||||||
|
| ISS-011 | Chat 诊断 StateGraph 编排改造 | 已归档 | [archived/ISS-011-chat-diagnosis-stategraph-orchestration.md](archived/ISS-011-chat-diagnosis-stategraph-orchestration.md) |
|
||||||
| ISS-001 | Executor 重复召回同一文档 | 已修复 | [archived/ISS-001-duplicate-retrieval.md](archived/ISS-001-duplicate-retrieval.md) |
|
| ISS-001 | Executor 重复召回同一文档 | 已修复 | [archived/ISS-001-duplicate-retrieval.md](archived/ISS-001-duplicate-retrieval.md) |
|
||||||
| ISS-002 | Executor 无约束重复调用 lookup_knowledge | 已修复 | [archived/ISS-002-executor-unconstrained-lookup.md](archived/ISS-002-executor-unconstrained-lookup.md) |
|
| ISS-002 | Executor 无约束重复调用 lookup_knowledge | 已修复 | [archived/ISS-002-executor-unconstrained-lookup.md](archived/ISS-002-executor-unconstrained-lookup.md) |
|
||||||
| ISS-005 | 证据链补齐与降级契约收敛 | 已归档 | [archived/ISS-005-evidence-trace-hardening.md](archived/ISS-005-evidence-trace-hardening.md) |
|
| ISS-005 | 证据链补齐与降级契约收敛 | 已归档 | [archived/ISS-005-evidence-trace-hardening.md](archived/ISS-005-evidence-trace-hardening.md) |
|
||||||
|
|||||||
@@ -569,7 +569,7 @@ tool_invocation(run_id, id)
|
|||||||
- `mvp/architecture/data-model.md`
|
- `mvp/architecture/data-model.md`
|
||||||
- `mvp/tables/聊天会话表-chat_session.md`
|
- `mvp/tables/聊天会话表-chat_session.md`
|
||||||
- `mvp/tables/诊断运行表-diagnosis_run.md`
|
- `mvp/tables/诊断运行表-diagnosis_run.md`
|
||||||
- `mvp/tables/诊断会话表-diagnosis_session.md`
|
- `mvp/archive/2026-07-20-doc-cleanup/tables/诊断会话表-diagnosis_session.md`
|
||||||
- `mvp/tables/Agent步骤表-agent_step.md`
|
- `mvp/tables/Agent步骤表-agent_step.md`
|
||||||
- `mvp/tables/工具调用表-tool_invocation.md`
|
- `mvp/tables/工具调用表-tool_invocation.md`
|
||||||
- `mvp/tables/案例库表-case_library.md`
|
- `mvp/tables/案例库表-case_library.md`
|
||||||
|
|||||||
File diff suppressed because it is too large
Load Diff
@@ -12,7 +12,7 @@
|
|||||||
| 字段 | 类型 | 必填 | 说明 |
|
| 字段 | 类型 | 必填 | 说明 |
|
||||||
|---|---|---|---|
|
|---|---|---|---|
|
||||||
| `id` | BIGINT | 是 | 自增主键 |
|
| `id` | BIGINT | 是 | 自增主键 |
|
||||||
| `session_id` | VARCHAR(64) | 是 | 所属会话目录 ID,保留用于粗粒度过滤和兼容 |
|
| `session_id` | VARCHAR(64) | 是 | 所属会话目录 ID,作为粗粒度冗余筛选字段 |
|
||||||
| `run_id` | VARCHAR(64) | 否 | 所属 `diagnosis_run.run_id`;新执行应写入 |
|
| `run_id` | VARCHAR(64) | 否 | 所属 `diagnosis_run.run_id`;新执行应写入 |
|
||||||
| `step_index` | INT | 是 | 步骤序号,从 0 开始 |
|
| `step_index` | INT | 是 | 步骤序号,从 0 开始 |
|
||||||
| `agent_name` | VARCHAR(32) | 是 | Agent 名称,例如 planner、executor、verifier、composer |
|
| `agent_name` | VARCHAR(32) | 是 | Agent 名称,例如 planner、executor、verifier、composer |
|
||||||
@@ -28,14 +28,14 @@
|
|||||||
|
|
||||||
| 索引 | 字段 | 用途 |
|
| 索引 | 字段 | 用途 |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| `idx_session_step` | `session_id, step_index` | 历史兼容和粗粒度排查 |
|
| `idx_session_step` | `session_id, step_index` | 粗粒度排查 |
|
||||||
| `idx_agent_step_run_step` | `run_id, step_index` | 按运行筛选步骤并辅助顺序查询 |
|
| `idx_agent_step_run_step` | `run_id, step_index` | 按运行筛选步骤并辅助顺序查询 |
|
||||||
| `idx_agent_name` | `agent_name` | 按 Agent 类型筛选 |
|
| `idx_agent_name` | `agent_name` | 按 Agent 类型筛选 |
|
||||||
|
|
||||||
## 关系
|
## 关系
|
||||||
|
|
||||||
- `agent_step.run_id` 逻辑关联 `diagnosis_run.run_id`。
|
- `agent_step.run_id` 逻辑关联 `diagnosis_run.run_id`。
|
||||||
- `agent_step.session_id` 保留为 `chat_session.session_id` 的冗余关联,便于粗粒度过滤和兼容查询。
|
- `agent_step.session_id` 是 `chat_session.session_id` 的冗余关联,只用于粗粒度过滤。
|
||||||
- `tool_invocation.step_id` 可关联 `agent_step.id`,但当前允许为空且不强制外键。
|
- `tool_invocation.step_id` 可关联 `agent_step.id`,但当前允许为空且不强制外键。
|
||||||
|
|
||||||
## 注意点
|
## 注意点
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
# MVP 数据表索引
|
# MVP 数据表索引
|
||||||
|
|
||||||
**更新日期**:2026-07-10
|
**更新日期**:2026-07-20
|
||||||
**状态**:当前表文档入口
|
**状态**:当前表文档入口
|
||||||
|
|
||||||
本目录保存当前 MVP 使用的数据表说明。详细结构以 Flyway migration 和实体类为准;本目录用于面试讲解、排查索引和快速理解数据流。
|
本目录保存当前 MVP 使用的数据表说明。详细结构以 Flyway migration 和实体类为准;本目录用于面试讲解、排查索引和快速理解数据流。
|
||||||
@@ -16,13 +16,13 @@
|
|||||||
| `api_document` | 知识库文档元数据,和向量库 chunk 通过 `doc_id` 关联 | [文档元数据表-api_document.md](文档元数据表-api_document.md) |
|
| `api_document` | 知识库文档元数据,和向量库 chunk 通过 `doc_id` 关联 | [文档元数据表-api_document.md](文档元数据表-api_document.md) |
|
||||||
| `knowledge_domain` | 知识域元数据,支撑 RAG domain hint 和检索策略 | [知识域表-knowledge_domain.md](知识域表-knowledge_domain.md) |
|
| `knowledge_domain` | 知识域元数据,支撑 RAG domain hint 和检索策略 | [知识域表-knowledge_domain.md](知识域表-knowledge_domain.md) |
|
||||||
| `case_library` | 用户反馈沉淀出的高质量诊断案例 | [案例库表-case_library.md](案例库表-case_library.md) |
|
| `case_library` | 用户反馈沉淀出的高质量诊断案例 | [案例库表-case_library.md](案例库表-case_library.md) |
|
||||||
| `diagnosis_session` | 历史兼容和回滚表,新执行写入不再依赖它 | [诊断会话表-diagnosis_session.md](诊断会话表-diagnosis_session.md) |
|
|
||||||
|
|
||||||
## 已归档表
|
## 已归档表
|
||||||
|
|
||||||
| 表 | 归档原因 | 文档 |
|
| 表 | 归档原因 | 文档 |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| `diagnosis_record` | 已由 `V007` 删除,历史上被 `diagnosis_session + agent_step + tool_invocation` 替代;当前新模型是 `chat_session + diagnosis_run + trace detail` | [archive/2026-07-09-doc-cleanup/旧诊断记录表-diagnosis_record.md](archive/2026-07-09-doc-cleanup/旧诊断记录表-diagnosis_record.md) |
|
| `diagnosis_record` | 已由 `V007` 删除,历史上被 `diagnosis_session + agent_step + tool_invocation` 替代;当前新模型是 `chat_session + diagnosis_run + trace detail` | [archive/2026-07-09-doc-cleanup/旧诊断记录表-diagnosis_record.md](archive/2026-07-09-doc-cleanup/旧诊断记录表-diagnosis_record.md) |
|
||||||
|
| `diagnosis_session` | 旧会话级诊断模型,当前 Java 运行时不再映射、读取或写入 | [../archive/2026-07-20-doc-cleanup/tables/诊断会话表-diagnosis_session.md](../archive/2026-07-20-doc-cleanup/tables/诊断会话表-diagnosis_session.md) |
|
||||||
|
|
||||||
## 核心关系
|
## 核心关系
|
||||||
|
|
||||||
@@ -33,9 +33,6 @@ chat_session.session_id
|
|||||||
-> tool_invocation.run_id
|
-> tool_invocation.run_id
|
||||||
-> case_library.diagnosis_id (new AUTO cases use run_id)
|
-> case_library.diagnosis_id (new AUTO cases use run_id)
|
||||||
|
|
||||||
diagnosis_session.session_id
|
|
||||||
-> historical compatibility / rollback only
|
|
||||||
|
|
||||||
api_document.doc_id
|
api_document.doc_id
|
||||||
-> vector chunk metadata.docId / doc_id
|
-> vector chunk metadata.docId / doc_id
|
||||||
|
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
# diagnosis_record - 旧诊断记录表
|
# diagnosis_record - 旧诊断记录表
|
||||||
|
|
||||||
> 归档说明:`diagnosis_record` 已在 `V007__drop_diagnosis_record.sql` 中删除,当前主模型是 `diagnosis_session + agent_step + tool_invocation`。本文只用于追溯早期设计。
|
> 归档说明:`diagnosis_record` 已在 `V007__drop_diagnosis_record.sql` 中删除,该阶段由 `diagnosis_session + agent_step + tool_invocation` 替代。本文只用于追溯早期设计,不代表当前 Run-only v2 模型。
|
||||||
|
|
||||||
## 表定位
|
## 表定位
|
||||||
|
|
||||||
|
|||||||
@@ -12,7 +12,7 @@
|
|||||||
| 字段 | 类型 | 必填 | 说明 |
|
| 字段 | 类型 | 必填 | 说明 |
|
||||||
|---|---|---|---|
|
|---|---|---|---|
|
||||||
| `id` | BIGINT | 是 | 自增主键 |
|
| `id` | BIGINT | 是 | 自增主键 |
|
||||||
| `session_id` | VARCHAR(64) | 是 | 所属会话目录 ID,保留用于粗粒度过滤和兼容 |
|
| `session_id` | VARCHAR(64) | 是 | 所属会话目录 ID,作为粗粒度冗余筛选字段 |
|
||||||
| `run_id` | VARCHAR(64) | 否 | 所属 `diagnosis_run.run_id`;新执行应写入 |
|
| `run_id` | VARCHAR(64) | 否 | 所属 `diagnosis_run.run_id`;新执行应写入 |
|
||||||
| `step_id` | BIGINT | 否 | 可关联 `agent_step.id` |
|
| `step_id` | BIGINT | 否 | 可关联 `agent_step.id` |
|
||||||
| `tool_name` | VARCHAR(64) | 是 | 工具名称,例如 `lookup_knowledge`、日志查询、指标查询 |
|
| `tool_name` | VARCHAR(64) | 是 | 工具名称,例如 `lookup_knowledge`、日志查询、指标查询 |
|
||||||
@@ -35,7 +35,7 @@
|
|||||||
|
|
||||||
| 索引 | 字段 | 用途 |
|
| 索引 | 字段 | 用途 |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| `idx_session_id` | `session_id` | 历史兼容和粗粒度排查 |
|
| `idx_session_id` | `session_id` | 粗粒度排查 |
|
||||||
| `idx_tool_invocation_run_id` | `run_id, id` | Trace、Verifier、评测按运行查询工具调用 |
|
| `idx_tool_invocation_run_id` | `run_id, id` | Trace、Verifier、评测按运行查询工具调用 |
|
||||||
| `idx_tool_name` | `tool_name` | 按工具类型排查 |
|
| `idx_tool_name` | `tool_name` | 按工具类型排查 |
|
||||||
| `idx_retrieval_layer` | `retrieval_layer` | 观察 RAG L0/L1 行为 |
|
| `idx_retrieval_layer` | `retrieval_layer` | 观察 RAG L0/L1 行为 |
|
||||||
@@ -43,7 +43,7 @@
|
|||||||
## 关系
|
## 关系
|
||||||
|
|
||||||
- `tool_invocation.run_id` 逻辑关联 `diagnosis_run.run_id`。
|
- `tool_invocation.run_id` 逻辑关联 `diagnosis_run.run_id`。
|
||||||
- `tool_invocation.session_id` 保留为 `chat_session.session_id` 的冗余关联,便于粗粒度过滤和兼容查询。
|
- `tool_invocation.session_id` 是 `chat_session.session_id` 的冗余关联,只用于粗粒度过滤。
|
||||||
- `tool_invocation.step_id` 可关联 `agent_step.id`,但当前不强制。
|
- `tool_invocation.step_id` 可关联 `agent_step.id`,但当前不强制。
|
||||||
|
|
||||||
## 关键 JSON
|
## 关键 JSON
|
||||||
|
|||||||
@@ -13,7 +13,7 @@
|
|||||||
|---|---|---|---|
|
|---|---|---|---|
|
||||||
| `id` | BIGINT | 是 | 自增主键 |
|
| `id` | BIGINT | 是 | 自增主键 |
|
||||||
| `case_id` | VARCHAR(64) | 是 | 案例唯一 ID |
|
| `case_id` | VARCHAR(64) | 是 | 案例唯一 ID |
|
||||||
| `diagnosis_id` | VARCHAR(64) | 否 | 关联诊断来源;新自动生成时存 `diagnosis_run.run_id`,历史数据可能是 `diagnosis_session.session_id` |
|
| `diagnosis_id` | VARCHAR(64) | 否 | 关联诊断来源;自动生成时存 `diagnosis_run.run_id` |
|
||||||
| `source_type` | VARCHAR(16) | 否 | 来源类型:`AUTO` 或 `MANUAL` |
|
| `source_type` | VARCHAR(16) | 否 | 来源类型:`AUTO` 或 `MANUAL` |
|
||||||
| `fault_category` | VARCHAR(32) | 否 | 故障类别,实体侧使用 `FaultCategory` |
|
| `fault_category` | VARCHAR(32) | 否 | 故障类别,实体侧使用 `FaultCategory` |
|
||||||
| `fault_source` | VARCHAR(128) | 否 | 故障源,例如服务、系统或省份 |
|
| `fault_source` | VARCHAR(128) | 否 | 故障源,例如服务、系统或省份 |
|
||||||
@@ -35,17 +35,17 @@
|
|||||||
| `idx_error_code` | `error_code` | 按错误码精确匹配 |
|
| `idx_error_code` | `error_code` | 按错误码精确匹配 |
|
||||||
| `idx_fault_source` | `fault_source` | 按故障源筛选 |
|
| `idx_fault_source` | `fault_source` | 按故障源筛选 |
|
||||||
| `idx_fault_target` | `fault_target(100)` | 按故障目标筛选 |
|
| `idx_fault_target` | `fault_target(100)` | 按故障目标筛选 |
|
||||||
| `idx_diagnosis_id` | `diagnosis_id` | 追溯来源运行或历史会话 |
|
| `idx_diagnosis_id` | `diagnosis_id` | 追溯来源运行 |
|
||||||
| `idx_reference_count` | `reference_count` | 推荐排序 |
|
| `idx_reference_count` | `reference_count` | 推荐排序 |
|
||||||
| `idx_created_at` | `created_at` | 时间排序 |
|
| `idx_created_at` | `created_at` | 时间排序 |
|
||||||
|
|
||||||
## 关系
|
## 关系
|
||||||
|
|
||||||
- `case_library.diagnosis_id` 是过渡字段:新自动案例逻辑关联 `diagnosis_run.run_id`,历史自动案例可能仍是 `diagnosis_session.session_id`。
|
- `case_library.diagnosis_id` 逻辑关联 `diagnosis_run.run_id`。
|
||||||
- 人工录入案例可以不填写 `diagnosis_id`。
|
- 人工录入案例可以不填写 `diagnosis_id`。
|
||||||
|
|
||||||
## 注意点
|
## 注意点
|
||||||
|
|
||||||
- 旧文档里提到的 `diagnosis_record` 已被 `V007` 删除,不再是当前主模型。
|
- 旧文档里提到的 `diagnosis_record` 已被 `V007` 删除,不再是当前主模型。
|
||||||
- 查询新自动案例时优先按 `run_id` 追溯;遇到旧值时再按历史 `session_id` 解释。
|
- 自动案例只按 `run_id` 追溯,不执行会话级兼容查询。
|
||||||
- 当前自动沉淀仍比较粗:`root_cause` 和 `solution` 都可能来自完整 answer。后续可从结构化结论中拆分根因、证据和修复建议。
|
- 当前自动沉淀仍比较粗:`root_cause` 和 `solution` 都可能来自完整 answer。后续可从结构化结论中拆分根因、证据和修复建议。
|
||||||
|
|||||||
@@ -5,7 +5,7 @@
|
|||||||
|
|
||||||
## 定位
|
## 定位
|
||||||
|
|
||||||
`diagnosis_run` 表示一次可回放的 Chat 或 AIOps 诊断执行。`run_id` 是运行级边界,Trace、反馈、自评估、案例沉淀和统计都应优先按 `run_id` 绑定。
|
`diagnosis_run` 表示一次可回放的 Chat 或 AIOps 诊断执行。`run_id` 是运行级边界,Trace、反馈、自评估、案例沉淀和统计都必须按 `run_id` 绑定。
|
||||||
|
|
||||||
## 字段
|
## 字段
|
||||||
|
|
||||||
@@ -32,7 +32,7 @@
|
|||||||
| 索引 | 字段 | 用途 |
|
| 索引 | 字段 | 用途 |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| `run_id` unique | `run_id` | 运行唯一约束 |
|
| `run_id` unique | `run_id` | 运行唯一约束 |
|
||||||
| `idx_diagnosis_run_session_created` | `session_id, created_at, id` | session 下最新运行解析和运行列表 |
|
| `idx_diagnosis_run_session_created` | `session_id, created_at, id` | session 下运行列表和排序 |
|
||||||
| `idx_diagnosis_run_session_run` | `session_id, run_id` | exact trace / feedback ownership 校验 |
|
| `idx_diagnosis_run_session_run` | `session_id, run_id` | exact trace / feedback ownership 校验 |
|
||||||
| `idx_diagnosis_run_status` | `status` | 状态筛选 |
|
| `idx_diagnosis_run_status` | `status` | 状态筛选 |
|
||||||
| `idx_diagnosis_run_agent_flow` | `agent_flow` | 区分 Chat / AIOps |
|
| `idx_diagnosis_run_agent_flow` | `agent_flow` | 区分 Chat / AIOps |
|
||||||
@@ -46,6 +46,6 @@
|
|||||||
|
|
||||||
## 注意点
|
## 注意点
|
||||||
|
|
||||||
- `GET /api/diagnosis/{sessionId}/trace` 未带 `runId` 时只为兼容解析 latest run;新 demo 和新客户端应传 `runId`。
|
- Trace 查询必须同时提供 `sessionId + runId`,并验证 Run 归属;缺少 `runId` 时直接拒绝。
|
||||||
- latest run 排序使用 `created_at DESC, id DESC`,避免 feedback 或自评估更新 `updated_at` 后改变回放目标。
|
- Run 列表按 `created_at DESC, id DESC` 排序,feedback 或自评估更新 `updated_at` 不改变运行顺序。
|
||||||
- 历史 `diagnosis_session` 会被迁移成兼容 run,但旧混合数据不能被还原成真实多轮边界。
|
- 当前 Java 运行时只使用 `chat_session + diagnosis_run`,不读取或写入旧会话级诊断表。
|
||||||
|
|||||||
+1
@@ -0,0 +1 @@
|
|||||||
|
|
||||||
+3
@@ -0,0 +1,3 @@
|
|||||||
|
committed_at: 2026-07-17
|
||||||
|
checkpoint: Commit
|
||||||
|
authorization: user-requested-direct-implementation
|
||||||
+135
@@ -0,0 +1,135 @@
|
|||||||
|
## Context
|
||||||
|
|
||||||
|
复杂 Chat 的公开调用链为 `ChatController -> ChatService.executeChatWithStrategy -> executeChatComplex`。阶段 2 已提供真实 Node adapters、显式 Gatekeeper、verified input、evidence retry、Composer/Fallback 和 `DiagnosisGraphFactory`,但 `executeChatComplex` 仍维护 SequentialAgent 两轮循环、VerifierInputHook/ThreadLocal 状态和私有 Composer 调用。阶段 3 要在不改变 `/api/chat` 的前提下切换唯一生产编排,并增加 Run 级 orchestration trace 数据契约。
|
||||||
|
|
||||||
|
当前 Run 通过 `SessionContextHolder(sessionId, runId)` 绑定 AgentStep/ToolInvocation;DiagnosisRun 使用字符串 JSON 保存 self evaluation;TraceService 将 Run 映射为 `run` 和兼容 `session` 投影。`DiagnosisOrchestrationTraceBuilder` 已能从有界 events 生成无 Prompt/raw output 的摘要。数据库由 Flyway 管理且 Hibernate 使用 validate,因此实体、V012 migration 和 Trace 映射必须同批对齐。
|
||||||
|
|
||||||
|
本 change 是 L4:内部状态机从固定顺序改为条件图,Trace run 对象和 DB 增加字段。`/api/chat`、sessionId/runId、Agent 输出协议保持兼容。用户已冻结单轨切换、安全 Fallback=SUCCESS、只有未处理失败=FAILED、阶段 5 才 live E2E。
|
||||||
|
|
||||||
|
## Goals / Non-Goals
|
||||||
|
|
||||||
|
**Goals:**
|
||||||
|
|
||||||
|
- 让复杂 Chat 一次执行真实 CompiledGraph,删除 ChatService 的细粒度 Sequential/round 状态机。
|
||||||
|
- 保持现有 Agent prompt/tool/skill/logging 组装能力,但 Graph Verifier 不注册 VerifierInputHook。
|
||||||
|
- 用 runId 同时作为 Graph threadId 和 Run 审计边界,用 sessionId/runId metadata 维持 AgentStep/ToolInvocation 归属。
|
||||||
|
- 将 final state 显式映射为 answer、verifier evaluation、composer audit、orchestration trace 和 Run lifecycle。
|
||||||
|
- 在正常/handled fallback 路径持久化非空、紧凑、安全的 orchestration trace,并只在 Trace `run` 对象暴露解析结果。
|
||||||
|
- 更新 Verifier Prompt/self-evaluation 到 verified-only 数据模型。
|
||||||
|
- 用阶段 3 focused tests证明生产 cutover、Run/Trace/API/Prompt 边界,且不留下已知失败测试。
|
||||||
|
|
||||||
|
**Non-Goals:**
|
||||||
|
|
||||||
|
- 不改简单 Chat/AIOps 编排。
|
||||||
|
- 不删除 VerifierInputHook/VerifierContextHolder 类型;它们在阶段 5 清理,但复杂 Chat 不再引用。
|
||||||
|
- 不增加持久 Graph checkpoint、恢复、并行分支或新 Run status。
|
||||||
|
- 不回填历史 run,不给历史 null orchestration trace 设计兼容伪值。
|
||||||
|
- 不在本阶段全面重命名/收敛测试夹具;阶段 4 完成测试体系替换。
|
||||||
|
- 不运行 Maven live E2E、日志或数据库验收。
|
||||||
|
|
||||||
|
## Decisions
|
||||||
|
|
||||||
|
### 1. 单轨生产 cutover,不增加 feature flag
|
||||||
|
|
||||||
|
`executeChatComplex` 只构造当前 Run、RunnableConfig 和 diagnosis context,然后调用真实 Graph runtime。旧 SequentialAgent loop、score/feature-flag retry 和 ChatService 私有 Verifier/Composer 路由方法从生产类移除;不保留双运行、shadow compare 或 runtime fallback 到 Sequential。
|
||||||
|
|
||||||
|
选择单轨而不是 feature flag,因为用户冻结的阶段门禁已经提供 Git commit 回滚边界,双轨会继续维护两套重试、Gatekeeper 和安全材料真理源。数据库新列 nullable,因此代码回滚时列可保留,无需破坏性 migration down。
|
||||||
|
|
||||||
|
### 2. Agent builder 只迁移职责,不复制实现
|
||||||
|
|
||||||
|
新增专用复杂 Chat Graph runtime/factory,复用当前 Planner/Executor/Verifier/Composer Prompt、knowledge map、history、method tools、ToolCallbacks、skill hooks 和 AgentLoggingHook 组装规则。ChatService 只向它提供本次 ChatModel、callbacks、history 和 Run 上下文。Graph Verifier hooks 只有 AgentLoggingHook,不包含 VerifierInputHook;Gatekeeper 只由显式 Node 调用。
|
||||||
|
|
||||||
|
如果为控制改动风险暂时保留简单 Chat 的工具/Hook builder,复杂 Agent builder 仍只存在一份。不得在 Graph Node 或 runtime 内复制 prompt 文本或第二套 tool catalog。
|
||||||
|
|
||||||
|
### 3. Graph 初始 state 和 RunnableConfig 使用最小白名单
|
||||||
|
|
||||||
|
初始 state 包含:
|
||||||
|
|
||||||
|
- `diagnosis_context={query, original_query}`;不把完整 history 放入 Graph State,history 仅注入 Planner/Executor system prompt。
|
||||||
|
- `planner_mode=NORMAL`。
|
||||||
|
- planner/verifier/composer/evidence retry count 全部为 0。
|
||||||
|
- `orchestration_events` 初始为空。
|
||||||
|
|
||||||
|
RunnableConfig 使用 `threadId(runId)`,metadata 包含 `sessionId` 和 `runId`。这同时满足 Graph 隔离、AgentLoggingHook/ToolCallback 的 Run ownership 和 Gatekeeper current-run validation。
|
||||||
|
|
||||||
|
### 4. 流式执行捕获最后真实 state
|
||||||
|
|
||||||
|
Graph runtime 使用 `CompiledGraph.stream(initialState, config)` 顺序消费 NodeOutput,并保存最后一个真实 `OverAllState`。正常 END 以最后 state 的非空 `final_answer` 为成功条件。流式消费与原同步 invoke 一样在当前请求线程阻塞,但允许在未处理异常时取得已真实产生的 partial state。
|
||||||
|
|
||||||
|
若异常前已有 events,failure handler 可以 best-effort 构造/persist partial trace;若没有 event,不生成虚假 node/transition。Agent invocation failures已经由 adapters 转为显式 status,正常预期失败都应走 Graph Fallback 而不是抛出。
|
||||||
|
|
||||||
|
### 5. Graph result mapper 是唯一运行结果翻译层
|
||||||
|
|
||||||
|
新增显式 mapper 从 final/partial state读取:
|
||||||
|
|
||||||
|
- `final_answer`。
|
||||||
|
- Verifier status/model/effective verdict、score、claim/fact checks、rationale、round。
|
||||||
|
- Gatekeeper raw audit、verified Executor output/evidence。
|
||||||
|
- Composer audit。
|
||||||
|
- `DiagnosisOrchestrationTraceBuilder` 结果。
|
||||||
|
|
||||||
|
ChatService 不再读取 VerifierContextHolder。self-evaluation 继续写 `verifier_evaluation` container 以兼容 Trace/Eval 消费者,但数据来自 Graph State:`executor_structured_output` 只保存 verified projection;新增 `verified_evidence` 和显式 statuses;不保存 raw Executor 或完整 tool_trace_summary。若 Verifier 从未完成,evaluation 只记录可用 status/audit/prompt 信息,不伪造 verdict。
|
||||||
|
|
||||||
|
### 6. Run lifecycle 和持久化顺序按执行结果分离
|
||||||
|
|
||||||
|
成功路径:Graph 返回非空安全 answer -> 构造 trace/evaluation -> 设置 Run SUCCESS、answer、orchestrationTrace、duration/metrics -> 保存 -> 调用 `evaluateRun`。
|
||||||
|
|
||||||
|
handled Fallback 与 Composer 正常路径使用相同成功顺序;可信度由 effective verdict 或 trace degraded 表达。失败路径:未处理异常、final state/answer 缺失、trace invariant 失败或成功结果持久化失败 -> Run FAILED,保存错误答案、duration、已有可构造 partial trace 和 metrics;不调用 success Eval。若 failure save 本身失败,只记录明确 error,不能声称持久化成功。
|
||||||
|
|
||||||
|
### 7. Orchestration trace 是独立 nullable JSON 数据契约
|
||||||
|
|
||||||
|
新增 V012,仅执行:
|
||||||
|
|
||||||
|
`ALTER TABLE diagnosis_run ADD COLUMN orchestration_trace JSON NULL ... AFTER self_evaluation`。
|
||||||
|
|
||||||
|
DiagnosisRun 使用 `@JdbcTypeCode(SqlTypes.JSON)` + `String orchestrationTrace`,与 selfEvaluation 写法一致。nullable 只服务无回填 migration/历史 run;每个成功的新 StateGraph Chat run 应用层必须非空。序列化使用 `DiagnosisOrchestrationTrace.toMap()`,字段保持 snake_case JSON。
|
||||||
|
|
||||||
|
### 8. Trace API 只在 RunTrace 增加解析对象
|
||||||
|
|
||||||
|
`DiagnosisTraceResponse.RunTrace` 新增 `Map<String,Object> orchestrationTrace`。DiagnosisTraceService 从 `DiagnosisRun.orchestrationTrace` 解析后映射;顶层 response、ChatSessionTrace、兼容 SessionTrace 和 raw string 均不新增字段。Run list 也不扩展该字段。
|
||||||
|
|
||||||
|
历史/AIOps run 的 null 值原样返回 null,不进行 fallback synthesis。解析非法 JSON 时 fail closed 为 null,并由数据/测试暴露问题;本阶段不增加历史兼容分支。
|
||||||
|
|
||||||
|
### 9. Verifier Prompt 与 verified-only input 对齐
|
||||||
|
|
||||||
|
`chat-verifier-prompt.md` 输入改为:`diagnosis_context`、`verified_executor_output`、`verified_evidence`、`gatekeeper_audit`、`verdict_ceiling`、可选 `retry_context`。删除 `executor_final_answer`、完整 `tool_trace_summary`、Hook parse status 和“再次执行 Gatekeeper”语义。
|
||||||
|
|
||||||
|
Verifier 仍输出原 JSON contract。`evidence_refs` 通过 claim_id/source_invocation_id/tool_name/raw_path 对 verified_evidence 建立审计关联,不要求不可用的 tool_trace_summary.trace_ref。Prompt audit catalog/version 随输入契约升级并在所有 handled paths 持久化。
|
||||||
|
|
||||||
|
### 10. 测试分阶段但不允许已知失败
|
||||||
|
|
||||||
|
阶段 3 新增最小生产 cutover测试:Graph runtime/final mapping、threadId/metadata、SUCCESS/Fallback/FAILED lifecycle、trace storage/API location、Prompt forbidden fields、no Verifier step on pre-verification failure。运行相关 Graph/Trace/Controller/Eval 回归和 test compilation。
|
||||||
|
|
||||||
|
旧 `ChatServiceSequentialAgentTest` 若因单轨语义失效,阶段 3 必须删除/改写冲突断言或由 focused Graph integration 替代,不能留到阶段 4 才让测试恢复绿色。阶段 4 继续完成命名、夹具、分支矩阵和旧 Hook tests 的全面清理。
|
||||||
|
|
||||||
|
## Interface Impact
|
||||||
|
|
||||||
|
- 等级:L4。
|
||||||
|
- 外部兼容:`/api/chat` request/response 与 run identity 不变。
|
||||||
|
- 加法协议:Trace `run.orchestrationTrace`;DB `diagnosis_run.orchestration_trace`。
|
||||||
|
- 内部行为:Executor/Planner/Gatekeeper failure 可提前 Fallback,不再保证固定 Agent 顺序;score flag 不再控制 LOW_CONFID retry。
|
||||||
|
- 消费者:Trace UI/demo/eval 可读取新 run 字段但不强制历史值;阶段 5 demo script 才增加最终 E2E 断言。
|
||||||
|
- 回滚:revert 本阶段代码/Prompt/spec;保留 nullable DB 列。无双轨开关、无数据回填回滚。
|
||||||
|
|
||||||
|
## Risks / Trade-offs
|
||||||
|
|
||||||
|
- [流式 API 与同步 invoke 的终止语义不同] → focused test断言最后 state、END、事件顺序和异常 partial state;不依赖实现私有字段。
|
||||||
|
- [self-evaluation 字段变化影响 Eval fixture] → 保留 container、verdict/claim/fact/gatekeeper/composer/prompt audit 兼容键;新增 verified 字段,移除不安全 full summary 前同步 specs/tests。
|
||||||
|
- [Agent factory 移动破坏 skill/tool hooks] → 复用现有 buildHooks/buildMethodTools 规则,并断言 Planner/Executor/Verifier/Composer hooks和当前 Run metadata。
|
||||||
|
- [JSON 字段三层漂移] → V012、entity、DTO/service 和 tests同批提交;Maven test compilation + strict specs。
|
||||||
|
- [无法取得 partial state] → 不伪造 trace;标记 FAILED 并记录明确原因。预期 Agent failures 全部由 Node adapters handled。
|
||||||
|
- [阶段 3/4 测试边界重叠] → 阶段 3只保证 cutover 可验收且现有 suite 不红;阶段 4负责全面测试架构替换。
|
||||||
|
|
||||||
|
## Migration Plan
|
||||||
|
|
||||||
|
1. 先加 V012、entity/Trace DTO/service 与 isolated mapping tests。
|
||||||
|
2. 实现 Graph runtime/result mapper和 Prompt verified-only 更新,先用 fake agents验证 final/partial state。
|
||||||
|
3. 将 `executeChatComplex` 单轨切换并删除旧私有 Sequential 状态机逻辑,保留 Run、metrics、Eval和简单 Chat路径。
|
||||||
|
4. 运行 production cutover、Graph、Trace、Controller/Eval 回归与 test compilation;证明 DB schema diff 只有一个字段。
|
||||||
|
5. 归档、提交阶段 3;阶段 4 再全面替换测试体系。
|
||||||
|
|
||||||
|
部署时 Flyway 先加 nullable 列,随后新代码写入。回滚为 Git revert;旧代码忽略新增列,数据库不删除列。
|
||||||
|
|
||||||
|
## Open Questions
|
||||||
|
|
||||||
|
无。所有产品/协议边界均由 ISS-011 与用户确认冻结;实现期若发现 Graph API 无法提供真实 partial state,只能回写本 design/tasks 后使用不伪造的 FAILED 处理,不能扩大协议。
|
||||||
+75
@@ -0,0 +1,75 @@
|
|||||||
|
# Chat Diagnosis StateGraph ChatService Cutover
|
||||||
|
|
||||||
|
## Why
|
||||||
|
|
||||||
|
阶段 2 已交付可构造、可测试的真实 Diagnosis StateGraph Nodes,但复杂 Chat 生产入口仍由 `ChatService.executeChatComplex(...)` 创建 `SequentialAgent`、维护 LOW_CONFID 外层循环,并依赖 `VerifierInputHook`/`VerifierContextHolder` 回传隐式状态。阶段 3 需要正式切换生产编排,让 ChatService 只管理 Run 生命周期、Agent 装配、Graph 调用和结果持久化,同时把 Graph 路由摘要作为当前 Run 的独立 Trace 维度保存和查询。
|
||||||
|
|
||||||
|
## What Changes
|
||||||
|
|
||||||
|
- 将复杂 Chat 从 `SequentialAgent` 外层循环切换为一次 `CompiledGraph` invocation,初始 state 使用显式 diagnosis context 和独立计数器。
|
||||||
|
- 为 Planner、Executor、Verifier、Composer 构造现有 ReactAgent 实例,经 `ReactAgentDiagnosisInvoker` 注入真实 Node actions;Graph Verifier 不注册旧 `VerifierInputHook`。
|
||||||
|
- 使用 `runId` 作为 Graph `threadId`,并把 `sessionId`、`runId` 放入 RunnableConfig metadata,继续复用 AgentStep/ToolInvocation 的 Run 归属。
|
||||||
|
- 将 Graph 最终 state 映射为现有 `ChatResult`、Verifier self-evaluation、Run 指标和生命周期状态。
|
||||||
|
- 新增 `diagnosis_run.orchestration_trace` nullable JSON 列、实体字段和紧凑序列化;新 StateGraph Chat run 在应用契约上必须写入非空摘要。
|
||||||
|
- Trace API 只在 `run.orchestrationTrace` 返回解析 JSON,不在顶层、兼容 `session` 投影或 raw 字段重复。
|
||||||
|
- 同步 Verifier Prompt 到 verified-only Graph 输入,移除完整 tool trace、raw Executor 和 Hook gatekeeper 输入说明。
|
||||||
|
- 增加生产 cutover、Run/Trace、safe Fallback、thread/metadata、Prompt 边界的必要单元/集成测试。
|
||||||
|
|
||||||
|
## Capabilities
|
||||||
|
|
||||||
|
### New Capabilities
|
||||||
|
|
||||||
|
- `chat-diagnosis-stategraph-chatservice-cutover`:规定复杂 Chat 的 StateGraph 生产调用、Run 生命周期、结果映射、编排摘要持久化与 Trace API 投影。
|
||||||
|
|
||||||
|
### Modified Capabilities
|
||||||
|
|
||||||
|
- `chat-diagnosis-stategraph-real-nodes`:阶段 2 的生产隔离约束改为阶段 3 正式切换,真实 Nodes 成为复杂 Chat 唯一生产编排。
|
||||||
|
- `chat-verifier-agent`:Verifier 输入改为 Gatekeeper passed binding 投影后的 verified-only payload;self-evaluation 以 Graph state 为来源。
|
||||||
|
- `chat-composer-agent`:Composer 由 Graph Node 调用,但 allowed-material 与审计契约保持不变。
|
||||||
|
- `session-run-trace-isolation`:Run 增加独立 orchestration trace,Trace API 只在精确 run 对象暴露解析结果。
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
### In Scope
|
||||||
|
|
||||||
|
- `ChatService.executeChatComplex(...)` 的完整生产 cutover 和不再使用 SequentialAgent 的私有编排逻辑清理。
|
||||||
|
- 现有 Prompt/Agent builder 的 Graph 适配;不复制第二套 Agent factory。
|
||||||
|
- DiagnosisRun、Flyway、Trace DTO/Service 的加法式 orchestration trace 支持。
|
||||||
|
- Graph final state 到 answer、verifier evaluation、composer audit、Run status/metrics 的映射。
|
||||||
|
- 阶段 3 风险所需的 focused tests 与现有 Controller/Trace/Eval 契约回归。
|
||||||
|
|
||||||
|
### Out of Scope
|
||||||
|
|
||||||
|
- 不删除尚未被其他历史测试引用的 `VerifierInputHook`、`VerifierContextHolder` 类;阶段 5 统一清理。
|
||||||
|
- 不在阶段 3 完成整个测试体系命名/夹具迁移;阶段 4 负责全面替换旧 Sequential 测试体系,但阶段 3 不允许留下已知失败测试。
|
||||||
|
- 不修改 `/api/chat` 请求/响应结构或 Executor/Verifier/Composer 输出协议。
|
||||||
|
- 不新增 Run status,不做历史 run 的 orchestration trace 回填或兼容读取分支。
|
||||||
|
- 不运行 Maven live E2E、`logs/` 或数据库查询;统一保留到阶段 5。
|
||||||
|
|
||||||
|
## Context Constraints
|
||||||
|
|
||||||
|
- 阶段 0–2 archives 和主 specs 是实现基线;Graph 的 Node、条件边、计数所有权和安全 Fallback 边界不得在本阶段复制或放宽。
|
||||||
|
- `diagnosis_run.status` 只表达执行生命周期;Composer、LOW_CONFID、Verifier REJECT 或固定安全 Fallback 只要生成安全答案均为 SUCCESS。
|
||||||
|
- 编排摘要必须由有界 `orchestration_events` 构造,不从日志反推,不保存 Prompt、reasoning、raw tool output 或 Graph State snapshot。
|
||||||
|
- self-evaluation 与 orchestration trace 分离;AgentStep/ToolInvocation 继续作为详细 Trace 数据源。
|
||||||
|
- 新 Graph Verifier 只接收 `verified_executor_output`、`verified_evidence`、Gatekeeper audit/ceiling、query 和 retry context。
|
||||||
|
- 运行失败时不得伪造未发生的 events;可取得的部分 state 才允许 best-effort 持久化。
|
||||||
|
|
||||||
|
## Acceptance
|
||||||
|
|
||||||
|
- 复杂 Chat 生产代码不再创建或调用 SequentialAgent,且只调用一次 CompiledGraph。
|
||||||
|
- RunnableConfig 的 threadId 等于 runId,metadata 同时包含当前 sessionId/runId。
|
||||||
|
- Graph final answer 非空时映射为现有 ChatResult,Run 为 SUCCESS;无法产生安全响应或未处理失败时 Run 为 FAILED。
|
||||||
|
- 每个新 StateGraph Chat run 都持久化非空、无敏感材料的 orchestration trace;Trace API 只在 `run.orchestrationTrace` 返回解析对象。
|
||||||
|
- AgentStep、ToolInvocation、self-evaluation、answer、duration、token/step/tool count 和 Eval 仍绑定当前 runId。
|
||||||
|
- Verifier Prompt 和实际 verified-only payload 一致,Graph Verifier 不注册旧 input Hook。
|
||||||
|
- Executor/Gatekeeper 等前置失败的 Graph 路径不产生 Verifier AgentStep。
|
||||||
|
- `/api/chat` 和证据协议保持不变;数据库 schema 只新增一个 nullable JSON 列。
|
||||||
|
- 阶段 3 focused tests、Trace/Controller/Eval 回归、test compilation 和 OpenSpec strict validation 通过;不运行 live E2E。
|
||||||
|
|
||||||
|
## Risks
|
||||||
|
|
||||||
|
- ChatService 当前同时承载 Agent factory、Run 生命周期和旧解析逻辑;cutover 必须做完整内部重构,避免留下双编排路径或死代码。
|
||||||
|
- self-evaluation 旧字段曾依赖 ThreadLocal/full tool summary;Graph 映射必须明确兼容字段和安全边界,不能重新泄漏未验真材料。
|
||||||
|
- Graph invoke 异常可能没有可用 final state;失败处理只能持久化真实可取得的 partial snapshot,不能编造路径。
|
||||||
|
- JSON entity/DDL/Trace DTO 必须保持同一字段语义,否则 JPA validate、数据库迁移或 API 解析会漂移。
|
||||||
+60
@@ -0,0 +1,60 @@
|
|||||||
|
## MODIFIED Requirements
|
||||||
|
|
||||||
|
### Requirement: Composer SHALL generate final user-facing Chat answers
|
||||||
|
The system SHALL invoke the Composer Graph Node after Verifier routing to generate the final user-facing Chat answer from Verifier-allowed material.
|
||||||
|
|
||||||
|
#### Scenario: Composer receives only filtered material
|
||||||
|
- **WHEN** the Composer Graph Node invokes its configured Agent
|
||||||
|
- **THEN** the Composer input SHALL contain `original_query`, `verdict`, `allowed_claims`, `allowed_hypotheses`, `missing_info`, `recommended_actions`, and `rationale`
|
||||||
|
- **AND** the Composer input SHALL NOT contain raw tool output
|
||||||
|
- **AND** the Composer input SHALL NOT contain the full unscreened Executor output
|
||||||
|
- **AND** the Composer input SHALL NOT contain Executor `user_facing_answer`
|
||||||
|
|
||||||
|
#### Scenario: Composer outputs strict JSON
|
||||||
|
- **WHEN** Composer completes
|
||||||
|
- **THEN** it SHALL output exactly one JSON object
|
||||||
|
- **AND** the JSON object SHALL include `answer_summary`, `recommended_actions`, and `user_facing_answer`
|
||||||
|
- **AND** it SHALL NOT output Markdown, code fences, or explanatory text outside the JSON object
|
||||||
|
|
||||||
|
#### Scenario: Composer does not introduce new facts
|
||||||
|
- **WHEN** Composer produces `answer_summary`, `recommended_actions`, or `user_facing_answer`
|
||||||
|
- **THEN** every service name, entity, timestamp, error code, metric value, root cause, and recommendation reason SHALL be derived from the Composer input
|
||||||
|
- **AND** Composer SHALL NOT add facts from model knowledge, raw tool history, or Executor raw text
|
||||||
|
|
||||||
|
### Requirement: Composer input SHALL honor Verifier claim checks
|
||||||
|
The Composer Graph Node SHALL construct Composer input by filtering verified Executor structured output through Verifier `claim_checks`.
|
||||||
|
|
||||||
|
#### Scenario: Passing claims become allowed claims
|
||||||
|
- **WHEN** a claim check verification is `direct_observation`
|
||||||
|
- **THEN** the Composer input builder SHALL include the matching verified Executor claim in `allowed_claims`
|
||||||
|
|
||||||
|
#### Scenario: Reasonable inferences remain bounded
|
||||||
|
- **WHEN** a claim check verification is `reasonable_inference`
|
||||||
|
- **THEN** the Composer input builder MAY include the matching verified Executor claim in `allowed_claims`
|
||||||
|
- **AND** the final answer SHALL NOT describe it as the sole confirmed root cause unless the allowed claim itself is a root-cause claim and the final verdict is `PASS`
|
||||||
|
|
||||||
|
#### Scenario: Overstated claims are not confirmed findings
|
||||||
|
- **WHEN** a claim check verification is `overstated`
|
||||||
|
- **THEN** the Composer input builder SHALL NOT include the matching claim as a confirmed item in `allowed_claims`
|
||||||
|
- **AND** it MAY include it as `allowed_hypotheses` or represent it in `missing_info`
|
||||||
|
|
||||||
|
#### Scenario: Unsupported or external claims are withheld
|
||||||
|
- **WHEN** a claim check verification is `unsupported`, `external_unknown`, or `contradicted`
|
||||||
|
- **THEN** the Composer input builder SHALL NOT include the matching claim in `allowed_claims`
|
||||||
|
- **AND** the final user-facing answer SHALL NOT present that claim as confirmed
|
||||||
|
|
||||||
|
### Requirement: Composer failures SHALL degrade safely
|
||||||
|
The system SHALL tolerate malformed or failed Composer execution without leaking raw JSON or unverified Executor material.
|
||||||
|
|
||||||
|
#### Scenario: malformed Composer output falls back safely
|
||||||
|
- **WHEN** Composer exhausts its fixed-input technical retry or returns a non-retryable failure
|
||||||
|
- **THEN** the deterministic Fallback Node SHALL produce a final answer using only filtered material
|
||||||
|
- **AND** the final answer SHALL NOT expose raw Composer output
|
||||||
|
- **AND** the final answer SHALL NOT expose raw Executor output
|
||||||
|
- **AND** the final answer SHALL NOT use Executor `user_facing_answer`
|
||||||
|
|
||||||
|
#### Scenario: Composer audit is persisted
|
||||||
|
- **WHEN** Graph result mapping persists verifier evaluation
|
||||||
|
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation.composer_output` SHALL record parsed Composer audit when available
|
||||||
|
- **AND** handled Composer fallback SHALL be observable through orchestration trace and available status/reason fields
|
||||||
|
- **AND** the audit SHALL remain compact and SHALL NOT store full raw tool output
|
||||||
+107
@@ -0,0 +1,107 @@
|
|||||||
|
## ADDED Requirements
|
||||||
|
|
||||||
|
### Requirement: Complex Chat SHALL use the real Diagnosis StateGraph as its only production orchestrator
|
||||||
|
|
||||||
|
The system SHALL execute each complex Chat request through one real Diagnosis CompiledGraph and SHALL NOT create, invoke, or fall back to a SequentialAgent workflow.
|
||||||
|
|
||||||
|
#### Scenario: Complex Chat starts
|
||||||
|
|
||||||
|
- **WHEN** `executeChatWithStrategy` classifies a valid question as complex
|
||||||
|
- **THEN** ChatService SHALL create the current Diagnosis Run and invoke one Diagnosis CompiledGraph
|
||||||
|
- **AND** the production path SHALL NOT maintain an outer score-based retry loop
|
||||||
|
|
||||||
|
#### Scenario: Graph dependencies are assembled
|
||||||
|
|
||||||
|
- **WHEN** the complex Chat Graph is constructed
|
||||||
|
- **THEN** Planner, Executor, Verifier, and Composer SHALL use the existing project Prompt, tool, skill, and AgentLoggingHook assembly rules
|
||||||
|
- **AND** the Graph Verifier SHALL NOT register VerifierInputHook
|
||||||
|
- **AND** Gatekeeper SHALL run only as the explicit Graph Node
|
||||||
|
|
||||||
|
### Requirement: Graph invocation SHALL preserve current Run ownership
|
||||||
|
|
||||||
|
The system SHALL use the current `runId` as Graph `threadId` and SHALL pass both current `sessionId` and `runId` in RunnableConfig metadata.
|
||||||
|
|
||||||
|
#### Scenario: Agent Node runs
|
||||||
|
|
||||||
|
- **WHEN** any Agent Node is invoked for a complex Chat run
|
||||||
|
- **THEN** its RunnableConfig threadId SHALL equal the current runId
|
||||||
|
- **AND** AgentStep and ToolInvocation writes SHALL retain the current sessionId and runId
|
||||||
|
|
||||||
|
#### Scenario: Gatekeeper validates Executor output
|
||||||
|
|
||||||
|
- **WHEN** the Gatekeeper Node runs
|
||||||
|
- **THEN** it SHALL validate only tool invocations belonging to the runId in RunnableConfig metadata
|
||||||
|
- **AND** data from another run in the same session SHALL NOT be considered
|
||||||
|
|
||||||
|
### Requirement: Initial Graph State SHALL contain only bounded diagnosis control data
|
||||||
|
|
||||||
|
ChatService SHALL initialize Diagnosis Graph State with the current query context, NORMAL Planner mode, zero independent retry counters, and an empty orchestration event list.
|
||||||
|
|
||||||
|
#### Scenario: Initial state is projected
|
||||||
|
|
||||||
|
- **WHEN** a complex Chat run enters Planner for the first time
|
||||||
|
- **THEN** `diagnosis_context` SHALL contain the current query/original query
|
||||||
|
- **AND** complete conversation history SHALL NOT be stored in parent Graph State
|
||||||
|
- **AND** history MAY remain in the Planner and Executor system Prompt assembled for this request
|
||||||
|
|
||||||
|
#### Scenario: Counters are initialized
|
||||||
|
|
||||||
|
- **WHEN** the Graph starts
|
||||||
|
- **THEN** planner, verifier, composer, and evidence retry counters SHALL each be zero
|
||||||
|
- **AND** Planner mode SHALL be NORMAL
|
||||||
|
|
||||||
|
### Requirement: Graph final state SHALL map to the existing Chat result and Run lifecycle
|
||||||
|
|
||||||
|
The system SHALL use non-empty `final_answer` from a handled Graph terminal state as the existing ChatResult answer. Run status SHALL express execution lifecycle rather than diagnosis quality.
|
||||||
|
|
||||||
|
#### Scenario: Composer completes
|
||||||
|
|
||||||
|
- **WHEN** Graph reaches Composer and produces a safe non-empty final answer
|
||||||
|
- **THEN** ChatResult SHALL preserve the current answer/sessionId/runId protocol
|
||||||
|
- **AND** DiagnosisRun SHALL be saved as SUCCESS with answer, duration, token count, step count, and tool count
|
||||||
|
- **AND** EvaluationService SHALL evaluate the current runId
|
||||||
|
|
||||||
|
#### Scenario: Handled Fallback completes
|
||||||
|
|
||||||
|
- **WHEN** Graph reaches deterministic Fallback and produces a safe non-empty answer
|
||||||
|
- **THEN** DiagnosisRun SHALL be SUCCESS
|
||||||
|
- **AND** diagnosis quality SHALL be expressed by verifier fields when available or `orchestrationTrace.degraded=true`
|
||||||
|
- **AND** no DEGRADED Run status SHALL be introduced
|
||||||
|
|
||||||
|
#### Scenario: Graph cannot produce a safe response
|
||||||
|
|
||||||
|
- **WHEN** Graph has an unhandled failure, final state is unavailable, final answer is blank, or required successful-result persistence fails
|
||||||
|
- **THEN** DiagnosisRun SHALL be marked FAILED when it can still be saved
|
||||||
|
- **AND** the system SHALL NOT report a successful Graph result
|
||||||
|
|
||||||
|
### Requirement: Graph state SHALL be the only source for verifier evaluation persistence
|
||||||
|
|
||||||
|
The system SHALL build `diagnosis_run.self_evaluation.verifier_evaluation` from explicit Graph State and SHALL NOT read VerifierContextHolder on the complex Chat path.
|
||||||
|
|
||||||
|
#### Scenario: Verifier completed
|
||||||
|
|
||||||
|
- **WHEN** final Graph State contains a completed Verifier result
|
||||||
|
- **THEN** verifier evaluation SHALL include execution status, model verdict, effective verdict, groundedness, claim/fact checks, rationale, round, Gatekeeper audit, verified Executor output/evidence, Prompt audit, and Composer audit when available
|
||||||
|
- **AND** the compatibility `verdict` field SHALL equal effective verdict
|
||||||
|
|
||||||
|
#### Scenario: Pre-verification Fallback completed
|
||||||
|
|
||||||
|
- **WHEN** Graph reaches Fallback before Verifier completes
|
||||||
|
- **THEN** verifier evaluation SHALL preserve available execution statuses, Gatekeeper audit, Prompt audit, and fallback context
|
||||||
|
- **AND** it SHALL NOT fabricate a model or effective verdict
|
||||||
|
|
||||||
|
#### Scenario: Evaluation payload is inspected
|
||||||
|
|
||||||
|
- **WHEN** verifier evaluation is persisted
|
||||||
|
- **THEN** it SHALL NOT contain raw Executor text or complete tool trace summary
|
||||||
|
- **AND** `executor_structured_output` SHALL contain at most the verified projection retained for compatibility
|
||||||
|
|
||||||
|
### Requirement: Stage 3 verification SHALL not run the final live E2E
|
||||||
|
|
||||||
|
The change SHALL use focused automated tests for production cutover and contracts while reserving Maven live startup, log inspection, and database querying for stage 5.
|
||||||
|
|
||||||
|
#### Scenario: Stage 3 is accepted
|
||||||
|
|
||||||
|
- **WHEN** stage 3 verification completes
|
||||||
|
- **THEN** Graph cutover, Run/Trace, Prompt, Controller/Eval regressions, Maven test compilation, and OpenSpec strict validation SHALL have passed
|
||||||
|
- **AND** live E2E, `logs/`, and `scripts/query_mysql.py` SHALL be recorded as intentionally deferred to stage 5
|
||||||
+7
@@ -0,0 +1,7 @@
|
|||||||
|
## REMOVED Requirements
|
||||||
|
|
||||||
|
### Requirement: Real Nodes SHALL remain isolated from the production Chat path in stage 2
|
||||||
|
|
||||||
|
**Reason**: Stage 2 production isolation has completed its migration purpose; stage 3 intentionally makes the real Diagnosis Graph the only complex Chat production orchestrator.
|
||||||
|
|
||||||
|
**Migration**: Replace `ChatService.executeChatComplex` SequentialAgent orchestration with real Graph assembly/invocation. Keep `/api/chat` compatible and use a full Git revert of stage 3 for runtime rollback rather than retaining dual paths.
|
||||||
+263
@@ -0,0 +1,263 @@
|
|||||||
|
## MODIFIED Requirements
|
||||||
|
|
||||||
|
### Requirement: Verifier SHALL fact-check Executor answers
|
||||||
|
The system SHALL have a Verifier Agent that reads only Gatekeeper-projected structured Executor claims and verified claim-local evidence, then produces a structured verdict based on claim derivability.
|
||||||
|
|
||||||
|
#### Scenario: PASS verdict when all claims have evidence
|
||||||
|
- **WHEN** all critical claims in `verified_executor_output.claims` have direct observation or reasonable inference support in `verified_evidence`
|
||||||
|
- **AND** at least one critical claim has direct observation
|
||||||
|
- **AND** no critical claim is contradicted, unsupported, external unknown, or overstated
|
||||||
|
- **AND** verdict ceiling is PASS
|
||||||
|
- **THEN** the Verifier MAY output model verdict="PASS" with groundedness_score ≥ 0.5
|
||||||
|
|
||||||
|
#### Scenario: LOW_CONFID verdict with partial evidence
|
||||||
|
- **WHEN** no critical claim contradicts verified evidence
|
||||||
|
- **AND** some critical claims are `unsupported`, `external_unknown`, or `overstated`
|
||||||
|
- **THEN** the Verifier SHALL output model verdict="LOW_CONFID"
|
||||||
|
|
||||||
|
#### Scenario: LOW_CONFID verdict with only inference support
|
||||||
|
- **WHEN** no critical claim contradicts verified evidence
|
||||||
|
- **AND** all critical claims are only `reasonable_inference`
|
||||||
|
- **THEN** the Verifier SHALL output model verdict="LOW_CONFID"
|
||||||
|
|
||||||
|
#### Scenario: REJECT verdict when claims contradict evidence
|
||||||
|
- **WHEN** any critical claim in `verified_executor_output.claims` contradicts verified evidence
|
||||||
|
- **OR** the claim fabricates a key entity, error code, or conclusion that does not exist in verified evidence
|
||||||
|
- **THEN** the Verifier SHALL output model verdict="REJECT"
|
||||||
|
|
||||||
|
#### Scenario: Verified structured claims are the only verification target
|
||||||
|
- **WHEN** `verified_executor_output.claims` is present
|
||||||
|
- **THEN** Verifier SHALL verify each structured claim against matching `verified_evidence` through `claim_checks`
|
||||||
|
- **AND** each claim's evidence references SHALL match existing claim/invocation/tool/path identifiers when available
|
||||||
|
- **AND** a claim without matching verified evidence SHALL NOT be classified as `direct_observation`
|
||||||
|
- **AND** Verifier SHALL NOT receive or add confirmed facts from raw Executor text
|
||||||
|
|
||||||
|
#### Scenario: Executor output is invalid
|
||||||
|
- **WHEN** Executor does not return a legal structured contract
|
||||||
|
- **THEN** the Graph SHALL route directly to pre-verification Fallback
|
||||||
|
- **AND** Verifier SHALL NOT execute or fabricate a diagnostic verdict
|
||||||
|
|
||||||
|
### Requirement: facts_checked SHALL use a fixed classification set
|
||||||
|
The system SHALL continue to expose compatibility `facts_checked` using its fixed verification classification set.
|
||||||
|
|
||||||
|
#### Scenario: claim checks are mapped to legacy facts
|
||||||
|
- **WHEN** Verifier output contains `claim_checks`
|
||||||
|
- **THEN** the shared Verifier protocol parser SHALL derive compatibility `facts_checked` when the model did not provide them
|
||||||
|
- **AND** `direct_observation` SHALL map to `direct_evidence`
|
||||||
|
- **AND** `reasonable_inference` and `overstated` SHALL map to `indirect_support`
|
||||||
|
- **AND** `unsupported` and `external_unknown` SHALL map to `no_evidence`
|
||||||
|
- **AND** `contradicted` SHALL map to `contradicted`
|
||||||
|
|
||||||
|
### Requirement: ChatService SHALL route based on Verifier verdict
|
||||||
|
The system SHALL use Diagnosis StateGraph conditional edges, rather than a ChatService outer loop, to route explicit Verifier execution status and effective verdict.
|
||||||
|
|
||||||
|
#### Scenario: PASS routes to Composer
|
||||||
|
- **WHEN** Verifier completes with effective verdict="PASS"
|
||||||
|
- **THEN** the Graph SHALL invoke Composer with filtered Verifier-allowed material
|
||||||
|
- **AND** the final user-facing answer SHALL NOT pass through raw Executor output
|
||||||
|
- **AND** the final user-facing answer SHALL NOT read Executor `user_facing_answer`
|
||||||
|
|
||||||
|
#### Scenario: LOW_CONFID does not qualify for evidence retry
|
||||||
|
- **WHEN** Verifier completes LOW_CONFID but ceiling is LOW_CONFID, no valid critical evidence gap exists, or evidence retry count is already one
|
||||||
|
- **THEN** the Graph SHALL route to Composer without another Planner cycle
|
||||||
|
- **AND** the final answer SHALL distinguish confirmed information, possible directions, and evidence gaps
|
||||||
|
|
||||||
|
#### Scenario: LOW_CONFID qualifies for evidence retry
|
||||||
|
- **WHEN** Verifier completes LOW_CONFID with ceiling PASS, at least one critical valid evidence gap, and evidence retry count zero
|
||||||
|
- **THEN** the Graph SHALL invoke one EVIDENCE_GAP_ONLY Planner cycle
|
||||||
|
- **AND** it SHALL NOT use groundedness threshold or a ChatService feature flag to decide the retry
|
||||||
|
|
||||||
|
#### Scenario: REJECT does not enter retry round
|
||||||
|
- **WHEN** Verifier completes with effective verdict="REJECT"
|
||||||
|
- **THEN** the Graph SHALL NOT start an evidence supplementation round
|
||||||
|
- **AND** it SHALL route to Composer-safe output
|
||||||
|
|
||||||
|
#### Scenario: REJECT produces bounded output
|
||||||
|
- **WHEN** effective verdict is REJECT
|
||||||
|
- **THEN** the system SHALL output a degraded result indicating current evidence cannot support a reliable conclusion
|
||||||
|
- **AND** it SHALL NOT pass through raw Executor answer
|
||||||
|
- **AND** it SHALL NOT include an unsupported root-cause conclusion
|
||||||
|
|
||||||
|
#### Scenario: Verifier execution fails
|
||||||
|
- **WHEN** Verifier exhausts technical retry or returns a non-retryable failure
|
||||||
|
- **THEN** the Graph SHALL route to pre-verification Fallback
|
||||||
|
- **AND** no execution status string SHALL be used as model or effective verdict
|
||||||
|
|
||||||
|
### Requirement: Verifier SHALL be observable
|
||||||
|
The Verifier execution, effective verdict, and downstream final-answer composition SHALL be persisted in the current Diagnosis Run self-evaluation container.
|
||||||
|
|
||||||
|
#### Scenario: claim checks written to self_evaluation
|
||||||
|
- **WHEN** a completed Verifier evaluation is persisted
|
||||||
|
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include `claim_checks`
|
||||||
|
- **AND** it SHALL continue to include compatibility `facts_checked`
|
||||||
|
- **AND** it SHALL include `verifier_status`, `model_verdict`, `effective_verdict`, `verdict`, `groundedness_score`, `rationale`, verified output/evidence, and Gatekeeper audit
|
||||||
|
|
||||||
|
#### Scenario: composer output written to self_evaluation
|
||||||
|
- **WHEN** final answer composition completes
|
||||||
|
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include compact `composer_output` when available
|
||||||
|
- **AND** handled Composer fallback SHALL remain observable through orchestration trace and status/reason fields
|
||||||
|
- **AND** existing claim/fact and Gatekeeper fields SHALL be preserved
|
||||||
|
|
||||||
|
#### Scenario: verdict written to self_evaluation
|
||||||
|
- **WHEN** Verifier completes
|
||||||
|
- **THEN** Graph result mapping SHALL write effective verdict under `diagnosis_run.self_evaluation.verifier_evaluation.verdict`
|
||||||
|
- **AND** existing `rule_evaluation` and `aiops_rule_evaluation` channels SHALL be preserved
|
||||||
|
|
||||||
|
#### Scenario: pre-verification fallback is persisted
|
||||||
|
- **WHEN** Graph reaches Fallback before Verifier completes
|
||||||
|
- **THEN** verifier evaluation SHALL include available status, Gatekeeper audit, failure reason, and Prompt audit
|
||||||
|
- **AND** it SHALL NOT fabricate `model_verdict` or `effective_verdict`
|
||||||
|
|
||||||
|
#### Scenario: gatekeeper result written to self_evaluation
|
||||||
|
- **WHEN** Graph result mapping persists available Gatekeeper state
|
||||||
|
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result`
|
||||||
|
- **AND** the result SHALL retain status, severity, checked bindings, rules, failed rules, warnings, and errors when provided by Gatekeeper
|
||||||
|
|
||||||
|
#### Scenario: prompt audit written to verifier evaluation
|
||||||
|
- **WHEN** a complex Chat Graph result is persisted
|
||||||
|
- **THEN** the system SHALL include a `prompt_audit` object under `diagnosis_run.self_evaluation.verifier_evaluation`
|
||||||
|
- **AND** `prompt_audit.version` SHALL identify the Chat Prompt audit catalog version
|
||||||
|
- **AND** `prompt_audit.prompts` SHALL include Planner, Executor, Verifier, and Composer Prompt names and versions
|
||||||
|
- **AND** full Prompt text SHALL NOT be persisted
|
||||||
|
|
||||||
|
#### Scenario: prompt audit available on fallback paths
|
||||||
|
- **WHEN** Planner, Executor, Gatekeeper, Verifier, or Composer reaches a handled Fallback
|
||||||
|
- **THEN** the persisted verifier evaluation SHALL still include `prompt_audit`
|
||||||
|
|
||||||
|
#### Scenario: evaluation payload is inspected
|
||||||
|
- **WHEN** Graph verifier evaluation is persisted
|
||||||
|
- **THEN** it SHALL NOT contain raw Executor text or complete `tool_trace_summary`
|
||||||
|
- **AND** compatibility `executor_structured_output` SHALL contain at most the verified projection
|
||||||
|
|
||||||
|
### Requirement: Verifier SHALL consume explicit verification inputs
|
||||||
|
The Verifier SHALL receive a Graph-built verified-only payload rather than inferring business inputs from conversation history, ThreadLocal state, raw Executor text, or complete tool history.
|
||||||
|
|
||||||
|
#### Scenario: explicit input blocks available to Verifier
|
||||||
|
- **WHEN** the Verifier Graph Node starts
|
||||||
|
- **THEN** the payload SHALL provide `diagnosis_context`, `verified_executor_output`, `verified_evidence`, `gatekeeper_audit`, and `verdict_ceiling`
|
||||||
|
- **AND** permitted structured `retry_context` SHALL be provided only after evidence retry preparation
|
||||||
|
|
||||||
|
#### Scenario: Verifier remains isolated from intermediate and raw material
|
||||||
|
- **WHEN** the Verifier input is serialized
|
||||||
|
- **THEN** it SHALL exclude Planner reasoning, Executor intermediate reasoning, raw Executor text, complete tool trace summary, Prompt text, and unrelated parent Graph State
|
||||||
|
|
||||||
|
#### Scenario: only passed bindings are available
|
||||||
|
- **WHEN** Gatekeeper returns mixed passed and failed checked bindings
|
||||||
|
- **THEN** `verified_executor_output` and `verified_evidence` SHALL contain only claims/material matching passed bindings
|
||||||
|
- **AND** the Verifier SHALL NOT receive failed or unreferenced tool material
|
||||||
|
|
||||||
|
#### Scenario: verified evidence preserves precise references
|
||||||
|
- **WHEN** the system prepares Verifier input
|
||||||
|
- **THEN** each verified evidence item SHALL preserve claim id, source invocation id, tool name, raw path, and matched text
|
||||||
|
- **AND** the item SHALL be traceable to current-run Gatekeeper validation
|
||||||
|
|
||||||
|
#### Scenario: gatekeeper audit and ceiling are available
|
||||||
|
- **WHEN** the system prepares Verifier input
|
||||||
|
- **THEN** the payload SHALL include raw Gatekeeper audit separately from normalized verdict ceiling
|
||||||
|
- **AND** a LOW_CONFID ceiling SHALL prevent effective PASS
|
||||||
|
|
||||||
|
#### Scenario: technical retry occurs
|
||||||
|
- **WHEN** the first Verifier attempt returns invalid output or a retryable invocation failure
|
||||||
|
- **THEN** the second attempt SHALL receive byte-identical serialized input
|
||||||
|
- **AND** Executor, Gatekeeper, and tools SHALL NOT rerun
|
||||||
|
|
||||||
|
### Requirement: Verifier facts SHALL be auditable
|
||||||
|
Verifier claims and facts SHALL be linkable to the verified binding projection used during verification.
|
||||||
|
|
||||||
|
#### Scenario: claim checks contain evidence refs
|
||||||
|
- **WHEN** the Verifier emits `claim_checks`
|
||||||
|
- **THEN** each check SHALL include an `evidence_refs` array
|
||||||
|
- **AND** any non-empty evidence ref SHALL correspond to existing verified evidence by claim id, source invocation id, tool name, or raw path
|
||||||
|
- **AND** it SHALL NOT reference a failed or unverified binding
|
||||||
|
|
||||||
|
#### Scenario: verifier evaluation persists traceability snapshot
|
||||||
|
- **WHEN** Graph result mapping persists verifier evaluation
|
||||||
|
- **THEN** it SHALL include `traceability_version`
|
||||||
|
- **AND** it SHALL include the bounded `verified_evidence` snapshot used by the Verifier
|
||||||
|
- **AND** it SHALL NOT persist a complete tool trace summary as Verifier input
|
||||||
|
|
||||||
|
### Requirement: Structured Executor output SHALL degrade safely
|
||||||
|
The StateGraph runtime SHALL tolerate malformed or absent structured Executor output without crashing the Chat flow or invoking Verifier with untrusted material.
|
||||||
|
|
||||||
|
#### Scenario: Malformed Executor JSON is classified
|
||||||
|
- **WHEN** Executor returns malformed JSON or text outside the expected contract
|
||||||
|
- **THEN** Executor Node SHALL set INVALID_OUTPUT
|
||||||
|
- **AND** the Graph SHALL route directly to deterministic pre-verification Fallback
|
||||||
|
- **AND** Gatekeeper, Verifier, and model Composer SHALL NOT execute
|
||||||
|
|
||||||
|
#### Scenario: Structured parse failure remains observable
|
||||||
|
- **WHEN** Executor output parsing fails
|
||||||
|
- **THEN** orchestration events and verifier evaluation status/failure fields SHALL make the parse failure visible
|
||||||
|
- **AND** the failure SHALL NOT be treated as a successful evidence-attribution contract or diagnostic verdict
|
||||||
|
|
||||||
|
### Requirement: Executor Gatekeeper SHALL validate deterministic structured-output failures
|
||||||
|
The system SHALL run deterministic Gatekeeper checks as an explicit Graph Node after legal Executor output parsing and before Verifier model execution.
|
||||||
|
|
||||||
|
#### Scenario: schema rule rejects removed fields
|
||||||
|
- **WHEN** Executor structured output contains `diagnosis_summary` or `user_facing_answer`
|
||||||
|
- **THEN** `gatekeeper_result.status` SHALL be `fail`
|
||||||
|
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
|
||||||
|
|
||||||
|
#### Scenario: schema rule rejects missing evidence bindings
|
||||||
|
- **WHEN** a confirmed claim has no `evidence_bindings`
|
||||||
|
- **THEN** `gatekeeper_result.status` SHALL be `fail`
|
||||||
|
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
|
||||||
|
|
||||||
|
#### Scenario: invocation rule rejects fabricated invocation ids
|
||||||
|
- **WHEN** a claim evidence binding references an invocation id absent from current-run `tool_invocation` rows
|
||||||
|
- **THEN** `gatekeeper_result.status` SHALL be `fail`
|
||||||
|
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
|
||||||
|
|
||||||
|
#### Scenario: invocation rule rejects tool name mismatch
|
||||||
|
- **WHEN** a claim evidence binding references an existing current-run invocation id
|
||||||
|
- **AND** binding `tool_name` does not match persisted invocation `tool_name`
|
||||||
|
- **THEN** `gatekeeper_result.status` SHALL be `fail`
|
||||||
|
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
|
||||||
|
|
||||||
|
#### Scenario: valid structured output passes initial gatekeeper rules
|
||||||
|
- **WHEN** Executor emits legal `executor_evidence_v2`
|
||||||
|
- **AND** each claim has evidence bindings pointing to current-run invocations with matching tool names and paths
|
||||||
|
- **THEN** `gatekeeper_result.status` SHALL be `pass`
|
||||||
|
- **AND** `gatekeeper_result.failed_rules` SHALL be empty
|
||||||
|
|
||||||
|
#### Scenario: gatekeeper reject bypasses Verifier
|
||||||
|
- **WHEN** normalized Gatekeeper status is REJECT
|
||||||
|
- **THEN** the Graph SHALL route directly to pre-verification Fallback
|
||||||
|
- **AND** Verifier SHALL NOT execute
|
||||||
|
|
||||||
|
#### Scenario: gatekeeper low confidence is bounded
|
||||||
|
- **WHEN** normalized Gatekeeper status is LOW_CONFID with at least one passed binding
|
||||||
|
- **THEN** verified input SHALL contain only passed bindings
|
||||||
|
- **AND** effective verdict SHALL NOT exceed LOW_CONFID
|
||||||
|
|
||||||
|
### Requirement: Verifier SHALL use verified claim-local evidence for derivability
|
||||||
|
Verifier SHALL judge structured claims only against Gatekeeper-verified claim-local evidence excerpts and their precise current-run references.
|
||||||
|
|
||||||
|
#### Scenario: Verified excerpt supports direct observation
|
||||||
|
- **WHEN** verdict ceiling is PASS
|
||||||
|
- **AND** a claim's verified evidence matched text directly contains the claim's concrete facts
|
||||||
|
- **THEN** Verifier MAY classify that claim as `direct_observation`
|
||||||
|
|
||||||
|
#### Scenario: Verified evidence is complete Verifier context
|
||||||
|
- **WHEN** verified claims and evidence are available
|
||||||
|
- **THEN** Verifier SHALL use them as its evidence context
|
||||||
|
- **AND** it SHALL NOT require or request a complete tool trace summary
|
||||||
|
- **AND** it SHALL NOT read raw Executor or unreferenced tool material
|
||||||
|
|
||||||
|
### Requirement: Gatekeeper severity SHALL constrain effective verdict
|
||||||
|
Runtime effective verdict calculation SHALL treat normalized Gatekeeper ceiling as a hard upper bound independent from Verifier model output.
|
||||||
|
|
||||||
|
#### Scenario: Reject severity bypasses Verifier
|
||||||
|
- **WHEN** `gatekeeper_result.severity=reject`
|
||||||
|
- **THEN** the Graph SHALL route to pre-verification Fallback without invoking Verifier
|
||||||
|
- **AND** it SHALL NOT fabricate an effective diagnostic verdict
|
||||||
|
|
||||||
|
#### Scenario: Low confidence severity prevents PASS
|
||||||
|
- **WHEN** `gatekeeper_result.severity=low_confid`
|
||||||
|
- **AND** the Verifier model returns `verdict=PASS`
|
||||||
|
- **THEN** deterministic effective-verdict calculation SHALL downgrade the result
|
||||||
|
- **AND** effective verdict SHALL be `LOW_CONFID`
|
||||||
|
|
||||||
|
#### Scenario: Gatekeeper audit includes severity
|
||||||
|
- **WHEN** Graph verifier evaluation is persisted
|
||||||
|
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation.gatekeeper_result` SHALL include available `status`, `severity`, `checked_bindings`, `failed_rules`, `warnings`, and `errors`
|
||||||
+56
@@ -0,0 +1,56 @@
|
|||||||
|
## ADDED Requirements
|
||||||
|
|
||||||
|
### Requirement: StateGraph Chat runs SHALL persist a compact orchestration trace
|
||||||
|
|
||||||
|
Each successful new StateGraph complex Chat run SHALL persist a non-empty compact orchestration summary derived from its bounded Graph events in `diagnosis_run.orchestration_trace`.
|
||||||
|
|
||||||
|
#### Scenario: Graph reaches Composer
|
||||||
|
|
||||||
|
- **WHEN** a complex Chat Graph terminates through Composer with a safe answer
|
||||||
|
- **THEN** the current DiagnosisRun SHALL store version, transitions, final node, termination reason, degraded flag, and evidence retry count
|
||||||
|
- **AND** the summary SHALL be derived from the current Run's actual orchestration events
|
||||||
|
|
||||||
|
#### Scenario: Graph reaches handled Fallback
|
||||||
|
|
||||||
|
- **WHEN** a complex Chat Graph terminates through deterministic Fallback with a safe answer
|
||||||
|
- **THEN** the current DiagnosisRun SHALL store a non-empty orchestration trace with `degraded=true`
|
||||||
|
- **AND** the Run status SHALL be SUCCESS
|
||||||
|
|
||||||
|
#### Scenario: Unhandled execution fails after events exist
|
||||||
|
|
||||||
|
- **WHEN** an unhandled failure occurs after one or more real Graph events are available
|
||||||
|
- **THEN** the service SHALL best-effort persist a partial orchestration summary for the current failed run
|
||||||
|
- **AND** it SHALL NOT add a node or transition that did not occur
|
||||||
|
|
||||||
|
#### Scenario: Orchestration trace content is inspected
|
||||||
|
|
||||||
|
- **WHEN** orchestration trace JSON is serialized
|
||||||
|
- **THEN** it SHALL NOT include Prompt text, model reasoning, raw tool output, raw Executor output, or Graph State snapshots
|
||||||
|
- **AND** it SHALL NOT contain data owned by another run
|
||||||
|
|
||||||
|
### Requirement: Trace API SHALL expose orchestration trace only on the run object
|
||||||
|
|
||||||
|
The Trace API SHALL parse the current DiagnosisRun orchestration JSON and expose it only as `run.orchestrationTrace`.
|
||||||
|
|
||||||
|
#### Scenario: Exact StateGraph run trace is queried
|
||||||
|
|
||||||
|
- **WHEN** a caller queries a successful new StateGraph Chat run
|
||||||
|
- **THEN** `run.orchestrationTrace` SHALL be a non-empty parsed JSON object
|
||||||
|
- **AND** the response top level and compatibility `session` projection SHALL NOT duplicate the field
|
||||||
|
- **AND** no raw orchestration trace field SHALL be added
|
||||||
|
|
||||||
|
#### Scenario: Historical or non-StateGraph run is queried
|
||||||
|
|
||||||
|
- **WHEN** the selected DiagnosisRun has null orchestration trace
|
||||||
|
- **THEN** `run.orchestrationTrace` MAY be null
|
||||||
|
- **AND** the service SHALL NOT synthesize historical events or read another run's trace
|
||||||
|
|
||||||
|
### Requirement: Orchestration trace migration SHALL be additive and nullable
|
||||||
|
|
||||||
|
The database migration SHALL add only one nullable JSON column named `orchestration_trace` to `diagnosis_run` for this change.
|
||||||
|
|
||||||
|
#### Scenario: Migration is applied
|
||||||
|
|
||||||
|
- **WHEN** Flyway applies the stage 3 migration
|
||||||
|
- **THEN** existing DiagnosisRun rows SHALL remain valid without backfill
|
||||||
|
- **AND** no other table or column SHALL be changed by the stage 3 schema migration
|
||||||
+43
@@ -0,0 +1,43 @@
|
|||||||
|
## 1. Run Orchestration Trace Contract
|
||||||
|
|
||||||
|
- [x] 1.1 Add V012 migration containing only nullable `diagnosis_run.orchestration_trace JSON` and document that historical rows are not backfilled.
|
||||||
|
- [x] 1.2 Add DiagnosisRun JSON field and `DiagnosisTraceResponse.RunTrace.orchestrationTrace` parsed-map field without adding top-level, session, summary, run-list, or raw duplicates.
|
||||||
|
- [x] 1.3 Update DiagnosisTraceService exact/latest Run mapping and tests for current-run parsed trace, historical null, invalid JSON fail-closed, and no cross-run/session projection.
|
||||||
|
- [x] 1.4 Add entity/repository or migration source checks proving stage 3 changes no schema object except the one nullable column.
|
||||||
|
|
||||||
|
## 2. Graph Runtime And Result Mapping
|
||||||
|
|
||||||
|
- [x] 2.1 Implement a dedicated complex Chat Graph runtime that assembles existing real actions, compiles the Graph once per request/runtime instance, and consumes `CompiledGraph.stream` while retaining only the last real state.
|
||||||
|
- [x] 2.2 Define minimal initial state with query-only diagnosis context, NORMAL mode, independent zero counters, and empty bounded events.
|
||||||
|
- [x] 2.3 Build RunnableConfig with `threadId=runId` and current `sessionId`/`runId` metadata; prove config identity at every real Agent invoker and Gatekeeper boundary.
|
||||||
|
- [x] 2.4 Implement explicit final/partial Graph result mapping for answer, statuses/verdicts, verified output/evidence, Gatekeeper audit, Composer audit, failure reason, and orchestration trace.
|
||||||
|
- [x] 2.5 Add runtime/mapper tests for Composer success, handled pre/post-verification Fallback, blank answer, no state, thrown execution with real partial events, and no fabricated event.
|
||||||
|
|
||||||
|
## 3. Agent Assembly And Verified-only Prompt
|
||||||
|
|
||||||
|
- [x] 3.1 Move or encapsulate the existing complex Planner/Executor/Verifier/Composer builders so Graph runtime reuses current Prompt, knowledge map, history, method tool, ToolCallback, skill, and AgentLoggingHook rules without a duplicate factory.
|
||||||
|
- [x] 3.2 Remove VerifierInputHook from the Graph Verifier while retaining AgentLoggingHook, and prove explicit Gatekeeper executes once per Executor round.
|
||||||
|
- [x] 3.3 Update `chat-verifier-prompt.md` to the verified-only input contract and remove raw Executor, full tool trace, Hook parse, and Gatekeeper re-execution language.
|
||||||
|
- [x] 3.4 Bump Prompt audit catalog/Verifier Prompt version and add static/behavior tests proving the runtime payload and Prompt declare the same allowed fields.
|
||||||
|
|
||||||
|
## 4. ChatService Production Cutover
|
||||||
|
|
||||||
|
- [x] 4.1 Replace `executeChatComplex` SequentialAgent/outer retry loop with one Graph runtime call while preserving session/run creation and the public ChatResult signature.
|
||||||
|
- [x] 4.2 Remove obsolete ChatService private Sequential workflow, score/feature-flag retry, ThreadLocal verifier parsing, private Composer routing, and dead imports without changing simple Chat behavior.
|
||||||
|
- [x] 4.3 Persist Graph-derived verifier evaluation with compatibility verdict/claim/fact/gatekeeper/composer/prompt-audit fields, verified-only evidence, and no fabricated verdict before Verifier completion.
|
||||||
|
- [x] 4.4 Persist SUCCESS only for a non-empty safe Graph answer with non-empty orchestration trace, then backfill current-run metrics and call Eval; persist FAILED for unhandled/no-answer/invariant failures with any real partial trace available.
|
||||||
|
- [x] 4.5 Ensure `SessionContextHolder` and retrieved-doc cleanup still run in finally, no complex Chat code reads `VerifierContextHolder`, and no production path creates SequentialAgent.
|
||||||
|
|
||||||
|
## 5. Stage 3 Focused Tests And Alignment
|
||||||
|
|
||||||
|
- [x] 5.1 Add production cutover integration tests for `/api/chat`-compatible ChatResult, `agent_flow=CHAT`, current Run ownership, Graph thread/metadata, answer/metrics/Eval persistence, and multi-run isolation.
|
||||||
|
- [x] 5.2 Add handled failure tests proving Executor invalid/failure and Gatekeeper REJECT do not create Verifier AgentStep, safe Fallback is SUCCESS/degraded, and unhandled/no-answer paths are FAILED.
|
||||||
|
- [x] 5.3 Update or replace stage-3-conflicting Sequential tests so no known test remains red; preserve Controller, Gatekeeper, Trace, Repository, Eval, no-evidence, REJECT, and Composer safety regressions.
|
||||||
|
- [x] 5.4 Confirm the first completed Trace slice and final production assembly against design/specs, including single-track cutover, no duplicate Agent factory, and no core TODO/placeholder.
|
||||||
|
|
||||||
|
## 6. Stage 3 Verification And Handoff
|
||||||
|
|
||||||
|
- [x] 6.1 Run Graph runtime/result mapper, ChatService cutover, DiagnosisTraceService, Controller, Repository, Gatekeeper, Composer, and Eval focused tests.
|
||||||
|
- [x] 6.2 Run Maven test compilation and any existing tests whose public contracts were touched; classify failures as OpenSpec gap, code deviation, or out-of-scope environment issue.
|
||||||
|
- [x] 6.3 Run OpenSpec strict validation, all main specs strict validation, `git diff --check`, source/reference isolation checks, and a schema whitelist check.
|
||||||
|
- [x] 6.4 Record that Maven live E2E, `logs/`, and `scripts/query_mysql.py` database verification were intentionally not run in stage 3 and remain reserved for stage 5.
|
||||||
+3
@@ -0,0 +1,3 @@
|
|||||||
|
ready_at: 2026-07-17
|
||||||
|
devflow: devflow/projects/2026-07-17-chat-diagnosis-stategraph-design-freeze
|
||||||
|
authorization: user-requested-per-stage-archive
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
committed_at: 2026-07-17
|
||||||
|
scope: iss-011-stage-0-design-freeze
|
||||||
|
validation: openspec-strict-pass
|
||||||
+226
@@ -0,0 +1,226 @@
|
|||||||
|
## Context
|
||||||
|
|
||||||
|
复杂 Chat 当前由 `ChatController -> ChatService.executeChatWithStrategy -> executeChatComplex` 驱动。`executeChatComplex` 同时负责 Run 生命周期、外层 LOW_CONFID round、`SequentialAgent` 构造、Verifier 解析、Composer、Fallback、持久化和评估;`VerifierInputHook` 又解析 Executor 输出、读取当前 Run 工具摘要、执行 `ExecutorGatekeeperService` 并通过 `VerifierContextHolder` 回传隐式状态。这些职责形成无法精确表达失败恢复位置、条件边和审计路径的隐式状态机。
|
||||||
|
|
||||||
|
阶段 0 不改变运行时,只冻结后续五个独立 change 必须遵守的架构契约。当前 change 的运行时影响为 L1;冻结目标最终涉及状态机语义、数据库与 Trace API,属于 L4。
|
||||||
|
|
||||||
|
| 区域 | 当前职责 | 冻结后的职责 |
|
||||||
|
|---|---|---|
|
||||||
|
| `ChatController` | 调用 ChatService | 请求/响应协议不变 |
|
||||||
|
| `ChatService` | Run 生命周期和细粒度隐式状态机 | 只维护 Run 生命周期并调用 Graph orchestrator |
|
||||||
|
| `SequentialAgent` | 固定跨 Agent 顺序 | 不再用于复杂 Chat 编排 |
|
||||||
|
| `VerifierInputHook` / `VerifierContextHolder` | 隐式 Gatekeeper、输入构造和跨层回传 | Gatekeeper/投影迁入显式 Node;旧隐式入口删除 |
|
||||||
|
| `ExecutorGatekeeperService` | 确定性 binding 校验 | 原规则复用,由显式 Gatekeeper Node 单次调用 |
|
||||||
|
| `DiagnosisRun` / `DiagnosisTraceService` | Run 与 self evaluation 聚合 | 增加 run-scoped orchestration trace |
|
||||||
|
|
||||||
|
## Goals / Non-Goals
|
||||||
|
|
||||||
|
**Goals:**
|
||||||
|
|
||||||
|
- 冻结最小 Graph State 的字段、所有权和更新策略。
|
||||||
|
- 冻结完整、有界且可终止的条件边与 retry 语义。
|
||||||
|
- 冻结 verified-input、Fallback、Run 状态与审计安全边界。
|
||||||
|
- 冻结旧测试替换范围和后续五个独立 change 的交付边界。
|
||||||
|
- 提供可被后续 OpenSpec Commit gate 机械对照的设计基线。
|
||||||
|
|
||||||
|
**Non-Goals:**
|
||||||
|
|
||||||
|
- 本 change 不实现或编译任何 Java/SQL/Prompt/配置变更。
|
||||||
|
- 不修改当前主运行时 specs 来宣称 StateGraph 已上线。
|
||||||
|
- 不运行 Maven E2E、日志或数据库验收。
|
||||||
|
- 不引入 SupervisorAgent、持久 Checkpointer、HITL、并行执行或 AIOps 公共 Graph。
|
||||||
|
- 不保留 Sequential/StateGraph 双轨开关。
|
||||||
|
|
||||||
|
## Decisions
|
||||||
|
|
||||||
|
### 1. StateGraph 只拥有跨节点控制状态
|
||||||
|
|
||||||
|
StateGraph 负责顺序、条件边、重试、终止和降级;现有 ReactAgent 继续完成语义任务;Gatekeeper、verified-input builder、evidence retry prepare 和固定 Fallback 使用确定性 Java Node。
|
||||||
|
|
||||||
|
替代方案:
|
||||||
|
|
||||||
|
- 继续在 ChatService 外层叠加 if/for:无法精确恢复失败节点和审计条件边,拒绝。
|
||||||
|
- 使用 SupervisorAgent:固定诊断 Pipeline 不需要动态选择专科 Agent,拒绝。
|
||||||
|
- 直接把父 State 交给 `ReactAgent.asNode(...)`:存在 messages/outputKey/私有状态泄漏风险,首版拒绝。
|
||||||
|
|
||||||
|
### 2. 最小 Graph State
|
||||||
|
|
||||||
|
默认 Replace;只有有界 `orchestration_events` 使用 Append。
|
||||||
|
|
||||||
|
| 字段 | 所有者/用途 | 策略 |
|
||||||
|
|---|---|---|
|
||||||
|
| `diagnosis_context` | query、history、sessionId、runId | Replace |
|
||||||
|
| `planner_plan` | Planner 结构化计划 | Replace |
|
||||||
|
| `planner_status` | COMPLETED / INVALID_OUTPUT / RETRYABLE_FAILED / NON_RETRYABLE_FAILED | Replace |
|
||||||
|
| `planner_retry_count` | 当前 Planner 阶段技术重试次数 | Replace |
|
||||||
|
| `planner_mode` | NORMAL / EVIDENCE_GAP_ONLY | Replace |
|
||||||
|
| `executor_output` | 完整 `executor_evidence_v2` | Replace |
|
||||||
|
| `executor_status` | COMPLETED / INVALID_OUTPUT / TOOL_BLOCKED / FAILED | Replace |
|
||||||
|
| `gatekeeper_result` | 原始确定性校验结果 | Replace |
|
||||||
|
| `gatekeeper_status` | PASS / LOW_CONFID / REJECT | Replace |
|
||||||
|
| `verified_executor_output` | 只保留通过 binding 的 claim 投影 | Replace |
|
||||||
|
| `verified_evidence` | 从通过 binding 的 `matched_text` 构造 | Replace |
|
||||||
|
| `verified_binding_count` | 当前可信 binding 数 | Replace |
|
||||||
|
| `verifier_verdict_ceiling` | PASS 或 LOW_CONFID | Replace |
|
||||||
|
| `verifier_output` | Verifier 结构化结果 | Replace |
|
||||||
|
| `verifier_status` | COMPLETED / INVALID_OUTPUT / RETRYABLE_FAILED / NON_RETRYABLE_FAILED | Replace |
|
||||||
|
| `verifier_retry_count` | 固定 verified input 技术重试次数 | Replace |
|
||||||
|
| `verifier_model_verdict` | 模型 PASS / LOW_CONFID / REJECT | Replace |
|
||||||
|
| `effective_verdict` | 应用 Gatekeeper ceiling 后的 verdict | Replace |
|
||||||
|
| `composer_output` | Composer 结构化输出 | Replace |
|
||||||
|
| `composer_status` | COMPLETED / INVALID_OUTPUT / RETRYABLE_FAILED / NON_RETRYABLE_FAILED | Replace |
|
||||||
|
| `composer_retry_count` | 固定安全输入技术重试次数 | Replace |
|
||||||
|
| `evidence_retry_count` | 诊断补证据轮次 | Replace |
|
||||||
|
| `retry_context` | evidence gaps、prior verified data、completed queries、约束 | Replace |
|
||||||
|
| `orchestration_events` | 每个 node attempt 的有界终态事件 | Append |
|
||||||
|
| `final_answer` | 最终安全用户表达 | Replace |
|
||||||
|
| `failure_reason` | 确定性失败 reason code | Replace |
|
||||||
|
|
||||||
|
不保存 Prompt、模型思考、完整工具原文、完整 Graph State 快照或其他 Run 数据。
|
||||||
|
|
||||||
|
### 3. 完整路由矩阵
|
||||||
|
|
||||||
|
| From | Outcome / Guard | To | 计数与约束 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| START | always | Planner | Planner 阶段 retry count=0 |
|
||||||
|
| Planner | COMPLETED | Executor | 不增加 retry |
|
||||||
|
| Planner | INVALID_OUTPUT / RETRYABLE_FAILED 且 count=0 | Planner | 同业务输入重试一次,count=1 |
|
||||||
|
| Planner | NON_RETRYABLE_FAILED 或技术重试耗尽 | Fallback | 不执行后续 Agent |
|
||||||
|
| Executor | COMPLETED | Gatekeeper | 合法 no-evidence 也属于 COMPLETED |
|
||||||
|
| Executor | INVALID_OUTPUT / TOOL_BLOCKED / FAILED | Fallback | Executor 永不重试 |
|
||||||
|
| Gatekeeper | PASS | Verified Input Builder | ceiling=PASS |
|
||||||
|
| Gatekeeper | LOW_CONFID 且 verified binding > 0 | Verified Input Builder | ceiling=LOW_CONFID |
|
||||||
|
| Gatekeeper | REJECT、未知状态或 LOW_CONFID 且 binding=0 | Fallback | 不执行 Verifier |
|
||||||
|
| Verified Input Builder | completed | Verifier | 只投影通过 binding |
|
||||||
|
| Verifier | INVALID_OUTPUT / RETRYABLE_FAILED 且 count=0 | Verifier | 相同 verified input 重试一次 |
|
||||||
|
| Verifier | NON_RETRYABLE_FAILED 或技术重试耗尽 | Fallback | 不输出 Executor claim |
|
||||||
|
| Verifier | PASS / REJECT | Composer | 使用 effective verdict |
|
||||||
|
| Verifier | LOW_CONFID,非 ceiling 导致,存在有效 gaps,evidence count=0 | Prepare Evidence Retry | evidence count 增加一次 |
|
||||||
|
| Verifier | LOW_CONFID 且任一补查条件不满足 | Composer | 不补查 |
|
||||||
|
| Prepare Evidence Retry | completed | Planner | mode=EVIDENCE_GAP_ONLY;新 Planner 阶段 retry count 重置 |
|
||||||
|
| Composer | COMPLETED | END | 返回 Composer 安全表达 |
|
||||||
|
| Composer | INVALID_OUTPUT / RETRYABLE_FAILED 且 count=0 | Composer | 相同 allowed material 重试一次 |
|
||||||
|
| Composer | NON_RETRYABLE_FAILED 或技术重试耗尽 | Fallback | 可使用 Verifier 已允许材料 |
|
||||||
|
| Fallback | deterministic answer produced | END | degraded=true |
|
||||||
|
| 未处理异常/持久化失败/无法生成安全答案 | failure | outer failure handling | Run=FAILED,best-effort 保存已有 events |
|
||||||
|
|
||||||
|
Graph compile 必须设置 recursion limit 作为最终保险;业务循环仍由独立计数显式限制。
|
||||||
|
|
||||||
|
### 4. 四类计数互相独立
|
||||||
|
|
||||||
|
- `planner_retry_count`:每次进入 Planner 阶段最多一次;evidence retry 重新进入时重置。
|
||||||
|
- `verifier_retry_count`:同一 verified input 最多一次。
|
||||||
|
- `composer_retry_count`:同一 allowed material 最多一次。
|
||||||
|
- `evidence_retry_count`:整个 Run 最多一次。
|
||||||
|
|
||||||
|
技术重试不增加 evidence retry,也不得重新执行前序节点或工具。
|
||||||
|
|
||||||
|
### 5. Executor 状态按输出契约定义
|
||||||
|
|
||||||
|
合法 `executor_evidence_v2`(包括 no-evidence、空查询结果或工具限制说明)均为 COMPLETED。只有工具层明确禁止且没有合法结构时为 TOOL_BLOCKED;空/非法结构为 INVALID_OUTPUT;其他未形成合法结构的异常为 FAILED。只有 COMPLETED 进入 Gatekeeper,其余直接 Fallback 且不重试。
|
||||||
|
|
||||||
|
补证据轮只执行增量查询,但输出包含应保留 prior verified claims 的完整快照;Java 不做 claim 文本语义合并,完整快照重新经过 Gatekeeper。
|
||||||
|
|
||||||
|
### 6. Gatekeeper 与 Verifier 输入边界
|
||||||
|
|
||||||
|
Gatekeeper 保留原始结果并 fail-closed 标准化为 PASS / LOW_CONFID / REJECT;未知、缺失或无法识别结果映射为 REJECT。
|
||||||
|
|
||||||
|
PASS 与可继续 LOW_CONFID 都经过 Verified Input Builder。Builder 复用 `checked_bindings`,仅从通过项投影 `claim_id`、`source_invocation_id`、`tool_name`、`raw_path`、`matched_text`;不得传递完整 `tool_trace_summary`、失败 binding 或 Executor 自由文本。LOW_CONFID ceiling 不得被模型 PASS 提升。
|
||||||
|
|
||||||
|
### 7. 两类安全 Fallback
|
||||||
|
|
||||||
|
**前置验证失败:** Planner/Executor/Verifier 无法形成可信材料,或 Gatekeeper REJECT/零可信 binding。固定表达只包含校验状态、工具执行概况、诊断限制和人工查看 Trace 建议,不含任何 Executor claim。
|
||||||
|
|
||||||
|
**Composer 后置降级:** 已有 Verifier 允许材料,但 Composer 技术重试耗尽。确定性模板可以使用 allowed claims/missing info/recommended actions,仍不得读取 raw output。
|
||||||
|
|
||||||
|
成功生成安全答案时 Run=SUCCESS,并用 verdict 或 `orchestrationTrace.degraded=true` 表达质量;只有未处理异常、持久化失败或无法生成安全答案时为 FAILED,不新增 DEGRADED 状态。
|
||||||
|
|
||||||
|
### 8. Run-scoped 编排审计
|
||||||
|
|
||||||
|
每个 Node attempt 最多追加一个 `{node,outcome,reason_code,attempt}` 终态 event。按事件顺序压缩为 version、transitions、final_node、termination_reason、degraded、evidence_retry_count。
|
||||||
|
|
||||||
|
只持久化到当前 `diagnosis_run.orchestration_trace`;`RunnableConfig.threadId=runId`。Trace API 只在 `run.orchestrationTrace` 返回解析对象,不复制到顶层/session/raw 字段。它不替代 self evaluation、AgentStep、ToolInvocation 或 Graph checkpoint。
|
||||||
|
|
||||||
|
### 9. 测试替换边界
|
||||||
|
|
||||||
|
| 测试 | 处理 |
|
||||||
|
|---|---|
|
||||||
|
| `ChatServiceSequentialAgentTest` | 用 Graph 路由、Node 契约和 Chat 集成测试替换;删除固定顺序断言 |
|
||||||
|
| `VerifierInputHookTest` | 删除或改写为显式 Gatekeeper/Input Builder 契约测试 |
|
||||||
|
| `ExecutorGatekeeperServiceTest` | 保留并扩展 checked binding 与未知状态 fail-closed |
|
||||||
|
| `ChatControllerTest` | 保留,证明 `/api/chat` 兼容 |
|
||||||
|
| `DiagnosisTraceServiceTest` / Repository tests | 保留并扩展 run-scoped orchestration trace |
|
||||||
|
| Eval/Composer 安全测试 | 保留,证明 allowed-material、no-evidence 和 REJECT 不退化 |
|
||||||
|
|
||||||
|
阶段 0 无运行时行为,不新增单元测试。阶段 1–4 按风险建立上述测试;阶段 5 统一执行 Maven E2E、日志和数据库验收。
|
||||||
|
|
||||||
|
### 10. 六个独立 OpenSpec change
|
||||||
|
|
||||||
|
| 阶段 | Change | 唯一交付边界 |
|
||||||
|
|---|---|---|
|
||||||
|
| 0 | `chat-diagnosis-stategraph-design-freeze` | 冻结本设计,不改运行时 |
|
||||||
|
| 1 | `chat-diagnosis-stategraph-routing-skeleton` | Graph State、拓扑、Fake Node 路由与 trace builder |
|
||||||
|
| 2 | `chat-diagnosis-stategraph-real-nodes` | Agent Adapter、Gatekeeper、verified input、retry prepare、Fallback |
|
||||||
|
| 3 | `chat-diagnosis-stategraph-chatservice-cutover` | ChatService 生产入口、DB migration、Run/Trace API |
|
||||||
|
| 4 | `chat-diagnosis-stategraph-test-suite` | 新测试体系与旧测试替换 |
|
||||||
|
| 5 | `chat-diagnosis-stategraph-cleanup-docs` | 旧编排清理、最终回归、Maven E2E、日志/DB、文档 |
|
||||||
|
|
||||||
|
后续 change 必须先读取阶段 0 archive;若发现设计不准,必须在当阶段 OpenSpec 中显式记录和解决。
|
||||||
|
|
||||||
|
## Interface Impact
|
||||||
|
|
||||||
|
### Classification
|
||||||
|
|
||||||
|
- 当前阶段 0 change:L1。只交付设计、术语和决策档案,不改变运行时。
|
||||||
|
- 最终冻结目标:L4。状态机语义和 Run 终态判断变化,Trace API 与数据库契约扩展,旧内部调用路径将被删除。
|
||||||
|
|
||||||
|
### Changed contracts and consumers
|
||||||
|
|
||||||
|
| 契约 | 目标变化 | 消费者 |
|
||||||
|
|---|---|---|
|
||||||
|
| `/api/chat` | 请求/响应结构保持不变,内部路径改由 Graph 驱动 | ChatController、前端/demo 客户端 |
|
||||||
|
| Agent 输出协议 | `executor_evidence_v2`、Verifier、Composer 输出保持不变 | Agent adapters、Eval |
|
||||||
|
| Verifier 输入 | 改为 `verified_executor_output + verified_evidence`,不再提供完整 tool trace | Verifier adapter、Prompt binding |
|
||||||
|
| Run 状态 | 安全 Fallback 为 SUCCESS + degraded;只有无法安全响应的未处理失败为 FAILED | Trace、Feedback、Eval、运维审计 |
|
||||||
|
| Trace API | 仅 `run.orchestrationTrace` 新增解析对象 | Trace DTO/Service、demo 脚本、Trace UI |
|
||||||
|
| 数据库 | 仅新增 nullable JSON `diagnosis_run.orchestration_trace` | Flyway、DiagnosisRun repository |
|
||||||
|
| 内部编排 | 删除复杂 Chat Sequential、外层 round、Gatekeeper-in-Hook/ThreadLocal 入口 | ChatService、Hook、测试 |
|
||||||
|
|
||||||
|
### Compatibility, migration, and rollback
|
||||||
|
|
||||||
|
- 兼容消费者不需要修改 `/api/chat` 调用;Trace 消费者按加法字段处理。
|
||||||
|
- 新 StateGraph Chat Run 必须写非空 orchestration trace;历史 Run 不回填、不增加兼容读取。
|
||||||
|
- 数据库迁移先加 nullable 列,再切换代码;nullable 只服务于安全部署,不降低新 Run 应用契约。
|
||||||
|
- 每阶段可 Git revert;阶段 3 后回滚代码时允许保留无害 nullable 列。
|
||||||
|
- 不以 feature flag 或配置恢复长期双轨;若切换失败,回滚完整阶段提交。
|
||||||
|
|
||||||
|
### Interface acceptance
|
||||||
|
|
||||||
|
- Controller 契约测试证明 `/api/chat` 不变。
|
||||||
|
- Node/集成测试证明 Verifier 输入与安全 Fallback 边界。
|
||||||
|
- Repository/Trace 测试证明 run ownership 和唯一 API 投影。
|
||||||
|
- 阶段 5 Maven E2E、`logs/` 和数据库查询证明跨层最终行为。
|
||||||
|
|
||||||
|
## Risks / Trade-offs
|
||||||
|
|
||||||
|
- [阶段性 spec 与运行时差异] 阶段 0 只新增设计基线 capability,不修改运行时 specs。
|
||||||
|
- [六个 change 漂移] 每个 Discover/Commit gate 强制引用阶段 0 archive。
|
||||||
|
- [Graph API 漂移] 后续只使用本地 `1.1.2.0` JAR 已验证签名,并由阶段 1 编译测试锁定。
|
||||||
|
- [Gatekeeper 双执行] 阶段 2 迁移显式 Node,阶段 5 删除旧 Hook/ThreadLocal,不形成长期双轨。
|
||||||
|
- [安全降级泄漏] Fallback 输入类型分离并用 Node/集成测试证明。
|
||||||
|
- [Trace 无限增长或泄密] event 每 attempt 一条,受业务次数与 recursion limit 限制,字段白名单禁原文。
|
||||||
|
|
||||||
|
## Migration Plan
|
||||||
|
|
||||||
|
1. 阶段 0 archive 本设计基线并提交。
|
||||||
|
2. 阶段 1 引入未接生产入口的 Graph 骨架和路由测试。
|
||||||
|
3. 阶段 2 接入真实节点但不切换 ChatService。
|
||||||
|
4. 阶段 3 切换 ChatService,增加 nullable JSON 列和 run Trace 投影。
|
||||||
|
5. 阶段 4 用新测试体系覆盖路由和安全契约。
|
||||||
|
6. 阶段 5 删除旧编排并执行最终 Maven E2E、`logs/` 与数据库验收。
|
||||||
|
|
||||||
|
回滚按阶段 Git revert。阶段 3 后回滚运行时代码时允许保留 nullable 列;不通过配置重新启用长期双轨。历史 Run 不回填。
|
||||||
|
|
||||||
|
## Open Questions
|
||||||
|
|
||||||
|
无。若后续发现本设计与锁定 API 或安全契约冲突,必须在当阶段 sm-flow 中分类并按门禁处理。
|
||||||
+77
@@ -0,0 +1,77 @@
|
|||||||
|
# Chat Diagnosis StateGraph Design Freeze
|
||||||
|
|
||||||
|
## Why
|
||||||
|
|
||||||
|
ISS-011 将复杂 Chat 诊断从 `SequentialAgent + ChatService` 隐式状态机迁移到 Spring AI Alibaba `StateGraph`。在进入任何运行时代码实现前,需要先把跨节点状态、全部条件边、有限重试、安全降级、审计边界和旧测试替换范围冻结为单一、可引用的设计契约,避免后续五个独立 sm-flow 对同一行为产生不同解释。
|
||||||
|
|
||||||
|
## What Changes
|
||||||
|
|
||||||
|
本 change 只交付阶段 0 的设计冻结:
|
||||||
|
|
||||||
|
- 冻结最小 Graph State 及 Replace/Append 更新策略。
|
||||||
|
- 冻结 Planner、Executor、Gatekeeper、Verifier、Composer、Evidence Retry、Fallback 的全部条件边和终止路径。
|
||||||
|
- 冻结 Planner/Verifier/Composer 技术重试与 evidence retry 的独立计数和最大次数。
|
||||||
|
- 冻结 Gatekeeper 后的 Verifier 白名单输入、安全 Fallback 和 Run 状态语义。
|
||||||
|
- 冻结 `orchestration_events -> diagnosis_run.orchestration_trace -> run.orchestrationTrace` 的审计边界。
|
||||||
|
- 冻结旧 Sequential 测试的替换范围和必须保留的安全回归边界。
|
||||||
|
- 形成供阶段 1–5 各自独立 OpenSpec change 引用的长期决策记录。
|
||||||
|
|
||||||
|
本 change 不修改 Java、SQL、Prompt、配置或运行时行为。
|
||||||
|
|
||||||
|
## Capabilities
|
||||||
|
|
||||||
|
### New Capabilities
|
||||||
|
|
||||||
|
- `chat-diagnosis-stategraph-design-freeze`:为阶段 1–5 提供版本化、可审计的 StateGraph 设计基线,覆盖状态、路由、重试、安全降级、审计和测试迁移边界。
|
||||||
|
|
||||||
|
### Modified Capabilities
|
||||||
|
|
||||||
|
- 无。当前运行时 capabilities 只在对应实现阶段修改,避免阶段 0 归档后主规格提前宣称代码已切换。
|
||||||
|
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
### In Scope
|
||||||
|
|
||||||
|
- `mvp/issues/active/ISS-011-chat-diagnosis-stategraph-orchestration.md` 中阶段 0 的设计结论。
|
||||||
|
- `devflow/glossary/CONTEXT.md` 中 Diagnosis Orchestration Trace 和 Verifier verified-evidence 边界。
|
||||||
|
- 阶段 0 的 OpenSpec 设计、设计验收 spec、可执行文档任务和 devflow/ADR 档案。
|
||||||
|
- 对最终目标的 L4 接口影响、迁移、回滚、兼容性与消费者边界做设计级记录。
|
||||||
|
|
||||||
|
### Out of Scope
|
||||||
|
|
||||||
|
- 创建 StateGraph、Node、Adapter、路由或测试代码。
|
||||||
|
- 修改 ChatService、Hook、Gatekeeper、Agent、Prompt 或工具。
|
||||||
|
- 数据库迁移和 Trace DTO/API 修改。
|
||||||
|
- 运行 Maven E2E、日志核验或数据库核验。
|
||||||
|
- 把阶段 1–5 的实现任务放入本 change。
|
||||||
|
- SupervisorAgent、持久 Checkpointer、HITL、并行 Agent/工具或 AIOps 公共 Graph。
|
||||||
|
|
||||||
|
## Context Constraints
|
||||||
|
|
||||||
|
- `runId` 是一次 Diagnosis Run 与 Trace 的所有权边界;编排摘要不得写入 session 级数据。
|
||||||
|
- Gatekeeper 引用真实性检查和 `checked_bindings` 是 verified evidence 的唯一事实来源,不得在 Verifier Input Builder 重复实现另一套验真规则。
|
||||||
|
- Composer 只能消费 Verifier 允许材料;前置验证失败的固定 Fallback 不得泄漏 Executor claim。
|
||||||
|
- `/api/chat`、`executor_evidence_v2`、Verifier 输出协议和 Composer 输出协议保持兼容。
|
||||||
|
- 最终 Trace API 只在 `run.orchestrationTrace` 增加解析对象;数据库只新增 nullable JSON `diagnosis_run.orchestration_trace`。
|
||||||
|
- 不保留长期 Sequential/StateGraph 双轨或运行时切换开关。
|
||||||
|
- 锁定依赖版本 `spring-ai-alibaba-graph-core:1.1.2.0` 已通过本地 JAR API 核验,后续实现不得使用未经锁定版本验证的示例签名。
|
||||||
|
|
||||||
|
## Acceptance
|
||||||
|
|
||||||
|
- 最小 Graph State 中每个字段都有所有权、用途和更新策略。
|
||||||
|
- 每个节点状态都有唯一条件边;所有循环均有显式上限和终止路径。
|
||||||
|
- Planner、Verifier、Composer 技术重试各自最多一次,且与最多一次 evidence retry 分开计数。
|
||||||
|
- Executor 不重试;Gatekeeper REJECT 与零可信 binding 直接进入不泄漏 claim 的固定 Fallback。
|
||||||
|
- PASS 与可继续的 LOW_CONFID 都经过 verified-input builder;Verifier 不接收完整 `tool_trace_summary`。
|
||||||
|
- `orchestration_trace`、`self_evaluation`、`agent_step`、`tool_invocation` 和 Graph checkpoint 的职责不重叠。
|
||||||
|
- 旧测试替换清单明确区分应替换的实现测试与必须保留/扩展的契约测试。
|
||||||
|
- 不存在“待讨论决策”、未定义循环或没有终止路径的分支。
|
||||||
|
- 本阶段无运行时变化,因此不新增或运行单元测试;使用 OpenSpec strict validation、结构检查与文档一致性检查验收。
|
||||||
|
|
||||||
|
## Risks
|
||||||
|
|
||||||
|
- 阶段 0 只冻结设计,主分支运行时仍是 Sequential;后续阶段不得把设计文档状态误认为已实现状态。
|
||||||
|
- 六个独立 changes 可能发生契约漂移;每个后续 Discover 必须读取本 change 的 archive 和 devflow 决策,并在 Commit gate 对照冻结契约。
|
||||||
|
- 现有主 OpenSpec 仍描述旧 `tool_trace_summary` 和外层 LOW_CONFID round;只有对应运行时切换阶段才能修改并归档这些运行时 specs。
|
||||||
|
- 最终设计属于 L4 影响;阶段 0 必须记录迁移/回滚,但不能以文档完成替代后续代码和最终 E2E 证据。
|
||||||
+121
@@ -0,0 +1,121 @@
|
|||||||
|
## ADDED Requirements
|
||||||
|
|
||||||
|
### Requirement: StateGraph design baseline SHALL be versioned and authoritative
|
||||||
|
|
||||||
|
The project SHALL maintain an archived ISS-011 design baseline that defines Graph State, routing, retry limits, fallback classes, audit boundaries, migration stages, and test replacement scope without claiming runtime cutover is complete.
|
||||||
|
|
||||||
|
#### Scenario: Later stage starts implementation
|
||||||
|
|
||||||
|
- **WHEN** an ISS-011 stage 1–5 OpenSpec change is proposed
|
||||||
|
- **THEN** its context and design SHALL reference the archived stage 0 baseline
|
||||||
|
- **AND** any intentional deviation SHALL be resolved through that stage's OpenSpec before implementation
|
||||||
|
|
||||||
|
#### Scenario: Stage zero is accepted
|
||||||
|
|
||||||
|
- **WHEN** the design-freeze change is archived
|
||||||
|
- **THEN** no Java, SQL, Prompt, configuration, or runtime behavior SHALL have been changed by this change
|
||||||
|
- **AND** runtime specs SHALL NOT claim StateGraph cutover is already implemented
|
||||||
|
|
||||||
|
### Requirement: Graph State design SHALL use explicit bounded control state
|
||||||
|
|
||||||
|
Every cross-node field SHALL have one purpose, owner, and update strategy. Fields SHALL use Replace semantics except bounded `orchestration_events`, which SHALL use Append.
|
||||||
|
|
||||||
|
#### Scenario: Node state is designed
|
||||||
|
|
||||||
|
- **WHEN** an Agent or Java Node reads or writes parent Graph State
|
||||||
|
- **THEN** its allowed input projection and output fields SHALL be explicit
|
||||||
|
- **AND** Prompt text, model reasoning, complete tool output, and complete State snapshots SHALL NOT be control state
|
||||||
|
|
||||||
|
#### Scenario: Orchestration event is designed
|
||||||
|
|
||||||
|
- **WHEN** a node attempt reaches a handled terminal outcome
|
||||||
|
- **THEN** at most one event SHALL be appended for that attempt
|
||||||
|
- **AND** it SHALL contain only stable node, outcome, reason code, and attempt data
|
||||||
|
|
||||||
|
### Requirement: Routing design SHALL be complete and terminating
|
||||||
|
|
||||||
|
Every Planner, Executor, Gatekeeper, Verifier, Composer, evidence-retry, and Fallback outcome SHALL map to one next node or terminal result. Every loop SHALL have an explicit business limit and the Graph SHALL use a recursion limit.
|
||||||
|
|
||||||
|
#### Scenario: Technical retry is eligible
|
||||||
|
|
||||||
|
- **WHEN** Planner, Verifier, or Composer first returns INVALID_OUTPUT or RETRYABLE_FAILED with the same allowed input
|
||||||
|
- **THEN** only that node SHALL be retried once
|
||||||
|
- **AND** no preceding Agent, Gatekeeper, or tool SHALL be rerun
|
||||||
|
|
||||||
|
#### Scenario: Technical retry is exhausted
|
||||||
|
|
||||||
|
- **WHEN** Planner, Verifier, or Composer returns NON_RETRYABLE_FAILED or exhausts its retry
|
||||||
|
- **THEN** the route SHALL terminate through the defined safe Fallback
|
||||||
|
- **AND** no unbounded loop SHALL remain
|
||||||
|
|
||||||
|
#### Scenario: Executor cannot produce a legal contract
|
||||||
|
|
||||||
|
- **WHEN** Executor returns INVALID_OUTPUT, TOOL_BLOCKED, or FAILED without legal `executor_evidence_v2`
|
||||||
|
- **THEN** the route SHALL go directly to pre-verification Fallback
|
||||||
|
- **AND** Executor SHALL NOT be retried
|
||||||
|
|
||||||
|
#### Scenario: Evidence retry is eligible
|
||||||
|
|
||||||
|
- **WHEN** effective verdict is LOW_CONFID, it was not caused by Gatekeeper ceiling, valid evidence gaps exist, and the Run has not retried evidence
|
||||||
|
- **THEN** one EVIDENCE_GAP_ONLY retry SHALL return to a new Planner stage
|
||||||
|
- **AND** technical retry counters SHALL remain independent from evidence retry count
|
||||||
|
|
||||||
|
### Requirement: Verified evidence and fallback boundaries SHALL fail closed
|
||||||
|
|
||||||
|
PASS and eligible LOW_CONFID Gatekeeper outcomes SHALL pass through a verified-input builder. Verifier SHALL receive only claims and evidence projected from passed checked bindings, not complete `tool_trace_summary` or unverified Executor text.
|
||||||
|
|
||||||
|
#### Scenario: Gatekeeper result is unsafe or unknown
|
||||||
|
|
||||||
|
- **WHEN** Gatekeeper returns REJECT, unknown, or LOW_CONFID with zero verified bindings
|
||||||
|
- **THEN** the route SHALL enter pre-verification Fallback without Verifier
|
||||||
|
- **AND** the final answer SHALL NOT contain any Executor claim
|
||||||
|
|
||||||
|
#### Scenario: Gatekeeper permits verification
|
||||||
|
|
||||||
|
- **WHEN** Gatekeeper returns PASS or eligible LOW_CONFID
|
||||||
|
- **THEN** the builder SHALL project evidence only from passed bindings
|
||||||
|
- **AND** a LOW_CONFID ceiling SHALL NOT be upgraded to PASS
|
||||||
|
|
||||||
|
#### Scenario: Composer fails after verification
|
||||||
|
|
||||||
|
- **WHEN** Composer exhausts retry after receiving Verifier-allowed material
|
||||||
|
- **THEN** deterministic fallback MAY use only that allowed material
|
||||||
|
- **AND** it SHALL NOT read raw Executor or tool output
|
||||||
|
|
||||||
|
### Requirement: Orchestration audit design SHALL preserve Run ownership
|
||||||
|
|
||||||
|
Orchestration audit SHALL remain separate from self evaluation, AgentStep, ToolInvocation, and Graph checkpoint data. A compact summary SHALL be derived from bounded events and persisted only to the current Run.
|
||||||
|
|
||||||
|
#### Scenario: A safe response is produced
|
||||||
|
|
||||||
|
- **WHEN** Graph reaches Composer or handled Fallback and produces a safe answer
|
||||||
|
- **THEN** Run status SHALL be SUCCESS
|
||||||
|
- **AND** degradation SHALL be represented by verdict or `orchestrationTrace.degraded`
|
||||||
|
|
||||||
|
#### Scenario: Trace is exposed
|
||||||
|
|
||||||
|
- **WHEN** an exact new StateGraph Chat Run is queried
|
||||||
|
- **THEN** parsed audit SHALL appear only at `run.orchestrationTrace`
|
||||||
|
- **AND** it SHALL NOT be duplicated at top level, session projection, or raw field
|
||||||
|
|
||||||
|
#### Scenario: Audit ownership is evaluated
|
||||||
|
|
||||||
|
- **WHEN** events or summaries are persisted
|
||||||
|
- **THEN** `runId` SHALL be the ownership and Graph thread boundary
|
||||||
|
- **AND** no data SHALL include Prompt, reasoning, complete tool output, or another Run
|
||||||
|
|
||||||
|
### Requirement: Test migration design SHALL preserve safety behavior
|
||||||
|
|
||||||
|
The baseline SHALL identify Sequential/Hook implementation tests to replace and public/security contract tests to retain or extend. Fixed Agent call order SHALL NOT remain a correctness criterion.
|
||||||
|
|
||||||
|
#### Scenario: Old tests are replaced
|
||||||
|
|
||||||
|
- **WHEN** StateGraph tests become authoritative
|
||||||
|
- **THEN** `ChatServiceSequentialAgentTest` SHALL be replaced by route, node-contract, and Chat integration coverage
|
||||||
|
- **AND** `VerifierInputHookTest` SHALL be removed or rewritten for explicit nodes
|
||||||
|
|
||||||
|
#### Scenario: Safety tests are retained
|
||||||
|
|
||||||
|
- **WHEN** the new suite is assembled
|
||||||
|
- **THEN** Gatekeeper, Controller, Run/Trace, Repository, Composer, no-evidence, REJECT, and Eval safety contracts SHALL remain covered
|
||||||
|
- **AND** fixed-order-only assertions SHALL be removed
|
||||||
@@ -0,0 +1,15 @@
|
|||||||
|
## 1. Freeze Source Documents
|
||||||
|
|
||||||
|
- [x] 1.1 Compare ISS-011 sections 3–14 against the committed state, routing, retry, fallback, audit, and test-migration design; resolve every drift without editing runtime code.
|
||||||
|
- [x] 1.2 Verify `devflow/glossary/CONTEXT.md` defines Diagnosis Orchestration Trace and the Verifier verified-evidence boundary without implementation-detail leakage.
|
||||||
|
|
||||||
|
## 2. Record Durable Architecture Decisions
|
||||||
|
|
||||||
|
- [x] 2.1 Create a project ADR that records the StateGraph control boundary, explicit Gatekeeper Node, run-scoped orchestration trace, rejected alternatives, compatibility, migration, and rollback decisions.
|
||||||
|
- [x] 2.2 Record the six independent sm-flow handoff boundaries and require stages 1–5 to reference the archived stage 0 baseline.
|
||||||
|
|
||||||
|
## 3. Validate The Design Freeze
|
||||||
|
|
||||||
|
- [x] 3.1 Run strict OpenSpec validation and cross-artifact checks for proposal, design, specs, and tasks; resolve every validation or alignment gap.
|
||||||
|
- [x] 3.2 Verify the stage 0 diff contains no Java, SQL, Prompt, configuration, or runtime behavior changes.
|
||||||
|
- [x] 3.3 Record that unit tests and Maven E2E are intentionally not run because stage 0 changes only design artifacts; reserve Maven E2E, logs, and database evidence for stage 5.
|
||||||
+3
@@ -0,0 +1,3 @@
|
|||||||
|
ready_at: 2026-07-17
|
||||||
|
devflow: devflow/projects/2026-07-17-chat-diagnosis-stategraph-real-nodes
|
||||||
|
authorization: user-requested-per-stage-archive
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
Committed by sm-flow after stage 2 Commit checkpoint validation on 2026-07-17.
|
||||||
@@ -0,0 +1,106 @@
|
|||||||
|
## Context
|
||||||
|
|
||||||
|
阶段 1 已交付未接生产入口的 Diagnosis StateGraph 状态、拓扑、路由和 trace builder。当前复杂 Chat 仍由 `ChatService`、`SequentialAgent`、`VerifierInputHook` 和 `VerifierContextHolder` 编排;Executor 解析在 Hook 内,Verifier/Composer 解析与安全渲染在 ChatService 内。阶段 2 只把真实 ReactAgent 和确定性 Java 服务接入 Graph action ports,并抽出可复用协议组件,生产切换留到阶段 3。
|
||||||
|
|
||||||
|
本 change 涉及 `graph.diagnosis`、Hook 和 ChatService 内部协作,接口影响为 L2。`/api/chat`、数据库、Trace API、Prompt 业务协议和当前生产路由均不改变。
|
||||||
|
|
||||||
|
## Goals / Non-Goals
|
||||||
|
|
||||||
|
**Goals:**
|
||||||
|
|
||||||
|
- 提供显式透传 `RunnableConfig` 的 Planner、Executor、Verifier、Composer adapters。
|
||||||
|
- 将 Executor、Verifier、Composer 协议解析和安全渲染抽为单一共享实现,并让旧路径委托以保持行为。
|
||||||
|
- 将 Gatekeeper、可信输入投影、evidence retry prepare 和两类 Fallback 实现为确定性 Java Node。
|
||||||
|
- 保证 Verifier 只接收通过 binding 投影的材料,并将执行状态、模型 verdict 和 effective verdict 分离。
|
||||||
|
- 修正 LOW_CONFID evidence retry 仅接受 critical gap 的阶段 1 偏差。
|
||||||
|
|
||||||
|
**Non-Goals:**
|
||||||
|
|
||||||
|
- 不切换 `ChatService.executeChatComplex` 到 StateGraph。
|
||||||
|
- 不修改 `/api/chat`、数据库、Run/Trace DTO 或持久化。
|
||||||
|
- 不删除 SequentialAgent、VerifierInputHook 或 VerifierContextHolder。
|
||||||
|
- 不修改共享 Verifier Prompt;其输入说明在阶段 3 切换时同步。
|
||||||
|
- 不运行 Maven live E2E、日志或数据库验收。
|
||||||
|
|
||||||
|
## Decisions
|
||||||
|
|
||||||
|
### 1. Agent 调用经最小 port 隔离
|
||||||
|
|
||||||
|
`DiagnosisAgentInvoker` 只暴露 `invoke(String, RunnableConfig) -> String`;`ReactAgentDiagnosisInvoker` 包装 `ReactAgent.call(input, config)` 并返回消息文本。各 Agent adapter 自己负责白名单输入序列化、输出解析、状态映射和 orchestration event,不读取完整父 Graph State。
|
||||||
|
|
||||||
|
选择最小 port 而不是直接使用 `ReactAgent.asNode(...)`,因为显式字符串输入便于证明输入白名单、固定技术重试输入和 run config 透传,也避免父 Graph messages/private state 泄漏。ReactAgent/invoker 通过构造注入;本阶段不复制 ChatService 的 Prompt 或 Agent factory,生产实例装配属于阶段 3。
|
||||||
|
|
||||||
|
### 2. 共享协议组件替代复制
|
||||||
|
|
||||||
|
在中立的 `com.superbiz.agent.diagnosis.protocol` 包抽取无状态的 `JsonPayloadSupport`、Executor/Verifier/Composer parser、Composer safe input builder 和 fallback renderer。旧 `VerifierInputHook`/`ChatService` 委托这些组件,新 Nodes 使用同一实现;抽取本身不得改变旧可观察行为。共享组件不得依赖 `graph.diagnosis`、Hook、ChatService 或 ThreadLocal。
|
||||||
|
|
||||||
|
选择共享组件而不是在 Graph 包复制旧私有方法,避免旧 Sequential 与新 Graph 对同一 JSON 契约产生两个真理源。阶段 2 保留旧控制流只是迁移顺序,不形成长期双轨。
|
||||||
|
|
||||||
|
### 3. Node failure 分类 fail closed
|
||||||
|
|
||||||
|
Adapter 将合法结构映射为 COMPLETED,将可识别的临时调用失败映射为 RETRYABLE_FAILED,将非法 JSON/结构映射为 INVALID_OUTPUT;未知或明确不可重试异常映射为 NON_RETRYABLE_FAILED/FAILED。Executor 不论失败类型都不重试,合法 no-evidence 结构仍为 COMPLETED。
|
||||||
|
|
||||||
|
分类器允许注入以便测试。无法确定的异常不猜测为可重试,防止无限或扩大副作用。
|
||||||
|
|
||||||
|
### 4. Gatekeeper 是 Graph 唯一验证入口
|
||||||
|
|
||||||
|
Graph Verifier ReactAgent 不注册旧 `VerifierInputHook`。显式 Gatekeeper Node 从 `RunnableConfig.metadata.runId` 和 `executor_output` 调用 `ExecutorGatekeeperService.validateRun` 恰好一次,同时保存 raw result 与 normalized status:pass 为 PASS;fail/low_confid 为 LOW_CONFID;fail/reject 为 REJECT;缺失、unknown、异常或自相矛盾均为 REJECT。
|
||||||
|
|
||||||
|
旧 Sequential 路径在阶段 3 前仍由 Hook 调用 Gatekeeper。两个入口服务于互斥的控制流,不允许同一次 Graph run 双执行。
|
||||||
|
|
||||||
|
### 5. Verified Input 按 binding 精确投影
|
||||||
|
|
||||||
|
Builder 只接受 Gatekeeper `checked_bindings.status=pass`。它以 `claim_id + source_invocation_id + tool_name + raw_path` 精确关联 Executor claim binding,输出过滤后的 `verified_executor_output` 和只含 `claim_id/source_invocation_id/tool_name/raw_path/matched_text` 的 `verified_evidence`。
|
||||||
|
|
||||||
|
失败 binding、hypotheses、未引用工具结果、完整 `tool_trace_summary` 和 raw Executor 文本不得进入 Verifier。合法零 claim 输出保留空列表,但 LOW_CONFID 零可信 binding 已由 Router 在 Builder 前阻断。
|
||||||
|
|
||||||
|
### 6. Verifier 状态与 verdict 分离
|
||||||
|
|
||||||
|
Verifier parser 产出执行状态和模型 verdict;Adapter 再应用 Gatekeeper ceiling 得到 effective verdict。技术失败只写 `verifier_status`,不得伪造诊断 verdict。ceiling=LOW_CONFID 时模型 PASS 最高只能得到 LOW_CONFID。
|
||||||
|
|
||||||
|
Verifier technical retry 复用首次生成的完全相同输入字符串;Composer technical retry 同样复用首次 allowed-material 输入。重试不得读取变化后的前序 raw state。
|
||||||
|
|
||||||
|
### 7. Evidence retry 使用共享 critical-gap extractor
|
||||||
|
|
||||||
|
`EvidenceGapExtractor` 同时被 Router guard 与 Retry Prepare Node 使用,唯一资格为 `is_critical=true` 且 `verification` 为 `no_evidence` 或 `indirect_support`。这样修复阶段 1 Router 漏检 critical 标记,又避免路由判断与 retry payload 漂移。
|
||||||
|
|
||||||
|
Retry context 包含 prior verified output/evidence、结构化 gaps、从 verified evidence 去重得到的 completed query refs,以及固定约束:最多一次、不重复成功查询、只做增量查询、保留 prior verified claims。第二轮 Executor 被要求执行增量查询但输出完整 `executor_evidence_v2` 快照;Java 不合并 claim 文本,完整快照重新经过 Gatekeeper。
|
||||||
|
|
||||||
|
### 8. 两类 Fallback 使用不同材料边界
|
||||||
|
|
||||||
|
前置验证 Fallback 只使用 reason code、校验状态、工具概况和人工查看 Trace 建议,绝不输出 Executor claim。Composer 后置 Fallback 只使用 Verifier 已允许的 claims、missing info 和 recommendations,不读取 raw Executor/tool output。
|
||||||
|
|
||||||
|
Fallback 由确定性 renderer 生成;若连安全答案都无法生成,异常交给阶段 3 的外层 Run failure handling。
|
||||||
|
|
||||||
|
### 9. 阶段内装配与测试边界
|
||||||
|
|
||||||
|
新增真实 Node action set/factory 装配入口,但不让 ChatService 成为消费者。单元测试使用 fake invoker 验证 Node 契约,使用真实 `ExecutorGatekeeperService` mock 边界验证单次调用,并以 CompiledGraph 场景验证 adapters、critical evidence retry 和 Fallback 路由。旧 Sequential/Hook/Gatekeeper focused tests作为共享抽取回归门禁。
|
||||||
|
|
||||||
|
## Interface Impact
|
||||||
|
|
||||||
|
- 等级:L2 内部接口。
|
||||||
|
- 新内部消费者:阶段 3 的 Graph orchestrator/Agent factory。
|
||||||
|
- 旧内部消费者:VerifierInputHook 与 ChatService 改为委托共享 parser/renderer,公开方法与外部响应不变。
|
||||||
|
- 外部 API、DB、Trace、Prompt 输出协议:无变化。
|
||||||
|
- 回滚:revert 本阶段提交;生产仍走旧 Sequential 路径。
|
||||||
|
|
||||||
|
## Risks / Trade-offs
|
||||||
|
|
||||||
|
- [共享解析器抽取造成旧行为漂移] → 保留原输入/输出形态并运行旧 focused tests。
|
||||||
|
- [binding 关联不精确导致证据泄漏] → 使用四元键精确匹配并覆盖 mixed pass/fail binding。
|
||||||
|
- [Gatekeeper 双执行] → Graph Verifier 不注册旧 Hook,并以调用次数测试锁定。
|
||||||
|
- [异常误判为可重试] → 未知异常默认 fail closed;分类器用显式用例覆盖。
|
||||||
|
- [阶段 2 Prompt 与 Node payload 说明暂不一致] → 本阶段不接生产入口;阶段 3 切换时同一 change 更新 Prompt。
|
||||||
|
|
||||||
|
## Migration Plan
|
||||||
|
|
||||||
|
1. 先抽共享协议组件并让旧 Hook/ChatService 委托,运行旧 focused tests。
|
||||||
|
2. 接入 invoker、Agent adapters、Gatekeeper、Verified Input、Retry Prepare 和 Fallback Nodes。
|
||||||
|
3. 修正 Router critical-gap guard,装配未接生产入口的真实 Graph 并运行 Node/Graph tests。
|
||||||
|
4. 归档并提交本阶段;阶段 3 再切换 ChatService、Prompt、DB 和 Trace。
|
||||||
|
|
||||||
|
回滚为完整 revert 本阶段提交;不需要数据库回滚,不存在外部协议迁移。
|
||||||
|
|
||||||
|
## Open Questions
|
||||||
|
|
||||||
|
无。阶段 3 前不得扩大范围到生产入口或 Prompt 切换。
|
||||||
@@ -0,0 +1,82 @@
|
|||||||
|
# Chat Diagnosis StateGraph Real Nodes
|
||||||
|
|
||||||
|
## Why
|
||||||
|
|
||||||
|
阶段 1 已提供可编译的 StateGraph 骨架和 35 个 Fake Node 路由测试,但所有 action ports 仍是测试脚本,尚不能调用现有 ReactAgent、Executor Gatekeeper 或构造安全的 Verifier/Composer 输入。阶段 2 需要把现有 Agent 协议和确定性服务接到 Graph ports,同时保持 ChatService 生产入口仍走旧链路,避免把真实 Node 接入与生产切换混成一个不可回滚阶段。
|
||||||
|
|
||||||
|
## What Changes
|
||||||
|
|
||||||
|
- 新增 `DiagnosisAgentInvoker` 及 ReactAgent adapter,所有 Agent Node 显式传递 `RunnableConfig`。
|
||||||
|
- 新增 Planner、Executor、Verifier、Composer Node Adapters 和可注入失败分类器。
|
||||||
|
- 抽取 Executor/Verifier/Composer 共享解析组件,旧 Hook/ChatService 委托它们,避免复制协议逻辑。
|
||||||
|
- 新增显式 Gatekeeper Node,调用 `ExecutorGatekeeperService.validateRun` 并 fail-closed 标准化状态。
|
||||||
|
- 新增 Verified Input Builder,只投影 Gatekeeper passed bindings 对应的 claims 与 `matched_text` evidence。
|
||||||
|
- 新增 Evidence Gap Extractor / Retry Prepare Node,仅处理关键 `no_evidence` / `indirect_support` facts。
|
||||||
|
- 新增 Composer Safe Input Builder 和两类固定 Fallback Node 表达。
|
||||||
|
- 修正阶段 1 Router:evidence retry guard 必须要求关键 evidence gap。
|
||||||
|
- 新增 Node 契约和真实 CompiledGraph 集成测试,并保留旧 Sequential/Hook/Gatekeeper 回归。
|
||||||
|
|
||||||
|
本 change 不切换 ChatService 到 StateGraph,不修改 DB/Trace API,也不删除旧 Hook/Sequential 结构。
|
||||||
|
|
||||||
|
## Capabilities
|
||||||
|
|
||||||
|
### New Capabilities
|
||||||
|
|
||||||
|
- `chat-diagnosis-stategraph-real-nodes`:提供真实 Agent/Java Node adapters、可信输入投影、关键证据补查上下文和安全 Fallback。
|
||||||
|
|
||||||
|
### Modified Capabilities
|
||||||
|
|
||||||
|
- `chat-diagnosis-stategraph-routing-skeleton`:将 LOW_CONFID evidence retry guard 收紧为仅接受 `is_critical=true` 的 `no_evidence` 或 `indirect_support` facts,与阶段 0 冻结基线一致。
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
### In Scope
|
||||||
|
|
||||||
|
- `com.superbiz.agent.graph.diagnosis` 下的 Agent adapters、Gatekeeper、Verified Input、Evidence Retry、Composer Input 和 Fallback。
|
||||||
|
- 可被旧 Hook/ChatService 与新 Nodes 共用的输出解析/安全渲染组件。
|
||||||
|
- 旧 Hook/ChatService 的行为保持型委托重构。
|
||||||
|
- `DiagnosisGraphRouter` 的 critical gap guard 修正。
|
||||||
|
- Node 单元测试、真实 CompiledGraph adapter tests、旧路径 focused regressions。
|
||||||
|
|
||||||
|
### Out of Scope
|
||||||
|
|
||||||
|
- 将 `ChatService.executeChatComplex` 替换为 Graph invocation。
|
||||||
|
- 修改 Agent 基础 Prompt 业务语义或输出协议。
|
||||||
|
- 移除 `VerifierInputHook` / `VerifierContextHolder` / SequentialAgent。
|
||||||
|
- Flyway、DiagnosisRun orchestration trace 字段、Trace DTO/API。
|
||||||
|
- 最终全量 Graph 测试体系替换。
|
||||||
|
- Maven E2E、`logs/` 和数据库核验。
|
||||||
|
|
||||||
|
## Context Constraints
|
||||||
|
|
||||||
|
- 阶段 0/1 archives 是执行基线;Node 不得改变条件边和计数所有权。
|
||||||
|
- Graph Verifier 只接收 `verified_executor_output`、`verified_evidence`、gatekeeper audit/ceiling、query 和 retry context,不接收完整 tool trace 或 Executor raw text。
|
||||||
|
- Gatekeeper validation 使用当前 `runId`,缺失/未知/异常一律 REJECT。
|
||||||
|
- Gatekeeper raw result 与 normalized status 分开保存。
|
||||||
|
- Verifier execution status、model verdict 和 effective verdict 分离;status 永远不写进 verdict。
|
||||||
|
- Planner/Verifier/Composer retry 输入由相同 state 白名单投影构造,技术 retry 不重跑前序节点。
|
||||||
|
- Executor 不重试;合法 no-evidence 仍为 COMPLETED。
|
||||||
|
- Evidence Retry Context 保留 prior verified output/evidence、关键 gaps、completed query refs 和固定约束;Java 不合并 claim 文本。
|
||||||
|
- 旧 Sequential path 在阶段 3 前继续工作;共享解析器重构必须通过旧 focused tests。
|
||||||
|
- 本阶段不修改 `/api/chat`、DB 或当前运行路由。
|
||||||
|
|
||||||
|
## Acceptance
|
||||||
|
|
||||||
|
- ReactAgent invoker 透传输入与 RunnableConfig,现有 Agent Hook/ToolCallback 可继续工作。
|
||||||
|
- Planner 合法 JSON 为 COMPLETED;非法 JSON/结构为 INVALID_OUTPUT;临时失败与不可重试失败分开。
|
||||||
|
- Executor 合法 `executor_evidence_v2`(包括 no-evidence)为 COMPLETED;非法结构、工具阻断、其他失败正确映射且不重试。
|
||||||
|
- Gatekeeper Node 每轮只调用一次 `validateRun`,unknown/error fail closed。
|
||||||
|
- PASS 与可继续 LOW_CONFID 都经过 Verified Input Builder;失败 binding、未引用工具结果和 raw Executor text 不进入 Verifier。
|
||||||
|
- Verifier input 只有白名单字段;输出分离 status/model/effective verdict,ceiling 正确限制 PASS。
|
||||||
|
- Verifier/Composer technical retry 使用相同序列化输入。
|
||||||
|
- Evidence retry 只从关键 gap 生成,包含 prior verified data/completed refs/约束,最多一次。
|
||||||
|
- 第二轮 Executor 输入明确要求增量查询和完整 `executor_evidence_v2` 快照,不由 Java 合并 claims。
|
||||||
|
- Pre-verification Fallback 不输出 Executor claim;post-verification Composer fallback 只使用 allowed material。
|
||||||
|
- 新 Node/Graph tests 和旧路径 focused regressions 通过;不运行 Maven E2E。
|
||||||
|
|
||||||
|
## Risks
|
||||||
|
|
||||||
|
- 共享解析器抽取可能改变旧 Sequential 行为;旧 ChatService/Hook tests 必须同批通过。
|
||||||
|
- Gatekeeper checked binding 与 claim binding 关联错误可能泄漏失败 evidence;使用 claim_id + invocation/tool/path 精确匹配并测试 mixed binding。
|
||||||
|
- 异常分类依赖 cause/message;未知异常默认不可重试或 FAILED,安全优先。
|
||||||
|
- Prompt 当前仍描述旧 Hook payload;本阶段只验证 Node input contract,阶段 3 在生产切换时同步 Verifier Prompt 输入说明,避免旧生产路径提前不兼容。
|
||||||
+132
@@ -0,0 +1,132 @@
|
|||||||
|
## ADDED Requirements
|
||||||
|
|
||||||
|
### Requirement: Agent adapters SHALL invoke real agents through explicit run config
|
||||||
|
|
||||||
|
The system SHALL provide Planner, Executor, Verifier, and Composer adapters that invoke their configured ReactAgent through a minimal invoker port with an explicit `RunnableConfig`. Each adapter SHALL serialize only its declared input fields and append one terminal orchestration event per attempt.
|
||||||
|
|
||||||
|
#### Scenario: Adapter invokes an agent
|
||||||
|
|
||||||
|
- **WHEN** an adapter receives valid Graph State and a RunnableConfig containing the current run metadata
|
||||||
|
- **THEN** it SHALL pass its projected input and the same RunnableConfig to the configured invoker
|
||||||
|
- **AND** existing Agent hooks and ToolCallbacks SHALL be able to observe the current run metadata
|
||||||
|
|
||||||
|
#### Scenario: Parent state is inspected
|
||||||
|
|
||||||
|
- **WHEN** an adapter builds an Agent input
|
||||||
|
- **THEN** it SHALL NOT serialize undeclared Graph State, raw prompts, other Agent private state, or orchestration events
|
||||||
|
|
||||||
|
### Requirement: Agent outputs SHALL be parsed into explicit execution statuses
|
||||||
|
|
||||||
|
The system SHALL use shared Executor, Verifier, and Composer protocol parsers for both Graph Nodes and the legacy path. Legal structured output SHALL map to COMPLETED; invalid JSON or contract shape SHALL map to INVALID_OUTPUT; recognized transient invocation failure SHALL map to RETRYABLE_FAILED where that Agent supports technical retry; unknown or permanent failure SHALL fail closed. A legal Executor no-evidence result SHALL be COMPLETED.
|
||||||
|
|
||||||
|
#### Scenario: Legal no-evidence Executor output
|
||||||
|
|
||||||
|
- **WHEN** Executor returns a valid `executor_evidence_v2` document containing a legal no-evidence result
|
||||||
|
- **THEN** Executor status SHALL be COMPLETED
|
||||||
|
- **AND** the Graph SHALL continue to Gatekeeper
|
||||||
|
|
||||||
|
#### Scenario: Invalid structured output
|
||||||
|
|
||||||
|
- **WHEN** an Agent returns malformed JSON or violates its required output structure
|
||||||
|
- **THEN** its adapter SHALL set INVALID_OUTPUT
|
||||||
|
- **AND** it SHALL NOT fabricate a diagnostic verdict or evidence
|
||||||
|
|
||||||
|
#### Scenario: Legacy parser behavior is exercised
|
||||||
|
|
||||||
|
- **WHEN** the existing Sequential path parses the same payloads after shared component extraction
|
||||||
|
- **THEN** its externally observable parser and safe-rendering behavior SHALL remain unchanged
|
||||||
|
|
||||||
|
### Requirement: Gatekeeper Node SHALL validate exactly once and fail closed
|
||||||
|
|
||||||
|
The Gatekeeper Node SHALL call `ExecutorGatekeeperService.validateRun` exactly once for the current run and Executor structured output. It SHALL preserve the raw result separately from normalized status. Raw pass SHALL normalize to PASS; fail with low-confid severity SHALL normalize to LOW_CONFID; fail with reject severity SHALL normalize to REJECT; missing, unknown, inconsistent, or exceptional results SHALL normalize to REJECT.
|
||||||
|
|
||||||
|
#### Scenario: Gatekeeper passes output
|
||||||
|
|
||||||
|
- **WHEN** `validateRun` returns a valid pass result
|
||||||
|
- **THEN** the Node SHALL store the raw result and status PASS
|
||||||
|
- **AND** validation SHALL have been called exactly once with the current runId
|
||||||
|
|
||||||
|
#### Scenario: Gatekeeper result cannot be trusted
|
||||||
|
|
||||||
|
- **WHEN** validation throws or returns a missing, unknown, or inconsistent result
|
||||||
|
- **THEN** the Node SHALL normalize status to REJECT
|
||||||
|
- **AND** Verifier SHALL NOT receive unverified Executor material
|
||||||
|
|
||||||
|
### Requirement: Verified Input Builder SHALL project only passed bindings
|
||||||
|
|
||||||
|
PASS and continuable LOW_CONFID results SHALL pass through a Verified Input Builder. The Builder SHALL match passed checked bindings to Executor claims by `claim_id`, `source_invocation_id`, `tool_name`, and `raw_path`, and SHALL produce only filtered `verified_executor_output` plus `verified_evidence` entries containing the matched binding fields and `matched_text`.
|
||||||
|
|
||||||
|
#### Scenario: Mixed checked bindings are projected
|
||||||
|
|
||||||
|
- **WHEN** Gatekeeper returns both passed and failed checked bindings
|
||||||
|
- **THEN** only claims and evidence matching passed bindings SHALL be projected
|
||||||
|
- **AND** failed bindings, hypotheses, unreferenced tool results, and raw Executor text SHALL be absent
|
||||||
|
|
||||||
|
#### Scenario: Verifier input is serialized
|
||||||
|
|
||||||
|
- **WHEN** the Verifier adapter builds its input
|
||||||
|
- **THEN** it SHALL include only diagnosis query context, verified Executor output, verified evidence, Gatekeeper audit/ceiling, and permitted retry context
|
||||||
|
- **AND** it SHALL NOT include complete `tool_trace_summary` or raw Executor output
|
||||||
|
|
||||||
|
### Requirement: Verifier SHALL separate execution status from diagnostic verdict
|
||||||
|
|
||||||
|
The Verifier adapter SHALL store `verifier_status`, `verifier_model_verdict`, and `effective_verdict` as separate values. A Gatekeeper LOW_CONFID ceiling SHALL prevent model PASS from producing effective PASS. Technical retry SHALL reuse the exact same serialized verified input and SHALL NOT rerun any preceding Node.
|
||||||
|
|
||||||
|
#### Scenario: Ceiling limits model verdict
|
||||||
|
|
||||||
|
- **WHEN** Gatekeeper ceiling is LOW_CONFID and the model verdict is PASS
|
||||||
|
- **THEN** effective verdict SHALL be LOW_CONFID
|
||||||
|
- **AND** verifier execution status SHALL remain COMPLETED
|
||||||
|
|
||||||
|
#### Scenario: Verifier technical retry occurs
|
||||||
|
|
||||||
|
- **WHEN** the first Verifier attempt returns INVALID_OUTPUT or RETRYABLE_FAILED
|
||||||
|
- **THEN** its single retry SHALL receive the same serialized input
|
||||||
|
- **AND** Executor, Gatekeeper, Verified Input, and tools SHALL NOT rerun
|
||||||
|
|
||||||
|
### Requirement: Evidence retry SHALL contain only structured critical gaps and incremental constraints
|
||||||
|
|
||||||
|
The system SHALL extract evidence gaps only from facts with `is_critical=true` and verification `no_evidence` or `indirect_support`. Retry Prepare SHALL include prior verified output/evidence, structured gaps, deduplicated completed query references, and fixed constraints requiring at most one incremental retry without repeating successful queries. The second Executor invocation SHALL be instructed to return a complete `executor_evidence_v2` snapshot; Java code SHALL NOT merge claim text.
|
||||||
|
|
||||||
|
#### Scenario: Critical gaps prepare a retry
|
||||||
|
|
||||||
|
- **WHEN** effective verdict is LOW_CONFID, ceiling is PASS, evidence retry count is zero, and at least one qualifying critical gap exists
|
||||||
|
- **THEN** Retry Prepare SHALL build the bounded retry context and increment evidence retry count once
|
||||||
|
- **AND** the new Planner stage SHALL use EVIDENCE_GAP_ONLY mode
|
||||||
|
|
||||||
|
#### Scenario: Non-critical gap is present
|
||||||
|
|
||||||
|
- **WHEN** facts contain only non-critical no-evidence or indirect-support items
|
||||||
|
- **THEN** no evidence retry SHALL occur
|
||||||
|
- **AND** the Graph SHALL continue to Composer
|
||||||
|
|
||||||
|
#### Scenario: Second Executor input is built
|
||||||
|
|
||||||
|
- **WHEN** Planner produces an evidence-gap-only incremental plan
|
||||||
|
- **THEN** Executor input SHALL prohibit repeating completed queries and require a complete output snapshot preserving prior verified claims
|
||||||
|
- **AND** the Java layer SHALL NOT semantically merge old and new claims
|
||||||
|
|
||||||
|
### Requirement: Composer and Fallback SHALL use only allowed material
|
||||||
|
|
||||||
|
Composer SHALL receive only effective verdict and Verifier-allowed claims, missing information, and recommendations. Composer technical retry SHALL reuse the exact same serialized input. A pre-verification Fallback SHALL never output Executor claims; a post-verification Composer Fallback SHALL use only Verifier-allowed material.
|
||||||
|
|
||||||
|
#### Scenario: Pre-verification path degrades
|
||||||
|
|
||||||
|
- **WHEN** Planner, Executor, Gatekeeper, Verified Input, or Verifier cannot establish trusted material
|
||||||
|
- **THEN** deterministic Fallback output SHALL contain no Executor claim or raw tool output
|
||||||
|
|
||||||
|
#### Scenario: Composer retry is exhausted
|
||||||
|
|
||||||
|
- **WHEN** Composer fails after its one technical retry and Verifier-allowed material exists
|
||||||
|
- **THEN** deterministic Fallback SHALL express only the allowed claims, missing information, and recommendations
|
||||||
|
- **AND** it SHALL NOT read raw Executor or tool output
|
||||||
|
|
||||||
|
### Requirement: Real Nodes SHALL remain isolated from the production Chat path in stage 2
|
||||||
|
|
||||||
|
The real Node action set and CompiledGraph SHALL be constructible and testable, but ChatService, database, Trace API, shared Agent prompts, and the current production routing SHALL remain unchanged until the stage 3 change.
|
||||||
|
|
||||||
|
#### Scenario: Stage 2 production isolation is inspected
|
||||||
|
|
||||||
|
- **WHEN** this change is accepted
|
||||||
|
- **THEN** no production ChatService code path SHALL invoke the real Diagnosis Graph
|
||||||
|
- **AND** no database migration, Trace API field, or shared Prompt contract SHALL be changed
|
||||||
+33
@@ -0,0 +1,33 @@
|
|||||||
|
## MODIFIED Requirements
|
||||||
|
|
||||||
|
### Requirement: Verifier routing SHALL separate technical retry from evidence retry
|
||||||
|
|
||||||
|
Verifier INVALID_OUTPUT and RETRYABLE_FAILED SHALL self-retry once with the same verified input. COMPLETED PASS or REJECT SHALL route to Composer. COMPLETED LOW_CONFID SHALL route to one Evidence Retry only when all frozen guards are true, including at least one critical evidence gap; otherwise it SHALL route to Composer. Other outcomes SHALL fail closed.
|
||||||
|
|
||||||
|
#### Scenario: Verifier first technical failure
|
||||||
|
|
||||||
|
- **WHEN** Verifier first returns INVALID_OUTPUT or RETRYABLE_FAILED
|
||||||
|
- **THEN** only Verifier SHALL run again
|
||||||
|
- **AND** Gatekeeper, Executor, and tools SHALL NOT rerun
|
||||||
|
|
||||||
|
#### Scenario: Verifier retry is exhausted
|
||||||
|
|
||||||
|
- **WHEN** Verifier returns a technical failure after `verifier_retry_count=1`
|
||||||
|
- **THEN** the Graph SHALL route to Fallback
|
||||||
|
|
||||||
|
#### Scenario: Verifier verdict reaches Composer
|
||||||
|
|
||||||
|
- **WHEN** Verifier completes with effective PASS or REJECT
|
||||||
|
- **THEN** Composer SHALL run
|
||||||
|
|
||||||
|
#### Scenario: LOW_CONFID qualifies for evidence retry
|
||||||
|
|
||||||
|
- **WHEN** Verifier completes LOW_CONFID with ceiling PASS, at least one fact having `is_critical=true` and verification `no_evidence` or `indirect_support`, and `evidence_retry_count=0`
|
||||||
|
- **THEN** Evidence Retry SHALL run once and return to a new Planner stage
|
||||||
|
- **AND** `planner_retry_count` SHALL reset to 0
|
||||||
|
- **AND** `evidence_retry_count` SHALL become 1
|
||||||
|
|
||||||
|
#### Scenario: LOW_CONFID does not qualify for evidence retry
|
||||||
|
|
||||||
|
- **WHEN** ceiling is LOW_CONFID, facts contain no critical valid gap, or evidence retry count is already 1
|
||||||
|
- **THEN** Composer SHALL run without another Planner cycle
|
||||||
@@ -0,0 +1,44 @@
|
|||||||
|
## 1. Shared Protocol Components
|
||||||
|
|
||||||
|
- [x] 1.1 Extract stateless JSON payload and Executor evidence parsing components into neutral `com.superbiz.agent.diagnosis.protocol`, preserving legal no-evidence behavior and introducing no dependency on Graph, Hook, ChatService, or ThreadLocal.
|
||||||
|
- [x] 1.2 Extract Verifier and Composer output parsers, effective-verdict calculation, Composer safe-input construction, and deterministic safe rendering from ChatService.
|
||||||
|
- [x] 1.3 Make VerifierInputHook and ChatService delegate to the shared components without changing their public methods or legacy payload/output behavior.
|
||||||
|
- [x] 1.4 Run `VerifierInputHookTest`, `ChatServiceSequentialAgentTest`, and `ExecutorGatekeeperServiceTest` after the extraction and fix only behavior regressions in scope.
|
||||||
|
|
||||||
|
## 2. Agent Invocation And Adapters
|
||||||
|
|
||||||
|
- [x] 2.1 Implement constructor-injected `DiagnosisAgentInvoker` and a ReactAgent wrapper that forwards the exact input and RunnableConfig and returns AssistantMessage text, without adding a second Prompt/Agent factory.
|
||||||
|
- [x] 2.2 Implement a fail-closed, injectable Node failure classifier for invalid output, retryable invocation failure, and non-retryable failure.
|
||||||
|
- [x] 2.3 Implement Planner adapter with NORMAL/EVIDENCE_GAP_ONLY input projection, structured plan parsing, status mapping, and one event per attempt.
|
||||||
|
- [x] 2.4 Implement Executor adapter with incremental retry instructions, legal no-evidence completion, no technical retry, and complete-snapshot output parsing.
|
||||||
|
- [x] 2.5 Add focused adapter tests proving input whitelists, RunnableConfig identity, failure classification, event shape, legal no-evidence handling, and absence of parent-state leakage.
|
||||||
|
|
||||||
|
## 3. Gatekeeper And Verified Projection
|
||||||
|
|
||||||
|
- [x] 3.1 Implement the explicit Gatekeeper Node using current runId and `ExecutorGatekeeperService.validateRun`, preserving raw result separately from normalized status.
|
||||||
|
- [x] 3.2 Normalize missing, unknown, inconsistent, and exceptional Gatekeeper results to REJECT and record a deterministic reason code.
|
||||||
|
- [x] 3.3 Implement Verified Input Builder matching passed bindings by claim/invocation/tool/path and projecting only filtered claims plus matched-text evidence.
|
||||||
|
- [x] 3.4 Add mixed-binding tests proving one Gatekeeper call, correct PASS/LOW_CONFID/REJECT normalization, precise pass projection, and exclusion of failed, unreferenced, hypothesis, tool-summary, and raw materials.
|
||||||
|
|
||||||
|
## 4. Verifier, Evidence Retry, Composer And Fallback
|
||||||
|
|
||||||
|
- [x] 4.1 Implement Verifier adapter with whitelisted verified input, separate execution/model/effective fields, ceiling enforcement, and byte-identical technical retry input.
|
||||||
|
- [x] 4.2 Implement shared `EvidenceGapExtractor` and update `DiagnosisGraphRouter` so only critical no-evidence/indirect-support facts qualify.
|
||||||
|
- [x] 4.3 Implement Evidence Retry Prepare Node with prior verified data, structured gaps, deduplicated completed query refs, bounded constraints, and planner retry reset.
|
||||||
|
- [x] 4.4 Implement Composer adapter and safe-input builder with byte-identical technical retry input and no raw Executor/tool access.
|
||||||
|
- [x] 4.5 Implement distinct deterministic pre-verification and post-verification Fallback inputs/rendering with their allowed-material boundaries.
|
||||||
|
- [x] 4.6 Add focused tests for ceiling, fixed retry inputs, non-critical gap rejection, critical gap context, no Java claim merge, and both Fallback safety boundaries.
|
||||||
|
|
||||||
|
## 5. Real Graph Assembly
|
||||||
|
|
||||||
|
- [x] 5.1 Extend the diagnosis action set/factory to construct a CompiledGraph from the real Agent and deterministic Node dependencies without registering legacy VerifierInputHook on the Graph Verifier.
|
||||||
|
- [x] 5.2 Add CompiledGraph integration tests for PASS, legal no-evidence, Gatekeeper REJECT/LOW_CONFID, critical evidence retry, exhausted Agent technical retry, and Composer fallback paths.
|
||||||
|
- [x] 5.3 Verify Graph events/transitions, per-node invocation counts, same-input retries, one full-snapshot Gatekeeper revalidation after evidence retry, and bounded termination.
|
||||||
|
- [x] 5.4 Confirm the first completed Node module and final assembly against design/specs, with no core TODO or placeholder implementation.
|
||||||
|
|
||||||
|
## 6. Stage 2 Verification And Handoff
|
||||||
|
|
||||||
|
- [x] 6.1 Run new Node/Graph focused tests plus `DiagnosisGraphRoutingTest` and `DiagnosisOrchestrationTraceBuilderTest`.
|
||||||
|
- [x] 6.2 Run legacy focused regressions `ChatServiceSequentialAgentTest`, `VerifierInputHookTest`, and `ExecutorGatekeeperServiceTest`, then run Maven test compilation.
|
||||||
|
- [x] 6.3 Run OpenSpec strict validation, `git diff --check`, and source/reference checks proving ChatService does not invoke the real Graph and no DB/Trace/Prompt contract changed.
|
||||||
|
- [x] 6.4 Record that Maven live E2E, `logs/`, and database verification were intentionally not run in stage 2 and are reserved for stage 5.
|
||||||
+3
@@ -0,0 +1,3 @@
|
|||||||
|
ready_at: 2026-07-17
|
||||||
|
devflow: devflow/projects/2026-07-17-chat-diagnosis-stategraph-routing-skeleton
|
||||||
|
authorization: user-requested-per-stage-archive
|
||||||
+3
@@ -0,0 +1,3 @@
|
|||||||
|
committed_at: 2026-07-17
|
||||||
|
scope: iss-011-stage-1-routing-skeleton
|
||||||
|
validation: openspec-strict-pass
|
||||||
+169
@@ -0,0 +1,169 @@
|
|||||||
|
## Context
|
||||||
|
|
||||||
|
阶段 0 已归档 StateGraph 设计基线,当前仓库仍没有 Graph 实现。阶段 1 只引入可编译、可用 Fake Node 执行的内部骨架,不连接 Spring Bean、ChatService、真实 Agent 或数据库。
|
||||||
|
|
||||||
|
锁定依赖 `spring-ai-alibaba-graph-core:1.1.2.0` 已通过本地 JAR 验证:
|
||||||
|
|
||||||
|
- `StateGraph.addNode(... AsyncNodeActionWithConfig)`
|
||||||
|
- `StateGraph.addConditionalEdges(... AsyncEdgeActionWithConfig, mappings)`
|
||||||
|
- `CompileConfig.builder().recursionLimit(...)`
|
||||||
|
- `KeyStrategyFactoryBuilder`、`ReplaceStrategy`、`AppendStrategy`
|
||||||
|
- `CompiledGraph.invoke(input, RunnableConfig)`
|
||||||
|
- `RunnableConfig.builder().threadId(...)`
|
||||||
|
|
||||||
|
阶段 2 将消费 Node action ports,阶段 3 将消费 compiled graph。阶段 1 自身没有生产调用方。
|
||||||
|
|
||||||
|
## Goals / Non-Goals
|
||||||
|
|
||||||
|
**Goals:**
|
||||||
|
|
||||||
|
- 建立明确的状态、状态枚举、节点、route 和 reason code 常量。
|
||||||
|
- 建立 config-aware Node action ports 和可编译 Graph 工厂。
|
||||||
|
- 以冻结计数和 guard 实现完整条件边。
|
||||||
|
- 验证 AppendStrategy 和实际 CompiledGraph 行为。
|
||||||
|
- 从有界 events 构造确定性的 orchestration trace。
|
||||||
|
- 用 Fake Node 测试所有路由和终止边界。
|
||||||
|
|
||||||
|
**Non-Goals:**
|
||||||
|
|
||||||
|
- 不接 ReactAgent、Gatekeeper Service、ToolTrace 或 Prompt。
|
||||||
|
- 不修改 ChatService、Hook、ThreadLocal 或生产 Spring 配置。
|
||||||
|
- 不持久化 trace,不修改 DB/DTO/API。
|
||||||
|
- 不删除旧 Sequential 实现或测试。
|
||||||
|
- 不运行 Maven E2E。
|
||||||
|
|
||||||
|
## Decisions
|
||||||
|
|
||||||
|
### 1. Package and class boundaries
|
||||||
|
|
||||||
|
新代码位于 `com.superbiz.agent.graph.diagnosis`:
|
||||||
|
|
||||||
|
| 类型 | 职责 |
|
||||||
|
|---|---|
|
||||||
|
| `DiagnosisGraphState` | 26 个 state key、默认 Replace + events Append strategy、typed reads |
|
||||||
|
| `DiagnosisGraphStatus` | Planner/Executor/Gatekeeper/Verifier/Composer status、Verdict、PlannerMode |
|
||||||
|
| `DiagnosisGraphTopology` | Node id、route key、reason code 常量 |
|
||||||
|
| `DiagnosisGraphActions` | 八个 `AsyncNodeActionWithConfig` ports |
|
||||||
|
| `DiagnosisGraphRouter` | 纯确定性 edge guard,不修改 state |
|
||||||
|
| `DiagnosisGraphFactory` | 注册 Node/Edge、retry wrapper、evidence retry wrapper、recursion limit |
|
||||||
|
| `OrchestrationEvent` | node/outcome/reasonCode/attempt |
|
||||||
|
| `OrchestrationTransition` | from/to/reasonCode/attempt |
|
||||||
|
| `DiagnosisOrchestrationTrace` | version/transitions/finalNode/terminationReason/degraded/evidenceRetryCount |
|
||||||
|
| `DiagnosisOrchestrationTraceBuilder` | events → transitions/summary |
|
||||||
|
|
||||||
|
Fake Node 和 script fixture 只存在于 test 源集。
|
||||||
|
|
||||||
|
替代方案:把 router 和状态读取放进 ChatService 或每个 Adapter。拒绝,因为会复制 guard 并阻碍 Fake Node 独立验证。
|
||||||
|
|
||||||
|
### 1.1 Direct dependency ownership
|
||||||
|
|
||||||
|
代码直接 import Graph Core API,因此 `pom.xml` 显式声明 `com.alibaba.cloud.ai:spring-ai-alibaba-graph-core`。版本继续由现有 Spring AI Alibaba BOM 管理为 1.1.2.0,不重复写版本。依赖 Agent Framework 的传递依赖虽可编译,但会让上游依赖图调整无意中破坏本模块,拒绝。
|
||||||
|
|
||||||
|
### 2. State strategy and typed reads
|
||||||
|
|
||||||
|
`KeyStrategyFactoryBuilder.defaultStrategy(new ReplaceStrategy())`,仅对 `orchestration_events` 使用 `new AppendStrategy()`。Node 每次返回 `List.of(event)`,Graph 合并后保持有序列表。
|
||||||
|
|
||||||
|
typed reads 对 absent、null、错误类型和未知 enum 返回安全默认,不抛出边路由异常。未知 status/verdict 的路由默认 Fallback。
|
||||||
|
|
||||||
|
### 3. Config-aware action ports
|
||||||
|
|
||||||
|
所有 port 使用 `AsyncNodeActionWithConfig`,即使 Fake Node 当前只需要 state。原因是阶段 2 必须读取 `RunnableConfig.threadId/metadata` 并向 Agent 传递 run context;现在锁定接口可以避免随后重写 Graph topology。
|
||||||
|
|
||||||
|
`DiagnosisGraphActions` 构造时对八个 action 做 non-null 校验。
|
||||||
|
|
||||||
|
### 4. Retry counter ownership
|
||||||
|
|
||||||
|
Edge 只选 route,不修改 state。Factory 用 wrapper 管理计数:
|
||||||
|
|
||||||
|
- Planner 重入前,如果上次 `planner_status` 是 INVALID_OUTPUT/RETRYABLE_FAILED,count + 1。
|
||||||
|
- Verifier/Composer 同理。
|
||||||
|
- 第一次失败后的 state count 为 0,允许 self-loop;重入执行后 count 为 1,再失败则 Fallback。
|
||||||
|
- Evidence Retry Node 完成时将 `evidence_retry_count + 1`、`planner_retry_count=0`、`planner_mode=EVIDENCE_GAP_ONLY`。
|
||||||
|
- Verified Input Node 完成时重置 `verifier_retry_count=0`,因为它产生新的 verified input。
|
||||||
|
- 技术重试不修改 evidence count。
|
||||||
|
|
||||||
|
替代方案:由每个 Fake/real action 自行维护计数。拒绝,因为遗漏会造成无限 self-loop,骨架必须拥有控制计数。
|
||||||
|
|
||||||
|
### 5. Conditional routing
|
||||||
|
|
||||||
|
| Source | Route keys |
|
||||||
|
|---|---|
|
||||||
|
| Planner | executor / retry_planner / fallback |
|
||||||
|
| Executor | gatekeeper / fallback |
|
||||||
|
| Gatekeeper | verified_input / fallback |
|
||||||
|
| Verifier | composer / retry_verifier / evidence_retry / fallback |
|
||||||
|
| Composer | end / retry_composer / fallback |
|
||||||
|
|
||||||
|
Verified Input 固定到 Verifier,Evidence Retry 固定到 Planner,Fallback 固定到 END。
|
||||||
|
|
||||||
|
Verifier evidence-retry guard 同时要求:
|
||||||
|
|
||||||
|
- `verifier_status=COMPLETED`
|
||||||
|
- `effective_verdict=LOW_CONFID`
|
||||||
|
- `verifier_verdict_ceiling=PASS`
|
||||||
|
- `verifier_output.facts_checked` 至少一项具有非空 fact 且 verification 为 `no_evidence` 或 `indirect_support`
|
||||||
|
- `evidence_retry_count < 1`
|
||||||
|
|
||||||
|
否则 LOW_CONFID 到 Composer。router 只检查现有 Verifier 输出,不选择工具或构造 retry context;后者属于阶段 2 Evidence Retry Node。
|
||||||
|
|
||||||
|
### 6. Recursion limit
|
||||||
|
|
||||||
|
合法最坏路径低于 20 次 Node 执行;compile recursion limit 固定 32。测试断言 compiled graph 的 max iterations/compile config,并覆盖第二次技术失败和第二次 LOW_CONFID 均终止。
|
||||||
|
|
||||||
|
recursion limit 是最后保险,不替代业务 guard。
|
||||||
|
|
||||||
|
### 7. Event and trace contract
|
||||||
|
|
||||||
|
`OrchestrationEvent` 构造时拒绝 blank node/outcome/reason 或 attempt<1。每个 Fake/真实 Node attempt 返回一个 event;AppendStrategy 累积实际路径。
|
||||||
|
|
||||||
|
Builder:
|
||||||
|
|
||||||
|
1. 要求 events 非空且类型正确。
|
||||||
|
2. 相邻 events 生成 transition,transition reason/attempt 取 source event。
|
||||||
|
3. 最后 event 决定 finalNode 和 terminationReason。
|
||||||
|
4. final node 为 Fallback 时 degraded=true;其他质量由 effective verdict 表达。
|
||||||
|
5. evidence retry count 从 state 读取。
|
||||||
|
6. `toMap()` 使用冻结 snake_case 字段,transition 同样提供 map。
|
||||||
|
|
||||||
|
不从应用日志推导路径,也不在 state 维护第二份 transitions。
|
||||||
|
|
||||||
|
### 8. Test architecture
|
||||||
|
|
||||||
|
`DiagnosisGraphRoutingTest` 使用真实 `CompiledGraph` + scriptable Fake actions:
|
||||||
|
|
||||||
|
- 每个 node 有 outcome 队列和调用计数。
|
||||||
|
- Fake action 只输出本 node 的状态字段、必要 guard payload 和一个 event。
|
||||||
|
- 每个用例使用唯一 `RunnableConfig.threadId`。
|
||||||
|
- 断言 node sequence、调用次数、最终 state、events、trace transitions。
|
||||||
|
- 参数化覆盖同类技术失败,独立用例覆盖 evidence retry/counter reset 和 unknown fail-closed。
|
||||||
|
|
||||||
|
`DiagnosisOrchestrationTraceBuilderTest` 覆盖空 events、非法 event、顺序、fallback degraded 和 map shape。
|
||||||
|
|
||||||
|
## Interface Impact
|
||||||
|
|
||||||
|
- 等级:L2 internal interface。
|
||||||
|
- 新接口消费者:阶段 2 adapters、阶段 3 orchestrator、阶段 4 test suite。
|
||||||
|
- 构建消费者:Maven 直接声明 Graph Core,版本仍由既有 BOM 统一管理。
|
||||||
|
- 当前生产消费者:无。
|
||||||
|
- 外部 API/DTO/数据库/状态:无变化。
|
||||||
|
- 兼容策略:后续 actions 实现 ports;Graph State 不直接传给 Agent。
|
||||||
|
- 回滚:revert 本阶段提交,无数据迁移。
|
||||||
|
|
||||||
|
## Risks / Trade-offs
|
||||||
|
|
||||||
|
- [Graph merge 行为与假设不同] → 使用真实 CompiledGraph 单元测试,不 mock StateGraph。
|
||||||
|
- [计数 wrapper 与 action 更新冲突] → wrapper 最后写入控制计数,actions 不拥有 retry counters。
|
||||||
|
- [Verifier gap 解析过早耦合] → router 只识别现有 facts_checked 最小字段,retry context 构造留到阶段 2。
|
||||||
|
- [未接生产入口被误认为完成] → proposal/spec/acceptance 明确 skeleton-only,阶段 3 才切换。
|
||||||
|
- [事件无限追加] → 所有业务循环有硬上限且 recursion limit=32。
|
||||||
|
|
||||||
|
## Migration Plan
|
||||||
|
|
||||||
|
1. 添加纯 Java 状态、router、actions、factory 和 trace types。
|
||||||
|
2. 先用 Fake actions 编译和执行完整 Graph。
|
||||||
|
3. 阶段 1 archive + commit 后,阶段 2 基于 ports 实现真实 nodes。
|
||||||
|
4. 如需回滚,revert 阶段 1 commit;当前生产路径不受影响。
|
||||||
|
|
||||||
|
## Open Questions
|
||||||
|
|
||||||
|
无。
|
||||||
+77
@@ -0,0 +1,77 @@
|
|||||||
|
# Chat Diagnosis StateGraph Routing Skeleton
|
||||||
|
|
||||||
|
## Why
|
||||||
|
|
||||||
|
阶段 0 已归档 ISS-011 的 StateGraph 设计基线,但仓库尚无可编译的 StateGraph 实现,也没有证据证明锁定的 Graph Core `1.1.2.0` 能按冻结条件边、追加事件策略和循环上限工作。阶段 1 需要用 Fake Node 建立不接真实模型、不会进入生产入口的路由骨架,为后续真实 Node 接入提供稳定内部接口。
|
||||||
|
|
||||||
|
## What Changes
|
||||||
|
|
||||||
|
- 新增 Diagnosis Graph State key、Node、route、status/verdict 常量和类型。
|
||||||
|
- 在现有 Spring AI Alibaba BOM 管理下显式声明 Graph Core 直接依赖,避免依赖 Agent Framework 的传递关系。
|
||||||
|
- 新增可注入 config-aware Node action ports 的 StateGraph 编译工厂。
|
||||||
|
- 实现阶段 0 冻结的完整条件边、三类技术重试上限、一次 evidence retry 和 recursion limit。
|
||||||
|
- 为 `orchestration_events` 配置 AppendStrategy,其余字段默认 ReplaceStrategy。
|
||||||
|
- 新增有界 event、transition 和 orchestration trace builder。
|
||||||
|
- 使用 Fake Node 单元测试覆盖正常、重试、失败、Gatekeeper、LOW_CONFID、Fallback、循环上限和 trace 顺序。
|
||||||
|
|
||||||
|
本 change 不接入 ChatService 或任何真实 Agent/Service。
|
||||||
|
|
||||||
|
## Capabilities
|
||||||
|
|
||||||
|
### New Capabilities
|
||||||
|
|
||||||
|
- `chat-diagnosis-stategraph-routing-skeleton`:提供未接生产入口的 StateGraph 状态、拓扑、路由、有限循环和编排 trace 构造能力。
|
||||||
|
|
||||||
|
### Modified Capabilities
|
||||||
|
|
||||||
|
- 无。阶段 0 设计基线保持不变;如果实现发现 API 冲突,必须先回写本 change,而不是静默改变基线语义。
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
### In Scope
|
||||||
|
|
||||||
|
- 包 `com.superbiz.agent.graph.diagnosis` 下的纯 Java Graph 骨架。
|
||||||
|
- `pom.xml` 中 Graph Core 的直接编译依赖;版本继续由现有 BOM 锁定为 1.1.2.0。
|
||||||
|
- Planner、Executor、Gatekeeper、Verified Input、Verifier、Evidence Retry、Composer、Fallback Node action ports。
|
||||||
|
- `orchestration_events` Append 策略和紧凑 trace builder。
|
||||||
|
- Fake Node route tests,使用 `RunnableConfig.threadId` 执行。
|
||||||
|
- recursion limit=32;业务计数仍是主要终止机制。
|
||||||
|
|
||||||
|
### Out of Scope
|
||||||
|
|
||||||
|
- ReactAgent Adapter、Prompt、ToolCallback、Skill 或模型调用。
|
||||||
|
- 调用 `ExecutorGatekeeperService` 或读取数据库 ToolInvocation。
|
||||||
|
- 修改 `ChatService`、`VerifierInputHook`、`VerifierContextHolder`。
|
||||||
|
- Flyway、`DiagnosisRun`、Trace DTO/API 或持久化。
|
||||||
|
- 删除旧 Sequential 实现或测试。
|
||||||
|
- Maven E2E、`logs/` 和数据库核验。
|
||||||
|
|
||||||
|
## Context Constraints
|
||||||
|
|
||||||
|
- 必须引用阶段 0 archive,不得改变其路由、安全或 run ownership 语义。
|
||||||
|
- 使用本地已验证的 Graph Core `1.1.2.0` API,不依赖未锁定版本示例。
|
||||||
|
- Graph action ports 使用 `AsyncNodeActionWithConfig`,为阶段 2 显式传递 run config 留出接口。
|
||||||
|
- Executor 不重试;Planner/Verifier/Composer 各最多一次技术重试;整个 Run 最多一次 evidence retry。
|
||||||
|
- Unknown/null status 必须 fail closed 到 Fallback。
|
||||||
|
- Fake nodes 只存在于测试代码,生产骨架不得包含模拟业务输出。
|
||||||
|
- 阶段 1 不改变任何外部 API 或当前生产路径。
|
||||||
|
|
||||||
|
## Acceptance
|
||||||
|
|
||||||
|
- Graph 可编译,默认 Replace、events Append,recursion limit 为 32。
|
||||||
|
- PASS 正常路径到 Composer/END,events 与 transitions 顺序一致。
|
||||||
|
- Planner INVALID_OUTPUT/RETRYABLE_FAILED 首次自重试,第二次或 NON_RETRYABLE_FAILED Fallback。
|
||||||
|
- Executor COMPLETED 进入 Gatekeeper;INVALID_OUTPUT/TOOL_BLOCKED/FAILED Fallback 且从不重试。
|
||||||
|
- Gatekeeper PASS 和有 binding 的 LOW_CONFID 进入 Verified Input;REJECT、unknown、零 binding LOW_CONFID Fallback。
|
||||||
|
- Verifier 技术失败首次自重试;PASS/REJECT 到 Composer;满足全部 guard 的 LOW_CONFID 只补证据一次;其余 LOW_CONFID 到 Composer。
|
||||||
|
- Composer 技术失败首次自重试,耗尽或不可重试进入 Fallback。
|
||||||
|
- Evidence retry 进入新 Planner 阶段时重置 planner retry count,且不影响其他技术计数。
|
||||||
|
- 每个测试路径能构造 bounded orchestration trace,无未定义条件边或无限循环。
|
||||||
|
- focused unit tests 和 Maven test-compile 通过;不运行 Maven E2E。
|
||||||
|
|
||||||
|
## Risks
|
||||||
|
|
||||||
|
- Graph Core 的 state merge/conditional-edge 细节可能与 API 签名表面不同;用真实 CompiledGraph Fake Node tests 验证。
|
||||||
|
- Node action 若忘记写 status/event,路由必须 fail closed,测试覆盖 null/unknown。
|
||||||
|
- 通用骨架若掺入真实 Agent 语义会污染阶段边界;本 change 只定义 ports 和确定性路由。
|
||||||
|
- 32 次 recursion limit 是保险,不替代显式计数。
|
||||||
+151
@@ -0,0 +1,151 @@
|
|||||||
|
## ADDED Requirements
|
||||||
|
|
||||||
|
### Requirement: Diagnosis Graph skeleton SHALL compile against the locked Graph API
|
||||||
|
|
||||||
|
The system SHALL provide an internal Diagnosis StateGraph skeleton using Graph Core 1.1.2.0 config-aware node actions, conditional edges, explicit key strategies, and a recursion limit of 32. The skeleton SHALL NOT be wired to the current production Chat path.
|
||||||
|
|
||||||
|
#### Scenario: Skeleton is compiled
|
||||||
|
|
||||||
|
- **WHEN** valid node action ports are supplied
|
||||||
|
- **THEN** the Graph SHALL compile with all frozen nodes and conditional edges
|
||||||
|
- **AND** `orchestration_events` SHALL use Append semantics while all other state uses Replace semantics
|
||||||
|
- **AND** the compiled Graph SHALL enforce recursion limit 32
|
||||||
|
|
||||||
|
#### Scenario: Run config is supplied
|
||||||
|
|
||||||
|
- **WHEN** a Fake Node Graph is invoked with a `RunnableConfig.threadId`
|
||||||
|
- **THEN** config-aware node ports SHALL receive that config
|
||||||
|
- **AND** the final state SHALL remain scoped to that invocation
|
||||||
|
|
||||||
|
#### Scenario: Production path is inspected
|
||||||
|
|
||||||
|
- **WHEN** stage 1 is accepted
|
||||||
|
- **THEN** ChatService, real Agents, Gatekeeper service, database, and Trace API SHALL NOT invoke the new skeleton
|
||||||
|
|
||||||
|
### Requirement: Planner routing SHALL allow one technical retry per Planner stage
|
||||||
|
|
||||||
|
The skeleton SHALL route Planner COMPLETED to Executor. INVALID_OUTPUT and RETRYABLE_FAILED SHALL self-retry only while `planner_retry_count=0`; NON_RETRYABLE_FAILED, unknown status, or a second technical failure SHALL route to Fallback.
|
||||||
|
|
||||||
|
#### Scenario: Planner first technical failure
|
||||||
|
|
||||||
|
- **WHEN** Planner first returns INVALID_OUTPUT or RETRYABLE_FAILED
|
||||||
|
- **THEN** only Planner SHALL run again
|
||||||
|
- **AND** the re-entered Planner state SHALL have `planner_retry_count=1`
|
||||||
|
|
||||||
|
#### Scenario: Planner retry is exhausted
|
||||||
|
|
||||||
|
- **WHEN** Planner returns a technical failure after retry count reaches 1
|
||||||
|
- **THEN** the Graph SHALL route to Fallback
|
||||||
|
- **AND** Executor SHALL NOT run
|
||||||
|
|
||||||
|
#### Scenario: Planner fails non-retryably
|
||||||
|
|
||||||
|
- **WHEN** Planner returns NON_RETRYABLE_FAILED or an unknown status
|
||||||
|
- **THEN** the Graph SHALL route directly to Fallback
|
||||||
|
|
||||||
|
### Requirement: Executor and Gatekeeper routing SHALL fail closed
|
||||||
|
|
||||||
|
Executor SHALL route only COMPLETED output to Gatekeeper and SHALL never retry. Gatekeeper SHALL route PASS and LOW_CONFID with verified bindings to Verified Input; REJECT, unknown status, or LOW_CONFID with zero bindings SHALL route to Fallback.
|
||||||
|
|
||||||
|
#### Scenario: Executor completes legal output
|
||||||
|
|
||||||
|
- **WHEN** Executor returns COMPLETED, including a legal no-evidence output or legal output after tool errors
|
||||||
|
- **THEN** Gatekeeper SHALL run exactly once
|
||||||
|
|
||||||
|
#### Scenario: Executor cannot complete contract
|
||||||
|
|
||||||
|
- **WHEN** Executor returns INVALID_OUTPUT, TOOL_BLOCKED, FAILED, or unknown status
|
||||||
|
- **THEN** the Graph SHALL route directly to Fallback
|
||||||
|
- **AND** Gatekeeper and Verifier SHALL NOT run
|
||||||
|
|
||||||
|
#### Scenario: Gatekeeper permits verification
|
||||||
|
|
||||||
|
- **WHEN** Gatekeeper returns PASS or LOW_CONFID with `verified_binding_count>0`
|
||||||
|
- **THEN** Verified Input SHALL run before Verifier
|
||||||
|
|
||||||
|
#### Scenario: Gatekeeper blocks verification
|
||||||
|
|
||||||
|
- **WHEN** Gatekeeper returns REJECT, unknown status, or LOW_CONFID with zero verified bindings
|
||||||
|
- **THEN** the Graph SHALL route to Fallback
|
||||||
|
- **AND** Verifier SHALL NOT run
|
||||||
|
|
||||||
|
### Requirement: Verifier routing SHALL separate technical retry from evidence retry
|
||||||
|
|
||||||
|
Verifier INVALID_OUTPUT and RETRYABLE_FAILED SHALL self-retry once with the same verified input. COMPLETED PASS or REJECT SHALL route to Composer. COMPLETED LOW_CONFID SHALL route to one Evidence Retry only when all frozen guards are true; otherwise it SHALL route to Composer. Other outcomes SHALL fail closed.
|
||||||
|
|
||||||
|
#### Scenario: Verifier first technical failure
|
||||||
|
|
||||||
|
- **WHEN** Verifier first returns INVALID_OUTPUT or RETRYABLE_FAILED
|
||||||
|
- **THEN** only Verifier SHALL run again
|
||||||
|
- **AND** Gatekeeper, Executor, and tools SHALL NOT rerun
|
||||||
|
|
||||||
|
#### Scenario: Verifier retry is exhausted
|
||||||
|
|
||||||
|
- **WHEN** Verifier returns a technical failure after `verifier_retry_count=1`
|
||||||
|
- **THEN** the Graph SHALL route to Fallback
|
||||||
|
|
||||||
|
#### Scenario: Verifier verdict reaches Composer
|
||||||
|
|
||||||
|
- **WHEN** Verifier completes with effective PASS or REJECT
|
||||||
|
- **THEN** Composer SHALL run
|
||||||
|
|
||||||
|
#### Scenario: LOW_CONFID qualifies for evidence retry
|
||||||
|
|
||||||
|
- **WHEN** Verifier completes LOW_CONFID with ceiling PASS, valid no_evidence or indirect_support facts, and `evidence_retry_count=0`
|
||||||
|
- **THEN** Evidence Retry SHALL run once and return to a new Planner stage
|
||||||
|
- **AND** `planner_retry_count` SHALL reset to 0
|
||||||
|
- **AND** `evidence_retry_count` SHALL become 1
|
||||||
|
|
||||||
|
#### Scenario: LOW_CONFID does not qualify for evidence retry
|
||||||
|
|
||||||
|
- **WHEN** ceiling is LOW_CONFID, facts contain no valid gap, or evidence retry count is already 1
|
||||||
|
- **THEN** Composer SHALL run without another Planner cycle
|
||||||
|
|
||||||
|
### Requirement: Composer routing SHALL allow one technical retry and then terminate safely
|
||||||
|
|
||||||
|
Composer COMPLETED SHALL terminate at END. INVALID_OUTPUT and RETRYABLE_FAILED SHALL self-retry once; NON_RETRYABLE_FAILED, unknown status, or a second technical failure SHALL route to Fallback and then END.
|
||||||
|
|
||||||
|
#### Scenario: Composer first technical failure
|
||||||
|
|
||||||
|
- **WHEN** Composer first returns INVALID_OUTPUT or RETRYABLE_FAILED
|
||||||
|
- **THEN** only Composer SHALL run again
|
||||||
|
- **AND** Verifier and preceding nodes SHALL NOT rerun
|
||||||
|
|
||||||
|
#### Scenario: Composer retry is exhausted
|
||||||
|
|
||||||
|
- **WHEN** Composer fails technically after `composer_retry_count=1`
|
||||||
|
- **THEN** Fallback SHALL run exactly once
|
||||||
|
- **AND** the Graph SHALL terminate
|
||||||
|
|
||||||
|
#### Scenario: Composer succeeds
|
||||||
|
|
||||||
|
- **WHEN** Composer returns COMPLETED
|
||||||
|
- **THEN** the Graph SHALL terminate without invoking Fallback
|
||||||
|
|
||||||
|
### Requirement: Orchestration events SHALL produce an exact bounded trace
|
||||||
|
|
||||||
|
Each node attempt SHALL append one terminal orchestration event. The trace builder SHALL derive transitions from adjacent events, final node and termination reason from the last event, degraded state from Fallback termination, and evidence retry count from state.
|
||||||
|
|
||||||
|
#### Scenario: Normal path trace is built
|
||||||
|
|
||||||
|
- **WHEN** Fake nodes execute Planner, Executor, Gatekeeper, Verified Input, Verifier, and Composer
|
||||||
|
- **THEN** events and transitions SHALL preserve that exact order
|
||||||
|
- **AND** the trace SHALL terminate at Composer without degradation
|
||||||
|
|
||||||
|
#### Scenario: Fallback path trace is built
|
||||||
|
|
||||||
|
- **WHEN** a route terminates through Fallback
|
||||||
|
- **THEN** the trace final node SHALL be Fallback
|
||||||
|
- **AND** `degraded` SHALL be true
|
||||||
|
- **AND** its termination reason SHALL come from the Fallback event
|
||||||
|
|
||||||
|
#### Scenario: Trace input is invalid
|
||||||
|
|
||||||
|
- **WHEN** no events exist or an event has blank node, outcome, reason code, or attempt below 1
|
||||||
|
- **THEN** trace construction SHALL fail explicitly rather than fabricate a path
|
||||||
|
|
||||||
|
#### Scenario: Event payload is inspected
|
||||||
|
|
||||||
|
- **WHEN** orchestration events and trace maps are produced
|
||||||
|
- **THEN** they SHALL contain only routing metadata
|
||||||
|
- **AND** they SHALL NOT contain Prompt, model reasoning, tool output, or Graph State snapshots
|
||||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user