feat(harness): add single diagnosis react agent
This commit is contained in:
@@ -159,6 +159,10 @@
|
|||||||
- 定义:围绕 Diagnosis Agent 提供确定性运行控制的边界,负责 Run、预算、取消、重试装配、Tool 调用记录、证据验真和最终释放,不承担业务诊断推理。
|
- 定义:围绕 Diagnosis Agent 提供确定性运行控制的边界,负责 Run、预算、取消、重试装配、Tool 调用记录、证据验真和最终释放,不承担业务诊断推理。
|
||||||
- 边界:Harness 不是工作流引擎,不实现 Planner/Executor/Composer 节点或自行编写 ReAct 循环。
|
- 边界:Harness 不是工作流引擎,不实现 Planner/Executor/Composer 节点或自行编写 ReAct 循环。
|
||||||
|
|
||||||
|
### Diagnosis Agent
|
||||||
|
- 定义:诊断链路中唯一拥有 ReAct 工具循环并生成 `DiagnosisDraft` 的 Agent,负责规划证据查询、判断证据充分性和撰写完整诊断草稿。
|
||||||
|
- 边界:不负责意图路由、Run/Session 生命周期、证据物理验真、独立语义审查或最终发布;证据不足时必须明确停止并保留限制。
|
||||||
|
|
||||||
### EvidenceGuard
|
### EvidenceGuard
|
||||||
- 定义:Harness 内部的确定性证据验真能力,校验 Draft 引用、当前 Run 所有权、Tool 调用状态和有界 Agent 投影。
|
- 定义:Harness 内部的确定性证据验真能力,校验 Draft 引用、当前 Run 所有权、Tool 调用状态和有界 Agent 投影。
|
||||||
- 边界:EvidenceGuard 不调用 LLM,也不判断证据是否足以推出业务结论。
|
- 边界:EvidenceGuard 不调用 LLM,也不判断证据是否足以推出业务结论。
|
||||||
|
|||||||
@@ -39,3 +39,4 @@
|
|||||||
| 2026-05-29 | chatmodel-abstraction | 抽象 ChatModel 和 EmbeddingModel,支持多模型路由。 | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | - | archived |
|
| 2026-05-29 | chatmodel-abstraction | 抽象 ChatModel 和 EmbeddingModel,支持多模型路由。 | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | - | archived |
|
||||||
| 2026-07-21 | single-react-rag-log-projections | RAG/log projection adapters through ToolBoundary | Harness/Tool projection | ISS-014, RAG, query_logs, projection, scope, redaction, MOCK, NO_EVIDENCE | openspec/changes/archive/2026-07-21-single-react-rag-log-projections | archived |
|
| 2026-07-21 | single-react-rag-log-projections | RAG/log projection adapters through ToolBoundary | Harness/Tool projection | ISS-014, RAG, query_logs, projection, scope, redaction, MOCK, NO_EVIDENCE | openspec/changes/archive/2026-07-21-single-react-rag-log-projections | archived |
|
||||||
| 2026-07-21 | single-react-mysql-readonly-tool | Fail-closed read-only MySQL evidence Tool with AST allowlist, JDBC controls and bounded projection | Harness/MySQL security | ISS-014, MySQL, JSqlParser, allowlist, PreparedStatement, timeout, projection | openspec/changes/archive/2026-07-21-single-react-mysql-readonly-tool | archived |
|
| 2026-07-21 | single-react-mysql-readonly-tool | Fail-closed read-only MySQL evidence Tool with AST allowlist, JDBC controls and bounded projection | Harness/MySQL security | ISS-014, MySQL, JSqlParser, allowlist, PreparedStatement, timeout, projection | openspec/changes/archive/2026-07-21-single-react-mysql-readonly-tool | archived |
|
||||||
|
| 2026-07-21 | single-react-diagnosis-agent | Single internal Diagnosis ReactAgent with Harness-controlled model/tool loop, bounded context and typed Draft | Harness/Diagnosis Agent/ReAct | ISS-014, ReactAgent, DiagnosisDraft, PreviousTurn, ToolInterceptor, ModelInterceptor, budget | openspec/changes/archive/2026-07-21-single-react-diagnosis-agent | archived |
|
||||||
|
|||||||
@@ -0,0 +1,65 @@
|
|||||||
|
# Single React Diagnosis Agent Acceptance
|
||||||
|
|
||||||
|
## 结果
|
||||||
|
|
||||||
|
已接受。阶段 4 Committed OpenSpec 的 11 项任务全部完成,新链路仅通过内部 Java/test 入口运行,公开 Chat 未切换。
|
||||||
|
|
||||||
|
## 验证
|
||||||
|
|
||||||
|
### 静态验证
|
||||||
|
|
||||||
|
- 检查:新 `harness.agent` 包与 Prompt 搜索 `ThreadLocal|SessionContextHolder|while|SequentialAgent|SupervisorAgent|StateGraph|Planner|Composer|Verifier|raw_response`。
|
||||||
|
- 结果:通过,无匹配。
|
||||||
|
- 检查:`git diff --name-only` 限定 Controller、ChatService、AiOpsService。
|
||||||
|
- 结果:通过,无 diff。
|
||||||
|
- 检查:OpenSpec artifact status、cross-artifact、L2 interface impact、`.committed`、`.archive-ready`。
|
||||||
|
- 结果:通过。
|
||||||
|
|
||||||
|
### 脚本验证
|
||||||
|
|
||||||
|
- 命令:`mvn -q -DskipTests compile`
|
||||||
|
- 结果:通过。
|
||||||
|
- 命令:`mvn -q '-Dtest=HarnessToolInterceptorTest,DiagnosisAgentUseCaseTest' test`
|
||||||
|
- 结果:通过,14 tests。
|
||||||
|
- 命令:`mvn -q '-Dtest=DiagnosisAgentUseCaseTest,HarnessToolInterceptorTest,DiagnosisHarnessCoreTest,ToolBoundaryTest,ToolAdapterTest,MysqlToolAdapterTest,CanonicalInvocationStoreTest,HarnessContractTest,RagToolContractTest,QueryLogsToolContractTest,MysqlToolContractTest' test`
|
||||||
|
- 结果:通过。
|
||||||
|
- 命令:`openspec validate single-react-diagnosis-agent --strict`
|
||||||
|
- 结果:通过。
|
||||||
|
- 命令:归档后 `openspec validate single-react-diagnosis-agent --type spec --strict`
|
||||||
|
- 结果:通过;仓库级 `openspec validate --specs --strict` 另发现前序 `mysql-readonly-tool`、`rag-log-projections` 主规格仍使用旧 scenario 格式,本阶段不跨归档边界修改。
|
||||||
|
|
||||||
|
### 浏览器/人工验证
|
||||||
|
|
||||||
|
- 不适用。本阶段无 UI、HTTP 或 SSE 行为变化。
|
||||||
|
|
||||||
|
### 未验证
|
||||||
|
|
||||||
|
- 未启动 Maven 应用、外部模型、MySQL、Redis、Milvus 或日志服务;按 ISS-014 门禁统一留到阶段 7 live E2E。
|
||||||
|
- 未验证 provider 生产响应始终携带 Token Usage;缺失 Usage 时模型调用次数仍强制,实际 Token 只能在 provider 返回时记录。
|
||||||
|
|
||||||
|
## 已完成范围
|
||||||
|
|
||||||
|
- 单一 Diagnosis ReactAgent、单 Prompt 和固定 Query/PreviousTurn 输入。
|
||||||
|
- 框架原生 tool loop、exact Tool Call ID 和三类 Harness adapter bridge。
|
||||||
|
- 模型/Token/Tool/上下文/Draft 预算与无自动 retry。
|
||||||
|
- 严格 DiagnosisDraft 输出和无证据停止规则。
|
||||||
|
- 显式 sessionId/runId audit metadata、可注入 AgentStep Hook、canonical Tool invocation。
|
||||||
|
- 公开 Chat/AiOps 与旧多 Agent 路径保持不变。
|
||||||
|
|
||||||
|
## 已知限制
|
||||||
|
|
||||||
|
- Draft 尚未经过 EvidenceGuard/SemanticGuard,不能公开发布;阶段 5 负责释放门禁。
|
||||||
|
- previous turn 由调用方传入,阶段 6A 才实现同 Session 安全 Run 选择和确定性组装。
|
||||||
|
- 真实业务 adapter/Spring Bean 装配和公开入口切换分别留给阶段 6A/6B。
|
||||||
|
- 仓库级全规格 strict validation 仍受阶段 3B/3C 主规格 scenario 标题格式阻塞,建议阶段 7 文档清理时统一修正并复验。
|
||||||
|
|
||||||
|
## Bug 修复和诊断
|
||||||
|
|
||||||
|
- 编译发现框架 `Interceptor.getName()` 必须实现,已补充稳定名称并回归。
|
||||||
|
- 首次测试修正了两个测试假设:callback 异常包装类型,以及 ToolResponseMessage 在 Prompt instructions 中的实际位置;生产规格与实现方向未变。
|
||||||
|
- 提交前自审发现默认 BeanOutputConverter schema 不允许 `conclusion=null`,与冻结的无证据契约冲突;已增加 schema post-process 和 focused regression,未放宽其他 Draft 字段。
|
||||||
|
|
||||||
|
## 交接
|
||||||
|
|
||||||
|
- 下一步:提交独立阶段 4 commit,再进入阶段 5。
|
||||||
|
- OpenSpec 归档确认:用户已对 ISS-014 每阶段 Apply、Archive 和 Git commit 提供持续授权;已归档至 `openspec/changes/archive/2026-07-21-single-react-diagnosis-agent`,主规格已同步至 `openspec/specs/single-react-diagnosis-agent/spec.md`。
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
# Single React Diagnosis Agent Brief
|
||||||
|
|
||||||
|
## 背景
|
||||||
|
|
||||||
|
- 用户目标:按 ISS-014 阶段 4 建立唯一的 Diagnosis ReAct Agent,并完成独立 sm-flow 归档与提交。
|
||||||
|
- 当前问题:Harness Core 和三类 evidence Tool 已就绪,但没有一个内部诊断执行链消费它们;公开 Chat 仍依赖旧的简单/多 Agent 路径。
|
||||||
|
- 关联 OpenSpec:`openspec/changes/archive/2026-07-21-single-react-diagnosis-agent/`
|
||||||
|
- devflow 分档:complex
|
||||||
|
- 需求真理源:`mvp/issues/active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md`,不重复创建独立 PRD。
|
||||||
|
|
||||||
|
## 范围
|
||||||
|
|
||||||
|
- 本次要做:单 Prompt、单 Diagnosis ReactAgent、固定 Query/PreviousTurn 输入、Harness Model/Tool interceptors、三类 Tool 注册、严格 DiagnosisDraft 解析、上下文/模型/Tool/Token 预算和内部审计注入。
|
||||||
|
- 本次不做:公开 Chat 切换、旧链路删除、EvidenceGuard、SemanticGuard、previous turn 选择、DiagnosisRun 持久化和 live E2E。
|
||||||
|
- 影响区域:`com.superbiz.agent.harness.agent`、`src/main/resources/prompts`、focused tests、OpenSpec/devflow/ISS-014 状态。
|
||||||
|
|
||||||
|
## OpenSpec 对齐
|
||||||
|
|
||||||
|
- proposal 覆盖状态:已覆盖
|
||||||
|
- specs 覆盖状态:已覆盖
|
||||||
|
- tasks 覆盖状态:已覆盖
|
||||||
@@ -0,0 +1,156 @@
|
|||||||
|
# Decisions: single-react-diagnosis-agent
|
||||||
|
|
||||||
|
## Discover status
|
||||||
|
|
||||||
|
- Checkpoint: Discover
|
||||||
|
- Capability source: `sm-flow`,使用 `grill-with-docs` 进行代码可证问题澄清,并使用 GitNexus、源码、依赖 sources jar 与 `javap` 核对调用链和框架 API。
|
||||||
|
- Scale: complex。变更新增内部 Agent/use case,跨越模型、Tool、预算、结构化输出和审计边界,但不切换公开协议。
|
||||||
|
- `devflow/index.md` 命中 design freeze、RunContext、ACI contracts、canonical store、RAG/log projection 和 MySQL Tool;没有与 OpenSpec 冲突的 ADR。
|
||||||
|
|
||||||
|
## Question Pool
|
||||||
|
|
||||||
|
| # | 维度 | 问题 | 模式 | 状态 |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| Q1 | 术语 | 单体 Diagnosis Agent 是否包含外层 Graph 或多个报告作者? | evidence-driven | 已解决 |
|
||||||
|
| Q2 | 边界 | 阶段 4 是否切换公开 Chat 或删除旧多 Agent 链路? | evidence-driven | 已解决 |
|
||||||
|
| Q3 | 验收 | 如何证明 ReAct 正常轮次不是 Harness retry? | evidence-driven | 已解决 |
|
||||||
|
| Q4 | 技术 | 如何取得并保留框架原始 tool_call_id? | evidence-driven | 已解决 |
|
||||||
|
| Q5 | 技术 | 如何在每个模型和 Tool 边界强制 Run 预算? | evidence-driven | 已解决 |
|
||||||
|
| Q6 | 输出 | 框架生成的 DiagnosisDraft schema 是否与可空 conclusion 契约一致,并会直接返回 DiagnosisDraft? | evidence-driven | 已解决 |
|
||||||
|
| Q7 | 上下文 | previous_turn 如何进入模型且保持固定 Schema 和有界? | evidence-driven | 已解决 |
|
||||||
|
| Q8 | 审计 | 如何保留 Run/AgentStep/ToolInvocation 而不引入 ThreadLocal? | evidence-driven | 已解决 |
|
||||||
|
| Q9 | 验收 | 无证据结果如何停止而不补造根因? | evidence-driven | 已解决 |
|
||||||
|
|
||||||
|
## Evidence-driven
|
||||||
|
|
||||||
|
| 结论 | 证据来源 | 是否已汇报用户 |
|
||||||
|
|---|---|---|
|
||||||
|
| 只新增一个拥有 tool loop 的 `ReactAgent`,无 SequentialAgent、SupervisorAgent 或业务 StateGraph。 | ISS-014 阶段 4;`ChatService` 旧链路反例;Spring AI Alibaba `ReactAgent` API | 已汇报 |
|
||||||
|
| 阶段 4 只提供内部用例,公开入口保持旧实现,接口影响 L2。 | ISS-014 1174-1197、阶段 0 design freeze、GitNexus `createReactAgent` 引用 | 已汇报 |
|
||||||
|
| `ToolCallRequest.getToolCallId()` 精确暴露 `AssistantMessage.ToolCall.id()`;ToolInterceptor 可在回调执行前取得该 ID。 | framework sources `ToolCallRequest`、`AgentToolNode` | 已汇报 |
|
||||||
|
| `ModelInterceptor` 包围每次真实模型调用,适合调用 `beforeModelCall` 并记录响应 Usage;ToolBoundary 已负责 Tool 预算,不能重复计数。 | framework `AgentLlmNode`/`InterceptorChain`;`ToolBoundary` | 已汇报 |
|
||||||
|
| BeanOutputConverter 可从 DiagnosisDraft 生成格式说明,但默认 schema 不允许冻结契约的 `conclusion=null`;必须 post-process nullable conclusion,且 `ReactAgent.call` 仍返回 AssistantMessage,需要 Harness 严格解析 JSON。 | framework `DefaultBuilder`、`ReactAgent`、本地生成 schema | 已汇报 |
|
||||||
|
| RunContext 可放入 `RunnableConfig` metadata,由框架控制地传播到 ToolInterceptor;不需要 ThreadLocal。 | `RunnableConfig.addMetadata(String,Object)`、`AgentToolNode` | 已汇报 |
|
||||||
|
| 现有 `AgentLoggingHook` 优先读取 config metadata 的 sessionId/runId,可作为可选审计 Hook 复用;ToolBoundary 保持 canonical Tool invocation。 | `AgentLoggingHook`、`ToolBoundary` | 已汇报 |
|
||||||
|
| 原始 Query 不截断;超限 fail closed。PreviousTurn 先按固定 record 序列化并受独立/总上下文字节限制。 | ISS-014 极简上下文与预算规则 | 已汇报 |
|
||||||
|
| `NO_EVIDENCE` 不是错误,只能形成限定范围的 NEGATIVE_OBSERVATION;Prompt 必须要求 conclusion=null、记录 missing_info 并停止。 | glossary、ACI spec、DiagnosisDraft contract | 已汇报 |
|
||||||
|
|
||||||
|
## User-interview
|
||||||
|
|
||||||
|
- 本阶段没有新增 user-interview 问题。方向、范围、阶段串行规则、框架 ID、状态语义以及 routine Apply/Archive/Commit 持续授权均已由用户在 ISS-014 评审与前序阶段确认。
|
||||||
|
|
||||||
|
## 关键取舍
|
||||||
|
|
||||||
|
- 决策:使用框架 `ToolInterceptor` 直接桥接已注册的 Tool 定义与 Harness adapter。
|
||||||
|
- 原因:Spring AI `ToolCallback` 的调用参数不包含 Tool Call ID,而 Alibaba interceptor 明确提供原始 ID 和运行 metadata。
|
||||||
|
- 影响:ToolCallback 负责模型可见定义,interceptor 负责受控执行;未知 Tool 仍交给框架 handler 并最终失败,不伪造结果。
|
||||||
|
- 决策:模型预算由 `ModelInterceptor` 执行,Tool 预算继续由 `ToolBoundary` 执行。
|
||||||
|
- 原因:避免在 Agent 层和 Tool boundary 双重 reserve。
|
||||||
|
- 决策:阶段 4 只严格反序列化 `DiagnosisDraft`,不提前实现 EvidenceGuard。
|
||||||
|
- 原因:字段引用真实性、唯一性和语义支持属于阶段 5;本阶段只验证 Agent 能生成冻结结构并携带框架 IDs。
|
||||||
|
- 决策:复用现有 Agent Hook 注入点,不复制 AgentStep 持久化实现。
|
||||||
|
- 原因:阶段 4 不接公开运行态,阶段 6A 再装配真实 Repository 和 Run 持久化。
|
||||||
|
|
||||||
|
## OpenSpec 回写
|
||||||
|
|
||||||
|
- 需进入 proposal/design/spec/tasks:单 Agent、内部入口、ToolInterceptor 原始 ID、ModelInterceptor 模型预算、ToolBoundary 单点 Tool 预算、严格 JSON 解析、上下文字节限制、审计 Hook 注入、无证据停止、不切换公开入口。
|
||||||
|
- 不创建 ADR:这些是 ISS-014 已冻结方向和当前阶段可逆的内部装配,不满足新的难逆转架构决策条件。
|
||||||
|
|
||||||
|
## Cross-artifact 对齐
|
||||||
|
|
||||||
|
| 上游 -> 下游 | 检查内容 | 状态 |
|
||||||
|
|---|---|---|
|
||||||
|
| brief/ISS-014 -> proposal | 单 Agent、内部入口、极简上下文、预算、审计、非目标和验收预期 | 已对齐 |
|
||||||
|
| proposal -> design | Tool/Model interceptor、严格 Draft、无重试、审计注入和公开隔离 | 已对齐 |
|
||||||
|
| design -> specs/tasks | ID 传播、预算单点、输入输出限制、生命周期所有权和风险缓解 | 已对齐 |
|
||||||
|
| specs -> tasks | 7 组可观察要求均有输入/Prompt、Tool bridge、Agent use case 和 focused test 切片 | 已对齐 |
|
||||||
|
|
||||||
|
## Architecture Audit
|
||||||
|
|
||||||
|
- 能力来源:`zoom-out`,使用 glossary 的 Diagnosis Agent、Diagnosis Harness、RunContext、Evidence Status 和 Invocation Status 术语。
|
||||||
|
- 链路为 `DiagnosisAgentInput + RunContext -> internal use case -> one ReactAgent -> model/tool interceptors -> ChatModel/ToolBoundary -> strict DiagnosisDraft`,没有外层业务 Graph。
|
||||||
|
- RunContext 拥有预算/取消/生命周期,ToolBoundary/store 拥有 canonical invocation,use case 拥有输入和 Draft 解析,Agent 只拥有诊断语义;数据所有权不重叠。
|
||||||
|
- 最大耦合风险是 Spring AI Alibaba interceptor/output schema API;设计通过本地 1.1.2.0 sources jar 和实际 schema 生成结果证实,并以 focused framework-loop/schema tests 固定。
|
||||||
|
- 最大阶段风险是 Draft 在 Guards 前被误用;类名、文档和零 Controller 消费者共同保持 internal/draft 边界,阶段 5 前不得发布。
|
||||||
|
|
||||||
|
## 接口影响
|
||||||
|
|
||||||
|
- 级别:L2 内部接口。
|
||||||
|
- 变更对象:新增内部 Java use case/factory/input/limits/interceptors/registry,不修改既有方法签名。
|
||||||
|
- 消费者:本阶段只有 focused tests;阶段 6A 将成为首个生产消费者。
|
||||||
|
- 兼容性:公开 HTTP/SSE、Controller DTO、数据库和旧 Chat/AiOps 路径不变,无迁移或回滚要求。
|
||||||
|
|
||||||
|
## Pre-apply Research
|
||||||
|
|
||||||
|
### 参考实现与框架源码
|
||||||
|
|
||||||
|
- `src/main/java/com/superbiz/agent/service/ChatService.java`:`createReactAgent`、`buildChatExecutorAgent` 和旧多 Agent 反例;只复用 Builder 形态,不复用运行职责。
|
||||||
|
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`:scripted `ChatModel` 测试模式。
|
||||||
|
- `src/main/java/com/superbiz/agent/harness/core/DiagnosisHarnessCore.java`:模型/Tool/Token/容量边界。
|
||||||
|
- `src/main/java/com/superbiz/agent/harness/tool/boundary/ToolBoundary.java`:Tool budget 和 canonical record 单点。
|
||||||
|
- `src/main/java/com/superbiz/agent/harness/tool/adapter/RagToolAdapter.java`、`QueryLogsToolAdapter.java`、`MysqlToolAdapter.java`:阶段 4 唯一允许的 evidence Tool 执行入口。
|
||||||
|
- `src/main/java/com/superbiz/agent/hook/AgentLoggingHook.java`:可注入 AgentStep audit,优先读取 RunnableConfig metadata。
|
||||||
|
- Spring AI Alibaba 1.1.2.0 sources:`ReactAgent`、`DefaultBuilder`、`AgentLlmNode`、`AgentToolNode`、`ToolCallRequest`、`InterceptorChain`。
|
||||||
|
|
||||||
|
### 技术栈清单
|
||||||
|
|
||||||
|
- `ReactAgent.builder()` + 从 `DiagnosisDraft` 生成并修正 nullable conclusion 的 `.outputSchema(...)`,不新建 Graph/SequentialAgent。
|
||||||
|
- `RunnableConfig.addMetadata(String,Object)` 显式携带 sessionId/runId/RunContext,并设置 `_stream_=false`。
|
||||||
|
- `ModelInterceptor` 对每次模型调用执行 Core budget;读取 `ChatResponseMetadata.Usage`。
|
||||||
|
- `ToolInterceptor` 读取 exact framework ID,调用 adapter bridge;`ToolBoundary` 保留唯一 Tool reserve/store 边界。
|
||||||
|
- Spring `FunctionToolCallback` 仅定义模型可见 Tool Schema/description;直接 callback 执行 fail closed。
|
||||||
|
- Jackson `ObjectMapper` 负责固定输入 JSON 和严格 DiagnosisDraft 反序列化;UTF-8 字节按 `StandardCharsets.UTF_8` 计算。
|
||||||
|
- 框架 Hook 列表作为 AgentStep audit 扩展点;阶段 4 不新增 JPA/Redis/Flyway。
|
||||||
|
|
||||||
|
### 新建基础设施
|
||||||
|
|
||||||
|
- `harness.agent`:Input、Limits、Tool registry/interceptor、Model interceptor、Factory、UseCase 和异常类型。
|
||||||
|
- `prompts/diagnosis-agent-prompt.md`:唯一 Diagnosis Agent Prompt。
|
||||||
|
- focused scripted model/tool-loop tests;无需新 Maven 依赖。
|
||||||
|
|
||||||
|
### ISS-014 PRD 复用
|
||||||
|
|
||||||
|
- 阶段 4 不新建独立 `prd.md`:ISS-014 已完整覆盖问题、用户价值、数据流、Draft Schema、预算、重试、阶段边界和验收;`brief.md` 只索引本切片,不复制总 Issue。
|
||||||
|
|
||||||
|
## Commit Gate Preflight
|
||||||
|
|
||||||
|
- proposal、design、specs、tasks 完整,`openspec status` 为 complete,`openspec validate single-react-diagnosis-agent --strict` 通过。
|
||||||
|
- question pool 全部为已汇报的 evidence-driven 结论;无未确认 user-interview、接口等级或风险接受问题。
|
||||||
|
- Cross-artifact 四段对齐无 gap;架构审计的数据所有权、阶段边界和框架耦合缓解已进入 design/spec/tasks。
|
||||||
|
- 接口影响为 L2,仅新增内部 Java API;公开 Controller/SSE/JPA/旧 ChatService 保持不变。
|
||||||
|
- Apply、Archive 和阶段 Git commit 使用用户对 ISS-014 各阶段的持续授权;实现必须严格限制为 Committed OpenSpec。
|
||||||
|
- `.committed` 已创建,Committed OpenSpec 可进入 Apply。
|
||||||
|
|
||||||
|
## Apply Progress
|
||||||
|
|
||||||
|
- 首模块对齐:Input/Limits/Prompt/Tool registry/Tool interceptor 已落地,任务 1.2、2.1、2.2 完成;1.1 等待 focused test 后完成。
|
||||||
|
- 编译首次发现 `ToolInterceptor` 必须实现 `Interceptor.getName()`;分类为代码偏离,已补充稳定名称并复编译通过,无需修改 OpenSpec。
|
||||||
|
- 单 Agent use case、Model interceptor、strict Draft parser 和 audit Hook 注入已完成;scripted ChatModel 真实执行框架 `model -> Tool -> model` loop。
|
||||||
|
- 首次测试的两处失败均为测试假设偏差:Spring 会将 callback 异常包装为 `ToolExecutionException`,Tool observation 位于 `Prompt.instructions` 的 `ToolResponseMessage` 而非 `Prompt.getContents()`;已按框架真实 API 修正测试,生产设计未变。
|
||||||
|
- 阶段 4 focused tests 共 14 个通过,任务 1.1-4.2 完成;剩余任务 4.3 为综合回归、静态范围和 OpenSpec 验证。
|
||||||
|
|
||||||
|
## Apply Result
|
||||||
|
|
||||||
|
- 新增一个且仅一个 `diagnosis_agent` Factory 和内部 `DiagnosisAgentUseCase`;没有外层业务 Graph、SequentialAgent、SupervisorAgent 或手写 ReAct loop。
|
||||||
|
- 新增单一 Prompt、固定 `DiagnosisAgentInput(query, previous_turn)`、可配置 UTF-8 限制和严格 `DiagnosisDraft` 解析;不接受完整历史,不执行结构修复或 Agent retry。
|
||||||
|
- 新增 `HarnessModelInterceptor`,对每个非流式模型轮次执行 Core model/Token budget,并在 late result 返回后复查 Run active 状态。
|
||||||
|
- 新增 `HarnessEvidenceTools` 和 `HarnessToolInterceptor`,注册三类冻结 Tool,精确传播框架 Tool Call ID,并通过阶段 3B/3C adapter/ToolBoundary 返回有界结果。
|
||||||
|
- AgentStep 使用可注入框架 Hook 保留,RunnableConfig 显式传播 sessionId/runId;Run 和 Tool canonical invocation 继续由既有 Harness 边界拥有。
|
||||||
|
- 公开 Controller、ChatService、AiOpsService 无 diff;旧多 Agent 主链路继续保留。
|
||||||
|
|
||||||
|
## Apply Verification
|
||||||
|
|
||||||
|
- 编译:`mvn -q -DskipTests compile` 通过。
|
||||||
|
- 阶段 4 focused:`mvn -q '-Dtest=HarnessToolInterceptorTest,DiagnosisAgentUseCaseTest' test`,14 tests 通过。
|
||||||
|
- 综合回归:`mvn -q '-Dtest=DiagnosisAgentUseCaseTest,HarnessToolInterceptorTest,DiagnosisHarnessCoreTest,ToolBoundaryTest,ToolAdapterTest,MysqlToolAdapterTest,CanonicalInvocationStoreTest,HarnessContractTest,RagToolContractTest,QueryLogsToolContractTest,MysqlToolContractTest' test` 通过。
|
||||||
|
- OpenSpec:`openspec validate single-react-diagnosis-agent --strict` 通过。
|
||||||
|
- 静态范围:新 Agent 包无 ThreadLocal、手写 while、外层 Graph、多 Agent 类型或 raw response;Controller/ChatService/AiOpsService diff 为空。
|
||||||
|
- 未执行 live E2E:按 ISS-014 阶段门禁统一留到阶段 7。
|
||||||
|
|
||||||
|
## Archive Result
|
||||||
|
|
||||||
|
- 11/11 OpenSpec tasks 完成,`.archive-ready` 已创建。
|
||||||
|
- 主规格已同步至 `openspec/specs/single-react-diagnosis-agent/spec.md`。
|
||||||
|
- Change 已归档至 `openspec/changes/archive/2026-07-21-single-react-diagnosis-agent`。
|
||||||
|
- `devflow/index.md` 和 ISS-014 阶段表已更新为阶段 0-4 archived,下一阶段为 5。
|
||||||
|
- 提交前 schema 自审发现框架默认 BeanOutputConverter 将 `conclusion` 限为 object,与冻结的无证据 `null` 语义冲突;分类为实现设计遗漏,已回写 archived design,新增 nullable schema post-process 与回归测试,规格行为未改变。
|
||||||
@@ -0,0 +1,29 @@
|
|||||||
|
# Single React Diagnosis Agent Evidence
|
||||||
|
|
||||||
|
## 代码与文档证据
|
||||||
|
|
||||||
|
| 来源 | 证据 | 结论 | 是否已汇报 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `mvp/issues/active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md` | 阶段 4 明确单 Agent、内部入口、极简上下文、结构化 Draft、无重试和预算验收 | 本 change 不得切换公开入口或删除旧链路 | 是 |
|
||||||
|
| `ChatService.java` + GitNexus references | `createReactAgent` 被旧策略入口调用;复杂路径仍创建 Planner/Executor/Verifier/Composer | 新实现必须是独立内部 use case,不能复用旧 Service 运行职责 | 是 |
|
||||||
|
| Spring AI Alibaba 1.1.2.0 `ReactAgent`/`AgentLlmNode` sources | 框架自带 ReAct loop;非流式 `ModelResponse` 保留 `ChatResponse` Usage | 不手写循环,使用 ModelInterceptor 强制模型/Token 预算 | 是 |
|
||||||
|
| Spring AI Alibaba `ToolCallRequest`/`AgentToolNode` sources | `ToolCallRequest.getToolCallId()` 来自 `AssistantMessage.ToolCall.id()`,Tool interceptor 在 callback 前执行 | 可以精确传播框架 ID,不生成第二套 ID | 是 |
|
||||||
|
| Spring AI Alibaba `DefaultBuilder` + 本地 BeanOutputConverter 输出 | 默认 schema 只允许 object conclusion,但冻结契约允许无证据时 conclusion=null | 生成 schema 必须 post-process nullable conclusion,UseCase 仍严格反序列化 DiagnosisDraft | 是 |
|
||||||
|
| `DiagnosisHarnessCore.java` / `ToolBoundary.java` | Core 提供 model/Token/capacity;ToolBoundary 已负责 Tool reserve/store | Agent 层只 reserve model/context,禁止双扣 Tool budget | 是 |
|
||||||
|
| 阶段 3B/3C adapters | RAG/log/MySQL 都以 `RunContext + ToolCallRequestEnvelope` 进入 ToolBoundary | Tool registry 可直接桥接既有 adapter,不复制投影或安全策略 | 是 |
|
||||||
|
| `AgentLoggingHook.java` | 优先从 RunnableConfig metadata 读取 sessionId/runId | Factory 可复用现有 Hook 注入点,不依赖 ThreadLocal | 是 |
|
||||||
|
|
||||||
|
## Evidence-driven 结论
|
||||||
|
|
||||||
|
- 单 Agent 的可验证边界是一个 `ReactAgent.call`,内部允许正常模型/Tool轮次,但外部没有 Agent retry 或第二个报告作者。
|
||||||
|
- ToolCallback 不直接提供 Tool Call ID,必须使用 Alibaba `ToolInterceptor`;已由真实 scripted framework loop 证明 ID 进入 ToolResponseMessage。
|
||||||
|
- `PreviousTurn` 只作为固定上下文,当前 Draft 只允许引用当前 Run Tool result 中的 ID。
|
||||||
|
- `NO_EVIDENCE` 必须以 `conclusion=null`、限定范围的负向观察和 missing info 结束,不能推导健康或排除根因。
|
||||||
|
- 阶段 4 的 L2 内部 API 不影响公开消费者;阶段 5/6A/6B 前 Draft 不可发布。
|
||||||
|
|
||||||
|
## 实现证据
|
||||||
|
|
||||||
|
- 新包 `com.superbiz.agent.harness.agent` 包含 Input/Limits、Prompt loader、Tool registry/interceptor、Model interceptor、Factory 和 UseCase。
|
||||||
|
- `diagnosis-agent-prompt.md` 是唯一新 Diagnosis Agent Prompt,包含 Tool、证据绑定、停止和精确 JSON 规则。
|
||||||
|
- 14 个阶段 4 tests 覆盖框架 model -> Tool -> model、ID、错误、未知 Tool、PreviousTurn、审计 metadata、nullable conclusion schema、no-evidence、解析和预算。
|
||||||
|
- 综合回归同时覆盖 Harness Core、ToolBoundary、3B/3C adapter、canonical store 和冻结 contracts。
|
||||||
@@ -1,6 +1,6 @@
|
|||||||
# ISS-014 单体 ReAct Agent、Harness 与 ACI 工具瘦身
|
# ISS-014 单体 ReAct Agent、Harness 与 ACI 工具瘦身
|
||||||
|
|
||||||
**状态**:实施中(阶段 0-3C 已归档,下一阶段 4)
|
**状态**:实施中(阶段 0-4 已归档,下一阶段 5)
|
||||||
**严重程度**:高
|
**严重程度**:高
|
||||||
**发现时间**:2026-07-20
|
**发现时间**:2026-07-20
|
||||||
**目标分支**:`refactor/chat-single-react-harness`
|
**目标分支**:`refactor/chat-single-react-harness`
|
||||||
@@ -1044,7 +1044,7 @@ ISS-014 是总设计 Issue,不创建跨阶段共享的 OpenSpec change。以
|
|||||||
| 3A | `single-react-tool-invocation-store` | Completed;已归档 |
|
| 3A | `single-react-tool-invocation-store` | Completed;已归档 |
|
||||||
| 3B | `single-react-rag-log-projections` | Completed;已归档 |
|
| 3B | `single-react-rag-log-projections` | Completed;已归档 |
|
||||||
| 3C | `single-react-mysql-readonly-tool` | Completed;已归档 |
|
| 3C | `single-react-mysql-readonly-tool` | Completed;已归档 |
|
||||||
| 4 | `single-react-diagnosis-agent` | Pending |
|
| 4 | `single-react-diagnosis-agent` | Completed;已归档 |
|
||||||
| 5 | `single-react-evidence-semantic-guards` | Pending |
|
| 5 | `single-react-evidence-semantic-guards` | Pending |
|
||||||
| 6A | `single-react-chat-application-usecase` | Pending |
|
| 6A | `single-react-chat-application-usecase` | Pending |
|
||||||
| 6B | `single-react-chat-sse-cutover` | Pending |
|
| 6B | `single-react-chat-sse-cutover` | Pending |
|
||||||
|
|||||||
@@ -0,0 +1 @@
|
|||||||
|
ready
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
committed
|
||||||
@@ -0,0 +1,2 @@
|
|||||||
|
schema: spec-driven
|
||||||
|
created: 2026-07-21
|
||||||
@@ -0,0 +1,106 @@
|
|||||||
|
## Context
|
||||||
|
|
||||||
|
阶段 2 的 `DiagnosisHarnessCore` 已提供显式 `RunContext`、模型/Tool/Token/字节预算、取消和唯一生命周期;阶段 3A-3C 已提供 canonical invocation、`ToolBoundary` 以及 RAG、日志、MySQL adapter。当前这些组件没有生产 Agent 消费者,公开 Chat 仍由 `ChatService` 创建单 Agent 简单路径或 Planner/Executor/Verifier/Composer 多 Agent 路径。
|
||||||
|
|
||||||
|
Spring AI Alibaba 1.1.2.0 的 `ReactAgent` 自带模型与 Tool 交替执行的 ReAct loop。`ModelInterceptor` 包围每次模型调用;`ToolInterceptor` 获得 `ToolCallRequest.getToolCallId()`、参数和 `RunnableConfig` metadata;`BeanOutputConverter` 可从输出类型生成格式提示,但 `call` 最终仍返回 `AssistantMessage`。因此本阶段可以使用框架扩展点接入 Harness,无需自行编写循环或修改框架。
|
||||||
|
|
||||||
|
## Goals / Non-Goals
|
||||||
|
|
||||||
|
**Goals:**
|
||||||
|
|
||||||
|
- 创建一个且仅一个拥有 Tool loop 的 Diagnosis ReactAgent。
|
||||||
|
- 以显式 `RunContext` 执行当前 Query 和可选 `PreviousTurn`,输入失败时 fail closed,不截断原始 Query。
|
||||||
|
- 将框架原始 Tool Call ID 传给阶段 3B/3C adapter,并只把有界 `agent_result` 返回模型。
|
||||||
|
- 对每次模型调用、Token、Tool 调用和输入/输出上下文字节实施确定性预算。
|
||||||
|
- 输出可严格反序列化的 `DiagnosisDraft`;无证据时明确停止,不补造结论。
|
||||||
|
- 保留可注入 AgentStep Hook、RunContext 生命周期和 canonical Tool invocation 审计边界。
|
||||||
|
|
||||||
|
**Non-Goals:**
|
||||||
|
|
||||||
|
- 不实现 EvidenceGuard、SemanticGuard、结构修复、引用展开或最终发布。
|
||||||
|
- 不创建 Intent Router、不选择 previous turn、不持久化 DiagnosisRun。
|
||||||
|
- 不修改 Controller、SSE、ChatService、AiOpsService 或旧多 Agent 链路。
|
||||||
|
- 不实现外层 Graph、手写 ReAct while、Agent 自动重跑或 Tool 自动重试。
|
||||||
|
- 不增加真实日志源或新的 evidence Tool。
|
||||||
|
|
||||||
|
## Decisions
|
||||||
|
|
||||||
|
### One per-run ReactAgent uses the framework loop
|
||||||
|
|
||||||
|
`DiagnosisAgentFactory` 为每个 `RunContext` 构建一个名为 `diagnosis_agent` 的 `ReactAgent`,关闭并行 Tool 执行并使用单一 Prompt。内部 `DiagnosisAgentUseCase` 对 Agent 只调用一次;一次调用内由框架决定正常的模型/Tool轮次。
|
||||||
|
|
||||||
|
替代方案是复用 `ChatService.createReactAgent`,但它注册旧 Tool、旧 Hook 和完整旧运行职责,无法保证 Harness boundary、极简上下文和不接公开入口。另一个替代方案是手写 while,直接违反 ISS-014。
|
||||||
|
|
||||||
|
### Model and Tool control use different framework boundaries
|
||||||
|
|
||||||
|
`HarnessModelInterceptor` 在每次模型调用前执行 `core.beforeModelCall(context)`,调用完成后读取 `ChatResponseMetadata.Usage` 并执行 `core.recordTokens`。内部用例设置 `_stream_=false`,保证 interceptor 可取得完整 `ChatResponse` 和 Token Usage。
|
||||||
|
|
||||||
|
`HarnessToolInterceptor` 读取框架原始 Tool Call ID,并按 Tool 名调用 `HarnessEvidenceTools` 中的 adapter bridge。bridge 创建 `ToolCallRequestEnvelope(runId, frameworkId, toolName, arguments, true, true)`;授权来自注册到当前 Diagnosis Agent 的固定 Tool 集合。Tool 成功时只返回 `agent_result`,失败时只返回稳定的 `evidence_status/tool_call_id/error_code`,不返回 raw response、内部异常或 invocation lifecycle。
|
||||||
|
|
||||||
|
Tool 调用预算只在 `ToolBoundary` 中 reserve;Agent interceptor 不重复调用 `beforeToolCall`。这保持阶段 3A 的 canonical record 与预算原子边界。
|
||||||
|
|
||||||
|
### Tool definitions and execution are registered together
|
||||||
|
|
||||||
|
`HarnessEvidenceTools` 固定暴露 `lookup_knowledge`、`query_logs`、`query_mysql` 三个 Spring `ToolCallback` 定义,输入类型分别复用已冻结的 Request records,描述复用 `AgentToolContracts`。callback 本体不允许绕过 interceptor 直接执行;实际执行映射与定义在同一 registry 中,构造时拒绝缺项或重复项。
|
||||||
|
|
||||||
|
替代方案是在 `ToolCallback.call` 中读取 ThreadLocal 或生成 ID,都会违反 RunContext 和 canonical ID 契约。
|
||||||
|
|
||||||
|
### Input and output remain typed and bounded
|
||||||
|
|
||||||
|
`DiagnosisAgentInput` 只包含非空原始 `query` 和可选 `PreviousTurn`。`DiagnosisAgentLimits` 配置 query、previous turn、总输入和 Draft 的 UTF-8 字节上限。内部用例先分别验证,再将固定 `{query, previous_turn}` JSON 作为唯一 User 输入,并通过 `core.reserveRunBytes` 计入 Run 容量;任何超限都在模型调用前失败,不做语义截断。
|
||||||
|
|
||||||
|
Agent 使用基于 `BeanOutputConverter<DiagnosisDraft>` 生成的格式提示,并通过 schema post-process 明确允许冻结契约中的 `conclusion=null`;Factory 以 `.outputSchema(...)` 注入该格式。最终文本先检查 UTF-8 上限并计入容量,再由注入的 `ObjectMapper` 严格解析为冻结 record。Markdown fence、前后说明、未知结构或空输出不做修复和自动重试;阶段 5 再实现一次显式无 Tool 结构修复。
|
||||||
|
|
||||||
|
### Previous turn is context, not current evidence
|
||||||
|
|
||||||
|
`PreviousTurn` 按固定 record 原样序列化,仅帮助理解追问。Prompt 明确禁止引用旧 Tool Call ID 或把 previous turn 当作当前 Run 的 evidence;当前 `DiagnosisDraft.analysis[*].tool_call_ids` 只能来自本 Run Tool result。
|
||||||
|
|
||||||
|
### Prompt owns diagnosis semantics, Harness owns control
|
||||||
|
|
||||||
|
唯一 classpath Prompt 要求结论先行、所有分析绑定 Tool Call IDs、`NO_EVIDENCE` 只能形成限定范围的 `NEGATIVE_OBSERVATION`。没有足够 `EVIDENCE_FOUND` 时 `conclusion=null`,在 `limitations` 记录范围和缺失信息并结束,不把 `NO_EVIDENCE` 推导为系统健康或根因排除。
|
||||||
|
|
||||||
|
阶段 4 不实现物理或语义 Guard,因此内部返回值叫 Draft,不能直接公开发布。
|
||||||
|
|
||||||
|
### Audit remains injectable and explicit
|
||||||
|
|
||||||
|
Factory 接收可选框架 `Hook` 列表。阶段 6A 可以注入现有 `AgentLoggingHook`;内部用例始终在 `RunnableConfig` metadata 写入 sessionId/runId,所以该 Hook 不需要读取 ThreadLocal fallback。模型/Tool 预算和生命周期仍绑定 RunContext,Tool canonical audit 仍由 `ToolBoundary`/store 负责。
|
||||||
|
|
||||||
|
## Architecture And Interface Impact
|
||||||
|
|
||||||
|
```text
|
||||||
|
DiagnosisAgentInput(query, previousTurn) + RunContext
|
||||||
|
-> DiagnosisAgentUseCase: validate/serialize/reserve context bytes
|
||||||
|
-> DiagnosisAgentFactory: one ReactAgent + one Prompt
|
||||||
|
-> HarnessModelInterceptor -> DiagnosisHarnessCore -> ChatModel
|
||||||
|
-> framework ReAct loop
|
||||||
|
-> HarnessToolInterceptor -> HarnessEvidenceTools -> 3B/3C adapters
|
||||||
|
-> ToolBoundary -> canonical invocation store
|
||||||
|
-> bounded AssistantMessage text
|
||||||
|
-> strict ObjectMapper -> DiagnosisDraft
|
||||||
|
```
|
||||||
|
|
||||||
|
- 数据所有权:use case 拥有本次输入和 Draft 解析;RunContext 拥有预算/取消/生命周期;ToolBoundary/store 拥有调用记录;Agent 只拥有诊断语义。
|
||||||
|
- 接口影响:L2 内部接口。新增阶段 6A 可消费的 Java 类型,不改变公开 HTTP/SSE、JPA、数据库或旧 Service 方法。
|
||||||
|
- 生命周期:本 use case 不把成功 Draft 标记为 Run SUCCESS,因为阶段 5 Guards 尚未执行;预算、取消和超时可由 Core 先行终止 Run,最终持久化映射留给阶段 6A。
|
||||||
|
|
||||||
|
## Risks / Trade-offs
|
||||||
|
|
||||||
|
- [Risk] Provider 返回 fenced JSON 或附加说明导致解析失败 → Mitigation:阶段 4 fail closed 且不重跑;测试固定这一行为,阶段 5 才允许一次无 Tool 结构修复。
|
||||||
|
- [Risk] Token Usage 缺失或为 null → Mitigation:模型调用次数仍被强制;仅在可用且非负时记录实际 Token,并在 acceptance 标记 provider 元数据依赖。
|
||||||
|
- [Risk] Tool interceptor 错误泄露内部异常 → Mitigation:只返回 allowlisted error code,不返回 message/stack/raw response。
|
||||||
|
- [Risk] 输入和 Tool 投影共同消耗 Run bytes,配置过小会过早耗尽 → Mitigation:所有限制可配置,usage 可观察,不在本阶段硬编码生产校准值。
|
||||||
|
- [Risk] 复用旧 AgentLoggingHook 仍包含 ThreadLocal fallback → Mitigation:新链路总是提供 metadata;阶段 7 物理清理前不修改旧 Hook,focused test 验证相同 sessionId/runId 传播。
|
||||||
|
- [Risk] ReactAgent 内部基于框架 StateGraph → Mitigation:这是框架原生实现细节;本项目不再包一层业务 Graph 或多 Agent workflow。
|
||||||
|
|
||||||
|
## Migration Plan
|
||||||
|
|
||||||
|
1. 新增内部 Agent 类型、Prompt、interceptors 和 Tool registry,不注册 Controller bean。
|
||||||
|
2. 使用 scripted ChatModel 和 fake/adapted Tool 完成内部 tool-loop、output、budget、no-evidence 和审计测试。
|
||||||
|
3. 保持公开 Chat 与旧多 Agent 代码 diff 为空,Archive 并提交阶段 4。
|
||||||
|
4. 阶段 5 在内部 Draft 后增加 Guards;阶段 6A 再创建运行应用用例并装配真实审计/previous turn;阶段 6B 原子切换公开入口。
|
||||||
|
|
||||||
|
Rollback:删除新增内部包、Prompt 和测试即可;不存在公开协议、数据迁移或运行时切换。
|
||||||
|
|
||||||
|
## Open Questions
|
||||||
|
|
||||||
|
无。生产预算默认值、公开切换和最终发布策略分别由配置校准、阶段 6B 和阶段 5 处理。
|
||||||
@@ -0,0 +1,31 @@
|
|||||||
|
## Why
|
||||||
|
|
||||||
|
阶段 2-3C 已建立显式 RunContext、预算、canonical Tool invocation 和三类有界 evidence Tool,但还没有使用这些边界的诊断执行者。需要在不影响公开 Chat 的前提下建立单一 Diagnosis ReAct Agent,证明框架原生 tool loop、冻结的 DiagnosisDraft 和 Harness 控制能够闭环,再进入阶段 5 的释放门禁。
|
||||||
|
|
||||||
|
## What Changes
|
||||||
|
|
||||||
|
- 新增唯一的 Diagnosis Agent Prompt,把诊断规划、证据查询、证据充分性判断和报告草稿生成合并到一个 ReactAgent。
|
||||||
|
- 新增独立内部 Diagnosis 用例,显式接收 RunContext、当前原始 Query 和可选的固定 PreviousTurn,不读取完整 Session 历史。
|
||||||
|
- 通过 Spring AI Alibaba 的 ModelInterceptor 和 ToolInterceptor 接入 Harness;使用框架原生 ReAct/tool loop 和原始 tool_call_id,不手写循环或第二套调用 ID。
|
||||||
|
- 将 lookup_knowledge、query_logs 和 query_mysql 的冻结 Tool 定义连接到阶段 3B/3C adapter,只向 Agent 返回有界 agent_result。
|
||||||
|
- 使用 DiagnosisDraft 作为结构化输出契约,并对输入上下文、模型轮次、Token、Tool 调用和输出字节执行确定性预算。
|
||||||
|
- 允许注入现有 AgentStep Hook,Run 审计继续由 RunContext/lifecycle 承载,ToolInvocation 审计继续由 canonical Tool boundary 承载。
|
||||||
|
- 证据不足时要求 Agent 输出 conclusion=null、明确 limitations 并结束当前 ReAct 执行,不自动重跑 Diagnosis Agent、模型调用或 Tool 调用。
|
||||||
|
- 不修改 ChatController、公开 /api/chat 或 /api/chat_stream,不删除旧 Planner/Executor/Verifier/Composer 链路。
|
||||||
|
|
||||||
|
## Capabilities
|
||||||
|
|
||||||
|
### New Capabilities
|
||||||
|
|
||||||
|
- `single-react-diagnosis-agent`: 定义单一 Diagnosis ReactAgent 的内部输入、Harness-controlled ReAct/tool loop、结构化 Draft、预算、无证据停止和审计边界。
|
||||||
|
|
||||||
|
### Modified Capabilities
|
||||||
|
|
||||||
|
- None. 既有 Harness、Tool 和公开 Chat capability 的需求语义不在本阶段改变。
|
||||||
|
|
||||||
|
## Impact
|
||||||
|
|
||||||
|
- 新增 `com.superbiz.agent.harness.agent` 内部包、一个 classpath Prompt 和 focused Agent/tool-loop tests。
|
||||||
|
- 复用 `DiagnosisHarnessCore`、`RunContext`、`DiagnosisDraft`、`PreviousTurn`、`AgentToolContracts`、三类 Tool adapter 和 Spring AI Alibaba ReactAgent/interceptor API。
|
||||||
|
- 内部接口影响为 L2:新增可供阶段 6A 调用的内部用例与工厂;现有公开 Controller、SSE、DTO、数据库契约和旧 ChatService 行为保持不变。
|
||||||
|
- 主要风险是框架结构化输出只提供格式提示而不负责 Java 反序列化,以及 Tool Callback 本身拿不到框架调用 ID;实现必须分别用严格 ObjectMapper 解析和 ToolInterceptor 解决,失败时 fail closed。
|
||||||
+78
@@ -0,0 +1,78 @@
|
|||||||
|
## ADDED Requirements
|
||||||
|
|
||||||
|
### Requirement: Diagnosis execution SHALL use one framework ReactAgent
|
||||||
|
The internal diagnosis path SHALL create exactly one `diagnosis_agent` with the framework-provided ReAct/tool loop. It SHALL NOT create Planner, Executor, Verifier, Composer, SequentialAgent, SupervisorAgent, an outer business Graph, or a hand-written model/Tool loop. One use-case invocation SHALL call the Diagnosis Agent once and SHALL NOT wrap the Agent or any individual model call in Harness retry.
|
||||||
|
|
||||||
|
#### Scenario: Diagnosis needs Tool evidence
|
||||||
|
- **WHEN** the model emits one or more Tool actions before its final response
|
||||||
|
- **THEN** the same Diagnosis ReactAgent executes those actions through its framework loop and returns one final Draft without invoking another Agent
|
||||||
|
|
||||||
|
#### Scenario: Agent execution fails
|
||||||
|
- **WHEN** the single Diagnosis Agent invocation throws or returns invalid structured output
|
||||||
|
- **THEN** the internal use case fails closed without automatically invoking the Agent or model again
|
||||||
|
|
||||||
|
### Requirement: Diagnosis input SHALL contain only current Query and optional PreviousTurn
|
||||||
|
The internal use case SHALL accept a non-blank current Query and an optional frozen `PreviousTurn`, serialize them as the fixed `query` and `previous_turn` input fields, and SHALL NOT load or accept complete Session history, Redis memory, prior raw Tool results, or model-generated history summaries. The current Query SHALL retain its original text and SHALL NOT be rewritten or silently truncated.
|
||||||
|
|
||||||
|
#### Scenario: Follow-up diagnosis has a safe previous turn
|
||||||
|
- **WHEN** a caller supplies a `PreviousTurn`
|
||||||
|
- **THEN** the Agent receives exactly the frozen previous-turn fields plus the current Query and no complete conversation history
|
||||||
|
|
||||||
|
#### Scenario: Query exceeds configured context budget
|
||||||
|
- **WHEN** the current Query exceeds its UTF-8 byte limit
|
||||||
|
- **THEN** the use case rejects it before any model call instead of truncating or rewriting it
|
||||||
|
|
||||||
|
### Requirement: Every model round SHALL be controlled by RunContext
|
||||||
|
The Diagnosis Agent SHALL receive `RunContext` explicitly and SHALL use a model interceptor to call `DiagnosisHarnessCore.beforeModelCall` for every framework ReAct model round. Non-streaming model response Usage SHALL be recorded into the same Run budget when available. Cancellation, deadline, model-call exhaustion, or Token exhaustion SHALL prevent subsequent controlled work.
|
||||||
|
|
||||||
|
#### Scenario: ReAct performs two model rounds
|
||||||
|
- **WHEN** one model round requests a Tool and the next produces the Draft
|
||||||
|
- **THEN** the same Run budget records two model calls and no Harness retry attempt
|
||||||
|
|
||||||
|
#### Scenario: Model-call budget is exhausted
|
||||||
|
- **WHEN** the framework attempts a model round beyond the configured maximum
|
||||||
|
- **THEN** the call is rejected before reaching ChatModel and the Run records budget exhaustion
|
||||||
|
|
||||||
|
### Requirement: Evidence Tools SHALL execute through the Harness boundary
|
||||||
|
The Diagnosis Agent SHALL expose only the frozen `lookup_knowledge`, `query_logs`, and `query_mysql` definitions. A Tool interceptor SHALL propagate the exact framework Tool Call ID, Run ID, Tool name, and raw JSON arguments into the corresponding stage 3B/3C adapter and `ToolBoundary`. Successful observations SHALL contain only the bounded `agent_result`; failed observations SHALL contain only stable error semantics and SHALL NOT contain raw responses, internal exceptions, credentials, or invocation lifecycle internals.
|
||||||
|
|
||||||
|
#### Scenario: Framework requests RAG evidence
|
||||||
|
- **WHEN** the model calls `lookup_knowledge` with framework ID `call-1`
|
||||||
|
- **THEN** the RAG adapter and Agent observation use exactly `call-1`, and the canonical invocation is owned by the current Run
|
||||||
|
|
||||||
|
#### Scenario: Tool execution fails
|
||||||
|
- **WHEN** a registered adapter returns an error result
|
||||||
|
- **THEN** the Agent receives `evidence_status=ERROR`, the framework Tool Call ID and a stable error code without automatic Tool retry or raw failure detail
|
||||||
|
|
||||||
|
#### Scenario: Unknown Tool is requested
|
||||||
|
- **WHEN** a model requests a Tool outside the three registered definitions
|
||||||
|
- **THEN** the Harness does not authorize or emulate it and does not create a canonical evidence record
|
||||||
|
|
||||||
|
### Requirement: Diagnosis output SHALL be a bounded DiagnosisDraft
|
||||||
|
The Agent SHALL receive the generated schema for `DiagnosisDraft` and SHALL return JSON that the internal use case strictly parses into the frozen record. The use case SHALL enforce configured UTF-8 limits for query, previous turn, total input and Draft output and account accepted input/output bytes against the Run capacity. It SHALL reject blank, fenced, prefixed, malformed, oversized or schema-incompatible output without repair or retry.
|
||||||
|
|
||||||
|
#### Scenario: Supported Draft is returned
|
||||||
|
- **WHEN** the Agent returns valid JSON containing Conclusion, Analysis items, Action Plan, Recommendations and Limitations
|
||||||
|
- **THEN** the use case returns a `DiagnosisDraft` whose Analysis Tool Call IDs are the exact strings emitted by the Agent
|
||||||
|
|
||||||
|
#### Scenario: Model returns prose around JSON
|
||||||
|
- **WHEN** the final response contains a Markdown fence or explanatory prefix around an otherwise valid object
|
||||||
|
- **THEN** strict parsing fails and the Diagnosis Agent is not invoked a second time
|
||||||
|
|
||||||
|
### Requirement: Insufficient evidence SHALL terminate without a fabricated conclusion
|
||||||
|
The single Prompt SHALL require every normal Analysis item to cite current-Run evidence Tool Call IDs and SHALL restrict `NO_EVIDENCE` to scoped `NEGATIVE_OBSERVATION`. If current evidence cannot support a diagnosis, the Agent SHALL stop the current ReAct execution with `conclusion=null`, describe the actual scope and missing information in `limitations`, and SHALL NOT infer that the problem does not exist or fabricate a root cause.
|
||||||
|
|
||||||
|
#### Scenario: Tool finds no evidence
|
||||||
|
- **WHEN** the only completed Tool observation has `evidence_status=NO_EVIDENCE`
|
||||||
|
- **THEN** the final Draft has no confirmed Conclusion, records the bounded negative observation and missing information, and makes no additional automatic retry
|
||||||
|
|
||||||
|
### Requirement: Stage-four execution SHALL remain internal and auditable
|
||||||
|
The new use case SHALL be callable only as an internal Java/test entry in this stage and SHALL NOT be wired into public Chat or AIOps controllers. It SHALL propagate the same sessionId/runId through RunnableConfig metadata, allow the existing AgentStep Hook to be injected, retain ToolBoundary canonical invocation audit, and leave final Run success/persistence ownership to later Guard/application stages.
|
||||||
|
|
||||||
|
#### Scenario: Internal audited execution completes
|
||||||
|
- **WHEN** an injected audit Hook observes a model round and a Tool executes
|
||||||
|
- **THEN** both observe the same RunContext sessionId/runId and the public Chat call path remains unchanged
|
||||||
|
|
||||||
|
#### Scenario: Stage four is archived
|
||||||
|
- **WHEN** focused Agent tests pass and the change is archived
|
||||||
|
- **THEN** `ChatController`, public SSE behavior and the old multi-Agent implementation remain available for stages 5, 6A and 6B
|
||||||
@@ -0,0 +1,22 @@
|
|||||||
|
## 1. Input And Prompt Contract
|
||||||
|
|
||||||
|
- [x] 1.1 Add immutable DiagnosisAgentInput and configurable UTF-8 context/output limits with validation and focused boundary tests.
|
||||||
|
- [x] 1.2 Add the single Diagnosis Agent classpath Prompt covering Tool selection, current-Run evidence binding, no-evidence stopping and exact DiagnosisDraft JSON output.
|
||||||
|
|
||||||
|
## 2. Harness Tool Integration
|
||||||
|
|
||||||
|
- [x] 2.1 Add HarnessEvidenceTools with the three frozen Tool definitions and adapter bridges for RAG, logs and MySQL.
|
||||||
|
- [x] 2.2 Add a ToolInterceptor that propagates the exact framework Tool Call ID and RunContext, returns only bounded agent_result or stable error JSON, and never retries or leaks raw results.
|
||||||
|
- [x] 2.3 Add focused tests for framework ID fidelity, successful Tool observations, error observations, unknown Tool rejection and single Tool budget accounting.
|
||||||
|
|
||||||
|
## 3. Single Diagnosis Agent Execution
|
||||||
|
|
||||||
|
- [x] 3.1 Add a ModelInterceptor that enforces every model boundary and records non-streaming Token Usage into the same Run budget.
|
||||||
|
- [x] 3.2 Add DiagnosisAgentFactory that creates one non-parallel ReactAgent with one Prompt, DiagnosisDraft output schema, Harness interceptors and injectable audit Hooks.
|
||||||
|
- [x] 3.3 Add the internal DiagnosisAgentUseCase that validates/serializes fixed input, enforces context/Draft byte limits, calls the Agent once and strictly parses DiagnosisDraft without repair or retry.
|
||||||
|
|
||||||
|
## 4. Focused Behavior Verification
|
||||||
|
|
||||||
|
- [x] 4.1 Add scripted ChatModel tests proving a single Agent performs framework model -> Tool -> model execution, preserves Query/PreviousTurn, returns typed Analysis/Conclusion/Tool IDs and propagates sessionId/runId to audit.
|
||||||
|
- [x] 4.2 Add no-evidence, invalid/fenced output, model-call budget, Token budget, context budget and no-automatic-retry tests.
|
||||||
|
- [x] 4.3 Verify no public Chat/AiOps runtime file changed, run focused Harness/Agent regressions, compile, and validate the OpenSpec strictly.
|
||||||
@@ -0,0 +1,81 @@
|
|||||||
|
# single-react-diagnosis-agent Specification
|
||||||
|
|
||||||
|
## Purpose
|
||||||
|
定义内部单体 Diagnosis ReactAgent 的极简输入、框架 ReAct/tool loop、Harness 模型与 Tool 边界、结构化 DiagnosisDraft、无证据停止、确定性预算和审计隔离要求。
|
||||||
|
## Requirements
|
||||||
|
### Requirement: Diagnosis execution SHALL use one framework ReactAgent
|
||||||
|
The internal diagnosis path SHALL create exactly one `diagnosis_agent` with the framework-provided ReAct/tool loop. It SHALL NOT create Planner, Executor, Verifier, Composer, SequentialAgent, SupervisorAgent, an outer business Graph, or a hand-written model/Tool loop. One use-case invocation SHALL call the Diagnosis Agent once and SHALL NOT wrap the Agent or any individual model call in Harness retry.
|
||||||
|
|
||||||
|
#### Scenario: Diagnosis needs Tool evidence
|
||||||
|
- **WHEN** the model emits one or more Tool actions before its final response
|
||||||
|
- **THEN** the same Diagnosis ReactAgent executes those actions through its framework loop and returns one final Draft without invoking another Agent
|
||||||
|
|
||||||
|
#### Scenario: Agent execution fails
|
||||||
|
- **WHEN** the single Diagnosis Agent invocation throws or returns invalid structured output
|
||||||
|
- **THEN** the internal use case fails closed without automatically invoking the Agent or model again
|
||||||
|
|
||||||
|
### Requirement: Diagnosis input SHALL contain only current Query and optional PreviousTurn
|
||||||
|
The internal use case SHALL accept a non-blank current Query and an optional frozen `PreviousTurn`, serialize them as the fixed `query` and `previous_turn` input fields, and SHALL NOT load or accept complete Session history, Redis memory, prior raw Tool results, or model-generated history summaries. The current Query SHALL retain its original text and SHALL NOT be rewritten or silently truncated.
|
||||||
|
|
||||||
|
#### Scenario: Follow-up diagnosis has a safe previous turn
|
||||||
|
- **WHEN** a caller supplies a `PreviousTurn`
|
||||||
|
- **THEN** the Agent receives exactly the frozen previous-turn fields plus the current Query and no complete conversation history
|
||||||
|
|
||||||
|
#### Scenario: Query exceeds configured context budget
|
||||||
|
- **WHEN** the current Query exceeds its UTF-8 byte limit
|
||||||
|
- **THEN** the use case rejects it before any model call instead of truncating or rewriting it
|
||||||
|
|
||||||
|
### Requirement: Every model round SHALL be controlled by RunContext
|
||||||
|
The Diagnosis Agent SHALL receive `RunContext` explicitly and SHALL use a model interceptor to call `DiagnosisHarnessCore.beforeModelCall` for every framework ReAct model round. Non-streaming model response Usage SHALL be recorded into the same Run budget when available. Cancellation, deadline, model-call exhaustion, or Token exhaustion SHALL prevent subsequent controlled work.
|
||||||
|
|
||||||
|
#### Scenario: ReAct performs two model rounds
|
||||||
|
- **WHEN** one model round requests a Tool and the next produces the Draft
|
||||||
|
- **THEN** the same Run budget records two model calls and no Harness retry attempt
|
||||||
|
|
||||||
|
#### Scenario: Model-call budget is exhausted
|
||||||
|
- **WHEN** the framework attempts a model round beyond the configured maximum
|
||||||
|
- **THEN** the call is rejected before reaching ChatModel and the Run records budget exhaustion
|
||||||
|
|
||||||
|
### Requirement: Evidence Tools SHALL execute through the Harness boundary
|
||||||
|
The Diagnosis Agent SHALL expose only the frozen `lookup_knowledge`, `query_logs`, and `query_mysql` definitions. A Tool interceptor SHALL propagate the exact framework Tool Call ID, Run ID, Tool name, and raw JSON arguments into the corresponding stage 3B/3C adapter and `ToolBoundary`. Successful observations SHALL contain only the bounded `agent_result`; failed observations SHALL contain only stable error semantics and SHALL NOT contain raw responses, internal exceptions, credentials, or invocation lifecycle internals.
|
||||||
|
|
||||||
|
#### Scenario: Framework requests RAG evidence
|
||||||
|
- **WHEN** the model calls `lookup_knowledge` with framework ID `call-1`
|
||||||
|
- **THEN** the RAG adapter and Agent observation use exactly `call-1`, and the canonical invocation is owned by the current Run
|
||||||
|
|
||||||
|
#### Scenario: Tool execution fails
|
||||||
|
- **WHEN** a registered adapter returns an error result
|
||||||
|
- **THEN** the Agent receives `evidence_status=ERROR`, the framework Tool Call ID and a stable error code without automatic Tool retry or raw failure detail
|
||||||
|
|
||||||
|
#### Scenario: Unknown Tool is requested
|
||||||
|
- **WHEN** a model requests a Tool outside the three registered definitions
|
||||||
|
- **THEN** the Harness does not authorize or emulate it and does not create a canonical evidence record
|
||||||
|
|
||||||
|
### Requirement: Diagnosis output SHALL be a bounded DiagnosisDraft
|
||||||
|
The Agent SHALL receive the generated schema for `DiagnosisDraft` and SHALL return JSON that the internal use case strictly parses into the frozen record. The use case SHALL enforce configured UTF-8 limits for query, previous turn, total input and Draft output and account accepted input/output bytes against the Run capacity. It SHALL reject blank, fenced, prefixed, malformed, oversized or schema-incompatible output without repair or retry.
|
||||||
|
|
||||||
|
#### Scenario: Supported Draft is returned
|
||||||
|
- **WHEN** the Agent returns valid JSON containing Conclusion, Analysis items, Action Plan, Recommendations and Limitations
|
||||||
|
- **THEN** the use case returns a `DiagnosisDraft` whose Analysis Tool Call IDs are the exact strings emitted by the Agent
|
||||||
|
|
||||||
|
#### Scenario: Model returns prose around JSON
|
||||||
|
- **WHEN** the final response contains a Markdown fence or explanatory prefix around an otherwise valid object
|
||||||
|
- **THEN** strict parsing fails and the Diagnosis Agent is not invoked a second time
|
||||||
|
|
||||||
|
### Requirement: Insufficient evidence SHALL terminate without a fabricated conclusion
|
||||||
|
The single Prompt SHALL require every normal Analysis item to cite current-Run evidence Tool Call IDs and SHALL restrict `NO_EVIDENCE` to scoped `NEGATIVE_OBSERVATION`. If current evidence cannot support a diagnosis, the Agent SHALL stop the current ReAct execution with `conclusion=null`, describe the actual scope and missing information in `limitations`, and SHALL NOT infer that the problem does not exist or fabricate a root cause.
|
||||||
|
|
||||||
|
#### Scenario: Tool finds no evidence
|
||||||
|
- **WHEN** the only completed Tool observation has `evidence_status=NO_EVIDENCE`
|
||||||
|
- **THEN** the final Draft has no confirmed Conclusion, records the bounded negative observation and missing information, and makes no additional automatic retry
|
||||||
|
|
||||||
|
### Requirement: Stage-four execution SHALL remain internal and auditable
|
||||||
|
The new use case SHALL be callable only as an internal Java/test entry in this stage and SHALL NOT be wired into public Chat or AIOps controllers. It SHALL propagate the same sessionId/runId through RunnableConfig metadata, allow the existing AgentStep Hook to be injected, retain ToolBoundary canonical invocation audit, and leave final Run success/persistence ownership to later Guard/application stages.
|
||||||
|
|
||||||
|
#### Scenario: Internal audited execution completes
|
||||||
|
- **WHEN** an injected audit Hook observes a model round and a Tool executes
|
||||||
|
- **THEN** both observe the same RunContext sessionId/runId and the public Chat call path remains unchanged
|
||||||
|
|
||||||
|
#### Scenario: Stage four is archived
|
||||||
|
- **WHEN** focused Agent tests pass and the change is archived
|
||||||
|
- **THEN** `ChatController`, public SSE behavior and the old multi-Agent implementation remain available for stages 5, 6A and 6B
|
||||||
@@ -0,0 +1,54 @@
|
|||||||
|
package com.superbiz.agent.harness.agent;
|
||||||
|
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.ReactAgent;
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.hook.Hook;
|
||||||
|
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||||
|
import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
|
||||||
|
import com.superbiz.agent.harness.core.RunContext;
|
||||||
|
import org.springframework.ai.chat.model.ChatModel;
|
||||||
|
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.Objects;
|
||||||
|
|
||||||
|
public final class DiagnosisAgentFactory {
|
||||||
|
|
||||||
|
public static final String AGENT_NAME = "diagnosis_agent";
|
||||||
|
|
||||||
|
private final ChatModel chatModel;
|
||||||
|
private final DiagnosisHarnessCore core;
|
||||||
|
private final HarnessEvidenceTools evidenceTools;
|
||||||
|
private final ObjectMapper objectMapper;
|
||||||
|
private final List<Hook> auditHooks;
|
||||||
|
private final String prompt;
|
||||||
|
|
||||||
|
public DiagnosisAgentFactory(ChatModel chatModel,
|
||||||
|
DiagnosisHarnessCore core,
|
||||||
|
HarnessEvidenceTools evidenceTools,
|
||||||
|
ObjectMapper objectMapper,
|
||||||
|
List<? extends Hook> auditHooks) {
|
||||||
|
this.chatModel = Objects.requireNonNull(chatModel, "chatModel must not be null");
|
||||||
|
this.core = Objects.requireNonNull(core, "core must not be null");
|
||||||
|
this.evidenceTools = Objects.requireNonNull(evidenceTools, "evidenceTools must not be null");
|
||||||
|
this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null");
|
||||||
|
this.auditHooks = List.copyOf(Objects.requireNonNull(auditHooks, "auditHooks must not be null"));
|
||||||
|
this.prompt = DiagnosisAgentPrompt.load();
|
||||||
|
}
|
||||||
|
|
||||||
|
public ReactAgent create(RunContext context) {
|
||||||
|
Objects.requireNonNull(context, "context must not be null");
|
||||||
|
return ReactAgent.builder()
|
||||||
|
.name(AGENT_NAME)
|
||||||
|
.description("Collects bounded evidence and authors one diagnosis draft")
|
||||||
|
.model(chatModel)
|
||||||
|
.systemPrompt(prompt)
|
||||||
|
.tools(evidenceTools.callbacks())
|
||||||
|
.interceptors(
|
||||||
|
new HarnessModelInterceptor(core, context),
|
||||||
|
new HarnessToolInterceptor(context, evidenceTools, objectMapper))
|
||||||
|
.hooks(auditHooks)
|
||||||
|
.outputSchema(new DiagnosisDraftOutputSchema(objectMapper).getFormat())
|
||||||
|
.parallelToolExecution(false)
|
||||||
|
.releaseThread(true)
|
||||||
|
.build();
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,15 @@
|
|||||||
|
package com.superbiz.agent.harness.agent;
|
||||||
|
|
||||||
|
import com.fasterxml.jackson.annotation.JsonProperty;
|
||||||
|
import com.superbiz.agent.harness.contract.PreviousTurn;
|
||||||
|
|
||||||
|
public record DiagnosisAgentInput(
|
||||||
|
@JsonProperty("query") String query,
|
||||||
|
@JsonProperty("previous_turn") PreviousTurn previousTurn) {
|
||||||
|
|
||||||
|
public DiagnosisAgentInput {
|
||||||
|
if (query == null || query.isBlank()) {
|
||||||
|
throw new IllegalArgumentException("query must not be blank");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,27 @@
|
|||||||
|
package com.superbiz.agent.harness.agent;
|
||||||
|
|
||||||
|
public final class DiagnosisAgentLimitException extends IllegalArgumentException {
|
||||||
|
|
||||||
|
private final String boundary;
|
||||||
|
private final long limit;
|
||||||
|
private final long actual;
|
||||||
|
|
||||||
|
public DiagnosisAgentLimitException(String boundary, long limit, long actual) {
|
||||||
|
super(boundary + " exceeds UTF-8 byte limit: limit=" + limit + ", actual=" + actual);
|
||||||
|
this.boundary = boundary;
|
||||||
|
this.limit = limit;
|
||||||
|
this.actual = actual;
|
||||||
|
}
|
||||||
|
|
||||||
|
public String boundary() {
|
||||||
|
return boundary;
|
||||||
|
}
|
||||||
|
|
||||||
|
public long limit() {
|
||||||
|
return limit;
|
||||||
|
}
|
||||||
|
|
||||||
|
public long actual() {
|
||||||
|
return actual;
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
package com.superbiz.agent.harness.agent;
|
||||||
|
|
||||||
|
public record DiagnosisAgentLimits(
|
||||||
|
long maxQueryBytes,
|
||||||
|
long maxPreviousTurnBytes,
|
||||||
|
long maxInputBytes,
|
||||||
|
long maxDraftBytes) {
|
||||||
|
|
||||||
|
public DiagnosisAgentLimits {
|
||||||
|
requirePositive(maxQueryBytes, "maxQueryBytes");
|
||||||
|
requirePositive(maxPreviousTurnBytes, "maxPreviousTurnBytes");
|
||||||
|
requirePositive(maxInputBytes, "maxInputBytes");
|
||||||
|
requirePositive(maxDraftBytes, "maxDraftBytes");
|
||||||
|
}
|
||||||
|
|
||||||
|
private static void requirePositive(long value, String name) {
|
||||||
|
if (value <= 0) {
|
||||||
|
throw new IllegalArgumentException(name + " must be positive");
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,12 @@
|
|||||||
|
package com.superbiz.agent.harness.agent;
|
||||||
|
|
||||||
|
public final class DiagnosisAgentOutputException extends RuntimeException {
|
||||||
|
|
||||||
|
public DiagnosisAgentOutputException(String message) {
|
||||||
|
super(message);
|
||||||
|
}
|
||||||
|
|
||||||
|
public DiagnosisAgentOutputException(String message, Throwable cause) {
|
||||||
|
super(message, cause);
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,24 @@
|
|||||||
|
package com.superbiz.agent.harness.agent;
|
||||||
|
|
||||||
|
import org.springframework.core.io.ClassPathResource;
|
||||||
|
|
||||||
|
import java.io.IOException;
|
||||||
|
import java.io.InputStream;
|
||||||
|
import java.nio.charset.StandardCharsets;
|
||||||
|
|
||||||
|
public final class DiagnosisAgentPrompt {
|
||||||
|
|
||||||
|
public static final String RESOURCE_PATH = "prompts/diagnosis-agent-prompt.md";
|
||||||
|
|
||||||
|
private DiagnosisAgentPrompt() {
|
||||||
|
}
|
||||||
|
|
||||||
|
public static String load() {
|
||||||
|
ClassPathResource resource = new ClassPathResource(RESOURCE_PATH);
|
||||||
|
try (InputStream input = resource.getInputStream()) {
|
||||||
|
return new String(input.readAllBytes(), StandardCharsets.UTF_8);
|
||||||
|
} catch (IOException e) {
|
||||||
|
throw new IllegalStateException("Failed to load Diagnosis Agent prompt", e);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,103 @@
|
|||||||
|
package com.superbiz.agent.harness.agent;
|
||||||
|
|
||||||
|
import com.alibaba.cloud.ai.graph.RunnableConfig;
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.ReactAgent;
|
||||||
|
import com.fasterxml.jackson.core.JsonProcessingException;
|
||||||
|
import com.fasterxml.jackson.databind.DeserializationFeature;
|
||||||
|
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||||
|
import com.fasterxml.jackson.databind.ObjectReader;
|
||||||
|
import com.superbiz.agent.harness.contract.DiagnosisDraft;
|
||||||
|
import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
|
||||||
|
import com.superbiz.agent.harness.core.RunContext;
|
||||||
|
import org.springframework.ai.chat.messages.AssistantMessage;
|
||||||
|
|
||||||
|
import java.nio.charset.StandardCharsets;
|
||||||
|
import java.util.Objects;
|
||||||
|
|
||||||
|
public final class DiagnosisAgentUseCase {
|
||||||
|
|
||||||
|
public static final String RUN_CONTEXT_METADATA = "runContext";
|
||||||
|
|
||||||
|
private final DiagnosisHarnessCore core;
|
||||||
|
private final DiagnosisAgentFactory agentFactory;
|
||||||
|
private final ObjectMapper objectMapper;
|
||||||
|
private final ObjectReader draftReader;
|
||||||
|
private final DiagnosisAgentLimits limits;
|
||||||
|
|
||||||
|
public DiagnosisAgentUseCase(DiagnosisHarnessCore core,
|
||||||
|
DiagnosisAgentFactory agentFactory,
|
||||||
|
ObjectMapper objectMapper,
|
||||||
|
DiagnosisAgentLimits limits) {
|
||||||
|
this.core = Objects.requireNonNull(core, "core must not be null");
|
||||||
|
this.agentFactory = Objects.requireNonNull(agentFactory, "agentFactory must not be null");
|
||||||
|
this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null");
|
||||||
|
this.draftReader = objectMapper.readerFor(DiagnosisDraft.class)
|
||||||
|
.with(DeserializationFeature.FAIL_ON_UNKNOWN_PROPERTIES)
|
||||||
|
.with(DeserializationFeature.FAIL_ON_TRAILING_TOKENS);
|
||||||
|
this.limits = Objects.requireNonNull(limits, "limits must not be null");
|
||||||
|
}
|
||||||
|
|
||||||
|
public DiagnosisDraft execute(RunContext context, DiagnosisAgentInput input) {
|
||||||
|
Objects.requireNonNull(context, "context must not be null");
|
||||||
|
Objects.requireNonNull(input, "input must not be null");
|
||||||
|
core.checkActive(context);
|
||||||
|
|
||||||
|
checkLimit("query", utf8Bytes(input.query()), limits.maxQueryBytes());
|
||||||
|
if (input.previousTurn() != null) {
|
||||||
|
checkLimit("previous_turn", utf8Bytes(writeJson(input.previousTurn())),
|
||||||
|
limits.maxPreviousTurnBytes());
|
||||||
|
}
|
||||||
|
|
||||||
|
String inputJson = writeJson(input);
|
||||||
|
long inputBytes = utf8Bytes(inputJson);
|
||||||
|
checkLimit("input", inputBytes, limits.maxInputBytes());
|
||||||
|
core.reserveRunBytes(context, inputBytes);
|
||||||
|
|
||||||
|
RunnableConfig config = RunnableConfig.builder()
|
||||||
|
.threadId(context.runId())
|
||||||
|
.addMetadata("sessionId", context.sessionId())
|
||||||
|
.addMetadata("runId", context.runId())
|
||||||
|
.addMetadata(RUN_CONTEXT_METADATA, context)
|
||||||
|
.addMetadata("_stream_", false)
|
||||||
|
.build();
|
||||||
|
ReactAgent agent = agentFactory.create(context);
|
||||||
|
AssistantMessage response;
|
||||||
|
try {
|
||||||
|
response = agent.call(inputJson, config);
|
||||||
|
} catch (Exception e) {
|
||||||
|
throw new DiagnosisAgentOutputException("Diagnosis Agent execution failed", e);
|
||||||
|
}
|
||||||
|
|
||||||
|
core.checkActive(context);
|
||||||
|
String output = response == null ? null : response.getText();
|
||||||
|
if (output == null || output.isBlank()) {
|
||||||
|
throw new DiagnosisAgentOutputException("Diagnosis Agent returned an empty Draft");
|
||||||
|
}
|
||||||
|
long draftBytes = utf8Bytes(output);
|
||||||
|
checkLimit("draft", draftBytes, limits.maxDraftBytes());
|
||||||
|
core.reserveRunBytes(context, draftBytes);
|
||||||
|
try {
|
||||||
|
return draftReader.readValue(output);
|
||||||
|
} catch (JsonProcessingException e) {
|
||||||
|
throw new DiagnosisAgentOutputException("Diagnosis Agent returned an invalid Draft", e);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private String writeJson(Object value) {
|
||||||
|
try {
|
||||||
|
return objectMapper.writeValueAsString(value);
|
||||||
|
} catch (JsonProcessingException e) {
|
||||||
|
throw new IllegalArgumentException("Diagnosis Agent input is not serializable", e);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private static long utf8Bytes(String value) {
|
||||||
|
return value.getBytes(StandardCharsets.UTF_8).length;
|
||||||
|
}
|
||||||
|
|
||||||
|
private static void checkLimit(String boundary, long actual, long limit) {
|
||||||
|
if (actual > limit) {
|
||||||
|
throw new DiagnosisAgentLimitException(boundary, limit, actual);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,29 @@
|
|||||||
|
package com.superbiz.agent.harness.agent;
|
||||||
|
|
||||||
|
import com.fasterxml.jackson.databind.JsonNode;
|
||||||
|
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||||
|
import com.fasterxml.jackson.databind.node.ArrayNode;
|
||||||
|
import com.fasterxml.jackson.databind.node.ObjectNode;
|
||||||
|
import com.superbiz.agent.harness.contract.DiagnosisDraft;
|
||||||
|
import org.springframework.ai.converter.BeanOutputConverter;
|
||||||
|
|
||||||
|
import java.util.Objects;
|
||||||
|
|
||||||
|
final class DiagnosisDraftOutputSchema extends BeanOutputConverter<DiagnosisDraft> {
|
||||||
|
|
||||||
|
DiagnosisDraftOutputSchema(ObjectMapper objectMapper) {
|
||||||
|
super(DiagnosisDraft.class, Objects.requireNonNull(objectMapper, "objectMapper must not be null"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
protected void postProcessSchema(JsonNode schema) {
|
||||||
|
JsonNode conclusion = schema.path("properties").path("conclusion");
|
||||||
|
if (!(conclusion instanceof ObjectNode conclusionSchema)) {
|
||||||
|
throw new IllegalStateException("DiagnosisDraft schema has no conclusion property");
|
||||||
|
}
|
||||||
|
ArrayNode allowedTypes = conclusionSchema.arrayNode();
|
||||||
|
allowedTypes.add("object");
|
||||||
|
allowedTypes.add("null");
|
||||||
|
conclusionSchema.set("type", allowedTypes);
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,10 @@
|
|||||||
|
package com.superbiz.agent.harness.agent;
|
||||||
|
|
||||||
|
import com.superbiz.agent.harness.core.RunContext;
|
||||||
|
import com.superbiz.agent.harness.tool.boundary.ToolBoundaryResult;
|
||||||
|
|
||||||
|
@FunctionalInterface
|
||||||
|
public interface EvidenceToolInvoker {
|
||||||
|
|
||||||
|
ToolBoundaryResult invoke(RunContext context, String toolCallId, String arguments);
|
||||||
|
}
|
||||||
@@ -0,0 +1,95 @@
|
|||||||
|
package com.superbiz.agent.harness.agent;
|
||||||
|
|
||||||
|
import com.superbiz.agent.harness.core.RunContext;
|
||||||
|
import com.superbiz.agent.harness.tool.adapter.MysqlToolAdapter;
|
||||||
|
import com.superbiz.agent.harness.tool.adapter.QueryLogsToolAdapter;
|
||||||
|
import com.superbiz.agent.harness.tool.adapter.RagToolAdapter;
|
||||||
|
import com.superbiz.agent.harness.tool.boundary.ToolBoundaryResult;
|
||||||
|
import com.superbiz.agent.harness.tool.boundary.ToolCallRequestEnvelope;
|
||||||
|
import com.superbiz.agent.harness.tool.contract.AgentToolContracts;
|
||||||
|
import com.superbiz.agent.harness.tool.contract.MysqlToolRequest;
|
||||||
|
import com.superbiz.agent.harness.tool.contract.QueryLogsRequest;
|
||||||
|
import com.superbiz.agent.harness.tool.contract.RagToolRequest;
|
||||||
|
import org.springframework.ai.tool.ToolCallback;
|
||||||
|
import org.springframework.ai.tool.function.FunctionToolCallback;
|
||||||
|
|
||||||
|
import java.util.LinkedHashMap;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.Objects;
|
||||||
|
|
||||||
|
public final class HarnessEvidenceTools {
|
||||||
|
|
||||||
|
private final List<ToolCallback> callbacks;
|
||||||
|
private final Map<String, EvidenceToolInvoker> invokers;
|
||||||
|
|
||||||
|
public HarnessEvidenceTools(EvidenceToolInvoker ragInvoker,
|
||||||
|
EvidenceToolInvoker logsInvoker,
|
||||||
|
EvidenceToolInvoker mysqlInvoker) {
|
||||||
|
Map<String, EvidenceToolInvoker> registered = new LinkedHashMap<>();
|
||||||
|
registered.put(AgentToolContracts.LOOKUP_KNOWLEDGE,
|
||||||
|
Objects.requireNonNull(ragInvoker, "ragInvoker must not be null"));
|
||||||
|
registered.put(AgentToolContracts.QUERY_LOGS,
|
||||||
|
Objects.requireNonNull(logsInvoker, "logsInvoker must not be null"));
|
||||||
|
registered.put(AgentToolContracts.QUERY_MYSQL,
|
||||||
|
Objects.requireNonNull(mysqlInvoker, "mysqlInvoker must not be null"));
|
||||||
|
this.invokers = Map.copyOf(registered);
|
||||||
|
this.callbacks = List.of(
|
||||||
|
definition(AgentToolContracts.LOOKUP_KNOWLEDGE,
|
||||||
|
AgentToolContracts.LOOKUP_KNOWLEDGE_DESCRIPTION, RagToolRequest.class),
|
||||||
|
definition(AgentToolContracts.QUERY_LOGS,
|
||||||
|
AgentToolContracts.QUERY_LOGS_DESCRIPTION, QueryLogsRequest.class),
|
||||||
|
definition(AgentToolContracts.QUERY_MYSQL,
|
||||||
|
AgentToolContracts.QUERY_MYSQL_DESCRIPTION, MysqlToolRequest.class));
|
||||||
|
}
|
||||||
|
|
||||||
|
public static HarnessEvidenceTools fromAdapters(RagToolAdapter ragAdapter,
|
||||||
|
QueryLogsToolAdapter logsAdapter,
|
||||||
|
MysqlToolAdapter mysqlAdapter) {
|
||||||
|
Objects.requireNonNull(ragAdapter, "ragAdapter must not be null");
|
||||||
|
Objects.requireNonNull(logsAdapter, "logsAdapter must not be null");
|
||||||
|
Objects.requireNonNull(mysqlAdapter, "mysqlAdapter must not be null");
|
||||||
|
return new HarnessEvidenceTools(
|
||||||
|
bridge(AgentToolContracts.LOOKUP_KNOWLEDGE, ragAdapter::execute),
|
||||||
|
bridge(AgentToolContracts.QUERY_LOGS, logsAdapter::execute),
|
||||||
|
bridge(AgentToolContracts.QUERY_MYSQL, mysqlAdapter::execute));
|
||||||
|
}
|
||||||
|
|
||||||
|
public List<ToolCallback> callbacks() {
|
||||||
|
return callbacks;
|
||||||
|
}
|
||||||
|
|
||||||
|
public boolean supports(String toolName) {
|
||||||
|
return invokers.containsKey(toolName);
|
||||||
|
}
|
||||||
|
|
||||||
|
public ToolBoundaryResult invoke(RunContext context, String toolName,
|
||||||
|
String toolCallId, String arguments) {
|
||||||
|
EvidenceToolInvoker invoker = invokers.get(toolName);
|
||||||
|
if (invoker == null) {
|
||||||
|
throw new IllegalArgumentException("Unsupported evidence Tool: " + toolName);
|
||||||
|
}
|
||||||
|
return invoker.invoke(context, toolCallId, arguments);
|
||||||
|
}
|
||||||
|
|
||||||
|
private static EvidenceToolInvoker bridge(String toolName, AdapterCall adapter) {
|
||||||
|
return (context, toolCallId, arguments) -> adapter.execute(
|
||||||
|
context,
|
||||||
|
new ToolCallRequestEnvelope(
|
||||||
|
context.runId(), toolCallId, toolName, arguments, true, true));
|
||||||
|
}
|
||||||
|
|
||||||
|
private static <I> ToolCallback definition(String name, String description, Class<I> inputType) {
|
||||||
|
return FunctionToolCallback.<I, String>builder(name, ignored -> {
|
||||||
|
throw new IllegalStateException("Harness evidence Tools require the framework Tool interceptor");
|
||||||
|
})
|
||||||
|
.description(description)
|
||||||
|
.inputType(inputType)
|
||||||
|
.build();
|
||||||
|
}
|
||||||
|
|
||||||
|
@FunctionalInterface
|
||||||
|
private interface AdapterCall {
|
||||||
|
ToolBoundaryResult execute(RunContext context, ToolCallRequestEnvelope envelope);
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,56 @@
|
|||||||
|
package com.superbiz.agent.harness.agent;
|
||||||
|
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.interceptor.ModelCallHandler;
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.interceptor.ModelInterceptor;
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.interceptor.ModelRequest;
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.interceptor.ModelResponse;
|
||||||
|
import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
|
||||||
|
import com.superbiz.agent.harness.core.RunContext;
|
||||||
|
import org.springframework.ai.chat.metadata.Usage;
|
||||||
|
import org.springframework.ai.chat.model.ChatResponse;
|
||||||
|
|
||||||
|
import java.util.Objects;
|
||||||
|
|
||||||
|
public final class HarnessModelInterceptor extends ModelInterceptor {
|
||||||
|
|
||||||
|
private final DiagnosisHarnessCore core;
|
||||||
|
private final RunContext context;
|
||||||
|
|
||||||
|
public HarnessModelInterceptor(DiagnosisHarnessCore core, RunContext context) {
|
||||||
|
this.core = Objects.requireNonNull(core, "core must not be null");
|
||||||
|
this.context = Objects.requireNonNull(context, "context must not be null");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public String getName() {
|
||||||
|
return "harness_model_budget_interceptor";
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public ModelResponse interceptModel(ModelRequest request, ModelCallHandler handler) {
|
||||||
|
Objects.requireNonNull(request, "request must not be null");
|
||||||
|
Objects.requireNonNull(handler, "handler must not be null");
|
||||||
|
core.beforeModelCall(context);
|
||||||
|
ModelResponse response = handler.call(request);
|
||||||
|
recordUsage(response == null ? null : response.getChatResponse());
|
||||||
|
core.checkActive(context);
|
||||||
|
return response;
|
||||||
|
}
|
||||||
|
|
||||||
|
private void recordUsage(ChatResponse response) {
|
||||||
|
if (response == null || response.getMetadata() == null) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
Usage usage = response.getMetadata().getUsage();
|
||||||
|
if (usage == null) {
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
long inputTokens = nonNegative(usage.getPromptTokens());
|
||||||
|
long outputTokens = nonNegative(usage.getCompletionTokens());
|
||||||
|
core.recordTokens(context, inputTokens, outputTokens);
|
||||||
|
}
|
||||||
|
|
||||||
|
private static long nonNegative(Integer value) {
|
||||||
|
return value == null || value < 0 ? 0L : value.longValue();
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,69 @@
|
|||||||
|
package com.superbiz.agent.harness.agent;
|
||||||
|
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.interceptor.ToolCallHandler;
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.interceptor.ToolCallRequest;
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.interceptor.ToolCallResponse;
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.interceptor.ToolInterceptor;
|
||||||
|
import com.fasterxml.jackson.core.JsonProcessingException;
|
||||||
|
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||||
|
import com.superbiz.agent.harness.contract.InvocationStatus;
|
||||||
|
import com.superbiz.agent.harness.core.RunContext;
|
||||||
|
import com.superbiz.agent.harness.tool.boundary.ToolBoundaryResult;
|
||||||
|
|
||||||
|
import java.util.LinkedHashMap;
|
||||||
|
import java.util.Map;
|
||||||
|
import java.util.Objects;
|
||||||
|
|
||||||
|
public final class HarnessToolInterceptor extends ToolInterceptor {
|
||||||
|
|
||||||
|
private final RunContext context;
|
||||||
|
private final HarnessEvidenceTools evidenceTools;
|
||||||
|
private final ObjectMapper objectMapper;
|
||||||
|
|
||||||
|
public HarnessToolInterceptor(RunContext context,
|
||||||
|
HarnessEvidenceTools evidenceTools,
|
||||||
|
ObjectMapper objectMapper) {
|
||||||
|
this.context = Objects.requireNonNull(context, "context must not be null");
|
||||||
|
this.evidenceTools = Objects.requireNonNull(evidenceTools, "evidenceTools must not be null");
|
||||||
|
this.objectMapper = Objects.requireNonNull(objectMapper, "objectMapper must not be null");
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public String getName() {
|
||||||
|
return "harness_evidence_tool_interceptor";
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public ToolCallResponse interceptToolCall(ToolCallRequest request, ToolCallHandler handler) {
|
||||||
|
Objects.requireNonNull(request, "request must not be null");
|
||||||
|
Objects.requireNonNull(handler, "handler must not be null");
|
||||||
|
if (!evidenceTools.supports(request.getToolName())) {
|
||||||
|
return handler.call(request);
|
||||||
|
}
|
||||||
|
|
||||||
|
ToolBoundaryResult result = evidenceTools.invoke(
|
||||||
|
context, request.getToolName(), request.getToolCallId(), request.getArguments());
|
||||||
|
if (result.status() == InvocationStatus.READY) {
|
||||||
|
return ToolCallResponse.of(request.getToolCallId(), request.getToolName(), result.agentResult());
|
||||||
|
}
|
||||||
|
return ToolCallResponse.builder()
|
||||||
|
.toolCallId(request.getToolCallId())
|
||||||
|
.toolName(request.getToolName())
|
||||||
|
.content(errorObservation(result))
|
||||||
|
.status("error")
|
||||||
|
.metadata(Map.of("error", true))
|
||||||
|
.build();
|
||||||
|
}
|
||||||
|
|
||||||
|
private String errorObservation(ToolBoundaryResult result) {
|
||||||
|
Map<String, Object> observation = new LinkedHashMap<>();
|
||||||
|
observation.put("evidence_status", result.evidenceStatus());
|
||||||
|
observation.put("tool_call_id", result.toolCallId());
|
||||||
|
observation.put("error_code", result.errorCode());
|
||||||
|
try {
|
||||||
|
return objectMapper.writeValueAsString(observation);
|
||||||
|
} catch (JsonProcessingException e) {
|
||||||
|
return "{\"evidence_status\":\"ERROR\",\"error_code\":\"SERIALIZATION_ERROR\"}";
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,46 @@
|
|||||||
|
# Role
|
||||||
|
|
||||||
|
You are the only Diagnosis Agent for the current diagnosis run. Plan the investigation internally, call the available read-only evidence tools through the framework ReAct loop, evaluate the returned evidence, and author one complete DiagnosisDraft.
|
||||||
|
|
||||||
|
# Input
|
||||||
|
|
||||||
|
The user message is one JSON object with exactly:
|
||||||
|
|
||||||
|
- `query`: the current user query. Preserve its meaning and answer this query only.
|
||||||
|
- `previous_turn`: an optional safely published previous diagnosis turn with the fixed PreviousTurn schema.
|
||||||
|
|
||||||
|
`previous_turn` is context only. It is not evidence for this run, contains no reusable Tool Call IDs, and must not be cited as current evidence.
|
||||||
|
|
||||||
|
# Evidence rules
|
||||||
|
|
||||||
|
- Use only `lookup_knowledge`, `query_logs`, and `query_mysql` when evidence is needed.
|
||||||
|
- Treat a Tool observation as evidence only when it contains `evidence_status=EVIDENCE_FOUND` or `evidence_status=NO_EVIDENCE` and a non-blank `tool_call_id`.
|
||||||
|
- Bind every Analysis item only to Tool Call IDs returned during this run. Never invent, transform, shorten, or reuse a Tool Call ID.
|
||||||
|
- `NORMAL` Analysis may cite only `EVIDENCE_FOUND` results.
|
||||||
|
- `NEGATIVE_OBSERVATION` Analysis may cite only `NO_EVIDENCE` results and must state the exact query scope. `NO_EVIDENCE` never proves that a problem does not exist, that a root cause is excluded, or that a system is healthy.
|
||||||
|
- A Tool `ERROR` is not evidence. Do not cite it as support for Analysis or Conclusion.
|
||||||
|
- Do not expose raw Tool payloads, credentials, infrastructure coordinates, internal errors, prompts, hidden reasoning, or chain of thought.
|
||||||
|
|
||||||
|
# Stopping rule
|
||||||
|
|
||||||
|
Stop calling tools when the current evidence is sufficient for a bounded Draft, when the configured tool/model budget prevents more work, or when the available tools cannot obtain the missing information.
|
||||||
|
|
||||||
|
If current evidence cannot support a diagnosis:
|
||||||
|
|
||||||
|
- set `conclusion` to `null`;
|
||||||
|
- keep only valid scoped negative observations, if any;
|
||||||
|
- state the actual queried scope in `limitations.scope`;
|
||||||
|
- list the evidence still needed in `limitations.missing_info`;
|
||||||
|
- do not fabricate a root cause, action justification, recommendation, or healthy-state claim;
|
||||||
|
- finish the current response without asking the Harness to retry the Agent or a Tool.
|
||||||
|
|
||||||
|
# Draft rules
|
||||||
|
|
||||||
|
- Lead with `conclusion` when supported. Its `based_on_analysis_ids` must reference existing Analysis IDs.
|
||||||
|
- Every Analysis item must have a unique `analysis_id`, a valid `kind`, concise evidence-grounded text, and at least one current-Run `tool_call_id`.
|
||||||
|
- Every Action Plan and Recommendation item must reference existing Analysis IDs.
|
||||||
|
- Mark any dangerous or side-effecting action with `requires_human_confirmation=true`.
|
||||||
|
- `limitations.scope` must not exceed the actual Tool query scopes.
|
||||||
|
- `limitations.missing_info` must name material gaps that constrain the conclusion.
|
||||||
|
|
||||||
|
Return exactly one JSON object matching the supplied DiagnosisDraft schema. Do not use Markdown fences, headings, prose before or after JSON, or a separate answer field.
|
||||||
@@ -0,0 +1,366 @@
|
|||||||
|
package com.superbiz.agent.harness.agent;
|
||||||
|
|
||||||
|
import com.alibaba.cloud.ai.graph.OverAllState;
|
||||||
|
import com.alibaba.cloud.ai.graph.RunnableConfig;
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.hook.HookPosition;
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.hook.HookPositions;
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.hook.ModelHook;
|
||||||
|
import com.fasterxml.jackson.databind.JsonNode;
|
||||||
|
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||||
|
import com.superbiz.agent.harness.contract.AnalysisKind;
|
||||||
|
import com.superbiz.agent.harness.contract.DiagnosisDraft;
|
||||||
|
import com.superbiz.agent.harness.contract.EvidenceStatus;
|
||||||
|
import com.superbiz.agent.harness.contract.PreviousTurn;
|
||||||
|
import com.superbiz.agent.harness.contract.SourceDocument;
|
||||||
|
import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
|
||||||
|
import com.superbiz.agent.harness.core.RunBudgetLimits;
|
||||||
|
import com.superbiz.agent.harness.core.RunContext;
|
||||||
|
import com.superbiz.agent.harness.core.RunState;
|
||||||
|
import com.superbiz.agent.harness.retry.HarnessRetryPolicies;
|
||||||
|
import com.superbiz.agent.harness.tool.boundary.ToolBoundaryErrorCode;
|
||||||
|
import com.superbiz.agent.harness.tool.boundary.ToolBoundaryResult;
|
||||||
|
import com.superbiz.agent.harness.tool.contract.AgentToolContracts;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.springframework.ai.chat.messages.AssistantMessage;
|
||||||
|
import org.springframework.ai.chat.messages.Message;
|
||||||
|
import org.springframework.ai.chat.messages.ToolResponseMessage;
|
||||||
|
import org.springframework.ai.chat.metadata.ChatResponseMetadata;
|
||||||
|
import org.springframework.ai.chat.metadata.DefaultUsage;
|
||||||
|
import org.springframework.ai.chat.model.ChatModel;
|
||||||
|
import org.springframework.ai.chat.model.ChatResponse;
|
||||||
|
import org.springframework.ai.chat.model.Generation;
|
||||||
|
import org.springframework.ai.chat.prompt.Prompt;
|
||||||
|
|
||||||
|
import java.time.Clock;
|
||||||
|
import java.time.Duration;
|
||||||
|
import java.util.ArrayList;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.concurrent.CompletableFuture;
|
||||||
|
import java.util.concurrent.atomic.AtomicInteger;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
|
||||||
|
class DiagnosisAgentUseCaseTest {
|
||||||
|
|
||||||
|
private static final DiagnosisAgentLimits LARGE_LIMITS =
|
||||||
|
new DiagnosisAgentLimits(20_000, 20_000, 40_000, 40_000);
|
||||||
|
|
||||||
|
private final ObjectMapper objectMapper = new ObjectMapper();
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void singleAgentRunsFrameworkToolLoopAndReturnsTypedDraft() {
|
||||||
|
DiagnosisHarnessCore core = core(new RunBudgetLimits(4, 3, 2,
|
||||||
|
100, 100, 200, 100_000));
|
||||||
|
RunContext context = core.startRun("session-agent", "run-agent");
|
||||||
|
AtomicInteger toolCalls = new AtomicInteger();
|
||||||
|
HarnessEvidenceTools tools = tools((runContext, id, arguments) -> {
|
||||||
|
core.beforeToolCall(runContext, AgentToolContracts.LOOKUP_KNOWLEDGE);
|
||||||
|
toolCalls.incrementAndGet();
|
||||||
|
return ToolBoundaryResult.ready(id, evidence(id), EvidenceStatus.EVIDENCE_FOUND);
|
||||||
|
});
|
||||||
|
ScriptedChatModel model = new ScriptedChatModel(3, 2,
|
||||||
|
toolCall("call-rag-1", "{\"query\":\"refund timeout\"}"),
|
||||||
|
new AssistantMessage(supportedDraft("call-rag-1")));
|
||||||
|
CapturingAuditHook audit = new CapturingAuditHook();
|
||||||
|
DiagnosisAgentUseCase useCase = useCase(core, model, tools, LARGE_LIMITS, audit);
|
||||||
|
PreviousTurn previous = new PreviousTurn(
|
||||||
|
"上一轮支付服务为什么超时?", "支付服务连接池已耗尽", "payment-service",
|
||||||
|
List.of("未覆盖退款服务"), List.of(new SourceDocument("doc-1", "连接池手册")));
|
||||||
|
|
||||||
|
DiagnosisDraft draft = useCase.execute(
|
||||||
|
context, new DiagnosisAgentInput("那退款服务呢?", previous));
|
||||||
|
|
||||||
|
assertEquals("退款服务出现连接池等待", draft.conclusion().text());
|
||||||
|
assertEquals(AnalysisKind.NORMAL, draft.analysis().get(0).kind());
|
||||||
|
assertEquals(List.of("call-rag-1"), draft.analysis().get(0).toolCallIds());
|
||||||
|
assertEquals(2, model.calls());
|
||||||
|
assertEquals(1, toolCalls.get());
|
||||||
|
assertEquals(2, context.budget().snapshot().modelCalls());
|
||||||
|
assertEquals(1, context.budget().snapshot().toolCalls());
|
||||||
|
assertEquals(10, context.budget().snapshot().totalTokens());
|
||||||
|
assertTrue(model.prompts().get(0).contains("那退款服务呢?"));
|
||||||
|
assertTrue(model.prompts().get(0).contains("支付服务连接池已耗尽"));
|
||||||
|
assertTrue(model.instructions().get(1).stream()
|
||||||
|
.filter(ToolResponseMessage.class::isInstance)
|
||||||
|
.map(ToolResponseMessage.class::cast)
|
||||||
|
.flatMap(message -> message.getResponses().stream())
|
||||||
|
.anyMatch(response -> response.id().equals("call-rag-1")
|
||||||
|
&& response.responseData().contains("call-rag-1")));
|
||||||
|
assertEquals(List.of("session-agent", "session-agent"), audit.sessionIds);
|
||||||
|
assertEquals(List.of("run-agent", "run-agent"), audit.runIds);
|
||||||
|
assertEquals(RunState.RUNNING, context.lifecycle().state());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void noEvidenceStopsWithNullConclusionAndNoRetry() {
|
||||||
|
DiagnosisHarnessCore core = core(defaultBudget());
|
||||||
|
RunContext context = core.startRun("session-no-evidence", "run-no-evidence");
|
||||||
|
AtomicInteger toolCalls = new AtomicInteger();
|
||||||
|
HarnessEvidenceTools tools = tools((runContext, id, arguments) -> {
|
||||||
|
core.beforeToolCall(runContext, AgentToolContracts.LOOKUP_KNOWLEDGE);
|
||||||
|
toolCalls.incrementAndGet();
|
||||||
|
return ToolBoundaryResult.ready(id, noEvidence(id), EvidenceStatus.NO_EVIDENCE);
|
||||||
|
});
|
||||||
|
ScriptedChatModel model = new ScriptedChatModel(1, 1,
|
||||||
|
toolCall("call-empty-1", "{\"query\":\"refund timeout\"}"),
|
||||||
|
new AssistantMessage(noEvidenceDraft("call-empty-1")));
|
||||||
|
|
||||||
|
DiagnosisDraft draft = useCase(core, model, tools, LARGE_LIMITS)
|
||||||
|
.execute(context, new DiagnosisAgentInput("退款为什么超时?", null));
|
||||||
|
|
||||||
|
assertNull(draft.conclusion());
|
||||||
|
assertEquals(AnalysisKind.NEGATIVE_OBSERVATION, draft.analysis().get(0).kind());
|
||||||
|
assertTrue(draft.limitations().missingInfo().contains("退款服务运行日志"));
|
||||||
|
assertEquals(2, model.calls());
|
||||||
|
assertEquals(1, toolCalls.get());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void fencedOutputFailsClosedWithoutAgentRetry() {
|
||||||
|
DiagnosisHarnessCore core = core(defaultBudget());
|
||||||
|
ScriptedChatModel model = new ScriptedChatModel(1, 1,
|
||||||
|
new AssistantMessage("```json\n" + supportedDraft("call-1") + "\n```"));
|
||||||
|
RunContext context = core.startRun("session-fenced", "run-fenced");
|
||||||
|
|
||||||
|
assertThrows(DiagnosisAgentOutputException.class, () ->
|
||||||
|
useCase(core, model, tools(errorInvoker()), LARGE_LIMITS)
|
||||||
|
.execute(context, new DiagnosisAgentInput("诊断超时", null)));
|
||||||
|
|
||||||
|
assertEquals(1, model.calls());
|
||||||
|
assertEquals(1, context.budget().snapshot().modelCalls());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void modelCallBudgetStopsNextReactRoundBeforeChatModel() {
|
||||||
|
DiagnosisHarnessCore core = core(new RunBudgetLimits(1, 2, 2,
|
||||||
|
100, 100, 200, 100_000));
|
||||||
|
RunContext context = core.startRun("session-model-budget", "run-model-budget");
|
||||||
|
HarnessEvidenceTools tools = tools((runContext, id, arguments) -> {
|
||||||
|
core.beforeToolCall(runContext, AgentToolContracts.LOOKUP_KNOWLEDGE);
|
||||||
|
return ToolBoundaryResult.ready(id, evidence(id), EvidenceStatus.EVIDENCE_FOUND);
|
||||||
|
});
|
||||||
|
ScriptedChatModel model = new ScriptedChatModel(1, 1,
|
||||||
|
toolCall("call-budget-1", "{\"query\":\"timeout\"}"),
|
||||||
|
new AssistantMessage(supportedDraft("call-budget-1")));
|
||||||
|
|
||||||
|
assertThrows(DiagnosisAgentOutputException.class, () ->
|
||||||
|
useCase(core, model, tools, LARGE_LIMITS)
|
||||||
|
.execute(context, new DiagnosisAgentInput("诊断超时", null)));
|
||||||
|
|
||||||
|
assertEquals(1, model.calls());
|
||||||
|
assertEquals(RunState.BUDGET_EXHAUSTED, context.lifecycle().state());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void actualTokenUsageExhaustionStopsDraftWithoutRetry() {
|
||||||
|
DiagnosisHarnessCore core = core(new RunBudgetLimits(2, 2, 2,
|
||||||
|
10, 10, 10, 100_000));
|
||||||
|
RunContext context = core.startRun("session-token-budget", "run-token-budget");
|
||||||
|
ScriptedChatModel model = new ScriptedChatModel(6, 6,
|
||||||
|
new AssistantMessage(supportedDraft("call-unused")));
|
||||||
|
|
||||||
|
assertThrows(DiagnosisAgentOutputException.class, () ->
|
||||||
|
useCase(core, model, tools(errorInvoker()), LARGE_LIMITS)
|
||||||
|
.execute(context, new DiagnosisAgentInput("诊断超时", null)));
|
||||||
|
|
||||||
|
assertEquals(1, model.calls());
|
||||||
|
assertEquals(12, context.budget().snapshot().totalTokens());
|
||||||
|
assertEquals(RunState.BUDGET_EXHAUSTED, context.lifecycle().state());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void contextLimitRejectsOriginalQueryBeforeModelCall() {
|
||||||
|
DiagnosisHarnessCore core = core(defaultBudget());
|
||||||
|
RunContext context = core.startRun("session-context-budget", "run-context-budget");
|
||||||
|
ScriptedChatModel model = new ScriptedChatModel(1, 1,
|
||||||
|
new AssistantMessage(supportedDraft("call-unused")));
|
||||||
|
DiagnosisAgentLimits limits = new DiagnosisAgentLimits(4, 100, 200, 1_000);
|
||||||
|
|
||||||
|
DiagnosisAgentLimitException failure = assertThrows(DiagnosisAgentLimitException.class, () ->
|
||||||
|
useCase(core, model, tools(errorInvoker()), limits)
|
||||||
|
.execute(context, new DiagnosisAgentInput("超过四个字节", null)));
|
||||||
|
|
||||||
|
assertEquals("query", failure.boundary());
|
||||||
|
assertEquals(0, model.calls());
|
||||||
|
assertEquals(0, context.budget().snapshot().modelCalls());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void draftLimitRejectsOversizedOutputWithoutSecondCall() {
|
||||||
|
DiagnosisHarnessCore core = core(defaultBudget());
|
||||||
|
RunContext context = core.startRun("session-draft-budget", "run-draft-budget");
|
||||||
|
ScriptedChatModel model = new ScriptedChatModel(1, 1,
|
||||||
|
new AssistantMessage(supportedDraft("call-unused")));
|
||||||
|
DiagnosisAgentLimits limits = new DiagnosisAgentLimits(1_000, 1_000, 2_000, 20);
|
||||||
|
|
||||||
|
DiagnosisAgentLimitException failure = assertThrows(DiagnosisAgentLimitException.class, () ->
|
||||||
|
useCase(core, model, tools(errorInvoker()), limits)
|
||||||
|
.execute(context, new DiagnosisAgentInput("诊断超时", null)));
|
||||||
|
|
||||||
|
assertEquals("draft", failure.boundary());
|
||||||
|
assertEquals(1, model.calls());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void rejectsInvalidLimitConfiguration() {
|
||||||
|
assertThrows(IllegalArgumentException.class,
|
||||||
|
() -> new DiagnosisAgentLimits(0, 1, 1, 1));
|
||||||
|
assertThrows(IllegalArgumentException.class,
|
||||||
|
() -> new DiagnosisAgentInput(" ", null));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void generatedDraftSchemaAllowsExplicitNullConclusion() throws Exception {
|
||||||
|
JsonNode schema = objectMapper.readTree(new DiagnosisDraftOutputSchema(objectMapper).getJsonSchema());
|
||||||
|
|
||||||
|
JsonNode conclusionTypes = schema.path("properties").path("conclusion").path("type");
|
||||||
|
assertTrue(conclusionTypes.isArray());
|
||||||
|
assertTrue(java.util.stream.StreamSupport.stream(conclusionTypes.spliterator(), false)
|
||||||
|
.map(JsonNode::asText)
|
||||||
|
.anyMatch("object"::equals));
|
||||||
|
assertTrue(java.util.stream.StreamSupport.stream(conclusionTypes.spliterator(), false)
|
||||||
|
.map(JsonNode::asText)
|
||||||
|
.anyMatch("null"::equals));
|
||||||
|
}
|
||||||
|
|
||||||
|
private DiagnosisAgentUseCase useCase(DiagnosisHarnessCore core,
|
||||||
|
ChatModel model,
|
||||||
|
HarnessEvidenceTools tools,
|
||||||
|
DiagnosisAgentLimits limits,
|
||||||
|
ModelHook... hooks) {
|
||||||
|
DiagnosisAgentFactory factory = new DiagnosisAgentFactory(
|
||||||
|
model, core, tools, objectMapper, List.of(hooks));
|
||||||
|
return new DiagnosisAgentUseCase(core, factory, objectMapper, limits);
|
||||||
|
}
|
||||||
|
|
||||||
|
private HarnessEvidenceTools tools(EvidenceToolInvoker ragInvoker) {
|
||||||
|
EvidenceToolInvoker unused = errorInvoker();
|
||||||
|
return new HarnessEvidenceTools(ragInvoker, unused, unused);
|
||||||
|
}
|
||||||
|
|
||||||
|
private EvidenceToolInvoker errorInvoker() {
|
||||||
|
return (context, id, arguments) ->
|
||||||
|
ToolBoundaryResult.error(id, ToolBoundaryErrorCode.INVALID_REQUEST);
|
||||||
|
}
|
||||||
|
|
||||||
|
private DiagnosisHarnessCore core(RunBudgetLimits budget) {
|
||||||
|
return new DiagnosisHarnessCore(
|
||||||
|
Clock.systemUTC(), () -> "unused-run", Duration.ofMinutes(5),
|
||||||
|
budget, HarnessRetryPolicies.strict());
|
||||||
|
}
|
||||||
|
|
||||||
|
private RunBudgetLimits defaultBudget() {
|
||||||
|
return new RunBudgetLimits(4, 3, 2, 100, 100, 200, 100_000);
|
||||||
|
}
|
||||||
|
|
||||||
|
private AssistantMessage toolCall(String id, String arguments) {
|
||||||
|
return AssistantMessage.builder()
|
||||||
|
.content("")
|
||||||
|
.toolCalls(List.of(new AssistantMessage.ToolCall(
|
||||||
|
id, "function", AgentToolContracts.LOOKUP_KNOWLEDGE, arguments)))
|
||||||
|
.build();
|
||||||
|
}
|
||||||
|
|
||||||
|
private String evidence(String id) {
|
||||||
|
return "{\"evidence_status\":\"EVIDENCE_FOUND\",\"tool_call_id\":\"" + id
|
||||||
|
+ "\",\"query\":\"refund timeout\",\"evidence\":[{\"document_id\":\"doc-1\","
|
||||||
|
+ "\"source\":\"runbook\",\"title\":\"退款连接池\",\"breadcrumb\":[\"运行\"],"
|
||||||
|
+ "\"excerpt\":\"active=50 max=50\"}],\"returned_count\":1,\"truncated\":false}";
|
||||||
|
}
|
||||||
|
|
||||||
|
private String noEvidence(String id) {
|
||||||
|
return "{\"evidence_status\":\"NO_EVIDENCE\",\"tool_call_id\":\"" + id
|
||||||
|
+ "\",\"query\":\"refund timeout\",\"evidence\":[],\"returned_count\":0,\"truncated\":false}";
|
||||||
|
}
|
||||||
|
|
||||||
|
private String supportedDraft(String toolCallId) {
|
||||||
|
return """
|
||||||
|
{
|
||||||
|
"conclusion":{"text":"退款服务出现连接池等待","based_on_analysis_ids":["a-1"]},
|
||||||
|
"analysis":[{"analysis_id":"a-1","kind":"NORMAL","text":"连接池 active 达到上限","tool_call_ids":["%s"]}],
|
||||||
|
"action_plan":[{"action":"检查连接归还路径","based_on_analysis_ids":["a-1"],"requires_human_confirmation":false}],
|
||||||
|
"recommendations":[{"text":"增加连接池饱和告警","based_on_analysis_ids":["a-1"]}],
|
||||||
|
"limitations":{"scope":"refund-service","missing_info":[]}
|
||||||
|
}
|
||||||
|
""".formatted(toolCallId);
|
||||||
|
}
|
||||||
|
|
||||||
|
private String noEvidenceDraft(String toolCallId) {
|
||||||
|
return """
|
||||||
|
{
|
||||||
|
"conclusion":null,
|
||||||
|
"analysis":[{"analysis_id":"a-1","kind":"NEGATIVE_OBSERVATION","text":"知识库查询范围内未找到退款超时证据","tool_call_ids":["%s"]}],
|
||||||
|
"action_plan":[],
|
||||||
|
"recommendations":[],
|
||||||
|
"limitations":{"scope":"lookup query=refund timeout","missing_info":["退款服务运行日志"]}
|
||||||
|
}
|
||||||
|
""".formatted(toolCallId);
|
||||||
|
}
|
||||||
|
|
||||||
|
private static final class ScriptedChatModel implements ChatModel {
|
||||||
|
|
||||||
|
private final int promptTokens;
|
||||||
|
private final int completionTokens;
|
||||||
|
private final List<AssistantMessage> responses;
|
||||||
|
private final List<String> prompts = new ArrayList<>();
|
||||||
|
private final List<List<Message>> instructions = new ArrayList<>();
|
||||||
|
private final AtomicInteger calls = new AtomicInteger();
|
||||||
|
|
||||||
|
private ScriptedChatModel(int promptTokens, int completionTokens,
|
||||||
|
AssistantMessage... responses) {
|
||||||
|
this.promptTokens = promptTokens;
|
||||||
|
this.completionTokens = completionTokens;
|
||||||
|
this.responses = List.of(responses);
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public ChatResponse call(Prompt prompt) {
|
||||||
|
prompts.add(prompt.getContents());
|
||||||
|
instructions.add(List.copyOf(prompt.getInstructions()));
|
||||||
|
int index = calls.getAndIncrement();
|
||||||
|
if (index >= responses.size()) {
|
||||||
|
throw new AssertionError("unexpected model retry");
|
||||||
|
}
|
||||||
|
ChatResponseMetadata metadata = ChatResponseMetadata.builder()
|
||||||
|
.usage(new DefaultUsage(promptTokens, completionTokens))
|
||||||
|
.build();
|
||||||
|
return new ChatResponse(List.of(new Generation(responses.get(index))), metadata);
|
||||||
|
}
|
||||||
|
|
||||||
|
int calls() {
|
||||||
|
return calls.get();
|
||||||
|
}
|
||||||
|
|
||||||
|
List<String> prompts() {
|
||||||
|
return prompts;
|
||||||
|
}
|
||||||
|
|
||||||
|
List<List<Message>> instructions() {
|
||||||
|
return instructions;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
@HookPositions({HookPosition.BEFORE_MODEL})
|
||||||
|
private static final class CapturingAuditHook extends ModelHook {
|
||||||
|
|
||||||
|
private final List<String> sessionIds = new ArrayList<>();
|
||||||
|
private final List<String> runIds = new ArrayList<>();
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public String getName() {
|
||||||
|
return "capturing_agent_step_audit";
|
||||||
|
}
|
||||||
|
|
||||||
|
@Override
|
||||||
|
public CompletableFuture<java.util.Map<String, Object>> beforeModel(
|
||||||
|
OverAllState state, RunnableConfig config) {
|
||||||
|
sessionIds.add(config.metadata("sessionId").map(Object::toString).orElseThrow());
|
||||||
|
runIds.add(config.metadata("runId").map(Object::toString).orElseThrow());
|
||||||
|
return CompletableFuture.completedFuture(java.util.Map.of());
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,184 @@
|
|||||||
|
package com.superbiz.agent.harness.agent;
|
||||||
|
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.interceptor.ToolCallRequest;
|
||||||
|
import com.alibaba.cloud.ai.graph.agent.interceptor.ToolCallResponse;
|
||||||
|
import com.fasterxml.jackson.databind.JsonNode;
|
||||||
|
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||||
|
import com.superbiz.agent.harness.contract.EvidenceStatus;
|
||||||
|
import com.superbiz.agent.harness.core.DiagnosisHarnessCore;
|
||||||
|
import com.superbiz.agent.harness.core.RunBudgetLimits;
|
||||||
|
import com.superbiz.agent.harness.core.RunContext;
|
||||||
|
import com.superbiz.agent.harness.retry.HarnessRetryPolicies;
|
||||||
|
import com.superbiz.agent.harness.tool.adapter.MysqlToolAdapter;
|
||||||
|
import com.superbiz.agent.harness.tool.adapter.QueryLogsToolAdapter;
|
||||||
|
import com.superbiz.agent.harness.tool.adapter.RagToolAdapter;
|
||||||
|
import com.superbiz.agent.harness.tool.boundary.ToolBoundaryErrorCode;
|
||||||
|
import com.superbiz.agent.harness.tool.boundary.ToolBoundaryResult;
|
||||||
|
import com.superbiz.agent.harness.tool.boundary.ToolCallRequestEnvelope;
|
||||||
|
import com.superbiz.agent.harness.tool.contract.AgentToolContracts;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
import org.mockito.ArgumentCaptor;
|
||||||
|
import org.springframework.ai.tool.ToolCallback;
|
||||||
|
import org.springframework.ai.tool.execution.ToolExecutionException;
|
||||||
|
|
||||||
|
import java.time.Clock;
|
||||||
|
import java.time.Duration;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.concurrent.atomic.AtomicInteger;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertThrows;
|
||||||
|
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||||
|
import static org.mockito.ArgumentMatchers.any;
|
||||||
|
import static org.mockito.ArgumentMatchers.eq;
|
||||||
|
import static org.mockito.Mockito.mock;
|
||||||
|
import static org.mockito.Mockito.verify;
|
||||||
|
import static org.mockito.Mockito.when;
|
||||||
|
|
||||||
|
class HarnessToolInterceptorTest {
|
||||||
|
|
||||||
|
private final ObjectMapper objectMapper = new ObjectMapper();
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void exposesOnlyFrozenToolDefinitionsWithoutAgentSuppliedCallId() {
|
||||||
|
HarnessEvidenceTools tools = fakeTools((context, id, arguments) -> ready(id));
|
||||||
|
|
||||||
|
List<ToolCallback> callbacks = tools.callbacks();
|
||||||
|
|
||||||
|
assertEquals(List.of(
|
||||||
|
AgentToolContracts.LOOKUP_KNOWLEDGE,
|
||||||
|
AgentToolContracts.QUERY_LOGS,
|
||||||
|
AgentToolContracts.QUERY_MYSQL),
|
||||||
|
callbacks.stream().map(callback -> callback.getToolDefinition().name()).toList());
|
||||||
|
callbacks.forEach(callback -> {
|
||||||
|
assertFalse(callback.getToolDefinition().inputSchema().contains("tool_call_id"));
|
||||||
|
assertFalse(callback.getToolDefinition().description().isBlank());
|
||||||
|
assertThrows(ToolExecutionException.class, () -> callback.call("{}"));
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void adapterBridgeUsesExactFrameworkIdAndRunIdentity() {
|
||||||
|
RagToolAdapter ragAdapter = mock(RagToolAdapter.class);
|
||||||
|
QueryLogsToolAdapter logsAdapter = mock(QueryLogsToolAdapter.class);
|
||||||
|
MysqlToolAdapter mysqlAdapter = mock(MysqlToolAdapter.class);
|
||||||
|
RunContext context = context(3, 3, 3);
|
||||||
|
when(ragAdapter.execute(eq(context), any())).thenAnswer(invocation -> {
|
||||||
|
ToolCallRequestEnvelope envelope = invocation.getArgument(1);
|
||||||
|
return ready(envelope.toolCallId());
|
||||||
|
});
|
||||||
|
HarnessEvidenceTools tools = HarnessEvidenceTools.fromAdapters(ragAdapter, logsAdapter, mysqlAdapter);
|
||||||
|
HarnessToolInterceptor interceptor = new HarnessToolInterceptor(context, tools, objectMapper);
|
||||||
|
|
||||||
|
ToolCallResponse response = interceptor.interceptToolCall(
|
||||||
|
request(AgentToolContracts.LOOKUP_KNOWLEDGE, "framework-call-7", "{\"query\":\"timeout\"}"),
|
||||||
|
ignored -> {
|
||||||
|
throw new AssertionError("registered Tool must not bypass Harness interceptor");
|
||||||
|
});
|
||||||
|
|
||||||
|
ArgumentCaptor<ToolCallRequestEnvelope> envelope = ArgumentCaptor.forClass(ToolCallRequestEnvelope.class);
|
||||||
|
verify(ragAdapter).execute(eq(context), envelope.capture());
|
||||||
|
assertEquals("framework-call-7", envelope.getValue().toolCallId());
|
||||||
|
assertEquals(context.runId(), envelope.getValue().runId());
|
||||||
|
assertEquals(AgentToolContracts.LOOKUP_KNOWLEDGE, envelope.getValue().toolName());
|
||||||
|
assertEquals("framework-call-7", response.getToolCallId());
|
||||||
|
assertTrue(response.getResult().contains("framework-call-7"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void returnsStableErrorObservationWithoutRawFailureDetail() throws Exception {
|
||||||
|
HarnessEvidenceTools tools = fakeTools((context, id, arguments) ->
|
||||||
|
ToolBoundaryResult.error(id, ToolBoundaryErrorCode.TOOL_EXECUTION_ERROR));
|
||||||
|
HarnessToolInterceptor interceptor = new HarnessToolInterceptor(
|
||||||
|
context(3, 3, 3), tools, objectMapper);
|
||||||
|
|
||||||
|
ToolCallResponse response = interceptor.interceptToolCall(
|
||||||
|
request(AgentToolContracts.LOOKUP_KNOWLEDGE, "call-error", "{\"query\":\"secret\"}"),
|
||||||
|
ignored -> null);
|
||||||
|
|
||||||
|
JsonNode observation = objectMapper.readTree(response.getResult());
|
||||||
|
assertTrue(response.isError());
|
||||||
|
assertEquals("ERROR", observation.path("evidence_status").asText());
|
||||||
|
assertEquals("call-error", observation.path("tool_call_id").asText());
|
||||||
|
assertEquals("TOOL_EXECUTION_ERROR", observation.path("error_code").asText());
|
||||||
|
assertFalse(response.getResult().contains("secret"));
|
||||||
|
assertFalse(response.getResult().contains("raw_response"));
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void delegatesUnknownToolWithoutCreatingEvidenceInvocation() {
|
||||||
|
AtomicInteger evidenceCalls = new AtomicInteger();
|
||||||
|
HarnessEvidenceTools tools = fakeTools((context, id, arguments) -> {
|
||||||
|
evidenceCalls.incrementAndGet();
|
||||||
|
return ready(id);
|
||||||
|
});
|
||||||
|
AtomicInteger handlerCalls = new AtomicInteger();
|
||||||
|
HarnessToolInterceptor interceptor = new HarnessToolInterceptor(
|
||||||
|
context(3, 3, 3), tools, objectMapper);
|
||||||
|
|
||||||
|
ToolCallResponse response = interceptor.interceptToolCall(
|
||||||
|
request("unknown_tool", "call-unknown", "{}"),
|
||||||
|
request -> {
|
||||||
|
handlerCalls.incrementAndGet();
|
||||||
|
return ToolCallResponse.error(request.getToolCallId(), request.getToolName(), "not registered");
|
||||||
|
});
|
||||||
|
|
||||||
|
assertTrue(response.isError());
|
||||||
|
assertEquals(1, handlerCalls.get());
|
||||||
|
assertEquals(0, evidenceCalls.get());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void oneFrameworkActionConsumesOneToolBudgetWithoutRetry() {
|
||||||
|
DiagnosisHarnessCore core = core(3, 1, 1, 10_000);
|
||||||
|
RunContext context = core.startRun("session-tool-budget", "run-tool-budget");
|
||||||
|
AtomicInteger invocations = new AtomicInteger();
|
||||||
|
HarnessEvidenceTools tools = fakeTools((runContext, id, arguments) -> {
|
||||||
|
core.beforeToolCall(runContext, AgentToolContracts.LOOKUP_KNOWLEDGE);
|
||||||
|
invocations.incrementAndGet();
|
||||||
|
return ready(id);
|
||||||
|
});
|
||||||
|
HarnessToolInterceptor interceptor = new HarnessToolInterceptor(context, tools, objectMapper);
|
||||||
|
|
||||||
|
interceptor.interceptToolCall(
|
||||||
|
request(AgentToolContracts.LOOKUP_KNOWLEDGE, "call-1", "{\"query\":\"one\"}"),
|
||||||
|
ignored -> null);
|
||||||
|
|
||||||
|
assertEquals(1, invocations.get());
|
||||||
|
assertEquals(1, context.budget().snapshot().toolCalls());
|
||||||
|
}
|
||||||
|
|
||||||
|
private HarnessEvidenceTools fakeTools(EvidenceToolInvoker ragInvoker) {
|
||||||
|
EvidenceToolInvoker unused = (context, id, arguments) ->
|
||||||
|
ToolBoundaryResult.error(id, ToolBoundaryErrorCode.INVALID_REQUEST);
|
||||||
|
return new HarnessEvidenceTools(ragInvoker, unused, unused);
|
||||||
|
}
|
||||||
|
|
||||||
|
private ToolCallRequest request(String toolName, String toolCallId, String arguments) {
|
||||||
|
return new ToolCallRequest(toolName, arguments, toolCallId, java.util.Map.of());
|
||||||
|
}
|
||||||
|
|
||||||
|
private ToolBoundaryResult ready(String toolCallId) {
|
||||||
|
return ToolBoundaryResult.ready(toolCallId,
|
||||||
|
"{\"evidence_status\":\"EVIDENCE_FOUND\",\"tool_call_id\":\""
|
||||||
|
+ toolCallId + "\",\"evidence\":[]}",
|
||||||
|
EvidenceStatus.EVIDENCE_FOUND);
|
||||||
|
}
|
||||||
|
|
||||||
|
private RunContext context(int maxModelCalls, int maxToolCalls, int maxCallsPerTool) {
|
||||||
|
return core(maxModelCalls, maxToolCalls, maxCallsPerTool, 10_000)
|
||||||
|
.startRun("session-tool", "run-tool");
|
||||||
|
}
|
||||||
|
|
||||||
|
private DiagnosisHarnessCore core(int maxModelCalls, int maxToolCalls,
|
||||||
|
int maxCallsPerTool, long maxTokens) {
|
||||||
|
return new DiagnosisHarnessCore(
|
||||||
|
Clock.systemUTC(),
|
||||||
|
() -> "unused-run",
|
||||||
|
Duration.ofMinutes(5),
|
||||||
|
new RunBudgetLimits(maxModelCalls, maxToolCalls, maxCallsPerTool,
|
||||||
|
maxTokens, maxTokens, maxTokens, 100_000),
|
||||||
|
HarnessRetryPolicies.strict());
|
||||||
|
}
|
||||||
|
}
|
||||||
Reference in New Issue
Block a user