Files

107 lines
8.8 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
## Context
阶段 2 的 `DiagnosisHarnessCore` 已提供显式 `RunContext`、模型/Tool/Token/字节预算、取消和唯一生命周期;阶段 3A-3C 已提供 canonical invocation、`ToolBoundary` 以及 RAG、日志、MySQL adapter。当前这些组件没有生产 Agent 消费者,公开 Chat 仍由 `ChatService` 创建单 Agent 简单路径或 Planner/Executor/Verifier/Composer 多 Agent 路径。
Spring AI Alibaba 1.1.2.0 的 `ReactAgent` 自带模型与 Tool 交替执行的 ReAct loop。`ModelInterceptor` 包围每次模型调用;`ToolInterceptor` 获得 `ToolCallRequest.getToolCallId()`、参数和 `RunnableConfig` metadata;`BeanOutputConverter` 可从输出类型生成格式提示,但 `call` 最终仍返回 `AssistantMessage`。因此本阶段可以使用框架扩展点接入 Harness,无需自行编写循环或修改框架。
## Goals / Non-Goals
**Goals:**
- 创建一个且仅一个拥有 Tool loop 的 Diagnosis ReactAgent。
- 以显式 `RunContext` 执行当前 Query 和可选 `PreviousTurn`,输入失败时 fail closed,不截断原始 Query。
- 将框架原始 Tool Call ID 传给阶段 3B/3C adapter,并只把有界 `agent_result` 返回模型。
- 对每次模型调用、Token、Tool 调用和输入/输出上下文字节实施确定性预算。
- 输出可严格反序列化的 `DiagnosisDraft`;无证据时明确停止,不补造结论。
- 保留可注入 AgentStep Hook、RunContext 生命周期和 canonical Tool invocation 审计边界。
**Non-Goals:**
- 不实现 EvidenceGuard、SemanticGuard、结构修复、引用展开或最终发布。
- 不创建 Intent Router、不选择 previous turn、不持久化 DiagnosisRun。
- 不修改 Controller、SSE、ChatService、AiOpsService 或旧多 Agent 链路。
- 不实现外层 Graph、手写 ReAct while、Agent 自动重跑或 Tool 自动重试。
- 不增加真实日志源或新的 evidence Tool。
## Decisions
### One per-run ReactAgent uses the framework loop
`DiagnosisAgentFactory` 为每个 `RunContext` 构建一个名为 `diagnosis_agent` 的 `ReactAgent`,关闭并行 Tool 执行并使用单一 Prompt。内部 `DiagnosisAgentUseCase` 对 Agent 只调用一次;一次调用内由框架决定正常的模型/Tool轮次。
替代方案是复用 `ChatService.createReactAgent`,但它注册旧 Tool、旧 Hook 和完整旧运行职责,无法保证 Harness boundary、极简上下文和不接公开入口。另一个替代方案是手写 while,直接违反 ISS-014。
### Model and Tool control use different framework boundaries
`HarnessModelInterceptor` 在每次模型调用前执行 `core.beforeModelCall(context)`,调用完成后读取 `ChatResponseMetadata.Usage` 并执行 `core.recordTokens`。内部用例设置 `_stream_=false`,保证 interceptor 可取得完整 `ChatResponse` 和 Token Usage。
`HarnessToolInterceptor` 读取框架原始 Tool Call ID,并按 Tool 名调用 `HarnessEvidenceTools` 中的 adapter bridge。bridge 创建 `ToolCallRequestEnvelope(runId, frameworkId, toolName, arguments, true, true)`;授权来自注册到当前 Diagnosis Agent 的固定 Tool 集合。Tool 成功时只返回 `agent_result`,失败时只返回稳定的 `evidence_status/tool_call_id/error_code`,不返回 raw response、内部异常或 invocation lifecycle。
Tool 调用预算只在 `ToolBoundary` 中 reserve;Agent interceptor 不重复调用 `beforeToolCall`。这保持阶段 3A 的 canonical record 与预算原子边界。
### Tool definitions and execution are registered together
`HarnessEvidenceTools` 固定暴露 `lookup_knowledge`、`query_logs`、`query_mysql` 三个 Spring `ToolCallback` 定义,输入类型分别复用已冻结的 Request records,描述复用 `AgentToolContracts`。callback 本体不允许绕过 interceptor 直接执行;实际执行映射与定义在同一 registry 中,构造时拒绝缺项或重复项。
替代方案是在 `ToolCallback.call` 中读取 ThreadLocal 或生成 ID,都会违反 RunContext 和 canonical ID 契约。
### Input and output remain typed and bounded
`DiagnosisAgentInput` 只包含非空原始 `query` 和可选 `PreviousTurn`。`DiagnosisAgentLimits` 配置 query、previous turn、总输入和 Draft 的 UTF-8 字节上限。内部用例先分别验证,再将固定 `{query, previous_turn}` JSON 作为唯一 User 输入,并通过 `core.reserveRunBytes` 计入 Run 容量;任何超限都在模型调用前失败,不做语义截断。
Agent 使用基于 `BeanOutputConverter<DiagnosisDraft>` 生成的格式提示,并通过 schema post-process 明确允许冻结契约中的 `conclusion=null`;Factory 以 `.outputSchema(...)` 注入该格式。最终文本先检查 UTF-8 上限并计入容量,再由注入的 `ObjectMapper` 严格解析为冻结 record。Markdown fence、前后说明、未知结构或空输出不做修复和自动重试;阶段 5 再实现一次显式无 Tool 结构修复。
### Previous turn is context, not current evidence
`PreviousTurn` 按固定 record 原样序列化,仅帮助理解追问。Prompt 明确禁止引用旧 Tool Call ID 或把 previous turn 当作当前 Run 的 evidence;当前 `DiagnosisDraft.analysis[*].tool_call_ids` 只能来自本 Run Tool result。
### Prompt owns diagnosis semantics, Harness owns control
唯一 classpath Prompt 要求结论先行、所有分析绑定 Tool Call IDs、`NO_EVIDENCE` 只能形成限定范围的 `NEGATIVE_OBSERVATION`。没有足够 `EVIDENCE_FOUND` 时 `conclusion=null`,在 `limitations` 记录范围和缺失信息并结束,不把 `NO_EVIDENCE` 推导为系统健康或根因排除。
阶段 4 不实现物理或语义 Guard,因此内部返回值叫 Draft,不能直接公开发布。
### Audit remains injectable and explicit
Factory 接收可选框架 `Hook` 列表。阶段 6A 可以注入现有 `AgentLoggingHook`;内部用例始终在 `RunnableConfig` metadata 写入 sessionId/runId,所以该 Hook 不需要读取 ThreadLocal fallback。模型/Tool 预算和生命周期仍绑定 RunContext,Tool canonical audit 仍由 `ToolBoundary`/store 负责。
## Architecture And Interface Impact
```text
DiagnosisAgentInput(query, previousTurn) + RunContext
-> DiagnosisAgentUseCase: validate/serialize/reserve context bytes
-> DiagnosisAgentFactory: one ReactAgent + one Prompt
-> HarnessModelInterceptor -> DiagnosisHarnessCore -> ChatModel
-> framework ReAct loop
-> HarnessToolInterceptor -> HarnessEvidenceTools -> 3B/3C adapters
-> ToolBoundary -> canonical invocation store
-> bounded AssistantMessage text
-> strict ObjectMapper -> DiagnosisDraft
```
- 数据所有权:use case 拥有本次输入和 Draft 解析;RunContext 拥有预算/取消/生命周期;ToolBoundary/store 拥有调用记录;Agent 只拥有诊断语义。
- 接口影响:L2 内部接口。新增阶段 6A 可消费的 Java 类型,不改变公开 HTTP/SSE、JPA、数据库或旧 Service 方法。
- 生命周期:本 use case 不把成功 Draft 标记为 Run SUCCESS,因为阶段 5 Guards 尚未执行;预算、取消和超时可由 Core 先行终止 Run,最终持久化映射留给阶段 6A。
## Risks / Trade-offs
- [Risk] Provider 返回 fenced JSON 或附加说明导致解析失败 → Mitigation:阶段 4 fail closed 且不重跑;测试固定这一行为,阶段 5 才允许一次无 Tool 结构修复。
- [Risk] Token Usage 缺失或为 null → Mitigation:模型调用次数仍被强制;仅在可用且非负时记录实际 Token,并在 acceptance 标记 provider 元数据依赖。
- [Risk] Tool interceptor 错误泄露内部异常 → Mitigation:只返回 allowlisted error code,不返回 message/stack/raw response。
- [Risk] 输入和 Tool 投影共同消耗 Run bytes,配置过小会过早耗尽 → Mitigation:所有限制可配置,usage 可观察,不在本阶段硬编码生产校准值。
- [Risk] 复用旧 AgentLoggingHook 仍包含 ThreadLocal fallback → Mitigation:新链路总是提供 metadata;阶段 7 物理清理前不修改旧 Hook,focused test 验证相同 sessionId/runId 传播。
- [Risk] ReactAgent 内部基于框架 StateGraph → Mitigation:这是框架原生实现细节;本项目不再包一层业务 Graph 或多 Agent workflow。
## Migration Plan
1. 新增内部 Agent 类型、Prompt、interceptors 和 Tool registry,不注册 Controller bean。
2. 使用 scripted ChatModel 和 fake/adapted Tool 完成内部 tool-loop、output、budget、no-evidence 和审计测试。
3. 保持公开 Chat 与旧多 Agent 代码 diff 为空,Archive 并提交阶段 4。
4. 阶段 5 在内部 Draft 后增加 Guards;阶段 6A 再创建运行应用用例并装配真实审计/previous turn;阶段 6B 原子切换公开入口。
Rollback:删除新增内部包、Prompt 和测试即可;不存在公开协议、数据迁移或运行时切换。
## Open Questions
无。生产预算默认值、公开切换和最终发布策略分别由配置校准、阶段 6B 和阶段 5 处理。