feat(harness): add single diagnosis react agent
This commit is contained in:
@@ -0,0 +1 @@
|
||||
ready
|
||||
@@ -0,0 +1 @@
|
||||
committed
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-21
|
||||
@@ -0,0 +1,106 @@
|
||||
## Context
|
||||
|
||||
阶段 2 的 `DiagnosisHarnessCore` 已提供显式 `RunContext`、模型/Tool/Token/字节预算、取消和唯一生命周期;阶段 3A-3C 已提供 canonical invocation、`ToolBoundary` 以及 RAG、日志、MySQL adapter。当前这些组件没有生产 Agent 消费者,公开 Chat 仍由 `ChatService` 创建单 Agent 简单路径或 Planner/Executor/Verifier/Composer 多 Agent 路径。
|
||||
|
||||
Spring AI Alibaba 1.1.2.0 的 `ReactAgent` 自带模型与 Tool 交替执行的 ReAct loop。`ModelInterceptor` 包围每次模型调用;`ToolInterceptor` 获得 `ToolCallRequest.getToolCallId()`、参数和 `RunnableConfig` metadata;`BeanOutputConverter` 可从输出类型生成格式提示,但 `call` 最终仍返回 `AssistantMessage`。因此本阶段可以使用框架扩展点接入 Harness,无需自行编写循环或修改框架。
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- 创建一个且仅一个拥有 Tool loop 的 Diagnosis ReactAgent。
|
||||
- 以显式 `RunContext` 执行当前 Query 和可选 `PreviousTurn`,输入失败时 fail closed,不截断原始 Query。
|
||||
- 将框架原始 Tool Call ID 传给阶段 3B/3C adapter,并只把有界 `agent_result` 返回模型。
|
||||
- 对每次模型调用、Token、Tool 调用和输入/输出上下文字节实施确定性预算。
|
||||
- 输出可严格反序列化的 `DiagnosisDraft`;无证据时明确停止,不补造结论。
|
||||
- 保留可注入 AgentStep Hook、RunContext 生命周期和 canonical Tool invocation 审计边界。
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- 不实现 EvidenceGuard、SemanticGuard、结构修复、引用展开或最终发布。
|
||||
- 不创建 Intent Router、不选择 previous turn、不持久化 DiagnosisRun。
|
||||
- 不修改 Controller、SSE、ChatService、AiOpsService 或旧多 Agent 链路。
|
||||
- 不实现外层 Graph、手写 ReAct while、Agent 自动重跑或 Tool 自动重试。
|
||||
- 不增加真实日志源或新的 evidence Tool。
|
||||
|
||||
## Decisions
|
||||
|
||||
### One per-run ReactAgent uses the framework loop
|
||||
|
||||
`DiagnosisAgentFactory` 为每个 `RunContext` 构建一个名为 `diagnosis_agent` 的 `ReactAgent`,关闭并行 Tool 执行并使用单一 Prompt。内部 `DiagnosisAgentUseCase` 对 Agent 只调用一次;一次调用内由框架决定正常的模型/Tool轮次。
|
||||
|
||||
替代方案是复用 `ChatService.createReactAgent`,但它注册旧 Tool、旧 Hook 和完整旧运行职责,无法保证 Harness boundary、极简上下文和不接公开入口。另一个替代方案是手写 while,直接违反 ISS-014。
|
||||
|
||||
### Model and Tool control use different framework boundaries
|
||||
|
||||
`HarnessModelInterceptor` 在每次模型调用前执行 `core.beforeModelCall(context)`,调用完成后读取 `ChatResponseMetadata.Usage` 并执行 `core.recordTokens`。内部用例设置 `_stream_=false`,保证 interceptor 可取得完整 `ChatResponse` 和 Token Usage。
|
||||
|
||||
`HarnessToolInterceptor` 读取框架原始 Tool Call ID,并按 Tool 名调用 `HarnessEvidenceTools` 中的 adapter bridge。bridge 创建 `ToolCallRequestEnvelope(runId, frameworkId, toolName, arguments, true, true)`;授权来自注册到当前 Diagnosis Agent 的固定 Tool 集合。Tool 成功时只返回 `agent_result`,失败时只返回稳定的 `evidence_status/tool_call_id/error_code`,不返回 raw response、内部异常或 invocation lifecycle。
|
||||
|
||||
Tool 调用预算只在 `ToolBoundary` 中 reserve;Agent interceptor 不重复调用 `beforeToolCall`。这保持阶段 3A 的 canonical record 与预算原子边界。
|
||||
|
||||
### Tool definitions and execution are registered together
|
||||
|
||||
`HarnessEvidenceTools` 固定暴露 `lookup_knowledge`、`query_logs`、`query_mysql` 三个 Spring `ToolCallback` 定义,输入类型分别复用已冻结的 Request records,描述复用 `AgentToolContracts`。callback 本体不允许绕过 interceptor 直接执行;实际执行映射与定义在同一 registry 中,构造时拒绝缺项或重复项。
|
||||
|
||||
替代方案是在 `ToolCallback.call` 中读取 ThreadLocal 或生成 ID,都会违反 RunContext 和 canonical ID 契约。
|
||||
|
||||
### Input and output remain typed and bounded
|
||||
|
||||
`DiagnosisAgentInput` 只包含非空原始 `query` 和可选 `PreviousTurn`。`DiagnosisAgentLimits` 配置 query、previous turn、总输入和 Draft 的 UTF-8 字节上限。内部用例先分别验证,再将固定 `{query, previous_turn}` JSON 作为唯一 User 输入,并通过 `core.reserveRunBytes` 计入 Run 容量;任何超限都在模型调用前失败,不做语义截断。
|
||||
|
||||
Agent 使用基于 `BeanOutputConverter<DiagnosisDraft>` 生成的格式提示,并通过 schema post-process 明确允许冻结契约中的 `conclusion=null`;Factory 以 `.outputSchema(...)` 注入该格式。最终文本先检查 UTF-8 上限并计入容量,再由注入的 `ObjectMapper` 严格解析为冻结 record。Markdown fence、前后说明、未知结构或空输出不做修复和自动重试;阶段 5 再实现一次显式无 Tool 结构修复。
|
||||
|
||||
### Previous turn is context, not current evidence
|
||||
|
||||
`PreviousTurn` 按固定 record 原样序列化,仅帮助理解追问。Prompt 明确禁止引用旧 Tool Call ID 或把 previous turn 当作当前 Run 的 evidence;当前 `DiagnosisDraft.analysis[*].tool_call_ids` 只能来自本 Run Tool result。
|
||||
|
||||
### Prompt owns diagnosis semantics, Harness owns control
|
||||
|
||||
唯一 classpath Prompt 要求结论先行、所有分析绑定 Tool Call IDs、`NO_EVIDENCE` 只能形成限定范围的 `NEGATIVE_OBSERVATION`。没有足够 `EVIDENCE_FOUND` 时 `conclusion=null`,在 `limitations` 记录范围和缺失信息并结束,不把 `NO_EVIDENCE` 推导为系统健康或根因排除。
|
||||
|
||||
阶段 4 不实现物理或语义 Guard,因此内部返回值叫 Draft,不能直接公开发布。
|
||||
|
||||
### Audit remains injectable and explicit
|
||||
|
||||
Factory 接收可选框架 `Hook` 列表。阶段 6A 可以注入现有 `AgentLoggingHook`;内部用例始终在 `RunnableConfig` metadata 写入 sessionId/runId,所以该 Hook 不需要读取 ThreadLocal fallback。模型/Tool 预算和生命周期仍绑定 RunContext,Tool canonical audit 仍由 `ToolBoundary`/store 负责。
|
||||
|
||||
## Architecture And Interface Impact
|
||||
|
||||
```text
|
||||
DiagnosisAgentInput(query, previousTurn) + RunContext
|
||||
-> DiagnosisAgentUseCase: validate/serialize/reserve context bytes
|
||||
-> DiagnosisAgentFactory: one ReactAgent + one Prompt
|
||||
-> HarnessModelInterceptor -> DiagnosisHarnessCore -> ChatModel
|
||||
-> framework ReAct loop
|
||||
-> HarnessToolInterceptor -> HarnessEvidenceTools -> 3B/3C adapters
|
||||
-> ToolBoundary -> canonical invocation store
|
||||
-> bounded AssistantMessage text
|
||||
-> strict ObjectMapper -> DiagnosisDraft
|
||||
```
|
||||
|
||||
- 数据所有权:use case 拥有本次输入和 Draft 解析;RunContext 拥有预算/取消/生命周期;ToolBoundary/store 拥有调用记录;Agent 只拥有诊断语义。
|
||||
- 接口影响:L2 内部接口。新增阶段 6A 可消费的 Java 类型,不改变公开 HTTP/SSE、JPA、数据库或旧 Service 方法。
|
||||
- 生命周期:本 use case 不把成功 Draft 标记为 Run SUCCESS,因为阶段 5 Guards 尚未执行;预算、取消和超时可由 Core 先行终止 Run,最终持久化映射留给阶段 6A。
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Provider 返回 fenced JSON 或附加说明导致解析失败 → Mitigation:阶段 4 fail closed 且不重跑;测试固定这一行为,阶段 5 才允许一次无 Tool 结构修复。
|
||||
- [Risk] Token Usage 缺失或为 null → Mitigation:模型调用次数仍被强制;仅在可用且非负时记录实际 Token,并在 acceptance 标记 provider 元数据依赖。
|
||||
- [Risk] Tool interceptor 错误泄露内部异常 → Mitigation:只返回 allowlisted error code,不返回 message/stack/raw response。
|
||||
- [Risk] 输入和 Tool 投影共同消耗 Run bytes,配置过小会过早耗尽 → Mitigation:所有限制可配置,usage 可观察,不在本阶段硬编码生产校准值。
|
||||
- [Risk] 复用旧 AgentLoggingHook 仍包含 ThreadLocal fallback → Mitigation:新链路总是提供 metadata;阶段 7 物理清理前不修改旧 Hook,focused test 验证相同 sessionId/runId 传播。
|
||||
- [Risk] ReactAgent 内部基于框架 StateGraph → Mitigation:这是框架原生实现细节;本项目不再包一层业务 Graph 或多 Agent workflow。
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. 新增内部 Agent 类型、Prompt、interceptors 和 Tool registry,不注册 Controller bean。
|
||||
2. 使用 scripted ChatModel 和 fake/adapted Tool 完成内部 tool-loop、output、budget、no-evidence 和审计测试。
|
||||
3. 保持公开 Chat 与旧多 Agent 代码 diff 为空,Archive 并提交阶段 4。
|
||||
4. 阶段 5 在内部 Draft 后增加 Guards;阶段 6A 再创建运行应用用例并装配真实审计/previous turn;阶段 6B 原子切换公开入口。
|
||||
|
||||
Rollback:删除新增内部包、Prompt 和测试即可;不存在公开协议、数据迁移或运行时切换。
|
||||
|
||||
## Open Questions
|
||||
|
||||
无。生产预算默认值、公开切换和最终发布策略分别由配置校准、阶段 6B 和阶段 5 处理。
|
||||
@@ -0,0 +1,31 @@
|
||||
## Why
|
||||
|
||||
阶段 2-3C 已建立显式 RunContext、预算、canonical Tool invocation 和三类有界 evidence Tool,但还没有使用这些边界的诊断执行者。需要在不影响公开 Chat 的前提下建立单一 Diagnosis ReAct Agent,证明框架原生 tool loop、冻结的 DiagnosisDraft 和 Harness 控制能够闭环,再进入阶段 5 的释放门禁。
|
||||
|
||||
## What Changes
|
||||
|
||||
- 新增唯一的 Diagnosis Agent Prompt,把诊断规划、证据查询、证据充分性判断和报告草稿生成合并到一个 ReactAgent。
|
||||
- 新增独立内部 Diagnosis 用例,显式接收 RunContext、当前原始 Query 和可选的固定 PreviousTurn,不读取完整 Session 历史。
|
||||
- 通过 Spring AI Alibaba 的 ModelInterceptor 和 ToolInterceptor 接入 Harness;使用框架原生 ReAct/tool loop 和原始 tool_call_id,不手写循环或第二套调用 ID。
|
||||
- 将 lookup_knowledge、query_logs 和 query_mysql 的冻结 Tool 定义连接到阶段 3B/3C adapter,只向 Agent 返回有界 agent_result。
|
||||
- 使用 DiagnosisDraft 作为结构化输出契约,并对输入上下文、模型轮次、Token、Tool 调用和输出字节执行确定性预算。
|
||||
- 允许注入现有 AgentStep Hook,Run 审计继续由 RunContext/lifecycle 承载,ToolInvocation 审计继续由 canonical Tool boundary 承载。
|
||||
- 证据不足时要求 Agent 输出 conclusion=null、明确 limitations 并结束当前 ReAct 执行,不自动重跑 Diagnosis Agent、模型调用或 Tool 调用。
|
||||
- 不修改 ChatController、公开 /api/chat 或 /api/chat_stream,不删除旧 Planner/Executor/Verifier/Composer 链路。
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `single-react-diagnosis-agent`: 定义单一 Diagnosis ReactAgent 的内部输入、Harness-controlled ReAct/tool loop、结构化 Draft、预算、无证据停止和审计边界。
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- None. 既有 Harness、Tool 和公开 Chat capability 的需求语义不在本阶段改变。
|
||||
|
||||
## Impact
|
||||
|
||||
- 新增 `com.superbiz.agent.harness.agent` 内部包、一个 classpath Prompt 和 focused Agent/tool-loop tests。
|
||||
- 复用 `DiagnosisHarnessCore`、`RunContext`、`DiagnosisDraft`、`PreviousTurn`、`AgentToolContracts`、三类 Tool adapter 和 Spring AI Alibaba ReactAgent/interceptor API。
|
||||
- 内部接口影响为 L2:新增可供阶段 6A 调用的内部用例与工厂;现有公开 Controller、SSE、DTO、数据库契约和旧 ChatService 行为保持不变。
|
||||
- 主要风险是框架结构化输出只提供格式提示而不负责 Java 反序列化,以及 Tool Callback 本身拿不到框架调用 ID;实现必须分别用严格 ObjectMapper 解析和 ToolInterceptor 解决,失败时 fail closed。
|
||||
+78
@@ -0,0 +1,78 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Diagnosis execution SHALL use one framework ReactAgent
|
||||
The internal diagnosis path SHALL create exactly one `diagnosis_agent` with the framework-provided ReAct/tool loop. It SHALL NOT create Planner, Executor, Verifier, Composer, SequentialAgent, SupervisorAgent, an outer business Graph, or a hand-written model/Tool loop. One use-case invocation SHALL call the Diagnosis Agent once and SHALL NOT wrap the Agent or any individual model call in Harness retry.
|
||||
|
||||
#### Scenario: Diagnosis needs Tool evidence
|
||||
- **WHEN** the model emits one or more Tool actions before its final response
|
||||
- **THEN** the same Diagnosis ReactAgent executes those actions through its framework loop and returns one final Draft without invoking another Agent
|
||||
|
||||
#### Scenario: Agent execution fails
|
||||
- **WHEN** the single Diagnosis Agent invocation throws or returns invalid structured output
|
||||
- **THEN** the internal use case fails closed without automatically invoking the Agent or model again
|
||||
|
||||
### Requirement: Diagnosis input SHALL contain only current Query and optional PreviousTurn
|
||||
The internal use case SHALL accept a non-blank current Query and an optional frozen `PreviousTurn`, serialize them as the fixed `query` and `previous_turn` input fields, and SHALL NOT load or accept complete Session history, Redis memory, prior raw Tool results, or model-generated history summaries. The current Query SHALL retain its original text and SHALL NOT be rewritten or silently truncated.
|
||||
|
||||
#### Scenario: Follow-up diagnosis has a safe previous turn
|
||||
- **WHEN** a caller supplies a `PreviousTurn`
|
||||
- **THEN** the Agent receives exactly the frozen previous-turn fields plus the current Query and no complete conversation history
|
||||
|
||||
#### Scenario: Query exceeds configured context budget
|
||||
- **WHEN** the current Query exceeds its UTF-8 byte limit
|
||||
- **THEN** the use case rejects it before any model call instead of truncating or rewriting it
|
||||
|
||||
### Requirement: Every model round SHALL be controlled by RunContext
|
||||
The Diagnosis Agent SHALL receive `RunContext` explicitly and SHALL use a model interceptor to call `DiagnosisHarnessCore.beforeModelCall` for every framework ReAct model round. Non-streaming model response Usage SHALL be recorded into the same Run budget when available. Cancellation, deadline, model-call exhaustion, or Token exhaustion SHALL prevent subsequent controlled work.
|
||||
|
||||
#### Scenario: ReAct performs two model rounds
|
||||
- **WHEN** one model round requests a Tool and the next produces the Draft
|
||||
- **THEN** the same Run budget records two model calls and no Harness retry attempt
|
||||
|
||||
#### Scenario: Model-call budget is exhausted
|
||||
- **WHEN** the framework attempts a model round beyond the configured maximum
|
||||
- **THEN** the call is rejected before reaching ChatModel and the Run records budget exhaustion
|
||||
|
||||
### Requirement: Evidence Tools SHALL execute through the Harness boundary
|
||||
The Diagnosis Agent SHALL expose only the frozen `lookup_knowledge`, `query_logs`, and `query_mysql` definitions. A Tool interceptor SHALL propagate the exact framework Tool Call ID, Run ID, Tool name, and raw JSON arguments into the corresponding stage 3B/3C adapter and `ToolBoundary`. Successful observations SHALL contain only the bounded `agent_result`; failed observations SHALL contain only stable error semantics and SHALL NOT contain raw responses, internal exceptions, credentials, or invocation lifecycle internals.
|
||||
|
||||
#### Scenario: Framework requests RAG evidence
|
||||
- **WHEN** the model calls `lookup_knowledge` with framework ID `call-1`
|
||||
- **THEN** the RAG adapter and Agent observation use exactly `call-1`, and the canonical invocation is owned by the current Run
|
||||
|
||||
#### Scenario: Tool execution fails
|
||||
- **WHEN** a registered adapter returns an error result
|
||||
- **THEN** the Agent receives `evidence_status=ERROR`, the framework Tool Call ID and a stable error code without automatic Tool retry or raw failure detail
|
||||
|
||||
#### Scenario: Unknown Tool is requested
|
||||
- **WHEN** a model requests a Tool outside the three registered definitions
|
||||
- **THEN** the Harness does not authorize or emulate it and does not create a canonical evidence record
|
||||
|
||||
### Requirement: Diagnosis output SHALL be a bounded DiagnosisDraft
|
||||
The Agent SHALL receive the generated schema for `DiagnosisDraft` and SHALL return JSON that the internal use case strictly parses into the frozen record. The use case SHALL enforce configured UTF-8 limits for query, previous turn, total input and Draft output and account accepted input/output bytes against the Run capacity. It SHALL reject blank, fenced, prefixed, malformed, oversized or schema-incompatible output without repair or retry.
|
||||
|
||||
#### Scenario: Supported Draft is returned
|
||||
- **WHEN** the Agent returns valid JSON containing Conclusion, Analysis items, Action Plan, Recommendations and Limitations
|
||||
- **THEN** the use case returns a `DiagnosisDraft` whose Analysis Tool Call IDs are the exact strings emitted by the Agent
|
||||
|
||||
#### Scenario: Model returns prose around JSON
|
||||
- **WHEN** the final response contains a Markdown fence or explanatory prefix around an otherwise valid object
|
||||
- **THEN** strict parsing fails and the Diagnosis Agent is not invoked a second time
|
||||
|
||||
### Requirement: Insufficient evidence SHALL terminate without a fabricated conclusion
|
||||
The single Prompt SHALL require every normal Analysis item to cite current-Run evidence Tool Call IDs and SHALL restrict `NO_EVIDENCE` to scoped `NEGATIVE_OBSERVATION`. If current evidence cannot support a diagnosis, the Agent SHALL stop the current ReAct execution with `conclusion=null`, describe the actual scope and missing information in `limitations`, and SHALL NOT infer that the problem does not exist or fabricate a root cause.
|
||||
|
||||
#### Scenario: Tool finds no evidence
|
||||
- **WHEN** the only completed Tool observation has `evidence_status=NO_EVIDENCE`
|
||||
- **THEN** the final Draft has no confirmed Conclusion, records the bounded negative observation and missing information, and makes no additional automatic retry
|
||||
|
||||
### Requirement: Stage-four execution SHALL remain internal and auditable
|
||||
The new use case SHALL be callable only as an internal Java/test entry in this stage and SHALL NOT be wired into public Chat or AIOps controllers. It SHALL propagate the same sessionId/runId through RunnableConfig metadata, allow the existing AgentStep Hook to be injected, retain ToolBoundary canonical invocation audit, and leave final Run success/persistence ownership to later Guard/application stages.
|
||||
|
||||
#### Scenario: Internal audited execution completes
|
||||
- **WHEN** an injected audit Hook observes a model round and a Tool executes
|
||||
- **THEN** both observe the same RunContext sessionId/runId and the public Chat call path remains unchanged
|
||||
|
||||
#### Scenario: Stage four is archived
|
||||
- **WHEN** focused Agent tests pass and the change is archived
|
||||
- **THEN** `ChatController`, public SSE behavior and the old multi-Agent implementation remain available for stages 5, 6A and 6B
|
||||
@@ -0,0 +1,22 @@
|
||||
## 1. Input And Prompt Contract
|
||||
|
||||
- [x] 1.1 Add immutable DiagnosisAgentInput and configurable UTF-8 context/output limits with validation and focused boundary tests.
|
||||
- [x] 1.2 Add the single Diagnosis Agent classpath Prompt covering Tool selection, current-Run evidence binding, no-evidence stopping and exact DiagnosisDraft JSON output.
|
||||
|
||||
## 2. Harness Tool Integration
|
||||
|
||||
- [x] 2.1 Add HarnessEvidenceTools with the three frozen Tool definitions and adapter bridges for RAG, logs and MySQL.
|
||||
- [x] 2.2 Add a ToolInterceptor that propagates the exact framework Tool Call ID and RunContext, returns only bounded agent_result or stable error JSON, and never retries or leaks raw results.
|
||||
- [x] 2.3 Add focused tests for framework ID fidelity, successful Tool observations, error observations, unknown Tool rejection and single Tool budget accounting.
|
||||
|
||||
## 3. Single Diagnosis Agent Execution
|
||||
|
||||
- [x] 3.1 Add a ModelInterceptor that enforces every model boundary and records non-streaming Token Usage into the same Run budget.
|
||||
- [x] 3.2 Add DiagnosisAgentFactory that creates one non-parallel ReactAgent with one Prompt, DiagnosisDraft output schema, Harness interceptors and injectable audit Hooks.
|
||||
- [x] 3.3 Add the internal DiagnosisAgentUseCase that validates/serializes fixed input, enforces context/Draft byte limits, calls the Agent once and strictly parses DiagnosisDraft without repair or retry.
|
||||
|
||||
## 4. Focused Behavior Verification
|
||||
|
||||
- [x] 4.1 Add scripted ChatModel tests proving a single Agent performs framework model -> Tool -> model execution, preserves Query/PreviousTurn, returns typed Analysis/Conclusion/Tool IDs and propagates sessionId/runId to audit.
|
||||
- [x] 4.2 Add no-evidence, invalid/fenced output, model-call budget, Token budget, context budget and no-automatic-retry tests.
|
||||
- [x] 4.3 Verify no public Chat/AiOps runtime file changed, run focused Harness/Agent regressions, compile, and validate the OpenSpec strictly.
|
||||
Reference in New Issue
Block a user