feat(harness): add canonical tool invocation boundary

This commit is contained in:
zhuyongxin
2026-07-21 19:36:25 +08:00
parent 6b74990f86
commit 0dbdd7d8d3
31 changed files with 1543 additions and 1 deletions
@@ -0,0 +1 @@
Archive-ready after implementation and focused verification on 2026-07-21.
@@ -0,0 +1 @@
Committed after strict validation on 2026-07-21.
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-21
@@ -0,0 +1,102 @@
## Context
当前 `ToolInvocationRecorder` 将 JPA `ToolInvocation` 作为旧链路的 durable audit,保存 `output_preview`(500 字符)和 retrieval details,并通过 `SessionContextHolder` 补 session/run。它不能作为 EvidenceGuard 的 canonical source:raw response 已被截断、生命周期没有 PROJECTING/READY/ERROR、没有框架 Tool Call ID,也没有单 Run 容量/TTL 门禁。
阶段 2 已提供 `RunContext`、budget、capacity 和 `ToolCallKeyFactory`。本阶段需把所有 Tool 共用的执行边界与 canonical invocation store 建起来,供阶段 3B/3C 的 projector 直接复用。
## Goals / Non-Goals
**Goals:**
- 在 Tool 调用前统一校验 Run 所有权、框架 ID、JSON object、授权、只读和 Tool/Run budget。
- 保存同一 canonical record 的完整 request/raw_response/agent_result 及状态、证据语义、时间和错误。
- 集中执行 PROJECTING -> READY/ERROR,防止未经投影的 raw 进入 Agent。
- 固定 TTL 不续期、单记录/Agent result/Run bytes 上限和 RESULT_TOO_LARGE 语义。
- 提供 Redis 实现和 store-independent Fake boundary tests。
**Non-Goals:**
- 不实现 RAG/log/MySQL specific projector 或 adapter。
- 不修改旧 ToolInvocationRecorder/JPA/数据库、Chat/AIOps、Controller/SSE。
- 不让 Agent 访问 Redis/client/key/raw record。
- 不实现并发 Lua/CAS 更新、永久 audit、脱敏或跨 Run 查询。
## Decisions
### 1. Internal ToolBoundary envelope and result
`ToolCallRequestEnvelope` 是 Harness 内部输入,包含 `runId`、框架 `toolCallId`、toolName、JSON request、`authorized` 和 `readOnly`。Agent-facing DTO 不暴露这些门禁字段;后续 Alibaba ToolInterceptor 负责组装 envelope。
`ToolBoundaryResult` 只返回 invocation status、evidence status、原始框架 ID、bounded `agentResult` 或安全 errorCode;raw 只进入 store,不返回 Agent。
### 2. Canonical record and lifecycle
`CanonicalToolInvocation` 是 JSON serializable record,字段包含 toolCallId、runId、toolName、request、rawResponse、agentResult、InvocationStatus、EvidenceStatus、errorCode、startedAt、completedAt。`begin` 只接受 PROJECTING;`markReady` 只接受 PROJECTING + 非空 bounded projection + FOUND/NO_EVIDENCE;`markError` 将 evidence status 固定为 ERROR。
状态转换和 duplicate 检查在 `CanonicalInvocationStore` 内集中执行。ToolBoundary 不直接写 Redis。
### 3. Redis value and TTL
Redis 实现复用现有 `RedisTemplate<String,Object>`,将 record 序列化为 JSON String。创建使用 `setIfAbsent(key,json,ttl)`,保证同一 `runId+toolCallId` 不覆盖;读取不调用 expire。更新先读取剩余毫秒 TTL,再用不大于该值的 TTL 写回,避免恢复初始 TTL。过期/缺失读取返回 empty。
替代方案是 Redis Hash;单 JSON value 能保证 request/raw/agent 原子同记录,并让 Fake/序列化 schema 与 canonical record 一致,故采用。并发 projector 更新窗口是已接受风险,后续需要时再升级 Lua/CAS。
### 4. Size and status rules
`CanonicalInvocationLimits` 由 caller 提供 TTL、maxRecordBytes 和 maxAgentResultBytes。request/raw/agent 使用 UTF-8 bytes 计数;raw 超过 record 或 agent projection 超过独立上限,均不截断,记录 ERROR/RESULT_TOO_LARGE。Run bytes 通过阶段 2 Core 再做单 Run 累计门禁。
PROJECTING 时 evidence status 仅作为内部未知/ERROR 占位;READY 只接受 `EVIDENCE_FOUND` 或 `NO_EVIDENCE`;ERROR 永远不可引用。NO_EVIDENCE 不触发重试或成功解释。
### 5. Preflight and projector boundary
ToolBoundary 顺序固定:
```text
RunContext active/deadline
-> runId + toolCallId + authorization + read-only + JSON object
-> Core.beforeToolCall + request/run capacity
-> store.begin(PROJECTING)
-> ToolExecutor(raw)
-> raw size/capacity
-> ToolResultProjector(agentResult,evidenceStatus)
-> agent size/capacity
-> store.markReady or markError
-> bounded ToolBoundaryResult
```
执行或投影异常都会写 ERROR;raw 已在可信边界且未超限时保留在 canonical record,但不返回 Agent。preflight/duplicate/cross-run 错误在 begin 前返回安全 ERROR。
### 6. Old audit separation
旧 recorder/JPA 继续接收旧 Tool 调用,阶段 3A 不改其字段和 ThreadLocal fallback。新 canonical store 没有旧消费者;阶段 3B/3C 接入时必须明确写新 boundary,并在需要 durable audit 时另行脱敏摘要。
## Module Flow
```text
Alibaba ToolInterceptor (future)
-> ToolCallRequestEnvelope
-> ToolBoundary
-> DiagnosisHarnessCore + ToolCallKeyFactory
-> CanonicalInvocationStore (Redis JSON / Fake)
-> ToolExecutor
-> ToolResultProjector (future RAG/log/MySQL)
-> bounded ToolBoundaryResult
```
## Risks / Trade-offs
- [Redis update read-TTL-write 存在并发窗口] -> 当前每个 invocation 只允许 boundary 顺序更新;后续并发需求升级 Lua/CAS。
- [canonical raw 可能敏感] -> 仅 Harness store 访问,TTL/ACL/容量受限;脱敏在 projector/durable audit 阶段处理。
- [旧 recorder 与新 store 短期并存] -> 包和接口隔离,spec 明确旧 preview 不能作为 canonical evidence。
- [preflight 失败可能没有 canonical record] -> 返回安全 ERROR 且不执行 Tool;阶段 3A 的可引用记录只针对已通过 begin 的调用。
## Migration Plan
1. 本阶段新增 boundary/store/Redis adapter 和 fake tests,旧运行链路不变。
2. 阶段 3B/3C 将各 Tool adapter/projector 包装到本 boundary。
3. 阶段 4 Diagnosis Agent 只接收 boundary 的 bounded result。
4. 阶段 6A/7 再决定 durable audit 如何从 canonical 摘要回填,并清理旧 recorder/ThreadLocal。
## Open Questions
无。真实 Redis 的 ACL、网络和 TTL 由最终运行/E2E 阶段验证;并发更新 Lua 化留作后续需求。
@@ -0,0 +1,45 @@
## Why
阶段 1 已冻结 Agent-facing Tool Contract,阶段 2 已提供 RunContext、预算、取消和 Tool Call Key 基础,但当前 `ToolInvocationRecorder` 仍把截断 preview 写入 JPA、依赖 ThreadLocal,并没有同一条记录中的 `request/raw_response/agent_result`、生命周期或当前 Run 所有权。RAG、日志和 MySQL 投影若各自保存调用,会重新复制状态机并让 EvidenceGuard 无法证明引用来自当前 Run。
## What Changes
- 新增统一 `ToolBoundary`,在每个 Tool 调用前执行 JSON Schema/只读/Run/预算/Tool Call ID 门禁,执行后统一处理 raw、投影、状态和错误。
- 新增 `CanonicalInvocationStore` 抽象与 Redis 实现,按阶段 2 Key Factory 保存一条完整 JSON 调用记录:`request`、`raw_response`、`agent_result`、`status`、`evidence_status`、时间和错误信息。
- `PROJECTING -> READY/ERROR` 生命周期和独立 EvidenceStatus 在 store 中集中执行;READY 才允许 `EVIDENCE_FOUND/NO_EVIDENCE`,ERROR 不可引用。
- 创建时设置 TTL,读取不刷新;更新只使用当前剩余 TTL,不延长生命周期;单记录、Agent projection 和单 Run 容量超限显式返回 `RESULT_TOO_LARGE`,不静默截断 raw。
- 拒绝缺失/非法/重复 Tool Call ID、跨 Run 引用、不可解析 JSON、非只读请求和已超预算调用;不生成第二套 ID。
- 使用 Fake Tool/Projector/In-memory Store 覆盖成功、no-evidence、projection error、execution error、duplicate/cross-run、TTL、容量和 raw oversize。
- 本阶段不实现 RAG/log/MySQL specific projector,不修改旧 `ToolInvocationRecorder`、JPA entity、Controller、ChatService 或公开协议。
## Capabilities
### New Capabilities
- `canonical-tool-invocation-store`: 提供统一 ToolBoundary、canonical invocation 生命周期、Run 所有权、容量/TTL 和可引用状态边界,供后续 RAG/log/MySQL 投影复用。
### Modified Capabilities
- None. 旧 JPA audit 记录继续服务旧链路;新 store 先作为零消费者 Harness foundation。
## Context Constraints
- canonical store 只能由 Harness/ToolBoundary 访问,Agent 不获得 Redis client/key/raw record。
- Redis key 固定由阶段 2 `ToolCallKeyFactory` 生成:`prefix:runId:toolCallId`。
- 同一调用的完整 request/raw/agent projection 必须在同一记录;raw 不能只保存 preview,也不能未经 projector 返回 Agent。
- 创建 TTL 默认配置由 caller 提供且必须大于 0;读取与更新不得续期。
- `PROJECTING` 时 evidence_status 只能是内部暂态 ERROR/unknown;只有 READY 才能成为 `EVIDENCE_FOUND` 或 `NO_EVIDENCE`。
- `NO_EVIDENCE` 仅作为结果语义,不可被 boundary 自动升级为成功事实或重试。
## Interface Impact
- 等级:L2(内部 Harness/Tool boundary)。新增接口会被阶段 3B/3C 直接消费,旧调用方不变。
- 不改变 JPA `tool_invocation`、数据库 Schema、旧审计 preview 或公开 HTTP/SSE。
- Redis 是新增运行时依赖使用既有 `RedisTemplate<String,Object>` bean;真实连接验证留给阶段 7,focused tests 使用 fake/mocks。
## Risks
- Redis JSON value 更新需要读取剩余 TTL 后再写回,存在并发更新窗口;当前单 Tool Call 只有 boundary 状态机写入,后续若并发 projector 必须升级 Lua/CAS。
- canonical raw 可包含敏感内容;本 Issue 保留阶段 0 已确认的 Harness-only ACL/TTL 约束,持久化脱敏和 durable audit 留给后续阶段。
- ToolBoundary 同时负责预算、store 状态和 projector 错误,若异常分类不清会产生错误状态;每个边界分支都有 Fake tests。
- 当前旧 recorder 继续运行,新旧两条 audit 链短期并存;proposal 明确禁止把旧 JPA 记录当 canonical evidence。
@@ -0,0 +1,74 @@
## ADDED Requirements
### Requirement: ToolBoundary SHALL enforce explicit preflight before execution
The Harness SHALL reject a Tool call before invoking the executor when the envelope has a blank/unsafe framework `tool_call_id`, a run ID different from RunContext, invalid JSON object input, unauthorized access, non-read-only access, an inactive/deadline-expired Run, duplicate canonical key, or exhausted Tool/Run budget.
#### Scenario: Cross-Run Tool Call is rejected
- **WHEN** an envelope run ID differs from the explicit RunContext run ID
- **THEN** ToolBoundary returns a safe `ERROR`, does not create a canonical record, and does not invoke the Tool
#### Scenario: Unauthorized or writable Tool is rejected
- **WHEN** `authorized=false` or `readOnly=false`
- **THEN** ToolBoundary returns `ERROR` before execution and does not expose the request to an Agent
### Requirement: Canonical invocation SHALL keep complete data in one record
The store SHALL save the complete request JSON, raw Tool response, bounded Agent result, framework `tool_call_id`, run ID, tool name, lifecycle status, evidence status, timestamps, and error code in one canonical record. The raw response SHALL NOT be replaced by a preview, and SHALL NOT be returned to the Agent.
#### Scenario: Tool begins execution
- **WHEN** preflight succeeds and the Tool is about to execute
- **THEN** one record is created with `status=PROJECTING`, the complete request, the exact framework ID, and no Agent result yet
#### Scenario: Projector succeeds
- **WHEN** the executor returns raw data and the projector returns a bounded result
- **THEN** the same record contains raw data and Agent result with `status=READY`, and ToolBoundary returns only the bounded Agent result
### Requirement: Invocation lifecycle and evidence semantics SHALL be enforced centrally
The store SHALL allow only `PROJECTING -> READY` or `PROJECTING -> ERROR`. READY SHALL require `EVIDENCE_FOUND` or `NO_EVIDENCE`; ERROR SHALL use `evidence_status=ERROR` and SHALL NOT be referencable. `NO_EVIDENCE` SHALL remain a scoped negative observation and SHALL NOT be upgraded to success or retried by the boundary.
#### Scenario: No evidence projection completes
- **WHEN** a projector returns a valid bounded result with `NO_EVIDENCE`
- **THEN** the record becomes `READY`, preserves the result scope, and remains eligible only for a negative observation
#### Scenario: Projection fails
- **WHEN** the projector throws or returns an invalid evidence status
- **THEN** the same record becomes `ERROR`, stores a safe error code, and no Agent result is returned
### Requirement: Tool Call ID and Run ownership SHALL be preserved
The boundary SHALL use the exact framework `tool_call_id` with the RunContext run ID to create the canonical key. It SHALL reject missing, unsafe, duplicate, and cross-Run references and SHALL never generate or replace a fallback ID.
#### Scenario: Duplicate Tool Call ID is submitted
- **WHEN** a second invocation uses the same valid run ID and framework Tool Call ID
- **THEN** the second Tool is not executed and returns `ERROR` without overwriting the first record
#### Scenario: Framework ID is preserved
- **WHEN** a valid envelope passes preflight
- **THEN** the key and canonical record contain the exact supplied `tool_call_id`
### Requirement: TTL and size limits SHALL fail closed
The Redis implementation SHALL set TTL only when a canonical record is created. Reads SHALL NOT refresh TTL; updates SHALL preserve only the current remaining TTL. Request/raw/record and Agent-result byte limits SHALL be enforced in UTF-8. Raw overflow SHALL produce `ERROR/RESULT_TOO_LARGE` without silent truncation; Agent projection overflow SHALL produce the same error.
#### Scenario: Read does not renew TTL
- **WHEN** a canonical record is read before expiration
- **THEN** its expiry remains at or before the original expiry and no expire/refresh operation is issued
#### Scenario: Raw response is too large
- **WHEN** the executor returns raw data beyond the record limit
- **THEN** the record becomes `ERROR` with `RESULT_TOO_LARGE`, the raw payload is not silently truncated, and the projector is not invoked
#### Scenario: Agent result is too large
- **WHEN** a projector returns a result beyond the Agent projection limit
- **THEN** the record becomes `ERROR/RESULT_TOO_LARGE` and the oversized result is not returned to the Agent
### Requirement: Tool and store failures SHALL return safe errors
Execution, serialization, Redis, duplicate, preflight, and projection failures SHALL return a bounded error code without raw payload, credentials, internal stack trace, or Redis key to the Agent. If raw data was safely stored before a projection failure, it remains Harness-only canonical data.
#### Scenario: Tool execution throws
- **WHEN** the executor raises an exception after PROJECTING begins
- **THEN** the record becomes `ERROR` with a stable execution error code and ToolBoundary returns no raw response
### Requirement: Stage 3A SHALL remain reusable and independent from legacy audit
The boundary and canonical store SHALL be usable by later RAG/log/MySQL projectors without modifying legacy `ToolInvocationRecorder`, JPA `ToolInvocation`, Chat/AIOps, Controller, or public protocol. Redis access SHALL remain inside the store implementation and no Agent-facing API SHALL expose the Redis client.
#### Scenario: Fake projector tests pass
- **WHEN** Fake Tool/Projector tests cover success, no evidence, error, duplicate, TTL, and capacity cases
- **THEN** later Tool-specific changes can depend on the same boundary without copying lifecycle or Redis logic, while legacy call sites remain unchanged
@@ -0,0 +1,24 @@
## 1. Canonical Model and Store
- [x] 1.1 Implement CanonicalToolInvocation, CanonicalInvocationLimits, store exceptions, and lifecycle/evidence transition validation.
- [x] 1.2 Implement CanonicalInvocationStore and Redis JSON-value adapter with set-if-absent creation, complete record persistence, and remaining-TTL updates.
- [x] 1.3 Add store tests for duplicate creation, PROJECTING/READY/ERROR transitions, evidence status rules, and read-without-TTL-refresh semantics.
## 2. Tool Boundary
- [x] 2.1 Implement ToolCallRequestEnvelope, ToolBoundaryResult, ProjectedToolResult, executor/projector interfaces, and safe error codes.
- [x] 2.2 Implement ToolBoundary preflight for Run ownership, framework ID, JSON object, authorization, read-only, budget, and duplicate gates.
- [x] 2.3 Implement execution -> raw size/capacity -> projection -> Agent size/capacity -> READY/ERROR flow without returning raw data.
- [x] 2.4 Add Fake Tool/Projector boundary tests for success, NO_EVIDENCE, execution/projection errors, duplicate/cross-run/unauthorized/writable requests.
## 3. Limits and Failure Semantics
- [x] 3.1 Enforce UTF-8 request/raw/record/Agent-result limits and `RESULT_TOO_LARGE` without silent raw truncation.
- [x] 3.2 Add tests proving raw overflow skips projector, Agent overflow is not returned, and Run capacity remains consistent.
- [x] 3.3 Verify ERROR records cannot become referencable READY evidence and NO_EVIDENCE remains scoped negative observation.
## 4. Verification and Isolation
- [x] 4.1 Run focused canonical store and ToolBoundary tests with fake store/Redis operations.
- [x] 4.2 Run stage 0/1/2 contract, Core, retry, key, and ChatController regression tests.
- [x] 4.3 Verify legacy ToolInvocationRecorder/JPA, Chat/AIOps, Controller, and public protocol files are unchanged; Redis access is confined to the new store adapter.
@@ -0,0 +1,78 @@
# canonical-tool-invocation-store Specification
## Purpose
定义 Harness ToolBoundary 与 canonical invocation store 的统一执行边界,包括 Run/Tool preflight、PROJECTING/READY/ERROR 生命周期、evidence status、框架 Tool Call ID、TTL、容量和有界 Agent 投影。
## Requirements
### Requirement: ToolBoundary SHALL enforce explicit preflight before execution
The Harness SHALL reject a Tool call before invoking the executor when the envelope has a blank/unsafe framework `tool_call_id`, a run ID different from RunContext, invalid JSON object input, unauthorized access, non-read-only access, an inactive/deadline-expired Run, duplicate canonical key, or exhausted Tool/Run budget.
#### Scenario: Cross-Run Tool Call is rejected
- **WHEN** an envelope run ID differs from the explicit RunContext run ID
- **THEN** ToolBoundary returns a safe `ERROR`, does not create a canonical record, and does not invoke the Tool
#### Scenario: Unauthorized or writable Tool is rejected
- **WHEN** `authorized=false` or `readOnly=false`
- **THEN** ToolBoundary returns `ERROR` before execution and does not expose the request to an Agent
### Requirement: Canonical invocation SHALL keep complete data in one record
The store SHALL save the complete request JSON, raw Tool response, bounded Agent result, framework `tool_call_id`, run ID, tool name, lifecycle status, evidence status, timestamps, and error code in one canonical record. The raw response SHALL NOT be replaced by a preview, and SHALL NOT be returned to the Agent.
#### Scenario: Tool begins execution
- **WHEN** preflight succeeds and the Tool is about to execute
- **THEN** one record is created with `status=PROJECTING`, the complete request, the exact framework ID, and no Agent result yet
#### Scenario: Projector succeeds
- **WHEN** the executor returns raw data and the projector returns a bounded result
- **THEN** the same record contains raw data and Agent result with `status=READY`, and ToolBoundary returns only the bounded Agent result
### Requirement: Invocation lifecycle and evidence semantics SHALL be enforced centrally
The store SHALL allow only `PROJECTING -> READY` or `PROJECTING -> ERROR`. READY SHALL require `EVIDENCE_FOUND` or `NO_EVIDENCE`; ERROR SHALL use `evidence_status=ERROR` and SHALL NOT be referencable. `NO_EVIDENCE` SHALL remain a scoped negative observation and SHALL NOT be upgraded to success or retried by the boundary.
#### Scenario: No evidence projection completes
- **WHEN** a projector returns a valid bounded result with `NO_EVIDENCE`
- **THEN** the record becomes `READY`, preserves the result scope, and remains eligible only for a negative observation
#### Scenario: Projection fails
- **WHEN** the projector throws or returns an invalid evidence status
- **THEN** the same record becomes `ERROR`, stores a safe error code, and no Agent result is returned
### Requirement: Tool Call ID and Run ownership SHALL be preserved
The boundary SHALL use the exact framework `tool_call_id` with the RunContext run ID to create the canonical key. It SHALL reject missing, unsafe, duplicate, and cross-Run references and SHALL never generate or replace a fallback ID.
#### Scenario: Duplicate Tool Call ID is submitted
- **WHEN** a second invocation uses the same valid run ID and framework Tool Call ID
- **THEN** the second Tool is not executed and returns `ERROR` without overwriting the first record
#### Scenario: Framework ID is preserved
- **WHEN** a valid envelope passes preflight
- **THEN** the key and canonical record contain the exact supplied `tool_call_id`
### Requirement: TTL and size limits SHALL fail closed
The Redis implementation SHALL set TTL only when a canonical record is created. Reads SHALL NOT refresh TTL; updates SHALL preserve only the current remaining TTL. Request/raw/record and Agent-result byte limits SHALL be enforced in UTF-8. Raw overflow SHALL produce `ERROR/RESULT_TOO_LARGE` without silent truncation; Agent projection overflow SHALL produce the same error.
#### Scenario: Read does not renew TTL
- **WHEN** a canonical record is read before expiration
- **THEN** its expiry remains at or before the original expiry and no expire/refresh operation is issued
#### Scenario: Raw response is too large
- **WHEN** the executor returns raw data beyond the record limit
- **THEN** the record becomes `ERROR` with `RESULT_TOO_LARGE`, the raw payload is not silently truncated, and the projector is not invoked
#### Scenario: Agent result is too large
- **WHEN** a projector returns a result beyond the Agent projection limit
- **THEN** the record becomes `ERROR/RESULT_TOO_LARGE` and the oversized result is not returned to the Agent
### Requirement: Tool and store failures SHALL return safe errors
Execution, serialization, Redis, duplicate, preflight, and projection failures SHALL return a bounded error code without raw payload, credentials, internal stack trace, or Redis key to the Agent. If raw data was safely stored before a projection failure, it remains Harness-only canonical data.
#### Scenario: Tool execution throws
- **WHEN** the executor raises an exception after PROJECTING begins
- **THEN** the record becomes `ERROR` with a stable execution error code and ToolBoundary returns no raw response
### Requirement: Stage 3A SHALL remain reusable and independent from legacy audit
The boundary and canonical store SHALL be usable by later RAG/log/MySQL projectors without modifying legacy `ToolInvocationRecorder`, JPA `ToolInvocation`, Chat/AIOps, Controller, or public protocol. Redis access SHALL remain inside the store implementation and no Agent-facing API SHALL expose the Redis client.
#### Scenario: Fake projector tests pass
- **WHEN** Fake Tool/Projector tests cover success, no evidence, error, duplicate, TTL, and capacity cases
- **THEN** later Tool-specific changes can depend on the same boundary without copying lifecycle or Redis logic, while legacy call sites remain unchanged