Add repairable INVALID_PROGRESS_PROTOCOL observations, independent PROGRESS_PROTOCOL_VIOLATED saturation, and controlled release paths. Archive the OpenSpec change after syncing main specs and devflow.
83 lines
6.8 KiB
Markdown
83 lines
6.8 KiB
Markdown
# canonical-tool-invocation-store Specification
|
|
|
|
## Purpose
|
|
定义 Harness ToolBoundary 与 canonical invocation store 的统一执行边界,包括 Run/Tool preflight、PROJECTING/READY/ERROR 生命周期、evidence status、框架 Tool Call ID、TTL、容量和有界 Agent 投影。
|
|
|
|
## Requirements
|
|
### Requirement: ToolBoundary SHALL enforce explicit preflight before execution
|
|
The Harness SHALL reject a Tool call before invoking the executor when the envelope has a blank/unsafe framework `tool_call_id`, a run ID different from RunContext, invalid JSON object input, unauthorized access, non-read-only access, an inactive/deadline-expired Run, duplicate canonical key, or exhausted Tool/Run budget.
|
|
|
|
#### Scenario: Cross-Run Tool Call is rejected
|
|
- **WHEN** an envelope run ID differs from the explicit RunContext run ID
|
|
- **THEN** ToolBoundary returns a safe `ERROR`, does not create a canonical record, and does not invoke the Tool
|
|
|
|
#### Scenario: Unauthorized or writable Tool is rejected
|
|
- **WHEN** `authorized=false` or `readOnly=false`
|
|
- **THEN** ToolBoundary returns `ERROR` before execution and does not expose the request to an Agent
|
|
### Requirement: Canonical invocation SHALL keep complete data in one record
|
|
The store SHALL save the complete request JSON, raw Tool response, bounded Agent result, framework `tool_call_id`, run ID, tool name, lifecycle status, evidence status, timestamps, and error code in one canonical record. The raw response SHALL NOT be replaced by a preview, and SHALL NOT be returned to the Agent.
|
|
|
|
#### Scenario: Tool begins execution
|
|
- **WHEN** preflight succeeds and the Tool is about to execute
|
|
- **THEN** one record is created with `status=PROJECTING`, the complete request, the exact framework ID, and no Agent result yet
|
|
|
|
#### Scenario: Projector succeeds
|
|
- **WHEN** the executor returns raw data and the projector returns a bounded result
|
|
- **THEN** the same record contains raw data and Agent result with `status=READY`, and ToolBoundary returns only the bounded Agent result
|
|
### Requirement: Invocation lifecycle and evidence semantics SHALL be enforced centrally
|
|
The store SHALL allow only `PROJECTING -> READY` or `PROJECTING -> ERROR`. READY SHALL require `EVIDENCE_FOUND` or `NO_EVIDENCE`; ERROR SHALL use `evidence_status=ERROR` and SHALL NOT be referencable. `NO_EVIDENCE` SHALL remain a scoped negative observation and SHALL NOT be upgraded to success or retried by the boundary.
|
|
|
|
#### Scenario: No evidence projection completes
|
|
- **WHEN** a projector returns a valid bounded result with `NO_EVIDENCE`
|
|
- **THEN** the record becomes `READY`, preserves the result scope, and remains eligible only for a negative observation
|
|
|
|
#### Scenario: Projection fails
|
|
- **WHEN** the projector throws or returns an invalid evidence status
|
|
- **THEN** the same record becomes `ERROR`, stores a safe error code, and no Agent result is returned
|
|
### Requirement: Tool Call ID and Run ownership SHALL be preserved
|
|
The boundary SHALL use the exact framework `tool_call_id` with the RunContext run ID to create the canonical key. It SHALL reject missing, unsafe, duplicate, and cross-Run references and SHALL never generate or replace a fallback ID.
|
|
|
|
#### Scenario: Duplicate Tool Call ID is submitted
|
|
- **WHEN** a second invocation uses the same valid run ID and framework Tool Call ID
|
|
- **THEN** the second Tool is not executed and returns `ERROR` without overwriting the first record
|
|
|
|
#### Scenario: Framework ID is preserved
|
|
- **WHEN** a valid envelope passes preflight
|
|
- **THEN** the key and canonical record contain the exact supplied `tool_call_id`
|
|
### Requirement: TTL and size limits SHALL fail closed
|
|
The Redis implementation SHALL set TTL only when a canonical record is created. Reads SHALL NOT refresh TTL; updates SHALL preserve only the current remaining TTL. Request/raw/record and Agent-result byte limits SHALL be enforced in UTF-8. Raw overflow SHALL produce `ERROR/RESULT_TOO_LARGE` without silent truncation; Agent projection overflow SHALL produce the same error.
|
|
|
|
#### Scenario: Read does not renew TTL
|
|
- **WHEN** a canonical record is read before expiration
|
|
- **THEN** its expiry remains at or before the original expiry and no expire/refresh operation is issued
|
|
|
|
#### Scenario: Raw response is too large
|
|
- **WHEN** the executor returns raw data beyond the record limit
|
|
- **THEN** the record becomes `ERROR` with `RESULT_TOO_LARGE`, the raw payload is not silently truncated, and the projector is not invoked
|
|
|
|
#### Scenario: Agent result is too large
|
|
- **WHEN** a projector returns a result beyond the Agent projection limit
|
|
- **THEN** the record becomes `ERROR/RESULT_TOO_LARGE` and the oversized result is not returned to the Agent
|
|
### Requirement: Tool and store failures SHALL return safe errors
|
|
Execution, serialization, Redis, duplicate, preflight, and projection failures SHALL return a bounded error code without raw payload, credentials, internal stack trace, or Redis key to the Agent. If raw data was safely stored before a projection failure, it remains Harness-only canonical data.
|
|
|
|
#### Scenario: Tool execution throws
|
|
- **WHEN** the executor raises an exception after PROJECTING begins
|
|
- **THEN** the record becomes `ERROR` with a stable execution error code and ToolBoundary returns no raw response
|
|
### Requirement: Stage 3A SHALL remain reusable and independent from legacy audit
|
|
The boundary and canonical store SHALL be usable by later RAG/log/MySQL projectors without modifying legacy `ToolInvocationRecorder`, JPA `ToolInvocation`, Chat/AIOps, Controller, or public protocol. Redis access SHALL remain inside the store implementation and no Agent-facing API SHALL expose the Redis client.
|
|
|
|
#### Scenario: Fake projector tests pass
|
|
- **WHEN** Fake Tool/Projector tests cover success, no evidence, error, duplicate, TTL, and capacity cases
|
|
- **THEN** later Tool-specific changes can depend on the same boundary without copying lifecycle or Redis logic, while legacy call sites remain unchanged
|
|
### Requirement: Run progress projection SHALL reference canonical records without duplicating truth
|
|
The Run progress tracker SHALL retain only ordered canonical identities for completed Tool calls. At Tool-loop completion, a projector SHALL resolve those identities through the existing canonical store and SHALL accept only READY records owned by the current Run. The tracker SHALL NOT store or reconstruct raw Tool responses, complete Agent results, or a second durable evidence record.
|
|
|
|
#### Scenario: Completed calls are projected
|
|
- **WHEN** a Run ends after multiple READY canonical Tool invocations
|
|
- **THEN** the progress projector reads each indexed canonical record in execution order and creates bounded observed facts
|
|
|
|
#### Scenario: Indexed identity is invalid
|
|
- **WHEN** an indexed canonical identity is missing, expired, cross-Run, PROJECTING, or ERROR
|
|
- **THEN** it is excluded from observed facts and cannot become verified evidence
|