refactor(harness): remove legacy agent architecture

This commit is contained in:
zhuyongxin
2026-07-22 18:02:01 +08:00
parent bc36248cd8
commit 8ee7cc0b70
148 changed files with 3091 additions and 13889 deletions
@@ -0,0 +1 @@
committed
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-22
@@ -0,0 +1,124 @@
## Context
阶段 6B 后,`POST /api/chat` 已经只调用新 `ChatApplicationUseCase`,但仓库仍有四类遗留:一是 bundled frontend 可达的 `/api/ai_ops` Supervisor/Planner/Executor 链;二是只有测试引用的 `ChatService`、Redis Session API、Verifier/Gatekeeper/ThreadLocal;三是被新 Harness adapter 复用但仍带旧 `@Tool` 和 DB recorder 副作用的 RAG/log backend;四是仍描述旧架构的文档和 demo。
新 Harness 已以 Redis canonical invocation 保存当前 Run 完整 Tool 调用并用 EvidenceGuard 验真,但 `ToolBoundary` 没有写长期 `tool_invocation` audit。当前 `AgentLoggingHook` 又会持久化/日志输出模型正文、Tool arguments 和 `thought`,且回退到 Session ThreadLocal。阶段 7 必须同时删除旧链和收紧审计,否则代码结构与最终 E2E 都不能满足 ISS-014。
这是 L4 协议/前端删除和跨模块物理重构。历史数据库表与历史数据保持不变;现有 `chat_session` JPA 元数据仍由 `JpaChatRunStore` 使用,不属于 Redis conversation Session 遗留。
## Goals / Non-Goals
**Goals:**
- 业务代码只剩一个拥有 Tool loop 的 Diagnosis ReAct Agent,公开诊断只剩 `/api/chat`。
- 删除旧 Chat/AiOps/Session 编排、Hook、ThreadLocal、prompts、Tool compatibility annotations 和过时测试。
- RAG/log backend 只通过 Harness adapters 调用,不自行写旧 ToolInvocation。
- AgentStep 与 ToolInvocation durable audit 按 exact Run 持久化有界脱敏 metadata。
- 文档、issues、OpenSpec strict 状态与最终代码一致。
- 真实启动应用并完成 SSE、日志、MySQL exact-run E2E。
**Non-Goals:**
- 不改变 Intent Router、Diagnosis Draft、EvidenceGuard、SemanticGuard、Fallback 或 SSE 五事件语义。
- 不删除历史表/数据,不做 schema destructive migration。
- 不实现真实 CLS 或生产业务 MySQL datasource live 连接。
- 不改 RAG 算法、Trace API response schema 或 UI 视觉设计。
## Decisions
### 1. 删除 legacy AiOps 和 Redis Session surface,不迁移第二条诊断链
删除 `AiOpsController`、`AiOpsService`、AIOps DTO/config/rule evaluation、三个 prompts、前端按钮/consumer 和相应 tests。删除 `ChatSessionController`、Redis `SessionManager`/`SessionContext` 和 tests。告警诊断通过 `/api/chat` 提交自然语言或结构化文本,由同一个 Intent Router 进入 Diagnosis。
替代方案是把 `/api/ai_ops` 适配到新 use case;拒绝,因为它保留第二个公开诊断协议和前端双轨,违反 ISS-014 唯一入口与无兼容分支目标。`chat_session` entity/repository 不删除,因为它是当前 Run 目录与 PreviousTurn 生命周期的一部分。
### 2. 旧 Chat orchestration 按闭包删除
删除 `ChatService` 及其 Sequential Planner/Executor/Verifier/Composer tests、旧 prompts、Verifier/Gatekeeper helpers、Skill metadata hook、Token ThreadLocal wrapper、旧 Tool trace summary/recorder。删除顺序为公开 caller -> service/orchestration -> helper/hook/context -> resources/tests,并在每层后编译/引用扫描。
`SelfEvaluationMergeService`、`DiagnosisTraceService`、repositories 和 eval 基础设施若仍有非旧链调用则保留。替代仅取消 Spring annotation 而保留死类;拒绝,因为阶段 7 明确要求物理清理。
### 3. RAG/log 是 backend,不是 Agent-facing Tool
保留 `LookupKnowledgeTool.lookupKnowledge` 的检索管线和 `QueryLogsTools.queryLogs` 的 Mock/边界实现,供 `RagToolAdapter`/`QueryLogsToolAdapter` 调用;移除所有 `@Tool/@ToolParam`、旧 topic discovery contract、Session ThreadLocal、session dedup 和 `ToolInvocationRecorder`。Agent-facing schema/description 只来自 `HarnessEvidenceTools` 与 `AgentToolContracts`。
替代完全重写 RAG/log;拒绝,因为阶段 3B 已验证 adapter/projector contract,阶段 7 只需清除双 contract 与副作用。
### 4. Agent audit 使用 Harness-native metadata-only Hook
新增 Harness audit hook,强制从 `RunnableConfig` 读取 exact sessionId/runId。before-model 只保存 message count/roles;after-model 只保存是否有文本、是否有 Tool call、Tool names 和 duration,不保存 message content、model output text、Tool arguments、Prompt 或 Thought。缺少 metadata 时跳过持久化并记录安全 warning,不回退 ThreadLocal。
替代继续修补 `AgentLoggingHook`;拒绝,因为旧类包含 verifier 特例、Token ThreadLocal、正文日志和 thought persistence,保留会让旧边界继续存在。
### 5. ToolBoundary 通过安全 port 写 durable audit
新增 `ToolInvocationAuditSink` 和不可变 audit event。`ToolBoundary.execute` 在 canonical 结果确定后调用 sink;event 只包含 sessionId、runId、toolCallId、toolName、InvocationStatus、EvidenceStatus、stable error code、duration、request/agent-result byte count。audit 调用 fail-open:JPA 审计失败写 warning,但不改变已经确定的 Tool observation;canonical Redis store 仍按原设计 fail-closed。
JPA adapter 复用现有 `tool_invocation` 表:`input_params` 只保存 tool_call_id/request_bytes,`output_preview` 只保存 status/evidence_status,`output_length` 保存 agent result bytes,`error_message` 只保存 stable error code。禁止 raw response、完整 request、SQL、日志正文和模型内容。
替代让旧 backend recorder 继续双写;拒绝,因为它依赖 ThreadLocal、不同 Tool 各自实现且可能保存大 payload。替代审计失败阻断业务;拒绝,因为 durable audit 不是 canonical evidence truth source。
### 6. 文档与 issue 以当前单链架构为准
重写当前 MVP、Agent、Harness、Trace/Demo 文档中的旧双入口与多 Agent 描述;ISS-012/ISS-013 标记由 ISS-014 吸收并归档。历史架构文档保留在 archive,不篡改历史。
### 7. 最终 E2E 使用 exact metadata identity
启动 Spring Boot 后向 `/api/chat` 发送固定 payment-timeout 诊断,严格解析 named SSE,取得 metadata sessionId/runId 并要求 `content|failure -> done`。检查 `logs/` 不含 Prompt/Thought/raw Tool payload/stack leakage;使用 `scripts/query_mysql.py` 按 exact runId 查询 `diagnosis_run`、`agent_step`、`tool_invocation`,核对 intent/outcome/status、唯一 Diagnosis Agent、Tool 名称/次数、同一 identity 和最终 safe content。
query_logs 使用 Mock,query_mysql 使用配置的隔离 datasource contract;不声称真实 CLS/生产业务库 live。若真实模型选择非 Diagnosis intent,则使用明确诊断 query 重试一次,不伪造数据库结果。
## Module and Ownership Audit
```text
Browser POST /api/chat
-> ChatController / ChatSseSession
-> ChatApplicationUseCase
-> Intent Router
-> System / Knowledge / Diagnosis executor
-> Diagnosis Harness
-> Diagnosis Agent (only Tool loop)
-> HarnessEvidenceTools
-> ToolBoundary
-> Redis canonical invocation (full, TTL)
-> Durable audit sink (bounded metadata)
-> EvidenceGuard / SemanticGuard / Release
-> named SSE safe content
-> diagnosis_run + agent_step + tool_invocation exact-run Trace
```
- Application owns Session/Run/PreviousTurn;Controller 不拥有业务状态。
- Core owns budget/cancel/lifecycle;Diagnosis Agent owns diagnosis authorship;Guards own verification/release。
- Redis canonical invocation owns short-lived complete Tool truth;MySQL ToolInvocation owns long-lived metadata audit。
- `chat_session` is current JPA run directory;Redis SessionContext is legacy conversation memory and is removed。
- 最大耦合风险是复用 backend 时旧 annotations/recorder 被 Spring 自动发现;static scan + context startup 覆盖。
## Interface Impact
- Level: L4 breaking HTTP/frontend contract。
- Removed: `POST /api/ai_ops`、`POST /api/chat/clear`、`GET /api/chat/session/{sessionId}`、`GET /api/chat/session/{sessionId}/runs`、legacy AiOps message-wrapper SSE。
- Retained: `POST /api/chat` named SSE、Trace/feedback/document/search APIs。
- Consumer migration: bundled frontend 同 commit 删除 AiOps 按钮和 consumer;外部调用方迁移到 `/api/chat`。
- Rollback: 整体回滚阶段 7 commit;不单独恢复旧 endpoint 或 old Tool discovery。
## Risks / Trade-offs
- [外部 AiOps caller 失败] -> L4 文档明确迁移到 `/api/chat`,不提供双轨。
- [删除 bean 导致 context 启动失败] -> compile、focused context test、全量 tests、真实 Spring startup 分层验证。
- [backend 仍被 ToolCallbackProvider 发现] -> 删除 annotations/imports 并静态扫描旧 Tool 名/description。
- [audit 泄漏敏感内容] -> typed metadata-only event + serialization tests + DB query preview inspection。
- [audit 写失败丢 Trace 明细] -> warning + Run 主记录仍完成;验收环境要求 audit row 存在,生产运行可观测 failure。
- [live 环境依赖不可用] -> 先检查服务与启动日志,分类环境/实现失败;不降低为假 E2E。
## Migration Plan
1. 删除 frontend/Controller/Service 可达旧入口和旧 orchestration 闭包。
2. 解耦 RAG/log backend,替换 Agent/Tool durable audit。
3. 删除剩余 dead resources/tests/config,跑 compile/focused/full regression 和 static scans。
4. 更新 docs/issues/OpenSpec,并修复已知 strict validation。
5. 启动应用,执行 final SSE/log/DB E2E,记录 exact identity evidence。
6. 同一 commit 部署所有删除与文档;回滚整体回滚该 commit。
## Open Questions
- None。live 默认值是否需要校准由 E2E 数据决定,但不会改变 architecture direction。
@@ -0,0 +1,55 @@
## Why
阶段 0-6B 已把公开 Chat 切换到单一 Diagnosis ReAct Agent + Harness,但仓库仍保留可达的 `/api/ai_ops` Supervisor/Planner/Executor 多 Agent 链、无生产调用方的旧 `ChatService`、Redis Session endpoint、ThreadLocal、Verifier/Gatekeeper Hook、旧 Agent-facing Tool annotations 和过时文档。新 Harness 的 ToolBoundary 也只写 Redis canonical invocation,尚缺最终 E2E 所需的脱敏 `tool_invocation` durable audit。
## What Changes
- 删除 `/api/ai_ops`、bundled frontend AI Ops 按钮/consumer,以及对应 AiOps service、DTO、prompt、rule evaluation 和测试;公开诊断只保留 `/api/chat`。
- 删除无生产调用方的旧 `ChatService`、Planner/Executor/Verifier/Composer prompts、Gatekeeper/Verifier helpers、旧 session controller/manager/context 和过时测试。
- 保留 RAG 与 Mock logs 的底层查询能力,但移除旧 `@Tool` contract、Session ThreadLocal、session dedup 和 `ToolInvocationRecorder` 副作用;Diagnosis Agent 只能看到 `HarnessEvidenceTools` 的三项 ACI contract。
- 用 Harness-native、metadata-only Agent trace hook 替代旧 `AgentLoggingHook`,禁止持久化 Thought、Prompt、Tool arguments 或模型正文。
- 为 ToolBoundary 增加 fail-open durable audit port 与 JPA adapter,只保存 exact session/run/tool_call identity、工具名、状态、耗时和有界脱敏摘要,不保存 raw response 或完整请求。
- 删除不再使用的配置、logback logger、资源和测试,修正阶段 3B/3C 已知 OpenSpec strict 格式问题。
- 更新当前架构、Agent/Harness、Trace、Demo 和 API 文档;将已被 ISS-014 吸收的 ISS-012/ISS-013 归档。
- 启动真实 Spring Boot 应用,执行最终 Chat SSE E2E,并用 `logs/` 与 `scripts/query_mysql.py` 核对 exact sessionId/runId、Run、AgentStep、ToolInvocation 和安全发布结果。
## Capabilities
### New Capabilities
- `single-react-cleanup-e2e`: 定义旧链物理删除、唯一公开 Chat surface、Harness durable audit、安全 Trace 和最终 live E2E 验收。
### Modified Capabilities
- `single-react-chat-sse-cutover`: 删除阶段 6B 暂时保留的 legacy AiOps 与 Session endpoints,使唯一 Chat surface 成为最终状态。
## Scope
- Controller/service/hook/util/tool/resource/test 的旧架构删除与 Harness-native audit replacement。
- bundled frontend AI Ops 路径清理。
- MVP architecture/demo/API/issue 文档收口。
- OpenSpec strict validation cleanup。
- Spring Boot live Chat SSE、日志和 MySQL exact-run 验收。
## Non-goals
- 不修改三类 Intent、Diagnosis Agent Draft、EvidenceGuard、SemanticGuard 或 Release Policy 业务语义。
- 不新增 Graph、工作流 DSL、兼容 endpoint、双轨开关或 Token streaming。
- 不删除历史数据库表或历史数据,不迁移生产数据。
- 不声称完成真实 CLS 或生产业务 MySQL datasource live E2E;query_logs 和 query_mysql 继续按 ISS-014 使用 Mock/隔离契约验收。
- 不重新设计 RAG 检索算法、Trace API 或前端视觉体验。
## Context Constraints
- Diagnosis Agent 是唯一拥有 Tool loop 的业务 Agent;项目业务代码不得保留 Supervisor/Sequential/Planner/Executor/Verifier/Composer 编排。
- Agent-facing Tools 只能来自 `HarnessEvidenceTools`,底层 RAG/log 实现不是 Agent contract。
- Durable audit 只能保存有界脱敏元数据;Redis canonical invocation 仍是当前 Run 完整调用记录真理源。
- live E2E 必须对齐 SSE metadata 的 exact sessionId/runId,不能用“最新一条”替代。
- L4 endpoint 删除不提供兼容分支;回滚只能整体回滚阶段 7 commit。
## Risks
- 旧类存在隐式 Spring bean 或动态 Tool discovery,删除不完整会继续把旧 Tools 暴露给模型。
- 旧 Tool recorder 与新 canonical store 并存会产生双写或缺少 exact run 的数据库审计。
- Agent trace hook 若保存模型正文或 Tool arguments,会违反不持久化 Thought/Prompt/raw payload 的边界。
- live 环境依赖模型、Redis、Milvus 和 MySQL;必须区分环境失败、实现失败与验收限制。
@@ -0,0 +1,6 @@
## REMOVED Requirements
### Requirement: AiOps public behavior SHALL remain isolated
**Reason**: Stage 6B temporarily preserved the legacy `/api/ai_ops` path only to make the Chat SSE cutover atomic. Stage 7 removes the remaining Supervisor/Planner/Executor orchestration so `/api/chat` is the single diagnosis entry required by ISS-014.
**Migration**: Bundled and external consumers MUST submit diagnosis requests to the named-event SSE `POST /api/chat`. There is no compatibility endpoint or legacy message-wrapper parser.
@@ -0,0 +1,85 @@
## ADDED Requirements
### Requirement: Legacy diagnosis orchestration SHALL be physically absent
Production and test source SHALL contain no legacy ChatService, AiOps Supervisor/Planner/Executor, Sequential Planner/Executor/Verifier/Composer, legacy Gatekeeper/Verifier helpers, ThreadLocal execution context, or their dedicated prompts and obsolete tests. The only business Agent with a Tool loop SHALL be the Diagnosis Agent constructed inside the Harness.
#### Scenario: Source inventory after cleanup
- **WHEN** production source, resources and tests are scanned
- **THEN** legacy orchestration classes, prompts, ThreadLocals and self-only tests are absent rather than disabled or commented out
#### Scenario: Agent construction inventory
- **WHEN** business Agent builders and framework orchestration types are inspected
- **THEN** only the Diagnosis Agent has evidence Tool callbacks and no Supervisor/Sequential business workflow remains
### Requirement: Public diagnosis SHALL use one endpoint
The application SHALL expose `POST /api/chat` as the only public diagnosis execution endpoint. `/api/ai_ops`, legacy Chat Session management endpoints and bundled frontend callers for those endpoints MUST be removed without a compatibility branch.
#### Scenario: Bundled frontend diagnosis
- **WHEN** a user submits a diagnosis from the bundled frontend
- **THEN** it uses the named-event SSE `/api/chat` consumer and exposes no AiOps mode or button
#### Scenario: Removed endpoint scan
- **WHEN** Controller mappings and frontend request targets are inspected
- **THEN** no `/api/ai_ops`, `/api/chat/clear` or `/api/chat/session` execution/management mapping remains
### Requirement: Evidence backends SHALL not be Agent contracts
RAG and log query implementations MAY be reused behind Harness adapters, but MUST NOT expose `@Tool`, `ToolCallbackProvider`, topic-discovery-first behavior, legacy Tool descriptions, Session ThreadLocal, session dedup or per-backend durable recorder side effects. Agent-facing Tool names, schemas and descriptions SHALL come only from `HarnessEvidenceTools` and `AgentToolContracts`.
#### Scenario: Tool discovery inspection
- **WHEN** Spring Tool annotations and Agent callback registration are inspected
- **THEN** only `lookup_knowledge`, `query_logs` and `query_mysql` Harness ACI callbacks are available to the Diagnosis Agent
#### Scenario: Backend invocation
- **WHEN** a Harness adapter calls RAG or Mock logs
- **THEN** the backend returns raw adapter input without reading ThreadLocal or writing a second ToolInvocation record
### Requirement: Agent durable audit SHALL be metadata-only
Every persisted Diagnosis AgentStep SHALL use the exact RunnableConfig sessionId/runId and MAY contain only bounded metadata such as message count/roles, Tool names, text presence, duration and budget counters. It MUST NOT persist or log Prompt text, message content, model output text, Tool arguments, raw evidence or Thought, and MUST NOT fall back to ThreadLocal identity.
#### Scenario: Model step persistence
- **WHEN** the Diagnosis Agent performs model calls
- **THEN** AgentStep rows use the SSE metadata identity and contain no user query, evidence body, Tool argument or chain-of-thought text
#### Scenario: Missing metadata
- **WHEN** an audit hook is invoked without exact sessionId/runId metadata
- **THEN** it skips persistence with a safe warning rather than inventing or reading implicit identity
### Requirement: ToolBoundary SHALL write safe durable audit
For each accepted Harness evidence Tool call, ToolBoundary SHALL attempt to write one durable audit row with exact sessionId/runId/toolCallId/toolName, invocation/evidence status, stable error code, duration and byte counts. Durable audit MUST NOT contain the complete request, SQL/log query body, raw response or Agent projection content. Audit persistence failure SHALL be observable but MUST NOT change the canonical Tool result.
#### Scenario: Successful Tool call
- **WHEN** ToolBoundary completes a READY invocation
- **THEN** Redis retains the canonical record and MySQL receives one metadata-only ToolInvocation row for the same Run and Tool call
#### Scenario: Failed Tool call
- **WHEN** ToolBoundary returns a stable error after an accepted request
- **THEN** the durable row records only stable status/error metadata and no internal exception or raw payload
#### Scenario: Audit database failure
- **WHEN** durable audit persistence throws after canonical result determination
- **THEN** ToolBoundary logs a safe warning and returns the unchanged canonical Tool result
### Requirement: Current documentation SHALL describe the single Harness architecture
Current MVP architecture, Agent/Harness, Trace, API and demo documents SHALL describe the single Diagnosis Agent, explicit Harness ownership, ACI Tools, named SSE and safe durable audit. ISS-012 and ISS-013 SHALL be recorded as absorbed by ISS-014; historical archived documents MAY retain historical descriptions.
#### Scenario: Current documentation scan
- **WHEN** non-archived current architecture and demo documents are inspected
- **THEN** they do not present Planner/Executor/Verifier/Composer or `/api/ai_ops` as current runtime behavior
### Requirement: Repository verification SHALL be clean
The final implementation SHALL compile, pass focused and relevant full regressions, pass JavaScript syntax and static legacy scans, and pass strict OpenSpec validation for all current specs. No dead imports, temporary files, hardcoded credentials, compatibility flags or unexplained legacy references may be introduced.
#### Scenario: Automated verification
- **WHEN** the stage 7 verification suite runs
- **THEN** all selected tests/build/static/OpenSpec checks pass with no legacy runtime token in production source
### Requirement: Final live E2E SHALL correlate exact Run evidence
The final acceptance SHALL start the Spring Boot application, send a real diagnosis request through `/api/chat`, strictly parse the named SSE sequence, inspect `logs/`, and query MySQL with `scripts/query_mysql.py` using the exact metadata sessionId/runId. It SHALL verify Run intent/outcome/status, one Diagnosis Agent identity, bounded model/Tool counts, ToolInvocation identity, safe final content/fallback and no prohibited content leakage.
#### Scenario: Live successful or safe fallback diagnosis
- **WHEN** the configured model and infrastructure process the fixed diagnosis request
- **THEN** SSE, logs, diagnosis_run, agent_step and tool_invocation evidence agree on the exact identity and final release is SUCCESS or documented safe FALLBACK
#### Scenario: External Tool boundary statement
- **WHEN** final acceptance is archived
- **THEN** it distinguishes Mock query_logs and isolated query_mysql contract evidence from unverified real CLS or production business MySQL integration
@@ -0,0 +1,28 @@
## 1. Remove legacy public and orchestration paths
- [x] 1.1 Remove `/api/ai_ops`, bundled frontend AiOps controls/consumer, AiOps DTO/config/service/rule evaluation/prompts/tests and obsolete logger/config references.
- [x] 1.2 Remove legacy `ChatService`, Sequential Planner/Executor/Verifier/Composer prompts, Gatekeeper/Verifier/Skill/Token helpers, ThreadLocals and self-only tests/resources.
- [x] 1.3 Remove legacy Chat Session controller, Redis SessionManager/SessionContext and tests while retaining current JPA `chat_session` Run metadata.
- [x] 1.4 Compile and run reference/static scans proving no production Supervisor/Sequential legacy orchestration or removed endpoint remains.
## 2. Make Tool and audit boundaries Harness-native
- [x] 2.1 Refactor `LookupKnowledgeTool` into a Harness-only backend by removing old Tool annotation, Session ThreadLocal, session dedup and legacy ToolInvocation recording while preserving retrieval behavior tests.
- [x] 2.2 Refactor `QueryLogsTools` into a Harness-only Mock/log backend by removing old Tool annotations, topic-discovery-first contract and legacy recorder side effects while preserving adapter contract tests.
- [x] 2.3 Replace `AgentLoggingHook` with a metadata-only Harness audit hook that requires exact RunnableConfig IDs and persists no Prompt, message/model content, Tool arguments or Thought.
- [x] 2.4 Add ToolBoundary audit event/sink and a fail-open JPA adapter that persists one bounded metadata-only ToolInvocation row per accepted Harness call.
- [x] 2.5 Add focused negative/safety tests for exact audit identity, no sensitive payload persistence, audit failure isolation and exclusive Harness Tool discovery.
## 3. Close documentation and repository contracts
- [x] 3.1 Update current MVP architecture, Agent/Harness, Trace/API and demo documents to the single Diagnosis Agent, ACI Tool and named SSE architecture.
- [x] 3.2 Mark ISS-012 and ISS-013 absorbed by ISS-014 and archive them; update ISS-014 stage/completion state only after final acceptance.
- [x] 3.3 Fix known `mysql-readonly-tool` and `rag-log-projections` main spec strict-validation defects without changing their behavior.
## 4. Final verification and live acceptance
- [x] 4.1 Run focused stage 7 tests, relevant stage 1-6B regressions, Maven compile/package, JavaScript syntax, static legacy/sensitive scans and strict OpenSpec validation.
- [x] 4.2 Start the Spring Boot application with Maven and verify clean context startup plus bounded Harness/Audit bean wiring in `logs/`.
- [x] 4.3 Execute a real `/api/chat` diagnosis, strictly validate named SSE order/outcome and capture exact metadata sessionId/runId.
- [x] 4.4 Use `logs/` and `scripts/query_mysql.py` to verify exact diagnosis_run, Diagnosis AgentStep, ToolInvocation identity/count/status/safe fields and released content/fallback; record Mock/isolated external Tool limits.
- [x] 4.5 Record final cleanup inventory, validation outputs, live E2E evidence, remaining risks and L4 migration/rollback in devflow acceptance before archive.