Compare commits
8
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
ee0949d464 | ||
|
|
2362665519 | ||
|
|
85029d96a7 | ||
|
|
3e602781d6 | ||
|
|
0dbdd7d8d3 | ||
|
|
6b74990f86 | ||
|
|
4274f3350b | ||
|
|
58c39107c5 |
@@ -148,7 +148,7 @@ openspec/changes/phase-1-infrastructure/
|
||||
|
||||
## 敏感信息(已编辑)
|
||||
|
||||
- MySQL 密码:已配置在 application.yml(`!Fucker123..`)
|
||||
- MySQL 密码:已从仓库移除,使用环境变量注入
|
||||
- Redis:无密码
|
||||
|
||||
---
|
||||
|
||||
@@ -97,7 +97,7 @@ spring:
|
||||
redis:
|
||||
host: 119.29.78.52
|
||||
port: 6379
|
||||
password: '!Fucker123..'
|
||||
password: ${SUPERBIZ_REDIS_PASSWORD}
|
||||
database: 0
|
||||
timeout: 3000
|
||||
```
|
||||
|
||||
@@ -140,7 +140,7 @@ Error Code: 1049
|
||||
datasource:
|
||||
url: jdbc:mysql://119.29.78.52:33306/superbiz_agent?...
|
||||
username: root
|
||||
password: '!Fucker123..'
|
||||
password: ${SUPERBIZ_MYSQL_PASSWORD}
|
||||
```
|
||||
|
||||
**Redis 配置**:
|
||||
@@ -149,7 +149,7 @@ data:
|
||||
redis:
|
||||
host: 119.29.78.52
|
||||
port: 6379
|
||||
password: '!Fucker123..'
|
||||
password: ${SUPERBIZ_REDIS_PASSWORD}
|
||||
```
|
||||
|
||||
**Flyway 配置**:
|
||||
|
||||
@@ -44,6 +44,13 @@ build/
|
||||
app.log
|
||||
logs/
|
||||
|
||||
### Local Secrets ###
|
||||
.env
|
||||
.env.*
|
||||
!.env.example
|
||||
application-local.yml
|
||||
application-*.local.yml
|
||||
|
||||
### Upload Files ###
|
||||
uploads/
|
||||
|
||||
|
||||
@@ -155,6 +155,46 @@
|
||||
- 使用场景:Executor 按 skill workflow 调用 evidence tools 收集事实,`tool_invocation` 记录这些事实证据。
|
||||
- 边界:最终诊断结论必须被 evidence tools 支撑,不能仅由 skill 正文支撑。
|
||||
|
||||
### Diagnosis Harness
|
||||
- 定义:围绕 Diagnosis Agent 提供确定性运行控制的边界,负责 Run、预算、取消、重试装配、Tool 调用记录、证据验真和最终释放,不承担业务诊断推理。
|
||||
- 边界:Harness 不是工作流引擎,不实现 Planner/Executor/Composer 节点或自行编写 ReAct 循环。
|
||||
|
||||
### Diagnosis Agent
|
||||
- 定义:诊断链路中唯一拥有 ReAct 工具循环并生成 `DiagnosisDraft` 的 Agent,负责规划证据查询、判断证据充分性和撰写完整诊断草稿。
|
||||
- 边界:不负责意图路由、Run/Session 生命周期、证据物理验真、独立语义审查或最终发布;证据不足时必须明确停止并保留限制。
|
||||
|
||||
### EvidenceGuard
|
||||
- 定义:Harness 内部的确定性证据验真能力,校验 Draft 引用、当前 Run 所有权、Tool 调用状态和有界 Agent 投影。
|
||||
- 边界:EvidenceGuard 不调用 LLM,也不判断证据是否足以推出业务结论。
|
||||
|
||||
### SemanticGuard
|
||||
- 定义:使用隔离上下文对完整诊断 Draft 与已验真证据做报告级语义审查的单轮 Agent。
|
||||
- 边界:无工具、无记忆、无 ReAct 循环,不访问 Redis,不生成或改写用户报告。
|
||||
|
||||
### Invocation Status
|
||||
- 定义:Tool 调用及结果投影的生命周期状态,固定为 `PROJECTING`、`READY`、`ERROR`。
|
||||
- 边界:它只说明调用记录是否完成,不说明结果是否包含证据。
|
||||
|
||||
### Evidence Status
|
||||
- 定义:证据 Tool 的结果语义,固定为 `EVIDENCE_FOUND`、`NO_EVIDENCE`、`ERROR`。
|
||||
- 边界:`NO_EVIDENCE` 只表示当前查询范围内没有匹配结果,不能解释为问题不存在、根因被排除或系统健康。
|
||||
|
||||
### RunContext
|
||||
- 定义:一次 Diagnosis Run 的显式执行上下文,结构不可变地携带 `sessionId`、`runId`、deadline,以及该 Run 独占的取消、预算、重试策略和生命周期状态句柄。
|
||||
- 边界:RunContext 通过方法参数或框架受控 context 显式传播,不依赖 ThreadLocal;结构不可变不等于内部计数和取消状态不能变化,这些变化由线程安全句柄管理。
|
||||
|
||||
### Run Lifecycle
|
||||
- 定义:Diagnosis Harness 对单次 Run 执行状态的内存控制,采用 first-terminal-wins 规则保证成功、失败、取消、超时和预算耗尽只能产生一个最终终态。
|
||||
- 边界:Run Lifecycle 不直接等同于数据库实体写入;应用用例负责把最终状态映射到 `diagnosis_run` 持久化。
|
||||
|
||||
### Run Budget
|
||||
- 定义:单次 Run 的模型调用、Tool 调用、单 Tool 调用、输入/输出/总 Token 和 canonical invocation 字节容量的线程安全消耗计数与门禁。
|
||||
- 边界:预算上限由 Harness 配置显式提供;实际 Token 在模型响应后记录,超限后保留真实消耗并阻止后续执行。
|
||||
|
||||
### Harness Retry Policy
|
||||
- 定义:Harness 对同一技术操作 attempt 数和可重试失败类型的显式策略。
|
||||
- 边界:Router 与 SemanticGuard 的技术失败最多两次 attempt;Diagnosis Agent、Tool 和 Evidence repair 只有一次 attempt。Agent 正常 ReAct 轮次不是 retry,`NO_EVIDENCE`、业务拒绝、取消和预算耗尽不可重试。
|
||||
|
||||
### Verifier Skill Isolation
|
||||
- 定义:Chat Verifier 与 skill 系统隔离,只校验 Executor 答案和 `tool_trace_summary`。
|
||||
- 使用场景:防止 Verifier 把 playbook 指令当作事实证据;Verifier 只判断已有证据是否支持结论。
|
||||
|
||||
@@ -4,6 +4,10 @@
|
||||
|
||||
| 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|
||||
|---|---|---|---|---|---|---|
|
||||
| 2026-07-21 | single-react-tool-invocation-store | 建立统一 ToolBoundary 与 Redis canonical invocation store,集中生命周期、证据状态、TTL、容量和 Run 所有权。 | Harness/Tool boundary/Canonical store | ISS-014, ToolBoundary, canonical invocation, PROJECTING, READY, ERROR, TTL, RESULT_TOO_LARGE | openspec/changes/archive/2026-07-21-single-react-tool-invocation-store | archived |
|
||||
| 2026-07-21 | single-react-harness-run-context | 建立显式 RunContext、Harness Core、预算、取消、类型化重试和 Tool Store 基础。 | Harness/Run lifecycle/Budget | ISS-014, RunContext, deadline, cancellation, budget, retry, ToolCallKey | openspec/changes/archive/2026-07-21-single-react-harness-run-context | archived |
|
||||
| 2026-07-21 | single-react-aci-tool-contracts | 冻结 RAG、日志和 MySQL evidence Tool 的 Agent-facing ACI Schema、状态、框架调用引用和描述边界。 | Harness/Agent Tool contract | ISS-014, ACI, tool_call_id, evidence_status, RAG, query_logs, query_mysql, MOCK | openspec/changes/archive/2026-07-21-single-react-aci-tool-contracts | archived |
|
||||
| 2026-07-21 | single-react-design-freeze | 冻结单体 Diagnosis Agent、Harness、Guard、工具证据与阶段门禁契约。 | Chat/Harness/Agent contract | ISS-014, single ReactAgent, Harness, EvidenceGuard, SemanticGuard, tool_call_id, evidence_status | openspec/changes/archive/2026-07-21-single-react-design-freeze | archived |
|
||||
| 2026-07-10 | session-run-trace-isolation | 拆分会话态和运行态,引入 runId 隔离 Trace、Feedback、AIOps 和 demo 链路。 | Trace/session/run isolation | chat_session, diagnosis_run, runId, trace exact run, feedback fallback, AIOps SSE metadata, baseline drift | openspec/changes/archive/2026-07-10-session-run-trace-isolation | archived |
|
||||
| 2026-07-09 | interview-demo-quality-audit | 增加面试演示前置质量审计,覆盖 prompt、Gatekeeper 和评测基线。 | Agent eval/demo/Prompt audit | interview demo preflight, prompt_audit, gatekeeper rules, diagnosis baseline, 12 fixtures | openspec/changes/archive/2026-07-09-interview-demo-quality-audit | archived |
|
||||
| 2026-07-08 | executor-composer-final-answer | 引入 Composer 生成最终回答,只使用 Verifier 允许的结论材料。 | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
|
||||
@@ -33,3 +37,7 @@
|
||||
| 2026-06-24 | lookup-knowledge-integration | 接入知识库检索,支持 L0 精确匹配和 L1 语义检索。 | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | - | archived |
|
||||
| 2026-06-23 | phase1-infrastructure | 搭建第一阶段基础设施,包括 MySQL、Redis、Milvus、Flyway 和 JPA。 | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | - | archived |
|
||||
| 2026-05-29 | chatmodel-abstraction | 抽象 ChatModel 和 EmbeddingModel,支持多模型路由。 | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | - | archived |
|
||||
| 2026-07-21 | single-react-rag-log-projections | RAG/log projection adapters through ToolBoundary | Harness/Tool projection | ISS-014, RAG, query_logs, projection, scope, redaction, MOCK, NO_EVIDENCE | openspec/changes/archive/2026-07-21-single-react-rag-log-projections | archived |
|
||||
| 2026-07-21 | single-react-mysql-readonly-tool | Fail-closed read-only MySQL evidence Tool with AST allowlist, JDBC controls and bounded projection | Harness/MySQL security | ISS-014, MySQL, JSqlParser, allowlist, PreparedStatement, timeout, projection | openspec/changes/archive/2026-07-21-single-react-mysql-readonly-tool | archived |
|
||||
| 2026-07-21 | single-react-diagnosis-agent | Single internal Diagnosis ReactAgent with Harness-controlled model/tool loop, bounded context and typed Draft | Harness/Diagnosis Agent/ReAct | ISS-014, ReactAgent, DiagnosisDraft, PreviousTurn, ToolInterceptor, ModelInterceptor, budget | openspec/changes/archive/2026-07-21-single-react-diagnosis-agent | archived |
|
||||
| 2026-07-21 | single-react-evidence-semantic-guards | Deterministic evidence validation, isolated semantic review and fail-closed diagnosis release | Harness/EvidenceGuard/SemanticGuard/Release | ISS-014, EvidenceGuard, verified snapshot, SemanticGuard, repair, fallback, release policy | openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards | archived |
|
||||
|
||||
@@ -0,0 +1,44 @@
|
||||
# Acceptance: single-react-aci-tool-contracts
|
||||
|
||||
## 实现结果
|
||||
|
||||
- RAG Contract:最小 query Request、bounded document evidence Result。
|
||||
- Log Contract:逻辑 Topic/Lookback Request、Mock provenance、Scope/Pattern/Event Result。
|
||||
- MySQL Contract:逻辑 data source、参数化 SQL Request、bounded structured rows Result。
|
||||
- 共享 Contract:snake_case Tool 名称、ACI 描述、不可变集合 helper,复用阶段 0 两套状态枚举。
|
||||
- 旧 Tool、Chat/AIOps、Controller、持久化、数据源和公开协议未修改。
|
||||
|
||||
## 静态验证
|
||||
|
||||
- `openspec validate single-react-aci-tool-contracts --strict`:通过。
|
||||
- `openspec instructions apply --change single-react-aci-tool-contracts --json`:12/12 tasks complete。
|
||||
- `rg` 引用检查:新 contract 生产包未接入旧运行链路。
|
||||
- 受保护文件 diff scope:旧 Tool、Chat/AIOps、Controller、Repository、resources 均为空。
|
||||
|
||||
## 脚本验证
|
||||
|
||||
- `mvn -q -DskipTests compile`:通过。
|
||||
- `mvn -q '-Dtest=RagToolContractTest,QueryLogsToolContractTest,MysqlToolContractTest' test`:通过。
|
||||
- `mvn -q '-Dtest=HarnessContractTest,LookupKnowledgeToolTest,QueryLogsToolsTest' test`:通过。
|
||||
|
||||
## 浏览器/人工验证
|
||||
|
||||
- 不适用。本阶段无 UI、Controller、SSE 或公开运行行为变化。
|
||||
|
||||
## 未验证
|
||||
|
||||
- 未执行 live LLM/Redis/CLS/MySQL E2E;本阶段没有接入这些运行路径,最终 live E2E 按 ISS-014 门禁留到阶段 7。
|
||||
- 未验证 ToolInterceptor 的真实 ID 传播;阶段 2/3A 必须以 `ToolCallRequest.getToolCallId()` 添加集成测试。
|
||||
- 未验证真实日志 adapter 的 `SourceKind` 扩展;本 Issue 首版明确只使用 Mock。
|
||||
- Provider 侧旧凭据轮换仍需凭据所有者完成,仓库只能证明明文已移除。
|
||||
|
||||
## 剩余风险与后续门禁
|
||||
|
||||
- 新旧 Contract 短期并存,阶段 3B/3C 接入前不得声称旧 Tool 已符合新 ACI 输出。
|
||||
- 下一阶段只能在本 change OpenSpec Archive 和 Git commit 完成后开始。
|
||||
|
||||
## 状态
|
||||
|
||||
- Stage acceptance: accepted
|
||||
- OpenSpec archive: archived at `openspec/changes/archive/2026-07-21-single-react-aci-tool-contracts`
|
||||
- Main spec sync: `openspec/specs/aci-evidence-tool-contracts/spec.md`(7 added requirements)
|
||||
@@ -0,0 +1,32 @@
|
||||
# Brief: single-react-aci-tool-contracts
|
||||
|
||||
## 背景
|
||||
|
||||
旧 RAG 和日志 Tool 暴露检索/基础设施细节与不一致状态,MySQL Tool 尚无 Agent-facing 类型。Harness 实现前需要先冻结三类最小 ACI 契约。
|
||||
|
||||
## 目标
|
||||
|
||||
- 冻结 RAG、日志、MySQL 的 Request/Result JSON Schema。
|
||||
- 统一 `evidence_status`,并保持其与 invocation lifecycle 独立。
|
||||
- 统一使用框架 `tool_call_id`,禁止模型传入或 Harness 生成第二套 ID。
|
||||
- 冻结简短 Tool 名称/描述和 Mock 日志来源边界。
|
||||
- 用三个独立契约测试锁定行为。
|
||||
|
||||
## 范围
|
||||
|
||||
- 新增 `com.superbiz.agent.harness.tool.contract` 值对象、枚举和描述常量。
|
||||
- 复用阶段 0 的 `InvocationStatus` 与 `EvidenceStatus`。
|
||||
- 验证 JSON、不可变集合、描述泄漏和新旧边界。
|
||||
|
||||
## 非目标
|
||||
|
||||
- 不切换旧 RAG/日志运行方法或 Chat/AIOps 注册。
|
||||
- 不实现投影、Redis store、真实日志适配器或 MySQL 执行。
|
||||
- 不修改 Controller/SSE 或公开协议。
|
||||
|
||||
## 元数据
|
||||
|
||||
- 分档:standard
|
||||
- 接口影响:L2 前置内部契约;当前运行行为无变化
|
||||
- 关联 Issue:ISS-014 阶段 1
|
||||
- 关联 OpenSpec:`openspec/changes/single-react-aci-tool-contracts`
|
||||
@@ -0,0 +1,116 @@
|
||||
# Decisions: single-react-aci-tool-contracts
|
||||
|
||||
## 规模与入口
|
||||
|
||||
- 分档:standard。
|
||||
- 入口:ISS-014 阶段 1,前置 `single-react-design-freeze` 已 Archive 并由 Git commit `58c3910` 固化。
|
||||
- 目标:冻结三类 evidence Tool 的 Agent-facing ACI 契约,不接入新运行链路。
|
||||
|
||||
## Context
|
||||
|
||||
- `devflow/index.md` 命中 `single-react-design-freeze`、`modular-rag-pipeline` 和 `evidence-trace-hardening`。
|
||||
- `devflow/glossary/CONTEXT.md` 已定义 Diagnosis Harness、Invocation Status、Evidence Status 与 Evidence Tools。
|
||||
- 阶段 0 已确认 `tool_call_id` 使用框架 ID、两套状态语义分离、阶段串行门禁和阶段 6B 才公开切换。
|
||||
- 未发现根目录旧 `CONTEXT.md` 与 glossary 冲突。
|
||||
|
||||
## Question Pool
|
||||
|
||||
| 维度 | 问题 | 模式 | 证据与结论 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| 术语 | invocation lifecycle 与 evidence result 是否使用同一状态? | evidence-driven | ISS-014 5.2、阶段 0 contract 明确分离;分别复用 `InvocationStatus` 与 `EvidenceStatus`。 | 已解决并汇报 |
|
||||
| 术语 | `tool_call_id` 由谁生成、从哪里取得? | evidence-driven | Spring AI `AssistantMessage.ToolCall.id()` 与 Alibaba `ToolCallRequest.getToolCallId()` 提供框架 ID;普通 `ToolContext` 不自动加入该 ID。Harness 不生成第二套 ID。 | 已解决并汇报 |
|
||||
| 边界 | 阶段 1 是否直接改旧 RAG/日志执行签名与返回值? | evidence-driven | ISS-014 阶段 1 只冻结 Contract,投影在阶段 3B、公开切换在阶段 6B;本阶段只新增契约代码和测试。 | 已解决并汇报 |
|
||||
| 边界 | 日志阶段是否实现真实 CLS/MCP 或保留 Topic discovery? | evidence-driven | ISS-014 8.1 明确继续 Mock、删除 Agent 侧 discovery、真实适配器不在本 Issue 提前设计。 | 已解决并汇报 |
|
||||
| 验收 | 如何证明契约已冻结且有界? | evidence-driven | 对三类独立 DTO 做精确 JSON、不可变集合、状态和描述泄漏测试;不以旧 Tool 集成测试代替。 | 已解决并汇报 |
|
||||
| 技术 | 框架 ID 能否在后续 Harness 边界取得? | evidence-driven | 本地依赖 Spring AI Alibaba 1.1.2.0 暴露 `ToolInterceptor.interceptToolCall(ToolCallRequest, ToolCallHandler)`,request 含 `toolCallId`。 | 已解决并汇报 |
|
||||
|
||||
## Grill 结论
|
||||
|
||||
- 术语、边界、验收三类问题均已由代码、依赖 API、阶段 0 档案和 ISS-014 证明。
|
||||
- 没有需要新增用户偏好或风险取舍的 `user-interview` 问题;不代理确认任何新方向。
|
||||
- `grill-with-docs` 要求的代码可证问题已先查证;结论已向用户汇报。
|
||||
- Proposal 已回写框架 ID 接入点、旧运行链路不切换、Mock 边界和 L2 接口影响。
|
||||
|
||||
## 已确认决策
|
||||
|
||||
- DTO 放在 Harness 的 Tool Contract 边界,复用阶段 0 的共享状态枚举,不在旧 `dto` 包继续堆叠协议。
|
||||
- 三类结果只携带 `evidence_status`,canonical invocation 的 `status` 保持独立;阶段 1 通过测试冻结枚举,不提前定义存储实现。
|
||||
- Tool description 使用代码常量冻结,后续 Tool adapter 注册时复用;旧 `@Tool` 注解本阶段不改,避免提前改变运行行为。
|
||||
- RAG 输入仅保留 `query`;日志输入仅保留逻辑 `topic/query/lookback_minutes`;MySQL 输入仅保留逻辑 `data_source/sql/params`。
|
||||
- 日志 `source_kind` 首版固定支持 `MOCK` 契约值,但保留 enum 扩展位置给后续真实适配器 change 审查。
|
||||
|
||||
## 能力与工具限制
|
||||
|
||||
- Discover 能力来源:`sm-flow` + `grill-with-docs`。
|
||||
- 仓库要求的 `codebase-retrieval` 和 LSP 工具在当前工具集中不可用;已用 `rg` 引用搜索、源码阅读和本地依赖 `javap` 补足事实核对。该限制不改变契约方向,但后续 Apply 仍需通过编译和引用测试验证。
|
||||
|
||||
## Cross-artifact 对齐
|
||||
|
||||
| 链路 | 状态 | 结论 |
|
||||
|---|---|---|
|
||||
| brief 目标/范围/非目标 -> proposal | 已对齐 | 三类 DTO、状态、框架 ID、短描述和不切旧运行链路均有对应。 |
|
||||
| proposal 范围/约束/承诺 -> design | 已对齐 | 包边界、ID 来源、record/defensive copy、三类 Schema 和迁移顺序均已设计。 |
|
||||
| design 决策/接口影响/风险 -> specs/tasks | 已对齐 | L2 边界、字段、状态、描述、Mock provenance 和运行不切换均有可验证 requirement 与任务。 |
|
||||
| specs 可观察行为 -> tasks | 已对齐 | 每类 Contract 都有实现与独立测试,另有状态、回归和 diff scope 验证。 |
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
- 能力来源:`zoom-out`,以项目 glossary 的 Diagnosis Harness、Evidence Tools、Invocation Status 和 Evidence Status 术语审计。
|
||||
- 当前输入到输出链路仍为 `ChatController/ChatService` 或 `AiOpsService -> ReactAgent -> LookupKnowledgeTool/QueryLogsTools -> 旧结果`,新契约没有运行消费者。
|
||||
- 未来链路为 `Diagnosis Agent -> Alibaba ToolInterceptor/Harness -> typed Request -> adapter/store/projector -> bounded Result -> Agent observation`,阶段 1 只占有 typed contract 边界。
|
||||
- 数据所有权保持明确:框架拥有 `tool_call_id`,canonical invocation 拥有生命周期,Tool-specific result 拥有证据语义与有界内容。
|
||||
- 主要耦合风险是新旧契约短期并存被误当成已迁移;通过独立包、无旧调用方修改和后续阶段门禁控制,无 ADR 冲突。
|
||||
|
||||
## Commit Gate Preflight
|
||||
|
||||
- proposal、design、specs、tasks 文件完整,OpenSpec CLI 状态为 complete。
|
||||
- strict validation:`openspec validate single-react-aci-tool-contracts --strict` 通过。
|
||||
- question pool 中没有未汇报的 evidence-driven 结论或未确认的 user-interview 问题。
|
||||
- 接口影响为 L2,当前运行消费者零变更;阶段 6B 的 L4 切换保持独立。
|
||||
- cross-artifact 四段对齐无 gap,架构审计未发现需要回写的新实现约束。
|
||||
- Apply 已由用户对 ISS-014 全阶段的持续授权覆盖;仍严格限制在本 Committed OpenSpec tasks 内。
|
||||
|
||||
## Pre-apply Research
|
||||
|
||||
### 参考实现
|
||||
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`:确认旧 RAG 暴露 `LookupResult`、ContextPack 和检索 Trace,本阶段不改。
|
||||
- `src/main/java/com/superbiz/agent/agent/tool/QueryLogsTools.java`:确认旧日志输入包含 region/logTopic/limit、存在 Topic discovery 和 mutable nested DTO,本阶段不改。
|
||||
- `src/test/java/com/superbiz/agent/tool/LookupKnowledgeToolTest.java`:现有 RAG 行为回归基线。
|
||||
- `src/test/java/com/superbiz/agent/agent/tool/QueryLogsToolsTest.java`:现有 Mock 日志行为回归基线。
|
||||
- `src/main/java/com/superbiz/agent/harness/contract/*.java`:Java 17 record、Jackson snake_case 和共享状态枚举风格。
|
||||
|
||||
### 技术栈清单
|
||||
|
||||
- 请求/响应标准:Java 17 record + Jackson `@JsonProperty`,集合构造时 defensive copy。
|
||||
- Tool 定义:当前使用 Spring AI `@Tool`,新名称/描述先以常量冻结,后续 adapter 注册复用。
|
||||
- 框架调用 ID:Spring AI Alibaba `ToolCallRequest.getToolCallId()`;普通 `ToolContext` 不作为 ID 来源。
|
||||
- 异常与校验:本阶段只冻结结构;Schema、ID、权限、范围和状态组合错误由后续 Pre-Tool/Projector 显式返回安全 `ERROR`。
|
||||
- MQ/Consumer/加密验签:本 change 不涉及。
|
||||
|
||||
### 新建类型
|
||||
|
||||
- `AgentToolContracts` 与 contract defensive-copy helper。
|
||||
- RAG Request/Result/Evidence。
|
||||
- Log Topic/SourceKind/Request/Result/Scope/Pattern/Event。
|
||||
- MySQL Request/Result。
|
||||
- 三个独立 contract test classes。
|
||||
|
||||
### 影响半径
|
||||
|
||||
- 新增包当前应无生产调用方;旧 Chat/AIOps、Tool、Controller、Repository 和配置文件均不修改。
|
||||
- 通过 focused compile/tests 和 `rg`/diff scope 证明边界。
|
||||
|
||||
## Apply 结果
|
||||
|
||||
- 冲突分类:未发现 OpenSpec 遗漏、代码偏离或方向不确定项。
|
||||
- 新增共享 ACI Tool 名称/描述、defensive-copy helper 和三类 typed Request/Result records。
|
||||
- 新增三个独立契约测试,覆盖精确 JSON、状态分离、框架 ID 原样保留、不可变集合、Mock provenance 和基础设施字段排除。
|
||||
- 首模块对齐:共享/RAG/日志/MySQL contract 与 design/tasks 全部完成;旧 runtime 接入保持 TODO,归属后续 3B/3C/6B changes。
|
||||
|
||||
## Apply 验证
|
||||
|
||||
- 编译:`mvn -q -DskipTests compile` 通过。
|
||||
- 新契约:`mvn -q '-Dtest=RagToolContractTest,QueryLogsToolContractTest,MysqlToolContractTest' test` 通过。
|
||||
- 旧行为回归:`mvn -q '-Dtest=HarnessContractTest,LookupKnowledgeToolTest,QueryLogsToolsTest' test` 通过。
|
||||
- 静态 scope:新 contract 生产类型当前无旧运行消费者;受保护的 Tool、Chat/AIOps、Controller、Repository 和配置文件 diff 为空。
|
||||
@@ -0,0 +1,27 @@
|
||||
# Evidence: single-react-aci-tool-contracts
|
||||
|
||||
## 文档证据
|
||||
|
||||
- ISS-014 5.2/5.3 冻结 `evidence_status` 与框架 `tool_call_id`;6-9 节冻结 RAG、日志和 MySQL Agent-facing Schema。
|
||||
- ISS-014 阶段 1 明确只冻结 ACI Contract,不实现 Agent 架构切换、真实 CLS/MCP 或 MySQL 执行。
|
||||
- 阶段 0 OpenSpec 与 devflow 已冻结 `InvocationStatus`、`EvidenceStatus`、框架 ID 真理源和阶段 6B 才公开切换。
|
||||
|
||||
## 代码证据
|
||||
|
||||
- `LookupKnowledgeTool` 仍返回包含 ContextPack/Trace 的旧 `LookupResult`,证明需要新的 bounded RAG Contract,也证明本阶段未提前切换。
|
||||
- `QueryLogsTools` 仍暴露 region/logTopic/limit、Topic discovery 和旧 mutable DTO,证明逻辑 Topic/Scope/Mock provenance 契约的必要性。
|
||||
- `ChatService` 与 `AiOpsService` 仍引用旧 `LookupKnowledgeTool`/`QueryLogsTools`,新 contract 生产包当前没有旧运行消费者。
|
||||
- 本地 Spring AI 1.1.7 `AssistantMessage.ToolCall` 提供 `id()`;Spring AI Alibaba 1.1.2.0 `ToolCallRequest` 提供 `getToolCallId()` 与 `ToolInterceptor` 边界。
|
||||
- Spring AI 1.1.7 普通 `ToolContext` 只传递调用方 context/history,不自动提供当前 Tool Call ID,因此后续必须从 Alibaba interceptor request 接入。
|
||||
|
||||
## Evidence-driven 结论
|
||||
|
||||
- lifecycle 与 evidence result 必须保持两套正交状态;已汇报并进入 OpenSpec/代码测试。
|
||||
- `tool_call_id` 可以从当前框架 API 取得,Harness 无需也不得生成第二套 ID;已汇报并进入 OpenSpec。
|
||||
- 阶段 1 不改旧运行签名/返回;已通过引用和 diff scope 验证。
|
||||
- 日志首版保留 Mock 数据源但必须显式 `source_kind=MOCK`,不实现真实适配器;已进入 Log Contract。
|
||||
- 三类 Contract 的可验证口径是精确 JSON、不可变结果、短描述和基础设施字段排除;三个独立测试均通过。
|
||||
|
||||
## 工具限制
|
||||
|
||||
- 当前会话未提供 `codebase-retrieval` 或 LSP;使用 `rg` 引用搜索、源码阅读、本地依赖 `javap`、Maven 编译和 focused tests 完成等价核对。
|
||||
@@ -0,0 +1,36 @@
|
||||
# Acceptance: single-react-design-freeze
|
||||
|
||||
## 实现结果
|
||||
|
||||
- 新增类型化 Harness contracts 和序列化契约测试。
|
||||
- ISS-014 已调整为 11 个串行 changes,并补齐 Tool ID、双状态、previous turn 和安全切换决策。
|
||||
- Python 查询脚本和 Spring 主配置改为环境变量 Secret。
|
||||
- 集成测试和历史 handoff 中的旧 Secret 副本已删除或脱敏。
|
||||
- devflow glossary 新增 Harness、EvidenceGuard、SemanticGuard、Invocation Status 和 Evidence Status。
|
||||
|
||||
## 静态验证
|
||||
|
||||
- OpenSpec strict validation:通过。
|
||||
- Secret 扫描:通过,已知真实 Secret 模式零匹配。
|
||||
- Diff whitespace 检查:通过。
|
||||
|
||||
## 脚本验证
|
||||
|
||||
- 4 个 contract tests:通过。
|
||||
- ToolInvocationRecorder、ExecutorGatekeeperService、ChatController focused baseline:通过。
|
||||
- 查询脚本缺失密码环境变量时 fail fast:通过。
|
||||
|
||||
## 浏览器/人工验证
|
||||
|
||||
- 不适用;阶段 0 未切换任何公开 UI/API。
|
||||
|
||||
## 未验证与后续门禁
|
||||
|
||||
- Provider 侧旧凭据轮换待凭据所有者完成。
|
||||
- Live E2E 留到阶段 7。
|
||||
- 阶段 1 只有在本 change Archive 和 Git commit 后才能开始。
|
||||
|
||||
## 状态
|
||||
|
||||
- Stage acceptance: accepted
|
||||
- OpenSpec archive: authorized by standing user instruction
|
||||
@@ -0,0 +1,31 @@
|
||||
# Brief: single-react-design-freeze
|
||||
|
||||
## 背景
|
||||
|
||||
ISS-014 将 Chat 诊断从多 Agent、Hook、ThreadLocal 和外层重试编排收敛为单体 Diagnosis ReAct Agent、确定性 Harness 和隔离 SemanticGuard。阶段 0 先冻结后续实施共同依赖的契约和安全边界。
|
||||
|
||||
## 目标
|
||||
|
||||
- 提供可复用的 Diagnosis Draft、Knowledge Answer、Fallback、published result 和 previous turn 类型。
|
||||
- 固定 framework Tool Call ID、调用生命周期和证据结果双状态。
|
||||
- 固定取消、重试、no-evidence、阶段门禁和安全凭据策略。
|
||||
- 保持当前公开 Chat 运行行为不变。
|
||||
|
||||
## 范围
|
||||
|
||||
- `com.superbiz.agent.harness.contract` 类型层和 focused tests。
|
||||
- ISS-014 的 11 个串行 OpenSpec changes 台账。
|
||||
- 主配置、查询脚本和集成测试中的明文 Secret 清理。
|
||||
- OpenSpec、devflow glossary 和阶段验收基线。
|
||||
|
||||
## 非目标
|
||||
|
||||
- 不接入新 Harness、Agent 或 Guards。
|
||||
- 不切换 `/api/chat`,不删除旧运行链。
|
||||
- 不执行 live E2E。
|
||||
|
||||
## 元数据
|
||||
|
||||
- Scale: complex
|
||||
- Parent issue: `ISS-014`
|
||||
- OpenSpec: `openspec/changes/single-react-design-freeze`
|
||||
@@ -0,0 +1,20 @@
|
||||
# Decisions: single-react-design-freeze
|
||||
|
||||
## 已确认决策
|
||||
|
||||
- ISS-014 保留为总 Issue,实施拆成 11 个独立、串行 sm-flow changes。
|
||||
- 每个 change 必须完成 Apply、阶段验收、OpenSpec Archive 和 Git commit 后才能进入下一项。
|
||||
- Apply、Archive 和阶段 Git commit 已获得用户对整个目标的持续授权;仅方向性决策需要暂停。
|
||||
- `tool_call_id` 使用框架协议 ID,Harness 不生成第二套 ID。
|
||||
- `status=PROJECTING/READY/ERROR` 表示 invocation lifecycle;`evidence_status=EVIDENCE_FOUND/NO_EVIDENCE/ERROR` 表示结果语义。
|
||||
- `NO_EVIDENCE` 只能支持限定范围的 `NEGATIVE_OBSERVATION`。
|
||||
- KNOWLEDGE_QUERY 使用 answer items 绑定 `tool_call_id + document_id`,首版不进入 SemanticGuard。
|
||||
- `diagnosis_run` 是 previous turn 的持久化真理源;只有最近的 `DIAGNOSIS + SUCCESS` 安全发布结果可用。
|
||||
- 取消采用分层、可观测语义;同步模型调用不承诺无法证明的立即硬取消。
|
||||
- 底层重试压为一次 attempt,Router/SemanticGuard 的允许重试只由 Harness 执行。
|
||||
- 阶段 4 和 6A 不接管公开入口,阶段 6B 才执行原子 SSE 切换。
|
||||
|
||||
## 权衡
|
||||
|
||||
- contract types 提前落地会增加少量文件,但后续 output type、validator、持久化和 SSE 可以复用,避免多阶段字符串协议漂移。
|
||||
- Provider 侧凭据轮换是外部前置,阶段 0 只证明工作树已清理,不伪造外部完成状态。
|
||||
@@ -0,0 +1,25 @@
|
||||
# Evidence: single-react-design-freeze
|
||||
|
||||
## 代码与框架证据
|
||||
|
||||
- `ChatController` 当前直接管理会话、模型、工具、执行和伪流式 SSE,证明公开切换必须延后到阶段 6B。
|
||||
- `ChatService` 当前构建 Planner/Executor/Verifier/Composer 并依赖 ThreadLocal,证明阶段 0 只能创建无运行依赖的 contract types。
|
||||
- Spring AI Alibaba `Builder` 提供 ToolInterceptor、output type、tool timeout 和调用限制扩展点;`ReactAgent` 提供 interrupt。
|
||||
- Spring AI 1.1.7 `SpringAiRetryProperties` 构造器默认 `maxAttempts=10`,与 Harness 单一重试所有权冲突,阶段 2 必须压为一次底层 attempt。
|
||||
- `DiagnosisRun` 当前只有通用 status 和文本 answer,不足以确定性恢复安全 previous turn,因此确认后续保存 `intent/release_outcome/published_result`。
|
||||
- 历史 ISS-009 和 verifier evidence 归档证明 `NO_EVIDENCE` 只能表达限定范围内无匹配结果。
|
||||
|
||||
## 验证证据
|
||||
|
||||
- `mvn -q '-Dtest=HarnessContractTest' test`:通过。
|
||||
- `mvn -q '-Dtest=ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,ChatControllerTest' test`:通过。
|
||||
- `mvn -q '-Dtest=HarnessContractTest,ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,ChatControllerTest' test`:通过。
|
||||
- `openspec validate single-react-design-freeze --strict`:通过。
|
||||
- `git diff --check`:通过,仅有 Windows line-ending warning。
|
||||
- 已知旧 Secret 模式扫描:零匹配。
|
||||
- 缺少 `SUPERBIZ_MYSQL_PASSWORD` 时运行查询脚本:在连接前 fail fast,符合预期。
|
||||
|
||||
## 未验证
|
||||
|
||||
- 未运行模型、Redis、MySQL 或 Milvus live E2E;按阶段门禁留到阶段 7。
|
||||
- Provider 侧旧凭据是否已经轮换无法由仓库证明,必须由凭据所有者完成,并在阶段 3C/7 前核验。
|
||||
@@ -0,0 +1,65 @@
|
||||
# Single React Diagnosis Agent Acceptance
|
||||
|
||||
## 结果
|
||||
|
||||
已接受。阶段 4 Committed OpenSpec 的 11 项任务全部完成,新链路仅通过内部 Java/test 入口运行,公开 Chat 未切换。
|
||||
|
||||
## 验证
|
||||
|
||||
### 静态验证
|
||||
|
||||
- 检查:新 `harness.agent` 包与 Prompt 搜索 `ThreadLocal|SessionContextHolder|while|SequentialAgent|SupervisorAgent|StateGraph|Planner|Composer|Verifier|raw_response`。
|
||||
- 结果:通过,无匹配。
|
||||
- 检查:`git diff --name-only` 限定 Controller、ChatService、AiOpsService。
|
||||
- 结果:通过,无 diff。
|
||||
- 检查:OpenSpec artifact status、cross-artifact、L2 interface impact、`.committed`、`.archive-ready`。
|
||||
- 结果:通过。
|
||||
|
||||
### 脚本验证
|
||||
|
||||
- 命令:`mvn -q -DskipTests compile`
|
||||
- 结果:通过。
|
||||
- 命令:`mvn -q '-Dtest=HarnessToolInterceptorTest,DiagnosisAgentUseCaseTest' test`
|
||||
- 结果:通过,14 tests。
|
||||
- 命令:`mvn -q '-Dtest=DiagnosisAgentUseCaseTest,HarnessToolInterceptorTest,DiagnosisHarnessCoreTest,ToolBoundaryTest,ToolAdapterTest,MysqlToolAdapterTest,CanonicalInvocationStoreTest,HarnessContractTest,RagToolContractTest,QueryLogsToolContractTest,MysqlToolContractTest' test`
|
||||
- 结果:通过。
|
||||
- 命令:`openspec validate single-react-diagnosis-agent --strict`
|
||||
- 结果:通过。
|
||||
- 命令:归档后 `openspec validate single-react-diagnosis-agent --type spec --strict`
|
||||
- 结果:通过;仓库级 `openspec validate --specs --strict` 另发现前序 `mysql-readonly-tool`、`rag-log-projections` 主规格仍使用旧 scenario 格式,本阶段不跨归档边界修改。
|
||||
|
||||
### 浏览器/人工验证
|
||||
|
||||
- 不适用。本阶段无 UI、HTTP 或 SSE 行为变化。
|
||||
|
||||
### 未验证
|
||||
|
||||
- 未启动 Maven 应用、外部模型、MySQL、Redis、Milvus 或日志服务;按 ISS-014 门禁统一留到阶段 7 live E2E。
|
||||
- 未验证 provider 生产响应始终携带 Token Usage;缺失 Usage 时模型调用次数仍强制,实际 Token 只能在 provider 返回时记录。
|
||||
|
||||
## 已完成范围
|
||||
|
||||
- 单一 Diagnosis ReactAgent、单 Prompt 和固定 Query/PreviousTurn 输入。
|
||||
- 框架原生 tool loop、exact Tool Call ID 和三类 Harness adapter bridge。
|
||||
- 模型/Token/Tool/上下文/Draft 预算与无自动 retry。
|
||||
- 严格 DiagnosisDraft 输出和无证据停止规则。
|
||||
- 显式 sessionId/runId audit metadata、可注入 AgentStep Hook、canonical Tool invocation。
|
||||
- 公开 Chat/AiOps 与旧多 Agent 路径保持不变。
|
||||
|
||||
## 已知限制
|
||||
|
||||
- Draft 尚未经过 EvidenceGuard/SemanticGuard,不能公开发布;阶段 5 负责释放门禁。
|
||||
- previous turn 由调用方传入,阶段 6A 才实现同 Session 安全 Run 选择和确定性组装。
|
||||
- 真实业务 adapter/Spring Bean 装配和公开入口切换分别留给阶段 6A/6B。
|
||||
- 仓库级全规格 strict validation 仍受阶段 3B/3C 主规格 scenario 标题格式阻塞,建议阶段 7 文档清理时统一修正并复验。
|
||||
|
||||
## Bug 修复和诊断
|
||||
|
||||
- 编译发现框架 `Interceptor.getName()` 必须实现,已补充稳定名称并回归。
|
||||
- 首次测试修正了两个测试假设:callback 异常包装类型,以及 ToolResponseMessage 在 Prompt instructions 中的实际位置;生产规格与实现方向未变。
|
||||
- 提交前自审发现默认 BeanOutputConverter schema 不允许 `conclusion=null`,与冻结的无证据契约冲突;已增加 schema post-process 和 focused regression,未放宽其他 Draft 字段。
|
||||
|
||||
## 交接
|
||||
|
||||
- 下一步:提交独立阶段 4 commit,再进入阶段 5。
|
||||
- OpenSpec 归档确认:用户已对 ISS-014 每阶段 Apply、Archive 和 Git commit 提供持续授权;已归档至 `openspec/changes/archive/2026-07-21-single-react-diagnosis-agent`,主规格已同步至 `openspec/specs/single-react-diagnosis-agent/spec.md`。
|
||||
@@ -0,0 +1,21 @@
|
||||
# Single React Diagnosis Agent Brief
|
||||
|
||||
## 背景
|
||||
|
||||
- 用户目标:按 ISS-014 阶段 4 建立唯一的 Diagnosis ReAct Agent,并完成独立 sm-flow 归档与提交。
|
||||
- 当前问题:Harness Core 和三类 evidence Tool 已就绪,但没有一个内部诊断执行链消费它们;公开 Chat 仍依赖旧的简单/多 Agent 路径。
|
||||
- 关联 OpenSpec:`openspec/changes/archive/2026-07-21-single-react-diagnosis-agent/`
|
||||
- devflow 分档:complex
|
||||
- 需求真理源:`mvp/issues/active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md`,不重复创建独立 PRD。
|
||||
|
||||
## 范围
|
||||
|
||||
- 本次要做:单 Prompt、单 Diagnosis ReactAgent、固定 Query/PreviousTurn 输入、Harness Model/Tool interceptors、三类 Tool 注册、严格 DiagnosisDraft 解析、上下文/模型/Tool/Token 预算和内部审计注入。
|
||||
- 本次不做:公开 Chat 切换、旧链路删除、EvidenceGuard、SemanticGuard、previous turn 选择、DiagnosisRun 持久化和 live E2E。
|
||||
- 影响区域:`com.superbiz.agent.harness.agent`、`src/main/resources/prompts`、focused tests、OpenSpec/devflow/ISS-014 状态。
|
||||
|
||||
## OpenSpec 对齐
|
||||
|
||||
- proposal 覆盖状态:已覆盖
|
||||
- specs 覆盖状态:已覆盖
|
||||
- tasks 覆盖状态:已覆盖
|
||||
@@ -0,0 +1,156 @@
|
||||
# Decisions: single-react-diagnosis-agent
|
||||
|
||||
## Discover status
|
||||
|
||||
- Checkpoint: Discover
|
||||
- Capability source: `sm-flow`,使用 `grill-with-docs` 进行代码可证问题澄清,并使用 GitNexus、源码、依赖 sources jar 与 `javap` 核对调用链和框架 API。
|
||||
- Scale: complex。变更新增内部 Agent/use case,跨越模型、Tool、预算、结构化输出和审计边界,但不切换公开协议。
|
||||
- `devflow/index.md` 命中 design freeze、RunContext、ACI contracts、canonical store、RAG/log projection 和 MySQL Tool;没有与 OpenSpec 冲突的 ADR。
|
||||
|
||||
## Question Pool
|
||||
|
||||
| # | 维度 | 问题 | 模式 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | 术语 | 单体 Diagnosis Agent 是否包含外层 Graph 或多个报告作者? | evidence-driven | 已解决 |
|
||||
| Q2 | 边界 | 阶段 4 是否切换公开 Chat 或删除旧多 Agent 链路? | evidence-driven | 已解决 |
|
||||
| Q3 | 验收 | 如何证明 ReAct 正常轮次不是 Harness retry? | evidence-driven | 已解决 |
|
||||
| Q4 | 技术 | 如何取得并保留框架原始 tool_call_id? | evidence-driven | 已解决 |
|
||||
| Q5 | 技术 | 如何在每个模型和 Tool 边界强制 Run 预算? | evidence-driven | 已解决 |
|
||||
| Q6 | 输出 | 框架生成的 DiagnosisDraft schema 是否与可空 conclusion 契约一致,并会直接返回 DiagnosisDraft? | evidence-driven | 已解决 |
|
||||
| Q7 | 上下文 | previous_turn 如何进入模型且保持固定 Schema 和有界? | evidence-driven | 已解决 |
|
||||
| Q8 | 审计 | 如何保留 Run/AgentStep/ToolInvocation 而不引入 ThreadLocal? | evidence-driven | 已解决 |
|
||||
| Q9 | 验收 | 无证据结果如何停止而不补造根因? | evidence-driven | 已解决 |
|
||||
|
||||
## Evidence-driven
|
||||
|
||||
| 结论 | 证据来源 | 是否已汇报用户 |
|
||||
|---|---|---|
|
||||
| 只新增一个拥有 tool loop 的 `ReactAgent`,无 SequentialAgent、SupervisorAgent 或业务 StateGraph。 | ISS-014 阶段 4;`ChatService` 旧链路反例;Spring AI Alibaba `ReactAgent` API | 已汇报 |
|
||||
| 阶段 4 只提供内部用例,公开入口保持旧实现,接口影响 L2。 | ISS-014 1174-1197、阶段 0 design freeze、GitNexus `createReactAgent` 引用 | 已汇报 |
|
||||
| `ToolCallRequest.getToolCallId()` 精确暴露 `AssistantMessage.ToolCall.id()`;ToolInterceptor 可在回调执行前取得该 ID。 | framework sources `ToolCallRequest`、`AgentToolNode` | 已汇报 |
|
||||
| `ModelInterceptor` 包围每次真实模型调用,适合调用 `beforeModelCall` 并记录响应 Usage;ToolBoundary 已负责 Tool 预算,不能重复计数。 | framework `AgentLlmNode`/`InterceptorChain`;`ToolBoundary` | 已汇报 |
|
||||
| BeanOutputConverter 可从 DiagnosisDraft 生成格式说明,但默认 schema 不允许冻结契约的 `conclusion=null`;必须 post-process nullable conclusion,且 `ReactAgent.call` 仍返回 AssistantMessage,需要 Harness 严格解析 JSON。 | framework `DefaultBuilder`、`ReactAgent`、本地生成 schema | 已汇报 |
|
||||
| RunContext 可放入 `RunnableConfig` metadata,由框架控制地传播到 ToolInterceptor;不需要 ThreadLocal。 | `RunnableConfig.addMetadata(String,Object)`、`AgentToolNode` | 已汇报 |
|
||||
| 现有 `AgentLoggingHook` 优先读取 config metadata 的 sessionId/runId,可作为可选审计 Hook 复用;ToolBoundary 保持 canonical Tool invocation。 | `AgentLoggingHook`、`ToolBoundary` | 已汇报 |
|
||||
| 原始 Query 不截断;超限 fail closed。PreviousTurn 先按固定 record 序列化并受独立/总上下文字节限制。 | ISS-014 极简上下文与预算规则 | 已汇报 |
|
||||
| `NO_EVIDENCE` 不是错误,只能形成限定范围的 NEGATIVE_OBSERVATION;Prompt 必须要求 conclusion=null、记录 missing_info 并停止。 | glossary、ACI spec、DiagnosisDraft contract | 已汇报 |
|
||||
|
||||
## User-interview
|
||||
|
||||
- 本阶段没有新增 user-interview 问题。方向、范围、阶段串行规则、框架 ID、状态语义以及 routine Apply/Archive/Commit 持续授权均已由用户在 ISS-014 评审与前序阶段确认。
|
||||
|
||||
## 关键取舍
|
||||
|
||||
- 决策:使用框架 `ToolInterceptor` 直接桥接已注册的 Tool 定义与 Harness adapter。
|
||||
- 原因:Spring AI `ToolCallback` 的调用参数不包含 Tool Call ID,而 Alibaba interceptor 明确提供原始 ID 和运行 metadata。
|
||||
- 影响:ToolCallback 负责模型可见定义,interceptor 负责受控执行;未知 Tool 仍交给框架 handler 并最终失败,不伪造结果。
|
||||
- 决策:模型预算由 `ModelInterceptor` 执行,Tool 预算继续由 `ToolBoundary` 执行。
|
||||
- 原因:避免在 Agent 层和 Tool boundary 双重 reserve。
|
||||
- 决策:阶段 4 只严格反序列化 `DiagnosisDraft`,不提前实现 EvidenceGuard。
|
||||
- 原因:字段引用真实性、唯一性和语义支持属于阶段 5;本阶段只验证 Agent 能生成冻结结构并携带框架 IDs。
|
||||
- 决策:复用现有 Agent Hook 注入点,不复制 AgentStep 持久化实现。
|
||||
- 原因:阶段 4 不接公开运行态,阶段 6A 再装配真实 Repository 和 Run 持久化。
|
||||
|
||||
## OpenSpec 回写
|
||||
|
||||
- 需进入 proposal/design/spec/tasks:单 Agent、内部入口、ToolInterceptor 原始 ID、ModelInterceptor 模型预算、ToolBoundary 单点 Tool 预算、严格 JSON 解析、上下文字节限制、审计 Hook 注入、无证据停止、不切换公开入口。
|
||||
- 不创建 ADR:这些是 ISS-014 已冻结方向和当前阶段可逆的内部装配,不满足新的难逆转架构决策条件。
|
||||
|
||||
## Cross-artifact 对齐
|
||||
|
||||
| 上游 -> 下游 | 检查内容 | 状态 |
|
||||
|---|---|---|
|
||||
| brief/ISS-014 -> proposal | 单 Agent、内部入口、极简上下文、预算、审计、非目标和验收预期 | 已对齐 |
|
||||
| proposal -> design | Tool/Model interceptor、严格 Draft、无重试、审计注入和公开隔离 | 已对齐 |
|
||||
| design -> specs/tasks | ID 传播、预算单点、输入输出限制、生命周期所有权和风险缓解 | 已对齐 |
|
||||
| specs -> tasks | 7 组可观察要求均有输入/Prompt、Tool bridge、Agent use case 和 focused test 切片 | 已对齐 |
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
- 能力来源:`zoom-out`,使用 glossary 的 Diagnosis Agent、Diagnosis Harness、RunContext、Evidence Status 和 Invocation Status 术语。
|
||||
- 链路为 `DiagnosisAgentInput + RunContext -> internal use case -> one ReactAgent -> model/tool interceptors -> ChatModel/ToolBoundary -> strict DiagnosisDraft`,没有外层业务 Graph。
|
||||
- RunContext 拥有预算/取消/生命周期,ToolBoundary/store 拥有 canonical invocation,use case 拥有输入和 Draft 解析,Agent 只拥有诊断语义;数据所有权不重叠。
|
||||
- 最大耦合风险是 Spring AI Alibaba interceptor/output schema API;设计通过本地 1.1.2.0 sources jar 和实际 schema 生成结果证实,并以 focused framework-loop/schema tests 固定。
|
||||
- 最大阶段风险是 Draft 在 Guards 前被误用;类名、文档和零 Controller 消费者共同保持 internal/draft 边界,阶段 5 前不得发布。
|
||||
|
||||
## 接口影响
|
||||
|
||||
- 级别:L2 内部接口。
|
||||
- 变更对象:新增内部 Java use case/factory/input/limits/interceptors/registry,不修改既有方法签名。
|
||||
- 消费者:本阶段只有 focused tests;阶段 6A 将成为首个生产消费者。
|
||||
- 兼容性:公开 HTTP/SSE、Controller DTO、数据库和旧 Chat/AiOps 路径不变,无迁移或回滚要求。
|
||||
|
||||
## Pre-apply Research
|
||||
|
||||
### 参考实现与框架源码
|
||||
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`:`createReactAgent`、`buildChatExecutorAgent` 和旧多 Agent 反例;只复用 Builder 形态,不复用运行职责。
|
||||
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`:scripted `ChatModel` 测试模式。
|
||||
- `src/main/java/com/superbiz/agent/harness/core/DiagnosisHarnessCore.java`:模型/Tool/Token/容量边界。
|
||||
- `src/main/java/com/superbiz/agent/harness/tool/boundary/ToolBoundary.java`:Tool budget 和 canonical record 单点。
|
||||
- `src/main/java/com/superbiz/agent/harness/tool/adapter/RagToolAdapter.java`、`QueryLogsToolAdapter.java`、`MysqlToolAdapter.java`:阶段 4 唯一允许的 evidence Tool 执行入口。
|
||||
- `src/main/java/com/superbiz/agent/hook/AgentLoggingHook.java`:可注入 AgentStep audit,优先读取 RunnableConfig metadata。
|
||||
- Spring AI Alibaba 1.1.2.0 sources:`ReactAgent`、`DefaultBuilder`、`AgentLlmNode`、`AgentToolNode`、`ToolCallRequest`、`InterceptorChain`。
|
||||
|
||||
### 技术栈清单
|
||||
|
||||
- `ReactAgent.builder()` + 从 `DiagnosisDraft` 生成并修正 nullable conclusion 的 `.outputSchema(...)`,不新建 Graph/SequentialAgent。
|
||||
- `RunnableConfig.addMetadata(String,Object)` 显式携带 sessionId/runId/RunContext,并设置 `_stream_=false`。
|
||||
- `ModelInterceptor` 对每次模型调用执行 Core budget;读取 `ChatResponseMetadata.Usage`。
|
||||
- `ToolInterceptor` 读取 exact framework ID,调用 adapter bridge;`ToolBoundary` 保留唯一 Tool reserve/store 边界。
|
||||
- Spring `FunctionToolCallback` 仅定义模型可见 Tool Schema/description;直接 callback 执行 fail closed。
|
||||
- Jackson `ObjectMapper` 负责固定输入 JSON 和严格 DiagnosisDraft 反序列化;UTF-8 字节按 `StandardCharsets.UTF_8` 计算。
|
||||
- 框架 Hook 列表作为 AgentStep audit 扩展点;阶段 4 不新增 JPA/Redis/Flyway。
|
||||
|
||||
### 新建基础设施
|
||||
|
||||
- `harness.agent`:Input、Limits、Tool registry/interceptor、Model interceptor、Factory、UseCase 和异常类型。
|
||||
- `prompts/diagnosis-agent-prompt.md`:唯一 Diagnosis Agent Prompt。
|
||||
- focused scripted model/tool-loop tests;无需新 Maven 依赖。
|
||||
|
||||
### ISS-014 PRD 复用
|
||||
|
||||
- 阶段 4 不新建独立 `prd.md`:ISS-014 已完整覆盖问题、用户价值、数据流、Draft Schema、预算、重试、阶段边界和验收;`brief.md` 只索引本切片,不复制总 Issue。
|
||||
|
||||
## Commit Gate Preflight
|
||||
|
||||
- proposal、design、specs、tasks 完整,`openspec status` 为 complete,`openspec validate single-react-diagnosis-agent --strict` 通过。
|
||||
- question pool 全部为已汇报的 evidence-driven 结论;无未确认 user-interview、接口等级或风险接受问题。
|
||||
- Cross-artifact 四段对齐无 gap;架构审计的数据所有权、阶段边界和框架耦合缓解已进入 design/spec/tasks。
|
||||
- 接口影响为 L2,仅新增内部 Java API;公开 Controller/SSE/JPA/旧 ChatService 保持不变。
|
||||
- Apply、Archive 和阶段 Git commit 使用用户对 ISS-014 各阶段的持续授权;实现必须严格限制为 Committed OpenSpec。
|
||||
- `.committed` 已创建,Committed OpenSpec 可进入 Apply。
|
||||
|
||||
## Apply Progress
|
||||
|
||||
- 首模块对齐:Input/Limits/Prompt/Tool registry/Tool interceptor 已落地,任务 1.2、2.1、2.2 完成;1.1 等待 focused test 后完成。
|
||||
- 编译首次发现 `ToolInterceptor` 必须实现 `Interceptor.getName()`;分类为代码偏离,已补充稳定名称并复编译通过,无需修改 OpenSpec。
|
||||
- 单 Agent use case、Model interceptor、strict Draft parser 和 audit Hook 注入已完成;scripted ChatModel 真实执行框架 `model -> Tool -> model` loop。
|
||||
- 首次测试的两处失败均为测试假设偏差:Spring 会将 callback 异常包装为 `ToolExecutionException`,Tool observation 位于 `Prompt.instructions` 的 `ToolResponseMessage` 而非 `Prompt.getContents()`;已按框架真实 API 修正测试,生产设计未变。
|
||||
- 阶段 4 focused tests 共 14 个通过,任务 1.1-4.2 完成;剩余任务 4.3 为综合回归、静态范围和 OpenSpec 验证。
|
||||
|
||||
## Apply Result
|
||||
|
||||
- 新增一个且仅一个 `diagnosis_agent` Factory 和内部 `DiagnosisAgentUseCase`;没有外层业务 Graph、SequentialAgent、SupervisorAgent 或手写 ReAct loop。
|
||||
- 新增单一 Prompt、固定 `DiagnosisAgentInput(query, previous_turn)`、可配置 UTF-8 限制和严格 `DiagnosisDraft` 解析;不接受完整历史,不执行结构修复或 Agent retry。
|
||||
- 新增 `HarnessModelInterceptor`,对每个非流式模型轮次执行 Core model/Token budget,并在 late result 返回后复查 Run active 状态。
|
||||
- 新增 `HarnessEvidenceTools` 和 `HarnessToolInterceptor`,注册三类冻结 Tool,精确传播框架 Tool Call ID,并通过阶段 3B/3C adapter/ToolBoundary 返回有界结果。
|
||||
- AgentStep 使用可注入框架 Hook 保留,RunnableConfig 显式传播 sessionId/runId;Run 和 Tool canonical invocation 继续由既有 Harness 边界拥有。
|
||||
- 公开 Controller、ChatService、AiOpsService 无 diff;旧多 Agent 主链路继续保留。
|
||||
|
||||
## Apply Verification
|
||||
|
||||
- 编译:`mvn -q -DskipTests compile` 通过。
|
||||
- 阶段 4 focused:`mvn -q '-Dtest=HarnessToolInterceptorTest,DiagnosisAgentUseCaseTest' test`,14 tests 通过。
|
||||
- 综合回归:`mvn -q '-Dtest=DiagnosisAgentUseCaseTest,HarnessToolInterceptorTest,DiagnosisHarnessCoreTest,ToolBoundaryTest,ToolAdapterTest,MysqlToolAdapterTest,CanonicalInvocationStoreTest,HarnessContractTest,RagToolContractTest,QueryLogsToolContractTest,MysqlToolContractTest' test` 通过。
|
||||
- OpenSpec:`openspec validate single-react-diagnosis-agent --strict` 通过。
|
||||
- 静态范围:新 Agent 包无 ThreadLocal、手写 while、外层 Graph、多 Agent 类型或 raw response;Controller/ChatService/AiOpsService diff 为空。
|
||||
- 未执行 live E2E:按 ISS-014 阶段门禁统一留到阶段 7。
|
||||
|
||||
## Archive Result
|
||||
|
||||
- 11/11 OpenSpec tasks 完成,`.archive-ready` 已创建。
|
||||
- 主规格已同步至 `openspec/specs/single-react-diagnosis-agent/spec.md`。
|
||||
- Change 已归档至 `openspec/changes/archive/2026-07-21-single-react-diagnosis-agent`。
|
||||
- `devflow/index.md` 和 ISS-014 阶段表已更新为阶段 0-4 archived,下一阶段为 5。
|
||||
- 提交前 schema 自审发现框架默认 BeanOutputConverter 将 `conclusion` 限为 object,与冻结的无证据 `null` 语义冲突;分类为实现设计遗漏,已回写 archived design,新增 nullable schema post-process 与回归测试,规格行为未改变。
|
||||
@@ -0,0 +1,29 @@
|
||||
# Single React Diagnosis Agent Evidence
|
||||
|
||||
## 代码与文档证据
|
||||
|
||||
| 来源 | 证据 | 结论 | 是否已汇报 |
|
||||
|---|---|---|---|
|
||||
| `mvp/issues/active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md` | 阶段 4 明确单 Agent、内部入口、极简上下文、结构化 Draft、无重试和预算验收 | 本 change 不得切换公开入口或删除旧链路 | 是 |
|
||||
| `ChatService.java` + GitNexus references | `createReactAgent` 被旧策略入口调用;复杂路径仍创建 Planner/Executor/Verifier/Composer | 新实现必须是独立内部 use case,不能复用旧 Service 运行职责 | 是 |
|
||||
| Spring AI Alibaba 1.1.2.0 `ReactAgent`/`AgentLlmNode` sources | 框架自带 ReAct loop;非流式 `ModelResponse` 保留 `ChatResponse` Usage | 不手写循环,使用 ModelInterceptor 强制模型/Token 预算 | 是 |
|
||||
| Spring AI Alibaba `ToolCallRequest`/`AgentToolNode` sources | `ToolCallRequest.getToolCallId()` 来自 `AssistantMessage.ToolCall.id()`,Tool interceptor 在 callback 前执行 | 可以精确传播框架 ID,不生成第二套 ID | 是 |
|
||||
| Spring AI Alibaba `DefaultBuilder` + 本地 BeanOutputConverter 输出 | 默认 schema 只允许 object conclusion,但冻结契约允许无证据时 conclusion=null | 生成 schema 必须 post-process nullable conclusion,UseCase 仍严格反序列化 DiagnosisDraft | 是 |
|
||||
| `DiagnosisHarnessCore.java` / `ToolBoundary.java` | Core 提供 model/Token/capacity;ToolBoundary 已负责 Tool reserve/store | Agent 层只 reserve model/context,禁止双扣 Tool budget | 是 |
|
||||
| 阶段 3B/3C adapters | RAG/log/MySQL 都以 `RunContext + ToolCallRequestEnvelope` 进入 ToolBoundary | Tool registry 可直接桥接既有 adapter,不复制投影或安全策略 | 是 |
|
||||
| `AgentLoggingHook.java` | 优先从 RunnableConfig metadata 读取 sessionId/runId | Factory 可复用现有 Hook 注入点,不依赖 ThreadLocal | 是 |
|
||||
|
||||
## Evidence-driven 结论
|
||||
|
||||
- 单 Agent 的可验证边界是一个 `ReactAgent.call`,内部允许正常模型/Tool轮次,但外部没有 Agent retry 或第二个报告作者。
|
||||
- ToolCallback 不直接提供 Tool Call ID,必须使用 Alibaba `ToolInterceptor`;已由真实 scripted framework loop 证明 ID 进入 ToolResponseMessage。
|
||||
- `PreviousTurn` 只作为固定上下文,当前 Draft 只允许引用当前 Run Tool result 中的 ID。
|
||||
- `NO_EVIDENCE` 必须以 `conclusion=null`、限定范围的负向观察和 missing info 结束,不能推导健康或排除根因。
|
||||
- 阶段 4 的 L2 内部 API 不影响公开消费者;阶段 5/6A/6B 前 Draft 不可发布。
|
||||
|
||||
## 实现证据
|
||||
|
||||
- 新包 `com.superbiz.agent.harness.agent` 包含 Input/Limits、Prompt loader、Tool registry/interceptor、Model interceptor、Factory 和 UseCase。
|
||||
- `diagnosis-agent-prompt.md` 是唯一新 Diagnosis Agent Prompt,包含 Tool、证据绑定、停止和精确 JSON 规则。
|
||||
- 14 个阶段 4 tests 覆盖框架 model -> Tool -> model、ID、错误、未知 Tool、PreviousTurn、审计 metadata、nullable conclusion schema、no-evidence、解析和预算。
|
||||
- 综合回归同时覆盖 Harness Core、ToolBoundary、3B/3C adapter、canonical store 和冻结 contracts。
|
||||
@@ -0,0 +1,44 @@
|
||||
# Acceptance: single-react-evidence-semantic-guards
|
||||
|
||||
## Result
|
||||
|
||||
- Status: archived
|
||||
- OpenSpec tasks: 14/14 complete
|
||||
- Interface impact: L2 internal
|
||||
- Public protocol: unchanged
|
||||
|
||||
## Static Verification
|
||||
|
||||
- `git diff --check`:阶段文件无 whitespace error;工作区用户已有 `AGENTS.md`/`CLAUDE.md` 仅报告 line-ending warning,未纳入阶段范围。
|
||||
- 公开 `controller`、`ChatService.java`、`AiOpsService.java` diff 为空。
|
||||
- SemanticGuard 包和 prompt 不含 Tool Call ID、raw response、ReactAgent、StateGraph、ThreadLocal 或手写 while。
|
||||
- 新 guard/release 包不包含业务 Agent loop。
|
||||
|
||||
## Script Verification
|
||||
|
||||
- `mvn -q -DskipTests compile`:通过。
|
||||
- Stage-focused `EvidenceGuardTest,SemanticGuardTest,DiagnosisReleaseUseCaseTest`:19 tests,0 failure/error。
|
||||
- Harness/Tool/Agent regression selection:18 suites / 70 tests,0 failure/error/skipped。
|
||||
- `openspec validate single-react-evidence-semantic-guards --strict`:通过。
|
||||
- `openspec validate --specs --strict`:16 passed / 2 failed;失败是前序 `mysql-readonly-tool`、`rag-log-projections` scenario 标题格式,不阻塞本 change,计划阶段 7 统一修复。
|
||||
|
||||
## Browser or Manual Verification
|
||||
|
||||
- Not applicable。阶段 5 没有 UI 或公开入口变化。
|
||||
|
||||
## Not Verified
|
||||
|
||||
- 未运行真实 LLM、Redis、日志和 MySQL live E2E;按 ISS-014 串行门禁统一留到阶段 7。
|
||||
- Provider 是否立即响应 Java thread interrupt 取决于 SDK;Harness 已保证 Future cancel、迟到 Token 记账和迟到结果不释放,阶段 7 需用真实模型观察取消延迟。
|
||||
|
||||
## Remaining Work
|
||||
|
||||
- 阶段 6A:Chat Application Use Case、Intent Router、PreviousTurn 和 Run persistence。
|
||||
- 阶段 6B:公开 SSE 原子切换。
|
||||
- 阶段 7:旧链路清理、全局 spec 格式修复和最终 live E2E。
|
||||
|
||||
## Archive
|
||||
|
||||
- `.archive-ready`: created
|
||||
- OpenSpec archive: `openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards`
|
||||
- Main spec: `openspec/specs/single-react-evidence-semantic-guards/spec.md`
|
||||
@@ -0,0 +1,32 @@
|
||||
# Brief: single-react-evidence-semantic-guards
|
||||
|
||||
## Background
|
||||
|
||||
阶段 4 已能生成带框架 Tool Call ID 的 `DiagnosisDraft`,但草稿在发布前还缺少当前 Run 证据验真、独立语义审查和 fail-closed 释放边界。
|
||||
|
||||
## Goal
|
||||
|
||||
在不恢复业务 Graph、不切换公开入口的前提下,实现确定性 EvidenceGuard、一次语义不变的结构修复、无 Tool/无记忆的单轮 SemanticGuard,以及只发布原 Draft 或固定 SafeFallback 的内部 release use case。
|
||||
|
||||
## Scope
|
||||
|
||||
- Draft 结构、Analysis/报告引用和 canonical Tool invocation 验真。
|
||||
- RAG/log/MySQL verified evidence snapshot。
|
||||
- Evidence repair 单次调用与语义不变约束。
|
||||
- SemanticGuard 模型/Token/字节预算、超时、取消、技术重试和严格二元输出。
|
||||
- SUPPORTED/UNSUPPORTED/unavailable/evidence-failed release policy。
|
||||
- focused tests、Harness/Tool/Agent 回归和静态范围验证。
|
||||
|
||||
## Non-goals
|
||||
|
||||
- 不切换 Chat/AiOps/SSE 公开协议。
|
||||
- 不实现 Intent Router、PreviousTurn、Run persistence 或最终报告渲染。
|
||||
- 不删除旧 Gatekeeper/Verifier/Composer/多 Agent 链路。
|
||||
- 不运行 live model/Redis/MySQL E2E;统一留到阶段 7。
|
||||
|
||||
## Metadata
|
||||
|
||||
- Scale: complex
|
||||
- Interface impact: L2 internal
|
||||
- OpenSpec: `single-react-evidence-semantic-guards`
|
||||
- Parent issue: `ISS-014`
|
||||
@@ -0,0 +1,144 @@
|
||||
# Decisions: single-react-evidence-semantic-guards
|
||||
|
||||
## Discover Status
|
||||
|
||||
- Checkpoint: Discover
|
||||
- Capability source: `sm-flow`,使用 `grill-with-docs` 做 evidence-driven 澄清;`codebase-retrieval`、LSP 和 GitNexus MCP 在当前会话不可用,降级为既有 GitNexus 研究结论、`rg` 调用点核对和逐文件源码阅读。
|
||||
- Scale: complex。变更跨 canonical store、三类 Tool projection、模型预算、超时/取消、结构修复、语义审查和释放边界,但不切换公开协议。
|
||||
- `devflow/index.md` 命中阶段 0-4 的设计冻结、RunContext、canonical store、Tool projection 和 Diagnosis Agent;没有与当前 OpenSpec 冲突的 ADR。
|
||||
|
||||
## Question Pool
|
||||
|
||||
| # | 维度 | 问题 | 模式 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | 术语 | EvidenceGuard 是否判断证据足以推出结论? | evidence-driven | 已解决 |
|
||||
| Q2 | 边界 | verified snapshot 能读取哪些 canonical 字段,并向 SemanticGuard 暴露哪些内容? | evidence-driven | 已解决 |
|
||||
| Q3 | 边界 | `NO_EVIDENCE` 能支持哪类 Analysis? | evidence-driven | 已解决 |
|
||||
| Q4 | 修复 | 首次 EvidenceGuard 失败允许如何修复,是否可重跑 Agent 或 Tool? | evidence-driven | 已解决 |
|
||||
| Q5 | 语义 | SemanticGuard 是否拥有 Tool、记忆、ReAct loop 或报告写作能力? | evidence-driven | 已解决 |
|
||||
| Q6 | 重试 | 哪些 SemanticGuard 结果允许第二次 attempt? | evidence-driven | 已解决 |
|
||||
| Q7 | 超时 | 单轮模型超时和 Run 取消如何终止后台调用? | evidence-driven | 已解决 |
|
||||
| Q8 | 释放 | 三种 Fallback 能否包含 Draft、reason 或未验真来源? | evidence-driven | 已解决 |
|
||||
| Q9 | 接口 | 阶段 5 是否切换公开入口或修改 SSE/持久化协议? | evidence-driven | 已解决 |
|
||||
| Q10 | 验收 | 如何证明不存在隐式 Agent/repair/retry loop? | evidence-driven | 已解决 |
|
||||
|
||||
## Evidence-driven
|
||||
|
||||
| 结论 | 证据来源 | 是否已汇报用户 |
|
||||
|---|---|---|
|
||||
| EvidenceGuard 只验证结构、引用和 canonical invocation 真实性,不判断 Analysis/Conclusion 的语义充分性。 | ISS-014 4.3、阶段 5;glossary `EvidenceGuard` | 已汇报 |
|
||||
| Store 只能按 `ToolCallKeyFactory.create(runId, toolCallId)` 查找;可引用记录必须是当前 Run、`READY`、含 `agent_result` 且状态为 `EVIDENCE_FOUND|NO_EVIDENCE`。 | `CanonicalInvocationStore`、`CanonicalToolInvocation.isReferencableBy` | 已汇报 |
|
||||
| Snapshot 解析 `agent_result` 和必要的有界 request 字段,不读取 `raw_response`;输出按 `analysis_id` 分组且不包含 Tool Call ID。 | ISS-014 4.4、10.2、已冻结 Tool contracts | 已汇报 |
|
||||
| `NORMAL` 只绑定 `EVIDENCE_FOUND`;`NEGATIVE_OBSERVATION` 只绑定 `NO_EVIDENCE`,且零结果范围来自投影/请求。 | ISS-014 4.3、10.2;`EvidenceStatus` glossary | 已汇报 |
|
||||
| Evidence repair 只在首次物理验真失败后执行一次无 Tool单轮模型调用,不重跑 Diagnosis Agent 或 Tool loop。 | ISS-014 10.3、阶段 5、重试边界决策 | 已汇报 |
|
||||
| SemanticGuard 复用同一 `ChatModel`,使用全新 `Prompt` 单轮调用,无 Tool/记忆/ReAct;只返回二元 verdict 和审计 reason。 | ISS-014 4.4、阶段 5;Spring AI `ChatModel.call(Prompt)` | 已汇报 |
|
||||
| 仅超时、传输、解析或 Schema 技术失败可重试一次;`UNSUPPORTED` 是有效业务结果,不重试。 | `HarnessRetryPolicies.strict().semanticGuard()`、ISS-014 重试策略 | 已汇报 |
|
||||
| `ChatResponseMetadata.Usage` 可复用 Core 的 model/Token 预算;受控 `Future.get(timeout)` 可在超时或 Run 取消时取消任务。 | `HarnessModelInterceptor`、`DiagnosisHarnessCore`、`RunCancellation` | 已汇报 |
|
||||
| `EVIDENCE_VALIDATION_FAILED` 的来源必须为空;其余 Fallback 只能包含已验真来源,不包含 Draft 或 SemanticGuard reason。 | ISS-014 10.3、阶段 5;`SafeFallback`/`FallbackType` | 已汇报 |
|
||||
| 阶段 5 只新增内部用例,公开入口切换留到阶段 6B。 | ISS-014 阶段顺序和 6B 门禁 | 已汇报 |
|
||||
|
||||
## User-interview
|
||||
|
||||
- 本阶段没有新增 user-interview 问题。上述方向、范围、修复次数、模型复用、超时/重试、Fallback 和阶段串行规则均已在 ISS-014 评审及前序对话中由用户确认。
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Snapshot 使用 Tool-specific adapter 严格解析冻结 projection;未知 Tool、ID 不一致、projection 反序列化失败或 projection 的 EvidenceStatus 与 canonical record 不一致均 fail closed。
|
||||
- MySQL 的逻辑数据源和查询范围来自 canonical `request`,行/列和值来自 `agent_result`;RAG/log 以 `agent_result` 自带的稳定来源和范围为准。
|
||||
- SemanticGuard 使用直接 `ChatModel.call(Prompt)`,不创建第二个 `ReactAgent`;模型调用由受控 Executor 执行,超时/取消时调用 `Future.cancel(true)`。
|
||||
- Evidence repair 与 SemanticGuard 共用同一模型抽象,但各自拥有独立 prompt、严格输出类型和预算;repair 策略固定一次 attempt,不能嵌套 retry executor。
|
||||
- Release use case 只在 EvidenceGuard 成功后调用 SemanticGuard;SUPPORTED 返回原 Draft 对象,所有其他路径只返回固定 `SafeFallback`。
|
||||
- 不创建 ADR:这些是 ISS-014 已冻结架构的阶段实现,不产生新的难逆转跨项目决策。
|
||||
|
||||
## OpenSpec Backfill
|
||||
|
||||
- 需进入 proposal/design/spec/tasks:EvidenceGuard 规则、Tool-specific snapshot、无 raw/ID 泄漏、repair 单次边界、SemanticGuard 隔离/预算/超时/重试、原样发布与三类 Fallback、L2/公开隔离。
|
||||
- 不进入本阶段:公开 Chat/SSE 装配、Run 持久化、旧多 Agent 删除和 live E2E。
|
||||
|
||||
## Cross-artifact Alignment
|
||||
|
||||
| 上游 -> 下游 | 检查内容 | 状态 |
|
||||
|---|---|---|
|
||||
| ISS-014/brief -> proposal | 物理验真、结构修复、隔离语义审查、固定 Fallback、阶段边界 | 已对齐 |
|
||||
| proposal -> design | Store ownership、Tool-specific snapshot、timeout/cancel/retry、唯一报告作者和 L2 影响 | 已对齐 |
|
||||
| design -> specs/tasks | 每个关键决策均有可观察 requirement 和对应实现/测试切片 | 已对齐 |
|
||||
| specs -> tasks | 9 组 requirements 覆盖为 Evidence、model boundary、repair/release 和回归验证任务 | 已对齐 |
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
- 能力来源:`zoom-out`,使用 glossary 的 Diagnosis Harness、EvidenceGuard、SemanticGuard、RunContext、Invocation Status 和 Evidence Status 术语。
|
||||
- 链路为 `query + Draft + RunContext -> EvidenceGuard/store -> optional repair/shared model boundary -> SemanticGuard/shared model boundary -> original Draft or fixed Fallback`,没有外层 Graph 或第二个 Tool loop。
|
||||
- Store 拥有物理调用记录,EvidenceGuard 只读并生成 snapshot;Diagnosis Agent 拥有报告语义,repair 只修 ID,SemanticGuard 只审查,release 只做二选一,数据所有权没有重叠。
|
||||
- 最大运行风险是 provider 忽略 interrupt;design 要求 Future cancel、迟到 Token 记录和 `core.checkActive` 丢弃迟到结果,阶段 6A 再负责 Run 终态。
|
||||
- 架构审计未发现与阶段 0-4 或 ADR 冲突;所有风险缓解已回写 design/spec/tasks。
|
||||
|
||||
## Interface Impact
|
||||
|
||||
- 级别:L2 内部接口。
|
||||
- 变更对象:新增 guard/release records、interfaces、use cases、prompts 和 tests;现有方法签名不变。
|
||||
- 消费者:本阶段只有 focused tests,阶段 6A 才接入 production application use case。
|
||||
- 兼容性:公开 HTTP/SSE、Controller DTO、数据库、Redis key schema 和旧 Chat/AiOps 路径不变。
|
||||
|
||||
## Commit Gate Preflight
|
||||
|
||||
- proposal、design、specs、tasks 完整,`openspec status` 为 complete,`openspec validate single-react-evidence-semantic-guards --strict` 通过。
|
||||
- Question pool 全部为已汇报的 evidence-driven 结论;无未确认 user-interview、接口等级或风险接受问题。
|
||||
- Cross-artifact 四段对齐无 gap;架构审计的数据所有权、模型取消和唯一报告作者约束已进入 design/spec/tasks。
|
||||
- 接口影响为 L2,仅新增内部 Java API;公开 Controller/SSE/JPA/Redis schema/旧 ChatService 保持不变。
|
||||
- Apply、Archive 和阶段 Git commit 使用用户对 ISS-014 各阶段的持续授权;实现必须严格限制为 Committed OpenSpec。
|
||||
- `.committed` 已创建,Committed OpenSpec 可进入 Apply。
|
||||
|
||||
## Pre-apply Research
|
||||
|
||||
### Reference Implementations
|
||||
|
||||
- `harness/tool/store/CanonicalInvocationStore.java`、`CanonicalToolInvocation.java`、`ToolCallKeyFactory.java`:当前 Run 物理验真和 lifecycle 真理源。
|
||||
- `harness/tool/projection/RagResultProjector.java`、`QueryLogsResultProjector.java`、`harness/tool/mysql/MysqlResultProjector.java`:三类 `agent_result` 的唯一生产者和有界字段来源。
|
||||
- `harness/agent/HarnessModelInterceptor.java`:模型调用前预算、响应 Usage 记账和 late-result active check 模式。
|
||||
- `harness/retry/HarnessRetryExecutor.java`、`HarnessRetryPolicies.java`:SemanticGuard 两次技术 attempt 与 repair 一次 attempt 的装配边界。
|
||||
- `harness/core/RunCancellation.java`:取消 callback 注册和 first-cancel 语义。
|
||||
- `service/ExecutorGatekeeperService.java`:仅参考报告内部引用检查思想,不复用旧 Map/JPA/session DTO 或规则目录。
|
||||
|
||||
### Technology Stack
|
||||
|
||||
- Jackson strict `ObjectReader` 解析 frozen Draft/Tool projections;Tool-specific adapter 生成 typed snapshot,不透传 JsonNode/raw response。
|
||||
- Spring AI `ChatModel.call(Prompt)` 执行无 Tool 单轮 guard;`ChatResponseMetadata.Usage` 接入现有 Core Token 预算。
|
||||
- Java 17 `ExecutorService/Future.get(timeout)` 提供 per-attempt timeout,`Future.cancel(true)` 接入 Run cancellation;Semantic 总时限由 monotonic elapsed time 控制。
|
||||
- JUnit 5 scripted `ChatModel` 和 in-memory canonical store 作为外部边界 fake;测试只通过 EvidenceGuard/SemanticGuard/Release use case 公共接口断言行为。
|
||||
- 无 Controller、MQ、JPA、Flyway 或新 Maven dependency;不需要新共享基础设施。
|
||||
|
||||
## Apply Progress
|
||||
|
||||
- 首模块对齐:EvidenceGuard typed contracts、Draft/reference checks、current-Run canonical validation 和 RAG/log/MySQL snapshot adapters 已完成,tasks 1.1-1.4 完成。
|
||||
- 8 个 `EvidenceGuardTest` 行为测试通过;快照序列化不含 Tool Call ID 或 `raw_response`,`NO_EVIDENCE` 仅在 `NEGATIVE_OBSERVATION` 下进入带 scope/zero-match 的快照。
|
||||
- TODO:tasks 2.1-4.2,尚未实现模型边界、SemanticGuard、repair/release 和综合回归。
|
||||
- REVIEW 发现 corrupted Store key 下 record ID 可能与 Draft reference 不同;分类为代码偏离,已增加 exact canonical ID 校验和回归测试,无需改变规格方向。
|
||||
- REVIEW 发现 fallback `DiagnosisReleaseResult` 携带完整 snapshot 会通过 `analysis_text` 间接泄漏 Draft;分类为规格安全边界细化,已回写 design/spec,并让 fallback result 强制使用空 snapshot,安全来源只保留在 `SafeFallback.verified_sources`。
|
||||
|
||||
## Apply Result
|
||||
|
||||
- 新增确定性 EvidenceGuard:校验 typed Draft、唯一 Analysis ID、报告引用、exact current-Run Tool ID、READY lifecycle、agent result、Evidence Status 与 Analysis Kind,并严格展开 RAG/log/MySQL projection。
|
||||
- verified snapshot 按 Analysis ID 分组,只含稳定 source/scope/timestamp/excerpt/values;SemanticGuard 输入不含 Tool Call ID、Redis key 或 raw response。
|
||||
- 新增共享 `GuardModelCall`:复用系统 ChatModel,执行 Core model/Token/Run byte budget、per-attempt timeout、total timeout、Future cancellation 和 late-result active check。
|
||||
- 新增无 Tool/无记忆/无 ReAct 的单轮 SemanticGuard,严格解析 `SUPPORTED|UNSUPPORTED + reason`;仅技术失败使用既有两次 attempt,业务 UNSUPPORTED 不重试。
|
||||
- 新增一次 EvidenceRepair,只允许 ID/reference 修复;任何可见文本、kind、顺序、human-confirmation 或 limitations 变化都直接 Evidence fallback。
|
||||
- 新增 DiagnosisReleaseUseCase 和固定 SafeFallbackFactory:SUPPORTED 原样返回 Draft;evidence failed、semantic unsupported/unavailable 均不返回 Draft、完整 snapshot 或审计 reason。
|
||||
- 公开 Controller、ChatService、AiOpsService、SSE、JPA/Flyway 和旧多 Agent 链路无修改。
|
||||
|
||||
## Apply Verification
|
||||
|
||||
- 编译:`mvn -q -DskipTests compile` 通过。
|
||||
- 阶段 5 focused:`EvidenceGuardTest` 9、`SemanticGuardTest` 4、`DiagnosisReleaseUseCaseTest` 6,共 19 tests,0 failure/error。
|
||||
- 综合回归:阶段 5 + Harness Core/Retry/Tool boundary/store/projections/contracts + Diagnosis Agent,共 18 suites / 70 tests,0 failure/error/skipped。
|
||||
- OpenSpec:`openspec validate single-react-evidence-semantic-guards --strict` 通过。
|
||||
- 静态范围:公开 Controller/ChatService/AiOpsService diff 为空;SemanticGuard 无 Tool ID/raw response/ReactAgent/StateGraph/ThreadLocal/手写 while;新 guard/release 包无业务 Agent loop。
|
||||
- 仓库级 `openspec validate --specs --strict` 为 16 passed / 2 failed;失败仍是前序 `mysql-readonly-tool`、`rag-log-projections` 的 scenario 标题格式,按计划阶段 7 统一修复。
|
||||
- 未执行 live E2E:按 ISS-014 阶段门禁统一留到阶段 7。
|
||||
|
||||
## Archive Result
|
||||
|
||||
- 14/14 OpenSpec tasks 完成,`.archive-ready` 已创建。
|
||||
- 主规格已同步至 `openspec/specs/single-react-evidence-semantic-guards/spec.md`,新增 9 个 requirements。
|
||||
- Change 已归档至 `openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards`。
|
||||
- `devflow/index.md` 和 ISS-014 阶段表已更新为阶段 0-5 archived,下一阶段为 6A。
|
||||
- 未创建 ADR/compound knowledge:本阶段落实 ISS-014 已冻结边界,没有新的跨项目难逆转决策。
|
||||
@@ -0,0 +1,29 @@
|
||||
# Evidence: single-react-evidence-semantic-guards
|
||||
|
||||
## Code and Contract Evidence
|
||||
|
||||
- `CanonicalInvocationStore` 只提供 key lookup;`ToolCallKeyFactory` 冻结 `runId + toolCallId` 隔离,`CanonicalToolInvocation` 冻结 READY/agent_result/evidence status 可引用条件。
|
||||
- `RagToolResult`、`QueryLogsToolResult`、`MysqlToolResult` 是 Agent-facing 有界 projection,足以构造 snapshot;只有 MySQL 的逻辑数据源/SQL/params 需要从 canonical request 补足。
|
||||
- `AnalysisKind.accepts` 已冻结 `NORMAL -> EVIDENCE_FOUND`、`NEGATIVE_OBSERVATION -> NO_EVIDENCE`。
|
||||
- `HarnessRetryPolicies.strict()` 已冻结 SemanticGuard 两次技术 attempt 和 Evidence repair 一次 attempt。
|
||||
- `ChatModel.call(Prompt)`、`ChatResponseMetadata.Usage` 和 `RunCancellation.onCancel` 支持直接单轮模型调用、Token 记账与 Future cancellation,无需 ReactAgent。
|
||||
|
||||
## Confirmed Boundaries
|
||||
|
||||
- EvidenceGuard 不判断证据是否支持结论,只验证结构、引用和物理真实性。
|
||||
- SemanticGuard 只接收原始 Query、移除 Tool ID 的完整 Draft view 和 verified snapshot,不访问 Redis/raw response。
|
||||
- Evidence repair 只修改 ID/reference;用户可见语义发生任何变化即失败。
|
||||
- `UNSUPPORTED` 是有效业务结果,不重试;只有 timeout/transport/parse/schema 技术失败可进行第二次 attempt。
|
||||
- Fallback 不含 Draft、完整 snapshot 或 SemanticGuard reason;Evidence failure 的 verified sources 为空。
|
||||
|
||||
## Review Findings
|
||||
|
||||
- 增加 canonical record ID 与 Draft reference 的 exact match,防止 corrupted key 映射被误信任。
|
||||
- Fallback release result 强制使用空 snapshot,防止 `analysis_text` 通过误序列化泄漏 Draft;安全来源只保留在 `SafeFallback.verified_sources`。
|
||||
|
||||
## Verification Evidence
|
||||
|
||||
- Stage-focused: 19 tests,覆盖 EvidenceGuard 9、SemanticGuard 4、DiagnosisReleaseUseCase 6。
|
||||
- Regression: 18 suites / 70 tests,0 failure/error/skipped。
|
||||
- Maven compile 和 change strict validation 通过。
|
||||
- 公开 Controller/ChatService/AiOpsService 零 diff;SemanticGuard 无 Tool ID/raw/ReactAgent/Graph/ThreadLocal/手写 loop。
|
||||
@@ -0,0 +1,45 @@
|
||||
# Acceptance: single-react-harness-run-context
|
||||
|
||||
## 实现结果
|
||||
|
||||
- 新增显式 `RunContext`、取消信号、deadline 检查、Run Lifecycle 和 first-terminal-wins。
|
||||
- 新增 caller-supplied limits、线程安全 Model/Tool/Token/Run bytes 预算和容量 CAS。
|
||||
- 新增 strict typed retry policies/executor,记录每个实际 attempt 并阻止取消/预算异常误重试。
|
||||
- 新增无 Redis 依赖的 ToolCallKeyFactory,精确保留框架 Tool Call ID。
|
||||
- Spring AI `spring.ai.retry.max-attempts` 设置为 1;未修改 provider/model routing。
|
||||
- 未接入旧 Chat/AIOps/Controller、ThreadLocal、JPA、Redis、Agent 或公开协议。
|
||||
|
||||
## 静态验证
|
||||
|
||||
- `openspec validate single-react-harness-run-context --strict`:通过。
|
||||
- 新 Harness 包 `rg`:无 ThreadLocal/current-holder/Redis 引用。
|
||||
- 旧 Chat/AIOps/Controller/JPA 调用链 diff:为空。
|
||||
- 受保护配置检查:仅新增 `spring.ai.retry.max-attempts: 1`,model routing/provider 保持不变。
|
||||
|
||||
## 脚本验证
|
||||
|
||||
- `mvn -q -DskipTests compile`:通过。
|
||||
- Core focused suite:通过。
|
||||
- 综合回归 suite(阶段 0/1 契约 + 阶段 2 Core + ChatController):通过。
|
||||
|
||||
## 浏览器/人工验证
|
||||
|
||||
- 不适用。本阶段没有 UI、Controller、SSE 或公开协议变化。
|
||||
|
||||
## 未验证
|
||||
|
||||
- 未执行真实模型调用、Redis、CLS、MySQL 或客户端断开 live E2E;这些属于后续 Tool store/adapter/最终 E2E 门禁。
|
||||
- 未将 RunContext 接入旧 ChatService;这是本阶段明确非目标,阶段 6A 才接入。
|
||||
- `RunContext` 取消对已进入的同步第三方调用仍是协作式;实际 HTTP/JDBC future 取消留给后续 adapter。
|
||||
- Provider 侧凭据轮换状态不由仓库证明。
|
||||
|
||||
## 剩余风险与后续门禁
|
||||
|
||||
- 关闭 SDK 隐式 retry 会使旧路径瞬时错误不再自动重试,直到后续 Router/SemanticGuard 接入 Harness;该变化已记录并可通过恢复配置回滚。
|
||||
- 下一阶段 3A 必须在 ToolInterceptor 接收 RunContext 和框架 Tool Call ID,直接复用 Key/Capacity/Cancellation 门禁。
|
||||
|
||||
## 状态
|
||||
|
||||
- Stage acceptance: accepted
|
||||
- OpenSpec archive: archived at `openspec/changes/archive/2026-07-21-single-react-harness-run-context`
|
||||
- Main spec sync: `openspec/specs/diagnosis-harness-run-context/spec.md`(8 added requirements)
|
||||
@@ -0,0 +1,32 @@
|
||||
# Brief: single-react-harness-run-context
|
||||
|
||||
## 背景
|
||||
|
||||
旧 Chat/AIOps 通过多个 ThreadLocal 和业务方法内状态机传播 session/run/token/retry,无法成为后续 Tool、Agent、Guard 和新入口的稳定共同边界。Spring AI 默认 10 attempts 还会制造未被 Harness 记录的隐藏重试。
|
||||
|
||||
## 目标
|
||||
|
||||
- 建立显式、结构不可变、可异步传播的 RunContext。
|
||||
- 集中实现 deadline、协作式取消、线程安全预算和唯一 Run 终态。
|
||||
- 实现类型化、最多两次 attempt、逐 attempt 记录的 Harness retry。
|
||||
- 提供阶段 3A 可直接使用的 Tool Call Key Factory 和单 Run 容量计数器。
|
||||
- 将 Spring AI 底层 retry 压为一次。
|
||||
|
||||
## 范围
|
||||
|
||||
- 新增 Harness Core/Retry/Tool Store foundation 类型和 focused Fake tests。
|
||||
- 更新 `application.yml` 的 Spring AI retry 配置。
|
||||
- 更新 glossary 和 OpenSpec/devflow 档案。
|
||||
|
||||
## 非目标
|
||||
|
||||
- 不接入旧 ChatService/AiOpsService/Controller/SSE。
|
||||
- 不读写现有 ThreadLocal,不删除旧实现。
|
||||
- 不写 `diagnosis_run` 或 Redis,不实现 Tool projection。
|
||||
|
||||
## 元数据
|
||||
|
||||
- 分档:complex
|
||||
- 接口影响:L2;另有关闭旧 SDK 隐式 retry 的有意内部行为变化
|
||||
- 关联 Issue:ISS-014 阶段 2
|
||||
- 关联 OpenSpec:`openspec/changes/single-react-harness-run-context`
|
||||
@@ -0,0 +1,110 @@
|
||||
# Decisions: single-react-harness-run-context
|
||||
|
||||
## 规模与入口
|
||||
|
||||
- 分档:complex。
|
||||
- 入口:ISS-014 阶段 2;阶段 0/1 已分别由 Git commit `58c3910`、`4274f33` 固化并 Archive。
|
||||
- 目标:建立后续 Tool、Agent、Guard 和入口共同依赖的显式 Harness 运行边界,不接旧 ChatService。
|
||||
|
||||
## Context
|
||||
|
||||
- `devflow/index.md` 命中 design freeze、ACI contracts 和 session/run trace isolation。
|
||||
- 当前 `SessionContextHolder`、`VerifierContextHolder`、`TokenUsageHolder` 使用 ThreadLocal;新 Harness 禁止复用。
|
||||
- 当前 ChatService 自行生成 runId、写 `diagnosis_run`、执行两轮 retry loop 并在 finally 清理 ThreadLocal,职责与新 Harness 边界冲突。
|
||||
- `DiagnosisRun.status` 当前是字符串 `PENDING/RUNNING/SUCCESS/FAILED`;本阶段不修改实体或迁移,由后续应用用例映射。
|
||||
- Spring AI 1.1.7 `SpringAiRetryProperties` 的配置前缀是 `spring.ai.retry`,默认 `maxAttempts=10`。
|
||||
|
||||
## Question Pool
|
||||
|
||||
| 维度 | 问题 | 模式 | 证据与结论 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| 术语 | “不可变 RunContext”是否意味着预算/取消也不能变化? | evidence-driven | ISS 要求上下文不可变同时要求计数/取消/终态;采用结构不可变 record + 线程安全单 Run 状态句柄。 | 已解决并汇报 |
|
||||
| 术语 | Run 与 DiagnosisRun 是否在阶段 2 直接持久化绑定? | evidence-driven | 阶段 2 只建 Core,阶段 6A 才接应用用例;本阶段生命周期为内存执行真理,不改 JPA。 | 已解决并汇报 |
|
||||
| 边界 | 是否迁移旧 ThreadLocal/ChatService 调用? | evidence-driven | ISS 1109-1111 明确禁止 ThreadLocal 和旧 ChatService 临时适配;只新增零消费者 Core。 | 已解决并汇报 |
|
||||
| 边界 | 取消是否承诺立即中断同步模型调用? | evidence-driven | 阶段 0 已确认分层取消;阻止后续边界并执行资源回调,不承诺不可证明的硬中断。 | 已解决并汇报 |
|
||||
| 验收 | 唯一终态如何证明? | evidence-driven | atomic first-terminal-wins lifecycle,并用取消/预算/异常/成功竞态测试证明后续终态不能覆盖。 | 已解决并汇报 |
|
||||
| 验收 | 异步传播如何证明不依赖 ThreadLocal? | evidence-driven | Fake Tool 在 `CompletableFuture` 线程只接收显式 RunContext,并验证相同 session/run/state handles。 | 已解决并汇报 |
|
||||
| 技术 | 隐藏重试如何关闭? | evidence-driven | 本地依赖 `SpringAiRetryProperties` 证实 `spring.ai.retry.max-attempts` 默认 10;配置改为 1 并加配置测试。 | 已解决并汇报 |
|
||||
| 技术 | 尚未校准的预算默认值如何处理? | evidence-driven | ISS 要求集中配置且不伪造数值;Core 接受显式 limits,不内置默认预算,后续 Spring wiring 决定配置值。 | 已解决并汇报 |
|
||||
| 技术 | Tool Call Key 是否生成 Tool ID? | evidence-driven | 阶段 0/1 冻结框架 ID;Factory 仅验证安全 segment 并拼接,不生成或改写。 | 已解决并汇报 |
|
||||
|
||||
## Grill 结论
|
||||
|
||||
- 术语、边界、验收和技术问题均可由已确认 ISS、现有代码和本地依赖 API 证明。
|
||||
- 没有新的产品偏好、公开协议或风险接受度问题需要 `user-interview`;短期关闭旧 SDK retry 的影响已在 proposal 明示。
|
||||
- `grill-with-docs` 的代码可证问题已先查证并向用户汇报;所有结论均已回写 proposal。
|
||||
|
||||
## 能力与工具限制
|
||||
|
||||
- Discover 能力来源:`sm-flow` + `grill-with-docs`。
|
||||
- 当前工具集没有 `codebase-retrieval` 和 LSP;使用 `rg`、源码阅读、本地依赖 `jar/javap`、编译和 focused tests 补足调用链与 API 核对。
|
||||
|
||||
## Cross-artifact 对齐
|
||||
|
||||
| 链路 | 状态 | 结论 |
|
||||
|---|---|---|
|
||||
| brief 目标/范围/非目标 -> proposal | 已对齐 | 显式 context、Core、预算/取消/终态、retry、key/capacity、隐藏 retry 和不接旧 runtime 全部覆盖。 |
|
||||
| proposal 范围/约束/承诺 -> design | 已对齐 | 类职责、并发语义、first-wins、配置变化、迁移和回滚均有明确设计。 |
|
||||
| design 决策/接口影响/风险 -> specs/tasks | 已对齐 | 每个状态边界都有 scenario,配置行为变化和 zero-consumer 边界有独立任务与验证。 |
|
||||
| specs 可观察行为 -> tasks | 已对齐 | 8 条 requirements 分解为 primitives、budget、Core、retry、key、配置和三层验证。 |
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
- 能力来源:`zoom-out`,使用 glossary 的 RunContext、Run Lifecycle、Run Budget、Harness Retry Policy 和 Diagnosis Run 术语。
|
||||
- 输入链路为未来 application use case 创建 RunContext,处理链路由 Core/Retry/Key/Capacity 通过显式参数消费,输出为唯一 RunTermination;当前 runtime 不接入。
|
||||
- RunContext 拥有单 Run 内存状态,应用用例拥有数据库映射,Tool store 拥有 Redis invocation;数据所有权没有重叠。
|
||||
- Core 不保存全局 Run map、不调用 Agent/数据库/Redis、不实现循环编排,因此不会演变为工作流引擎。
|
||||
- 最大风险是关闭 SDK retry 对旧路径的短期行为影响;已作为显式配置变更进入 proposal/spec/test 和回滚说明,无 ADR 冲突。
|
||||
|
||||
## Commit Gate Preflight
|
||||
|
||||
- proposal、design、specs、tasks 完整,OpenSpec status complete,strict validation 通过。
|
||||
- question pool 无未汇报 evidence-driven 结论、无未确认 user-interview 问题。
|
||||
- 接口影响 L2;`spring.ai.retry.max-attempts=1` 的有意内部行为变化已明确影响和回滚边界。
|
||||
- cross-artifact 四段对齐无 gap,架构审计约束已进入 design/spec/tasks。
|
||||
- Apply 持续授权已存在;执行范围严格限制为新 Harness foundation、配置 override 和 focused tests。
|
||||
|
||||
## Pre-apply Research
|
||||
|
||||
### 参考实现与反例
|
||||
|
||||
- `SessionContextHolder`:ThreadLocal session/run fallback,新 Harness 明确禁止复用。
|
||||
- `TokenUsageHolder` / `TokenTrackingChatModel`:当前只记录 total token 且依赖 ThreadLocal,后续 model boundary 应改为显式 RunBudget;本阶段不改旧类。
|
||||
- `ChatService.executeChatComplex`:当前业务方法内创建 Run、两轮 retry、写终态和清理 ThreadLocal,是后续替换对象,不是 Core 参考实现。
|
||||
- `DiagnosisRun`:现有持久化字段和字符串状态;本阶段只确认映射边界,不修改实体或 Repository。
|
||||
- Spring AI 1.1.7 `SpringAiRetryProperties`:`spring.ai.retry` 前缀、默认 `maxAttempts=10`,支持精确配置 override。
|
||||
|
||||
### 技术栈清单
|
||||
|
||||
- Java 17 record 表达结构不可变 context/limits/snapshot/policy/attempt。
|
||||
- `AtomicReference` 实现 first-reason/first-terminal-wins;`AtomicLong` 实现 capacity CAS;同步临界区维护复合预算一致性。
|
||||
- `Clock` 和 `Supplier<String>` 注入保证 deadline/ID 可测试,不引入 scheduler 或全局 registry。
|
||||
- SLF4J 只记录取消 callback 异常,不记录用户输入、Tool payload 或凭据。
|
||||
- SnakeYAML 直接解析 classpath `application.yml` 验证 retry override,不启动外部 MySQL/Redis/Milvus/模型。
|
||||
|
||||
### 新建基础设施
|
||||
|
||||
- `harness.core`:RunContext、Cancellation、Lifecycle、Budget、Capacity、Core 和类型化异常/状态。
|
||||
- `harness.retry`:RetryFailure/Policy/Policies/Attempt/Executor/Exception 与函数接口。
|
||||
- `harness.tool.store.ToolCallKeyFactory`:纯 Key 构造,不访问 Redis。
|
||||
- focused unit tests 与 Fake Model/Tool;无需新 Maven 依赖。
|
||||
|
||||
### 影响半径
|
||||
|
||||
- 新生产包在本阶段保持零消费者。
|
||||
- 唯一现有运行配置变化为 `spring.ai.retry.max-attempts=1`;Model 路由、provider、Controller、JPA 和 Redis 配置保持不变。
|
||||
|
||||
## Apply 结果
|
||||
|
||||
- 冲突分类:未发现 OpenSpec 遗漏、代码偏离或方向不确定项;一次自审发现 RetryExecutor 需要无条件拦截预算/取消异常,已回写代码并通过回归测试。
|
||||
- 新增 `RunContext`、Cancellation、Lifecycle、Budget、Capacity、DiagnosisHarnessCore、typed Retry 和 ToolCallKeyFactory;未接旧 Chat/AIOps/Controller/Redis/JPA。
|
||||
- Spring AI 全局 retry 已由默认 10 压为 1;Harness strict policies 只允许 Router/SemanticGuard 技术失败一次显式重试。
|
||||
- 首模块对齐:Run state/budget/Core/retry/key/config 与 design/tasks 全部完成;Tool interceptor/store/Agent/应用用例仍留给后续阶段。
|
||||
|
||||
## Apply 验证
|
||||
|
||||
- 编译:`mvn -q -DskipTests compile`:通过。
|
||||
- Core focused:`mvn -q '-Dtest=RunContextTest,RunBudgetTest,DiagnosisHarnessCoreTest,HarnessRetryExecutorTest,ToolCallKeyFactoryTest,SpringAiRetryConfigurationTest' test`:通过(加固后复跑通过)。
|
||||
- 综合回归:`mvn -q '-Dtest=HarnessContractTest,RagToolContractTest,QueryLogsToolContractTest,MysqlToolContractTest,RunContextTest,RunBudgetTest,DiagnosisHarnessCoreTest,HarnessRetryExecutorTest,ToolCallKeyFactoryTest,SpringAiRetryConfigurationTest,ChatControllerTest' test`:通过。
|
||||
- 静态 scope:新 Harness 包无 ThreadLocal/current-holder/Redis 引用;旧 Chat/AIOps/Controller/JPA 调用链 diff 为空;模型路由/provider 未改。
|
||||
- OpenSpec:`openspec validate single-react-harness-run-context --strict`:通过。
|
||||
@@ -0,0 +1,25 @@
|
||||
# Evidence: single-react-harness-run-context
|
||||
|
||||
## 文档与依赖证据
|
||||
|
||||
- ISS-014 4.2/阶段 2 要求显式 `RunContext`、deadline、取消、预算、retry、Key Factory 和 no-ThreadLocal 边界。
|
||||
- 阶段 0/1 OpenSpec 已冻结框架 `tool_call_id`、两套状态语义和后续阶段串行门禁。
|
||||
- 本地 Spring AI 1.1.7 `SpringAiRetryProperties` 的 `@ConfigurationProperties("spring.ai.retry")` 默认 `maxAttempts=10`;配置已覆盖为 1。
|
||||
|
||||
## 代码证据
|
||||
|
||||
- `SessionContextHolder`、`TokenUsageHolder`、`VerifierContextHolder` 当前是旧链路 ThreadLocal;新 `com.superbiz.agent.harness` 包无任何 holder/ThreadLocal/Redis 引用。
|
||||
- `ChatService.executeChatComplex` 当前自行创建 run、执行两轮 retry、写 `diagnosis_run` 和清理 ThreadLocal;新 Core 不接入该方法,后续应用用例负责迁移。
|
||||
- `DiagnosisRun` 仍保留现有字符串状态和 JPA Schema;阶段 2 未修改实体、Repository 或数据库。
|
||||
- `DiagnosisHarnessCore` 不保存全局 Run map;RunContext 结构不可变,Cancellation/Budget/Lifecycle 为同一 Run 的线程安全句柄。
|
||||
|
||||
## Evidence-driven 结论
|
||||
|
||||
- first-reason-wins cancellation + first-terminal-wins lifecycle 可以用 AtomicReference 实现并跨异步边界共享。
|
||||
- 复合 Tool/Token 预算需要同步一致性;Run bytes 用 CAS 预留避免并发超限或部分增长。
|
||||
- SDK 隐藏 retry 压为一次后,Harness 才能记录 Router/SemanticGuard 的显式 attempt;Tool/Diagnosis/Evidence repair 固定一次。
|
||||
- Key Factory 可直接复用阶段 3A,但不生成或改写框架 Tool Call ID,也不访问 Redis。
|
||||
|
||||
## 工具限制
|
||||
|
||||
- `codebase-retrieval` 和 LSP 不在当前工具集中;使用 `rg`、源码阅读、`jar/javap`、Maven 编译、YAML 解析和 focused tests 补足核对。
|
||||
@@ -0,0 +1,26 @@
|
||||
# Acceptance: single-react-mysql-readonly-tool
|
||||
|
||||
## Commit preflight
|
||||
|
||||
- OpenSpec strict validation: passed.
|
||||
- Scope: fail-closed SQL validator, exact allowlist, JDBC read-only executor, bounded MySQL projection, ToolBoundary adapter and query-script safety cleanup.
|
||||
- Non-goals: public Agent/Chat cutover, metadata discovery, Agent persistence database access, dynamic/tenant authorization and live production datasource provisioning.
|
||||
- Security prerequisite: script write branch/default external connection values are explicitly included in this change.
|
||||
|
||||
## Apply acceptance
|
||||
|
||||
- Implemented and verified.
|
||||
- Static, Maven and script evidence is recorded in `evidence.md`.
|
||||
- No browser/manual verification applies; this stage adds no UI or public protocol change.
|
||||
- Residual risk is limited to later live datasource provisioning/driver behavior and stage 4 integration.
|
||||
|
||||
## Archive acceptance
|
||||
|
||||
- All OpenSpec tasks are complete.
|
||||
- `.archive-ready` marker is created after focused verification.
|
||||
- OpenSpec is ready to move to the dated archive directory.
|
||||
- Archive completed at `openspec/changes/archive/2026-07-21-single-react-mysql-readonly-tool`.
|
||||
|
||||
## Remaining work
|
||||
|
||||
- Stage 4 Diagnosis Agent integration and later live datasource/driver E2E remain outside this archive.
|
||||
@@ -0,0 +1,26 @@
|
||||
# Brief: single-react-mysql-readonly-tool
|
||||
|
||||
## Background
|
||||
|
||||
阶段 3A 提供了统一 ToolBoundary 和 canonical invocation store,阶段 3B 提供了 RAG/日志投影。3C 需要独立实现安全敏感的只读 MySQL Tool,避免 Agent 生成的 SQL 直接进入数据库。
|
||||
|
||||
## Goals
|
||||
|
||||
- 使用 JSqlParser 对保守 SELECT 子集进行 fail-closed AST 校验。
|
||||
- 使用逻辑数据源和 schema/table/column 精确 allowlist 授权。
|
||||
- 使用参数绑定、只读 JDBC、超时、取消和结果预算。
|
||||
- 通过阶段 3A boundary 投影为冻结的 `MysqlToolResult`。
|
||||
- 清理查询脚本的写入分支和默认连接风险。
|
||||
|
||||
## Non-goals
|
||||
|
||||
- 不接入公开 Agent/Chat 入口。
|
||||
- 不提供元数据发现、动态授权、租户/行级权限或生产 datasource provisioning。
|
||||
- 不查询 Agent 自身持久化数据库。
|
||||
|
||||
## Classification
|
||||
|
||||
- Scale: complex
|
||||
- Interface impact: L2 internal Harness tool/adapter, plus build dependency and script safety behavior
|
||||
- Issue: ISS-014 stage 3C
|
||||
- Change slug: `single-react-mysql-readonly-tool`
|
||||
@@ -0,0 +1,108 @@
|
||||
# Decisions: single-react-mysql-readonly-tool
|
||||
|
||||
## Discover status
|
||||
|
||||
- Checkpoint: Discover
|
||||
- Capability source: `sm-flow` with local ISS-014, OpenSpec contracts, existing JDBC dependency/configuration and JSqlParser 4.6 already present in the local Maven cache.
|
||||
- Scale: complex, because this stage combines AST policy, authorization, JDBC resource limits, projection, and security cleanup.
|
||||
|
||||
## Evidence-driven findings
|
||||
|
||||
1. `MysqlToolRequest` and `MysqlToolResult` are already frozen under `harness.tool.contract`; no public DTO change is needed.
|
||||
2. The project already has MySQL JDBC/JPA dependencies, but no Agent-facing external read-only executor or SQL policy.
|
||||
3. JSqlParser 4.6 is available in the local Maven cache and exposes `CCJSqlParserUtil`, `Select`, `PlainSelect`, `Table`, `Column`, `Function`, `JdbcParameter` and visitor adapters compatible with Java 17.
|
||||
4. The current application datasource points to the Agent persistence database; the new Tool must use an independently configured logical datasource map and must not reuse that datasource implicitly.
|
||||
5. `scripts/query_mysql.py` currently defaults host/port/user values and contains a non-SELECT commit branch. This violates the ISS-014 security prerequisite and will be changed to read-only, environment-only behavior.
|
||||
|
||||
## Question pool
|
||||
|
||||
| Dimension | Question | Mode | Conclusion | Status |
|
||||
|---|---|---|---|---|
|
||||
| SQL language | Which SQL subset is executable? | evidence-driven | One SELECT, explicit columns, INNER/LEFT JOIN, predicates/group/order, parameter placeholders and allowlisted aggregates. | resolved |
|
||||
| Security | How is authorization decided? | evidence-driven | Independent exact schema/table/column allowlist; parser acceptance alone is insufficient. | resolved |
|
||||
| Data source | Can the Agent pass JDBC coordinates? | evidence-driven | No. Only logical data_source IDs are accepted; connection properties remain configuration/Secret data. | resolved |
|
||||
| Execution | Which JDBC controls are mandatory? | evidence-driven | PreparedStatement, readOnly connection, setMaxRows, query timeout and Run cancellation. | resolved |
|
||||
| Metadata | Can the Tool discover tables/columns? | evidence-driven | No. SHOW/DESCRIBE/information_schema are rejected. | resolved |
|
||||
| Compatibility | Does this cut over public runtime now? | evidence-driven | No. Add internal adapter/executor; Diagnosis Agent integration is stage 4. | resolved |
|
||||
|
||||
## User-confirmed direction
|
||||
|
||||
- Use the frozen `MysqlToolRequest`/`MysqlToolResult` contract.
|
||||
- Reuse the existing ToolBoundary and canonical invocation store.
|
||||
- Keep stage boundaries serial: archive and commit 3C before stage 4.
|
||||
- Do not pause for routine apply/archive/commit confirmation.
|
||||
|
||||
## Pre-apply research
|
||||
|
||||
### Existing implementations and dependencies
|
||||
|
||||
- `src/main/java/com/superbiz/agent/harness/tool/contract/MysqlToolRequest.java`
|
||||
- `src/main/java/com/superbiz/agent/harness/tool/contract/MysqlToolResult.java`
|
||||
- `src/main/java/com/superbiz/agent/harness/tool/boundary/ToolBoundary.java`
|
||||
- `src/main/java/com/superbiz/agent/harness/tool/adapter/QueryLogsToolAdapter.java`
|
||||
- `src/main/resources/application.yml`
|
||||
- `pom.xml` (`mysql-connector-j` already present; add JSqlParser 4.6)
|
||||
- `scripts/query_mysql.py`
|
||||
|
||||
### New classes
|
||||
|
||||
- `MysqlToolLimits`
|
||||
- `MysqlDataSourceDefinition` / allowlist value objects
|
||||
- `MysqlQueryPlan`
|
||||
- `MysqlSqlValidator`
|
||||
- `MysqlReadOnlyExecutor` and JDBC implementation
|
||||
- `MysqlResultProjector`
|
||||
- `MysqlToolAdapter`
|
||||
|
||||
### Risk controls
|
||||
|
||||
- Do not create a generic plugin/DSL layer.
|
||||
- Do not use regex or `startsWith` as SQL authorization.
|
||||
- Fail closed on parser/visitor uncertainty.
|
||||
- Keep raw result canonical-only and expose only bounded projection.
|
||||
|
||||
## Commit checkpoint preparation
|
||||
|
||||
- Proposal scope, design choices, frozen SQL subset, security script cleanup and acceptance scenarios are ready for Commit artifact generation.
|
||||
|
||||
## Commit audit
|
||||
|
||||
- Capability source: `sm-flow` and local OpenSpec CLI.
|
||||
- OpenSpec strict validation: passed for `single-react-mysql-readonly-tool`.
|
||||
- Cross-artifact alignment:
|
||||
- brief goals/non-goals -> proposal scope: aligned.
|
||||
- proposal SQL/security boundaries -> design architecture: aligned.
|
||||
- design validator/executor/projector decisions -> spec requirements: aligned.
|
||||
- spec scenarios -> tasks for dependency, policy, JDBC, projection, adapter and verification: aligned.
|
||||
- Interface impact: L2 internal Harness tool/adapter plus JSqlParser dependency and query-script behavior; no public protocol changes.
|
||||
- Preflight risk accepted: parser ambiguity, datasource isolation, driver cancellation behavior and sensitive result values all fail closed or remain Harness-only.
|
||||
|
||||
## Commit gate
|
||||
|
||||
- [x] proposal, design, specs and tasks exist.
|
||||
- [x] strict OpenSpec validation passes.
|
||||
- [x] all evidence-driven questions are resolved.
|
||||
- [x] no unresolved interface decision remains.
|
||||
- [x] `.committed` marker created for Apply.
|
||||
|
||||
## Apply result
|
||||
|
||||
- Added JSqlParser 4.6 and fail-closed `MysqlSqlValidator`.
|
||||
- Added immutable logical datasource/allowlist/limit/query-plan/raw-result models and independent `MysqlToolProperties` binding.
|
||||
- Added `JdbcMysqlReadOnlyExecutor` with read-only connection, PreparedStatement binding, query timeout, max rows, cell/result limits and Run cancellation callback.
|
||||
- Added `MysqlResultProjector` with sensitive-column redaction, row/cell/total UTF-8 bounds and `NO_EVIDENCE`.
|
||||
- Added `MysqlToolAdapter` through the existing ToolBoundary; invalid SQL is rejected before database execution.
|
||||
- Replaced `scripts/query_mysql.py` with environment-only, read-only transaction behavior and pre-connect write/metadata rejection.
|
||||
|
||||
## Apply conflicts and corrections
|
||||
|
||||
- `COUNT(*)` is represented by JSqlParser as an `AllColumns` parameter in this version; the visitor was corrected to permit only the explicit `COUNT(*)` exception.
|
||||
- JDBC metadata access cannot be used as a checked-exception stream method reference; the implementation uses an explicit column loop.
|
||||
- No OpenSpec/design conflict was found; both corrections were implementation details.
|
||||
|
||||
## Archive result
|
||||
|
||||
- Apply tasks complete and `.archive-ready` created.
|
||||
- OpenSpec archived at `openspec/changes/archive/2026-07-21-single-react-mysql-readonly-tool`.
|
||||
- Main capability specification added at `openspec/specs/mysql-readonly-tool/spec.md`.
|
||||
- Stage 4 may consume the internal adapter only after this stage is committed.
|
||||
@@ -0,0 +1,29 @@
|
||||
# Evidence: single-react-mysql-readonly-tool
|
||||
|
||||
## Static verification
|
||||
|
||||
- `git diff --check`: passed before archive.
|
||||
- Security scan confirms the query helper has no default external host/port/root user and no `commit()` write path.
|
||||
- New MySQL Harness code receives only injected logical DataSources and does not reference `spring.datasource` or the application persistence datasource.
|
||||
- OpenSpec strict validation passed for `single-react-mysql-readonly-tool`.
|
||||
|
||||
## Script/build verification
|
||||
|
||||
- `mvn -q -DskipTests compile`: passed.
|
||||
- Focused suite passed: `MysqlSqlValidatorTest`, `MysqlResultProjectorTest`, `JdbcMysqlReadOnlyExecutorTest`, `MysqlToolAdapterTest`, `MysqlToolContractTest`, `ToolBoundaryTest`, `CanonicalInvocationStoreTest`.
|
||||
- Python syntax compilation passed for `scripts/query_mysql.py`.
|
||||
- Missing connection environment variables exit before connection with code 2.
|
||||
- A write SQL invocation is rejected before connection with code 3.
|
||||
|
||||
## Security coverage
|
||||
|
||||
- Allowed: explicit allowlisted SELECT, parameter placeholders, qualified INNER JOIN and `COUNT(*)`.
|
||||
- Rejected: write, WITH, subquery, UNION, wildcard projection, unknown table/column, ambiguous column, dangerous function, inline literal, CASE, FOR UPDATE, multi-statement and placeholder mismatch.
|
||||
- JDBC controls verified: `setReadOnly(true)`, `PreparedStatement`, `setQueryTimeout`, `setMaxRows`, ordered parameter binding and cancellation-before-execution.
|
||||
- Projection controls verified: max rows, max cell chars, total UTF-8 bytes, sensitive-column redaction, valid bounded JSON and `NO_EVIDENCE`.
|
||||
|
||||
## Not verified in this stage
|
||||
|
||||
- No live production business datasource was provisioned or queried; ISS-014 explicitly assigns live E2E to a later issue/stage.
|
||||
- No public Diagnosis Agent/Chat integration was performed; stage 4 will consume the adapter internally.
|
||||
- JDBC driver timeout/cancel behavior against a real remote MySQL server remains an operational integration risk.
|
||||
@@ -0,0 +1,25 @@
|
||||
# Acceptance: single-react-rag-log-projections
|
||||
|
||||
## Commit preflight
|
||||
|
||||
- OpenSpec strict validation: passed.
|
||||
- Scope: RAG and Mock query-log projectors/adapters through the existing ToolBoundary.
|
||||
- Explicit non-goals: real CLS/MCP, MySQL, Diagnosis Agent cutover, public Chat/AIOps/SSE changes, legacy recorder cleanup.
|
||||
- Main residual risk: legacy log payloads contain sensitive values; projector tests must prove redaction before Agent serialization.
|
||||
|
||||
## Apply acceptance
|
||||
|
||||
- Implemented and verified.
|
||||
- Static verification and focused/regression Maven tests are recorded in `evidence.md`.
|
||||
- No browser/manual verification applies; this stage adds no UI or public protocol change.
|
||||
- Residual risks: live Redis/CLS integration and Agent cutover remain later stages.
|
||||
|
||||
## Archive acceptance
|
||||
|
||||
- `.archive-ready` marker is created after all tasks and checks pass.
|
||||
- OpenSpec is ready to move to the dated archive directory.
|
||||
- Archive completed at `openspec/changes/archive/2026-07-21-single-react-rag-log-projections`.
|
||||
|
||||
## Remaining work
|
||||
|
||||
- Real Redis/CLS integration, MySQL projection, Diagnosis Agent cutover and final E2E remain later stages.
|
||||
@@ -0,0 +1,24 @@
|
||||
# Brief: single-react-rag-log-projections
|
||||
|
||||
## Background
|
||||
|
||||
阶段 3A 已经统一 ToolBoundary 和 canonical invocation store。下一阶段需要把现有 RAG 与 Mock 日志工具接入该边界,并把旧工具输出投影为冻结的 Agent-facing ACI 结果。
|
||||
|
||||
## Goal
|
||||
|
||||
- 提供 bounded RAG evidence projection。
|
||||
- 提供带完整 logical scope 和 Mock provenance 的 query-log projection。
|
||||
- 复用阶段 3A 生命周期、Run ownership、tool_call_id、预算、错误和 canonical record。
|
||||
|
||||
## Non-goals
|
||||
|
||||
- 不接入真实 CLS/MCP。
|
||||
- 不实现 MySQL projection。
|
||||
- 不切换 Diagnosis Agent、Chat/AIOps、SSE 或旧 recorder。
|
||||
|
||||
## Classification
|
||||
|
||||
- Scale: complex
|
||||
- Interface impact: L2 internal Harness adapter/projector
|
||||
- Issue: ISS-014 stage 3B
|
||||
- Change slug: `single-react-rag-log-projections`
|
||||
@@ -0,0 +1,72 @@
|
||||
# Decisions: single-react-rag-log-projections
|
||||
|
||||
## Discover status
|
||||
|
||||
- Checkpoint: Discover
|
||||
- Capability source: `sm-flow` with `grill-with-docs` codebase evidence; no external service integration required.
|
||||
- Scale: complex, because two tool adapters share a lifecycle boundary and define bounded Agent-facing output semantics.
|
||||
|
||||
## Evidence-driven findings
|
||||
|
||||
1. `ToolBoundary` currently accepts `project(String rawResponse)` and already owns Run/ID/authorization/read-only/JSON/budget/lifecycle enforcement.
|
||||
2. `LookupKnowledgeTool` emits `evidenceBlocks`, `contextPack`, `retrievalTrace`, `rerankTrace`, session domains, and message; these are internal retrieval/audit fields and must not be projected.
|
||||
3. `QueryLogsTools` emits region, physical log topic, result limit, instance and metrics; the frozen contract requires logical topic/query/lookback scope and `source_kind=MOCK` instead.
|
||||
4. Existing ACI records already define the required snake_case fields and immutable collections.
|
||||
5. The request scope must be passed to the log projector through a typed adapter method rather than inferred from raw output.
|
||||
|
||||
## User-confirmed direction
|
||||
|
||||
- Use the framework-provided `tool_call_id` only.
|
||||
- Keep lifecycle status and evidence status separate.
|
||||
- Implement the stage in phases and complete sm-flow archive plus Git commit before the next stage.
|
||||
- Adopt the request-aware projector adapter for log scope preservation.
|
||||
|
||||
## Question pool
|
||||
|
||||
| Dimension | Question | Mode | Conclusion | Status |
|
||||
|---|---|---|---|---|
|
||||
| Terminology | Are RAG traces and context packs Agent evidence? | evidence-driven | No. They are internal retrieval/audit details and are excluded from projection. | resolved |
|
||||
| Boundary | How is log scope preserved when the generic projector has no request? | evidence-driven | Typed adapter carries `QueryLogsRequest` into a request-aware projector method. | resolved |
|
||||
| Provenance | Which log source is implemented now? | evidence-driven | Existing Mock source only; result always records `source_kind=MOCK`. | resolved |
|
||||
| Negative result | What does an empty query mean? | evidence-driven | `NO_EVIDENCE` for the recorded scope, with no health/problem inference. | resolved |
|
||||
| Compatibility | Should legacy tools and public paths be changed now? | evidence-driven | No. Add adapters/projectors only; cutover is later. | resolved |
|
||||
|
||||
## Risks
|
||||
|
||||
- Existing mock messages contain hostnames, pod IDs, SQL literals, and stack-like text; sanitization must happen before projection.
|
||||
- Collection limits and excerpt limits can make the Agent result incomplete; `truncated` must be explicit.
|
||||
- The generic boundary API should remain reusable for stage 3C, so request-aware behavior belongs in an adapter or specialized projector interface.
|
||||
|
||||
## Discover checkpoint
|
||||
|
||||
- Proposal created: `openspec/changes/single-react-rag-log-projections/proposal.md`
|
||||
- Context and issue evidence recorded.
|
||||
- No unresolved user-interview question remains for this bounded stage; implementation direction was explicitly accepted in the conversation.
|
||||
|
||||
## Commit audit
|
||||
|
||||
- Capability source: `sm-flow` and local OpenSpec CLI.
|
||||
- OpenSpec strict validation: passed for `single-react-rag-log-projections`.
|
||||
- Cross-artifact alignment:
|
||||
- brief goals/non-goals -> proposal scope: aligned.
|
||||
- proposal boundaries and request-aware projector decision -> design: aligned.
|
||||
- design projection bounds, redaction, scope and adapter ownership -> spec requirements: aligned.
|
||||
- spec scenarios -> tasks for limits, RAG, logs, boundary integration and verification: aligned.
|
||||
- Interface impact: L2 internal Harness adapter/projector only; no public protocol or legacy runtime cutover.
|
||||
- Preflight risks accepted: legacy payload drift fails closed; sensitive log fields are redacted; total projection budget is explicit.
|
||||
|
||||
## Commit gate
|
||||
|
||||
- [x] proposal, design, specs and tasks exist.
|
||||
- [x] strict OpenSpec validation passes.
|
||||
- [x] all evidence-driven questions are resolved.
|
||||
- [x] no unresolved interface decision remains.
|
||||
- [x] `.committed` marker created for Apply.
|
||||
|
||||
## Archive result
|
||||
|
||||
- Apply tasks complete.
|
||||
- `.archive-ready` marker created.
|
||||
- OpenSpec archived at `openspec/changes/archive/2026-07-21-single-react-rag-log-projections`.
|
||||
- Main capability specification added at `openspec/specs/rag-log-projections/spec.md`.
|
||||
- Next stage remains 3C MySQL projection; no Agent cutover is implied by this archive.
|
||||
@@ -0,0 +1,25 @@
|
||||
# Evidence: single-react-rag-log-projections
|
||||
|
||||
## Static verification
|
||||
|
||||
- `git diff --check`: passed.
|
||||
- Existing legacy files were checked and have no diff: `LookupKnowledgeTool`, `QueryLogsTools`, `ChatService`, `AiOpsService`, and `ToolInvocationRecorder`.
|
||||
- OpenSpec strict validation: passed for `single-react-rag-log-projections`.
|
||||
|
||||
## Script/build verification
|
||||
|
||||
- `mvn -q -DskipTests compile`: passed.
|
||||
- Focused projection tests: passed (`RagResultProjectorTest`, `QueryLogsResultProjectorTest`).
|
||||
- Adapter tests: passed (`ToolAdapterTest`).
|
||||
- Stage regression suite: passed (`CanonicalInvocationStoreTest`, `ToolBoundaryTest`, ACI/Core/retry/key/config/ChatController tests plus the new projection tests).
|
||||
|
||||
## Coverage
|
||||
|
||||
- RAG: bounded exact excerpts, duplicate document IDs, internal field exclusion, `NO_EVIDENCE`, truncation and framework ID.
|
||||
- Logs: logical scope, Mock provenance, pattern aggregation, timeline sampling, sensitive value redaction, empty result and bounded output.
|
||||
- Boundary: adapters use the existing canonical lifecycle and do not return raw payloads.
|
||||
|
||||
## Not verified in this stage
|
||||
|
||||
- Real Redis connectivity and live CLS/MCP integration.
|
||||
- Diagnosis Agent/application cutover and end-to-end Maven runtime flow.
|
||||
@@ -0,0 +1,43 @@
|
||||
# Acceptance: single-react-tool-invocation-store
|
||||
|
||||
## 实现结果
|
||||
|
||||
- 新增 canonical record/limits/exceptions/store interface/Redis JSON adapter。
|
||||
- 新增 ToolBoundary envelope/result、executor/projector interfaces 和稳定错误码。
|
||||
- 实现 preflight、Run/Tool budget、Run capacity、PROJECTING、READY/ERROR、evidence semantics、TTL 和 UTF-8 limits。
|
||||
- 旧 ToolInvocationRecorder、JPA、Chat/AIOps、Controller、Repository 和公开协议未修改。
|
||||
|
||||
## 静态验证
|
||||
|
||||
- `openspec validate single-react-tool-invocation-store --strict`:通过。
|
||||
- 新 Harness 包 Redis 引用仅为 `RedisCanonicalInvocationStore`。
|
||||
- legacy recorder/JPA/Chat/AIOps/Controller/repository/resources diff:为空。
|
||||
- staged diff check 与 Secret scan:提交前执行并通过。
|
||||
|
||||
## 脚本验证
|
||||
|
||||
- `mvn -q -DskipTests compile`:通过。
|
||||
- `mvn -q '-Dtest=CanonicalInvocationStoreTest,ToolBoundaryTest' test`:通过。
|
||||
- 综合阶段 0/1/2/3A suite(含 Harness/ACI/Core/retry/key/config/ChatController):通过。
|
||||
|
||||
## 浏览器/人工验证
|
||||
|
||||
- 不适用。本阶段没有 UI、Controller、SSE 或公开协议变化。
|
||||
|
||||
## 未验证
|
||||
|
||||
- 未连接真实 Redis,未做 ACL/network/TTL live 验证;最终 E2E 阶段执行。
|
||||
- 未接入 Alibaba ToolInterceptor,真实框架 ID 传播留给后续 Agent/application stage。
|
||||
- 未实现 RAG/log/MySQL projector,留给 3B/3C。
|
||||
- canonical update 的 read-TTL-write 并发窗口已记录为风险,尚未 Lua/CAS 化。
|
||||
|
||||
## 剩余风险与后续门禁
|
||||
|
||||
- 新旧 JPA audit 与 canonical store 短期并存,后续 projector 必须只以新 boundary 的 READY record 作为引用来源。
|
||||
- 下一阶段 3B/3C 必须复用本 ToolBoundary,不复制 Redis 状态机。
|
||||
|
||||
## 状态
|
||||
|
||||
- Stage acceptance: accepted
|
||||
- OpenSpec archive: archived at `openspec/changes/archive/2026-07-21-single-react-tool-invocation-store`
|
||||
- Main spec sync: `openspec/specs/canonical-tool-invocation-store/spec.md`(7 added requirements)
|
||||
@@ -0,0 +1,29 @@
|
||||
# Brief: single-react-tool-invocation-store
|
||||
|
||||
## 背景
|
||||
|
||||
旧 `ToolInvocationRecorder` 依赖 ThreadLocal 和 JPA preview,不能证明完整 Tool 结果、生命周期和当前 Run 所有权。阶段 2 已提供 RunContext/Key/Capacity,需要统一 canonical ToolBoundary 和 Redis store 供后续 projector 复用。
|
||||
|
||||
## 目标
|
||||
|
||||
- 统一 Pre-Tool 门禁、PROJECTING/READY/ERROR 状态和 evidence semantics。
|
||||
- 在同一 Redis record 保存 request/raw_response/agent_result、框架 ID、Run、时间和错误。
|
||||
- 固定 TTL 不续期、容量/结果大小 fail-closed、raw 不静默截断。
|
||||
- 用 Fake Tool/Projector/Store 覆盖 duplicate、cross-run、no-evidence、error、TTL 和 oversize。
|
||||
|
||||
## 范围
|
||||
|
||||
- Canonical invocation model/store、Redis JSON adapter、ToolBoundary 和 focused tests。
|
||||
- 复用阶段 2 Core、budget、capacity、ToolCallKeyFactory。
|
||||
|
||||
## 非目标
|
||||
|
||||
- 不实现 RAG/log/MySQL projector。
|
||||
- 不修改旧 recorder/JPA、Chat/AIOps、Controller/SSE 或公开协议。
|
||||
|
||||
## 元数据
|
||||
|
||||
- 分档:complex
|
||||
- 接口影响:L2 内部 Harness boundary/store
|
||||
- 关联 Issue:ISS-014 阶段 3A
|
||||
- 关联 OpenSpec:`openspec/changes/single-react-tool-invocation-store`
|
||||
@@ -0,0 +1,106 @@
|
||||
# Decisions: single-react-tool-invocation-store
|
||||
|
||||
## 规模与入口
|
||||
|
||||
- 分档:complex。
|
||||
- 入口:ISS-014 阶段 3A;阶段 0/1/2 已 Archive 并由 `58c3910`、`4274f33`、`6b74990` 提交。
|
||||
- 目标:统一 ToolBoundary 和 Redis canonical invocation store,不实现 Tool-specific projection。
|
||||
|
||||
## Context
|
||||
|
||||
- 阶段 1 主规格已冻结 RAG/log/MySQL Agent-facing Request/Result 和 `EvidenceStatus`。
|
||||
- 阶段 2 主规格已冻结 `RunContext`、预算、取消、Key Factory 和 strict retry。
|
||||
- 旧 `ToolInvocationRecorder` 依赖 `SessionContextHolder`、JPA `ToolInvocation` 和 500 字符 preview,属于 durable audit 兼容路径,不是 canonical store。
|
||||
- 既有 `SessionConfiguration` 提供 `RedisTemplate<String,Object>` + JSON serializer;新 store 复用 bean,不新增连接配置。
|
||||
|
||||
## Question Pool
|
||||
|
||||
| 维度 | 问题 | 模式 | 证据与结论 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| 术语 | canonical invocation 与旧 JPA ToolInvocation 是否同一记录? | evidence-driven | ISS-014 数据分层明确 Redis canonical 保存完整 request/raw/agent,JPA 只做 durable audit;两者分离。 | 已解决并汇报 |
|
||||
| 术语 | `PROJECTING/READY/ERROR` 与 evidence status 如何组合? | evidence-driven | 阶段 0/1 规格:只有 READY 可为 FOUND/NO_EVIDENCE,ERROR 不可引用;PROJECTING 是内部暂态。 | 已解决并汇报 |
|
||||
| 边界 | 阶段 3A 是否实现 RAG/log projector? | evidence-driven | ISS-014 3A 明确不实现 Tool-specific projection,3B/3C 单独接入。 | 已解决并汇报 |
|
||||
| 边界 | Redis 读取是否刷新 TTL? | evidence-driven | ISS-014 固定“创建设置、读取/更新不续期”;更新采用当前剩余 TTL,不恢复初始 TTL。 | 已解决并汇报 |
|
||||
| 验收 | raw 超限是否静默截断? | evidence-driven | ISS-014 明确 `RESULT_TOO_LARGE` ERROR,raw 不静默截断;agent projection 才可按 projector 预算截断并标记。 | 已解决并汇报 |
|
||||
| 验收 | 缺失/重复/cross-run ID 如何处理? | evidence-driven | 阶段 3A 任务明确覆盖;Key Factory 保留框架 ID,ToolBoundary 在 store 创建前校验 Run/ID 和 duplicate。 | 已解决并汇报 |
|
||||
| 技术 | Redis 如何避免 update 重置 TTL? | evidence-driven | 既有 RedisTemplate;begin 使用 setIfAbsent + TTL,update 先读取剩余 TTL 再写回相同/更短 TTL,get 不调用 expire。 | 已解决并汇报 |
|
||||
| 技术 | 是否修改旧 recorder 以复用新 store? | evidence-driven | 旧链路大量测试依赖 JPA preview/evidence_refs;本阶段零消费者,保持旧 recorder 不变,避免行为回归。 | 已解决并汇报 |
|
||||
|
||||
## Grill 结论
|
||||
|
||||
- 所有术语、边界、验收和技术问题均由 ISS、阶段规格、旧代码和 Redis 配置事实证明。
|
||||
- 没有新增产品偏好或兼容性取舍需要 user-interview;并发更新窗口作为已接受风险记录。
|
||||
- `grill-with-docs` 的代码可证结论已回写 proposal;没有未确认问题。
|
||||
|
||||
## 能力与工具限制
|
||||
|
||||
- Discover 能力来源:`sm-flow` + `grill-with-docs`。
|
||||
- 当前无 `codebase-retrieval`/LSP;使用 `rg`、源码、既有测试、本地 Redis 配置和 focused fake tests 进行等价核对。
|
||||
|
||||
## Cross-artifact 对齐
|
||||
|
||||
| 链路 | 状态 | 结论 |
|
||||
|---|---|---|
|
||||
| brief 目标/范围/非目标 -> proposal | 已对齐 | Boundary、canonical record、TTL/size、Fake tests 和 legacy isolation 全部覆盖。 |
|
||||
| proposal 范围/约束/承诺 -> design | 已对齐 | Redis value、状态机、preflight 顺序、raw/projection 和失败处理均已设计。 |
|
||||
| design 决策/接口影响/风险 -> specs/tasks | 已对齐 | L2 boundary、TTL 更新窗口、oversize/error、ID/Run 所有权有对应要求和任务。 |
|
||||
| specs 可观察行为 -> tasks | 已对齐 | 7 条 requirements 分解为 store、boundary、limits/failure 和隔离验证纵向切片。 |
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
- 能力来源:`zoom-out`,按 RunContext、Invocation Status、Evidence Status、canonical invocation 和 durable audit 术语审计。
|
||||
- ToolBoundary 只编排一次调用;CanonicalInvocationStore 独占 Redis 状态转换;DiagnosisHarnessCore 独占 Run budget/cancellation;旧 recorder 只写 JPA audit。
|
||||
- request/raw/agent result 归同一 canonical record,Agent 只获得 ToolBoundaryResult,不存在 raw 旁路。
|
||||
- Redis adapter 是唯一 Redis 访问点,接口/Fake 不依赖 Redis;后续 3B/3C 可直接复用而不复制状态机。
|
||||
- 风险集中在 read-TTL-write 并发窗口和 canonical raw 敏感性,已进入 design/spec/limits,无架构或 ADR 冲突。
|
||||
|
||||
## Commit Gate Preflight
|
||||
|
||||
- proposal、design、specs、tasks 完整,OpenSpec status complete,strict validation 通过。
|
||||
- question pool 无未汇报 evidence-driven 或未确认 user-interview 项。
|
||||
- L2 内部接口影响已记录;旧 JPA/Chat/Controller/协议不改。
|
||||
- cross-artifact 无 gap,所有错误、TTL、size、ID/Run 所有权和隔离要求可由 Fake tests 验证。
|
||||
- Apply 已获持续授权,范围只包括新 store/boundary 与 focused tests。
|
||||
|
||||
## Pre-apply Research
|
||||
|
||||
### 参考实现与复用
|
||||
|
||||
- 复用 `DiagnosisHarnessCore` 的 active/deadline/Tool budget/Run bytes 门禁。
|
||||
- 复用 `ToolCallKeyFactory` 精确保留框架 Tool Call ID 并隔离 Run key。
|
||||
- 复用阶段 0 `InvocationStatus` / `EvidenceStatus`,不创建字符串状态副本。
|
||||
- 复用 `SessionConfiguration` 提供的 `RedisTemplate<String,Object>` 和 Spring Boot ObjectMapper。
|
||||
- 旧 `ToolInvocationRecorder`/JPA preview 仅作为 durable audit 反例,本阶段不修改或调用。
|
||||
|
||||
### 技术栈清单
|
||||
|
||||
- canonical record:Java 17 record + Jackson JSON String,所有状态组合在 record transition 方法中校验。
|
||||
- Redis create:`ValueOperations.setIfAbsent` + TTL;update:读取当前 remaining TTL 后写回;get 不调用 expire。
|
||||
- 大小:UTF-8 bytes;record/agent result/store limits 与 RunContext 累计 capacity 双重门禁。
|
||||
- 测试:Mockito RedisTemplate/ValueOperations 验证 TTL API;In-memory fake store + Fake Tool/Projector 验证 boundary,不连接外部 Redis。
|
||||
|
||||
### 新建类型
|
||||
|
||||
- canonical model/limits/store/exceptions/Redis adapter。
|
||||
- ToolCallRequestEnvelope、ToolBoundaryResult、ProjectedToolResult、ToolExecutor、ToolResultProjector、ToolBoundary 和稳定 error codes。
|
||||
- Redis adapter tests 与 boundary fake tests。
|
||||
|
||||
### 影响半径
|
||||
|
||||
- 新 package 在阶段 3A 保持零现有消费者。
|
||||
- Redis 访问只允许出现在 `RedisCanonicalInvocationStore`;旧 Chat/AIOps/Controller/JPA/Recorder 不修改。
|
||||
|
||||
## Apply 结果
|
||||
|
||||
- 冲突分类:一次 focused test 断言错误(duplicate 场景应调用 `setIfAbsent` 两次)已修正并复跑通过;无规格偏离。
|
||||
- 新增 canonical record/limits/store exceptions、Redis JSON adapter、ToolBoundary envelope/result/interfaces 和 stable error codes。
|
||||
- ToolBoundary 已实现 preflight、Core budget/capacity、PROJECTING、raw/projection size、READY/ERROR 和安全返回边界;未接具体 projector。
|
||||
- 首模块对齐:store/boundary/limits/failure 与 design/tasks 全部完成;RAG/log/MySQL adapters 仍留给后续阶段。
|
||||
|
||||
## Apply 验证
|
||||
|
||||
- 编译:`mvn -q -DskipTests compile`:通过。
|
||||
- Store/Boundary focused:`mvn -q '-Dtest=CanonicalInvocationStoreTest,ToolBoundaryTest' test`:通过。
|
||||
- 综合回归:`mvn -q '-Dtest=CanonicalInvocationStoreTest,ToolBoundaryTest,HarnessContractTest,RagToolContractTest,QueryLogsToolContractTest,MysqlToolContractTest,RunContextTest,RunBudgetTest,DiagnosisHarnessCoreTest,HarnessRetryExecutorTest,ToolCallKeyFactoryTest,SpringAiRetryConfigurationTest,ChatControllerTest' test`:通过。
|
||||
- 静态隔离:新 Harness 包仅 `RedisCanonicalInvocationStore` 引用 Redis;legacy recorder/JPA/Chat/AIOps/Controller/repository/resources diff 为空。
|
||||
- OpenSpec:`openspec validate single-react-tool-invocation-store --strict`:通过。
|
||||
@@ -0,0 +1,20 @@
|
||||
# Evidence: single-react-tool-invocation-store
|
||||
|
||||
## 文档与代码证据
|
||||
|
||||
- ISS-014 阶段 3A 明确要求统一 ToolBoundary、canonical invocation、PROJECTING/READY/ERROR、TTL/容量、ID/Run 所有权,且不实现 3B/3C projector。
|
||||
- 阶段 2 已提供 `RunContext`、Run bytes capacity 和 `ToolCallKeyFactory`,本阶段直接复用。
|
||||
- 旧 `ToolInvocationRecorder` 使用 JPA preview 和 ThreadLocal fallback;新 canonical record 独立保存完整 request/raw/agent,不修改旧 recorder/JPA。
|
||||
- 既有 `SessionConfiguration` 提供 `RedisTemplate<String,Object>` JSON bean;`RedisCanonicalInvocationStore` 是新 Harness 包唯一 Redis 引用。
|
||||
|
||||
## Evidence-driven 结论
|
||||
|
||||
- `setIfAbsent` 确保同一 `runId+toolCallId` 不覆盖;读取不调用 expire;更新使用剩余 TTL。
|
||||
- canonical record transition 只允许 PROJECTING -> READY/ERROR;READY 只接受 FOUND/NO_EVIDENCE,ERROR 不可引用。
|
||||
- ToolBoundary 在执行前校验 Run、ID、JSON、授权、只读和预算;raw 只在可信 store 保存,不返回 Agent。
|
||||
- UTF-8 record/Agent limits 与 Run capacity 双门禁;raw oversize 跳过 projector,Agent oversize 不返回,均产生 RESULT_TOO_LARGE。
|
||||
- Fake store/Redis mock tests 已覆盖 duplicate、cross-run、unauthorized、writable、execution/projection error、NO_EVIDENCE、TTL、raw/agent oversize。
|
||||
|
||||
## 工具限制
|
||||
|
||||
- 当前无 `codebase-retrieval`/LSP;使用 `rg`、源码、Maven 编译、Mockito Redis API 和 in-memory fake 完成等价验证。
|
||||
@@ -31,7 +31,7 @@
|
||||
**配置信息**(已完成):
|
||||
- **MySQL**: 119.29.78.52:33306/superbiz_agent
|
||||
- 用户: root
|
||||
- 密码: !Fucker123..
|
||||
- 密码: 已从仓库移除,使用环境变量注入
|
||||
- driver: com.mysql.cj.jdbc.Driver
|
||||
- URL参数: useUnicode=true&serverTimezone=Asia/Shanghai&allowPublicKeyRetrieval=true
|
||||
- **Redis**: 119.29.78.52:6379
|
||||
@@ -165,7 +165,7 @@ openspec/changes/phase-1-infrastructure/
|
||||
|
||||
## 敏感信息(已编辑)
|
||||
|
||||
- MySQL 密码:已配置在 application.yml(`!Fucker123..`)
|
||||
- MySQL 密码:已从仓库移除,使用环境变量注入
|
||||
- Redis:无密码
|
||||
|
||||
---
|
||||
|
||||
@@ -0,0 +1,56 @@
|
||||
结合你在前几轮对话中梳理出的“过度设计”痛点,既然你已经决定回归单 Agent(ReAct)+ 强 Harness 架构,接下来你需要做一次彻底的“架构物理重构”。
|
||||
以下是为你量身定制的5步落地行动指南,按优先级从高到低执行:
|
||||
第一步:链路合并,砍掉“伪 Graph”节点
|
||||
目标:把被拆散的推理逻辑还给单一的 ReAct 循环。
|
||||
|
||||
删除节点:直接在 Graph/状态机中抹掉 Planner、Executor、Composer 以及串联它们的边。
|
||||
合并为单一 Agent:创建一个 DiagnosisAgent。在它的 System Prompt 中明确:“你需要自行规划排查路径,调用工具获取证据,并在证据充分后输出最终的结构化诊断报告。”
|
||||
保留机制:该 Agent 内部运行一个带最大步数限制的 While 循环(如 max_iterations=10),防止死循环。
|
||||
|
||||
第二步:实施“上下文卸载”与“工具清洗”(解决7次调用膨胀问题)
|
||||
目标:确保单 Agent 在多轮工具调用后,上下文依然干净,从源头切断幻觉。
|
||||
|
||||
工具输出清洗(RTK 机制):改造你的所有工具(如 queryLogs, queryDB)。在工具的 Callback 层加代码,将原始的大段返回结果(如 1000 行日志)强制过滤、聚合,只把最核心的 5-10 行 ERROR 或统计指标返回给 Agent。
|
||||
上下文卸载:如果某些原始数据必须保留,在工具返回时,将全量数据写入本地文件(如 refs/log_001.md),上下文里只注入一行:[发现5个504错误,详见 refs/log_001.md]。
|
||||
|
||||
第三步:下沉确定性逻辑,用代码替代 LLM 节点
|
||||
目标:将你之前用 Gatekeeper 和 VerifiedInput 做的事,降级为零 LLM 调用的代码拦截器。
|
||||
|
||||
前置拦截(Pre-Tool Hook):Agent 发起工具调用时,Harness 代码用 JSON Schema 校验参数格式。不合法直接报错打回,不执行工具。
|
||||
后置断言(Post-Tool Hook):Agent 输出最终诊断报告时,代码层强制校验报告中引用的 evidence_id 和 raw_path 是否真实存在于历史记录或文件系统中。不合法直接拒绝输出,发回重试。
|
||||
|
||||
第四步:锁死输出契约
|
||||
目标:防止 Executor 过度输出和发散。
|
||||
|
||||
在 System Prompt 中强制规定 Agent 的中间思考步和最终输出步必须符合严格的 JSON 结构。
|
||||
例如,中间步必须是 {"thought": "<不超过50字>", "action": "queryLogs", "parameters": {...}}。Harness 代码检查字数,超长直接打回。
|
||||
|
||||
第五步:剥离异步验证(保留你最初的“防幻觉”初衷)
|
||||
目标:在不增加主链路复杂度的前提下,保留交叉验证能力。
|
||||
|
||||
主 ReAct Agent 输出报告后,不要在主 Graph 里串行接一个 Verifier 节点。
|
||||
改为异步触发一个轻量级 LLM(或小模型),只传入“压缩后的证据摘要 + 草稿结论”。让它判断时间线与逻辑是否一致。如果不一致,在最终输出上加“低置信度警告”;如果一致,直接放行。
|
||||
|
||||
总结:你的重构后架构全景图
|
||||
重构后,你的代码结构应该极其清爽,大致如下:
|
||||
[用户输入]
|
||||
│
|
||||
▼
|
||||
[Diagnosis ReAct Agent] (唯一的 LLM 推理节点,自带规划、执行、总结)
|
||||
│
|
||||
├── Tool: queryLogs
|
||||
│ └── [Harness 代码]: 过滤 INFO,提取 ERROR,写入 refs,返回摘要
|
||||
├── Tool: queryMetrics
|
||||
│ └── [Harness 代码]: 聚合统计值,返回 3 行核心指标
|
||||
│
|
||||
▼ (Agent 输出最终 JSON 报告)
|
||||
[Output Schema Linter] (纯代码层,0 LLM)
|
||||
│
|
||||
├─ 校验失败 ──> 返回错误给 [Agent] 重新生成
|
||||
│
|
||||
▼ (校验通过)
|
||||
[Async Verifier] (异步轻量 LLM,只做时间线/逻辑一致性校验)
|
||||
│
|
||||
▼
|
||||
[最终输出 / 带警告输出]
|
||||
现在你应该做的第一件事:打开你的代码,把主链路上除 Diagnosis Agent 以外的所有 LLM 编排节点全部注释掉,然后按照上面的结构,给工具加上 Pre-Tool 和 Post-Tool 的代码拦截器。把精力从“画 Graph”转移到“写工具清洗代码”上。
|
||||
@@ -1,6 +1,6 @@
|
||||
# MVP Issues 索引
|
||||
|
||||
**更新日期**:2026-07-10
|
||||
**更新日期**:2026-07-20
|
||||
**状态**:按活跃问题、设计笔记、RAG 问题集和已归档问题整理
|
||||
|
||||
## 目录约定
|
||||
@@ -18,6 +18,9 @@
|
||||
|---|---|---|---|---|
|
||||
| ISS-003 | MVP 设计与实现 Review 收敛 | 高 | 待规划 | [active/ISS-003-mvp-design-implementation-review.md](active/ISS-003-mvp-design-implementation-review.md) |
|
||||
| ISS-004 | Executor 域级检索水位控制 | 低 | 待规划 | [active/ISS-004-executor-domain-hard-limit.md](active/ISS-004-executor-domain-hard-limit.md) |
|
||||
| ISS-012 | Executor Token 预算与上下文膨胀 | 高 | 待规划 | [active/ISS-012-executor-token-budget-and-context-growth.md](active/ISS-012-executor-token-budget-and-context-growth.md) |
|
||||
| ISS-013 | Chat 入口解耦与真正 SSE 收敛 | 高 | 待规划 | [active/ISS-013-chat-entry-decoupling-and-sse.md](active/ISS-013-chat-entry-decoupling-and-sse.md) |
|
||||
| ISS-014 | 单体 ReAct Agent、Harness 与 ACI 工具瘦身 | 高 | 待阶段 0 冻结 | [active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md](active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md) |
|
||||
| executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [active/executor-evidence-attribution-hallucination.md](active/executor-evidence-attribution-hallucination.md) |
|
||||
| rag-refactor-plan | RAG 检索重构计划 | 高 | 待规划 | [active/rag-refactor-plan.md](active/rag-refactor-plan.md) |
|
||||
|
||||
|
||||
@@ -0,0 +1,119 @@
|
||||
# ISS-012 Executor Token 预算与上下文膨胀
|
||||
|
||||
**状态**:待规划
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-20
|
||||
**关联**:ISS-002、ISS-004、ISS-011
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
ISS-011 完成 StateGraph 切换后,复杂 Chat 链路已经具备显式 Node、Gatekeeper、Verifier、Composer 和 Run 级 Trace。但最终 live E2E 暴露出 Executor 的 Token 成本和上下文增长问题:一次只返回安全 Fallback 的请求消耗了超过 11 万 Token。
|
||||
|
||||
## 现象
|
||||
|
||||
最终验收 Run:
|
||||
|
||||
- `runId`:`run-808ac38f-3ad0-4462-a6d0-ed50d8686473`
|
||||
- 总耗时:`75964ms`
|
||||
- 最终答案:109 字符,Run 为 `CHAT/SUCCESS`,`degraded=true`
|
||||
- AgentStep:8 条,其中 Planner 1 次、Executor 7 次
|
||||
- ToolInvocation:12 次
|
||||
- Run `total_token_count`:`111802`
|
||||
|
||||
Executor 每次模型调用的 Token 逐步上升:
|
||||
|
||||
```text
|
||||
6396 -> 11265 -> 12786 -> 16774 -> 17965 -> 19340 -> 24236
|
||||
```
|
||||
|
||||
工具调用包括 4 次 `lookup_knowledge`、6 次 `query_logs`、1 次 `query_metrics` 和 1 次 `get_available_log_topics`。最终因模型输出缺少 `source_invocation_id`,Gatekeeper 将结果降为 `LOW_CONFID` 并进入安全 Fallback。
|
||||
|
||||
## 已确认事实
|
||||
|
||||
1. 数据库中 8 条 `agent_step` 均为不同记录,不存在重复插入;`111802` 等于各步骤 `token_count` 的求和。
|
||||
2. `TokenTrackingChatModel` 当前只保存供应商返回的 `usage.totalTokens`,没有拆分输入、输出、缓存和推理 Token。
|
||||
3. `ChatService.backfillRunMetrics` 直接累加每条 AgentStep 的 `token_count`。
|
||||
4. `AgentLoggingHook` 只把模型输入截断为 500 字符写入审计表,无法从当前 Trace 还原模型实际发送的完整 Prompt。
|
||||
5. Executor 当前没有独立的模型调用次数、工具调用次数或 Token 预算;Graph recursion limit 不能限制 ReactAgent 内部工具循环。
|
||||
|
||||
## 初步根因假设
|
||||
|
||||
- 每次 Executor 模型调用都会重新携带 Planner 结果、历史消息和之前的工具返回,导致输入上下文随工具循环增长。
|
||||
- 工具返回内容包含较多日志、知识库结果和检索明细,完整结果被反复带入后续模型请求。
|
||||
- 工具结果没有稳定返回 `source_invocation_id`,模型无法可靠生成精确证据引用,导致高成本检索后仍然进入 Fallback。
|
||||
- 当前只能看到 `totalTokens`,尚未确认供应商 usage 中 input/output/cached/reasoning 的精确占比。
|
||||
|
||||
## 影响
|
||||
|
||||
- 单次诊断成本和延迟不可控,复杂问题可能继续超过模型上下文窗口。
|
||||
- Token 消耗与最终答案质量不匹配,出现“高成本检索 + 安全降级”的低收益路径。
|
||||
- 缺少 Token 分项指标,无法建立成本预算、P95 延迟和 degraded rate 门禁。
|
||||
- Executor 可能重复查询相同或相近的知识域、日志主题和指标。
|
||||
|
||||
## 目标
|
||||
|
||||
1. 建立按 Run/AgentStep 的 input、output、cached、reasoning Token 可观测性。
|
||||
2. 为 Executor 增加硬性模型轮数、工具调用和 Token 预算。
|
||||
3. 将完整工具结果留在 Run Trace/数据库中,模型上下文只接收有界证据投影。
|
||||
4. 让工具结果直接携带可引用的 `source_invocation_id` 和紧凑 `evidence_refs`。
|
||||
5. 在预算耗尽时安全结束并明确记录原因,不绕过 Gatekeeper、Verifier 或 Run Trace。
|
||||
|
||||
## 建议方案
|
||||
|
||||
### 1. Token 统计拆分
|
||||
|
||||
- 从 ChatModel usage 中记录 `input_tokens`、`output_tokens`、`cached_tokens`、`reasoning_tokens`(供应商提供时)。
|
||||
- 保留 `total_token_count` 作为汇总字段,但明确其计算口径。
|
||||
- 在 `orchestration_trace` 中记录每个 Agent 的累计 Token 和预算命中情况。
|
||||
|
||||
### 2. Executor 硬预算
|
||||
|
||||
初版建议从以下上限开始,并通过固定 E2E 调整:
|
||||
|
||||
- Executor 模型调用最多 4 次。
|
||||
- 工具调用最多 8 次。
|
||||
- `lookup_knowledge` 最多 2 次。
|
||||
- `query_logs` 默认最多返回 5 条日志,并限制单次输出长度。
|
||||
- 达到预算后停止扩展检索,基于已验真证据输出,或进入带原因的安全 Fallback。
|
||||
|
||||
### 3. 有界证据上下文
|
||||
|
||||
- 工具完整原始结果继续写入 `tool_invocation`,不直接作为下一轮完整上下文。
|
||||
- 返回模型的工具视图只保留 invocation ID、工具名、查询条件、有限 evidence refs、excerpt 和 no-evidence 状态。
|
||||
- 同一工具数组项禁止重复绑定;相同知识域和日志主题不重复查询。
|
||||
|
||||
### 4. 证据引用闭环
|
||||
|
||||
- 每次 evidence tool 返回结果时直接包含 `source_invocation_id`。
|
||||
- Executor 输出必须引用该 ID;Gatekeeper 不再依赖事后猜测或唯一候选补全。
|
||||
- 由于引用失败进入 Fallback 时,Trace 必须记录具体缺失字段和预算消耗。
|
||||
|
||||
## 验收标准
|
||||
|
||||
- [ ] 每个 AgentStep 可查看 input/output/total Token,供应商支持时可查看 cached/reasoning Token。
|
||||
- [ ] 固定 `payment-timeout` E2E 的 Token 上限、工具调用上限和最大延迟已定义并通过回归。
|
||||
- [ ] 连续至少 10 次相同 fixture 运行,Token 和延迟 P95 不超过定义的预算。
|
||||
- [ ] Executor 预算耗尽时只走安全 Fallback,不绕过 Gatekeeper、Verifier 或 Trace 持久化。
|
||||
- [ ] 工具返回包含真实 `source_invocation_id`;正常证据链不再因缺少该字段而无谓降级。
|
||||
- [ ] 12 个 diagnosis eval fixture、Graph workflow/node contract、Trace ownership 回归全部通过。
|
||||
- [ ] E2E 日志和数据库能按 exact `sessionId + runId` 对齐 Token、工具调用、Fallback 原因和最终状态。
|
||||
|
||||
## 非目标
|
||||
|
||||
- 不删除 Gatekeeper、Verified Input 或 Verifier。
|
||||
- 不以降低模型 `maxTokens` 代替上下文治理。
|
||||
- 不恢复 Sequential/StateGraph 双轨或旧兼容协议。
|
||||
- 不在本 Issue 中物理删除数据库中的历史 `diagnosis_session` 表。
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `src/main/java/com/superbiz/agent/hook/TokenTrackingChatModel.java`
|
||||
- `src/main/java/com/superbiz/agent/hook/AgentLoggingHook.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- `src/main/java/com/superbiz/agent/graph/diagnosis/ReactAgentDiagnosisInvoker.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
|
||||
- `src/main/resources/prompts/chat-executor-prompt.md`
|
||||
- `mvp/architecture/stategraph-runtime-architecture.md`
|
||||
- `devflow/projects/2026-07-17-chat-diagnosis-stategraph-cleanup-docs/evidence.md`
|
||||
@@ -0,0 +1,81 @@
|
||||
# ISS-013 Chat 入口解耦与真正 SSE 收敛
|
||||
|
||||
**状态**:待规划
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-20
|
||||
**关联**:ISS-011、ISS-012
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
当前 Chat 入口同时提供 `/api/chat` 和 `/api/chat_stream`。`ChatController` 不仅处理 HTTP/SSE 协议,还直接承担会话创建、历史读取与回写、模型和工具获取、执行策略调用以及异常响应组装,入口职责已经明显超出协议适配层。
|
||||
|
||||
现有 `/api/chat_stream` 会等待完整答案生成后再按固定长度切片发送,并不是真正的流式生成。同步和伪流式入口还复制了大部分业务流程,增加了维护成本和行为不一致风险。
|
||||
|
||||
## 已确认问题
|
||||
|
||||
1. Chat 的会话、历史、模型、工具、执行和结果回写功能耦合在 `ChatController` 中,HTTP 层与应用用例边界不清晰。
|
||||
2. `/api/chat` 与 `/api/chat_stream` 重复编排同一套 Chat 流程。
|
||||
3. `/api/chat_stream` 只是对完整答案做事后分块,不具备模型生成过程中的真实增量输出能力。
|
||||
4. Controller 直接获取 `ChatModel` 和 `ToolCallbackProvider`,将模型基础设施细节暴露到入口层。
|
||||
5. SSE 使用 Controller 自建的无界缓存线程池,缺少统一生命周期和容量治理。
|
||||
6. 当前入口错误响应存在 HTTP 状态、外层 `ApiResponse` 与内层 `ChatResponse` 状态不一致的问题。
|
||||
|
||||
## 目标
|
||||
|
||||
1. 将会话生命周期、历史管理、执行调用和结果回写从 Controller 分离,形成单一 Chat 应用用例入口。
|
||||
2. 只保留一个 `/api/chat` 接口,并将其协议改为真正的 SSE。
|
||||
3. SSE 在模型或诊断链路产生内容时增量发送,而不是等待完整答案后再切片。
|
||||
4. 保留 `sessionId + runId` 作为一次 Chat Run 的稳定关联契约。
|
||||
5. 在满足入口职责分离的前提下使用最少组件,不引入没有实际职责的接口、工厂或适配层。
|
||||
|
||||
## 设计约束
|
||||
|
||||
- Controller 只负责请求校验、协议转换和 SSE 生命周期,不负责选择模型、组装工具、管理历史或编排诊断流程。
|
||||
- 同一次请求只能进入一个应用用例入口,禁止同步和流式路径各自维护一套业务逻辑。
|
||||
- 真正 SSE 至少需要区分元数据、内容增量、完成和错误事件。
|
||||
- `sessionId`、`runId` 必须在内容事件之前可获得,并用于日志、数据库和 Trace 对齐。
|
||||
- 客户端断开、超时和执行失败必须显式终止后台执行并完成 Run 状态记录。
|
||||
- 不保留旧 `/api/chat_stream` 或同步 `/api/chat` 的兼容分支,直接以新协议为准。
|
||||
- 优先使用 Spring 管理的执行设施和现有服务能力,不创建无界线程池。
|
||||
|
||||
## 建议的最小边界
|
||||
|
||||
```text
|
||||
POST /api/chat (SSE)
|
||||
-> ChatController:请求与 SSE 协议
|
||||
-> Chat 应用用例:会话、Run、历史和执行生命周期
|
||||
-> 现有 Chat 执行能力:简单回答或诊断编排
|
||||
```
|
||||
|
||||
这里的“应用用例”是职责边界,不要求预先拆出多层接口。只有出现独立变化原因或明确复用需求时才增加新组件。
|
||||
|
||||
## 验收标准
|
||||
|
||||
- [ ] 对外只保留一个 `POST /api/chat`,响应类型为 `text/event-stream`。
|
||||
- [ ] 删除 `/api/chat_stream` 及同步 Chat 兼容路径。
|
||||
- [ ] 首个内容事件在完整答案生成完成前发送,禁止通过固定字符切片伪造流式输出。
|
||||
- [ ] SSE 事件包含稳定的 metadata、content、error、done 契约。
|
||||
- [ ] Controller 不再直接依赖 `ChatModel`、`ToolCallbackProvider`,也不管理会话历史和 Run 持久化。
|
||||
- [ ] 同一请求的 `sessionId + runId` 在 SSE、应用日志、`diagnosis_run`、`agent_step` 和 `tool_invocation` 中一致。
|
||||
- [ ] 客户端断开、超时、模型失败和工具失败都有明确的资源清理与 Run 终态。
|
||||
- [ ] 不存在 Controller 自建的无界线程池。
|
||||
- [ ] 单元测试覆盖入口校验和 SSE 事件契约;端到端测试验证真实增量输出、断开清理及 Trace 对齐。
|
||||
|
||||
## 非目标
|
||||
|
||||
- 不在本 Issue 中重新设计 StateGraph 节点、Gatekeeper、Verifier 或证据协议。
|
||||
- 不为未来可能出现的其他传输协议预建通用框架。
|
||||
- 不引入多套 Command、Handler、Adapter、Factory 只为形式上的分层。
|
||||
- 不保留旧同步接口或 `/api/chat_stream` 的兼容逻辑。
|
||||
- 不以“完整答案分块发送”作为 SSE 验收通过条件。
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `src/main/java/com/superbiz/agent/controller/ChatController.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/session/SessionManager.java`
|
||||
- `src/test/java/com/superbiz/agent/controller/ChatControllerTest.java`
|
||||
- `src/test/java/com/superbiz/agent/service/ChatServiceGraphIntegrationTest.java`
|
||||
- `mvp/architecture/current-mvp-architecture.md`
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1 @@
|
||||
Archive-ready after implementation and focused verification on 2026-07-21.
|
||||
@@ -0,0 +1 @@
|
||||
Committed after strict validation on 2026-07-21.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-21
|
||||
@@ -0,0 +1,89 @@
|
||||
## Context
|
||||
|
||||
当前 `LookupKnowledgeTool` 返回 `LookupResult`,包含 `ContextPack`、`RetrievalTrace`、`RerankTrace` 等内部检索信息;`QueryLogsTools` 暴露 region、TopicId、limit、Topic discovery 和旧 `success/total/logs` 结构。MySQL evidence Tool 尚不存在。阶段 0 已创建共享 `InvocationStatus` 与 `EvidenceStatus`,但尚无三类 Agent-facing DTO 或稳定 Tool 描述。
|
||||
|
||||
本地依赖验证显示,Spring AI 1.1.7 的 `AssistantMessage.ToolCall` 持有 `id`,Spring AI Alibaba 1.1.2.0 的 `ToolCallRequest` 将其暴露为 `getToolCallId()` 并允许通过 `ToolInterceptor` 包装调用;普通 Spring AI `ToolContext` 只包含调用方传入的 context 和 history,不自动提供当前 Tool Call ID。因此后续 Harness 必须以 Alibaba interceptor request 为 ID 接入边界。
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- 用 Java 类型冻结 RAG、日志和 MySQL 的最小 Agent Request/Result JSON Schema。
|
||||
- 复用统一证据状态,并保持 invocation lifecycle 与 evidence result 两套语义正交。
|
||||
- 保证所有可引用结果只携带框架 Tool Call ID,不生成替代 ID。
|
||||
- 冻结简短、面向行动的 Tool 名称和描述,禁止基础设施/审计实现泄漏。
|
||||
- 通过三个独立契约测试锁定字段、集合不可变性、描述边界和 Mock 来源标识。
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- 不修改旧 `@Tool` 方法、当前 Chat/AIOps 工具注册或返回行为。
|
||||
- 不实现 `ToolInterceptor`、Pre/Post Tool、Redis invocation store 或 ToolResultProjector。
|
||||
- 不实现真实 CLS/MCP、MySQL 连接、JSqlParser 校验、allowlist 或脱敏。
|
||||
- 不删除旧 DTO、Topic discovery 或审计服务;这些在后续切片接入和清理。
|
||||
|
||||
## Decisions
|
||||
|
||||
### 1. 契约放在独立 Harness Tool 包
|
||||
|
||||
新增类型放在 `com.superbiz.agent.harness.tool.contract`,共享状态继续引用 `com.superbiz.agent.harness.contract`。这使新 Harness 边界与旧 `dto`、旧工具内部模型明确隔离。
|
||||
|
||||
替代方案是在旧工具类内新增嵌套 DTO;这会继续把运行实现与 Agent Contract 绑定,并妨碍三个 projector 复用,故不采用。
|
||||
|
||||
### 2. Tool Call ID 是结果字段,不是 Agent 输入字段
|
||||
|
||||
三类 Request 均不包含 `tool_call_id`。后续 Harness 从 `ToolCallRequest.getToolCallId()` 取得 ID,在投影 Result 时写入;模型不能选择或覆盖该值。`EVIDENCE_FOUND` 和 `NO_EVIDENCE` 结果必须具备合法 ID;如果 Pre-Tool 失败原因就是 ID 缺失/非法,`ERROR` 可以没有可引用 ID。
|
||||
|
||||
替代方案是让模型把 ID 作为参数回传,或由 Harness 生成 UUID;两者都会形成不可信输入或第二套标识,违反阶段 0 决策。
|
||||
|
||||
### 3. Result 直接携带 evidence status,不携带 invocation status
|
||||
|
||||
`RagToolResult`、`QueryLogsToolResult` 和 `MysqlToolResult` 只包含 `EvidenceStatus`。`InvocationStatus` 继续描述 Redis canonical invocation 的 `PROJECTING/READY/ERROR` 生命周期,由阶段 3A 的存储模型承载。契约测试分别锁定两个枚举,防止把 `READY` 当成“找到证据”或把 `NO_EVIDENCE` 当成生命周期状态。
|
||||
|
||||
### 4. 使用 record 和 defensive copy
|
||||
|
||||
Request/Result 使用 Java 17 record,并用 Jackson `@JsonProperty` 固定 snake_case。所有 list 和 row map 在构造时 defensive copy,避免后续投影、存储或测试在对象创建后改变 Agent-visible 结果。
|
||||
|
||||
替代方案是沿用 Lombok mutable bean;它更贴近旧代码,但无法天然表达冻结后的值对象边界。
|
||||
|
||||
### 5. 三类最小 Schema
|
||||
|
||||
- RAG Request 只有 `query`;Result 只有查询回显、有界 evidence、计数和截断标记。
|
||||
- Logs Request 只有逻辑 `topic`、`query` 和可选 `lookback_minutes`;Result 保留 `source_kind`、完整 scope、聚合 patterns、少量 timeline events、计数和截断标记。
|
||||
- MySQL Request 只有逻辑 `data_source`、参数化 `sql` 和 `params`;Result 保留 columns、结构化 rows、计数和截断标记。
|
||||
|
||||
日志逻辑 Topic 首版冻结为 `APPLICATION`、`DATABASE_SLOW_QUERY`、`SYSTEM_EVENTS`。旧 `system-metrics` 不纳入日志 Contract;指标 Tool 不属于本 Issue。`SourceKind` 首版只有 `MOCK`,真实适配器以后在保持字段语义的前提下扩展。
|
||||
|
||||
### 6. Tool 描述作为常量契约
|
||||
|
||||
使用单一 `AgentToolContracts` 定义三个 snake_case Tool 名称和 ACI 描述。测试禁止描述中出现 Milvus、L0/L1、rerank、CLS region/TopicId、连接、凭据、topK、limit、Redis、Trace 等实现词。旧 `@Tool` 注解暂不引用这些常量,以免本阶段提前改变模型可观察行为。
|
||||
|
||||
## Module Flow
|
||||
|
||||
```text
|
||||
ReactAgent tool call
|
||||
-> [later] Alibaba ToolInterceptor obtains framework tool_call_id
|
||||
-> typed Request contract
|
||||
-> [later] data adapter + canonical persistence + ToolResultProjector
|
||||
-> typed bounded Result contract
|
||||
-> Agent observation
|
||||
```
|
||||
|
||||
阶段 1 只实现图中的 typed contracts 和描述常量。旧 `ChatService/AiOpsService -> LookupKnowledgeTool/QueryLogsTools` 调用链保持不变。
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [新旧 Contract 短期并存,容易误判迁移已完成] -> 契约放独立包,proposal/spec/tasks 明确禁止本阶段改旧工具,并用 `rg` 验证旧注册路径未切换。
|
||||
- [DTO 不能自行保证 ID 来自框架] -> spec 明确 provenance,阶段 2/3 在 `ToolInterceptor` 测试真实 `ToolCallRequest.getToolCallId()` 传播。
|
||||
- [record 只做结构冻结,不做业务校验] -> 参数范围、SQL allowlist、ID 格式和状态组合由后续 Pre-Tool/Projector validator 实现;本阶段不在构造器复制安全策略。
|
||||
- [逻辑 Topic 与旧 TopicId 需要映射] -> mapping 属于日志 adapter/projector 阶段,Agent Contract 不接受 region 或物理 TopicId。
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. 本阶段新增契约和测试,不切换消费者;回滚只需删除新增类型。
|
||||
2. 阶段 2/3A 从 Alibaba `ToolInterceptor` 接入框架 ID 和 invocation lifecycle。
|
||||
3. 阶段 3B/3C 分别让 RAG/日志/MySQL adapter 输出本契约。
|
||||
4. 阶段 6B 原子切换公开 Chat;旧 Contract 最终由阶段 7 清理。
|
||||
|
||||
## Open Questions
|
||||
|
||||
无。真实日志 adapter 的 `SourceKind` 扩展和值映射在后续接入 change 决策,不阻塞本阶段。
|
||||
@@ -0,0 +1,43 @@
|
||||
## Why
|
||||
|
||||
现有 `lookup_knowledge` 和 `query_logs` 将检索实现、基础设施参数、审计数据和不一致的成功语义暴露给 Agent;新增 MySQL Tool 也缺少可复用的 Agent-facing 类型。进入 Harness 和投影实现前,需要先把三类 evidence Tool 的最小输入、稳定有界输出、状态语义与框架 Tool Call 引用冻结为代码契约,避免后续阶段继续依赖字符串和旧 DTO。
|
||||
|
||||
## What Changes
|
||||
|
||||
- 新增 RAG、日志和只读 MySQL 三类 Agent-facing Request/Result 契约,JSON 字段严格使用 ISS-014 已确认的 snake_case Schema。
|
||||
- 三类结果统一复用 `EvidenceStatus`,并携带框架提供的 `tool_call_id`;契约代码不生成、替换或推导第二套调用 ID。
|
||||
- 冻结 `InvocationStatus=PROJECTING/READY/ERROR` 与 `EvidenceStatus=EVIDENCE_FOUND/NO_EVIDENCE/ERROR` 的独立语义,并通过测试禁止混用。
|
||||
- 冻结三类 Tool 的短名称与 ACI 描述,描述只说明用途、输入和禁用场景,不泄露 L0/L1、Milvus、CLS region/TopicId、连接、凭据、topK、limit、rerank 或审计实现。
|
||||
- 冻结日志 `source_kind=MOCK` 及逻辑 Topic 边界,为后续真实适配器保留同一 Contract;本阶段不新增 CLS/MCP 适配器。
|
||||
- 添加三类独立契约测试,覆盖序列化字段、不可变集合、状态与描述边界。
|
||||
- 本阶段不修改旧 Tool 的执行签名、返回值或 Chat/AIOps 注册路径,不实现 ResultProjector、Harness invocation store、MySQL SQL 校验或数据库访问。
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `aci-evidence-tool-contracts`: 定义 RAG、日志和 MySQL evidence Tool 的 Agent-facing ACI Schema、状态语义、框架调用引用和描述边界。
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- None. 当前公开运行链路仍使用旧 Tool Contract;新契约将在后续 Harness/Tool 投影和 SSE 切换 change 中接入。
|
||||
|
||||
## Context Constraints
|
||||
|
||||
- `tool_call_id` 的真理源是 Spring AI Alibaba `ToolCallRequest.getToolCallId()`;普通 Spring AI `ToolContext` 不保证包含该 ID。
|
||||
- `NO_EVIDENCE` 只表示当前查询范围没有匹配证据,只能支持 `NEGATIVE_OBSERVATION`,不能表达工具失败或系统健康。
|
||||
- 生命周期 `status` 属于 canonical invocation,`evidence_status` 属于 Agent-facing 查询结果,两者不得互相替代。
|
||||
- Agent 不可控制日志 region/TopicId/limit、RAG topK/filter 或 MySQL 连接与资源上限。
|
||||
- Mock 日志必须显式保留 `source_kind=MOCK`,不得被后续 Harness 表述为生产实时事实。
|
||||
|
||||
## Interface Impact
|
||||
|
||||
- 等级:L2(内部接口,前置冻结)。本 change 新增未来 Harness 内部使用的 Agent-facing DTO 和描述常量,所有消费者均在 ISS-014 后续实施范围内。
|
||||
- 当前运行中的旧 Tool 方法、Controller、ChatService 和 AiOpsService 不切换,因此本阶段没有对外可观察行为变化。
|
||||
- 阶段 6B 切换公开 `/api/chat` 时属于独立的 L4 破坏性变更,必须使用该阶段自己的迁移和回滚规格。
|
||||
|
||||
## Risks
|
||||
|
||||
- 仅有 DTO 不能证明框架 ID 已贯穿执行;阶段 2/3 必须在 Alibaba `ToolInterceptor` 边界接收并校验 `ToolCallRequest.getToolCallId()`。
|
||||
- 新旧 Contract 会短期并存;旧运行工具不得被误认为已符合新 ACI 输出,真正接入留给 RAG/日志投影和 MySQL Tool 阶段。
|
||||
- Provider 侧旧凭据轮换仍是外部安全前置,不因本阶段契约完成而视为关闭。
|
||||
+70
@@ -0,0 +1,70 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Evidence result semantics SHALL be distinct from invocation lifecycle
|
||||
The system SHALL expose `EVIDENCE_FOUND`, `NO_EVIDENCE`, and `ERROR` as Agent-facing evidence result semantics while retaining `PROJECTING`, `READY`, and `ERROR` only for canonical invocation lifecycle. `NO_EVIDENCE` SHALL mean that the executed query found no evidence within its recorded scope and SHALL NOT mean tool failure or system health.
|
||||
|
||||
#### Scenario: Successful empty query result
|
||||
- **WHEN** an evidence Tool completes successfully with no matching evidence in its recorded scope
|
||||
- **THEN** its Agent-facing result uses `evidence_status=NO_EVIDENCE` and the invocation lifecycle may independently reach `status=READY`
|
||||
|
||||
#### Scenario: Tool execution fails
|
||||
- **WHEN** schema, authorization, execution, or projection fails
|
||||
- **THEN** the Agent-facing result uses `evidence_status=ERROR` and SHALL NOT report `NO_EVIDENCE`
|
||||
|
||||
### Requirement: Referencable Tool results SHALL use the framework Tool Call ID
|
||||
Each `EVIDENCE_FOUND` or `NO_EVIDENCE` result SHALL contain the non-blank `tool_call_id` supplied by the framework Tool Call request. Agent inputs SHALL NOT contain `tool_call_id`, and contract code SHALL NOT generate, replace, or derive a second call ID.
|
||||
|
||||
#### Scenario: Framework requests a Tool call
|
||||
- **WHEN** Spring AI Alibaba exposes an `AssistantMessage.ToolCall.id` through `ToolCallRequest.getToolCallId()`
|
||||
- **THEN** the bounded Agent result carries exactly that ID as `tool_call_id`
|
||||
|
||||
#### Scenario: Framework ID is invalid
|
||||
- **WHEN** the Tool request has a missing or invalid framework Tool Call ID
|
||||
- **THEN** the call returns an `ERROR` result without inventing a referencable ID
|
||||
|
||||
### Requirement: RAG Tool contract SHALL expose only bounded document evidence
|
||||
The RAG Request SHALL contain only `query`. The RAG Result SHALL contain `evidence_status`, `tool_call_id`, `query`, bounded `evidence`, `returned_count`, and `truncated`; each evidence item SHALL contain only `document_id`, `source`, `title`, `breadcrumb`, and an exact `excerpt`.
|
||||
|
||||
#### Scenario: RAG evidence is serialized
|
||||
- **WHEN** a RAG result contains a matching document excerpt
|
||||
- **THEN** its JSON matches the frozen snake_case fields and excludes ContextPack, RetrievalTrace, RerankTrace, raw scores, fallback attempts, metadata, and full document bodies
|
||||
|
||||
#### Scenario: RAG query has no evidence
|
||||
- **WHEN** RAG executes successfully without a usable document excerpt
|
||||
- **THEN** it returns `NO_EVIDENCE`, preserves the original query and framework Tool Call ID, and returns an empty evidence list
|
||||
|
||||
### Requirement: Log Tool contract SHALL use logical scope and retain Mock provenance
|
||||
The log Request SHALL contain logical `topic`, `query`, and optional `lookback_minutes` only. The log Result SHALL contain `evidence_status`, `tool_call_id`, `source_kind`, complete query `scope`, `match_count`, `returned_count`, bounded `patterns`, bounded timeline `events`, and `truncated`. The initial logical topics SHALL be `APPLICATION`, `DATABASE_SLOW_QUERY`, and `SYSTEM_EVENTS`, and the initial source kind SHALL be `MOCK`.
|
||||
|
||||
#### Scenario: Mock log evidence is serialized
|
||||
- **WHEN** the existing Mock source returns matching application events
|
||||
- **THEN** the Agent result records `source_kind=MOCK`, the logical query scope, aggregate patterns, bounded timeline events, and distinct match and returned counts
|
||||
|
||||
#### Scenario: Agent creates a log request
|
||||
- **WHEN** the Agent requests log evidence
|
||||
- **THEN** it selects a logical topic and lookback window without supplying region, physical TopicId, credentials, or result limit and without calling a Topic discovery Tool first
|
||||
|
||||
### Requirement: MySQL Tool contract SHALL expose a logical read-only query interface
|
||||
The MySQL Request SHALL contain only logical `data_source`, parameterized `sql`, and `params`. The MySQL Result SHALL contain `evidence_status`, `tool_call_id`, `columns`, bounded structured `rows`, `returned_count`, and `truncated`; it SHALL NOT expose connection details, credentials, internal stack traces, or resource-limit controls.
|
||||
|
||||
#### Scenario: MySQL evidence is serialized
|
||||
- **WHEN** a future read-only adapter returns authorized rows
|
||||
- **THEN** the Agent result preserves column order, structured row values, the framework Tool Call ID, returned count, and truncation state using the frozen JSON fields
|
||||
|
||||
#### Scenario: Agent creates a MySQL request
|
||||
- **WHEN** the Agent requests business database evidence
|
||||
- **THEN** it supplies a logical data source, SQL placeholders, and parameter values without supplying JDBC connection information or security policy
|
||||
|
||||
### Requirement: Tool descriptions SHALL be concise and implementation-neutral
|
||||
The system SHALL define stable snake_case names and concise descriptions for `lookup_knowledge`, `query_logs`, and `query_mysql`. Each description SHALL state when to call the Tool, its minimal input, and what it cannot query, and SHALL NOT describe retrieval internals, infrastructure coordinates, credentials, audit storage, retries, result limits, or ranking implementation.
|
||||
|
||||
#### Scenario: Agent receives Tool definitions
|
||||
- **WHEN** a future Agent adapter registers the frozen Tool definitions
|
||||
- **THEN** the definitions describe available actions and boundaries without exposing Milvus, L0/L1, rerank, CLS region/TopicId, Redis, JDBC credentials, topK, limit, or Trace internals
|
||||
|
||||
### Requirement: Contract freeze SHALL NOT cut over the current runtime
|
||||
This change SHALL add contract types, descriptions, and tests without changing the current `LookupKnowledgeTool`, `QueryLogsTools`, ChatService, AiOpsService, Controller, Tool registration, or Agent-visible runtime results.
|
||||
|
||||
#### Scenario: Stage 1 tests pass
|
||||
- **WHEN** all ACI contract tests pass
|
||||
- **THEN** the current public Chat and AIOps paths still execute the old Tool implementations until their later projector and cutover changes
|
||||
@@ -0,0 +1,26 @@
|
||||
## 1. Shared ACI Contract
|
||||
|
||||
- [x] 1.1 Add stable snake_case Tool names and concise implementation-neutral descriptions for RAG, logs, and MySQL.
|
||||
- [x] 1.2 Reuse and lock the independent invocation lifecycle and evidence result enums without adding a second Tool Call ID type.
|
||||
- [x] 1.3 Add a shared defensive-copy helper for immutable Agent-facing list and row values.
|
||||
|
||||
## 2. RAG Contract
|
||||
|
||||
- [x] 2.1 Implement the RAG Request, Result, and document evidence records with the frozen snake_case fields.
|
||||
- [x] 2.2 Add an independent RAG contract test covering exact JSON fields, framework ID preservation, immutability, and description leakage.
|
||||
|
||||
## 3. Log Contract
|
||||
|
||||
- [x] 3.1 Implement logical log Topic, source kind, Request, Result, Scope, Pattern, and Event records.
|
||||
- [x] 3.2 Add an independent log contract test covering Mock provenance, scope/count semantics, exact JSON fields, immutability, and excluded infrastructure inputs.
|
||||
|
||||
## 4. MySQL Contract
|
||||
|
||||
- [x] 4.1 Implement the logical MySQL Request and bounded Result records with immutable params, columns, and rows.
|
||||
- [x] 4.2 Add an independent MySQL contract test covering parameterized input, structured output, exact JSON fields, deep immutability, and secret/connection exclusion.
|
||||
|
||||
## 5. Verification
|
||||
|
||||
- [x] 5.1 Run the three ACI contract test classes together and confirm lifecycle/evidence status separation.
|
||||
- [x] 5.2 Run the stage 0 Harness contract and existing RAG/log focused tests to prove the old runtime path remains unchanged.
|
||||
- [x] 5.3 Verify references and diff scope show no changes to old Tool methods, Chat/AIOps registration, Controller, persistence, datasource, or public protocol.
|
||||
@@ -0,0 +1 @@
|
||||
Devflow archive prepared and stage verification passed on 2026-07-21.
|
||||
@@ -0,0 +1 @@
|
||||
Committed after strict validation on 2026-07-21.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-21
|
||||
@@ -0,0 +1,29 @@
|
||||
# Stage 0 Acceptance Evidence
|
||||
|
||||
## Static Verification
|
||||
|
||||
- Contract package is dependency-free from Redis, JPA, Controller, Agent state and existing Hook classes.
|
||||
- Repository secret scan covers tracked worktree files and reports no known plaintext credential matches.
|
||||
- `scripts/query_mysql.py` requires `SUPERBIZ_MYSQL_PASSWORD` and exits before connecting when it is absent.
|
||||
- Spring AI 1.1.7 `SpringAiRetryProperties` bytecode shows a default `maxAttempts` value of 10; stage 2 must set underlying retries to one attempt and keep retry ownership in Harness.
|
||||
|
||||
## Script Verification
|
||||
|
||||
- `mvn -q '-Dtest=HarnessContractTest' test` - passed.
|
||||
- `mvn -q '-Dtest=ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,ChatControllerTest' test` - passed.
|
||||
- `mvn -q '-Dtest=HarnessContractTest,ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,ChatControllerTest' test` - passed.
|
||||
- `openspec validate single-react-design-freeze --strict` - passed.
|
||||
- `git diff --check` - passed; only existing Windows line-ending warnings were reported.
|
||||
- Secret scan for known committed key/password patterns - zero matches.
|
||||
- `python scripts/query_mysql.py "SELECT 1"` without `SUPERBIZ_MYSQL_PASSWORD` - exited before connecting with the expected missing-variable error.
|
||||
|
||||
## Runtime Behavior
|
||||
|
||||
- Public Chat runtime was not switched in stage zero.
|
||||
- No live model, Redis, MySQL or Milvus E2E was run; full live E2E remains stage 7 scope.
|
||||
|
||||
## External Security Prerequisite
|
||||
|
||||
- Plaintext credentials previously present in the repository must be rotated in their respective MySQL, Redis, DeepSeek, SiliconFlow and Milvus systems by the credential owner.
|
||||
- Repository changes can prove removal but cannot prove provider-side rotation.
|
||||
- Stage 3C and stage 7 must not claim live security/E2E acceptance until required environment variables contain rotated credentials.
|
||||
@@ -0,0 +1,25 @@
|
||||
# Brief: single-react-design-freeze
|
||||
|
||||
## Background
|
||||
|
||||
ISS-014 将当前 Chat 多 Agent/Hook/ThreadLocal 主链路重构为一个 Diagnosis ReAct Agent、一个确定性 Harness 和一个隔离 SemanticGuard。阶段 0 先冻结后续 10 个实施 change 共同依赖的契约和安全边界。
|
||||
|
||||
## Goal
|
||||
|
||||
产出可执行、可测试、可归档的 contract types、失败语义、安全前置和阶段门禁,同时保持现有公开 Chat 运行行为不变。
|
||||
|
||||
## Scope
|
||||
|
||||
- 类型化 Draft、Knowledge Answer、Fallback、previous turn 和状态枚举。
|
||||
- Tool ID、双状态、取消、重试、Redis canonical record 和 MySQL 安全设计冻结。
|
||||
- 明文脚本凭据清理和 focused baseline。
|
||||
- ISS-014、OpenSpec、devflow 对齐。
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- 不实现或接入新 Harness/Agent/Guard。
|
||||
- 不切换 `/api/chat`、不删除旧链路、不运行 live E2E。
|
||||
|
||||
## Source PRD
|
||||
|
||||
`mvp/issues/active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md` 是本 change 的完整 PRD 和总设计来源,不复制为第二份 PRD。
|
||||
@@ -0,0 +1,80 @@
|
||||
# Decisions
|
||||
|
||||
## Entry
|
||||
|
||||
- Parent issue: `ISS-014`
|
||||
- Change: `single-react-design-freeze`
|
||||
- Scale: `complex`
|
||||
- Interface impact: future L4; this change freezes contracts without switching runtime behavior.
|
||||
- Capability sources: sm-flow built-in clarify/context/propose, `grill-with-docs`, `openspec-propose`, `zoom-out`, `openspec-apply-change`, `openspec-archive-change`.
|
||||
|
||||
## Context Evidence
|
||||
|
||||
- `session-run-trace-isolation` established Chat Session and Diagnosis Run as separate lifecycles and made `runId` the Trace ownership key.
|
||||
- `verifier-evidence-reference-fidelity` established that no-evidence is a scoped negative observation, not proof that a problem does not exist.
|
||||
- `executor-composer-final-answer` established deterministic safe fallback boundaries and prohibited unfiltered raw output from reaching users.
|
||||
- `modular-rag-pipeline` established that retrieval trace and context packing are audit details rather than direct facts.
|
||||
- Current Spring AI Alibaba `ToolCallRequest` already provides `tool_call_id`; Harness must validate and persist it rather than create a second identity.
|
||||
- Current Spring AI `ChatModel.call(Prompt)` has no cancellation token, so cancellation must be expressed as layered, observable semantics rather than an unsupported absolute guarantee.
|
||||
|
||||
## Question Pool
|
||||
|
||||
| ID | Dimension | Question | Mode | Status |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | Terminology | Does `tool_call_id` use the framework ID or a Harness-generated ID? | user-interview | confirmed: framework ID |
|
||||
| Q2 | Terminology | Are invocation lifecycle and evidence outcome separate fields? | user-interview | confirmed: `status` + `evidence_status` |
|
||||
| Q3 | Boundary | Is ISS-014 one umbrella Issue with independent OpenSpec changes? | user-interview | confirmed: one Issue, 11 changes |
|
||||
| Q4 | Boundary | May stage 4 publish before Guards exist? | user-interview | confirmed: no; public cutover only in 6B |
|
||||
| Q5 | Lifecycle | What is the durable source of truth for `previous_turn` and `last_intent`? | user-interview | confirmed: `diagnosis_run` safe published result |
|
||||
| Q6 | Contract | How is KNOWLEDGE_QUERY citation validation represented? | evidence-driven | resolved: structured answer items with exact RAG bindings |
|
||||
| Q7 | Cancellation | What cancellation guarantees are technically enforceable? | evidence-driven | resolved: layered cancellation, no false hard-cancel claim |
|
||||
| Q8 | Acceptance | Are Apply, Archive and phase Git commit pre-authorized? | user-interview | confirmed: yes, for all phases |
|
||||
|
||||
## Confirmed Decisions
|
||||
|
||||
- `tool_call_id` is the framework Tool Call protocol ID. Harness validates non-empty, bounded, safe characters and Run-local uniqueness; duplicate/invalid/missing IDs fail closed.
|
||||
- Redis `status=PROJECTING/READY/ERROR` represents invocation/projector lifecycle.
|
||||
- `evidence_status=EVIDENCE_FOUND/NO_EVIDENCE/ERROR` represents result semantics.
|
||||
- `NO_EVIDENCE` may only support `NEGATIVE_OBSERVATION` within the exact query scope; it cannot prove absence, exclusion or health.
|
||||
- Stage 4 and 6A remain internal. Stage 6B performs the only public Chat cutover after stage 5 release gates pass.
|
||||
- ISS-014 remains the umbrella Issue. Eleven independent changes run serially; each must complete sm-flow, OpenSpec archive and Git commit before the next starts.
|
||||
- Full live E2E is deferred to stage 7; earlier stages run focused verification proportional to their change.
|
||||
- KNOWLEDGE_QUERY uses a dedicated structured draft: each answer item binds the single lookup `tool_call_id` and one or more returned `document_id` values; Harness validates exact membership before rendering `answer + references + limitations`. It does not reuse the Diagnosis Analysis schema and does not enter SemanticGuard in the first version.
|
||||
- Cancellation is layered: mark cancellation requested, prevent new model/Tool rounds and any final Draft release, invoke framework interruption, cancel owned Tool/JDBC work where supported, and rely on configured HTTP timeouts for an already-blocking synchronous model call. Run finalization is atomic and late results are discarded.
|
||||
- Public SSE `done` is not emitted after the client has disconnected; internal Run state still reaches `CANCELLED`.
|
||||
- `diagnosis_run` is the durable source for `intent`, `release_outcome` and `published_result`. Only the latest same-session `DIAGNOSIS + SUCCESS` record with a non-null safe published result may become `previous_turn`; `FALLBACK/FAILED/CANCELLED` remain auditable but are excluded.
|
||||
- `published_result` stores only `user_query/published_conclusion/scope/limitations/source_documents`; it excludes Tool Call IDs, raw evidence, full Draft and SemanticGuard audit reasons.
|
||||
|
||||
## Evidence-Driven Findings To Report
|
||||
|
||||
- The framework already exposes `ToolInterceptor`, structured output types, tool execution timeout, model/tool call limit hooks and `ReactAgent.interrupt`; later Harness stages should reuse these extension points.
|
||||
- Spring AI model dependencies include retry support, while current application configuration does not explicitly freeze all retry layers; stage 0 must define a retry inventory and stage 2 must enforce it.
|
||||
- Current `DiagnosisRun` stores a text answer and generic status but has no explicit `intent`, `release_outcome` or structured published result; Q5 must be resolved before the previous-turn contract is executable.
|
||||
- Current KNOWLEDGE_QUERY target behavior promises citation validation, but the issue only defines the Diagnosis Draft binding schema; the committed spec must add the dedicated answer-item contract described above.
|
||||
- Spring AI retry auto-configuration defaults `maxAttempts` to 10. The target Harness retry matrix requires underlying model/HTTP retries to be set to one attempt, with Router and SemanticGuard retries performed only by Harness.
|
||||
|
||||
## OpenSpec Backfill
|
||||
|
||||
- All confirmed decisions above must appear in design/specs/tasks before `.committed` is created.
|
||||
- All user-interview questions are confirmed; no pending decision blocks Commit.
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
Current input flows from `ChatController` into `ChatService`, which owns routing, ReactAgent construction, multi-Agent orchestration and final rendering; tools persist evidence through `ToolInvocationRecorder`, while Hooks and ThreadLocal bridge Run and verifier state. Stage zero introduces only dependency-free contract types under `harness.contract`; those types must not depend on Controller, Redis, JPA, Spring Agent state or current Hook classes. Later stages move ownership in order: RunContext, invocation store, Tool-specific projection, Diagnosis Agent, Guards, application use case and finally the public SSE adapter. `diagnosis_run` remains durable Run ownership, Redis canonical invocation remains short-lived Harness ownership, and `agent_step/tool_invocation` remain durable audit detail. The principal risk is spec/runtime drift, mitigated by archiving only the stage-zero contract capability now and delaying modifications to existing runtime capabilities until their implementation changes.
|
||||
|
||||
## Cross-Artifact Alignment
|
||||
|
||||
| Chain | Status | Evidence |
|
||||
|---|---|---|
|
||||
| ISS-014/brief goals, scope and non-goals → proposal | aligned | Proposal limits stage zero to contracts, security and baseline with no public cutover. |
|
||||
| proposal commitments → design | aligned | Design records every ID, status, Draft, fallback, previous-turn, retry, cancellation and phase-gate commitment. |
|
||||
| design decisions → specs | aligned | The single stage-zero capability has testable requirements for every stable contract boundary. |
|
||||
| specs observable behavior → tasks | aligned | Tasks create reusable types/tests, remove the secret, align artifacts and verify without switching runtime behavior. |
|
||||
|
||||
## Commit Gate Result
|
||||
|
||||
- Question pool covers terminology, boundary, lifecycle, contract, cancellation and acceptance.
|
||||
- All user-interview items are explicitly confirmed.
|
||||
- Evidence-driven conclusions were reported and written into design/spec/tasks.
|
||||
- Interface impact is recorded as future L4; this change itself does not switch the public API.
|
||||
- No devflow/OpenSpec conflict remains.
|
||||
@@ -0,0 +1,77 @@
|
||||
## Context
|
||||
|
||||
ISS-014 是一次 L4 Chat 重构的总设计来源,但实施被拆成 11 个必须串行归档的 OpenSpec changes。阶段 0 不切换公开协议或 Agent 运行链,只创建后续阶段复用的类型化契约、安全前置、失败语义和 focused baseline。
|
||||
|
||||
现有代码已经具备 `runId` Trace、工具调用审计、no-evidence 精确引用、Verifier/Composer fallback 和模块化 RAG,但这些能力分散在 `ChatService`、Hook、ThreadLocal、Tool 和 JSON 字符串中。Spring AI Alibaba 已提供 `ToolInterceptor`、结构化输出类型、工具执行超时、调用限额 Hook 和 `ReactAgent.interrupt`;Spring AI 底层 retry 默认最多 10 次,不能直接满足 ISS-014 的显式 Harness 重试矩阵。
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- 生成后续阶段可直接复用的 Java contract types 和枚举,不实现新 Agent 执行链。
|
||||
- 冻结 Tool ID、双状态、Draft、Knowledge Answer、Fallback、previous turn、SSE、重试和取消语义。
|
||||
- 冻结 MySQL fail-closed 允许子集和安全前置。
|
||||
- 移除仓库脚本和主配置中的明文凭据并建立改造前 focused baseline。
|
||||
- 保证 ISS-014、OpenSpec、devflow 术语和 11 个阶段门禁一致。
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- 不接入 Harness、Diagnosis Agent、EvidenceGuard 或 SemanticGuard 运行时。
|
||||
- 不修改 Controller 协议、旧 ChatService 行为、Redis invocation store 或数据库表。
|
||||
- 不实现 RAG/日志/MySQL Tool 投影。
|
||||
- 不运行完整 live E2E。
|
||||
|
||||
## Decisions
|
||||
|
||||
### Contract types are reusable runtime inputs
|
||||
|
||||
阶段 0 创建位于 `com.superbiz.agent.harness.contract` 的轻量 record/enum,而不是只写文档或引入 JSON Schema 引擎。后续 ReactAgent `outputType`、Harness validator、持久化和 SSE DTO 可以直接复用这些类型,减少同一字段在多个阶段重复定义。
|
||||
|
||||
### Framework Tool Call ID is canonical
|
||||
|
||||
`tool_call_id` 使用框架协议 ID。Harness 后续只校验非空、长度/字符安全和 Run 内唯一性,不生成第二套 ID。Redis Key 仍按 `runId + toolCallId` 隔离,真实性来自当前 Run 的 canonical record,而不是 ID 本身。
|
||||
|
||||
### Invocation lifecycle and evidence outcome are orthogonal
|
||||
|
||||
`InvocationStatus=PROJECTING/READY/ERROR` 只描述调用与投影生命周期;`EvidenceStatus=EVIDENCE_FOUND/NO_EVIDENCE/ERROR` 描述结果语义。`NO_EVIDENCE` 只允许绑定 `AnalysisKind=NEGATIVE_OBSERVATION`,并必须保留查询范围和零匹配信息。
|
||||
|
||||
### Diagnosis and knowledge answer contracts stay separate
|
||||
|
||||
Diagnosis Draft 使用 Analysis ID 与 Tool Call IDs;KNOWLEDGE_QUERY 使用 answer items,每项绑定一次 lookup 的 Tool Call ID 和返回的 document IDs。Harness 后续验证精确成员关系并确定性展开引用。首版 KNOWLEDGE_QUERY 不进入 SemanticGuard。
|
||||
|
||||
### Safe published context belongs to Diagnosis Run
|
||||
|
||||
`diagnosis_run` 是 `intent/release_outcome/published_result` 的持久化真理源。只有同 Session 最近一个 `DIAGNOSIS + SUCCESS` 且 published result 非空的 Run 可形成 previous turn。Published result 只含用户查询、已发布结论、范围、限制和 RAG 文档元数据。
|
||||
|
||||
### Cancellation is observable and layered
|
||||
|
||||
取消请求立即阻止新模型/Tool 轮次和最终 Draft 释放;框架中断、可控 Future/JDBC 取消尽力执行;已进入同步 `ChatModel.call` 的请求依靠底层 HTTP timeout。Run 终态通过原子状态转换保证唯一,晚到结果被丢弃,不宣称无法证明的底层硬取消。
|
||||
|
||||
### Harness owns all retries
|
||||
|
||||
底层 SDK/HTTP/数据库 retry 必须关闭或压为一次 attempt。Intent Router 和 SemanticGuard 的第二次 attempt 由 Harness 显式执行并审计。阶段 0 只冻结矩阵;阶段 2 实现执行器和配置。
|
||||
|
||||
### Stage boundaries are release boundaries
|
||||
|
||||
阶段 4 和 6A 仅内部运行;阶段 6B 在 Guards 已完成后原子切换公开入口。每个 change 必须 Archive 并 Git commit 后才能进入下一项,阶段 7 才执行统一 live E2E。
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Contract records过早绑定实现细节 → Mitigation:只包含跨阶段稳定字段,不包含 Redis、JPA 或框架对象。
|
||||
- [Risk] Spring AI provider 对 Tool Call ID 行为不同 → Mitigation:阶段 3A 使用 Fake Model 和当前 DeepSeek 路径验证,缺失或重复时 fail closed。
|
||||
- [Risk] 同步模型调用不能立即取消 → Mitigation:明确 layered semantics、HTTP timeout 和晚到结果丢弃,不把状态更新等同底层资源已终止。
|
||||
- [Risk] 设计冻结测试增加维护成本 → Mitigation:只保留 focused serialization/validation tests,不复制完整 E2E。
|
||||
- [Risk] 明文凭据可能已泄露 → Mitigation:仓库中移除并要求外部轮换;轮换证据记录在 acceptance。
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. 创建 contract types、契约测试和阶段台账,不接运行链。
|
||||
2. 清理查询脚本凭据,改为环境变量注入。
|
||||
3. 记录 focused baseline;Archive 本 change。
|
||||
4. 后续 10 个 change 逐步实现,并在各自 Archive 时同步对应运行 capability specs。
|
||||
|
||||
Rollback:阶段 0 没有公开行为变化;可删除新增 contract package/tests 并恢复文档。已经轮换的凭据不得回滚为旧值。
|
||||
|
||||
## Open Questions
|
||||
|
||||
无。所有影响实现的用户决策已经在 `decisions.md` 中确认。
|
||||
@@ -0,0 +1,32 @@
|
||||
## Why
|
||||
|
||||
当前 Chat 诊断把一个 ReAct 生命周期拆成多个 Agent、Hook、ThreadLocal 和重试分支,导致证据契约、失败语义、上下文预算和公开释放边界分散。进入分阶段重构前,需要先把 ISS-014 的跨阶段契约、安全前置和验收基线冻结为唯一可执行规格,避免后续 change 各自解释同一概念。
|
||||
|
||||
## What Changes
|
||||
|
||||
- 冻结单体 Diagnosis ReAct Agent、确定性 Harness、EvidenceGuard 和隔离 SemanticGuard 的职责边界。
|
||||
- 冻结 Diagnosis Draft、Analysis、Conclusion、Fallback、Intent Router 和 SSE 事件契约。
|
||||
- 冻结 Tool Call ID、调用生命周期 `status`、结果语义 `evidence_status`、Redis canonical invocation 和 ToolResultProjector 命名。
|
||||
- 冻结预算、取消、重试、隐藏重试禁用和 Run 终态语义,但不在本阶段实现新运行链路。
|
||||
- 冻结只读 MySQL Tool 的 JSqlParser 允许子集、静态 allowlist 和安全前置。
|
||||
- 移除仓库脚本和主配置中的明文数据库、Redis、模型及向量服务凭据,并记录必须完成外部轮换。
|
||||
- 建立改造前 focused baseline 和后续 11 个串行 sm-flow change 的阶段台账。
|
||||
- **BREAKING(后续阶段实施)**:最终仅保留 `POST /api/chat` SSE、移除旧多 Agent/Graph Chat 主链路和旧 Tool Contract;本 change 不执行公开协议切换。
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `single-react-diagnosis-harness`: 冻结单体诊断 Agent、Harness、证据状态、Guard、Fallback、预算、重试和阶段门禁的跨阶段基础契约。
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- None. 本阶段不声明旧运行能力已经迁移;后续 change 在实现对应行为时再修改现有 capability specs。
|
||||
|
||||
## Impact
|
||||
|
||||
- 设计与规格:ISS-014、OpenSpec 主规格、devflow 词汇表和阶段归档台账。
|
||||
- 契约测试:Draft/Fallback、状态语义、SSE、Router、预算、重试和 MySQL 安全基线。
|
||||
- 安全:`scripts/query_mysql.py` 和 `application.yml` 中的明文凭据必须移除并在外部轮换。
|
||||
- 后续代码范围:ChatController、ChatService、Agent/Hook、Tool、Redis、JPA/Flyway、静态前端、Trace/Eval fixtures。
|
||||
- 本 change 不切换 Controller 协议、不接入新 Agent、不修改公开运行行为。
|
||||
+75
@@ -0,0 +1,75 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Design contracts SHALL separate deterministic control from diagnosis reasoning
|
||||
The frozen contract set SHALL define Diagnosis Agent as the only diagnosis report author, Harness as deterministic execution control, EvidenceGuard as deterministic evidence validation, and SemanticGuard as an isolated single-turn semantic reviewer.
|
||||
|
||||
#### Scenario: Contract ownership is inspected
|
||||
- **WHEN** a later phase reads the stage-zero contracts
|
||||
- **THEN** no Harness contract assigns Planner, Executor, Composer, workflow routing, or diagnosis reasoning responsibilities to Harness
|
||||
|
||||
### Requirement: Tool invocation identity SHALL use the framework Tool Call ID
|
||||
The contract SHALL use the framework-provided `tool_call_id` as the sole Tool invocation reference and SHALL require later Harness implementations to reject missing, invalid, or duplicate IDs within a Run.
|
||||
|
||||
#### Scenario: Duplicate Tool Call ID is proposed
|
||||
- **WHEN** two Tool actions in one Run present the same framework Tool Call ID
|
||||
- **THEN** the contract classifies the second action as an error and prohibits overwriting the first canonical invocation
|
||||
|
||||
### Requirement: Invocation status and evidence status SHALL be independent
|
||||
The contract SHALL define `PROJECTING/READY/ERROR` as invocation lifecycle states and `EVIDENCE_FOUND/NO_EVIDENCE/ERROR` as evidence result states.
|
||||
|
||||
#### Scenario: Successful query returns no evidence
|
||||
- **WHEN** a Tool executes successfully and its bounded projection contains zero matching evidence
|
||||
- **THEN** invocation status is `READY` and evidence status is `NO_EVIDENCE`
|
||||
|
||||
#### Scenario: No-evidence result is cited
|
||||
- **WHEN** a Diagnosis Draft cites a `NO_EVIDENCE` Tool result
|
||||
- **THEN** the Analysis kind MUST be `NEGATIVE_OBSERVATION` and MUST remain bounded to the Tool query scope
|
||||
|
||||
### Requirement: Diagnosis Draft SHALL expose typed report structure
|
||||
The contract SHALL define conclusion, analysis items, action plan, recommendations, limitations and Tool Call bindings without exposing chain-of-thought or raw Tool payloads.
|
||||
|
||||
#### Scenario: Draft contains an analysis item
|
||||
- **WHEN** the Diagnosis Agent emits a structured Draft
|
||||
- **THEN** every Analysis has a unique analysis ID, a fixed Analysis kind and at least one Tool Call ID
|
||||
|
||||
### Requirement: Knowledge answers SHALL use exact RAG bindings
|
||||
The KNOWLEDGE_QUERY contract SHALL represent the answer as bounded answer items whose references identify the single lookup Tool Call and returned document IDs.
|
||||
|
||||
#### Scenario: Knowledge answer cites an unknown document
|
||||
- **WHEN** an answer item references a document ID absent from the bounded lookup result
|
||||
- **THEN** later Harness validation rejects the answer instead of publishing the fabricated citation
|
||||
|
||||
### Requirement: Safe fallback SHALL use a fixed schema
|
||||
The contract SHALL define stable fallback types for evidence validation failure, semantic unsupported and semantic unavailable outcomes, and SHALL exclude unvalidated Draft content and internal errors.
|
||||
|
||||
#### Scenario: Evidence validation fails twice
|
||||
- **WHEN** initial validation and the single no-Tool structural repair both fail
|
||||
- **THEN** the fallback type is `EVIDENCE_VALIDATION_FAILED` and verified sources are empty
|
||||
|
||||
### Requirement: Previous turn SHALL come from a safe durable Run result
|
||||
The contract SHALL define `diagnosis_run` as the durable source of intent, release outcome and safe published result, and SHALL exclude fallback, failed and cancelled Runs from Diagnosis previous-turn selection.
|
||||
|
||||
#### Scenario: Latest Run is a fallback
|
||||
- **WHEN** the latest same-session Run ended with `FALLBACK`
|
||||
- **THEN** it is not used as Diagnosis previous turn and selection continues to the latest eligible `DIAGNOSIS + SUCCESS` Run
|
||||
|
||||
### Requirement: Cancellation SHALL be layered and observable
|
||||
The contract SHALL distinguish cancellation request, prevention of new work, framework interruption, cancellable Tool work and HTTP timeout for already-blocking synchronous model calls.
|
||||
|
||||
#### Scenario: Client disconnects during a model call
|
||||
- **WHEN** an SSE client disconnects while a synchronous model call is in flight
|
||||
- **THEN** the system prevents later Draft release, requests interruption, records an internal cancelled terminal state and discards any late model result
|
||||
|
||||
### Requirement: Retry attempts SHALL be owned by Harness
|
||||
The contract SHALL require underlying SDK, HTTP and database retry layers to execute one attempt, while Harness explicitly owns any allowed Router or SemanticGuard retry.
|
||||
|
||||
#### Scenario: SemanticGuard returns an invalid schema
|
||||
- **WHEN** the first SemanticGuard attempt returns an invalid structured result
|
||||
- **THEN** Harness may execute one second attempt with the same verified snapshot and records both attempts
|
||||
|
||||
### Requirement: Phase gates SHALL remain serial
|
||||
The implementation plan SHALL contain eleven independent OpenSpec changes and SHALL prohibit starting a change before its predecessor is archived and committed.
|
||||
|
||||
#### Scenario: Stage 4 completes internal Agent tests
|
||||
- **WHEN** stage 4 passes its focused tests but stage 5 Guards are not implemented
|
||||
- **THEN** the public Chat entry remains on the old path and stage 6B cutover is prohibited
|
||||
@@ -0,0 +1,21 @@
|
||||
## 1. Contract Model
|
||||
|
||||
- [x] 1.1 Add typed enums for intent, release outcome, invocation status, evidence status, analysis kind, semantic verdict, fallback type and SSE outcome.
|
||||
- [x] 1.2 Add reusable records for Diagnosis Draft, Knowledge Answer Draft, safe fallback, published result and previous turn.
|
||||
- [x] 1.3 Add focused serialization and contract-shape tests covering positive evidence, no-evidence negative observation and forbidden raw/internal fields.
|
||||
|
||||
## 2. Security And Configuration Baseline
|
||||
|
||||
- [x] 2.1 Remove tracked plaintext credentials from `scripts/query_mysql.py` and `application.yml`, require environment-based secrets, and ignore local secret files.
|
||||
- [x] 2.2 Record the required external credential rotation and the Spring AI hidden-retry baseline without changing the public Chat runtime in this stage.
|
||||
|
||||
## 3. Design Freeze Alignment
|
||||
|
||||
- [x] 3.1 Keep ISS-014, OpenSpec artifacts and the devflow glossary aligned on the 11 serial changes, framework Tool Call ID, dual status fields and safe previous-turn source.
|
||||
- [x] 3.2 Add an architecture audit and cross-artifact alignment result to `decisions.md`.
|
||||
- [x] 3.3 Create the sm-flow `.committed` marker after proposal/design/specs/tasks and all decision gates pass.
|
||||
|
||||
## 4. Verification
|
||||
|
||||
- [x] 4.1 Run focused contract tests and the smallest existing Chat/evidence baseline needed to prove stage zero did not switch runtime behavior.
|
||||
- [x] 4.2 Run strict OpenSpec validation and record commands, results and unverified external rotation in acceptance evidence.
|
||||
@@ -0,0 +1 @@
|
||||
ready
|
||||
@@ -0,0 +1 @@
|
||||
committed
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-21
|
||||
@@ -0,0 +1,106 @@
|
||||
## Context
|
||||
|
||||
阶段 2 的 `DiagnosisHarnessCore` 已提供显式 `RunContext`、模型/Tool/Token/字节预算、取消和唯一生命周期;阶段 3A-3C 已提供 canonical invocation、`ToolBoundary` 以及 RAG、日志、MySQL adapter。当前这些组件没有生产 Agent 消费者,公开 Chat 仍由 `ChatService` 创建单 Agent 简单路径或 Planner/Executor/Verifier/Composer 多 Agent 路径。
|
||||
|
||||
Spring AI Alibaba 1.1.2.0 的 `ReactAgent` 自带模型与 Tool 交替执行的 ReAct loop。`ModelInterceptor` 包围每次模型调用;`ToolInterceptor` 获得 `ToolCallRequest.getToolCallId()`、参数和 `RunnableConfig` metadata;`BeanOutputConverter` 可从输出类型生成格式提示,但 `call` 最终仍返回 `AssistantMessage`。因此本阶段可以使用框架扩展点接入 Harness,无需自行编写循环或修改框架。
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- 创建一个且仅一个拥有 Tool loop 的 Diagnosis ReactAgent。
|
||||
- 以显式 `RunContext` 执行当前 Query 和可选 `PreviousTurn`,输入失败时 fail closed,不截断原始 Query。
|
||||
- 将框架原始 Tool Call ID 传给阶段 3B/3C adapter,并只把有界 `agent_result` 返回模型。
|
||||
- 对每次模型调用、Token、Tool 调用和输入/输出上下文字节实施确定性预算。
|
||||
- 输出可严格反序列化的 `DiagnosisDraft`;无证据时明确停止,不补造结论。
|
||||
- 保留可注入 AgentStep Hook、RunContext 生命周期和 canonical Tool invocation 审计边界。
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- 不实现 EvidenceGuard、SemanticGuard、结构修复、引用展开或最终发布。
|
||||
- 不创建 Intent Router、不选择 previous turn、不持久化 DiagnosisRun。
|
||||
- 不修改 Controller、SSE、ChatService、AiOpsService 或旧多 Agent 链路。
|
||||
- 不实现外层 Graph、手写 ReAct while、Agent 自动重跑或 Tool 自动重试。
|
||||
- 不增加真实日志源或新的 evidence Tool。
|
||||
|
||||
## Decisions
|
||||
|
||||
### One per-run ReactAgent uses the framework loop
|
||||
|
||||
`DiagnosisAgentFactory` 为每个 `RunContext` 构建一个名为 `diagnosis_agent` 的 `ReactAgent`,关闭并行 Tool 执行并使用单一 Prompt。内部 `DiagnosisAgentUseCase` 对 Agent 只调用一次;一次调用内由框架决定正常的模型/Tool轮次。
|
||||
|
||||
替代方案是复用 `ChatService.createReactAgent`,但它注册旧 Tool、旧 Hook 和完整旧运行职责,无法保证 Harness boundary、极简上下文和不接公开入口。另一个替代方案是手写 while,直接违反 ISS-014。
|
||||
|
||||
### Model and Tool control use different framework boundaries
|
||||
|
||||
`HarnessModelInterceptor` 在每次模型调用前执行 `core.beforeModelCall(context)`,调用完成后读取 `ChatResponseMetadata.Usage` 并执行 `core.recordTokens`。内部用例设置 `_stream_=false`,保证 interceptor 可取得完整 `ChatResponse` 和 Token Usage。
|
||||
|
||||
`HarnessToolInterceptor` 读取框架原始 Tool Call ID,并按 Tool 名调用 `HarnessEvidenceTools` 中的 adapter bridge。bridge 创建 `ToolCallRequestEnvelope(runId, frameworkId, toolName, arguments, true, true)`;授权来自注册到当前 Diagnosis Agent 的固定 Tool 集合。Tool 成功时只返回 `agent_result`,失败时只返回稳定的 `evidence_status/tool_call_id/error_code`,不返回 raw response、内部异常或 invocation lifecycle。
|
||||
|
||||
Tool 调用预算只在 `ToolBoundary` 中 reserve;Agent interceptor 不重复调用 `beforeToolCall`。这保持阶段 3A 的 canonical record 与预算原子边界。
|
||||
|
||||
### Tool definitions and execution are registered together
|
||||
|
||||
`HarnessEvidenceTools` 固定暴露 `lookup_knowledge`、`query_logs`、`query_mysql` 三个 Spring `ToolCallback` 定义,输入类型分别复用已冻结的 Request records,描述复用 `AgentToolContracts`。callback 本体不允许绕过 interceptor 直接执行;实际执行映射与定义在同一 registry 中,构造时拒绝缺项或重复项。
|
||||
|
||||
替代方案是在 `ToolCallback.call` 中读取 ThreadLocal 或生成 ID,都会违反 RunContext 和 canonical ID 契约。
|
||||
|
||||
### Input and output remain typed and bounded
|
||||
|
||||
`DiagnosisAgentInput` 只包含非空原始 `query` 和可选 `PreviousTurn`。`DiagnosisAgentLimits` 配置 query、previous turn、总输入和 Draft 的 UTF-8 字节上限。内部用例先分别验证,再将固定 `{query, previous_turn}` JSON 作为唯一 User 输入,并通过 `core.reserveRunBytes` 计入 Run 容量;任何超限都在模型调用前失败,不做语义截断。
|
||||
|
||||
Agent 使用基于 `BeanOutputConverter<DiagnosisDraft>` 生成的格式提示,并通过 schema post-process 明确允许冻结契约中的 `conclusion=null`;Factory 以 `.outputSchema(...)` 注入该格式。最终文本先检查 UTF-8 上限并计入容量,再由注入的 `ObjectMapper` 严格解析为冻结 record。Markdown fence、前后说明、未知结构或空输出不做修复和自动重试;阶段 5 再实现一次显式无 Tool 结构修复。
|
||||
|
||||
### Previous turn is context, not current evidence
|
||||
|
||||
`PreviousTurn` 按固定 record 原样序列化,仅帮助理解追问。Prompt 明确禁止引用旧 Tool Call ID 或把 previous turn 当作当前 Run 的 evidence;当前 `DiagnosisDraft.analysis[*].tool_call_ids` 只能来自本 Run Tool result。
|
||||
|
||||
### Prompt owns diagnosis semantics, Harness owns control
|
||||
|
||||
唯一 classpath Prompt 要求结论先行、所有分析绑定 Tool Call IDs、`NO_EVIDENCE` 只能形成限定范围的 `NEGATIVE_OBSERVATION`。没有足够 `EVIDENCE_FOUND` 时 `conclusion=null`,在 `limitations` 记录范围和缺失信息并结束,不把 `NO_EVIDENCE` 推导为系统健康或根因排除。
|
||||
|
||||
阶段 4 不实现物理或语义 Guard,因此内部返回值叫 Draft,不能直接公开发布。
|
||||
|
||||
### Audit remains injectable and explicit
|
||||
|
||||
Factory 接收可选框架 `Hook` 列表。阶段 6A 可以注入现有 `AgentLoggingHook`;内部用例始终在 `RunnableConfig` metadata 写入 sessionId/runId,所以该 Hook 不需要读取 ThreadLocal fallback。模型/Tool 预算和生命周期仍绑定 RunContext,Tool canonical audit 仍由 `ToolBoundary`/store 负责。
|
||||
|
||||
## Architecture And Interface Impact
|
||||
|
||||
```text
|
||||
DiagnosisAgentInput(query, previousTurn) + RunContext
|
||||
-> DiagnosisAgentUseCase: validate/serialize/reserve context bytes
|
||||
-> DiagnosisAgentFactory: one ReactAgent + one Prompt
|
||||
-> HarnessModelInterceptor -> DiagnosisHarnessCore -> ChatModel
|
||||
-> framework ReAct loop
|
||||
-> HarnessToolInterceptor -> HarnessEvidenceTools -> 3B/3C adapters
|
||||
-> ToolBoundary -> canonical invocation store
|
||||
-> bounded AssistantMessage text
|
||||
-> strict ObjectMapper -> DiagnosisDraft
|
||||
```
|
||||
|
||||
- 数据所有权:use case 拥有本次输入和 Draft 解析;RunContext 拥有预算/取消/生命周期;ToolBoundary/store 拥有调用记录;Agent 只拥有诊断语义。
|
||||
- 接口影响:L2 内部接口。新增阶段 6A 可消费的 Java 类型,不改变公开 HTTP/SSE、JPA、数据库或旧 Service 方法。
|
||||
- 生命周期:本 use case 不把成功 Draft 标记为 Run SUCCESS,因为阶段 5 Guards 尚未执行;预算、取消和超时可由 Core 先行终止 Run,最终持久化映射留给阶段 6A。
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Provider 返回 fenced JSON 或附加说明导致解析失败 → Mitigation:阶段 4 fail closed 且不重跑;测试固定这一行为,阶段 5 才允许一次无 Tool 结构修复。
|
||||
- [Risk] Token Usage 缺失或为 null → Mitigation:模型调用次数仍被强制;仅在可用且非负时记录实际 Token,并在 acceptance 标记 provider 元数据依赖。
|
||||
- [Risk] Tool interceptor 错误泄露内部异常 → Mitigation:只返回 allowlisted error code,不返回 message/stack/raw response。
|
||||
- [Risk] 输入和 Tool 投影共同消耗 Run bytes,配置过小会过早耗尽 → Mitigation:所有限制可配置,usage 可观察,不在本阶段硬编码生产校准值。
|
||||
- [Risk] 复用旧 AgentLoggingHook 仍包含 ThreadLocal fallback → Mitigation:新链路总是提供 metadata;阶段 7 物理清理前不修改旧 Hook,focused test 验证相同 sessionId/runId 传播。
|
||||
- [Risk] ReactAgent 内部基于框架 StateGraph → Mitigation:这是框架原生实现细节;本项目不再包一层业务 Graph 或多 Agent workflow。
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. 新增内部 Agent 类型、Prompt、interceptors 和 Tool registry,不注册 Controller bean。
|
||||
2. 使用 scripted ChatModel 和 fake/adapted Tool 完成内部 tool-loop、output、budget、no-evidence 和审计测试。
|
||||
3. 保持公开 Chat 与旧多 Agent 代码 diff 为空,Archive 并提交阶段 4。
|
||||
4. 阶段 5 在内部 Draft 后增加 Guards;阶段 6A 再创建运行应用用例并装配真实审计/previous turn;阶段 6B 原子切换公开入口。
|
||||
|
||||
Rollback:删除新增内部包、Prompt 和测试即可;不存在公开协议、数据迁移或运行时切换。
|
||||
|
||||
## Open Questions
|
||||
|
||||
无。生产预算默认值、公开切换和最终发布策略分别由配置校准、阶段 6B 和阶段 5 处理。
|
||||
@@ -0,0 +1,31 @@
|
||||
## Why
|
||||
|
||||
阶段 2-3C 已建立显式 RunContext、预算、canonical Tool invocation 和三类有界 evidence Tool,但还没有使用这些边界的诊断执行者。需要在不影响公开 Chat 的前提下建立单一 Diagnosis ReAct Agent,证明框架原生 tool loop、冻结的 DiagnosisDraft 和 Harness 控制能够闭环,再进入阶段 5 的释放门禁。
|
||||
|
||||
## What Changes
|
||||
|
||||
- 新增唯一的 Diagnosis Agent Prompt,把诊断规划、证据查询、证据充分性判断和报告草稿生成合并到一个 ReactAgent。
|
||||
- 新增独立内部 Diagnosis 用例,显式接收 RunContext、当前原始 Query 和可选的固定 PreviousTurn,不读取完整 Session 历史。
|
||||
- 通过 Spring AI Alibaba 的 ModelInterceptor 和 ToolInterceptor 接入 Harness;使用框架原生 ReAct/tool loop 和原始 tool_call_id,不手写循环或第二套调用 ID。
|
||||
- 将 lookup_knowledge、query_logs 和 query_mysql 的冻结 Tool 定义连接到阶段 3B/3C adapter,只向 Agent 返回有界 agent_result。
|
||||
- 使用 DiagnosisDraft 作为结构化输出契约,并对输入上下文、模型轮次、Token、Tool 调用和输出字节执行确定性预算。
|
||||
- 允许注入现有 AgentStep Hook,Run 审计继续由 RunContext/lifecycle 承载,ToolInvocation 审计继续由 canonical Tool boundary 承载。
|
||||
- 证据不足时要求 Agent 输出 conclusion=null、明确 limitations 并结束当前 ReAct 执行,不自动重跑 Diagnosis Agent、模型调用或 Tool 调用。
|
||||
- 不修改 ChatController、公开 /api/chat 或 /api/chat_stream,不删除旧 Planner/Executor/Verifier/Composer 链路。
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `single-react-diagnosis-agent`: 定义单一 Diagnosis ReactAgent 的内部输入、Harness-controlled ReAct/tool loop、结构化 Draft、预算、无证据停止和审计边界。
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- None. 既有 Harness、Tool 和公开 Chat capability 的需求语义不在本阶段改变。
|
||||
|
||||
## Impact
|
||||
|
||||
- 新增 `com.superbiz.agent.harness.agent` 内部包、一个 classpath Prompt 和 focused Agent/tool-loop tests。
|
||||
- 复用 `DiagnosisHarnessCore`、`RunContext`、`DiagnosisDraft`、`PreviousTurn`、`AgentToolContracts`、三类 Tool adapter 和 Spring AI Alibaba ReactAgent/interceptor API。
|
||||
- 内部接口影响为 L2:新增可供阶段 6A 调用的内部用例与工厂;现有公开 Controller、SSE、DTO、数据库契约和旧 ChatService 行为保持不变。
|
||||
- 主要风险是框架结构化输出只提供格式提示而不负责 Java 反序列化,以及 Tool Callback 本身拿不到框架调用 ID;实现必须分别用严格 ObjectMapper 解析和 ToolInterceptor 解决,失败时 fail closed。
|
||||
+78
@@ -0,0 +1,78 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Diagnosis execution SHALL use one framework ReactAgent
|
||||
The internal diagnosis path SHALL create exactly one `diagnosis_agent` with the framework-provided ReAct/tool loop. It SHALL NOT create Planner, Executor, Verifier, Composer, SequentialAgent, SupervisorAgent, an outer business Graph, or a hand-written model/Tool loop. One use-case invocation SHALL call the Diagnosis Agent once and SHALL NOT wrap the Agent or any individual model call in Harness retry.
|
||||
|
||||
#### Scenario: Diagnosis needs Tool evidence
|
||||
- **WHEN** the model emits one or more Tool actions before its final response
|
||||
- **THEN** the same Diagnosis ReactAgent executes those actions through its framework loop and returns one final Draft without invoking another Agent
|
||||
|
||||
#### Scenario: Agent execution fails
|
||||
- **WHEN** the single Diagnosis Agent invocation throws or returns invalid structured output
|
||||
- **THEN** the internal use case fails closed without automatically invoking the Agent or model again
|
||||
|
||||
### Requirement: Diagnosis input SHALL contain only current Query and optional PreviousTurn
|
||||
The internal use case SHALL accept a non-blank current Query and an optional frozen `PreviousTurn`, serialize them as the fixed `query` and `previous_turn` input fields, and SHALL NOT load or accept complete Session history, Redis memory, prior raw Tool results, or model-generated history summaries. The current Query SHALL retain its original text and SHALL NOT be rewritten or silently truncated.
|
||||
|
||||
#### Scenario: Follow-up diagnosis has a safe previous turn
|
||||
- **WHEN** a caller supplies a `PreviousTurn`
|
||||
- **THEN** the Agent receives exactly the frozen previous-turn fields plus the current Query and no complete conversation history
|
||||
|
||||
#### Scenario: Query exceeds configured context budget
|
||||
- **WHEN** the current Query exceeds its UTF-8 byte limit
|
||||
- **THEN** the use case rejects it before any model call instead of truncating or rewriting it
|
||||
|
||||
### Requirement: Every model round SHALL be controlled by RunContext
|
||||
The Diagnosis Agent SHALL receive `RunContext` explicitly and SHALL use a model interceptor to call `DiagnosisHarnessCore.beforeModelCall` for every framework ReAct model round. Non-streaming model response Usage SHALL be recorded into the same Run budget when available. Cancellation, deadline, model-call exhaustion, or Token exhaustion SHALL prevent subsequent controlled work.
|
||||
|
||||
#### Scenario: ReAct performs two model rounds
|
||||
- **WHEN** one model round requests a Tool and the next produces the Draft
|
||||
- **THEN** the same Run budget records two model calls and no Harness retry attempt
|
||||
|
||||
#### Scenario: Model-call budget is exhausted
|
||||
- **WHEN** the framework attempts a model round beyond the configured maximum
|
||||
- **THEN** the call is rejected before reaching ChatModel and the Run records budget exhaustion
|
||||
|
||||
### Requirement: Evidence Tools SHALL execute through the Harness boundary
|
||||
The Diagnosis Agent SHALL expose only the frozen `lookup_knowledge`, `query_logs`, and `query_mysql` definitions. A Tool interceptor SHALL propagate the exact framework Tool Call ID, Run ID, Tool name, and raw JSON arguments into the corresponding stage 3B/3C adapter and `ToolBoundary`. Successful observations SHALL contain only the bounded `agent_result`; failed observations SHALL contain only stable error semantics and SHALL NOT contain raw responses, internal exceptions, credentials, or invocation lifecycle internals.
|
||||
|
||||
#### Scenario: Framework requests RAG evidence
|
||||
- **WHEN** the model calls `lookup_knowledge` with framework ID `call-1`
|
||||
- **THEN** the RAG adapter and Agent observation use exactly `call-1`, and the canonical invocation is owned by the current Run
|
||||
|
||||
#### Scenario: Tool execution fails
|
||||
- **WHEN** a registered adapter returns an error result
|
||||
- **THEN** the Agent receives `evidence_status=ERROR`, the framework Tool Call ID and a stable error code without automatic Tool retry or raw failure detail
|
||||
|
||||
#### Scenario: Unknown Tool is requested
|
||||
- **WHEN** a model requests a Tool outside the three registered definitions
|
||||
- **THEN** the Harness does not authorize or emulate it and does not create a canonical evidence record
|
||||
|
||||
### Requirement: Diagnosis output SHALL be a bounded DiagnosisDraft
|
||||
The Agent SHALL receive the generated schema for `DiagnosisDraft` and SHALL return JSON that the internal use case strictly parses into the frozen record. The use case SHALL enforce configured UTF-8 limits for query, previous turn, total input and Draft output and account accepted input/output bytes against the Run capacity. It SHALL reject blank, fenced, prefixed, malformed, oversized or schema-incompatible output without repair or retry.
|
||||
|
||||
#### Scenario: Supported Draft is returned
|
||||
- **WHEN** the Agent returns valid JSON containing Conclusion, Analysis items, Action Plan, Recommendations and Limitations
|
||||
- **THEN** the use case returns a `DiagnosisDraft` whose Analysis Tool Call IDs are the exact strings emitted by the Agent
|
||||
|
||||
#### Scenario: Model returns prose around JSON
|
||||
- **WHEN** the final response contains a Markdown fence or explanatory prefix around an otherwise valid object
|
||||
- **THEN** strict parsing fails and the Diagnosis Agent is not invoked a second time
|
||||
|
||||
### Requirement: Insufficient evidence SHALL terminate without a fabricated conclusion
|
||||
The single Prompt SHALL require every normal Analysis item to cite current-Run evidence Tool Call IDs and SHALL restrict `NO_EVIDENCE` to scoped `NEGATIVE_OBSERVATION`. If current evidence cannot support a diagnosis, the Agent SHALL stop the current ReAct execution with `conclusion=null`, describe the actual scope and missing information in `limitations`, and SHALL NOT infer that the problem does not exist or fabricate a root cause.
|
||||
|
||||
#### Scenario: Tool finds no evidence
|
||||
- **WHEN** the only completed Tool observation has `evidence_status=NO_EVIDENCE`
|
||||
- **THEN** the final Draft has no confirmed Conclusion, records the bounded negative observation and missing information, and makes no additional automatic retry
|
||||
|
||||
### Requirement: Stage-four execution SHALL remain internal and auditable
|
||||
The new use case SHALL be callable only as an internal Java/test entry in this stage and SHALL NOT be wired into public Chat or AIOps controllers. It SHALL propagate the same sessionId/runId through RunnableConfig metadata, allow the existing AgentStep Hook to be injected, retain ToolBoundary canonical invocation audit, and leave final Run success/persistence ownership to later Guard/application stages.
|
||||
|
||||
#### Scenario: Internal audited execution completes
|
||||
- **WHEN** an injected audit Hook observes a model round and a Tool executes
|
||||
- **THEN** both observe the same RunContext sessionId/runId and the public Chat call path remains unchanged
|
||||
|
||||
#### Scenario: Stage four is archived
|
||||
- **WHEN** focused Agent tests pass and the change is archived
|
||||
- **THEN** `ChatController`, public SSE behavior and the old multi-Agent implementation remain available for stages 5, 6A and 6B
|
||||
@@ -0,0 +1,22 @@
|
||||
## 1. Input And Prompt Contract
|
||||
|
||||
- [x] 1.1 Add immutable DiagnosisAgentInput and configurable UTF-8 context/output limits with validation and focused boundary tests.
|
||||
- [x] 1.2 Add the single Diagnosis Agent classpath Prompt covering Tool selection, current-Run evidence binding, no-evidence stopping and exact DiagnosisDraft JSON output.
|
||||
|
||||
## 2. Harness Tool Integration
|
||||
|
||||
- [x] 2.1 Add HarnessEvidenceTools with the three frozen Tool definitions and adapter bridges for RAG, logs and MySQL.
|
||||
- [x] 2.2 Add a ToolInterceptor that propagates the exact framework Tool Call ID and RunContext, returns only bounded agent_result or stable error JSON, and never retries or leaks raw results.
|
||||
- [x] 2.3 Add focused tests for framework ID fidelity, successful Tool observations, error observations, unknown Tool rejection and single Tool budget accounting.
|
||||
|
||||
## 3. Single Diagnosis Agent Execution
|
||||
|
||||
- [x] 3.1 Add a ModelInterceptor that enforces every model boundary and records non-streaming Token Usage into the same Run budget.
|
||||
- [x] 3.2 Add DiagnosisAgentFactory that creates one non-parallel ReactAgent with one Prompt, DiagnosisDraft output schema, Harness interceptors and injectable audit Hooks.
|
||||
- [x] 3.3 Add the internal DiagnosisAgentUseCase that validates/serializes fixed input, enforces context/Draft byte limits, calls the Agent once and strictly parses DiagnosisDraft without repair or retry.
|
||||
|
||||
## 4. Focused Behavior Verification
|
||||
|
||||
- [x] 4.1 Add scripted ChatModel tests proving a single Agent performs framework model -> Tool -> model execution, preserves Query/PreviousTurn, returns typed Analysis/Conclusion/Tool IDs and propagates sessionId/runId to audit.
|
||||
- [x] 4.2 Add no-evidence, invalid/fenced output, model-call budget, Token budget, context budget and no-automatic-retry tests.
|
||||
- [x] 4.3 Verify no public Chat/AiOps runtime file changed, run focused Harness/Agent regressions, compile, and validate the OpenSpec strictly.
|
||||
+1
@@ -0,0 +1 @@
|
||||
ready
|
||||
@@ -0,0 +1 @@
|
||||
committed
|
||||
+2
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-21
|
||||
@@ -0,0 +1,118 @@
|
||||
## Context
|
||||
|
||||
阶段 2-4 已提供显式 `RunContext`、类型化重试策略、canonical invocation store、三类有界 Tool projection 和内部 `DiagnosisAgentUseCase`。当前 `DiagnosisDraft` 仍只是未发布草稿:类型反序列化不保证 ID 唯一、报告内部引用闭合,也不证明 Tool Call 属于当前 Run 或 projection 可引用;即使物理证据真实,也还需要独立判断整份报告是否被证据支持。
|
||||
|
||||
本阶段只建立内部 release boundary。公开 `ChatController`、SSE、Run persistence 和旧 Planner/Executor/Verifier/Composer 链路由阶段 6A/6B/7 处理。接口影响为 L2,新增内部 Java use case 供后续阶段消费,不改变当前外部协议。
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- 用确定性代码验证 Draft 结构、Analysis 引用和当前 Run canonical invocation。
|
||||
- 只从 `agent_result` 和必要的 canonical request 生成无 Tool ID、无 raw response 的 verified evidence snapshot。
|
||||
- 首次验真失败时允许一次保持用户可见语义不变的结构修复。
|
||||
- 使用同一 `ChatModel` 执行无 Tool、无记忆、无 ReAct 的隔离 SemanticGuard,并由 Harness 控制预算、超时、取消和技术重试。
|
||||
- 只发布通过两层门禁的原 Draft;所有其他可降级路径返回固定 `SafeFallback`。
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- 不切换 `/api/chat`、`/api/chat_stream` 或 `/api/ai_ops`。
|
||||
- 不实现 SSE 事件、Run 持久化、PreviousTurn、Intent Router 或最终报告渲染。
|
||||
- 不删除或改造旧 Gatekeeper、Verifier、Composer 和多 Agent 生产链路。
|
||||
- 不访问 canonical `raw_response`,不新增 Redis 索引或历史证据召回。
|
||||
- 不执行 live model/Redis/MySQL E2E;最终 live 验收留到阶段 7。
|
||||
|
||||
## Decisions
|
||||
|
||||
### 1. EvidenceGuard 是 fail-closed 的纯确定性边界
|
||||
|
||||
`EvidenceGuard.validate(RunContext, DiagnosisDraft)` 分两步执行:先验证 Draft 顶层/嵌套结构、非空文本、Analysis ID 唯一性和 `based_on_analysis_ids` 闭合;再对每个 Analysis 的非空 Tool Call ID 使用 `ToolCallKeyFactory.create(runId, id)` 查询 `CanonicalInvocationStore`。
|
||||
|
||||
可引用记录必须同时满足:record `runId` 等于当前 Run、`status=READY`、`agent_result` 非空、`evidence_status=EVIDENCE_FOUND|NO_EVIDENCE`。`AnalysisKind.accepts` 固定 `NORMAL -> EVIDENCE_FOUND`、`NEGATIVE_OBSERVATION -> NO_EVIDENCE`。任何非法 key、缺失记录、跨 Run、错误/投影中记录、空 projection、状态不一致或解析错误都产生稳定 `EvidenceViolation`,不会抛出原始存储内容或进入 SemanticGuard。
|
||||
|
||||
替代方案是复用旧 `ExecutorGatekeeperService`;拒绝该方案,因为旧实现绑定 JPA `ToolInvocation`、Session/Run 双入口和 Map DTO,不能表达 Redis canonical lifecycle、`NO_EVIDENCE` kind 约束或新的 typed Draft。
|
||||
|
||||
### 2. verified snapshot 使用 Tool-specific 严格 projection adapter
|
||||
|
||||
Guard 根据冻结的 `AgentToolContracts` 严格解析 `RagToolResult`、`QueryLogsToolResult`、`MysqlToolResult`,并校验 projection 内 `tool_call_id`、`evidence_status` 与 canonical record 一致。未知 Tool 一律拒绝。
|
||||
|
||||
快照结构为 `analysis_id + analysis_text + analysis_kind + verified_evidence[]`。单条 evidence 只包含 `source_type/source/scope/timestamp/excerpt/values`:RAG 展开稳定 document metadata/excerpt;日志展开 pattern/event 或带精确 query/time window 的零结果;MySQL 从 canonical request 取得逻辑 data source/SQL/params scope,从 projection 取得 columns/有界 rows 或零结果。快照不包含 Tool Call ID、Redis key、raw response 或未被 Draft 引用的调用。
|
||||
|
||||
Snapshot 同时可确定性提取去重的 `SafeFallback.VerifiedSource`,因此 release policy 无需读取 Store。替代通用 JsonNode 透传;拒绝该方案,因为它会扩大模型输入面并削弱字段级测试。
|
||||
|
||||
### 3. Evidence repair 只能改引用结构,不能成为报告作者
|
||||
|
||||
首次 Guard 失败后,`EvidenceRepair` 使用直接单轮 `ChatModel.call(Prompt)`,输入为原始 query、完整 Draft 和稳定 violation code/target;没有 Tool callback、memory 或 ReAct Agent。调用由 `context.retryPolicies().evidenceRepair()` 执行,其 `maxAttempts=1`,因此不能形成隐式 retry loop。
|
||||
|
||||
修复输出经过严格 `DiagnosisDraft` 解析后,转换为移除所有标识/引用字段的 `SemanticDraftView` 并与原 Draft 比较。Conclusion/Analysis/Action/Recommendation/Limitations 的文本、kind、顺序和人工确认标志有任何变化都拒绝;只允许 `analysis_id`、`tool_call_ids` 和 `based_on_analysis_ids` 修正。修复后完整重跑 EvidenceGuard,第二次仍失败直接 `EVIDENCE_VALIDATION_FAILED`。
|
||||
|
||||
替代方案是允许模型删除无效 Analysis 或改写结论;拒绝该方案,因为 ISS-014 固定 Diagnosis Agent 是唯一报告作者,Harness/repair 不能改变用户报告语义。
|
||||
|
||||
### 4. SemanticGuard 使用直接 ChatModel 单轮调用
|
||||
|
||||
`SemanticGuardInput` 包含原始 query、`SemanticDraftView` 和 verified snapshot。View 保留完整用户可见 Draft 语义和 Analysis IDs,但删除所有 Tool Call IDs。每个 attempt 都构造只含一个 system message 和一个 user JSON message 的新 `Prompt`,直接调用系统注入的同一 `ChatModel`;不构建 `ReactAgent`,因此没有 Tool、memory、checkpoint、Graph 或 Agent 回调。
|
||||
|
||||
输出只允许精确 JSON object `{verdict, reason}`,字段集合固定,verdict 只允许 `SUPPORTED|UNSUPPORTED`,reason 必须非空。SemanticGuard 不返回 corrected report。`UNSUPPORTED` 正常返回给 release policy,不能进入 retry classifier。
|
||||
|
||||
### 5. 单轮模型执行由共享受控边界管理
|
||||
|
||||
`GuardModelCall` 在每次调用前执行 `DiagnosisHarnessCore.beforeModelCall`,在模型实际返回后从 `ChatResponseMetadata.Usage` 记录 input/output Token,并再次检查 Run active。调用提交到注入的 `ExecutorService`,等待时间为 `min(singleAttemptTimeout, totalTimeoutRemaining)`;attempt 超时或 Run cancellation 时执行 `Future.cancel(true)`。模型即使忽略中断而迟到返回,后台任务仍记录真实 Token,并在 `core.checkActive` 处阻止迟到结果被使用。
|
||||
|
||||
SemanticGuard 通过既有 `HarnessRetryExecutor` 和 `context.retryPolicies().semanticGuard()` 执行。只有 timeout、transport、parse error、schema invalid 可进行第二次 attempt;两次使用完全相同的序列化输入。Run cancelled/budget exhausted 不转换为 Fallback,而是继续抛给阶段 6A 生命周期边界;其他最终技术失败映射 `SEMANTIC_UNAVAILABLE`。
|
||||
|
||||
### 6. Release policy 不携带未验证语义
|
||||
|
||||
`DiagnosisReleaseUseCase.execute(context, query, draft)` 固定顺序为 `EvidenceGuard -> optional repair -> EvidenceGuard -> SemanticGuard -> decision`。成功结果保留与 Guard 通过的同一个 Draft 实例/修复实例,不总结、裁剪或重排。
|
||||
|
||||
- `SUPPORTED`:`ReleaseOutcome.SUCCESS`,返回完整 Draft 和 verified snapshot。
|
||||
- `UNSUPPORTED`:`ReleaseOutcome.FALLBACK` + `SEMANTIC_UNSUPPORTED`。
|
||||
- SemanticGuard 最终技术失败:`ReleaseOutcome.FALLBACK` + `SEMANTIC_UNAVAILABLE`。
|
||||
- 首次修复调用失败、语义被改写或第二次 Guard 失败:`ReleaseOutcome.FALLBACK` + `EVIDENCE_VALIDATION_FAILED`。
|
||||
|
||||
`SafeFallbackFactory` 使用固定模板。Evidence failure 的 `verified_sources=[]`;Semantic fallback 只从 snapshot 提取来源。所有 Fallback 的 `conclusion=null`,且不接收 Draft 文本、SemanticGuard reason、供应商错误或异常作为模板参数。
|
||||
Fallback release result 不保留完整 snapshot,因为 snapshot 含 Analysis text;已验真来源只以 `SafeFallback.verified_sources` 的安全子集返回,防止误序列化内部 result 泄漏草稿。
|
||||
|
||||
## Module and Ownership Audit
|
||||
|
||||
```text
|
||||
query + DiagnosisDraft + RunContext
|
||||
-> DiagnosisReleaseUseCase
|
||||
-> EvidenceGuard -> CanonicalInvocationStore + Tool projection parsers
|
||||
-> optional EvidenceRepair -> GuardModelCall -> shared ChatModel
|
||||
-> SemanticGuardInput(snapshot + ID-free Draft)
|
||||
-> SemanticGuard -> HarnessRetryExecutor -> GuardModelCall -> shared ChatModel
|
||||
-> original Draft OR SafeFallbackFactory
|
||||
```
|
||||
|
||||
- `RunContext` owns deadline、cancellation、budget、retry policy 和 lifecycle;本阶段不创建第二套 Run 状态。
|
||||
- canonical store owns Tool 调用真实性;EvidenceGuard 只读并生成隔离 snapshot,不持有 Redis client。
|
||||
- Diagnosis Agent owns report semantics;repair 只修标识结构,SemanticGuard 只审查,release 只选择 Draft 或模板。
|
||||
- 最大耦合风险是阻塞式 `ChatModel` 的取消能力;受控 Future 可发出 interrupt,但 provider 是否立即终止取决于 SDK,故迟到结果必须由 Core active check 丢弃。
|
||||
- 无新 ADR:模块边界均来自 ISS-014 已冻结设计,且公开消费者尚未切换。
|
||||
|
||||
## Interface Impact
|
||||
|
||||
- Level: L2 internal interface。
|
||||
- 新增内部 records/interfaces/use cases;不修改现有 Java 方法签名。
|
||||
- 当前消费者仅 focused tests;阶段 6A 将装配真实 `DiagnosisAgentUseCase -> DiagnosisReleaseUseCase`。
|
||||
- HTTP/SSE、Controller DTO、JPA/Flyway、Redis key schema、旧 Chat/AiOps 行为均不变,无迁移或回滚数据操作。
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Projection 字段遗漏导致误判] -> 三类显式 adapter、ID/status 双重校验和 contract fixture tests。
|
||||
- [Repair 改写报告] -> `SemanticDraftView` 确定性相等检查,任何可见差异直接 Evidence fallback。
|
||||
- [模型忽略 interrupt] -> Future cancel、迟到 Token 记录、Core active check;不允许迟到结果进入 release。
|
||||
- [业务 `UNSUPPORTED` 被重试] -> verdict 作为正常返回值,classifier 只处理异常。
|
||||
- [Fallback 泄漏草稿或内部 reason] -> 固定 factory 不接受这些参数,并对序列化结果做负向断言。
|
||||
- [输入快照过大] -> repair/SemanticGuard 独立 UTF-8 input/output limits,并占用 Run byte/model/token budget。
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. 新增内部 guard/release 包和聚焦测试,不连接公开入口。
|
||||
2. 阶段 6A 将内部 Chat application use case 装配到这些接口。
|
||||
3. 阶段 6B 在新 release boundary 通过后原子切换公开 SSE。
|
||||
4. 如阶段 5 回滚,只需移除新增内部包;阶段 4 和旧生产链路不受影响。
|
||||
|
||||
## Open Questions
|
||||
|
||||
- None。单次/总超时与字节预算由构造配置提供可调默认值,具体生产数值按阶段 7 Trace 校准,不构成本阶段方向问题。
|
||||
@@ -0,0 +1,30 @@
|
||||
## Why
|
||||
|
||||
阶段 4 已能生成带框架 Tool Call ID 的 `DiagnosisDraft`,但草稿尚未经过当前 Run 证据验真和隔离语义审查,不能安全发布。阶段 5 需要补齐确定性 EvidenceGuard、单轮 SemanticGuard 和 fail-closed 释放策略,作为阶段 6A/6B 接入公开链路的前置门禁。
|
||||
|
||||
## What Changes
|
||||
|
||||
- 新增确定性 EvidenceGuard,校验 Draft 结构、Analysis ID、报告内部引用,以及当前 Run canonical Tool invocation 的生命周期、证据状态和 `agent_result`。
|
||||
- 从三类冻结 Tool projection 确定性生成按 Analysis 分组的 verified evidence snapshot;不读取 `raw_response`,不向 SemanticGuard 暴露 Tool Call ID 或 Redis 细节。
|
||||
- 首次 EvidenceGuard 失败时允许一次无 Tool、无 ReAct 的结构修复;修复后仍失败直接返回 `EVIDENCE_VALIDATION_FAILED`。
|
||||
- 新增复用系统同一 `ChatModel` 的隔离单轮 SemanticGuard,校验完整报告语义,只输出 `SUPPORTED|UNSUPPORTED + reason`,不生成或修改报告。
|
||||
- Harness 对 SemanticGuard 执行输入预算、模型/Token 预算、单次超时、取消、严格输出解析和最多两次技术 attempt;`UNSUPPORTED` 不重试。
|
||||
- 新增 release use case:`SUPPORTED` 原样释放 Draft,业务不支持或技术不可用返回固定 SafeFallback,任何失败均不泄漏 Draft 或 SemanticGuard reason。
|
||||
- 不切换公开 Chat/AiOps,不删除旧 Gatekeeper/Verifier/Composer,不修改 SSE 或持久化协议。
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `single-react-evidence-semantic-guards`: 定义 DiagnosisDraft 的物理验真、verified snapshot、单次结构修复、隔离语义审查和安全释放行为。
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- None. 既有 Agent、Tool、Harness Core 和公开 Chat capability 的需求语义不在本阶段改变。
|
||||
|
||||
## Impact
|
||||
|
||||
- 新增 `com.superbiz.agent.harness.guard.evidence`、`guard.semantic` 和 `release` 内部包,以及 SemanticGuard prompt 和 focused tests。
|
||||
- 复用 `CanonicalInvocationStore`、`ToolCallKeyFactory`、三类 Tool contract、`DiagnosisHarnessCore`、`HarnessRetryExecutor`、`RunContext`、`DiagnosisDraft`、`SafeFallback` 和系统 `ChatModel`。
|
||||
- 内部接口影响为 L2:阶段 6A 将消费新的 release use case;本阶段无公开 Controller、SSE、DTO、数据库或旧 ChatService 行为变化。
|
||||
- 主要风险是快照字段遗漏、超时任务残留、错误重试业务 `UNSUPPORTED`,以及 Fallback 意外携带未验证内容;设计和测试必须逐项 fail closed。
|
||||
+112
@@ -0,0 +1,112 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Deterministic Draft and evidence validation
|
||||
The Harness SHALL deterministically reject a DiagnosisDraft unless every Analysis has a unique non-blank Analysis ID, a supported kind, non-blank text, and at least one Tool Call ID, and every non-null Conclusion, Action Plan item, and Recommendation has non-empty references to existing Analysis IDs.
|
||||
|
||||
#### Scenario: Duplicate or missing Analysis ID
|
||||
- **WHEN** a Draft contains a blank or duplicate Analysis ID
|
||||
- **THEN** EvidenceGuard returns violations and SemanticGuard is not invoked
|
||||
|
||||
#### Scenario: Broken report reference
|
||||
- **WHEN** a Conclusion, Action Plan item, or Recommendation has an empty or unknown Analysis reference
|
||||
- **THEN** EvidenceGuard rejects the Draft before semantic review
|
||||
|
||||
### Requirement: Current Run canonical invocation ownership
|
||||
EvidenceGuard SHALL resolve each referenced Tool Call through `runId + toolCallId` and SHALL accept only an invocation owned by the current Run with lifecycle `READY`, a non-empty `agent_result`, and evidence status `EVIDENCE_FOUND` or `NO_EVIDENCE`.
|
||||
|
||||
#### Scenario: Fabricated or cross-Run Tool Call
|
||||
- **WHEN** a Draft references a missing Tool Call or the resolved record belongs to another Run
|
||||
- **THEN** EvidenceGuard rejects the reference and does not expose any record content to SemanticGuard
|
||||
|
||||
#### Scenario: Failed or incomplete invocation
|
||||
- **WHEN** a referenced invocation is `PROJECTING`, `ERROR`, lacks `agent_result`, or has `evidence_status=ERROR`
|
||||
- **THEN** EvidenceGuard rejects the Draft
|
||||
|
||||
### Requirement: Analysis kind matches evidence semantics
|
||||
EvidenceGuard SHALL permit `NORMAL` Analysis only with `EVIDENCE_FOUND` calls and SHALL permit `NEGATIVE_OBSERVATION` Analysis only with `NO_EVIDENCE` calls.
|
||||
|
||||
#### Scenario: Valid negative observation
|
||||
- **WHEN** a `NEGATIVE_OBSERVATION` references a READY `NO_EVIDENCE` projection with its query scope and zero-match data
|
||||
- **THEN** EvidenceGuard accepts the binding without interpreting it as proof of system health or root-cause exclusion
|
||||
|
||||
#### Scenario: Positive claim uses no-evidence result
|
||||
- **WHEN** a `NORMAL` Analysis references a `NO_EVIDENCE` invocation
|
||||
- **THEN** EvidenceGuard rejects the binding
|
||||
|
||||
### Requirement: Verified evidence snapshot is minimal and deterministic
|
||||
The Harness SHALL strictly parse only supported Tool projections and SHALL construct evidence grouped by Analysis ID from referenced `agent_result` and required bounded request scope. The snapshot MUST NOT contain Tool Call IDs, Redis keys, raw responses, or unreferenced invocations.
|
||||
|
||||
#### Scenario: Supported RAG, log, and MySQL projections
|
||||
- **WHEN** a Draft references valid RAG, log, or MySQL calls
|
||||
- **THEN** the snapshot contains the corresponding stable source, scope, timestamp, exact excerpt or bounded values grouped under the referencing Analysis
|
||||
|
||||
#### Scenario: Projection contract mismatch
|
||||
- **WHEN** a projection has an unknown Tool name, invalid JSON, mismatched Tool Call ID, or evidence status inconsistent with its canonical record
|
||||
- **THEN** EvidenceGuard fails closed
|
||||
|
||||
### Requirement: Evidence repair is single-turn and semantics-preserving
|
||||
On the first EvidenceGuard failure, the Harness SHALL allow exactly one direct no-Tool model call to repair identifier and reference structure. It MUST NOT rerun the Diagnosis Agent or any Tool, and MUST reject a repair that changes user-visible report semantics.
|
||||
|
||||
#### Scenario: Structural repair succeeds
|
||||
- **WHEN** the one repair attempt changes only IDs/references and the repaired Draft passes EvidenceGuard
|
||||
- **THEN** the repaired Draft proceeds to SemanticGuard
|
||||
|
||||
#### Scenario: Repair changes report text
|
||||
- **WHEN** the repair changes Conclusion, Analysis, Action Plan, Recommendation, Limitation text, kind, order, or human-confirmation flag
|
||||
- **THEN** the Harness returns `EVIDENCE_VALIDATION_FAILED`
|
||||
|
||||
#### Scenario: Second validation fails
|
||||
- **WHEN** the repaired Draft still fails EvidenceGuard
|
||||
- **THEN** the Harness returns `EVIDENCE_VALIDATION_FAILED` with empty verified sources and does not invoke SemanticGuard
|
||||
|
||||
### Requirement: Isolated single-turn SemanticGuard
|
||||
SemanticGuard SHALL reuse the system ChatModel through a fresh single-turn Prompt containing only the original Query, the complete user-visible Draft without Tool Call IDs, and the verified evidence snapshot. It MUST have no Tool, memory, ReAct loop, Redis access, raw response, or callback to the Diagnosis Agent.
|
||||
|
||||
#### Scenario: Semantic input isolation
|
||||
- **WHEN** a verified Draft enters SemanticGuard
|
||||
- **THEN** the model sees the original Query, all report sections and verified evidence, but no Tool Call ID, Redis key, raw response, diagnosis history, or Tool definition
|
||||
|
||||
#### Scenario: Binary review output
|
||||
- **WHEN** SemanticGuard completes normally
|
||||
- **THEN** it returns only `SUPPORTED` or `UNSUPPORTED` with a non-blank audit reason and cannot return a corrected report
|
||||
|
||||
### Requirement: Semantic model budgets timeout cancellation and retry
|
||||
The Harness SHALL enforce input/output byte limits, Run byte/model/token budgets, per-attempt timeout, total SemanticGuard timeout, Run cancellation, strict JSON parsing, and the configured two-attempt technical retry policy. It SHALL retry only timeout, transport, parse, or schema failures and SHALL use the exact same input for both attempts.
|
||||
|
||||
#### Scenario: Technical failure then success
|
||||
- **WHEN** the first SemanticGuard attempt times out or returns invalid output and the second attempt returns a valid verdict
|
||||
- **THEN** exactly two model attempts are recorded and the second verdict controls release
|
||||
|
||||
#### Scenario: Unsupported is not retried
|
||||
- **WHEN** SemanticGuard returns valid `UNSUPPORTED`
|
||||
- **THEN** the Harness records one attempt and immediately applies the unsupported fallback
|
||||
|
||||
#### Scenario: Run cancellation during model call
|
||||
- **WHEN** the Run is cancelled while a guard model call is pending
|
||||
- **THEN** the Future is cancelled, no late model result is released, and cancellation is not converted into a normal Fallback
|
||||
|
||||
### Requirement: Fail-closed release policy
|
||||
The release use case SHALL publish the unchanged verified Draft only for `SUPPORTED`. It SHALL publish fixed `SafeFallback` content for evidence failure, semantic unsupported, or final semantic technical failure, and MUST NOT include the Draft, full verified snapshot, or SemanticGuard reason in a fallback release result.
|
||||
|
||||
#### Scenario: Supported report release
|
||||
- **WHEN** EvidenceGuard succeeds and SemanticGuard returns `SUPPORTED`
|
||||
- **THEN** release outcome is `SUCCESS` and the same verified Draft semantics are returned without summarization or partial editing
|
||||
|
||||
#### Scenario: Unsupported report fallback
|
||||
- **WHEN** SemanticGuard returns `UNSUPPORTED`
|
||||
- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNSUPPORTED`, and verified sources are derived only from the snapshot
|
||||
|
||||
#### Scenario: Semantic review remains unavailable
|
||||
- **WHEN** all permitted technical attempts fail
|
||||
- **THEN** release outcome is `FALLBACK`, type is `SEMANTIC_UNAVAILABLE`, and no Draft or internal failure reason is exposed
|
||||
|
||||
#### Scenario: Evidence validation fallback sources
|
||||
- **WHEN** evidence repair fails or the second EvidenceGuard rejects the Draft
|
||||
- **THEN** release outcome is `FALLBACK`, type is `EVIDENCE_VALIDATION_FAILED`, and `verified_sources` is empty
|
||||
|
||||
### Requirement: Stage-five public isolation
|
||||
The stage-five implementation SHALL remain internal and MUST NOT switch public Chat, AiOps, SSE, persistence, or legacy multi-Agent behavior.
|
||||
|
||||
#### Scenario: Focused implementation scope
|
||||
- **WHEN** stage-five changes are inspected
|
||||
- **THEN** only internal guard/release code, prompts, tests, OpenSpec and devflow artifacts have changed
|
||||
@@ -0,0 +1,25 @@
|
||||
## 1. Evidence validation and snapshot
|
||||
|
||||
- [x] 1.1 Add typed EvidenceViolation, EvidenceGuardResult and verified snapshot contracts with immutable source extraction.
|
||||
- [x] 1.2 Implement Draft structure/reference validation and current-Run canonical invocation checks.
|
||||
- [x] 1.3 Implement strict RAG/log/MySQL projection adapters that build ID-free, raw-free evidence grouped by Analysis ID.
|
||||
- [x] 1.4 Add focused tests for duplicate/missing IDs, empty/broken references, fabricated/cross-Run calls, invalid lifecycle/result, kind/status mismatch and valid negative observations.
|
||||
|
||||
## 2. Guard model boundary and semantic review
|
||||
|
||||
- [x] 2.1 Add the shared single-turn GuardModelCall with Core model/Token budgets, UTF-8 limits, per-attempt timeout, total timeout support and Future cancellation.
|
||||
- [x] 2.2 Add SemanticDraftView and SemanticGuardInput that preserve all user-visible report semantics while excluding Tool Call IDs.
|
||||
- [x] 2.3 Implement strict binary SemanticGuard output parsing and HarnessRetryExecutor integration with stable attempt auditing.
|
||||
- [x] 2.4 Add tests proving input isolation, same-model direct calls, retryable technical failures, non-retried UNSUPPORTED, timeout cancellation and no Tool/ReAct loop.
|
||||
|
||||
## 3. Evidence repair and release policy
|
||||
|
||||
- [x] 3.1 Implement one-attempt no-Tool EvidenceRepair and reject any repair that changes SemanticDraftView.
|
||||
- [x] 3.2 Implement fixed SafeFallbackFactory and release result contracts without Draft/reason parameters on fallback paths.
|
||||
- [x] 3.3 Implement DiagnosisReleaseUseCase sequencing Guard, optional repair, SemanticGuard and fail-closed outcomes while propagating cancellation/budget exhaustion.
|
||||
- [x] 3.4 Add release tests for successful repair, second validation failure, semantic support/unsupported/unavailable, fallback types, empty evidence-failure sources and no Draft/reason leakage.
|
||||
|
||||
## 4. Verification and scope
|
||||
|
||||
- [x] 4.1 Run stage-five focused tests plus Harness/Tool/Agent regression tests and Maven compile.
|
||||
- [x] 4.2 Verify strict OpenSpec validation, no public Controller/ChatService/AiOps diff, no raw response/Tool ID in semantic input, and no hidden retry or handwritten Agent loop.
|
||||
@@ -0,0 +1 @@
|
||||
Archive-ready after implementation and focused verification on 2026-07-21.
|
||||
@@ -0,0 +1 @@
|
||||
Committed after strict validation on 2026-07-21.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-21
|
||||
@@ -0,0 +1,109 @@
|
||||
## Context
|
||||
|
||||
现有 Chat/AIOps 在进入 Agent 前自行生成 session/run 标识,通过多个 ThreadLocal 和 RunnableConfig metadata 混合传播,并在业务方法内直接处理 retry loop、运行持久化和成功/失败终态。`TokenTrackingChatModel` 也只把 total token 写到 ThreadLocal,无法在异步或并发边界可靠归属 Run。后续 Tool interceptor、ResultProjector、Diagnosis Agent、EvidenceGuard、SemanticGuard 和 Chat application use case 需要共同使用一个不依赖旧 ChatService 的显式执行上下文。
|
||||
|
||||
Spring AI 1.1.7 的全局 retry properties 前缀为 `spring.ai.retry`,默认 `maxAttempts=10`。ISS-014 已确认所有底层模型 attempt 必须压为一次,由 Harness 对 Router/SemanticGuard 的特定技术失败显式执行最多一次重试并记录每次 attempt。
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- 提供显式、结构不可变、可跨同步/异步边界传递的 RunContext。
|
||||
- 在内存执行层统一 deadline、取消、预算和唯一终态语义。
|
||||
- 提供无隐藏循环的类型化 retry policy/executor 和 attempt 记录。
|
||||
- 提供阶段 3A 可直接复用的 Tool Call Key Factory 与单 Run 字节容量计数器。
|
||||
- 关闭 Spring AI 默认隐藏重试,保证实际模型 attempt 可由 Harness 观察。
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- 不接入 Controller、SSE、旧 ChatService、AiOpsService、ReactAgent 或现有 Tool。
|
||||
- 不替换或删除旧 ThreadLocal;阶段 7 在新入口切换后清理。
|
||||
- 不修改 `diagnosis_run` 实体、表结构或 Repository,不写数据库终态。
|
||||
- 不实现 Redis invocation store、ToolInterceptor、ToolResultProjector 或 Tool Call ID 校验全套策略。
|
||||
- 不实现定时调度器、线程中断、工作流引擎、动态配置中心、Retry DSL 或退避算法。
|
||||
|
||||
## Decisions
|
||||
|
||||
### 1. RunContext 是结构不可变的共享状态句柄集合
|
||||
|
||||
`RunContext` 使用 Java 17 record,字段固定为 `sessionId`、`runId`、`deadline`、`RunCancellation`、`RunBudget`、`HarnessRetryPolicies` 和 `RunLifecycle`。record 不提供 setter;同一 Run 的同步/异步消费者必须显式传递同一个 context,因此共享相同的取消、预算和生命周期句柄。
|
||||
|
||||
替代方案是把所有字段做成纯值并在每次变化时复制 context;这会造成并发分支状态分叉,无法保证唯一终态和原子预算,故不采用。ThreadLocal/RunnableConfig fallback 也不进入新 Core。
|
||||
|
||||
### 2. DiagnosisHarnessCore 是唯一执行门禁与状态转换入口
|
||||
|
||||
Core 由 `Clock`、Run ID supplier、最大 Run duration、显式 `RunBudgetLimits` 和 `HarnessRetryPolicies` 构造。它负责创建 context、在每个模型/Tool/容量边界检查 active/deadline、预留预算、记录 Token、显式取消以及完成成功/失败终态。
|
||||
|
||||
Core 不保存全局 Run map;生命周期完全随 RunContext 所有,避免跨请求泄漏。调用方可以传入既有 runId 进行确定性测试/恢复,但 Core 不生成 sessionId,也不修改业务持久化。
|
||||
|
||||
### 3. Deadline 与取消采用协作式、可观察语义
|
||||
|
||||
`RunCancellation` 用 atomic first-reason-wins 保存 `CLIENT_DISCONNECTED`、`USER_REQUESTED`、`DEADLINE_EXCEEDED`、`BUDGET_EXHAUSTED` 或 `INTERNAL_FAILURE`,并允许注册资源取消回调。Core 在每个受控边界比较 `Clock.instant()` 与 deadline;到达或超过 deadline 时先写 `TIMED_OUT` 终态,再触发取消信号并拒绝后续执行。
|
||||
|
||||
取消回调异常会记录并继续通知其他回调,不能阻止取消传播。此模型不承诺已进入的同步第三方调用立即停止;具体 HTTP/JDBC future 取消由后续 adapter 注册回调实现。
|
||||
|
||||
### 4. RunLifecycle 使用 first-terminal-wins
|
||||
|
||||
`RunLifecycle` 初始为 `RUNNING`,终态固定为 `SUCCESS`、`FAILED`、`CANCELLED`、`TIMED_OUT`、`BUDGET_EXHAUSTED`。内部使用 atomic compare-and-set 保存唯一 `RunTermination(state, reason, completedAt)`;任何后续 finish 都返回 false 且不得覆盖首个终态。
|
||||
|
||||
这只是执行内存真理;阶段 6A 的应用用例负责把终态映射到 `diagnosis_run.status/release_outcome`。本阶段不创建第二套数据库状态机。
|
||||
|
||||
### 5. RunBudget 使用显式 limits 和一致性更新
|
||||
|
||||
`RunBudgetLimits` 必须由调用方显式提供,包含模型调用、总 Tool 调用、单 Tool 调用、input/output/total Token 和 max Run bytes,所有值必须为正。`RunBudget` 使用同步临界区保证总 Tool/单 Tool 计数和三类 Token 计数要么一致提交,要么明确抛出 `BudgetExceededException`;模型返回后的实际 Token 即使超限也会先记入 usage,再终止后续执行。
|
||||
|
||||
`RunCapacityCounter` 用 CAS 原子预留 UTF-8 JSON 字节,失败时不部分增长。它被 RunBudget 持有并可由阶段 3A 直接调用。Core 在预算失败时先写 `BUDGET_EXHAUSTED` 终态,再触发取消。
|
||||
|
||||
### 6. Retry policy 只描述 attempt 上限与类型
|
||||
|
||||
`RetryPolicy(maxAttempts, retryableFailures)` 只允许 1 或 2 attempts。`HarnessRetryPolicies.strict()` 固定:Router=2(timeout/transport/invalid output)、Diagnosis=1、Tool=1、SemanticGuard=2(timeout/transport/parse/schema)、EvidenceRepair=1。
|
||||
|
||||
`HarnessRetryExecutor` 在每次 attempt 前调用 Core active check,操作成功/失败都发送 `RetryAttempt` 给 recorder。只有 policy 允许且尚有 attempt 时继续;取消、deadline、预算、`NO_EVIDENCE`、业务拒绝和未知失败不得重试。Executor 不 sleep、不退避、不递归、不调用 Agent。
|
||||
|
||||
### 7. 底层 Spring AI retry 固定为一次
|
||||
|
||||
`application.yml` 设置 `spring.ai.retry.max-attempts: 1`,并通过直接读取 YAML 的单元测试锁定,避免启动完整外部基础设施。该配置会立即影响旧运行路径:瞬时模型错误不再由 SDK 隐式重试;这是保证 attempt 可观测的有意内部行为变化。
|
||||
|
||||
### 8. ToolCallKeyFactory 不拥有 ID
|
||||
|
||||
Factory 接收可配置前缀,验证 `runId/toolCallId` 为非空、安全长度和安全字符 segment 后构造 `{prefix}:{runId}:{toolCallId}`。它不生成、规范化或哈希框架 Tool Call ID,不访问 Redis。严格的 duplicate/cross-run invocation 语义属于阶段 3A store。
|
||||
|
||||
## Module Flow
|
||||
|
||||
```text
|
||||
future Chat application use case
|
||||
-> DiagnosisHarnessCore.startRun(...)
|
||||
-> RunContext
|
||||
-> RunCancellation
|
||||
-> RunBudget -> RunCapacityCounter
|
||||
-> HarnessRetryPolicies
|
||||
-> RunLifecycle
|
||||
-> future Model/Tool/Guard boundary explicitly receives RunContext
|
||||
-> Core check/reserve/record/finish
|
||||
-> HarnessRetryExecutor for allowed technical operations only
|
||||
-> future application use case maps the one RunTermination to diagnosis_run
|
||||
```
|
||||
|
||||
阶段 3A 直接依赖 RunContext、Core、Key Factory 和 capacity counter;它不需要调用旧 ChatService 或读取 ThreadLocal。
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [结构不可变但句柄可变容易被误用] -> 命名、Javadoc 和并发/异步测试明确同一 Run 必须共享同一 context 实例。
|
||||
- [关闭 SDK retry 后旧路径瞬时失败率可能上升] -> 明确列为行为变化,配置测试锁定;后续 Router/SemanticGuard 只按已确认策略补回可观测重试。
|
||||
- [实际 Token 在响应后才知道,可能超预算] -> usage 保留实际值并立即进入预算终态,不丢失消耗、不再执行下一步。
|
||||
- [取消回调由 adapter 决定能否硬取消] -> Core 只承诺协作式取消和后续门禁,后续 JDBC/HTTP adapter 必须注册资源回调。
|
||||
- [纯内存终态无法替代审计] -> 阶段 6A 持久化映射;本阶段 focused tests 只验证执行语义。
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. 本阶段新增零消费者 Core 类型和 tests,同时把 Spring AI retry 压为一次。
|
||||
2. 阶段 3A 的 Tool interceptor/store 显式接收 RunContext,复用 Key/容量门禁。
|
||||
3. 阶段 4/5 的 Agent/Guard 使用 Core model/retry/deadline 边界。
|
||||
4. 阶段 6A 的 Chat application use case 创建 RunContext 并映射终态到数据库。
|
||||
5. 阶段 6B 切换入口;阶段 7 删除 ThreadLocal 和旧重试循环。
|
||||
|
||||
回滚本阶段代码可删除新包;`spring.ai.retry.max-attempts` 若回滚到默认 10 会恢复隐藏重试,但会再次失去 attempt 可观测性,因此只允许在整体重构回滚时明确执行。
|
||||
|
||||
## Open Questions
|
||||
|
||||
无。具体预算数值、Run ID 格式和持久化状态映射由后续 wiring/应用用例在不改变本 Core 语义的前提下配置。
|
||||
@@ -0,0 +1,46 @@
|
||||
## Why
|
||||
|
||||
当前 Chat/AIOps 通过 `SessionContextHolder`、`VerifierContextHolder` 和 `TokenUsageHolder` 等 ThreadLocal 隐式传播 session/run/token/retry 状态,且 Run 终态、取消、预算和重试分散在业务循环与 SDK 默认行为中。阶段 3A 的 Tool boundary、后续 Agent/Guard 和最终入口需要先共同依赖一个显式、可跨同步/异步边界传递的 Harness RunContext,否则会继续耦合旧 ChatService 并产生重复状态机。
|
||||
|
||||
## What Changes
|
||||
|
||||
- 新增结构不可变的 `RunContext`,显式携带 `sessionId`、`runId`、deadline、取消信号、线程安全预算、类型化重试策略和 first-terminal-wins 生命周期。
|
||||
- 新增 `DiagnosisHarnessCore` 负责 Run 创建、deadline 检查、模型/Tool/Token/容量预算门禁、取消传播和唯一终态,不承担业务推理或持久化。
|
||||
- 新增 caller-supplied `RunBudgetLimits`、线程安全 `RunBudget` 与单 Run 字节容量计数器;Core 不硬编码尚未校准的预算默认值。
|
||||
- 新增 `HarnessRetryPolicies` 与 `HarnessRetryExecutor`,只允许 Router/SemanticGuard 的类型化技术失败最多两次 attempt;Diagnosis Agent、Tool 和 Evidence repair 固定一次 attempt。
|
||||
- 将 Spring AI 1.1.7 全局底层重试 `spring.ai.retry.max-attempts` 从默认 10 压为 1,避免与 Harness 形成隐藏嵌套重试。
|
||||
- 新增 Redis Tool Call Key Factory,只负责安全构造固定前缀下的 `runId + toolCallId` Key,不访问 Redis、不生成 Tool Call ID。
|
||||
- 使用 Fake Model/Tool 和可控 Clock 验证显式异步传播、deadline、客户端取消、预算耗尽、重试记录、Key/容量边界和唯一 Run 终态。
|
||||
- 本阶段不改 Controller/HTTP/SSE,不接入或临时适配旧 ChatService,不修改 `diagnosis_run` 持久化,也不实现 Tool-specific 投影或 Redis store。
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `diagnosis-harness-run-context`: 提供显式 RunContext、Harness Core、预算、取消、deadline、类型化重试、唯一终态和后续 Tool store 所需的 Key/容量基础。
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- None. 本阶段不声明旧 Chat/AIOps 运行路径已迁移;公开运行和数据库状态映射在后续 change 接入。
|
||||
|
||||
## Context Constraints
|
||||
|
||||
- 新代码不得读取或写入任何 ThreadLocal;RunContext 只能通过参数或框架受控 context 显式传播。
|
||||
- RunContext 的结构不可变,但其 cancellation、budget 和 lifecycle 是线程安全的单 Run 状态句柄。
|
||||
- 同一 Run 只允许第一个终态生效;取消、deadline、预算和异常之间的竞态不得覆盖先到终态。
|
||||
- 同步模型/Tool 调用的取消只保证阻止后续执行并触发已注册资源取消回调,不虚假承诺无法证明的立即线程中断。
|
||||
- `NO_EVIDENCE`、业务拒绝、取消和预算耗尽不是可重试技术失败。
|
||||
- Tool Call Key 使用框架 `tool_call_id`;Key Factory 不生成、替换或回退到其他 ID。
|
||||
|
||||
## Interface Impact
|
||||
|
||||
- 等级:L2(内部 Harness 基础接口)。新增 Core 类型将被后续 3A-7 阶段消费,但本 change 不修改现有调用方。
|
||||
- `application.yml` 的 Spring AI retry 从默认 10 attempts 变为 1 attempt,属于有意的内部运行配置变化;当前旧调用若遇到瞬时模型失败将不再由 SDK 隐式重试,避免未记录 attempts。Harness 允许的 Router/SemanticGuard 重试要到后续接入后显式执行和记录。
|
||||
- 不改变 Controller、SSE、DTO、数据库 Schema 或公开错误码。
|
||||
|
||||
## Risks
|
||||
|
||||
- 旧运行路径在切换前会失去 SDK 隐式重试但尚未使用 Harness retry;这是为保证“所有 attempt 可观测”接受的短期行为变化,focused tests 需证明启动配置正确。
|
||||
- RunContext 内含线程安全可变句柄,若被误解为纯值对象可能错误复制;设计和测试必须明确同一 Run 共享同一状态句柄。
|
||||
- Budget 在模型调用后才能取得实际 Token,可能在记录实际消耗时才发现超限;Core 必须保留实际 usage 并立即终止后续执行。
|
||||
- 本阶段不写数据库,内存生命周期只服务一次调用链;持久化终态映射由 Chat application use case 阶段负责。
|
||||
+89
@@ -0,0 +1,89 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: RunContext SHALL be explicit and structurally immutable
|
||||
The Harness SHALL create a structurally immutable RunContext containing non-blank `sessionId`, non-blank `runId`, deadline, cancellation, budget, retry policies, and lifecycle. Every new Model, Tool, Guard, and application boundary SHALL receive the RunContext explicitly by parameter or framework-controlled context and SHALL NOT read or write ThreadLocal state.
|
||||
|
||||
#### Scenario: Context crosses an asynchronous boundary
|
||||
- **WHEN** a Fake Tool runs on another thread with an explicitly supplied RunContext
|
||||
- **THEN** it observes the same session ID, run ID, deadline, cancellation, budget, retry policies, and lifecycle handles without copying ThreadLocal state
|
||||
|
||||
#### Scenario: Invalid identity starts a Run
|
||||
- **WHEN** Run creation receives a blank session ID or run ID
|
||||
- **THEN** the Core rejects creation before any budget, cancellation, or lifecycle state is published
|
||||
|
||||
### Requirement: Deadline and cancellation SHALL stop subsequent controlled work
|
||||
The Core SHALL compare the current Clock instant with the Run deadline before each controlled Model, Tool, retry, and capacity boundary. At or after the deadline it SHALL produce `TIMED_OUT`, signal `DEADLINE_EXCEEDED`, and reject subsequent work. Explicit cancellation SHALL preserve the first cancellation reason, notify all registered resource callbacks, and produce `CANCELLED` unless an earlier terminal state exists.
|
||||
|
||||
#### Scenario: Deadline expires before a Tool call
|
||||
- **WHEN** the controlled Clock reaches the Run deadline before the Fake Tool boundary
|
||||
- **THEN** the Tool is not invoked, the Run terminates as `TIMED_OUT`, and later checks return the same terminal outcome
|
||||
|
||||
#### Scenario: Client disconnects
|
||||
- **WHEN** the Core receives `CLIENT_DISCONNECTED`
|
||||
- **THEN** registered cancellation callbacks run, subsequent Model/Tool boundaries are rejected, and the Run has one `CANCELLED` terminal outcome
|
||||
|
||||
### Requirement: Run budgets SHALL be explicit, thread-safe, and observable
|
||||
The Core SHALL require caller-supplied positive limits for Model calls, total Tool calls, per-Tool calls, input Tokens, output Tokens, total Tokens, and Run bytes. Budget updates SHALL be thread-safe and preserve a usage snapshot. Exceeding any limit SHALL record the actual applicable usage, produce `BUDGET_EXHAUSTED`, signal cancellation, and prevent subsequent controlled work.
|
||||
|
||||
#### Scenario: Per-Tool budget is exhausted
|
||||
- **WHEN** a Fake Tool attempts one call beyond its configured per-Tool limit
|
||||
- **THEN** the extra call is not reserved, the total and per-Tool usage remain internally consistent, and the Run terminates as `BUDGET_EXHAUSTED`
|
||||
|
||||
#### Scenario: Actual Token usage exceeds the limit
|
||||
- **WHEN** a completed Fake Model response reports Token usage above the configured limit
|
||||
- **THEN** the actual input/output/total usage remains recorded and the Core rejects the next controlled operation
|
||||
|
||||
### Requirement: A Run SHALL have exactly one terminal outcome
|
||||
The lifecycle SHALL start as `RUNNING` and accept only the first terminal outcome from `SUCCESS`, `FAILED`, `CANCELLED`, `TIMED_OUT`, or `BUDGET_EXHAUSTED`. Later success, failure, cancellation, deadline, or budget events SHALL NOT replace the first `RunTermination` state, reason, or timestamp.
|
||||
|
||||
#### Scenario: Success wins a terminal race
|
||||
- **WHEN** success is recorded before a later failure and cancellation
|
||||
- **THEN** the lifecycle remains `SUCCESS` and both later finish attempts report that they did not change the terminal outcome
|
||||
|
||||
#### Scenario: Budget exhaustion wins a terminal race
|
||||
- **WHEN** budget exhaustion is recorded before an exception handler reports failure
|
||||
- **THEN** the lifecycle remains `BUDGET_EXHAUSTED` with its original reason and completion time
|
||||
|
||||
### Requirement: Harness retry SHALL be typed, bounded, and recorded
|
||||
The Harness SHALL provide immutable policies and an executor that records every actual attempt. Intent Router and SemanticGuard SHALL allow at most two attempts for their listed technical failures; Diagnosis Agent, Tool calls, and Evidence repair SHALL allow one attempt. Cancellation, deadline, budget exhaustion, `NO_EVIDENCE`, business rejection, and unknown failures SHALL NOT be retried.
|
||||
|
||||
#### Scenario: Router transport failure succeeds on retry
|
||||
- **WHEN** a Fake Router fails its first attempt with a retryable transport failure and succeeds on the second
|
||||
- **THEN** the executor returns the second result and records exactly one failed and one successful attempt
|
||||
|
||||
#### Scenario: Tool call fails
|
||||
- **WHEN** a Fake Tool operation fails on its first attempt
|
||||
- **THEN** the Tool policy records one failed attempt and throws without invoking the Tool again
|
||||
|
||||
#### Scenario: Run is cancelled between attempts
|
||||
- **WHEN** cancellation is signalled after a retryable first failure
|
||||
- **THEN** the executor rejects the next attempt and does not invoke the operation again
|
||||
|
||||
### Requirement: Spring AI hidden retry SHALL be disabled
|
||||
Application configuration SHALL set `spring.ai.retry.max-attempts=1`, overriding the Spring AI 1.1.7 default of 10. Harness-permitted retries SHALL occur only in HarnessRetryExecutor and SHALL be observable as separate attempts.
|
||||
|
||||
#### Scenario: Retry configuration is loaded
|
||||
- **WHEN** the repository application YAML is parsed in a focused configuration test
|
||||
- **THEN** `spring.ai.retry.max-attempts` equals 1
|
||||
|
||||
### Requirement: Tool invocation key and Run capacity primitives SHALL be safe and store-independent
|
||||
The Harness SHALL provide a ToolCallKeyFactory that constructs `{prefix}:{runId}:{toolCallId}` only from valid safe segments and never generates or rewrites the framework Tool Call ID. The Run capacity counter SHALL atomically reserve positive bytes up to its configured limit and SHALL leave usage unchanged when a reservation is rejected.
|
||||
|
||||
#### Scenario: Framework Tool Call ID is used in a key
|
||||
- **WHEN** a valid run ID and framework Tool Call ID are supplied
|
||||
- **THEN** the factory returns the configured prefix followed by the exact run ID and exact Tool Call ID
|
||||
|
||||
#### Scenario: Unsafe key segment is supplied
|
||||
- **WHEN** either ID is blank, too long, or contains a separator or unsafe character
|
||||
- **THEN** the factory rejects it and does not produce a Redis key
|
||||
|
||||
#### Scenario: Run byte limit would be exceeded
|
||||
- **WHEN** a reservation would exceed max Run bytes
|
||||
- **THEN** the counter rejects it and retains the usage from previously successful reservations
|
||||
|
||||
### Requirement: Harness Core SHALL remain independent from current runtime and persistence
|
||||
This change SHALL NOT modify Controller/HTTP/SSE contracts, connect RunContext to the old ChatService or AiOpsService, use current ThreadLocal holders, persist `diagnosis_run`, access Redis, or implement Tool-specific projection. The new Core SHALL compile and pass focused Fake Model/Tool tests as a zero-consumer foundation for stage 3A.
|
||||
|
||||
#### Scenario: Stage 2 implementation completes
|
||||
- **WHEN** focused Core, retry, budget, key, capacity, and configuration tests pass
|
||||
- **THEN** current Chat/AIOps call sites remain unchanged and stage 3A can depend directly on the new Harness types
|
||||
@@ -0,0 +1,38 @@
|
||||
## 1. Run State Primitives
|
||||
|
||||
- [x] 1.1 Implement structurally immutable RunContext with explicit identity, deadline, cancellation, budget, retry policy, and lifecycle handles.
|
||||
- [x] 1.2 Implement first-reason-wins cancellation with resource callbacks and first-terminal-wins Run lifecycle.
|
||||
- [x] 1.3 Add Run identity, cancellation, and terminal race tests with a controlled Clock.
|
||||
|
||||
## 2. Budget and Capacity
|
||||
|
||||
- [x] 2.1 Implement validated caller-supplied RunBudgetLimits and thread-safe Model/Tool/per-Tool/Token accounting.
|
||||
- [x] 2.2 Implement atomic single-Run byte capacity reservation with no partial update on rejection.
|
||||
- [x] 2.3 Add budget tests for consistent counters, actual Token recording, concurrent reservations, and exhaustion details.
|
||||
|
||||
## 3. Harness Core
|
||||
|
||||
- [x] 3.1 Implement DiagnosisHarnessCore Run creation, active/deadline checks, budget gates, explicit cancellation, and success/failure completion.
|
||||
- [x] 3.2 Add Fake Model/Tool tests proving explicit synchronous/asynchronous RunContext propagation and cancellation/deadline/budget terminal outcomes.
|
||||
|
||||
## 4. Retry Boundary
|
||||
|
||||
- [x] 4.1 Implement typed RetryFailure, bounded RetryPolicy, strict HarnessRetryPolicies, RetryAttempt, and RetryExecutionException.
|
||||
- [x] 4.2 Implement HarnessRetryExecutor with per-attempt active checks and attempt recording, without backoff, recursion, or hidden loops.
|
||||
- [x] 4.3 Add Fake Router/Tool/SemanticGuard retry tests for allowed technical retry, one-attempt policies, cancellation between attempts, and non-retryable outcomes.
|
||||
|
||||
## 5. Tool Store Foundations
|
||||
|
||||
- [x] 5.1 Implement configurable ToolCallKeyFactory with exact framework ID preservation and safe-segment validation.
|
||||
- [x] 5.2 Add key factory tests covering exact key format, blank/unsafe/oversized segments, and no generated fallback ID.
|
||||
|
||||
## 6. Hidden Retry Configuration
|
||||
|
||||
- [x] 6.1 Set `spring.ai.retry.max-attempts` to 1 without changing Model routing or provider configuration.
|
||||
- [x] 6.2 Add a focused YAML configuration test proving the hidden Spring AI retry override.
|
||||
|
||||
## 7. Verification
|
||||
|
||||
- [x] 7.1 Run focused Harness Core, retry, budget, key/capacity, and retry configuration tests.
|
||||
- [x] 7.2 Run stage 0/1 Harness and ACI contract tests plus existing ChatController test as regression coverage.
|
||||
- [x] 7.3 Verify new production code has no ThreadLocal/current-holder usage and current Chat/AIOps/Controller/JPA/Redis call sites remain unchanged.
|
||||
@@ -0,0 +1,84 @@
|
||||
# Design: single-react-mysql-readonly-tool
|
||||
|
||||
## Architecture
|
||||
|
||||
```text
|
||||
MysqlToolRequest + ToolCallRequestEnvelope
|
||||
|
|
||||
v
|
||||
MysqlToolAdapter
|
||||
- parse typed request
|
||||
- MysqlSqlValidator -> MysqlQueryPlan
|
||||
- ToolBoundary.execute
|
||||
|
|
||||
+--> MysqlReadOnlyExecutor (JDBC or isolated fake)
|
||||
| - read-only connection
|
||||
| - PreparedStatement params
|
||||
| - timeout/max rows/cancel
|
||||
|
|
||||
+--> MysqlResultProjector
|
||||
- bounded cells/rows/bytes
|
||||
- sensitive column redaction
|
||||
- MysqlToolResult JSON
|
||||
```
|
||||
|
||||
The adapter validates the typed request before entering ToolBoundary. Once validation succeeds, the executor and projector run inside the existing canonical lifecycle. Validation failures have no canonical record because no database call is authorized; execution/projection failures after `PROJECTING` become canonical `ERROR` records.
|
||||
|
||||
## Configuration model
|
||||
|
||||
`MysqlDataSourceDefinition` is an immutable logical definition:
|
||||
|
||||
- logical ID
|
||||
- JDBC URL, username and password supplied by configuration/Secret
|
||||
- default schema
|
||||
- exact schema/table/column allowlist
|
||||
- query timeout seconds, max rows, max cell characters and max result bytes
|
||||
|
||||
The Agent sees only the logical ID. The stage does not create a dynamic datasource registry or reuse the application persistence datasource. A caller supplies a `Map<String, DataSource>` to the JDBC executor, allowing isolated test data sources and later production wiring.
|
||||
|
||||
## SQL validator
|
||||
|
||||
`MysqlSqlValidator` uses `CCJSqlParserUtil.parseStatements` and rejects unless there is exactly one `Select` with a `PlainSelect` body and no CTE/compound body. It rejects `SubSelect`, `SetOperationList`, `ValuesStatement`, `ForUpdate`, metadata statements, wildcard projection except `COUNT(*)`, unsupported joins, unsupported functions, and unknown table/from items.
|
||||
|
||||
The validator builds an alias map for every base table and checks every table/column reference in select items, joins, where, group by, having and order by. Unqualified columns must resolve to exactly one allowlisted table; qualified columns must resolve through a declared alias/table and exact allowlist. It counts `JdbcParameter` nodes and requires an exact match with `params`.
|
||||
|
||||
The function allowlist starts with `COUNT`, `SUM`, `AVG`, `MIN`, and `MAX`. `COUNT(*)` is represented as a function-level exception; a projection `*` or `table.*` is rejected. Any parser exception or policy ambiguity produces a `MysqlSecurityException` before execution.
|
||||
|
||||
## JDBC executor
|
||||
|
||||
`JdbcMysqlReadOnlyExecutor` resolves the validated logical ID to a caller-supplied DataSource, opens a connection, calls `setReadOnly(true)`, prepares the validated SQL, binds parameters in order, applies `setQueryTimeout` and `setMaxRows`, and reads column labels plus structured values. It checks `RunContext` cancellation/deadline while iterating and cancels/closes the statement on abort.
|
||||
|
||||
The executor returns a Harness-only raw record. It never returns a JDBC connection, SQL exception details, credentials or stack traces to the Agent. JDBC errors are mapped by ToolBoundary to a stable safe error.
|
||||
|
||||
## Result projection
|
||||
|
||||
`MysqlResultProjector` parses only the executor raw record and emits `MysqlToolResult`:
|
||||
|
||||
- immutable ordered columns
|
||||
- bounded ordered rows
|
||||
- cell values converted to JSON-safe scalar/text values
|
||||
- redaction for sensitive column names (`password`, `token`, `secret`, `api_key`, etc.)
|
||||
- `returned_count` equal to projected rows
|
||||
- `truncated=true` for row/cell/byte removal
|
||||
- `NO_EVIDENCE` for a successful empty result
|
||||
|
||||
Projection never exposes raw JDBC metadata, connection coordinates, internal error messages or the unbounded raw result.
|
||||
|
||||
## Script safety cleanup
|
||||
|
||||
`scripts/query_mysql.py` will require `SUPERBIZ_MYSQL_HOST`, `SUPERBIZ_MYSQL_PORT`, `SUPERBIZ_MYSQL_USERNAME`, `SUPERBIZ_MYSQL_PASSWORD` and `SUPERBIZ_MYSQL_DATABASE` (with no external defaults), accept only one SELECT statement, and reject all non-SELECT/metadata/write SQL before opening a connection. It will call `rollback`/close defensively and remove the interactive write path.
|
||||
|
||||
## Interface and compatibility impact
|
||||
|
||||
- L2 internal Harness classes plus the JSqlParser build dependency.
|
||||
- No public HTTP/SSE/Chat contract changes.
|
||||
- No changes to legacy tools, recorder, JPA entities, or current application datasource.
|
||||
- The main capability spec will be synchronized after archive.
|
||||
|
||||
## Risks and mitigations
|
||||
|
||||
- JSqlParser AST API drift: lock version 4.6 and run parser security fixtures.
|
||||
- Alias/column ambiguity: fail closed rather than guessing.
|
||||
- JDBC cancellation is driver-dependent: check Run state before/while iteration and call `Statement.cancel()` on abort.
|
||||
- Sensitive result values: redact by column name before JSON serialization and keep raw only in canonical Harness storage.
|
||||
- A malicious query can still be expensive within SELECT: timeout, max rows, read-only connection and configured byte limits remain mandatory.
|
||||
@@ -0,0 +1,54 @@
|
||||
# Proposal: single-react-mysql-readonly-tool
|
||||
|
||||
## Problem
|
||||
|
||||
阶段 1 已冻结 `query_mysql` 的逻辑请求和有界结果契约,但当前仓库没有一个能把 SQL 安全地转换为可执行查询的实现。直接把 Agent 生成的 SQL 交给 JDBC 会允许写操作、元数据探测、通配符泄漏、allowlist 绕过和无界结果。
|
||||
|
||||
## Proposed change
|
||||
|
||||
- 引入 JSqlParser 4.6,解析单条 SQL 并对保守的 SELECT 子集执行 fail-closed AST 校验。
|
||||
- 增加逻辑数据源定义和静态 `schema -> table -> column` 精确 allowlist;Agent 只能传逻辑数据源 ID。
|
||||
- 增加参数绑定、占位符数量校验、函数 allowlist、显式列校验、JOIN/过滤/分组/排序字段校验和受控 LIMIT。
|
||||
- 增加 JDBC 只读执行器:只使用配置映射的 DataSource,设置 `readOnly`、`PreparedStatement`、`setMaxRows`、查询超时,并在 Run 取消时中止执行。
|
||||
- 增加 MySQL 结果 projector:限制行数、单元格和总 UTF-8 字节,脱敏敏感列,产生 `MysqlToolResult` 和 `NO_EVIDENCE`。
|
||||
- 通过阶段 3A `ToolBoundary` 接入 canonical invocation store、Run budget、生命周期、`evidence_status` 和框架 `tool_call_id`。
|
||||
- 把 `scripts/query_mysql.py` 收敛为仅允许 SELECT/受控 SHOW 的只读查询脚本,移除默认外部连接参数和 commit 分支;凭据与连接信息全部从环境变量读取。
|
||||
|
||||
## Scope
|
||||
|
||||
### In scope
|
||||
|
||||
- SQL AST validator and validated query plan.
|
||||
- Static logical data-source/allowlist model.
|
||||
- JDBC read-only executor abstraction and implementation.
|
||||
- MySQL result projector and ToolBoundary adapter.
|
||||
- Security, truncation, timeout/cancellation, no-evidence and boundary tests.
|
||||
- Query script safety cleanup required by the ISS-014 security prerequisite.
|
||||
|
||||
### Out of scope
|
||||
|
||||
- Diagnosis Agent cutover or public Chat/AIOps/SSE changes.
|
||||
- Querying the application's own persistence database through the Agent-facing Tool.
|
||||
- Metadata discovery (`SHOW TABLES`, `SHOW COLUMNS`, `DESCRIBE`, `information_schema`).
|
||||
- Tenant/row-level authorization, dynamic allowlists, arbitrary SQL functions, real production business datasource provisioning.
|
||||
|
||||
## Frozen SQL subset
|
||||
|
||||
- One `SELECT` statement only.
|
||||
- Explicit projection columns; `COUNT(*)` is the only star exception.
|
||||
- `INNER JOIN` and `LEFT JOIN`, normal predicates, `GROUP BY`, `HAVING`, `ORDER BY` and parameter placeholders.
|
||||
- No `WITH`, subquery, `UNION`, window function, `CROSS JOIN`, write statement, metadata query, transaction control, dangerous function, or unknown AST node.
|
||||
|
||||
## Constraints and risks
|
||||
|
||||
- Parser acceptance is not authorization: every table and column used by projection, predicates, joins, grouping and ordering must pass the independent allowlist.
|
||||
- Unknown data source, schema, table, column, function, placeholder mismatch or parser failure fails closed before JDBC execution.
|
||||
- JDBC row/cell/byte limits are enforced in addition to ToolBoundary limits; oversized results are projected with `truncated=true` or become a safe error when no valid bounded result can be produced.
|
||||
- The adapter is not wired into the existing public runtime in this stage; later Diagnosis Agent work will select it through the internal Harness application use case.
|
||||
|
||||
## Acceptance direction
|
||||
|
||||
- Valid explicit-column SELECT and `COUNT(*)` pass AST and allowlist validation.
|
||||
- Write statements, metadata, wildcard projection, nested/compound queries, unknown AST, allowlist bypass and placeholder mismatch are rejected without executor invocation.
|
||||
- JDBC executor uses read-only prepared statements, max rows, timeout and cancellation.
|
||||
- Agent receives only bounded `MysqlToolResult`; raw JDBC rows remain Harness-only canonical data.
|
||||
+96
@@ -0,0 +1,96 @@
|
||||
# mysql-readonly-tool Specification
|
||||
|
||||
## Purpose
|
||||
|
||||
Define a fail-closed, parameterized, read-only MySQL evidence Tool that reuses the stage 3A ToolBoundary and exposes only bounded ACI results.
|
||||
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: SQL validation SHALL fail closed on a conservative SELECT subset
|
||||
|
||||
The Tool SHALL parse exactly one SQL statement with JSqlParser and SHALL accept only a single `SELECT` with explicit projection columns, supported predicates/grouping/ordering, `INNER JOIN` or `LEFT JOIN`, parameter placeholders and allowlisted functions. It SHALL reject writes, CTEs, subqueries, set operations, wildcard projections except `COUNT(*)`, metadata discovery, unsupported joins/functions, `FOR UPDATE`, multiple statements and unknown/ambiguous AST structures.
|
||||
|
||||
#### Scenario: Valid explicit-column SELECT
|
||||
|
||||
- WHEN a query selects allowlisted columns from an allowlisted table with matching `?` parameters
|
||||
- THEN validation returns a query plan and the executor may be invoked
|
||||
|
||||
#### Scenario: COUNT star exception
|
||||
|
||||
- WHEN a query uses `SELECT COUNT(*)` against an allowlisted table
|
||||
- THEN validation succeeds without treating the projection as an unrestricted wildcard
|
||||
|
||||
#### Scenario: Security query is rejected
|
||||
|
||||
- WHEN SQL contains INSERT/UPDATE/DELETE, `WITH`, a subquery, `UNION`, `SELECT *`, metadata discovery, `FOR UPDATE`, a dangerous function, multiple statements or an unknown AST node
|
||||
- THEN validation fails before executor invocation
|
||||
|
||||
### Requirement: Data-source and identifier authorization SHALL use exact independent allowlists
|
||||
|
||||
The Tool SHALL accept only a logical `data_source` ID and SHALL authorize every schema, table and column used in projection, join, predicate, grouping and ordering against the configured exact allowlist. It SHALL reject unknown data sources, schemas, tables, columns, aliases and ambiguous unqualified columns. Agent input SHALL NOT provide JDBC coordinates or authorization controls.
|
||||
|
||||
#### Scenario: Allowlisted query
|
||||
|
||||
- WHEN every referenced identifier resolves to one configured schema/table/column
|
||||
- THEN validation succeeds and retains the logical data-source ID only
|
||||
|
||||
#### Scenario: Allowlist bypass
|
||||
|
||||
- WHEN a query references an unconfigured table, column, schema, alias or ambiguous unqualified column
|
||||
- THEN validation fails closed and the executor is not called
|
||||
|
||||
### Requirement: Parameter binding and JDBC execution SHALL be read-only and bounded
|
||||
|
||||
The executor SHALL use a configured logical datasource, a read-only JDBC connection, `PreparedStatement` parameter binding, query timeout, max rows and Run cancellation/deadline checks. Placeholder count SHALL exactly match `params`. The executor SHALL not expose connection details or raw JDBC failures to the Agent.
|
||||
|
||||
#### Scenario: Bound read-only execution
|
||||
|
||||
- WHEN a validated plan has matching parameters and an active Run
|
||||
- THEN the executor binds values in order, sets read-only/timeout/max rows, and returns structured raw rows
|
||||
|
||||
#### Scenario: Timeout or cancellation
|
||||
|
||||
- WHEN the query exceeds its timeout or the Run is cancelled/deadline-expired
|
||||
- THEN the statement is cancelled/closed and ToolBoundary returns a safe error without an Agent result
|
||||
|
||||
### Requirement: MySQL projection SHALL be bounded and evidence-aware
|
||||
|
||||
The projector SHALL expose only ordered columns, bounded JSON-safe rows, returned count and truncation. It SHALL enforce row, cell and total UTF-8 limits, redact sensitive column values, return `NO_EVIDENCE` for a successful empty result, and never expose raw JDBC metadata or credentials.
|
||||
|
||||
#### Scenario: Bounded rows are projected
|
||||
|
||||
- WHEN the executor returns rows within configured limits
|
||||
- THEN the Agent receives `EVIDENCE_FOUND` with ordered columns and structured rows
|
||||
|
||||
#### Scenario: Result exceeds a limit
|
||||
|
||||
- WHEN row count, cell length or total bytes exceed a configured bound
|
||||
- THEN the Agent receives valid bounded JSON with `truncated=true` and no oversized value
|
||||
|
||||
#### Scenario: Empty result
|
||||
|
||||
- WHEN a valid query returns zero rows
|
||||
- THEN the boundary reaches `READY` with `evidence_status=NO_EVIDENCE` and an empty row list
|
||||
|
||||
### Requirement: MySQL Tool SHALL reuse canonical Harness ownership
|
||||
|
||||
The adapter SHALL pass the exact framework `tool_call_id` and RunContext through the existing ToolBoundary and canonical invocation store. It SHALL not create a second ID, use a parallel store, return raw SQL results, or modify legacy audit/public runtime paths.
|
||||
|
||||
#### Scenario: Successful adapter call
|
||||
|
||||
- WHEN a valid request passes validation and execution
|
||||
- THEN the canonical record contains request/raw/bounded result under the exact framework ID and the Agent receives only the bounded result
|
||||
|
||||
#### Scenario: Validation or boundary failure
|
||||
|
||||
- WHEN validation, authorization, duplicate, budget or lifecycle preflight fails
|
||||
- THEN no database call is made and the Agent receives a safe bounded error
|
||||
|
||||
### Requirement: The query helper script SHALL be read-only and secret-free by default
|
||||
|
||||
The repository query helper SHALL require connection values from environment variables, reject non-SELECT and metadata discovery SQL before connection, and SHALL NOT commit writes or expose hardcoded external connection defaults.
|
||||
|
||||
#### Scenario: Unsafe script query
|
||||
|
||||
- WHEN a caller passes a write, multi-statement, metadata or transaction command
|
||||
- THEN the script exits with a safe validation error without opening a connection
|
||||
@@ -0,0 +1,24 @@
|
||||
# Tasks: single-react-mysql-readonly-tool
|
||||
|
||||
## 1. Security prerequisite and dependency
|
||||
|
||||
- [x] 1.1 Add JSqlParser 4.6 and record the locked parser version.
|
||||
- [x] 1.2 Make `scripts/query_mysql.py` environment-only and read-only; remove commit and unsafe interactive paths.
|
||||
|
||||
## 2. SQL policy and allowlist model
|
||||
|
||||
- [x] 2.1 Implement immutable data-source/allowlist/limit models and validated query plan.
|
||||
- [x] 2.2 Implement JSqlParser single-SELECT validator, identifier resolution, function/wildcard/join policy and placeholder count checks.
|
||||
- [x] 2.3 Add parser and allowlist security fixtures for accepted/rejected SQL.
|
||||
|
||||
## 3. JDBC executor and projection
|
||||
|
||||
- [x] 3.1 Implement read-only PreparedStatement executor with timeout, max rows, Run cancellation and safe error mapping.
|
||||
- [x] 3.2 Implement `MysqlResultProjector` with row/cell/byte bounds, redaction, JSON-safe values and `NO_EVIDENCE`.
|
||||
- [x] 3.3 Add isolated executor/projector tests for valid rows, empty rows, truncation, sensitive columns and timeout/cancel behavior.
|
||||
|
||||
## 4. ToolBoundary adapter and verification
|
||||
|
||||
- [x] 4.1 Implement typed MySQL adapter through the existing ToolBoundary and canonical invocation store.
|
||||
- [x] 4.2 Add adapter tests proving exact framework ID, raw isolation, validation-before-execution and safe errors.
|
||||
- [x] 4.3 Run focused compile/tests, validate OpenSpec, update devflow evidence, archive and commit before stage 4.
|
||||
@@ -0,0 +1,69 @@
|
||||
# Design: single-react-rag-log-projections
|
||||
|
||||
## Architecture
|
||||
|
||||
```text
|
||||
typed ACI request + framework envelope
|
||||
|
|
||||
v
|
||||
RagToolAdapter / QueryLogsToolAdapter
|
||||
| (legacy executor is internal and request controls are injected here)
|
||||
v
|
||||
ToolBoundary.execute(..., raw -> requestAwareProjector.project(...))
|
||||
|
|
||||
+--> CanonicalInvocationStore (complete raw + bounded agent result)
|
||||
+--> ToolBoundaryResult (bounded result only)
|
||||
```
|
||||
|
||||
The adapters are the only place that knows how the current legacy tools are called. They parse the typed ACI request before invoking the boundary, serialize the legacy result as raw JSON, and bind the typed request and framework `tool_call_id` to the projector closure. `ToolBoundary` remains responsible for all lifecycle, ownership, budget, size and error rules.
|
||||
|
||||
## Request-aware projection decision
|
||||
|
||||
The generic `ToolResultProjector.project(String rawResponse)` remains unchanged for stage 3A compatibility. Stage 3B introduces specialized projector methods:
|
||||
|
||||
```java
|
||||
ProjectedToolResult project(RagToolRequest request, String toolCallId, String rawResponse);
|
||||
ProjectedToolResult project(QueryLogsRequest request, String toolCallId,
|
||||
LogQueryScope scope, String rawResponse);
|
||||
```
|
||||
|
||||
Adapters pass these methods through a lambda to the generic boundary. This preserves request scope without adding request fields to `ProjectedToolResult`, changing the generic boundary API, or relying on raw JSON conventions.
|
||||
|
||||
## RAG projection
|
||||
|
||||
- Parse only `evidenceBlocks` from the legacy response.
|
||||
- Accept an evidence block only when its content/excerpt is non-blank.
|
||||
- Use `source` as the stable `document_id`; fall back to title, then a deterministic ordinal only for malformed legacy records.
|
||||
- Deduplicate by document ID while retaining the first exact excerpt.
|
||||
- Project only `document_id`, `source`, `title`, `breadcrumb`, and bounded `excerpt`.
|
||||
- Ignore `contextPack`, `retrievalTrace`, `rerankTrace`, scores, hit reasons, domains, and messages.
|
||||
- Set `NO_EVIDENCE` when no usable excerpt remains.
|
||||
|
||||
## Query-log projection
|
||||
|
||||
- Parse only the legacy `logs` array and success/total fields needed for semantics.
|
||||
- Map `LogTopic` to the existing Mock topic names inside the adapter (`APPLICATION -> application-logs`, `DATABASE_SLOW_QUERY -> database-slow-query`, `SYSTEM_EVENTS -> system-events`).
|
||||
- Inject the legacy default region and an internal source limit; neither is part of the ACI request or result.
|
||||
- Build `LogQueryScope` from the logical request and `lookback_minutes`, using the adapter clock for `end_time`.
|
||||
- Always record `source_kind=MOCK` for this stage.
|
||||
- Exclude `instance` and `metrics`; redact credentials, host/pod identifiers, PIDs, IPs, SQL literals and stack-like suffixes from messages.
|
||||
- Aggregate patterns by sanitized level, service and normalized message, with count and first/last timestamps.
|
||||
- Sample timeline events deterministically when the source exceeds the event bound.
|
||||
- Set `NO_EVIDENCE` for a successful empty `logs` array; legacy error responses become projection errors.
|
||||
|
||||
## Bounds
|
||||
|
||||
`ToolProjectionLimits` defines maximum evidence/items, excerpt/message characters, patterns, events and total Agent projection UTF-8 bytes. Collection bounds set `truncated=true`. If the serialized projection still exceeds the total budget, the projector removes the last timeline/pattern/evidence items until it fits; if no valid bounded result can be produced, it throws and the boundary records `PROJECTION_ERROR`.
|
||||
|
||||
## Compatibility and ownership
|
||||
|
||||
- No changes to `LookupKnowledgeTool`, `QueryLogsTools`, `ToolInvocationRecorder`, JPA entities, ChatService, AiOpsService, Controller or public HTTP/SSE payloads.
|
||||
- Redis access remains in `RedisCanonicalInvocationStore`.
|
||||
- Legacy tools remain the source of raw data until a later Diagnosis Agent cutover.
|
||||
|
||||
## Risks and mitigations
|
||||
|
||||
- Legacy output shape drift: strict JSON parsing and focused malformed-response tests fail closed.
|
||||
- Redaction can remove useful details: only sensitive tokens/identifiers are replaced; service, level and sanitized message context remain.
|
||||
- Scope clock skew: adapter uses an injected `Clock` and tests use a fixed clock.
|
||||
- Projection size pressure: deterministic collection bounds and explicit `truncated` prevent silent overflow.
|
||||
@@ -0,0 +1,51 @@
|
||||
# Proposal: single-react-rag-log-projections
|
||||
|
||||
## Problem
|
||||
|
||||
阶段 3A 已提供统一 `ToolBoundary` 和 canonical invocation store,但 RAG 与 `query_logs` 仍只有旧工具输出。旧输出包含检索轨迹、上下文打包、基础设施字段、实例信息和未受约束的日志内容,不能直接作为单体 Diagnosis Agent 的 ACI evidence projection。若分别实现状态机或存储,会重新引入重复的生命周期、Run ownership 和错误处理逻辑。
|
||||
|
||||
## Proposed change
|
||||
|
||||
- 新增 RAG result projector,将旧 `LookupKnowledgeTool` JSON 转换为冻结的 `RagToolResult`。
|
||||
- 新增 query-log result projector,将现有 Mock `QueryLogsTools` JSON 转换为冻结的 `QueryLogsToolResult`。
|
||||
- projector 负责精确摘录、文档去重、证据/事件/模式数量上限、整体字符预算、时间线抽样和脱敏。
|
||||
- projector 由请求感知适配层调用,使 `topic`、`query` 和 `lookback_minutes` 保留在完整 `scope` 中;不扩展 Agent request,不让 raw response 直接进入 Agent。
|
||||
- 两个适配器均通过阶段 3A `ToolBoundary` 执行,沿用框架 `tool_call_id`、RunContext、canonical store、`PROJECTING/READY/ERROR` 与 `EVIDENCE_FOUND/NO_EVIDENCE/ERROR` 语义。
|
||||
- 保留旧工具、旧 recorder、Chat/AIOps、Controller 和公开协议不变;本阶段只新增投影/适配层与 focused tests。
|
||||
|
||||
## Scope
|
||||
|
||||
### In scope
|
||||
|
||||
- RAG projector and adapter.
|
||||
- Query-log projector and adapter for existing Mock source.
|
||||
- Request-aware projection entry point required to preserve log scope.
|
||||
- Contract, sanitization, deduplication, truncation, no-evidence and boundary integration tests.
|
||||
|
||||
### Out of scope
|
||||
|
||||
- Real CLS/MCP log integration.
|
||||
- MySQL projector (stage 3C).
|
||||
- Diagnosis Agent cutover, Guards, Chat use-case cutover, SSE changes, or legacy recorder cleanup.
|
||||
- Changes to the frozen ACI DTO field names.
|
||||
|
||||
## Constraints and risks
|
||||
|
||||
- Raw payload remains Harness-only canonical data and is never returned to the Agent.
|
||||
- `NO_EVIDENCE` is scoped to the recorded query and time window; it is not a health claim.
|
||||
- Log messages may contain credentials, hostnames, pod IDs, SQL literals, or stack traces; projection must redact these before Agent exposure.
|
||||
- A projector must not silently truncate raw data. It may bound Agent-facing collections and excerpts while setting `truncated=true`.
|
||||
- The existing generic `ToolResultProjector` API has no request argument. The adapter will carry the typed request alongside the projector invocation, keeping the generic boundary reusable and avoiding a raw JSON scope convention.
|
||||
|
||||
## Acceptance direction
|
||||
|
||||
- Serialized RAG output contains only the frozen fields and bounded exact excerpts.
|
||||
- Serialized log output contains complete logical scope, `source_kind=MOCK`, bounded patterns/events, distinct match and returned counts, and no infrastructure controls.
|
||||
- Both projectors produce `NO_EVIDENCE` for successful empty results and `ERROR` for malformed/unsafe input.
|
||||
- Boundary tests prove framework ID preservation, canonical lifecycle reuse, raw isolation, and safe errors.
|
||||
|
||||
## Context sources
|
||||
|
||||
- ISS-014 stage 3B requirements.
|
||||
- `aci-evidence-tool-contracts` and `canonical-tool-invocation-store` specifications.
|
||||
- Existing `LookupKnowledgeTool`, `QueryLogsTools`, and frozen contract tests.
|
||||
+73
@@ -0,0 +1,73 @@
|
||||
# rag-log-projections Specification
|
||||
|
||||
## Purpose
|
||||
|
||||
Define bounded RAG and Mock query-log projections that execute through the stage 3A ToolBoundary and expose only the frozen ACI contracts to the Agent.
|
||||
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: RAG projection SHALL expose bounded document evidence only
|
||||
|
||||
The RAG adapter SHALL accept the logical `query` request, execute the existing knowledge tool through ToolBoundary, and project only `RagToolResult` fields. Context packs, retrieval traces, rerank traces, scores, hit reasons, domains, messages and full document bodies SHALL NOT appear in the Agent result.
|
||||
|
||||
#### Scenario: RAG evidence is projected
|
||||
|
||||
- WHEN the legacy response contains usable evidence blocks
|
||||
- THEN the result contains the original query, framework `tool_call_id`, deduplicated evidence with exact bounded excerpts, a returned count, and an explicit truncation flag
|
||||
|
||||
#### Scenario: RAG result has no usable excerpt
|
||||
|
||||
- WHEN the legacy response has no usable evidence block
|
||||
- THEN the boundary returns `READY` with `evidence_status=NO_EVIDENCE`, the original query and an empty evidence list
|
||||
|
||||
#### Scenario: RAG duplicate documents
|
||||
|
||||
- WHEN multiple blocks identify the same source/document
|
||||
- THEN only the first document is returned and the result remains deterministic
|
||||
|
||||
### Requirement: Query-log projection SHALL preserve logical scope and Mock provenance
|
||||
|
||||
The query-log adapter SHALL accept only logical topic, query and optional lookback minutes, execute the existing Mock source through ToolBoundary, and project `source_kind=MOCK`, complete scope, match count, returned count, bounded patterns, bounded timeline events and truncation.
|
||||
|
||||
#### Scenario: Mock logs are projected
|
||||
|
||||
- WHEN the legacy Mock response contains log entries
|
||||
- THEN the result includes the logical topic/query/time window, sanitized pattern aggregates, sanitized timeline events, distinct match and returned counts, and `source_kind=MOCK`
|
||||
|
||||
#### Scenario: Empty Mock result
|
||||
|
||||
- WHEN the legacy Mock response is successful with an empty log array
|
||||
- THEN the boundary returns `READY` with `evidence_status=NO_EVIDENCE` and preserves the full query scope
|
||||
|
||||
#### Scenario: Legacy log error
|
||||
|
||||
- WHEN the legacy response indicates failure or is malformed
|
||||
- THEN the boundary returns `ERROR` with a bounded projection error and no Agent result
|
||||
|
||||
### Requirement: Projection SHALL redact and bound sensitive log content
|
||||
|
||||
The log projector SHALL exclude instance and metrics fields and redact credentials, token-like values, host/pod identifiers, PIDs, IP addresses, SQL literals and stack-like suffixes from Agent-facing messages. It SHALL enforce per-item, collection and total UTF-8 bounds and set `truncated=true` when any bound removes data.
|
||||
|
||||
#### Scenario: Sensitive fields are present
|
||||
|
||||
- WHEN a log entry includes instance, metrics, a secret assignment, a pod identifier or a SQL literal
|
||||
- THEN none of those raw values are present in the projected Agent result
|
||||
|
||||
#### Scenario: Projection exceeds collection budget
|
||||
|
||||
- WHEN evidence, patterns, events or excerpts exceed configured bounds
|
||||
- THEN the result is valid JSON, contains only bounded collections, and sets `truncated=true`
|
||||
|
||||
### Requirement: Adapters SHALL reuse canonical boundary ownership
|
||||
|
||||
Both adapters SHALL pass the framework `tool_call_id` and RunContext to the existing ToolBoundary and SHALL NOT create a second ID, write a parallel store, return raw responses, or modify legacy audit paths.
|
||||
|
||||
#### Scenario: Framework ID and Run are valid
|
||||
|
||||
- WHEN an adapter executes a valid request
|
||||
- THEN the canonical record and bounded result use the exact framework ID and the current Run ID
|
||||
|
||||
#### Scenario: Boundary rejects the request
|
||||
|
||||
- WHEN Run ownership, authorization, read-only, duplicate, budget or size preflight fails
|
||||
- THEN the adapter returns the boundary's safe error without invoking the legacy tool or projector
|
||||
@@ -0,0 +1,24 @@
|
||||
# Tasks: single-react-rag-log-projections
|
||||
|
||||
## 1. Projection limits and request-aware interfaces
|
||||
|
||||
- [x] 1.1 Add `ToolProjectionLimits` with explicit item, excerpt/message, collection and total UTF-8 limits.
|
||||
- [x] 1.2 Add request-aware projector methods and adapter-facing typed execution helpers without changing the generic ToolBoundary API.
|
||||
|
||||
## 2. RAG projection and adapter
|
||||
|
||||
- [x] 2.1 Implement JSON parsing, usable excerpt filtering, source/document ID fallback, deduplication, exact excerpt bounding and `RagToolResult` serialization.
|
||||
- [x] 2.2 Implement the RAG adapter that maps typed request JSON to the legacy knowledge tool and invokes ToolBoundary with the framework ID.
|
||||
- [x] 2.3 Add tests for field exclusion, no evidence, duplicate documents, excerpt/total bounds, malformed raw response and canonical ID preservation.
|
||||
|
||||
## 3. Query-log projection and adapter
|
||||
|
||||
- [x] 3.1 Implement logical topic mapping, injected legacy Mock invocation, lookback scope construction and `source_kind=MOCK`.
|
||||
- [x] 3.2 Implement redaction, pattern aggregation, deterministic timeline sampling, counts and bounded serialization.
|
||||
- [x] 3.3 Add tests for scope, provenance, pattern/timeline bounds, redaction, empty result, malformed/error response and no topic discovery call.
|
||||
|
||||
## 4. Boundary integration and verification
|
||||
|
||||
- [x] 4.1 Route both adapters through the existing ToolBoundary and verify raw data remains canonical-only.
|
||||
- [x] 4.2 Run focused projector/adapter/boundary tests and the prior stage regression suite.
|
||||
- [x] 4.3 Validate OpenSpec, update devflow acceptance/evidence, archive the change and commit before stage 3C.
|
||||
@@ -0,0 +1 @@
|
||||
Archive-ready after implementation and focused verification on 2026-07-21.
|
||||
@@ -0,0 +1 @@
|
||||
Committed after strict validation on 2026-07-21.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-21
|
||||
@@ -0,0 +1,102 @@
|
||||
## Context
|
||||
|
||||
当前 `ToolInvocationRecorder` 将 JPA `ToolInvocation` 作为旧链路的 durable audit,保存 `output_preview`(500 字符)和 retrieval details,并通过 `SessionContextHolder` 补 session/run。它不能作为 EvidenceGuard 的 canonical source:raw response 已被截断、生命周期没有 PROJECTING/READY/ERROR、没有框架 Tool Call ID,也没有单 Run 容量/TTL 门禁。
|
||||
|
||||
阶段 2 已提供 `RunContext`、budget、capacity 和 `ToolCallKeyFactory`。本阶段需把所有 Tool 共用的执行边界与 canonical invocation store 建起来,供阶段 3B/3C 的 projector 直接复用。
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- 在 Tool 调用前统一校验 Run 所有权、框架 ID、JSON object、授权、只读和 Tool/Run budget。
|
||||
- 保存同一 canonical record 的完整 request/raw_response/agent_result 及状态、证据语义、时间和错误。
|
||||
- 集中执行 PROJECTING -> READY/ERROR,防止未经投影的 raw 进入 Agent。
|
||||
- 固定 TTL 不续期、单记录/Agent result/Run bytes 上限和 RESULT_TOO_LARGE 语义。
|
||||
- 提供 Redis 实现和 store-independent Fake boundary tests。
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- 不实现 RAG/log/MySQL specific projector 或 adapter。
|
||||
- 不修改旧 ToolInvocationRecorder/JPA/数据库、Chat/AIOps、Controller/SSE。
|
||||
- 不让 Agent 访问 Redis/client/key/raw record。
|
||||
- 不实现并发 Lua/CAS 更新、永久 audit、脱敏或跨 Run 查询。
|
||||
|
||||
## Decisions
|
||||
|
||||
### 1. Internal ToolBoundary envelope and result
|
||||
|
||||
`ToolCallRequestEnvelope` 是 Harness 内部输入,包含 `runId`、框架 `toolCallId`、toolName、JSON request、`authorized` 和 `readOnly`。Agent-facing DTO 不暴露这些门禁字段;后续 Alibaba ToolInterceptor 负责组装 envelope。
|
||||
|
||||
`ToolBoundaryResult` 只返回 invocation status、evidence status、原始框架 ID、bounded `agentResult` 或安全 errorCode;raw 只进入 store,不返回 Agent。
|
||||
|
||||
### 2. Canonical record and lifecycle
|
||||
|
||||
`CanonicalToolInvocation` 是 JSON serializable record,字段包含 toolCallId、runId、toolName、request、rawResponse、agentResult、InvocationStatus、EvidenceStatus、errorCode、startedAt、completedAt。`begin` 只接受 PROJECTING;`markReady` 只接受 PROJECTING + 非空 bounded projection + FOUND/NO_EVIDENCE;`markError` 将 evidence status 固定为 ERROR。
|
||||
|
||||
状态转换和 duplicate 检查在 `CanonicalInvocationStore` 内集中执行。ToolBoundary 不直接写 Redis。
|
||||
|
||||
### 3. Redis value and TTL
|
||||
|
||||
Redis 实现复用现有 `RedisTemplate<String,Object>`,将 record 序列化为 JSON String。创建使用 `setIfAbsent(key,json,ttl)`,保证同一 `runId+toolCallId` 不覆盖;读取不调用 expire。更新先读取剩余毫秒 TTL,再用不大于该值的 TTL 写回,避免恢复初始 TTL。过期/缺失读取返回 empty。
|
||||
|
||||
替代方案是 Redis Hash;单 JSON value 能保证 request/raw/agent 原子同记录,并让 Fake/序列化 schema 与 canonical record 一致,故采用。并发 projector 更新窗口是已接受风险,后续需要时再升级 Lua/CAS。
|
||||
|
||||
### 4. Size and status rules
|
||||
|
||||
`CanonicalInvocationLimits` 由 caller 提供 TTL、maxRecordBytes 和 maxAgentResultBytes。request/raw/agent 使用 UTF-8 bytes 计数;raw 超过 record 或 agent projection 超过独立上限,均不截断,记录 ERROR/RESULT_TOO_LARGE。Run bytes 通过阶段 2 Core 再做单 Run 累计门禁。
|
||||
|
||||
PROJECTING 时 evidence status 仅作为内部未知/ERROR 占位;READY 只接受 `EVIDENCE_FOUND` 或 `NO_EVIDENCE`;ERROR 永远不可引用。NO_EVIDENCE 不触发重试或成功解释。
|
||||
|
||||
### 5. Preflight and projector boundary
|
||||
|
||||
ToolBoundary 顺序固定:
|
||||
|
||||
```text
|
||||
RunContext active/deadline
|
||||
-> runId + toolCallId + authorization + read-only + JSON object
|
||||
-> Core.beforeToolCall + request/run capacity
|
||||
-> store.begin(PROJECTING)
|
||||
-> ToolExecutor(raw)
|
||||
-> raw size/capacity
|
||||
-> ToolResultProjector(agentResult,evidenceStatus)
|
||||
-> agent size/capacity
|
||||
-> store.markReady or markError
|
||||
-> bounded ToolBoundaryResult
|
||||
```
|
||||
|
||||
执行或投影异常都会写 ERROR;raw 已在可信边界且未超限时保留在 canonical record,但不返回 Agent。preflight/duplicate/cross-run 错误在 begin 前返回安全 ERROR。
|
||||
|
||||
### 6. Old audit separation
|
||||
|
||||
旧 recorder/JPA 继续接收旧 Tool 调用,阶段 3A 不改其字段和 ThreadLocal fallback。新 canonical store 没有旧消费者;阶段 3B/3C 接入时必须明确写新 boundary,并在需要 durable audit 时另行脱敏摘要。
|
||||
|
||||
## Module Flow
|
||||
|
||||
```text
|
||||
Alibaba ToolInterceptor (future)
|
||||
-> ToolCallRequestEnvelope
|
||||
-> ToolBoundary
|
||||
-> DiagnosisHarnessCore + ToolCallKeyFactory
|
||||
-> CanonicalInvocationStore (Redis JSON / Fake)
|
||||
-> ToolExecutor
|
||||
-> ToolResultProjector (future RAG/log/MySQL)
|
||||
-> bounded ToolBoundaryResult
|
||||
```
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Redis update read-TTL-write 存在并发窗口] -> 当前每个 invocation 只允许 boundary 顺序更新;后续并发需求升级 Lua/CAS。
|
||||
- [canonical raw 可能敏感] -> 仅 Harness store 访问,TTL/ACL/容量受限;脱敏在 projector/durable audit 阶段处理。
|
||||
- [旧 recorder 与新 store 短期并存] -> 包和接口隔离,spec 明确旧 preview 不能作为 canonical evidence。
|
||||
- [preflight 失败可能没有 canonical record] -> 返回安全 ERROR 且不执行 Tool;阶段 3A 的可引用记录只针对已通过 begin 的调用。
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. 本阶段新增 boundary/store/Redis adapter 和 fake tests,旧运行链路不变。
|
||||
2. 阶段 3B/3C 将各 Tool adapter/projector 包装到本 boundary。
|
||||
3. 阶段 4 Diagnosis Agent 只接收 boundary 的 bounded result。
|
||||
4. 阶段 6A/7 再决定 durable audit 如何从 canonical 摘要回填,并清理旧 recorder/ThreadLocal。
|
||||
|
||||
## Open Questions
|
||||
|
||||
无。真实 Redis 的 ACL、网络和 TTL 由最终运行/E2E 阶段验证;并发更新 Lua 化留作后续需求。
|
||||
@@ -0,0 +1,45 @@
|
||||
## Why
|
||||
|
||||
阶段 1 已冻结 Agent-facing Tool Contract,阶段 2 已提供 RunContext、预算、取消和 Tool Call Key 基础,但当前 `ToolInvocationRecorder` 仍把截断 preview 写入 JPA、依赖 ThreadLocal,并没有同一条记录中的 `request/raw_response/agent_result`、生命周期或当前 Run 所有权。RAG、日志和 MySQL 投影若各自保存调用,会重新复制状态机并让 EvidenceGuard 无法证明引用来自当前 Run。
|
||||
|
||||
## What Changes
|
||||
|
||||
- 新增统一 `ToolBoundary`,在每个 Tool 调用前执行 JSON Schema/只读/Run/预算/Tool Call ID 门禁,执行后统一处理 raw、投影、状态和错误。
|
||||
- 新增 `CanonicalInvocationStore` 抽象与 Redis 实现,按阶段 2 Key Factory 保存一条完整 JSON 调用记录:`request`、`raw_response`、`agent_result`、`status`、`evidence_status`、时间和错误信息。
|
||||
- `PROJECTING -> READY/ERROR` 生命周期和独立 EvidenceStatus 在 store 中集中执行;READY 才允许 `EVIDENCE_FOUND/NO_EVIDENCE`,ERROR 不可引用。
|
||||
- 创建时设置 TTL,读取不刷新;更新只使用当前剩余 TTL,不延长生命周期;单记录、Agent projection 和单 Run 容量超限显式返回 `RESULT_TOO_LARGE`,不静默截断 raw。
|
||||
- 拒绝缺失/非法/重复 Tool Call ID、跨 Run 引用、不可解析 JSON、非只读请求和已超预算调用;不生成第二套 ID。
|
||||
- 使用 Fake Tool/Projector/In-memory Store 覆盖成功、no-evidence、projection error、execution error、duplicate/cross-run、TTL、容量和 raw oversize。
|
||||
- 本阶段不实现 RAG/log/MySQL specific projector,不修改旧 `ToolInvocationRecorder`、JPA entity、Controller、ChatService 或公开协议。
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `canonical-tool-invocation-store`: 提供统一 ToolBoundary、canonical invocation 生命周期、Run 所有权、容量/TTL 和可引用状态边界,供后续 RAG/log/MySQL 投影复用。
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- None. 旧 JPA audit 记录继续服务旧链路;新 store 先作为零消费者 Harness foundation。
|
||||
|
||||
## Context Constraints
|
||||
|
||||
- canonical store 只能由 Harness/ToolBoundary 访问,Agent 不获得 Redis client/key/raw record。
|
||||
- Redis key 固定由阶段 2 `ToolCallKeyFactory` 生成:`prefix:runId:toolCallId`。
|
||||
- 同一调用的完整 request/raw/agent projection 必须在同一记录;raw 不能只保存 preview,也不能未经 projector 返回 Agent。
|
||||
- 创建 TTL 默认配置由 caller 提供且必须大于 0;读取与更新不得续期。
|
||||
- `PROJECTING` 时 evidence_status 只能是内部暂态 ERROR/unknown;只有 READY 才能成为 `EVIDENCE_FOUND` 或 `NO_EVIDENCE`。
|
||||
- `NO_EVIDENCE` 仅作为结果语义,不可被 boundary 自动升级为成功事实或重试。
|
||||
|
||||
## Interface Impact
|
||||
|
||||
- 等级:L2(内部 Harness/Tool boundary)。新增接口会被阶段 3B/3C 直接消费,旧调用方不变。
|
||||
- 不改变 JPA `tool_invocation`、数据库 Schema、旧审计 preview 或公开 HTTP/SSE。
|
||||
- Redis 是新增运行时依赖使用既有 `RedisTemplate<String,Object>` bean;真实连接验证留给阶段 7,focused tests 使用 fake/mocks。
|
||||
|
||||
## Risks
|
||||
|
||||
- Redis JSON value 更新需要读取剩余 TTL 后再写回,存在并发更新窗口;当前单 Tool Call 只有 boundary 状态机写入,后续若并发 projector 必须升级 Lua/CAS。
|
||||
- canonical raw 可包含敏感内容;本 Issue 保留阶段 0 已确认的 Harness-only ACL/TTL 约束,持久化脱敏和 durable audit 留给后续阶段。
|
||||
- ToolBoundary 同时负责预算、store 状态和 projector 错误,若异常分类不清会产生错误状态;每个边界分支都有 Fake tests。
|
||||
- 当前旧 recorder 继续运行,新旧两条 audit 链短期并存;proposal 明确禁止把旧 JPA 记录当 canonical evidence。
|
||||
+74
@@ -0,0 +1,74 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: ToolBoundary SHALL enforce explicit preflight before execution
|
||||
The Harness SHALL reject a Tool call before invoking the executor when the envelope has a blank/unsafe framework `tool_call_id`, a run ID different from RunContext, invalid JSON object input, unauthorized access, non-read-only access, an inactive/deadline-expired Run, duplicate canonical key, or exhausted Tool/Run budget.
|
||||
|
||||
#### Scenario: Cross-Run Tool Call is rejected
|
||||
- **WHEN** an envelope run ID differs from the explicit RunContext run ID
|
||||
- **THEN** ToolBoundary returns a safe `ERROR`, does not create a canonical record, and does not invoke the Tool
|
||||
|
||||
#### Scenario: Unauthorized or writable Tool is rejected
|
||||
- **WHEN** `authorized=false` or `readOnly=false`
|
||||
- **THEN** ToolBoundary returns `ERROR` before execution and does not expose the request to an Agent
|
||||
|
||||
### Requirement: Canonical invocation SHALL keep complete data in one record
|
||||
The store SHALL save the complete request JSON, raw Tool response, bounded Agent result, framework `tool_call_id`, run ID, tool name, lifecycle status, evidence status, timestamps, and error code in one canonical record. The raw response SHALL NOT be replaced by a preview, and SHALL NOT be returned to the Agent.
|
||||
|
||||
#### Scenario: Tool begins execution
|
||||
- **WHEN** preflight succeeds and the Tool is about to execute
|
||||
- **THEN** one record is created with `status=PROJECTING`, the complete request, the exact framework ID, and no Agent result yet
|
||||
|
||||
#### Scenario: Projector succeeds
|
||||
- **WHEN** the executor returns raw data and the projector returns a bounded result
|
||||
- **THEN** the same record contains raw data and Agent result with `status=READY`, and ToolBoundary returns only the bounded Agent result
|
||||
|
||||
### Requirement: Invocation lifecycle and evidence semantics SHALL be enforced centrally
|
||||
The store SHALL allow only `PROJECTING -> READY` or `PROJECTING -> ERROR`. READY SHALL require `EVIDENCE_FOUND` or `NO_EVIDENCE`; ERROR SHALL use `evidence_status=ERROR` and SHALL NOT be referencable. `NO_EVIDENCE` SHALL remain a scoped negative observation and SHALL NOT be upgraded to success or retried by the boundary.
|
||||
|
||||
#### Scenario: No evidence projection completes
|
||||
- **WHEN** a projector returns a valid bounded result with `NO_EVIDENCE`
|
||||
- **THEN** the record becomes `READY`, preserves the result scope, and remains eligible only for a negative observation
|
||||
|
||||
#### Scenario: Projection fails
|
||||
- **WHEN** the projector throws or returns an invalid evidence status
|
||||
- **THEN** the same record becomes `ERROR`, stores a safe error code, and no Agent result is returned
|
||||
|
||||
### Requirement: Tool Call ID and Run ownership SHALL be preserved
|
||||
The boundary SHALL use the exact framework `tool_call_id` with the RunContext run ID to create the canonical key. It SHALL reject missing, unsafe, duplicate, and cross-Run references and SHALL never generate or replace a fallback ID.
|
||||
|
||||
#### Scenario: Duplicate Tool Call ID is submitted
|
||||
- **WHEN** a second invocation uses the same valid run ID and framework Tool Call ID
|
||||
- **THEN** the second Tool is not executed and returns `ERROR` without overwriting the first record
|
||||
|
||||
#### Scenario: Framework ID is preserved
|
||||
- **WHEN** a valid envelope passes preflight
|
||||
- **THEN** the key and canonical record contain the exact supplied `tool_call_id`
|
||||
|
||||
### Requirement: TTL and size limits SHALL fail closed
|
||||
The Redis implementation SHALL set TTL only when a canonical record is created. Reads SHALL NOT refresh TTL; updates SHALL preserve only the current remaining TTL. Request/raw/record and Agent-result byte limits SHALL be enforced in UTF-8. Raw overflow SHALL produce `ERROR/RESULT_TOO_LARGE` without silent truncation; Agent projection overflow SHALL produce the same error.
|
||||
|
||||
#### Scenario: Read does not renew TTL
|
||||
- **WHEN** a canonical record is read before expiration
|
||||
- **THEN** its expiry remains at or before the original expiry and no expire/refresh operation is issued
|
||||
|
||||
#### Scenario: Raw response is too large
|
||||
- **WHEN** the executor returns raw data beyond the record limit
|
||||
- **THEN** the record becomes `ERROR` with `RESULT_TOO_LARGE`, the raw payload is not silently truncated, and the projector is not invoked
|
||||
|
||||
#### Scenario: Agent result is too large
|
||||
- **WHEN** a projector returns a result beyond the Agent projection limit
|
||||
- **THEN** the record becomes `ERROR/RESULT_TOO_LARGE` and the oversized result is not returned to the Agent
|
||||
|
||||
### Requirement: Tool and store failures SHALL return safe errors
|
||||
Execution, serialization, Redis, duplicate, preflight, and projection failures SHALL return a bounded error code without raw payload, credentials, internal stack trace, or Redis key to the Agent. If raw data was safely stored before a projection failure, it remains Harness-only canonical data.
|
||||
|
||||
#### Scenario: Tool execution throws
|
||||
- **WHEN** the executor raises an exception after PROJECTING begins
|
||||
- **THEN** the record becomes `ERROR` with a stable execution error code and ToolBoundary returns no raw response
|
||||
|
||||
### Requirement: Stage 3A SHALL remain reusable and independent from legacy audit
|
||||
The boundary and canonical store SHALL be usable by later RAG/log/MySQL projectors without modifying legacy `ToolInvocationRecorder`, JPA `ToolInvocation`, Chat/AIOps, Controller, or public protocol. Redis access SHALL remain inside the store implementation and no Agent-facing API SHALL expose the Redis client.
|
||||
|
||||
#### Scenario: Fake projector tests pass
|
||||
- **WHEN** Fake Tool/Projector tests cover success, no evidence, error, duplicate, TTL, and capacity cases
|
||||
- **THEN** later Tool-specific changes can depend on the same boundary without copying lifecycle or Redis logic, while legacy call sites remain unchanged
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user