refactor(harness): freeze single-agent contracts

This commit is contained in:
zhuyongxin
2026-07-21 17:33:25 +08:00
parent 30d3296043
commit 58c39107c5
47 changed files with 2607 additions and 24 deletions
@@ -148,7 +148,7 @@ openspec/changes/phase-1-infrastructure/
## 敏感信息(已编辑) ## 敏感信息(已编辑)
- MySQL 密码:已配置在 application.yml(`!Fucker123..`) - MySQL 密码:已从仓库移除,使用环境变量注入
- Redis:无密码 - Redis:无密码
--- ---
+1 -1
View File
@@ -97,7 +97,7 @@ spring:
redis: redis:
host: 119.29.78.52 host: 119.29.78.52
port: 6379 port: 6379
password: '!Fucker123..' password: ${SUPERBIZ_REDIS_PASSWORD}
database: 0 database: 0
timeout: 3000 timeout: 3000
``` ```
+2 -2
View File
@@ -140,7 +140,7 @@ Error Code: 1049
datasource: datasource:
url: jdbc:mysql://119.29.78.52:33306/superbiz_agent?... url: jdbc:mysql://119.29.78.52:33306/superbiz_agent?...
username: root username: root
password: '!Fucker123..' password: ${SUPERBIZ_MYSQL_PASSWORD}
``` ```
**Redis 配置**: **Redis 配置**:
@@ -149,7 +149,7 @@ data:
redis: redis:
host: 119.29.78.52 host: 119.29.78.52
port: 6379 port: 6379
password: '!Fucker123..' password: ${SUPERBIZ_REDIS_PASSWORD}
``` ```
**Flyway 配置**: **Flyway 配置**:
+7
View File
@@ -44,6 +44,13 @@ build/
app.log app.log
logs/ logs/
### Local Secrets ###
.env
.env.*
!.env.example
application-local.yml
application-*.local.yml
### Upload Files ### ### Upload Files ###
uploads/ uploads/
+20
View File
@@ -155,6 +155,26 @@
- 使用场景:Executor 按 skill workflow 调用 evidence tools 收集事实,`tool_invocation` 记录这些事实证据。 - 使用场景:Executor 按 skill workflow 调用 evidence tools 收集事实,`tool_invocation` 记录这些事实证据。
- 边界:最终诊断结论必须被 evidence tools 支撑,不能仅由 skill 正文支撑。 - 边界:最终诊断结论必须被 evidence tools 支撑,不能仅由 skill 正文支撑。
### Diagnosis Harness
- 定义:围绕 Diagnosis Agent 提供确定性运行控制的边界,负责 Run、预算、取消、重试装配、Tool 调用记录、证据验真和最终释放,不承担业务诊断推理。
- 边界:Harness 不是工作流引擎,不实现 Planner/Executor/Composer 节点或自行编写 ReAct 循环。
### EvidenceGuard
- 定义:Harness 内部的确定性证据验真能力,校验 Draft 引用、当前 Run 所有权、Tool 调用状态和有界 Agent 投影。
- 边界:EvidenceGuard 不调用 LLM,也不判断证据是否足以推出业务结论。
### SemanticGuard
- 定义:使用隔离上下文对完整诊断 Draft 与已验真证据做报告级语义审查的单轮 Agent。
- 边界:无工具、无记忆、无 ReAct 循环,不访问 Redis,不生成或改写用户报告。
### Invocation Status
- 定义:Tool 调用及结果投影的生命周期状态,固定为 `PROJECTING`、`READY`、`ERROR`。
- 边界:它只说明调用记录是否完成,不说明结果是否包含证据。
### Evidence Status
- 定义:证据 Tool 的结果语义,固定为 `EVIDENCE_FOUND`、`NO_EVIDENCE`、`ERROR`。
- 边界:`NO_EVIDENCE` 只表示当前查询范围内没有匹配结果,不能解释为问题不存在、根因被排除或系统健康。
### Verifier Skill Isolation ### Verifier Skill Isolation
- 定义:Chat Verifier 与 skill 系统隔离,只校验 Executor 答案和 `tool_trace_summary`。 - 定义:Chat Verifier 与 skill 系统隔离,只校验 Executor 答案和 `tool_trace_summary`。
- 使用场景:防止 Verifier 把 playbook 指令当作事实证据;Verifier 只判断已有证据是否支持结论。 - 使用场景:防止 Verifier 把 playbook 指令当作事实证据;Verifier 只判断已有证据是否支持结论。
+1
View File
@@ -4,6 +4,7 @@
| 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 | | 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|---|---|---|---|---|---|---| |---|---|---|---|---|---|---|
| 2026-07-21 | single-react-design-freeze | 冻结单体 Diagnosis Agent、Harness、Guard、工具证据与阶段门禁契约。 | Chat/Harness/Agent contract | ISS-014, single ReactAgent, Harness, EvidenceGuard, SemanticGuard, tool_call_id, evidence_status | openspec/changes/archive/2026-07-21-single-react-design-freeze | archived |
| 2026-07-10 | session-run-trace-isolation | 拆分会话态和运行态,引入 runId 隔离 Trace、Feedback、AIOps 和 demo 链路。 | Trace/session/run isolation | chat_session, diagnosis_run, runId, trace exact run, feedback fallback, AIOps SSE metadata, baseline drift | openspec/changes/archive/2026-07-10-session-run-trace-isolation | archived | | 2026-07-10 | session-run-trace-isolation | 拆分会话态和运行态,引入 runId 隔离 Trace、Feedback、AIOps 和 demo 链路。 | Trace/session/run isolation | chat_session, diagnosis_run, runId, trace exact run, feedback fallback, AIOps SSE metadata, baseline drift | openspec/changes/archive/2026-07-10-session-run-trace-isolation | archived |
| 2026-07-09 | interview-demo-quality-audit | 增加面试演示前置质量审计,覆盖 prompt、Gatekeeper 和评测基线。 | Agent eval/demo/Prompt audit | interview demo preflight, prompt_audit, gatekeeper rules, diagnosis baseline, 12 fixtures | openspec/changes/archive/2026-07-09-interview-demo-quality-audit | archived | | 2026-07-09 | interview-demo-quality-audit | 增加面试演示前置质量审计,覆盖 prompt、Gatekeeper 和评测基线。 | Agent eval/demo/Prompt audit | interview demo preflight, prompt_audit, gatekeeper rules, diagnosis baseline, 12 fixtures | openspec/changes/archive/2026-07-09-interview-demo-quality-audit | archived |
| 2026-07-08 | executor-composer-final-answer | 引入 Composer 生成最终回答,只使用 Verifier 允许的结论材料。 | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived | | 2026-07-08 | executor-composer-final-answer | 引入 Composer 生成最终回答,只使用 Verifier 允许的结论材料。 | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
@@ -0,0 +1,36 @@
# Acceptance: single-react-design-freeze
## 实现结果
- 新增类型化 Harness contracts 和序列化契约测试。
- ISS-014 已调整为 11 个串行 changes,并补齐 Tool ID、双状态、previous turn 和安全切换决策。
- Python 查询脚本和 Spring 主配置改为环境变量 Secret。
- 集成测试和历史 handoff 中的旧 Secret 副本已删除或脱敏。
- devflow glossary 新增 Harness、EvidenceGuard、SemanticGuard、Invocation Status 和 Evidence Status。
## 静态验证
- OpenSpec strict validation:通过。
- Secret 扫描:通过,已知真实 Secret 模式零匹配。
- Diff whitespace 检查:通过。
## 脚本验证
- 4 个 contract tests:通过。
- ToolInvocationRecorder、ExecutorGatekeeperService、ChatController focused baseline:通过。
- 查询脚本缺失密码环境变量时 fail fast:通过。
## 浏览器/人工验证
- 不适用;阶段 0 未切换任何公开 UI/API。
## 未验证与后续门禁
- Provider 侧旧凭据轮换待凭据所有者完成。
- Live E2E 留到阶段 7。
- 阶段 1 只有在本 change Archive 和 Git commit 后才能开始。
## 状态
- Stage acceptance: accepted
- OpenSpec archive: authorized by standing user instruction
@@ -0,0 +1,31 @@
# Brief: single-react-design-freeze
## 背景
ISS-014 将 Chat 诊断从多 Agent、Hook、ThreadLocal 和外层重试编排收敛为单体 Diagnosis ReAct Agent、确定性 Harness 和隔离 SemanticGuard。阶段 0 先冻结后续实施共同依赖的契约和安全边界。
## 目标
- 提供可复用的 Diagnosis Draft、Knowledge Answer、Fallback、published result 和 previous turn 类型。
- 固定 framework Tool Call ID、调用生命周期和证据结果双状态。
- 固定取消、重试、no-evidence、阶段门禁和安全凭据策略。
- 保持当前公开 Chat 运行行为不变。
## 范围
- `com.superbiz.agent.harness.contract` 类型层和 focused tests。
- ISS-014 的 11 个串行 OpenSpec changes 台账。
- 主配置、查询脚本和集成测试中的明文 Secret 清理。
- OpenSpec、devflow glossary 和阶段验收基线。
## 非目标
- 不接入新 Harness、Agent 或 Guards。
- 不切换 `/api/chat`,不删除旧运行链。
- 不执行 live E2E。
## 元数据
- Scale: complex
- Parent issue: `ISS-014`
- OpenSpec: `openspec/changes/single-react-design-freeze`
@@ -0,0 +1,20 @@
# Decisions: single-react-design-freeze
## 已确认决策
- ISS-014 保留为总 Issue,实施拆成 11 个独立、串行 sm-flow changes。
- 每个 change 必须完成 Apply、阶段验收、OpenSpec Archive 和 Git commit 后才能进入下一项。
- Apply、Archive 和阶段 Git commit 已获得用户对整个目标的持续授权;仅方向性决策需要暂停。
- `tool_call_id` 使用框架协议 ID,Harness 不生成第二套 ID。
- `status=PROJECTING/READY/ERROR` 表示 invocation lifecycle;`evidence_status=EVIDENCE_FOUND/NO_EVIDENCE/ERROR` 表示结果语义。
- `NO_EVIDENCE` 只能支持限定范围的 `NEGATIVE_OBSERVATION`。
- KNOWLEDGE_QUERY 使用 answer items 绑定 `tool_call_id + document_id`,首版不进入 SemanticGuard。
- `diagnosis_run` 是 previous turn 的持久化真理源;只有最近的 `DIAGNOSIS + SUCCESS` 安全发布结果可用。
- 取消采用分层、可观测语义;同步模型调用不承诺无法证明的立即硬取消。
- 底层重试压为一次 attempt,Router/SemanticGuard 的允许重试只由 Harness 执行。
- 阶段 4 和 6A 不接管公开入口,阶段 6B 才执行原子 SSE 切换。
## 权衡
- contract types 提前落地会增加少量文件,但后续 output type、validator、持久化和 SSE 可以复用,避免多阶段字符串协议漂移。
- Provider 侧凭据轮换是外部前置,阶段 0 只证明工作树已清理,不伪造外部完成状态。
@@ -0,0 +1,25 @@
# Evidence: single-react-design-freeze
## 代码与框架证据
- `ChatController` 当前直接管理会话、模型、工具、执行和伪流式 SSE,证明公开切换必须延后到阶段 6B。
- `ChatService` 当前构建 Planner/Executor/Verifier/Composer 并依赖 ThreadLocal,证明阶段 0 只能创建无运行依赖的 contract types。
- Spring AI Alibaba `Builder` 提供 ToolInterceptor、output type、tool timeout 和调用限制扩展点;`ReactAgent` 提供 interrupt。
- Spring AI 1.1.7 `SpringAiRetryProperties` 构造器默认 `maxAttempts=10`,与 Harness 单一重试所有权冲突,阶段 2 必须压为一次底层 attempt。
- `DiagnosisRun` 当前只有通用 status 和文本 answer,不足以确定性恢复安全 previous turn,因此确认后续保存 `intent/release_outcome/published_result`。
- 历史 ISS-009 和 verifier evidence 归档证明 `NO_EVIDENCE` 只能表达限定范围内无匹配结果。
## 验证证据
- `mvn -q '-Dtest=HarnessContractTest' test`:通过。
- `mvn -q '-Dtest=ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,ChatControllerTest' test`:通过。
- `mvn -q '-Dtest=HarnessContractTest,ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,ChatControllerTest' test`:通过。
- `openspec validate single-react-design-freeze --strict`:通过。
- `git diff --check`:通过,仅有 Windows line-ending warning。
- 已知旧 Secret 模式扫描:零匹配。
- 缺少 `SUPERBIZ_MYSQL_PASSWORD` 时运行查询脚本:在连接前 fail fast,符合预期。
## 未验证
- 未运行模型、Redis、MySQL 或 Milvus live E2E;按阶段门禁留到阶段 7。
- Provider 侧旧凭据是否已经轮换无法由仓库证明,必须由凭据所有者完成,并在阶段 3C/7 前核验。
+2 -2
View File
@@ -31,7 +31,7 @@
**配置信息**(已完成): **配置信息**(已完成):
- **MySQL**: 119.29.78.52:33306/superbiz_agent - **MySQL**: 119.29.78.52:33306/superbiz_agent
- 用户: root - 用户: root
- 密码: !Fucker123.. - 密码: 已从仓库移除,使用环境变量注入
- driver: com.mysql.cj.jdbc.Driver - driver: com.mysql.cj.jdbc.Driver
- URL参数: useUnicode=true&serverTimezone=Asia/Shanghai&allowPublicKeyRetrieval=true - URL参数: useUnicode=true&serverTimezone=Asia/Shanghai&allowPublicKeyRetrieval=true
- **Redis**: 119.29.78.52:6379 - **Redis**: 119.29.78.52:6379
@@ -165,7 +165,7 @@ openspec/changes/phase-1-infrastructure/
## 敏感信息(已编辑) ## 敏感信息(已编辑)
- MySQL 密码:已配置在 application.yml(`!Fucker123..`) - MySQL 密码:已从仓库移除,使用环境变量注入
- Redis:无密码 - Redis:无密码
--- ---
+56
View File
@@ -0,0 +1,56 @@
结合你在前几轮对话中梳理出的“过度设计”痛点,既然你已经决定回归单 Agent(ReAct)+ 强 Harness 架构,接下来你需要做一次彻底的“架构物理重构”。
以下是为你量身定制的5步落地行动指南,按优先级从高到低执行:
第一步:链路合并,砍掉“伪 Graph”节点
目标:把被拆散的推理逻辑还给单一的 ReAct 循环。
删除节点:直接在 Graph/状态机中抹掉 Planner、Executor、Composer 以及串联它们的边。
合并为单一 Agent:创建一个 DiagnosisAgent。在它的 System Prompt 中明确:“你需要自行规划排查路径,调用工具获取证据,并在证据充分后输出最终的结构化诊断报告。”
保留机制:该 Agent 内部运行一个带最大步数限制的 While 循环(如 max_iterations=10),防止死循环。
第二步:实施“上下文卸载”与“工具清洗”(解决7次调用膨胀问题)
目标:确保单 Agent 在多轮工具调用后,上下文依然干净,从源头切断幻觉。
工具输出清洗(RTK 机制):改造你的所有工具(如 queryLogs, queryDB)。在工具的 Callback 层加代码,将原始的大段返回结果(如 1000 行日志)强制过滤、聚合,只把最核心的 5-10 行 ERROR 或统计指标返回给 Agent。
上下文卸载:如果某些原始数据必须保留,在工具返回时,将全量数据写入本地文件(如 refs/log_001.md),上下文里只注入一行:[发现5个504错误,详见 refs/log_001.md]。
第三步:下沉确定性逻辑,用代码替代 LLM 节点
目标:将你之前用 Gatekeeper 和 VerifiedInput 做的事,降级为零 LLM 调用的代码拦截器。
前置拦截(Pre-Tool Hook):Agent 发起工具调用时,Harness 代码用 JSON Schema 校验参数格式。不合法直接报错打回,不执行工具。
后置断言(Post-Tool Hook):Agent 输出最终诊断报告时,代码层强制校验报告中引用的 evidence_id 和 raw_path 是否真实存在于历史记录或文件系统中。不合法直接拒绝输出,发回重试。
第四步:锁死输出契约
目标:防止 Executor 过度输出和发散。
在 System Prompt 中强制规定 Agent 的中间思考步和最终输出步必须符合严格的 JSON 结构。
例如,中间步必须是 {"thought": "<不超过50字>", "action": "queryLogs", "parameters": {...}}。Harness 代码检查字数,超长直接打回。
第五步:剥离异步验证(保留你最初的“防幻觉”初衷)
目标:在不增加主链路复杂度的前提下,保留交叉验证能力。
主 ReAct Agent 输出报告后,不要在主 Graph 里串行接一个 Verifier 节点。
改为异步触发一个轻量级 LLM(或小模型),只传入“压缩后的证据摘要 + 草稿结论”。让它判断时间线与逻辑是否一致。如果不一致,在最终输出上加“低置信度警告”;如果一致,直接放行。
总结:你的重构后架构全景图
重构后,你的代码结构应该极其清爽,大致如下:
[用户输入]
│
▼
[Diagnosis ReAct Agent] (唯一的 LLM 推理节点,自带规划、执行、总结)
│
├── Tool: queryLogs
│ └── [Harness 代码]: 过滤 INFO,提取 ERROR,写入 refs,返回摘要
├── Tool: queryMetrics
│ └── [Harness 代码]: 聚合统计值,返回 3 行核心指标
│
▼ (Agent 输出最终 JSON 报告)
[Output Schema Linter] (纯代码层,0 LLM)
│
├─ 校验失败 ──> 返回错误给 [Agent] 重新生成
│
▼ (校验通过)
[Async Verifier] (异步轻量 LLM,只做时间线/逻辑一致性校验)
│
▼
[最终输出 / 带警告输出]
现在你应该做的第一件事:打开你的代码,把主链路上除 Diagnosis Agent 以外的所有 LLM 编排节点全部注释掉,然后按照上面的结构,给工具加上 Pre-Tool 和 Post-Tool 的代码拦截器。把精力从“画 Graph”转移到“写工具清洗代码”上。
+4 -1
View File
@@ -1,6 +1,6 @@
# MVP Issues 索引 # MVP Issues 索引
**更新日期**:2026-07-10 **更新日期**:2026-07-20
**状态**:按活跃问题、设计笔记、RAG 问题集和已归档问题整理 **状态**:按活跃问题、设计笔记、RAG 问题集和已归档问题整理
## 目录约定 ## 目录约定
@@ -18,6 +18,9 @@
|---|---|---|---|---| |---|---|---|---|---|
| ISS-003 | MVP 设计与实现 Review 收敛 | 高 | 待规划 | [active/ISS-003-mvp-design-implementation-review.md](active/ISS-003-mvp-design-implementation-review.md) | | ISS-003 | MVP 设计与实现 Review 收敛 | 高 | 待规划 | [active/ISS-003-mvp-design-implementation-review.md](active/ISS-003-mvp-design-implementation-review.md) |
| ISS-004 | Executor 域级检索水位控制 | 低 | 待规划 | [active/ISS-004-executor-domain-hard-limit.md](active/ISS-004-executor-domain-hard-limit.md) | | ISS-004 | Executor 域级检索水位控制 | 低 | 待规划 | [active/ISS-004-executor-domain-hard-limit.md](active/ISS-004-executor-domain-hard-limit.md) |
| ISS-012 | Executor Token 预算与上下文膨胀 | 高 | 待规划 | [active/ISS-012-executor-token-budget-and-context-growth.md](active/ISS-012-executor-token-budget-and-context-growth.md) |
| ISS-013 | Chat 入口解耦与真正 SSE 收敛 | 高 | 待规划 | [active/ISS-013-chat-entry-decoupling-and-sse.md](active/ISS-013-chat-entry-decoupling-and-sse.md) |
| ISS-014 | 单体 ReAct Agent、Harness 与 ACI 工具瘦身 | 高 | 待阶段 0 冻结 | [active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md](active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md) |
| executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [active/executor-evidence-attribution-hallucination.md](active/executor-evidence-attribution-hallucination.md) | | executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [active/executor-evidence-attribution-hallucination.md](active/executor-evidence-attribution-hallucination.md) |
| rag-refactor-plan | RAG 检索重构计划 | 高 | 待规划 | [active/rag-refactor-plan.md](active/rag-refactor-plan.md) | | rag-refactor-plan | RAG 检索重构计划 | 高 | 待规划 | [active/rag-refactor-plan.md](active/rag-refactor-plan.md) |
@@ -0,0 +1,119 @@
# ISS-012 Executor Token 预算与上下文膨胀
**状态**:待规划
**严重程度**:高
**发现时间**:2026-07-20
**关联**:ISS-002、ISS-004、ISS-011
---
## 背景
ISS-011 完成 StateGraph 切换后,复杂 Chat 链路已经具备显式 Node、Gatekeeper、Verifier、Composer 和 Run 级 Trace。但最终 live E2E 暴露出 Executor 的 Token 成本和上下文增长问题:一次只返回安全 Fallback 的请求消耗了超过 11 万 Token。
## 现象
最终验收 Run:
- `runId`:`run-808ac38f-3ad0-4462-a6d0-ed50d8686473`
- 总耗时:`75964ms`
- 最终答案:109 字符,Run 为 `CHAT/SUCCESS`,`degraded=true`
- AgentStep:8 条,其中 Planner 1 次、Executor 7 次
- ToolInvocation:12 次
- Run `total_token_count`:`111802`
Executor 每次模型调用的 Token 逐步上升:
```text
6396 -> 11265 -> 12786 -> 16774 -> 17965 -> 19340 -> 24236
```
工具调用包括 4 次 `lookup_knowledge`、6 次 `query_logs`、1 次 `query_metrics` 和 1 次 `get_available_log_topics`。最终因模型输出缺少 `source_invocation_id`,Gatekeeper 将结果降为 `LOW_CONFID` 并进入安全 Fallback。
## 已确认事实
1. 数据库中 8 条 `agent_step` 均为不同记录,不存在重复插入;`111802` 等于各步骤 `token_count` 的求和。
2. `TokenTrackingChatModel` 当前只保存供应商返回的 `usage.totalTokens`,没有拆分输入、输出、缓存和推理 Token。
3. `ChatService.backfillRunMetrics` 直接累加每条 AgentStep 的 `token_count`。
4. `AgentLoggingHook` 只把模型输入截断为 500 字符写入审计表,无法从当前 Trace 还原模型实际发送的完整 Prompt。
5. Executor 当前没有独立的模型调用次数、工具调用次数或 Token 预算;Graph recursion limit 不能限制 ReactAgent 内部工具循环。
## 初步根因假设
- 每次 Executor 模型调用都会重新携带 Planner 结果、历史消息和之前的工具返回,导致输入上下文随工具循环增长。
- 工具返回内容包含较多日志、知识库结果和检索明细,完整结果被反复带入后续模型请求。
- 工具结果没有稳定返回 `source_invocation_id`,模型无法可靠生成精确证据引用,导致高成本检索后仍然进入 Fallback。
- 当前只能看到 `totalTokens`,尚未确认供应商 usage 中 input/output/cached/reasoning 的精确占比。
## 影响
- 单次诊断成本和延迟不可控,复杂问题可能继续超过模型上下文窗口。
- Token 消耗与最终答案质量不匹配,出现“高成本检索 + 安全降级”的低收益路径。
- 缺少 Token 分项指标,无法建立成本预算、P95 延迟和 degraded rate 门禁。
- Executor 可能重复查询相同或相近的知识域、日志主题和指标。
## 目标
1. 建立按 Run/AgentStep 的 input、output、cached、reasoning Token 可观测性。
2. 为 Executor 增加硬性模型轮数、工具调用和 Token 预算。
3. 将完整工具结果留在 Run Trace/数据库中,模型上下文只接收有界证据投影。
4. 让工具结果直接携带可引用的 `source_invocation_id` 和紧凑 `evidence_refs`。
5. 在预算耗尽时安全结束并明确记录原因,不绕过 Gatekeeper、Verifier 或 Run Trace。
## 建议方案
### 1. Token 统计拆分
- 从 ChatModel usage 中记录 `input_tokens`、`output_tokens`、`cached_tokens`、`reasoning_tokens`(供应商提供时)。
- 保留 `total_token_count` 作为汇总字段,但明确其计算口径。
- 在 `orchestration_trace` 中记录每个 Agent 的累计 Token 和预算命中情况。
### 2. Executor 硬预算
初版建议从以下上限开始,并通过固定 E2E 调整:
- Executor 模型调用最多 4 次。
- 工具调用最多 8 次。
- `lookup_knowledge` 最多 2 次。
- `query_logs` 默认最多返回 5 条日志,并限制单次输出长度。
- 达到预算后停止扩展检索,基于已验真证据输出,或进入带原因的安全 Fallback。
### 3. 有界证据上下文
- 工具完整原始结果继续写入 `tool_invocation`,不直接作为下一轮完整上下文。
- 返回模型的工具视图只保留 invocation ID、工具名、查询条件、有限 evidence refs、excerpt 和 no-evidence 状态。
- 同一工具数组项禁止重复绑定;相同知识域和日志主题不重复查询。
### 4. 证据引用闭环
- 每次 evidence tool 返回结果时直接包含 `source_invocation_id`。
- Executor 输出必须引用该 ID;Gatekeeper 不再依赖事后猜测或唯一候选补全。
- 由于引用失败进入 Fallback 时,Trace 必须记录具体缺失字段和预算消耗。
## 验收标准
- [ ] 每个 AgentStep 可查看 input/output/total Token,供应商支持时可查看 cached/reasoning Token。
- [ ] 固定 `payment-timeout` E2E 的 Token 上限、工具调用上限和最大延迟已定义并通过回归。
- [ ] 连续至少 10 次相同 fixture 运行,Token 和延迟 P95 不超过定义的预算。
- [ ] Executor 预算耗尽时只走安全 Fallback,不绕过 Gatekeeper、Verifier 或 Trace 持久化。
- [ ] 工具返回包含真实 `source_invocation_id`;正常证据链不再因缺少该字段而无谓降级。
- [ ] 12 个 diagnosis eval fixture、Graph workflow/node contract、Trace ownership 回归全部通过。
- [ ] E2E 日志和数据库能按 exact `sessionId + runId` 对齐 Token、工具调用、Fallback 原因和最终状态。
## 非目标
- 不删除 Gatekeeper、Verified Input 或 Verifier。
- 不以降低模型 `maxTokens` 代替上下文治理。
- 不恢复 Sequential/StateGraph 双轨或旧兼容协议。
- 不在本 Issue 中物理删除数据库中的历史 `diagnosis_session` 表。
## 相关文件
- `src/main/java/com/superbiz/agent/hook/TokenTrackingChatModel.java`
- `src/main/java/com/superbiz/agent/hook/AgentLoggingHook.java`
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- `src/main/java/com/superbiz/agent/graph/diagnosis/ReactAgentDiagnosisInvoker.java`
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
- `src/main/resources/prompts/chat-executor-prompt.md`
- `mvp/architecture/stategraph-runtime-architecture.md`
- `devflow/projects/2026-07-17-chat-diagnosis-stategraph-cleanup-docs/evidence.md`
@@ -0,0 +1,81 @@
# ISS-013 Chat 入口解耦与真正 SSE 收敛
**状态**:待规划
**严重程度**:高
**发现时间**:2026-07-20
**关联**:ISS-011、ISS-012
---
## 背景
当前 Chat 入口同时提供 `/api/chat` 和 `/api/chat_stream`。`ChatController` 不仅处理 HTTP/SSE 协议,还直接承担会话创建、历史读取与回写、模型和工具获取、执行策略调用以及异常响应组装,入口职责已经明显超出协议适配层。
现有 `/api/chat_stream` 会等待完整答案生成后再按固定长度切片发送,并不是真正的流式生成。同步和伪流式入口还复制了大部分业务流程,增加了维护成本和行为不一致风险。
## 已确认问题
1. Chat 的会话、历史、模型、工具、执行和结果回写功能耦合在 `ChatController` 中,HTTP 层与应用用例边界不清晰。
2. `/api/chat` 与 `/api/chat_stream` 重复编排同一套 Chat 流程。
3. `/api/chat_stream` 只是对完整答案做事后分块,不具备模型生成过程中的真实增量输出能力。
4. Controller 直接获取 `ChatModel` 和 `ToolCallbackProvider`,将模型基础设施细节暴露到入口层。
5. SSE 使用 Controller 自建的无界缓存线程池,缺少统一生命周期和容量治理。
6. 当前入口错误响应存在 HTTP 状态、外层 `ApiResponse` 与内层 `ChatResponse` 状态不一致的问题。
## 目标
1. 将会话生命周期、历史管理、执行调用和结果回写从 Controller 分离,形成单一 Chat 应用用例入口。
2. 只保留一个 `/api/chat` 接口,并将其协议改为真正的 SSE。
3. SSE 在模型或诊断链路产生内容时增量发送,而不是等待完整答案后再切片。
4. 保留 `sessionId + runId` 作为一次 Chat Run 的稳定关联契约。
5. 在满足入口职责分离的前提下使用最少组件,不引入没有实际职责的接口、工厂或适配层。
## 设计约束
- Controller 只负责请求校验、协议转换和 SSE 生命周期,不负责选择模型、组装工具、管理历史或编排诊断流程。
- 同一次请求只能进入一个应用用例入口,禁止同步和流式路径各自维护一套业务逻辑。
- 真正 SSE 至少需要区分元数据、内容增量、完成和错误事件。
- `sessionId`、`runId` 必须在内容事件之前可获得,并用于日志、数据库和 Trace 对齐。
- 客户端断开、超时和执行失败必须显式终止后台执行并完成 Run 状态记录。
- 不保留旧 `/api/chat_stream` 或同步 `/api/chat` 的兼容分支,直接以新协议为准。
- 优先使用 Spring 管理的执行设施和现有服务能力,不创建无界线程池。
## 建议的最小边界
```text
POST /api/chat (SSE)
-> ChatController:请求与 SSE 协议
-> Chat 应用用例:会话、Run、历史和执行生命周期
-> 现有 Chat 执行能力:简单回答或诊断编排
```
这里的“应用用例”是职责边界,不要求预先拆出多层接口。只有出现独立变化原因或明确复用需求时才增加新组件。
## 验收标准
- [ ] 对外只保留一个 `POST /api/chat`,响应类型为 `text/event-stream`。
- [ ] 删除 `/api/chat_stream` 及同步 Chat 兼容路径。
- [ ] 首个内容事件在完整答案生成完成前发送,禁止通过固定字符切片伪造流式输出。
- [ ] SSE 事件包含稳定的 metadata、content、error、done 契约。
- [ ] Controller 不再直接依赖 `ChatModel`、`ToolCallbackProvider`,也不管理会话历史和 Run 持久化。
- [ ] 同一请求的 `sessionId + runId` 在 SSE、应用日志、`diagnosis_run`、`agent_step` 和 `tool_invocation` 中一致。
- [ ] 客户端断开、超时、模型失败和工具失败都有明确的资源清理与 Run 终态。
- [ ] 不存在 Controller 自建的无界线程池。
- [ ] 单元测试覆盖入口校验和 SSE 事件契约;端到端测试验证真实增量输出、断开清理及 Trace 对齐。
## 非目标
- 不在本 Issue 中重新设计 StateGraph 节点、Gatekeeper、Verifier 或证据协议。
- 不为未来可能出现的其他传输协议预建通用框架。
- 不引入多套 Command、Handler、Adapter、Factory 只为形式上的分层。
- 不保留旧同步接口或 `/api/chat_stream` 的兼容逻辑。
- 不以“完整答案分块发送”作为 SSE 验收通过条件。
## 相关文件
- `src/main/java/com/superbiz/agent/controller/ChatController.java`
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- `src/main/java/com/superbiz/agent/service/session/SessionManager.java`
- `src/test/java/com/superbiz/agent/controller/ChatControllerTest.java`
- `src/test/java/com/superbiz/agent/service/ChatServiceGraphIntegrationTest.java`
- `mvp/architecture/current-mvp-architecture.md`
File diff suppressed because it is too large Load Diff
@@ -0,0 +1 @@
Devflow archive prepared and stage verification passed on 2026-07-21.
@@ -0,0 +1 @@
Committed after strict validation on 2026-07-21.
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-21
@@ -0,0 +1,29 @@
# Stage 0 Acceptance Evidence
## Static Verification
- Contract package is dependency-free from Redis, JPA, Controller, Agent state and existing Hook classes.
- Repository secret scan covers tracked worktree files and reports no known plaintext credential matches.
- `scripts/query_mysql.py` requires `SUPERBIZ_MYSQL_PASSWORD` and exits before connecting when it is absent.
- Spring AI 1.1.7 `SpringAiRetryProperties` bytecode shows a default `maxAttempts` value of 10; stage 2 must set underlying retries to one attempt and keep retry ownership in Harness.
## Script Verification
- `mvn -q '-Dtest=HarnessContractTest' test` - passed.
- `mvn -q '-Dtest=ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,ChatControllerTest' test` - passed.
- `mvn -q '-Dtest=HarnessContractTest,ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,ChatControllerTest' test` - passed.
- `openspec validate single-react-design-freeze --strict` - passed.
- `git diff --check` - passed; only existing Windows line-ending warnings were reported.
- Secret scan for known committed key/password patterns - zero matches.
- `python scripts/query_mysql.py "SELECT 1"` without `SUPERBIZ_MYSQL_PASSWORD` - exited before connecting with the expected missing-variable error.
## Runtime Behavior
- Public Chat runtime was not switched in stage zero.
- No live model, Redis, MySQL or Milvus E2E was run; full live E2E remains stage 7 scope.
## External Security Prerequisite
- Plaintext credentials previously present in the repository must be rotated in their respective MySQL, Redis, DeepSeek, SiliconFlow and Milvus systems by the credential owner.
- Repository changes can prove removal but cannot prove provider-side rotation.
- Stage 3C and stage 7 must not claim live security/E2E acceptance until required environment variables contain rotated credentials.
@@ -0,0 +1,25 @@
# Brief: single-react-design-freeze
## Background
ISS-014 将当前 Chat 多 Agent/Hook/ThreadLocal 主链路重构为一个 Diagnosis ReAct Agent、一个确定性 Harness 和一个隔离 SemanticGuard。阶段 0 先冻结后续 10 个实施 change 共同依赖的契约和安全边界。
## Goal
产出可执行、可测试、可归档的 contract types、失败语义、安全前置和阶段门禁,同时保持现有公开 Chat 运行行为不变。
## Scope
- 类型化 Draft、Knowledge Answer、Fallback、previous turn 和状态枚举。
- Tool ID、双状态、取消、重试、Redis canonical record 和 MySQL 安全设计冻结。
- 明文脚本凭据清理和 focused baseline。
- ISS-014、OpenSpec、devflow 对齐。
## Non-Goals
- 不实现或接入新 Harness/Agent/Guard。
- 不切换 `/api/chat`、不删除旧链路、不运行 live E2E。
## Source PRD
`mvp/issues/active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md` 是本 change 的完整 PRD 和总设计来源,不复制为第二份 PRD。
@@ -0,0 +1,80 @@
# Decisions
## Entry
- Parent issue: `ISS-014`
- Change: `single-react-design-freeze`
- Scale: `complex`
- Interface impact: future L4; this change freezes contracts without switching runtime behavior.
- Capability sources: sm-flow built-in clarify/context/propose, `grill-with-docs`, `openspec-propose`, `zoom-out`, `openspec-apply-change`, `openspec-archive-change`.
## Context Evidence
- `session-run-trace-isolation` established Chat Session and Diagnosis Run as separate lifecycles and made `runId` the Trace ownership key.
- `verifier-evidence-reference-fidelity` established that no-evidence is a scoped negative observation, not proof that a problem does not exist.
- `executor-composer-final-answer` established deterministic safe fallback boundaries and prohibited unfiltered raw output from reaching users.
- `modular-rag-pipeline` established that retrieval trace and context packing are audit details rather than direct facts.
- Current Spring AI Alibaba `ToolCallRequest` already provides `tool_call_id`; Harness must validate and persist it rather than create a second identity.
- Current Spring AI `ChatModel.call(Prompt)` has no cancellation token, so cancellation must be expressed as layered, observable semantics rather than an unsupported absolute guarantee.
## Question Pool
| ID | Dimension | Question | Mode | Status |
|---|---|---|---|---|
| Q1 | Terminology | Does `tool_call_id` use the framework ID or a Harness-generated ID? | user-interview | confirmed: framework ID |
| Q2 | Terminology | Are invocation lifecycle and evidence outcome separate fields? | user-interview | confirmed: `status` + `evidence_status` |
| Q3 | Boundary | Is ISS-014 one umbrella Issue with independent OpenSpec changes? | user-interview | confirmed: one Issue, 11 changes |
| Q4 | Boundary | May stage 4 publish before Guards exist? | user-interview | confirmed: no; public cutover only in 6B |
| Q5 | Lifecycle | What is the durable source of truth for `previous_turn` and `last_intent`? | user-interview | confirmed: `diagnosis_run` safe published result |
| Q6 | Contract | How is KNOWLEDGE_QUERY citation validation represented? | evidence-driven | resolved: structured answer items with exact RAG bindings |
| Q7 | Cancellation | What cancellation guarantees are technically enforceable? | evidence-driven | resolved: layered cancellation, no false hard-cancel claim |
| Q8 | Acceptance | Are Apply, Archive and phase Git commit pre-authorized? | user-interview | confirmed: yes, for all phases |
## Confirmed Decisions
- `tool_call_id` is the framework Tool Call protocol ID. Harness validates non-empty, bounded, safe characters and Run-local uniqueness; duplicate/invalid/missing IDs fail closed.
- Redis `status=PROJECTING/READY/ERROR` represents invocation/projector lifecycle.
- `evidence_status=EVIDENCE_FOUND/NO_EVIDENCE/ERROR` represents result semantics.
- `NO_EVIDENCE` may only support `NEGATIVE_OBSERVATION` within the exact query scope; it cannot prove absence, exclusion or health.
- Stage 4 and 6A remain internal. Stage 6B performs the only public Chat cutover after stage 5 release gates pass.
- ISS-014 remains the umbrella Issue. Eleven independent changes run serially; each must complete sm-flow, OpenSpec archive and Git commit before the next starts.
- Full live E2E is deferred to stage 7; earlier stages run focused verification proportional to their change.
- KNOWLEDGE_QUERY uses a dedicated structured draft: each answer item binds the single lookup `tool_call_id` and one or more returned `document_id` values; Harness validates exact membership before rendering `answer + references + limitations`. It does not reuse the Diagnosis Analysis schema and does not enter SemanticGuard in the first version.
- Cancellation is layered: mark cancellation requested, prevent new model/Tool rounds and any final Draft release, invoke framework interruption, cancel owned Tool/JDBC work where supported, and rely on configured HTTP timeouts for an already-blocking synchronous model call. Run finalization is atomic and late results are discarded.
- Public SSE `done` is not emitted after the client has disconnected; internal Run state still reaches `CANCELLED`.
- `diagnosis_run` is the durable source for `intent`, `release_outcome` and `published_result`. Only the latest same-session `DIAGNOSIS + SUCCESS` record with a non-null safe published result may become `previous_turn`; `FALLBACK/FAILED/CANCELLED` remain auditable but are excluded.
- `published_result` stores only `user_query/published_conclusion/scope/limitations/source_documents`; it excludes Tool Call IDs, raw evidence, full Draft and SemanticGuard audit reasons.
## Evidence-Driven Findings To Report
- The framework already exposes `ToolInterceptor`, structured output types, tool execution timeout, model/tool call limit hooks and `ReactAgent.interrupt`; later Harness stages should reuse these extension points.
- Spring AI model dependencies include retry support, while current application configuration does not explicitly freeze all retry layers; stage 0 must define a retry inventory and stage 2 must enforce it.
- Current `DiagnosisRun` stores a text answer and generic status but has no explicit `intent`, `release_outcome` or structured published result; Q5 must be resolved before the previous-turn contract is executable.
- Current KNOWLEDGE_QUERY target behavior promises citation validation, but the issue only defines the Diagnosis Draft binding schema; the committed spec must add the dedicated answer-item contract described above.
- Spring AI retry auto-configuration defaults `maxAttempts` to 10. The target Harness retry matrix requires underlying model/HTTP retries to be set to one attempt, with Router and SemanticGuard retries performed only by Harness.
## OpenSpec Backfill
- All confirmed decisions above must appear in design/specs/tasks before `.committed` is created.
- All user-interview questions are confirmed; no pending decision blocks Commit.
## Architecture Audit
Current input flows from `ChatController` into `ChatService`, which owns routing, ReactAgent construction, multi-Agent orchestration and final rendering; tools persist evidence through `ToolInvocationRecorder`, while Hooks and ThreadLocal bridge Run and verifier state. Stage zero introduces only dependency-free contract types under `harness.contract`; those types must not depend on Controller, Redis, JPA, Spring Agent state or current Hook classes. Later stages move ownership in order: RunContext, invocation store, Tool-specific projection, Diagnosis Agent, Guards, application use case and finally the public SSE adapter. `diagnosis_run` remains durable Run ownership, Redis canonical invocation remains short-lived Harness ownership, and `agent_step/tool_invocation` remain durable audit detail. The principal risk is spec/runtime drift, mitigated by archiving only the stage-zero contract capability now and delaying modifications to existing runtime capabilities until their implementation changes.
## Cross-Artifact Alignment
| Chain | Status | Evidence |
|---|---|---|
| ISS-014/brief goals, scope and non-goals → proposal | aligned | Proposal limits stage zero to contracts, security and baseline with no public cutover. |
| proposal commitments → design | aligned | Design records every ID, status, Draft, fallback, previous-turn, retry, cancellation and phase-gate commitment. |
| design decisions → specs | aligned | The single stage-zero capability has testable requirements for every stable contract boundary. |
| specs observable behavior → tasks | aligned | Tasks create reusable types/tests, remove the secret, align artifacts and verify without switching runtime behavior. |
## Commit Gate Result
- Question pool covers terminology, boundary, lifecycle, contract, cancellation and acceptance.
- All user-interview items are explicitly confirmed.
- Evidence-driven conclusions were reported and written into design/spec/tasks.
- Interface impact is recorded as future L4; this change itself does not switch the public API.
- No devflow/OpenSpec conflict remains.
@@ -0,0 +1,77 @@
## Context
ISS-014 是一次 L4 Chat 重构的总设计来源,但实施被拆成 11 个必须串行归档的 OpenSpec changes。阶段 0 不切换公开协议或 Agent 运行链,只创建后续阶段复用的类型化契约、安全前置、失败语义和 focused baseline。
现有代码已经具备 `runId` Trace、工具调用审计、no-evidence 精确引用、Verifier/Composer fallback 和模块化 RAG,但这些能力分散在 `ChatService`、Hook、ThreadLocal、Tool 和 JSON 字符串中。Spring AI Alibaba 已提供 `ToolInterceptor`、结构化输出类型、工具执行超时、调用限额 Hook 和 `ReactAgent.interrupt`;Spring AI 底层 retry 默认最多 10 次,不能直接满足 ISS-014 的显式 Harness 重试矩阵。
## Goals / Non-Goals
**Goals:**
- 生成后续阶段可直接复用的 Java contract types 和枚举,不实现新 Agent 执行链。
- 冻结 Tool ID、双状态、Draft、Knowledge Answer、Fallback、previous turn、SSE、重试和取消语义。
- 冻结 MySQL fail-closed 允许子集和安全前置。
- 移除仓库脚本和主配置中的明文凭据并建立改造前 focused baseline。
- 保证 ISS-014、OpenSpec、devflow 术语和 11 个阶段门禁一致。
**Non-Goals:**
- 不接入 Harness、Diagnosis Agent、EvidenceGuard 或 SemanticGuard 运行时。
- 不修改 Controller 协议、旧 ChatService 行为、Redis invocation store 或数据库表。
- 不实现 RAG/日志/MySQL Tool 投影。
- 不运行完整 live E2E。
## Decisions
### Contract types are reusable runtime inputs
阶段 0 创建位于 `com.superbiz.agent.harness.contract` 的轻量 record/enum,而不是只写文档或引入 JSON Schema 引擎。后续 ReactAgent `outputType`、Harness validator、持久化和 SSE DTO 可以直接复用这些类型,减少同一字段在多个阶段重复定义。
### Framework Tool Call ID is canonical
`tool_call_id` 使用框架协议 ID。Harness 后续只校验非空、长度/字符安全和 Run 内唯一性,不生成第二套 ID。Redis Key 仍按 `runId + toolCallId` 隔离,真实性来自当前 Run 的 canonical record,而不是 ID 本身。
### Invocation lifecycle and evidence outcome are orthogonal
`InvocationStatus=PROJECTING/READY/ERROR` 只描述调用与投影生命周期;`EvidenceStatus=EVIDENCE_FOUND/NO_EVIDENCE/ERROR` 描述结果语义。`NO_EVIDENCE` 只允许绑定 `AnalysisKind=NEGATIVE_OBSERVATION`,并必须保留查询范围和零匹配信息。
### Diagnosis and knowledge answer contracts stay separate
Diagnosis Draft 使用 Analysis ID 与 Tool Call IDs;KNOWLEDGE_QUERY 使用 answer items,每项绑定一次 lookup 的 Tool Call ID 和返回的 document IDs。Harness 后续验证精确成员关系并确定性展开引用。首版 KNOWLEDGE_QUERY 不进入 SemanticGuard。
### Safe published context belongs to Diagnosis Run
`diagnosis_run` 是 `intent/release_outcome/published_result` 的持久化真理源。只有同 Session 最近一个 `DIAGNOSIS + SUCCESS` 且 published result 非空的 Run 可形成 previous turn。Published result 只含用户查询、已发布结论、范围、限制和 RAG 文档元数据。
### Cancellation is observable and layered
取消请求立即阻止新模型/Tool 轮次和最终 Draft 释放;框架中断、可控 Future/JDBC 取消尽力执行;已进入同步 `ChatModel.call` 的请求依靠底层 HTTP timeout。Run 终态通过原子状态转换保证唯一,晚到结果被丢弃,不宣称无法证明的底层硬取消。
### Harness owns all retries
底层 SDK/HTTP/数据库 retry 必须关闭或压为一次 attempt。Intent Router 和 SemanticGuard 的第二次 attempt 由 Harness 显式执行并审计。阶段 0 只冻结矩阵;阶段 2 实现执行器和配置。
### Stage boundaries are release boundaries
阶段 4 和 6A 仅内部运行;阶段 6B 在 Guards 已完成后原子切换公开入口。每个 change 必须 Archive 并 Git commit 后才能进入下一项,阶段 7 才执行统一 live E2E。
## Risks / Trade-offs
- [Risk] Contract records过早绑定实现细节 → Mitigation:只包含跨阶段稳定字段,不包含 Redis、JPA 或框架对象。
- [Risk] Spring AI provider 对 Tool Call ID 行为不同 → Mitigation:阶段 3A 使用 Fake Model 和当前 DeepSeek 路径验证,缺失或重复时 fail closed。
- [Risk] 同步模型调用不能立即取消 → Mitigation:明确 layered semantics、HTTP timeout 和晚到结果丢弃,不把状态更新等同底层资源已终止。
- [Risk] 设计冻结测试增加维护成本 → Mitigation:只保留 focused serialization/validation tests,不复制完整 E2E。
- [Risk] 明文凭据可能已泄露 → Mitigation:仓库中移除并要求外部轮换;轮换证据记录在 acceptance。
## Migration Plan
1. 创建 contract types、契约测试和阶段台账,不接运行链。
2. 清理查询脚本凭据,改为环境变量注入。
3. 记录 focused baseline;Archive 本 change。
4. 后续 10 个 change 逐步实现,并在各自 Archive 时同步对应运行 capability specs。
Rollback:阶段 0 没有公开行为变化;可删除新增 contract package/tests 并恢复文档。已经轮换的凭据不得回滚为旧值。
## Open Questions
无。所有影响实现的用户决策已经在 `decisions.md` 中确认。
@@ -0,0 +1,32 @@
## Why
当前 Chat 诊断把一个 ReAct 生命周期拆成多个 Agent、Hook、ThreadLocal 和重试分支,导致证据契约、失败语义、上下文预算和公开释放边界分散。进入分阶段重构前,需要先把 ISS-014 的跨阶段契约、安全前置和验收基线冻结为唯一可执行规格,避免后续 change 各自解释同一概念。
## What Changes
- 冻结单体 Diagnosis ReAct Agent、确定性 Harness、EvidenceGuard 和隔离 SemanticGuard 的职责边界。
- 冻结 Diagnosis Draft、Analysis、Conclusion、Fallback、Intent Router 和 SSE 事件契约。
- 冻结 Tool Call ID、调用生命周期 `status`、结果语义 `evidence_status`、Redis canonical invocation 和 ToolResultProjector 命名。
- 冻结预算、取消、重试、隐藏重试禁用和 Run 终态语义,但不在本阶段实现新运行链路。
- 冻结只读 MySQL Tool 的 JSqlParser 允许子集、静态 allowlist 和安全前置。
- 移除仓库脚本和主配置中的明文数据库、Redis、模型及向量服务凭据,并记录必须完成外部轮换。
- 建立改造前 focused baseline 和后续 11 个串行 sm-flow change 的阶段台账。
- **BREAKING(后续阶段实施)**:最终仅保留 `POST /api/chat` SSE、移除旧多 Agent/Graph Chat 主链路和旧 Tool Contract;本 change 不执行公开协议切换。
## Capabilities
### New Capabilities
- `single-react-diagnosis-harness`: 冻结单体诊断 Agent、Harness、证据状态、Guard、Fallback、预算、重试和阶段门禁的跨阶段基础契约。
### Modified Capabilities
- None. 本阶段不声明旧运行能力已经迁移;后续 change 在实现对应行为时再修改现有 capability specs。
## Impact
- 设计与规格:ISS-014、OpenSpec 主规格、devflow 词汇表和阶段归档台账。
- 契约测试:Draft/Fallback、状态语义、SSE、Router、预算、重试和 MySQL 安全基线。
- 安全:`scripts/query_mysql.py` 和 `application.yml` 中的明文凭据必须移除并在外部轮换。
- 后续代码范围:ChatController、ChatService、Agent/Hook、Tool、Redis、JPA/Flyway、静态前端、Trace/Eval fixtures。
- 本 change 不切换 Controller 协议、不接入新 Agent、不修改公开运行行为。
@@ -0,0 +1,75 @@
## ADDED Requirements
### Requirement: Design contracts SHALL separate deterministic control from diagnosis reasoning
The frozen contract set SHALL define Diagnosis Agent as the only diagnosis report author, Harness as deterministic execution control, EvidenceGuard as deterministic evidence validation, and SemanticGuard as an isolated single-turn semantic reviewer.
#### Scenario: Contract ownership is inspected
- **WHEN** a later phase reads the stage-zero contracts
- **THEN** no Harness contract assigns Planner, Executor, Composer, workflow routing, or diagnosis reasoning responsibilities to Harness
### Requirement: Tool invocation identity SHALL use the framework Tool Call ID
The contract SHALL use the framework-provided `tool_call_id` as the sole Tool invocation reference and SHALL require later Harness implementations to reject missing, invalid, or duplicate IDs within a Run.
#### Scenario: Duplicate Tool Call ID is proposed
- **WHEN** two Tool actions in one Run present the same framework Tool Call ID
- **THEN** the contract classifies the second action as an error and prohibits overwriting the first canonical invocation
### Requirement: Invocation status and evidence status SHALL be independent
The contract SHALL define `PROJECTING/READY/ERROR` as invocation lifecycle states and `EVIDENCE_FOUND/NO_EVIDENCE/ERROR` as evidence result states.
#### Scenario: Successful query returns no evidence
- **WHEN** a Tool executes successfully and its bounded projection contains zero matching evidence
- **THEN** invocation status is `READY` and evidence status is `NO_EVIDENCE`
#### Scenario: No-evidence result is cited
- **WHEN** a Diagnosis Draft cites a `NO_EVIDENCE` Tool result
- **THEN** the Analysis kind MUST be `NEGATIVE_OBSERVATION` and MUST remain bounded to the Tool query scope
### Requirement: Diagnosis Draft SHALL expose typed report structure
The contract SHALL define conclusion, analysis items, action plan, recommendations, limitations and Tool Call bindings without exposing chain-of-thought or raw Tool payloads.
#### Scenario: Draft contains an analysis item
- **WHEN** the Diagnosis Agent emits a structured Draft
- **THEN** every Analysis has a unique analysis ID, a fixed Analysis kind and at least one Tool Call ID
### Requirement: Knowledge answers SHALL use exact RAG bindings
The KNOWLEDGE_QUERY contract SHALL represent the answer as bounded answer items whose references identify the single lookup Tool Call and returned document IDs.
#### Scenario: Knowledge answer cites an unknown document
- **WHEN** an answer item references a document ID absent from the bounded lookup result
- **THEN** later Harness validation rejects the answer instead of publishing the fabricated citation
### Requirement: Safe fallback SHALL use a fixed schema
The contract SHALL define stable fallback types for evidence validation failure, semantic unsupported and semantic unavailable outcomes, and SHALL exclude unvalidated Draft content and internal errors.
#### Scenario: Evidence validation fails twice
- **WHEN** initial validation and the single no-Tool structural repair both fail
- **THEN** the fallback type is `EVIDENCE_VALIDATION_FAILED` and verified sources are empty
### Requirement: Previous turn SHALL come from a safe durable Run result
The contract SHALL define `diagnosis_run` as the durable source of intent, release outcome and safe published result, and SHALL exclude fallback, failed and cancelled Runs from Diagnosis previous-turn selection.
#### Scenario: Latest Run is a fallback
- **WHEN** the latest same-session Run ended with `FALLBACK`
- **THEN** it is not used as Diagnosis previous turn and selection continues to the latest eligible `DIAGNOSIS + SUCCESS` Run
### Requirement: Cancellation SHALL be layered and observable
The contract SHALL distinguish cancellation request, prevention of new work, framework interruption, cancellable Tool work and HTTP timeout for already-blocking synchronous model calls.
#### Scenario: Client disconnects during a model call
- **WHEN** an SSE client disconnects while a synchronous model call is in flight
- **THEN** the system prevents later Draft release, requests interruption, records an internal cancelled terminal state and discards any late model result
### Requirement: Retry attempts SHALL be owned by Harness
The contract SHALL require underlying SDK, HTTP and database retry layers to execute one attempt, while Harness explicitly owns any allowed Router or SemanticGuard retry.
#### Scenario: SemanticGuard returns an invalid schema
- **WHEN** the first SemanticGuard attempt returns an invalid structured result
- **THEN** Harness may execute one second attempt with the same verified snapshot and records both attempts
### Requirement: Phase gates SHALL remain serial
The implementation plan SHALL contain eleven independent OpenSpec changes and SHALL prohibit starting a change before its predecessor is archived and committed.
#### Scenario: Stage 4 completes internal Agent tests
- **WHEN** stage 4 passes its focused tests but stage 5 Guards are not implemented
- **THEN** the public Chat entry remains on the old path and stage 6B cutover is prohibited
@@ -0,0 +1,21 @@
## 1. Contract Model
- [x] 1.1 Add typed enums for intent, release outcome, invocation status, evidence status, analysis kind, semantic verdict, fallback type and SSE outcome.
- [x] 1.2 Add reusable records for Diagnosis Draft, Knowledge Answer Draft, safe fallback, published result and previous turn.
- [x] 1.3 Add focused serialization and contract-shape tests covering positive evidence, no-evidence negative observation and forbidden raw/internal fields.
## 2. Security And Configuration Baseline
- [x] 2.1 Remove tracked plaintext credentials from `scripts/query_mysql.py` and `application.yml`, require environment-based secrets, and ignore local secret files.
- [x] 2.2 Record the required external credential rotation and the Spring AI hidden-retry baseline without changing the public Chat runtime in this stage.
## 3. Design Freeze Alignment
- [x] 3.1 Keep ISS-014, OpenSpec artifacts and the devflow glossary aligned on the 11 serial changes, framework Tool Call ID, dual status fields and safe previous-turn source.
- [x] 3.2 Add an architecture audit and cross-artifact alignment result to `decisions.md`.
- [x] 3.3 Create the sm-flow `.committed` marker after proposal/design/specs/tasks and all decision gates pass.
## 4. Verification
- [x] 4.1 Run focused contract tests and the smallest existing Chat/evidence baseline needed to prove stage zero did not switch runtime behavior.
- [x] 4.2 Run strict OpenSpec validation and record commands, results and unverified external rotation in acceptance evidence.
@@ -0,0 +1,78 @@
# single-react-diagnosis-harness Specification
## Purpose
TBD - created by archiving change single-react-design-freeze. Update Purpose after archive.
## Requirements
### Requirement: Design contracts SHALL separate deterministic control from diagnosis reasoning
The frozen contract set SHALL define Diagnosis Agent as the only diagnosis report author, Harness as deterministic execution control, EvidenceGuard as deterministic evidence validation, and SemanticGuard as an isolated single-turn semantic reviewer.
#### Scenario: Contract ownership is inspected
- **WHEN** a later phase reads the stage-zero contracts
- **THEN** no Harness contract assigns Planner, Executor, Composer, workflow routing, or diagnosis reasoning responsibilities to Harness
### Requirement: Tool invocation identity SHALL use the framework Tool Call ID
The contract SHALL use the framework-provided `tool_call_id` as the sole Tool invocation reference and SHALL require later Harness implementations to reject missing, invalid, or duplicate IDs within a Run.
#### Scenario: Duplicate Tool Call ID is proposed
- **WHEN** two Tool actions in one Run present the same framework Tool Call ID
- **THEN** the contract classifies the second action as an error and prohibits overwriting the first canonical invocation
### Requirement: Invocation status and evidence status SHALL be independent
The contract SHALL define `PROJECTING/READY/ERROR` as invocation lifecycle states and `EVIDENCE_FOUND/NO_EVIDENCE/ERROR` as evidence result states.
#### Scenario: Successful query returns no evidence
- **WHEN** a Tool executes successfully and its bounded projection contains zero matching evidence
- **THEN** invocation status is `READY` and evidence status is `NO_EVIDENCE`
#### Scenario: No-evidence result is cited
- **WHEN** a Diagnosis Draft cites a `NO_EVIDENCE` Tool result
- **THEN** the Analysis kind MUST be `NEGATIVE_OBSERVATION` and MUST remain bounded to the Tool query scope
### Requirement: Diagnosis Draft SHALL expose typed report structure
The contract SHALL define conclusion, analysis items, action plan, recommendations, limitations and Tool Call bindings without exposing chain-of-thought or raw Tool payloads.
#### Scenario: Draft contains an analysis item
- **WHEN** the Diagnosis Agent emits a structured Draft
- **THEN** every Analysis has a unique analysis ID, a fixed Analysis kind and at least one Tool Call ID
### Requirement: Knowledge answers SHALL use exact RAG bindings
The KNOWLEDGE_QUERY contract SHALL represent the answer as bounded answer items whose references identify the single lookup Tool Call and returned document IDs.
#### Scenario: Knowledge answer cites an unknown document
- **WHEN** an answer item references a document ID absent from the bounded lookup result
- **THEN** later Harness validation rejects the answer instead of publishing the fabricated citation
### Requirement: Safe fallback SHALL use a fixed schema
The contract SHALL define stable fallback types for evidence validation failure, semantic unsupported and semantic unavailable outcomes, and SHALL exclude unvalidated Draft content and internal errors.
#### Scenario: Evidence validation fails twice
- **WHEN** initial validation and the single no-Tool structural repair both fail
- **THEN** the fallback type is `EVIDENCE_VALIDATION_FAILED` and verified sources are empty
### Requirement: Previous turn SHALL come from a safe durable Run result
The contract SHALL define `diagnosis_run` as the durable source of intent, release outcome and safe published result, and SHALL exclude fallback, failed and cancelled Runs from Diagnosis previous-turn selection.
#### Scenario: Latest Run is a fallback
- **WHEN** the latest same-session Run ended with `FALLBACK`
- **THEN** it is not used as Diagnosis previous turn and selection continues to the latest eligible `DIAGNOSIS + SUCCESS` Run
### Requirement: Cancellation SHALL be layered and observable
The contract SHALL distinguish cancellation request, prevention of new work, framework interruption, cancellable Tool work and HTTP timeout for already-blocking synchronous model calls.
#### Scenario: Client disconnects during a model call
- **WHEN** an SSE client disconnects while a synchronous model call is in flight
- **THEN** the system prevents later Draft release, requests interruption, records an internal cancelled terminal state and discards any late model result
### Requirement: Retry attempts SHALL be owned by Harness
The contract SHALL require underlying SDK, HTTP and database retry layers to execute one attempt, while Harness explicitly owns any allowed Router or SemanticGuard retry.
#### Scenario: SemanticGuard returns an invalid schema
- **WHEN** the first SemanticGuard attempt returns an invalid structured result
- **THEN** Harness may execute one second attempt with the same verified snapshot and records both attempts
### Requirement: Phase gates SHALL remain serial
The implementation plan SHALL contain eleven independent OpenSpec changes and SHALL prohibit starting a change before its predecessor is archived and committed.
#### Scenario: Stage 4 completes internal Agent tests
- **WHEN** stage 4 passes its focused tests but stage 5 Guards are not implemented
- **THEN** the public Chat entry remains on the old path and stage 6B cutover is prohibited
+13 -6
View File
@@ -22,13 +22,20 @@ except ImportError:
print("缺少依赖,请先执行: pip install pymysql") print("缺少依赖,请先执行: pip install pymysql")
sys.exit(1) sys.exit(1)
# 从 application.yml 读取的连接信息 def required_env(name: str) -> str:
value = os.getenv(name)
if value is None or not value.strip():
print(f"缺少必需环境变量: {name}")
sys.exit(2)
return value.strip()
DB_CONFIG = { DB_CONFIG = {
"host": "119.29.78.52", "host": os.getenv("SUPERBIZ_MYSQL_HOST", "119.29.78.52"),
"port": 33306, "port": int(os.getenv("SUPERBIZ_MYSQL_PORT", "33306")),
"user": "root", "user": os.getenv("SUPERBIZ_MYSQL_USERNAME", "root"),
"password": "!Fucker123..", "password": required_env("SUPERBIZ_MYSQL_PASSWORD"),
"database": "superbiz_agent", "database": os.getenv("SUPERBIZ_MYSQL_DATABASE", "superbiz_agent"),
"charset": "utf8mb4", "charset": "utf8mb4",
"cursorclass": pymysql.cursors.DictCursor, "cursorclass": pymysql.cursors.DictCursor,
} }
@@ -0,0 +1,13 @@
package com.superbiz.agent.harness.contract;
public enum AnalysisKind {
NORMAL,
NEGATIVE_OBSERVATION;
public boolean accepts(EvidenceStatus evidenceStatus) {
return switch (this) {
case NORMAL -> evidenceStatus == EvidenceStatus.EVIDENCE_FOUND;
case NEGATIVE_OBSERVATION -> evidenceStatus == EvidenceStatus.NO_EVIDENCE;
};
}
}
@@ -0,0 +1,13 @@
package com.superbiz.agent.harness.contract;
import java.util.List;
final class ContractCollections {
private ContractCollections() {
}
static <T> List<T> immutable(List<T> values) {
return values == null ? List.of() : List.copyOf(values);
}
}
@@ -0,0 +1,67 @@
package com.superbiz.agent.harness.contract;
import com.fasterxml.jackson.annotation.JsonProperty;
import java.util.List;
public record DiagnosisDraft(
@JsonProperty("conclusion") Conclusion conclusion,
@JsonProperty("analysis") List<AnalysisItem> analysis,
@JsonProperty("action_plan") List<ActionPlanItem> actionPlan,
@JsonProperty("recommendations") List<Recommendation> recommendations,
@JsonProperty("limitations") Limitations limitations) {
public DiagnosisDraft {
analysis = ContractCollections.immutable(analysis);
actionPlan = ContractCollections.immutable(actionPlan);
recommendations = ContractCollections.immutable(recommendations);
}
public record Conclusion(
@JsonProperty("text") String text,
@JsonProperty("based_on_analysis_ids") List<String> basedOnAnalysisIds) {
public Conclusion {
basedOnAnalysisIds = ContractCollections.immutable(basedOnAnalysisIds);
}
}
public record AnalysisItem(
@JsonProperty("analysis_id") String analysisId,
@JsonProperty("kind") AnalysisKind kind,
@JsonProperty("text") String text,
@JsonProperty("tool_call_ids") List<String> toolCallIds) {
public AnalysisItem {
toolCallIds = ContractCollections.immutable(toolCallIds);
}
}
public record ActionPlanItem(
@JsonProperty("action") String action,
@JsonProperty("based_on_analysis_ids") List<String> basedOnAnalysisIds,
@JsonProperty("requires_human_confirmation") boolean requiresHumanConfirmation) {
public ActionPlanItem {
basedOnAnalysisIds = ContractCollections.immutable(basedOnAnalysisIds);
}
}
public record Recommendation(
@JsonProperty("text") String text,
@JsonProperty("based_on_analysis_ids") List<String> basedOnAnalysisIds) {
public Recommendation {
basedOnAnalysisIds = ContractCollections.immutable(basedOnAnalysisIds);
}
}
public record Limitations(
@JsonProperty("scope") String scope,
@JsonProperty("missing_info") List<String> missingInfo) {
public Limitations {
missingInfo = ContractCollections.immutable(missingInfo);
}
}
}
@@ -0,0 +1,7 @@
package com.superbiz.agent.harness.contract;
public enum EvidenceStatus {
EVIDENCE_FOUND,
NO_EVIDENCE,
ERROR
}
@@ -0,0 +1,7 @@
package com.superbiz.agent.harness.contract;
public enum FallbackType {
EVIDENCE_VALIDATION_FAILED,
SEMANTIC_UNSUPPORTED,
SEMANTIC_UNAVAILABLE
}
@@ -0,0 +1,7 @@
package com.superbiz.agent.harness.contract;
public enum IntentType {
SYSTEM_CHAT,
KNOWLEDGE_QUERY,
DIAGNOSIS
}
@@ -0,0 +1,7 @@
package com.superbiz.agent.harness.contract;
public enum InvocationStatus {
PROJECTING,
READY,
ERROR
}
@@ -0,0 +1,25 @@
package com.superbiz.agent.harness.contract;
import com.fasterxml.jackson.annotation.JsonProperty;
import java.util.List;
public record KnowledgeAnswerDraft(
@JsonProperty("answer_items") List<AnswerItem> answerItems,
@JsonProperty("limitations") List<String> limitations) {
public KnowledgeAnswerDraft {
answerItems = ContractCollections.immutable(answerItems);
limitations = ContractCollections.immutable(limitations);
}
public record AnswerItem(
@JsonProperty("text") String text,
@JsonProperty("tool_call_id") String toolCallId,
@JsonProperty("document_ids") List<String> documentIds) {
public AnswerItem {
documentIds = ContractCollections.immutable(documentIds);
}
}
}
@@ -0,0 +1,27 @@
package com.superbiz.agent.harness.contract;
import com.fasterxml.jackson.annotation.JsonProperty;
import java.util.List;
public record PreviousTurn(
@JsonProperty("user_query") String userQuery,
@JsonProperty("published_conclusion") String publishedConclusion,
@JsonProperty("scope") String scope,
@JsonProperty("limitations") List<String> limitations,
@JsonProperty("source_documents") List<SourceDocument> sourceDocuments) {
public PreviousTurn {
limitations = ContractCollections.immutable(limitations);
sourceDocuments = ContractCollections.immutable(sourceDocuments);
}
public static PreviousTurn from(PublishedResult result) {
return new PreviousTurn(
result.userQuery(),
result.publishedConclusion(),
result.scope(),
result.limitations(),
result.sourceDocuments());
}
}
@@ -0,0 +1,18 @@
package com.superbiz.agent.harness.contract;
import com.fasterxml.jackson.annotation.JsonProperty;
import java.util.List;
public record PublishedResult(
@JsonProperty("user_query") String userQuery,
@JsonProperty("published_conclusion") String publishedConclusion,
@JsonProperty("scope") String scope,
@JsonProperty("limitations") List<String> limitations,
@JsonProperty("source_documents") List<SourceDocument> sourceDocuments) {
public PublishedResult {
limitations = ContractCollections.immutable(limitations);
sourceDocuments = ContractCollections.immutable(sourceDocuments);
}
}
@@ -0,0 +1,8 @@
package com.superbiz.agent.harness.contract;
public enum ReleaseOutcome {
SUCCESS,
FALLBACK,
FAILED,
CANCELLED
}
@@ -0,0 +1,26 @@
package com.superbiz.agent.harness.contract;
import com.fasterxml.jackson.annotation.JsonProperty;
import java.util.List;
public record SafeFallback(
@JsonProperty("type") FallbackType type,
@JsonProperty("conclusion") String conclusion,
@JsonProperty("message") String message,
@JsonProperty("verified_sources") List<VerifiedSource> verifiedSources,
@JsonProperty("limitations") List<String> limitations,
@JsonProperty("next_steps") List<String> nextSteps) {
public SafeFallback {
verifiedSources = ContractCollections.immutable(verifiedSources);
limitations = ContractCollections.immutable(limitations);
nextSteps = ContractCollections.immutable(nextSteps);
}
public record VerifiedSource(
@JsonProperty("source_type") String sourceType,
@JsonProperty("source") String source,
@JsonProperty("scope") String scope) {
}
}
@@ -0,0 +1,6 @@
package com.superbiz.agent.harness.contract;
public enum SemanticVerdict {
SUPPORTED,
UNSUPPORTED
}
@@ -0,0 +1,8 @@
package com.superbiz.agent.harness.contract;
import com.fasterxml.jackson.annotation.JsonProperty;
public record SourceDocument(
@JsonProperty("document_id") String documentId,
@JsonProperty("title") String title) {
}
@@ -0,0 +1,7 @@
package com.superbiz.agent.harness.contract;
public enum SseOutcome {
SUCCESS,
FALLBACK,
FAILED
}
+5 -5
View File
@@ -22,7 +22,7 @@ milvus:
password: "" password: ""
database: db_4a578da0f27ce9d database: db_4a578da0f27ce9d
timeout: 10000 timeout: 10000
token: d246a77f43a109685596e3c68ecfd359e1cd8b29d35c41d708160f02b97ae623d2ab392738df740d0c77b58ca1bbadfa7c412140 token: ${MILVUS_TOKEN}
secure: true secure: true
vector-dim: 1024 # BGE-M3 = 1024,换模型时同步改 vector-dim: 1024 # BGE-M3 = 1024,换模型时同步改
@@ -45,7 +45,7 @@ spring:
datasource: datasource:
url: jdbc:mysql://119.29.78.52:33306/superbiz_agent?useUnicode=true&serverTimezone=Asia/Shanghai&allowPublicKeyRetrieval=true url: jdbc:mysql://119.29.78.52:33306/superbiz_agent?useUnicode=true&serverTimezone=Asia/Shanghai&allowPublicKeyRetrieval=true
username: root username: root
password: '!Fucker123..' password: ${SUPERBIZ_MYSQL_PASSWORD}
driver-class-name: com.mysql.cj.jdbc.Driver driver-class-name: com.mysql.cj.jdbc.Driver
hikari: hikari:
maximum-pool-size: 5 maximum-pool-size: 5
@@ -83,7 +83,7 @@ spring:
redis: redis:
host: 119.29.78.52 host: 119.29.78.52
port: 33308 port: 33308
password: '!Fucker123..' password: ${SUPERBIZ_REDIS_PASSWORD}
database: 0 database: 0
timeout: 3000 timeout: 3000
lettuce: lettuce:
@@ -119,7 +119,7 @@ spring:
# --- Chat: DeepSeek (原生) --- # --- Chat: DeepSeek (原生) ---
deepseek: deepseek:
api-key: sk-1f44696abe644bd684f09cc43f12c557 api-key: ${DEEPSEEK_API_KEY}
base-url: https://api.deepseek.com base-url: https://api.deepseek.com
chat: chat:
options: options:
@@ -136,7 +136,7 @@ spring:
# --- Embedding: SiliconFlow BGE-M3 --- # --- Embedding: SiliconFlow BGE-M3 ---
siliconflow: siliconflow:
api-key: sk-rlxqcnlohjqwkzoffollthmzzfiohngdrabrmmqhcgtewnzx api-key: ${SILICONFLOW_API_KEY}
base-url: https://api.siliconflow.cn base-url: https://api.siliconflow.cn
embedding: embedding:
model: BAAI/bge-m3 model: BAAI/bge-m3
@@ -0,0 +1,96 @@
package com.superbiz.agent.harness.contract;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import org.junit.jupiter.api.Test;
import java.util.List;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
class HarnessContractTest {
private final ObjectMapper objectMapper = new ObjectMapper();
@Test
void diagnosisDraftUsesStableBoundedShape() throws Exception {
DiagnosisDraft draft = new DiagnosisDraft(
new DiagnosisDraft.Conclusion("Pool exhausted", List.of("a1")),
List.of(new DiagnosisDraft.AnalysisItem(
"a1",
AnalysisKind.NORMAL,
"active=50 max=50",
List.of("call-1"))),
List.of(new DiagnosisDraft.ActionPlanItem(
"Inspect long transactions",
List.of("a1"),
false)),
List.of(new DiagnosisDraft.Recommendation(
"Add pool wait alerts",
List.of("a1"))),
new DiagnosisDraft.Limitations("order-service, last 30 minutes", List.of("No slow SQL data")));
JsonNode json = objectMapper.valueToTree(draft);
assertEquals("call-1", json.path("analysis").get(0).path("tool_call_ids").get(0).asText());
assertEquals("NORMAL", json.path("analysis").get(0).path("kind").asText());
assertTrue(json.has("action_plan"));
assertFalse(json.toString().contains("raw_response"));
assertFalse(json.toString().contains("thought"));
}
@Test
void evidenceStatusRemainsSeparateFromInvocationStatus() {
assertTrue(AnalysisKind.NORMAL.accepts(EvidenceStatus.EVIDENCE_FOUND));
assertFalse(AnalysisKind.NORMAL.accepts(EvidenceStatus.NO_EVIDENCE));
assertTrue(AnalysisKind.NEGATIVE_OBSERVATION.accepts(EvidenceStatus.NO_EVIDENCE));
assertFalse(AnalysisKind.NEGATIVE_OBSERVATION.accepts(EvidenceStatus.ERROR));
assertEquals(InvocationStatus.READY, InvocationStatus.valueOf("READY"));
}
@Test
void knowledgeAnswerBindsExactToolAndDocuments() throws Exception {
KnowledgeAnswerDraft draft = new KnowledgeAnswerDraft(
List.of(new KnowledgeAnswerDraft.AnswerItem(
"Check the timeout code first.",
"call-rag-1",
List.of("payment-timeout-guide"))),
List.of("Runtime state was not queried"));
JsonNode json = objectMapper.valueToTree(draft);
assertEquals("call-rag-1", json.path("answer_items").get(0).path("tool_call_id").asText());
assertEquals("payment-timeout-guide",
json.path("answer_items").get(0).path("document_ids").get(0).asText());
}
@Test
void fallbackAndPreviousTurnExcludeInternalEvidence() throws Exception {
SafeFallback fallback = new SafeFallback(
FallbackType.SEMANTIC_UNSUPPORTED,
null,
"Current evidence is insufficient",
List.of(new SafeFallback.VerifiedSource("LOG", "APPLICATION", "last 30 minutes")),
List.of("Semantic validation did not pass"),
List.of("Collect the missing data and retry"));
PublishedResult result = new PublishedResult(
"Check payment timeout",
"Connection pool exhaustion",
"order-service, last 30 minutes",
List.of("No slow SQL data"),
List.of(new SourceDocument("payment-timeout-guide", "Payment timeout guide")));
JsonNode fallbackJson = objectMapper.valueToTree(fallback);
JsonNode previousJson = objectMapper.valueToTree(PreviousTurn.from(result));
assertNull(fallback.conclusion());
assertEquals("SEMANTIC_UNSUPPORTED", fallbackJson.path("type").asText());
assertEquals("payment-timeout-guide",
previousJson.path("source_documents").get(0).path("document_id").asText());
assertFalse(previousJson.toString().contains("tool_call_id"));
assertFalse(previousJson.toString().contains("raw_response"));
}
}
@@ -4,6 +4,7 @@ import io.milvus.client.MilvusServiceClient;
import io.milvus.param.ConnectParam; import io.milvus.param.ConnectParam;
import io.milvus.param.R; import io.milvus.param.R;
import io.milvus.param.collection.HasCollectionParam; import io.milvus.param.collection.HasCollectionParam;
import org.junit.jupiter.api.Assumptions;
import org.junit.jupiter.api.Test; import org.junit.jupiter.api.Test;
/** /**
@@ -13,9 +14,12 @@ public class SimpleMilvusTest {
@Test @Test
public void testConnection() { public void testConnection() {
String host = "in03-4a578da0f27ce9d.serverless.aws-eu-central-1.cloud.zilliz.com"; String host = System.getenv().getOrDefault(
int port = 443; "MILVUS_HOST",
String token = "d246a77f43a109685596e3c68ecfd359e1cd8b29d35c41d708160f02b97ae623d2ab392738df740d0c77b58ca1bbadfa7c412140"; "in03-4a578da0f27ce9d.serverless.aws-eu-central-1.cloud.zilliz.com");
int port = Integer.parseInt(System.getenv().getOrDefault("MILVUS_PORT", "443"));
String token = System.getenv("MILVUS_TOKEN");
Assumptions.assumeTrue(token != null && !token.isBlank(), "MILVUS_TOKEN is required");
System.out.println("尝试连接 Milvus..."); System.out.println("尝试连接 Milvus...");
System.out.println("Host: " + host); System.out.println("Host: " + host);
@@ -4,6 +4,7 @@ import com.superbiz.agent.domain.model.SessionContext;
import com.superbiz.agent.domain.model.ToolCall; import com.superbiz.agent.domain.model.ToolCall;
import org.junit.jupiter.api.BeforeEach; import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test; import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.condition.EnabledIfEnvironmentVariable;
import org.springframework.beans.factory.annotation.Autowired; import org.springframework.beans.factory.annotation.Autowired;
import org.springframework.boot.test.context.SpringBootTest; import org.springframework.boot.test.context.SpringBootTest;
import org.springframework.test.context.TestPropertySource; import org.springframework.test.context.TestPropertySource;
@@ -20,10 +21,11 @@ import static org.junit.jupiter.api.Assertions.*;
* RedisSessionManager 单元测试 * RedisSessionManager 单元测试
*/ */
@SpringBootTest(webEnvironment = SpringBootTest.WebEnvironment.NONE) @SpringBootTest(webEnvironment = SpringBootTest.WebEnvironment.NONE)
@EnabledIfEnvironmentVariable(named = "SUPERBIZ_REDIS_PASSWORD", matches = ".+")
@TestPropertySource(properties = { @TestPropertySource(properties = {
"spring.redis.host=119.29.78.52", "spring.data.redis.host=${SUPERBIZ_REDIS_HOST:119.29.78.52}",
"spring.redis.port=6379", "spring.data.redis.port=${SUPERBIZ_REDIS_PORT:33308}",
"spring.redis.password=!Fucker123.." "spring.data.redis.password=${SUPERBIZ_REDIS_PASSWORD}"
}) })
class RedisSessionManagerTest { class RedisSessionManagerTest {