refactor(harness): freeze single-agent contracts
This commit is contained in:
@@ -148,7 +148,7 @@ openspec/changes/phase-1-infrastructure/
|
||||
|
||||
## 敏感信息(已编辑)
|
||||
|
||||
- MySQL 密码:已配置在 application.yml(`!Fucker123..`)
|
||||
- MySQL 密码:已从仓库移除,使用环境变量注入
|
||||
- Redis:无密码
|
||||
|
||||
---
|
||||
|
||||
@@ -97,7 +97,7 @@ spring:
|
||||
redis:
|
||||
host: 119.29.78.52
|
||||
port: 6379
|
||||
password: '!Fucker123..'
|
||||
password: ${SUPERBIZ_REDIS_PASSWORD}
|
||||
database: 0
|
||||
timeout: 3000
|
||||
```
|
||||
|
||||
@@ -140,7 +140,7 @@ Error Code: 1049
|
||||
datasource:
|
||||
url: jdbc:mysql://119.29.78.52:33306/superbiz_agent?...
|
||||
username: root
|
||||
password: '!Fucker123..'
|
||||
password: ${SUPERBIZ_MYSQL_PASSWORD}
|
||||
```
|
||||
|
||||
**Redis 配置**:
|
||||
@@ -149,7 +149,7 @@ data:
|
||||
redis:
|
||||
host: 119.29.78.52
|
||||
port: 6379
|
||||
password: '!Fucker123..'
|
||||
password: ${SUPERBIZ_REDIS_PASSWORD}
|
||||
```
|
||||
|
||||
**Flyway 配置**:
|
||||
|
||||
@@ -44,6 +44,13 @@ build/
|
||||
app.log
|
||||
logs/
|
||||
|
||||
### Local Secrets ###
|
||||
.env
|
||||
.env.*
|
||||
!.env.example
|
||||
application-local.yml
|
||||
application-*.local.yml
|
||||
|
||||
### Upload Files ###
|
||||
uploads/
|
||||
|
||||
|
||||
@@ -155,6 +155,26 @@
|
||||
- 使用场景:Executor 按 skill workflow 调用 evidence tools 收集事实,`tool_invocation` 记录这些事实证据。
|
||||
- 边界:最终诊断结论必须被 evidence tools 支撑,不能仅由 skill 正文支撑。
|
||||
|
||||
### Diagnosis Harness
|
||||
- 定义:围绕 Diagnosis Agent 提供确定性运行控制的边界,负责 Run、预算、取消、重试装配、Tool 调用记录、证据验真和最终释放,不承担业务诊断推理。
|
||||
- 边界:Harness 不是工作流引擎,不实现 Planner/Executor/Composer 节点或自行编写 ReAct 循环。
|
||||
|
||||
### EvidenceGuard
|
||||
- 定义:Harness 内部的确定性证据验真能力,校验 Draft 引用、当前 Run 所有权、Tool 调用状态和有界 Agent 投影。
|
||||
- 边界:EvidenceGuard 不调用 LLM,也不判断证据是否足以推出业务结论。
|
||||
|
||||
### SemanticGuard
|
||||
- 定义:使用隔离上下文对完整诊断 Draft 与已验真证据做报告级语义审查的单轮 Agent。
|
||||
- 边界:无工具、无记忆、无 ReAct 循环,不访问 Redis,不生成或改写用户报告。
|
||||
|
||||
### Invocation Status
|
||||
- 定义:Tool 调用及结果投影的生命周期状态,固定为 `PROJECTING`、`READY`、`ERROR`。
|
||||
- 边界:它只说明调用记录是否完成,不说明结果是否包含证据。
|
||||
|
||||
### Evidence Status
|
||||
- 定义:证据 Tool 的结果语义,固定为 `EVIDENCE_FOUND`、`NO_EVIDENCE`、`ERROR`。
|
||||
- 边界:`NO_EVIDENCE` 只表示当前查询范围内没有匹配结果,不能解释为问题不存在、根因被排除或系统健康。
|
||||
|
||||
### Verifier Skill Isolation
|
||||
- 定义:Chat Verifier 与 skill 系统隔离,只校验 Executor 答案和 `tool_trace_summary`。
|
||||
- 使用场景:防止 Verifier 把 playbook 指令当作事实证据;Verifier 只判断已有证据是否支持结论。
|
||||
|
||||
@@ -4,6 +4,7 @@
|
||||
|
||||
| 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|
||||
|---|---|---|---|---|---|---|
|
||||
| 2026-07-21 | single-react-design-freeze | 冻结单体 Diagnosis Agent、Harness、Guard、工具证据与阶段门禁契约。 | Chat/Harness/Agent contract | ISS-014, single ReactAgent, Harness, EvidenceGuard, SemanticGuard, tool_call_id, evidence_status | openspec/changes/archive/2026-07-21-single-react-design-freeze | archived |
|
||||
| 2026-07-10 | session-run-trace-isolation | 拆分会话态和运行态,引入 runId 隔离 Trace、Feedback、AIOps 和 demo 链路。 | Trace/session/run isolation | chat_session, diagnosis_run, runId, trace exact run, feedback fallback, AIOps SSE metadata, baseline drift | openspec/changes/archive/2026-07-10-session-run-trace-isolation | archived |
|
||||
| 2026-07-09 | interview-demo-quality-audit | 增加面试演示前置质量审计,覆盖 prompt、Gatekeeper 和评测基线。 | Agent eval/demo/Prompt audit | interview demo preflight, prompt_audit, gatekeeper rules, diagnosis baseline, 12 fixtures | openspec/changes/archive/2026-07-09-interview-demo-quality-audit | archived |
|
||||
| 2026-07-08 | executor-composer-final-answer | 引入 Composer 生成最终回答,只使用 Verifier 允许的结论材料。 | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
|
||||
|
||||
@@ -0,0 +1,36 @@
|
||||
# Acceptance: single-react-design-freeze
|
||||
|
||||
## 实现结果
|
||||
|
||||
- 新增类型化 Harness contracts 和序列化契约测试。
|
||||
- ISS-014 已调整为 11 个串行 changes,并补齐 Tool ID、双状态、previous turn 和安全切换决策。
|
||||
- Python 查询脚本和 Spring 主配置改为环境变量 Secret。
|
||||
- 集成测试和历史 handoff 中的旧 Secret 副本已删除或脱敏。
|
||||
- devflow glossary 新增 Harness、EvidenceGuard、SemanticGuard、Invocation Status 和 Evidence Status。
|
||||
|
||||
## 静态验证
|
||||
|
||||
- OpenSpec strict validation:通过。
|
||||
- Secret 扫描:通过,已知真实 Secret 模式零匹配。
|
||||
- Diff whitespace 检查:通过。
|
||||
|
||||
## 脚本验证
|
||||
|
||||
- 4 个 contract tests:通过。
|
||||
- ToolInvocationRecorder、ExecutorGatekeeperService、ChatController focused baseline:通过。
|
||||
- 查询脚本缺失密码环境变量时 fail fast:通过。
|
||||
|
||||
## 浏览器/人工验证
|
||||
|
||||
- 不适用;阶段 0 未切换任何公开 UI/API。
|
||||
|
||||
## 未验证与后续门禁
|
||||
|
||||
- Provider 侧旧凭据轮换待凭据所有者完成。
|
||||
- Live E2E 留到阶段 7。
|
||||
- 阶段 1 只有在本 change Archive 和 Git commit 后才能开始。
|
||||
|
||||
## 状态
|
||||
|
||||
- Stage acceptance: accepted
|
||||
- OpenSpec archive: authorized by standing user instruction
|
||||
@@ -0,0 +1,31 @@
|
||||
# Brief: single-react-design-freeze
|
||||
|
||||
## 背景
|
||||
|
||||
ISS-014 将 Chat 诊断从多 Agent、Hook、ThreadLocal 和外层重试编排收敛为单体 Diagnosis ReAct Agent、确定性 Harness 和隔离 SemanticGuard。阶段 0 先冻结后续实施共同依赖的契约和安全边界。
|
||||
|
||||
## 目标
|
||||
|
||||
- 提供可复用的 Diagnosis Draft、Knowledge Answer、Fallback、published result 和 previous turn 类型。
|
||||
- 固定 framework Tool Call ID、调用生命周期和证据结果双状态。
|
||||
- 固定取消、重试、no-evidence、阶段门禁和安全凭据策略。
|
||||
- 保持当前公开 Chat 运行行为不变。
|
||||
|
||||
## 范围
|
||||
|
||||
- `com.superbiz.agent.harness.contract` 类型层和 focused tests。
|
||||
- ISS-014 的 11 个串行 OpenSpec changes 台账。
|
||||
- 主配置、查询脚本和集成测试中的明文 Secret 清理。
|
||||
- OpenSpec、devflow glossary 和阶段验收基线。
|
||||
|
||||
## 非目标
|
||||
|
||||
- 不接入新 Harness、Agent 或 Guards。
|
||||
- 不切换 `/api/chat`,不删除旧运行链。
|
||||
- 不执行 live E2E。
|
||||
|
||||
## 元数据
|
||||
|
||||
- Scale: complex
|
||||
- Parent issue: `ISS-014`
|
||||
- OpenSpec: `openspec/changes/single-react-design-freeze`
|
||||
@@ -0,0 +1,20 @@
|
||||
# Decisions: single-react-design-freeze
|
||||
|
||||
## 已确认决策
|
||||
|
||||
- ISS-014 保留为总 Issue,实施拆成 11 个独立、串行 sm-flow changes。
|
||||
- 每个 change 必须完成 Apply、阶段验收、OpenSpec Archive 和 Git commit 后才能进入下一项。
|
||||
- Apply、Archive 和阶段 Git commit 已获得用户对整个目标的持续授权;仅方向性决策需要暂停。
|
||||
- `tool_call_id` 使用框架协议 ID,Harness 不生成第二套 ID。
|
||||
- `status=PROJECTING/READY/ERROR` 表示 invocation lifecycle;`evidence_status=EVIDENCE_FOUND/NO_EVIDENCE/ERROR` 表示结果语义。
|
||||
- `NO_EVIDENCE` 只能支持限定范围的 `NEGATIVE_OBSERVATION`。
|
||||
- KNOWLEDGE_QUERY 使用 answer items 绑定 `tool_call_id + document_id`,首版不进入 SemanticGuard。
|
||||
- `diagnosis_run` 是 previous turn 的持久化真理源;只有最近的 `DIAGNOSIS + SUCCESS` 安全发布结果可用。
|
||||
- 取消采用分层、可观测语义;同步模型调用不承诺无法证明的立即硬取消。
|
||||
- 底层重试压为一次 attempt,Router/SemanticGuard 的允许重试只由 Harness 执行。
|
||||
- 阶段 4 和 6A 不接管公开入口,阶段 6B 才执行原子 SSE 切换。
|
||||
|
||||
## 权衡
|
||||
|
||||
- contract types 提前落地会增加少量文件,但后续 output type、validator、持久化和 SSE 可以复用,避免多阶段字符串协议漂移。
|
||||
- Provider 侧凭据轮换是外部前置,阶段 0 只证明工作树已清理,不伪造外部完成状态。
|
||||
@@ -0,0 +1,25 @@
|
||||
# Evidence: single-react-design-freeze
|
||||
|
||||
## 代码与框架证据
|
||||
|
||||
- `ChatController` 当前直接管理会话、模型、工具、执行和伪流式 SSE,证明公开切换必须延后到阶段 6B。
|
||||
- `ChatService` 当前构建 Planner/Executor/Verifier/Composer 并依赖 ThreadLocal,证明阶段 0 只能创建无运行依赖的 contract types。
|
||||
- Spring AI Alibaba `Builder` 提供 ToolInterceptor、output type、tool timeout 和调用限制扩展点;`ReactAgent` 提供 interrupt。
|
||||
- Spring AI 1.1.7 `SpringAiRetryProperties` 构造器默认 `maxAttempts=10`,与 Harness 单一重试所有权冲突,阶段 2 必须压为一次底层 attempt。
|
||||
- `DiagnosisRun` 当前只有通用 status 和文本 answer,不足以确定性恢复安全 previous turn,因此确认后续保存 `intent/release_outcome/published_result`。
|
||||
- 历史 ISS-009 和 verifier evidence 归档证明 `NO_EVIDENCE` 只能表达限定范围内无匹配结果。
|
||||
|
||||
## 验证证据
|
||||
|
||||
- `mvn -q '-Dtest=HarnessContractTest' test`:通过。
|
||||
- `mvn -q '-Dtest=ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,ChatControllerTest' test`:通过。
|
||||
- `mvn -q '-Dtest=HarnessContractTest,ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,ChatControllerTest' test`:通过。
|
||||
- `openspec validate single-react-design-freeze --strict`:通过。
|
||||
- `git diff --check`:通过,仅有 Windows line-ending warning。
|
||||
- 已知旧 Secret 模式扫描:零匹配。
|
||||
- 缺少 `SUPERBIZ_MYSQL_PASSWORD` 时运行查询脚本:在连接前 fail fast,符合预期。
|
||||
|
||||
## 未验证
|
||||
|
||||
- 未运行模型、Redis、MySQL 或 Milvus live E2E;按阶段门禁留到阶段 7。
|
||||
- Provider 侧旧凭据是否已经轮换无法由仓库证明,必须由凭据所有者完成,并在阶段 3C/7 前核验。
|
||||
@@ -31,7 +31,7 @@
|
||||
**配置信息**(已完成):
|
||||
- **MySQL**: 119.29.78.52:33306/superbiz_agent
|
||||
- 用户: root
|
||||
- 密码: !Fucker123..
|
||||
- 密码: 已从仓库移除,使用环境变量注入
|
||||
- driver: com.mysql.cj.jdbc.Driver
|
||||
- URL参数: useUnicode=true&serverTimezone=Asia/Shanghai&allowPublicKeyRetrieval=true
|
||||
- **Redis**: 119.29.78.52:6379
|
||||
@@ -165,7 +165,7 @@ openspec/changes/phase-1-infrastructure/
|
||||
|
||||
## 敏感信息(已编辑)
|
||||
|
||||
- MySQL 密码:已配置在 application.yml(`!Fucker123..`)
|
||||
- MySQL 密码:已从仓库移除,使用环境变量注入
|
||||
- Redis:无密码
|
||||
|
||||
---
|
||||
|
||||
@@ -0,0 +1,56 @@
|
||||
结合你在前几轮对话中梳理出的“过度设计”痛点,既然你已经决定回归单 Agent(ReAct)+ 强 Harness 架构,接下来你需要做一次彻底的“架构物理重构”。
|
||||
以下是为你量身定制的5步落地行动指南,按优先级从高到低执行:
|
||||
第一步:链路合并,砍掉“伪 Graph”节点
|
||||
目标:把被拆散的推理逻辑还给单一的 ReAct 循环。
|
||||
|
||||
删除节点:直接在 Graph/状态机中抹掉 Planner、Executor、Composer 以及串联它们的边。
|
||||
合并为单一 Agent:创建一个 DiagnosisAgent。在它的 System Prompt 中明确:“你需要自行规划排查路径,调用工具获取证据,并在证据充分后输出最终的结构化诊断报告。”
|
||||
保留机制:该 Agent 内部运行一个带最大步数限制的 While 循环(如 max_iterations=10),防止死循环。
|
||||
|
||||
第二步:实施“上下文卸载”与“工具清洗”(解决7次调用膨胀问题)
|
||||
目标:确保单 Agent 在多轮工具调用后,上下文依然干净,从源头切断幻觉。
|
||||
|
||||
工具输出清洗(RTK 机制):改造你的所有工具(如 queryLogs, queryDB)。在工具的 Callback 层加代码,将原始的大段返回结果(如 1000 行日志)强制过滤、聚合,只把最核心的 5-10 行 ERROR 或统计指标返回给 Agent。
|
||||
上下文卸载:如果某些原始数据必须保留,在工具返回时,将全量数据写入本地文件(如 refs/log_001.md),上下文里只注入一行:[发现5个504错误,详见 refs/log_001.md]。
|
||||
|
||||
第三步:下沉确定性逻辑,用代码替代 LLM 节点
|
||||
目标:将你之前用 Gatekeeper 和 VerifiedInput 做的事,降级为零 LLM 调用的代码拦截器。
|
||||
|
||||
前置拦截(Pre-Tool Hook):Agent 发起工具调用时,Harness 代码用 JSON Schema 校验参数格式。不合法直接报错打回,不执行工具。
|
||||
后置断言(Post-Tool Hook):Agent 输出最终诊断报告时,代码层强制校验报告中引用的 evidence_id 和 raw_path 是否真实存在于历史记录或文件系统中。不合法直接拒绝输出,发回重试。
|
||||
|
||||
第四步:锁死输出契约
|
||||
目标:防止 Executor 过度输出和发散。
|
||||
|
||||
在 System Prompt 中强制规定 Agent 的中间思考步和最终输出步必须符合严格的 JSON 结构。
|
||||
例如,中间步必须是 {"thought": "<不超过50字>", "action": "queryLogs", "parameters": {...}}。Harness 代码检查字数,超长直接打回。
|
||||
|
||||
第五步:剥离异步验证(保留你最初的“防幻觉”初衷)
|
||||
目标:在不增加主链路复杂度的前提下,保留交叉验证能力。
|
||||
|
||||
主 ReAct Agent 输出报告后,不要在主 Graph 里串行接一个 Verifier 节点。
|
||||
改为异步触发一个轻量级 LLM(或小模型),只传入“压缩后的证据摘要 + 草稿结论”。让它判断时间线与逻辑是否一致。如果不一致,在最终输出上加“低置信度警告”;如果一致,直接放行。
|
||||
|
||||
总结:你的重构后架构全景图
|
||||
重构后,你的代码结构应该极其清爽,大致如下:
|
||||
[用户输入]
|
||||
│
|
||||
▼
|
||||
[Diagnosis ReAct Agent] (唯一的 LLM 推理节点,自带规划、执行、总结)
|
||||
│
|
||||
├── Tool: queryLogs
|
||||
│ └── [Harness 代码]: 过滤 INFO,提取 ERROR,写入 refs,返回摘要
|
||||
├── Tool: queryMetrics
|
||||
│ └── [Harness 代码]: 聚合统计值,返回 3 行核心指标
|
||||
│
|
||||
▼ (Agent 输出最终 JSON 报告)
|
||||
[Output Schema Linter] (纯代码层,0 LLM)
|
||||
│
|
||||
├─ 校验失败 ──> 返回错误给 [Agent] 重新生成
|
||||
│
|
||||
▼ (校验通过)
|
||||
[Async Verifier] (异步轻量 LLM,只做时间线/逻辑一致性校验)
|
||||
│
|
||||
▼
|
||||
[最终输出 / 带警告输出]
|
||||
现在你应该做的第一件事:打开你的代码,把主链路上除 Diagnosis Agent 以外的所有 LLM 编排节点全部注释掉,然后按照上面的结构,给工具加上 Pre-Tool 和 Post-Tool 的代码拦截器。把精力从“画 Graph”转移到“写工具清洗代码”上。
|
||||
@@ -1,6 +1,6 @@
|
||||
# MVP Issues 索引
|
||||
|
||||
**更新日期**:2026-07-10
|
||||
**更新日期**:2026-07-20
|
||||
**状态**:按活跃问题、设计笔记、RAG 问题集和已归档问题整理
|
||||
|
||||
## 目录约定
|
||||
@@ -18,6 +18,9 @@
|
||||
|---|---|---|---|---|
|
||||
| ISS-003 | MVP 设计与实现 Review 收敛 | 高 | 待规划 | [active/ISS-003-mvp-design-implementation-review.md](active/ISS-003-mvp-design-implementation-review.md) |
|
||||
| ISS-004 | Executor 域级检索水位控制 | 低 | 待规划 | [active/ISS-004-executor-domain-hard-limit.md](active/ISS-004-executor-domain-hard-limit.md) |
|
||||
| ISS-012 | Executor Token 预算与上下文膨胀 | 高 | 待规划 | [active/ISS-012-executor-token-budget-and-context-growth.md](active/ISS-012-executor-token-budget-and-context-growth.md) |
|
||||
| ISS-013 | Chat 入口解耦与真正 SSE 收敛 | 高 | 待规划 | [active/ISS-013-chat-entry-decoupling-and-sse.md](active/ISS-013-chat-entry-decoupling-and-sse.md) |
|
||||
| ISS-014 | 单体 ReAct Agent、Harness 与 ACI 工具瘦身 | 高 | 待阶段 0 冻结 | [active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md](active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md) |
|
||||
| executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [active/executor-evidence-attribution-hallucination.md](active/executor-evidence-attribution-hallucination.md) |
|
||||
| rag-refactor-plan | RAG 检索重构计划 | 高 | 待规划 | [active/rag-refactor-plan.md](active/rag-refactor-plan.md) |
|
||||
|
||||
|
||||
@@ -0,0 +1,119 @@
|
||||
# ISS-012 Executor Token 预算与上下文膨胀
|
||||
|
||||
**状态**:待规划
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-20
|
||||
**关联**:ISS-002、ISS-004、ISS-011
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
ISS-011 完成 StateGraph 切换后,复杂 Chat 链路已经具备显式 Node、Gatekeeper、Verifier、Composer 和 Run 级 Trace。但最终 live E2E 暴露出 Executor 的 Token 成本和上下文增长问题:一次只返回安全 Fallback 的请求消耗了超过 11 万 Token。
|
||||
|
||||
## 现象
|
||||
|
||||
最终验收 Run:
|
||||
|
||||
- `runId`:`run-808ac38f-3ad0-4462-a6d0-ed50d8686473`
|
||||
- 总耗时:`75964ms`
|
||||
- 最终答案:109 字符,Run 为 `CHAT/SUCCESS`,`degraded=true`
|
||||
- AgentStep:8 条,其中 Planner 1 次、Executor 7 次
|
||||
- ToolInvocation:12 次
|
||||
- Run `total_token_count`:`111802`
|
||||
|
||||
Executor 每次模型调用的 Token 逐步上升:
|
||||
|
||||
```text
|
||||
6396 -> 11265 -> 12786 -> 16774 -> 17965 -> 19340 -> 24236
|
||||
```
|
||||
|
||||
工具调用包括 4 次 `lookup_knowledge`、6 次 `query_logs`、1 次 `query_metrics` 和 1 次 `get_available_log_topics`。最终因模型输出缺少 `source_invocation_id`,Gatekeeper 将结果降为 `LOW_CONFID` 并进入安全 Fallback。
|
||||
|
||||
## 已确认事实
|
||||
|
||||
1. 数据库中 8 条 `agent_step` 均为不同记录,不存在重复插入;`111802` 等于各步骤 `token_count` 的求和。
|
||||
2. `TokenTrackingChatModel` 当前只保存供应商返回的 `usage.totalTokens`,没有拆分输入、输出、缓存和推理 Token。
|
||||
3. `ChatService.backfillRunMetrics` 直接累加每条 AgentStep 的 `token_count`。
|
||||
4. `AgentLoggingHook` 只把模型输入截断为 500 字符写入审计表,无法从当前 Trace 还原模型实际发送的完整 Prompt。
|
||||
5. Executor 当前没有独立的模型调用次数、工具调用次数或 Token 预算;Graph recursion limit 不能限制 ReactAgent 内部工具循环。
|
||||
|
||||
## 初步根因假设
|
||||
|
||||
- 每次 Executor 模型调用都会重新携带 Planner 结果、历史消息和之前的工具返回,导致输入上下文随工具循环增长。
|
||||
- 工具返回内容包含较多日志、知识库结果和检索明细,完整结果被反复带入后续模型请求。
|
||||
- 工具结果没有稳定返回 `source_invocation_id`,模型无法可靠生成精确证据引用,导致高成本检索后仍然进入 Fallback。
|
||||
- 当前只能看到 `totalTokens`,尚未确认供应商 usage 中 input/output/cached/reasoning 的精确占比。
|
||||
|
||||
## 影响
|
||||
|
||||
- 单次诊断成本和延迟不可控,复杂问题可能继续超过模型上下文窗口。
|
||||
- Token 消耗与最终答案质量不匹配,出现“高成本检索 + 安全降级”的低收益路径。
|
||||
- 缺少 Token 分项指标,无法建立成本预算、P95 延迟和 degraded rate 门禁。
|
||||
- Executor 可能重复查询相同或相近的知识域、日志主题和指标。
|
||||
|
||||
## 目标
|
||||
|
||||
1. 建立按 Run/AgentStep 的 input、output、cached、reasoning Token 可观测性。
|
||||
2. 为 Executor 增加硬性模型轮数、工具调用和 Token 预算。
|
||||
3. 将完整工具结果留在 Run Trace/数据库中,模型上下文只接收有界证据投影。
|
||||
4. 让工具结果直接携带可引用的 `source_invocation_id` 和紧凑 `evidence_refs`。
|
||||
5. 在预算耗尽时安全结束并明确记录原因,不绕过 Gatekeeper、Verifier 或 Run Trace。
|
||||
|
||||
## 建议方案
|
||||
|
||||
### 1. Token 统计拆分
|
||||
|
||||
- 从 ChatModel usage 中记录 `input_tokens`、`output_tokens`、`cached_tokens`、`reasoning_tokens`(供应商提供时)。
|
||||
- 保留 `total_token_count` 作为汇总字段,但明确其计算口径。
|
||||
- 在 `orchestration_trace` 中记录每个 Agent 的累计 Token 和预算命中情况。
|
||||
|
||||
### 2. Executor 硬预算
|
||||
|
||||
初版建议从以下上限开始,并通过固定 E2E 调整:
|
||||
|
||||
- Executor 模型调用最多 4 次。
|
||||
- 工具调用最多 8 次。
|
||||
- `lookup_knowledge` 最多 2 次。
|
||||
- `query_logs` 默认最多返回 5 条日志,并限制单次输出长度。
|
||||
- 达到预算后停止扩展检索,基于已验真证据输出,或进入带原因的安全 Fallback。
|
||||
|
||||
### 3. 有界证据上下文
|
||||
|
||||
- 工具完整原始结果继续写入 `tool_invocation`,不直接作为下一轮完整上下文。
|
||||
- 返回模型的工具视图只保留 invocation ID、工具名、查询条件、有限 evidence refs、excerpt 和 no-evidence 状态。
|
||||
- 同一工具数组项禁止重复绑定;相同知识域和日志主题不重复查询。
|
||||
|
||||
### 4. 证据引用闭环
|
||||
|
||||
- 每次 evidence tool 返回结果时直接包含 `source_invocation_id`。
|
||||
- Executor 输出必须引用该 ID;Gatekeeper 不再依赖事后猜测或唯一候选补全。
|
||||
- 由于引用失败进入 Fallback 时,Trace 必须记录具体缺失字段和预算消耗。
|
||||
|
||||
## 验收标准
|
||||
|
||||
- [ ] 每个 AgentStep 可查看 input/output/total Token,供应商支持时可查看 cached/reasoning Token。
|
||||
- [ ] 固定 `payment-timeout` E2E 的 Token 上限、工具调用上限和最大延迟已定义并通过回归。
|
||||
- [ ] 连续至少 10 次相同 fixture 运行,Token 和延迟 P95 不超过定义的预算。
|
||||
- [ ] Executor 预算耗尽时只走安全 Fallback,不绕过 Gatekeeper、Verifier 或 Trace 持久化。
|
||||
- [ ] 工具返回包含真实 `source_invocation_id`;正常证据链不再因缺少该字段而无谓降级。
|
||||
- [ ] 12 个 diagnosis eval fixture、Graph workflow/node contract、Trace ownership 回归全部通过。
|
||||
- [ ] E2E 日志和数据库能按 exact `sessionId + runId` 对齐 Token、工具调用、Fallback 原因和最终状态。
|
||||
|
||||
## 非目标
|
||||
|
||||
- 不删除 Gatekeeper、Verified Input 或 Verifier。
|
||||
- 不以降低模型 `maxTokens` 代替上下文治理。
|
||||
- 不恢复 Sequential/StateGraph 双轨或旧兼容协议。
|
||||
- 不在本 Issue 中物理删除数据库中的历史 `diagnosis_session` 表。
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `src/main/java/com/superbiz/agent/hook/TokenTrackingChatModel.java`
|
||||
- `src/main/java/com/superbiz/agent/hook/AgentLoggingHook.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- `src/main/java/com/superbiz/agent/graph/diagnosis/ReactAgentDiagnosisInvoker.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
|
||||
- `src/main/resources/prompts/chat-executor-prompt.md`
|
||||
- `mvp/architecture/stategraph-runtime-architecture.md`
|
||||
- `devflow/projects/2026-07-17-chat-diagnosis-stategraph-cleanup-docs/evidence.md`
|
||||
@@ -0,0 +1,81 @@
|
||||
# ISS-013 Chat 入口解耦与真正 SSE 收敛
|
||||
|
||||
**状态**:待规划
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-20
|
||||
**关联**:ISS-011、ISS-012
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
当前 Chat 入口同时提供 `/api/chat` 和 `/api/chat_stream`。`ChatController` 不仅处理 HTTP/SSE 协议,还直接承担会话创建、历史读取与回写、模型和工具获取、执行策略调用以及异常响应组装,入口职责已经明显超出协议适配层。
|
||||
|
||||
现有 `/api/chat_stream` 会等待完整答案生成后再按固定长度切片发送,并不是真正的流式生成。同步和伪流式入口还复制了大部分业务流程,增加了维护成本和行为不一致风险。
|
||||
|
||||
## 已确认问题
|
||||
|
||||
1. Chat 的会话、历史、模型、工具、执行和结果回写功能耦合在 `ChatController` 中,HTTP 层与应用用例边界不清晰。
|
||||
2. `/api/chat` 与 `/api/chat_stream` 重复编排同一套 Chat 流程。
|
||||
3. `/api/chat_stream` 只是对完整答案做事后分块,不具备模型生成过程中的真实增量输出能力。
|
||||
4. Controller 直接获取 `ChatModel` 和 `ToolCallbackProvider`,将模型基础设施细节暴露到入口层。
|
||||
5. SSE 使用 Controller 自建的无界缓存线程池,缺少统一生命周期和容量治理。
|
||||
6. 当前入口错误响应存在 HTTP 状态、外层 `ApiResponse` 与内层 `ChatResponse` 状态不一致的问题。
|
||||
|
||||
## 目标
|
||||
|
||||
1. 将会话生命周期、历史管理、执行调用和结果回写从 Controller 分离,形成单一 Chat 应用用例入口。
|
||||
2. 只保留一个 `/api/chat` 接口,并将其协议改为真正的 SSE。
|
||||
3. SSE 在模型或诊断链路产生内容时增量发送,而不是等待完整答案后再切片。
|
||||
4. 保留 `sessionId + runId` 作为一次 Chat Run 的稳定关联契约。
|
||||
5. 在满足入口职责分离的前提下使用最少组件,不引入没有实际职责的接口、工厂或适配层。
|
||||
|
||||
## 设计约束
|
||||
|
||||
- Controller 只负责请求校验、协议转换和 SSE 生命周期,不负责选择模型、组装工具、管理历史或编排诊断流程。
|
||||
- 同一次请求只能进入一个应用用例入口,禁止同步和流式路径各自维护一套业务逻辑。
|
||||
- 真正 SSE 至少需要区分元数据、内容增量、完成和错误事件。
|
||||
- `sessionId`、`runId` 必须在内容事件之前可获得,并用于日志、数据库和 Trace 对齐。
|
||||
- 客户端断开、超时和执行失败必须显式终止后台执行并完成 Run 状态记录。
|
||||
- 不保留旧 `/api/chat_stream` 或同步 `/api/chat` 的兼容分支,直接以新协议为准。
|
||||
- 优先使用 Spring 管理的执行设施和现有服务能力,不创建无界线程池。
|
||||
|
||||
## 建议的最小边界
|
||||
|
||||
```text
|
||||
POST /api/chat (SSE)
|
||||
-> ChatController:请求与 SSE 协议
|
||||
-> Chat 应用用例:会话、Run、历史和执行生命周期
|
||||
-> 现有 Chat 执行能力:简单回答或诊断编排
|
||||
```
|
||||
|
||||
这里的“应用用例”是职责边界,不要求预先拆出多层接口。只有出现独立变化原因或明确复用需求时才增加新组件。
|
||||
|
||||
## 验收标准
|
||||
|
||||
- [ ] 对外只保留一个 `POST /api/chat`,响应类型为 `text/event-stream`。
|
||||
- [ ] 删除 `/api/chat_stream` 及同步 Chat 兼容路径。
|
||||
- [ ] 首个内容事件在完整答案生成完成前发送,禁止通过固定字符切片伪造流式输出。
|
||||
- [ ] SSE 事件包含稳定的 metadata、content、error、done 契约。
|
||||
- [ ] Controller 不再直接依赖 `ChatModel`、`ToolCallbackProvider`,也不管理会话历史和 Run 持久化。
|
||||
- [ ] 同一请求的 `sessionId + runId` 在 SSE、应用日志、`diagnosis_run`、`agent_step` 和 `tool_invocation` 中一致。
|
||||
- [ ] 客户端断开、超时、模型失败和工具失败都有明确的资源清理与 Run 终态。
|
||||
- [ ] 不存在 Controller 自建的无界线程池。
|
||||
- [ ] 单元测试覆盖入口校验和 SSE 事件契约;端到端测试验证真实增量输出、断开清理及 Trace 对齐。
|
||||
|
||||
## 非目标
|
||||
|
||||
- 不在本 Issue 中重新设计 StateGraph 节点、Gatekeeper、Verifier 或证据协议。
|
||||
- 不为未来可能出现的其他传输协议预建通用框架。
|
||||
- 不引入多套 Command、Handler、Adapter、Factory 只为形式上的分层。
|
||||
- 不保留旧同步接口或 `/api/chat_stream` 的兼容逻辑。
|
||||
- 不以“完整答案分块发送”作为 SSE 验收通过条件。
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `src/main/java/com/superbiz/agent/controller/ChatController.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/session/SessionManager.java`
|
||||
- `src/test/java/com/superbiz/agent/controller/ChatControllerTest.java`
|
||||
- `src/test/java/com/superbiz/agent/service/ChatServiceGraphIntegrationTest.java`
|
||||
- `mvp/architecture/current-mvp-architecture.md`
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1 @@
|
||||
Devflow archive prepared and stage verification passed on 2026-07-21.
|
||||
@@ -0,0 +1 @@
|
||||
Committed after strict validation on 2026-07-21.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-21
|
||||
@@ -0,0 +1,29 @@
|
||||
# Stage 0 Acceptance Evidence
|
||||
|
||||
## Static Verification
|
||||
|
||||
- Contract package is dependency-free from Redis, JPA, Controller, Agent state and existing Hook classes.
|
||||
- Repository secret scan covers tracked worktree files and reports no known plaintext credential matches.
|
||||
- `scripts/query_mysql.py` requires `SUPERBIZ_MYSQL_PASSWORD` and exits before connecting when it is absent.
|
||||
- Spring AI 1.1.7 `SpringAiRetryProperties` bytecode shows a default `maxAttempts` value of 10; stage 2 must set underlying retries to one attempt and keep retry ownership in Harness.
|
||||
|
||||
## Script Verification
|
||||
|
||||
- `mvn -q '-Dtest=HarnessContractTest' test` - passed.
|
||||
- `mvn -q '-Dtest=ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,ChatControllerTest' test` - passed.
|
||||
- `mvn -q '-Dtest=HarnessContractTest,ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,ChatControllerTest' test` - passed.
|
||||
- `openspec validate single-react-design-freeze --strict` - passed.
|
||||
- `git diff --check` - passed; only existing Windows line-ending warnings were reported.
|
||||
- Secret scan for known committed key/password patterns - zero matches.
|
||||
- `python scripts/query_mysql.py "SELECT 1"` without `SUPERBIZ_MYSQL_PASSWORD` - exited before connecting with the expected missing-variable error.
|
||||
|
||||
## Runtime Behavior
|
||||
|
||||
- Public Chat runtime was not switched in stage zero.
|
||||
- No live model, Redis, MySQL or Milvus E2E was run; full live E2E remains stage 7 scope.
|
||||
|
||||
## External Security Prerequisite
|
||||
|
||||
- Plaintext credentials previously present in the repository must be rotated in their respective MySQL, Redis, DeepSeek, SiliconFlow and Milvus systems by the credential owner.
|
||||
- Repository changes can prove removal but cannot prove provider-side rotation.
|
||||
- Stage 3C and stage 7 must not claim live security/E2E acceptance until required environment variables contain rotated credentials.
|
||||
@@ -0,0 +1,25 @@
|
||||
# Brief: single-react-design-freeze
|
||||
|
||||
## Background
|
||||
|
||||
ISS-014 将当前 Chat 多 Agent/Hook/ThreadLocal 主链路重构为一个 Diagnosis ReAct Agent、一个确定性 Harness 和一个隔离 SemanticGuard。阶段 0 先冻结后续 10 个实施 change 共同依赖的契约和安全边界。
|
||||
|
||||
## Goal
|
||||
|
||||
产出可执行、可测试、可归档的 contract types、失败语义、安全前置和阶段门禁,同时保持现有公开 Chat 运行行为不变。
|
||||
|
||||
## Scope
|
||||
|
||||
- 类型化 Draft、Knowledge Answer、Fallback、previous turn 和状态枚举。
|
||||
- Tool ID、双状态、取消、重试、Redis canonical record 和 MySQL 安全设计冻结。
|
||||
- 明文脚本凭据清理和 focused baseline。
|
||||
- ISS-014、OpenSpec、devflow 对齐。
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- 不实现或接入新 Harness/Agent/Guard。
|
||||
- 不切换 `/api/chat`、不删除旧链路、不运行 live E2E。
|
||||
|
||||
## Source PRD
|
||||
|
||||
`mvp/issues/active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md` 是本 change 的完整 PRD 和总设计来源,不复制为第二份 PRD。
|
||||
@@ -0,0 +1,80 @@
|
||||
# Decisions
|
||||
|
||||
## Entry
|
||||
|
||||
- Parent issue: `ISS-014`
|
||||
- Change: `single-react-design-freeze`
|
||||
- Scale: `complex`
|
||||
- Interface impact: future L4; this change freezes contracts without switching runtime behavior.
|
||||
- Capability sources: sm-flow built-in clarify/context/propose, `grill-with-docs`, `openspec-propose`, `zoom-out`, `openspec-apply-change`, `openspec-archive-change`.
|
||||
|
||||
## Context Evidence
|
||||
|
||||
- `session-run-trace-isolation` established Chat Session and Diagnosis Run as separate lifecycles and made `runId` the Trace ownership key.
|
||||
- `verifier-evidence-reference-fidelity` established that no-evidence is a scoped negative observation, not proof that a problem does not exist.
|
||||
- `executor-composer-final-answer` established deterministic safe fallback boundaries and prohibited unfiltered raw output from reaching users.
|
||||
- `modular-rag-pipeline` established that retrieval trace and context packing are audit details rather than direct facts.
|
||||
- Current Spring AI Alibaba `ToolCallRequest` already provides `tool_call_id`; Harness must validate and persist it rather than create a second identity.
|
||||
- Current Spring AI `ChatModel.call(Prompt)` has no cancellation token, so cancellation must be expressed as layered, observable semantics rather than an unsupported absolute guarantee.
|
||||
|
||||
## Question Pool
|
||||
|
||||
| ID | Dimension | Question | Mode | Status |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | Terminology | Does `tool_call_id` use the framework ID or a Harness-generated ID? | user-interview | confirmed: framework ID |
|
||||
| Q2 | Terminology | Are invocation lifecycle and evidence outcome separate fields? | user-interview | confirmed: `status` + `evidence_status` |
|
||||
| Q3 | Boundary | Is ISS-014 one umbrella Issue with independent OpenSpec changes? | user-interview | confirmed: one Issue, 11 changes |
|
||||
| Q4 | Boundary | May stage 4 publish before Guards exist? | user-interview | confirmed: no; public cutover only in 6B |
|
||||
| Q5 | Lifecycle | What is the durable source of truth for `previous_turn` and `last_intent`? | user-interview | confirmed: `diagnosis_run` safe published result |
|
||||
| Q6 | Contract | How is KNOWLEDGE_QUERY citation validation represented? | evidence-driven | resolved: structured answer items with exact RAG bindings |
|
||||
| Q7 | Cancellation | What cancellation guarantees are technically enforceable? | evidence-driven | resolved: layered cancellation, no false hard-cancel claim |
|
||||
| Q8 | Acceptance | Are Apply, Archive and phase Git commit pre-authorized? | user-interview | confirmed: yes, for all phases |
|
||||
|
||||
## Confirmed Decisions
|
||||
|
||||
- `tool_call_id` is the framework Tool Call protocol ID. Harness validates non-empty, bounded, safe characters and Run-local uniqueness; duplicate/invalid/missing IDs fail closed.
|
||||
- Redis `status=PROJECTING/READY/ERROR` represents invocation/projector lifecycle.
|
||||
- `evidence_status=EVIDENCE_FOUND/NO_EVIDENCE/ERROR` represents result semantics.
|
||||
- `NO_EVIDENCE` may only support `NEGATIVE_OBSERVATION` within the exact query scope; it cannot prove absence, exclusion or health.
|
||||
- Stage 4 and 6A remain internal. Stage 6B performs the only public Chat cutover after stage 5 release gates pass.
|
||||
- ISS-014 remains the umbrella Issue. Eleven independent changes run serially; each must complete sm-flow, OpenSpec archive and Git commit before the next starts.
|
||||
- Full live E2E is deferred to stage 7; earlier stages run focused verification proportional to their change.
|
||||
- KNOWLEDGE_QUERY uses a dedicated structured draft: each answer item binds the single lookup `tool_call_id` and one or more returned `document_id` values; Harness validates exact membership before rendering `answer + references + limitations`. It does not reuse the Diagnosis Analysis schema and does not enter SemanticGuard in the first version.
|
||||
- Cancellation is layered: mark cancellation requested, prevent new model/Tool rounds and any final Draft release, invoke framework interruption, cancel owned Tool/JDBC work where supported, and rely on configured HTTP timeouts for an already-blocking synchronous model call. Run finalization is atomic and late results are discarded.
|
||||
- Public SSE `done` is not emitted after the client has disconnected; internal Run state still reaches `CANCELLED`.
|
||||
- `diagnosis_run` is the durable source for `intent`, `release_outcome` and `published_result`. Only the latest same-session `DIAGNOSIS + SUCCESS` record with a non-null safe published result may become `previous_turn`; `FALLBACK/FAILED/CANCELLED` remain auditable but are excluded.
|
||||
- `published_result` stores only `user_query/published_conclusion/scope/limitations/source_documents`; it excludes Tool Call IDs, raw evidence, full Draft and SemanticGuard audit reasons.
|
||||
|
||||
## Evidence-Driven Findings To Report
|
||||
|
||||
- The framework already exposes `ToolInterceptor`, structured output types, tool execution timeout, model/tool call limit hooks and `ReactAgent.interrupt`; later Harness stages should reuse these extension points.
|
||||
- Spring AI model dependencies include retry support, while current application configuration does not explicitly freeze all retry layers; stage 0 must define a retry inventory and stage 2 must enforce it.
|
||||
- Current `DiagnosisRun` stores a text answer and generic status but has no explicit `intent`, `release_outcome` or structured published result; Q5 must be resolved before the previous-turn contract is executable.
|
||||
- Current KNOWLEDGE_QUERY target behavior promises citation validation, but the issue only defines the Diagnosis Draft binding schema; the committed spec must add the dedicated answer-item contract described above.
|
||||
- Spring AI retry auto-configuration defaults `maxAttempts` to 10. The target Harness retry matrix requires underlying model/HTTP retries to be set to one attempt, with Router and SemanticGuard retries performed only by Harness.
|
||||
|
||||
## OpenSpec Backfill
|
||||
|
||||
- All confirmed decisions above must appear in design/specs/tasks before `.committed` is created.
|
||||
- All user-interview questions are confirmed; no pending decision blocks Commit.
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
Current input flows from `ChatController` into `ChatService`, which owns routing, ReactAgent construction, multi-Agent orchestration and final rendering; tools persist evidence through `ToolInvocationRecorder`, while Hooks and ThreadLocal bridge Run and verifier state. Stage zero introduces only dependency-free contract types under `harness.contract`; those types must not depend on Controller, Redis, JPA, Spring Agent state or current Hook classes. Later stages move ownership in order: RunContext, invocation store, Tool-specific projection, Diagnosis Agent, Guards, application use case and finally the public SSE adapter. `diagnosis_run` remains durable Run ownership, Redis canonical invocation remains short-lived Harness ownership, and `agent_step/tool_invocation` remain durable audit detail. The principal risk is spec/runtime drift, mitigated by archiving only the stage-zero contract capability now and delaying modifications to existing runtime capabilities until their implementation changes.
|
||||
|
||||
## Cross-Artifact Alignment
|
||||
|
||||
| Chain | Status | Evidence |
|
||||
|---|---|---|
|
||||
| ISS-014/brief goals, scope and non-goals → proposal | aligned | Proposal limits stage zero to contracts, security and baseline with no public cutover. |
|
||||
| proposal commitments → design | aligned | Design records every ID, status, Draft, fallback, previous-turn, retry, cancellation and phase-gate commitment. |
|
||||
| design decisions → specs | aligned | The single stage-zero capability has testable requirements for every stable contract boundary. |
|
||||
| specs observable behavior → tasks | aligned | Tasks create reusable types/tests, remove the secret, align artifacts and verify without switching runtime behavior. |
|
||||
|
||||
## Commit Gate Result
|
||||
|
||||
- Question pool covers terminology, boundary, lifecycle, contract, cancellation and acceptance.
|
||||
- All user-interview items are explicitly confirmed.
|
||||
- Evidence-driven conclusions were reported and written into design/spec/tasks.
|
||||
- Interface impact is recorded as future L4; this change itself does not switch the public API.
|
||||
- No devflow/OpenSpec conflict remains.
|
||||
@@ -0,0 +1,77 @@
|
||||
## Context
|
||||
|
||||
ISS-014 是一次 L4 Chat 重构的总设计来源,但实施被拆成 11 个必须串行归档的 OpenSpec changes。阶段 0 不切换公开协议或 Agent 运行链,只创建后续阶段复用的类型化契约、安全前置、失败语义和 focused baseline。
|
||||
|
||||
现有代码已经具备 `runId` Trace、工具调用审计、no-evidence 精确引用、Verifier/Composer fallback 和模块化 RAG,但这些能力分散在 `ChatService`、Hook、ThreadLocal、Tool 和 JSON 字符串中。Spring AI Alibaba 已提供 `ToolInterceptor`、结构化输出类型、工具执行超时、调用限额 Hook 和 `ReactAgent.interrupt`;Spring AI 底层 retry 默认最多 10 次,不能直接满足 ISS-014 的显式 Harness 重试矩阵。
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- 生成后续阶段可直接复用的 Java contract types 和枚举,不实现新 Agent 执行链。
|
||||
- 冻结 Tool ID、双状态、Draft、Knowledge Answer、Fallback、previous turn、SSE、重试和取消语义。
|
||||
- 冻结 MySQL fail-closed 允许子集和安全前置。
|
||||
- 移除仓库脚本和主配置中的明文凭据并建立改造前 focused baseline。
|
||||
- 保证 ISS-014、OpenSpec、devflow 术语和 11 个阶段门禁一致。
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- 不接入 Harness、Diagnosis Agent、EvidenceGuard 或 SemanticGuard 运行时。
|
||||
- 不修改 Controller 协议、旧 ChatService 行为、Redis invocation store 或数据库表。
|
||||
- 不实现 RAG/日志/MySQL Tool 投影。
|
||||
- 不运行完整 live E2E。
|
||||
|
||||
## Decisions
|
||||
|
||||
### Contract types are reusable runtime inputs
|
||||
|
||||
阶段 0 创建位于 `com.superbiz.agent.harness.contract` 的轻量 record/enum,而不是只写文档或引入 JSON Schema 引擎。后续 ReactAgent `outputType`、Harness validator、持久化和 SSE DTO 可以直接复用这些类型,减少同一字段在多个阶段重复定义。
|
||||
|
||||
### Framework Tool Call ID is canonical
|
||||
|
||||
`tool_call_id` 使用框架协议 ID。Harness 后续只校验非空、长度/字符安全和 Run 内唯一性,不生成第二套 ID。Redis Key 仍按 `runId + toolCallId` 隔离,真实性来自当前 Run 的 canonical record,而不是 ID 本身。
|
||||
|
||||
### Invocation lifecycle and evidence outcome are orthogonal
|
||||
|
||||
`InvocationStatus=PROJECTING/READY/ERROR` 只描述调用与投影生命周期;`EvidenceStatus=EVIDENCE_FOUND/NO_EVIDENCE/ERROR` 描述结果语义。`NO_EVIDENCE` 只允许绑定 `AnalysisKind=NEGATIVE_OBSERVATION`,并必须保留查询范围和零匹配信息。
|
||||
|
||||
### Diagnosis and knowledge answer contracts stay separate
|
||||
|
||||
Diagnosis Draft 使用 Analysis ID 与 Tool Call IDs;KNOWLEDGE_QUERY 使用 answer items,每项绑定一次 lookup 的 Tool Call ID 和返回的 document IDs。Harness 后续验证精确成员关系并确定性展开引用。首版 KNOWLEDGE_QUERY 不进入 SemanticGuard。
|
||||
|
||||
### Safe published context belongs to Diagnosis Run
|
||||
|
||||
`diagnosis_run` 是 `intent/release_outcome/published_result` 的持久化真理源。只有同 Session 最近一个 `DIAGNOSIS + SUCCESS` 且 published result 非空的 Run 可形成 previous turn。Published result 只含用户查询、已发布结论、范围、限制和 RAG 文档元数据。
|
||||
|
||||
### Cancellation is observable and layered
|
||||
|
||||
取消请求立即阻止新模型/Tool 轮次和最终 Draft 释放;框架中断、可控 Future/JDBC 取消尽力执行;已进入同步 `ChatModel.call` 的请求依靠底层 HTTP timeout。Run 终态通过原子状态转换保证唯一,晚到结果被丢弃,不宣称无法证明的底层硬取消。
|
||||
|
||||
### Harness owns all retries
|
||||
|
||||
底层 SDK/HTTP/数据库 retry 必须关闭或压为一次 attempt。Intent Router 和 SemanticGuard 的第二次 attempt 由 Harness 显式执行并审计。阶段 0 只冻结矩阵;阶段 2 实现执行器和配置。
|
||||
|
||||
### Stage boundaries are release boundaries
|
||||
|
||||
阶段 4 和 6A 仅内部运行;阶段 6B 在 Guards 已完成后原子切换公开入口。每个 change 必须 Archive 并 Git commit 后才能进入下一项,阶段 7 才执行统一 live E2E。
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Contract records过早绑定实现细节 → Mitigation:只包含跨阶段稳定字段,不包含 Redis、JPA 或框架对象。
|
||||
- [Risk] Spring AI provider 对 Tool Call ID 行为不同 → Mitigation:阶段 3A 使用 Fake Model 和当前 DeepSeek 路径验证,缺失或重复时 fail closed。
|
||||
- [Risk] 同步模型调用不能立即取消 → Mitigation:明确 layered semantics、HTTP timeout 和晚到结果丢弃,不把状态更新等同底层资源已终止。
|
||||
- [Risk] 设计冻结测试增加维护成本 → Mitigation:只保留 focused serialization/validation tests,不复制完整 E2E。
|
||||
- [Risk] 明文凭据可能已泄露 → Mitigation:仓库中移除并要求外部轮换;轮换证据记录在 acceptance。
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. 创建 contract types、契约测试和阶段台账,不接运行链。
|
||||
2. 清理查询脚本凭据,改为环境变量注入。
|
||||
3. 记录 focused baseline;Archive 本 change。
|
||||
4. 后续 10 个 change 逐步实现,并在各自 Archive 时同步对应运行 capability specs。
|
||||
|
||||
Rollback:阶段 0 没有公开行为变化;可删除新增 contract package/tests 并恢复文档。已经轮换的凭据不得回滚为旧值。
|
||||
|
||||
## Open Questions
|
||||
|
||||
无。所有影响实现的用户决策已经在 `decisions.md` 中确认。
|
||||
@@ -0,0 +1,32 @@
|
||||
## Why
|
||||
|
||||
当前 Chat 诊断把一个 ReAct 生命周期拆成多个 Agent、Hook、ThreadLocal 和重试分支,导致证据契约、失败语义、上下文预算和公开释放边界分散。进入分阶段重构前,需要先把 ISS-014 的跨阶段契约、安全前置和验收基线冻结为唯一可执行规格,避免后续 change 各自解释同一概念。
|
||||
|
||||
## What Changes
|
||||
|
||||
- 冻结单体 Diagnosis ReAct Agent、确定性 Harness、EvidenceGuard 和隔离 SemanticGuard 的职责边界。
|
||||
- 冻结 Diagnosis Draft、Analysis、Conclusion、Fallback、Intent Router 和 SSE 事件契约。
|
||||
- 冻结 Tool Call ID、调用生命周期 `status`、结果语义 `evidence_status`、Redis canonical invocation 和 ToolResultProjector 命名。
|
||||
- 冻结预算、取消、重试、隐藏重试禁用和 Run 终态语义,但不在本阶段实现新运行链路。
|
||||
- 冻结只读 MySQL Tool 的 JSqlParser 允许子集、静态 allowlist 和安全前置。
|
||||
- 移除仓库脚本和主配置中的明文数据库、Redis、模型及向量服务凭据,并记录必须完成外部轮换。
|
||||
- 建立改造前 focused baseline 和后续 11 个串行 sm-flow change 的阶段台账。
|
||||
- **BREAKING(后续阶段实施)**:最终仅保留 `POST /api/chat` SSE、移除旧多 Agent/Graph Chat 主链路和旧 Tool Contract;本 change 不执行公开协议切换。
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `single-react-diagnosis-harness`: 冻结单体诊断 Agent、Harness、证据状态、Guard、Fallback、预算、重试和阶段门禁的跨阶段基础契约。
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- None. 本阶段不声明旧运行能力已经迁移;后续 change 在实现对应行为时再修改现有 capability specs。
|
||||
|
||||
## Impact
|
||||
|
||||
- 设计与规格:ISS-014、OpenSpec 主规格、devflow 词汇表和阶段归档台账。
|
||||
- 契约测试:Draft/Fallback、状态语义、SSE、Router、预算、重试和 MySQL 安全基线。
|
||||
- 安全:`scripts/query_mysql.py` 和 `application.yml` 中的明文凭据必须移除并在外部轮换。
|
||||
- 后续代码范围:ChatController、ChatService、Agent/Hook、Tool、Redis、JPA/Flyway、静态前端、Trace/Eval fixtures。
|
||||
- 本 change 不切换 Controller 协议、不接入新 Agent、不修改公开运行行为。
|
||||
+75
@@ -0,0 +1,75 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Design contracts SHALL separate deterministic control from diagnosis reasoning
|
||||
The frozen contract set SHALL define Diagnosis Agent as the only diagnosis report author, Harness as deterministic execution control, EvidenceGuard as deterministic evidence validation, and SemanticGuard as an isolated single-turn semantic reviewer.
|
||||
|
||||
#### Scenario: Contract ownership is inspected
|
||||
- **WHEN** a later phase reads the stage-zero contracts
|
||||
- **THEN** no Harness contract assigns Planner, Executor, Composer, workflow routing, or diagnosis reasoning responsibilities to Harness
|
||||
|
||||
### Requirement: Tool invocation identity SHALL use the framework Tool Call ID
|
||||
The contract SHALL use the framework-provided `tool_call_id` as the sole Tool invocation reference and SHALL require later Harness implementations to reject missing, invalid, or duplicate IDs within a Run.
|
||||
|
||||
#### Scenario: Duplicate Tool Call ID is proposed
|
||||
- **WHEN** two Tool actions in one Run present the same framework Tool Call ID
|
||||
- **THEN** the contract classifies the second action as an error and prohibits overwriting the first canonical invocation
|
||||
|
||||
### Requirement: Invocation status and evidence status SHALL be independent
|
||||
The contract SHALL define `PROJECTING/READY/ERROR` as invocation lifecycle states and `EVIDENCE_FOUND/NO_EVIDENCE/ERROR` as evidence result states.
|
||||
|
||||
#### Scenario: Successful query returns no evidence
|
||||
- **WHEN** a Tool executes successfully and its bounded projection contains zero matching evidence
|
||||
- **THEN** invocation status is `READY` and evidence status is `NO_EVIDENCE`
|
||||
|
||||
#### Scenario: No-evidence result is cited
|
||||
- **WHEN** a Diagnosis Draft cites a `NO_EVIDENCE` Tool result
|
||||
- **THEN** the Analysis kind MUST be `NEGATIVE_OBSERVATION` and MUST remain bounded to the Tool query scope
|
||||
|
||||
### Requirement: Diagnosis Draft SHALL expose typed report structure
|
||||
The contract SHALL define conclusion, analysis items, action plan, recommendations, limitations and Tool Call bindings without exposing chain-of-thought or raw Tool payloads.
|
||||
|
||||
#### Scenario: Draft contains an analysis item
|
||||
- **WHEN** the Diagnosis Agent emits a structured Draft
|
||||
- **THEN** every Analysis has a unique analysis ID, a fixed Analysis kind and at least one Tool Call ID
|
||||
|
||||
### Requirement: Knowledge answers SHALL use exact RAG bindings
|
||||
The KNOWLEDGE_QUERY contract SHALL represent the answer as bounded answer items whose references identify the single lookup Tool Call and returned document IDs.
|
||||
|
||||
#### Scenario: Knowledge answer cites an unknown document
|
||||
- **WHEN** an answer item references a document ID absent from the bounded lookup result
|
||||
- **THEN** later Harness validation rejects the answer instead of publishing the fabricated citation
|
||||
|
||||
### Requirement: Safe fallback SHALL use a fixed schema
|
||||
The contract SHALL define stable fallback types for evidence validation failure, semantic unsupported and semantic unavailable outcomes, and SHALL exclude unvalidated Draft content and internal errors.
|
||||
|
||||
#### Scenario: Evidence validation fails twice
|
||||
- **WHEN** initial validation and the single no-Tool structural repair both fail
|
||||
- **THEN** the fallback type is `EVIDENCE_VALIDATION_FAILED` and verified sources are empty
|
||||
|
||||
### Requirement: Previous turn SHALL come from a safe durable Run result
|
||||
The contract SHALL define `diagnosis_run` as the durable source of intent, release outcome and safe published result, and SHALL exclude fallback, failed and cancelled Runs from Diagnosis previous-turn selection.
|
||||
|
||||
#### Scenario: Latest Run is a fallback
|
||||
- **WHEN** the latest same-session Run ended with `FALLBACK`
|
||||
- **THEN** it is not used as Diagnosis previous turn and selection continues to the latest eligible `DIAGNOSIS + SUCCESS` Run
|
||||
|
||||
### Requirement: Cancellation SHALL be layered and observable
|
||||
The contract SHALL distinguish cancellation request, prevention of new work, framework interruption, cancellable Tool work and HTTP timeout for already-blocking synchronous model calls.
|
||||
|
||||
#### Scenario: Client disconnects during a model call
|
||||
- **WHEN** an SSE client disconnects while a synchronous model call is in flight
|
||||
- **THEN** the system prevents later Draft release, requests interruption, records an internal cancelled terminal state and discards any late model result
|
||||
|
||||
### Requirement: Retry attempts SHALL be owned by Harness
|
||||
The contract SHALL require underlying SDK, HTTP and database retry layers to execute one attempt, while Harness explicitly owns any allowed Router or SemanticGuard retry.
|
||||
|
||||
#### Scenario: SemanticGuard returns an invalid schema
|
||||
- **WHEN** the first SemanticGuard attempt returns an invalid structured result
|
||||
- **THEN** Harness may execute one second attempt with the same verified snapshot and records both attempts
|
||||
|
||||
### Requirement: Phase gates SHALL remain serial
|
||||
The implementation plan SHALL contain eleven independent OpenSpec changes and SHALL prohibit starting a change before its predecessor is archived and committed.
|
||||
|
||||
#### Scenario: Stage 4 completes internal Agent tests
|
||||
- **WHEN** stage 4 passes its focused tests but stage 5 Guards are not implemented
|
||||
- **THEN** the public Chat entry remains on the old path and stage 6B cutover is prohibited
|
||||
@@ -0,0 +1,21 @@
|
||||
## 1. Contract Model
|
||||
|
||||
- [x] 1.1 Add typed enums for intent, release outcome, invocation status, evidence status, analysis kind, semantic verdict, fallback type and SSE outcome.
|
||||
- [x] 1.2 Add reusable records for Diagnosis Draft, Knowledge Answer Draft, safe fallback, published result and previous turn.
|
||||
- [x] 1.3 Add focused serialization and contract-shape tests covering positive evidence, no-evidence negative observation and forbidden raw/internal fields.
|
||||
|
||||
## 2. Security And Configuration Baseline
|
||||
|
||||
- [x] 2.1 Remove tracked plaintext credentials from `scripts/query_mysql.py` and `application.yml`, require environment-based secrets, and ignore local secret files.
|
||||
- [x] 2.2 Record the required external credential rotation and the Spring AI hidden-retry baseline without changing the public Chat runtime in this stage.
|
||||
|
||||
## 3. Design Freeze Alignment
|
||||
|
||||
- [x] 3.1 Keep ISS-014, OpenSpec artifacts and the devflow glossary aligned on the 11 serial changes, framework Tool Call ID, dual status fields and safe previous-turn source.
|
||||
- [x] 3.2 Add an architecture audit and cross-artifact alignment result to `decisions.md`.
|
||||
- [x] 3.3 Create the sm-flow `.committed` marker after proposal/design/specs/tasks and all decision gates pass.
|
||||
|
||||
## 4. Verification
|
||||
|
||||
- [x] 4.1 Run focused contract tests and the smallest existing Chat/evidence baseline needed to prove stage zero did not switch runtime behavior.
|
||||
- [x] 4.2 Run strict OpenSpec validation and record commands, results and unverified external rotation in acceptance evidence.
|
||||
@@ -0,0 +1,78 @@
|
||||
# single-react-diagnosis-harness Specification
|
||||
|
||||
## Purpose
|
||||
TBD - created by archiving change single-react-design-freeze. Update Purpose after archive.
|
||||
## Requirements
|
||||
### Requirement: Design contracts SHALL separate deterministic control from diagnosis reasoning
|
||||
The frozen contract set SHALL define Diagnosis Agent as the only diagnosis report author, Harness as deterministic execution control, EvidenceGuard as deterministic evidence validation, and SemanticGuard as an isolated single-turn semantic reviewer.
|
||||
|
||||
#### Scenario: Contract ownership is inspected
|
||||
- **WHEN** a later phase reads the stage-zero contracts
|
||||
- **THEN** no Harness contract assigns Planner, Executor, Composer, workflow routing, or diagnosis reasoning responsibilities to Harness
|
||||
|
||||
### Requirement: Tool invocation identity SHALL use the framework Tool Call ID
|
||||
The contract SHALL use the framework-provided `tool_call_id` as the sole Tool invocation reference and SHALL require later Harness implementations to reject missing, invalid, or duplicate IDs within a Run.
|
||||
|
||||
#### Scenario: Duplicate Tool Call ID is proposed
|
||||
- **WHEN** two Tool actions in one Run present the same framework Tool Call ID
|
||||
- **THEN** the contract classifies the second action as an error and prohibits overwriting the first canonical invocation
|
||||
|
||||
### Requirement: Invocation status and evidence status SHALL be independent
|
||||
The contract SHALL define `PROJECTING/READY/ERROR` as invocation lifecycle states and `EVIDENCE_FOUND/NO_EVIDENCE/ERROR` as evidence result states.
|
||||
|
||||
#### Scenario: Successful query returns no evidence
|
||||
- **WHEN** a Tool executes successfully and its bounded projection contains zero matching evidence
|
||||
- **THEN** invocation status is `READY` and evidence status is `NO_EVIDENCE`
|
||||
|
||||
#### Scenario: No-evidence result is cited
|
||||
- **WHEN** a Diagnosis Draft cites a `NO_EVIDENCE` Tool result
|
||||
- **THEN** the Analysis kind MUST be `NEGATIVE_OBSERVATION` and MUST remain bounded to the Tool query scope
|
||||
|
||||
### Requirement: Diagnosis Draft SHALL expose typed report structure
|
||||
The contract SHALL define conclusion, analysis items, action plan, recommendations, limitations and Tool Call bindings without exposing chain-of-thought or raw Tool payloads.
|
||||
|
||||
#### Scenario: Draft contains an analysis item
|
||||
- **WHEN** the Diagnosis Agent emits a structured Draft
|
||||
- **THEN** every Analysis has a unique analysis ID, a fixed Analysis kind and at least one Tool Call ID
|
||||
|
||||
### Requirement: Knowledge answers SHALL use exact RAG bindings
|
||||
The KNOWLEDGE_QUERY contract SHALL represent the answer as bounded answer items whose references identify the single lookup Tool Call and returned document IDs.
|
||||
|
||||
#### Scenario: Knowledge answer cites an unknown document
|
||||
- **WHEN** an answer item references a document ID absent from the bounded lookup result
|
||||
- **THEN** later Harness validation rejects the answer instead of publishing the fabricated citation
|
||||
|
||||
### Requirement: Safe fallback SHALL use a fixed schema
|
||||
The contract SHALL define stable fallback types for evidence validation failure, semantic unsupported and semantic unavailable outcomes, and SHALL exclude unvalidated Draft content and internal errors.
|
||||
|
||||
#### Scenario: Evidence validation fails twice
|
||||
- **WHEN** initial validation and the single no-Tool structural repair both fail
|
||||
- **THEN** the fallback type is `EVIDENCE_VALIDATION_FAILED` and verified sources are empty
|
||||
|
||||
### Requirement: Previous turn SHALL come from a safe durable Run result
|
||||
The contract SHALL define `diagnosis_run` as the durable source of intent, release outcome and safe published result, and SHALL exclude fallback, failed and cancelled Runs from Diagnosis previous-turn selection.
|
||||
|
||||
#### Scenario: Latest Run is a fallback
|
||||
- **WHEN** the latest same-session Run ended with `FALLBACK`
|
||||
- **THEN** it is not used as Diagnosis previous turn and selection continues to the latest eligible `DIAGNOSIS + SUCCESS` Run
|
||||
|
||||
### Requirement: Cancellation SHALL be layered and observable
|
||||
The contract SHALL distinguish cancellation request, prevention of new work, framework interruption, cancellable Tool work and HTTP timeout for already-blocking synchronous model calls.
|
||||
|
||||
#### Scenario: Client disconnects during a model call
|
||||
- **WHEN** an SSE client disconnects while a synchronous model call is in flight
|
||||
- **THEN** the system prevents later Draft release, requests interruption, records an internal cancelled terminal state and discards any late model result
|
||||
|
||||
### Requirement: Retry attempts SHALL be owned by Harness
|
||||
The contract SHALL require underlying SDK, HTTP and database retry layers to execute one attempt, while Harness explicitly owns any allowed Router or SemanticGuard retry.
|
||||
|
||||
#### Scenario: SemanticGuard returns an invalid schema
|
||||
- **WHEN** the first SemanticGuard attempt returns an invalid structured result
|
||||
- **THEN** Harness may execute one second attempt with the same verified snapshot and records both attempts
|
||||
|
||||
### Requirement: Phase gates SHALL remain serial
|
||||
The implementation plan SHALL contain eleven independent OpenSpec changes and SHALL prohibit starting a change before its predecessor is archived and committed.
|
||||
|
||||
#### Scenario: Stage 4 completes internal Agent tests
|
||||
- **WHEN** stage 4 passes its focused tests but stage 5 Guards are not implemented
|
||||
- **THEN** the public Chat entry remains on the old path and stage 6B cutover is prohibited
|
||||
+13
-6
@@ -22,13 +22,20 @@ except ImportError:
|
||||
print("缺少依赖,请先执行: pip install pymysql")
|
||||
sys.exit(1)
|
||||
|
||||
# 从 application.yml 读取的连接信息
|
||||
def required_env(name: str) -> str:
|
||||
value = os.getenv(name)
|
||||
if value is None or not value.strip():
|
||||
print(f"缺少必需环境变量: {name}")
|
||||
sys.exit(2)
|
||||
return value.strip()
|
||||
|
||||
|
||||
DB_CONFIG = {
|
||||
"host": "119.29.78.52",
|
||||
"port": 33306,
|
||||
"user": "root",
|
||||
"password": "!Fucker123..",
|
||||
"database": "superbiz_agent",
|
||||
"host": os.getenv("SUPERBIZ_MYSQL_HOST", "119.29.78.52"),
|
||||
"port": int(os.getenv("SUPERBIZ_MYSQL_PORT", "33306")),
|
||||
"user": os.getenv("SUPERBIZ_MYSQL_USERNAME", "root"),
|
||||
"password": required_env("SUPERBIZ_MYSQL_PASSWORD"),
|
||||
"database": os.getenv("SUPERBIZ_MYSQL_DATABASE", "superbiz_agent"),
|
||||
"charset": "utf8mb4",
|
||||
"cursorclass": pymysql.cursors.DictCursor,
|
||||
}
|
||||
|
||||
@@ -0,0 +1,13 @@
|
||||
package com.superbiz.agent.harness.contract;
|
||||
|
||||
public enum AnalysisKind {
|
||||
NORMAL,
|
||||
NEGATIVE_OBSERVATION;
|
||||
|
||||
public boolean accepts(EvidenceStatus evidenceStatus) {
|
||||
return switch (this) {
|
||||
case NORMAL -> evidenceStatus == EvidenceStatus.EVIDENCE_FOUND;
|
||||
case NEGATIVE_OBSERVATION -> evidenceStatus == EvidenceStatus.NO_EVIDENCE;
|
||||
};
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,13 @@
|
||||
package com.superbiz.agent.harness.contract;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
final class ContractCollections {
|
||||
|
||||
private ContractCollections() {
|
||||
}
|
||||
|
||||
static <T> List<T> immutable(List<T> values) {
|
||||
return values == null ? List.of() : List.copyOf(values);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,67 @@
|
||||
package com.superbiz.agent.harness.contract;
|
||||
|
||||
import com.fasterxml.jackson.annotation.JsonProperty;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
public record DiagnosisDraft(
|
||||
@JsonProperty("conclusion") Conclusion conclusion,
|
||||
@JsonProperty("analysis") List<AnalysisItem> analysis,
|
||||
@JsonProperty("action_plan") List<ActionPlanItem> actionPlan,
|
||||
@JsonProperty("recommendations") List<Recommendation> recommendations,
|
||||
@JsonProperty("limitations") Limitations limitations) {
|
||||
|
||||
public DiagnosisDraft {
|
||||
analysis = ContractCollections.immutable(analysis);
|
||||
actionPlan = ContractCollections.immutable(actionPlan);
|
||||
recommendations = ContractCollections.immutable(recommendations);
|
||||
}
|
||||
|
||||
public record Conclusion(
|
||||
@JsonProperty("text") String text,
|
||||
@JsonProperty("based_on_analysis_ids") List<String> basedOnAnalysisIds) {
|
||||
|
||||
public Conclusion {
|
||||
basedOnAnalysisIds = ContractCollections.immutable(basedOnAnalysisIds);
|
||||
}
|
||||
}
|
||||
|
||||
public record AnalysisItem(
|
||||
@JsonProperty("analysis_id") String analysisId,
|
||||
@JsonProperty("kind") AnalysisKind kind,
|
||||
@JsonProperty("text") String text,
|
||||
@JsonProperty("tool_call_ids") List<String> toolCallIds) {
|
||||
|
||||
public AnalysisItem {
|
||||
toolCallIds = ContractCollections.immutable(toolCallIds);
|
||||
}
|
||||
}
|
||||
|
||||
public record ActionPlanItem(
|
||||
@JsonProperty("action") String action,
|
||||
@JsonProperty("based_on_analysis_ids") List<String> basedOnAnalysisIds,
|
||||
@JsonProperty("requires_human_confirmation") boolean requiresHumanConfirmation) {
|
||||
|
||||
public ActionPlanItem {
|
||||
basedOnAnalysisIds = ContractCollections.immutable(basedOnAnalysisIds);
|
||||
}
|
||||
}
|
||||
|
||||
public record Recommendation(
|
||||
@JsonProperty("text") String text,
|
||||
@JsonProperty("based_on_analysis_ids") List<String> basedOnAnalysisIds) {
|
||||
|
||||
public Recommendation {
|
||||
basedOnAnalysisIds = ContractCollections.immutable(basedOnAnalysisIds);
|
||||
}
|
||||
}
|
||||
|
||||
public record Limitations(
|
||||
@JsonProperty("scope") String scope,
|
||||
@JsonProperty("missing_info") List<String> missingInfo) {
|
||||
|
||||
public Limitations {
|
||||
missingInfo = ContractCollections.immutable(missingInfo);
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,7 @@
|
||||
package com.superbiz.agent.harness.contract;
|
||||
|
||||
public enum EvidenceStatus {
|
||||
EVIDENCE_FOUND,
|
||||
NO_EVIDENCE,
|
||||
ERROR
|
||||
}
|
||||
@@ -0,0 +1,7 @@
|
||||
package com.superbiz.agent.harness.contract;
|
||||
|
||||
public enum FallbackType {
|
||||
EVIDENCE_VALIDATION_FAILED,
|
||||
SEMANTIC_UNSUPPORTED,
|
||||
SEMANTIC_UNAVAILABLE
|
||||
}
|
||||
@@ -0,0 +1,7 @@
|
||||
package com.superbiz.agent.harness.contract;
|
||||
|
||||
public enum IntentType {
|
||||
SYSTEM_CHAT,
|
||||
KNOWLEDGE_QUERY,
|
||||
DIAGNOSIS
|
||||
}
|
||||
@@ -0,0 +1,7 @@
|
||||
package com.superbiz.agent.harness.contract;
|
||||
|
||||
public enum InvocationStatus {
|
||||
PROJECTING,
|
||||
READY,
|
||||
ERROR
|
||||
}
|
||||
@@ -0,0 +1,25 @@
|
||||
package com.superbiz.agent.harness.contract;
|
||||
|
||||
import com.fasterxml.jackson.annotation.JsonProperty;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
public record KnowledgeAnswerDraft(
|
||||
@JsonProperty("answer_items") List<AnswerItem> answerItems,
|
||||
@JsonProperty("limitations") List<String> limitations) {
|
||||
|
||||
public KnowledgeAnswerDraft {
|
||||
answerItems = ContractCollections.immutable(answerItems);
|
||||
limitations = ContractCollections.immutable(limitations);
|
||||
}
|
||||
|
||||
public record AnswerItem(
|
||||
@JsonProperty("text") String text,
|
||||
@JsonProperty("tool_call_id") String toolCallId,
|
||||
@JsonProperty("document_ids") List<String> documentIds) {
|
||||
|
||||
public AnswerItem {
|
||||
documentIds = ContractCollections.immutable(documentIds);
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,27 @@
|
||||
package com.superbiz.agent.harness.contract;
|
||||
|
||||
import com.fasterxml.jackson.annotation.JsonProperty;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
public record PreviousTurn(
|
||||
@JsonProperty("user_query") String userQuery,
|
||||
@JsonProperty("published_conclusion") String publishedConclusion,
|
||||
@JsonProperty("scope") String scope,
|
||||
@JsonProperty("limitations") List<String> limitations,
|
||||
@JsonProperty("source_documents") List<SourceDocument> sourceDocuments) {
|
||||
|
||||
public PreviousTurn {
|
||||
limitations = ContractCollections.immutable(limitations);
|
||||
sourceDocuments = ContractCollections.immutable(sourceDocuments);
|
||||
}
|
||||
|
||||
public static PreviousTurn from(PublishedResult result) {
|
||||
return new PreviousTurn(
|
||||
result.userQuery(),
|
||||
result.publishedConclusion(),
|
||||
result.scope(),
|
||||
result.limitations(),
|
||||
result.sourceDocuments());
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,18 @@
|
||||
package com.superbiz.agent.harness.contract;
|
||||
|
||||
import com.fasterxml.jackson.annotation.JsonProperty;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
public record PublishedResult(
|
||||
@JsonProperty("user_query") String userQuery,
|
||||
@JsonProperty("published_conclusion") String publishedConclusion,
|
||||
@JsonProperty("scope") String scope,
|
||||
@JsonProperty("limitations") List<String> limitations,
|
||||
@JsonProperty("source_documents") List<SourceDocument> sourceDocuments) {
|
||||
|
||||
public PublishedResult {
|
||||
limitations = ContractCollections.immutable(limitations);
|
||||
sourceDocuments = ContractCollections.immutable(sourceDocuments);
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,8 @@
|
||||
package com.superbiz.agent.harness.contract;
|
||||
|
||||
public enum ReleaseOutcome {
|
||||
SUCCESS,
|
||||
FALLBACK,
|
||||
FAILED,
|
||||
CANCELLED
|
||||
}
|
||||
@@ -0,0 +1,26 @@
|
||||
package com.superbiz.agent.harness.contract;
|
||||
|
||||
import com.fasterxml.jackson.annotation.JsonProperty;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
public record SafeFallback(
|
||||
@JsonProperty("type") FallbackType type,
|
||||
@JsonProperty("conclusion") String conclusion,
|
||||
@JsonProperty("message") String message,
|
||||
@JsonProperty("verified_sources") List<VerifiedSource> verifiedSources,
|
||||
@JsonProperty("limitations") List<String> limitations,
|
||||
@JsonProperty("next_steps") List<String> nextSteps) {
|
||||
|
||||
public SafeFallback {
|
||||
verifiedSources = ContractCollections.immutable(verifiedSources);
|
||||
limitations = ContractCollections.immutable(limitations);
|
||||
nextSteps = ContractCollections.immutable(nextSteps);
|
||||
}
|
||||
|
||||
public record VerifiedSource(
|
||||
@JsonProperty("source_type") String sourceType,
|
||||
@JsonProperty("source") String source,
|
||||
@JsonProperty("scope") String scope) {
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,6 @@
|
||||
package com.superbiz.agent.harness.contract;
|
||||
|
||||
public enum SemanticVerdict {
|
||||
SUPPORTED,
|
||||
UNSUPPORTED
|
||||
}
|
||||
@@ -0,0 +1,8 @@
|
||||
package com.superbiz.agent.harness.contract;
|
||||
|
||||
import com.fasterxml.jackson.annotation.JsonProperty;
|
||||
|
||||
public record SourceDocument(
|
||||
@JsonProperty("document_id") String documentId,
|
||||
@JsonProperty("title") String title) {
|
||||
}
|
||||
@@ -0,0 +1,7 @@
|
||||
package com.superbiz.agent.harness.contract;
|
||||
|
||||
public enum SseOutcome {
|
||||
SUCCESS,
|
||||
FALLBACK,
|
||||
FAILED
|
||||
}
|
||||
@@ -22,7 +22,7 @@ milvus:
|
||||
password: ""
|
||||
database: db_4a578da0f27ce9d
|
||||
timeout: 10000
|
||||
token: d246a77f43a109685596e3c68ecfd359e1cd8b29d35c41d708160f02b97ae623d2ab392738df740d0c77b58ca1bbadfa7c412140
|
||||
token: ${MILVUS_TOKEN}
|
||||
secure: true
|
||||
vector-dim: 1024 # BGE-M3 = 1024,换模型时同步改
|
||||
|
||||
@@ -45,7 +45,7 @@ spring:
|
||||
datasource:
|
||||
url: jdbc:mysql://119.29.78.52:33306/superbiz_agent?useUnicode=true&serverTimezone=Asia/Shanghai&allowPublicKeyRetrieval=true
|
||||
username: root
|
||||
password: '!Fucker123..'
|
||||
password: ${SUPERBIZ_MYSQL_PASSWORD}
|
||||
driver-class-name: com.mysql.cj.jdbc.Driver
|
||||
hikari:
|
||||
maximum-pool-size: 5
|
||||
@@ -83,7 +83,7 @@ spring:
|
||||
redis:
|
||||
host: 119.29.78.52
|
||||
port: 33308
|
||||
password: '!Fucker123..'
|
||||
password: ${SUPERBIZ_REDIS_PASSWORD}
|
||||
database: 0
|
||||
timeout: 3000
|
||||
lettuce:
|
||||
@@ -119,7 +119,7 @@ spring:
|
||||
|
||||
# --- Chat: DeepSeek (原生) ---
|
||||
deepseek:
|
||||
api-key: sk-1f44696abe644bd684f09cc43f12c557
|
||||
api-key: ${DEEPSEEK_API_KEY}
|
||||
base-url: https://api.deepseek.com
|
||||
chat:
|
||||
options:
|
||||
@@ -136,7 +136,7 @@ spring:
|
||||
|
||||
# --- Embedding: SiliconFlow BGE-M3 ---
|
||||
siliconflow:
|
||||
api-key: sk-rlxqcnlohjqwkzoffollthmzzfiohngdrabrmmqhcgtewnzx
|
||||
api-key: ${SILICONFLOW_API_KEY}
|
||||
base-url: https://api.siliconflow.cn
|
||||
embedding:
|
||||
model: BAAI/bge-m3
|
||||
|
||||
@@ -0,0 +1,96 @@
|
||||
package com.superbiz.agent.harness.contract;
|
||||
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
import static org.junit.jupiter.api.Assertions.assertEquals;
|
||||
import static org.junit.jupiter.api.Assertions.assertFalse;
|
||||
import static org.junit.jupiter.api.Assertions.assertNull;
|
||||
import static org.junit.jupiter.api.Assertions.assertTrue;
|
||||
|
||||
class HarnessContractTest {
|
||||
|
||||
private final ObjectMapper objectMapper = new ObjectMapper();
|
||||
|
||||
@Test
|
||||
void diagnosisDraftUsesStableBoundedShape() throws Exception {
|
||||
DiagnosisDraft draft = new DiagnosisDraft(
|
||||
new DiagnosisDraft.Conclusion("Pool exhausted", List.of("a1")),
|
||||
List.of(new DiagnosisDraft.AnalysisItem(
|
||||
"a1",
|
||||
AnalysisKind.NORMAL,
|
||||
"active=50 max=50",
|
||||
List.of("call-1"))),
|
||||
List.of(new DiagnosisDraft.ActionPlanItem(
|
||||
"Inspect long transactions",
|
||||
List.of("a1"),
|
||||
false)),
|
||||
List.of(new DiagnosisDraft.Recommendation(
|
||||
"Add pool wait alerts",
|
||||
List.of("a1"))),
|
||||
new DiagnosisDraft.Limitations("order-service, last 30 minutes", List.of("No slow SQL data")));
|
||||
|
||||
JsonNode json = objectMapper.valueToTree(draft);
|
||||
|
||||
assertEquals("call-1", json.path("analysis").get(0).path("tool_call_ids").get(0).asText());
|
||||
assertEquals("NORMAL", json.path("analysis").get(0).path("kind").asText());
|
||||
assertTrue(json.has("action_plan"));
|
||||
assertFalse(json.toString().contains("raw_response"));
|
||||
assertFalse(json.toString().contains("thought"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void evidenceStatusRemainsSeparateFromInvocationStatus() {
|
||||
assertTrue(AnalysisKind.NORMAL.accepts(EvidenceStatus.EVIDENCE_FOUND));
|
||||
assertFalse(AnalysisKind.NORMAL.accepts(EvidenceStatus.NO_EVIDENCE));
|
||||
assertTrue(AnalysisKind.NEGATIVE_OBSERVATION.accepts(EvidenceStatus.NO_EVIDENCE));
|
||||
assertFalse(AnalysisKind.NEGATIVE_OBSERVATION.accepts(EvidenceStatus.ERROR));
|
||||
assertEquals(InvocationStatus.READY, InvocationStatus.valueOf("READY"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void knowledgeAnswerBindsExactToolAndDocuments() throws Exception {
|
||||
KnowledgeAnswerDraft draft = new KnowledgeAnswerDraft(
|
||||
List.of(new KnowledgeAnswerDraft.AnswerItem(
|
||||
"Check the timeout code first.",
|
||||
"call-rag-1",
|
||||
List.of("payment-timeout-guide"))),
|
||||
List.of("Runtime state was not queried"));
|
||||
|
||||
JsonNode json = objectMapper.valueToTree(draft);
|
||||
|
||||
assertEquals("call-rag-1", json.path("answer_items").get(0).path("tool_call_id").asText());
|
||||
assertEquals("payment-timeout-guide",
|
||||
json.path("answer_items").get(0).path("document_ids").get(0).asText());
|
||||
}
|
||||
|
||||
@Test
|
||||
void fallbackAndPreviousTurnExcludeInternalEvidence() throws Exception {
|
||||
SafeFallback fallback = new SafeFallback(
|
||||
FallbackType.SEMANTIC_UNSUPPORTED,
|
||||
null,
|
||||
"Current evidence is insufficient",
|
||||
List.of(new SafeFallback.VerifiedSource("LOG", "APPLICATION", "last 30 minutes")),
|
||||
List.of("Semantic validation did not pass"),
|
||||
List.of("Collect the missing data and retry"));
|
||||
PublishedResult result = new PublishedResult(
|
||||
"Check payment timeout",
|
||||
"Connection pool exhaustion",
|
||||
"order-service, last 30 minutes",
|
||||
List.of("No slow SQL data"),
|
||||
List.of(new SourceDocument("payment-timeout-guide", "Payment timeout guide")));
|
||||
|
||||
JsonNode fallbackJson = objectMapper.valueToTree(fallback);
|
||||
JsonNode previousJson = objectMapper.valueToTree(PreviousTurn.from(result));
|
||||
|
||||
assertNull(fallback.conclusion());
|
||||
assertEquals("SEMANTIC_UNSUPPORTED", fallbackJson.path("type").asText());
|
||||
assertEquals("payment-timeout-guide",
|
||||
previousJson.path("source_documents").get(0).path("document_id").asText());
|
||||
assertFalse(previousJson.toString().contains("tool_call_id"));
|
||||
assertFalse(previousJson.toString().contains("raw_response"));
|
||||
}
|
||||
}
|
||||
@@ -4,6 +4,7 @@ import io.milvus.client.MilvusServiceClient;
|
||||
import io.milvus.param.ConnectParam;
|
||||
import io.milvus.param.R;
|
||||
import io.milvus.param.collection.HasCollectionParam;
|
||||
import org.junit.jupiter.api.Assumptions;
|
||||
import org.junit.jupiter.api.Test;
|
||||
|
||||
/**
|
||||
@@ -13,9 +14,12 @@ public class SimpleMilvusTest {
|
||||
|
||||
@Test
|
||||
public void testConnection() {
|
||||
String host = "in03-4a578da0f27ce9d.serverless.aws-eu-central-1.cloud.zilliz.com";
|
||||
int port = 443;
|
||||
String token = "d246a77f43a109685596e3c68ecfd359e1cd8b29d35c41d708160f02b97ae623d2ab392738df740d0c77b58ca1bbadfa7c412140";
|
||||
String host = System.getenv().getOrDefault(
|
||||
"MILVUS_HOST",
|
||||
"in03-4a578da0f27ce9d.serverless.aws-eu-central-1.cloud.zilliz.com");
|
||||
int port = Integer.parseInt(System.getenv().getOrDefault("MILVUS_PORT", "443"));
|
||||
String token = System.getenv("MILVUS_TOKEN");
|
||||
Assumptions.assumeTrue(token != null && !token.isBlank(), "MILVUS_TOKEN is required");
|
||||
|
||||
System.out.println("尝试连接 Milvus...");
|
||||
System.out.println("Host: " + host);
|
||||
|
||||
@@ -4,6 +4,7 @@ import com.superbiz.agent.domain.model.SessionContext;
|
||||
import com.superbiz.agent.domain.model.ToolCall;
|
||||
import org.junit.jupiter.api.BeforeEach;
|
||||
import org.junit.jupiter.api.Test;
|
||||
import org.junit.jupiter.api.condition.EnabledIfEnvironmentVariable;
|
||||
import org.springframework.beans.factory.annotation.Autowired;
|
||||
import org.springframework.boot.test.context.SpringBootTest;
|
||||
import org.springframework.test.context.TestPropertySource;
|
||||
@@ -20,10 +21,11 @@ import static org.junit.jupiter.api.Assertions.*;
|
||||
* RedisSessionManager 单元测试
|
||||
*/
|
||||
@SpringBootTest(webEnvironment = SpringBootTest.WebEnvironment.NONE)
|
||||
@EnabledIfEnvironmentVariable(named = "SUPERBIZ_REDIS_PASSWORD", matches = ".+")
|
||||
@TestPropertySource(properties = {
|
||||
"spring.redis.host=119.29.78.52",
|
||||
"spring.redis.port=6379",
|
||||
"spring.redis.password=!Fucker123.."
|
||||
"spring.data.redis.host=${SUPERBIZ_REDIS_HOST:119.29.78.52}",
|
||||
"spring.data.redis.port=${SUPERBIZ_REDIS_PORT:33308}",
|
||||
"spring.data.redis.password=${SUPERBIZ_REDIS_PASSWORD}"
|
||||
})
|
||||
class RedisSessionManagerTest {
|
||||
|
||||
|
||||
Reference in New Issue
Block a user