refactor(harness): remove legacy agent architecture
This commit is contained in:
@@ -1,237 +1,54 @@
|
||||
# Agent 编排架构
|
||||
# Diagnosis Agent 执行架构
|
||||
|
||||
**更新日期**:2026-07-08
|
||||
**更新日期**:2026-07-22
|
||||
**状态**:当前可运行架构
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
||||
|
||||
## 1. 设计定位
|
||||
## 1. 单 Agent 原则
|
||||
|
||||
旧版 Agent 架构把系统描述为 Supervisor、Planner、SubAgent、Verifier 的团队协作。当前 MVP 保留这个核心思想,但实现更收敛:
|
||||
当前业务诊断只有一个 `Diagnosis Agent`。它使用框架 `ReactAgent` 完成规划、行动、观察和最终 Draft,但项目不在外层复制 ReAct 状态机,也不使用业务 Graph 或多角色协作链。
|
||||
|
||||
- Chat 链路使用固定顺序工作流:`Planner -> Executor -> Gatekeeper -> Verifier -> Composer`。
|
||||
- AIOps 链路使用 `SupervisorAgent` 调度 `Planner + Executor`,最终由规则评估器做轻量验证。
|
||||
- 当前没有拆分 ExternalApiSubAgent、InternalErrorSubAgent、DatabaseSubAgent;这些作为后续演进方向保留。
|
||||
- 证据工具不直接散落在各个 Agent 里,而是通过 Spring AI ToolCallback / `@Tool` 统一暴露。
|
||||
## 2. 职责
|
||||
|
||||
## 2. 当前 Agent 全景
|
||||
Diagnosis Agent:
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph Chat["Chat diagnosis"]
|
||||
ChatIn["POST /api/chat"] --> ChatService["ChatService"]
|
||||
ChatService --> ChatPlanner["chat_planner"]
|
||||
ChatPlanner --> ChatExecutor["chat_executor"]
|
||||
ChatExecutor --> ChatTools["evidence tools"]
|
||||
ChatTools --> ChatExecutor
|
||||
ChatExecutor --> ChatGatekeeper["ExecutorGatekeeperService"]
|
||||
ChatGatekeeper --> ChatVerifier["chat_verifier"]
|
||||
ChatVerifier --> ChatDecision{"PASS / LOW_CONFID / REJECT"}
|
||||
ChatDecision --> ChatComposer["chat_composer"]
|
||||
ChatComposer --> ChatAnswer["final answer"]
|
||||
end
|
||||
- 接收当前 query 与可选、受限的安全 PreviousTurn。
|
||||
- 自主选择只读 evidence Tool。
|
||||
- 根据 Agent projection 判断是否需要继续查询。
|
||||
- 输出结构化 `DiagnosisDraft`,每条 analysis 绑定 framework `tool_call_id`。
|
||||
- 证据不足时明确限制,不补造事实。
|
||||
|
||||
subgraph AiOps["AIOps diagnosis"]
|
||||
AiOpsIn["POST /api/ai_ops"] --> AiOpsService["AiOpsService"]
|
||||
AiOpsService --> Supervisor["ai_ops_supervisor"]
|
||||
Supervisor --> AiOpsPlanner["planner_agent"]
|
||||
Supervisor --> AiOpsExecutor["executor_agent"]
|
||||
AiOpsPlanner --> AiOpsExecutor
|
||||
AiOpsExecutor --> AiOpsTools["Prometheus / logs / lookup_knowledge"]
|
||||
AiOpsTools --> AiOpsReport["alert report"]
|
||||
AiOpsReport --> AiOpsRule["AiOpsRuleEvaluationService"]
|
||||
end
|
||||
Diagnosis Agent 不负责:
|
||||
|
||||
subgraph Trace["Trace persistence"]
|
||||
ChatSession["chat_session"]
|
||||
Run["diagnosis_run"]
|
||||
Step["agent_step"]
|
||||
Invocation["tool_invocation"]
|
||||
SelfEval["self_evaluation"]
|
||||
end
|
||||
- HTTP/SSE、Session/Run 生命周期和持久化。
|
||||
- 模型/Tool/Token/timeout/cancel 预算。
|
||||
- Tool 参数授权、raw response 投影或证据物理验真。
|
||||
- SemanticGuard 与最终发布决定。
|
||||
|
||||
ChatService --> ChatSession
|
||||
ChatService --> Run
|
||||
ChatPlanner --> Step
|
||||
ChatExecutor --> Step
|
||||
ChatGatekeeper --> SelfEval
|
||||
ChatVerifier --> Step
|
||||
ChatTools --> Invocation
|
||||
ChatDecision --> SelfEval
|
||||
ChatComposer --> Step
|
||||
|
||||
AiOpsService --> ChatSession
|
||||
AiOpsService --> Run
|
||||
AiOpsPlanner --> Step
|
||||
AiOpsExecutor --> Step
|
||||
AiOpsTools --> Invocation
|
||||
AiOpsRule --> SelfEval
|
||||
```
|
||||
|
||||
## 3. Chat 编排
|
||||
|
||||
Chat 复杂诊断采用 `SequentialAgent`,顺序固定:
|
||||
|
||||
```text
|
||||
chat_planner
|
||||
-> chat_executor
|
||||
-> lookup_knowledge / query_logs / query_metrics / date_time
|
||||
-> outputs executor_evidence_v2
|
||||
-> VerifierInputHook / ExecutorGatekeeperService
|
||||
-> validates source_invocation_id / raw_path / evidence_excerpt
|
||||
-> chat_verifier
|
||||
-> judges whether verified evidence can derive claims
|
||||
-> chat_composer
|
||||
-> writes final user-facing answer
|
||||
```
|
||||
|
||||
关键行为:
|
||||
|
||||
| 角色 | 当前职责 | 输出 |
|
||||
|---|---|---|
|
||||
| `chat_planner` | 拆解问题,注入知识域地图和对话历史,给出排查方向 | `planner_plan` |
|
||||
| `chat_executor` | 按计划调用证据工具,抽取带 `source_invocation_id + raw_path + evidence_excerpt` 的微观事实 | `executor_evidence_v2` |
|
||||
| `ExecutorGatekeeperService` | 在 Verifier 前做代码级引用验真,拒绝伪造 ID、错配 raw_path、错配 excerpt | `gatekeeper_result` |
|
||||
| `chat_verifier` | 只判断已验真 evidence excerpt 是否能推出 claim,不做新检索 | `verifier_output` |
|
||||
| `chat_composer` | 只表达 Verifier 允许输出的 claims、缺口和建议,生成最终用户答复 | `composer_output` |
|
||||
|
||||
Chat 链路最多支持两轮验证:
|
||||
## 3. 执行序列
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant C as ChatService
|
||||
participant P as chat_planner
|
||||
participant E as chat_executor
|
||||
participant T as tools
|
||||
participant G as gatekeeper
|
||||
participant V as chat_verifier
|
||||
participant M as chat_composer
|
||||
participant R as diagnosis_run
|
||||
participant App as Chat Application
|
||||
participant Core as Harness Core
|
||||
participant Agent as Diagnosis Agent
|
||||
participant Tool as ACI Tool Boundary
|
||||
participant EG as EvidenceGuard
|
||||
participant SG as SemanticGuard
|
||||
participant Release as Release Policy
|
||||
|
||||
C->>P: 原始问题 + history + retry_context
|
||||
P-->>C: planner_plan
|
||||
C->>E: planner_plan + 上下文
|
||||
E->>T: 调用证据工具
|
||||
T-->>E: 证据结果
|
||||
E-->>C: executor_evidence_v2
|
||||
C->>G: executor_structured_output + tool_invocation.evidence_refs
|
||||
G-->>C: gatekeeper_result
|
||||
C->>V: executor_structured_output + gatekeeper_result + tool_trace_summary
|
||||
V-->>C: PASS / LOW_CONFID / REJECT
|
||||
C->>R: 写入 verifier_evaluation
|
||||
alt LOW_CONFID 且允许补证据
|
||||
C->>P: retry_context: 仅补缺失证据
|
||||
else PASS 或 REJECT
|
||||
C->>M: allowed_claims + missing_info + recommended_actions
|
||||
M-->>C: composer_output
|
||||
C->>R: 保存 Composer 最终 answer
|
||||
end
|
||||
App->>Core: start RunContext
|
||||
App->>Agent: query + safe previous_turn
|
||||
Agent->>Tool: tool name + framework tool_call_id + typed args
|
||||
Tool-->>Agent: bounded agent_result
|
||||
Agent-->>App: DiagnosisDraft
|
||||
App->>EG: Draft + current Run canonical invocations
|
||||
EG-->>App: verified snapshot or deterministic failure
|
||||
App->>SG: query + full Draft + verified snapshot
|
||||
SG-->>App: SUPPORTED / UNSUPPORTED
|
||||
App->>Release: decide public content
|
||||
Release-->>App: report or fixed fallback
|
||||
```
|
||||
|
||||
决策语义:
|
||||
## 4. PreviousTurn
|
||||
|
||||
| Verdict | 行为 |
|
||||
|---|---|
|
||||
| `PASS` | 把 Verifier 允许表达的 claims 交给 Composer 输出 |
|
||||
| `LOW_CONFID` | 如果分数低于阈值且仍有轮次,构造 `retry_context` 补证据;否则输出低置信提示 |
|
||||
| `REJECT` | 输出降级答复,只保留已确认信息和下一步建议 |
|
||||
|
||||
## 4. AIOps 编排
|
||||
|
||||
AIOps 使用 `SupervisorAgent` 调度两个子 Agent:
|
||||
|
||||
```text
|
||||
ai_ops_supervisor
|
||||
-> planner_agent
|
||||
-> executor_agent
|
||||
-> final report
|
||||
-> AiOpsRuleEvaluationService
|
||||
```
|
||||
|
||||
与 Chat 的差异:
|
||||
|
||||
- AIOps 的输入可能是结构化告警 payload。
|
||||
- payload 模式会进入 `PAYLOAD_TARGETED`,最终报告必须聚焦输入告警。
|
||||
- 无 payload 时进入 `AUTO_DISCOVERY`,先通过告警工具发现活跃告警。
|
||||
- 当前 AIOps 不使用 LLM Verifier,而使用轻量规则评估器写入 `self_evaluation.aiops_rule_evaluation`。
|
||||
|
||||
## 5. 工具边界
|
||||
|
||||
当前 Executor 可用工具来自两类:
|
||||
|
||||
```text
|
||||
methodTools
|
||||
-> dateTimeTools
|
||||
-> lookupKnowledgeTool
|
||||
-> queryMetricsTools
|
||||
-> queryLogsTools when mock enabled
|
||||
|
||||
ToolCallbackProvider
|
||||
-> framework-discovered tools
|
||||
```
|
||||
|
||||
工具调用必须写入 `tool_invocation`。其中 `lookup_knowledge` 额外记录:
|
||||
|
||||
- L0/L1 命中数量。
|
||||
- 检索层。
|
||||
- relevance level。
|
||||
- retrieved domains。
|
||||
- dedup reason。
|
||||
|
||||
## 6. Skill / Playbook 流程
|
||||
|
||||
当前 Skill 是诊断流程编排提示,不是事实证据来源。Planner 只能看到 `SkillRegistry.listAll()` 暴露的 name/description 元数据;Executor 才能通过 Spring AI Alibaba 官方 `SkillsAgentHook` 使用 `read_skill` 读取完整 `SKILL.md`。
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Registry["SkillRegistry<br/>active skill metadata"] --> PlannerHook["PlannerSkillMetadataHook"]
|
||||
PlannerHook --> Planner["Planner<br/>metadata only"]
|
||||
Planner --> Plan["planner_plan<br/>selected_skill + steps"]
|
||||
|
||||
Registry --> ExecutorHook["SkillsAgentHook"]
|
||||
ExecutorHook --> ReadSkill["read_skill"]
|
||||
Plan --> Executor["Executor"]
|
||||
Executor --> ReadSkill
|
||||
ReadSkill --> SkillBody["SKILL.md workflow"]
|
||||
SkillBody --> Executor
|
||||
Executor --> EvidenceTools["lookup_knowledge / logs / metrics"]
|
||||
EvidenceTools --> ToolTrace["tool_invocation evidence"]
|
||||
Executor --> Gatekeeper["Gatekeeper"]
|
||||
Gatekeeper --> Verifier["Verifier"]
|
||||
ToolTrace --> Verifier
|
||||
Verifier --> Composer["Composer"]
|
||||
```
|
||||
|
||||
| 角色 | Skill 可见性 | 工具权限 |
|
||||
|---|---|---|
|
||||
| Planner | 只看 skill name / description,并输出 `selected_skill` | 不暴露 `read_skill` |
|
||||
| Executor | 读取 Planner 选中的 skill 正文 | 暴露官方 `read_skill` 和证据工具 |
|
||||
| Gatekeeper | 不看 skill catalog,也不读 skill 正文 | 只读取 Executor 输出和 `tool_invocation.retrieval_details.evidence_refs` |
|
||||
| Verifier | 不看 skill catalog,也不读 skill 正文 | 只读取 Gatekeeper 结果、结构化 claims 和 trace summary |
|
||||
| Composer | 不看 skill catalog,也不读 skill 正文 | 只读取 Verifier 允许表达的内容 |
|
||||
|
||||
## 7. 与旧版设计的差异
|
||||
|
||||
| 旧版设想 | 当前实现 |
|
||||
|---|---|
|
||||
| Supervisor + Planner + 多个专科 SubAgent + Verifier | Chat: Planner + Executor + Gatekeeper + Verifier + Composer;AIOps: Supervisor + Planner + Executor |
|
||||
| ExternalApiSubAgent / InternalErrorSubAgent / DatabaseSubAgent | 暂未拆分,能力通过通用 Executor + 工具 + Prompt 约束实现 |
|
||||
| 每个 SubAgent 专属工具集 | 当前 Executor 持有统一证据工具集合 |
|
||||
| Verifier 支持 PASS / REVISE / REJECT | 当前 Chat Verifier 输出 PASS / LOW_CONFID / REJECT |
|
||||
| Skill 驱动不同诊断流程 | 当前以 Planner 元数据选择 + Executor 读取 playbook 的方式接入 |
|
||||
|
||||
## 8. 后续演进
|
||||
|
||||
当诊断场景和工具复杂度继续上升时,再考虑拆分:
|
||||
|
||||
- `ExternalApiSubAgent`:接口文档、错误码、请求参数、第三方日志。
|
||||
- `DatabaseSubAgent`:连接池、慢 SQL、死锁、索引建议。
|
||||
- `CacheSubAgent`:Redis 超时、连接、热点 key、内存风险。
|
||||
- `GenericDiagnosisSubAgent`:专项 Agent 失败后的兜底。
|
||||
|
||||
拆分前提:
|
||||
|
||||
- 当前 Executor prompt 已难以维护。
|
||||
- 不同故障类型的工具权限明显不同。
|
||||
- Trace 能证明某类问题需要独立的推理策略。
|
||||
- 评测集能覆盖拆分前后的行为差异。
|
||||
PreviousTurn 只来自同一 Session 最近一个 `DIAGNOSIS + SUCCESS + published_result`。Fallback、失败、取消、raw evidence 和完整历史都不能进入下一轮;字段与字节上限由 Harness 配置控制。
|
||||
|
||||
Reference in New Issue
Block a user