refactor(harness): remove legacy agent architecture
This commit is contained in:
@@ -1,258 +1,48 @@
|
||||
# Harness 与质量门禁架构
|
||||
# Harness 与质量门禁
|
||||
|
||||
**更新日期**:2026-07-08
|
||||
**状态**:当前可运行架构 + 后续门禁规划
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
||||
**更新日期**:2026-07-22
|
||||
**状态**:当前可运行架构
|
||||
|
||||
## 1. 设计目标
|
||||
## 1. Harness 定位
|
||||
|
||||
Agent 系统的核心风险不是“没有答案”,而是:
|
||||
Harness 是确定性执行边界,不承担业务推理。它统一管理:
|
||||
|
||||
- 答案引用了不存在的证据。
|
||||
- 工具调用失败后仍然编造结论。
|
||||
- 检索结果相关性不足但被当作强证据。
|
||||
- 多轮诊断重复检索同一文档,浪费上下文。
|
||||
- 最终报告无法回放执行过程。
|
||||
- RunContext、deadline、first-terminal-wins lifecycle 与客户端取消。
|
||||
- 模型调用、Tool 调用、Token、字节数和单 Tool 次数预算。
|
||||
- 类型化 retry policy;Diagnosis Agent 和 Tool 调用不自动重试。
|
||||
- ToolBoundary、canonical invocation 与 Agent projection。
|
||||
- EvidenceGuard、Evidence repair、SemanticGuard 与 Release Policy。
|
||||
- metadata-only durable audit。
|
||||
|
||||
因此当前 MVP 的 Harness 不是单个组件,而是一组约束:
|
||||
## 2. ToolBoundary
|
||||
|
||||
```text
|
||||
Prompt contract
|
||||
+ Tool boundary
|
||||
+ Agent hooks
|
||||
+ Trace persistence
|
||||
+ Gatekeeper deterministic validation
|
||||
+ Verifier / rule evaluation
|
||||
+ Eval baseline
|
||||
framework tool_call_id
|
||||
-> exact Run / schema / authorization / read-only / budget
|
||||
-> backend execution
|
||||
-> raw response -> Redis canonical invocation
|
||||
-> projector -> bounded agent_result
|
||||
-> ToolInvocation durable metadata audit
|
||||
-> Agent observation
|
||||
```
|
||||
|
||||
## 2. Harness 总图
|
||||
Redis canonical invocation 可在 TTL 内保存完整 request/raw_response/agent_result,受独立前缀、容量和 Harness-only 访问保护。Durable audit 只保存 identity、Tool 名、状态、耗时和字节数;audit 写入失败可观测但不改变 canonical Tool 结果。
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
Input["User / AIOps input"] --> Prompt["Prompt contract"]
|
||||
Prompt --> Agent["Planner / Executor / Verifier / Composer"]
|
||||
Agent --> Tools["Evidence tools"]
|
||||
Tools --> Invocation["tool_invocation"]
|
||||
Agent --> StepHook["AgentLoggingHook"]
|
||||
StepHook --> Step["agent_step"]
|
||||
Agent --> Run["diagnosis_run"]
|
||||
## 3. EvidenceGuard
|
||||
|
||||
Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
|
||||
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
Agent --> Gatekeeper
|
||||
Invocation --> TraceSummary["ToolTraceSummaryService"]
|
||||
Gatekeeper --> Verifier["chat_verifier"]
|
||||
TraceSummary --> Verifier
|
||||
Verifier --> SelfEval["self_evaluation.verifier_evaluation"]
|
||||
EvidenceGuard 不调用模型。它校验 Draft schema、analysis ID、当前 Run Tool ownership、READY 状态、evidence status 和每条结论的引用闭包,并生成只包含 Agent projection 的 verified snapshot。
|
||||
|
||||
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
|
||||
AiOpsRule --> AiOpsEval["self_evaluation.aiops_rule_evaluation"]
|
||||
## 4. SemanticGuard
|
||||
|
||||
Run --> TraceAPI["DiagnosisTraceService"]
|
||||
Step --> TraceAPI
|
||||
Invocation --> TraceAPI
|
||||
SelfEval --> TraceAPI
|
||||
AiOpsEval --> TraceAPI
|
||||
SemanticGuard 使用隔离的单轮模型调用,只接收原始 query、完整 Draft 和 verified snapshot。它无 Tool、无记忆、不访问 Redis、不改写报告;技术失败最多按相同输入重试一次,仍失败则安全降级。
|
||||
|
||||
TraceAPI --> Eval["diagnosis eval / RAG eval"]
|
||||
```
|
||||
## 5. Release Policy
|
||||
|
||||
## 3. Prompt Contract
|
||||
- `SUPPORTED`:发布 Diagnosis Agent 原始安全 Draft 的 typed report。
|
||||
- `UNSUPPORTED` 或 evidence failure:发布固定 SAFE_FALLBACK。
|
||||
- technical failure:发布 stable failure,不泄漏内部异常。
|
||||
- cancel/timeout:结束 exact Run,禁止 late content。
|
||||
|
||||
当前 Prompt 按角色拆分:
|
||||
## 6. Audit 安全
|
||||
|
||||
| Prompt | 用途 |
|
||||
|---|---|
|
||||
| `supervisor-prompt.md` | AIOps Supervisor 调度 Planner / Executor |
|
||||
| `planner-prompt.md` | AIOps Planner 规划、再规划、输出告警报告 |
|
||||
| `executor-prompt.md` | AIOps Executor 按步骤调用工具 |
|
||||
| `chat-planner-prompt.md` | Chat 复杂问题规划 |
|
||||
| `chat-executor-prompt.md` | Chat 执行工具并输出 `executor_evidence_v2` 微观事实 |
|
||||
| `chat-verifier-prompt.md` | 基于 Gatekeeper 已验真的证据判断 claims 是否可推出 |
|
||||
| `chat-composer-prompt.md` | 基于 Verifier 允许表达的内容生成最终用户答复 |
|
||||
|
||||
Prompt 层当前承担的门禁:
|
||||
|
||||
- 禁止凭记忆回答错误码、接口定义、排障步骤。
|
||||
- 需要外部信息时必须调用工具。
|
||||
- 工具连续失败或返回空结果时,最终报告必须诚实说明。
|
||||
- Chat Executor 不允许在窄范围问题中扩展根因、风险或修复建议。
|
||||
- Chat Verifier 不允许做新检索,只能判断已验真证据是否可推出 claims。
|
||||
- Chat Composer 不允许补事实,尤其不能把 `$.no_evidence` 表达为“已排除/确认没有”。
|
||||
- AIOps payload 模式必须聚焦输入告警。
|
||||
|
||||
Chat 链路还会在 `verifier_evaluation.prompt_audit` 中持久化紧凑 Prompt 审计快照:
|
||||
|
||||
```json
|
||||
{
|
||||
"version": "chat-prompts-v1",
|
||||
"prompts": [
|
||||
{
|
||||
"name": "chat_executor",
|
||||
"version": "chat-executor-v2",
|
||||
"resource": "prompts/chat-executor-prompt.md"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
该快照只保存版本和资源路径,不保存完整 Prompt 文本。它用于面试演示、trace 回放和离线 baseline 解释“本次诊断使用了哪套 Prompt 契约”。
|
||||
|
||||
## 4. Trace Hooks
|
||||
|
||||
`AgentLoggingHook` 是当前 Agent step 可观测性的核心。
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant A as Agent
|
||||
participant H as AgentLoggingHook
|
||||
participant DB as agent_step
|
||||
|
||||
A->>H: before_model(messages, sessionId, runId)
|
||||
H->>DB: 写入 session_id / run_id / model_input / step_index / agent_name
|
||||
A-->>A: LLM 推理
|
||||
A->>H: after_model(messages, sessionId, runId)
|
||||
H->>DB: 回填 model_output / thought / has_tool_call / duration / token_count
|
||||
```
|
||||
|
||||
记录内容:
|
||||
|
||||
- 最近输入消息摘要。
|
||||
- Agent 输出摘要。
|
||||
- 是否包含 tool call。
|
||||
- duration。
|
||||
- token count。
|
||||
- Verifier 的 JSON 输出摘要。
|
||||
|
||||
新写入必须带 `run_id`;`session_id` 仍保留用于粗粒度排查和历史兼容。
|
||||
|
||||
## 5. Tool Invocation 门禁
|
||||
|
||||
工具调用记录由 `ToolInvocationRecorder` 和具体工具共同完成。
|
||||
|
||||
核心记录:
|
||||
|
||||
```text
|
||||
tool_name
|
||||
input_params
|
||||
output_preview
|
||||
retrieval_layer
|
||||
l0_match_count
|
||||
l1_match_count
|
||||
retrieval_details
|
||||
-> evidence_refs
|
||||
relevance_level
|
||||
dedup_reason
|
||||
duration_ms
|
||||
success
|
||||
error_message
|
||||
```
|
||||
|
||||
对 `lookup_knowledge` 的质量约束:
|
||||
|
||||
- L0 只作为 hint,不绕过 L1。
|
||||
- 检索结果归一化为 `PRECISE`、`HIGHLY_RELEVANT`、`REFERENCE`。
|
||||
- 同 session 内重复文档会被 `RetrievedDocTracker` 去重。
|
||||
- dedup、no evidence、failed 等状态进入 `retrieval_details.evidence_status`。
|
||||
- `retrieval_details.evidence_refs` 记录可被 Executor 引用的最小证据文本,格式为 `raw_path + text`。
|
||||
- no-hit / no-evidence 工具结果会生成 `raw_path=$.no_evidence` 的负向证据引用,语义仅限“本次查询未检索到匹配证据”。
|
||||
|
||||
## 6. Gatekeeper 与 Verifier 门禁
|
||||
|
||||
Chat Verifier 前置一层 Gatekeeper。Gatekeeper 不调用 LLM,只用代码检查 Executor 输出的证据引用是否真实存在。
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Invocation["tool_invocation"] --> EvidenceRefs["retrieval_details.evidence_refs"]
|
||||
ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
EvidenceRefs --> Gatekeeper
|
||||
Gatekeeper --> GateResult["gatekeeper_result"]
|
||||
Invocation --> Summary["ToolTraceSummaryService"]
|
||||
Summary --> EvidenceIndex["tool_trace_summary"]
|
||||
GateResult --> Verifier["chat_verifier"]
|
||||
ExecutorOutput --> Verifier
|
||||
EvidenceIndex --> Verifier
|
||||
Verifier --> Verdict{"verdict"}
|
||||
Verdict -->|PASS| Composer["chat_composer"]
|
||||
Composer --> Pass["输出最终答复"]
|
||||
Verdict -->|LOW_CONFID| Low["补证据或低置信输出"]
|
||||
Verdict -->|REJECT| Reject["降级输出"]
|
||||
```
|
||||
|
||||
Gatekeeper 检查:
|
||||
|
||||
| 检查 | 失败语义 |
|
||||
|---|---|
|
||||
| `answer_version=executor_evidence_v2` | 非结构化或旧结构输出降为低置信 |
|
||||
| `source_invocation_id` 真实存在 | 伪造 ID 直接拒绝 |
|
||||
| `tool_name` 与 invocation 对齐 | 张冠李戴直接拒绝 |
|
||||
| `raw_path` 存在于 `evidence_refs` | 无中生有直接拒绝 |
|
||||
| `evidence_excerpt` 由 `evidence_refs[].text` 支撑 | excerpt 编造或错配直接拒绝 |
|
||||
| `negative_observation` 只能引用 `$.no_evidence` | 用正向日志证明“没查到”直接拒绝 |
|
||||
|
||||
Gatekeeper 审计还会记录 `rule_set_version` 和已启用规则元数据摘要。当前规则元数据来自本地 `gatekeeper-rules.json`,规则执行仍是确定性 Java 代码。
|
||||
|
||||
Verifier 输出:
|
||||
|
||||
```json
|
||||
{
|
||||
"verdict": "PASS|LOW_CONFID|REJECT",
|
||||
"groundedness_score": 0.8,
|
||||
"critical_fact_count": 2,
|
||||
"claim_checks": [],
|
||||
"facts_checked": [],
|
||||
"rationale": "..."
|
||||
}
|
||||
```
|
||||
|
||||
Verifier 不再逐字核验 excerpt 真伪;这由 Gatekeeper 完成。Verifier 只回答一个问题:`claim_text` 是否能由已经验真的 `evidence_excerpt` 推导出来。
|
||||
|
||||
结果写入:
|
||||
|
||||
```text
|
||||
diagnosis_run.self_evaluation.verifier_evaluation
|
||||
```
|
||||
|
||||
其中同时持久化 `executor_structured_output`、`gatekeeper_result`、`tool_trace_summary`、`prompt_audit` 和 `composer_output`,用于 Trace 回放。
|
||||
|
||||
## 7. AIOps 规则门禁
|
||||
|
||||
AIOps 当前不走 Chat Verifier,而是用 `AiOpsRuleEvaluationService` 做轻量检查。
|
||||
|
||||
检查重点:
|
||||
|
||||
- 最终报告是否存在。
|
||||
- payload 模式是否围绕输入告警展开。
|
||||
- 是否调用证据工具,尤其是 `lookup_knowledge`、日志、指标。
|
||||
- 是否把无关活跃告警扩展成主诊断对象。
|
||||
|
||||
结果写入:
|
||||
|
||||
```text
|
||||
diagnosis_run.self_evaluation.aiops_rule_evaluation
|
||||
```
|
||||
|
||||
## 8. Eval Baseline
|
||||
|
||||
当前质量门禁还包括离线评测资产:
|
||||
|
||||
| 评测 | 位置 | 作用 |
|
||||
|---|---|---|
|
||||
| Diagnosis eval | `mvp/eval/` | 检查诊断 trace、报告和证据行为 |
|
||||
| RAG retrieval eval | `eval/rag-retrieval/` | 检查固定检索 query 的召回稳定性 |
|
||||
| Live RAG acceptance | `scripts/eval_rag_live_acceptance.py` | 检查运行环境中真实 `/api/search/similar` 行为 |
|
||||
|
||||
## 9. 后续门禁规划
|
||||
|
||||
从旧版设计继承但尚未完整实现的门禁:
|
||||
|
||||
- 工具参数 schema 校验。
|
||||
- 同一工具调用次数上限。
|
||||
- 工具超时的统一熔断。
|
||||
- Gatekeeper 规则远程化或三层分离:索引层、元数据层、规则实现层。
|
||||
- Prompt 版本回滚和更细粒度变更审计。
|
||||
- Verifier 对 AIOps 报告的 LLM 级事实校验。
|
||||
|
||||
这些应在评测集扩大后逐步加入,避免一次性把诊断流程卡得过死。
|
||||
AgentStep 不保存 Prompt、消息正文、模型正文、Tool arguments 或 Thought。ToolInvocation 不保存完整 request、SQL/日志 query、raw response 或 Agent projection。应用日志不得打印这些字段。
|
||||
|
||||
Reference in New Issue
Block a user