7.8 KiB
7.8 KiB
Harness 与质量门禁架构
更新日期:2026-07-08
状态:当前可运行架构 + 后续门禁规划
参考历史文档:archive/2026-07-05-legacy/agent-architecture.md
1. 设计目标
Agent 系统的核心风险不是“没有答案”,而是:
- 答案引用了不存在的证据。
- 工具调用失败后仍然编造结论。
- 检索结果相关性不足但被当作强证据。
- 多轮诊断重复检索同一文档,浪费上下文。
- 最终报告无法回放执行过程。
因此当前 MVP 的 Harness 不是单个组件,而是一组约束:
Prompt contract
+ Tool boundary
+ Agent hooks
+ Trace persistence
+ Gatekeeper deterministic validation
+ Verifier / rule evaluation
+ Eval baseline
2. Harness 总图
flowchart TB
Input["User / AIOps input"] --> Prompt["Prompt contract"]
Prompt --> Agent["Planner / Executor / Verifier / Composer"]
Agent --> Tools["Evidence tools"]
Tools --> Invocation["tool_invocation"]
Agent --> StepHook["AgentLoggingHook"]
StepHook --> Step["agent_step"]
Agent --> Session["diagnosis_session"]
Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
Agent --> Gatekeeper
Invocation --> TraceSummary["ToolTraceSummaryService"]
Gatekeeper --> Verifier["chat_verifier"]
TraceSummary --> Verifier
Verifier --> SelfEval["self_evaluation.verifier_evaluation"]
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
AiOpsRule --> AiOpsEval["self_evaluation.aiops_rule_evaluation"]
Session --> TraceAPI["DiagnosisTraceService"]
Step --> TraceAPI
Invocation --> TraceAPI
SelfEval --> TraceAPI
AiOpsEval --> TraceAPI
TraceAPI --> Eval["diagnosis eval / RAG eval"]
3. Prompt Contract
当前 Prompt 按角色拆分:
| Prompt | 用途 |
|---|---|
supervisor-prompt.md |
AIOps Supervisor 调度 Planner / Executor |
planner-prompt.md |
AIOps Planner 规划、再规划、输出告警报告 |
executor-prompt.md |
AIOps Executor 按步骤调用工具 |
chat-planner-prompt.md |
Chat 复杂问题规划 |
chat-executor-prompt.md |
Chat 执行工具并输出 executor_evidence_v2 微观事实 |
chat-verifier-prompt.md |
基于 Gatekeeper 已验真的证据判断 claims 是否可推出 |
chat-composer-prompt.md |
基于 Verifier 允许表达的内容生成最终用户答复 |
Prompt 层当前承担的门禁:
- 禁止凭记忆回答错误码、接口定义、排障步骤。
- 需要外部信息时必须调用工具。
- 工具连续失败或返回空结果时,最终报告必须诚实说明。
- Chat Executor 不允许在窄范围问题中扩展根因、风险或修复建议。
- Chat Verifier 不允许做新检索,只能判断已验真证据是否可推出 claims。
- Chat Composer 不允许补事实,尤其不能把
$.no_evidence表达为“已排除/确认没有”。 - AIOps payload 模式必须聚焦输入告警。
4. Trace Hooks
AgentLoggingHook 是当前 Agent step 可观测性的核心。
sequenceDiagram
autonumber
participant A as Agent
participant H as AgentLoggingHook
participant DB as agent_step
A->>H: before_model(messages, sessionId)
H->>DB: 写入 model_input / step_index / agent_name
A-->>A: LLM 推理
A->>H: after_model(messages, sessionId)
H->>DB: 回填 model_output / thought / has_tool_call / duration / token_count
记录内容:
- 最近输入消息摘要。
- Agent 输出摘要。
- 是否包含 tool call。
- duration。
- token count。
- Verifier 的 JSON 输出摘要。
5. Tool Invocation 门禁
工具调用记录由 ToolInvocationRecorder 和具体工具共同完成。
核心记录:
tool_name
input_params
output_preview
retrieval_layer
l0_match_count
l1_match_count
retrieval_details
-> evidence_refs
relevance_level
dedup_reason
duration_ms
success
error_message
对 lookup_knowledge 的质量约束:
- L0 只作为 hint,不绕过 L1。
- 检索结果归一化为
PRECISE、HIGHLY_RELEVANT、REFERENCE。 - 同 session 内重复文档会被
RetrievedDocTracker去重。 - dedup、no evidence、failed 等状态进入
retrieval_details.evidence_status。 retrieval_details.evidence_refs记录可被 Executor 引用的最小证据文本,格式为raw_path + text。- no-hit / no-evidence 工具结果会生成
raw_path=$.no_evidence的负向证据引用,语义仅限“本次查询未检索到匹配证据”。
6. Gatekeeper 与 Verifier 门禁
Chat Verifier 前置一层 Gatekeeper。Gatekeeper 不调用 LLM,只用代码检查 Executor 输出的证据引用是否真实存在。
flowchart LR
Invocation["tool_invocation"] --> EvidenceRefs["retrieval_details.evidence_refs"]
ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
EvidenceRefs --> Gatekeeper
Gatekeeper --> GateResult["gatekeeper_result"]
Invocation --> Summary["ToolTraceSummaryService"]
Summary --> EvidenceIndex["tool_trace_summary"]
GateResult --> Verifier["chat_verifier"]
ExecutorOutput --> Verifier
EvidenceIndex --> Verifier
Verifier --> Verdict{"verdict"}
Verdict -->|PASS| Composer["chat_composer"]
Composer --> Pass["输出最终答复"]
Verdict -->|LOW_CONFID| Low["补证据或低置信输出"]
Verdict -->|REJECT| Reject["降级输出"]
Gatekeeper 检查:
| 检查 | 失败语义 |
|---|---|
answer_version=executor_evidence_v2 |
非结构化或旧结构输出降为低置信 |
source_invocation_id 真实存在 |
伪造 ID 直接拒绝 |
tool_name 与 invocation 对齐 |
张冠李戴直接拒绝 |
raw_path 存在于 evidence_refs |
无中生有直接拒绝 |
evidence_excerpt 由 evidence_refs[].text 支撑 |
excerpt 编造或错配直接拒绝 |
negative_observation 只能引用 $.no_evidence |
用正向日志证明“没查到”直接拒绝 |
Verifier 输出:
{
"verdict": "PASS|LOW_CONFID|REJECT",
"groundedness_score": 0.8,
"critical_fact_count": 2,
"claim_checks": [],
"facts_checked": [],
"rationale": "..."
}
Verifier 不再逐字核验 excerpt 真伪;这由 Gatekeeper 完成。Verifier 只回答一个问题:claim_text 是否能由已经验真的 evidence_excerpt 推导出来。
结果写入:
diagnosis_session.self_evaluation.verifier_evaluation
其中同时持久化 executor_structured_output、gatekeeper_result、tool_trace_summary 和 composer_output,用于 Trace 回放。
7. AIOps 规则门禁
AIOps 当前不走 Chat Verifier,而是用 AiOpsRuleEvaluationService 做轻量检查。
检查重点:
- 最终报告是否存在。
- payload 模式是否围绕输入告警展开。
- 是否调用证据工具,尤其是
lookup_knowledge、日志、指标。 - 是否把无关活跃告警扩展成主诊断对象。
结果写入:
diagnosis_session.self_evaluation.aiops_rule_evaluation
8. Eval Baseline
当前质量门禁还包括离线评测资产:
| 评测 | 位置 | 作用 |
|---|---|---|
| Diagnosis eval | mvp/eval/ |
检查诊断 trace、报告和证据行为 |
| RAG retrieval eval | eval/rag-retrieval/ |
检查固定检索 query 的召回稳定性 |
| Live RAG acceptance | scripts/eval_rag_live_acceptance.py |
检查运行环境中真实 /api/search/similar 行为 |
9. 后续门禁规划
从旧版设计继承但尚未完整实现的门禁:
- 工具参数 schema 校验。
- 同一工具调用次数上限。
- 工具超时的统一熔断。
- Gatekeeper 规则三层分离:索引层、元数据层、规则实现层。
- Prompt 版本记录和回滚。
- Verifier 对 AIOps 报告的 LLM 级事实校验。
这些应在评测集扩大后逐步加入,避免一次性把诊断流程卡得过死。