259 lines
8.6 KiB
Markdown
259 lines
8.6 KiB
Markdown
# Harness 与质量门禁架构
|
||
|
||
**更新日期**:2026-07-08
|
||
**状态**:当前可运行架构 + 后续门禁规划
|
||
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
||
|
||
## 1. 设计目标
|
||
|
||
Agent 系统的核心风险不是“没有答案”,而是:
|
||
|
||
- 答案引用了不存在的证据。
|
||
- 工具调用失败后仍然编造结论。
|
||
- 检索结果相关性不足但被当作强证据。
|
||
- 多轮诊断重复检索同一文档,浪费上下文。
|
||
- 最终报告无法回放执行过程。
|
||
|
||
因此当前 MVP 的 Harness 不是单个组件,而是一组约束:
|
||
|
||
```text
|
||
Prompt contract
|
||
+ Tool boundary
|
||
+ Agent hooks
|
||
+ Trace persistence
|
||
+ Gatekeeper deterministic validation
|
||
+ Verifier / rule evaluation
|
||
+ Eval baseline
|
||
```
|
||
|
||
## 2. Harness 总图
|
||
|
||
```mermaid
|
||
flowchart TB
|
||
Input["User / AIOps input"] --> Prompt["Prompt contract"]
|
||
Prompt --> Agent["Planner / Executor / Verifier / Composer"]
|
||
Agent --> Tools["Evidence tools"]
|
||
Tools --> Invocation["tool_invocation"]
|
||
Agent --> StepHook["AgentLoggingHook"]
|
||
StepHook --> Step["agent_step"]
|
||
Agent --> Run["diagnosis_run"]
|
||
|
||
Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
|
||
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
|
||
Agent --> Gatekeeper
|
||
Invocation --> TraceSummary["ToolTraceSummaryService"]
|
||
Gatekeeper --> Verifier["chat_verifier"]
|
||
TraceSummary --> Verifier
|
||
Verifier --> SelfEval["self_evaluation.verifier_evaluation"]
|
||
|
||
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
|
||
AiOpsRule --> AiOpsEval["self_evaluation.aiops_rule_evaluation"]
|
||
|
||
Run --> TraceAPI["DiagnosisTraceService"]
|
||
Step --> TraceAPI
|
||
Invocation --> TraceAPI
|
||
SelfEval --> TraceAPI
|
||
AiOpsEval --> TraceAPI
|
||
|
||
TraceAPI --> Eval["diagnosis eval / RAG eval"]
|
||
```
|
||
|
||
## 3. Prompt Contract
|
||
|
||
当前 Prompt 按角色拆分:
|
||
|
||
| Prompt | 用途 |
|
||
|---|---|
|
||
| `supervisor-prompt.md` | AIOps Supervisor 调度 Planner / Executor |
|
||
| `planner-prompt.md` | AIOps Planner 规划、再规划、输出告警报告 |
|
||
| `executor-prompt.md` | AIOps Executor 按步骤调用工具 |
|
||
| `chat-planner-prompt.md` | Chat 复杂问题规划 |
|
||
| `chat-executor-prompt.md` | Chat 执行工具并输出 `executor_evidence_v2` 微观事实 |
|
||
| `chat-verifier-prompt.md` | 基于 Gatekeeper 已验真的证据判断 claims 是否可推出 |
|
||
| `chat-composer-prompt.md` | 基于 Verifier 允许表达的内容生成最终用户答复 |
|
||
|
||
Prompt 层当前承担的门禁:
|
||
|
||
- 禁止凭记忆回答错误码、接口定义、排障步骤。
|
||
- 需要外部信息时必须调用工具。
|
||
- 工具连续失败或返回空结果时,最终报告必须诚实说明。
|
||
- Chat Executor 不允许在窄范围问题中扩展根因、风险或修复建议。
|
||
- Chat Verifier 不允许做新检索,只能判断已验真证据是否可推出 claims。
|
||
- Chat Composer 不允许补事实,尤其不能把 `$.no_evidence` 表达为“已排除/确认没有”。
|
||
- AIOps payload 模式必须聚焦输入告警。
|
||
|
||
Chat 链路还会在 `verifier_evaluation.prompt_audit` 中持久化紧凑 Prompt 审计快照:
|
||
|
||
```json
|
||
{
|
||
"version": "chat-prompts-v1",
|
||
"prompts": [
|
||
{
|
||
"name": "chat_executor",
|
||
"version": "chat-executor-v2",
|
||
"resource": "prompts/chat-executor-prompt.md"
|
||
}
|
||
]
|
||
}
|
||
```
|
||
|
||
该快照只保存版本和资源路径,不保存完整 Prompt 文本。它用于面试演示、trace 回放和离线 baseline 解释“本次诊断使用了哪套 Prompt 契约”。
|
||
|
||
## 4. Trace Hooks
|
||
|
||
`AgentLoggingHook` 是当前 Agent step 可观测性的核心。
|
||
|
||
```mermaid
|
||
sequenceDiagram
|
||
autonumber
|
||
participant A as Agent
|
||
participant H as AgentLoggingHook
|
||
participant DB as agent_step
|
||
|
||
A->>H: before_model(messages, sessionId, runId)
|
||
H->>DB: 写入 session_id / run_id / model_input / step_index / agent_name
|
||
A-->>A: LLM 推理
|
||
A->>H: after_model(messages, sessionId, runId)
|
||
H->>DB: 回填 model_output / thought / has_tool_call / duration / token_count
|
||
```
|
||
|
||
记录内容:
|
||
|
||
- 最近输入消息摘要。
|
||
- Agent 输出摘要。
|
||
- 是否包含 tool call。
|
||
- duration。
|
||
- token count。
|
||
- Verifier 的 JSON 输出摘要。
|
||
|
||
新写入必须带 `run_id`;`session_id` 仍保留用于粗粒度排查和历史兼容。
|
||
|
||
## 5. Tool Invocation 门禁
|
||
|
||
工具调用记录由 `ToolInvocationRecorder` 和具体工具共同完成。
|
||
|
||
核心记录:
|
||
|
||
```text
|
||
tool_name
|
||
input_params
|
||
output_preview
|
||
retrieval_layer
|
||
l0_match_count
|
||
l1_match_count
|
||
retrieval_details
|
||
-> evidence_refs
|
||
relevance_level
|
||
dedup_reason
|
||
duration_ms
|
||
success
|
||
error_message
|
||
```
|
||
|
||
对 `lookup_knowledge` 的质量约束:
|
||
|
||
- L0 只作为 hint,不绕过 L1。
|
||
- 检索结果归一化为 `PRECISE`、`HIGHLY_RELEVANT`、`REFERENCE`。
|
||
- 同 session 内重复文档会被 `RetrievedDocTracker` 去重。
|
||
- dedup、no evidence、failed 等状态进入 `retrieval_details.evidence_status`。
|
||
- `retrieval_details.evidence_refs` 记录可被 Executor 引用的最小证据文本,格式为 `raw_path + text`。
|
||
- no-hit / no-evidence 工具结果会生成 `raw_path=$.no_evidence` 的负向证据引用,语义仅限“本次查询未检索到匹配证据”。
|
||
|
||
## 6. Gatekeeper 与 Verifier 门禁
|
||
|
||
Chat Verifier 前置一层 Gatekeeper。Gatekeeper 不调用 LLM,只用代码检查 Executor 输出的证据引用是否真实存在。
|
||
|
||
```mermaid
|
||
flowchart LR
|
||
Invocation["tool_invocation"] --> EvidenceRefs["retrieval_details.evidence_refs"]
|
||
ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
|
||
EvidenceRefs --> Gatekeeper
|
||
Gatekeeper --> GateResult["gatekeeper_result"]
|
||
Invocation --> Summary["ToolTraceSummaryService"]
|
||
Summary --> EvidenceIndex["tool_trace_summary"]
|
||
GateResult --> Verifier["chat_verifier"]
|
||
ExecutorOutput --> Verifier
|
||
EvidenceIndex --> Verifier
|
||
Verifier --> Verdict{"verdict"}
|
||
Verdict -->|PASS| Composer["chat_composer"]
|
||
Composer --> Pass["输出最终答复"]
|
||
Verdict -->|LOW_CONFID| Low["补证据或低置信输出"]
|
||
Verdict -->|REJECT| Reject["降级输出"]
|
||
```
|
||
|
||
Gatekeeper 检查:
|
||
|
||
| 检查 | 失败语义 |
|
||
|---|---|
|
||
| `answer_version=executor_evidence_v2` | 非结构化或旧结构输出降为低置信 |
|
||
| `source_invocation_id` 真实存在 | 伪造 ID 直接拒绝 |
|
||
| `tool_name` 与 invocation 对齐 | 张冠李戴直接拒绝 |
|
||
| `raw_path` 存在于 `evidence_refs` | 无中生有直接拒绝 |
|
||
| `evidence_excerpt` 由 `evidence_refs[].text` 支撑 | excerpt 编造或错配直接拒绝 |
|
||
| `negative_observation` 只能引用 `$.no_evidence` | 用正向日志证明“没查到”直接拒绝 |
|
||
|
||
Gatekeeper 审计还会记录 `rule_set_version` 和已启用规则元数据摘要。当前规则元数据来自本地 `gatekeeper-rules.json`,规则执行仍是确定性 Java 代码。
|
||
|
||
Verifier 输出:
|
||
|
||
```json
|
||
{
|
||
"verdict": "PASS|LOW_CONFID|REJECT",
|
||
"groundedness_score": 0.8,
|
||
"critical_fact_count": 2,
|
||
"claim_checks": [],
|
||
"facts_checked": [],
|
||
"rationale": "..."
|
||
}
|
||
```
|
||
|
||
Verifier 不再逐字核验 excerpt 真伪;这由 Gatekeeper 完成。Verifier 只回答一个问题:`claim_text` 是否能由已经验真的 `evidence_excerpt` 推导出来。
|
||
|
||
结果写入:
|
||
|
||
```text
|
||
diagnosis_run.self_evaluation.verifier_evaluation
|
||
```
|
||
|
||
其中同时持久化 `executor_structured_output`、`gatekeeper_result`、`tool_trace_summary`、`prompt_audit` 和 `composer_output`,用于 Trace 回放。
|
||
|
||
## 7. AIOps 规则门禁
|
||
|
||
AIOps 当前不走 Chat Verifier,而是用 `AiOpsRuleEvaluationService` 做轻量检查。
|
||
|
||
检查重点:
|
||
|
||
- 最终报告是否存在。
|
||
- payload 模式是否围绕输入告警展开。
|
||
- 是否调用证据工具,尤其是 `lookup_knowledge`、日志、指标。
|
||
- 是否把无关活跃告警扩展成主诊断对象。
|
||
|
||
结果写入:
|
||
|
||
```text
|
||
diagnosis_run.self_evaluation.aiops_rule_evaluation
|
||
```
|
||
|
||
## 8. Eval Baseline
|
||
|
||
当前质量门禁还包括离线评测资产:
|
||
|
||
| 评测 | 位置 | 作用 |
|
||
|---|---|---|
|
||
| Diagnosis eval | `mvp/eval/` | 检查诊断 trace、报告和证据行为 |
|
||
| RAG retrieval eval | `eval/rag-retrieval/` | 检查固定检索 query 的召回稳定性 |
|
||
| Live RAG acceptance | `scripts/eval_rag_live_acceptance.py` | 检查运行环境中真实 `/api/search/similar` 行为 |
|
||
|
||
## 9. 后续门禁规划
|
||
|
||
从旧版设计继承但尚未完整实现的门禁:
|
||
|
||
- 工具参数 schema 校验。
|
||
- 同一工具调用次数上限。
|
||
- 工具超时的统一熔断。
|
||
- Gatekeeper 规则远程化或三层分离:索引层、元数据层、规则实现层。
|
||
- Prompt 版本回滚和更细粒度变更审计。
|
||
- Verifier 对 AIOps 报告的 LLM 级事实校验。
|
||
|
||
这些应在评测集扩大后逐步加入,避免一次性把诊断流程卡得过死。
|