docs: reorganize MVP interview documentation

This commit is contained in:
aruo
2026-07-05 15:29:28 +08:00
parent b22f2d22c8
commit 88e0a6c944
51 changed files with 4352 additions and 1318 deletions
+205
View File
@@ -0,0 +1,205 @@
# Harness 与质量门禁架构
**更新日期**:2026-07-05
**状态**:当前可运行架构 + 后续门禁规划
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
## 1. 设计目标
Agent 系统的核心风险不是“没有答案”,而是:
- 答案引用了不存在的证据。
- 工具调用失败后仍然编造结论。
- 检索结果相关性不足但被当作强证据。
- 多轮诊断重复检索同一文档,浪费上下文。
- 最终报告无法回放执行过程。
因此当前 MVP 的 Harness 不是单个组件,而是一组约束:
```text
Prompt contract
+ Tool boundary
+ Agent hooks
+ Trace persistence
+ Verifier / rule evaluation
+ Eval baseline
```
## 2. Harness 总图
```mermaid
flowchart TB
Input["User / AIOps input"] --> Prompt["Prompt contract"]
Prompt --> Agent["Planner / Executor / Verifier"]
Agent --> Tools["Evidence tools"]
Tools --> Invocation["tool_invocation"]
Agent --> StepHook["AgentLoggingHook"]
StepHook --> Step["agent_step"]
Agent --> Session["diagnosis_session"]
Invocation --> TraceSummary["ToolTraceSummaryService"]
TraceSummary --> Verifier["chat_verifier"]
Verifier --> SelfEval["self_evaluation.verifier_evaluation"]
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
AiOpsRule --> AiOpsEval["self_evaluation.aiops_rule_evaluation"]
Session --> TraceAPI["DiagnosisTraceService"]
Step --> TraceAPI
Invocation --> TraceAPI
SelfEval --> TraceAPI
AiOpsEval --> TraceAPI
TraceAPI --> Eval["diagnosis eval / RAG eval"]
```
## 3. Prompt Contract
当前 Prompt 按角色拆分:
| Prompt | 用途 |
|---|---|
| `supervisor-prompt.md` | AIOps Supervisor 调度 Planner / Executor |
| `planner-prompt.md` | AIOps Planner 规划、再规划、输出告警报告 |
| `executor-prompt.md` | AIOps Executor 按步骤调用工具 |
| `chat-planner-prompt.md` | Chat 复杂问题规划 |
| `chat-executor-prompt.md` | Chat 执行工具并形成诊断答复 |
| `chat-verifier-prompt.md` | 校验 Executor 答案是否被工具证据支撑 |
Prompt 层当前承担的门禁:
- 禁止凭记忆回答错误码、接口定义、排障步骤。
- 需要外部信息时必须调用工具。
- 工具连续失败或返回空结果时,最终报告必须诚实说明。
- Chat Verifier 不允许做新检索,只能校验已有证据。
- AIOps payload 模式必须聚焦输入告警。
## 4. Trace Hooks
`AgentLoggingHook` 是当前 Agent step 可观测性的核心。
```mermaid
sequenceDiagram
autonumber
participant A as Agent
participant H as AgentLoggingHook
participant DB as agent_step
A->>H: before_model(messages, sessionId)
H->>DB: 写入 model_input / step_index / agent_name
A-->>A: LLM 推理
A->>H: after_model(messages, sessionId)
H->>DB: 回填 model_output / thought / has_tool_call / duration / token_count
```
记录内容:
- 最近输入消息摘要。
- Agent 输出摘要。
- 是否包含 tool call。
- duration。
- token count。
- Verifier 的 JSON 输出摘要。
## 5. Tool Invocation 门禁
工具调用记录由 `ToolInvocationRecorder` 和具体工具共同完成。
核心记录:
```text
tool_name
input_params
output_preview
retrieval_layer
l0_match_count
l1_match_count
retrieval_details
relevance_level
dedup_reason
duration_ms
success
error_message
```
对 `lookup_knowledge` 的质量约束:
- L0 只作为 hint,不绕过 L1。
- 检索结果归一化为 `PRECISE`、`HIGHLY_RELEVANT`、`REFERENCE`。
- 同 session 内重复文档会被 `RetrievedDocTracker` 去重。
- dedup、no evidence、failed 等状态进入 `retrieval_details.evidence_status`。
## 6. Verifier 门禁
Chat Verifier 的输入不是原始工具日志,而是 `ToolTraceSummaryService` 构造的证据索引。
```mermaid
flowchart LR
Invocation["tool_invocation"] --> Summary["ToolTraceSummaryService"]
Summary --> EvidenceIndex["tool_trace_summary"]
EvidenceIndex --> Verifier["chat_verifier"]
ExecutorAnswer["executor_final_answer"] --> Verifier
Verifier --> Verdict{"verdict"}
Verdict -->|PASS| Pass["输出原答案"]
Verdict -->|LOW_CONFID| Low["补证据或低置信输出"]
Verdict -->|REJECT| Reject["降级输出"]
```
Verifier 输出:
```json
{
"verdict": "PASS|LOW_CONFID|REJECT",
"groundedness_score": 0.8,
"critical_fact_count": 2,
"facts_checked": [],
"rationale": "..."
}
```
结果写入:
```text
diagnosis_session.self_evaluation.verifier_evaluation
```
## 7. AIOps 规则门禁
AIOps 当前不走 Chat Verifier,而是用 `AiOpsRuleEvaluationService` 做轻量检查。
检查重点:
- 最终报告是否存在。
- payload 模式是否围绕输入告警展开。
- 是否调用证据工具,尤其是 `lookup_knowledge`、日志、指标。
- 是否把无关活跃告警扩展成主诊断对象。
结果写入:
```text
diagnosis_session.self_evaluation.aiops_rule_evaluation
```
## 8. Eval Baseline
当前质量门禁还包括离线评测资产:
| 评测 | 位置 | 作用 |
|---|---|---|
| Diagnosis eval | `mvp/eval/` | 检查诊断 trace、报告和证据行为 |
| RAG retrieval eval | `eval/rag-retrieval/` | 检查固定检索 query 的召回稳定性 |
| Live RAG acceptance | `scripts/eval_rag_live_acceptance.py` | 检查运行环境中真实 `/api/search/similar` 行为 |
## 9. 后续门禁规划
从旧版设计继承但尚未完整实现的门禁:
- 工具参数 schema 校验。
- 同一工具调用次数上限。
- 工具超时的统一熔断。
- 报告中的数值与工具返回值自动对齐校验。
- Prompt 版本记录和回滚。
- Verifier 对 AIOps 报告的 LLM 级事实校验。
这些应在评测集扩大后逐步加入,避免一次性把诊断流程卡得过死。