docs: reorganize MVP interview documentation

This commit is contained in:
aruo
2026-07-05 15:29:28 +08:00
parent b22f2d22c8
commit 88e0a6c944
51 changed files with 4352 additions and 1318 deletions
+44 -31
View File
@@ -1,62 +1,75 @@
# AIOps Query Augmentation
# AIOps 查询增强说明
## What Changed
## 1. 改动是什么
Payload-targeted AIOps prompts now include a deterministic recommended knowledge query.
AIOps 在 `PAYLOAD_TARGETED` 模式下,会从告警 payload 中稳定生成一条推荐知识库检索 query。
The query is built from the non-blank payload fields:
参与拼接的非空字段:
```text
alertName service severity description timeRange userRequest
```
Example:
示例:
```text
HighCPUUsage payment-service P1 CPU usage is above 80% last_15m
HighCPUUsage payment-service P1 CPU 使用率超过 80% last_15m
```
## Why This Matters
AIOps payload fields contain high-value retrieval terms:
- alert name
- service name
- severity
- symptom description
- time range
- operator request
Before this change, the Agent still had to invent its own `lookup_knowledge` query from the full prompt. That can work, but it may omit important terms such as the service name or alert name.
The new prompt makes the retrieval seed explicit:
最终会进入 Prompt:
```text
Recommended lookup_knowledge query: ...
```
## Design Choice
## 2. 为什么重要
This is prompt-level query augmentation, not hidden retrieval.
AIOps payload 里包含高价值检索词:
I intentionally did not call `lookup_knowledge` automatically before the Agent runs. The project values traceability: tool calls should appear as Agent actions, with their inputs and outputs recorded in `tool_invocation`.
- 告警名称。
- 服务名。
- 严重等级。
- 症状描述。
- 时间范围。
- 用户补充请求。
So the design is:
如果完全让 Agent 从长 Prompt 里自己组织检索 query,可能遗漏服务名或告警名。推荐 query 让检索种子更稳定。
## 3. 设计取舍
这是 Prompt 层 query augmentation,不是隐藏检索。
我没有在 Agent 运行前自动调用 `lookup_knowledge`,原因是项目强调可追踪性:工具调用应该由 Agent 显式发起,并记录到 `tool_invocation`。
当前设计:
```text
AIOps payload
-> deterministic recommended retrieval query
-> Agent prompt
-> Agent may call lookup_knowledge explicitly
-> tool_invocation records the real retrieval action
-> Agent 显式调用 lookup_knowledge
-> tool_invocation 记录真实检索行为
```
## Interview Answer
## 4. 面试回答
If asked how AIOps payload improves RAG retrieval:
如果被问:AIOps payload 怎么提升 RAG 检索?
> I do not replace the user query with a broad domain. I extract the high-signal alert terms from the payload, such as alertName, service, severity, symptom, and time range, and put them into a compact recommended lookup query. The Agent still calls `lookup_knowledge` explicitly, so the trace remains auditable, but the retrieval query is less dependent on model improvisation.
```text
我没有把告警 payload 粗暴替换成一个宽泛领域,而是提取 alertName、service、severity、description、timeRange 等高信号字段,拼成推荐的 lookup_knowledge query。
Agent 仍然显式调用工具,所以 trace 仍然能看到真实检索行为,但 query 不再完全依赖模型临场发挥。
```
If asked why not auto-call retrieval:
如果被问:为什么不自动检索?
```text
自动检索会在 Agent 真正决策前制造一份隐藏证据。
这个项目的重点是可观测 Agent 执行,所以我选择 Prompt 层增强:给 Agent 一个更好的 query seed,但不改变工具调用必须显式可追踪的契约。
```
## 5. 后续增强
- 将 recommended query 写入 trace 的结构化字段,便于对比 Agent 实际 query。
- 对 payload 字段加权,例如 alertName/service 权重大于 timeRange。
- 后续接入 Query Transformer 时,保留原始 query、推荐 query、改写 query 三者的可追踪关系。
> Auto-calling retrieval would create hidden evidence before the Agent actually decides to use a tool. For this project, explicit tool invocation is more important because the interview story is about observable Agent execution. Prompt-level augmentation gives the Agent a better query seed without changing the trace contract.