feat: add aiops payload query augmentation

This commit is contained in:
aruo
2026-07-05 12:56:20 +08:00
parent 674dd27a48
commit 72a3dbf8c5
8 changed files with 219 additions and 0 deletions
+62
View File
@@ -0,0 +1,62 @@
# AIOps Query Augmentation
## What Changed
Payload-targeted AIOps prompts now include a deterministic recommended knowledge query.
The query is built from the non-blank payload fields:
```text
alertName service severity description timeRange userRequest
```
Example:
```text
HighCPUUsage payment-service P1 CPU usage is above 80% last_15m
```
## Why This Matters
AIOps payload fields contain high-value retrieval terms:
- alert name
- service name
- severity
- symptom description
- time range
- operator request
Before this change, the Agent still had to invent its own `lookup_knowledge` query from the full prompt. That can work, but it may omit important terms such as the service name or alert name.
The new prompt makes the retrieval seed explicit:
```text
Recommended lookup_knowledge query: ...
```
## Design Choice
This is prompt-level query augmentation, not hidden retrieval.
I intentionally did not call `lookup_knowledge` automatically before the Agent runs. The project values traceability: tool calls should appear as Agent actions, with their inputs and outputs recorded in `tool_invocation`.
So the design is:
```text
AIOps payload
-> deterministic recommended retrieval query
-> Agent prompt
-> Agent may call lookup_knowledge explicitly
-> tool_invocation records the real retrieval action
```
## Interview Answer
If asked how AIOps payload improves RAG retrieval:
> I do not replace the user query with a broad domain. I extract the high-signal alert terms from the payload, such as alertName, service, severity, symptom, and time range, and put them into a compact recommended lookup query. The Agent still calls `lookup_knowledge` explicitly, so the trace remains auditable, but the retrieval query is less dependent on model improvisation.
If asked why not auto-call retrieval:
> Auto-calling retrieval would create hidden evidence before the Agent actually decides to use a tool. For this project, explicit tool invocation is more important because the interview story is about observable Agent execution. Prompt-level augmentation gives the Agent a better query seed without changing the trace contract.
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-05
@@ -0,0 +1,53 @@
## Context
`AiOpsService.buildTaskPrompt(...)` already distinguishes two modes:
- `PAYLOAD_TARGETED`: diagnose the supplied alert payload.
- `AUTO_DISCOVERY`: discover active alerts first.
In payload-targeted mode, the prompt includes alert fields, but it does not provide a normalized retrieval query for `lookup_knowledge`. The Agent may still call the tool, but the exact query is left to model behavior.
## Goals / Non-Goals
**Goals:**
- Build a deterministic retrieval query from AIOps payload fields.
- Preserve the original payload fields in the prompt.
- Make the recommended knowledge query visible in prompt text for trace/debugging.
- Keep the Agent responsible for deciding when to call `lookup_knowledge`.
**Non-Goals:**
- Do not add automatic pre-Agent retrieval.
- Do not add verifier logic.
- Do not change tool invocation schema.
- Do not change L0/L1 retrieval internals.
## Decisions
### Decision 1: Prompt-Level Query Augmentation
Add a recommended knowledge query to the payload-targeted prompt instead of calling `lookup_knowledge` directly.
Rationale:
- The current AIOps flow is Agent-driven; tools remain explicit.
- Prompt-level augmentation is low risk and easy to inspect.
- It avoids introducing another hidden retrieval path that would complicate trace semantics.
Alternative considered: automatically call `lookup_knowledge` before invoking the Supervisor. This was rejected because it changes execution behavior and may create evidence that the Agent did not request.
### Decision 2: Preserve Original Query Terms
The generated query includes raw alert/service/symptom terms rather than replacing them with broad domains.
Rationale:
- Alert name, service name, severity, and symptom are high-value retrieval terms.
- Broad categories such as `infrastructure` are useful hints but should not replace concrete terms.
## Risks / Trade-offs
- [Risk] Prompt grows slightly longer. -> Mitigation: keep the query compact and skip blank fields.
- [Risk] The model may ignore the recommendation. -> Mitigation: make the instruction explicit and test prompt inclusion.
- [Risk] Query construction duplicates some summary fields. -> Mitigation: treat the retrieval query as a compact, tool-oriented view of the payload.
@@ -0,0 +1,25 @@
## Why
AIOps payload-targeted diagnosis already scopes the Agent to the supplied alert, but the prompt does not provide a deterministic knowledge-retrieval query. This leaves the Agent to invent lookup terms from the full prompt, which can omit high-value alert fields such as alert name, service, severity, and symptom.
## What Changes
- Build a stable knowledge retrieval query from AIOps payload fields.
- Include the generated retrieval query in payload-targeted prompts as the recommended `lookup_knowledge` query.
- Keep retrieval explicit through the Agent tool; do not automatically call `lookup_knowledge` before the Agent runs.
- Add focused tests for query construction and prompt inclusion.
## Capabilities
### New Capabilities
None.
### Modified Capabilities
- `aiops-traceable-diagnosis-entry`: Payload-targeted AIOps prompts include a deterministic knowledge retrieval query derived from alert payload fields.
## Impact
- Affects `AiOpsService` prompt construction only.
- Does not change the `lookup_knowledge` tool signature, VectorStore retrieval, AIOps API contract, or trace schema.
@@ -0,0 +1,18 @@
## ADDED Requirements
### Requirement: AIOps payload prompts SHALL include a recommended knowledge query
When an AIOps request includes alert payload fields, the system SHALL include a deterministic recommended knowledge retrieval query in the prompt sent to the Agent flow.
#### Scenario: Payload-targeted prompt includes knowledge query
- **WHEN** an AIOps request contains alert name, service, severity, or description
- **THEN** the generated task prompt SHALL include a recommended `lookup_knowledge` query derived from the supplied payload fields
#### Scenario: Query skips blank fields
- **WHEN** some AIOps payload fields are blank
- **THEN** the recommended knowledge query SHALL omit those blank fields
- **AND** it SHALL preserve the non-blank alert-specific terms
#### Scenario: Auto-discovery prompt does not invent payload query
- **WHEN** an AIOps request does not include alert payload fields
- **THEN** the generated task prompt SHALL remain in auto-discovery mode
- **AND** it SHALL not include a payload-derived recommended knowledge query
@@ -0,0 +1,14 @@
## 1. Prompt Query Construction
- [x] 1.1 Add a deterministic AIOps knowledge query builder from payload fields.
- [x] 1.2 Include the recommended query in payload-targeted task prompts.
## 2. Tests And Docs
- [x] 2.1 Add unit tests for query construction and prompt inclusion.
- [x] 2.2 Update interview/RAG notes to reflect AIOps payload query augmentation.
## 3. Verification
- [x] 3.1 Run focused AIOps service tests.
- [x] 3.2 Validate the OpenSpec change and review git scope.
@@ -205,15 +205,33 @@ public class AiOpsService {
|| !isBlank(request.getTimeRange());
}
String buildKnowledgeRetrievalQuery(AIOpsRequest request) {
if (request == null || !hasAlertPayload(request)) {
return "";
}
StringBuilder query = new StringBuilder();
appendQueryTerm(query, request.getAlertName());
appendQueryTerm(query, request.getService());
appendQueryTerm(query, request.getSeverity());
appendQueryTerm(query, request.getDescription());
appendQueryTerm(query, request.getTimeRange());
appendQueryTerm(query, request.getUserRequest());
return query.toString();
}
String buildTaskPrompt(AIOpsRequest request) {
StringBuilder prompt = new StringBuilder();
prompt.append("你是企业级 SRE,接到了自动化告警排查任务。请结合工具调用,执行**规划→执行→再规划**的闭环,并最终按照固定模板输出《告警分析报告》。禁止编造虚假数据,如连续多次查询失败需诚实反馈无法完成的原因。");
prompt.append("\n\n本次告警输入:\n");
prompt.append(buildQuerySummary(request));
if (hasAlertPayload(request)) {
String knowledgeQuery = buildKnowledgeRetrievalQuery(request);
prompt.append("\n\nAIOps scope mode: PAYLOAD_TARGETED\n");
prompt.append("- The request includes an alert payload. Treat the supplied alert payload as the primary and only main diagnosis target.\n");
prompt.append("- The final report must focus on the supplied alert fields such as alertName, service, severity, description, and timeRange.\n");
prompt.append("- Recommended lookup_knowledge query: ").append(knowledgeQuery).append("\n");
prompt.append("- If knowledge-base evidence is needed, call lookup_knowledge with the recommended query or a narrower query that preserves alertName and service.\n");
prompt.append("- You may call queryPrometheusAlerts only to verify whether the supplied alert is still active or to identify related risk/context.\n");
prompt.append("- If queryPrometheusAlerts returns unrelated active alerts, do not create full root-cause or remediation sections for them.\n");
prompt.append("- Mention unrelated active alerts only briefly in a Related Risk section when they help explain the supplied alert.\n");
@@ -316,6 +334,15 @@ public class AiOpsService {
}
}
private void appendQueryTerm(StringBuilder builder, String value) {
if (!isBlank(value)) {
if (!builder.isEmpty()) {
builder.append(' ');
}
builder.append(value.trim());
}
}
private boolean isBlank(String value) {
return value == null || value.trim().isEmpty();
}
@@ -95,11 +95,28 @@ class AiOpsServiceTest {
assertTrue(prompt.contains("queryPrometheusAlerts only to verify"));
assertTrue(prompt.contains("do not create full root-cause or remediation sections"));
assertTrue(prompt.contains("Related Risk"));
assertTrue(prompt.contains("Recommended lookup_knowledge query: HighCPUUsage payment-service P1 CPU usage is above 80% last_15m"));
assertTrue(prompt.contains("preserves alertName and service"));
assertTrue(prompt.contains("告警: HighCPUUsage"));
assertTrue(prompt.contains("服务: payment-service"));
assertFalse(prompt.contains("AIOps scope mode: AUTO_DISCOVERY"));
}
@Test
void buildKnowledgeRetrievalQueryUsesPayloadFieldsAndSkipsBlankValues() {
AIOpsRequest request = new AIOpsRequest();
request.setAlertName("HighLatency");
request.setService(" payment-service ");
request.setSeverity(" ");
request.setDescription("P95 latency above threshold");
request.setTimeRange("last_10m");
request.setUserRequest("结合日志和指标排查");
String query = service.buildKnowledgeRetrievalQuery(request);
assertEquals("HighLatency payment-service P95 latency above threshold last_10m 结合日志和指标排查", query);
}
@Test
void buildTaskPromptUsesAutoDiscoveryModeWhenAlertPayloadIsMissing() {
String nullRequestPrompt = service.buildTaskPrompt(null);
@@ -116,6 +133,7 @@ class AiOpsServiceTest {
assertTrue(userRequestOnlyPrompt.contains("AIOps scope mode: AUTO_DISCOVERY"));
assertTrue(userRequestOnlyPrompt.contains("First call queryPrometheusAlerts"));
assertFalse(userRequestOnlyPrompt.contains("Recommended lookup_knowledge query"));
}
@Test