feat: add aiops payload query augmentation
This commit is contained in:
@@ -0,0 +1,62 @@
|
||||
# AIOps Query Augmentation
|
||||
|
||||
## What Changed
|
||||
|
||||
Payload-targeted AIOps prompts now include a deterministic recommended knowledge query.
|
||||
|
||||
The query is built from the non-blank payload fields:
|
||||
|
||||
```text
|
||||
alertName service severity description timeRange userRequest
|
||||
```
|
||||
|
||||
Example:
|
||||
|
||||
```text
|
||||
HighCPUUsage payment-service P1 CPU usage is above 80% last_15m
|
||||
```
|
||||
|
||||
## Why This Matters
|
||||
|
||||
AIOps payload fields contain high-value retrieval terms:
|
||||
|
||||
- alert name
|
||||
- service name
|
||||
- severity
|
||||
- symptom description
|
||||
- time range
|
||||
- operator request
|
||||
|
||||
Before this change, the Agent still had to invent its own `lookup_knowledge` query from the full prompt. That can work, but it may omit important terms such as the service name or alert name.
|
||||
|
||||
The new prompt makes the retrieval seed explicit:
|
||||
|
||||
```text
|
||||
Recommended lookup_knowledge query: ...
|
||||
```
|
||||
|
||||
## Design Choice
|
||||
|
||||
This is prompt-level query augmentation, not hidden retrieval.
|
||||
|
||||
I intentionally did not call `lookup_knowledge` automatically before the Agent runs. The project values traceability: tool calls should appear as Agent actions, with their inputs and outputs recorded in `tool_invocation`.
|
||||
|
||||
So the design is:
|
||||
|
||||
```text
|
||||
AIOps payload
|
||||
-> deterministic recommended retrieval query
|
||||
-> Agent prompt
|
||||
-> Agent may call lookup_knowledge explicitly
|
||||
-> tool_invocation records the real retrieval action
|
||||
```
|
||||
|
||||
## Interview Answer
|
||||
|
||||
If asked how AIOps payload improves RAG retrieval:
|
||||
|
||||
> I do not replace the user query with a broad domain. I extract the high-signal alert terms from the payload, such as alertName, service, severity, symptom, and time range, and put them into a compact recommended lookup query. The Agent still calls `lookup_knowledge` explicitly, so the trace remains auditable, but the retrieval query is less dependent on model improvisation.
|
||||
|
||||
If asked why not auto-call retrieval:
|
||||
|
||||
> Auto-calling retrieval would create hidden evidence before the Agent actually decides to use a tool. For this project, explicit tool invocation is more important because the interview story is about observable Agent execution. Prompt-level augmentation gives the Agent a better query seed without changing the trace contract.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-05
|
||||
@@ -0,0 +1,53 @@
|
||||
## Context
|
||||
|
||||
`AiOpsService.buildTaskPrompt(...)` already distinguishes two modes:
|
||||
|
||||
- `PAYLOAD_TARGETED`: diagnose the supplied alert payload.
|
||||
- `AUTO_DISCOVERY`: discover active alerts first.
|
||||
|
||||
In payload-targeted mode, the prompt includes alert fields, but it does not provide a normalized retrieval query for `lookup_knowledge`. The Agent may still call the tool, but the exact query is left to model behavior.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Build a deterministic retrieval query from AIOps payload fields.
|
||||
- Preserve the original payload fields in the prompt.
|
||||
- Make the recommended knowledge query visible in prompt text for trace/debugging.
|
||||
- Keep the Agent responsible for deciding when to call `lookup_knowledge`.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not add automatic pre-Agent retrieval.
|
||||
- Do not add verifier logic.
|
||||
- Do not change tool invocation schema.
|
||||
- Do not change L0/L1 retrieval internals.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Decision 1: Prompt-Level Query Augmentation
|
||||
|
||||
Add a recommended knowledge query to the payload-targeted prompt instead of calling `lookup_knowledge` directly.
|
||||
|
||||
Rationale:
|
||||
|
||||
- The current AIOps flow is Agent-driven; tools remain explicit.
|
||||
- Prompt-level augmentation is low risk and easy to inspect.
|
||||
- It avoids introducing another hidden retrieval path that would complicate trace semantics.
|
||||
|
||||
Alternative considered: automatically call `lookup_knowledge` before invoking the Supervisor. This was rejected because it changes execution behavior and may create evidence that the Agent did not request.
|
||||
|
||||
### Decision 2: Preserve Original Query Terms
|
||||
|
||||
The generated query includes raw alert/service/symptom terms rather than replacing them with broad domains.
|
||||
|
||||
Rationale:
|
||||
|
||||
- Alert name, service name, severity, and symptom are high-value retrieval terms.
|
||||
- Broad categories such as `infrastructure` are useful hints but should not replace concrete terms.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Prompt grows slightly longer. -> Mitigation: keep the query compact and skip blank fields.
|
||||
- [Risk] The model may ignore the recommendation. -> Mitigation: make the instruction explicit and test prompt inclusion.
|
||||
- [Risk] Query construction duplicates some summary fields. -> Mitigation: treat the retrieval query as a compact, tool-oriented view of the payload.
|
||||
@@ -0,0 +1,25 @@
|
||||
## Why
|
||||
|
||||
AIOps payload-targeted diagnosis already scopes the Agent to the supplied alert, but the prompt does not provide a deterministic knowledge-retrieval query. This leaves the Agent to invent lookup terms from the full prompt, which can omit high-value alert fields such as alert name, service, severity, and symptom.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Build a stable knowledge retrieval query from AIOps payload fields.
|
||||
- Include the generated retrieval query in payload-targeted prompts as the recommended `lookup_knowledge` query.
|
||||
- Keep retrieval explicit through the Agent tool; do not automatically call `lookup_knowledge` before the Agent runs.
|
||||
- Add focused tests for query construction and prompt inclusion.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `aiops-traceable-diagnosis-entry`: Payload-targeted AIOps prompts include a deterministic knowledge retrieval query derived from alert payload fields.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affects `AiOpsService` prompt construction only.
|
||||
- Does not change the `lookup_knowledge` tool signature, VectorStore retrieval, AIOps API contract, or trace schema.
|
||||
+18
@@ -0,0 +1,18 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: AIOps payload prompts SHALL include a recommended knowledge query
|
||||
When an AIOps request includes alert payload fields, the system SHALL include a deterministic recommended knowledge retrieval query in the prompt sent to the Agent flow.
|
||||
|
||||
#### Scenario: Payload-targeted prompt includes knowledge query
|
||||
- **WHEN** an AIOps request contains alert name, service, severity, or description
|
||||
- **THEN** the generated task prompt SHALL include a recommended `lookup_knowledge` query derived from the supplied payload fields
|
||||
|
||||
#### Scenario: Query skips blank fields
|
||||
- **WHEN** some AIOps payload fields are blank
|
||||
- **THEN** the recommended knowledge query SHALL omit those blank fields
|
||||
- **AND** it SHALL preserve the non-blank alert-specific terms
|
||||
|
||||
#### Scenario: Auto-discovery prompt does not invent payload query
|
||||
- **WHEN** an AIOps request does not include alert payload fields
|
||||
- **THEN** the generated task prompt SHALL remain in auto-discovery mode
|
||||
- **AND** it SHALL not include a payload-derived recommended knowledge query
|
||||
@@ -0,0 +1,14 @@
|
||||
## 1. Prompt Query Construction
|
||||
|
||||
- [x] 1.1 Add a deterministic AIOps knowledge query builder from payload fields.
|
||||
- [x] 1.2 Include the recommended query in payload-targeted task prompts.
|
||||
|
||||
## 2. Tests And Docs
|
||||
|
||||
- [x] 2.1 Add unit tests for query construction and prompt inclusion.
|
||||
- [x] 2.2 Update interview/RAG notes to reflect AIOps payload query augmentation.
|
||||
|
||||
## 3. Verification
|
||||
|
||||
- [x] 3.1 Run focused AIOps service tests.
|
||||
- [x] 3.2 Validate the OpenSpec change and review git scope.
|
||||
@@ -205,15 +205,33 @@ public class AiOpsService {
|
||||
|| !isBlank(request.getTimeRange());
|
||||
}
|
||||
|
||||
String buildKnowledgeRetrievalQuery(AIOpsRequest request) {
|
||||
if (request == null || !hasAlertPayload(request)) {
|
||||
return "";
|
||||
}
|
||||
|
||||
StringBuilder query = new StringBuilder();
|
||||
appendQueryTerm(query, request.getAlertName());
|
||||
appendQueryTerm(query, request.getService());
|
||||
appendQueryTerm(query, request.getSeverity());
|
||||
appendQueryTerm(query, request.getDescription());
|
||||
appendQueryTerm(query, request.getTimeRange());
|
||||
appendQueryTerm(query, request.getUserRequest());
|
||||
return query.toString();
|
||||
}
|
||||
|
||||
String buildTaskPrompt(AIOpsRequest request) {
|
||||
StringBuilder prompt = new StringBuilder();
|
||||
prompt.append("你是企业级 SRE,接到了自动化告警排查任务。请结合工具调用,执行**规划→执行→再规划**的闭环,并最终按照固定模板输出《告警分析报告》。禁止编造虚假数据,如连续多次查询失败需诚实反馈无法完成的原因。");
|
||||
prompt.append("\n\n本次告警输入:\n");
|
||||
prompt.append(buildQuerySummary(request));
|
||||
if (hasAlertPayload(request)) {
|
||||
String knowledgeQuery = buildKnowledgeRetrievalQuery(request);
|
||||
prompt.append("\n\nAIOps scope mode: PAYLOAD_TARGETED\n");
|
||||
prompt.append("- The request includes an alert payload. Treat the supplied alert payload as the primary and only main diagnosis target.\n");
|
||||
prompt.append("- The final report must focus on the supplied alert fields such as alertName, service, severity, description, and timeRange.\n");
|
||||
prompt.append("- Recommended lookup_knowledge query: ").append(knowledgeQuery).append("\n");
|
||||
prompt.append("- If knowledge-base evidence is needed, call lookup_knowledge with the recommended query or a narrower query that preserves alertName and service.\n");
|
||||
prompt.append("- You may call queryPrometheusAlerts only to verify whether the supplied alert is still active or to identify related risk/context.\n");
|
||||
prompt.append("- If queryPrometheusAlerts returns unrelated active alerts, do not create full root-cause or remediation sections for them.\n");
|
||||
prompt.append("- Mention unrelated active alerts only briefly in a Related Risk section when they help explain the supplied alert.\n");
|
||||
@@ -316,6 +334,15 @@ public class AiOpsService {
|
||||
}
|
||||
}
|
||||
|
||||
private void appendQueryTerm(StringBuilder builder, String value) {
|
||||
if (!isBlank(value)) {
|
||||
if (!builder.isEmpty()) {
|
||||
builder.append(' ');
|
||||
}
|
||||
builder.append(value.trim());
|
||||
}
|
||||
}
|
||||
|
||||
private boolean isBlank(String value) {
|
||||
return value == null || value.trim().isEmpty();
|
||||
}
|
||||
|
||||
@@ -95,11 +95,28 @@ class AiOpsServiceTest {
|
||||
assertTrue(prompt.contains("queryPrometheusAlerts only to verify"));
|
||||
assertTrue(prompt.contains("do not create full root-cause or remediation sections"));
|
||||
assertTrue(prompt.contains("Related Risk"));
|
||||
assertTrue(prompt.contains("Recommended lookup_knowledge query: HighCPUUsage payment-service P1 CPU usage is above 80% last_15m"));
|
||||
assertTrue(prompt.contains("preserves alertName and service"));
|
||||
assertTrue(prompt.contains("告警: HighCPUUsage"));
|
||||
assertTrue(prompt.contains("服务: payment-service"));
|
||||
assertFalse(prompt.contains("AIOps scope mode: AUTO_DISCOVERY"));
|
||||
}
|
||||
|
||||
@Test
|
||||
void buildKnowledgeRetrievalQueryUsesPayloadFieldsAndSkipsBlankValues() {
|
||||
AIOpsRequest request = new AIOpsRequest();
|
||||
request.setAlertName("HighLatency");
|
||||
request.setService(" payment-service ");
|
||||
request.setSeverity(" ");
|
||||
request.setDescription("P95 latency above threshold");
|
||||
request.setTimeRange("last_10m");
|
||||
request.setUserRequest("结合日志和指标排查");
|
||||
|
||||
String query = service.buildKnowledgeRetrievalQuery(request);
|
||||
|
||||
assertEquals("HighLatency payment-service P95 latency above threshold last_10m 结合日志和指标排查", query);
|
||||
}
|
||||
|
||||
@Test
|
||||
void buildTaskPromptUsesAutoDiscoveryModeWhenAlertPayloadIsMissing() {
|
||||
String nullRequestPrompt = service.buildTaskPrompt(null);
|
||||
@@ -116,6 +133,7 @@ class AiOpsServiceTest {
|
||||
|
||||
assertTrue(userRequestOnlyPrompt.contains("AIOps scope mode: AUTO_DISCOVERY"));
|
||||
assertTrue(userRequestOnlyPrompt.contains("First call queryPrometheusAlerts"));
|
||||
assertFalse(userRequestOnlyPrompt.contains("Recommended lookup_knowledge query"));
|
||||
}
|
||||
|
||||
@Test
|
||||
|
||||
Reference in New Issue
Block a user