feat: add aiops lightweight verifier

This commit is contained in:
aruo
2026-07-05 13:44:30 +08:00
parent 2658742119
commit ed267d753d
15 changed files with 417 additions and 5 deletions
+53
View File
@@ -0,0 +1,53 @@
# AIOps Lightweight Verifier
## What Changed
AIOps now has a deterministic post-run quality gate.
After the final AIOps report is persisted, the service evaluates:
- whether the final report exists and is not trivially short
- whether a payload-targeted report mentions the supplied alert and service
- whether evidence tools such as `lookup_knowledge`, `query_metrics`, or `query_logs` were persisted
The result is stored under:
```text
diagnosis_session.self_evaluation.aiops_rule_evaluation
```
The trace API returns this payload through the existing session self-evaluation field.
## Why Rule-Based First
This is not a full LLM verifier yet.
The first AIOps quality risks are concrete and easy to check with rules:
- Did the report stay focused on the payload?
- Did the run use evidence tools?
- Did the system produce a usable final report?
Rule evaluation is stable, cheap, and easy to explain. It also avoids adding another hidden model call to the AIOps flow before the current trace contract is mature.
## Verdicts
The evaluator emits:
```text
PASS
WARN
FAIL
```
`FAIL` is reserved for critical issues such as a missing or too-short report. Missing payload focus terms or missing evidence tools currently produce `WARN`, because valid reports may use slightly different wording or evidence may be unavailable in a mock/demo environment.
## Interview Answer
If asked why AIOps has a verifier now:
> Chat already has an LLM verifier because the user questions are open-ended. For AIOps, I started with a lighter rule-based verifier because the first quality checks are very concrete: payload focus, evidence coverage, and report completeness. The evaluation is persisted into `self_evaluation`, so the trace can show not only what the Agent did, but also whether the output passed basic quality gates.
If asked why not use the Chat verifier directly:
> AIOps verification is different from Chat verification. It needs to check alert scope, evidence tool coverage, and whether unrelated active alerts were over-expanded. Reusing the Chat verifier directly would blur those semantics. The rule-based evaluator gives us a stable first quality gate; a later AIOps LLM verifier can build on the same trace contract.
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-05
@@ -0,0 +1,59 @@
## Context
Chat diagnosis has a Verifier Agent that writes structured evaluation into `diagnosis_session.self_evaluation`. AIOps currently focuses on payload scoping, evidence tools, and trace persistence, but it has no quality gate that checks whether the final report stayed on target or used evidence.
The next stage should add a low-risk quality gate before considering a full AIOps LLM verifier.
## Goals / Non-Goals
**Goals:**
- Evaluate AIOps final reports with deterministic rules.
- Persist the evaluation under a dedicated `aiops_rule_evaluation` self-evaluation key.
- Keep trace replay able to show whether AIOps output passed, warned, or failed basic quality checks.
- Add focused unit tests without requiring live LLMs or external tools.
**Non-Goals:**
- Do not add an AIOps Verifier Agent yet.
- Do not route/retry AIOps execution based on the evaluation result.
- Do not change `tool_invocation` schema.
- Do not require new database migrations.
## Decisions
### Decision 1: Rule-Based Before LLM Verifier
The first AIOps verifier is a deterministic evaluator, not an LLM agent.
Rationale:
- AIOps quality risks are concrete at this stage: payload focus, evidence coverage, and report presence.
- Rule evaluation is cheap, stable, and easy to explain in an interview.
- A full verifier agent can be added later once AIOps trace expectations are stable.
### Decision 2: Dedicated Self-Evaluation Channel
Persist under `aiops_rule_evaluation` instead of reusing `rule_evaluation` or `verifier_evaluation`.
Rationale:
- `verifier_evaluation` is already associated with Chat's LLM verifier.
- `rule_evaluation` may be used by generic diagnosis evaluation.
- A dedicated key avoids conflating AIOps-specific checks with other evaluation channels.
### Decision 3: Evaluate After Final Report Persistence
Run the evaluator when `persistFinalReport(...)` is called.
Rationale:
- It has access to the final report and session id.
- It can read persisted tool invocations for the same session.
- It does not disturb the Agent execution path.
## Risks / Trade-offs
- [Risk] Rule evaluation can miss semantic hallucinations. -> Mitigation: position it as lightweight AIOps quality gate, not full groundedness verification.
- [Risk] Strict keyword checks may warn on valid reports with different wording. -> Mitigation: use WARN for missing soft signals and FAIL only for critical absence.
- [Risk] Evaluation after report persistence does not trigger retries. -> Mitigation: keep routing unchanged in this phase; later changes can consume the verdict.
@@ -0,0 +1,26 @@
## Why
AIOps now has traceable payload scope control and improved RAG retrieval, but it still lacks a quality gate comparable to Chat's verifier. A lightweight rule-based verifier can check the most important AIOps risks without introducing another LLM agent.
## What Changes
- Add a rule-based AIOps evaluation service that checks final report quality after the AIOps flow completes.
- Persist the evaluation under `diagnosis_session.self_evaluation.aiops_rule_evaluation`.
- Evaluate payload focus, evidence-tool coverage, and basic report completeness.
- Expose the evaluation through the existing trace API self-evaluation payload.
## Capabilities
### New Capabilities
None.
### Modified Capabilities
- `aiops-traceable-diagnosis-entry`: AIOps sessions include a lightweight rule evaluation for trace replay.
## Impact
- Affects AIOps session finalization and trace self-evaluation.
- Does not change AIOps API input, Agent flow topology, tool signatures, or database schema.
- Does not add an LLM verifier agent.
@@ -0,0 +1,22 @@
## ADDED Requirements
### Requirement: AIOps sessions SHALL persist lightweight rule evaluation
When an AIOps final report is persisted, the system SHALL evaluate it with deterministic AIOps-specific quality rules and store the result in session self-evaluation.
#### Scenario: Payload-focused report is evaluated
- **WHEN** an AIOps session has alert payload fields and a final report is persisted
- **THEN** the system SHALL evaluate whether the report mentions the supplied alert and service
- **AND** it SHALL store the result under `self_evaluation.aiops_rule_evaluation`
#### Scenario: Evidence coverage is evaluated
- **WHEN** an AIOps final report is evaluated
- **THEN** the system SHALL check whether evidence tool invocations such as `lookup_knowledge`, `query_metrics`, or `query_logs` were persisted for the session
#### Scenario: Evaluation is traceable
- **WHEN** the diagnosis trace API returns an AIOps session
- **THEN** the session self-evaluation payload SHALL include `aiops_rule_evaluation` when it has been generated
#### Scenario: Evaluation uses stable verdicts
- **WHEN** AIOps rule evaluation completes
- **THEN** it SHALL produce a verdict from `PASS`, `WARN`, or `FAIL`
- **AND** it SHALL include check details and a human-readable rationale
@@ -0,0 +1,19 @@
## 1. Rule Evaluation
- [x] 1.1 Add an AIOps rule evaluation service with PASS/WARN/FAIL verdicts.
- [x] 1.2 Check payload focus, evidence-tool coverage, and report completeness.
## 2. AIOps Integration
- [x] 2.1 Persist AIOps rule evaluation when the final AIOps report is saved.
- [x] 2.2 Make trace summary indicate that AIOps rule evaluation exists.
## 3. Tests And Docs
- [x] 3.1 Add focused unit tests for the evaluator and AIOps integration.
- [x] 3.2 Add interview notes for the lightweight AIOps verifier.
## 4. Verification
- [x] 4.1 Run focused service tests.
- [x] 4.2 Validate the OpenSpec change and review git scope.
@@ -248,7 +248,7 @@ public class ChatController {
if (finalReportOptional.isPresent()) {
String finalReportText = finalReportOptional.get();
logger.info("提取到 Planner 最终报告,长度: {}", finalReportText.length());
aiOpsService.persistFinalReport(sessionId, finalReportText);
aiOpsService.persistFinalReport(sessionId, finalReportText, request);
// 发送分隔线
emitter.send(SseEmitter.event().name("message")
@@ -97,6 +97,7 @@ public class DiagnosisTraceResponse {
private int persistedToolCallCount;
private int returnedToolCallCount;
private boolean hasVerifierEvaluation;
private boolean hasAiOpsRuleEvaluation;
private boolean hasFeedback;
}
}
@@ -0,0 +1,146 @@
package com.superbiz.agent.service;
import com.superbiz.agent.domain.entity.ToolInvocation;
import com.superbiz.agent.dto.AIOpsRequest;
import org.springframework.stereotype.Service;
import java.util.ArrayList;
import java.util.LinkedHashMap;
import java.util.List;
import java.util.Locale;
import java.util.Map;
@Service
public class AiOpsRuleEvaluationService {
public static final String PASS = "PASS";
public static final String WARN = "WARN";
public static final String FAIL = "FAIL";
public Map<String, Object> evaluate(AIOpsRequest request,
String finalReport,
List<ToolInvocation> toolInvocations) {
List<Map<String, Object>> checks = new ArrayList<>();
checks.add(checkReportPresent(finalReport));
checks.add(checkPayloadFocus(request, finalReport));
checks.add(checkEvidenceCoverage(toolInvocations));
String verdict = aggregateVerdict(checks);
Map<String, Object> evaluation = new LinkedHashMap<>();
evaluation.put("verdict", verdict);
evaluation.put("checks", checks);
evaluation.put("rationale", buildRationale(verdict, checks));
evaluation.put("traceability_version", "aiops-rule-v1");
return evaluation;
}
private Map<String, Object> checkReportPresent(String finalReport) {
boolean passed = finalReport != null && finalReport.trim().length() >= 40;
return check(
"final_report_present",
passed ? PASS : FAIL,
passed ? "Final report is present." : "Final report is missing or too short."
);
}
private Map<String, Object> checkPayloadFocus(AIOpsRequest request, String finalReport) {
if (!hasAlertPayload(request)) {
return check("payload_focus", PASS, "No alert payload was supplied; payload focus is not required.");
}
String report = lower(finalReport);
List<String> missing = new ArrayList<>();
if (!contains(report, request.getAlertName())) {
missing.add("alertName");
}
if (!contains(report, request.getService())) {
missing.add("service");
}
if (missing.isEmpty()) {
return check("payload_focus", PASS, "Final report mentions the supplied alert and service.");
}
return check(
"payload_focus",
WARN,
"Final report is missing payload focus terms: " + String.join(", ", missing)
);
}
private Map<String, Object> checkEvidenceCoverage(List<ToolInvocation> toolInvocations) {
List<String> evidenceTools = safeTools(toolInvocations).stream()
.filter(tool -> tool.equals("lookup_knowledge")
|| tool.equals("query_metrics")
|| tool.equals("query_logs"))
.distinct()
.toList();
if (evidenceTools.isEmpty()) {
return check("evidence_tool_coverage", WARN, "No persisted AIOps evidence tool calls were found.");
}
return check(
"evidence_tool_coverage",
PASS,
"Persisted evidence tools: " + String.join(", ", evidenceTools)
);
}
private List<String> safeTools(List<ToolInvocation> toolInvocations) {
if (toolInvocations == null) {
return List.of();
}
return toolInvocations.stream()
.map(ToolInvocation::getToolName)
.filter(name -> name != null && !name.isBlank())
.map(name -> name.trim().toLowerCase(Locale.ROOT))
.toList();
}
private String aggregateVerdict(List<Map<String, Object>> checks) {
boolean hasFail = checks.stream().anyMatch(check -> FAIL.equals(check.get("verdict")));
if (hasFail) {
return FAIL;
}
boolean hasWarn = checks.stream().anyMatch(check -> WARN.equals(check.get("verdict")));
return hasWarn ? WARN : PASS;
}
private String buildRationale(String verdict, List<Map<String, Object>> checks) {
long passCount = checks.stream().filter(check -> PASS.equals(check.get("verdict"))).count();
long warnCount = checks.stream().filter(check -> WARN.equals(check.get("verdict"))).count();
long failCount = checks.stream().filter(check -> FAIL.equals(check.get("verdict"))).count();
return "AIOps rule evaluation %s: pass=%d, warn=%d, fail=%d"
.formatted(verdict, passCount, warnCount, failCount);
}
private Map<String, Object> check(String name, String verdict, String detail) {
Map<String, Object> result = new LinkedHashMap<>();
result.put("name", name);
result.put("verdict", verdict);
result.put("detail", detail);
return result;
}
private boolean hasAlertPayload(AIOpsRequest request) {
if (request == null) {
return false;
}
return !isBlank(request.getAlertName())
|| !isBlank(request.getService())
|| !isBlank(request.getSeverity())
|| !isBlank(request.getDescription())
|| !isBlank(request.getTimeRange());
}
private boolean contains(String lowerText, String value) {
return isBlank(value) || lowerText.contains(value.trim().toLowerCase(Locale.ROOT));
}
private String lower(String value) {
return value == null ? "" : value.toLowerCase(Locale.ROOT);
}
private boolean isBlank(String value) {
return value == null || value.trim().isEmpty();
}
}
@@ -27,6 +27,7 @@ import com.superbiz.agent.config.AiOpsPromptProperties;
import com.superbiz.agent.tool.LookupKnowledgeTool;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.UUID;
@@ -66,6 +67,12 @@ public class AiOpsService {
@Autowired
private ToolInvocationRepository toolInvocationRepository;
@Autowired
private AiOpsRuleEvaluationService aiOpsRuleEvaluationService;
@Autowired
private SelfEvaluationMergeService selfEvaluationMergeService;
/**
* 执行 AI Ops 告警分析流程
*
@@ -170,11 +177,20 @@ public class AiOpsService {
}
public void persistFinalReport(String sessionId, String finalReport) {
persistFinalReport(sessionId, finalReport, null);
}
public void persistFinalReport(String sessionId, String finalReport, AIOpsRequest request) {
if (isBlank(sessionId) || isBlank(finalReport)) {
return;
}
diagnosisSessionRepository.findBySessionId(sessionId.trim()).ifPresent(session -> {
session.setAnswer(finalReport);
List<com.superbiz.agent.domain.entity.ToolInvocation> invocations =
toolInvocationRepository.findBySessionIdOrderByIdAsc(session.getSessionId());
Map<String, Object> evaluation = aiOpsRuleEvaluationService.evaluate(request, finalReport, invocations);
session.setSelfEvaluation(selfEvaluationMergeService.mergeAiOpsRuleEvaluation(
session.getSelfEvaluation(), evaluation));
diagnosisSessionRepository.save(session);
});
}
@@ -115,6 +115,7 @@ public class DiagnosisTraceService {
.persistedToolCallCount(defaultInt(session.getToolCallCount()))
.returnedToolCallCount(toolInvocations.size())
.hasVerifierEvaluation(selfEvaluation != null && selfEvaluation.containsKey("verifier_evaluation"))
.hasAiOpsRuleEvaluation(selfEvaluation != null && selfEvaluation.containsKey("aiops_rule_evaluation"))
.hasFeedback(session.getFeedback() != null && !session.getFeedback().isBlank())
.build();
}
@@ -27,6 +27,10 @@ public class SelfEvaluationMergeService {
return merge(existingJson, "verifier_evaluation", verifierEvaluation);
}
public String mergeAiOpsRuleEvaluation(String existingJson, Map<String, Object> aiOpsRuleEvaluation) {
return merge(existingJson, "aiops_rule_evaluation", aiOpsRuleEvaluation);
}
private String merge(String existingJson, String key, Map<String, Object> value) {
try {
Map<String, Object> root = parseRoot(existingJson);
@@ -44,7 +48,9 @@ public class SelfEvaluationMergeService {
}
Map<String, Object> parsed = objectMapper.readValue(existingJson, MAP_TYPE);
if (parsed.containsKey("rule_evaluation") || parsed.containsKey("verifier_evaluation")) {
if (parsed.containsKey("rule_evaluation")
|| parsed.containsKey("verifier_evaluation")
|| parsed.containsKey("aiops_rule_evaluation")) {
return new LinkedHashMap<>(parsed);
}
@@ -0,0 +1,56 @@
package com.superbiz.agent.service;
import com.superbiz.agent.domain.entity.ToolInvocation;
import com.superbiz.agent.dto.AIOpsRequest;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.Map;
import static org.junit.jupiter.api.Assertions.assertEquals;
class AiOpsRuleEvaluationServiceTest {
private final AiOpsRuleEvaluationService service = new AiOpsRuleEvaluationService();
@Test
void evaluatePassesWhenReportFocusesPayloadAndHasEvidenceTools() {
AIOpsRequest request = new AIOpsRequest();
request.setAlertName("HighCPUUsage");
request.setService("payment-service");
ToolInvocation invocation = ToolInvocation.builder()
.toolName("lookup_knowledge")
.build();
Map<String, Object> evaluation = service.evaluate(
request,
"HighCPUUsage alert on payment-service was diagnosed using metrics and knowledge evidence.",
List.of(invocation)
);
assertEquals("PASS", evaluation.get("verdict"));
}
@Test
void evaluateWarnsWhenEvidenceToolsAreMissing() {
AIOpsRequest request = new AIOpsRequest();
request.setAlertName("HighCPUUsage");
request.setService("payment-service");
Map<String, Object> evaluation = service.evaluate(
request,
"HighCPUUsage alert on payment-service has a likely resource saturation issue.",
List.of()
);
assertEquals("WARN", evaluation.get("verdict"));
}
@Test
void evaluateFailsWhenReportIsMissing() {
Map<String, Object> evaluation = service.evaluate(null, "too short", List.of());
assertEquals("FAIL", evaluation.get("verdict"));
}
}
@@ -28,6 +28,8 @@ class AiOpsServiceTest {
ReflectionTestUtils.setField(service, "diagnosisSessionRepository", diagnosisSessionRepository);
ReflectionTestUtils.setField(service, "agentStepRepository", agentStepRepository);
ReflectionTestUtils.setField(service, "toolInvocationRepository", toolInvocationRepository);
ReflectionTestUtils.setField(service, "aiOpsRuleEvaluationService", new AiOpsRuleEvaluationService());
ReflectionTestUtils.setField(service, "selfEvaluationMergeService", new SelfEvaluationMergeService());
}
@Test
@@ -145,10 +147,12 @@ class AiOpsServiceTest {
.agentFlow("AI_OPS")
.build();
when(diagnosisSessionRepository.findBySessionId("aiops-session-001")).thenReturn(Optional.of(session));
when(toolInvocationRepository.findBySessionIdOrderByIdAsc("aiops-session-001")).thenReturn(List.of());
service.persistFinalReport("aiops-session-001", "# 告警分析报告");
service.persistFinalReport("aiops-session-001", "# 告警分析报告\nHighCPUUsage payment-service analysis with evidence summary.");
assertEquals("# 告警分析报告", session.getAnswer());
assertEquals("# 告警分析报告\nHighCPUUsage payment-service analysis with evidence summary.", session.getAnswer());
assertTrue(session.getSelfEvaluation().contains("aiops_rule_evaluation"));
verify(diagnosisSessionRepository).save(session);
}
@@ -45,7 +45,7 @@ class DiagnosisTraceServiceTest {
.stepCount(2)
.toolCallCount(1)
.answer("restart payment gateway pool")
.selfEvaluation("{\"verifier_evaluation\":{\"verdict\":\"PASS\"}}")
.selfEvaluation("{\"verifier_evaluation\":{\"verdict\":\"PASS\"},\"aiops_rule_evaluation\":{\"verdict\":\"WARN\"}}")
.feedback("useful")
.createdAt(now)
.updatedAt(now)
@@ -103,6 +103,7 @@ class DiagnosisTraceServiceTest {
assertEquals(1, response.getSummary().getPersistedToolCallCount());
assertEquals(1, response.getSummary().getReturnedToolCallCount());
assertTrue(response.getSummary().isHasVerifierEvaluation());
assertTrue(response.getSummary().isHasAiOpsRuleEvaluation());
assertTrue(response.getSummary().isHasFeedback());
}