feat(trace): isolate aiops runs

This commit is contained in:
zhuyongxin
2026-07-10 20:57:46 +08:00
parent d928a1968a
commit 78c1477198
7 changed files with 373 additions and 71 deletions
@@ -2,7 +2,7 @@
## sm-flow State
- Checkpoint: Apply / Phase 5 ready
- Checkpoint: Apply / Phase 6 ready
- Scale: complex
- Capability source: sm-flow built-in protocol for context/proposal; grill decisions are recorded from the confirmed user discussion in the issue thread.
- Change slug: `session-run-trace-isolation`
@@ -179,3 +179,13 @@ Audit conclusions:
- Added `CaseLibraryService.createFromRun`, using `diagnosis_run.run_id` as the new automatic `case_library.diagnosis_id`; `createFromSession` remains the legacy session-id path.
- Phase 4 evidence is recorded in `phase-4-evidence.md`.
- Document review follow-up: clarified that the legacy `DiagnosisSession` feedback path only applies when no `diagnosis_run` exists for the session. It returns no bound `runId` and is not the same as latest-run fallback.
## Phase 5 Apply Notes
- Capability source: `openspec-apply-change` + sm-flow apply protocol. `codebase-retrieval` and LSP tools remain unavailable; call-chain confirmation used OpenSpec context, `rg`, targeted file reads, dependency method inspection with `javap`, focused tests, E2E, DB inspection, and logs.
- Changed `/api/ai_ops` to allocate a `runId` before execution and emit a JSON `SseMessage` with `type=metadata`, `sessionId`, and `runId` on SSE event name `message`.
- Changed `AiOpsService` to create one `diagnosis_run` with `agent_flow=AI_OPS` for each valid execution instead of writing new execution state to `diagnosis_session`.
- Changed AIOps execution context propagation to pass `sessionId/runId` through both `RunnableConfig.metadata` and `SessionContextHolder`.
- Changed AIOps final report, metrics, status, and `aiops_rule_evaluation` persistence to update the current `diagnosis_run`.
- Preserved `persistFinalReport(sessionId, finalReport, request)` as a historical `DiagnosisSession` compatibility path.
- Added focused tests for run-scoped final report evaluation, run-scoped metrics, distinct runs under the same AIOps session, and SSE metadata shape.
@@ -0,0 +1,157 @@
# Phase 5 Evidence: AIOps Run Isolation
## Scope
Phase 5 implements AIOps run isolation:
- `/api/ai_ops` allocates a `runId` before execution.
- The SSE stream keeps event name `message` and emits a JSON `SseMessage` with `type=metadata`, `sessionId`, and `runId` before content.
- `AiOpsService` creates `diagnosis_run` rows with `agent_flow=AI_OPS`.
- AIOps Agent hooks and tool recording receive `sessionId + runId` through `RunnableConfig.metadata` and `SessionContextHolder`.
- AIOps status, final report, metrics, and `self_evaluation.aiops_rule_evaluation` write to the current `diagnosis_run`.
- Historical `DiagnosisSession` final-report persistence remains available only through the legacy overload.
## Focused Tests
Focused tests:
```text
mvn -q "-Dtest=AiOpsServiceTest,ChatControllerTest,AgentLoggingHookTest,ToolInvocationRecorderTest,AiOpsRuleEvaluationServiceTest" test
```
Result: passed.
Coverage:
- AIOps final report updates `diagnosis_run.answer` and `diagnosis_run.self_evaluation`.
- Rule evaluation reads `tool_invocation` rows by `run_id`.
- AIOps run metrics count `agent_step` and `tool_invocation` rows by `run_id`.
- The same AIOps `sessionId` can create distinct `runId` values.
- SSE metadata uses `type=metadata` and carries `sessionId/runId`.
- Existing hook and tool-recorder tests cover `runId` propagation into `agent_step` and `tool_invocation`.
## E2E Runtime
Maven startup:
```text
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
```
Captured artifacts:
- `target/e2e/phase5-mvn-20260710-204732.out.log`
- `target/e2e/phase5-mvn-20260710-204732.err.log`
- `target/e2e/phase5-aiops-request.json`
- `target/e2e/phase5-aiops-sse-response.txt`
- `target/e2e/phase5-mvn-20260710-205204.out.log`
- `target/e2e/phase5-mvn-20260710-205204.err.log`
- `target/e2e/phase5-aiops-request-2053.json`
- `target/e2e/phase5-aiops-sse-response-2053.txt`
- `target/e2e/phase5-aiops-trace-2053.json`
The Maven process was stopped after evidence collection.
First E2E attempt exposed an existing AIOps runtime integration bug:
```text
AI Ops 流程失败: mainAgent (ReactAgent) must be provided for supervisor agent
```
Diagnosis result:
- feedback loop: fixed `/api/ai_ops` request with `mvp-demo` profile;
- root cause: current `SupervisorAgent` dependency validates that `mainAgent(ReactAgent)` is set;
- fix: `AiOpsService.buildSupervisorAgent(...)` now sets Planner as `mainAgent` and Executor as sub-agent;
- regression coverage: `AiOpsServiceTest.buildSupervisorAgentSetsPlannerAsMainAgent`.
Successful E2E:
```text
sessionId = e2e-phase5-aiops-codex-20260710-2053
runId = run-84b8c02b-4d1e-4b24-ac88-222c88c8db26
```
SSE response:
```text
event:message
data:{"type":"metadata","data":null,"sessionId":"e2e-phase5-aiops-codex-20260710-2053","runId":"run-84b8c02b-4d1e-4b24-ac88-222c88c8db26"}
...
event:message
data:{"type":"done","data":null,"sessionId":null,"runId":null}
```
Exact trace check:
```text
GET /api/diagnosis/e2e-phase5-aiops-codex-20260710-2053/trace?runId=run-84b8c02b-4d1e-4b24-ac88-222c88c8db26
```
Observed:
```text
code=200
runId=run-84b8c02b-4d1e-4b24-ac88-222c88c8db26
agentFlow=AI_OPS
summary.hasAiOpsRuleEvaluation=true
summary.persistedStepCount=1
summary.persistedToolCallCount=0
```
The E2E model did not call evidence tools. This is captured as rule-evaluation WARN rather than a run-isolation failure.
## Database Inspection
Queried through `scripts/query_mysql.py`.
`diagnosis_run`:
```text
run_id=run-84b8c02b-4d1e-4b24-ac88-222c88c8db26
session_id=e2e-phase5-aiops-codex-20260710-2053
status=SUCCESS
agent_flow=AI_OPS
has_answer=1
has_aiops_eval=1
step_count=1
tool_call_count=0
```
`agent_step` grouped by run:
```text
run-84b8c02b-4d1e-4b24-ac88-222c88c8db26 | step_rows=1
```
`tool_invocation` grouped by run:
```text
(empty; the successful E2E did not invoke evidence tools)
```
Log evidence from `phase5-mvn-20260710-205204.out.log`:
- AIOps request was received with the expected `sessionId` and `runId`.
- `AiOpsService` started analysis and invoked the supervisor agent.
- AIOps orchestration completed and final report extraction ran.
- Exact trace was queried with the same `sessionId + runId`.
## Final Gate
Final Phase 5 gate:
```text
mvn -q clean test-compile
mvn -q "-Dtest=AiOpsServiceTest,ChatControllerTest,AgentLoggingHookTest,ToolInvocationRecorderTest,AiOpsRuleEvaluationServiceTest,DiagnosisTraceServiceTest" test
openspec validate session-run-trace-isolation --strict
git diff --check
```
Result: passed.
## Conclusion
Phase 5 satisfies AIOps run creation, SSE run metadata, run-scoped execution context propagation, run-scoped final report/evaluation/metrics writes, focused tests, and Maven E2E DB/log verification.
@@ -42,12 +42,12 @@
## 5. AIOps Run Isolation
- [ ] 5.1 Change valid `/api/ai_ops` executions to create `diagnosis_run` with `agent_flow=AI_OPS`.
- [ ] 5.2 Expose `runId` in the AIOps SSE-compatible metadata stream while preserving existing report streaming.
- [ ] 5.3 Propagate `runId` through AIOps Agent hooks and tool recording.
- [ ] 5.4 Change AIOps final report, status, counts, and `diagnosis_run.self_evaluation.aiops_rule_evaluation` writes to the current run.
- [ ] 5.5 Add tests for repeated AIOps executions with the same `sessionId` and run-scoped rule evaluation.
- [ ] 5.6 Phase 5 gate: run focused AIOps tests and E2E when needed, inspect DB/logs, update task status, archive phase evidence, and commit before starting Phase 6.
- [x] 5.1 Change valid `/api/ai_ops` executions to create `diagnosis_run` with `agent_flow=AI_OPS`.
- [x] 5.2 Expose `runId` in the AIOps SSE-compatible metadata stream while preserving existing report streaming.
- [x] 5.3 Propagate `runId` through AIOps Agent hooks and tool recording.
- [x] 5.4 Change AIOps final report, status, counts, and `diagnosis_run.self_evaluation.aiops_rule_evaluation` writes to the current run.
- [x] 5.5 Add tests for repeated AIOps executions with the same `sessionId` and run-scoped rule evaluation.
- [x] 5.6 Phase 5 gate: run focused AIOps tests and E2E when needed, inspect DB/logs, update task status, archive phase evidence, and commit before starting Phase 6.
## 6. Demo, Trace UI, Documentation, and Verification
@@ -222,20 +222,21 @@ public class ChatController {
public SseEmitter aiOps(@RequestBody(required = false) AIOpsRequest request) {
SseEmitter emitter = new SseEmitter(600000L); // 10分钟超时(告警分析可能较慢)
String sessionId = aiOpsService.resolveSessionId(request);
String runId = aiOpsService.newRunId();
executor.execute(() -> {
try {
logger.info("收到 AI 智能运维请求 - SessionId: {}, 启动多 Agent 协作流程", sessionId);
logger.info("收到 AI 智能运维请求 - SessionId: {}, RunId: {}, 启动多 Agent 协作流程", sessionId, runId);
ChatModel chatModel = chatService.getChatModel();
ToolCallback[] toolCallbacks = tools != null ? tools.getToolCallbacks() : new ToolCallback[0];
emitter.send(SseEmitter.event().name("message").data(SseMessage.session(sessionId), MediaType.APPLICATION_JSON));
emitter.send(SseEmitter.event().name("message").data(SseMessage.metadata(sessionId, runId), MediaType.APPLICATION_JSON));
emitter.send(SseEmitter.event().name("message").data(SseMessage.content("正在读取告警并拆解任务...\n")));
// 调用 AiOpsService 执行分析流程
Optional<OverAllState> overAllStateOptional = aiOpsService.executeAiOpsAnalysis(chatModel, toolCallbacks, request, sessionId);
Optional<OverAllState> overAllStateOptional = aiOpsService.executeAiOpsAnalysis(chatModel, toolCallbacks, request, sessionId, runId);
if (overAllStateOptional.isEmpty()) {
emitter.send(SseEmitter.event().name("message")
@@ -254,7 +255,7 @@ public class ChatController {
if (finalReportOptional.isPresent()) {
String finalReportText = finalReportOptional.get();
logger.info("提取到 Planner 最终报告,长度: {}", finalReportText.length());
aiOpsService.persistFinalReport(sessionId, finalReportText, request);
aiOpsService.persistFinalReport(sessionId, runId, finalReportText, request);
// 发送分隔线
emitter.send(SseEmitter.event().name("message")
@@ -451,8 +452,10 @@ public class ChatController {
@Setter
@Getter
public static class SseMessage {
private String type; // content: 内容块, error: 错误, done: 完成
private String type; // metadata: 元数据, content: 内容块, error: 错误, done: 完成
private String data;
private String sessionId;
private String runId;
public static SseMessage content(String data) {
SseMessage message = new SseMessage();
@@ -468,6 +471,14 @@ public class ChatController {
return message;
}
public static SseMessage metadata(String sessionId, String runId) {
SseMessage message = new SseMessage();
message.setType("metadata");
message.setSessionId(sessionId);
message.setRunId(runId);
return message;
}
public static SseMessage error(String errorMessage) {
SseMessage message = new SseMessage();
message.setType("error");
@@ -1,6 +1,7 @@
package com.superbiz.agent.service;
import org.springframework.ai.chat.model.ChatModel;
import com.alibaba.cloud.ai.graph.RunnableConfig;
import com.alibaba.cloud.ai.graph.OverAllState;
import com.alibaba.cloud.ai.graph.agent.ReactAgent;
import com.alibaba.cloud.ai.graph.agent.flow.agent.SupervisorAgent;
@@ -13,11 +14,13 @@ import com.superbiz.agent.agent.tool.InternalDocsTools;
import com.superbiz.agent.agent.tool.QueryLogsTools;
import com.superbiz.agent.agent.tool.QueryMetricsTools;
import com.superbiz.agent.domain.entity.AgentStep;
import com.superbiz.agent.domain.entity.DiagnosisRun;
import com.superbiz.agent.domain.entity.DiagnosisSession;
import com.superbiz.agent.dto.AIOpsRequest;
import com.superbiz.agent.hook.AgentLoggingHook;
import com.superbiz.agent.hook.PlannerSkillMetadataHook;
import com.superbiz.agent.repository.AgentStepRepository;
import com.superbiz.agent.repository.DiagnosisRunRepository;
import com.superbiz.agent.repository.DiagnosisSessionRepository;
import com.superbiz.agent.repository.ToolInvocationRepository;
import com.superbiz.agent.util.SessionContextHolder;
@@ -65,6 +68,9 @@ public class AiOpsService {
@Autowired
private DiagnosisSessionRepository diagnosisSessionRepository;
@Autowired
private DiagnosisRunRepository diagnosisRunRepository;
@Autowired
private AgentStepRepository agentStepRepository;
@@ -89,46 +95,48 @@ public class AiOpsService {
* @throws GraphRunnerException Agent 图执行失败时抛出
*/
public Optional<OverAllState> executeAiOpsAnalysis(ChatModel chatModel, ToolCallback[] toolCallbacks) throws GraphRunnerException {
return executeAiOpsAnalysis(chatModel, toolCallbacks, null, resolveSessionId(null));
return executeAiOpsAnalysis(chatModel, toolCallbacks, null, resolveSessionId(null), newRunId());
}
public Optional<OverAllState> executeAiOpsAnalysis(ChatModel chatModel, ToolCallback[] toolCallbacks,
AIOpsRequest request, String sessionId) throws GraphRunnerException {
logger.info("Starting AI Ops multi-agent analysis");
return executeAiOpsAnalysis(chatModel, toolCallbacks, request, sessionId, newRunId());
}
public Optional<OverAllState> executeAiOpsAnalysis(ChatModel chatModel, ToolCallback[] toolCallbacks,
AIOpsRequest request, String sessionId, String runId) throws GraphRunnerException {
logger.info("Starting AI Ops multi-agent analysis");
String resolvedSessionId = isBlank(sessionId) ? resolveSessionId(request) : sessionId.trim();
String resolvedRunId = isBlank(runId) ? newRunId() : runId.trim();
long startTime = System.currentTimeMillis();
DiagnosisSession session = startDiagnosisSession(resolvedSessionId, request);
diagnosisSessionRepository.save(session);
DiagnosisRun run = startDiagnosisRun(resolvedSessionId, resolvedRunId, request);
// 让工具调用、Hook 和知识库检索能够拿到当前诊断会话 ID。
SessionContextHolder.setSessionId(resolvedSessionId);
// 让工具调用、Hook 和知识库检索能够拿到当前诊断会话和运行 ID。
SessionContextHolder.setContext(resolvedSessionId, resolvedRunId);
try {
ReactAgent plannerAgent = buildPlannerAgent(chatModel, toolCallbacks);
ReactAgent executorAgent = buildExecutorAgent(chatModel, toolCallbacks);
SupervisorAgent supervisorAgent = SupervisorAgent.builder()
.name("ai_ops_supervisor")
.description("Coordinates Planner and Executor agents")
.model(chatModel)
.systemPrompt(promptProperties.getSupervisor())
.subAgents(List.of(plannerAgent, executorAgent))
.build();
SupervisorAgent supervisorAgent = buildSupervisorAgent(chatModel, plannerAgent, executorAgent);
String taskPrompt = buildTaskPrompt(request);
RunnableConfig config = RunnableConfig.builder()
.addMetadata("sessionId", resolvedSessionId)
.addMetadata("runId", resolvedRunId)
.build();
logger.info("Invoking AI Ops supervisor agent");
Optional<OverAllState> stateOptional = supervisorAgent.invoke(taskPrompt);
Optional<OverAllState> stateOptional = supervisorAgent.invoke(taskPrompt, config);
long duration = System.currentTimeMillis() - startTime;
session.setStatus(stateOptional.isPresent() ? "SUCCESS" : "FAILED");
session.setTotalDurationMs((int) duration);
backfillSessionMetrics(session);
diagnosisSessionRepository.save(session);
run.setStatus(stateOptional.isPresent() ? "SUCCESS" : "FAILED");
run.setTotalDurationMs((int) duration);
backfillRunMetrics(run);
diagnosisRunRepository.save(run);
if (stateOptional.isPresent()) {
OverAllState state = stateOptional.get();
@@ -139,8 +147,10 @@ public class AiOpsService {
return stateOptional;
} catch (Exception e) {
session.setStatus("FAILED");
diagnosisSessionRepository.save(session);
run.setStatus("FAILED");
run.setTotalDurationMs((int) (System.currentTimeMillis() - startTime));
backfillRunMetrics(run);
diagnosisRunRepository.save(run);
throw e;
} finally {
SessionContextHolder.clear();
@@ -177,11 +187,34 @@ public class AiOpsService {
return UUID.randomUUID().toString();
}
public String newRunId() {
return "run-" + UUID.randomUUID();
}
public void persistFinalReport(String sessionId, String finalReport) {
persistFinalReport(sessionId, finalReport, null);
}
public void persistFinalReport(String sessionId, String finalReport, AIOpsRequest request) {
persistLegacyFinalReport(sessionId, finalReport, request);
}
public void persistFinalReport(String sessionId, String runId, String finalReport, AIOpsRequest request) {
if (isBlank(sessionId) || isBlank(runId) || isBlank(finalReport)) {
return;
}
diagnosisRunRepository.findBySessionIdAndRunId(sessionId.trim(), runId.trim()).ifPresent(run -> {
run.setAnswer(finalReport);
List<com.superbiz.agent.domain.entity.ToolInvocation> invocations =
toolInvocationRepository.findByRunIdOrderByIdAsc(run.getRunId());
Map<String, Object> evaluation = aiOpsRuleEvaluationService.evaluate(request, finalReport, invocations);
run.setSelfEvaluation(selfEvaluationMergeService.mergeAiOpsRuleEvaluation(
run.getSelfEvaluation(), evaluation));
diagnosisRunRepository.save(run);
});
}
private void persistLegacyFinalReport(String sessionId, String finalReport, AIOpsRequest request) {
if (isBlank(sessionId) || isBlank(finalReport)) {
return;
}
@@ -261,21 +294,15 @@ public class AiOpsService {
return prompt.toString();
}
private DiagnosisSession startDiagnosisSession(String sessionId, AIOpsRequest request) {
DiagnosisSession session = diagnosisSessionRepository.findBySessionId(sessionId)
.orElseGet(() -> DiagnosisSession.builder()
.sessionId(sessionId)
.agentFlow("AI_OPS")
.build());
session.setQuery(buildQuerySummary(request));
session.setStatus("RUNNING");
session.setAgentFlow("AI_OPS");
session.setAnswer(null);
session.setTotalDurationMs(null);
session.setTotalTokenCount(null);
session.setStepCount(null);
session.setToolCallCount(null);
return session;
private DiagnosisRun startDiagnosisRun(String sessionId, String runId, AIOpsRequest request) {
DiagnosisRun run = DiagnosisRun.builder()
.runId(runId)
.sessionId(sessionId)
.query(buildQuerySummary(request))
.status("RUNNING")
.agentFlow("AI_OPS")
.build();
return diagnosisRunRepository.save(run);
}
/**
@@ -334,16 +361,27 @@ public class AiOpsService {
loggingHook,
SkillsAgentHook.builder()
.skillRegistry(skillRegistry)
.build()
.build()
};
}
SupervisorAgent buildSupervisorAgent(ChatModel chatModel, ReactAgent plannerAgent, ReactAgent executorAgent) {
return SupervisorAgent.builder()
.name("ai_ops_supervisor")
.description("Coordinates Planner and Executor agents")
.model(chatModel)
.mainAgent(plannerAgent)
.systemPrompt(promptProperties.getSupervisor())
.subAgents(List.of(executorAgent))
.build();
}
/**
* 从 agent_step 和 tool_invocation 回填 diagnosis_session 的汇总指标。
* 从 agent_step 和 tool_invocation 回填 diagnosis_run 的汇总指标。
*/
private void backfillSessionMetrics(DiagnosisSession session) {
private void backfillRunMetrics(DiagnosisRun run) {
try {
List<AgentStep> steps = agentStepRepository.findBySessionIdOrderByStepIndex(session.getSessionId());
List<AgentStep> steps = agentStepRepository.findByRunIdOrderByStepIndex(run.getRunId());
int totalTokens = 0;
int stepCount = 0;
@@ -351,12 +389,12 @@ public class AiOpsService {
stepCount++;
if (s.getTokenCount() != null) totalTokens += s.getTokenCount();
}
long toolCallCount = toolInvocationRepository.countBySessionId(session.getSessionId());
session.setTotalTokenCount(totalTokens);
session.setStepCount(stepCount);
session.setToolCallCount(Math.toIntExact(toolCallCount));
long toolCallCount = toolInvocationRepository.countByRunId(run.getRunId());
run.setTotalTokenCount(totalTokens);
run.setStepCount(stepCount);
run.setToolCallCount(Math.toIntExact(toolCallCount));
} catch (Exception e) {
logger.warn("Failed to backfill AI Ops session metrics, sessionId={}", session.getSessionId(), e);
logger.warn("Failed to backfill AI Ops run metrics, runId={}", run.getRunId(), e);
}
}
@@ -29,4 +29,13 @@ class ChatControllerTest {
assertEquals("问题内容不能为空", body.getErrorMessage());
verifyNoInteractions(chatService);
}
@Test
void aiOpsMetadataMessageCarriesSessionAndRunId() {
ChatController.SseMessage message = ChatController.SseMessage.metadata("session-1", "run-1");
assertEquals("metadata", message.getType());
assertEquals("session-1", message.getSessionId());
assertEquals("run-1", message.getRunId());
}
}
@@ -1,13 +1,20 @@
package com.superbiz.agent.service;
import com.alibaba.cloud.ai.graph.agent.ReactAgent;
import com.alibaba.cloud.ai.graph.agent.flow.agent.SupervisorAgent;
import com.superbiz.agent.config.AiOpsPromptProperties;
import com.superbiz.agent.domain.entity.AgentStep;
import com.superbiz.agent.domain.entity.DiagnosisRun;
import com.superbiz.agent.domain.entity.DiagnosisSession;
import com.superbiz.agent.domain.entity.ToolInvocation;
import com.superbiz.agent.dto.AIOpsRequest;
import com.superbiz.agent.repository.AgentStepRepository;
import com.superbiz.agent.repository.DiagnosisRunRepository;
import com.superbiz.agent.repository.DiagnosisSessionRepository;
import com.superbiz.agent.repository.ToolInvocationRepository;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import org.springframework.ai.chat.model.ChatModel;
import org.springframework.test.util.ReflectionTestUtils;
import java.util.Optional;
@@ -19,17 +26,22 @@ import static org.mockito.Mockito.*;
class AiOpsServiceTest {
private final DiagnosisSessionRepository diagnosisSessionRepository = mock(DiagnosisSessionRepository.class);
private final DiagnosisRunRepository diagnosisRunRepository = mock(DiagnosisRunRepository.class);
private final AgentStepRepository agentStepRepository = mock(AgentStepRepository.class);
private final ToolInvocationRepository toolInvocationRepository = mock(ToolInvocationRepository.class);
private final AiOpsPromptProperties promptProperties = mock(AiOpsPromptProperties.class);
private final AiOpsService service = new AiOpsService();
@BeforeEach
void setUp() {
ReflectionTestUtils.setField(service, "diagnosisSessionRepository", diagnosisSessionRepository);
ReflectionTestUtils.setField(service, "diagnosisRunRepository", diagnosisRunRepository);
ReflectionTestUtils.setField(service, "agentStepRepository", agentStepRepository);
ReflectionTestUtils.setField(service, "toolInvocationRepository", toolInvocationRepository);
ReflectionTestUtils.setField(service, "aiOpsRuleEvaluationService", new AiOpsRuleEvaluationService());
ReflectionTestUtils.setField(service, "selfEvaluationMergeService", new SelfEvaluationMergeService());
ReflectionTestUtils.setField(service, "promptProperties", promptProperties);
when(promptProperties.getSupervisor()).thenReturn("supervisor prompt");
}
@Test
@@ -139,19 +151,48 @@ class AiOpsServiceTest {
}
@Test
void persistFinalReportUpdatesDiagnosisSessionAnswer() {
DiagnosisSession session = DiagnosisSession.builder()
void persistFinalReportUpdatesDiagnosisRunAnswerAndEvaluationByRun() {
DiagnosisRun run = DiagnosisRun.builder()
.sessionId("aiops-session-001")
.runId("run-aiops-001")
.query("AI Ops alert analysis")
.status("SUCCESS")
.agentFlow("AI_OPS")
.build();
when(diagnosisSessionRepository.findBySessionId("aiops-session-001")).thenReturn(Optional.of(session));
when(toolInvocationRepository.findBySessionIdOrderByIdAsc("aiops-session-001")).thenReturn(List.of());
ToolInvocation invocation = ToolInvocation.builder()
.sessionId("aiops-session-001")
.runId("run-aiops-001")
.toolName("query_logs")
.success(true)
.build();
when(diagnosisRunRepository.findBySessionIdAndRunId("aiops-session-001", "run-aiops-001"))
.thenReturn(Optional.of(run));
when(toolInvocationRepository.findByRunIdOrderByIdAsc("run-aiops-001")).thenReturn(List.of(invocation));
service.persistFinalReport("aiops-session-001", "# 告警分析报告\nHighCPUUsage payment-service analysis with evidence summary.");
service.persistFinalReport("aiops-session-001", "run-aiops-001",
"# 告警分析报告\nHighCPUUsage payment-service analysis with evidence summary.", null);
assertEquals("# 告警分析报告\nHighCPUUsage payment-service analysis with evidence summary.", session.getAnswer());
assertEquals("# 告警分析报告\nHighCPUUsage payment-service analysis with evidence summary.", run.getAnswer());
assertTrue(run.getSelfEvaluation().contains("aiops_rule_evaluation"));
verify(toolInvocationRepository).findByRunIdOrderByIdAsc("run-aiops-001");
verify(diagnosisRunRepository).save(run);
verify(diagnosisSessionRepository, never()).save(any());
}
@Test
void legacyPersistFinalReportStillUpdatesHistoricalDiagnosisSession() {
DiagnosisSession session = DiagnosisSession.builder()
.sessionId("legacy-aiops-session")
.query("AI Ops alert analysis")
.status("SUCCESS")
.agentFlow("AI_OPS")
.build();
when(diagnosisSessionRepository.findBySessionId("legacy-aiops-session")).thenReturn(Optional.of(session));
when(toolInvocationRepository.findBySessionIdOrderByIdAsc("legacy-aiops-session")).thenReturn(List.of());
service.persistFinalReport("legacy-aiops-session", "# 告警分析报告\nLegacy analysis with evidence summary.");
assertEquals("# 告警分析报告\nLegacy analysis with evidence summary.", session.getAnswer());
assertTrue(session.getSelfEvaluation().contains("aiops_rule_evaluation"));
verify(diagnosisSessionRepository).save(session);
}
@@ -160,32 +201,68 @@ class AiOpsServiceTest {
void persistFinalReportSkipsBlankInput() {
service.persistFinalReport("aiops-session-001", " ");
verifyNoInteractions(diagnosisSessionRepository);
verifyNoInteractions(diagnosisSessionRepository, diagnosisRunRepository);
}
@Test
void backfillSessionMetricsUsesRealToolInvocationCount() {
DiagnosisSession session = DiagnosisSession.builder()
void backfillRunMetricsUsesRunScopedRows() {
DiagnosisRun run = DiagnosisRun.builder()
.sessionId("aiops-session-002")
.runId("run-aiops-002")
.build();
AgentStep stepWithTool = AgentStep.builder()
.sessionId("aiops-session-002")
.runId("run-aiops-002")
.hasToolCall(true)
.tokenCount(10)
.build();
AgentStep stepWithoutTool = AgentStep.builder()
.sessionId("aiops-session-002")
.runId("run-aiops-002")
.hasToolCall(false)
.tokenCount(20)
.build();
when(agentStepRepository.findBySessionIdOrderByStepIndex("aiops-session-002"))
when(agentStepRepository.findByRunIdOrderByStepIndex("run-aiops-002"))
.thenReturn(List.of(stepWithTool, stepWithoutTool));
when(toolInvocationRepository.countBySessionId("aiops-session-002")).thenReturn(11L);
when(toolInvocationRepository.countByRunId("run-aiops-002")).thenReturn(11L);
ReflectionTestUtils.invokeMethod(service, "backfillSessionMetrics", session);
ReflectionTestUtils.invokeMethod(service, "backfillRunMetrics", run);
assertEquals(2, session.getStepCount());
assertEquals(30, session.getTotalTokenCount());
assertEquals(11, session.getToolCallCount());
assertEquals(2, run.getStepCount());
assertEquals(30, run.getTotalTokenCount());
assertEquals(11, run.getToolCallCount());
verify(agentStepRepository).findByRunIdOrderByStepIndex("run-aiops-002");
verify(toolInvocationRepository).countByRunId("run-aiops-002");
}
@Test
void sameAiOpsSessionCanStartDistinctRuns() {
AIOpsRequest request = new AIOpsRequest();
request.setAlertName("HighCPUUsage");
ReflectionTestUtils.invokeMethod(service, "startDiagnosisRun", "same-session", "run-aiops-a", request);
ReflectionTestUtils.invokeMethod(service, "startDiagnosisRun", "same-session", "run-aiops-b", request);
verify(diagnosisRunRepository).save(argThat(run ->
"same-session".equals(run.getSessionId())
&& "run-aiops-a".equals(run.getRunId())
&& "AI_OPS".equals(run.getAgentFlow())
&& "RUNNING".equals(run.getStatus())));
verify(diagnosisRunRepository).save(argThat(run ->
"same-session".equals(run.getSessionId())
&& "run-aiops-b".equals(run.getRunId())
&& "AI_OPS".equals(run.getAgentFlow())
&& "RUNNING".equals(run.getStatus())));
}
@Test
void buildSupervisorAgentSetsPlannerAsMainAgent() {
ChatModel chatModel = mock(ChatModel.class);
ReactAgent planner = mock(ReactAgent.class);
ReactAgent executor = mock(ReactAgent.class);
SupervisorAgent supervisor = service.buildSupervisorAgent(chatModel, planner, executor);
assertSame(planner, supervisor.getMainAgent());
}
}