Compare commits
11
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
79feed3314 | ||
|
|
98155ae1d8 | ||
|
|
2609c5a5ab | ||
|
|
bf5286c8f4 | ||
|
|
26e12a8d6b | ||
|
|
cbef3ddd3c | ||
|
|
69deb15330 | ||
|
|
4c7c53b024 | ||
|
|
ca5c61fabf | ||
|
|
23ee05c7c3 | ||
|
|
dc6cd32a67 |
@@ -60,3 +60,7 @@ uploads/
|
||||
### Windows / Runtime Artifacts
|
||||
*.stackdump
|
||||
NUL
|
||||
|
||||
### MVP Demo Generated Outputs
|
||||
mvp/demo/output/*.json
|
||||
!mvp/demo/output/README.md
|
||||
|
||||
@@ -4,7 +4,14 @@
|
||||
|
||||
| 日期 | slug | 领域 | 关键词 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
|
||||
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
|
||||
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
|
||||
| 2026-07-04 | diagnosis-eval-harness | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
|
||||
| 2026-07-04 | evidence-trace-hardening | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
|
||||
| 2026-07-03 | mvp-demo-trace-acceptance | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
|
||||
| 2026-07-04 | aiops-traceable-diagnosis-entry | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
|
||||
| 2026-07-04 | aiops-alert-scope-control | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
|
||||
| 2026-05-29 | chatmodel-abstraction | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | archived |
|
||||
| 2026-06-23 | phase1-infrastructure | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | archived |
|
||||
| 2026-06-24 | lookup-knowledge-integration | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | archived |
|
||||
|
||||
@@ -0,0 +1,14 @@
|
||||
# Acceptance: aiops-alert-scope-control
|
||||
|
||||
## Verification
|
||||
|
||||
- [x] Payload-mode prompt focuses the final report on the supplied alert.
|
||||
- [x] No-payload prompt requires active-alert discovery first.
|
||||
- [x] Targeted tests pass.
|
||||
- [x] Compile passes.
|
||||
- [x] OpenSpec validates.
|
||||
|
||||
## Known Limits
|
||||
|
||||
- Prompt-only scope control may still require runtime observation.
|
||||
- AIOps Verifier remains deferred.
|
||||
@@ -0,0 +1,28 @@
|
||||
# Brief: aiops-alert-scope-control
|
||||
|
||||
## Background
|
||||
|
||||
After `aiops-traceable-diagnosis-entry`, AIOps can be triggered by payload and replayed through trace. Runtime verification showed one semantic gap: payload mode still produced a broad report over all active mock alerts.
|
||||
|
||||
## Goal
|
||||
|
||||
Make AIOps scope explicit:
|
||||
|
||||
- Payload present -> targeted diagnosis for the supplied alert.
|
||||
- Payload absent -> automatic active-alert discovery and diagnosis.
|
||||
|
||||
## Scope
|
||||
|
||||
- In scope:
|
||||
- `AiOpsService.buildTaskPrompt(...)` scope rules.
|
||||
- Focused tests.
|
||||
- Demo acceptance wording.
|
||||
- Out of scope:
|
||||
- Verifier integration.
|
||||
- Java-side filtering of tool results.
|
||||
- API shape changes.
|
||||
- Database changes.
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/aiops-alert-scope-control/`
|
||||
@@ -0,0 +1,42 @@
|
||||
# Decisions: aiops-alert-scope-control
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: tighten AIOps report scope after runtime verification showed payload mode still analyzes all active alerts.
|
||||
- Slug: `aiops-alert-scope-control`
|
||||
- Scale: standard-light.
|
||||
|
||||
## Context
|
||||
|
||||
- AIOps traceability is implemented and verified.
|
||||
- Mock Prometheus returns multiple active alerts.
|
||||
- Payload demo supplies `HighCPUUsage/payment-service`, but previous report expanded to `HighMemoryUsage` and `SlowResponse`.
|
||||
|
||||
## Grill Question Pool
|
||||
|
||||
| # | Dimension | Question | Mode | Status |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | Product Boundary | What makes `/api/ai_ops` different from `/api/chat` when payload exists? | evidence-driven | Payload is alert-event driven and should be scoped to that event. |
|
||||
| Q2 | Scope | Should payload mode ignore all other active alerts? | user-interview | No; mention only as related risk/context. |
|
||||
| Q3 | Compatibility | Should no-payload mode keep old "query active alerts" behavior? | evidence-driven | Yes. |
|
||||
| Q4 | Enforcement | Should Java filter unrelated tool results now? | evidence-driven | No; prompt-only is sufficient for this small change. |
|
||||
| Q5 | Verifier | Should this change add AIOps Verifier? | user-interview | No; keep deferred. |
|
||||
|
||||
## Evidence-Driven Conclusions
|
||||
|
||||
| Conclusion | Evidence Source | Result |
|
||||
|---|---|---|
|
||||
| Scope issue is prompt-level. | `/api_ ai_ops` trace showed all mock alerts analyzed despite payload. | Update task prompt. |
|
||||
| No API or persistence changes are needed. | `AIOpsRequest` already carries payload and trace works. | Keep endpoint unchanged. |
|
||||
| Blast radius is low. | `buildTaskPrompt(...)` is internal to `AiOpsService`. | Add tests for prompt content. |
|
||||
|
||||
## GitNexus
|
||||
|
||||
GitNexus remains skipped by prior user decision and because tools are not exposed in this session. Local impact analysis is recorded instead.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Payload mode is detected when any alert field is present.
|
||||
- Payload mode final report must focus on the supplied alert.
|
||||
- No-payload mode must first call `queryPrometheusAlerts`.
|
||||
- Other active alerts in payload mode can appear only as related risk, not as separate root-cause sections.
|
||||
@@ -0,0 +1,44 @@
|
||||
# Evidence: aiops-alert-scope-control
|
||||
|
||||
## Local Impact Analysis
|
||||
|
||||
- `AiOpsService.buildTaskPrompt(...)` is used by `executeAiOpsAnalysis(...)`.
|
||||
- No controller, DTO, repository, or database changes are required.
|
||||
- Existing `AiOpsServiceTest` already exercises request summary helpers and can be extended for scope prompt rules.
|
||||
|
||||
## Verification Results
|
||||
|
||||
- `mvn -q "-Dtest=AiOpsServiceTest" test` passed.
|
||||
- `mvn -q -DskipTests compile` passed.
|
||||
- `openspec.cmd validate aiops-alert-scope-control --strict` passed.
|
||||
|
||||
## Runtime Verification
|
||||
|
||||
- Runtime session: `mvp-demo-aiops-payment-cpu-codex-scope-003`.
|
||||
- `/api/ai_ops` SSE emitted the requested `session` message and finished with `done`.
|
||||
- `diagnosis_session` persisted:
|
||||
- `agent_flow = AI_OPS`
|
||||
- `status = SUCCESS`
|
||||
- `total_duration_ms = 69875`
|
||||
- `step_count = 5`
|
||||
- `tool_call_count = 8`
|
||||
- Tool invocation counts:
|
||||
- `query_metrics = 1`
|
||||
- `lookup_knowledge = 1`
|
||||
- `query_logs = 6`
|
||||
- Report scope check:
|
||||
- `告警根因分析 - HighCPUUsage` exists.
|
||||
- `告警根因分析 - HighMemoryUsage` does not exist.
|
||||
- `告警根因分析 - SlowResponse` does not exist.
|
||||
- `相关风险告警` exists.
|
||||
|
||||
## Runtime Fix
|
||||
|
||||
- Added Hikari settings in `src/main/resources/application.yml` after the first runtime attempt failed on stale MySQL pool connections:
|
||||
- `maximum-pool-size: 5`
|
||||
- `minimum-idle: 1`
|
||||
- `connection-timeout: 10000`
|
||||
- `validation-timeout: 5000`
|
||||
- `idle-timeout: 60000`
|
||||
- `max-lifetime: 120000`
|
||||
- `keepalive-time: 30000`
|
||||
@@ -0,0 +1,18 @@
|
||||
# Acceptance: aiops-traceable-diagnosis-entry
|
||||
|
||||
## Verification
|
||||
|
||||
- [x] OpenSpec validates for `aiops-traceable-diagnosis-entry`.
|
||||
- [x] Targeted AIOps service tests pass.
|
||||
- [x] Compile verification passes.
|
||||
- [x] Demo docs describe AIOps request -> session id -> trace query.
|
||||
|
||||
## Result
|
||||
|
||||
Accepted for implementation scope.
|
||||
|
||||
## Known Limits
|
||||
|
||||
- AIOps Verifier integration is deferred.
|
||||
- Runtime still depends on configured model and infrastructure.
|
||||
- Full browser/SSE runtime verification is not guaranteed in this coding pass.
|
||||
@@ -0,0 +1,28 @@
|
||||
# Brief: aiops-traceable-diagnosis-entry
|
||||
|
||||
## Background
|
||||
|
||||
The MVP chat diagnosis path is now traceable through `diagnosis_session`, `agent_step`, `tool_invocation`, and `GET /api/diagnosis/{sessionId}/trace`. The older `/api/ai_ops` endpoint still acts like a standalone SSE demo: it accepts no alert payload, generates an internal session id, and does not make trace replay obvious to callers.
|
||||
|
||||
## Goal
|
||||
|
||||
Turn AIOps into an alert-triggered diagnosis entry point that shares the same evidence and trace story as the main MVP, without rewriting the whole AIOps flow.
|
||||
|
||||
## Scope
|
||||
|
||||
- In scope:
|
||||
- Optional AIOps alert request body.
|
||||
- Stable request/session id propagation.
|
||||
- Persisted AIOps query summary and final answer.
|
||||
- SSE session id event.
|
||||
- Demo documentation and focused tests.
|
||||
- Out of scope:
|
||||
- Full AIOps and ChatService unification.
|
||||
- AIOps Verifier integration.
|
||||
- Database schema changes.
|
||||
- Sensitive configuration cleanup.
|
||||
- Fully offline runtime.
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/aiops-traceable-diagnosis-entry/`
|
||||
@@ -0,0 +1,68 @@
|
||||
# Decisions: aiops-traceable-diagnosis-entry
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: make the legacy AIOps SSE endpoint a traceable alert diagnosis entry for the Agent Engineer interview MVP.
|
||||
- Slug: `aiops-traceable-diagnosis-entry`
|
||||
- Scale: standard-light, because this extends one public endpoint and reuses existing persistence/trace infrastructure.
|
||||
|
||||
## Context
|
||||
|
||||
- `mvp-demo-trace-acceptance` already added `GET /api/diagnosis/{sessionId}/trace`.
|
||||
- `chat-verifier-agent` made the chat path stronger than the older AIOps path.
|
||||
- Current AIOps value is as a second entry point: system alert -> automated diagnosis -> evidence trace.
|
||||
|
||||
## Grill Question Pool
|
||||
|
||||
| # | Dimension | Question | Mode | Status |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | Positioning | Is AIOps an independent product path or an alert-triggered sibling of Chat Diagnosis? | user-interview | Resolved: sibling entry, unified trace story |
|
||||
| Q2 | API | Should we keep `/api/ai_ops` or add a new endpoint? | evidence-driven | Resolved: keep existing endpoint and extend optional body |
|
||||
| Q3 | Input | What is the minimum alert payload? | user-interview | Resolved: `sessionId`, `alertName`, `service`, `severity`, `description`, `timeRange`, plus `userRequest` fallback |
|
||||
| Q4 | Output | How does the caller learn the trace session id? | evidence-driven | Resolved: first SSE event uses type `session` |
|
||||
| Q5 | Trace | Must AIOps be replayable with existing trace API? | evidence-driven | Resolved: yes, this is the main acceptance criterion |
|
||||
| Q6 | Verifier | Must this slice add AIOps Verifier? | user-interview | Resolved: no, defer as follow-up |
|
||||
| Q7 | Compatibility | Should no-body calls still work? | evidence-driven | Resolved: yes, preserve old demo behavior |
|
||||
| Q8 | GitNexus | Should unavailable GitNexus block implementation? | user-interview | Resolved: skip GitNexus by user decision |
|
||||
|
||||
## Evidence-Driven Conclusions
|
||||
|
||||
| Conclusion | Evidence Source | Result |
|
||||
|---|---|---|
|
||||
| AIOps is currently isolated from request-driven trace replay. | `ChatController.aiOps()` has no request body; `AiOpsService` creates its own random session id. | Extend endpoint and service. |
|
||||
| No schema change is needed. | `DiagnosisSession` already has `query`, `agentFlow`, `answer`, counts, and status. | Reuse existing table. |
|
||||
| Trace API can already replay AIOps if session id and answer are persisted. | `DiagnosisTraceService` loads by session id and is flow-agnostic. | Keep trace API unchanged. |
|
||||
| Blast radius is moderate and local. | `rg` shows only `ChatController` calls `executeAiOpsAnalysis` and `extractFinalReport`. | Change service/controller carefully and add tests. |
|
||||
|
||||
## User-Interview Confirmations
|
||||
|
||||
| Topic | User Words | Decision |
|
||||
|---|---|---|
|
||||
| Use sm-flow | "可以,改造一下AIOps 接口,用sm-flow流程看看" | Use OpenSpec + devflow. |
|
||||
| GitNexus | "跳过gitnexus把" | Record skip and use local impact analysis. |
|
||||
| Proceed after Grill | "可以" | Continue with lightweight Grill conclusions. |
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Keep `/api/ai_ops` and make its body optional.
|
||||
- Emit `SseMessage.type=session` before long-running analysis starts.
|
||||
- Store AIOps request summary in `diagnosis_session.query`.
|
||||
- Store final report in `diagnosis_session.answer`.
|
||||
- Defer AIOps Verifier to a later change so this slice stays focused.
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
```text
|
||||
POST /api/ai_ops
|
||||
-> optional AIOpsRequest
|
||||
-> resolve sessionId
|
||||
-> create diagnosis_session(agentFlow=AI_OPS)
|
||||
-> run ai_ops_supervisor(planner, executor)
|
||||
-> AgentLoggingHook persists steps
|
||||
-> tools persist invocations under SessionContextHolder
|
||||
-> extract final report
|
||||
-> persist answer
|
||||
-> GET /api/diagnosis/{sessionId}/trace replays the run
|
||||
```
|
||||
|
||||
Risk level: medium. The endpoint is public and SSE-based, but the change is additive and does not change the chat diagnosis path or database schema.
|
||||
@@ -0,0 +1,37 @@
|
||||
# Evidence: aiops-traceable-diagnosis-entry
|
||||
|
||||
## Local Impact Analysis
|
||||
|
||||
- `ChatController.aiOps()` is the only caller of `AiOpsService.executeAiOpsAnalysis(...)`.
|
||||
- `ChatController.aiOps()` is the only caller of `AiOpsService.extractFinalReport(...)`.
|
||||
- `AIOpsRequest` exists but only has `userRequest`; no current controller consumes it.
|
||||
- `DiagnosisTraceService` is flow-agnostic and reads persisted session/step/tool records by `sessionId`.
|
||||
|
||||
## GitNexus
|
||||
|
||||
GitNexus MCP tools were not exposed in this session. The user explicitly approved skipping GitNexus for this change. Local impact analysis and targeted tests are used instead.
|
||||
|
||||
## Expected Verification
|
||||
|
||||
- Focused unit tests for AIOps request/session/report helper behavior.
|
||||
- Compile verification.
|
||||
- OpenSpec validation if CLI is available.
|
||||
|
||||
## Verification Results
|
||||
|
||||
- `openspec.cmd validate aiops-traceable-diagnosis-entry --strict`: passed.
|
||||
- `mvn -q "-Dtest=AiOpsServiceTest,DiagnosisTraceServiceTest" test`: passed after rerun with approved Maven access.
|
||||
- `mvn -q -DskipTests compile`: passed.
|
||||
|
||||
## Demo Alignment
|
||||
|
||||
- Added `knowledge_base/troubleshooting/aiops-alert-runbook.md` so mock AIOps alerts have matching knowledge-base guidance.
|
||||
- Aligned the documented AIOps demo with mock data: `HighCPUUsage` on `payment-service`, using `system-metrics` evidence.
|
||||
|
||||
## Metric Alignment Follow-up
|
||||
|
||||
- Runtime verification showed `diagnosis_session.tool_call_count` counted agent steps with tool calls, while trace returned actual `tool_invocation` records.
|
||||
- Updated `ChatService` and `AiOpsService` metric backfill to use `ToolInvocationRepository.countBySessionId(sessionId)`.
|
||||
- Targeted verification:
|
||||
- `mvn -q "-Dtest=AiOpsServiceTest,ChatServiceSequentialAgentTest,DiagnosisTraceServiceTest" test`: passed.
|
||||
- `mvn -q -DskipTests compile`: passed.
|
||||
@@ -0,0 +1,36 @@
|
||||
# Acceptance: diagnosis-eval-harness
|
||||
|
||||
## Classification
|
||||
|
||||
standard-light
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Issue and OpenSpec setup | Done | `ISS-006` and initial OpenSpec artifacts were created. |
|
||||
| Implementation | Done | Added fixed cases, fixture-mode trace evaluation, aggregate metrics, and JSON / Markdown report writer. |
|
||||
| Verification | Done | Targeted evaluator tests, compile verification, and OpenSpec validation passed. |
|
||||
|
||||
## Current State
|
||||
|
||||
- First implementation uses fixture-mode evaluation.
|
||||
- Live trace API polling remains a follow-up option.
|
||||
|
||||
## Verification
|
||||
|
||||
### Script Verification
|
||||
|
||||
- Command: `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
|
||||
- Result: passed
|
||||
- Notes: Covers fixed case loading, fixture evaluation, missing fixture reporting, reject degraded-output validation, and report writing.
|
||||
|
||||
### Static Verification
|
||||
|
||||
- Command: `mvn -q -DskipTests compile`
|
||||
- Result: passed
|
||||
|
||||
### OpenSpec Verification
|
||||
|
||||
- Command: `openspec validate diagnosis-eval-harness --strict`
|
||||
- Result: passed
|
||||
@@ -0,0 +1,31 @@
|
||||
# Brief: diagnosis-eval-harness
|
||||
|
||||
## Background
|
||||
|
||||
The MVP has a runnable demo and hardened evidence trace semantics, but it still lacks a fixed regression baseline for Agent diagnosis quality. P1-B creates a small evaluation harness that can validate diagnosis traces against fixed cases and produce repeatable reports.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Define fixed diagnosis cases for the MVP demo domain.
|
||||
2. Validate trace evidence, verifier verdicts, answer keywords, and degraded-output behavior.
|
||||
3. Produce JSON and Markdown reports for interview and regression use.
|
||||
4. Keep the first version offline by supporting trace fixtures.
|
||||
|
||||
## Scope
|
||||
|
||||
- Evaluation case definitions
|
||||
- Trace fixture shape
|
||||
- Rule-based evaluator
|
||||
- JSON / Markdown report output
|
||||
- Focused offline tests and docs
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No LLM-as-judge
|
||||
- No live end-to-end runtime requirement
|
||||
- No production API
|
||||
- No chat or verifier runtime change
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/diagnosis-eval-harness/`
|
||||
@@ -0,0 +1,28 @@
|
||||
# Diagnosis Eval Harness Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: build P1-B fixed case evaluation after evidence trace hardening.
|
||||
- Slug: `diagnosis-eval-harness`
|
||||
- Devflow scale: standard-light
|
||||
|
||||
## Context
|
||||
|
||||
- P1-A `evidence-trace-hardening` created stable evidence semantics for supported, no-evidence, deduped, and failed tool calls.
|
||||
- The MVP demo trace API already provides an aggregate trace shape suitable for evaluation.
|
||||
- The first evaluator should avoid depending on external infrastructure so it can run in regular development.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Start with rule-based trace validation instead of LLM-as-judge.
|
||||
- Reason: The first regression signal should be deterministic and tied to trace contracts.
|
||||
|
||||
- Decision: Support offline fixture traces first.
|
||||
- Reason: This makes the harness usable without MySQL, Redis, Milvus, or a real LLM.
|
||||
|
||||
- Decision: Output both JSON and Markdown.
|
||||
- Reason: JSON supports automation; Markdown is easier to discuss in interviews.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Whether live trace API polling belongs in this change or a follow-up after fixture mode lands.
|
||||
@@ -0,0 +1,10 @@
|
||||
# Diagnosis Eval Harness Evidence
|
||||
|
||||
## Evidence
|
||||
|
||||
| Source | Evidence | Conclusion | Reported |
|
||||
|---|---|---|---|
|
||||
| `openspec/specs/evidence-trace-hardening/spec.md` | Defines stable evidence states and summary behavior | Evaluation can rely on trace semantics rather than ad hoc log parsing | Yes |
|
||||
| `mvp/demo/README.md` | Documents an end-to-end demo flow with chat, trace, and feedback | Existing demo flow provides the runtime story, but not a reusable evaluation baseline | Yes |
|
||||
| `DiagnosisTraceService` | Aggregates session, steps, tools, and self-evaluation | Trace response shape can be reused as evaluation input | Yes |
|
||||
| `ToolTraceSummaryService` | Builds verifier-facing evidence summaries from persisted tool rows | Evaluator can check evidence coverage through persisted trace artifacts | Yes |
|
||||
@@ -0,0 +1,32 @@
|
||||
# Acceptance: evidence-trace-hardening
|
||||
|
||||
## Classification
|
||||
|
||||
standard-light
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Issue and OpenSpec setup | Done | `ISS-005` and the initial OpenSpec artifacts were created. |
|
||||
| Implementation | Done | Recorder contract, lookup persistence path, evidence summary semantics, and degraded-path tests were implemented. |
|
||||
| Verification | Done | Targeted offline tests and compile verification passed. |
|
||||
|
||||
## Verification
|
||||
|
||||
### Script Verification
|
||||
|
||||
- Command: `mvn -q "-Dtest=ToolInvocationRecorderTest,ToolTraceSummaryServiceTest,ChatServiceSequentialAgentTest,LookupKnowledgeToolTest" test`
|
||||
- Result: passed
|
||||
- Notes: Covers recorder contract, summary semantics for success/failure/no-evidence, and `ChatService` fallback / degraded paths.
|
||||
|
||||
### Static Verification
|
||||
|
||||
- Command: `mvn -q -DskipTests compile`
|
||||
- Result: passed
|
||||
|
||||
## Open Questions
|
||||
|
||||
| Question | Current position |
|
||||
| --- | --- |
|
||||
| Should deduped retrievals be counted separately from generic no-hit events in future evaluation metrics? | Deferred to P1-B; this change preserves enough structure to decide later. |
|
||||
@@ -0,0 +1,32 @@
|
||||
# Brief: evidence-trace-hardening
|
||||
|
||||
## Background
|
||||
|
||||
The MVP already has persisted tool traces and a verifier, but the evidence contract is still only partially standardized. For interview-focused hardening, the project now needs a tighter contract for evidence persistence, no-evidence / failure semantics, and degraded-output behavior.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Standardize the persisted evidence-tool contract across `lookup_knowledge`, `query_logs`, and `query_metrics`.
|
||||
2. Make verifier-facing summaries distinguish failed calls, no-hit calls, deduped retrievals, and actual supporting evidence.
|
||||
3. Add offline tests for verifier fallback and degraded-output paths.
|
||||
|
||||
## Scope
|
||||
|
||||
- `ToolInvocationRecorder`
|
||||
- `LookupKnowledgeTool`
|
||||
- `QueryLogsTools`
|
||||
- `QueryMetricsTools`
|
||||
- `ToolTraceSummaryService`
|
||||
- `ChatService`
|
||||
- Focused offline tests
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No new API or schema
|
||||
- No evaluation harness yet
|
||||
- No trace UI
|
||||
- No security/config cleanup
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/evidence-trace-hardening/`
|
||||
@@ -0,0 +1,24 @@
|
||||
# Evidence Trace Hardening Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: harden the MVP evidence contract before building the P1-B evaluation harness.
|
||||
- Slug: `evidence-trace-hardening`
|
||||
- Devflow scale: standard-light
|
||||
|
||||
## Context
|
||||
|
||||
- `ISS-003` raised verifier traceability and failure-path concerns.
|
||||
- Current code inspection shows `QueryLogsTools` and `QueryMetricsTools` already use `ToolInvocationRecorder`, while `LookupKnowledgeTool` still persists rows through a local helper.
|
||||
- `ChatService` already contains fallback behavior for missing/invalid `verifier_output`, but coverage is narrow.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Treat this as a contract-hardening change, not a new feature change.
|
||||
- Reason: The project already has the necessary runtime pieces; the gap is semantic consistency and testability.
|
||||
|
||||
- Decision: Keep the scope before P1-B.
|
||||
- Reason: The evaluation harness will rely on stable evidence semantics, so this contract slice should land first.
|
||||
|
||||
- Decision: Preserve schema and API stability.
|
||||
- Reason: The interview value here is engineering rigor, not more surface area.
|
||||
@@ -0,0 +1,11 @@
|
||||
# Evidence Trace Hardening Evidence
|
||||
|
||||
## Evidence
|
||||
|
||||
| Source | Evidence | Conclusion | Reported |
|
||||
|---|---|---|---|
|
||||
| `ToolInvocationRecorder` | Provides a common persistence seam for evidence tools | Contract hardening should build on the existing recorder instead of introducing a new store path | Yes |
|
||||
| `LookupKnowledgeTool` | Still constructs `ToolInvocation` rows through a local helper | Retrieval-aware evidence persistence is not yet unified with the recorder contract | Yes |
|
||||
| `QueryLogsTools` / `QueryMetricsTools` | Already record evidence invocations through `recordEvidenceTool(...)` | Current gap is semantic alignment, not missing persistence | Yes |
|
||||
| `ToolTraceSummaryService` | Merges rows by tool and topic domain and infers evidence level heuristically | Summary rules need explicit handling for failure, no-hit, and dedup cases | Yes |
|
||||
| `ChatService` | Falls back to `LOW_CONFID` when verifier output is missing or invalid | These degraded paths exist and should now be covered by focused offline tests | Yes |
|
||||
@@ -0,0 +1,37 @@
|
||||
# Acceptance: expand-diagnosis-eval-fixtures
|
||||
|
||||
## Classification
|
||||
|
||||
standard-light
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Issue and OpenSpec setup | Done | Created slug-based issue and initial OpenSpec artifacts. |
|
||||
| Implementation | Done | Added remaining fixtures, full baseline reports, and documentation updates. |
|
||||
| Verification | Done | Evaluator tests, compile verification, and OpenSpec validation passed. |
|
||||
|
||||
## Current State
|
||||
|
||||
- Fixture coverage is complete for the five fixed diagnosis cases.
|
||||
- Baseline reports are saved under `mvp/eval/reports`.
|
||||
- No production runtime behavior has been changed.
|
||||
|
||||
## Verification
|
||||
|
||||
### Script Verification
|
||||
|
||||
- Command: `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
|
||||
- Result: passed
|
||||
- Notes: Covers full fixture coverage, baseline report matching, reject degraded-output validation, and report writing.
|
||||
|
||||
### Static Verification
|
||||
|
||||
- Command: `mvn -q -DskipTests compile`
|
||||
- Result: passed
|
||||
|
||||
### OpenSpec Verification
|
||||
|
||||
- Command: `openspec validate expand-diagnosis-eval-fixtures --strict`
|
||||
- Result: passed
|
||||
@@ -0,0 +1,31 @@
|
||||
# Brief: expand-diagnosis-eval-fixtures
|
||||
|
||||
## Background
|
||||
|
||||
The diagnosis eval harness is implemented and archived, but the fixed baseline is incomplete because three of the five diagnosis cases still reference missing fixtures.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Add representative trace fixtures for all remaining fixed diagnosis cases.
|
||||
2. Save a reproducible baseline report in JSON and Markdown.
|
||||
3. Document how to regenerate and interpret the baseline.
|
||||
4. Keep evaluation offline and deterministic.
|
||||
|
||||
## Scope
|
||||
|
||||
- Redis timeout fixture
|
||||
- Slow response fixture
|
||||
- JVM memory risk fixture
|
||||
- Baseline reports under `mvp/eval/reports`
|
||||
- Focused tests for full fixture coverage and report generation
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No new diagnosis cases
|
||||
- No production Agent runtime changes
|
||||
- No LLM-as-judge
|
||||
- No live infrastructure requirement
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/expand-diagnosis-eval-fixtures/`
|
||||
@@ -0,0 +1,28 @@
|
||||
# Expand Diagnosis Eval Fixtures Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: complete the fixed diagnosis eval baseline after the harness is in place.
|
||||
- Slug: `expand-diagnosis-eval-fixtures`
|
||||
- Devflow scale: standard-light
|
||||
|
||||
## Context
|
||||
|
||||
- `diagnosis-eval-harness` created the evaluator, case file, fixture mode, and report writer.
|
||||
- The first baseline still has missing fixtures by design.
|
||||
- This follow-up turns that partial baseline into a full fixed-case baseline.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Keep this change data-focused.
|
||||
- Reason: the evaluator rules already landed; this change should not blur fixture expansion with harness behavior changes.
|
||||
|
||||
- Decision: Save baseline reports in the repository.
|
||||
- Reason: interview review and future diffs are easier when the expected baseline is visible.
|
||||
|
||||
- Decision: Use deterministic fixture traces instead of live trace generation.
|
||||
- Reason: this baseline should run without infrastructure or external model calls.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Whether a future change should add a CLI or Maven goal for report regeneration.
|
||||
@@ -0,0 +1,11 @@
|
||||
# Evidence: expand-diagnosis-eval-fixtures
|
||||
|
||||
## Evidence Log
|
||||
|
||||
- 2026-07-04: Created slug-based issue `expand-diagnosis-eval-fixtures.md`.
|
||||
- 2026-07-04: Created OpenSpec change `expand-diagnosis-eval-fixtures`.
|
||||
- 2026-07-04: Added Redis timeout, slow response, and JVM memory risk fixtures.
|
||||
- 2026-07-04: Added baseline JSON and Markdown reports under `mvp/eval/reports`.
|
||||
- 2026-07-04: Verification passed with `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`.
|
||||
- 2026-07-04: Verification passed with `mvn -q -DskipTests compile`.
|
||||
- 2026-07-04: Verification passed with `openspec validate expand-diagnosis-eval-fixtures --strict`.
|
||||
@@ -0,0 +1,37 @@
|
||||
# Acceptance: diagnosis-eval-baseline-diff
|
||||
|
||||
## Classification
|
||||
|
||||
standard-light
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Issue and OpenSpec setup | Done | Created slug-based issue and initial OpenSpec artifacts. |
|
||||
| Implementation | Done | Added diff model, comparator, writer, docs, sample outputs, and focused tests. |
|
||||
| Verification | Done | Diff/evaluator tests, compile verification, and OpenSpec validation passed. |
|
||||
|
||||
## Current State
|
||||
|
||||
- Baseline diff is implemented for aggregate metrics, verdict distribution, case-level state, keyword coverage, evidence coverage, missing cases, and new cases.
|
||||
- JSON and Markdown diff output are available.
|
||||
- No production runtime behavior has been changed.
|
||||
|
||||
## Verification
|
||||
|
||||
### Script Verification
|
||||
|
||||
- Command: `mvn -q "-Dtest=DiagnosisEvalBaselineDiffTest" test`
|
||||
- Result: passed
|
||||
- Notes: Also verified with `DiagnosisTraceEvaluatorTest`.
|
||||
|
||||
### Static Verification
|
||||
|
||||
- Command: `mvn -q -DskipTests compile`
|
||||
- Result: passed
|
||||
|
||||
### OpenSpec Verification
|
||||
|
||||
- Command: `openspec validate diagnosis-eval-baseline-diff --strict`
|
||||
- Result: passed
|
||||
@@ -0,0 +1,30 @@
|
||||
# Brief: diagnosis-eval-baseline-diff
|
||||
|
||||
## Background
|
||||
|
||||
The eval harness now has a complete saved baseline. This change adds the comparison layer that turns the baseline into an actionable regression signal.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Compare baseline and current `DiagnosisEvalReport` objects.
|
||||
2. Detect aggregate and per-case regressions.
|
||||
3. Output JSON and Markdown diff reports.
|
||||
4. Document how to read the diff in interview and engineering terms.
|
||||
|
||||
## Scope
|
||||
|
||||
- Diff data structures
|
||||
- Deterministic report comparison
|
||||
- JSON / Markdown diff output
|
||||
- Focused tests and eval docs
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No live Agent execution
|
||||
- No LLM-as-judge
|
||||
- No evaluator scoring rule changes
|
||||
- No production API changes
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/diagnosis-eval-baseline-diff/`
|
||||
@@ -0,0 +1,28 @@
|
||||
# Diagnosis Eval Baseline Diff Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: add report diffing on top of the completed diagnosis eval baseline.
|
||||
- Slug: `diagnosis-eval-baseline-diff`
|
||||
- Devflow scale: standard-light
|
||||
|
||||
## Context
|
||||
|
||||
- `diagnosis-eval-harness` created deterministic fixture evaluation.
|
||||
- `expand-diagnosis-eval-fixtures` created a complete saved baseline.
|
||||
- This change compares new reports against that baseline.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Diff report DTOs instead of raw traces.
|
||||
- Reason: the report is the stable contract for regression review.
|
||||
|
||||
- Decision: Use deterministic code rules instead of LLM-as-judge.
|
||||
- Reason: baseline regression checks should be repeatable and explainable.
|
||||
|
||||
- Decision: Output both JSON and Markdown.
|
||||
- Reason: JSON supports automation; Markdown is useful in reviews and interviews.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Whether a future change should expose this through a CLI or Maven goal.
|
||||
@@ -0,0 +1,11 @@
|
||||
# Evidence: diagnosis-eval-baseline-diff
|
||||
|
||||
## Evidence Log
|
||||
|
||||
- 2026-07-05: Created slug-based issue `diagnosis-eval-baseline-diff.md`.
|
||||
- 2026-07-05: Created OpenSpec change `diagnosis-eval-baseline-diff`.
|
||||
- 2026-07-05: Added baseline diff DTOs, deterministic comparer, and JSON / Markdown writer.
|
||||
- 2026-07-05: Added sample baseline diff JSON and Markdown reports.
|
||||
- 2026-07-05: Verification passed with `mvn -q "-Dtest=DiagnosisEvalBaselineDiffTest,DiagnosisTraceEvaluatorTest" test`.
|
||||
- 2026-07-05: Verification passed with `mvn -q -DskipTests compile`.
|
||||
- 2026-07-05: Verification passed with `openspec validate diagnosis-eval-baseline-diff --strict`.
|
||||
@@ -0,0 +1,25 @@
|
||||
# Acceptance: mvp-demo-interview-runbook
|
||||
|
||||
## Classification
|
||||
|
||||
standard-light
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Issue and OpenSpec setup | Done | Created slug-based issue and OpenSpec artifacts. |
|
||||
| Implementation | Done | Added request payload, runnable script, output directory docs, interview walkthrough, and trace checklist. |
|
||||
| Verification | Done | OpenSpec validation passed. |
|
||||
|
||||
## Current State
|
||||
|
||||
- No backend runtime behavior has been changed.
|
||||
- Demo is packaged under `mvp/demo` for interview use.
|
||||
|
||||
## Verification
|
||||
|
||||
### OpenSpec Verification
|
||||
|
||||
- Command: `openspec validate mvp-demo-interview-runbook --strict`
|
||||
- Result: passed
|
||||
@@ -0,0 +1,29 @@
|
||||
# Brief: mvp-demo-interview-runbook
|
||||
|
||||
## Background
|
||||
|
||||
Plan C is the interview-facing demo package. The project has the engineering pieces, but needs a single place to run and explain the MVP flow.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Provide a fixed payment-timeout request payload.
|
||||
2. Provide a PowerShell script that runs chat, trace, and feedback.
|
||||
3. Save demo responses under `mvp/demo/output`.
|
||||
4. Add interview walkthrough and trace checklist.
|
||||
|
||||
## Scope
|
||||
|
||||
- Demo docs and scripts only
|
||||
- Existing local APIs only
|
||||
- Existing `mvp-demo` profile only
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No backend code changes
|
||||
- No eval extension
|
||||
- No secret cleanup
|
||||
- No full offline runtime
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/mvp-demo-interview-runbook/`
|
||||
@@ -0,0 +1,27 @@
|
||||
# MVP Demo Interview Runbook Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: package existing MVP capabilities into a repeatable interview demo.
|
||||
- Slug: `mvp-demo-interview-runbook`
|
||||
- Devflow scale: standard-light
|
||||
|
||||
## Context
|
||||
|
||||
- Evidence trace and eval baseline work are already done.
|
||||
- The next useful step is not more eval tooling, but a runnable demo path.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Keep this change documentation/script-only.
|
||||
- Reason: Plan C is about demo packaging, not new runtime capability.
|
||||
|
||||
- Decision: Use a stable session id.
|
||||
- Reason: it makes trace lookup and saved output predictable.
|
||||
|
||||
- Decision: Save outputs to `mvp/demo/output`.
|
||||
- Reason: generated artifacts should be easy to review without mixing into source fixtures.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Whether a later change should add a truly offline stubbed demo mode.
|
||||
@@ -0,0 +1,9 @@
|
||||
# Evidence: mvp-demo-interview-runbook
|
||||
|
||||
## Evidence Log
|
||||
|
||||
- 2026-07-05: Created Plan C demo packaging issue and OpenSpec change.
|
||||
- 2026-07-05: Added fixed payment-timeout request payload.
|
||||
- 2026-07-05: Added PowerShell demo script for chat, trace, and feedback.
|
||||
- 2026-07-05: Added interview walkthrough and trace inspection checklist.
|
||||
- 2026-07-05: Verification passed with `openspec validate mvp-demo-interview-runbook --strict`.
|
||||
@@ -0,0 +1,59 @@
|
||||
# SuperBizAgent Interview Guide
|
||||
|
||||
## 一句话定位
|
||||
|
||||
SuperBizAgent 是一个面向企业故障诊断场景的 Agent Engineering 项目:它把用户问题或告警事件转成可追踪的多 Agent 执行链路,并把工具证据、模型步骤、最终答案和反馈统一落到诊断 trace 中。
|
||||
|
||||
## 面试重点
|
||||
|
||||
- **多 Agent 编排**:普通 Chat 的复杂问题走 `Planner -> Executor -> Verifier`;AIOps 告警入口走 Supervisor 调度 Planner/Executor。
|
||||
- **工具证据链**:知识库、日志、指标和 Prometheus 告警都通过工具调用进入链路,并记录到 `tool_invocation`。
|
||||
- **可追踪诊断**:一次会话对应一个 `sessionId`,最终可以通过 `GET /api/diagnosis/{sessionId}/trace` 回放。
|
||||
- **质量门**:Chat 链路包含 Verifier,把 groundedness、facts checked 和 evidence refs 写回 `diagnosis_session.self_evaluation`。
|
||||
- **AIOps 产品边界**:有告警 payload 时聚焦该告警;没有 payload 时先自动发现 active alerts。
|
||||
- **可复现 Demo**:`mvp-demo` profile 使用 mock Prometheus 和 mock CLS,让面试演示不依赖真实线上故障。
|
||||
|
||||
## 推荐阅读顺序
|
||||
|
||||
1. `interview/demo-script.md`:面试现场怎么讲、怎么演示。
|
||||
2. `interview/architecture.md`:系统架构和两条主链路。
|
||||
3. `interview/design-tradeoffs.md`:关键设计取舍和可被追问的问题。
|
||||
4. `interview/acceptance-checklist.md`:面试前验证清单。
|
||||
5. `mvp/demo/README.md`:更细的 MVP 可执行 runbook。
|
||||
|
||||
## 核心 Demo
|
||||
|
||||
### Chat Diagnosis
|
||||
|
||||
```text
|
||||
POST /api/chat
|
||||
-> ChatService.executeChatWithStrategy(...)
|
||||
-> simple ReactAgent or Planner -> Executor -> Verifier
|
||||
-> lookup_knowledge / query_logs / query_metrics
|
||||
-> diagnosis_session + agent_step + tool_invocation
|
||||
-> GET /api/diagnosis/{sessionId}/trace
|
||||
```
|
||||
|
||||
### AIOps Alert Diagnosis
|
||||
|
||||
```text
|
||||
POST /api/ai_ops
|
||||
-> AiOpsService.executeAiOpsAnalysis(...)
|
||||
-> ai_ops_supervisor
|
||||
-> planner_agent / executor_agent
|
||||
-> queryPrometheusAlerts + logs + knowledge
|
||||
-> scoped alert report
|
||||
-> GET /api/diagnosis/{sessionId}/trace
|
||||
```
|
||||
|
||||
## 当前完成度
|
||||
|
||||
- Chat 诊断链路:可运行、可追踪、有 Verifier。
|
||||
- AIOps 告警链路:可运行、可追踪、支持 payload scope control。
|
||||
- Trace API:统一返回 session、agent steps、tool invocations 和 summary。
|
||||
- Demo 文档:`mvp/demo/README.md` 和 `mvp/demo/aiops-alert-acceptance.md`。
|
||||
- Devflow 沉淀:`devflow/index.md` 记录了 MVP、Verifier、AIOps trace 和 AIOps scope-control 的演进。
|
||||
|
||||
## 面试时的主叙事
|
||||
|
||||
这个项目不是简单调用大模型,而是在做一个可审计的 Agent 诊断系统。核心价值是:模型可以规划和推理,但每一步工具证据、最终结论和质量评估都能被 trace API 回放。面试时重点展示“从问题到证据到答案到验证”的完整闭环。
|
||||
@@ -0,0 +1,170 @@
|
||||
# Acceptance Checklist
|
||||
|
||||
## 面试前环境检查
|
||||
|
||||
- 当前分支包含最新 AIOps trace/scope 变更。
|
||||
- MySQL 可连接。
|
||||
- Redis 可连接。
|
||||
- Milvus/Zilliz 可连接。
|
||||
- 模型 API key 可用。
|
||||
- `mvp-demo` profile 开启 mock Prometheus 和 mock CLS。
|
||||
|
||||
启动:
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||
```
|
||||
|
||||
编译检查:
|
||||
|
||||
```powershell
|
||||
mvn -q -DskipTests compile
|
||||
```
|
||||
|
||||
目标测试:
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=AiOpsServiceTest,ChatServiceSequentialAgentTest,DiagnosisTraceServiceTest" test
|
||||
```
|
||||
|
||||
## Chat Demo 验收
|
||||
|
||||
请求:
|
||||
|
||||
```powershell
|
||||
$sessionId = "interview-chat-payment-timeout-001"
|
||||
$body = @{
|
||||
Id = $sessionId
|
||||
Question = "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/chat" `
|
||||
-ContentType "application/json" `
|
||||
-Body $body
|
||||
```
|
||||
|
||||
验收:
|
||||
|
||||
- 返回 `data.success = true`。
|
||||
- 返回 `data.sessionId = interview-chat-payment-timeout-001`。
|
||||
- `diagnosis_session.agent_flow = CHAT`。
|
||||
- trace API 返回 session、steps、toolInvocations。
|
||||
- 复杂问题下 trace 中能看到 verifier 相关数据。
|
||||
|
||||
SQL:
|
||||
|
||||
```powershell
|
||||
python scripts/query_mysql.py "SELECT session_id, agent_flow, status, step_count, tool_call_count FROM diagnosis_session WHERE session_id='interview-chat-payment-timeout-001'"
|
||||
```
|
||||
|
||||
## AIOps Demo 验收
|
||||
|
||||
请求:
|
||||
|
||||
```powershell
|
||||
$aiopsSessionId = "interview-aiops-payment-cpu-001"
|
||||
$aiopsBody = @{
|
||||
sessionId = $aiopsSessionId
|
||||
alertName = "HighCPUUsage"
|
||||
service = "payment-service"
|
||||
severity = "P1"
|
||||
description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
|
||||
timeRange = "last_15m"
|
||||
userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-WebRequest `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/ai_ops" `
|
||||
-ContentType "application/json" `
|
||||
-Body $aiopsBody
|
||||
```
|
||||
|
||||
验收:
|
||||
|
||||
- SSE 首条包含 `type=session`。
|
||||
- SSE 最后包含 `type=done`。
|
||||
- `diagnosis_session.agent_flow = AI_OPS`。
|
||||
- `diagnosis_session.status = SUCCESS`。
|
||||
- `diagnosis_session.answer` 有最终报告。
|
||||
- trace API 返回 AIOps steps 和 tool invocations。
|
||||
- 报告主章节聚焦 `HighCPUUsage/payment-service`。
|
||||
- 无 `告警根因分析 - HighMemoryUsage` 独立章节。
|
||||
- 无 `告警根因分析 - SlowResponse` 独立章节。
|
||||
- 有“相关风险告警”或类似上下文说明。
|
||||
|
||||
SQL:
|
||||
|
||||
```powershell
|
||||
python scripts/query_mysql.py "SELECT session_id, agent_flow, status, total_duration_ms, step_count, tool_call_count FROM diagnosis_session WHERE session_id='interview-aiops-payment-cpu-001'"
|
||||
```
|
||||
|
||||
```powershell
|
||||
python scripts/query_mysql.py "SELECT tool_name, COUNT(*) AS cnt FROM tool_invocation WHERE session_id='interview-aiops-payment-cpu-001' GROUP BY tool_name ORDER BY tool_name"
|
||||
```
|
||||
|
||||
Scope 检查:
|
||||
|
||||
```powershell
|
||||
python scripts/query_mysql.py "SELECT (answer LIKE '%告警根因分析 - HighCPUUsage%') AS has_main_root_cause, (answer LIKE '%告警根因分析 - HighMemoryUsage%') AS has_memory_root_cause, (answer LIKE '%告警根因分析 - SlowResponse%') AS has_slow_root_cause, (answer LIKE '%相关风险告警%') AS has_related_risk FROM diagnosis_session WHERE session_id='interview-aiops-payment-cpu-001'"
|
||||
```
|
||||
|
||||
期望:
|
||||
|
||||
```text
|
||||
has_main_root_cause = 1
|
||||
has_memory_root_cause = 0
|
||||
has_slow_root_cause = 0
|
||||
has_related_risk = 1
|
||||
```
|
||||
|
||||
## Trace API 验收
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/interview-aiops-payment-cpu-001/trace"
|
||||
```
|
||||
|
||||
若 PowerShell 对长 JSON 或特殊字符不稳定,可以用:
|
||||
|
||||
```powershell
|
||||
curl.exe --silent --show-error --max-time 60 "http://localhost:9900/api/diagnosis/interview-aiops-payment-cpu-001/trace"
|
||||
```
|
||||
|
||||
## 常见问题
|
||||
|
||||
### MySQL stale connection
|
||||
|
||||
现象:
|
||||
|
||||
```text
|
||||
HikariPool - Connection is not available
|
||||
No operations allowed after connection closed
|
||||
```
|
||||
|
||||
当前已在 `application.yml` 配置:
|
||||
|
||||
- `maximum-pool-size: 5`
|
||||
- `minimum-idle: 1`
|
||||
- `connection-timeout: 10000`
|
||||
- `validation-timeout: 5000`
|
||||
- `idle-timeout: 60000`
|
||||
- `max-lifetime: 120000`
|
||||
- `keepalive-time: 30000`
|
||||
|
||||
处理:
|
||||
|
||||
- 重新编译或重启服务。
|
||||
- 确认日志中新的 HikariPool 启动成功。
|
||||
- 再跑 trace 或 AIOps 请求。
|
||||
|
||||
### SSE 客户端显示异常
|
||||
|
||||
PowerShell `Invoke-WebRequest` 有时对 SSE 或长 JSON 处理不稳定。可以改用 `curl.exe` 或直接查询 MySQL 和 trace API 验证结果。
|
||||
|
||||
### OpenSpec 全量校验失败
|
||||
|
||||
`openspec validate --all --strict` 可能因为历史未完成 change 失败。面试材料主要依赖已归档的 AIOps spec 和 MVP trace spec,可以单独验证相关 spec。
|
||||
@@ -0,0 +1,147 @@
|
||||
# Architecture
|
||||
|
||||
## 系统分层
|
||||
|
||||
```text
|
||||
API Layer
|
||||
-> ChatController / DiagnosisTraceController
|
||||
|
||||
Agent Orchestration
|
||||
-> ChatService / AiOpsService
|
||||
|
||||
Tools
|
||||
-> lookupKnowledgeTool / queryLogs / queryMetrics / queryPrometheusAlerts
|
||||
|
||||
Persistence
|
||||
-> diagnosis_session / agent_step / tool_invocation
|
||||
|
||||
Trace
|
||||
-> GET /api/diagnosis/{sessionId}/trace
|
||||
```
|
||||
|
||||
## Chat 链路
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
User[User Question] --> ChatAPI[POST /api/chat]
|
||||
ChatAPI --> Strategy[ChatService.executeChatWithStrategy]
|
||||
Strategy --> Complexity{QuestionComplexity}
|
||||
Complexity -->|simple| Single[ReactAgent]
|
||||
Complexity -->|complex| Planner[Planner Agent]
|
||||
Planner --> Executor[Executor Agent]
|
||||
Executor --> Tools[Evidence Tools]
|
||||
Tools --> Executor
|
||||
Executor --> Verifier[Verifier Agent]
|
||||
Verifier --> Answer[Final Answer]
|
||||
Answer --> Session[diagnosis_session]
|
||||
Planner --> Steps[agent_step]
|
||||
Executor --> Steps
|
||||
Verifier --> Steps
|
||||
Tools --> Invocations[tool_invocation]
|
||||
Session --> Trace[GET /api/diagnosis/{sessionId}/trace]
|
||||
Steps --> Trace
|
||||
Invocations --> Trace
|
||||
```
|
||||
|
||||
关键代码:
|
||||
|
||||
- `ChatController.chat(...)`
|
||||
- `ChatService.executeChatWithStrategy(...)`
|
||||
- `ChatService.executeChatComplex(...)`
|
||||
- `AgentLoggingHook`
|
||||
- `ToolInvocationRecorder`
|
||||
- `DiagnosisTraceService.getTrace(...)`
|
||||
|
||||
## AIOps 链路
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Alert[Alert Payload or Empty Request] --> AiOpsAPI[POST /api/ai_ops]
|
||||
AiOpsAPI --> SessionEvent[SSE session event]
|
||||
AiOpsAPI --> AiOpsService[AiOpsService.executeAiOpsAnalysis]
|
||||
AiOpsService --> PromptMode{Payload?}
|
||||
PromptMode -->|yes| Targeted[PAYLOAD_TARGETED]
|
||||
PromptMode -->|no| Discovery[AUTO_DISCOVERY]
|
||||
Targeted --> Supervisor[ai_ops_supervisor]
|
||||
Discovery --> Supervisor
|
||||
Supervisor --> Planner[planner_agent]
|
||||
Supervisor --> Executor[executor_agent]
|
||||
Planner --> Tools[Prometheus / Logs / Knowledge]
|
||||
Executor --> Tools
|
||||
Tools --> Report[Alert Report]
|
||||
Report --> Persist[diagnosis_session.answer]
|
||||
Planner --> Steps[agent_step]
|
||||
Executor --> Steps
|
||||
Tools --> Invocations[tool_invocation]
|
||||
Persist --> Trace[GET /api/diagnosis/{sessionId}/trace]
|
||||
Steps --> Trace
|
||||
Invocations --> Trace
|
||||
```
|
||||
|
||||
关键代码:
|
||||
|
||||
- `ChatController.aiOps(...)`
|
||||
- `AIOpsRequest`
|
||||
- `AiOpsService.resolveSessionId(...)`
|
||||
- `AiOpsService.buildTaskPrompt(...)`
|
||||
- `AiOpsService.hasAlertPayload(...)`
|
||||
- `AiOpsService.persistFinalReport(...)`
|
||||
|
||||
## Trace 数据模型
|
||||
|
||||
### `diagnosis_session`
|
||||
|
||||
记录一次诊断会话的主信息:
|
||||
|
||||
- `session_id`
|
||||
- `query`
|
||||
- `status`
|
||||
- `agent_flow`
|
||||
- `total_duration_ms`
|
||||
- `total_token_count`
|
||||
- `step_count`
|
||||
- `tool_call_count`
|
||||
- `answer`
|
||||
- `self_evaluation`
|
||||
- `feedback`
|
||||
|
||||
### `agent_step`
|
||||
|
||||
记录 Agent 模型调用过程:
|
||||
|
||||
- `session_id`
|
||||
- `step_index`
|
||||
- `agent_name`
|
||||
- `model_input`
|
||||
- `model_output`
|
||||
- `thought`
|
||||
- `has_tool_call`
|
||||
- `duration_ms`
|
||||
- `token_count`
|
||||
|
||||
### `tool_invocation`
|
||||
|
||||
记录真实工具调用:
|
||||
|
||||
- `session_id`
|
||||
- `tool_name`
|
||||
- `input_params`
|
||||
- `output_preview`
|
||||
- `output_length`
|
||||
- `retrieval_layer`
|
||||
- `relevance_level`
|
||||
- `duration_ms`
|
||||
- `success`
|
||||
- `error_message`
|
||||
|
||||
## 为什么 trace 是核心
|
||||
|
||||
Agent 系统的风险不只是“答案错”,还包括“答案看起来对但无法解释”。这个项目把执行链路拆成 session、step、tool 三层,让面试官可以看到:
|
||||
|
||||
- 模型为什么这么答
|
||||
- 调了哪些工具
|
||||
- 工具返回了什么证据
|
||||
- Verifier 如何判断答案可信度
|
||||
- 用户反馈如何回写到同一个 session
|
||||
|
||||
这就是项目区别于普通 Chatbot 的地方。
|
||||
@@ -0,0 +1,132 @@
|
||||
# Interview Demo Script
|
||||
|
||||
## 30 秒开场
|
||||
|
||||
这是一个 Agent Engineering 项目,场景是企业故障诊断。它支持两类入口:用户主动提问的 Chat 诊断,以及告警事件驱动的 AIOps 诊断。项目重点不是单次回答,而是把多 Agent 执行、工具证据、Verifier 评估、最终报告和反馈都沉淀成可回放的 trace。
|
||||
|
||||
## Demo 准备
|
||||
|
||||
启动服务:
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||
```
|
||||
|
||||
确认服务地址:
|
||||
|
||||
```text
|
||||
http://localhost:9900
|
||||
```
|
||||
|
||||
`mvp-demo` profile 下:
|
||||
|
||||
- Prometheus 告警使用 mock 数据。
|
||||
- CLS 日志使用 mock 数据。
|
||||
- MySQL、Redis、Milvus/Zilliz 和模型配置仍使用当前项目配置。
|
||||
|
||||
## Demo 1: Chat 诊断
|
||||
|
||||
目标:展示普通用户问题如何进入多 Agent 诊断、调用工具、经过 Verifier,并生成 trace。
|
||||
|
||||
请求:
|
||||
|
||||
```powershell
|
||||
$sessionId = "interview-chat-payment-timeout-001"
|
||||
$body = @{
|
||||
Id = $sessionId
|
||||
Question = "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/chat" `
|
||||
-ContentType "application/json" `
|
||||
-Body $body
|
||||
```
|
||||
|
||||
讲解点:
|
||||
|
||||
- `ChatController` 把请求交给 `ChatService.executeChatWithStrategy(...)`。
|
||||
- 简单问题走单 ReactAgent,复杂问题走 `Planner -> Executor -> Verifier`。
|
||||
- Executor 可以调用知识库、日志、指标等工具。
|
||||
- Verifier 会基于工具证据生成 groundedness 评估。
|
||||
- 最终会写入 `diagnosis_session`、`agent_step`、`tool_invocation`。
|
||||
|
||||
查询 trace:
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace"
|
||||
```
|
||||
|
||||
展示点:
|
||||
|
||||
- `data.session.agentFlow = CHAT`
|
||||
- `data.steps` 中能看到 planner/executor/verifier
|
||||
- `data.toolInvocations` 中能看到证据工具
|
||||
- `data.session.selfEvaluation` 中有 verifier 结果
|
||||
|
||||
## Demo 2: AIOps 告警诊断
|
||||
|
||||
目标:展示告警 payload 如何触发 AIOps 入口,并且报告只聚焦目标告警。
|
||||
|
||||
请求:
|
||||
|
||||
```powershell
|
||||
$aiopsSessionId = "interview-aiops-payment-cpu-001"
|
||||
$aiopsBody = @{
|
||||
sessionId = $aiopsSessionId
|
||||
alertName = "HighCPUUsage"
|
||||
service = "payment-service"
|
||||
severity = "P1"
|
||||
description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
|
||||
timeRange = "last_15m"
|
||||
userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-WebRequest `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/ai_ops" `
|
||||
-ContentType "application/json" `
|
||||
-Body $aiopsBody
|
||||
```
|
||||
|
||||
讲解点:
|
||||
|
||||
- `/api/ai_ops` 接受可选 `AIOpsRequest`。
|
||||
- 首条 SSE 消息会返回 `type=session`。
|
||||
- `AiOpsService` 根据 payload 判断模式:
|
||||
- `PAYLOAD_TARGETED`:聚焦传入告警。
|
||||
- `AUTO_DISCOVERY`:没有 payload 时先查 active alerts。
|
||||
- AIOps 暂时不加 Verifier,先保证告警入口、证据工具和 trace 可用。
|
||||
|
||||
查询 trace:
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace"
|
||||
```
|
||||
|
||||
展示点:
|
||||
|
||||
- `data.session.agentFlow = AI_OPS`
|
||||
- `data.session.answer` 有最终告警报告
|
||||
- `data.toolInvocations` 有 `query_metrics`、`query_logs`、`lookup_knowledge`
|
||||
- 报告有 `HighCPUUsage/payment-service` 的完整根因分析
|
||||
- 其他 active alerts 只作为相关风险出现,不展开成独立根因章节
|
||||
|
||||
## MySQL 验证
|
||||
|
||||
```powershell
|
||||
python scripts/query_mysql.py "SELECT session_id, agent_flow, status, step_count, tool_call_count FROM diagnosis_session ORDER BY id DESC LIMIT 5"
|
||||
```
|
||||
|
||||
```powershell
|
||||
python scripts/query_mysql.py "SELECT tool_name, COUNT(*) AS cnt FROM tool_invocation WHERE session_id='interview-aiops-payment-cpu-001' GROUP BY tool_name"
|
||||
```
|
||||
|
||||
## 收尾总结
|
||||
|
||||
这套 Demo 展示的是一个完整 Agent 系统,而不是一次模型问答:入口有明确场景边界,Agent 负责规划和执行,工具提供证据,Verifier 提供质量门,trace API 提供审计和复盘能力。AIOps 入口进一步证明它可以从用户问答扩展到事件驱动诊断。
|
||||
@@ -0,0 +1,99 @@
|
||||
# Design Tradeoffs
|
||||
|
||||
## 1. 为什么要做 trace,而不是只返回答案
|
||||
|
||||
普通 Chatbot 只关注最终回答,但故障诊断更需要可审计性。一次诊断至少要回答三件事:
|
||||
|
||||
- 结论是什么
|
||||
- 证据来自哪里
|
||||
- 哪些步骤由哪个 Agent 完成
|
||||
|
||||
因此项目把一次会话拆成:
|
||||
|
||||
- `diagnosis_session`:会话级摘要、最终答案、质量评估、反馈。
|
||||
- `agent_step`:Agent 模型输入输出、耗时、token 和工具调用标记。
|
||||
- `tool_invocation`:真实工具调用参数、输出预览、成功状态和检索元数据。
|
||||
|
||||
这个设计牺牲了一些实现复杂度,但换来了可回放、可调试、可演示。
|
||||
|
||||
## 2. 为什么 Chat 有 Verifier,AIOps 暂时没有
|
||||
|
||||
Chat 入口的问题更开放,用户可能要求复杂推理或跨领域结论,所以 Verifier 是必要的质量门。当前 Chat 链路通过 `Planner -> Executor -> Verifier` 固定流程,把 groundedness 和 facts checked 写入 `self_evaluation`。
|
||||
|
||||
AIOps 当前阶段先不加 Verifier,原因是:
|
||||
|
||||
- AIOps 刚完成从“自动跑告警”到“可追踪告警入口”的改造。
|
||||
- 先要确认告警 payload、工具证据、最终报告和 trace 能闭环。
|
||||
- AIOps Verifier 的规则不同于 Chat Verifier,需要检查告警 scope、证据覆盖和处置建议,不宜直接复用。
|
||||
|
||||
后续可以做 lightweight AIOps Verifier,检查报告是否聚焦 payload、是否引用工具证据、是否误展开无关告警。
|
||||
|
||||
## 3. 为什么 AIOps payload scope 先用 prompt 控制
|
||||
|
||||
运行验证发现:传入 `HighCPUUsage/payment-service` 后,Agent 仍可能把 mock Prometheus 返回的所有 active alerts 都展开分析。这个问题的本质是任务边界不清晰。
|
||||
|
||||
当前选择 prompt-level scope control:
|
||||
|
||||
- 有 payload:`PAYLOAD_TARGETED`,最终报告围绕传入告警。
|
||||
- 无 payload:`AUTO_DISCOVERY`,先调用 `queryPrometheusAlerts` 自动发现告警。
|
||||
|
||||
没有先做 Java 侧过滤,是因为:
|
||||
|
||||
- 过滤工具结果会降低 Agent 发现关联风险的能力。
|
||||
- 目前需要的是报告主线聚焦,而不是完全屏蔽上下文。
|
||||
- Prompt 改动小,风险低,能保留 Agent 灵活性。
|
||||
|
||||
已验证结果:主报告有 `HighCPUUsage/payment-service` 的完整根因分析,`HighMemoryUsage` 和 `SlowResponse` 只作为相关风险出现。
|
||||
|
||||
## 4. 为什么用 `tool_invocation` 统计真实工具调用次数
|
||||
|
||||
早期可以通过 `agent_step.hasToolCall` 粗略判断是否调用工具,但它统计的是“哪些模型步骤包含工具调用”,不是“真实调用了几次工具”。
|
||||
|
||||
现在 `tool_call_count` 来自:
|
||||
|
||||
```text
|
||||
ToolInvocationRepository.countBySessionId(sessionId)
|
||||
```
|
||||
|
||||
这样更符合 trace 语义:
|
||||
|
||||
- 一个 step 可能调用多个工具。
|
||||
- 工具可能来自不同来源:知识库、日志、指标、Prometheus。
|
||||
- 面试时可以把 `tool_call_count` 和 trace 中返回的工具明细对上。
|
||||
|
||||
## 5. 为什么保留 mock Prometheus 和 mock CLS
|
||||
|
||||
面试 Demo 最怕不稳定。真实 Prometheus、日志平台和线上故障都有不可控因素,所以 MVP profile 保留 mock 工具:
|
||||
|
||||
- `prometheus.mock-enabled=true`
|
||||
- `cls.mock-enabled=true`
|
||||
|
||||
这样可以稳定复现:
|
||||
|
||||
- `HighCPUUsage/payment-service`
|
||||
- `HighMemoryUsage/order-service`
|
||||
- `SlowResponse/user-service`
|
||||
- system-metrics、application-logs、database-slow-query 等日志证据
|
||||
|
||||
这不是逃避真实集成,而是把“Agent 编排和证据追踪”作为面试演示的主目标。
|
||||
|
||||
## 6. 为什么把面试材料单独放 `interview/`
|
||||
|
||||
`mvp/` 是持续迭代现场,包含过程文档、验收记录和 runbook。面试材料的目标不同,它应该是可讲、可演示、可评估的展示层。
|
||||
|
||||
因此:
|
||||
|
||||
- `mvp/` 保留真实演进材料。
|
||||
- `devflow/` 保留决策沉淀。
|
||||
- `interview/` 只组织面试叙事和演示脚本。
|
||||
|
||||
这样后续继续做 AIOps Verifier、UI、更多工具集成时,不会污染面试讲稿。
|
||||
|
||||
## 7. 可以主动承认的限制
|
||||
|
||||
- AIOps 还没有 Verifier。
|
||||
- Prompt-level scope control 不能做到强约束,只能通过 trace 和测试观察遵循情况。
|
||||
- 当前 mock 数据适合 demo,不代表生产接入已经完成。
|
||||
- Hikari 连接池已经加了短生命周期和 keepalive,但真实生产还需要按数据库 wait_timeout 和连接数预算调优。
|
||||
|
||||
主动讲清这些限制,反而能体现工程判断:先把可追踪闭环打通,再逐步增强质量门和生产可靠性。
|
||||
@@ -0,0 +1,81 @@
|
||||
---
|
||||
title: AIOps 告警排障 Runbook
|
||||
keywords: [AIOps, 告警, HighCPUUsage, SlowResponse, payment-service, system-metrics, application-logs]
|
||||
summary: 面向 AIOps 告警诊断的排障步骤,覆盖 Prometheus 活动告警、CLS 日志主题和处理建议。
|
||||
category: troubleshooting
|
||||
---
|
||||
|
||||
# AIOps 告警排障 Runbook
|
||||
|
||||
## 1. 告警输入处理原则
|
||||
|
||||
AIOps 诊断入口有两种触发方式:
|
||||
|
||||
- **有告警 payload**:将 payload 视为已触发告警,围绕 `alertName`、`service`、`severity`、`timeRange` 查询指标、日志和知识库。
|
||||
- **无告警 payload**:先调用 `queryPrometheusAlerts` 获取当前 firing 告警,再选择 P0/P1 或持续时间最长的告警进入诊断。
|
||||
|
||||
最终报告必须基于工具证据,不得凭空编造指标、日志或处理结果。
|
||||
|
||||
## 2. Mock 告警与日志主题映射
|
||||
|
||||
| 告警名 | 典型服务 | 优先日志主题 | 推荐查询 |
|
||||
|---|---|---|---|
|
||||
| HighCPUUsage | payment-service | system-metrics | `cpu_usage:>80 AND service:payment-service` |
|
||||
| HighMemoryUsage | order-service | system-metrics, system-events | `memory_usage:>85` |
|
||||
| SlowResponse | user-service | application-logs, database-slow-query | `duration:>3000 OR slow request` |
|
||||
| ServiceUnavailable | 任意核心服务 | application-logs, system-events | `level:ERROR OR container crash` |
|
||||
|
||||
## 3. HighCPUUsage / payment-service 排障步骤
|
||||
|
||||
### 3.1 现象确认
|
||||
|
||||
先确认 Prometheus 活动告警中是否存在:
|
||||
|
||||
- `alert_name = HighCPUUsage`
|
||||
- `service = payment-service`
|
||||
- CPU 使用率超过 80%
|
||||
- 状态为 firing
|
||||
|
||||
如果 payload 已经提供该告警,也仍需通过指标或日志工具验证。
|
||||
|
||||
### 3.2 指标与日志取证
|
||||
|
||||
推荐工具调用顺序:
|
||||
|
||||
1. `queryPrometheusAlerts`:确认当前活动告警。
|
||||
2. `queryLogs(region=ap-guangzhou, logTopic=system-metrics, query=cpu_usage:>80 AND service:payment-service)`:确认 CPU 使用率、实例和持续时间。
|
||||
3. 如报告中提到 Redis、数据库或下游依赖,再查询 `application-logs` 或对应主题交叉验证。
|
||||
|
||||
### 3.3 根因判断
|
||||
|
||||
可接受的根因结论必须至少满足一项:
|
||||
|
||||
- system-metrics 显示 payment-service 实例 CPU 使用率持续高于阈值。
|
||||
- application-logs 显示与 CPU 飙高同时出现的慢请求、线程池耗尽或依赖超时。
|
||||
- 告警持续时间与日志时间线一致。
|
||||
|
||||
如果只有活动告警,没有日志或指标明细,应输出低置信结论并建议人工确认。
|
||||
|
||||
## 4. 处理建议
|
||||
|
||||
### 临时止血
|
||||
|
||||
- 对 payment-service 做水平扩容,优先扩容受影响实例所在 Deployment。
|
||||
- 对高耗时接口开启限流或降级非核心功能。
|
||||
- 如果近期有发布,检查变更窗口并准备回滚。
|
||||
|
||||
### 根因修复
|
||||
|
||||
- 分析 CPU 热点线程、慢请求接口和依赖调用耗时。
|
||||
- 检查连接池、线程池、缓存穿透和批量任务是否导致 CPU 飙高。
|
||||
- 补充针对 `payment-service` 的 CPU、P95/P99 延迟、错误率和依赖超时联动告警。
|
||||
|
||||
## 5. 报告要求
|
||||
|
||||
告警分析报告至少包含:
|
||||
|
||||
- 活跃告警清单。
|
||||
- 告警根因分析。
|
||||
- 使用过的工具证据:Prometheus 告警、system-metrics 日志、application-logs 或知识库。
|
||||
- 已执行或建议执行的处理方案。
|
||||
- 置信度说明:哪些结论有直接证据,哪些需要人工进一步确认。
|
||||
@@ -2,6 +2,13 @@
|
||||
|
||||
This demo proves the MVP flow from user question to persisted diagnosis trace.
|
||||
|
||||
For interview use, start with:
|
||||
|
||||
- `interview-walkthrough.md` for the talk track
|
||||
- `trace-inspection-checklist.md` for fields to inspect
|
||||
- `scripts/run-payment-timeout-demo.ps1` for the runnable local demo
|
||||
- `requests/payment-timeout-chat.json` for the fixed request payload
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- MySQL, Redis, Milvus/Zilliz, and LLM/embedding configuration are available through the current project configuration.
|
||||
@@ -22,6 +29,22 @@ http://localhost:9900
|
||||
|
||||
## 1. Run Chat Diagnosis
|
||||
|
||||
Fast path:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
||||
```
|
||||
|
||||
This writes:
|
||||
|
||||
```text
|
||||
mvp/demo/output/chat-response.json
|
||||
mvp/demo/output/trace-response.json
|
||||
mvp/demo/output/feedback-response.json
|
||||
```
|
||||
|
||||
Manual path:
|
||||
|
||||
```powershell
|
||||
$sessionId = "mvp-demo-payment-timeout-001"
|
||||
$body = @{
|
||||
@@ -78,6 +101,43 @@ Expected result:
|
||||
- `success` is `true`.
|
||||
- A later trace query shows `data.session.feedback` as `useful`.
|
||||
|
||||
## 4. Run AIOps Alert Diagnosis
|
||||
|
||||
```powershell
|
||||
$aiopsSessionId = "mvp-demo-aiops-payment-cpu-001"
|
||||
$aiopsBody = @{
|
||||
sessionId = $aiopsSessionId
|
||||
alertName = "HighCPUUsage"
|
||||
service = "payment-service"
|
||||
severity = "P1"
|
||||
description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
|
||||
timeRange = "last_15m"
|
||||
userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-WebRequest `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/ai_ops" `
|
||||
-ContentType "application/json" `
|
||||
-Body $aiopsBody
|
||||
```
|
||||
|
||||
Expected result:
|
||||
|
||||
- The SSE stream starts with a `session` message containing `mvp-demo-aiops-payment-cpu-001`.
|
||||
- The stream later contains an AIOps alert analysis report focused on the supplied `HighCPUUsage/payment-service` payload.
|
||||
- A trace query for the same session id returns `data.session.agentFlow` as `AI_OPS`.
|
||||
- `data.session.answer` contains the final alert analysis report when a report is generated.
|
||||
- `data.toolInvocations` contains evidence tools such as `lookup_knowledge`, `query_logs`, or `query_metrics` when the runtime uses them.
|
||||
|
||||
Query the AIOps trace:
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace"
|
||||
```
|
||||
|
||||
## Demo Story
|
||||
|
||||
The important interview story is:
|
||||
@@ -92,3 +152,14 @@ one session id
|
||||
-> feedback
|
||||
-> trace API for replay and audit
|
||||
```
|
||||
|
||||
The AIOps story uses the same audit spine:
|
||||
|
||||
```text
|
||||
one session id
|
||||
-> alert payload
|
||||
-> AIOps planner/executor execution
|
||||
-> evidence tools
|
||||
-> alert analysis report
|
||||
-> trace API for replay and audit
|
||||
```
|
||||
|
||||
@@ -0,0 +1,38 @@
|
||||
# AIOps Alert Acceptance Case
|
||||
|
||||
## Goal
|
||||
|
||||
Validate that the legacy AIOps endpoint can act as a traceable alert-triggered diagnosis entry.
|
||||
|
||||
## Input
|
||||
|
||||
- Session id: `mvp-demo-aiops-payment-cpu-001`
|
||||
- Endpoint: `POST /api/ai_ops`
|
||||
- Profile: `mvp-demo`
|
||||
- Alert:
|
||||
|
||||
```json
|
||||
{
|
||||
"sessionId": "mvp-demo-aiops-payment-cpu-001",
|
||||
"alertName": "HighCPUUsage",
|
||||
"service": "payment-service",
|
||||
"severity": "P1",
|
||||
"description": "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。",
|
||||
"timeRange": "last_15m",
|
||||
"userRequest": "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
|
||||
}
|
||||
```
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
1. The SSE stream emits a `session` message containing the requested session id.
|
||||
2. The AIOps run creates or updates `diagnosis_session` with `agent_flow = AI_OPS`.
|
||||
3. The persisted session query contains the alert name, service, severity, time range, and description.
|
||||
4. If a final report is generated, `diagnosis_session.answer` contains that report.
|
||||
5. `GET /api/diagnosis/{sessionId}/trace` returns the AIOps session, ordered agent steps, and ordered tool invocations.
|
||||
6. In payload mode, the report focuses on `HighCPUUsage/payment-service`; unrelated active alerts may appear only as related risk or context, not as separate full root-cause sections.
|
||||
|
||||
## Known Limits
|
||||
|
||||
- This slice does not add a Verifier Agent to AIOps.
|
||||
- Full runtime verification still depends on valid DB, Redis, Milvus/Zilliz, model, and embedding configuration.
|
||||
@@ -0,0 +1,146 @@
|
||||
# Interview Walkthrough: MVP Diagnosis Agent
|
||||
|
||||
This walkthrough is the Plan C demo story. It is meant for a short Agent Engineer interview, not as exhaustive system documentation.
|
||||
|
||||
## 30-Second Summary
|
||||
|
||||
```text
|
||||
This is an enterprise diagnosis Agent MVP.
|
||||
It takes a payment-timeout question, plans the investigation, calls evidence tools,
|
||||
checks the answer through a verifier, persists the full trace, and accepts feedback.
|
||||
```
|
||||
|
||||
The important claim is not "the model answered once." The claim is:
|
||||
|
||||
```text
|
||||
The system can show what evidence was used, how the answer was checked, and how to replay the session.
|
||||
```
|
||||
|
||||
## Demo Flow
|
||||
|
||||
1. Start the service with the `mvp-demo` profile.
|
||||
2. Run the fixed payment-timeout request.
|
||||
3. Open `mvp/demo/output/chat-response.json`.
|
||||
4. Open `mvp/demo/output/trace-response.json`.
|
||||
5. Point to evidence tools and verifier evaluation.
|
||||
6. Submit feedback and show it is attached to the same session.
|
||||
|
||||
## Commands
|
||||
|
||||
Start service:
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||
```
|
||||
|
||||
Run the demo from another terminal:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
||||
```
|
||||
|
||||
Optional custom session:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1 -SessionId "mvp-demo-payment-timeout-002"
|
||||
```
|
||||
|
||||
## What To Show
|
||||
|
||||
### 1. User-Facing Answer
|
||||
|
||||
File:
|
||||
|
||||
```text
|
||||
mvp/demo/output/chat-response.json
|
||||
```
|
||||
|
||||
Say:
|
||||
|
||||
```text
|
||||
This is the answer the user sees. The session id is stable, so I can trace this exact answer later.
|
||||
```
|
||||
|
||||
### 2. Evidence Trace
|
||||
|
||||
File:
|
||||
|
||||
```text
|
||||
mvp/demo/output/trace-response.json
|
||||
```
|
||||
|
||||
Say:
|
||||
|
||||
```text
|
||||
This is the important Agent engineering part.
|
||||
I can inspect which tools were called, what inputs they received,
|
||||
whether they succeeded, and what evidence preview was persisted.
|
||||
```
|
||||
|
||||
Point to:
|
||||
|
||||
- `data.toolInvocations[*].toolName`
|
||||
- `data.toolInvocations[*].inputParams`
|
||||
- `data.toolInvocations[*].outputPreview`
|
||||
- `data.toolInvocations[*].success`
|
||||
|
||||
### 3. Verifier / Self-Evaluation
|
||||
|
||||
Point to:
|
||||
|
||||
- `data.session.selfEvaluation`
|
||||
- `data.summary.hasVerifierEvaluation`
|
||||
|
||||
Say:
|
||||
|
||||
```text
|
||||
The final answer is not just raw Executor output.
|
||||
It is checked by a verifier or self-evaluation layer using the persisted trace.
|
||||
That lets the system return PASS, LOW_CONFID, or REJECT-style behavior instead of pretending all answers are equally certain.
|
||||
```
|
||||
|
||||
### 4. Feedback Loop
|
||||
|
||||
File:
|
||||
|
||||
```text
|
||||
mvp/demo/output/feedback-response.json
|
||||
```
|
||||
|
||||
Then re-query trace if needed.
|
||||
|
||||
Say:
|
||||
|
||||
```text
|
||||
Feedback is attached to the same diagnosis session.
|
||||
That makes it possible to mine useful / not useful cases later.
|
||||
```
|
||||
|
||||
### 5. Regression Story
|
||||
|
||||
Mention, do not deep dive unless asked:
|
||||
|
||||
```text
|
||||
For repeatability, I also built an offline eval baseline.
|
||||
The demo proves the runtime trace; the eval baseline proves fixed-case regression.
|
||||
The two are separate on purpose: demo for human review, eval for automated signal.
|
||||
```
|
||||
|
||||
## Strong Interview Framing
|
||||
|
||||
Use this phrasing:
|
||||
|
||||
```text
|
||||
I focused on the Agent engineering surface:
|
||||
traceability, evidence persistence, verifier gating, feedback, and regression checks.
|
||||
The model answer is only one part of the system.
|
||||
The more important part is whether we can audit and improve the answer after it is produced.
|
||||
```
|
||||
|
||||
## Known Limits To Say Proactively
|
||||
|
||||
```text
|
||||
This MVP still depends on configured MySQL, Redis, Milvus, and model credentials.
|
||||
The mvp-demo profile mocks logs and metrics, but not the full application runtime.
|
||||
Secret cleanup and fully isolated default tests are separate production-hardening tasks.
|
||||
```
|
||||
@@ -0,0 +1,11 @@
|
||||
# Demo Output
|
||||
|
||||
This directory is the default output location for local demo responses.
|
||||
|
||||
Generated files are intentionally ignored by Git:
|
||||
|
||||
- `chat-response.json`
|
||||
- `trace-response.json`
|
||||
- `feedback-response.json`
|
||||
|
||||
Keep this README so the directory exists in the repository.
|
||||
@@ -0,0 +1,4 @@
|
||||
{
|
||||
"Id": "mvp-demo-payment-timeout-001",
|
||||
"Question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
|
||||
}
|
||||
@@ -0,0 +1,54 @@
|
||||
param(
|
||||
[string]$BaseUrl = "http://localhost:9900",
|
||||
[string]$SessionId = "mvp-demo-payment-timeout-001",
|
||||
[string]$RequestFile = "$PSScriptRoot/../requests/payment-timeout-chat.json",
|
||||
[string]$OutputDir = "$PSScriptRoot/../output"
|
||||
)
|
||||
|
||||
$ErrorActionPreference = "Stop"
|
||||
|
||||
New-Item -ItemType Directory -Force -Path $OutputDir | Out-Null
|
||||
|
||||
$request = Get-Content -Raw -Encoding UTF8 -Path $RequestFile | ConvertFrom-Json
|
||||
$request.Id = $SessionId
|
||||
$body = $request | ConvertTo-Json -Depth 8
|
||||
|
||||
Write-Host "Running payment-timeout chat demo..."
|
||||
Write-Host "BaseUrl: $BaseUrl"
|
||||
Write-Host "SessionId: $SessionId"
|
||||
|
||||
$chat = Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "$BaseUrl/api/chat" `
|
||||
-ContentType "application/json; charset=utf-8" `
|
||||
-Body $body
|
||||
|
||||
$chat | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-response.json"
|
||||
Write-Host "Saved chat response: $OutputDir/chat-response.json"
|
||||
|
||||
$trace = Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "$BaseUrl/api/diagnosis/$SessionId/trace"
|
||||
|
||||
$trace | ConvertTo-Json -Depth 50 | Set-Content -Encoding UTF8 -Path "$OutputDir/trace-response.json"
|
||||
Write-Host "Saved trace response: $OutputDir/trace-response.json"
|
||||
|
||||
$feedbackBody = @{
|
||||
sessionId = $SessionId
|
||||
feedback = "useful"
|
||||
} | ConvertTo-Json
|
||||
|
||||
$feedback = Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "$BaseUrl/api/feedback" `
|
||||
-ContentType "application/json; charset=utf-8" `
|
||||
-Body $feedbackBody
|
||||
|
||||
$feedback | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/feedback-response.json"
|
||||
Write-Host "Saved feedback response: $OutputDir/feedback-response.json"
|
||||
|
||||
Write-Host ""
|
||||
Write-Host "Demo completed. Review:"
|
||||
Write-Host "- mvp/demo/output/chat-response.json"
|
||||
Write-Host "- mvp/demo/output/trace-response.json"
|
||||
Write-Host "- mvp/demo/output/feedback-response.json"
|
||||
@@ -0,0 +1,52 @@
|
||||
# Trace Inspection Checklist
|
||||
|
||||
Use this checklist after running `scripts/run-payment-timeout-demo.ps1`.
|
||||
|
||||
## Session
|
||||
|
||||
| JSON path | What to check | Interview point |
|
||||
| --- | --- | --- |
|
||||
| `data.session.sessionId` | Matches `mvp-demo-payment-timeout-001` | One session id connects chat, tools, verifier, feedback, and trace. |
|
||||
| `data.session.query` | Contains the payment-timeout question | The trace records the original user intent. |
|
||||
| `data.session.answer` | Contains the final diagnosis answer | The final answer is not detached from the trace. |
|
||||
| `data.session.selfEvaluation` | Contains verifier or rule evaluation | The answer has a quality gate, not just raw model output. |
|
||||
| `data.session.feedback` | Becomes `useful` after feedback submission | User feedback is attached to the same diagnosis session. |
|
||||
|
||||
## Agent Steps
|
||||
|
||||
| JSON path | What to check | Interview point |
|
||||
| --- | --- | --- |
|
||||
| `data.steps[*].agentName` | Planner / Executor / Verifier or equivalent step names | The flow is decomposed into inspectable Agent steps. |
|
||||
| `data.steps[*].thought` | High-level step reasoning where available | Internal reasoning is auditable without relying only on final text. |
|
||||
| `data.steps[*].durationMs` | Step duration | The trace can support cost and latency review. |
|
||||
| `data.steps[*].tokenCount` | Token count where available | The trace can support model-cost review. |
|
||||
|
||||
## Tool Evidence
|
||||
|
||||
| JSON path | What to check | Interview point |
|
||||
| --- | --- | --- |
|
||||
| `data.toolInvocations[*].toolName` | Includes evidence tools such as `lookup_knowledge`, `query_logs`, `query_metrics` | The Agent uses tools, not unsupported guesses. |
|
||||
| `data.toolInvocations[*].inputParams` | Shows what each tool was asked | Inputs are inspectable for debugging and audit. |
|
||||
| `data.toolInvocations[*].outputPreview` | Shows a bounded preview of evidence | Evidence is preserved without dumping huge payloads. |
|
||||
| `data.toolInvocations[*].success` | Distinguishes success from failure | Tool failure is visible to verifier and reviewers. |
|
||||
| `data.toolInvocations[*].retrievalDetails` | Shows retrieval metadata when available | Retrieval quality can be reviewed after the fact. |
|
||||
|
||||
## Summary
|
||||
|
||||
| JSON path | What to check | Interview point |
|
||||
| --- | --- | --- |
|
||||
| `data.summary.persistedStepCount` | Step rows were persisted | The trace is backed by storage, not only response memory. |
|
||||
| `data.summary.persistedToolCallCount` | Tool rows were persisted | Evidence survives the request. |
|
||||
| `data.summary.hasVerifierEvaluation` | Verifier evaluation exists | The final answer passed through a quality gate. |
|
||||
| `data.summary.hasFeedback` | Feedback exists after feedback step | Human feedback closes the loop. |
|
||||
|
||||
## What Good Looks Like
|
||||
|
||||
```text
|
||||
same session id
|
||||
-> final answer
|
||||
-> persisted agent steps
|
||||
-> persisted evidence tool calls
|
||||
-> verifier/self-evaluation
|
||||
-> feedback attached to the same session
|
||||
```
|
||||
@@ -0,0 +1,61 @@
|
||||
# Diagnosis Eval Harness
|
||||
|
||||
This folder contains the first fixed-case evaluation set for the MVP diagnosis Agent.
|
||||
|
||||
## Scope
|
||||
|
||||
- Case definitions: `cases/diagnosis-cases.json`
|
||||
- Offline trace fixtures: `fixtures/*.json`
|
||||
- Field definitions: `schema.md`
|
||||
- Baseline reports: `reports/baseline-report.json` and `reports/baseline-report.md`
|
||||
- Baseline diff sample: `reports/baseline-diff-sample.json` and `reports/baseline-diff-sample.md`
|
||||
- Evaluator implementation: `DiagnosisTraceEvaluator`
|
||||
- Report writer: `DiagnosisEvalReportWriter`
|
||||
|
||||
## Current Mode
|
||||
|
||||
The first version evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.
|
||||
|
||||
## Verification
|
||||
|
||||
Run the focused evaluator test:
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test
|
||||
```
|
||||
|
||||
The committed baseline report represents the current fixed fixture set:
|
||||
|
||||
```text
|
||||
5 fixed cases
|
||||
5 passing fixture evaluations
|
||||
2 PASS verdicts
|
||||
3 LOW_CONFID verdicts
|
||||
```
|
||||
|
||||
When fixtures or evaluator rules change, regenerate the report from the same case file and fixture directory, then update both JSON and Markdown outputs together.
|
||||
|
||||
## Interview Story
|
||||
|
||||
The harness gives the MVP a repeatable baseline:
|
||||
|
||||
```text
|
||||
fixed diagnosis case
|
||||
-> saved or runtime trace
|
||||
-> rule-based trace validation
|
||||
-> JSON / Markdown report
|
||||
-> regression signal for prompts, tools, retrieval, and verifier behavior
|
||||
```
|
||||
|
||||
## Baseline Diff
|
||||
|
||||
Baseline diff compares a current report against `reports/baseline-report.json`.
|
||||
|
||||
```text
|
||||
baseline report
|
||||
current report
|
||||
-> deterministic diff
|
||||
-> regressions, improvements, and changed signals
|
||||
```
|
||||
|
||||
Use it to answer: did a prompt, tool, retrieval, or verifier change make the Agent worse than the fixed baseline?
|
||||
@@ -0,0 +1,57 @@
|
||||
[
|
||||
{
|
||||
"id": "payment-timeout",
|
||||
"title": "Payment API timeout",
|
||||
"question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
|
||||
"traceFixture": "payment-timeout-pass.json",
|
||||
"expectedRootCauseKeywords": ["支付", "超时", "连接池"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["lookup_knowledge", "query_logs", "query_metrics"],
|
||||
"allowedVerdicts": ["PASS", "LOW_CONFID"],
|
||||
"forbiddenAnswerKeywords": ["无证据确定"]
|
||||
},
|
||||
{
|
||||
"id": "mysql-pool-exhausted",
|
||||
"title": "MySQL connection pool exhausted",
|
||||
"question": "订单服务大量请求超时,请判断是否和 MySQL 连接池有关。",
|
||||
"traceFixture": "mysql-pool-low-confid.json",
|
||||
"expectedRootCauseKeywords": ["mysql", "连接池", "超时"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["lookup_knowledge", "query_logs"],
|
||||
"allowedVerdicts": ["LOW_CONFID", "PASS"],
|
||||
"forbiddenAnswerKeywords": ["已经完全确认"]
|
||||
},
|
||||
{
|
||||
"id": "redis-timeout",
|
||||
"title": "Redis timeout",
|
||||
"question": "支付服务出现 Redis 连接超时,请定位可能原因。",
|
||||
"traceFixture": "redis-timeout-low-confid.json",
|
||||
"expectedRootCauseKeywords": ["redis", "超时"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_logs"],
|
||||
"allowedVerdicts": ["LOW_CONFID", "PASS"],
|
||||
"forbiddenAnswerKeywords": ["无需进一步排查"]
|
||||
},
|
||||
{
|
||||
"id": "slow-response",
|
||||
"title": "Slow response",
|
||||
"question": "用户服务 P99 响应时间升高,请结合指标和日志分析。",
|
||||
"traceFixture": "slow-response-pass.json",
|
||||
"expectedRootCauseKeywords": ["p99", "慢响应"],
|
||||
"minKeywordMatches": 1,
|
||||
"requiredEvidenceTools": ["query_metrics", "query_logs"],
|
||||
"allowedVerdicts": ["LOW_CONFID", "PASS"],
|
||||
"forbiddenAnswerKeywords": ["没有风险"]
|
||||
},
|
||||
{
|
||||
"id": "jvm-memory-risk",
|
||||
"title": "JVM memory risk",
|
||||
"question": "订单服务内存使用率过高,请判断是否存在 OOM 风险。",
|
||||
"traceFixture": "jvm-memory-risk-low-confid.json",
|
||||
"expectedRootCauseKeywords": ["jvm", "内存", "oom"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_metrics", "query_logs"],
|
||||
"allowedVerdicts": ["LOW_CONFID", "PASS"],
|
||||
"forbiddenAnswerKeywords": ["可以忽略"]
|
||||
}
|
||||
]
|
||||
@@ -0,0 +1,52 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-jvm-memory-risk",
|
||||
"query": "订单服务内存使用率过高,请判断是否存在 OOM 风险。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 53000,
|
||||
"toolCallCount": 2,
|
||||
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n订单服务存在 JVM 内存风险,但还不能完全确认会发生 OOM。指标显示 heap 使用率持续高于 88%,日志出现多次 Full GC 和 allocation pressure,需要继续观察对象增长来源并检查最近发布。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.52,
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_metrics",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
},
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "indirect"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-jvm-memory-risk",
|
||||
"toolName": "query_metrics",
|
||||
"success": true
|
||||
},
|
||||
{
|
||||
"id": 2,
|
||||
"sessionId": "eval-jvm-memory-risk",
|
||||
"toolName": "query_logs",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 2,
|
||||
"returnedToolCallCount": 2,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,53 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-mysql-pool",
|
||||
"query": "订单服务大量请求超时,请判断是否和 MySQL 连接池有关。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 51000,
|
||||
"toolCallCount": 2,
|
||||
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\nMySQL 连接池可能参与了本次超时问题。日志中出现 connection pool exhausted,但当前缺少完整指标证据,因此只能作为低置信结论处理。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.48,
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "lookup_knowledge",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
},
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-mysql-pool",
|
||||
"toolName": "lookup_knowledge",
|
||||
"success": true,
|
||||
"relevanceLevel": "PRECISE"
|
||||
},
|
||||
{
|
||||
"id": 2,
|
||||
"sessionId": "eval-mysql-pool",
|
||||
"toolName": "query_logs",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 2,
|
||||
"returnedToolCallCount": 2,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,64 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-payment-timeout",
|
||||
"query": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 42000,
|
||||
"toolCallCount": 3,
|
||||
"answer": "支付接口超时与连接池等待有关。知识库说明支付超时需要同时检查连接池、日志和指标;日志出现 connection pool exhausted;指标显示支付服务延迟升高。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 0.86,
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "lookup_knowledge",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
},
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
},
|
||||
{
|
||||
"tool_name": "query_metrics",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-payment-timeout",
|
||||
"toolName": "lookup_knowledge",
|
||||
"success": true,
|
||||
"relevanceLevel": "PRECISE"
|
||||
},
|
||||
{
|
||||
"id": 2,
|
||||
"sessionId": "eval-payment-timeout",
|
||||
"toolName": "query_logs",
|
||||
"success": true
|
||||
},
|
||||
{
|
||||
"id": 3,
|
||||
"sessionId": "eval-payment-timeout",
|
||||
"toolName": "query_metrics",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 3,
|
||||
"returnedToolCallCount": 3,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,41 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-redis-timeout",
|
||||
"query": "支付服务出现 Redis 连接超时,请定位可能原因。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 36000,
|
||||
"toolCallCount": 1,
|
||||
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\nRedis 连接超时可能和支付服务到 Redis 的网络抖动或连接池等待有关。日志中出现 redis timeout 和 command timeout 记录,但当前缺少指标侧证据,因此只能作为低置信结论处理。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.46,
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-redis-timeout",
|
||||
"toolName": "query_logs",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 2,
|
||||
"returnedStepCount": 2,
|
||||
"persistedToolCallCount": 1,
|
||||
"returnedToolCallCount": 1,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,52 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-slow-response",
|
||||
"query": "用户服务 P99 响应时间升高,请结合指标和日志分析。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 47000,
|
||||
"toolCallCount": 2,
|
||||
"answer": "用户服务 P99 升高主要表现为慢响应。指标显示 P99 latency 从 280ms 上升到 1800ms,日志中同时出现 slow request 和 downstream timeout,因此优先排查下游依赖耗时和线程池排队。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 0.78,
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_metrics",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
},
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-slow-response",
|
||||
"toolName": "query_metrics",
|
||||
"success": true
|
||||
},
|
||||
{
|
||||
"id": 2,
|
||||
"sessionId": "eval-slow-response",
|
||||
"toolName": "query_logs",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 2,
|
||||
"returnedToolCallCount": 2,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,85 @@
|
||||
{
|
||||
"baselineTotalCases" : 5,
|
||||
"currentTotalCases" : 5,
|
||||
"baselinePassedCases" : 5,
|
||||
"currentPassedCases" : 4,
|
||||
"baselinePassRate" : 1.0,
|
||||
"currentPassRate" : 0.8,
|
||||
"regressionCount" : 6,
|
||||
"improvementCount" : 0,
|
||||
"changedCount" : 2,
|
||||
"hasRegression" : true,
|
||||
"items" : [ {
|
||||
"type" : "REGRESSION",
|
||||
"scope" : "aggregate",
|
||||
"caseId" : null,
|
||||
"metric" : "passRate",
|
||||
"baselineValue" : "1.0",
|
||||
"currentValue" : "0.8",
|
||||
"delta" : -0.19999999999999996,
|
||||
"message" : "passRate changed"
|
||||
}, {
|
||||
"type" : "REGRESSION",
|
||||
"scope" : "aggregate",
|
||||
"caseId" : null,
|
||||
"metric" : "averageToolCallCount",
|
||||
"baselineValue" : "2.0",
|
||||
"currentValue" : "3.0",
|
||||
"delta" : 1.0,
|
||||
"message" : "averageToolCallCount changed"
|
||||
}, {
|
||||
"type" : "CHANGED",
|
||||
"scope" : "aggregate",
|
||||
"caseId" : null,
|
||||
"metric" : "verdictDistribution.LOW_CONFID",
|
||||
"baselineValue" : "3",
|
||||
"currentValue" : "2",
|
||||
"delta" : -1.0,
|
||||
"message" : "verdict count changed for LOW_CONFID"
|
||||
}, {
|
||||
"type" : "CHANGED",
|
||||
"scope" : "aggregate",
|
||||
"caseId" : null,
|
||||
"metric" : "verdictDistribution.REJECT",
|
||||
"baselineValue" : "0",
|
||||
"currentValue" : "1",
|
||||
"delta" : 1.0,
|
||||
"message" : "verdict count changed for REJECT"
|
||||
}, {
|
||||
"type" : "REGRESSION",
|
||||
"scope" : "case",
|
||||
"caseId" : "redis-timeout",
|
||||
"metric" : "passed",
|
||||
"baselineValue" : "true",
|
||||
"currentValue" : "false",
|
||||
"delta" : null,
|
||||
"message" : "redis-timeout pass state changed"
|
||||
}, {
|
||||
"type" : "REGRESSION",
|
||||
"scope" : "case",
|
||||
"caseId" : "redis-timeout",
|
||||
"metric" : "verdict",
|
||||
"baselineValue" : "LOW_CONFID",
|
||||
"currentValue" : "REJECT",
|
||||
"delta" : -1.0,
|
||||
"message" : "redis-timeout verdict changed"
|
||||
}, {
|
||||
"type" : "REGRESSION",
|
||||
"scope" : "case",
|
||||
"caseId" : "redis-timeout",
|
||||
"metric" : "matchedKeywordCount",
|
||||
"baselineValue" : "2",
|
||||
"currentValue" : "1",
|
||||
"delta" : -1.0,
|
||||
"message" : "redis-timeout matchedKeywordCount changed"
|
||||
}, {
|
||||
"type" : "REGRESSION",
|
||||
"scope" : "case",
|
||||
"caseId" : "redis-timeout",
|
||||
"metric" : "evidenceCoverage.query_logs",
|
||||
"baselineValue" : "true",
|
||||
"currentValue" : "false",
|
||||
"delta" : null,
|
||||
"message" : "redis-timeout evidence coverage changed for query_logs"
|
||||
} ]
|
||||
}
|
||||
@@ -0,0 +1,22 @@
|
||||
# Diagnosis Eval Baseline Diff
|
||||
|
||||
- Baseline pass rate: 100.00%
|
||||
- Current pass rate: 80.00%
|
||||
- Baseline passed cases: 5/5
|
||||
- Current passed cases: 4/5
|
||||
- Regressions: 6
|
||||
- Improvements: 0
|
||||
- Other changes: 2
|
||||
|
||||
## Diff Items
|
||||
|
||||
| Type | Scope | Case | Metric | Baseline | Current | Delta | Message |
|
||||
| --- | --- | --- | --- | --- | --- | ---: | --- |
|
||||
| REGRESSION | aggregate | - | passRate | 1.0 | 0.8 | -0.200 | passRate changed |
|
||||
| REGRESSION | aggregate | - | averageToolCallCount | 2.0 | 3.0 | 1.000 | averageToolCallCount changed |
|
||||
| CHANGED | aggregate | - | verdictDistribution.LOW_CONFID | 3 | 2 | -1.000 | verdict count changed for LOW_CONFID |
|
||||
| CHANGED | aggregate | - | verdictDistribution.REJECT | 0 | 1 | 1.000 | verdict count changed for REJECT |
|
||||
| REGRESSION | case | redis-timeout | passed | true | false | - | redis-timeout pass state changed |
|
||||
| REGRESSION | case | redis-timeout | verdict | LOW_CONFID | REJECT | -1.000 | redis-timeout verdict changed |
|
||||
| REGRESSION | case | redis-timeout | matchedKeywordCount | 2 | 1 | -1.000 | redis-timeout matchedKeywordCount changed |
|
||||
| REGRESSION | case | redis-timeout | evidenceCoverage.query_logs | true | false | - | redis-timeout evidence coverage changed for query_logs |
|
||||
@@ -0,0 +1,82 @@
|
||||
{
|
||||
"totalCases" : 5,
|
||||
"passedCases" : 5,
|
||||
"passRate" : 1.0,
|
||||
"verdictDistribution" : {
|
||||
"PASS" : 2,
|
||||
"LOW_CONFID" : 3
|
||||
},
|
||||
"averageToolCallCount" : 2.0,
|
||||
"averageDurationMs" : 45800.0,
|
||||
"results" : [ {
|
||||
"caseId" : "payment-timeout",
|
||||
"title" : "Payment API timeout",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "PASS",
|
||||
"matchedKeywordCount" : 3,
|
||||
"requiredKeywordCount" : 3,
|
||||
"evidenceCoverage" : {
|
||||
"lookup_knowledge" : true,
|
||||
"query_logs" : true,
|
||||
"query_metrics" : true
|
||||
},
|
||||
"toolCallCount" : 3,
|
||||
"durationMs" : 42000
|
||||
}, {
|
||||
"caseId" : "mysql-pool-exhausted",
|
||||
"title" : "MySQL connection pool exhausted",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "LOW_CONFID",
|
||||
"matchedKeywordCount" : 3,
|
||||
"requiredKeywordCount" : 3,
|
||||
"evidenceCoverage" : {
|
||||
"lookup_knowledge" : true,
|
||||
"query_logs" : true
|
||||
},
|
||||
"toolCallCount" : 2,
|
||||
"durationMs" : 51000
|
||||
}, {
|
||||
"caseId" : "redis-timeout",
|
||||
"title" : "Redis timeout",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "LOW_CONFID",
|
||||
"matchedKeywordCount" : 2,
|
||||
"requiredKeywordCount" : 2,
|
||||
"evidenceCoverage" : {
|
||||
"query_logs" : true
|
||||
},
|
||||
"toolCallCount" : 1,
|
||||
"durationMs" : 36000
|
||||
}, {
|
||||
"caseId" : "slow-response",
|
||||
"title" : "Slow response",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "PASS",
|
||||
"matchedKeywordCount" : 2,
|
||||
"requiredKeywordCount" : 2,
|
||||
"evidenceCoverage" : {
|
||||
"query_metrics" : true,
|
||||
"query_logs" : true
|
||||
},
|
||||
"toolCallCount" : 2,
|
||||
"durationMs" : 47000
|
||||
}, {
|
||||
"caseId" : "jvm-memory-risk",
|
||||
"title" : "JVM memory risk",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "LOW_CONFID",
|
||||
"matchedKeywordCount" : 3,
|
||||
"requiredKeywordCount" : 3,
|
||||
"evidenceCoverage" : {
|
||||
"query_metrics" : true,
|
||||
"query_logs" : true
|
||||
},
|
||||
"toolCallCount" : 2,
|
||||
"durationMs" : 53000
|
||||
} ]
|
||||
}
|
||||
@@ -0,0 +1,22 @@
|
||||
# Diagnosis Eval Report
|
||||
|
||||
- Total cases: 5
|
||||
- Passed cases: 5
|
||||
- Pass rate: 100.00%
|
||||
- Average tool calls: 2.00
|
||||
- Average duration ms: 45800.00
|
||||
|
||||
## Verdict Distribution
|
||||
|
||||
- PASS: 2
|
||||
- LOW_CONFID: 3
|
||||
|
||||
## Cases
|
||||
|
||||
| Case | Result | Verdict | Keywords | Tool Calls | Duration ms | Failed Checks |
|
||||
| --- | --- | --- | --- | ---: | ---: | --- |
|
||||
| payment-timeout | PASS | PASS | 3/3 | 3 | 42000 | - |
|
||||
| mysql-pool-exhausted | PASS | LOW_CONFID | 3/3 | 2 | 51000 | - |
|
||||
| redis-timeout | PASS | LOW_CONFID | 2/2 | 1 | 36000 | - |
|
||||
| slow-response | PASS | PASS | 2/2 | 2 | 47000 | - |
|
||||
| jvm-memory-risk | PASS | LOW_CONFID | 3/3 | 2 | 53000 | - |
|
||||
@@ -0,0 +1,202 @@
|
||||
# Diagnosis Eval Data Schema
|
||||
|
||||
这份文档记录评测基准里的数据结构。口语化理解就是:
|
||||
|
||||
```text
|
||||
用例文件说“我要考什么”
|
||||
trace 文件说“Agent 实际做了什么”
|
||||
评测结果说“这次有没有跑偏”
|
||||
汇总报告说“整体稳定性怎么样”
|
||||
```
|
||||
|
||||
当前这套评测是代码规则判断,不是再调用一个 LLM 来打分。
|
||||
|
||||
## 1. 用例定义
|
||||
|
||||
文件:`mvp/eval/cases/diagnosis-cases.json`
|
||||
|
||||
每一条 case 是一个固定考题,告诉评测器“这个问题应该看哪些点、需要哪些证据、哪些结论可以接受”。
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "payment-timeout",
|
||||
"title": "Payment API timeout",
|
||||
"question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
|
||||
"traceFixture": "payment-timeout-pass.json",
|
||||
"expectedRootCauseKeywords": ["支付", "超时", "连接池"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["lookup_knowledge", "query_logs", "query_metrics"],
|
||||
"allowedVerdicts": ["PASS", "LOW_CONFID"],
|
||||
"forbiddenAnswerKeywords": ["无证据确定"]
|
||||
}
|
||||
```
|
||||
|
||||
字段说明:
|
||||
|
||||
| 字段 | 意思 | 评测器怎么用 |
|
||||
| --- | --- | --- |
|
||||
| `id` | 这条用例的唯一名字 | 出现在报告里,方便定位是哪条 case 挂了 |
|
||||
| `title` | 给人看的标题 | 出现在结果里,方便快速理解场景 |
|
||||
| `question` | 要问 Agent 的问题 | fixture 模式下不会真的发送给 Agent,但它记录了这条 case 的原始输入 |
|
||||
| `traceFixture` | 对应的 trace 文件名 | 评测器会去 `fixtures/` 目录加载这个文件 |
|
||||
| `expectedRootCauseKeywords` | 最终回答里希望看到的关键点 | 评测器会在 `session.answer` 里做关键词命中检查 |
|
||||
| `minKeywordMatches` | 至少要命中几个关键词 | 命中数低于这个值,就认为根因覆盖不够 |
|
||||
| `requiredEvidenceTools` | 这条 case 至少应该用到哪些证据工具 | 评测器会检查 trace 里是否出现这些工具 |
|
||||
| `allowedVerdicts` | Verifier 允许给出的结论 | 比如 `PASS` 或 `LOW_CONFID`,不在列表里就失败 |
|
||||
| `forbiddenAnswerKeywords` | 回答里不应该出现的危险说法 | 命中这些词,说明回答可能过度自信或不符合降级策略 |
|
||||
|
||||
## 2. Trace Fixture
|
||||
|
||||
目录:`mvp/eval/fixtures/*.json`
|
||||
|
||||
trace fixture 是一次 Agent 运行后的“留痕快照”。评测器不会关心整个 trace 的所有字段,只读取当前能支撑基准判断的字段。
|
||||
|
||||
当前会读取这些字段:
|
||||
|
||||
| Trace 字段 | 意思 | 评测器怎么用 |
|
||||
| --- | --- | --- |
|
||||
| `session.answer` | Agent 最终给用户的回答 | 用来检查根因关键词和禁用词 |
|
||||
| `session.totalDurationMs` | 这次运行耗时 | 进入报告,帮助观察性能是否明显变差 |
|
||||
| `session.selfEvaluation.verifier_evaluation.verdict` | Verifier 对最终回答的判断 | 必须存在,并且要落在 case 的 `allowedVerdicts` 里 |
|
||||
| `session.selfEvaluation.verifier_evaluation.tool_trace_summary[*].tool_name` | Verifier 总结里看到的工具证据 | 用来补充判断证据工具是否出现 |
|
||||
| `toolInvocations[*].toolName` | Agent 实际调用过的工具名 | 用来检查 `requiredEvidenceTools` 是否满足 |
|
||||
| `toolInvocations[*].success` | 工具调用是否成功 | 当前主要保留在 trace 里,后续可以升级成更严格的成功率检查 |
|
||||
|
||||
简单说,trace 里最重要的是三类信息:
|
||||
|
||||
```text
|
||||
最终回答:它说了什么
|
||||
工具证据:它查了什么
|
||||
Verifier:它自己有没有承认这个结论可靠
|
||||
```
|
||||
|
||||
## 3. 单条评测结果
|
||||
|
||||
Java 类型:`DiagnosisEvalResult`
|
||||
|
||||
这是每条 case 跑完之后的判断结果。
|
||||
|
||||
| 字段 | 意思 |
|
||||
| --- | --- |
|
||||
| `caseId` | 对应的 case id |
|
||||
| `title` | case 标题 |
|
||||
| `passed` | 这条 case 是否通过 |
|
||||
| `failedChecks` | 没通过的具体原因,比如缺工具、关键词不够、verdict 不允许 |
|
||||
| `verdict` | 从 trace 里读出来的 Verifier verdict |
|
||||
| `matchedKeywordCount` | 最终回答命中的关键词数量 |
|
||||
| `requiredKeywordCount` | case 定义里一共有多少个关键词 |
|
||||
| `evidenceCoverage` | 每个必需工具是否出现,例如 `{ "query_logs": true }` |
|
||||
| `toolCallCount` | 本次 trace 里工具调用总数 |
|
||||
| `durationMs` | 本次 trace 的耗时 |
|
||||
|
||||
判断通过的口语化规则:
|
||||
|
||||
```text
|
||||
回答要说到关键点
|
||||
该查的证据工具要查到
|
||||
Verifier 的结论要在可接受范围内
|
||||
回答不能出现危险的过度自信表达
|
||||
如果是 REJECT,就必须走降级模板
|
||||
```
|
||||
|
||||
## 4. 汇总报告
|
||||
|
||||
Java 类型:`DiagnosisEvalReport`
|
||||
|
||||
这是整个基准集跑完之后的总结果。
|
||||
|
||||
| 字段 | 意思 |
|
||||
| --- | --- |
|
||||
| `totalCases` | 总共评测了多少条 case |
|
||||
| `passedCases` | 通过了多少条 |
|
||||
| `passRate` | 通过率,范围是 `0.0` 到 `1.0` |
|
||||
| `verdictDistribution` | Verifier verdict 的分布,比如有几个 `PASS`、几个 `LOW_CONFID` |
|
||||
| `averageToolCallCount` | 平均每条 case 调用了多少次工具 |
|
||||
| `averageDurationMs` | 平均耗时 |
|
||||
| `results` | 每条 case 的详细结果列表 |
|
||||
|
||||
## 5. 怎么看这个基准
|
||||
|
||||
这套结构不是为了证明 Agent 永远正确,而是为了在每次改 prompt、工具、检索、Verifier 之后,有一个固定尺子能回答:
|
||||
|
||||
```text
|
||||
以前能过的诊断题,现在还过不过?
|
||||
它是不是少查了某些证据?
|
||||
它是不是变得更自信但证据不足?
|
||||
它是不是开始输出不该说的话?
|
||||
它是不是明显变慢了?
|
||||
```
|
||||
|
||||
所以面试里可以这样讲:
|
||||
|
||||
```text
|
||||
我没有只看一次 demo 效果,而是把典型诊断场景固化成 case。
|
||||
每条 case 都定义预期关键点、必需证据工具和可接受的 verifier 结论。
|
||||
Agent 每次运行会留下 trace,评测器用代码规则读取 trace,输出结构化报告。
|
||||
这样我改 Agent 的时候,可以用同一套基准判断有没有行为回退。
|
||||
```
|
||||
|
||||
## 6. Baseline Diff
|
||||
|
||||
Baseline diff 是拿两份 report 做对比:
|
||||
|
||||
```text
|
||||
baseline report:以前认可的基准结果
|
||||
current report:这次改动后跑出来的新结果
|
||||
diff report:告诉你哪里变好了、哪里变差了、哪里只是变了
|
||||
```
|
||||
|
||||
Java 类型:
|
||||
|
||||
- `DiagnosisEvalDiffReport`
|
||||
- `DiagnosisEvalDiffItem`
|
||||
|
||||
`DiagnosisEvalDiffReport` 字段:
|
||||
|
||||
| 字段 | 意思 |
|
||||
| --- | --- |
|
||||
| `baselineTotalCases` | baseline 里有多少条 case |
|
||||
| `currentTotalCases` | current 里有多少条 case |
|
||||
| `baselinePassedCases` | baseline 通过了多少条 |
|
||||
| `currentPassedCases` | current 通过了多少条 |
|
||||
| `baselinePassRate` | baseline 通过率 |
|
||||
| `currentPassRate` | current 通过率 |
|
||||
| `regressionCount` | 退化项数量 |
|
||||
| `improvementCount` | 改善项数量 |
|
||||
| `changedCount` | 普通变化项数量 |
|
||||
| `hasRegression` | 是否存在退化 |
|
||||
| `items` | 具体 diff 明细 |
|
||||
|
||||
`DiagnosisEvalDiffItem` 字段:
|
||||
|
||||
| 字段 | 意思 |
|
||||
| --- | --- |
|
||||
| `type` | `REGRESSION`、`IMPROVEMENT` 或 `CHANGED` |
|
||||
| `scope` | `aggregate` 表示整体指标,`case` 表示单条 case |
|
||||
| `caseId` | 如果是单条 case 变化,这里记录 case id |
|
||||
| `metric` | 哪个指标变了,比如 `passRate` 或 `evidenceCoverage.query_logs` |
|
||||
| `baselineValue` | baseline 里的值 |
|
||||
| `currentValue` | current 里的值 |
|
||||
| `delta` | 数值变化量;非数值变化为空 |
|
||||
| `message` | 给人看的变化说明 |
|
||||
|
||||
口语化判断规则:
|
||||
|
||||
```text
|
||||
pass rate 下降:退化
|
||||
case 从通过变失败:退化
|
||||
证据工具从有变没有:退化
|
||||
关键词命中变少:退化
|
||||
工具调用或耗时升高:成本上升,记为退化信号
|
||||
verdict 分布变化:记录变化,供人工判断是否符合预期
|
||||
```
|
||||
|
||||
面试里可以这样讲:
|
||||
|
||||
```text
|
||||
我把 baseline report 和当前 report 做结构化 diff。
|
||||
它不是再问 LLM,而是用代码比较固定字段。
|
||||
如果某个 case 从 PASS 变 FAIL,或者 query_logs 证据没了,
|
||||
diff 会直接标成 regression。
|
||||
这样 Agent 改动可以用固定基准做回归判断。
|
||||
```
|
||||
@@ -0,0 +1,137 @@
|
||||
# ISS-005 证据链补齐与降级契约收敛
|
||||
|
||||
**状态**:进行中(sm-flow)
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-04
|
||||
**来源**:P1-A 面试打磨项 / 基于 ISS-003 的当前实现复核
|
||||
**关联**:ISS-003(Verifier 证据链、失败路径可验证性)、`chat-verifier-agent`、`mvp-demo-trace-acceptance`
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
当前 MVP 已具备:
|
||||
|
||||
- `lookup_knowledge`、`query_logs`、`query_metrics` 的工具调用落库
|
||||
- Verifier 基于 `tool_trace_summary` 做事实核查
|
||||
- `LOW_CONFID` / `REJECT` 的用户侧降级输出
|
||||
- trace API 可回放 session、agent_step、tool_invocation 和 self_evaluation
|
||||
|
||||
但如果目标是拿这个项目去面试 Agent 工程师,当前实现仍有一个明显短板:
|
||||
|
||||
**证据链已经“有了”,但还没有被收敛成清晰、稳定、可测试的工程契约。**
|
||||
|
||||
这会直接影响三个面试问题的回答质量:
|
||||
|
||||
1. 工具失败时系统会怎样降级?
|
||||
2. Verifier 看到的 evidence 到底是否一致、可审计?
|
||||
3. 这些失败路径和降级行为有没有稳定测试,而不是只靠 runtime 演示?
|
||||
|
||||
---
|
||||
|
||||
## 当前现状复核
|
||||
|
||||
### 1. 工具落库入口已经存在,但契约不统一
|
||||
|
||||
- `QueryLogsTools` 和 `QueryMetricsTools` 通过 `ToolInvocationRecorder.recordEvidenceTool(...)` 记录 evidence tool 调用。
|
||||
- `LookupKnowledgeTool` 仍保留独立的 `saveToolInvocation(...)` 路径,自己构造 `ToolInvocation` 实体。
|
||||
|
||||
这意味着:
|
||||
|
||||
- evidence tool 的公共字段有一套约定
|
||||
- knowledge retrieval 又有一套定制字段拼装
|
||||
|
||||
两者都能工作,但**没有形成统一的“证据调用记录契约”**。
|
||||
|
||||
### 2. 失败 / 无结果 / 去重命中的语义不够显式
|
||||
|
||||
当前实现里:
|
||||
|
||||
- `query_logs` 未命中时会返回 `success=false` + `"未找到匹配的日志"`
|
||||
- `query_metrics` 失败时会返回 `success=false`
|
||||
- `lookup_knowledge` 去重命中时会返回 `found=false`,但 `tool_invocation.success=true`
|
||||
- `ToolTraceSummaryService` 通过 `success`、`relevanceLevel`、`dedupReason` 等字段做启发式摘要
|
||||
|
||||
这些行为在代码里是分散成立的,但**没有被定义成统一契约**,导致:
|
||||
|
||||
- Verifier 能看到的“失败”和“无证据”边界不够稳定
|
||||
- 评测时难以明确统计哪些是“调用失败”、哪些是“无命中”、哪些是“已检索过”
|
||||
|
||||
### 3. ChatService 的降级路径有实现,但测试矩阵不完整
|
||||
|
||||
`ChatService` 已处理:
|
||||
|
||||
- `verifier_output` 缺失或无法解析 → fallback `LOW_CONFID`
|
||||
- `REJECT` → degraded output
|
||||
- `LOW_CONFID` → disclaimer output
|
||||
|
||||
但目前缺少成体系的专项验证,尤其是:
|
||||
|
||||
- Verifier 输出非法 JSON
|
||||
- evidence tool 查询失败
|
||||
- knowledge lookup 无有效证据
|
||||
- fallback 文案是否只基于 verifier 缺口拼装
|
||||
|
||||
---
|
||||
|
||||
## 影响
|
||||
|
||||
- **面试表达弱化**:你能讲“我有 trace”,但还不能很硬地讲“我的失败路径是有契约和测试保护的”。
|
||||
- **评测基础不稳**:后续 P1-B 做 case-based harness 时,统计口径会受 evidence 语义不一致影响。
|
||||
- **Verifier 可审计性打折**:当前实现可用,但 still relies on code convention,而不是一份明确收敛后的工程协议。
|
||||
|
||||
---
|
||||
|
||||
## 本 issue 目标
|
||||
|
||||
P1-A 只做三件事:
|
||||
|
||||
1. 收敛 evidence tool 的落库契约,让 `lookup_knowledge`、`query_logs`、`query_metrics` 的公共语义一致。
|
||||
2. 明确失败 / 无证据 / 去重 / verifier 非法输出等降级契约,让 `ToolTraceSummaryService` 和 `ChatService` 面向统一状态工作。
|
||||
3. 增加专项离线测试,覆盖证据摘要与关键降级路径。
|
||||
|
||||
---
|
||||
|
||||
## 范围
|
||||
|
||||
### In scope
|
||||
|
||||
- `ToolInvocationRecorder` 契约增强
|
||||
- `LookupKnowledgeTool` 入库路径收敛
|
||||
- `QueryLogsTools` / `QueryMetricsTools` evidence 语义对齐
|
||||
- `ToolTraceSummaryService` 对失败 / no-hit / mixed evidence 的摘要规则收敛
|
||||
- `ChatService` 对 verifier 非法输出与降级输出的专项测试
|
||||
- 与该 change 直接相关的文档、OpenSpec、devflow 记录
|
||||
|
||||
### Out of scope
|
||||
|
||||
- 不引入新的数据库表或 schema 变更
|
||||
- 不扩展新的 evidence tool
|
||||
- 不做 P1-B 评测集 / harness
|
||||
- 不做前端 trace UI
|
||||
- 不处理敏感配置和默认 `mvn test` 离线化
|
||||
|
||||
---
|
||||
|
||||
## 预期结果
|
||||
|
||||
完成后,项目在面试里应能更清楚地表述为:
|
||||
|
||||
```text
|
||||
我不仅把 Agent 的工具调用落到了库里,
|
||||
还把 evidence trace、失败语义和 verifier 降级路径收敛成了稳定契约,
|
||||
并用离线测试覆盖了这些关键失败场景。
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
- `src/main/java/com/superbiz/agent/agent/tool/QueryLogsTools.java`
|
||||
- `src/main/java/com/superbiz/agent/agent/tool/QueryMetricsTools.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ToolTraceSummaryService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`
|
||||
- `src/test/java/com/superbiz/agent/tool/LookupKnowledgeToolTest.java`
|
||||
@@ -0,0 +1,99 @@
|
||||
# ISS-006 固定诊断评测集与回归 Harness
|
||||
|
||||
**状态**:进行中(sm-flow)
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-04
|
||||
**来源**:P1-B 面试打磨项
|
||||
**依赖**:ISS-005 / `evidence-trace-hardening`
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
MVP 已经具备可追溯证据链、Verifier 质量门禁、trace API 和固定 demo 流程。上一阶段 `evidence-trace-hardening` 进一步统一了 evidence tool 的状态语义,让系统能稳定区分:
|
||||
|
||||
- `supported`
|
||||
- `no_evidence`
|
||||
- `deduped`
|
||||
- `failed`
|
||||
|
||||
下一步需要证明 Agent 在一组固定诊断场景下的表现,而不是只依赖单次 demo。
|
||||
|
||||
---
|
||||
|
||||
## 问题
|
||||
|
||||
当前项目能演示一次支付超时诊断,但还缺少稳定的评测基线:
|
||||
|
||||
- 每次改 prompt、工具、Verifier 或检索逻辑后,无法快速判断是否退化。
|
||||
- 只能人工看 trace,缺少结构化通过 / 失败结果。
|
||||
- 缺少面试时能展示的指标,如 evidence coverage、verdict 分布、工具调用数量和耗时。
|
||||
|
||||
---
|
||||
|
||||
## 目标
|
||||
|
||||
建立一个轻量的固定 case 评测 harness,用于验证 MVP Agent 的诊断质量和证据链完整性。
|
||||
|
||||
第一版不做 LLM-as-judge,优先做规则化校验:
|
||||
|
||||
- 固定 5 个 MVP 诊断 case
|
||||
- 每个 case 定义 expected root-cause keywords、required evidence tools、allowed verdicts
|
||||
- 基于 trace 结果校验 evidence coverage、verifier evaluation、tool invocation、final answer shape
|
||||
- 输出 JSON 和 Markdown 报告
|
||||
|
||||
---
|
||||
|
||||
## 范围
|
||||
|
||||
### In scope
|
||||
|
||||
- 评测 case 定义文件
|
||||
- trace 规则校验器
|
||||
- eval runner 或测试入口
|
||||
- JSON / Markdown 报告输出
|
||||
- demo 文档和 devflow 记录
|
||||
|
||||
### Out of scope
|
||||
|
||||
- 不引入 LLM-as-judge
|
||||
- 不要求完整离线 LLM runtime
|
||||
- 不新增生产 API
|
||||
- 不修改 Chat 主链路
|
||||
- 不修改 evidence trace 运行时语义
|
||||
|
||||
---
|
||||
|
||||
## 预期面试表达
|
||||
|
||||
完成后可以这样描述:
|
||||
|
||||
```text
|
||||
我不仅有一个可演示的 Agent,还给它建立了固定 case 的回归评测。
|
||||
每次修改 prompt、工具或 verifier 后,都可以跑同一批诊断 case,
|
||||
检查证据覆盖、verdict 分布、工具调用成本和关键结论是否退化。
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 初始候选 case
|
||||
|
||||
| Case | 目标 |
|
||||
| --- | --- |
|
||||
| payment-timeout | 支付接口超时,验证知识库 + 日志 + 指标证据 |
|
||||
| mysql-pool-exhausted | 数据库连接池耗尽,验证日志和知识库证据 |
|
||||
| redis-timeout | Redis 连接超时,验证日志依赖证据 |
|
||||
| slow-response | P99 响应时间过高,验证指标 + 慢请求日志 |
|
||||
| jvm-memory-risk | JVM 内存 / OOM 风险,验证指标 + 系统事件日志 |
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `mvp/demo/README.md`
|
||||
- `mvp/demo/payment-timeout-acceptance.md`
|
||||
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ToolTraceSummaryService.java`
|
||||
- `src/main/java/com/superbiz/agent/domain/entity/DiagnosisSession.java`
|
||||
- `src/main/java/com/superbiz/agent/domain/entity/ToolInvocation.java`
|
||||
- `openspec/specs/evidence-trace-hardening/spec.md`
|
||||
@@ -6,3 +6,37 @@
|
||||
| ISS-002 | Executor 无约束重复调用 lookup_knowledge | 中 | 已修复 | [ISS-002-executor-unconstrained-lookup.md](ISS-002-executor-unconstrained-lookup.md) |
|
||||
| ISS-003 | MVP 设计与实现 Review 收敛 | 高 | 待规划 | [ISS-003-mvp-design-implementation-review.md](ISS-003-mvp-design-implementation-review.md) |
|
||||
| ISS-004 | Executor 域级检索水位控制(Phase 2) | 低 | 待规划 | [ISS-004-executor-domain-hard-limit.md](ISS-004-executor-domain-hard-limit.md) |
|
||||
| ISS-005 | 证据链补齐与降级契约收敛 | 高 | 已归档 | [ISS-005-evidence-trace-hardening.md](ISS-005-evidence-trace-hardening.md) |
|
||||
| ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) |
|
||||
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) |
|
||||
| diagnosis-eval-baseline-diff | 诊断评测 baseline diff 与回归判断 | 中 | 已归档 | [diagnosis-eval-baseline-diff.md](diagnosis-eval-baseline-diff.md) |
|
||||
| mvp-demo-interview-runbook | Plan C 面试可复现 Demo 包 | 中 | 已归档 | [mvp-demo-interview-runbook.md](mvp-demo-interview-runbook.md) |
|
||||
|
||||
## RAG 重构计划
|
||||
|
||||
| 名称 | 标题 | 严重程度 | 状态 | 文件 |
|
||||
|---|---|---|---|---|
|
||||
| rag-refactor-plan | RAG 检索重构计划 | 高 | 待规划 | [rag-refactor-plan.md](rag-refactor-plan.md) |
|
||||
|
||||
## RAG 检索问题
|
||||
|
||||
| 名称 | 标题 | 严重程度 | 状态 | 文件 |
|
||||
|---|---|---|---|---|
|
||||
| chunk-context-reconstruction | RAG 切片上下文重建缺失 | 高 | 已合并到重构计划 | [rag-chunk-context-reconstruction.md](rag-chunk-context-reconstruction.md) |
|
||||
| breadcrumb-embedding-gap | RAG breadcrumb 未参与向量语义 | 高 | 已合并到重构计划 | [rag-breadcrumb-embedding-gap.md](rag-breadcrumb-embedding-gap.md) |
|
||||
| l0-l1-fusion-ranking | RAG L0 和 L1 未真正融合排序 | 中 | 已合并到重构计划 | [rag-l0-l1-fusion-ranking.md](rag-l0-l1-fusion-ranking.md) |
|
||||
| l0-keyword-matching-quality | RAG L0 关键词匹配质量不足 | 中 | 已合并到重构计划 | [rag-l0-keyword-matching-quality.md](rag-l0-keyword-matching-quality.md) |
|
||||
| l1-score-calibration | RAG L1 分数阈值未校准 | 中 | 已合并到重构计划 | [rag-l1-score-calibration.md](rag-l1-score-calibration.md) |
|
||||
| context-packing-and-reranking | RAG 缺少上下文打包和 Rerank | 中 | 已合并到重构计划 | [rag-context-packing-and-reranking.md](rag-context-packing-and-reranking.md) |
|
||||
| upload-chunk-parameter-drift | RAG 上传切片参数未真正生效 | 低 | 已合并到重构计划 | [rag-upload-chunk-parameter-drift.md](rag-upload-chunk-parameter-drift.md) |
|
||||
| query-rewrite-gap | RAG 查询改写能力薄弱 | 中 | 已合并到重构计划 | [rag-query-rewrite-gap.md](rag-query-rewrite-gap.md) |
|
||||
|
||||
## RAG 框架化改造
|
||||
|
||||
| 名称 | 标题 | 严重程度 | 状态 | 文件 |
|
||||
|---|---|---|---|---|
|
||||
| spring-ai-vectorstore-migration | RAG 迁移到 Spring AI VectorStore 检索抽象 | 高 | 已合并到重构计划 | [rag-spring-ai-vectorstore-migration.md](rag-spring-ai-vectorstore-migration.md) |
|
||||
| spring-ai-query-transformer | RAG 接入 Spring AI Query Transformer | 中 | 已合并到重构计划 | [rag-spring-ai-query-transformer.md](rag-spring-ai-query-transformer.md) |
|
||||
| spring-ai-document-postprocessor | RAG 使用 DocumentPostProcessor 做后处理 | 中 | 已合并到重构计划 | [rag-spring-ai-document-postprocessor.md](rag-spring-ai-document-postprocessor.md) |
|
||||
| l0-domain-entity-hint | RAG 将 L0 降级为领域和实体 Hint | 中 | 已合并到重构计划 | [rag-l0-domain-entity-hint.md](rag-l0-domain-entity-hint.md) |
|
||||
| spring-ai-advisor-boundary | RAG 明确 Spring AI Advisor 与 Agent Tool 的边界 | 中 | 已合并到重构计划 | [rag-spring-ai-advisor-boundary.md](rag-spring-ai-advisor-boundary.md) |
|
||||
|
||||
@@ -0,0 +1,73 @@
|
||||
# Diagnosis Eval Baseline Diff
|
||||
|
||||
**状态**:已归档
|
||||
**严重程度**:中
|
||||
**发现时间**:2026-07-05
|
||||
**来源**:P1-B follow-up
|
||||
**依赖**:`diagnosis-eval-harness`, `expand-diagnosis-eval-fixtures`
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
现在项目已经有固定诊断 case、完整 fixture 和 baseline report。下一步需要把 baseline 真正用起来:每次改 Agent 后,把新的 report 和 baseline report 做对比。
|
||||
|
||||
---
|
||||
|
||||
## 问题
|
||||
|
||||
当前 baseline 只能告诉我们“标准状态是什么”,但还不能自动告诉我们“这次改动有没有变差”。
|
||||
|
||||
典型问题包括:
|
||||
|
||||
- pass rate 是否下降。
|
||||
- 某个 case 是否从通过变失败。
|
||||
- 某个 evidence tool 是否从覆盖变成缺失。
|
||||
- verifier verdict 分布是否异常变化。
|
||||
- 平均工具调用数和耗时是否明显上升。
|
||||
|
||||
---
|
||||
|
||||
## 目标
|
||||
|
||||
新增一个 deterministic baseline diff 能力,用代码比较两份 `DiagnosisEvalReport`。
|
||||
|
||||
完成后应该做到:
|
||||
|
||||
- 输入 baseline report 和 current report。
|
||||
- 输出结构化 diff。
|
||||
- 标出 regression、improvement 和普通 changed。
|
||||
- 支持 JSON 和 Markdown 输出。
|
||||
- 文档说明面试时怎么解释这套回归判断。
|
||||
|
||||
---
|
||||
|
||||
## 范围
|
||||
|
||||
### In scope
|
||||
|
||||
- report-level diff 数据结构。
|
||||
- aggregate 指标比较。
|
||||
- case-level 指标比较。
|
||||
- JSON / Markdown diff writer。
|
||||
- focused tests 和 eval 文档。
|
||||
|
||||
### Out of scope
|
||||
|
||||
- 不运行真实 Agent。
|
||||
- 不生成新 trace。
|
||||
- 不引入 LLM-as-judge。
|
||||
- 不改现有 evaluator 评分规则。
|
||||
|
||||
---
|
||||
|
||||
## 面试表达
|
||||
|
||||
可以这样讲:
|
||||
|
||||
```text
|
||||
我不是只保存了一份 baseline,而是加了 baseline diff。
|
||||
每次改 prompt、tool、retrieval 或 verifier 后,
|
||||
我都能把新 report 和 baseline 比较,
|
||||
直接看到哪些 case 退化、哪些证据缺失、成本有没有上升。
|
||||
```
|
||||
@@ -0,0 +1,83 @@
|
||||
# Expand Diagnosis Eval Fixtures
|
||||
|
||||
**状态**:已归档
|
||||
**严重程度**:中
|
||||
**发现时间**:2026-07-04
|
||||
**来源**:P1-B follow-up
|
||||
**依赖**:`diagnosis-eval-harness`
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
`diagnosis-eval-harness` 已经把固定 case、trace evaluator、JSON / Markdown report 和字段文档搭起来了。
|
||||
|
||||
现在还差一步:5 条固定诊断 case 里,只有 2 条有 fixture,另外 3 条还是 missing 状态。这个状态可以验证 evaluator 的错误报告能力,但还不能作为完整 baseline 展示。
|
||||
|
||||
---
|
||||
|
||||
## 问题
|
||||
|
||||
当前 baseline 还不够完整:
|
||||
|
||||
- `redis-timeout` 没有对应 trace fixture。
|
||||
- `slow-response` 没有对应 trace fixture。
|
||||
- `jvm-memory-risk` 没有对应 trace fixture。
|
||||
- 仓库里还没有一份固定的 baseline JSON / Markdown 报告可供对比。
|
||||
|
||||
---
|
||||
|
||||
## 目标
|
||||
|
||||
补齐固定诊断评测集,让它从“框架可跑”变成“基准可用”。
|
||||
|
||||
完成后应该做到:
|
||||
|
||||
- 5 条固定 case 都能加载到对应 fixture。
|
||||
- evaluator 能输出完整 baseline report。
|
||||
- baseline report 被保存到仓库,后续 Agent 改动可以拿它做对比。
|
||||
- 文档说明怎么重新生成和怎么看报告。
|
||||
|
||||
---
|
||||
|
||||
## 范围
|
||||
|
||||
### In scope
|
||||
|
||||
- 补齐 3 个缺失 fixture。
|
||||
- 保存 baseline JSON / Markdown 报告。
|
||||
- 更新 eval 文档。
|
||||
- 补充测试,确保 case 文件引用的 fixture 都存在。
|
||||
|
||||
### Out of scope
|
||||
|
||||
- 不新增 case 数量。
|
||||
- 不改生产 Agent 主链路。
|
||||
- 不引入 LLM-as-judge。
|
||||
- 不启动真实 MySQL、Redis、Milvus 或 LLM。
|
||||
|
||||
---
|
||||
|
||||
## 面试表达
|
||||
|
||||
可以这样讲:
|
||||
|
||||
```text
|
||||
我先搭了评测 harness,然后把固定 case 的 trace fixture 补齐,
|
||||
生成一份可复现的 baseline report。
|
||||
这样以后每次改 prompt、tool 或 verifier,
|
||||
都能看固定诊断集有没有行为回退,而不是只靠人工感觉。
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `mvp/eval/cases/diagnosis-cases.json`
|
||||
- `mvp/eval/fixtures/`
|
||||
- `mvp/eval/reports/`
|
||||
- `mvp/eval/README.md`
|
||||
- `mvp/eval/schema.md`
|
||||
- `src/main/java/com/superbiz/agent/eval/DiagnosisTraceEvaluator.java`
|
||||
- `src/test/java/com/superbiz/agent/eval/DiagnosisTraceEvaluatorTest.java`
|
||||
- `openspec/specs/diagnosis-eval-harness/spec.md`
|
||||
@@ -0,0 +1,53 @@
|
||||
# MVP Demo Interview Runbook
|
||||
|
||||
**状态**:已归档
|
||||
**严重程度**:中
|
||||
**发现时间**:2026-07-05
|
||||
**来源**:Plan C
|
||||
**依赖**:`mvp-demo-trace-acceptance`, `evidence-trace-hardening`, `diagnosis-eval-harness`
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
项目已经有 Agent 主链路、证据 trace、Verifier、反馈、eval baseline,但这些材料分散在不同目录。面试时真正需要的是一个能快速跑、快速讲清楚的 demo 入口。
|
||||
|
||||
---
|
||||
|
||||
## 问题
|
||||
|
||||
当前 demo 还不够“面试友好”:
|
||||
|
||||
- 启动、请求、trace、反馈步骤分散在文档里。
|
||||
- 没有固定请求 payload 文件。
|
||||
- 没有一键跑 payment-timeout demo 的脚本。
|
||||
- 没有把 trace 字段和面试讲法对应起来的 walkthrough。
|
||||
|
||||
---
|
||||
|
||||
## 目标
|
||||
|
||||
把 Plan C 落地成 `mvp/demo` 下的可复现 demo 包:
|
||||
|
||||
- 固定支付超时请求。
|
||||
- 一键执行 chat、trace、feedback。
|
||||
- 保存 demo 输出,便于复盘。
|
||||
- 提供面试讲解稿和 trace 检查清单。
|
||||
|
||||
---
|
||||
|
||||
## 范围
|
||||
|
||||
### In scope
|
||||
|
||||
- `mvp/demo` 文档。
|
||||
- `mvp/demo/requests` 请求文件。
|
||||
- `mvp/demo/scripts` PowerShell 脚本。
|
||||
- `mvp/demo/output` 目录说明。
|
||||
|
||||
### Out of scope
|
||||
|
||||
- 不新增后端 API。
|
||||
- 不改 Agent prompt。
|
||||
- 不扩 eval harness。
|
||||
- 不处理密钥外置和完整离线化。
|
||||
@@ -0,0 +1,58 @@
|
||||
# RAG breadcrumb 未参与向量语义
|
||||
|
||||
**状态**:待规划
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-04
|
||||
**范围**:向量化输入、检索相关性、知识库 metadata 使用
|
||||
|
||||
---
|
||||
|
||||
## 现象
|
||||
|
||||
当前 chunk metadata 中保存了 `title` 和 `breadcrumb`,但向量化时主要使用 `chunk.getContent()`。这意味着标题层级、所属模块、章节路径没有进入 embedding 语义空间。
|
||||
|
||||
当用户问题依赖章节语境时,例如“诊断流程里的验证步骤是什么”,如果 chunk 正文里没有重复出现完整标题语义,向量召回可能无法稳定命中正确片段。
|
||||
|
||||
---
|
||||
|
||||
## 当前实现
|
||||
|
||||
- `DocumentChunkService` 会生成 `breadcrumb`。
|
||||
- `VectorIndexService` 会把 `breadcrumb` 写入 metadata。
|
||||
- `VectorEmbeddingService` 接收的 embedding 内容来自 chunk 正文。
|
||||
- `VectorSearchService` 只基于 query embedding 和 chunk embedding 做向量搜索。
|
||||
|
||||
metadata 目前更像是展示和追踪字段,不是检索相关性的一部分。
|
||||
|
||||
---
|
||||
|
||||
## 影响
|
||||
|
||||
- 标题语义丢失,尤其影响短段落、步骤列表、配置表格类 chunk。
|
||||
- 同名概念出现在不同章节时,缺少章节路径帮助 disambiguation。
|
||||
- 用户问的是“某个模块下的问题”,检索可能只看正文关键词,忽略模块归属。
|
||||
|
||||
---
|
||||
|
||||
## 建议修复
|
||||
|
||||
构建面向 embedding 的增强文本:
|
||||
|
||||
```text
|
||||
标题: {title}
|
||||
路径: {breadcrumb}
|
||||
正文:
|
||||
{content}
|
||||
```
|
||||
|
||||
落库时仍保留原始 `content`,避免展示内容被污染。可以新增 `embeddingText` 构造逻辑,只用于向量化。
|
||||
|
||||
后续还可以在 rerank 阶段把 `breadcrumb` 作为加权信号,例如同域、同章节、同文档优先。
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `src/main/java/com/superbiz/agent/service/DocumentChunkService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/VectorIndexService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/VectorEmbeddingService.java`
|
||||
@@ -0,0 +1,55 @@
|
||||
# RAG 切片上下文重建缺失
|
||||
|
||||
**状态**:待规划
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-04
|
||||
**范围**:知识库上传、切片、向量召回、Agent 上下文组装
|
||||
|
||||
---
|
||||
|
||||
## 现象
|
||||
|
||||
同一个 Markdown 章节在内容较长时会被拆成多个 chunk。当前检索命中其中一个 chunk 后,返回给 Agent 的主要是单个 chunk 内容,不会自动把同章节的前后片段、章节标题链路、相邻 chunk 一起恢复出来。
|
||||
|
||||
这会导致两个问题:
|
||||
|
||||
1. 命中片段只包含局部语义,缺少前置定义、约束条件或后续步骤。
|
||||
2. 同章节被分段后,检索结果之间缺少可追溯的关联,Agent 不一定知道它们属于同一章节。
|
||||
|
||||
---
|
||||
|
||||
## 当前实现
|
||||
|
||||
- `DocumentChunkService` 会按 Markdown 标题建立 `title` 和 `breadcrumb`,再按段落累积切片。
|
||||
- 超过阈值时仍会切断同一章节,只是尽量避免打断代码块和列表。
|
||||
- `VectorIndexService` 会把 `chunkIndex`、`totalChunks`、`title`、`breadcrumb` 放入 metadata。
|
||||
- `VectorSearchService` 查询 Milvus 后直接返回命中的 chunk,没有做相邻 chunk 扩展或 section 级聚合。
|
||||
- `LookupKnowledgeTool` 消费 L1 结果时,也没有根据 `docId + chunkIndex + breadcrumb` 回补上下文。
|
||||
|
||||
---
|
||||
|
||||
## 影响
|
||||
|
||||
- RAG 回答容易漏掉同章节中的约束条件。
|
||||
- 长流程类文档会被拆散,Agent 看到的是“片段证据”,不是“完整流程”。
|
||||
- 面试解释中需要承认:当前系统有 metadata 基础,但还没有把它用于上下文重建。
|
||||
|
||||
---
|
||||
|
||||
## 建议修复
|
||||
|
||||
优先做命中后的上下文扩展:
|
||||
|
||||
1. L1 命中 chunk 后,按 `docId + chunkIndex` 拉取前后 N 个相邻 chunk。
|
||||
2. 如果 metadata 中 `breadcrumb` 相同,允许扩展到同章节的多个 chunk。
|
||||
3. 上下文打包时标记 `命中片段`、`前文`、`后文`,避免 Agent 把扩展内容误认为全部都是高置信命中。
|
||||
4. 增加 token budget 控制,超过预算时优先保留命中 chunk 和标题链路。
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `src/main/java/com/superbiz/agent/service/DocumentChunkService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/VectorIndexService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
@@ -0,0 +1,48 @@
|
||||
# RAG 缺少上下文打包和 Rerank
|
||||
|
||||
**状态**:待规划
|
||||
**严重程度**:中
|
||||
**发现时间**:2026-07-04
|
||||
**范围**:检索后处理、证据排序、Agent 输入质量
|
||||
|
||||
---
|
||||
|
||||
## 现象
|
||||
|
||||
当前 RAG 检索主要依赖 L0/L1 的原始召回顺序,没有独立的 reranker、cross-encoder 或 LLM rerank 阶段。召回结果进入 Agent 前,也缺少统一的上下文打包策略。
|
||||
|
||||
这意味着“检索到”不等于“以最适合推理的形式喂给 Agent”。
|
||||
|
||||
---
|
||||
|
||||
## 当前实现
|
||||
|
||||
- L0 和 L1 结果由 `LookupKnowledgeTool` 拼装后返回。
|
||||
- 没有候选级 rerank。
|
||||
- 没有明确的 token budget 分配策略,例如每个文档最多占多少、命中片段和扩展片段如何排序。
|
||||
- 没有把 `title`、`breadcrumb`、score、source 统一包装成证据块。
|
||||
|
||||
---
|
||||
|
||||
## 影响
|
||||
|
||||
- 相关结果可能被排在不理想的位置。
|
||||
- 多个候选内容相近时,Agent 可能读到重复信息。
|
||||
- 证据结构不清晰,后续 verifier 或 trace 解释成本较高。
|
||||
|
||||
---
|
||||
|
||||
## 建议修复
|
||||
|
||||
1. 引入 `RetrievedEvidence` 这样的内部结构,统一承载 source、title、breadcrumb、score、hitReason、content。
|
||||
2. 做简单 rerank:关键词命中、向量分、breadcrumb 匹配、文档去重、相邻片段扩展一起排序。
|
||||
3. 上下文打包时按证据块输出,明确来源和置信度。
|
||||
4. 面试版可以先实现规则 rerank,后续再替换为 cross-encoder 或 LLM rerank。
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java`
|
||||
@@ -0,0 +1,71 @@
|
||||
# RAG 将 L0 降级为领域和实体 Hint
|
||||
|
||||
**状态**:待规划
|
||||
**严重程度**:中
|
||||
**发现时间**:2026-07-05
|
||||
**范围**:L0 检索、metadata filter、业务可解释性
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
当前 L0 是基于 frontmatter / keyword 的轻量检索。它具备可解释性,但不适合作为最终相关性判断。
|
||||
|
||||
在引入 Spring AI VectorStore / Retriever 后,L0 更适合从“召回主链路”调整为“检索前处理和解释信号”。
|
||||
|
||||
---
|
||||
|
||||
## 问题
|
||||
|
||||
当前 L0 如果唯一命中,容易被过度信任:
|
||||
|
||||
```text
|
||||
L0 unique hit -> 直接返回 / 优先采信
|
||||
```
|
||||
|
||||
这会带来误召回风险,尤其是关键词过泛、frontmatter 质量不稳定时。
|
||||
|
||||
---
|
||||
|
||||
## 改造方向
|
||||
|
||||
L0 保留,但职责调整为:
|
||||
|
||||
1. **Domain detector**
|
||||
- 识别 query 所属 category/domain。
|
||||
- 用于 Spring AI retriever metadata filter。
|
||||
|
||||
2. **Entity extractor**
|
||||
- 识别组件名、服务名、指标名、错误码、接口名。
|
||||
- 用于 query augmentation。
|
||||
|
||||
3. **Explainability signal**
|
||||
- 记录 matched keywords。
|
||||
- 解释为什么进入某个知识域。
|
||||
|
||||
目标链路:
|
||||
|
||||
```text
|
||||
query / payload
|
||||
-> L0 domain/entity hint
|
||||
-> metadata filter + query augmentation
|
||||
-> vector retriever
|
||||
-> post processor
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 验收标准
|
||||
|
||||
- L0 不再默认作为最终检索结果直接返回。
|
||||
- L0 命中的 domain/category 能传给 retriever filter。
|
||||
- L0 命中的实体能进入增强 query 或工具调用记录。
|
||||
- `tool_invocation` 能展示 L0 matched keywords 和使用方式。
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java`
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
|
||||
@@ -0,0 +1,53 @@
|
||||
# RAG L0 关键词匹配质量不足
|
||||
|
||||
**状态**:待规划
|
||||
**严重程度**:中
|
||||
**发现时间**:2026-07-04
|
||||
**范围**:L0 索引、frontmatter、精确召回
|
||||
|
||||
---
|
||||
|
||||
## 现象
|
||||
|
||||
L0 当前依赖文档 frontmatter 中的关键词,并使用较粗的字符串包含逻辑做匹配。关键词质量越依赖人工维护,召回稳定性越容易波动。
|
||||
|
||||
如果 frontmatter 填写不完整、同义词缺失、关键词过短或过泛,L0 就可能误召回或漏召回。
|
||||
|
||||
---
|
||||
|
||||
## 当前实现
|
||||
|
||||
- `KnowledgeIndexService` 从 `ApiDocument` metadata/frontmatter 加载关键词。
|
||||
- exact match 的判断类似:
|
||||
|
||||
```java
|
||||
query.contains(keywordLower) || keywordLower.contains(query)
|
||||
```
|
||||
|
||||
- 没有分词、同义词归一、字段权重、关键词质量校验。
|
||||
|
||||
---
|
||||
|
||||
## 影响
|
||||
|
||||
- 短关键词容易误命中。
|
||||
- 用户换一种说法时,L0 无法命中。
|
||||
- 文档 frontmatter 质量变成检索质量的隐性前提。
|
||||
|
||||
---
|
||||
|
||||
## 建议修复
|
||||
|
||||
1. 为关键词增加最小长度、停用词、领域前缀等基础规则。
|
||||
2. 区分 `exactKeywords`、`aliases`、`domainTags`,避免所有词混在一个匹配池。
|
||||
3. 引入轻量中文分词或归一化策略,先不必上复杂搜索引擎。
|
||||
4. 上传文档时校验 frontmatter 质量,缺失关键词时给出警告。
|
||||
5. 在 issue 修复前,至少补一份知识库文档 frontmatter 编写规范。
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/DocumentManagementService.java`
|
||||
- `src/main/java/com/superbiz/agent/controller/DocumentController.java`
|
||||
@@ -0,0 +1,55 @@
|
||||
# RAG L0 和 L1 未真正融合排序
|
||||
|
||||
**状态**:待规划
|
||||
**严重程度**:中
|
||||
**发现时间**:2026-07-04
|
||||
**范围**:知识检索工具、召回排序、Agent 证据质量
|
||||
|
||||
---
|
||||
|
||||
## 现象
|
||||
|
||||
当前 `lookup_knowledge` 的 L0 和 L1 更像是串行兜底关系,不是真正的多路召回融合:
|
||||
|
||||
- L0 命中唯一结果时,直接返回 L0。
|
||||
- L0 不唯一或不足时,才进入 L1。
|
||||
- L1 查询 `topK=3`,但最终主要把第一条结果作为补充证据。
|
||||
|
||||
这会导致关键词召回和语义召回没有充分互补。
|
||||
|
||||
---
|
||||
|
||||
## 当前实现
|
||||
|
||||
- `LookupKnowledgeTool` 先调用 `KnowledgeIndexService.exactMatch` 做 L0。
|
||||
- 再按条件调用 `VectorSearchService.search` 做 L1。
|
||||
- L0 和 L1 结果没有统一进入候选池做 fusion ranking。
|
||||
- L1 多结果没有充分利用,相关性接近的候选可能被丢弃。
|
||||
|
||||
---
|
||||
|
||||
## 影响
|
||||
|
||||
- L0 命中但质量一般时,会压过更好的 L1 语义结果。
|
||||
- L1 找到多个相近片段时,只有 top1 被 Agent 看到,降低召回覆盖率。
|
||||
- 难以解释检索排序,因为当前更像规则分支,不是可调的排序模型。
|
||||
|
||||
---
|
||||
|
||||
## 建议修复
|
||||
|
||||
建立统一候选池:
|
||||
|
||||
1. L0 和 L1 都返回候选列表。
|
||||
2. 按 `docId/chunkId` 去重。
|
||||
3. 为候选计算综合分:`keywordScore`、`vectorScore`、`domainScore`、`freshness`、`breadcrumbMatch`。
|
||||
4. 取 topN 进入上下文打包,而不是只取 L1 top1。
|
||||
5. 在 `tool_invocation` 中记录每个候选的分数组成,方便调试。
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
- `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
|
||||
@@ -0,0 +1,49 @@
|
||||
# RAG L1 分数阈值未校准
|
||||
|
||||
**状态**:待规划
|
||||
**严重程度**:中
|
||||
**发现时间**:2026-07-04
|
||||
**范围**:向量搜索、相关性判断、工具调用记录
|
||||
|
||||
---
|
||||
|
||||
## 现象
|
||||
|
||||
L1 语义检索使用 Milvus 向量距离后,会做相关性归一和阈值判断。但当前阈值更偏经验值,没有基于真实查询集和真实分数分布做校准。
|
||||
|
||||
由于当前使用 L2 距离,不同 embedding 模型、不同语料密度、不同 query 长度都会影响分数分布。
|
||||
|
||||
---
|
||||
|
||||
## 当前实现
|
||||
|
||||
- `VectorSearchService` 使用 query embedding 搜索 Milvus。
|
||||
- Milvus metric type 为 `L2`。
|
||||
- `LookupKnowledgeTool` 会把 L2 score 转成 normalized relevance。
|
||||
- 阈值没有配套评测集或分布统计。
|
||||
|
||||
---
|
||||
|
||||
## 影响
|
||||
|
||||
- 阈值过松时,低相关片段会进入 Agent 上下文。
|
||||
- 阈值过紧时,正确片段可能被过滤掉。
|
||||
- 面试中如果被追问“为什么这个阈值合理”,当前只能回答是 MVP 经验值。
|
||||
|
||||
---
|
||||
|
||||
## 建议修复
|
||||
|
||||
1. 固化一组 RAG 回归查询集,覆盖告警、数据库、流程规范、AIOps 诊断等场景。
|
||||
2. 记录每次 topK 的原始 L2 score、归一化分数、最终是否采纳。
|
||||
3. 统计正例和负例分布,确定阈值区间。
|
||||
4. 将阈值配置化,并在 README 或 issue 中记录选择依据。
|
||||
5. 后续引入 reranker 后,L1 阈值可以从“最终判断”退化为“粗召回过滤”。
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
- `src/main/resources/application.yml`
|
||||
@@ -0,0 +1,48 @@
|
||||
# RAG 查询改写能力薄弱
|
||||
|
||||
**状态**:待规划
|
||||
**严重程度**:中
|
||||
**发现时间**:2026-07-04
|
||||
**范围**:检索工具、Agent 查询生成、召回稳定性
|
||||
|
||||
---
|
||||
|
||||
## 现象
|
||||
|
||||
当前检索主要使用 Agent 传入 `lookup_knowledge` 的原始 query。工具层没有显式的 query rewrite、同义词扩展、领域词补全或多 query 检索。
|
||||
|
||||
当用户问题口语化、上下文依赖强,或缺少领域关键词时,L0 和 L1 的召回都可能不稳定。
|
||||
|
||||
---
|
||||
|
||||
## 当前实现
|
||||
|
||||
- Agent 决定何时调用 `lookup_knowledge` 和传入什么 query。
|
||||
- `LookupKnowledgeTool` 接收 query 后直接进入 L0/L1 检索。
|
||||
- 工具层没有把用户问题改写成多个检索 query。
|
||||
- 也没有把当前任务域、Planner step、告警 payload 等上下文显式拼入检索 query。
|
||||
|
||||
---
|
||||
|
||||
## 影响
|
||||
|
||||
- Agent query 写得好时召回正常,query 写得差时检索链路缺少兜底。
|
||||
- AIOps 场景里,告警名称、服务名、指标名、故障类型之间的别名关系没有被充分利用。
|
||||
- 很难稳定复现同一类问题的检索质量。
|
||||
|
||||
---
|
||||
|
||||
## 建议修复
|
||||
|
||||
1. 在工具层增加轻量 query rewrite:原始问题、领域词增强问题、关键词查询并行召回。
|
||||
2. 对 AIOps 场景,把 alertName、service、metric、symptom 显式构造成检索 query。
|
||||
3. 记录 rewrite 前后的 query 到 `tool_invocation`,便于分析。
|
||||
4. 后续可以引入 LLM query rewrite,但 MVP 先用规则模板更可控。
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
- `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
|
||||
@@ -0,0 +1,376 @@
|
||||
# RAG 检索重构计划
|
||||
|
||||
**状态**:待规划
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-05
|
||||
**范围**:RAG、检索、知识库、Agent Tool、AIOps 诊断证据链
|
||||
|
||||
---
|
||||
|
||||
## 目标
|
||||
|
||||
将当前自研 RAG MVP 重构为“成熟框架能力 + 业务可观测编排”的架构:
|
||||
|
||||
```text
|
||||
Agent
|
||||
-> lookup_knowledge Tool
|
||||
-> L0 domain/entity hint
|
||||
-> query augmentation / transformer
|
||||
-> Spring AI Retriever / VectorStore
|
||||
-> metadata filter
|
||||
-> document postprocess
|
||||
-> neighbor / section expansion
|
||||
-> evidence packing
|
||||
-> tool_invocation record
|
||||
```
|
||||
|
||||
核心原则:
|
||||
|
||||
1. 通用 RAG 基础设施尽量交给 Spring AI / Spring AI Alibaba。
|
||||
2. Agent 工具入口、AIOps 业务语义、证据追踪继续保留在项目内。
|
||||
3. 不把系统改成隐式 Chat RAG,仍然保留显式 `lookup_knowledge` 工具调用。
|
||||
4. 分阶段迁移,避免一次性推倒当前可运行链路。
|
||||
|
||||
---
|
||||
|
||||
## 当前问题汇总
|
||||
|
||||
当前 RAG 已经打通上传、切片、向量化、L0/L1 召回和工具调用记录,但主要问题集中在:
|
||||
|
||||
1. **检索基础设施偏自研**
|
||||
- Milvus 写入和查询直接使用 SDK。
|
||||
- topK、threshold、metadata filter、结果结构由业务代码维护。
|
||||
- 后续接入 Spring AI RAG 能力会有重复适配成本。
|
||||
|
||||
2. **L0 职责过重**
|
||||
- 当前 L0 可能被当成最终召回决策。
|
||||
- 关键词质量不稳定时容易误召回。
|
||||
- 更适合作为 domain/entity hint,而不是最终答案来源。
|
||||
|
||||
3. **query 构造不稳定**
|
||||
- 主要依赖 Agent 传入原始 query。
|
||||
- AIOps payload 中的 alertName、service、metric、symptom 没有稳定进入检索 query。
|
||||
|
||||
4. **上下文重建不足**
|
||||
- 同章节被切成多个 chunk 后,命中片段不会自动扩展前后文。
|
||||
- `breadcrumb` 存在 metadata 中,但没有充分参与 embedding、filter 或 context packing。
|
||||
|
||||
5. **缺少检索后处理**
|
||||
- 缺少统一 evidence block。
|
||||
- 缺少去重、token budget、hitReason、source 结构化输出。
|
||||
|
||||
6. **缺少量化评测**
|
||||
- 目前主要靠接口回放、日志和 `tool_invocation` 人工判断。
|
||||
- 还没有 golden query set、Recall@K、MRR、NDCG 等检索评测。
|
||||
|
||||
---
|
||||
|
||||
## 保留设计
|
||||
|
||||
这些设计值得保留,并作为重构后的项目亮点:
|
||||
|
||||
### 1. `lookup_knowledge` 显式 Agent Tool
|
||||
|
||||
保留显式工具调用,不直接用隐式 Advisor 取代。
|
||||
|
||||
原因:
|
||||
|
||||
- 面试项目重点是 Agent 工程,不是普通 Chat RAG。
|
||||
- 显式工具调用能展示 Agent 何时检索、检索了什么、证据如何支撑诊断。
|
||||
- `tool_invocation`、evidence score、diagnosis session 都依赖这条链路。
|
||||
|
||||
### 2. L0
|
||||
|
||||
保留 L0,但降级为:
|
||||
|
||||
- domain detector
|
||||
- entity extractor
|
||||
- metadata filter generator
|
||||
- explainability signal
|
||||
|
||||
不再默认执行:
|
||||
|
||||
```text
|
||||
L0 unique hit -> 直接返回
|
||||
```
|
||||
|
||||
目标职责:
|
||||
|
||||
```text
|
||||
query / payload
|
||||
-> L0 matched keywords/entities/domain
|
||||
-> metadata filter + query augmentation
|
||||
-> retriever
|
||||
```
|
||||
|
||||
### 3. metadata
|
||||
|
||||
保留并加强 metadata:
|
||||
|
||||
```text
|
||||
docId
|
||||
chunkIndex
|
||||
totalChunks
|
||||
title
|
||||
breadcrumb
|
||||
category
|
||||
source
|
||||
```
|
||||
|
||||
后续可扩展:
|
||||
|
||||
```text
|
||||
sectionId
|
||||
parentSection
|
||||
documentType
|
||||
domain
|
||||
tags
|
||||
version
|
||||
```
|
||||
|
||||
metadata 是 filter、上下文扩展、证据追踪、可解释性的基础。
|
||||
|
||||
### 4. Markdown-aware chunking
|
||||
|
||||
保留当前 Markdown 结构化切片思路:
|
||||
|
||||
- 识别标题层级
|
||||
- 生成 `title`
|
||||
- 生成 `breadcrumb`
|
||||
- 保留 `chunkIndex`
|
||||
- 尽量不打断列表和代码块
|
||||
|
||||
可以替换或复用框架能力的是底层 token 长度控制和 overlap 策略,而不是完全抛弃结构化切片。
|
||||
|
||||
### 5. `tool_invocation` 证据追踪
|
||||
|
||||
保留并增强:
|
||||
|
||||
```text
|
||||
sessionId
|
||||
query
|
||||
rewrittenQuery
|
||||
matchedKeywords
|
||||
domain/entities
|
||||
retrievedDocs
|
||||
scores
|
||||
hitReasons
|
||||
evidence
|
||||
duration
|
||||
relevanceLevel
|
||||
```
|
||||
|
||||
这是后续检索评测、诊断质量评估、面试讲解的基础。
|
||||
|
||||
### 6. AIOps payload 到 query 的业务映射
|
||||
|
||||
保留 AIOps 场景逻辑:
|
||||
|
||||
- alertName
|
||||
- service
|
||||
- metric
|
||||
- symptom
|
||||
- category/domain
|
||||
|
||||
这些是业务语义,不能完全交给通用框架隐式处理。
|
||||
|
||||
---
|
||||
|
||||
## 替换设计
|
||||
|
||||
这些能力适合逐步交给 Spring AI / Spring AI Alibaba:
|
||||
|
||||
| 当前能力 | 目标能力 | 说明 |
|
||||
|---|---|---|
|
||||
| Milvus SDK 直接写入/查询 | Spring AI `VectorStore` | 减少基础设施代码 |
|
||||
| 自研 `VectorSearchService` 检索细节 | `VectorStoreDocumentRetriever` | 标准化 topK、threshold、filter |
|
||||
| 手写 query 拼接 | Query Transformer / 模板化 query augmentation | 先规则化,后框架化 |
|
||||
| 手写结果拼接 | DocumentPostProcessor / evidence postprocess | 做去重、压缩、证据块 |
|
||||
| L0 最终召回判断 | L0 domain/entity hint | 降低误召回风险 |
|
||||
|
||||
---
|
||||
|
||||
## 分阶段计划
|
||||
|
||||
### Phase 0:重构前基线
|
||||
|
||||
目标:先固定当前行为,避免重构后不知道是否变好。
|
||||
|
||||
任务:
|
||||
|
||||
- 固化 10-20 条 golden queries。
|
||||
- 覆盖 Chat 和 AIOps 场景。
|
||||
- 每条 query 标注 expected doc、breadcrumb、关键 chunk 或 evidence。
|
||||
- 用当前链路跑一遍,记录 baseline。
|
||||
|
||||
验收:
|
||||
|
||||
- 有可重复运行的检索回放清单。
|
||||
- 能记录当前 Recall@K、first hit rank 或人工 hit level。
|
||||
|
||||
### Phase 1:L0 降级为 domain/entity hint
|
||||
|
||||
目标:保留 L0 价值,降低 L0 误决策风险。
|
||||
|
||||
任务:
|
||||
|
||||
- `KnowledgeIndexService` 输出 matched keywords、domain、entities。
|
||||
- `LookupKnowledgeTool` 不再把 L0 unique hit 作为默认最终结果。
|
||||
- 将 L0 结果用于 query augmentation 和 metadata filter。
|
||||
- `tool_invocation` 记录 L0 hit reason。
|
||||
|
||||
验收:
|
||||
|
||||
- L0 命中不会绕过向量检索直接返回。
|
||||
- 检索记录能看到 domain/entities/matchedKeywords。
|
||||
- AIOps payload 能生成稳定领域 hint。
|
||||
|
||||
### Phase 2:Evidence Postprocess 和上下文打包
|
||||
|
||||
目标:先提升 Agent 实际拿到的证据质量。
|
||||
|
||||
任务:
|
||||
|
||||
- 定义 evidence block:
|
||||
|
||||
```text
|
||||
source
|
||||
docId
|
||||
chunkIndex
|
||||
title
|
||||
breadcrumb
|
||||
score
|
||||
hitReason
|
||||
content
|
||||
expandedFrom
|
||||
```
|
||||
|
||||
- 对检索结果做去重。
|
||||
- 支持命中 chunk 的相邻 chunk / 同章节扩展。
|
||||
- 加 token 或字符预算控制。
|
||||
- 返回给 Agent 的内容按 evidence block 组织。
|
||||
|
||||
验收:
|
||||
|
||||
- 同一 docId/chunkIndex 不重复进入上下文。
|
||||
- 命中 chunk 可以补充前后文。
|
||||
- `tool_invocation` 记录 postprocess 前后候选数量和最终 evidence 数量。
|
||||
|
||||
### Phase 3:Spring AI VectorStore 旁路验证
|
||||
|
||||
目标:验证框架能力,不直接替换主链路。
|
||||
|
||||
任务:
|
||||
|
||||
- 引入 Spring AI Milvus VectorStore。
|
||||
- 建立旁路 `SpringAiVectorSearchService` 或适配层。
|
||||
- 同一批 golden queries 同时跑旧链路和新链路。
|
||||
- 对比 topK、metadata、score、filter 行为。
|
||||
|
||||
验收:
|
||||
|
||||
- 旁路检索可跑通。
|
||||
- metadata 不丢失。
|
||||
- 查询结果与当前链路差异可解释。
|
||||
- 不影响现有 Chat / AIOps 主链路。
|
||||
|
||||
### Phase 4:替换底层 VectorSearchService
|
||||
|
||||
目标:对外接口不变,内部检索切到 Spring AI VectorStore / Retriever。
|
||||
|
||||
任务:
|
||||
|
||||
- 保持 `LookupKnowledgeTool` 调用方式不变。
|
||||
- `VectorSearchService` 内部迁移到 Spring AI 检索抽象。
|
||||
- 支持 topK、similarity threshold、category metadata filter。
|
||||
- 保留旧实现一段时间作为 fallback。
|
||||
|
||||
验收:
|
||||
|
||||
- Chat / AIOps 检索链路行为兼容。
|
||||
- golden queries 不低于 baseline。
|
||||
- 检索结果仍能完整记录到 `tool_invocation`。
|
||||
|
||||
### Phase 5:Query Transformer 和框架化 PostProcessor
|
||||
|
||||
目标:在稳定的 VectorStore 基础上接入更成熟 RAG 能力。
|
||||
|
||||
任务:
|
||||
|
||||
- AIOps 场景优先使用模板化 query augmentation。
|
||||
- 需要时接入 Spring AI Query Transformer / MultiQuery。
|
||||
- 将现有 evidence postprocess 抽象成 DocumentPostProcessor 风格。
|
||||
- 可选接入 rerank,但不作为第一优先级。
|
||||
|
||||
验收:
|
||||
|
||||
- 原始 query 和 rewritten query 都可追踪。
|
||||
- query rewrite 失败可以 fallback。
|
||||
- postprocess 行为可配置、可记录、可回放。
|
||||
|
||||
---
|
||||
|
||||
## 暂不做
|
||||
|
||||
以下能力暂不进入近期重构:
|
||||
|
||||
1. 不做完整自研 RRF 框架。
|
||||
2. 不直接把 `lookup_knowledge` 替换成隐式 Advisor。
|
||||
3. 不一口气迁移所有 RAG ETL。
|
||||
4. 不先引入 Elasticsearch / OpenSearch,除非评测证明 BM25 必须。
|
||||
5. 不先上 cross-encoder / LLM rerank,先做规则型 evidence postprocess。
|
||||
|
||||
---
|
||||
|
||||
## 风险
|
||||
|
||||
### 1. Milvus schema 兼容风险
|
||||
|
||||
当前 collection 是项目自建,Spring AI VectorStore 可能有自己的 schema 假设。需要旁路验证。
|
||||
|
||||
### 2. 检索行为变化风险
|
||||
|
||||
框架检索分数和当前 L2 score 可能不完全一致,需要 golden queries 对比。
|
||||
|
||||
### 3. 可观测性丢失风险
|
||||
|
||||
如果迁移到隐式 Advisor,可能丢失工具调用证据链。因此 Spring AI RAG 能力应优先封装在 `lookup_knowledge` 内部。
|
||||
|
||||
### 4. 重构范围膨胀风险
|
||||
|
||||
RAG、Agent、AIOps、数据库记录互相关联,必须分阶段推进,每阶段都保持可运行。
|
||||
|
||||
---
|
||||
|
||||
## 合并来源
|
||||
|
||||
本计划合并以下问题和改造方向:
|
||||
|
||||
- [rag-chunk-context-reconstruction.md](rag-chunk-context-reconstruction.md)
|
||||
- [rag-breadcrumb-embedding-gap.md](rag-breadcrumb-embedding-gap.md)
|
||||
- [rag-l0-l1-fusion-ranking.md](rag-l0-l1-fusion-ranking.md)
|
||||
- [rag-l0-keyword-matching-quality.md](rag-l0-keyword-matching-quality.md)
|
||||
- [rag-l1-score-calibration.md](rag-l1-score-calibration.md)
|
||||
- [rag-context-packing-and-reranking.md](rag-context-packing-and-reranking.md)
|
||||
- [rag-upload-chunk-parameter-drift.md](rag-upload-chunk-parameter-drift.md)
|
||||
- [rag-query-rewrite-gap.md](rag-query-rewrite-gap.md)
|
||||
- [rag-spring-ai-vectorstore-migration.md](rag-spring-ai-vectorstore-migration.md)
|
||||
- [rag-spring-ai-query-transformer.md](rag-spring-ai-query-transformer.md)
|
||||
- [rag-spring-ai-document-postprocessor.md](rag-spring-ai-document-postprocessor.md)
|
||||
- [rag-l0-domain-entity-hint.md](rag-l0-domain-entity-hint.md)
|
||||
- [rag-spring-ai-advisor-boundary.md](rag-spring-ai-advisor-boundary.md)
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
- `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/VectorIndexService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/DocumentChunkService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/DocumentManagementService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
|
||||
- `src/main/resources/application.yml`
|
||||
- `pom.xml`
|
||||
@@ -0,0 +1,66 @@
|
||||
# RAG 明确 Spring AI Advisor 与 Agent Tool 的边界
|
||||
|
||||
**状态**:待规划
|
||||
**严重程度**:中
|
||||
**发现时间**:2026-07-05
|
||||
**范围**:Agent 编排、RAG Advisor、工具调用可观测性
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
Spring AI 提供 `QuestionAnswerAdvisor`、`RetrievalAugmentationAdvisor` 等 RAG Advisor 能力,可以把检索增强直接挂到模型调用流程中。
|
||||
|
||||
但当前项目是 Agent 工程项目,知识检索不是普通聊天增强,而是 Agent 在诊断流程中显式调用的工具。系统还依赖 `tool_invocation` 记录检索事实,用于 evidence score 和诊断追踪。
|
||||
|
||||
---
|
||||
|
||||
## 问题
|
||||
|
||||
如果直接把 RAG 全部迁到 Advisor,可能会损失当前项目已有的显式工具链路:
|
||||
|
||||
1. Agent 是否调用知识库不够透明。
|
||||
2. `tool_invocation` 记录可能变弱。
|
||||
3. AIOps 诊断步骤和知识证据之间的对应关系不清晰。
|
||||
4. 面试项目中“Agent 如何使用工具”的展示价值下降。
|
||||
|
||||
---
|
||||
|
||||
## 改造方向
|
||||
|
||||
不要把 `lookup_knowledge` 完全替换成隐式 Advisor,而是分层使用:
|
||||
|
||||
```text
|
||||
Agent Tool 层:
|
||||
lookup_knowledge
|
||||
sessionId
|
||||
traceId
|
||||
tool_invocation
|
||||
evidence score
|
||||
|
||||
Spring AI RAG 层:
|
||||
query transformer
|
||||
retriever
|
||||
vector store
|
||||
document post processor
|
||||
```
|
||||
|
||||
也就是说,Advisor / Retriever 可以作为工具内部实现,而不是取代工具本身。
|
||||
|
||||
---
|
||||
|
||||
## 验收标准
|
||||
|
||||
- Agent 仍然通过显式 `lookup_knowledge` 使用知识库。
|
||||
- Spring AI RAG 能力被封装在工具内部或服务内部。
|
||||
- 每次检索仍能落 `tool_invocation`。
|
||||
- Chat / AIOps 两条链路都能追踪检索输入、输出和证据来源。
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
|
||||
@@ -0,0 +1,60 @@
|
||||
# RAG 使用 DocumentPostProcessor 做后处理
|
||||
|
||||
**状态**:待规划
|
||||
**严重程度**:中
|
||||
**发现时间**:2026-07-05
|
||||
**范围**:检索后处理、去重、上下文打包、轻量 rerank
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
当前检索结果返回给 Agent 前,主要依赖 `LookupKnowledgeTool` 自己拼接内容。系统还缺少统一的后处理阶段。
|
||||
|
||||
Spring AI RAG 流程中可以使用 DocumentPostProcessor 类能力,在文档进入模型上下文前做过滤、去重、压缩或 rerank。
|
||||
|
||||
---
|
||||
|
||||
## 问题
|
||||
|
||||
当前检索后处理不足:
|
||||
|
||||
1. L1 topK 候选没有被充分利用。
|
||||
2. 同文档或同章节结果可能重复。
|
||||
3. 命中 chunk 后没有统一处理前后文扩展。
|
||||
4. 证据块缺少统一格式,后续 verifier / evaluator 不容易复用。
|
||||
|
||||
---
|
||||
|
||||
## 改造方向
|
||||
|
||||
建立一个轻量后处理链:
|
||||
|
||||
```text
|
||||
retrieved documents
|
||||
-> deduplicate
|
||||
-> optional neighbor / section expansion
|
||||
-> score / reason annotation
|
||||
-> token budget packing
|
||||
-> evidence blocks
|
||||
```
|
||||
|
||||
优先做规则型后处理,不急于引入 cross-encoder 或 LLM rerank。
|
||||
|
||||
---
|
||||
|
||||
## 验收标准
|
||||
|
||||
- 同一个 docId/chunkIndex 不重复进入 Agent 上下文。
|
||||
- 最终返回内容包含 source、title、breadcrumb、score、hitReason。
|
||||
- 可以限制单次工具调用返回的最大 token 或最大字符数。
|
||||
- 后处理前后的候选数量、去重数量、最终 evidence 数量记录到 `tool_invocation`。
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/DocumentChunkService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
|
||||
@@ -0,0 +1,70 @@
|
||||
# RAG 接入 Spring AI Query Transformer
|
||||
|
||||
**状态**:待规划
|
||||
**严重程度**:中
|
||||
**发现时间**:2026-07-05
|
||||
**范围**:查询改写、多查询扩展、AIOps 检索稳定性
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
当前 `lookup_knowledge` 主要使用 Agent 传入的原始 query 进行 L0/L1 检索。query 的质量高度依赖 Agent 当次生成结果。
|
||||
|
||||
Spring AI 提供 Query Transformer / Query Expander 类能力,可以把用户问题或 Agent 子任务改写成更适合检索的查询。
|
||||
|
||||
---
|
||||
|
||||
## 问题
|
||||
|
||||
当前检索 query 存在几个风险:
|
||||
|
||||
1. 用户问题口语化时,缺少领域关键词。
|
||||
2. AIOps payload 中的 alertName、service、metric 没有稳定拼入检索 query。
|
||||
3. 同义表达没有扩展,例如“连接耗尽”和“连接池打满”。
|
||||
4. 工具层无法复用框架提供的 rewrite / expansion 能力。
|
||||
|
||||
---
|
||||
|
||||
## 改造方向
|
||||
|
||||
在 `lookup_knowledge` 前增加查询改写层:
|
||||
|
||||
```text
|
||||
raw query / alert payload
|
||||
-> query transformer
|
||||
-> rewritten query / expanded queries
|
||||
-> retriever
|
||||
```
|
||||
|
||||
优先支持两类场景:
|
||||
|
||||
1. **AIOps 模板化改写**
|
||||
- alertName
|
||||
- service
|
||||
- metric
|
||||
- symptom
|
||||
- domain/category
|
||||
|
||||
2. **Spring AI Query Transformer**
|
||||
- rewrite 原始 query
|
||||
- multi-query expansion
|
||||
- 必要时做 query compression
|
||||
|
||||
---
|
||||
|
||||
## 验收标准
|
||||
|
||||
- `tool_invocation` 中记录原始 query 和改写后的 query。
|
||||
- AIOps payload 存在时,检索 query 能稳定带上告警和服务上下文。
|
||||
- 对同一个测试问题,改写前后 topK 命中结果可对比。
|
||||
- 未配置 transformer 时,可以回退到原始 query。
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
|
||||
- `src/main/java/com/superbiz/agent/domain/entity/ToolInvocation.java`
|
||||
@@ -0,0 +1,80 @@
|
||||
# RAG 迁移到 Spring AI VectorStore 检索抽象
|
||||
|
||||
**状态**:待规划
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-05
|
||||
**范围**:向量检索、Milvus 接入、RAG 框架化改造
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
当前系统的向量检索链路主要由项目手写实现:
|
||||
|
||||
- `VectorIndexService` 负责向量化和写入 Milvus。
|
||||
- `VectorSearchService` 直接使用 Milvus SDK 查询。
|
||||
- `LookupKnowledgeTool` 自己组织 L0/L1 检索结果。
|
||||
|
||||
这能满足 MVP 打通链路,但继续扩展 RAG 能力时,容易把项目变成自研搜索框架。
|
||||
|
||||
项目当前已引入 Spring AI / Spring AI Alibaba 依赖,可以考虑迁移到 Spring AI 的 `VectorStore`、`VectorStoreDocumentRetriever` 等标准抽象。
|
||||
|
||||
---
|
||||
|
||||
## 问题
|
||||
|
||||
当前手写 Milvus 检索存在几个成本:
|
||||
|
||||
1. topK、similarity threshold、metadata filter 等逻辑分散在业务代码中。
|
||||
2. 检索结果结构和 Spring AI RAG Advisor 生态不兼容。
|
||||
3. 后续接入 query transformer、post processor、advisor 时需要重复适配。
|
||||
4. Milvus SDK 直接调用让业务层承担了过多基础设施细节。
|
||||
|
||||
---
|
||||
|
||||
## 改造方向
|
||||
|
||||
优先引入 Spring AI 的 Milvus VectorStore 能力:
|
||||
|
||||
```text
|
||||
当前:
|
||||
VectorSearchService -> Milvus SDK
|
||||
|
||||
目标:
|
||||
LookupKnowledgeTool / RAG Service
|
||||
-> VectorStoreDocumentRetriever
|
||||
-> Spring AI VectorStore
|
||||
-> Milvus
|
||||
```
|
||||
|
||||
业务层保留:
|
||||
|
||||
- `lookup_knowledge` 工具入口
|
||||
- `tool_invocation` 记录
|
||||
- sessionId / category / domain 等业务上下文
|
||||
|
||||
底层检索交给框架:
|
||||
|
||||
- topK
|
||||
- similarity threshold
|
||||
- metadata filter
|
||||
- vector search options
|
||||
|
||||
---
|
||||
|
||||
## 验收标准
|
||||
|
||||
- `VectorSearchService` 不再直接散落 Milvus 查询细节,至少封装到 Spring AI `VectorStore` 适配层。
|
||||
- 支持按 `category` 或其他 metadata filter 检索。
|
||||
- 检索结果仍能记录到 `tool_invocation`。
|
||||
- 现有 AIOps / Chat 检索链路行为保持兼容。
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `pom.xml`
|
||||
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/VectorIndexService.java`
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
- `src/main/resources/application.yml`
|
||||
@@ -0,0 +1,49 @@
|
||||
# RAG 上传切片参数未真正生效
|
||||
|
||||
**状态**:待规划
|
||||
**严重程度**:低
|
||||
**发现时间**:2026-07-04
|
||||
**范围**:文档上传接口、切片配置、API 行为一致性
|
||||
|
||||
---
|
||||
|
||||
## 现象
|
||||
|
||||
上传接口暴露了 `chunkSize` 和 `chunkOverlap` 参数,但实际切片主要使用全局 `DocumentChunkConfig`。这会造成 API 表面能力和真实行为不一致。
|
||||
|
||||
---
|
||||
|
||||
## 当前实现
|
||||
|
||||
- `DocumentController` 的上传接口接收 `chunkSize` 和 `chunkOverlap`。
|
||||
- 这些字段会进入 `DocumentUploadRequest`。
|
||||
- `DocumentChunkService` 的切片阈值主要来自 `DocumentChunkConfig`。
|
||||
- 单次上传请求中的参数没有真正覆盖切片配置。
|
||||
|
||||
---
|
||||
|
||||
## 影响
|
||||
|
||||
- 调用方以为可以控制切片大小,但实际无法影响结果。
|
||||
- 测试时容易误判“参数调优无效”的原因。
|
||||
- 面试中如果展示 API,会被追问参数是否真实生效。
|
||||
|
||||
---
|
||||
|
||||
## 建议修复
|
||||
|
||||
两个方向二选一:
|
||||
|
||||
1. 如果 MVP 不需要请求级切片参数,就从接口中移除或标记为暂不支持。
|
||||
2. 如果需要支持,就让 `DocumentChunkService` 接收 per-request chunk options,并记录到文档 metadata 中。
|
||||
|
||||
建议面试项目中优先选择第二种,因为它更能体现工程闭环:API、配置、落库、追踪一致。
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `src/main/java/com/superbiz/agent/controller/DocumentController.java`
|
||||
- `src/main/java/com/superbiz/agent/service/DocumentManagementService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/DocumentChunkService.java`
|
||||
- `src/main/java/com/superbiz/agent/config/DocumentChunkConfig.java`
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,54 @@
|
||||
## Context
|
||||
|
||||
The AIOps endpoint has two natural modes:
|
||||
|
||||
- **Payload mode**: caller supplies `alertName`, `service`, or other alert fields. The caller is asking for targeted diagnosis of that alert.
|
||||
- **Auto-discovery mode**: caller omits alert fields. The system should discover active alerts first, then analyze them.
|
||||
|
||||
The current task prompt does not distinguish these modes, so the agent may query all active alerts and produce a broad report even when a specific alert payload was supplied.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Make AIOps payload mode single-alert focused.
|
||||
- Keep no-payload mode compatible with the original "query active alerts then diagnose" behavior.
|
||||
- Keep the change prompt-only and low risk.
|
||||
- Add tests for prompt scope rules.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not add a Verifier Agent.
|
||||
- Do not force tool calls in Java code.
|
||||
- Do not change `/api/ai_ops` request/response contracts.
|
||||
- Do not modify mock alert data.
|
||||
|
||||
## Decisions
|
||||
|
||||
| Decision | Choice | Alternative Considered | Rationale |
|
||||
|---|---|---|---|
|
||||
| Scope detection | Treat non-empty alert fields as payload mode | Add explicit `mode` field | Existing payload already carries enough intent; no API change needed. |
|
||||
| Payload mode behavior | Final report focuses only on supplied alert | Filter tool results in Java | Prompt-level rule is the smallest change and preserves agent flexibility. |
|
||||
| Auto mode behavior | Require active-alert discovery first | Always analyze only one alert | Original AIOps value is automated alert discovery when no payload exists. |
|
||||
| Other active alerts in payload mode | Mention only as related risk | Ignore entirely | Some context can be useful, but not enough to expand the report. |
|
||||
|
||||
## Prompt Rules
|
||||
|
||||
Payload mode MUST instruct the agent:
|
||||
|
||||
- Treat supplied payload as the primary and only report target.
|
||||
- Use `queryPrometheusAlerts` only to verify the supplied alert state or identify related risk.
|
||||
- Do not create root-cause sections for unrelated active alerts.
|
||||
- Report unrelated alerts only in a brief "关联风险" note if they appear relevant.
|
||||
|
||||
Auto-discovery mode MUST instruct the agent:
|
||||
|
||||
- First call `queryPrometheusAlerts`.
|
||||
- Select P0/P1 or longest-running firing alerts.
|
||||
- Analyze one or more active alerts based on severity and evidence.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Prompt-only control may not be perfectly followed by the LLM. -> Mitigation: tests lock prompt wording; runtime can be reviewed through trace.
|
||||
- [Risk] Payload mode may miss broader incidents. -> Mitigation: related active alerts may be mentioned as risk, but not expanded into full sections.
|
||||
- [Risk] Future stronger enforcement may be needed. -> Mitigation: a later change can filter tool summaries or add AIOps Verifier.
|
||||
@@ -0,0 +1,27 @@
|
||||
## Why
|
||||
|
||||
Runtime verification showed that AIOps now correctly accepts an alert payload and persists a trace, but the generated report still expands to every active mock Prometheus alert. That weakens the product boundary between `/api/chat` and `/api/ai_ops`: an alert payload should mean targeted alert diagnosis, while an empty payload should mean automatic active-alert discovery.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Tighten the AIOps task prompt so payload mode focuses the final report on the supplied alert.
|
||||
- Preserve full active-alert discovery when no payload is supplied.
|
||||
- Allow Prometheus active-alert lookup in payload mode only as supporting evidence, not as permission to expand the report to unrelated alerts.
|
||||
- Update tests and demo acceptance wording to lock the new behavior.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `aiops-alert-scope-control`: Defines AIOps diagnosis scope rules for payload mode versus auto-discovery mode.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `aiops-traceable-diagnosis-entry`: Keeps the same API and trace behavior but clarifies how AIOps should scope its diagnosis.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected code: `AiOpsService.buildTaskPrompt(...)`, focused tests, demo documentation, devflow records.
|
||||
- Affected API: no endpoint or request/response shape change.
|
||||
- Affected persistence: no schema change.
|
||||
- Non-goals: no Verifier integration, no tool implementation change, no prompt rewrite for Chat.
|
||||
+25
@@ -0,0 +1,25 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: AIOps payload mode focuses on supplied alert
|
||||
When an AIOps request includes alert payload fields, the system SHALL instruct the agent to focus the final alert analysis report on the supplied alert.
|
||||
|
||||
#### Scenario: Request includes alertName and service
|
||||
- **WHEN** a caller posts to `/api/ai_ops` with `alertName` and `service`
|
||||
- **THEN** the AIOps task prompt identifies payload mode
|
||||
- **AND** the prompt instructs the agent not to create full root-cause sections for unrelated active alerts
|
||||
|
||||
### Requirement: AIOps auto-discovery mode queries active alerts first
|
||||
When an AIOps request omits alert payload fields, the system SHALL instruct the agent to first discover active Prometheus alerts.
|
||||
|
||||
#### Scenario: Request body is omitted
|
||||
- **WHEN** a caller posts to `/api/ai_ops` without alert fields
|
||||
- **THEN** the AIOps task prompt identifies auto-discovery mode
|
||||
- **AND** the prompt instructs the agent to call `queryPrometheusAlerts` first
|
||||
|
||||
### Requirement: Payload mode may use active alerts as supporting context
|
||||
Payload mode SHALL allow active-alert lookup as supporting evidence, but SHALL keep unrelated alerts out of the main report sections.
|
||||
|
||||
#### Scenario: Prometheus returns multiple active alerts
|
||||
- **WHEN** payload mode is active and `queryPrometheusAlerts` returns unrelated active alerts
|
||||
- **THEN** the prompt permits mentioning those alerts only as related risk or context
|
||||
- **AND** the final report target remains the supplied alert
|
||||
@@ -0,0 +1,21 @@
|
||||
## 1. Flow Records
|
||||
|
||||
- [x] 1.1 Add devflow brief, decisions with lightweight Grill, evidence, and acceptance records.
|
||||
- [x] 1.2 Record local impact analysis and GitNexus skip context.
|
||||
|
||||
## 2. Prompt Scope Control
|
||||
|
||||
- [x] 2.1 Add payload detection helper in `AiOpsService`.
|
||||
- [x] 2.2 Update `buildTaskPrompt(...)` with payload-mode and auto-discovery-mode rules.
|
||||
|
||||
## 3. Tests And Docs
|
||||
|
||||
- [x] 3.1 Add tests for payload-mode prompt rules.
|
||||
- [x] 3.2 Add tests for no-payload auto-discovery prompt rules.
|
||||
- [x] 3.3 Update AIOps demo acceptance wording for single-alert payload mode.
|
||||
|
||||
## 4. Verification
|
||||
|
||||
- [x] 4.1 Run targeted tests.
|
||||
- [x] 4.2 Run compile verification.
|
||||
- [x] 4.3 Run OpenSpec validation.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,90 @@
|
||||
## Context
|
||||
|
||||
The AIOps endpoint is currently useful as a standalone alert-analysis demo, but it is not aligned with the MVP trace story:
|
||||
|
||||
- `ChatController.aiOps()` accepts no request body.
|
||||
- `AiOpsService.executeAiOpsAnalysis(...)` creates a random 8-character session id internally.
|
||||
- The caller cannot reliably discover that id and query `GET /api/diagnosis/{sessionId}/trace`.
|
||||
- `extractFinalReport(...)` returns the report text but does not persist it to `diagnosis_session.answer`.
|
||||
|
||||
The existing trace API already aggregates `diagnosis_session`, `agent_step`, and `tool_invocation`, so this change should reuse that storage rather than introduce new persistence.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Make `/api/ai_ops` usable as an alert-triggered diagnosis entry point.
|
||||
- Preserve backward compatibility for callers that post with no request body.
|
||||
- Return the resolved `sessionId` through SSE.
|
||||
- Persist the final report into the existing diagnosis session.
|
||||
- Keep the AIOps path observable through the existing trace API.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not merge AIOps into `ChatService`.
|
||||
- Do not add a new Verifier Agent to AIOps in this slice.
|
||||
- Do not change `GET /api/diagnosis/{sessionId}/trace`.
|
||||
- Do not add database migrations.
|
||||
- Do not clean up sensitive configuration.
|
||||
|
||||
## Decisions
|
||||
|
||||
| Decision | Choice | Alternative Considered | Rationale |
|
||||
|---|---|---|---|
|
||||
| API compatibility | Keep `POST /api/ai_ops` SSE and make body optional | Add a new `/api/ai_ops/v2` endpoint | Optional body keeps existing demo callers working while enabling traceable alert input. |
|
||||
| Session identity | Accept request `sessionId`, otherwise generate UUID | Continue internal-only 8-char id | Reviewers need the id to query trace and submit feedback. |
|
||||
| Query persistence | Build a concise alert diagnosis query from request fields | Store only "AI Ops 告警分析" | Trace should show what alert was diagnosed. |
|
||||
| Final answer persistence | Save extracted final report to `diagnosis_session.answer` | Only stream the report | Trace replay must include the final answer without relying on SSE logs. |
|
||||
| Verifier scope | Defer AIOps Verifier integration | Add Chat Verifier now | The minimum interview value is traceability; Verifier unification can be a follow-up after this entry point is stable. |
|
||||
| GitNexus | Skip by user decision | Block until MCP available | GitNexus tools are not exposed in this session, and the user explicitly requested skipping GitNexus. Local impact analysis and tests cover this slice. |
|
||||
|
||||
## Interface Impact
|
||||
|
||||
- Level: L3 API behavior extension.
|
||||
- Endpoint: `POST /api/ai_ops`
|
||||
- Compatibility: callers may still omit a body. New callers may send:
|
||||
|
||||
```json
|
||||
{
|
||||
"sessionId": "mvp-demo-aiops-payment-latency-001",
|
||||
"alertName": "payment-service-latency-high",
|
||||
"service": "payment-service",
|
||||
"severity": "P1",
|
||||
"description": "支付服务 P95 延迟升高并伴随超时错误",
|
||||
"timeRange": "last_15m"
|
||||
}
|
||||
```
|
||||
|
||||
The SSE stream emits a first content message containing the resolved session id:
|
||||
|
||||
```text
|
||||
sessionId: mvp-demo-aiops-payment-latency-001
|
||||
```
|
||||
|
||||
## Data Flow
|
||||
|
||||
```text
|
||||
POST /api/ai_ops
|
||||
-> ChatController resolves request body and tools
|
||||
-> AiOpsService.executeAiOpsAnalysis(chatModel, tools, request)
|
||||
-> create diagnosis_session(agentFlow=AI_OPS, query=<alert summary>)
|
||||
-> set SessionContextHolder(sessionId)
|
||||
-> ai_ops_supervisor -> planner_agent -> executor_agent
|
||||
-> persist agent_step and tool_invocation through existing hooks/tools
|
||||
-> extract final report
|
||||
-> persist diagnosis_session.answer/status/counts
|
||||
-> caller queries GET /api/diagnosis/{sessionId}/trace
|
||||
```
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] AIOps still lacks the Chat Verifier quality gate. -> Mitigation: document as follow-up and keep this slice focused on traceability.
|
||||
- [Risk] SSE clients may not parse the new first message. -> Mitigation: message is additive content; existing clients still receive the final report.
|
||||
- [Risk] Optional request body in Spring MVC can be easy to mishandle. -> Mitigation: use `@RequestBody(required = false)` and default request values in service code.
|
||||
- [Risk] AIOps generated reports may still depend on real infrastructure. -> Mitigation: demo profile already enables mock logs/metrics where available; full offline mode remains out of scope.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
- No database migration.
|
||||
- Deploy with application restart.
|
||||
- Rollback by reverting controller/service/DTO changes; existing persisted sessions remain valid.
|
||||
@@ -0,0 +1,29 @@
|
||||
## Why
|
||||
|
||||
The MVP already has a strong traceable chat diagnosis path, but the legacy `/api/ai_ops` endpoint still behaves like an early standalone demo: it accepts no alert payload, generates an internal session id that callers cannot reuse, and streams a report without reliably persisting the final answer for trace replay. For an Agent Engineer interview project, AIOps should become a second entry point into the same observable diagnosis story rather than a disconnected legacy path.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Allow `/api/ai_ops` to accept an optional alert diagnosis request body.
|
||||
- Resolve a stable session id from the request or generate one when omitted.
|
||||
- Persist the AIOps alert query and final report into `diagnosis_session`.
|
||||
- Emit the resolved session id in the SSE stream so reviewers can call `GET /api/diagnosis/{sessionId}/trace`.
|
||||
- Keep the existing AIOps planner/executor flow and evidence tools; do not replace it with the chat flow in this slice.
|
||||
- Document the AIOps demo path beside the existing MVP demo trace flow.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `aiops-traceable-diagnosis-entry`: Makes the AIOps alert endpoint traceable by session id and replayable through the existing diagnosis trace API.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- Existing `/api/ai_ops` behavior is extended from a no-input SSE trigger into an optional request-body alert diagnosis endpoint.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected code: `ChatController`, `AiOpsService`, `AIOpsRequest`, focused tests, MVP demo documentation, devflow records.
|
||||
- Affected API: `POST /api/ai_ops` remains SSE, but now accepts an optional JSON body and streams a first message containing `sessionId`.
|
||||
- Affected persistence: no schema migration; writes existing `diagnosis_session.query`, `answer`, `status`, timing, and aggregate counts.
|
||||
- Non-goals: no full AIOps/Chat service unification, no new database table, no production security cleanup, no full offline fake runtime, no mandatory Verifier integration for AIOps in this slice.
|
||||
+41
@@ -0,0 +1,41 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: AIOps endpoint accepts optional alert input
|
||||
The system SHALL allow `POST /api/ai_ops` to accept an optional JSON request body describing the alert diagnosis request.
|
||||
|
||||
#### Scenario: Caller supplies alert input
|
||||
- **WHEN** a caller posts to `/api/ai_ops` with alert fields such as session id, alert name, service, severity, description, and time range
|
||||
- **THEN** the AIOps analysis uses those fields to build the diagnosis task prompt and persisted session query
|
||||
|
||||
#### Scenario: Caller omits alert input
|
||||
- **WHEN** a caller posts to `/api/ai_ops` without a body
|
||||
- **THEN** the system still starts the default AIOps alert-analysis flow
|
||||
|
||||
### Requirement: AIOps session id is traceable
|
||||
The system SHALL resolve a stable AIOps session id from the request when provided, otherwise generate one, and SHALL expose that session id to the SSE caller.
|
||||
|
||||
#### Scenario: Request includes session id
|
||||
- **WHEN** a caller posts to `/api/ai_ops` with `sessionId`
|
||||
- **THEN** the created `diagnosis_session.session_id` equals that value
|
||||
- **AND** the SSE stream includes the same session id
|
||||
|
||||
#### Scenario: Request omits session id
|
||||
- **WHEN** a caller posts to `/api/ai_ops` without `sessionId`
|
||||
- **THEN** the system generates a session id
|
||||
- **AND** the SSE stream includes the generated session id
|
||||
|
||||
### Requirement: AIOps report is persisted for trace replay
|
||||
The system SHALL persist the final AIOps report into the existing `diagnosis_session.answer` field when a report is available.
|
||||
|
||||
#### Scenario: AIOps report is generated
|
||||
- **WHEN** the AIOps planner/executor flow returns a final report
|
||||
- **THEN** the corresponding diagnosis session is marked successful
|
||||
- **AND** `diagnosis_session.answer` stores the final report
|
||||
- **AND** `GET /api/diagnosis/{sessionId}/trace` can include that answer
|
||||
|
||||
### Requirement: AIOps trace uses existing evidence tables
|
||||
The system SHALL continue using existing `agent_step` and `tool_invocation` persistence for AIOps trace evidence.
|
||||
|
||||
#### Scenario: AIOps uses evidence tools
|
||||
- **WHEN** the AIOps flow calls available evidence tools
|
||||
- **THEN** existing hooks and recorders persist agent steps and tool invocations under the resolved AIOps session id
|
||||
@@ -0,0 +1,22 @@
|
||||
## 1. Flow Records
|
||||
|
||||
- [x] 1.1 Add devflow brief, decisions, evidence, and acceptance records for `aiops-traceable-diagnosis-entry`.
|
||||
- [x] 1.2 Record user-approved GitNexus skip and local impact analysis.
|
||||
|
||||
## 2. AIOps API And Service
|
||||
|
||||
- [x] 2.1 Extend `AIOpsRequest` with optional session id and alert fields.
|
||||
- [x] 2.2 Change `/api/ai_ops` to accept an optional request body and emit the resolved session id in SSE.
|
||||
- [x] 2.3 Change `AiOpsService` to accept the request, resolve session id, build a request-specific prompt, and persist the request summary.
|
||||
- [x] 2.4 Persist the final AIOps report to `diagnosis_session.answer`.
|
||||
|
||||
## 3. Demo Documentation
|
||||
|
||||
- [x] 3.1 Add an AIOps alert demo section to `mvp/demo/README.md`.
|
||||
- [x] 3.2 Add a concrete AIOps acceptance case under `mvp/demo`.
|
||||
|
||||
## 4. Verification
|
||||
|
||||
- [x] 4.1 Add focused unit coverage for AIOps request/session/report persistence behavior where practical.
|
||||
- [x] 4.2 Run targeted tests.
|
||||
- [x] 4.3 Run compile verification.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,54 @@
|
||||
## Context
|
||||
|
||||
The project now has the pieces needed for trace-based evaluation:
|
||||
|
||||
- `diagnosis_session` stores final answer, status, duration, counts, feedback, and `self_evaluation`
|
||||
- `agent_step` stores ordered agent execution records
|
||||
- `tool_invocation` stores evidence tool calls with normalized evidence semantics
|
||||
- `GET /api/diagnosis/{sessionId}/trace` can aggregate one diagnosis trace for demo review
|
||||
- `evidence-trace-hardening` defined stable `supported`, `no_evidence`, `deduped`, and `failed` semantics
|
||||
|
||||
P1-B should not add another runtime agent. It should create a repeatable evaluation surface that can be used after changing prompts, retrieval behavior, tools, or verifier logic.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
- Define fixed MVP diagnosis cases with expected evidence and verdict rules.
|
||||
- Build a deterministic evaluator that can validate a diagnosis trace against a case definition.
|
||||
- Produce JSON and Markdown reports with pass/fail status and key metrics.
|
||||
- Keep the first version usable without a real LLM by allowing fixture trace inputs.
|
||||
- Leave room for a later runtime mode that queries the trace API after a demo run.
|
||||
|
||||
**Non-Goals:**
|
||||
- No LLM-as-judge in this slice.
|
||||
- No automatic prompt optimization.
|
||||
- No new production API.
|
||||
- No change to chat, verifier, retrieval, upload, or feedback behavior.
|
||||
- No requirement to start MySQL/Redis/Milvus/LLM for the first offline evaluator.
|
||||
|
||||
## Decisions
|
||||
|
||||
| Decision | Choice | Alternative Considered | Rationale |
|
||||
|---|---|---|---|
|
||||
| Evaluation source | Start with fixture / persisted trace JSON input | Always run live `/api/chat` first | Keeps the first harness deterministic and avoids mixing quality checks with external infrastructure availability. |
|
||||
| Judging strategy | Rule-based trace validation | LLM-as-judge | The immediate goal is regression signal for evidence coverage and degraded behavior, not subjective answer scoring. |
|
||||
| Case format | Static JSON/YAML case definitions | Hard-coded Java tests only | Case files are easier to inspect and explain in interviews. |
|
||||
| Report format | JSON plus Markdown | Console-only output | JSON supports automation; Markdown supports quick human review. |
|
||||
| Metrics | Evidence coverage, verdict distribution, tool-call count, duration, answer keyword coverage | Full semantic correctness | These metrics are available from existing trace data and align with the MVP's observable contract. |
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Rule-based keyword checks can be brittle. -> Mitigation: keep checks focused on required evidence, verdicts, and high-signal root-cause terms rather than exact answer text.
|
||||
- [Risk] Fixture-only evaluation may drift from runtime behavior. -> Mitigation: design the evaluator around the same trace response shape so runtime traces can be fed in later.
|
||||
- [Risk] Metrics may encourage gaming tool counts. -> Mitigation: report tool counts as cost/efficiency signals, not the sole pass/fail criterion.
|
||||
- [Risk] Too many cases can slow iteration. -> Mitigation: start with 5 MVP cases and keep each case small.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
- No deployment migration is required.
|
||||
- The harness is additive and can be run locally as a test or script.
|
||||
- Rollback is deleting the eval case files, runner, and report docs.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Should runtime trace API polling be included in the first implementation, or left as a follow-up after the fixture validator lands?
|
||||
@@ -0,0 +1,27 @@
|
||||
## Why
|
||||
|
||||
The MVP can now run a traceable diagnosis flow, but it still lacks a repeatable way to evaluate whether changes to prompts, tools, retrieval, or verifier behavior improve or regress agent quality. A fixed diagnosis evaluation harness gives the project an interview-ready quality baseline instead of relying on a single manual demo.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add a small fixed evaluation set for representative MVP diagnosis scenarios.
|
||||
- Define expected assertions per case: root-cause keywords, required evidence tools, allowed verifier verdicts, and forbidden behavior.
|
||||
- Add a trace-based evaluator that checks persisted diagnosis traces for evidence coverage, verifier output, final answer shape, tool-call count, and duration.
|
||||
- Add JSON and Markdown report output for quick review after a run.
|
||||
- Add documentation that explains how this evaluation harness should be used during prompt/tool/verifier iteration.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- `diagnosis-eval-harness`: Defines fixed diagnosis cases, trace-based validation rules, and evaluation report output for MVP Agent regression checks.
|
||||
|
||||
### Modified Capabilities
|
||||
- None.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected areas: evaluation resources/scripts/tests, MVP demo documentation, and devflow records.
|
||||
- Affected runtime behavior: none. This change reads persisted trace data or fixture trace data and does not modify the chat execution path.
|
||||
- Affected APIs: none.
|
||||
- Dependencies: relies on the evidence semantics from `evidence-trace-hardening`, especially `tool_invocation`, `tool_trace_summary`, `verifier_evaluation`, and evidence status conventions.
|
||||
- Non-goals: no LLM-as-judge, no full offline LLM runtime, no new production endpoint, no schema migration.
|
||||
+59
@@ -0,0 +1,59 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Evaluation harness SHALL define fixed diagnosis cases
|
||||
The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria.
|
||||
|
||||
#### Scenario: Case definition includes expected evidence
|
||||
- **WHEN** an evaluation case is defined
|
||||
- **THEN** it SHALL include a case id, user question, expected root-cause keywords, required evidence tools, and allowed verifier verdicts
|
||||
|
||||
#### Scenario: Case definition can express forbidden behavior
|
||||
- **WHEN** a case has known unsafe behavior
|
||||
- **THEN** the case definition SHALL be able to declare forbidden answer keywords or forbidden verdicts
|
||||
|
||||
### Requirement: Evaluation harness SHALL validate diagnosis traces
|
||||
The system SHALL validate a diagnosis trace against the corresponding case definition using deterministic rules.
|
||||
|
||||
#### Scenario: Evidence coverage validation
|
||||
- **WHEN** a trace is evaluated
|
||||
- **THEN** the evaluator SHALL verify that required evidence tools appear in `toolInvocations` or verifier trace summaries
|
||||
|
||||
#### Scenario: Verifier evaluation validation
|
||||
- **WHEN** a trace is evaluated
|
||||
- **THEN** the evaluator SHALL verify that `selfEvaluation.verifier_evaluation.verdict` exists
|
||||
- **AND** the verdict SHALL be one of the case's allowed verdicts
|
||||
|
||||
#### Scenario: Answer keyword validation
|
||||
- **WHEN** a trace is evaluated
|
||||
- **THEN** the evaluator SHALL verify that the final answer includes at least the configured minimum root-cause keyword coverage
|
||||
|
||||
#### Scenario: Degraded output validation
|
||||
- **WHEN** a trace verdict is `REJECT`
|
||||
- **THEN** the evaluator SHALL verify that the final answer follows the degraded-output contract rather than leaking an unverified executor answer
|
||||
|
||||
### Requirement: Evaluation harness SHALL report quality and cost signals
|
||||
The system SHALL produce a report that summarizes pass/fail results and key trace metrics.
|
||||
|
||||
#### Scenario: JSON report output
|
||||
- **WHEN** an evaluation run completes
|
||||
- **THEN** the evaluator SHALL output a JSON report with per-case result, failed checks, verdict, evidence coverage, tool-call count, and duration
|
||||
|
||||
#### Scenario: Markdown report output
|
||||
- **WHEN** an evaluation run completes
|
||||
- **THEN** the evaluator SHALL output a Markdown report suitable for review in the repository
|
||||
|
||||
#### Scenario: Aggregate metrics
|
||||
- **WHEN** multiple cases are evaluated
|
||||
- **THEN** the report SHALL include aggregate pass rate, verdict distribution, average tool-call count, and average duration when available
|
||||
|
||||
### Requirement: Evaluation harness SHALL support offline fixture mode
|
||||
The first evaluator version SHALL be runnable without live MySQL, Redis, Milvus, or LLM services by evaluating saved trace fixtures.
|
||||
|
||||
#### Scenario: Fixture trace evaluation
|
||||
- **WHEN** the evaluator is run against a directory of trace fixture files
|
||||
- **THEN** it SHALL evaluate each trace file against its matching case definition
|
||||
- **AND** it SHALL not require a running application service
|
||||
|
||||
#### Scenario: Missing fixture is reported clearly
|
||||
- **WHEN** a case has no matching trace fixture
|
||||
- **THEN** the evaluator SHALL mark the case as not run or failed with a clear reason
|
||||
@@ -0,0 +1,25 @@
|
||||
## 1. Case Definitions
|
||||
|
||||
- [x] 1.1 Add evaluation case definition format for fixed MVP diagnosis scenarios.
|
||||
- [x] 1.2 Add the first 5 case definitions: payment timeout, MySQL pool exhausted, Redis timeout, slow response, and JVM memory risk.
|
||||
- [x] 1.3 Document the meaning of expected keywords, required evidence tools, allowed verdicts, and forbidden behavior.
|
||||
|
||||
## 2. Trace Fixtures
|
||||
|
||||
- [x] 2.1 Add fixture trace schema or DTOs that match `DiagnosisTraceResponse` enough for offline evaluation.
|
||||
- [x] 2.2 Add at least one representative trace fixture for a passing case.
|
||||
- [x] 2.3 Add at least one fixture covering low-confidence or degraded behavior.
|
||||
|
||||
## 3. Evaluator
|
||||
|
||||
- [x] 3.1 Implement trace validation rules for evidence coverage, verifier verdict, answer keyword coverage, and degraded-output contract.
|
||||
- [x] 3.2 Implement aggregate metrics: pass rate, verdict distribution, average tool-call count, and average duration.
|
||||
- [x] 3.3 Implement JSON report output.
|
||||
- [x] 3.4 Implement Markdown report output.
|
||||
|
||||
## 4. Tests And Documentation
|
||||
|
||||
- [x] 4.1 Add focused offline tests for the evaluator.
|
||||
- [x] 4.2 Add run instructions under `mvp/demo` or `mvp/notes`.
|
||||
- [x] 4.3 Run targeted tests for the evaluator.
|
||||
- [x] 4.4 Run compile verification.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,64 @@
|
||||
## Context
|
||||
|
||||
The current MVP already has the core pieces required for traceable agent execution:
|
||||
|
||||
- `LookupKnowledgeTool` writes rich retrieval metadata into `tool_invocation`
|
||||
- `QueryLogsTools` and `QueryMetricsTools` use `ToolInvocationRecorder`
|
||||
- `ToolTraceSummaryService` turns persisted rows into verifier-facing evidence summaries
|
||||
- `ChatService` already contains fallback behavior for missing or invalid `verifier_output`
|
||||
|
||||
The gap is no longer “there is no evidence trace”. The gap is that the evidence trace contract is split across two persistence paths and several implicit conventions:
|
||||
|
||||
- `LookupKnowledgeTool` builds `ToolInvocation` rows itself
|
||||
- the other evidence tools use `ToolInvocationRecorder.recordEvidenceTool(...)`
|
||||
- “failed”, “no evidence”, “deduped”, and “successful but weak” are inferred differently across tools
|
||||
- degraded output behavior exists in code but is only lightly covered by tests
|
||||
|
||||
For interview-facing hardening, this slice should make those semantics explicit and testable without changing the database schema or the overall multi-agent workflow.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
- Centralize the common persistence contract for evidence-bearing tools.
|
||||
- Preserve `lookup_knowledge`-specific retrieval fields while removing ad hoc duplication in how evidence rows are created.
|
||||
- Define stable summarization semantics for:
|
||||
- successful evidence
|
||||
- no-hit / no-usable-evidence
|
||||
- deduped retrievals
|
||||
- failed evidence queries
|
||||
- Make `ChatService` fallback and degraded-output paths testable as explicit product behavior.
|
||||
- Keep the scope small enough to unblock the next P1-B evaluation harness.
|
||||
|
||||
**Non-Goals:**
|
||||
- No new table, column, or Flyway migration.
|
||||
- No new public API.
|
||||
- No new verifier verdict type beyond `PASS` / `LOW_CONFID` / `REJECT`.
|
||||
- No attempt to redesign planner/executor routing.
|
||||
- No full offline runtime or end-to-end benchmark harness in this slice.
|
||||
|
||||
## Decisions
|
||||
|
||||
| Decision | Choice | Alternative Considered | Rationale |
|
||||
|---|---|---|---|
|
||||
| Evidence persistence ownership | Keep `ToolInvocationRecorder` as the single common entry point | Let each tool continue building `ToolInvocation` rows ad hoc | The recorder already exists and is the right seam for contract hardening. |
|
||||
| `lookup_knowledge` integration style | Add a richer recorder entry path for retrieval-aware calls | Force `lookup_knowledge` into the same minimal method used by logs/metrics | `lookup_knowledge` carries domain-specific fields such as L0/L1 counts, relevance, dedup reason, and retrieval details that should stay structured. |
|
||||
| No-evidence semantics | Distinguish failed calls from successful calls that yield no usable evidence | Collapse all non-successful evidence into one bucket | Verifier and future evaluation harnesses need to separate “tool broke” from “tool succeeded but found nothing useful”. |
|
||||
| Degraded-path hardening | Add focused unit tests around verifier fallback and output shaping | Rely on runtime demo only | Interview value comes from proving the system fails predictably, not just that the happy path ran once. |
|
||||
| Scope boundary | Keep changes additive and contract-oriented | Expand into P1-B evaluation harness in the same change | This keeps the slice reviewable and avoids mixing infrastructure hardening with evaluation product work. |
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Tightening persistence semantics could subtly change existing trace summaries. -> Mitigation: keep field names stable and add regression tests around summary output.
|
||||
- [Risk] Over-generalizing the recorder could make retrieval-specific rows less informative. -> Mitigation: keep a retrieval-aware recording path rather than flattening all tools to the same minimal payload.
|
||||
- [Risk] Tests may lock in the current fallback copy too aggressively. -> Mitigation: assert protocol-level behavior and key phrases, not brittle full-string snapshots.
|
||||
- [Risk] `lookup_knowledge` dedup semantics are product-specific and may not fit generic “success/failure” labels cleanly. -> Mitigation: preserve `dedupReason` and treat dedup as a first-class no-new-evidence case in summary logic.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
- No deployment migration is required beyond shipping the code changes.
|
||||
- Existing `tool_invocation` rows remain valid because this change reuses the same schema.
|
||||
- Rollback is code-only: revert the recorder/summary/fallback hardening and keep the persisted rows as-is.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Should P1-B metrics count deduped retrievals as “no-evidence”, or report them as a separate category? This change will preserve enough structure to decide later without another schema change.
|
||||
@@ -0,0 +1,26 @@
|
||||
## Why
|
||||
|
||||
The MVP already persists evidence tool invocations and uses a Verifier to judge answer quality, but the current evidence trace semantics are still only partially standardized. For interview-grade agent engineering, the system needs a tighter contract for evidence persistence, no-evidence/failure states, and degraded output behavior, plus focused tests that prove those paths work offline.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Standardize the persisted evidence-tool contract across `lookup_knowledge`, `query_logs`, and `query_metrics`.
|
||||
- Align how evidence tools represent success, no-hit, deduped, and failed calls so `ToolTraceSummaryService` can summarize them consistently.
|
||||
- Harden `ChatService` fallback behavior for invalid or missing verifier output and make the degraded-output paths explicitly testable.
|
||||
- Add focused offline tests for evidence recording, trace summarization, and verifier fallback / degraded output behavior.
|
||||
- Record this slice as a dedicated P1-A change tied to the interview-focused MVP hardening track.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
- `evidence-trace-hardening`: Covers standardized evidence invocation persistence, verifier-facing evidence summary semantics, and explicit degraded-output contracts for evidence gaps and verifier failures.
|
||||
|
||||
### Modified Capabilities
|
||||
- `chat-verifier-agent`: Tightens verifier input evidence semantics and fallback guarantees without changing the high-level planner/executor/verifier workflow.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected code: `ToolInvocationRecorder`, `LookupKnowledgeTool`, `QueryLogsTools`, `QueryMetricsTools`, `ToolTraceSummaryService`, `ChatService`, and focused test classes.
|
||||
- Affected runtime behavior: evidence-bearing tools will persist more consistent invocation semantics; verifier fallback and degraded outputs remain additive hardening, not a product-flow rewrite.
|
||||
- Affected APIs: none. No new endpoint or schema is introduced.
|
||||
- Non-goals: no new evidence tools, no database migration, no evaluation harness, no trace UI, no security/config cleanup in this slice.
|
||||
+14
@@ -0,0 +1,14 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Verifier inputs SHALL tolerate hardened no-evidence semantics
|
||||
The verifier integration SHALL continue to work when evidence summaries distinguish failed calls, no-hit calls, and deduped retrievals more explicitly.
|
||||
|
||||
#### Scenario: Failed evidence remains a verifier-visible gap
|
||||
- **WHEN** an evidence-bearing tool invocation fails
|
||||
- **THEN** the verifier-facing trace summary SHALL preserve that failure as a gap
|
||||
- **AND** the verifier flow SHALL continue without crashing
|
||||
|
||||
#### Scenario: Deduped retrievals do not count as fresh support
|
||||
- **WHEN** the verifier-facing trace summary contains deduped `lookup_knowledge` entries
|
||||
- **THEN** those entries SHALL be treated as no-new-evidence
|
||||
- **AND** they SHALL NOT be interpreted as fresh direct support for the answer
|
||||
+82
@@ -0,0 +1,82 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Evidence-bearing tools SHALL persist a unified invocation contract
|
||||
The system SHALL persist evidence-bearing tool calls through a unified contract that guarantees the same baseline fields across `lookup_knowledge`, `query_logs`, and `query_metrics`.
|
||||
|
||||
#### Scenario: Common evidence fields are always persisted
|
||||
- **WHEN** an evidence-bearing tool finishes a call
|
||||
- **THEN** the persisted `tool_invocation` row SHALL include `session_id`, `tool_name`, `input_params`, `duration_ms`, `success`, and an output preview or explicit no-output state
|
||||
|
||||
#### Scenario: Retrieval-aware tools preserve structured retrieval fields
|
||||
- **WHEN** `lookup_knowledge` persists a tool invocation
|
||||
- **THEN** the row SHALL preserve retrieval-specific fields such as retrieval layer, L0/L1 counts, relevance level, dedup reason, and retrieval details
|
||||
- **AND** those fields SHALL be written through the same recorder contract rather than by ad hoc row construction in unrelated code paths
|
||||
|
||||
### Requirement: Evidence trace semantics SHALL distinguish failure from no-evidence outcomes
|
||||
The system SHALL keep failed calls separate from successful calls that return no usable evidence.
|
||||
|
||||
#### Scenario: Tool failure is preserved as failure
|
||||
- **WHEN** an evidence-bearing tool throws, times out, or returns an execution error
|
||||
- **THEN** the persisted row SHALL set `success=false`
|
||||
- **AND** it SHALL preserve an `error_message` explaining the failure
|
||||
|
||||
#### Scenario: No usable evidence is preserved without pretending success
|
||||
- **WHEN** an evidence-bearing tool completes normally but yields no usable evidence for the verifier
|
||||
- **THEN** the persisted contract SHALL preserve that the call completed
|
||||
- **AND** the verifier-facing summary SHALL describe it as a no-evidence outcome rather than direct support
|
||||
|
||||
#### Scenario: Deduped retrieval remains auditable
|
||||
- **WHEN** `lookup_knowledge` is blocked by session-level deduplication
|
||||
- **THEN** the persisted row SHALL preserve the dedup reason
|
||||
- **AND** the verifier-facing summary SHALL treat that row as no-new-evidence rather than as a fresh supporting hit
|
||||
|
||||
### Requirement: Verifier-facing evidence summaries SHALL use stable no-evidence rules
|
||||
The system SHALL summarize persisted tool rows into verifier-facing evidence entries using stable rules for success, failure, no-hit, and deduped outcomes.
|
||||
|
||||
#### Scenario: Failed evidence calls remain visible in the summary
|
||||
- **WHEN** `ToolTraceSummaryService` processes failed evidence-bearing tool rows
|
||||
- **THEN** the summary SHALL retain them
|
||||
- **AND** it SHALL mark them as unsuccessful evidence with an output summary that explains the gap
|
||||
|
||||
#### Scenario: No-hit and deduped calls do not upgrade evidence level
|
||||
- **WHEN** `ToolTraceSummaryService` processes rows that returned no usable evidence or were deduped
|
||||
- **THEN** those rows SHALL NOT be promoted to direct or indirect evidence
|
||||
- **AND** their counts SHALL still be reflected in the merged summary entry
|
||||
|
||||
#### Scenario: Successful evidence keeps the strongest available support
|
||||
- **WHEN** multiple rows for the same tool and topic domain are merged
|
||||
- **THEN** the summary SHALL preserve the strongest successful evidence level among them
|
||||
- **AND** it SHALL also retain repeated-call, failed-call, and no-hit counts for auditability
|
||||
|
||||
### Requirement: ChatService SHALL degrade predictably on verifier output failures
|
||||
The system SHALL treat missing or invalid verifier output as a bounded degraded path instead of an unstructured runtime error.
|
||||
|
||||
#### Scenario: Missing verifier output falls back to LOW_CONFID
|
||||
- **WHEN** the verifier step completes without a usable `verifier_output`
|
||||
- **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision
|
||||
- **AND** the final user-facing output SHALL use the fixed low-confidence protocol
|
||||
|
||||
#### Scenario: Invalid verifier JSON falls back to LOW_CONFID
|
||||
- **WHEN** the verifier returns malformed or non-parseable JSON
|
||||
- **THEN** `ChatService` SHALL fall back to a `LOW_CONFID` decision
|
||||
- **AND** the fallback SHALL still persist a verifier-evaluation record
|
||||
|
||||
#### Scenario: REJECT output hides unverified raw answer text
|
||||
- **WHEN** the final verifier decision is `REJECT`
|
||||
- **THEN** the user-facing output SHALL use the degraded template
|
||||
- **AND** it SHALL NOT pass through the raw executor answer
|
||||
|
||||
### Requirement: Evidence-trace hardening SHALL be covered by focused offline tests
|
||||
The system SHALL provide offline tests for the hardened evidence contract and degraded-output behavior.
|
||||
|
||||
#### Scenario: Evidence recorder contract is tested offline
|
||||
- **WHEN** the test suite runs the focused recorder tests
|
||||
- **THEN** it SHALL verify the persisted semantics for successful, failed, and no-evidence evidence-tool calls without requiring external infrastructure
|
||||
|
||||
#### Scenario: Trace summary hardening is tested offline
|
||||
- **WHEN** the test suite runs the focused trace-summary tests
|
||||
- **THEN** it SHALL verify merged summary behavior for mixed success, failure, no-hit, and deduped tool rows
|
||||
|
||||
#### Scenario: Verifier fallback behavior is tested offline
|
||||
- **WHEN** the test suite runs the focused `ChatService` fallback tests
|
||||
- **THEN** it SHALL verify the missing-output, invalid-JSON, `LOW_CONFID`, and `REJECT` degraded paths without requiring a real LLM or database
|
||||
@@ -0,0 +1,22 @@
|
||||
## 1. Evidence Persistence Contract
|
||||
|
||||
- [x] 1.1 Extend `ToolInvocationRecorder` with a richer evidence-recording path that can preserve retrieval-aware fields as well as common evidence fields.
|
||||
- [x] 1.2 Refactor `LookupKnowledgeTool` to persist `tool_invocation` rows through `ToolInvocationRecorder` instead of its own ad hoc row-construction path.
|
||||
- [x] 1.3 Align `QueryLogsTools` and `QueryMetricsTools` no-hit / failure payloads with the hardened evidence contract.
|
||||
|
||||
## 2. Verifier-Facing Summary Semantics
|
||||
|
||||
- [x] 2.1 Harden `ToolTraceSummaryService` so failed, no-hit, and deduped evidence rows are summarized with stable no-evidence semantics.
|
||||
- [x] 2.2 Preserve merged-call counts for repeated hits, failures, and no-new-evidence rows without overstating evidence strength.
|
||||
|
||||
## 3. Chat Degraded Paths
|
||||
|
||||
- [x] 3.1 Add focused `ChatService` tests for missing verifier output fallback to `LOW_CONFID`.
|
||||
- [x] 3.2 Add focused `ChatService` tests for invalid verifier JSON fallback to `LOW_CONFID`.
|
||||
- [x] 3.3 Add focused `ChatService` tests that `REJECT` output uses the degraded template and does not leak raw executor answer content.
|
||||
|
||||
## 4. Verification
|
||||
|
||||
- [x] 4.1 Add focused offline tests for the recorder contract and `ToolTraceSummaryService`.
|
||||
- [x] 4.2 Run targeted test commands for the new/updated offline tests.
|
||||
- [x] 4.3 Run compile verification.
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-04
|
||||
@@ -0,0 +1,39 @@
|
||||
## Context
|
||||
|
||||
The eval harness now has a complete five-case fixture baseline and saved JSON / Markdown baseline reports. The missing piece is a deterministic comparison step that explains whether a new report is better, worse, or just different from the baseline.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Compare two `DiagnosisEvalReport` objects without requiring external services.
|
||||
- Surface aggregate regressions such as pass-rate drops, verdict distribution shifts, and cost increases.
|
||||
- Surface per-case regressions such as pass-to-fail changes, missing evidence coverage, verdict changes, keyword coverage loss, and missing cases.
|
||||
- Write JSON and Markdown diff outputs for review.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not run the Agent or regenerate traces.
|
||||
- Do not introduce LLM-as-judge.
|
||||
- Do not change evaluator scoring rules.
|
||||
- Do not block on performance thresholds beyond simple numeric diff signals.
|
||||
|
||||
## Decisions
|
||||
|
||||
- Decision: Compare report DTOs instead of raw traces.
|
||||
- Reason: `DiagnosisEvalReport` is already the stable structured output of the evaluator and is cheaper to diff than trace internals.
|
||||
- Alternative considered: compare raw trace fixtures. That would expose more detail but duplicate evaluator responsibilities.
|
||||
|
||||
- Decision: Classify each diff item as `REGRESSION`, `IMPROVEMENT`, or `CHANGED`.
|
||||
- Reason: interview and CI usage both need a quick answer to "did this get worse?" while still preserving neutral changes.
|
||||
- Alternative considered: only output numeric deltas. That is harder to scan and less actionable.
|
||||
|
||||
- Decision: Keep thresholds explicit and conservative.
|
||||
- Reason: pass/fail and missing evidence are hard regressions; tool calls and duration are cost signals that should be visible even if not always blocking.
|
||||
- Alternative considered: fail only on pass-rate drop. That misses cases where quality stays green but cost or confidence behavior changes.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- Report comparison can only see fields already captured by `DiagnosisEvalReport`. Mitigation: use this as the first regression layer and add richer report fields later if needed.
|
||||
- Duration may fluctuate in live runs. Mitigation: fixture baseline uses stable durations; live-mode thresholds can be added later.
|
||||
- Verdict distribution changes can be intentional. Mitigation: classify them as `CHANGED` unless they coincide with per-case regressions.
|
||||
@@ -0,0 +1,28 @@
|
||||
## Why
|
||||
|
||||
The evaluation baseline is now complete, but developers still need a repeatable way to decide whether a new Agent run regressed against that baseline. A deterministic baseline diff turns saved reports into an actionable regression signal instead of a static artifact.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add a baseline diff model that compares two `DiagnosisEvalReport` objects.
|
||||
- Detect aggregate changes such as pass-rate drops, verdict distribution shifts, tool-call cost changes, and duration changes.
|
||||
- Detect per-case changes such as pass/fail regression, verdict changes, keyword coverage changes, evidence coverage loss, and missing/new cases.
|
||||
- Add JSON and Markdown diff output suitable for review.
|
||||
- Document how to interpret the diff in the eval docs.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `diagnosis-eval-harness`: Extend the existing evaluation harness so a current report can be compared against the saved baseline report.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affects eval-only Java code under `src/main/java/com/superbiz/agent/eval`.
|
||||
- Adds focused tests under `src/test/java/com/superbiz/agent/eval`.
|
||||
- Updates `mvp/eval` documentation and may add sample diff output.
|
||||
- No production Agent runtime, API, database schema, or external dependency changes are expected.
|
||||
+46
@@ -0,0 +1,46 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Evaluation harness SHALL compare reports against a baseline
|
||||
The system SHALL compare a current diagnosis evaluation report against a saved baseline report using deterministic rules.
|
||||
|
||||
#### Scenario: Aggregate regression detection
|
||||
- **WHEN** the current report has a lower pass rate than the baseline report
|
||||
- **THEN** the diff SHALL record a regression with the old value, new value, and delta
|
||||
|
||||
#### Scenario: Cost signal detection
|
||||
- **WHEN** average tool-call count or average duration changes between reports
|
||||
- **THEN** the diff SHALL record the baseline value, current value, and delta
|
||||
|
||||
#### Scenario: Verdict distribution comparison
|
||||
- **WHEN** verdict counts differ between reports
|
||||
- **THEN** the diff SHALL record the verdict distribution changes
|
||||
|
||||
### Requirement: Evaluation harness SHALL compare case-level report results
|
||||
The system SHALL compare case results by case id and report actionable per-case changes.
|
||||
|
||||
#### Scenario: Case pass/fail regression
|
||||
- **WHEN** a case changes from passing in the baseline to failing in the current report
|
||||
- **THEN** the diff SHALL record a regression for that case
|
||||
|
||||
#### Scenario: Evidence coverage regression
|
||||
- **WHEN** a required evidence tool changes from covered to uncovered for a case
|
||||
- **THEN** the diff SHALL record a regression naming the case and tool
|
||||
|
||||
#### Scenario: Missing case detection
|
||||
- **WHEN** a baseline case is absent from the current report
|
||||
- **THEN** the diff SHALL record a regression for the missing case
|
||||
|
||||
#### Scenario: New case detection
|
||||
- **WHEN** a current report contains a case absent from the baseline
|
||||
- **THEN** the diff SHALL record the case as a non-regression change
|
||||
|
||||
### Requirement: Evaluation harness SHALL report baseline diff results
|
||||
The system SHALL expose baseline diff output in structured JSON and reviewable Markdown formats.
|
||||
|
||||
#### Scenario: JSON diff output
|
||||
- **WHEN** a baseline diff is written as JSON
|
||||
- **THEN** it SHALL include aggregate summary fields and detailed diff items
|
||||
|
||||
#### Scenario: Markdown diff output
|
||||
- **WHEN** a baseline diff is written as Markdown
|
||||
- **THEN** it SHALL include a readable summary and a table of diff items
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user