Compare commits

...
Author SHA1 Message Date
aruo ed267d753d feat: add aiops lightweight verifier 2026-07-05 13:44:30 +08:00
aruo 2658742119 docs: archive rag and aiops query changes 2026-07-05 13:12:47 +08:00
aruo 72a3dbf8c5 feat: add aiops payload query augmentation 2026-07-05 12:56:20 +08:00
aruo 674dd27a48 feat: add rag post-reindex acceptance 2026-07-05 12:29:24 +08:00
aruo c7e2fc2ee2 feat: include breadcrumb in embedding text 2026-07-05 12:15:27 +08:00
aruo 1bfe1a17b4 docs: add rag refactor story 2026-07-05 12:09:32 +08:00
aruo 9dd6823fe7 docs: add rag retrieval quality report 2026-07-05 11:29:16 +08:00
aruo 9376448804 docs: add rag vectorstore interview notes 2026-07-05 11:13:50 +08:00
aruo f2bae0382c fix: align vectorstore live retrieval 2026-07-05 10:53:45 +08:00
aruo 5c71f5fc79 feat: integrate spring ai vectorstore fallback 2026-07-05 10:20:29 +08:00
aruo b9ec07de57 feat: add spring ai retrieval sidecar 2026-07-05 03:22:03 +08:00
aruo 5197712719 feat: add rag evidence postprocess blocks 2026-07-05 03:03:18 +08:00
aruo 4a94c14feb feat: treat l0 retrieval as domain hint 2026-07-05 02:18:40 +08:00
aruo 9a2a44d1b5 test: add rag retrieval baseline 2026-07-05 02:02:27 +08:00
aruo 79feed3314 Merge branch 'emdash/shy-items-fry-f4zze' into refactor/mvp1.0
# Conflicts:
#	mvp/issues/README.md
2026-07-05 01:42:42 +08:00
aruo 98155ae1d8 Merge branch 'aiops-trace-scope' into refactor/mvp1.0 2026-07-05 01:42:23 +08:00
aruo 2609c5a5ab docs: consolidate rag refactor issues 2026-07-05 01:40:08 +08:00
aruo bf5286c8f4 docs: add interview project materials 2026-07-05 01:39:37 +08:00
aruo 26e12a8d6b Archive MVP demo interview runbook 2026-07-05 01:34:04 +08:00
aruo cbef3ddd3c Add MVP demo interview runbook 2026-07-05 01:25:20 +08:00
aruo 69deb15330 Add diagnosis eval baseline diff 2026-07-05 00:59:53 +08:00
aruo 4c7c53b024 Expand diagnosis eval fixtures 2026-07-05 00:27:57 +08:00
aruo ca5c61fabf Add diagnosis eval harness 2026-07-04 23:51:43 +08:00
aruo 23ee05c7c3 feat: add traceable scoped AIOps diagnosis 2026-07-04 22:57:28 +08:00
aruo dc6cd32a67 Harden evidence trace semantics 2026-07-04 22:36:30 +08:00
235 changed files with 13072 additions and 283 deletions
+4
View File
@@ -60,3 +60,7 @@ uploads/
### Windows / Runtime Artifacts
*.stackdump
NUL
### MVP Demo Generated Outputs
mvp/demo/output/*.json
!mvp/demo/output/README.md
+7
View File
@@ -4,7 +4,14 @@
| 日期 | slug | 领域 | 关键词 | 状态 |
|---|---|---|---|---|
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
| 2026-07-04 | diagnosis-eval-harness | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
| 2026-07-04 | evidence-trace-hardening | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
| 2026-07-03 | mvp-demo-trace-acceptance | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
| 2026-07-04 | aiops-traceable-diagnosis-entry | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
| 2026-07-04 | aiops-alert-scope-control | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
| 2026-05-29 | chatmodel-abstraction | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | archived |
| 2026-06-23 | phase1-infrastructure | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | archived |
| 2026-06-24 | lookup-knowledge-integration | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | archived |
@@ -0,0 +1,14 @@
# Acceptance: aiops-alert-scope-control
## Verification
- [x] Payload-mode prompt focuses the final report on the supplied alert.
- [x] No-payload prompt requires active-alert discovery first.
- [x] Targeted tests pass.
- [x] Compile passes.
- [x] OpenSpec validates.
## Known Limits
- Prompt-only scope control may still require runtime observation.
- AIOps Verifier remains deferred.
@@ -0,0 +1,28 @@
# Brief: aiops-alert-scope-control
## Background
After `aiops-traceable-diagnosis-entry`, AIOps can be triggered by payload and replayed through trace. Runtime verification showed one semantic gap: payload mode still produced a broad report over all active mock alerts.
## Goal
Make AIOps scope explicit:
- Payload present -> targeted diagnosis for the supplied alert.
- Payload absent -> automatic active-alert discovery and diagnosis.
## Scope
- In scope:
- `AiOpsService.buildTaskPrompt(...)` scope rules.
- Focused tests.
- Demo acceptance wording.
- Out of scope:
- Verifier integration.
- Java-side filtering of tool results.
- API shape changes.
- Database changes.
## Related OpenSpec
`openspec/changes/aiops-alert-scope-control/`
@@ -0,0 +1,42 @@
# Decisions: aiops-alert-scope-control
## Clarify
- Entry summary: tighten AIOps report scope after runtime verification showed payload mode still analyzes all active alerts.
- Slug: `aiops-alert-scope-control`
- Scale: standard-light.
## Context
- AIOps traceability is implemented and verified.
- Mock Prometheus returns multiple active alerts.
- Payload demo supplies `HighCPUUsage/payment-service`, but previous report expanded to `HighMemoryUsage` and `SlowResponse`.
## Grill Question Pool
| # | Dimension | Question | Mode | Status |
|---|---|---|---|---|
| Q1 | Product Boundary | What makes `/api/ai_ops` different from `/api/chat` when payload exists? | evidence-driven | Payload is alert-event driven and should be scoped to that event. |
| Q2 | Scope | Should payload mode ignore all other active alerts? | user-interview | No; mention only as related risk/context. |
| Q3 | Compatibility | Should no-payload mode keep old "query active alerts" behavior? | evidence-driven | Yes. |
| Q4 | Enforcement | Should Java filter unrelated tool results now? | evidence-driven | No; prompt-only is sufficient for this small change. |
| Q5 | Verifier | Should this change add AIOps Verifier? | user-interview | No; keep deferred. |
## Evidence-Driven Conclusions
| Conclusion | Evidence Source | Result |
|---|---|---|
| Scope issue is prompt-level. | `/api_ ai_ops` trace showed all mock alerts analyzed despite payload. | Update task prompt. |
| No API or persistence changes are needed. | `AIOpsRequest` already carries payload and trace works. | Keep endpoint unchanged. |
| Blast radius is low. | `buildTaskPrompt(...)` is internal to `AiOpsService`. | Add tests for prompt content. |
## GitNexus
GitNexus remains skipped by prior user decision and because tools are not exposed in this session. Local impact analysis is recorded instead.
## Key Decisions
- Payload mode is detected when any alert field is present.
- Payload mode final report must focus on the supplied alert.
- No-payload mode must first call `queryPrometheusAlerts`.
- Other active alerts in payload mode can appear only as related risk, not as separate root-cause sections.
@@ -0,0 +1,44 @@
# Evidence: aiops-alert-scope-control
## Local Impact Analysis
- `AiOpsService.buildTaskPrompt(...)` is used by `executeAiOpsAnalysis(...)`.
- No controller, DTO, repository, or database changes are required.
- Existing `AiOpsServiceTest` already exercises request summary helpers and can be extended for scope prompt rules.
## Verification Results
- `mvn -q "-Dtest=AiOpsServiceTest" test` passed.
- `mvn -q -DskipTests compile` passed.
- `openspec.cmd validate aiops-alert-scope-control --strict` passed.
## Runtime Verification
- Runtime session: `mvp-demo-aiops-payment-cpu-codex-scope-003`.
- `/api/ai_ops` SSE emitted the requested `session` message and finished with `done`.
- `diagnosis_session` persisted:
- `agent_flow = AI_OPS`
- `status = SUCCESS`
- `total_duration_ms = 69875`
- `step_count = 5`
- `tool_call_count = 8`
- Tool invocation counts:
- `query_metrics = 1`
- `lookup_knowledge = 1`
- `query_logs = 6`
- Report scope check:
- `告警根因分析 - HighCPUUsage` exists.
- `告警根因分析 - HighMemoryUsage` does not exist.
- `告警根因分析 - SlowResponse` does not exist.
- `相关风险告警` exists.
## Runtime Fix
- Added Hikari settings in `src/main/resources/application.yml` after the first runtime attempt failed on stale MySQL pool connections:
- `maximum-pool-size: 5`
- `minimum-idle: 1`
- `connection-timeout: 10000`
- `validation-timeout: 5000`
- `idle-timeout: 60000`
- `max-lifetime: 120000`
- `keepalive-time: 30000`
@@ -0,0 +1,18 @@
# Acceptance: aiops-traceable-diagnosis-entry
## Verification
- [x] OpenSpec validates for `aiops-traceable-diagnosis-entry`.
- [x] Targeted AIOps service tests pass.
- [x] Compile verification passes.
- [x] Demo docs describe AIOps request -> session id -> trace query.
## Result
Accepted for implementation scope.
## Known Limits
- AIOps Verifier integration is deferred.
- Runtime still depends on configured model and infrastructure.
- Full browser/SSE runtime verification is not guaranteed in this coding pass.
@@ -0,0 +1,28 @@
# Brief: aiops-traceable-diagnosis-entry
## Background
The MVP chat diagnosis path is now traceable through `diagnosis_session`, `agent_step`, `tool_invocation`, and `GET /api/diagnosis/{sessionId}/trace`. The older `/api/ai_ops` endpoint still acts like a standalone SSE demo: it accepts no alert payload, generates an internal session id, and does not make trace replay obvious to callers.
## Goal
Turn AIOps into an alert-triggered diagnosis entry point that shares the same evidence and trace story as the main MVP, without rewriting the whole AIOps flow.
## Scope
- In scope:
- Optional AIOps alert request body.
- Stable request/session id propagation.
- Persisted AIOps query summary and final answer.
- SSE session id event.
- Demo documentation and focused tests.
- Out of scope:
- Full AIOps and ChatService unification.
- AIOps Verifier integration.
- Database schema changes.
- Sensitive configuration cleanup.
- Fully offline runtime.
## Related OpenSpec
`openspec/changes/aiops-traceable-diagnosis-entry/`
@@ -0,0 +1,68 @@
# Decisions: aiops-traceable-diagnosis-entry
## Clarify
- Entry summary: make the legacy AIOps SSE endpoint a traceable alert diagnosis entry for the Agent Engineer interview MVP.
- Slug: `aiops-traceable-diagnosis-entry`
- Scale: standard-light, because this extends one public endpoint and reuses existing persistence/trace infrastructure.
## Context
- `mvp-demo-trace-acceptance` already added `GET /api/diagnosis/{sessionId}/trace`.
- `chat-verifier-agent` made the chat path stronger than the older AIOps path.
- Current AIOps value is as a second entry point: system alert -> automated diagnosis -> evidence trace.
## Grill Question Pool
| # | Dimension | Question | Mode | Status |
|---|---|---|---|---|
| Q1 | Positioning | Is AIOps an independent product path or an alert-triggered sibling of Chat Diagnosis? | user-interview | Resolved: sibling entry, unified trace story |
| Q2 | API | Should we keep `/api/ai_ops` or add a new endpoint? | evidence-driven | Resolved: keep existing endpoint and extend optional body |
| Q3 | Input | What is the minimum alert payload? | user-interview | Resolved: `sessionId`, `alertName`, `service`, `severity`, `description`, `timeRange`, plus `userRequest` fallback |
| Q4 | Output | How does the caller learn the trace session id? | evidence-driven | Resolved: first SSE event uses type `session` |
| Q5 | Trace | Must AIOps be replayable with existing trace API? | evidence-driven | Resolved: yes, this is the main acceptance criterion |
| Q6 | Verifier | Must this slice add AIOps Verifier? | user-interview | Resolved: no, defer as follow-up |
| Q7 | Compatibility | Should no-body calls still work? | evidence-driven | Resolved: yes, preserve old demo behavior |
| Q8 | GitNexus | Should unavailable GitNexus block implementation? | user-interview | Resolved: skip GitNexus by user decision |
## Evidence-Driven Conclusions
| Conclusion | Evidence Source | Result |
|---|---|---|
| AIOps is currently isolated from request-driven trace replay. | `ChatController.aiOps()` has no request body; `AiOpsService` creates its own random session id. | Extend endpoint and service. |
| No schema change is needed. | `DiagnosisSession` already has `query`, `agentFlow`, `answer`, counts, and status. | Reuse existing table. |
| Trace API can already replay AIOps if session id and answer are persisted. | `DiagnosisTraceService` loads by session id and is flow-agnostic. | Keep trace API unchanged. |
| Blast radius is moderate and local. | `rg` shows only `ChatController` calls `executeAiOpsAnalysis` and `extractFinalReport`. | Change service/controller carefully and add tests. |
## User-Interview Confirmations
| Topic | User Words | Decision |
|---|---|---|
| Use sm-flow | "可以,改造一下AIOps 接口,用sm-flow流程看看" | Use OpenSpec + devflow. |
| GitNexus | "跳过gitnexus把" | Record skip and use local impact analysis. |
| Proceed after Grill | "可以" | Continue with lightweight Grill conclusions. |
## Key Decisions
- Keep `/api/ai_ops` and make its body optional.
- Emit `SseMessage.type=session` before long-running analysis starts.
- Store AIOps request summary in `diagnosis_session.query`.
- Store final report in `diagnosis_session.answer`.
- Defer AIOps Verifier to a later change so this slice stays focused.
## Architecture Audit
```text
POST /api/ai_ops
-> optional AIOpsRequest
-> resolve sessionId
-> create diagnosis_session(agentFlow=AI_OPS)
-> run ai_ops_supervisor(planner, executor)
-> AgentLoggingHook persists steps
-> tools persist invocations under SessionContextHolder
-> extract final report
-> persist answer
-> GET /api/diagnosis/{sessionId}/trace replays the run
```
Risk level: medium. The endpoint is public and SSE-based, but the change is additive and does not change the chat diagnosis path or database schema.
@@ -0,0 +1,37 @@
# Evidence: aiops-traceable-diagnosis-entry
## Local Impact Analysis
- `ChatController.aiOps()` is the only caller of `AiOpsService.executeAiOpsAnalysis(...)`.
- `ChatController.aiOps()` is the only caller of `AiOpsService.extractFinalReport(...)`.
- `AIOpsRequest` exists but only has `userRequest`; no current controller consumes it.
- `DiagnosisTraceService` is flow-agnostic and reads persisted session/step/tool records by `sessionId`.
## GitNexus
GitNexus MCP tools were not exposed in this session. The user explicitly approved skipping GitNexus for this change. Local impact analysis and targeted tests are used instead.
## Expected Verification
- Focused unit tests for AIOps request/session/report helper behavior.
- Compile verification.
- OpenSpec validation if CLI is available.
## Verification Results
- `openspec.cmd validate aiops-traceable-diagnosis-entry --strict`: passed.
- `mvn -q "-Dtest=AiOpsServiceTest,DiagnosisTraceServiceTest" test`: passed after rerun with approved Maven access.
- `mvn -q -DskipTests compile`: passed.
## Demo Alignment
- Added `knowledge_base/troubleshooting/aiops-alert-runbook.md` so mock AIOps alerts have matching knowledge-base guidance.
- Aligned the documented AIOps demo with mock data: `HighCPUUsage` on `payment-service`, using `system-metrics` evidence.
## Metric Alignment Follow-up
- Runtime verification showed `diagnosis_session.tool_call_count` counted agent steps with tool calls, while trace returned actual `tool_invocation` records.
- Updated `ChatService` and `AiOpsService` metric backfill to use `ToolInvocationRepository.countBySessionId(sessionId)`.
- Targeted verification:
- `mvn -q "-Dtest=AiOpsServiceTest,ChatServiceSequentialAgentTest,DiagnosisTraceServiceTest" test`: passed.
- `mvn -q -DskipTests compile`: passed.
@@ -0,0 +1,36 @@
# Acceptance: diagnosis-eval-harness
## Classification
standard-light
## Task Status
| Task | Status | Notes |
| --- | --- | --- |
| Issue and OpenSpec setup | Done | `ISS-006` and initial OpenSpec artifacts were created. |
| Implementation | Done | Added fixed cases, fixture-mode trace evaluation, aggregate metrics, and JSON / Markdown report writer. |
| Verification | Done | Targeted evaluator tests, compile verification, and OpenSpec validation passed. |
## Current State
- First implementation uses fixture-mode evaluation.
- Live trace API polling remains a follow-up option.
## Verification
### Script Verification
- Command: `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
- Result: passed
- Notes: Covers fixed case loading, fixture evaluation, missing fixture reporting, reject degraded-output validation, and report writing.
### Static Verification
- Command: `mvn -q -DskipTests compile`
- Result: passed
### OpenSpec Verification
- Command: `openspec validate diagnosis-eval-harness --strict`
- Result: passed
@@ -0,0 +1,31 @@
# Brief: diagnosis-eval-harness
## Background
The MVP has a runnable demo and hardened evidence trace semantics, but it still lacks a fixed regression baseline for Agent diagnosis quality. P1-B creates a small evaluation harness that can validate diagnosis traces against fixed cases and produce repeatable reports.
## Goals
1. Define fixed diagnosis cases for the MVP demo domain.
2. Validate trace evidence, verifier verdicts, answer keywords, and degraded-output behavior.
3. Produce JSON and Markdown reports for interview and regression use.
4. Keep the first version offline by supporting trace fixtures.
## Scope
- Evaluation case definitions
- Trace fixture shape
- Rule-based evaluator
- JSON / Markdown report output
- Focused offline tests and docs
## Non-Goals
- No LLM-as-judge
- No live end-to-end runtime requirement
- No production API
- No chat or verifier runtime change
## Related OpenSpec
`openspec/changes/diagnosis-eval-harness/`
@@ -0,0 +1,28 @@
# Diagnosis Eval Harness Decisions
## Clarify
- Entry summary: build P1-B fixed case evaluation after evidence trace hardening.
- Slug: `diagnosis-eval-harness`
- Devflow scale: standard-light
## Context
- P1-A `evidence-trace-hardening` created stable evidence semantics for supported, no-evidence, deduped, and failed tool calls.
- The MVP demo trace API already provides an aggregate trace shape suitable for evaluation.
- The first evaluator should avoid depending on external infrastructure so it can run in regular development.
## Key Decisions
- Decision: Start with rule-based trace validation instead of LLM-as-judge.
- Reason: The first regression signal should be deterministic and tied to trace contracts.
- Decision: Support offline fixture traces first.
- Reason: This makes the harness usable without MySQL, Redis, Milvus, or a real LLM.
- Decision: Output both JSON and Markdown.
- Reason: JSON supports automation; Markdown is easier to discuss in interviews.
## Open Questions
- Whether live trace API polling belongs in this change or a follow-up after fixture mode lands.
@@ -0,0 +1,10 @@
# Diagnosis Eval Harness Evidence
## Evidence
| Source | Evidence | Conclusion | Reported |
|---|---|---|---|
| `openspec/specs/evidence-trace-hardening/spec.md` | Defines stable evidence states and summary behavior | Evaluation can rely on trace semantics rather than ad hoc log parsing | Yes |
| `mvp/demo/README.md` | Documents an end-to-end demo flow with chat, trace, and feedback | Existing demo flow provides the runtime story, but not a reusable evaluation baseline | Yes |
| `DiagnosisTraceService` | Aggregates session, steps, tools, and self-evaluation | Trace response shape can be reused as evaluation input | Yes |
| `ToolTraceSummaryService` | Builds verifier-facing evidence summaries from persisted tool rows | Evaluator can check evidence coverage through persisted trace artifacts | Yes |
@@ -0,0 +1,32 @@
# Acceptance: evidence-trace-hardening
## Classification
standard-light
## Task Status
| Task | Status | Notes |
| --- | --- | --- |
| Issue and OpenSpec setup | Done | `ISS-005` and the initial OpenSpec artifacts were created. |
| Implementation | Done | Recorder contract, lookup persistence path, evidence summary semantics, and degraded-path tests were implemented. |
| Verification | Done | Targeted offline tests and compile verification passed. |
## Verification
### Script Verification
- Command: `mvn -q "-Dtest=ToolInvocationRecorderTest,ToolTraceSummaryServiceTest,ChatServiceSequentialAgentTest,LookupKnowledgeToolTest" test`
- Result: passed
- Notes: Covers recorder contract, summary semantics for success/failure/no-evidence, and `ChatService` fallback / degraded paths.
### Static Verification
- Command: `mvn -q -DskipTests compile`
- Result: passed
## Open Questions
| Question | Current position |
| --- | --- |
| Should deduped retrievals be counted separately from generic no-hit events in future evaluation metrics? | Deferred to P1-B; this change preserves enough structure to decide later. |
@@ -0,0 +1,32 @@
# Brief: evidence-trace-hardening
## Background
The MVP already has persisted tool traces and a verifier, but the evidence contract is still only partially standardized. For interview-focused hardening, the project now needs a tighter contract for evidence persistence, no-evidence / failure semantics, and degraded-output behavior.
## Goals
1. Standardize the persisted evidence-tool contract across `lookup_knowledge`, `query_logs`, and `query_metrics`.
2. Make verifier-facing summaries distinguish failed calls, no-hit calls, deduped retrievals, and actual supporting evidence.
3. Add offline tests for verifier fallback and degraded-output paths.
## Scope
- `ToolInvocationRecorder`
- `LookupKnowledgeTool`
- `QueryLogsTools`
- `QueryMetricsTools`
- `ToolTraceSummaryService`
- `ChatService`
- Focused offline tests
## Non-Goals
- No new API or schema
- No evaluation harness yet
- No trace UI
- No security/config cleanup
## Related OpenSpec
`openspec/changes/evidence-trace-hardening/`
@@ -0,0 +1,24 @@
# Evidence Trace Hardening Decisions
## Clarify
- Entry summary: harden the MVP evidence contract before building the P1-B evaluation harness.
- Slug: `evidence-trace-hardening`
- Devflow scale: standard-light
## Context
- `ISS-003` raised verifier traceability and failure-path concerns.
- Current code inspection shows `QueryLogsTools` and `QueryMetricsTools` already use `ToolInvocationRecorder`, while `LookupKnowledgeTool` still persists rows through a local helper.
- `ChatService` already contains fallback behavior for missing/invalid `verifier_output`, but coverage is narrow.
## Key Decisions
- Decision: Treat this as a contract-hardening change, not a new feature change.
- Reason: The project already has the necessary runtime pieces; the gap is semantic consistency and testability.
- Decision: Keep the scope before P1-B.
- Reason: The evaluation harness will rely on stable evidence semantics, so this contract slice should land first.
- Decision: Preserve schema and API stability.
- Reason: The interview value here is engineering rigor, not more surface area.
@@ -0,0 +1,11 @@
# Evidence Trace Hardening Evidence
## Evidence
| Source | Evidence | Conclusion | Reported |
|---|---|---|---|
| `ToolInvocationRecorder` | Provides a common persistence seam for evidence tools | Contract hardening should build on the existing recorder instead of introducing a new store path | Yes |
| `LookupKnowledgeTool` | Still constructs `ToolInvocation` rows through a local helper | Retrieval-aware evidence persistence is not yet unified with the recorder contract | Yes |
| `QueryLogsTools` / `QueryMetricsTools` | Already record evidence invocations through `recordEvidenceTool(...)` | Current gap is semantic alignment, not missing persistence | Yes |
| `ToolTraceSummaryService` | Merges rows by tool and topic domain and infers evidence level heuristically | Summary rules need explicit handling for failure, no-hit, and dedup cases | Yes |
| `ChatService` | Falls back to `LOW_CONFID` when verifier output is missing or invalid | These degraded paths exist and should now be covered by focused offline tests | Yes |
@@ -0,0 +1,37 @@
# Acceptance: expand-diagnosis-eval-fixtures
## Classification
standard-light
## Task Status
| Task | Status | Notes |
| --- | --- | --- |
| Issue and OpenSpec setup | Done | Created slug-based issue and initial OpenSpec artifacts. |
| Implementation | Done | Added remaining fixtures, full baseline reports, and documentation updates. |
| Verification | Done | Evaluator tests, compile verification, and OpenSpec validation passed. |
## Current State
- Fixture coverage is complete for the five fixed diagnosis cases.
- Baseline reports are saved under `mvp/eval/reports`.
- No production runtime behavior has been changed.
## Verification
### Script Verification
- Command: `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
- Result: passed
- Notes: Covers full fixture coverage, baseline report matching, reject degraded-output validation, and report writing.
### Static Verification
- Command: `mvn -q -DskipTests compile`
- Result: passed
### OpenSpec Verification
- Command: `openspec validate expand-diagnosis-eval-fixtures --strict`
- Result: passed
@@ -0,0 +1,31 @@
# Brief: expand-diagnosis-eval-fixtures
## Background
The diagnosis eval harness is implemented and archived, but the fixed baseline is incomplete because three of the five diagnosis cases still reference missing fixtures.
## Goals
1. Add representative trace fixtures for all remaining fixed diagnosis cases.
2. Save a reproducible baseline report in JSON and Markdown.
3. Document how to regenerate and interpret the baseline.
4. Keep evaluation offline and deterministic.
## Scope
- Redis timeout fixture
- Slow response fixture
- JVM memory risk fixture
- Baseline reports under `mvp/eval/reports`
- Focused tests for full fixture coverage and report generation
## Non-Goals
- No new diagnosis cases
- No production Agent runtime changes
- No LLM-as-judge
- No live infrastructure requirement
## Related OpenSpec
`openspec/changes/expand-diagnosis-eval-fixtures/`
@@ -0,0 +1,28 @@
# Expand Diagnosis Eval Fixtures Decisions
## Clarify
- Entry summary: complete the fixed diagnosis eval baseline after the harness is in place.
- Slug: `expand-diagnosis-eval-fixtures`
- Devflow scale: standard-light
## Context
- `diagnosis-eval-harness` created the evaluator, case file, fixture mode, and report writer.
- The first baseline still has missing fixtures by design.
- This follow-up turns that partial baseline into a full fixed-case baseline.
## Key Decisions
- Decision: Keep this change data-focused.
- Reason: the evaluator rules already landed; this change should not blur fixture expansion with harness behavior changes.
- Decision: Save baseline reports in the repository.
- Reason: interview review and future diffs are easier when the expected baseline is visible.
- Decision: Use deterministic fixture traces instead of live trace generation.
- Reason: this baseline should run without infrastructure or external model calls.
## Open Questions
- Whether a future change should add a CLI or Maven goal for report regeneration.
@@ -0,0 +1,11 @@
# Evidence: expand-diagnosis-eval-fixtures
## Evidence Log
- 2026-07-04: Created slug-based issue `expand-diagnosis-eval-fixtures.md`.
- 2026-07-04: Created OpenSpec change `expand-diagnosis-eval-fixtures`.
- 2026-07-04: Added Redis timeout, slow response, and JVM memory risk fixtures.
- 2026-07-04: Added baseline JSON and Markdown reports under `mvp/eval/reports`.
- 2026-07-04: Verification passed with `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`.
- 2026-07-04: Verification passed with `mvn -q -DskipTests compile`.
- 2026-07-04: Verification passed with `openspec validate expand-diagnosis-eval-fixtures --strict`.
@@ -0,0 +1,37 @@
# Acceptance: diagnosis-eval-baseline-diff
## Classification
standard-light
## Task Status
| Task | Status | Notes |
| --- | --- | --- |
| Issue and OpenSpec setup | Done | Created slug-based issue and initial OpenSpec artifacts. |
| Implementation | Done | Added diff model, comparator, writer, docs, sample outputs, and focused tests. |
| Verification | Done | Diff/evaluator tests, compile verification, and OpenSpec validation passed. |
## Current State
- Baseline diff is implemented for aggregate metrics, verdict distribution, case-level state, keyword coverage, evidence coverage, missing cases, and new cases.
- JSON and Markdown diff output are available.
- No production runtime behavior has been changed.
## Verification
### Script Verification
- Command: `mvn -q "-Dtest=DiagnosisEvalBaselineDiffTest" test`
- Result: passed
- Notes: Also verified with `DiagnosisTraceEvaluatorTest`.
### Static Verification
- Command: `mvn -q -DskipTests compile`
- Result: passed
### OpenSpec Verification
- Command: `openspec validate diagnosis-eval-baseline-diff --strict`
- Result: passed
@@ -0,0 +1,30 @@
# Brief: diagnosis-eval-baseline-diff
## Background
The eval harness now has a complete saved baseline. This change adds the comparison layer that turns the baseline into an actionable regression signal.
## Goals
1. Compare baseline and current `DiagnosisEvalReport` objects.
2. Detect aggregate and per-case regressions.
3. Output JSON and Markdown diff reports.
4. Document how to read the diff in interview and engineering terms.
## Scope
- Diff data structures
- Deterministic report comparison
- JSON / Markdown diff output
- Focused tests and eval docs
## Non-Goals
- No live Agent execution
- No LLM-as-judge
- No evaluator scoring rule changes
- No production API changes
## Related OpenSpec
`openspec/changes/diagnosis-eval-baseline-diff/`
@@ -0,0 +1,28 @@
# Diagnosis Eval Baseline Diff Decisions
## Clarify
- Entry summary: add report diffing on top of the completed diagnosis eval baseline.
- Slug: `diagnosis-eval-baseline-diff`
- Devflow scale: standard-light
## Context
- `diagnosis-eval-harness` created deterministic fixture evaluation.
- `expand-diagnosis-eval-fixtures` created a complete saved baseline.
- This change compares new reports against that baseline.
## Key Decisions
- Decision: Diff report DTOs instead of raw traces.
- Reason: the report is the stable contract for regression review.
- Decision: Use deterministic code rules instead of LLM-as-judge.
- Reason: baseline regression checks should be repeatable and explainable.
- Decision: Output both JSON and Markdown.
- Reason: JSON supports automation; Markdown is useful in reviews and interviews.
## Open Questions
- Whether a future change should expose this through a CLI or Maven goal.
@@ -0,0 +1,11 @@
# Evidence: diagnosis-eval-baseline-diff
## Evidence Log
- 2026-07-05: Created slug-based issue `diagnosis-eval-baseline-diff.md`.
- 2026-07-05: Created OpenSpec change `diagnosis-eval-baseline-diff`.
- 2026-07-05: Added baseline diff DTOs, deterministic comparer, and JSON / Markdown writer.
- 2026-07-05: Added sample baseline diff JSON and Markdown reports.
- 2026-07-05: Verification passed with `mvn -q "-Dtest=DiagnosisEvalBaselineDiffTest,DiagnosisTraceEvaluatorTest" test`.
- 2026-07-05: Verification passed with `mvn -q -DskipTests compile`.
- 2026-07-05: Verification passed with `openspec validate diagnosis-eval-baseline-diff --strict`.
@@ -0,0 +1,25 @@
# Acceptance: mvp-demo-interview-runbook
## Classification
standard-light
## Task Status
| Task | Status | Notes |
| --- | --- | --- |
| Issue and OpenSpec setup | Done | Created slug-based issue and OpenSpec artifacts. |
| Implementation | Done | Added request payload, runnable script, output directory docs, interview walkthrough, and trace checklist. |
| Verification | Done | OpenSpec validation passed. |
## Current State
- No backend runtime behavior has been changed.
- Demo is packaged under `mvp/demo` for interview use.
## Verification
### OpenSpec Verification
- Command: `openspec validate mvp-demo-interview-runbook --strict`
- Result: passed
@@ -0,0 +1,29 @@
# Brief: mvp-demo-interview-runbook
## Background
Plan C is the interview-facing demo package. The project has the engineering pieces, but needs a single place to run and explain the MVP flow.
## Goals
1. Provide a fixed payment-timeout request payload.
2. Provide a PowerShell script that runs chat, trace, and feedback.
3. Save demo responses under `mvp/demo/output`.
4. Add interview walkthrough and trace checklist.
## Scope
- Demo docs and scripts only
- Existing local APIs only
- Existing `mvp-demo` profile only
## Non-Goals
- No backend code changes
- No eval extension
- No secret cleanup
- No full offline runtime
## Related OpenSpec
`openspec/changes/mvp-demo-interview-runbook/`
@@ -0,0 +1,27 @@
# MVP Demo Interview Runbook Decisions
## Clarify
- Entry summary: package existing MVP capabilities into a repeatable interview demo.
- Slug: `mvp-demo-interview-runbook`
- Devflow scale: standard-light
## Context
- Evidence trace and eval baseline work are already done.
- The next useful step is not more eval tooling, but a runnable demo path.
## Key Decisions
- Decision: Keep this change documentation/script-only.
- Reason: Plan C is about demo packaging, not new runtime capability.
- Decision: Use a stable session id.
- Reason: it makes trace lookup and saved output predictable.
- Decision: Save outputs to `mvp/demo/output`.
- Reason: generated artifacts should be easy to review without mixing into source fixtures.
## Open Questions
- Whether a later change should add a truly offline stubbed demo mode.
@@ -0,0 +1,9 @@
# Evidence: mvp-demo-interview-runbook
## Evidence Log
- 2026-07-05: Created Plan C demo packaging issue and OpenSpec change.
- 2026-07-05: Added fixed payment-timeout request payload.
- 2026-07-05: Added PowerShell demo script for chat, trace, and feedback.
- 2026-07-05: Added interview walkthrough and trace inspection checklist.
- 2026-07-05: Verification passed with `openspec validate mvp-demo-interview-runbook --strict`.
+86
View File
@@ -0,0 +1,86 @@
# RAG Retrieval Baseline
This directory contains the offline retrieval baseline for the RAG refactor.
The baseline is intentionally narrower than full diagnosis evaluation. It checks
whether fixed retrieval queries can recover expected documents, breadcrumbs, and
evidence keywords before changing L0 behavior, query augmentation, evidence
post-processing, or Spring AI VectorStore integration.
## Layout
```text
eval/rag-retrieval/
cases/golden-cases.json Fixed retrieval golden cases
fixtures/*.json Saved retrieval candidates for each case
reports/baseline.json Machine-readable baseline report
reports/baseline.md Human-readable baseline report
reports/live-post-reindex.* Optional live acceptance reports
```
## Run
From the repository root:
```bash
python scripts/eval_rag_retrieval.py
```
Custom paths are also supported:
```bash
python scripts/eval_rag_retrieval.py \
--cases eval/rag-retrieval/cases/golden-cases.json \
--fixtures eval/rag-retrieval/fixtures \
--json-report eval/rag-retrieval/reports/baseline.json \
--markdown-report eval/rag-retrieval/reports/baseline.md
```
## Hit Levels
- `strong`: expected document is found and breadcrumb or evidence keyword coverage is satisfied.
- `medium`: expected document is found, but breadcrumb or keyword coverage is incomplete.
- `weak`: expected evidence keyword is found, but expected document is missing.
- `miss`: expected document and expected evidence are not found.
`Recall@K` counts `strong` and `medium` as retrieved.
## Scope
This baseline runs fully offline and does not call MySQL, Redis, Milvus, an LLM,
or the Spring Boot application. It is a regression harness for retrieval behavior,
not a claim that live production retrieval accuracy is complete.
## Live Post-Reindex Acceptance
When embedding input changes, existing vectors do not update by themselves. For
example, after adding `title` and `breadcrumb` to the embedding text, the live
Milvus/Zilliz collection must be reindexed before retrieval can reflect that new
semantic signal.
Use this optional live acceptance flow after the application is running and the
knowledge base has been reindexed:
```bash
python scripts/eval_rag_live_acceptance.py
```
Custom service URL and output paths are supported:
```bash
python scripts/eval_rag_live_acceptance.py \
--base-url http://127.0.0.1:9900 \
--json-report eval/rag-retrieval/reports/live-post-reindex.json \
--markdown-report eval/rag-retrieval/reports/live-post-reindex.md
```
The script calls:
```text
GET /api/search/similar
```
It writes JSON and Markdown reports with query, topK, result count, top
candidates, breadcrumb, score labels, and raw response fields. This is a live
smoke check for environment readiness and post-reindex behavior; it does not
replace the deterministic offline baseline above.
@@ -0,0 +1,61 @@
{
"version": 1,
"description": "Offline golden retrieval cases for RAG refactor baseline.",
"topK": 5,
"cases": [
{
"caseId": "chat-mysql-connection-pool",
"scenario": "chat",
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
"expectedDocIds": ["mysql-connection-pool"],
"expectedBreadcrumbs": ["Database > MySQL > Connection Pool"],
"expectedKeywords": ["connection pool", "max_connections", "HikariCP"],
"notes": "Covers precise database troubleshooting retrieval."
},
{
"caseId": "chat-diagnosis-flow",
"scenario": "chat",
"query": "What is the standard troubleshooting flow for an application incident?",
"expectedDocIds": ["incident-diagnosis-flow"],
"expectedBreadcrumbs": ["AIOps > Diagnosis Flow"],
"expectedKeywords": ["collect evidence", "verify", "remediation"],
"notes": "Covers process-style knowledge where breadcrumb matters."
},
{
"caseId": "aiops-payment-latency-alert",
"scenario": "aiops",
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
"expectedDocIds": ["payment-service-latency"],
"expectedBreadcrumbs": ["AIOps > Service Alerts > Payment Latency"],
"expectedKeywords": ["p95 latency", "payment-service", "downstream dependency"],
"notes": "Covers alert payload terms that should become retrieval hints."
},
{
"caseId": "aiops-prometheus-alert-scope",
"scenario": "aiops",
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
"expectedDocIds": ["aiops-alert-scope-control"],
"expectedBreadcrumbs": ["AIOps > Alert Scope Control"],
"expectedKeywords": ["payload", "unrelated active alerts", "scope"],
"notes": "Covers scoped alert diagnosis behavior."
},
{
"caseId": "chat-rag-chunk-context",
"scenario": "chat",
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
"expectedDocIds": ["rag-chunk-context-reconstruction"],
"expectedBreadcrumbs": ["RAG > Chunking > Context Reconstruction"],
"expectedKeywords": ["neighbor chunk", "same section", "breadcrumb"],
"notes": "Covers the known RAG refactor issue around context reconstruction."
},
{
"caseId": "chat-l0-domain-hint",
"scenario": "chat",
"query": "Should L0 keyword matching decide the final retrieval result?",
"expectedDocIds": ["rag-l0-domain-entity-hint"],
"expectedBreadcrumbs": ["RAG > L0 > Domain Entity Hint"],
"expectedKeywords": ["domain detector", "entity extractor", "metadata filter"],
"notes": "Covers the target L0 role after refactor."
}
]
}
@@ -0,0 +1,25 @@
{
"caseId": "aiops-payment-latency-alert",
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
"retrievedAt": "2026-07-05T00:00:00Z",
"candidates": [
{
"rank": 1,
"docId": "payment-service-latency",
"title": "Payment Service Latency Alert Playbook",
"breadcrumb": "AIOps > Service Alerts > Payment Latency",
"content": "For payment-service p95 latency alerts, check downstream dependency latency, thread pool saturation, gateway retries, and recent deployment changes.",
"score": 0.84,
"retrievalLayer": "L1"
},
{
"rank": 2,
"docId": "mysql-connection-pool",
"title": "MySQL Connection Pool Troubleshooting",
"breadcrumb": "Database > MySQL > Connection Pool",
"content": "Database connection pool saturation can increase payment latency when checkout paths wait for connections.",
"score": 0.68,
"retrievalLayer": "L1"
}
]
}
@@ -0,0 +1,16 @@
{
"caseId": "aiops-prometheus-alert-scope",
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
"retrievedAt": "2026-07-05T00:00:00Z",
"candidates": [
{
"rank": 1,
"docId": "aiops-alert-scope-control",
"title": "AIOps Alert Scope Control",
"breadcrumb": "AIOps > Alert Scope Control",
"content": "When payload mode is active, queryPrometheusAlerts can verify the supplied alert, but unrelated active alerts must remain scoped context and should not become full diagnoses.",
"score": 0.9,
"retrievalLayer": "L0+L1"
}
]
}
@@ -0,0 +1,25 @@
{
"caseId": "chat-diagnosis-flow",
"query": "What is the standard troubleshooting flow for an application incident?",
"retrievedAt": "2026-07-05T00:00:00Z",
"candidates": [
{
"rank": 1,
"docId": "incident-diagnosis-flow",
"title": "Incident Diagnosis Flow",
"breadcrumb": "AIOps > Diagnosis Flow",
"content": "The standard flow is to collect evidence, identify the suspected fault domain, verify the hypothesis, apply remediation, and confirm recovery.",
"score": 0.82,
"retrievalLayer": "L1"
},
{
"rank": 2,
"docId": "rag-chunk-context-reconstruction",
"title": "RAG Chunk Context Reconstruction",
"breadcrumb": "RAG > Chunking > Context Reconstruction",
"content": "Long sections may require neighbor chunk expansion and breadcrumb-aware packing.",
"score": 0.55,
"retrievalLayer": "L1"
}
]
}
@@ -0,0 +1,25 @@
{
"caseId": "chat-l0-domain-hint",
"query": "Should L0 keyword matching decide the final retrieval result?",
"retrievedAt": "2026-07-05T00:00:00Z",
"candidates": [
{
"rank": 1,
"docId": "rag-l0-domain-entity-hint",
"title": "RAG L0 Domain Entity Hint",
"breadcrumb": "RAG > L0 > Domain Entity Hint",
"content": "L0 should be retained as a domain detector, entity extractor, metadata filter generator, and explainability signal, not as the final retrieval decision.",
"score": 0.88,
"retrievalLayer": "L0"
},
{
"rank": 2,
"docId": "rag-l0-l1-fusion-ranking",
"title": "RAG L0 L1 Fusion Ranking",
"breadcrumb": "RAG > Ranking > Fusion",
"content": "L0 and L1 candidates should eventually be fused rather than handled as an early-return branch.",
"score": 0.75,
"retrievalLayer": "L1"
}
]
}
@@ -0,0 +1,25 @@
{
"caseId": "chat-mysql-connection-pool",
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
"retrievedAt": "2026-07-05T00:00:00Z",
"candidates": [
{
"rank": 1,
"docId": "mysql-connection-pool",
"title": "MySQL Connection Pool Troubleshooting",
"breadcrumb": "Database > MySQL > Connection Pool",
"content": "When the connection pool is exhausted, inspect HikariCP active connections, max_connections, slow SQL, leak detection, and database wait events.",
"score": 0.86,
"retrievalLayer": "L0+L1"
},
{
"rank": 2,
"docId": "incident-diagnosis-flow",
"title": "Incident Diagnosis Flow",
"breadcrumb": "AIOps > Diagnosis Flow",
"content": "Collect evidence, compare metrics and logs, then verify remediation before closing the incident.",
"score": 0.61,
"retrievalLayer": "L1"
}
]
}
@@ -0,0 +1,25 @@
{
"caseId": "chat-rag-chunk-context",
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
"retrievedAt": "2026-07-05T00:00:00Z",
"candidates": [
{
"rank": 1,
"docId": "rag-chunk-context-reconstruction",
"title": "RAG Chunk Context Reconstruction",
"breadcrumb": "RAG > Chunking > Context Reconstruction",
"content": "After a chunk hit, expand to neighbor chunk candidates from the same section and preserve breadcrumb metadata in the evidence pack.",
"score": 0.79,
"retrievalLayer": "L1"
},
{
"rank": 2,
"docId": "rag-breadcrumb-embedding-gap",
"title": "RAG Breadcrumb Embedding Gap",
"breadcrumb": "RAG > Embedding > Breadcrumb",
"content": "Embedding title and breadcrumb with content helps recover section semantics.",
"score": 0.72,
"retrievalLayer": "L1"
}
]
}
+1
View File
@@ -0,0 +1 @@
+131
View File
@@ -0,0 +1,131 @@
{
"generatedAt": "2026-07-04T17:59:52.172759+00:00",
"caseFile": "eval/rag-retrieval/cases/golden-cases.json",
"fixtureDir": "eval/rag-retrieval/fixtures",
"aggregate": {
"caseCount": 6,
"topK": 5,
"strongHitCount": 6,
"mediumHitCount": 0,
"weakHitCount": 0,
"missCount": 0,
"recallAtK": 1.0,
"strongHitRate": 1.0,
"averageFirstHitRank": 1.0
},
"results": [
{
"caseId": "chat-mysql-connection-pool",
"scenario": "chat",
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
"hitLevel": "strong",
"passed": true,
"firstExpectedRank": 1,
"topCandidates": [
"1:mysql-connection-pool",
"2:incident-diagnosis-flow"
],
"matchedKeywords": [
"connection pool",
"max_connections",
"hikaricp"
],
"breadcrumbMatched": true,
"failedChecks": []
},
{
"caseId": "chat-diagnosis-flow",
"scenario": "chat",
"query": "What is the standard troubleshooting flow for an application incident?",
"hitLevel": "strong",
"passed": true,
"firstExpectedRank": 1,
"topCandidates": [
"1:incident-diagnosis-flow",
"2:rag-chunk-context-reconstruction"
],
"matchedKeywords": [
"collect evidence",
"verify",
"remediation"
],
"breadcrumbMatched": true,
"failedChecks": []
},
{
"caseId": "aiops-payment-latency-alert",
"scenario": "aiops",
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
"hitLevel": "strong",
"passed": true,
"firstExpectedRank": 1,
"topCandidates": [
"1:payment-service-latency",
"2:mysql-connection-pool"
],
"matchedKeywords": [
"p95 latency",
"payment-service",
"downstream dependency"
],
"breadcrumbMatched": true,
"failedChecks": []
},
{
"caseId": "aiops-prometheus-alert-scope",
"scenario": "aiops",
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
"hitLevel": "strong",
"passed": true,
"firstExpectedRank": 1,
"topCandidates": [
"1:aiops-alert-scope-control"
],
"matchedKeywords": [
"payload",
"unrelated active alerts",
"scope"
],
"breadcrumbMatched": true,
"failedChecks": []
},
{
"caseId": "chat-rag-chunk-context",
"scenario": "chat",
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
"hitLevel": "strong",
"passed": true,
"firstExpectedRank": 1,
"topCandidates": [
"1:rag-chunk-context-reconstruction",
"2:rag-breadcrumb-embedding-gap"
],
"matchedKeywords": [
"neighbor chunk",
"same section",
"breadcrumb"
],
"breadcrumbMatched": true,
"failedChecks": []
},
{
"caseId": "chat-l0-domain-hint",
"scenario": "chat",
"query": "Should L0 keyword matching decide the final retrieval result?",
"hitLevel": "strong",
"passed": true,
"firstExpectedRank": 1,
"topCandidates": [
"1:rag-l0-domain-entity-hint",
"2:rag-l0-l1-fusion-ranking"
],
"matchedKeywords": [
"domain detector",
"entity extractor",
"metadata filter"
],
"breadcrumbMatched": true,
"failedChecks": []
}
]
}
+28
View File
@@ -0,0 +1,28 @@
# RAG Retrieval Baseline
Generated at: `2026-07-04T17:59:52.172759+00:00`
## Aggregate
| Metric | Value |
|---|---:|
| Cases | 6 |
| Top K | 5 |
| Recall@K | 1.0 |
| Strong hit rate | 1.0 |
| Strong hits | 6 |
| Medium hits | 0 |
| Weak hits | 0 |
| Misses | 0 |
| Average first hit rank | 1.0 |
## Cases
| Case | Scenario | Hit | First Expected Rank | Top Candidates | Failed Checks |
|---|---|---|---:|---|---|
| chat-mysql-connection-pool | chat | strong | 1 | 1:mysql-connection-pool<br>2:incident-diagnosis-flow | |
| chat-diagnosis-flow | chat | strong | 1 | 1:incident-diagnosis-flow<br>2:rag-chunk-context-reconstruction | |
| aiops-payment-latency-alert | aiops | strong | 1 | 1:payment-service-latency<br>2:mysql-connection-pool | |
| aiops-prometheus-alert-scope | aiops | strong | 1 | 1:aiops-alert-scope-control | |
| chat-rag-chunk-context | chat | strong | 1 | 1:rag-chunk-context-reconstruction<br>2:rag-breadcrumb-embedding-gap | |
| chat-l0-domain-hint | chat | strong | 1 | 1:rag-l0-domain-entity-hint<br>2:rag-l0-l1-fusion-ranking | |
+59
View File
@@ -0,0 +1,59 @@
# SuperBizAgent Interview Guide
## 一句话定位
SuperBizAgent 是一个面向企业故障诊断场景的 Agent Engineering 项目:它把用户问题或告警事件转成可追踪的多 Agent 执行链路,并把工具证据、模型步骤、最终答案和反馈统一落到诊断 trace 中。
## 面试重点
- **多 Agent 编排**:普通 Chat 的复杂问题走 `Planner -> Executor -> Verifier`;AIOps 告警入口走 Supervisor 调度 Planner/Executor。
- **工具证据链**:知识库、日志、指标和 Prometheus 告警都通过工具调用进入链路,并记录到 `tool_invocation`。
- **可追踪诊断**:一次会话对应一个 `sessionId`,最终可以通过 `GET /api/diagnosis/{sessionId}/trace` 回放。
- **质量门**:Chat 链路包含 Verifier,把 groundedness、facts checked 和 evidence refs 写回 `diagnosis_session.self_evaluation`。
- **AIOps 产品边界**:有告警 payload 时聚焦该告警;没有 payload 时先自动发现 active alerts。
- **可复现 Demo**:`mvp-demo` profile 使用 mock Prometheus 和 mock CLS,让面试演示不依赖真实线上故障。
## 推荐阅读顺序
1. `interview/demo-script.md`:面试现场怎么讲、怎么演示。
2. `interview/architecture.md`:系统架构和两条主链路。
3. `interview/design-tradeoffs.md`:关键设计取舍和可被追问的问题。
4. `interview/acceptance-checklist.md`:面试前验证清单。
5. `mvp/demo/README.md`:更细的 MVP 可执行 runbook。
## 核心 Demo
### Chat Diagnosis
```text
POST /api/chat
-> ChatService.executeChatWithStrategy(...)
-> simple ReactAgent or Planner -> Executor -> Verifier
-> lookup_knowledge / query_logs / query_metrics
-> diagnosis_session + agent_step + tool_invocation
-> GET /api/diagnosis/{sessionId}/trace
```
### AIOps Alert Diagnosis
```text
POST /api/ai_ops
-> AiOpsService.executeAiOpsAnalysis(...)
-> ai_ops_supervisor
-> planner_agent / executor_agent
-> queryPrometheusAlerts + logs + knowledge
-> scoped alert report
-> GET /api/diagnosis/{sessionId}/trace
```
## 当前完成度
- Chat 诊断链路:可运行、可追踪、有 Verifier。
- AIOps 告警链路:可运行、可追踪、支持 payload scope control。
- Trace API:统一返回 session、agent steps、tool invocations 和 summary。
- Demo 文档:`mvp/demo/README.md` 和 `mvp/demo/aiops-alert-acceptance.md`。
- Devflow 沉淀:`devflow/index.md` 记录了 MVP、Verifier、AIOps trace 和 AIOps scope-control 的演进。
## 面试时的主叙事
这个项目不是简单调用大模型,而是在做一个可审计的 Agent 诊断系统。核心价值是:模型可以规划和推理,但每一步工具证据、最终结论和质量评估都能被 trace API 回放。面试时重点展示“从问题到证据到答案到验证”的完整闭环。
+170
View File
@@ -0,0 +1,170 @@
# Acceptance Checklist
## 面试前环境检查
- 当前分支包含最新 AIOps trace/scope 变更。
- MySQL 可连接。
- Redis 可连接。
- Milvus/Zilliz 可连接。
- 模型 API key 可用。
- `mvp-demo` profile 开启 mock Prometheus 和 mock CLS。
启动:
```powershell
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
```
编译检查:
```powershell
mvn -q -DskipTests compile
```
目标测试:
```powershell
mvn -q "-Dtest=AiOpsServiceTest,ChatServiceSequentialAgentTest,DiagnosisTraceServiceTest" test
```
## Chat Demo 验收
请求:
```powershell
$sessionId = "interview-chat-payment-timeout-001"
$body = @{
Id = $sessionId
Question = "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
} | ConvertTo-Json
Invoke-RestMethod `
-Method Post `
-Uri "http://localhost:9900/api/chat" `
-ContentType "application/json" `
-Body $body
```
验收:
- 返回 `data.success = true`。
- 返回 `data.sessionId = interview-chat-payment-timeout-001`。
- `diagnosis_session.agent_flow = CHAT`。
- trace API 返回 session、steps、toolInvocations。
- 复杂问题下 trace 中能看到 verifier 相关数据。
SQL:
```powershell
python scripts/query_mysql.py "SELECT session_id, agent_flow, status, step_count, tool_call_count FROM diagnosis_session WHERE session_id='interview-chat-payment-timeout-001'"
```
## AIOps Demo 验收
请求:
```powershell
$aiopsSessionId = "interview-aiops-payment-cpu-001"
$aiopsBody = @{
sessionId = $aiopsSessionId
alertName = "HighCPUUsage"
service = "payment-service"
severity = "P1"
description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
timeRange = "last_15m"
userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
} | ConvertTo-Json
Invoke-WebRequest `
-Method Post `
-Uri "http://localhost:9900/api/ai_ops" `
-ContentType "application/json" `
-Body $aiopsBody
```
验收:
- SSE 首条包含 `type=session`。
- SSE 最后包含 `type=done`。
- `diagnosis_session.agent_flow = AI_OPS`。
- `diagnosis_session.status = SUCCESS`。
- `diagnosis_session.answer` 有最终报告。
- trace API 返回 AIOps steps 和 tool invocations。
- 报告主章节聚焦 `HighCPUUsage/payment-service`。
- 无 `告警根因分析 - HighMemoryUsage` 独立章节。
- 无 `告警根因分析 - SlowResponse` 独立章节。
- 有“相关风险告警”或类似上下文说明。
SQL:
```powershell
python scripts/query_mysql.py "SELECT session_id, agent_flow, status, total_duration_ms, step_count, tool_call_count FROM diagnosis_session WHERE session_id='interview-aiops-payment-cpu-001'"
```
```powershell
python scripts/query_mysql.py "SELECT tool_name, COUNT(*) AS cnt FROM tool_invocation WHERE session_id='interview-aiops-payment-cpu-001' GROUP BY tool_name ORDER BY tool_name"
```
Scope 检查:
```powershell
python scripts/query_mysql.py "SELECT (answer LIKE '%告警根因分析 - HighCPUUsage%') AS has_main_root_cause, (answer LIKE '%告警根因分析 - HighMemoryUsage%') AS has_memory_root_cause, (answer LIKE '%告警根因分析 - SlowResponse%') AS has_slow_root_cause, (answer LIKE '%相关风险告警%') AS has_related_risk FROM diagnosis_session WHERE session_id='interview-aiops-payment-cpu-001'"
```
期望:
```text
has_main_root_cause = 1
has_memory_root_cause = 0
has_slow_root_cause = 0
has_related_risk = 1
```
## Trace API 验收
```powershell
Invoke-RestMethod `
-Method Get `
-Uri "http://localhost:9900/api/diagnosis/interview-aiops-payment-cpu-001/trace"
```
若 PowerShell 对长 JSON 或特殊字符不稳定,可以用:
```powershell
curl.exe --silent --show-error --max-time 60 "http://localhost:9900/api/diagnosis/interview-aiops-payment-cpu-001/trace"
```
## 常见问题
### MySQL stale connection
现象:
```text
HikariPool - Connection is not available
No operations allowed after connection closed
```
当前已在 `application.yml` 配置:
- `maximum-pool-size: 5`
- `minimum-idle: 1`
- `connection-timeout: 10000`
- `validation-timeout: 5000`
- `idle-timeout: 60000`
- `max-lifetime: 120000`
- `keepalive-time: 30000`
处理:
- 重新编译或重启服务。
- 确认日志中新的 HikariPool 启动成功。
- 再跑 trace 或 AIOps 请求。
### SSE 客户端显示异常
PowerShell `Invoke-WebRequest` 有时对 SSE 或长 JSON 处理不稳定。可以改用 `curl.exe` 或直接查询 MySQL 和 trace API 验证结果。
### OpenSpec 全量校验失败
`openspec validate --all --strict` 可能因为历史未完成 change 失败。面试材料主要依赖已归档的 AIOps spec 和 MVP trace spec,可以单独验证相关 spec。
+53
View File
@@ -0,0 +1,53 @@
# AIOps Lightweight Verifier
## What Changed
AIOps now has a deterministic post-run quality gate.
After the final AIOps report is persisted, the service evaluates:
- whether the final report exists and is not trivially short
- whether a payload-targeted report mentions the supplied alert and service
- whether evidence tools such as `lookup_knowledge`, `query_metrics`, or `query_logs` were persisted
The result is stored under:
```text
diagnosis_session.self_evaluation.aiops_rule_evaluation
```
The trace API returns this payload through the existing session self-evaluation field.
## Why Rule-Based First
This is not a full LLM verifier yet.
The first AIOps quality risks are concrete and easy to check with rules:
- Did the report stay focused on the payload?
- Did the run use evidence tools?
- Did the system produce a usable final report?
Rule evaluation is stable, cheap, and easy to explain. It also avoids adding another hidden model call to the AIOps flow before the current trace contract is mature.
## Verdicts
The evaluator emits:
```text
PASS
WARN
FAIL
```
`FAIL` is reserved for critical issues such as a missing or too-short report. Missing payload focus terms or missing evidence tools currently produce `WARN`, because valid reports may use slightly different wording or evidence may be unavailable in a mock/demo environment.
## Interview Answer
If asked why AIOps has a verifier now:
> Chat already has an LLM verifier because the user questions are open-ended. For AIOps, I started with a lighter rule-based verifier because the first quality checks are very concrete: payload focus, evidence coverage, and report completeness. The evaluation is persisted into `self_evaluation`, so the trace can show not only what the Agent did, but also whether the output passed basic quality gates.
If asked why not use the Chat verifier directly:
> AIOps verification is different from Chat verification. It needs to check alert scope, evidence tool coverage, and whether unrelated active alerts were over-expanded. Reusing the Chat verifier directly would blur those semantics. The rule-based evaluator gives us a stable first quality gate; a later AIOps LLM verifier can build on the same trace contract.
+62
View File
@@ -0,0 +1,62 @@
# AIOps Query Augmentation
## What Changed
Payload-targeted AIOps prompts now include a deterministic recommended knowledge query.
The query is built from the non-blank payload fields:
```text
alertName service severity description timeRange userRequest
```
Example:
```text
HighCPUUsage payment-service P1 CPU usage is above 80% last_15m
```
## Why This Matters
AIOps payload fields contain high-value retrieval terms:
- alert name
- service name
- severity
- symptom description
- time range
- operator request
Before this change, the Agent still had to invent its own `lookup_knowledge` query from the full prompt. That can work, but it may omit important terms such as the service name or alert name.
The new prompt makes the retrieval seed explicit:
```text
Recommended lookup_knowledge query: ...
```
## Design Choice
This is prompt-level query augmentation, not hidden retrieval.
I intentionally did not call `lookup_knowledge` automatically before the Agent runs. The project values traceability: tool calls should appear as Agent actions, with their inputs and outputs recorded in `tool_invocation`.
So the design is:
```text
AIOps payload
-> deterministic recommended retrieval query
-> Agent prompt
-> Agent may call lookup_knowledge explicitly
-> tool_invocation records the real retrieval action
```
## Interview Answer
If asked how AIOps payload improves RAG retrieval:
> I do not replace the user query with a broad domain. I extract the high-signal alert terms from the payload, such as alertName, service, severity, symptom, and time range, and put them into a compact recommended lookup query. The Agent still calls `lookup_knowledge` explicitly, so the trace remains auditable, but the retrieval query is less dependent on model improvisation.
If asked why not auto-call retrieval:
> Auto-calling retrieval would create hidden evidence before the Agent actually decides to use a tool. For this project, explicit tool invocation is more important because the interview story is about observable Agent execution. Prompt-level augmentation gives the Agent a better query seed without changing the trace contract.
+147
View File
@@ -0,0 +1,147 @@
# Architecture
## 系统分层
```text
API Layer
-> ChatController / DiagnosisTraceController
Agent Orchestration
-> ChatService / AiOpsService
Tools
-> lookupKnowledgeTool / queryLogs / queryMetrics / queryPrometheusAlerts
Persistence
-> diagnosis_session / agent_step / tool_invocation
Trace
-> GET /api/diagnosis/{sessionId}/trace
```
## Chat 链路
```mermaid
flowchart TD
User[User Question] --> ChatAPI[POST /api/chat]
ChatAPI --> Strategy[ChatService.executeChatWithStrategy]
Strategy --> Complexity{QuestionComplexity}
Complexity -->|simple| Single[ReactAgent]
Complexity -->|complex| Planner[Planner Agent]
Planner --> Executor[Executor Agent]
Executor --> Tools[Evidence Tools]
Tools --> Executor
Executor --> Verifier[Verifier Agent]
Verifier --> Answer[Final Answer]
Answer --> Session[diagnosis_session]
Planner --> Steps[agent_step]
Executor --> Steps
Verifier --> Steps
Tools --> Invocations[tool_invocation]
Session --> Trace[GET /api/diagnosis/{sessionId}/trace]
Steps --> Trace
Invocations --> Trace
```
关键代码:
- `ChatController.chat(...)`
- `ChatService.executeChatWithStrategy(...)`
- `ChatService.executeChatComplex(...)`
- `AgentLoggingHook`
- `ToolInvocationRecorder`
- `DiagnosisTraceService.getTrace(...)`
## AIOps 链路
```mermaid
flowchart TD
Alert[Alert Payload or Empty Request] --> AiOpsAPI[POST /api/ai_ops]
AiOpsAPI --> SessionEvent[SSE session event]
AiOpsAPI --> AiOpsService[AiOpsService.executeAiOpsAnalysis]
AiOpsService --> PromptMode{Payload?}
PromptMode -->|yes| Targeted[PAYLOAD_TARGETED]
PromptMode -->|no| Discovery[AUTO_DISCOVERY]
Targeted --> Supervisor[ai_ops_supervisor]
Discovery --> Supervisor
Supervisor --> Planner[planner_agent]
Supervisor --> Executor[executor_agent]
Planner --> Tools[Prometheus / Logs / Knowledge]
Executor --> Tools
Tools --> Report[Alert Report]
Report --> Persist[diagnosis_session.answer]
Planner --> Steps[agent_step]
Executor --> Steps
Tools --> Invocations[tool_invocation]
Persist --> Trace[GET /api/diagnosis/{sessionId}/trace]
Steps --> Trace
Invocations --> Trace
```
关键代码:
- `ChatController.aiOps(...)`
- `AIOpsRequest`
- `AiOpsService.resolveSessionId(...)`
- `AiOpsService.buildTaskPrompt(...)`
- `AiOpsService.hasAlertPayload(...)`
- `AiOpsService.persistFinalReport(...)`
## Trace 数据模型
### `diagnosis_session`
记录一次诊断会话的主信息:
- `session_id`
- `query`
- `status`
- `agent_flow`
- `total_duration_ms`
- `total_token_count`
- `step_count`
- `tool_call_count`
- `answer`
- `self_evaluation`
- `feedback`
### `agent_step`
记录 Agent 模型调用过程:
- `session_id`
- `step_index`
- `agent_name`
- `model_input`
- `model_output`
- `thought`
- `has_tool_call`
- `duration_ms`
- `token_count`
### `tool_invocation`
记录真实工具调用:
- `session_id`
- `tool_name`
- `input_params`
- `output_preview`
- `output_length`
- `retrieval_layer`
- `relevance_level`
- `duration_ms`
- `success`
- `error_message`
## 为什么 trace 是核心
Agent 系统的风险不只是“答案错”,还包括“答案看起来对但无法解释”。这个项目把执行链路拆成 session、step、tool 三层,让面试官可以看到:
- 模型为什么这么答
- 调了哪些工具
- 工具返回了什么证据
- Verifier 如何判断答案可信度
- 用户反馈如何回写到同一个 session
这就是项目区别于普通 Chatbot 的地方。
+132
View File
@@ -0,0 +1,132 @@
# Interview Demo Script
## 30 秒开场
这是一个 Agent Engineering 项目,场景是企业故障诊断。它支持两类入口:用户主动提问的 Chat 诊断,以及告警事件驱动的 AIOps 诊断。项目重点不是单次回答,而是把多 Agent 执行、工具证据、Verifier 评估、最终报告和反馈都沉淀成可回放的 trace。
## Demo 准备
启动服务:
```powershell
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
```
确认服务地址:
```text
http://localhost:9900
```
`mvp-demo` profile 下:
- Prometheus 告警使用 mock 数据。
- CLS 日志使用 mock 数据。
- MySQL、Redis、Milvus/Zilliz 和模型配置仍使用当前项目配置。
## Demo 1: Chat 诊断
目标:展示普通用户问题如何进入多 Agent 诊断、调用工具、经过 Verifier,并生成 trace。
请求:
```powershell
$sessionId = "interview-chat-payment-timeout-001"
$body = @{
Id = $sessionId
Question = "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
} | ConvertTo-Json
Invoke-RestMethod `
-Method Post `
-Uri "http://localhost:9900/api/chat" `
-ContentType "application/json" `
-Body $body
```
讲解点:
- `ChatController` 把请求交给 `ChatService.executeChatWithStrategy(...)`。
- 简单问题走单 ReactAgent,复杂问题走 `Planner -> Executor -> Verifier`。
- Executor 可以调用知识库、日志、指标等工具。
- Verifier 会基于工具证据生成 groundedness 评估。
- 最终会写入 `diagnosis_session`、`agent_step`、`tool_invocation`。
查询 trace:
```powershell
Invoke-RestMethod `
-Method Get `
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace"
```
展示点:
- `data.session.agentFlow = CHAT`
- `data.steps` 中能看到 planner/executor/verifier
- `data.toolInvocations` 中能看到证据工具
- `data.session.selfEvaluation` 中有 verifier 结果
## Demo 2: AIOps 告警诊断
目标:展示告警 payload 如何触发 AIOps 入口,并且报告只聚焦目标告警。
请求:
```powershell
$aiopsSessionId = "interview-aiops-payment-cpu-001"
$aiopsBody = @{
sessionId = $aiopsSessionId
alertName = "HighCPUUsage"
service = "payment-service"
severity = "P1"
description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
timeRange = "last_15m"
userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
} | ConvertTo-Json
Invoke-WebRequest `
-Method Post `
-Uri "http://localhost:9900/api/ai_ops" `
-ContentType "application/json" `
-Body $aiopsBody
```
讲解点:
- `/api/ai_ops` 接受可选 `AIOpsRequest`。
- 首条 SSE 消息会返回 `type=session`。
- `AiOpsService` 根据 payload 判断模式:
- `PAYLOAD_TARGETED`:聚焦传入告警。
- `AUTO_DISCOVERY`:没有 payload 时先查 active alerts。
- AIOps 暂时不加 Verifier,先保证告警入口、证据工具和 trace 可用。
查询 trace:
```powershell
Invoke-RestMethod `
-Method Get `
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace"
```
展示点:
- `data.session.agentFlow = AI_OPS`
- `data.session.answer` 有最终告警报告
- `data.toolInvocations` 有 `query_metrics`、`query_logs`、`lookup_knowledge`
- 报告有 `HighCPUUsage/payment-service` 的完整根因分析
- 其他 active alerts 只作为相关风险出现,不展开成独立根因章节
## MySQL 验证
```powershell
python scripts/query_mysql.py "SELECT session_id, agent_flow, status, step_count, tool_call_count FROM diagnosis_session ORDER BY id DESC LIMIT 5"
```
```powershell
python scripts/query_mysql.py "SELECT tool_name, COUNT(*) AS cnt FROM tool_invocation WHERE session_id='interview-aiops-payment-cpu-001' GROUP BY tool_name"
```
## 收尾总结
这套 Demo 展示的是一个完整 Agent 系统,而不是一次模型问答:入口有明确场景边界,Agent 负责规划和执行,工具提供证据,Verifier 提供质量门,trace API 提供审计和复盘能力。AIOps 入口进一步证明它可以从用户问答扩展到事件驱动诊断。
+99
View File
@@ -0,0 +1,99 @@
# Design Tradeoffs
## 1. 为什么要做 trace,而不是只返回答案
普通 Chatbot 只关注最终回答,但故障诊断更需要可审计性。一次诊断至少要回答三件事:
- 结论是什么
- 证据来自哪里
- 哪些步骤由哪个 Agent 完成
因此项目把一次会话拆成:
- `diagnosis_session`:会话级摘要、最终答案、质量评估、反馈。
- `agent_step`:Agent 模型输入输出、耗时、token 和工具调用标记。
- `tool_invocation`:真实工具调用参数、输出预览、成功状态和检索元数据。
这个设计牺牲了一些实现复杂度,但换来了可回放、可调试、可演示。
## 2. 为什么 Chat 有 Verifier,AIOps 暂时没有
Chat 入口的问题更开放,用户可能要求复杂推理或跨领域结论,所以 Verifier 是必要的质量门。当前 Chat 链路通过 `Planner -> Executor -> Verifier` 固定流程,把 groundedness 和 facts checked 写入 `self_evaluation`。
AIOps 当前阶段先不加 Verifier,原因是:
- AIOps 刚完成从“自动跑告警”到“可追踪告警入口”的改造。
- 先要确认告警 payload、工具证据、最终报告和 trace 能闭环。
- AIOps Verifier 的规则不同于 Chat Verifier,需要检查告警 scope、证据覆盖和处置建议,不宜直接复用。
后续可以做 lightweight AIOps Verifier,检查报告是否聚焦 payload、是否引用工具证据、是否误展开无关告警。
## 3. 为什么 AIOps payload scope 先用 prompt 控制
运行验证发现:传入 `HighCPUUsage/payment-service` 后,Agent 仍可能把 mock Prometheus 返回的所有 active alerts 都展开分析。这个问题的本质是任务边界不清晰。
当前选择 prompt-level scope control:
- 有 payload:`PAYLOAD_TARGETED`,最终报告围绕传入告警。
- 无 payload:`AUTO_DISCOVERY`,先调用 `queryPrometheusAlerts` 自动发现告警。
没有先做 Java 侧过滤,是因为:
- 过滤工具结果会降低 Agent 发现关联风险的能力。
- 目前需要的是报告主线聚焦,而不是完全屏蔽上下文。
- Prompt 改动小,风险低,能保留 Agent 灵活性。
已验证结果:主报告有 `HighCPUUsage/payment-service` 的完整根因分析,`HighMemoryUsage` 和 `SlowResponse` 只作为相关风险出现。
## 4. 为什么用 `tool_invocation` 统计真实工具调用次数
早期可以通过 `agent_step.hasToolCall` 粗略判断是否调用工具,但它统计的是“哪些模型步骤包含工具调用”,不是“真实调用了几次工具”。
现在 `tool_call_count` 来自:
```text
ToolInvocationRepository.countBySessionId(sessionId)
```
这样更符合 trace 语义:
- 一个 step 可能调用多个工具。
- 工具可能来自不同来源:知识库、日志、指标、Prometheus。
- 面试时可以把 `tool_call_count` 和 trace 中返回的工具明细对上。
## 5. 为什么保留 mock Prometheus 和 mock CLS
面试 Demo 最怕不稳定。真实 Prometheus、日志平台和线上故障都有不可控因素,所以 MVP profile 保留 mock 工具:
- `prometheus.mock-enabled=true`
- `cls.mock-enabled=true`
这样可以稳定复现:
- `HighCPUUsage/payment-service`
- `HighMemoryUsage/order-service`
- `SlowResponse/user-service`
- system-metrics、application-logs、database-slow-query 等日志证据
这不是逃避真实集成,而是把“Agent 编排和证据追踪”作为面试演示的主目标。
## 6. 为什么把面试材料单独放 `interview/`
`mvp/` 是持续迭代现场,包含过程文档、验收记录和 runbook。面试材料的目标不同,它应该是可讲、可演示、可评估的展示层。
因此:
- `mvp/` 保留真实演进材料。
- `devflow/` 保留决策沉淀。
- `interview/` 只组织面试叙事和演示脚本。
这样后续继续做 AIOps Verifier、UI、更多工具集成时,不会污染面试讲稿。
## 7. 可以主动承认的限制
- AIOps 还没有 Verifier。
- Prompt-level scope control 不能做到强约束,只能通过 trace 和测试观察遵循情况。
- 当前 mock 数据适合 demo,不代表生产接入已经完成。
- Hikari 连接池已经加了短生命周期和 keepalive,但真实生产还需要按数据库 wait_timeout 和连接数预算调优。
主动讲清这些限制,反而能体现工程判断:先把可追踪闭环打通,再逐步增强质量门和生产可靠性。
@@ -0,0 +1,66 @@
# RAG Breadcrumb Embedding Acceptance
## What Changed
The indexing path now builds embedding text from chunk structure plus content:
```text
Title: {title}
Path: {breadcrumb}
Content:
{content}
```
The stored Milvus `content` field remains the original chunk content. This keeps display and evidence output clean while allowing the vector to carry section-level semantics.
## Why Reindex Is Required
Embeddings are materialized at index time. Existing vectors were generated from the previous content-only text, so they cannot benefit from `title` and `breadcrumb` until the knowledge base is reindexed.
This is the key acceptance point:
```text
code change alone != live retrieval changed
code change + reindex + live query report = accepted behavior
```
## How To Validate
1. Start the Spring Boot application.
2. Reindex the knowledge base through the existing indexing path.
3. Run:
```bash
python scripts/eval_rag_live_acceptance.py
```
The script writes:
```text
eval/rag-retrieval/reports/live-post-reindex.json
eval/rag-retrieval/reports/live-post-reindex.md
```
The default cases cover:
- RAG chunk context questions where breadcrumb matters.
- Diagnosis flow questions where section path matters.
- `ERR_TIMEOUT` exact error-code retrieval.
- MySQL connection pool troubleshooting.
- AIOps payment-service latency alert retrieval.
## What To Look For
For breadcrumb-sensitive cases, inspect whether top candidates expose expected `title` and `breadcrumb` values in the report.
For core troubleshooting cases, check that result counts and top candidates remain stable. The goal is not to prove a full benchmark; it is to prove that reindexing did not obviously break important demo retrieval paths.
## Interview Answer
If asked how I verified the breadcrumb embedding change:
> I separated deterministic regression from live acceptance. The offline fixture baseline still runs without services. But because embedding changes only affect newly indexed vectors, I added a live post-reindex acceptance script. It calls the real `/api/search/similar` endpoint against representative breadcrumb-sensitive, troubleshooting, and AIOps queries, then writes JSON and Markdown reports. This lets me prove both that the code changed and that the live vector collection was refreshed.
If asked why the script does not reindex automatically:
> Reindexing mutates the vector store and depends on environment-specific data. I kept mutation explicit and made the script validation-only. That makes failures easier to diagnose: if retrieval does not improve, I can distinguish code changes, reindex state, and runtime retrieval behavior.
+209
View File
@@ -0,0 +1,209 @@
# RAG Refactor Story
## The Starting Point
The original RAG implementation was already usable for the MVP:
- Documents could be uploaded, chunked, embedded, and written to Milvus/Zilliz.
- The Agent could call `lookup_knowledge` as an explicit tool.
- AIOps diagnosis could retrieve troubleshooting knowledge during an alert workflow.
- Tool invocations were persisted, so the retrieval step was visible in the execution trace.
But the design had several engineering problems:
- Retrieval was too SDK-specific. The business code directly owned many Milvus search details.
- L0 and L1 responsibilities were blurry. L0 keyword matching could look like a final retrieval decision instead of a hint.
- Chunk-level retrieval could lose section context when one section was split into multiple chunks.
- Metadata such as `breadcrumb` existed, but it was not fully used in retrieval, filtering, or context reconstruction.
- Retrieval quality was mostly checked by manual API calls and logs, not by repeatable cases.
So the refactor goal was not "replace everything with a framework." The goal was to move generic RAG infrastructure toward Spring AI while keeping the project-specific Agent evidence chain.
## How I Broke The Problem Down
I treated this as a staged migration, because RAG touches the Agent tool layer, AIOps diagnosis, vector retrieval, evidence packing, and database traces.
The first step was to establish a baseline. I added retrieval evaluation cases under `eval/rag-retrieval/` so future changes could be compared against known queries instead of judged only by intuition.
Then I clarified the retrieval roles:
```text
L0 = domain/entity hint
L1 = semantic retrieval
postprocess = evidence shaping and trace-friendly output
```
That means L0 is still valuable, but it should not bypass semantic retrieval as the default path. It is better used to extract service names, alert names, error codes, domains, and metadata hints.
After that, I added evidence postprocessing. The Agent should not just receive raw chunks; it should receive structured evidence with source, title, breadcrumb, score, hit reason, and content. This makes the result easier to inspect and easier to explain in an interview.
Finally, I integrated Spring AI `VectorStore` as the main read path while preserving the original Milvus SDK implementation as fallback.
## Current Architecture
The current retrieval path is:
```text
Agent / API
-> lookup_knowledge or /api/search/similar
-> L0 domain/entity hint
-> VectorSearchService
-> Spring AI VectorStore
-> Milvus SDK fallback
-> evidence postprocess
-> tool_invocation trace
```
`VectorSearchService` is still the public retrieval facade. This is deliberate: the Agent tool layer does not need to know whether the underlying retrieval engine is SDK-based or Spring AI-based.
The supported retrieval modes are:
```text
auto -> try Spring AI VectorStore, fallback to SDK
spring-ai -> force Spring AI VectorStore
sdk -> force Milvus SDK
```
This keeps the migration reversible and testable.
## Key Tradeoffs
### Keep The Explicit Tool
I did not hide retrieval inside a Spring AI Advisor.
For this project, `lookup_knowledge` is part of the Agent execution story. It records what query was used, which evidence was retrieved, how relevant it looked, and how it supported diagnosis. If retrieval is hidden inside an advisor, the answer may still work, but the audit trail becomes harder to show.
### Keep SDK Fallback
The SDK path is not dead code. It is a safety net during migration.
This proved useful during live validation. The first VectorStore run pointed at the wrong collection name, but `auto` mode fell back to SDK and still returned results. After the collection was corrected to `biz`, the Spring AI path worked as the main path.
### Keep L0, But Reduce Its Authority
L0 is worth keeping because production incidents often contain exact identifiers:
- error code
- alert name
- service name
- metric name
- domain tag
But L0 should not be the final judge of retrieval quality. Its role is now closer to domain hint, entity extraction, metadata filtering, and explainability signal.
### Split Score Semantics
The old SDK path used L2 distance. Spring AI exposes similarity. Treating those as the same number would quietly break relevance normalization.
So the result separates:
```text
score -> compatibility score used by existing logic
rawScore -> raw score from the retrieval implementation
scoreLabel -> semantic label for rawScore
```
For SDK:
```text
score = L2 distance
rawScore = L2 distance
scoreLabel = l2_distance
```
For VectorStore:
```text
score = Milvus metadata.distance when available
rawScore = Spring AI similarity
scoreLabel = similarity
```
This makes the migration inspectable instead of hiding score changes behind one overloaded field.
### Do Not Migrate Writes Yet
Writes and indexing still use the SDK path.
That is intentional. Migrating reads and writes at the same time would make debugging harder. The read path can be validated first; write-path migration can happen later if Spring AI `VectorStore.add(...)` fits the existing metadata and chunk model.
## Validation Story
I validated the refactor at multiple levels.
Unit tests cover:
- SDK mode.
- Spring AI mode.
- `auto` fallback.
- category filter behavior.
- distance metadata mapping.
Live API verification used:
```text
GET /api/search/similar?query=ERR_TIMEOUT&topK=3
```
Logs confirmed when the Spring AI VectorStore path was used and when fallback happened.
Then I compared SDK and VectorStore retrieval quality on representative queries:
| Query Type | Result |
| --- | --- |
| exact error code | same top3 |
| payment-service timeout | same top3 |
| MySQL connection pool | same top3 |
| AIOps alert-style query | same top3 |
| abstract RAG design query | same top1, VectorStore returned fewer tail results |
| category filter | both returned zero because metadata taxonomy did not match |
The acceptance decision was that Spring AI VectorStore is good enough for the current MVP read path, with SDK fallback preserved.
## Known Gaps
The refactor improved the architecture, but it did not solve every retrieval-quality problem.
Known gaps:
- Metadata taxonomy still needs cleanup, for example `database` vs `infrastructure`.
- Abstract design questions may need query rewriting or better indexed interview/devflow documents.
- Chunk context reconstruction is still limited when one logical section spans multiple chunks.
- `breadcrumb` now participates in embedding text, but it can still be used more strongly in context expansion, rerank, and evidence packing.
- Rerank, RRF, BM25, and hybrid retrieval are not implemented yet.
- Indexing writes still use SDK.
These are good follow-up issues because they are retrieval-quality improvements, not blockers for the VectorStore migration.
## How I Present This In An Interview
My short version would be:
> This RAG system started as a self-built MVP around Milvus SDK retrieval. It worked, but too much infrastructure logic lived in business code, and L0/L1 responsibilities were unclear. I refactored it in stages: first I added baseline retrieval cases, then made L0 a domain/entity hint instead of a final decision layer, then added evidence postprocessing, and finally moved the main read path to Spring AI VectorStore with SDK fallback. I kept `lookup_knowledge` as an explicit Agent tool because the project values traceability: the interviewer can see when retrieval happened, what evidence was found, and how it supported the diagnosis. The result is closer to standard Spring AI RAG while still preserving business-specific observability.
If asked why this is not a full framework migration:
> I intentionally did not migrate everything at once. Reads moved first because they are easier to compare using golden queries. Writes/indexing stayed on SDK to avoid mixing schema and retrieval behavior changes in one step. Advisors were not used as the main interface because hidden retrieval would weaken the Agent trace.
If asked what I would improve next:
> I would add query transformation for AIOps payloads, improve metadata taxonomy, use breadcrumb and section metadata for context expansion, and then evaluate whether hybrid retrieval or rerank is necessary based on measured recall and topK overlap.
## Interview Follow-Up Questions
### Why introduce Spring AI VectorStore if the SDK path already worked?
Because SDK-only retrieval made the project own too much low-level RAG infrastructure. `VectorStore` gives a standard abstraction for retrieval and makes future Spring AI features easier to adopt, while the facade keeps the Agent layer stable.
### Why keep custom code at all?
The custom code is where the Agent engineering value lives: AIOps payload mapping, L0 hints, evidence packing, score compatibility, and tool invocation tracing. Those are domain-specific and should remain visible.
### How do you know quality did not regress?
I compared SDK and VectorStore modes on representative live queries. Core troubleshooting and AIOps cases returned the same top3 documents in the same order. The differences were isolated to abstract design queries and metadata taxonomy, which are documented follow-up work.
### What is the most important design decision?
Keeping a stable boundary: `lookup_knowledge` calls `VectorSearchService`, and `VectorSearchService` decides whether to use Spring AI or SDK. That boundary made the migration small enough to validate and explain.
+211
View File
@@ -0,0 +1,211 @@
# RAG Retrieval Quality Report
## Purpose
This report compares the live retrieval behavior of the original Milvus SDK path and the new Spring AI VectorStore path.
The goal is to answer an interview-critical question:
> After moving retrieval to Spring AI VectorStore, how do we know retrieval quality did not regress?
This is not a full benchmark yet. It is a focused live smoke comparison using representative RAG queries against the current Milvus/Zilliz collection.
## Setup
Service endpoint:
```text
GET http://127.0.0.1:9900/api/search/similar
```
Collection:
```text
biz
```
Compared modes:
```text
retrieval.vector-store.mode=sdk
retrieval.vector-store.mode=spring-ai
```
Each case used:
```text
topK=3
```
The application was restarted once per mode using command-line configuration so no repository config file had to be changed.
## Cases
| Case | Query | Purpose |
| --- | --- | --- |
| `err-timeout` | `ERR_TIMEOUT` | Exact error-code retrieval |
| `payment-service-timeout` | `payment-service timeout` | Service timeout troubleshooting |
| `mysql-connection-pool` | `MySQL connection pool is exhausted. How should I diagnose it?` | Database troubleshooting |
| `high-cpu-payment` | `HighCPUUsage payment-service` | AIOps alert-style retrieval |
| `rag-l0-l1` | `Should L0 keyword matching decide the final retrieval result?` | Abstract RAG design query |
| `database-filter` | `mysql timeout`, category=`database` | Metadata filter behavior |
## Summary
| Case | SDK Count | VectorStore Count | Top1 Same | TopK Overlap | Notes |
| --- | ---: | ---: | --- | ---: | --- |
| `err-timeout` | 3 | 3 | Yes | 3/3 | Same ordering and same documents |
| `payment-service-timeout` | 3 | 3 | Yes | 3/3 | Same ordering and same documents |
| `mysql-connection-pool` | 3 | 3 | Yes | 3/3 | Same ordering and same documents |
| `high-cpu-payment` | 3 | 3 | Yes | 3/3 | Same ordering and same documents |
| `rag-l0-l1` | 3 | 1 | Yes | 1/3 | VectorStore returned only the strongest candidate |
| `database-filter` | 0 | 0 | N/A | N/A | Both paths applied the filter consistently; no live docs matched `category=database` |
## Representative Results
### `ERR_TIMEOUT`
SDK:
```text
1. ERR_TIMEOUT score=0.5659486 label=l2_distance
2. ERR_GATEWAY_TIMEOUT score=0.6048740 label=l2_distance
3. Error handling score=0.7735061 label=l2_distance
```
VectorStore:
```text
1. ERR_TIMEOUT score=0.5659486 rawScore=0.4340513 label=similarity
2. ERR_GATEWAY_TIMEOUT score=0.6048740 rawScore=0.3951259 label=similarity
3. Error handling score=0.7735061 rawScore=0.2264938 label=similarity
```
Interpretation:
- Document ordering is identical.
- Compatibility `score` is identical to SDK L2 distance.
- VectorStore `rawScore` exposes Spring AI similarity separately.
### `MySQL connection pool`
Both paths returned:
```text
1. MySQL connection pool config
2. wait_timeout timeout
3. idle-timeout
```
Interpretation:
- The migration preserves a precise infrastructure troubleshooting retrieval case.
- Metadata fields such as title, category, and source remain available.
### `HighCPUUsage payment-service`
Both paths returned:
```text
1. 3. HighCPUUsage / payment-service troubleshooting steps
2. evidence mapping table row for HighCPUUsage/payment-service
3. 3.1 Symptom confirmation
```
Interpretation:
- AIOps-style alert terms still retrieve the expected troubleshooting document.
- This is important because AIOps diagnosis depends on knowledge retrieval plus metrics/log evidence.
### `rag-l0-l1`
SDK returned three results, while VectorStore returned one:
```text
Top1: Return error information
```
Interpretation:
- Top1 did not regress.
- VectorStore appears stricter for low-similarity tail results because the Spring AI path uses `similarityThresholdAll()`.
- This is acceptable for current read-path migration, but it is worth tracking because abstract design questions may need query rewriting, better indexed docs, or adjusted threshold behavior.
### `database-filter`
Both paths returned zero results for:
```text
query=mysql timeout
category=database
```
Interpretation:
- The filter path is consistent.
- The live indexed MySQL docs are categorized as `infrastructure`, not `database`.
- This highlights a metadata taxonomy issue rather than a VectorStore migration regression.
## Score Compatibility
The comparison validates the score design:
```text
SDK:
score = L2 distance
rawScore = L2 distance
scoreLabel = l2_distance
VectorStore:
score = Milvus metadata.distance
rawScore = Spring AI similarity
scoreLabel = similarity
```
This keeps `lookup_knowledge` relevance normalization stable while still exposing the VectorStore score semantics for trace/debugging.
## Findings
### Finding 1: Main live cases are equivalent
For exact error code, service timeout, MySQL troubleshooting, and AIOps alert-style retrieval, SDK and VectorStore returned identical top3 documents in identical order.
This is strong evidence that the read-path migration did not regress the most important demo and troubleshooting cases.
### Finding 2: Abstract RAG design queries need better retrieval support
The `rag-l0-l1` query only returned one VectorStore candidate. The top result matched SDK top1, but the tail differed.
This suggests the next quality work should focus on:
- Query transformation for abstract design questions.
- Better indexing of interview/devflow RAG design docs.
- Context expansion around same-section chunks.
- Possibly tuning VectorStore threshold behavior.
### Finding 3: Metadata taxonomy matters
The category filter case returned zero results in both modes because the relevant MySQL docs are categorized as `infrastructure`, not `database`.
This supports a previous RAG issue: category/domain metadata should be normalized before it is used as a hard filter.
## Acceptance Decision
The Spring AI VectorStore read path is accepted for current MVP/interview use:
- Core troubleshooting cases match SDK behavior.
- Score compatibility is preserved.
- The VectorStore path exposes better score semantics without changing the `lookup_knowledge` API.
- SDK fallback remains available for runtime safety.
The next retrieval-quality improvements should not block this migration. They should be handled as separate RAG quality work.
## Next Work
Recommended next steps:
- Add a small automated live comparison script if repeated validation becomes common.
- Add topK overlap and top1 hit metrics to the offline evaluator.
- Normalize metadata categories such as `database` vs `infrastructure`.
- Add query rewriting for abstract RAG questions.
- Decide later whether to migrate indexing writes to Spring AI `VectorStore.add(...)`.
@@ -0,0 +1,167 @@
# RAG VectorStore Interview Notes
## 60-Second Explanation
I refactored the RAG retrieval path from a direct Milvus SDK-only implementation to a Spring AI `VectorStore` main path, while keeping the SDK path as a fallback.
The important part is not just the dependency change. I kept `VectorSearchService` as the boundary, so `lookup_knowledge` and the Agent workflow did not need to change. The system now supports three modes:
```text
auto -> try Spring AI VectorStore, fallback to SDK
spring-ai -> force VectorStore
sdk -> force SDK
```
During live verification, the first run found a real config mismatch: VectorStore was pointed at `business_knowledge`, but the real Zilliz collection was `biz`. The fallback worked, so the system still returned results through SDK. After aligning the collection name, the same query went through Spring AI VectorStore successfully.
I also fixed score compatibility. Spring AI Milvus exposes similarity as the document score, but the old `lookup_knowledge` logic expects L2 distance. So I preserve `rawScore` and `scoreLabel`, and use Milvus `metadata.distance` as the compatibility `score` when available.
## Architecture Answer
```text
Agent / API
-> lookup_knowledge or /api/search/similar
-> VectorSearchService
-> Spring AI VectorStore
-> Milvus SDK fallback
-> Milvus/Zilliz collection: biz
```
The key design choice is that `VectorSearchService` remains the retrieval facade. This avoids spreading framework-specific code into the Agent tool layer.
## Why Keep The SDK Path?
I kept SDK fallback for three reasons:
- Migration safety: the existing SDK path was already proven against the live collection.
- Runtime resilience: if VectorStore schema mapping or filtering fails, retrieval still works.
- Interview/demo stability: a retrieval abstraction change should not break the main Agent diagnosis demo.
This was validated in practice. When VectorStore pointed at the wrong collection, `auto` mode fell back to SDK and still returned results.
## Why Use Spring AI VectorStore At All?
Using Spring AI `VectorStore` moves the project closer to a standard RAG abstraction:
- Retrieval code no longer needs to own all Milvus-specific search details.
- Later features such as query transformers, document postprocessors, advisors, or retrievers can be introduced more naturally.
- The code becomes easier to compare with common Spring AI RAG patterns in an interview.
But I did not blindly replace everything. Writes/indexing still use SDK because changing read and write paths at the same time would make failures harder to isolate.
## Why Keep L0?
L0 is no longer treated as the final source of truth. It is a deterministic hint layer:
- It extracts domain/entity hints from indexed metadata.
- It helps constrain L1 retrieval by category when possible.
- It gives the Agent a stable clue even when semantic retrieval is weak.
The current design is:
```text
L0 = domain/entity hint
L1 = semantic retrieval through VectorStore/SDK
postprocess = evidence trace and relevance normalization
```
This is easier to defend than saying "we only use vector search." Real incident diagnosis often has exact identifiers, error codes, service names, and alert names. L0 is useful for those.
## Why Not Use Hidden Spring AI Advisors Directly?
For this project, `lookup_knowledge` remains an explicit tool.
Reason:
- The Agent trace needs to show when knowledge was retrieved.
- `tool_invocation` records input, output preview, relevance level, and evidence metadata.
- The interview story is about auditable Agent execution, not only answer quality.
Spring AI Advisors may be useful later, but hiding retrieval inside an advisor would make the evidence chain less visible unless we rebuild trace hooks around it.
## Score Design
The result object intentionally separates these fields:
```text
score -> compatibility score used by old relevance normalization
rawScore -> raw score from the retrieval implementation
scoreLabel -> semantic label for rawScore
```
For SDK:
```text
score = L2 distance
rawScore = L2 distance
scoreLabel = l2_distance
```
For VectorStore:
```text
score = metadata.distance if present
rawScore = Spring AI document score
scoreLabel = similarity
```
This prevents a subtle bug: if we treat Spring AI similarity as L2 distance, relevance becomes wrong. If we only expose distance, we lose the ability to compare Spring AI behavior. Keeping both makes the migration inspectable.
## How I Verified It
I verified at three levels:
- Unit tests: SDK mode, auto VectorStore mode, fallback mode, category filter, distance metadata mapping.
- Live API: `/api/search/similar?query=ERR_TIMEOUT&topK=3`.
- Logs: confirmed whether the path was VectorStore success or SDK fallback.
The live API returned:
```text
scoreLabel = similarity
rawScore = Spring AI similarity
score = Milvus distance metadata
```
That means the main path was Spring AI VectorStore and compatibility scoring remained stable.
## What I Would Do Next
I would not immediately migrate indexing writes. The next responsible steps are:
- Add a small live acceptance report for several golden queries.
- Compare `sdk` and `spring-ai` mode side by side for topK overlap.
- Decide whether `VectorIndexService` should move to `VectorStore.add(...)`.
- Add query transformation or hybrid retrieval only after we have baseline metrics.
This staged approach is intentional: first stabilize the read path, then evaluate retrieval quality, then migrate writes if the abstraction proves reliable.
## Interview Questions And Short Answers
### Why did you not remove the SDK?
Because this is a migration, not a rewrite. SDK fallback gives rollback safety and proved useful when VectorStore config was initially wrong.
### What changed for `lookup_knowledge`?
The public contract did not change. It still calls `VectorSearchService.searchSimilarDocuments(...)`. The implementation behind that facade changed.
### How do you know VectorStore is actually used?
The logs show `Starting Spring AI VectorStore search` followed by `Spring AI VectorStore search complete`. The API response also has `scoreLabel=similarity`, which only comes from the VectorStore path.
### What was the main bug found during live validation?
The configured collection name was wrong. Spring AI looked for `business_knowledge`, but the actual Milvus collection was `biz`.
### What did fallback prove?
It proved that `auto` mode is resilient: VectorStore failed, SDK search still returned valid results, and the API did not fail.
### Why is `metadata.distance` important?
Because `lookup_knowledge` uses L2 distance normalization. Spring AI returns similarity as the main document score, but the Milvus distance is available in metadata. Using it preserves old relevance behavior.
### Is this full Spring AI RAG now?
Not yet. It uses Spring AI VectorStore for the main read path, but keeps explicit tools, custom evidence trace, L0 hints, and SDK indexing. That is deliberate because the project values auditability and staged migration.
@@ -0,0 +1,199 @@
# RAG VectorStore Live Acceptance
## Purpose
This note records the live acceptance result for the RAG retrieval refactor.
The goal of this refactor was not only to add a Spring AI abstraction, but to prove that the production retrieval path can:
- Prefer Spring AI `VectorStore` for Milvus retrieval.
- Preserve the existing Milvus SDK path as fallback.
- Keep the `lookup_knowledge` tool contract stable.
- Keep L2-distance based relevance normalization compatible.
## Current Retrieval Shape
```text
lookup_knowledge / /api/search/similar
-> VectorSearchService.searchSimilarDocuments(...)
-> retrieval.vector-store.mode
-> auto
-> Spring AI VectorStore
-> fallback to Milvus SDK if VectorStore fails
-> spring-ai
-> Spring AI VectorStore only
-> sdk
-> Milvus SDK only
```
## Configuration Verified
The live Milvus/Zilliz database contains the collection:
```text
biz
```
The Spring AI VectorStore configuration was aligned with the existing SDK collection:
```yaml
spring:
ai:
vectorstore:
type: milvus
milvus:
initialize-schema: false
database-name: ${milvus.database}
collection-name: biz
embedding-dimension: ${milvus.vector-dim}
metric-type: L2
id-field-name: id
content-field-name: content
metadata-field-name: metadata
embedding-field-name: vector
```
Why this matters: the earlier config used `business_knowledge`, but the SDK path and real collection use `biz`. That mismatch proved the fallback worked, but it also meant VectorStore was not the successful main path until the config was corrected.
## Commands Used
Health check:
```powershell
Invoke-RestMethod `
-Uri "http://127.0.0.1:9900/milvus/health" `
-Method Get
```
Observed result:
```json
{
"collections": ["biz"],
"message": "ok"
}
```
Direct retrieval check:
```powershell
Invoke-RestMethod `
-Uri "http://127.0.0.1:9900/api/search/similar?query=ERR_TIMEOUT&topK=3" `
-Method Get
```
Observed result shape:
```json
{
"code": 200,
"message": "success",
"data": [
{
"id": "f7dff7c8-5665-3145-9f75-ef741528b914",
"content": "### ERR_TIMEOUT ...",
"score": 0.5662,
"rawScore": 0.4337,
"scoreLabel": "similarity",
"metadata": {
"distance": 0.5662,
"title": "ERR_TIMEOUT",
"category": "api"
}
}
]
}
```
## What The Logs Proved
Before collection alignment:
```text
Starting Spring AI VectorStore search
SearchRequest collectionName:business_knowledge failed
Spring AI VectorStore retrieval failed, falling back to Milvus SDK
Starting Milvus SDK search
```
After collection alignment:
```text
Starting Spring AI VectorStore search: query=ERR_TIMEOUT
Spring AI VectorStore search complete, candidates=3
```
This proves:
- `auto` mode really attempts VectorStore first.
- The fallback is functional when VectorStore fails.
- After config alignment, the main path is Spring AI VectorStore rather than SDK fallback.
## Score Semantics
The project keeps three score fields intentionally:
```text
rawScore -> the raw score from the active retrieval implementation
scoreLabel -> the semantic meaning of rawScore
score -> compatibility score used by existing lookup relevance normalization
```
For SDK retrieval:
```text
rawScore = L2 distance
scoreLabel = l2_distance
score = L2 distance
```
For Spring AI VectorStore retrieval:
```text
rawScore = Spring AI similarity score
scoreLabel = similarity
score = Milvus distance metadata when available
```
Why use `metadata.distance` for `score`: `LookupKnowledgeTool` already normalizes relevance from L2 distance. Spring AI Milvus returns similarity as the document score, but also includes the Milvus distance in metadata. Using distance preserves the old relevance behavior while still exposing the new VectorStore score semantics through `rawScore` and `scoreLabel`.
## Regression Checks
Targeted tests:
```powershell
mvn -q "-Dtest=VectorSearchServiceTest,LookupKnowledgeToolTest" test
```
Spec validation:
```powershell
openspec.cmd validate rag-knowledge-retrieval --specs
openspec.cmd validate rag-retrieval-evaluation --specs
```
Whitespace check:
```powershell
git diff --check
```
Observed result:
```text
All targeted tests passed.
All related specs passed.
No diff-check errors.
```
## Acceptance Conclusion
The VectorStore refactor is accepted for the read path:
- Spring AI VectorStore is integrated and selected in `auto` mode.
- The SDK path remains available and was proven by fallback behavior.
- The live collection configuration is aligned with the existing Milvus collection.
- The `lookup_knowledge` public contract remains stable.
- Existing L2-based relevance normalization remains compatible.
The write/indexing path still uses the Milvus SDK. That is an intentional staged migration decision, not a failed acceptance item.
@@ -0,0 +1,81 @@
---
title: AIOps 告警排障 Runbook
keywords: [AIOps, 告警, HighCPUUsage, SlowResponse, payment-service, system-metrics, application-logs]
summary: 面向 AIOps 告警诊断的排障步骤,覆盖 Prometheus 活动告警、CLS 日志主题和处理建议。
category: troubleshooting
---
# AIOps 告警排障 Runbook
## 1. 告警输入处理原则
AIOps 诊断入口有两种触发方式:
- **有告警 payload**:将 payload 视为已触发告警,围绕 `alertName`、`service`、`severity`、`timeRange` 查询指标、日志和知识库。
- **无告警 payload**:先调用 `queryPrometheusAlerts` 获取当前 firing 告警,再选择 P0/P1 或持续时间最长的告警进入诊断。
最终报告必须基于工具证据,不得凭空编造指标、日志或处理结果。
## 2. Mock 告警与日志主题映射
| 告警名 | 典型服务 | 优先日志主题 | 推荐查询 |
|---|---|---|---|
| HighCPUUsage | payment-service | system-metrics | `cpu_usage:>80 AND service:payment-service` |
| HighMemoryUsage | order-service | system-metrics, system-events | `memory_usage:>85` |
| SlowResponse | user-service | application-logs, database-slow-query | `duration:>3000 OR slow request` |
| ServiceUnavailable | 任意核心服务 | application-logs, system-events | `level:ERROR OR container crash` |
## 3. HighCPUUsage / payment-service 排障步骤
### 3.1 现象确认
先确认 Prometheus 活动告警中是否存在:
- `alert_name = HighCPUUsage`
- `service = payment-service`
- CPU 使用率超过 80%
- 状态为 firing
如果 payload 已经提供该告警,也仍需通过指标或日志工具验证。
### 3.2 指标与日志取证
推荐工具调用顺序:
1. `queryPrometheusAlerts`:确认当前活动告警。
2. `queryLogs(region=ap-guangzhou, logTopic=system-metrics, query=cpu_usage:>80 AND service:payment-service)`:确认 CPU 使用率、实例和持续时间。
3. 如报告中提到 Redis、数据库或下游依赖,再查询 `application-logs` 或对应主题交叉验证。
### 3.3 根因判断
可接受的根因结论必须至少满足一项:
- system-metrics 显示 payment-service 实例 CPU 使用率持续高于阈值。
- application-logs 显示与 CPU 飙高同时出现的慢请求、线程池耗尽或依赖超时。
- 告警持续时间与日志时间线一致。
如果只有活动告警,没有日志或指标明细,应输出低置信结论并建议人工确认。
## 4. 处理建议
### 临时止血
- 对 payment-service 做水平扩容,优先扩容受影响实例所在 Deployment。
- 对高耗时接口开启限流或降级非核心功能。
- 如果近期有发布,检查变更窗口并准备回滚。
### 根因修复
- 分析 CPU 热点线程、慢请求接口和依赖调用耗时。
- 检查连接池、线程池、缓存穿透和批量任务是否导致 CPU 飙高。
- 补充针对 `payment-service` 的 CPU、P95/P99 延迟、错误率和依赖超时联动告警。
## 5. 报告要求
告警分析报告至少包含:
- 活跃告警清单。
- 告警根因分析。
- 使用过的工具证据:Prometheus 告警、system-metrics 日志、application-logs 或知识库。
- 已执行或建议执行的处理方案。
- 置信度说明:哪些结论有直接证据,哪些需要人工进一步确认。
+71
View File
@@ -2,6 +2,13 @@
This demo proves the MVP flow from user question to persisted diagnosis trace.
For interview use, start with:
- `interview-walkthrough.md` for the talk track
- `trace-inspection-checklist.md` for fields to inspect
- `scripts/run-payment-timeout-demo.ps1` for the runnable local demo
- `requests/payment-timeout-chat.json` for the fixed request payload
## Prerequisites
- MySQL, Redis, Milvus/Zilliz, and LLM/embedding configuration are available through the current project configuration.
@@ -22,6 +29,22 @@ http://localhost:9900
## 1. Run Chat Diagnosis
Fast path:
```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
```
This writes:
```text
mvp/demo/output/chat-response.json
mvp/demo/output/trace-response.json
mvp/demo/output/feedback-response.json
```
Manual path:
```powershell
$sessionId = "mvp-demo-payment-timeout-001"
$body = @{
@@ -78,6 +101,43 @@ Expected result:
- `success` is `true`.
- A later trace query shows `data.session.feedback` as `useful`.
## 4. Run AIOps Alert Diagnosis
```powershell
$aiopsSessionId = "mvp-demo-aiops-payment-cpu-001"
$aiopsBody = @{
sessionId = $aiopsSessionId
alertName = "HighCPUUsage"
service = "payment-service"
severity = "P1"
description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
timeRange = "last_15m"
userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
} | ConvertTo-Json
Invoke-WebRequest `
-Method Post `
-Uri "http://localhost:9900/api/ai_ops" `
-ContentType "application/json" `
-Body $aiopsBody
```
Expected result:
- The SSE stream starts with a `session` message containing `mvp-demo-aiops-payment-cpu-001`.
- The stream later contains an AIOps alert analysis report focused on the supplied `HighCPUUsage/payment-service` payload.
- A trace query for the same session id returns `data.session.agentFlow` as `AI_OPS`.
- `data.session.answer` contains the final alert analysis report when a report is generated.
- `data.toolInvocations` contains evidence tools such as `lookup_knowledge`, `query_logs`, or `query_metrics` when the runtime uses them.
Query the AIOps trace:
```powershell
Invoke-RestMethod `
-Method Get `
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace"
```
## Demo Story
The important interview story is:
@@ -92,3 +152,14 @@ one session id
-> feedback
-> trace API for replay and audit
```
The AIOps story uses the same audit spine:
```text
one session id
-> alert payload
-> AIOps planner/executor execution
-> evidence tools
-> alert analysis report
-> trace API for replay and audit
```
+38
View File
@@ -0,0 +1,38 @@
# AIOps Alert Acceptance Case
## Goal
Validate that the legacy AIOps endpoint can act as a traceable alert-triggered diagnosis entry.
## Input
- Session id: `mvp-demo-aiops-payment-cpu-001`
- Endpoint: `POST /api/ai_ops`
- Profile: `mvp-demo`
- Alert:
```json
{
"sessionId": "mvp-demo-aiops-payment-cpu-001",
"alertName": "HighCPUUsage",
"service": "payment-service",
"severity": "P1",
"description": "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。",
"timeRange": "last_15m",
"userRequest": "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
}
```
## Acceptance Criteria
1. The SSE stream emits a `session` message containing the requested session id.
2. The AIOps run creates or updates `diagnosis_session` with `agent_flow = AI_OPS`.
3. The persisted session query contains the alert name, service, severity, time range, and description.
4. If a final report is generated, `diagnosis_session.answer` contains that report.
5. `GET /api/diagnosis/{sessionId}/trace` returns the AIOps session, ordered agent steps, and ordered tool invocations.
6. In payload mode, the report focuses on `HighCPUUsage/payment-service`; unrelated active alerts may appear only as related risk or context, not as separate full root-cause sections.
## Known Limits
- This slice does not add a Verifier Agent to AIOps.
- Full runtime verification still depends on valid DB, Redis, Milvus/Zilliz, model, and embedding configuration.
+146
View File
@@ -0,0 +1,146 @@
# Interview Walkthrough: MVP Diagnosis Agent
This walkthrough is the Plan C demo story. It is meant for a short Agent Engineer interview, not as exhaustive system documentation.
## 30-Second Summary
```text
This is an enterprise diagnosis Agent MVP.
It takes a payment-timeout question, plans the investigation, calls evidence tools,
checks the answer through a verifier, persists the full trace, and accepts feedback.
```
The important claim is not "the model answered once." The claim is:
```text
The system can show what evidence was used, how the answer was checked, and how to replay the session.
```
## Demo Flow
1. Start the service with the `mvp-demo` profile.
2. Run the fixed payment-timeout request.
3. Open `mvp/demo/output/chat-response.json`.
4. Open `mvp/demo/output/trace-response.json`.
5. Point to evidence tools and verifier evaluation.
6. Submit feedback and show it is attached to the same session.
## Commands
Start service:
```powershell
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
```
Run the demo from another terminal:
```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
```
Optional custom session:
```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1 -SessionId "mvp-demo-payment-timeout-002"
```
## What To Show
### 1. User-Facing Answer
File:
```text
mvp/demo/output/chat-response.json
```
Say:
```text
This is the answer the user sees. The session id is stable, so I can trace this exact answer later.
```
### 2. Evidence Trace
File:
```text
mvp/demo/output/trace-response.json
```
Say:
```text
This is the important Agent engineering part.
I can inspect which tools were called, what inputs they received,
whether they succeeded, and what evidence preview was persisted.
```
Point to:
- `data.toolInvocations[*].toolName`
- `data.toolInvocations[*].inputParams`
- `data.toolInvocations[*].outputPreview`
- `data.toolInvocations[*].success`
### 3. Verifier / Self-Evaluation
Point to:
- `data.session.selfEvaluation`
- `data.summary.hasVerifierEvaluation`
Say:
```text
The final answer is not just raw Executor output.
It is checked by a verifier or self-evaluation layer using the persisted trace.
That lets the system return PASS, LOW_CONFID, or REJECT-style behavior instead of pretending all answers are equally certain.
```
### 4. Feedback Loop
File:
```text
mvp/demo/output/feedback-response.json
```
Then re-query trace if needed.
Say:
```text
Feedback is attached to the same diagnosis session.
That makes it possible to mine useful / not useful cases later.
```
### 5. Regression Story
Mention, do not deep dive unless asked:
```text
For repeatability, I also built an offline eval baseline.
The demo proves the runtime trace; the eval baseline proves fixed-case regression.
The two are separate on purpose: demo for human review, eval for automated signal.
```
## Strong Interview Framing
Use this phrasing:
```text
I focused on the Agent engineering surface:
traceability, evidence persistence, verifier gating, feedback, and regression checks.
The model answer is only one part of the system.
The more important part is whether we can audit and improve the answer after it is produced.
```
## Known Limits To Say Proactively
```text
This MVP still depends on configured MySQL, Redis, Milvus, and model credentials.
The mvp-demo profile mocks logs and metrics, but not the full application runtime.
Secret cleanup and fully isolated default tests are separate production-hardening tasks.
```
+11
View File
@@ -0,0 +1,11 @@
# Demo Output
This directory is the default output location for local demo responses.
Generated files are intentionally ignored by Git:
- `chat-response.json`
- `trace-response.json`
- `feedback-response.json`
Keep this README so the directory exists in the repository.
@@ -0,0 +1,4 @@
{
"Id": "mvp-demo-payment-timeout-001",
"Question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
}
@@ -0,0 +1,54 @@
param(
[string]$BaseUrl = "http://localhost:9900",
[string]$SessionId = "mvp-demo-payment-timeout-001",
[string]$RequestFile = "$PSScriptRoot/../requests/payment-timeout-chat.json",
[string]$OutputDir = "$PSScriptRoot/../output"
)
$ErrorActionPreference = "Stop"
New-Item -ItemType Directory -Force -Path $OutputDir | Out-Null
$request = Get-Content -Raw -Encoding UTF8 -Path $RequestFile | ConvertFrom-Json
$request.Id = $SessionId
$body = $request | ConvertTo-Json -Depth 8
Write-Host "Running payment-timeout chat demo..."
Write-Host "BaseUrl: $BaseUrl"
Write-Host "SessionId: $SessionId"
$chat = Invoke-RestMethod `
-Method Post `
-Uri "$BaseUrl/api/chat" `
-ContentType "application/json; charset=utf-8" `
-Body $body
$chat | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-response.json"
Write-Host "Saved chat response: $OutputDir/chat-response.json"
$trace = Invoke-RestMethod `
-Method Get `
-Uri "$BaseUrl/api/diagnosis/$SessionId/trace"
$trace | ConvertTo-Json -Depth 50 | Set-Content -Encoding UTF8 -Path "$OutputDir/trace-response.json"
Write-Host "Saved trace response: $OutputDir/trace-response.json"
$feedbackBody = @{
sessionId = $SessionId
feedback = "useful"
} | ConvertTo-Json
$feedback = Invoke-RestMethod `
-Method Post `
-Uri "$BaseUrl/api/feedback" `
-ContentType "application/json; charset=utf-8" `
-Body $feedbackBody
$feedback | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/feedback-response.json"
Write-Host "Saved feedback response: $OutputDir/feedback-response.json"
Write-Host ""
Write-Host "Demo completed. Review:"
Write-Host "- mvp/demo/output/chat-response.json"
Write-Host "- mvp/demo/output/trace-response.json"
Write-Host "- mvp/demo/output/feedback-response.json"
+52
View File
@@ -0,0 +1,52 @@
# Trace Inspection Checklist
Use this checklist after running `scripts/run-payment-timeout-demo.ps1`.
## Session
| JSON path | What to check | Interview point |
| --- | --- | --- |
| `data.session.sessionId` | Matches `mvp-demo-payment-timeout-001` | One session id connects chat, tools, verifier, feedback, and trace. |
| `data.session.query` | Contains the payment-timeout question | The trace records the original user intent. |
| `data.session.answer` | Contains the final diagnosis answer | The final answer is not detached from the trace. |
| `data.session.selfEvaluation` | Contains verifier or rule evaluation | The answer has a quality gate, not just raw model output. |
| `data.session.feedback` | Becomes `useful` after feedback submission | User feedback is attached to the same diagnosis session. |
## Agent Steps
| JSON path | What to check | Interview point |
| --- | --- | --- |
| `data.steps[*].agentName` | Planner / Executor / Verifier or equivalent step names | The flow is decomposed into inspectable Agent steps. |
| `data.steps[*].thought` | High-level step reasoning where available | Internal reasoning is auditable without relying only on final text. |
| `data.steps[*].durationMs` | Step duration | The trace can support cost and latency review. |
| `data.steps[*].tokenCount` | Token count where available | The trace can support model-cost review. |
## Tool Evidence
| JSON path | What to check | Interview point |
| --- | --- | --- |
| `data.toolInvocations[*].toolName` | Includes evidence tools such as `lookup_knowledge`, `query_logs`, `query_metrics` | The Agent uses tools, not unsupported guesses. |
| `data.toolInvocations[*].inputParams` | Shows what each tool was asked | Inputs are inspectable for debugging and audit. |
| `data.toolInvocations[*].outputPreview` | Shows a bounded preview of evidence | Evidence is preserved without dumping huge payloads. |
| `data.toolInvocations[*].success` | Distinguishes success from failure | Tool failure is visible to verifier and reviewers. |
| `data.toolInvocations[*].retrievalDetails` | Shows retrieval metadata when available | Retrieval quality can be reviewed after the fact. |
## Summary
| JSON path | What to check | Interview point |
| --- | --- | --- |
| `data.summary.persistedStepCount` | Step rows were persisted | The trace is backed by storage, not only response memory. |
| `data.summary.persistedToolCallCount` | Tool rows were persisted | Evidence survives the request. |
| `data.summary.hasVerifierEvaluation` | Verifier evaluation exists | The final answer passed through a quality gate. |
| `data.summary.hasFeedback` | Feedback exists after feedback step | Human feedback closes the loop. |
## What Good Looks Like
```text
same session id
-> final answer
-> persisted agent steps
-> persisted evidence tool calls
-> verifier/self-evaluation
-> feedback attached to the same session
```
+61
View File
@@ -0,0 +1,61 @@
# Diagnosis Eval Harness
This folder contains the first fixed-case evaluation set for the MVP diagnosis Agent.
## Scope
- Case definitions: `cases/diagnosis-cases.json`
- Offline trace fixtures: `fixtures/*.json`
- Field definitions: `schema.md`
- Baseline reports: `reports/baseline-report.json` and `reports/baseline-report.md`
- Baseline diff sample: `reports/baseline-diff-sample.json` and `reports/baseline-diff-sample.md`
- Evaluator implementation: `DiagnosisTraceEvaluator`
- Report writer: `DiagnosisEvalReportWriter`
## Current Mode
The first version evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.
## Verification
Run the focused evaluator test:
```powershell
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test
```
The committed baseline report represents the current fixed fixture set:
```text
5 fixed cases
5 passing fixture evaluations
2 PASS verdicts
3 LOW_CONFID verdicts
```
When fixtures or evaluator rules change, regenerate the report from the same case file and fixture directory, then update both JSON and Markdown outputs together.
## Interview Story
The harness gives the MVP a repeatable baseline:
```text
fixed diagnosis case
-> saved or runtime trace
-> rule-based trace validation
-> JSON / Markdown report
-> regression signal for prompts, tools, retrieval, and verifier behavior
```
## Baseline Diff
Baseline diff compares a current report against `reports/baseline-report.json`.
```text
baseline report
current report
-> deterministic diff
-> regressions, improvements, and changed signals
```
Use it to answer: did a prompt, tool, retrieval, or verifier change make the Agent worse than the fixed baseline?
+57
View File
@@ -0,0 +1,57 @@
[
{
"id": "payment-timeout",
"title": "Payment API timeout",
"question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
"traceFixture": "payment-timeout-pass.json",
"expectedRootCauseKeywords": ["支付", "超时", "连接池"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["lookup_knowledge", "query_logs", "query_metrics"],
"allowedVerdicts": ["PASS", "LOW_CONFID"],
"forbiddenAnswerKeywords": ["无证据确定"]
},
{
"id": "mysql-pool-exhausted",
"title": "MySQL connection pool exhausted",
"question": "订单服务大量请求超时,请判断是否和 MySQL 连接池有关。",
"traceFixture": "mysql-pool-low-confid.json",
"expectedRootCauseKeywords": ["mysql", "连接池", "超时"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["lookup_knowledge", "query_logs"],
"allowedVerdicts": ["LOW_CONFID", "PASS"],
"forbiddenAnswerKeywords": ["已经完全确认"]
},
{
"id": "redis-timeout",
"title": "Redis timeout",
"question": "支付服务出现 Redis 连接超时,请定位可能原因。",
"traceFixture": "redis-timeout-low-confid.json",
"expectedRootCauseKeywords": ["redis", "超时"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["query_logs"],
"allowedVerdicts": ["LOW_CONFID", "PASS"],
"forbiddenAnswerKeywords": ["无需进一步排查"]
},
{
"id": "slow-response",
"title": "Slow response",
"question": "用户服务 P99 响应时间升高,请结合指标和日志分析。",
"traceFixture": "slow-response-pass.json",
"expectedRootCauseKeywords": ["p99", "慢响应"],
"minKeywordMatches": 1,
"requiredEvidenceTools": ["query_metrics", "query_logs"],
"allowedVerdicts": ["LOW_CONFID", "PASS"],
"forbiddenAnswerKeywords": ["没有风险"]
},
{
"id": "jvm-memory-risk",
"title": "JVM memory risk",
"question": "订单服务内存使用率过高,请判断是否存在 OOM 风险。",
"traceFixture": "jvm-memory-risk-low-confid.json",
"expectedRootCauseKeywords": ["jvm", "内存", "oom"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["query_metrics", "query_logs"],
"allowedVerdicts": ["LOW_CONFID", "PASS"],
"forbiddenAnswerKeywords": ["可以忽略"]
}
]
@@ -0,0 +1,52 @@
{
"session": {
"sessionId": "eval-jvm-memory-risk",
"query": "订单服务内存使用率过高,请判断是否存在 OOM 风险。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 53000,
"toolCallCount": 2,
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n订单服务存在 JVM 内存风险,但还不能完全确认会发生 OOM。指标显示 heap 使用率持续高于 88%,日志出现多次 Full GC 和 allocation pressure,需要继续观察对象增长来源并检查最近发布。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "LOW_CONFID",
"groundedness_score": 0.52,
"tool_trace_summary": [
{
"tool_name": "query_metrics",
"success": true,
"evidence_level": "direct"
},
{
"tool_name": "query_logs",
"success": true,
"evidence_level": "indirect"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 1,
"sessionId": "eval-jvm-memory-risk",
"toolName": "query_metrics",
"success": true
},
{
"id": 2,
"sessionId": "eval-jvm-memory-risk",
"toolName": "query_logs",
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 2,
"returnedToolCallCount": 2,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,53 @@
{
"session": {
"sessionId": "eval-mysql-pool",
"query": "订单服务大量请求超时,请判断是否和 MySQL 连接池有关。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 51000,
"toolCallCount": 2,
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\nMySQL 连接池可能参与了本次超时问题。日志中出现 connection pool exhausted,但当前缺少完整指标证据,因此只能作为低置信结论处理。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "LOW_CONFID",
"groundedness_score": 0.48,
"tool_trace_summary": [
{
"tool_name": "lookup_knowledge",
"success": true,
"evidence_level": "direct"
},
{
"tool_name": "query_logs",
"success": true,
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 1,
"sessionId": "eval-mysql-pool",
"toolName": "lookup_knowledge",
"success": true,
"relevanceLevel": "PRECISE"
},
{
"id": 2,
"sessionId": "eval-mysql-pool",
"toolName": "query_logs",
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 2,
"returnedToolCallCount": 2,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,64 @@
{
"session": {
"sessionId": "eval-payment-timeout",
"query": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 42000,
"toolCallCount": 3,
"answer": "支付接口超时与连接池等待有关。知识库说明支付超时需要同时检查连接池、日志和指标;日志出现 connection pool exhausted;指标显示支付服务延迟升高。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "PASS",
"groundedness_score": 0.86,
"tool_trace_summary": [
{
"tool_name": "lookup_knowledge",
"success": true,
"evidence_level": "direct"
},
{
"tool_name": "query_logs",
"success": true,
"evidence_level": "direct"
},
{
"tool_name": "query_metrics",
"success": true,
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 1,
"sessionId": "eval-payment-timeout",
"toolName": "lookup_knowledge",
"success": true,
"relevanceLevel": "PRECISE"
},
{
"id": 2,
"sessionId": "eval-payment-timeout",
"toolName": "query_logs",
"success": true
},
{
"id": 3,
"sessionId": "eval-payment-timeout",
"toolName": "query_metrics",
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 3,
"returnedToolCallCount": 3,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,41 @@
{
"session": {
"sessionId": "eval-redis-timeout",
"query": "支付服务出现 Redis 连接超时,请定位可能原因。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 36000,
"toolCallCount": 1,
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\nRedis 连接超时可能和支付服务到 Redis 的网络抖动或连接池等待有关。日志中出现 redis timeout 和 command timeout 记录,但当前缺少指标侧证据,因此只能作为低置信结论处理。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "LOW_CONFID",
"groundedness_score": 0.46,
"tool_trace_summary": [
{
"tool_name": "query_logs",
"success": true,
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 1,
"sessionId": "eval-redis-timeout",
"toolName": "query_logs",
"success": true
}
],
"summary": {
"persistedStepCount": 2,
"returnedStepCount": 2,
"persistedToolCallCount": 1,
"returnedToolCallCount": 1,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
+52
View File
@@ -0,0 +1,52 @@
{
"session": {
"sessionId": "eval-slow-response",
"query": "用户服务 P99 响应时间升高,请结合指标和日志分析。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 47000,
"toolCallCount": 2,
"answer": "用户服务 P99 升高主要表现为慢响应。指标显示 P99 latency 从 280ms 上升到 1800ms,日志中同时出现 slow request 和 downstream timeout,因此优先排查下游依赖耗时和线程池排队。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "PASS",
"groundedness_score": 0.78,
"tool_trace_summary": [
{
"tool_name": "query_metrics",
"success": true,
"evidence_level": "direct"
},
{
"tool_name": "query_logs",
"success": true,
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 1,
"sessionId": "eval-slow-response",
"toolName": "query_metrics",
"success": true
},
{
"id": 2,
"sessionId": "eval-slow-response",
"toolName": "query_logs",
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 2,
"returnedToolCallCount": 2,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,85 @@
{
"baselineTotalCases" : 5,
"currentTotalCases" : 5,
"baselinePassedCases" : 5,
"currentPassedCases" : 4,
"baselinePassRate" : 1.0,
"currentPassRate" : 0.8,
"regressionCount" : 6,
"improvementCount" : 0,
"changedCount" : 2,
"hasRegression" : true,
"items" : [ {
"type" : "REGRESSION",
"scope" : "aggregate",
"caseId" : null,
"metric" : "passRate",
"baselineValue" : "1.0",
"currentValue" : "0.8",
"delta" : -0.19999999999999996,
"message" : "passRate changed"
}, {
"type" : "REGRESSION",
"scope" : "aggregate",
"caseId" : null,
"metric" : "averageToolCallCount",
"baselineValue" : "2.0",
"currentValue" : "3.0",
"delta" : 1.0,
"message" : "averageToolCallCount changed"
}, {
"type" : "CHANGED",
"scope" : "aggregate",
"caseId" : null,
"metric" : "verdictDistribution.LOW_CONFID",
"baselineValue" : "3",
"currentValue" : "2",
"delta" : -1.0,
"message" : "verdict count changed for LOW_CONFID"
}, {
"type" : "CHANGED",
"scope" : "aggregate",
"caseId" : null,
"metric" : "verdictDistribution.REJECT",
"baselineValue" : "0",
"currentValue" : "1",
"delta" : 1.0,
"message" : "verdict count changed for REJECT"
}, {
"type" : "REGRESSION",
"scope" : "case",
"caseId" : "redis-timeout",
"metric" : "passed",
"baselineValue" : "true",
"currentValue" : "false",
"delta" : null,
"message" : "redis-timeout pass state changed"
}, {
"type" : "REGRESSION",
"scope" : "case",
"caseId" : "redis-timeout",
"metric" : "verdict",
"baselineValue" : "LOW_CONFID",
"currentValue" : "REJECT",
"delta" : -1.0,
"message" : "redis-timeout verdict changed"
}, {
"type" : "REGRESSION",
"scope" : "case",
"caseId" : "redis-timeout",
"metric" : "matchedKeywordCount",
"baselineValue" : "2",
"currentValue" : "1",
"delta" : -1.0,
"message" : "redis-timeout matchedKeywordCount changed"
}, {
"type" : "REGRESSION",
"scope" : "case",
"caseId" : "redis-timeout",
"metric" : "evidenceCoverage.query_logs",
"baselineValue" : "true",
"currentValue" : "false",
"delta" : null,
"message" : "redis-timeout evidence coverage changed for query_logs"
} ]
}
+22
View File
@@ -0,0 +1,22 @@
# Diagnosis Eval Baseline Diff
- Baseline pass rate: 100.00%
- Current pass rate: 80.00%
- Baseline passed cases: 5/5
- Current passed cases: 4/5
- Regressions: 6
- Improvements: 0
- Other changes: 2
## Diff Items
| Type | Scope | Case | Metric | Baseline | Current | Delta | Message |
| --- | --- | --- | --- | --- | --- | ---: | --- |
| REGRESSION | aggregate | - | passRate | 1.0 | 0.8 | -0.200 | passRate changed |
| REGRESSION | aggregate | - | averageToolCallCount | 2.0 | 3.0 | 1.000 | averageToolCallCount changed |
| CHANGED | aggregate | - | verdictDistribution.LOW_CONFID | 3 | 2 | -1.000 | verdict count changed for LOW_CONFID |
| CHANGED | aggregate | - | verdictDistribution.REJECT | 0 | 1 | 1.000 | verdict count changed for REJECT |
| REGRESSION | case | redis-timeout | passed | true | false | - | redis-timeout pass state changed |
| REGRESSION | case | redis-timeout | verdict | LOW_CONFID | REJECT | -1.000 | redis-timeout verdict changed |
| REGRESSION | case | redis-timeout | matchedKeywordCount | 2 | 1 | -1.000 | redis-timeout matchedKeywordCount changed |
| REGRESSION | case | redis-timeout | evidenceCoverage.query_logs | true | false | - | redis-timeout evidence coverage changed for query_logs |
+82
View File
@@ -0,0 +1,82 @@
{
"totalCases" : 5,
"passedCases" : 5,
"passRate" : 1.0,
"verdictDistribution" : {
"PASS" : 2,
"LOW_CONFID" : 3
},
"averageToolCallCount" : 2.0,
"averageDurationMs" : 45800.0,
"results" : [ {
"caseId" : "payment-timeout",
"title" : "Payment API timeout",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "PASS",
"matchedKeywordCount" : 3,
"requiredKeywordCount" : 3,
"evidenceCoverage" : {
"lookup_knowledge" : true,
"query_logs" : true,
"query_metrics" : true
},
"toolCallCount" : 3,
"durationMs" : 42000
}, {
"caseId" : "mysql-pool-exhausted",
"title" : "MySQL connection pool exhausted",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "LOW_CONFID",
"matchedKeywordCount" : 3,
"requiredKeywordCount" : 3,
"evidenceCoverage" : {
"lookup_knowledge" : true,
"query_logs" : true
},
"toolCallCount" : 2,
"durationMs" : 51000
}, {
"caseId" : "redis-timeout",
"title" : "Redis timeout",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "LOW_CONFID",
"matchedKeywordCount" : 2,
"requiredKeywordCount" : 2,
"evidenceCoverage" : {
"query_logs" : true
},
"toolCallCount" : 1,
"durationMs" : 36000
}, {
"caseId" : "slow-response",
"title" : "Slow response",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "PASS",
"matchedKeywordCount" : 2,
"requiredKeywordCount" : 2,
"evidenceCoverage" : {
"query_metrics" : true,
"query_logs" : true
},
"toolCallCount" : 2,
"durationMs" : 47000
}, {
"caseId" : "jvm-memory-risk",
"title" : "JVM memory risk",
"passed" : true,
"failedChecks" : [ ],
"verdict" : "LOW_CONFID",
"matchedKeywordCount" : 3,
"requiredKeywordCount" : 3,
"evidenceCoverage" : {
"query_metrics" : true,
"query_logs" : true
},
"toolCallCount" : 2,
"durationMs" : 53000
} ]
}
+22
View File
@@ -0,0 +1,22 @@
# Diagnosis Eval Report
- Total cases: 5
- Passed cases: 5
- Pass rate: 100.00%
- Average tool calls: 2.00
- Average duration ms: 45800.00
## Verdict Distribution
- PASS: 2
- LOW_CONFID: 3
## Cases
| Case | Result | Verdict | Keywords | Tool Calls | Duration ms | Failed Checks |
| --- | --- | --- | --- | ---: | ---: | --- |
| payment-timeout | PASS | PASS | 3/3 | 3 | 42000 | - |
| mysql-pool-exhausted | PASS | LOW_CONFID | 3/3 | 2 | 51000 | - |
| redis-timeout | PASS | LOW_CONFID | 2/2 | 1 | 36000 | - |
| slow-response | PASS | PASS | 2/2 | 2 | 47000 | - |
| jvm-memory-risk | PASS | LOW_CONFID | 3/3 | 2 | 53000 | - |
+202
View File
@@ -0,0 +1,202 @@
# Diagnosis Eval Data Schema
这份文档记录评测基准里的数据结构。口语化理解就是:
```text
用例文件说“我要考什么”
trace 文件说“Agent 实际做了什么”
评测结果说“这次有没有跑偏”
汇总报告说“整体稳定性怎么样”
```
当前这套评测是代码规则判断,不是再调用一个 LLM 来打分。
## 1. 用例定义
文件:`mvp/eval/cases/diagnosis-cases.json`
每一条 case 是一个固定考题,告诉评测器“这个问题应该看哪些点、需要哪些证据、哪些结论可以接受”。
```json
{
"id": "payment-timeout",
"title": "Payment API timeout",
"question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
"traceFixture": "payment-timeout-pass.json",
"expectedRootCauseKeywords": ["支付", "超时", "连接池"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["lookup_knowledge", "query_logs", "query_metrics"],
"allowedVerdicts": ["PASS", "LOW_CONFID"],
"forbiddenAnswerKeywords": ["无证据确定"]
}
```
字段说明:
| 字段 | 意思 | 评测器怎么用 |
| --- | --- | --- |
| `id` | 这条用例的唯一名字 | 出现在报告里,方便定位是哪条 case 挂了 |
| `title` | 给人看的标题 | 出现在结果里,方便快速理解场景 |
| `question` | 要问 Agent 的问题 | fixture 模式下不会真的发送给 Agent,但它记录了这条 case 的原始输入 |
| `traceFixture` | 对应的 trace 文件名 | 评测器会去 `fixtures/` 目录加载这个文件 |
| `expectedRootCauseKeywords` | 最终回答里希望看到的关键点 | 评测器会在 `session.answer` 里做关键词命中检查 |
| `minKeywordMatches` | 至少要命中几个关键词 | 命中数低于这个值,就认为根因覆盖不够 |
| `requiredEvidenceTools` | 这条 case 至少应该用到哪些证据工具 | 评测器会检查 trace 里是否出现这些工具 |
| `allowedVerdicts` | Verifier 允许给出的结论 | 比如 `PASS` 或 `LOW_CONFID`,不在列表里就失败 |
| `forbiddenAnswerKeywords` | 回答里不应该出现的危险说法 | 命中这些词,说明回答可能过度自信或不符合降级策略 |
## 2. Trace Fixture
目录:`mvp/eval/fixtures/*.json`
trace fixture 是一次 Agent 运行后的“留痕快照”。评测器不会关心整个 trace 的所有字段,只读取当前能支撑基准判断的字段。
当前会读取这些字段:
| Trace 字段 | 意思 | 评测器怎么用 |
| --- | --- | --- |
| `session.answer` | Agent 最终给用户的回答 | 用来检查根因关键词和禁用词 |
| `session.totalDurationMs` | 这次运行耗时 | 进入报告,帮助观察性能是否明显变差 |
| `session.selfEvaluation.verifier_evaluation.verdict` | Verifier 对最终回答的判断 | 必须存在,并且要落在 case 的 `allowedVerdicts` 里 |
| `session.selfEvaluation.verifier_evaluation.tool_trace_summary[*].tool_name` | Verifier 总结里看到的工具证据 | 用来补充判断证据工具是否出现 |
| `toolInvocations[*].toolName` | Agent 实际调用过的工具名 | 用来检查 `requiredEvidenceTools` 是否满足 |
| `toolInvocations[*].success` | 工具调用是否成功 | 当前主要保留在 trace 里,后续可以升级成更严格的成功率检查 |
简单说,trace 里最重要的是三类信息:
```text
最终回答:它说了什么
工具证据:它查了什么
Verifier:它自己有没有承认这个结论可靠
```
## 3. 单条评测结果
Java 类型:`DiagnosisEvalResult`
这是每条 case 跑完之后的判断结果。
| 字段 | 意思 |
| --- | --- |
| `caseId` | 对应的 case id |
| `title` | case 标题 |
| `passed` | 这条 case 是否通过 |
| `failedChecks` | 没通过的具体原因,比如缺工具、关键词不够、verdict 不允许 |
| `verdict` | 从 trace 里读出来的 Verifier verdict |
| `matchedKeywordCount` | 最终回答命中的关键词数量 |
| `requiredKeywordCount` | case 定义里一共有多少个关键词 |
| `evidenceCoverage` | 每个必需工具是否出现,例如 `{ "query_logs": true }` |
| `toolCallCount` | 本次 trace 里工具调用总数 |
| `durationMs` | 本次 trace 的耗时 |
判断通过的口语化规则:
```text
回答要说到关键点
该查的证据工具要查到
Verifier 的结论要在可接受范围内
回答不能出现危险的过度自信表达
如果是 REJECT,就必须走降级模板
```
## 4. 汇总报告
Java 类型:`DiagnosisEvalReport`
这是整个基准集跑完之后的总结果。
| 字段 | 意思 |
| --- | --- |
| `totalCases` | 总共评测了多少条 case |
| `passedCases` | 通过了多少条 |
| `passRate` | 通过率,范围是 `0.0` 到 `1.0` |
| `verdictDistribution` | Verifier verdict 的分布,比如有几个 `PASS`、几个 `LOW_CONFID` |
| `averageToolCallCount` | 平均每条 case 调用了多少次工具 |
| `averageDurationMs` | 平均耗时 |
| `results` | 每条 case 的详细结果列表 |
## 5. 怎么看这个基准
这套结构不是为了证明 Agent 永远正确,而是为了在每次改 prompt、工具、检索、Verifier 之后,有一个固定尺子能回答:
```text
以前能过的诊断题,现在还过不过?
它是不是少查了某些证据?
它是不是变得更自信但证据不足?
它是不是开始输出不该说的话?
它是不是明显变慢了?
```
所以面试里可以这样讲:
```text
我没有只看一次 demo 效果,而是把典型诊断场景固化成 case。
每条 case 都定义预期关键点、必需证据工具和可接受的 verifier 结论。
Agent 每次运行会留下 trace,评测器用代码规则读取 trace,输出结构化报告。
这样我改 Agent 的时候,可以用同一套基准判断有没有行为回退。
```
## 6. Baseline Diff
Baseline diff 是拿两份 report 做对比:
```text
baseline report:以前认可的基准结果
current report:这次改动后跑出来的新结果
diff report:告诉你哪里变好了、哪里变差了、哪里只是变了
```
Java 类型:
- `DiagnosisEvalDiffReport`
- `DiagnosisEvalDiffItem`
`DiagnosisEvalDiffReport` 字段:
| 字段 | 意思 |
| --- | --- |
| `baselineTotalCases` | baseline 里有多少条 case |
| `currentTotalCases` | current 里有多少条 case |
| `baselinePassedCases` | baseline 通过了多少条 |
| `currentPassedCases` | current 通过了多少条 |
| `baselinePassRate` | baseline 通过率 |
| `currentPassRate` | current 通过率 |
| `regressionCount` | 退化项数量 |
| `improvementCount` | 改善项数量 |
| `changedCount` | 普通变化项数量 |
| `hasRegression` | 是否存在退化 |
| `items` | 具体 diff 明细 |
`DiagnosisEvalDiffItem` 字段:
| 字段 | 意思 |
| --- | --- |
| `type` | `REGRESSION`、`IMPROVEMENT` 或 `CHANGED` |
| `scope` | `aggregate` 表示整体指标,`case` 表示单条 case |
| `caseId` | 如果是单条 case 变化,这里记录 case id |
| `metric` | 哪个指标变了,比如 `passRate` 或 `evidenceCoverage.query_logs` |
| `baselineValue` | baseline 里的值 |
| `currentValue` | current 里的值 |
| `delta` | 数值变化量;非数值变化为空 |
| `message` | 给人看的变化说明 |
口语化判断规则:
```text
pass rate 下降:退化
case 从通过变失败:退化
证据工具从有变没有:退化
关键词命中变少:退化
工具调用或耗时升高:成本上升,记为退化信号
verdict 分布变化:记录变化,供人工判断是否符合预期
```
面试里可以这样讲:
```text
我把 baseline report 和当前 report 做结构化 diff。
它不是再问 LLM,而是用代码比较固定字段。
如果某个 case 从 PASS 变 FAIL,或者 query_logs 证据没了,
diff 会直接标成 regression。
这样 Agent 改动可以用固定基准做回归判断。
```
@@ -0,0 +1,137 @@
# ISS-005 证据链补齐与降级契约收敛
**状态**:进行中(sm-flow)
**严重程度**:高
**发现时间**:2026-07-04
**来源**:P1-A 面试打磨项 / 基于 ISS-003 的当前实现复核
**关联**:ISS-003(Verifier 证据链、失败路径可验证性)、`chat-verifier-agent`、`mvp-demo-trace-acceptance`
---
## 背景
当前 MVP 已具备:
- `lookup_knowledge`、`query_logs`、`query_metrics` 的工具调用落库
- Verifier 基于 `tool_trace_summary` 做事实核查
- `LOW_CONFID` / `REJECT` 的用户侧降级输出
- trace API 可回放 session、agent_step、tool_invocation 和 self_evaluation
但如果目标是拿这个项目去面试 Agent 工程师,当前实现仍有一个明显短板:
**证据链已经“有了”,但还没有被收敛成清晰、稳定、可测试的工程契约。**
这会直接影响三个面试问题的回答质量:
1. 工具失败时系统会怎样降级?
2. Verifier 看到的 evidence 到底是否一致、可审计?
3. 这些失败路径和降级行为有没有稳定测试,而不是只靠 runtime 演示?
---
## 当前现状复核
### 1. 工具落库入口已经存在,但契约不统一
- `QueryLogsTools` 和 `QueryMetricsTools` 通过 `ToolInvocationRecorder.recordEvidenceTool(...)` 记录 evidence tool 调用。
- `LookupKnowledgeTool` 仍保留独立的 `saveToolInvocation(...)` 路径,自己构造 `ToolInvocation` 实体。
这意味着:
- evidence tool 的公共字段有一套约定
- knowledge retrieval 又有一套定制字段拼装
两者都能工作,但**没有形成统一的“证据调用记录契约”**。
### 2. 失败 / 无结果 / 去重命中的语义不够显式
当前实现里:
- `query_logs` 未命中时会返回 `success=false` + `"未找到匹配的日志"`
- `query_metrics` 失败时会返回 `success=false`
- `lookup_knowledge` 去重命中时会返回 `found=false`,但 `tool_invocation.success=true`
- `ToolTraceSummaryService` 通过 `success`、`relevanceLevel`、`dedupReason` 等字段做启发式摘要
这些行为在代码里是分散成立的,但**没有被定义成统一契约**,导致:
- Verifier 能看到的“失败”和“无证据”边界不够稳定
- 评测时难以明确统计哪些是“调用失败”、哪些是“无命中”、哪些是“已检索过”
### 3. ChatService 的降级路径有实现,但测试矩阵不完整
`ChatService` 已处理:
- `verifier_output` 缺失或无法解析 → fallback `LOW_CONFID`
- `REJECT` → degraded output
- `LOW_CONFID` → disclaimer output
但目前缺少成体系的专项验证,尤其是:
- Verifier 输出非法 JSON
- evidence tool 查询失败
- knowledge lookup 无有效证据
- fallback 文案是否只基于 verifier 缺口拼装
---
## 影响
- **面试表达弱化**:你能讲“我有 trace”,但还不能很硬地讲“我的失败路径是有契约和测试保护的”。
- **评测基础不稳**:后续 P1-B 做 case-based harness 时,统计口径会受 evidence 语义不一致影响。
- **Verifier 可审计性打折**:当前实现可用,但 still relies on code convention,而不是一份明确收敛后的工程协议。
---
## 本 issue 目标
P1-A 只做三件事:
1. 收敛 evidence tool 的落库契约,让 `lookup_knowledge`、`query_logs`、`query_metrics` 的公共语义一致。
2. 明确失败 / 无证据 / 去重 / verifier 非法输出等降级契约,让 `ToolTraceSummaryService` 和 `ChatService` 面向统一状态工作。
3. 增加专项离线测试,覆盖证据摘要与关键降级路径。
---
## 范围
### In scope
- `ToolInvocationRecorder` 契约增强
- `LookupKnowledgeTool` 入库路径收敛
- `QueryLogsTools` / `QueryMetricsTools` evidence 语义对齐
- `ToolTraceSummaryService` 对失败 / no-hit / mixed evidence 的摘要规则收敛
- `ChatService` 对 verifier 非法输出与降级输出的专项测试
- 与该 change 直接相关的文档、OpenSpec、devflow 记录
### Out of scope
- 不引入新的数据库表或 schema 变更
- 不扩展新的 evidence tool
- 不做 P1-B 评测集 / harness
- 不做前端 trace UI
- 不处理敏感配置和默认 `mvn test` 离线化
---
## 预期结果
完成后,项目在面试里应能更清楚地表述为:
```text
我不仅把 Agent 的工具调用落到了库里,
还把 evidence trace、失败语义和 verifier 降级路径收敛成了稳定契约,
并用离线测试覆盖了这些关键失败场景。
```
---
## 相关文件
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
- `src/main/java/com/superbiz/agent/agent/tool/QueryLogsTools.java`
- `src/main/java/com/superbiz/agent/agent/tool/QueryMetricsTools.java`
- `src/main/java/com/superbiz/agent/service/ToolTraceSummaryService.java`
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`
- `src/test/java/com/superbiz/agent/tool/LookupKnowledgeToolTest.java`
@@ -0,0 +1,99 @@
# ISS-006 固定诊断评测集与回归 Harness
**状态**:进行中(sm-flow)
**严重程度**:高
**发现时间**:2026-07-04
**来源**:P1-B 面试打磨项
**依赖**:ISS-005 / `evidence-trace-hardening`
---
## 背景
MVP 已经具备可追溯证据链、Verifier 质量门禁、trace API 和固定 demo 流程。上一阶段 `evidence-trace-hardening` 进一步统一了 evidence tool 的状态语义,让系统能稳定区分:
- `supported`
- `no_evidence`
- `deduped`
- `failed`
下一步需要证明 Agent 在一组固定诊断场景下的表现,而不是只依赖单次 demo。
---
## 问题
当前项目能演示一次支付超时诊断,但还缺少稳定的评测基线:
- 每次改 prompt、工具、Verifier 或检索逻辑后,无法快速判断是否退化。
- 只能人工看 trace,缺少结构化通过 / 失败结果。
- 缺少面试时能展示的指标,如 evidence coverage、verdict 分布、工具调用数量和耗时。
---
## 目标
建立一个轻量的固定 case 评测 harness,用于验证 MVP Agent 的诊断质量和证据链完整性。
第一版不做 LLM-as-judge,优先做规则化校验:
- 固定 5 个 MVP 诊断 case
- 每个 case 定义 expected root-cause keywords、required evidence tools、allowed verdicts
- 基于 trace 结果校验 evidence coverage、verifier evaluation、tool invocation、final answer shape
- 输出 JSON 和 Markdown 报告
---
## 范围
### In scope
- 评测 case 定义文件
- trace 规则校验器
- eval runner 或测试入口
- JSON / Markdown 报告输出
- demo 文档和 devflow 记录
### Out of scope
- 不引入 LLM-as-judge
- 不要求完整离线 LLM runtime
- 不新增生产 API
- 不修改 Chat 主链路
- 不修改 evidence trace 运行时语义
---
## 预期面试表达
完成后可以这样描述:
```text
我不仅有一个可演示的 Agent,还给它建立了固定 case 的回归评测。
每次修改 prompt、工具或 verifier 后,都可以跑同一批诊断 case,
检查证据覆盖、verdict 分布、工具调用成本和关键结论是否退化。
```
---
## 初始候选 case
| Case | 目标 |
| --- | --- |
| payment-timeout | 支付接口超时,验证知识库 + 日志 + 指标证据 |
| mysql-pool-exhausted | 数据库连接池耗尽,验证日志和知识库证据 |
| redis-timeout | Redis 连接超时,验证日志依赖证据 |
| slow-response | P99 响应时间过高,验证指标 + 慢请求日志 |
| jvm-memory-risk | JVM 内存 / OOM 风险,验证指标 + 系统事件日志 |
---
## 相关文件
- `mvp/demo/README.md`
- `mvp/demo/payment-timeout-acceptance.md`
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
- `src/main/java/com/superbiz/agent/service/ToolTraceSummaryService.java`
- `src/main/java/com/superbiz/agent/domain/entity/DiagnosisSession.java`
- `src/main/java/com/superbiz/agent/domain/entity/ToolInvocation.java`
- `openspec/specs/evidence-trace-hardening/spec.md`
+34
View File
@@ -6,3 +6,37 @@
| ISS-002 | Executor 无约束重复调用 lookup_knowledge | 中 | 已修复 | [ISS-002-executor-unconstrained-lookup.md](ISS-002-executor-unconstrained-lookup.md) |
| ISS-003 | MVP 设计与实现 Review 收敛 | 高 | 待规划 | [ISS-003-mvp-design-implementation-review.md](ISS-003-mvp-design-implementation-review.md) |
| ISS-004 | Executor 域级检索水位控制(Phase 2) | 低 | 待规划 | [ISS-004-executor-domain-hard-limit.md](ISS-004-executor-domain-hard-limit.md) |
| ISS-005 | 证据链补齐与降级契约收敛 | 高 | 已归档 | [ISS-005-evidence-trace-hardening.md](ISS-005-evidence-trace-hardening.md) |
| ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) |
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) |
| diagnosis-eval-baseline-diff | 诊断评测 baseline diff 与回归判断 | 中 | 已归档 | [diagnosis-eval-baseline-diff.md](diagnosis-eval-baseline-diff.md) |
| mvp-demo-interview-runbook | Plan C 面试可复现 Demo 包 | 中 | 已归档 | [mvp-demo-interview-runbook.md](mvp-demo-interview-runbook.md) |
## RAG 重构计划
| 名称 | 标题 | 严重程度 | 状态 | 文件 |
|---|---|---|---|---|
| rag-refactor-plan | RAG 检索重构计划 | 高 | 待规划 | [rag-refactor-plan.md](rag-refactor-plan.md) |
## RAG 检索问题
| 名称 | 标题 | 严重程度 | 状态 | 文件 |
|---|---|---|---|---|
| chunk-context-reconstruction | RAG 切片上下文重建缺失 | 高 | 已合并到重构计划 | [rag-chunk-context-reconstruction.md](rag-chunk-context-reconstruction.md) |
| breadcrumb-embedding-gap | RAG breadcrumb 未参与向量语义 | 高 | 已合并到重构计划 | [rag-breadcrumb-embedding-gap.md](rag-breadcrumb-embedding-gap.md) |
| l0-l1-fusion-ranking | RAG L0 和 L1 未真正融合排序 | 中 | 已合并到重构计划 | [rag-l0-l1-fusion-ranking.md](rag-l0-l1-fusion-ranking.md) |
| l0-keyword-matching-quality | RAG L0 关键词匹配质量不足 | 中 | 已合并到重构计划 | [rag-l0-keyword-matching-quality.md](rag-l0-keyword-matching-quality.md) |
| l1-score-calibration | RAG L1 分数阈值未校准 | 中 | 已合并到重构计划 | [rag-l1-score-calibration.md](rag-l1-score-calibration.md) |
| context-packing-and-reranking | RAG 缺少上下文打包和 Rerank | 中 | 已合并到重构计划 | [rag-context-packing-and-reranking.md](rag-context-packing-and-reranking.md) |
| upload-chunk-parameter-drift | RAG 上传切片参数未真正生效 | 低 | 已合并到重构计划 | [rag-upload-chunk-parameter-drift.md](rag-upload-chunk-parameter-drift.md) |
| query-rewrite-gap | RAG 查询改写能力薄弱 | 中 | 已合并到重构计划 | [rag-query-rewrite-gap.md](rag-query-rewrite-gap.md) |
## RAG 框架化改造
| 名称 | 标题 | 严重程度 | 状态 | 文件 |
|---|---|---|---|---|
| spring-ai-vectorstore-migration | RAG 迁移到 Spring AI VectorStore 检索抽象 | 高 | 已合并到重构计划 | [rag-spring-ai-vectorstore-migration.md](rag-spring-ai-vectorstore-migration.md) |
| spring-ai-query-transformer | RAG 接入 Spring AI Query Transformer | 中 | 已合并到重构计划 | [rag-spring-ai-query-transformer.md](rag-spring-ai-query-transformer.md) |
| spring-ai-document-postprocessor | RAG 使用 DocumentPostProcessor 做后处理 | 中 | 已合并到重构计划 | [rag-spring-ai-document-postprocessor.md](rag-spring-ai-document-postprocessor.md) |
| l0-domain-entity-hint | RAG 将 L0 降级为领域和实体 Hint | 中 | 已合并到重构计划 | [rag-l0-domain-entity-hint.md](rag-l0-domain-entity-hint.md) |
| spring-ai-advisor-boundary | RAG 明确 Spring AI Advisor 与 Agent Tool 的边界 | 中 | 已合并到重构计划 | [rag-spring-ai-advisor-boundary.md](rag-spring-ai-advisor-boundary.md) |
@@ -0,0 +1,73 @@
# Diagnosis Eval Baseline Diff
**状态**:已归档
**严重程度**:中
**发现时间**:2026-07-05
**来源**:P1-B follow-up
**依赖**:`diagnosis-eval-harness`, `expand-diagnosis-eval-fixtures`
---
## 背景
现在项目已经有固定诊断 case、完整 fixture 和 baseline report。下一步需要把 baseline 真正用起来:每次改 Agent 后,把新的 report 和 baseline report 做对比。
---
## 问题
当前 baseline 只能告诉我们“标准状态是什么”,但还不能自动告诉我们“这次改动有没有变差”。
典型问题包括:
- pass rate 是否下降。
- 某个 case 是否从通过变失败。
- 某个 evidence tool 是否从覆盖变成缺失。
- verifier verdict 分布是否异常变化。
- 平均工具调用数和耗时是否明显上升。
---
## 目标
新增一个 deterministic baseline diff 能力,用代码比较两份 `DiagnosisEvalReport`。
完成后应该做到:
- 输入 baseline report 和 current report。
- 输出结构化 diff。
- 标出 regression、improvement 和普通 changed。
- 支持 JSON 和 Markdown 输出。
- 文档说明面试时怎么解释这套回归判断。
---
## 范围
### In scope
- report-level diff 数据结构。
- aggregate 指标比较。
- case-level 指标比较。
- JSON / Markdown diff writer。
- focused tests 和 eval 文档。
### Out of scope
- 不运行真实 Agent。
- 不生成新 trace。
- 不引入 LLM-as-judge。
- 不改现有 evaluator 评分规则。
---
## 面试表达
可以这样讲:
```text
我不是只保存了一份 baseline,而是加了 baseline diff。
每次改 prompt、tool、retrieval 或 verifier 后,
我都能把新 report 和 baseline 比较,
直接看到哪些 case 退化、哪些证据缺失、成本有没有上升。
```
@@ -0,0 +1,83 @@
# Expand Diagnosis Eval Fixtures
**状态**:已归档
**严重程度**:中
**发现时间**:2026-07-04
**来源**:P1-B follow-up
**依赖**:`diagnosis-eval-harness`
---
## 背景
`diagnosis-eval-harness` 已经把固定 case、trace evaluator、JSON / Markdown report 和字段文档搭起来了。
现在还差一步:5 条固定诊断 case 里,只有 2 条有 fixture,另外 3 条还是 missing 状态。这个状态可以验证 evaluator 的错误报告能力,但还不能作为完整 baseline 展示。
---
## 问题
当前 baseline 还不够完整:
- `redis-timeout` 没有对应 trace fixture。
- `slow-response` 没有对应 trace fixture。
- `jvm-memory-risk` 没有对应 trace fixture。
- 仓库里还没有一份固定的 baseline JSON / Markdown 报告可供对比。
---
## 目标
补齐固定诊断评测集,让它从“框架可跑”变成“基准可用”。
完成后应该做到:
- 5 条固定 case 都能加载到对应 fixture。
- evaluator 能输出完整 baseline report。
- baseline report 被保存到仓库,后续 Agent 改动可以拿它做对比。
- 文档说明怎么重新生成和怎么看报告。
---
## 范围
### In scope
- 补齐 3 个缺失 fixture。
- 保存 baseline JSON / Markdown 报告。
- 更新 eval 文档。
- 补充测试,确保 case 文件引用的 fixture 都存在。
### Out of scope
- 不新增 case 数量。
- 不改生产 Agent 主链路。
- 不引入 LLM-as-judge。
- 不启动真实 MySQL、Redis、Milvus 或 LLM。
---
## 面试表达
可以这样讲:
```text
我先搭了评测 harness,然后把固定 case 的 trace fixture 补齐,
生成一份可复现的 baseline report。
这样以后每次改 prompt、tool 或 verifier,
都能看固定诊断集有没有行为回退,而不是只靠人工感觉。
```
---
## 相关文件
- `mvp/eval/cases/diagnosis-cases.json`
- `mvp/eval/fixtures/`
- `mvp/eval/reports/`
- `mvp/eval/README.md`
- `mvp/eval/schema.md`
- `src/main/java/com/superbiz/agent/eval/DiagnosisTraceEvaluator.java`
- `src/test/java/com/superbiz/agent/eval/DiagnosisTraceEvaluatorTest.java`
- `openspec/specs/diagnosis-eval-harness/spec.md`
+53
View File
@@ -0,0 +1,53 @@
# MVP Demo Interview Runbook
**状态**:已归档
**严重程度**:中
**发现时间**:2026-07-05
**来源**:Plan C
**依赖**:`mvp-demo-trace-acceptance`, `evidence-trace-hardening`, `diagnosis-eval-harness`
---
## 背景
项目已经有 Agent 主链路、证据 trace、Verifier、反馈、eval baseline,但这些材料分散在不同目录。面试时真正需要的是一个能快速跑、快速讲清楚的 demo 入口。
---
## 问题
当前 demo 还不够“面试友好”:
- 启动、请求、trace、反馈步骤分散在文档里。
- 没有固定请求 payload 文件。
- 没有一键跑 payment-timeout demo 的脚本。
- 没有把 trace 字段和面试讲法对应起来的 walkthrough。
---
## 目标
把 Plan C 落地成 `mvp/demo` 下的可复现 demo 包:
- 固定支付超时请求。
- 一键执行 chat、trace、feedback。
- 保存 demo 输出,便于复盘。
- 提供面试讲解稿和 trace 检查清单。
---
## 范围
### In scope
- `mvp/demo` 文档。
- `mvp/demo/requests` 请求文件。
- `mvp/demo/scripts` PowerShell 脚本。
- `mvp/demo/output` 目录说明。
### Out of scope
- 不新增后端 API。
- 不改 Agent prompt。
- 不扩 eval harness。
- 不处理密钥外置和完整离线化。
@@ -0,0 +1,58 @@
# RAG breadcrumb 未参与向量语义
**状态**:待规划
**严重程度**:高
**发现时间**:2026-07-04
**范围**:向量化输入、检索相关性、知识库 metadata 使用
---
## 现象
当前 chunk metadata 中保存了 `title` 和 `breadcrumb`,但向量化时主要使用 `chunk.getContent()`。这意味着标题层级、所属模块、章节路径没有进入 embedding 语义空间。
当用户问题依赖章节语境时,例如“诊断流程里的验证步骤是什么”,如果 chunk 正文里没有重复出现完整标题语义,向量召回可能无法稳定命中正确片段。
---
## 当前实现
- `DocumentChunkService` 会生成 `breadcrumb`。
- `VectorIndexService` 会把 `breadcrumb` 写入 metadata。
- `VectorEmbeddingService` 接收的 embedding 内容来自 chunk 正文。
- `VectorSearchService` 只基于 query embedding 和 chunk embedding 做向量搜索。
metadata 目前更像是展示和追踪字段,不是检索相关性的一部分。
---
## 影响
- 标题语义丢失,尤其影响短段落、步骤列表、配置表格类 chunk。
- 同名概念出现在不同章节时,缺少章节路径帮助 disambiguation。
- 用户问的是“某个模块下的问题”,检索可能只看正文关键词,忽略模块归属。
---
## 建议修复
构建面向 embedding 的增强文本:
```text
标题: {title}
路径: {breadcrumb}
正文:
{content}
```
落库时仍保留原始 `content`,避免展示内容被污染。可以新增 `embeddingText` 构造逻辑,只用于向量化。
后续还可以在 rerank 阶段把 `breadcrumb` 作为加权信号,例如同域、同章节、同文档优先。
---
## 相关文件
- `src/main/java/com/superbiz/agent/service/DocumentChunkService.java`
- `src/main/java/com/superbiz/agent/service/VectorIndexService.java`
- `src/main/java/com/superbiz/agent/service/VectorEmbeddingService.java`
@@ -0,0 +1,55 @@
# RAG 切片上下文重建缺失
**状态**:待规划
**严重程度**:高
**发现时间**:2026-07-04
**范围**:知识库上传、切片、向量召回、Agent 上下文组装
---
## 现象
同一个 Markdown 章节在内容较长时会被拆成多个 chunk。当前检索命中其中一个 chunk 后,返回给 Agent 的主要是单个 chunk 内容,不会自动把同章节的前后片段、章节标题链路、相邻 chunk 一起恢复出来。
这会导致两个问题:
1. 命中片段只包含局部语义,缺少前置定义、约束条件或后续步骤。
2. 同章节被分段后,检索结果之间缺少可追溯的关联,Agent 不一定知道它们属于同一章节。
---
## 当前实现
- `DocumentChunkService` 会按 Markdown 标题建立 `title` 和 `breadcrumb`,再按段落累积切片。
- 超过阈值时仍会切断同一章节,只是尽量避免打断代码块和列表。
- `VectorIndexService` 会把 `chunkIndex`、`totalChunks`、`title`、`breadcrumb` 放入 metadata。
- `VectorSearchService` 查询 Milvus 后直接返回命中的 chunk,没有做相邻 chunk 扩展或 section 级聚合。
- `LookupKnowledgeTool` 消费 L1 结果时,也没有根据 `docId + chunkIndex + breadcrumb` 回补上下文。
---
## 影响
- RAG 回答容易漏掉同章节中的约束条件。
- 长流程类文档会被拆散,Agent 看到的是“片段证据”,不是“完整流程”。
- 面试解释中需要承认:当前系统有 metadata 基础,但还没有把它用于上下文重建。
---
## 建议修复
优先做命中后的上下文扩展:
1. L1 命中 chunk 后,按 `docId + chunkIndex` 拉取前后 N 个相邻 chunk。
2. 如果 metadata 中 `breadcrumb` 相同,允许扩展到同章节的多个 chunk。
3. 上下文打包时标记 `命中片段`、`前文`、`后文`,避免 Agent 把扩展内容误认为全部都是高置信命中。
4. 增加 token budget 控制,超过预算时优先保留命中 chunk 和标题链路。
---
## 相关文件
- `src/main/java/com/superbiz/agent/service/DocumentChunkService.java`
- `src/main/java/com/superbiz/agent/service/VectorIndexService.java`
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
@@ -0,0 +1,48 @@
# RAG 缺少上下文打包和 Rerank
**状态**:待规划
**严重程度**:中
**发现时间**:2026-07-04
**范围**:检索后处理、证据排序、Agent 输入质量
---
## 现象
当前 RAG 检索主要依赖 L0/L1 的原始召回顺序,没有独立的 reranker、cross-encoder 或 LLM rerank 阶段。召回结果进入 Agent 前,也缺少统一的上下文打包策略。
这意味着“检索到”不等于“以最适合推理的形式喂给 Agent”。
---
## 当前实现
- L0 和 L1 结果由 `LookupKnowledgeTool` 拼装后返回。
- 没有候选级 rerank。
- 没有明确的 token budget 分配策略,例如每个文档最多占多少、命中片段和扩展片段如何排序。
- 没有把 `title`、`breadcrumb`、score、source 统一包装成证据块。
---
## 影响
- 相关结果可能被排在不理想的位置。
- 多个候选内容相近时,Agent 可能读到重复信息。
- 证据结构不清晰,后续 verifier 或 trace 解释成本较高。
---
## 建议修复
1. 引入 `RetrievedEvidence` 这样的内部结构,统一承载 source、title、breadcrumb、score、hitReason、content。
2. 做简单 rerank:关键词命中、向量分、breadcrumb 匹配、文档去重、相邻片段扩展一起排序。
3. 上下文打包时按证据块输出,明确来源和置信度。
4. 面试版可以先实现规则 rerank,后续再替换为 cross-encoder 或 LLM rerank。
---
## 相关文件
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
- `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java`
+71
View File
@@ -0,0 +1,71 @@
# RAG 将 L0 降级为领域和实体 Hint
**状态**:待规划
**严重程度**:中
**发现时间**:2026-07-05
**范围**:L0 检索、metadata filter、业务可解释性
---
## 背景
当前 L0 是基于 frontmatter / keyword 的轻量检索。它具备可解释性,但不适合作为最终相关性判断。
在引入 Spring AI VectorStore / Retriever 后,L0 更适合从“召回主链路”调整为“检索前处理和解释信号”。
---
## 问题
当前 L0 如果唯一命中,容易被过度信任:
```text
L0 unique hit -> 直接返回 / 优先采信
```
这会带来误召回风险,尤其是关键词过泛、frontmatter 质量不稳定时。
---
## 改造方向
L0 保留,但职责调整为:
1. **Domain detector**
- 识别 query 所属 category/domain。
- 用于 Spring AI retriever metadata filter。
2. **Entity extractor**
- 识别组件名、服务名、指标名、错误码、接口名。
- 用于 query augmentation。
3. **Explainability signal**
- 记录 matched keywords。
- 解释为什么进入某个知识域。
目标链路:
```text
query / payload
-> L0 domain/entity hint
-> metadata filter + query augmentation
-> vector retriever
-> post processor
```
---
## 验收标准
- L0 不再默认作为最终检索结果直接返回。
- L0 命中的 domain/category 能传给 retriever filter。
- L0 命中的实体能进入增强 query 或工具调用记录。
- `tool_invocation` 能展示 L0 matched keywords 和使用方式。
---
## 相关文件
- `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java`
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
@@ -0,0 +1,53 @@
# RAG L0 关键词匹配质量不足
**状态**:待规划
**严重程度**:中
**发现时间**:2026-07-04
**范围**:L0 索引、frontmatter、精确召回
---
## 现象
L0 当前依赖文档 frontmatter 中的关键词,并使用较粗的字符串包含逻辑做匹配。关键词质量越依赖人工维护,召回稳定性越容易波动。
如果 frontmatter 填写不完整、同义词缺失、关键词过短或过泛,L0 就可能误召回或漏召回。
---
## 当前实现
- `KnowledgeIndexService` 从 `ApiDocument` metadata/frontmatter 加载关键词。
- exact match 的判断类似:
```java
query.contains(keywordLower) || keywordLower.contains(query)
```
- 没有分词、同义词归一、字段权重、关键词质量校验。
---
## 影响
- 短关键词容易误命中。
- 用户换一种说法时,L0 无法命中。
- 文档 frontmatter 质量变成检索质量的隐性前提。
---
## 建议修复
1. 为关键词增加最小长度、停用词、领域前缀等基础规则。
2. 区分 `exactKeywords`、`aliases`、`domainTags`,避免所有词混在一个匹配池。
3. 引入轻量中文分词或归一化策略,先不必上复杂搜索引擎。
4. 上传文档时校验 frontmatter 质量,缺失关键词时给出警告。
5. 在 issue 修复前,至少补一份知识库文档 frontmatter 编写规范。
---
## 相关文件
- `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java`
- `src/main/java/com/superbiz/agent/service/DocumentManagementService.java`
- `src/main/java/com/superbiz/agent/controller/DocumentController.java`
+55
View File
@@ -0,0 +1,55 @@
# RAG L0 和 L1 未真正融合排序
**状态**:待规划
**严重程度**:中
**发现时间**:2026-07-04
**范围**:知识检索工具、召回排序、Agent 证据质量
---
## 现象
当前 `lookup_knowledge` 的 L0 和 L1 更像是串行兜底关系,不是真正的多路召回融合:
- L0 命中唯一结果时,直接返回 L0。
- L0 不唯一或不足时,才进入 L1。
- L1 查询 `topK=3`,但最终主要把第一条结果作为补充证据。
这会导致关键词召回和语义召回没有充分互补。
---
## 当前实现
- `LookupKnowledgeTool` 先调用 `KnowledgeIndexService.exactMatch` 做 L0。
- 再按条件调用 `VectorSearchService.search` 做 L1。
- L0 和 L1 结果没有统一进入候选池做 fusion ranking。
- L1 多结果没有充分利用,相关性接近的候选可能被丢弃。
---
## 影响
- L0 命中但质量一般时,会压过更好的 L1 语义结果。
- L1 找到多个相近片段时,只有 top1 被 Agent 看到,降低召回覆盖率。
- 难以解释检索排序,因为当前更像规则分支,不是可调的排序模型。
---
## 建议修复
建立统一候选池:
1. L0 和 L1 都返回候选列表。
2. 按 `docId/chunkId` 去重。
3. 为候选计算综合分:`keywordScore`、`vectorScore`、`domainScore`、`freshness`、`breadcrumbMatch`。
4. 取 topN 进入上下文打包,而不是只取 L1 top1。
5. 在 `tool_invocation` 中记录每个候选的分数组成,方便调试。
---
## 相关文件
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
- `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java`
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
+49
View File
@@ -0,0 +1,49 @@
# RAG L1 分数阈值未校准
**状态**:待规划
**严重程度**:中
**发现时间**:2026-07-04
**范围**:向量搜索、相关性判断、工具调用记录
---
## 现象
L1 语义检索使用 Milvus 向量距离后,会做相关性归一和阈值判断。但当前阈值更偏经验值,没有基于真实查询集和真实分数分布做校准。
由于当前使用 L2 距离,不同 embedding 模型、不同语料密度、不同 query 长度都会影响分数分布。
---
## 当前实现
- `VectorSearchService` 使用 query embedding 搜索 Milvus。
- Milvus metric type 为 `L2`。
- `LookupKnowledgeTool` 会把 L2 score 转成 normalized relevance。
- 阈值没有配套评测集或分布统计。
---
## 影响
- 阈值过松时,低相关片段会进入 Agent 上下文。
- 阈值过紧时,正确片段可能被过滤掉。
- 面试中如果被追问“为什么这个阈值合理”,当前只能回答是 MVP 经验值。
---
## 建议修复
1. 固化一组 RAG 回归查询集,覆盖告警、数据库、流程规范、AIOps 诊断等场景。
2. 记录每次 topK 的原始 L2 score、归一化分数、最终是否采纳。
3. 统计正例和负例分布,确定阈值区间。
4. 将阈值配置化,并在 README 或 issue 中记录选择依据。
5. 后续引入 reranker 后,L1 阈值可以从“最终判断”退化为“粗召回过滤”。
---
## 相关文件
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
- `src/main/resources/application.yml`
+48
View File
@@ -0,0 +1,48 @@
# RAG 查询改写能力薄弱
**状态**:待规划
**严重程度**:中
**发现时间**:2026-07-04
**范围**:检索工具、Agent 查询生成、召回稳定性
---
## 现象
当前检索主要使用 Agent 传入 `lookup_knowledge` 的原始 query。工具层没有显式的 query rewrite、同义词扩展、领域词补全或多 query 检索。
当用户问题口语化、上下文依赖强,或缺少领域关键词时,L0 和 L1 的召回都可能不稳定。
---
## 当前实现
- Agent 决定何时调用 `lookup_knowledge` 和传入什么 query。
- `LookupKnowledgeTool` 接收 query 后直接进入 L0/L1 检索。
- 工具层没有把用户问题改写成多个检索 query。
- 也没有把当前任务域、Planner step、告警 payload 等上下文显式拼入检索 query。
---
## 影响
- Agent query 写得好时召回正常,query 写得差时检索链路缺少兜底。
- AIOps 场景里,告警名称、服务名、指标名、故障类型之间的别名关系没有被充分利用。
- 很难稳定复现同一类问题的检索质量。
---
## 建议修复
1. 在工具层增加轻量 query rewrite:原始问题、领域词增强问题、关键词查询并行召回。
2. 对 AIOps 场景,把 alertName、service、metric、symptom 显式构造成检索 query。
3. 记录 rewrite 前后的 query 到 `tool_invocation`,便于分析。
4. 后续可以引入 LLM query rewrite,但 MVP 先用规则模板更可控。
---
## 相关文件
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
- `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java`
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
+377
View File
@@ -0,0 +1,377 @@
# RAG 检索重构计划
**状态**:待规划
**严重程度**:高
**发现时间**:2026-07-05
**范围**:RAG、检索、知识库、Agent Tool、AIOps 诊断证据链
---
## 目标
将当前自研 RAG MVP 重构为“成熟框架能力 + 业务可观测编排”的架构:
```text
Agent
-> lookup_knowledge Tool
-> L0 domain/entity hint
-> query augmentation / transformer
-> Spring AI Retriever / VectorStore
-> metadata filter
-> document postprocess
-> neighbor / section expansion
-> evidence packing
-> tool_invocation record
```
核心原则:
1. 通用 RAG 基础设施尽量交给 Spring AI / Spring AI Alibaba。
2. Agent 工具入口、AIOps 业务语义、证据追踪继续保留在项目内。
3. 不把系统改成隐式 Chat RAG,仍然保留显式 `lookup_knowledge` 工具调用。
4. 分阶段迁移,避免一次性推倒当前可运行链路。
---
## 当前问题汇总
当前 RAG 已经打通上传、切片、向量化、L0/L1 召回和工具调用记录,但主要问题集中在:
1. **检索基础设施偏自研**
- Milvus 写入和查询直接使用 SDK。
- topK、threshold、metadata filter、结果结构由业务代码维护。
- 后续接入 Spring AI RAG 能力会有重复适配成本。
2. **L0 职责过重**
- 当前 L0 可能被当成最终召回决策。
- 关键词质量不稳定时容易误召回。
- 更适合作为 domain/entity hint,而不是最终答案来源。
3. **query 构造不稳定**
- 主要依赖 Agent 传入原始 query。
- AIOps payload 中的 alertName、service、metric、symptom 没有稳定进入检索 query。
4. **上下文重建不足**
- 同章节被切成多个 chunk 后,命中片段不会自动扩展前后文。
- `breadcrumb` 存在 metadata 中,但没有充分参与 embedding、filter 或 context packing。
5. **缺少检索后处理**
- 缺少统一 evidence block。
- 缺少去重、token budget、hitReason、source 结构化输出。
6. **缺少量化评测**
- 目前主要靠接口回放、日志和 `tool_invocation` 人工判断。
- 还没有 golden query set、Recall@K、MRR、NDCG 等检索评测。
---
## 保留设计
这些设计值得保留,并作为重构后的项目亮点:
### 1. `lookup_knowledge` 显式 Agent Tool
保留显式工具调用,不直接用隐式 Advisor 取代。
原因:
- 面试项目重点是 Agent 工程,不是普通 Chat RAG。
- 显式工具调用能展示 Agent 何时检索、检索了什么、证据如何支撑诊断。
- `tool_invocation`、evidence score、diagnosis session 都依赖这条链路。
### 2. L0
保留 L0,但降级为:
- domain detector
- entity extractor
- metadata filter generator
- explainability signal
不再默认执行:
```text
L0 unique hit -> 直接返回
```
目标职责:
```text
query / payload
-> L0 matched keywords/entities/domain
-> metadata filter + query augmentation
-> retriever
```
### 3. metadata
保留并加强 metadata:
```text
docId
chunkIndex
totalChunks
title
breadcrumb
category
source
```
后续可扩展:
```text
sectionId
parentSection
documentType
domain
tags
version
```
metadata 是 filter、上下文扩展、证据追踪、可解释性的基础。
### 4. Markdown-aware chunking
保留当前 Markdown 结构化切片思路:
- 识别标题层级
- 生成 `title`
- 生成 `breadcrumb`
- 保留 `chunkIndex`
- 尽量不打断列表和代码块
可以替换或复用框架能力的是底层 token 长度控制和 overlap 策略,而不是完全抛弃结构化切片。
### 5. `tool_invocation` 证据追踪
保留并增强:
```text
sessionId
query
rewrittenQuery
matchedKeywords
domain/entities
retrievedDocs
scores
hitReasons
evidence
duration
relevanceLevel
```
这是后续检索评测、诊断质量评估、面试讲解的基础。
### 6. AIOps payload 到 query 的业务映射
保留 AIOps 场景逻辑:
- alertName
- service
- metric
- symptom
- category/domain
这些是业务语义,不能完全交给通用框架隐式处理。
---
## 替换设计
这些能力适合逐步交给 Spring AI / Spring AI Alibaba:
| 当前能力 | 目标能力 | 说明 |
|---|---|---|
| Milvus SDK 直接写入/查询 | Spring AI `VectorStore` | 减少基础设施代码 |
| 自研 `VectorSearchService` 检索细节 | `VectorStoreDocumentRetriever` | 标准化 topK、threshold、filter |
| 手写 query 拼接 | Query Transformer / 模板化 query augmentation | 先规则化,后框架化 |
| 手写结果拼接 | DocumentPostProcessor / evidence postprocess | 做去重、压缩、证据块 |
| L0 最终召回判断 | L0 domain/entity hint | 降低误召回风险 |
---
## 分阶段计划
### Phase 0:重构前基线
目标:先固定当前行为,避免重构后不知道是否变好。
任务:
- 固化 10-20 条 golden queries。
- 覆盖 Chat 和 AIOps 场景。
- 每条 query 标注 expected doc、breadcrumb、关键 chunk 或 evidence。
- 用当前链路跑一遍,记录 baseline。
- 初始离线基线落在 `eval/rag-retrieval/`,用于后续 change 对比。
验收:
- 有可重复运行的检索回放清单。
- 能记录当前 Recall@K、first hit rank 或人工 hit level。
### Phase 1:L0 降级为 domain/entity hint
目标:保留 L0 价值,降低 L0 误决策风险。
任务:
- `KnowledgeIndexService` 输出 matched keywords、domain、entities。
- `LookupKnowledgeTool` 不再把 L0 unique hit 作为默认最终结果。
- 将 L0 结果用于 query augmentation 和 metadata filter。
- `tool_invocation` 记录 L0 hit reason。
验收:
- L0 命中不会绕过向量检索直接返回。
- 检索记录能看到 domain/entities/matchedKeywords。
- AIOps payload 能生成稳定领域 hint。
### Phase 2:Evidence Postprocess 和上下文打包
目标:先提升 Agent 实际拿到的证据质量。
任务:
- 定义 evidence block:
```text
source
docId
chunkIndex
title
breadcrumb
score
hitReason
content
expandedFrom
```
- 对检索结果做去重。
- 支持命中 chunk 的相邻 chunk / 同章节扩展。
- 加 token 或字符预算控制。
- 返回给 Agent 的内容按 evidence block 组织。
验收:
- 同一 docId/chunkIndex 不重复进入上下文。
- 命中 chunk 可以补充前后文。
- `tool_invocation` 记录 postprocess 前后候选数量和最终 evidence 数量。
### Phase 3:Spring AI VectorStore 旁路验证
目标:验证框架能力,不直接替换主链路。
任务:
- 引入 Spring AI Milvus VectorStore。
- 建立旁路 `SpringAiVectorSearchService` 或适配层。
- 同一批 golden queries 同时跑旧链路和新链路。
- 对比 topK、metadata、score、filter 行为。
验收:
- 旁路检索可跑通。
- metadata 不丢失。
- 查询结果与当前链路差异可解释。
- 不影响现有 Chat / AIOps 主链路。
### Phase 4:替换底层 VectorSearchService
目标:对外接口不变,内部检索切到 Spring AI VectorStore / Retriever。
任务:
- 保持 `LookupKnowledgeTool` 调用方式不变。
- `VectorSearchService` 内部迁移到 Spring AI 检索抽象。
- 支持 topK、similarity threshold、category metadata filter。
- 保留旧实现一段时间作为 fallback。
验收:
- Chat / AIOps 检索链路行为兼容。
- golden queries 不低于 baseline。
- 检索结果仍能完整记录到 `tool_invocation`。
### Phase 5:Query Transformer 和框架化 PostProcessor
目标:在稳定的 VectorStore 基础上接入更成熟 RAG 能力。
任务:
- AIOps 场景优先使用模板化 query augmentation。
- 需要时接入 Spring AI Query Transformer / MultiQuery。
- 将现有 evidence postprocess 抽象成 DocumentPostProcessor 风格。
- 可选接入 rerank,但不作为第一优先级。
验收:
- 原始 query 和 rewritten query 都可追踪。
- query rewrite 失败可以 fallback。
- postprocess 行为可配置、可记录、可回放。
---
## 暂不做
以下能力暂不进入近期重构:
1. 不做完整自研 RRF 框架。
2. 不直接把 `lookup_knowledge` 替换成隐式 Advisor。
3. 不一口气迁移所有 RAG ETL。
4. 不先引入 Elasticsearch / OpenSearch,除非评测证明 BM25 必须。
5. 不先上 cross-encoder / LLM rerank,先做规则型 evidence postprocess。
---
## 风险
### 1. Milvus schema 兼容风险
当前 collection 是项目自建,Spring AI VectorStore 可能有自己的 schema 假设。需要旁路验证。
### 2. 检索行为变化风险
框架检索分数和当前 L2 score 可能不完全一致,需要 golden queries 对比。
### 3. 可观测性丢失风险
如果迁移到隐式 Advisor,可能丢失工具调用证据链。因此 Spring AI RAG 能力应优先封装在 `lookup_knowledge` 内部。
### 4. 重构范围膨胀风险
RAG、Agent、AIOps、数据库记录互相关联,必须分阶段推进,每阶段都保持可运行。
---
## 合并来源
本计划合并以下问题和改造方向:
- [rag-chunk-context-reconstruction.md](rag-chunk-context-reconstruction.md)
- [rag-breadcrumb-embedding-gap.md](rag-breadcrumb-embedding-gap.md)
- [rag-l0-l1-fusion-ranking.md](rag-l0-l1-fusion-ranking.md)
- [rag-l0-keyword-matching-quality.md](rag-l0-keyword-matching-quality.md)
- [rag-l1-score-calibration.md](rag-l1-score-calibration.md)
- [rag-context-packing-and-reranking.md](rag-context-packing-and-reranking.md)
- [rag-upload-chunk-parameter-drift.md](rag-upload-chunk-parameter-drift.md)
- [rag-query-rewrite-gap.md](rag-query-rewrite-gap.md)
- [rag-spring-ai-vectorstore-migration.md](rag-spring-ai-vectorstore-migration.md)
- [rag-spring-ai-query-transformer.md](rag-spring-ai-query-transformer.md)
- [rag-spring-ai-document-postprocessor.md](rag-spring-ai-document-postprocessor.md)
- [rag-l0-domain-entity-hint.md](rag-l0-domain-entity-hint.md)
- [rag-spring-ai-advisor-boundary.md](rag-spring-ai-advisor-boundary.md)
---
## 相关文件
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
- `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java`
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
- `src/main/java/com/superbiz/agent/service/VectorIndexService.java`
- `src/main/java/com/superbiz/agent/service/DocumentChunkService.java`
- `src/main/java/com/superbiz/agent/service/DocumentManagementService.java`
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
- `src/main/resources/application.yml`
- `pom.xml`
@@ -0,0 +1,66 @@
# RAG 明确 Spring AI Advisor 与 Agent Tool 的边界
**状态**:待规划
**严重程度**:中
**发现时间**:2026-07-05
**范围**:Agent 编排、RAG Advisor、工具调用可观测性
---
## 背景
Spring AI 提供 `QuestionAnswerAdvisor`、`RetrievalAugmentationAdvisor` 等 RAG Advisor 能力,可以把检索增强直接挂到模型调用流程中。
但当前项目是 Agent 工程项目,知识检索不是普通聊天增强,而是 Agent 在诊断流程中显式调用的工具。系统还依赖 `tool_invocation` 记录检索事实,用于 evidence score 和诊断追踪。
---
## 问题
如果直接把 RAG 全部迁到 Advisor,可能会损失当前项目已有的显式工具链路:
1. Agent 是否调用知识库不够透明。
2. `tool_invocation` 记录可能变弱。
3. AIOps 诊断步骤和知识证据之间的对应关系不清晰。
4. 面试项目中“Agent 如何使用工具”的展示价值下降。
---
## 改造方向
不要把 `lookup_knowledge` 完全替换成隐式 Advisor,而是分层使用:
```text
Agent Tool 层:
lookup_knowledge
sessionId
traceId
tool_invocation
evidence score
Spring AI RAG 层:
query transformer
retriever
vector store
document post processor
```
也就是说,Advisor / Retriever 可以作为工具内部实现,而不是取代工具本身。
---
## 验收标准
- Agent 仍然通过显式 `lookup_knowledge` 使用知识库。
- Spring AI RAG 能力被封装在工具内部或服务内部。
- 每次检索仍能落 `tool_invocation`。
- Chat / AIOps 两条链路都能追踪检索输入、输出和证据来源。
---
## 相关文件
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
@@ -0,0 +1,60 @@
# RAG 使用 DocumentPostProcessor 做后处理
**状态**:待规划
**严重程度**:中
**发现时间**:2026-07-05
**范围**:检索后处理、去重、上下文打包、轻量 rerank
---
## 背景
当前检索结果返回给 Agent 前,主要依赖 `LookupKnowledgeTool` 自己拼接内容。系统还缺少统一的后处理阶段。
Spring AI RAG 流程中可以使用 DocumentPostProcessor 类能力,在文档进入模型上下文前做过滤、去重、压缩或 rerank。
---
## 问题
当前检索后处理不足:
1. L1 topK 候选没有被充分利用。
2. 同文档或同章节结果可能重复。
3. 命中 chunk 后没有统一处理前后文扩展。
4. 证据块缺少统一格式,后续 verifier / evaluator 不容易复用。
---
## 改造方向
建立一个轻量后处理链:
```text
retrieved documents
-> deduplicate
-> optional neighbor / section expansion
-> score / reason annotation
-> token budget packing
-> evidence blocks
```
优先做规则型后处理,不急于引入 cross-encoder 或 LLM rerank。
---
## 验收标准
- 同一个 docId/chunkIndex 不重复进入 Agent 上下文。
- 最终返回内容包含 source、title、breadcrumb、score、hitReason。
- 可以限制单次工具调用返回的最大 token 或最大字符数。
- 后处理前后的候选数量、去重数量、最终 evidence 数量记录到 `tool_invocation`。
---
## 相关文件
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
- `src/main/java/com/superbiz/agent/service/DocumentChunkService.java`
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
@@ -0,0 +1,70 @@
# RAG 接入 Spring AI Query Transformer
**状态**:待规划
**严重程度**:中
**发现时间**:2026-07-05
**范围**:查询改写、多查询扩展、AIOps 检索稳定性
---
## 背景
当前 `lookup_knowledge` 主要使用 Agent 传入的原始 query 进行 L0/L1 检索。query 的质量高度依赖 Agent 当次生成结果。
Spring AI 提供 Query Transformer / Query Expander 类能力,可以把用户问题或 Agent 子任务改写成更适合检索的查询。
---
## 问题
当前检索 query 存在几个风险:
1. 用户问题口语化时,缺少领域关键词。
2. AIOps payload 中的 alertName、service、metric 没有稳定拼入检索 query。
3. 同义表达没有扩展,例如“连接耗尽”和“连接池打满”。
4. 工具层无法复用框架提供的 rewrite / expansion 能力。
---
## 改造方向
在 `lookup_knowledge` 前增加查询改写层:
```text
raw query / alert payload
-> query transformer
-> rewritten query / expanded queries
-> retriever
```
优先支持两类场景:
1. **AIOps 模板化改写**
- alertName
- service
- metric
- symptom
- domain/category
2. **Spring AI Query Transformer**
- rewrite 原始 query
- multi-query expansion
- 必要时做 query compression
---
## 验收标准
- `tool_invocation` 中记录原始 query 和改写后的 query。
- AIOps payload 存在时,检索 query 能稳定带上告警和服务上下文。
- 对同一个测试问题,改写前后 topK 命中结果可对比。
- 未配置 transformer 时,可以回退到原始 query。
---
## 相关文件
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
- `src/main/java/com/superbiz/agent/domain/entity/ToolInvocation.java`
@@ -0,0 +1,80 @@
# RAG 迁移到 Spring AI VectorStore 检索抽象
**状态**:待规划
**严重程度**:高
**发现时间**:2026-07-05
**范围**:向量检索、Milvus 接入、RAG 框架化改造
---
## 背景
当前系统的向量检索链路主要由项目手写实现:
- `VectorIndexService` 负责向量化和写入 Milvus。
- `VectorSearchService` 直接使用 Milvus SDK 查询。
- `LookupKnowledgeTool` 自己组织 L0/L1 检索结果。
这能满足 MVP 打通链路,但继续扩展 RAG 能力时,容易把项目变成自研搜索框架。
项目当前已引入 Spring AI / Spring AI Alibaba 依赖,可以考虑迁移到 Spring AI 的 `VectorStore`、`VectorStoreDocumentRetriever` 等标准抽象。
---
## 问题
当前手写 Milvus 检索存在几个成本:
1. topK、similarity threshold、metadata filter 等逻辑分散在业务代码中。
2. 检索结果结构和 Spring AI RAG Advisor 生态不兼容。
3. 后续接入 query transformer、post processor、advisor 时需要重复适配。
4. Milvus SDK 直接调用让业务层承担了过多基础设施细节。
---
## 改造方向
优先引入 Spring AI 的 Milvus VectorStore 能力:
```text
当前:
VectorSearchService -> Milvus SDK
目标:
LookupKnowledgeTool / RAG Service
-> VectorStoreDocumentRetriever
-> Spring AI VectorStore
-> Milvus
```
业务层保留:
- `lookup_knowledge` 工具入口
- `tool_invocation` 记录
- sessionId / category / domain 等业务上下文
底层检索交给框架:
- topK
- similarity threshold
- metadata filter
- vector search options
---
## 验收标准
- `VectorSearchService` 不再直接散落 Milvus 查询细节,至少封装到 Spring AI `VectorStore` 适配层。
- 支持按 `category` 或其他 metadata filter 检索。
- 检索结果仍能记录到 `tool_invocation`。
- 现有 AIOps / Chat 检索链路行为保持兼容。
---
## 相关文件
- `pom.xml`
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
- `src/main/java/com/superbiz/agent/service/VectorIndexService.java`
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
- `src/main/resources/application.yml`
@@ -0,0 +1,49 @@
# RAG 上传切片参数未真正生效
**状态**:待规划
**严重程度**:低
**发现时间**:2026-07-04
**范围**:文档上传接口、切片配置、API 行为一致性
---
## 现象
上传接口暴露了 `chunkSize` 和 `chunkOverlap` 参数,但实际切片主要使用全局 `DocumentChunkConfig`。这会造成 API 表面能力和真实行为不一致。
---
## 当前实现
- `DocumentController` 的上传接口接收 `chunkSize` 和 `chunkOverlap`。
- 这些字段会进入 `DocumentUploadRequest`。
- `DocumentChunkService` 的切片阈值主要来自 `DocumentChunkConfig`。
- 单次上传请求中的参数没有真正覆盖切片配置。
---
## 影响
- 调用方以为可以控制切片大小,但实际无法影响结果。
- 测试时容易误判“参数调优无效”的原因。
- 面试中如果展示 API,会被追问参数是否真实生效。
---
## 建议修复
两个方向二选一:
1. 如果 MVP 不需要请求级切片参数,就从接口中移除或标记为暂不支持。
2. 如果需要支持,就让 `DocumentChunkService` 接收 per-request chunk options,并记录到文档 metadata 中。
建议面试项目中优先选择第二种,因为它更能体现工程闭环:API、配置、落库、追踪一致。
---
## 相关文件
- `src/main/java/com/superbiz/agent/controller/DocumentController.java`
- `src/main/java/com/superbiz/agent/service/DocumentManagementService.java`
- `src/main/java/com/superbiz/agent/service/DocumentChunkService.java`
- `src/main/java/com/superbiz/agent/config/DocumentChunkConfig.java`
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-05
@@ -0,0 +1,59 @@
## Context
Chat diagnosis has a Verifier Agent that writes structured evaluation into `diagnosis_session.self_evaluation`. AIOps currently focuses on payload scoping, evidence tools, and trace persistence, but it has no quality gate that checks whether the final report stayed on target or used evidence.
The next stage should add a low-risk quality gate before considering a full AIOps LLM verifier.
## Goals / Non-Goals
**Goals:**
- Evaluate AIOps final reports with deterministic rules.
- Persist the evaluation under a dedicated `aiops_rule_evaluation` self-evaluation key.
- Keep trace replay able to show whether AIOps output passed, warned, or failed basic quality checks.
- Add focused unit tests without requiring live LLMs or external tools.
**Non-Goals:**
- Do not add an AIOps Verifier Agent yet.
- Do not route/retry AIOps execution based on the evaluation result.
- Do not change `tool_invocation` schema.
- Do not require new database migrations.
## Decisions
### Decision 1: Rule-Based Before LLM Verifier
The first AIOps verifier is a deterministic evaluator, not an LLM agent.
Rationale:
- AIOps quality risks are concrete at this stage: payload focus, evidence coverage, and report presence.
- Rule evaluation is cheap, stable, and easy to explain in an interview.
- A full verifier agent can be added later once AIOps trace expectations are stable.
### Decision 2: Dedicated Self-Evaluation Channel
Persist under `aiops_rule_evaluation` instead of reusing `rule_evaluation` or `verifier_evaluation`.
Rationale:
- `verifier_evaluation` is already associated with Chat's LLM verifier.
- `rule_evaluation` may be used by generic diagnosis evaluation.
- A dedicated key avoids conflating AIOps-specific checks with other evaluation channels.
### Decision 3: Evaluate After Final Report Persistence
Run the evaluator when `persistFinalReport(...)` is called.
Rationale:
- It has access to the final report and session id.
- It can read persisted tool invocations for the same session.
- It does not disturb the Agent execution path.
## Risks / Trade-offs
- [Risk] Rule evaluation can miss semantic hallucinations. -> Mitigation: position it as lightweight AIOps quality gate, not full groundedness verification.
- [Risk] Strict keyword checks may warn on valid reports with different wording. -> Mitigation: use WARN for missing soft signals and FAIL only for critical absence.
- [Risk] Evaluation after report persistence does not trigger retries. -> Mitigation: keep routing unchanged in this phase; later changes can consume the verdict.
@@ -0,0 +1,26 @@
## Why
AIOps now has traceable payload scope control and improved RAG retrieval, but it still lacks a quality gate comparable to Chat's verifier. A lightweight rule-based verifier can check the most important AIOps risks without introducing another LLM agent.
## What Changes
- Add a rule-based AIOps evaluation service that checks final report quality after the AIOps flow completes.
- Persist the evaluation under `diagnosis_session.self_evaluation.aiops_rule_evaluation`.
- Evaluate payload focus, evidence-tool coverage, and basic report completeness.
- Expose the evaluation through the existing trace API self-evaluation payload.
## Capabilities
### New Capabilities
None.
### Modified Capabilities
- `aiops-traceable-diagnosis-entry`: AIOps sessions include a lightweight rule evaluation for trace replay.
## Impact
- Affects AIOps session finalization and trace self-evaluation.
- Does not change AIOps API input, Agent flow topology, tool signatures, or database schema.
- Does not add an LLM verifier agent.
@@ -0,0 +1,22 @@
## ADDED Requirements
### Requirement: AIOps sessions SHALL persist lightweight rule evaluation
When an AIOps final report is persisted, the system SHALL evaluate it with deterministic AIOps-specific quality rules and store the result in session self-evaluation.
#### Scenario: Payload-focused report is evaluated
- **WHEN** an AIOps session has alert payload fields and a final report is persisted
- **THEN** the system SHALL evaluate whether the report mentions the supplied alert and service
- **AND** it SHALL store the result under `self_evaluation.aiops_rule_evaluation`
#### Scenario: Evidence coverage is evaluated
- **WHEN** an AIOps final report is evaluated
- **THEN** the system SHALL check whether evidence tool invocations such as `lookup_knowledge`, `query_metrics`, or `query_logs` were persisted for the session
#### Scenario: Evaluation is traceable
- **WHEN** the diagnosis trace API returns an AIOps session
- **THEN** the session self-evaluation payload SHALL include `aiops_rule_evaluation` when it has been generated
#### Scenario: Evaluation uses stable verdicts
- **WHEN** AIOps rule evaluation completes
- **THEN** it SHALL produce a verdict from `PASS`, `WARN`, or `FAIL`
- **AND** it SHALL include check details and a human-readable rationale
@@ -0,0 +1,19 @@
## 1. Rule Evaluation
- [x] 1.1 Add an AIOps rule evaluation service with PASS/WARN/FAIL verdicts.
- [x] 1.2 Check payload focus, evidence-tool coverage, and report completeness.
## 2. AIOps Integration
- [x] 2.1 Persist AIOps rule evaluation when the final AIOps report is saved.
- [x] 2.2 Make trace summary indicate that AIOps rule evaluation exists.
## 3. Tests And Docs
- [x] 3.1 Add focused unit tests for the evaluator and AIOps integration.
- [x] 3.2 Add interview notes for the lightweight AIOps verifier.
## 4. Verification
- [x] 4.1 Run focused service tests.
- [x] 4.2 Validate the OpenSpec change and review git scope.
@@ -0,0 +1,2 @@
schema: spec-driven
created: 2026-07-04
@@ -0,0 +1,54 @@
## Context
The AIOps endpoint has two natural modes:
- **Payload mode**: caller supplies `alertName`, `service`, or other alert fields. The caller is asking for targeted diagnosis of that alert.
- **Auto-discovery mode**: caller omits alert fields. The system should discover active alerts first, then analyze them.
The current task prompt does not distinguish these modes, so the agent may query all active alerts and produce a broad report even when a specific alert payload was supplied.
## Goals / Non-Goals
**Goals:**
- Make AIOps payload mode single-alert focused.
- Keep no-payload mode compatible with the original "query active alerts then diagnose" behavior.
- Keep the change prompt-only and low risk.
- Add tests for prompt scope rules.
**Non-Goals:**
- Do not add a Verifier Agent.
- Do not force tool calls in Java code.
- Do not change `/api/ai_ops` request/response contracts.
- Do not modify mock alert data.
## Decisions
| Decision | Choice | Alternative Considered | Rationale |
|---|---|---|---|
| Scope detection | Treat non-empty alert fields as payload mode | Add explicit `mode` field | Existing payload already carries enough intent; no API change needed. |
| Payload mode behavior | Final report focuses only on supplied alert | Filter tool results in Java | Prompt-level rule is the smallest change and preserves agent flexibility. |
| Auto mode behavior | Require active-alert discovery first | Always analyze only one alert | Original AIOps value is automated alert discovery when no payload exists. |
| Other active alerts in payload mode | Mention only as related risk | Ignore entirely | Some context can be useful, but not enough to expand the report. |
## Prompt Rules
Payload mode MUST instruct the agent:
- Treat supplied payload as the primary and only report target.
- Use `queryPrometheusAlerts` only to verify the supplied alert state or identify related risk.
- Do not create root-cause sections for unrelated active alerts.
- Report unrelated alerts only in a brief "关联风险" note if they appear relevant.
Auto-discovery mode MUST instruct the agent:
- First call `queryPrometheusAlerts`.
- Select P0/P1 or longest-running firing alerts.
- Analyze one or more active alerts based on severity and evidence.
## Risks / Trade-offs
- [Risk] Prompt-only control may not be perfectly followed by the LLM. -> Mitigation: tests lock prompt wording; runtime can be reviewed through trace.
- [Risk] Payload mode may miss broader incidents. -> Mitigation: related active alerts may be mentioned as risk, but not expanded into full sections.
- [Risk] Future stronger enforcement may be needed. -> Mitigation: a later change can filter tool summaries or add AIOps Verifier.

Some files were not shown because too many files have changed in this diff Show More