Compare commits

...
Author SHA1 Message Date
aruo 88e0a6c944 docs: reorganize MVP interview documentation 2026-07-05 15:29:28 +08:00
aruo b22f2d22c8 docs: archive historical openspec changes 2026-07-05 14:04:32 +08:00
aruo 63b62b28a2 docs: archive aiops lightweight verifier change 2026-07-05 14:01:33 +08:00
aruo d5902a0499 docs: update mvp architecture snapshot 2026-07-05 13:56:51 +08:00
aruo ed267d753d feat: add aiops lightweight verifier 2026-07-05 13:44:30 +08:00
aruo 2658742119 docs: archive rag and aiops query changes 2026-07-05 13:12:47 +08:00
aruo 72a3dbf8c5 feat: add aiops payload query augmentation 2026-07-05 12:56:20 +08:00
aruo 674dd27a48 feat: add rag post-reindex acceptance 2026-07-05 12:29:24 +08:00
aruo c7e2fc2ee2 feat: include breadcrumb in embedding text 2026-07-05 12:15:27 +08:00
aruo 1bfe1a17b4 docs: add rag refactor story 2026-07-05 12:09:32 +08:00
aruo 9dd6823fe7 docs: add rag retrieval quality report 2026-07-05 11:29:16 +08:00
aruo 9376448804 docs: add rag vectorstore interview notes 2026-07-05 11:13:50 +08:00
aruo f2bae0382c fix: align vectorstore live retrieval 2026-07-05 10:53:45 +08:00
aruo 5c71f5fc79 feat: integrate spring ai vectorstore fallback 2026-07-05 10:20:29 +08:00
aruo b9ec07de57 feat: add spring ai retrieval sidecar 2026-07-05 03:22:03 +08:00
aruo 5197712719 feat: add rag evidence postprocess blocks 2026-07-05 03:03:18 +08:00
aruo 4a94c14feb feat: treat l0 retrieval as domain hint 2026-07-05 02:18:40 +08:00
aruo 9a2a44d1b5 test: add rag retrieval baseline 2026-07-05 02:02:27 +08:00
aruo 79feed3314 Merge branch 'emdash/shy-items-fry-f4zze' into refactor/mvp1.0
# Conflicts:
#	mvp/issues/README.md
2026-07-05 01:42:42 +08:00
aruo 98155ae1d8 Merge branch 'aiops-trace-scope' into refactor/mvp1.0 2026-07-05 01:42:23 +08:00
aruo 2609c5a5ab docs: consolidate rag refactor issues 2026-07-05 01:40:08 +08:00
aruo bf5286c8f4 docs: add interview project materials 2026-07-05 01:39:37 +08:00
aruo 26e12a8d6b Archive MVP demo interview runbook 2026-07-05 01:34:04 +08:00
aruo cbef3ddd3c Add MVP demo interview runbook 2026-07-05 01:25:20 +08:00
aruo 69deb15330 Add diagnosis eval baseline diff 2026-07-05 00:59:53 +08:00
aruo 4c7c53b024 Expand diagnosis eval fixtures 2026-07-05 00:27:57 +08:00
aruo ca5c61fabf Add diagnosis eval harness 2026-07-04 23:51:43 +08:00
aruo 23ee05c7c3 feat: add traceable scoped AIOps diagnosis 2026-07-04 22:57:28 +08:00
aruo dc6cd32a67 Harden evidence trace semantics 2026-07-04 22:36:30 +08:00
279 changed files with 16546 additions and 480 deletions
+4
View File
@@ -60,3 +60,7 @@ uploads/
### Windows / Runtime Artifacts
*.stackdump
NUL
### MVP Demo Generated Outputs
mvp/demo/output/*.json
!mvp/demo/output/README.md
+1 -1
View File
@@ -1,7 +1,7 @@
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **SuperBizAgent-java** (1528 symbols, 2828 relationships, 87 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **SuperBizAgent-java** (7988 symbols, 12713 relationships, 297 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
+1 -1
View File
@@ -115,7 +115,7 @@ trailing off into the following information in 99% of cases:
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **SuperBizAgent-java** (1001 symbols, 2043 relationships, 78 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **SuperBizAgent-java** (7988 symbols, 12713 relationships, 297 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
+7
View File
@@ -4,7 +4,14 @@
| 日期 | slug | 领域 | 关键词 | 状态 |
|---|---|---|---|---|
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
| 2026-07-04 | diagnosis-eval-harness | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
| 2026-07-04 | evidence-trace-hardening | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
| 2026-07-03 | mvp-demo-trace-acceptance | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
| 2026-07-04 | aiops-traceable-diagnosis-entry | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
| 2026-07-04 | aiops-alert-scope-control | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
| 2026-05-29 | chatmodel-abstraction | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | archived |
| 2026-06-23 | phase1-infrastructure | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | archived |
| 2026-06-24 | lookup-knowledge-integration | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | archived |
@@ -0,0 +1,14 @@
# Acceptance: aiops-alert-scope-control
## Verification
- [x] Payload-mode prompt focuses the final report on the supplied alert.
- [x] No-payload prompt requires active-alert discovery first.
- [x] Targeted tests pass.
- [x] Compile passes.
- [x] OpenSpec validates.
## Known Limits
- Prompt-only scope control may still require runtime observation.
- AIOps Verifier remains deferred.
@@ -0,0 +1,28 @@
# Brief: aiops-alert-scope-control
## Background
After `aiops-traceable-diagnosis-entry`, AIOps can be triggered by payload and replayed through trace. Runtime verification showed one semantic gap: payload mode still produced a broad report over all active mock alerts.
## Goal
Make AIOps scope explicit:
- Payload present -> targeted diagnosis for the supplied alert.
- Payload absent -> automatic active-alert discovery and diagnosis.
## Scope
- In scope:
- `AiOpsService.buildTaskPrompt(...)` scope rules.
- Focused tests.
- Demo acceptance wording.
- Out of scope:
- Verifier integration.
- Java-side filtering of tool results.
- API shape changes.
- Database changes.
## Related OpenSpec
`openspec/changes/aiops-alert-scope-control/`
@@ -0,0 +1,42 @@
# Decisions: aiops-alert-scope-control
## Clarify
- Entry summary: tighten AIOps report scope after runtime verification showed payload mode still analyzes all active alerts.
- Slug: `aiops-alert-scope-control`
- Scale: standard-light.
## Context
- AIOps traceability is implemented and verified.
- Mock Prometheus returns multiple active alerts.
- Payload demo supplies `HighCPUUsage/payment-service`, but previous report expanded to `HighMemoryUsage` and `SlowResponse`.
## Grill Question Pool
| # | Dimension | Question | Mode | Status |
|---|---|---|---|---|
| Q1 | Product Boundary | What makes `/api/ai_ops` different from `/api/chat` when payload exists? | evidence-driven | Payload is alert-event driven and should be scoped to that event. |
| Q2 | Scope | Should payload mode ignore all other active alerts? | user-interview | No; mention only as related risk/context. |
| Q3 | Compatibility | Should no-payload mode keep old "query active alerts" behavior? | evidence-driven | Yes. |
| Q4 | Enforcement | Should Java filter unrelated tool results now? | evidence-driven | No; prompt-only is sufficient for this small change. |
| Q5 | Verifier | Should this change add AIOps Verifier? | user-interview | No; keep deferred. |
## Evidence-Driven Conclusions
| Conclusion | Evidence Source | Result |
|---|---|---|
| Scope issue is prompt-level. | `/api_ ai_ops` trace showed all mock alerts analyzed despite payload. | Update task prompt. |
| No API or persistence changes are needed. | `AIOpsRequest` already carries payload and trace works. | Keep endpoint unchanged. |
| Blast radius is low. | `buildTaskPrompt(...)` is internal to `AiOpsService`. | Add tests for prompt content. |
## GitNexus
GitNexus remains skipped by prior user decision and because tools are not exposed in this session. Local impact analysis is recorded instead.
## Key Decisions
- Payload mode is detected when any alert field is present.
- Payload mode final report must focus on the supplied alert.
- No-payload mode must first call `queryPrometheusAlerts`.
- Other active alerts in payload mode can appear only as related risk, not as separate root-cause sections.
@@ -0,0 +1,44 @@
# Evidence: aiops-alert-scope-control
## Local Impact Analysis
- `AiOpsService.buildTaskPrompt(...)` is used by `executeAiOpsAnalysis(...)`.
- No controller, DTO, repository, or database changes are required.
- Existing `AiOpsServiceTest` already exercises request summary helpers and can be extended for scope prompt rules.
## Verification Results
- `mvn -q "-Dtest=AiOpsServiceTest" test` passed.
- `mvn -q -DskipTests compile` passed.
- `openspec.cmd validate aiops-alert-scope-control --strict` passed.
## Runtime Verification
- Runtime session: `mvp-demo-aiops-payment-cpu-codex-scope-003`.
- `/api/ai_ops` SSE emitted the requested `session` message and finished with `done`.
- `diagnosis_session` persisted:
- `agent_flow = AI_OPS`
- `status = SUCCESS`
- `total_duration_ms = 69875`
- `step_count = 5`
- `tool_call_count = 8`
- Tool invocation counts:
- `query_metrics = 1`
- `lookup_knowledge = 1`
- `query_logs = 6`
- Report scope check:
- `告警根因分析 - HighCPUUsage` exists.
- `告警根因分析 - HighMemoryUsage` does not exist.
- `告警根因分析 - SlowResponse` does not exist.
- `相关风险告警` exists.
## Runtime Fix
- Added Hikari settings in `src/main/resources/application.yml` after the first runtime attempt failed on stale MySQL pool connections:
- `maximum-pool-size: 5`
- `minimum-idle: 1`
- `connection-timeout: 10000`
- `validation-timeout: 5000`
- `idle-timeout: 60000`
- `max-lifetime: 120000`
- `keepalive-time: 30000`
@@ -0,0 +1,18 @@
# Acceptance: aiops-traceable-diagnosis-entry
## Verification
- [x] OpenSpec validates for `aiops-traceable-diagnosis-entry`.
- [x] Targeted AIOps service tests pass.
- [x] Compile verification passes.
- [x] Demo docs describe AIOps request -> session id -> trace query.
## Result
Accepted for implementation scope.
## Known Limits
- AIOps Verifier integration is deferred.
- Runtime still depends on configured model and infrastructure.
- Full browser/SSE runtime verification is not guaranteed in this coding pass.
@@ -0,0 +1,28 @@
# Brief: aiops-traceable-diagnosis-entry
## Background
The MVP chat diagnosis path is now traceable through `diagnosis_session`, `agent_step`, `tool_invocation`, and `GET /api/diagnosis/{sessionId}/trace`. The older `/api/ai_ops` endpoint still acts like a standalone SSE demo: it accepts no alert payload, generates an internal session id, and does not make trace replay obvious to callers.
## Goal
Turn AIOps into an alert-triggered diagnosis entry point that shares the same evidence and trace story as the main MVP, without rewriting the whole AIOps flow.
## Scope
- In scope:
- Optional AIOps alert request body.
- Stable request/session id propagation.
- Persisted AIOps query summary and final answer.
- SSE session id event.
- Demo documentation and focused tests.
- Out of scope:
- Full AIOps and ChatService unification.
- AIOps Verifier integration.
- Database schema changes.
- Sensitive configuration cleanup.
- Fully offline runtime.
## Related OpenSpec
`openspec/changes/aiops-traceable-diagnosis-entry/`
@@ -0,0 +1,68 @@
# Decisions: aiops-traceable-diagnosis-entry
## Clarify
- Entry summary: make the legacy AIOps SSE endpoint a traceable alert diagnosis entry for the Agent Engineer interview MVP.
- Slug: `aiops-traceable-diagnosis-entry`
- Scale: standard-light, because this extends one public endpoint and reuses existing persistence/trace infrastructure.
## Context
- `mvp-demo-trace-acceptance` already added `GET /api/diagnosis/{sessionId}/trace`.
- `chat-verifier-agent` made the chat path stronger than the older AIOps path.
- Current AIOps value is as a second entry point: system alert -> automated diagnosis -> evidence trace.
## Grill Question Pool
| # | Dimension | Question | Mode | Status |
|---|---|---|---|---|
| Q1 | Positioning | Is AIOps an independent product path or an alert-triggered sibling of Chat Diagnosis? | user-interview | Resolved: sibling entry, unified trace story |
| Q2 | API | Should we keep `/api/ai_ops` or add a new endpoint? | evidence-driven | Resolved: keep existing endpoint and extend optional body |
| Q3 | Input | What is the minimum alert payload? | user-interview | Resolved: `sessionId`, `alertName`, `service`, `severity`, `description`, `timeRange`, plus `userRequest` fallback |
| Q4 | Output | How does the caller learn the trace session id? | evidence-driven | Resolved: first SSE event uses type `session` |
| Q5 | Trace | Must AIOps be replayable with existing trace API? | evidence-driven | Resolved: yes, this is the main acceptance criterion |
| Q6 | Verifier | Must this slice add AIOps Verifier? | user-interview | Resolved: no, defer as follow-up |
| Q7 | Compatibility | Should no-body calls still work? | evidence-driven | Resolved: yes, preserve old demo behavior |
| Q8 | GitNexus | Should unavailable GitNexus block implementation? | user-interview | Resolved: skip GitNexus by user decision |
## Evidence-Driven Conclusions
| Conclusion | Evidence Source | Result |
|---|---|---|
| AIOps is currently isolated from request-driven trace replay. | `ChatController.aiOps()` has no request body; `AiOpsService` creates its own random session id. | Extend endpoint and service. |
| No schema change is needed. | `DiagnosisSession` already has `query`, `agentFlow`, `answer`, counts, and status. | Reuse existing table. |
| Trace API can already replay AIOps if session id and answer are persisted. | `DiagnosisTraceService` loads by session id and is flow-agnostic. | Keep trace API unchanged. |
| Blast radius is moderate and local. | `rg` shows only `ChatController` calls `executeAiOpsAnalysis` and `extractFinalReport`. | Change service/controller carefully and add tests. |
## User-Interview Confirmations
| Topic | User Words | Decision |
|---|---|---|
| Use sm-flow | "可以,改造一下AIOps 接口,用sm-flow流程看看" | Use OpenSpec + devflow. |
| GitNexus | "跳过gitnexus把" | Record skip and use local impact analysis. |
| Proceed after Grill | "可以" | Continue with lightweight Grill conclusions. |
## Key Decisions
- Keep `/api/ai_ops` and make its body optional.
- Emit `SseMessage.type=session` before long-running analysis starts.
- Store AIOps request summary in `diagnosis_session.query`.
- Store final report in `diagnosis_session.answer`.
- Defer AIOps Verifier to a later change so this slice stays focused.
## Architecture Audit
```text
POST /api/ai_ops
-> optional AIOpsRequest
-> resolve sessionId
-> create diagnosis_session(agentFlow=AI_OPS)
-> run ai_ops_supervisor(planner, executor)
-> AgentLoggingHook persists steps
-> tools persist invocations under SessionContextHolder
-> extract final report
-> persist answer
-> GET /api/diagnosis/{sessionId}/trace replays the run
```
Risk level: medium. The endpoint is public and SSE-based, but the change is additive and does not change the chat diagnosis path or database schema.
@@ -0,0 +1,37 @@
# Evidence: aiops-traceable-diagnosis-entry
## Local Impact Analysis
- `ChatController.aiOps()` is the only caller of `AiOpsService.executeAiOpsAnalysis(...)`.
- `ChatController.aiOps()` is the only caller of `AiOpsService.extractFinalReport(...)`.
- `AIOpsRequest` exists but only has `userRequest`; no current controller consumes it.
- `DiagnosisTraceService` is flow-agnostic and reads persisted session/step/tool records by `sessionId`.
## GitNexus
GitNexus MCP tools were not exposed in this session. The user explicitly approved skipping GitNexus for this change. Local impact analysis and targeted tests are used instead.
## Expected Verification
- Focused unit tests for AIOps request/session/report helper behavior.
- Compile verification.
- OpenSpec validation if CLI is available.
## Verification Results
- `openspec.cmd validate aiops-traceable-diagnosis-entry --strict`: passed.
- `mvn -q "-Dtest=AiOpsServiceTest,DiagnosisTraceServiceTest" test`: passed after rerun with approved Maven access.
- `mvn -q -DskipTests compile`: passed.
## Demo Alignment
- Added `knowledge_base/troubleshooting/aiops-alert-runbook.md` so mock AIOps alerts have matching knowledge-base guidance.
- Aligned the documented AIOps demo with mock data: `HighCPUUsage` on `payment-service`, using `system-metrics` evidence.
## Metric Alignment Follow-up
- Runtime verification showed `diagnosis_session.tool_call_count` counted agent steps with tool calls, while trace returned actual `tool_invocation` records.
- Updated `ChatService` and `AiOpsService` metric backfill to use `ToolInvocationRepository.countBySessionId(sessionId)`.
- Targeted verification:
- `mvn -q "-Dtest=AiOpsServiceTest,ChatServiceSequentialAgentTest,DiagnosisTraceServiceTest" test`: passed.
- `mvn -q -DskipTests compile`: passed.
@@ -0,0 +1,36 @@
# Acceptance: diagnosis-eval-harness
## Classification
standard-light
## Task Status
| Task | Status | Notes |
| --- | --- | --- |
| Issue and OpenSpec setup | Done | `ISS-006` and initial OpenSpec artifacts were created. |
| Implementation | Done | Added fixed cases, fixture-mode trace evaluation, aggregate metrics, and JSON / Markdown report writer. |
| Verification | Done | Targeted evaluator tests, compile verification, and OpenSpec validation passed. |
## Current State
- First implementation uses fixture-mode evaluation.
- Live trace API polling remains a follow-up option.
## Verification
### Script Verification
- Command: `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
- Result: passed
- Notes: Covers fixed case loading, fixture evaluation, missing fixture reporting, reject degraded-output validation, and report writing.
### Static Verification
- Command: `mvn -q -DskipTests compile`
- Result: passed
### OpenSpec Verification
- Command: `openspec validate diagnosis-eval-harness --strict`
- Result: passed
@@ -0,0 +1,31 @@
# Brief: diagnosis-eval-harness
## Background
The MVP has a runnable demo and hardened evidence trace semantics, but it still lacks a fixed regression baseline for Agent diagnosis quality. P1-B creates a small evaluation harness that can validate diagnosis traces against fixed cases and produce repeatable reports.
## Goals
1. Define fixed diagnosis cases for the MVP demo domain.
2. Validate trace evidence, verifier verdicts, answer keywords, and degraded-output behavior.
3. Produce JSON and Markdown reports for interview and regression use.
4. Keep the first version offline by supporting trace fixtures.
## Scope
- Evaluation case definitions
- Trace fixture shape
- Rule-based evaluator
- JSON / Markdown report output
- Focused offline tests and docs
## Non-Goals
- No LLM-as-judge
- No live end-to-end runtime requirement
- No production API
- No chat or verifier runtime change
## Related OpenSpec
`openspec/changes/diagnosis-eval-harness/`
@@ -0,0 +1,28 @@
# Diagnosis Eval Harness Decisions
## Clarify
- Entry summary: build P1-B fixed case evaluation after evidence trace hardening.
- Slug: `diagnosis-eval-harness`
- Devflow scale: standard-light
## Context
- P1-A `evidence-trace-hardening` created stable evidence semantics for supported, no-evidence, deduped, and failed tool calls.
- The MVP demo trace API already provides an aggregate trace shape suitable for evaluation.
- The first evaluator should avoid depending on external infrastructure so it can run in regular development.
## Key Decisions
- Decision: Start with rule-based trace validation instead of LLM-as-judge.
- Reason: The first regression signal should be deterministic and tied to trace contracts.
- Decision: Support offline fixture traces first.
- Reason: This makes the harness usable without MySQL, Redis, Milvus, or a real LLM.
- Decision: Output both JSON and Markdown.
- Reason: JSON supports automation; Markdown is easier to discuss in interviews.
## Open Questions
- Whether live trace API polling belongs in this change or a follow-up after fixture mode lands.
@@ -0,0 +1,10 @@
# Diagnosis Eval Harness Evidence
## Evidence
| Source | Evidence | Conclusion | Reported |
|---|---|---|---|
| `openspec/specs/evidence-trace-hardening/spec.md` | Defines stable evidence states and summary behavior | Evaluation can rely on trace semantics rather than ad hoc log parsing | Yes |
| `mvp/demo/README.md` | Documents an end-to-end demo flow with chat, trace, and feedback | Existing demo flow provides the runtime story, but not a reusable evaluation baseline | Yes |
| `DiagnosisTraceService` | Aggregates session, steps, tools, and self-evaluation | Trace response shape can be reused as evaluation input | Yes |
| `ToolTraceSummaryService` | Builds verifier-facing evidence summaries from persisted tool rows | Evaluator can check evidence coverage through persisted trace artifacts | Yes |
@@ -0,0 +1,32 @@
# Acceptance: evidence-trace-hardening
## Classification
standard-light
## Task Status
| Task | Status | Notes |
| --- | --- | --- |
| Issue and OpenSpec setup | Done | `ISS-005` and the initial OpenSpec artifacts were created. |
| Implementation | Done | Recorder contract, lookup persistence path, evidence summary semantics, and degraded-path tests were implemented. |
| Verification | Done | Targeted offline tests and compile verification passed. |
## Verification
### Script Verification
- Command: `mvn -q "-Dtest=ToolInvocationRecorderTest,ToolTraceSummaryServiceTest,ChatServiceSequentialAgentTest,LookupKnowledgeToolTest" test`
- Result: passed
- Notes: Covers recorder contract, summary semantics for success/failure/no-evidence, and `ChatService` fallback / degraded paths.
### Static Verification
- Command: `mvn -q -DskipTests compile`
- Result: passed
## Open Questions
| Question | Current position |
| --- | --- |
| Should deduped retrievals be counted separately from generic no-hit events in future evaluation metrics? | Deferred to P1-B; this change preserves enough structure to decide later. |
@@ -0,0 +1,32 @@
# Brief: evidence-trace-hardening
## Background
The MVP already has persisted tool traces and a verifier, but the evidence contract is still only partially standardized. For interview-focused hardening, the project now needs a tighter contract for evidence persistence, no-evidence / failure semantics, and degraded-output behavior.
## Goals
1. Standardize the persisted evidence-tool contract across `lookup_knowledge`, `query_logs`, and `query_metrics`.
2. Make verifier-facing summaries distinguish failed calls, no-hit calls, deduped retrievals, and actual supporting evidence.
3. Add offline tests for verifier fallback and degraded-output paths.
## Scope
- `ToolInvocationRecorder`
- `LookupKnowledgeTool`
- `QueryLogsTools`
- `QueryMetricsTools`
- `ToolTraceSummaryService`
- `ChatService`
- Focused offline tests
## Non-Goals
- No new API or schema
- No evaluation harness yet
- No trace UI
- No security/config cleanup
## Related OpenSpec
`openspec/changes/evidence-trace-hardening/`
@@ -0,0 +1,24 @@
# Evidence Trace Hardening Decisions
## Clarify
- Entry summary: harden the MVP evidence contract before building the P1-B evaluation harness.
- Slug: `evidence-trace-hardening`
- Devflow scale: standard-light
## Context
- `ISS-003` raised verifier traceability and failure-path concerns.
- Current code inspection shows `QueryLogsTools` and `QueryMetricsTools` already use `ToolInvocationRecorder`, while `LookupKnowledgeTool` still persists rows through a local helper.
- `ChatService` already contains fallback behavior for missing/invalid `verifier_output`, but coverage is narrow.
## Key Decisions
- Decision: Treat this as a contract-hardening change, not a new feature change.
- Reason: The project already has the necessary runtime pieces; the gap is semantic consistency and testability.
- Decision: Keep the scope before P1-B.
- Reason: The evaluation harness will rely on stable evidence semantics, so this contract slice should land first.
- Decision: Preserve schema and API stability.
- Reason: The interview value here is engineering rigor, not more surface area.
@@ -0,0 +1,11 @@
# Evidence Trace Hardening Evidence
## Evidence
| Source | Evidence | Conclusion | Reported |
|---|---|---|---|
| `ToolInvocationRecorder` | Provides a common persistence seam for evidence tools | Contract hardening should build on the existing recorder instead of introducing a new store path | Yes |
| `LookupKnowledgeTool` | Still constructs `ToolInvocation` rows through a local helper | Retrieval-aware evidence persistence is not yet unified with the recorder contract | Yes |
| `QueryLogsTools` / `QueryMetricsTools` | Already record evidence invocations through `recordEvidenceTool(...)` | Current gap is semantic alignment, not missing persistence | Yes |
| `ToolTraceSummaryService` | Merges rows by tool and topic domain and infers evidence level heuristically | Summary rules need explicit handling for failure, no-hit, and dedup cases | Yes |
| `ChatService` | Falls back to `LOW_CONFID` when verifier output is missing or invalid | These degraded paths exist and should now be covered by focused offline tests | Yes |
@@ -0,0 +1,37 @@
# Acceptance: expand-diagnosis-eval-fixtures
## Classification
standard-light
## Task Status
| Task | Status | Notes |
| --- | --- | --- |
| Issue and OpenSpec setup | Done | Created slug-based issue and initial OpenSpec artifacts. |
| Implementation | Done | Added remaining fixtures, full baseline reports, and documentation updates. |
| Verification | Done | Evaluator tests, compile verification, and OpenSpec validation passed. |
## Current State
- Fixture coverage is complete for the five fixed diagnosis cases.
- Baseline reports are saved under `mvp/eval/reports`.
- No production runtime behavior has been changed.
## Verification
### Script Verification
- Command: `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
- Result: passed
- Notes: Covers full fixture coverage, baseline report matching, reject degraded-output validation, and report writing.
### Static Verification
- Command: `mvn -q -DskipTests compile`
- Result: passed
### OpenSpec Verification
- Command: `openspec validate expand-diagnosis-eval-fixtures --strict`
- Result: passed
@@ -0,0 +1,31 @@
# Brief: expand-diagnosis-eval-fixtures
## Background
The diagnosis eval harness is implemented and archived, but the fixed baseline is incomplete because three of the five diagnosis cases still reference missing fixtures.
## Goals
1. Add representative trace fixtures for all remaining fixed diagnosis cases.
2. Save a reproducible baseline report in JSON and Markdown.
3. Document how to regenerate and interpret the baseline.
4. Keep evaluation offline and deterministic.
## Scope
- Redis timeout fixture
- Slow response fixture
- JVM memory risk fixture
- Baseline reports under `mvp/eval/reports`
- Focused tests for full fixture coverage and report generation
## Non-Goals
- No new diagnosis cases
- No production Agent runtime changes
- No LLM-as-judge
- No live infrastructure requirement
## Related OpenSpec
`openspec/changes/expand-diagnosis-eval-fixtures/`
@@ -0,0 +1,28 @@
# Expand Diagnosis Eval Fixtures Decisions
## Clarify
- Entry summary: complete the fixed diagnosis eval baseline after the harness is in place.
- Slug: `expand-diagnosis-eval-fixtures`
- Devflow scale: standard-light
## Context
- `diagnosis-eval-harness` created the evaluator, case file, fixture mode, and report writer.
- The first baseline still has missing fixtures by design.
- This follow-up turns that partial baseline into a full fixed-case baseline.
## Key Decisions
- Decision: Keep this change data-focused.
- Reason: the evaluator rules already landed; this change should not blur fixture expansion with harness behavior changes.
- Decision: Save baseline reports in the repository.
- Reason: interview review and future diffs are easier when the expected baseline is visible.
- Decision: Use deterministic fixture traces instead of live trace generation.
- Reason: this baseline should run without infrastructure or external model calls.
## Open Questions
- Whether a future change should add a CLI or Maven goal for report regeneration.
@@ -0,0 +1,11 @@
# Evidence: expand-diagnosis-eval-fixtures
## Evidence Log
- 2026-07-04: Created slug-based issue `expand-diagnosis-eval-fixtures.md`.
- 2026-07-04: Created OpenSpec change `expand-diagnosis-eval-fixtures`.
- 2026-07-04: Added Redis timeout, slow response, and JVM memory risk fixtures.
- 2026-07-04: Added baseline JSON and Markdown reports under `mvp/eval/reports`.
- 2026-07-04: Verification passed with `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`.
- 2026-07-04: Verification passed with `mvn -q -DskipTests compile`.
- 2026-07-04: Verification passed with `openspec validate expand-diagnosis-eval-fixtures --strict`.
@@ -0,0 +1,37 @@
# Acceptance: diagnosis-eval-baseline-diff
## Classification
standard-light
## Task Status
| Task | Status | Notes |
| --- | --- | --- |
| Issue and OpenSpec setup | Done | Created slug-based issue and initial OpenSpec artifacts. |
| Implementation | Done | Added diff model, comparator, writer, docs, sample outputs, and focused tests. |
| Verification | Done | Diff/evaluator tests, compile verification, and OpenSpec validation passed. |
## Current State
- Baseline diff is implemented for aggregate metrics, verdict distribution, case-level state, keyword coverage, evidence coverage, missing cases, and new cases.
- JSON and Markdown diff output are available.
- No production runtime behavior has been changed.
## Verification
### Script Verification
- Command: `mvn -q "-Dtest=DiagnosisEvalBaselineDiffTest" test`
- Result: passed
- Notes: Also verified with `DiagnosisTraceEvaluatorTest`.
### Static Verification
- Command: `mvn -q -DskipTests compile`
- Result: passed
### OpenSpec Verification
- Command: `openspec validate diagnosis-eval-baseline-diff --strict`
- Result: passed
@@ -0,0 +1,30 @@
# Brief: diagnosis-eval-baseline-diff
## Background
The eval harness now has a complete saved baseline. This change adds the comparison layer that turns the baseline into an actionable regression signal.
## Goals
1. Compare baseline and current `DiagnosisEvalReport` objects.
2. Detect aggregate and per-case regressions.
3. Output JSON and Markdown diff reports.
4. Document how to read the diff in interview and engineering terms.
## Scope
- Diff data structures
- Deterministic report comparison
- JSON / Markdown diff output
- Focused tests and eval docs
## Non-Goals
- No live Agent execution
- No LLM-as-judge
- No evaluator scoring rule changes
- No production API changes
## Related OpenSpec
`openspec/changes/diagnosis-eval-baseline-diff/`
@@ -0,0 +1,28 @@
# Diagnosis Eval Baseline Diff Decisions
## Clarify
- Entry summary: add report diffing on top of the completed diagnosis eval baseline.
- Slug: `diagnosis-eval-baseline-diff`
- Devflow scale: standard-light
## Context
- `diagnosis-eval-harness` created deterministic fixture evaluation.
- `expand-diagnosis-eval-fixtures` created a complete saved baseline.
- This change compares new reports against that baseline.
## Key Decisions
- Decision: Diff report DTOs instead of raw traces.
- Reason: the report is the stable contract for regression review.
- Decision: Use deterministic code rules instead of LLM-as-judge.
- Reason: baseline regression checks should be repeatable and explainable.
- Decision: Output both JSON and Markdown.
- Reason: JSON supports automation; Markdown is useful in reviews and interviews.
## Open Questions
- Whether a future change should expose this through a CLI or Maven goal.
@@ -0,0 +1,11 @@
# Evidence: diagnosis-eval-baseline-diff
## Evidence Log
- 2026-07-05: Created slug-based issue `diagnosis-eval-baseline-diff.md`.
- 2026-07-05: Created OpenSpec change `diagnosis-eval-baseline-diff`.
- 2026-07-05: Added baseline diff DTOs, deterministic comparer, and JSON / Markdown writer.
- 2026-07-05: Added sample baseline diff JSON and Markdown reports.
- 2026-07-05: Verification passed with `mvn -q "-Dtest=DiagnosisEvalBaselineDiffTest,DiagnosisTraceEvaluatorTest" test`.
- 2026-07-05: Verification passed with `mvn -q -DskipTests compile`.
- 2026-07-05: Verification passed with `openspec validate diagnosis-eval-baseline-diff --strict`.
@@ -0,0 +1,25 @@
# Acceptance: mvp-demo-interview-runbook
## Classification
standard-light
## Task Status
| Task | Status | Notes |
| --- | --- | --- |
| Issue and OpenSpec setup | Done | Created slug-based issue and OpenSpec artifacts. |
| Implementation | Done | Added request payload, runnable script, output directory docs, interview walkthrough, and trace checklist. |
| Verification | Done | OpenSpec validation passed. |
## Current State
- No backend runtime behavior has been changed.
- Demo is packaged under `mvp/demo` for interview use.
## Verification
### OpenSpec Verification
- Command: `openspec validate mvp-demo-interview-runbook --strict`
- Result: passed
@@ -0,0 +1,29 @@
# Brief: mvp-demo-interview-runbook
## Background
Plan C is the interview-facing demo package. The project has the engineering pieces, but needs a single place to run and explain the MVP flow.
## Goals
1. Provide a fixed payment-timeout request payload.
2. Provide a PowerShell script that runs chat, trace, and feedback.
3. Save demo responses under `mvp/demo/output`.
4. Add interview walkthrough and trace checklist.
## Scope
- Demo docs and scripts only
- Existing local APIs only
- Existing `mvp-demo` profile only
## Non-Goals
- No backend code changes
- No eval extension
- No secret cleanup
- No full offline runtime
## Related OpenSpec
`openspec/changes/mvp-demo-interview-runbook/`
@@ -0,0 +1,27 @@
# MVP Demo Interview Runbook Decisions
## Clarify
- Entry summary: package existing MVP capabilities into a repeatable interview demo.
- Slug: `mvp-demo-interview-runbook`
- Devflow scale: standard-light
## Context
- Evidence trace and eval baseline work are already done.
- The next useful step is not more eval tooling, but a runnable demo path.
## Key Decisions
- Decision: Keep this change documentation/script-only.
- Reason: Plan C is about demo packaging, not new runtime capability.
- Decision: Use a stable session id.
- Reason: it makes trace lookup and saved output predictable.
- Decision: Save outputs to `mvp/demo/output`.
- Reason: generated artifacts should be easy to review without mixing into source fixtures.
## Open Questions
- Whether a later change should add a truly offline stubbed demo mode.
@@ -0,0 +1,9 @@
# Evidence: mvp-demo-interview-runbook
## Evidence Log
- 2026-07-05: Created Plan C demo packaging issue and OpenSpec change.
- 2026-07-05: Added fixed payment-timeout request payload.
- 2026-07-05: Added PowerShell demo script for chat, trace, and feedback.
- 2026-07-05: Added interview walkthrough and trace inspection checklist.
- 2026-07-05: Verification passed with `openspec validate mvp-demo-interview-runbook --strict`.
+86
View File
@@ -0,0 +1,86 @@
# RAG Retrieval Baseline
This directory contains the offline retrieval baseline for the RAG refactor.
The baseline is intentionally narrower than full diagnosis evaluation. It checks
whether fixed retrieval queries can recover expected documents, breadcrumbs, and
evidence keywords before changing L0 behavior, query augmentation, evidence
post-processing, or Spring AI VectorStore integration.
## Layout
```text
eval/rag-retrieval/
cases/golden-cases.json Fixed retrieval golden cases
fixtures/*.json Saved retrieval candidates for each case
reports/baseline.json Machine-readable baseline report
reports/baseline.md Human-readable baseline report
reports/live-post-reindex.* Optional live acceptance reports
```
## Run
From the repository root:
```bash
python scripts/eval_rag_retrieval.py
```
Custom paths are also supported:
```bash
python scripts/eval_rag_retrieval.py \
--cases eval/rag-retrieval/cases/golden-cases.json \
--fixtures eval/rag-retrieval/fixtures \
--json-report eval/rag-retrieval/reports/baseline.json \
--markdown-report eval/rag-retrieval/reports/baseline.md
```
## Hit Levels
- `strong`: expected document is found and breadcrumb or evidence keyword coverage is satisfied.
- `medium`: expected document is found, but breadcrumb or keyword coverage is incomplete.
- `weak`: expected evidence keyword is found, but expected document is missing.
- `miss`: expected document and expected evidence are not found.
`Recall@K` counts `strong` and `medium` as retrieved.
## Scope
This baseline runs fully offline and does not call MySQL, Redis, Milvus, an LLM,
or the Spring Boot application. It is a regression harness for retrieval behavior,
not a claim that live production retrieval accuracy is complete.
## Live Post-Reindex Acceptance
When embedding input changes, existing vectors do not update by themselves. For
example, after adding `title` and `breadcrumb` to the embedding text, the live
Milvus/Zilliz collection must be reindexed before retrieval can reflect that new
semantic signal.
Use this optional live acceptance flow after the application is running and the
knowledge base has been reindexed:
```bash
python scripts/eval_rag_live_acceptance.py
```
Custom service URL and output paths are supported:
```bash
python scripts/eval_rag_live_acceptance.py \
--base-url http://127.0.0.1:9900 \
--json-report eval/rag-retrieval/reports/live-post-reindex.json \
--markdown-report eval/rag-retrieval/reports/live-post-reindex.md
```
The script calls:
```text
GET /api/search/similar
```
It writes JSON and Markdown reports with query, topK, result count, top
candidates, breadcrumb, score labels, and raw response fields. This is a live
smoke check for environment readiness and post-reindex behavior; it does not
replace the deterministic offline baseline above.
@@ -0,0 +1,61 @@
{
"version": 1,
"description": "Offline golden retrieval cases for RAG refactor baseline.",
"topK": 5,
"cases": [
{
"caseId": "chat-mysql-connection-pool",
"scenario": "chat",
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
"expectedDocIds": ["mysql-connection-pool"],
"expectedBreadcrumbs": ["Database > MySQL > Connection Pool"],
"expectedKeywords": ["connection pool", "max_connections", "HikariCP"],
"notes": "Covers precise database troubleshooting retrieval."
},
{
"caseId": "chat-diagnosis-flow",
"scenario": "chat",
"query": "What is the standard troubleshooting flow for an application incident?",
"expectedDocIds": ["incident-diagnosis-flow"],
"expectedBreadcrumbs": ["AIOps > Diagnosis Flow"],
"expectedKeywords": ["collect evidence", "verify", "remediation"],
"notes": "Covers process-style knowledge where breadcrumb matters."
},
{
"caseId": "aiops-payment-latency-alert",
"scenario": "aiops",
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
"expectedDocIds": ["payment-service-latency"],
"expectedBreadcrumbs": ["AIOps > Service Alerts > Payment Latency"],
"expectedKeywords": ["p95 latency", "payment-service", "downstream dependency"],
"notes": "Covers alert payload terms that should become retrieval hints."
},
{
"caseId": "aiops-prometheus-alert-scope",
"scenario": "aiops",
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
"expectedDocIds": ["aiops-alert-scope-control"],
"expectedBreadcrumbs": ["AIOps > Alert Scope Control"],
"expectedKeywords": ["payload", "unrelated active alerts", "scope"],
"notes": "Covers scoped alert diagnosis behavior."
},
{
"caseId": "chat-rag-chunk-context",
"scenario": "chat",
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
"expectedDocIds": ["rag-chunk-context-reconstruction"],
"expectedBreadcrumbs": ["RAG > Chunking > Context Reconstruction"],
"expectedKeywords": ["neighbor chunk", "same section", "breadcrumb"],
"notes": "Covers the known RAG refactor issue around context reconstruction."
},
{
"caseId": "chat-l0-domain-hint",
"scenario": "chat",
"query": "Should L0 keyword matching decide the final retrieval result?",
"expectedDocIds": ["rag-l0-domain-entity-hint"],
"expectedBreadcrumbs": ["RAG > L0 > Domain Entity Hint"],
"expectedKeywords": ["domain detector", "entity extractor", "metadata filter"],
"notes": "Covers the target L0 role after refactor."
}
]
}
@@ -0,0 +1,25 @@
{
"caseId": "aiops-payment-latency-alert",
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
"retrievedAt": "2026-07-05T00:00:00Z",
"candidates": [
{
"rank": 1,
"docId": "payment-service-latency",
"title": "Payment Service Latency Alert Playbook",
"breadcrumb": "AIOps > Service Alerts > Payment Latency",
"content": "For payment-service p95 latency alerts, check downstream dependency latency, thread pool saturation, gateway retries, and recent deployment changes.",
"score": 0.84,
"retrievalLayer": "L1"
},
{
"rank": 2,
"docId": "mysql-connection-pool",
"title": "MySQL Connection Pool Troubleshooting",
"breadcrumb": "Database > MySQL > Connection Pool",
"content": "Database connection pool saturation can increase payment latency when checkout paths wait for connections.",
"score": 0.68,
"retrievalLayer": "L1"
}
]
}
@@ -0,0 +1,16 @@
{
"caseId": "aiops-prometheus-alert-scope",
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
"retrievedAt": "2026-07-05T00:00:00Z",
"candidates": [
{
"rank": 1,
"docId": "aiops-alert-scope-control",
"title": "AIOps Alert Scope Control",
"breadcrumb": "AIOps > Alert Scope Control",
"content": "When payload mode is active, queryPrometheusAlerts can verify the supplied alert, but unrelated active alerts must remain scoped context and should not become full diagnoses.",
"score": 0.9,
"retrievalLayer": "L0+L1"
}
]
}
@@ -0,0 +1,25 @@
{
"caseId": "chat-diagnosis-flow",
"query": "What is the standard troubleshooting flow for an application incident?",
"retrievedAt": "2026-07-05T00:00:00Z",
"candidates": [
{
"rank": 1,
"docId": "incident-diagnosis-flow",
"title": "Incident Diagnosis Flow",
"breadcrumb": "AIOps > Diagnosis Flow",
"content": "The standard flow is to collect evidence, identify the suspected fault domain, verify the hypothesis, apply remediation, and confirm recovery.",
"score": 0.82,
"retrievalLayer": "L1"
},
{
"rank": 2,
"docId": "rag-chunk-context-reconstruction",
"title": "RAG Chunk Context Reconstruction",
"breadcrumb": "RAG > Chunking > Context Reconstruction",
"content": "Long sections may require neighbor chunk expansion and breadcrumb-aware packing.",
"score": 0.55,
"retrievalLayer": "L1"
}
]
}
@@ -0,0 +1,25 @@
{
"caseId": "chat-l0-domain-hint",
"query": "Should L0 keyword matching decide the final retrieval result?",
"retrievedAt": "2026-07-05T00:00:00Z",
"candidates": [
{
"rank": 1,
"docId": "rag-l0-domain-entity-hint",
"title": "RAG L0 Domain Entity Hint",
"breadcrumb": "RAG > L0 > Domain Entity Hint",
"content": "L0 should be retained as a domain detector, entity extractor, metadata filter generator, and explainability signal, not as the final retrieval decision.",
"score": 0.88,
"retrievalLayer": "L0"
},
{
"rank": 2,
"docId": "rag-l0-l1-fusion-ranking",
"title": "RAG L0 L1 Fusion Ranking",
"breadcrumb": "RAG > Ranking > Fusion",
"content": "L0 and L1 candidates should eventually be fused rather than handled as an early-return branch.",
"score": 0.75,
"retrievalLayer": "L1"
}
]
}
@@ -0,0 +1,25 @@
{
"caseId": "chat-mysql-connection-pool",
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
"retrievedAt": "2026-07-05T00:00:00Z",
"candidates": [
{
"rank": 1,
"docId": "mysql-connection-pool",
"title": "MySQL Connection Pool Troubleshooting",
"breadcrumb": "Database > MySQL > Connection Pool",
"content": "When the connection pool is exhausted, inspect HikariCP active connections, max_connections, slow SQL, leak detection, and database wait events.",
"score": 0.86,
"retrievalLayer": "L0+L1"
},
{
"rank": 2,
"docId": "incident-diagnosis-flow",
"title": "Incident Diagnosis Flow",
"breadcrumb": "AIOps > Diagnosis Flow",
"content": "Collect evidence, compare metrics and logs, then verify remediation before closing the incident.",
"score": 0.61,
"retrievalLayer": "L1"
}
]
}
@@ -0,0 +1,25 @@
{
"caseId": "chat-rag-chunk-context",
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
"retrievedAt": "2026-07-05T00:00:00Z",
"candidates": [
{
"rank": 1,
"docId": "rag-chunk-context-reconstruction",
"title": "RAG Chunk Context Reconstruction",
"breadcrumb": "RAG > Chunking > Context Reconstruction",
"content": "After a chunk hit, expand to neighbor chunk candidates from the same section and preserve breadcrumb metadata in the evidence pack.",
"score": 0.79,
"retrievalLayer": "L1"
},
{
"rank": 2,
"docId": "rag-breadcrumb-embedding-gap",
"title": "RAG Breadcrumb Embedding Gap",
"breadcrumb": "RAG > Embedding > Breadcrumb",
"content": "Embedding title and breadcrumb with content helps recover section semantics.",
"score": 0.72,
"retrievalLayer": "L1"
}
]
}
+1
View File
@@ -0,0 +1 @@
+131
View File
@@ -0,0 +1,131 @@
{
"generatedAt": "2026-07-04T17:59:52.172759+00:00",
"caseFile": "eval/rag-retrieval/cases/golden-cases.json",
"fixtureDir": "eval/rag-retrieval/fixtures",
"aggregate": {
"caseCount": 6,
"topK": 5,
"strongHitCount": 6,
"mediumHitCount": 0,
"weakHitCount": 0,
"missCount": 0,
"recallAtK": 1.0,
"strongHitRate": 1.0,
"averageFirstHitRank": 1.0
},
"results": [
{
"caseId": "chat-mysql-connection-pool",
"scenario": "chat",
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
"hitLevel": "strong",
"passed": true,
"firstExpectedRank": 1,
"topCandidates": [
"1:mysql-connection-pool",
"2:incident-diagnosis-flow"
],
"matchedKeywords": [
"connection pool",
"max_connections",
"hikaricp"
],
"breadcrumbMatched": true,
"failedChecks": []
},
{
"caseId": "chat-diagnosis-flow",
"scenario": "chat",
"query": "What is the standard troubleshooting flow for an application incident?",
"hitLevel": "strong",
"passed": true,
"firstExpectedRank": 1,
"topCandidates": [
"1:incident-diagnosis-flow",
"2:rag-chunk-context-reconstruction"
],
"matchedKeywords": [
"collect evidence",
"verify",
"remediation"
],
"breadcrumbMatched": true,
"failedChecks": []
},
{
"caseId": "aiops-payment-latency-alert",
"scenario": "aiops",
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
"hitLevel": "strong",
"passed": true,
"firstExpectedRank": 1,
"topCandidates": [
"1:payment-service-latency",
"2:mysql-connection-pool"
],
"matchedKeywords": [
"p95 latency",
"payment-service",
"downstream dependency"
],
"breadcrumbMatched": true,
"failedChecks": []
},
{
"caseId": "aiops-prometheus-alert-scope",
"scenario": "aiops",
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
"hitLevel": "strong",
"passed": true,
"firstExpectedRank": 1,
"topCandidates": [
"1:aiops-alert-scope-control"
],
"matchedKeywords": [
"payload",
"unrelated active alerts",
"scope"
],
"breadcrumbMatched": true,
"failedChecks": []
},
{
"caseId": "chat-rag-chunk-context",
"scenario": "chat",
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
"hitLevel": "strong",
"passed": true,
"firstExpectedRank": 1,
"topCandidates": [
"1:rag-chunk-context-reconstruction",
"2:rag-breadcrumb-embedding-gap"
],
"matchedKeywords": [
"neighbor chunk",
"same section",
"breadcrumb"
],
"breadcrumbMatched": true,
"failedChecks": []
},
{
"caseId": "chat-l0-domain-hint",
"scenario": "chat",
"query": "Should L0 keyword matching decide the final retrieval result?",
"hitLevel": "strong",
"passed": true,
"firstExpectedRank": 1,
"topCandidates": [
"1:rag-l0-domain-entity-hint",
"2:rag-l0-l1-fusion-ranking"
],
"matchedKeywords": [
"domain detector",
"entity extractor",
"metadata filter"
],
"breadcrumbMatched": true,
"failedChecks": []
}
]
}
+28
View File
@@ -0,0 +1,28 @@
# RAG Retrieval Baseline
Generated at: `2026-07-04T17:59:52.172759+00:00`
## Aggregate
| Metric | Value |
|---|---:|
| Cases | 6 |
| Top K | 5 |
| Recall@K | 1.0 |
| Strong hit rate | 1.0 |
| Strong hits | 6 |
| Medium hits | 0 |
| Weak hits | 0 |
| Misses | 0 |
| Average first hit rank | 1.0 |
## Cases
| Case | Scenario | Hit | First Expected Rank | Top Candidates | Failed Checks |
|---|---|---|---:|---|---|
| chat-mysql-connection-pool | chat | strong | 1 | 1:mysql-connection-pool<br>2:incident-diagnosis-flow | |
| chat-diagnosis-flow | chat | strong | 1 | 1:incident-diagnosis-flow<br>2:rag-chunk-context-reconstruction | |
| aiops-payment-latency-alert | aiops | strong | 1 | 1:payment-service-latency<br>2:mysql-connection-pool | |
| aiops-prometheus-alert-scope | aiops | strong | 1 | 1:aiops-alert-scope-control | |
| chat-rag-chunk-context | chat | strong | 1 | 1:rag-chunk-context-reconstruction<br>2:rag-breadcrumb-embedding-gap | |
| chat-l0-domain-hint | chat | strong | 1 | 1:rag-l0-domain-entity-hint<br>2:rag-l0-l1-fusion-ranking | |
+66
View File
@@ -0,0 +1,66 @@
# SuperBizAgent 面试资料包
## 一句话定位
SuperBizAgent 是一个面向企业故障诊断场景的 Agent 工程项目。它把用户问题或 AIOps 告警转换成可追踪的 Agent 执行链路,并把工具证据、模型步骤、最终答案、自评估和用户反馈统一沉淀到诊断 Trace 中。
## 面试重点
- **Agent 编排**:Chat 复杂问题走 `Planner -> Executor -> Verifier`;AIOps 告警入口走 `Supervisor -> Planner / Executor`。
- **工具证据链**:知识库、日志、指标、Prometheus 告警都通过显式工具调用进入链路,并记录到 `tool_invocation`。
- **可追踪诊断**:一次诊断对应一个 `sessionId`,可通过 `GET /api/diagnosis/{sessionId}/trace` 回放。
- **质量门禁**:Chat Verifier 校验 groundedness;AIOps 规则评估检查报告完整性、payload 聚焦和证据工具覆盖。
- **RAG 工程化**:`lookup_knowledge` 是显式 Agent Tool,底层通过 Spring AI VectorStore 主路径 + Milvus SDK fallback。
- **反馈闭环**:用户反馈 `useful` 会沉淀 `case_library`,`not_useful` 保留 bad case 信号。
## 推荐阅读顺序
1. `mvp/architecture/interview-one-pager.md`:一页式架构图和 2-5 分钟讲解。
2. `mvp/demo/ten-minute-interview-demo.md`:10 分钟现场演示脚本。
3. `interview/story-cases.md`:可复用的面试故事案例。
4. `interview/architecture.md`:面试版系统架构。
5. `interview/design-tradeoffs.md`:关键设计取舍。
6. `interview/demo-script.md`:更细的命令式演示脚本。
7. `interview/acceptance-checklist.md`:面试前验收清单。
8. RAG 专题文档:`rag-refactor-story.md`、`rag-vectorstore-interview-notes.md`、`rag-retrieval-quality-report.md`。
## 核心演示链路
### Chat 诊断
```text
POST /api/chat
-> ChatService
-> Planner -> Executor -> Verifier
-> lookup_knowledge / query_logs / query_metrics
-> diagnosis_session + agent_step + tool_invocation
-> GET /api/diagnosis/{sessionId}/trace
-> POST /api/feedback
```
### AIOps 告警诊断
```text
POST /api/ai_ops
-> AiOpsService
-> PAYLOAD_TARGETED / AUTO_DISCOVERY
-> ai_ops_supervisor
-> planner_agent / executor_agent
-> queryPrometheusAlerts + logs + metrics + lookup_knowledge
-> alert report
-> aiops_rule_evaluation
-> GET /api/diagnosis/{sessionId}/trace
```
## 当前完成度
- Chat 诊断链路:可运行、可追踪、有 Verifier。
- AIOps 告警链路:可运行、可追踪、支持 payload scope control。
- RAG 检索链路:Spring AI VectorStore 主路径、Milvus SDK fallback、L0 hint、检索评测 baseline。
- Trace API:统一返回 session、agent steps、tool invocations 和 summary。
- Demo 材料:`mvp/demo/README.md`、`mvp/demo/ten-minute-interview-demo.md`。
## 主叙事
这个项目不是简单调用大模型,而是在做一个可审计、可验证、可回归的 Agent 诊断系统。模型可以规划和推理,但每一步工具证据、最终结论、Verifier 结果和用户反馈都能被 Trace API 回放。面试时重点展示“从问题到证据到答案到验证再到反馈”的闭环。
+166
View File
@@ -0,0 +1,166 @@
# 面试前验收清单
## 1. 环境检查
- 当前分支包含最新架构文档和面试材料。
- MySQL 可连接。
- Redis 可连接。
- Milvus/Zilliz 可连接。
- 模型 API key 可用。
- `mvp-demo` profile 开启 mock Prometheus 和 mock CLS。
启动服务:
```powershell
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
```
编译检查:
```powershell
mvn -q -DskipTests compile
```
目标测试:
```powershell
mvn -q "-Dtest=AiOpsServiceTest,ChatServiceSequentialAgentTest,DiagnosisTraceServiceTest,VectorSearchServiceTest,LookupKnowledgeToolTest" test
```
## 2. Chat Demo 验收
请求:
```powershell
$sessionId = "interview-chat-payment-timeout-001"
$body = @{
Id = $sessionId
Question = "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
} | ConvertTo-Json
Invoke-RestMethod `
-Method Post `
-Uri "http://localhost:9900/api/chat" `
-ContentType "application/json" `
-Body $body
```
验收:
- 返回 `data.success = true`。
- 返回 `data.sessionId = interview-chat-payment-timeout-001`。
- `diagnosis_session.agent_flow = CHAT`。
- Trace API 返回 session、steps、toolInvocations。
- 复杂问题下 trace 中能看到 verifier 相关数据。
SQL:
```powershell
python scripts/query_mysql.py "SELECT session_id, agent_flow, status, step_count, tool_call_count FROM diagnosis_session WHERE session_id='interview-chat-payment-timeout-001'"
```
## 3. AIOps Demo 验收
请求:
```powershell
$aiopsSessionId = "interview-aiops-payment-cpu-001"
$aiopsBody = @{
sessionId = $aiopsSessionId
alertName = "HighCPUUsage"
service = "payment-service"
severity = "P1"
description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
timeRange = "last_15m"
userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
} | ConvertTo-Json
Invoke-WebRequest `
-Method Post `
-Uri "http://localhost:9900/api/ai_ops" `
-ContentType "application/json" `
-Body $aiopsBody
```
验收:
- SSE 首条包含 `type=session`。
- SSE 最后包含 `type=done`。
- `diagnosis_session.agent_flow = AI_OPS`。
- `diagnosis_session.status = SUCCESS`。
- `diagnosis_session.answer` 有最终报告。
- Trace API 返回 AIOps steps 和 tool invocations。
- 报告主章节聚焦 `HighCPUUsage/payment-service`。
- 其他 active alerts 不应展开成独立主根因章节。
- `self_evaluation.aiops_rule_evaluation` 存在。
SQL:
```powershell
python scripts/query_mysql.py "SELECT session_id, agent_flow, status, total_duration_ms, step_count, tool_call_count FROM diagnosis_session WHERE session_id='interview-aiops-payment-cpu-001'"
```
```powershell
python scripts/query_mysql.py "SELECT tool_name, COUNT(*) AS cnt FROM tool_invocation WHERE session_id='interview-aiops-payment-cpu-001' GROUP BY tool_name ORDER BY tool_name"
```
Scope 检查:
```powershell
python scripts/query_mysql.py "SELECT (answer LIKE '%HighCPUUsage%') AS has_main_alert, (answer LIKE '%payment-service%') AS has_service FROM diagnosis_session WHERE session_id='interview-aiops-payment-cpu-001'"
```
## 4. Trace API 验收
```powershell
Invoke-RestMethod `
-Method Get `
-Uri "http://localhost:9900/api/diagnosis/interview-aiops-payment-cpu-001/trace"
```
如果 PowerShell 对长 JSON 或特殊字符不稳定,可以用:
```powershell
curl.exe --silent --show-error --max-time 60 "http://localhost:9900/api/diagnosis/interview-aiops-payment-cpu-001/trace"
```
## 5. RAG 验收
```powershell
Invoke-RestMethod `
-Uri "http://127.0.0.1:9900/api/search/similar?query=ERR_TIMEOUT&topK=3" `
-Method Get
```
验收:
- 返回 `code = 200`。
- top candidates 中包含 `ERR_TIMEOUT` 相关文档。
- `scoreLabel` 能体现当前检索路径语义。
- 如果走 VectorStore,日志应出现 Spring AI VectorStore search。
## 6. 常见问题
### MySQL stale connection
现象:
```text
HikariPool - Connection is not available
No operations allowed after connection closed
```
处理:
- 重启服务。
- 确认 HikariPool 使用当前配置启动成功。
- 再跑 trace 或 AIOps 请求。
### SSE 客户端显示异常
PowerShell `Invoke-WebRequest` 有时对 SSE 或长 JSON 处理不稳定。可以改用 `curl.exe` 或直接查询 MySQL 和 Trace API 验证结果。
### OpenSpec 全量校验失败
`openspec validate --all --strict` 可能因为历史未完成 change 失败。面试演示主要依赖已归档 spec、MVP trace 和 RAG 验收材料,可以单独验证相关 spec。
+74
View File
@@ -0,0 +1,74 @@
# AIOps 轻量规则验证器
## 1. 改动是什么
AIOps 现在有一个确定性的后置质量门禁。最终告警报告持久化后,`AiOpsRuleEvaluationService` 会检查:
- 最终报告是否存在,且不是明显过短。
- payload 模式下,报告是否提到输入的告警和服务。
- 是否有证据工具调用,例如 `lookup_knowledge`、`query_metrics`、`query_logs`。
结果写入:
```text
diagnosis_session.self_evaluation.aiops_rule_evaluation
```
Trace API 会通过 session self-evaluation 展示这个结果。
## 2. 为什么先做规则型
这还不是完整 LLM Verifier。
AIOps 第一阶段质量风险比较具体,适合先用规则:
- 报告有没有生成。
- 报告有没有聚焦 payload。
- 有没有使用证据工具。
- 有没有把无关告警展开成主诊断对象。
规则验证稳定、便宜、容易解释,也不会在当前链路里额外引入一次隐藏模型调用。
## 3. 判定结果
当前评估器输出:
```text
PASS
WARN
FAIL
```
含义:
- `PASS`:核心检查通过。
- `WARN`:报告存在,但可能缺少 payload 关键词或证据工具。
- `FAIL`:缺少最终报告、报告过短等关键问题。
缺少 payload 关键词或证据工具先给 `WARN`,因为 demo/mock 环境下证据可能不可用,且报告措辞可能与 payload 字段不完全一致。
## 4. 面试回答
如果被问:为什么 AIOps 也需要验证器?
```text
Chat 已经有 LLM Verifier,因为用户问题开放度高。
AIOps 的第一阶段质量风险更明确:报告是否聚焦输入告警、是否使用证据工具、报告是否完整。
所以我先做了轻量规则验证器,把结果写入 self_evaluation,让 Trace 不只展示 Agent 做了什么,也展示输出是否通过基础质量门。
```
如果被问:为什么不直接复用 Chat Verifier?
```text
AIOps 验证语义和 Chat 不一样。
它要检查 alert scope、payload focus、证据工具覆盖,以及是否过度展开无关 active alerts。
直接复用 Chat Verifier 会混淆这些语义。
规则评估先提供稳定质量门,后续 AIOps LLM Verifier 可以基于同一套 trace contract 扩展。
```
## 5. 后续增强
- 引入 AIOps LLM Verifier,逐条校验根因和建议是否有 evidence refs。
- 把 rule evaluation 的 checks 在 Trace API 中结构化展示。
- 将 payload scope violation 沉淀为 bad case。
+75
View File
@@ -0,0 +1,75 @@
# AIOps 查询增强说明
## 1. 改动是什么
AIOps 在 `PAYLOAD_TARGETED` 模式下,会从告警 payload 中稳定生成一条推荐知识库检索 query。
参与拼接的非空字段:
```text
alertName service severity description timeRange userRequest
```
示例:
```text
HighCPUUsage payment-service P1 CPU 使用率超过 80% last_15m
```
最终会进入 Prompt:
```text
Recommended lookup_knowledge query: ...
```
## 2. 为什么重要
AIOps payload 里包含高价值检索词:
- 告警名称。
- 服务名。
- 严重等级。
- 症状描述。
- 时间范围。
- 用户补充请求。
如果完全让 Agent 从长 Prompt 里自己组织检索 query,可能遗漏服务名或告警名。推荐 query 让检索种子更稳定。
## 3. 设计取舍
这是 Prompt 层 query augmentation,不是隐藏检索。
我没有在 Agent 运行前自动调用 `lookup_knowledge`,原因是项目强调可追踪性:工具调用应该由 Agent 显式发起,并记录到 `tool_invocation`。
当前设计:
```text
AIOps payload
-> deterministic recommended retrieval query
-> Agent prompt
-> Agent 显式调用 lookup_knowledge
-> tool_invocation 记录真实检索行为
```
## 4. 面试回答
如果被问:AIOps payload 怎么提升 RAG 检索?
```text
我没有把告警 payload 粗暴替换成一个宽泛领域,而是提取 alertName、service、severity、description、timeRange 等高信号字段,拼成推荐的 lookup_knowledge query。
Agent 仍然显式调用工具,所以 trace 仍然能看到真实检索行为,但 query 不再完全依赖模型临场发挥。
```
如果被问:为什么不自动检索?
```text
自动检索会在 Agent 真正决策前制造一份隐藏证据。
这个项目的重点是可观测 Agent 执行,所以我选择 Prompt 层增强:给 Agent 一个更好的 query seed,但不改变工具调用必须显式可追踪的契约。
```
## 5. 后续增强
- 将 recommended query 写入 trace 的结构化字段,便于对比 Agent 实际 query。
- 对 payload 字段加权,例如 alertName/service 权重大于 timeRange。
- 后续接入 Query Transformer 时,保留原始 query、推荐 query、改写 query 三者的可追踪关系。
+97
View File
@@ -0,0 +1,97 @@
# 面试版系统架构
## 1. 系统分层
```mermaid
flowchart TB
API["API 层\nChatController / DiagnosisTraceController / SearchController"] --> Service["应用服务层\nChatService / AiOpsService / DiagnosisTraceService"]
Service --> Agent["Agent 编排层\nPlanner / Executor / Verifier / Supervisor"]
Agent --> Tools["工具层\nlookup_knowledge / query_logs / query_metrics / Prometheus"]
Tools --> RAG["RAG 检索\nL0 hint + VectorSearchService"]
RAG --> VectorStore["Spring AI VectorStore"]
RAG --> SDK["Milvus SDK fallback"]
Agent --> Trace["Trace 持久化"]
Tools --> Trace
Trace --> Session["diagnosis_session"]
Trace --> Step["agent_step"]
Trace --> Invocation["tool_invocation"]
Session --> TraceAPI["GET /api/diagnosis/{sessionId}/trace"]
Step --> TraceAPI
Invocation --> TraceAPI
```
## 2. Chat 链路
```mermaid
flowchart TD
User["用户问题"] --> ChatAPI["POST /api/chat"]
ChatAPI --> Strategy["ChatService.executeChatWithStrategy"]
Strategy --> Complexity{"复杂问题?"}
Complexity -->|否| Single["单 ReactAgent 快速回答"]
Complexity -->|是| Planner["chat_planner"]
Planner --> Executor["chat_executor"]
Executor --> Tools["证据工具"]
Tools --> Executor
Executor --> Verifier["chat_verifier"]
Verifier --> Decision{"PASS / LOW_CONFID / REJECT"}
Decision --> Answer["最终答复"]
Planner --> Step["agent_step"]
Executor --> Step
Verifier --> Step
Tools --> Invocation["tool_invocation"]
Answer --> Session["diagnosis_session"]
```
讲解重点:
- Planner 拆解问题和排查方向。
- Executor 必须通过工具收集证据。
- Verifier 只基于 `tool_trace_summary` 校验答案,不做新检索。
- Trace API 能回放模型步骤和工具证据。
## 3. AIOps 链路
```mermaid
flowchart TD
Alert["告警 payload 或空请求"] --> API["POST /api/ai_ops"]
API --> AiOps["AiOpsService"]
AiOps --> Mode{"是否有 payload?"}
Mode -->|有| Targeted["PAYLOAD_TARGETED\n聚焦输入告警"]
Mode -->|无| Discovery["AUTO_DISCOVERY\n先发现活跃告警"]
Targeted --> Supervisor["ai_ops_supervisor"]
Discovery --> Supervisor
Supervisor --> Planner["planner_agent"]
Supervisor --> Executor["executor_agent"]
Planner --> Tools["Prometheus / 日志 / 知识库"]
Executor --> Tools
Tools --> Report["告警分析报告"]
Report --> Eval["AiOpsRuleEvaluationService"]
Eval --> SelfEval["self_evaluation.aiops_rule_evaluation"]
```
讲解重点:
- AIOps 有明确产品边界:有 payload 时必须聚焦该告警。
- payload 字段会生成 recommended `lookup_knowledge` query。
- 当前 AIOps 先用规则评估做质量门,后续再扩展 LLM Verifier。
## 4. Trace 数据模型
| 表 | 作用 |
|---|---|
| `diagnosis_session` | 一次诊断的主记录:问题、状态、答案、自评估、反馈 |
| `agent_step` | Agent 模型调用记录:输入、输出、耗时、token、是否有工具调用 |
| `tool_invocation` | 工具调用事实:工具名、入参、输出预览、检索层、相关性、成功状态 |
## 5. 为什么 Trace 是核心
故障诊断系统的风险不只是“答案错”,还包括“答案看起来对但无法解释”。这个项目把执行链路拆成 session、step、tool 三层,让面试官可以看到:
- 模型为什么这么答。
- 调了哪些工具。
- 工具返回了什么证据。
- Verifier 如何判断答案可信度。
- 用户反馈如何回写到同一个 session。
这就是它区别于普通 Chatbot 的地方。
+145
View File
@@ -0,0 +1,145 @@
# 面试演示脚本
## 1. 30 秒开场
```text
这是一个 Agent 工程项目,场景是企业故障诊断。
它支持两类入口:用户主动提问的 Chat 诊断,以及告警事件驱动的 AIOps 诊断。
项目重点不是单次模型回答,而是把 Agent 编排、工具证据、Verifier 评估、最终报告和反馈都沉淀成可回放的 Trace。
```
## 2. 启动服务
```powershell
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
```
服务地址:
```text
http://localhost:9900
```
`mvp-demo` profile 下:
- Prometheus 告警使用 mock 数据。
- CLS 日志使用 mock 数据。
- MySQL、Redis、Milvus/Zilliz 和模型配置仍使用当前项目配置。
## 3. Demo 1:Chat 诊断
目标:展示用户问题如何进入多 Agent 诊断、调用工具、经过 Verifier,并生成 Trace。
```powershell
$sessionId = "interview-chat-payment-timeout-001"
$body = @{
Id = $sessionId
Question = "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
} | ConvertTo-Json
Invoke-RestMethod `
-Method Post `
-Uri "http://localhost:9900/api/chat" `
-ContentType "application/json" `
-Body $body
```
讲解点:
- `ChatService` 会根据问题复杂度选择轻量回答或复杂 Agent 流程。
- 复杂问题走 `Planner -> Executor -> Verifier`。
- Executor 调用知识库、日志、指标等证据工具。
- Verifier 基于 `tool_trace_summary` 生成 groundedness 评估。
- 最终写入 `diagnosis_session`、`agent_step`、`tool_invocation`。
查询 Trace:
```powershell
Invoke-RestMethod `
-Method Get `
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace"
```
展示点:
- `data.session.agentFlow = CHAT`
- `data.steps` 中能看到 planner/executor/verifier
- `data.toolInvocations` 中能看到证据工具
- `data.session.selfEvaluation` 中有 verifier 结果
## 4. Demo 2:AIOps 告警诊断
目标:展示告警 payload 如何触发 AIOps,并且报告聚焦目标告警。
```powershell
$aiopsSessionId = "interview-aiops-payment-cpu-001"
$aiopsBody = @{
sessionId = $aiopsSessionId
alertName = "HighCPUUsage"
service = "payment-service"
severity = "P1"
description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
timeRange = "last_15m"
userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
} | ConvertTo-Json
Invoke-WebRequest `
-Method Post `
-Uri "http://localhost:9900/api/ai_ops" `
-ContentType "application/json" `
-Body $aiopsBody
```
讲解点:
- `/api/ai_ops` 接受可选 `AIOpsRequest`。
- 首条 SSE 消息会返回 `type=session`。
- `AiOpsService` 根据 payload 判断模式:
- `PAYLOAD_TARGETED`:聚焦传入告警。
- `AUTO_DISCOVERY`:没有 payload 时先查 active alerts。
- AIOps 当前用 rule evaluation 检查报告完整性、payload 聚焦和证据工具覆盖。
查询 Trace:
```powershell
Invoke-RestMethod `
-Method Get `
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace"
```
展示点:
- `data.session.agentFlow = AI_OPS`
- `data.session.answer` 有最终告警报告
- `data.toolInvocations` 有 `query_metrics`、`query_logs`、`lookup_knowledge`
- 报告主线聚焦 `HighCPUUsage/payment-service`
## 5. Demo 3:反馈闭环
```powershell
$feedback = @{
sessionId = $sessionId
feedback = "useful"
} | ConvertTo-Json
Invoke-RestMethod `
-Method Post `
-Uri "http://localhost:9900/api/feedback" `
-ContentType "application/json" `
-Body $feedback
```
讲解点:
- feedback 写回同一个 `diagnosis_session`。
- `useful` 会沉淀 `case_library`。
- `not_useful` 不改变 `status`,只作为质量信号。
## 6. 收尾总结
```text
这个 Demo 展示的是完整 Agent 闭环:
用户问题或告警 -> Agent 编排 -> 工具证据 -> 自评估 -> Trace 回放 -> 用户反馈 -> 案例沉淀。
我关注的不是一次回答,而是这个回答能否被审计、验证和持续改进。
```
+107
View File
@@ -0,0 +1,107 @@
# 关键设计取舍
## 1. 为什么先做 Trace,而不是只返回答案
普通 Chatbot 只关注最终回答,但故障诊断更需要可审计性。一次诊断至少要回答:
- 结论是什么。
- 证据来自哪里。
- 哪些步骤由哪个 Agent 完成。
- 如果答案不可靠,系统怎么降级。
因此项目把一次会话拆成:
- `diagnosis_session`:会话摘要、最终答案、自评估、用户反馈。
- `agent_step`:模型输入输出、耗时、token 和工具调用标记。
- `tool_invocation`:真实工具调用参数、输出预览、成功状态和检索元数据。
代价是实现复杂度上升,收益是可回放、可调试、可演示。
## 2. 为什么 RAG 不直接隐藏在 Advisor 里
Spring AI Advisor 可以让 RAG 更隐式,但本项目的核心是 Agent 证据链。`lookup_knowledge` 必须作为显式工具调用出现,这样 Trace 里才能看到:
- Agent 什么时候决定检索。
- 用了什么 query。
- 命中了哪些文档。
- 相关性等级是什么。
- 证据如何支撑最终答案。
所以当前设计是:
```text
Executor -> lookup_knowledge -> VectorSearchService -> VectorStore / SDK fallback
```
这牺牲了一点框架自动化,但保留了可审计性。
## 3. 为什么 L0 只做 hint,不直接返回
旧版 L0 关键词唯一命中时可能直接跳过 L1。这个策略速度快,但风险是:关键词子串命中不等于最终语义相关。
当前改成:
```text
L0 = domain/entity hint
L1 = semantic retrieval
postprocess = evidence shaping + trace
```
L0 仍然有价值:错误码、服务名、告警名、指标名都很适合做精确 hint。但最终证据仍需要 L1 和后处理支撑。
## 4. 为什么保留 Milvus SDK fallback
Spring AI VectorStore 是当前读路径主方向,但 SDK fallback 没有删除,原因有三点:
- 迁移安全:旧 SDK 路径已经被验证过。
- 运行韧性:VectorStore 配置、schema、collection 出问题时可以回退。
- 面试稳定:检索抽象迁移不应该破坏主 demo。
这不是“没有迁完”,而是分阶段迁移:先稳定读路径,再决定是否迁移写入和索引。
## 5. 为什么 Chat 有 Verifier,AIOps 先用规则评估
Chat 问题更开放,容易出现跨领域推理,所以需要 LLM Verifier 做 groundedness 校验。
AIOps 当前优先解决更具体的问题:
- 最终报告是否存在。
- payload 模式是否聚焦输入告警。
- 是否使用了证据工具。
- 是否把无关活跃告警展开成主根因。
这些用规则就能稳定检查。后续可以在同一个 `self_evaluation` 容器下增加 AIOps LLM Verifier。
## 6. 为什么 AIOps payload scope 先用 Prompt + Rule
真实告警环境里可能同时有多个 active alerts。用户传入 `HighCPUUsage/payment-service` 时,Agent 如果把所有告警都展开分析,报告会跑偏。
当前选择:
- Prompt 中加入 `PAYLOAD_TARGETED`。
- 从 payload 生成 recommended `lookup_knowledge` query。
- 用 `AiOpsRuleEvaluationService` 检查报告是否聚焦输入告警。
没有先做硬过滤,是因为有些相关告警可以作为风险背景。目标不是屏蔽上下文,而是控制主诊断对象。
## 7. 为什么反馈不改 status
`status` 表示执行状态,`feedback` 表示用户评价。一个执行成功但用户觉得没用的诊断,应该是:
```text
status = SUCCESS
feedback = not_useful
```
这样才能区分系统异常和质量问题。`useful` 反馈会沉淀 `case_library`,`not_useful` 作为 bad case 信号保留。
## 8. 可以主动承认的限制
- AIOps 还没有完整 LLM Verifier。
- RAG 还没有 hybrid search、rerank、邻居 chunk 扩展。
- `case_library` 的 rootCause/solution 仍需要结构化抽取。
- `tool_invocation.step_id` 关联还可以更严格。
- `mvp-demo` profile 使用 mock 日志和指标,主要服务稳定面试演示。
主动讲清这些限制,能体现工程判断:先把可追踪闭环打通,再逐步增强质量门和生产可靠性。
@@ -0,0 +1,91 @@
# RAG Breadcrumb Embedding 验收说明
## 1. 改动是什么
索引路径现在构造 embedding 文本时,不只使用 chunk 内容,还会把结构上下文拼进去:
```text
Title: {title}
Path: {breadcrumb}
Content:
{content}
```
Milvus 中存储的 `content` 字段仍然保留原始 chunk 内容。这样展示和证据输出保持干净,而向量本身携带章节语义。
## 2. 为什么必须重新索引
Embedding 是索引时物化的。已有向量是用旧的 content-only 文本生成的,所以只有代码变化并不会改变线上检索结果。
验收关键点:
```text
只改代码 != live retrieval 已变化
代码改动 + 重新索引 + live query report = 行为验收完成
```
## 3. 如何验证
1. 启动 Spring Boot 应用。
2. 通过现有索引路径重新索引知识库。
3. 运行:
```bash
python scripts/eval_rag_live_acceptance.py
```
脚本输出:
```text
eval/rag-retrieval/reports/live-post-reindex.json
eval/rag-retrieval/reports/live-post-reindex.md
```
默认覆盖:
- breadcrumb 敏感的 RAG chunk context query。
- 需要章节路径的诊断流程问题。
- `ERR_TIMEOUT` 精确错误码检索。
- MySQL 连接池排障。
- AIOps payment-service 延迟告警检索。
## 4. 看什么结果
对 breadcrumb 敏感 case:
- top candidates 是否暴露预期 `title`。
- top candidates 是否暴露预期 `breadcrumb`。
- 命中内容是否能看出所属章节。
对核心排障 case:
- 结果数量是否稳定。
- top candidates 是否仍然命中核心文档。
- 没有因为拼接 title/breadcrumb 导致核心检索退化。
## 5. 面试回答
如果被问:你怎么验证 breadcrumb 参与 embedding 后真的生效?
```text
我把 deterministic regression 和 live acceptance 分开。
离线 fixture baseline 不依赖服务,可以做稳定回归。
但 embedding 改动只会影响新生成的向量,所以我另外加了 live post-reindex acceptance 脚本。
脚本会调用真实 /api/search/similar,对 breadcrumb 敏感、排障和 AIOps query 生成 JSON/Markdown 报告。
这样能证明代码改了,也能证明 live vector collection 已经刷新。
```
如果被问:为什么脚本不自动 reindex?
```text
reindex 会修改向量库,而且依赖环境中的知识库数据。
我把 reindex 保持为显式动作,验收脚本只做读取验证。
这样如果检索没有改善,我能区分是代码问题、索引未刷新,还是运行时检索行为问题。
```
## 6. 后续增强
- 将 live acceptance 结果加入面试 Demo 输出。
- 增加 breadcrumb hit rate 统计。
- 对同章节 chunk 做邻居扩展,进一步利用 breadcrumb。
+127
View File
@@ -0,0 +1,127 @@
# RAG 重构故事
## 1. 起点
原始 RAG 实现已经能支撑 MVP:
- 文档可以上传、切片、向量化,并写入 Milvus/Zilliz。
- Agent 可以显式调用 `lookup_knowledge`。
- AIOps 诊断能在告警流程里检索排障知识。
- 工具调用会落到 `tool_invocation`,检索步骤可见。
但它有几个工程问题:
- 检索实现过于依赖 Milvus SDK,业务代码承担了太多底层搜索细节。
- L0 和 L1 职责不清,L0 关键词命中容易被当作最终召回决策。
- chunk 级检索容易丢失章节上下文。
- `breadcrumb` 存在 metadata 中,但没有充分参与 embedding、filter 和上下文重建。
- 检索质量主要靠手工接口和日志判断,缺少可重复的 golden cases。
所以重构目标不是“全盘替换成框架”,而是:
```text
通用 RAG 基础设施交给 Spring AI,
业务可观测链路保留在项目里。
```
## 2. 我如何拆解问题
我把迁移拆成几个阶段,因为 RAG 同时影响 Agent 工具层、AIOps、向量检索、证据打包和 Trace。
第一步是建立 baseline。`eval/rag-retrieval/` 中的 golden cases 用来对比后续改动,而不是只靠直觉判断检索有没有变好。
第二步是明确职责:
```text
L0 = domain/entity hint
L1 = semantic retrieval
postprocess = evidence shaping + trace-friendly output
```
L0 仍然有价值,但不再默认绕过语义检索。它更适合提取服务名、告警名、错误码、领域和 metadata filter。
第三步是增强 evidence 输出。Agent 不应该只拿到 raw chunk,而应该拿到带 source、title、breadcrumb、score、hit reason 的证据块。
最后,我把 Spring AI `VectorStore` 接入为读取主路径,同时保留原 Milvus SDK 作为 fallback。
## 3. 当前架构
```text
Agent / API
-> lookup_knowledge or /api/search/similar
-> L0 domain/entity hint
-> VectorSearchService
-> Spring AI VectorStore
-> Milvus SDK fallback
-> relevance normalization
-> tool_invocation trace
```
`VectorSearchService` 仍然是公共检索门面。Agent 工具层不需要知道底层是 SDK 还是 Spring AI。
支持三种模式:
```text
auto -> 优先 Spring AI VectorStore,失败后 fallback 到 SDK
spring-ai -> 强制 Spring AI VectorStore
sdk -> 强制 Milvus SDK
```
## 4. 关键取舍
### 保留显式工具
我没有把检索藏进 Spring AI Advisor。原因是这个项目强调 Agent 执行可见性:`lookup_knowledge` 的 query、命中文档、相关性和证据预览都要进入 Trace。
### 保留 SDK fallback
SDK fallback 不是废代码,而是迁移安全网。实际验证时,第一次 VectorStore 指向了错误 collection,`auto` 模式 fallback 到 SDK 后仍能返回结果。修正 collection 后,Spring AI 路径成为主路径。
### L0 降权
生产事故中经常有精确标识:错误码、告警名、服务名、指标名。L0 适合做 hint,但不应该做最终裁判。
### 分数语义拆开
SDK 使用 L2 distance,Spring AI 暴露 similarity。混在一个字段里会让 relevance normalization 出错。
当前拆成:
```text
score -> 兼容旧逻辑的距离型分数
rawScore -> 底层原始分数
scoreLabel -> rawScore 的语义
```
### 暂不迁移写入
写入和索引仍走 SDK。这是有意分阶段:先验证读路径,再评估 `VectorStore.add(...)` 是否适合现有 metadata 和 chunk 模型。
## 5. 验证方式
我用了三层验证:
- 单元测试:SDK mode、Spring AI mode、auto fallback、category filter、distance metadata mapping。
- Live API:`GET /api/search/similar?query=ERR_TIMEOUT&topK=3`。
- 代表性 query 对比:错误码、支付超时、MySQL 连接池、AIOps 告警式 query、抽象 RAG 设计问题。
核心排障和 AIOps query 在 SDK 与 VectorStore 下 top3 一致。差异主要集中在抽象设计类问题和 metadata taxonomy,这些被记录为后续质量工作。
## 6. 面试短版
```text
这个 RAG 系统最初是基于 Milvus SDK 的自研 MVP。它能跑,但底层检索细节过多地散落在业务代码里,L0/L1 职责也不够清晰。
我按阶段重构:先加 retrieval baseline,再把 L0 降级为 domain/entity hint,再增强 evidence postprocess,最后把读取主路径切到 Spring AI VectorStore,并保留 SDK fallback。
我没有把 lookup_knowledge 替换成隐式 Advisor,因为这个项目的核心是可追踪 Agent:面试官可以看到什么时候检索、检索了什么、证据如何支撑诊断。
```
## 7. 可主动承认的不足
- metadata taxonomy 还需要清理,例如 `database` 与 `infrastructure`。
- 抽象设计问题可能需要 query rewrite 或更好的文档索引。
- 邻居 chunk / 同章节上下文扩展还不完整。
- rerank、RRF、BM25、hybrid retrieval 还没有接入。
- 写入路径仍使用 SDK。
这些不是当前迁移阻塞项,而是后续检索质量优化方向。
+148
View File
@@ -0,0 +1,148 @@
# RAG 检索质量报告
## 1. 目的
这份报告回答一个面试关键问题:
```text
迁移到 Spring AI VectorStore 后,怎么证明检索质量没有退化?
```
这不是完整 benchmark,而是针对当前 Milvus/Zilliz collection 的代表性 live smoke comparison。
## 2. 验证设置
服务端点:
```text
GET http://127.0.0.1:9900/api/search/similar
```
collection:
```text
biz
```
对比模式:
```text
retrieval.vector-store.mode=sdk
retrieval.vector-store.mode=spring-ai
```
每个 case:
```text
topK=3
```
## 3. 测试案例
| Case | Query | 目的 |
|---|---|---|
| `err-timeout` | `ERR_TIMEOUT` | 精确错误码检索 |
| `payment-service-timeout` | `payment-service timeout` | 服务超时排障 |
| `mysql-connection-pool` | `MySQL connection pool is exhausted. How should I diagnose it?` | 数据库排障 |
| `high-cpu-payment` | `HighCPUUsage payment-service` | AIOps 告警式检索 |
| `rag-l0-l1` | `Should L0 keyword matching decide the final retrieval result?` | 抽象 RAG 设计问题 |
| `database-filter` | `mysql timeout`, category=`database` | metadata filter 行为 |
## 4. 对比摘要
| Case | SDK 数量 | VectorStore 数量 | Top1 一致 | TopK 重叠 | 结论 |
|---|---:|---:|---|---:|---|
| `err-timeout` | 3 | 3 | 是 | 3/3 | 文档和顺序一致 |
| `payment-service-timeout` | 3 | 3 | 是 | 3/3 | 文档和顺序一致 |
| `mysql-connection-pool` | 3 | 3 | 是 | 3/3 | 文档和顺序一致 |
| `high-cpu-payment` | 3 | 3 | 是 | 3/3 | AIOps 核心 query 一致 |
| `rag-l0-l1` | 3 | 1 | 是 | 1/3 | VectorStore 尾部结果更少 |
| `database-filter` | 0 | 0 | 不适用 | 不适用 | filter 行为一致,taxonomy 有问题 |
## 5. 代表性结果
### ERR_TIMEOUT
SDK:
```text
1. ERR_TIMEOUT score=0.5659486 label=l2_distance
2. ERR_GATEWAY_TIMEOUT score=0.6048740 label=l2_distance
3. Error handling score=0.7735061 label=l2_distance
```
VectorStore:
```text
1. ERR_TIMEOUT score=0.5659486 rawScore=0.4340513 label=similarity
2. ERR_GATEWAY_TIMEOUT score=0.6048740 rawScore=0.3951259 label=similarity
3. Error handling score=0.7735061 rawScore=0.2264938 label=similarity
```
解释:
- 排序一致。
- 兼容 `score` 与 SDK L2 distance 一致。
- `rawScore` 暴露 Spring AI similarity。
### MySQL connection pool
两条路径都返回:
```text
1. MySQL connection pool config
2. wait_timeout timeout
3. idle-timeout
```
说明迁移保留了核心基础设施排障检索能力。
### HighCPUUsage payment-service
两条路径都返回 payment-service 高 CPU 相关排障文档,说明 AIOps 告警式 query 没有退化。
### rag-l0-l1
VectorStore 只返回一个候选,但 Top1 与 SDK 一致。这说明抽象设计类 query 需要后续 query rewrite、补充索引或 threshold 调整。
### database-filter
两条路径都返回 0,因为相关 MySQL 文档当前分类是 `infrastructure`,不是 `database`。这是 metadata taxonomy 问题,不是 VectorStore 回归。
## 6. 分数兼容结论
对比验证了当前分数设计:
```text
SDK:
score = L2 distance
rawScore = L2 distance
scoreLabel = l2_distance
VectorStore:
score = Milvus metadata.distance
rawScore = Spring AI similarity
scoreLabel = similarity
```
这样既保持 `lookup_knowledge` 原有归一化逻辑,又能暴露 VectorStore 语义。
## 7. 验收结论
Spring AI VectorStore 读路径可以接受用于当前 MVP/面试:
- 核心排障和 AIOps case 与 SDK top3 一致。
- 分数兼容性保留。
- VectorStore 语义通过 `rawScore` 和 `scoreLabel` 可观察。
- SDK fallback 仍保留运行安全。
后续检索质量工作不阻塞这次迁移,应作为独立优化继续推进。
## 8. 下一步
- 增加自动 live comparison 脚本。
- 在 offline evaluator 中加入 topK overlap、top1 hit、MRR。
- 规范 metadata category,例如 `database` 与 `infrastructure`。
- 为抽象设计类 query 增加 query rewriting。
- 后续再评估是否迁移写入路径到 `VectorStore.add(...)`。
@@ -0,0 +1,119 @@
# RAG VectorStore 面试要点
## 1. 60 秒讲法
```text
我把 RAG 检索从 Milvus SDK-only 重构为 Spring AI VectorStore 主路径,同时保留 SDK fallback。
关键不是换了一个依赖,而是保留 VectorSearchService 作为边界,所以 lookup_knowledge 和 Agent workflow 不需要改。
现在支持 auto、spring-ai、sdk 三种模式。auto 会优先尝试 VectorStore,失败后 fallback 到 SDK。
```
现场验证时,第一次发现 VectorStore 指向了错误 collection:`business_knowledge`,而实际 Zilliz collection 是 `biz`。fallback 生效,所以系统仍能通过 SDK 返回结果。修正 collection 后,同一个 query 成功走 Spring AI VectorStore。
## 2. 架构回答
```text
Agent / API
-> lookup_knowledge or /api/search/similar
-> VectorSearchService
-> Spring AI VectorStore
-> Milvus SDK fallback
-> Milvus/Zilliz collection: biz
```
关键设计:`VectorSearchService` 是检索门面,避免 Spring AI 或 SDK 细节扩散到 Agent 工具层。
## 3. 为什么保留 SDK
- 迁移安全:原 SDK 路径已验证可用。
- 运行韧性:VectorStore schema、filter 或配置失败时,检索仍可用。
- Demo 稳定:检索抽象变化不应该破坏主诊断演示。
这在实际验证中发挥了作用:VectorStore 配置错时,`auto` 模式 fallback 到 SDK,API 没有失败。
## 4. 为什么引入 Spring AI VectorStore
使用 `VectorStore` 可以让项目更接近标准 RAG 抽象:
- 业务代码不再持有全部 Milvus search 细节。
- 后续 QueryTransformer、DocumentPostProcessor、Retriever 等能力更容易接入。
- 面试中也更容易解释和 Spring AI 生态的关系。
但我没有一次性迁移写入,因为读写同时迁移会让问题难定位。当前先稳定读路径。
## 5. 为什么保留 L0
L0 现在不是最终答案来源,而是确定性 hint 层:
- 提取 domain/entity。
- 在可能时生成 category filter。
- 给 trace 提供解释信号。
当前职责:
```text
L0 = domain/entity hint
L1 = semantic retrieval through VectorStore/SDK
postprocess = evidence trace + relevance normalization
```
真实故障诊断里有很多精确标识,完全只靠向量检索并不稳。
## 6. 为什么不用隐藏 Advisor
`lookup_knowledge` 保持显式工具,因为:
- Trace 要展示什么时候检索。
- `tool_invocation` 要记录输入、输出预览、相关性和 metadata。
- 面试故事是可审计 Agent 执行,而不只是答案质量。
Advisor 后续可以接入,但需要先解决可观测性。
## 7. 分数设计
当前结果故意拆成:
```text
score -> 兼容旧 relevance normalization 的分数
rawScore -> 当前检索实现原始分数
scoreLabel -> rawScore 的语义
```
SDK:
```text
score = L2 distance
rawScore = L2 distance
scoreLabel = l2_distance
```
VectorStore:
```text
score = metadata.distance if present
rawScore = Spring AI document score
scoreLabel = similarity
```
这样避免把 similarity 当成 L2 distance 的隐蔽 bug。
## 8. 如何证明 VectorStore 被使用
- 日志出现 `Starting Spring AI VectorStore search` 和 `Spring AI VectorStore search complete`。
- API 响应中 `scoreLabel=similarity`。
- `rawScore` 是 Spring AI similarity,`score` 仍是兼容 distance。
## 9. 常见追问
### 为什么不删 SDK?
这是迁移,不是重写。fallback 提供回滚安全,并且已经证明配置错误时仍能保证主链路可用。
### `lookup_knowledge` 变了吗?
外部契约没变。它仍然调用 `VectorSearchService.searchSimilarDocuments(...)`,变化在门面背后的实现。
### 这是完整 Spring AI RAG 了吗?
还不是。当前是 Spring AI VectorStore 读路径 + 显式工具 + 自定义 evidence trace + SDK 写入。这样做是为了保留审计能力和分阶段迁移安全。
@@ -0,0 +1,198 @@
# RAG VectorStore Live 验收说明
## 1. 目的
本文记录 RAG 检索重构的 live 验收结论。
这次重构的目标不只是接入 Spring AI 抽象,而是证明线上读路径能够:
- 优先使用 Spring AI `VectorStore` 做 Milvus 检索。
- 保留原 Milvus SDK 作为 fallback。
- 保持 `lookup_knowledge` 工具契约稳定。
- 保持基于 L2 distance 的相关性归一化兼容。
## 2. 当前检索形态
```text
lookup_knowledge / /api/search/similar
-> VectorSearchService.searchSimilarDocuments(...)
-> retrieval.vector-store.mode
-> auto
-> Spring AI VectorStore
-> VectorStore 失败时 fallback 到 Milvus SDK
-> spring-ai
-> 只走 Spring AI VectorStore
-> sdk
-> 只走 Milvus SDK
```
## 3. 已验证配置
live Milvus/Zilliz 数据库中存在 collection:
```text
biz
```
Spring AI VectorStore 配置与 SDK 使用的 collection 对齐:
```yaml
spring:
ai:
vectorstore:
type: milvus
milvus:
initialize-schema: false
database-name: ${milvus.database}
collection-name: biz
embedding-dimension: ${milvus.vector-dim}
metric-type: L2
id-field-name: id
content-field-name: content
metadata-field-name: metadata
embedding-field-name: vector
```
为什么重要:早期配置使用 `business_knowledge`,而真实 collection 是 `biz`。这个错配证明了 fallback 生效,但也说明修正前 VectorStore 不是成功主路径。
## 4. 验收命令
健康检查:
```powershell
Invoke-RestMethod `
-Uri "http://127.0.0.1:9900/milvus/health" `
-Method Get
```
期望:
```json
{
"collections": ["biz"],
"message": "ok"
}
```
直接检索:
```powershell
Invoke-RestMethod `
-Uri "http://127.0.0.1:9900/api/search/similar?query=ERR_TIMEOUT&topK=3" `
-Method Get
```
期望结果形态:
```json
{
"code": 200,
"message": "success",
"data": [
{
"content": "### ERR_TIMEOUT ...",
"score": 0.5662,
"rawScore": 0.4337,
"scoreLabel": "similarity",
"metadata": {
"distance": 0.5662,
"title": "ERR_TIMEOUT",
"category": "api"
}
}
]
}
```
## 5. 日志证明了什么
collection 修正前:
```text
Starting Spring AI VectorStore search
Spring AI VectorStore retrieval failed, falling back to Milvus SDK
Starting Milvus SDK search
```
collection 修正后:
```text
Starting Spring AI VectorStore search: query=ERR_TIMEOUT
Spring AI VectorStore search complete, candidates=3
```
这证明:
- `auto` 模式确实先尝试 VectorStore。
- VectorStore 失败时 fallback 可用。
- 配置对齐后,主路径是 Spring AI VectorStore,而不是 SDK fallback。
## 6. 分数语义
项目保留三个分数字段:
```text
rawScore -> 当前检索实现的原始分数
scoreLabel -> rawScore 的语义
score -> lookup relevance normalization 使用的兼容分数
```
SDK:
```text
rawScore = L2 distance
scoreLabel = l2_distance
score = L2 distance
```
Spring AI VectorStore:
```text
rawScore = Spring AI similarity score
scoreLabel = similarity
score = Milvus distance metadata when available
```
使用 `metadata.distance` 的原因:`LookupKnowledgeTool` 已经基于 L2 distance 做相关性归一化。Spring AI Milvus 主分数是 similarity,但 metadata 中仍有 Milvus distance。用 distance 保持旧逻辑稳定,同时通过 `rawScore` 暴露新语义。
## 7. 回归检查
目标测试:
```powershell
mvn -q "-Dtest=VectorSearchServiceTest,LookupKnowledgeToolTest" test
```
相关 spec:
```powershell
openspec.cmd validate rag-knowledge-retrieval --specs
openspec.cmd validate rag-retrieval-evaluation --specs
```
diff 检查:
```powershell
git diff --check
```
验收结论:
```text
目标测试通过。
相关 spec 通过。
diff-check 无错误。
```
## 8. 验收结论
VectorStore 读路径可以接受:
- Spring AI VectorStore 已集成,并在 `auto` 模式中优先使用。
- SDK fallback 保留且已被实际验证。
- live collection 配置与现有 Milvus collection 对齐。
- `lookup_knowledge` 对外契约保持稳定。
- 旧的 L2 relevance normalization 仍兼容。
写入和索引路径仍使用 Milvus SDK。这是有意的分阶段迁移,不是验收失败项。
+225
View File
@@ -0,0 +1,225 @@
# 面试故事案例
**用途**:把项目能力讲成可被面试官理解的工程故事
**使用方式**:按问题选择一个故事,不需要从头到尾背诵
## 故事 1:从黑盒 Chatbot 到可追踪 Agent
### 面试官问题
```text
这个项目和普通调用大模型有什么区别?
```
### 30 秒回答
```text
普通 Chatbot 只给最终答案,出了问题很难解释答案怎么来的。
我这个项目把诊断过程拆成 Planner、Executor、Verifier,并把每个 Agent 步骤和每次工具调用落库。
最后通过 Trace API 可以回放:模型怎么规划、调用了哪些工具、工具返回了什么证据、Verifier 怎么判断答案可信。
```
### 展开讲法
一开始最容易做的是:用户问题进来,直接让模型回答。但故障诊断场景不能只看答案,因为答案可能看起来合理却没有证据支撑。
所以我把系统拆成三层:
- `diagnosis_session` 记录一次诊断的主状态和最终答案。
- `agent_step` 记录 Planner、Executor、Verifier 的模型调用。
- `tool_invocation` 记录知识库、日志、指标等真实工具证据。
这样就能做到:答案不是孤立文本,而是一条可审计的执行链。
### 可展示文件
- `mvp/architecture/interview-one-pager.md`
- `mvp/architecture/session-trace-lifecycle.md`
- `mvp/demo/output/trace-response.json`
### 主动说不足
```text
当前 tool_invocation.step_id 还不是每次都强绑定具体 agent_step,后续可以加 runId 和更严格的 step 关联,让多轮同 session 诊断更清晰。
```
## 故事 2:RAG 从自研 SDK 检索迁移到 Spring AI VectorStore
### 面试官问题
```text
你的 RAG 是怎么设计的?为什么不用框架全包?
```
### 30 秒回答
```text
我把 RAG 分成两部分:通用检索基础设施尽量交给 Spring AI VectorStore,业务可观测链路留在项目里。
所以 Agent 仍然显式调用 lookup_knowledge,底层通过 VectorSearchService 走 Spring AI VectorStore,失败时 fallback 到原 Milvus SDK。
这样既能减少自研检索代码,又不会丢失工具调用 trace。
```
### 展开讲法
旧实现里,Milvus SDK 查询、topK、filter、score 映射都在业务代码里。它能跑,但后续扩展成本高。
我没有直接把 RAG 隐藏进 Advisor,因为这个项目的核心是 Agent 工程,需要知道 Agent 何时检索、检索了什么、证据怎么支撑诊断。
于是我保留了边界:
```text
Executor -> lookup_knowledge -> VectorSearchService -> VectorStore / SDK fallback
```
同时把 L0 从“直接返回结果”降级为 domain/entity hint,降低关键词误召回的风险。
### 可展示文件
- `mvp/architecture/rag-architecture.md`
- `mvp/architecture/retrieval-observability.md`
- `interview/rag-refactor-story.md`
### 主动说不足
```text
当前还没有完整 hybrid search 和 rerank。
我先做 golden cases、VectorStore 主路径和 SDK fallback,是为了让每一步迁移都能被验证。
```
## 故事 3:Verifier 如何降低幻觉风险
### 面试官问题
```text
Agent 怎么保证不胡说?
```
### 30 秒回答
```text
我没有假设模型天然可靠,而是加了 Verifier。
Executor 给出答案后,Verifier 只拿 executor_final_answer 和 tool_trace_summary,不允许做新检索。
它把答案里的关键事实逐条校验,输出 PASS、LOW_CONFID 或 REJECT。
这个结果会写回 self_evaluation,Trace API 可以看到。
```
### 展开讲法
Verifier 的关键不是再问一次模型“你觉得对吗”,而是让它基于真实工具调用做 groundedness 检查。
`ToolTraceSummaryService` 会从 `tool_invocation` 里整理证据索引,包含:
- 工具名。
- 输入摘要。
- 输出摘要。
- evidence level。
- source invocation ids。
Verifier 输出结构化 JSON,ChatService 根据 verdict 决定是否输出、补证据或降级。
### 可展示文件
- `mvp/architecture/harness-quality-gates.md`
- `mvp/architecture/feedback-architecture.md`
- `src/main/resources/prompts/chat-verifier-prompt.md`
### 主动说不足
```text
AIOps 当前还是轻量 rule evaluation,不是完整 LLM Verifier。
这是有意收敛:先用规则保证 payload 聚焦和工具证据使用,后续再加 AIOps LLM Verifier。
```
## 故事 4:AIOps 告警为什么要做 payload scope control
### 面试官问题
```text
AIOps 场景和普通 Chat 有什么区别?
```
### 30 秒回答
```text
AIOps 告警有一个很关键的问题:环境里可能同时有很多活跃告警,Agent 容易跑偏。
所以我把 AIOps 分成 PAYLOAD_TARGETED 和 AUTO_DISCOVERY。
如果请求带 alert payload,最终报告必须聚焦输入告警,并且会把 alertName、service、severity、description 拼成 recommended lookup_knowledge query。
```
### 展开讲法
没有 payload 时,Agent 可以先查询活跃告警,再选择目标排查。
但有 payload 时,用户已经告诉系统“我要查这个告警”。这时如果 Agent 把其他活跃告警写成主根因,产品体验会很差。
所以我做了两件事:
- Prompt 中明确 `PAYLOAD_TARGETED` 范围。
- `AiOpsRuleEvaluationService` 检查最终报告是否聚焦输入告警,以及是否使用证据工具。
### 可展示文件
- `mvp/architecture/agent-orchestration.md`
- `mvp/architecture/current-mvp-architecture.md`
- `interview/aiops-query-augmentation.md`
- `interview/aiops-lightweight-verifier.md`
### 主动说不足
```text
当前 scope control 主要靠 prompt 和规则评估。
后续可以把 AIOps 也接入类似 Chat Verifier 的事实校验,让告警报告的每个根因和建议都有 evidence refs。
```
## 故事 5:反馈不是点赞按钮,而是案例沉淀入口
### 面试官问题
```text
用户反馈在系统里有什么用?
```
### 30 秒回答
```text
反馈不只是前端按钮。
用户提交 useful 后,系统会把同一个 diagnosis_session 沉淀为 case_library。
not_useful 不会改执行状态,而是作为 bad case 信号保留。
这样 status、self_evaluation、feedback 三个维度是分开的。
```
### 展开讲法
我刻意没有把 `not_useful` 写成 `FAILED`。因为失败表示系统执行异常,而用户觉得不好用是质量标签。
当前设计里:
```text
status -> 执行是否成功
self_evaluation -> 系统自己判断证据和事实支撑度
feedback -> 用户是否认可
```
`useful` 会进入 `CaseLibraryService.createFromSession`,生成可复用案例。后续可以做相似案例推荐或高质量样本积累。
### 可展示文件
- `mvp/architecture/feedback-architecture.md`
- `mvp/architecture/data-model.md`
- `mvp/demo/output/feedback-response.json`
### 主动说不足
```text
当前 case_library 的 rootCause 和 solution 还直接使用完整 answer。
后续应该从报告中结构化抽取 rootCause、solution、errorCode 和 service,提高案例复用质量。
```
## 结尾万能总结
```text
这个项目我最想展示的不是某一个模型效果,而是 Agent 工程化能力:
一个诊断答案从哪里来、用了什么证据、是否被验证、用户是否认可、后续怎么沉淀和回归。
这些链路都被结构化记录下来,所以它可以继续演进,而不是一次性 demo。
```
@@ -0,0 +1,81 @@
---
title: AIOps 告警排障 Runbook
keywords: [AIOps, 告警, HighCPUUsage, SlowResponse, payment-service, system-metrics, application-logs]
summary: 面向 AIOps 告警诊断的排障步骤,覆盖 Prometheus 活动告警、CLS 日志主题和处理建议。
category: troubleshooting
---
# AIOps 告警排障 Runbook
## 1. 告警输入处理原则
AIOps 诊断入口有两种触发方式:
- **有告警 payload**:将 payload 视为已触发告警,围绕 `alertName`、`service`、`severity`、`timeRange` 查询指标、日志和知识库。
- **无告警 payload**:先调用 `queryPrometheusAlerts` 获取当前 firing 告警,再选择 P0/P1 或持续时间最长的告警进入诊断。
最终报告必须基于工具证据,不得凭空编造指标、日志或处理结果。
## 2. Mock 告警与日志主题映射
| 告警名 | 典型服务 | 优先日志主题 | 推荐查询 |
|---|---|---|---|
| HighCPUUsage | payment-service | system-metrics | `cpu_usage:>80 AND service:payment-service` |
| HighMemoryUsage | order-service | system-metrics, system-events | `memory_usage:>85` |
| SlowResponse | user-service | application-logs, database-slow-query | `duration:>3000 OR slow request` |
| ServiceUnavailable | 任意核心服务 | application-logs, system-events | `level:ERROR OR container crash` |
## 3. HighCPUUsage / payment-service 排障步骤
### 3.1 现象确认
先确认 Prometheus 活动告警中是否存在:
- `alert_name = HighCPUUsage`
- `service = payment-service`
- CPU 使用率超过 80%
- 状态为 firing
如果 payload 已经提供该告警,也仍需通过指标或日志工具验证。
### 3.2 指标与日志取证
推荐工具调用顺序:
1. `queryPrometheusAlerts`:确认当前活动告警。
2. `queryLogs(region=ap-guangzhou, logTopic=system-metrics, query=cpu_usage:>80 AND service:payment-service)`:确认 CPU 使用率、实例和持续时间。
3. 如报告中提到 Redis、数据库或下游依赖,再查询 `application-logs` 或对应主题交叉验证。
### 3.3 根因判断
可接受的根因结论必须至少满足一项:
- system-metrics 显示 payment-service 实例 CPU 使用率持续高于阈值。
- application-logs 显示与 CPU 飙高同时出现的慢请求、线程池耗尽或依赖超时。
- 告警持续时间与日志时间线一致。
如果只有活动告警,没有日志或指标明细,应输出低置信结论并建议人工确认。
## 4. 处理建议
### 临时止血
- 对 payment-service 做水平扩容,优先扩容受影响实例所在 Deployment。
- 对高耗时接口开启限流或降级非核心功能。
- 如果近期有发布,检查变更窗口并准备回滚。
### 根因修复
- 分析 CPU 热点线程、慢请求接口和依赖调用耗时。
- 检查连接池、线程池、缓存穿透和批量任务是否导致 CPU 飙高。
- 补充针对 `payment-service` 的 CPU、P95/P99 延迟、错误率和依赖超时联动告警。
## 5. 报告要求
告警分析报告至少包含:
- 活跃告警清单。
- 告警根因分析。
- 使用过的工具证据:Prometheus 告警、system-metrics 日志、application-logs 或知识库。
- 已执行或建议执行的处理方案。
- 置信度说明:哪些结论有直接证据,哪些需要人工进一步确认。
+96 -140
View File
@@ -1,160 +1,116 @@
# 数据库设计文档
# SuperBizAgent MVP 文档
## 📚 文档导航
**更新日期**:2026-07-05
### 核心表设计
- [diagnosis_record](tables/diagnosis_record.md) - 诊断记录表(核心)
- [case_library](tables/case_library.md) - 案例库表
- [api_document](tables/api_document.md) - 文档元数据表
本目录保存 MVP 阶段的架构、问题、演示、评测和数据表说明。当前架构入口已经整理到 `mvp/architecture/`,旧版架构材料已归档,避免继续把历史方案当成当前实现。
### 架构设计
- [Agent 架构设计](architecture/agent-architecture.md) - Agent 协作 + Skill + Harness
- [知识库检索架构](architecture/knowledge-retrieval-architecture.md) - L0+L1 混合检索架构 ⭐新增
- [知识库检索使用指南](architecture/knowledge-retrieval-usage.md) - 文档编写和使用说明 ⭐新增
- [会话管理](architecture/session-management.md) - Redis + MySQL 会话管理
- [实施规划](architecture/implementation-plan.md) - 分阶段实施计划
- [会话级去重与知识域地图](architecture/session-dedup-knowledge-map.md) - 文档级去重 + Planner 知识域地图注入解决 ISS-001 ⭐新增
- [证据评分与用户反馈](architecture/confidence-feedback.md) - evidence_score 规则引擎 + feedback API ⭐新增
- [行动记忆与检索归一化](architecture/action-memory-relevance.md) - Executor 行动记忆 + 归一化质量等级解决 ISS-002 ⭐新增
## 当前入口
---
| 目录/文档 | 用途 |
|---|---|
| [architecture/README.md](architecture/README.md) | 当前 MVP 架构入口 |
| [architecture/current-mvp-architecture.md](architecture/current-mvp-architecture.md) | 当前可运行系统架构 |
| [architecture/interview-one-pager.md](architecture/interview-one-pager.md) | 面试一页式架构讲解 |
| [architecture/agent-orchestration.md](architecture/agent-orchestration.md) | Agent 编排架构 |
| [architecture/harness-quality-gates.md](architecture/harness-quality-gates.md) | Harness 与质量门禁 |
| [architecture/rag-architecture.md](architecture/rag-architecture.md) | RAG/知识检索新架构 |
| [architecture/retrieval-observability.md](architecture/retrieval-observability.md) | 检索与可观测性架构 |
| [architecture/feedback-architecture.md](architecture/feedback-architecture.md) | 反馈与自评估架构 |
| [architecture/session-trace-lifecycle.md](architecture/session-trace-lifecycle.md) | 会话与 Trace 生命周期 |
| [architecture/knowledge-base-authoring.md](architecture/knowledge-base-authoring.md) | 知识库文档编写与维护 |
| [architecture/data-model.md](architecture/data-model.md) | 数据模型总览 |
| [architecture/evolution-roadmap.md](architecture/evolution-roadmap.md) | Agent 架构演进路线 |
| [issues/rag-refactor-plan.md](issues/rag-refactor-plan.md) | RAG 重构计划和阶段拆解 |
| [demo/README.md](demo/README.md) | Demo 运行和面试演示材料 |
| [demo/ten-minute-interview-demo.md](demo/ten-minute-interview-demo.md) | 10 分钟面试演示脚本 |
| [eval/README.md](eval/README.md) | 诊断评测材料 |
| [issues/README.md](issues/README.md) | MVP issue 索引 |
## 一、设计原则
## 当前系统一句话
### 1.1 核心原则
- ✅ **简单优先**:满足诊断流程需要,避免过度设计
- ✅ **渐进增强**:先实现核心功能,再逐步扩展
- ✅ **数据分离**:诊断结果持久化(MySQL),会话上下文临时化(Redis)
- ✅ **适度冗余**:避免过度范式化,适当冗余提升查询性能
SuperBizAgent MVP 是一个可追踪的故障诊断 Agent:Chat 和 AIOps 入口进入 Agent 编排,Executor 显式调用知识库、日志、指标等工具收集证据,诊断过程落到 `diagnosis_session`、`agent_step`、`tool_invocation`,最终通过 Trace API、Verifier 和评测脚本证明结果可解释、可回放、可对比。
### 1.2 系统定位
**自动化诊断系统**
- 核心:一键诊断 → 返回完整报告
- 辅助:支持追问,但不是主要场景
- 特点:大部分用户单次诊断即结束,少数用户会追问细节
## 文档结构
---
## 二、表结构总览
### 2.1 核心表关系
```
┌─────────────────────┐
│ diagnosis_record │ 诊断记录(核心)
│ - 每次诊断一条 │
└──────────┬──────────┘
│ 1:1
↓
┌─────────────────────┐
│ case_library │ 案例库(知识沉淀)
│ - 诊断成功→案例 │
└─────────────────────┘
┌─────────────────────┐
│ api_document │ 文档元数据(管理层)
│ - 状态追踪/去重 │
└──────────┬──────────┘
│ doc_id
↓
┌─────────────────────┐
│ Milvus │ 文档内容(检索层)
│ - 向量检索 │
└─────────────────────┘
┌─────────────────────┐
│ Redis Session │ 会话管理(临时)
│ - 30分钟过期 │
│ - 支持追问 │
└─────────────────────┘
```text
mvp/
architecture/
README.md
current-mvp-architecture.md
interview-one-pager.md
agent-orchestration.md
harness-quality-gates.md
rag-architecture.md
retrieval-observability.md
feedback-architecture.md
session-trace-lifecycle.md
knowledge-base-authoring.md
data-model.md
evolution-roadmap.md
archive/2026-07-05-legacy/
issues/
README.md
rag-refactor-plan.md
ISS-*.md
rag-*.md
demo/
README.md
ten-minute-interview-demo.md
requests/
scripts/
output/
eval/
README.md
schema.md
cases/
fixtures/
reports/
notes/
plan/
tables/
```
### 2.2 表统计
## 当前核心设计
| 表名 | 类型 | 预估数据量 | 用途 |
|------|------|-----------|------|
| diagnosis_record | 核心 | 3.6万/年 | 诊断记录 |
| case_library | 核心 | 500-1000 | 案例库 |
| api_document | 核心 | 100-200 | 文档管理 |
- `lookup_knowledge` 保持显式 Agent Tool,不隐藏到 Chat Advisor。
- L0 降级为 domain/entity hint,不再默认承担最终召回决策。
- `VectorSearchService` 是检索稳定门面。
- Spring AI VectorStore 是当前读取主路径,Milvus SDK 保留为 fallback。
- AIOps payload 会生成推荐知识库 query,保留业务语义。
- Trace API 聚合 session、step、tool invocation 和 self evaluation。
- RAG 行为通过 offline baseline 和 live acceptance 脚本做回归验证。
---
## 关键运行链路
## 三、技术栈
```text
Chat
-> ChatService
-> Planner / Executor / Verifier
-> evidence tools
-> diagnosis_session / agent_step / tool_invocation
-> DiagnosisTraceService
### 3.1 数据存储
```
MySQL 8.0+
├─ 元数据管理
├─ 事务支持
└─ JSON 字段支持
AIOps
-> AiOpsService
-> PAYLOAD_TARGETED or AUTO_DISCOVERY
-> Planner / Executor
-> Prometheus / logs / lookup_knowledge
-> AiOpsRuleEvaluationService
-> DiagnosisTraceService
Redis 6.0+
├─ 会话存储
├─ 缓存
└─ TTL 自动过期
Milvus 2.6+
├─ 向量存储
├─ 语义检索
└─ 混合检索
RAG
-> lookup_knowledge
-> L0 domain/entity hint
-> VectorSearchService
-> Spring AI VectorStore / Milvus SDK fallback
-> relevance normalization
-> tool_invocation
```
### 3.2 开发框架
```
Spring Boot 3.2
Spring AI Alibaba 1.1.0
Milvus SDK Java 2.6.10
DashScope SDK
```
## 旧文档说明
---
旧版架构文档已移动到:
## 四、快速开始
- [architecture/archive/2026-07-05-legacy/](architecture/archive/2026-07-05-legacy/)
### 4.1 创建数据库
```sql
-- 1. 创建数据库
CREATE DATABASE diagnosis_system CHARACTER SET utf8mb4 COLLATE utf8mb4_unicode_ci;
-- 2. 执行建表脚本(按顺序)
SOURCE tables/diagnosis_record.sql;
SOURCE tables/case_library.sql;
SOURCE tables/api_document.sql;
```
### 4.2 初始化 Milvus
```java
// 创建 Collection
MilvusClientFactory.createCollection();
```
### 4.3 配置 Redis
```yaml
spring:
redis:
host: localhost
port: 6379
database: 0
```
---
## 五、版本历史
| 版本 | 日期 | 变更内容 |
|------|------|---------|
| v1.0 | 2024-06-15 | 初版,定义核心表结构 |
| v2.0 | 2024-06-15 | diagnosis_record 字段泛化,支持多种故障类型 |
| v2.1 | 2024-06-22 | 文档拆分,增加 api_document 表 |
---
## 六、维护说明
- 每个表的详细设计在 `tables/` 目录下
- 架构设计文档在 `architecture/` 目录下
- 修改表结构时,同步更新对应的 Markdown 文档
- 重大变更需记录在版本历史中
归档文档只用于追溯设计历史。当前实现和后续规划以 `architecture/current-mvp-architecture.md` 与 `architecture/rag-architecture.md` 为准。
+42
View File
@@ -0,0 +1,42 @@
# MVP 架构文档
**更新日期**:2026-07-05
这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到:
- `mvp/architecture/archive/2026-07-05-legacy/`
归档材料只作为设计历史阅读,不再作为当前实现依据。
## 当前文档
| 文档 | 用途 |
|---|---|
| [current-mvp-architecture.md](current-mvp-architecture.md) | 当前可运行 MVP 的总体架构、链路、持久化和质量门禁 |
| [interview-one-pager.md](interview-one-pager.md) | 面试一页式架构讲解,包含总图、亮点、取舍和追问回答 |
| [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat SequentialAgent、AIOps SupervisorAgent、工具边界 |
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Verifier、评测基线组成的质量门禁 |
| [rag-architecture.md](rag-architecture.md) | RAG/知识检索新架构,覆盖 L0 hint、VectorStore 主路径、SDK fallback、证据追踪 |
| [retrieval-observability.md](retrieval-observability.md) | 检索运行细节和可观测性,覆盖 L0/L1、去重、分数归一、评测 |
| [feedback-architecture.md](feedback-architecture.md) | 反馈与自评估闭环,覆盖 rule evaluation、Verifier、AIOps rule、用户反馈和案例沉淀 |
| [session-trace-lifecycle.md](session-trace-lifecycle.md) | 会话和 Trace 生命周期,覆盖 sessionId、状态流转、agent_step、tool_invocation、Trace API |
| [knowledge-base-authoring.md](knowledge-base-authoring.md) | 知识库文档编写与维护规范,覆盖 frontmatter、category、chunk、reindex |
| [data-model.md](data-model.md) | 数据模型总览,覆盖 Trace、知识库、反馈沉淀和 Milvus metadata |
| [evolution-roadmap.md](evolution-roadmap.md) | 从旧版 Agent 蓝图继承的后续演进路线,不代表当前已实现 |
## 当前架构一句话
SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,执行过程落到 `diagnosis_session`、`agent_step`、`tool_invocation`,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
## 阅读顺序
1. 先读 [current-mvp-architecture.md](current-mvp-architecture.md),理解系统边界和主链路。
2. 面试前读 [interview-one-pager.md](interview-one-pager.md),准备 2-5 分钟讲解。
3. 再读 [agent-orchestration.md](agent-orchestration.md),理解当前 Agent 如何协作。
4. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。
5. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。
6. 继续读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
7. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
8. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
9. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
10. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
+187
View File
@@ -0,0 +1,187 @@
# Agent 编排架构
**更新日期**:2026-07-05
**状态**:当前可运行架构
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
## 1. 设计定位
旧版 Agent 架构把系统描述为 Supervisor、Planner、SubAgent、Verifier 的团队协作。当前 MVP 保留这个核心思想,但实现更收敛:
- Chat 链路使用固定顺序工作流:`Planner -> Executor -> Verifier`。
- AIOps 链路使用 `SupervisorAgent` 调度 `Planner + Executor`,最终由规则评估器做轻量验证。
- 当前没有拆分 ExternalApiSubAgent、InternalErrorSubAgent、DatabaseSubAgent;这些作为后续演进方向保留。
- 证据工具不直接散落在各个 Agent 里,而是通过 Spring AI ToolCallback / `@Tool` 统一暴露。
## 2. 当前 Agent 全景
```mermaid
flowchart TB
subgraph Chat["Chat diagnosis"]
ChatIn["POST /api/chat"] --> ChatService["ChatService"]
ChatService --> ChatPlanner["chat_planner"]
ChatPlanner --> ChatExecutor["chat_executor"]
ChatExecutor --> ChatTools["evidence tools"]
ChatTools --> ChatExecutor
ChatExecutor --> ChatVerifier["chat_verifier"]
ChatVerifier --> ChatDecision{"PASS / LOW_CONFID / REJECT"}
ChatDecision --> ChatAnswer["final answer"]
end
subgraph AiOps["AIOps diagnosis"]
AiOpsIn["POST /api/ai_ops"] --> AiOpsService["AiOpsService"]
AiOpsService --> Supervisor["ai_ops_supervisor"]
Supervisor --> AiOpsPlanner["planner_agent"]
Supervisor --> AiOpsExecutor["executor_agent"]
AiOpsPlanner --> AiOpsExecutor
AiOpsExecutor --> AiOpsTools["Prometheus / logs / lookup_knowledge"]
AiOpsTools --> AiOpsReport["alert report"]
AiOpsReport --> AiOpsRule["AiOpsRuleEvaluationService"]
end
subgraph Trace["Trace persistence"]
Session["diagnosis_session"]
Step["agent_step"]
Invocation["tool_invocation"]
SelfEval["self_evaluation"]
end
ChatService --> Session
ChatPlanner --> Step
ChatExecutor --> Step
ChatVerifier --> Step
ChatTools --> Invocation
ChatDecision --> SelfEval
AiOpsService --> Session
AiOpsPlanner --> Step
AiOpsExecutor --> Step
AiOpsTools --> Invocation
AiOpsRule --> SelfEval
```
## 3. Chat 编排
Chat 复杂诊断采用 `SequentialAgent`,顺序固定:
```text
chat_planner
-> chat_executor
-> lookup_knowledge / query_logs / query_metrics / date_time
-> chat_verifier
-> reads tool_trace_summary
-> outputs verifier JSON
```
关键行为:
| 角色 | 当前职责 | 输出 |
|---|---|---|
| `chat_planner` | 拆解问题,注入知识域地图和对话历史,给出排查方向 | `planner_plan` |
| `chat_executor` | 按计划调用证据工具,组合工具返回形成诊断答复 | `executor_feedback` |
| `chat_verifier` | 只基于已有证据校验 Executor 答案,不做新检索 | `verifier_output` |
Chat 链路最多支持两轮验证:
```mermaid
sequenceDiagram
autonumber
participant C as ChatService
participant P as chat_planner
participant E as chat_executor
participant T as tools
participant V as chat_verifier
participant S as diagnosis_session
C->>P: 原始问题 + history + retry_context
P-->>C: planner_plan
C->>E: planner_plan + 上下文
E->>T: 调用证据工具
T-->>E: 证据结果
E-->>C: executor_feedback
C->>V: executor_final_answer + tool_trace_summary
V-->>C: PASS / LOW_CONFID / REJECT
C->>S: 写入 verifier_evaluation
alt LOW_CONFID 且允许补证据
C->>P: retry_context: 仅补缺失证据
else PASS 或 REJECT
C-->>S: 保存最终 answer
end
```
决策语义:
| Verdict | 行为 |
|---|---|
| `PASS` | 输出 Executor 答案 |
| `LOW_CONFID` | 如果分数低于阈值且仍有轮次,构造 `retry_context` 补证据;否则输出低置信提示 |
| `REJECT` | 输出降级答复,只保留已确认信息和下一步建议 |
## 4. AIOps 编排
AIOps 使用 `SupervisorAgent` 调度两个子 Agent:
```text
ai_ops_supervisor
-> planner_agent
-> executor_agent
-> final report
-> AiOpsRuleEvaluationService
```
与 Chat 的差异:
- AIOps 的输入可能是结构化告警 payload。
- payload 模式会进入 `PAYLOAD_TARGETED`,最终报告必须聚焦输入告警。
- 无 payload 时进入 `AUTO_DISCOVERY`,先通过告警工具发现活跃告警。
- 当前 AIOps 不使用 LLM Verifier,而使用轻量规则评估器写入 `self_evaluation.aiops_rule_evaluation`。
## 5. 工具边界
当前 Executor 可用工具来自两类:
```text
methodTools
-> dateTimeTools
-> lookupKnowledgeTool
-> queryMetricsTools
-> queryLogsTools when mock enabled
ToolCallbackProvider
-> framework-discovered tools
```
工具调用必须写入 `tool_invocation`。其中 `lookup_knowledge` 额外记录:
- L0/L1 命中数量。
- 检索层。
- relevance level。
- retrieved domains。
- dedup reason。
## 6. 与旧版设计的差异
| 旧版设想 | 当前实现 |
|---|---|
| Supervisor + Planner + 多个专科 SubAgent + Verifier | Chat: Planner + Executor + Verifier;AIOps: Supervisor + Planner + Executor |
| ExternalApiSubAgent / InternalErrorSubAgent / DatabaseSubAgent | 暂未拆分,能力通过通用 Executor + 工具 + Prompt 约束实现 |
| 每个 SubAgent 专属工具集 | 当前 Executor 持有统一证据工具集合 |
| Verifier 支持 PASS / REVISE / REJECT | 当前 Chat Verifier 输出 PASS / LOW_CONFID / REJECT |
| Skill 驱动不同诊断流程 | 当前以 Prompt、知识域地图、工具调用和评测 baseline 控制 |
## 7. 后续演进
当诊断场景和工具复杂度继续上升时,再考虑拆分:
- `ExternalApiSubAgent`:接口文档、错误码、请求参数、第三方日志。
- `DatabaseSubAgent`:连接池、慢 SQL、死锁、索引建议。
- `CacheSubAgent`:Redis 超时、连接、热点 key、内存风险。
- `GenericDiagnosisSubAgent`:专项 Agent 失败后的兜底。
拆分前提:
- 当前 Executor prompt 已难以维护。
- 不同故障类型的工具权限明显不同。
- Trace 能证明某类问题需要独立的推理策略。
- 评测集能覆盖拆分前后的行为差异。
@@ -0,0 +1,28 @@
# 旧版架构文档归档
**归档日期**:2026-07-05
本目录保存 `mvp/architecture` 下的旧版架构文档。它们包含早期 MVP 设计、旧 RAG 方案、会话存储设计、行动记忆和实施计划等历史材料。
这些文档不再作为当前实现依据。当前架构请阅读:
- `mvp/architecture/README.md`
- `mvp/architecture/current-mvp-architecture.md`
- `mvp/architecture/rag-architecture.md`
## 归档文件
| 文件 | 说明 |
|---|---|
| `agent-architecture.md` | 早期完整 Agent 设想,包含较多超出当前 MVP 的 SubAgent 设计 |
| `agent-architecture-mvp.md` | 早期 MVP Agent 设计 |
| `knowledge-retrieval-architecture.md` | 旧版 L0 + L1 检索架构,包含 L0 唯一命中跳过 L1 的旧逻辑 |
| `knowledge-retrieval-usage.md` | 旧版知识库检索使用说明 |
| `current-mvp-architecture.md` | 归档前的当前架构快照 |
| `implementation-plan.md` | 早期实施计划 |
| `implementation-detail.md` | 早期完整实施计划 |
| `session-management.md` | 会话管理旧设计 |
| `session-dedup-knowledge-map.md` | 会话去重和知识域地图设计 |
| `confidence-feedback.md` | 证据评分和用户反馈旧设计 |
| `action-memory-relevance.md` | 行动记忆和检索质量归一化旧设计 |
@@ -0,0 +1,220 @@
# Current MVP Architecture Snapshot
**Updated**: 2026-07-05
This document records the current runnable MVP architecture. Older architecture notes in this folder still represent design history; this file should be read as the current snapshot for demos, interviews, and next-step planning.
## 1. Positioning
The MVP is an Agent engineering project for traceable troubleshooting, not a generic chatbot.
Core goals:
- Support normal chat-based diagnosis.
- Support AIOps alert-triggered diagnosis.
- Keep tool calls explicit and traceable.
- Keep RAG retrieval observable through `lookup_knowledge`.
- Persist enough execution evidence for replay, evaluation, and interview explanation.
## 2. Runtime Architecture
```text
HTTP API
-> ChatService / AiOpsService
-> Agent orchestration
-> Supervisor / Planner / Executor / Verifier
-> Tools
-> lookup_knowledge
-> query_logs
-> query_metrics
-> other diagnosis tools
-> Persistence
-> diagnosis_session
-> agent_step
-> tool_invocation
-> Trace API
-> DiagnosisTraceService
```
Current entry points:
- `ChatService`: user-driven troubleshooting and follow-up diagnosis.
- `AiOpsService`: alert-driven diagnosis, including payload mode and auto-discovery mode.
- `DiagnosisTraceService`: trace view of session, steps, tool calls, and self-evaluation.
## 3. Chat Diagnosis Flow
```text
User question
-> ChatService
-> simple response or diagnosis flow
-> Planner creates investigation direction
-> Executor calls tools for evidence
-> lookup_knowledge
-> query_logs
-> query_metrics
-> Verifier checks final diagnosis quality
-> self_evaluation.verifier_evaluation
-> diagnosis trace
```
The chat path uses the LLM verifier as the main quality gate. The verifier result is persisted under `diagnosis_session.self_evaluation.verifier_evaluation`.
## 4. AIOps Diagnosis Flow
```text
AIOps request
-> AiOpsService
-> payload mode or auto-discovery mode
-> build alert-focused diagnosis prompt
-> append recommended lookup_knowledge query when payload exists
-> Agent diagnosis flow
-> Supervisor / Planner / Executor
-> evidence tools
-> final report
-> AiOpsRuleEvaluationService
-> self_evaluation.aiops_rule_evaluation
-> diagnosis trace
```
AIOps keeps two modes:
- Payload mode: the request already contains alert fields such as alert name, service, metric, severity, and symptom. The system builds a recommended knowledge query from these fields.
- Auto-discovery mode: the system follows the original alert-discovery behavior and lets the Agent collect alert context through tools.
The AIOps verifier is currently lightweight and rule-based. It checks:
- Whether the final report exists.
- Whether the result stays focused on the alert payload when payload exists.
- Whether evidence tools were used, especially `lookup_knowledge`, `query_logs`, and `query_metrics`.
## 5. RAG Architecture
```text
lookup_knowledge
-> L0 domain/entity hint
-> matched domain
-> matched keywords/entities
-> metadata filter signal
-> VectorSearchService
-> Spring AI VectorStore path
-> Milvus SDK fallback path
-> evidence post-processing
-> score / rawScore / scoreLabel
-> source metadata
-> title / breadcrumb / content evidence block
-> tool_invocation record
```
Important decisions:
- `lookup_knowledge` remains an explicit Agent tool. It is not replaced by an implicit chat Advisor because the project needs visible Agent decision-making.
- L0 is retained but downgraded. It is a domain/entity hint and explainability signal, not the final recall decision.
- L1 retrieval now goes through `VectorSearchService`.
- Spring AI `VectorStore` is the preferred retrieval path.
- The original Milvus SDK path is retained as fallback and compatibility path.
- `title`, `breadcrumb`, and `content` participate in embedding text so chunk context is less likely to be lost.
- Retrieval output keeps compatibility fields: `score`, `rawScore`, and `scoreLabel`.
Vector retrieval modes:
```text
retrieval.vector-store.mode=auto # Prefer Spring AI VectorStore, fallback to SDK
retrieval.vector-store.mode=spring-ai # Use Spring AI VectorStore only
retrieval.vector-store.mode=sdk # Use original Milvus SDK path
```
## 6. Persistence And Trace
Current trace-related persistence:
```text
diagnosis_session
-> final_report
-> self_evaluation
-> verifier_evaluation
-> aiops_rule_evaluation
agent_step
-> role
-> step input/output
-> execution order
tool_invocation
-> tool_name
-> query
-> retrieval_layer
-> retrieval_details
-> evidence blocks
-> duration
```
Trace API aggregates these records into a session-level view:
- Agent step sequence.
- Tool calls and retrieval details.
- Final diagnosis report.
- Chat verifier status.
- AIOps rule verifier status.
## 7. Quality Gates
Current quality gates:
- Chat verifier: LLM-based final answer verification for normal diagnosis.
- AIOps rule verifier: lightweight deterministic checks for alert-focused diagnosis.
- Diagnosis eval baseline: fixture-based evaluation for trace and evidence behavior.
- RAG retrieval baseline: golden query set with offline baseline report.
- Live RAG acceptance: post-reindex script for validating retrieval against the running stack.
These gates are intentionally layered. The MVP proves the Agent chain can produce evidence, persist it, and be inspected after execution.
## 8. Current Completion State
Completed for the current MVP stage:
- Explicit `lookup_knowledge` Agent tool.
- L0 + L1 retrieval shape retained.
- L0 downgraded to domain/entity hint.
- Spring AI VectorStore retrieval path integrated.
- Milvus SDK fallback retained.
- RAG evidence post-processing added.
- Breadcrumb/title/content embedding text improved.
- RAG offline baseline and live acceptance script added.
- AIOps payload query augmentation added.
- AIOps lightweight verifier added.
- Trace summary includes both chat verifier and AIOps verifier signals.
Deferred future enhancements:
- LLM QueryTransformer / MultiQuery.
- BM25, RRF, and reranker.
- Neighbor chunk or section-level context expansion.
- VectorStore write path migration.
- Full LLM-based AIOps verifier.
- More complete golden set for recall, MRR, and nDCG metrics.
## 9. Key Code References
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
- `src/main/java/com/superbiz/agent/service/AiOpsRuleEvaluationService.java`
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
- `src/main/java/com/superbiz/agent/service/VectorIndexService.java`
- `src/main/java/com/superbiz/agent/service/SpringAiVectorStoreSidecarService.java`
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
- `src/main/java/com/superbiz/agent/service/SelfEvaluationMergeService.java`
## 10. Supporting Materials
- `mvp/issues/rag-refactor-plan.md`
- `eval/rag-retrieval/README.md`
- `scripts/eval_rag_live_acceptance.py`
- `interview/rag-refactor-story.md`
- `interview/rag-vectorstore-interview-notes.md`
- `interview/rag-retrieval-quality-report.md`
- `interview/rag-breadcrumb-embedding-acceptance.md`
- `interview/aiops-query-augmentation.md`
- `interview/aiops-lightweight-verifier.md`
@@ -0,0 +1,391 @@
# 当前 MVP 架构
**更新日期**:2026-07-05
**状态**:当前可运行架构
**适用范围**:Demo、面试讲解、后续迭代规划
## 1. 系统定位
SuperBizAgent MVP 不是通用 Chatbot,而是面向故障诊断的 Agent 工程项目。
核心目标:
- 支持用户主动发起的 Chat 诊断。
- 支持 AIOps 告警触发的自动诊断。
- 保留 Agent 的规划、执行、验证过程。
- 工具调用必须显式、可追踪、可回放。
- RAG 检索必须通过 `lookup_knowledge` 暴露证据链。
- 每次诊断都沉淀 session、step、tool invocation 和 self evaluation。
## 2. 总体分层
```mermaid
flowchart TB
subgraph API["API Layer"]
ChatController["ChatController"]
TraceController["DiagnosisTraceController"]
SearchController["SearchController"]
DocumentController["DocumentController"]
end
subgraph App["Application Service"]
ChatService["ChatService"]
AiOpsService["AiOpsService"]
TraceService["DiagnosisTraceService"]
end
subgraph Agent["Agent Orchestration"]
Supervisor["Supervisor"]
Planner["Planner"]
Executor["Executor"]
Verifier["Verifier"]
end
subgraph Tools["Evidence Tools"]
KnowledgeTool["lookup_knowledge"]
LogsTool["query_logs"]
MetricsTool["query_metrics"]
AlertsTool["queryPrometheusAlerts"]
end
subgraph RAG["RAG Retrieval"]
L0["KnowledgeIndexService"]
VectorSearch["VectorSearchService"]
VectorStore["Spring AI VectorStore"]
SdkFallback["Milvus SDK fallback"]
end
subgraph Store["Persistence and Trace"]
Session["diagnosis_session"]
Step["agent_step"]
Invocation["tool_invocation"]
ApiDoc["api_document"]
Milvus["Milvus/Zilliz"]
end
API --> App
ChatService --> Agent
AiOpsService --> Agent
Agent --> Tools
KnowledgeTool --> RAG
RAG --> Store
Tools --> Invocation
Agent --> Step
App --> Session
TraceService --> Session
TraceService --> Step
TraceService --> Invocation
```
```text
API Layer
-> ChatController
-> DiagnosisTraceController
-> SearchController
-> DocumentController
Application Service
-> ChatService
-> AiOpsService
-> DiagnosisTraceService
Agent Orchestration
-> Supervisor
-> Planner
-> Executor
-> Verifier
Evidence Tools
-> lookup_knowledge
-> query_logs
-> query_metrics
-> queryPrometheusAlerts
RAG Retrieval
-> KnowledgeIndexService
-> VectorSearchService
-> Spring AI VectorStore
-> Milvus SDK fallback
Persistence
-> diagnosis_session
-> agent_step
-> tool_invocation
-> api_document
-> Milvus/Zilliz collection
Quality Gates
-> chat verifier
-> AIOps rule evaluation
-> diagnosis eval baseline
-> RAG retrieval baseline
```
## 3. Chat 诊断链路
```mermaid
sequenceDiagram
autonumber
actor User as 用户
participant API as POST /api/chat
participant Chat as ChatService
participant Planner as Planner Agent
participant Executor as Executor Agent
participant Tool as Evidence Tools
participant Verifier as Verifier Agent
participant DB as Trace Tables
participant Trace as Trace API
User->>API: 提交诊断问题
API->>Chat: execute chat strategy
Chat->>Planner: 复杂问题进入规划
Planner->>DB: 写入 agent_step
Planner->>Executor: 下发排查方向
Executor->>Tool: lookup_knowledge / logs / metrics
Tool->>DB: 写入 tool_invocation
Tool-->>Executor: 返回证据
Executor->>Verifier: 生成候选诊断并校验
Verifier->>DB: 合并 self_evaluation.verifier_evaluation
Chat->>DB: 保存 diagnosis_session.answer
User->>Trace: GET /api/diagnosis/{sessionId}/trace
Trace->>DB: 聚合 session / step / tool
Trace-->>User: 返回可回放诊断链路
```
```text
POST /api/chat
-> ChatService
-> 简单问题:轻量回答
-> 复杂诊断:Agent 编排
-> Planner 制定排查方向
-> Executor 调用证据工具
-> lookup_knowledge
-> query_logs
-> query_metrics
-> Verifier 校验最终诊断
-> 保存 diagnosis_session
-> 保存 agent_step
-> 保存 tool_invocation
-> 合并 self_evaluation.verifier_evaluation
```
Chat 链路的质量门禁是 LLM Verifier。Verifier 输出合并到 `diagnosis_session.self_evaluation.verifier_evaluation`,Trace API 会展示该验证结果。
Agent 编排细节见 [agent-orchestration.md](agent-orchestration.md)。
关键代码:
- `src/main/java/com/superbiz/agent/controller/ChatController.java`
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- `src/main/java/com/superbiz/agent/service/SelfEvaluationMergeService.java`
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
## 4. AIOps 诊断链路
```mermaid
flowchart TD
Request["POST /api/ai_ops"] --> Payload{"包含告警 payload?"}
Payload -->|是| Targeted["PAYLOAD_TARGETED"]
Payload -->|否| Discovery["AUTO_DISCOVERY"]
Targeted --> BuildPrompt["构造聚焦 payload 的诊断 prompt"]
Targeted --> QueryAug["生成 recommended lookup_knowledge query"]
Discovery --> DiscoverAlert["通过 queryPrometheusAlerts 发现活跃告警"]
BuildPrompt --> Plan["Planner 规划排查"]
QueryAug --> Plan
DiscoverAlert --> Plan
Plan --> Execute["Executor 收集证据"]
Execute --> Knowledge["lookup_knowledge"]
Execute --> Metrics["query_metrics / Prometheus"]
Execute --> Logs["query_logs"]
Knowledge --> Report["告警分析报告"]
Metrics --> Report
Logs --> Report
Report --> RuleEval["AiOpsRuleEvaluationService"]
RuleEval --> SelfEval["self_evaluation.aiops_rule_evaluation"]
Report --> Trace["DiagnosisTraceService"]
SelfEval --> Trace
```
```text
POST /api/ai_ops
-> AiOpsService
-> 判断是否有告警 payload
-> PAYLOAD_TARGETED
-> AUTO_DISCOVERY
-> 构造 AIOps 诊断 prompt
-> payload 模式补充 recommended lookup_knowledge query
-> Agent 编排
-> Planner / Executor
-> Prometheus / logs / knowledge tools
-> 生成告警分析报告
-> AiOpsRuleEvaluationService
-> 合并 self_evaluation.aiops_rule_evaluation
-> Trace API 可查看全链路
```
AIOps 保留两种模式:
| 模式 | 触发条件 | 行为 |
|---|---|---|
| `PAYLOAD_TARGETED` | 请求包含 alertName、service、severity、description、timeRange 等字段 | 以 payload 为唯一主诊断对象,并生成推荐知识库 query |
| `AUTO_DISCOVERY` | 请求没有明确告警 payload | 先查询当前活跃告警,再选择目标排查 |
AIOps 当前使用轻量规则验证器,重点检查:
- 最终报告是否存在。
- payload 模式是否聚焦输入告警。
- 是否使用关键证据工具,例如 `lookup_knowledge`、日志、指标。
关键代码:
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
- `src/main/java/com/superbiz/agent/service/AiOpsRuleEvaluationService.java`
## 5. RAG 位置
RAG 不是隐藏在 Chat Advisor 里的隐式能力,而是 Executor 可以显式调用的工具:
```mermaid
flowchart LR
Executor["Executor Agent"] --> Tool["lookup_knowledge Tool"]
Tool --> L0["L0 domain/entity hint"]
Tool --> Search["VectorSearchService"]
L0 --> Search
Search --> VectorStore["Spring AI VectorStore"]
Search --> Fallback["Milvus SDK fallback"]
VectorStore --> Normalize["score/rawScore/scoreLabel"]
Fallback --> Normalize
Normalize --> Evidence["evidence output"]
Evidence --> Invocation["tool_invocation"]
Evidence --> Executor
```
```text
Executor
-> lookup_knowledge(query)
-> L0 domain/entity hint
-> VectorSearchService
-> Spring AI VectorStore
-> Milvus SDK fallback
-> evidence shaping
-> tool_invocation
```
保留显式工具的原因:
- Agent 何时检索、检索什么、证据是什么,必须能在 trace 中解释。
- AIOps payload 到 query 的业务映射需要项目内控制。
- `tool_invocation` 是后续评测、回放和面试讲解的核心材料。
RAG 总体设计见 [rag-architecture.md](rag-architecture.md),检索运行细节见 [retrieval-observability.md](retrieval-observability.md)。
## 6. 持久化模型
当前诊断持久化以三张表为核心:
```text
diagnosis_session
-> 一次诊断会话的主记录
-> query / status / agent_flow / answer
-> self_evaluation
-> step_count / tool_call_count / duration
agent_step
-> Agent 模型调用步骤
-> step_index / agent_name
-> model_input / model_output / thought
-> duration / token_count
tool_invocation
-> 工具调用事实
-> tool_name / input_params / output_preview
-> retrieval_layer / retrieval_details
-> relevance_level / dedup_reason
-> duration / success
```
说明:
- 旧的 `diagnosis_record` 已不是当前主模型,迁移脚本中已经由 `diagnosis_session + agent_step + tool_invocation` 取代。
- `api_document` 仍用于文档元数据管理。
- 文档向量内容存放在 Milvus/Zilliz collection 中。
会话和 Trace 生命周期见 [session-trace-lifecycle.md](session-trace-lifecycle.md),完整数据关系见 [data-model.md](data-model.md)。
## 7. Trace API
```text
GET /api/diagnosis/{sessionId}/trace
```
Trace API 聚合:
- 会话状态和最终报告。
- Agent step 序列。
- 工具调用和检索细节。
- Chat verifier 结果。
- AIOps rule evaluation 结果。
Trace 是本项目区别于普通问答系统的关键:答案不是孤立文本,而是可以追溯到 Agent 决策、工具调用和证据来源。
Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gates.md](harness-quality-gates.md),用户反馈与 `self_evaluation` 闭环见 [feedback-architecture.md](feedback-architecture.md)。
## 8. 质量门禁
当前质量门禁分层如下:
| 门禁 | 位置 | 作用 |
|---|---|---|
| Chat Verifier | `ChatService` | 校验普通诊断回答质量 |
| AIOps Rule Evaluation | `AiOpsRuleEvaluationService` | 校验告警诊断是否聚焦 payload 并使用证据 |
| Diagnosis Eval Baseline | `mvp/eval/` | 固化诊断 trace 和报告行为 |
| RAG Retrieval Baseline | `eval/rag-retrieval/` | 固化检索召回行为,避免 RAG 重构回退 |
| Live RAG Acceptance | `scripts/eval_rag_live_acceptance.py` | 在运行环境中验证重建索引后的真实检索 |
## 9. 当前完成状态
已经完成:
- Chat 和 AIOps 两条入口链路。
- 显式 `lookup_knowledge` Agent Tool。
- L0 从最终决策降级为 domain/entity hint。
- `VectorSearchService` 作为稳定检索门面。
- Spring AI VectorStore 读取路径。
- Milvus SDK fallback。
- `score` / `rawScore` / `scoreLabel` 分数语义拆分。
- `title`、`breadcrumb`、`content` 参与 embedding 文本。
- `tool_invocation` 记录检索层、relevance level、dedup reason。
- Chat verifier 和 AIOps rule evaluation 合并进 `self_evaluation`。
- RAG offline baseline 和 live acceptance 脚本。
暂不作为当前已完成能力声明:
- 完整 QueryTransformer / MultiQuery。
- BM25、RRF、cross-encoder rerank。
- 完整邻居 chunk / section context expansion。
- VectorStore 写入路径全面迁移。
- 完整 LLM-based AIOps verifier。
后续 Agent 拆分、Skill/Playbook、MCP 工具协议化和进程隔离等方向见 [evolution-roadmap.md](evolution-roadmap.md)。
## 10. 关键代码索引
| 能力 | 代码 |
|---|---|
| Chat 入口与编排 | `ChatController`, `ChatService` |
| AIOps 入口与编排 | `ChatController.aiOps`, `AiOpsService` |
| AIOps 规则验证 | `AiOpsRuleEvaluationService` |
| 知识库工具 | `LookupKnowledgeTool` |
| L0 hint | `KnowledgeIndexService` |
| 向量检索门面 | `VectorSearchService` |
| 文档切片 | `DocumentChunkService` |
| 向量写入 | `VectorIndexService` |
| Spring AI VectorStore 配置辅助 | `SpringAiVectorStoreSidecarService` |
| Trace 聚合 | `DiagnosisTraceService` |
| 工具调用记录 | `ToolInvocationRecorder` |
| self_evaluation 合并 | `SelfEvaluationMergeService` |
+258
View File
@@ -0,0 +1,258 @@
# 数据模型总览
**更新日期**:2026-07-05
**状态**:当前可运行架构
## 1. 定位
本文从架构角度说明当前 MVP 的核心数据模型。详细字段仍以 Flyway migration 和 `mvp/tables/` 为准。
核心数据分三组:
- 诊断 Trace:`diagnosis_session`、`agent_step`、`tool_invocation`
- 知识库:`api_document`、`knowledge_domain`、Milvus/Zilliz metadata
- 反馈沉淀:`case_library`
## 2. 总体关系
```mermaid
erDiagram
diagnosis_session ||--o{ agent_step : has
diagnosis_session ||--o{ tool_invocation : has
diagnosis_session ||--o| case_library : creates_when_useful
api_document ||--o{ milvus_chunk : indexed_as
knowledge_domain ||--o{ api_document : groups
diagnosis_session {
bigint id
varchar session_id
text query
varchar status
varchar agent_flow
longtext answer
json self_evaluation
varchar feedback
}
agent_step {
bigint id
varchar session_id
int step_index
varchar agent_name
text model_input
text model_output
text thought
boolean has_tool_call
}
tool_invocation {
bigint id
varchar session_id
varchar tool_name
json input_params
text output_preview
varchar retrieval_layer
json retrieval_details
varchar relevance_level
varchar dedup_reason
}
api_document {
bigint id
varchar doc_id
varchar file_name
varchar file_path
varchar status
int chunk_count
text metadata
}
knowledge_domain {
bigint id
varchar domain_id
varchar description
text when_to_retrieve
int document_count
}
case_library {
bigint id
varchar case_id
varchar diagnosis_id
varchar source_type
varchar fault_category
text root_cause
text solution
}
milvus_chunk {
varchar id
text content
json metadata
vector vector
}
```
说明:Milvus/Zilliz collection 不是 MySQL 表,图中的 `milvus_chunk` 是逻辑模型。
## 3. 诊断 Trace 模型
### diagnosis_session
会话级主记录。
关键字段:
| 字段 | 说明 |
|---|---|
| `session_id` | 外部关联键,Trace 和 Feedback 都使用它 |
| `query` | 用户原始问题或 AIOps 输入摘要 |
| `status` | 执行状态 |
| `agent_flow` | `CHAT` / `AI_OPS` |
| `answer` | 最终答复或告警报告 |
| `self_evaluation` | rule/verifier/aiops 自评估容器 |
| `feedback` | 用户反馈 |
### agent_step
记录模型调用步骤。
用途:
- 回放 Agent 推理过程。
- 查看 Planner / Executor / Verifier 的输入输出摘要。
- 统计 step count、duration、token count。
### tool_invocation
记录工具调用事实。
用途:
- 给 Trace API 展示证据。
- 给 Verifier 构造 `tool_trace_summary`。
- 给 `EvaluationService` 计算 evidence score。
- 给 RAG eval 和人工排查提供检索细节。
## 4. 知识库模型
### api_document
MySQL 中的文档元数据表。
职责:
- 管理上传文件。
- 保存 file hash,用于去重。
- 记录索引状态和 chunk 数量。
- 保存 frontmatter JSON。
### knowledge_domain
领域级元数据。
职责:
- 按 category 聚合文档。
- 存储领域描述。
- 存储 `when_to_retrieve`,辅助 Planner/Executor 判断什么时候检索该领域。
### Milvus/Zilliz metadata
向量 collection 中每个 chunk 的 metadata 主要包括:
```text
docId
_source
chunkIndex
totalChunks
title
breadcrumb
category
```
这些字段支撑:
- category filter。
- source 展示。
- breadcrumb 上下文。
- docId 删除和重建索引。
- evidence block 构造。
## 5. 反馈沉淀模型
### case_library
`useful` 反馈会触发 `CaseLibraryService.createFromSession`。
当前自动映射:
| 字段 | 来源 |
|---|---|
| `case_id` | UUID |
| `diagnosis_id` | `diagnosis_session.session_id` |
| `source_type` | `AUTO` |
| `fault_category` | 当前默认 `GENERAL` |
| `title` | session query 前 100 字符 |
| `root_cause` | session answer |
| `solution` | session answer |
| `created_by` | `system` |
## 6. self_evaluation 结构
`diagnosis_session.self_evaluation` 是 JSON 容器:
```json
{
"rule_evaluation": {},
"verifier_evaluation": {},
"aiops_rule_evaluation": {}
}
```
边界:
- `rule_evaluation` 评估证据收集充分度。
- `verifier_evaluation` 评估 Chat 答案关键事实是否有证据支撑。
- `aiops_rule_evaluation` 评估 AIOps 报告是否聚焦告警并使用证据。
## 7. 数据写入时序
```mermaid
sequenceDiagram
autonumber
participant API as API
participant Svc as ChatService/AiOpsService
participant Session as diagnosis_session
participant Agent as Agent
participant Step as agent_step
participant Tool as tool_invocation
participant Eval as self_evaluation
participant Feedback as case_library
API->>Svc: request
Svc->>Session: create/update RUNNING
Agent->>Step: before/after model
Agent->>Tool: tool call record
Svc->>Session: SUCCESS/FAILED + answer
Svc->>Eval: merge evaluation
API->>Svc: feedback useful
Svc->>Feedback: create case
```
## 8. 当前边界和后续
当前边界:
- `agent_step.session_id` 和 `tool_invocation.session_id` 通过 sessionId 关联,不强制外键。
- `tool_invocation.step_id` 可为空。
- Milvus chunk 与 `api_document` 通过 metadata.docId 逻辑关联。
- `case_library` 与 session 通过 `diagnosis_id=session_id` 关联。
后续可增强:
1. 增加 run id,支持同 session 多次独立诊断。
2. 强化 `tool_invocation.step_id` 关联。
3. 将 evidence block 结构化保存。
4. 将 `case_library` 的 rootCause/solution 从完整 answer 中结构化抽取。
+179
View File
@@ -0,0 +1,179 @@
# Agent 架构演进路线
**更新日期**:2026-07-05
**状态**:后续演进设计,不代表当前已实现
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
## 1. 为什么需要演进路线
旧版 `agent-architecture.md` 包含很多生产级设想:专科 SubAgent、Skill 体系、进程隔离、回退路由、MCP 工具协议化、进化引擎。它们不应作为当前 MVP 事实写入主架构,但可以作为后续扩展路线。
当前原则:
- 当前文档只声明已经可运行或明确落地的能力。
- 演进路线记录未来方向和触发条件。
- 每个演进项必须有可验证收益,不能只因为“架构更炫”就拆。
## 2. 演进总图
```mermaid
flowchart TD
MVP["Current MVP: Planner + Executor + Verifier"] --> Split{"Executor 是否过载?"}
Split -->|是| SubAgents["专科 SubAgent"]
Split -->|否| Keep["继续强化通用 Executor"]
SubAgents --> Skills["Skill / Playbook 体系"]
Skills --> Fallback["回退路由"]
Fallback --> Isolation["进程或 Pod 隔离"]
MVP --> ToolGrowth{"工具数量和来源是否增长?"}
ToolGrowth -->|是| MCP["MCP / Tool Server 协议化"]
ToolGrowth -->|否| ToolCallbacks["继续使用 @Tool / ToolCallback"]
MVP --> EvalGrowth{"评测数据是否足够?"}
EvalGrowth -->|是| Evolution["Prompt / Skill 进化引擎"]
EvalGrowth -->|否| Baseline["先扩大 baseline"]
```
## 3. 专科 SubAgent
### 触发条件
- Executor prompt 变得臃肿,难以同时覆盖接口、数据库、缓存、网络等场景。
- 不同故障类型需要明显不同的工具权限。
- Trace 显示某些场景经常走错排查路径。
- 评测集已经能衡量拆分前后的收益。
### 候选 SubAgent
| SubAgent | 场景 | 工具倾向 |
|---|---|---|
| `ExternalApiSubAgent` | 错误码、接口参数、第三方调用失败 | `lookup_knowledge`, logs, trace |
| `DatabaseSubAgent` | 连接池、慢 SQL、死锁、数据库不可用 | metrics, logs, knowledge |
| `CacheSubAgent` | Redis 超时、热点 key、内存风险 | metrics, logs, knowledge |
| `GenericDiagnosisSubAgent` | 兜底诊断 | 全量只读证据工具 |
### 不立即拆分的原因
- 当前 MVP 的工具规模还可由通用 Executor 管理。
- 过早拆分会增加 Prompt、评测和 trace 分析成本。
- 没有足够分类评测前,拆分可能只是移动复杂度。
## 4. Skill / Playbook 体系
旧版设计中的 Skill 可以在当前项目中演进为可版本化的诊断 Playbook。
```text
fault_category
-> playbook
-> required evidence
-> tool sequence
-> stop condition
-> report template
-> evaluation checks
```
优先落地方向:
- AIOps 告警处理 Playbook。
- 支付超时 Playbook。
- MySQL 连接池风险 Playbook。
- Redis timeout Playbook。
落地前提:
- 每个 Playbook 至少有 3-5 个 eval case。
- Playbook 失败时可以回退到通用 Executor。
- Trace 中能标记使用了哪个 Playbook 和哪个版本。
## 5. 回退路由
当前 Chat 已有低置信补证据和 REJECT 降级输出。后续如果引入 SubAgent,可扩展为:
```text
Specialized SubAgent
-> failed / low confidence
-> another specialized SubAgent
-> GenericDiagnosisSubAgent
-> degraded answer with confirmed facts only
```
回退依据:
- 工具连续失败。
- Verifier `REJECT`。
- Verifier `LOW_CONFID` 且补证据失败。
- Agent 输出缺失关键报告字段。
## 6. 进程隔离
当前所有 Agent 在同一 JVM 内运行。生产级隔离可以考虑:
```text
API service
-> Supervisor service
-> Planner service
-> SubAgent services
-> Verifier service
```
触发条件:
- 某类 Agent 需要独立扩缩容。
- 某类工具依赖不稳定,可能拖垮主应用。
- 不同 Agent 需要不同权限和网络访问策略。
- 单 JVM 内资源隔离不足。
MVP 阶段暂不拆分进程,优先保证 trace、评测和工具边界清晰。
## 7. MCP / Tool Server 协议化
当前工具主要通过 `@Tool`、`methodTools` 和 `ToolCallbackProvider` 暴露。工具数量增加后,可演进为:
```text
Agent
-> Tool registry
-> MCP / tool server
-> log server
-> metrics server
-> knowledge server
-> ticket/change server
```
收益:
- 工具独立部署。
- 新工具上线不必重发主应用。
- 不同 Agent 可获得不同工具子集。
- 工具调用协议统一,更利于审计。
风险:
- 调用链更长。
- 权限和超时治理更复杂。
- 本地开发和 Demo 成本上升。
## 8. 进化引擎
旧版文档提到从诊断中学习。当前可以拆成更务实的步骤:
1. 先扩大 diagnosis eval 和 RAG eval。
2. 从失败 trace 中标注 bad case。
3. 将高频失败沉淀为 Playbook 或 Prompt 规则。
4. 对 Prompt 版本做离线对比。
5. 足够稳定后再考虑线上 A/B。
不建议 MVP 直接做自动 Prompt 自优化。没有可靠评测和回滚机制时,自动优化更容易引入不可解释变化。
## 9. 演进优先级
| 优先级 | 项目 | 原因 |
|---|---|---|
| P0 | 扩大 eval baseline | 没有评测,拆任何架构都难以证明收益 |
| P1 | Playbook 化高频故障 | 可控、可解释、比拆 SubAgent 更轻 |
| P1 | 完整 evidence block | 提升 Verifier 和 Trace 质量 |
| P2 | 专科 SubAgent | 等问题类型和工具权限差异足够明显 |
| P2 | AIOps LLM Verifier | 规则门禁不足时再引入 |
| P3 | MCP 工具协议化 | 工具来源复杂后再做 |
| P3 | 进程隔离 | 生产负载和权限隔离需要明确后再做 |
+251
View File
@@ -0,0 +1,251 @@
# 反馈与自评估架构
**更新日期**:2026-07-05
**状态**:当前可运行架构
**参考历史文档**:`archive/2026-07-05-legacy/confidence-feedback.md`
## 1. 定位
反馈架构包含两条闭环:
1. 系统自评估:基于工具调用、Verifier、AIOps 规则检查,写入 `diagnosis_session.self_evaluation`。
2. 用户反馈:用户标记 `useful` 或 `not_useful`,写入 `diagnosis_session.feedback`,其中 `useful` 会沉淀案例。
当前重要边界:
- `status` 表示执行状态,不表示答案质量。
- `feedback` 表示用户反馈,不覆盖 `status`。
- `self_evaluation` 是 JSON 容器,内部按来源分层,不再把所有评分字段平铺在根节点。
## 2. 总体闭环
```mermaid
flowchart TD
Answer["Chat / AIOps final answer"] --> Session["diagnosis_session.answer"]
subgraph SelfEval["Self evaluation"]
Invocation["tool_invocation"] --> RuleEval["EvaluationService: rule_evaluation"]
Invocation --> TraceSummary["ToolTraceSummaryService"]
TraceSummary --> Verifier["chat_verifier"]
Verifier --> VerifierEval["verifier_evaluation"]
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
AiOpsRule --> AiOpsEval["aiops_rule_evaluation"]
end
RuleEval --> Merge["SelfEvaluationMergeService"]
VerifierEval --> Merge
AiOpsEval --> Merge
Merge --> SelfJson["diagnosis_session.self_evaluation"]
subgraph UserFeedback["User feedback"]
UI["Feedback bar"] --> API["POST /api/feedback"]
API --> FeedbackService["FeedbackService"]
FeedbackService --> FeedbackField["diagnosis_session.feedback"]
FeedbackService --> Useful{"feedback == useful?"}
Useful -->|yes| CaseService["CaseLibraryService.createFromSession"]
CaseService --> Case["case_library"]
Useful -->|no| BadCase["Bad case by feedback=not_useful"]
end
Session --> UI
```
## 3. self_evaluation JSON
`SelfEvaluationMergeService` 统一维护 `diagnosis_session.self_evaluation`。
当前结构:
```json
{
"rule_evaluation": {
"evidence_score": 65,
"source": "rule",
"factors": []
},
"verifier_evaluation": {
"verdict": "PASS",
"groundedness_score": 0.8,
"critical_fact_count": 2,
"facts_checked": [],
"rationale": "...",
"round": 1,
"traceability_version": "v1",
"tool_trace_summary": []
},
"aiops_rule_evaluation": {
"verdict": "...",
"checks": []
}
}
```
兼容逻辑:
- 如果旧 JSON 根节点包含 `evidence_score`,会被包进 `rule_evaluation`。
- 如果旧 JSON 根节点包含 `verdict` / `groundedness_score`,会被包进 `verifier_evaluation`。
## 4. 规则评分
`EvaluationService` 只消费 `tool_invocation` 和 session 状态,输出 `rule_evaluation`。
定位:
- 衡量证据收集充分度。
- 不直接证明答案是否推理正确。
- 不依赖 LLM。
规则:
| 规则名 | 条件 | 分数变化 |
|---|---|---|
| `execution_failed` | session status = `FAILED` | 直接 0 |
| `no_tool_call` | 没有工具调用 | 直接 0 |
| `has_successful_tool_call` | 至少一次工具成功 | +30 |
| `l0_exact_match` | 任意工具调用有 L0 命中 | +35 |
| `l1_semantic_match` | 无 L0 命中但有 L1 命中 | +20 |
| `retrieval_no_hit` | 有检索调用但无命中 | -10 |
| `all_tool_calls_failed` | 工具全部失败 | -20 |
最终分数裁剪到 `[0, 100]`。
说明:
- 当前 `rule_evaluation` 是异步写入,失败时 `self_evaluation` 可能暂时为空或缺少该节点。
- L0/L1 分支互斥:有 L0 命中时优先记 L0。
- 更强的答案真实性校验由 Chat Verifier 承担。
## 5. Chat Verifier 自评估
Chat Verifier 校验 Executor 的最终答案是否被证据支撑。
```mermaid
flowchart LR
Answer["executor_final_answer"] --> Verifier["chat_verifier"]
Invocation["tool_invocation"] --> Summary["ToolTraceSummaryService"]
Summary --> Evidence["tool_trace_summary"]
Evidence --> Verifier
Verifier --> Output["verifier_output JSON"]
Output --> Merge["SelfEvaluationMergeService.mergeVerifierEvaluation"]
Merge --> Session["diagnosis_session.self_evaluation.verifier_evaluation"]
```
Verifier 输出:
| 字段 | 说明 |
|---|---|
| `verdict` | `PASS` / `LOW_CONFID` / `REJECT` |
| `groundedness_score` | 关键事实证据支撑度 |
| `critical_fact_count` | 关键事实数量 |
| `facts_checked` | 逐条事实校验 |
| `rationale` | 判定原因 |
| `tool_trace_summary` | 本次校验使用的证据索引 |
ChatService 根据 verdict 决定:
- `PASS`:输出 Executor 答案。
- `LOW_CONFID`:必要时构造 `retry_context` 补证据;否则输出低置信提示。
- `REJECT`:降级输出,只保留已确认信息。
## 6. AIOps 规则自评估
AIOps 当前使用 `AiOpsRuleEvaluationService`,结果写入 `aiops_rule_evaluation`。
检查重点:
- 是否有最终报告。
- payload 模式是否聚焦输入告警。
- 是否调用 `lookup_knowledge`、日志、指标等证据工具。
- 是否把无关活跃告警扩展成主诊断对象。
这是轻量规则检查,不等价于完整 LLM Verifier。完整 AIOps Verifier 是后续增强项。
## 7. 用户反馈 API
```text
POST /api/feedback
Content-Type: application/json
{
"sessionId": "xxx",
"feedback": "useful" | "not_useful"
}
```
响应:
```json
{
"success": true,
"message": "反馈已记录",
"caseId": "uuid 或 null"
}
```
后端行为:
| feedback | 行为 |
|---|---|
| `useful` | 写入 `DiagnosisSession.feedback`,调用 `CaseLibraryService.createFromSession` |
| `not_useful` | 写入 `DiagnosisSession.feedback`,不改变 session status |
| 其他值 | 返回 HTTP 400 |
## 8. 案例沉淀
`useful` 反馈会生成或复用 `case_library` 记录。
字段映射:
| CaseLibrary 字段 | 来源 |
|---|---|
| `caseId` | UUID |
| `diagnosisId` | `DiagnosisSession.sessionId` |
| `sourceType` | `AUTO` |
| `faultCategory` | 当前固定为 `GENERAL` |
| `title` | `query` 前 100 字符 |
| `rootCause` | `answer` |
| `solution` | `answer` |
| `createdBy` | `system` |
幂等性:
```text
case_library.diagnosisId == sessionId
-> existing case: return existing
-> missing case: create new
```
## 9. Trace 呈现
Trace API 会展示:
- `feedback`
- `hasFeedback`
- `hasVerifierEvaluation`
- `hasAiOpsRuleEvaluation`
- session、step、tool invocation 明细
这让一次诊断可以被分成三种视角查看:
| 视角 | 数据来源 |
|---|---|
| 执行是否成功 | `diagnosis_session.status` |
| 证据是否充分 | `self_evaluation.rule_evaluation` / `verifier_evaluation` |
| 用户是否认可 | `diagnosis_session.feedback` |
## 10. 后续增强
近期优先:
1. 将 `rule_evaluation` 与 `verifier_evaluation` 在 Trace API 中结构化展示。
2. `not_useful` 反馈沉淀 bad case,而不是只写字段。
3. useful 案例自动提取 faultCategory、errorCode、service、rootCause、solution。
4. AIOps 引入 LLM Verifier。
5. 把反馈和 eval baseline 打通,形成可回归的质量改进闭环。
暂不优先:
- 用用户反馈直接修改 session status。
- 仅凭 `evidence_score` 判断答案正确。
- 在没有人工审核时自动把 bad case 反向写入 Prompt。
+205
View File
@@ -0,0 +1,205 @@
# Harness 与质量门禁架构
**更新日期**:2026-07-05
**状态**:当前可运行架构 + 后续门禁规划
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
## 1. 设计目标
Agent 系统的核心风险不是“没有答案”,而是:
- 答案引用了不存在的证据。
- 工具调用失败后仍然编造结论。
- 检索结果相关性不足但被当作强证据。
- 多轮诊断重复检索同一文档,浪费上下文。
- 最终报告无法回放执行过程。
因此当前 MVP 的 Harness 不是单个组件,而是一组约束:
```text
Prompt contract
+ Tool boundary
+ Agent hooks
+ Trace persistence
+ Verifier / rule evaluation
+ Eval baseline
```
## 2. Harness 总图
```mermaid
flowchart TB
Input["User / AIOps input"] --> Prompt["Prompt contract"]
Prompt --> Agent["Planner / Executor / Verifier"]
Agent --> Tools["Evidence tools"]
Tools --> Invocation["tool_invocation"]
Agent --> StepHook["AgentLoggingHook"]
StepHook --> Step["agent_step"]
Agent --> Session["diagnosis_session"]
Invocation --> TraceSummary["ToolTraceSummaryService"]
TraceSummary --> Verifier["chat_verifier"]
Verifier --> SelfEval["self_evaluation.verifier_evaluation"]
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
AiOpsRule --> AiOpsEval["self_evaluation.aiops_rule_evaluation"]
Session --> TraceAPI["DiagnosisTraceService"]
Step --> TraceAPI
Invocation --> TraceAPI
SelfEval --> TraceAPI
AiOpsEval --> TraceAPI
TraceAPI --> Eval["diagnosis eval / RAG eval"]
```
## 3. Prompt Contract
当前 Prompt 按角色拆分:
| Prompt | 用途 |
|---|---|
| `supervisor-prompt.md` | AIOps Supervisor 调度 Planner / Executor |
| `planner-prompt.md` | AIOps Planner 规划、再规划、输出告警报告 |
| `executor-prompt.md` | AIOps Executor 按步骤调用工具 |
| `chat-planner-prompt.md` | Chat 复杂问题规划 |
| `chat-executor-prompt.md` | Chat 执行工具并形成诊断答复 |
| `chat-verifier-prompt.md` | 校验 Executor 答案是否被工具证据支撑 |
Prompt 层当前承担的门禁:
- 禁止凭记忆回答错误码、接口定义、排障步骤。
- 需要外部信息时必须调用工具。
- 工具连续失败或返回空结果时,最终报告必须诚实说明。
- Chat Verifier 不允许做新检索,只能校验已有证据。
- AIOps payload 模式必须聚焦输入告警。
## 4. Trace Hooks
`AgentLoggingHook` 是当前 Agent step 可观测性的核心。
```mermaid
sequenceDiagram
autonumber
participant A as Agent
participant H as AgentLoggingHook
participant DB as agent_step
A->>H: before_model(messages, sessionId)
H->>DB: 写入 model_input / step_index / agent_name
A-->>A: LLM 推理
A->>H: after_model(messages, sessionId)
H->>DB: 回填 model_output / thought / has_tool_call / duration / token_count
```
记录内容:
- 最近输入消息摘要。
- Agent 输出摘要。
- 是否包含 tool call。
- duration。
- token count。
- Verifier 的 JSON 输出摘要。
## 5. Tool Invocation 门禁
工具调用记录由 `ToolInvocationRecorder` 和具体工具共同完成。
核心记录:
```text
tool_name
input_params
output_preview
retrieval_layer
l0_match_count
l1_match_count
retrieval_details
relevance_level
dedup_reason
duration_ms
success
error_message
```
对 `lookup_knowledge` 的质量约束:
- L0 只作为 hint,不绕过 L1。
- 检索结果归一化为 `PRECISE`、`HIGHLY_RELEVANT`、`REFERENCE`。
- 同 session 内重复文档会被 `RetrievedDocTracker` 去重。
- dedup、no evidence、failed 等状态进入 `retrieval_details.evidence_status`。
## 6. Verifier 门禁
Chat Verifier 的输入不是原始工具日志,而是 `ToolTraceSummaryService` 构造的证据索引。
```mermaid
flowchart LR
Invocation["tool_invocation"] --> Summary["ToolTraceSummaryService"]
Summary --> EvidenceIndex["tool_trace_summary"]
EvidenceIndex --> Verifier["chat_verifier"]
ExecutorAnswer["executor_final_answer"] --> Verifier
Verifier --> Verdict{"verdict"}
Verdict -->|PASS| Pass["输出原答案"]
Verdict -->|LOW_CONFID| Low["补证据或低置信输出"]
Verdict -->|REJECT| Reject["降级输出"]
```
Verifier 输出:
```json
{
"verdict": "PASS|LOW_CONFID|REJECT",
"groundedness_score": 0.8,
"critical_fact_count": 2,
"facts_checked": [],
"rationale": "..."
}
```
结果写入:
```text
diagnosis_session.self_evaluation.verifier_evaluation
```
## 7. AIOps 规则门禁
AIOps 当前不走 Chat Verifier,而是用 `AiOpsRuleEvaluationService` 做轻量检查。
检查重点:
- 最终报告是否存在。
- payload 模式是否围绕输入告警展开。
- 是否调用证据工具,尤其是 `lookup_knowledge`、日志、指标。
- 是否把无关活跃告警扩展成主诊断对象。
结果写入:
```text
diagnosis_session.self_evaluation.aiops_rule_evaluation
```
## 8. Eval Baseline
当前质量门禁还包括离线评测资产:
| 评测 | 位置 | 作用 |
|---|---|---|
| Diagnosis eval | `mvp/eval/` | 检查诊断 trace、报告和证据行为 |
| RAG retrieval eval | `eval/rag-retrieval/` | 检查固定检索 query 的召回稳定性 |
| Live RAG acceptance | `scripts/eval_rag_live_acceptance.py` | 检查运行环境中真实 `/api/search/similar` 行为 |
## 9. 后续门禁规划
从旧版设计继承但尚未完整实现的门禁:
- 工具参数 schema 校验。
- 同一工具调用次数上限。
- 工具超时的统一熔断。
- 报告中的数值与工具返回值自动对齐校验。
- Prompt 版本记录和回滚。
- Verifier 对 AIOps 报告的 LLM 级事实校验。
这些应在评测集扩大后逐步加入,避免一次性把诊断流程卡得过死。
+113
View File
@@ -0,0 +1,113 @@
# 面试一页式架构讲解
**用途**:面试现场 2-5 分钟讲清项目
**适合场景**:开场介绍、架构追问、Demo 前铺垫
## 1. 一句话
SuperBizAgent 是一个面向企业故障诊断的可追踪 Agent 系统:它把用户问题或 AIOps 告警转换成 Planner、Executor、Verifier 的诊断链路,所有工具证据、模型步骤、最终答案、自评估和用户反馈都能通过同一个 `sessionId` 回放。
## 2. 一张图
```mermaid
flowchart TB
User["用户问题 / AIOps 告警"] --> API["API Layer"]
API --> Chat["ChatService"]
API --> AiOps["AiOpsService"]
Chat --> ChatFlow["Chat: Planner -> Executor -> Verifier"]
AiOps --> AiOpsFlow["AIOps: Supervisor -> Planner / Executor"]
ChatFlow --> Tools["Evidence Tools"]
AiOpsFlow --> Tools
Tools --> Knowledge["lookup_knowledge"]
Tools --> Logs["query_logs"]
Tools --> Metrics["query_metrics / Prometheus"]
Knowledge --> RAG["RAG: L0 hint + VectorSearchService"]
RAG --> VectorStore["Spring AI VectorStore"]
RAG --> SDK["Milvus SDK fallback"]
ChatFlow --> Trace["Trace Persistence"]
AiOpsFlow --> Trace
Tools --> Trace
Trace --> Session["diagnosis_session"]
Trace --> Step["agent_step"]
Trace --> Invocation["tool_invocation"]
Invocation --> Verifier["Verifier / Rule Evaluation"]
Verifier --> SelfEval["self_evaluation"]
Session --> TraceAPI["GET /api/diagnosis/{sessionId}/trace"]
Step --> TraceAPI
Invocation --> TraceAPI
SelfEval --> TraceAPI
TraceAPI --> Feedback["POST /api/feedback"]
Feedback --> Case["useful -> case_library"]
```
## 3. 面试讲法
```text
这个项目不是把问题直接丢给大模型,而是把诊断拆成可审计的执行链路。
Chat 复杂问题走 Planner -> Executor -> Verifier:
Planner 负责拆解,Executor 负责调用知识库、日志和指标工具,Verifier 只基于已有工具证据校验最终答案。
AIOps 告警入口走 Supervisor 调度 Planner/Executor:
如果请求里有 alert payload,系统会进入 PAYLOAD_TARGETED 模式,报告必须聚焦这个告警,而不是被当前环境中的其他活跃告警带偏。
所有过程都会落到 diagnosis_session、agent_step、tool_invocation。
所以我可以用一个 sessionId 回放:模型怎么规划、调了哪些工具、工具返回什么、Verifier 怎么判定、用户最后是否反馈有用。
```
## 4. 五个亮点
| 亮点 | 怎么讲 |
|---|---|
| 可追踪 Agent | 每次诊断都有 `sessionId`,Trace API 可以回放 session、step、tool |
| 显式工具证据链 | `lookup_knowledge`、日志、指标都记录到 `tool_invocation` |
| RAG 工程化 | L0 降级为 hint,Spring AI VectorStore 做主检索,SDK fallback 保底 |
| 质量门禁 | Chat Verifier 校验 groundedness,AIOps rule evaluation 控制告警聚焦 |
| 反馈闭环 | useful 反馈沉淀 `case_library`,not_useful 保留 bad case 信号 |
## 5. 三个关键取舍
### 取舍 1:为什么不用隐式 Advisor 做 RAG?
因为这个项目强调 Agent 决策可见性。`lookup_knowledge` 必须作为显式工具调用被记录,这样才能解释“什么时候检索、检索了什么、证据如何支撑结论”。
### 取舍 2:为什么保留 Milvus SDK fallback?
因为迁移到 Spring AI VectorStore 期间,schema、collection、score 语义都可能变化。`auto` 模式先走 VectorStore,失败时 fallback 到 SDK,保证 MVP 主链路可运行,也方便对比新旧检索质量。
### 取舍 3:为什么 self_evaluation 分三层?
因为三类评估回答的问题不同:
```text
rule_evaluation -> 工具证据是否充分
verifier_evaluation -> Chat 答案关键事实是否有证据支撑
aiops_rule_evaluation -> AIOps 报告是否聚焦告警并使用证据
```
## 6. 面试官可能追问
| 追问 | 回答方向 |
|---|---|
| 怎么防止幻觉? | Executor 必须用工具;Verifier 只基于 `tool_trace_summary` 校验;LOW_CONFID/REJECT 会降级输出 |
| RAG 质量怎么保证? | offline golden cases + live acceptance + trace inspection 三层验证 |
| 为什么 L0 不直接返回? | L0 子串命中不等于语义相关,当前只做 domain/entity hint 和 metadata filter |
| AIOps 如何避免跑偏? | payload 模式生成 recommended query,并用 rule evaluation 检查报告聚焦输入告警 |
| 下一步怎么演进? | evidence block、邻居 chunk、Playbook、AIOps LLM Verifier、MCP 工具协议化 |
## 7. 现场演示入口
- Demo 脚本:`mvp/demo/ten-minute-interview-demo.md`
- 故事案例:`interview/story-cases.md`
- 架构细节:`mvp/architecture/README.md`
@@ -0,0 +1,240 @@
# 知识库文档编写与维护
**更新日期**:2026-07-05
**状态**:当前建议规范
**参考历史文档**:`archive/2026-07-05-legacy/knowledge-retrieval-usage.md`
## 1. 定位
知识库文档不是普通 Markdown 资料堆叠,而是 RAG 检索的输入资产。写得好的文档会提升:
- L0 hint 的关键词和领域识别。
- L1 向量召回质量。
- `breadcrumb` 上下文恢复能力。
- Verifier 可引用的证据质量。
当前推荐写法:结构化 Markdown + frontmatter + 明确分类 + 可检索关键词。
## 2. 文档进入系统的链路
```mermaid
flowchart TD
Markdown["Markdown file"] --> Upload["POST /api/documents/upload"]
Upload --> Parse["FrontmatterParser"]
Parse --> Enrich["DocumentFieldEnricher"]
Enrich --> Metadata["api_document.metadata"]
Upload --> Chunk["DocumentChunkService"]
Chunk --> Breadcrumb["title / breadcrumb / chunkIndex"]
Breadcrumb --> Embedding["VectorIndexService embedding text"]
Embedding --> Milvus["Milvus/Zilliz"]
Metadata --> L0["KnowledgeIndexService L0 index"]
Milvus --> L1["VectorSearchService L1 retrieval"]
```
## 3. Frontmatter
推荐模板:
```markdown
---
title: 支付网关错误码定义
keywords: [ERR_TIMEOUT, 支付超时, payment timeout, 支付网关]
summary: 记录支付网关核心错误码的含义、常见原因和排查步骤
category: api
version: 1.0
author: sre-team
---
# 支付网关错误码定义
...
```
字段说明:
| 字段 | 必填 | 用途 |
|---|---:|---|
| `title` | 是 | 文档标题,进入 L0 索引和 embedding 上下文 |
| `keywords` | 是 | L0 hint 的主要来源 |
| `summary` | 是 | 文档摘要,进入知识域描述和 Agent 上下文 |
| `category` | 建议 | 知识域、metadata filter、上传目录 |
| `version` | 可选 | 文档版本 |
| `author` | 可选 | 维护人 |
当前解析器会提示缺少 `title`、`keywords`、`summary` 的情况;缺失不一定阻断上传,但会降低检索质量。
## 4. category 建议
`category` 会影响:
- 上传文件本地目录。
- Milvus metadata。
- L0 domain hint。
- `knowledge_domain` 聚合。
- VectorStore / SDK category filter。
推荐保持稳定,不要频繁换名。
| category | 用途 |
|---|---|
| `api` | 接口、错误码、请求/响应协议 |
| `infrastructure` | MySQL、Redis、JVM、网络、中间件 |
| `troubleshooting` | 通用排障流程、Runbook |
| `domain` | 业务领域规则 |
| `spring-ai` | Spring AI / Agent / 工具最佳实践 |
注意:分类过细会导致 filter 召回不足;分类过粗会降低 L0 hint 解释力。
## 5. 关键词写法
好的关键词应该覆盖:
- 精确实体:错误码、服务名、指标名。
- 常用中文说法。
- 英文别名。
- 组合词。
示例:
```yaml
keywords: [ERR_TIMEOUT, timeout, 支付超时, 支付网关超时, payment-service, gateway timeout]
```
避免:
```yaml
keywords: [错误, 问题, 系统]
```
原因:过宽关键词会让 L0 hint 变脏,多个文档同时命中,影响解释性和 category filter。
## 6. Markdown 结构
推荐结构:
```markdown
# 文档总标题
## 场景或错误码
### 含义
### 常见原因
### 排查步骤
### 处理方案
### 日志示例
```
为什么这样写:
- `DocumentChunkService` 会按 Markdown 标题切分。
- 标题层级会生成 `breadcrumb`。
- `title + breadcrumb + content` 会一起进入 embedding 文本。
- 命中 chunk 时,Agent 更容易知道证据属于哪个章节。
## 7. 内容建议
每个可诊断条目尽量包含:
- 现象。
- 判断条件。
- 可能原因。
- 证据来源。
- 排查步骤。
- 处理建议。
- 日志或配置示例。
示例:
```markdown
## ERR_TIMEOUT
### 含义
支付网关请求超过本地或上游超时时间。
### 常见原因
1. 第三方支付服务响应慢。
2. 本地 timeout 配置过短。
3. 网络链路抖动。
### 排查步骤
1. 查询 payment-service 日志中的请求耗时。
2. 查看网关 5xx 和 timeout 指标。
3. 对比当前 timeout 配置。
### 处理建议
- 短期:重试受影响订单。
- 长期:调整 timeout 和重试策略,并监控上游延迟。
```
## 8. 上传与索引
上传接口:
```text
POST /api/documents/upload
Content-Type: multipart/form-data
file=<markdown file>
category=<category>
```
系统处理:
1. 计算文件 hash,避免重复上传。
2. 保存原始文件。
3. 解析 frontmatter。
4. 补全文档字段。
5. 写入 `api_document`。
6. Markdown-aware chunking。
7. 写入 Milvus/Zilliz。
8. 更新 L0 索引和 `knowledge_domain`。
## 9. 重建索引注意事项
当以下内容变化时,需要重新索引:
- 正文内容。
- 标题层级。
- `category`。
- `title`、`summary`、`keywords`。
- embedding 输入策略,例如加入 `breadcrumb`。
特别注意:
```text
修改 Markdown 或 embedding 输入策略,不会自动改变已有向量。
必须重新上传或重建索引后,live retrieval 才能体现变化。
```
可用 live 验收:
```bash
python scripts/eval_rag_live_acceptance.py
```
## 10. 维护 checklist
新增文档前检查:
- frontmatter 是否包含 `title`、`keywords`、`summary`。
- `category` 是否属于现有稳定分类。
- 关键词是否既有精确词也有常用表达。
- Markdown 标题层级是否清晰。
- 每个故障条目是否包含可执行排查步骤。
- 日志/配置示例是否脱敏。
更新文档后检查:
- `api_document.status` 是否为 `INDEXED`。
- `/api/search/similar` 是否能搜到目标文档。
- `eval/rag-retrieval` 是否需要新增 golden case。
- Trace 中 `tool_invocation` 是否记录到正确 source 和 breadcrumb。
+414
View File
@@ -0,0 +1,414 @@
# RAG 新架构
**更新日期**:2026-07-05
**状态**:当前主架构 + 后续演进边界
**关联计划**:`mvp/issues/rag-refactor-plan.md`
## 1. 架构目标
RAG 重构的目标不是把所有能力交给框架,也不是继续维护一套完全自研检索框架,而是形成:
```text
成熟框架能力 + 业务可观测编排
```
具体原则:
- 通用向量检索能力交给 Spring AI `VectorStore`。
- 项目保留 Agent Tool 入口、AIOps 业务 query 映射、证据打包、trace 记录。
- `lookup_knowledge` 继续是显式工具,不替换成隐式 Advisor。
- Spring AI 读取路径作为主路径,Milvus SDK 作为 fallback。
- 所有检索行为必须可评测、可回放、可解释。
## 2. 当前主链路
```mermaid
flowchart TD
Agent["Agent Executor"] --> Tool["lookup_knowledge(query)"]
Tool --> L0["KnowledgeIndexService.analyzeQuery"]
L0 --> Hint["L0 hint: domain / entities / matchedKeywords"]
Hint --> Filter["category filter candidate"]
Tool --> Search["VectorSearchService.searchSimilarDocuments"]
Filter --> Search
Search --> Mode{"retrieval.vector-store.mode"}
Mode -->|auto| SpringTry["try Spring AI VectorStore"]
SpringTry -->|success| Results["SearchResult list"]
SpringTry -->|failure| SdkFallback["Milvus SDK fallback"]
Mode -->|spring-ai| SpringOnly["Spring AI VectorStore only"]
Mode -->|sdk| SdkOnly["Milvus SDK only"]
SpringOnly --> Results
SdkFallback --> Results
SdkOnly --> Results
Results --> Normalize["relevance normalization"]
Normalize --> Dedup["session dedup: RetrievedDocTracker"]
Dedup --> Output["LookupResult"]
Output --> Record["tool_invocation record"]
Output --> Agent
```
```text
Agent Executor
-> lookup_knowledge(query)
-> KnowledgeIndexService.analyzeQuery
-> L0 domain/entity hint
-> matchedKeywords
-> category filter candidate
-> VectorSearchService.searchSimilarDocuments
-> mode=auto
-> Spring AI VectorStore
-> fallback: Milvus SDK
-> mode=spring-ai
-> Spring AI VectorStore only
-> mode=sdk
-> Milvus SDK only
-> result normalization
-> relevanceLevel
-> completenessHint
-> score/rawScore/scoreLabel
-> session dedup
-> RetrievedDocTracker
-> tool_invocation record
```
运行配置:
```properties
retrieval.vector-store.mode=auto
retrieval.normalization.max-l2-distance=2.0
retrieval.normalization.highly-relevant-threshold=0.75
retrieval.normalization.reference-threshold=0.5
```
## 3. 稳定边界
```mermaid
flowchart LR
subgraph AgentBoundary["Agent boundary"]
Executor["Executor Agent"]
Tool["LookupKnowledgeTool"]
end
subgraph RetrievalBoundary["Retrieval boundary"]
Search["VectorSearchService"]
Spring["Spring AI VectorStore"]
SDK["Milvus SDK"]
end
subgraph ObservabilityBoundary["Observability boundary"]
Invocation["tool_invocation"]
Eval["RAG baseline / trace inspection"]
end
Executor --> Tool
Tool --> Search
Search --> Spring
Search --> SDK
Tool --> Invocation
Invocation --> Eval
```
### 3.1 Agent 边界
Agent 只知道自己可以调用 `lookup_knowledge`,不直接关心底层是 Spring AI VectorStore 还是 Milvus SDK。
```text
Executor -> LookupKnowledgeTool -> VectorSearchService
```
这个边界让 RAG 底层迁移不影响 Agent prompt、工具声明和 trace 数据结构。
### 3.2 检索边界
`VectorSearchService` 是当前检索门面:
- `auto`:优先 Spring AI VectorStore,失败后 fallback 到 SDK。
- `spring-ai`:只走 Spring AI VectorStore。
- `sdk`:只走原 Milvus SDK。
这样可以在不改 Agent 工具的情况下切换检索实现,并支持线上验证和回退。
### 3.3 可观测边界
无论底层检索路径如何变化,都必须写入 `tool_invocation`:
```text
sessionId
toolName
inputParams
outputPreview
retrievalLayer
l0MatchCount
l1MatchCount
retrievalDetails
relevanceLevel
dedupReason
duration
success
```
## 4. L0 的新职责
旧版 L0 容易承担过重职责,例如唯一匹配后直接跳过 L1。当前架构中 L0 被降级为 hint 层。
```mermaid
flowchart TD
Input["query / AIOps payload"] --> L0["L0 hint analysis"]
L0 --> Domain["domain detector"]
L0 --> Entity["entity extractor"]
L0 --> Keyword["matched keyword explanation"]
L0 --> Filter["metadata/category filter candidate"]
Domain --> Retrieval["L1 semantic retrieval"]
Entity --> Retrieval
Keyword --> Trace["hit reason in tool_invocation"]
Filter --> Retrieval
Retrieval --> Normalize["relevance normalization"]
Normalize --> Evidence["evidence returned to Agent"]
```
L0 负责:
- domain detector
- entity extractor
- matched keyword explanation
- metadata/category filter candidate
- trace 中的 hit reason
L0 不再默认负责:
```text
L0 unique hit -> 直接作为最终检索结果
```
当前职责是:
```text
query / AIOps payload
-> L0 matched keywords / domains / entities
-> category filter candidate
-> L1 semantic retrieval
-> relevance normalization
```
这样既保留精确关键词和领域 hint 的价值,也避免 L0 误召回直接污染最终证据。
## 5. L1 向量检索
L1 语义检索通过 `VectorSearchService` 调度。
```mermaid
flowchart TD
Search["VectorSearchService"] --> Request["SearchRequest: query / topK / threshold / filter"]
Request --> VectorStore["Spring AI VectorStore"]
VectorStore --> Docs["Document results"]
Docs --> Map["map to SearchResult"]
Map --> Score["score compatibility mapping"]
Search --> SDK["Milvus SDK fallback"]
SDK --> SdkRows["id / content / metadata / L2 distance"]
SdkRows --> Map
Score --> Output["id / content / metadata / score / rawScore / scoreLabel"]
```
### Spring AI VectorStore 路径
```text
SearchRequest
-> query
-> topK
-> similarityThresholdAll
-> optional filterExpression: category == '...'
-> VectorStore.similaritySearch
```
返回结果会映射为项目兼容结构:
```text
id
content
metadata
score
rawScore
scoreLabel
```
### Milvus SDK fallback
SDK 路径仍保留:
- 用于 `auto` 模式兜底。
- 用于与旧链路对比。
- 用于 VectorStore 配置或 collection schema 异常时保证 MVP 可运行。
## 6. 分数语义
旧 SDK 使用 L2 distance,Spring AI 返回 similarity。两者不能混用为同一个含义。
当前统一输出:
| 字段 | 含义 |
|---|---|
| `score` | 兼容旧逻辑的距离型分数,越小越近 |
| `rawScore` | 底层实现的原始分数 |
| `scoreLabel` | `l2_distance` 或 `similarity` |
SDK 路径:
```text
score = L2 distance
rawScore = L2 distance
scoreLabel = l2_distance
```
VectorStore 路径:
```text
rawScore = Spring AI similarity
scoreLabel = similarity
score = metadata.distance if available else compatible distance
```
## 7. 文档切片和 embedding 输入
当前保留 Markdown-aware chunking:
- 识别 Markdown 标题层级。
- 生成 `title`。
- 生成 `breadcrumb`。
- 保留 `chunkIndex`。
- 使用 token 估算和软/硬上限控制 chunk 大小。
- 尽量不打断列表和代码块。
embedding 输入中已经加强:
```text
title + breadcrumb + content
```
这样可以降低单个 chunk 脱离章节上下文后的召回损失。
## 8. AIOps query 增强
AIOps payload 中的业务字段不能完全交给通用检索框架隐式理解。
payload 模式会把以下字段拼成推荐知识库 query:
- `alertName`
- `service`
- `severity`
- `description`
- `timeRange`
- `userRequest`
Prompt 会明确要求 Agent 在需要知识库证据时,优先使用推荐 query 或保留 alertName/service 的更窄 query。
```text
AIOps payload
-> buildKnowledgeRetrievalQuery
-> Recommended lookup_knowledge query
-> lookup_knowledge
-> tool_invocation
```
## 9. Evidence 与去重
当前 evidence 输出仍以 `LookupResult` 和工具返回文本为主,已经具备:
- L0/L1 命中数量。
- 检索层记录。
- relevance level。
- completeness hint。
- session 级文档去重。
- domain 行动记忆。
- `tool_invocation` 明细记录。
后续更完整的 evidence block 目标:
```text
source
docId
chunkIndex
title
breadcrumb
score
rawScore
scoreLabel
hitReason
content
expandedFrom
```
这部分应作为下一阶段增强,而不是当前已完全完成能力。
## 10. 评测与验收
RAG 架构变更必须先过评测,再认为可合入主链路。
当前评测资产:
- `eval/rag-retrieval/cases/golden-cases.json`
- `eval/rag-retrieval/fixtures/`
- `eval/rag-retrieval/reports/baseline.json`
- `eval/rag-retrieval/reports/baseline.md`
- `scripts/eval_rag_retrieval.py`
- `scripts/eval_rag_live_acceptance.py`
评测层次:
| 层次 | 作用 |
|---|---|
| Offline baseline | 不依赖 MySQL、Redis、Milvus、LLM,用固定 fixtures 检查召回行为 |
| Live acceptance | 应用运行并重建索引后,调用 `/api/search/similar` 验证真实检索 |
| Trace inspection | 通过 `tool_invocation` 检查 Agent 是否真的使用了证据 |
## 11. 当前已完成
- `lookup_knowledge` 保持显式 Agent Tool。
- L0 降级为 domain/entity hint。
- L1 默认执行语义检索。
- `VectorSearchService` 支持 `auto`、`spring-ai`、`sdk` 三种模式。
- Spring AI VectorStore 成为读取主路径。
- Milvus SDK fallback 保留。
- 分数语义拆成 `score`、`rawScore`、`scoreLabel`。
- Markdown chunk 保留 `title` 和 `breadcrumb`。
- embedding 输入包含 `title`、`breadcrumb` 和 `content`。
- AIOps payload 生成推荐知识库 query。
- `tool_invocation` 记录 relevance level 和 dedup reason。
- RAG offline baseline 和 live acceptance 脚本已补齐。
## 12. 后续演进
近期优先:
1. 完整 evidence block 结构化输出。
2. 命中 chunk 的相邻 chunk / 同章节上下文扩展。
3. metadata taxonomy 清理,例如 `database` 与 `infrastructure` 的分类边界。
4. Query Transformer / MultiQuery 的可回退接入。
5. VectorStore 写入路径评估。
暂不优先:
- 把 `lookup_knowledge` 替换成隐式 Advisor。
- 完整自研 RRF 框架。
- 立即引入 Elasticsearch / OpenSearch。
- 立即引入 cross-encoder 或 LLM rerank。
## 13. 关键代码索引
| 能力 | 代码 |
|---|---|
| Agent 工具入口 | `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java` |
| L0 hint | `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java` |
| 向量检索门面 | `src/main/java/com/superbiz/agent/service/VectorSearchService.java` |
| 文档切片 | `src/main/java/com/superbiz/agent/service/DocumentChunkService.java` |
| 文档管理 | `src/main/java/com/superbiz/agent/service/DocumentManagementService.java` |
| 向量写入 | `src/main/java/com/superbiz/agent/service/VectorIndexService.java` |
| AIOps query 增强 | `src/main/java/com/superbiz/agent/service/AiOpsService.java` |
| 工具调用记录 | `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java` |
+266
View File
@@ -0,0 +1,266 @@
# 检索与可观测性架构
**更新日期**:2026-07-05
**状态**:当前可运行架构
**参考历史文档**:`archive/2026-07-05-legacy/knowledge-retrieval-architecture.md`
## 1. 定位
本文补充 [rag-architecture.md](rag-architecture.md) 中的检索细节,重点回答:
- 查询如何进入 `lookup_knowledge`。
- L0 和 L1 当前分别承担什么职责。
- 检索结果如何归一化、去重、记录。
- 如何通过 trace 和 eval 判断检索质量。
当前架构与旧版最大的差异是:L0 不再因为唯一命中而默认跳过 L1。L0 是 hint 和解释信号,L1 语义检索是默认召回路径。
## 2. 检索总图
```mermaid
flowchart TD
Query["Agent query / AIOps recommended query"] --> Tool["LookupKnowledgeTool"]
Tool --> L0["KnowledgeIndexService.analyzeQuery"]
L0 --> L0Result["L0 hint: matches / domains / keywords"]
L0Result --> Filter["singleDomainOrNull -> category filter"]
Tool --> L1["VectorSearchService.searchSimilarDocuments"]
Filter --> L1
L1 --> Mode{"retrieval.vector-store.mode"}
Mode -->|auto| Spring["Spring AI VectorStore"]
Spring -->|failure| SDK["Milvus SDK fallback"]
Mode -->|spring-ai| Spring
Mode -->|sdk| SDK
Spring --> Candidates["L1 candidates"]
SDK --> Candidates
Candidates --> Normalize["relevance normalization"]
L0Result --> Normalize
Normalize --> Result["LookupResult"]
Result --> Dedup["RetrievedDocTracker session dedup"]
Dedup --> Final["final tool output"]
Final --> Invocation["tool_invocation"]
Final --> Agent["Agent Executor"]
```
## 3. L0 Hint 层
L0 的输入是原始 query,输出是解释性结构:
```text
matches
matchedKeywords
domains
singleDomainOrNull
```
当前职责:
| 职责 | 说明 |
|---|---|
| domain hint | 判断 query 可能属于哪个知识域 |
| entity / keyword hint | 记录命中的关键词、错误码、服务名等 |
| category filter candidate | 当只有单一领域时,给 L1 一个 metadata filter 候选 |
| trace explanation | 写入 `tool_invocation.retrieval_details`,用于解释检索为什么这么走 |
不再承担:
```text
matches=1 -> skip L1 -> 直接返回 L0 文档正文
```
原因:
- 子串命中不等价于最终相关性。
- L0 没有稳定排序和语义相似度。
- AIOps query 往往包含多个字段,单点关键词命中容易误导。
## 4. L1 语义检索层
L1 通过 `VectorSearchService` 调度,支持三种模式:
| 模式 | 行为 | 用途 |
|---|---|---|
| `auto` | 优先 Spring AI VectorStore,失败 fallback 到 SDK | 默认运行模式 |
| `spring-ai` | 只走 Spring AI VectorStore | 验证框架路径 |
| `sdk` | 只走 Milvus SDK | 对比旧链路或临时回退 |
### Spring AI VectorStore 路径
```text
SearchRequest
-> query
-> topK
-> similarityThresholdAll
-> optional filterExpression
-> VectorStore.similaritySearch
```
### Milvus SDK fallback
```text
query
-> VectorEmbeddingService.generateQueryVector
-> Milvus search(vector, topK, L2)
-> id / content / metadata
```
SDK fallback 保留的价值:
- VectorStore bean 缺失时不让 MVP 主链路中断。
- Spring AI collection/schema 配置异常时可回退。
- 便于 SDK 与 VectorStore 的结果对比。
## 5. 分数与相关性归一化
检索结果输出三类分数字段:
| 字段 | 说明 |
|---|---|
| `score` | 兼容旧逻辑的距离型分数 |
| `rawScore` | 底层检索实现原始分数 |
| `scoreLabel` | 原始分数语义,例如 `similarity` 或 `l2_distance` |
工具层再把 L0/L1 情况归一为:
| relevanceLevel | 含义 |
|---|---|
| `PRECISE` | L0 单命中且 L1 相似度高 |
| `HIGHLY_RELEVANT` | L1 相似度高,或 L0 多命中且 L1 支撑强 |
| `REFERENCE` | 可作为参考,但不足以声明强证据 |
| `DEDUPED` | 同 session 中已检索过,不重复注入上下文 |
归一化结果用于:
- 给 Agent 输出 completeness hint。
- 写入 `tool_invocation.relevance_level`。
- 给 Verifier 构造 `tool_trace_summary`。
- 供 EvaluationService 计算 evidence score。
## 6. 文档切片和 metadata
当前保留 Markdown-aware chunking。
关键 metadata:
```text
docId
chunkIndex
totalChunks
title
breadcrumb
category
source
```
embedding 输入已经增强为:
```text
title + breadcrumb + content
```
这解决旧版检索中的一个主要问题:单个 chunk 被召回后,LLM 不知道它属于哪个文档、哪个章节。
## 7. 输出和记录
`lookup_knowledge` 的输出会进入两条路径:
```mermaid
flowchart LR
LookupResult["LookupResult"] --> Agent["Agent context"]
LookupResult --> Recorder["ToolInvocationRecorder"]
Recorder --> Invocation["tool_invocation"]
Invocation --> Trace["DiagnosisTraceService"]
Invocation --> Summary["ToolTraceSummaryService"]
Summary --> Verifier["chat_verifier"]
Invocation --> Eval["EvaluationService / RAG eval"]
```
`tool_invocation` 中与检索相关的字段:
```text
retrieval_layer
l0_match_count
l1_match_count
retrieval_details
relevance_level
dedup_reason
output_preview
duration_ms
success
```
`retrieval_details` 承载更细信息,例如:
- L0 命中文档标题和路径。
- L1 分数。
- retrieved domains。
- evidence status。
- dedup reason。
## 8. 去重与行动记忆
当前 session 级去重由 `RetrievedDocTracker` 负责。
```text
sessionId + docKey
-> already retrieved?
-> yes: return dedup message and record dedup_reason
-> no: mark retrieved and return evidence
```
去重目的:
- 避免同一文档反复进入上下文。
- 降低 token 浪费。
- 给 Executor 一个“这个方向已经查过”的行动记忆。
注意:去重不是全局缓存,只在当前诊断 session 内生效。
## 9. 检索质量评测
检索质量不能只看一次接口返回,需要用固定 query 回归。
当前评测资产:
| 资产 | 用途 |
|---|---|
| `eval/rag-retrieval/cases/golden-cases.json` | 固定 query 和期望证据 |
| `eval/rag-retrieval/fixtures/` | 离线候选结果 |
| `eval/rag-retrieval/reports/baseline.md` | 人类可读基线 |
| `scripts/eval_rag_retrieval.py` | 离线回归 |
| `scripts/eval_rag_live_acceptance.py` | 运行环境验收 |
评测层次:
```text
offline baseline
-> 不依赖服务和外部组件
live acceptance
-> 调用 /api/search/similar
-> 验证重建索引后的真实检索
trace inspection
-> 检查 Agent 是否真的调用 lookup_knowledge
-> 检查 tool_invocation 证据是否完整
```
## 10. 后续增强
近期优先:
1. 完整 evidence block 输出。
2. 邻居 chunk / 同章节上下文扩展。
3. metadata taxonomy 清理。
4. Query Transformer / MultiQuery 可回退接入。
5. 更完整的 Recall@K、MRR、nDCG 报告。
暂不优先:
- 重新引入 L0 直接返回。
- 一次性迁移所有写入路径。
- 在没有评测收益前引入 rerank / RRF / BM25。
+196
View File
@@ -0,0 +1,196 @@
# 会话与 Trace 生命周期
**更新日期**:2026-07-05
**状态**:当前可运行架构
**参考历史文档**:`archive/2026-07-05-legacy/session-management.md`
## 1. 定位
旧版会话设计以 Redis 会话为主,MySQL 作为可选长期沉淀。当前 MVP 的可追踪诊断已经转为 MySQL Trace 三表为主:
```text
diagnosis_session
-> agent_step
-> tool_invocation
```
因此本文描述的是当前可运行链路:
- `sessionId` 是一次诊断和后续 trace/feedback 的关联键。
- `diagnosis_session` 保存会话级状态、问题、答案、自评估和反馈。
- `agent_step` 保存每个 Agent 模型调用。
- `tool_invocation` 保存工具调用事实。
- `DiagnosisTraceService` 聚合三类记录,形成可回放 trace。
## 2. 生命周期总图
```mermaid
flowchart TD
Start["request: chat / ai_ops"] --> Resolve["resolve sessionId"]
Resolve --> Create["create or reset diagnosis_session"]
Create --> Running["status = RUNNING"]
Running --> Agent["Agent workflow"]
Agent --> StepHook["AgentLoggingHook"]
StepHook --> Step["agent_step"]
Agent --> Tool["Evidence tools"]
Tool --> Invocation["tool_invocation"]
Agent --> Final{"workflow result"}
Final -->|success| Success["status = SUCCESS, answer saved"]
Final -->|failed| Failed["status = FAILED"]
Success --> Evaluation["self_evaluation merge"]
Failed --> Evaluation
Evaluation --> Trace["GET /api/diagnosis/{sessionId}/trace"]
Success --> Feedback["POST /api/feedback"]
Feedback --> Case["useful -> case_library"]
```
## 3. sessionId 规则
| 链路 | sessionId 来源 |
|---|---|
| Chat | 如果请求带 sessionId,则复用;否则生成短 UUID |
| AIOps | 如果 payload 带 sessionId,则复用;否则生成 UUID |
| Trace | URL path 中的 `{sessionId}` |
| Feedback | request body 中的 `sessionId` |
设计含义:
- 同一个 `sessionId` 可以贯穿诊断、trace 查询和用户反馈。
- 当前诊断开始时会重置当前 session 的运行态字段,例如 answer、duration、step/tool count。
- `sessionId` 是业务关联键,不依赖数据库自增 ID 暴露给外部。
## 4. 状态流转
```mermaid
stateDiagram-v2
[*] --> PENDING
PENDING --> RUNNING: start diagnosis
RUNNING --> SUCCESS: workflow completed
RUNNING --> FAILED: exception / empty state
SUCCESS --> SUCCESS: feedback submitted
FAILED --> FAILED: feedback submitted
```
字段边界:
| 字段 | 含义 |
|---|---|
| `status` | 执行状态:`PENDING` / `RUNNING` / `SUCCESS` / `FAILED` |
| `answer` | Agent 最终返回给用户的报告或答复 |
| `self_evaluation` | 系统自评估 JSON |
| `feedback` | 用户反馈:`useful` / `not_useful` / null |
`feedback` 不修改 `status`。一个执行成功但用户标记 `not_useful` 的 session,仍然应该是 `SUCCESS + feedback=not_useful`。
## 5. agent_step 写入
`AgentLoggingHook` 在模型调用前后写入和回填 `agent_step`。
```mermaid
sequenceDiagram
autonumber
participant Agent as ReactAgent
participant Hook as AgentLoggingHook
participant DB as agent_step
Agent->>Hook: before_model(messages, sessionId)
Hook->>DB: insert step_index / agent_name / model_input
Agent-->>Agent: model call
Agent->>Hook: after_model(messages, sessionId)
Hook->>DB: update model_output / thought / has_tool_call / duration / token_count
```
当前记录:
- `session_id`
- `step_index`
- `agent_name`
- `model_input`
- `model_output`
- `thought`
- `has_tool_call`
- `duration_ms`
- `token_count`
## 6. tool_invocation 写入
工具调用记录真实工具事实,不记录模型猜测。
关键字段:
```text
session_id
step_id
tool_name
input_params
output_preview
output_length
retrieval_layer
l0_match_count
l1_match_count
retrieval_details
relevance_level
dedup_reason
duration_ms
success
error_message
```
对 `lookup_knowledge`,`retrieval_details` 会承载 L0/L1、领域、证据状态、去重等检索细节。对非检索工具,检索字段可以为空。
## 7. Trace API 聚合
```text
GET /api/diagnosis/{sessionId}/trace
```
聚合逻辑:
```text
diagnosis_session by sessionId
+ agent_step ordered by step_index
+ tool_invocation ordered by id
-> DiagnosisTraceResponse
```
Trace 视图回答的问题:
- 这次诊断是否成功?
- 哪些 Agent 参与了?
- 每一步模型输入输出是什么摘要?
- 调用了哪些工具?
- 工具返回了什么证据?
- Verifier / AIOps rule 是否通过?
- 用户是否反馈有用?
## 8. Chat 与 AIOps 差异
| 维度 | Chat | AIOps |
|---|---|---|
| `agent_flow` | `CHAT` | `AI_OPS` |
| 编排方式 | `SequentialAgent`: Planner -> Executor -> Verifier | `SupervisorAgent`: Planner + Executor |
| 自评估 | `rule_evaluation` + `verifier_evaluation` | `aiops_rule_evaluation` |
| 答案字段 | Chat 最终答复 | 告警分析报告 |
| payload | 用户自然语言 + history | alert payload 或 auto-discovery |
## 9. 清理与边界
当前会话持久化边界:
- MySQL Trace 记录是主要可回放来源。
- Chat 历史仍可作为请求上下文传入 Agent,但不是本文档的主持久化模型。
- Redis 主会话存储是历史设计,不作为当前架构事实。
- `RetrievedDocTracker` 是 session 级运行时去重状态,诊断结束后清理。
## 10. 后续增强
可考虑:
1. Trace API 增加更结构化的 `self_evaluation` 展示。
2. `agent_step` 与 `tool_invocation.step_id` 建立更严格关联。
3. 对多轮同 session 诊断增加 run id,避免复用 session 时历史记录混杂。
4. 为 Trace 增加导出能力,服务面试演示和回归分析。
+109 -34
View File
@@ -1,26 +1,50 @@
# MVP Demo Runbook
# MVP 演示手册
This demo proves the MVP flow from user question to persisted diagnosis trace.
本目录用于演示 MVP 从用户问题到诊断 Trace 的完整闭环。
## Prerequisites
面试时建议先读:
- MySQL, Redis, Milvus/Zilliz, and LLM/embedding configuration are available through the current project configuration.
- Security and secret cleanup are intentionally out of scope for this MVP slice.
- The `mvp-demo` profile enables mock Prometheus and CLS providers so log and metric tools can return repeatable evidence.
- `ten-minute-interview-demo.md`:10 分钟现场演示脚本。
- `interview-walkthrough.md`:面试讲解话术。
- `trace-inspection-checklist.md`:Trace 字段检查清单。
- `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。
- `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。
## Start
## 1. 前置条件
- MySQL、Redis、Milvus/Zilliz、LLM 和 embedding 配置可用。
- 安全和密钥清理不属于当前 MVP 演示范围。
- `mvp-demo` profile 会启用 mock Prometheus 和 mock CLS,让日志和指标工具返回可复现证据。
## 2. 启动服务
```powershell
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
```
The service listens on:
服务地址:
```text
http://localhost:9900
```
## 1. Run Chat Diagnosis
## 3. Chat 诊断 Demo
最快方式:
```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
```
脚本会生成:
```text
mvp/demo/output/chat-response.json
mvp/demo/output/trace-response.json
mvp/demo/output/feedback-response.json
```
手动请求:
```powershell
$sessionId = "mvp-demo-payment-timeout-001"
@@ -36,13 +60,13 @@ Invoke-RestMethod `
-Body $body
```
Expected result:
期望结果:
- `data.success` is `true`.
- `data.sessionId` equals `mvp-demo-payment-timeout-001`.
- `data.answer` contains a diagnosis answer.
- `data.success = true`
- `data.sessionId = mvp-demo-payment-timeout-001`
- `data.answer` 包含诊断答复
## 2. Query Trace
## 4. 查询 Trace
```powershell
Invoke-RestMethod `
@@ -50,15 +74,15 @@ Invoke-RestMethod `
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace"
```
Expected result:
期望结果:
- `code` is `200`.
- `data.session.sessionId` equals the chat session id.
- `data.steps` contains planner/executor/verifier records for complex questions.
- `data.toolInvocations` contains evidence tool calls such as `lookup_knowledge`, `query_logs`, or `query_metrics`.
- `data.session.selfEvaluation` contains verifier or rule evaluation when available.
- `code = 200`
- `data.session.sessionId` 等于 Chat session id
- `data.steps` 包含 planner / executor / verifier 等步骤
- `data.toolInvocations` 包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具
- `data.session.selfEvaluation` 包含 verifier 或 rule evaluation
## 3. Submit Feedback
## 5. 提交反馈
```powershell
$feedback = @{
@@ -73,22 +97,73 @@ Invoke-RestMethod `
-Body $feedback
```
Expected result:
期望结果:
- `success` is `true`.
- A later trace query shows `data.session.feedback` as `useful`.
- `success = true`
- 后续 Trace 中 `data.session.feedback = useful`
- useful 反馈会尝试沉淀 `case_library`
## Demo Story
## 6. AIOps 告警诊断 Demo
The important interview story is:
```powershell
$aiopsSessionId = "mvp-demo-aiops-payment-cpu-001"
$aiopsBody = @{
sessionId = $aiopsSessionId
alertName = "HighCPUUsage"
service = "payment-service"
severity = "P1"
description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
timeRange = "last_15m"
userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
} | ConvertTo-Json
Invoke-WebRequest `
-Method Post `
-Uri "http://localhost:9900/api/ai_ops" `
-ContentType "application/json" `
-Body $aiopsBody
```
期望结果:
- SSE 首条包含 `session` 消息,sessionId 为 `mvp-demo-aiops-payment-cpu-001`
- 后续流式输出包含 AIOps 告警分析报告
- 报告聚焦输入的 `HighCPUUsage/payment-service`
- 同一 session 的 Trace 中 `data.session.agentFlow = AI_OPS`
- `data.session.answer` 包含最终告警报告
- `data.toolInvocations` 包含证据工具调用
查询 AIOps Trace:
```powershell
Invoke-RestMethod `
-Method Get `
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace"
```
## 7. Demo 主线
Chat 主线:
```text
one session id
-> user question
-> multi-agent execution
-> evidence tools
-> verifier/self-evaluation
-> final answer
-> feedback
-> trace API for replay and audit
一个 session id
-> 用户问题
-> 多 Agent 执行
-> 证据工具
-> Verifier / self_evaluation
-> 最终答案
-> 用户反馈
-> Trace API 回放
```
AIOps 主线:
```text
一个 session id
-> 告警 payload
-> AIOps Planner / Executor
-> 证据工具
-> 告警分析报告
-> AIOps rule evaluation
-> Trace API 回放
```
+42
View File
@@ -0,0 +1,42 @@
# AIOps 告警验收用例
## 1. 目标
验证旧版 `/api/ai_ops` 入口可以作为可追踪的告警触发诊断入口,并且 payload 模式下报告聚焦输入告警。
## 2. 输入
- Session id:`mvp-demo-aiops-payment-cpu-001`
- Endpoint:`POST /api/ai_ops`
- Profile:`mvp-demo`
- 告警 payload:
```json
{
"sessionId": "mvp-demo-aiops-payment-cpu-001",
"alertName": "HighCPUUsage",
"service": "payment-service",
"severity": "P1",
"description": "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。",
"timeRange": "last_15m",
"userRequest": "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
}
```
## 3. 验收标准
1. SSE 流输出 `session` 消息,且包含请求中的 session id。
2. AIOps 执行创建或更新 `diagnosis_session`,并写入 `agent_flow = AI_OPS`。
3. 持久化的 session query 包含告警名、服务名、等级、时间范围和描述。
4. 如果生成最终报告,`diagnosis_session.answer` 包含该报告。
5. `GET /api/diagnosis/{sessionId}/trace` 返回 AIOps session、按顺序排列的 agent steps 和 tool invocations。
6. payload 模式下,报告主线聚焦 `HighCPUUsage/payment-service`。
7. 其他活跃告警最多作为相关风险或上下文出现,不应展开成完整独立根因章节。
8. `self_evaluation.aiops_rule_evaluation` 存在,并能反映报告完整性、payload 聚焦和证据工具覆盖情况。
## 4. 已知边界
- 当前 AIOps 使用轻量规则评估器,不是完整 LLM Verifier。
- 完整运行仍依赖有效的 DB、Redis、Milvus/Zilliz、模型和 embedding 配置。
- `mvp-demo` profile 使用 mock Prometheus 和 mock CLS,主要用于稳定演示。
+144
View File
@@ -0,0 +1,144 @@
# 面试演示讲解稿
这是一份短时间 Agent 工程面试用讲解稿,不是完整系统文档。
## 1. 30 秒摘要
```text
这是一个企业故障诊断 Agent MVP。
它接收支付超时问题,规划排查步骤,调用证据工具,
用 Verifier 检查答案,把完整 Trace 持久化,并支持用户反馈。
```
关键主张不是“模型回答了一次”,而是:
```text
系统能展示用了什么证据、答案如何被检查、如何用 sessionId 回放整次诊断。
```
## 2. Demo 流程
1. 用 `mvp-demo` profile 启动服务。
2. 运行固定的支付超时请求。
3. 打开 `mvp/demo/output/chat-response.json`。
4. 打开 `mvp/demo/output/trace-response.json`。
5. 指出证据工具和 verifier evaluation。
6. 提交 feedback,并展示它挂在同一个 session 上。
## 3. 命令
启动服务:
```powershell
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
```
另开终端运行 Demo:
```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
```
可选自定义 session:
```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1 -SessionId "mvp-demo-payment-timeout-002"
```
## 4. 展示什么
### 4.1 用户侧答案
文件:
```text
mvp/demo/output/chat-response.json
```
话术:
```text
这是用户看到的答案。这里的 sessionId 是稳定的,所以我后面可以追踪这一次回答是怎么来的。
```
### 4.2 证据 Trace
文件:
```text
mvp/demo/output/trace-response.json
```
话术:
```text
这才是 Agent 工程最重要的部分。
我可以检查 Agent 调用了哪些工具、每个工具拿到什么入参、是否成功、返回了什么证据预览。
```
重点字段:
- `data.toolInvocations[*].toolName`
- `data.toolInvocations[*].inputParams`
- `data.toolInvocations[*].outputPreview`
- `data.toolInvocations[*].success`
### 4.3 Verifier / 自评估
重点字段:
- `data.session.selfEvaluation`
- `data.summary.hasVerifierEvaluation`
话术:
```text
最终答案不是 Executor 原始输出直接返回。
系统会基于持久化的工具 trace 做 Verifier 或规则自评估。
这样系统可以区分 PASS、LOW_CONFID、REJECT,而不是假装每个答案都同样可信。
```
### 4.4 反馈闭环
文件:
```text
mvp/demo/output/feedback-response.json
```
必要时重新查询 Trace。
话术:
```text
feedback 会挂在同一个 diagnosis session 上。
这让后续挖掘 useful case 或 not_useful bad case 成为可能。
```
### 4.5 回归故事
如果被问到稳定性,可以补充:
```text
我把运行时 Demo 和离线 eval 分开。
Demo 证明真实链路能跑通,offline eval baseline 证明固定 case 可以回归。
这两者分开是有意的:Demo 面向人类审阅,eval 面向自动化信号。
```
## 5. 强面试表达
```text
我关注的是 Agent 工程表面:
traceability、evidence persistence、verifier gating、feedback 和 regression checks。
模型答案只是系统的一部分。
更重要的是答案产出后,能否被审计、验证和持续改进。
```
## 6. 主动说明限制
```text
这个 MVP 仍依赖 MySQL、Redis、Milvus 和模型凭证。
mvp-demo profile mock 了日志和指标,但不是完整生产运行环境。
密钥清理、默认隔离测试和生产可靠性是后续 hardening 工作。
```
+12
View File
@@ -0,0 +1,12 @@
# Demo 输出目录
本目录是本地 Demo 响应的默认输出位置。
生成文件会被 Git 忽略:
- `chat-response.json`
- `trace-response.json`
- `feedback-response.json`
保留此 README 是为了让目录存在于仓库中。
+19 -18
View File
@@ -1,24 +1,24 @@
# Payment Timeout Acceptance Case
# 支付超时诊断验收用例
## Goal
## 1. 目标
Validate that the MVP can diagnose a payment timeout incident and expose the complete trace for replay.
验证 MVP 能诊断支付超时问题,并暴露完整 Trace 供回放。
## Input
## 2. 输入
- Session id: `mvp-demo-payment-timeout-001`
- Question: `支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。`
- Profile: `mvp-demo`
- Session id:`mvp-demo-payment-timeout-001`
- 问题:`支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。`
- Profile:`mvp-demo`
## Acceptance Criteria
## 3. 验收标准
1. Chat returns a successful answer with the same session id.
2. Trace API returns session metadata, final answer, ordered agent steps, and ordered tool invocations.
3. Trace contains enough evidence to explain which tools were used and whether verifier/self-evaluation was persisted.
4. Feedback can be submitted for the same session id.
5. A follow-up trace query shows the persisted feedback value.
1. Chat 返回成功答复,且 session id 与请求一致。
2. Trace API 返回 session 元数据、最终答案、按顺序排列的 agent steps 和 tool invocations。
3. Trace 中有足够证据说明用了哪些工具,以及 verifier / self-evaluation 是否已持久化。
4. 可以使用同一个 session id 提交反馈。
5. 后续 Trace 查询能看到已持久化的 feedback 值。
## Trace Fields To Inspect
## 4. 需要检查的 Trace 字段
- `data.session.query`
- `data.session.answer`
@@ -32,8 +32,9 @@ Validate that the MVP can diagnose a payment timeout incident and expose the com
- `data.toolInvocations[*].retrievalDetails`
- `data.summary`
## Known Limits
## 5. 已知边界
- 这不是完整离线测试,仍需要有效的 chat、持久化、向量检索和模型调用环境。
- `mvp-demo` profile 启用 mock 日志和指标,让证据工具返回更稳定。
- 敏感配置清理不属于当前 MVP 优先级。
- This case is not a full offline test. It still requires valid infrastructure for chat, persistence, vector search, and model calls.
- Mock logs and metrics are enabled by the `mvp-demo` profile to make those evidence tools repeatable.
- Sensitive configuration cleanup is deferred by current MVP priority.
@@ -0,0 +1,4 @@
{
"Id": "mvp-demo-payment-timeout-001",
"Question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
}
@@ -0,0 +1,54 @@
param(
[string]$BaseUrl = "http://localhost:9900",
[string]$SessionId = "mvp-demo-payment-timeout-001",
[string]$RequestFile = "$PSScriptRoot/../requests/payment-timeout-chat.json",
[string]$OutputDir = "$PSScriptRoot/../output"
)
$ErrorActionPreference = "Stop"
New-Item -ItemType Directory -Force -Path $OutputDir | Out-Null
$request = Get-Content -Raw -Encoding UTF8 -Path $RequestFile | ConvertFrom-Json
$request.Id = $SessionId
$body = $request | ConvertTo-Json -Depth 8
Write-Host "正在运行支付超时 Chat 诊断 Demo..."
Write-Host "BaseUrl: $BaseUrl"
Write-Host "SessionId: $SessionId"
$chat = Invoke-RestMethod `
-Method Post `
-Uri "$BaseUrl/api/chat" `
-ContentType "application/json; charset=utf-8" `
-Body $body
$chat | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-response.json"
Write-Host "已保存 Chat 响应: $OutputDir/chat-response.json"
$trace = Invoke-RestMethod `
-Method Get `
-Uri "$BaseUrl/api/diagnosis/$SessionId/trace"
$trace | ConvertTo-Json -Depth 50 | Set-Content -Encoding UTF8 -Path "$OutputDir/trace-response.json"
Write-Host "已保存 Trace 响应: $OutputDir/trace-response.json"
$feedbackBody = @{
sessionId = $SessionId
feedback = "useful"
} | ConvertTo-Json
$feedback = Invoke-RestMethod `
-Method Post `
-Uri "$BaseUrl/api/feedback" `
-ContentType "application/json; charset=utf-8" `
-Body $feedbackBody
$feedback | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/feedback-response.json"
Write-Host "已保存反馈响应: $OutputDir/feedback-response.json"
Write-Host ""
Write-Host "Demo 已完成,请检查:"
Write-Host "- mvp/demo/output/chat-response.json"
Write-Host "- mvp/demo/output/trace-response.json"
Write-Host "- mvp/demo/output/feedback-response.json"
+237
View File
@@ -0,0 +1,237 @@
# 10 分钟面试演示脚本
**用途**:面试现场按步骤演示
**目标**:展示从问题到证据、验证、Trace、反馈的闭环
**前置条件**:服务以 `mvp-demo` profile 启动
更完整的 runbook 见 [README.md](README.md),字段检查见 [trace-inspection-checklist.md](trace-inspection-checklist.md)。
## 0. 开场话术
```text
我会演示一个支付超时诊断。
重点不是看模型给出一段答案,而是看这个答案背后的 Agent 执行链路:
Planner 怎么拆解,Executor 调了哪些工具,Verifier 如何判断证据是否支撑答案,以及最终如何通过 sessionId 回放。
```
## 1. 启动服务
```powershell
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
```
服务地址:
```text
http://localhost:9900
```
说明:
- `mvp-demo` profile 使用 mock Prometheus 和 mock CLS。
- 演示不依赖真实线上故障。
- MySQL、Redis、Milvus/Zilliz 和模型配置仍需要可用。
## 2. 演示 Chat 诊断
推荐使用固定脚本:
```powershell
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
```
脚本会写出:
```text
mvp/demo/output/chat-response.json
mvp/demo/output/trace-response.json
mvp/demo/output/feedback-response.json
```
现场话术:
```text
这里我用固定 sessionId 跑一个支付接口超时问题。
固定 sessionId 的好处是,后面 trace 和 feedback 都能关联到同一次诊断。
```
## 3. 展示用户答案
打开:
```text
mvp/demo/output/chat-response.json
```
重点看:
```text
data.sessionId
data.answer
```
现场话术:
```text
这是用户看到的答案。
但这个项目的重点不是这段文字,而是这段文字是否有证据链。
接下来我用同一个 sessionId 查 trace。
```
## 4. 展示 Trace
打开:
```text
mvp/demo/output/trace-response.json
```
重点看:
```text
data.session.sessionId
data.session.agentFlow
data.steps[*].agentName
data.toolInvocations[*].toolName
data.toolInvocations[*].inputParams
data.toolInvocations[*].outputPreview
data.toolInvocations[*].retrievalLayer
data.toolInvocations[*].relevanceLevel
data.summary.hasVerifierEvaluation
```
现场话术:
```text
这里能看到三个层次:
第一,session 记录了这次诊断的问题、答案、耗时和自评估。
第二,agent_step 记录 Planner、Executor、Verifier 的模型步骤。
第三,tool_invocation 记录真实工具调用,包括 lookup_knowledge、日志和指标。
所以这不是一个黑盒 Chatbot,而是一条可以回放的诊断链路。
```
## 5. 展示知识库检索
在 trace 中找到 `lookup_knowledge`。
重点看:
```text
toolName = lookup_knowledge
inputParams.query
retrievalLayer
l0MatchCount
l1MatchCount
relevanceLevel
retrievalDetails
outputPreview
```
现场话术:
```text
知识库检索保留为显式工具,而不是藏在 Advisor 里。
这样面试官或线上排查人员能看到:Agent 查了什么 query,命中了哪个知识域,检索层是 L0/L1 还是混合,相关性等级是什么。
底层检索现在走 VectorSearchService,优先 Spring AI VectorStore,失败时 fallback 到 Milvus SDK。
```
## 6. 展示 Verifier
在 trace 中查看:
```text
data.session.selfEvaluation
data.summary.hasVerifierEvaluation
```
现场话术:
```text
Verifier 不做新检索,只看工具 trace 汇总。
它会把 Executor 答案里的关键事实拆出来,判断每条事实是 direct_evidence、indirect_support、no_evidence 还是 contradicted。
如果 PASS,就输出原答案。
如果 LOW_CONFID,可以补证据或加低置信提示。
如果 REJECT,就降级输出,只保留已确认信息。
```
## 7. 展示反馈闭环
打开:
```text
mvp/demo/output/feedback-response.json
```
重点看:
```text
success
caseId
```
现场话术:
```text
用户反馈 useful 会写回同一个 diagnosis_session。
后端会把这次诊断自动沉淀到 case_library,后续可以做案例检索或 bad case 分析。
这里 status 和 feedback 是分开的:
status 表示执行是否成功,feedback 表示用户是否认可。
```
## 8. 可选演示 AIOps
如果时间允许,再演示 AIOps payload。
请求示例见:
```text
mvp/demo/README.md
```
现场话术:
```text
AIOps 有两个模式。
有 payload 时进入 PAYLOAD_TARGETED,报告必须聚焦这个告警。
没有 payload 时进入 AUTO_DISCOVERY,先发现活跃告警再排查。
我专门加了 recommended lookup_knowledge query,把 alertName、service、severity、description 等字段稳定送入知识库检索,避免 Agent 随意扩展问题范围。
```
## 9. 结束总结
```text
这个 Demo 展示的是一个完整闭环:
用户问题
-> Agent 规划和执行
-> 显式工具证据
-> Verifier / self_evaluation
-> Trace 回放
-> 用户反馈
-> 案例沉淀
我把重点放在 Agent 工程能力:可追踪、可验证、可回归、可演进。
```
## 10. 如果现场失败
如果模型或外部组件不可用,不要硬跑。可以直接打开上一次输出:
```text
mvp/demo/output/chat-response.json
mvp/demo/output/trace-response.json
mvp/demo/output/feedback-response.json
```
降级话术:
```text
现场环境依赖 MySQL、Redis、Milvus 和模型服务。
如果外部服务不可用,我会用固定输出讲 trace 结构。
因为这个项目的核心不是一次在线请求,而是诊断链路如何被记录、检查和回放。
```
+53
View File
@@ -0,0 +1,53 @@
# Trace 检查清单
运行 `scripts/run-payment-timeout-demo.ps1` 后,用这份清单检查 `trace-response.json`。
## 1. Session
| JSON path | 检查点 | 面试讲点 |
|---|---|---|
| `data.session.sessionId` | 是否等于 `mvp-demo-payment-timeout-001` | 一个 session id 串起 chat、工具、verifier、feedback 和 trace |
| `data.session.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
| `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
| `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在同一次诊断上 |
## 2. Agent 步骤
| JSON path | 检查点 | 面试讲点 |
|---|---|---|
| `data.steps[*].agentName` | 是否有 Planner / Executor / Verifier 或等价步骤 | 流程被拆成可检查的 Agent 步骤 |
| `data.steps[*].thought` | 是否有高层步骤摘要 | 内部过程可审计,不只看最终文本 |
| `data.steps[*].durationMs` | 是否有步骤耗时 | Trace 可用于耗时分析 |
| `data.steps[*].tokenCount` | 如可用,是否记录 token | Trace 可用于模型成本分析 |
## 3. 工具证据
| JSON path | 检查点 | 面试讲点 |
|---|---|---|
| `data.toolInvocations[*].toolName` | 是否包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具 | Agent 通过工具收集证据,而不是无依据猜测 |
| `data.toolInvocations[*].inputParams` | 是否能看到每个工具的入参 | 工具输入可审计、可调试 |
| `data.toolInvocations[*].outputPreview` | 是否有受控长度的证据预览 | 保留证据但不倾倒巨大 payload |
| `data.toolInvocations[*].success` | 是否区分成功和失败 | 工具失败对 Verifier 和 reviewer 可见 |
| `data.toolInvocations[*].retrievalDetails` | 是否包含检索 metadata | 检索质量可事后检查 |
| `data.toolInvocations[*].relevanceLevel` | 是否有相关性等级 | 可解释检索结果强弱 |
## 4. Summary
| JSON path | 检查点 | 面试讲点 |
|---|---|---|
| `data.summary.persistedStepCount` | step 行是否持久化 | Trace 来自存储,不是响应内存 |
| `data.summary.persistedToolCallCount` | tool 行是否持久化 | 工具证据在请求结束后仍可回放 |
| `data.summary.hasVerifierEvaluation` | 是否存在 Verifier 结果 | 最终答案经过质量门 |
| `data.summary.hasFeedback` | 提交反馈后是否为 true | 人类反馈闭环完成 |
## 5. 好的结果长什么样
```text
同一个 session id
-> 最终答案
-> 持久化 agent steps
-> 持久化 evidence tool calls
-> verifier / self-evaluation
-> feedback attached to the same session
```
+61
View File
@@ -0,0 +1,61 @@
# Diagnosis Eval Harness
This folder contains the first fixed-case evaluation set for the MVP diagnosis Agent.
## Scope
- Case definitions: `cases/diagnosis-cases.json`
- Offline trace fixtures: `fixtures/*.json`
- Field definitions: `schema.md`
- Baseline reports: `reports/baseline-report.json` and `reports/baseline-report.md`
- Baseline diff sample: `reports/baseline-diff-sample.json` and `reports/baseline-diff-sample.md`
- Evaluator implementation: `DiagnosisTraceEvaluator`
- Report writer: `DiagnosisEvalReportWriter`
## Current Mode
The first version evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.
## Verification
Run the focused evaluator test:
```powershell
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test
```
The committed baseline report represents the current fixed fixture set:
```text
5 fixed cases
5 passing fixture evaluations
2 PASS verdicts
3 LOW_CONFID verdicts
```
When fixtures or evaluator rules change, regenerate the report from the same case file and fixture directory, then update both JSON and Markdown outputs together.
## Interview Story
The harness gives the MVP a repeatable baseline:
```text
fixed diagnosis case
-> saved or runtime trace
-> rule-based trace validation
-> JSON / Markdown report
-> regression signal for prompts, tools, retrieval, and verifier behavior
```
## Baseline Diff
Baseline diff compares a current report against `reports/baseline-report.json`.
```text
baseline report
current report
-> deterministic diff
-> regressions, improvements, and changed signals
```
Use it to answer: did a prompt, tool, retrieval, or verifier change make the Agent worse than the fixed baseline?
+57
View File
@@ -0,0 +1,57 @@
[
{
"id": "payment-timeout",
"title": "Payment API timeout",
"question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
"traceFixture": "payment-timeout-pass.json",
"expectedRootCauseKeywords": ["支付", "超时", "连接池"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["lookup_knowledge", "query_logs", "query_metrics"],
"allowedVerdicts": ["PASS", "LOW_CONFID"],
"forbiddenAnswerKeywords": ["无证据确定"]
},
{
"id": "mysql-pool-exhausted",
"title": "MySQL connection pool exhausted",
"question": "订单服务大量请求超时,请判断是否和 MySQL 连接池有关。",
"traceFixture": "mysql-pool-low-confid.json",
"expectedRootCauseKeywords": ["mysql", "连接池", "超时"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["lookup_knowledge", "query_logs"],
"allowedVerdicts": ["LOW_CONFID", "PASS"],
"forbiddenAnswerKeywords": ["已经完全确认"]
},
{
"id": "redis-timeout",
"title": "Redis timeout",
"question": "支付服务出现 Redis 连接超时,请定位可能原因。",
"traceFixture": "redis-timeout-low-confid.json",
"expectedRootCauseKeywords": ["redis", "超时"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["query_logs"],
"allowedVerdicts": ["LOW_CONFID", "PASS"],
"forbiddenAnswerKeywords": ["无需进一步排查"]
},
{
"id": "slow-response",
"title": "Slow response",
"question": "用户服务 P99 响应时间升高,请结合指标和日志分析。",
"traceFixture": "slow-response-pass.json",
"expectedRootCauseKeywords": ["p99", "慢响应"],
"minKeywordMatches": 1,
"requiredEvidenceTools": ["query_metrics", "query_logs"],
"allowedVerdicts": ["LOW_CONFID", "PASS"],
"forbiddenAnswerKeywords": ["没有风险"]
},
{
"id": "jvm-memory-risk",
"title": "JVM memory risk",
"question": "订单服务内存使用率过高,请判断是否存在 OOM 风险。",
"traceFixture": "jvm-memory-risk-low-confid.json",
"expectedRootCauseKeywords": ["jvm", "内存", "oom"],
"minKeywordMatches": 2,
"requiredEvidenceTools": ["query_metrics", "query_logs"],
"allowedVerdicts": ["LOW_CONFID", "PASS"],
"forbiddenAnswerKeywords": ["可以忽略"]
}
]
@@ -0,0 +1,52 @@
{
"session": {
"sessionId": "eval-jvm-memory-risk",
"query": "订单服务内存使用率过高,请判断是否存在 OOM 风险。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 53000,
"toolCallCount": 2,
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n订单服务存在 JVM 内存风险,但还不能完全确认会发生 OOM。指标显示 heap 使用率持续高于 88%,日志出现多次 Full GC 和 allocation pressure,需要继续观察对象增长来源并检查最近发布。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "LOW_CONFID",
"groundedness_score": 0.52,
"tool_trace_summary": [
{
"tool_name": "query_metrics",
"success": true,
"evidence_level": "direct"
},
{
"tool_name": "query_logs",
"success": true,
"evidence_level": "indirect"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 1,
"sessionId": "eval-jvm-memory-risk",
"toolName": "query_metrics",
"success": true
},
{
"id": 2,
"sessionId": "eval-jvm-memory-risk",
"toolName": "query_logs",
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 2,
"returnedToolCallCount": 2,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,53 @@
{
"session": {
"sessionId": "eval-mysql-pool",
"query": "订单服务大量请求超时,请判断是否和 MySQL 连接池有关。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 51000,
"toolCallCount": 2,
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\nMySQL 连接池可能参与了本次超时问题。日志中出现 connection pool exhausted,但当前缺少完整指标证据,因此只能作为低置信结论处理。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "LOW_CONFID",
"groundedness_score": 0.48,
"tool_trace_summary": [
{
"tool_name": "lookup_knowledge",
"success": true,
"evidence_level": "direct"
},
{
"tool_name": "query_logs",
"success": true,
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 1,
"sessionId": "eval-mysql-pool",
"toolName": "lookup_knowledge",
"success": true,
"relevanceLevel": "PRECISE"
},
{
"id": 2,
"sessionId": "eval-mysql-pool",
"toolName": "query_logs",
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 2,
"returnedToolCallCount": 2,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,64 @@
{
"session": {
"sessionId": "eval-payment-timeout",
"query": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 42000,
"toolCallCount": 3,
"answer": "支付接口超时与连接池等待有关。知识库说明支付超时需要同时检查连接池、日志和指标;日志出现 connection pool exhausted;指标显示支付服务延迟升高。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "PASS",
"groundedness_score": 0.86,
"tool_trace_summary": [
{
"tool_name": "lookup_knowledge",
"success": true,
"evidence_level": "direct"
},
{
"tool_name": "query_logs",
"success": true,
"evidence_level": "direct"
},
{
"tool_name": "query_metrics",
"success": true,
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 1,
"sessionId": "eval-payment-timeout",
"toolName": "lookup_knowledge",
"success": true,
"relevanceLevel": "PRECISE"
},
{
"id": 2,
"sessionId": "eval-payment-timeout",
"toolName": "query_logs",
"success": true
},
{
"id": 3,
"sessionId": "eval-payment-timeout",
"toolName": "query_metrics",
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 3,
"returnedToolCallCount": 3,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,41 @@
{
"session": {
"sessionId": "eval-redis-timeout",
"query": "支付服务出现 Redis 连接超时,请定位可能原因。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 36000,
"toolCallCount": 1,
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\nRedis 连接超时可能和支付服务到 Redis 的网络抖动或连接池等待有关。日志中出现 redis timeout 和 command timeout 记录,但当前缺少指标侧证据,因此只能作为低置信结论处理。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "LOW_CONFID",
"groundedness_score": 0.46,
"tool_trace_summary": [
{
"tool_name": "query_logs",
"success": true,
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 1,
"sessionId": "eval-redis-timeout",
"toolName": "query_logs",
"success": true
}
],
"summary": {
"persistedStepCount": 2,
"returnedStepCount": 2,
"persistedToolCallCount": 1,
"returnedToolCallCount": 1,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
+52
View File
@@ -0,0 +1,52 @@
{
"session": {
"sessionId": "eval-slow-response",
"query": "用户服务 P99 响应时间升高,请结合指标和日志分析。",
"status": "SUCCESS",
"agentFlow": "CHAT",
"totalDurationMs": 47000,
"toolCallCount": 2,
"answer": "用户服务 P99 升高主要表现为慢响应。指标显示 P99 latency 从 280ms 上升到 1800ms,日志中同时出现 slow request 和 downstream timeout,因此优先排查下游依赖耗时和线程池排队。",
"selfEvaluation": {
"verifier_evaluation": {
"verdict": "PASS",
"groundedness_score": 0.78,
"tool_trace_summary": [
{
"tool_name": "query_metrics",
"success": true,
"evidence_level": "direct"
},
{
"tool_name": "query_logs",
"success": true,
"evidence_level": "direct"
}
]
}
}
},
"steps": [],
"toolInvocations": [
{
"id": 1,
"sessionId": "eval-slow-response",
"toolName": "query_metrics",
"success": true
},
{
"id": 2,
"sessionId": "eval-slow-response",
"toolName": "query_logs",
"success": true
}
],
"summary": {
"persistedStepCount": 3,
"returnedStepCount": 3,
"persistedToolCallCount": 2,
"returnedToolCallCount": 2,
"hasVerifierEvaluation": true,
"hasFeedback": false
}
}
@@ -0,0 +1,85 @@
{
"baselineTotalCases" : 5,
"currentTotalCases" : 5,
"baselinePassedCases" : 5,
"currentPassedCases" : 4,
"baselinePassRate" : 1.0,
"currentPassRate" : 0.8,
"regressionCount" : 6,
"improvementCount" : 0,
"changedCount" : 2,
"hasRegression" : true,
"items" : [ {
"type" : "REGRESSION",
"scope" : "aggregate",
"caseId" : null,
"metric" : "passRate",
"baselineValue" : "1.0",
"currentValue" : "0.8",
"delta" : -0.19999999999999996,
"message" : "passRate changed"
}, {
"type" : "REGRESSION",
"scope" : "aggregate",
"caseId" : null,
"metric" : "averageToolCallCount",
"baselineValue" : "2.0",
"currentValue" : "3.0",
"delta" : 1.0,
"message" : "averageToolCallCount changed"
}, {
"type" : "CHANGED",
"scope" : "aggregate",
"caseId" : null,
"metric" : "verdictDistribution.LOW_CONFID",
"baselineValue" : "3",
"currentValue" : "2",
"delta" : -1.0,
"message" : "verdict count changed for LOW_CONFID"
}, {
"type" : "CHANGED",
"scope" : "aggregate",
"caseId" : null,
"metric" : "verdictDistribution.REJECT",
"baselineValue" : "0",
"currentValue" : "1",
"delta" : 1.0,
"message" : "verdict count changed for REJECT"
}, {
"type" : "REGRESSION",
"scope" : "case",
"caseId" : "redis-timeout",
"metric" : "passed",
"baselineValue" : "true",
"currentValue" : "false",
"delta" : null,
"message" : "redis-timeout pass state changed"
}, {
"type" : "REGRESSION",
"scope" : "case",
"caseId" : "redis-timeout",
"metric" : "verdict",
"baselineValue" : "LOW_CONFID",
"currentValue" : "REJECT",
"delta" : -1.0,
"message" : "redis-timeout verdict changed"
}, {
"type" : "REGRESSION",
"scope" : "case",
"caseId" : "redis-timeout",
"metric" : "matchedKeywordCount",
"baselineValue" : "2",
"currentValue" : "1",
"delta" : -1.0,
"message" : "redis-timeout matchedKeywordCount changed"
}, {
"type" : "REGRESSION",
"scope" : "case",
"caseId" : "redis-timeout",
"metric" : "evidenceCoverage.query_logs",
"baselineValue" : "true",
"currentValue" : "false",
"delta" : null,
"message" : "redis-timeout evidence coverage changed for query_logs"
} ]
}
+22
View File
@@ -0,0 +1,22 @@
# Diagnosis Eval Baseline Diff
- Baseline pass rate: 100.00%
- Current pass rate: 80.00%
- Baseline passed cases: 5/5
- Current passed cases: 4/5
- Regressions: 6
- Improvements: 0
- Other changes: 2
## Diff Items
| Type | Scope | Case | Metric | Baseline | Current | Delta | Message |
| --- | --- | --- | --- | --- | --- | ---: | --- |
| REGRESSION | aggregate | - | passRate | 1.0 | 0.8 | -0.200 | passRate changed |
| REGRESSION | aggregate | - | averageToolCallCount | 2.0 | 3.0 | 1.000 | averageToolCallCount changed |
| CHANGED | aggregate | - | verdictDistribution.LOW_CONFID | 3 | 2 | -1.000 | verdict count changed for LOW_CONFID |
| CHANGED | aggregate | - | verdictDistribution.REJECT | 0 | 1 | 1.000 | verdict count changed for REJECT |
| REGRESSION | case | redis-timeout | passed | true | false | - | redis-timeout pass state changed |
| REGRESSION | case | redis-timeout | verdict | LOW_CONFID | REJECT | -1.000 | redis-timeout verdict changed |
| REGRESSION | case | redis-timeout | matchedKeywordCount | 2 | 1 | -1.000 | redis-timeout matchedKeywordCount changed |
| REGRESSION | case | redis-timeout | evidenceCoverage.query_logs | true | false | - | redis-timeout evidence coverage changed for query_logs |

Some files were not shown because too many files have changed in this diff Show More