chore(rag): add eval knowledge base mirror

This commit is contained in:
zhuyongxin
2026-07-06 21:48:21 +08:00
parent ed7efc58b7
commit 64adb998cf
10 changed files with 231 additions and 0 deletions
+18
View File
@@ -0,0 +1,18 @@
# RAG Eval Knowledge Base Mirror
This folder stores the committed knowledge-base copy of the canonical RAG eval
documents.
The source of truth for the eval importer remains:
```text
eval/rag-retrieval/seed-docs/
```
The documents are kept under `knowledge_base/rag-eval/` so eval/test knowledge
does not mix with the normal business knowledge folders such as `api`,
`infrastructure`, or `troubleshooting`.
Each document keeps its original frontmatter `category` and `kb_scope`. The
category is still the retrieval category used by L0/L1, while `kb_scope:
rag-eval` isolates these documents during eval runs.
@@ -0,0 +1,26 @@
---
title: AIOps Alert Scope Control
keywords: [alert payload, unrelated active alerts, scope control]
summary: Keep diagnosis scoped to the request payload and avoid diagnosing unrelated active alerts.
category: aiops
source: aiops-alert-scope-control
breadcrumb: AIOps > Alert Scope Control
kb_scope: rag-eval
covers: [alert scope, payload, active alerts]
when_to_retrieve: Use when an AIOps request includes a concrete alert payload and scope boundaries matter.
---
# AIOps
## Alert Scope Control
When an AIOps request already includes an alert payload, the agent should diagnose that payload first.
It must not expand the task into unrelated active alerts unless the user asks for broad alert triage.
Scope rules:
1. Treat the provided payload as the primary incident boundary.
2. Use unrelated active alerts only as correlation evidence when they share service, dependency, time window, or trace context.
3. Do not replace the requested alert with a louder but unrelated alert.
This runbook anchors payload, unrelated active alerts, and scope behavior.
@@ -0,0 +1,28 @@
---
title: Payment Service Latency Alert
keywords: [HighLatency, payment-service, p95 latency, downstream dependency]
summary: Diagnose payment-service p95 latency alerts and identify downstream dependency bottlenecks.
category: aiops
source: payment-service-latency
breadcrumb: AIOps > Service Alerts > Payment Latency
kb_scope: rag-eval
covers: [payment-service, latency, downstream dependency]
when_to_retrieve: Use when an alert mentions payment-service, HighLatency, or elevated p95 latency.
---
# AIOps
## Service Alerts
### Payment Latency
For `HighLatency` alerts on `payment-service`, treat p95 latency as the primary symptom.
Diagnosis steps:
1. Confirm whether p95 latency is isolated to payment-service or shared across upstream callers.
2. Compare payment-service latency with downstream dependency latency for gateway, risk, and order services.
3. Check connection pool wait time, retry spikes, and timeout rates.
4. If downstream dependency latency increased first, classify payment-service as affected rather than root cause.
The expected evidence terms are p95 latency, payment-service, and downstream dependency.
@@ -0,0 +1,29 @@
---
title: MySQL Connection Pool Runbook
keywords: [MySQL connection pool, pool exhausted, max_connections, HikariCP]
summary: Diagnose exhausted MySQL connection pools and distinguish application leaks from database limits.
category: database
source: mysql-connection-pool
breadcrumb: Database > MySQL > Connection Pool
kb_scope: rag-eval
covers: [mysql, connection pool, database capacity]
when_to_retrieve: Use when MySQL clients report exhausted pools, connection acquisition timeout, max_connections pressure, or HikariCP saturation.
---
# Database
## MySQL
### Connection Pool
When MySQL connection pool is exhausted, first compare application pool usage with database `max_connections`.
For HikariCP, check `active`, `idle`, `pending`, and connection acquisition timeout metrics.
Recommended diagnosis:
1. Verify whether HikariCP active connections stay near maximum while pending threads grow.
2. Check MySQL `Threads_connected`, `Threads_running`, and `max_connections`.
3. Inspect slow SQL and long transactions that keep connections checked out.
4. If the database is healthy, look for application connection leaks or missing transaction boundaries.
Use this runbook as evidence for connection pool, max_connections, and HikariCP incidents.
@@ -0,0 +1,25 @@
---
title: RAG L0 Filter Fallback
keywords: [golden retry contract, second pass retrieval]
summary: Retry the raw query without the L0 category filter when filtered vector evidence is missing or low quality.
category: fallback
source: rag-l0-filter-fallback
breadcrumb: RAG > Fallback > Unfiltered Retry
kb_scope: rag-eval
covers: [fallback, unfiltered retry, retrieval quality]
when_to_retrieve: Use when validating the fallback contract for low-quality filtered vector retrieval.
---
# RAG
## Fallback
### Unfiltered Retry
If the first vector search is over-constrained by an L0 metadata filter and returns low quality evidence,
the retriever should skip the L0 filter and run an unfiltered vector retry with the original query.
The fallback reason should be `filtered_vector_low_quality` when the filtered candidate exists but is below the
reference threshold. If there is no usable evidence at all, use `filtered_vector_no_evidence`.
This document is the expected evidence for skip the L0 filter, unfiltered vector retry, and low quality behavior.
@@ -0,0 +1,27 @@
---
title: Incident Diagnosis Flow
keywords: [standard troubleshooting flow, application incident, collect evidence, verify, remediation]
summary: Standard flow for diagnosing application incidents with evidence, hypothesis verification, and remediation.
category: ops
source: incident-diagnosis-flow
breadcrumb: AIOps > Diagnosis Flow
kb_scope: rag-eval
covers: [incident diagnosis, evidence collection, remediation]
when_to_retrieve: Use when the user asks for a standard troubleshooting flow or incident diagnosis sequence.
---
# AIOps
## Diagnosis Flow
The standard troubleshooting flow is evidence first, hypothesis second, remediation last.
Recommended sequence:
1. Collect evidence from alerts, metrics, logs, traces, deployments, and recent configuration changes.
2. Define a small hypothesis that explains the observed symptoms.
3. Verify the hypothesis with a targeted metric, log query, or reproduction step.
4. Choose remediation that directly addresses the verified cause.
5. Record the outcome and the evidence used to make the decision.
Do not skip collect evidence, verify, and remediation ordering during an application incident.
@@ -0,0 +1,19 @@
---
title: RAG L0 Filter Decoy
keywords: [over-filtered by L0, filtered vector search, low quality evidence]
summary: Decoy document used to force the first filtered retrieval attempt into a low-quality category.
category: overfilter-decoy
source: rag-l0-filter-decoy
breadcrumb: RAG > Fallback > Decoy
kb_scope: rag-eval
covers: [fallback test decoy]
when_to_retrieve: Use only as a controlled eval decoy for over-filter fallback testing.
---
# Release Calendar
## Approval Window
This document describes an unrelated release calendar approval window.
It intentionally avoids the real fallback instructions so the filtered retrieval
attempt is low quality and the retriever must retry without the L0 category filter.
@@ -0,0 +1,28 @@
---
title: RAG Chunk Context Reconstruction
keywords: [split into multiple chunks, retrieval context, neighbor chunk, same section, breadcrumb context]
summary: Preserve context when long RAG sections are split into multiple retrievable chunks.
category: rag
source: rag-chunk-context-reconstruction
breadcrumb: RAG > Chunking > Context Reconstruction
kb_scope: rag-eval
covers: [rag chunking, context packing, breadcrumbs]
when_to_retrieve: Use when a retrieval question asks how to preserve context across split chunks.
---
# RAG
## Chunking
### Context Reconstruction
When a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.
Recommended behavior:
1. Store the breadcrumb with every chunk.
2. Preserve the same section identity across adjacent chunks.
3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.
4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.
The key concepts are neighbor chunk, same section, and breadcrumb.
@@ -0,0 +1,29 @@
---
title: RAG L0 Domain Entity Hint
keywords: [L0 keyword matching, final retrieval result, domain detector, entity extractor, metadata filter]
summary: Define L0 as a query transformation hint layer instead of final retrieval evidence.
category: rag
source: rag-l0-domain-entity-hint
breadcrumb: RAG > L0 > Domain Entity Hint
kb_scope: rag-eval
covers: [l0 hint, query transformation, metadata filter]
when_to_retrieve: Use when a question asks whether L0 should decide final retrieval or only provide hints.
---
# RAG
## L0
### Domain Entity Hint
L0 keyword matching should not decide the final retrieval result.
In the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.
The output can provide:
1. Candidate domain hints.
2. Matched entities and keywords.
3. An optional metadata filter for the first vector retrieval attempt.
Final evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.
The important terms are domain detector, entity extractor, and metadata filter.
+2
View File
@@ -155,6 +155,8 @@ eval/rag-retrieval/seed-docs/*.md
-> api_document metadata + L0 index + Milvus chunks
```
仓库内还保留一份 `knowledge_base/rag-eval/` 镜像,方便直接查看和提交 eval 知识库文档。它们放在单独目录下,避免和 `knowledge_base/api`、`knowledge_base/infrastructure` 等业务知识目录混在一起;检索 category 仍由 frontmatter 中的 `category` 决定。
Seed frontmatter includes:
```yaml