feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
This commit is contained in:
+89
-143
@@ -1,38 +1,65 @@
|
||||
# RAG Retrieval Baseline
|
||||
|
||||
This directory contains the offline retrieval baseline for the RAG refactor.
|
||||
Offline regression harness for `lookup_knowledge` **after** hybrid retrieval + qualityScore post-process.
|
||||
|
||||
The baseline is intentionally narrower than full diagnosis evaluation. It checks
|
||||
whether fixed retrieval queries can recover expected documents, breadcrumbs, and
|
||||
evidence keywords before changing L0 behavior, query augmentation, evidence
|
||||
post-processing, or Spring AI VectorStore integration.
|
||||
It checks whether fixed queries still recover expected documents, breadcrumbs, keywords, and pipeline behaviors (filter / unfiltered retry). It is **not** a full diagnosis-agent E2E.
|
||||
|
||||
Production knowledge path: `MilvusHybridKnowledgeStore` with `retrieval.search.mode=hybrid` (dense+BM25+RRF).
|
||||
`mode=dense` remains a same-collection baseline for recall comparison (not a second index).
|
||||
|
||||
Related design notes:
|
||||
|
||||
- `docs/RAG-Hybrid质量分与后处理.md`
|
||||
- `docs/RAG-Agent如何读relevance_level.md`
|
||||
- `mvp/architecture/RAG知识检索架构.md` §6
|
||||
|
||||
## Offline vs live
|
||||
|
||||
| Layer | What | Needs live stack? |
|
||||
|-------|------|-------------------|
|
||||
| **Offline** | `fixtures/*.json` × `golden-cases.json` → pass/fail + baseline diff | **No** (no Milvus/LLM/Boot) |
|
||||
| **Snapshot generate** | Real `LookupKnowledgeTool` writes fixtures | **Yes** (embedding + Milvus + DB/L0 as configured) |
|
||||
| **Live smoke** | optional `eval_rag_live_acceptance.py` | Yes (running app) |
|
||||
|
||||
Daily CI / local quick check: **offline only**.
|
||||
After changing retrieval, indexing, or search mode: **regenerate fixtures**, then offline eval, then update baseline if the diff is intentional.
|
||||
|
||||
## Layout
|
||||
|
||||
```text
|
||||
eval/rag-retrieval/
|
||||
cases/golden-cases.json Fixed retrieval golden cases
|
||||
seed-docs/*.md Canonical docs imported into the live KB for real-tool eval
|
||||
fixtures/*.json Saved retrieval fixtures for each case
|
||||
reports/baseline.json Machine-readable baseline report
|
||||
reports/baseline.md Human-readable baseline report
|
||||
reports/baseline-diff.* Optional diff reports
|
||||
reports/live-post-reindex.* Optional live acceptance reports
|
||||
cases/golden-cases.json Fixed queries + expectations
|
||||
seed-docs/*.md Canonical docs for live snapshot (kb_scope: rag-eval)
|
||||
fixtures/*.json Frozen lookupResult snapshots (+ searchMode meta)
|
||||
reports/baseline.json|md Last accepted offline report
|
||||
reports/baseline-diff.* Optional diff vs previous report
|
||||
```
|
||||
|
||||
## Seed Docs + Import/Reindex
|
||||
## Fixture shape (minimum)
|
||||
|
||||
The live-tool eval uses canonical seed documents so the real
|
||||
`LookupKnowledgeTool` can retrieve stable evidence from MySQL/Milvus instead of
|
||||
whatever ad hoc documents happen to exist in the local knowledge base.
|
||||
```text
|
||||
caseId
|
||||
query
|
||||
retrievedAt
|
||||
searchMode # hybrid | dense (required on newly generated fixtures)
|
||||
kbScope # e.g. rag-eval when generation used a scope
|
||||
lookupResult # found, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, …
|
||||
```
|
||||
|
||||
Seed documents live in:
|
||||
Offline eval **ignores unknown top-level meta** and does **not** full-JSON-compare.
|
||||
It asserts golden key fields only (doc/source, keywords, attempt, fallback, …). Raw scores are not pass criteria.
|
||||
|
||||
Older fixtures may omit `searchMode`; regenerate to attach meta.
|
||||
|
||||
## Seed docs + import
|
||||
|
||||
Live snapshot generation should use seed docs so results do not depend on ad-hoc local KB junk:
|
||||
|
||||
```text
|
||||
eval/rag-retrieval/seed-docs/*.md
|
||||
```
|
||||
|
||||
Each seed doc uses frontmatter fields that are propagated into vector metadata:
|
||||
Frontmatter example:
|
||||
|
||||
```yaml
|
||||
source: mysql-connection-pool
|
||||
@@ -40,40 +67,29 @@ breadcrumb: Database > MySQL > Connection Pool
|
||||
kb_scope: rag-eval
|
||||
```
|
||||
|
||||
Import or reindex the seed docs through the real upload pipeline:
|
||||
Import via real upload pipeline:
|
||||
|
||||
```powershell
|
||||
.\scripts\prepare_rag_eval_seed.ps1
|
||||
```
|
||||
|
||||
The script runs `RagEvalSeedImporterTest` with `rag.seed.enabled=true`. It
|
||||
deletes the existing document with the same `source`/`docId`, uploads the seed
|
||||
doc through `DocumentManagementService`, updates DB metadata and L0, and rebuilds
|
||||
Milvus chunks.
|
||||
Isolation:
|
||||
|
||||
`kb_scope` isolates eval data:
|
||||
- App default may leave `retrieval.kb-scope` empty (all docs).
|
||||
- Eval generation passes `-Dretrieval.kb-scope=rag-eval`.
|
||||
- Category-filter fallback retries without L0 category filter only; **kb_scope still applies**.
|
||||
|
||||
- default application config leaves `retrieval.kb-scope` empty, so legacy docs
|
||||
without `kb_scope` remain searchable;
|
||||
- eval scripts pass `-Dretrieval.kb-scope=rag-eval`, so L0 query hints and L1
|
||||
vector retrieval both use only the canonical eval seed docs;
|
||||
- the fallback retry skips only the L0 category filter, not the `kb_scope`
|
||||
boundary.
|
||||
Body is chunked/embedded; frontmatter feeds metadata/L0 (decoy keywords in frontmatter alone should not become dense content).
|
||||
|
||||
Frontmatter is not embedded as chunk content during upload. It feeds metadata,
|
||||
L0, and document enrichment; only the Markdown body is chunked and embedded.
|
||||
This keeps controlled L0 decoys from becoming semantically relevant just because
|
||||
their frontmatter keywords matched the query.
|
||||
Seeds must live in the **current hybrid collection schema** (`milvus.collection`, default `biz`). If the collection was recreated for BM25 hybrid, re-import seeds after rebuild.
|
||||
|
||||
## Run
|
||||
|
||||
From the repository root:
|
||||
## Offline run (no live stack)
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_retrieval.py
|
||||
```
|
||||
|
||||
Custom paths are also supported:
|
||||
Custom paths:
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_retrieval.py \
|
||||
@@ -83,83 +99,53 @@ python scripts/eval_rag_retrieval.py \
|
||||
--markdown-report eval/rag-retrieval/reports/baseline.md
|
||||
```
|
||||
|
||||
## Generate Fixtures From LookupKnowledgeTool
|
||||
|
||||
Use the snapshot generator when fixtures should reflect the real
|
||||
`LookupKnowledgeTool` pipeline:
|
||||
|
||||
```powershell
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1
|
||||
```
|
||||
|
||||
For the intended live loop, run seed import first:
|
||||
## Generate fixtures (live stack)
|
||||
|
||||
```powershell
|
||||
.\scripts\prepare_rag_eval_seed.ps1
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1
|
||||
python scripts\eval_rag_retrieval.py
|
||||
# default: SearchMode=hybrid, KbScope=rag-eval, then offline eval
|
||||
```
|
||||
|
||||
The script runs a Spring test harness:
|
||||
|
||||
```text
|
||||
mvn -q -Dtest=RagLookupSnapshotGeneratorTest -Drag.snapshot.enabled=true -Dretrieval.kb-scope=rag-eval -Dretrieval.vector-store.mode=spring test
|
||||
```
|
||||
|
||||
The generator reads `golden-cases.json`, injects the real `LookupKnowledgeTool`
|
||||
bean, calls `lookupKnowledge(query)` for each case, writes
|
||||
`fixtures/{caseId}.json`, and then runs `eval_rag_retrieval.py` unless
|
||||
`-SkipEval` is provided. It defaults to Spring AI VectorStore mode; pass
|
||||
`-VectorStoreMode sdk` only when intentionally comparing the legacy SDK path.
|
||||
|
||||
Custom paths are supported:
|
||||
Dense baseline snapshot (same seed, comparison only):
|
||||
|
||||
```powershell
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1 `
|
||||
-Cases eval\rag-retrieval\cases\golden-cases.json `
|
||||
-Fixtures eval\rag-retrieval\fixtures `
|
||||
-RetrievedAt 2026-07-06T00:00:00Z
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode dense -Fixtures eval\rag-retrieval\fixtures-dense -SkipEval
|
||||
```
|
||||
|
||||
The generator is disabled in normal test runs. It only executes when
|
||||
`rag.snapshot.enabled=true` is provided because it writes repository files and
|
||||
depends on the configured runtime retrieval stack.
|
||||
(Dual-directory comparison reports are optional / future; knife-1 only documents the override.)
|
||||
|
||||
If generated fixtures fail the offline baseline, treat that as a real alignment
|
||||
signal: either the golden expectations need to be adjusted to the current
|
||||
knowledge base, or the knowledge base/indexing path needs to be fixed.
|
||||
|
||||
## Modular RAG Contract
|
||||
|
||||
Fixtures must use the current `lookupResult` shape, which mirrors the
|
||||
`lookup_knowledge` output:
|
||||
Maven equivalent:
|
||||
|
||||
```text
|
||||
lookupResult.evidenceBlocks
|
||||
lookupResult.contextPack
|
||||
lookupResult.retrievalTrace
|
||||
lookupResult.rerankTrace
|
||||
mvn -q -Dtest=RagLookupSnapshotGeneratorTest \
|
||||
-Drag.snapshot.enabled=true \
|
||||
-Dretrieval.kb-scope=rag-eval \
|
||||
-Dretrieval.search.mode=hybrid \
|
||||
test
|
||||
```
|
||||
|
||||
Golden cases can assert both retrieval quality and pipeline behavior:
|
||||
Generator is **off** in normal tests; only runs when `rag.snapshot.enabled=true` (writes files).
|
||||
|
||||
If generated fixtures fail offline golden checks: either fix retrieval/index, or update golden/baseline **with an explicit reason** — do not silently overwrite.
|
||||
|
||||
## Golden assertions
|
||||
|
||||
Supported expectation fields include:
|
||||
|
||||
- `expectedSources` / `expectedDocIds`
|
||||
- `expectedBreadcrumbs`
|
||||
- `expectedKeywords`
|
||||
- `expectedBreadcrumbs` / `expectedKeywords`
|
||||
- `expectedSelectedAttempt`
|
||||
- `expectedFallbackReason`
|
||||
- `expectedFallbackReasons`
|
||||
- `expectedFallbackReason` / `expectedFallbackReasons`
|
||||
- `expectedEvidenceStatus`
|
||||
- `expectedContextSources`
|
||||
- `expectedRerankTopSource`
|
||||
|
||||
This lets the baseline catch regressions such as losing the expected evidence
|
||||
source, skipping context packing, changing the selected retrieval attempt, or
|
||||
breaking the filtered-vector to unfiltered-retry fallback.
|
||||
Catch regressions such as missing expected source, broken context pack sources, wrong selected attempt, or broken filtered → unfiltered retry.
|
||||
|
||||
## Baseline Diff
|
||||
**Note:** `relevance_level` is not a hard golden gate here (hybrid quality is rank-ordinal; see agent relevance-level doc).
|
||||
|
||||
To compare a freshly generated report against an existing baseline:
|
||||
## Baseline diff
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_retrieval.py \
|
||||
@@ -170,64 +156,24 @@ python scripts/eval_rag_retrieval.py \
|
||||
--diff-markdown-report eval/rag-retrieval/reports/baseline-diff.md
|
||||
```
|
||||
|
||||
The diff reports aggregate regressions and case-level changes for:
|
||||
Diff covers pass rate, recall@K, hit level, first expected rank, attempt, fallback, evidence status, rerank top source.
|
||||
Non-zero exit on case failure or regression in diff mode.
|
||||
|
||||
- pass rate, recall@K, strong hit rate, miss count
|
||||
- pass state
|
||||
- hit level
|
||||
- first expected rank
|
||||
- selected attempt
|
||||
- fallback reason
|
||||
- evidence status
|
||||
- rerank top source
|
||||
## Hit levels
|
||||
|
||||
The command exits non-zero when a case fails or the diff contains a regression.
|
||||
- `strong`: expected document found **and** breadcrumb or keyword coverage OK
|
||||
- `medium`: expected document found, coverage incomplete
|
||||
- `weak`: keyword hit without expected document
|
||||
- `miss`: neither
|
||||
|
||||
## Hit Levels
|
||||
`Recall@K` counts `strong` + `medium`.
|
||||
|
||||
- `strong`: expected document is found and breadcrumb or evidence keyword coverage is satisfied.
|
||||
- `medium`: expected document is found, but breadcrumb or keyword coverage is incomplete.
|
||||
- `weak`: expected evidence keyword is found, but expected document is missing.
|
||||
- `miss`: expected document and expected evidence are not found.
|
||||
## Optional live smoke (post-reindex)
|
||||
|
||||
`Recall@K` counts `strong` and `medium` as retrieved.
|
||||
|
||||
## Scope
|
||||
|
||||
This baseline runs fully offline and does not call MySQL, Redis, Milvus, an LLM,
|
||||
or the Spring Boot application. It is a regression harness for retrieval behavior,
|
||||
not a claim that live production retrieval accuracy is complete.
|
||||
|
||||
## Live Post-Reindex Acceptance
|
||||
|
||||
When embedding input changes, existing vectors do not update by themselves. For
|
||||
example, after adding `title` and `breadcrumb` to the embedding text, the live
|
||||
Milvus/Zilliz collection must be reindexed before retrieval can reflect that new
|
||||
semantic signal.
|
||||
|
||||
Use this optional live acceptance flow after the application is running and the
|
||||
knowledge base has been reindexed:
|
||||
After reindex, with app up:
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_live_acceptance.py
|
||||
python scripts/eval_rag_live_acceptance.py --base-url http://127.0.0.1:9900
|
||||
```
|
||||
|
||||
Custom service URL and output paths are supported:
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_live_acceptance.py \
|
||||
--base-url http://127.0.0.1:9900 \
|
||||
--json-report eval/rag-retrieval/reports/live-post-reindex.json \
|
||||
--markdown-report eval/rag-retrieval/reports/live-post-reindex.md
|
||||
```
|
||||
|
||||
The script calls:
|
||||
|
||||
```text
|
||||
GET /api/search/similar
|
||||
```
|
||||
|
||||
It writes JSON and Markdown reports with query, topK, result count, top
|
||||
results, breadcrumb, score labels, and raw response fields. This is a live
|
||||
smoke check for environment readiness and post-reindex behavior; it does not
|
||||
replace the deterministic offline baseline above.
|
||||
Calls `GET /api/search/similar`. Environment smoke only — does **not** replace offline baseline.
|
||||
|
||||
@@ -1,81 +1,122 @@
|
||||
{
|
||||
"caseId": "aiops-payment-latency-alert",
|
||||
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "payment-service-latency",
|
||||
"title": "Payment Service Latency Alert Playbook",
|
||||
"breadcrumb": "AIOps > Service Alerts > Payment Latency",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "For payment-service p95 latency alerts, check downstream dependency latency, thread pool saturation, gateway retries, and recent deployment changes.",
|
||||
"score": 0.84,
|
||||
"hitReasons": ["domain_match:+0.15", "entity_match:+0.20", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"source": "mysql-connection-pool",
|
||||
"title": "MySQL Connection Pool Troubleshooting",
|
||||
"breadcrumb": "Database > MySQL > Connection Pool",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "Database connection pool saturation can increase payment latency when checkout paths wait for connections.",
|
||||
"score": 0.68,
|
||||
"hitReasons": ["keyword_match:+0.10"]
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] Payment Service Latency Alert Playbook\nAIOps > Service Alerts > Payment Latency\nFor payment-service p95 latency alerts, check downstream dependency latency, thread pool saturation, gateway retries, and recent deployment changes.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 236,
|
||||
"includedSources": ["payment-service-latency", "mysql-connection-pool"],
|
||||
"omittedSources": []
|
||||
"caseId" : "aiops-payment-latency-alert",
|
||||
"query" : "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"retrievedAt" : "2026-07-28T06:54:51.843450400Z",
|
||||
"searchMode" : "hybrid",
|
||||
"kbScope" : "rag-eval",
|
||||
"lookupResult" : {
|
||||
"found" : true,
|
||||
"evidenceBlocks" : [ {
|
||||
"docId" : "payment-service-latency",
|
||||
"chunkIndex" : 2,
|
||||
"evidenceKey" : "payment-service-latency#chunk-2",
|
||||
"source" : "payment-service-latency",
|
||||
"title" : "Payment Latency",
|
||||
"breadcrumb" : "AIOps > Service Alerts > Payment Latency",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "### Payment Latency\n\nFor `HighLatency` alerts on `payment-service`, treat p95 latency as the primary symptom.\n\nDiagnosis steps:\n\n1. Confirm whether p95 latency is isolated to payment-service or shared across upstream callers.\n2. Compare payment-service latency with downstream dependency latency for gateway, risk, and order services.\n3. Check connection pool wait time, retry spikes, and timeout rates.\n4. If downstream dependency latency increased first, classify payment-service as affected rather than root cause.\n\nThe expected evidence terms are p95 latency, payment-service, and downstream dependency.",
|
||||
"score" : 0.032786883413791656,
|
||||
"hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"docId" : "aiops-alert-scope-control",
|
||||
"chunkIndex" : 1,
|
||||
"evidenceKey" : "aiops-alert-scope-control#chunk-1",
|
||||
"source" : "aiops-alert-scope-control",
|
||||
"title" : "Alert Scope Control",
|
||||
"breadcrumb" : "AIOps > Alert Scope Control",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "## Alert Scope Control\n\nWhen an AIOps request already includes an alert payload, the agent should diagnose that payload first.\nIt must not expand the task into unrelated active alerts unless the user asks for broad alert triage.\n\nScope rules:\n\n1. Treat the provided payload as the primary incident boundary.\n2. Use unrelated active alerts only as correlation evidence when they share service, dependency, time window, or trace context.\n3. Do not replace the requested alert with a louder but unrelated alert.\n\nThis runbook anchors payload, unrelated active alerts, and scope behavior.",
|
||||
"score" : 0.0320020467042923,
|
||||
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
}, {
|
||||
"docId" : "payment-service-latency",
|
||||
"chunkIndex" : 1,
|
||||
"evidenceKey" : "payment-service-latency#chunk-1",
|
||||
"source" : "payment-service-latency",
|
||||
"title" : "Service Alerts",
|
||||
"breadcrumb" : "AIOps > Service Alerts > Payment Latency",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "## Service Alerts",
|
||||
"score" : 0.0320020467042923,
|
||||
"hitReasons" : [ "semantic_rank:3", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"docId" : "aiops-alert-scope-control",
|
||||
"chunkIndex" : 0,
|
||||
"evidenceKey" : "aiops-alert-scope-control#chunk-0",
|
||||
"source" : "aiops-alert-scope-control",
|
||||
"title" : "AIOps",
|
||||
"breadcrumb" : "AIOps > Alert Scope Control",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "# AIOps",
|
||||
"score" : 0.015384615398943424,
|
||||
"hitReasons" : [ "semantic_rank:5", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
} ],
|
||||
"contextPack" : {
|
||||
"packedText" : "[Evidence 1]\nsource: payment-service-latency\ntitle: Payment Latency\nbreadcrumb: AIOps > Service Alerts > Payment Latency\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n### Payment Latency\n\nFor `HighLatency` alerts on `payment-service`, treat p95 latency as the primary symptom.\n\nDiagnosis steps:\n\n1. Confirm whether p95 latency is isolated to payment-service or shared across upstream callers.\n2. Compare payment-service latency with downstream dependency latency for gateway, risk, and order services.\n3. Check connection pool wait time, retry spikes, and timeout rates.\n4. If downstream dependency latency increased first, classify payment-service as affected rather than root cause.\n\nThe expected evidence terms are p95 latency, payment-service, and downstream dependency.\n\n[Evidence 2]\nsource: aiops-alert-scope-control\ntitle: Alert Scope Control\nbreadcrumb: AIOps > Alert Scope Control\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n## Alert Scope Control\n\nWhen an AIOps request already includes an alert payload, the agent should diagnose that payload first.\nIt must not expand the task into unrelated active alerts unless the user asks for broad alert triage.\n\nScope rules:\n\n1. Treat the provided payload as the primary incident boundary.\n2. Use unrelated active alerts only as correlation evidence when they share service, dependency, time window, or trace context.\n3. Do not replace the requested alert with a louder but unrelated alert.\n\nThis runbook anchors payload, unrelated active alerts, and scope behavior.\n\n[Evidence 3]\nsource: payment-service-latency\ntitle: Service Alerts\nbreadcrumb: AIOps > Service Alerts > Payment Latency\nlayer: L1\nreasons: semantic_rank:3, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n## Service Alerts\n\n[Evidence 4]\nsource: aiops-alert-scope-control\ntitle: AIOps\nbreadcrumb: AIOps > Alert Scope Control\nlayer: L1\nreasons: semantic_rank:5, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n# AIOps",
|
||||
"strategy" : "ranked_evidence_char_budget",
|
||||
"charBudget" : 4000,
|
||||
"usedChars" : 2108,
|
||||
"includedSources" : [ "payment-service-latency", "aiops-alert-scope-control", "payment-service-latency", "aiops-alert-scope-control" ],
|
||||
"omittedSources" : [ ]
|
||||
},
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"rewrittenQuery": "HighLatency payment-service p95 latency alert downstream dependency diagnosis",
|
||||
"categoryFilter": "AIOps",
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["AIOps"],
|
||||
"matched_keywords": ["p95 latency", "payment-service", "downstream dependency"],
|
||||
"entities": ["payment-service", "HighLatency"],
|
||||
"l0_titles": ["Payment Service Latency Alert Playbook"],
|
||||
"l0_match_count": 1
|
||||
"retrievalTrace" : {
|
||||
"originalQuery" : "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"rewrittenQuery" : "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"categoryFilter" : "aiops",
|
||||
"selectedAttempt" : "FILTERED_VECTOR",
|
||||
"fallbackReason" : null,
|
||||
"evidenceStatus" : "supported",
|
||||
"queryHints" : {
|
||||
"domains" : [ "aiops" ],
|
||||
"matched_keywords" : [ "HighLatency", "payment-service", "p95 latency" ],
|
||||
"entities" : [ "HighLatency", "payment-service", "p95 latency" ],
|
||||
"l0_titles" : [ "Payment Service Latency Alert" ],
|
||||
"l0_match_count" : 1
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "HighLatency payment-service p95 latency alert downstream dependency diagnosis",
|
||||
"categoryFilter": "AIOps",
|
||||
"candidateCount": 2,
|
||||
"usable": true,
|
||||
"durationMs": 11,
|
||||
"topScore": 0.84,
|
||||
"topSimilarity": 0.84
|
||||
}
|
||||
]
|
||||
"attempts" : [ {
|
||||
"name" : "FILTERED_VECTOR",
|
||||
"query" : "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"categoryFilter" : "aiops",
|
||||
"candidateCount" : 5,
|
||||
"usable" : true,
|
||||
"errorMessage" : null,
|
||||
"durationMs" : 969,
|
||||
"topScore" : 0.032786883413791656,
|
||||
"topSimilarity" : 0.7736010700464249
|
||||
} ]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "payment-service-latency",
|
||||
"baseScore": 0.84,
|
||||
"finalScore": 1.29,
|
||||
"boostReasons": ["domain_match:+0.15", "entity_match:+0.20", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"finalRank": 2,
|
||||
"source": "mysql-connection-pool",
|
||||
"baseScore": 0.68,
|
||||
"finalScore": 0.78,
|
||||
"boostReasons": ["keyword_match:+0.10"]
|
||||
}
|
||||
]
|
||||
}
|
||||
"rerankTrace" : {
|
||||
"items" : [ {
|
||||
"finalRank" : 1,
|
||||
"source" : "payment-service-latency",
|
||||
"baseScore" : 0.7736010700464249,
|
||||
"finalScore" : 0.7736010700464249,
|
||||
"boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 2,
|
||||
"source" : "aiops-alert-scope-control",
|
||||
"baseScore" : 0.4683566689491272,
|
||||
"finalScore" : 0.4683566689491272,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 3,
|
||||
"source" : "payment-service-latency",
|
||||
"baseScore" : 0.4981400966644287,
|
||||
"finalScore" : 0.4981400966644287,
|
||||
"boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 4,
|
||||
"source" : "aiops-alert-scope-control",
|
||||
"baseScore" : 0.3814886808395386,
|
||||
"finalScore" : 0.3814886808395386,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
} ]
|
||||
},
|
||||
"evidenceCandidateCount" : 5,
|
||||
"evidenceBlockCount" : 4,
|
||||
"relevanceLevel" : "PRECISE",
|
||||
"completenessHint" : "知识库中不存在比上述结果更精准的文档",
|
||||
"retrievedDomainsThisSession" : null,
|
||||
"message" : null
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1,65 +1,122 @@
|
||||
{
|
||||
"caseId": "aiops-prometheus-alert-scope",
|
||||
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "aiops-alert-scope-control",
|
||||
"title": "AIOps Alert Scope Control",
|
||||
"breadcrumb": "AIOps > Alert Scope Control",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "When payload mode is used, diagnose the input alert payload and do not expand unrelated active alerts into the main diagnosis scope.",
|
||||
"score": 0.88,
|
||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] AIOps Alert Scope Control\nAIOps > Alert Scope Control\nWhen payload mode is used, diagnose the input alert payload and do not expand unrelated active alerts into the main diagnosis scope.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 188,
|
||||
"includedSources": ["aiops-alert-scope-control"],
|
||||
"omittedSources": []
|
||||
"caseId" : "aiops-prometheus-alert-scope",
|
||||
"query" : "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"retrievedAt" : "2026-07-28T06:54:51.843450400Z",
|
||||
"searchMode" : "hybrid",
|
||||
"kbScope" : "rag-eval",
|
||||
"lookupResult" : {
|
||||
"found" : true,
|
||||
"evidenceBlocks" : [ {
|
||||
"docId" : "aiops-alert-scope-control",
|
||||
"chunkIndex" : 1,
|
||||
"evidenceKey" : "aiops-alert-scope-control#chunk-1",
|
||||
"source" : "aiops-alert-scope-control",
|
||||
"title" : "Alert Scope Control",
|
||||
"breadcrumb" : "AIOps > Alert Scope Control",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "## Alert Scope Control\n\nWhen an AIOps request already includes an alert payload, the agent should diagnose that payload first.\nIt must not expand the task into unrelated active alerts unless the user asks for broad alert triage.\n\nScope rules:\n\n1. Treat the provided payload as the primary incident boundary.\n2. Use unrelated active alerts only as correlation evidence when they share service, dependency, time window, or trace context.\n3. Do not replace the requested alert with a louder but unrelated alert.\n\nThis runbook anchors payload, unrelated active alerts, and scope behavior.",
|
||||
"score" : 0.032786883413791656,
|
||||
"hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"docId" : "payment-service-latency",
|
||||
"chunkIndex" : 1,
|
||||
"evidenceKey" : "payment-service-latency#chunk-1",
|
||||
"source" : "payment-service-latency",
|
||||
"title" : "Service Alerts",
|
||||
"breadcrumb" : "AIOps > Service Alerts > Payment Latency",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "## Service Alerts",
|
||||
"score" : 0.0320020467042923,
|
||||
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
}, {
|
||||
"docId" : "payment-service-latency",
|
||||
"chunkIndex" : 2,
|
||||
"evidenceKey" : "payment-service-latency#chunk-2",
|
||||
"source" : "payment-service-latency",
|
||||
"title" : "Payment Latency",
|
||||
"breadcrumb" : "AIOps > Service Alerts > Payment Latency",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "### Payment Latency\n\nFor `HighLatency` alerts on `payment-service`, treat p95 latency as the primary symptom.\n\nDiagnosis steps:\n\n1. Confirm whether p95 latency is isolated to payment-service or shared across upstream callers.\n2. Compare payment-service latency with downstream dependency latency for gateway, risk, and order services.\n3. Check connection pool wait time, retry spikes, and timeout rates.\n4. If downstream dependency latency increased first, classify payment-service as affected rather than root cause.\n\nThe expected evidence terms are p95 latency, payment-service, and downstream dependency.",
|
||||
"score" : 0.0320020467042923,
|
||||
"hitReasons" : [ "semantic_rank:3", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
}, {
|
||||
"docId" : "aiops-alert-scope-control",
|
||||
"chunkIndex" : 0,
|
||||
"evidenceKey" : "aiops-alert-scope-control#chunk-0",
|
||||
"source" : "aiops-alert-scope-control",
|
||||
"title" : "AIOps",
|
||||
"breadcrumb" : "AIOps > Alert Scope Control",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "# AIOps",
|
||||
"score" : 0.03076923079788685,
|
||||
"hitReasons" : [ "semantic_rank:5", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
} ],
|
||||
"contextPack" : {
|
||||
"packedText" : "[Evidence 1]\nsource: aiops-alert-scope-control\ntitle: Alert Scope Control\nbreadcrumb: AIOps > Alert Scope Control\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n## Alert Scope Control\n\nWhen an AIOps request already includes an alert payload, the agent should diagnose that payload first.\nIt must not expand the task into unrelated active alerts unless the user asks for broad alert triage.\n\nScope rules:\n\n1. Treat the provided payload as the primary incident boundary.\n2. Use unrelated active alerts only as correlation evidence when they share service, dependency, time window, or trace context.\n3. Do not replace the requested alert with a louder but unrelated alert.\n\nThis runbook anchors payload, unrelated active alerts, and scope behavior.\n\n[Evidence 2]\nsource: payment-service-latency\ntitle: Service Alerts\nbreadcrumb: AIOps > Service Alerts > Payment Latency\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n## Service Alerts\n\n[Evidence 3]\nsource: payment-service-latency\ntitle: Payment Latency\nbreadcrumb: AIOps > Service Alerts > Payment Latency\nlayer: L1\nreasons: semantic_rank:3, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n### Payment Latency\n\nFor `HighLatency` alerts on `payment-service`, treat p95 latency as the primary symptom.\n\nDiagnosis steps:\n\n1. Confirm whether p95 latency is isolated to payment-service or shared across upstream callers.\n2. Compare payment-service latency with downstream dependency latency for gateway, risk, and order services.\n3. Check connection pool wait time, retry spikes, and timeout rates.\n4. If downstream dependency latency increased first, classify payment-service as affected rather than root cause.\n\nThe expected evidence terms are p95 latency, payment-service, and downstream dependency.\n\n[Evidence 4]\nsource: aiops-alert-scope-control\ntitle: AIOps\nbreadcrumb: AIOps > Alert Scope Control\nlayer: L1\nreasons: semantic_rank:5, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n# AIOps",
|
||||
"strategy" : "ranked_evidence_char_budget",
|
||||
"charBudget" : 4000,
|
||||
"usedChars" : 2069,
|
||||
"includedSources" : [ "aiops-alert-scope-control", "payment-service-latency", "payment-service-latency", "aiops-alert-scope-control" ],
|
||||
"omittedSources" : [ ]
|
||||
},
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"rewrittenQuery": "AIOps alert payload scope unrelated active alerts diagnosis",
|
||||
"categoryFilter": "AIOps",
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["AIOps"],
|
||||
"matched_keywords": ["payload", "unrelated active alerts", "scope"],
|
||||
"entities": ["alert payload"],
|
||||
"l0_titles": ["AIOps Alert Scope Control"],
|
||||
"l0_match_count": 1
|
||||
"retrievalTrace" : {
|
||||
"originalQuery" : "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"rewrittenQuery" : "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"categoryFilter" : "aiops",
|
||||
"selectedAttempt" : "FILTERED_VECTOR",
|
||||
"fallbackReason" : null,
|
||||
"evidenceStatus" : "supported",
|
||||
"queryHints" : {
|
||||
"domains" : [ "aiops" ],
|
||||
"matched_keywords" : [ "alert payload", "unrelated active alerts" ],
|
||||
"entities" : [ "alert payload", "unrelated active alerts" ],
|
||||
"l0_titles" : [ "AIOps Alert Scope Control" ],
|
||||
"l0_match_count" : 1
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "AIOps alert payload scope unrelated active alerts diagnosis",
|
||||
"categoryFilter": "AIOps",
|
||||
"candidateCount": 1,
|
||||
"usable": true,
|
||||
"durationMs": 8,
|
||||
"topScore": 0.88,
|
||||
"topSimilarity": 0.88
|
||||
}
|
||||
]
|
||||
"attempts" : [ {
|
||||
"name" : "FILTERED_VECTOR",
|
||||
"query" : "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"categoryFilter" : "aiops",
|
||||
"candidateCount" : 5,
|
||||
"usable" : true,
|
||||
"errorMessage" : null,
|
||||
"durationMs" : 1623,
|
||||
"topScore" : 0.032786883413791656,
|
||||
"topSimilarity" : 0.7561411112546921
|
||||
} ]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "aiops-alert-scope-control",
|
||||
"baseScore": 0.88,
|
||||
"finalScore": 1.13,
|
||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
}
|
||||
]
|
||||
}
|
||||
"rerankTrace" : {
|
||||
"items" : [ {
|
||||
"finalRank" : 1,
|
||||
"source" : "aiops-alert-scope-control",
|
||||
"baseScore" : 0.7561411112546921,
|
||||
"finalScore" : 0.7561411112546921,
|
||||
"boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 2,
|
||||
"source" : "payment-service-latency",
|
||||
"baseScore" : 0.503810703754425,
|
||||
"finalScore" : 0.503810703754425,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 3,
|
||||
"source" : "payment-service-latency",
|
||||
"baseScore" : 0.5772626996040344,
|
||||
"finalScore" : 0.5772626996040344,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 4,
|
||||
"source" : "aiops-alert-scope-control",
|
||||
"baseScore" : 0.36770421266555786,
|
||||
"finalScore" : 0.36770421266555786,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
} ]
|
||||
},
|
||||
"evidenceCandidateCount" : 5,
|
||||
"evidenceBlockCount" : 4,
|
||||
"relevanceLevel" : "PRECISE",
|
||||
"completenessHint" : "知识库中不存在比上述结果更精准的文档",
|
||||
"retrievedDomainsThisSession" : null,
|
||||
"message" : null
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1,81 +1,88 @@
|
||||
{
|
||||
"caseId": "chat-diagnosis-flow",
|
||||
"query": "What is the standard troubleshooting flow for an application incident?",
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "incident-diagnosis-flow",
|
||||
"title": "Incident Diagnosis Flow",
|
||||
"breadcrumb": "AIOps > Diagnosis Flow",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "The standard flow is to collect evidence, identify the suspected fault domain, verify the hypothesis, apply remediation, and confirm recovery.",
|
||||
"score": 0.82,
|
||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"source": "rag-chunk-context-reconstruction",
|
||||
"title": "RAG Chunk Context Reconstruction",
|
||||
"breadcrumb": "RAG > Chunking > Context Reconstruction",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "Long sections may require neighbor chunk expansion and breadcrumb-aware packing.",
|
||||
"score": 0.55,
|
||||
"hitReasons": []
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] Incident Diagnosis Flow\nAIOps > Diagnosis Flow\nThe standard flow is to collect evidence, identify the suspected fault domain, verify the hypothesis, apply remediation, and confirm recovery.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 192,
|
||||
"includedSources": ["incident-diagnosis-flow", "rag-chunk-context-reconstruction"],
|
||||
"omittedSources": []
|
||||
"caseId" : "chat-diagnosis-flow",
|
||||
"query" : "What is the standard troubleshooting flow for an application incident?",
|
||||
"retrievedAt" : "2026-07-28T06:54:51.843450400Z",
|
||||
"searchMode" : "hybrid",
|
||||
"kbScope" : "rag-eval",
|
||||
"lookupResult" : {
|
||||
"found" : true,
|
||||
"evidenceBlocks" : [ {
|
||||
"docId" : "incident-diagnosis-flow",
|
||||
"chunkIndex" : 1,
|
||||
"evidenceKey" : "incident-diagnosis-flow#chunk-1",
|
||||
"source" : "incident-diagnosis-flow",
|
||||
"title" : "Diagnosis Flow",
|
||||
"breadcrumb" : "AIOps > Diagnosis Flow",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "## Diagnosis Flow\n\nThe standard troubleshooting flow is evidence first, hypothesis second, remediation last.\n\nRecommended sequence:\n\n1. Collect evidence from alerts, metrics, logs, traces, deployments, and recent configuration changes.\n2. Define a small hypothesis that explains the observed symptoms.\n3. Verify the hypothesis with a targeted metric, log query, or reproduction step.\n4. Choose remediation that directly addresses the verified cause.\n5. Record the outcome and the evidence used to make the decision.\n\nDo not skip collect evidence, verify, and remediation ordering during an application incident.",
|
||||
"score" : 0.032786883413791656,
|
||||
"hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"docId" : "incident-diagnosis-flow",
|
||||
"chunkIndex" : 0,
|
||||
"evidenceKey" : "incident-diagnosis-flow#chunk-0",
|
||||
"source" : "incident-diagnosis-flow",
|
||||
"title" : "AIOps",
|
||||
"breadcrumb" : "AIOps > Diagnosis Flow",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "# AIOps",
|
||||
"score" : 0.016129031777381897,
|
||||
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
} ],
|
||||
"contextPack" : {
|
||||
"packedText" : "[Evidence 1]\nsource: incident-diagnosis-flow\ntitle: Diagnosis Flow\nbreadcrumb: AIOps > Diagnosis Flow\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n## Diagnosis Flow\n\nThe standard troubleshooting flow is evidence first, hypothesis second, remediation last.\n\nRecommended sequence:\n\n1. Collect evidence from alerts, metrics, logs, traces, deployments, and recent configuration changes.\n2. Define a small hypothesis that explains the observed symptoms.\n3. Verify the hypothesis with a targeted metric, log query, or reproduction step.\n4. Choose remediation that directly addresses the verified cause.\n5. Record the outcome and the evidence used to make the decision.\n\nDo not skip collect evidence, verify, and remediation ordering during an application incident.\n\n[Evidence 2]\nsource: incident-diagnosis-flow\ntitle: AIOps\nbreadcrumb: AIOps > Diagnosis Flow\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n# AIOps",
|
||||
"strategy" : "ranked_evidence_char_budget",
|
||||
"charBudget" : 4000,
|
||||
"usedChars" : 1032,
|
||||
"includedSources" : [ "incident-diagnosis-flow", "incident-diagnosis-flow" ],
|
||||
"omittedSources" : [ ]
|
||||
},
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "What is the standard troubleshooting flow for an application incident?",
|
||||
"rewrittenQuery": "standard application incident troubleshooting flow collect evidence verify remediation",
|
||||
"categoryFilter": "AIOps",
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["AIOps"],
|
||||
"matched_keywords": ["collect evidence", "verify", "remediation"],
|
||||
"entities": ["application incident"],
|
||||
"l0_titles": ["Incident Diagnosis Flow"],
|
||||
"l0_match_count": 1
|
||||
"retrievalTrace" : {
|
||||
"originalQuery" : "What is the standard troubleshooting flow for an application incident?",
|
||||
"rewrittenQuery" : "What is the standard troubleshooting flow for an application incident?",
|
||||
"categoryFilter" : "ops",
|
||||
"selectedAttempt" : "FILTERED_VECTOR",
|
||||
"fallbackReason" : null,
|
||||
"evidenceStatus" : "supported",
|
||||
"queryHints" : {
|
||||
"domains" : [ "ops" ],
|
||||
"matched_keywords" : [ "standard troubleshooting flow", "application incident" ],
|
||||
"entities" : [ "standard troubleshooting flow", "application incident" ],
|
||||
"l0_titles" : [ "Incident Diagnosis Flow" ],
|
||||
"l0_match_count" : 1
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "standard application incident troubleshooting flow collect evidence verify remediation",
|
||||
"categoryFilter": "AIOps",
|
||||
"candidateCount": 2,
|
||||
"usable": true,
|
||||
"durationMs": 10,
|
||||
"topScore": 0.82,
|
||||
"topSimilarity": 0.82
|
||||
}
|
||||
]
|
||||
"attempts" : [ {
|
||||
"name" : "FILTERED_VECTOR",
|
||||
"query" : "What is the standard troubleshooting flow for an application incident?",
|
||||
"categoryFilter" : "ops",
|
||||
"candidateCount" : 2,
|
||||
"usable" : true,
|
||||
"errorMessage" : null,
|
||||
"durationMs" : 1540,
|
||||
"topScore" : 0.032786883413791656,
|
||||
"topSimilarity" : 0.6828859150409698
|
||||
} ]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "incident-diagnosis-flow",
|
||||
"baseScore": 0.82,
|
||||
"finalScore": 1.07,
|
||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"finalRank": 2,
|
||||
"source": "rag-chunk-context-reconstruction",
|
||||
"baseScore": 0.55,
|
||||
"finalScore": 0.55,
|
||||
"boostReasons": []
|
||||
}
|
||||
]
|
||||
}
|
||||
"rerankTrace" : {
|
||||
"items" : [ {
|
||||
"finalRank" : 1,
|
||||
"source" : "incident-diagnosis-flow",
|
||||
"baseScore" : 0.6828859150409698,
|
||||
"finalScore" : 0.6828859150409698,
|
||||
"boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 2,
|
||||
"source" : "incident-diagnosis-flow",
|
||||
"baseScore" : 0.3166210651397705,
|
||||
"finalScore" : 0.3166210651397705,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
} ]
|
||||
},
|
||||
"evidenceCandidateCount" : 2,
|
||||
"evidenceBlockCount" : 2,
|
||||
"relevanceLevel" : "REFERENCE",
|
||||
"completenessHint" : "当前结果为相关参考,如需更精准信息请明确缺少的具体维度",
|
||||
"retrievedDomainsThisSession" : null,
|
||||
"message" : null
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1,81 +1,122 @@
|
||||
{
|
||||
"caseId": "chat-l0-domain-hint",
|
||||
"query": "Should L0 keyword matching decide the final retrieval result?",
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "rag-l0-domain-entity-hint",
|
||||
"title": "RAG L0 Domain Entity Hint",
|
||||
"breadcrumb": "RAG > L0 > Domain Entity Hint",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "L0 should be retained as a domain detector, entity extractor, metadata filter generator, and explainability signal, not as the final retrieval decision.",
|
||||
"score": 0.88,
|
||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"source": "rag-l0-l1-fusion-ranking",
|
||||
"title": "RAG L0 L1 Fusion Ranking",
|
||||
"breadcrumb": "RAG > Ranking > Fusion",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "L0 and L1 candidates should eventually be fused rather than handled as an early-return branch.",
|
||||
"score": 0.75,
|
||||
"hitReasons": ["domain_match:+0.15"]
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] RAG L0 Domain Entity Hint\nRAG > L0 > Domain Entity Hint\nL0 should be retained as a domain detector, entity extractor, metadata filter generator, and explainability signal, not as the final retrieval decision.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 219,
|
||||
"includedSources": ["rag-l0-domain-entity-hint", "rag-l0-l1-fusion-ranking"],
|
||||
"omittedSources": []
|
||||
"caseId" : "chat-l0-domain-hint",
|
||||
"query" : "Should L0 keyword matching decide the final retrieval result?",
|
||||
"retrievedAt" : "2026-07-28T06:54:51.843450400Z",
|
||||
"searchMode" : "hybrid",
|
||||
"kbScope" : "rag-eval",
|
||||
"lookupResult" : {
|
||||
"found" : true,
|
||||
"evidenceBlocks" : [ {
|
||||
"docId" : "rag-l0-domain-entity-hint",
|
||||
"chunkIndex" : 2,
|
||||
"evidenceKey" : "rag-l0-domain-entity-hint#chunk-2",
|
||||
"source" : "rag-l0-domain-entity-hint",
|
||||
"title" : "Domain Entity Hint",
|
||||
"breadcrumb" : "RAG > L0 > Domain Entity Hint",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "### Domain Entity Hint\n\nL0 keyword matching should not decide the final retrieval result.\nIn the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.\n\nThe output can provide:\n\n1. Candidate domain hints.\n2. Matched entities and keywords.\n3. An optional metadata filter for the first vector retrieval attempt.\n\nFinal evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.\nThe important terms are domain detector, entity extractor, and metadata filter.",
|
||||
"score" : 0.032786883413791656,
|
||||
"hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"docId" : "rag-chunk-context-reconstruction",
|
||||
"chunkIndex" : 2,
|
||||
"evidenceKey" : "rag-chunk-context-reconstruction#chunk-2",
|
||||
"source" : "rag-chunk-context-reconstruction",
|
||||
"title" : "Context Reconstruction",
|
||||
"breadcrumb" : "RAG > Chunking > Context Reconstruction",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "### Context Reconstruction\n\nWhen a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.\n\nRecommended behavior:\n\n1. Store the breadcrumb with every chunk.\n2. Preserve the same section identity across adjacent chunks.\n3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.\n4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.\n\nThe key concepts are neighbor chunk, same section, and breadcrumb.",
|
||||
"score" : 0.032258063554763794,
|
||||
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
}, {
|
||||
"docId" : "rag-l0-domain-entity-hint",
|
||||
"chunkIndex" : 1,
|
||||
"evidenceKey" : "rag-l0-domain-entity-hint#chunk-1",
|
||||
"source" : "rag-l0-domain-entity-hint",
|
||||
"title" : "L0",
|
||||
"breadcrumb" : "RAG > L0 > Domain Entity Hint",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "## L0",
|
||||
"score" : 0.0317460335791111,
|
||||
"hitReasons" : [ "semantic_rank:3", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
}, {
|
||||
"docId" : "rag-chunk-context-reconstruction",
|
||||
"chunkIndex" : 1,
|
||||
"evidenceKey" : "rag-chunk-context-reconstruction#chunk-1",
|
||||
"source" : "rag-chunk-context-reconstruction",
|
||||
"title" : "Chunking",
|
||||
"breadcrumb" : "RAG > Chunking > Context Reconstruction",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "## Chunking",
|
||||
"score" : 0.015625,
|
||||
"hitReasons" : [ "semantic_rank:4", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
} ],
|
||||
"contextPack" : {
|
||||
"packedText" : "[Evidence 1]\nsource: rag-l0-domain-entity-hint\ntitle: Domain Entity Hint\nbreadcrumb: RAG > L0 > Domain Entity Hint\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n### Domain Entity Hint\n\nL0 keyword matching should not decide the final retrieval result.\nIn the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.\n\nThe output can provide:\n\n1. Candidate domain hints.\n2. Matched entities and keywords.\n3. An optional metadata filter for the first vector retrieval attempt.\n\nFinal evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.\nThe important terms are domain detector, entity extractor, and metadata filter.\n\n[Evidence 2]\nsource: rag-chunk-context-reconstruction\ntitle: Context Reconstruction\nbreadcrumb: RAG > Chunking > Context Reconstruction\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n### Context Reconstruction\n\nWhen a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.\n\nRecommended behavior:\n\n1. Store the breadcrumb with every chunk.\n2. Preserve the same section identity across adjacent chunks.\n3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.\n4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.\n\nThe key concepts are neighbor chunk, same section, and breadcrumb.\n\n[Evidence 3]\nsource: rag-l0-domain-entity-hint\ntitle: L0\nbreadcrumb: RAG > L0 > Domain Entity Hint\nlayer: L1\nreasons: semantic_rank:3, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n## L0\n\n[Evidence 4]\nsource: rag-chunk-context-reconstruction\ntitle: Chunking\nbreadcrumb: RAG > Chunking > Context Reconstruction\nlayer: L1\nreasons: semantic_rank:4, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n## Chunking",
|
||||
"strategy" : "ranked_evidence_char_budget",
|
||||
"charBudget" : 4000,
|
||||
"usedChars" : 1965,
|
||||
"includedSources" : [ "rag-l0-domain-entity-hint", "rag-chunk-context-reconstruction", "rag-l0-domain-entity-hint", "rag-chunk-context-reconstruction" ],
|
||||
"omittedSources" : [ ]
|
||||
},
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "Should L0 keyword matching decide the final retrieval result?",
|
||||
"rewrittenQuery": "RAG L0 keyword matching domain entity hint final retrieval decision",
|
||||
"categoryFilter": "RAG",
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["RAG"],
|
||||
"matched_keywords": ["domain detector", "entity extractor", "metadata filter"],
|
||||
"entities": ["L0"],
|
||||
"l0_titles": ["RAG L0 Domain Entity Hint"],
|
||||
"l0_match_count": 1
|
||||
"retrievalTrace" : {
|
||||
"originalQuery" : "Should L0 keyword matching decide the final retrieval result?",
|
||||
"rewrittenQuery" : "Should L0 keyword matching decide the final retrieval result?",
|
||||
"categoryFilter" : "rag",
|
||||
"selectedAttempt" : "FILTERED_VECTOR",
|
||||
"fallbackReason" : null,
|
||||
"evidenceStatus" : "supported",
|
||||
"queryHints" : {
|
||||
"domains" : [ "rag" ],
|
||||
"matched_keywords" : [ "L0 keyword matching", "final retrieval result" ],
|
||||
"entities" : [ "L0 keyword matching", "final retrieval result" ],
|
||||
"l0_titles" : [ "RAG L0 Domain Entity Hint" ],
|
||||
"l0_match_count" : 1
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "RAG L0 keyword matching domain entity hint final retrieval decision",
|
||||
"categoryFilter": "RAG",
|
||||
"candidateCount": 2,
|
||||
"usable": true,
|
||||
"durationMs": 9,
|
||||
"topScore": 0.88,
|
||||
"topSimilarity": 0.88
|
||||
}
|
||||
]
|
||||
"attempts" : [ {
|
||||
"name" : "FILTERED_VECTOR",
|
||||
"query" : "Should L0 keyword matching decide the final retrieval result?",
|
||||
"categoryFilter" : "rag",
|
||||
"candidateCount" : 6,
|
||||
"usable" : true,
|
||||
"errorMessage" : null,
|
||||
"durationMs" : 850,
|
||||
"topScore" : 0.032786883413791656,
|
||||
"topSimilarity" : 0.6438122987747192
|
||||
} ]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "rag-l0-domain-entity-hint",
|
||||
"baseScore": 0.88,
|
||||
"finalScore": 1.13,
|
||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"finalRank": 2,
|
||||
"source": "rag-l0-l1-fusion-ranking",
|
||||
"baseScore": 0.75,
|
||||
"finalScore": 0.9,
|
||||
"boostReasons": ["domain_match:+0.15"]
|
||||
}
|
||||
]
|
||||
}
|
||||
"rerankTrace" : {
|
||||
"items" : [ {
|
||||
"finalRank" : 1,
|
||||
"source" : "rag-l0-domain-entity-hint",
|
||||
"baseScore" : 0.6438122987747192,
|
||||
"finalScore" : 0.6438122987747192,
|
||||
"boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 2,
|
||||
"source" : "rag-chunk-context-reconstruction",
|
||||
"baseScore" : 0.41426247358322144,
|
||||
"finalScore" : 0.41426247358322144,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 3,
|
||||
"source" : "rag-l0-domain-entity-hint",
|
||||
"baseScore" : 0.3964804410934448,
|
||||
"finalScore" : 0.3964804410934448,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 4,
|
||||
"source" : "rag-chunk-context-reconstruction",
|
||||
"baseScore" : 0.28678786754608154,
|
||||
"finalScore" : 0.28678786754608154,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
} ]
|
||||
},
|
||||
"evidenceCandidateCount" : 6,
|
||||
"evidenceBlockCount" : 4,
|
||||
"relevanceLevel" : "REFERENCE",
|
||||
"completenessHint" : "当前结果为相关参考,如需更精准信息请明确缺少的具体维度",
|
||||
"retrievedDomainsThisSession" : null,
|
||||
"message" : null
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1,91 +1,149 @@
|
||||
{
|
||||
"caseId": "chat-l0-filter-fallback",
|
||||
"query": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "rag-l0-filter-fallback",
|
||||
"title": "RAG L0 Filter Fallback",
|
||||
"breadcrumb": "RAG > Fallback > Unfiltered Retry",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "When filtered vector retrieval is low quality, skip the L0 filter and run an unfiltered vector retry with the raw query before returning no evidence.",
|
||||
"score": 0.83,
|
||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"source": "rag-l0-domain-entity-hint",
|
||||
"title": "RAG L0 Domain Entity Hint",
|
||||
"breadcrumb": "RAG > L0 > Domain Entity Hint",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "L0 supplies hints for metadata filtering and explanation, but it should not be treated as final fact evidence.",
|
||||
"score": 0.66,
|
||||
"hitReasons": ["domain_match:+0.15"]
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] RAG L0 Filter Fallback\nRAG > Fallback > Unfiltered Retry\nWhen filtered vector retrieval is low quality, skip the L0 filter and run an unfiltered vector retry with the raw query before returning no evidence.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 214,
|
||||
"includedSources": ["rag-l0-filter-fallback", "rag-l0-domain-entity-hint"],
|
||||
"omittedSources": []
|
||||
"caseId" : "chat-l0-filter-fallback",
|
||||
"query" : "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"retrievedAt" : "2026-07-28T06:54:51.843450400Z",
|
||||
"searchMode" : "hybrid",
|
||||
"kbScope" : "rag-eval",
|
||||
"lookupResult" : {
|
||||
"found" : true,
|
||||
"evidenceBlocks" : [ {
|
||||
"docId" : "rag-l0-filter-fallback",
|
||||
"chunkIndex" : 2,
|
||||
"evidenceKey" : "rag-l0-filter-fallback#chunk-2",
|
||||
"source" : "rag-l0-filter-fallback",
|
||||
"title" : "Unfiltered Retry",
|
||||
"breadcrumb" : "RAG > Fallback > Unfiltered Retry",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "### Unfiltered Retry\n\nIf the first vector search is over-constrained by an L0 metadata filter and returns low quality evidence,\nthe retriever should skip the L0 filter and run an unfiltered vector retry with the original query.\n\nThe fallback reason should be `filtered_vector_low_quality` when the filtered candidate exists but is below the\nreference threshold. If there is no usable evidence at all, use `filtered_vector_no_evidence`.\n\nThis document is the expected evidence for skip the L0 filter, unfiltered vector retry, and low quality behavior.",
|
||||
"score" : 0.032786883413791656,
|
||||
"hitReasons" : [ "semantic_rank:1", "attempt:UNFILTERED_VECTOR_RETRY", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"docId" : "rag-l0-filter-decoy",
|
||||
"chunkIndex" : 1,
|
||||
"evidenceKey" : "rag-l0-filter-decoy#chunk-1",
|
||||
"source" : "rag-l0-filter-decoy",
|
||||
"title" : "Approval Window",
|
||||
"breadcrumb" : "RAG > Fallback > Decoy",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "## Approval Window\n\nThis document describes an unrelated release calendar approval window.\nIt intentionally avoids the real fallback instructions so the filtered retrieval\nattempt is low quality and the retriever must retry without the L0 category filter.",
|
||||
"score" : 0.0320020467042923,
|
||||
"hitReasons" : [ "semantic_rank:2", "attempt:UNFILTERED_VECTOR_RETRY", "l0_domain_overlap" ]
|
||||
}, {
|
||||
"docId" : "rag-l0-domain-entity-hint",
|
||||
"chunkIndex" : 2,
|
||||
"evidenceKey" : "rag-l0-domain-entity-hint#chunk-2",
|
||||
"source" : "rag-l0-domain-entity-hint",
|
||||
"title" : "Domain Entity Hint",
|
||||
"breadcrumb" : "RAG > L0 > Domain Entity Hint",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "### Domain Entity Hint\n\nL0 keyword matching should not decide the final retrieval result.\nIn the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.\n\nThe output can provide:\n\n1. Candidate domain hints.\n2. Matched entities and keywords.\n3. An optional metadata filter for the first vector retrieval attempt.\n\nFinal evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.\nThe important terms are domain detector, entity extractor, and metadata filter.",
|
||||
"score" : 0.0320020467042923,
|
||||
"hitReasons" : [ "semantic_rank:3", "attempt:UNFILTERED_VECTOR_RETRY" ]
|
||||
}, {
|
||||
"docId" : "rag-l0-domain-entity-hint",
|
||||
"chunkIndex" : 1,
|
||||
"evidenceKey" : "rag-l0-domain-entity-hint#chunk-1",
|
||||
"source" : "rag-l0-domain-entity-hint",
|
||||
"title" : "L0",
|
||||
"breadcrumb" : "RAG > L0 > Domain Entity Hint",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "## L0",
|
||||
"score" : 0.03125,
|
||||
"hitReasons" : [ "semantic_rank:4", "attempt:UNFILTERED_VECTOR_RETRY" ]
|
||||
}, {
|
||||
"docId" : "rag-chunk-context-reconstruction",
|
||||
"chunkIndex" : 2,
|
||||
"evidenceKey" : "rag-chunk-context-reconstruction#chunk-2",
|
||||
"source" : "rag-chunk-context-reconstruction",
|
||||
"title" : "Context Reconstruction",
|
||||
"breadcrumb" : "RAG > Chunking > Context Reconstruction",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "### Context Reconstruction\n\nWhen a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.\n\nRecommended behavior:\n\n1. Store the breadcrumb with every chunk.\n2. Preserve the same section identity across adjacent chunks.\n3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.\n4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.\n\nThe key concepts are neighbor chunk, same section, and breadcrumb.",
|
||||
"score" : 0.03053613007068634,
|
||||
"hitReasons" : [ "semantic_rank:5", "attempt:UNFILTERED_VECTOR_RETRY" ]
|
||||
} ],
|
||||
"contextPack" : {
|
||||
"packedText" : "[Evidence 1]\nsource: rag-l0-filter-fallback\ntitle: Unfiltered Retry\nbreadcrumb: RAG > Fallback > Unfiltered Retry\nlayer: L1\nreasons: semantic_rank:1, attempt:UNFILTERED_VECTOR_RETRY, l0_entity_overlap, l0_keyword_overlap\ncontent:\n### Unfiltered Retry\n\nIf the first vector search is over-constrained by an L0 metadata filter and returns low quality evidence,\nthe retriever should skip the L0 filter and run an unfiltered vector retry with the original query.\n\nThe fallback reason should be `filtered_vector_low_quality` when the filtered candidate exists but is below the\nreference threshold. If there is no usable evidence at all, use `filtered_vector_no_evidence`.\n\nThis document is the expected evidence for skip the L0 filter, unfiltered vector retry, and low quality behavior.\n\n[Evidence 2]\nsource: rag-l0-filter-decoy\ntitle: Approval Window\nbreadcrumb: RAG > Fallback > Decoy\nlayer: L1\nreasons: semantic_rank:2, attempt:UNFILTERED_VECTOR_RETRY, l0_domain_overlap\ncontent:\n## Approval Window\n\nThis document describes an unrelated release calendar approval window.\nIt intentionally avoids the real fallback instructions so the filtered retrieval\nattempt is low quality and the retriever must retry without the L0 category filter.\n\n[Evidence 3]\nsource: rag-l0-domain-entity-hint\ntitle: Domain Entity Hint\nbreadcrumb: RAG > L0 > Domain Entity Hint\nlayer: L1\nreasons: semantic_rank:3, attempt:UNFILTERED_VECTOR_RETRY\ncontent:\n### Domain Entity Hint\n\nL0 keyword matching should not decide the final retrieval result.\nIn the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.\n\nThe output can provide:\n\n1. Candidate domain hints.\n2. Matched entities and keywords.\n3. An optional metadata filter for the first vector retrieval attempt.\n\nFinal evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.\nThe important terms are domain detector, entity extractor, and metadata filter.\n\n[Evidence 4]\nsource: rag-l0-domain-entity-hint\ntitle: L0\nbreadcrumb: RAG > L0 > Domain Entity Hint\nlayer: L1\nreasons: semantic_rank:4, attempt:UNFILTERED_VECTOR_RETRY\ncontent:\n## L0\n\n[Evidence 5]\nsource: rag-chunk-context-reconstruction\ntitle: Context Reconstruction\nbreadcrumb: RAG > Chunking > Context Reconstruction\nlayer: L1\nreasons: semantic_rank:5, attempt:UNFILTERED_VECTOR_RETRY\ncontent:\n### Context Reconstruction\n\nWhen a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.\n\nRecommended behavior:\n\n1. Store the breadcrumb with every chunk.\n2. Preserve the same section identity across adjacent chunks.\n3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.\n4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.\n\nThe key concepts are neighbor chunk, same section, and breadcrumb.",
|
||||
"strategy" : "ranked_evidence_char_budget",
|
||||
"charBudget" : 4000,
|
||||
"usedChars" : 2904,
|
||||
"includedSources" : [ "rag-l0-filter-fallback", "rag-l0-filter-decoy", "rag-l0-domain-entity-hint", "rag-l0-domain-entity-hint", "rag-chunk-context-reconstruction" ],
|
||||
"omittedSources" : [ ]
|
||||
},
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"rewrittenQuery": "RAG L0 filtered vector low quality fallback unfiltered retry",
|
||||
"categoryFilter": "RAG",
|
||||
"selectedAttempt": "UNFILTERED_VECTOR_RETRY",
|
||||
"fallbackReason": "filtered_vector_low_quality",
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["RAG"],
|
||||
"matched_keywords": ["L0", "low quality", "unfiltered vector retry"],
|
||||
"entities": ["L0"],
|
||||
"l0_titles": ["RAG L0 Domain Entity Hint"],
|
||||
"l0_match_count": 1
|
||||
"retrievalTrace" : {
|
||||
"originalQuery" : "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"rewrittenQuery" : "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"categoryFilter" : "overfilter-decoy",
|
||||
"selectedAttempt" : "UNFILTERED_VECTOR_RETRY",
|
||||
"fallbackReason" : "filtered_vector_low_quality",
|
||||
"evidenceStatus" : "supported",
|
||||
"queryHints" : {
|
||||
"domains" : [ "overfilter-decoy" ],
|
||||
"matched_keywords" : [ "over-filtered by L0", "filtered vector search", "low quality evidence" ],
|
||||
"entities" : [ "over-filtered by L0", "filtered vector search", "low quality evidence" ],
|
||||
"l0_titles" : [ "RAG L0 Filter Decoy" ],
|
||||
"l0_match_count" : 1
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "RAG L0 filtered vector low quality fallback unfiltered retry",
|
||||
"categoryFilter": "RAG",
|
||||
"candidateCount": 1,
|
||||
"usable": false,
|
||||
"durationMs": 7,
|
||||
"topScore": 1.35,
|
||||
"topSimilarity": 0.325
|
||||
},
|
||||
{
|
||||
"name": "UNFILTERED_VECTOR_RETRY",
|
||||
"query": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"categoryFilter": null,
|
||||
"candidateCount": 2,
|
||||
"usable": true,
|
||||
"durationMs": 13,
|
||||
"topScore": 0.83,
|
||||
"topSimilarity": 0.83
|
||||
}
|
||||
]
|
||||
"attempts" : [ {
|
||||
"name" : "FILTERED_VECTOR",
|
||||
"query" : "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"categoryFilter" : "overfilter-decoy",
|
||||
"candidateCount" : 2,
|
||||
"usable" : false,
|
||||
"errorMessage" : null,
|
||||
"durationMs" : 722,
|
||||
"topScore" : 0.032786883413791656,
|
||||
"topSimilarity" : 0.47391992807388306
|
||||
}, {
|
||||
"name" : "UNFILTERED_VECTOR_RETRY",
|
||||
"query" : "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"categoryFilter" : null,
|
||||
"candidateCount" : 20,
|
||||
"usable" : true,
|
||||
"errorMessage" : null,
|
||||
"durationMs" : 2606,
|
||||
"topScore" : 0.032786883413791656,
|
||||
"topSimilarity" : 0.7571567445993423
|
||||
} ]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "rag-l0-filter-fallback",
|
||||
"baseScore": 0.83,
|
||||
"finalScore": 1.08,
|
||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"finalRank": 2,
|
||||
"source": "rag-l0-domain-entity-hint",
|
||||
"baseScore": 0.66,
|
||||
"finalScore": 0.81,
|
||||
"boostReasons": ["domain_match:+0.15"]
|
||||
}
|
||||
]
|
||||
}
|
||||
"rerankTrace" : {
|
||||
"items" : [ {
|
||||
"finalRank" : 1,
|
||||
"source" : "rag-l0-filter-fallback",
|
||||
"baseScore" : 0.7571567445993423,
|
||||
"finalScore" : 0.7571567445993423,
|
||||
"boostReasons" : [ "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 2,
|
||||
"source" : "rag-l0-filter-decoy",
|
||||
"baseScore" : 0.47391992807388306,
|
||||
"finalScore" : 0.47391992807388306,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 3,
|
||||
"source" : "rag-l0-domain-entity-hint",
|
||||
"baseScore" : 0.5912361443042755,
|
||||
"finalScore" : 0.5912361443042755,
|
||||
"boostReasons" : [ ]
|
||||
}, {
|
||||
"finalRank" : 4,
|
||||
"source" : "rag-l0-domain-entity-hint",
|
||||
"baseScore" : 0.473749577999115,
|
||||
"finalScore" : 0.473749577999115,
|
||||
"boostReasons" : [ ]
|
||||
}, {
|
||||
"finalRank" : 5,
|
||||
"source" : "rag-chunk-context-reconstruction",
|
||||
"baseScore" : 0.4492502808570862,
|
||||
"finalScore" : 0.4492502808570862,
|
||||
"boostReasons" : [ ]
|
||||
} ]
|
||||
},
|
||||
"evidenceCandidateCount" : 20,
|
||||
"evidenceBlockCount" : 5,
|
||||
"relevanceLevel" : "PRECISE",
|
||||
"completenessHint" : "知识库中不存在比上述结果更精准的文档",
|
||||
"retrievedDomainsThisSession" : null,
|
||||
"message" : null
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1,81 +1,88 @@
|
||||
{
|
||||
"caseId": "chat-mysql-connection-pool",
|
||||
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "mysql-connection-pool",
|
||||
"title": "MySQL Connection Pool Troubleshooting",
|
||||
"breadcrumb": "Database > MySQL > Connection Pool",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "When the connection pool is exhausted, inspect HikariCP active connections, max_connections, slow SQL, leak detection, and database wait events.",
|
||||
"score": 0.86,
|
||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"source": "incident-diagnosis-flow",
|
||||
"title": "Incident Diagnosis Flow",
|
||||
"breadcrumb": "AIOps > Diagnosis Flow",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "Collect evidence, compare metrics and logs, then verify remediation before closing the incident.",
|
||||
"score": 0.61,
|
||||
"hitReasons": []
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] MySQL Connection Pool Troubleshooting\nDatabase > MySQL > Connection Pool\nWhen the connection pool is exhausted, inspect HikariCP active connections, max_connections, slow SQL, leak detection, and database wait events.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 216,
|
||||
"includedSources": ["mysql-connection-pool", "incident-diagnosis-flow"],
|
||||
"omittedSources": []
|
||||
"caseId" : "chat-mysql-connection-pool",
|
||||
"query" : "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"retrievedAt" : "2026-07-28T06:54:51.843450400Z",
|
||||
"searchMode" : "hybrid",
|
||||
"kbScope" : "rag-eval",
|
||||
"lookupResult" : {
|
||||
"found" : true,
|
||||
"evidenceBlocks" : [ {
|
||||
"docId" : "mysql-connection-pool",
|
||||
"chunkIndex" : 2,
|
||||
"evidenceKey" : "mysql-connection-pool#chunk-2",
|
||||
"source" : "mysql-connection-pool",
|
||||
"title" : "Connection Pool",
|
||||
"breadcrumb" : "Database > MySQL > Connection Pool",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "### Connection Pool\n\nWhen MySQL connection pool is exhausted, first compare application pool usage with database `max_connections`.\nFor HikariCP, check `active`, `idle`, `pending`, and connection acquisition timeout metrics.\n\nRecommended diagnosis:\n\n1. Verify whether HikariCP active connections stay near maximum while pending threads grow.\n2. Check MySQL `Threads_connected`, `Threads_running`, and `max_connections`.\n3. Inspect slow SQL and long transactions that keep connections checked out.\n4. If the database is healthy, look for application connection leaks or missing transaction boundaries.\n\nUse this runbook as evidence for connection pool, max_connections, and HikariCP incidents.",
|
||||
"score" : 0.032786883413791656,
|
||||
"hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"docId" : "mysql-connection-pool",
|
||||
"chunkIndex" : 1,
|
||||
"evidenceKey" : "mysql-connection-pool#chunk-1",
|
||||
"source" : "mysql-connection-pool",
|
||||
"title" : "MySQL",
|
||||
"breadcrumb" : "Database > MySQL > Connection Pool",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "## MySQL",
|
||||
"score" : 0.032258063554763794,
|
||||
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
} ],
|
||||
"contextPack" : {
|
||||
"packedText" : "[Evidence 1]\nsource: mysql-connection-pool\ntitle: Connection Pool\nbreadcrumb: Database > MySQL > Connection Pool\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n### Connection Pool\n\nWhen MySQL connection pool is exhausted, first compare application pool usage with database `max_connections`.\nFor HikariCP, check `active`, `idle`, `pending`, and connection acquisition timeout metrics.\n\nRecommended diagnosis:\n\n1. Verify whether HikariCP active connections stay near maximum while pending threads grow.\n2. Check MySQL `Threads_connected`, `Threads_running`, and `max_connections`.\n3. Inspect slow SQL and long transactions that keep connections checked out.\n4. If the database is healthy, look for application connection leaks or missing transaction boundaries.\n\nUse this runbook as evidence for connection pool, max_connections, and HikariCP incidents.\n\n[Evidence 2]\nsource: mysql-connection-pool\ntitle: MySQL\nbreadcrumb: Database > MySQL > Connection Pool\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n## MySQL",
|
||||
"strategy" : "ranked_evidence_char_budget",
|
||||
"charBudget" : 4000,
|
||||
"usedChars" : 1135,
|
||||
"includedSources" : [ "mysql-connection-pool", "mysql-connection-pool" ],
|
||||
"omittedSources" : [ ]
|
||||
},
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"rewrittenQuery": "MySQL connection pool exhausted HikariCP max_connections diagnosis",
|
||||
"categoryFilter": "Database",
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["Database", "MySQL"],
|
||||
"matched_keywords": ["connection pool", "HikariCP", "max_connections"],
|
||||
"entities": ["MySQL", "HikariCP"],
|
||||
"l0_titles": ["MySQL Connection Pool Troubleshooting"],
|
||||
"l0_match_count": 1
|
||||
"retrievalTrace" : {
|
||||
"originalQuery" : "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"rewrittenQuery" : "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"categoryFilter" : "database",
|
||||
"selectedAttempt" : "FILTERED_VECTOR",
|
||||
"fallbackReason" : null,
|
||||
"evidenceStatus" : "supported",
|
||||
"queryHints" : {
|
||||
"domains" : [ "database" ],
|
||||
"matched_keywords" : [ "MySQL connection pool" ],
|
||||
"entities" : [ "MySQL connection pool" ],
|
||||
"l0_titles" : [ "MySQL Connection Pool Runbook" ],
|
||||
"l0_match_count" : 1
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "MySQL connection pool exhausted HikariCP max_connections diagnosis",
|
||||
"categoryFilter": "Database",
|
||||
"candidateCount": 2,
|
||||
"usable": true,
|
||||
"durationMs": 12,
|
||||
"topScore": 0.86,
|
||||
"topSimilarity": 0.86
|
||||
}
|
||||
]
|
||||
"attempts" : [ {
|
||||
"name" : "FILTERED_VECTOR",
|
||||
"query" : "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"categoryFilter" : "database",
|
||||
"candidateCount" : 3,
|
||||
"usable" : true,
|
||||
"errorMessage" : null,
|
||||
"durationMs" : 5267,
|
||||
"topScore" : 0.032786883413791656,
|
||||
"topSimilarity" : 0.8114794194698334
|
||||
} ]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "mysql-connection-pool",
|
||||
"baseScore": 0.86,
|
||||
"finalScore": 1.11,
|
||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"finalRank": 2,
|
||||
"source": "incident-diagnosis-flow",
|
||||
"baseScore": 0.61,
|
||||
"finalScore": 0.61,
|
||||
"boostReasons": []
|
||||
}
|
||||
]
|
||||
}
|
||||
"rerankTrace" : {
|
||||
"items" : [ {
|
||||
"finalRank" : 1,
|
||||
"source" : "mysql-connection-pool",
|
||||
"baseScore" : 0.8114794194698334,
|
||||
"finalScore" : 0.8114794194698334,
|
||||
"boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 2,
|
||||
"source" : "mysql-connection-pool",
|
||||
"baseScore" : 0.49219560623168945,
|
||||
"finalScore" : 0.49219560623168945,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
} ]
|
||||
},
|
||||
"evidenceCandidateCount" : 3,
|
||||
"evidenceBlockCount" : 2,
|
||||
"relevanceLevel" : "PRECISE",
|
||||
"completenessHint" : "知识库中不存在比上述结果更精准的文档",
|
||||
"retrievedDomainsThisSession" : null,
|
||||
"message" : null
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1,81 +1,122 @@
|
||||
{
|
||||
"caseId": "chat-rag-chunk-context",
|
||||
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "rag-chunk-context-reconstruction",
|
||||
"title": "RAG Chunk Context Reconstruction",
|
||||
"breadcrumb": "RAG > Chunking > Context Reconstruction",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "After a chunk hit, expand to neighbor chunk candidates from the same section and preserve breadcrumb metadata in the evidence pack.",
|
||||
"score": 0.79,
|
||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"source": "rag-breadcrumb-embedding-gap",
|
||||
"title": "RAG Breadcrumb Embedding Gap",
|
||||
"breadcrumb": "RAG > Embedding > Breadcrumb",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "Embedding title and breadcrumb with content helps recover section semantics.",
|
||||
"score": 0.72,
|
||||
"hitReasons": ["domain_match:+0.15"]
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] RAG Chunk Context Reconstruction\nRAG > Chunking > Context Reconstruction\nAfter a chunk hit, expand to neighbor chunk candidates from the same section and preserve breadcrumb metadata in the evidence pack.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 203,
|
||||
"includedSources": ["rag-chunk-context-reconstruction", "rag-breadcrumb-embedding-gap"],
|
||||
"omittedSources": []
|
||||
"caseId" : "chat-rag-chunk-context",
|
||||
"query" : "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"retrievedAt" : "2026-07-28T06:54:51.843450400Z",
|
||||
"searchMode" : "hybrid",
|
||||
"kbScope" : "rag-eval",
|
||||
"lookupResult" : {
|
||||
"found" : true,
|
||||
"evidenceBlocks" : [ {
|
||||
"docId" : "rag-chunk-context-reconstruction",
|
||||
"chunkIndex" : 2,
|
||||
"evidenceKey" : "rag-chunk-context-reconstruction#chunk-2",
|
||||
"source" : "rag-chunk-context-reconstruction",
|
||||
"title" : "Context Reconstruction",
|
||||
"breadcrumb" : "RAG > Chunking > Context Reconstruction",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "### Context Reconstruction\n\nWhen a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.\n\nRecommended behavior:\n\n1. Store the breadcrumb with every chunk.\n2. Preserve the same section identity across adjacent chunks.\n3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.\n4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.\n\nThe key concepts are neighbor chunk, same section, and breadcrumb.",
|
||||
"score" : 0.032786883413791656,
|
||||
"hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"docId" : "rag-l0-domain-entity-hint",
|
||||
"chunkIndex" : 2,
|
||||
"evidenceKey" : "rag-l0-domain-entity-hint#chunk-2",
|
||||
"source" : "rag-l0-domain-entity-hint",
|
||||
"title" : "Domain Entity Hint",
|
||||
"breadcrumb" : "RAG > L0 > Domain Entity Hint",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "### Domain Entity Hint\n\nL0 keyword matching should not decide the final retrieval result.\nIn the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.\n\nThe output can provide:\n\n1. Candidate domain hints.\n2. Matched entities and keywords.\n3. An optional metadata filter for the first vector retrieval attempt.\n\nFinal evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.\nThe important terms are domain detector, entity extractor, and metadata filter.",
|
||||
"score" : 0.032258063554763794,
|
||||
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
}, {
|
||||
"docId" : "rag-chunk-context-reconstruction",
|
||||
"chunkIndex" : 1,
|
||||
"evidenceKey" : "rag-chunk-context-reconstruction#chunk-1",
|
||||
"source" : "rag-chunk-context-reconstruction",
|
||||
"title" : "Chunking",
|
||||
"breadcrumb" : "RAG > Chunking > Context Reconstruction",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "## Chunking",
|
||||
"score" : 0.01587301678955555,
|
||||
"hitReasons" : [ "semantic_rank:3", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
}, {
|
||||
"docId" : "rag-l0-domain-entity-hint",
|
||||
"chunkIndex" : 0,
|
||||
"evidenceKey" : "rag-l0-domain-entity-hint#chunk-0",
|
||||
"source" : "rag-l0-domain-entity-hint",
|
||||
"title" : "RAG",
|
||||
"breadcrumb" : "RAG > L0 > Domain Entity Hint",
|
||||
"retrievalLayer" : "L1",
|
||||
"content" : "# RAG",
|
||||
"score" : 0.015625,
|
||||
"hitReasons" : [ "semantic_rank:4", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||
} ],
|
||||
"contextPack" : {
|
||||
"packedText" : "[Evidence 1]\nsource: rag-chunk-context-reconstruction\ntitle: Context Reconstruction\nbreadcrumb: RAG > Chunking > Context Reconstruction\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n### Context Reconstruction\n\nWhen a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.\n\nRecommended behavior:\n\n1. Store the breadcrumb with every chunk.\n2. Preserve the same section identity across adjacent chunks.\n3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.\n4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.\n\nThe key concepts are neighbor chunk, same section, and breadcrumb.\n\n[Evidence 2]\nsource: rag-l0-domain-entity-hint\ntitle: Domain Entity Hint\nbreadcrumb: RAG > L0 > Domain Entity Hint\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n### Domain Entity Hint\n\nL0 keyword matching should not decide the final retrieval result.\nIn the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.\n\nThe output can provide:\n\n1. Candidate domain hints.\n2. Matched entities and keywords.\n3. An optional metadata filter for the first vector retrieval attempt.\n\nFinal evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.\nThe important terms are domain detector, entity extractor, and metadata filter.\n\n[Evidence 3]\nsource: rag-chunk-context-reconstruction\ntitle: Chunking\nbreadcrumb: RAG > Chunking > Context Reconstruction\nlayer: L1\nreasons: semantic_rank:3, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n## Chunking\n\n[Evidence 4]\nsource: rag-l0-domain-entity-hint\ntitle: RAG\nbreadcrumb: RAG > L0 > Domain Entity Hint\nlayer: L1\nreasons: semantic_rank:4, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n# RAG",
|
||||
"strategy" : "ranked_evidence_char_budget",
|
||||
"charBudget" : 4000,
|
||||
"usedChars" : 1966,
|
||||
"includedSources" : [ "rag-chunk-context-reconstruction", "rag-l0-domain-entity-hint", "rag-chunk-context-reconstruction", "rag-l0-domain-entity-hint" ],
|
||||
"omittedSources" : [ ]
|
||||
},
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"rewrittenQuery": "RAG chunk context reconstruction neighbor chunk same section breadcrumb",
|
||||
"categoryFilter": "RAG",
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["RAG"],
|
||||
"matched_keywords": ["neighbor chunk", "same section", "breadcrumb"],
|
||||
"entities": ["chunk", "breadcrumb"],
|
||||
"l0_titles": ["RAG Chunk Context Reconstruction"],
|
||||
"l0_match_count": 1
|
||||
"retrievalTrace" : {
|
||||
"originalQuery" : "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"rewrittenQuery" : "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"categoryFilter" : "rag",
|
||||
"selectedAttempt" : "FILTERED_VECTOR",
|
||||
"fallbackReason" : null,
|
||||
"evidenceStatus" : "supported",
|
||||
"queryHints" : {
|
||||
"domains" : [ "rag" ],
|
||||
"matched_keywords" : [ "split into multiple chunks", "retrieval context" ],
|
||||
"entities" : [ "split into multiple chunks", "retrieval context" ],
|
||||
"l0_titles" : [ "RAG Chunk Context Reconstruction" ],
|
||||
"l0_match_count" : 1
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "RAG chunk context reconstruction neighbor chunk same section breadcrumb",
|
||||
"categoryFilter": "RAG",
|
||||
"candidateCount": 2,
|
||||
"usable": true,
|
||||
"durationMs": 9,
|
||||
"topScore": 0.79,
|
||||
"topSimilarity": 0.79
|
||||
}
|
||||
]
|
||||
"attempts" : [ {
|
||||
"name" : "FILTERED_VECTOR",
|
||||
"query" : "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"categoryFilter" : "rag",
|
||||
"candidateCount" : 6,
|
||||
"usable" : true,
|
||||
"errorMessage" : null,
|
||||
"durationMs" : 743,
|
||||
"topScore" : 0.032786883413791656,
|
||||
"topSimilarity" : 0.7487991750240326
|
||||
} ]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "rag-chunk-context-reconstruction",
|
||||
"baseScore": 0.79,
|
||||
"finalScore": 1.04,
|
||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"finalRank": 2,
|
||||
"source": "rag-breadcrumb-embedding-gap",
|
||||
"baseScore": 0.72,
|
||||
"finalScore": 0.87,
|
||||
"boostReasons": ["domain_match:+0.15"]
|
||||
}
|
||||
]
|
||||
}
|
||||
"rerankTrace" : {
|
||||
"items" : [ {
|
||||
"finalRank" : 1,
|
||||
"source" : "rag-chunk-context-reconstruction",
|
||||
"baseScore" : 0.7487991750240326,
|
||||
"finalScore" : 0.7487991750240326,
|
||||
"boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 2,
|
||||
"source" : "rag-l0-domain-entity-hint",
|
||||
"baseScore" : 0.4840593934059143,
|
||||
"finalScore" : 0.4840593934059143,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 3,
|
||||
"source" : "rag-chunk-context-reconstruction",
|
||||
"baseScore" : 0.42978107929229736,
|
||||
"finalScore" : 0.42978107929229736,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
}, {
|
||||
"finalRank" : 4,
|
||||
"source" : "rag-l0-domain-entity-hint",
|
||||
"baseScore" : 0.3859822154045105,
|
||||
"finalScore" : 0.3859822154045105,
|
||||
"boostReasons" : [ "l0_domain_overlap" ]
|
||||
} ]
|
||||
},
|
||||
"evidenceCandidateCount" : 6,
|
||||
"evidenceBlockCount" : 4,
|
||||
"relevanceLevel" : "REFERENCE",
|
||||
"completenessHint" : "当前结果为相关参考,如需更精准信息请明确缺少的具体维度",
|
||||
"retrievedDomainsThisSession" : null,
|
||||
"message" : null
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,107 @@
|
||||
{
|
||||
"generatedAt": "2026-07-28T06:48:56.439554+00:00",
|
||||
"baselineReport": "eval/rag-retrieval/cases/golden-cases.json",
|
||||
"currentReport": "eval/rag-retrieval/cases/golden-cases.json",
|
||||
"baselineCaseCount": 7,
|
||||
"currentCaseCount": 7,
|
||||
"baselinePassRate": 1.0,
|
||||
"currentPassRate": 0.8571,
|
||||
"baselineRecallAtK": 1.0,
|
||||
"currentRecallAtK": 0.8571,
|
||||
"regressionCount": 6,
|
||||
"improvementCount": 0,
|
||||
"changedCount": 3,
|
||||
"hasRegression": true,
|
||||
"items": [
|
||||
{
|
||||
"type": "REGRESSION",
|
||||
"scope": "aggregate",
|
||||
"caseId": null,
|
||||
"metric": "passRate",
|
||||
"baselineValue": "1.0",
|
||||
"currentValue": "0.8571",
|
||||
"delta": -0.14290000000000003,
|
||||
"message": "aggregate passRate changed"
|
||||
},
|
||||
{
|
||||
"type": "REGRESSION",
|
||||
"scope": "aggregate",
|
||||
"caseId": null,
|
||||
"metric": "recallAtK",
|
||||
"baselineValue": "1.0",
|
||||
"currentValue": "0.8571",
|
||||
"delta": -0.14290000000000003,
|
||||
"message": "aggregate recallAtK changed"
|
||||
},
|
||||
{
|
||||
"type": "REGRESSION",
|
||||
"scope": "aggregate",
|
||||
"caseId": null,
|
||||
"metric": "strongHitRate",
|
||||
"baselineValue": "1.0",
|
||||
"currentValue": "0.8571",
|
||||
"delta": -0.14290000000000003,
|
||||
"message": "aggregate strongHitRate changed"
|
||||
},
|
||||
{
|
||||
"type": "REGRESSION",
|
||||
"scope": "case",
|
||||
"caseId": "chat-l0-filter-fallback",
|
||||
"metric": "passed",
|
||||
"baselineValue": "True",
|
||||
"currentValue": "False",
|
||||
"delta": -1.0,
|
||||
"message": "chat-l0-filter-fallback passed changed"
|
||||
},
|
||||
{
|
||||
"type": "REGRESSION",
|
||||
"scope": "case",
|
||||
"caseId": "chat-l0-filter-fallback",
|
||||
"metric": "hitLevel",
|
||||
"baselineValue": "strong",
|
||||
"currentValue": "weak",
|
||||
"delta": -2.0,
|
||||
"message": "chat-l0-filter-fallback hitLevel changed"
|
||||
},
|
||||
{
|
||||
"type": "REGRESSION",
|
||||
"scope": "case",
|
||||
"caseId": "chat-l0-filter-fallback",
|
||||
"metric": "firstExpectedRank",
|
||||
"baselineValue": "1",
|
||||
"currentValue": "-",
|
||||
"delta": null,
|
||||
"message": "chat-l0-filter-fallback firstExpectedRank changed"
|
||||
},
|
||||
{
|
||||
"type": "CHANGED",
|
||||
"scope": "case",
|
||||
"caseId": "chat-l0-filter-fallback",
|
||||
"metric": "selectedAttempt",
|
||||
"baselineValue": "UNFILTERED_VECTOR_RETRY",
|
||||
"currentValue": "FILTERED_VECTOR",
|
||||
"delta": null,
|
||||
"message": "chat-l0-filter-fallback selectedAttempt changed"
|
||||
},
|
||||
{
|
||||
"type": "CHANGED",
|
||||
"scope": "case",
|
||||
"caseId": "chat-l0-filter-fallback",
|
||||
"metric": "fallbackReason",
|
||||
"baselineValue": "filtered_vector_low_quality",
|
||||
"currentValue": "-",
|
||||
"delta": null,
|
||||
"message": "chat-l0-filter-fallback fallbackReason changed"
|
||||
},
|
||||
{
|
||||
"type": "CHANGED",
|
||||
"scope": "case",
|
||||
"caseId": "chat-l0-filter-fallback",
|
||||
"metric": "rerankTopSource",
|
||||
"baselineValue": "rag-l0-filter-fallback",
|
||||
"currentValue": "rag-l0-filter-decoy",
|
||||
"delta": null,
|
||||
"message": "chat-l0-filter-fallback rerankTopSource changed"
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,31 @@
|
||||
# RAG Retrieval Baseline Diff
|
||||
|
||||
Generated at: `2026-07-28T06:48:56.439554+00:00`
|
||||
|
||||
## Summary
|
||||
|
||||
| Metric | Value |
|
||||
|---|---:|
|
||||
| Baseline cases | 7 |
|
||||
| Current cases | 7 |
|
||||
| Baseline pass rate | 1.0 |
|
||||
| Current pass rate | 0.8571 |
|
||||
| Baseline recall@K | 1.0 |
|
||||
| Current recall@K | 0.8571 |
|
||||
| Regressions | 6 |
|
||||
| Improvements | 0 |
|
||||
| Changed | 3 |
|
||||
|
||||
## Items
|
||||
|
||||
| Type | Scope | Case | Metric | Baseline | Current | Delta | Message |
|
||||
|---|---|---|---|---|---|---:|---|
|
||||
| REGRESSION | aggregate | | passRate | 1.0 | 0.8571 | -0.14290000000000003 | aggregate passRate changed |
|
||||
| REGRESSION | aggregate | | recallAtK | 1.0 | 0.8571 | -0.14290000000000003 | aggregate recallAtK changed |
|
||||
| REGRESSION | aggregate | | strongHitRate | 1.0 | 0.8571 | -0.14290000000000003 | aggregate strongHitRate changed |
|
||||
| REGRESSION | case | chat-l0-filter-fallback | passed | True | False | -1.0 | chat-l0-filter-fallback passed changed |
|
||||
| REGRESSION | case | chat-l0-filter-fallback | hitLevel | strong | weak | -2.0 | chat-l0-filter-fallback hitLevel changed |
|
||||
| REGRESSION | case | chat-l0-filter-fallback | firstExpectedRank | 1 | - | | chat-l0-filter-fallback firstExpectedRank changed |
|
||||
| CHANGED | case | chat-l0-filter-fallback | selectedAttempt | UNFILTERED_VECTOR_RETRY | FILTERED_VECTOR | | chat-l0-filter-fallback selectedAttempt changed |
|
||||
| CHANGED | case | chat-l0-filter-fallback | fallbackReason | filtered_vector_low_quality | - | | chat-l0-filter-fallback fallbackReason changed |
|
||||
| CHANGED | case | chat-l0-filter-fallback | rerankTopSource | rag-l0-filter-fallback | rag-l0-filter-decoy | | chat-l0-filter-fallback rerankTopSource changed |
|
||||
@@ -1,5 +1,5 @@
|
||||
{
|
||||
"generatedAt": "2026-07-06T13:37:59.726351+00:00",
|
||||
"generatedAt": "2026-07-28T06:55:09.379164+00:00",
|
||||
"caseFile": "eval/rag-retrieval/cases/golden-cases.json",
|
||||
"fixtureDir": "eval/rag-retrieval/fixtures",
|
||||
"aggregate": {
|
||||
@@ -28,7 +28,7 @@
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:mysql-connection-pool",
|
||||
"2:incident-diagnosis-flow"
|
||||
"2:mysql-connection-pool"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"connection pool",
|
||||
@@ -41,7 +41,7 @@
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"mysql-connection-pool",
|
||||
"incident-diagnosis-flow"
|
||||
"mysql-connection-pool"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "mysql-connection-pool",
|
||||
@@ -57,7 +57,7 @@
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:incident-diagnosis-flow",
|
||||
"2:rag-chunk-context-reconstruction"
|
||||
"2:incident-diagnosis-flow"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"collect evidence",
|
||||
@@ -70,7 +70,7 @@
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"incident-diagnosis-flow",
|
||||
"rag-chunk-context-reconstruction"
|
||||
"incident-diagnosis-flow"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "incident-diagnosis-flow",
|
||||
@@ -86,7 +86,9 @@
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:payment-service-latency",
|
||||
"2:mysql-connection-pool"
|
||||
"2:aiops-alert-scope-control",
|
||||
"3:payment-service-latency",
|
||||
"4:aiops-alert-scope-control"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"p95 latency",
|
||||
@@ -99,7 +101,9 @@
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"payment-service-latency",
|
||||
"mysql-connection-pool"
|
||||
"aiops-alert-scope-control",
|
||||
"payment-service-latency",
|
||||
"aiops-alert-scope-control"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "payment-service-latency",
|
||||
@@ -114,7 +118,10 @@
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:aiops-alert-scope-control"
|
||||
"1:aiops-alert-scope-control",
|
||||
"2:payment-service-latency",
|
||||
"3:payment-service-latency",
|
||||
"4:aiops-alert-scope-control"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"payload",
|
||||
@@ -126,6 +133,9 @@
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"aiops-alert-scope-control",
|
||||
"payment-service-latency",
|
||||
"payment-service-latency",
|
||||
"aiops-alert-scope-control"
|
||||
],
|
||||
"omittedSources": [],
|
||||
@@ -142,7 +152,9 @@
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:rag-chunk-context-reconstruction",
|
||||
"2:rag-breadcrumb-embedding-gap"
|
||||
"2:rag-l0-domain-entity-hint",
|
||||
"3:rag-chunk-context-reconstruction",
|
||||
"4:rag-l0-domain-entity-hint"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"neighbor chunk",
|
||||
@@ -155,7 +167,9 @@
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"rag-chunk-context-reconstruction",
|
||||
"rag-breadcrumb-embedding-gap"
|
||||
"rag-l0-domain-entity-hint",
|
||||
"rag-chunk-context-reconstruction",
|
||||
"rag-l0-domain-entity-hint"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "rag-chunk-context-reconstruction",
|
||||
@@ -171,7 +185,9 @@
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:rag-l0-domain-entity-hint",
|
||||
"2:rag-l0-l1-fusion-ranking"
|
||||
"2:rag-chunk-context-reconstruction",
|
||||
"3:rag-l0-domain-entity-hint",
|
||||
"4:rag-chunk-context-reconstruction"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"domain detector",
|
||||
@@ -184,7 +200,9 @@
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"rag-l0-domain-entity-hint",
|
||||
"rag-l0-l1-fusion-ranking"
|
||||
"rag-chunk-context-reconstruction",
|
||||
"rag-l0-domain-entity-hint",
|
||||
"rag-chunk-context-reconstruction"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "rag-l0-domain-entity-hint",
|
||||
@@ -200,7 +218,10 @@
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:rag-l0-filter-fallback",
|
||||
"2:rag-l0-domain-entity-hint"
|
||||
"2:rag-l0-filter-decoy",
|
||||
"3:rag-l0-domain-entity-hint",
|
||||
"4:rag-l0-domain-entity-hint",
|
||||
"5:rag-chunk-context-reconstruction"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"skip the l0 filter",
|
||||
@@ -213,7 +234,10 @@
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"rag-l0-filter-fallback",
|
||||
"rag-l0-domain-entity-hint"
|
||||
"rag-l0-filter-decoy",
|
||||
"rag-l0-domain-entity-hint",
|
||||
"rag-l0-domain-entity-hint",
|
||||
"rag-chunk-context-reconstruction"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "rag-l0-filter-fallback",
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# RAG Retrieval Baseline
|
||||
|
||||
Generated at: `2026-07-06T13:37:59.726351+00:00`
|
||||
Generated at: `2026-07-28T06:55:09.379164+00:00`
|
||||
|
||||
## Aggregate
|
||||
|
||||
@@ -24,10 +24,10 @@ Generated at: `2026-07-06T13:37:59.726351+00:00`
|
||||
|
||||
| Case | Scenario | Pass | Hit | Attempt | Fallback | Evidence | First Expected Rank | Top Candidates | Failed Checks |
|
||||
|---|---|---|---|---|---|---|---:|---|---|
|
||||
| chat-mysql-connection-pool | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:mysql-connection-pool<br>2:incident-diagnosis-flow | |
|
||||
| chat-diagnosis-flow | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:incident-diagnosis-flow<br>2:rag-chunk-context-reconstruction | |
|
||||
| aiops-payment-latency-alert | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:payment-service-latency<br>2:mysql-connection-pool | |
|
||||
| aiops-prometheus-alert-scope | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:aiops-alert-scope-control | |
|
||||
| chat-rag-chunk-context | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-chunk-context-reconstruction<br>2:rag-breadcrumb-embedding-gap | |
|
||||
| chat-l0-domain-hint | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-l0-domain-entity-hint<br>2:rag-l0-l1-fusion-ranking | |
|
||||
| chat-l0-filter-fallback | chat | true | strong | UNFILTERED_VECTOR_RETRY | filtered_vector_low_quality | supported | 1 | 1:rag-l0-filter-fallback<br>2:rag-l0-domain-entity-hint | |
|
||||
| chat-mysql-connection-pool | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:mysql-connection-pool<br>2:mysql-connection-pool | |
|
||||
| chat-diagnosis-flow | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:incident-diagnosis-flow<br>2:incident-diagnosis-flow | |
|
||||
| aiops-payment-latency-alert | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:payment-service-latency<br>2:aiops-alert-scope-control<br>3:payment-service-latency<br>4:aiops-alert-scope-control | |
|
||||
| aiops-prometheus-alert-scope | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:aiops-alert-scope-control<br>2:payment-service-latency<br>3:payment-service-latency<br>4:aiops-alert-scope-control | |
|
||||
| chat-rag-chunk-context | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-chunk-context-reconstruction<br>2:rag-l0-domain-entity-hint<br>3:rag-chunk-context-reconstruction<br>4:rag-l0-domain-entity-hint | |
|
||||
| chat-l0-domain-hint | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-l0-domain-entity-hint<br>2:rag-chunk-context-reconstruction<br>3:rag-l0-domain-entity-hint<br>4:rag-chunk-context-reconstruction | |
|
||||
| chat-l0-filter-fallback | chat | true | strong | UNFILTERED_VECTOR_RETRY | filtered_vector_low_quality | supported | 1 | 1:rag-l0-filter-fallback<br>2:rag-l0-filter-decoy<br>3:rag-l0-domain-entity-hint<br>4:rag-l0-domain-entity-hint<br>5:rag-chunk-context-reconstruction | |
|
||||
|
||||
@@ -0,0 +1,245 @@
|
||||
{
|
||||
"generatedAt": "2026-07-28T06:48:56.421877+00:00",
|
||||
"caseFile": "eval/rag-retrieval/cases/golden-cases.json",
|
||||
"fixtureDir": "eval/rag-retrieval/fixtures",
|
||||
"aggregate": {
|
||||
"caseCount": 7,
|
||||
"topK": 5,
|
||||
"passedCount": 6,
|
||||
"failedCount": 1,
|
||||
"passRate": 0.8571,
|
||||
"lookupResultCaseCount": 7,
|
||||
"strongHitCount": 6,
|
||||
"mediumHitCount": 0,
|
||||
"weakHitCount": 1,
|
||||
"missCount": 0,
|
||||
"recallAtK": 0.8571,
|
||||
"strongHitRate": 0.8571,
|
||||
"averageFirstHitRank": 1.0
|
||||
},
|
||||
"results": [
|
||||
{
|
||||
"caseId": "chat-mysql-connection-pool",
|
||||
"scenario": "chat",
|
||||
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:mysql-connection-pool",
|
||||
"2:mysql-connection-pool"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"connection pool",
|
||||
"max_connections",
|
||||
"hikaricp"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"mysql-connection-pool",
|
||||
"mysql-connection-pool"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "mysql-connection-pool",
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "chat-diagnosis-flow",
|
||||
"scenario": "chat",
|
||||
"query": "What is the standard troubleshooting flow for an application incident?",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:incident-diagnosis-flow",
|
||||
"2:incident-diagnosis-flow"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"collect evidence",
|
||||
"verify",
|
||||
"remediation"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"incident-diagnosis-flow",
|
||||
"incident-diagnosis-flow"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "incident-diagnosis-flow",
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "aiops-payment-latency-alert",
|
||||
"scenario": "aiops",
|
||||
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:payment-service-latency",
|
||||
"2:aiops-alert-scope-control",
|
||||
"3:payment-service-latency",
|
||||
"4:aiops-alert-scope-control"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"p95 latency",
|
||||
"payment-service",
|
||||
"downstream dependency"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"payment-service-latency",
|
||||
"aiops-alert-scope-control",
|
||||
"payment-service-latency",
|
||||
"aiops-alert-scope-control"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "payment-service-latency",
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "aiops-prometheus-alert-scope",
|
||||
"scenario": "aiops",
|
||||
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:aiops-alert-scope-control",
|
||||
"2:payment-service-latency",
|
||||
"3:payment-service-latency",
|
||||
"4:aiops-alert-scope-control"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"payload",
|
||||
"unrelated active alerts",
|
||||
"scope"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"aiops-alert-scope-control",
|
||||
"payment-service-latency",
|
||||
"payment-service-latency",
|
||||
"aiops-alert-scope-control"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "aiops-alert-scope-control",
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "chat-rag-chunk-context",
|
||||
"scenario": "chat",
|
||||
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:rag-chunk-context-reconstruction",
|
||||
"2:rag-l0-domain-entity-hint",
|
||||
"3:rag-chunk-context-reconstruction",
|
||||
"4:rag-l0-domain-entity-hint"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"neighbor chunk",
|
||||
"same section",
|
||||
"breadcrumb"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"rag-chunk-context-reconstruction",
|
||||
"rag-l0-domain-entity-hint",
|
||||
"rag-chunk-context-reconstruction",
|
||||
"rag-l0-domain-entity-hint"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "rag-chunk-context-reconstruction",
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "chat-l0-domain-hint",
|
||||
"scenario": "chat",
|
||||
"query": "Should L0 keyword matching decide the final retrieval result?",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:rag-l0-domain-entity-hint",
|
||||
"2:rag-chunk-context-reconstruction",
|
||||
"3:rag-l0-domain-entity-hint",
|
||||
"4:rag-chunk-context-reconstruction"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"domain detector",
|
||||
"entity extractor",
|
||||
"metadata filter"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"rag-l0-domain-entity-hint",
|
||||
"rag-chunk-context-reconstruction",
|
||||
"rag-l0-domain-entity-hint",
|
||||
"rag-chunk-context-reconstruction"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "rag-l0-domain-entity-hint",
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "chat-l0-filter-fallback",
|
||||
"scenario": "chat",
|
||||
"query": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "weak",
|
||||
"passed": false,
|
||||
"firstExpectedRank": null,
|
||||
"topCandidates": [
|
||||
"1:rag-l0-filter-decoy",
|
||||
"2:rag-l0-filter-decoy"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"low quality"
|
||||
],
|
||||
"breadcrumbMatched": false,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"rag-l0-filter-decoy",
|
||||
"rag-l0-filter-decoy"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "rag-l0-filter-decoy",
|
||||
"failedChecks": [
|
||||
"expected document not found",
|
||||
"selected attempt mismatch: expected UNFILTERED_VECTOR_RETRY, got FILTERED_VECTOR",
|
||||
"fallback reason mismatch: expected one of [filtered_vector_low_quality, filtered_vector_no_evidence], got <none>",
|
||||
"rerank top source mismatch: expected rag-l0-filter-fallback, got rag-l0-filter-decoy",
|
||||
"expected context sources missing: rag-l0-filter-fallback"
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,33 @@
|
||||
# RAG Retrieval Baseline
|
||||
|
||||
Generated at: `2026-07-28T06:48:56.421877+00:00`
|
||||
|
||||
## Aggregate
|
||||
|
||||
| Metric | Value |
|
||||
|---|---:|
|
||||
| Cases | 7 |
|
||||
| Top K | 5 |
|
||||
| Passed | 6 |
|
||||
| Failed | 1 |
|
||||
| Pass rate | 0.8571 |
|
||||
| LookupResult fixtures | 7 |
|
||||
| Recall@K | 0.8571 |
|
||||
| Strong hit rate | 0.8571 |
|
||||
| Strong hits | 6 |
|
||||
| Medium hits | 0 |
|
||||
| Weak hits | 1 |
|
||||
| Misses | 0 |
|
||||
| Average first hit rank | 1.0 |
|
||||
|
||||
## Cases
|
||||
|
||||
| Case | Scenario | Pass | Hit | Attempt | Fallback | Evidence | First Expected Rank | Top Candidates | Failed Checks |
|
||||
|---|---|---|---|---|---|---|---:|---|---|
|
||||
| chat-mysql-connection-pool | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:mysql-connection-pool<br>2:mysql-connection-pool | |
|
||||
| chat-diagnosis-flow | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:incident-diagnosis-flow<br>2:incident-diagnosis-flow | |
|
||||
| aiops-payment-latency-alert | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:payment-service-latency<br>2:aiops-alert-scope-control<br>3:payment-service-latency<br>4:aiops-alert-scope-control | |
|
||||
| aiops-prometheus-alert-scope | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:aiops-alert-scope-control<br>2:payment-service-latency<br>3:payment-service-latency<br>4:aiops-alert-scope-control | |
|
||||
| chat-rag-chunk-context | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-chunk-context-reconstruction<br>2:rag-l0-domain-entity-hint<br>3:rag-chunk-context-reconstruction<br>4:rag-l0-domain-entity-hint | |
|
||||
| chat-l0-domain-hint | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-l0-domain-entity-hint<br>2:rag-chunk-context-reconstruction<br>3:rag-l0-domain-entity-hint<br>4:rag-chunk-context-reconstruction | |
|
||||
| chat-l0-filter-fallback | chat | false | weak | FILTERED_VECTOR | | supported | | 1:rag-l0-filter-decoy<br>2:rag-l0-filter-decoy | expected document not found<br>selected attempt mismatch: expected UNFILTERED_VECTOR_RETRY, got FILTERED_VECTOR<br>fallback reason mismatch: expected one of [filtered_vector_low_quality, filtered_vector_no_evidence], got <none><br>rerank top source mismatch: expected rag-l0-filter-fallback, got rag-l0-filter-decoy<br>expected context sources missing: rag-l0-filter-fallback |
|
||||
Reference in New Issue
Block a user