feat(rag): close eval pipeline with live snapshots
This commit is contained in:
@@ -12,12 +12,59 @@ post-processing, or Spring AI VectorStore integration.
|
||||
```text
|
||||
eval/rag-retrieval/
|
||||
cases/golden-cases.json Fixed retrieval golden cases
|
||||
fixtures/*.json Saved retrieval candidates for each case
|
||||
seed-docs/*.md Canonical docs imported into the live KB for real-tool eval
|
||||
fixtures/*.json Saved retrieval fixtures for each case
|
||||
reports/baseline.json Machine-readable baseline report
|
||||
reports/baseline.md Human-readable baseline report
|
||||
reports/baseline-diff.* Optional diff reports
|
||||
reports/live-post-reindex.* Optional live acceptance reports
|
||||
```
|
||||
|
||||
## Seed Docs + Import/Reindex
|
||||
|
||||
The live-tool eval uses canonical seed documents so the real
|
||||
`LookupKnowledgeTool` can retrieve stable evidence from MySQL/Milvus instead of
|
||||
whatever ad hoc documents happen to exist in the local knowledge base.
|
||||
|
||||
Seed documents live in:
|
||||
|
||||
```text
|
||||
eval/rag-retrieval/seed-docs/*.md
|
||||
```
|
||||
|
||||
Each seed doc uses frontmatter fields that are propagated into vector metadata:
|
||||
|
||||
```yaml
|
||||
source: mysql-connection-pool
|
||||
breadcrumb: Database > MySQL > Connection Pool
|
||||
kb_scope: rag-eval
|
||||
```
|
||||
|
||||
Import or reindex the seed docs through the real upload pipeline:
|
||||
|
||||
```powershell
|
||||
.\scripts\prepare_rag_eval_seed.ps1
|
||||
```
|
||||
|
||||
The script runs `RagEvalSeedImporterTest` with `rag.seed.enabled=true`. It
|
||||
deletes the existing document with the same `source`/`docId`, uploads the seed
|
||||
doc through `DocumentManagementService`, updates DB metadata and L0, and rebuilds
|
||||
Milvus chunks.
|
||||
|
||||
`kb_scope` isolates eval data:
|
||||
|
||||
- default application config leaves `retrieval.kb-scope` empty, so legacy docs
|
||||
without `kb_scope` remain searchable;
|
||||
- eval scripts pass `-Dretrieval.kb-scope=rag-eval`, so L0 query hints and L1
|
||||
vector retrieval both use only the canonical eval seed docs;
|
||||
- the fallback retry skips only the L0 category filter, not the `kb_scope`
|
||||
boundary.
|
||||
|
||||
Frontmatter is not embedded as chunk content during upload. It feeds metadata,
|
||||
L0, and document enrichment; only the Markdown body is chunked and embedded.
|
||||
This keeps controlled L0 decoys from becoming semantically relevant just because
|
||||
their frontmatter keywords matched the query.
|
||||
|
||||
## Run
|
||||
|
||||
From the repository root:
|
||||
@@ -36,6 +83,106 @@ python scripts/eval_rag_retrieval.py \
|
||||
--markdown-report eval/rag-retrieval/reports/baseline.md
|
||||
```
|
||||
|
||||
## Generate Fixtures From LookupKnowledgeTool
|
||||
|
||||
Use the snapshot generator when fixtures should reflect the real
|
||||
`LookupKnowledgeTool` pipeline:
|
||||
|
||||
```powershell
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1
|
||||
```
|
||||
|
||||
For the intended live loop, run seed import first:
|
||||
|
||||
```powershell
|
||||
.\scripts\prepare_rag_eval_seed.ps1
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1
|
||||
python scripts\eval_rag_retrieval.py
|
||||
```
|
||||
|
||||
The script runs a Spring test harness:
|
||||
|
||||
```text
|
||||
mvn -q -Dtest=RagLookupSnapshotGeneratorTest -Drag.snapshot.enabled=true -Dretrieval.kb-scope=rag-eval -Dretrieval.vector-store.mode=spring test
|
||||
```
|
||||
|
||||
The generator reads `golden-cases.json`, injects the real `LookupKnowledgeTool`
|
||||
bean, calls `lookupKnowledge(query)` for each case, writes
|
||||
`fixtures/{caseId}.json`, and then runs `eval_rag_retrieval.py` unless
|
||||
`-SkipEval` is provided. It defaults to Spring AI VectorStore mode; pass
|
||||
`-VectorStoreMode sdk` only when intentionally comparing the legacy SDK path.
|
||||
|
||||
Custom paths are supported:
|
||||
|
||||
```powershell
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1 `
|
||||
-Cases eval\rag-retrieval\cases\golden-cases.json `
|
||||
-Fixtures eval\rag-retrieval\fixtures `
|
||||
-RetrievedAt 2026-07-06T00:00:00Z
|
||||
```
|
||||
|
||||
The generator is disabled in normal test runs. It only executes when
|
||||
`rag.snapshot.enabled=true` is provided because it writes repository files and
|
||||
depends on the configured runtime retrieval stack.
|
||||
|
||||
If generated fixtures fail the offline baseline, treat that as a real alignment
|
||||
signal: either the golden expectations need to be adjusted to the current
|
||||
knowledge base, or the knowledge base/indexing path needs to be fixed.
|
||||
|
||||
## Modular RAG Contract
|
||||
|
||||
Fixtures must use the current `lookupResult` shape, which mirrors the
|
||||
`lookup_knowledge` output:
|
||||
|
||||
```text
|
||||
lookupResult.evidenceBlocks
|
||||
lookupResult.contextPack
|
||||
lookupResult.retrievalTrace
|
||||
lookupResult.rerankTrace
|
||||
```
|
||||
|
||||
Golden cases can assert both retrieval quality and pipeline behavior:
|
||||
|
||||
- `expectedSources` / `expectedDocIds`
|
||||
- `expectedBreadcrumbs`
|
||||
- `expectedKeywords`
|
||||
- `expectedSelectedAttempt`
|
||||
- `expectedFallbackReason`
|
||||
- `expectedFallbackReasons`
|
||||
- `expectedEvidenceStatus`
|
||||
- `expectedContextSources`
|
||||
- `expectedRerankTopSource`
|
||||
|
||||
This lets the baseline catch regressions such as losing the expected evidence
|
||||
source, skipping context packing, changing the selected retrieval attempt, or
|
||||
breaking the filtered-vector to unfiltered-retry fallback.
|
||||
|
||||
## Baseline Diff
|
||||
|
||||
To compare a freshly generated report against an existing baseline:
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_retrieval.py \
|
||||
--json-report eval/rag-retrieval/reports/current.json \
|
||||
--markdown-report eval/rag-retrieval/reports/current.md \
|
||||
--compare-to eval/rag-retrieval/reports/baseline.json \
|
||||
--diff-json-report eval/rag-retrieval/reports/baseline-diff.json \
|
||||
--diff-markdown-report eval/rag-retrieval/reports/baseline-diff.md
|
||||
```
|
||||
|
||||
The diff reports aggregate regressions and case-level changes for:
|
||||
|
||||
- pass rate, recall@K, strong hit rate, miss count
|
||||
- pass state
|
||||
- hit level
|
||||
- first expected rank
|
||||
- selected attempt
|
||||
- fallback reason
|
||||
- evidence status
|
||||
- rerank top source
|
||||
|
||||
The command exits non-zero when a case fails or the diff contains a regression.
|
||||
|
||||
## Hit Levels
|
||||
|
||||
- `strong`: expected document is found and breadcrumb or evidence keyword coverage is satisfied.
|
||||
@@ -81,6 +228,6 @@ GET /api/search/similar
|
||||
```
|
||||
|
||||
It writes JSON and Markdown reports with query, topK, result count, top
|
||||
candidates, breadcrumb, score labels, and raw response fields. This is a live
|
||||
results, breadcrumb, score labels, and raw response fields. This is a live
|
||||
smoke check for environment readiness and post-reindex behavior; it does not
|
||||
replace the deterministic offline baseline above.
|
||||
|
||||
@@ -8,8 +8,14 @@
|
||||
"scenario": "chat",
|
||||
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"expectedDocIds": ["mysql-connection-pool"],
|
||||
"expectedSources": ["mysql-connection-pool"],
|
||||
"expectedBreadcrumbs": ["Database > MySQL > Connection Pool"],
|
||||
"expectedKeywords": ["connection pool", "max_connections", "HikariCP"],
|
||||
"expectedSelectedAttempt": "FILTERED_VECTOR",
|
||||
"expectedFallbackReason": null,
|
||||
"expectedEvidenceStatus": "supported",
|
||||
"expectedContextSources": ["mysql-connection-pool"],
|
||||
"expectedRerankTopSource": "mysql-connection-pool",
|
||||
"notes": "Covers precise database troubleshooting retrieval."
|
||||
},
|
||||
{
|
||||
@@ -17,8 +23,14 @@
|
||||
"scenario": "chat",
|
||||
"query": "What is the standard troubleshooting flow for an application incident?",
|
||||
"expectedDocIds": ["incident-diagnosis-flow"],
|
||||
"expectedSources": ["incident-diagnosis-flow"],
|
||||
"expectedBreadcrumbs": ["AIOps > Diagnosis Flow"],
|
||||
"expectedKeywords": ["collect evidence", "verify", "remediation"],
|
||||
"expectedSelectedAttempt": "FILTERED_VECTOR",
|
||||
"expectedFallbackReason": null,
|
||||
"expectedEvidenceStatus": "supported",
|
||||
"expectedContextSources": ["incident-diagnosis-flow"],
|
||||
"expectedRerankTopSource": "incident-diagnosis-flow",
|
||||
"notes": "Covers process-style knowledge where breadcrumb matters."
|
||||
},
|
||||
{
|
||||
@@ -26,8 +38,14 @@
|
||||
"scenario": "aiops",
|
||||
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"expectedDocIds": ["payment-service-latency"],
|
||||
"expectedSources": ["payment-service-latency"],
|
||||
"expectedBreadcrumbs": ["AIOps > Service Alerts > Payment Latency"],
|
||||
"expectedKeywords": ["p95 latency", "payment-service", "downstream dependency"],
|
||||
"expectedSelectedAttempt": "FILTERED_VECTOR",
|
||||
"expectedFallbackReason": null,
|
||||
"expectedEvidenceStatus": "supported",
|
||||
"expectedContextSources": ["payment-service-latency"],
|
||||
"expectedRerankTopSource": "payment-service-latency",
|
||||
"notes": "Covers alert payload terms that should become retrieval hints."
|
||||
},
|
||||
{
|
||||
@@ -35,8 +53,14 @@
|
||||
"scenario": "aiops",
|
||||
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"expectedDocIds": ["aiops-alert-scope-control"],
|
||||
"expectedSources": ["aiops-alert-scope-control"],
|
||||
"expectedBreadcrumbs": ["AIOps > Alert Scope Control"],
|
||||
"expectedKeywords": ["payload", "unrelated active alerts", "scope"],
|
||||
"expectedSelectedAttempt": "FILTERED_VECTOR",
|
||||
"expectedFallbackReason": null,
|
||||
"expectedEvidenceStatus": "supported",
|
||||
"expectedContextSources": ["aiops-alert-scope-control"],
|
||||
"expectedRerankTopSource": "aiops-alert-scope-control",
|
||||
"notes": "Covers scoped alert diagnosis behavior."
|
||||
},
|
||||
{
|
||||
@@ -44,8 +68,14 @@
|
||||
"scenario": "chat",
|
||||
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"expectedDocIds": ["rag-chunk-context-reconstruction"],
|
||||
"expectedSources": ["rag-chunk-context-reconstruction"],
|
||||
"expectedBreadcrumbs": ["RAG > Chunking > Context Reconstruction"],
|
||||
"expectedKeywords": ["neighbor chunk", "same section", "breadcrumb"],
|
||||
"expectedSelectedAttempt": "FILTERED_VECTOR",
|
||||
"expectedFallbackReason": null,
|
||||
"expectedEvidenceStatus": "supported",
|
||||
"expectedContextSources": ["rag-chunk-context-reconstruction"],
|
||||
"expectedRerankTopSource": "rag-chunk-context-reconstruction",
|
||||
"notes": "Covers the known RAG refactor issue around context reconstruction."
|
||||
},
|
||||
{
|
||||
@@ -53,9 +83,30 @@
|
||||
"scenario": "chat",
|
||||
"query": "Should L0 keyword matching decide the final retrieval result?",
|
||||
"expectedDocIds": ["rag-l0-domain-entity-hint"],
|
||||
"expectedSources": ["rag-l0-domain-entity-hint"],
|
||||
"expectedBreadcrumbs": ["RAG > L0 > Domain Entity Hint"],
|
||||
"expectedKeywords": ["domain detector", "entity extractor", "metadata filter"],
|
||||
"expectedSelectedAttempt": "FILTERED_VECTOR",
|
||||
"expectedFallbackReason": null,
|
||||
"expectedEvidenceStatus": "supported",
|
||||
"expectedContextSources": ["rag-l0-domain-entity-hint"],
|
||||
"expectedRerankTopSource": "rag-l0-domain-entity-hint",
|
||||
"notes": "Covers the target L0 role after refactor."
|
||||
},
|
||||
{
|
||||
"caseId": "chat-l0-filter-fallback",
|
||||
"scenario": "chat",
|
||||
"query": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"expectedDocIds": ["rag-l0-filter-fallback"],
|
||||
"expectedSources": ["rag-l0-filter-fallback"],
|
||||
"expectedBreadcrumbs": ["RAG > Fallback > Unfiltered Retry"],
|
||||
"expectedKeywords": ["skip the L0 filter", "unfiltered vector retry", "low quality"],
|
||||
"expectedSelectedAttempt": "UNFILTERED_VECTOR_RETRY",
|
||||
"expectedFallbackReasons": ["filtered_vector_low_quality", "filtered_vector_no_evidence"],
|
||||
"expectedEvidenceStatus": "supported",
|
||||
"expectedContextSources": ["rag-l0-filter-fallback"],
|
||||
"expectedRerankTopSource": "rag-l0-filter-fallback",
|
||||
"notes": "Covers the MVP fallback rule: if filtered L1 is low quality, retry raw query without L0 filter."
|
||||
}
|
||||
]
|
||||
}
|
||||
|
||||
@@ -1,25 +1,81 @@
|
||||
{
|
||||
"caseId": "aiops-payment-latency-alert",
|
||||
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"retrievedAt": "2026-07-05T00:00:00Z",
|
||||
"candidates": [
|
||||
{
|
||||
"rank": 1,
|
||||
"docId": "payment-service-latency",
|
||||
"title": "Payment Service Latency Alert Playbook",
|
||||
"breadcrumb": "AIOps > Service Alerts > Payment Latency",
|
||||
"content": "For payment-service p95 latency alerts, check downstream dependency latency, thread pool saturation, gateway retries, and recent deployment changes.",
|
||||
"score": 0.84,
|
||||
"retrievalLayer": "L1"
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "payment-service-latency",
|
||||
"title": "Payment Service Latency Alert Playbook",
|
||||
"breadcrumb": "AIOps > Service Alerts > Payment Latency",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "For payment-service p95 latency alerts, check downstream dependency latency, thread pool saturation, gateway retries, and recent deployment changes.",
|
||||
"score": 0.84,
|
||||
"hitReasons": ["domain_match:+0.15", "entity_match:+0.20", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"source": "mysql-connection-pool",
|
||||
"title": "MySQL Connection Pool Troubleshooting",
|
||||
"breadcrumb": "Database > MySQL > Connection Pool",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "Database connection pool saturation can increase payment latency when checkout paths wait for connections.",
|
||||
"score": 0.68,
|
||||
"hitReasons": ["keyword_match:+0.10"]
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] Payment Service Latency Alert Playbook\nAIOps > Service Alerts > Payment Latency\nFor payment-service p95 latency alerts, check downstream dependency latency, thread pool saturation, gateway retries, and recent deployment changes.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 236,
|
||||
"includedSources": ["payment-service-latency", "mysql-connection-pool"],
|
||||
"omittedSources": []
|
||||
},
|
||||
{
|
||||
"rank": 2,
|
||||
"docId": "mysql-connection-pool",
|
||||
"title": "MySQL Connection Pool Troubleshooting",
|
||||
"breadcrumb": "Database > MySQL > Connection Pool",
|
||||
"content": "Database connection pool saturation can increase payment latency when checkout paths wait for connections.",
|
||||
"score": 0.68,
|
||||
"retrievalLayer": "L1"
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"rewrittenQuery": "HighLatency payment-service p95 latency alert downstream dependency diagnosis",
|
||||
"categoryFilter": "AIOps",
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["AIOps"],
|
||||
"matched_keywords": ["p95 latency", "payment-service", "downstream dependency"],
|
||||
"entities": ["payment-service", "HighLatency"],
|
||||
"l0_titles": ["Payment Service Latency Alert Playbook"],
|
||||
"l0_match_count": 1
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "HighLatency payment-service p95 latency alert downstream dependency diagnosis",
|
||||
"categoryFilter": "AIOps",
|
||||
"candidateCount": 2,
|
||||
"usable": true,
|
||||
"durationMs": 11,
|
||||
"topScore": 0.84,
|
||||
"topSimilarity": 0.84
|
||||
}
|
||||
]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "payment-service-latency",
|
||||
"baseScore": 0.84,
|
||||
"finalScore": 1.29,
|
||||
"boostReasons": ["domain_match:+0.15", "entity_match:+0.20", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"finalRank": 2,
|
||||
"source": "mysql-connection-pool",
|
||||
"baseScore": 0.68,
|
||||
"finalScore": 0.78,
|
||||
"boostReasons": ["keyword_match:+0.10"]
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,16 +1,65 @@
|
||||
{
|
||||
"caseId": "aiops-prometheus-alert-scope",
|
||||
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"retrievedAt": "2026-07-05T00:00:00Z",
|
||||
"candidates": [
|
||||
{
|
||||
"rank": 1,
|
||||
"docId": "aiops-alert-scope-control",
|
||||
"title": "AIOps Alert Scope Control",
|
||||
"breadcrumb": "AIOps > Alert Scope Control",
|
||||
"content": "When payload mode is active, queryPrometheusAlerts can verify the supplied alert, but unrelated active alerts must remain scoped context and should not become full diagnoses.",
|
||||
"score": 0.9,
|
||||
"retrievalLayer": "L0+L1"
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "aiops-alert-scope-control",
|
||||
"title": "AIOps Alert Scope Control",
|
||||
"breadcrumb": "AIOps > Alert Scope Control",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "When payload mode is used, diagnose the input alert payload and do not expand unrelated active alerts into the main diagnosis scope.",
|
||||
"score": 0.88,
|
||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] AIOps Alert Scope Control\nAIOps > Alert Scope Control\nWhen payload mode is used, diagnose the input alert payload and do not expand unrelated active alerts into the main diagnosis scope.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 188,
|
||||
"includedSources": ["aiops-alert-scope-control"],
|
||||
"omittedSources": []
|
||||
},
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"rewrittenQuery": "AIOps alert payload scope unrelated active alerts diagnosis",
|
||||
"categoryFilter": "AIOps",
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["AIOps"],
|
||||
"matched_keywords": ["payload", "unrelated active alerts", "scope"],
|
||||
"entities": ["alert payload"],
|
||||
"l0_titles": ["AIOps Alert Scope Control"],
|
||||
"l0_match_count": 1
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "AIOps alert payload scope unrelated active alerts diagnosis",
|
||||
"categoryFilter": "AIOps",
|
||||
"candidateCount": 1,
|
||||
"usable": true,
|
||||
"durationMs": 8,
|
||||
"topScore": 0.88,
|
||||
"topSimilarity": 0.88
|
||||
}
|
||||
]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "aiops-alert-scope-control",
|
||||
"baseScore": 0.88,
|
||||
"finalScore": 1.13,
|
||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,25 +1,81 @@
|
||||
{
|
||||
"caseId": "chat-diagnosis-flow",
|
||||
"query": "What is the standard troubleshooting flow for an application incident?",
|
||||
"retrievedAt": "2026-07-05T00:00:00Z",
|
||||
"candidates": [
|
||||
{
|
||||
"rank": 1,
|
||||
"docId": "incident-diagnosis-flow",
|
||||
"title": "Incident Diagnosis Flow",
|
||||
"breadcrumb": "AIOps > Diagnosis Flow",
|
||||
"content": "The standard flow is to collect evidence, identify the suspected fault domain, verify the hypothesis, apply remediation, and confirm recovery.",
|
||||
"score": 0.82,
|
||||
"retrievalLayer": "L1"
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "incident-diagnosis-flow",
|
||||
"title": "Incident Diagnosis Flow",
|
||||
"breadcrumb": "AIOps > Diagnosis Flow",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "The standard flow is to collect evidence, identify the suspected fault domain, verify the hypothesis, apply remediation, and confirm recovery.",
|
||||
"score": 0.82,
|
||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"source": "rag-chunk-context-reconstruction",
|
||||
"title": "RAG Chunk Context Reconstruction",
|
||||
"breadcrumb": "RAG > Chunking > Context Reconstruction",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "Long sections may require neighbor chunk expansion and breadcrumb-aware packing.",
|
||||
"score": 0.55,
|
||||
"hitReasons": []
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] Incident Diagnosis Flow\nAIOps > Diagnosis Flow\nThe standard flow is to collect evidence, identify the suspected fault domain, verify the hypothesis, apply remediation, and confirm recovery.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 192,
|
||||
"includedSources": ["incident-diagnosis-flow", "rag-chunk-context-reconstruction"],
|
||||
"omittedSources": []
|
||||
},
|
||||
{
|
||||
"rank": 2,
|
||||
"docId": "rag-chunk-context-reconstruction",
|
||||
"title": "RAG Chunk Context Reconstruction",
|
||||
"breadcrumb": "RAG > Chunking > Context Reconstruction",
|
||||
"content": "Long sections may require neighbor chunk expansion and breadcrumb-aware packing.",
|
||||
"score": 0.55,
|
||||
"retrievalLayer": "L1"
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "What is the standard troubleshooting flow for an application incident?",
|
||||
"rewrittenQuery": "standard application incident troubleshooting flow collect evidence verify remediation",
|
||||
"categoryFilter": "AIOps",
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["AIOps"],
|
||||
"matched_keywords": ["collect evidence", "verify", "remediation"],
|
||||
"entities": ["application incident"],
|
||||
"l0_titles": ["Incident Diagnosis Flow"],
|
||||
"l0_match_count": 1
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "standard application incident troubleshooting flow collect evidence verify remediation",
|
||||
"categoryFilter": "AIOps",
|
||||
"candidateCount": 2,
|
||||
"usable": true,
|
||||
"durationMs": 10,
|
||||
"topScore": 0.82,
|
||||
"topSimilarity": 0.82
|
||||
}
|
||||
]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "incident-diagnosis-flow",
|
||||
"baseScore": 0.82,
|
||||
"finalScore": 1.07,
|
||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"finalRank": 2,
|
||||
"source": "rag-chunk-context-reconstruction",
|
||||
"baseScore": 0.55,
|
||||
"finalScore": 0.55,
|
||||
"boostReasons": []
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,25 +1,81 @@
|
||||
{
|
||||
"caseId": "chat-l0-domain-hint",
|
||||
"query": "Should L0 keyword matching decide the final retrieval result?",
|
||||
"retrievedAt": "2026-07-05T00:00:00Z",
|
||||
"candidates": [
|
||||
{
|
||||
"rank": 1,
|
||||
"docId": "rag-l0-domain-entity-hint",
|
||||
"title": "RAG L0 Domain Entity Hint",
|
||||
"breadcrumb": "RAG > L0 > Domain Entity Hint",
|
||||
"content": "L0 should be retained as a domain detector, entity extractor, metadata filter generator, and explainability signal, not as the final retrieval decision.",
|
||||
"score": 0.88,
|
||||
"retrievalLayer": "L0"
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "rag-l0-domain-entity-hint",
|
||||
"title": "RAG L0 Domain Entity Hint",
|
||||
"breadcrumb": "RAG > L0 > Domain Entity Hint",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "L0 should be retained as a domain detector, entity extractor, metadata filter generator, and explainability signal, not as the final retrieval decision.",
|
||||
"score": 0.88,
|
||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"source": "rag-l0-l1-fusion-ranking",
|
||||
"title": "RAG L0 L1 Fusion Ranking",
|
||||
"breadcrumb": "RAG > Ranking > Fusion",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "L0 and L1 candidates should eventually be fused rather than handled as an early-return branch.",
|
||||
"score": 0.75,
|
||||
"hitReasons": ["domain_match:+0.15"]
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] RAG L0 Domain Entity Hint\nRAG > L0 > Domain Entity Hint\nL0 should be retained as a domain detector, entity extractor, metadata filter generator, and explainability signal, not as the final retrieval decision.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 219,
|
||||
"includedSources": ["rag-l0-domain-entity-hint", "rag-l0-l1-fusion-ranking"],
|
||||
"omittedSources": []
|
||||
},
|
||||
{
|
||||
"rank": 2,
|
||||
"docId": "rag-l0-l1-fusion-ranking",
|
||||
"title": "RAG L0 L1 Fusion Ranking",
|
||||
"breadcrumb": "RAG > Ranking > Fusion",
|
||||
"content": "L0 and L1 candidates should eventually be fused rather than handled as an early-return branch.",
|
||||
"score": 0.75,
|
||||
"retrievalLayer": "L1"
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "Should L0 keyword matching decide the final retrieval result?",
|
||||
"rewrittenQuery": "RAG L0 keyword matching domain entity hint final retrieval decision",
|
||||
"categoryFilter": "RAG",
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["RAG"],
|
||||
"matched_keywords": ["domain detector", "entity extractor", "metadata filter"],
|
||||
"entities": ["L0"],
|
||||
"l0_titles": ["RAG L0 Domain Entity Hint"],
|
||||
"l0_match_count": 1
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "RAG L0 keyword matching domain entity hint final retrieval decision",
|
||||
"categoryFilter": "RAG",
|
||||
"candidateCount": 2,
|
||||
"usable": true,
|
||||
"durationMs": 9,
|
||||
"topScore": 0.88,
|
||||
"topSimilarity": 0.88
|
||||
}
|
||||
]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "rag-l0-domain-entity-hint",
|
||||
"baseScore": 0.88,
|
||||
"finalScore": 1.13,
|
||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"finalRank": 2,
|
||||
"source": "rag-l0-l1-fusion-ranking",
|
||||
"baseScore": 0.75,
|
||||
"finalScore": 0.9,
|
||||
"boostReasons": ["domain_match:+0.15"]
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,91 @@
|
||||
{
|
||||
"caseId": "chat-l0-filter-fallback",
|
||||
"query": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "rag-l0-filter-fallback",
|
||||
"title": "RAG L0 Filter Fallback",
|
||||
"breadcrumb": "RAG > Fallback > Unfiltered Retry",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "When filtered vector retrieval is low quality, skip the L0 filter and run an unfiltered vector retry with the raw query before returning no evidence.",
|
||||
"score": 0.83,
|
||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"source": "rag-l0-domain-entity-hint",
|
||||
"title": "RAG L0 Domain Entity Hint",
|
||||
"breadcrumb": "RAG > L0 > Domain Entity Hint",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "L0 supplies hints for metadata filtering and explanation, but it should not be treated as final fact evidence.",
|
||||
"score": 0.66,
|
||||
"hitReasons": ["domain_match:+0.15"]
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] RAG L0 Filter Fallback\nRAG > Fallback > Unfiltered Retry\nWhen filtered vector retrieval is low quality, skip the L0 filter and run an unfiltered vector retry with the raw query before returning no evidence.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 214,
|
||||
"includedSources": ["rag-l0-filter-fallback", "rag-l0-domain-entity-hint"],
|
||||
"omittedSources": []
|
||||
},
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"rewrittenQuery": "RAG L0 filtered vector low quality fallback unfiltered retry",
|
||||
"categoryFilter": "RAG",
|
||||
"selectedAttempt": "UNFILTERED_VECTOR_RETRY",
|
||||
"fallbackReason": "filtered_vector_low_quality",
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["RAG"],
|
||||
"matched_keywords": ["L0", "low quality", "unfiltered vector retry"],
|
||||
"entities": ["L0"],
|
||||
"l0_titles": ["RAG L0 Domain Entity Hint"],
|
||||
"l0_match_count": 1
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "RAG L0 filtered vector low quality fallback unfiltered retry",
|
||||
"categoryFilter": "RAG",
|
||||
"candidateCount": 1,
|
||||
"usable": false,
|
||||
"durationMs": 7,
|
||||
"topScore": 1.35,
|
||||
"topSimilarity": 0.325
|
||||
},
|
||||
{
|
||||
"name": "UNFILTERED_VECTOR_RETRY",
|
||||
"query": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"categoryFilter": null,
|
||||
"candidateCount": 2,
|
||||
"usable": true,
|
||||
"durationMs": 13,
|
||||
"topScore": 0.83,
|
||||
"topSimilarity": 0.83
|
||||
}
|
||||
]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "rag-l0-filter-fallback",
|
||||
"baseScore": 0.83,
|
||||
"finalScore": 1.08,
|
||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"finalRank": 2,
|
||||
"source": "rag-l0-domain-entity-hint",
|
||||
"baseScore": 0.66,
|
||||
"finalScore": 0.81,
|
||||
"boostReasons": ["domain_match:+0.15"]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1,25 +1,81 @@
|
||||
{
|
||||
"caseId": "chat-mysql-connection-pool",
|
||||
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"retrievedAt": "2026-07-05T00:00:00Z",
|
||||
"candidates": [
|
||||
{
|
||||
"rank": 1,
|
||||
"docId": "mysql-connection-pool",
|
||||
"title": "MySQL Connection Pool Troubleshooting",
|
||||
"breadcrumb": "Database > MySQL > Connection Pool",
|
||||
"content": "When the connection pool is exhausted, inspect HikariCP active connections, max_connections, slow SQL, leak detection, and database wait events.",
|
||||
"score": 0.86,
|
||||
"retrievalLayer": "L0+L1"
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "mysql-connection-pool",
|
||||
"title": "MySQL Connection Pool Troubleshooting",
|
||||
"breadcrumb": "Database > MySQL > Connection Pool",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "When the connection pool is exhausted, inspect HikariCP active connections, max_connections, slow SQL, leak detection, and database wait events.",
|
||||
"score": 0.86,
|
||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"source": "incident-diagnosis-flow",
|
||||
"title": "Incident Diagnosis Flow",
|
||||
"breadcrumb": "AIOps > Diagnosis Flow",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "Collect evidence, compare metrics and logs, then verify remediation before closing the incident.",
|
||||
"score": 0.61,
|
||||
"hitReasons": []
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] MySQL Connection Pool Troubleshooting\nDatabase > MySQL > Connection Pool\nWhen the connection pool is exhausted, inspect HikariCP active connections, max_connections, slow SQL, leak detection, and database wait events.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 216,
|
||||
"includedSources": ["mysql-connection-pool", "incident-diagnosis-flow"],
|
||||
"omittedSources": []
|
||||
},
|
||||
{
|
||||
"rank": 2,
|
||||
"docId": "incident-diagnosis-flow",
|
||||
"title": "Incident Diagnosis Flow",
|
||||
"breadcrumb": "AIOps > Diagnosis Flow",
|
||||
"content": "Collect evidence, compare metrics and logs, then verify remediation before closing the incident.",
|
||||
"score": 0.61,
|
||||
"retrievalLayer": "L1"
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"rewrittenQuery": "MySQL connection pool exhausted HikariCP max_connections diagnosis",
|
||||
"categoryFilter": "Database",
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["Database", "MySQL"],
|
||||
"matched_keywords": ["connection pool", "HikariCP", "max_connections"],
|
||||
"entities": ["MySQL", "HikariCP"],
|
||||
"l0_titles": ["MySQL Connection Pool Troubleshooting"],
|
||||
"l0_match_count": 1
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "MySQL connection pool exhausted HikariCP max_connections diagnosis",
|
||||
"categoryFilter": "Database",
|
||||
"candidateCount": 2,
|
||||
"usable": true,
|
||||
"durationMs": 12,
|
||||
"topScore": 0.86,
|
||||
"topSimilarity": 0.86
|
||||
}
|
||||
]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "mysql-connection-pool",
|
||||
"baseScore": 0.86,
|
||||
"finalScore": 1.11,
|
||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"finalRank": 2,
|
||||
"source": "incident-diagnosis-flow",
|
||||
"baseScore": 0.61,
|
||||
"finalScore": 0.61,
|
||||
"boostReasons": []
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,25 +1,81 @@
|
||||
{
|
||||
"caseId": "chat-rag-chunk-context",
|
||||
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"retrievedAt": "2026-07-05T00:00:00Z",
|
||||
"candidates": [
|
||||
{
|
||||
"rank": 1,
|
||||
"docId": "rag-chunk-context-reconstruction",
|
||||
"title": "RAG Chunk Context Reconstruction",
|
||||
"breadcrumb": "RAG > Chunking > Context Reconstruction",
|
||||
"content": "After a chunk hit, expand to neighbor chunk candidates from the same section and preserve breadcrumb metadata in the evidence pack.",
|
||||
"score": 0.79,
|
||||
"retrievalLayer": "L1"
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "rag-chunk-context-reconstruction",
|
||||
"title": "RAG Chunk Context Reconstruction",
|
||||
"breadcrumb": "RAG > Chunking > Context Reconstruction",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "After a chunk hit, expand to neighbor chunk candidates from the same section and preserve breadcrumb metadata in the evidence pack.",
|
||||
"score": 0.79,
|
||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"source": "rag-breadcrumb-embedding-gap",
|
||||
"title": "RAG Breadcrumb Embedding Gap",
|
||||
"breadcrumb": "RAG > Embedding > Breadcrumb",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "Embedding title and breadcrumb with content helps recover section semantics.",
|
||||
"score": 0.72,
|
||||
"hitReasons": ["domain_match:+0.15"]
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] RAG Chunk Context Reconstruction\nRAG > Chunking > Context Reconstruction\nAfter a chunk hit, expand to neighbor chunk candidates from the same section and preserve breadcrumb metadata in the evidence pack.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 203,
|
||||
"includedSources": ["rag-chunk-context-reconstruction", "rag-breadcrumb-embedding-gap"],
|
||||
"omittedSources": []
|
||||
},
|
||||
{
|
||||
"rank": 2,
|
||||
"docId": "rag-breadcrumb-embedding-gap",
|
||||
"title": "RAG Breadcrumb Embedding Gap",
|
||||
"breadcrumb": "RAG > Embedding > Breadcrumb",
|
||||
"content": "Embedding title and breadcrumb with content helps recover section semantics.",
|
||||
"score": 0.72,
|
||||
"retrievalLayer": "L1"
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"rewrittenQuery": "RAG chunk context reconstruction neighbor chunk same section breadcrumb",
|
||||
"categoryFilter": "RAG",
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["RAG"],
|
||||
"matched_keywords": ["neighbor chunk", "same section", "breadcrumb"],
|
||||
"entities": ["chunk", "breadcrumb"],
|
||||
"l0_titles": ["RAG Chunk Context Reconstruction"],
|
||||
"l0_match_count": 1
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "RAG chunk context reconstruction neighbor chunk same section breadcrumb",
|
||||
"categoryFilter": "RAG",
|
||||
"candidateCount": 2,
|
||||
"usable": true,
|
||||
"durationMs": 9,
|
||||
"topScore": 0.79,
|
||||
"topSimilarity": 0.79
|
||||
}
|
||||
]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "rag-chunk-context-reconstruction",
|
||||
"baseScore": 0.79,
|
||||
"finalScore": 1.04,
|
||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"finalRank": 2,
|
||||
"source": "rag-breadcrumb-embedding-gap",
|
||||
"baseScore": 0.72,
|
||||
"finalScore": 0.87,
|
||||
"boostReasons": ["domain_match:+0.15"]
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,11 +1,15 @@
|
||||
{
|
||||
"generatedAt": "2026-07-04T17:59:52.172759+00:00",
|
||||
"generatedAt": "2026-07-06T13:37:59.726351+00:00",
|
||||
"caseFile": "eval/rag-retrieval/cases/golden-cases.json",
|
||||
"fixtureDir": "eval/rag-retrieval/fixtures",
|
||||
"aggregate": {
|
||||
"caseCount": 6,
|
||||
"caseCount": 7,
|
||||
"topK": 5,
|
||||
"strongHitCount": 6,
|
||||
"passedCount": 7,
|
||||
"failedCount": 0,
|
||||
"passRate": 1.0,
|
||||
"lookupResultCaseCount": 7,
|
||||
"strongHitCount": 7,
|
||||
"mediumHitCount": 0,
|
||||
"weakHitCount": 0,
|
||||
"missCount": 0,
|
||||
@@ -18,6 +22,7 @@
|
||||
"caseId": "chat-mysql-connection-pool",
|
||||
"scenario": "chat",
|
||||
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
@@ -31,12 +36,22 @@
|
||||
"hikaricp"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"mysql-connection-pool",
|
||||
"incident-diagnosis-flow"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "mysql-connection-pool",
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "chat-diagnosis-flow",
|
||||
"scenario": "chat",
|
||||
"query": "What is the standard troubleshooting flow for an application incident?",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
@@ -50,12 +65,22 @@
|
||||
"remediation"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"incident-diagnosis-flow",
|
||||
"rag-chunk-context-reconstruction"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "incident-diagnosis-flow",
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "aiops-payment-latency-alert",
|
||||
"scenario": "aiops",
|
||||
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
@@ -69,12 +94,22 @@
|
||||
"downstream dependency"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"payment-service-latency",
|
||||
"mysql-connection-pool"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "payment-service-latency",
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "aiops-prometheus-alert-scope",
|
||||
"scenario": "aiops",
|
||||
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
@@ -87,12 +122,21 @@
|
||||
"scope"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"aiops-alert-scope-control"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "aiops-alert-scope-control",
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "chat-rag-chunk-context",
|
||||
"scenario": "chat",
|
||||
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
@@ -106,12 +150,22 @@
|
||||
"breadcrumb"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"rag-chunk-context-reconstruction",
|
||||
"rag-breadcrumb-embedding-gap"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "rag-chunk-context-reconstruction",
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "chat-l0-domain-hint",
|
||||
"scenario": "chat",
|
||||
"query": "Should L0 keyword matching decide the final retrieval result?",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
@@ -125,6 +179,44 @@
|
||||
"metadata filter"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"rag-l0-domain-entity-hint",
|
||||
"rag-l0-l1-fusion-ranking"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "rag-l0-domain-entity-hint",
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "chat-l0-filter-fallback",
|
||||
"scenario": "chat",
|
||||
"query": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:rag-l0-filter-fallback",
|
||||
"2:rag-l0-domain-entity-hint"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"skip the l0 filter",
|
||||
"unfiltered vector retry",
|
||||
"low quality"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "UNFILTERED_VECTOR_RETRY",
|
||||
"fallbackReason": "filtered_vector_low_quality",
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"rag-l0-filter-fallback",
|
||||
"rag-l0-domain-entity-hint"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "rag-l0-filter-fallback",
|
||||
"failedChecks": []
|
||||
}
|
||||
]
|
||||
|
||||
@@ -1,16 +1,20 @@
|
||||
# RAG Retrieval Baseline
|
||||
|
||||
Generated at: `2026-07-04T17:59:52.172759+00:00`
|
||||
Generated at: `2026-07-06T13:37:59.726351+00:00`
|
||||
|
||||
## Aggregate
|
||||
|
||||
| Metric | Value |
|
||||
|---|---:|
|
||||
| Cases | 6 |
|
||||
| Cases | 7 |
|
||||
| Top K | 5 |
|
||||
| Passed | 7 |
|
||||
| Failed | 0 |
|
||||
| Pass rate | 1.0 |
|
||||
| LookupResult fixtures | 7 |
|
||||
| Recall@K | 1.0 |
|
||||
| Strong hit rate | 1.0 |
|
||||
| Strong hits | 6 |
|
||||
| Strong hits | 7 |
|
||||
| Medium hits | 0 |
|
||||
| Weak hits | 0 |
|
||||
| Misses | 0 |
|
||||
@@ -18,11 +22,12 @@ Generated at: `2026-07-04T17:59:52.172759+00:00`
|
||||
|
||||
## Cases
|
||||
|
||||
| Case | Scenario | Hit | First Expected Rank | Top Candidates | Failed Checks |
|
||||
|---|---|---|---:|---|---|
|
||||
| chat-mysql-connection-pool | chat | strong | 1 | 1:mysql-connection-pool<br>2:incident-diagnosis-flow | |
|
||||
| chat-diagnosis-flow | chat | strong | 1 | 1:incident-diagnosis-flow<br>2:rag-chunk-context-reconstruction | |
|
||||
| aiops-payment-latency-alert | aiops | strong | 1 | 1:payment-service-latency<br>2:mysql-connection-pool | |
|
||||
| aiops-prometheus-alert-scope | aiops | strong | 1 | 1:aiops-alert-scope-control | |
|
||||
| chat-rag-chunk-context | chat | strong | 1 | 1:rag-chunk-context-reconstruction<br>2:rag-breadcrumb-embedding-gap | |
|
||||
| chat-l0-domain-hint | chat | strong | 1 | 1:rag-l0-domain-entity-hint<br>2:rag-l0-l1-fusion-ranking | |
|
||||
| Case | Scenario | Pass | Hit | Attempt | Fallback | Evidence | First Expected Rank | Top Candidates | Failed Checks |
|
||||
|---|---|---|---|---|---|---|---:|---|---|
|
||||
| chat-mysql-connection-pool | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:mysql-connection-pool<br>2:incident-diagnosis-flow | |
|
||||
| chat-diagnosis-flow | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:incident-diagnosis-flow<br>2:rag-chunk-context-reconstruction | |
|
||||
| aiops-payment-latency-alert | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:payment-service-latency<br>2:mysql-connection-pool | |
|
||||
| aiops-prometheus-alert-scope | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:aiops-alert-scope-control | |
|
||||
| chat-rag-chunk-context | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-chunk-context-reconstruction<br>2:rag-breadcrumb-embedding-gap | |
|
||||
| chat-l0-domain-hint | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-l0-domain-entity-hint<br>2:rag-l0-l1-fusion-ranking | |
|
||||
| chat-l0-filter-fallback | chat | true | strong | UNFILTERED_VECTOR_RETRY | filtered_vector_low_quality | supported | 1 | 1:rag-l0-filter-fallback<br>2:rag-l0-domain-entity-hint | |
|
||||
|
||||
@@ -0,0 +1,26 @@
|
||||
---
|
||||
title: AIOps Alert Scope Control
|
||||
keywords: [alert payload, unrelated active alerts, scope control]
|
||||
summary: Keep diagnosis scoped to the request payload and avoid diagnosing unrelated active alerts.
|
||||
category: aiops
|
||||
source: aiops-alert-scope-control
|
||||
breadcrumb: AIOps > Alert Scope Control
|
||||
kb_scope: rag-eval
|
||||
covers: [alert scope, payload, active alerts]
|
||||
when_to_retrieve: Use when an AIOps request includes a concrete alert payload and scope boundaries matter.
|
||||
---
|
||||
|
||||
# AIOps
|
||||
|
||||
## Alert Scope Control
|
||||
|
||||
When an AIOps request already includes an alert payload, the agent should diagnose that payload first.
|
||||
It must not expand the task into unrelated active alerts unless the user asks for broad alert triage.
|
||||
|
||||
Scope rules:
|
||||
|
||||
1. Treat the provided payload as the primary incident boundary.
|
||||
2. Use unrelated active alerts only as correlation evidence when they share service, dependency, time window, or trace context.
|
||||
3. Do not replace the requested alert with a louder but unrelated alert.
|
||||
|
||||
This runbook anchors payload, unrelated active alerts, and scope behavior.
|
||||
@@ -0,0 +1,27 @@
|
||||
---
|
||||
title: Incident Diagnosis Flow
|
||||
keywords: [standard troubleshooting flow, application incident, collect evidence, verify, remediation]
|
||||
summary: Standard flow for diagnosing application incidents with evidence, hypothesis verification, and remediation.
|
||||
category: ops
|
||||
source: incident-diagnosis-flow
|
||||
breadcrumb: AIOps > Diagnosis Flow
|
||||
kb_scope: rag-eval
|
||||
covers: [incident diagnosis, evidence collection, remediation]
|
||||
when_to_retrieve: Use when the user asks for a standard troubleshooting flow or incident diagnosis sequence.
|
||||
---
|
||||
|
||||
# AIOps
|
||||
|
||||
## Diagnosis Flow
|
||||
|
||||
The standard troubleshooting flow is evidence first, hypothesis second, remediation last.
|
||||
|
||||
Recommended sequence:
|
||||
|
||||
1. Collect evidence from alerts, metrics, logs, traces, deployments, and recent configuration changes.
|
||||
2. Define a small hypothesis that explains the observed symptoms.
|
||||
3. Verify the hypothesis with a targeted metric, log query, or reproduction step.
|
||||
4. Choose remediation that directly addresses the verified cause.
|
||||
5. Record the outcome and the evidence used to make the decision.
|
||||
|
||||
Do not skip collect evidence, verify, and remediation ordering during an application incident.
|
||||
@@ -0,0 +1,29 @@
|
||||
---
|
||||
title: MySQL Connection Pool Runbook
|
||||
keywords: [MySQL connection pool, pool exhausted, max_connections, HikariCP]
|
||||
summary: Diagnose exhausted MySQL connection pools and distinguish application leaks from database limits.
|
||||
category: database
|
||||
source: mysql-connection-pool
|
||||
breadcrumb: Database > MySQL > Connection Pool
|
||||
kb_scope: rag-eval
|
||||
covers: [mysql, connection pool, database capacity]
|
||||
when_to_retrieve: Use when MySQL clients report exhausted pools, connection acquisition timeout, max_connections pressure, or HikariCP saturation.
|
||||
---
|
||||
|
||||
# Database
|
||||
|
||||
## MySQL
|
||||
|
||||
### Connection Pool
|
||||
|
||||
When MySQL connection pool is exhausted, first compare application pool usage with database `max_connections`.
|
||||
For HikariCP, check `active`, `idle`, `pending`, and connection acquisition timeout metrics.
|
||||
|
||||
Recommended diagnosis:
|
||||
|
||||
1. Verify whether HikariCP active connections stay near maximum while pending threads grow.
|
||||
2. Check MySQL `Threads_connected`, `Threads_running`, and `max_connections`.
|
||||
3. Inspect slow SQL and long transactions that keep connections checked out.
|
||||
4. If the database is healthy, look for application connection leaks or missing transaction boundaries.
|
||||
|
||||
Use this runbook as evidence for connection pool, max_connections, and HikariCP incidents.
|
||||
@@ -0,0 +1,28 @@
|
||||
---
|
||||
title: Payment Service Latency Alert
|
||||
keywords: [HighLatency, payment-service, p95 latency, downstream dependency]
|
||||
summary: Diagnose payment-service p95 latency alerts and identify downstream dependency bottlenecks.
|
||||
category: aiops
|
||||
source: payment-service-latency
|
||||
breadcrumb: AIOps > Service Alerts > Payment Latency
|
||||
kb_scope: rag-eval
|
||||
covers: [payment-service, latency, downstream dependency]
|
||||
when_to_retrieve: Use when an alert mentions payment-service, HighLatency, or elevated p95 latency.
|
||||
---
|
||||
|
||||
# AIOps
|
||||
|
||||
## Service Alerts
|
||||
|
||||
### Payment Latency
|
||||
|
||||
For `HighLatency` alerts on `payment-service`, treat p95 latency as the primary symptom.
|
||||
|
||||
Diagnosis steps:
|
||||
|
||||
1. Confirm whether p95 latency is isolated to payment-service or shared across upstream callers.
|
||||
2. Compare payment-service latency with downstream dependency latency for gateway, risk, and order services.
|
||||
3. Check connection pool wait time, retry spikes, and timeout rates.
|
||||
4. If downstream dependency latency increased first, classify payment-service as affected rather than root cause.
|
||||
|
||||
The expected evidence terms are p95 latency, payment-service, and downstream dependency.
|
||||
@@ -0,0 +1,28 @@
|
||||
---
|
||||
title: RAG Chunk Context Reconstruction
|
||||
keywords: [split into multiple chunks, retrieval context, neighbor chunk, same section, breadcrumb context]
|
||||
summary: Preserve context when long RAG sections are split into multiple retrievable chunks.
|
||||
category: rag
|
||||
source: rag-chunk-context-reconstruction
|
||||
breadcrumb: RAG > Chunking > Context Reconstruction
|
||||
kb_scope: rag-eval
|
||||
covers: [rag chunking, context packing, breadcrumbs]
|
||||
when_to_retrieve: Use when a retrieval question asks how to preserve context across split chunks.
|
||||
---
|
||||
|
||||
# RAG
|
||||
|
||||
## Chunking
|
||||
|
||||
### Context Reconstruction
|
||||
|
||||
When a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.
|
||||
|
||||
Recommended behavior:
|
||||
|
||||
1. Store the breadcrumb with every chunk.
|
||||
2. Preserve the same section identity across adjacent chunks.
|
||||
3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.
|
||||
4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.
|
||||
|
||||
The key concepts are neighbor chunk, same section, and breadcrumb.
|
||||
@@ -0,0 +1,29 @@
|
||||
---
|
||||
title: RAG L0 Domain Entity Hint
|
||||
keywords: [L0 keyword matching, final retrieval result, domain detector, entity extractor, metadata filter]
|
||||
summary: Define L0 as a query transformation hint layer instead of final retrieval evidence.
|
||||
category: rag
|
||||
source: rag-l0-domain-entity-hint
|
||||
breadcrumb: RAG > L0 > Domain Entity Hint
|
||||
kb_scope: rag-eval
|
||||
covers: [l0 hint, query transformation, metadata filter]
|
||||
when_to_retrieve: Use when a question asks whether L0 should decide final retrieval or only provide hints.
|
||||
---
|
||||
|
||||
# RAG
|
||||
|
||||
## L0
|
||||
|
||||
### Domain Entity Hint
|
||||
|
||||
L0 keyword matching should not decide the final retrieval result.
|
||||
In the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.
|
||||
|
||||
The output can provide:
|
||||
|
||||
1. Candidate domain hints.
|
||||
2. Matched entities and keywords.
|
||||
3. An optional metadata filter for the first vector retrieval attempt.
|
||||
|
||||
Final evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.
|
||||
The important terms are domain detector, entity extractor, and metadata filter.
|
||||
@@ -0,0 +1,19 @@
|
||||
---
|
||||
title: RAG L0 Filter Decoy
|
||||
keywords: [over-filtered by L0, filtered vector search, low quality evidence]
|
||||
summary: Decoy document used to force the first filtered retrieval attempt into a low-quality category.
|
||||
category: overfilter-decoy
|
||||
source: rag-l0-filter-decoy
|
||||
breadcrumb: RAG > Fallback > Decoy
|
||||
kb_scope: rag-eval
|
||||
covers: [fallback test decoy]
|
||||
when_to_retrieve: Use only as a controlled eval decoy for over-filter fallback testing.
|
||||
---
|
||||
|
||||
# Release Calendar
|
||||
|
||||
## Approval Window
|
||||
|
||||
This document describes an unrelated release calendar approval window.
|
||||
It intentionally avoids the real fallback instructions so the filtered retrieval
|
||||
attempt is low quality and the retriever must retry without the L0 category filter.
|
||||
@@ -0,0 +1,25 @@
|
||||
---
|
||||
title: RAG L0 Filter Fallback
|
||||
keywords: [golden retry contract, second pass retrieval]
|
||||
summary: Retry the raw query without the L0 category filter when filtered vector evidence is missing or low quality.
|
||||
category: fallback
|
||||
source: rag-l0-filter-fallback
|
||||
breadcrumb: RAG > Fallback > Unfiltered Retry
|
||||
kb_scope: rag-eval
|
||||
covers: [fallback, unfiltered retry, retrieval quality]
|
||||
when_to_retrieve: Use when validating the fallback contract for low-quality filtered vector retrieval.
|
||||
---
|
||||
|
||||
# RAG
|
||||
|
||||
## Fallback
|
||||
|
||||
### Unfiltered Retry
|
||||
|
||||
If the first vector search is over-constrained by an L0 metadata filter and returns low quality evidence,
|
||||
the retriever should skip the L0 filter and run an unfiltered vector retry with the original query.
|
||||
|
||||
The fallback reason should be `filtered_vector_low_quality` when the filtered candidate exists but is below the
|
||||
reference threshold. If there is no usable evidence at all, use `filtered_vector_no_evidence`.
|
||||
|
||||
This document is the expected evidence for skip the L0 filter, unfiltered vector retry, and low quality behavior.
|
||||
Reference in New Issue
Block a user