# Design: rag-quality-score-unify ## Context - Knowledge path already uses single `MilvusHybridKnowledgeStore` (MilvusClientV2) with dense + BM25 + RRF. - Chunk-level `evidenceKey` dedup and `retrieve-k` / `return-n` are landed. - Gap: hybrid ordering is RRF, but post-process still pretends scores are L2 and re-ranks with L0 keyword contains boosts. - Prior design in `rag-bm25-hybrid-drop-sdk` required dense L2 enrichment for threshold compatibility — **this change supersedes that decision**. Stakeholders: `lookup_knowledge` internal pipeline; Agent ACI field *names* unchanged; operators comparing `retrieval.search.mode=dense|hybrid`. ## Goals / Non-Goals **Goals:** 1. First-class `scoreLabel` values: only `dense` | `hybrid` (aliases canonicalize). 2. Single `toQualityScore` adapter; post-process is label-agnostic. 3. Preserve retrieval `originalRank` as sort authority; remove keyword/domain boost re-ranking. 4. Hybrid quality = pure rank mapping over the current candidate batch (confirmed). 5. Keep `mode=dense` for offline recall comparison; production default remains hybrid. **Non-Goals:** - Cross-encoder / query rewrite / neighbor chunks. - Schema rebuild or collection rename. - Changing Agent-facing ACI JSON field names. - Configurable `max(rank, denseSim)` quality (future). ## Decisions ### D1 — Two labels only | label | `score` meaning | quality mapping | |---|---|---| | `dense` | L2 distance (smaller better) | `1 - clamp(l2)/maxL2Distance` | | `hybrid` | engine fused score optional in `rawScore`; **not** used as L2 | `rankToQuality(originalRank, batchSize)` | Canonicalize legacy strings: `l2_distance`→dense; `rrf_fused` / `bm25_only_*`→hybrid. **Why not keep `bm25_only`:** it is not a search mode; it was a L2-fake patch. Hybrid path hits are all `hybrid`. ### D2 — Stop dense L2 overwrite on hybrid hits `searchHybrid` SHALL: 1. Run `hybridSearch` + RRFRanker. 2. Emit hits in RRF order with `scoreLabel=hybrid`, `originalRank=1..n`. 3. Set `rawScore` from engine when present; `score` MAY equal raw fused score or rank placeholder — MUST NOT be replaced by dense L2 for post-process consumption. 4. MUST NOT set `bm25_only_no_dense` or force `score=maxL2Distance` for threshold faking. 5. MUST NOT run a parallel dense search solely to rewrite scores (slice-1 pure rank quality). `searchDense` SHALL emit `scoreLabel=dense` and L2 in `score`. ### D3 — `RetrievalScoreNormalizer` is the only label branch ```text qualityScore = RetrievalScoreNormalizer.toQualityScore( scoreLabel, score, originalRank, batchSize, maxL2Distance) ``` - Hybrid: linear rank map — rank 1 → 1.0; rank n → ~1/n floor so last item > 0. - Dense: existing L2 formula (behavior parity for dense mode). Post-processor, `isLowQuality`, and `relevance_level` consume **only** `qualityScore` (exposed today as `baseScore` / `topSimilarity` fields for minimal DTO churn). ### D4 — Post-process sort and boosts ```text order = originalRank ASC, then stable evidenceKey // NO finalScore = base + 0.15 domain + ... ``` - Remove additive boosts from sort key and from `finalScore` used for ordering. - Optional: if L0 hint string matches, append explanatory `hitReasons` only (e.g. `l0_keyword_overlap`) — zero score delta. - Keep: evidenceKey dedup, max-chunks-per-document, return-n, excerpt truncate. - `relevance_level`: compare top `qualityScore` to existing thresholds; **remove** `hasHintSupport` gate for PRECISE. - `isLowQuality` / category unfiltered retry: unchanged control flow, new score semantics. ### D5 — Trace fields - `RerankTrace` may keep `baseScore`/`finalScore` names but both equal `qualityScore` when boosts are zero; `boostReasons` empty or explanation-only reasons without `:+0.xx` score deltas. - Prefer renaming comments to “quality trace”; no Agent contract change required. ### D6 — WIP files Workspace may contain draft `RetrievalScoreLabels` / `RetrievalScoreNormalizer`. Apply MUST align them to this design (or replace) and wire call sites; drafts alone are not done. ## Interface impact - **Level: L2** — internal DTO semantics (`score`, `scoreLabel`), post-process ordering and relevance distribution. - Agent ACI: field names stable; `relevance_level` *values distribution* may change (accepted behavioral change). ## Data flow (target) ```text VectorSearchService (mode dense|hybrid) -> hits{ rank, score, scoreLabel=dense|hybrid, rawScore? } -> KnowledgeDocumentRetriever candidates -> KnowledgeEvidencePostProcessor for each: qualityScore = toQualityScore(...) sort by originalRank dedup / caps / return-n relevance + topSimilarity from qualityScore -> pack / assemble / project ``` ## Risks / Trade-offs | Risk | Mitigation | |---|---| | Rank→quality not comparable across queries | Document; thresholds may need later tune; dense mode still L2-absolute | | Hybrid top quality always high if batch small | batchSize = candidate list size after retrieve; rank1 always 1.0 by design for “best of this round” | | PRECISE without hint support more often | Accepted; semantic/rank quality no longer gated on contains | | Tests assert boost re-order | Update `LookupKnowledgeToolTest.rerankUsesHintMatches...` | | Old label strings in sidecar/eval | canonicalize in normalizer | ## Migration Plan 1. Deploy code; no Milvus schema migration. 2. Default `retrieval.search.mode=hybrid` unchanged. 3. Rollback: revert change; old L2-enrichment behavior returns. 4. Optional ops: A/B dense vs hybrid recall using mode switch (unchanged capability). ## Open Questions - None for slice-1 (Q8 confirmed: pure rank). - Follow-up: threshold calibration after live traces; optional denseSim blend. ## Audit notes (inline) Module chain: Tool → Retriever → Store → PostProcessor → Packer → Projector. Ownership: retrieval scores owned by store+normalizer; evidence assembly by post-processor; Agent view by projector. No new cross-module lifecycle. Couples only internal RAG pipeline. Supersedes hybrid L2-enrichment ADR-equivalent decision from bm25-hybrid change.