# rag-retrieval-quality-score Specification ## Purpose Unify knowledge retrieval score labels and quality normalization so dense and hybrid modes share one post-process pipeline without L2 faking or keyword boost re-ranking. ## Requirements ### Requirement: Primary score labels SHALL be only dense or hybrid Knowledge search hits used by `lookup_knowledge` SHALL set `scoreLabel` to `dense` or `hybrid` (after any legacy alias canonicalization). The system SHALL NOT treat `bm25_only_no_dense`, `rrf_fused`, or `l2_distance` as distinct first-class labels in new emissions. #### Scenario: Dense mode labels hits as dense - **WHEN** `retrieval.search.mode` is `dense` and search returns hits - **THEN** each hit SHALL have `scoreLabel` canonicalizing to `dense` - **AND** `score` SHALL be the dense L2 distance #### Scenario: Hybrid mode labels hits as hybrid - **WHEN** `retrieval.search.mode` is `hybrid` and search returns hits - **THEN** each hit SHALL have `scoreLabel` canonicalizing to `hybrid` - **AND** hit order SHALL follow the hybrid/RRF result order via `originalRank` #### Scenario: Legacy aliases canonicalize - **WHEN** a candidate carries a legacy label such as `l2_distance` or `rrf_fused` or `bm25_only_no_dense` - **THEN** quality normalization SHALL canonicalize it to `dense` or `hybrid` respectively before computing quality ### Requirement: Hybrid search SHALL NOT overwrite scores with dense L2 for post-process When hybrid search runs, the store SHALL NOT replace hybrid hit scores with a parallel dense L2 map for the purpose of post-process thresholds, and SHALL NOT assign max-L2 weak placeholders under a `bm25_only_*` primary label. #### Scenario: No L2 enrichment overwrite - **WHEN** hybrid search completes - **THEN** post-process input scores SHALL NOT be forced to dense L2 solely for threshold compatibility - **AND** the system SHALL NOT emit `bm25_only_no_dense` as the primary score label on new hits ### Requirement: Quality score SHALL be produced by a single normalizer The pipeline SHALL compute a `qualityScore` in `[0, 1]` (higher is better) using one normalizer API that branches only on canonical label. #### Scenario: Dense quality from L2 - **WHEN** label is `dense` and score is L2 distance `d` with configured `maxL2Distance` - **THEN** `qualityScore` SHALL equal `max(0, 1 - min(d, maxL2Distance) / maxL2Distance)` (null score → 0) #### Scenario: Hybrid quality from rank - **WHEN** label is `hybrid` and `originalRank` is `r` within a candidate batch of size `n` (`n >= 1`) - **THEN** `qualityScore` SHALL be a monotonically non-increasing function of `r` over that batch - **AND** rank `1` SHALL map to `1.0` when `n >= 1` - **AND** the mapping SHALL NOT require dense L2 or RRF raw magnitude ### Requirement: Post-process SHALL preserve retrieval rank order Evidence post-processing SHALL order candidates by `originalRank` ascending (stable tie-break allowed). It SHALL NOT re-order primarily by L0 domain/entity/keyword/source_type additive boosts. #### Scenario: Keyword overlap does not promote lower rank - **WHEN** candidate A has `originalRank=1` and candidate B has `originalRank=2` - **AND** B matches more L0 keywords via string contains than A - **THEN** after post-process acceptance order, A SHALL appear before B among accepted blocks (subject only to dedup/caps removing one of them) #### Scenario: Structural caps still apply after rank order - **WHEN** more than `rag.max-chunks-per-document` chunks share a docId - **THEN** only the best-ranked (lowest `originalRank`) up to the cap SHALL remain - **AND** `rag.return-n` SHALL still bound total blocks ### Requirement: Relevance and low-quality gates SHALL use qualityScore only `relevance_level`, completeness hints, `topSimilarity` (or equivalent top quality field), and category-filter low-quality retry decisions SHALL use the normalized `qualityScore`, not raw L2-under-hybrid fakes and not keyword-boosted final scores. #### Scenario: Low quality uses top qualityScore - **WHEN** post-process finishes with at least one evidence block - **THEN** low-quality detection SHALL compare the top `qualityScore` to the configured reference threshold - **AND** SHALL NOT require L0 hint contains-match to treat the result as usable when quality meets threshold #### Scenario: PRECISE does not require hint support - **WHEN** top `qualityScore` is at or above the highly-relevant threshold - **THEN** the system MAY assign `PRECISE` or `HIGHLY_RELEVANT` without requiring domain/entity/keyword contains support ### Requirement: Dense mode remains available for recall comparison Configuration `retrieval.search.mode=dense` SHALL remain supported as a same-collection baseline that runs dense ANN only, without removing hybrid as the default production mode. #### Scenario: Mode dense still callable - **WHEN** mode is `dense` - **THEN** search SHALL call dense ANN only and label hits `dense` ### Requirement: Hybrid quality gates MAY use optional dense distance without reordering When hybrid search attaches an optional dense L2 distance for a hit, quality normalization and low-quality gates MAY use that distance for absolute similarity. Retrieval order SHALL remain the hybrid/RRF order and the primary scoreLabel SHALL remain hybrid. #### Scenario: Dense distance does not replace hybrid label - **WHEN** a hybrid hit includes denseDistance - **THEN** scoreLabel SHALL still canonicalize to hybrid - **AND** post-process sort order SHALL still follow originalRank from hybrid results