Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
5.4 KiB
rag-retrieval-quality-score Specification
Purpose
Unify knowledge retrieval score labels and quality normalization so dense and hybrid modes share one post-process pipeline without L2 faking or keyword boost re-ranking.
Requirements
Requirement: Primary score labels SHALL be only dense or hybrid
Knowledge search hits used by lookup_knowledge SHALL set scoreLabel to dense or hybrid (after any legacy alias canonicalization). The system SHALL NOT treat bm25_only_no_dense, rrf_fused, or l2_distance as distinct first-class labels in new emissions.
Scenario: Dense mode labels hits as dense
- WHEN
retrieval.search.modeisdenseand search returns hits - THEN each hit SHALL have
scoreLabelcanonicalizing todense - AND
scoreSHALL be the dense L2 distance
Scenario: Hybrid mode labels hits as hybrid
- WHEN
retrieval.search.modeishybridand search returns hits - THEN each hit SHALL have
scoreLabelcanonicalizing tohybrid - AND hit order SHALL follow the hybrid/RRF result order via
originalRank
Scenario: Legacy aliases canonicalize
- WHEN a candidate carries a legacy label such as
l2_distanceorrrf_fusedorbm25_only_no_dense - THEN quality normalization SHALL canonicalize it to
denseorhybridrespectively before computing quality
Requirement: Hybrid search SHALL NOT overwrite scores with dense L2 for post-process
When hybrid search runs, the store SHALL NOT replace hybrid hit scores with a parallel dense L2 map for the purpose of post-process thresholds, and SHALL NOT assign max-L2 weak placeholders under a bm25_only_* primary label.
Scenario: No L2 enrichment overwrite
- WHEN hybrid search completes
- THEN post-process input scores SHALL NOT be forced to dense L2 solely for threshold compatibility
- AND the system SHALL NOT emit
bm25_only_no_denseas the primary score label on new hits
Requirement: Quality score SHALL be produced by a single normalizer
The pipeline SHALL compute a qualityScore in [0, 1] (higher is better) using one normalizer API that branches only on canonical label.
Scenario: Dense quality from L2
- WHEN label is
denseand score is L2 distancedwith configuredmaxL2Distance - THEN
qualityScoreSHALL equalmax(0, 1 - min(d, maxL2Distance) / maxL2Distance)(null score → 0)
Scenario: Hybrid quality from rank
- WHEN label is
hybridandoriginalRankisrwithin a candidate batch of sizen(n >= 1) - THEN
qualityScoreSHALL be a monotonically non-increasing function ofrover that batch - AND rank
1SHALL map to1.0whenn >= 1 - AND the mapping SHALL NOT require dense L2 or RRF raw magnitude
Requirement: Post-process SHALL preserve retrieval rank order
Evidence post-processing SHALL order candidates by originalRank ascending (stable tie-break allowed). It SHALL NOT re-order primarily by L0 domain/entity/keyword/source_type additive boosts.
Scenario: Keyword overlap does not promote lower rank
- WHEN candidate A has
originalRank=1and candidate B hasoriginalRank=2 - AND B matches more L0 keywords via string contains than A
- THEN after post-process acceptance order, A SHALL appear before B among accepted blocks (subject only to dedup/caps removing one of them)
Scenario: Structural caps still apply after rank order
- WHEN more than
rag.max-chunks-per-documentchunks share a docId - THEN only the best-ranked (lowest
originalRank) up to the cap SHALL remain - AND
rag.return-nSHALL still bound total blocks
Requirement: Relevance and low-quality gates SHALL use qualityScore only
relevance_level, completeness hints, topSimilarity (or equivalent top quality field), and category-filter low-quality retry decisions SHALL use the normalized qualityScore, not raw L2-under-hybrid fakes and not keyword-boosted final scores.
Scenario: Low quality uses top qualityScore
- WHEN post-process finishes with at least one evidence block
- THEN low-quality detection SHALL compare the top
qualityScoreto the configured reference threshold - AND SHALL NOT require L0 hint contains-match to treat the result as usable when quality meets threshold
Scenario: PRECISE does not require hint support
- WHEN top
qualityScoreis at or above the highly-relevant threshold - THEN the system MAY assign
PRECISEorHIGHLY_RELEVANTwithout requiring domain/entity/keyword contains support
Requirement: Dense mode remains available for recall comparison
Configuration retrieval.search.mode=dense SHALL remain supported as a same-collection baseline that runs dense ANN only, without removing hybrid as the default production mode.
Scenario: Mode dense still callable
- WHEN mode is
dense - THEN search SHALL call dense ANN only and label hits
dense
Requirement: Hybrid quality gates MAY use optional dense distance without reordering
When hybrid search attaches an optional dense L2 distance for a hit, quality normalization and low-quality gates MAY use that distance for absolute similarity. Retrieval order SHALL remain the hybrid/RRF order and the primary scoreLabel SHALL remain hybrid.
Scenario: Dense distance does not replace hybrid label
- WHEN a hybrid hit includes denseDistance
- THEN scoreLabel SHALL still canonicalize to hybrid
- AND post-process sort order SHALL still follow originalRank from hybrid results