Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
105 lines
5.4 KiB
Markdown
105 lines
5.4 KiB
Markdown
# rag-retrieval-quality-score Specification
|
|
|
|
## Purpose
|
|
|
|
Unify knowledge retrieval score labels and quality normalization so dense and hybrid modes share one post-process pipeline without L2 faking or keyword boost re-ranking.
|
|
|
|
## Requirements
|
|
|
|
### Requirement: Primary score labels SHALL be only dense or hybrid
|
|
|
|
Knowledge search hits used by `lookup_knowledge` SHALL set `scoreLabel` to `dense` or `hybrid` (after any legacy alias canonicalization). The system SHALL NOT treat `bm25_only_no_dense`, `rrf_fused`, or `l2_distance` as distinct first-class labels in new emissions.
|
|
|
|
#### Scenario: Dense mode labels hits as dense
|
|
|
|
- **WHEN** `retrieval.search.mode` is `dense` and search returns hits
|
|
- **THEN** each hit SHALL have `scoreLabel` canonicalizing to `dense`
|
|
- **AND** `score` SHALL be the dense L2 distance
|
|
|
|
#### Scenario: Hybrid mode labels hits as hybrid
|
|
|
|
- **WHEN** `retrieval.search.mode` is `hybrid` and search returns hits
|
|
- **THEN** each hit SHALL have `scoreLabel` canonicalizing to `hybrid`
|
|
- **AND** hit order SHALL follow the hybrid/RRF result order via `originalRank`
|
|
|
|
#### Scenario: Legacy aliases canonicalize
|
|
|
|
- **WHEN** a candidate carries a legacy label such as `l2_distance` or `rrf_fused` or `bm25_only_no_dense`
|
|
- **THEN** quality normalization SHALL canonicalize it to `dense` or `hybrid` respectively before computing quality
|
|
|
|
### Requirement: Hybrid search SHALL NOT overwrite scores with dense L2 for post-process
|
|
|
|
When hybrid search runs, the store SHALL NOT replace hybrid hit scores with a parallel dense L2 map for the purpose of post-process thresholds, and SHALL NOT assign max-L2 weak placeholders under a `bm25_only_*` primary label.
|
|
|
|
#### Scenario: No L2 enrichment overwrite
|
|
|
|
- **WHEN** hybrid search completes
|
|
- **THEN** post-process input scores SHALL NOT be forced to dense L2 solely for threshold compatibility
|
|
- **AND** the system SHALL NOT emit `bm25_only_no_dense` as the primary score label on new hits
|
|
|
|
### Requirement: Quality score SHALL be produced by a single normalizer
|
|
|
|
The pipeline SHALL compute a `qualityScore` in `[0, 1]` (higher is better) using one normalizer API that branches only on canonical label.
|
|
|
|
#### Scenario: Dense quality from L2
|
|
|
|
- **WHEN** label is `dense` and score is L2 distance `d` with configured `maxL2Distance`
|
|
- **THEN** `qualityScore` SHALL equal `max(0, 1 - min(d, maxL2Distance) / maxL2Distance)` (null score → 0)
|
|
|
|
#### Scenario: Hybrid quality from rank
|
|
|
|
- **WHEN** label is `hybrid` and `originalRank` is `r` within a candidate batch of size `n` (`n >= 1`)
|
|
- **THEN** `qualityScore` SHALL be a monotonically non-increasing function of `r` over that batch
|
|
- **AND** rank `1` SHALL map to `1.0` when `n >= 1`
|
|
- **AND** the mapping SHALL NOT require dense L2 or RRF raw magnitude
|
|
|
|
### Requirement: Post-process SHALL preserve retrieval rank order
|
|
|
|
Evidence post-processing SHALL order candidates by `originalRank` ascending (stable tie-break allowed). It SHALL NOT re-order primarily by L0 domain/entity/keyword/source_type additive boosts.
|
|
|
|
#### Scenario: Keyword overlap does not promote lower rank
|
|
|
|
- **WHEN** candidate A has `originalRank=1` and candidate B has `originalRank=2`
|
|
- **AND** B matches more L0 keywords via string contains than A
|
|
- **THEN** after post-process acceptance order, A SHALL appear before B among accepted blocks (subject only to dedup/caps removing one of them)
|
|
|
|
#### Scenario: Structural caps still apply after rank order
|
|
|
|
- **WHEN** more than `rag.max-chunks-per-document` chunks share a docId
|
|
- **THEN** only the best-ranked (lowest `originalRank`) up to the cap SHALL remain
|
|
- **AND** `rag.return-n` SHALL still bound total blocks
|
|
|
|
### Requirement: Relevance and low-quality gates SHALL use qualityScore only
|
|
|
|
`relevance_level`, completeness hints, `topSimilarity` (or equivalent top quality field), and category-filter low-quality retry decisions SHALL use the normalized `qualityScore`, not raw L2-under-hybrid fakes and not keyword-boosted final scores.
|
|
|
|
#### Scenario: Low quality uses top qualityScore
|
|
|
|
- **WHEN** post-process finishes with at least one evidence block
|
|
- **THEN** low-quality detection SHALL compare the top `qualityScore` to the configured reference threshold
|
|
- **AND** SHALL NOT require L0 hint contains-match to treat the result as usable when quality meets threshold
|
|
|
|
#### Scenario: PRECISE does not require hint support
|
|
|
|
- **WHEN** top `qualityScore` is at or above the highly-relevant threshold
|
|
- **THEN** the system MAY assign `PRECISE` or `HIGHLY_RELEVANT` without requiring domain/entity/keyword contains support
|
|
|
|
### Requirement: Dense mode remains available for recall comparison
|
|
|
|
Configuration `retrieval.search.mode=dense` SHALL remain supported as a same-collection baseline that runs dense ANN only, without removing hybrid as the default production mode.
|
|
|
|
#### Scenario: Mode dense still callable
|
|
|
|
- **WHEN** mode is `dense`
|
|
- **THEN** search SHALL call dense ANN only and label hits `dense`
|
|
|
|
### Requirement: Hybrid quality gates MAY use optional dense distance without reordering
|
|
|
|
When hybrid search attaches an optional dense L2 distance for a hit, quality normalization and low-quality gates MAY use that distance for absolute similarity. Retrieval order SHALL remain the hybrid/RRF order and the primary scoreLabel SHALL remain hybrid.
|
|
|
|
#### Scenario: Dense distance does not replace hybrid label
|
|
|
|
- **WHEN** a hybrid hit includes denseDistance
|
|
- **THEN** scoreLabel SHALL still canonicalize to hybrid
|
|
- **AND** post-process sort order SHALL still follow originalRank from hybrid results
|