Files
SuperBizAgent-java/openspec/specs/rag-retrieval-quality-score/spec.md
T
zhuyongxin 7ae9707a3b feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit
(DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG
tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools,
and align MVP docs after live E2E verification.
2026-07-28 19:43:13 +08:00

5.4 KiB

rag-retrieval-quality-score Specification

Purpose

Unify knowledge retrieval score labels and quality normalization so dense and hybrid modes share one post-process pipeline without L2 faking or keyword boost re-ranking.

Requirements

Requirement: Primary score labels SHALL be only dense or hybrid

Knowledge search hits used by lookup_knowledge SHALL set scoreLabel to dense or hybrid (after any legacy alias canonicalization). The system SHALL NOT treat bm25_only_no_dense, rrf_fused, or l2_distance as distinct first-class labels in new emissions.

Scenario: Dense mode labels hits as dense

  • WHEN retrieval.search.mode is dense and search returns hits
  • THEN each hit SHALL have scoreLabel canonicalizing to dense
  • AND score SHALL be the dense L2 distance

Scenario: Hybrid mode labels hits as hybrid

  • WHEN retrieval.search.mode is hybrid and search returns hits
  • THEN each hit SHALL have scoreLabel canonicalizing to hybrid
  • AND hit order SHALL follow the hybrid/RRF result order via originalRank

Scenario: Legacy aliases canonicalize

  • WHEN a candidate carries a legacy label such as l2_distance or rrf_fused or bm25_only_no_dense
  • THEN quality normalization SHALL canonicalize it to dense or hybrid respectively before computing quality

Requirement: Hybrid search SHALL NOT overwrite scores with dense L2 for post-process

When hybrid search runs, the store SHALL NOT replace hybrid hit scores with a parallel dense L2 map for the purpose of post-process thresholds, and SHALL NOT assign max-L2 weak placeholders under a bm25_only_* primary label.

Scenario: No L2 enrichment overwrite

  • WHEN hybrid search completes
  • THEN post-process input scores SHALL NOT be forced to dense L2 solely for threshold compatibility
  • AND the system SHALL NOT emit bm25_only_no_dense as the primary score label on new hits

Requirement: Quality score SHALL be produced by a single normalizer

The pipeline SHALL compute a qualityScore in [0, 1] (higher is better) using one normalizer API that branches only on canonical label.

Scenario: Dense quality from L2

  • WHEN label is dense and score is L2 distance d with configured maxL2Distance
  • THEN qualityScore SHALL equal max(0, 1 - min(d, maxL2Distance) / maxL2Distance) (null score → 0)

Scenario: Hybrid quality from rank

  • WHEN label is hybrid and originalRank is r within a candidate batch of size n (n >= 1)
  • THEN qualityScore SHALL be a monotonically non-increasing function of r over that batch
  • AND rank 1 SHALL map to 1.0 when n >= 1
  • AND the mapping SHALL NOT require dense L2 or RRF raw magnitude

Requirement: Post-process SHALL preserve retrieval rank order

Evidence post-processing SHALL order candidates by originalRank ascending (stable tie-break allowed). It SHALL NOT re-order primarily by L0 domain/entity/keyword/source_type additive boosts.

Scenario: Keyword overlap does not promote lower rank

  • WHEN candidate A has originalRank=1 and candidate B has originalRank=2
  • AND B matches more L0 keywords via string contains than A
  • THEN after post-process acceptance order, A SHALL appear before B among accepted blocks (subject only to dedup/caps removing one of them)

Scenario: Structural caps still apply after rank order

  • WHEN more than rag.max-chunks-per-document chunks share a docId
  • THEN only the best-ranked (lowest originalRank) up to the cap SHALL remain
  • AND rag.return-n SHALL still bound total blocks

Requirement: Relevance and low-quality gates SHALL use qualityScore only

relevance_level, completeness hints, topSimilarity (or equivalent top quality field), and category-filter low-quality retry decisions SHALL use the normalized qualityScore, not raw L2-under-hybrid fakes and not keyword-boosted final scores.

Scenario: Low quality uses top qualityScore

  • WHEN post-process finishes with at least one evidence block
  • THEN low-quality detection SHALL compare the top qualityScore to the configured reference threshold
  • AND SHALL NOT require L0 hint contains-match to treat the result as usable when quality meets threshold

Scenario: PRECISE does not require hint support

  • WHEN top qualityScore is at or above the highly-relevant threshold
  • THEN the system MAY assign PRECISE or HIGHLY_RELEVANT without requiring domain/entity/keyword contains support

Requirement: Dense mode remains available for recall comparison

Configuration retrieval.search.mode=dense SHALL remain supported as a same-collection baseline that runs dense ANN only, without removing hybrid as the default production mode.

Scenario: Mode dense still callable

  • WHEN mode is dense
  • THEN search SHALL call dense ANN only and label hits dense

Requirement: Hybrid quality gates MAY use optional dense distance without reordering

When hybrid search attaches an optional dense L2 distance for a hit, quality normalization and low-quality gates MAY use that distance for absolute similarity. Retrieval order SHALL remain the hybrid/RRF order and the primary scoreLabel SHALL remain hybrid.

Scenario: Dense distance does not replace hybrid label

  • WHEN a hybrid hit includes denseDistance
  • THEN scoreLabel SHALL still canonicalize to hybrid
  • AND post-process sort order SHALL still follow originalRank from hybrid results