Files
SuperBizAgent-java/openspec/changes/archive/2026-07-28-rag-quality-score-unify/design.md
T
zhuyongxin 7ae9707a3b feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit
(DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG
tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools,
and align MVP docs after live E2E verification.
2026-07-28 19:43:13 +08:00

6.0 KiB

Design: rag-quality-score-unify

Context

  • Knowledge path already uses single MilvusHybridKnowledgeStore (MilvusClientV2) with dense + BM25 + RRF.
  • Chunk-level evidenceKey dedup and retrieve-k / return-n are landed.
  • Gap: hybrid ordering is RRF, but post-process still pretends scores are L2 and re-ranks with L0 keyword contains boosts.
  • Prior design in rag-bm25-hybrid-drop-sdk required dense L2 enrichment for threshold compatibility — this change supersedes that decision.

Stakeholders: lookup_knowledge internal pipeline; Agent ACI field names unchanged; operators comparing retrieval.search.mode=dense|hybrid.

Goals / Non-Goals

Goals:

  1. First-class scoreLabel values: only dense | hybrid (aliases canonicalize).
  2. Single toQualityScore adapter; post-process is label-agnostic.
  3. Preserve retrieval originalRank as sort authority; remove keyword/domain boost re-ranking.
  4. Hybrid quality = pure rank mapping over the current candidate batch (confirmed).
  5. Keep mode=dense for offline recall comparison; production default remains hybrid.

Non-Goals:

  • Cross-encoder / query rewrite / neighbor chunks.
  • Schema rebuild or collection rename.
  • Changing Agent-facing ACI JSON field names.
  • Configurable max(rank, denseSim) quality (future).

Decisions

D1 — Two labels only

label score meaning quality mapping
dense L2 distance (smaller better) 1 - clamp(l2)/maxL2Distance
hybrid engine fused score optional in rawScore; not used as L2 rankToQuality(originalRank, batchSize)

Canonicalize legacy strings: l2_distance→dense; rrf_fused / bm25_only_*→hybrid.

Why not keep bm25_only: it is not a search mode; it was a L2-fake patch. Hybrid path hits are all hybrid.

D2 — Stop dense L2 overwrite on hybrid hits

searchHybrid SHALL:

  1. Run hybridSearch + RRFRanker.
  2. Emit hits in RRF order with scoreLabel=hybrid, originalRank=1..n.
  3. Set rawScore from engine when present; score MAY equal raw fused score or rank placeholder — MUST NOT be replaced by dense L2 for post-process consumption.
  4. MUST NOT set bm25_only_no_dense or force score=maxL2Distance for threshold faking.
  5. MUST NOT run a parallel dense search solely to rewrite scores (slice-1 pure rank quality).

searchDense SHALL emit scoreLabel=dense and L2 in score.

D3 — RetrievalScoreNormalizer is the only label branch

qualityScore = RetrievalScoreNormalizer.toQualityScore(
  scoreLabel, score, originalRank, batchSize, maxL2Distance)
  • Hybrid: linear rank map — rank 1 → 1.0; rank n → ~1/n floor so last item > 0.
  • Dense: existing L2 formula (behavior parity for dense mode).

Post-processor, isLowQuality, and relevance_level consume only qualityScore (exposed today as baseScore / topSimilarity fields for minimal DTO churn).

D4 — Post-process sort and boosts

order = originalRank ASC, then stable evidenceKey
// NO finalScore = base + 0.15 domain + ...
  • Remove additive boosts from sort key and from finalScore used for ordering.
  • Optional: if L0 hint string matches, append explanatory hitReasons only (e.g. l0_keyword_overlap) — zero score delta.
  • Keep: evidenceKey dedup, max-chunks-per-document, return-n, excerpt truncate.
  • relevance_level: compare top qualityScore to existing thresholds; remove hasHintSupport gate for PRECISE.
  • isLowQuality / category unfiltered retry: unchanged control flow, new score semantics.

D5 — Trace fields

  • RerankTrace may keep baseScore/finalScore names but both equal qualityScore when boosts are zero; boostReasons empty or explanation-only reasons without :+0.xx score deltas.
  • Prefer renaming comments to “quality trace”; no Agent contract change required.

D6 — WIP files

Workspace may contain draft RetrievalScoreLabels / RetrievalScoreNormalizer. Apply MUST align them to this design (or replace) and wire call sites; drafts alone are not done.

Interface impact

  • Level: L2 — internal DTO semantics (score, scoreLabel), post-process ordering and relevance distribution.
  • Agent ACI: field names stable; relevance_level values distribution may change (accepted behavioral change).

Data flow (target)

VectorSearchService (mode dense|hybrid)
  -> hits{ rank, score, scoreLabel=dense|hybrid, rawScore? }
  -> KnowledgeDocumentRetriever candidates
  -> KnowledgeEvidencePostProcessor
       for each: qualityScore = toQualityScore(...)
       sort by originalRank
       dedup / caps / return-n
       relevance + topSimilarity from qualityScore
  -> pack / assemble / project

Risks / Trade-offs

Risk Mitigation
Rank→quality not comparable across queries Document; thresholds may need later tune; dense mode still L2-absolute
Hybrid top quality always high if batch small batchSize = candidate list size after retrieve; rank1 always 1.0 by design for “best of this round”
PRECISE without hint support more often Accepted; semantic/rank quality no longer gated on contains
Tests assert boost re-order Update LookupKnowledgeToolTest.rerankUsesHintMatches...
Old label strings in sidecar/eval canonicalize in normalizer

Migration Plan

  1. Deploy code; no Milvus schema migration.
  2. Default retrieval.search.mode=hybrid unchanged.
  3. Rollback: revert change; old L2-enrichment behavior returns.
  4. Optional ops: A/B dense vs hybrid recall using mode switch (unchanged capability).

Open Questions

  • None for slice-1 (Q8 confirmed: pure rank).
  • Follow-up: threshold calibration after live traces; optional denseSim blend.

Audit notes (inline)

Module chain: Tool → Retriever → Store → PostProcessor → Packer → Projector.
Ownership: retrieval scores owned by store+normalizer; evidence assembly by post-processor; Agent view by projector.
No new cross-module lifecycle. Couples only internal RAG pipeline.
Supersedes hybrid L2-enrichment ADR-equivalent decision from bm25-hybrid change.