Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
6.0 KiB
Design: rag-quality-score-unify
Context
- Knowledge path already uses single
MilvusHybridKnowledgeStore(MilvusClientV2) with dense + BM25 + RRF. - Chunk-level
evidenceKeydedup andretrieve-k/return-nare landed. - Gap: hybrid ordering is RRF, but post-process still pretends scores are L2 and re-ranks with L0 keyword contains boosts.
- Prior design in
rag-bm25-hybrid-drop-sdkrequired dense L2 enrichment for threshold compatibility — this change supersedes that decision.
Stakeholders: lookup_knowledge internal pipeline; Agent ACI field names unchanged; operators comparing retrieval.search.mode=dense|hybrid.
Goals / Non-Goals
Goals:
- First-class
scoreLabelvalues: onlydense|hybrid(aliases canonicalize). - Single
toQualityScoreadapter; post-process is label-agnostic. - Preserve retrieval
originalRankas sort authority; remove keyword/domain boost re-ranking. - Hybrid quality = pure rank mapping over the current candidate batch (confirmed).
- Keep
mode=densefor offline recall comparison; production default remains hybrid.
Non-Goals:
- Cross-encoder / query rewrite / neighbor chunks.
- Schema rebuild or collection rename.
- Changing Agent-facing ACI JSON field names.
- Configurable
max(rank, denseSim)quality (future).
Decisions
D1 — Two labels only
| label | score meaning |
quality mapping |
|---|---|---|
dense |
L2 distance (smaller better) | 1 - clamp(l2)/maxL2Distance |
hybrid |
engine fused score optional in rawScore; not used as L2 |
rankToQuality(originalRank, batchSize) |
Canonicalize legacy strings: l2_distance→dense; rrf_fused / bm25_only_*→hybrid.
Why not keep bm25_only: it is not a search mode; it was a L2-fake patch. Hybrid path hits are all hybrid.
D2 — Stop dense L2 overwrite on hybrid hits
searchHybrid SHALL:
- Run
hybridSearch+ RRFRanker. - Emit hits in RRF order with
scoreLabel=hybrid,originalRank=1..n. - Set
rawScorefrom engine when present;scoreMAY equal raw fused score or rank placeholder — MUST NOT be replaced by dense L2 for post-process consumption. - MUST NOT set
bm25_only_no_denseor forcescore=maxL2Distancefor threshold faking. - MUST NOT run a parallel dense search solely to rewrite scores (slice-1 pure rank quality).
searchDense SHALL emit scoreLabel=dense and L2 in score.
D3 — RetrievalScoreNormalizer is the only label branch
qualityScore = RetrievalScoreNormalizer.toQualityScore(
scoreLabel, score, originalRank, batchSize, maxL2Distance)
- Hybrid: linear rank map — rank 1 → 1.0; rank n → ~1/n floor so last item > 0.
- Dense: existing L2 formula (behavior parity for dense mode).
Post-processor, isLowQuality, and relevance_level consume only qualityScore (exposed today as baseScore / topSimilarity fields for minimal DTO churn).
D4 — Post-process sort and boosts
order = originalRank ASC, then stable evidenceKey
// NO finalScore = base + 0.15 domain + ...
- Remove additive boosts from sort key and from
finalScoreused for ordering. - Optional: if L0 hint string matches, append explanatory
hitReasonsonly (e.g.l0_keyword_overlap) — zero score delta. - Keep: evidenceKey dedup, max-chunks-per-document, return-n, excerpt truncate.
relevance_level: compare topqualityScoreto existing thresholds; removehasHintSupportgate for PRECISE.isLowQuality/ category unfiltered retry: unchanged control flow, new score semantics.
D5 — Trace fields
RerankTracemay keepbaseScore/finalScorenames but both equalqualityScorewhen boosts are zero;boostReasonsempty or explanation-only reasons without:+0.xxscore deltas.- Prefer renaming comments to “quality trace”; no Agent contract change required.
D6 — WIP files
Workspace may contain draft RetrievalScoreLabels / RetrievalScoreNormalizer. Apply MUST align them to this design (or replace) and wire call sites; drafts alone are not done.
Interface impact
- Level: L2 — internal DTO semantics (
score,scoreLabel), post-process ordering and relevance distribution. - Agent ACI: field names stable;
relevance_levelvalues distribution may change (accepted behavioral change).
Data flow (target)
VectorSearchService (mode dense|hybrid)
-> hits{ rank, score, scoreLabel=dense|hybrid, rawScore? }
-> KnowledgeDocumentRetriever candidates
-> KnowledgeEvidencePostProcessor
for each: qualityScore = toQualityScore(...)
sort by originalRank
dedup / caps / return-n
relevance + topSimilarity from qualityScore
-> pack / assemble / project
Risks / Trade-offs
| Risk | Mitigation |
|---|---|
| Rank→quality not comparable across queries | Document; thresholds may need later tune; dense mode still L2-absolute |
| Hybrid top quality always high if batch small | batchSize = candidate list size after retrieve; rank1 always 1.0 by design for “best of this round” |
| PRECISE without hint support more often | Accepted; semantic/rank quality no longer gated on contains |
| Tests assert boost re-order | Update LookupKnowledgeToolTest.rerankUsesHintMatches... |
| Old label strings in sidecar/eval | canonicalize in normalizer |
Migration Plan
- Deploy code; no Milvus schema migration.
- Default
retrieval.search.mode=hybridunchanged. - Rollback: revert change; old L2-enrichment behavior returns.
- Optional ops: A/B dense vs hybrid recall using mode switch (unchanged capability).
Open Questions
- None for slice-1 (Q8 confirmed: pure rank).
- Follow-up: threshold calibration after live traces; optional denseSim blend.
Audit notes (inline)
Module chain: Tool → Retriever → Store → PostProcessor → Packer → Projector.
Ownership: retrieval scores owned by store+normalizer; evidence assembly by post-processor; Agent view by projector.
No new cross-module lifecycle. Couples only internal RAG pipeline.
Supersedes hybrid L2-enrichment ADR-equivalent decision from bm25-hybrid change.