Files
SuperBizAgent-java/openspec/changes/archive/2026-07-28-rag-quality-score-unify/tasks.md
T
zhuyongxin 7ae9707a3b feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit
(DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG
tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools,
and align MVP docs after live E2E verification.
2026-07-28 19:43:13 +08:00

2.3 KiB

Tasks: rag-quality-score-unify

1. Score contract utilities

  • 1.1 Finalize RetrievalScoreLabels (dense / hybrid + canonicalize legacy aliases)
  • 1.2 Finalize RetrievalScoreNormalizer.toQualityScore (dense L2 formula; hybrid pure rank map with batchSize)
  • 1.3 Unit tests for normalizer: dense L2 edges; hybrid rank monotonicity; alias canonicalize

2. Store / search emission

  • 2.1 searchDense: emit scoreLabel=dense, L2 score, stable rank order
  • 2.2 searchHybrid: emit scoreLabel=hybrid; keep RRF order as originalRank; stop dense L2 overwrite and bm25_only_* labels; no parallel dense probe for score rewrite
  • 2.3 Update VectorSearchService.SearchResult / adapter comments so score+scoreLabel contract matches design
  • 2.4 Ensure KnowledgeDocumentRetriever / VectorKnowledgeSearchAdapter propagate scoreLabel, score, rawScore, originalRank unchanged

3. Post-process

  • 3.1 KnowledgeEvidencePostProcessor: compute quality via normalizer; sort by originalRank ASC (stable tie-break)
  • 3.2 Remove domain/entity/keyword/source_type additive boosts from ordering/finalScore
  • 3.3 Optional: L0 overlap only as explanatory hitReasons (no score delta)
  • 3.4 relevance_level / isLowQuality / topSimilarity use qualityScore only; drop hasHintSupport gate for PRECISE
  • 3.5 Keep evidenceKey dedup, max-chunks-per-document, return-n, excerpt truncate

4. Tests

  • 4.1 Update KnowledgeEvidencePostProcessorTest for rank order + caps under new scoring
  • 4.2 Update LookupKnowledgeToolTest.rerankUsesHintMatchesAndContextPackPreservesMetadata (no boost re-order; metadata/context pack still ok)
  • 4.3 Adjust any tests asserting l2_distance / boost reasons :+0.xx as needed
  • 4.4 Run targeted unit tests for touched classes

5. Docs

  • 5.1 Update mvp/architecture/RAG知识检索架构.md §6 score/post-process (replace L2-enrichment narrative)
  • 5.2 Align application.yml comments if still describing L2-only post-process for hybrid

6. Verify

  • 6.1 Confirm no production path still sets bm25_only_no_dense or overwrites hybrid scores with L2 for thresholds
  • 6.2 Note known limitation: hybrid quality is ordinal within batch; thresholds may need later calibration