feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
This commit is contained in:
@@ -0,0 +1,39 @@
|
||||
# Tasks: rag-quality-score-unify
|
||||
|
||||
## 1. Score contract utilities
|
||||
|
||||
- [x] 1.1 Finalize `RetrievalScoreLabels` (`dense` / `hybrid` + canonicalize legacy aliases)
|
||||
- [x] 1.2 Finalize `RetrievalScoreNormalizer.toQualityScore` (dense L2 formula; hybrid pure rank map with batchSize)
|
||||
- [x] 1.3 Unit tests for normalizer: dense L2 edges; hybrid rank monotonicity; alias canonicalize
|
||||
|
||||
## 2. Store / search emission
|
||||
|
||||
- [x] 2.1 `searchDense`: emit `scoreLabel=dense`, L2 `score`, stable rank order
|
||||
- [x] 2.2 `searchHybrid`: emit `scoreLabel=hybrid`; keep RRF order as `originalRank`; stop dense L2 overwrite and `bm25_only_*` labels; no parallel dense probe for score rewrite
|
||||
- [x] 2.3 Update `VectorSearchService.SearchResult` / adapter comments so `score`+`scoreLabel` contract matches design
|
||||
- [x] 2.4 Ensure `KnowledgeDocumentRetriever` / `VectorKnowledgeSearchAdapter` propagate `scoreLabel`, `score`, `rawScore`, `originalRank` unchanged
|
||||
|
||||
## 3. Post-process
|
||||
|
||||
- [x] 3.1 `KnowledgeEvidencePostProcessor`: compute quality via normalizer; sort by `originalRank` ASC (stable tie-break)
|
||||
- [x] 3.2 Remove domain/entity/keyword/source_type additive boosts from ordering/`finalScore`
|
||||
- [x] 3.3 Optional: L0 overlap only as explanatory `hitReasons` (no score delta)
|
||||
- [x] 3.4 `relevance_level` / `isLowQuality` / `topSimilarity` use qualityScore only; drop `hasHintSupport` gate for PRECISE
|
||||
- [x] 3.5 Keep evidenceKey dedup, max-chunks-per-document, return-n, excerpt truncate
|
||||
|
||||
## 4. Tests
|
||||
|
||||
- [x] 4.1 Update `KnowledgeEvidencePostProcessorTest` for rank order + caps under new scoring
|
||||
- [x] 4.2 Update `LookupKnowledgeToolTest.rerankUsesHintMatchesAndContextPackPreservesMetadata` (no boost re-order; metadata/context pack still ok)
|
||||
- [x] 4.3 Adjust any tests asserting `l2_distance` / boost reasons `:+0.xx` as needed
|
||||
- [x] 4.4 Run targeted unit tests for touched classes
|
||||
|
||||
## 5. Docs
|
||||
|
||||
- [x] 5.1 Update `mvp/architecture/RAG知识检索架构.md` §6 score/post-process (replace L2-enrichment narrative)
|
||||
- [x] 5.2 Align `application.yml` comments if still describing L2-only post-process for hybrid
|
||||
|
||||
## 6. Verify
|
||||
|
||||
- [x] 6.1 Confirm no production path still sets `bm25_only_no_dense` or overwrites hybrid scores with L2 for thresholds
|
||||
- [x] 6.2 Note known limitation: hybrid quality is ordinal within batch; thresholds may need later calibration
|
||||
Reference in New Issue
Block a user