feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
This commit is contained in:
@@ -0,0 +1,134 @@
|
||||
# Design: rag-quality-score-unify
|
||||
|
||||
## Context
|
||||
|
||||
- Knowledge path already uses single `MilvusHybridKnowledgeStore` (MilvusClientV2) with dense + BM25 + RRF.
|
||||
- Chunk-level `evidenceKey` dedup and `retrieve-k` / `return-n` are landed.
|
||||
- Gap: hybrid ordering is RRF, but post-process still pretends scores are L2 and re-ranks with L0 keyword contains boosts.
|
||||
- Prior design in `rag-bm25-hybrid-drop-sdk` required dense L2 enrichment for threshold compatibility — **this change supersedes that decision**.
|
||||
|
||||
Stakeholders: `lookup_knowledge` internal pipeline; Agent ACI field *names* unchanged; operators comparing `retrieval.search.mode=dense|hybrid`.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
1. First-class `scoreLabel` values: only `dense` | `hybrid` (aliases canonicalize).
|
||||
2. Single `toQualityScore` adapter; post-process is label-agnostic.
|
||||
3. Preserve retrieval `originalRank` as sort authority; remove keyword/domain boost re-ranking.
|
||||
4. Hybrid quality = pure rank mapping over the current candidate batch (confirmed).
|
||||
5. Keep `mode=dense` for offline recall comparison; production default remains hybrid.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Cross-encoder / query rewrite / neighbor chunks.
|
||||
- Schema rebuild or collection rename.
|
||||
- Changing Agent-facing ACI JSON field names.
|
||||
- Configurable `max(rank, denseSim)` quality (future).
|
||||
|
||||
## Decisions
|
||||
|
||||
### D1 — Two labels only
|
||||
|
||||
| label | `score` meaning | quality mapping |
|
||||
|---|---|---|
|
||||
| `dense` | L2 distance (smaller better) | `1 - clamp(l2)/maxL2Distance` |
|
||||
| `hybrid` | engine fused score optional in `rawScore`; **not** used as L2 | `rankToQuality(originalRank, batchSize)` |
|
||||
|
||||
Canonicalize legacy strings: `l2_distance`→dense; `rrf_fused` / `bm25_only_*`→hybrid.
|
||||
|
||||
**Why not keep `bm25_only`:** it is not a search mode; it was a L2-fake patch. Hybrid path hits are all `hybrid`.
|
||||
|
||||
### D2 — Stop dense L2 overwrite on hybrid hits
|
||||
|
||||
`searchHybrid` SHALL:
|
||||
|
||||
1. Run `hybridSearch` + RRFRanker.
|
||||
2. Emit hits in RRF order with `scoreLabel=hybrid`, `originalRank=1..n`.
|
||||
3. Set `rawScore` from engine when present; `score` MAY equal raw fused score or rank placeholder — MUST NOT be replaced by dense L2 for post-process consumption.
|
||||
4. MUST NOT set `bm25_only_no_dense` or force `score=maxL2Distance` for threshold faking.
|
||||
5. MUST NOT run a parallel dense search solely to rewrite scores (slice-1 pure rank quality).
|
||||
|
||||
`searchDense` SHALL emit `scoreLabel=dense` and L2 in `score`.
|
||||
|
||||
### D3 — `RetrievalScoreNormalizer` is the only label branch
|
||||
|
||||
```text
|
||||
qualityScore = RetrievalScoreNormalizer.toQualityScore(
|
||||
scoreLabel, score, originalRank, batchSize, maxL2Distance)
|
||||
```
|
||||
|
||||
- Hybrid: linear rank map — rank 1 → 1.0; rank n → ~1/n floor so last item > 0.
|
||||
- Dense: existing L2 formula (behavior parity for dense mode).
|
||||
|
||||
Post-processor, `isLowQuality`, and `relevance_level` consume **only** `qualityScore` (exposed today as `baseScore` / `topSimilarity` fields for minimal DTO churn).
|
||||
|
||||
### D4 — Post-process sort and boosts
|
||||
|
||||
```text
|
||||
order = originalRank ASC, then stable evidenceKey
|
||||
// NO finalScore = base + 0.15 domain + ...
|
||||
```
|
||||
|
||||
- Remove additive boosts from sort key and from `finalScore` used for ordering.
|
||||
- Optional: if L0 hint string matches, append explanatory `hitReasons` only (e.g. `l0_keyword_overlap`) — zero score delta.
|
||||
- Keep: evidenceKey dedup, max-chunks-per-document, return-n, excerpt truncate.
|
||||
- `relevance_level`: compare top `qualityScore` to existing thresholds; **remove** `hasHintSupport` gate for PRECISE.
|
||||
- `isLowQuality` / category unfiltered retry: unchanged control flow, new score semantics.
|
||||
|
||||
### D5 — Trace fields
|
||||
|
||||
- `RerankTrace` may keep `baseScore`/`finalScore` names but both equal `qualityScore` when boosts are zero; `boostReasons` empty or explanation-only reasons without `:+0.xx` score deltas.
|
||||
- Prefer renaming comments to “quality trace”; no Agent contract change required.
|
||||
|
||||
### D6 — WIP files
|
||||
|
||||
Workspace may contain draft `RetrievalScoreLabels` / `RetrievalScoreNormalizer`. Apply MUST align them to this design (or replace) and wire call sites; drafts alone are not done.
|
||||
|
||||
## Interface impact
|
||||
|
||||
- **Level: L2** — internal DTO semantics (`score`, `scoreLabel`), post-process ordering and relevance distribution.
|
||||
- Agent ACI: field names stable; `relevance_level` *values distribution* may change (accepted behavioral change).
|
||||
|
||||
## Data flow (target)
|
||||
|
||||
```text
|
||||
VectorSearchService (mode dense|hybrid)
|
||||
-> hits{ rank, score, scoreLabel=dense|hybrid, rawScore? }
|
||||
-> KnowledgeDocumentRetriever candidates
|
||||
-> KnowledgeEvidencePostProcessor
|
||||
for each: qualityScore = toQualityScore(...)
|
||||
sort by originalRank
|
||||
dedup / caps / return-n
|
||||
relevance + topSimilarity from qualityScore
|
||||
-> pack / assemble / project
|
||||
```
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
| Risk | Mitigation |
|
||||
|---|---|
|
||||
| Rank→quality not comparable across queries | Document; thresholds may need later tune; dense mode still L2-absolute |
|
||||
| Hybrid top quality always high if batch small | batchSize = candidate list size after retrieve; rank1 always 1.0 by design for “best of this round” |
|
||||
| PRECISE without hint support more often | Accepted; semantic/rank quality no longer gated on contains |
|
||||
| Tests assert boost re-order | Update `LookupKnowledgeToolTest.rerankUsesHintMatches...` |
|
||||
| Old label strings in sidecar/eval | canonicalize in normalizer |
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Deploy code; no Milvus schema migration.
|
||||
2. Default `retrieval.search.mode=hybrid` unchanged.
|
||||
3. Rollback: revert change; old L2-enrichment behavior returns.
|
||||
4. Optional ops: A/B dense vs hybrid recall using mode switch (unchanged capability).
|
||||
|
||||
## Open Questions
|
||||
|
||||
- None for slice-1 (Q8 confirmed: pure rank).
|
||||
- Follow-up: threshold calibration after live traces; optional denseSim blend.
|
||||
|
||||
## Audit notes (inline)
|
||||
|
||||
Module chain: Tool → Retriever → Store → PostProcessor → Packer → Projector.
|
||||
Ownership: retrieval scores owned by store+normalizer; evidence assembly by post-processor; Agent view by projector.
|
||||
No new cross-module lifecycle. Couples only internal RAG pipeline.
|
||||
Supersedes hybrid L2-enrichment ADR-equivalent decision from bm25-hybrid change.
|
||||
Reference in New Issue
Block a user