feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality

Persist provider reasoning and assistant text separately on agent_reasoning_audit
(DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG
tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools,
and align MVP docs after live E2E verification.
This commit is contained in:
zhuyongxin
2026-07-28 19:43:13 +08:00
parent 2f40536248
commit 7ae9707a3b
116 changed files with 8364 additions and 1141 deletions
@@ -0,0 +1,50 @@
# rag-eval-offline-baseline Specification
## Purpose
Keep the offline RAG retrieval baseline aligned with the hybrid search main path and fixture metadata needed for reproducible regression.
## ADDED Requirements
### Requirement: Snapshot generation SHALL use retrieval search mode
RAG lookup fixture generation SHALL configure `retrieval.search.mode` and SHALL NOT require `retrieval.vector-store.mode` for knowledge snapshot generation.
#### Scenario: Default hybrid generation
- **WHEN** the snapshot generator is invoked with default parameters
- **THEN** it SHALL run with `retrieval.search.mode=hybrid` (or equivalent default)
- **AND** it SHALL NOT pass `retrieval.vector-store.mode` as a required generation setting
#### Scenario: Dense mode override for comparison runs
- **WHEN** the operator sets search mode to `dense`
- **THEN** fixture generation SHALL use dense retrieval for that run
### Requirement: Generated fixtures SHALL record search meta
Each generated fixture file SHALL include stable identity and generation meta in addition to the lookup payload.
#### Scenario: Meta fields present
- **WHEN** a fixture is written for a golden case
- **THEN** the fixture SHALL contain `caseId`, `query`, `retrievedAt`, and `searchMode`
- **AND** when kb scope is configured non-empty, the fixture SHOULD contain `kbScope`
### Requirement: Offline evaluation SHALL remain dependency-free
The offline baseline checker SHALL evaluate golden cases against fixture files without calling Milvus, embedding APIs, or starting the full application.
#### Scenario: Offline eval without live stack
- **WHEN** `eval_rag_retrieval.py` (or successor) runs against cases and fixtures
- **THEN** it SHALL produce pass/fail results using fixture contents only
### Requirement: Eval documentation SHALL describe the hybrid-era loop
Project eval README SHALL document seed import, snapshot generation with `search.mode`, offline eval, and baseline diff, without presenting Spring vector-store mode as the knowledge main path.
#### Scenario: README main path
- **WHEN** an engineer follows the eval README happy path
- **THEN** the documented default generation mode SHALL be hybrid search mode
@@ -0,0 +1,104 @@
# rag-retrieval-quality-score Specification
## Purpose
Unify knowledge retrieval score labels and quality normalization so dense and hybrid modes share one post-process pipeline without L2 faking or keyword boost re-ranking.
## Requirements
### Requirement: Primary score labels SHALL be only dense or hybrid
Knowledge search hits used by `lookup_knowledge` SHALL set `scoreLabel` to `dense` or `hybrid` (after any legacy alias canonicalization). The system SHALL NOT treat `bm25_only_no_dense`, `rrf_fused`, or `l2_distance` as distinct first-class labels in new emissions.
#### Scenario: Dense mode labels hits as dense
- **WHEN** `retrieval.search.mode` is `dense` and search returns hits
- **THEN** each hit SHALL have `scoreLabel` canonicalizing to `dense`
- **AND** `score` SHALL be the dense L2 distance
#### Scenario: Hybrid mode labels hits as hybrid
- **WHEN** `retrieval.search.mode` is `hybrid` and search returns hits
- **THEN** each hit SHALL have `scoreLabel` canonicalizing to `hybrid`
- **AND** hit order SHALL follow the hybrid/RRF result order via `originalRank`
#### Scenario: Legacy aliases canonicalize
- **WHEN** a candidate carries a legacy label such as `l2_distance` or `rrf_fused` or `bm25_only_no_dense`
- **THEN** quality normalization SHALL canonicalize it to `dense` or `hybrid` respectively before computing quality
### Requirement: Hybrid search SHALL NOT overwrite scores with dense L2 for post-process
When hybrid search runs, the store SHALL NOT replace hybrid hit scores with a parallel dense L2 map for the purpose of post-process thresholds, and SHALL NOT assign max-L2 weak placeholders under a `bm25_only_*` primary label.
#### Scenario: No L2 enrichment overwrite
- **WHEN** hybrid search completes
- **THEN** post-process input scores SHALL NOT be forced to dense L2 solely for threshold compatibility
- **AND** the system SHALL NOT emit `bm25_only_no_dense` as the primary score label on new hits
### Requirement: Quality score SHALL be produced by a single normalizer
The pipeline SHALL compute a `qualityScore` in `[0, 1]` (higher is better) using one normalizer API that branches only on canonical label.
#### Scenario: Dense quality from L2
- **WHEN** label is `dense` and score is L2 distance `d` with configured `maxL2Distance`
- **THEN** `qualityScore` SHALL equal `max(0, 1 - min(d, maxL2Distance) / maxL2Distance)` (null score → 0)
#### Scenario: Hybrid quality from rank
- **WHEN** label is `hybrid` and `originalRank` is `r` within a candidate batch of size `n` (`n >= 1`)
- **THEN** `qualityScore` SHALL be a monotonically non-increasing function of `r` over that batch
- **AND** rank `1` SHALL map to `1.0` when `n >= 1`
- **AND** the mapping SHALL NOT require dense L2 or RRF raw magnitude
### Requirement: Post-process SHALL preserve retrieval rank order
Evidence post-processing SHALL order candidates by `originalRank` ascending (stable tie-break allowed). It SHALL NOT re-order primarily by L0 domain/entity/keyword/source_type additive boosts.
#### Scenario: Keyword overlap does not promote lower rank
- **WHEN** candidate A has `originalRank=1` and candidate B has `originalRank=2`
- **AND** B matches more L0 keywords via string contains than A
- **THEN** after post-process acceptance order, A SHALL appear before B among accepted blocks (subject only to dedup/caps removing one of them)
#### Scenario: Structural caps still apply after rank order
- **WHEN** more than `rag.max-chunks-per-document` chunks share a docId
- **THEN** only the best-ranked (lowest `originalRank`) up to the cap SHALL remain
- **AND** `rag.return-n` SHALL still bound total blocks
### Requirement: Relevance and low-quality gates SHALL use qualityScore only
`relevance_level`, completeness hints, `topSimilarity` (or equivalent top quality field), and category-filter low-quality retry decisions SHALL use the normalized `qualityScore`, not raw L2-under-hybrid fakes and not keyword-boosted final scores.
#### Scenario: Low quality uses top qualityScore
- **WHEN** post-process finishes with at least one evidence block
- **THEN** low-quality detection SHALL compare the top `qualityScore` to the configured reference threshold
- **AND** SHALL NOT require L0 hint contains-match to treat the result as usable when quality meets threshold
#### Scenario: PRECISE does not require hint support
- **WHEN** top `qualityScore` is at or above the highly-relevant threshold
- **THEN** the system MAY assign `PRECISE` or `HIGHLY_RELEVANT` without requiring domain/entity/keyword contains support
### Requirement: Dense mode remains available for recall comparison
Configuration `retrieval.search.mode=dense` SHALL remain supported as a same-collection baseline that runs dense ANN only, without removing hybrid as the default production mode.
#### Scenario: Mode dense still callable
- **WHEN** mode is `dense`
- **THEN** search SHALL call dense ANN only and label hits `dense`
### Requirement: Hybrid quality gates MAY use optional dense distance without reordering
When hybrid search attaches an optional dense L2 distance for a hit, quality normalization and low-quality gates MAY use that distance for absolute similarity. Retrieval order SHALL remain the hybrid/RRF order and the primary scoreLabel SHALL remain hybrid.
#### Scenario: Dense distance does not replace hybrid label
- **WHEN** a hybrid hit includes denseDistance
- **THEN** scoreLabel SHALL still canonicalize to hybrid
- **AND** post-process sort order SHALL still follow originalRank from hybrid results