Persist provider reasoning and assistant text separately on agent_reasoning_audit (DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools, and align MVP docs after live E2E verification.
867 B
867 B
Decisions: rag-eval-hybrid-baseline(最终版)
Process
sm-flow standard-lean: Discover → Commit → Apply → Archive.
Key decisions
- Replace eval generator
vector-store.modewithretrieval.search.mode(default hybrid). - Fixture meta:
searchMode,kbScopewhen set. - Knife-2 (dual fixtures / mustNot golden) deferred.
- Live refresh succeeded in apply env; baseline updated to hybrid snapshots.
- Quality gate refinement (apply-found): hybrid absolute quality for
isLowQuality/ relevance uses optional dense L2 (denseDistance); does not overwrite hybrid scoreLabel or RRF order. Rank mapping remains fallback when dense missing.
Trade-offs
- Extra dense ANN on hybrid path for gate calibration (latency) vs correct filter-fallback behavior.
- relevance_level still not a hard golden assertion (ordinal vs absolute mix).