Files
SuperBizAgent-java/mvp/architecture/RAG检索可观测性与审计.md
zhuyongxin 584639fa2a docs(mvp): move engineering notes under mvp/engineering
Relocate RAG and diagnosis decision/E2E writeups from docs/ root into
mvp/engineering so architecture, issues, and engineering narrative stay
together. Update indexes and cross-links; leave docs/learning as legacy.
2026-07-29 10:49:45 +08:00

13 KiB
Raw Permalink Blame History

RAG 检索可观测性、审计与 Trace(现行)

更新日期:2026-07-28
状态:当前可运行
关联:lookup_knowledge、Harness ToolBoundary、tool_invocation、DiagnosisTraceService、离线 eval


1. 三层边界

flowchart TB
    subgraph A["A. 请求内 Trace"]
        LR[LookupResult<br/>retrievalTrace / rerankTrace / relevanceLevel]
    end

    subgraph B["B. 持久化审计 + Trace API"]
        TI[tool_invocation 表]
        DT[diagnosis_trace 事件摘要]
        API["GET /api/diagnosis/{sessionId}/trace"]
    end

    subgraph C["C. 质量回归"]
        EV[eval/rag-retrieval offline baseline]
    end

    subgraph agent["Agent 可见(非审计)"]
        RT[RagToolResult<br/>evidence + optional relevance_level]
    end

    LK[LookupKnowledgeTool] --> LR
    LR --> PROJ[RagResultProjector]
    PROJ --> RT
    LR --> BOUND[ToolBoundary audit]
    BOUND --> TI
    BOUND --> DT
    TI --> API
    DT --> API
    EV -.->|不替代运行时 Trace| LK
层 完善度 说明
A 请求内 高 attempt / fallback / quality 齐全
B 持久化 + Trace API 中高 RAG 富字段入 tool_invocation,经 Trace API 回放
C 离线 eval 高 hybrid fixtures 回归

Agent 看到的不是完整 Trace。 完整检索轨迹在 A/B;Agent 只拿投影后的证据契约。


2. 端到端:从 lookup 到 Trace API

sequenceDiagram
    participant Agent
    participant Adapter as RagToolAdapter
    participant Bound as ToolBoundary
    participant Tool as LookupKnowledgeTool
    participant Sink as JpaToolInvocationAuditSink
    participant DB as tool_invocation
    participant Trace as DiagnosisTraceService
    participant API as GET .../trace

    Agent->>Adapter: lookup_knowledge(query)
    Adapter->>Bound: execute(legacy, projector)
    Bound->>Tool: execute(query)
    Tool-->>Bound: raw LookupResult JSON
    Note over Tool: 内含 retrievalTrace / rerankTrace / evidenceBlocks
    Bound->>Bound: project → RagToolResult
    Bound->>Sink: AuditEvent + rawResultJson + agentResultJson
    Sink->>Sink: RagLookupAuditEnricher
    Sink->>DB: 富字段行
    Bound-->>Agent: 投影后 agent_result(无完整 trace)

    API->>Trace: sessionId + optional runId
    Trace->>DB: find tool_invocation by run/session
    Trace-->>API: DiagnosisTraceResponse.toolInvocations[]

3. A 层:请求内 Trace(LookupResult)

一次成功的 lookup_knowledge 内部出口是 LookupResult(比 Agent 契约更富)。

3.1 结构总览

flowchart TB
    LR[LookupResult]
    LR --> F[found]
    LR --> EB[evidenceBlocks[]]
    LR --> CP[contextPack]
    LR --> RT[retrievalTrace]
    LR --> RR[rerankTrace]
    LR --> RL[relevanceLevel]
    LR --> CH[completenessHint]
    LR --> CNT[evidenceCandidateCount / evidenceBlockCount]

    RT --> ATT[attempts[]]
    RT --> SEL[selectedAttempt]
    RT --> FB[fallbackReason]
    RT --> HINT[queryHints L0]
字段 含义
found 是否有可用证据块
evidenceBlocks 后处理后的 chunk 级证据(含 evidenceKey、source、content…)
contextPack 字符预算打包文本(内部/审计用)
retrievalTrace 检索路径 Trace(见下)
rerankTrace 后处理排序/quality 痕迹(现多为保序后的 quality)
relevanceLevel PRECISE / REFERENCE / null
completenessHint 给模型的天花板提示文案

3.2 retrievalTrace(检索路径)

字段 含义
originalQuery 原始查询
rewrittenQuery L0/变换后用于检索的 query
categoryFilter 首次过滤的 category(可 null)
selectedAttempt 最终采用的 attempt 名
fallbackReason 如 filtered_vector_low_quality;未降级为 null
evidenceStatus 内部:supported / no_evidence 等
queryHints L0:domains、keywords、entities、l0_match_count…
attempts[] 每次检索尝试快照

常见 selectedAttempt:

值 含义
FILTERED_VECTOR 带 category 的首次检索即采用
UNFILTERED_VECTOR 无 category,直接全库检索
UNFILTERED_VECTOR_RETRY filtered 低质/无证据后去掉 category 重试

单次 attempts[] 元素:

字段 含义
name attempt 名
query 该次实际检索句
categoryFilter 该次 filter
candidateCount 召回候选数
usable 后处理阈值后是否可用
topScore / topSimilarity 引擎分 / 归一化 quality(0~1)
durationMs 耗时
errorMessage 失败时

3.3 一次典型路径(含 filter fallback)

flowchart TB
    Q[query] --> L0[L0 hint → 可选 categoryFilter]
    L0 --> A1[attempt FILTERED_VECTOR]
    A1 --> PQ{isLowQuality?}
    PQ -->|否| USE1[selectedAttempt = FILTERED_VECTOR]
    PQ -->|是| A2[attempt UNFILTERED_VECTOR_RETRY]
    A2 --> USE2[selectedAttempt = RETRY<br/>fallbackReason = low_quality / no_evidence]
    USE1 --> POST[PostProcess · evidenceBlocks · relevanceLevel]
    USE2 --> POST
    POST --> LR[LookupResult 完整 Trace]

3.4 与 Agent 投影的关系

flowchart LR
    LR[LookupResult 全量 Trace] --> PROJ[RagResultProjector]
    PROJ --> AG[RagToolResult]
    AG --> F1[evidence_status]
    AG --> F2[evidence excerpt]
    AG --> F3[relevance_level 可选]
    AG --> F4[truncated / returned_count]

    LR -.->|不投影| X1[retrievalTrace]
    LR -.->|不投影| X2[rerankTrace]
    LR -.->|不投影| X3[raw scores / contextPack 全文]

人/系统要「为什么这样检索」→ 看 A 全量 或 B 落库摘要,不要只看 Agent 字段。


4. B 层:持久化 + Trace API

4.1 写入路径

组件 职责
ToolBoundary 执行后发 ToolInvocationAuditEvent(含 raw LookupResult JSON + agent JSON)
RagLookupAuditEnricher 从 LookupResult 抽有界 RAG 字段
JpaToolInvocationAuditSink 写入 tool_invocation
TraceAuditEvents.toolInvocation 另写一条 diagnosis_trace 摘要事件(不含全文 LookupResult)

4.2 tool_invocation 列(RAG)

列 lookup_knowledge 其它工具
tool_name lookup_knowledge 各自工具名
retrieval_layer 通常 L1 HARNESS
relevance_level PRECISE / REFERENCE / … null(不再写 evidence_status)
l0_match_count queryHints null
l1_match_count evidence 块数等 null
is_truncated 投影 truncated false
retrieval_details JSON rag_lookup_v1 通用 status 元数据
output_preview level/attempt 摘要 status=…
duration_ms / success 有 有

4.3 retrieval_details(rag_lookup_v1)示例

{
  "audit_schema": "rag_lookup_v1",
  "search_mode": "hybrid",
  "selected_attempt": "UNFILTERED_VECTOR_RETRY",
  "fallback_reason": "filtered_vector_low_quality",
  "category_filter": "overfilter-decoy",
  "evidence_keys": ["doc#chunk-0"],
  "sources": ["doc"],
  "evidence_candidate_count": 8,
  "evidence_block_count": 2,
  "l0_hints": { "domains": ["mysql"], "matched_keywords": ["pool"] },
  "attempts": [
    {
      "name": "FILTERED_VECTOR",
      "category_filter": "overfilter-decoy",
      "candidate_count": 2,
      "usable": false,
      "top_similarity": 0.3,
      "duration_ms": 12
    },
    {
      "name": "UNFILTERED_VECTOR_RETRY",
      "candidate_count": 5,
      "usable": true,
      "top_similarity": 0.9,
      "duration_ms": 20
    }
  ],
  "truncated": false,
  "returned_count": 2,
  "evidence_status": "EVIDENCE_FOUND",
  "invocation_status": "READY",
  "tool_call_id": "call-…"
}

默认不落库: 原始 query 全文、chunk 正文 excerpt、完整 rerankTrace(体积与隐私)。

4.4 Trace API:人怎么读 RAG

接口:

GET /api/diagnosis/{sessionId}/trace
GET /api/diagnosis/{sessionId}/trace?runId={runId}

实现: DiagnosisTraceController → DiagnosisTraceService.getTrace
按 sessionId(可选精确 runId)拉 run、steps、toolInvocations、摘要等。

响应中与 RAG 相关的核心块: DiagnosisTraceResponse.toolInvocations[]

API 字段 来源列 读法
toolName tool_name 是否为 lookup_knowledge
retrievalLayer retrieval_layer L1 / HARNESS
relevanceLevel relevance_level RAG 粗相关度(非 evidence_status)
l0MatchCount / l1MatchCount 同名列 L0/L1 规模提示
truncated is_truncated 证据是否被投影截断
outputPreview output_preview 一行摘要(level/attempt…)
retrievalDetails 解析自 retrieval_details RAG Trace 主阵地
retrievalDetailsRaw 原始 JSON 字符串 调试
durationMs / success / errorMessage 同名列 耗时与成败
inputParams 通常仅 tool_call_id、request_bytes 不含完整 query(有意)
flowchart TB
    API["GET /api/diagnosis/{sessionId}/trace"] --> SVC[DiagnosisTraceService]
    SVC --> ROW[tool_invocation 行]
    ROW --> T1[列: relevanceLevel, L0/L1 count, layer…]
    ROW --> T2[retrievalDetails Map]
    T2 --> D1[search_mode]
    T2 --> D2[selected_attempt / fallback_reason]
    T2 --> D3[attempts[] / evidence_keys]
    T2 --> D4[evidence_status 契约状态]

4.5 读 Trace 的推荐顺序(排查「这次知识库怎么检的」)

flowchart TB
    S1[找到 toolName=lookup_knowledge 的 invocation] --> S2{success?}
    S2 -->|否| E[看 errorMessage / evidence_status]
    S2 -->|是| S3[看 retrievalDetails.search_mode]
    S3 --> S4[看 selected_attempt + fallback_reason]
    S4 --> S5[看 attempts[] 每次 candidate_count / top_similarity / usable]
    S5 --> S6[看 evidence_keys / sources]
    S6 --> S7[看 relevanceLevel 列]
    S7 --> S8[需要原文?看 Agent 侧 evidence 或当时 canonical 存储 · 审计默认无 excerpt]
现象 优先看
为何走了 retry fallback_reason + 两次 attempts
是否 hybrid search_mode
滤错域 category_filter + L0 domains
相关度档 列 relevanceLevel(PRECISE/REFERENCE)
返回了哪些块 evidence_keys / sources(无正文)
Agent 是否被截断 truncated / returned_count

4.6 diagnosis_trace 事件 vs tool_invocation 行

通道 内容 用途
tool_invocation 行 RAG 富字段完整摘要 主审计/回放
diagnosis_trace 中 TOOL_INVOCATION tool_call_id、status、字节数、has_raw_result 等薄摘要 时间线事件,不含完整 retrieval_details

查 RAG 细节以 toolInvocations[].retrievalDetails 为准。


5. 与 Agent / Eval 的边界

flowchart LR
    subgraph human["人 / 运维 / 评测"]
        TRACE[Trace API]
        EVAL[Offline eval]
    end

    subgraph model["模型"]
        AGENT[RagToolResult only]
    end

    TI[(tool_invocation)] --> TRACE
    FX[fixtures] --> EVAL
    PROJ[Projector] --> AGENT
消费者 能看到
Agent evidence + 可选 relevance_level,无 attempt 细节
Trace API 落库摘要:mode/attempt/fallback/keys/level…
Offline eval 冻结 fixture 全量 LookupResult(含 trace),与 golden 比对

6. 与旧文档差异

旧(archive retrieval-observability) 现
vector-store.mode 多后端 search_mode dense|hybrid,单一 V2 store
sink 理想化未落地 RagLookupAuditEnricher + 列回填
relevance_level 混用 evidence_status 列仅 RAG 等级;契约状态在 details
未写清 Trace API 读法 本文 §4.4–4.5

7. 代码锚点

职责 类 / 路径
内建 Trace LookupKnowledgeTool、RetrievalTrace、LookupResult
投影 RagResultProjector、RagToolResult
审计事件 ToolInvocationAuditEvent、ToolBoundary
富化 RagLookupAuditEnricher
落库 JpaToolInvocationAuditSink、ToolInvocation
Trace API DiagnosisTraceController、DiagnosisTraceService、DiagnosisTraceResponse.ToolInvocationTrace
离线回归 eval/rag-retrieval/、scripts/eval_rag_retrieval.py

8. 已知限制

  • 持久化 不存 完整 query/excerpt(有意);要正文需 Agent 侧证据或其它存储
  • dedup_reason 列可能仍为空
  • 非 lookup_knowledge 工具仍为薄审计
  • 历史 tool_invocation 行可能仍把 evidence_status 写进 relevance_level(旧 sink)
  • diagnosis_trace 时间线事件不替代 retrieval_details

9. 相关文档

文档 内容
mvp/architecture/RAG知识检索架构.md 检索主架构
../engineering/rag/RAG-Agent如何读relevance_level.md Agent 如何读 level
../engineering/rag/RAG-Hybrid质量分与后处理.md quality / 排序闸门
../engineering/rag/RAG离线评测-基线设计.md 离线评测(非运行时 Trace)
mvp/architecture/session-trace-lifecycle.md 会话/run Trace 总览(若存在)