Files
SuperBizAgent-java/openspec/changes/archive/2026-07-06-modular-rag-pipeline/proposal.md
T

173 lines
7.0 KiB
Markdown

# Modular RAG Pipeline Proposal
## Problem
`lookup_knowledge` already exposes structured retrieval evidence, but the runtime flow is still concentrated inside `LookupKnowledgeTool`. Query understanding, vector retrieval, relevance normalization, evidence block construction, session deduplication, and trace recording are tightly coupled. This makes the RAG path harder to explain, test, evolve, and present as a modular Agent engineering design.
The current result contract also still carries the old `primary` / `supplement` model, where `primary` means L0 exact match and `supplement` means L1 semantic retrieval. That contract no longer matches the intended architecture: L0 should be a query understanding and retrieval-control signal, while L1 vector retrieval should be the main evidence source.
## Proposed Change
Refactor `lookup_knowledge` into a modular RAG pipeline while preserving the explicit Agent tool boundary and `tool_invocation` evidence trace.
Target pipeline:
```text
LookupKnowledgeTool
-> KnowledgeQueryTransformer
-> KnowledgeDocumentRetriever
-> KnowledgeEvidencePostProcessor
-> KnowledgeContextPacker
-> LookupResultAssembler
-> ToolInvocationRecorder
```
### Query Transformation
Introduce a query transformer around the current L0 analysis.
L0 SHALL provide:
- domain/category hints
- matched keywords
- extracted entities
- optional metadata filter candidate
- traceable query understanding data
L0 SHALL NOT act as a main evidence retrieval path in the normal flow.
### Retrieval
L1 vector retrieval remains the main document retrieval path through `VectorSearchService`, which already supports Spring AI `VectorStore` as the preferred path and Milvus SDK fallback.
MVP fallback strategy:
```text
1. Run filtered vector retrieval with the L0-derived category filter when unambiguous.
2. If filtered retrieval returns no usable evidence or low-quality evidence, retry raw query through unfiltered vector retrieval.
3. If unfiltered retrieval also fails, return no_evidence.
```
The fallback SHALL skip the L0-derived filter rather than returning L0 documents as fact evidence.
### Post-Retrieval Processing
Move evidence post-processing out of `LookupKnowledgeTool`.
The post-processor SHALL handle:
- relevance normalization
- evidence block creation
- source-level deduplication
- rule-based lightweight rerank
- retrieval trace and rerank trace generation
The first rerank implementation should be rule-based, using available signals such as vector score, domain match, entity match, keyword match, source type, and whether evidence aligns with query hints.
### Context Packing
Add a context packing step that converts final evidence blocks into an Agent-facing context package.
The packer SHALL:
- keep source/title/breadcrumb visible
- obey a configurable character budget in the MVP
- prioritize reranked evidence order
- avoid duplicate source content
- produce a compact summary of included and omitted evidence
### Result Contract
This change intentionally updates the `lookup_knowledge` return contract.
New preferred contract:
```text
found
evidenceBlocks
contextPack
retrievalTrace
rerankTrace
relevanceLevel
completenessHint
retrievedDomainsThisSession
message
```
The old `primary` and `supplement` fields may be removed as part of this change, because they encode the outdated assumption that L0 is the primary evidence source and L1 is supplemental evidence.
## Scope
In scope:
- Refactor `LookupKnowledgeTool` into a thin tool boundary and pipeline orchestrator.
- Add local pipeline classes and DTOs for query transformation, retrieval result normalization, post-processing, context packing, and traces.
- Update `LookupResult` to prefer `evidenceBlocks`, `contextPack`, `retrievalTrace`, and `rerankTrace`.
- Remove or deprecate `primary` / `supplement` according to the final spec.
- Update `ToolInvocationRecorder` to persist compact summaries for evidence blocks, context pack, retrieval trace, fallback reason, and rerank trace.
- Update tool description / prompt wording so Agent behavior matches the new contract.
- Update tests for filtered retrieval, unfiltered retry, rerank, context packing, evidence persistence, and result contract changes.
- Update `rag-knowledge-retrieval` spec to remove L0 primary fallback and old compatibility-field requirements.
Out of scope:
- Replacing the explicit `lookup_knowledge` tool with an implicit Advisor.
- Introducing a model-based reranker or cross-encoder.
- Introducing BM25, RRF, Elasticsearch, or OpenSearch.
- Migrating document upload, chunking, embedding writes, or Milvus schema.
- Changing the Agent decision of when to call `lookup_knowledge`.
## Interface Impact
Impact level: L4 breaking interface change.
Reason:
- `LookupResult.primary` and `LookupResult.supplement` may be removed.
- The JSON returned by the `lookup_knowledge` Agent tool changes shape.
- Tests and internal consumers that read `primary` / `supplement` must migrate to `evidenceBlocks` and `contextPack`.
Known affected areas:
- `LookupKnowledgeTool`
- `LookupResult`
- `PrimaryResult` / `SupplementResult`
- `ToolInvocationRecorder`
- `LookupKnowledgeToolTest`
- `ToolInvocationRecorderTest`
- Agent tool prompt / executor prompt references
- `rag-knowledge-retrieval` OpenSpec requirements
- RAG docs under `mvp/architecture/`
Migration approach:
- Update all in-repo consumers in the same change.
- Keep `lookup_knowledge` tool name and input argument unchanged.
- Keep `tool_invocation` persistence compatible at table level while enriching `retrieval_details`.
- Record fallback and no-evidence semantics explicitly so Verifier and Eval do not treat hint-only data as fact evidence.
## Context Constraints
Relevant project decisions:
- `lookup_knowledge` remains an explicit Agent evidence tool.
- L0 is already documented as a hint layer, not a final decision layer.
- Spring AI `VectorStore` is already the preferred retrieval path behind `VectorSearchService`.
- Milvus SDK fallback remains valuable for MVP resilience.
- `tool_invocation` is the stable evidence trace used by trace inspection, Verifier, and Eval.
- Existing `rag-knowledge-retrieval` spec still contains legacy fallback and compatibility-field requirements that must be changed.
## Risks
- Breaking result contract may affect prompt behavior because the tool JSON changes.
- Removing `primary` / `supplement` requires updating tests and recorder preview logic.
- Rule-based rerank may create unexpected ordering changes if score semantics are not handled carefully.
- Context packing can hide useful evidence if the budget is too small.
- Existing OpenSpec requirements conflict with the new L0 fallback policy and must be updated before implementation.
## Open Questions
- What exact minimum fields should `ContextPack`, `RetrievalTrace`, and `RerankTrace` expose in the committed spec?
- Should `PrimaryResult` and `SupplementResult` classes be deleted immediately or left deprecated for one change cycle?
- What threshold defines "filtered retrieval low quality" for triggering unfiltered retry?