7.0 KiB
Modular RAG Pipeline Proposal
Problem
lookup_knowledge already exposes structured retrieval evidence, but the runtime flow is still concentrated inside LookupKnowledgeTool. Query understanding, vector retrieval, relevance normalization, evidence block construction, session deduplication, and trace recording are tightly coupled. This makes the RAG path harder to explain, test, evolve, and present as a modular Agent engineering design.
The current result contract also still carries the old primary / supplement model, where primary means L0 exact match and supplement means L1 semantic retrieval. That contract no longer matches the intended architecture: L0 should be a query understanding and retrieval-control signal, while L1 vector retrieval should be the main evidence source.
Proposed Change
Refactor lookup_knowledge into a modular RAG pipeline while preserving the explicit Agent tool boundary and tool_invocation evidence trace.
Target pipeline:
LookupKnowledgeTool
-> KnowledgeQueryTransformer
-> KnowledgeDocumentRetriever
-> KnowledgeEvidencePostProcessor
-> KnowledgeContextPacker
-> LookupResultAssembler
-> ToolInvocationRecorder
Query Transformation
Introduce a query transformer around the current L0 analysis.
L0 SHALL provide:
- domain/category hints
- matched keywords
- extracted entities
- optional metadata filter candidate
- traceable query understanding data
L0 SHALL NOT act as a main evidence retrieval path in the normal flow.
Retrieval
L1 vector retrieval remains the main document retrieval path through VectorSearchService, which already supports Spring AI VectorStore as the preferred path and Milvus SDK fallback.
MVP fallback strategy:
1. Run filtered vector retrieval with the L0-derived category filter when unambiguous.
2. If filtered retrieval returns no usable evidence or low-quality evidence, retry raw query through unfiltered vector retrieval.
3. If unfiltered retrieval also fails, return no_evidence.
The fallback SHALL skip the L0-derived filter rather than returning L0 documents as fact evidence.
Post-Retrieval Processing
Move evidence post-processing out of LookupKnowledgeTool.
The post-processor SHALL handle:
- relevance normalization
- evidence block creation
- source-level deduplication
- rule-based lightweight rerank
- retrieval trace and rerank trace generation
The first rerank implementation should be rule-based, using available signals such as vector score, domain match, entity match, keyword match, source type, and whether evidence aligns with query hints.
Context Packing
Add a context packing step that converts final evidence blocks into an Agent-facing context package.
The packer SHALL:
- keep source/title/breadcrumb visible
- obey a configurable character budget in the MVP
- prioritize reranked evidence order
- avoid duplicate source content
- produce a compact summary of included and omitted evidence
Result Contract
This change intentionally updates the lookup_knowledge return contract.
New preferred contract:
found
evidenceBlocks
contextPack
retrievalTrace
rerankTrace
relevanceLevel
completenessHint
retrievedDomainsThisSession
message
The old primary and supplement fields may be removed as part of this change, because they encode the outdated assumption that L0 is the primary evidence source and L1 is supplemental evidence.
Scope
In scope:
- Refactor
LookupKnowledgeToolinto a thin tool boundary and pipeline orchestrator. - Add local pipeline classes and DTOs for query transformation, retrieval result normalization, post-processing, context packing, and traces.
- Update
LookupResultto preferevidenceBlocks,contextPack,retrievalTrace, andrerankTrace. - Remove or deprecate
primary/supplementaccording to the final spec. - Update
ToolInvocationRecorderto persist compact summaries for evidence blocks, context pack, retrieval trace, fallback reason, and rerank trace. - Update tool description / prompt wording so Agent behavior matches the new contract.
- Update tests for filtered retrieval, unfiltered retry, rerank, context packing, evidence persistence, and result contract changes.
- Update
rag-knowledge-retrievalspec to remove L0 primary fallback and old compatibility-field requirements.
Out of scope:
- Replacing the explicit
lookup_knowledgetool with an implicit Advisor. - Introducing a model-based reranker or cross-encoder.
- Introducing BM25, RRF, Elasticsearch, or OpenSearch.
- Migrating document upload, chunking, embedding writes, or Milvus schema.
- Changing the Agent decision of when to call
lookup_knowledge.
Interface Impact
Impact level: L4 breaking interface change.
Reason:
LookupResult.primaryandLookupResult.supplementmay be removed.- The JSON returned by the
lookup_knowledgeAgent tool changes shape. - Tests and internal consumers that read
primary/supplementmust migrate toevidenceBlocksandcontextPack.
Known affected areas:
LookupKnowledgeToolLookupResultPrimaryResult/SupplementResultToolInvocationRecorderLookupKnowledgeToolTestToolInvocationRecorderTest- Agent tool prompt / executor prompt references
rag-knowledge-retrievalOpenSpec requirements- RAG docs under
mvp/architecture/
Migration approach:
- Update all in-repo consumers in the same change.
- Keep
lookup_knowledgetool name and input argument unchanged. - Keep
tool_invocationpersistence compatible at table level while enrichingretrieval_details. - Record fallback and no-evidence semantics explicitly so Verifier and Eval do not treat hint-only data as fact evidence.
Context Constraints
Relevant project decisions:
lookup_knowledgeremains an explicit Agent evidence tool.- L0 is already documented as a hint layer, not a final decision layer.
- Spring AI
VectorStoreis already the preferred retrieval path behindVectorSearchService. - Milvus SDK fallback remains valuable for MVP resilience.
tool_invocationis the stable evidence trace used by trace inspection, Verifier, and Eval.- Existing
rag-knowledge-retrievalspec still contains legacy fallback and compatibility-field requirements that must be changed.
Risks
- Breaking result contract may affect prompt behavior because the tool JSON changes.
- Removing
primary/supplementrequires updating tests and recorder preview logic. - Rule-based rerank may create unexpected ordering changes if score semantics are not handled carefully.
- Context packing can hide useful evidence if the budget is too small.
- Existing OpenSpec requirements conflict with the new L0 fallback policy and must be updated before implementation.
Open Questions
- What exact minimum fields should
ContextPack,RetrievalTrace, andRerankTraceexpose in the committed spec? - Should
PrimaryResultandSupplementResultclasses be deleted immediately or left deprecated for one change cycle? - What threshold defines "filtered retrieval low quality" for triggering unfiltered retry?