# Modular RAG Pipeline Proposal ## Problem `lookup_knowledge` already exposes structured retrieval evidence, but the runtime flow is still concentrated inside `LookupKnowledgeTool`. Query understanding, vector retrieval, relevance normalization, evidence block construction, session deduplication, and trace recording are tightly coupled. This makes the RAG path harder to explain, test, evolve, and present as a modular Agent engineering design. The current result contract also still carries the old `primary` / `supplement` model, where `primary` means L0 exact match and `supplement` means L1 semantic retrieval. That contract no longer matches the intended architecture: L0 should be a query understanding and retrieval-control signal, while L1 vector retrieval should be the main evidence source. ## Proposed Change Refactor `lookup_knowledge` into a modular RAG pipeline while preserving the explicit Agent tool boundary and `tool_invocation` evidence trace. Target pipeline: ```text LookupKnowledgeTool -> KnowledgeQueryTransformer -> KnowledgeDocumentRetriever -> KnowledgeEvidencePostProcessor -> KnowledgeContextPacker -> LookupResultAssembler -> ToolInvocationRecorder ``` ### Query Transformation Introduce a query transformer around the current L0 analysis. L0 SHALL provide: - domain/category hints - matched keywords - extracted entities - optional metadata filter candidate - traceable query understanding data L0 SHALL NOT act as a main evidence retrieval path in the normal flow. ### Retrieval L1 vector retrieval remains the main document retrieval path through `VectorSearchService`, which already supports Spring AI `VectorStore` as the preferred path and Milvus SDK fallback. MVP fallback strategy: ```text 1. Run filtered vector retrieval with the L0-derived category filter when unambiguous. 2. If filtered retrieval returns no usable evidence or low-quality evidence, retry raw query through unfiltered vector retrieval. 3. If unfiltered retrieval also fails, return no_evidence. ``` The fallback SHALL skip the L0-derived filter rather than returning L0 documents as fact evidence. ### Post-Retrieval Processing Move evidence post-processing out of `LookupKnowledgeTool`. The post-processor SHALL handle: - relevance normalization - evidence block creation - source-level deduplication - rule-based lightweight rerank - retrieval trace and rerank trace generation The first rerank implementation should be rule-based, using available signals such as vector score, domain match, entity match, keyword match, source type, and whether evidence aligns with query hints. ### Context Packing Add a context packing step that converts final evidence blocks into an Agent-facing context package. The packer SHALL: - keep source/title/breadcrumb visible - obey a configurable character budget in the MVP - prioritize reranked evidence order - avoid duplicate source content - produce a compact summary of included and omitted evidence ### Result Contract This change intentionally updates the `lookup_knowledge` return contract. New preferred contract: ```text found evidenceBlocks contextPack retrievalTrace rerankTrace relevanceLevel completenessHint retrievedDomainsThisSession message ``` The old `primary` and `supplement` fields may be removed as part of this change, because they encode the outdated assumption that L0 is the primary evidence source and L1 is supplemental evidence. ## Scope In scope: - Refactor `LookupKnowledgeTool` into a thin tool boundary and pipeline orchestrator. - Add local pipeline classes and DTOs for query transformation, retrieval result normalization, post-processing, context packing, and traces. - Update `LookupResult` to prefer `evidenceBlocks`, `contextPack`, `retrievalTrace`, and `rerankTrace`. - Remove or deprecate `primary` / `supplement` according to the final spec. - Update `ToolInvocationRecorder` to persist compact summaries for evidence blocks, context pack, retrieval trace, fallback reason, and rerank trace. - Update tool description / prompt wording so Agent behavior matches the new contract. - Update tests for filtered retrieval, unfiltered retry, rerank, context packing, evidence persistence, and result contract changes. - Update `rag-knowledge-retrieval` spec to remove L0 primary fallback and old compatibility-field requirements. Out of scope: - Replacing the explicit `lookup_knowledge` tool with an implicit Advisor. - Introducing a model-based reranker or cross-encoder. - Introducing BM25, RRF, Elasticsearch, or OpenSearch. - Migrating document upload, chunking, embedding writes, or Milvus schema. - Changing the Agent decision of when to call `lookup_knowledge`. ## Interface Impact Impact level: L4 breaking interface change. Reason: - `LookupResult.primary` and `LookupResult.supplement` may be removed. - The JSON returned by the `lookup_knowledge` Agent tool changes shape. - Tests and internal consumers that read `primary` / `supplement` must migrate to `evidenceBlocks` and `contextPack`. Known affected areas: - `LookupKnowledgeTool` - `LookupResult` - `PrimaryResult` / `SupplementResult` - `ToolInvocationRecorder` - `LookupKnowledgeToolTest` - `ToolInvocationRecorderTest` - Agent tool prompt / executor prompt references - `rag-knowledge-retrieval` OpenSpec requirements - RAG docs under `mvp/architecture/` Migration approach: - Update all in-repo consumers in the same change. - Keep `lookup_knowledge` tool name and input argument unchanged. - Keep `tool_invocation` persistence compatible at table level while enriching `retrieval_details`. - Record fallback and no-evidence semantics explicitly so Verifier and Eval do not treat hint-only data as fact evidence. ## Context Constraints Relevant project decisions: - `lookup_knowledge` remains an explicit Agent evidence tool. - L0 is already documented as a hint layer, not a final decision layer. - Spring AI `VectorStore` is already the preferred retrieval path behind `VectorSearchService`. - Milvus SDK fallback remains valuable for MVP resilience. - `tool_invocation` is the stable evidence trace used by trace inspection, Verifier, and Eval. - Existing `rag-knowledge-retrieval` spec still contains legacy fallback and compatibility-field requirements that must be changed. ## Risks - Breaking result contract may affect prompt behavior because the tool JSON changes. - Removing `primary` / `supplement` requires updating tests and recorder preview logic. - Rule-based rerank may create unexpected ordering changes if score semantics are not handled carefully. - Context packing can hide useful evidence if the budget is too small. - Existing OpenSpec requirements conflict with the new L0 fallback policy and must be updated before implementation. ## Open Questions - What exact minimum fields should `ContextPack`, `RetrievalTrace`, and `RerankTrace` expose in the committed spec? - Should `PrimaryResult` and `SupplementResult` classes be deleted immediately or left deprecated for one change cycle? - What threshold defines "filtered retrieval low quality" for triggering unfiltered retry?