Files
SuperBizAgent-java/openspec/changes/archive/2026-07-06-modular-rag-pipeline/proposal.md
T

7.0 KiB

Modular RAG Pipeline Proposal

Problem

lookup_knowledge already exposes structured retrieval evidence, but the runtime flow is still concentrated inside LookupKnowledgeTool. Query understanding, vector retrieval, relevance normalization, evidence block construction, session deduplication, and trace recording are tightly coupled. This makes the RAG path harder to explain, test, evolve, and present as a modular Agent engineering design.

The current result contract also still carries the old primary / supplement model, where primary means L0 exact match and supplement means L1 semantic retrieval. That contract no longer matches the intended architecture: L0 should be a query understanding and retrieval-control signal, while L1 vector retrieval should be the main evidence source.

Proposed Change

Refactor lookup_knowledge into a modular RAG pipeline while preserving the explicit Agent tool boundary and tool_invocation evidence trace.

Target pipeline:

LookupKnowledgeTool
  -> KnowledgeQueryTransformer
  -> KnowledgeDocumentRetriever
  -> KnowledgeEvidencePostProcessor
  -> KnowledgeContextPacker
  -> LookupResultAssembler
  -> ToolInvocationRecorder

Query Transformation

Introduce a query transformer around the current L0 analysis.

L0 SHALL provide:

  • domain/category hints
  • matched keywords
  • extracted entities
  • optional metadata filter candidate
  • traceable query understanding data

L0 SHALL NOT act as a main evidence retrieval path in the normal flow.

Retrieval

L1 vector retrieval remains the main document retrieval path through VectorSearchService, which already supports Spring AI VectorStore as the preferred path and Milvus SDK fallback.

MVP fallback strategy:

1. Run filtered vector retrieval with the L0-derived category filter when unambiguous.
2. If filtered retrieval returns no usable evidence or low-quality evidence, retry raw query through unfiltered vector retrieval.
3. If unfiltered retrieval also fails, return no_evidence.

The fallback SHALL skip the L0-derived filter rather than returning L0 documents as fact evidence.

Post-Retrieval Processing

Move evidence post-processing out of LookupKnowledgeTool.

The post-processor SHALL handle:

  • relevance normalization
  • evidence block creation
  • source-level deduplication
  • rule-based lightweight rerank
  • retrieval trace and rerank trace generation

The first rerank implementation should be rule-based, using available signals such as vector score, domain match, entity match, keyword match, source type, and whether evidence aligns with query hints.

Context Packing

Add a context packing step that converts final evidence blocks into an Agent-facing context package.

The packer SHALL:

  • keep source/title/breadcrumb visible
  • obey a configurable character budget in the MVP
  • prioritize reranked evidence order
  • avoid duplicate source content
  • produce a compact summary of included and omitted evidence

Result Contract

This change intentionally updates the lookup_knowledge return contract.

New preferred contract:

found
evidenceBlocks
contextPack
retrievalTrace
rerankTrace
relevanceLevel
completenessHint
retrievedDomainsThisSession
message

The old primary and supplement fields may be removed as part of this change, because they encode the outdated assumption that L0 is the primary evidence source and L1 is supplemental evidence.

Scope

In scope:

  • Refactor LookupKnowledgeTool into a thin tool boundary and pipeline orchestrator.
  • Add local pipeline classes and DTOs for query transformation, retrieval result normalization, post-processing, context packing, and traces.
  • Update LookupResult to prefer evidenceBlocks, contextPack, retrievalTrace, and rerankTrace.
  • Remove or deprecate primary / supplement according to the final spec.
  • Update ToolInvocationRecorder to persist compact summaries for evidence blocks, context pack, retrieval trace, fallback reason, and rerank trace.
  • Update tool description / prompt wording so Agent behavior matches the new contract.
  • Update tests for filtered retrieval, unfiltered retry, rerank, context packing, evidence persistence, and result contract changes.
  • Update rag-knowledge-retrieval spec to remove L0 primary fallback and old compatibility-field requirements.

Out of scope:

  • Replacing the explicit lookup_knowledge tool with an implicit Advisor.
  • Introducing a model-based reranker or cross-encoder.
  • Introducing BM25, RRF, Elasticsearch, or OpenSearch.
  • Migrating document upload, chunking, embedding writes, or Milvus schema.
  • Changing the Agent decision of when to call lookup_knowledge.

Interface Impact

Impact level: L4 breaking interface change.

Reason:

  • LookupResult.primary and LookupResult.supplement may be removed.
  • The JSON returned by the lookup_knowledge Agent tool changes shape.
  • Tests and internal consumers that read primary / supplement must migrate to evidenceBlocks and contextPack.

Known affected areas:

  • LookupKnowledgeTool
  • LookupResult
  • PrimaryResult / SupplementResult
  • ToolInvocationRecorder
  • LookupKnowledgeToolTest
  • ToolInvocationRecorderTest
  • Agent tool prompt / executor prompt references
  • rag-knowledge-retrieval OpenSpec requirements
  • RAG docs under mvp/architecture/

Migration approach:

  • Update all in-repo consumers in the same change.
  • Keep lookup_knowledge tool name and input argument unchanged.
  • Keep tool_invocation persistence compatible at table level while enriching retrieval_details.
  • Record fallback and no-evidence semantics explicitly so Verifier and Eval do not treat hint-only data as fact evidence.

Context Constraints

Relevant project decisions:

  • lookup_knowledge remains an explicit Agent evidence tool.
  • L0 is already documented as a hint layer, not a final decision layer.
  • Spring AI VectorStore is already the preferred retrieval path behind VectorSearchService.
  • Milvus SDK fallback remains valuable for MVP resilience.
  • tool_invocation is the stable evidence trace used by trace inspection, Verifier, and Eval.
  • Existing rag-knowledge-retrieval spec still contains legacy fallback and compatibility-field requirements that must be changed.

Risks

  • Breaking result contract may affect prompt behavior because the tool JSON changes.
  • Removing primary / supplement requires updating tests and recorder preview logic.
  • Rule-based rerank may create unexpected ordering changes if score semantics are not handled carefully.
  • Context packing can hide useful evidence if the budget is too small.
  • Existing OpenSpec requirements conflict with the new L0 fallback policy and must be updated before implementation.

Open Questions

  • What exact minimum fields should ContextPack, RetrievalTrace, and RerankTrace expose in the committed spec?
  • Should PrimaryResult and SupplementResult classes be deleted immediately or left deprecated for one change cycle?
  • What threshold defines "filtered retrieval low quality" for triggering unfiltered retry?