## Context The current `lookup_knowledge` implementation is operational but still organized around a legacy L0/L1 result model. `LookupKnowledgeTool` currently performs query analysis, vector retrieval, relevance normalization, evidence block construction, session deduplication, and trace recording in one class. Evidence blocks already exist, but post-retrieval processing is not a first-class pipeline boundary. The project direction is already documented as: - keep `lookup_knowledge` as an explicit Agent evidence tool; - treat L0 as domain/entity/keyword hint data; - use L1 vector retrieval as the main evidence source; - keep Spring AI `VectorStore` behind `VectorSearchService`; - preserve `tool_invocation` as the evidence trace for Verifier, Eval, and trace APIs. This change turns that architecture into code structure and updates the tool result contract so future Agent, Verifier, and Eval flows consume structured evidence instead of the old `primary` / `supplement` split. ## Goals / Non-Goals **Goals:** - Refactor `lookup_knowledge` into a modular RAG pipeline. - Keep L0 as query understanding and retrieval-control signal. - Keep L1 vector retrieval as the main document retrieval path. - Add a simple MVP fallback: retry raw query through unfiltered L1 when filtered L1 is low quality. - Move relevance normalization, evidence block creation, deduplication, and lightweight rerank into post-retrieval processing. - Add context packing as a first-class output. - Replace the old `primary` / `supplement` result contract with `evidenceBlocks`, `contextPack`, `retrievalTrace`, and `rerankTrace`. - Keep `lookup_knowledge` tool name and input argument unchanged. - Keep `tool_invocation` table schema stable and enrich `retrieval_details`. **Non-Goals:** - Do not replace `lookup_knowledge` with an implicit Advisor. - Do not add a model-based reranker, cross-encoder, BM25, RRF, Elasticsearch, or OpenSearch. - Do not migrate document upload, chunking, embedding write path, Milvus schema, or vector collection layout. - Do not change when the Agent decides to call `lookup_knowledge`. - Do not introduce multi-query expansion in this change. ## Decisions ### Decision: Use explicit pipeline components Create a local pipeline behind `LookupKnowledgeTool`: ```text LookupKnowledgeTool -> KnowledgeQueryTransformer -> KnowledgeDocumentRetriever -> KnowledgeEvidencePostProcessor -> KnowledgeContextPacker -> LookupResultAssembler -> ToolInvocationRecorder ``` Rationale: this keeps the Agent tool boundary stable while making the RAG flow easy to test and explain. It also maps cleanly to Spring AI modular RAG concepts without hiding business observability inside an Advisor. Alternative considered: keep all logic in `LookupKnowledgeTool` and only add fields. Rejected because the class would continue to mix query transformation, retrieval, post-processing, and trace responsibilities. ### Decision: L0 is a query transformer signal, not a main retriever `KnowledgeQueryTransformer` will wrap the existing `KnowledgeIndexService.analyzeQuery` behavior and produce a transformed query object containing: - original query - rewritten query, initially equal to the raw query unless a future rule rewrites it - domain hints - matched keywords - entities - optional category filter - L0 titles for trace only L0 matches must not be converted into normal evidence candidates in the main path. Rationale: L0 keyword/frontmatter matching is useful for controlling retrieval, but it is not reliable enough to be treated as fact evidence when L1 cannot support it. Alternative considered: combine L0 and L1 into one candidate list. Rejected because it makes L0 an equal retrieval layer again and conflicts with the desired architecture. ### Decision: Use filtered L1 first, then unfiltered L1 retry `KnowledgeDocumentRetriever` will perform: ```text attempt 1: vector search with L0-derived category filter, when unambiguous attempt 2: raw query vector search without the L0-derived filter, when attempt 1 is low quality ``` Filtered retrieval is low quality when any of the following is true: - no candidates are returned; - post-processing would produce zero evidence blocks; - top candidate normalized similarity is below `retrieval.normalization.reference-threshold`. Rationale: the most likely MVP failure mode is an over-strict or wrong metadata filter. A raw unfiltered vector retry addresses that without adding a complex multi-stage fallback policy. Alternative considered: return L0 documents as weak fallback evidence. Rejected for this MVP because it can let keyword hints masquerade as factual evidence. ### Decision: Keep rerank rule-based `KnowledgeEvidencePostProcessor` will rerank with deterministic signals: - vector score / normalized similarity; - domain match with query hints; - entity match; - keyword match; - source type or metadata priority when available; - retrieval attempt, with filtered hits not automatically preferred over stronger unfiltered hits. The rerank trace should explain major score contributions per final evidence block. Rationale: a rule-based reranker is explainable, cheap, testable, and enough for the interview-oriented MVP. It also avoids introducing model latency and new dependencies. Alternative considered: model-based rerank. Rejected as out of scope until evaluation shows a need. ### Decision: Context pack becomes the Agent-facing content `KnowledgeContextPacker` will turn final evidence blocks into a compact context package: ```text packedText strategy charBudget usedChars includedSources omittedSources ``` MVP packing uses a character budget rather than exact token counting. The packer preserves source/title/breadcrumb/hit reasons and truncates content only after preserving metadata. Rationale: the Agent should consume curated evidence context instead of inferring semantics from `primary` and `supplement`. Alternative considered: keep `primary` and `supplement` as the Agent-facing fields. Rejected because those names preserve the outdated L0 primary / L1 supplement model. ### Decision: Break the old result contract deliberately `LookupResult` will move to the preferred contract: ```text found evidenceBlocks contextPack retrievalTrace rerankTrace relevanceLevel completenessHint retrievedDomainsThisSession message ``` `primary` and `supplement` may be removed in this change, and `PrimaryResult` / `SupplementResult` may be deleted if no production code remains after migration. Rationale: this is an MVP project intended to demonstrate clean modular RAG. Keeping obsolete fields would force later code to preserve misleading semantics. Alternative considered: add new fields while keeping old fields deprecated. Rejected because the user explicitly accepted removing old fields if the later flow becomes cleaner. ### Decision: Preserve persistence compatibility at table level `ToolInvocationRecorder` will stop depending on `result.getPrimary()` for output preview. Preview should come from `contextPack.packedText` or the top evidence block content. `retrieval_details` will include compact summaries: - query transform summary; - retrieval trace with attempts and fallback reason; - rerank trace summary; - context pack summary; - evidence block summaries. Rationale: trace consumers already read `tool_invocation` rows. The database schema can remain stable while JSON details evolve. Alternative considered: add columns for each new trace object. Rejected because the current trace model already stores retrieval-specific details in JSON and does not need schema churn for this change. ## Risks / Trade-offs - [Risk] Breaking tool output shape can affect Agent prompt behavior. -> Mitigation: update tool description and executor prompt to prefer `contextPack` and `evidenceBlocks`. - [Risk] Existing tests assert `primary` / `supplement`. -> Mitigation: migrate tests to evidence/context/traces in the same change. - [Risk] Rule-based rerank can reorder evidence unexpectedly. -> Mitigation: persist `rerankTrace` and cover ordering behavior with tests. - [Risk] Context packing can omit useful evidence under a small budget. -> Mitigation: record included and omitted sources and keep the character budget configurable. - [Risk] Existing OpenSpec requirements conflict with the new fallback policy. -> Mitigation: update `rag-knowledge-retrieval` delta before apply and validate the change. - [Risk] Unfiltered retry may increase latency. -> Mitigation: retry only when filtered result is empty or below reference quality, and record attempt counts/duration in trace. ## Migration Plan 1. Add new internal DTOs and pipeline components while keeping `LookupKnowledgeTool` as the public tool bean. 2. Migrate `LookupKnowledgeTool` orchestration to the pipeline. 3. Update `LookupResult` to the new result contract and remove `primary` / `supplement` usages. 4. Update `ToolInvocationRecorder` preview and retrieval details to use context/evidence/traces. 5. Update prompt text and tool description. 6. Update tests for the new contract. 7. Run targeted unit tests and OpenSpec validation. Rollback strategy: - Revert the change as one unit if Agent tool behavior regresses. - Because the table schema remains stable and the tool input is unchanged, rollback does not require database migration. ## Open Questions - The exact default context pack character budget should be finalized during implementation; recommended MVP default is 3000 to 5000 characters. - Whether to delete `PrimaryResult` and `SupplementResult` immediately depends on final compile-time references after migration.