feat(rag): modularize knowledge retrieval pipeline

This commit is contained in:
zhuyongxin
2026-07-06 17:06:05 +08:00
parent a375daead7
commit cf3333d607
38 changed files with 2981 additions and 1033 deletions
@@ -0,0 +1 @@
committed
@@ -0,0 +1,396 @@
# Modular RAG Pipeline Decisions
## Discover Summary
- Capability source: sm-flow Discover using local repository evidence and existing OpenSpec/devflow context.
- Slug: `modular-rag-pipeline`.
- Scale: standard.
- Goal: turn `lookup_knowledge` into a modular RAG pipeline suitable for Agent engineering interview use, while keeping the explicit Agent tool and evidence trace.
## Context Evidence
### Existing Architecture
- `mvp/architecture/rag-architecture.md` documents the desired boundary: mature framework retrieval plus business-observable orchestration.
- `mvp/architecture/retrieval-observability.md` says L0 is a hint/explainability layer and L1 semantic retrieval is the default recall path.
- `VectorSearchService` is already the retrieval facade and supports Spring AI `VectorStore` with SDK fallback.
- `LookupKnowledgeTool` currently still owns query analysis, L1 invocation, relevance normalization, result assembly, evidence block construction, session dedup, and recorder calls.
### Existing Evidence Blocks
- `EvidenceBlock` already exists.
- `LookupResult` already has `evidenceBlocks`, `evidenceCandidateCount`, and `evidenceBlockCount`.
- `LookupKnowledgeTool` currently builds evidence blocks internally.
- `ToolInvocationRecorder` already persists compact evidence block summaries in `retrieval_details`.
Conclusion: evidence blocks are partially implemented, but the post-retrieval module boundary is not.
### Current Gaps
- Context packing is not a first-class module.
- Rerank is not a first-class module; only original semantic rank and trace summary sorting exist.
- The `primary` / `supplement` result model still encodes old L0/L1 semantics.
- `lookup_knowledge` tool description still describes the old two-stage retrieval model.
## Question Pool
| ID | Dimension | Question | Mode | Status |
| --- | --- | --- | --- | --- |
| Q1 | Boundary | Should this be internal-only refactor or allow return contract changes? | user-interview | confirmed |
| Q2 | Fallback | If filtered L1 fails, should L0 provide weak fallback evidence? | user-interview | confirmed |
| Q3 | Interface | Can `LookupResult` add new fields and remove old `primary` / `supplement` if simpler? | user-interview | confirmed |
| Q4 | Architecture | Does current repo already have evidence blocks, context packing, and rerank? | evidence-driven | resolved |
| Q5 | Compatibility | Which in-repo consumers reference `primary` / `supplement`? | evidence-driven | resolved |
| Q6 | Commit detail | What are the exact fields and thresholds for traces/context pack/low quality? | user-interview or specify | pending |
## Confirmed User Decisions
### D1: Prefer one-shot modular RAG refactor
User confirmed that the change can be done "一次到位" instead of only doing a compatibility-preserving internal refactor.
Implementation implication:
- Create full pipeline modules now.
- Do not leave `LookupKnowledgeTool` as a large procedural class.
### D2: L0 is not a normal evidence retrieval path
User challenged the first design because it made L0 participate in too many flows.
Confirmed direction:
```text
L0 -> query understanding / category filter / domain/entity/keyword signal
L1 -> main vector retrieval
```
Implementation implication:
- Do not model L0 and L1 as equal retrievers in the normal path.
- L0 output can influence filter, rerank, and trace.
### D3: MVP fallback is unfiltered L1 retry
User proposed a simpler MVP fallback:
```text
Filtered L1 using L0 category filter
-> if inaccurate or empty
-> retry raw query with L1 and no L0 filter
```
Confirmed direction:
- Use filtered vector retrieval first when L0 provides an unambiguous category.
- If filtered retrieval is low-quality, retry unfiltered vector retrieval with raw query.
- Do not return L0 documents as fact evidence fallback in this MVP design.
### D4: Result contract can change
User confirmed new fields can be added and old fields can be removed if the later flow becomes cleaner.
Implementation implication:
- `LookupResult.primary` and `LookupResult.supplement` may be removed.
- Preferred contract becomes `evidenceBlocks + contextPack + retrievalTrace + rerankTrace`.
- This is a breaking interface change and must be treated as L4.
## Evidence-Driven Findings
### E1: Existing spec conflict
`openspec/specs/rag-knowledge-retrieval/spec.md` currently says L1 no-result or failure should return an L0-based primary result. This conflicts with the confirmed design.
Required OpenSpec update:
- Replace L0 primary fallback with unfiltered vector retry.
- Define no-evidence behavior when both filtered and unfiltered L1 fail.
### E2: Existing compatibility-field requirement conflict
`openspec/specs/rag-knowledge-retrieval/spec.md` currently requires `primary` and `supplement` compatibility fields to remain when evidence blocks exist.
Required OpenSpec update:
- Remove compatibility-field requirement.
- Define evidence blocks and context pack as the preferred tool result contract.
### E3: Primary/supplement references are localized
Search found concrete Java references in:
- `LookupKnowledgeTool`
- `ToolInvocationRecorder`
- `LookupKnowledgeToolTest`
- `ToolInvocationRecorderTest`
- `LookupResult`
- `PrimaryResult`
- `SupplementResult`
No broad in-repo service usage was found beyond tool implementation, recorder, tests, prompts, and historical docs.
Implementation implication:
- One-shot migration is feasible if tests and prompts are updated in the same change.
### E4: Existing trace contract must be preserved
`tool_invocation` is used by diagnosis trace, verifier, and evaluation code. The database table does not need to change for this design if new details remain inside `retrieval_details`.
Implementation implication:
- Keep table-level fields stable.
- Enrich JSON `retrieval_details` with `retrieval_trace`, `rerank_trace`, `context_pack_summary`, and fallback reason.
## Interface Impact
Level: L4 breaking interface.
Reason:
- Removes or changes old result fields consumed by current tests and possibly by Agent prompt behavior.
- Changes `lookup_knowledge` tool JSON shape.
Mitigation:
- Keep tool name and input signature unchanged.
- Update all in-repo consumers in the same change.
- Keep `tool_invocation` table schema stable.
- Add tests for the new result contract.
- Update prompt text to teach Agent to use `contextPack` and `evidenceBlocks`.
## Proposed Implementation Shape
Pipeline classes:
- `KnowledgeQueryTransformer`
- `KnowledgeDocumentRetriever`
- `KnowledgeEvidencePostProcessor`
- `KnowledgeContextPacker`
- `LookupResultAssembler`
New or updated DTOs:
- `KnowledgeQuery`
- `RetrievedEvidenceCandidate`
- `EvidencePostprocessResult`
- `ContextPack`
- `RetrievalTrace`
- `RerankTrace`
- `LookupResult`
Policy:
- L0-derived filter is optional and only used when unambiguous.
- Filtered L1 low-quality result triggers raw unfiltered L1 retry.
- Rule-based rerank is sufficient for MVP.
- Context packing uses character budget first, not exact token counting.
## Pending For Commit
- Specify exact fields for `ContextPack`, `RetrievalTrace`, and `RerankTrace`.
- Specify low-quality trigger for unfiltered retry.
- Decide whether `PrimaryResult` / `SupplementResult` classes are deleted or deprecated during the first apply.
- Write OpenSpec `design.md`, specs, and executable `tasks.md`.
## Specify Results
Created committed-design artifacts:
- `design.md`
- `specs/rag-knowledge-retrieval/spec.md`
- `tasks.md`
Resolved pending items:
- `ContextPack` minimum fields: `packedText`, `strategy`, `charBudget`, `usedChars`, `includedSources`, `omittedSources`.
- `RetrievalTrace` minimum behavior: record filtered attempt, unfiltered retry when used, fallback reason, and no-evidence paths.
- `RerankTrace` minimum behavior: record final rank, source, base retrieval score when available, and major boost reasons for top evidence blocks.
- Low-quality trigger: empty candidates, empty final evidence, or top normalized similarity below `retrieval.normalization.reference-threshold`.
- `PrimaryResult` / `SupplementResult`: may be removed during apply if all compile-time usages are migrated.
## Cross-Artifact Alignment
| Check | Result |
| --- | --- |
| proposal goals/scope -> design decisions | aligned |
| design module boundaries -> specs behavior | aligned |
| specs observable behavior -> tasks | aligned |
| interface impact -> design/tasks migration work | aligned |
No cross-artifact gaps remain for Commit.
## Architecture Audit
Input -> processing -> output chain:
```text
lookup_knowledge(query)
-> KnowledgeQueryTransformer
-> KnowledgeDocumentRetriever
-> KnowledgeEvidencePostProcessor
-> KnowledgeContextPacker
-> LookupResultAssembler
-> ToolInvocationRecorder
```
Risk assessment:
- The architecture keeps the explicit Agent tool boundary and does not move retrieval into an implicit Advisor.
- Data ownership is clearer: query hints belong to transformer, vector candidates to retriever, evidence/context/traces to post-retrieval pipeline, persistence summaries to recorder.
- The largest risk is the L4 result contract change; design and tasks require prompt/test/recorder migration in the same apply.
- Database migration risk is low because `tool_invocation` table fields remain stable and new trace details stay in JSON.
- Latency risk from unfiltered retry is accepted for MVP because retry only happens below reference quality.
## Commit Gate
OpenSpec validation:
```text
openspec validate modular-rag-pipeline --strict
Change 'modular-rag-pipeline' is valid
```
File integrity:
- proposal exists and states problem, proposed change, scope, non-goals, risks, and interface impact.
- design exists and records module boundaries, decisions, migration, rollback, and risks.
- specs exist and define observable behavior for modular pipeline, unfiltered retry, evidence-first result, context pack, rerank, and L0 hint boundaries.
- tasks exist and are executable vertical slices.
Consistency:
- proposal core concepts are represented in design.
- design decisions are represented in specs and tasks.
- tasks have verifiable implementation and test steps.
- L4 interface impact is recorded and mapped to migration tasks.
Status: ready to mark as Committed OpenSpec.
## Pre-apply Research
Capability source: `openspec-apply-change` + sm-flow apply protocol. `codebase-retrieval` and LSP tools were not available in this session, so call-chain confirmation used OpenSpec context, `rg`, targeted file reads, and tests.
Reference implementation and affected files inspected:
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
- `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java`
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
- `src/main/java/com/superbiz/agent/tool/RetrievedDocTracker.java`
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
- `src/main/java/com/superbiz/agent/dto/LookupResult.java`
- `src/test/java/com/superbiz/agent/tool/LookupKnowledgeToolTest.java`
- `src/test/java/com/superbiz/agent/service/ToolInvocationRecorderTest.java`
- `src/main/resources/prompts/chat-executor-prompt.md`
- `src/main/resources/prompts/executor-prompt.md`
- `src/main/resources/application.yml`
Technical stack checklist:
- Request/response structure: `lookup_knowledge` returns a Java DTO serialized as Agent tool JSON.
- Retrieval facade: `VectorSearchService.searchSimilarDocuments(query, topK, category)` is the stable retrieval API and already hides Spring AI VectorStore vs SDK fallback.
- L0 query hints: `KnowledgeIndexService.analyzeQuery` returns `L0Hint(matches, matchedKeywords, domains, entities, titles)` and `singleDomainOrNull()`.
- Session dedup: `RetrievedDocTracker` stores `sessionId -> domain -> filePath` and exposes backward-compatible `isAlreadyRetrieved`.
- Trace persistence: `ToolInvocationRecorder.recordLookupKnowledge` writes stable table columns and JSON `retrieval_details`.
- Prompt consumers: executor prompts still describe old L0/L1 and `primary.content`; these must be migrated.
- Tests: `LookupKnowledgeToolTest` is the main consumer of old `primary` / `supplement` assertions; recorder tests verify retrieval details.
Implementation decision:
- Add pipeline DTOs under `com.superbiz.agent.dto`.
- Add pipeline services under `com.superbiz.agent.service`.
- Keep `VectorSearchService` unchanged.
- Keep `lookup_knowledge` tool name and query argument unchanged.
- Delete `PrimaryResult` / `SupplementResult` only after production and tests stop referencing them.
## Apply Results
Completed tasks: 31/31.
Implemented:
- Added modular RAG DTOs: `KnowledgeQuery`, `RetrievedEvidenceCandidate`, `ContextPack`, `RetrievalTrace`, `RerankTrace`, `EvidencePostprocessResult`.
- Added pipeline services: `KnowledgeQueryTransformer`, `KnowledgeDocumentRetriever`, `KnowledgeEvidencePostProcessor`, `KnowledgeContextPacker`, `LookupResultAssembler`.
- Refactored `LookupKnowledgeTool` into a thin orchestrator.
- Migrated `LookupResult` to evidence-first fields and removed `primary` / `supplement`.
- Deleted `PrimaryResult` and `SupplementResult`.
- Updated `ToolInvocationRecorder` to persist query transform, retrieval trace, context pack summary, rerank trace, fallback reason, and evidence summaries in `retrieval_details`.
- Updated executor prompt text for `contextPack` / `evidenceBlocks`.
- Updated RAG architecture docs.
- Rewrote lookup and recorder tests for the new contract.
Verification:
```text
mvn -q -DskipTests compile
PASS
mvn -q "-Dtest=LookupKnowledgeToolTest,ToolInvocationRecorderTest" test
PASS
openspec validate modular-rag-pipeline --strict
PASS: Change 'modular-rag-pipeline' is valid
```
Full suite attempt:
```text
mvn -q test
FAIL
```
The full suite failed on pre-existing/environment-dependent tests:
- `MilvusConnectionTest.connect`: `MILVUS_TOKEN` not set.
- `RedisSessionManagerTest`: Redis JSON contains legacy `messagePairCount`, not accepted by current `SessionContext`.
- Spring context / repository tests attempted MySQL/Flyway and failed when database connectivity was unavailable in the first sandboxed run.
The full suite was retried outside the sandbox after approval. It still failed for the Milvus token and Redis serialization issues above, so these failures are not attributed to the modular RAG change.
Diff scope reviewed:
- RAG implementation: DTOs, pipeline services, `LookupKnowledgeTool`, `ToolInvocationRecorder`.
- Contract cleanup: removed `PrimaryResult` / `SupplementResult`, updated `LookupResult`.
- Tests: lookup tool and recorder tests.
- Prompts: executor and chat executor guidance.
- Docs/OpenSpec: modular RAG change files and architecture docs.
Known remaining risk:
- Runtime Agent prompt behavior should be demo-tested manually because the tool JSON contract changed from `primary/supplement` to `evidenceBlocks/contextPack/traces`.
## Review Results
Review finding:
- Session dedup returned `found=false` but still carried `evidenceBlocks` and `contextPack`, allowing the Agent to consume duplicate evidence despite the dedup message.
Fix:
- `LookupResultAssembler.deduped` now returns an empty evidence/context payload while preserving trace, relevance hint, and retrieved-domain memory.
- Added `LookupKnowledgeToolTest.sessionDedupDoesNotReturnConsumableEvidenceAgain`.
Additional verification:
```text
mvn -q -DskipTests compile
PASS
mvn -q "-Dtest=LookupKnowledgeToolTest,ToolInvocationRecorderTest" test
PASS
MILVUS_TOKEN=<application.yml milvus.token> mvn -q test
PASS
openspec validate modular-rag-pipeline --strict
PASS: Change 'modular-rag-pipeline' is valid
git diff --check
PASS with LF/CRLF warnings only
```
Updated full-suite note:
- `MilvusConnectionTest` passes when `MILVUS_TOKEN` is injected into the Maven process from `application.yml`.
- `RedisSessionManagerTest` now passes after marking computed `SessionContext` getters as non-serialized JSON properties and ignoring unknown legacy Redis fields.
@@ -0,0 +1,193 @@
## Context
The current `lookup_knowledge` implementation is operational but still organized around a legacy L0/L1 result model. `LookupKnowledgeTool` currently performs query analysis, vector retrieval, relevance normalization, evidence block construction, session deduplication, and trace recording in one class. Evidence blocks already exist, but post-retrieval processing is not a first-class pipeline boundary.
The project direction is already documented as:
- keep `lookup_knowledge` as an explicit Agent evidence tool;
- treat L0 as domain/entity/keyword hint data;
- use L1 vector retrieval as the main evidence source;
- keep Spring AI `VectorStore` behind `VectorSearchService`;
- preserve `tool_invocation` as the evidence trace for Verifier, Eval, and trace APIs.
This change turns that architecture into code structure and updates the tool result contract so future Agent, Verifier, and Eval flows consume structured evidence instead of the old `primary` / `supplement` split.
## Goals / Non-Goals
**Goals:**
- Refactor `lookup_knowledge` into a modular RAG pipeline.
- Keep L0 as query understanding and retrieval-control signal.
- Keep L1 vector retrieval as the main document retrieval path.
- Add a simple MVP fallback: retry raw query through unfiltered L1 when filtered L1 is low quality.
- Move relevance normalization, evidence block creation, deduplication, and lightweight rerank into post-retrieval processing.
- Add context packing as a first-class output.
- Replace the old `primary` / `supplement` result contract with `evidenceBlocks`, `contextPack`, `retrievalTrace`, and `rerankTrace`.
- Keep `lookup_knowledge` tool name and input argument unchanged.
- Keep `tool_invocation` table schema stable and enrich `retrieval_details`.
**Non-Goals:**
- Do not replace `lookup_knowledge` with an implicit Advisor.
- Do not add a model-based reranker, cross-encoder, BM25, RRF, Elasticsearch, or OpenSearch.
- Do not migrate document upload, chunking, embedding write path, Milvus schema, or vector collection layout.
- Do not change when the Agent decides to call `lookup_knowledge`.
- Do not introduce multi-query expansion in this change.
## Decisions
### Decision: Use explicit pipeline components
Create a local pipeline behind `LookupKnowledgeTool`:
```text
LookupKnowledgeTool
-> KnowledgeQueryTransformer
-> KnowledgeDocumentRetriever
-> KnowledgeEvidencePostProcessor
-> KnowledgeContextPacker
-> LookupResultAssembler
-> ToolInvocationRecorder
```
Rationale: this keeps the Agent tool boundary stable while making the RAG flow easy to test and explain. It also maps cleanly to Spring AI modular RAG concepts without hiding business observability inside an Advisor.
Alternative considered: keep all logic in `LookupKnowledgeTool` and only add fields. Rejected because the class would continue to mix query transformation, retrieval, post-processing, and trace responsibilities.
### Decision: L0 is a query transformer signal, not a main retriever
`KnowledgeQueryTransformer` will wrap the existing `KnowledgeIndexService.analyzeQuery` behavior and produce a transformed query object containing:
- original query
- rewritten query, initially equal to the raw query unless a future rule rewrites it
- domain hints
- matched keywords
- entities
- optional category filter
- L0 titles for trace only
L0 matches must not be converted into normal evidence candidates in the main path.
Rationale: L0 keyword/frontmatter matching is useful for controlling retrieval, but it is not reliable enough to be treated as fact evidence when L1 cannot support it.
Alternative considered: combine L0 and L1 into one candidate list. Rejected because it makes L0 an equal retrieval layer again and conflicts with the desired architecture.
### Decision: Use filtered L1 first, then unfiltered L1 retry
`KnowledgeDocumentRetriever` will perform:
```text
attempt 1: vector search with L0-derived category filter, when unambiguous
attempt 2: raw query vector search without the L0-derived filter, when attempt 1 is low quality
```
Filtered retrieval is low quality when any of the following is true:
- no candidates are returned;
- post-processing would produce zero evidence blocks;
- top candidate normalized similarity is below `retrieval.normalization.reference-threshold`.
Rationale: the most likely MVP failure mode is an over-strict or wrong metadata filter. A raw unfiltered vector retry addresses that without adding a complex multi-stage fallback policy.
Alternative considered: return L0 documents as weak fallback evidence. Rejected for this MVP because it can let keyword hints masquerade as factual evidence.
### Decision: Keep rerank rule-based
`KnowledgeEvidencePostProcessor` will rerank with deterministic signals:
- vector score / normalized similarity;
- domain match with query hints;
- entity match;
- keyword match;
- source type or metadata priority when available;
- retrieval attempt, with filtered hits not automatically preferred over stronger unfiltered hits.
The rerank trace should explain major score contributions per final evidence block.
Rationale: a rule-based reranker is explainable, cheap, testable, and enough for the interview-oriented MVP. It also avoids introducing model latency and new dependencies.
Alternative considered: model-based rerank. Rejected as out of scope until evaluation shows a need.
### Decision: Context pack becomes the Agent-facing content
`KnowledgeContextPacker` will turn final evidence blocks into a compact context package:
```text
packedText
strategy
charBudget
usedChars
includedSources
omittedSources
```
MVP packing uses a character budget rather than exact token counting. The packer preserves source/title/breadcrumb/hit reasons and truncates content only after preserving metadata.
Rationale: the Agent should consume curated evidence context instead of inferring semantics from `primary` and `supplement`.
Alternative considered: keep `primary` and `supplement` as the Agent-facing fields. Rejected because those names preserve the outdated L0 primary / L1 supplement model.
### Decision: Break the old result contract deliberately
`LookupResult` will move to the preferred contract:
```text
found
evidenceBlocks
contextPack
retrievalTrace
rerankTrace
relevanceLevel
completenessHint
retrievedDomainsThisSession
message
```
`primary` and `supplement` may be removed in this change, and `PrimaryResult` / `SupplementResult` may be deleted if no production code remains after migration.
Rationale: this is an MVP project intended to demonstrate clean modular RAG. Keeping obsolete fields would force later code to preserve misleading semantics.
Alternative considered: add new fields while keeping old fields deprecated. Rejected because the user explicitly accepted removing old fields if the later flow becomes cleaner.
### Decision: Preserve persistence compatibility at table level
`ToolInvocationRecorder` will stop depending on `result.getPrimary()` for output preview. Preview should come from `contextPack.packedText` or the top evidence block content. `retrieval_details` will include compact summaries:
- query transform summary;
- retrieval trace with attempts and fallback reason;
- rerank trace summary;
- context pack summary;
- evidence block summaries.
Rationale: trace consumers already read `tool_invocation` rows. The database schema can remain stable while JSON details evolve.
Alternative considered: add columns for each new trace object. Rejected because the current trace model already stores retrieval-specific details in JSON and does not need schema churn for this change.
## Risks / Trade-offs
- [Risk] Breaking tool output shape can affect Agent prompt behavior. -> Mitigation: update tool description and executor prompt to prefer `contextPack` and `evidenceBlocks`.
- [Risk] Existing tests assert `primary` / `supplement`. -> Mitigation: migrate tests to evidence/context/traces in the same change.
- [Risk] Rule-based rerank can reorder evidence unexpectedly. -> Mitigation: persist `rerankTrace` and cover ordering behavior with tests.
- [Risk] Context packing can omit useful evidence under a small budget. -> Mitigation: record included and omitted sources and keep the character budget configurable.
- [Risk] Existing OpenSpec requirements conflict with the new fallback policy. -> Mitigation: update `rag-knowledge-retrieval` delta before apply and validate the change.
- [Risk] Unfiltered retry may increase latency. -> Mitigation: retry only when filtered result is empty or below reference quality, and record attempt counts/duration in trace.
## Migration Plan
1. Add new internal DTOs and pipeline components while keeping `LookupKnowledgeTool` as the public tool bean.
2. Migrate `LookupKnowledgeTool` orchestration to the pipeline.
3. Update `LookupResult` to the new result contract and remove `primary` / `supplement` usages.
4. Update `ToolInvocationRecorder` preview and retrieval details to use context/evidence/traces.
5. Update prompt text and tool description.
6. Update tests for the new contract.
7. Run targeted unit tests and OpenSpec validation.
Rollback strategy:
- Revert the change as one unit if Agent tool behavior regresses.
- Because the table schema remains stable and the tool input is unchanged, rollback does not require database migration.
## Open Questions
- The exact default context pack character budget should be finalized during implementation; recommended MVP default is 3000 to 5000 characters.
- Whether to delete `PrimaryResult` and `SupplementResult` immediately depends on final compile-time references after migration.
@@ -0,0 +1,172 @@
# Modular RAG Pipeline Proposal
## Problem
`lookup_knowledge` already exposes structured retrieval evidence, but the runtime flow is still concentrated inside `LookupKnowledgeTool`. Query understanding, vector retrieval, relevance normalization, evidence block construction, session deduplication, and trace recording are tightly coupled. This makes the RAG path harder to explain, test, evolve, and present as a modular Agent engineering design.
The current result contract also still carries the old `primary` / `supplement` model, where `primary` means L0 exact match and `supplement` means L1 semantic retrieval. That contract no longer matches the intended architecture: L0 should be a query understanding and retrieval-control signal, while L1 vector retrieval should be the main evidence source.
## Proposed Change
Refactor `lookup_knowledge` into a modular RAG pipeline while preserving the explicit Agent tool boundary and `tool_invocation` evidence trace.
Target pipeline:
```text
LookupKnowledgeTool
-> KnowledgeQueryTransformer
-> KnowledgeDocumentRetriever
-> KnowledgeEvidencePostProcessor
-> KnowledgeContextPacker
-> LookupResultAssembler
-> ToolInvocationRecorder
```
### Query Transformation
Introduce a query transformer around the current L0 analysis.
L0 SHALL provide:
- domain/category hints
- matched keywords
- extracted entities
- optional metadata filter candidate
- traceable query understanding data
L0 SHALL NOT act as a main evidence retrieval path in the normal flow.
### Retrieval
L1 vector retrieval remains the main document retrieval path through `VectorSearchService`, which already supports Spring AI `VectorStore` as the preferred path and Milvus SDK fallback.
MVP fallback strategy:
```text
1. Run filtered vector retrieval with the L0-derived category filter when unambiguous.
2. If filtered retrieval returns no usable evidence or low-quality evidence, retry raw query through unfiltered vector retrieval.
3. If unfiltered retrieval also fails, return no_evidence.
```
The fallback SHALL skip the L0-derived filter rather than returning L0 documents as fact evidence.
### Post-Retrieval Processing
Move evidence post-processing out of `LookupKnowledgeTool`.
The post-processor SHALL handle:
- relevance normalization
- evidence block creation
- source-level deduplication
- rule-based lightweight rerank
- retrieval trace and rerank trace generation
The first rerank implementation should be rule-based, using available signals such as vector score, domain match, entity match, keyword match, source type, and whether evidence aligns with query hints.
### Context Packing
Add a context packing step that converts final evidence blocks into an Agent-facing context package.
The packer SHALL:
- keep source/title/breadcrumb visible
- obey a configurable character budget in the MVP
- prioritize reranked evidence order
- avoid duplicate source content
- produce a compact summary of included and omitted evidence
### Result Contract
This change intentionally updates the `lookup_knowledge` return contract.
New preferred contract:
```text
found
evidenceBlocks
contextPack
retrievalTrace
rerankTrace
relevanceLevel
completenessHint
retrievedDomainsThisSession
message
```
The old `primary` and `supplement` fields may be removed as part of this change, because they encode the outdated assumption that L0 is the primary evidence source and L1 is supplemental evidence.
## Scope
In scope:
- Refactor `LookupKnowledgeTool` into a thin tool boundary and pipeline orchestrator.
- Add local pipeline classes and DTOs for query transformation, retrieval result normalization, post-processing, context packing, and traces.
- Update `LookupResult` to prefer `evidenceBlocks`, `contextPack`, `retrievalTrace`, and `rerankTrace`.
- Remove or deprecate `primary` / `supplement` according to the final spec.
- Update `ToolInvocationRecorder` to persist compact summaries for evidence blocks, context pack, retrieval trace, fallback reason, and rerank trace.
- Update tool description / prompt wording so Agent behavior matches the new contract.
- Update tests for filtered retrieval, unfiltered retry, rerank, context packing, evidence persistence, and result contract changes.
- Update `rag-knowledge-retrieval` spec to remove L0 primary fallback and old compatibility-field requirements.
Out of scope:
- Replacing the explicit `lookup_knowledge` tool with an implicit Advisor.
- Introducing a model-based reranker or cross-encoder.
- Introducing BM25, RRF, Elasticsearch, or OpenSearch.
- Migrating document upload, chunking, embedding writes, or Milvus schema.
- Changing the Agent decision of when to call `lookup_knowledge`.
## Interface Impact
Impact level: L4 breaking interface change.
Reason:
- `LookupResult.primary` and `LookupResult.supplement` may be removed.
- The JSON returned by the `lookup_knowledge` Agent tool changes shape.
- Tests and internal consumers that read `primary` / `supplement` must migrate to `evidenceBlocks` and `contextPack`.
Known affected areas:
- `LookupKnowledgeTool`
- `LookupResult`
- `PrimaryResult` / `SupplementResult`
- `ToolInvocationRecorder`
- `LookupKnowledgeToolTest`
- `ToolInvocationRecorderTest`
- Agent tool prompt / executor prompt references
- `rag-knowledge-retrieval` OpenSpec requirements
- RAG docs under `mvp/architecture/`
Migration approach:
- Update all in-repo consumers in the same change.
- Keep `lookup_knowledge` tool name and input argument unchanged.
- Keep `tool_invocation` persistence compatible at table level while enriching `retrieval_details`.
- Record fallback and no-evidence semantics explicitly so Verifier and Eval do not treat hint-only data as fact evidence.
## Context Constraints
Relevant project decisions:
- `lookup_knowledge` remains an explicit Agent evidence tool.
- L0 is already documented as a hint layer, not a final decision layer.
- Spring AI `VectorStore` is already the preferred retrieval path behind `VectorSearchService`.
- Milvus SDK fallback remains valuable for MVP resilience.
- `tool_invocation` is the stable evidence trace used by trace inspection, Verifier, and Eval.
- Existing `rag-knowledge-retrieval` spec still contains legacy fallback and compatibility-field requirements that must be changed.
## Risks
- Breaking result contract may affect prompt behavior because the tool JSON changes.
- Removing `primary` / `supplement` requires updating tests and recorder preview logic.
- Rule-based rerank may create unexpected ordering changes if score semantics are not handled carefully.
- Context packing can hide useful evidence if the budget is too small.
- Existing OpenSpec requirements conflict with the new L0 fallback policy and must be updated before implementation.
## Open Questions
- What exact minimum fields should `ContextPack`, `RetrievalTrace`, and `RerankTrace` expose in the committed spec?
- Should `PrimaryResult` and `SupplementResult` classes be deleted immediately or left deprecated for one change cycle?
- What threshold defines "filtered retrieval low quality" for triggering unfiltered retry?
@@ -0,0 +1,121 @@
## ADDED Requirements
### Requirement: Knowledge retrieval SHALL use a modular RAG pipeline
The `lookup_knowledge` tool SHALL route each request through explicit query transformation, vector retrieval, post-retrieval processing, context packing, result assembly, and trace recording components.
#### Scenario: Pipeline components execute in order
- **WHEN** `lookup_knowledge` receives a query
- **THEN** the system SHALL transform the query before retrieval
- **AND** it SHALL retrieve vector candidates before post-processing
- **AND** it SHALL build evidence blocks before context packing
- **AND** it SHALL record trace details after result assembly
#### Scenario: Tool boundary remains explicit
- **WHEN** the modular pipeline is used
- **THEN** the Agent SHALL still call the explicit `lookup_knowledge` tool with the same query argument
- **AND** the implementation SHALL NOT require an implicit Advisor to inject knowledge into every chat response
### Requirement: Knowledge retrieval SHALL retry without L0 filter when filtered L1 is low quality
The retrieval flow SHALL treat L0-derived category filtering as an optimization, not as a hard dependency for final recall.
#### Scenario: Filtered retrieval succeeds
- **WHEN** L0 provides an unambiguous category filter
- **AND** filtered L1 retrieval returns usable evidence at or above the configured reference threshold
- **THEN** the tool SHALL use the filtered L1 candidates without running an unfiltered retry
#### Scenario: Filtered retrieval returns no evidence
- **WHEN** L0 provides a category filter
- **AND** filtered L1 retrieval returns no candidates or no final evidence blocks
- **THEN** the tool SHALL retry L1 retrieval with the raw query and no L0-derived category filter
- **AND** the retrieval trace SHALL record fallback reason `filtered_vector_no_evidence`
#### Scenario: Filtered retrieval is below reference quality
- **WHEN** L0 provides a category filter
- **AND** filtered L1 retrieval returns candidates whose top normalized similarity is below the configured reference threshold
- **THEN** the tool SHALL retry L1 retrieval with the raw query and no L0-derived category filter
- **AND** the retrieval trace SHALL record fallback reason `filtered_vector_low_quality`
#### Scenario: Both retrieval attempts fail
- **WHEN** filtered L1 retrieval and unfiltered L1 retry both produce no usable evidence
- **THEN** the tool SHALL return `found=false`
- **AND** the tool SHALL set evidence status to `no_evidence`
- **AND** the tool SHALL NOT return L0 documents as fact evidence
### Requirement: Knowledge retrieval SHALL return an evidence-first result contract
The `lookup_knowledge` result SHALL expose structured evidence and packed context as the preferred contract.
#### Scenario: Evidence result contains context and traces
- **WHEN** `lookup_knowledge` returns usable evidence
- **THEN** the result SHALL include `evidenceBlocks`
- **AND** it SHALL include `contextPack`
- **AND** it SHALL include `retrievalTrace`
- **AND** it SHALL include `rerankTrace`
- **AND** it SHALL include `relevanceLevel` and `completenessHint`
#### Scenario: No-evidence result keeps traceability
- **WHEN** `lookup_knowledge` returns no usable evidence
- **THEN** the result SHALL include `found=false`
- **AND** it SHALL include a message explaining that no knowledge evidence was found
- **AND** it SHALL include retrieval trace details for attempted retrieval paths
### Requirement: Knowledge retrieval SHALL pack evidence context for Agent consumption
The post-retrieval flow SHALL convert final evidence blocks into a compact context package for the Agent.
#### Scenario: Context pack preserves source metadata
- **WHEN** evidence blocks are packed
- **THEN** the packed context SHALL preserve source, title when available, breadcrumb when available, and hit reasons for included evidence
#### Scenario: Context pack respects budget
- **WHEN** final evidence content exceeds the configured context budget
- **THEN** the packer SHALL truncate content rather than source metadata
- **AND** it SHALL record included and omitted sources in the context pack summary
### Requirement: Knowledge retrieval SHALL rerank evidence with traceable rule signals
The post-retrieval flow SHALL rerank vector candidates using deterministic rule-based signals and expose the explanation.
#### Scenario: Rerank trace records score contributions
- **WHEN** candidates are reranked
- **THEN** the rerank trace SHALL record final rank, source, base retrieval score when available, and major boost reasons for top evidence blocks
#### Scenario: Query hints influence rerank without becoming evidence
- **WHEN** L0 query hints match candidate metadata or content
- **THEN** the reranker MAY boost the candidate
- **AND** the evidence block SHALL record the hint as a hit reason
- **AND** the system SHALL NOT treat the L0 hint itself as fact evidence
## MODIFIED Requirements
### Requirement: Knowledge retrieval SHALL keep L0 as a hint provider
The `lookup_knowledge` retrieval flow SHALL retain L0 keyword/frontmatter matching but use it only as query understanding, filtering, rerank, and explainability hint data rather than as a final evidence retrieval decision.
#### Scenario: L0 produces traceable hint data
- **WHEN** L0 matches one or more indexed knowledge entries
- **THEN** the retrieval flow SHALL expose matched titles, matched keywords, domains or categories, and entity terms as structured hint data
#### Scenario: L0 does not bypass semantic retrieval
- **WHEN** L0 returns exactly one match
- **THEN** the retrieval flow SHALL still attempt semantic L1 retrieval unless L1 is explicitly disabled by configuration
#### Scenario: L0 hints do not become normal evidence
- **WHEN** L1 retrieval returns no usable evidence
- **THEN** L0 matched documents SHALL NOT be returned as fact evidence blocks
- **AND** L0 hint data MAY still be recorded in retrieval trace details
### Requirement: Knowledge retrieval SHALL return structured evidence blocks
The `lookup_knowledge` retrieval flow SHALL expose retrieved evidence as structured evidence blocks.
#### Scenario: Evidence block contains source metadata
- **WHEN** a `lookup_knowledge` call returns evidence
- **THEN** each evidence block SHALL include source, title when available, breadcrumb when available, retrieval layer, content, and hit reasons
#### Scenario: Evidence blocks are the primary evidence contract
- **WHEN** evidence blocks are returned
- **THEN** Agent-facing knowledge content SHALL be derived from evidence blocks and context pack
- **AND** the result SHALL NOT rely on legacy `primary` or `supplement` fields for L0/L1 meaning
## REMOVED Requirements
### Requirement: Knowledge retrieval SHALL preserve fallback evidence
**Reason**: The new MVP fallback strategy retries unfiltered L1 retrieval when L0-derived filtering appears to suppress usable semantic evidence. Returning L0 documents as fallback fact evidence can let keyword hints masquerade as verified knowledge.
**Migration**: Use the new requirement "Knowledge retrieval SHALL retry without L0 filter when filtered L1 is low quality". If filtered and unfiltered L1 both fail, return `no_evidence` while preserving L0 hint data in trace details only.
@@ -0,0 +1,51 @@
## 1. Pipeline Models
- [x] 1.1 Add query transformation model for original query, rewritten query, domain hints, matched keywords, entities, category filter, and trace-only L0 titles.
- [x] 1.2 Add retrieved evidence candidate model that normalizes vector result metadata, score semantics, source, title, breadcrumb, content, retrieval attempt, and original rank.
- [x] 1.3 Add context pack, retrieval trace, and rerank trace DTOs for the new `LookupResult` contract.
- [x] 1.4 Update `LookupResult` to expose evidence blocks, context pack, retrieval trace, rerank trace, relevance level, completeness hint, retrieved domains, and message.
## 2. Query And Retrieval Pipeline
- [x] 2.1 Implement `KnowledgeQueryTransformer` by reusing `KnowledgeIndexService.analyzeQuery` and mapping L0 output to query hints and optional category filter.
- [x] 2.2 Implement `KnowledgeDocumentRetriever` as a wrapper around `VectorSearchService` for filtered and unfiltered vector retrieval attempts.
- [x] 2.3 Implement low-quality detection using empty candidates, empty final evidence, or top normalized similarity below `retrieval.normalization.reference-threshold`.
- [x] 2.4 Implement unfiltered raw-query retry when filtered retrieval is low quality and record fallback reason in retrieval trace.
## 3. Post-Retrieval Processing
- [x] 3.1 Move relevance normalization out of `LookupKnowledgeTool` into `KnowledgeEvidencePostProcessor`.
- [x] 3.2 Move evidence block creation and source-level deduplication out of `LookupKnowledgeTool` into the post-processor.
- [x] 3.3 Implement rule-based lightweight rerank using vector similarity, domain match, entity match, keyword match, and metadata/source-type signals.
- [x] 3.4 Ensure L0 hints influence filter/rerank/trace only and are not returned as standalone fact evidence when L1 has no usable evidence.
## 4. Context Packing And Result Assembly
- [x] 4.1 Implement `KnowledgeContextPacker` with a configurable MVP character budget.
- [x] 4.2 Pack evidence blocks while preserving source, title, breadcrumb, and hit reasons before truncating content.
- [x] 4.3 Implement `LookupResultAssembler` to build evidence-first results for usable evidence, no-evidence, and session dedup cases.
- [x] 4.4 Remove or migrate all `primary` and `supplement` result usage from production code.
## 5. Tool Boundary And Trace Recording
- [x] 5.1 Refactor `LookupKnowledgeTool` into a thin orchestrator that invokes the pipeline and handles tool boundary concerns.
- [x] 5.2 Update `ToolInvocationRecorder.LookupKnowledgeRecord` to summarize context pack, retrieval trace, rerank trace, fallback reason, and evidence blocks without relying on `result.getPrimary()`.
- [x] 5.3 Preserve stable `tool_invocation` table fields and store new retrieval details in JSON.
- [x] 5.4 Update `@Tool` description and relevant executor prompt text to describe evidence blocks, context pack, and no-realtime-data boundaries.
## 6. Tests And Evaluation
- [x] 6.1 Update `LookupKnowledgeToolTest` for evidence-first result contract and removal of `primary` / `supplement`.
- [x] 6.2 Add tests for filtered L1 success without retry.
- [x] 6.3 Add tests for filtered low-quality retrieval triggering raw unfiltered L1 retry.
- [x] 6.4 Add tests proving L0 hint data does not become standalone fact evidence when L1 fails.
- [x] 6.5 Add tests for rerank ordering, rerank trace, context pack budget behavior, and source metadata preservation.
- [x] 6.6 Update `ToolInvocationRecorderTest` for new retrieval detail summaries and output preview source.
- [x] 6.7 Run targeted Java tests for lookup knowledge and recorder changes.
- [x] 6.8 Run OpenSpec validation for `modular-rag-pipeline`.
## 7. Documentation Cleanup
- [x] 7.1 Update RAG architecture docs to reflect modular pipeline, unfiltered vector retry, and evidence-first contract.
- [x] 7.2 Update retrieval observability docs to remove L0 primary fallback and `primary` / `supplement` compatibility language.
- [x] 7.3 Review git diff to confirm only expected RAG, prompt, test, and spec files changed.
+95 -19
View File
@@ -4,15 +4,20 @@
Define the runtime contract for the explicit `lookup_knowledge` Agent tool, including how L0 keyword/frontmatter hints cooperate with L1 semantic retrieval while preserving metadata filters, fallback evidence, and traceable retrieval details.
## Requirements
### Requirement: Knowledge retrieval SHALL keep L0 as a hint provider
The `lookup_knowledge` retrieval flow SHALL retain L0 keyword/frontmatter matching but use it as domain, entity, and explainability hint data rather than as the sole final retrieval decision.
The `lookup_knowledge` retrieval flow SHALL retain L0 keyword/frontmatter matching but use it only as query understanding, filtering, rerank, and explainability hint data rather than as a final evidence retrieval decision.
#### Scenario: L0 produces traceable hint data
- **WHEN** L0 matches one or more indexed knowledge entries
- **THEN** the retrieval flow SHALL expose matched titles, matched keywords, domains or categories, and entity terms as structured hint data
#### Scenario: L0 does not bypass semantic retrieval by default
#### Scenario: L0 does not bypass semantic retrieval
- **WHEN** L0 returns exactly one match
- **THEN** the retrieval flow SHALL still attempt semantic L1 retrieval unless L1 is unavailable or explicitly disabled by configuration
- **THEN** the retrieval flow SHALL still attempt semantic L1 retrieval unless L1 is explicitly disabled by configuration
#### Scenario: L0 hints do not become normal evidence
- **WHEN** L1 retrieval returns no usable evidence
- **THEN** L0 matched documents SHALL NOT be returned as fact evidence blocks
- **AND** L0 hint data MAY still be recorded in retrieval trace details
### Requirement: Knowledge retrieval SHALL use L0 domain as optional L1 filter
The retrieval flow SHALL use L0 domain/category information as an optional metadata filter for L1 retrieval when the domain is unambiguous.
@@ -25,19 +30,6 @@ The retrieval flow SHALL use L0 domain/category information as an optional metad
- **WHEN** L0 hint data contains zero domains or multiple domains
- **THEN** the L1 retrieval request SHALL run without an L0-derived category filter
### Requirement: Knowledge retrieval SHALL preserve fallback evidence
The retrieval flow SHALL still return useful L0 evidence when L1 produces no usable result.
#### Scenario: L1 has no results
- **WHEN** L0 has at least one match and L1 returns no candidates
- **THEN** the tool SHALL return an L0-based primary result
- **AND** the relevance assessment SHALL not claim semantic support from L1
#### Scenario: L1 fails
- **WHEN** L0 has at least one match and L1 retrieval throws or fails
- **THEN** the tool SHALL return an L0-based primary result
- **AND** the tool invocation record SHALL preserve the L0 hint details
### Requirement: Knowledge retrieval SHALL persist L0 hints
The system SHALL persist L0 hint details in `tool_invocation.retrieval_details` for `lookup_knowledge` calls.
@@ -50,15 +42,16 @@ The system SHALL persist L0 hint details in `tool_invocation.retrieval_details`
- **THEN** the recorded retrieval layer SHALL be `L0+L1`
### Requirement: Knowledge retrieval SHALL return structured evidence blocks
The `lookup_knowledge` retrieval flow SHALL expose retrieved evidence as structured evidence blocks in addition to the existing compatibility fields.
The `lookup_knowledge` retrieval flow SHALL expose retrieved evidence as structured evidence blocks.
#### Scenario: Evidence block contains source metadata
- **WHEN** a `lookup_knowledge` call returns evidence
- **THEN** each evidence block SHALL include source, title when available, breadcrumb when available, retrieval layer, content, and hit reasons
#### Scenario: Compatibility fields remain available
#### Scenario: Evidence blocks are the primary evidence contract
- **WHEN** evidence blocks are returned
- **THEN** the existing `primary` and `supplement` result fields SHALL remain available when their source evidence exists
- **THEN** Agent-facing knowledge content SHALL be derived from evidence blocks and context pack
- **AND** the result SHALL NOT rely on legacy `primary` or `supplement` fields for L0/L1 meaning
### Requirement: Knowledge retrieval SHALL deduplicate evidence blocks
The retrieval flow SHALL remove duplicate evidence blocks before returning them to the Agent.
@@ -151,3 +144,86 @@ Spring AI Milvus integration SHALL be configured to use the existing collection
- **WHEN** Spring AI Milvus VectorStore is configured
- **THEN** it SHALL use the existing id, content, vector, and metadata field names
- **AND** it SHALL use the configured embedding dimension and metric type compatible with existing vectors
### Requirement: Knowledge retrieval SHALL use a modular RAG pipeline
The `lookup_knowledge` tool SHALL route each request through explicit query transformation, vector retrieval, post-retrieval processing, context packing, result assembly, and trace recording components.
#### Scenario: Pipeline components execute in order
- **WHEN** `lookup_knowledge` receives a query
- **THEN** the system SHALL transform the query before retrieval
- **AND** it SHALL retrieve vector candidates before post-processing
- **AND** it SHALL build evidence blocks before context packing
- **AND** it SHALL record trace details after result assembly
#### Scenario: Tool boundary remains explicit
- **WHEN** the modular pipeline is used
- **THEN** the Agent SHALL still call the explicit `lookup_knowledge` tool with the same query argument
- **AND** the implementation SHALL NOT require an implicit Advisor to inject knowledge into every chat response
### Requirement: Knowledge retrieval SHALL retry without L0 filter when filtered L1 is low quality
The retrieval flow SHALL treat L0-derived category filtering as an optimization, not as a hard dependency for final recall.
#### Scenario: Filtered retrieval succeeds
- **WHEN** L0 provides an unambiguous category filter
- **AND** filtered L1 retrieval returns usable evidence at or above the configured reference threshold
- **THEN** the tool SHALL use the filtered L1 candidates without running an unfiltered retry
#### Scenario: Filtered retrieval returns no evidence
- **WHEN** L0 provides a category filter
- **AND** filtered L1 retrieval returns no candidates or no final evidence blocks
- **THEN** the tool SHALL retry L1 retrieval with the raw query and no L0-derived category filter
- **AND** the retrieval trace SHALL record fallback reason `filtered_vector_no_evidence`
#### Scenario: Filtered retrieval is below reference quality
- **WHEN** L0 provides a category filter
- **AND** filtered L1 retrieval returns candidates whose top normalized similarity is below the configured reference threshold
- **THEN** the tool SHALL retry L1 retrieval with the raw query and no L0-derived category filter
- **AND** the retrieval trace SHALL record fallback reason `filtered_vector_low_quality`
#### Scenario: Both retrieval attempts fail
- **WHEN** filtered L1 retrieval and unfiltered L1 retry both produce no usable evidence
- **THEN** the tool SHALL return `found=false`
- **AND** the tool SHALL set evidence status to `no_evidence`
- **AND** the tool SHALL NOT return L0 documents as fact evidence
### Requirement: Knowledge retrieval SHALL return an evidence-first result contract
The `lookup_knowledge` result SHALL expose structured evidence and packed context as the preferred contract.
#### Scenario: Evidence result contains context and traces
- **WHEN** `lookup_knowledge` returns usable evidence
- **THEN** the result SHALL include `evidenceBlocks`
- **AND** it SHALL include `contextPack`
- **AND** it SHALL include `retrievalTrace`
- **AND** it SHALL include `rerankTrace`
- **AND** it SHALL include `relevanceLevel` and `completenessHint`
#### Scenario: No-evidence result keeps traceability
- **WHEN** `lookup_knowledge` returns no usable evidence
- **THEN** the result SHALL include `found=false`
- **AND** it SHALL include a message explaining that no knowledge evidence was found
- **AND** it SHALL include retrieval trace details for attempted retrieval paths
### Requirement: Knowledge retrieval SHALL pack evidence context for Agent consumption
The post-retrieval flow SHALL convert final evidence blocks into a compact context package for the Agent.
#### Scenario: Context pack preserves source metadata
- **WHEN** evidence blocks are packed
- **THEN** the packed context SHALL preserve source, title when available, breadcrumb when available, and hit reasons for included evidence
#### Scenario: Context pack respects budget
- **WHEN** final evidence content exceeds the configured context budget
- **THEN** the packer SHALL truncate content rather than source metadata
- **AND** it SHALL record included and omitted sources in the context pack summary
### Requirement: Knowledge retrieval SHALL rerank evidence with traceable rule signals
The post-retrieval flow SHALL rerank vector candidates using deterministic rule-based signals and expose the explanation.
#### Scenario: Rerank trace records score contributions
- **WHEN** candidates are reranked
- **THEN** the rerank trace SHALL record final rank, source, base retrieval score when available, and major boost reasons for top evidence blocks
#### Scenario: Query hints influence rerank without becoming evidence
- **WHEN** L0 query hints match candidate metadata or content
- **THEN** the reranker MAY boost the candidate
- **AND** the evidence block SHALL record the hint as a hit reason
- **AND** the system SHALL NOT treat the L0 hint itself as fact evidence