Compare commits
7
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
64adb998cf | ||
|
|
ed7efc58b7 | ||
|
|
cf3333d607 | ||
|
|
a375daead7 | ||
|
|
a5a0e0c6be | ||
|
|
9e8e20b3b5 | ||
|
|
37083fc92a |
@@ -0,0 +1,117 @@
|
||||
---
|
||||
name: diagnose
|
||||
description: Disciplined diagnosis loop for hard bugs and performance regressions. Reproduce → minimise → hypothesise → instrument → fix → regression-test. Use when user says "diagnose this" / "debug this", reports a bug, says something is broken/throwing/failing, or describes a performance regression.
|
||||
---
|
||||
|
||||
# Diagnose
|
||||
|
||||
A discipline for hard bugs. Skip phases only when explicitly justified.
|
||||
|
||||
When exploring the codebase, use the project's domain glossary to get a clear mental model of the relevant modules, and check ADRs in the area you're touching.
|
||||
|
||||
## Phase 1 — Build a feedback loop
|
||||
|
||||
**This is the skill.** Everything else is mechanical. If you have a fast, deterministic, agent-runnable pass/fail signal for the bug, you will find the cause — bisection, hypothesis-testing, and instrumentation all just consume that signal. If you don't have one, no amount of staring at code will save you.
|
||||
|
||||
Spend disproportionate effort here. **Be aggressive. Be creative. Refuse to give up.**
|
||||
|
||||
### Ways to construct one — try them in roughly this order
|
||||
|
||||
1. **Failing test** at whatever seam reaches the bug — unit, integration, e2e.
|
||||
2. **Curl / HTTP script** against a running dev server.
|
||||
3. **CLI invocation** with a fixture input, diffing stdout against a known-good snapshot.
|
||||
4. **Headless browser script** (Playwright / Puppeteer) — drives the UI, asserts on DOM/console/network.
|
||||
5. **Replay a captured trace.** Save a real network request / payload / event log to disk; replay it through the code path in isolation.
|
||||
6. **Throwaway harness.** Spin up a minimal subset of the system (one service, mocked deps) that exercises the bug code path with a single function call.
|
||||
7. **Property / fuzz loop.** If the bug is "sometimes wrong output", run 1000 random inputs and look for the failure mode.
|
||||
8. **Bisection harness.** If the bug appeared between two known states (commit, dataset, version), automate "boot at state X, check, repeat" so you can `git bisect run` it.
|
||||
9. **Differential loop.** Run the same input through old-version vs new-version (or two configs) and diff outputs.
|
||||
10. **HITL bash script.** Last resort. If a human must click, drive _them_ with `scripts/hitl-loop.template.sh` so the loop is still structured. Captured output feeds back to you.
|
||||
|
||||
Build the right feedback loop, and the bug is 90% fixed.
|
||||
|
||||
### Iterate on the loop itself
|
||||
|
||||
Treat the loop as a product. Once you have _a_ loop, ask:
|
||||
|
||||
- Can I make it faster? (Cache setup, skip unrelated init, narrow the test scope.)
|
||||
- Can I make the signal sharper? (Assert on the specific symptom, not "didn't crash".)
|
||||
- Can I make it more deterministic? (Pin time, seed RNG, isolate filesystem, freeze network.)
|
||||
|
||||
A 30-second flaky loop is barely better than no loop. A 2-second deterministic loop is a debugging superpower.
|
||||
|
||||
### Non-deterministic bugs
|
||||
|
||||
The goal is not a clean repro but a **higher reproduction rate**. Loop the trigger 100×, parallelise, add stress, narrow timing windows, inject sleeps. A 50%-flake bug is debuggable; 1% is not — keep raising the rate until it's debuggable.
|
||||
|
||||
### When you genuinely cannot build a loop
|
||||
|
||||
Stop and say so explicitly. List what you tried. Ask the user for: (a) access to whatever environment reproduces it, (b) a captured artifact (HAR file, log dump, core dump, screen recording with timestamps), or (c) permission to add temporary production instrumentation. Do **not** proceed to hypothesise without a loop.
|
||||
|
||||
Do not proceed to Phase 2 until you have a loop you believe in.
|
||||
|
||||
## Phase 2 — Reproduce
|
||||
|
||||
Run the loop. Watch the bug appear.
|
||||
|
||||
Confirm:
|
||||
|
||||
- [ ] The loop produces the failure mode the **user** described — not a different failure that happens to be nearby. Wrong bug = wrong fix.
|
||||
- [ ] The failure is reproducible across multiple runs (or, for non-deterministic bugs, reproducible at a high enough rate to debug against).
|
||||
- [ ] You have captured the exact symptom (error message, wrong output, slow timing) so later phases can verify the fix actually addresses it.
|
||||
|
||||
Do not proceed until you reproduce the bug.
|
||||
|
||||
## Phase 3 — Hypothesise
|
||||
|
||||
Generate **3–5 ranked hypotheses** before testing any of them. Single-hypothesis generation anchors on the first plausible idea.
|
||||
|
||||
Each hypothesis must be **falsifiable**: state the prediction it makes.
|
||||
|
||||
> Format: "If <X> is the cause, then <changing Y> will make the bug disappear / <changing Z> will make it worse."
|
||||
|
||||
If you cannot state the prediction, the hypothesis is a vibe — discard or sharpen it.
|
||||
|
||||
**Show the ranked list to the user before testing.** They often have domain knowledge that re-ranks instantly ("we just deployed a change to #3"), or know hypotheses they've already ruled out. Cheap checkpoint, big time saver. Don't block on it — proceed with your ranking if the user is AFK.
|
||||
|
||||
## Phase 4 — Instrument
|
||||
|
||||
Each probe must map to a specific prediction from Phase 3. **Change one variable at a time.**
|
||||
|
||||
Tool preference:
|
||||
|
||||
1. **Debugger / REPL inspection** if the env supports it. One breakpoint beats ten logs.
|
||||
2. **Targeted logs** at the boundaries that distinguish hypotheses.
|
||||
3. Never "log everything and grep".
|
||||
|
||||
**Tag every debug log** with a unique prefix, e.g. `[DEBUG-a4f2]`. Cleanup at the end becomes a single grep. Untagged logs survive; tagged logs die.
|
||||
|
||||
**Perf branch.** For performance regressions, logs are usually wrong. Instead: establish a baseline measurement (timing harness, `performance.now()`, profiler, query plan), then bisect. Measure first, fix second.
|
||||
|
||||
## Phase 5 — Fix + regression test
|
||||
|
||||
Write the regression test **before the fix** — but only if there is a **correct seam** for it.
|
||||
|
||||
A correct seam is one where the test exercises the **real bug pattern** as it occurs at the call site. If the only available seam is too shallow (single-caller test when the bug needs multiple callers, unit test that can't replicate the chain that triggered the bug), a regression test there gives false confidence.
|
||||
|
||||
**If no correct seam exists, that itself is the finding.** Note it. The codebase architecture is preventing the bug from being locked down. Flag this for the next phase.
|
||||
|
||||
If a correct seam exists:
|
||||
|
||||
1. Turn the minimised repro into a failing test at that seam.
|
||||
2. Watch it fail.
|
||||
3. Apply the fix.
|
||||
4. Watch it pass.
|
||||
5. Re-run the Phase 1 feedback loop against the original (un-minimised) scenario.
|
||||
|
||||
## Phase 6 — Cleanup + post-mortem
|
||||
|
||||
Required before declaring done:
|
||||
|
||||
- [ ] Original repro no longer reproduces (re-run the Phase 1 loop)
|
||||
- [ ] Regression test passes (or absence of seam is documented)
|
||||
- [ ] All `[DEBUG-...]` instrumentation removed (`grep` the prefix)
|
||||
- [ ] Throwaway prototypes deleted (or moved to a clearly-marked debug location)
|
||||
- [ ] The hypothesis that turned out correct is stated in the commit / PR message — so the next debugger learns
|
||||
|
||||
**Then ask: what would have prevented this bug?** If the answer involves architectural change (no good test seam, tangled callers, hidden coupling) hand off to the `/improve-codebase-architecture` skill with the specifics. Make the recommendation **after** the fix is in, not before — you have more information now than when you started.
|
||||
@@ -0,0 +1,47 @@
|
||||
# ADR Format
|
||||
|
||||
ADRs live in `docs/adr/` and use sequential numbering: `0001-slug.md`, `0002-slug.md`, etc.
|
||||
|
||||
Create the `docs/adr/` directory lazily — only when the first ADR is needed.
|
||||
|
||||
## Template
|
||||
|
||||
```md
|
||||
# {Short title of the decision}
|
||||
|
||||
{1-3 sentences: what's the context, what did we decide, and why.}
|
||||
```
|
||||
|
||||
That's it. An ADR can be a single paragraph. The value is in recording *that* a decision was made and *why* — not in filling out sections.
|
||||
|
||||
## Optional sections
|
||||
|
||||
Only include these when they add genuine value. Most ADRs won't need them.
|
||||
|
||||
- **Status** frontmatter (`proposed | accepted | deprecated | superseded by ADR-NNNN`) — useful when decisions are revisited
|
||||
- **Considered Options** — only when the rejected alternatives are worth remembering
|
||||
- **Consequences** — only when non-obvious downstream effects need to be called out
|
||||
|
||||
## Numbering
|
||||
|
||||
Scan `docs/adr/` for the highest existing number and increment by one.
|
||||
|
||||
## When to offer an ADR
|
||||
|
||||
All three of these must be true:
|
||||
|
||||
1. **Hard to reverse** — the cost of changing your mind later is meaningful
|
||||
2. **Surprising without context** — a future reader will look at the code and wonder "why on earth did they do it this way?"
|
||||
3. **The result of a real trade-off** — there were genuine alternatives and you picked one for specific reasons
|
||||
|
||||
If a decision is easy to reverse, skip it — you'll just reverse it. If it's not surprising, nobody will wonder why. If there was no real alternative, there's nothing to record beyond "we did the obvious thing."
|
||||
|
||||
### What qualifies
|
||||
|
||||
- **Architectural shape.** "We're using a monorepo." "The write model is event-sourced, the read model is projected into Postgres."
|
||||
- **Integration patterns between contexts.** "Ordering and Billing communicate via domain events, not synchronous HTTP."
|
||||
- **Technology choices that carry lock-in.** Database, message bus, auth provider, deployment target. Not every library — just the ones that would take a quarter to swap out.
|
||||
- **Boundary and scope decisions.** "Customer data is owned by the Customer context; other contexts reference it by ID only." The explicit no-s are as valuable as the yes-s.
|
||||
- **Deliberate deviations from the obvious path.** "We're using manual SQL instead of an ORM because X." Anything where a reasonable reader would assume the opposite. These stop the next engineer from "fixing" something that was deliberate.
|
||||
- **Constraints not visible in the code.** "We can't use AWS because of compliance requirements." "Response times must be under 200ms because of the partner API contract."
|
||||
- **Rejected alternatives when the rejection is non-obvious.** If you considered GraphQL and picked REST for subtle reasons, record it — otherwise someone will suggest GraphQL again in six months.
|
||||
@@ -0,0 +1,77 @@
|
||||
# CONTEXT.md Format
|
||||
|
||||
## Structure
|
||||
|
||||
```md
|
||||
# {Context Name}
|
||||
|
||||
{One or two sentence description of what this context is and why it exists.}
|
||||
|
||||
## Language
|
||||
|
||||
**Order**:
|
||||
{A concise description of the term}
|
||||
_Avoid_: Purchase, transaction
|
||||
|
||||
**Invoice**:
|
||||
A request for payment sent to a customer after delivery.
|
||||
_Avoid_: Bill, payment request
|
||||
|
||||
**Customer**:
|
||||
A person or organization that places orders.
|
||||
_Avoid_: Client, buyer, account
|
||||
|
||||
## Relationships
|
||||
|
||||
- An **Order** produces one or more **Invoices**
|
||||
- An **Invoice** belongs to exactly one **Customer**
|
||||
|
||||
## Example dialogue
|
||||
|
||||
> **Dev:** "When a **Customer** places an **Order**, do we create the **Invoice** immediately?"
|
||||
> **Domain expert:** "No — an **Invoice** is only generated once a **Fulfillment** is confirmed."
|
||||
|
||||
## Flagged ambiguities
|
||||
|
||||
- "account" was used to mean both **Customer** and **User** — resolved: these are distinct concepts.
|
||||
```
|
||||
|
||||
## Rules
|
||||
|
||||
- **Be opinionated.** When multiple words exist for the same concept, pick the best one and list the others as aliases to avoid.
|
||||
- **Flag conflicts explicitly.** If a term is used ambiguously, call it out in "Flagged ambiguities" with a clear resolution.
|
||||
- **Keep definitions tight.** One sentence max. Define what it IS, not what it does.
|
||||
- **Show relationships.** Use bold term names and express cardinality where obvious.
|
||||
- **Only include terms specific to this project's context.** General programming concepts (timeouts, error types, utility patterns) don't belong even if the project uses them extensively. Before adding a term, ask: is this a concept unique to this context, or a general programming concept? Only the former belongs.
|
||||
- **Group terms under subheadings** when natural clusters emerge. If all terms belong to a single cohesive area, a flat list is fine.
|
||||
- **Write an example dialogue.** A conversation between a dev and a domain expert that demonstrates how the terms interact naturally and clarifies boundaries between related concepts.
|
||||
|
||||
## Single vs multi-context repos
|
||||
|
||||
**Single context (most repos):** One `CONTEXT.md` at the repo root.
|
||||
|
||||
**Multiple contexts:** A `CONTEXT-MAP.md` at the repo root lists the contexts, where they live, and how they relate to each other:
|
||||
|
||||
```md
|
||||
# Context Map
|
||||
|
||||
## Contexts
|
||||
|
||||
- [Ordering](./src/ordering/CONTEXT.md) — receives and tracks customer orders
|
||||
- [Billing](./src/billing/CONTEXT.md) — generates invoices and processes payments
|
||||
- [Fulfillment](./src/fulfillment/CONTEXT.md) — manages warehouse picking and shipping
|
||||
|
||||
## Relationships
|
||||
|
||||
- **Ordering → Fulfillment**: Ordering emits `OrderPlaced` events; Fulfillment consumes them to start picking
|
||||
- **Fulfillment → Billing**: Fulfillment emits `ShipmentDispatched` events; Billing consumes them to generate invoices
|
||||
- **Ordering ↔ Billing**: Shared types for `CustomerId` and `Money`
|
||||
```
|
||||
|
||||
The skill infers which structure applies:
|
||||
|
||||
- If `CONTEXT-MAP.md` exists, read it to find contexts
|
||||
- If only a root `CONTEXT.md` exists, single context
|
||||
- If neither exists, create a root `CONTEXT.md` lazily when the first term is resolved
|
||||
|
||||
When multiple contexts exist, infer which one the current topic relates to. If unclear, ask.
|
||||
@@ -0,0 +1,88 @@
|
||||
---
|
||||
name: grill-with-docs
|
||||
description: Grilling session that challenges your plan against the existing domain model, sharpens terminology, and updates documentation (CONTEXT.md, ADRs) inline as decisions crystallise. Use when user wants to stress-test a plan against their project's language and documented decisions.
|
||||
---
|
||||
|
||||
<what-to-do>
|
||||
|
||||
Interview me relentlessly about every aspect of this plan until we reach a shared understanding. Walk down each branch of the design tree, resolving dependencies between decisions one-by-one. For each question, provide your recommended answer.
|
||||
|
||||
Ask the questions one at a time, waiting for feedback on each question before continuing.
|
||||
|
||||
If a question can be answered by exploring the codebase, explore the codebase instead.
|
||||
|
||||
</what-to-do>
|
||||
|
||||
<supporting-info>
|
||||
|
||||
## Domain awareness
|
||||
|
||||
During codebase exploration, also look for existing documentation:
|
||||
|
||||
### File structure
|
||||
|
||||
Most repos have a single context:
|
||||
|
||||
```
|
||||
/
|
||||
├── CONTEXT.md
|
||||
├── docs/
|
||||
│ └── adr/
|
||||
│ ├── 0001-event-sourced-orders.md
|
||||
│ └── 0002-postgres-for-write-model.md
|
||||
└── src/
|
||||
```
|
||||
|
||||
If a `CONTEXT-MAP.md` exists at the root, the repo has multiple contexts. The map points to where each one lives:
|
||||
|
||||
```
|
||||
/
|
||||
├── CONTEXT-MAP.md
|
||||
├── docs/
|
||||
│ └── adr/ ← system-wide decisions
|
||||
├── src/
|
||||
│ ├── ordering/
|
||||
│ │ ├── CONTEXT.md
|
||||
│ │ └── docs/adr/ ← context-specific decisions
|
||||
│ └── billing/
|
||||
│ ├── CONTEXT.md
|
||||
│ └── docs/adr/
|
||||
```
|
||||
|
||||
Create files lazily — only when you have something to write. If no `CONTEXT.md` exists, create one when the first term is resolved. If no `docs/adr/` exists, create it when the first ADR is needed.
|
||||
|
||||
## During the session
|
||||
|
||||
### Challenge against the glossary
|
||||
|
||||
When the user uses a term that conflicts with the existing language in `CONTEXT.md`, call it out immediately. "Your glossary defines 'cancellation' as X, but you seem to mean Y — which is it?"
|
||||
|
||||
### Sharpen fuzzy language
|
||||
|
||||
When the user uses vague or overloaded terms, propose a precise canonical term. "You're saying 'account' — do you mean the Customer or the User? Those are different things."
|
||||
|
||||
### Discuss concrete scenarios
|
||||
|
||||
When domain relationships are being discussed, stress-test them with specific scenarios. Invent scenarios that probe edge cases and force the user to be precise about the boundaries between concepts.
|
||||
|
||||
### Cross-reference with code
|
||||
|
||||
When the user states how something works, check whether the code agrees. If you find a contradiction, surface it: "Your code cancels entire Orders, but you just said partial cancellation is possible — which is right?"
|
||||
|
||||
### Update CONTEXT.md inline
|
||||
|
||||
When a term is resolved, update `CONTEXT.md` right there. Don't batch these up — capture them as they happen. Use the format in [CONTEXT-FORMAT.md](./CONTEXT-FORMAT.md).
|
||||
|
||||
`CONTEXT.md` should be totally devoid of implementation details. Do not treat `CONTEXT.md` as a spec, a scratch pad, or a repository for implementation decisions. It is a glossary and nothing else.
|
||||
|
||||
### Offer ADRs sparingly
|
||||
|
||||
Only offer to create an ADR when all three are true:
|
||||
|
||||
1. **Hard to reverse** — the cost of changing your mind later is meaningful
|
||||
2. **Surprising without context** — a future reader will wonder "why did they do it this way?"
|
||||
3. **The result of a real trade-off** — there were genuine alternatives and you picked one for specific reasons
|
||||
|
||||
If any of the three is missing, skip the ADR. Use the format in [ADR-FORMAT.md](./ADR-FORMAT.md).
|
||||
|
||||
</supporting-info>
|
||||
@@ -0,0 +1,109 @@
|
||||
---
|
||||
name: tdd
|
||||
description: Test-driven development with red-green-refactor loop. Use when user wants to build features or fix bugs using TDD, mentions "red-green-refactor", wants integration tests, or asks for test-first development.
|
||||
---
|
||||
|
||||
# Test-Driven Development
|
||||
|
||||
## Philosophy
|
||||
|
||||
**Core principle**: Tests should verify behavior through public interfaces, not implementation details. Code can change entirely; tests shouldn't.
|
||||
|
||||
**Good tests** are integration-style: they exercise real code paths through public APIs. They describe _what_ the system does, not _how_ it does it. A good test reads like a specification - "user can checkout with valid cart" tells you exactly what capability exists. These tests survive refactors because they don't care about internal structure.
|
||||
|
||||
**Bad tests** are coupled to implementation. They mock internal collaborators, test private methods, or verify through external means (like querying a database directly instead of using the interface). The warning sign: your test breaks when you refactor, but behavior hasn't changed. If you rename an internal function and tests fail, those tests were testing implementation, not behavior.
|
||||
|
||||
See [tests.md](tests.md) for examples and [mocking.md](mocking.md) for mocking guidelines.
|
||||
|
||||
## Anti-Pattern: Horizontal Slices
|
||||
|
||||
**DO NOT write all tests first, then all implementation.** This is "horizontal slicing" - treating RED as "write all tests" and GREEN as "write all code."
|
||||
|
||||
This produces **crap tests**:
|
||||
|
||||
- Tests written in bulk test _imagined_ behavior, not _actual_ behavior
|
||||
- You end up testing the _shape_ of things (data structures, function signatures) rather than user-facing behavior
|
||||
- Tests become insensitive to real changes - they pass when behavior breaks, fail when behavior is fine
|
||||
- You outrun your headlights, committing to test structure before understanding the implementation
|
||||
|
||||
**Correct approach**: Vertical slices via tracer bullets. One test → one implementation → repeat. Each test responds to what you learned from the previous cycle. Because you just wrote the code, you know exactly what behavior matters and how to verify it.
|
||||
|
||||
```
|
||||
WRONG (horizontal):
|
||||
RED: test1, test2, test3, test4, test5
|
||||
GREEN: impl1, impl2, impl3, impl4, impl5
|
||||
|
||||
RIGHT (vertical):
|
||||
RED→GREEN: test1→impl1
|
||||
RED→GREEN: test2→impl2
|
||||
RED→GREEN: test3→impl3
|
||||
...
|
||||
```
|
||||
|
||||
## Workflow
|
||||
|
||||
### 1. Planning
|
||||
|
||||
When exploring the codebase, use the project's domain glossary so that test names and interface vocabulary match the project's language, and respect ADRs in the area you're touching.
|
||||
|
||||
Before writing any code:
|
||||
|
||||
- [ ] Confirm with user what interface changes are needed
|
||||
- [ ] Confirm with user which behaviors to test (prioritize)
|
||||
- [ ] Identify opportunities for [deep modules](deep-modules.md) (small interface, deep implementation)
|
||||
- [ ] Design interfaces for [testability](interface-design.md)
|
||||
- [ ] List the behaviors to test (not implementation steps)
|
||||
- [ ] Get user approval on the plan
|
||||
|
||||
Ask: "What should the public interface look like? Which behaviors are most important to test?"
|
||||
|
||||
**You can't test everything.** Confirm with the user exactly which behaviors matter most. Focus testing effort on critical paths and complex logic, not every possible edge case.
|
||||
|
||||
### 2. Tracer Bullet
|
||||
|
||||
Write ONE test that confirms ONE thing about the system:
|
||||
|
||||
```
|
||||
RED: Write test for first behavior → test fails
|
||||
GREEN: Write minimal code to pass → test passes
|
||||
```
|
||||
|
||||
This is your tracer bullet - proves the path works end-to-end.
|
||||
|
||||
### 3. Incremental Loop
|
||||
|
||||
For each remaining behavior:
|
||||
|
||||
```
|
||||
RED: Write next test → fails
|
||||
GREEN: Minimal code to pass → passes
|
||||
```
|
||||
|
||||
Rules:
|
||||
|
||||
- One test at a time
|
||||
- Only enough code to pass current test
|
||||
- Don't anticipate future tests
|
||||
- Keep tests focused on observable behavior
|
||||
|
||||
### 4. Refactor
|
||||
|
||||
After all tests pass, look for [refactor candidates](refactoring.md):
|
||||
|
||||
- [ ] Extract duplication
|
||||
- [ ] Deepen modules (move complexity behind simple interfaces)
|
||||
- [ ] Apply SOLID principles where natural
|
||||
- [ ] Consider what new code reveals about existing code
|
||||
- [ ] Run tests after each refactor step
|
||||
|
||||
**Never refactor while RED.** Get to GREEN first.
|
||||
|
||||
## Checklist Per Cycle
|
||||
|
||||
```
|
||||
[ ] Test describes behavior, not implementation
|
||||
[ ] Test uses public interface only
|
||||
[ ] Test would survive internal refactor
|
||||
[ ] Code is minimal for this test
|
||||
[ ] No speculative features added
|
||||
```
|
||||
@@ -0,0 +1,33 @@
|
||||
# Deep Modules
|
||||
|
||||
From "A Philosophy of Software Design":
|
||||
|
||||
**Deep module** = small interface + lots of implementation
|
||||
|
||||
```
|
||||
┌─────────────────────┐
|
||||
│ Small Interface │ ← Few methods, simple params
|
||||
├─────────────────────┤
|
||||
│ │
|
||||
│ │
|
||||
│ Deep Implementation│ ← Complex logic hidden
|
||||
│ │
|
||||
│ │
|
||||
└─────────────────────┘
|
||||
```
|
||||
|
||||
**Shallow module** = large interface + little implementation (avoid)
|
||||
|
||||
```
|
||||
┌─────────────────────────────────┐
|
||||
│ Large Interface │ ← Many methods, complex params
|
||||
├─────────────────────────────────┤
|
||||
│ Thin Implementation │ ← Just passes through
|
||||
└─────────────────────────────────┘
|
||||
```
|
||||
|
||||
When designing interfaces, ask:
|
||||
|
||||
- Can I reduce the number of methods?
|
||||
- Can I simplify the parameters?
|
||||
- Can I hide more complexity inside?
|
||||
@@ -0,0 +1,31 @@
|
||||
# Interface Design for Testability
|
||||
|
||||
Good interfaces make testing natural:
|
||||
|
||||
1. **Accept dependencies, don't create them**
|
||||
|
||||
```typescript
|
||||
// Testable
|
||||
function processOrder(order, paymentGateway) {}
|
||||
|
||||
// Hard to test
|
||||
function processOrder(order) {
|
||||
const gateway = new StripeGateway();
|
||||
}
|
||||
```
|
||||
|
||||
2. **Return results, don't produce side effects**
|
||||
|
||||
```typescript
|
||||
// Testable
|
||||
function calculateDiscount(cart): Discount {}
|
||||
|
||||
// Hard to test
|
||||
function applyDiscount(cart): void {
|
||||
cart.total -= discount;
|
||||
}
|
||||
```
|
||||
|
||||
3. **Small surface area**
|
||||
- Fewer methods = fewer tests needed
|
||||
- Fewer params = simpler test setup
|
||||
@@ -0,0 +1,59 @@
|
||||
# When to Mock
|
||||
|
||||
Mock at **system boundaries** only:
|
||||
|
||||
- External APIs (payment, email, etc.)
|
||||
- Databases (sometimes - prefer test DB)
|
||||
- Time/randomness
|
||||
- File system (sometimes)
|
||||
|
||||
Don't mock:
|
||||
|
||||
- Your own classes/modules
|
||||
- Internal collaborators
|
||||
- Anything you control
|
||||
|
||||
## Designing for Mockability
|
||||
|
||||
At system boundaries, design interfaces that are easy to mock:
|
||||
|
||||
**1. Use dependency injection**
|
||||
|
||||
Pass external dependencies in rather than creating them internally:
|
||||
|
||||
```typescript
|
||||
// Easy to mock
|
||||
function processPayment(order, paymentClient) {
|
||||
return paymentClient.charge(order.total);
|
||||
}
|
||||
|
||||
// Hard to mock
|
||||
function processPayment(order) {
|
||||
const client = new StripeClient(process.env.STRIPE_KEY);
|
||||
return client.charge(order.total);
|
||||
}
|
||||
```
|
||||
|
||||
**2. Prefer SDK-style interfaces over generic fetchers**
|
||||
|
||||
Create specific functions for each external operation instead of one generic function with conditional logic:
|
||||
|
||||
```typescript
|
||||
// GOOD: Each function is independently mockable
|
||||
const api = {
|
||||
getUser: (id) => fetch(`/users/${id}`),
|
||||
getOrders: (userId) => fetch(`/users/${userId}/orders`),
|
||||
createOrder: (data) => fetch('/orders', { method: 'POST', body: data }),
|
||||
};
|
||||
|
||||
// BAD: Mocking requires conditional logic inside the mock
|
||||
const api = {
|
||||
fetch: (endpoint, options) => fetch(endpoint, options),
|
||||
};
|
||||
```
|
||||
|
||||
The SDK approach means:
|
||||
- Each mock returns one specific shape
|
||||
- No conditional logic in test setup
|
||||
- Easier to see which endpoints a test exercises
|
||||
- Type safety per endpoint
|
||||
@@ -0,0 +1,10 @@
|
||||
# Refactor Candidates
|
||||
|
||||
After TDD cycle, look for:
|
||||
|
||||
- **Duplication** → Extract function/class
|
||||
- **Long methods** → Break into private helpers (keep tests on public interface)
|
||||
- **Shallow modules** → Combine or deepen
|
||||
- **Feature envy** → Move logic to where data lives
|
||||
- **Primitive obsession** → Introduce value objects
|
||||
- **Existing code** the new code reveals as problematic
|
||||
@@ -0,0 +1,61 @@
|
||||
# Good and Bad Tests
|
||||
|
||||
## Good Tests
|
||||
|
||||
**Integration-style**: Test through real interfaces, not mocks of internal parts.
|
||||
|
||||
```typescript
|
||||
// GOOD: Tests observable behavior
|
||||
test("user can checkout with valid cart", async () => {
|
||||
const cart = createCart();
|
||||
cart.add(product);
|
||||
const result = await checkout(cart, paymentMethod);
|
||||
expect(result.status).toBe("confirmed");
|
||||
});
|
||||
```
|
||||
|
||||
Characteristics:
|
||||
|
||||
- Tests behavior users/callers care about
|
||||
- Uses public API only
|
||||
- Survives internal refactors
|
||||
- Describes WHAT, not HOW
|
||||
- One logical assertion per test
|
||||
|
||||
## Bad Tests
|
||||
|
||||
**Implementation-detail tests**: Coupled to internal structure.
|
||||
|
||||
```typescript
|
||||
// BAD: Tests implementation details
|
||||
test("checkout calls paymentService.process", async () => {
|
||||
const mockPayment = jest.mock(paymentService);
|
||||
await checkout(cart, payment);
|
||||
expect(mockPayment.process).toHaveBeenCalledWith(cart.total);
|
||||
});
|
||||
```
|
||||
|
||||
Red flags:
|
||||
|
||||
- Mocking internal collaborators
|
||||
- Testing private methods
|
||||
- Asserting on call counts/order
|
||||
- Test breaks when refactoring without behavior change
|
||||
- Test name describes HOW not WHAT
|
||||
- Verifying through external means instead of interface
|
||||
|
||||
```typescript
|
||||
// BAD: Bypasses interface to verify
|
||||
test("createUser saves to database", async () => {
|
||||
await createUser({ name: "Alice" });
|
||||
const row = await db.query("SELECT * FROM users WHERE name = ?", ["Alice"]);
|
||||
expect(row).toBeDefined();
|
||||
});
|
||||
|
||||
// GOOD: Verifies through interface
|
||||
test("createUser makes user retrievable", async () => {
|
||||
const user = await createUser({ name: "Alice" });
|
||||
const retrieved = await getUser(user.id);
|
||||
expect(retrieved.name).toBe("Alice");
|
||||
});
|
||||
```
|
||||
@@ -0,0 +1,76 @@
|
||||
---
|
||||
name: to-prd
|
||||
description: Turn the current conversation context into a PRD and publish it to the project issue tracker. Use when user wants to create a PRD from the current context.
|
||||
---
|
||||
|
||||
This skill takes the current conversation context and codebase understanding and produces a PRD. Do NOT interview the user — just synthesize what you already know.
|
||||
|
||||
The issue tracker and triage label vocabulary should have been provided to you — run `/setup-matt-pocock-skills` if not.
|
||||
|
||||
## Process
|
||||
|
||||
1. Explore the repo to understand the current state of the codebase, if you haven't already. Use the project's domain glossary vocabulary throughout the PRD, and respect any ADRs in the area you're touching.
|
||||
|
||||
2. Sketch out the major modules you will need to build or modify to complete the implementation. Actively look for opportunities to extract deep modules that can be tested in isolation.
|
||||
|
||||
A deep module (as opposed to a shallow module) is one which encapsulates a lot of functionality in a simple, testable interface which rarely changes.
|
||||
|
||||
Check with the user that these modules match their expectations. Check with the user which modules they want tests written for.
|
||||
|
||||
3. Write the PRD using the template below, then publish it to the project issue tracker. Apply the `ready-for-agent` triage label - no need for additional triage.
|
||||
|
||||
<prd-template>
|
||||
|
||||
## Problem Statement
|
||||
|
||||
The problem that the user is facing, from the user's perspective.
|
||||
|
||||
## Solution
|
||||
|
||||
The solution to the problem, from the user's perspective.
|
||||
|
||||
## User Stories
|
||||
|
||||
A LONG, numbered list of user stories. Each user story should be in the format of:
|
||||
|
||||
1. As an <actor>, I want a <feature>, so that <benefit>
|
||||
|
||||
<user-story-example>
|
||||
1. As a mobile bank customer, I want to see balance on my accounts, so that I can make better informed decisions about my spending
|
||||
</user-story-example>
|
||||
|
||||
This list of user stories should be extremely extensive and cover all aspects of the feature.
|
||||
|
||||
## Implementation Decisions
|
||||
|
||||
A list of implementation decisions that were made. This can include:
|
||||
|
||||
- The modules that will be built/modified
|
||||
- The interfaces of those modules that will be modified
|
||||
- Technical clarifications from the developer
|
||||
- Architectural decisions
|
||||
- Schema changes
|
||||
- API contracts
|
||||
- Specific interactions
|
||||
|
||||
Do NOT include specific file paths or code snippets. They may end up being outdated very quickly.
|
||||
|
||||
Exception: if a prototype produced a snippet that encodes a decision more precisely than prose can (state machine, reducer, schema, type shape), inline it within the relevant decision and note briefly that it came from a prototype. Trim to the decision-rich parts — not a working demo, just the important bits.
|
||||
|
||||
## Testing Decisions
|
||||
|
||||
A list of testing decisions that were made. Include:
|
||||
|
||||
- A description of what makes a good test (only test external behavior, not implementation details)
|
||||
- Which modules will be tested
|
||||
- Prior art for the tests (i.e. similar types of tests in the codebase)
|
||||
|
||||
## Out of Scope
|
||||
|
||||
A description of the things that are out of scope for this PRD.
|
||||
|
||||
## Further Notes
|
||||
|
||||
Any further notes about the feature.
|
||||
|
||||
</prd-template>
|
||||
@@ -0,0 +1,7 @@
|
||||
---
|
||||
name: zoom-out
|
||||
description: Tell the agent to zoom out and give broader context or a higher-level perspective. Use when you're unfamiliar with a section of code or need to understand how it fits into the bigger picture.
|
||||
disable-model-invocation: true
|
||||
---
|
||||
|
||||
I don't know this area of code well. Go up a layer of abstraction. Give me a map of all the relevant modules and callers, using the project's domain glossary vocabulary.
|
||||
@@ -1,43 +1,107 @@
|
||||
<!-- gitnexus:start -->
|
||||
# GitNexus — Code Intelligence
|
||||
# CLAUDE.md
|
||||
|
||||
This project is indexed by GitNexus as **SuperBizAgent-java** (7988 symbols, 12713 relationships, 297 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
|
||||
## Defaults
|
||||
|
||||
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
|
||||
- Reply in **Chinese** unless I explicitly ask for English.
|
||||
- No emojis.
|
||||
- Do not truncate important outputs (logs, diffs, stack traces, commands, or critical reasoning that affects
|
||||
safety/correctness).
|
||||
|
||||
## Always Do
|
||||
## Refactor policy (legacy code)
|
||||
|
||||
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
|
||||
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
|
||||
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
|
||||
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
|
||||
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
|
||||
- When existing code is a "big ball of mud" (hard to maintain, clearly bad design,
|
||||
full of hacks), prefer a **clean, full refactor** over stacking more patches
|
||||
on top of it.
|
||||
- A refactor may completely replace internal structure
|
||||
(functions, modules, classes, data flow).
|
||||
- By default, try to preserve externally observable behaviour.
|
||||
If you intentionally change behaviour or protocols, you MUST:
|
||||
- Call out clearly that this is a **behaviour/protocol change**.
|
||||
- Explain why the change is necessary and which code paths/consumers are affected.
|
||||
- Update or add tests to cover the new behaviour.
|
||||
|
||||
## Never Do
|
||||
## Before touching code (mandatory)
|
||||
|
||||
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
|
||||
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
|
||||
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
|
||||
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
|
||||
Find reuse opportunities + Trace the call/dependency chain and impact radius:
|
||||
|
||||
## Resources
|
||||
- Use semantic code search first via `codebase-retrieval` tool.
|
||||
- Confirm understanding with LSP: `goToDefinition`, `findReferences`.
|
||||
- Use Grep/Glob for verifying and understanding additional code snippets.
|
||||
|
||||
| Resource | Use for |
|
||||
|----------|---------|
|
||||
| `gitnexus://repo/SuperBizAgent-java/context` | Codebase overview, check index freshness |
|
||||
| `gitnexus://repo/SuperBizAgent-java/clusters` | All functional areas |
|
||||
| `gitnexus://repo/SuperBizAgent-java/processes` | All execution flows |
|
||||
| `gitnexus://repo/SuperBizAgent-java/process/{name}` | Step-by-step execution trace |
|
||||
## Red lines
|
||||
|
||||
## CLI
|
||||
- No copy-paste duplication.
|
||||
- Do not break existing externally observable behaviour **unless**:
|
||||
- It is part of a deliberate refactor as described in the refactor policy, and
|
||||
- You clearly document the behavioural change and its impact.
|
||||
- Do not proceed with a known-wrong approach.
|
||||
- Critical paths must have explicit error handling.
|
||||
- Never implement "blindly": always confirm understanding via code reading + references.
|
||||
|
||||
| Task | Read this skill file |
|
||||
|------|---------------------|
|
||||
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
|
||||
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
|
||||
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
|
||||
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
|
||||
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
|
||||
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
|
||||
## Task sizing
|
||||
|
||||
- **Simple**
|
||||
- Criteria — single file, clear requirement, < 20 lines changed,
|
||||
clearly local impact.
|
||||
- Handling — after doing the "Before touching code" steps
|
||||
(research + impact analysis + internal three-question checklist),
|
||||
you may execute directly with minimal explanation.
|
||||
- A very short context line is enough;
|
||||
a full breakdown of the checklist is not required.
|
||||
|
||||
- **Medium**
|
||||
- Criteria — 2–5 files, or requires some research, or impact is not obviously local.
|
||||
- Handling — write a short plan (bullet points) → then implement.
|
||||
- Briefly surface the checklist result in the reply
|
||||
(1–3 short lines describing real issue, key reuse, and main impact).
|
||||
|
||||
- **Complex**
|
||||
- Criteria — architecture changes, multiple modules, high uncertainty or risk.
|
||||
- Handling — follow this workflow:
|
||||
1. **RESEARCH**: inspect code and facts only (no proposals yet).
|
||||
2. **PLAN**: present options + tradeoffs + recommendation;
|
||||
use `AskUserQuestion` actively to align with the user;
|
||||
wait for user's confirmation.
|
||||
3. **EXECUTE**: implement exactly the approved plan.
|
||||
4. **REVIEW**: self-check (tests, edge cases, cleanup).
|
||||
|
||||
## Git
|
||||
|
||||
- Do not commit unless I explicitly ask.
|
||||
- Do not push unless I explicitly ask.
|
||||
- Before writing a commit message, glance at a few recent commits and match the repo's style:
|
||||
- `git log -n 5 --oneline`
|
||||
- If there is no obvious existing style, use this default format:
|
||||
- `<type>(<scope>): <description>`
|
||||
- Before any commit: run `git diff` and confirm the exact scope of changes.
|
||||
- Never force-push to `main` / `master` unless the user approves.
|
||||
- Do not add attribution lines in commit messages.
|
||||
|
||||
## Security
|
||||
|
||||
- Never hardcode secrets (keys/passwords/tokens).
|
||||
- Never commit `.env` files or any credentials.
|
||||
- Validate user input at trust boundaries (APIs, CLIs, external data sources).
|
||||
|
||||
## Quality & cleanup
|
||||
|
||||
- Prefer clarity and simplicity first (KISS); apply DRY to remove obvious
|
||||
copy-paste duplication when it does not hurt readability.
|
||||
- If you change a function signature, update **all** call sites.
|
||||
- After changes:
|
||||
- Remove temporary files.
|
||||
- Remove dead/commented-out code.
|
||||
- Remove unused imports.
|
||||
- Remove debug logging that is no longer needed.
|
||||
- Run the smallest meaningful verification (lint/test/build) for the parts you touched.
|
||||
|
||||
## Windows / PowerShell (if used)
|
||||
|
||||
- PowerShell does not support `&&`; use `;` to chain commands.
|
||||
- Quote paths that contain spaces or non-ASCII characters.
|
||||
|
||||
## Baisc Infos
|
||||
|
||||
Unless directly relevant to the user's current question, you should avoid proactively mentioning, illustrating, or
|
||||
trailing off into the following information in 99% of cases:
|
||||
|
||||
<!-- gitnexus:end -->
|
||||
|
||||
@@ -111,47 +111,3 @@ trailing off into the following information in 99% of cases:
|
||||
- 文档目录结构:
|
||||
- 不要将文档放到用户目录(如 `C:\Users\EDY\.claude\`)中
|
||||
|
||||
|
||||
<!-- gitnexus:start -->
|
||||
# GitNexus — Code Intelligence
|
||||
|
||||
This project is indexed by GitNexus as **SuperBizAgent-java** (7988 symbols, 12713 relationships, 297 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
|
||||
|
||||
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
|
||||
|
||||
## Always Do
|
||||
|
||||
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
|
||||
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
|
||||
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
|
||||
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
|
||||
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
|
||||
|
||||
## Never Do
|
||||
|
||||
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
|
||||
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
|
||||
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
|
||||
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
|
||||
|
||||
## Resources
|
||||
|
||||
| Resource | Use for |
|
||||
|----------|---------|
|
||||
| `gitnexus://repo/SuperBizAgent-java/context` | Codebase overview, check index freshness |
|
||||
| `gitnexus://repo/SuperBizAgent-java/clusters` | All functional areas |
|
||||
| `gitnexus://repo/SuperBizAgent-java/processes` | All execution flows |
|
||||
| `gitnexus://repo/SuperBizAgent-java/process/{name}` | Step-by-step execution trace |
|
||||
|
||||
## CLI
|
||||
|
||||
| Task | Read this skill file |
|
||||
|------|---------------------|
|
||||
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
|
||||
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
|
||||
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
|
||||
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
|
||||
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
|
||||
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
|
||||
|
||||
<!-- gitnexus:end -->
|
||||
|
||||
@@ -5,6 +5,8 @@
|
||||
| 日期 | slug | 领域 | 关键词 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| 2026-07-05 | diagnosis-playbook-skills | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented |
|
||||
| 2026-07-06 | rag-eval-pipeline-closure | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
|
||||
| 2026-07-06 | modular-rag-pipeline | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
|
||||
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
|
||||
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
|
||||
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
|
||||
|
||||
@@ -0,0 +1,101 @@
|
||||
# Modular RAG Pipeline — Acceptance
|
||||
|
||||
## 验收状态
|
||||
|
||||
状态:通过,OpenSpec 已归档。
|
||||
|
||||
任务完成:
|
||||
|
||||
- OpenSpec tasks:31/31 完成。
|
||||
- Review 后新增去重边界修复和回归测试。
|
||||
- OpenSpec archive:`openspec/changes/archive/2026-07-06-modular-rag-pipeline`。
|
||||
|
||||
## 静态验证
|
||||
|
||||
```powershell
|
||||
openspec validate modular-rag-pipeline --strict
|
||||
```
|
||||
|
||||
结果:
|
||||
|
||||
```text
|
||||
Change 'modular-rag-pipeline' is valid
|
||||
```
|
||||
|
||||
```powershell
|
||||
git diff --check
|
||||
```
|
||||
|
||||
结果:
|
||||
|
||||
```text
|
||||
PASS
|
||||
```
|
||||
|
||||
说明:仅出现 Windows LF/CRLF warning,无 whitespace error。
|
||||
|
||||
## 脚本验证
|
||||
|
||||
```powershell
|
||||
mvn -q -DskipTests compile
|
||||
```
|
||||
|
||||
结果:PASS。
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=LookupKnowledgeToolTest,ToolInvocationRecorderTest" test
|
||||
```
|
||||
|
||||
结果:PASS。
|
||||
|
||||
覆盖:
|
||||
|
||||
- filtered L1 成功不 retry。
|
||||
- filtered L1 低质量触发 raw unfiltered retry。
|
||||
- filtered L1 无 evidence 触发 raw unfiltered retry。
|
||||
- L0 hint 不作为 standalone fact evidence。
|
||||
- 无 L0 hint 时直接 unfiltered vector search。
|
||||
- rerank 使用 hint match 并记录 trace。
|
||||
- context pack 保留 source/title/breadcrumb/hit reasons。
|
||||
- evidence blocks 按 source 去重。
|
||||
- session dedup 不再返回可消费 evidence/context。
|
||||
- recorder 记录 retrieval trace、rerank trace、context pack summary 和 evidence summaries。
|
||||
|
||||
```powershell
|
||||
$env:MILVUS_TOKEN = <application.yml 中的 milvus.token>; mvn -q test
|
||||
```
|
||||
|
||||
结果:PASS。
|
||||
|
||||
说明:
|
||||
|
||||
- `MilvusConnectionTest` 需要 `MILVUS_TOKEN` 环境变量,直接读 `System.getenv`,不会自动读 `application.yml`。
|
||||
- 注入该环境变量后完整测试通过。
|
||||
|
||||
## 实现验收
|
||||
|
||||
已验证行为:
|
||||
|
||||
- `LookupKnowledgeTool` 已变为 pipeline orchestrator。
|
||||
- `LookupResult` 新契约包含 `evidenceBlocks`、`contextPack`、`retrievalTrace`、`rerankTrace`。
|
||||
- 旧 `primary/supplement` 字段和 DTO 已删除。
|
||||
- `ToolInvocationRecorder` 不再依赖 `result.getPrimary()`。
|
||||
- filtered retrieval 失败时会记录 `filtered_vector_no_evidence` 或 `filtered_vector_low_quality`。
|
||||
- no-evidence 情况返回 `found=false` 且保留 retrieval trace。
|
||||
- session dedup 情况返回 `found=false` 且 evidence/context 为空。
|
||||
|
||||
## 未验证项
|
||||
|
||||
人工 Demo 未执行:
|
||||
|
||||
- 还没有通过真实 Chat/AIOps 会话观察 Agent 是否稳定按 `contextPack.packedText` 和 `evidenceBlocks` 引用证据。
|
||||
|
||||
风险:
|
||||
|
||||
- 工具 JSON 契约是 L4 breaking change,prompt 已更新,但真实对话行为仍建议做一次端到端 demo。
|
||||
|
||||
## 后续建议
|
||||
|
||||
- 增加一组 RAG eval cases,固定 query、期望 evidence source、期望 fallback path。
|
||||
- 将 `MilvusConnectionTest` 改成 Spring 配置驱动或 integration profile,避免配置源混用。
|
||||
- 后续可在评测数据足够后再考虑 model-based rerank 或 hybrid retrieval。
|
||||
@@ -0,0 +1,54 @@
|
||||
# Modular RAG Pipeline — Brief
|
||||
|
||||
## 背景
|
||||
|
||||
`lookup_knowledge` 已经能返回知识库证据,但实现集中在 `LookupKnowledgeTool` 内部:L0 查询分析、L1 向量召回、相关性归一化、证据组装、会话去重和 trace 入库耦合在一起。
|
||||
|
||||
旧返回契约 `primary/supplement` 也延续了“L0 是主结果、L1 是补充”的语义,和当前设计目标不一致。新的目标是让 L0 只作为 query understanding / filter / rerank / trace 信号,让 L1 向量检索成为事实证据来源。
|
||||
|
||||
## 目标
|
||||
|
||||
- 将 `lookup_knowledge` 改造成模块化 RAG pipeline。
|
||||
- 保留显式 Agent tool 边界,不改工具名和 query 参数。
|
||||
- L0 只提供领域、关键词、实体、category filter 和 trace hint。
|
||||
- L1 filtered vector retrieval 失败或低质量时,降级为 raw query unfiltered L1 retry。
|
||||
- 输出 evidence-first contract:`evidenceBlocks`、`contextPack`、`retrievalTrace`、`rerankTrace`。
|
||||
- 保持 `tool_invocation` 表结构稳定,把新 trace 写入 `retrieval_details` JSON。
|
||||
|
||||
## 范围
|
||||
|
||||
已完成:
|
||||
|
||||
- 新增 pipeline DTO:`KnowledgeQuery`、`RetrievedEvidenceCandidate`、`ContextPack`、`RetrievalTrace`、`RerankTrace`、`EvidencePostprocessResult`。
|
||||
- 新增 pipeline service:`KnowledgeQueryTransformer`、`KnowledgeDocumentRetriever`、`KnowledgeEvidencePostProcessor`、`KnowledgeContextPacker`、`LookupResultAssembler`。
|
||||
- 重构 `LookupKnowledgeTool` 为薄 orchestration 层。
|
||||
- 迁移 `LookupResult`,删除 `primary/supplement` 字段和 `PrimaryResult` / `SupplementResult` 类。
|
||||
- 更新 `ToolInvocationRecorder`,记录 query transform、retrieval trace、context pack summary、rerank trace、fallback reason 和 evidence summaries。
|
||||
- 更新 executor prompt 和 RAG 架构文档。
|
||||
- 补充 lookup、recorder、fallback、rerank、context pack、session dedup 测试。
|
||||
|
||||
非目标:
|
||||
|
||||
- 不引入 implicit Advisor。
|
||||
- 不引入 cross-encoder、BM25、RRF、Elasticsearch、OpenSearch。
|
||||
- 不改文档上传、chunk、embedding 写入、Milvus schema。
|
||||
- 不改变 Agent 何时调用 `lookup_knowledge`。
|
||||
|
||||
## 关联 OpenSpec
|
||||
|
||||
- `openspec/changes/archive/2026-07-06-modular-rag-pipeline`
|
||||
|
||||
## 接口影响
|
||||
|
||||
级别:L4 breaking interface。
|
||||
|
||||
原因:
|
||||
|
||||
- 删除旧 `LookupResult.primary` / `LookupResult.supplement`。
|
||||
- `lookup_knowledge` tool JSON 输出形状变化。
|
||||
|
||||
缓解:
|
||||
|
||||
- 工具名和输入参数保持不变。
|
||||
- in-repo 消费方、测试和 prompt 同步迁移。
|
||||
- `tool_invocation` 表结构不变。
|
||||
@@ -0,0 +1,72 @@
|
||||
# Modular RAG Pipeline — Decisions
|
||||
|
||||
## D1: `lookup_knowledge` 保持显式工具
|
||||
|
||||
不把知识检索做成隐式 Advisor。Agent 仍显式调用 `lookup_knowledge(query)`,这样 trace、Verifier、Eval 都能看到工具调用边界。
|
||||
|
||||
## D2: L0 只做 query understanding
|
||||
|
||||
L0 产出:
|
||||
|
||||
- `domainHints`
|
||||
- `matchedKeywords`
|
||||
- `entities`
|
||||
- `categoryFilter`
|
||||
- `l0Titles`
|
||||
- `l0MatchCount`
|
||||
|
||||
L0 不再直接转成 fact evidence。L0 hint 可以影响 filter、rerank、trace,但不能在 L1 失败时冒充知识证据。
|
||||
|
||||
## D3: MVP 降级策略采用 unfiltered L1 retry
|
||||
|
||||
流程:
|
||||
|
||||
```text
|
||||
filtered L1 with L0 category filter
|
||||
-> empty / no final evidence / below reference threshold
|
||||
-> raw query unfiltered L1 retry
|
||||
-> still no evidence => no_evidence
|
||||
```
|
||||
|
||||
取舍:
|
||||
|
||||
- 简单、可解释、适合 MVP。
|
||||
- 避免引入 BM25/RRF/multi-query/cross-encoder 的复杂度。
|
||||
- 代价是低质量场景多一次向量查询,已通过 trace 记录 attempt duration。
|
||||
|
||||
## D4: 删除 `primary/supplement`
|
||||
|
||||
这是一次 L4 breaking interface change。
|
||||
|
||||
删除原因:
|
||||
|
||||
- `primary/supplement` 绑定旧语义:L0 primary、L1 supplement。
|
||||
- 新设计中事实证据来自 `evidenceBlocks/contextPack`。
|
||||
|
||||
迁移结果:
|
||||
|
||||
- `LookupResult` 暴露 evidence-first 字段。
|
||||
- `PrimaryResult` / `SupplementResult` 已删除。
|
||||
- 生产代码和测试不再引用 `getPrimary()` / `getSupplement()`。
|
||||
|
||||
## D5: Trace 表结构保持稳定
|
||||
|
||||
`tool_invocation` 表不新增列。新增信息写入 `retrieval_details` JSON:
|
||||
|
||||
- `query_transform`
|
||||
- `retrieval_trace`
|
||||
- `context_pack_summary`
|
||||
- `rerank_trace`
|
||||
- `fallback_reason`
|
||||
- `evidence_blocks`
|
||||
|
||||
原因:当前 trace、Verifier、Eval 已经以 `tool_invocation` 为证据入口,JSON details 足够承载 RAG 细节,避免 schema churn。
|
||||
|
||||
## D6: 会话去重不返回可消费证据
|
||||
|
||||
Review 后修正:
|
||||
|
||||
- dedup result 的 `found=false` 必须和 evidence/context 语义一致。
|
||||
- 返回消息说明文档已检索过。
|
||||
- 不再返回 `evidenceBlocks/contextPack`,避免 Agent 重复使用同一证据。
|
||||
- 保留 `retrievalTrace` 和 `retrievedDomainsThisSession` 便于可观测。
|
||||
@@ -0,0 +1,56 @@
|
||||
# Modular RAG Pipeline — Evidence
|
||||
|
||||
## 代码证据
|
||||
|
||||
关键入口:
|
||||
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
|
||||
- `src/main/java/com/superbiz/agent/dto/LookupResult.java`
|
||||
|
||||
新增模块:
|
||||
|
||||
- `KnowledgeQueryTransformer`:复用 `KnowledgeIndexService.analyzeQuery`,把 L0 转成 query hints 和可选 `categoryFilter`。
|
||||
- `KnowledgeDocumentRetriever`:封装 `VectorSearchService.searchSimilarDocuments(query, topK, category)`,统一 filtered / unfiltered attempt。
|
||||
- `KnowledgeEvidencePostProcessor`:归一化 L2、创建 evidence blocks、source dedup、规则 rerank、输出 `RerankTrace`。
|
||||
- `KnowledgeContextPacker`:按字符预算打包 evidence,保留 source/title/breadcrumb/hit reasons。
|
||||
- `LookupResultAssembler`:统一组装 evidence-first result、no-evidence result、session dedup result。
|
||||
|
||||
## 设计证据
|
||||
|
||||
已有文档约束:
|
||||
|
||||
- `mvp/architecture/rag-architecture.md`:RAG 应表达为可解释 pipeline,而不是一坨工具逻辑。
|
||||
- `mvp/architecture/retrieval-observability.md`:L0 是 hint/explainability 层,L1 是语义检索主路径。
|
||||
- `devflow/glossary/CONTEXT.md`:`lookup_knowledge` 是显式 Agent evidence tool,`tool_invocation` 是 trace / verifier / eval 的证据来源。
|
||||
|
||||
OpenSpec 对齐:
|
||||
|
||||
- `openspec/changes/modular-rag-pipeline/proposal.md`
|
||||
- `openspec/changes/modular-rag-pipeline/design.md`
|
||||
- `openspec/changes/modular-rag-pipeline/specs/rag-knowledge-retrieval/spec.md`
|
||||
- `openspec/changes/modular-rag-pipeline/tasks.md`
|
||||
|
||||
## 用户确认
|
||||
|
||||
- 一次到位做模块化 RAG,而不是只做小补丁。
|
||||
- L0 不再作为事实证据兜底。
|
||||
- filtered L1 不准时,MVP 降级为 raw query unfiltered L1 retry。
|
||||
- 可以新增字段,并删除旧字段以换取后续流程清晰。
|
||||
|
||||
## Review 发现
|
||||
|
||||
Review 中发现一个非阻塞但应修复的问题:
|
||||
|
||||
- 会话去重命中时,返回 `found=false` 但仍带 `evidenceBlocks/contextPack`,可能导致 Agent 重复消费同一份证据。
|
||||
|
||||
修复:
|
||||
|
||||
- `LookupResultAssembler.deduped` 清空可消费 evidence/context,只保留 message、trace、relevance hint 和 session domain memory。
|
||||
- 新增 `LookupKnowledgeToolTest.sessionDedupDoesNotReturnConsumableEvidenceAgain`。
|
||||
|
||||
## 非阻塞观察
|
||||
|
||||
- `MilvusConnectionTest` 仍直接依赖 `MILVUS_TOKEN` 环境变量;主配置中已有 token,但测试不读 Spring 配置。
|
||||
- 测试日志仍有 ANTLR 版本 warning,不影响测试通过。
|
||||
- 控制台在部分命令输出中仍会出现中文编码显示问题,但源码按 UTF-8 读取时关键用户提示文本正常。
|
||||
@@ -0,0 +1,78 @@
|
||||
# Acceptance: rag-eval-pipeline-closure
|
||||
|
||||
## Status
|
||||
|
||||
Archived.
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
| Item | Status | Notes |
|
||||
|---|---|---|
|
||||
| Modular fixture support | Done | Evaluator reads `lookupResult.evidenceBlocks/contextPack/retrievalTrace/rerankTrace`. |
|
||||
| LookupResult-only contract | Done | Evaluator fails fixtures that do not expose `lookupResult`. |
|
||||
| Golden modular assertions | Done | Cases assert selected attempt, accepted fallback reason, evidence status, context sources, and rerank top source. |
|
||||
| Fallback coverage | Done | Added `chat-l0-filter-fallback` for filtered low-quality/no-evidence to unfiltered retry. |
|
||||
| Real tool snapshot generation | Done | Added `RagLookupSnapshotGeneratorTest` and `generate_rag_lookup_snapshots.ps1`, defaulting to Spring AI VectorStore mode. |
|
||||
| Seed docs import/reindex | Done | Added canonical seed docs, `RagEvalSeedImporterTest`, and `prepare_rag_eval_seed.ps1`. |
|
||||
| Eval metadata isolation | Done | Added `kb_scope` metadata and `retrieval.kb-scope` filtering for L0 and L1. |
|
||||
| Frontmatter body split | Done | Upload chunking embeds Markdown body, while frontmatter feeds metadata and L0. |
|
||||
| Baseline diff | Done | `--compare-to` writes JSON/Markdown diff and exits non-zero on regression. |
|
||||
| Documentation | Done | Updated RAG eval README and added `mvp/architecture/rag-eval-closure.md`. |
|
||||
|
||||
## Verification
|
||||
|
||||
```powershell
|
||||
python scripts\eval_rag_retrieval.py
|
||||
```
|
||||
|
||||
Result: passed. 7 cases, passRate=1.0, recall@5=1.0.
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=RagLookupSnapshotGeneratorTest" test
|
||||
```
|
||||
|
||||
Result: passed. The snapshot generator stays disabled unless `rag.snapshot.enabled=true` is provided.
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=RagLookupSnapshotGeneratorTest" "-Drag.snapshot.enabled=true" "-Drag.snapshot.fixtures=<temp-fixtures>" "-Drag.snapshot.retrievedAt=2026-07-06T00:00:00Z" "-Dretrieval.kb-scope=rag-eval" "-Dretrieval.vector-store.mode=spring" test
|
||||
python scripts\eval_rag_retrieval.py --fixtures <temp-fixtures> --json-report <temp-current.json> --markdown-report <temp-current.md>
|
||||
```
|
||||
|
||||
Result: passed. Spring AI VectorStore live snapshot produced 7 cases, passRate=1.0, recall@5=1.0. The fallback case used `selectedAttempt=UNFILTERED_VECTOR_RETRY` and `fallbackReason=filtered_vector_no_evidence`; the expected source `rag-l0-filter-fallback` remained rank 1.
|
||||
|
||||
```powershell
|
||||
python scripts\eval_rag_retrieval.py --json-report <temp-current.json> --markdown-report <temp-current.md> --compare-to eval\rag-retrieval\reports\baseline.json --diff-json-report <temp-diff.json> --diff-markdown-report <temp-diff.md>
|
||||
```
|
||||
|
||||
Result: passed. regressions=0.
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=FrontmatterParserTest,VectorIndexServiceTest,VectorSearchServiceTest,DocumentManagementServiceTest,RagLookupSnapshotGeneratorTest,RagEvalSeedImporterTest" test
|
||||
```
|
||||
|
||||
Result: passed. The seed importer and snapshot generator remain disabled unless their system properties are explicitly enabled.
|
||||
|
||||
```powershell
|
||||
$null = [scriptblock]::Create((Get-Content -Raw scripts\prepare_rag_eval_seed.ps1))
|
||||
$null = [scriptblock]::Create((Get-Content -Raw scripts\generate_rag_lookup_snapshots.ps1))
|
||||
```
|
||||
|
||||
Result: PowerShell syntax OK.
|
||||
|
||||
```powershell
|
||||
.\scripts\prepare_rag_eval_seed.ps1
|
||||
```
|
||||
|
||||
Result: passed. Seed docs were imported through `DocumentManagementService` and reindexed into the configured runtime DB/vector stack.
|
||||
|
||||
```powershell
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1 -Fixtures <temp-fixtures> -RetrievedAt 2026-07-06T00:00:00Z -SkipEval
|
||||
```
|
||||
|
||||
Result: passed after defaulting the script to `retrieval.vector-store.mode=spring`. The script generated live fixtures through the real `LookupKnowledgeTool` and then the offline evaluator reported 7 cases, passRate=1.0, recall@5=1.0.
|
||||
|
||||
```powershell
|
||||
git diff --check
|
||||
```
|
||||
|
||||
Result: no whitespace errors. Git reported only LF/CRLF conversion warnings.
|
||||
@@ -0,0 +1,21 @@
|
||||
# Brief: rag-eval-pipeline-closure
|
||||
|
||||
## Background
|
||||
|
||||
The modular RAG pipeline now returns `LookupResult` with `evidenceBlocks`, `contextPack`, `retrievalTrace`, and `rerankTrace`. The RAG retrieval baseline must validate that full contract, so it can detect regressions in fallback behavior, context packing, or rerank trace.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Reuse the existing offline RAG retrieval baseline.
|
||||
2. Extend it to support modular `LookupResult` fixtures.
|
||||
3. Add golden assertions for selected attempt, fallback reason, evidence status, context sources, and rerank top source.
|
||||
4. Add a RAG baseline diff path for regression detection.
|
||||
5. Add a snapshot generator that calls the real `LookupKnowledgeTool`.
|
||||
6. Document how RAG baseline and diagnosis baseline form a quality loop.
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No new production API.
|
||||
- No LLM-as-judge scoring.
|
||||
- No production API behavior changes.
|
||||
- No replacement for diagnosis eval.
|
||||
@@ -0,0 +1,57 @@
|
||||
# Decisions: rag-eval-pipeline-closure
|
||||
|
||||
## D1. Reuse the existing evaluator
|
||||
|
||||
Decision: extend `scripts/eval_rag_retrieval.py` instead of creating a second evaluator.
|
||||
|
||||
Reason: the old evaluator already owns golden cases, fixtures, hit-level classification, and Markdown/JSON reports. Extending it keeps one RAG baseline path.
|
||||
|
||||
## D2. Use LookupResult as the only fixture contract
|
||||
|
||||
Decision: support `lookupResult` only.
|
||||
|
||||
Reason: the MVP has moved to evidence-first RAG. Keeping an older fixture contract would weaken the baseline and let incomplete fixtures bypass context packing, retrieval trace, and rerank checks.
|
||||
|
||||
## D3. Make modular assertions opt-in per case
|
||||
|
||||
Decision: use fields such as `expectedSelectedAttempt`, `expectedFallbackReason`/`expectedFallbackReasons`, `expectedEvidenceStatus`, `expectedContextSources`, and `expectedRerankTopSource`.
|
||||
|
||||
Reason: golden cases can be strict where the pipeline path matters without forcing every historical case to assert every new field.
|
||||
|
||||
## D4. Diff remains deterministic
|
||||
|
||||
Decision: RAG diff compares report fields only and does not call live services or models.
|
||||
|
||||
Reason: this keeps it suitable for local regression checks and CI-style gates.
|
||||
|
||||
## D5. Isolate live eval docs with kb_scope
|
||||
|
||||
Decision: add `kb_scope` metadata and use `rag-eval` for canonical eval seed documents.
|
||||
|
||||
Reason: local production documents are not stable enough for golden retrieval expectations. Scope isolation lets real `LookupKnowledgeTool` snapshots use the same MySQL/Milvus stack while avoiding accidental matches from unrelated local data.
|
||||
|
||||
Default runtime keeps `retrieval.kb-scope` empty so legacy documents without `kb_scope` remain searchable. Eval scripts pass `-Dretrieval.kb-scope=rag-eval`. The same scope applies to L0 query hints and L1 vector retrieval.
|
||||
|
||||
## D6. Import seed docs through the real upload pipeline
|
||||
|
||||
Decision: seed docs are imported by `RagEvalSeedImporterTest` through `DocumentManagementService.uploadDocument`.
|
||||
|
||||
Reason: this updates DB metadata, L0 index state, local knowledge files, and Milvus chunks in the same way as normal document ingestion. A direct Milvus-only seed would make the live eval less representative.
|
||||
|
||||
## D7. Strip frontmatter before chunk embedding
|
||||
|
||||
Decision: uploaded Markdown frontmatter feeds metadata/L0 but is stripped before chunking and embedding.
|
||||
|
||||
Reason: frontmatter is a control plane, not evidence text. Keeping it in chunks lets L0-only keywords artificially improve vector similarity, especially for fallback decoy cases.
|
||||
|
||||
## D8. Treat retry behavior as the stable fallback contract
|
||||
|
||||
Decision: the fallback golden case accepts both `filtered_vector_low_quality` and `filtered_vector_no_evidence`, while still requiring `selectedAttempt=UNFILTERED_VECTOR_RETRY`, expected evidence source, context packing, and rerank top source.
|
||||
|
||||
Reason: Spring AI VectorStore and the Milvus SDK can differ on whether an over-filtered first pass returns a weak candidate or no candidate. The MVP contract is that the retriever skips only the L0 category filter, keeps `kb_scope`, retries the original query, and returns the correct evidence.
|
||||
|
||||
## D9. Default live snapshots to Spring AI VectorStore
|
||||
|
||||
Decision: `generate_rag_lookup_snapshots.ps1` defaults to `retrieval.vector-store.mode=spring`.
|
||||
|
||||
Reason: Spring AI VectorStore is the current framework path for the project and should be the default live verification route. SDK mode remains available through `-VectorStoreMode sdk` for comparison.
|
||||
@@ -0,0 +1,87 @@
|
||||
# Evidence: rag-eval-pipeline-closure
|
||||
|
||||
## Changed Assets
|
||||
|
||||
- `scripts/eval_rag_retrieval.py`
|
||||
- `eval/rag-retrieval/cases/golden-cases.json`
|
||||
- `eval/rag-retrieval/fixtures/*.json`
|
||||
- `eval/rag-retrieval/reports/baseline.json`
|
||||
- `eval/rag-retrieval/reports/baseline.md`
|
||||
- `eval/rag-retrieval/README.md`
|
||||
- `mvp/architecture/rag-eval-closure.md`
|
||||
- `src/test/java/com/superbiz/agent/eval/RagLookupSnapshotGeneratorTest.java`
|
||||
- `src/test/java/com/superbiz/agent/eval/RagEvalSeedImporterTest.java`
|
||||
- `scripts/generate_rag_lookup_snapshots.ps1`
|
||||
- `scripts/prepare_rag_eval_seed.ps1`
|
||||
- `eval/rag-retrieval/seed-docs/*.md`
|
||||
- `src/main/java/com/superbiz/agent/dto/Frontmatter.java`
|
||||
- `src/main/java/com/superbiz/agent/dto/KnowledgeEntry.java`
|
||||
- `src/main/java/com/superbiz/agent/service/FrontmatterParser.java`
|
||||
- `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/DocumentManagementService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/VectorIndexService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/SpringAiVectorStoreSidecarService.java`
|
||||
- `src/main/resources/application.yml`
|
||||
|
||||
## Baseline Result
|
||||
|
||||
```text
|
||||
Evaluated 7 cases: passRate=1.0, recall@5=1.0, failed=0
|
||||
```
|
||||
|
||||
## Regression Signals
|
||||
|
||||
The evaluator now fails on:
|
||||
|
||||
- missing expected source
|
||||
- missing breadcrumb or evidence keyword
|
||||
- non-`lookupResult` fixture
|
||||
- selected attempt mismatch
|
||||
- fallback reason mismatch
|
||||
- fallback reason outside accepted values
|
||||
- evidence status mismatch
|
||||
- missing context source
|
||||
- rerank top source mismatch
|
||||
|
||||
The snapshot generator now provides:
|
||||
|
||||
- real `LookupKnowledgeTool` invocation
|
||||
- one fixture per golden case
|
||||
- explicit opt-in through `rag.snapshot.enabled=true`
|
||||
- optional post-generation baseline evaluation
|
||||
- scoped retrieval through `retrieval.kb-scope=rag-eval`
|
||||
- Spring AI VectorStore by default through `retrieval.vector-store.mode=spring`
|
||||
- scoped L0 hints through the same `retrieval.kb-scope`
|
||||
|
||||
The seed importer now provides:
|
||||
|
||||
- canonical eval docs under `eval/rag-retrieval/seed-docs`
|
||||
- real `DocumentManagementService` import/reindex
|
||||
- stable `source`/`docId` metadata
|
||||
- `kb_scope=rag-eval` isolation from local non-eval documents
|
||||
- frontmatter stripping before chunk embedding
|
||||
- an over-filter decoy seed doc for fallback-path evaluation
|
||||
|
||||
The diff now detects:
|
||||
|
||||
- aggregate pass/recall regression
|
||||
- case pass regression
|
||||
- hit-level regression
|
||||
- first-rank regression
|
||||
- selected attempt/fallback/evidence/rerank changes
|
||||
|
||||
## Final Spring Live Snapshot
|
||||
|
||||
```text
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1 -Fixtures <temp-fixtures> -RetrievedAt 2026-07-06T00:00:00Z
|
||||
Evaluated 7 cases: passRate=1.0, recall@5=1.0, failed=0
|
||||
```
|
||||
|
||||
Key fallback trace:
|
||||
|
||||
```text
|
||||
selectedAttempt=UNFILTERED_VECTOR_RETRY
|
||||
fallbackReason=filtered_vector_no_evidence
|
||||
rerankTopSource=rag-l0-filter-fallback
|
||||
```
|
||||
@@ -12,12 +12,59 @@ post-processing, or Spring AI VectorStore integration.
|
||||
```text
|
||||
eval/rag-retrieval/
|
||||
cases/golden-cases.json Fixed retrieval golden cases
|
||||
fixtures/*.json Saved retrieval candidates for each case
|
||||
seed-docs/*.md Canonical docs imported into the live KB for real-tool eval
|
||||
fixtures/*.json Saved retrieval fixtures for each case
|
||||
reports/baseline.json Machine-readable baseline report
|
||||
reports/baseline.md Human-readable baseline report
|
||||
reports/baseline-diff.* Optional diff reports
|
||||
reports/live-post-reindex.* Optional live acceptance reports
|
||||
```
|
||||
|
||||
## Seed Docs + Import/Reindex
|
||||
|
||||
The live-tool eval uses canonical seed documents so the real
|
||||
`LookupKnowledgeTool` can retrieve stable evidence from MySQL/Milvus instead of
|
||||
whatever ad hoc documents happen to exist in the local knowledge base.
|
||||
|
||||
Seed documents live in:
|
||||
|
||||
```text
|
||||
eval/rag-retrieval/seed-docs/*.md
|
||||
```
|
||||
|
||||
Each seed doc uses frontmatter fields that are propagated into vector metadata:
|
||||
|
||||
```yaml
|
||||
source: mysql-connection-pool
|
||||
breadcrumb: Database > MySQL > Connection Pool
|
||||
kb_scope: rag-eval
|
||||
```
|
||||
|
||||
Import or reindex the seed docs through the real upload pipeline:
|
||||
|
||||
```powershell
|
||||
.\scripts\prepare_rag_eval_seed.ps1
|
||||
```
|
||||
|
||||
The script runs `RagEvalSeedImporterTest` with `rag.seed.enabled=true`. It
|
||||
deletes the existing document with the same `source`/`docId`, uploads the seed
|
||||
doc through `DocumentManagementService`, updates DB metadata and L0, and rebuilds
|
||||
Milvus chunks.
|
||||
|
||||
`kb_scope` isolates eval data:
|
||||
|
||||
- default application config leaves `retrieval.kb-scope` empty, so legacy docs
|
||||
without `kb_scope` remain searchable;
|
||||
- eval scripts pass `-Dretrieval.kb-scope=rag-eval`, so L0 query hints and L1
|
||||
vector retrieval both use only the canonical eval seed docs;
|
||||
- the fallback retry skips only the L0 category filter, not the `kb_scope`
|
||||
boundary.
|
||||
|
||||
Frontmatter is not embedded as chunk content during upload. It feeds metadata,
|
||||
L0, and document enrichment; only the Markdown body is chunked and embedded.
|
||||
This keeps controlled L0 decoys from becoming semantically relevant just because
|
||||
their frontmatter keywords matched the query.
|
||||
|
||||
## Run
|
||||
|
||||
From the repository root:
|
||||
@@ -36,6 +83,106 @@ python scripts/eval_rag_retrieval.py \
|
||||
--markdown-report eval/rag-retrieval/reports/baseline.md
|
||||
```
|
||||
|
||||
## Generate Fixtures From LookupKnowledgeTool
|
||||
|
||||
Use the snapshot generator when fixtures should reflect the real
|
||||
`LookupKnowledgeTool` pipeline:
|
||||
|
||||
```powershell
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1
|
||||
```
|
||||
|
||||
For the intended live loop, run seed import first:
|
||||
|
||||
```powershell
|
||||
.\scripts\prepare_rag_eval_seed.ps1
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1
|
||||
python scripts\eval_rag_retrieval.py
|
||||
```
|
||||
|
||||
The script runs a Spring test harness:
|
||||
|
||||
```text
|
||||
mvn -q -Dtest=RagLookupSnapshotGeneratorTest -Drag.snapshot.enabled=true -Dretrieval.kb-scope=rag-eval -Dretrieval.vector-store.mode=spring test
|
||||
```
|
||||
|
||||
The generator reads `golden-cases.json`, injects the real `LookupKnowledgeTool`
|
||||
bean, calls `lookupKnowledge(query)` for each case, writes
|
||||
`fixtures/{caseId}.json`, and then runs `eval_rag_retrieval.py` unless
|
||||
`-SkipEval` is provided. It defaults to Spring AI VectorStore mode; pass
|
||||
`-VectorStoreMode sdk` only when intentionally comparing the legacy SDK path.
|
||||
|
||||
Custom paths are supported:
|
||||
|
||||
```powershell
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1 `
|
||||
-Cases eval\rag-retrieval\cases\golden-cases.json `
|
||||
-Fixtures eval\rag-retrieval\fixtures `
|
||||
-RetrievedAt 2026-07-06T00:00:00Z
|
||||
```
|
||||
|
||||
The generator is disabled in normal test runs. It only executes when
|
||||
`rag.snapshot.enabled=true` is provided because it writes repository files and
|
||||
depends on the configured runtime retrieval stack.
|
||||
|
||||
If generated fixtures fail the offline baseline, treat that as a real alignment
|
||||
signal: either the golden expectations need to be adjusted to the current
|
||||
knowledge base, or the knowledge base/indexing path needs to be fixed.
|
||||
|
||||
## Modular RAG Contract
|
||||
|
||||
Fixtures must use the current `lookupResult` shape, which mirrors the
|
||||
`lookup_knowledge` output:
|
||||
|
||||
```text
|
||||
lookupResult.evidenceBlocks
|
||||
lookupResult.contextPack
|
||||
lookupResult.retrievalTrace
|
||||
lookupResult.rerankTrace
|
||||
```
|
||||
|
||||
Golden cases can assert both retrieval quality and pipeline behavior:
|
||||
|
||||
- `expectedSources` / `expectedDocIds`
|
||||
- `expectedBreadcrumbs`
|
||||
- `expectedKeywords`
|
||||
- `expectedSelectedAttempt`
|
||||
- `expectedFallbackReason`
|
||||
- `expectedFallbackReasons`
|
||||
- `expectedEvidenceStatus`
|
||||
- `expectedContextSources`
|
||||
- `expectedRerankTopSource`
|
||||
|
||||
This lets the baseline catch regressions such as losing the expected evidence
|
||||
source, skipping context packing, changing the selected retrieval attempt, or
|
||||
breaking the filtered-vector to unfiltered-retry fallback.
|
||||
|
||||
## Baseline Diff
|
||||
|
||||
To compare a freshly generated report against an existing baseline:
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_retrieval.py \
|
||||
--json-report eval/rag-retrieval/reports/current.json \
|
||||
--markdown-report eval/rag-retrieval/reports/current.md \
|
||||
--compare-to eval/rag-retrieval/reports/baseline.json \
|
||||
--diff-json-report eval/rag-retrieval/reports/baseline-diff.json \
|
||||
--diff-markdown-report eval/rag-retrieval/reports/baseline-diff.md
|
||||
```
|
||||
|
||||
The diff reports aggregate regressions and case-level changes for:
|
||||
|
||||
- pass rate, recall@K, strong hit rate, miss count
|
||||
- pass state
|
||||
- hit level
|
||||
- first expected rank
|
||||
- selected attempt
|
||||
- fallback reason
|
||||
- evidence status
|
||||
- rerank top source
|
||||
|
||||
The command exits non-zero when a case fails or the diff contains a regression.
|
||||
|
||||
## Hit Levels
|
||||
|
||||
- `strong`: expected document is found and breadcrumb or evidence keyword coverage is satisfied.
|
||||
@@ -81,6 +228,6 @@ GET /api/search/similar
|
||||
```
|
||||
|
||||
It writes JSON and Markdown reports with query, topK, result count, top
|
||||
candidates, breadcrumb, score labels, and raw response fields. This is a live
|
||||
results, breadcrumb, score labels, and raw response fields. This is a live
|
||||
smoke check for environment readiness and post-reindex behavior; it does not
|
||||
replace the deterministic offline baseline above.
|
||||
|
||||
@@ -8,8 +8,14 @@
|
||||
"scenario": "chat",
|
||||
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"expectedDocIds": ["mysql-connection-pool"],
|
||||
"expectedSources": ["mysql-connection-pool"],
|
||||
"expectedBreadcrumbs": ["Database > MySQL > Connection Pool"],
|
||||
"expectedKeywords": ["connection pool", "max_connections", "HikariCP"],
|
||||
"expectedSelectedAttempt": "FILTERED_VECTOR",
|
||||
"expectedFallbackReason": null,
|
||||
"expectedEvidenceStatus": "supported",
|
||||
"expectedContextSources": ["mysql-connection-pool"],
|
||||
"expectedRerankTopSource": "mysql-connection-pool",
|
||||
"notes": "Covers precise database troubleshooting retrieval."
|
||||
},
|
||||
{
|
||||
@@ -17,8 +23,14 @@
|
||||
"scenario": "chat",
|
||||
"query": "What is the standard troubleshooting flow for an application incident?",
|
||||
"expectedDocIds": ["incident-diagnosis-flow"],
|
||||
"expectedSources": ["incident-diagnosis-flow"],
|
||||
"expectedBreadcrumbs": ["AIOps > Diagnosis Flow"],
|
||||
"expectedKeywords": ["collect evidence", "verify", "remediation"],
|
||||
"expectedSelectedAttempt": "FILTERED_VECTOR",
|
||||
"expectedFallbackReason": null,
|
||||
"expectedEvidenceStatus": "supported",
|
||||
"expectedContextSources": ["incident-diagnosis-flow"],
|
||||
"expectedRerankTopSource": "incident-diagnosis-flow",
|
||||
"notes": "Covers process-style knowledge where breadcrumb matters."
|
||||
},
|
||||
{
|
||||
@@ -26,8 +38,14 @@
|
||||
"scenario": "aiops",
|
||||
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"expectedDocIds": ["payment-service-latency"],
|
||||
"expectedSources": ["payment-service-latency"],
|
||||
"expectedBreadcrumbs": ["AIOps > Service Alerts > Payment Latency"],
|
||||
"expectedKeywords": ["p95 latency", "payment-service", "downstream dependency"],
|
||||
"expectedSelectedAttempt": "FILTERED_VECTOR",
|
||||
"expectedFallbackReason": null,
|
||||
"expectedEvidenceStatus": "supported",
|
||||
"expectedContextSources": ["payment-service-latency"],
|
||||
"expectedRerankTopSource": "payment-service-latency",
|
||||
"notes": "Covers alert payload terms that should become retrieval hints."
|
||||
},
|
||||
{
|
||||
@@ -35,8 +53,14 @@
|
||||
"scenario": "aiops",
|
||||
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"expectedDocIds": ["aiops-alert-scope-control"],
|
||||
"expectedSources": ["aiops-alert-scope-control"],
|
||||
"expectedBreadcrumbs": ["AIOps > Alert Scope Control"],
|
||||
"expectedKeywords": ["payload", "unrelated active alerts", "scope"],
|
||||
"expectedSelectedAttempt": "FILTERED_VECTOR",
|
||||
"expectedFallbackReason": null,
|
||||
"expectedEvidenceStatus": "supported",
|
||||
"expectedContextSources": ["aiops-alert-scope-control"],
|
||||
"expectedRerankTopSource": "aiops-alert-scope-control",
|
||||
"notes": "Covers scoped alert diagnosis behavior."
|
||||
},
|
||||
{
|
||||
@@ -44,8 +68,14 @@
|
||||
"scenario": "chat",
|
||||
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"expectedDocIds": ["rag-chunk-context-reconstruction"],
|
||||
"expectedSources": ["rag-chunk-context-reconstruction"],
|
||||
"expectedBreadcrumbs": ["RAG > Chunking > Context Reconstruction"],
|
||||
"expectedKeywords": ["neighbor chunk", "same section", "breadcrumb"],
|
||||
"expectedSelectedAttempt": "FILTERED_VECTOR",
|
||||
"expectedFallbackReason": null,
|
||||
"expectedEvidenceStatus": "supported",
|
||||
"expectedContextSources": ["rag-chunk-context-reconstruction"],
|
||||
"expectedRerankTopSource": "rag-chunk-context-reconstruction",
|
||||
"notes": "Covers the known RAG refactor issue around context reconstruction."
|
||||
},
|
||||
{
|
||||
@@ -53,9 +83,30 @@
|
||||
"scenario": "chat",
|
||||
"query": "Should L0 keyword matching decide the final retrieval result?",
|
||||
"expectedDocIds": ["rag-l0-domain-entity-hint"],
|
||||
"expectedSources": ["rag-l0-domain-entity-hint"],
|
||||
"expectedBreadcrumbs": ["RAG > L0 > Domain Entity Hint"],
|
||||
"expectedKeywords": ["domain detector", "entity extractor", "metadata filter"],
|
||||
"expectedSelectedAttempt": "FILTERED_VECTOR",
|
||||
"expectedFallbackReason": null,
|
||||
"expectedEvidenceStatus": "supported",
|
||||
"expectedContextSources": ["rag-l0-domain-entity-hint"],
|
||||
"expectedRerankTopSource": "rag-l0-domain-entity-hint",
|
||||
"notes": "Covers the target L0 role after refactor."
|
||||
},
|
||||
{
|
||||
"caseId": "chat-l0-filter-fallback",
|
||||
"scenario": "chat",
|
||||
"query": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"expectedDocIds": ["rag-l0-filter-fallback"],
|
||||
"expectedSources": ["rag-l0-filter-fallback"],
|
||||
"expectedBreadcrumbs": ["RAG > Fallback > Unfiltered Retry"],
|
||||
"expectedKeywords": ["skip the L0 filter", "unfiltered vector retry", "low quality"],
|
||||
"expectedSelectedAttempt": "UNFILTERED_VECTOR_RETRY",
|
||||
"expectedFallbackReasons": ["filtered_vector_low_quality", "filtered_vector_no_evidence"],
|
||||
"expectedEvidenceStatus": "supported",
|
||||
"expectedContextSources": ["rag-l0-filter-fallback"],
|
||||
"expectedRerankTopSource": "rag-l0-filter-fallback",
|
||||
"notes": "Covers the MVP fallback rule: if filtered L1 is low quality, retry raw query without L0 filter."
|
||||
}
|
||||
]
|
||||
}
|
||||
|
||||
@@ -1,25 +1,81 @@
|
||||
{
|
||||
"caseId": "aiops-payment-latency-alert",
|
||||
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"retrievedAt": "2026-07-05T00:00:00Z",
|
||||
"candidates": [
|
||||
{
|
||||
"rank": 1,
|
||||
"docId": "payment-service-latency",
|
||||
"title": "Payment Service Latency Alert Playbook",
|
||||
"breadcrumb": "AIOps > Service Alerts > Payment Latency",
|
||||
"content": "For payment-service p95 latency alerts, check downstream dependency latency, thread pool saturation, gateway retries, and recent deployment changes.",
|
||||
"score": 0.84,
|
||||
"retrievalLayer": "L1"
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "payment-service-latency",
|
||||
"title": "Payment Service Latency Alert Playbook",
|
||||
"breadcrumb": "AIOps > Service Alerts > Payment Latency",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "For payment-service p95 latency alerts, check downstream dependency latency, thread pool saturation, gateway retries, and recent deployment changes.",
|
||||
"score": 0.84,
|
||||
"hitReasons": ["domain_match:+0.15", "entity_match:+0.20", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"source": "mysql-connection-pool",
|
||||
"title": "MySQL Connection Pool Troubleshooting",
|
||||
"breadcrumb": "Database > MySQL > Connection Pool",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "Database connection pool saturation can increase payment latency when checkout paths wait for connections.",
|
||||
"score": 0.68,
|
||||
"hitReasons": ["keyword_match:+0.10"]
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] Payment Service Latency Alert Playbook\nAIOps > Service Alerts > Payment Latency\nFor payment-service p95 latency alerts, check downstream dependency latency, thread pool saturation, gateway retries, and recent deployment changes.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 236,
|
||||
"includedSources": ["payment-service-latency", "mysql-connection-pool"],
|
||||
"omittedSources": []
|
||||
},
|
||||
{
|
||||
"rank": 2,
|
||||
"docId": "mysql-connection-pool",
|
||||
"title": "MySQL Connection Pool Troubleshooting",
|
||||
"breadcrumb": "Database > MySQL > Connection Pool",
|
||||
"content": "Database connection pool saturation can increase payment latency when checkout paths wait for connections.",
|
||||
"score": 0.68,
|
||||
"retrievalLayer": "L1"
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"rewrittenQuery": "HighLatency payment-service p95 latency alert downstream dependency diagnosis",
|
||||
"categoryFilter": "AIOps",
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["AIOps"],
|
||||
"matched_keywords": ["p95 latency", "payment-service", "downstream dependency"],
|
||||
"entities": ["payment-service", "HighLatency"],
|
||||
"l0_titles": ["Payment Service Latency Alert Playbook"],
|
||||
"l0_match_count": 1
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "HighLatency payment-service p95 latency alert downstream dependency diagnosis",
|
||||
"categoryFilter": "AIOps",
|
||||
"candidateCount": 2,
|
||||
"usable": true,
|
||||
"durationMs": 11,
|
||||
"topScore": 0.84,
|
||||
"topSimilarity": 0.84
|
||||
}
|
||||
]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "payment-service-latency",
|
||||
"baseScore": 0.84,
|
||||
"finalScore": 1.29,
|
||||
"boostReasons": ["domain_match:+0.15", "entity_match:+0.20", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"finalRank": 2,
|
||||
"source": "mysql-connection-pool",
|
||||
"baseScore": 0.68,
|
||||
"finalScore": 0.78,
|
||||
"boostReasons": ["keyword_match:+0.10"]
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,16 +1,65 @@
|
||||
{
|
||||
"caseId": "aiops-prometheus-alert-scope",
|
||||
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"retrievedAt": "2026-07-05T00:00:00Z",
|
||||
"candidates": [
|
||||
{
|
||||
"rank": 1,
|
||||
"docId": "aiops-alert-scope-control",
|
||||
"title": "AIOps Alert Scope Control",
|
||||
"breadcrumb": "AIOps > Alert Scope Control",
|
||||
"content": "When payload mode is active, queryPrometheusAlerts can verify the supplied alert, but unrelated active alerts must remain scoped context and should not become full diagnoses.",
|
||||
"score": 0.9,
|
||||
"retrievalLayer": "L0+L1"
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "aiops-alert-scope-control",
|
||||
"title": "AIOps Alert Scope Control",
|
||||
"breadcrumb": "AIOps > Alert Scope Control",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "When payload mode is used, diagnose the input alert payload and do not expand unrelated active alerts into the main diagnosis scope.",
|
||||
"score": 0.88,
|
||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] AIOps Alert Scope Control\nAIOps > Alert Scope Control\nWhen payload mode is used, diagnose the input alert payload and do not expand unrelated active alerts into the main diagnosis scope.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 188,
|
||||
"includedSources": ["aiops-alert-scope-control"],
|
||||
"omittedSources": []
|
||||
},
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"rewrittenQuery": "AIOps alert payload scope unrelated active alerts diagnosis",
|
||||
"categoryFilter": "AIOps",
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["AIOps"],
|
||||
"matched_keywords": ["payload", "unrelated active alerts", "scope"],
|
||||
"entities": ["alert payload"],
|
||||
"l0_titles": ["AIOps Alert Scope Control"],
|
||||
"l0_match_count": 1
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "AIOps alert payload scope unrelated active alerts diagnosis",
|
||||
"categoryFilter": "AIOps",
|
||||
"candidateCount": 1,
|
||||
"usable": true,
|
||||
"durationMs": 8,
|
||||
"topScore": 0.88,
|
||||
"topSimilarity": 0.88
|
||||
}
|
||||
]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "aiops-alert-scope-control",
|
||||
"baseScore": 0.88,
|
||||
"finalScore": 1.13,
|
||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,25 +1,81 @@
|
||||
{
|
||||
"caseId": "chat-diagnosis-flow",
|
||||
"query": "What is the standard troubleshooting flow for an application incident?",
|
||||
"retrievedAt": "2026-07-05T00:00:00Z",
|
||||
"candidates": [
|
||||
{
|
||||
"rank": 1,
|
||||
"docId": "incident-diagnosis-flow",
|
||||
"title": "Incident Diagnosis Flow",
|
||||
"breadcrumb": "AIOps > Diagnosis Flow",
|
||||
"content": "The standard flow is to collect evidence, identify the suspected fault domain, verify the hypothesis, apply remediation, and confirm recovery.",
|
||||
"score": 0.82,
|
||||
"retrievalLayer": "L1"
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "incident-diagnosis-flow",
|
||||
"title": "Incident Diagnosis Flow",
|
||||
"breadcrumb": "AIOps > Diagnosis Flow",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "The standard flow is to collect evidence, identify the suspected fault domain, verify the hypothesis, apply remediation, and confirm recovery.",
|
||||
"score": 0.82,
|
||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"source": "rag-chunk-context-reconstruction",
|
||||
"title": "RAG Chunk Context Reconstruction",
|
||||
"breadcrumb": "RAG > Chunking > Context Reconstruction",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "Long sections may require neighbor chunk expansion and breadcrumb-aware packing.",
|
||||
"score": 0.55,
|
||||
"hitReasons": []
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] Incident Diagnosis Flow\nAIOps > Diagnosis Flow\nThe standard flow is to collect evidence, identify the suspected fault domain, verify the hypothesis, apply remediation, and confirm recovery.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 192,
|
||||
"includedSources": ["incident-diagnosis-flow", "rag-chunk-context-reconstruction"],
|
||||
"omittedSources": []
|
||||
},
|
||||
{
|
||||
"rank": 2,
|
||||
"docId": "rag-chunk-context-reconstruction",
|
||||
"title": "RAG Chunk Context Reconstruction",
|
||||
"breadcrumb": "RAG > Chunking > Context Reconstruction",
|
||||
"content": "Long sections may require neighbor chunk expansion and breadcrumb-aware packing.",
|
||||
"score": 0.55,
|
||||
"retrievalLayer": "L1"
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "What is the standard troubleshooting flow for an application incident?",
|
||||
"rewrittenQuery": "standard application incident troubleshooting flow collect evidence verify remediation",
|
||||
"categoryFilter": "AIOps",
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["AIOps"],
|
||||
"matched_keywords": ["collect evidence", "verify", "remediation"],
|
||||
"entities": ["application incident"],
|
||||
"l0_titles": ["Incident Diagnosis Flow"],
|
||||
"l0_match_count": 1
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "standard application incident troubleshooting flow collect evidence verify remediation",
|
||||
"categoryFilter": "AIOps",
|
||||
"candidateCount": 2,
|
||||
"usable": true,
|
||||
"durationMs": 10,
|
||||
"topScore": 0.82,
|
||||
"topSimilarity": 0.82
|
||||
}
|
||||
]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "incident-diagnosis-flow",
|
||||
"baseScore": 0.82,
|
||||
"finalScore": 1.07,
|
||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"finalRank": 2,
|
||||
"source": "rag-chunk-context-reconstruction",
|
||||
"baseScore": 0.55,
|
||||
"finalScore": 0.55,
|
||||
"boostReasons": []
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,25 +1,81 @@
|
||||
{
|
||||
"caseId": "chat-l0-domain-hint",
|
||||
"query": "Should L0 keyword matching decide the final retrieval result?",
|
||||
"retrievedAt": "2026-07-05T00:00:00Z",
|
||||
"candidates": [
|
||||
{
|
||||
"rank": 1,
|
||||
"docId": "rag-l0-domain-entity-hint",
|
||||
"title": "RAG L0 Domain Entity Hint",
|
||||
"breadcrumb": "RAG > L0 > Domain Entity Hint",
|
||||
"content": "L0 should be retained as a domain detector, entity extractor, metadata filter generator, and explainability signal, not as the final retrieval decision.",
|
||||
"score": 0.88,
|
||||
"retrievalLayer": "L0"
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "rag-l0-domain-entity-hint",
|
||||
"title": "RAG L0 Domain Entity Hint",
|
||||
"breadcrumb": "RAG > L0 > Domain Entity Hint",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "L0 should be retained as a domain detector, entity extractor, metadata filter generator, and explainability signal, not as the final retrieval decision.",
|
||||
"score": 0.88,
|
||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"source": "rag-l0-l1-fusion-ranking",
|
||||
"title": "RAG L0 L1 Fusion Ranking",
|
||||
"breadcrumb": "RAG > Ranking > Fusion",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "L0 and L1 candidates should eventually be fused rather than handled as an early-return branch.",
|
||||
"score": 0.75,
|
||||
"hitReasons": ["domain_match:+0.15"]
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] RAG L0 Domain Entity Hint\nRAG > L0 > Domain Entity Hint\nL0 should be retained as a domain detector, entity extractor, metadata filter generator, and explainability signal, not as the final retrieval decision.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 219,
|
||||
"includedSources": ["rag-l0-domain-entity-hint", "rag-l0-l1-fusion-ranking"],
|
||||
"omittedSources": []
|
||||
},
|
||||
{
|
||||
"rank": 2,
|
||||
"docId": "rag-l0-l1-fusion-ranking",
|
||||
"title": "RAG L0 L1 Fusion Ranking",
|
||||
"breadcrumb": "RAG > Ranking > Fusion",
|
||||
"content": "L0 and L1 candidates should eventually be fused rather than handled as an early-return branch.",
|
||||
"score": 0.75,
|
||||
"retrievalLayer": "L1"
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "Should L0 keyword matching decide the final retrieval result?",
|
||||
"rewrittenQuery": "RAG L0 keyword matching domain entity hint final retrieval decision",
|
||||
"categoryFilter": "RAG",
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["RAG"],
|
||||
"matched_keywords": ["domain detector", "entity extractor", "metadata filter"],
|
||||
"entities": ["L0"],
|
||||
"l0_titles": ["RAG L0 Domain Entity Hint"],
|
||||
"l0_match_count": 1
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "RAG L0 keyword matching domain entity hint final retrieval decision",
|
||||
"categoryFilter": "RAG",
|
||||
"candidateCount": 2,
|
||||
"usable": true,
|
||||
"durationMs": 9,
|
||||
"topScore": 0.88,
|
||||
"topSimilarity": 0.88
|
||||
}
|
||||
]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "rag-l0-domain-entity-hint",
|
||||
"baseScore": 0.88,
|
||||
"finalScore": 1.13,
|
||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"finalRank": 2,
|
||||
"source": "rag-l0-l1-fusion-ranking",
|
||||
"baseScore": 0.75,
|
||||
"finalScore": 0.9,
|
||||
"boostReasons": ["domain_match:+0.15"]
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,91 @@
|
||||
{
|
||||
"caseId": "chat-l0-filter-fallback",
|
||||
"query": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "rag-l0-filter-fallback",
|
||||
"title": "RAG L0 Filter Fallback",
|
||||
"breadcrumb": "RAG > Fallback > Unfiltered Retry",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "When filtered vector retrieval is low quality, skip the L0 filter and run an unfiltered vector retry with the raw query before returning no evidence.",
|
||||
"score": 0.83,
|
||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"source": "rag-l0-domain-entity-hint",
|
||||
"title": "RAG L0 Domain Entity Hint",
|
||||
"breadcrumb": "RAG > L0 > Domain Entity Hint",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "L0 supplies hints for metadata filtering and explanation, but it should not be treated as final fact evidence.",
|
||||
"score": 0.66,
|
||||
"hitReasons": ["domain_match:+0.15"]
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] RAG L0 Filter Fallback\nRAG > Fallback > Unfiltered Retry\nWhen filtered vector retrieval is low quality, skip the L0 filter and run an unfiltered vector retry with the raw query before returning no evidence.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 214,
|
||||
"includedSources": ["rag-l0-filter-fallback", "rag-l0-domain-entity-hint"],
|
||||
"omittedSources": []
|
||||
},
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"rewrittenQuery": "RAG L0 filtered vector low quality fallback unfiltered retry",
|
||||
"categoryFilter": "RAG",
|
||||
"selectedAttempt": "UNFILTERED_VECTOR_RETRY",
|
||||
"fallbackReason": "filtered_vector_low_quality",
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["RAG"],
|
||||
"matched_keywords": ["L0", "low quality", "unfiltered vector retry"],
|
||||
"entities": ["L0"],
|
||||
"l0_titles": ["RAG L0 Domain Entity Hint"],
|
||||
"l0_match_count": 1
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "RAG L0 filtered vector low quality fallback unfiltered retry",
|
||||
"categoryFilter": "RAG",
|
||||
"candidateCount": 1,
|
||||
"usable": false,
|
||||
"durationMs": 7,
|
||||
"topScore": 1.35,
|
||||
"topSimilarity": 0.325
|
||||
},
|
||||
{
|
||||
"name": "UNFILTERED_VECTOR_RETRY",
|
||||
"query": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"categoryFilter": null,
|
||||
"candidateCount": 2,
|
||||
"usable": true,
|
||||
"durationMs": 13,
|
||||
"topScore": 0.83,
|
||||
"topSimilarity": 0.83
|
||||
}
|
||||
]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "rag-l0-filter-fallback",
|
||||
"baseScore": 0.83,
|
||||
"finalScore": 1.08,
|
||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"finalRank": 2,
|
||||
"source": "rag-l0-domain-entity-hint",
|
||||
"baseScore": 0.66,
|
||||
"finalScore": 0.81,
|
||||
"boostReasons": ["domain_match:+0.15"]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1,25 +1,81 @@
|
||||
{
|
||||
"caseId": "chat-mysql-connection-pool",
|
||||
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"retrievedAt": "2026-07-05T00:00:00Z",
|
||||
"candidates": [
|
||||
{
|
||||
"rank": 1,
|
||||
"docId": "mysql-connection-pool",
|
||||
"title": "MySQL Connection Pool Troubleshooting",
|
||||
"breadcrumb": "Database > MySQL > Connection Pool",
|
||||
"content": "When the connection pool is exhausted, inspect HikariCP active connections, max_connections, slow SQL, leak detection, and database wait events.",
|
||||
"score": 0.86,
|
||||
"retrievalLayer": "L0+L1"
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "mysql-connection-pool",
|
||||
"title": "MySQL Connection Pool Troubleshooting",
|
||||
"breadcrumb": "Database > MySQL > Connection Pool",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "When the connection pool is exhausted, inspect HikariCP active connections, max_connections, slow SQL, leak detection, and database wait events.",
|
||||
"score": 0.86,
|
||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"source": "incident-diagnosis-flow",
|
||||
"title": "Incident Diagnosis Flow",
|
||||
"breadcrumb": "AIOps > Diagnosis Flow",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "Collect evidence, compare metrics and logs, then verify remediation before closing the incident.",
|
||||
"score": 0.61,
|
||||
"hitReasons": []
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] MySQL Connection Pool Troubleshooting\nDatabase > MySQL > Connection Pool\nWhen the connection pool is exhausted, inspect HikariCP active connections, max_connections, slow SQL, leak detection, and database wait events.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 216,
|
||||
"includedSources": ["mysql-connection-pool", "incident-diagnosis-flow"],
|
||||
"omittedSources": []
|
||||
},
|
||||
{
|
||||
"rank": 2,
|
||||
"docId": "incident-diagnosis-flow",
|
||||
"title": "Incident Diagnosis Flow",
|
||||
"breadcrumb": "AIOps > Diagnosis Flow",
|
||||
"content": "Collect evidence, compare metrics and logs, then verify remediation before closing the incident.",
|
||||
"score": 0.61,
|
||||
"retrievalLayer": "L1"
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"rewrittenQuery": "MySQL connection pool exhausted HikariCP max_connections diagnosis",
|
||||
"categoryFilter": "Database",
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["Database", "MySQL"],
|
||||
"matched_keywords": ["connection pool", "HikariCP", "max_connections"],
|
||||
"entities": ["MySQL", "HikariCP"],
|
||||
"l0_titles": ["MySQL Connection Pool Troubleshooting"],
|
||||
"l0_match_count": 1
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "MySQL connection pool exhausted HikariCP max_connections diagnosis",
|
||||
"categoryFilter": "Database",
|
||||
"candidateCount": 2,
|
||||
"usable": true,
|
||||
"durationMs": 12,
|
||||
"topScore": 0.86,
|
||||
"topSimilarity": 0.86
|
||||
}
|
||||
]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "mysql-connection-pool",
|
||||
"baseScore": 0.86,
|
||||
"finalScore": 1.11,
|
||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"finalRank": 2,
|
||||
"source": "incident-diagnosis-flow",
|
||||
"baseScore": 0.61,
|
||||
"finalScore": 0.61,
|
||||
"boostReasons": []
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,25 +1,81 @@
|
||||
{
|
||||
"caseId": "chat-rag-chunk-context",
|
||||
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"retrievedAt": "2026-07-05T00:00:00Z",
|
||||
"candidates": [
|
||||
{
|
||||
"rank": 1,
|
||||
"docId": "rag-chunk-context-reconstruction",
|
||||
"title": "RAG Chunk Context Reconstruction",
|
||||
"breadcrumb": "RAG > Chunking > Context Reconstruction",
|
||||
"content": "After a chunk hit, expand to neighbor chunk candidates from the same section and preserve breadcrumb metadata in the evidence pack.",
|
||||
"score": 0.79,
|
||||
"retrievalLayer": "L1"
|
||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
||||
"lookupResult": {
|
||||
"found": true,
|
||||
"evidenceBlocks": [
|
||||
{
|
||||
"source": "rag-chunk-context-reconstruction",
|
||||
"title": "RAG Chunk Context Reconstruction",
|
||||
"breadcrumb": "RAG > Chunking > Context Reconstruction",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "After a chunk hit, expand to neighbor chunk candidates from the same section and preserve breadcrumb metadata in the evidence pack.",
|
||||
"score": 0.79,
|
||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"source": "rag-breadcrumb-embedding-gap",
|
||||
"title": "RAG Breadcrumb Embedding Gap",
|
||||
"breadcrumb": "RAG > Embedding > Breadcrumb",
|
||||
"retrievalLayer": "L1",
|
||||
"content": "Embedding title and breadcrumb with content helps recover section semantics.",
|
||||
"score": 0.72,
|
||||
"hitReasons": ["domain_match:+0.15"]
|
||||
}
|
||||
],
|
||||
"contextPack": {
|
||||
"packedText": "[1] RAG Chunk Context Reconstruction\nRAG > Chunking > Context Reconstruction\nAfter a chunk hit, expand to neighbor chunk candidates from the same section and preserve breadcrumb metadata in the evidence pack.",
|
||||
"strategy": "top_evidence_blocks",
|
||||
"charBudget": 3500,
|
||||
"usedChars": 203,
|
||||
"includedSources": ["rag-chunk-context-reconstruction", "rag-breadcrumb-embedding-gap"],
|
||||
"omittedSources": []
|
||||
},
|
||||
{
|
||||
"rank": 2,
|
||||
"docId": "rag-breadcrumb-embedding-gap",
|
||||
"title": "RAG Breadcrumb Embedding Gap",
|
||||
"breadcrumb": "RAG > Embedding > Breadcrumb",
|
||||
"content": "Embedding title and breadcrumb with content helps recover section semantics.",
|
||||
"score": 0.72,
|
||||
"retrievalLayer": "L1"
|
||||
"retrievalTrace": {
|
||||
"originalQuery": "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"rewrittenQuery": "RAG chunk context reconstruction neighbor chunk same section breadcrumb",
|
||||
"categoryFilter": "RAG",
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"queryHints": {
|
||||
"domains": ["RAG"],
|
||||
"matched_keywords": ["neighbor chunk", "same section", "breadcrumb"],
|
||||
"entities": ["chunk", "breadcrumb"],
|
||||
"l0_titles": ["RAG Chunk Context Reconstruction"],
|
||||
"l0_match_count": 1
|
||||
},
|
||||
"attempts": [
|
||||
{
|
||||
"name": "FILTERED_VECTOR",
|
||||
"query": "RAG chunk context reconstruction neighbor chunk same section breadcrumb",
|
||||
"categoryFilter": "RAG",
|
||||
"candidateCount": 2,
|
||||
"usable": true,
|
||||
"durationMs": 9,
|
||||
"topScore": 0.79,
|
||||
"topSimilarity": 0.79
|
||||
}
|
||||
]
|
||||
},
|
||||
"rerankTrace": {
|
||||
"items": [
|
||||
{
|
||||
"finalRank": 1,
|
||||
"source": "rag-chunk-context-reconstruction",
|
||||
"baseScore": 0.79,
|
||||
"finalScore": 1.04,
|
||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
||||
},
|
||||
{
|
||||
"finalRank": 2,
|
||||
"source": "rag-breadcrumb-embedding-gap",
|
||||
"baseScore": 0.72,
|
||||
"finalScore": 0.87,
|
||||
"boostReasons": ["domain_match:+0.15"]
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,11 +1,15 @@
|
||||
{
|
||||
"generatedAt": "2026-07-04T17:59:52.172759+00:00",
|
||||
"generatedAt": "2026-07-06T13:37:59.726351+00:00",
|
||||
"caseFile": "eval/rag-retrieval/cases/golden-cases.json",
|
||||
"fixtureDir": "eval/rag-retrieval/fixtures",
|
||||
"aggregate": {
|
||||
"caseCount": 6,
|
||||
"caseCount": 7,
|
||||
"topK": 5,
|
||||
"strongHitCount": 6,
|
||||
"passedCount": 7,
|
||||
"failedCount": 0,
|
||||
"passRate": 1.0,
|
||||
"lookupResultCaseCount": 7,
|
||||
"strongHitCount": 7,
|
||||
"mediumHitCount": 0,
|
||||
"weakHitCount": 0,
|
||||
"missCount": 0,
|
||||
@@ -18,6 +22,7 @@
|
||||
"caseId": "chat-mysql-connection-pool",
|
||||
"scenario": "chat",
|
||||
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
@@ -31,12 +36,22 @@
|
||||
"hikaricp"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"mysql-connection-pool",
|
||||
"incident-diagnosis-flow"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "mysql-connection-pool",
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "chat-diagnosis-flow",
|
||||
"scenario": "chat",
|
||||
"query": "What is the standard troubleshooting flow for an application incident?",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
@@ -50,12 +65,22 @@
|
||||
"remediation"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"incident-diagnosis-flow",
|
||||
"rag-chunk-context-reconstruction"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "incident-diagnosis-flow",
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "aiops-payment-latency-alert",
|
||||
"scenario": "aiops",
|
||||
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
@@ -69,12 +94,22 @@
|
||||
"downstream dependency"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"payment-service-latency",
|
||||
"mysql-connection-pool"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "payment-service-latency",
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "aiops-prometheus-alert-scope",
|
||||
"scenario": "aiops",
|
||||
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
@@ -87,12 +122,21 @@
|
||||
"scope"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"aiops-alert-scope-control"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "aiops-alert-scope-control",
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "chat-rag-chunk-context",
|
||||
"scenario": "chat",
|
||||
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
@@ -106,12 +150,22 @@
|
||||
"breadcrumb"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"rag-chunk-context-reconstruction",
|
||||
"rag-breadcrumb-embedding-gap"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "rag-chunk-context-reconstruction",
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "chat-l0-domain-hint",
|
||||
"scenario": "chat",
|
||||
"query": "Should L0 keyword matching decide the final retrieval result?",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
@@ -125,6 +179,44 @@
|
||||
"metadata filter"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "FILTERED_VECTOR",
|
||||
"fallbackReason": null,
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"rag-l0-domain-entity-hint",
|
||||
"rag-l0-l1-fusion-ranking"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "rag-l0-domain-entity-hint",
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "chat-l0-filter-fallback",
|
||||
"scenario": "chat",
|
||||
"query": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||
"dataShape": "lookupResult",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:rag-l0-filter-fallback",
|
||||
"2:rag-l0-domain-entity-hint"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"skip the l0 filter",
|
||||
"unfiltered vector retry",
|
||||
"low quality"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"selectedAttempt": "UNFILTERED_VECTOR_RETRY",
|
||||
"fallbackReason": "filtered_vector_low_quality",
|
||||
"evidenceStatus": "supported",
|
||||
"includedSources": [
|
||||
"rag-l0-filter-fallback",
|
||||
"rag-l0-domain-entity-hint"
|
||||
],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": "rag-l0-filter-fallback",
|
||||
"failedChecks": []
|
||||
}
|
||||
]
|
||||
|
||||
@@ -1,16 +1,20 @@
|
||||
# RAG Retrieval Baseline
|
||||
|
||||
Generated at: `2026-07-04T17:59:52.172759+00:00`
|
||||
Generated at: `2026-07-06T13:37:59.726351+00:00`
|
||||
|
||||
## Aggregate
|
||||
|
||||
| Metric | Value |
|
||||
|---|---:|
|
||||
| Cases | 6 |
|
||||
| Cases | 7 |
|
||||
| Top K | 5 |
|
||||
| Passed | 7 |
|
||||
| Failed | 0 |
|
||||
| Pass rate | 1.0 |
|
||||
| LookupResult fixtures | 7 |
|
||||
| Recall@K | 1.0 |
|
||||
| Strong hit rate | 1.0 |
|
||||
| Strong hits | 6 |
|
||||
| Strong hits | 7 |
|
||||
| Medium hits | 0 |
|
||||
| Weak hits | 0 |
|
||||
| Misses | 0 |
|
||||
@@ -18,11 +22,12 @@ Generated at: `2026-07-04T17:59:52.172759+00:00`
|
||||
|
||||
## Cases
|
||||
|
||||
| Case | Scenario | Hit | First Expected Rank | Top Candidates | Failed Checks |
|
||||
|---|---|---|---:|---|---|
|
||||
| chat-mysql-connection-pool | chat | strong | 1 | 1:mysql-connection-pool<br>2:incident-diagnosis-flow | |
|
||||
| chat-diagnosis-flow | chat | strong | 1 | 1:incident-diagnosis-flow<br>2:rag-chunk-context-reconstruction | |
|
||||
| aiops-payment-latency-alert | aiops | strong | 1 | 1:payment-service-latency<br>2:mysql-connection-pool | |
|
||||
| aiops-prometheus-alert-scope | aiops | strong | 1 | 1:aiops-alert-scope-control | |
|
||||
| chat-rag-chunk-context | chat | strong | 1 | 1:rag-chunk-context-reconstruction<br>2:rag-breadcrumb-embedding-gap | |
|
||||
| chat-l0-domain-hint | chat | strong | 1 | 1:rag-l0-domain-entity-hint<br>2:rag-l0-l1-fusion-ranking | |
|
||||
| Case | Scenario | Pass | Hit | Attempt | Fallback | Evidence | First Expected Rank | Top Candidates | Failed Checks |
|
||||
|---|---|---|---|---|---|---|---:|---|---|
|
||||
| chat-mysql-connection-pool | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:mysql-connection-pool<br>2:incident-diagnosis-flow | |
|
||||
| chat-diagnosis-flow | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:incident-diagnosis-flow<br>2:rag-chunk-context-reconstruction | |
|
||||
| aiops-payment-latency-alert | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:payment-service-latency<br>2:mysql-connection-pool | |
|
||||
| aiops-prometheus-alert-scope | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:aiops-alert-scope-control | |
|
||||
| chat-rag-chunk-context | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-chunk-context-reconstruction<br>2:rag-breadcrumb-embedding-gap | |
|
||||
| chat-l0-domain-hint | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-l0-domain-entity-hint<br>2:rag-l0-l1-fusion-ranking | |
|
||||
| chat-l0-filter-fallback | chat | true | strong | UNFILTERED_VECTOR_RETRY | filtered_vector_low_quality | supported | 1 | 1:rag-l0-filter-fallback<br>2:rag-l0-domain-entity-hint | |
|
||||
|
||||
@@ -0,0 +1,26 @@
|
||||
---
|
||||
title: AIOps Alert Scope Control
|
||||
keywords: [alert payload, unrelated active alerts, scope control]
|
||||
summary: Keep diagnosis scoped to the request payload and avoid diagnosing unrelated active alerts.
|
||||
category: aiops
|
||||
source: aiops-alert-scope-control
|
||||
breadcrumb: AIOps > Alert Scope Control
|
||||
kb_scope: rag-eval
|
||||
covers: [alert scope, payload, active alerts]
|
||||
when_to_retrieve: Use when an AIOps request includes a concrete alert payload and scope boundaries matter.
|
||||
---
|
||||
|
||||
# AIOps
|
||||
|
||||
## Alert Scope Control
|
||||
|
||||
When an AIOps request already includes an alert payload, the agent should diagnose that payload first.
|
||||
It must not expand the task into unrelated active alerts unless the user asks for broad alert triage.
|
||||
|
||||
Scope rules:
|
||||
|
||||
1. Treat the provided payload as the primary incident boundary.
|
||||
2. Use unrelated active alerts only as correlation evidence when they share service, dependency, time window, or trace context.
|
||||
3. Do not replace the requested alert with a louder but unrelated alert.
|
||||
|
||||
This runbook anchors payload, unrelated active alerts, and scope behavior.
|
||||
@@ -0,0 +1,27 @@
|
||||
---
|
||||
title: Incident Diagnosis Flow
|
||||
keywords: [standard troubleshooting flow, application incident, collect evidence, verify, remediation]
|
||||
summary: Standard flow for diagnosing application incidents with evidence, hypothesis verification, and remediation.
|
||||
category: ops
|
||||
source: incident-diagnosis-flow
|
||||
breadcrumb: AIOps > Diagnosis Flow
|
||||
kb_scope: rag-eval
|
||||
covers: [incident diagnosis, evidence collection, remediation]
|
||||
when_to_retrieve: Use when the user asks for a standard troubleshooting flow or incident diagnosis sequence.
|
||||
---
|
||||
|
||||
# AIOps
|
||||
|
||||
## Diagnosis Flow
|
||||
|
||||
The standard troubleshooting flow is evidence first, hypothesis second, remediation last.
|
||||
|
||||
Recommended sequence:
|
||||
|
||||
1. Collect evidence from alerts, metrics, logs, traces, deployments, and recent configuration changes.
|
||||
2. Define a small hypothesis that explains the observed symptoms.
|
||||
3. Verify the hypothesis with a targeted metric, log query, or reproduction step.
|
||||
4. Choose remediation that directly addresses the verified cause.
|
||||
5. Record the outcome and the evidence used to make the decision.
|
||||
|
||||
Do not skip collect evidence, verify, and remediation ordering during an application incident.
|
||||
@@ -0,0 +1,29 @@
|
||||
---
|
||||
title: MySQL Connection Pool Runbook
|
||||
keywords: [MySQL connection pool, pool exhausted, max_connections, HikariCP]
|
||||
summary: Diagnose exhausted MySQL connection pools and distinguish application leaks from database limits.
|
||||
category: database
|
||||
source: mysql-connection-pool
|
||||
breadcrumb: Database > MySQL > Connection Pool
|
||||
kb_scope: rag-eval
|
||||
covers: [mysql, connection pool, database capacity]
|
||||
when_to_retrieve: Use when MySQL clients report exhausted pools, connection acquisition timeout, max_connections pressure, or HikariCP saturation.
|
||||
---
|
||||
|
||||
# Database
|
||||
|
||||
## MySQL
|
||||
|
||||
### Connection Pool
|
||||
|
||||
When MySQL connection pool is exhausted, first compare application pool usage with database `max_connections`.
|
||||
For HikariCP, check `active`, `idle`, `pending`, and connection acquisition timeout metrics.
|
||||
|
||||
Recommended diagnosis:
|
||||
|
||||
1. Verify whether HikariCP active connections stay near maximum while pending threads grow.
|
||||
2. Check MySQL `Threads_connected`, `Threads_running`, and `max_connections`.
|
||||
3. Inspect slow SQL and long transactions that keep connections checked out.
|
||||
4. If the database is healthy, look for application connection leaks or missing transaction boundaries.
|
||||
|
||||
Use this runbook as evidence for connection pool, max_connections, and HikariCP incidents.
|
||||
@@ -0,0 +1,28 @@
|
||||
---
|
||||
title: Payment Service Latency Alert
|
||||
keywords: [HighLatency, payment-service, p95 latency, downstream dependency]
|
||||
summary: Diagnose payment-service p95 latency alerts and identify downstream dependency bottlenecks.
|
||||
category: aiops
|
||||
source: payment-service-latency
|
||||
breadcrumb: AIOps > Service Alerts > Payment Latency
|
||||
kb_scope: rag-eval
|
||||
covers: [payment-service, latency, downstream dependency]
|
||||
when_to_retrieve: Use when an alert mentions payment-service, HighLatency, or elevated p95 latency.
|
||||
---
|
||||
|
||||
# AIOps
|
||||
|
||||
## Service Alerts
|
||||
|
||||
### Payment Latency
|
||||
|
||||
For `HighLatency` alerts on `payment-service`, treat p95 latency as the primary symptom.
|
||||
|
||||
Diagnosis steps:
|
||||
|
||||
1. Confirm whether p95 latency is isolated to payment-service or shared across upstream callers.
|
||||
2. Compare payment-service latency with downstream dependency latency for gateway, risk, and order services.
|
||||
3. Check connection pool wait time, retry spikes, and timeout rates.
|
||||
4. If downstream dependency latency increased first, classify payment-service as affected rather than root cause.
|
||||
|
||||
The expected evidence terms are p95 latency, payment-service, and downstream dependency.
|
||||
@@ -0,0 +1,28 @@
|
||||
---
|
||||
title: RAG Chunk Context Reconstruction
|
||||
keywords: [split into multiple chunks, retrieval context, neighbor chunk, same section, breadcrumb context]
|
||||
summary: Preserve context when long RAG sections are split into multiple retrievable chunks.
|
||||
category: rag
|
||||
source: rag-chunk-context-reconstruction
|
||||
breadcrumb: RAG > Chunking > Context Reconstruction
|
||||
kb_scope: rag-eval
|
||||
covers: [rag chunking, context packing, breadcrumbs]
|
||||
when_to_retrieve: Use when a retrieval question asks how to preserve context across split chunks.
|
||||
---
|
||||
|
||||
# RAG
|
||||
|
||||
## Chunking
|
||||
|
||||
### Context Reconstruction
|
||||
|
||||
When a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.
|
||||
|
||||
Recommended behavior:
|
||||
|
||||
1. Store the breadcrumb with every chunk.
|
||||
2. Preserve the same section identity across adjacent chunks.
|
||||
3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.
|
||||
4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.
|
||||
|
||||
The key concepts are neighbor chunk, same section, and breadcrumb.
|
||||
@@ -0,0 +1,29 @@
|
||||
---
|
||||
title: RAG L0 Domain Entity Hint
|
||||
keywords: [L0 keyword matching, final retrieval result, domain detector, entity extractor, metadata filter]
|
||||
summary: Define L0 as a query transformation hint layer instead of final retrieval evidence.
|
||||
category: rag
|
||||
source: rag-l0-domain-entity-hint
|
||||
breadcrumb: RAG > L0 > Domain Entity Hint
|
||||
kb_scope: rag-eval
|
||||
covers: [l0 hint, query transformation, metadata filter]
|
||||
when_to_retrieve: Use when a question asks whether L0 should decide final retrieval or only provide hints.
|
||||
---
|
||||
|
||||
# RAG
|
||||
|
||||
## L0
|
||||
|
||||
### Domain Entity Hint
|
||||
|
||||
L0 keyword matching should not decide the final retrieval result.
|
||||
In the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.
|
||||
|
||||
The output can provide:
|
||||
|
||||
1. Candidate domain hints.
|
||||
2. Matched entities and keywords.
|
||||
3. An optional metadata filter for the first vector retrieval attempt.
|
||||
|
||||
Final evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.
|
||||
The important terms are domain detector, entity extractor, and metadata filter.
|
||||
@@ -0,0 +1,19 @@
|
||||
---
|
||||
title: RAG L0 Filter Decoy
|
||||
keywords: [over-filtered by L0, filtered vector search, low quality evidence]
|
||||
summary: Decoy document used to force the first filtered retrieval attempt into a low-quality category.
|
||||
category: overfilter-decoy
|
||||
source: rag-l0-filter-decoy
|
||||
breadcrumb: RAG > Fallback > Decoy
|
||||
kb_scope: rag-eval
|
||||
covers: [fallback test decoy]
|
||||
when_to_retrieve: Use only as a controlled eval decoy for over-filter fallback testing.
|
||||
---
|
||||
|
||||
# Release Calendar
|
||||
|
||||
## Approval Window
|
||||
|
||||
This document describes an unrelated release calendar approval window.
|
||||
It intentionally avoids the real fallback instructions so the filtered retrieval
|
||||
attempt is low quality and the retriever must retry without the L0 category filter.
|
||||
@@ -0,0 +1,25 @@
|
||||
---
|
||||
title: RAG L0 Filter Fallback
|
||||
keywords: [golden retry contract, second pass retrieval]
|
||||
summary: Retry the raw query without the L0 category filter when filtered vector evidence is missing or low quality.
|
||||
category: fallback
|
||||
source: rag-l0-filter-fallback
|
||||
breadcrumb: RAG > Fallback > Unfiltered Retry
|
||||
kb_scope: rag-eval
|
||||
covers: [fallback, unfiltered retry, retrieval quality]
|
||||
when_to_retrieve: Use when validating the fallback contract for low-quality filtered vector retrieval.
|
||||
---
|
||||
|
||||
# RAG
|
||||
|
||||
## Fallback
|
||||
|
||||
### Unfiltered Retry
|
||||
|
||||
If the first vector search is over-constrained by an L0 metadata filter and returns low quality evidence,
|
||||
the retriever should skip the L0 filter and run an unfiltered vector retry with the original query.
|
||||
|
||||
The fallback reason should be `filtered_vector_low_quality` when the filtered candidate exists but is below the
|
||||
reference threshold. If there is no usable evidence at all, use `filtered_vector_no_evidence`.
|
||||
|
||||
This document is the expected evidence for skip the L0 filter, unfiltered vector retry, and low quality behavior.
|
||||
@@ -0,0 +1,18 @@
|
||||
# RAG Eval Knowledge Base Mirror
|
||||
|
||||
This folder stores the committed knowledge-base copy of the canonical RAG eval
|
||||
documents.
|
||||
|
||||
The source of truth for the eval importer remains:
|
||||
|
||||
```text
|
||||
eval/rag-retrieval/seed-docs/
|
||||
```
|
||||
|
||||
The documents are kept under `knowledge_base/rag-eval/` so eval/test knowledge
|
||||
does not mix with the normal business knowledge folders such as `api`,
|
||||
`infrastructure`, or `troubleshooting`.
|
||||
|
||||
Each document keeps its original frontmatter `category` and `kb_scope`. The
|
||||
category is still the retrieval category used by L0/L1, while `kb_scope:
|
||||
rag-eval` isolates these documents during eval runs.
|
||||
@@ -0,0 +1,26 @@
|
||||
---
|
||||
title: AIOps Alert Scope Control
|
||||
keywords: [alert payload, unrelated active alerts, scope control]
|
||||
summary: Keep diagnosis scoped to the request payload and avoid diagnosing unrelated active alerts.
|
||||
category: aiops
|
||||
source: aiops-alert-scope-control
|
||||
breadcrumb: AIOps > Alert Scope Control
|
||||
kb_scope: rag-eval
|
||||
covers: [alert scope, payload, active alerts]
|
||||
when_to_retrieve: Use when an AIOps request includes a concrete alert payload and scope boundaries matter.
|
||||
---
|
||||
|
||||
# AIOps
|
||||
|
||||
## Alert Scope Control
|
||||
|
||||
When an AIOps request already includes an alert payload, the agent should diagnose that payload first.
|
||||
It must not expand the task into unrelated active alerts unless the user asks for broad alert triage.
|
||||
|
||||
Scope rules:
|
||||
|
||||
1. Treat the provided payload as the primary incident boundary.
|
||||
2. Use unrelated active alerts only as correlation evidence when they share service, dependency, time window, or trace context.
|
||||
3. Do not replace the requested alert with a louder but unrelated alert.
|
||||
|
||||
This runbook anchors payload, unrelated active alerts, and scope behavior.
|
||||
@@ -0,0 +1,28 @@
|
||||
---
|
||||
title: Payment Service Latency Alert
|
||||
keywords: [HighLatency, payment-service, p95 latency, downstream dependency]
|
||||
summary: Diagnose payment-service p95 latency alerts and identify downstream dependency bottlenecks.
|
||||
category: aiops
|
||||
source: payment-service-latency
|
||||
breadcrumb: AIOps > Service Alerts > Payment Latency
|
||||
kb_scope: rag-eval
|
||||
covers: [payment-service, latency, downstream dependency]
|
||||
when_to_retrieve: Use when an alert mentions payment-service, HighLatency, or elevated p95 latency.
|
||||
---
|
||||
|
||||
# AIOps
|
||||
|
||||
## Service Alerts
|
||||
|
||||
### Payment Latency
|
||||
|
||||
For `HighLatency` alerts on `payment-service`, treat p95 latency as the primary symptom.
|
||||
|
||||
Diagnosis steps:
|
||||
|
||||
1. Confirm whether p95 latency is isolated to payment-service or shared across upstream callers.
|
||||
2. Compare payment-service latency with downstream dependency latency for gateway, risk, and order services.
|
||||
3. Check connection pool wait time, retry spikes, and timeout rates.
|
||||
4. If downstream dependency latency increased first, classify payment-service as affected rather than root cause.
|
||||
|
||||
The expected evidence terms are p95 latency, payment-service, and downstream dependency.
|
||||
@@ -0,0 +1,29 @@
|
||||
---
|
||||
title: MySQL Connection Pool Runbook
|
||||
keywords: [MySQL connection pool, pool exhausted, max_connections, HikariCP]
|
||||
summary: Diagnose exhausted MySQL connection pools and distinguish application leaks from database limits.
|
||||
category: database
|
||||
source: mysql-connection-pool
|
||||
breadcrumb: Database > MySQL > Connection Pool
|
||||
kb_scope: rag-eval
|
||||
covers: [mysql, connection pool, database capacity]
|
||||
when_to_retrieve: Use when MySQL clients report exhausted pools, connection acquisition timeout, max_connections pressure, or HikariCP saturation.
|
||||
---
|
||||
|
||||
# Database
|
||||
|
||||
## MySQL
|
||||
|
||||
### Connection Pool
|
||||
|
||||
When MySQL connection pool is exhausted, first compare application pool usage with database `max_connections`.
|
||||
For HikariCP, check `active`, `idle`, `pending`, and connection acquisition timeout metrics.
|
||||
|
||||
Recommended diagnosis:
|
||||
|
||||
1. Verify whether HikariCP active connections stay near maximum while pending threads grow.
|
||||
2. Check MySQL `Threads_connected`, `Threads_running`, and `max_connections`.
|
||||
3. Inspect slow SQL and long transactions that keep connections checked out.
|
||||
4. If the database is healthy, look for application connection leaks or missing transaction boundaries.
|
||||
|
||||
Use this runbook as evidence for connection pool, max_connections, and HikariCP incidents.
|
||||
@@ -0,0 +1,25 @@
|
||||
---
|
||||
title: RAG L0 Filter Fallback
|
||||
keywords: [golden retry contract, second pass retrieval]
|
||||
summary: Retry the raw query without the L0 category filter when filtered vector evidence is missing or low quality.
|
||||
category: fallback
|
||||
source: rag-l0-filter-fallback
|
||||
breadcrumb: RAG > Fallback > Unfiltered Retry
|
||||
kb_scope: rag-eval
|
||||
covers: [fallback, unfiltered retry, retrieval quality]
|
||||
when_to_retrieve: Use when validating the fallback contract for low-quality filtered vector retrieval.
|
||||
---
|
||||
|
||||
# RAG
|
||||
|
||||
## Fallback
|
||||
|
||||
### Unfiltered Retry
|
||||
|
||||
If the first vector search is over-constrained by an L0 metadata filter and returns low quality evidence,
|
||||
the retriever should skip the L0 filter and run an unfiltered vector retry with the original query.
|
||||
|
||||
The fallback reason should be `filtered_vector_low_quality` when the filtered candidate exists but is below the
|
||||
reference threshold. If there is no usable evidence at all, use `filtered_vector_no_evidence`.
|
||||
|
||||
This document is the expected evidence for skip the L0 filter, unfiltered vector retry, and low quality behavior.
|
||||
@@ -0,0 +1,27 @@
|
||||
---
|
||||
title: Incident Diagnosis Flow
|
||||
keywords: [standard troubleshooting flow, application incident, collect evidence, verify, remediation]
|
||||
summary: Standard flow for diagnosing application incidents with evidence, hypothesis verification, and remediation.
|
||||
category: ops
|
||||
source: incident-diagnosis-flow
|
||||
breadcrumb: AIOps > Diagnosis Flow
|
||||
kb_scope: rag-eval
|
||||
covers: [incident diagnosis, evidence collection, remediation]
|
||||
when_to_retrieve: Use when the user asks for a standard troubleshooting flow or incident diagnosis sequence.
|
||||
---
|
||||
|
||||
# AIOps
|
||||
|
||||
## Diagnosis Flow
|
||||
|
||||
The standard troubleshooting flow is evidence first, hypothesis second, remediation last.
|
||||
|
||||
Recommended sequence:
|
||||
|
||||
1. Collect evidence from alerts, metrics, logs, traces, deployments, and recent configuration changes.
|
||||
2. Define a small hypothesis that explains the observed symptoms.
|
||||
3. Verify the hypothesis with a targeted metric, log query, or reproduction step.
|
||||
4. Choose remediation that directly addresses the verified cause.
|
||||
5. Record the outcome and the evidence used to make the decision.
|
||||
|
||||
Do not skip collect evidence, verify, and remediation ordering during an application incident.
|
||||
@@ -0,0 +1,19 @@
|
||||
---
|
||||
title: RAG L0 Filter Decoy
|
||||
keywords: [over-filtered by L0, filtered vector search, low quality evidence]
|
||||
summary: Decoy document used to force the first filtered retrieval attempt into a low-quality category.
|
||||
category: overfilter-decoy
|
||||
source: rag-l0-filter-decoy
|
||||
breadcrumb: RAG > Fallback > Decoy
|
||||
kb_scope: rag-eval
|
||||
covers: [fallback test decoy]
|
||||
when_to_retrieve: Use only as a controlled eval decoy for over-filter fallback testing.
|
||||
---
|
||||
|
||||
# Release Calendar
|
||||
|
||||
## Approval Window
|
||||
|
||||
This document describes an unrelated release calendar approval window.
|
||||
It intentionally avoids the real fallback instructions so the filtered retrieval
|
||||
attempt is low quality and the retriever must retry without the L0 category filter.
|
||||
@@ -0,0 +1,28 @@
|
||||
---
|
||||
title: RAG Chunk Context Reconstruction
|
||||
keywords: [split into multiple chunks, retrieval context, neighbor chunk, same section, breadcrumb context]
|
||||
summary: Preserve context when long RAG sections are split into multiple retrievable chunks.
|
||||
category: rag
|
||||
source: rag-chunk-context-reconstruction
|
||||
breadcrumb: RAG > Chunking > Context Reconstruction
|
||||
kb_scope: rag-eval
|
||||
covers: [rag chunking, context packing, breadcrumbs]
|
||||
when_to_retrieve: Use when a retrieval question asks how to preserve context across split chunks.
|
||||
---
|
||||
|
||||
# RAG
|
||||
|
||||
## Chunking
|
||||
|
||||
### Context Reconstruction
|
||||
|
||||
When a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.
|
||||
|
||||
Recommended behavior:
|
||||
|
||||
1. Store the breadcrumb with every chunk.
|
||||
2. Preserve the same section identity across adjacent chunks.
|
||||
3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.
|
||||
4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.
|
||||
|
||||
The key concepts are neighbor chunk, same section, and breadcrumb.
|
||||
@@ -0,0 +1,29 @@
|
||||
---
|
||||
title: RAG L0 Domain Entity Hint
|
||||
keywords: [L0 keyword matching, final retrieval result, domain detector, entity extractor, metadata filter]
|
||||
summary: Define L0 as a query transformation hint layer instead of final retrieval evidence.
|
||||
category: rag
|
||||
source: rag-l0-domain-entity-hint
|
||||
breadcrumb: RAG > L0 > Domain Entity Hint
|
||||
kb_scope: rag-eval
|
||||
covers: [l0 hint, query transformation, metadata filter]
|
||||
when_to_retrieve: Use when a question asks whether L0 should decide final retrieval or only provide hints.
|
||||
---
|
||||
|
||||
# RAG
|
||||
|
||||
## L0
|
||||
|
||||
### Domain Entity Hint
|
||||
|
||||
L0 keyword matching should not decide the final retrieval result.
|
||||
In the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.
|
||||
|
||||
The output can provide:
|
||||
|
||||
1. Candidate domain hints.
|
||||
2. Matched entities and keywords.
|
||||
3. An optional metadata filter for the first vector retrieval attempt.
|
||||
|
||||
Final evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.
|
||||
The important terms are domain detector, entity extractor, and metadata filter.
|
||||
@@ -1,6 +1,6 @@
|
||||
# MVP 架构文档
|
||||
|
||||
**更新日期**:2026-07-05
|
||||
**更新日期**:2026-07-06
|
||||
|
||||
这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到:
|
||||
|
||||
@@ -17,6 +17,8 @@
|
||||
| [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat SequentialAgent、AIOps SupervisorAgent、工具边界 |
|
||||
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Verifier、评测基线组成的质量门禁 |
|
||||
| [rag-architecture.md](rag-architecture.md) | RAG/知识检索新架构,覆盖 L0 hint、VectorStore 主路径、SDK fallback、证据追踪 |
|
||||
| [modular-rag-pipeline.md](modular-rag-pipeline.md) | `lookup_knowledge` 模块化 RAG 落地架构,覆盖 pipeline、fallback、evidence-first contract、trace |
|
||||
| [rag-eval-closure.md](rag-eval-closure.md) | RAG 评测闭环,覆盖 offline baseline、baseline diff、diagnosis eval 和 live acceptance |
|
||||
| [retrieval-observability.md](retrieval-observability.md) | 检索运行细节和可观测性,覆盖 L0/L1、去重、分数归一、评测 |
|
||||
| [feedback-architecture.md](feedback-architecture.md) | 反馈与自评估闭环,覆盖 rule evaluation、Verifier、AIOps rule、用户反馈和案例沉淀 |
|
||||
| [session-trace-lifecycle.md](session-trace-lifecycle.md) | 会话和 Trace 生命周期,覆盖 sessionId、状态流转、agent_step、tool_invocation、Trace API |
|
||||
@@ -35,8 +37,10 @@ SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat
|
||||
3. 再读 [agent-orchestration.md](agent-orchestration.md),理解当前 Agent 如何协作。
|
||||
4. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。
|
||||
5. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。
|
||||
6. 继续读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
|
||||
7. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
|
||||
8. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
|
||||
9. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
|
||||
10. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
|
||||
6. 继续读 [modular-rag-pipeline.md](modular-rag-pipeline.md),看 `lookup_knowledge` 的模块化落地和 evidence-first contract。
|
||||
7. 再读 [rag-eval-closure.md](rag-eval-closure.md),看 RAG baseline 如何形成质量闭环。
|
||||
8. 然后读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
|
||||
9. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
|
||||
10. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
|
||||
11. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
|
||||
12. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
|
||||
|
||||
@@ -0,0 +1,201 @@
|
||||
# 模块化 RAG Pipeline 架构
|
||||
|
||||
**更新日期**:2026-07-06
|
||||
**状态**:当前已实现架构
|
||||
**关联 OpenSpec**:`openspec/changes/archive/2026-07-06-modular-rag-pipeline`
|
||||
|
||||
## 1. 定位
|
||||
|
||||
本文记录 `lookup_knowledge` 的当前模块化 RAG 实现。它是 [rag-architecture.md](rag-architecture.md) 的落地版,重点说明代码模块、数据契约、降级策略和可观测性边界。
|
||||
|
||||
核心目标:
|
||||
|
||||
- 保留显式 Agent Tool:`lookup_knowledge(query)`。
|
||||
- L0 只作为 query understanding / filter / rerank / trace hint。
|
||||
- L1 向量检索作为事实证据来源。
|
||||
- filtered L1 低质量时,降级为 raw query unfiltered L1 retry。
|
||||
- 输出 evidence-first contract,替代旧 `primary/supplement`。
|
||||
|
||||
## 2. 当前链路
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Agent["Executor Agent"] --> Tool["LookupKnowledgeTool.lookupKnowledge(query)"]
|
||||
|
||||
Tool --> Transform["KnowledgeQueryTransformer"]
|
||||
Transform --> KQ["KnowledgeQuery"]
|
||||
|
||||
KQ --> Retriever["KnowledgeDocumentRetriever"]
|
||||
Retriever --> Attempt1["FILTERED_VECTOR or UNFILTERED_VECTOR"]
|
||||
Attempt1 --> Post1["KnowledgeEvidencePostProcessor"]
|
||||
Post1 --> Quality{"usable evidence?"}
|
||||
Quality -->|yes| Pack
|
||||
Quality -->|no and categoryFilter exists| Retry["UNFILTERED_VECTOR_RETRY"]
|
||||
Retry --> Post2["KnowledgeEvidencePostProcessor"]
|
||||
Post2 --> Pack["KnowledgeContextPacker"]
|
||||
|
||||
Pack --> Assemble["LookupResultAssembler"]
|
||||
Assemble --> Result["LookupResult"]
|
||||
Result --> Dedup["RetrievedDocTracker session dedup"]
|
||||
Dedup --> Recorder["ToolInvocationRecorder"]
|
||||
Recorder --> Trace["tool_invocation.retrieval_details"]
|
||||
Result --> Agent
|
||||
```
|
||||
|
||||
对应代码:
|
||||
|
||||
| 阶段 | 类 | 职责 |
|
||||
|---|---|---|
|
||||
| Tool Boundary | `LookupKnowledgeTool` | 接收 Agent 工具调用,编排 pipeline,处理 session dedup 和 recorder |
|
||||
| Query Transformation | `KnowledgeQueryTransformer` | 复用 L0,输出 query hints 和可选 category filter |
|
||||
| Retrieval | `KnowledgeDocumentRetriever` | 调用 `VectorSearchService`,统一 filtered / unfiltered attempt |
|
||||
| Post-Retrieval | `KnowledgeEvidencePostProcessor` | L2 归一化、证据块构建、source dedup、规则 rerank |
|
||||
| Context Packing | `KnowledgeContextPacker` | 按字符预算打包 Agent 可消费 context |
|
||||
| Result Assembly | `LookupResultAssembler` | 统一 evidence result、no-evidence result、dedup result |
|
||||
| Observability | `ToolInvocationRecorder` | 写入 query transform、retrieval trace、rerank trace、context pack summary |
|
||||
|
||||
## 3. L0 与 L1 边界
|
||||
|
||||
L0 来源于 `KnowledgeIndexService.analyzeQuery`,输出进入 `KnowledgeQuery`:
|
||||
|
||||
```text
|
||||
originalQuery
|
||||
rewrittenQuery
|
||||
domainHints
|
||||
matchedKeywords
|
||||
entities
|
||||
categoryFilter
|
||||
l0Titles
|
||||
l0MatchCount
|
||||
```
|
||||
|
||||
L0 可以做:
|
||||
|
||||
- 给 L1 提供单一 category filter。
|
||||
- 给 rerank 提供 domain / keyword / entity boost 信号。
|
||||
- 给 trace 提供解释信息。
|
||||
|
||||
L0 不再做:
|
||||
|
||||
- 不因唯一命中直接返回文档正文。
|
||||
- 不在 L1 无结果时作为事实证据兜底。
|
||||
- 不进入 `evidenceBlocks`,除非未来明确引入新的 evidence source 规则。
|
||||
|
||||
L1 通过 `VectorSearchService.searchSimilarDocuments(query, topK, category)` 执行,内部仍保留 Spring AI VectorStore 优先和 Milvus SDK fallback。
|
||||
|
||||
## 4. 降级策略
|
||||
|
||||
MVP 降级策略保持简单:
|
||||
|
||||
```text
|
||||
if categoryFilter exists:
|
||||
run FILTERED_VECTOR
|
||||
if empty / no evidence / top similarity < referenceThreshold:
|
||||
run UNFILTERED_VECTOR_RETRY with original query
|
||||
else:
|
||||
run UNFILTERED_VECTOR
|
||||
```
|
||||
|
||||
fallback reason:
|
||||
|
||||
| reason | 含义 |
|
||||
|---|---|
|
||||
| `filtered_vector_no_evidence` | filtered L1 无候选或 post-processing 后无 evidence |
|
||||
| `filtered_vector_low_quality` | filtered L1 有候选,但 top normalized similarity 低于 `retrieval.normalization.reference-threshold` |
|
||||
|
||||
当 retry 后仍无证据:
|
||||
|
||||
- `found=false`
|
||||
- `evidenceStatus=no_evidence`
|
||||
- 保留 `retrievalTrace`
|
||||
- 不返回 L0 文档作为事实证据
|
||||
|
||||
## 5. Evidence-First Contract
|
||||
|
||||
`LookupResult` 当前核心字段:
|
||||
|
||||
```text
|
||||
found
|
||||
evidenceBlocks
|
||||
contextPack
|
||||
retrievalTrace
|
||||
rerankTrace
|
||||
relevanceLevel
|
||||
completenessHint
|
||||
retrievedDomainsThisSession
|
||||
message
|
||||
```
|
||||
|
||||
旧字段已删除:
|
||||
|
||||
```text
|
||||
primary
|
||||
supplement
|
||||
```
|
||||
|
||||
这是一项 L4 breaking interface change。项目内已同步迁移:
|
||||
|
||||
- `LookupKnowledgeTool`
|
||||
- `ToolInvocationRecorder`
|
||||
- executor prompts
|
||||
- lookup / recorder tests
|
||||
- RAG architecture docs
|
||||
- OpenSpec 主 spec
|
||||
|
||||
## 6. Trace 结构
|
||||
|
||||
`retrieval_details` 保持 JSON 扩展,不改表结构。关键内容:
|
||||
|
||||
```json
|
||||
{
|
||||
"query_transform": {},
|
||||
"retrieval_trace": {
|
||||
"selected_attempt": "UNFILTERED_VECTOR_RETRY",
|
||||
"fallback_reason": "filtered_vector_no_evidence",
|
||||
"attempts": []
|
||||
},
|
||||
"context_pack_summary": {},
|
||||
"rerank_trace": {},
|
||||
"evidence_blocks": []
|
||||
}
|
||||
```
|
||||
|
||||
这样 Trace API、Verifier、Eval 可以继续从 `tool_invocation` 读取证据链。
|
||||
|
||||
## 7. Review 修正
|
||||
|
||||
归档前 review 发现:session dedup 命中时返回 `found=false`,但仍携带 `evidenceBlocks/contextPack`,可能导致 Agent 重复消费证据。
|
||||
|
||||
当前行为已修正:
|
||||
|
||||
- dedup result 不再返回可消费 evidence/context。
|
||||
- 保留 message、retrieval trace、relevance hint 和 retrieved domains。
|
||||
- 测试覆盖:`LookupKnowledgeToolTest.sessionDedupDoesNotReturnConsumableEvidenceAgain`。
|
||||
|
||||
## 8. 验证
|
||||
|
||||
已执行:
|
||||
|
||||
```powershell
|
||||
mvn -q -DskipTests compile
|
||||
mvn -q "-Dtest=LookupKnowledgeToolTest,ToolInvocationRecorderTest" test
|
||||
$env:MILVUS_TOKEN = <application.yml 中的 milvus.token>; mvn -q test
|
||||
openspec validate --all --strict
|
||||
git diff --check
|
||||
```
|
||||
|
||||
结果:全部通过。
|
||||
|
||||
注意:
|
||||
|
||||
- `MilvusConnectionTest` 直接读 `MILVUS_TOKEN` 环境变量,不读 Spring 配置。
|
||||
- 完整测试需要在 Maven 进程里注入该环境变量。
|
||||
|
||||
## 9. 后续演进
|
||||
|
||||
建议后续按评测结果推进,而不是先堆复杂能力:
|
||||
|
||||
- 增加 RAG eval cases:固定 query、期望 source、期望 fallback path。
|
||||
- 引入更严格的 evidence grounding 检查。
|
||||
- 当规则 rerank 不足时,再考虑 model-based rerank。
|
||||
- 当召回覆盖率不足时,再考虑 BM25/RRF/hybrid retrieval。
|
||||
@@ -1,6 +1,6 @@
|
||||
# RAG 新架构
|
||||
|
||||
**更新日期**:2026-07-05
|
||||
**更新日期**:2026-07-06
|
||||
**状态**:当前主架构 + 后续演进边界
|
||||
**关联计划**:`mvp/issues/rag-refactor-plan.md`
|
||||
|
||||
@@ -37,16 +37,20 @@ flowchart TD
|
||||
Mode -->|auto| SpringTry["try Spring AI VectorStore"]
|
||||
SpringTry -->|success| Results["SearchResult list"]
|
||||
SpringTry -->|failure| SdkFallback["Milvus SDK fallback"]
|
||||
Mode -->|spring-ai| SpringOnly["Spring AI VectorStore only"]
|
||||
Mode -->|spring / spring-ai| SpringOnly["Spring AI VectorStore only"]
|
||||
Mode -->|sdk| SdkOnly["Milvus SDK only"]
|
||||
|
||||
SpringOnly --> Results
|
||||
SdkFallback --> Results
|
||||
SdkOnly --> Results
|
||||
|
||||
Results --> Normalize["relevance normalization"]
|
||||
Normalize --> Dedup["session dedup: RetrievedDocTracker"]
|
||||
Dedup --> Output["LookupResult"]
|
||||
Results --> Retry{"filtered result usable?"}
|
||||
Retry -->|no| RetryL1["raw query unfiltered L1 retry"]
|
||||
Retry -->|yes| Post["post-retrieval processing"]
|
||||
RetryL1 --> Post
|
||||
Post --> Pack["context packing"]
|
||||
Pack --> Dedup["session dedup: RetrievedDocTracker"]
|
||||
Dedup --> Output["LookupResult: evidenceBlocks / contextPack / traces"]
|
||||
Output --> Record["tool_invocation record"]
|
||||
Output --> Agent
|
||||
```
|
||||
@@ -62,14 +66,17 @@ Agent Executor
|
||||
-> mode=auto
|
||||
-> Spring AI VectorStore
|
||||
-> fallback: Milvus SDK
|
||||
-> mode=spring-ai
|
||||
-> mode=spring / spring-ai
|
||||
-> Spring AI VectorStore only
|
||||
-> mode=sdk
|
||||
-> Milvus SDK only
|
||||
-> result normalization
|
||||
-> post-retrieval processing
|
||||
-> relevanceLevel
|
||||
-> completenessHint
|
||||
-> score/rawScore/scoreLabel
|
||||
-> evidenceBlocks
|
||||
-> rerankTrace
|
||||
-> context packing
|
||||
-> contextPack
|
||||
-> session dedup
|
||||
-> RetrievedDocTracker
|
||||
-> tool_invocation record
|
||||
@@ -127,7 +134,7 @@ Executor -> LookupKnowledgeTool -> VectorSearchService
|
||||
`VectorSearchService` 是当前检索门面:
|
||||
|
||||
- `auto`:优先 Spring AI VectorStore,失败后 fallback 到 SDK。
|
||||
- `spring-ai`:只走 Spring AI VectorStore。
|
||||
- `spring` / `spring-ai`:只走 Spring AI VectorStore。
|
||||
- `sdk`:只走原 Milvus SDK。
|
||||
|
||||
这样可以在不改 Agent 工具的情况下切换检索实现,并支持线上验证和回退。
|
||||
@@ -180,10 +187,11 @@ L0 负责:
|
||||
- metadata/category filter candidate
|
||||
- trace 中的 hit reason
|
||||
|
||||
L0 不再默认负责:
|
||||
L0 不再负责:
|
||||
|
||||
```text
|
||||
L0 unique hit -> 直接作为最终检索结果
|
||||
L1 no result -> 返回 L0 文档作为事实证据
|
||||
```
|
||||
|
||||
当前职责是:
|
||||
@@ -192,8 +200,10 @@ L0 unique hit -> 直接作为最终检索结果
|
||||
query / AIOps payload
|
||||
-> L0 matched keywords / domains / entities
|
||||
-> category filter candidate
|
||||
-> L1 semantic retrieval
|
||||
-> relevance normalization
|
||||
-> filtered L1 semantic retrieval
|
||||
-> low-quality? raw query unfiltered L1 retry
|
||||
-> post-retrieval processing
|
||||
-> context packing
|
||||
```
|
||||
|
||||
这样既保留精确关键词和领域 hint 的价值,也避免 L0 误召回直接污染最终证据。
|
||||
@@ -319,33 +329,24 @@ AIOps payload
|
||||
|
||||
## 9. Evidence 与去重
|
||||
|
||||
当前 evidence 输出仍以 `LookupResult` 和工具返回文本为主,已经具备:
|
||||
当前 evidence 输出已从旧 `primary/supplement` 迁移为 evidence-first contract,核心字段包括:
|
||||
|
||||
- L0/L1 命中数量。
|
||||
- 检索层记录。
|
||||
- relevance level。
|
||||
- completeness hint。
|
||||
- session 级文档去重。
|
||||
- domain 行动记忆。
|
||||
- `tool_invocation` 明细记录。
|
||||
- `evidenceBlocks`
|
||||
- `contextPack`
|
||||
- `retrievalTrace`
|
||||
- `rerankTrace`
|
||||
- `relevanceLevel`
|
||||
- `completenessHint`
|
||||
- `retrievedDomainsThisSession`
|
||||
- `tool_invocation.retrieval_details`
|
||||
|
||||
后续更完整的 evidence block 目标:
|
||||
evidence block 结构:
|
||||
|
||||
```text
|
||||
source
|
||||
docId
|
||||
chunkIndex
|
||||
title
|
||||
breadcrumb
|
||||
score
|
||||
rawScore
|
||||
scoreLabel
|
||||
hitReason
|
||||
content
|
||||
expandedFrom
|
||||
source / title / breadcrumb / retrievalLayer / content / score / hitReasons
|
||||
```
|
||||
|
||||
这部分应作为下一阶段增强,而不是当前已完全完成能力。
|
||||
context pack 会按重排后的证据顺序生成 Agent 可消费的紧凑上下文,并保留 included/omitted sources 供 trace 检查。
|
||||
|
||||
## 10. 评测与验收
|
||||
|
||||
@@ -373,25 +374,25 @@ RAG 架构变更必须先过评测,再认为可合入主链路。
|
||||
- `lookup_knowledge` 保持显式 Agent Tool。
|
||||
- L0 降级为 domain/entity hint。
|
||||
- L1 默认执行语义检索。
|
||||
- `VectorSearchService` 支持 `auto`、`spring-ai`、`sdk` 三种模式。
|
||||
- `VectorSearchService` 支持 `auto`、`spring`/`spring-ai`、`sdk` 三种模式。
|
||||
- Spring AI VectorStore 成为读取主路径。
|
||||
- Milvus SDK fallback 保留。
|
||||
- 分数语义拆成 `score`、`rawScore`、`scoreLabel`。
|
||||
- Markdown chunk 保留 `title` 和 `breadcrumb`。
|
||||
- embedding 输入包含 `title`、`breadcrumb` 和 `content`。
|
||||
- AIOps payload 生成推荐知识库 query。
|
||||
- `tool_invocation` 记录 relevance level 和 dedup reason。
|
||||
- `tool_invocation` 记录 relevance level、dedup reason、evidence summaries、retrieval trace、rerank trace 和 context pack summary。
|
||||
- `lookup_knowledge` 输出使用 evidence-first contract,不再暴露旧 `primary/supplement` 字段。
|
||||
- RAG offline baseline 和 live acceptance 脚本已补齐。
|
||||
|
||||
## 12. 后续演进
|
||||
|
||||
近期优先:
|
||||
|
||||
1. 完整 evidence block 结构化输出。
|
||||
2. 命中 chunk 的相邻 chunk / 同章节上下文扩展。
|
||||
3. metadata taxonomy 清理,例如 `database` 与 `infrastructure` 的分类边界。
|
||||
4. Query Transformer / MultiQuery 的可回退接入。
|
||||
5. VectorStore 写入路径评估。
|
||||
1. 命中 chunk 的相邻 chunk / 同章节上下文扩展。
|
||||
2. metadata taxonomy 清理,例如 `database` 与 `infrastructure` 的分类边界。
|
||||
3. Query Transformer / MultiQuery 的可回退接入。
|
||||
4. VectorStore 写入路径评估。
|
||||
|
||||
暂不优先:
|
||||
|
||||
|
||||
@@ -0,0 +1,183 @@
|
||||
# RAG 评测闭环架构
|
||||
|
||||
**更新日期**:2026-07-06
|
||||
|
||||
本文记录当前 RAG 质量闭环。它的目标不是证明检索“永远正确”,而是让每次改 `lookup_knowledge`、L0 hint、向量召回、post-retrieval、rerank 或 context packing 时,都能得到可重复的回归信号。
|
||||
|
||||
## 1. 闭环分层
|
||||
|
||||
```text
|
||||
RAG pipeline change
|
||||
-> LookupKnowledgeTool snapshot generation
|
||||
-> offline RAG retrieval baseline
|
||||
-> RAG baseline diff
|
||||
-> diagnosis eval baseline
|
||||
-> diagnosis baseline diff
|
||||
-> accept / fix / archive
|
||||
```
|
||||
|
||||
| 层级 | 位置 | 作用 |
|
||||
|---|---|---|
|
||||
| RAG retrieval baseline | `eval/rag-retrieval/` | 检查固定 query 是否命中期望证据、路径和 fallback |
|
||||
| RAG baseline diff | `scripts/eval_rag_retrieval.py --compare-to ...` | 对比当前报告和旧基线,输出 regression/change |
|
||||
| Diagnosis eval baseline | `mvp/eval/` | 检查 Agent 最终诊断 trace、报告和证据行为 |
|
||||
| Live acceptance | `scripts/eval_rag_live_acceptance.py` | 在应用和向量库运行后做真实环境 smoke check |
|
||||
|
||||
## 2. Offline RAG Baseline
|
||||
|
||||
核心资产:
|
||||
|
||||
```text
|
||||
eval/rag-retrieval/cases/golden-cases.json
|
||||
eval/rag-retrieval/fixtures/*.json
|
||||
eval/rag-retrieval/reports/baseline.json
|
||||
eval/rag-retrieval/reports/baseline.md
|
||||
scripts/eval_rag_retrieval.py
|
||||
scripts/generate_rag_lookup_snapshots.ps1
|
||||
src/test/java/com/superbiz/agent/eval/RagLookupSnapshotGeneratorTest.java
|
||||
```
|
||||
|
||||
运行:
|
||||
|
||||
```powershell
|
||||
python scripts\eval_rag_retrieval.py
|
||||
```
|
||||
|
||||
该 baseline 完全离线,不依赖 MySQL、Redis、Milvus、LLM 或 Spring Boot。它适合在改 RAG 代码后快速判断:
|
||||
|
||||
- 期望 source 是否仍在 topK 内。
|
||||
- breadcrumb 和 evidence keyword 是否仍能覆盖。
|
||||
- `LookupResult` 是否仍包含 `evidenceBlocks/contextPack/retrievalTrace/rerankTrace`。
|
||||
- selected attempt 是否符合预期。
|
||||
- fallback reason 是否符合预期。
|
||||
- context pack 是否包含期望 source。
|
||||
- rerank top source 是否稳定。
|
||||
|
||||
## 3. 模块化输出契约
|
||||
|
||||
fixture 必须使用当前模块化格式:
|
||||
|
||||
```json
|
||||
{
|
||||
"lookupResult": {
|
||||
"evidenceBlocks": [],
|
||||
"contextPack": {},
|
||||
"retrievalTrace": {},
|
||||
"rerankTrace": {}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
当前 golden cases 直接以模块化格式为唯一契约,因为这个版本的目标是验证完整 RAG pipeline,而不只是验证候选召回。
|
||||
|
||||
## 4. Fallback Case
|
||||
|
||||
当前 baseline 增加了 `chat-l0-filter-fallback`:
|
||||
|
||||
```text
|
||||
FILTERED_VECTOR low quality or no evidence
|
||||
-> UNFILTERED_VECTOR_RETRY
|
||||
-> fallbackReason = filtered_vector_low_quality | filtered_vector_no_evidence
|
||||
```
|
||||
|
||||
这个 case 固化了 MVP 版本的降级策略:如果经过 L0 filter 后 L1 低质量或没有证据,就跳过 L0 filter,用原始 query 再做一次无过滤向量检索。不同向量后端对“低质量候选”和“无候选”的边界可能不同,所以 golden case 允许两个 fallback reason,但强制要求 retry 行为和最终证据正确。
|
||||
|
||||
## 5. Diff 闭环
|
||||
|
||||
生成当前报告并与旧基线对比:
|
||||
|
||||
```powershell
|
||||
python scripts\eval_rag_retrieval.py `
|
||||
--json-report eval\rag-retrieval\reports\current.json `
|
||||
--markdown-report eval\rag-retrieval\reports\current.md `
|
||||
--compare-to eval\rag-retrieval\reports\baseline.json `
|
||||
--diff-json-report eval\rag-retrieval\reports\baseline-diff.json `
|
||||
--diff-markdown-report eval\rag-retrieval\reports\baseline-diff.md
|
||||
```
|
||||
|
||||
diff 会检查:
|
||||
|
||||
- pass rate
|
||||
- recall@K
|
||||
- strong hit rate
|
||||
- miss count
|
||||
- case pass state
|
||||
- hit level
|
||||
- first expected rank
|
||||
- selected attempt
|
||||
- fallback reason
|
||||
- evidence status
|
||||
- rerank top source
|
||||
|
||||
当 case 失败或 diff 出现 regression 时,脚本会返回非 0 退出码,可作为本地质量门禁或 CI 门禁。
|
||||
|
||||
## 6. 与 Diagnosis Eval 的关系
|
||||
|
||||
RAG baseline 解决的是“证据有没有被正确检索、处理和打包”。
|
||||
|
||||
Diagnosis eval 解决的是“Agent 有没有把证据用于最终诊断,并保持 trace 可解释”。
|
||||
|
||||
两者不是替代关系:
|
||||
|
||||
- 改 RAG pipeline:先跑 RAG baseline,再跑相关 Agent 测试。
|
||||
- 改 prompt、Agent 编排、Verifier:重点跑 diagnosis eval。
|
||||
- 改 embedding 输入、reindex、向量库配置:跑 RAG baseline + live acceptance。
|
||||
|
||||
## 7. 面试表达
|
||||
|
||||
可以概括为:
|
||||
|
||||
> 我没有只做一个 RAG 调用,而是把 RAG 拆成 Query Transform、Retrieval、Post-Retrieval、Rerank、Context Packing,并为它建设了离线 golden cases、baseline report、baseline diff 和上层 diagnosis eval,形成可回放、可对比、可回归的 Agent 质量闭环。
|
||||
|
||||
## 8. LookupKnowledgeTool Snapshot
|
||||
|
||||
真实工具快照生成命令:
|
||||
|
||||
```powershell
|
||||
.\scripts\generate_rag_lookup_snapshots.ps1
|
||||
```
|
||||
|
||||
该命令默认使用 `retrieval.vector-store.mode=spring`,通过 `RagLookupSnapshotGeneratorTest` 启动 Spring test context,注入真实 `LookupKnowledgeTool` bean,对 `golden-cases.json` 中每个 query 调用 `lookupKnowledge(query)`,并把返回的 `LookupResult` 写入 `eval/rag-retrieval/fixtures/{caseId}.json`。
|
||||
|
||||
普通测试不会执行快照生成器;只有显式传入 `rag.snapshot.enabled=true` 时才会写 fixture。
|
||||
|
||||
## 9. Seed Docs And Scope Isolation
|
||||
|
||||
Live `LookupKnowledgeTool` snapshots are only stable if the expected documents
|
||||
exist in the real knowledge base and vector index. The eval loop therefore adds
|
||||
a canonical seed layer:
|
||||
|
||||
```text
|
||||
eval/rag-retrieval/seed-docs/*.md
|
||||
-> scripts/prepare_rag_eval_seed.ps1
|
||||
-> RagEvalSeedImporterTest
|
||||
-> DocumentManagementService.uploadDocument
|
||||
-> api_document metadata + L0 index + Milvus chunks
|
||||
```
|
||||
|
||||
仓库内还保留一份 `knowledge_base/rag-eval/` 镜像,方便直接查看和提交 eval 知识库文档。它们放在单独目录下,避免和 `knowledge_base/api`、`knowledge_base/infrastructure` 等业务知识目录混在一起;检索 category 仍由 frontmatter 中的 `category` 决定。
|
||||
|
||||
Seed frontmatter includes:
|
||||
|
||||
```yaml
|
||||
source: mysql-connection-pool
|
||||
breadcrumb: Database > MySQL > Connection Pool
|
||||
kb_scope: rag-eval
|
||||
```
|
||||
|
||||
`source` becomes the stable `docId` when it fits the DB column, and is also
|
||||
written to vector metadata as `_source` and `source`. `breadcrumb` is copied into
|
||||
chunk metadata so evidence blocks can keep a stable path. `kb_scope` isolates
|
||||
eval documents from local production documents.
|
||||
|
||||
Default runtime behavior keeps `retrieval.kb-scope` empty, so existing documents
|
||||
without `kb_scope` are still searchable. Eval scripts pass
|
||||
`-Dretrieval.kb-scope=rag-eval`, so L0 query hints, the filtered attempt, and
|
||||
the unfiltered retry stay inside the eval corpus while the retry still skips the
|
||||
L0 category filter.
|
||||
|
||||
Frontmatter is used for DB metadata, L0 hints, and vector metadata. It is
|
||||
stripped before document chunking so embedding content represents the Markdown
|
||||
body, not the YAML control plane. This is important for fallback eval: a decoy
|
||||
document may intentionally match L0 keywords, but its body should remain low
|
||||
quality evidence so the retry path can be exercised.
|
||||
@@ -1,6 +1,6 @@
|
||||
# 检索与可观测性架构
|
||||
|
||||
**更新日期**:2026-07-05
|
||||
**更新日期**:2026-07-06
|
||||
**状态**:当前可运行架构
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/knowledge-retrieval-architecture.md`
|
||||
|
||||
@@ -13,7 +13,7 @@
|
||||
- 检索结果如何归一化、去重、记录。
|
||||
- 如何通过 trace 和 eval 判断检索质量。
|
||||
|
||||
当前架构与旧版最大的差异是:L0 不再因为唯一命中而默认跳过 L1。L0 是 hint 和解释信号,L1 语义检索是默认召回路径。
|
||||
当前架构与旧版最大的差异是:L0 不再因为唯一命中而默认跳过 L1,也不在 L1 失败时作为事实证据兜底。L0 是 hint 和解释信号,L1 语义检索是默认召回路径。
|
||||
|
||||
## 2. 检索总图
|
||||
|
||||
@@ -30,14 +30,18 @@ flowchart TD
|
||||
L1 --> Mode{"retrieval.vector-store.mode"}
|
||||
Mode -->|auto| Spring["Spring AI VectorStore"]
|
||||
Spring -->|failure| SDK["Milvus SDK fallback"]
|
||||
Mode -->|spring-ai| Spring
|
||||
Mode -->|spring / spring-ai| Spring
|
||||
Mode -->|sdk| SDK
|
||||
|
||||
Spring --> Candidates["L1 candidates"]
|
||||
SDK --> Candidates
|
||||
Candidates --> Normalize["relevance normalization"]
|
||||
L0Result --> Normalize
|
||||
Normalize --> Result["LookupResult"]
|
||||
Candidates --> Quality{"filtered L1 usable?"}
|
||||
Quality -->|no| Retry["raw query unfiltered L1 retry"]
|
||||
Quality -->|yes| Post["post-retrieval processing"]
|
||||
Retry --> Post
|
||||
L0Result --> Post
|
||||
Post --> Pack["context packing"]
|
||||
Pack --> Result["LookupResult evidenceBlocks/contextPack/traces"]
|
||||
|
||||
Result --> Dedup["RetrievedDocTracker session dedup"]
|
||||
Dedup --> Final["final tool output"]
|
||||
@@ -69,6 +73,7 @@ singleDomainOrNull
|
||||
|
||||
```text
|
||||
matches=1 -> skip L1 -> 直接返回 L0 文档正文
|
||||
L1 无可用证据 -> 返回 L0 文档正文
|
||||
```
|
||||
|
||||
原因:
|
||||
@@ -84,7 +89,7 @@ L1 通过 `VectorSearchService` 调度,支持三种模式:
|
||||
| 模式 | 行为 | 用途 |
|
||||
|---|---|---|
|
||||
| `auto` | 优先 Spring AI VectorStore,失败 fallback 到 SDK | 默认运行模式 |
|
||||
| `spring-ai` | 只走 Spring AI VectorStore | 验证框架路径 |
|
||||
| `spring` / `spring-ai` | 只走 Spring AI VectorStore | 验证框架路径 |
|
||||
| `sdk` | 只走 Milvus SDK | 对比旧链路或临时回退 |
|
||||
|
||||
### Spring AI VectorStore 路径
|
||||
@@ -123,12 +128,12 @@ SDK fallback 保留的价值:
|
||||
| `rawScore` | 底层检索实现原始分数 |
|
||||
| `scoreLabel` | 原始分数语义,例如 `similarity` 或 `l2_distance` |
|
||||
|
||||
工具层再把 L0/L1 情况归一为:
|
||||
post-retrieval 层再把检索候选归一为:
|
||||
|
||||
| relevanceLevel | 含义 |
|
||||
|---|---|
|
||||
| `PRECISE` | L0 单命中且 L1 相似度高 |
|
||||
| `HIGHLY_RELEVANT` | L1 相似度高,或 L0 多命中且 L1 支撑强 |
|
||||
| `PRECISE` | L1 相似度高且 query hint 与候选证据互相支撑 |
|
||||
| `HIGHLY_RELEVANT` | L1 相似度高 |
|
||||
| `REFERENCE` | 可作为参考,但不足以声明强证据 |
|
||||
| `DEDUPED` | 同 session 中已检索过,不重复注入上下文 |
|
||||
|
||||
@@ -169,7 +174,7 @@ title + breadcrumb + content
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
LookupResult["LookupResult"] --> Agent["Agent context"]
|
||||
LookupResult["LookupResult: evidenceBlocks/contextPack/traces"] --> Agent["Agent context"]
|
||||
LookupResult --> Recorder["ToolInvocationRecorder"]
|
||||
Recorder --> Invocation["tool_invocation"]
|
||||
Invocation --> Trace["DiagnosisTraceService"]
|
||||
@@ -195,10 +200,13 @@ success
|
||||
`retrieval_details` 承载更细信息,例如:
|
||||
|
||||
- L0 命中文档标题和路径。
|
||||
- L1 分数。
|
||||
- L1 attempts、fallback reason、分数和 similarity。
|
||||
- retrieved domains。
|
||||
- evidence status。
|
||||
- dedup reason。
|
||||
- evidence block summaries。
|
||||
- context pack summary。
|
||||
- rerank trace。
|
||||
|
||||
## 8. 去重与行动记忆
|
||||
|
||||
@@ -252,15 +260,15 @@ trace inspection
|
||||
|
||||
近期优先:
|
||||
|
||||
1. 完整 evidence block 输出。
|
||||
2. 邻居 chunk / 同章节上下文扩展。
|
||||
3. metadata taxonomy 清理。
|
||||
4. Query Transformer / MultiQuery 可回退接入。
|
||||
5. 更完整的 Recall@K、MRR、nDCG 报告。
|
||||
1. 邻居 chunk / 同章节上下文扩展。
|
||||
2. metadata taxonomy 清理。
|
||||
3. Query Transformer / MultiQuery 可回退接入。
|
||||
4. 更完整的 Recall@K、MRR、nDCG 报告。
|
||||
|
||||
暂不优先:
|
||||
|
||||
- 重新引入 L0 直接返回。
|
||||
- 重新引入 L0 文档作为 L1 失败时的事实证据兜底。
|
||||
- 一次性迁移所有写入路径。
|
||||
- 在没有评测收益前引入 rerank / RRF / BM25。
|
||||
- 在没有评测收益前引入模型 rerank / RRF / BM25。
|
||||
|
||||
|
||||
@@ -0,0 +1 @@
|
||||
ready
|
||||
@@ -0,0 +1 @@
|
||||
committed
|
||||
@@ -0,0 +1,396 @@
|
||||
# Modular RAG Pipeline Decisions
|
||||
|
||||
## Discover Summary
|
||||
|
||||
- Capability source: sm-flow Discover using local repository evidence and existing OpenSpec/devflow context.
|
||||
- Slug: `modular-rag-pipeline`.
|
||||
- Scale: standard.
|
||||
- Goal: turn `lookup_knowledge` into a modular RAG pipeline suitable for Agent engineering interview use, while keeping the explicit Agent tool and evidence trace.
|
||||
|
||||
## Context Evidence
|
||||
|
||||
### Existing Architecture
|
||||
|
||||
- `mvp/architecture/rag-architecture.md` documents the desired boundary: mature framework retrieval plus business-observable orchestration.
|
||||
- `mvp/architecture/retrieval-observability.md` says L0 is a hint/explainability layer and L1 semantic retrieval is the default recall path.
|
||||
- `VectorSearchService` is already the retrieval facade and supports Spring AI `VectorStore` with SDK fallback.
|
||||
- `LookupKnowledgeTool` currently still owns query analysis, L1 invocation, relevance normalization, result assembly, evidence block construction, session dedup, and recorder calls.
|
||||
|
||||
### Existing Evidence Blocks
|
||||
|
||||
- `EvidenceBlock` already exists.
|
||||
- `LookupResult` already has `evidenceBlocks`, `evidenceCandidateCount`, and `evidenceBlockCount`.
|
||||
- `LookupKnowledgeTool` currently builds evidence blocks internally.
|
||||
- `ToolInvocationRecorder` already persists compact evidence block summaries in `retrieval_details`.
|
||||
|
||||
Conclusion: evidence blocks are partially implemented, but the post-retrieval module boundary is not.
|
||||
|
||||
### Current Gaps
|
||||
|
||||
- Context packing is not a first-class module.
|
||||
- Rerank is not a first-class module; only original semantic rank and trace summary sorting exist.
|
||||
- The `primary` / `supplement` result model still encodes old L0/L1 semantics.
|
||||
- `lookup_knowledge` tool description still describes the old two-stage retrieval model.
|
||||
|
||||
## Question Pool
|
||||
|
||||
| ID | Dimension | Question | Mode | Status |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| Q1 | Boundary | Should this be internal-only refactor or allow return contract changes? | user-interview | confirmed |
|
||||
| Q2 | Fallback | If filtered L1 fails, should L0 provide weak fallback evidence? | user-interview | confirmed |
|
||||
| Q3 | Interface | Can `LookupResult` add new fields and remove old `primary` / `supplement` if simpler? | user-interview | confirmed |
|
||||
| Q4 | Architecture | Does current repo already have evidence blocks, context packing, and rerank? | evidence-driven | resolved |
|
||||
| Q5 | Compatibility | Which in-repo consumers reference `primary` / `supplement`? | evidence-driven | resolved |
|
||||
| Q6 | Commit detail | What are the exact fields and thresholds for traces/context pack/low quality? | user-interview or specify | pending |
|
||||
|
||||
## Confirmed User Decisions
|
||||
|
||||
### D1: Prefer one-shot modular RAG refactor
|
||||
|
||||
User confirmed that the change can be done "一次到位" instead of only doing a compatibility-preserving internal refactor.
|
||||
|
||||
Implementation implication:
|
||||
|
||||
- Create full pipeline modules now.
|
||||
- Do not leave `LookupKnowledgeTool` as a large procedural class.
|
||||
|
||||
### D2: L0 is not a normal evidence retrieval path
|
||||
|
||||
User challenged the first design because it made L0 participate in too many flows.
|
||||
|
||||
Confirmed direction:
|
||||
|
||||
```text
|
||||
L0 -> query understanding / category filter / domain/entity/keyword signal
|
||||
L1 -> main vector retrieval
|
||||
```
|
||||
|
||||
Implementation implication:
|
||||
|
||||
- Do not model L0 and L1 as equal retrievers in the normal path.
|
||||
- L0 output can influence filter, rerank, and trace.
|
||||
|
||||
### D3: MVP fallback is unfiltered L1 retry
|
||||
|
||||
User proposed a simpler MVP fallback:
|
||||
|
||||
```text
|
||||
Filtered L1 using L0 category filter
|
||||
-> if inaccurate or empty
|
||||
-> retry raw query with L1 and no L0 filter
|
||||
```
|
||||
|
||||
Confirmed direction:
|
||||
|
||||
- Use filtered vector retrieval first when L0 provides an unambiguous category.
|
||||
- If filtered retrieval is low-quality, retry unfiltered vector retrieval with raw query.
|
||||
- Do not return L0 documents as fact evidence fallback in this MVP design.
|
||||
|
||||
### D4: Result contract can change
|
||||
|
||||
User confirmed new fields can be added and old fields can be removed if the later flow becomes cleaner.
|
||||
|
||||
Implementation implication:
|
||||
|
||||
- `LookupResult.primary` and `LookupResult.supplement` may be removed.
|
||||
- Preferred contract becomes `evidenceBlocks + contextPack + retrievalTrace + rerankTrace`.
|
||||
- This is a breaking interface change and must be treated as L4.
|
||||
|
||||
## Evidence-Driven Findings
|
||||
|
||||
### E1: Existing spec conflict
|
||||
|
||||
`openspec/specs/rag-knowledge-retrieval/spec.md` currently says L1 no-result or failure should return an L0-based primary result. This conflicts with the confirmed design.
|
||||
|
||||
Required OpenSpec update:
|
||||
|
||||
- Replace L0 primary fallback with unfiltered vector retry.
|
||||
- Define no-evidence behavior when both filtered and unfiltered L1 fail.
|
||||
|
||||
### E2: Existing compatibility-field requirement conflict
|
||||
|
||||
`openspec/specs/rag-knowledge-retrieval/spec.md` currently requires `primary` and `supplement` compatibility fields to remain when evidence blocks exist.
|
||||
|
||||
Required OpenSpec update:
|
||||
|
||||
- Remove compatibility-field requirement.
|
||||
- Define evidence blocks and context pack as the preferred tool result contract.
|
||||
|
||||
### E3: Primary/supplement references are localized
|
||||
|
||||
Search found concrete Java references in:
|
||||
|
||||
- `LookupKnowledgeTool`
|
||||
- `ToolInvocationRecorder`
|
||||
- `LookupKnowledgeToolTest`
|
||||
- `ToolInvocationRecorderTest`
|
||||
- `LookupResult`
|
||||
- `PrimaryResult`
|
||||
- `SupplementResult`
|
||||
|
||||
No broad in-repo service usage was found beyond tool implementation, recorder, tests, prompts, and historical docs.
|
||||
|
||||
Implementation implication:
|
||||
|
||||
- One-shot migration is feasible if tests and prompts are updated in the same change.
|
||||
|
||||
### E4: Existing trace contract must be preserved
|
||||
|
||||
`tool_invocation` is used by diagnosis trace, verifier, and evaluation code. The database table does not need to change for this design if new details remain inside `retrieval_details`.
|
||||
|
||||
Implementation implication:
|
||||
|
||||
- Keep table-level fields stable.
|
||||
- Enrich JSON `retrieval_details` with `retrieval_trace`, `rerank_trace`, `context_pack_summary`, and fallback reason.
|
||||
|
||||
## Interface Impact
|
||||
|
||||
Level: L4 breaking interface.
|
||||
|
||||
Reason:
|
||||
|
||||
- Removes or changes old result fields consumed by current tests and possibly by Agent prompt behavior.
|
||||
- Changes `lookup_knowledge` tool JSON shape.
|
||||
|
||||
Mitigation:
|
||||
|
||||
- Keep tool name and input signature unchanged.
|
||||
- Update all in-repo consumers in the same change.
|
||||
- Keep `tool_invocation` table schema stable.
|
||||
- Add tests for the new result contract.
|
||||
- Update prompt text to teach Agent to use `contextPack` and `evidenceBlocks`.
|
||||
|
||||
## Proposed Implementation Shape
|
||||
|
||||
Pipeline classes:
|
||||
|
||||
- `KnowledgeQueryTransformer`
|
||||
- `KnowledgeDocumentRetriever`
|
||||
- `KnowledgeEvidencePostProcessor`
|
||||
- `KnowledgeContextPacker`
|
||||
- `LookupResultAssembler`
|
||||
|
||||
New or updated DTOs:
|
||||
|
||||
- `KnowledgeQuery`
|
||||
- `RetrievedEvidenceCandidate`
|
||||
- `EvidencePostprocessResult`
|
||||
- `ContextPack`
|
||||
- `RetrievalTrace`
|
||||
- `RerankTrace`
|
||||
- `LookupResult`
|
||||
|
||||
Policy:
|
||||
|
||||
- L0-derived filter is optional and only used when unambiguous.
|
||||
- Filtered L1 low-quality result triggers raw unfiltered L1 retry.
|
||||
- Rule-based rerank is sufficient for MVP.
|
||||
- Context packing uses character budget first, not exact token counting.
|
||||
|
||||
## Pending For Commit
|
||||
|
||||
- Specify exact fields for `ContextPack`, `RetrievalTrace`, and `RerankTrace`.
|
||||
- Specify low-quality trigger for unfiltered retry.
|
||||
- Decide whether `PrimaryResult` / `SupplementResult` classes are deleted or deprecated during the first apply.
|
||||
- Write OpenSpec `design.md`, specs, and executable `tasks.md`.
|
||||
|
||||
## Specify Results
|
||||
|
||||
Created committed-design artifacts:
|
||||
|
||||
- `design.md`
|
||||
- `specs/rag-knowledge-retrieval/spec.md`
|
||||
- `tasks.md`
|
||||
|
||||
Resolved pending items:
|
||||
|
||||
- `ContextPack` minimum fields: `packedText`, `strategy`, `charBudget`, `usedChars`, `includedSources`, `omittedSources`.
|
||||
- `RetrievalTrace` minimum behavior: record filtered attempt, unfiltered retry when used, fallback reason, and no-evidence paths.
|
||||
- `RerankTrace` minimum behavior: record final rank, source, base retrieval score when available, and major boost reasons for top evidence blocks.
|
||||
- Low-quality trigger: empty candidates, empty final evidence, or top normalized similarity below `retrieval.normalization.reference-threshold`.
|
||||
- `PrimaryResult` / `SupplementResult`: may be removed during apply if all compile-time usages are migrated.
|
||||
|
||||
## Cross-Artifact Alignment
|
||||
|
||||
| Check | Result |
|
||||
| --- | --- |
|
||||
| proposal goals/scope -> design decisions | aligned |
|
||||
| design module boundaries -> specs behavior | aligned |
|
||||
| specs observable behavior -> tasks | aligned |
|
||||
| interface impact -> design/tasks migration work | aligned |
|
||||
|
||||
No cross-artifact gaps remain for Commit.
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
Input -> processing -> output chain:
|
||||
|
||||
```text
|
||||
lookup_knowledge(query)
|
||||
-> KnowledgeQueryTransformer
|
||||
-> KnowledgeDocumentRetriever
|
||||
-> KnowledgeEvidencePostProcessor
|
||||
-> KnowledgeContextPacker
|
||||
-> LookupResultAssembler
|
||||
-> ToolInvocationRecorder
|
||||
```
|
||||
|
||||
Risk assessment:
|
||||
|
||||
- The architecture keeps the explicit Agent tool boundary and does not move retrieval into an implicit Advisor.
|
||||
- Data ownership is clearer: query hints belong to transformer, vector candidates to retriever, evidence/context/traces to post-retrieval pipeline, persistence summaries to recorder.
|
||||
- The largest risk is the L4 result contract change; design and tasks require prompt/test/recorder migration in the same apply.
|
||||
- Database migration risk is low because `tool_invocation` table fields remain stable and new trace details stay in JSON.
|
||||
- Latency risk from unfiltered retry is accepted for MVP because retry only happens below reference quality.
|
||||
|
||||
## Commit Gate
|
||||
|
||||
OpenSpec validation:
|
||||
|
||||
```text
|
||||
openspec validate modular-rag-pipeline --strict
|
||||
Change 'modular-rag-pipeline' is valid
|
||||
```
|
||||
|
||||
File integrity:
|
||||
|
||||
- proposal exists and states problem, proposed change, scope, non-goals, risks, and interface impact.
|
||||
- design exists and records module boundaries, decisions, migration, rollback, and risks.
|
||||
- specs exist and define observable behavior for modular pipeline, unfiltered retry, evidence-first result, context pack, rerank, and L0 hint boundaries.
|
||||
- tasks exist and are executable vertical slices.
|
||||
|
||||
Consistency:
|
||||
|
||||
- proposal core concepts are represented in design.
|
||||
- design decisions are represented in specs and tasks.
|
||||
- tasks have verifiable implementation and test steps.
|
||||
- L4 interface impact is recorded and mapped to migration tasks.
|
||||
|
||||
Status: ready to mark as Committed OpenSpec.
|
||||
|
||||
## Pre-apply Research
|
||||
|
||||
Capability source: `openspec-apply-change` + sm-flow apply protocol. `codebase-retrieval` and LSP tools were not available in this session, so call-chain confirmation used OpenSpec context, `rg`, targeted file reads, and tests.
|
||||
|
||||
Reference implementation and affected files inspected:
|
||||
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
- `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
|
||||
- `src/main/java/com/superbiz/agent/tool/RetrievedDocTracker.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
|
||||
- `src/main/java/com/superbiz/agent/dto/LookupResult.java`
|
||||
- `src/test/java/com/superbiz/agent/tool/LookupKnowledgeToolTest.java`
|
||||
- `src/test/java/com/superbiz/agent/service/ToolInvocationRecorderTest.java`
|
||||
- `src/main/resources/prompts/chat-executor-prompt.md`
|
||||
- `src/main/resources/prompts/executor-prompt.md`
|
||||
- `src/main/resources/application.yml`
|
||||
|
||||
Technical stack checklist:
|
||||
|
||||
- Request/response structure: `lookup_knowledge` returns a Java DTO serialized as Agent tool JSON.
|
||||
- Retrieval facade: `VectorSearchService.searchSimilarDocuments(query, topK, category)` is the stable retrieval API and already hides Spring AI VectorStore vs SDK fallback.
|
||||
- L0 query hints: `KnowledgeIndexService.analyzeQuery` returns `L0Hint(matches, matchedKeywords, domains, entities, titles)` and `singleDomainOrNull()`.
|
||||
- Session dedup: `RetrievedDocTracker` stores `sessionId -> domain -> filePath` and exposes backward-compatible `isAlreadyRetrieved`.
|
||||
- Trace persistence: `ToolInvocationRecorder.recordLookupKnowledge` writes stable table columns and JSON `retrieval_details`.
|
||||
- Prompt consumers: executor prompts still describe old L0/L1 and `primary.content`; these must be migrated.
|
||||
- Tests: `LookupKnowledgeToolTest` is the main consumer of old `primary` / `supplement` assertions; recorder tests verify retrieval details.
|
||||
|
||||
Implementation decision:
|
||||
|
||||
- Add pipeline DTOs under `com.superbiz.agent.dto`.
|
||||
- Add pipeline services under `com.superbiz.agent.service`.
|
||||
- Keep `VectorSearchService` unchanged.
|
||||
- Keep `lookup_knowledge` tool name and query argument unchanged.
|
||||
- Delete `PrimaryResult` / `SupplementResult` only after production and tests stop referencing them.
|
||||
|
||||
## Apply Results
|
||||
|
||||
Completed tasks: 31/31.
|
||||
|
||||
Implemented:
|
||||
|
||||
- Added modular RAG DTOs: `KnowledgeQuery`, `RetrievedEvidenceCandidate`, `ContextPack`, `RetrievalTrace`, `RerankTrace`, `EvidencePostprocessResult`.
|
||||
- Added pipeline services: `KnowledgeQueryTransformer`, `KnowledgeDocumentRetriever`, `KnowledgeEvidencePostProcessor`, `KnowledgeContextPacker`, `LookupResultAssembler`.
|
||||
- Refactored `LookupKnowledgeTool` into a thin orchestrator.
|
||||
- Migrated `LookupResult` to evidence-first fields and removed `primary` / `supplement`.
|
||||
- Deleted `PrimaryResult` and `SupplementResult`.
|
||||
- Updated `ToolInvocationRecorder` to persist query transform, retrieval trace, context pack summary, rerank trace, fallback reason, and evidence summaries in `retrieval_details`.
|
||||
- Updated executor prompt text for `contextPack` / `evidenceBlocks`.
|
||||
- Updated RAG architecture docs.
|
||||
- Rewrote lookup and recorder tests for the new contract.
|
||||
|
||||
Verification:
|
||||
|
||||
```text
|
||||
mvn -q -DskipTests compile
|
||||
PASS
|
||||
|
||||
mvn -q "-Dtest=LookupKnowledgeToolTest,ToolInvocationRecorderTest" test
|
||||
PASS
|
||||
|
||||
openspec validate modular-rag-pipeline --strict
|
||||
PASS: Change 'modular-rag-pipeline' is valid
|
||||
```
|
||||
|
||||
Full suite attempt:
|
||||
|
||||
```text
|
||||
mvn -q test
|
||||
FAIL
|
||||
```
|
||||
|
||||
The full suite failed on pre-existing/environment-dependent tests:
|
||||
|
||||
- `MilvusConnectionTest.connect`: `MILVUS_TOKEN` not set.
|
||||
- `RedisSessionManagerTest`: Redis JSON contains legacy `messagePairCount`, not accepted by current `SessionContext`.
|
||||
- Spring context / repository tests attempted MySQL/Flyway and failed when database connectivity was unavailable in the first sandboxed run.
|
||||
|
||||
The full suite was retried outside the sandbox after approval. It still failed for the Milvus token and Redis serialization issues above, so these failures are not attributed to the modular RAG change.
|
||||
|
||||
Diff scope reviewed:
|
||||
|
||||
- RAG implementation: DTOs, pipeline services, `LookupKnowledgeTool`, `ToolInvocationRecorder`.
|
||||
- Contract cleanup: removed `PrimaryResult` / `SupplementResult`, updated `LookupResult`.
|
||||
- Tests: lookup tool and recorder tests.
|
||||
- Prompts: executor and chat executor guidance.
|
||||
- Docs/OpenSpec: modular RAG change files and architecture docs.
|
||||
|
||||
Known remaining risk:
|
||||
|
||||
- Runtime Agent prompt behavior should be demo-tested manually because the tool JSON contract changed from `primary/supplement` to `evidenceBlocks/contextPack/traces`.
|
||||
|
||||
## Review Results
|
||||
|
||||
Review finding:
|
||||
|
||||
- Session dedup returned `found=false` but still carried `evidenceBlocks` and `contextPack`, allowing the Agent to consume duplicate evidence despite the dedup message.
|
||||
|
||||
Fix:
|
||||
|
||||
- `LookupResultAssembler.deduped` now returns an empty evidence/context payload while preserving trace, relevance hint, and retrieved-domain memory.
|
||||
- Added `LookupKnowledgeToolTest.sessionDedupDoesNotReturnConsumableEvidenceAgain`.
|
||||
|
||||
Additional verification:
|
||||
|
||||
```text
|
||||
mvn -q -DskipTests compile
|
||||
PASS
|
||||
|
||||
mvn -q "-Dtest=LookupKnowledgeToolTest,ToolInvocationRecorderTest" test
|
||||
PASS
|
||||
|
||||
MILVUS_TOKEN=<application.yml milvus.token> mvn -q test
|
||||
PASS
|
||||
|
||||
openspec validate modular-rag-pipeline --strict
|
||||
PASS: Change 'modular-rag-pipeline' is valid
|
||||
|
||||
git diff --check
|
||||
PASS with LF/CRLF warnings only
|
||||
```
|
||||
|
||||
Updated full-suite note:
|
||||
|
||||
- `MilvusConnectionTest` passes when `MILVUS_TOKEN` is injected into the Maven process from `application.yml`.
|
||||
- `RedisSessionManagerTest` now passes after marking computed `SessionContext` getters as non-serialized JSON properties and ignoring unknown legacy Redis fields.
|
||||
@@ -0,0 +1,193 @@
|
||||
## Context
|
||||
|
||||
The current `lookup_knowledge` implementation is operational but still organized around a legacy L0/L1 result model. `LookupKnowledgeTool` currently performs query analysis, vector retrieval, relevance normalization, evidence block construction, session deduplication, and trace recording in one class. Evidence blocks already exist, but post-retrieval processing is not a first-class pipeline boundary.
|
||||
|
||||
The project direction is already documented as:
|
||||
|
||||
- keep `lookup_knowledge` as an explicit Agent evidence tool;
|
||||
- treat L0 as domain/entity/keyword hint data;
|
||||
- use L1 vector retrieval as the main evidence source;
|
||||
- keep Spring AI `VectorStore` behind `VectorSearchService`;
|
||||
- preserve `tool_invocation` as the evidence trace for Verifier, Eval, and trace APIs.
|
||||
|
||||
This change turns that architecture into code structure and updates the tool result contract so future Agent, Verifier, and Eval flows consume structured evidence instead of the old `primary` / `supplement` split.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Refactor `lookup_knowledge` into a modular RAG pipeline.
|
||||
- Keep L0 as query understanding and retrieval-control signal.
|
||||
- Keep L1 vector retrieval as the main document retrieval path.
|
||||
- Add a simple MVP fallback: retry raw query through unfiltered L1 when filtered L1 is low quality.
|
||||
- Move relevance normalization, evidence block creation, deduplication, and lightweight rerank into post-retrieval processing.
|
||||
- Add context packing as a first-class output.
|
||||
- Replace the old `primary` / `supplement` result contract with `evidenceBlocks`, `contextPack`, `retrievalTrace`, and `rerankTrace`.
|
||||
- Keep `lookup_knowledge` tool name and input argument unchanged.
|
||||
- Keep `tool_invocation` table schema stable and enrich `retrieval_details`.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not replace `lookup_knowledge` with an implicit Advisor.
|
||||
- Do not add a model-based reranker, cross-encoder, BM25, RRF, Elasticsearch, or OpenSearch.
|
||||
- Do not migrate document upload, chunking, embedding write path, Milvus schema, or vector collection layout.
|
||||
- Do not change when the Agent decides to call `lookup_knowledge`.
|
||||
- Do not introduce multi-query expansion in this change.
|
||||
|
||||
## Decisions
|
||||
|
||||
### Decision: Use explicit pipeline components
|
||||
|
||||
Create a local pipeline behind `LookupKnowledgeTool`:
|
||||
|
||||
```text
|
||||
LookupKnowledgeTool
|
||||
-> KnowledgeQueryTransformer
|
||||
-> KnowledgeDocumentRetriever
|
||||
-> KnowledgeEvidencePostProcessor
|
||||
-> KnowledgeContextPacker
|
||||
-> LookupResultAssembler
|
||||
-> ToolInvocationRecorder
|
||||
```
|
||||
|
||||
Rationale: this keeps the Agent tool boundary stable while making the RAG flow easy to test and explain. It also maps cleanly to Spring AI modular RAG concepts without hiding business observability inside an Advisor.
|
||||
|
||||
Alternative considered: keep all logic in `LookupKnowledgeTool` and only add fields. Rejected because the class would continue to mix query transformation, retrieval, post-processing, and trace responsibilities.
|
||||
|
||||
### Decision: L0 is a query transformer signal, not a main retriever
|
||||
|
||||
`KnowledgeQueryTransformer` will wrap the existing `KnowledgeIndexService.analyzeQuery` behavior and produce a transformed query object containing:
|
||||
|
||||
- original query
|
||||
- rewritten query, initially equal to the raw query unless a future rule rewrites it
|
||||
- domain hints
|
||||
- matched keywords
|
||||
- entities
|
||||
- optional category filter
|
||||
- L0 titles for trace only
|
||||
|
||||
L0 matches must not be converted into normal evidence candidates in the main path.
|
||||
|
||||
Rationale: L0 keyword/frontmatter matching is useful for controlling retrieval, but it is not reliable enough to be treated as fact evidence when L1 cannot support it.
|
||||
|
||||
Alternative considered: combine L0 and L1 into one candidate list. Rejected because it makes L0 an equal retrieval layer again and conflicts with the desired architecture.
|
||||
|
||||
### Decision: Use filtered L1 first, then unfiltered L1 retry
|
||||
|
||||
`KnowledgeDocumentRetriever` will perform:
|
||||
|
||||
```text
|
||||
attempt 1: vector search with L0-derived category filter, when unambiguous
|
||||
attempt 2: raw query vector search without the L0-derived filter, when attempt 1 is low quality
|
||||
```
|
||||
|
||||
Filtered retrieval is low quality when any of the following is true:
|
||||
|
||||
- no candidates are returned;
|
||||
- post-processing would produce zero evidence blocks;
|
||||
- top candidate normalized similarity is below `retrieval.normalization.reference-threshold`.
|
||||
|
||||
Rationale: the most likely MVP failure mode is an over-strict or wrong metadata filter. A raw unfiltered vector retry addresses that without adding a complex multi-stage fallback policy.
|
||||
|
||||
Alternative considered: return L0 documents as weak fallback evidence. Rejected for this MVP because it can let keyword hints masquerade as factual evidence.
|
||||
|
||||
### Decision: Keep rerank rule-based
|
||||
|
||||
`KnowledgeEvidencePostProcessor` will rerank with deterministic signals:
|
||||
|
||||
- vector score / normalized similarity;
|
||||
- domain match with query hints;
|
||||
- entity match;
|
||||
- keyword match;
|
||||
- source type or metadata priority when available;
|
||||
- retrieval attempt, with filtered hits not automatically preferred over stronger unfiltered hits.
|
||||
|
||||
The rerank trace should explain major score contributions per final evidence block.
|
||||
|
||||
Rationale: a rule-based reranker is explainable, cheap, testable, and enough for the interview-oriented MVP. It also avoids introducing model latency and new dependencies.
|
||||
|
||||
Alternative considered: model-based rerank. Rejected as out of scope until evaluation shows a need.
|
||||
|
||||
### Decision: Context pack becomes the Agent-facing content
|
||||
|
||||
`KnowledgeContextPacker` will turn final evidence blocks into a compact context package:
|
||||
|
||||
```text
|
||||
packedText
|
||||
strategy
|
||||
charBudget
|
||||
usedChars
|
||||
includedSources
|
||||
omittedSources
|
||||
```
|
||||
|
||||
MVP packing uses a character budget rather than exact token counting. The packer preserves source/title/breadcrumb/hit reasons and truncates content only after preserving metadata.
|
||||
|
||||
Rationale: the Agent should consume curated evidence context instead of inferring semantics from `primary` and `supplement`.
|
||||
|
||||
Alternative considered: keep `primary` and `supplement` as the Agent-facing fields. Rejected because those names preserve the outdated L0 primary / L1 supplement model.
|
||||
|
||||
### Decision: Break the old result contract deliberately
|
||||
|
||||
`LookupResult` will move to the preferred contract:
|
||||
|
||||
```text
|
||||
found
|
||||
evidenceBlocks
|
||||
contextPack
|
||||
retrievalTrace
|
||||
rerankTrace
|
||||
relevanceLevel
|
||||
completenessHint
|
||||
retrievedDomainsThisSession
|
||||
message
|
||||
```
|
||||
|
||||
`primary` and `supplement` may be removed in this change, and `PrimaryResult` / `SupplementResult` may be deleted if no production code remains after migration.
|
||||
|
||||
Rationale: this is an MVP project intended to demonstrate clean modular RAG. Keeping obsolete fields would force later code to preserve misleading semantics.
|
||||
|
||||
Alternative considered: add new fields while keeping old fields deprecated. Rejected because the user explicitly accepted removing old fields if the later flow becomes cleaner.
|
||||
|
||||
### Decision: Preserve persistence compatibility at table level
|
||||
|
||||
`ToolInvocationRecorder` will stop depending on `result.getPrimary()` for output preview. Preview should come from `contextPack.packedText` or the top evidence block content. `retrieval_details` will include compact summaries:
|
||||
|
||||
- query transform summary;
|
||||
- retrieval trace with attempts and fallback reason;
|
||||
- rerank trace summary;
|
||||
- context pack summary;
|
||||
- evidence block summaries.
|
||||
|
||||
Rationale: trace consumers already read `tool_invocation` rows. The database schema can remain stable while JSON details evolve.
|
||||
|
||||
Alternative considered: add columns for each new trace object. Rejected because the current trace model already stores retrieval-specific details in JSON and does not need schema churn for this change.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Breaking tool output shape can affect Agent prompt behavior. -> Mitigation: update tool description and executor prompt to prefer `contextPack` and `evidenceBlocks`.
|
||||
- [Risk] Existing tests assert `primary` / `supplement`. -> Mitigation: migrate tests to evidence/context/traces in the same change.
|
||||
- [Risk] Rule-based rerank can reorder evidence unexpectedly. -> Mitigation: persist `rerankTrace` and cover ordering behavior with tests.
|
||||
- [Risk] Context packing can omit useful evidence under a small budget. -> Mitigation: record included and omitted sources and keep the character budget configurable.
|
||||
- [Risk] Existing OpenSpec requirements conflict with the new fallback policy. -> Mitigation: update `rag-knowledge-retrieval` delta before apply and validate the change.
|
||||
- [Risk] Unfiltered retry may increase latency. -> Mitigation: retry only when filtered result is empty or below reference quality, and record attempt counts/duration in trace.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
1. Add new internal DTOs and pipeline components while keeping `LookupKnowledgeTool` as the public tool bean.
|
||||
2. Migrate `LookupKnowledgeTool` orchestration to the pipeline.
|
||||
3. Update `LookupResult` to the new result contract and remove `primary` / `supplement` usages.
|
||||
4. Update `ToolInvocationRecorder` preview and retrieval details to use context/evidence/traces.
|
||||
5. Update prompt text and tool description.
|
||||
6. Update tests for the new contract.
|
||||
7. Run targeted unit tests and OpenSpec validation.
|
||||
|
||||
Rollback strategy:
|
||||
|
||||
- Revert the change as one unit if Agent tool behavior regresses.
|
||||
- Because the table schema remains stable and the tool input is unchanged, rollback does not require database migration.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- The exact default context pack character budget should be finalized during implementation; recommended MVP default is 3000 to 5000 characters.
|
||||
- Whether to delete `PrimaryResult` and `SupplementResult` immediately depends on final compile-time references after migration.
|
||||
@@ -0,0 +1,172 @@
|
||||
# Modular RAG Pipeline Proposal
|
||||
|
||||
## Problem
|
||||
|
||||
`lookup_knowledge` already exposes structured retrieval evidence, but the runtime flow is still concentrated inside `LookupKnowledgeTool`. Query understanding, vector retrieval, relevance normalization, evidence block construction, session deduplication, and trace recording are tightly coupled. This makes the RAG path harder to explain, test, evolve, and present as a modular Agent engineering design.
|
||||
|
||||
The current result contract also still carries the old `primary` / `supplement` model, where `primary` means L0 exact match and `supplement` means L1 semantic retrieval. That contract no longer matches the intended architecture: L0 should be a query understanding and retrieval-control signal, while L1 vector retrieval should be the main evidence source.
|
||||
|
||||
## Proposed Change
|
||||
|
||||
Refactor `lookup_knowledge` into a modular RAG pipeline while preserving the explicit Agent tool boundary and `tool_invocation` evidence trace.
|
||||
|
||||
Target pipeline:
|
||||
|
||||
```text
|
||||
LookupKnowledgeTool
|
||||
-> KnowledgeQueryTransformer
|
||||
-> KnowledgeDocumentRetriever
|
||||
-> KnowledgeEvidencePostProcessor
|
||||
-> KnowledgeContextPacker
|
||||
-> LookupResultAssembler
|
||||
-> ToolInvocationRecorder
|
||||
```
|
||||
|
||||
### Query Transformation
|
||||
|
||||
Introduce a query transformer around the current L0 analysis.
|
||||
|
||||
L0 SHALL provide:
|
||||
|
||||
- domain/category hints
|
||||
- matched keywords
|
||||
- extracted entities
|
||||
- optional metadata filter candidate
|
||||
- traceable query understanding data
|
||||
|
||||
L0 SHALL NOT act as a main evidence retrieval path in the normal flow.
|
||||
|
||||
### Retrieval
|
||||
|
||||
L1 vector retrieval remains the main document retrieval path through `VectorSearchService`, which already supports Spring AI `VectorStore` as the preferred path and Milvus SDK fallback.
|
||||
|
||||
MVP fallback strategy:
|
||||
|
||||
```text
|
||||
1. Run filtered vector retrieval with the L0-derived category filter when unambiguous.
|
||||
2. If filtered retrieval returns no usable evidence or low-quality evidence, retry raw query through unfiltered vector retrieval.
|
||||
3. If unfiltered retrieval also fails, return no_evidence.
|
||||
```
|
||||
|
||||
The fallback SHALL skip the L0-derived filter rather than returning L0 documents as fact evidence.
|
||||
|
||||
### Post-Retrieval Processing
|
||||
|
||||
Move evidence post-processing out of `LookupKnowledgeTool`.
|
||||
|
||||
The post-processor SHALL handle:
|
||||
|
||||
- relevance normalization
|
||||
- evidence block creation
|
||||
- source-level deduplication
|
||||
- rule-based lightweight rerank
|
||||
- retrieval trace and rerank trace generation
|
||||
|
||||
The first rerank implementation should be rule-based, using available signals such as vector score, domain match, entity match, keyword match, source type, and whether evidence aligns with query hints.
|
||||
|
||||
### Context Packing
|
||||
|
||||
Add a context packing step that converts final evidence blocks into an Agent-facing context package.
|
||||
|
||||
The packer SHALL:
|
||||
|
||||
- keep source/title/breadcrumb visible
|
||||
- obey a configurable character budget in the MVP
|
||||
- prioritize reranked evidence order
|
||||
- avoid duplicate source content
|
||||
- produce a compact summary of included and omitted evidence
|
||||
|
||||
### Result Contract
|
||||
|
||||
This change intentionally updates the `lookup_knowledge` return contract.
|
||||
|
||||
New preferred contract:
|
||||
|
||||
```text
|
||||
found
|
||||
evidenceBlocks
|
||||
contextPack
|
||||
retrievalTrace
|
||||
rerankTrace
|
||||
relevanceLevel
|
||||
completenessHint
|
||||
retrievedDomainsThisSession
|
||||
message
|
||||
```
|
||||
|
||||
The old `primary` and `supplement` fields may be removed as part of this change, because they encode the outdated assumption that L0 is the primary evidence source and L1 is supplemental evidence.
|
||||
|
||||
## Scope
|
||||
|
||||
In scope:
|
||||
|
||||
- Refactor `LookupKnowledgeTool` into a thin tool boundary and pipeline orchestrator.
|
||||
- Add local pipeline classes and DTOs for query transformation, retrieval result normalization, post-processing, context packing, and traces.
|
||||
- Update `LookupResult` to prefer `evidenceBlocks`, `contextPack`, `retrievalTrace`, and `rerankTrace`.
|
||||
- Remove or deprecate `primary` / `supplement` according to the final spec.
|
||||
- Update `ToolInvocationRecorder` to persist compact summaries for evidence blocks, context pack, retrieval trace, fallback reason, and rerank trace.
|
||||
- Update tool description / prompt wording so Agent behavior matches the new contract.
|
||||
- Update tests for filtered retrieval, unfiltered retry, rerank, context packing, evidence persistence, and result contract changes.
|
||||
- Update `rag-knowledge-retrieval` spec to remove L0 primary fallback and old compatibility-field requirements.
|
||||
|
||||
Out of scope:
|
||||
|
||||
- Replacing the explicit `lookup_knowledge` tool with an implicit Advisor.
|
||||
- Introducing a model-based reranker or cross-encoder.
|
||||
- Introducing BM25, RRF, Elasticsearch, or OpenSearch.
|
||||
- Migrating document upload, chunking, embedding writes, or Milvus schema.
|
||||
- Changing the Agent decision of when to call `lookup_knowledge`.
|
||||
|
||||
## Interface Impact
|
||||
|
||||
Impact level: L4 breaking interface change.
|
||||
|
||||
Reason:
|
||||
|
||||
- `LookupResult.primary` and `LookupResult.supplement` may be removed.
|
||||
- The JSON returned by the `lookup_knowledge` Agent tool changes shape.
|
||||
- Tests and internal consumers that read `primary` / `supplement` must migrate to `evidenceBlocks` and `contextPack`.
|
||||
|
||||
Known affected areas:
|
||||
|
||||
- `LookupKnowledgeTool`
|
||||
- `LookupResult`
|
||||
- `PrimaryResult` / `SupplementResult`
|
||||
- `ToolInvocationRecorder`
|
||||
- `LookupKnowledgeToolTest`
|
||||
- `ToolInvocationRecorderTest`
|
||||
- Agent tool prompt / executor prompt references
|
||||
- `rag-knowledge-retrieval` OpenSpec requirements
|
||||
- RAG docs under `mvp/architecture/`
|
||||
|
||||
Migration approach:
|
||||
|
||||
- Update all in-repo consumers in the same change.
|
||||
- Keep `lookup_knowledge` tool name and input argument unchanged.
|
||||
- Keep `tool_invocation` persistence compatible at table level while enriching `retrieval_details`.
|
||||
- Record fallback and no-evidence semantics explicitly so Verifier and Eval do not treat hint-only data as fact evidence.
|
||||
|
||||
## Context Constraints
|
||||
|
||||
Relevant project decisions:
|
||||
|
||||
- `lookup_knowledge` remains an explicit Agent evidence tool.
|
||||
- L0 is already documented as a hint layer, not a final decision layer.
|
||||
- Spring AI `VectorStore` is already the preferred retrieval path behind `VectorSearchService`.
|
||||
- Milvus SDK fallback remains valuable for MVP resilience.
|
||||
- `tool_invocation` is the stable evidence trace used by trace inspection, Verifier, and Eval.
|
||||
- Existing `rag-knowledge-retrieval` spec still contains legacy fallback and compatibility-field requirements that must be changed.
|
||||
|
||||
## Risks
|
||||
|
||||
- Breaking result contract may affect prompt behavior because the tool JSON changes.
|
||||
- Removing `primary` / `supplement` requires updating tests and recorder preview logic.
|
||||
- Rule-based rerank may create unexpected ordering changes if score semantics are not handled carefully.
|
||||
- Context packing can hide useful evidence if the budget is too small.
|
||||
- Existing OpenSpec requirements conflict with the new L0 fallback policy and must be updated before implementation.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- What exact minimum fields should `ContextPack`, `RetrievalTrace`, and `RerankTrace` expose in the committed spec?
|
||||
- Should `PrimaryResult` and `SupplementResult` classes be deleted immediately or left deprecated for one change cycle?
|
||||
- What threshold defines "filtered retrieval low quality" for triggering unfiltered retry?
|
||||
+121
@@ -0,0 +1,121 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL use a modular RAG pipeline
|
||||
The `lookup_knowledge` tool SHALL route each request through explicit query transformation, vector retrieval, post-retrieval processing, context packing, result assembly, and trace recording components.
|
||||
|
||||
#### Scenario: Pipeline components execute in order
|
||||
- **WHEN** `lookup_knowledge` receives a query
|
||||
- **THEN** the system SHALL transform the query before retrieval
|
||||
- **AND** it SHALL retrieve vector candidates before post-processing
|
||||
- **AND** it SHALL build evidence blocks before context packing
|
||||
- **AND** it SHALL record trace details after result assembly
|
||||
|
||||
#### Scenario: Tool boundary remains explicit
|
||||
- **WHEN** the modular pipeline is used
|
||||
- **THEN** the Agent SHALL still call the explicit `lookup_knowledge` tool with the same query argument
|
||||
- **AND** the implementation SHALL NOT require an implicit Advisor to inject knowledge into every chat response
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL retry without L0 filter when filtered L1 is low quality
|
||||
The retrieval flow SHALL treat L0-derived category filtering as an optimization, not as a hard dependency for final recall.
|
||||
|
||||
#### Scenario: Filtered retrieval succeeds
|
||||
- **WHEN** L0 provides an unambiguous category filter
|
||||
- **AND** filtered L1 retrieval returns usable evidence at or above the configured reference threshold
|
||||
- **THEN** the tool SHALL use the filtered L1 candidates without running an unfiltered retry
|
||||
|
||||
#### Scenario: Filtered retrieval returns no evidence
|
||||
- **WHEN** L0 provides a category filter
|
||||
- **AND** filtered L1 retrieval returns no candidates or no final evidence blocks
|
||||
- **THEN** the tool SHALL retry L1 retrieval with the raw query and no L0-derived category filter
|
||||
- **AND** the retrieval trace SHALL record fallback reason `filtered_vector_no_evidence`
|
||||
|
||||
#### Scenario: Filtered retrieval is below reference quality
|
||||
- **WHEN** L0 provides a category filter
|
||||
- **AND** filtered L1 retrieval returns candidates whose top normalized similarity is below the configured reference threshold
|
||||
- **THEN** the tool SHALL retry L1 retrieval with the raw query and no L0-derived category filter
|
||||
- **AND** the retrieval trace SHALL record fallback reason `filtered_vector_low_quality`
|
||||
|
||||
#### Scenario: Both retrieval attempts fail
|
||||
- **WHEN** filtered L1 retrieval and unfiltered L1 retry both produce no usable evidence
|
||||
- **THEN** the tool SHALL return `found=false`
|
||||
- **AND** the tool SHALL set evidence status to `no_evidence`
|
||||
- **AND** the tool SHALL NOT return L0 documents as fact evidence
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL return an evidence-first result contract
|
||||
The `lookup_knowledge` result SHALL expose structured evidence and packed context as the preferred contract.
|
||||
|
||||
#### Scenario: Evidence result contains context and traces
|
||||
- **WHEN** `lookup_knowledge` returns usable evidence
|
||||
- **THEN** the result SHALL include `evidenceBlocks`
|
||||
- **AND** it SHALL include `contextPack`
|
||||
- **AND** it SHALL include `retrievalTrace`
|
||||
- **AND** it SHALL include `rerankTrace`
|
||||
- **AND** it SHALL include `relevanceLevel` and `completenessHint`
|
||||
|
||||
#### Scenario: No-evidence result keeps traceability
|
||||
- **WHEN** `lookup_knowledge` returns no usable evidence
|
||||
- **THEN** the result SHALL include `found=false`
|
||||
- **AND** it SHALL include a message explaining that no knowledge evidence was found
|
||||
- **AND** it SHALL include retrieval trace details for attempted retrieval paths
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL pack evidence context for Agent consumption
|
||||
The post-retrieval flow SHALL convert final evidence blocks into a compact context package for the Agent.
|
||||
|
||||
#### Scenario: Context pack preserves source metadata
|
||||
- **WHEN** evidence blocks are packed
|
||||
- **THEN** the packed context SHALL preserve source, title when available, breadcrumb when available, and hit reasons for included evidence
|
||||
|
||||
#### Scenario: Context pack respects budget
|
||||
- **WHEN** final evidence content exceeds the configured context budget
|
||||
- **THEN** the packer SHALL truncate content rather than source metadata
|
||||
- **AND** it SHALL record included and omitted sources in the context pack summary
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL rerank evidence with traceable rule signals
|
||||
The post-retrieval flow SHALL rerank vector candidates using deterministic rule-based signals and expose the explanation.
|
||||
|
||||
#### Scenario: Rerank trace records score contributions
|
||||
- **WHEN** candidates are reranked
|
||||
- **THEN** the rerank trace SHALL record final rank, source, base retrieval score when available, and major boost reasons for top evidence blocks
|
||||
|
||||
#### Scenario: Query hints influence rerank without becoming evidence
|
||||
- **WHEN** L0 query hints match candidate metadata or content
|
||||
- **THEN** the reranker MAY boost the candidate
|
||||
- **AND** the evidence block SHALL record the hint as a hit reason
|
||||
- **AND** the system SHALL NOT treat the L0 hint itself as fact evidence
|
||||
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL keep L0 as a hint provider
|
||||
The `lookup_knowledge` retrieval flow SHALL retain L0 keyword/frontmatter matching but use it only as query understanding, filtering, rerank, and explainability hint data rather than as a final evidence retrieval decision.
|
||||
|
||||
#### Scenario: L0 produces traceable hint data
|
||||
- **WHEN** L0 matches one or more indexed knowledge entries
|
||||
- **THEN** the retrieval flow SHALL expose matched titles, matched keywords, domains or categories, and entity terms as structured hint data
|
||||
|
||||
#### Scenario: L0 does not bypass semantic retrieval
|
||||
- **WHEN** L0 returns exactly one match
|
||||
- **THEN** the retrieval flow SHALL still attempt semantic L1 retrieval unless L1 is explicitly disabled by configuration
|
||||
|
||||
#### Scenario: L0 hints do not become normal evidence
|
||||
- **WHEN** L1 retrieval returns no usable evidence
|
||||
- **THEN** L0 matched documents SHALL NOT be returned as fact evidence blocks
|
||||
- **AND** L0 hint data MAY still be recorded in retrieval trace details
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL return structured evidence blocks
|
||||
The `lookup_knowledge` retrieval flow SHALL expose retrieved evidence as structured evidence blocks.
|
||||
|
||||
#### Scenario: Evidence block contains source metadata
|
||||
- **WHEN** a `lookup_knowledge` call returns evidence
|
||||
- **THEN** each evidence block SHALL include source, title when available, breadcrumb when available, retrieval layer, content, and hit reasons
|
||||
|
||||
#### Scenario: Evidence blocks are the primary evidence contract
|
||||
- **WHEN** evidence blocks are returned
|
||||
- **THEN** Agent-facing knowledge content SHALL be derived from evidence blocks and context pack
|
||||
- **AND** the result SHALL NOT rely on legacy `primary` or `supplement` fields for L0/L1 meaning
|
||||
|
||||
## REMOVED Requirements
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL preserve fallback evidence
|
||||
**Reason**: The new MVP fallback strategy retries unfiltered L1 retrieval when L0-derived filtering appears to suppress usable semantic evidence. Returning L0 documents as fallback fact evidence can let keyword hints masquerade as verified knowledge.
|
||||
|
||||
**Migration**: Use the new requirement "Knowledge retrieval SHALL retry without L0 filter when filtered L1 is low quality". If filtered and unfiltered L1 both fail, return `no_evidence` while preserving L0 hint data in trace details only.
|
||||
@@ -0,0 +1,51 @@
|
||||
## 1. Pipeline Models
|
||||
|
||||
- [x] 1.1 Add query transformation model for original query, rewritten query, domain hints, matched keywords, entities, category filter, and trace-only L0 titles.
|
||||
- [x] 1.2 Add retrieved evidence candidate model that normalizes vector result metadata, score semantics, source, title, breadcrumb, content, retrieval attempt, and original rank.
|
||||
- [x] 1.3 Add context pack, retrieval trace, and rerank trace DTOs for the new `LookupResult` contract.
|
||||
- [x] 1.4 Update `LookupResult` to expose evidence blocks, context pack, retrieval trace, rerank trace, relevance level, completeness hint, retrieved domains, and message.
|
||||
|
||||
## 2. Query And Retrieval Pipeline
|
||||
|
||||
- [x] 2.1 Implement `KnowledgeQueryTransformer` by reusing `KnowledgeIndexService.analyzeQuery` and mapping L0 output to query hints and optional category filter.
|
||||
- [x] 2.2 Implement `KnowledgeDocumentRetriever` as a wrapper around `VectorSearchService` for filtered and unfiltered vector retrieval attempts.
|
||||
- [x] 2.3 Implement low-quality detection using empty candidates, empty final evidence, or top normalized similarity below `retrieval.normalization.reference-threshold`.
|
||||
- [x] 2.4 Implement unfiltered raw-query retry when filtered retrieval is low quality and record fallback reason in retrieval trace.
|
||||
|
||||
## 3. Post-Retrieval Processing
|
||||
|
||||
- [x] 3.1 Move relevance normalization out of `LookupKnowledgeTool` into `KnowledgeEvidencePostProcessor`.
|
||||
- [x] 3.2 Move evidence block creation and source-level deduplication out of `LookupKnowledgeTool` into the post-processor.
|
||||
- [x] 3.3 Implement rule-based lightweight rerank using vector similarity, domain match, entity match, keyword match, and metadata/source-type signals.
|
||||
- [x] 3.4 Ensure L0 hints influence filter/rerank/trace only and are not returned as standalone fact evidence when L1 has no usable evidence.
|
||||
|
||||
## 4. Context Packing And Result Assembly
|
||||
|
||||
- [x] 4.1 Implement `KnowledgeContextPacker` with a configurable MVP character budget.
|
||||
- [x] 4.2 Pack evidence blocks while preserving source, title, breadcrumb, and hit reasons before truncating content.
|
||||
- [x] 4.3 Implement `LookupResultAssembler` to build evidence-first results for usable evidence, no-evidence, and session dedup cases.
|
||||
- [x] 4.4 Remove or migrate all `primary` and `supplement` result usage from production code.
|
||||
|
||||
## 5. Tool Boundary And Trace Recording
|
||||
|
||||
- [x] 5.1 Refactor `LookupKnowledgeTool` into a thin orchestrator that invokes the pipeline and handles tool boundary concerns.
|
||||
- [x] 5.2 Update `ToolInvocationRecorder.LookupKnowledgeRecord` to summarize context pack, retrieval trace, rerank trace, fallback reason, and evidence blocks without relying on `result.getPrimary()`.
|
||||
- [x] 5.3 Preserve stable `tool_invocation` table fields and store new retrieval details in JSON.
|
||||
- [x] 5.4 Update `@Tool` description and relevant executor prompt text to describe evidence blocks, context pack, and no-realtime-data boundaries.
|
||||
|
||||
## 6. Tests And Evaluation
|
||||
|
||||
- [x] 6.1 Update `LookupKnowledgeToolTest` for evidence-first result contract and removal of `primary` / `supplement`.
|
||||
- [x] 6.2 Add tests for filtered L1 success without retry.
|
||||
- [x] 6.3 Add tests for filtered low-quality retrieval triggering raw unfiltered L1 retry.
|
||||
- [x] 6.4 Add tests proving L0 hint data does not become standalone fact evidence when L1 fails.
|
||||
- [x] 6.5 Add tests for rerank ordering, rerank trace, context pack budget behavior, and source metadata preservation.
|
||||
- [x] 6.6 Update `ToolInvocationRecorderTest` for new retrieval detail summaries and output preview source.
|
||||
- [x] 6.7 Run targeted Java tests for lookup knowledge and recorder changes.
|
||||
- [x] 6.8 Run OpenSpec validation for `modular-rag-pipeline`.
|
||||
|
||||
## 7. Documentation Cleanup
|
||||
|
||||
- [x] 7.1 Update RAG architecture docs to reflect modular pipeline, unfiltered vector retry, and evidence-first contract.
|
||||
- [x] 7.2 Update retrieval observability docs to remove L0 primary fallback and `primary` / `supplement` compatibility language.
|
||||
- [x] 7.3 Review git diff to confirm only expected RAG, prompt, test, and spec files changed.
|
||||
@@ -4,15 +4,20 @@
|
||||
Define the runtime contract for the explicit `lookup_knowledge` Agent tool, including how L0 keyword/frontmatter hints cooperate with L1 semantic retrieval while preserving metadata filters, fallback evidence, and traceable retrieval details.
|
||||
## Requirements
|
||||
### Requirement: Knowledge retrieval SHALL keep L0 as a hint provider
|
||||
The `lookup_knowledge` retrieval flow SHALL retain L0 keyword/frontmatter matching but use it as domain, entity, and explainability hint data rather than as the sole final retrieval decision.
|
||||
The `lookup_knowledge` retrieval flow SHALL retain L0 keyword/frontmatter matching but use it only as query understanding, filtering, rerank, and explainability hint data rather than as a final evidence retrieval decision.
|
||||
|
||||
#### Scenario: L0 produces traceable hint data
|
||||
- **WHEN** L0 matches one or more indexed knowledge entries
|
||||
- **THEN** the retrieval flow SHALL expose matched titles, matched keywords, domains or categories, and entity terms as structured hint data
|
||||
|
||||
#### Scenario: L0 does not bypass semantic retrieval by default
|
||||
#### Scenario: L0 does not bypass semantic retrieval
|
||||
- **WHEN** L0 returns exactly one match
|
||||
- **THEN** the retrieval flow SHALL still attempt semantic L1 retrieval unless L1 is unavailable or explicitly disabled by configuration
|
||||
- **THEN** the retrieval flow SHALL still attempt semantic L1 retrieval unless L1 is explicitly disabled by configuration
|
||||
|
||||
#### Scenario: L0 hints do not become normal evidence
|
||||
- **WHEN** L1 retrieval returns no usable evidence
|
||||
- **THEN** L0 matched documents SHALL NOT be returned as fact evidence blocks
|
||||
- **AND** L0 hint data MAY still be recorded in retrieval trace details
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL use L0 domain as optional L1 filter
|
||||
The retrieval flow SHALL use L0 domain/category information as an optional metadata filter for L1 retrieval when the domain is unambiguous.
|
||||
@@ -25,19 +30,6 @@ The retrieval flow SHALL use L0 domain/category information as an optional metad
|
||||
- **WHEN** L0 hint data contains zero domains or multiple domains
|
||||
- **THEN** the L1 retrieval request SHALL run without an L0-derived category filter
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL preserve fallback evidence
|
||||
The retrieval flow SHALL still return useful L0 evidence when L1 produces no usable result.
|
||||
|
||||
#### Scenario: L1 has no results
|
||||
- **WHEN** L0 has at least one match and L1 returns no candidates
|
||||
- **THEN** the tool SHALL return an L0-based primary result
|
||||
- **AND** the relevance assessment SHALL not claim semantic support from L1
|
||||
|
||||
#### Scenario: L1 fails
|
||||
- **WHEN** L0 has at least one match and L1 retrieval throws or fails
|
||||
- **THEN** the tool SHALL return an L0-based primary result
|
||||
- **AND** the tool invocation record SHALL preserve the L0 hint details
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL persist L0 hints
|
||||
The system SHALL persist L0 hint details in `tool_invocation.retrieval_details` for `lookup_knowledge` calls.
|
||||
|
||||
@@ -50,15 +42,16 @@ The system SHALL persist L0 hint details in `tool_invocation.retrieval_details`
|
||||
- **THEN** the recorded retrieval layer SHALL be `L0+L1`
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL return structured evidence blocks
|
||||
The `lookup_knowledge` retrieval flow SHALL expose retrieved evidence as structured evidence blocks in addition to the existing compatibility fields.
|
||||
The `lookup_knowledge` retrieval flow SHALL expose retrieved evidence as structured evidence blocks.
|
||||
|
||||
#### Scenario: Evidence block contains source metadata
|
||||
- **WHEN** a `lookup_knowledge` call returns evidence
|
||||
- **THEN** each evidence block SHALL include source, title when available, breadcrumb when available, retrieval layer, content, and hit reasons
|
||||
|
||||
#### Scenario: Compatibility fields remain available
|
||||
#### Scenario: Evidence blocks are the primary evidence contract
|
||||
- **WHEN** evidence blocks are returned
|
||||
- **THEN** the existing `primary` and `supplement` result fields SHALL remain available when their source evidence exists
|
||||
- **THEN** Agent-facing knowledge content SHALL be derived from evidence blocks and context pack
|
||||
- **AND** the result SHALL NOT rely on legacy `primary` or `supplement` fields for L0/L1 meaning
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL deduplicate evidence blocks
|
||||
The retrieval flow SHALL remove duplicate evidence blocks before returning them to the Agent.
|
||||
@@ -151,3 +144,86 @@ Spring AI Milvus integration SHALL be configured to use the existing collection
|
||||
- **WHEN** Spring AI Milvus VectorStore is configured
|
||||
- **THEN** it SHALL use the existing id, content, vector, and metadata field names
|
||||
- **AND** it SHALL use the configured embedding dimension and metric type compatible with existing vectors
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL use a modular RAG pipeline
|
||||
The `lookup_knowledge` tool SHALL route each request through explicit query transformation, vector retrieval, post-retrieval processing, context packing, result assembly, and trace recording components.
|
||||
|
||||
#### Scenario: Pipeline components execute in order
|
||||
- **WHEN** `lookup_knowledge` receives a query
|
||||
- **THEN** the system SHALL transform the query before retrieval
|
||||
- **AND** it SHALL retrieve vector candidates before post-processing
|
||||
- **AND** it SHALL build evidence blocks before context packing
|
||||
- **AND** it SHALL record trace details after result assembly
|
||||
|
||||
#### Scenario: Tool boundary remains explicit
|
||||
- **WHEN** the modular pipeline is used
|
||||
- **THEN** the Agent SHALL still call the explicit `lookup_knowledge` tool with the same query argument
|
||||
- **AND** the implementation SHALL NOT require an implicit Advisor to inject knowledge into every chat response
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL retry without L0 filter when filtered L1 is low quality
|
||||
The retrieval flow SHALL treat L0-derived category filtering as an optimization, not as a hard dependency for final recall.
|
||||
|
||||
#### Scenario: Filtered retrieval succeeds
|
||||
- **WHEN** L0 provides an unambiguous category filter
|
||||
- **AND** filtered L1 retrieval returns usable evidence at or above the configured reference threshold
|
||||
- **THEN** the tool SHALL use the filtered L1 candidates without running an unfiltered retry
|
||||
|
||||
#### Scenario: Filtered retrieval returns no evidence
|
||||
- **WHEN** L0 provides a category filter
|
||||
- **AND** filtered L1 retrieval returns no candidates or no final evidence blocks
|
||||
- **THEN** the tool SHALL retry L1 retrieval with the raw query and no L0-derived category filter
|
||||
- **AND** the retrieval trace SHALL record fallback reason `filtered_vector_no_evidence`
|
||||
|
||||
#### Scenario: Filtered retrieval is below reference quality
|
||||
- **WHEN** L0 provides a category filter
|
||||
- **AND** filtered L1 retrieval returns candidates whose top normalized similarity is below the configured reference threshold
|
||||
- **THEN** the tool SHALL retry L1 retrieval with the raw query and no L0-derived category filter
|
||||
- **AND** the retrieval trace SHALL record fallback reason `filtered_vector_low_quality`
|
||||
|
||||
#### Scenario: Both retrieval attempts fail
|
||||
- **WHEN** filtered L1 retrieval and unfiltered L1 retry both produce no usable evidence
|
||||
- **THEN** the tool SHALL return `found=false`
|
||||
- **AND** the tool SHALL set evidence status to `no_evidence`
|
||||
- **AND** the tool SHALL NOT return L0 documents as fact evidence
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL return an evidence-first result contract
|
||||
The `lookup_knowledge` result SHALL expose structured evidence and packed context as the preferred contract.
|
||||
|
||||
#### Scenario: Evidence result contains context and traces
|
||||
- **WHEN** `lookup_knowledge` returns usable evidence
|
||||
- **THEN** the result SHALL include `evidenceBlocks`
|
||||
- **AND** it SHALL include `contextPack`
|
||||
- **AND** it SHALL include `retrievalTrace`
|
||||
- **AND** it SHALL include `rerankTrace`
|
||||
- **AND** it SHALL include `relevanceLevel` and `completenessHint`
|
||||
|
||||
#### Scenario: No-evidence result keeps traceability
|
||||
- **WHEN** `lookup_knowledge` returns no usable evidence
|
||||
- **THEN** the result SHALL include `found=false`
|
||||
- **AND** it SHALL include a message explaining that no knowledge evidence was found
|
||||
- **AND** it SHALL include retrieval trace details for attempted retrieval paths
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL pack evidence context for Agent consumption
|
||||
The post-retrieval flow SHALL convert final evidence blocks into a compact context package for the Agent.
|
||||
|
||||
#### Scenario: Context pack preserves source metadata
|
||||
- **WHEN** evidence blocks are packed
|
||||
- **THEN** the packed context SHALL preserve source, title when available, breadcrumb when available, and hit reasons for included evidence
|
||||
|
||||
#### Scenario: Context pack respects budget
|
||||
- **WHEN** final evidence content exceeds the configured context budget
|
||||
- **THEN** the packer SHALL truncate content rather than source metadata
|
||||
- **AND** it SHALL record included and omitted sources in the context pack summary
|
||||
|
||||
### Requirement: Knowledge retrieval SHALL rerank evidence with traceable rule signals
|
||||
The post-retrieval flow SHALL rerank vector candidates using deterministic rule-based signals and expose the explanation.
|
||||
|
||||
#### Scenario: Rerank trace records score contributions
|
||||
- **WHEN** candidates are reranked
|
||||
- **THEN** the rerank trace SHALL record final rank, source, base retrieval score when available, and major boost reasons for top evidence blocks
|
||||
|
||||
#### Scenario: Query hints influence rerank without becoming evidence
|
||||
- **WHEN** L0 query hints match candidate metadata or content
|
||||
- **THEN** the reranker MAY boost the candidate
|
||||
- **AND** the evidence block SHALL record the hint as a hit reason
|
||||
- **AND** the system SHALL NOT treat the L0 hint itself as fact evidence
|
||||
|
||||
+499
-22
@@ -2,7 +2,8 @@
|
||||
"""Offline evaluator for RAG retrieval golden cases.
|
||||
|
||||
The evaluator reads fixed golden cases and saved retrieval fixtures. It does not
|
||||
call the running application or any external service.
|
||||
call the running application or any external service. Fixtures must use the
|
||||
modular `LookupResult` shape produced by lookup_knowledge.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -19,6 +20,15 @@ DEFAULT_CASES = Path("eval/rag-retrieval/cases/golden-cases.json")
|
||||
DEFAULT_FIXTURES = Path("eval/rag-retrieval/fixtures")
|
||||
DEFAULT_JSON_REPORT = Path("eval/rag-retrieval/reports/baseline.json")
|
||||
DEFAULT_MD_REPORT = Path("eval/rag-retrieval/reports/baseline.md")
|
||||
DEFAULT_DIFF_JSON_REPORT = Path("eval/rag-retrieval/reports/baseline-diff.json")
|
||||
DEFAULT_DIFF_MD_REPORT = Path("eval/rag-retrieval/reports/baseline-diff.md")
|
||||
|
||||
HIT_LEVEL_RANK = {
|
||||
"miss": 0,
|
||||
"weak": 1,
|
||||
"medium": 2,
|
||||
"strong": 3,
|
||||
}
|
||||
|
||||
|
||||
@dataclass
|
||||
@@ -34,12 +44,16 @@ class Candidate:
|
||||
@classmethod
|
||||
def from_json(cls, raw: dict[str, Any], fallback_rank: int) -> "Candidate":
|
||||
return cls(
|
||||
rank=int(raw.get("rank") or fallback_rank),
|
||||
doc_id=str(raw.get("docId") or raw.get("id") or ""),
|
||||
rank=int(raw.get("rank") or raw.get("finalRank") or fallback_rank),
|
||||
doc_id=str(raw.get("docId") or raw.get("source") or raw.get("id") or ""),
|
||||
title=str(raw.get("title") or ""),
|
||||
breadcrumb=str(raw.get("breadcrumb") or ""),
|
||||
content=str(raw.get("content") or ""),
|
||||
score=_optional_float(raw.get("score")),
|
||||
content=str(raw.get("content") or raw.get("contentPreview") or ""),
|
||||
score=_optional_float(
|
||||
raw.get("score")
|
||||
if raw.get("score") is not None
|
||||
else raw.get("finalScore")
|
||||
),
|
||||
retrieval_layer=(
|
||||
str(raw.get("retrievalLayer"))
|
||||
if raw.get("retrievalLayer") is not None
|
||||
@@ -57,6 +71,18 @@ class Candidate:
|
||||
return f"{self.rank}:{label}"
|
||||
|
||||
|
||||
@dataclass
|
||||
class NormalizedFixture:
|
||||
data_shape: str
|
||||
candidates: list[Candidate]
|
||||
selected_attempt: str | None
|
||||
fallback_reason: str | None
|
||||
evidence_status: str | None
|
||||
included_sources: list[str]
|
||||
omitted_sources: list[str]
|
||||
rerank_top_source: str | None
|
||||
|
||||
|
||||
def _optional_float(value: Any) -> float | None:
|
||||
if value is None:
|
||||
return None
|
||||
@@ -88,10 +114,76 @@ def normalize_terms(values: list[Any]) -> list[str]:
|
||||
return [str(value).lower() for value in values if str(value).strip()]
|
||||
|
||||
|
||||
def normalize_sources(values: list[Any]) -> list[str]:
|
||||
return [str(value) for value in values if str(value).strip()]
|
||||
|
||||
|
||||
def get_lookup_result(fixture: dict[str, Any]) -> dict[str, Any] | None:
|
||||
lookup = fixture.get("lookupResult")
|
||||
if isinstance(lookup, dict):
|
||||
return lookup
|
||||
if "evidenceBlocks" in fixture or "retrievalTrace" in fixture:
|
||||
return fixture
|
||||
return None
|
||||
|
||||
|
||||
def normalize_fixture(fixture: dict[str, Any], top_k: int) -> NormalizedFixture:
|
||||
lookup_result = get_lookup_result(fixture)
|
||||
if lookup_result is None:
|
||||
return NormalizedFixture(
|
||||
data_shape="invalid",
|
||||
candidates=[],
|
||||
selected_attempt=None,
|
||||
fallback_reason=None,
|
||||
evidence_status=None,
|
||||
included_sources=[],
|
||||
omitted_sources=[],
|
||||
rerank_top_source=None,
|
||||
)
|
||||
|
||||
raw_blocks = lookup_result.get("evidenceBlocks") or []
|
||||
candidates = [
|
||||
Candidate.from_json(raw, index + 1)
|
||||
for index, raw in enumerate(raw_blocks[:top_k])
|
||||
if isinstance(raw, dict)
|
||||
]
|
||||
retrieval_trace = lookup_result.get("retrievalTrace") or {}
|
||||
context_pack = lookup_result.get("contextPack") or {}
|
||||
rerank_trace = lookup_result.get("rerankTrace") or {}
|
||||
rerank_items = [
|
||||
item for item in rerank_trace.get("items", [])
|
||||
if isinstance(item, dict)
|
||||
]
|
||||
rerank_items.sort(key=lambda item: int(item.get("finalRank") or 999999))
|
||||
return NormalizedFixture(
|
||||
data_shape="lookupResult",
|
||||
candidates=candidates,
|
||||
selected_attempt=optional_string(retrieval_trace.get("selectedAttempt")),
|
||||
fallback_reason=optional_string(retrieval_trace.get("fallbackReason")),
|
||||
evidence_status=optional_string(retrieval_trace.get("evidenceStatus")),
|
||||
included_sources=normalize_sources(context_pack.get("includedSources") or []),
|
||||
omitted_sources=normalize_sources(context_pack.get("omittedSources") or []),
|
||||
rerank_top_source=(
|
||||
optional_string(rerank_items[0].get("source"))
|
||||
if rerank_items
|
||||
else None
|
||||
),
|
||||
)
|
||||
|
||||
|
||||
def optional_string(value: Any) -> str | None:
|
||||
if value is None:
|
||||
return None
|
||||
text = str(value)
|
||||
return text if text else None
|
||||
|
||||
|
||||
def evaluate_case(case: dict[str, Any], fixture_dir: Path, top_k: int) -> dict[str, Any]:
|
||||
case_id = str(case["caseId"])
|
||||
fixture_path = fixture_dir / f"{case_id}.json"
|
||||
expected_doc_ids = normalize_terms(case.get("expectedDocIds", []))
|
||||
expected_sources = normalize_terms(case.get("expectedSources", []))
|
||||
expected_documents = expected_doc_ids or expected_sources
|
||||
expected_breadcrumbs = normalize_terms(case.get("expectedBreadcrumbs", []))
|
||||
expected_keywords = normalize_terms(case.get("expectedKeywords", []))
|
||||
|
||||
@@ -100,25 +192,30 @@ def evaluate_case(case: dict[str, Any], fixture_dir: Path, top_k: int) -> dict[s
|
||||
"caseId": case_id,
|
||||
"scenario": case.get("scenario"),
|
||||
"query": case.get("query"),
|
||||
"dataShape": None,
|
||||
"hitLevel": "miss",
|
||||
"passed": False,
|
||||
"firstExpectedRank": None,
|
||||
"topCandidates": [],
|
||||
"matchedKeywords": [],
|
||||
"breadcrumbMatched": False,
|
||||
"selectedAttempt": None,
|
||||
"fallbackReason": None,
|
||||
"evidenceStatus": None,
|
||||
"includedSources": [],
|
||||
"omittedSources": [],
|
||||
"rerankTopSource": None,
|
||||
"failedChecks": [f"missing fixture: {fixture_path.as_posix()}"],
|
||||
}
|
||||
|
||||
fixture = load_json(fixture_path)
|
||||
raw_candidates = fixture.get("candidates", [])
|
||||
candidates = [
|
||||
Candidate.from_json(raw, index + 1)
|
||||
for index, raw in enumerate(raw_candidates[:top_k])
|
||||
]
|
||||
fixture = normalize_fixture(load_json(fixture_path), top_k)
|
||||
candidates = fixture.candidates
|
||||
|
||||
first_expected = None
|
||||
expected_doc_candidate = None
|
||||
for candidate in candidates:
|
||||
candidate_doc = candidate.doc_id.lower()
|
||||
if any(expected == candidate_doc for expected in expected_doc_ids):
|
||||
if any(expected == candidate_doc for expected in expected_documents):
|
||||
first_expected = candidate.rank
|
||||
expected_doc_candidate = candidate
|
||||
break
|
||||
@@ -149,6 +246,8 @@ def evaluate_case(case: dict[str, Any], fixture_dir: Path, top_k: int) -> dict[s
|
||||
if expected_keywords and not keyword_matches:
|
||||
failed_checks.append("expected evidence keywords not found")
|
||||
|
||||
failed_checks.extend(check_modular_contract(case, fixture))
|
||||
|
||||
if expected_doc_candidate is not None and (
|
||||
breadcrumb_match or bool(keyword_matches)
|
||||
):
|
||||
@@ -164,16 +263,119 @@ def evaluate_case(case: dict[str, Any], fixture_dir: Path, top_k: int) -> dict[s
|
||||
"caseId": case_id,
|
||||
"scenario": case.get("scenario"),
|
||||
"query": case.get("query"),
|
||||
"dataShape": fixture.data_shape,
|
||||
"hitLevel": hit_level,
|
||||
"passed": hit_level in {"strong", "medium"},
|
||||
"passed": hit_level in {"strong", "medium"} and not failed_checks,
|
||||
"firstExpectedRank": first_expected,
|
||||
"topCandidates": [candidate.label() for candidate in candidates],
|
||||
"matchedKeywords": keyword_matches,
|
||||
"breadcrumbMatched": breadcrumb_match,
|
||||
"selectedAttempt": fixture.selected_attempt,
|
||||
"fallbackReason": fixture.fallback_reason,
|
||||
"evidenceStatus": fixture.evidence_status,
|
||||
"includedSources": fixture.included_sources,
|
||||
"omittedSources": fixture.omitted_sources,
|
||||
"rerankTopSource": fixture.rerank_top_source,
|
||||
"failedChecks": failed_checks,
|
||||
}
|
||||
|
||||
|
||||
def check_modular_contract(case: dict[str, Any], fixture: NormalizedFixture) -> list[str]:
|
||||
failed: list[str] = []
|
||||
if fixture.data_shape != "lookupResult":
|
||||
failed.append("fixture must use lookupResult shape")
|
||||
compare_expected(
|
||||
failed,
|
||||
case,
|
||||
"expectedSelectedAttempt",
|
||||
fixture.selected_attempt,
|
||||
"selected attempt mismatch",
|
||||
)
|
||||
if "expectedFallbackReasons" in case:
|
||||
compare_expected_any(
|
||||
failed,
|
||||
case,
|
||||
"expectedFallbackReasons",
|
||||
fixture.fallback_reason,
|
||||
"fallback reason mismatch",
|
||||
)
|
||||
else:
|
||||
compare_expected(
|
||||
failed,
|
||||
case,
|
||||
"expectedFallbackReason",
|
||||
fixture.fallback_reason,
|
||||
"fallback reason mismatch",
|
||||
)
|
||||
compare_expected(
|
||||
failed,
|
||||
case,
|
||||
"expectedEvidenceStatus",
|
||||
fixture.evidence_status,
|
||||
"evidence status mismatch",
|
||||
)
|
||||
compare_expected(
|
||||
failed,
|
||||
case,
|
||||
"expectedRerankTopSource",
|
||||
fixture.rerank_top_source,
|
||||
"rerank top source mismatch",
|
||||
)
|
||||
|
||||
expected_context_sources = normalize_sources(case.get("expectedContextSources", []))
|
||||
if expected_context_sources:
|
||||
included = set(fixture.included_sources)
|
||||
missing = [
|
||||
source for source in expected_context_sources
|
||||
if source not in included
|
||||
]
|
||||
if missing:
|
||||
failed.append("expected context sources missing: " + ", ".join(missing))
|
||||
return failed
|
||||
|
||||
|
||||
def compare_expected(
|
||||
failed: list[str],
|
||||
case: dict[str, Any],
|
||||
field: str,
|
||||
actual: str | None,
|
||||
message: str,
|
||||
) -> None:
|
||||
if field not in case:
|
||||
return
|
||||
expected = case.get(field)
|
||||
if expected is None:
|
||||
if actual is not None:
|
||||
failed.append(f"{message}: expected <none>, got {actual}")
|
||||
return
|
||||
if str(expected) != str(actual):
|
||||
failed.append(f"{message}: expected {expected}, got {actual or '<none>'}")
|
||||
|
||||
|
||||
def compare_expected_any(
|
||||
failed: list[str],
|
||||
case: dict[str, Any],
|
||||
field: str,
|
||||
actual: str | None,
|
||||
message: str,
|
||||
) -> None:
|
||||
expected_values = case.get(field)
|
||||
if not isinstance(expected_values, list):
|
||||
failed.append(f"{field} must be a list")
|
||||
return
|
||||
normalized_expected = [
|
||||
None if value is None else str(value)
|
||||
for value in expected_values
|
||||
]
|
||||
normalized_actual = None if actual is None else str(actual)
|
||||
if normalized_actual not in normalized_expected:
|
||||
expected_text = ", ".join(
|
||||
"<none>" if value is None else value
|
||||
for value in normalized_expected
|
||||
)
|
||||
failed.append(f"{message}: expected one of [{expected_text}], got {actual or '<none>'}")
|
||||
|
||||
|
||||
def aggregate(results: list[dict[str, Any]], top_k: int) -> dict[str, Any]:
|
||||
total = len(results)
|
||||
counts = {
|
||||
@@ -187,15 +389,23 @@ def aggregate(results: list[dict[str, Any]], top_k: int) -> dict[str, Any]:
|
||||
for item in results
|
||||
if item.get("firstExpectedRank") is not None
|
||||
]
|
||||
passed = counts["strong"] + counts["medium"]
|
||||
retrieved = counts["strong"] + counts["medium"]
|
||||
passed = sum(1 for item in results if item["passed"])
|
||||
lookup_result_cases = sum(
|
||||
1 for item in results if item.get("dataShape") == "lookupResult"
|
||||
)
|
||||
return {
|
||||
"caseCount": total,
|
||||
"topK": top_k,
|
||||
"passedCount": passed,
|
||||
"failedCount": total - passed,
|
||||
"passRate": round(passed / total, 4) if total else 0,
|
||||
"lookupResultCaseCount": lookup_result_cases,
|
||||
"strongHitCount": counts["strong"],
|
||||
"mediumHitCount": counts["medium"],
|
||||
"weakHitCount": counts["weak"],
|
||||
"missCount": counts["miss"],
|
||||
"recallAtK": round(passed / total, 4) if total else 0,
|
||||
"recallAtK": round(retrieved / total, 4) if total else 0,
|
||||
"strongHitRate": round(counts["strong"] / total, 4) if total else 0,
|
||||
"averageFirstHitRank": (
|
||||
round(sum(expected_ranks) / len(expected_ranks), 4)
|
||||
@@ -218,6 +428,10 @@ def render_markdown(report: dict[str, Any]) -> str:
|
||||
"|---|---:|",
|
||||
f"| Cases | {metrics['caseCount']} |",
|
||||
f"| Top K | {metrics['topK']} |",
|
||||
f"| Passed | {metrics['passedCount']} |",
|
||||
f"| Failed | {metrics['failedCount']} |",
|
||||
f"| Pass rate | {metrics['passRate']} |",
|
||||
f"| LookupResult fixtures | {metrics['lookupResultCaseCount']} |",
|
||||
f"| Recall@K | {metrics['recallAtK']} |",
|
||||
f"| Strong hit rate | {metrics['strongHitRate']} |",
|
||||
f"| Strong hits | {metrics['strongHitCount']} |",
|
||||
@@ -228,18 +442,22 @@ def render_markdown(report: dict[str, Any]) -> str:
|
||||
"",
|
||||
"## Cases",
|
||||
"",
|
||||
"| Case | Scenario | Hit | First Expected Rank | Top Candidates | Failed Checks |",
|
||||
"|---|---|---|---:|---|---|",
|
||||
"| Case | Scenario | Pass | Hit | Attempt | Fallback | Evidence | First Expected Rank | Top Candidates | Failed Checks |",
|
||||
"|---|---|---|---|---|---|---|---:|---|---|",
|
||||
]
|
||||
for item in report["results"]:
|
||||
failed = "<br>".join(item["failedChecks"]) if item["failedChecks"] else ""
|
||||
top = "<br>".join(item["topCandidates"])
|
||||
first_rank = item["firstExpectedRank"]
|
||||
lines.append(
|
||||
"| {case} | {scenario} | {hit} | {rank} | {top} | {failed} |".format(
|
||||
"| {case} | {scenario} | {passed} | {hit} | {attempt} | {fallback} | {evidence} | {rank} | {top} | {failed} |".format(
|
||||
case=item["caseId"],
|
||||
scenario=item.get("scenario") or "",
|
||||
passed=str(item["passed"]).lower(),
|
||||
hit=item["hitLevel"],
|
||||
attempt=item.get("selectedAttempt") or "",
|
||||
fallback=item.get("fallbackReason") or "",
|
||||
evidence=item.get("evidenceStatus") or "",
|
||||
rank=first_rank if first_rank is not None else "",
|
||||
top=top,
|
||||
failed=failed,
|
||||
@@ -249,12 +467,255 @@ def render_markdown(report: dict[str, Any]) -> str:
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def compare_reports(baseline: dict[str, Any], current: dict[str, Any]) -> dict[str, Any]:
|
||||
items: list[dict[str, Any]] = []
|
||||
compare_metric(items, "aggregate", None, "passRate", baseline, current, higher_is_better=True)
|
||||
compare_metric(items, "aggregate", None, "recallAtK", baseline, current, higher_is_better=True)
|
||||
compare_metric(items, "aggregate", None, "strongHitRate", baseline, current, higher_is_better=True)
|
||||
compare_metric(items, "aggregate", None, "missCount", baseline, current, higher_is_better=False)
|
||||
compare_cases(items, baseline.get("results") or [], current.get("results") or [])
|
||||
|
||||
regression_count = count_type(items, "REGRESSION")
|
||||
improvement_count = count_type(items, "IMPROVEMENT")
|
||||
changed_count = count_type(items, "CHANGED")
|
||||
return {
|
||||
"generatedAt": datetime.now(timezone.utc).isoformat(),
|
||||
"baselineReport": baseline.get("caseFile"),
|
||||
"currentReport": current.get("caseFile"),
|
||||
"baselineCaseCount": get_aggregate_value(baseline, "caseCount"),
|
||||
"currentCaseCount": get_aggregate_value(current, "caseCount"),
|
||||
"baselinePassRate": get_aggregate_value(baseline, "passRate"),
|
||||
"currentPassRate": get_aggregate_value(current, "passRate"),
|
||||
"baselineRecallAtK": get_aggregate_value(baseline, "recallAtK"),
|
||||
"currentRecallAtK": get_aggregate_value(current, "recallAtK"),
|
||||
"regressionCount": regression_count,
|
||||
"improvementCount": improvement_count,
|
||||
"changedCount": changed_count,
|
||||
"hasRegression": regression_count > 0,
|
||||
"items": items,
|
||||
}
|
||||
|
||||
|
||||
def compare_metric(
|
||||
items: list[dict[str, Any]],
|
||||
scope: str,
|
||||
case_id: str | None,
|
||||
metric: str,
|
||||
baseline: dict[str, Any],
|
||||
current: dict[str, Any],
|
||||
higher_is_better: bool,
|
||||
) -> None:
|
||||
baseline_value = get_aggregate_value(baseline, metric)
|
||||
current_value = get_aggregate_value(current, metric)
|
||||
if baseline_value == current_value:
|
||||
return
|
||||
if not isinstance(baseline_value, (int, float)) or not isinstance(current_value, (int, float)):
|
||||
items.append(diff_item("CHANGED", scope, case_id, metric, baseline_value, current_value, None))
|
||||
return
|
||||
delta = float(current_value) - float(baseline_value)
|
||||
items.append(diff_item(classify_delta(delta, higher_is_better), scope, case_id, metric,
|
||||
baseline_value, current_value, delta))
|
||||
|
||||
|
||||
def compare_cases(
|
||||
items: list[dict[str, Any]],
|
||||
baseline_results: list[dict[str, Any]],
|
||||
current_results: list[dict[str, Any]],
|
||||
) -> None:
|
||||
baseline_by_id = by_case_id(baseline_results)
|
||||
current_by_id = by_case_id(current_results)
|
||||
case_ids = sorted(set(baseline_by_id) | set(current_by_id))
|
||||
for case_id in case_ids:
|
||||
baseline = baseline_by_id.get(case_id)
|
||||
current = current_by_id.get(case_id)
|
||||
if baseline is None:
|
||||
items.append(diff_item("CHANGED", "case", case_id, "casePresence", "missing", "present", None))
|
||||
continue
|
||||
if current is None:
|
||||
items.append(diff_item("REGRESSION", "case", case_id, "casePresence", "present", "missing", None))
|
||||
continue
|
||||
compare_case_bool(items, case_id, "passed", baseline, current, higher_is_better=True)
|
||||
compare_hit_level(items, case_id, baseline, current)
|
||||
compare_case_rank(items, case_id, baseline, current)
|
||||
compare_case_value(items, case_id, "selectedAttempt", baseline, current)
|
||||
compare_case_value(items, case_id, "fallbackReason", baseline, current)
|
||||
compare_case_value(items, case_id, "evidenceStatus", baseline, current)
|
||||
compare_case_value(items, case_id, "rerankTopSource", baseline, current)
|
||||
|
||||
|
||||
def compare_case_bool(
|
||||
items: list[dict[str, Any]],
|
||||
case_id: str,
|
||||
metric: str,
|
||||
baseline: dict[str, Any],
|
||||
current: dict[str, Any],
|
||||
higher_is_better: bool,
|
||||
) -> None:
|
||||
baseline_value = bool(baseline.get(metric))
|
||||
current_value = bool(current.get(metric))
|
||||
if baseline_value == current_value:
|
||||
return
|
||||
delta = int(current_value) - int(baseline_value)
|
||||
items.append(diff_item(classify_delta(delta, higher_is_better), "case", case_id, metric,
|
||||
baseline_value, current_value, float(delta)))
|
||||
|
||||
|
||||
def compare_hit_level(
|
||||
items: list[dict[str, Any]],
|
||||
case_id: str,
|
||||
baseline: dict[str, Any],
|
||||
current: dict[str, Any],
|
||||
) -> None:
|
||||
baseline_value = baseline.get("hitLevel")
|
||||
current_value = current.get("hitLevel")
|
||||
if baseline_value == current_value:
|
||||
return
|
||||
delta = HIT_LEVEL_RANK.get(str(current_value), 0) - HIT_LEVEL_RANK.get(str(baseline_value), 0)
|
||||
items.append(diff_item(classify_delta(delta, True), "case", case_id, "hitLevel",
|
||||
baseline_value, current_value, float(delta)))
|
||||
|
||||
|
||||
def compare_case_rank(
|
||||
items: list[dict[str, Any]],
|
||||
case_id: str,
|
||||
baseline: dict[str, Any],
|
||||
current: dict[str, Any],
|
||||
) -> None:
|
||||
baseline_value = baseline.get("firstExpectedRank")
|
||||
current_value = current.get("firstExpectedRank")
|
||||
if baseline_value == current_value:
|
||||
return
|
||||
if baseline_value is None or current_value is None:
|
||||
change_type = "REGRESSION" if current_value is None else "IMPROVEMENT"
|
||||
items.append(diff_item(change_type, "case", case_id, "firstExpectedRank",
|
||||
baseline_value, current_value, None))
|
||||
return
|
||||
delta = int(current_value) - int(baseline_value)
|
||||
items.append(diff_item(classify_delta(delta, False), "case", case_id, "firstExpectedRank",
|
||||
baseline_value, current_value, float(delta)))
|
||||
|
||||
|
||||
def compare_case_value(
|
||||
items: list[dict[str, Any]],
|
||||
case_id: str,
|
||||
metric: str,
|
||||
baseline: dict[str, Any],
|
||||
current: dict[str, Any],
|
||||
) -> None:
|
||||
baseline_value = baseline.get(metric)
|
||||
current_value = current.get(metric)
|
||||
if baseline_value == current_value:
|
||||
return
|
||||
items.append(diff_item("CHANGED", "case", case_id, metric, baseline_value, current_value, None))
|
||||
|
||||
|
||||
def diff_item(
|
||||
change_type: str,
|
||||
scope: str,
|
||||
case_id: str | None,
|
||||
metric: str,
|
||||
baseline_value: Any,
|
||||
current_value: Any,
|
||||
delta: float | None,
|
||||
) -> dict[str, Any]:
|
||||
target = case_id or scope
|
||||
return {
|
||||
"type": change_type,
|
||||
"scope": scope,
|
||||
"caseId": case_id,
|
||||
"metric": metric,
|
||||
"baselineValue": value_label(baseline_value),
|
||||
"currentValue": value_label(current_value),
|
||||
"delta": delta,
|
||||
"message": f"{target} {metric} changed",
|
||||
}
|
||||
|
||||
|
||||
def render_diff_markdown(report: dict[str, Any]) -> str:
|
||||
lines = [
|
||||
"# RAG Retrieval Baseline Diff",
|
||||
"",
|
||||
f"Generated at: `{report['generatedAt']}`",
|
||||
"",
|
||||
"## Summary",
|
||||
"",
|
||||
"| Metric | Value |",
|
||||
"|---|---:|",
|
||||
f"| Baseline cases | {report['baselineCaseCount']} |",
|
||||
f"| Current cases | {report['currentCaseCount']} |",
|
||||
f"| Baseline pass rate | {report['baselinePassRate']} |",
|
||||
f"| Current pass rate | {report['currentPassRate']} |",
|
||||
f"| Baseline recall@K | {report['baselineRecallAtK']} |",
|
||||
f"| Current recall@K | {report['currentRecallAtK']} |",
|
||||
f"| Regressions | {report['regressionCount']} |",
|
||||
f"| Improvements | {report['improvementCount']} |",
|
||||
f"| Changed | {report['changedCount']} |",
|
||||
"",
|
||||
"## Items",
|
||||
"",
|
||||
"| Type | Scope | Case | Metric | Baseline | Current | Delta | Message |",
|
||||
"|---|---|---|---|---|---|---:|---|",
|
||||
]
|
||||
for item in report["items"]:
|
||||
delta = "" if item.get("delta") is None else item["delta"]
|
||||
lines.append(
|
||||
"| {type} | {scope} | {case} | {metric} | {baseline} | {current} | {delta} | {message} |".format(
|
||||
type=item["type"],
|
||||
scope=item["scope"],
|
||||
case=item.get("caseId") or "",
|
||||
metric=item["metric"],
|
||||
baseline=item["baselineValue"],
|
||||
current=item["currentValue"],
|
||||
delta=delta,
|
||||
message=item["message"],
|
||||
)
|
||||
)
|
||||
lines.append("")
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def get_aggregate_value(report: dict[str, Any], metric: str) -> Any:
|
||||
return (report.get("aggregate") or {}).get(metric)
|
||||
|
||||
|
||||
def by_case_id(results: list[dict[str, Any]]) -> dict[str, dict[str, Any]]:
|
||||
return {
|
||||
str(item.get("caseId")): item
|
||||
for item in sorted(results, key=lambda item: str(item.get("caseId")))
|
||||
}
|
||||
|
||||
|
||||
def classify_delta(delta: float, higher_is_better: bool) -> str:
|
||||
if delta == 0.0:
|
||||
return "CHANGED"
|
||||
improved = delta > 0 if higher_is_better else delta < 0
|
||||
return "IMPROVEMENT" if improved else "REGRESSION"
|
||||
|
||||
|
||||
def count_type(items: list[dict[str, Any]], change_type: str) -> int:
|
||||
return sum(1 for item in items if item["type"] == change_type)
|
||||
|
||||
|
||||
def value_label(value: Any) -> str:
|
||||
if value is None:
|
||||
return "-"
|
||||
return str(value)
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("--cases", type=Path, default=DEFAULT_CASES)
|
||||
parser.add_argument("--fixtures", type=Path, default=DEFAULT_FIXTURES)
|
||||
parser.add_argument("--json-report", type=Path, default=DEFAULT_JSON_REPORT)
|
||||
parser.add_argument("--markdown-report", type=Path, default=DEFAULT_MD_REPORT)
|
||||
parser.add_argument(
|
||||
"--compare-to",
|
||||
type=Path,
|
||||
default=None,
|
||||
help="Optional baseline report to diff against the newly generated report.",
|
||||
)
|
||||
parser.add_argument("--diff-json-report", type=Path, default=DEFAULT_DIFF_JSON_REPORT)
|
||||
parser.add_argument("--diff-markdown-report", type=Path, default=DEFAULT_DIFF_MD_REPORT)
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
@@ -274,16 +735,32 @@ def main() -> int:
|
||||
write_json(args.json_report, report)
|
||||
write_text(args.markdown_report, render_markdown(report))
|
||||
|
||||
failed = [item for item in results if item["hitLevel"] == "miss"]
|
||||
failed = [item for item in results if not item["passed"]]
|
||||
diff_failed = False
|
||||
if args.compare_to is not None:
|
||||
baseline = load_json(args.compare_to)
|
||||
diff = compare_reports(baseline, report)
|
||||
write_json(args.diff_json_report, diff)
|
||||
write_text(args.diff_markdown_report, render_diff_markdown(diff))
|
||||
diff_failed = bool(diff["hasRegression"])
|
||||
print(
|
||||
"Diffed against {baseline}: regressions={regressions}, changes={changes}".format(
|
||||
baseline=args.compare_to.as_posix(),
|
||||
regressions=diff["regressionCount"],
|
||||
changes=diff["changedCount"],
|
||||
)
|
||||
)
|
||||
|
||||
print(
|
||||
"Evaluated {total} cases: recall@{top_k}={recall}, misses={misses}".format(
|
||||
"Evaluated {total} cases: passRate={pass_rate}, recall@{top_k}={recall}, failed={failed}".format(
|
||||
total=len(results),
|
||||
pass_rate=report["aggregate"]["passRate"],
|
||||
top_k=top_k,
|
||||
recall=report["aggregate"]["recallAtK"],
|
||||
misses=len(failed),
|
||||
failed=len(failed),
|
||||
)
|
||||
)
|
||||
return 1 if failed else 0
|
||||
return 1 if failed or diff_failed else 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
|
||||
@@ -0,0 +1,34 @@
|
||||
param(
|
||||
[string]$Cases = "eval\rag-retrieval\cases\golden-cases.json",
|
||||
[string]$Fixtures = "eval\rag-retrieval\fixtures",
|
||||
[string]$RetrievedAt = "",
|
||||
[string]$KbScope = "rag-eval",
|
||||
[string]$VectorStoreMode = "spring",
|
||||
[switch]$SkipEval
|
||||
)
|
||||
|
||||
$ErrorActionPreference = "Stop"
|
||||
|
||||
$mavenArgs = @(
|
||||
"-q",
|
||||
"-Dtest=RagLookupSnapshotGeneratorTest",
|
||||
"-Drag.snapshot.enabled=true",
|
||||
"-Drag.snapshot.cases=$Cases",
|
||||
"-Drag.snapshot.fixtures=$Fixtures",
|
||||
"-Dretrieval.kb-scope=$KbScope",
|
||||
"-Dretrieval.vector-store.mode=$VectorStoreMode"
|
||||
)
|
||||
|
||||
if ($RetrievedAt -ne "") {
|
||||
$mavenArgs += "-Drag.snapshot.retrievedAt=$RetrievedAt"
|
||||
}
|
||||
|
||||
$mavenArgs += "test"
|
||||
|
||||
Write-Host "Generating RAG lookupResult fixtures from real LookupKnowledgeTool..."
|
||||
& mvn @mavenArgs
|
||||
|
||||
if (-not $SkipEval) {
|
||||
Write-Host "Running offline RAG retrieval baseline..."
|
||||
& python scripts\eval_rag_retrieval.py --cases $Cases --fixtures $Fixtures
|
||||
}
|
||||
@@ -0,0 +1,18 @@
|
||||
param(
|
||||
[string]$SeedDocs = "eval\rag-retrieval\seed-docs",
|
||||
[string]$KbScope = "rag-eval"
|
||||
)
|
||||
|
||||
$ErrorActionPreference = "Stop"
|
||||
|
||||
$mavenArgs = @(
|
||||
"-q",
|
||||
"-Dtest=RagEvalSeedImporterTest",
|
||||
"-Drag.seed.enabled=true",
|
||||
"-Drag.seed.docs=$SeedDocs",
|
||||
"-Dretrieval.kb-scope=$KbScope",
|
||||
"test"
|
||||
)
|
||||
|
||||
Write-Host "Importing RAG eval seed docs through DocumentManagementService..."
|
||||
& mvn @mavenArgs
|
||||
@@ -1,5 +1,7 @@
|
||||
package com.superbiz.agent.domain.model;
|
||||
|
||||
import com.fasterxml.jackson.annotation.JsonIgnore;
|
||||
import com.fasterxml.jackson.annotation.JsonIgnoreProperties;
|
||||
import lombok.AllArgsConstructor;
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
@@ -20,6 +22,7 @@ import java.util.Map;
|
||||
@Builder
|
||||
@NoArgsConstructor
|
||||
@AllArgsConstructor
|
||||
@JsonIgnoreProperties(ignoreUnknown = true)
|
||||
public class SessionContext implements Serializable {
|
||||
|
||||
private static final long serialVersionUID = 1L;
|
||||
@@ -118,6 +121,7 @@ public class SessionContext implements Serializable {
|
||||
/**
|
||||
* 获取聊天历史副本,避免调用方直接修改内部列表。
|
||||
*/
|
||||
@JsonIgnore
|
||||
public List<Map<String, String>> getMessageHistorySnapshot() {
|
||||
if (this.messageHistory == null || this.messageHistory.isEmpty()) {
|
||||
return new ArrayList<>();
|
||||
@@ -141,6 +145,7 @@ public class SessionContext implements Serializable {
|
||||
this.lastActiveAt = LocalDateTime.now();
|
||||
}
|
||||
|
||||
@JsonIgnore
|
||||
public int getMessagePairCount() {
|
||||
return this.messageHistory == null ? 0 : this.messageHistory.size() / 2;
|
||||
}
|
||||
|
||||
@@ -0,0 +1,26 @@
|
||||
package com.superbiz.agent.dto;
|
||||
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
/**
|
||||
* Agent-facing packed context assembled from final evidence blocks.
|
||||
*/
|
||||
@Data
|
||||
@Builder
|
||||
public class ContextPack {
|
||||
|
||||
private String packedText;
|
||||
|
||||
private String strategy;
|
||||
|
||||
private Integer charBudget;
|
||||
|
||||
private Integer usedChars;
|
||||
|
||||
private List<String> includedSources;
|
||||
|
||||
private List<String> omittedSources;
|
||||
}
|
||||
@@ -0,0 +1,32 @@
|
||||
package com.superbiz.agent.dto;
|
||||
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
/**
|
||||
* Output of post-retrieval processing before context packing.
|
||||
*/
|
||||
@Data
|
||||
@Builder
|
||||
public class EvidencePostprocessResult {
|
||||
|
||||
private Integer candidateCount;
|
||||
|
||||
private Integer evidenceBlockCount;
|
||||
|
||||
private List<EvidenceBlock> evidenceBlocks;
|
||||
|
||||
private String relevanceLevel;
|
||||
|
||||
private String completenessHint;
|
||||
|
||||
private RerankTrace rerankTrace;
|
||||
|
||||
private Double topSimilarity;
|
||||
|
||||
public boolean hasUsableEvidence() {
|
||||
return evidenceBlocks != null && !evidenceBlocks.isEmpty();
|
||||
}
|
||||
}
|
||||
@@ -1,5 +1,6 @@
|
||||
package com.superbiz.agent.dto;
|
||||
|
||||
import com.fasterxml.jackson.annotation.JsonProperty;
|
||||
import lombok.AllArgsConstructor;
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
@@ -39,6 +40,13 @@ public class Frontmatter {
|
||||
*/
|
||||
private String category;
|
||||
|
||||
private String source;
|
||||
|
||||
private String breadcrumb;
|
||||
|
||||
@JsonProperty("kb_scope")
|
||||
private String kbScope;
|
||||
|
||||
/**
|
||||
* 章节锚点(预留字段,MVP 不使用)
|
||||
* Key: 章节标题,Value: 章节 Markdown 标题
|
||||
|
||||
@@ -39,6 +39,8 @@ public class KnowledgeEntry {
|
||||
*/
|
||||
private String category;
|
||||
|
||||
private String kbScope;
|
||||
|
||||
/**
|
||||
* 章节锚点(预留字段,MVP 不使用)
|
||||
*/
|
||||
|
||||
@@ -0,0 +1,30 @@
|
||||
package com.superbiz.agent.dto;
|
||||
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
/**
|
||||
* Query understanding output used by the knowledge retrieval pipeline.
|
||||
*/
|
||||
@Data
|
||||
@Builder
|
||||
public class KnowledgeQuery {
|
||||
|
||||
private String originalQuery;
|
||||
|
||||
private String rewrittenQuery;
|
||||
|
||||
private List<String> domainHints;
|
||||
|
||||
private List<String> matchedKeywords;
|
||||
|
||||
private List<String> entities;
|
||||
|
||||
private String categoryFilter;
|
||||
|
||||
private List<String> l0Titles;
|
||||
|
||||
private Integer l0MatchCount;
|
||||
}
|
||||
@@ -17,21 +17,26 @@ public class LookupResult {
|
||||
*/
|
||||
private boolean found;
|
||||
|
||||
/**
|
||||
* 主要结果(L0 精确匹配)
|
||||
*/
|
||||
private PrimaryResult primary;
|
||||
|
||||
/**
|
||||
* 补充结果(L1 语义检索)
|
||||
*/
|
||||
private SupplementResult supplement;
|
||||
|
||||
/**
|
||||
* Structured evidence blocks after retrieval post-processing.
|
||||
*/
|
||||
private List<EvidenceBlock> evidenceBlocks;
|
||||
|
||||
/**
|
||||
* Packed Agent-facing context assembled from evidence blocks.
|
||||
*/
|
||||
private ContextPack contextPack;
|
||||
|
||||
/**
|
||||
* Retrieval attempts and fallback trace.
|
||||
*/
|
||||
private RetrievalTrace retrievalTrace;
|
||||
|
||||
/**
|
||||
* Rule-based rerank explanation.
|
||||
*/
|
||||
private RerankTrace rerankTrace;
|
||||
|
||||
/**
|
||||
* Candidate count before evidence deduplication.
|
||||
*/
|
||||
|
||||
@@ -1,39 +0,0 @@
|
||||
package com.superbiz.agent.dto;
|
||||
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
/**
|
||||
* L0 精确匹配结果
|
||||
*/
|
||||
@Data
|
||||
@Builder
|
||||
public class PrimaryResult {
|
||||
|
||||
/**
|
||||
* 文档内容(前 2000 字符)
|
||||
*/
|
||||
private String content;
|
||||
|
||||
/**
|
||||
* 文档来源路径
|
||||
*/
|
||||
private String source;
|
||||
|
||||
/**
|
||||
* 匹配类型(exact_L0)
|
||||
*/
|
||||
private String matchType;
|
||||
|
||||
/**
|
||||
* 置信度(high / low)
|
||||
*/
|
||||
private String confidence;
|
||||
|
||||
/**
|
||||
* 可用的章节列表(预留字段,MVP 返回 null)
|
||||
*/
|
||||
private List<String> availableSections;
|
||||
}
|
||||
@@ -0,0 +1,26 @@
|
||||
package com.superbiz.agent.dto;
|
||||
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
/**
|
||||
* Rule-based rerank explanation for final evidence blocks.
|
||||
*/
|
||||
@Data
|
||||
@Builder
|
||||
public class RerankTrace {
|
||||
|
||||
private List<Item> items;
|
||||
|
||||
@Data
|
||||
@Builder
|
||||
public static class Item {
|
||||
private Integer finalRank;
|
||||
private String source;
|
||||
private Double baseScore;
|
||||
private Double finalScore;
|
||||
private List<String> boostReasons;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,45 @@
|
||||
package com.superbiz.agent.dto;
|
||||
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
/**
|
||||
* Trace of retrieval attempts used by lookup_knowledge.
|
||||
*/
|
||||
@Data
|
||||
@Builder
|
||||
public class RetrievalTrace {
|
||||
|
||||
private String originalQuery;
|
||||
|
||||
private String rewrittenQuery;
|
||||
|
||||
private String categoryFilter;
|
||||
|
||||
private String selectedAttempt;
|
||||
|
||||
private String fallbackReason;
|
||||
|
||||
private String evidenceStatus;
|
||||
|
||||
private Map<String, Object> queryHints;
|
||||
|
||||
private List<Attempt> attempts;
|
||||
|
||||
@Data
|
||||
@Builder
|
||||
public static class Attempt {
|
||||
private String name;
|
||||
private String query;
|
||||
private String categoryFilter;
|
||||
private Integer candidateCount;
|
||||
private Boolean usable;
|
||||
private String errorMessage;
|
||||
private Integer durationMs;
|
||||
private Double topScore;
|
||||
private Double topSimilarity;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,41 @@
|
||||
package com.superbiz.agent.dto;
|
||||
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
/**
|
||||
* Normalized vector retrieval candidate before evidence post-processing.
|
||||
*/
|
||||
@Data
|
||||
@Builder
|
||||
public class RetrievedEvidenceCandidate {
|
||||
|
||||
private String id;
|
||||
|
||||
private String source;
|
||||
|
||||
private String title;
|
||||
|
||||
private String breadcrumb;
|
||||
|
||||
private String content;
|
||||
|
||||
private String retrievalLayer;
|
||||
|
||||
private String retrievalAttempt;
|
||||
|
||||
private Double score;
|
||||
|
||||
private Double rawScore;
|
||||
|
||||
private String scoreLabel;
|
||||
|
||||
private Integer originalRank;
|
||||
|
||||
private Map<String, String> metadata;
|
||||
|
||||
private List<String> hitReasons;
|
||||
}
|
||||
@@ -1,27 +0,0 @@
|
||||
package com.superbiz.agent.dto;
|
||||
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
|
||||
/**
|
||||
* L1 语义检索补充结果
|
||||
*/
|
||||
@Data
|
||||
@Builder
|
||||
public class SupplementResult {
|
||||
|
||||
/**
|
||||
* 文档内容片段
|
||||
*/
|
||||
private String content;
|
||||
|
||||
/**
|
||||
* 文档来源
|
||||
*/
|
||||
private String source;
|
||||
|
||||
/**
|
||||
* 匹配类型(semantic_L1)
|
||||
*/
|
||||
private String matchType;
|
||||
}
|
||||
@@ -36,8 +36,8 @@ import java.util.Optional;
|
||||
import java.util.UUID;
|
||||
|
||||
/**
|
||||
* AI Ops 闂備礁鎼幊妯肩磽濮樿泛绀傛俊顖滅帛娴溿倝鏌熼柇锕€鏋熸俊顖氾躬閺岋繝宕煎┑鍩裤垹鈹?
|
||||
* 闂佽崵濮甸崝妤呭窗閺囥垺鍎楁俊銈勭缁?Agent 闂備礁鎲¢〃鍛崲鐎n剛绀婇柡鍐ㄧ墛閸庡秹鏌涢弴銊ヤ簼闁哥喓鍋ら幃褰掑焵椤掑嫭鏅濋柍褜鍓熷畷瑙勬償閵娿儱鍤戦梺褰掑亰閸橀箖濡堕敂鍓х<?
|
||||
* AI Ops 智能运维服务。
|
||||
* 负责构建 Planner、Executor、Supervisor 多 Agent 编排流程,并持久化诊断会话与执行指标。
|
||||
*/
|
||||
@Service
|
||||
public class AiOpsService {
|
||||
@@ -53,7 +53,7 @@ public class AiOpsService {
|
||||
@Autowired
|
||||
private QueryMetricsTools queryMetricsTools;
|
||||
|
||||
@Autowired(required = false) // Mock 婵犵妲呴崹顏堝焵椤掆偓绾绢厾娑甸埀顒佺箾閹寸偞灏い鎴濇閺呭爼鎮╁ù瀣亙闂侀潧顭堥崕閬嶅棘閳?
|
||||
@Autowired(required = false) // Mock 模式下才注册本地日志查询工具。
|
||||
private QueryLogsTools queryLogsTools;
|
||||
|
||||
@Autowired
|
||||
@@ -81,12 +81,12 @@ public class AiOpsService {
|
||||
private SkillRegistry skillRegistry;
|
||||
|
||||
/**
|
||||
* 闂備礁婀遍悷鎶藉幢閳哄倹鏉?AI Ops 闂備礁鎲$粙鎴︽晝閵娾晩鏁嗛柣鏃傚帶缁€鍡涙煕閳╁喚鐒介柍褜鍓濆Λ鍕箒婵炶揪缍€閵嗏偓闁?
|
||||
* 执行 AI Ops 告警分析流程。
|
||||
*
|
||||
* @param chatModel 濠电姰鍨归悥銏ゅ礋閳ь剚绗熼埀顒€鐣烽崷顓涘亾閿濆簼绨介柡澶庢閵?
|
||||
* @param toolCallbacks 闁诲氦顫夐幃鍫曞磿闁秴鐭楅柛褎顨呴悙濠囨煟閹邦剙顣虫繛鍫濈埣閺屸剝鎷呴悷閭︽缂?
|
||||
* @return 闂備礁鎲$敮鎺懳涘▎鎾村€甸柣锝呯灱绾惧ジ鏌熼幆褜鍤熷ù婊庡灦閺岋絽螣閸喚鍘梺?
|
||||
* @throws GraphRunnerException 濠电姷顣介埀顒€鍟块埀顒€缍婇幃?Agent 闂備礁婀遍悷鎶藉幢閳哄倹鏉稿┑鐘灪閸庤偐鍒掗崜褎鍠?
|
||||
* @param chatModel 大模型实例
|
||||
* @param toolCallbacks Spring AI 工具回调
|
||||
* @return 多 Agent 编排后的最终状态
|
||||
* @throws GraphRunnerException Agent 图执行失败时抛出
|
||||
*/
|
||||
public Optional<OverAllState> executeAiOpsAnalysis(ChatModel chatModel, ToolCallback[] toolCallbacks) throws GraphRunnerException {
|
||||
return executeAiOpsAnalysis(chatModel, toolCallbacks, null, resolveSessionId(null));
|
||||
@@ -99,19 +99,16 @@ public class AiOpsService {
|
||||
String resolvedSessionId = isBlank(sessionId) ? resolveSessionId(request) : sessionId.trim();
|
||||
long startTime = System.currentTimeMillis();
|
||||
|
||||
// 闂備礁鎲$敮妤冪矙閹寸姷纾介柟鎹愵嚙缁狅綁鏌″鍐ㄥ缂佺虎鍨堕弻锟犲磼濞戞﹩鈧粓鏌i敂鐣屽⒌鐎殿噮鍓熼、妯衡攽閸垻宕堕梺?
|
||||
DiagnosisSession session = startDiagnosisSession(resolvedSessionId, request);
|
||||
diagnosisSessionRepository.save(session);
|
||||
|
||||
// 闂佽崵濮崇粈浣规櫠娴犲鍋?ThreadLocal 濠电偞鍨堕幐鎼佹晝閿濆洨绠旈柛娑欐綑濡﹢鏌涢妷鈺婃缂佲偓閸戯箰okupKnowledgeTool 闂傚倷绶¢崑鍛┍閾忚宕查柛鎰电厛濞间即鏌曢崼婵堝缂佺媭鍨堕弻?sessionId闂?
|
||||
// 让工具调用、Hook 和知识库检索能够拿到当前诊断会话 ID。
|
||||
SessionContextHolder.setSessionId(resolvedSessionId);
|
||||
|
||||
try {
|
||||
// 闂備礁鎼鍛偓姘煎墰缁?Planner 闂?Executor Agent闂備焦瀵х粙鎴︽偋婵犲洦鍎婇柍鈺佸暟閳?Agent 闂備礁鎲¢懝鍓р偓姘煎灦瀹曢潧顭ㄩ崨顔芥?Hook闂?
|
||||
ReactAgent plannerAgent = buildPlannerAgent(chatModel, toolCallbacks);
|
||||
ReactAgent executorAgent = buildExecutorAgent(chatModel, toolCallbacks);
|
||||
|
||||
// 闂備礁鎼鍛偓姘煎墰缁?Supervisor Agent闂備焦瀵х粙鎴︽偋閸涱垳绠斿鑸靛姇缁€?Hook闂?
|
||||
SupervisorAgent supervisorAgent = SupervisorAgent.builder()
|
||||
.name("ai_ops_supervisor")
|
||||
.description("Coordinates Planner and Executor agents")
|
||||
@@ -122,19 +119,17 @@ public class AiOpsService {
|
||||
|
||||
String taskPrompt = buildTaskPrompt(request);
|
||||
|
||||
logger.info("闂佽崵濮撮鍛村疮娴兼潙鏋?Supervisor Agent 闁诲孩顔栭崰鎺楀磻閹炬枼鏀芥い鏃傗拡閸庢垹绱掓鏍﹂偗妤?..");
|
||||
logger.info("Invoking AI Ops supervisor agent");
|
||||
|
||||
Optional<OverAllState> stateOptional = supervisorAgent.invoke(taskPrompt);
|
||||
|
||||
long duration = System.currentTimeMillis() - startTime;
|
||||
|
||||
// 闂備礁鎼ú銈夋偤閵娾晛钃熷┑鐘插鐎氭艾鈹戦悩鎻掓殲闁绘帟妫勯湁闁稿繘妫挎禍銏ゆ煟?
|
||||
session.setStatus(stateOptional.isPresent() ? "SUCCESS" : "FAILED");
|
||||
session.setTotalDurationMs((int) duration);
|
||||
backfillSessionMetrics(session);
|
||||
diagnosisSessionRepository.save(session);
|
||||
|
||||
// 婵犵數鍎戠紞鈧い鏇嗗嫭鍙忛柣鎰仛鐎氼剟鏌涢幇闈涘箻婵¤尙顭堥湁闁绘瑥鎳愰幃濂告煟?
|
||||
if (stateOptional.isPresent()) {
|
||||
OverAllState state = stateOptional.get();
|
||||
logger.debug("Final State Keys: {}", state.data().keySet());
|
||||
@@ -153,22 +148,21 @@ public class AiOpsService {
|
||||
}
|
||||
|
||||
/**
|
||||
* 濠电偛顕慨瀵糕偓娑掓櫆閺呭爼鎮剧仦鎯т粧閻庡厜鍋撻柍褜鍓涢崚鎺楀Ω閳轰礁鍤戝┑鐘才堥崑鎾绘煠閸偄鐏存鐐存崌楠炲洭顢楅埀顒傚緤閸ф鐓涢柛顐h壘娴滃墽绱撻崒娆戭槮闁绘锕ラ幈銊╁Χ婢跺﹤绐涙繝鐢靛Т閸燁垶鎮楅鈧弻?
|
||||
* 从多 Agent 执行状态中提取最终报告文本。
|
||||
*
|
||||
* @param state 闂備礁婀遍悷鎶藉幢閳哄倹鏉搁梻浣虹帛椤牓宕洪弽顓炵劦?
|
||||
* @return 闂備胶顢婄紙浼村磿闁秴绠熼柨鐔哄Т濡﹢鏌涢妷锝呭闁圭兘浜堕弻銊モ槈濡厧顣哄銈傛暘閸パ冨殤濠电姴锕ら崯浼村箺閻樼粯鐓曢柨鏃囧吹閸樻粎绱?
|
||||
* @param state Agent 图执行状态
|
||||
* @return Planner 输出中的最终报告
|
||||
*/
|
||||
public Optional<String> extractFinalReport(OverAllState state) {
|
||||
logger.info("闁诲孩顔栭崰鎺楀磻閹炬枼鏀芥い鏃傗拡閸庢劗鎲告0浣虹獢鐎规洩缍佸浠嬪Ω閿旇法甯涚紓鍌氬€风粈渚€鎮ф繝鍐╁弿闁靛牆顦?..");
|
||||
logger.info("Extracting final AI Ops report");
|
||||
|
||||
// 闂備礁婀辩划顖炲礉閺嚶颁汗?Planner 闂備礁鎼悧鍐磻閹惧墎纾藉ù锝呮憸婢э絿绱掓0婵嗗籍鐎规洘鐟╅幃顔锯偓闈涙憸椤︹晠姊洪崨濠勫ⅹ闁瑰啿閰i獮鍡涘醇閳垛晛浜鹃柣鐔哄濠€浼存煛閸☆厾绉柟顖氬暣瀹曠喖顢曢敐鍛畼闂佽崵濮崑鎾绘煥閺囨浜鹃梺鎼炲妼闁帮絽顕i幖浣哥疀妞ゆ挾鍊幘缁樼厱婵炴垶锕╅悡顓犵磼?
|
||||
Optional<AssistantMessage> plannerFinalOutput = state.value("planner_plan")
|
||||
.filter(AssistantMessage.class::isInstance)
|
||||
.map(AssistantMessage.class::cast);
|
||||
|
||||
if (plannerFinalOutput.isPresent()) {
|
||||
String reportText = plannerFinalOutput.get().getText();
|
||||
logger.info("闂備胶鎳撻悺銊╁礉閺囩喐鍙忔繛鎴欏灩缁犵敻鏌熼柇锕€澧紒鎻掓健閺?Planner 闂備礁鎼悧鍐磻閹惧墎纾藉ù锝呮憸婢ф稑鈹戦鍝勨偓婵嗙暦閵婏妇绡€闊洦娲滈ˇ顕€姊婚崒妤€浜鹃梺鍓茬厛閸犳牠顢? {}", reportText.length());
|
||||
logger.info("Extracted Planner final report, length: {}", reportText.length());
|
||||
return Optional.of(reportText);
|
||||
} else {
|
||||
logger.warn("Unable to extract Planner final report");
|
||||
@@ -285,7 +279,7 @@ public class AiOpsService {
|
||||
}
|
||||
|
||||
/**
|
||||
* 闂備礁鎼鍛偓姘煎墰缁?Planner Agent
|
||||
* 构建 Planner Agent。
|
||||
*/
|
||||
private ReactAgent buildPlannerAgent(ChatModel chatModel, ToolCallback[] toolCallbacks) {
|
||||
return ReactAgent.builder()
|
||||
@@ -299,7 +293,7 @@ public class AiOpsService {
|
||||
}
|
||||
|
||||
/**
|
||||
* 闂備礁鎼鍛偓姘煎墰缁?Executor Agent
|
||||
* 构建 Executor Agent。
|
||||
*/
|
||||
private ReactAgent buildExecutorAgent(ChatModel chatModel, ToolCallback[] toolCallbacks) {
|
||||
return ReactAgent.builder()
|
||||
@@ -315,16 +309,13 @@ public class AiOpsService {
|
||||
}
|
||||
|
||||
/**
|
||||
* 闂備礁鎲¢弻锝夊礉瀹ュ鐒垫い鎴f硶閸斿秹鏌f惔顔肩仩妞ゆ洘鐟╅幃婊兾熼懡銈呭箥婵犵數鍋涢ˇ鏉棵洪弽銊ヮ嚤闁圭増婢樼粈鍌炴⒑閸噮鍎愭繛鍫濆缁?
|
||||
* 闂備礁鎼粔鐑斤綖婢跺﹦鏆?cls.mock-enabled 闂備礁鎲¢崝鏇㈠疮閸ф鍋╁Δ锝呭暙閸欏﹥銇勯弽銊ь暡闁稿骸锕弻娑㈠冀瑜庨崳褰掓煙?QueryLogsTools
|
||||
* 闁诲氦顫夐幃鍫曞磿闁秴鐭楅柟绋跨昂娴滄粓鏌涘┑鍡楊伀缁炬澘绉归弻銊モ槈濞嗘劗娈ら梺缁樻惈缁辨洟骞忛悩璇插耿婵°倕鍟惃鎴︽⒑閸濆嫯顫﹂柛搴㈡尦椤㈡艾螖娴e壊鍤ゅ┑鈽嗗灠閹碱偆鏁妷鈺傜叆婵炴垶蓱濠€鐗堜繆椤愮喐娅堢紒鐘崇☉铻栧ù锝呮惈瀵劑鏌i悩鍙夊偍闁搞劍妞介、鏇㈠礂閼测斁鏋欓柣搴到婢у海绮堟径灞稿亾濞堝灝鏋涢柛鐔跺嵆瀵偊濡堕崪浣告櫊闂侀潧顦崕鍗烆嚗閺冨牊鐓涢柛顐h壘娴滈箖姊?
|
||||
* 根据运行模式构建方法工具数组。
|
||||
* Mock 模式注入本地 QueryLogsTools;真实模式下日志查询由外部 MCP 工具提供。
|
||||
*/
|
||||
private Object[] buildMethodToolsArray() {
|
||||
if (queryLogsTools != null) {
|
||||
// Mock 婵犵妲呴崹顏堝焵椤掆偓绾绢厾娑甸埀顒勬⒑閹稿海鈯曢柤鐟板⒔閳ь剙鐏氶敃銏犵暦?QueryLogsTools
|
||||
return new Object[]{dateTimeTools, lookupKnowledgeTool, queryMetricsTools, queryLogsTools};
|
||||
}
|
||||
// Real mode excludes local QueryLogsTools because logs are provided by MCP.
|
||||
return new Object[]{dateTimeTools, lookupKnowledgeTool, queryMetricsTools};
|
||||
}
|
||||
|
||||
@@ -347,7 +338,9 @@ public class AiOpsService {
|
||||
};
|
||||
}
|
||||
|
||||
/** 濠?agent_step 闂?tool_invocation 婵犳鍠氶幊鎾趁洪敃鍌氱劦妞ゆ帒鍊荤敮娑㈡倵閸偄鍝虹€殿喕绮欏畷鎯邦槼缂佲偓閳ь剚绻?diagnosis_session */
|
||||
/**
|
||||
* 从 agent_step 和 tool_invocation 回填 diagnosis_session 的汇总指标。
|
||||
*/
|
||||
private void backfillSessionMetrics(DiagnosisSession session) {
|
||||
try {
|
||||
List<AgentStep> steps = agentStepRepository.findBySessionIdOrderByStepIndex(session.getSessionId());
|
||||
@@ -363,7 +356,7 @@ public class AiOpsService {
|
||||
session.setStepCount(stepCount);
|
||||
session.setToolCallCount(Math.toIntExact(toolCallCount));
|
||||
} catch (Exception e) {
|
||||
logger.warn("闂備焦鎮堕崕鎶藉磻濞戙垺鏅查柣鎰綑椤曡鲸鎱ㄥΟ铏癸紞婵☆垰鐗撻弻鐔虹矙閹稿骸顦╅梺缁樼壄缁叉儳顕ラ崟顒佺秶妞ゆ劑鍎? sessionId={}", session.getSessionId(), e);
|
||||
logger.warn("Failed to backfill AI Ops session metrics, sessionId={}", session.getSessionId(), e);
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -126,11 +126,13 @@ public class DocumentManagementService {
|
||||
// 5. 解析 frontmatter
|
||||
long frontmatterStart = System.currentTimeMillis();
|
||||
Frontmatter frontmatter = null;
|
||||
String bodyText = text;
|
||||
if (frontmatterParser.hasFrontmatter(text)) {
|
||||
frontmatter = frontmatterParser.parse(text);
|
||||
if (frontmatter != null) {
|
||||
// LLM 补全 covers / whenToRetrieve(已有值则跳过)
|
||||
documentFieldEnricher.enrich(frontmatter, text, category);
|
||||
bodyText = frontmatterParser.stripFrontmatter(text);
|
||||
documentFieldEnricher.enrich(frontmatter, bodyText, category);
|
||||
log.info("解析到frontmatter: title={}, keywords={}, time={}ms",
|
||||
frontmatter.getTitle(), frontmatter.getKeywords(), System.currentTimeMillis() - frontmatterStart);
|
||||
} else {
|
||||
@@ -142,7 +144,7 @@ public class DocumentManagementService {
|
||||
|
||||
// 6. 分块
|
||||
long chunkStart = System.currentTimeMillis();
|
||||
List<DocumentChunk> chunks = documentChunkService.chunkDocument(text, fileName);
|
||||
List<DocumentChunk> chunks = documentChunkService.chunkDocument(bodyText, fileName);
|
||||
if (chunks.isEmpty()) {
|
||||
throw new DocumentProcessException(fileName, "upload", "文档分块失败");
|
||||
}
|
||||
@@ -150,7 +152,7 @@ public class DocumentManagementService {
|
||||
fileName, chunks.size(), System.currentTimeMillis() - chunkStart);
|
||||
|
||||
// 7. 创建文档元数据
|
||||
String docId = UUID.randomUUID().toString();
|
||||
String docId = resolveDocumentId(frontmatter);
|
||||
String metadataJson = null;
|
||||
if (frontmatter != null) {
|
||||
try {
|
||||
@@ -181,7 +183,7 @@ public class DocumentManagementService {
|
||||
// 8. 向量化并索引
|
||||
try {
|
||||
long vectorStart = System.currentTimeMillis();
|
||||
vectorIndexService.indexDocumentChunks(docId, chunks, category);
|
||||
vectorIndexService.indexDocumentChunks(docId, chunks, category, frontmatter);
|
||||
document.setStatus("INDEXED");
|
||||
document.setIndexedAt(LocalDateTime.now());
|
||||
apiDocumentRepository.save(document);
|
||||
@@ -203,6 +205,7 @@ public class DocumentManagementService {
|
||||
.keywords(frontmatter.getKeywords())
|
||||
.summary(frontmatter.getSummary())
|
||||
.category(category)
|
||||
.kbScope(frontmatter.getKbScope())
|
||||
.sections(frontmatter.getSections())
|
||||
.covers(frontmatter.getCovers())
|
||||
.whenToRetrieve(frontmatter.getWhenToRetrieve())
|
||||
@@ -311,6 +314,16 @@ public class DocumentManagementService {
|
||||
}
|
||||
}
|
||||
|
||||
private String resolveDocumentId(Frontmatter frontmatter) {
|
||||
if (frontmatter != null && frontmatter.getSource() != null) {
|
||||
String source = frontmatter.getSource().trim();
|
||||
if (!source.isEmpty() && source.length() <= 64) {
|
||||
return source;
|
||||
}
|
||||
}
|
||||
return UUID.randomUUID().toString();
|
||||
}
|
||||
|
||||
/**
|
||||
* 根据 docId 查询文档
|
||||
*/
|
||||
|
||||
@@ -62,6 +62,9 @@ public class FrontmatterParser {
|
||||
.keywords((java.util.List<String>) map.get("keywords"))
|
||||
.summary((String) map.get("summary"))
|
||||
.category((String) map.get("category"))
|
||||
.source((String) map.get("source"))
|
||||
.breadcrumb((String) map.get("breadcrumb"))
|
||||
.kbScope(firstString(map, "kb_scope", "kbScope"))
|
||||
.sections((Map<String, String>) map.get("sections"))
|
||||
.version((String) map.get("version"))
|
||||
.author((String) map.get("author"))
|
||||
@@ -87,6 +90,35 @@ public class FrontmatterParser {
|
||||
}
|
||||
}
|
||||
|
||||
public String stripFrontmatter(String content) {
|
||||
if (!hasFrontmatter(content)) {
|
||||
return content;
|
||||
}
|
||||
|
||||
String trimmed = content.trim();
|
||||
int secondDelimiter = trimmed.indexOf("\n---", 3);
|
||||
int delimiterLength = 4;
|
||||
if (secondDelimiter == -1) {
|
||||
secondDelimiter = trimmed.indexOf("\r\n---", 3);
|
||||
delimiterLength = 5;
|
||||
}
|
||||
if (secondDelimiter == -1) {
|
||||
return content;
|
||||
}
|
||||
|
||||
int bodyStart = secondDelimiter + delimiterLength;
|
||||
if (bodyStart < trimmed.length()) {
|
||||
char next = trimmed.charAt(bodyStart);
|
||||
if (next == '\r') {
|
||||
bodyStart++;
|
||||
}
|
||||
if (bodyStart < trimmed.length() && trimmed.charAt(bodyStart) == '\n') {
|
||||
bodyStart++;
|
||||
}
|
||||
}
|
||||
return trimmed.substring(Math.min(bodyStart, trimmed.length())).stripLeading();
|
||||
}
|
||||
|
||||
/**
|
||||
* 提取 frontmatter 文本(两个 --- 之间的内容)
|
||||
*
|
||||
@@ -115,4 +147,14 @@ public class FrontmatterParser {
|
||||
// 提取 frontmatter(不包含 --- 标记)
|
||||
return content.substring(3, secondDelimiter).trim();
|
||||
}
|
||||
|
||||
private String firstString(Map<String, Object> map, String... keys) {
|
||||
for (String key : keys) {
|
||||
Object value = map.get(key);
|
||||
if (value instanceof String text && !text.isBlank()) {
|
||||
return text;
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,70 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.superbiz.agent.dto.ContextPack;
|
||||
import com.superbiz.agent.dto.EvidenceBlock;
|
||||
import org.springframework.beans.factory.annotation.Value;
|
||||
import org.springframework.stereotype.Service;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.List;
|
||||
|
||||
/**
|
||||
* Packs final evidence blocks into compact Agent-facing context.
|
||||
*/
|
||||
@Service
|
||||
public class KnowledgeContextPacker {
|
||||
|
||||
@Value("${rag.context-pack.char-budget:4000}")
|
||||
private int charBudget = 4000;
|
||||
|
||||
public ContextPack pack(List<EvidenceBlock> blocks) {
|
||||
List<EvidenceBlock> safeBlocks = blocks == null ? List.of() : blocks;
|
||||
StringBuilder packed = new StringBuilder();
|
||||
List<String> included = new ArrayList<>();
|
||||
List<String> omitted = new ArrayList<>();
|
||||
|
||||
int rank = 1;
|
||||
for (EvidenceBlock block : safeBlocks) {
|
||||
String header = buildHeader(rank, block);
|
||||
String content = block.getContent() == null ? "" : block.getContent();
|
||||
int remaining = charBudget - packed.length() - header.length();
|
||||
if (remaining <= 0) {
|
||||
omitted.add(block.getSource());
|
||||
continue;
|
||||
}
|
||||
String body = content.length() <= remaining ? content : content.substring(0, Math.max(0, remaining)) + "...";
|
||||
packed.append(header).append(body).append("\n\n");
|
||||
included.add(block.getSource());
|
||||
rank++;
|
||||
}
|
||||
|
||||
return ContextPack.builder()
|
||||
.packedText(packed.toString().trim())
|
||||
.strategy("ranked_evidence_char_budget")
|
||||
.charBudget(charBudget)
|
||||
.usedChars(packed.length())
|
||||
.includedSources(included)
|
||||
.omittedSources(omitted)
|
||||
.build();
|
||||
}
|
||||
|
||||
private String buildHeader(int rank, EvidenceBlock block) {
|
||||
StringBuilder header = new StringBuilder();
|
||||
header.append("[Evidence ").append(rank).append("]\n");
|
||||
appendLine(header, "source", block.getSource());
|
||||
appendLine(header, "title", block.getTitle());
|
||||
appendLine(header, "breadcrumb", block.getBreadcrumb());
|
||||
appendLine(header, "layer", block.getRetrievalLayer());
|
||||
if (block.getHitReasons() != null && !block.getHitReasons().isEmpty()) {
|
||||
appendLine(header, "reasons", String.join(", ", block.getHitReasons()));
|
||||
}
|
||||
header.append("content:\n");
|
||||
return header.toString();
|
||||
}
|
||||
|
||||
private void appendLine(StringBuilder builder, String key, String value) {
|
||||
if (value != null && !value.isBlank()) {
|
||||
builder.append(key).append(": ").append(value).append("\n");
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,139 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.dto.RetrievalTrace;
|
||||
import com.superbiz.agent.dto.RetrievedEvidenceCandidate;
|
||||
import org.springframework.stereotype.Service;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
/**
|
||||
* Vector retrieval adapter for the modular knowledge pipeline.
|
||||
*/
|
||||
@Service
|
||||
public class KnowledgeDocumentRetriever {
|
||||
|
||||
private final VectorSearchService vectorSearchService;
|
||||
private final ObjectMapper objectMapper;
|
||||
|
||||
public KnowledgeDocumentRetriever(VectorSearchService vectorSearchService, ObjectMapper objectMapper) {
|
||||
this.vectorSearchService = vectorSearchService;
|
||||
this.objectMapper = objectMapper;
|
||||
}
|
||||
|
||||
public RetrievalAttemptResult retrieve(String attemptName, String query, String categoryFilter, int topK) {
|
||||
long start = System.currentTimeMillis();
|
||||
try {
|
||||
List<VectorSearchService.SearchResult> results =
|
||||
vectorSearchService.searchSimilarDocuments(query, topK, categoryFilter);
|
||||
List<RetrievedEvidenceCandidate> candidates = toCandidates(attemptName, results);
|
||||
return new RetrievalAttemptResult(
|
||||
attempt(attemptName, query, categoryFilter, candidates.size(), null,
|
||||
(int) (System.currentTimeMillis() - start), topScore(results)),
|
||||
candidates
|
||||
);
|
||||
} catch (Exception e) {
|
||||
return new RetrievalAttemptResult(
|
||||
attempt(attemptName, query, categoryFilter, 0, e.getMessage(),
|
||||
(int) (System.currentTimeMillis() - start), null),
|
||||
List.of()
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
private List<RetrievedEvidenceCandidate> toCandidates(String attemptName,
|
||||
List<VectorSearchService.SearchResult> results) {
|
||||
if (results == null || results.isEmpty()) {
|
||||
return List.of();
|
||||
}
|
||||
List<RetrievedEvidenceCandidate> candidates = new ArrayList<>();
|
||||
for (int i = 0; i < results.size(); i++) {
|
||||
VectorSearchService.SearchResult result = results.get(i);
|
||||
Map<String, String> metadata = parseMetadata(result.getMetadata());
|
||||
String source = firstNonBlank(
|
||||
metadata.get("_source"),
|
||||
metadata.get("source"),
|
||||
metadata.get("filePath"),
|
||||
metadata.get("docId"),
|
||||
result.getMetadata(),
|
||||
result.getId()
|
||||
);
|
||||
candidates.add(RetrievedEvidenceCandidate.builder()
|
||||
.id(result.getId())
|
||||
.source(source)
|
||||
.title(metadata.get("title"))
|
||||
.breadcrumb(metadata.get("breadcrumb"))
|
||||
.content(result.getContent())
|
||||
.retrievalLayer("L1")
|
||||
.retrievalAttempt(attemptName)
|
||||
.score((double) result.getScore())
|
||||
.rawScore(result.getRawScore())
|
||||
.scoreLabel(result.getScoreLabel())
|
||||
.originalRank(i + 1)
|
||||
.metadata(metadata)
|
||||
.hitReasons(List.of("semantic_rank:" + (i + 1), "attempt:" + attemptName))
|
||||
.build());
|
||||
}
|
||||
return candidates;
|
||||
}
|
||||
|
||||
private RetrievalTrace.Attempt attempt(String name,
|
||||
String query,
|
||||
String categoryFilter,
|
||||
int candidateCount,
|
||||
String errorMessage,
|
||||
int durationMs,
|
||||
Double topScore) {
|
||||
return RetrievalTrace.Attempt.builder()
|
||||
.name(name)
|
||||
.query(query)
|
||||
.categoryFilter(categoryFilter)
|
||||
.candidateCount(candidateCount)
|
||||
.usable(errorMessage == null && candidateCount > 0)
|
||||
.errorMessage(errorMessage)
|
||||
.durationMs(durationMs)
|
||||
.topScore(topScore)
|
||||
.build();
|
||||
}
|
||||
|
||||
private Double topScore(List<VectorSearchService.SearchResult> results) {
|
||||
if (results == null || results.isEmpty()) {
|
||||
return null;
|
||||
}
|
||||
return (double) results.get(0).getScore();
|
||||
}
|
||||
|
||||
private Map<String, String> parseMetadata(String metadata) {
|
||||
if (metadata == null || metadata.isBlank()) {
|
||||
return Map.of();
|
||||
}
|
||||
try {
|
||||
Map<?, ?> raw = objectMapper.readValue(metadata, Map.class);
|
||||
Map<String, String> parsed = new LinkedHashMap<>();
|
||||
for (Map.Entry<?, ?> entry : raw.entrySet()) {
|
||||
if (entry.getKey() != null && entry.getValue() != null) {
|
||||
parsed.put(String.valueOf(entry.getKey()), String.valueOf(entry.getValue()));
|
||||
}
|
||||
}
|
||||
return parsed;
|
||||
} catch (Exception ignored) {
|
||||
return Map.of();
|
||||
}
|
||||
}
|
||||
|
||||
private String firstNonBlank(String... values) {
|
||||
for (String value : values) {
|
||||
if (value != null && !value.isBlank()) {
|
||||
return value;
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
public record RetrievalAttemptResult(RetrievalTrace.Attempt attempt,
|
||||
List<RetrievedEvidenceCandidate> candidates) {
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,251 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.superbiz.agent.dto.EvidenceBlock;
|
||||
import com.superbiz.agent.dto.EvidencePostprocessResult;
|
||||
import com.superbiz.agent.dto.KnowledgeQuery;
|
||||
import com.superbiz.agent.dto.RerankTrace;
|
||||
import com.superbiz.agent.dto.RetrievedEvidenceCandidate;
|
||||
import org.springframework.beans.factory.annotation.Value;
|
||||
import org.springframework.stereotype.Service;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.Comparator;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.LinkedHashSet;
|
||||
import java.util.List;
|
||||
import java.util.Locale;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
|
||||
/**
|
||||
* Post-retrieval evidence normalization, rerank, and evidence block assembly.
|
||||
*/
|
||||
@Service
|
||||
public class KnowledgeEvidencePostProcessor {
|
||||
|
||||
private static final String LEVEL_PRECISE = "PRECISE";
|
||||
private static final String LEVEL_HIGHLY_RELEVANT = "HIGHLY_RELEVANT";
|
||||
private static final String LEVEL_REFERENCE = "REFERENCE";
|
||||
|
||||
private static final String HINT_PRECISE = "知识库中不存在比上述结果更精准的文档";
|
||||
private static final String HINT_HIGHLY_RELEVANT = "当前结果已高度相关,继续检索不太可能找到更精准的文档";
|
||||
private static final String HINT_REFERENCE = "当前结果为相关参考,如需更精准信息请明确缺少的具体维度";
|
||||
|
||||
@Value("${retrieval.normalization.max-l2-distance:2.0}")
|
||||
private double maxL2Distance = 2.0;
|
||||
|
||||
@Value("${retrieval.normalization.highly-relevant-threshold:0.75}")
|
||||
private double highlyRelevantThreshold = 0.75;
|
||||
|
||||
@Value("${retrieval.normalization.reference-threshold:0.5}")
|
||||
private double referenceThreshold = 0.5;
|
||||
|
||||
public EvidencePostprocessResult process(KnowledgeQuery query, List<RetrievedEvidenceCandidate> candidates) {
|
||||
List<RetrievedEvidenceCandidate> safeCandidates = candidates == null ? List.of() : candidates;
|
||||
List<ScoredCandidate> ranked = safeCandidates.stream()
|
||||
.map(candidate -> score(query, candidate))
|
||||
.sorted(Comparator.comparingDouble(ScoredCandidate::finalScore).reversed())
|
||||
.toList();
|
||||
|
||||
Map<String, EvidenceBlock> deduped = new LinkedHashMap<>();
|
||||
List<RerankTrace.Item> traceItems = new ArrayList<>();
|
||||
int finalRank = 1;
|
||||
for (ScoredCandidate scored : ranked) {
|
||||
RetrievedEvidenceCandidate candidate = scored.candidate();
|
||||
EvidenceBlock block = EvidenceBlock.builder()
|
||||
.source(candidate.getSource())
|
||||
.title(candidate.getTitle())
|
||||
.breadcrumb(candidate.getBreadcrumb())
|
||||
.retrievalLayer(candidate.getRetrievalLayer())
|
||||
.content(truncate(candidate.getContent(), 800))
|
||||
.score(candidate.getScore())
|
||||
.hitReasons(mergeReasons(candidate.getHitReasons(), scored.boostReasons()))
|
||||
.build();
|
||||
String key = sourceKey(block, "candidate-" + candidate.getOriginalRank());
|
||||
if (!deduped.containsKey(key)) {
|
||||
deduped.put(key, block);
|
||||
traceItems.add(RerankTrace.Item.builder()
|
||||
.finalRank(finalRank++)
|
||||
.source(candidate.getSource())
|
||||
.baseScore(scored.baseScore())
|
||||
.finalScore(scored.finalScore())
|
||||
.boostReasons(scored.boostReasons())
|
||||
.build());
|
||||
} else {
|
||||
mergeEvidence(deduped.get(key), block);
|
||||
}
|
||||
}
|
||||
|
||||
List<EvidenceBlock> blocks = new ArrayList<>(deduped.values());
|
||||
Double topSimilarity = ranked.isEmpty() ? null : ranked.get(0).baseScore();
|
||||
RelevanceAssessment assessment = computeRelevance(query, ranked);
|
||||
return EvidencePostprocessResult.builder()
|
||||
.candidateCount(safeCandidates.size())
|
||||
.evidenceBlockCount(blocks.size())
|
||||
.evidenceBlocks(blocks)
|
||||
.relevanceLevel(assessment.level())
|
||||
.completenessHint(assessment.hint())
|
||||
.rerankTrace(RerankTrace.builder().items(traceItems).build())
|
||||
.topSimilarity(topSimilarity)
|
||||
.build();
|
||||
}
|
||||
|
||||
public boolean isLowQuality(EvidencePostprocessResult result) {
|
||||
if (result == null || !result.hasUsableEvidence()) {
|
||||
return true;
|
||||
}
|
||||
Double topSimilarity = result.getTopSimilarity();
|
||||
return topSimilarity == null || topSimilarity < referenceThreshold;
|
||||
}
|
||||
|
||||
public double normalizeL2(Double l2Score) {
|
||||
if (l2Score == null) {
|
||||
return 0.0;
|
||||
}
|
||||
double clamped = Math.min(l2Score, maxL2Distance);
|
||||
return Math.max(0.0, 1.0 - clamped / maxL2Distance);
|
||||
}
|
||||
|
||||
public double getReferenceThreshold() {
|
||||
return referenceThreshold;
|
||||
}
|
||||
|
||||
private ScoredCandidate score(KnowledgeQuery query, RetrievedEvidenceCandidate candidate) {
|
||||
double baseScore = normalizeL2(candidate.getScore());
|
||||
double finalScore = baseScore;
|
||||
List<String> boosts = new ArrayList<>();
|
||||
|
||||
if (matchesAny(candidate, query.getDomainHints())) {
|
||||
finalScore += 0.15;
|
||||
boosts.add("domain_match:+0.15");
|
||||
}
|
||||
if (matchesAny(candidate, query.getEntities())) {
|
||||
finalScore += 0.20;
|
||||
boosts.add("entity_match:+0.20");
|
||||
}
|
||||
if (matchesAny(candidate, query.getMatchedKeywords())) {
|
||||
finalScore += 0.10;
|
||||
boosts.add("keyword_match:+0.10");
|
||||
}
|
||||
if (isPreferredSourceType(candidate)) {
|
||||
finalScore += 0.05;
|
||||
boosts.add("source_type:+0.05");
|
||||
}
|
||||
|
||||
return new ScoredCandidate(candidate, baseScore, finalScore, boosts);
|
||||
}
|
||||
|
||||
private boolean matchesAny(RetrievedEvidenceCandidate candidate, List<String> hints) {
|
||||
if (hints == null || hints.isEmpty()) {
|
||||
return false;
|
||||
}
|
||||
String haystack = String.join(" ",
|
||||
nullToEmpty(candidate.getSource()),
|
||||
nullToEmpty(candidate.getTitle()),
|
||||
nullToEmpty(candidate.getBreadcrumb()),
|
||||
nullToEmpty(candidate.getContent()),
|
||||
candidate.getMetadata() == null ? "" : candidate.getMetadata().toString()
|
||||
).toLowerCase(Locale.ROOT);
|
||||
for (String hint : hints) {
|
||||
if (hint != null && !hint.isBlank() && haystack.contains(hint.toLowerCase(Locale.ROOT))) {
|
||||
return true;
|
||||
}
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
private boolean isPreferredSourceType(RetrievedEvidenceCandidate candidate) {
|
||||
Map<String, String> metadata = candidate.getMetadata();
|
||||
if (metadata == null || metadata.isEmpty()) {
|
||||
return false;
|
||||
}
|
||||
String type = firstNonBlank(metadata.get("source_type"), metadata.get("documentType"), metadata.get("type"));
|
||||
if (type == null) {
|
||||
return false;
|
||||
}
|
||||
String normalized = type.toLowerCase(Locale.ROOT);
|
||||
return normalized.contains("runbook") || normalized.contains("guide") || normalized.contains("case");
|
||||
}
|
||||
|
||||
private RelevanceAssessment computeRelevance(KnowledgeQuery query, List<ScoredCandidate> ranked) {
|
||||
if (ranked.isEmpty()) {
|
||||
return new RelevanceAssessment(null, null);
|
||||
}
|
||||
ScoredCandidate top = ranked.get(0);
|
||||
if (top.baseScore() >= highlyRelevantThreshold && hasHintSupport(query, top)) {
|
||||
return new RelevanceAssessment(LEVEL_PRECISE, HINT_PRECISE);
|
||||
}
|
||||
if (top.baseScore() >= highlyRelevantThreshold) {
|
||||
return new RelevanceAssessment(LEVEL_HIGHLY_RELEVANT, HINT_HIGHLY_RELEVANT);
|
||||
}
|
||||
if (top.baseScore() >= referenceThreshold) {
|
||||
return new RelevanceAssessment(LEVEL_REFERENCE, HINT_REFERENCE);
|
||||
}
|
||||
return new RelevanceAssessment(null, null);
|
||||
}
|
||||
|
||||
private boolean hasHintSupport(KnowledgeQuery query, ScoredCandidate top) {
|
||||
return matchesAny(top.candidate(), query.getDomainHints())
|
||||
|| matchesAny(top.candidate(), query.getEntities())
|
||||
|| matchesAny(top.candidate(), query.getMatchedKeywords());
|
||||
}
|
||||
|
||||
private void mergeEvidence(EvidenceBlock existing, EvidenceBlock incoming) {
|
||||
Set<String> reasons = new LinkedHashSet<>();
|
||||
if (existing.getHitReasons() != null) {
|
||||
reasons.addAll(existing.getHitReasons());
|
||||
}
|
||||
if (incoming.getHitReasons() != null) {
|
||||
reasons.addAll(incoming.getHitReasons());
|
||||
}
|
||||
existing.setHitReasons(new ArrayList<>(reasons));
|
||||
if ((existing.getBreadcrumb() == null || existing.getBreadcrumb().isBlank())
|
||||
&& incoming.getBreadcrumb() != null) {
|
||||
existing.setBreadcrumb(incoming.getBreadcrumb());
|
||||
}
|
||||
}
|
||||
|
||||
private List<String> mergeReasons(List<String> base, List<String> boosts) {
|
||||
Set<String> merged = new LinkedHashSet<>();
|
||||
if (base != null) {
|
||||
merged.addAll(base);
|
||||
}
|
||||
if (boosts != null) {
|
||||
merged.addAll(boosts);
|
||||
}
|
||||
return new ArrayList<>(merged);
|
||||
}
|
||||
|
||||
private String sourceKey(EvidenceBlock block, String fallback) {
|
||||
return firstNonBlank(block.getSource(), block.getTitle(), block.getBreadcrumb(), fallback);
|
||||
}
|
||||
|
||||
private String truncate(String text, int maxLength) {
|
||||
if (text == null || text.length() <= maxLength) {
|
||||
return text;
|
||||
}
|
||||
return text.substring(0, maxLength) + "...";
|
||||
}
|
||||
|
||||
private String firstNonBlank(String... values) {
|
||||
for (String value : values) {
|
||||
if (value != null && !value.isBlank()) {
|
||||
return value;
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
private String nullToEmpty(String value) {
|
||||
return value == null ? "" : value;
|
||||
}
|
||||
|
||||
private record ScoredCandidate(RetrievedEvidenceCandidate candidate,
|
||||
double baseScore,
|
||||
double finalScore,
|
||||
List<String> boostReasons) {
|
||||
}
|
||||
|
||||
private record RelevanceAssessment(String level, String hint) {
|
||||
}
|
||||
}
|
||||
@@ -36,6 +36,9 @@ public class KnowledgeIndexService {
|
||||
@Value("${knowledge.base-path:knowledge_base}")
|
||||
private String knowledgeBasePath;
|
||||
|
||||
@Value("${retrieval.kb-scope:}")
|
||||
private String kbScope = "";
|
||||
|
||||
@Autowired
|
||||
private ApiDocumentRepository apiDocumentRepository;
|
||||
|
||||
@@ -114,6 +117,7 @@ public class KnowledgeIndexService {
|
||||
.keywords(frontmatter.getKeywords())
|
||||
.summary(frontmatter.getSummary())
|
||||
.category(frontmatter.getCategory())
|
||||
.kbScope(frontmatter.getKbScope())
|
||||
.covers(frontmatter.getCovers())
|
||||
.whenToRetrieve(frontmatter.getWhenToRetrieve())
|
||||
.build();
|
||||
@@ -144,6 +148,9 @@ public class KnowledgeIndexService {
|
||||
Set<String> titles = new LinkedHashSet<>();
|
||||
|
||||
for (KnowledgeEntry entry : knowledgeIndex) {
|
||||
if (!matchesConfiguredScope(entry)) {
|
||||
continue;
|
||||
}
|
||||
List<String> entryMatchedKeywords = matchedKeywords(entry, queryLower);
|
||||
if (entryMatchedKeywords.isEmpty()) {
|
||||
continue;
|
||||
@@ -178,6 +185,21 @@ public class KnowledgeIndexService {
|
||||
return !matchedKeywords(entry, query).isEmpty();
|
||||
}
|
||||
|
||||
private boolean matchesConfiguredScope(KnowledgeEntry entry) {
|
||||
String scope = trimToNull(kbScope);
|
||||
if (scope == null) {
|
||||
return true;
|
||||
}
|
||||
return scope.equals(trimToNull(entry.getKbScope()));
|
||||
}
|
||||
|
||||
private String trimToNull(String value) {
|
||||
if (value == null || value.isBlank()) {
|
||||
return null;
|
||||
}
|
||||
return value.trim();
|
||||
}
|
||||
|
||||
private List<String> matchedKeywords(KnowledgeEntry entry, String query) {
|
||||
if (entry.getKeywords() == null || entry.getKeywords().isEmpty()) {
|
||||
return List.of();
|
||||
|
||||
@@ -0,0 +1,38 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.superbiz.agent.dto.KnowledgeQuery;
|
||||
import org.springframework.stereotype.Service;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
/**
|
||||
* Converts a raw Agent query into retrieval-control hints.
|
||||
*/
|
||||
@Service
|
||||
public class KnowledgeQueryTransformer {
|
||||
|
||||
private final KnowledgeIndexService knowledgeIndexService;
|
||||
|
||||
public KnowledgeQueryTransformer(KnowledgeIndexService knowledgeIndexService) {
|
||||
this.knowledgeIndexService = knowledgeIndexService;
|
||||
}
|
||||
|
||||
public KnowledgeQuery transform(String rawQuery) {
|
||||
String normalized = rawQuery == null ? "" : rawQuery.trim();
|
||||
KnowledgeIndexService.L0Hint hint = knowledgeIndexService.analyzeQuery(normalized);
|
||||
return KnowledgeQuery.builder()
|
||||
.originalQuery(normalized)
|
||||
.rewrittenQuery(normalized)
|
||||
.domainHints(safeList(hint.domains()))
|
||||
.matchedKeywords(safeList(hint.matchedKeywords()))
|
||||
.entities(safeList(hint.entities()))
|
||||
.categoryFilter(hint.singleDomainOrNull())
|
||||
.l0Titles(safeList(hint.titles()))
|
||||
.l0MatchCount(hint.matches() == null ? 0 : hint.matches().size())
|
||||
.build();
|
||||
}
|
||||
|
||||
private List<String> safeList(List<String> values) {
|
||||
return values == null ? List.of() : values;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,49 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.superbiz.agent.dto.ContextPack;
|
||||
import com.superbiz.agent.dto.EvidencePostprocessResult;
|
||||
import com.superbiz.agent.dto.LookupResult;
|
||||
import com.superbiz.agent.dto.RetrievalTrace;
|
||||
import org.springframework.stereotype.Service;
|
||||
|
||||
import java.util.List;
|
||||
|
||||
/**
|
||||
* Assembles evidence-first LookupResult instances.
|
||||
*/
|
||||
@Service
|
||||
public class LookupResultAssembler {
|
||||
|
||||
public LookupResult assemble(EvidencePostprocessResult evidence,
|
||||
ContextPack contextPack,
|
||||
RetrievalTrace retrievalTrace) {
|
||||
boolean found = evidence != null && evidence.hasUsableEvidence();
|
||||
return LookupResult.builder()
|
||||
.found(found)
|
||||
.evidenceBlocks(evidence != null ? evidence.getEvidenceBlocks() : List.of())
|
||||
.evidenceCandidateCount(evidence != null ? evidence.getCandidateCount() : 0)
|
||||
.evidenceBlockCount(evidence != null ? evidence.getEvidenceBlockCount() : 0)
|
||||
.contextPack(contextPack)
|
||||
.retrievalTrace(retrievalTrace)
|
||||
.rerankTrace(evidence != null ? evidence.getRerankTrace() : null)
|
||||
.relevanceLevel(evidence != null ? evidence.getRelevanceLevel() : null)
|
||||
.completenessHint(evidence != null ? evidence.getCompletenessHint() : null)
|
||||
.message(found ? null : "知识库未检索到可用证据,请结合日志、指标、告警继续排查")
|
||||
.build();
|
||||
}
|
||||
|
||||
public LookupResult deduped(LookupResult original, List<String> retrievedDomains, String docKey) {
|
||||
return LookupResult.builder()
|
||||
.found(false)
|
||||
.message("文档已在本会话中检索过,无需重复召回: " + docKey)
|
||||
.evidenceBlocks(List.of())
|
||||
.evidenceCandidateCount(0)
|
||||
.evidenceBlockCount(0)
|
||||
.retrievalTrace(original.getRetrievalTrace())
|
||||
.rerankTrace(original.getRerankTrace())
|
||||
.relevanceLevel(original.getRelevanceLevel())
|
||||
.completenessHint(original.getCompletenessHint())
|
||||
.retrievedDomainsThisSession(retrievedDomains)
|
||||
.build();
|
||||
}
|
||||
}
|
||||
@@ -8,6 +8,7 @@ import org.springframework.ai.document.Document;
|
||||
import org.springframework.ai.vectorstore.SearchRequest;
|
||||
import org.springframework.ai.vectorstore.VectorStore;
|
||||
import org.springframework.beans.factory.ObjectProvider;
|
||||
import org.springframework.beans.factory.annotation.Value;
|
||||
import org.springframework.stereotype.Service;
|
||||
|
||||
import java.util.ArrayList;
|
||||
@@ -21,6 +22,9 @@ public class SpringAiVectorStoreSidecarService {
|
||||
private final ObjectProvider<VectorStore> vectorStoreProvider;
|
||||
private final RetrievalResultNormalizer normalizer;
|
||||
|
||||
@Value("${retrieval.kb-scope:}")
|
||||
private String kbScope = "";
|
||||
|
||||
public SpringAiVectorStoreSidecarService(RagSidecarProperties properties,
|
||||
ObjectProvider<VectorStore> vectorStoreProvider,
|
||||
RetrievalResultNormalizer normalizer) {
|
||||
@@ -44,8 +48,9 @@ public class SpringAiVectorStoreSidecarService {
|
||||
.query(query)
|
||||
.topK(topK)
|
||||
.similarityThresholdAll();
|
||||
if (category != null && !category.isBlank()) {
|
||||
builder.filterExpression("category == '" + escapeFilterValue(category) + "'");
|
||||
String filterExpression = buildFilterExpression(category);
|
||||
if (filterExpression != null) {
|
||||
builder.filterExpression(filterExpression);
|
||||
}
|
||||
|
||||
List<Document> documents = vectorStore.similaritySearch(builder.build());
|
||||
@@ -78,4 +83,24 @@ public class SpringAiVectorStoreSidecarService {
|
||||
private String escapeFilterValue(String value) {
|
||||
return value.replace("'", "\\'");
|
||||
}
|
||||
|
||||
String buildFilterExpression(String category) {
|
||||
List<String> parts = new ArrayList<>();
|
||||
String categoryFilter = trimToNull(category);
|
||||
if (categoryFilter != null) {
|
||||
parts.add("category == '" + escapeFilterValue(categoryFilter) + "'");
|
||||
}
|
||||
String scopeFilter = trimToNull(kbScope);
|
||||
if (scopeFilter != null) {
|
||||
parts.add("kb_scope == '" + escapeFilterValue(scopeFilter) + "'");
|
||||
}
|
||||
return parts.isEmpty() ? null : String.join(" && ", parts);
|
||||
}
|
||||
|
||||
private String trimToNull(String value) {
|
||||
if (value == null || value.isBlank()) {
|
||||
return null;
|
||||
}
|
||||
return value.trim();
|
||||
}
|
||||
}
|
||||
|
||||
@@ -3,10 +3,13 @@ package com.superbiz.agent.service;
|
||||
import com.fasterxml.jackson.core.JsonProcessingException;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.domain.entity.ToolInvocation;
|
||||
import com.superbiz.agent.dto.ContextPack;
|
||||
import com.superbiz.agent.dto.EvidenceBlock;
|
||||
import com.superbiz.agent.dto.KnowledgeQuery;
|
||||
import com.superbiz.agent.dto.LookupResult;
|
||||
import com.superbiz.agent.dto.RerankTrace;
|
||||
import com.superbiz.agent.dto.RetrievalTrace;
|
||||
import com.superbiz.agent.repository.ToolInvocationRepository;
|
||||
import com.superbiz.agent.dto.KnowledgeEntry;
|
||||
import com.superbiz.agent.util.SessionContextHolder;
|
||||
import lombok.Builder;
|
||||
import lombok.extern.slf4j.Slf4j;
|
||||
@@ -138,6 +141,21 @@ public class ToolInvocationRecorder {
|
||||
if (record.evidenceBlocks() != null && !record.evidenceBlocks().isEmpty()) {
|
||||
details.put("evidence_blocks", record.evidenceBlocks());
|
||||
}
|
||||
if (record.queryTransform() != null && !record.queryTransform().isEmpty()) {
|
||||
details.put("query_transform", record.queryTransform());
|
||||
}
|
||||
if (record.retrievalTrace() != null && !record.retrievalTrace().isEmpty()) {
|
||||
details.put("retrieval_trace", record.retrievalTrace());
|
||||
}
|
||||
if (record.contextPack() != null && !record.contextPack().isEmpty()) {
|
||||
details.put("context_pack_summary", record.contextPack());
|
||||
}
|
||||
if (record.rerankTrace() != null && !record.rerankTrace().isEmpty()) {
|
||||
details.put("rerank_trace", record.rerankTrace());
|
||||
}
|
||||
if (record.fallbackReason() != null && !record.fallbackReason().isBlank()) {
|
||||
details.put("fallback_reason", record.fallbackReason());
|
||||
}
|
||||
if (record.relevanceLevel() != null) {
|
||||
details.put("relevance_level", record.relevanceLevel());
|
||||
}
|
||||
@@ -222,40 +240,33 @@ public class ToolInvocationRecorder {
|
||||
List<Double> l1Scores,
|
||||
Integer evidenceCandidateCount,
|
||||
Integer evidenceBlockCount,
|
||||
List<Map<String, Object>> evidenceBlocks
|
||||
List<Map<String, Object>> evidenceBlocks,
|
||||
Map<String, Object> queryTransform,
|
||||
Map<String, Object> retrievalTrace,
|
||||
Map<String, Object> contextPack,
|
||||
Map<String, Object> rerankTrace,
|
||||
String fallbackReason
|
||||
) {
|
||||
public static LookupKnowledgeRecord from(String query,
|
||||
KnowledgeIndexService.L0Hint l0Hint,
|
||||
List<VectorSearchService.SearchResult> l1Results,
|
||||
boolean highConfidence,
|
||||
public static LookupKnowledgeRecord from(KnowledgeQuery query,
|
||||
LookupResult result,
|
||||
String domain,
|
||||
String dedupReason,
|
||||
int durationMs,
|
||||
double l1TopSimilarity) {
|
||||
List<KnowledgeEntry> l0Matches = l0Hint != null ? l0Hint.matches() : List.of();
|
||||
boolean hasL0 = l0Matches != null && !l0Matches.isEmpty();
|
||||
boolean hasL1 = l1Results != null && !l1Results.isEmpty();
|
||||
String layer;
|
||||
if (hasL0 && !highConfidence) {
|
||||
layer = "L0+L1";
|
||||
} else if (hasL0) {
|
||||
layer = "L0";
|
||||
} else if (hasL1) {
|
||||
layer = "L1";
|
||||
} else {
|
||||
layer = null;
|
||||
}
|
||||
int durationMs) {
|
||||
RetrievalTrace trace = result != null ? result.getRetrievalTrace() : null;
|
||||
String layer = trace != null ? trace.getSelectedAttempt() : null;
|
||||
|
||||
String outputPreview = null;
|
||||
int outputLength = 0;
|
||||
boolean truncated = false;
|
||||
if (result != null && result.getPrimary() != null && result.getPrimary().getContent() != null) {
|
||||
outputPreview = result.getPrimary().getContent();
|
||||
if (result != null && result.getContextPack() != null
|
||||
&& result.getContextPack().getPackedText() != null) {
|
||||
outputPreview = result.getContextPack().getPackedText();
|
||||
outputLength = outputPreview.length();
|
||||
truncated = outputLength > OUTPUT_PREVIEW_LIMIT;
|
||||
} else if (hasL1 && l1Results.get(0).getContent() != null) {
|
||||
outputPreview = l1Results.get(0).getContent();
|
||||
} else if (result != null && result.getEvidenceBlocks() != null
|
||||
&& !result.getEvidenceBlocks().isEmpty()
|
||||
&& result.getEvidenceBlocks().get(0).getContent() != null) {
|
||||
outputPreview = result.getEvidenceBlocks().get(0).getContent();
|
||||
outputLength = outputPreview.length();
|
||||
truncated = outputLength > OUTPUT_PREVIEW_LIMIT;
|
||||
}
|
||||
@@ -265,29 +276,19 @@ public class ToolInvocationRecorder {
|
||||
evidenceStatus = EVIDENCE_STATUS_DEDUPED;
|
||||
} else if (result == null || !result.isFound()) {
|
||||
evidenceStatus = EVIDENCE_STATUS_NO_EVIDENCE;
|
||||
} else if (trace != null && trace.getEvidenceStatus() != null) {
|
||||
evidenceStatus = trace.getEvidenceStatus();
|
||||
}
|
||||
|
||||
List<String> l0Titles = new ArrayList<>();
|
||||
if (hasL0) {
|
||||
for (int i = 0; i < Math.min(3, l0Matches.size()); i++) {
|
||||
l0Titles.add(l0Matches.get(i).getTitle());
|
||||
}
|
||||
}
|
||||
|
||||
List<Double> l1Scores = new ArrayList<>();
|
||||
if (hasL1) {
|
||||
for (int i = 0; i < Math.min(3, l1Results.size()); i++) {
|
||||
l1Scores.add((double) l1Results.get(i).getScore());
|
||||
}
|
||||
}
|
||||
List<Double> l1Scores = collectAttemptScores(trace);
|
||||
|
||||
return LookupKnowledgeRecord.builder()
|
||||
.query(query)
|
||||
.query(query != null ? query.getOriginalQuery() : null)
|
||||
.outputPreview(outputPreview)
|
||||
.outputLength(outputLength)
|
||||
.retrievalLayer(layer)
|
||||
.l0MatchCount(hasL0 ? l0Matches.size() : null)
|
||||
.l1MatchCount(hasL1 ? l1Results.size() : null)
|
||||
.l0MatchCount(query != null ? query.getL0MatchCount() : null)
|
||||
.l1MatchCount(totalCandidateCount(trace))
|
||||
.truncated(truncated)
|
||||
.relevanceLevel(result != null ? result.getRelevanceLevel() : null)
|
||||
.completenessHint(result != null ? result.getCompletenessHint() : null)
|
||||
@@ -296,16 +297,21 @@ public class ToolInvocationRecorder {
|
||||
.durationMs(durationMs)
|
||||
.success(true)
|
||||
.evidenceStatus(evidenceStatus)
|
||||
.l0Titles(l0Titles)
|
||||
.l0MatchedKeywords(l0Hint != null ? l0Hint.matchedKeywords() : List.of())
|
||||
.l0Domains(l0Hint != null ? l0Hint.domains() : List.of())
|
||||
.l0Entities(l0Hint != null ? l0Hint.entities() : List.of())
|
||||
.l1TopScore(hasL1 ? (double) l1Results.get(0).getScore() : null)
|
||||
.l1TopSimilarity(hasL1 ? l1TopSimilarity : null)
|
||||
.l0Titles(query != null ? query.getL0Titles() : List.of())
|
||||
.l0MatchedKeywords(query != null ? query.getMatchedKeywords() : List.of())
|
||||
.l0Domains(query != null ? query.getDomainHints() : List.of())
|
||||
.l0Entities(query != null ? query.getEntities() : List.of())
|
||||
.l1TopScore(firstAttemptScore(trace))
|
||||
.l1TopSimilarity(firstAttemptSimilarity(trace))
|
||||
.l1Scores(l1Scores)
|
||||
.evidenceCandidateCount(result != null ? result.getEvidenceCandidateCount() : null)
|
||||
.evidenceBlockCount(result != null ? result.getEvidenceBlockCount() : null)
|
||||
.evidenceBlocks(result != null ? summarizeEvidenceBlocks(result.getEvidenceBlocks()) : List.of())
|
||||
.queryTransform(summarizeQueryTransform(query))
|
||||
.retrievalTrace(summarizeRetrievalTrace(trace))
|
||||
.contextPack(summarizeContextPack(result != null ? result.getContextPack() : null))
|
||||
.rerankTrace(summarizeRerankTrace(result != null ? result.getRerankTrace() : null))
|
||||
.fallbackReason(trace != null ? trace.getFallbackReason() : null)
|
||||
.build();
|
||||
}
|
||||
|
||||
@@ -331,5 +337,130 @@ public class ToolInvocationRecorder {
|
||||
}
|
||||
return summaries;
|
||||
}
|
||||
|
||||
private static Map<String, Object> summarizeQueryTransform(KnowledgeQuery query) {
|
||||
if (query == null) {
|
||||
return Map.of();
|
||||
}
|
||||
Map<String, Object> summary = new LinkedHashMap<>();
|
||||
summary.put("original_query", query.getOriginalQuery());
|
||||
summary.put("rewritten_query", query.getRewrittenQuery());
|
||||
summary.put("category_filter", query.getCategoryFilter());
|
||||
summary.put("domain_hints", query.getDomainHints());
|
||||
summary.put("matched_keywords", query.getMatchedKeywords());
|
||||
summary.put("entities", query.getEntities());
|
||||
summary.put("l0_titles", query.getL0Titles());
|
||||
summary.put("l0_match_count", query.getL0MatchCount());
|
||||
return summary;
|
||||
}
|
||||
|
||||
private static Map<String, Object> summarizeRetrievalTrace(RetrievalTrace trace) {
|
||||
if (trace == null) {
|
||||
return Map.of();
|
||||
}
|
||||
Map<String, Object> summary = new LinkedHashMap<>();
|
||||
summary.put("selected_attempt", trace.getSelectedAttempt());
|
||||
summary.put("fallback_reason", trace.getFallbackReason());
|
||||
summary.put("evidence_status", trace.getEvidenceStatus());
|
||||
summary.put("category_filter", trace.getCategoryFilter());
|
||||
List<Map<String, Object>> attempts = new ArrayList<>();
|
||||
if (trace.getAttempts() != null) {
|
||||
for (RetrievalTrace.Attempt attempt : trace.getAttempts()) {
|
||||
Map<String, Object> item = new LinkedHashMap<>();
|
||||
item.put("name", attempt.getName());
|
||||
item.put("category_filter", attempt.getCategoryFilter());
|
||||
item.put("candidate_count", attempt.getCandidateCount());
|
||||
item.put("usable", attempt.getUsable());
|
||||
item.put("duration_ms", attempt.getDurationMs());
|
||||
item.put("top_score", attempt.getTopScore());
|
||||
item.put("top_similarity", attempt.getTopSimilarity());
|
||||
item.put("error_message", attempt.getErrorMessage());
|
||||
attempts.add(item);
|
||||
}
|
||||
}
|
||||
summary.put("attempts", attempts);
|
||||
return summary;
|
||||
}
|
||||
|
||||
private static Map<String, Object> summarizeContextPack(ContextPack contextPack) {
|
||||
if (contextPack == null) {
|
||||
return Map.of();
|
||||
}
|
||||
Map<String, Object> summary = new LinkedHashMap<>();
|
||||
summary.put("strategy", contextPack.getStrategy());
|
||||
summary.put("char_budget", contextPack.getCharBudget());
|
||||
summary.put("used_chars", contextPack.getUsedChars());
|
||||
summary.put("included_sources", contextPack.getIncludedSources());
|
||||
summary.put("omitted_sources", contextPack.getOmittedSources());
|
||||
return summary;
|
||||
}
|
||||
|
||||
private static Map<String, Object> summarizeRerankTrace(RerankTrace trace) {
|
||||
if (trace == null || trace.getItems() == null || trace.getItems().isEmpty()) {
|
||||
return Map.of();
|
||||
}
|
||||
List<Map<String, Object>> items = new ArrayList<>();
|
||||
for (int i = 0; i < Math.min(5, trace.getItems().size()); i++) {
|
||||
RerankTrace.Item item = trace.getItems().get(i);
|
||||
Map<String, Object> summary = new LinkedHashMap<>();
|
||||
summary.put("final_rank", item.getFinalRank());
|
||||
summary.put("source", item.getSource());
|
||||
summary.put("base_score", item.getBaseScore());
|
||||
summary.put("final_score", item.getFinalScore());
|
||||
summary.put("boost_reasons", item.getBoostReasons());
|
||||
items.add(summary);
|
||||
}
|
||||
return Map.of("items", items);
|
||||
}
|
||||
|
||||
private static Integer totalCandidateCount(RetrievalTrace trace) {
|
||||
if (trace == null || trace.getAttempts() == null || trace.getAttempts().isEmpty()) {
|
||||
return null;
|
||||
}
|
||||
int total = 0;
|
||||
for (RetrievalTrace.Attempt attempt : trace.getAttempts()) {
|
||||
if (attempt.getCandidateCount() != null) {
|
||||
total += attempt.getCandidateCount();
|
||||
}
|
||||
}
|
||||
return total;
|
||||
}
|
||||
|
||||
private static Double firstAttemptScore(RetrievalTrace trace) {
|
||||
if (trace == null || trace.getAttempts() == null || trace.getAttempts().isEmpty()) {
|
||||
return null;
|
||||
}
|
||||
for (RetrievalTrace.Attempt attempt : trace.getAttempts()) {
|
||||
if (attempt.getTopScore() != null) {
|
||||
return attempt.getTopScore();
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
private static Double firstAttemptSimilarity(RetrievalTrace trace) {
|
||||
if (trace == null || trace.getAttempts() == null || trace.getAttempts().isEmpty()) {
|
||||
return null;
|
||||
}
|
||||
for (RetrievalTrace.Attempt attempt : trace.getAttempts()) {
|
||||
if (attempt.getTopSimilarity() != null) {
|
||||
return attempt.getTopSimilarity();
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
private static List<Double> collectAttemptScores(RetrievalTrace trace) {
|
||||
if (trace == null || trace.getAttempts() == null || trace.getAttempts().isEmpty()) {
|
||||
return List.of();
|
||||
}
|
||||
List<Double> scores = new ArrayList<>();
|
||||
for (RetrievalTrace.Attempt attempt : trace.getAttempts()) {
|
||||
if (attempt.getTopScore() != null) {
|
||||
scores.add(attempt.getTopScore());
|
||||
}
|
||||
}
|
||||
return scores;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -11,6 +11,7 @@ import lombok.Getter;
|
||||
import lombok.Setter;
|
||||
import com.superbiz.agent.constant.MilvusConstants;
|
||||
import com.superbiz.agent.dto.DocumentChunk;
|
||||
import com.superbiz.agent.dto.Frontmatter;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
import org.springframework.beans.factory.annotation.Autowired;
|
||||
@@ -176,6 +177,10 @@ public class VectorIndexService {
|
||||
* @throws Exception 索引失败时抛出异常
|
||||
*/
|
||||
public void indexDocumentChunks(String docId, List<DocumentChunk> chunks, String category) throws Exception {
|
||||
indexDocumentChunks(docId, chunks, category, null);
|
||||
}
|
||||
|
||||
public void indexDocumentChunks(String docId, List<DocumentChunk> chunks, String category, Frontmatter frontmatter) throws Exception {
|
||||
if (chunks == null || chunks.isEmpty()) {
|
||||
throw new IllegalArgumentException("文档分块列表为空");
|
||||
}
|
||||
@@ -194,7 +199,7 @@ public class VectorIndexService {
|
||||
List<Float> vector = embeddingService.generateEmbedding(buildEmbeddingText(chunk));
|
||||
|
||||
// 构建元数据(使用 docId 和 category)
|
||||
Map<String, Object> metadata = buildDocumentMetadata(docId, chunk, chunks.size(), category);
|
||||
Map<String, Object> metadata = buildDocumentMetadata(docId, chunk, chunks.size(), category, frontmatter);
|
||||
|
||||
// 插入到 Milvus
|
||||
insertToMilvus(chunk.getContent(), vector, metadata, chunk.getChunkIndex());
|
||||
@@ -253,30 +258,47 @@ public class VectorIndexService {
|
||||
/**
|
||||
* 构建文档元数据(用于上传文档)
|
||||
*/
|
||||
private Map<String, Object> buildDocumentMetadata(String docId, DocumentChunk chunk, int totalChunks, String category) {
|
||||
static Map<String, Object> buildDocumentMetadata(String docId, DocumentChunk chunk, int totalChunks, String category) {
|
||||
return buildDocumentMetadata(docId, chunk, totalChunks, category, null);
|
||||
}
|
||||
|
||||
static Map<String, Object> buildDocumentMetadata(String docId,
|
||||
DocumentChunk chunk,
|
||||
int totalChunks,
|
||||
String category,
|
||||
Frontmatter frontmatter) {
|
||||
Map<String, Object> metadata = new HashMap<>();
|
||||
|
||||
// 文档标识
|
||||
String source = firstNonBlank(frontmatter != null ? frontmatter.getSource() : null, "upload:" + docId);
|
||||
metadata.put("docId", docId);
|
||||
metadata.put("_source", "upload:" + docId); // 区分文件索引和上传文档
|
||||
metadata.put("_source", source); // 区分文件索引和上传文档
|
||||
metadata.put("source", source);
|
||||
|
||||
// 分片信息
|
||||
metadata.put("chunkIndex", chunk.getChunkIndex());
|
||||
metadata.put("totalChunks", totalChunks);
|
||||
|
||||
// 标题信息
|
||||
if (chunk.getTitle() != null && !chunk.getTitle().isEmpty()) {
|
||||
metadata.put("title", chunk.getTitle());
|
||||
String title = firstNonBlank(chunk.getTitle(), frontmatter != null ? frontmatter.getTitle() : null);
|
||||
if (title != null) {
|
||||
metadata.put("title", title);
|
||||
}
|
||||
|
||||
// 面包屑导航(完整标题层级路径)
|
||||
if (chunk.getBreadcrumb() != null && !chunk.getBreadcrumb().isEmpty()) {
|
||||
metadata.put("breadcrumb", chunk.getBreadcrumb());
|
||||
String breadcrumb = firstNonBlank(frontmatter != null ? frontmatter.getBreadcrumb() : null, chunk.getBreadcrumb());
|
||||
if (breadcrumb != null) {
|
||||
metadata.put("breadcrumb", breadcrumb);
|
||||
}
|
||||
|
||||
// 文档类别
|
||||
metadata.put("category", category != null && !category.isBlank() ? category : "upload");
|
||||
|
||||
String kbScope = trimToNull(frontmatter != null ? frontmatter.getKbScope() : null);
|
||||
if (kbScope != null) {
|
||||
metadata.put("kb_scope", kbScope);
|
||||
}
|
||||
|
||||
return metadata;
|
||||
}
|
||||
|
||||
@@ -304,6 +326,23 @@ public class VectorIndexService {
|
||||
return value == null ? "" : value.trim();
|
||||
}
|
||||
|
||||
private static String trimToNull(String value) {
|
||||
if (value == null || value.isBlank()) {
|
||||
return null;
|
||||
}
|
||||
return value.trim();
|
||||
}
|
||||
|
||||
private static String firstNonBlank(String... values) {
|
||||
for (String value : values) {
|
||||
String trimmed = trimToNull(value);
|
||||
if (trimmed != null) {
|
||||
return trimmed;
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
/**
|
||||
* 删除文件的旧数据(根据 metadata._source)
|
||||
*/
|
||||
|
||||
@@ -54,6 +54,9 @@ public class VectorSearchService {
|
||||
@Value("${retrieval.normalization.max-l2-distance:2.0}")
|
||||
private double maxL2Distance = 2.0;
|
||||
|
||||
@Value("${retrieval.kb-scope:}")
|
||||
private String kbScope = "";
|
||||
|
||||
public List<SearchResult> searchSimilarDocuments(String query, int topK) {
|
||||
return searchSimilarDocuments(query, topK, null);
|
||||
}
|
||||
@@ -62,7 +65,7 @@ public class VectorSearchService {
|
||||
String mode = vectorStoreMode == null ? "auto" : vectorStoreMode.trim().toLowerCase();
|
||||
return switch (mode) {
|
||||
case "sdk" -> searchSimilarDocumentsWithSdk(query, topK, category);
|
||||
case "spring-ai" -> searchSimilarDocumentsWithVectorStore(query, topK, category);
|
||||
case "spring", "spring-ai" -> searchSimilarDocumentsWithVectorStore(query, topK, category);
|
||||
case "auto" -> searchWithAutoFallback(query, topK, category);
|
||||
default -> {
|
||||
logger.warn("Unknown retrieval.vector-store.mode={}, using auto mode", vectorStoreMode);
|
||||
@@ -86,15 +89,16 @@ public class VectorSearchService {
|
||||
throw new IllegalStateException("Spring AI VectorStore bean is unavailable");
|
||||
}
|
||||
|
||||
logger.info("Starting Spring AI VectorStore search: query={}, topK={}, category={}", query, topK, category);
|
||||
logger.info("Starting Spring AI VectorStore search: query={}, topK={}, category={}, kbScope={}",
|
||||
query, topK, category, effectiveKbScope());
|
||||
SearchRequest.Builder builder = SearchRequest.builder()
|
||||
.query(query)
|
||||
.topK(topK)
|
||||
.similarityThresholdAll();
|
||||
if (category != null && !category.trim().isEmpty()) {
|
||||
String filterExpression = "category == '" + escapeFilterValue(category.trim()) + "'";
|
||||
String filterExpression = buildSpringAiFilterExpression(category);
|
||||
if (filterExpression != null) {
|
||||
builder.filterExpression(filterExpression);
|
||||
logger.info("Spring AI VectorStore category filter: {}", filterExpression);
|
||||
logger.info("Spring AI VectorStore metadata filter: {}", filterExpression);
|
||||
}
|
||||
|
||||
List<Document> documents = vectorStore.similaritySearch(builder.build());
|
||||
@@ -115,7 +119,8 @@ public class VectorSearchService {
|
||||
|
||||
List<SearchResult> searchSimilarDocumentsWithSdk(String query, int topK, String category) {
|
||||
try {
|
||||
logger.info("Starting Milvus SDK search: query={}, topK={}, category={}", query, topK, category);
|
||||
logger.info("Starting Milvus SDK search: query={}, topK={}, category={}, kbScope={}",
|
||||
query, topK, category, effectiveKbScope());
|
||||
|
||||
List<Float> queryVector = embeddingService.generateQueryVector(query);
|
||||
logger.debug("Query vector generated, dimension={}", queryVector.size());
|
||||
@@ -129,10 +134,10 @@ public class VectorSearchService {
|
||||
.withOutFields(List.of("id", "content", "metadata"))
|
||||
.withParams("{\"nprobe\":10}");
|
||||
|
||||
if (category != null && !category.trim().isEmpty()) {
|
||||
String expr = String.format("metadata[\"category\"] == \"%s\"", category);
|
||||
String expr = buildSdkFilterExpression(category);
|
||||
if (expr != null) {
|
||||
searchParamBuilder.withExpr(expr);
|
||||
logger.info("Milvus SDK category filter: {}", expr);
|
||||
logger.info("Milvus SDK metadata filter: {}", expr);
|
||||
}
|
||||
|
||||
R<SearchResults> searchResponse = milvusClient.search(searchParamBuilder.build());
|
||||
@@ -215,6 +220,47 @@ public class VectorSearchService {
|
||||
return value.replace("'", "\\'");
|
||||
}
|
||||
|
||||
String buildSpringAiFilterExpression(String category) {
|
||||
List<String> parts = new ArrayList<>();
|
||||
String categoryFilter = trimToNull(category);
|
||||
if (categoryFilter != null) {
|
||||
parts.add("category == '" + escapeFilterValue(categoryFilter) + "'");
|
||||
}
|
||||
String scopeFilter = effectiveKbScope();
|
||||
if (scopeFilter != null) {
|
||||
parts.add("kb_scope == '" + escapeFilterValue(scopeFilter) + "'");
|
||||
}
|
||||
return parts.isEmpty() ? null : String.join(" && ", parts);
|
||||
}
|
||||
|
||||
String buildSdkFilterExpression(String category) {
|
||||
List<String> parts = new ArrayList<>();
|
||||
String categoryFilter = trimToNull(category);
|
||||
if (categoryFilter != null) {
|
||||
parts.add("metadata[\"category\"] == \"" + escapeMilvusString(categoryFilter) + "\"");
|
||||
}
|
||||
String scopeFilter = effectiveKbScope();
|
||||
if (scopeFilter != null) {
|
||||
parts.add("metadata[\"kb_scope\"] == \"" + escapeMilvusString(scopeFilter) + "\"");
|
||||
}
|
||||
return parts.isEmpty() ? null : String.join(" && ", parts);
|
||||
}
|
||||
|
||||
private String effectiveKbScope() {
|
||||
return trimToNull(kbScope);
|
||||
}
|
||||
|
||||
private String trimToNull(String value) {
|
||||
if (value == null || value.isBlank()) {
|
||||
return null;
|
||||
}
|
||||
return value.trim();
|
||||
}
|
||||
|
||||
private String escapeMilvusString(String value) {
|
||||
return value.replace("\\", "\\\\").replace("\"", "\\\"");
|
||||
}
|
||||
|
||||
@Setter
|
||||
@Getter
|
||||
public static class SearchResult {
|
||||
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user