Compare commits
68
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
b22f2d22c8 | ||
|
|
63b62b28a2 | ||
|
|
d5902a0499 | ||
|
|
ed267d753d | ||
|
|
2658742119 | ||
|
|
72a3dbf8c5 | ||
|
|
674dd27a48 | ||
|
|
c7e2fc2ee2 | ||
|
|
1bfe1a17b4 | ||
|
|
9dd6823fe7 | ||
|
|
9376448804 | ||
|
|
f2bae0382c | ||
|
|
5c71f5fc79 | ||
|
|
b9ec07de57 | ||
|
|
5197712719 | ||
|
|
4a94c14feb | ||
|
|
9a2a44d1b5 | ||
|
|
79feed3314 | ||
|
|
98155ae1d8 | ||
|
|
2609c5a5ab | ||
|
|
bf5286c8f4 | ||
|
|
26e12a8d6b | ||
|
|
cbef3ddd3c | ||
|
|
69deb15330 | ||
|
|
4c7c53b024 | ||
|
|
ca5c61fabf | ||
|
|
23ee05c7c3 | ||
|
|
dc6cd32a67 | ||
|
|
246c99b954 | ||
|
|
f01866c1a2 | ||
|
|
6919092b83 | ||
|
|
5b827fe90e | ||
|
|
b0f288ae36 | ||
|
|
1ff7f09d25 | ||
|
|
fd89d84fc0 | ||
|
|
9050487307 | ||
|
|
4f5316d473 | ||
|
|
2a7164288f | ||
|
|
a1c896ebda | ||
|
|
e438df4355 | ||
|
|
e4f37cb9e6 | ||
|
|
354ffc1947 | ||
|
|
bb44140901 | ||
|
|
2a796da490 | ||
|
|
3ffa5cc366 | ||
|
|
e3f20b1f06 | ||
|
|
9b52afce07 | ||
|
|
a3abe3f7a2 | ||
|
|
0d9cce75f9 | ||
|
|
a74ccea5be | ||
|
|
a1876286fd | ||
|
|
934d8eee29 | ||
|
|
7c8758d7fa | ||
|
|
b3ea6e202d | ||
|
|
8890cd2806 | ||
|
|
f4f0c63325 | ||
|
|
f02a1389c8 | ||
|
|
91931363d4 | ||
|
|
4e3502a51b | ||
|
|
b01f133efb | ||
|
|
3ed48e38cd | ||
|
|
dec587959c | ||
|
|
f002571629 | ||
|
|
553d1d1faf | ||
|
|
c88b287f83 | ||
|
|
463d8b817b | ||
|
|
363767d3e7 | ||
|
|
c4d23c3bd8 |
@@ -0,0 +1,57 @@
|
||||
# Frontend Design — Complete Guidance
|
||||
|
||||
This document provides a comprehensive framework for creating visually distinctive, non-templated UI designs. Here's the full breakdown:
|
||||
|
||||
## Foundational Approach
|
||||
|
||||
Act as the design lead for a studio known for unique client identities — the client has already turned down template-like proposals. Every choice about palette, typography, and layout must be specific to the brief, including "one real aesthetic risk you can justify."
|
||||
|
||||
## Grounding in Subject Matter
|
||||
|
||||
If the brief is vague about the product or subject, pin it down yourself: name the subject, its audience, and the page's single job. Draw inspiration from "the subject's own world, its materials, instruments, artifacts, and vernacular." Use any known context about the human's preferences or past designs as hints.
|
||||
|
||||
## Design Principles
|
||||
|
||||
- **Hero as thesis**: Open with "the most characteristic thing in the subject's world" — avoid default choices like a big number with a small label and gradient accent unless truly optimal.
|
||||
- **Typography**: Pair display and body faces deliberately, not from your usual repertoire. Set a clear type scale with intentional weights, widths, and spacing. "Make the type treatment itself a memorable part of the design."
|
||||
- **Structure as information**: Numbering, eyebrows, dividers must encode something true about the content. Question whether numbered markers (01/02/03) actually make sense before using them — only appropriate for real sequences.
|
||||
- **Motion**: Consider where animation serves the subject. "An orchestrated moment usually lands harder than scattered effects." Sometimes less is better to avoid an AI-generated feel.
|
||||
- **Complexity**: Match execution to the vision — maximalist needs elaborate execution, minimal needs precision.
|
||||
- **Content**: Come up with copy if the brief lacks it. Poor copy makes a design feel as templated as poor layout.
|
||||
|
||||
## AI-Generated Design Traps
|
||||
|
||||
Three common AI-default looks to watch for: (1) warm cream background (~#F4F1EA) with serif display and terracotta accent; (2) near-black with bright acid-green or vermilion; (3) broadsheet layout with hairline rules, zero border-radius, and dense columns. "All three are legitimate for some briefs, but they are defaults rather than choices." Where the brief leaves an axis free, don't spend that freedom on a default.
|
||||
|
||||
## Two-Pass Process
|
||||
|
||||
**Pass 1 — Plan**: Create a compact token system:
|
||||
|
||||
1. **Color**: 4–6 named hex values
|
||||
2. **Type**: Characterful display face (used with restraint), complementary body face, utility face for captions/data
|
||||
3. **Layout**: One-sentence prose descriptions + ASCII wireframes
|
||||
4. **Signature**: The single unique element the page will be remembered by
|
||||
|
||||
Review the plan against the brief. If any part reads like what you'd produce for any similar page, revise it. Only then write code.
|
||||
|
||||
**Pass 2 — Build**: Follow the revised plan exactly. Watch for CSS selector specificity conflicts (e.g., `.section` and `.cta` fighting over padding/margins). Do most planning internally, only sharing ideas when confident.
|
||||
|
||||
## Restraint & Self-Critique
|
||||
|
||||
"Spend your boldness in one place" — let the signature element be the one memorable thing; keep everything else quiet. "Not taking a risk can be a risk itself!" Build responsively down to mobile, with visible keyboard focus and reduced motion respected. Critique as you build. Follow Chanel's advice: before finishing, remove one accessory. Jot notes about what you've tried to avoid repeating yourself.
|
||||
|
||||
## Writing in Design
|
||||
|
||||
Words exist to make the design understandable and usable — they're "design material, not decoration." Write from the end user's perspective, naming things by what people control and recognize, never by how the system is built.
|
||||
|
||||
- Use active voice as default
|
||||
- A control should say exactly what happens: "Save changes," not "Submit"
|
||||
- Maintain consistent vocabulary throughout flows (button says "Publish," toast says "Published")
|
||||
- Treat errors as guidance, not mood — explain what went wrong and how to fix it
|
||||
- Empty screens are invitations to act
|
||||
- Keep the register conversational: "plain verbs, sentence case, no filler"
|
||||
- Let each element do exactly one job — "a label labels, an example demonstrates"
|
||||
|
||||
## License
|
||||
|
||||
Apache License 2.0 — see LICENSE.txt
|
||||
@@ -0,0 +1,156 @@
|
||||
---
|
||||
name: openspec-apply-change
|
||||
description: Implement tasks from an OpenSpec change. Use when the user wants to start implementing, continue implementation, or work through tasks.
|
||||
license: MIT
|
||||
compatibility: Requires openspec CLI.
|
||||
metadata:
|
||||
author: openspec
|
||||
version: "1.0"
|
||||
generatedBy: "1.3.1"
|
||||
---
|
||||
|
||||
Implement tasks from an OpenSpec change.
|
||||
|
||||
**Input**: Optionally specify a change name. If omitted, check if it can be inferred from conversation context. If vague or ambiguous you MUST prompt for available changes.
|
||||
|
||||
**Steps**
|
||||
|
||||
1. **Select the change**
|
||||
|
||||
If a name is provided, use it. Otherwise:
|
||||
- Infer from conversation context if the user mentioned a change
|
||||
- Auto-select if only one active change exists
|
||||
- If ambiguous, run `openspec list --json` to get available changes and use the **AskUserQuestion tool** to let the user select
|
||||
|
||||
Always announce: "Using change: <name>" and how to override (e.g., `/opsx:apply <other>`).
|
||||
|
||||
2. **Check status to understand the schema**
|
||||
```bash
|
||||
openspec status --change "<name>" --json
|
||||
```
|
||||
Parse the JSON to understand:
|
||||
- `schemaName`: The workflow being used (e.g., "spec-driven")
|
||||
- Which artifact contains the tasks (typically "tasks" for spec-driven, check status for others)
|
||||
|
||||
3. **Get apply instructions**
|
||||
|
||||
```bash
|
||||
openspec instructions apply --change "<name>" --json
|
||||
```
|
||||
|
||||
This returns:
|
||||
- `contextFiles`: artifact ID -> array of concrete file paths (varies by schema - could be proposal/specs/design/tasks or spec/tests/implementation/docs)
|
||||
- Progress (total, complete, remaining)
|
||||
- Task list with status
|
||||
- Dynamic instruction based on current state
|
||||
|
||||
**Handle states:**
|
||||
- If `state: "blocked"` (missing artifacts): show message, suggest using openspec-continue-change
|
||||
- If `state: "all_done"`: congratulate, suggest archive
|
||||
- Otherwise: proceed to implementation
|
||||
|
||||
4. **Read context files**
|
||||
|
||||
Read every file path listed under `contextFiles` from the apply instructions output.
|
||||
The files depend on the schema being used:
|
||||
- **spec-driven**: proposal, specs, design, tasks
|
||||
- Other schemas: follow the contextFiles from CLI output
|
||||
|
||||
5. **Show current progress**
|
||||
|
||||
Display:
|
||||
- Schema being used
|
||||
- Progress: "N/M tasks complete"
|
||||
- Remaining tasks overview
|
||||
- Dynamic instruction from CLI
|
||||
|
||||
6. **Implement tasks (loop until done or blocked)**
|
||||
|
||||
For each pending task:
|
||||
- Show which task is being worked on
|
||||
- Make the code changes required
|
||||
- Keep changes minimal and focused
|
||||
- Mark task complete in the tasks file: `- [ ]` → `- [x]`
|
||||
- Continue to next task
|
||||
|
||||
**Pause if:**
|
||||
- Task is unclear → ask for clarification
|
||||
- Implementation reveals a design issue → suggest updating artifacts
|
||||
- Error or blocker encountered → report and wait for guidance
|
||||
- User interrupts
|
||||
|
||||
7. **On completion or pause, show status**
|
||||
|
||||
Display:
|
||||
- Tasks completed this session
|
||||
- Overall progress: "N/M tasks complete"
|
||||
- If all done: suggest archive
|
||||
- If paused: explain why and wait for guidance
|
||||
|
||||
**Output During Implementation**
|
||||
|
||||
```
|
||||
## Implementing: <change-name> (schema: <schema-name>)
|
||||
|
||||
Working on task 3/7: <task description>
|
||||
[...implementation happening...]
|
||||
✓ Task complete
|
||||
|
||||
Working on task 4/7: <task description>
|
||||
[...implementation happening...]
|
||||
✓ Task complete
|
||||
```
|
||||
|
||||
**Output On Completion**
|
||||
|
||||
```
|
||||
## Implementation Complete
|
||||
|
||||
**Change:** <change-name>
|
||||
**Schema:** <schema-name>
|
||||
**Progress:** 7/7 tasks complete ✓
|
||||
|
||||
### Completed This Session
|
||||
- [x] Task 1
|
||||
- [x] Task 2
|
||||
...
|
||||
|
||||
All tasks complete! Ready to archive this change.
|
||||
```
|
||||
|
||||
**Output On Pause (Issue Encountered)**
|
||||
|
||||
```
|
||||
## Implementation Paused
|
||||
|
||||
**Change:** <change-name>
|
||||
**Schema:** <schema-name>
|
||||
**Progress:** 4/7 tasks complete
|
||||
|
||||
### Issue Encountered
|
||||
<description of the issue>
|
||||
|
||||
**Options:**
|
||||
1. <option 1>
|
||||
2. <option 2>
|
||||
3. Other approach
|
||||
|
||||
What would you like to do?
|
||||
```
|
||||
|
||||
**Guardrails**
|
||||
- Keep going through tasks until done or blocked
|
||||
- Always read context files before starting (from the apply instructions output)
|
||||
- If task is ambiguous, pause and ask before implementing
|
||||
- If implementation reveals issues, pause and suggest artifact updates
|
||||
- Keep code changes minimal and scoped to each task
|
||||
- Update task checkbox immediately after completing each task
|
||||
- Pause on errors, blockers, or unclear requirements - don't guess
|
||||
- Use contextFiles from CLI output, don't assume specific file names
|
||||
|
||||
**Fluid Workflow Integration**
|
||||
|
||||
This skill supports the "actions on a change" model:
|
||||
|
||||
- **Can be invoked anytime**: Before all artifacts are done (if tasks exist), after partial implementation, interleaved with other actions
|
||||
- **Allows artifact updates**: If implementation reveals design issues, suggest updating artifacts - not phase-locked, work fluidly
|
||||
@@ -0,0 +1,114 @@
|
||||
---
|
||||
name: openspec-archive-change
|
||||
description: Archive a completed change in the experimental workflow. Use when the user wants to finalize and archive a change after implementation is complete.
|
||||
license: MIT
|
||||
compatibility: Requires openspec CLI.
|
||||
metadata:
|
||||
author: openspec
|
||||
version: "1.0"
|
||||
generatedBy: "1.3.1"
|
||||
---
|
||||
|
||||
Archive a completed change in the experimental workflow.
|
||||
|
||||
**Input**: Optionally specify a change name. If omitted, check if it can be inferred from conversation context. If vague or ambiguous you MUST prompt for available changes.
|
||||
|
||||
**Steps**
|
||||
|
||||
1. **If no change name provided, prompt for selection**
|
||||
|
||||
Run `openspec list --json` to get available changes. Use the **AskUserQuestion tool** to let the user select.
|
||||
|
||||
Show only active changes (not already archived).
|
||||
Include the schema used for each change if available.
|
||||
|
||||
**IMPORTANT**: Do NOT guess or auto-select a change. Always let the user choose.
|
||||
|
||||
2. **Check artifact completion status**
|
||||
|
||||
Run `openspec status --change "<name>" --json` to check artifact completion.
|
||||
|
||||
Parse the JSON to understand:
|
||||
- `schemaName`: The workflow being used
|
||||
- `artifacts`: List of artifacts with their status (`done` or other)
|
||||
|
||||
**If any artifacts are not `done`:**
|
||||
- Display warning listing incomplete artifacts
|
||||
- Use **AskUserQuestion tool** to confirm user wants to proceed
|
||||
- Proceed if user confirms
|
||||
|
||||
3. **Check task completion status**
|
||||
|
||||
Read the tasks file (typically `tasks.md`) to check for incomplete tasks.
|
||||
|
||||
Count tasks marked with `- [ ]` (incomplete) vs `- [x]` (complete).
|
||||
|
||||
**If incomplete tasks found:**
|
||||
- Display warning showing count of incomplete tasks
|
||||
- Use **AskUserQuestion tool** to confirm user wants to proceed
|
||||
- Proceed if user confirms
|
||||
|
||||
**If no tasks file exists:** Proceed without task-related warning.
|
||||
|
||||
4. **Assess delta spec sync state**
|
||||
|
||||
Check for delta specs at `openspec/changes/<name>/specs/`. If none exist, proceed without sync prompt.
|
||||
|
||||
**If delta specs exist:**
|
||||
- Compare each delta spec with its corresponding main spec at `openspec/specs/<capability>/spec.md`
|
||||
- Determine what changes would be applied (adds, modifications, removals, renames)
|
||||
- Show a combined summary before prompting
|
||||
|
||||
**Prompt options:**
|
||||
- If changes needed: "Sync now (recommended)", "Archive without syncing"
|
||||
- If already synced: "Archive now", "Sync anyway", "Cancel"
|
||||
|
||||
If user chooses sync, use Task tool (subagent_type: "general-purpose", prompt: "Use Skill tool to invoke openspec-sync-specs for change '<name>'. Delta spec analysis: <include the analyzed delta spec summary>"). Proceed to archive regardless of choice.
|
||||
|
||||
5. **Perform the archive**
|
||||
|
||||
Create the archive directory if it doesn't exist:
|
||||
```bash
|
||||
mkdir -p openspec/changes/archive
|
||||
```
|
||||
|
||||
Generate target name using current date: `YYYY-MM-DD-<change-name>`
|
||||
|
||||
**Check if target already exists:**
|
||||
- If yes: Fail with error, suggest renaming existing archive or using different date
|
||||
- If no: Move the change directory to archive
|
||||
|
||||
```bash
|
||||
mv openspec/changes/<name> openspec/changes/archive/YYYY-MM-DD-<name>
|
||||
```
|
||||
|
||||
6. **Display summary**
|
||||
|
||||
Show archive completion summary including:
|
||||
- Change name
|
||||
- Schema that was used
|
||||
- Archive location
|
||||
- Whether specs were synced (if applicable)
|
||||
- Note about any warnings (incomplete artifacts/tasks)
|
||||
|
||||
**Output On Success**
|
||||
|
||||
```
|
||||
## Archive Complete
|
||||
|
||||
**Change:** <change-name>
|
||||
**Schema:** <schema-name>
|
||||
**Archived to:** openspec/changes/archive/YYYY-MM-DD-<name>/
|
||||
**Specs:** ✓ Synced to main specs (or "No delta specs" or "Sync skipped")
|
||||
|
||||
All artifacts complete. All tasks complete.
|
||||
```
|
||||
|
||||
**Guardrails**
|
||||
- Always prompt for change selection if not provided
|
||||
- Use artifact graph (openspec status --json) for completion checking
|
||||
- Don't block archive on warnings - just inform and confirm
|
||||
- Preserve .openspec.yaml when moving to archive (it moves with the directory)
|
||||
- Show clear summary of what happened
|
||||
- If sync is requested, use openspec-sync-specs approach (agent-driven)
|
||||
- If delta specs exist, always run the sync assessment and show the combined summary before prompting
|
||||
@@ -0,0 +1,288 @@
|
||||
---
|
||||
name: openspec-explore
|
||||
description: Enter explore mode - a thinking partner for exploring ideas, investigating problems, and clarifying requirements. Use when the user wants to think through something before or during a change.
|
||||
license: MIT
|
||||
compatibility: Requires openspec CLI.
|
||||
metadata:
|
||||
author: openspec
|
||||
version: "1.0"
|
||||
generatedBy: "1.3.1"
|
||||
---
|
||||
|
||||
Enter explore mode. Think deeply. Visualize freely. Follow the conversation wherever it goes.
|
||||
|
||||
**IMPORTANT: Explore mode is for thinking, not implementing.** You may read files, search code, and investigate the codebase, but you must NEVER write code or implement features. If the user asks you to implement something, remind them to exit explore mode first and create a change proposal. You MAY create OpenSpec artifacts (proposals, designs, specs) if the user asks—that's capturing thinking, not implementing.
|
||||
|
||||
**This is a stance, not a workflow.** There are no fixed steps, no required sequence, no mandatory outputs. You're a thinking partner helping the user explore.
|
||||
|
||||
---
|
||||
|
||||
## The Stance
|
||||
|
||||
- **Curious, not prescriptive** - Ask questions that emerge naturally, don't follow a script
|
||||
- **Open threads, not interrogations** - Surface multiple interesting directions and let the user follow what resonates. Don't funnel them through a single path of questions.
|
||||
- **Visual** - Use ASCII diagrams liberally when they'd help clarify thinking
|
||||
- **Adaptive** - Follow interesting threads, pivot when new information emerges
|
||||
- **Patient** - Don't rush to conclusions, let the shape of the problem emerge
|
||||
- **Grounded** - Explore the actual codebase when relevant, don't just theorize
|
||||
|
||||
---
|
||||
|
||||
## What You Might Do
|
||||
|
||||
Depending on what the user brings, you might:
|
||||
|
||||
**Explore the problem space**
|
||||
- Ask clarifying questions that emerge from what they said
|
||||
- Challenge assumptions
|
||||
- Reframe the problem
|
||||
- Find analogies
|
||||
|
||||
**Investigate the codebase**
|
||||
- Map existing architecture relevant to the discussion
|
||||
- Find integration points
|
||||
- Identify patterns already in use
|
||||
- Surface hidden complexity
|
||||
|
||||
**Compare options**
|
||||
- Brainstorm multiple approaches
|
||||
- Build comparison tables
|
||||
- Sketch tradeoffs
|
||||
- Recommend a path (if asked)
|
||||
|
||||
**Visualize**
|
||||
```
|
||||
┌─────────────────────────────────────────┐
|
||||
│ Use ASCII diagrams liberally │
|
||||
├─────────────────────────────────────────┤
|
||||
│ │
|
||||
│ ┌────────┐ ┌────────┐ │
|
||||
│ │ State │────────▶│ State │ │
|
||||
│ │ A │ │ B │ │
|
||||
│ └────────┘ └────────┘ │
|
||||
│ │
|
||||
│ System diagrams, state machines, │
|
||||
│ data flows, architecture sketches, │
|
||||
│ dependency graphs, comparison tables │
|
||||
│ │
|
||||
└─────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
**Surface risks and unknowns**
|
||||
- Identify what could go wrong
|
||||
- Find gaps in understanding
|
||||
- Suggest spikes or investigations
|
||||
|
||||
---
|
||||
|
||||
## OpenSpec Awareness
|
||||
|
||||
You have full context of the OpenSpec system. Use it naturally, don't force it.
|
||||
|
||||
### Check for context
|
||||
|
||||
At the start, quickly check what exists:
|
||||
```bash
|
||||
openspec list --json
|
||||
```
|
||||
|
||||
This tells you:
|
||||
- If there are active changes
|
||||
- Their names, schemas, and status
|
||||
- What the user might be working on
|
||||
|
||||
### When no change exists
|
||||
|
||||
Think freely. When insights crystallize, you might offer:
|
||||
|
||||
- "This feels solid enough to start a change. Want me to create a proposal?"
|
||||
- Or keep exploring - no pressure to formalize
|
||||
|
||||
### When a change exists
|
||||
|
||||
If the user mentions a change or you detect one is relevant:
|
||||
|
||||
1. **Read existing artifacts for context**
|
||||
- `openspec/changes/<name>/proposal.md`
|
||||
- `openspec/changes/<name>/design.md`
|
||||
- `openspec/changes/<name>/tasks.md`
|
||||
- etc.
|
||||
|
||||
2. **Reference them naturally in conversation**
|
||||
- "Your design mentions using Redis, but we just realized SQLite fits better..."
|
||||
- "The proposal scopes this to premium users, but we're now thinking everyone..."
|
||||
|
||||
3. **Offer to capture when decisions are made**
|
||||
|
||||
| Insight Type | Where to Capture |
|
||||
|----------------------------|--------------------------------|
|
||||
| New requirement discovered | `specs/<capability>/spec.md` |
|
||||
| Requirement changed | `specs/<capability>/spec.md` |
|
||||
| Design decision made | `design.md` |
|
||||
| Scope changed | `proposal.md` |
|
||||
| New work identified | `tasks.md` |
|
||||
| Assumption invalidated | Relevant artifact |
|
||||
|
||||
Example offers:
|
||||
- "That's a design decision. Capture it in design.md?"
|
||||
- "This is a new requirement. Add it to specs?"
|
||||
- "This changes scope. Update the proposal?"
|
||||
|
||||
4. **The user decides** - Offer and move on. Don't pressure. Don't auto-capture.
|
||||
|
||||
---
|
||||
|
||||
## What You Don't Have To Do
|
||||
|
||||
- Follow a script
|
||||
- Ask the same questions every time
|
||||
- Produce a specific artifact
|
||||
- Reach a conclusion
|
||||
- Stay on topic if a tangent is valuable
|
||||
- Be brief (this is thinking time)
|
||||
|
||||
---
|
||||
|
||||
## Handling Different Entry Points
|
||||
|
||||
**User brings a vague idea:**
|
||||
```
|
||||
User: I'm thinking about adding real-time collaboration
|
||||
|
||||
You: Real-time collab is a big space. Let me think about this...
|
||||
|
||||
COLLABORATION SPECTRUM
|
||||
════════════════════════════════════════════
|
||||
|
||||
Awareness Coordination Sync
|
||||
│ │ │
|
||||
▼ ▼ ▼
|
||||
┌────────┐ ┌────────┐ ┌────────┐
|
||||
│Presence│ │Cursors │ │ CRDT │
|
||||
│ "3 │ │ Multi │ │Conflict│
|
||||
│online" │ │ select │ │ free │
|
||||
└────────┘ └────────┘ └────────┘
|
||||
│ │ │
|
||||
trivial moderate complex
|
||||
|
||||
Where's your head at?
|
||||
```
|
||||
|
||||
**User brings a specific problem:**
|
||||
```
|
||||
User: The auth system is a mess
|
||||
|
||||
You: [reads codebase]
|
||||
|
||||
┌─────────────────────────────────────────────┐
|
||||
│ CURRENT AUTH FLOW │
|
||||
└─────────────────────────────────────────────┘
|
||||
│
|
||||
┌─────────────┼─────────────┐
|
||||
▼ ▼ ▼
|
||||
┌─────────┐ ┌─────────┐ ┌─────────┐
|
||||
│ Google │ │ GitHub │ │ Email │
|
||||
│ OAuth │ │ OAuth │ │ Magic │
|
||||
└────┬────┘ └────┬────┘ └────┬────┘
|
||||
│ │ │
|
||||
└─────────────┼─────────────┘
|
||||
▼
|
||||
┌───────────┐
|
||||
│ Session │
|
||||
└─────┬─────┘
|
||||
│
|
||||
▼
|
||||
┌───────────┐
|
||||
│ Perms │
|
||||
└───────────┘
|
||||
|
||||
I see three tangles. Which one's burning?
|
||||
```
|
||||
|
||||
**User is stuck mid-implementation:**
|
||||
```
|
||||
User: /opsx:explore add-auth-system
|
||||
The OAuth integration is more complex than expected
|
||||
|
||||
You: [reads change artifacts]
|
||||
|
||||
You're on task 4: "Implement OAuth flow"
|
||||
|
||||
Let me trace what's involved...
|
||||
|
||||
[draws diagram, explores options, suggests paths]
|
||||
|
||||
Want to update the design to reflect this?
|
||||
Or add a spike task to investigate?
|
||||
```
|
||||
|
||||
**User wants to compare options:**
|
||||
```
|
||||
User: Should we use Postgres or SQLite?
|
||||
|
||||
You: Generic answer is boring. What's the context?
|
||||
|
||||
User: A CLI tool that tracks local dev environments
|
||||
|
||||
You: That changes everything.
|
||||
|
||||
┌─────────────────────────────────────────────────┐
|
||||
│ CLI TOOL DATA STORAGE │
|
||||
└─────────────────────────────────────────────────┘
|
||||
|
||||
Key constraints:
|
||||
• No daemon running
|
||||
• Must work offline
|
||||
• Single user
|
||||
|
||||
SQLite Postgres
|
||||
Deployment embedded ✓ needs server ✗
|
||||
Offline yes ✓ no ✗
|
||||
Single file yes ✓ no ✗
|
||||
|
||||
SQLite. Not even close.
|
||||
|
||||
Unless... is there a sync component?
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Ending Discovery
|
||||
|
||||
There's no required ending. Discovery might:
|
||||
|
||||
- **Flow into a proposal**: "Ready to start? I can create a change proposal."
|
||||
- **Result in artifact updates**: "Updated design.md with these decisions"
|
||||
- **Just provide clarity**: User has what they need, moves on
|
||||
- **Continue later**: "We can pick this up anytime"
|
||||
|
||||
When it feels like things are crystallizing, you might summarize:
|
||||
|
||||
```
|
||||
## What We Figured Out
|
||||
|
||||
**The problem**: [crystallized understanding]
|
||||
|
||||
**The approach**: [if one emerged]
|
||||
|
||||
**Open questions**: [if any remain]
|
||||
|
||||
**Next steps** (if ready):
|
||||
- Create a change proposal
|
||||
- Keep exploring: just keep talking
|
||||
```
|
||||
|
||||
But this summary is optional. Sometimes the thinking IS the value.
|
||||
|
||||
---
|
||||
|
||||
## Guardrails
|
||||
|
||||
- **Don't implement** - Never write code or implement features. Creating OpenSpec artifacts is fine, writing application code is not.
|
||||
- **Don't fake understanding** - If something is unclear, dig deeper
|
||||
- **Don't rush** - Discovery is thinking time, not task time
|
||||
- **Don't force structure** - Let patterns emerge naturally
|
||||
- **Don't auto-capture** - Offer to save insights, don't just do it
|
||||
- **Do visualize** - A good diagram is worth many paragraphs
|
||||
- **Do explore the codebase** - Ground discussions in reality
|
||||
- **Do question assumptions** - Including the user's and your own
|
||||
@@ -0,0 +1,110 @@
|
||||
---
|
||||
name: openspec-propose
|
||||
description: Propose a new change with all artifacts generated in one step. Use when the user wants to quickly describe what they want to build and get a complete proposal with design, specs, and tasks ready for implementation.
|
||||
license: MIT
|
||||
compatibility: Requires openspec CLI.
|
||||
metadata:
|
||||
author: openspec
|
||||
version: "1.0"
|
||||
generatedBy: "1.3.1"
|
||||
---
|
||||
|
||||
Propose a new change - create the change and generate all artifacts in one step.
|
||||
|
||||
I'll create a change with artifacts:
|
||||
- proposal.md (what & why)
|
||||
- design.md (how)
|
||||
- tasks.md (implementation steps)
|
||||
|
||||
When ready to implement, run /opsx:apply
|
||||
|
||||
---
|
||||
|
||||
**Input**: The user's request should include a change name (kebab-case) OR a description of what they want to build.
|
||||
|
||||
**Steps**
|
||||
|
||||
1. **If no clear input provided, ask what they want to build**
|
||||
|
||||
Use the **AskUserQuestion tool** (open-ended, no preset options) to ask:
|
||||
> "What change do you want to work on? Describe what you want to build or fix."
|
||||
|
||||
From their description, derive a kebab-case name (e.g., "add user authentication" → `add-user-auth`).
|
||||
|
||||
**IMPORTANT**: Do NOT proceed without understanding what the user wants to build.
|
||||
|
||||
2. **Create the change directory**
|
||||
```bash
|
||||
openspec new change "<name>"
|
||||
```
|
||||
This creates a scaffolded change at `openspec/changes/<name>/` with `.openspec.yaml`.
|
||||
|
||||
3. **Get the artifact build order**
|
||||
```bash
|
||||
openspec status --change "<name>" --json
|
||||
```
|
||||
Parse the JSON to get:
|
||||
- `applyRequires`: array of artifact IDs needed before implementation (e.g., `["tasks"]`)
|
||||
- `artifacts`: list of all artifacts with their status and dependencies
|
||||
|
||||
4. **Create artifacts in sequence until apply-ready**
|
||||
|
||||
Use the **TodoWrite tool** to track progress through the artifacts.
|
||||
|
||||
Loop through artifacts in dependency order (artifacts with no pending dependencies first):
|
||||
|
||||
a. **For each artifact that is `ready` (dependencies satisfied)**:
|
||||
- Get instructions:
|
||||
```bash
|
||||
openspec instructions <artifact-id> --change "<name>" --json
|
||||
```
|
||||
- The instructions JSON includes:
|
||||
- `context`: Project background (constraints for you - do NOT include in output)
|
||||
- `rules`: Artifact-specific rules (constraints for you - do NOT include in output)
|
||||
- `template`: The structure to use for your output file
|
||||
- `instruction`: Schema-specific guidance for this artifact type
|
||||
- `outputPath`: Where to write the artifact
|
||||
- `dependencies`: Completed artifacts to read for context
|
||||
- Read any completed dependency files for context
|
||||
- Create the artifact file using `template` as the structure
|
||||
- Apply `context` and `rules` as constraints - but do NOT copy them into the file
|
||||
- Show brief progress: "Created <artifact-id>"
|
||||
|
||||
b. **Continue until all `applyRequires` artifacts are complete**
|
||||
- After creating each artifact, re-run `openspec status --change "<name>" --json`
|
||||
- Check if every artifact ID in `applyRequires` has `status: "done"` in the artifacts array
|
||||
- Stop when all `applyRequires` artifacts are done
|
||||
|
||||
c. **If an artifact requires user input** (unclear context):
|
||||
- Use **AskUserQuestion tool** to clarify
|
||||
- Then continue with creation
|
||||
|
||||
5. **Show final status**
|
||||
```bash
|
||||
openspec status --change "<name>"
|
||||
```
|
||||
|
||||
**Output**
|
||||
|
||||
After completing all artifacts, summarize:
|
||||
- Change name and location
|
||||
- List of artifacts created with brief descriptions
|
||||
- What's ready: "All artifacts created! Ready for implementation."
|
||||
- Prompt: "Run `/opsx:apply` or ask me to implement to start working on the tasks."
|
||||
|
||||
**Artifact Creation Guidelines**
|
||||
|
||||
- Follow the `instruction` field from `openspec instructions` for each artifact type
|
||||
- The schema defines what each artifact should contain - follow it
|
||||
- Read dependency artifacts for context before creating new ones
|
||||
- Use `template` as the structure for your output file - fill in its sections
|
||||
- **IMPORTANT**: `context` and `rules` are constraints for YOU, not content for the file
|
||||
- Do NOT copy `<context>`, `<rules>`, `<project_context>` blocks into the artifact
|
||||
- These guide what you write, but should never appear in the output
|
||||
|
||||
**Guardrails**
|
||||
- Create ALL artifacts needed for implementation (as defined by schema's `apply.requires`)
|
||||
- Always read dependency artifacts before creating a new one
|
||||
- If context is critically unclear, ask the user - but prefer making reasonable decisions to keep momentum
|
||||
- If a change with that name already exists, ask if user wants to continue it or create a new one
|
||||
- Verify each artifact file exists after writing before proceeding to next
|
||||
@@ -0,0 +1,233 @@
|
||||
# AI Ops Prompt 配置化 & LookupKnowledgeTool 集成
|
||||
|
||||
**日期**: 2026-06-24
|
||||
**类型**: 功能增强 + 架构优化
|
||||
**影响范围**: AI Ops 服务
|
||||
|
||||
---
|
||||
|
||||
## 一、变更背景
|
||||
|
||||
### 1.1 问题
|
||||
|
||||
- **硬编码 Prompt**:Planner、Executor、Supervisor 的系统提示词硬编码在 `AiOpsService.java` 中,难以维护和版本控制
|
||||
- **缺少知识库精确检索**:现有 `InternalDocsTools` 只支持 L1 语义检索(200-500ms),对于错误码、配置项等精确关键词查询效率较低
|
||||
|
||||
### 1.2 解决方案
|
||||
|
||||
1. **Prompt 配置化**:将所有 Agent 的 Prompt 抽取到 `prompts/ai-ops-prompts.yml` 配置文件
|
||||
2. **集成 L0+L1 混合检索**:引入 `LookupKnowledgeTool`,支持精确关键词匹配(< 10ms)+ 语义检索补充
|
||||
|
||||
---
|
||||
|
||||
## 二、架构变更
|
||||
|
||||
### 2.1 Prompt 配置化架构
|
||||
|
||||
```
|
||||
AiOpsService
|
||||
↓ 注入
|
||||
AiOpsPromptProperties (配置类)
|
||||
↓ @PostConstruct 加载
|
||||
ClassPathResource 读取 Markdown 文件
|
||||
↓ 读取
|
||||
prompts/
|
||||
├── planner-prompt.md
|
||||
├── executor-prompt.md
|
||||
└── supervisor-prompt.md
|
||||
```
|
||||
|
||||
**优点**:
|
||||
- 易于维护:Prompt 修改不需要重新编译
|
||||
- 格式友好:Markdown 格式支持代码块、表格,无 YAML 转义问题
|
||||
- 版本控制:配置文件独立管理
|
||||
- 易于扩展:后续可按环境区分(dev/prod)
|
||||
|
||||
### 2.2 工具层增强
|
||||
|
||||
```
|
||||
原有工具:
|
||||
- queryInternalDocs (纯 L1 语义检索,200-500ms)
|
||||
|
||||
新增工具:
|
||||
- lookup_knowledge (L0 精确匹配 + L1 补充,< 10ms 高置信度)
|
||||
```
|
||||
|
||||
**使用策略**:
|
||||
- 精确关键词(错误码、配置项)→ `lookup_knowledge`,未找到时降级到 `queryInternalDocs`
|
||||
- 模糊概念、故障流程 → 直接使用 `queryInternalDocs`
|
||||
|
||||
---
|
||||
|
||||
## 三、核心改动
|
||||
|
||||
### 3.1 新增文件
|
||||
|
||||
#### `AiOpsPromptProperties.java`
|
||||
```java
|
||||
@Configuration
|
||||
public class AiOpsPromptProperties {
|
||||
private String planner;
|
||||
private String executor;
|
||||
private String supervisor;
|
||||
|
||||
@PostConstruct
|
||||
public void loadPrompts() {
|
||||
planner = loadPromptFromFile("prompts/planner-prompt.md");
|
||||
executor = loadPromptFromFile("prompts/executor-prompt.md");
|
||||
supervisor = loadPromptFromFile("prompts/supervisor-prompt.md");
|
||||
}
|
||||
|
||||
private String loadPromptFromFile(String path) throws IOException {
|
||||
ClassPathResource resource = new ClassPathResource(path);
|
||||
return new String(resource.getInputStream().readAllBytes(), StandardCharsets.UTF_8);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### `prompts/*.md`
|
||||
三个独立的 Markdown 文件,包含 Agent 的完整系统提示词:
|
||||
- `planner-prompt.md` - Planner Agent 系统提示词
|
||||
- `executor-prompt.md` - Executor Agent 系统提示词(含工具选择指南)
|
||||
- `supervisor-prompt.md` - Supervisor Agent 系统提示词
|
||||
|
||||
### 3.2 修改文件
|
||||
|
||||
#### `AiOpsService.java`
|
||||
|
||||
**注入新组件**:
|
||||
```java
|
||||
@Autowired
|
||||
private LookupKnowledgeTool lookupKnowledgeTool;
|
||||
|
||||
@Autowired
|
||||
private AiOpsPromptProperties promptProperties;
|
||||
```
|
||||
|
||||
**使用配置化 Prompt**:
|
||||
```java
|
||||
// 原来
|
||||
.systemPrompt(buildPlannerPrompt())
|
||||
|
||||
// 改为
|
||||
.systemPrompt(promptProperties.getPlanner())
|
||||
```
|
||||
|
||||
**添加工具到工具数组**:
|
||||
```java
|
||||
return new Object[]{
|
||||
dateTimeTools,
|
||||
internalDocsTools,
|
||||
queryMetricsTools,
|
||||
lookupKnowledgeTool // 新增
|
||||
};
|
||||
```
|
||||
|
||||
**删除方法**:
|
||||
- `buildPlannerPrompt()`
|
||||
- `buildExecutorPrompt()`
|
||||
- `buildSupervisorSystemPrompt()`
|
||||
|
||||
---
|
||||
|
||||
## 四、Executor Prompt 变更详情
|
||||
|
||||
### 4.1 新增工具选择指南
|
||||
|
||||
```yaml
|
||||
- 根据查询内容选择合适的工具:
|
||||
* 精确关键词(错误码、配置项名称)→ 优先使用 lookup_knowledge,未找到时降级到 queryInternalDocs
|
||||
* 模糊概念、故障流程 → 直接使用 queryInternalDocs
|
||||
* 告警数据 → queryPrometheusAlerts
|
||||
* 日志数据 → queryLogs
|
||||
```
|
||||
|
||||
### 4.2 降级策略
|
||||
|
||||
关键改进:明确了 `lookup_knowledge` 未找到时的降级策略。
|
||||
|
||||
**流程**:
|
||||
```
|
||||
1. Planner: "查询 ERR_TIMEOUT 定义"
|
||||
2. Executor: 调用 lookup_knowledge("ERR_TIMEOUT")
|
||||
3a. 如果 found=true, confidence=high → 使用 primary.content
|
||||
3b. 如果 found=false → 自动降级到 queryInternalDocs("ERR_TIMEOUT 超时错误")
|
||||
4. 返回 feedback 给 Planner
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 五、兼容性说明
|
||||
|
||||
### 5.1 向后兼容
|
||||
|
||||
✅ **完全兼容**:
|
||||
- 现有工具调用逻辑不变
|
||||
- 3-Agent 协同模式不变
|
||||
- Planner/Executor/Supervisor 的职责边界不变
|
||||
|
||||
### 5.2 新增依赖
|
||||
|
||||
- `LookupKnowledgeTool` 依赖 `KnowledgeIndexService` 和 `VectorSearchService`
|
||||
- 需要 `knowledge_base/` 目录存在(已在 `application.yml` 中配置)
|
||||
|
||||
---
|
||||
|
||||
## 六、验证清单
|
||||
|
||||
### 6.1 编译验证
|
||||
|
||||
```bash
|
||||
mvn clean compile -DskipTests
|
||||
```
|
||||
|
||||
✅ **结果**: BUILD SUCCESS
|
||||
|
||||
### 6.2 运行时验证(待完成)
|
||||
|
||||
- [ ] 启动应用,验证 Prompt 配置加载成功
|
||||
- [ ] 触发 AI Ops 流程,验证 `lookup_knowledge` 工具可调用
|
||||
- [ ] 测试精确关键词查询(如 "ERR_TIMEOUT")
|
||||
- [ ] 测试降级策略(查询不存在的关键词)
|
||||
|
||||
---
|
||||
|
||||
## 七、后续工作
|
||||
|
||||
### 7.1 知识库内容准备
|
||||
|
||||
当前 `knowledge_base/` 目录需要补充文档:
|
||||
- 错误码定义(支付网关、订单系统等)
|
||||
- 配置最佳实践(Redis、HikariCP、Flyway 等)
|
||||
- 故障排查流程
|
||||
|
||||
**文档格式示例**:
|
||||
```markdown
|
||||
---
|
||||
title: 支付网关错误码定义
|
||||
keywords: [ERR_TIMEOUT, 超时, 支付网关]
|
||||
summary: 记录了支付网关所有核心错误码的含义及排查方向
|
||||
category: api
|
||||
---
|
||||
|
||||
# 支付网关错误码定义
|
||||
|
||||
## ERR_TIMEOUT
|
||||
...
|
||||
```
|
||||
|
||||
### 7.2 Prompt 优化
|
||||
|
||||
基于实际运行反馈,持续优化 `prompts/ai-ops-prompts.yml` 中的提示词。
|
||||
|
||||
### 7.3 可观测性增强
|
||||
|
||||
- 监控 `lookup_knowledge` 的调用频率和命中率
|
||||
- 记录降级场景(L0 未找到 → L1 补充)
|
||||
|
||||
---
|
||||
|
||||
## 八、参考文档
|
||||
|
||||
- [知识库检索架构说明](../mvp/architecture/knowledge-retrieval-architecture.md)
|
||||
- [AI Ops 核心设计 Essence 报告](../docs/learning/01-AI-Ops-核心设计-Essence报告.md)
|
||||
@@ -0,0 +1,100 @@
|
||||
# Prompt 配置化改进总结
|
||||
|
||||
**日期**: 2026-06-24
|
||||
**改进**: 从 YAML 配置改为 Markdown 文件
|
||||
|
||||
---
|
||||
|
||||
## 改进原因
|
||||
|
||||
YAML 格式存在以下问题:
|
||||
1. **多行字符串缩进敏感**:容易出现格式错误
|
||||
2. **转义字符复杂**:代码块、表格需要转义处理
|
||||
3. **可读性差**:长文本在 YAML 中难以阅读和维护
|
||||
|
||||
Markdown 格式优势:
|
||||
- ✅ 原生支持代码块、表格、列表
|
||||
- ✅ 无需转义,所见即所得
|
||||
- ✅ 版本控制 diff 更清晰
|
||||
- ✅ 编辑器语法高亮支持好
|
||||
|
||||
---
|
||||
|
||||
## 最终方案
|
||||
|
||||
### 文件结构
|
||||
```
|
||||
src/main/resources/prompts/
|
||||
├── planner-prompt.md # Planner Agent 系统提示词
|
||||
├── executor-prompt.md # Executor Agent 系统提示词
|
||||
└── supervisor-prompt.md # Supervisor Agent 系统提示词
|
||||
```
|
||||
|
||||
### 加载方式
|
||||
```java
|
||||
@Configuration
|
||||
public class AiOpsPromptProperties {
|
||||
|
||||
@PostConstruct
|
||||
public void loadPrompts() {
|
||||
planner = loadPromptFromFile("prompts/planner-prompt.md");
|
||||
executor = loadPromptFromFile("prompts/executor-prompt.md");
|
||||
supervisor = loadPromptFromFile("prompts/supervisor-prompt.md");
|
||||
}
|
||||
|
||||
private String loadPromptFromFile(String path) throws IOException {
|
||||
ClassPathResource resource = new ClassPathResource(path);
|
||||
return new String(resource.getInputStream().readAllBytes(), StandardCharsets.UTF_8);
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 使用方式
|
||||
```java
|
||||
@Autowired
|
||||
private AiOpsPromptProperties promptProperties;
|
||||
|
||||
// 直接使用
|
||||
.systemPrompt(promptProperties.getPlanner())
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 编译验证
|
||||
|
||||
```bash
|
||||
mvn clean compile -DskipTests
|
||||
```
|
||||
|
||||
✅ **结果**: BUILD SUCCESS
|
||||
|
||||
---
|
||||
|
||||
## 完整改动清单
|
||||
|
||||
| 文件 | 改动 |
|
||||
|------|------|
|
||||
| `AiOpsService.java` | 注入 `LookupKnowledgeTool` + `AiOpsPromptProperties` |
|
||||
| `AiOpsPromptProperties.java` | 从 Markdown 文件加载 Prompt(使用 `@PostConstruct`)|
|
||||
| `prompts/planner-prompt.md` | 新增:Planner 系统提示词 |
|
||||
| `prompts/executor-prompt.md` | 新增:Executor 系统提示词(含工具选择指南)|
|
||||
| `prompts/supervisor-prompt.md` | 新增:Supervisor 系统提示词 |
|
||||
| ~~`YamlPropertySourceFactory.java`~~ | 已删除(不再需要)|
|
||||
| ~~`prompts/ai-ops-prompts.yml`~~ | 已删除(改用 Markdown)|
|
||||
|
||||
---
|
||||
|
||||
## Executor Prompt 关键改进
|
||||
|
||||
新增工具选择指南:
|
||||
```markdown
|
||||
- 根据查询内容选择合适的工具:
|
||||
* 精确关键词(错误码、配置项名称)→ 优先使用 lookup_knowledge,未找到时降级到 queryInternalDocs
|
||||
* 模糊概念、故障流程 → 直接使用 queryInternalDocs
|
||||
* 告警数据 → queryPrometheusAlerts
|
||||
* 日志数据 → queryLogs
|
||||
```
|
||||
|
||||
降级策略:
|
||||
- `lookup_knowledge` 未找到 → 自动降级到 `queryInternalDocs`
|
||||
- 确保查询不会因为知识库缺少内容而失败
|
||||
@@ -0,0 +1,469 @@
|
||||
# 知识库初始化 API 使用文档
|
||||
|
||||
## 概述
|
||||
|
||||
提供了知识库批量初始化接口,用于将 `knowledge_base` 目录下的所有 Markdown 文档导入到数据库和向量索引(L0 + L1)。
|
||||
|
||||
**功能特点**:
|
||||
1. ✅ **批量扫描**:递归扫描 knowledge_base 目录下所有 .md 文件
|
||||
2. ✅ **自动去重**:基于文件路径检查,避免重复导入
|
||||
3. ✅ **数据入库**:保存文档元数据到 MySQL
|
||||
4. ✅ **L0 索引**:自动加入内存精确匹配索引
|
||||
5. ✅ **L1 索引**:文档分块并上传到 Milvus 向量数据库
|
||||
|
||||
---
|
||||
|
||||
## API 接口
|
||||
|
||||
### 1. 初始化知识库
|
||||
|
||||
**端点**:
|
||||
```
|
||||
POST /api/knowledge/init?force=false
|
||||
```
|
||||
|
||||
**参数**:
|
||||
- `force`(可选):是否强制重新导入,跳过去重检查
|
||||
- `false`(默认):跳过已存在的文档
|
||||
- `true`:强制重新导入所有文档
|
||||
|
||||
**请求示例**:
|
||||
```bash
|
||||
# 首次导入(去重模式)
|
||||
curl -X POST http://localhost:9900/api/knowledge/init
|
||||
|
||||
# 强制重新导入
|
||||
curl -X POST http://localhost:9900/api/knowledge/init?force=true
|
||||
```
|
||||
|
||||
**响应示例**:
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"message": "知识库初始化完成",
|
||||
"scanned": 6,
|
||||
"skipped": 0,
|
||||
"inserted": 6,
|
||||
"failed": 0,
|
||||
"details": {
|
||||
"api/payment-errors.md": "导入成功(L0+L1)",
|
||||
"domain/spring-ai-tool-best-practices.md": "导入成功(L0+L1)",
|
||||
"infrastructure/flyway-best-practices.md": "导入成功(L0+L1)",
|
||||
"infrastructure/mysql-connection-pool.md": "导入成功(L0+L1)",
|
||||
"infrastructure/redis-config.md": "导入成功(L0+L1)",
|
||||
"troubleshooting/fault-diagnosis-process.md": "导入成功(L0+L1)"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**字段说明**:
|
||||
- `scanned`:扫描到的文件总数
|
||||
- `skipped`:跳过的文件数量(已存在)
|
||||
- `inserted`:成功导入的文件数量
|
||||
- `failed`:失败的文件数量
|
||||
- `details`:每个文件的处理结果详情
|
||||
|
||||
---
|
||||
|
||||
### 2. 查询知识库统计
|
||||
|
||||
**端点**:
|
||||
```
|
||||
GET /api/knowledge/stats
|
||||
```
|
||||
|
||||
**请求示例**:
|
||||
```bash
|
||||
curl http://localhost:9900/api/knowledge/stats
|
||||
```
|
||||
|
||||
**响应示例**:
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"totalDocuments": 6,
|
||||
"totalVectors": 48,
|
||||
"categories": {
|
||||
"api": 1,
|
||||
"domain": 1,
|
||||
"infrastructure": 3,
|
||||
"troubleshooting": 1
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**字段说明**:
|
||||
- `totalDocuments`:数据库中的文档总数
|
||||
- `totalVectors`:Milvus 中的向量总数(chunk 数量)
|
||||
- `categories`:按分类统计的文档数量
|
||||
|
||||
---
|
||||
|
||||
## 使用场景
|
||||
|
||||
### 场景 1:项目启动时初始化
|
||||
|
||||
```bash
|
||||
# 1. 启动应用
|
||||
mvn spring-boot:run
|
||||
|
||||
# 2. 等待应用启动完成(约 10 秒)
|
||||
|
||||
# 3. 调用初始化接口
|
||||
curl -X POST http://localhost:9900/api/knowledge/init
|
||||
|
||||
# 4. 查看结果
|
||||
# 日志输出:知识库初始化完成: 扫描=6, 跳过=0, 新增=6, 失败=0
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 场景 2:添加新文档后重新初始化
|
||||
|
||||
```bash
|
||||
# 1. 添加新文档到 knowledge_base 目录
|
||||
echo "---
|
||||
title: 新文档
|
||||
keywords: [测试, test]
|
||||
summary: 这是一个测试文档
|
||||
category: test
|
||||
---
|
||||
|
||||
# 新文档内容
|
||||
" > knowledge_base/test/new-doc.md
|
||||
|
||||
# 2. 调用初始化接口(去重模式)
|
||||
curl -X POST http://localhost:9900/api/knowledge/init
|
||||
|
||||
# 3. 查看结果
|
||||
# 只会导入新文档,跳过已存在的 6 个文档
|
||||
# 响应: scanned=7, skipped=6, inserted=1, failed=0
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 场景 3:强制重新导入所有文档
|
||||
|
||||
```bash
|
||||
# 适用场景:
|
||||
# - 数据库被清空,需要重新导入
|
||||
# - 文档内容有更新,需要刷新
|
||||
# - 索引损坏,需要重建
|
||||
|
||||
curl -X POST http://localhost:9900/api/knowledge/init?force=true
|
||||
|
||||
# 响应: scanned=6, skipped=0, inserted=6, failed=0
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 去重机制
|
||||
|
||||
### 去重依据
|
||||
- **文件路径**:相对于 `knowledge_base` 目录的相对路径
|
||||
- 示例:`api/payment-errors.md`
|
||||
|
||||
### 去重逻辑
|
||||
```
|
||||
if (!force && existingFilePaths.contains(relativePath)) {
|
||||
跳过该文档
|
||||
} else {
|
||||
导入该文档
|
||||
}
|
||||
```
|
||||
|
||||
### 注意事项
|
||||
1. **文件移动会被视为新文档**:
|
||||
```bash
|
||||
# 移动前:api/payment-errors.md
|
||||
# 移动后:errors/payment-errors.md
|
||||
# 结果:会被当作两个不同的文档
|
||||
```
|
||||
|
||||
2. **文件重命名会被视为新文档**:
|
||||
```bash
|
||||
# 重命名前:payment-errors.md
|
||||
# 重命名后:payment-error-codes.md
|
||||
# 结果:会被当作两个不同的文档
|
||||
```
|
||||
|
||||
3. **内容更新不触发重新导入**(非 force 模式):
|
||||
```bash
|
||||
# 修改文件内容后调用 init(非 force)
|
||||
# 结果:跳过该文档,数据库中仍是旧内容
|
||||
# 解决:使用 force=true 强制重新导入
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 数据存储
|
||||
|
||||
### 完整的数据流
|
||||
|
||||
```
|
||||
knowledge_base/*.md
|
||||
↓ 1. 扫描
|
||||
KnowledgeBaseInitService
|
||||
↓ 2. 解析 frontmatter
|
||||
Frontmatter (title, keywords, summary)
|
||||
↓ 3. 保存到数据库
|
||||
MySQL (api_document)
|
||||
↓ 4. 提取正文 & 分块
|
||||
DocumentChunkService
|
||||
↓ 5. 生成向量
|
||||
VectorEmbeddingService
|
||||
↓ 6. 索引到 Milvus
|
||||
Milvus (L1 向量索引)
|
||||
↓ 7. 加入内存索引
|
||||
KnowledgeIndexService (L0)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 数据库表结构(api_document)
|
||||
|
||||
| 字段 | 类型 | 说明 | 示例 |
|
||||
|------|------|------|------|
|
||||
| `id` | BIGINT | 主键 | 1 |
|
||||
| `doc_id` | VARCHAR(64) | 文档唯一标识 | uuid |
|
||||
| `file_name` | VARCHAR(256) | 文件名 | payment-errors.md |
|
||||
| `file_path` | VARCHAR(512) | 相对路径 | api/payment-errors.md |
|
||||
| `api_name` | VARCHAR(128) | 文档标题 | 支付网关错误码定义 |
|
||||
| `status` | VARCHAR(16) | 状态 | INDEXED / FAILED |
|
||||
| `chunk_count` | INT | 分块数量 | 8 |
|
||||
| `error_message` | TEXT | 错误信息 | null |
|
||||
| `metadata` | TEXT | Frontmatter JSON | {"title":"...","keywords":[...]} |
|
||||
| `file_size` | BIGINT | 文件大小(字节) | 2048 |
|
||||
| `indexed_at` | DATETIME | 索引时间 | 2026-06-25 10:00:00 |
|
||||
|
||||
### metadata JSON 结构
|
||||
|
||||
```json
|
||||
{
|
||||
"title": "支付网关错误码定义",
|
||||
"summary": "记录了支付网关所有核心错误码的含义及排查方向",
|
||||
"category": "api",
|
||||
"keywords": ["ERR_TIMEOUT","超时","支付网关"]
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Milvus 向量索引
|
||||
|
||||
每个文档会被分块(chunk)并生成向量,存储到 Milvus 集合中:
|
||||
|
||||
**Collection**: `knowledge_base_collection`
|
||||
|
||||
**字段**:
|
||||
- `doc_id`:文档 ID
|
||||
- `chunk_id`:分块 ID
|
||||
- `chunk_text`:分块文本内容
|
||||
- `embedding`:768 维向量
|
||||
- `category`:文档分类
|
||||
- `file_path`:文件路径
|
||||
|
||||
**分块策略**:
|
||||
- Chunk Size:根据 `DocumentChunkConfig` 配置(默认 500 token)
|
||||
- Overlap:重叠区域(默认 50 token)
|
||||
|
||||
---
|
||||
|
||||
## L0 内存索引
|
||||
|
||||
导入过程会自动将文档加入 `KnowledgeIndexService` 的内存索引:
|
||||
|
||||
```java
|
||||
KnowledgeEntry entry = KnowledgeEntry.builder()
|
||||
.filePath(relativePath)
|
||||
.title(title)
|
||||
.keywords(keywords)
|
||||
.summary(summary)
|
||||
.category(category)
|
||||
.build();
|
||||
knowledgeIndexService.addToIndex(entry);
|
||||
```
|
||||
|
||||
**验证 L0 索引**:
|
||||
```bash
|
||||
# 应用启动后查看日志
|
||||
grep "知识库索引加载完成" logs/application.log
|
||||
|
||||
# 输出示例:
|
||||
# [INFO] 知识库索引加载完成,共 6 个文档
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 错误处理
|
||||
|
||||
### 常见错误
|
||||
|
||||
#### 1. 目录不存在
|
||||
```json
|
||||
{
|
||||
"success": false,
|
||||
"message": "初始化失败: 知识库目录不存在: knowledge_base"
|
||||
}
|
||||
```
|
||||
|
||||
**解决**:
|
||||
```bash
|
||||
mkdir -p knowledge_base/api
|
||||
mkdir -p knowledge_base/infrastructure
|
||||
mkdir -p knowledge_base/domain
|
||||
mkdir -p knowledge_base/troubleshooting
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
#### 2. 文档格式无效
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"scanned": 6,
|
||||
"inserted": 5,
|
||||
"failed": 1,
|
||||
"details": {
|
||||
"test/invalid.md": "格式无效: frontmatter 解析失败"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**原因**:
|
||||
- 缺少 frontmatter
|
||||
- YAML 格式错误
|
||||
- 缺少必填字段(title, keywords, summary)
|
||||
|
||||
**解决**:
|
||||
```markdown
|
||||
---
|
||||
title: 文档标题
|
||||
keywords: [关键词1, 关键词2]
|
||||
summary: 文档摘要
|
||||
category: api
|
||||
---
|
||||
|
||||
# 正文内容
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 问题 4: Milvus 连接失败
|
||||
|
||||
**症状**:
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"scanned": 6,
|
||||
"inserted": 0,
|
||||
"failed": 6,
|
||||
"details": {
|
||||
"api/payment-errors.md": "Milvus 索引失败: Connection refused"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**原因**:
|
||||
- Milvus 服务未启动
|
||||
- 网络连接问题
|
||||
- 配置错误
|
||||
|
||||
**解决**:
|
||||
```bash
|
||||
# 检查 Milvus 是否运行
|
||||
docker ps | grep milvus
|
||||
|
||||
# 检查配置
|
||||
grep milvus application.yml
|
||||
|
||||
# 启动 Milvus
|
||||
docker-compose up -d milvus-standalone
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 问题 5: 文档分块失败
|
||||
|
||||
**症状**:
|
||||
```json
|
||||
{
|
||||
"details": {
|
||||
"test/large-doc.md": "Milvus 索引失败: Document too large"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**原因**:
|
||||
- 文档内容过大
|
||||
- 分块配置不当
|
||||
|
||||
**解决**:
|
||||
- 检查 `DocumentChunkConfig` 配置
|
||||
- 调整 chunk size 和 overlap
|
||||
|
||||
---
|
||||
|
||||
#### 3. 文档缺少标题
|
||||
```json
|
||||
{
|
||||
"details": {
|
||||
"test/no-title.md": "缺少标题"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**解决**:在 frontmatter 中添加 `title` 字段。
|
||||
|
||||
---
|
||||
|
||||
## 最佳实践
|
||||
|
||||
### ✅ 推荐做法
|
||||
|
||||
1. **首次启动后立即初始化**:
|
||||
```bash
|
||||
mvn spring-boot:run
|
||||
sleep 15 # 等待启动完成
|
||||
curl -X POST http://localhost:9900/api/knowledge/init
|
||||
```
|
||||
|
||||
2. **新增文档后增量导入**:
|
||||
```bash
|
||||
# 不使用 force,只导入新文档
|
||||
curl -X POST http://localhost:9900/api/knowledge/init
|
||||
```
|
||||
|
||||
3. **定期检查统计信息**:
|
||||
```bash
|
||||
curl http://localhost:9900/api/knowledge/stats
|
||||
```
|
||||
|
||||
4. **更新文档内容后强制刷新**:
|
||||
```bash
|
||||
curl -X POST http://localhost:9900/api/knowledge/init?force=true
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### ❌ 避免做法
|
||||
|
||||
1. **不检查响应就认为成功**:
|
||||
- 始终检查 `failed` 字段
|
||||
- 查看 `details` 了解具体失败原因
|
||||
|
||||
2. **频繁使用 force=true**:
|
||||
- 会重复插入数据(违反唯一约束)
|
||||
- 建议先清理数据库,再使用 force
|
||||
|
||||
3. **不检查文档格式就导入**:
|
||||
- 先手动验证 frontmatter 格式
|
||||
- 确保必填字段完整
|
||||
|
||||
---
|
||||
|
||||
## 相关文档
|
||||
|
||||
- **知识库使用指南**:`mvp/architecture/knowledge-retrieval-usage.md`
|
||||
- **知识库架构**:`mvp/architecture/knowledge-retrieval-architecture.md`
|
||||
- **Executor Prompt**:`src/main/resources/prompts/executor-prompt.md`
|
||||
@@ -55,3 +55,12 @@ uploads/
|
||||
/volumes
|
||||
/server.pid
|
||||
.claude/settings.local.json
|
||||
.opencode/plugins/emdash-notifications.js
|
||||
|
||||
### Windows / Runtime Artifacts
|
||||
*.stackdump
|
||||
NUL
|
||||
|
||||
### MVP Demo Generated Outputs
|
||||
mvp/demo/output/*.json
|
||||
!mvp/demo/output/README.md
|
||||
|
||||
@@ -1,29 +0,0 @@
|
||||
Stack trace:
|
||||
Frame Function Args
|
||||
0007FFFFB920 00021005FE8E (000210285F68, 00021026AB6E, 000000000000, 0007FFFFA820) msys-2.0.dll+0x1FE8E
|
||||
0007FFFFB920 0002100467F9 (000000000000, 000000000000, 000000000000, 0007FFFFBBF8) msys-2.0.dll+0x67F9
|
||||
0007FFFFB920 000210046832 (000210286019, 0007FFFFB7D8, 000000000000, 000000000000) msys-2.0.dll+0x6832
|
||||
0007FFFFB920 000210068CF6 (000000000000, 000000000000, 000000000000, 000000000000) msys-2.0.dll+0x28CF6
|
||||
0007FFFFB920 000210068E24 (0007FFFFB930, 000000000000, 000000000000, 000000000000) msys-2.0.dll+0x28E24
|
||||
0007FFFFBC00 00021006A225 (0007FFFFB930, 000000000000, 000000000000, 000000000000) msys-2.0.dll+0x2A225
|
||||
End of stack trace
|
||||
Loaded modules:
|
||||
000100400000 bash.exe
|
||||
7FF9B93D0000 ntdll.dll
|
||||
7FF9B79A0000 KERNEL32.DLL
|
||||
7FF9B6860000 KERNELBASE.dll
|
||||
7FF9B8740000 USER32.dll
|
||||
7FF9B6830000 win32u.dll
|
||||
7FF9B84F0000 GDI32.dll
|
||||
7FF9B6CD0000 gdi32full.dll
|
||||
7FF9B6790000 msvcp_win.dll
|
||||
7FF9B7000000 ucrtbase.dll
|
||||
000210040000 msys-2.0.dll
|
||||
7FF9B7370000 advapi32.dll
|
||||
7FF9B8E40000 msvcrt.dll
|
||||
7FF9B85B0000 sechost.dll
|
||||
7FF9B6FD0000 bcrypt.dll
|
||||
7FF9B90F0000 RPCRT4.dll
|
||||
7FF9B5F20000 CRYPTBASE.DLL
|
||||
7FF9B6710000 bcryptPrimitives.dll
|
||||
7FF9B86E0000 IMM32.DLL
|
||||
@@ -4,7 +4,20 @@
|
||||
|
||||
| 日期 | slug | 领域 | 关键词 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
|
||||
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
|
||||
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
|
||||
| 2026-07-04 | diagnosis-eval-harness | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
|
||||
| 2026-07-04 | evidence-trace-hardening | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
|
||||
| 2026-07-03 | mvp-demo-trace-acceptance | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
|
||||
| 2026-07-04 | aiops-traceable-diagnosis-entry | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
|
||||
| 2026-07-04 | aiops-alert-scope-control | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
|
||||
| 2026-05-29 | chatmodel-abstraction | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | archived |
|
||||
| 2026-06-23 | phase1-infrastructure | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | archived |
|
||||
| 2026-06-24 | lookup-knowledge-integration | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | archived |
|
||||
| 2026-06-25 | doc-management-ui | 前端开发/文档管理 | 文档管理页面, CRUD, 状态监控, 纯静态页面, API集成 | archived |
|
||||
| 2026-06-26 | session-storage | 会话存储/可观测 | diagnosis_session, agent_step, tool_invocation, token追踪, 多Agent路由 | openspec/changes/session-storage | archived |
|
||||
| 2026-06-29 | confidence-feedback | 质量评估/反馈机制 | evidence_score, selfEvaluation, feedback, useful, not_useful, case_library, BAD_CASE, tool_invocation规则引擎, 反馈按钮, sessionId回传 | openspec/changes/confidence-feedback | archived |
|
||||
| 2026-06-30 | session-dedup-knowledge-map | 去重/知识图谱 | RetrievedDocTracker, KnowledgeDomainService, knowledge_domain, covers, whenToRetrieve, Planner注入, ISS-001 | openspec/changes/archive/2026-06-30-session-dedup-knowledge-map | archived |
|
||||
| 2026-07-01 | executor-action-memory-relevance | 检索质量/行动记忆 | relevanceLevel, completenessHint, Min-Max归一化, RetrievedDocTracker域级记录, Executor检索约束, ISS-002 | openspec/changes/archive/2026-07-01-executor-action-memory-relevance | archived |
|
||||
| 2026-07-02 | chat-verifier-agent | Chat质量门禁/可追溯验证 | Verifier, groundedness_score, facts_checked, evidence_refs, tool_trace_summary, self_evaluation | openspec/changes/archive/2026-07-03-chat-verifier-agent | archived |
|
||||
|
||||
@@ -0,0 +1,28 @@
|
||||
# 验收记录
|
||||
|
||||
## 验证情况
|
||||
|
||||
### 静态验证
|
||||
- [x] 编译通过(`mvn compile`)
|
||||
- [x] 42 个测试全部通过(DocumentChunkService / LookupKnowledgeTool / Repository)
|
||||
- [x] 三张新表通过 Flyway 成功创建
|
||||
|
||||
### 脚本验证
|
||||
- [x] `/api/chat` — 单 Agent 正常响应,agent_step 记录正确
|
||||
- [x] `/api/chat` — 复杂问题路由到多 Agent(Planner + Executor)
|
||||
- [x] `/api/ai_ops` — 多 Agent 流程正常,planner 步骤写入 agent_step
|
||||
- [x] Tool_invocation L0/L1 检索质量明细正确
|
||||
- [x] diagnosis_session 汇总指标(total_token_count / step_count / tool_call_count)正确
|
||||
- [x] TokenTrackingChatModel 捕获实际 token 数(已验证 total=827)
|
||||
- [x] 旧 diagnosis_record 表删除成功
|
||||
|
||||
### 未验证
|
||||
- `/api/chat_stream`(SSE 流式)— 未接入 session 存储,不在本次范围,后续覆盖
|
||||
- `self_evaluation` / `feedback` — 无前端交互入口
|
||||
|
||||
## 剩余风险
|
||||
|
||||
| 风险 | 说明 |
|
||||
|------|------|
|
||||
| Token 累加 | 当前每步独立记录,汇总在 `backfillSessionMetrics`,未在 Hook 层累加 |
|
||||
| Async 优化 | 同步写 DB 在低并发下无问题,后续可引入 @Async |
|
||||
@@ -0,0 +1,21 @@
|
||||
# 会话存储体系
|
||||
|
||||
## 背景
|
||||
当前 `diagnosis_record` 单表字段耦合在"告警分析"领域,无法支撑通用会话存储。缺少 Agent 决策链维度、检索质量明细、Token 消耗等可观测指标。
|
||||
|
||||
## 目标
|
||||
将单表拆分为三表体系,覆盖 ChatService 和 AiOpsService 两个 Agent 的完整决策链记录,支撑可观测和评估。
|
||||
|
||||
## 范围
|
||||
- 新建 3 张表(diagnosis_session / agent_step / tool_invocation)
|
||||
- Flyway 迁移 + JPA Entity + Repository
|
||||
- 改造 AgentLoggingHook 持久化 agent_step
|
||||
- 改造 LookupKnowledgeTool 写入 tool_invocation
|
||||
- ChatService / AiOpsService 支持 diagnosis_session 生命周期
|
||||
- Token 用量追踪(TokenTrackingChatModel)
|
||||
- 意图识别路由(单 Agent / 多 Agent)
|
||||
- 删除旧 diagnosis_record 表
|
||||
|
||||
## 非目标
|
||||
- 不涉及 UI 层面的会话展示
|
||||
- 不涉及历史数据迁移
|
||||
@@ -0,0 +1,22 @@
|
||||
# 会话存储 — 决策记录
|
||||
|
||||
## 关键决策
|
||||
|
||||
| 决策 | 选择 | 理由 |
|
||||
|------|------|------|
|
||||
| AgentLoggingHook 创建方式 | POJO(构造注入),非 @Component | 需为 ChatService/AiOpsService 创建多个实例(不同 agentName) |
|
||||
| AiOpsService 记录粒度 | 只记子 Agent(Planner/Executor),不记 Supervisor | Supervisor 编排日志已有体现,单独记录增加噪音 |
|
||||
| sessionId 传递 | RunnableConfig.metadata(优先)+ ThreadLocal(兜底) | RunnableConfig 线程安全,异步兼容 |
|
||||
| Tool 获取 sessionId | SessionContextHolder(ThreadLocal) | Tool 不在调用链中,无法通过 RunnableConfig 获取 |
|
||||
| Token 追踪 | TokenTrackingChatModel 包装器拦截 ChatModel.call() | 框架 _TOKEN_USAGE_ 仅 stream 路径可用 |
|
||||
| Chat 复杂度路由 | 关键词 + 长度判断 | MVP 简化实现 |
|
||||
| 多 Agent Planner 无工具 | 不注入 methodTools/tools | 防止 Planner 自己执行,强制通过 Executor 执行 |
|
||||
| 旧表处理 | V007 Flyway 迁移删除 diagnosis_record | 被三表替代,不再使用 |
|
||||
|
||||
## 风险
|
||||
|
||||
| 风险 | 等级 | 说明 |
|
||||
|------|:----:|------|
|
||||
| Hook 同步写 DB | 低 | MVP 阶段数据量小,后续可异步化 |
|
||||
| token_count 依赖 ChatResponse.usage | 低 | DeepSeek 已确认返回实际用量 |
|
||||
| stream 路径 session 记录 | 低 | 当前 call 路径正常,stream 需确认 RunnableConfig 传播 |
|
||||
@@ -0,0 +1,23 @@
|
||||
# 证据记录
|
||||
|
||||
## Evidence-Driven 查证
|
||||
|
||||
### E1: AgentLoggingHook 创建方式
|
||||
- **发现**: ChatService 通过 `new AgentLoggingHook()` 创建,非 Spring 管理,无法注入 Repository
|
||||
- **结论**: 需要改造为可注入的 POJO(构造注入)
|
||||
- **影响**: Hook 重构为构造注入 Repository + agentName
|
||||
|
||||
### E2: AiOpsService 未使用 Hook
|
||||
- **发现**: AiOpsService 的 Planner / Executor / Supervisor 均未配置 AgentLoggingHook
|
||||
- **结论**: 需要补齐,每个子 Agent 加 Hook
|
||||
- **影响**: Planner 和 Executor 各加 Hook,Supervisor 不加
|
||||
|
||||
### E3: 项目无异步基础设施
|
||||
- **发现**: 全局搜索 `@Async` / `@EnableAsync` 均无匹配
|
||||
- **结论**: MVP 阶段同步写 DB,后续优化
|
||||
- **影响**: 标记为技术债
|
||||
|
||||
### E4: RunnableConfig 支持 metadata
|
||||
- **发现**: `RunnableConfig` 的 `metadata` 为 `ConcurrentMap`,可在构建时设置
|
||||
- **结论**: sessionId 通过 `config.addMetadata("sessionId", id)` 传递,线程安全
|
||||
- **影响**: 取代 ThreadLocal 方案
|
||||
@@ -0,0 +1,64 @@
|
||||
# acceptance.md — confidence-feedback
|
||||
|
||||
## 实现清单
|
||||
|
||||
| 任务 | 文件 | 状态 |
|
||||
|---|---|---|
|
||||
| T0:Flyway V008 + answer 字段 | `V008__add_answer_to_diagnosis_session.sql`、`DiagnosisSession.java` | 完成 |
|
||||
| T1:EvaluationService(规则引擎) | `EvaluationService.java` | 完成 |
|
||||
| T2:ChatService 后置调用 | `ChatService.java` | 完成 |
|
||||
| T3:FeedbackController + FeedbackService | `FeedbackController.java`、`FeedbackService.java`、`FeedbackRequest.java`、`FeedbackResponse.java` | 完成 |
|
||||
| T4:CaseLibraryService | `CaseLibraryService.java` | 完成 |
|
||||
| T5:AsyncConfig | `AsyncConfig.java` | 完成 |
|
||||
|
||||
## 验证记录
|
||||
|
||||
### 静态验证(已通过)
|
||||
|
||||
- `mvn compile` BUILD SUCCESS(2026-06-30)
|
||||
- 无新增 ERROR,存量 WARNING 与本次改动无关
|
||||
- import 完整性人工检查通过
|
||||
|
||||
### 脚本验证(已通过,2026-06-30)
|
||||
|
||||
验证工具:`scripts/query_mysql.py`(本次新建)
|
||||
|
||||
| 步骤 | 操作 | 结果 |
|
||||
|---|---|---|
|
||||
| 1 | POST /api/chat 发送问题 | 200,answer 有值 |
|
||||
| 2 | 等 5 秒查 diagnosis_session | self_evaluation 写入规则引擎结果,answer 写入完整回答 |
|
||||
| 3 | POST /api/feedback useful | 200,返回 caseId;case_library 新增一行,feedback=useful,status=SUCCESS |
|
||||
| 4 | POST /api/feedback not_useful | 200,feedback=not_useful,status 仍为 SUCCESS(未被改写) |
|
||||
| 5(边界)| 重复提交 useful | 返回同一 caseId,case_library 无重复插入 |
|
||||
| 6(边界)| 非法 feedback 值 | HTTP 400 |
|
||||
|
||||
### Flyway V008 迁移
|
||||
|
||||
- 服务启动后 diagnosis_session 表存在 answer 列,验证通过(步骤 2 能写入 answer)
|
||||
|
||||
### 浏览器/人工验证(已通过,2026-06-30)
|
||||
|
||||
| 步骤 | 操作 | 结果 |
|
||||
|---|---|---|
|
||||
| 1 | 发送"今天天气怎么样" | AI 回复下方出现"有用/无用"按钮 |
|
||||
| 2 | 点击"有用" | 按钮区域替换为"已标记为有用" |
|
||||
| 3 | 网络请求确认 | POST /api/feedback 返回 HTTP 200,`success: true` |
|
||||
|
||||
### 前端反馈按钮(追加,2026-06-30)
|
||||
|
||||
**改动文件**:`app.js`、`styles.css`
|
||||
|
||||
关键设计:
|
||||
- `ChatResult` record 新增(`ChatService`),`ChatResponse` 增加 `sessionId` 字段(`ChatController`)
|
||||
- `sendQuickMessage` 读取 `chatResponse.sessionId` 存为 `this.lastSessionId`
|
||||
- `createFeedbackBar(sessionId)` 闭包绑定 sessionId,避免多轮对话时 sessionId 错位
|
||||
- `submitFeedback(feedback, barElement, sessionId)` 直接用传入参数,不依赖全局状态
|
||||
- 流式模式(`/api/chat_stream`)反馈按钮会渲染,但 sessionId 为空,点击不生效(已知限制)
|
||||
|
||||
## 已知限制
|
||||
|
||||
- 非检索工具(DateTimeTools 等)不写 tool_invocation,evidence_score = 0(已接受,符合"证据充分度"定义)
|
||||
- `@Async` 失败时 selfEvaluation 为 null,前端需处理 null(已接受)
|
||||
- CaseLibrary 的 faultCategory 固定为 GENERAL,需人工补充(已接受,Phase 2 优化)
|
||||
- LLM 观点层未实现,selfEvaluation JSON 预留 llm_opinion 扩展位(Phase 2)
|
||||
- 流式模式反馈按钮 sessionId 缺失,暂不处理(已知,后续处理流式接口时一并解决)
|
||||
@@ -0,0 +1,37 @@
|
||||
# brief.md — confidence-feedback
|
||||
|
||||
## 背景
|
||||
|
||||
DiagnosisSession 已预留 `selfEvaluation`(JSON)和 `feedback`(VARCHAR 16)两个字段,但完全为空。Agent 完成对话后不计算证据评分,也没有接收用户反馈的 API,无法支撑报告质量评估和 BadCase 追踪。
|
||||
|
||||
## 目标
|
||||
|
||||
1. 给每次对话结果自动打一个基于事实的证据充分度评分(evidence_score)
|
||||
2. 提供用户反馈 API(useful/not_useful),useful 触发案例自动沉淀,not_useful 标记 BadCase
|
||||
|
||||
## 范围
|
||||
|
||||
- `DiagnosisSession` 加 `answer` 字段(Flyway V008)
|
||||
- `EvaluationService`:基于 tool_invocation 的规则引擎,@Async 写 selfEvaluation
|
||||
- `FeedbackController` + `FeedbackService`:POST /api/feedback
|
||||
- `CaseLibraryService.createFromSession`:幂等案例沉淀
|
||||
- `AsyncConfig`:@EnableAsync
|
||||
- `ChatService`:SUCCESS 分支写 answer + 触发 evaluate;新增 `ChatResult` record 回传 sessionId
|
||||
- `ChatController.ChatResponse` 增加 `sessionId` 字段
|
||||
- 前端 `app.js`:AI 回复下方反馈按钮,点击调用 `/api/feedback`,闭包绑定 sessionId
|
||||
- 前端 `styles.css`:反馈栏样式
|
||||
|
||||
## 非目标
|
||||
|
||||
- 不实现 Verifier Agent 完整链路
|
||||
- 不实现 LLM 自评(预留扩展位,Phase 2 再做)
|
||||
- 不实现案例结构化字段自动填充(faultCategory 等暂时填 GENERAL)
|
||||
- 不实现 BadCase 自动分析或 Prompt 优化
|
||||
|
||||
## 分档
|
||||
|
||||
standard
|
||||
|
||||
## 关联 OpenSpec
|
||||
|
||||
`openspec/changes/confidence-feedback/`
|
||||
@@ -0,0 +1,115 @@
|
||||
# decisions.md — confidence-feedback
|
||||
|
||||
## Question Pool(grill 阶段)
|
||||
|
||||
| # | 问题 | 模式 | 状态 |
|
||||
|---|---|---|---|
|
||||
| Q1 | 置信度由谁计算 | user-interview | 已确认 |
|
||||
| Q2 | 反馈触发哪些后端操作 | user-interview | 已确认 |
|
||||
| Q3 | CaseLibrary 结构化字段从哪里填 | evidence-driven | 已确认(方案变更) |
|
||||
| Q4 | 验收口径 | user-interview | 已确认 |
|
||||
|
||||
---
|
||||
|
||||
## Evidence-Driven 结论
|
||||
|
||||
### Q3:CaseLibrary 内容来源
|
||||
|
||||
**初始结论**:从 `agent_step.thought` 提取(grill 阶段)
|
||||
|
||||
**修正(apply 阶段讨论后)**:
|
||||
- 代码证据:`agent_step.thought` 截断为 2000 字符,`modelOutput` 截断为 500 字符,均不是完整答案
|
||||
- `ChatService.executeChat` 第 269 行已有完整答案 `answer = response.getText()`,但未持久化
|
||||
- 决策:给 `DiagnosisSession` 加 `answer TEXT` 字段,Flyway V008 迁移,案例内容直接从 `session.answer` 取
|
||||
|
||||
---
|
||||
|
||||
## User-Interview 确认记录
|
||||
|
||||
### Q1 — 置信度由谁评估
|
||||
- 用户原话(grill):"两者都要:规则兜底 + Verifier 主打分"
|
||||
- **apply 后修正**:讨论后决定去掉 LLM 自评,仅用规则引擎(见"apply 阶段决策")
|
||||
- 最终实现:`EvaluationService` 纯规则,预留 `llm_opinion` 扩展位
|
||||
|
||||
### Q2 — 反馈触发操作
|
||||
- 用户原话:"写入 DiagnosisSession.feedback 字段, not_useful → 打 BAD_CASE 标记"
|
||||
- **apply 后修正**:BAD_CASE 不改 status,feedback 字段本身即为标记(见"apply 阶段决策")
|
||||
- 最终实现:`FeedbackService` 只写 feedback + 可选写 case_library,不改 status
|
||||
|
||||
### Q4 — 验收口径
|
||||
- 用户原话:"端到端可验证:发一次 chat → 查 DB 看 selfEvaluation 有值 → 提交 feedback → 查 DB 看 feedback + case_library"
|
||||
- 确认状态:已确认,未变化
|
||||
|
||||
---
|
||||
|
||||
## Apply 阶段决策(post-grill 重要变更)
|
||||
|
||||
### 决策 A:DiagnosisSession 加 answer 字段
|
||||
|
||||
- **问题**:案例沉淀需要完整答案,agent_step.thought 被截断,不可用
|
||||
- **决策**:新增 `answer LONGTEXT` 字段,ChatService SUCCESS 分支写入
|
||||
- **影响**:V008 Flyway 迁移,CaseLibraryService 直接读 session.answer
|
||||
|
||||
### 决策 B:去掉 LLM 自评,只用规则引擎
|
||||
|
||||
- **问题**:LLM 评估自己的答案系统性偏高分;多一次调用消耗 token;Verifier Agent 当前未实现
|
||||
- **决策**:MVP 阶段仅用基于 tool_invocation 的规则引擎
|
||||
- **理由**:规则可解释、可复现、不撒谎;Verifier 留待诊断全链路实现时再做
|
||||
- **预留**:`selfEvaluation` JSON 结构保留 `llm_opinion` 扩展位,代码底部注释说明接入点
|
||||
|
||||
### 决策 C:BAD_CASE 不改 status 字段
|
||||
|
||||
- **问题**:status 是执行状态语义(RUNNING/SUCCESS/FAILED),BAD_CASE 是质量标签,两个维度不同;覆盖 status 会破坏统计
|
||||
- **决策**:`not_useful` 通过 `feedback` 字段本身标识,查 BadCase 用 `WHERE feedback = 'not_useful'`
|
||||
|
||||
### 决策 D:评分字段重命名为 evidence_score
|
||||
|
||||
- **问题**:原名 confidence 容易误解为"答案准确性",实际衡量的是"证据收集充分度"
|
||||
- **决策**:重命名为 `evidence_score`,明确语义边界
|
||||
- **边界说明**:工具调用能证明 Agent 有尝试收集证据,但无法证明答案无幻觉;这个分数过滤最差情况(无工具调用就给答案),不能识别"调用了工具但结论仍错误"
|
||||
|
||||
### 决策 E:规则输入来源仅限 tool_invocation 事实
|
||||
|
||||
- **问题**:DateTimeTools、QueryMetricsTools 等非检索工具调用未写入 tool_invocation
|
||||
- **接受**:evidence_score 定义本来就是检索证据充分度,非检索工具排除在外是合理的,不是 bug
|
||||
- **已知限制**:调用了时间工具但 evidence_score = 0 的 session 存在
|
||||
|
||||
---
|
||||
|
||||
## 架构审计记录
|
||||
|
||||
- 接口影响:`POST /api/feedback` 是新接口(L2);ChatService 主流程返回值不变(L1)
|
||||
- 时序验证:tool_invocation 在工具执行时同步写入,evaluate @Async 在 Agent 完成后触发,无竞态问题
|
||||
- 已接受风险:
|
||||
- `@Async` 失败时 selfEvaluation 保持 null,前端需处理 null
|
||||
- 案例结构化字段(faultCategory 等)暂时填 GENERAL,后续可人工补充
|
||||
- LLM 自评预留但未实现,Phase 2 再迭代
|
||||
|
||||
### 决策 F:ChatResult record + ChatResponse.sessionId 回传
|
||||
|
||||
- **问题**:`ChatService` 内部生成 8 位 sessionId,但从不返回给前端;前端用自己的 sessionId 调 feedback 接口,后端查不到 session(400)
|
||||
- **决策**:新增 `ChatResult(answer, sessionId)` record,`executeChatWithStrategy` 链路全部返回 `ChatResult`;`ChatResponse` 增加 `sessionId` 字段;前端读取并闭包绑定至对应消息的反馈按钮
|
||||
- **影响**:`ChatService` 三个方法签名变更(内部链路),`ChatController` 调用方更新,前端 `app.js` 读取新字段
|
||||
|
||||
### 决策 G:反馈 sessionId 闭包绑定而非全局变量
|
||||
|
||||
- **问题**:最初实现用 `this.lastSessionId` 全局变量,多轮对话时点击早期消息的反馈按钮会提交最新 sessionId
|
||||
- **决策**:`createFeedbackBar(sessionId)` 接收 sessionId 参数,`submitFeedback(feedback, bar, sessionId)` 直接用传入值,不读全局状态
|
||||
- **效果**:每条 AI 回复绑定自己那轮的 sessionId,多轮对话下行为正确
|
||||
|
||||
### 项目技术栈清单
|
||||
|
||||
- ChatModel 注入:`@Autowired ChatModel chatModel`,通过 `ModelRoutingConfig` 路由
|
||||
- Repository:Spring Data JPA,`Optional<T>` 返回,方法命名约定
|
||||
- DTO:独立文件放 `dto/` 包
|
||||
- 异步:新建 `AsyncConfig.java` 加 `@EnableAsync`(项目原无此配置)
|
||||
- 无 MQ,无加密,工具类直接用 UUID.randomUUID()
|
||||
- 日志:SLF4J Logger,`LoggerFactory.getLogger()`
|
||||
- `ToolInvocationRepository.findBySessionId` 已有,可直接用
|
||||
|
||||
### 参考实现文件
|
||||
|
||||
- `ChatService.java`:executeChat/executeChatComplex 流程
|
||||
- `CaseLibraryRepository.findByDiagnosisId`:幂等检查用
|
||||
- `DiagnosisSessionRepository.findBySessionId`
|
||||
- `ToolInvocationRepository.findBySessionId`
|
||||
@@ -0,0 +1,52 @@
|
||||
# evidence.md — confidence-feedback
|
||||
|
||||
## 代码证据
|
||||
|
||||
### agent_step.thought 不可作为案例内容
|
||||
|
||||
- 文件:`AgentLoggingHook.java:135`
|
||||
- 证据:`thought` 在写入前截断为 2000 字符,`modelOutput` 截断为 500 字符
|
||||
- 结论:两者均不是返回给用户的完整答案,案例质量低
|
||||
|
||||
### ChatService 已有完整答案未持久化
|
||||
|
||||
- 文件:`ChatService.java:269`(executeChat)、`ChatService.java:353`(executeChatComplex)
|
||||
- 证据:`String answer = response.getText()` 只用于返回前端,未写入任何持久化存储
|
||||
- 结论:加 `DiagnosisSession.answer` 字段是最干净的方案
|
||||
|
||||
### ToolInvocationRepository 已有 findBySessionId
|
||||
|
||||
- 文件:`ToolInvocationRepository.java`
|
||||
- 证据:`findBySessionId(String sessionId)` 已实现,返回 `List<ToolInvocation>`
|
||||
- 结论:规则引擎可直接读取 tool_invocation 事实,无需新增查询方法
|
||||
|
||||
### tool_invocation 写入时序安全
|
||||
|
||||
- 文件:`LookupKnowledgeTool.java:144`
|
||||
- 证据:`saveToolInvocation` 在工具执行时同步调用,早于 ChatService 的 SUCCESS 分支
|
||||
- 结论:@Async evaluate 触发时 tool_invocation 数据已在库,无竞态
|
||||
|
||||
### 项目原无 @EnableAsync
|
||||
|
||||
- 证据:`grep -rn "EnableAsync"` 无任何命中(apply 前)
|
||||
- 结论:需要新建 `AsyncConfig.java`
|
||||
|
||||
### CaseLibraryRepository.findByDiagnosisId 已有幂等检查支持
|
||||
|
||||
- 文件:`CaseLibraryRepository.java`
|
||||
- 证据:`findByDiagnosisId(String diagnosisId)` 已实现
|
||||
- 结论:useful 重复提交时可用此方法检查,不重复插入
|
||||
|
||||
## 设计推导
|
||||
|
||||
### evidence_score vs confidence 命名
|
||||
|
||||
- 基于工具调用的分数衡量的是证据收集充分度,不是答案准确性
|
||||
- "confidence" 容易误解,改为 "evidence_score" 更准确
|
||||
- LLM 自评才适合叫 confidence,但当前未实现
|
||||
|
||||
### BAD_CASE 不应混入 status
|
||||
|
||||
- status 有明确执行状态语义(RUNNING/SUCCESS/FAILED)
|
||||
- 一个 SUCCESS 的 session 被标为 BAD_CASE 后,按 status 做的统计会失真
|
||||
- feedback 字段本身就够,`WHERE feedback = 'not_useful'` 即可查 BadCase
|
||||
@@ -0,0 +1,58 @@
|
||||
# Acceptance: session-dedup-knowledge-map
|
||||
|
||||
## 静态验证
|
||||
|
||||
| 项目 | 结果 | 说明 |
|
||||
|------|------|------|
|
||||
| 编译检查 | PASS | `mvn compile -q` exit code 0,所有 17 个变更文件无编译错误 |
|
||||
| 代码结构检查 | PASS | 6 个新文件(RetrievedDocTracker, DocumentFieldEnricher, KnowledgeDomainService, KnowledgeDomain, KnowledgeDomainRepository, V009 迁移)均存在且路径正确 |
|
||||
| Prompt 外部化 | PASS | `doc-field-enricher-prompt.md` 和 `domain-summary-prompt.md` 位于 `src/main/resources/prompts/`,Java 代码通过 `@PostConstruct` + `ClassPathResource` 加载 |
|
||||
| Flyway 迁移脚本 | PASS | `V009__add_knowledge_domain.sql` 存在,表结构完整 |
|
||||
| DTO 字段 | PASS | Frontmatter / KnowledgeEntry / LookupResult 新增字段均已添加 |
|
||||
| 解析器扩展 | PASS | FrontmatterParser 解析 `covers` 和 `when_to_retrieve` |
|
||||
| Jackson 替换 | PASS | KnowledgeIndexService 不再包含 extractJsonValue/extractJsonArray,改用 objectMapper.readValue |
|
||||
| Prompt 检索规则 | PASS | chat-planner-prompt.md 新增"知识库检索规则"区块(4 条规则) |
|
||||
|
||||
## 脚本验证
|
||||
|
||||
| 项目 | 结果 | 说明 |
|
||||
|------|------|------|
|
||||
| 单元测试 | 未运行 | 项目当前无针对本 change 的单元测试 |
|
||||
| 集成测试 | 未运行 | 需启动应用 + Milvus + MySQL 验证完整链路 |
|
||||
|
||||
## 浏览器/人工验证
|
||||
|
||||
| 项目 | 结果 | 说明 |
|
||||
|------|------|------|
|
||||
| V009 迁移 | PASS | Flyway 日志:`Successfully applied 1 migration to schema superbiz_agent, now at version v009` |
|
||||
| knowledge_domain 表数据 | PASS | 4 个域全部 LLM 生成 when_to_retrieve 成功(api/domain/infrastructure/troubleshooting),内容包含跨域边界引用 |
|
||||
| knowledge map 注入 Planner | PASS | 多 Agent 路径正常触发 `Supervisor → chat_planner → chat_executor`,Planner 能按域做检索规划 |
|
||||
| session 级去重 | PASS | 两个 session 均验证去重生效:session `7c517329` 去 4 次重拦截,session `9693b9fb` 6 次去重拦截 |
|
||||
| LLM 字段生成 | 未验证 | 需上传新文档后检查 metadata JSON 中是否包含 covers 和 whenToRetrieve |
|
||||
|
||||
## 未验证项
|
||||
|
||||
| 项目 | 风险 | 建议补验步骤 |
|
||||
|------|------|-------------|
|
||||
| LLM 字段生成 | 中 — 依赖外部 LLM 服务 | 上传新文档,检查 metadata JSON 中是否包含 covers 和 whenToRetrieve |
|
||||
|
||||
## 启动问题修复
|
||||
|
||||
| 问题 | 修复 | 状态 |
|
||||
|------|------|------|
|
||||
| `@PostConstruct` 中调用 `knowledgeDomainService.onDocumentChange()` 导致循环依赖 | 将域级生成从 `@PostConstruct` 移到 `@EventListener(ApplicationReadyEvent.class)` | 已修复,编译通过 |
|
||||
|
||||
## 任务完成状态
|
||||
|
||||
14/14 任务全部完成 (T1-1 ~ T6-2)。
|
||||
|
||||
## 遗留问题
|
||||
|
||||
ISS-002:Executor 无约束重复调用 `lookup_knowledge`(单会话 20+ 次),knowledge map 和检索约束只注入了 Planner 未注入 Executor。详见 `mvp/issues/ISS-002-executor-unconstrained-lookup.md`。
|
||||
|
||||
## 已知限制
|
||||
|
||||
1. **RetrievedDocTracker 为 JVM 内存存储**:应用重启后去重状态丢失,同一会话内重启无法继续去重(可接受,会话通常短于重启间隔)
|
||||
2. **Planner 只看域级 when_to_retrieve**:文档级细粒度筛选留 Phase 2
|
||||
3. **文档级 prompt 依赖同域其他文档**:首个上传到某域的文档无法获得同域参照(此时 prompt 输出"无同域其他文档")
|
||||
4. **域级 prompt 依赖其他域已入库**:首次启动且 DB 为空时,其他域信息从 L0 索引 category 列表兜底
|
||||
@@ -0,0 +1,33 @@
|
||||
# Brief: session-dedup-knowledge-map
|
||||
|
||||
## 背景
|
||||
|
||||
ISS-001:Executor 在单次对话中重复调用 `lookup_knowledge` 多达 20 次,同一文档被召回 13 次。原因是工具层无状态、Planner 无知识边界感知。
|
||||
|
||||
## 目标
|
||||
|
||||
1. 彻底消除 session 内重复文档召回(Part A)
|
||||
2. 给 Planner 注入知识图谱,让其在规划阶段就能判断需要检索哪个域、只检索一次(Part B)
|
||||
|
||||
## 范围
|
||||
|
||||
- `LookupKnowledgeTool`:session 级去重
|
||||
- `Frontmatter` / `KnowledgeEntry`:新增 covers + whenToRetrieve
|
||||
- `DocumentManagementService`:上传时 LLM 生成文档级字段
|
||||
- `KnowledgeDomainService`(新):域级聚合与 DB 存储
|
||||
- `knowledge_domain` 表(新)
|
||||
- `ChatService` + `chat-planner-prompt.md`:注入 knowledge map
|
||||
|
||||
## 非目标(Phase 2)
|
||||
|
||||
- Executor 文档级 when_to_retrieve 细粒度筛选
|
||||
- RRF 混合重排
|
||||
- 文档 frontmatter 自动生成(手动覆盖 LLM 优先已支持)
|
||||
|
||||
## 分档
|
||||
|
||||
standard
|
||||
|
||||
## 关联 OpenSpec
|
||||
|
||||
openspec/changes/session-dedup-knowledge-map/
|
||||
@@ -0,0 +1,60 @@
|
||||
# decisions.md — session-dedup-knowledge-map
|
||||
|
||||
## Question Pool
|
||||
|
||||
| # | 问题 | 类型 | 状态 |
|
||||
|---|---|---|---|
|
||||
| Q1 | domain.when_to_retrieve 来源(手动/自动聚合/LLM上传时生成) | user-interview | 已确认 |
|
||||
| Q2 | LLM 生成时机(同步上传 vs 异步补全) | user-interview | 已确认 |
|
||||
| Q3 | knowledge map 结构(域级平铺 vs 两层) | user-interview | 已确认 |
|
||||
| Q4 | domain.when_to_retrieve 存储(内存 vs DB) | user-interview | 已确认 |
|
||||
| Q5 | Executor 文档级细粒度筛选是否进 MVP | user-interview | 已确认 |
|
||||
| E1 | ThreadLocal 在多 Agent 路径是否安全 | evidence-driven | 已汇报 |
|
||||
| E2 | 6 个文档是否全部有 category 字段 | evidence-driven | 已汇报 |
|
||||
| E3 | 去重 key 设计 | evidence-driven | 已汇报 |
|
||||
| E4 | Planner prompt token 增量是否可接受 | evidence-driven | 已汇报 |
|
||||
| E5 | EvaluationService.tool_call_count 影响 | evidence-driven | 已汇报 |
|
||||
|
||||
## Evidence-Driven 结论
|
||||
|
||||
- **E1**:`AsyncConfig` 只启用 `@EnableAsync`,无 TaskDecorator。`SupervisorAgent.invoke()` 是同步阻塞调用,工具调用与主线程同线程,ThreadLocal 当前路径安全。异步扩展时需补 TaskDecorator。
|
||||
- **E2**:全部 6 个文档均有 `category` 字段:api(1)、domain(1)、infrastructure(3)、troubleshooting(1)。
|
||||
- **E3**:`KnowledgeEntry.filePath` 在 L0 内唯一,L1 `_source` 字段也是 filePath,统一用 filePath 作去重 key。
|
||||
- **E4**:当前 planner prompt 21 行,注入 knowledge map 约增加 200-400 字符,可接受。
|
||||
- **E5**:去重后 `agent_step.has_tool_call` 减少,`tool_call_count` 降低,这是修复效果,`EvaluationService` 评分规则无需改动。
|
||||
|
||||
## User-Interview 确认记录
|
||||
|
||||
**Q1** — doc.when_to_retrieve 来源
|
||||
用户原话:选 C(上传时 LLM 自动生成)
|
||||
确认状态:已确认
|
||||
|
||||
**Q2** — LLM 生成时机
|
||||
用户原话:选 X(同步,上传时当场生成)
|
||||
确认状态:已确认
|
||||
|
||||
**Q3** — knowledge map 结构
|
||||
用户原话:认可两层结构(domain → documents[])
|
||||
确认状态:已确认
|
||||
补充:Planner 只注入域级 when_to_retrieve,文档级 when_to_retrieve 留 Executor 筛选(Phase 2)
|
||||
|
||||
**Q4** — domain.when_to_retrieve 存储
|
||||
用户原话:存 DB,这样每次启动都不用让 LLM 再总结一次
|
||||
确认状态:已确认 → 新建 knowledge_domain 表,Flyway 迁移脚本
|
||||
|
||||
**Q5** — Executor 文档级细粒度筛选
|
||||
用户原话:留 Phase 2
|
||||
确认状态:已确认,MVP 不做
|
||||
|
||||
## Pre-apply 补充决策
|
||||
|
||||
- **P1:KnowledgeIndexService.parseDocumentToEntry 替换为 Jackson**:`extractJsonValue` / `extractJsonArray` 手写解析器遇到含逗号、引号的自然语言字段(whenToRetrieve)会截断。全量替换为 `objectMapper.readValue(metadata, Frontmatter.class)`,影响范围仅 `KnowledgeIndexService`,行为更健壮。(用户确认)
|
||||
- **P2:LookupResult 新增 message 字段**:去重命中时 `found=false` + `message="文档已在本会话中检索过:xxx"`,不复用 `primary.content`。语义清晰,LLM 能理解原因不会重试。(用户确认)
|
||||
|
||||
## 关键设计决策
|
||||
|
||||
1. **两级 when_to_retrieve**:文档级(upload 时 LLM 生成,存 metadata)+ 域级(文档变更时 LLM 聚合,存 knowledge_domain 表)
|
||||
2. **域级重算触发**:文档上传后、文档删除后,只重算受影响的域(不是全量);`loadIndex()` 时如果某域在 DB 没有记录,则触发生成
|
||||
3. **注入 Planner 只给域级**:knowledge map 只包含域级 when_to_retrieve + documents[](title + covers),不暴露文档级 when_to_retrieve
|
||||
4. **去重 key**:filePath(L0+L1 统一)
|
||||
5. **去重状态存储**:JVM 内 `ConcurrentHashMap<sessionId, Set<filePath>>`,`SessionContextHolder.clear()` 时同步清理
|
||||
@@ -0,0 +1,87 @@
|
||||
# Evidence: session-dedup-knowledge-map
|
||||
|
||||
## E1: ThreadLocal 在多 Agent 路径是否安全
|
||||
|
||||
**问题**:`SessionContextHolder` 基于 ThreadLocal,多 Agent 异步路径可能导致 sessionId 丢失。
|
||||
|
||||
**证据**:
|
||||
- `AsyncConfig` 只启用 `@EnableAsync`,无 `TaskDecorator`
|
||||
- `SupervisorAgent.invoke()` 是同步阻塞调用,工具调用与主线程同线程
|
||||
- 当前路径下 ThreadLocal 安全
|
||||
|
||||
**结论**:当前同步路径安全。未来引入异步扩展时需补 `TaskDecorator` 传递 ThreadLocal。
|
||||
|
||||
---
|
||||
|
||||
## E2: 6 个文档是否全部有 category 字段
|
||||
|
||||
**问题**:域聚合依赖 `category` 字段分组,需确认现有文档是否都有值。
|
||||
|
||||
**证据**:
|
||||
- 全部 6 个文档均有 `category` 字段:api(1)、domain(1)、infrastructure(3)、troubleshooting(1)
|
||||
|
||||
**结论**:现有文档无需修补,category 覆盖率 100%。
|
||||
|
||||
---
|
||||
|
||||
## E3: 去重 key 设计
|
||||
|
||||
**问题**:用什么字段唯一标识一个文档用于去重。
|
||||
|
||||
**证据**:
|
||||
- `KnowledgeEntry.filePath` 在 L0 索引内唯一
|
||||
- L1 向量索引的 `_source` 字段也是 filePath
|
||||
- 上传时 `saveToLocal()` 生成 `knowledge_base/{category}/{fileName}` 路径
|
||||
|
||||
**结论**:统一用 `filePath` 作去重 key,L0 和 L1 一致。
|
||||
|
||||
---
|
||||
|
||||
## E4: Planner prompt token 增量是否可接受
|
||||
|
||||
**问题**:knowledge map YAML 注入 Planner prompt 会增加固定 token 开销。
|
||||
|
||||
**证据**:
|
||||
- 当前 planner prompt 21 行
|
||||
- 注入 knowledge map 约增加 200-400 字符(6 个文档场景)
|
||||
- 相比 Planner 整体 prompt + 历史消息,增量占比 < 5%
|
||||
|
||||
**结论**:可接受,不构成性能瓶颈。
|
||||
|
||||
---
|
||||
|
||||
## E5: EvaluationService.tool_call_count 影响
|
||||
|
||||
**问题**:去重后 `tool_call_count` 降低,是否影响 `EvaluationService` 评分逻辑。
|
||||
|
||||
**证据**:
|
||||
- `EvaluationService` 使用 `tool_call_count` 作为评分因子
|
||||
- 去重导致重复调用被过滤,`tool_call_count` 下降
|
||||
- 这是修复效果(消除了无意义的重复调用),不是回归
|
||||
|
||||
**结论**:`EvaluationService` 评分规则无需改动。下降的 `tool_call_count` 反映了真实效率提升。
|
||||
|
||||
---
|
||||
|
||||
## P1: 手写 JSON 解析器脆弱性
|
||||
|
||||
**问题**:`KnowledgeIndexService.extractJsonValue` / `extractJsonArray` 在遇到含逗号、引号的自然语言字段时会截断。
|
||||
|
||||
**证据**:
|
||||
- `whenToRetrieve` 字段由 LLM 生成,内容为自然语言(含逗号、分号等标点)
|
||||
- 手写解析器以 `"` 和 `,` 作分隔符,自然语言中的标点会导致提前截断
|
||||
- Jackson `ObjectMapper.readValue(metadata, Frontmatter.class)` 是项目已有依赖
|
||||
|
||||
**结论**:全量替换为 Jackson,影响范围仅 `KnowledgeIndexService.parseDocumentToEntry()`,行为更健壮。
|
||||
|
||||
---
|
||||
|
||||
## P2: LookupResult 去重提示字段
|
||||
|
||||
**问题**:去重命中时如何向 LLM 返回"不要重试"的信号。
|
||||
|
||||
**证据**:
|
||||
- 复用 `primary.content` 语义不清,LLM 可能理解为正常检索结果
|
||||
- 独立 `message` 字段 + `found=false` 语义明确,LLM 能理解"已检索过"不再重试
|
||||
|
||||
**结论**:`LookupResult` 新增 `String message` 字段,去重时填入提示文本。
|
||||
@@ -0,0 +1,70 @@
|
||||
# Acceptance: executor-action-memory-relevance
|
||||
|
||||
## 分档
|
||||
|
||||
standard
|
||||
|
||||
## 任务完成状态
|
||||
|
||||
| 任务 | 状态 | 说明 |
|
||||
|------|------|------|
|
||||
| T1: RetrievedDocTracker 域级升级 | ✅ 完成 | 双层 Map 结构,域级+文档级记录 |
|
||||
| T2: LookupResult 新增字段 | ✅ 完成 | relevanceLevel / completenessHint / retrievedDomainsThisSession |
|
||||
| T3: 归一化计算逻辑 | ✅ 完成 | Min-Max 归一化 + 三等级判定 |
|
||||
| T4: LookupKnowledgeTool 集成 | ✅ 完成 | 归一化层 + 行动记忆注入 + 域拦截 |
|
||||
| T5: Executor Prompt 重写 | ✅ 完成 | 4 条检索约束,无 knowledge map |
|
||||
| T6: 入库可观测性 | ✅ 完成 | V010 + Entity + JSON 扩展 |
|
||||
| T7: BGE-M3 归一化验证测试 | ✅ 完成 | 范数=1.00000002,测试通过 |
|
||||
|
||||
## 静态验证
|
||||
|
||||
- [x] **语法/编译检查**: 所有 Java 文件编译通过
|
||||
- [x] **Impact Analysis**: LookupKnowledgeTool、RetrievedDocTracker 变更范围经 `gitnexus_impact` 检查,均为 L2 内部接口影响
|
||||
- [x] **Cross-artifact 对齐检查**: brief → proposal → design → specs → tasks 闭环,无 gap
|
||||
- [x] **Prompt 约束检查**: chat-executor-prompt.md 不包含 knowledge map,包含 4 条检索约束
|
||||
|
||||
## 脚本验证
|
||||
|
||||
- [x] **V010 Flyway 迁移**: 迁移成功,`relevance_level` 和 `dedup_reason` 列已添加
|
||||
```sql
|
||||
ALTER TABLE tool_invocation
|
||||
ADD COLUMN relevance_level VARCHAR(20),
|
||||
ADD COLUMN dedup_reason VARCHAR(32);
|
||||
```
|
||||
- [x] **FullPipelineSmokeTest**: BGE-M3 归一化测试通过(范数=1.00000002)
|
||||
- [x] **数据库数据校验**:
|
||||
- `relevance_level` 列已写入 HIGHLY_RELEVANT / REFERENCE
|
||||
- `dedup_reason` 列已写入 doc_retrieved / null
|
||||
- `retrieval_details` JSON 包含 l1_top_similarity、completeness_hint、retrieved_domains、dedup_reason
|
||||
|
||||
## 浏览器/人工验证
|
||||
|
||||
- [x] **应用启动验证**: Spring Boot 应用正常启动,端口 9900
|
||||
- [x] **Chat API 调用验证**: 通过 curl 测试 chat 接口,lookup_knowledge 调用链完整
|
||||
```
|
||||
curl -X POST "http://localhost:9900/api/chat/send" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"sessionId": "b66d799e", "question": "..."}'
|
||||
```
|
||||
- [x] **日志验证**: 应用日志可观察到 relevanceLevel、retrievedDomainsThisSession 输出
|
||||
- [x] **归一化数学验证**: l1_top_score=0.383 → l1_top_similarity=0.8085(`1 - 0.383/2.0 = 0.8085`)✅
|
||||
- [x] **域追踪验证**: `[infrastructure]` → `[infrastructure, api]` 域列表正常扩展
|
||||
|
||||
## 未验证
|
||||
|
||||
| 场景 | 原因 | 风险 | 补验建议 |
|
||||
|------|------|------|---------|
|
||||
| PRECISE 等级(L0 唯一精确匹配) | 测试会话无精确匹配场景 | 低 — L0 matchCount=1 的判断逻辑与 HIGHLY_RELEVANT 共用,实现确定性强 | 构造一条 L0 精确匹配的知识库文档后测试 |
|
||||
| domain_retrieved 域级去重 | 需要同一域全部文档已检索再查该域才触发 | 低 — isDomainRetrieved 逻辑简单,与 isDocRetrieved 等价 | Phase 2 启用域级硬限流时测试 |
|
||||
| DEDUPED 等级 | 当前 code path 去重时仍写 REFERENCE,DEDUPED 未被使用 | 低 — 设计预留,当前未启用 | Phase 2 若启用 DEDUPED 等级时验证 |
|
||||
| Phase 2 域级硬限流 | 非本次范围 | 中 — 当前仅有软约束(prompt),LLM 仍可能在 REFERENCE 下继续检索 | 实测观察,如果 lookup 调用仍偏高,启动 Phase 2 |
|
||||
|
||||
## 剩余风险
|
||||
|
||||
1. **Prompt 软约束局限性**:实测 10 次调用中 9 次为 REFERENCE,说明 LLM 仍倾向于继续检索。如果 prompt 约束效果不足,需启用 Phase 2 域级硬限流。
|
||||
2. **L1 Metadata 解析兼容性**:L1 domain 兜底路径解析 metadata JSON,如果知识库文档 frontmatter 格式不一致可能解析失败,已有 try-catch 兜底。
|
||||
|
||||
## 归档状态
|
||||
|
||||
- [ ] OpenSpec change 尚未归档
|
||||
- [ ] devflow/index.md 状态为 `implemented`,待改为 `archived`
|
||||
@@ -0,0 +1,35 @@
|
||||
# Brief: executor-action-memory-relevance
|
||||
|
||||
## 背景
|
||||
|
||||
ISS-002:Executor 在单次会话中调用 `lookup_knowledge` 20+ 次,大部分是同域换变体的冗余调用。前序 change `session-dedup-knowledge-map` 解决了文档级重复召回(ISS-001),但未解决 Executor 重复调用问题。
|
||||
|
||||
## 目标
|
||||
|
||||
- Executor 获得行动记忆(知道自己本次会话已检索了哪些域)
|
||||
- 检索结果提供归一化质量等级(PRECISE/HIGHLY_RELEVANT/REFERENCE)+ 兜底信号
|
||||
- Executor prompt 提供明确的检索约束和"放弃检索"的合法出口
|
||||
- 原始分数入库保留可观测性,但不暴露给 LLM
|
||||
|
||||
## 范围
|
||||
|
||||
- `RetrievedDocTracker`:域级 + 文档级双层记录
|
||||
- `LookupKnowledgeTool`:归一化层 + 行动记忆注入
|
||||
- `LookupResult`:新增 relevanceLevel / completenessHint / retrievedDomainsThisSession
|
||||
- `chat-executor-prompt.md`:检索约束重写
|
||||
- `ToolInvocation` + V010:入库可观测性
|
||||
|
||||
## 非目标
|
||||
|
||||
- 不给 Executor 注入 knowledge map(保持 Agent 边界)
|
||||
- 不修改 Planner prompt 或 Planner 逻辑
|
||||
- 不修改 PrimaryResult / SupplementResult 的字段(不暴露原始分数)
|
||||
- Phase 2 域级硬限制暂不实施
|
||||
|
||||
## 分档
|
||||
|
||||
standard
|
||||
|
||||
## 关联 OpenSpec change
|
||||
|
||||
openspec/changes/executor-action-memory-relevance
|
||||
@@ -0,0 +1,83 @@
|
||||
# Decisions: executor-action-memory-relevance
|
||||
|
||||
## 过程日志
|
||||
|
||||
### Clarify 阶段
|
||||
|
||||
**入口摘要**:ISS-002 Executor 无约束重复调用 lookup_knowledge(单会话 20+ 次),需要行动记忆 + 归一化质量等级 + prompt 约束来解决。
|
||||
|
||||
**slug**: `executor-action-memory-relevance`
|
||||
|
||||
**规模分档**: `standard`(涉及 7 个文件,跨 DTO/工具层/持久化/Prompt,有设计决策需澄清)
|
||||
|
||||
### Context 阶段
|
||||
|
||||
**devflow/index.md 使用状态**: 已命中。前序 change `session-dedup-knowledge-map`(archived)提供了 RetrievedDocTracker、KnowledgeDomainService、ISS-002 文档。
|
||||
|
||||
**相关 ADR**: 无直接 ADR,但 `session-dedup-knowledge-map` 的 decisions.md 和 evidence.md 记录了文档级去重和 knowledge map 注入的决策。
|
||||
|
||||
**不能违反的历史决策**:
|
||||
1. RetrievedDocTracker 的文档级去重必须保留
|
||||
2. knowledge map 只注入 Planner,不注入 Executor(本次讨论确认)
|
||||
3. L0/L1 原始分数不暴露给 LLM,只在归一化层内部使用(本次讨论确认)
|
||||
|
||||
**需进入 OpenSpec 的上下文点**:
|
||||
1. L1 score 是 L2 距离(值域 [0,+∞)),不是归一化分数——阈值设计需基于实际分布
|
||||
2. L0 的 category 可从 KnowledgeEntry.getCategory() 直接获取;L1 需解析 metadata JSON
|
||||
3. ReactAgent 是自主决策工具调用的 Agent,Prompt 约束是软约束
|
||||
|
||||
### Grill 阶段 — Question Pool
|
||||
|
||||
**维度:术语**
|
||||
1. [evidence-driven] `relevanceLevel` 三个等级(PRECISE/HIGHLY_RELEVANT/REFERENCE)的边界是否清晰,是否存在 LLM 误解的可能? → **已查证**:三个等级语义明确,PRECISE=唯一匹配、HIGHLY_RELEVANT=高分命中、REFERENCE=低置信度参考。LLM 理解风险低。
|
||||
|
||||
**维度:边界**
|
||||
2. [evidence-driven] L1 score 是 L2 距离(值域 [0,+∞)),当前代码无阈值判断。归一化阈值如何设计? → **已查证**:L2 距离典型范围取决于 BGE-M3 1024 维 embedding 的尺度,需从 `tool_invocation.retrieval_details` 中查询实际 `l1_scores` 分布才能定阈值。当前先以常量定义,标记为"需实测校准"。
|
||||
3. [evidence-driven] L1 结果的 category 提取需要解析 metadata JSON 字符串,当前 `SearchResult.metadata` 是 `toString()` 的结果。归一化层是否需要 L1 的 domain? → **已查证**:L1 的 domain 主要用于 RetrievedDocTracker 的域级记录。如果 L0 已命中且包含 category,可直接用 L0 的 category;如果仅 L1 命中,需解析 metadata 提取 category。当前知识库中 L0 大概率先命中,L1 domain 提取作为兜底路径。
|
||||
4. [user-interview] 归一化阈值(L1 score 分界线)在实测数据不足时,是否接受先用保守初始值 + 后续调优的策略? → **用户待确认**
|
||||
|
||||
**维度:验收**
|
||||
5. [evidence-driven] 现有 `tool_invocation` 表 `retrieval_details` JSON 中 `l1_scores` 存的是 L2 距离原始值,新增的 `relevance_level` 和 `completeness_hint` 入库后是否需要回填历史数据? → **已查证**:不需要回填历史数据,新列 nullable 即可,历史记录 relevance_level=null。
|
||||
|
||||
### Grill 结论
|
||||
|
||||
**evidence-driven 汇报**:
|
||||
- E1: relevanceLevel 三等级语义清晰,LLM 误解风险低
|
||||
- E2: L1 score 是 L2 距离,值域不固定,阈值需实测校准
|
||||
- E3: L0 category 直接可用,L1 category 需解析 metadata(兜底路径)
|
||||
- E4: 历史数据不回填,新列 nullable
|
||||
|
||||
**user-interview 已确认**:
|
||||
- Q4: 归一化阈值先用保守初始值 + 后续调优 → **用户已确认**,并建议用 Min-Max 归一化到 [0,1]
|
||||
|
||||
### Specify 阶段补充
|
||||
|
||||
**BGE-M3 L2 归一化实测验证**:
|
||||
- FullPipelineSmokeTest.embeddingBgeM3Works() 新增 L2 范数断言
|
||||
- 结果:范数=1.00000002,误差 < 0.01,测试通过
|
||||
- 结论:BGE-M3 输出为 L2 归一化单位向量,L2 距离数学硬上界 = 2.0
|
||||
- Min-Max 归一化公式:`similarity = 1 - min(l2Score, 2.0) / 2.0`
|
||||
|
||||
**Cross-artifact 对齐检查**:
|
||||
|
||||
| 对齐项 | 状态 |
|
||||
|--------|------|
|
||||
| brief 目标/范围/非目标 → proposal 覆盖 | 已对齐 |
|
||||
| proposal 范围/约束 → design 覆盖 | 已对齐 |
|
||||
| design 归一化/行动记忆/接口影响 → specs 覆盖 | 已对齐 |
|
||||
| specs 可观察行为 → tasks 覆盖 | 已对齐 |
|
||||
|
||||
**接口影响分级**:
|
||||
- RetrievedDocTracker 数据结构升级 → L2(内部接口,消费者只有 LookupKnowledgeTool)
|
||||
- LookupResult 新增 3 字段 → L2(工具返回值,无跨模块调用方)
|
||||
- tool_invocation 新增 2 列 → L2(Flyway nullable,不影响现有查询)
|
||||
- chat-executor-prompt.md 更新 → L1(Prompt 文本变更)
|
||||
|
||||
### Audit 阶段
|
||||
|
||||
**架构风险评估**(5 句以内):
|
||||
1. 归一化层嵌入 LookupKnowledgeTool 内部(静态方法),无跨模块耦合风险。
|
||||
2. RetrievedDocTracker 升级为双层结构,数据量级不变(文档数 × session 数),内存无风险。
|
||||
3. L1 metadata 解析 category 是兜底路径,如果 JSON 格式不一致可能解析失败——已有 try-catch 兜底。
|
||||
4. 归一化阈值 yml 配置化,运行时调优不需要改代码和重启——运维友好。
|
||||
5. Prompt 约束仍依赖 LLM 遵守——如果 Phase 1 效果不足,Phase 2 域级硬限制的 isDomainRetrieved 已就绪,无需额外改造。
|
||||
@@ -0,0 +1,65 @@
|
||||
# Evidence: executor-action-memory-relevance
|
||||
|
||||
## Evidence-driven 结论
|
||||
|
||||
### E1: relevanceLevel 三等级语义清晰度
|
||||
|
||||
- **来源**: Grill 阶段 Question Pool #1
|
||||
- **查证结果**: 三个等级语义明确,边界清晰:
|
||||
- PRECISE:L0 唯一精确匹配,LLM 应直接使用
|
||||
- HIGHLY_RELEVANT:归一化 similarity ≥ 0.75,高度相关
|
||||
- REFERENCE:归一化 similarity ≥ 0.5,相关参考
|
||||
- **结论**: LLM 误解风险低,语义边界足够清晰
|
||||
|
||||
### E2: L1 Score 值域与归一化阈值
|
||||
|
||||
- **来源**: Grill 阶段 Question Pool #2
|
||||
- **查证结果**:
|
||||
- L1 score 是 L2 距离,值域 [0, +∞)
|
||||
- BGE-M3 输出为 L2 归一化单位向量(实测范数=1.00000002),L2 距离数学硬上界 = 2.0
|
||||
- Min-Max 归一化公式:`similarity = 1 - min(l2Score, 2.0) / 2.0`
|
||||
- **结论**: 使用 `maxL2Distance=2.0` 作为归一化上界,阈值 yml 可配置
|
||||
|
||||
### E3: L1 Domain 提取兜底路径
|
||||
|
||||
- **来源**: Grill 阶段 Question Pool #3
|
||||
- **查证结果**:
|
||||
- L0 的 domain 可从 `KnowledgeEntry.getCategory()` 直接获取
|
||||
- L1 结果的 domain 需解析 `SearchResult.metadata` JSON 字符串
|
||||
- 当前知识库设计下 L0 大概率先命中,L1 domain 提取作为兜底
|
||||
- **结论**: 先尝试 L0 category,失败时解析 L1 metadata JSON(try-catch 兜底)
|
||||
|
||||
### E4: 历史数据不回填
|
||||
|
||||
- **来源**: Grill 阶段 Question Pool #5
|
||||
- **查证结果**: 新列 `relevance_level` 和 `dedup_reason` 均为 nullable,不影响现有查询
|
||||
- **结论**: 历史记录保持 null,不需要回填迁移
|
||||
|
||||
### E5: BGE-M3 L2 归一化实测验证
|
||||
|
||||
- **来源**: Specify 阶段 + FullPipelineSmokeTest
|
||||
- **查证结果**:
|
||||
- embeddingBgeM3Works() 测试新增 L2 范数断言
|
||||
- 实测范数 = 1.00000002,误差 < 0.01
|
||||
- 测试通过,BGE-M3 输出确认为 L2 归一化单位向量
|
||||
- **结论**: L2 距离上界 = 2.0 的数学依据成立
|
||||
|
||||
### E6: V010 迁移验证
|
||||
|
||||
- **来源**: Apply 阶段运行时验证
|
||||
- **查证结果**:
|
||||
- Flyway V010 迁移成功执行
|
||||
- `relevance_level` VARCHAR(20) 列可空,已正确写入
|
||||
- `dedup_reason` VARCHAR(32) 列可空,已正确写入
|
||||
- `retrieval_details` JSON 扩展字段(l1_top_similarity、relevance_level、completeness_hint、retrieved_domains、dedup_reason)全部写入
|
||||
- **结论**: 入库可观测性符合设计
|
||||
|
||||
### E7: 数据库数据校验
|
||||
|
||||
- **来源**: Apply 阶段运行时验证
|
||||
- **查证结果**:
|
||||
- session `b66d799e` 共 10 条 lookup_knowledge 调用
|
||||
- id=138: L2=0.383 → similarity=0.8085 → HIGHLY_RELEVANT(符合预期)
|
||||
- id=139-147: 主要为 REFERENCE,doc_retrieved 去重正常触发
|
||||
- retrieved_domains 域追踪:`[infrastructure]` → `[infrastructure, api]` 正常扩展
|
||||
- **结论**: 归一化、行动记忆、去重机制数据层面全部验证通过
|
||||
@@ -0,0 +1,58 @@
|
||||
# Acceptance: chat-verifier-agent
|
||||
|
||||
## Classification
|
||||
|
||||
standard
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Verifier prompt | Done | Strict JSON schema, verdict matrix, fact classifications, and `evidence_refs` are defined. |
|
||||
| VerifierInputHook | Done | Explicit verifier payload replaces raw conversation history. |
|
||||
| ChatService integration | Done | Planner, executor, and verifier are called explicitly with max two rounds. |
|
||||
| Verdict routing | Done | PASS, LOW_CONFID, and REJECT paths are handled in code. |
|
||||
| Trace summary | Done | Evidence summaries include `trace_ref` and `source_invocation_ids`. |
|
||||
| self_evaluation merge | Done | `rule_evaluation` and `verifier_evaluation` are preserved independently. |
|
||||
| Verifier observability | Done | `verifier_evaluation` persists facts, evidence refs, trace summary, rationale, score, and round. |
|
||||
|
||||
## Static Verification
|
||||
|
||||
- [x] OpenSpec artifacts exist: `proposal.md`, `design.md`, `specs/chat-verifier-agent/spec.md`, `tasks.md`, `.committed`.
|
||||
- [x] `change.json` exists and has `metadata.status = committed`.
|
||||
- [x] `.archive-ready` exists.
|
||||
- [x] devflow archive-prep files exist: `brief.md`, `evidence.md`, `decisions.md`, `acceptance.md`.
|
||||
- [x] `devflow/index.md` contains `chat-verifier-agent` with status `archived`.
|
||||
|
||||
## Script Verification
|
||||
|
||||
- [x] `mvn -q -DskipTests compile` passed.
|
||||
|
||||
## Runtime Verification
|
||||
|
||||
- [x] POST `/api/chat` with a complex question returned successfully.
|
||||
- [x] Runtime session `9138f064` showed planner, executor, and verifier execution in logs.
|
||||
- [x] Runtime session `9138f064` wrote `verifier_evaluation.verdict = LOW_CONFID`.
|
||||
- [x] Runtime session `9138f064` wrote `facts_checked[*].evidence_refs`.
|
||||
- [x] Runtime session `9138f064` wrote `tool_trace_summary[*].source_invocation_ids`.
|
||||
- [x] LOW_CONFID final answer included disclaimer and verifier-derived evidence gaps.
|
||||
|
||||
## Unverified
|
||||
|
||||
| Scenario | Reason | Risk | Follow-up |
|
||||
| --- | --- | --- | --- |
|
||||
| PASS runtime path | The exercised complex runtime case produced LOW_CONFID. | Low; PASS routing is simple pass-through after parsed verifier decision. | Add a fixture or deterministic verifier test if this becomes product-critical. |
|
||||
| REJECT runtime path | No forced contradiction case was run after traceability changes. | Medium; REJECT is the safety-critical degraded path. | Add a targeted test with a fabricated claim and evidence contradiction. |
|
||||
| Document-path-level evidence mapping | Current implementation records invocation ids and source document labels, not guaranteed canonical document paths for every retrieval mode. | Low for current audit need; medium for future UI drill-down. | Extend retrieval details with canonical document paths in a later change. |
|
||||
|
||||
## Remaining Risks
|
||||
|
||||
1. Verifier output still depends on model compliance with JSON schema; code falls back to LOW_CONFID on missing or invalid output.
|
||||
2. `AgentLoggingHook` is shared by several agent paths; current changes preserve compile and runtime behavior but should be watched in AiOps flows.
|
||||
3. `SupervisorAgent` construction remains as legacy residue in `ChatService`; runtime orchestration is explicit, but a later cleanup should remove unused supervisor construction.
|
||||
|
||||
## Archive State
|
||||
|
||||
- [x] OpenSpec change is archive-ready.
|
||||
- [x] OpenSpec change has been moved to `openspec/changes/archive/2026-07-03-chat-verifier-agent/`.
|
||||
- [x] Main spec exists at `openspec/specs/chat-verifier-agent/spec.md`.
|
||||
@@ -0,0 +1,36 @@
|
||||
# Brief: chat-verifier-agent
|
||||
|
||||
## Background
|
||||
|
||||
The complex Chat path previously returned Executor answers without a synchronous quality gate. Existing rule scoring was asynchronous and post-hoc, so it could not prevent unsupported answers from reaching users.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Add a Verifier Agent after Executor in the complex chat path.
|
||||
2. Require structured verifier output with `PASS`, `LOW_CONFID`, or `REJECT`.
|
||||
3. Route final user output in code based on verifier verdict.
|
||||
4. Persist verifier results under `diagnosis_session.self_evaluation.verifier_evaluation`.
|
||||
5. Preserve rule scoring under `rule_evaluation`.
|
||||
6. Make verifier decisions traceable to real tool invocations through `evidence_refs` and `source_invocation_ids`.
|
||||
|
||||
## Scope
|
||||
|
||||
- `ChatService`: explicit `planner -> executor -> verifier` orchestration, max two rounds, verdict routing, retry context, verifier persistence.
|
||||
- `VerifierInputHook`: explicit verifier input payload.
|
||||
- `ToolTraceSummaryService`: evidence summary from persisted tool calls.
|
||||
- `VerifierContextHolder`: round-local verifier context.
|
||||
- `SelfEvaluationMergeService`: safe JSON merge for evaluation channels.
|
||||
- `AgentLoggingHook`: concise verifier thought and fuller structured output retention.
|
||||
- `chat-verifier-prompt.md`: verifier contract, verdict matrix, and traceability schema.
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- Verifier does not call tools.
|
||||
- Verifier does not rewrite Executor output.
|
||||
- Single-agent chat path remains outside this change.
|
||||
- No database schema migration is included.
|
||||
- Document-path-level evidence attribution is deferred; current traceability is invocation-level with source document labels.
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/archive/2026-07-03-chat-verifier-agent/`
|
||||
@@ -0,0 +1,123 @@
|
||||
# Decisions: chat-verifier-agent
|
||||
|
||||
## 过程日志
|
||||
|
||||
### Clarify 阶段
|
||||
|
||||
**入口摘要**: 在 Chat 多 Agent 链路中新增 Verifier Agent,作为 Executor 输出后的质量门禁,做事实核查。
|
||||
|
||||
**slug**: `chat-verifier-agent`
|
||||
|
||||
**规模分档**: standard
|
||||
|
||||
### Context 阶段
|
||||
|
||||
**devflow/index.md 使用状态**: 已命中。前序 change `executor-action-memory-relevance`(archived)提供了 Chat 多 Agent 当前链路(Supervisor → Planner → Executor)。
|
||||
|
||||
**不能违反的历史决策**:
|
||||
1. Executor 已有完整的行动记忆和归一化质量等级,Verifier 不需要重复验证检索质量
|
||||
2. Chat Supervisor 的职责是调度,Verifier 作为子 Agent 加入后不改变 Supervisor 的定位
|
||||
3. 已有 evidence_score 做事后评分,Verifier 是事前门禁,两者不冲突
|
||||
|
||||
**需进入 OpenSpec 的上下文点**:
|
||||
1. Verifier 不需要工具调用,只是一个质量核查 Agent
|
||||
2. Verifier 需要访问 Executor 的输出 + 工具调用记录
|
||||
3. Supervisor prompt 需要重写以包含 Verifier 调度规则
|
||||
4. groundedness_score 的阈值需要在代码中定义
|
||||
|
||||
### Grill 阶段 — Question Pool
|
||||
|
||||
| # | 维度 | 问题 | 模式 | 状态 |
|
||||
|---|------|------|------|------|
|
||||
| Q1 | 术语 | evidence_score(事后评分)与 Verifier(事前门禁)职责是否冲突? | evidence-driven | 已解决 |
|
||||
| Q2 | 边界 | Verifier 需要的"工具调用记录"在 SupervisorAgent 中是否自动传递? | evidence-driven | 已解决 |
|
||||
| Q3 | 边界 | LOW_CONFID < 0.5 回调 Planner 后的新输出是否再次走 Verifier?循环上限多少? | user-interview | 已解决 |
|
||||
| Q4 | 验收 | Verifier 判决结果如何可观测?是否写入 agent_step 或 tool_invocation? | user-interview | 已解决 |
|
||||
| Q5 | 验收 | 当前 Supervisor 硬编码 prompt 是否支持多 Agent 路由变更? | evidence-driven | 已解决 |
|
||||
| Q6 | 技术 | Verifier 如何隔离 Executor 的中间推理过程,只看到干净的 query + tool 记录 + 最终答案? | user-interview | 已解决 |
|
||||
| Q7 | 验收 | groundedness_score 阈值(0.5)是否需要配置化? | user-interview | 已解决 |
|
||||
|
||||
### Evidence-driven 结论
|
||||
|
||||
| 结论 | 证据来源 | 是否已汇报用户 |
|
||||
|------|---------|-------------|
|
||||
| evidence_score(异步事后)与 Verifier(同步事前门禁)不冲突 | EvaluationService.java: @Async 注解 | 已汇报 |
|
||||
| SupervisorAgent 自动传递完整对话状态,Verifier 无需额外传递工具记录 | Spring AI Alibaba SupervisorAgent 实现 | 已汇报 |
|
||||
| Supervisor prompt 为字符串字面量,直接修改即可 | ChatService.java:353 .systemPrompt("...") | 已汇报 |
|
||||
|
||||
### User-interview 记录
|
||||
|
||||
| 问题 | 用户原话 | 确认状态 | OpenSpec 回写 |
|
||||
|------|---------|---------|-------------|
|
||||
| Q3: LOW_CONFID < 0.5 回调 Planner 循环上限? | "可以,回调一次" | 已确认 | 已回写 proposal |
|
||||
| Q4: Verifier 判决写入哪里做可观测? | "可以"(写入 diagnosis_session.self_evaluation JSON) | 已确认 | 已回写 proposal |
|
||||
| Q6: Verifier 如何隔离 Executor 中间推理? | "用 MessagesModelHook 过滤 messages" | 已确认 | 已回写 design |
|
||||
| Q7: groundedness_score 阈值是否需要配置化? | "需要配置化" | 已确认 | 已回写 design |
|
||||
|
||||
### Specify 阶段 — Cross-Artifact 对齐检查
|
||||
|
||||
| 上游 → 下游 | 检查内容 | 状态 |
|
||||
|---|---|---|
|
||||
| proposal → design | 范围、约束、关键承诺是否进入 design | 已对齐 |
|
||||
| design → specs | 关键决策、模块地图是否进入 specs | 已对齐 |
|
||||
| specs → tasks | 可观察行为是否被 tasks 覆盖为可执行切片 | 已对齐 |
|
||||
|
||||
**接口影响分级**:
|
||||
- buildChatVerifierAgent() 新增方法 → L1(内部方法,无外部消费者)
|
||||
- VerifierInputHook 类 → L1(内部 Hook,无外部消费者)
|
||||
- Supervisor prompt 重写 → L1(仅影响 Chat 多 Agent 内部调度)
|
||||
- subAgents 列表变更 → L1(Supervisor 内部配置)
|
||||
- verifier.low-confidence-threshold 配置 → L1(新增配置项,不改已有配置)
|
||||
|
||||
### Audit 阶段
|
||||
|
||||
**模块链路**:
|
||||
|
||||
```
|
||||
用户 → Supervisor → Planner(步骤) → Executor(答案+工具记录)
|
||||
│
|
||||
Supervisor 调用 Verifier
|
||||
│
|
||||
[VerifierInputHook BEFORE_MODEL]
|
||||
├─ 保留:system prompt + user query
|
||||
├─ 保留:tool call 记录(输入+返回)
|
||||
├─ 保留:Executor 最终答案
|
||||
└─ 去除:Executor 中间推理、Planner 规划过程
|
||||
│
|
||||
Verifier 判决
|
||||
│
|
||||
┌─── PASS ───→ 直接输出
|
||||
├─── LOW_CONFID≥0.5 → 带声明输出
|
||||
├─── LOW_CONFID<0.5 → 回调 Planner(一次)
|
||||
└─── REJECT → 降级输出
|
||||
│
|
||||
写入 self_evaluation JSON
|
||||
```
|
||||
|
||||
**架构风险评估**(5 句以内):
|
||||
1. Verifier 是轻量 Agent(无工具、无外部依赖),架构风险低。
|
||||
2. MessagesModelHook 纯过滤逻辑,不引入新数据源。
|
||||
3. LOW_CONFID 分级处理 + 回调仅一次的设计,避免无限循环风险。
|
||||
4. REJECT 降级确保编造内容不到达用户。
|
||||
5. 审计结论不影响现有 design/tasks,无需回写。
|
||||
|
||||
### 关键取舍
|
||||
|
||||
- 决策:LOW_CONFID < 0.5 回调 Planner 一次
|
||||
- 原因:给系统一次修正机会,但避免无限循环
|
||||
- 影响:Supervisor prompt 需维护"已回调"状态
|
||||
- 风险接受:用户已确认
|
||||
|
||||
- 决策:Verifier 判决写入 diagnosis_session.self_evaluation JSON
|
||||
- 原因:不改表结构,与 evidence_score 统一可观测体系
|
||||
- 影响:ChatService 后处理需追加 JSON
|
||||
- 风险接受:用户已确认
|
||||
|
||||
### Archive-Ready Update
|
||||
|
||||
- 实现调整:最终运行链路由 `ChatService` 显式调用 `planner -> executor -> verifier`,不再依赖 Supervisor prompt 保证 verifier 被调用。
|
||||
- 可追溯性补充:`tool_trace_summary` 增加 `trace_ref`、`source_invocation_ids`、查询样本、检索层级、相关性等级和来源文档标签。
|
||||
- 可追溯性补充:`facts_checked[*].evidence_refs` 被 prompt 要求、代码解析并持久化。
|
||||
- 验证记录:`mvn -q -DskipTests compile` 通过。
|
||||
- 验证记录:运行会话 `9138f064` 走通 planner、executor、verifier,并持久化 `verifier_evaluation.facts_checked[*].evidence_refs` 与 `tool_trace_summary[*].source_invocation_ids`。
|
||||
- 当前状态:OpenSpec change 已归档到 `openspec/changes/archive/2026-07-03-chat-verifier-agent/`,主规格已同步到 `openspec/specs/chat-verifier-agent/spec.md`。
|
||||
@@ -0,0 +1,52 @@
|
||||
# Evidence: chat-verifier-agent
|
||||
|
||||
## Code Evidence
|
||||
|
||||
### Complex chat path now invokes verifier deterministically
|
||||
|
||||
- File: `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- Evidence: `executeChatComplex` calls planner, executor, then verifier directly through `callAgent(...)`.
|
||||
- Conclusion: runtime no longer depends on prompt-only Supervisor behavior to call verifier.
|
||||
|
||||
### Verifier receives explicit inputs
|
||||
|
||||
- File: `src/main/java/com/superbiz/agent/hook/VerifierInputHook.java`
|
||||
- Evidence: the hook builds a JSON payload with `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
|
||||
- Conclusion: verifier input is stable and does not depend on guessing the last assistant message from raw history.
|
||||
|
||||
### Tool evidence is traceable to persisted invocations
|
||||
|
||||
- File: `src/main/java/com/superbiz/agent/service/ToolTraceSummaryService.java`
|
||||
- Evidence: summaries include `trace_ref`, `source_invocation_ids`, `query_samples`, `retrieval_layers`, `relevance_levels`, and `source_documents`.
|
||||
- Conclusion: verifier facts can be correlated with actual `tool_invocation` rows.
|
||||
|
||||
### Verifier facts preserve evidence references
|
||||
|
||||
- File: `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- Evidence: verifier parsing preserves `facts_checked[*].evidence_refs` and persists `tool_trace_summary` under `verifier_evaluation`.
|
||||
- Conclusion: `self_evaluation` now contains both verifier judgments and the evidence index used to form them.
|
||||
|
||||
### Evaluation channels no longer overwrite each other
|
||||
|
||||
- File: `src/main/java/com/superbiz/agent/service/SelfEvaluationMergeService.java`
|
||||
- Evidence: rule and verifier evaluations are merged into separate keys.
|
||||
- Conclusion: asynchronous rule scoring preserves verifier output.
|
||||
|
||||
### Verifier logging is less noisy
|
||||
|
||||
- File: `src/main/java/com/superbiz/agent/hook/AgentLoggingHook.java`
|
||||
- Evidence: verifier `thought` stores a concise verdict summary, while fuller model output remains available in structured storage.
|
||||
- Conclusion: `agent_step.thought` is no longer a misleading place for full verifier JSON.
|
||||
|
||||
## Runtime Evidence
|
||||
|
||||
- Compile verification passed: `mvn -q -DskipTests compile`.
|
||||
- Runtime session `9138f064` executed `planner -> executor -> verifier`.
|
||||
- Runtime session `9138f064` persisted `verifier_evaluation.facts_checked[*].evidence_refs`.
|
||||
- Runtime session `9138f064` persisted `verifier_evaluation.tool_trace_summary[*].source_invocation_ids`.
|
||||
|
||||
## Design Evidence
|
||||
|
||||
- `LOW_CONFID` returns a fixed disclaimer and verifier-derived gaps.
|
||||
- `REJECT` returns degraded output and does not pass through the raw Executor answer.
|
||||
- `retry_context` is derived from verifier-identified missing evidence facts.
|
||||
@@ -0,0 +1,65 @@
|
||||
# MVP Demo Trace Acceptance
|
||||
|
||||
## Result
|
||||
|
||||
Accepted for implementation scope.
|
||||
|
||||
## Verification
|
||||
|
||||
### Static Verification
|
||||
|
||||
- Command: `mvn -q -DskipTests compile`
|
||||
- Result: passed
|
||||
- Notes: New trace controller, service, DTO, profile, verifier fallback, and test sources compile with the project.
|
||||
|
||||
### Script Verification
|
||||
|
||||
- Command: `mvn -q "-Dtest=DiagnosisTraceServiceTest,ChatServiceSupervisorAgentTest" test`
|
||||
- Result: passed
|
||||
- Notes: Covers successful trace aggregation, missing-session 404 path via `SessionNotFoundException`, low-confidence no-retry behavior, method-tool injection, and verifier fallback when Supervisor skips `chat_verifier`.
|
||||
|
||||
### OpenSpec Verification
|
||||
|
||||
- Command: `openspec validate mvp-demo-trace-acceptance --strict`
|
||||
- Result: passed
|
||||
|
||||
### GitNexus Verification
|
||||
|
||||
- Result: skipped by user decision
|
||||
- Notes: User requested subsequent project flow to bypass GitNexus.
|
||||
|
||||
### Manual / Runtime Verification
|
||||
|
||||
- Steps: Follow `mvp/demo/README.md` with `--spring.profiles.active=mvp-demo`.
|
||||
- Result: passed
|
||||
- Notes:
|
||||
- Session `mvp-demo-payment-timeout-20260703-rerun2` completed as `SUCCESS`.
|
||||
- Chat request returned `code=200`, `success=true`, and the same `sessionId`.
|
||||
- Chat duration was `96316 ms`; persisted session duration was `95028 ms`.
|
||||
- Trace API returned `code=200`, `returnedSteps=13`, `returnedTools=12`, `hasVerifier=true`, and `verifierVerdict=LOW_CONFID`.
|
||||
- Trace agents included `planner,executor,verifier`.
|
||||
- Trace tools included `lookup_knowledge,query_logs,query_metrics`.
|
||||
- Feedback submission returned success, and a follow-up trace query showed `feedback=useful`.
|
||||
- MySQL verification confirmed `agent_step` count `13` with agents `executor,planner,verifier`.
|
||||
- MySQL verification confirmed `tool_invocation` count `12` with tools `lookup_knowledge,query_logs,query_metrics`.
|
||||
|
||||
## Completed Scope
|
||||
|
||||
- Added `GET /api/diagnosis/{sessionId}/trace`.
|
||||
- Added read-only trace aggregation from persisted diagnosis tables.
|
||||
- Added `mvp-demo` profile overlay.
|
||||
- Added payment-timeout demo acceptance documentation.
|
||||
- Added MVP note for interview storytelling.
|
||||
- Added verifier fallback so runtime trace remains complete when Supervisor returns without `verifier_output`.
|
||||
|
||||
## Known Limits
|
||||
|
||||
- `mvp-demo` is not a fully offline mock runtime.
|
||||
- Runtime still depends on available MySQL, Redis, Milvus/Zilliz, model, and embedding configuration.
|
||||
- Sensitive configuration cleanup remains intentionally deferred.
|
||||
- Supervisor can still make inefficient routing choices inside a single round; `ChatService` now invokes `chat_verifier` as a fallback when Supervisor returns without `verifier_output`, so trace completeness is preserved for the MVP demo.
|
||||
|
||||
## Handoff
|
||||
|
||||
- Runtime demo passed with current infrastructure.
|
||||
- OpenSpec archive confirmation: requested by user after successful rerun.
|
||||
@@ -0,0 +1,35 @@
|
||||
# MVP Demo Trace Acceptance Brief
|
||||
|
||||
## Background
|
||||
|
||||
- User goal: make the MVP runnable, observable, and explainable for an Agent Engineer interview.
|
||||
- Current problem: the system can execute diagnosis, but reviewers need a simple way to replay one session from final answer back to agent steps and tool evidence.
|
||||
- Associated OpenSpec: `openspec/changes/mvp-demo-trace-acceptance/`
|
||||
- Devflow scale: standard-light.
|
||||
|
||||
## Scope
|
||||
|
||||
- In scope:
|
||||
- `mvp-demo` Spring profile overlay.
|
||||
- `GET /api/diagnosis/{sessionId}/trace` read-only API.
|
||||
- Trace aggregation DTO/service/controller.
|
||||
- Focused service tests.
|
||||
- Demo and acceptance documentation.
|
||||
- Out of scope:
|
||||
- Sensitive configuration cleanup.
|
||||
- Full offline LLM/vector/database mock runtime.
|
||||
- Database schema migration.
|
||||
- Changes to chat execution, verifier routing, upload, or feedback behavior.
|
||||
- Impact area:
|
||||
- `src/main/java/com/superbiz/agent/controller`
|
||||
- `src/main/java/com/superbiz/agent/service`
|
||||
- `src/main/java/com/superbiz/agent/dto`
|
||||
- `src/main/resources/application-mvp-demo.yml`
|
||||
- `mvp/demo`
|
||||
- `mvp/notes`
|
||||
|
||||
## OpenSpec Alignment
|
||||
|
||||
- proposal coverage: covered
|
||||
- specs coverage: covered
|
||||
- tasks coverage: covered
|
||||
@@ -0,0 +1,87 @@
|
||||
# MVP Demo Trace Acceptance Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: continue the MVP toward a runnable and explainable demo by adding an `mvp-demo` profile, an end-to-end acceptance case, and a trace query API.
|
||||
- Slug: `mvp-demo-trace-acceptance`
|
||||
- Devflow scale: standard-light. The change adds a public read-only API and documentation, but does not alter core chat execution or persistence schemas.
|
||||
|
||||
## Context
|
||||
|
||||
- `devflow/index.md` was checked. Relevant history includes `session-storage`, `confidence-feedback`, `executor-action-memory-relevance`, and `chat-verifier-agent`.
|
||||
- `mvp/notes/agent-engineering-decisions.md` already recommends the next phase as "可复现 MVP Demo", including `mvp-demo` profile, fixed diagnosis case, one-click request, and `GET /api/diagnosis/{sessionId}/trace`.
|
||||
- `mvp/issues/ISS-003-mvp-design-implementation-review.md` identifies test stability, session traceability, verifier evidence chain, upload path, and SupervisorAgent consistency as recent MVP concerns. Security cleanup is intentionally deferred by user decision.
|
||||
|
||||
## Question Pool
|
||||
|
||||
| # | Dimension | Question | Mode | Status |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | Terminology | Should "trace" mean persisted diagnosis execution evidence instead of transient frontend chat history? | evidence-driven | Resolved |
|
||||
| Q2 | Boundary | Should this change modify chat execution or only expose existing persisted evidence? | evidence-driven | Resolved |
|
||||
| Q3 | Acceptance | What proves the MVP flow is end-to-end enough for demo/interview use? | evidence-driven | Resolved |
|
||||
| Q4 | Interface | What is the API impact level for `GET /api/diagnosis/{sessionId}/trace`? | evidence-driven | Resolved |
|
||||
|
||||
## Evidence-driven
|
||||
|
||||
| Conclusion | Evidence Source | Reported To User |
|
||||
|---|---|---|
|
||||
| Trace should aggregate persisted diagnosis evidence, not Redis-only chat history. | `DiagnosisSession`, `AgentStep`, `ToolInvocation` entities and repositories | Reported in progress update |
|
||||
| Core chat execution does not need to change for this slice. | Existing unified chat path and SupervisorAgent commits; requested scope is demo/profile/trace/acceptance | Reported in progress update |
|
||||
| End-to-end acceptance should cover start -> chat -> trace -> feedback. | `ChatController`, `FeedbackController`, traceable session id decision in MVP notes | Reported in progress update |
|
||||
| Trace API is additive L3 because it is a new HTTP API for frontend/demo consumers. | sm-flow interface impact rules | Recorded in OpenSpec design |
|
||||
|
||||
## User-interview
|
||||
|
||||
| Question | User Words | Confirmation | OpenSpec Writeback |
|
||||
|---|---|---|---|
|
||||
| Should security/sensitive config cleanup be included? | "安全问题先不考虑"; "敏感配置先不做" | Confirmed | Non-goal |
|
||||
| Should this be implemented under sm-flow? | "按照 sm-flow 的流程来实现吧" | Confirmed | This change follows sm-flow artifacts |
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Add a new trace API instead of embedding trace details in `/api/chat`.
|
||||
- Reason: Chat execution and observability should stay decoupled.
|
||||
- Impact: Demo can query trace after any successful chat request using the same session id.
|
||||
- Risk accepted: Response shape is new and should be treated as demo-facing contract.
|
||||
|
||||
- Decision: Keep `mvp-demo` profile as configuration overlay, not a fully mocked standalone runtime.
|
||||
- Reason: The current MVP still depends on real DB/Redis/Milvus/LLM for full chat execution; this change avoids inventing a fake runtime that hides integration behavior.
|
||||
- Impact: Demo profile improves repeatability for logs/metrics, while docs remain explicit about required external services.
|
||||
- Risk accepted: End-to-end acceptance may still require valid infrastructure and keys.
|
||||
|
||||
## Cross-Artifact Alignment
|
||||
|
||||
| Upstream -> Downstream | Check | Status |
|
||||
|---|---|---|
|
||||
| brief/prd -> proposal | Goal, scope, non-goals, and acceptance expectation are in proposal | Aligned |
|
||||
| proposal -> design | Scope, constraints, and API impact are in design | Aligned |
|
||||
| design -> specs/tasks | Trace DTO, controller/service, demo profile, and docs are represented | Aligned |
|
||||
| specs -> tasks | Observable behavior is covered by executable tasks | Aligned |
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
- Data path: HTTP trace request -> controller -> trace service -> repositories -> aggregate DTO -> `Result.success`.
|
||||
- The service is read-only and does not mutate diagnosis, step, tool, or feedback state.
|
||||
- No schema change is needed because all required fields already exist in `diagnosis_session`, `agent_step`, and `tool_invocation`.
|
||||
- Main risk is response size for large sessions; MVP mitigates by returning previews already persisted by tools rather than raw external logs.
|
||||
- The additive API is acceptable for MVP because old callers remain unaffected.
|
||||
|
||||
## Pre-apply Research
|
||||
|
||||
- Reference implementations read:
|
||||
- `ChatController` for `/api` controller conventions.
|
||||
- `FeedbackController` for simple API controller shape.
|
||||
- `GlobalExceptionHandler` and `SessionNotFoundException` for 404 handling.
|
||||
- `DiagnosisSessionRepository`, `AgentStepRepository`, `ToolInvocationRepository` for available queries.
|
||||
- `DiagnosisSession`, `AgentStep`, `ToolInvocation` for fields.
|
||||
- Impact analysis:
|
||||
- `DiagnosisSessionRepository`: LOW, direct imports in service/controller paths.
|
||||
- `AgentStepRepository`: HIGH because it participates in chat/AiOps flows. This change only consumes existing query methods and does not modify the repository.
|
||||
- `ToolInvocationRepository`: LOW.
|
||||
|
||||
## Commit Gate
|
||||
|
||||
- OpenSpec proposal/design/specs/tasks exist.
|
||||
- API impact: L3 additive collaboration API, documented in design and spec.
|
||||
- User-confirmed non-goal: sensitive configuration cleanup remains out of scope.
|
||||
- No unresolved user-interview questions remain for this slice.
|
||||
@@ -0,0 +1,25 @@
|
||||
# MVP Demo Trace Acceptance Evidence
|
||||
|
||||
## Evidence
|
||||
|
||||
| Source | Evidence | Conclusion | Reported |
|
||||
|---|---|---|---|
|
||||
| `DiagnosisSessionRepository` | Existing `findBySessionId(String)` query | Trace can locate the session without new repository methods | Yes |
|
||||
| `AgentStepRepository` | Existing `findBySessionIdOrderByStepIndex(String)` query | Agent steps can be returned in execution order | Yes |
|
||||
| `ToolInvocationRepository` | Existing `findBySessionIdOrderByIdAsc(String)` query | Tool evidence can be returned in persisted order | Yes |
|
||||
| `GlobalExceptionHandler` | Handles `SessionNotFoundException` as HTTP 404 with `Result.error(404, ...)` | Missing trace can reuse existing error contract | Yes |
|
||||
| `mvn -q "-Dtest=DiagnosisTraceServiceTest" test` | Command passed | Trace aggregation behavior is covered offline | Yes |
|
||||
| `mvn -q -DskipTests compile` | Command passed | New code compiles with the full project | Yes |
|
||||
| `gitnexus detect-changes --repo SuperBizAgent-java` | Command completed with `No changes detected` and line-ending warnings | Required GitNexus check ran; output likely does not capture newly added files | Yes |
|
||||
|
||||
## Evidence-driven Conclusions
|
||||
|
||||
- Conclusion: No database migration is required.
|
||||
- Evidence: All trace fields are available from existing `diagnosis_session`, `agent_step`, and `tool_invocation` entities.
|
||||
- Risk: Response shape becomes a new API contract.
|
||||
- User confirmation: Not required; additive L3 API recorded in OpenSpec.
|
||||
|
||||
- Conclusion: Trace aggregation can be tested without external infrastructure.
|
||||
- Evidence: `DiagnosisTraceServiceTest` uses mocked repositories and an `ObjectMapper`.
|
||||
- Risk: Runtime integration still depends on configured infrastructure.
|
||||
- User confirmation: Not required; limitation recorded in acceptance docs.
|
||||
@@ -0,0 +1,14 @@
|
||||
# Acceptance: aiops-alert-scope-control
|
||||
|
||||
## Verification
|
||||
|
||||
- [x] Payload-mode prompt focuses the final report on the supplied alert.
|
||||
- [x] No-payload prompt requires active-alert discovery first.
|
||||
- [x] Targeted tests pass.
|
||||
- [x] Compile passes.
|
||||
- [x] OpenSpec validates.
|
||||
|
||||
## Known Limits
|
||||
|
||||
- Prompt-only scope control may still require runtime observation.
|
||||
- AIOps Verifier remains deferred.
|
||||
@@ -0,0 +1,28 @@
|
||||
# Brief: aiops-alert-scope-control
|
||||
|
||||
## Background
|
||||
|
||||
After `aiops-traceable-diagnosis-entry`, AIOps can be triggered by payload and replayed through trace. Runtime verification showed one semantic gap: payload mode still produced a broad report over all active mock alerts.
|
||||
|
||||
## Goal
|
||||
|
||||
Make AIOps scope explicit:
|
||||
|
||||
- Payload present -> targeted diagnosis for the supplied alert.
|
||||
- Payload absent -> automatic active-alert discovery and diagnosis.
|
||||
|
||||
## Scope
|
||||
|
||||
- In scope:
|
||||
- `AiOpsService.buildTaskPrompt(...)` scope rules.
|
||||
- Focused tests.
|
||||
- Demo acceptance wording.
|
||||
- Out of scope:
|
||||
- Verifier integration.
|
||||
- Java-side filtering of tool results.
|
||||
- API shape changes.
|
||||
- Database changes.
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/aiops-alert-scope-control/`
|
||||
@@ -0,0 +1,42 @@
|
||||
# Decisions: aiops-alert-scope-control
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: tighten AIOps report scope after runtime verification showed payload mode still analyzes all active alerts.
|
||||
- Slug: `aiops-alert-scope-control`
|
||||
- Scale: standard-light.
|
||||
|
||||
## Context
|
||||
|
||||
- AIOps traceability is implemented and verified.
|
||||
- Mock Prometheus returns multiple active alerts.
|
||||
- Payload demo supplies `HighCPUUsage/payment-service`, but previous report expanded to `HighMemoryUsage` and `SlowResponse`.
|
||||
|
||||
## Grill Question Pool
|
||||
|
||||
| # | Dimension | Question | Mode | Status |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | Product Boundary | What makes `/api/ai_ops` different from `/api/chat` when payload exists? | evidence-driven | Payload is alert-event driven and should be scoped to that event. |
|
||||
| Q2 | Scope | Should payload mode ignore all other active alerts? | user-interview | No; mention only as related risk/context. |
|
||||
| Q3 | Compatibility | Should no-payload mode keep old "query active alerts" behavior? | evidence-driven | Yes. |
|
||||
| Q4 | Enforcement | Should Java filter unrelated tool results now? | evidence-driven | No; prompt-only is sufficient for this small change. |
|
||||
| Q5 | Verifier | Should this change add AIOps Verifier? | user-interview | No; keep deferred. |
|
||||
|
||||
## Evidence-Driven Conclusions
|
||||
|
||||
| Conclusion | Evidence Source | Result |
|
||||
|---|---|---|
|
||||
| Scope issue is prompt-level. | `/api_ ai_ops` trace showed all mock alerts analyzed despite payload. | Update task prompt. |
|
||||
| No API or persistence changes are needed. | `AIOpsRequest` already carries payload and trace works. | Keep endpoint unchanged. |
|
||||
| Blast radius is low. | `buildTaskPrompt(...)` is internal to `AiOpsService`. | Add tests for prompt content. |
|
||||
|
||||
## GitNexus
|
||||
|
||||
GitNexus remains skipped by prior user decision and because tools are not exposed in this session. Local impact analysis is recorded instead.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Payload mode is detected when any alert field is present.
|
||||
- Payload mode final report must focus on the supplied alert.
|
||||
- No-payload mode must first call `queryPrometheusAlerts`.
|
||||
- Other active alerts in payload mode can appear only as related risk, not as separate root-cause sections.
|
||||
@@ -0,0 +1,44 @@
|
||||
# Evidence: aiops-alert-scope-control
|
||||
|
||||
## Local Impact Analysis
|
||||
|
||||
- `AiOpsService.buildTaskPrompt(...)` is used by `executeAiOpsAnalysis(...)`.
|
||||
- No controller, DTO, repository, or database changes are required.
|
||||
- Existing `AiOpsServiceTest` already exercises request summary helpers and can be extended for scope prompt rules.
|
||||
|
||||
## Verification Results
|
||||
|
||||
- `mvn -q "-Dtest=AiOpsServiceTest" test` passed.
|
||||
- `mvn -q -DskipTests compile` passed.
|
||||
- `openspec.cmd validate aiops-alert-scope-control --strict` passed.
|
||||
|
||||
## Runtime Verification
|
||||
|
||||
- Runtime session: `mvp-demo-aiops-payment-cpu-codex-scope-003`.
|
||||
- `/api/ai_ops` SSE emitted the requested `session` message and finished with `done`.
|
||||
- `diagnosis_session` persisted:
|
||||
- `agent_flow = AI_OPS`
|
||||
- `status = SUCCESS`
|
||||
- `total_duration_ms = 69875`
|
||||
- `step_count = 5`
|
||||
- `tool_call_count = 8`
|
||||
- Tool invocation counts:
|
||||
- `query_metrics = 1`
|
||||
- `lookup_knowledge = 1`
|
||||
- `query_logs = 6`
|
||||
- Report scope check:
|
||||
- `告警根因分析 - HighCPUUsage` exists.
|
||||
- `告警根因分析 - HighMemoryUsage` does not exist.
|
||||
- `告警根因分析 - SlowResponse` does not exist.
|
||||
- `相关风险告警` exists.
|
||||
|
||||
## Runtime Fix
|
||||
|
||||
- Added Hikari settings in `src/main/resources/application.yml` after the first runtime attempt failed on stale MySQL pool connections:
|
||||
- `maximum-pool-size: 5`
|
||||
- `minimum-idle: 1`
|
||||
- `connection-timeout: 10000`
|
||||
- `validation-timeout: 5000`
|
||||
- `idle-timeout: 60000`
|
||||
- `max-lifetime: 120000`
|
||||
- `keepalive-time: 30000`
|
||||
@@ -0,0 +1,18 @@
|
||||
# Acceptance: aiops-traceable-diagnosis-entry
|
||||
|
||||
## Verification
|
||||
|
||||
- [x] OpenSpec validates for `aiops-traceable-diagnosis-entry`.
|
||||
- [x] Targeted AIOps service tests pass.
|
||||
- [x] Compile verification passes.
|
||||
- [x] Demo docs describe AIOps request -> session id -> trace query.
|
||||
|
||||
## Result
|
||||
|
||||
Accepted for implementation scope.
|
||||
|
||||
## Known Limits
|
||||
|
||||
- AIOps Verifier integration is deferred.
|
||||
- Runtime still depends on configured model and infrastructure.
|
||||
- Full browser/SSE runtime verification is not guaranteed in this coding pass.
|
||||
@@ -0,0 +1,28 @@
|
||||
# Brief: aiops-traceable-diagnosis-entry
|
||||
|
||||
## Background
|
||||
|
||||
The MVP chat diagnosis path is now traceable through `diagnosis_session`, `agent_step`, `tool_invocation`, and `GET /api/diagnosis/{sessionId}/trace`. The older `/api/ai_ops` endpoint still acts like a standalone SSE demo: it accepts no alert payload, generates an internal session id, and does not make trace replay obvious to callers.
|
||||
|
||||
## Goal
|
||||
|
||||
Turn AIOps into an alert-triggered diagnosis entry point that shares the same evidence and trace story as the main MVP, without rewriting the whole AIOps flow.
|
||||
|
||||
## Scope
|
||||
|
||||
- In scope:
|
||||
- Optional AIOps alert request body.
|
||||
- Stable request/session id propagation.
|
||||
- Persisted AIOps query summary and final answer.
|
||||
- SSE session id event.
|
||||
- Demo documentation and focused tests.
|
||||
- Out of scope:
|
||||
- Full AIOps and ChatService unification.
|
||||
- AIOps Verifier integration.
|
||||
- Database schema changes.
|
||||
- Sensitive configuration cleanup.
|
||||
- Fully offline runtime.
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/aiops-traceable-diagnosis-entry/`
|
||||
@@ -0,0 +1,68 @@
|
||||
# Decisions: aiops-traceable-diagnosis-entry
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: make the legacy AIOps SSE endpoint a traceable alert diagnosis entry for the Agent Engineer interview MVP.
|
||||
- Slug: `aiops-traceable-diagnosis-entry`
|
||||
- Scale: standard-light, because this extends one public endpoint and reuses existing persistence/trace infrastructure.
|
||||
|
||||
## Context
|
||||
|
||||
- `mvp-demo-trace-acceptance` already added `GET /api/diagnosis/{sessionId}/trace`.
|
||||
- `chat-verifier-agent` made the chat path stronger than the older AIOps path.
|
||||
- Current AIOps value is as a second entry point: system alert -> automated diagnosis -> evidence trace.
|
||||
|
||||
## Grill Question Pool
|
||||
|
||||
| # | Dimension | Question | Mode | Status |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | Positioning | Is AIOps an independent product path or an alert-triggered sibling of Chat Diagnosis? | user-interview | Resolved: sibling entry, unified trace story |
|
||||
| Q2 | API | Should we keep `/api/ai_ops` or add a new endpoint? | evidence-driven | Resolved: keep existing endpoint and extend optional body |
|
||||
| Q3 | Input | What is the minimum alert payload? | user-interview | Resolved: `sessionId`, `alertName`, `service`, `severity`, `description`, `timeRange`, plus `userRequest` fallback |
|
||||
| Q4 | Output | How does the caller learn the trace session id? | evidence-driven | Resolved: first SSE event uses type `session` |
|
||||
| Q5 | Trace | Must AIOps be replayable with existing trace API? | evidence-driven | Resolved: yes, this is the main acceptance criterion |
|
||||
| Q6 | Verifier | Must this slice add AIOps Verifier? | user-interview | Resolved: no, defer as follow-up |
|
||||
| Q7 | Compatibility | Should no-body calls still work? | evidence-driven | Resolved: yes, preserve old demo behavior |
|
||||
| Q8 | GitNexus | Should unavailable GitNexus block implementation? | user-interview | Resolved: skip GitNexus by user decision |
|
||||
|
||||
## Evidence-Driven Conclusions
|
||||
|
||||
| Conclusion | Evidence Source | Result |
|
||||
|---|---|---|
|
||||
| AIOps is currently isolated from request-driven trace replay. | `ChatController.aiOps()` has no request body; `AiOpsService` creates its own random session id. | Extend endpoint and service. |
|
||||
| No schema change is needed. | `DiagnosisSession` already has `query`, `agentFlow`, `answer`, counts, and status. | Reuse existing table. |
|
||||
| Trace API can already replay AIOps if session id and answer are persisted. | `DiagnosisTraceService` loads by session id and is flow-agnostic. | Keep trace API unchanged. |
|
||||
| Blast radius is moderate and local. | `rg` shows only `ChatController` calls `executeAiOpsAnalysis` and `extractFinalReport`. | Change service/controller carefully and add tests. |
|
||||
|
||||
## User-Interview Confirmations
|
||||
|
||||
| Topic | User Words | Decision |
|
||||
|---|---|---|
|
||||
| Use sm-flow | "可以,改造一下AIOps 接口,用sm-flow流程看看" | Use OpenSpec + devflow. |
|
||||
| GitNexus | "跳过gitnexus把" | Record skip and use local impact analysis. |
|
||||
| Proceed after Grill | "可以" | Continue with lightweight Grill conclusions. |
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Keep `/api/ai_ops` and make its body optional.
|
||||
- Emit `SseMessage.type=session` before long-running analysis starts.
|
||||
- Store AIOps request summary in `diagnosis_session.query`.
|
||||
- Store final report in `diagnosis_session.answer`.
|
||||
- Defer AIOps Verifier to a later change so this slice stays focused.
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
```text
|
||||
POST /api/ai_ops
|
||||
-> optional AIOpsRequest
|
||||
-> resolve sessionId
|
||||
-> create diagnosis_session(agentFlow=AI_OPS)
|
||||
-> run ai_ops_supervisor(planner, executor)
|
||||
-> AgentLoggingHook persists steps
|
||||
-> tools persist invocations under SessionContextHolder
|
||||
-> extract final report
|
||||
-> persist answer
|
||||
-> GET /api/diagnosis/{sessionId}/trace replays the run
|
||||
```
|
||||
|
||||
Risk level: medium. The endpoint is public and SSE-based, but the change is additive and does not change the chat diagnosis path or database schema.
|
||||
@@ -0,0 +1,37 @@
|
||||
# Evidence: aiops-traceable-diagnosis-entry
|
||||
|
||||
## Local Impact Analysis
|
||||
|
||||
- `ChatController.aiOps()` is the only caller of `AiOpsService.executeAiOpsAnalysis(...)`.
|
||||
- `ChatController.aiOps()` is the only caller of `AiOpsService.extractFinalReport(...)`.
|
||||
- `AIOpsRequest` exists but only has `userRequest`; no current controller consumes it.
|
||||
- `DiagnosisTraceService` is flow-agnostic and reads persisted session/step/tool records by `sessionId`.
|
||||
|
||||
## GitNexus
|
||||
|
||||
GitNexus MCP tools were not exposed in this session. The user explicitly approved skipping GitNexus for this change. Local impact analysis and targeted tests are used instead.
|
||||
|
||||
## Expected Verification
|
||||
|
||||
- Focused unit tests for AIOps request/session/report helper behavior.
|
||||
- Compile verification.
|
||||
- OpenSpec validation if CLI is available.
|
||||
|
||||
## Verification Results
|
||||
|
||||
- `openspec.cmd validate aiops-traceable-diagnosis-entry --strict`: passed.
|
||||
- `mvn -q "-Dtest=AiOpsServiceTest,DiagnosisTraceServiceTest" test`: passed after rerun with approved Maven access.
|
||||
- `mvn -q -DskipTests compile`: passed.
|
||||
|
||||
## Demo Alignment
|
||||
|
||||
- Added `knowledge_base/troubleshooting/aiops-alert-runbook.md` so mock AIOps alerts have matching knowledge-base guidance.
|
||||
- Aligned the documented AIOps demo with mock data: `HighCPUUsage` on `payment-service`, using `system-metrics` evidence.
|
||||
|
||||
## Metric Alignment Follow-up
|
||||
|
||||
- Runtime verification showed `diagnosis_session.tool_call_count` counted agent steps with tool calls, while trace returned actual `tool_invocation` records.
|
||||
- Updated `ChatService` and `AiOpsService` metric backfill to use `ToolInvocationRepository.countBySessionId(sessionId)`.
|
||||
- Targeted verification:
|
||||
- `mvn -q "-Dtest=AiOpsServiceTest,ChatServiceSequentialAgentTest,DiagnosisTraceServiceTest" test`: passed.
|
||||
- `mvn -q -DskipTests compile`: passed.
|
||||
@@ -0,0 +1,36 @@
|
||||
# Acceptance: diagnosis-eval-harness
|
||||
|
||||
## Classification
|
||||
|
||||
standard-light
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Issue and OpenSpec setup | Done | `ISS-006` and initial OpenSpec artifacts were created. |
|
||||
| Implementation | Done | Added fixed cases, fixture-mode trace evaluation, aggregate metrics, and JSON / Markdown report writer. |
|
||||
| Verification | Done | Targeted evaluator tests, compile verification, and OpenSpec validation passed. |
|
||||
|
||||
## Current State
|
||||
|
||||
- First implementation uses fixture-mode evaluation.
|
||||
- Live trace API polling remains a follow-up option.
|
||||
|
||||
## Verification
|
||||
|
||||
### Script Verification
|
||||
|
||||
- Command: `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
|
||||
- Result: passed
|
||||
- Notes: Covers fixed case loading, fixture evaluation, missing fixture reporting, reject degraded-output validation, and report writing.
|
||||
|
||||
### Static Verification
|
||||
|
||||
- Command: `mvn -q -DskipTests compile`
|
||||
- Result: passed
|
||||
|
||||
### OpenSpec Verification
|
||||
|
||||
- Command: `openspec validate diagnosis-eval-harness --strict`
|
||||
- Result: passed
|
||||
@@ -0,0 +1,31 @@
|
||||
# Brief: diagnosis-eval-harness
|
||||
|
||||
## Background
|
||||
|
||||
The MVP has a runnable demo and hardened evidence trace semantics, but it still lacks a fixed regression baseline for Agent diagnosis quality. P1-B creates a small evaluation harness that can validate diagnosis traces against fixed cases and produce repeatable reports.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Define fixed diagnosis cases for the MVP demo domain.
|
||||
2. Validate trace evidence, verifier verdicts, answer keywords, and degraded-output behavior.
|
||||
3. Produce JSON and Markdown reports for interview and regression use.
|
||||
4. Keep the first version offline by supporting trace fixtures.
|
||||
|
||||
## Scope
|
||||
|
||||
- Evaluation case definitions
|
||||
- Trace fixture shape
|
||||
- Rule-based evaluator
|
||||
- JSON / Markdown report output
|
||||
- Focused offline tests and docs
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No LLM-as-judge
|
||||
- No live end-to-end runtime requirement
|
||||
- No production API
|
||||
- No chat or verifier runtime change
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/diagnosis-eval-harness/`
|
||||
@@ -0,0 +1,28 @@
|
||||
# Diagnosis Eval Harness Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: build P1-B fixed case evaluation after evidence trace hardening.
|
||||
- Slug: `diagnosis-eval-harness`
|
||||
- Devflow scale: standard-light
|
||||
|
||||
## Context
|
||||
|
||||
- P1-A `evidence-trace-hardening` created stable evidence semantics for supported, no-evidence, deduped, and failed tool calls.
|
||||
- The MVP demo trace API already provides an aggregate trace shape suitable for evaluation.
|
||||
- The first evaluator should avoid depending on external infrastructure so it can run in regular development.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Start with rule-based trace validation instead of LLM-as-judge.
|
||||
- Reason: The first regression signal should be deterministic and tied to trace contracts.
|
||||
|
||||
- Decision: Support offline fixture traces first.
|
||||
- Reason: This makes the harness usable without MySQL, Redis, Milvus, or a real LLM.
|
||||
|
||||
- Decision: Output both JSON and Markdown.
|
||||
- Reason: JSON supports automation; Markdown is easier to discuss in interviews.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Whether live trace API polling belongs in this change or a follow-up after fixture mode lands.
|
||||
@@ -0,0 +1,10 @@
|
||||
# Diagnosis Eval Harness Evidence
|
||||
|
||||
## Evidence
|
||||
|
||||
| Source | Evidence | Conclusion | Reported |
|
||||
|---|---|---|---|
|
||||
| `openspec/specs/evidence-trace-hardening/spec.md` | Defines stable evidence states and summary behavior | Evaluation can rely on trace semantics rather than ad hoc log parsing | Yes |
|
||||
| `mvp/demo/README.md` | Documents an end-to-end demo flow with chat, trace, and feedback | Existing demo flow provides the runtime story, but not a reusable evaluation baseline | Yes |
|
||||
| `DiagnosisTraceService` | Aggregates session, steps, tools, and self-evaluation | Trace response shape can be reused as evaluation input | Yes |
|
||||
| `ToolTraceSummaryService` | Builds verifier-facing evidence summaries from persisted tool rows | Evaluator can check evidence coverage through persisted trace artifacts | Yes |
|
||||
@@ -0,0 +1,32 @@
|
||||
# Acceptance: evidence-trace-hardening
|
||||
|
||||
## Classification
|
||||
|
||||
standard-light
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Issue and OpenSpec setup | Done | `ISS-005` and the initial OpenSpec artifacts were created. |
|
||||
| Implementation | Done | Recorder contract, lookup persistence path, evidence summary semantics, and degraded-path tests were implemented. |
|
||||
| Verification | Done | Targeted offline tests and compile verification passed. |
|
||||
|
||||
## Verification
|
||||
|
||||
### Script Verification
|
||||
|
||||
- Command: `mvn -q "-Dtest=ToolInvocationRecorderTest,ToolTraceSummaryServiceTest,ChatServiceSequentialAgentTest,LookupKnowledgeToolTest" test`
|
||||
- Result: passed
|
||||
- Notes: Covers recorder contract, summary semantics for success/failure/no-evidence, and `ChatService` fallback / degraded paths.
|
||||
|
||||
### Static Verification
|
||||
|
||||
- Command: `mvn -q -DskipTests compile`
|
||||
- Result: passed
|
||||
|
||||
## Open Questions
|
||||
|
||||
| Question | Current position |
|
||||
| --- | --- |
|
||||
| Should deduped retrievals be counted separately from generic no-hit events in future evaluation metrics? | Deferred to P1-B; this change preserves enough structure to decide later. |
|
||||
@@ -0,0 +1,32 @@
|
||||
# Brief: evidence-trace-hardening
|
||||
|
||||
## Background
|
||||
|
||||
The MVP already has persisted tool traces and a verifier, but the evidence contract is still only partially standardized. For interview-focused hardening, the project now needs a tighter contract for evidence persistence, no-evidence / failure semantics, and degraded-output behavior.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Standardize the persisted evidence-tool contract across `lookup_knowledge`, `query_logs`, and `query_metrics`.
|
||||
2. Make verifier-facing summaries distinguish failed calls, no-hit calls, deduped retrievals, and actual supporting evidence.
|
||||
3. Add offline tests for verifier fallback and degraded-output paths.
|
||||
|
||||
## Scope
|
||||
|
||||
- `ToolInvocationRecorder`
|
||||
- `LookupKnowledgeTool`
|
||||
- `QueryLogsTools`
|
||||
- `QueryMetricsTools`
|
||||
- `ToolTraceSummaryService`
|
||||
- `ChatService`
|
||||
- Focused offline tests
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No new API or schema
|
||||
- No evaluation harness yet
|
||||
- No trace UI
|
||||
- No security/config cleanup
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/evidence-trace-hardening/`
|
||||
@@ -0,0 +1,24 @@
|
||||
# Evidence Trace Hardening Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: harden the MVP evidence contract before building the P1-B evaluation harness.
|
||||
- Slug: `evidence-trace-hardening`
|
||||
- Devflow scale: standard-light
|
||||
|
||||
## Context
|
||||
|
||||
- `ISS-003` raised verifier traceability and failure-path concerns.
|
||||
- Current code inspection shows `QueryLogsTools` and `QueryMetricsTools` already use `ToolInvocationRecorder`, while `LookupKnowledgeTool` still persists rows through a local helper.
|
||||
- `ChatService` already contains fallback behavior for missing/invalid `verifier_output`, but coverage is narrow.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Treat this as a contract-hardening change, not a new feature change.
|
||||
- Reason: The project already has the necessary runtime pieces; the gap is semantic consistency and testability.
|
||||
|
||||
- Decision: Keep the scope before P1-B.
|
||||
- Reason: The evaluation harness will rely on stable evidence semantics, so this contract slice should land first.
|
||||
|
||||
- Decision: Preserve schema and API stability.
|
||||
- Reason: The interview value here is engineering rigor, not more surface area.
|
||||
@@ -0,0 +1,11 @@
|
||||
# Evidence Trace Hardening Evidence
|
||||
|
||||
## Evidence
|
||||
|
||||
| Source | Evidence | Conclusion | Reported |
|
||||
|---|---|---|---|
|
||||
| `ToolInvocationRecorder` | Provides a common persistence seam for evidence tools | Contract hardening should build on the existing recorder instead of introducing a new store path | Yes |
|
||||
| `LookupKnowledgeTool` | Still constructs `ToolInvocation` rows through a local helper | Retrieval-aware evidence persistence is not yet unified with the recorder contract | Yes |
|
||||
| `QueryLogsTools` / `QueryMetricsTools` | Already record evidence invocations through `recordEvidenceTool(...)` | Current gap is semantic alignment, not missing persistence | Yes |
|
||||
| `ToolTraceSummaryService` | Merges rows by tool and topic domain and infers evidence level heuristically | Summary rules need explicit handling for failure, no-hit, and dedup cases | Yes |
|
||||
| `ChatService` | Falls back to `LOW_CONFID` when verifier output is missing or invalid | These degraded paths exist and should now be covered by focused offline tests | Yes |
|
||||
@@ -0,0 +1,37 @@
|
||||
# Acceptance: expand-diagnosis-eval-fixtures
|
||||
|
||||
## Classification
|
||||
|
||||
standard-light
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Issue and OpenSpec setup | Done | Created slug-based issue and initial OpenSpec artifacts. |
|
||||
| Implementation | Done | Added remaining fixtures, full baseline reports, and documentation updates. |
|
||||
| Verification | Done | Evaluator tests, compile verification, and OpenSpec validation passed. |
|
||||
|
||||
## Current State
|
||||
|
||||
- Fixture coverage is complete for the five fixed diagnosis cases.
|
||||
- Baseline reports are saved under `mvp/eval/reports`.
|
||||
- No production runtime behavior has been changed.
|
||||
|
||||
## Verification
|
||||
|
||||
### Script Verification
|
||||
|
||||
- Command: `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
|
||||
- Result: passed
|
||||
- Notes: Covers full fixture coverage, baseline report matching, reject degraded-output validation, and report writing.
|
||||
|
||||
### Static Verification
|
||||
|
||||
- Command: `mvn -q -DskipTests compile`
|
||||
- Result: passed
|
||||
|
||||
### OpenSpec Verification
|
||||
|
||||
- Command: `openspec validate expand-diagnosis-eval-fixtures --strict`
|
||||
- Result: passed
|
||||
@@ -0,0 +1,31 @@
|
||||
# Brief: expand-diagnosis-eval-fixtures
|
||||
|
||||
## Background
|
||||
|
||||
The diagnosis eval harness is implemented and archived, but the fixed baseline is incomplete because three of the five diagnosis cases still reference missing fixtures.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Add representative trace fixtures for all remaining fixed diagnosis cases.
|
||||
2. Save a reproducible baseline report in JSON and Markdown.
|
||||
3. Document how to regenerate and interpret the baseline.
|
||||
4. Keep evaluation offline and deterministic.
|
||||
|
||||
## Scope
|
||||
|
||||
- Redis timeout fixture
|
||||
- Slow response fixture
|
||||
- JVM memory risk fixture
|
||||
- Baseline reports under `mvp/eval/reports`
|
||||
- Focused tests for full fixture coverage and report generation
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No new diagnosis cases
|
||||
- No production Agent runtime changes
|
||||
- No LLM-as-judge
|
||||
- No live infrastructure requirement
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/expand-diagnosis-eval-fixtures/`
|
||||
@@ -0,0 +1,28 @@
|
||||
# Expand Diagnosis Eval Fixtures Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: complete the fixed diagnosis eval baseline after the harness is in place.
|
||||
- Slug: `expand-diagnosis-eval-fixtures`
|
||||
- Devflow scale: standard-light
|
||||
|
||||
## Context
|
||||
|
||||
- `diagnosis-eval-harness` created the evaluator, case file, fixture mode, and report writer.
|
||||
- The first baseline still has missing fixtures by design.
|
||||
- This follow-up turns that partial baseline into a full fixed-case baseline.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Keep this change data-focused.
|
||||
- Reason: the evaluator rules already landed; this change should not blur fixture expansion with harness behavior changes.
|
||||
|
||||
- Decision: Save baseline reports in the repository.
|
||||
- Reason: interview review and future diffs are easier when the expected baseline is visible.
|
||||
|
||||
- Decision: Use deterministic fixture traces instead of live trace generation.
|
||||
- Reason: this baseline should run without infrastructure or external model calls.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Whether a future change should add a CLI or Maven goal for report regeneration.
|
||||
@@ -0,0 +1,11 @@
|
||||
# Evidence: expand-diagnosis-eval-fixtures
|
||||
|
||||
## Evidence Log
|
||||
|
||||
- 2026-07-04: Created slug-based issue `expand-diagnosis-eval-fixtures.md`.
|
||||
- 2026-07-04: Created OpenSpec change `expand-diagnosis-eval-fixtures`.
|
||||
- 2026-07-04: Added Redis timeout, slow response, and JVM memory risk fixtures.
|
||||
- 2026-07-04: Added baseline JSON and Markdown reports under `mvp/eval/reports`.
|
||||
- 2026-07-04: Verification passed with `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`.
|
||||
- 2026-07-04: Verification passed with `mvn -q -DskipTests compile`.
|
||||
- 2026-07-04: Verification passed with `openspec validate expand-diagnosis-eval-fixtures --strict`.
|
||||
@@ -0,0 +1,37 @@
|
||||
# Acceptance: diagnosis-eval-baseline-diff
|
||||
|
||||
## Classification
|
||||
|
||||
standard-light
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Issue and OpenSpec setup | Done | Created slug-based issue and initial OpenSpec artifacts. |
|
||||
| Implementation | Done | Added diff model, comparator, writer, docs, sample outputs, and focused tests. |
|
||||
| Verification | Done | Diff/evaluator tests, compile verification, and OpenSpec validation passed. |
|
||||
|
||||
## Current State
|
||||
|
||||
- Baseline diff is implemented for aggregate metrics, verdict distribution, case-level state, keyword coverage, evidence coverage, missing cases, and new cases.
|
||||
- JSON and Markdown diff output are available.
|
||||
- No production runtime behavior has been changed.
|
||||
|
||||
## Verification
|
||||
|
||||
### Script Verification
|
||||
|
||||
- Command: `mvn -q "-Dtest=DiagnosisEvalBaselineDiffTest" test`
|
||||
- Result: passed
|
||||
- Notes: Also verified with `DiagnosisTraceEvaluatorTest`.
|
||||
|
||||
### Static Verification
|
||||
|
||||
- Command: `mvn -q -DskipTests compile`
|
||||
- Result: passed
|
||||
|
||||
### OpenSpec Verification
|
||||
|
||||
- Command: `openspec validate diagnosis-eval-baseline-diff --strict`
|
||||
- Result: passed
|
||||
@@ -0,0 +1,30 @@
|
||||
# Brief: diagnosis-eval-baseline-diff
|
||||
|
||||
## Background
|
||||
|
||||
The eval harness now has a complete saved baseline. This change adds the comparison layer that turns the baseline into an actionable regression signal.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Compare baseline and current `DiagnosisEvalReport` objects.
|
||||
2. Detect aggregate and per-case regressions.
|
||||
3. Output JSON and Markdown diff reports.
|
||||
4. Document how to read the diff in interview and engineering terms.
|
||||
|
||||
## Scope
|
||||
|
||||
- Diff data structures
|
||||
- Deterministic report comparison
|
||||
- JSON / Markdown diff output
|
||||
- Focused tests and eval docs
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No live Agent execution
|
||||
- No LLM-as-judge
|
||||
- No evaluator scoring rule changes
|
||||
- No production API changes
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/diagnosis-eval-baseline-diff/`
|
||||
@@ -0,0 +1,28 @@
|
||||
# Diagnosis Eval Baseline Diff Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: add report diffing on top of the completed diagnosis eval baseline.
|
||||
- Slug: `diagnosis-eval-baseline-diff`
|
||||
- Devflow scale: standard-light
|
||||
|
||||
## Context
|
||||
|
||||
- `diagnosis-eval-harness` created deterministic fixture evaluation.
|
||||
- `expand-diagnosis-eval-fixtures` created a complete saved baseline.
|
||||
- This change compares new reports against that baseline.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Diff report DTOs instead of raw traces.
|
||||
- Reason: the report is the stable contract for regression review.
|
||||
|
||||
- Decision: Use deterministic code rules instead of LLM-as-judge.
|
||||
- Reason: baseline regression checks should be repeatable and explainable.
|
||||
|
||||
- Decision: Output both JSON and Markdown.
|
||||
- Reason: JSON supports automation; Markdown is useful in reviews and interviews.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Whether a future change should expose this through a CLI or Maven goal.
|
||||
@@ -0,0 +1,11 @@
|
||||
# Evidence: diagnosis-eval-baseline-diff
|
||||
|
||||
## Evidence Log
|
||||
|
||||
- 2026-07-05: Created slug-based issue `diagnosis-eval-baseline-diff.md`.
|
||||
- 2026-07-05: Created OpenSpec change `diagnosis-eval-baseline-diff`.
|
||||
- 2026-07-05: Added baseline diff DTOs, deterministic comparer, and JSON / Markdown writer.
|
||||
- 2026-07-05: Added sample baseline diff JSON and Markdown reports.
|
||||
- 2026-07-05: Verification passed with `mvn -q "-Dtest=DiagnosisEvalBaselineDiffTest,DiagnosisTraceEvaluatorTest" test`.
|
||||
- 2026-07-05: Verification passed with `mvn -q -DskipTests compile`.
|
||||
- 2026-07-05: Verification passed with `openspec validate diagnosis-eval-baseline-diff --strict`.
|
||||
@@ -0,0 +1,25 @@
|
||||
# Acceptance: mvp-demo-interview-runbook
|
||||
|
||||
## Classification
|
||||
|
||||
standard-light
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Issue and OpenSpec setup | Done | Created slug-based issue and OpenSpec artifacts. |
|
||||
| Implementation | Done | Added request payload, runnable script, output directory docs, interview walkthrough, and trace checklist. |
|
||||
| Verification | Done | OpenSpec validation passed. |
|
||||
|
||||
## Current State
|
||||
|
||||
- No backend runtime behavior has been changed.
|
||||
- Demo is packaged under `mvp/demo` for interview use.
|
||||
|
||||
## Verification
|
||||
|
||||
### OpenSpec Verification
|
||||
|
||||
- Command: `openspec validate mvp-demo-interview-runbook --strict`
|
||||
- Result: passed
|
||||
@@ -0,0 +1,29 @@
|
||||
# Brief: mvp-demo-interview-runbook
|
||||
|
||||
## Background
|
||||
|
||||
Plan C is the interview-facing demo package. The project has the engineering pieces, but needs a single place to run and explain the MVP flow.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Provide a fixed payment-timeout request payload.
|
||||
2. Provide a PowerShell script that runs chat, trace, and feedback.
|
||||
3. Save demo responses under `mvp/demo/output`.
|
||||
4. Add interview walkthrough and trace checklist.
|
||||
|
||||
## Scope
|
||||
|
||||
- Demo docs and scripts only
|
||||
- Existing local APIs only
|
||||
- Existing `mvp-demo` profile only
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No backend code changes
|
||||
- No eval extension
|
||||
- No secret cleanup
|
||||
- No full offline runtime
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/mvp-demo-interview-runbook/`
|
||||
@@ -0,0 +1,27 @@
|
||||
# MVP Demo Interview Runbook Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: package existing MVP capabilities into a repeatable interview demo.
|
||||
- Slug: `mvp-demo-interview-runbook`
|
||||
- Devflow scale: standard-light
|
||||
|
||||
## Context
|
||||
|
||||
- Evidence trace and eval baseline work are already done.
|
||||
- The next useful step is not more eval tooling, but a runnable demo path.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Keep this change documentation/script-only.
|
||||
- Reason: Plan C is about demo packaging, not new runtime capability.
|
||||
|
||||
- Decision: Use a stable session id.
|
||||
- Reason: it makes trace lookup and saved output predictable.
|
||||
|
||||
- Decision: Save outputs to `mvp/demo/output`.
|
||||
- Reason: generated artifacts should be easy to review without mixing into source fixtures.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Whether a later change should add a truly offline stubbed demo mode.
|
||||
@@ -0,0 +1,9 @@
|
||||
# Evidence: mvp-demo-interview-runbook
|
||||
|
||||
## Evidence Log
|
||||
|
||||
- 2026-07-05: Created Plan C demo packaging issue and OpenSpec change.
|
||||
- 2026-07-05: Added fixed payment-timeout request payload.
|
||||
- 2026-07-05: Added PowerShell demo script for chat, trace, and feedback.
|
||||
- 2026-07-05: Added interview walkthrough and trace inspection checklist.
|
||||
- 2026-07-05: Verification passed with `openspec validate mvp-demo-interview-runbook --strict`.
|
||||
@@ -0,0 +1,86 @@
|
||||
# RAG Retrieval Baseline
|
||||
|
||||
This directory contains the offline retrieval baseline for the RAG refactor.
|
||||
|
||||
The baseline is intentionally narrower than full diagnosis evaluation. It checks
|
||||
whether fixed retrieval queries can recover expected documents, breadcrumbs, and
|
||||
evidence keywords before changing L0 behavior, query augmentation, evidence
|
||||
post-processing, or Spring AI VectorStore integration.
|
||||
|
||||
## Layout
|
||||
|
||||
```text
|
||||
eval/rag-retrieval/
|
||||
cases/golden-cases.json Fixed retrieval golden cases
|
||||
fixtures/*.json Saved retrieval candidates for each case
|
||||
reports/baseline.json Machine-readable baseline report
|
||||
reports/baseline.md Human-readable baseline report
|
||||
reports/live-post-reindex.* Optional live acceptance reports
|
||||
```
|
||||
|
||||
## Run
|
||||
|
||||
From the repository root:
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_retrieval.py
|
||||
```
|
||||
|
||||
Custom paths are also supported:
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_retrieval.py \
|
||||
--cases eval/rag-retrieval/cases/golden-cases.json \
|
||||
--fixtures eval/rag-retrieval/fixtures \
|
||||
--json-report eval/rag-retrieval/reports/baseline.json \
|
||||
--markdown-report eval/rag-retrieval/reports/baseline.md
|
||||
```
|
||||
|
||||
## Hit Levels
|
||||
|
||||
- `strong`: expected document is found and breadcrumb or evidence keyword coverage is satisfied.
|
||||
- `medium`: expected document is found, but breadcrumb or keyword coverage is incomplete.
|
||||
- `weak`: expected evidence keyword is found, but expected document is missing.
|
||||
- `miss`: expected document and expected evidence are not found.
|
||||
|
||||
`Recall@K` counts `strong` and `medium` as retrieved.
|
||||
|
||||
## Scope
|
||||
|
||||
This baseline runs fully offline and does not call MySQL, Redis, Milvus, an LLM,
|
||||
or the Spring Boot application. It is a regression harness for retrieval behavior,
|
||||
not a claim that live production retrieval accuracy is complete.
|
||||
|
||||
## Live Post-Reindex Acceptance
|
||||
|
||||
When embedding input changes, existing vectors do not update by themselves. For
|
||||
example, after adding `title` and `breadcrumb` to the embedding text, the live
|
||||
Milvus/Zilliz collection must be reindexed before retrieval can reflect that new
|
||||
semantic signal.
|
||||
|
||||
Use this optional live acceptance flow after the application is running and the
|
||||
knowledge base has been reindexed:
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_live_acceptance.py
|
||||
```
|
||||
|
||||
Custom service URL and output paths are supported:
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_live_acceptance.py \
|
||||
--base-url http://127.0.0.1:9900 \
|
||||
--json-report eval/rag-retrieval/reports/live-post-reindex.json \
|
||||
--markdown-report eval/rag-retrieval/reports/live-post-reindex.md
|
||||
```
|
||||
|
||||
The script calls:
|
||||
|
||||
```text
|
||||
GET /api/search/similar
|
||||
```
|
||||
|
||||
It writes JSON and Markdown reports with query, topK, result count, top
|
||||
candidates, breadcrumb, score labels, and raw response fields. This is a live
|
||||
smoke check for environment readiness and post-reindex behavior; it does not
|
||||
replace the deterministic offline baseline above.
|
||||
@@ -0,0 +1,61 @@
|
||||
{
|
||||
"version": 1,
|
||||
"description": "Offline golden retrieval cases for RAG refactor baseline.",
|
||||
"topK": 5,
|
||||
"cases": [
|
||||
{
|
||||
"caseId": "chat-mysql-connection-pool",
|
||||
"scenario": "chat",
|
||||
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"expectedDocIds": ["mysql-connection-pool"],
|
||||
"expectedBreadcrumbs": ["Database > MySQL > Connection Pool"],
|
||||
"expectedKeywords": ["connection pool", "max_connections", "HikariCP"],
|
||||
"notes": "Covers precise database troubleshooting retrieval."
|
||||
},
|
||||
{
|
||||
"caseId": "chat-diagnosis-flow",
|
||||
"scenario": "chat",
|
||||
"query": "What is the standard troubleshooting flow for an application incident?",
|
||||
"expectedDocIds": ["incident-diagnosis-flow"],
|
||||
"expectedBreadcrumbs": ["AIOps > Diagnosis Flow"],
|
||||
"expectedKeywords": ["collect evidence", "verify", "remediation"],
|
||||
"notes": "Covers process-style knowledge where breadcrumb matters."
|
||||
},
|
||||
{
|
||||
"caseId": "aiops-payment-latency-alert",
|
||||
"scenario": "aiops",
|
||||
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"expectedDocIds": ["payment-service-latency"],
|
||||
"expectedBreadcrumbs": ["AIOps > Service Alerts > Payment Latency"],
|
||||
"expectedKeywords": ["p95 latency", "payment-service", "downstream dependency"],
|
||||
"notes": "Covers alert payload terms that should become retrieval hints."
|
||||
},
|
||||
{
|
||||
"caseId": "aiops-prometheus-alert-scope",
|
||||
"scenario": "aiops",
|
||||
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"expectedDocIds": ["aiops-alert-scope-control"],
|
||||
"expectedBreadcrumbs": ["AIOps > Alert Scope Control"],
|
||||
"expectedKeywords": ["payload", "unrelated active alerts", "scope"],
|
||||
"notes": "Covers scoped alert diagnosis behavior."
|
||||
},
|
||||
{
|
||||
"caseId": "chat-rag-chunk-context",
|
||||
"scenario": "chat",
|
||||
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"expectedDocIds": ["rag-chunk-context-reconstruction"],
|
||||
"expectedBreadcrumbs": ["RAG > Chunking > Context Reconstruction"],
|
||||
"expectedKeywords": ["neighbor chunk", "same section", "breadcrumb"],
|
||||
"notes": "Covers the known RAG refactor issue around context reconstruction."
|
||||
},
|
||||
{
|
||||
"caseId": "chat-l0-domain-hint",
|
||||
"scenario": "chat",
|
||||
"query": "Should L0 keyword matching decide the final retrieval result?",
|
||||
"expectedDocIds": ["rag-l0-domain-entity-hint"],
|
||||
"expectedBreadcrumbs": ["RAG > L0 > Domain Entity Hint"],
|
||||
"expectedKeywords": ["domain detector", "entity extractor", "metadata filter"],
|
||||
"notes": "Covers the target L0 role after refactor."
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,25 @@
|
||||
{
|
||||
"caseId": "aiops-payment-latency-alert",
|
||||
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"retrievedAt": "2026-07-05T00:00:00Z",
|
||||
"candidates": [
|
||||
{
|
||||
"rank": 1,
|
||||
"docId": "payment-service-latency",
|
||||
"title": "Payment Service Latency Alert Playbook",
|
||||
"breadcrumb": "AIOps > Service Alerts > Payment Latency",
|
||||
"content": "For payment-service p95 latency alerts, check downstream dependency latency, thread pool saturation, gateway retries, and recent deployment changes.",
|
||||
"score": 0.84,
|
||||
"retrievalLayer": "L1"
|
||||
},
|
||||
{
|
||||
"rank": 2,
|
||||
"docId": "mysql-connection-pool",
|
||||
"title": "MySQL Connection Pool Troubleshooting",
|
||||
"breadcrumb": "Database > MySQL > Connection Pool",
|
||||
"content": "Database connection pool saturation can increase payment latency when checkout paths wait for connections.",
|
||||
"score": 0.68,
|
||||
"retrievalLayer": "L1"
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,16 @@
|
||||
{
|
||||
"caseId": "aiops-prometheus-alert-scope",
|
||||
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"retrievedAt": "2026-07-05T00:00:00Z",
|
||||
"candidates": [
|
||||
{
|
||||
"rank": 1,
|
||||
"docId": "aiops-alert-scope-control",
|
||||
"title": "AIOps Alert Scope Control",
|
||||
"breadcrumb": "AIOps > Alert Scope Control",
|
||||
"content": "When payload mode is active, queryPrometheusAlerts can verify the supplied alert, but unrelated active alerts must remain scoped context and should not become full diagnoses.",
|
||||
"score": 0.9,
|
||||
"retrievalLayer": "L0+L1"
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,25 @@
|
||||
{
|
||||
"caseId": "chat-diagnosis-flow",
|
||||
"query": "What is the standard troubleshooting flow for an application incident?",
|
||||
"retrievedAt": "2026-07-05T00:00:00Z",
|
||||
"candidates": [
|
||||
{
|
||||
"rank": 1,
|
||||
"docId": "incident-diagnosis-flow",
|
||||
"title": "Incident Diagnosis Flow",
|
||||
"breadcrumb": "AIOps > Diagnosis Flow",
|
||||
"content": "The standard flow is to collect evidence, identify the suspected fault domain, verify the hypothesis, apply remediation, and confirm recovery.",
|
||||
"score": 0.82,
|
||||
"retrievalLayer": "L1"
|
||||
},
|
||||
{
|
||||
"rank": 2,
|
||||
"docId": "rag-chunk-context-reconstruction",
|
||||
"title": "RAG Chunk Context Reconstruction",
|
||||
"breadcrumb": "RAG > Chunking > Context Reconstruction",
|
||||
"content": "Long sections may require neighbor chunk expansion and breadcrumb-aware packing.",
|
||||
"score": 0.55,
|
||||
"retrievalLayer": "L1"
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,25 @@
|
||||
{
|
||||
"caseId": "chat-l0-domain-hint",
|
||||
"query": "Should L0 keyword matching decide the final retrieval result?",
|
||||
"retrievedAt": "2026-07-05T00:00:00Z",
|
||||
"candidates": [
|
||||
{
|
||||
"rank": 1,
|
||||
"docId": "rag-l0-domain-entity-hint",
|
||||
"title": "RAG L0 Domain Entity Hint",
|
||||
"breadcrumb": "RAG > L0 > Domain Entity Hint",
|
||||
"content": "L0 should be retained as a domain detector, entity extractor, metadata filter generator, and explainability signal, not as the final retrieval decision.",
|
||||
"score": 0.88,
|
||||
"retrievalLayer": "L0"
|
||||
},
|
||||
{
|
||||
"rank": 2,
|
||||
"docId": "rag-l0-l1-fusion-ranking",
|
||||
"title": "RAG L0 L1 Fusion Ranking",
|
||||
"breadcrumb": "RAG > Ranking > Fusion",
|
||||
"content": "L0 and L1 candidates should eventually be fused rather than handled as an early-return branch.",
|
||||
"score": 0.75,
|
||||
"retrievalLayer": "L1"
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,25 @@
|
||||
{
|
||||
"caseId": "chat-mysql-connection-pool",
|
||||
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"retrievedAt": "2026-07-05T00:00:00Z",
|
||||
"candidates": [
|
||||
{
|
||||
"rank": 1,
|
||||
"docId": "mysql-connection-pool",
|
||||
"title": "MySQL Connection Pool Troubleshooting",
|
||||
"breadcrumb": "Database > MySQL > Connection Pool",
|
||||
"content": "When the connection pool is exhausted, inspect HikariCP active connections, max_connections, slow SQL, leak detection, and database wait events.",
|
||||
"score": 0.86,
|
||||
"retrievalLayer": "L0+L1"
|
||||
},
|
||||
{
|
||||
"rank": 2,
|
||||
"docId": "incident-diagnosis-flow",
|
||||
"title": "Incident Diagnosis Flow",
|
||||
"breadcrumb": "AIOps > Diagnosis Flow",
|
||||
"content": "Collect evidence, compare metrics and logs, then verify remediation before closing the incident.",
|
||||
"score": 0.61,
|
||||
"retrievalLayer": "L1"
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,25 @@
|
||||
{
|
||||
"caseId": "chat-rag-chunk-context",
|
||||
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"retrievedAt": "2026-07-05T00:00:00Z",
|
||||
"candidates": [
|
||||
{
|
||||
"rank": 1,
|
||||
"docId": "rag-chunk-context-reconstruction",
|
||||
"title": "RAG Chunk Context Reconstruction",
|
||||
"breadcrumb": "RAG > Chunking > Context Reconstruction",
|
||||
"content": "After a chunk hit, expand to neighbor chunk candidates from the same section and preserve breadcrumb metadata in the evidence pack.",
|
||||
"score": 0.79,
|
||||
"retrievalLayer": "L1"
|
||||
},
|
||||
{
|
||||
"rank": 2,
|
||||
"docId": "rag-breadcrumb-embedding-gap",
|
||||
"title": "RAG Breadcrumb Embedding Gap",
|
||||
"breadcrumb": "RAG > Embedding > Breadcrumb",
|
||||
"content": "Embedding title and breadcrumb with content helps recover section semantics.",
|
||||
"score": 0.72,
|
||||
"retrievalLayer": "L1"
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
|
||||
@@ -0,0 +1,131 @@
|
||||
{
|
||||
"generatedAt": "2026-07-04T17:59:52.172759+00:00",
|
||||
"caseFile": "eval/rag-retrieval/cases/golden-cases.json",
|
||||
"fixtureDir": "eval/rag-retrieval/fixtures",
|
||||
"aggregate": {
|
||||
"caseCount": 6,
|
||||
"topK": 5,
|
||||
"strongHitCount": 6,
|
||||
"mediumHitCount": 0,
|
||||
"weakHitCount": 0,
|
||||
"missCount": 0,
|
||||
"recallAtK": 1.0,
|
||||
"strongHitRate": 1.0,
|
||||
"averageFirstHitRank": 1.0
|
||||
},
|
||||
"results": [
|
||||
{
|
||||
"caseId": "chat-mysql-connection-pool",
|
||||
"scenario": "chat",
|
||||
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:mysql-connection-pool",
|
||||
"2:incident-diagnosis-flow"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"connection pool",
|
||||
"max_connections",
|
||||
"hikaricp"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "chat-diagnosis-flow",
|
||||
"scenario": "chat",
|
||||
"query": "What is the standard troubleshooting flow for an application incident?",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:incident-diagnosis-flow",
|
||||
"2:rag-chunk-context-reconstruction"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"collect evidence",
|
||||
"verify",
|
||||
"remediation"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "aiops-payment-latency-alert",
|
||||
"scenario": "aiops",
|
||||
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:payment-service-latency",
|
||||
"2:mysql-connection-pool"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"p95 latency",
|
||||
"payment-service",
|
||||
"downstream dependency"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "aiops-prometheus-alert-scope",
|
||||
"scenario": "aiops",
|
||||
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:aiops-alert-scope-control"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"payload",
|
||||
"unrelated active alerts",
|
||||
"scope"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "chat-rag-chunk-context",
|
||||
"scenario": "chat",
|
||||
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:rag-chunk-context-reconstruction",
|
||||
"2:rag-breadcrumb-embedding-gap"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"neighbor chunk",
|
||||
"same section",
|
||||
"breadcrumb"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"failedChecks": []
|
||||
},
|
||||
{
|
||||
"caseId": "chat-l0-domain-hint",
|
||||
"scenario": "chat",
|
||||
"query": "Should L0 keyword matching decide the final retrieval result?",
|
||||
"hitLevel": "strong",
|
||||
"passed": true,
|
||||
"firstExpectedRank": 1,
|
||||
"topCandidates": [
|
||||
"1:rag-l0-domain-entity-hint",
|
||||
"2:rag-l0-l1-fusion-ranking"
|
||||
],
|
||||
"matchedKeywords": [
|
||||
"domain detector",
|
||||
"entity extractor",
|
||||
"metadata filter"
|
||||
],
|
||||
"breadcrumbMatched": true,
|
||||
"failedChecks": []
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,28 @@
|
||||
# RAG Retrieval Baseline
|
||||
|
||||
Generated at: `2026-07-04T17:59:52.172759+00:00`
|
||||
|
||||
## Aggregate
|
||||
|
||||
| Metric | Value |
|
||||
|---|---:|
|
||||
| Cases | 6 |
|
||||
| Top K | 5 |
|
||||
| Recall@K | 1.0 |
|
||||
| Strong hit rate | 1.0 |
|
||||
| Strong hits | 6 |
|
||||
| Medium hits | 0 |
|
||||
| Weak hits | 0 |
|
||||
| Misses | 0 |
|
||||
| Average first hit rank | 1.0 |
|
||||
|
||||
## Cases
|
||||
|
||||
| Case | Scenario | Hit | First Expected Rank | Top Candidates | Failed Checks |
|
||||
|---|---|---|---:|---|---|
|
||||
| chat-mysql-connection-pool | chat | strong | 1 | 1:mysql-connection-pool<br>2:incident-diagnosis-flow | |
|
||||
| chat-diagnosis-flow | chat | strong | 1 | 1:incident-diagnosis-flow<br>2:rag-chunk-context-reconstruction | |
|
||||
| aiops-payment-latency-alert | aiops | strong | 1 | 1:payment-service-latency<br>2:mysql-connection-pool | |
|
||||
| aiops-prometheus-alert-scope | aiops | strong | 1 | 1:aiops-alert-scope-control | |
|
||||
| chat-rag-chunk-context | chat | strong | 1 | 1:rag-chunk-context-reconstruction<br>2:rag-breadcrumb-embedding-gap | |
|
||||
| chat-l0-domain-hint | chat | strong | 1 | 1:rag-l0-domain-entity-hint<br>2:rag-l0-l1-fusion-ranking | |
|
||||
@@ -0,0 +1,59 @@
|
||||
# SuperBizAgent Interview Guide
|
||||
|
||||
## 一句话定位
|
||||
|
||||
SuperBizAgent 是一个面向企业故障诊断场景的 Agent Engineering 项目:它把用户问题或告警事件转成可追踪的多 Agent 执行链路,并把工具证据、模型步骤、最终答案和反馈统一落到诊断 trace 中。
|
||||
|
||||
## 面试重点
|
||||
|
||||
- **多 Agent 编排**:普通 Chat 的复杂问题走 `Planner -> Executor -> Verifier`;AIOps 告警入口走 Supervisor 调度 Planner/Executor。
|
||||
- **工具证据链**:知识库、日志、指标和 Prometheus 告警都通过工具调用进入链路,并记录到 `tool_invocation`。
|
||||
- **可追踪诊断**:一次会话对应一个 `sessionId`,最终可以通过 `GET /api/diagnosis/{sessionId}/trace` 回放。
|
||||
- **质量门**:Chat 链路包含 Verifier,把 groundedness、facts checked 和 evidence refs 写回 `diagnosis_session.self_evaluation`。
|
||||
- **AIOps 产品边界**:有告警 payload 时聚焦该告警;没有 payload 时先自动发现 active alerts。
|
||||
- **可复现 Demo**:`mvp-demo` profile 使用 mock Prometheus 和 mock CLS,让面试演示不依赖真实线上故障。
|
||||
|
||||
## 推荐阅读顺序
|
||||
|
||||
1. `interview/demo-script.md`:面试现场怎么讲、怎么演示。
|
||||
2. `interview/architecture.md`:系统架构和两条主链路。
|
||||
3. `interview/design-tradeoffs.md`:关键设计取舍和可被追问的问题。
|
||||
4. `interview/acceptance-checklist.md`:面试前验证清单。
|
||||
5. `mvp/demo/README.md`:更细的 MVP 可执行 runbook。
|
||||
|
||||
## 核心 Demo
|
||||
|
||||
### Chat Diagnosis
|
||||
|
||||
```text
|
||||
POST /api/chat
|
||||
-> ChatService.executeChatWithStrategy(...)
|
||||
-> simple ReactAgent or Planner -> Executor -> Verifier
|
||||
-> lookup_knowledge / query_logs / query_metrics
|
||||
-> diagnosis_session + agent_step + tool_invocation
|
||||
-> GET /api/diagnosis/{sessionId}/trace
|
||||
```
|
||||
|
||||
### AIOps Alert Diagnosis
|
||||
|
||||
```text
|
||||
POST /api/ai_ops
|
||||
-> AiOpsService.executeAiOpsAnalysis(...)
|
||||
-> ai_ops_supervisor
|
||||
-> planner_agent / executor_agent
|
||||
-> queryPrometheusAlerts + logs + knowledge
|
||||
-> scoped alert report
|
||||
-> GET /api/diagnosis/{sessionId}/trace
|
||||
```
|
||||
|
||||
## 当前完成度
|
||||
|
||||
- Chat 诊断链路:可运行、可追踪、有 Verifier。
|
||||
- AIOps 告警链路:可运行、可追踪、支持 payload scope control。
|
||||
- Trace API:统一返回 session、agent steps、tool invocations 和 summary。
|
||||
- Demo 文档:`mvp/demo/README.md` 和 `mvp/demo/aiops-alert-acceptance.md`。
|
||||
- Devflow 沉淀:`devflow/index.md` 记录了 MVP、Verifier、AIOps trace 和 AIOps scope-control 的演进。
|
||||
|
||||
## 面试时的主叙事
|
||||
|
||||
这个项目不是简单调用大模型,而是在做一个可审计的 Agent 诊断系统。核心价值是:模型可以规划和推理,但每一步工具证据、最终结论和质量评估都能被 trace API 回放。面试时重点展示“从问题到证据到答案到验证”的完整闭环。
|
||||
@@ -0,0 +1,170 @@
|
||||
# Acceptance Checklist
|
||||
|
||||
## 面试前环境检查
|
||||
|
||||
- 当前分支包含最新 AIOps trace/scope 变更。
|
||||
- MySQL 可连接。
|
||||
- Redis 可连接。
|
||||
- Milvus/Zilliz 可连接。
|
||||
- 模型 API key 可用。
|
||||
- `mvp-demo` profile 开启 mock Prometheus 和 mock CLS。
|
||||
|
||||
启动:
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||
```
|
||||
|
||||
编译检查:
|
||||
|
||||
```powershell
|
||||
mvn -q -DskipTests compile
|
||||
```
|
||||
|
||||
目标测试:
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=AiOpsServiceTest,ChatServiceSequentialAgentTest,DiagnosisTraceServiceTest" test
|
||||
```
|
||||
|
||||
## Chat Demo 验收
|
||||
|
||||
请求:
|
||||
|
||||
```powershell
|
||||
$sessionId = "interview-chat-payment-timeout-001"
|
||||
$body = @{
|
||||
Id = $sessionId
|
||||
Question = "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/chat" `
|
||||
-ContentType "application/json" `
|
||||
-Body $body
|
||||
```
|
||||
|
||||
验收:
|
||||
|
||||
- 返回 `data.success = true`。
|
||||
- 返回 `data.sessionId = interview-chat-payment-timeout-001`。
|
||||
- `diagnosis_session.agent_flow = CHAT`。
|
||||
- trace API 返回 session、steps、toolInvocations。
|
||||
- 复杂问题下 trace 中能看到 verifier 相关数据。
|
||||
|
||||
SQL:
|
||||
|
||||
```powershell
|
||||
python scripts/query_mysql.py "SELECT session_id, agent_flow, status, step_count, tool_call_count FROM diagnosis_session WHERE session_id='interview-chat-payment-timeout-001'"
|
||||
```
|
||||
|
||||
## AIOps Demo 验收
|
||||
|
||||
请求:
|
||||
|
||||
```powershell
|
||||
$aiopsSessionId = "interview-aiops-payment-cpu-001"
|
||||
$aiopsBody = @{
|
||||
sessionId = $aiopsSessionId
|
||||
alertName = "HighCPUUsage"
|
||||
service = "payment-service"
|
||||
severity = "P1"
|
||||
description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
|
||||
timeRange = "last_15m"
|
||||
userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-WebRequest `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/ai_ops" `
|
||||
-ContentType "application/json" `
|
||||
-Body $aiopsBody
|
||||
```
|
||||
|
||||
验收:
|
||||
|
||||
- SSE 首条包含 `type=session`。
|
||||
- SSE 最后包含 `type=done`。
|
||||
- `diagnosis_session.agent_flow = AI_OPS`。
|
||||
- `diagnosis_session.status = SUCCESS`。
|
||||
- `diagnosis_session.answer` 有最终报告。
|
||||
- trace API 返回 AIOps steps 和 tool invocations。
|
||||
- 报告主章节聚焦 `HighCPUUsage/payment-service`。
|
||||
- 无 `告警根因分析 - HighMemoryUsage` 独立章节。
|
||||
- 无 `告警根因分析 - SlowResponse` 独立章节。
|
||||
- 有“相关风险告警”或类似上下文说明。
|
||||
|
||||
SQL:
|
||||
|
||||
```powershell
|
||||
python scripts/query_mysql.py "SELECT session_id, agent_flow, status, total_duration_ms, step_count, tool_call_count FROM diagnosis_session WHERE session_id='interview-aiops-payment-cpu-001'"
|
||||
```
|
||||
|
||||
```powershell
|
||||
python scripts/query_mysql.py "SELECT tool_name, COUNT(*) AS cnt FROM tool_invocation WHERE session_id='interview-aiops-payment-cpu-001' GROUP BY tool_name ORDER BY tool_name"
|
||||
```
|
||||
|
||||
Scope 检查:
|
||||
|
||||
```powershell
|
||||
python scripts/query_mysql.py "SELECT (answer LIKE '%告警根因分析 - HighCPUUsage%') AS has_main_root_cause, (answer LIKE '%告警根因分析 - HighMemoryUsage%') AS has_memory_root_cause, (answer LIKE '%告警根因分析 - SlowResponse%') AS has_slow_root_cause, (answer LIKE '%相关风险告警%') AS has_related_risk FROM diagnosis_session WHERE session_id='interview-aiops-payment-cpu-001'"
|
||||
```
|
||||
|
||||
期望:
|
||||
|
||||
```text
|
||||
has_main_root_cause = 1
|
||||
has_memory_root_cause = 0
|
||||
has_slow_root_cause = 0
|
||||
has_related_risk = 1
|
||||
```
|
||||
|
||||
## Trace API 验收
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/interview-aiops-payment-cpu-001/trace"
|
||||
```
|
||||
|
||||
若 PowerShell 对长 JSON 或特殊字符不稳定,可以用:
|
||||
|
||||
```powershell
|
||||
curl.exe --silent --show-error --max-time 60 "http://localhost:9900/api/diagnosis/interview-aiops-payment-cpu-001/trace"
|
||||
```
|
||||
|
||||
## 常见问题
|
||||
|
||||
### MySQL stale connection
|
||||
|
||||
现象:
|
||||
|
||||
```text
|
||||
HikariPool - Connection is not available
|
||||
No operations allowed after connection closed
|
||||
```
|
||||
|
||||
当前已在 `application.yml` 配置:
|
||||
|
||||
- `maximum-pool-size: 5`
|
||||
- `minimum-idle: 1`
|
||||
- `connection-timeout: 10000`
|
||||
- `validation-timeout: 5000`
|
||||
- `idle-timeout: 60000`
|
||||
- `max-lifetime: 120000`
|
||||
- `keepalive-time: 30000`
|
||||
|
||||
处理:
|
||||
|
||||
- 重新编译或重启服务。
|
||||
- 确认日志中新的 HikariPool 启动成功。
|
||||
- 再跑 trace 或 AIOps 请求。
|
||||
|
||||
### SSE 客户端显示异常
|
||||
|
||||
PowerShell `Invoke-WebRequest` 有时对 SSE 或长 JSON 处理不稳定。可以改用 `curl.exe` 或直接查询 MySQL 和 trace API 验证结果。
|
||||
|
||||
### OpenSpec 全量校验失败
|
||||
|
||||
`openspec validate --all --strict` 可能因为历史未完成 change 失败。面试材料主要依赖已归档的 AIOps spec 和 MVP trace spec,可以单独验证相关 spec。
|
||||
@@ -0,0 +1,53 @@
|
||||
# AIOps Lightweight Verifier
|
||||
|
||||
## What Changed
|
||||
|
||||
AIOps now has a deterministic post-run quality gate.
|
||||
|
||||
After the final AIOps report is persisted, the service evaluates:
|
||||
|
||||
- whether the final report exists and is not trivially short
|
||||
- whether a payload-targeted report mentions the supplied alert and service
|
||||
- whether evidence tools such as `lookup_knowledge`, `query_metrics`, or `query_logs` were persisted
|
||||
|
||||
The result is stored under:
|
||||
|
||||
```text
|
||||
diagnosis_session.self_evaluation.aiops_rule_evaluation
|
||||
```
|
||||
|
||||
The trace API returns this payload through the existing session self-evaluation field.
|
||||
|
||||
## Why Rule-Based First
|
||||
|
||||
This is not a full LLM verifier yet.
|
||||
|
||||
The first AIOps quality risks are concrete and easy to check with rules:
|
||||
|
||||
- Did the report stay focused on the payload?
|
||||
- Did the run use evidence tools?
|
||||
- Did the system produce a usable final report?
|
||||
|
||||
Rule evaluation is stable, cheap, and easy to explain. It also avoids adding another hidden model call to the AIOps flow before the current trace contract is mature.
|
||||
|
||||
## Verdicts
|
||||
|
||||
The evaluator emits:
|
||||
|
||||
```text
|
||||
PASS
|
||||
WARN
|
||||
FAIL
|
||||
```
|
||||
|
||||
`FAIL` is reserved for critical issues such as a missing or too-short report. Missing payload focus terms or missing evidence tools currently produce `WARN`, because valid reports may use slightly different wording or evidence may be unavailable in a mock/demo environment.
|
||||
|
||||
## Interview Answer
|
||||
|
||||
If asked why AIOps has a verifier now:
|
||||
|
||||
> Chat already has an LLM verifier because the user questions are open-ended. For AIOps, I started with a lighter rule-based verifier because the first quality checks are very concrete: payload focus, evidence coverage, and report completeness. The evaluation is persisted into `self_evaluation`, so the trace can show not only what the Agent did, but also whether the output passed basic quality gates.
|
||||
|
||||
If asked why not use the Chat verifier directly:
|
||||
|
||||
> AIOps verification is different from Chat verification. It needs to check alert scope, evidence tool coverage, and whether unrelated active alerts were over-expanded. Reusing the Chat verifier directly would blur those semantics. The rule-based evaluator gives us a stable first quality gate; a later AIOps LLM verifier can build on the same trace contract.
|
||||
@@ -0,0 +1,62 @@
|
||||
# AIOps Query Augmentation
|
||||
|
||||
## What Changed
|
||||
|
||||
Payload-targeted AIOps prompts now include a deterministic recommended knowledge query.
|
||||
|
||||
The query is built from the non-blank payload fields:
|
||||
|
||||
```text
|
||||
alertName service severity description timeRange userRequest
|
||||
```
|
||||
|
||||
Example:
|
||||
|
||||
```text
|
||||
HighCPUUsage payment-service P1 CPU usage is above 80% last_15m
|
||||
```
|
||||
|
||||
## Why This Matters
|
||||
|
||||
AIOps payload fields contain high-value retrieval terms:
|
||||
|
||||
- alert name
|
||||
- service name
|
||||
- severity
|
||||
- symptom description
|
||||
- time range
|
||||
- operator request
|
||||
|
||||
Before this change, the Agent still had to invent its own `lookup_knowledge` query from the full prompt. That can work, but it may omit important terms such as the service name or alert name.
|
||||
|
||||
The new prompt makes the retrieval seed explicit:
|
||||
|
||||
```text
|
||||
Recommended lookup_knowledge query: ...
|
||||
```
|
||||
|
||||
## Design Choice
|
||||
|
||||
This is prompt-level query augmentation, not hidden retrieval.
|
||||
|
||||
I intentionally did not call `lookup_knowledge` automatically before the Agent runs. The project values traceability: tool calls should appear as Agent actions, with their inputs and outputs recorded in `tool_invocation`.
|
||||
|
||||
So the design is:
|
||||
|
||||
```text
|
||||
AIOps payload
|
||||
-> deterministic recommended retrieval query
|
||||
-> Agent prompt
|
||||
-> Agent may call lookup_knowledge explicitly
|
||||
-> tool_invocation records the real retrieval action
|
||||
```
|
||||
|
||||
## Interview Answer
|
||||
|
||||
If asked how AIOps payload improves RAG retrieval:
|
||||
|
||||
> I do not replace the user query with a broad domain. I extract the high-signal alert terms from the payload, such as alertName, service, severity, symptom, and time range, and put them into a compact recommended lookup query. The Agent still calls `lookup_knowledge` explicitly, so the trace remains auditable, but the retrieval query is less dependent on model improvisation.
|
||||
|
||||
If asked why not auto-call retrieval:
|
||||
|
||||
> Auto-calling retrieval would create hidden evidence before the Agent actually decides to use a tool. For this project, explicit tool invocation is more important because the interview story is about observable Agent execution. Prompt-level augmentation gives the Agent a better query seed without changing the trace contract.
|
||||
@@ -0,0 +1,147 @@
|
||||
# Architecture
|
||||
|
||||
## 系统分层
|
||||
|
||||
```text
|
||||
API Layer
|
||||
-> ChatController / DiagnosisTraceController
|
||||
|
||||
Agent Orchestration
|
||||
-> ChatService / AiOpsService
|
||||
|
||||
Tools
|
||||
-> lookupKnowledgeTool / queryLogs / queryMetrics / queryPrometheusAlerts
|
||||
|
||||
Persistence
|
||||
-> diagnosis_session / agent_step / tool_invocation
|
||||
|
||||
Trace
|
||||
-> GET /api/diagnosis/{sessionId}/trace
|
||||
```
|
||||
|
||||
## Chat 链路
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
User[User Question] --> ChatAPI[POST /api/chat]
|
||||
ChatAPI --> Strategy[ChatService.executeChatWithStrategy]
|
||||
Strategy --> Complexity{QuestionComplexity}
|
||||
Complexity -->|simple| Single[ReactAgent]
|
||||
Complexity -->|complex| Planner[Planner Agent]
|
||||
Planner --> Executor[Executor Agent]
|
||||
Executor --> Tools[Evidence Tools]
|
||||
Tools --> Executor
|
||||
Executor --> Verifier[Verifier Agent]
|
||||
Verifier --> Answer[Final Answer]
|
||||
Answer --> Session[diagnosis_session]
|
||||
Planner --> Steps[agent_step]
|
||||
Executor --> Steps
|
||||
Verifier --> Steps
|
||||
Tools --> Invocations[tool_invocation]
|
||||
Session --> Trace[GET /api/diagnosis/{sessionId}/trace]
|
||||
Steps --> Trace
|
||||
Invocations --> Trace
|
||||
```
|
||||
|
||||
关键代码:
|
||||
|
||||
- `ChatController.chat(...)`
|
||||
- `ChatService.executeChatWithStrategy(...)`
|
||||
- `ChatService.executeChatComplex(...)`
|
||||
- `AgentLoggingHook`
|
||||
- `ToolInvocationRecorder`
|
||||
- `DiagnosisTraceService.getTrace(...)`
|
||||
|
||||
## AIOps 链路
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Alert[Alert Payload or Empty Request] --> AiOpsAPI[POST /api/ai_ops]
|
||||
AiOpsAPI --> SessionEvent[SSE session event]
|
||||
AiOpsAPI --> AiOpsService[AiOpsService.executeAiOpsAnalysis]
|
||||
AiOpsService --> PromptMode{Payload?}
|
||||
PromptMode -->|yes| Targeted[PAYLOAD_TARGETED]
|
||||
PromptMode -->|no| Discovery[AUTO_DISCOVERY]
|
||||
Targeted --> Supervisor[ai_ops_supervisor]
|
||||
Discovery --> Supervisor
|
||||
Supervisor --> Planner[planner_agent]
|
||||
Supervisor --> Executor[executor_agent]
|
||||
Planner --> Tools[Prometheus / Logs / Knowledge]
|
||||
Executor --> Tools
|
||||
Tools --> Report[Alert Report]
|
||||
Report --> Persist[diagnosis_session.answer]
|
||||
Planner --> Steps[agent_step]
|
||||
Executor --> Steps
|
||||
Tools --> Invocations[tool_invocation]
|
||||
Persist --> Trace[GET /api/diagnosis/{sessionId}/trace]
|
||||
Steps --> Trace
|
||||
Invocations --> Trace
|
||||
```
|
||||
|
||||
关键代码:
|
||||
|
||||
- `ChatController.aiOps(...)`
|
||||
- `AIOpsRequest`
|
||||
- `AiOpsService.resolveSessionId(...)`
|
||||
- `AiOpsService.buildTaskPrompt(...)`
|
||||
- `AiOpsService.hasAlertPayload(...)`
|
||||
- `AiOpsService.persistFinalReport(...)`
|
||||
|
||||
## Trace 数据模型
|
||||
|
||||
### `diagnosis_session`
|
||||
|
||||
记录一次诊断会话的主信息:
|
||||
|
||||
- `session_id`
|
||||
- `query`
|
||||
- `status`
|
||||
- `agent_flow`
|
||||
- `total_duration_ms`
|
||||
- `total_token_count`
|
||||
- `step_count`
|
||||
- `tool_call_count`
|
||||
- `answer`
|
||||
- `self_evaluation`
|
||||
- `feedback`
|
||||
|
||||
### `agent_step`
|
||||
|
||||
记录 Agent 模型调用过程:
|
||||
|
||||
- `session_id`
|
||||
- `step_index`
|
||||
- `agent_name`
|
||||
- `model_input`
|
||||
- `model_output`
|
||||
- `thought`
|
||||
- `has_tool_call`
|
||||
- `duration_ms`
|
||||
- `token_count`
|
||||
|
||||
### `tool_invocation`
|
||||
|
||||
记录真实工具调用:
|
||||
|
||||
- `session_id`
|
||||
- `tool_name`
|
||||
- `input_params`
|
||||
- `output_preview`
|
||||
- `output_length`
|
||||
- `retrieval_layer`
|
||||
- `relevance_level`
|
||||
- `duration_ms`
|
||||
- `success`
|
||||
- `error_message`
|
||||
|
||||
## 为什么 trace 是核心
|
||||
|
||||
Agent 系统的风险不只是“答案错”,还包括“答案看起来对但无法解释”。这个项目把执行链路拆成 session、step、tool 三层,让面试官可以看到:
|
||||
|
||||
- 模型为什么这么答
|
||||
- 调了哪些工具
|
||||
- 工具返回了什么证据
|
||||
- Verifier 如何判断答案可信度
|
||||
- 用户反馈如何回写到同一个 session
|
||||
|
||||
这就是项目区别于普通 Chatbot 的地方。
|
||||
@@ -0,0 +1,132 @@
|
||||
# Interview Demo Script
|
||||
|
||||
## 30 秒开场
|
||||
|
||||
这是一个 Agent Engineering 项目,场景是企业故障诊断。它支持两类入口:用户主动提问的 Chat 诊断,以及告警事件驱动的 AIOps 诊断。项目重点不是单次回答,而是把多 Agent 执行、工具证据、Verifier 评估、最终报告和反馈都沉淀成可回放的 trace。
|
||||
|
||||
## Demo 准备
|
||||
|
||||
启动服务:
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||
```
|
||||
|
||||
确认服务地址:
|
||||
|
||||
```text
|
||||
http://localhost:9900
|
||||
```
|
||||
|
||||
`mvp-demo` profile 下:
|
||||
|
||||
- Prometheus 告警使用 mock 数据。
|
||||
- CLS 日志使用 mock 数据。
|
||||
- MySQL、Redis、Milvus/Zilliz 和模型配置仍使用当前项目配置。
|
||||
|
||||
## Demo 1: Chat 诊断
|
||||
|
||||
目标:展示普通用户问题如何进入多 Agent 诊断、调用工具、经过 Verifier,并生成 trace。
|
||||
|
||||
请求:
|
||||
|
||||
```powershell
|
||||
$sessionId = "interview-chat-payment-timeout-001"
|
||||
$body = @{
|
||||
Id = $sessionId
|
||||
Question = "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/chat" `
|
||||
-ContentType "application/json" `
|
||||
-Body $body
|
||||
```
|
||||
|
||||
讲解点:
|
||||
|
||||
- `ChatController` 把请求交给 `ChatService.executeChatWithStrategy(...)`。
|
||||
- 简单问题走单 ReactAgent,复杂问题走 `Planner -> Executor -> Verifier`。
|
||||
- Executor 可以调用知识库、日志、指标等工具。
|
||||
- Verifier 会基于工具证据生成 groundedness 评估。
|
||||
- 最终会写入 `diagnosis_session`、`agent_step`、`tool_invocation`。
|
||||
|
||||
查询 trace:
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace"
|
||||
```
|
||||
|
||||
展示点:
|
||||
|
||||
- `data.session.agentFlow = CHAT`
|
||||
- `data.steps` 中能看到 planner/executor/verifier
|
||||
- `data.toolInvocations` 中能看到证据工具
|
||||
- `data.session.selfEvaluation` 中有 verifier 结果
|
||||
|
||||
## Demo 2: AIOps 告警诊断
|
||||
|
||||
目标:展示告警 payload 如何触发 AIOps 入口,并且报告只聚焦目标告警。
|
||||
|
||||
请求:
|
||||
|
||||
```powershell
|
||||
$aiopsSessionId = "interview-aiops-payment-cpu-001"
|
||||
$aiopsBody = @{
|
||||
sessionId = $aiopsSessionId
|
||||
alertName = "HighCPUUsage"
|
||||
service = "payment-service"
|
||||
severity = "P1"
|
||||
description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
|
||||
timeRange = "last_15m"
|
||||
userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-WebRequest `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/ai_ops" `
|
||||
-ContentType "application/json" `
|
||||
-Body $aiopsBody
|
||||
```
|
||||
|
||||
讲解点:
|
||||
|
||||
- `/api/ai_ops` 接受可选 `AIOpsRequest`。
|
||||
- 首条 SSE 消息会返回 `type=session`。
|
||||
- `AiOpsService` 根据 payload 判断模式:
|
||||
- `PAYLOAD_TARGETED`:聚焦传入告警。
|
||||
- `AUTO_DISCOVERY`:没有 payload 时先查 active alerts。
|
||||
- AIOps 暂时不加 Verifier,先保证告警入口、证据工具和 trace 可用。
|
||||
|
||||
查询 trace:
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace"
|
||||
```
|
||||
|
||||
展示点:
|
||||
|
||||
- `data.session.agentFlow = AI_OPS`
|
||||
- `data.session.answer` 有最终告警报告
|
||||
- `data.toolInvocations` 有 `query_metrics`、`query_logs`、`lookup_knowledge`
|
||||
- 报告有 `HighCPUUsage/payment-service` 的完整根因分析
|
||||
- 其他 active alerts 只作为相关风险出现,不展开成独立根因章节
|
||||
|
||||
## MySQL 验证
|
||||
|
||||
```powershell
|
||||
python scripts/query_mysql.py "SELECT session_id, agent_flow, status, step_count, tool_call_count FROM diagnosis_session ORDER BY id DESC LIMIT 5"
|
||||
```
|
||||
|
||||
```powershell
|
||||
python scripts/query_mysql.py "SELECT tool_name, COUNT(*) AS cnt FROM tool_invocation WHERE session_id='interview-aiops-payment-cpu-001' GROUP BY tool_name"
|
||||
```
|
||||
|
||||
## 收尾总结
|
||||
|
||||
这套 Demo 展示的是一个完整 Agent 系统,而不是一次模型问答:入口有明确场景边界,Agent 负责规划和执行,工具提供证据,Verifier 提供质量门,trace API 提供审计和复盘能力。AIOps 入口进一步证明它可以从用户问答扩展到事件驱动诊断。
|
||||
@@ -0,0 +1,99 @@
|
||||
# Design Tradeoffs
|
||||
|
||||
## 1. 为什么要做 trace,而不是只返回答案
|
||||
|
||||
普通 Chatbot 只关注最终回答,但故障诊断更需要可审计性。一次诊断至少要回答三件事:
|
||||
|
||||
- 结论是什么
|
||||
- 证据来自哪里
|
||||
- 哪些步骤由哪个 Agent 完成
|
||||
|
||||
因此项目把一次会话拆成:
|
||||
|
||||
- `diagnosis_session`:会话级摘要、最终答案、质量评估、反馈。
|
||||
- `agent_step`:Agent 模型输入输出、耗时、token 和工具调用标记。
|
||||
- `tool_invocation`:真实工具调用参数、输出预览、成功状态和检索元数据。
|
||||
|
||||
这个设计牺牲了一些实现复杂度,但换来了可回放、可调试、可演示。
|
||||
|
||||
## 2. 为什么 Chat 有 Verifier,AIOps 暂时没有
|
||||
|
||||
Chat 入口的问题更开放,用户可能要求复杂推理或跨领域结论,所以 Verifier 是必要的质量门。当前 Chat 链路通过 `Planner -> Executor -> Verifier` 固定流程,把 groundedness 和 facts checked 写入 `self_evaluation`。
|
||||
|
||||
AIOps 当前阶段先不加 Verifier,原因是:
|
||||
|
||||
- AIOps 刚完成从“自动跑告警”到“可追踪告警入口”的改造。
|
||||
- 先要确认告警 payload、工具证据、最终报告和 trace 能闭环。
|
||||
- AIOps Verifier 的规则不同于 Chat Verifier,需要检查告警 scope、证据覆盖和处置建议,不宜直接复用。
|
||||
|
||||
后续可以做 lightweight AIOps Verifier,检查报告是否聚焦 payload、是否引用工具证据、是否误展开无关告警。
|
||||
|
||||
## 3. 为什么 AIOps payload scope 先用 prompt 控制
|
||||
|
||||
运行验证发现:传入 `HighCPUUsage/payment-service` 后,Agent 仍可能把 mock Prometheus 返回的所有 active alerts 都展开分析。这个问题的本质是任务边界不清晰。
|
||||
|
||||
当前选择 prompt-level scope control:
|
||||
|
||||
- 有 payload:`PAYLOAD_TARGETED`,最终报告围绕传入告警。
|
||||
- 无 payload:`AUTO_DISCOVERY`,先调用 `queryPrometheusAlerts` 自动发现告警。
|
||||
|
||||
没有先做 Java 侧过滤,是因为:
|
||||
|
||||
- 过滤工具结果会降低 Agent 发现关联风险的能力。
|
||||
- 目前需要的是报告主线聚焦,而不是完全屏蔽上下文。
|
||||
- Prompt 改动小,风险低,能保留 Agent 灵活性。
|
||||
|
||||
已验证结果:主报告有 `HighCPUUsage/payment-service` 的完整根因分析,`HighMemoryUsage` 和 `SlowResponse` 只作为相关风险出现。
|
||||
|
||||
## 4. 为什么用 `tool_invocation` 统计真实工具调用次数
|
||||
|
||||
早期可以通过 `agent_step.hasToolCall` 粗略判断是否调用工具,但它统计的是“哪些模型步骤包含工具调用”,不是“真实调用了几次工具”。
|
||||
|
||||
现在 `tool_call_count` 来自:
|
||||
|
||||
```text
|
||||
ToolInvocationRepository.countBySessionId(sessionId)
|
||||
```
|
||||
|
||||
这样更符合 trace 语义:
|
||||
|
||||
- 一个 step 可能调用多个工具。
|
||||
- 工具可能来自不同来源:知识库、日志、指标、Prometheus。
|
||||
- 面试时可以把 `tool_call_count` 和 trace 中返回的工具明细对上。
|
||||
|
||||
## 5. 为什么保留 mock Prometheus 和 mock CLS
|
||||
|
||||
面试 Demo 最怕不稳定。真实 Prometheus、日志平台和线上故障都有不可控因素,所以 MVP profile 保留 mock 工具:
|
||||
|
||||
- `prometheus.mock-enabled=true`
|
||||
- `cls.mock-enabled=true`
|
||||
|
||||
这样可以稳定复现:
|
||||
|
||||
- `HighCPUUsage/payment-service`
|
||||
- `HighMemoryUsage/order-service`
|
||||
- `SlowResponse/user-service`
|
||||
- system-metrics、application-logs、database-slow-query 等日志证据
|
||||
|
||||
这不是逃避真实集成,而是把“Agent 编排和证据追踪”作为面试演示的主目标。
|
||||
|
||||
## 6. 为什么把面试材料单独放 `interview/`
|
||||
|
||||
`mvp/` 是持续迭代现场,包含过程文档、验收记录和 runbook。面试材料的目标不同,它应该是可讲、可演示、可评估的展示层。
|
||||
|
||||
因此:
|
||||
|
||||
- `mvp/` 保留真实演进材料。
|
||||
- `devflow/` 保留决策沉淀。
|
||||
- `interview/` 只组织面试叙事和演示脚本。
|
||||
|
||||
这样后续继续做 AIOps Verifier、UI、更多工具集成时,不会污染面试讲稿。
|
||||
|
||||
## 7. 可以主动承认的限制
|
||||
|
||||
- AIOps 还没有 Verifier。
|
||||
- Prompt-level scope control 不能做到强约束,只能通过 trace 和测试观察遵循情况。
|
||||
- 当前 mock 数据适合 demo,不代表生产接入已经完成。
|
||||
- Hikari 连接池已经加了短生命周期和 keepalive,但真实生产还需要按数据库 wait_timeout 和连接数预算调优。
|
||||
|
||||
主动讲清这些限制,反而能体现工程判断:先把可追踪闭环打通,再逐步增强质量门和生产可靠性。
|
||||
@@ -0,0 +1,66 @@
|
||||
# RAG Breadcrumb Embedding Acceptance
|
||||
|
||||
## What Changed
|
||||
|
||||
The indexing path now builds embedding text from chunk structure plus content:
|
||||
|
||||
```text
|
||||
Title: {title}
|
||||
Path: {breadcrumb}
|
||||
Content:
|
||||
{content}
|
||||
```
|
||||
|
||||
The stored Milvus `content` field remains the original chunk content. This keeps display and evidence output clean while allowing the vector to carry section-level semantics.
|
||||
|
||||
## Why Reindex Is Required
|
||||
|
||||
Embeddings are materialized at index time. Existing vectors were generated from the previous content-only text, so they cannot benefit from `title` and `breadcrumb` until the knowledge base is reindexed.
|
||||
|
||||
This is the key acceptance point:
|
||||
|
||||
```text
|
||||
code change alone != live retrieval changed
|
||||
code change + reindex + live query report = accepted behavior
|
||||
```
|
||||
|
||||
## How To Validate
|
||||
|
||||
1. Start the Spring Boot application.
|
||||
2. Reindex the knowledge base through the existing indexing path.
|
||||
3. Run:
|
||||
|
||||
```bash
|
||||
python scripts/eval_rag_live_acceptance.py
|
||||
```
|
||||
|
||||
The script writes:
|
||||
|
||||
```text
|
||||
eval/rag-retrieval/reports/live-post-reindex.json
|
||||
eval/rag-retrieval/reports/live-post-reindex.md
|
||||
```
|
||||
|
||||
The default cases cover:
|
||||
|
||||
- RAG chunk context questions where breadcrumb matters.
|
||||
- Diagnosis flow questions where section path matters.
|
||||
- `ERR_TIMEOUT` exact error-code retrieval.
|
||||
- MySQL connection pool troubleshooting.
|
||||
- AIOps payment-service latency alert retrieval.
|
||||
|
||||
## What To Look For
|
||||
|
||||
For breadcrumb-sensitive cases, inspect whether top candidates expose expected `title` and `breadcrumb` values in the report.
|
||||
|
||||
For core troubleshooting cases, check that result counts and top candidates remain stable. The goal is not to prove a full benchmark; it is to prove that reindexing did not obviously break important demo retrieval paths.
|
||||
|
||||
## Interview Answer
|
||||
|
||||
If asked how I verified the breadcrumb embedding change:
|
||||
|
||||
> I separated deterministic regression from live acceptance. The offline fixture baseline still runs without services. But because embedding changes only affect newly indexed vectors, I added a live post-reindex acceptance script. It calls the real `/api/search/similar` endpoint against representative breadcrumb-sensitive, troubleshooting, and AIOps queries, then writes JSON and Markdown reports. This lets me prove both that the code changed and that the live vector collection was refreshed.
|
||||
|
||||
If asked why the script does not reindex automatically:
|
||||
|
||||
> Reindexing mutates the vector store and depends on environment-specific data. I kept mutation explicit and made the script validation-only. That makes failures easier to diagnose: if retrieval does not improve, I can distinguish code changes, reindex state, and runtime retrieval behavior.
|
||||
@@ -0,0 +1,209 @@
|
||||
# RAG Refactor Story
|
||||
|
||||
## The Starting Point
|
||||
|
||||
The original RAG implementation was already usable for the MVP:
|
||||
|
||||
- Documents could be uploaded, chunked, embedded, and written to Milvus/Zilliz.
|
||||
- The Agent could call `lookup_knowledge` as an explicit tool.
|
||||
- AIOps diagnosis could retrieve troubleshooting knowledge during an alert workflow.
|
||||
- Tool invocations were persisted, so the retrieval step was visible in the execution trace.
|
||||
|
||||
But the design had several engineering problems:
|
||||
|
||||
- Retrieval was too SDK-specific. The business code directly owned many Milvus search details.
|
||||
- L0 and L1 responsibilities were blurry. L0 keyword matching could look like a final retrieval decision instead of a hint.
|
||||
- Chunk-level retrieval could lose section context when one section was split into multiple chunks.
|
||||
- Metadata such as `breadcrumb` existed, but it was not fully used in retrieval, filtering, or context reconstruction.
|
||||
- Retrieval quality was mostly checked by manual API calls and logs, not by repeatable cases.
|
||||
|
||||
So the refactor goal was not "replace everything with a framework." The goal was to move generic RAG infrastructure toward Spring AI while keeping the project-specific Agent evidence chain.
|
||||
|
||||
## How I Broke The Problem Down
|
||||
|
||||
I treated this as a staged migration, because RAG touches the Agent tool layer, AIOps diagnosis, vector retrieval, evidence packing, and database traces.
|
||||
|
||||
The first step was to establish a baseline. I added retrieval evaluation cases under `eval/rag-retrieval/` so future changes could be compared against known queries instead of judged only by intuition.
|
||||
|
||||
Then I clarified the retrieval roles:
|
||||
|
||||
```text
|
||||
L0 = domain/entity hint
|
||||
L1 = semantic retrieval
|
||||
postprocess = evidence shaping and trace-friendly output
|
||||
```
|
||||
|
||||
That means L0 is still valuable, but it should not bypass semantic retrieval as the default path. It is better used to extract service names, alert names, error codes, domains, and metadata hints.
|
||||
|
||||
After that, I added evidence postprocessing. The Agent should not just receive raw chunks; it should receive structured evidence with source, title, breadcrumb, score, hit reason, and content. This makes the result easier to inspect and easier to explain in an interview.
|
||||
|
||||
Finally, I integrated Spring AI `VectorStore` as the main read path while preserving the original Milvus SDK implementation as fallback.
|
||||
|
||||
## Current Architecture
|
||||
|
||||
The current retrieval path is:
|
||||
|
||||
```text
|
||||
Agent / API
|
||||
-> lookup_knowledge or /api/search/similar
|
||||
-> L0 domain/entity hint
|
||||
-> VectorSearchService
|
||||
-> Spring AI VectorStore
|
||||
-> Milvus SDK fallback
|
||||
-> evidence postprocess
|
||||
-> tool_invocation trace
|
||||
```
|
||||
|
||||
`VectorSearchService` is still the public retrieval facade. This is deliberate: the Agent tool layer does not need to know whether the underlying retrieval engine is SDK-based or Spring AI-based.
|
||||
|
||||
The supported retrieval modes are:
|
||||
|
||||
```text
|
||||
auto -> try Spring AI VectorStore, fallback to SDK
|
||||
spring-ai -> force Spring AI VectorStore
|
||||
sdk -> force Milvus SDK
|
||||
```
|
||||
|
||||
This keeps the migration reversible and testable.
|
||||
|
||||
## Key Tradeoffs
|
||||
|
||||
### Keep The Explicit Tool
|
||||
|
||||
I did not hide retrieval inside a Spring AI Advisor.
|
||||
|
||||
For this project, `lookup_knowledge` is part of the Agent execution story. It records what query was used, which evidence was retrieved, how relevant it looked, and how it supported diagnosis. If retrieval is hidden inside an advisor, the answer may still work, but the audit trail becomes harder to show.
|
||||
|
||||
### Keep SDK Fallback
|
||||
|
||||
The SDK path is not dead code. It is a safety net during migration.
|
||||
|
||||
This proved useful during live validation. The first VectorStore run pointed at the wrong collection name, but `auto` mode fell back to SDK and still returned results. After the collection was corrected to `biz`, the Spring AI path worked as the main path.
|
||||
|
||||
### Keep L0, But Reduce Its Authority
|
||||
|
||||
L0 is worth keeping because production incidents often contain exact identifiers:
|
||||
|
||||
- error code
|
||||
- alert name
|
||||
- service name
|
||||
- metric name
|
||||
- domain tag
|
||||
|
||||
But L0 should not be the final judge of retrieval quality. Its role is now closer to domain hint, entity extraction, metadata filtering, and explainability signal.
|
||||
|
||||
### Split Score Semantics
|
||||
|
||||
The old SDK path used L2 distance. Spring AI exposes similarity. Treating those as the same number would quietly break relevance normalization.
|
||||
|
||||
So the result separates:
|
||||
|
||||
```text
|
||||
score -> compatibility score used by existing logic
|
||||
rawScore -> raw score from the retrieval implementation
|
||||
scoreLabel -> semantic label for rawScore
|
||||
```
|
||||
|
||||
For SDK:
|
||||
|
||||
```text
|
||||
score = L2 distance
|
||||
rawScore = L2 distance
|
||||
scoreLabel = l2_distance
|
||||
```
|
||||
|
||||
For VectorStore:
|
||||
|
||||
```text
|
||||
score = Milvus metadata.distance when available
|
||||
rawScore = Spring AI similarity
|
||||
scoreLabel = similarity
|
||||
```
|
||||
|
||||
This makes the migration inspectable instead of hiding score changes behind one overloaded field.
|
||||
|
||||
### Do Not Migrate Writes Yet
|
||||
|
||||
Writes and indexing still use the SDK path.
|
||||
|
||||
That is intentional. Migrating reads and writes at the same time would make debugging harder. The read path can be validated first; write-path migration can happen later if Spring AI `VectorStore.add(...)` fits the existing metadata and chunk model.
|
||||
|
||||
## Validation Story
|
||||
|
||||
I validated the refactor at multiple levels.
|
||||
|
||||
Unit tests cover:
|
||||
|
||||
- SDK mode.
|
||||
- Spring AI mode.
|
||||
- `auto` fallback.
|
||||
- category filter behavior.
|
||||
- distance metadata mapping.
|
||||
|
||||
Live API verification used:
|
||||
|
||||
```text
|
||||
GET /api/search/similar?query=ERR_TIMEOUT&topK=3
|
||||
```
|
||||
|
||||
Logs confirmed when the Spring AI VectorStore path was used and when fallback happened.
|
||||
|
||||
Then I compared SDK and VectorStore retrieval quality on representative queries:
|
||||
|
||||
| Query Type | Result |
|
||||
| --- | --- |
|
||||
| exact error code | same top3 |
|
||||
| payment-service timeout | same top3 |
|
||||
| MySQL connection pool | same top3 |
|
||||
| AIOps alert-style query | same top3 |
|
||||
| abstract RAG design query | same top1, VectorStore returned fewer tail results |
|
||||
| category filter | both returned zero because metadata taxonomy did not match |
|
||||
|
||||
The acceptance decision was that Spring AI VectorStore is good enough for the current MVP read path, with SDK fallback preserved.
|
||||
|
||||
## Known Gaps
|
||||
|
||||
The refactor improved the architecture, but it did not solve every retrieval-quality problem.
|
||||
|
||||
Known gaps:
|
||||
|
||||
- Metadata taxonomy still needs cleanup, for example `database` vs `infrastructure`.
|
||||
- Abstract design questions may need query rewriting or better indexed interview/devflow documents.
|
||||
- Chunk context reconstruction is still limited when one logical section spans multiple chunks.
|
||||
- `breadcrumb` now participates in embedding text, but it can still be used more strongly in context expansion, rerank, and evidence packing.
|
||||
- Rerank, RRF, BM25, and hybrid retrieval are not implemented yet.
|
||||
- Indexing writes still use SDK.
|
||||
|
||||
These are good follow-up issues because they are retrieval-quality improvements, not blockers for the VectorStore migration.
|
||||
|
||||
## How I Present This In An Interview
|
||||
|
||||
My short version would be:
|
||||
|
||||
> This RAG system started as a self-built MVP around Milvus SDK retrieval. It worked, but too much infrastructure logic lived in business code, and L0/L1 responsibilities were unclear. I refactored it in stages: first I added baseline retrieval cases, then made L0 a domain/entity hint instead of a final decision layer, then added evidence postprocessing, and finally moved the main read path to Spring AI VectorStore with SDK fallback. I kept `lookup_knowledge` as an explicit Agent tool because the project values traceability: the interviewer can see when retrieval happened, what evidence was found, and how it supported the diagnosis. The result is closer to standard Spring AI RAG while still preserving business-specific observability.
|
||||
|
||||
If asked why this is not a full framework migration:
|
||||
|
||||
> I intentionally did not migrate everything at once. Reads moved first because they are easier to compare using golden queries. Writes/indexing stayed on SDK to avoid mixing schema and retrieval behavior changes in one step. Advisors were not used as the main interface because hidden retrieval would weaken the Agent trace.
|
||||
|
||||
If asked what I would improve next:
|
||||
|
||||
> I would add query transformation for AIOps payloads, improve metadata taxonomy, use breadcrumb and section metadata for context expansion, and then evaluate whether hybrid retrieval or rerank is necessary based on measured recall and topK overlap.
|
||||
|
||||
## Interview Follow-Up Questions
|
||||
|
||||
### Why introduce Spring AI VectorStore if the SDK path already worked?
|
||||
|
||||
Because SDK-only retrieval made the project own too much low-level RAG infrastructure. `VectorStore` gives a standard abstraction for retrieval and makes future Spring AI features easier to adopt, while the facade keeps the Agent layer stable.
|
||||
|
||||
### Why keep custom code at all?
|
||||
|
||||
The custom code is where the Agent engineering value lives: AIOps payload mapping, L0 hints, evidence packing, score compatibility, and tool invocation tracing. Those are domain-specific and should remain visible.
|
||||
|
||||
### How do you know quality did not regress?
|
||||
|
||||
I compared SDK and VectorStore modes on representative live queries. Core troubleshooting and AIOps cases returned the same top3 documents in the same order. The differences were isolated to abstract design queries and metadata taxonomy, which are documented follow-up work.
|
||||
|
||||
### What is the most important design decision?
|
||||
|
||||
Keeping a stable boundary: `lookup_knowledge` calls `VectorSearchService`, and `VectorSearchService` decides whether to use Spring AI or SDK. That boundary made the migration small enough to validate and explain.
|
||||
@@ -0,0 +1,211 @@
|
||||
# RAG Retrieval Quality Report
|
||||
|
||||
## Purpose
|
||||
|
||||
This report compares the live retrieval behavior of the original Milvus SDK path and the new Spring AI VectorStore path.
|
||||
|
||||
The goal is to answer an interview-critical question:
|
||||
|
||||
> After moving retrieval to Spring AI VectorStore, how do we know retrieval quality did not regress?
|
||||
|
||||
This is not a full benchmark yet. It is a focused live smoke comparison using representative RAG queries against the current Milvus/Zilliz collection.
|
||||
|
||||
## Setup
|
||||
|
||||
Service endpoint:
|
||||
|
||||
```text
|
||||
GET http://127.0.0.1:9900/api/search/similar
|
||||
```
|
||||
|
||||
Collection:
|
||||
|
||||
```text
|
||||
biz
|
||||
```
|
||||
|
||||
Compared modes:
|
||||
|
||||
```text
|
||||
retrieval.vector-store.mode=sdk
|
||||
retrieval.vector-store.mode=spring-ai
|
||||
```
|
||||
|
||||
Each case used:
|
||||
|
||||
```text
|
||||
topK=3
|
||||
```
|
||||
|
||||
The application was restarted once per mode using command-line configuration so no repository config file had to be changed.
|
||||
|
||||
## Cases
|
||||
|
||||
| Case | Query | Purpose |
|
||||
| --- | --- | --- |
|
||||
| `err-timeout` | `ERR_TIMEOUT` | Exact error-code retrieval |
|
||||
| `payment-service-timeout` | `payment-service timeout` | Service timeout troubleshooting |
|
||||
| `mysql-connection-pool` | `MySQL connection pool is exhausted. How should I diagnose it?` | Database troubleshooting |
|
||||
| `high-cpu-payment` | `HighCPUUsage payment-service` | AIOps alert-style retrieval |
|
||||
| `rag-l0-l1` | `Should L0 keyword matching decide the final retrieval result?` | Abstract RAG design query |
|
||||
| `database-filter` | `mysql timeout`, category=`database` | Metadata filter behavior |
|
||||
|
||||
## Summary
|
||||
|
||||
| Case | SDK Count | VectorStore Count | Top1 Same | TopK Overlap | Notes |
|
||||
| --- | ---: | ---: | --- | ---: | --- |
|
||||
| `err-timeout` | 3 | 3 | Yes | 3/3 | Same ordering and same documents |
|
||||
| `payment-service-timeout` | 3 | 3 | Yes | 3/3 | Same ordering and same documents |
|
||||
| `mysql-connection-pool` | 3 | 3 | Yes | 3/3 | Same ordering and same documents |
|
||||
| `high-cpu-payment` | 3 | 3 | Yes | 3/3 | Same ordering and same documents |
|
||||
| `rag-l0-l1` | 3 | 1 | Yes | 1/3 | VectorStore returned only the strongest candidate |
|
||||
| `database-filter` | 0 | 0 | N/A | N/A | Both paths applied the filter consistently; no live docs matched `category=database` |
|
||||
|
||||
## Representative Results
|
||||
|
||||
### `ERR_TIMEOUT`
|
||||
|
||||
SDK:
|
||||
|
||||
```text
|
||||
1. ERR_TIMEOUT score=0.5659486 label=l2_distance
|
||||
2. ERR_GATEWAY_TIMEOUT score=0.6048740 label=l2_distance
|
||||
3. Error handling score=0.7735061 label=l2_distance
|
||||
```
|
||||
|
||||
VectorStore:
|
||||
|
||||
```text
|
||||
1. ERR_TIMEOUT score=0.5659486 rawScore=0.4340513 label=similarity
|
||||
2. ERR_GATEWAY_TIMEOUT score=0.6048740 rawScore=0.3951259 label=similarity
|
||||
3. Error handling score=0.7735061 rawScore=0.2264938 label=similarity
|
||||
```
|
||||
|
||||
Interpretation:
|
||||
|
||||
- Document ordering is identical.
|
||||
- Compatibility `score` is identical to SDK L2 distance.
|
||||
- VectorStore `rawScore` exposes Spring AI similarity separately.
|
||||
|
||||
### `MySQL connection pool`
|
||||
|
||||
Both paths returned:
|
||||
|
||||
```text
|
||||
1. MySQL connection pool config
|
||||
2. wait_timeout timeout
|
||||
3. idle-timeout
|
||||
```
|
||||
|
||||
Interpretation:
|
||||
|
||||
- The migration preserves a precise infrastructure troubleshooting retrieval case.
|
||||
- Metadata fields such as title, category, and source remain available.
|
||||
|
||||
### `HighCPUUsage payment-service`
|
||||
|
||||
Both paths returned:
|
||||
|
||||
```text
|
||||
1. 3. HighCPUUsage / payment-service troubleshooting steps
|
||||
2. evidence mapping table row for HighCPUUsage/payment-service
|
||||
3. 3.1 Symptom confirmation
|
||||
```
|
||||
|
||||
Interpretation:
|
||||
|
||||
- AIOps-style alert terms still retrieve the expected troubleshooting document.
|
||||
- This is important because AIOps diagnosis depends on knowledge retrieval plus metrics/log evidence.
|
||||
|
||||
### `rag-l0-l1`
|
||||
|
||||
SDK returned three results, while VectorStore returned one:
|
||||
|
||||
```text
|
||||
Top1: Return error information
|
||||
```
|
||||
|
||||
Interpretation:
|
||||
|
||||
- Top1 did not regress.
|
||||
- VectorStore appears stricter for low-similarity tail results because the Spring AI path uses `similarityThresholdAll()`.
|
||||
- This is acceptable for current read-path migration, but it is worth tracking because abstract design questions may need query rewriting, better indexed docs, or adjusted threshold behavior.
|
||||
|
||||
### `database-filter`
|
||||
|
||||
Both paths returned zero results for:
|
||||
|
||||
```text
|
||||
query=mysql timeout
|
||||
category=database
|
||||
```
|
||||
|
||||
Interpretation:
|
||||
|
||||
- The filter path is consistent.
|
||||
- The live indexed MySQL docs are categorized as `infrastructure`, not `database`.
|
||||
- This highlights a metadata taxonomy issue rather than a VectorStore migration regression.
|
||||
|
||||
## Score Compatibility
|
||||
|
||||
The comparison validates the score design:
|
||||
|
||||
```text
|
||||
SDK:
|
||||
score = L2 distance
|
||||
rawScore = L2 distance
|
||||
scoreLabel = l2_distance
|
||||
|
||||
VectorStore:
|
||||
score = Milvus metadata.distance
|
||||
rawScore = Spring AI similarity
|
||||
scoreLabel = similarity
|
||||
```
|
||||
|
||||
This keeps `lookup_knowledge` relevance normalization stable while still exposing the VectorStore score semantics for trace/debugging.
|
||||
|
||||
## Findings
|
||||
|
||||
### Finding 1: Main live cases are equivalent
|
||||
|
||||
For exact error code, service timeout, MySQL troubleshooting, and AIOps alert-style retrieval, SDK and VectorStore returned identical top3 documents in identical order.
|
||||
|
||||
This is strong evidence that the read-path migration did not regress the most important demo and troubleshooting cases.
|
||||
|
||||
### Finding 2: Abstract RAG design queries need better retrieval support
|
||||
|
||||
The `rag-l0-l1` query only returned one VectorStore candidate. The top result matched SDK top1, but the tail differed.
|
||||
|
||||
This suggests the next quality work should focus on:
|
||||
|
||||
- Query transformation for abstract design questions.
|
||||
- Better indexing of interview/devflow RAG design docs.
|
||||
- Context expansion around same-section chunks.
|
||||
- Possibly tuning VectorStore threshold behavior.
|
||||
|
||||
### Finding 3: Metadata taxonomy matters
|
||||
|
||||
The category filter case returned zero results in both modes because the relevant MySQL docs are categorized as `infrastructure`, not `database`.
|
||||
|
||||
This supports a previous RAG issue: category/domain metadata should be normalized before it is used as a hard filter.
|
||||
|
||||
## Acceptance Decision
|
||||
|
||||
The Spring AI VectorStore read path is accepted for current MVP/interview use:
|
||||
|
||||
- Core troubleshooting cases match SDK behavior.
|
||||
- Score compatibility is preserved.
|
||||
- The VectorStore path exposes better score semantics without changing the `lookup_knowledge` API.
|
||||
- SDK fallback remains available for runtime safety.
|
||||
|
||||
The next retrieval-quality improvements should not block this migration. They should be handled as separate RAG quality work.
|
||||
|
||||
## Next Work
|
||||
|
||||
Recommended next steps:
|
||||
|
||||
- Add a small automated live comparison script if repeated validation becomes common.
|
||||
- Add topK overlap and top1 hit metrics to the offline evaluator.
|
||||
- Normalize metadata categories such as `database` vs `infrastructure`.
|
||||
- Add query rewriting for abstract RAG questions.
|
||||
- Decide later whether to migrate indexing writes to Spring AI `VectorStore.add(...)`.
|
||||
@@ -0,0 +1,167 @@
|
||||
# RAG VectorStore Interview Notes
|
||||
|
||||
## 60-Second Explanation
|
||||
|
||||
I refactored the RAG retrieval path from a direct Milvus SDK-only implementation to a Spring AI `VectorStore` main path, while keeping the SDK path as a fallback.
|
||||
|
||||
The important part is not just the dependency change. I kept `VectorSearchService` as the boundary, so `lookup_knowledge` and the Agent workflow did not need to change. The system now supports three modes:
|
||||
|
||||
```text
|
||||
auto -> try Spring AI VectorStore, fallback to SDK
|
||||
spring-ai -> force VectorStore
|
||||
sdk -> force SDK
|
||||
```
|
||||
|
||||
During live verification, the first run found a real config mismatch: VectorStore was pointed at `business_knowledge`, but the real Zilliz collection was `biz`. The fallback worked, so the system still returned results through SDK. After aligning the collection name, the same query went through Spring AI VectorStore successfully.
|
||||
|
||||
I also fixed score compatibility. Spring AI Milvus exposes similarity as the document score, but the old `lookup_knowledge` logic expects L2 distance. So I preserve `rawScore` and `scoreLabel`, and use Milvus `metadata.distance` as the compatibility `score` when available.
|
||||
|
||||
## Architecture Answer
|
||||
|
||||
```text
|
||||
Agent / API
|
||||
-> lookup_knowledge or /api/search/similar
|
||||
-> VectorSearchService
|
||||
-> Spring AI VectorStore
|
||||
-> Milvus SDK fallback
|
||||
-> Milvus/Zilliz collection: biz
|
||||
```
|
||||
|
||||
The key design choice is that `VectorSearchService` remains the retrieval facade. This avoids spreading framework-specific code into the Agent tool layer.
|
||||
|
||||
## Why Keep The SDK Path?
|
||||
|
||||
I kept SDK fallback for three reasons:
|
||||
|
||||
- Migration safety: the existing SDK path was already proven against the live collection.
|
||||
- Runtime resilience: if VectorStore schema mapping or filtering fails, retrieval still works.
|
||||
- Interview/demo stability: a retrieval abstraction change should not break the main Agent diagnosis demo.
|
||||
|
||||
This was validated in practice. When VectorStore pointed at the wrong collection, `auto` mode fell back to SDK and still returned results.
|
||||
|
||||
## Why Use Spring AI VectorStore At All?
|
||||
|
||||
Using Spring AI `VectorStore` moves the project closer to a standard RAG abstraction:
|
||||
|
||||
- Retrieval code no longer needs to own all Milvus-specific search details.
|
||||
- Later features such as query transformers, document postprocessors, advisors, or retrievers can be introduced more naturally.
|
||||
- The code becomes easier to compare with common Spring AI RAG patterns in an interview.
|
||||
|
||||
But I did not blindly replace everything. Writes/indexing still use SDK because changing read and write paths at the same time would make failures harder to isolate.
|
||||
|
||||
## Why Keep L0?
|
||||
|
||||
L0 is no longer treated as the final source of truth. It is a deterministic hint layer:
|
||||
|
||||
- It extracts domain/entity hints from indexed metadata.
|
||||
- It helps constrain L1 retrieval by category when possible.
|
||||
- It gives the Agent a stable clue even when semantic retrieval is weak.
|
||||
|
||||
The current design is:
|
||||
|
||||
```text
|
||||
L0 = domain/entity hint
|
||||
L1 = semantic retrieval through VectorStore/SDK
|
||||
postprocess = evidence trace and relevance normalization
|
||||
```
|
||||
|
||||
This is easier to defend than saying "we only use vector search." Real incident diagnosis often has exact identifiers, error codes, service names, and alert names. L0 is useful for those.
|
||||
|
||||
## Why Not Use Hidden Spring AI Advisors Directly?
|
||||
|
||||
For this project, `lookup_knowledge` remains an explicit tool.
|
||||
|
||||
Reason:
|
||||
|
||||
- The Agent trace needs to show when knowledge was retrieved.
|
||||
- `tool_invocation` records input, output preview, relevance level, and evidence metadata.
|
||||
- The interview story is about auditable Agent execution, not only answer quality.
|
||||
|
||||
Spring AI Advisors may be useful later, but hiding retrieval inside an advisor would make the evidence chain less visible unless we rebuild trace hooks around it.
|
||||
|
||||
## Score Design
|
||||
|
||||
The result object intentionally separates these fields:
|
||||
|
||||
```text
|
||||
score -> compatibility score used by old relevance normalization
|
||||
rawScore -> raw score from the retrieval implementation
|
||||
scoreLabel -> semantic label for rawScore
|
||||
```
|
||||
|
||||
For SDK:
|
||||
|
||||
```text
|
||||
score = L2 distance
|
||||
rawScore = L2 distance
|
||||
scoreLabel = l2_distance
|
||||
```
|
||||
|
||||
For VectorStore:
|
||||
|
||||
```text
|
||||
score = metadata.distance if present
|
||||
rawScore = Spring AI document score
|
||||
scoreLabel = similarity
|
||||
```
|
||||
|
||||
This prevents a subtle bug: if we treat Spring AI similarity as L2 distance, relevance becomes wrong. If we only expose distance, we lose the ability to compare Spring AI behavior. Keeping both makes the migration inspectable.
|
||||
|
||||
## How I Verified It
|
||||
|
||||
I verified at three levels:
|
||||
|
||||
- Unit tests: SDK mode, auto VectorStore mode, fallback mode, category filter, distance metadata mapping.
|
||||
- Live API: `/api/search/similar?query=ERR_TIMEOUT&topK=3`.
|
||||
- Logs: confirmed whether the path was VectorStore success or SDK fallback.
|
||||
|
||||
The live API returned:
|
||||
|
||||
```text
|
||||
scoreLabel = similarity
|
||||
rawScore = Spring AI similarity
|
||||
score = Milvus distance metadata
|
||||
```
|
||||
|
||||
That means the main path was Spring AI VectorStore and compatibility scoring remained stable.
|
||||
|
||||
## What I Would Do Next
|
||||
|
||||
I would not immediately migrate indexing writes. The next responsible steps are:
|
||||
|
||||
- Add a small live acceptance report for several golden queries.
|
||||
- Compare `sdk` and `spring-ai` mode side by side for topK overlap.
|
||||
- Decide whether `VectorIndexService` should move to `VectorStore.add(...)`.
|
||||
- Add query transformation or hybrid retrieval only after we have baseline metrics.
|
||||
|
||||
This staged approach is intentional: first stabilize the read path, then evaluate retrieval quality, then migrate writes if the abstraction proves reliable.
|
||||
|
||||
## Interview Questions And Short Answers
|
||||
|
||||
### Why did you not remove the SDK?
|
||||
|
||||
Because this is a migration, not a rewrite. SDK fallback gives rollback safety and proved useful when VectorStore config was initially wrong.
|
||||
|
||||
### What changed for `lookup_knowledge`?
|
||||
|
||||
The public contract did not change. It still calls `VectorSearchService.searchSimilarDocuments(...)`. The implementation behind that facade changed.
|
||||
|
||||
### How do you know VectorStore is actually used?
|
||||
|
||||
The logs show `Starting Spring AI VectorStore search` followed by `Spring AI VectorStore search complete`. The API response also has `scoreLabel=similarity`, which only comes from the VectorStore path.
|
||||
|
||||
### What was the main bug found during live validation?
|
||||
|
||||
The configured collection name was wrong. Spring AI looked for `business_knowledge`, but the actual Milvus collection was `biz`.
|
||||
|
||||
### What did fallback prove?
|
||||
|
||||
It proved that `auto` mode is resilient: VectorStore failed, SDK search still returned valid results, and the API did not fail.
|
||||
|
||||
### Why is `metadata.distance` important?
|
||||
|
||||
Because `lookup_knowledge` uses L2 distance normalization. Spring AI returns similarity as the main document score, but the Milvus distance is available in metadata. Using it preserves old relevance behavior.
|
||||
|
||||
### Is this full Spring AI RAG now?
|
||||
|
||||
Not yet. It uses Spring AI VectorStore for the main read path, but keeps explicit tools, custom evidence trace, L0 hints, and SDK indexing. That is deliberate because the project values auditability and staged migration.
|
||||
@@ -0,0 +1,199 @@
|
||||
# RAG VectorStore Live Acceptance
|
||||
|
||||
## Purpose
|
||||
|
||||
This note records the live acceptance result for the RAG retrieval refactor.
|
||||
|
||||
The goal of this refactor was not only to add a Spring AI abstraction, but to prove that the production retrieval path can:
|
||||
|
||||
- Prefer Spring AI `VectorStore` for Milvus retrieval.
|
||||
- Preserve the existing Milvus SDK path as fallback.
|
||||
- Keep the `lookup_knowledge` tool contract stable.
|
||||
- Keep L2-distance based relevance normalization compatible.
|
||||
|
||||
## Current Retrieval Shape
|
||||
|
||||
```text
|
||||
lookup_knowledge / /api/search/similar
|
||||
-> VectorSearchService.searchSimilarDocuments(...)
|
||||
-> retrieval.vector-store.mode
|
||||
-> auto
|
||||
-> Spring AI VectorStore
|
||||
-> fallback to Milvus SDK if VectorStore fails
|
||||
-> spring-ai
|
||||
-> Spring AI VectorStore only
|
||||
-> sdk
|
||||
-> Milvus SDK only
|
||||
```
|
||||
|
||||
## Configuration Verified
|
||||
|
||||
The live Milvus/Zilliz database contains the collection:
|
||||
|
||||
```text
|
||||
biz
|
||||
```
|
||||
|
||||
The Spring AI VectorStore configuration was aligned with the existing SDK collection:
|
||||
|
||||
```yaml
|
||||
spring:
|
||||
ai:
|
||||
vectorstore:
|
||||
type: milvus
|
||||
milvus:
|
||||
initialize-schema: false
|
||||
database-name: ${milvus.database}
|
||||
collection-name: biz
|
||||
embedding-dimension: ${milvus.vector-dim}
|
||||
metric-type: L2
|
||||
id-field-name: id
|
||||
content-field-name: content
|
||||
metadata-field-name: metadata
|
||||
embedding-field-name: vector
|
||||
```
|
||||
|
||||
Why this matters: the earlier config used `business_knowledge`, but the SDK path and real collection use `biz`. That mismatch proved the fallback worked, but it also meant VectorStore was not the successful main path until the config was corrected.
|
||||
|
||||
## Commands Used
|
||||
|
||||
Health check:
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Uri "http://127.0.0.1:9900/milvus/health" `
|
||||
-Method Get
|
||||
```
|
||||
|
||||
Observed result:
|
||||
|
||||
```json
|
||||
{
|
||||
"collections": ["biz"],
|
||||
"message": "ok"
|
||||
}
|
||||
```
|
||||
|
||||
Direct retrieval check:
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Uri "http://127.0.0.1:9900/api/search/similar?query=ERR_TIMEOUT&topK=3" `
|
||||
-Method Get
|
||||
```
|
||||
|
||||
Observed result shape:
|
||||
|
||||
```json
|
||||
{
|
||||
"code": 200,
|
||||
"message": "success",
|
||||
"data": [
|
||||
{
|
||||
"id": "f7dff7c8-5665-3145-9f75-ef741528b914",
|
||||
"content": "### ERR_TIMEOUT ...",
|
||||
"score": 0.5662,
|
||||
"rawScore": 0.4337,
|
||||
"scoreLabel": "similarity",
|
||||
"metadata": {
|
||||
"distance": 0.5662,
|
||||
"title": "ERR_TIMEOUT",
|
||||
"category": "api"
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
## What The Logs Proved
|
||||
|
||||
Before collection alignment:
|
||||
|
||||
```text
|
||||
Starting Spring AI VectorStore search
|
||||
SearchRequest collectionName:business_knowledge failed
|
||||
Spring AI VectorStore retrieval failed, falling back to Milvus SDK
|
||||
Starting Milvus SDK search
|
||||
```
|
||||
|
||||
After collection alignment:
|
||||
|
||||
```text
|
||||
Starting Spring AI VectorStore search: query=ERR_TIMEOUT
|
||||
Spring AI VectorStore search complete, candidates=3
|
||||
```
|
||||
|
||||
This proves:
|
||||
|
||||
- `auto` mode really attempts VectorStore first.
|
||||
- The fallback is functional when VectorStore fails.
|
||||
- After config alignment, the main path is Spring AI VectorStore rather than SDK fallback.
|
||||
|
||||
## Score Semantics
|
||||
|
||||
The project keeps three score fields intentionally:
|
||||
|
||||
```text
|
||||
rawScore -> the raw score from the active retrieval implementation
|
||||
scoreLabel -> the semantic meaning of rawScore
|
||||
score -> compatibility score used by existing lookup relevance normalization
|
||||
```
|
||||
|
||||
For SDK retrieval:
|
||||
|
||||
```text
|
||||
rawScore = L2 distance
|
||||
scoreLabel = l2_distance
|
||||
score = L2 distance
|
||||
```
|
||||
|
||||
For Spring AI VectorStore retrieval:
|
||||
|
||||
```text
|
||||
rawScore = Spring AI similarity score
|
||||
scoreLabel = similarity
|
||||
score = Milvus distance metadata when available
|
||||
```
|
||||
|
||||
Why use `metadata.distance` for `score`: `LookupKnowledgeTool` already normalizes relevance from L2 distance. Spring AI Milvus returns similarity as the document score, but also includes the Milvus distance in metadata. Using distance preserves the old relevance behavior while still exposing the new VectorStore score semantics through `rawScore` and `scoreLabel`.
|
||||
|
||||
## Regression Checks
|
||||
|
||||
Targeted tests:
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=VectorSearchServiceTest,LookupKnowledgeToolTest" test
|
||||
```
|
||||
|
||||
Spec validation:
|
||||
|
||||
```powershell
|
||||
openspec.cmd validate rag-knowledge-retrieval --specs
|
||||
openspec.cmd validate rag-retrieval-evaluation --specs
|
||||
```
|
||||
|
||||
Whitespace check:
|
||||
|
||||
```powershell
|
||||
git diff --check
|
||||
```
|
||||
|
||||
Observed result:
|
||||
|
||||
```text
|
||||
All targeted tests passed.
|
||||
All related specs passed.
|
||||
No diff-check errors.
|
||||
```
|
||||
|
||||
## Acceptance Conclusion
|
||||
|
||||
The VectorStore refactor is accepted for the read path:
|
||||
|
||||
- Spring AI VectorStore is integrated and selected in `auto` mode.
|
||||
- The SDK path remains available and was proven by fallback behavior.
|
||||
- The live collection configuration is aligned with the existing Milvus collection.
|
||||
- The `lookup_knowledge` public contract remains stable.
|
||||
- Existing L2-based relevance normalization remains compatible.
|
||||
|
||||
The write/indexing path still uses the Milvus SDK. That is an intentional staged migration decision, not a failed acceptance item.
|
||||
@@ -0,0 +1,81 @@
|
||||
---
|
||||
title: AIOps 告警排障 Runbook
|
||||
keywords: [AIOps, 告警, HighCPUUsage, SlowResponse, payment-service, system-metrics, application-logs]
|
||||
summary: 面向 AIOps 告警诊断的排障步骤,覆盖 Prometheus 活动告警、CLS 日志主题和处理建议。
|
||||
category: troubleshooting
|
||||
---
|
||||
|
||||
# AIOps 告警排障 Runbook
|
||||
|
||||
## 1. 告警输入处理原则
|
||||
|
||||
AIOps 诊断入口有两种触发方式:
|
||||
|
||||
- **有告警 payload**:将 payload 视为已触发告警,围绕 `alertName`、`service`、`severity`、`timeRange` 查询指标、日志和知识库。
|
||||
- **无告警 payload**:先调用 `queryPrometheusAlerts` 获取当前 firing 告警,再选择 P0/P1 或持续时间最长的告警进入诊断。
|
||||
|
||||
最终报告必须基于工具证据,不得凭空编造指标、日志或处理结果。
|
||||
|
||||
## 2. Mock 告警与日志主题映射
|
||||
|
||||
| 告警名 | 典型服务 | 优先日志主题 | 推荐查询 |
|
||||
|---|---|---|---|
|
||||
| HighCPUUsage | payment-service | system-metrics | `cpu_usage:>80 AND service:payment-service` |
|
||||
| HighMemoryUsage | order-service | system-metrics, system-events | `memory_usage:>85` |
|
||||
| SlowResponse | user-service | application-logs, database-slow-query | `duration:>3000 OR slow request` |
|
||||
| ServiceUnavailable | 任意核心服务 | application-logs, system-events | `level:ERROR OR container crash` |
|
||||
|
||||
## 3. HighCPUUsage / payment-service 排障步骤
|
||||
|
||||
### 3.1 现象确认
|
||||
|
||||
先确认 Prometheus 活动告警中是否存在:
|
||||
|
||||
- `alert_name = HighCPUUsage`
|
||||
- `service = payment-service`
|
||||
- CPU 使用率超过 80%
|
||||
- 状态为 firing
|
||||
|
||||
如果 payload 已经提供该告警,也仍需通过指标或日志工具验证。
|
||||
|
||||
### 3.2 指标与日志取证
|
||||
|
||||
推荐工具调用顺序:
|
||||
|
||||
1. `queryPrometheusAlerts`:确认当前活动告警。
|
||||
2. `queryLogs(region=ap-guangzhou, logTopic=system-metrics, query=cpu_usage:>80 AND service:payment-service)`:确认 CPU 使用率、实例和持续时间。
|
||||
3. 如报告中提到 Redis、数据库或下游依赖,再查询 `application-logs` 或对应主题交叉验证。
|
||||
|
||||
### 3.3 根因判断
|
||||
|
||||
可接受的根因结论必须至少满足一项:
|
||||
|
||||
- system-metrics 显示 payment-service 实例 CPU 使用率持续高于阈值。
|
||||
- application-logs 显示与 CPU 飙高同时出现的慢请求、线程池耗尽或依赖超时。
|
||||
- 告警持续时间与日志时间线一致。
|
||||
|
||||
如果只有活动告警,没有日志或指标明细,应输出低置信结论并建议人工确认。
|
||||
|
||||
## 4. 处理建议
|
||||
|
||||
### 临时止血
|
||||
|
||||
- 对 payment-service 做水平扩容,优先扩容受影响实例所在 Deployment。
|
||||
- 对高耗时接口开启限流或降级非核心功能。
|
||||
- 如果近期有发布,检查变更窗口并准备回滚。
|
||||
|
||||
### 根因修复
|
||||
|
||||
- 分析 CPU 热点线程、慢请求接口和依赖调用耗时。
|
||||
- 检查连接池、线程池、缓存穿透和批量任务是否导致 CPU 飙高。
|
||||
- 补充针对 `payment-service` 的 CPU、P95/P99 延迟、错误率和依赖超时联动告警。
|
||||
|
||||
## 5. 报告要求
|
||||
|
||||
告警分析报告至少包含:
|
||||
|
||||
- 活跃告警清单。
|
||||
- 告警根因分析。
|
||||
- 使用过的工具证据:Prometheus 告警、system-metrics 日志、application-logs 或知识库。
|
||||
- 已执行或建议执行的处理方案。
|
||||
- 置信度说明:哪些结论有直接证据,哪些需要人工进一步确认。
|
||||
@@ -1,5 +1,7 @@
|
||||
# 数据库设计文档
|
||||
|
||||
> 当前架构快照:[mvp/architecture/current-mvp-architecture.md](architecture/current-mvp-architecture.md)
|
||||
|
||||
## 📚 文档导航
|
||||
|
||||
### 核心表设计
|
||||
@@ -13,6 +15,9 @@
|
||||
- [知识库检索使用指南](architecture/knowledge-retrieval-usage.md) - 文档编写和使用说明 ⭐新增
|
||||
- [会话管理](architecture/session-management.md) - Redis + MySQL 会话管理
|
||||
- [实施规划](architecture/implementation-plan.md) - 分阶段实施计划
|
||||
- [会话级去重与知识域地图](architecture/session-dedup-knowledge-map.md) - 文档级去重 + Planner 知识域地图注入解决 ISS-001 ⭐新增
|
||||
- [证据评分与用户反馈](architecture/confidence-feedback.md) - evidence_score 规则引擎 + feedback API ⭐新增
|
||||
- [行动记忆与检索归一化](architecture/action-memory-relevance.md) - Executor 行动记忆 + 归一化质量等级解决 ISS-002 ⭐新增
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -0,0 +1,275 @@
|
||||
# 行动记忆与检索质量归一化
|
||||
|
||||
Executor 行动记忆 + 归一化质量等级设计,解决 ISS-002 Executor 无约束重复检索问题。
|
||||
|
||||
---
|
||||
|
||||
## 一、问题背景
|
||||
|
||||
ISS-001 修复文档级去重后,Executor 在单次会话中仍调用 `lookup_knowledge` 20+ 次。根因:
|
||||
|
||||
1. **行动记忆缺失**:Executor 不知道自己已检索过哪些域
|
||||
2. **质量信号缺失**:检索结果没有给 LLM 判断"结果够不够"的信号
|
||||
3. **Prompt 缺少合法出口**:原 prompt 要求"所有外部信息都必须调用工具",LLM 不敢停止检索
|
||||
|
||||
---
|
||||
|
||||
## 二、整体架构
|
||||
|
||||
```
|
||||
lookup_knowledge(query)
|
||||
│
|
||||
├─ Step 1: L0 精确匹配(keywords 索引)
|
||||
├─ Step 2: L1 语义检索(Milvus 向量)
|
||||
├─ Step 3: computeRelevance()
|
||||
│ ├─ 归一化:L2 → similarity [0,1]
|
||||
│ └─ 判定:PRECISE / HIGHLY_RELEVANT / REFERENCE
|
||||
├─ Step 4: RetrievedDocTracker 检查
|
||||
│ ├─ 文档级去重 → isDocRetrieved(sessionId, docKey)
|
||||
│ ├─ 域级检查 → isDomainRetrieved(sessionId, domain)
|
||||
│ └─ 记录 → markRetrieved(sessionId, domain, docKey)
|
||||
└─ Step 5: 返回 LookupResult
|
||||
├─ primary / supplement(原始内容,不含分数)
|
||||
├─ relevanceLevel(PRECISE / HIGHLY_RELEVANT / REFERENCE)
|
||||
├─ completenessHint(兜底信号)
|
||||
└─ retrievedDomainsThisSession(行动记忆)
|
||||
```
|
||||
|
||||
### 设计原则
|
||||
|
||||
| 原则 | 说明 |
|
||||
|------|------|
|
||||
| **Agent 边界清晰** | 不给 Executor 注入 knowledge map,Executor 只知道做了什么,不用知道有什么 |
|
||||
| **分数封装** | L0/L1 原始分数不在 LookupResult 中返回 LLM,只在归一化层内部使用 |
|
||||
| **原始分数只入库** | 原始 L2 距离写进 `tool_invocation.retrieval_details` JSON 用于可观测 |
|
||||
| **软约束 + 硬拦截** | Prompt 约束(软)+ 工具层域级去重(硬)两层防御 |
|
||||
|
||||
---
|
||||
|
||||
## 三、归一化质量等级
|
||||
|
||||
### L2 距离归一化
|
||||
|
||||
BGE-M3 输出为 L2 归一化单位向量(实测范数=1.00000002),L2 距离数学硬上界 = 2.0。
|
||||
|
||||
```
|
||||
similarity = 1 - min(l2Score, maxL2Distance) / maxL2Distance
|
||||
```
|
||||
|
||||
| L2 距离 | similarity | 等级 |
|
||||
|---------|-----------|------|
|
||||
| 0.0 | 1.0 | PRECISE |
|
||||
| 0.383 | 0.8085 | HIGHLY_RELEVANT |
|
||||
| 0.5 | 0.75 | HIGHLY_RELEVANT |
|
||||
| 0.6031 | 0.6984 | REFERENCE |
|
||||
| 1.0 | 0.5 | REFERENCE 边界 |
|
||||
| 2.0+ | 0.0 | 不视为有效结果 |
|
||||
|
||||
### 三等级判定
|
||||
|
||||
| 等级 | 条件 | completenessHint | LLM 行为 |
|
||||
|------|------|-----------------|---------|
|
||||
| PRECISE | L0 matchCount == 1 | "知识库中不存在比上述结果更精准的文档" | 直接使用,禁止再检索 |
|
||||
| HIGHLY_RELEVANT | L0 命中 + similarity ≥ 0.75,或仅 L1 similarity ≥ 0.75 | "当前结果已高度相关,继续检索不太可能找到更精准的文档" | 可综合推理,大概率不需要继续查 |
|
||||
| REFERENCE | 其余命中(similarity ≥ 0.5) | "当前结果为相关参考,如需更精准信息请明确缺少的具体维度" | 可参考,如需更精准请指出缺少的维度后定向补充 |
|
||||
|
||||
### 阈值配置
|
||||
|
||||
```yaml
|
||||
retrieval:
|
||||
normalization:
|
||||
max-l2-distance: 2.0 # L2 距离上界
|
||||
highly-relevant-threshold: 0.75 # similarity ≥ 0.75 → HIGHLY_RELEVANT
|
||||
reference-threshold: 0.5 # similarity ≥ 0.5 → REFERENCE
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 四、行动记忆
|
||||
|
||||
### RetrievedDocTracker 数据结构
|
||||
|
||||
```java
|
||||
// 从单层升级为双层:session → domain → filePath 集合
|
||||
ConcurrentHashMap<String, Map<String, Set<String>>> sessionRetrievals;
|
||||
```
|
||||
|
||||
### API
|
||||
|
||||
| 方法 | 作用 |
|
||||
|------|------|
|
||||
| `markRetrieved(sessionId, domain, filePath)` | 记录一次检索 |
|
||||
| `isDocRetrieved(sessionId, filePath)` | 文档级去重 |
|
||||
| `isDomainRetrieved(sessionId, domain)` | 域级检查 |
|
||||
| `getRetrievedDomains(sessionId)` | 获取已检索域列表 |
|
||||
| `clearSession(sessionId)` | 清理会话记录 |
|
||||
|
||||
### LookupResult 返回
|
||||
|
||||
```java
|
||||
LookupResult.builder()
|
||||
.found(true)
|
||||
.primary(primaryResult)
|
||||
.supplement(supplementResult)
|
||||
.relevanceLevel("HIGHLY_RELEVANT") // PRECISE / HIGHLY_RELEVANT / REFERENCE
|
||||
.completenessHint("当前结果已高度相关...") // 兜底信号
|
||||
.retrievedDomainsThisSession(["infrastructure", "api"]) // 行动记忆
|
||||
.message("...")
|
||||
.build();
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 五、Executor Prompt 约束
|
||||
|
||||
### 4 条检索约束
|
||||
|
||||
1. **判断重复**:基于 `retrievedDomainsThisSession` 判断语义重叠
|
||||
2. **重复了怎么办**:禁止换关键词重查;先指缺少的维度,再定向补充
|
||||
3. **合法出口**:"不查全不会被追责,重复检索才会被惩罚"
|
||||
4. **利用质量信号**:PRECISE → 停止;HIGHLY_RELEVANT + 域已检索 → 禁止;REFERENCE → 指出缺少维度
|
||||
|
||||
### 关键变化
|
||||
|
||||
原有 prompt:"所有需要外部信息的地方,都必须调用对应的工具"
|
||||
→ 改为:"需要外部信息时调用工具,但须遵守下方的检索约束"
|
||||
|
||||
---
|
||||
|
||||
## 六、数据库变更
|
||||
|
||||
### V010
|
||||
|
||||
```sql
|
||||
ALTER TABLE tool_invocation
|
||||
ADD COLUMN relevance_level VARCHAR(20) COMMENT 'PRECISE/HIGHLY_RELEVANT/REFERENCE/DEDUPED',
|
||||
ADD COLUMN dedup_reason VARCHAR(32) COMMENT 'doc_retrieved/domain_retrieved/null';
|
||||
```
|
||||
|
||||
### retrieval_details JSON 扩展
|
||||
|
||||
```json
|
||||
{
|
||||
"l0_titles": ["MySQL 数据库连接池配置", "Redis 缓存配置指南"],
|
||||
"l1_scores": [0.383, 0.4502, 0.7011],
|
||||
"l1_top_score": 0.383,
|
||||
"l1_top_similarity": 0.8085,
|
||||
"relevance_level": "HIGHLY_RELEVANT",
|
||||
"completeness_hint": "当前结果已高度相关,继续检索不太可能找到更精准的文档",
|
||||
"retrieved_domains": ["infrastructure"]
|
||||
}
|
||||
```
|
||||
|
||||
扩展字段使用方式:
|
||||
|
||||
| 字段 | 用途 |
|
||||
|------|------|
|
||||
| `l1_top_score` | 原始 L2 距离最小值(可观测性) |
|
||||
| `l1_top_similarity` | 归一化后的相似度 [0,1] |
|
||||
| `relevance_level` | 归一化质量等级 |
|
||||
| `completeness_hint` | 兜底信号 |
|
||||
| `retrieved_domains` | 已检索域列表 |
|
||||
| `dedup_reason` | 去重原因(如有) |
|
||||
|
||||
---
|
||||
|
||||
## 七、使用场景
|
||||
|
||||
### 场景 1:正常检索
|
||||
|
||||
```
|
||||
用户:数据库连接池怎么配置?
|
||||
|
||||
Executor 内部:
|
||||
1. lookup_knowledge("数据库连接池配置")
|
||||
→ relevanceLevel=HIGHLY_RELEVANT (similarity=0.8085)
|
||||
→ completenessHint="当前结果已高度相关..."
|
||||
→ retrievedDomainsThisSession=["infrastructure"]
|
||||
2. 基于已有信息直接回答,不再检索
|
||||
```
|
||||
|
||||
### 场景 2:行动记忆阻止重复
|
||||
|
||||
```
|
||||
Executor 步骤列表:
|
||||
- 查数据库连接池配置
|
||||
- 查 HikariCP 参数
|
||||
- 查连接池耗尽排查
|
||||
|
||||
实际行为:
|
||||
1. lookup("数据库连接池") → relevance=HIGHLY_RELEVANT, domains=["infrastructure"]
|
||||
2. lookup("HikariCP 参数") → retrievedDomainsThisSession=["infrastructure"]
|
||||
LLM 判断:infrastructure 域已检索过,禁止换关键词重查
|
||||
→ 基于已有信息回答,指出缺少的具体维度
|
||||
3. lookup("连接池耗尽") → 同域,被 prompt 约束拦截或工具层去重拦截
|
||||
```
|
||||
|
||||
### 场景 3:PRECISE 精确匹配
|
||||
|
||||
```
|
||||
用户:ERR_TIMEOUT 是什么?
|
||||
|
||||
Executor 内部:
|
||||
1. lookup_knowledge("ERR_TIMEOUT")
|
||||
→ L0 matchCount=1(唯一精确匹配)
|
||||
→ relevanceLevel=PRECISE
|
||||
→ completenessHint="知识库中不存在比上述结果更精准的文档"
|
||||
2. 直接使用,不再检索
|
||||
```
|
||||
|
||||
### 场景 4:REFERENCE + 定向补充
|
||||
|
||||
```
|
||||
用户:如何排查生产故障?
|
||||
|
||||
Executor 内部:
|
||||
1. lookup_knowledge("故障排查")
|
||||
→ relevanceLevel=REFERENCE (similarity=0.6)
|
||||
→ retrievedDomainsThisSession=["troubleshooting"]
|
||||
2. LLM 判断:信息不足,缺少"日志分析"维度的具体步骤
|
||||
3. lookup_knowledge("日志分析步骤")
|
||||
→ 定向补充,不盲目换关键词
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 八、可观测性
|
||||
|
||||
### 查询质量分布
|
||||
|
||||
```sql
|
||||
SELECT relevance_level, COUNT(*) AS cnt
|
||||
FROM tool_invocation
|
||||
WHERE tool_name = 'lookup_knowledge'
|
||||
GROUP BY relevance_level;
|
||||
```
|
||||
|
||||
### 去重原因分布
|
||||
|
||||
```sql
|
||||
SELECT dedup_reason, COUNT(*) AS cnt
|
||||
FROM tool_invocation
|
||||
WHERE tool_name = 'lookup_knowledge'
|
||||
GROUP BY dedup_reason;
|
||||
```
|
||||
|
||||
### 归一化分数分布
|
||||
|
||||
```sql
|
||||
SELECT
|
||||
JSON_EXTRACT(retrieval_details, '$.l1_top_similarity') AS similarity,
|
||||
COUNT(*) AS cnt
|
||||
FROM tool_invocation
|
||||
WHERE tool_name = 'lookup_knowledge'
|
||||
AND retrieval_details IS NOT NULL
|
||||
GROUP BY similarity
|
||||
ORDER BY similarity;
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 九、扩展方向(Phase 2)
|
||||
|
||||
- **域级硬限流**:`isDomainRetrieved` 已就绪,在 LookupKnowledgeTool 入口直接拦截同域调用,不依赖 LLM 遵守 prompt
|
||||
- **DEDUPED 等级**:去重时单独标记为 DEDUPED 等级,与 REFERENCE 区分
|
||||
- **分数反馈调优**:基于 feedback 数据优化归一化阈值
|
||||
@@ -0,0 +1,157 @@
|
||||
# 证据评分与用户反馈架构
|
||||
|
||||
## 一、整体架构
|
||||
|
||||
```
|
||||
用户对话
|
||||
↓
|
||||
ChatService.executeChat / executeChatComplex
|
||||
↓ SUCCESS 后写入 answer,异步触发
|
||||
EvaluationService.evaluate(sessionId, answer)
|
||||
└─ 读取 tool_invocation 事实 → 规则引擎 → 写 selfEvaluation
|
||||
|
||||
用户提交反馈
|
||||
↓
|
||||
POST /api/feedback { sessionId, feedback: "useful" | "not_useful" }
|
||||
↓
|
||||
FeedbackService.submitFeedback
|
||||
├─ 写 DiagnosisSession.feedback
|
||||
├─ useful → CaseLibraryService.createFromSession → 写 case_library
|
||||
└─ not_useful → 仅写 feedback,status 不变
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 二、评分规则(evidence_score)
|
||||
|
||||
### 定位
|
||||
|
||||
`evidence_score` 衡量的是**证据收集充分度**,不是答案准确性。
|
||||
|
||||
- 能证明的:Agent 是否有尝试收集证据、检索是否命中
|
||||
- 不能证明的:答案是否有幻觉、推理是否正确
|
||||
|
||||
### 数据来源
|
||||
|
||||
规则引擎只消费 `tool_invocation` 表的事实记录,不依赖 LLM 判断。
|
||||
|
||||
### 规则定义
|
||||
|
||||
| 规则名 | 条件 | delta |
|
||||
|---|---|---|
|
||||
| `no_tool_call` | 无任何工具调用 | 直接 0 分,不参与加权 |
|
||||
| `execution_failed` | status = FAILED | 直接 0 分,不参与加权 |
|
||||
| `has_successful_tool_call` | 至少 1 次成功调用 | +30 |
|
||||
| `l0_exact_match` | 任意调用有 L0 精确匹配命中 | +35 |
|
||||
| `l1_semantic_match` | 无 L0 命中但有 L1 语义匹配 | +20 |
|
||||
| `retrieval_no_hit` | 有检索调用但无任何命中 | -10 |
|
||||
| `all_tool_calls_failed` | 全部调用失败 | -20 |
|
||||
|
||||
> L0 和 L1 互斥取高优先级(L0 命中时跳过 L1 分支)。
|
||||
|
||||
### selfEvaluation 字段格式
|
||||
|
||||
```json
|
||||
{
|
||||
"evidence_score": 65,
|
||||
"source": "rule",
|
||||
"factors": [
|
||||
{"name": "has_successful_tool_call", "delta": 30, "description": "有成功的工具调用(20次)"},
|
||||
{"name": "l0_exact_match", "delta": 35, "description": "L0 精确匹配命中"}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 说明 |
|
||||
|---|---|
|
||||
| `evidence_score` | 0-100 整数 |
|
||||
| `source` | 当前固定为 `"rule"`;预留 `"llm"` 供后续扩展 |
|
||||
| `factors` | 命中的规则列表,含 name / delta / description |
|
||||
| `llm_opinion` | 预留字段(未实现),LLM 观点叠加时在此扩展 |
|
||||
|
||||
### 已知边界
|
||||
|
||||
- 非检索工具(DateTimeTools、QueryMetricsTools 等)不写 `tool_invocation`,这类 session 的 evidence_score = 0,属于设计边界
|
||||
- 评分为异步写入(`@Async`),失败时 `selfEvaluation` 保持 null,前端需处理 null
|
||||
|
||||
---
|
||||
|
||||
## 三、反馈机制
|
||||
|
||||
### API
|
||||
|
||||
```
|
||||
POST /api/feedback
|
||||
Content-Type: application/json
|
||||
|
||||
{
|
||||
"sessionId": "xxx",
|
||||
"feedback": "useful" | "not_useful"
|
||||
}
|
||||
```
|
||||
|
||||
**响应**
|
||||
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"message": "反馈已记录",
|
||||
"caseId": "uuid 或 null"
|
||||
}
|
||||
```
|
||||
|
||||
### 后端行为
|
||||
|
||||
| feedback 值 | 操作 |
|
||||
|---|---|
|
||||
| `useful` | 写 `DiagnosisSession.feedback = "useful"`,生成 `CaseLibrary` 记录,返回 caseId |
|
||||
| `not_useful` | 写 `DiagnosisSession.feedback = "not_useful"`,status 不变 |
|
||||
| 其他值 | 返回 HTTP 400 |
|
||||
|
||||
### 重要设计决策
|
||||
|
||||
**BAD_CASE 不改 status 字段**
|
||||
|
||||
`status` 表示执行状态(RUNNING/SUCCESS/FAILED),是独立维度,不能被质量标签覆盖。
|
||||
查询 BadCase 使用:`WHERE feedback = 'not_useful'`
|
||||
|
||||
**useful 触发案例沉淀规则**
|
||||
|
||||
| CaseLibrary 字段 | 来源 |
|
||||
|---|---|
|
||||
| caseId | UUID |
|
||||
| diagnosisId | DiagnosisSession.sessionId |
|
||||
| sourceType | AUTO |
|
||||
| faultCategory | GENERAL(暂时,后续人工补充) |
|
||||
| title | query 前 100 字符 |
|
||||
| rootCause / solution | DiagnosisSession.answer(完整答案) |
|
||||
| createdBy | "system" |
|
||||
|
||||
**幂等性**:同一 sessionId 重复提交 useful,返回已有 caseId,不重复插入 case_library。
|
||||
|
||||
---
|
||||
|
||||
## 四、数据库变更
|
||||
|
||||
### V008(新增)
|
||||
|
||||
```sql
|
||||
ALTER TABLE diagnosis_session ADD COLUMN answer LONGTEXT COMMENT 'Agent 返回给用户的完整答案';
|
||||
```
|
||||
|
||||
### diagnosis_session 关键字段
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
|---|---|---|
|
||||
| `answer` | LONGTEXT | Agent 完整回答,useful 案例沉淀的内容来源 |
|
||||
| `self_evaluation` | JSON | 证据评分结果,格式见上 |
|
||||
| `feedback` | VARCHAR(16) | useful / not_useful / null |
|
||||
| `status` | VARCHAR(16) | 执行状态,不受 feedback 影响 |
|
||||
|
||||
---
|
||||
|
||||
## 五、扩展方向(Phase 2)
|
||||
|
||||
- **LLM 观点层**:在 `selfEvaluation` 的 `llm_opinion` 字段叠加 LLM 结构化观点(has_root_cause、has_solution 等),作为独立 factors,不改变现有规则逻辑
|
||||
- **案例结构化字段**:useful 触发时自动提取 faultCategory / errorCode,替代暂时的 GENERAL
|
||||
- **重复召回问题**:Executor Prompt 约束或工具层 session 维度去重(见 [ISS-001](../issues/ISS-001-duplicate-retrieval.md))
|
||||
@@ -0,0 +1,220 @@
|
||||
# Current MVP Architecture Snapshot
|
||||
|
||||
**Updated**: 2026-07-05
|
||||
|
||||
This document records the current runnable MVP architecture. Older architecture notes in this folder still represent design history; this file should be read as the current snapshot for demos, interviews, and next-step planning.
|
||||
|
||||
## 1. Positioning
|
||||
|
||||
The MVP is an Agent engineering project for traceable troubleshooting, not a generic chatbot.
|
||||
|
||||
Core goals:
|
||||
|
||||
- Support normal chat-based diagnosis.
|
||||
- Support AIOps alert-triggered diagnosis.
|
||||
- Keep tool calls explicit and traceable.
|
||||
- Keep RAG retrieval observable through `lookup_knowledge`.
|
||||
- Persist enough execution evidence for replay, evaluation, and interview explanation.
|
||||
|
||||
## 2. Runtime Architecture
|
||||
|
||||
```text
|
||||
HTTP API
|
||||
-> ChatService / AiOpsService
|
||||
-> Agent orchestration
|
||||
-> Supervisor / Planner / Executor / Verifier
|
||||
-> Tools
|
||||
-> lookup_knowledge
|
||||
-> query_logs
|
||||
-> query_metrics
|
||||
-> other diagnosis tools
|
||||
-> Persistence
|
||||
-> diagnosis_session
|
||||
-> agent_step
|
||||
-> tool_invocation
|
||||
-> Trace API
|
||||
-> DiagnosisTraceService
|
||||
```
|
||||
|
||||
Current entry points:
|
||||
|
||||
- `ChatService`: user-driven troubleshooting and follow-up diagnosis.
|
||||
- `AiOpsService`: alert-driven diagnosis, including payload mode and auto-discovery mode.
|
||||
- `DiagnosisTraceService`: trace view of session, steps, tool calls, and self-evaluation.
|
||||
|
||||
## 3. Chat Diagnosis Flow
|
||||
|
||||
```text
|
||||
User question
|
||||
-> ChatService
|
||||
-> simple response or diagnosis flow
|
||||
-> Planner creates investigation direction
|
||||
-> Executor calls tools for evidence
|
||||
-> lookup_knowledge
|
||||
-> query_logs
|
||||
-> query_metrics
|
||||
-> Verifier checks final diagnosis quality
|
||||
-> self_evaluation.verifier_evaluation
|
||||
-> diagnosis trace
|
||||
```
|
||||
|
||||
The chat path uses the LLM verifier as the main quality gate. The verifier result is persisted under `diagnosis_session.self_evaluation.verifier_evaluation`.
|
||||
|
||||
## 4. AIOps Diagnosis Flow
|
||||
|
||||
```text
|
||||
AIOps request
|
||||
-> AiOpsService
|
||||
-> payload mode or auto-discovery mode
|
||||
-> build alert-focused diagnosis prompt
|
||||
-> append recommended lookup_knowledge query when payload exists
|
||||
-> Agent diagnosis flow
|
||||
-> Supervisor / Planner / Executor
|
||||
-> evidence tools
|
||||
-> final report
|
||||
-> AiOpsRuleEvaluationService
|
||||
-> self_evaluation.aiops_rule_evaluation
|
||||
-> diagnosis trace
|
||||
```
|
||||
|
||||
AIOps keeps two modes:
|
||||
|
||||
- Payload mode: the request already contains alert fields such as alert name, service, metric, severity, and symptom. The system builds a recommended knowledge query from these fields.
|
||||
- Auto-discovery mode: the system follows the original alert-discovery behavior and lets the Agent collect alert context through tools.
|
||||
|
||||
The AIOps verifier is currently lightweight and rule-based. It checks:
|
||||
|
||||
- Whether the final report exists.
|
||||
- Whether the result stays focused on the alert payload when payload exists.
|
||||
- Whether evidence tools were used, especially `lookup_knowledge`, `query_logs`, and `query_metrics`.
|
||||
|
||||
## 5. RAG Architecture
|
||||
|
||||
```text
|
||||
lookup_knowledge
|
||||
-> L0 domain/entity hint
|
||||
-> matched domain
|
||||
-> matched keywords/entities
|
||||
-> metadata filter signal
|
||||
-> VectorSearchService
|
||||
-> Spring AI VectorStore path
|
||||
-> Milvus SDK fallback path
|
||||
-> evidence post-processing
|
||||
-> score / rawScore / scoreLabel
|
||||
-> source metadata
|
||||
-> title / breadcrumb / content evidence block
|
||||
-> tool_invocation record
|
||||
```
|
||||
|
||||
Important decisions:
|
||||
|
||||
- `lookup_knowledge` remains an explicit Agent tool. It is not replaced by an implicit chat Advisor because the project needs visible Agent decision-making.
|
||||
- L0 is retained but downgraded. It is a domain/entity hint and explainability signal, not the final recall decision.
|
||||
- L1 retrieval now goes through `VectorSearchService`.
|
||||
- Spring AI `VectorStore` is the preferred retrieval path.
|
||||
- The original Milvus SDK path is retained as fallback and compatibility path.
|
||||
- `title`, `breadcrumb`, and `content` participate in embedding text so chunk context is less likely to be lost.
|
||||
- Retrieval output keeps compatibility fields: `score`, `rawScore`, and `scoreLabel`.
|
||||
|
||||
Vector retrieval modes:
|
||||
|
||||
```text
|
||||
retrieval.vector-store.mode=auto # Prefer Spring AI VectorStore, fallback to SDK
|
||||
retrieval.vector-store.mode=spring-ai # Use Spring AI VectorStore only
|
||||
retrieval.vector-store.mode=sdk # Use original Milvus SDK path
|
||||
```
|
||||
|
||||
## 6. Persistence And Trace
|
||||
|
||||
Current trace-related persistence:
|
||||
|
||||
```text
|
||||
diagnosis_session
|
||||
-> final_report
|
||||
-> self_evaluation
|
||||
-> verifier_evaluation
|
||||
-> aiops_rule_evaluation
|
||||
|
||||
agent_step
|
||||
-> role
|
||||
-> step input/output
|
||||
-> execution order
|
||||
|
||||
tool_invocation
|
||||
-> tool_name
|
||||
-> query
|
||||
-> retrieval_layer
|
||||
-> retrieval_details
|
||||
-> evidence blocks
|
||||
-> duration
|
||||
```
|
||||
|
||||
Trace API aggregates these records into a session-level view:
|
||||
|
||||
- Agent step sequence.
|
||||
- Tool calls and retrieval details.
|
||||
- Final diagnosis report.
|
||||
- Chat verifier status.
|
||||
- AIOps rule verifier status.
|
||||
|
||||
## 7. Quality Gates
|
||||
|
||||
Current quality gates:
|
||||
|
||||
- Chat verifier: LLM-based final answer verification for normal diagnosis.
|
||||
- AIOps rule verifier: lightweight deterministic checks for alert-focused diagnosis.
|
||||
- Diagnosis eval baseline: fixture-based evaluation for trace and evidence behavior.
|
||||
- RAG retrieval baseline: golden query set with offline baseline report.
|
||||
- Live RAG acceptance: post-reindex script for validating retrieval against the running stack.
|
||||
|
||||
These gates are intentionally layered. The MVP proves the Agent chain can produce evidence, persist it, and be inspected after execution.
|
||||
|
||||
## 8. Current Completion State
|
||||
|
||||
Completed for the current MVP stage:
|
||||
|
||||
- Explicit `lookup_knowledge` Agent tool.
|
||||
- L0 + L1 retrieval shape retained.
|
||||
- L0 downgraded to domain/entity hint.
|
||||
- Spring AI VectorStore retrieval path integrated.
|
||||
- Milvus SDK fallback retained.
|
||||
- RAG evidence post-processing added.
|
||||
- Breadcrumb/title/content embedding text improved.
|
||||
- RAG offline baseline and live acceptance script added.
|
||||
- AIOps payload query augmentation added.
|
||||
- AIOps lightweight verifier added.
|
||||
- Trace summary includes both chat verifier and AIOps verifier signals.
|
||||
|
||||
Deferred future enhancements:
|
||||
|
||||
- LLM QueryTransformer / MultiQuery.
|
||||
- BM25, RRF, and reranker.
|
||||
- Neighbor chunk or section-level context expansion.
|
||||
- VectorStore write path migration.
|
||||
- Full LLM-based AIOps verifier.
|
||||
- More complete golden set for recall, MRR, and nDCG metrics.
|
||||
|
||||
## 9. Key Code References
|
||||
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/AiOpsRuleEvaluationService.java`
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/VectorIndexService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/SpringAiVectorStoreSidecarService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
|
||||
- `src/main/java/com/superbiz/agent/service/SelfEvaluationMergeService.java`
|
||||
|
||||
## 10. Supporting Materials
|
||||
|
||||
- `mvp/issues/rag-refactor-plan.md`
|
||||
- `eval/rag-retrieval/README.md`
|
||||
- `scripts/eval_rag_live_acceptance.py`
|
||||
- `interview/rag-refactor-story.md`
|
||||
- `interview/rag-vectorstore-interview-notes.md`
|
||||
- `interview/rag-retrieval-quality-report.md`
|
||||
- `interview/rag-breadcrumb-embedding-acceptance.md`
|
||||
- `interview/aiops-query-augmentation.md`
|
||||
- `interview/aiops-lightweight-verifier.md`
|
||||
@@ -1,409 +1,421 @@
|
||||
# 知识库检索架构说明
|
||||
# 知识库检索架构(L0 + L1)
|
||||
|
||||
## 一、架构位置
|
||||
**更新日期**: 2026-06-25
|
||||
|
||||
知识库检索是 Agent 工具层的一部分,为所有 Agent 提供知识查询能力。
|
||||
---
|
||||
|
||||
## 一、概述
|
||||
|
||||
`LookupKnowledgeTool` 实现两阶段混合检索:
|
||||
|
||||
- **L0 精确匹配**:基于内存索引的关键词匹配(< 10ms),索引从数据库加载
|
||||
- **L1 语义检索**:基于 Milvus 向量数据库的相似度搜索(200-500ms)
|
||||
|
||||
---
|
||||
|
||||
## 二、完整流程
|
||||
|
||||
```
|
||||
Agent 层
|
||||
├── Supervisor Agent
|
||||
├── Planner Agent
|
||||
├── SubAgents (ExternalApi, InternalError, Database...)
|
||||
└── Verifier Agent
|
||||
↓ 调用
|
||||
工具层 (Tools)
|
||||
├── searchDoc (文档检索 - L1 向量检索)
|
||||
├── lookup_knowledge (混合检索 - L0+L1) ← 新增
|
||||
├── queryLogs (日志查询)
|
||||
├── queryTrace (链路追踪)
|
||||
└── queryOrder (订单查询)
|
||||
↓ 依赖
|
||||
服务层 (Services)
|
||||
├── VectorSearchService (L1 语义检索 - Milvus)
|
||||
├── KnowledgeIndexService (L0 精确匹配 - 内存) ← 新增
|
||||
├── FrontmatterParser (元数据解析) ← 新增
|
||||
└── DocumentManagementService (文档管理)
|
||||
↓ 持久化
|
||||
数据层
|
||||
├── MySQL (api_document + metadata 字段) ← 增强
|
||||
├── Milvus (向量索引)
|
||||
└── Local Files (knowledge_base/) ← 新增
|
||||
用户查询
|
||||
│
|
||||
▼
|
||||
┌─────────────────────────────┐
|
||||
│ L0: 关键词精确匹配 │ (< 10ms)
|
||||
│ • 从内存索引做关键词匹配 │
|
||||
│ • 索引来源: ApiDocument DB│
|
||||
└──────────┬──────────────────┘
|
||||
│
|
||||
▼
|
||||
┌──────┴──────┐
|
||||
│ matches=1 │ ← 唯一匹配(高置信度)
|
||||
└──────┬──────┘
|
||||
│
|
||||
▼
|
||||
┌──────────────┐ ┌──────────────────┐
|
||||
│ 跳过 L1 │ │ L0 返回正文摘要 │
|
||||
│ 置信度: high │ │ buildCompactSummary│
|
||||
└──────────────┘ └──────────────────┘
|
||||
|
||||
|
||||
┌──────┴──────┐
|
||||
│ matches=0 │ ← 无匹配
|
||||
└──────┬──────┘
|
||||
│
|
||||
▼
|
||||
┌──────────────┐ ┌──────────────────┐
|
||||
│ 触发 L1 │ │ L0 无结果 │
|
||||
│ L1 语义检索 │ │ 仅有 L1 补充结果 │
|
||||
└──────────────┘ └──────────────────┘
|
||||
|
||||
|
||||
┌──────┴──────┐
|
||||
│ matches>=2 │ ← 多匹配
|
||||
└──────┬──────┘
|
||||
│
|
||||
▼
|
||||
┌──────────────┐
|
||||
│ 触发 L1 │
|
||||
│ L1 语义检索 │
|
||||
└──────┬──────┘
|
||||
│
|
||||
┌─────┴─────┐
|
||||
│ ║ │
|
||||
▼ ▼
|
||||
L1 有结果 L1 无结果
|
||||
│ │
|
||||
▼ ▼
|
||||
元数据摘要 正文摘要
|
||||
(不读文件) (读文件)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 二、L0+L1 混合检索架构
|
||||
## 三、L0 返回内容策略
|
||||
|
||||
### 2.1 检索流程
|
||||
根据匹配场景决定 L0 返回给 LLM 的上下文内容量。
|
||||
|
||||
### 3.1 唯一匹配(高置信度,matches=1)
|
||||
|
||||
**策略**: `buildCompactSummary()`
|
||||
|
||||
L1 被跳过,LLM 只有 L0 信息来源,需要提供足够的正文内容。
|
||||
|
||||
```
|
||||
Agent 调用 lookup_knowledge(query)
|
||||
↓
|
||||
┌─────────────────────────────────────────┐
|
||||
│ LookupKnowledgeTool │
|
||||
│ (工具入口) │
|
||||
└────────────┬────────────────────────────┘
|
||||
│
|
||||
↓
|
||||
┌────────────────┐
|
||||
│ Step 1: L0 精确匹配 │ < 10ms
|
||||
│ (内存索引) │
|
||||
└────────┬───────────┘
|
||||
│
|
||||
┌───────┴────────┐
|
||||
│ │
|
||||
唯一匹配 多个/零个匹配
|
||||
│ │
|
||||
↓ ↓
|
||||
高置信度 低置信度
|
||||
(不调用L1) (调用L1补充)
|
||||
│ │
|
||||
│ ┌──────────────────┐
|
||||
│ │ Step 2: L1 语义检索 │ 200-500ms
|
||||
│ │ (Milvus) │
|
||||
│ └──────────┬─────────┘
|
||||
│ │
|
||||
└────────┬───────────┘
|
||||
↓
|
||||
┌─────────────────────┐
|
||||
│ Step 3: 组装结果 │
|
||||
│ primary + supplement │
|
||||
└─────────────────────┘
|
||||
↓
|
||||
返回给 Agent
|
||||
文档: 支付网关错误码定义
|
||||
摘要: 记录了支付网关所有核心错误码的含义及排查方向
|
||||
章节:
|
||||
- 超时类错误
|
||||
- 业务类错误
|
||||
- 签名类错误
|
||||
---
|
||||
**含义**:支付网关请求超时
|
||||
**常见原因**:网络延迟、第三方服务响应慢
|
||||
...
|
||||
```
|
||||
|
||||
### 2.2 数据流
|
||||
| 组成部分 | 说明 | 大小 |
|
||||
|---------|------|------|
|
||||
| title + summary | 从内存索引获取 | ~50-100 字符 |
|
||||
| 章节标题列表 | 从文件解析 `##` 标题 | ~50-200 字符 |
|
||||
| 正文片段 | 去 frontmatter/标题行/空行,短文档 800/长文档 500 字符截断 | ~300-800 字符 |
|
||||
| **总计** | | **~400-1000 字符** |
|
||||
|
||||
### 3.2 多匹配 + L1 有结果
|
||||
|
||||
**策略**: `buildMetadataOnlySummary()`
|
||||
|
||||
L1 已有语义内容片段,L0 仅需告知 LLM 命中了哪些文档。**不读文件**,仅用内存索引。
|
||||
|
||||
```
|
||||
文档上传流程:
|
||||
POST /api/documents/upload
|
||||
↓
|
||||
DocumentManagementService.uploadDocument()
|
||||
↓
|
||||
1. 文本提取
|
||||
2. 保存原始文件 → knowledge_base/{category}/{filename}
|
||||
3. 解析 frontmatter (FrontmatterParser)
|
||||
4. 分块 → 向量化 → Milvus 索引 (L1)
|
||||
5. 元数据存 MySQL (metadata 字段 JSON)
|
||||
6. 更新 L0 内存索引 (KnowledgeIndexService)
|
||||
↓
|
||||
完成
|
||||
|
||||
文档查询流程:
|
||||
Agent 调用 lookup_knowledge("ERR_TIMEOUT")
|
||||
↓
|
||||
KnowledgeIndexService.exactMatch()
|
||||
↓
|
||||
遍历内存索引 (keywords 精确匹配)
|
||||
↓
|
||||
找到唯一匹配 → 读取本地文件 (前 2000 字符)
|
||||
↓
|
||||
返回 primary (高置信度)
|
||||
文档: 支付网关错误码定义
|
||||
摘要: 记录了支付网关所有核心错误码的含义及排查方向
|
||||
关键词: ERR_TIMEOUT, 超时, 支付网关
|
||||
来源: api/payment-errors.md
|
||||
```
|
||||
|
||||
| 组成部分 | 说明 | 大小 |
|
||||
|---------|------|------|
|
||||
| title + summary + keywords | 全部从内存索引获取 | ~100-200 字符 |
|
||||
| **总计** | | **~100-200 字符** |
|
||||
|
||||
### 3.3 多匹配 + L1 无结果
|
||||
|
||||
**策略**: `buildCompactSummary()`(同 3.1)
|
||||
|
||||
L1 未返回结果, L0 作为兜底提供正文内容。
|
||||
|
||||
---
|
||||
|
||||
## 三、核心组件说明
|
||||
## 四、决策矩阵
|
||||
|
||||
### 3.1 FrontmatterParser
|
||||
|
||||
**职责**:解析 Markdown 文件头的 YAML frontmatter
|
||||
|
||||
**输入**:
|
||||
```markdown
|
||||
---
|
||||
title: 支付网关错误码定义
|
||||
keywords: [ERR_TIMEOUT, 超时, 支付网关]
|
||||
summary: 记录了支付网关所有核心错误码的含义及排查方向
|
||||
category: api
|
||||
---
|
||||
|
||||
# 正文内容
|
||||
```
|
||||
needFullContent = highConfidence || !hasL1
|
||||
```
|
||||
|
||||
**输出**:
|
||||
| 场景 | matches | L1 结果 | needFullContent | L0 策略 | 是否读文件 | 上下文大小 |
|
||||
|------|:-------:|:--------:|:---------------:|---------|:---------:|:--------:|
|
||||
| 唯一匹配 | 1 | 未执行 | true | `buildCompactSummary` | 是 | ~600 字符 |
|
||||
| 多匹配 + L1 有结果 | 2+ | 有 | false | `buildMetadataOnlySummary` | **否** | ~150 字符 |
|
||||
| 多匹配 + L1 无结果 | 2+ | 无 | true | `buildCompactSummary` | 是 | ~600 字符 |
|
||||
| 无匹配 | 0 | 有 | — | 无 L0,仅 L1 | 否 | 0 |
|
||||
|
||||
---
|
||||
|
||||
## 五、代码结构
|
||||
|
||||
```
|
||||
LookupKnowledgeTool
|
||||
├── lookupKnowledge(query) # 入口:编排 L0 + L1
|
||||
├── buildResult(l0, l1, confidence) # 组装结果,选择摘要策略
|
||||
├── buildCompactSummary(entry) # 元数据 + 章节 + 正文片段(读文件)
|
||||
├── buildMetadataOnlySummary(entry) # 仅元数据(不读文件)
|
||||
├── countMdHeadings(content) # 统计章节数(日志用)
|
||||
└── extractFirstMeaningfulLine(...) # 提取首个有意义文本行(日志用)
|
||||
```
|
||||
|
||||
### 关键逻辑(buildResult)
|
||||
|
||||
```java
|
||||
Frontmatter {
|
||||
title: "支付网关错误码定义",
|
||||
keywords: ["ERR_TIMEOUT", "超时", "支付网关"],
|
||||
summary: "...",
|
||||
category: "api"
|
||||
}
|
||||
boolean needFullContent = highConfidence || !hasL1;
|
||||
String content = needFullContent
|
||||
? buildCompactSummary(first)
|
||||
: buildMetadataOnlySummary(first);
|
||||
```
|
||||
|
||||
### 3.2 KnowledgeIndexService
|
||||
---
|
||||
|
||||
**职责**:维护 L0 内存索引,提供精确关键词匹配
|
||||
## 六、日志输出示例
|
||||
|
||||
**核心方法**:
|
||||
- `@PostConstruct loadIndex()` - 启动时扫描 knowledge_base/
|
||||
- `exactMatch(String query)` - 精确匹配(不区分大小写)
|
||||
- `readDocument(String filePath, int maxChars)` - 读取文档内容
|
||||
- `addToIndex(KnowledgeEntry entry)` - 添加到索引
|
||||
- `removeFromIndex(String filePath)` - 从索引移除
|
||||
### 多匹配场景(matches=2, L1 有结果)
|
||||
|
||||
**数据结构**:
|
||||
```java
|
||||
List<KnowledgeEntry> knowledgeIndex = new CopyOnWriteArrayList<>();
|
||||
|
||||
KnowledgeEntry {
|
||||
filePath: "knowledge_base/api/payment-errors.md",
|
||||
title: "支付网关错误码定义",
|
||||
keywords: ["ERR_TIMEOUT", "超时", "支付网关"],
|
||||
summary: "...",
|
||||
category: "api"
|
||||
}
|
||||
```
|
||||
[L0 精确匹配] 完成: matches=2, time=3ms
|
||||
[置信度判断] highConfidence=false, reason=多个或零个匹配
|
||||
[L1 语义检索] L0非唯一匹配,触发L1语义检索...
|
||||
[L1 语义检索] 完成: matches=1, time=245ms
|
||||
----------------------------------------
|
||||
<<< [工具返回] lookup_knowledge
|
||||
<<< [L0 主结果] 标题: 支付网关错误码定义
|
||||
<<< [L0 主结果] 摘要: 记录了支付网关所有核心错误码的含义及排查方向 ← 仅元数据
|
||||
<<< [L0 主结果] 内容: 126 字符, 0 个章节 ← 约150字符
|
||||
<<< [L1 补充] 相似度: 0.8234
|
||||
<<< [L1 补充] 内容片段: 支付网关请求超时... ← L1 提供具体内容
|
||||
```
|
||||
|
||||
### 3.3 LookupKnowledgeTool
|
||||
### 唯一匹配场景(matches=1, 跳过 L1)
|
||||
|
||||
**职责**:L0+L1 混合检索工具,Agent 可调用
|
||||
|
||||
**工具定义**:
|
||||
```java
|
||||
@Tool(description = "查询知识库文档。优先精确匹配关键词,未命中或多个匹配时自动补充语义相关片段。" +
|
||||
"参数 query: 查询关键词,例如 'ERR_TIMEOUT'、'支付网关超时'")
|
||||
public LookupResult lookupKnowledge(String query)
|
||||
```
|
||||
[L0 精确匹配] 完成: matches=1, time=2ms
|
||||
[置信度判断] highConfidence=true, reason=唯一匹配
|
||||
[L1 语义检索] L0唯一匹配,跳过L1检索
|
||||
----------------------------------------
|
||||
<<< [工具返回] lookup_knowledge
|
||||
<<< [L0 主结果] 标题: 支付网关错误码定义
|
||||
<<< [L0 主结果] 摘要: 记录了支付网关所有核心错误码的含义及排查方向
|
||||
<<< [L0 主结果] 内容: 725 字符, 3 个章节 ← 约700字符
|
||||
```
|
||||
|
||||
**返回格式**:
|
||||
---
|
||||
|
||||
## 七、MVP 效率评估 & 改进方向
|
||||
|
||||
### 7.1 当前效率评估
|
||||
|
||||
| 维度 | 评分 | 说明 |
|
||||
|------|:----:|------|
|
||||
| L0 匹配速度 | ★★★★★ | 内存索引,< 10ms,几乎没有优化空间 |
|
||||
| L1 检索速度 | ★★★★☆ | Milvus 向量检索,200-500ms,取决于数据量 |
|
||||
| L0 匹配准确率 | ★★☆☆☆ | 子串匹配,无排序无评分,匹配即返回 |
|
||||
| L1 检索准确率 | ★★★☆☆ | 语义相似度,但分块缺少上下文信息 |
|
||||
| 召回率(查全) | ★★★☆☆ | L0+L1 两阶段覆盖大多数场景,但缺乏融合重排 |
|
||||
| 上下文利用率 | ★★★★☆ | 根据场景动态控制 L0 内容量,已优化 |
|
||||
| **综合** | **★★★☆☆** | **MVP 可用,但检索质量有提升空间** |
|
||||
|
||||
### 7.2 关键瓶颈
|
||||
|
||||
#### 瓶颈 1:分块丢失上下文(✅ 已修复—见下方 7.5)
|
||||
|
||||
当前每个 Chunk 只记录最近的 `##` 标题:
|
||||
|
||||
```json
|
||||
{
|
||||
"found": true,
|
||||
"primary": {
|
||||
"content": "文档内容(前 2000 字符)",
|
||||
"source": "knowledge_base/api/payment-errors.md",
|
||||
"matchType": "exact_L0",
|
||||
"confidence": "high"
|
||||
},
|
||||
"supplement": {
|
||||
"content": "语义相关片段(L1)",
|
||||
"source": "metadata",
|
||||
"matchType": "semantic_L1"
|
||||
}
|
||||
"content": "**含义**:支付网关请求超时\n**常见原因**:网络延迟",
|
||||
"title": "超时类错误",
|
||||
"chunkIndex": 2
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
LLM 收到这个片段时**不知道**它属于"支付网关错误码定义"这个文档,也不知道具体错误码名称是 ERR_TIMEOUT。如果同时检索了多个文档的片段,LLM 容易混淆。
|
||||
|
||||
## 四、与现有架构的集成
|
||||
#### 瓶颈 2:L0 关键词匹配过于简单
|
||||
|
||||
### 4.1 Agent 使用场景
|
||||
当前 `KnowledgeIndexService.matchesKeywords()` 只做子串包含匹配,没有:
|
||||
- 排序/评分(多个匹配时按什么顺序?)
|
||||
- 权重(标题匹配 > 正文匹配)
|
||||
- 部分匹配("timeout" 匹配 "ERR_TIMEOUT")
|
||||
|
||||
**ExternalApiSubAgent** (接口专家):
|
||||
```
|
||||
诊断步骤:
|
||||
1. 提取错误码(如 "ERR_TIMEOUT")
|
||||
2. 调用 lookup_knowledge("ERR_TIMEOUT")
|
||||
3. 获得完整错误码定义和排查方向
|
||||
4. 结合日志/链路追踪进行分析
|
||||
```
|
||||
#### 瓶颈 3:L0 和 L1 无交叉融合
|
||||
|
||||
**DatabaseSubAgent** (数据库专家):
|
||||
```
|
||||
诊断步骤:
|
||||
1. 识别数据库问题(如 "连接池满")
|
||||
2. 调用 lookup_knowledge("HikariCP")
|
||||
3. 获得连接池配置最佳实践
|
||||
4. 提供优化建议
|
||||
```
|
||||
|
||||
**Planner Agent** (规划者):
|
||||
```
|
||||
规划阶段:
|
||||
1. 分析问题类型
|
||||
2. 调用 lookup_knowledge("故障诊断")
|
||||
3. 获得标准诊断流程
|
||||
4. 制定排查策略
|
||||
```
|
||||
|
||||
### 4.2 与现有工具对比
|
||||
|
||||
| 工具 | 检索方式 | 响应时间 | 适用场景 | 置信度 |
|
||||
|------|---------|---------|---------|--------|
|
||||
| searchDoc | L1 语义检索 | 200-500ms | 模糊查询、语义理解 | 依赖相似度 |
|
||||
| lookup_knowledge | L0+L1 混合 | < 10ms (高置信) | 精确关键词 + 语义补充 | high/low |
|
||||
|
||||
**推荐使用策略**:
|
||||
- 已知精确关键词(错误码、配置项)→ `lookup_knowledge`
|
||||
- 模糊描述、需要语义理解 → `searchDoc`
|
||||
两阶段检索结果只是简单的"1位L0 + 1位L1"拼接,没有:
|
||||
- RRF 或加权融合重排
|
||||
- 重复内容去重
|
||||
- 根据相关性选择 top-K
|
||||
|
||||
---
|
||||
|
||||
## 五、数据库变更
|
||||
### 7.3 改进方向分析
|
||||
|
||||
### 5.1 api_document 表增强
|
||||
#### 方向 A:面包屑导航(Chunk 携带层级上下文)
|
||||
|
||||
**新增字段**:
|
||||
```sql
|
||||
ALTER TABLE api_document
|
||||
ADD COLUMN metadata TEXT COMMENT 'Frontmatter 元数据 (JSON)';
|
||||
**做法**:分块时记录完整的标题层级路径作为 `breadcrumb`。
|
||||
|
||||
当前分块 metadata:
|
||||
```json
|
||||
{ "title": "超时类错误" }
|
||||
```
|
||||
|
||||
**字段说明**:
|
||||
- 类型:TEXT(最大 64KB)
|
||||
- 格式:JSON 字符串
|
||||
- 内容:frontmatter 解析结果
|
||||
|
||||
**示例数据**:
|
||||
改进后:
|
||||
```json
|
||||
{
|
||||
"title": "支付网关错误码定义",
|
||||
"keywords": ["ERR_TIMEOUT", "超时", "支付网关"],
|
||||
"summary": "记录了支付网关所有核心错误码的含义及排查方向",
|
||||
"title": "超时类错误",
|
||||
"breadcrumb": "支付网关错误码定义 > 超时类错误 > ERR_TIMEOUT",
|
||||
"heading_h1": "支付网关错误码定义",
|
||||
"heading_h2": "超时类错误",
|
||||
"heading_h3": "ERR_TIMEOUT"
|
||||
}
|
||||
```
|
||||
|
||||
**收益评估**:
|
||||
|
||||
| 场景 | 无面包屑的问题 | 有面包屑的改善 | 提升幅度 |
|
||||
|------|---------------|---------------|:--------:|
|
||||
| 单文档多分块 | LLM 知道标题但不知道层级关系 | 清楚"文档>章节>条目"归属 | 中等 |
|
||||
| 跨文档混合结果 | 分块看不出源文档 | breadcrumb 第一段就是文档标题 | 大 |
|
||||
| 深层嵌套文档(3+ 级) | 分块内容难以定位 | 完整路径一目了然 | 显著 |
|
||||
| 向量检索相关性 | 只对 chunk content 做 embedding | breadcrumb 可拼入 content 做 embedding 或单独索引 | 中等 |
|
||||
|
||||
**MVP 阶段价值**:当前文档结构较浅(2-3级),breadcrumb 对 LLM 理解帮助中等。但如果后续文档层级加深(像你提到的"排障指南 > 支付网关 > 502错误处理"),价值会显著提升。
|
||||
|
||||
**实现成本**:低。修改 `DocumentChunkService` 的分块逻辑,积累当前标题栈,写入 `DocumentChunk` 和 Milvus metadata。
|
||||
|
||||
#### 方向 B:混合检索 + RRF 重排
|
||||
|
||||
**做法**:L0 关键词和 L1 向量检索并行执行 → 结果用 Reciprocal Rank Fusion 统一排序 → 取 top-K。
|
||||
|
||||
```
|
||||
用户查询 → 并行的:
|
||||
├── L0 关键词匹配 → 得分向量 S₀
|
||||
└── L1 向量检索 → 得分向量 S₁
|
||||
↓
|
||||
RRF 融合重排
|
||||
↓
|
||||
top-K 统一结果
|
||||
```
|
||||
|
||||
RRF 公式:对每个文档 d,`score(d) = Σ 1/(k + rank_r(d))`,其中 k=60(常数)。
|
||||
|
||||
**收益评估**:
|
||||
|
||||
| 场景 | 当前的问题 | 混合 + RRF | 提升幅度 |
|
||||
|------|-----------|-----------|:--------:|
|
||||
| 精确关键词("ERR_TIMEOUT") | L0 匹配但不排序,L1 可能不匹配 | L0 高排名 → RRF 拉到顶部 | 大 |
|
||||
| 语义查询("支付超时如何处理") | L0 可能不匹配,全靠 L1 | L1 兜底不受影响 | 无变化 |
|
||||
| 混合查询("ERR_TIMEOUT 支付网关超时") | L0 匹配一个、L1 匹配一个,无融合 | RRF 统一排序,更合理 | 中等 |
|
||||
| 多文档匹配 | L0 返回无序列表 + L1 独立结果 | 统一排序、去重 | 大 |
|
||||
|
||||
**MVP 阶段价值**:RRF 的实现成本和维护成本较高,而当前 MVP 数据量小(6 个文档),人工检查即可确定哪些匹配是好的。**建议数据量 > 50 个文档时引入**。
|
||||
|
||||
#### 方向 C:Breadcrumb + Embedding 增强
|
||||
|
||||
**做法**:将 breadcrumb 拼入 chunk content 后再做 embedding,让向量包含层级语义。
|
||||
|
||||
```java
|
||||
// 当前
|
||||
embeddingService.generateEmbedding(chunk.getContent())
|
||||
|
||||
// 改进
|
||||
String augmentedContent = chunk.getBreadcrumb() + "\n" + chunk.getContent();
|
||||
embeddingService.generateEmbedding(augmentedContent);
|
||||
```
|
||||
|
||||
这样搜索"ERR_TIMEOUT"时,"支付网关错误码定义 > 超时类错误 > ERR_TIMEOUT" 也会匹配到,而不只是 chunk 正文。
|
||||
|
||||
| 场景 | 当前 | Breadcrumb + Embedding | 提升 |
|
||||
|------|------|------------------------|:----:|
|
||||
| 搜索"支付网关超时" | 匹配到正文含"超时"和"支付网关"的 chunk | breadcrumb 直接含"支付网关",匹配更准 | 中等 |
|
||||
| 搜索"错误码定义" | 可能匹配不到具体错误内容的 chunk | breadcrumb 含"错误码定义",相关性更高 | 大 |
|
||||
|
||||
---
|
||||
|
||||
### 7.4 实施优先级建议
|
||||
|
||||
| 优先级 | 改进项 | 复杂度 | 收益 | 状态 |
|
||||
|:------:|--------|:------:|:----:|:----:|
|
||||
| P0 | **Breadcrumb 上下文**(方向 A) | 低 | 中 | **✅ 已实现 (2026-06-26)** |
|
||||
| P1 | 下个版本 | 低 | 中-大 | 待定 |
|
||||
| P1 | L0 排序(匹配评分 + 排序) | 低 | 中 | 待定 |
|
||||
| P2 | 混合检索 + RRF 重排 | 高 | 大 | 数据量 > 50 文档时引入 |
|
||||
|
||||
### 7.5 Breadcrumb 实现说明
|
||||
|
||||
已于 2026-06-26 实现。改动范围:
|
||||
|
||||
| 文件 | 改动 |
|
||||
|------|------|
|
||||
| `DocumentChunk.java` | 新增 `breadcrumb` 字段 |
|
||||
| `DocumentChunkService.java` | `splitByHeadings()` 维护标题层级栈,`Section` 新增 `level`/`breadcrumb`,`chunkSection()` 和 `saveChunkAndGetNextStart()` 透传 Breadcrumb |
|
||||
| `VectorIndexService.java` | `buildMetadata()` 和 `buildDocumentMetadata()` 将 breadcrumb 写入 Milvus metadata |
|
||||
|
||||
#### 层级栈算法
|
||||
|
||||
```java
|
||||
// 在 splitByHeadings() 中,每次匹配到标题时:
|
||||
while (!headingStack.isEmpty() && headingStack.size() >= level) {
|
||||
headingStack.remove(headingStack.size() - 1); // 弹出同级或更高级
|
||||
}
|
||||
headingStack.add(title); // 追加当前标题
|
||||
currentBreadcrumb = String.join(" > ", headingStack);
|
||||
```
|
||||
|
||||
示例:处理 `fault-diagnosis-process.md` 的完整面包屑路径──
|
||||
|
||||
```json
|
||||
// 分块 "应急响应流程 > 1. 初步评估"
|
||||
{ "breadcrumb": "故障诊断流程规范 > 应急响应流程 > 1. 初步评估" }
|
||||
|
||||
// 分块 "根因分析方法 > 5-Why 分析法"
|
||||
{ "breadcrumb": "故障诊断流程规范 > 根因分析方法 > 5-Why 分析法" }
|
||||
```
|
||||
|
||||
#### 当前 metadata 结构(Milvus)
|
||||
|
||||
```json
|
||||
{
|
||||
"_source": "knowledge_base/api/payment-errors.md",
|
||||
"_file_name": "payment-errors.md",
|
||||
"category": "api",
|
||||
"version": "1.0",
|
||||
"author": "zhangsan"
|
||||
"chunkIndex": 2,
|
||||
"totalChunks": 5,
|
||||
"title": "超时类错误",
|
||||
"breadcrumb": "支付网关错误码定义 > 超时类错误 > ERR_TIMEOUT"
|
||||
}
|
||||
```
|
||||
|
||||
### 5.2 filePath 字段用途变更
|
||||
以 `fault-diagnosis-process.md` 为例:
|
||||
|
||||
**原用途**:存储相对路径或 URL
|
||||
```markdown
|
||||
# 故障诊断流程规范 ← heading_h1
|
||||
|
||||
**新用途**:存储本地文件绝对路径
|
||||
```
|
||||
knowledge_base/api/payment-errors.md
|
||||
knowledge_base/infrastructure/redis-config.md
|
||||
## 应急响应流程 ← heading_h2
|
||||
|
||||
### 1. 初步评估 ← heading_h3(分块1)
|
||||
内容...
|
||||
|
||||
### 2. 快速止血 ← heading_h3(分块2)
|
||||
内容...
|
||||
|
||||
## 根因分析方法 ← heading_h2
|
||||
|
||||
### 5-Why 分析法 ← heading_h3(分块3)
|
||||
内容...
|
||||
```
|
||||
|
||||
**用途**:
|
||||
1. L0 索引读取完整文档
|
||||
2. 支持未来的章节锚点功能
|
||||
|
||||
---
|
||||
|
||||
## 六、配置说明
|
||||
|
||||
### 6.1 application.yml 新增配置
|
||||
|
||||
```yaml
|
||||
knowledge:
|
||||
base-path: knowledge_base/
|
||||
```
|
||||
|
||||
**说明**:
|
||||
- 相对于项目根目录
|
||||
- 启动时递归扫描此目录
|
||||
- 建议按 category 组织子目录
|
||||
|
||||
### 6.2 目录结构规范
|
||||
改造后每个分块的 metadata:
|
||||
|
||||
```
|
||||
knowledge_base/
|
||||
├── api/ # API 相关文档
|
||||
│ └── payment-errors.md
|
||||
├── infrastructure/ # 基础设施配置
|
||||
│ ├── redis-config.md
|
||||
│ ├── mysql-connection-pool.md
|
||||
│ └── flyway-best-practices.md
|
||||
├── domain/ # 领域知识
|
||||
│ └── spring-ai-tool-best-practices.md
|
||||
└── troubleshooting/ # 故障排查
|
||||
└── fault-diagnosis-process.md
|
||||
分块1: breadcrumb = "故障诊断流程规范 > 应急响应流程 > 1. 初步评估"
|
||||
分块2: breadcrumb = "故障诊断流程规范 > 应急响应流程 > 2. 快速止血"
|
||||
分块3: breadcrumb = "故障诊断流程规范 > 根因分析方法 > 5-Why 分析法"
|
||||
```
|
||||
|
||||
---
|
||||
LLM 视角受益:当检索到 "2. 快速止血" 时,LLM 立刻知道它属于"故障诊断流程规范 > 应急响应流程"体系,不需要额外读取其他分块来推断上下文。
|
||||
|
||||
## 七、性能指标
|
||||
|
||||
### 7.1 查询性能
|
||||
|
||||
| 场景 | L0 耗时 | L1 耗时 | 总耗时 |
|
||||
|------|---------|---------|--------|
|
||||
| 唯一匹配(高置信) | < 5ms | 0 (不调用) | < 10ms |
|
||||
| 多个匹配(低置信) | < 5ms | 200-500ms | < 500ms |
|
||||
| 未匹配(仅L1) | < 5ms | 200-500ms | < 500ms |
|
||||
|
||||
### 7.2 索引性能
|
||||
|
||||
| 指标 | 实测值 | 目标值 |
|
||||
|------|--------|--------|
|
||||
| 启动扫描时间 | < 20ms (6 个文档) | < 1s (500 个文档) |
|
||||
| 内存占用 | < 1MB (6 个文档) | < 5MB (500 个文档) |
|
||||
| L0 匹配时间 | < 5ms | < 10ms |
|
||||
|
||||
---
|
||||
|
||||
## 八、可观测性
|
||||
|
||||
### 8.1 日志追踪
|
||||
|
||||
所有查询都带 requestId(8 位 UUID),可追踪完整流程:
|
||||
|
||||
```
|
||||
[a1b2c3d4] 收到知识库查询请求: query=ERR_TIMEOUT
|
||||
[a1b2c3d4] L0精确匹配完成: matches=1, time=2ms
|
||||
[a1b2c3d4] 置信度判断: highConfidence=true, reason=唯一匹配
|
||||
[a1b2c3d4] L0唯一匹配,跳过L1检索
|
||||
[a1b2c3d4] 查询完成: found=true, confidence=high, totalTime=5ms
|
||||
```
|
||||
|
||||
### 8.2 关键指标
|
||||
|
||||
**监控指标**:
|
||||
- L0 查询耗时(P50/P95/P99)
|
||||
- L1 调用频率(低置信度比例)
|
||||
- 查询总耗时(端到端)
|
||||
- 高置信度命中率
|
||||
|
||||
**告警阈值**:
|
||||
- 查询总耗时 > 2s
|
||||
- L0 索引加载失败
|
||||
- 高置信度命中率 < 20%
|
||||
|
||||
---
|
||||
|
||||
## 九、限制与注意事项
|
||||
|
||||
### 9.1 MVP 阶段限制
|
||||
|
||||
1. **L0 索引无持久化**
|
||||
- 应用重启需要重新扫描
|
||||
- 缓解:启动扫描通常 < 1s
|
||||
|
||||
2. **章节锚点未实现**
|
||||
- sectionTitle 参数预留
|
||||
- availableSections 返回 null
|
||||
|
||||
3. **批量导入不支持**
|
||||
- 当前仅支持单文件上传
|
||||
|
||||
### 9.2 最佳实践
|
||||
|
||||
1. **编写高质量 frontmatter**
|
||||
- keywords 精准且全面
|
||||
- 避免关键词重复(导致多匹配)
|
||||
|
||||
2. **知识库目录组织**
|
||||
- 按 category 分类
|
||||
- 文件命名语义化
|
||||
|
||||
3. **监控告警配置**
|
||||
- 慢查询告警
|
||||
- L0 索引加载失败告警
|
||||
|
||||
---
|
||||
|
||||
## 十、后续增强方向(Phase 2)
|
||||
|
||||
1. **章节锚点**
|
||||
- 支持 sectionTitle 参数
|
||||
- 直接定位到文档特定章节
|
||||
|
||||
2. **L0 索引持久化**
|
||||
- 序列化到文件
|
||||
- 避免重启扫描
|
||||
|
||||
3. **批量导入工具**
|
||||
- 支持目录批量导入
|
||||
- 进度监控
|
||||
|
||||
4. **知识库管理 API**
|
||||
- CRUD 接口
|
||||
- 在线编辑
|
||||
|
||||
5. **向量化元数据**
|
||||
- title/summary 也参与 L1 检索
|
||||
- 提升语义检索准确度
|
||||
| 文件 | 说明 |
|
||||
|------|------|
|
||||
| `LookupKnowledgeTool.java` | 检索工具入口 |
|
||||
| `KnowledgeIndexService.java` | L0 内存索引管理 |
|
||||
| `VectorSearchService.java` | L1 向量检索(Milvus) |
|
||||
| `KnowledgeEntry.java` | 索引条目 DTO(含 title, summary, keywords) |
|
||||
| `LookupResult.java` | 查询结果 DTO |
|
||||
| `PrimaryResult.java` | L0 结果 DTO |
|
||||
| `SupplementResult.java` | L1 结果 DTO |
|
||||
|
||||
@@ -0,0 +1,188 @@
|
||||
# 会话级去重与知识域地图
|
||||
|
||||
文档级去重 + 知识域地图注入 Planner,解决 ISS-001 Executor 重复召回同一文档问题。
|
||||
|
||||
---
|
||||
|
||||
## 一、整体架构
|
||||
|
||||
本 change 包含两个独立但互补的部分:
|
||||
|
||||
```
|
||||
Part A: 工具层去重
|
||||
LookupKnowledgeTool
|
||||
├── 维护 ConcurrentHashMap<sessionId, Set<filePath>>(JVM 内)
|
||||
├── 每次检索前过滤已召回文档
|
||||
└── SessionContextHolder.clear() 时同步清理
|
||||
|
||||
Part B: 知识域地图
|
||||
┌─────────────────────────────────────────────┐
|
||||
│ 文档上传 (DocumentManagementService) │
|
||||
│ → LLM 生成 doc.covers + doc.when_to_retrieve │
|
||||
│ → 存入 api_document.metadata │
|
||||
│ → 触发域级重算 (KnowledgeDomainService) │
|
||||
└─────────────────────────────────────────────┘
|
||||
↓
|
||||
┌─────────────────────────────────────────────┐
|
||||
│ 域级聚合 (KnowledgeDomainService) │
|
||||
│ → 读取同域所有文档的 when_to_retrieve │
|
||||
│ → LLM 生成 domain.when_to_retrieve │
|
||||
│ → 存入 knowledge_domain 表 │
|
||||
└─────────────────────────────────────────────┘
|
||||
↓
|
||||
┌─────────────────────────────────────────────┐
|
||||
│ 启动 (KnowledgeIndexService.loadIndex) │
|
||||
│ → 加载 knowledge_domain 表 │
|
||||
│ → 某域无记录则触发域级生成 │
|
||||
└─────────────────────────────────────────────┘
|
||||
↓
|
||||
┌─────────────────────────────────────────────┐
|
||||
│ Planner prompt (ChatService) │
|
||||
│ → 注入 knowledge map(域级) │
|
||||
│ → Planner 做粗粒度检索决策 │
|
||||
└─────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 二、Part A:工具层去重
|
||||
|
||||
### RetrievedDocTracker
|
||||
|
||||
session 级已召回文档追踪组件,将去重责任从 LLM 移交到工具层。
|
||||
|
||||
```java
|
||||
ConcurrentHashMap<String, Set<String>> retrieved
|
||||
key: sessionId
|
||||
value: Set<filePath>
|
||||
```
|
||||
|
||||
| 方法 | 作用 |
|
||||
|------|------|
|
||||
| `isAlreadyRetrieved(sessionId, filePath)` | 检查文档是否已召回 |
|
||||
| `markRetrieved(sessionId, filePath)` | 记录已召回文档 |
|
||||
| `clearSession(sessionId)` | 清理会话记录(SessionContextHolder.clear 触发) |
|
||||
|
||||
### 去重流程
|
||||
|
||||
```
|
||||
lookup_knowledge(query)
|
||||
→ L0 检索 → 命中一批文档
|
||||
→ 遍历结果,过滤 isAlreadyRetrieved=true 的文档
|
||||
→ 剩余文档作为 primary/supplement 返回
|
||||
→ 实际返回的文档调用 markRetrieved
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 三、Part B:知识域地图
|
||||
|
||||
### Frontmatter 新增字段
|
||||
|
||||
文档上传时 LLM 自动生成以下两个字段:
|
||||
|
||||
```yaml
|
||||
covers: ["支付失败排查", "扣款无回调"] # 业务场景标签
|
||||
when_to_retrieve: "用户描述支付失败、超时时" # 文档级检索时机
|
||||
```
|
||||
|
||||
### knowledge_domain 表
|
||||
|
||||
```sql
|
||||
CREATE TABLE knowledge_domain (
|
||||
id BIGINT AUTO_INCREMENT PRIMARY KEY,
|
||||
domain_id VARCHAR(64) NOT NULL UNIQUE,
|
||||
description VARCHAR(256),
|
||||
when_to_retrieve TEXT,
|
||||
document_count INT DEFAULT 0,
|
||||
updated_at DATETIME,
|
||||
created_at DATETIME
|
||||
);
|
||||
```
|
||||
|
||||
### Knowledge Map(注入 Planner 的 YAML)
|
||||
|
||||
```yaml
|
||||
available_knowledge_domains:
|
||||
- domain_id: "payment"
|
||||
description: "支付链路问题排查"
|
||||
when_to_retrieve: "用户问题涉及支付、退款、对账时检索;优先检索一次,勿重复"
|
||||
documents:
|
||||
- title: "支付失败排查手册"
|
||||
covers: ["支付超时", "扣款无回调"]
|
||||
- title: "退款处理指南"
|
||||
covers: ["退款未到账", "退款状态异常"]
|
||||
- domain_id: "infrastructure"
|
||||
...
|
||||
```
|
||||
|
||||
### 注入链路
|
||||
|
||||
```
|
||||
文档上传/删除
|
||||
→ KnowledgeDomainService.onDocumentChange(category)
|
||||
→ 读取同域所有文档的 when_to_retrieve
|
||||
→ LLM 聚合为 domain.when_to_retrieve
|
||||
→ 写入 knowledge_domain 表
|
||||
|
||||
应用启动
|
||||
→ KnowledgeIndexService.loadIndex()
|
||||
→ 加载 knowledge_domain → 无记录则触发聚合
|
||||
→ ChatService.buildChatPlannerAgent() 注入 prompt
|
||||
|
||||
Planner prompt 中包含知识域地图
|
||||
→ Planner 做粗粒度检索决策("查 payment 域")
|
||||
→ Executor 收到步骤后执行具体检索
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 四、关键设计决策
|
||||
|
||||
| 决策 | 方案 | 原因 |
|
||||
|------|------|------|
|
||||
| 域级 when_to_retrieve 存 DB | 持久化 | 避免每次重启调 LLM,文档变更时只重算受影响域 |
|
||||
| 文档级 when_to_retrieve 存 metadata JSON | 沿用现有路径 | 无需新增数据库字段 |
|
||||
| RetrievedDocTracker 独立于 SessionContextHolder | 职责分离 | SessionContextHolder 只持有 sessionId,Tracker 是业务状态 |
|
||||
| Planner 只看域级 | 分层决策 | 文档级 when_to_retrieve 留 Executor 筛选(Phase 2) |
|
||||
| LLM 调用同步执行 | 上传时即时生成 | 接受约 1-2s 延迟,保证数据库和 L0 索引立即一致 |
|
||||
|
||||
---
|
||||
|
||||
## 五、Agent 边界
|
||||
|
||||
```
|
||||
Planner 角色:知道"有什么域"
|
||||
└─ 知识域地图:选定要检索的域(一次规划)
|
||||
|
||||
Executor 角色:知道"做了什么"
|
||||
└─ 行动记忆:域级 + 文档级去重(ISS-002 升级为双层记忆)
|
||||
```
|
||||
|
||||
Part B(知识域地图)只注入 Planner prompt,**不注入 Executor prompt**。Executor 只通过 RetrievedDocTracker 知道自己已检索了哪些文档,不需要知道全局域有哪些。
|
||||
|
||||
---
|
||||
|
||||
## 六、数据库变更
|
||||
|
||||
### V009
|
||||
|
||||
```sql
|
||||
CREATE TABLE knowledge_domain (
|
||||
id BIGINT AUTO_INCREMENT PRIMARY KEY,
|
||||
domain_id VARCHAR(64) NOT NULL UNIQUE,
|
||||
description VARCHAR(256),
|
||||
when_to_retrieve TEXT,
|
||||
document_count INT DEFAULT 0,
|
||||
updated_at DATETIME,
|
||||
created_at DATETIME
|
||||
);
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 七、参考资料
|
||||
|
||||
- **架构文档**:`mvp/architecture/knowledge-retrieval-architecture.md`
|
||||
- **使用指南**:`mvp/architecture/knowledge-retrieval-usage.md`
|
||||
- **OpenSpec**:`openspec/changes/archive/2026-06-30-session-dedup-knowledge-map/`
|
||||
@@ -0,0 +1,165 @@
|
||||
# MVP Demo Runbook
|
||||
|
||||
This demo proves the MVP flow from user question to persisted diagnosis trace.
|
||||
|
||||
For interview use, start with:
|
||||
|
||||
- `interview-walkthrough.md` for the talk track
|
||||
- `trace-inspection-checklist.md` for fields to inspect
|
||||
- `scripts/run-payment-timeout-demo.ps1` for the runnable local demo
|
||||
- `requests/payment-timeout-chat.json` for the fixed request payload
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- MySQL, Redis, Milvus/Zilliz, and LLM/embedding configuration are available through the current project configuration.
|
||||
- Security and secret cleanup are intentionally out of scope for this MVP slice.
|
||||
- The `mvp-demo` profile enables mock Prometheus and CLS providers so log and metric tools can return repeatable evidence.
|
||||
|
||||
## Start
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||
```
|
||||
|
||||
The service listens on:
|
||||
|
||||
```text
|
||||
http://localhost:9900
|
||||
```
|
||||
|
||||
## 1. Run Chat Diagnosis
|
||||
|
||||
Fast path:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
||||
```
|
||||
|
||||
This writes:
|
||||
|
||||
```text
|
||||
mvp/demo/output/chat-response.json
|
||||
mvp/demo/output/trace-response.json
|
||||
mvp/demo/output/feedback-response.json
|
||||
```
|
||||
|
||||
Manual path:
|
||||
|
||||
```powershell
|
||||
$sessionId = "mvp-demo-payment-timeout-001"
|
||||
$body = @{
|
||||
Id = $sessionId
|
||||
Question = "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/chat" `
|
||||
-ContentType "application/json" `
|
||||
-Body $body
|
||||
```
|
||||
|
||||
Expected result:
|
||||
|
||||
- `data.success` is `true`.
|
||||
- `data.sessionId` equals `mvp-demo-payment-timeout-001`.
|
||||
- `data.answer` contains a diagnosis answer.
|
||||
|
||||
## 2. Query Trace
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace"
|
||||
```
|
||||
|
||||
Expected result:
|
||||
|
||||
- `code` is `200`.
|
||||
- `data.session.sessionId` equals the chat session id.
|
||||
- `data.steps` contains planner/executor/verifier records for complex questions.
|
||||
- `data.toolInvocations` contains evidence tool calls such as `lookup_knowledge`, `query_logs`, or `query_metrics`.
|
||||
- `data.session.selfEvaluation` contains verifier or rule evaluation when available.
|
||||
|
||||
## 3. Submit Feedback
|
||||
|
||||
```powershell
|
||||
$feedback = @{
|
||||
sessionId = $sessionId
|
||||
feedback = "useful"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/feedback" `
|
||||
-ContentType "application/json" `
|
||||
-Body $feedback
|
||||
```
|
||||
|
||||
Expected result:
|
||||
|
||||
- `success` is `true`.
|
||||
- A later trace query shows `data.session.feedback` as `useful`.
|
||||
|
||||
## 4. Run AIOps Alert Diagnosis
|
||||
|
||||
```powershell
|
||||
$aiopsSessionId = "mvp-demo-aiops-payment-cpu-001"
|
||||
$aiopsBody = @{
|
||||
sessionId = $aiopsSessionId
|
||||
alertName = "HighCPUUsage"
|
||||
service = "payment-service"
|
||||
severity = "P1"
|
||||
description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
|
||||
timeRange = "last_15m"
|
||||
userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-WebRequest `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/ai_ops" `
|
||||
-ContentType "application/json" `
|
||||
-Body $aiopsBody
|
||||
```
|
||||
|
||||
Expected result:
|
||||
|
||||
- The SSE stream starts with a `session` message containing `mvp-demo-aiops-payment-cpu-001`.
|
||||
- The stream later contains an AIOps alert analysis report focused on the supplied `HighCPUUsage/payment-service` payload.
|
||||
- A trace query for the same session id returns `data.session.agentFlow` as `AI_OPS`.
|
||||
- `data.session.answer` contains the final alert analysis report when a report is generated.
|
||||
- `data.toolInvocations` contains evidence tools such as `lookup_knowledge`, `query_logs`, or `query_metrics` when the runtime uses them.
|
||||
|
||||
Query the AIOps trace:
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace"
|
||||
```
|
||||
|
||||
## Demo Story
|
||||
|
||||
The important interview story is:
|
||||
|
||||
```text
|
||||
one session id
|
||||
-> user question
|
||||
-> multi-agent execution
|
||||
-> evidence tools
|
||||
-> verifier/self-evaluation
|
||||
-> final answer
|
||||
-> feedback
|
||||
-> trace API for replay and audit
|
||||
```
|
||||
|
||||
The AIOps story uses the same audit spine:
|
||||
|
||||
```text
|
||||
one session id
|
||||
-> alert payload
|
||||
-> AIOps planner/executor execution
|
||||
-> evidence tools
|
||||
-> alert analysis report
|
||||
-> trace API for replay and audit
|
||||
```
|
||||
@@ -0,0 +1,38 @@
|
||||
# AIOps Alert Acceptance Case
|
||||
|
||||
## Goal
|
||||
|
||||
Validate that the legacy AIOps endpoint can act as a traceable alert-triggered diagnosis entry.
|
||||
|
||||
## Input
|
||||
|
||||
- Session id: `mvp-demo-aiops-payment-cpu-001`
|
||||
- Endpoint: `POST /api/ai_ops`
|
||||
- Profile: `mvp-demo`
|
||||
- Alert:
|
||||
|
||||
```json
|
||||
{
|
||||
"sessionId": "mvp-demo-aiops-payment-cpu-001",
|
||||
"alertName": "HighCPUUsage",
|
||||
"service": "payment-service",
|
||||
"severity": "P1",
|
||||
"description": "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。",
|
||||
"timeRange": "last_15m",
|
||||
"userRequest": "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
|
||||
}
|
||||
```
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
1. The SSE stream emits a `session` message containing the requested session id.
|
||||
2. The AIOps run creates or updates `diagnosis_session` with `agent_flow = AI_OPS`.
|
||||
3. The persisted session query contains the alert name, service, severity, time range, and description.
|
||||
4. If a final report is generated, `diagnosis_session.answer` contains that report.
|
||||
5. `GET /api/diagnosis/{sessionId}/trace` returns the AIOps session, ordered agent steps, and ordered tool invocations.
|
||||
6. In payload mode, the report focuses on `HighCPUUsage/payment-service`; unrelated active alerts may appear only as related risk or context, not as separate full root-cause sections.
|
||||
|
||||
## Known Limits
|
||||
|
||||
- This slice does not add a Verifier Agent to AIOps.
|
||||
- Full runtime verification still depends on valid DB, Redis, Milvus/Zilliz, model, and embedding configuration.
|
||||
@@ -0,0 +1,146 @@
|
||||
# Interview Walkthrough: MVP Diagnosis Agent
|
||||
|
||||
This walkthrough is the Plan C demo story. It is meant for a short Agent Engineer interview, not as exhaustive system documentation.
|
||||
|
||||
## 30-Second Summary
|
||||
|
||||
```text
|
||||
This is an enterprise diagnosis Agent MVP.
|
||||
It takes a payment-timeout question, plans the investigation, calls evidence tools,
|
||||
checks the answer through a verifier, persists the full trace, and accepts feedback.
|
||||
```
|
||||
|
||||
The important claim is not "the model answered once." The claim is:
|
||||
|
||||
```text
|
||||
The system can show what evidence was used, how the answer was checked, and how to replay the session.
|
||||
```
|
||||
|
||||
## Demo Flow
|
||||
|
||||
1. Start the service with the `mvp-demo` profile.
|
||||
2. Run the fixed payment-timeout request.
|
||||
3. Open `mvp/demo/output/chat-response.json`.
|
||||
4. Open `mvp/demo/output/trace-response.json`.
|
||||
5. Point to evidence tools and verifier evaluation.
|
||||
6. Submit feedback and show it is attached to the same session.
|
||||
|
||||
## Commands
|
||||
|
||||
Start service:
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||
```
|
||||
|
||||
Run the demo from another terminal:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
||||
```
|
||||
|
||||
Optional custom session:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1 -SessionId "mvp-demo-payment-timeout-002"
|
||||
```
|
||||
|
||||
## What To Show
|
||||
|
||||
### 1. User-Facing Answer
|
||||
|
||||
File:
|
||||
|
||||
```text
|
||||
mvp/demo/output/chat-response.json
|
||||
```
|
||||
|
||||
Say:
|
||||
|
||||
```text
|
||||
This is the answer the user sees. The session id is stable, so I can trace this exact answer later.
|
||||
```
|
||||
|
||||
### 2. Evidence Trace
|
||||
|
||||
File:
|
||||
|
||||
```text
|
||||
mvp/demo/output/trace-response.json
|
||||
```
|
||||
|
||||
Say:
|
||||
|
||||
```text
|
||||
This is the important Agent engineering part.
|
||||
I can inspect which tools were called, what inputs they received,
|
||||
whether they succeeded, and what evidence preview was persisted.
|
||||
```
|
||||
|
||||
Point to:
|
||||
|
||||
- `data.toolInvocations[*].toolName`
|
||||
- `data.toolInvocations[*].inputParams`
|
||||
- `data.toolInvocations[*].outputPreview`
|
||||
- `data.toolInvocations[*].success`
|
||||
|
||||
### 3. Verifier / Self-Evaluation
|
||||
|
||||
Point to:
|
||||
|
||||
- `data.session.selfEvaluation`
|
||||
- `data.summary.hasVerifierEvaluation`
|
||||
|
||||
Say:
|
||||
|
||||
```text
|
||||
The final answer is not just raw Executor output.
|
||||
It is checked by a verifier or self-evaluation layer using the persisted trace.
|
||||
That lets the system return PASS, LOW_CONFID, or REJECT-style behavior instead of pretending all answers are equally certain.
|
||||
```
|
||||
|
||||
### 4. Feedback Loop
|
||||
|
||||
File:
|
||||
|
||||
```text
|
||||
mvp/demo/output/feedback-response.json
|
||||
```
|
||||
|
||||
Then re-query trace if needed.
|
||||
|
||||
Say:
|
||||
|
||||
```text
|
||||
Feedback is attached to the same diagnosis session.
|
||||
That makes it possible to mine useful / not useful cases later.
|
||||
```
|
||||
|
||||
### 5. Regression Story
|
||||
|
||||
Mention, do not deep dive unless asked:
|
||||
|
||||
```text
|
||||
For repeatability, I also built an offline eval baseline.
|
||||
The demo proves the runtime trace; the eval baseline proves fixed-case regression.
|
||||
The two are separate on purpose: demo for human review, eval for automated signal.
|
||||
```
|
||||
|
||||
## Strong Interview Framing
|
||||
|
||||
Use this phrasing:
|
||||
|
||||
```text
|
||||
I focused on the Agent engineering surface:
|
||||
traceability, evidence persistence, verifier gating, feedback, and regression checks.
|
||||
The model answer is only one part of the system.
|
||||
The more important part is whether we can audit and improve the answer after it is produced.
|
||||
```
|
||||
|
||||
## Known Limits To Say Proactively
|
||||
|
||||
```text
|
||||
This MVP still depends on configured MySQL, Redis, Milvus, and model credentials.
|
||||
The mvp-demo profile mocks logs and metrics, but not the full application runtime.
|
||||
Secret cleanup and fully isolated default tests are separate production-hardening tasks.
|
||||
```
|
||||
@@ -0,0 +1,11 @@
|
||||
# Demo Output
|
||||
|
||||
This directory is the default output location for local demo responses.
|
||||
|
||||
Generated files are intentionally ignored by Git:
|
||||
|
||||
- `chat-response.json`
|
||||
- `trace-response.json`
|
||||
- `feedback-response.json`
|
||||
|
||||
Keep this README so the directory exists in the repository.
|
||||
@@ -0,0 +1,39 @@
|
||||
# Payment Timeout Acceptance Case
|
||||
|
||||
## Goal
|
||||
|
||||
Validate that the MVP can diagnose a payment timeout incident and expose the complete trace for replay.
|
||||
|
||||
## Input
|
||||
|
||||
- Session id: `mvp-demo-payment-timeout-001`
|
||||
- Question: `支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。`
|
||||
- Profile: `mvp-demo`
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
1. Chat returns a successful answer with the same session id.
|
||||
2. Trace API returns session metadata, final answer, ordered agent steps, and ordered tool invocations.
|
||||
3. Trace contains enough evidence to explain which tools were used and whether verifier/self-evaluation was persisted.
|
||||
4. Feedback can be submitted for the same session id.
|
||||
5. A follow-up trace query shows the persisted feedback value.
|
||||
|
||||
## Trace Fields To Inspect
|
||||
|
||||
- `data.session.query`
|
||||
- `data.session.answer`
|
||||
- `data.session.selfEvaluation`
|
||||
- `data.session.feedback`
|
||||
- `data.steps[*].agentName`
|
||||
- `data.steps[*].thought`
|
||||
- `data.toolInvocations[*].toolName`
|
||||
- `data.toolInvocations[*].inputParams`
|
||||
- `data.toolInvocations[*].outputPreview`
|
||||
- `data.toolInvocations[*].retrievalDetails`
|
||||
- `data.summary`
|
||||
|
||||
## Known Limits
|
||||
|
||||
- This case is not a full offline test. It still requires valid infrastructure for chat, persistence, vector search, and model calls.
|
||||
- Mock logs and metrics are enabled by the `mvp-demo` profile to make those evidence tools repeatable.
|
||||
- Sensitive configuration cleanup is deferred by current MVP priority.
|
||||
@@ -0,0 +1,4 @@
|
||||
{
|
||||
"Id": "mvp-demo-payment-timeout-001",
|
||||
"Question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
|
||||
}
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user