Compare commits
16
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
246c99b954 | ||
|
|
f01866c1a2 | ||
|
|
6919092b83 | ||
|
|
5b827fe90e | ||
|
|
b0f288ae36 | ||
|
|
1ff7f09d25 | ||
|
|
fd89d84fc0 | ||
|
|
9050487307 | ||
|
|
4f5316d473 | ||
|
|
2a7164288f | ||
|
|
a1c896ebda | ||
|
|
e438df4355 | ||
|
|
e4f37cb9e6 | ||
|
|
354ffc1947 | ||
|
|
bb44140901 | ||
|
|
2a796da490 |
@@ -0,0 +1,156 @@
|
||||
---
|
||||
name: openspec-apply-change
|
||||
description: Implement tasks from an OpenSpec change. Use when the user wants to start implementing, continue implementation, or work through tasks.
|
||||
license: MIT
|
||||
compatibility: Requires openspec CLI.
|
||||
metadata:
|
||||
author: openspec
|
||||
version: "1.0"
|
||||
generatedBy: "1.3.1"
|
||||
---
|
||||
|
||||
Implement tasks from an OpenSpec change.
|
||||
|
||||
**Input**: Optionally specify a change name. If omitted, check if it can be inferred from conversation context. If vague or ambiguous you MUST prompt for available changes.
|
||||
|
||||
**Steps**
|
||||
|
||||
1. **Select the change**
|
||||
|
||||
If a name is provided, use it. Otherwise:
|
||||
- Infer from conversation context if the user mentioned a change
|
||||
- Auto-select if only one active change exists
|
||||
- If ambiguous, run `openspec list --json` to get available changes and use the **AskUserQuestion tool** to let the user select
|
||||
|
||||
Always announce: "Using change: <name>" and how to override (e.g., `/opsx:apply <other>`).
|
||||
|
||||
2. **Check status to understand the schema**
|
||||
```bash
|
||||
openspec status --change "<name>" --json
|
||||
```
|
||||
Parse the JSON to understand:
|
||||
- `schemaName`: The workflow being used (e.g., "spec-driven")
|
||||
- Which artifact contains the tasks (typically "tasks" for spec-driven, check status for others)
|
||||
|
||||
3. **Get apply instructions**
|
||||
|
||||
```bash
|
||||
openspec instructions apply --change "<name>" --json
|
||||
```
|
||||
|
||||
This returns:
|
||||
- `contextFiles`: artifact ID -> array of concrete file paths (varies by schema - could be proposal/specs/design/tasks or spec/tests/implementation/docs)
|
||||
- Progress (total, complete, remaining)
|
||||
- Task list with status
|
||||
- Dynamic instruction based on current state
|
||||
|
||||
**Handle states:**
|
||||
- If `state: "blocked"` (missing artifacts): show message, suggest using openspec-continue-change
|
||||
- If `state: "all_done"`: congratulate, suggest archive
|
||||
- Otherwise: proceed to implementation
|
||||
|
||||
4. **Read context files**
|
||||
|
||||
Read every file path listed under `contextFiles` from the apply instructions output.
|
||||
The files depend on the schema being used:
|
||||
- **spec-driven**: proposal, specs, design, tasks
|
||||
- Other schemas: follow the contextFiles from CLI output
|
||||
|
||||
5. **Show current progress**
|
||||
|
||||
Display:
|
||||
- Schema being used
|
||||
- Progress: "N/M tasks complete"
|
||||
- Remaining tasks overview
|
||||
- Dynamic instruction from CLI
|
||||
|
||||
6. **Implement tasks (loop until done or blocked)**
|
||||
|
||||
For each pending task:
|
||||
- Show which task is being worked on
|
||||
- Make the code changes required
|
||||
- Keep changes minimal and focused
|
||||
- Mark task complete in the tasks file: `- [ ]` → `- [x]`
|
||||
- Continue to next task
|
||||
|
||||
**Pause if:**
|
||||
- Task is unclear → ask for clarification
|
||||
- Implementation reveals a design issue → suggest updating artifacts
|
||||
- Error or blocker encountered → report and wait for guidance
|
||||
- User interrupts
|
||||
|
||||
7. **On completion or pause, show status**
|
||||
|
||||
Display:
|
||||
- Tasks completed this session
|
||||
- Overall progress: "N/M tasks complete"
|
||||
- If all done: suggest archive
|
||||
- If paused: explain why and wait for guidance
|
||||
|
||||
**Output During Implementation**
|
||||
|
||||
```
|
||||
## Implementing: <change-name> (schema: <schema-name>)
|
||||
|
||||
Working on task 3/7: <task description>
|
||||
[...implementation happening...]
|
||||
✓ Task complete
|
||||
|
||||
Working on task 4/7: <task description>
|
||||
[...implementation happening...]
|
||||
✓ Task complete
|
||||
```
|
||||
|
||||
**Output On Completion**
|
||||
|
||||
```
|
||||
## Implementation Complete
|
||||
|
||||
**Change:** <change-name>
|
||||
**Schema:** <schema-name>
|
||||
**Progress:** 7/7 tasks complete ✓
|
||||
|
||||
### Completed This Session
|
||||
- [x] Task 1
|
||||
- [x] Task 2
|
||||
...
|
||||
|
||||
All tasks complete! Ready to archive this change.
|
||||
```
|
||||
|
||||
**Output On Pause (Issue Encountered)**
|
||||
|
||||
```
|
||||
## Implementation Paused
|
||||
|
||||
**Change:** <change-name>
|
||||
**Schema:** <schema-name>
|
||||
**Progress:** 4/7 tasks complete
|
||||
|
||||
### Issue Encountered
|
||||
<description of the issue>
|
||||
|
||||
**Options:**
|
||||
1. <option 1>
|
||||
2. <option 2>
|
||||
3. Other approach
|
||||
|
||||
What would you like to do?
|
||||
```
|
||||
|
||||
**Guardrails**
|
||||
- Keep going through tasks until done or blocked
|
||||
- Always read context files before starting (from the apply instructions output)
|
||||
- If task is ambiguous, pause and ask before implementing
|
||||
- If implementation reveals issues, pause and suggest artifact updates
|
||||
- Keep code changes minimal and scoped to each task
|
||||
- Update task checkbox immediately after completing each task
|
||||
- Pause on errors, blockers, or unclear requirements - don't guess
|
||||
- Use contextFiles from CLI output, don't assume specific file names
|
||||
|
||||
**Fluid Workflow Integration**
|
||||
|
||||
This skill supports the "actions on a change" model:
|
||||
|
||||
- **Can be invoked anytime**: Before all artifacts are done (if tasks exist), after partial implementation, interleaved with other actions
|
||||
- **Allows artifact updates**: If implementation reveals design issues, suggest updating artifacts - not phase-locked, work fluidly
|
||||
@@ -0,0 +1,114 @@
|
||||
---
|
||||
name: openspec-archive-change
|
||||
description: Archive a completed change in the experimental workflow. Use when the user wants to finalize and archive a change after implementation is complete.
|
||||
license: MIT
|
||||
compatibility: Requires openspec CLI.
|
||||
metadata:
|
||||
author: openspec
|
||||
version: "1.0"
|
||||
generatedBy: "1.3.1"
|
||||
---
|
||||
|
||||
Archive a completed change in the experimental workflow.
|
||||
|
||||
**Input**: Optionally specify a change name. If omitted, check if it can be inferred from conversation context. If vague or ambiguous you MUST prompt for available changes.
|
||||
|
||||
**Steps**
|
||||
|
||||
1. **If no change name provided, prompt for selection**
|
||||
|
||||
Run `openspec list --json` to get available changes. Use the **AskUserQuestion tool** to let the user select.
|
||||
|
||||
Show only active changes (not already archived).
|
||||
Include the schema used for each change if available.
|
||||
|
||||
**IMPORTANT**: Do NOT guess or auto-select a change. Always let the user choose.
|
||||
|
||||
2. **Check artifact completion status**
|
||||
|
||||
Run `openspec status --change "<name>" --json` to check artifact completion.
|
||||
|
||||
Parse the JSON to understand:
|
||||
- `schemaName`: The workflow being used
|
||||
- `artifacts`: List of artifacts with their status (`done` or other)
|
||||
|
||||
**If any artifacts are not `done`:**
|
||||
- Display warning listing incomplete artifacts
|
||||
- Use **AskUserQuestion tool** to confirm user wants to proceed
|
||||
- Proceed if user confirms
|
||||
|
||||
3. **Check task completion status**
|
||||
|
||||
Read the tasks file (typically `tasks.md`) to check for incomplete tasks.
|
||||
|
||||
Count tasks marked with `- [ ]` (incomplete) vs `- [x]` (complete).
|
||||
|
||||
**If incomplete tasks found:**
|
||||
- Display warning showing count of incomplete tasks
|
||||
- Use **AskUserQuestion tool** to confirm user wants to proceed
|
||||
- Proceed if user confirms
|
||||
|
||||
**If no tasks file exists:** Proceed without task-related warning.
|
||||
|
||||
4. **Assess delta spec sync state**
|
||||
|
||||
Check for delta specs at `openspec/changes/<name>/specs/`. If none exist, proceed without sync prompt.
|
||||
|
||||
**If delta specs exist:**
|
||||
- Compare each delta spec with its corresponding main spec at `openspec/specs/<capability>/spec.md`
|
||||
- Determine what changes would be applied (adds, modifications, removals, renames)
|
||||
- Show a combined summary before prompting
|
||||
|
||||
**Prompt options:**
|
||||
- If changes needed: "Sync now (recommended)", "Archive without syncing"
|
||||
- If already synced: "Archive now", "Sync anyway", "Cancel"
|
||||
|
||||
If user chooses sync, use Task tool (subagent_type: "general-purpose", prompt: "Use Skill tool to invoke openspec-sync-specs for change '<name>'. Delta spec analysis: <include the analyzed delta spec summary>"). Proceed to archive regardless of choice.
|
||||
|
||||
5. **Perform the archive**
|
||||
|
||||
Create the archive directory if it doesn't exist:
|
||||
```bash
|
||||
mkdir -p openspec/changes/archive
|
||||
```
|
||||
|
||||
Generate target name using current date: `YYYY-MM-DD-<change-name>`
|
||||
|
||||
**Check if target already exists:**
|
||||
- If yes: Fail with error, suggest renaming existing archive or using different date
|
||||
- If no: Move the change directory to archive
|
||||
|
||||
```bash
|
||||
mv openspec/changes/<name> openspec/changes/archive/YYYY-MM-DD-<name>
|
||||
```
|
||||
|
||||
6. **Display summary**
|
||||
|
||||
Show archive completion summary including:
|
||||
- Change name
|
||||
- Schema that was used
|
||||
- Archive location
|
||||
- Whether specs were synced (if applicable)
|
||||
- Note about any warnings (incomplete artifacts/tasks)
|
||||
|
||||
**Output On Success**
|
||||
|
||||
```
|
||||
## Archive Complete
|
||||
|
||||
**Change:** <change-name>
|
||||
**Schema:** <schema-name>
|
||||
**Archived to:** openspec/changes/archive/YYYY-MM-DD-<name>/
|
||||
**Specs:** ✓ Synced to main specs (or "No delta specs" or "Sync skipped")
|
||||
|
||||
All artifacts complete. All tasks complete.
|
||||
```
|
||||
|
||||
**Guardrails**
|
||||
- Always prompt for change selection if not provided
|
||||
- Use artifact graph (openspec status --json) for completion checking
|
||||
- Don't block archive on warnings - just inform and confirm
|
||||
- Preserve .openspec.yaml when moving to archive (it moves with the directory)
|
||||
- Show clear summary of what happened
|
||||
- If sync is requested, use openspec-sync-specs approach (agent-driven)
|
||||
- If delta specs exist, always run the sync assessment and show the combined summary before prompting
|
||||
@@ -0,0 +1,288 @@
|
||||
---
|
||||
name: openspec-explore
|
||||
description: Enter explore mode - a thinking partner for exploring ideas, investigating problems, and clarifying requirements. Use when the user wants to think through something before or during a change.
|
||||
license: MIT
|
||||
compatibility: Requires openspec CLI.
|
||||
metadata:
|
||||
author: openspec
|
||||
version: "1.0"
|
||||
generatedBy: "1.3.1"
|
||||
---
|
||||
|
||||
Enter explore mode. Think deeply. Visualize freely. Follow the conversation wherever it goes.
|
||||
|
||||
**IMPORTANT: Explore mode is for thinking, not implementing.** You may read files, search code, and investigate the codebase, but you must NEVER write code or implement features. If the user asks you to implement something, remind them to exit explore mode first and create a change proposal. You MAY create OpenSpec artifacts (proposals, designs, specs) if the user asks—that's capturing thinking, not implementing.
|
||||
|
||||
**This is a stance, not a workflow.** There are no fixed steps, no required sequence, no mandatory outputs. You're a thinking partner helping the user explore.
|
||||
|
||||
---
|
||||
|
||||
## The Stance
|
||||
|
||||
- **Curious, not prescriptive** - Ask questions that emerge naturally, don't follow a script
|
||||
- **Open threads, not interrogations** - Surface multiple interesting directions and let the user follow what resonates. Don't funnel them through a single path of questions.
|
||||
- **Visual** - Use ASCII diagrams liberally when they'd help clarify thinking
|
||||
- **Adaptive** - Follow interesting threads, pivot when new information emerges
|
||||
- **Patient** - Don't rush to conclusions, let the shape of the problem emerge
|
||||
- **Grounded** - Explore the actual codebase when relevant, don't just theorize
|
||||
|
||||
---
|
||||
|
||||
## What You Might Do
|
||||
|
||||
Depending on what the user brings, you might:
|
||||
|
||||
**Explore the problem space**
|
||||
- Ask clarifying questions that emerge from what they said
|
||||
- Challenge assumptions
|
||||
- Reframe the problem
|
||||
- Find analogies
|
||||
|
||||
**Investigate the codebase**
|
||||
- Map existing architecture relevant to the discussion
|
||||
- Find integration points
|
||||
- Identify patterns already in use
|
||||
- Surface hidden complexity
|
||||
|
||||
**Compare options**
|
||||
- Brainstorm multiple approaches
|
||||
- Build comparison tables
|
||||
- Sketch tradeoffs
|
||||
- Recommend a path (if asked)
|
||||
|
||||
**Visualize**
|
||||
```
|
||||
┌─────────────────────────────────────────┐
|
||||
│ Use ASCII diagrams liberally │
|
||||
├─────────────────────────────────────────┤
|
||||
│ │
|
||||
│ ┌────────┐ ┌────────┐ │
|
||||
│ │ State │────────▶│ State │ │
|
||||
│ │ A │ │ B │ │
|
||||
│ └────────┘ └────────┘ │
|
||||
│ │
|
||||
│ System diagrams, state machines, │
|
||||
│ data flows, architecture sketches, │
|
||||
│ dependency graphs, comparison tables │
|
||||
│ │
|
||||
└─────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
**Surface risks and unknowns**
|
||||
- Identify what could go wrong
|
||||
- Find gaps in understanding
|
||||
- Suggest spikes or investigations
|
||||
|
||||
---
|
||||
|
||||
## OpenSpec Awareness
|
||||
|
||||
You have full context of the OpenSpec system. Use it naturally, don't force it.
|
||||
|
||||
### Check for context
|
||||
|
||||
At the start, quickly check what exists:
|
||||
```bash
|
||||
openspec list --json
|
||||
```
|
||||
|
||||
This tells you:
|
||||
- If there are active changes
|
||||
- Their names, schemas, and status
|
||||
- What the user might be working on
|
||||
|
||||
### When no change exists
|
||||
|
||||
Think freely. When insights crystallize, you might offer:
|
||||
|
||||
- "This feels solid enough to start a change. Want me to create a proposal?"
|
||||
- Or keep exploring - no pressure to formalize
|
||||
|
||||
### When a change exists
|
||||
|
||||
If the user mentions a change or you detect one is relevant:
|
||||
|
||||
1. **Read existing artifacts for context**
|
||||
- `openspec/changes/<name>/proposal.md`
|
||||
- `openspec/changes/<name>/design.md`
|
||||
- `openspec/changes/<name>/tasks.md`
|
||||
- etc.
|
||||
|
||||
2. **Reference them naturally in conversation**
|
||||
- "Your design mentions using Redis, but we just realized SQLite fits better..."
|
||||
- "The proposal scopes this to premium users, but we're now thinking everyone..."
|
||||
|
||||
3. **Offer to capture when decisions are made**
|
||||
|
||||
| Insight Type | Where to Capture |
|
||||
|----------------------------|--------------------------------|
|
||||
| New requirement discovered | `specs/<capability>/spec.md` |
|
||||
| Requirement changed | `specs/<capability>/spec.md` |
|
||||
| Design decision made | `design.md` |
|
||||
| Scope changed | `proposal.md` |
|
||||
| New work identified | `tasks.md` |
|
||||
| Assumption invalidated | Relevant artifact |
|
||||
|
||||
Example offers:
|
||||
- "That's a design decision. Capture it in design.md?"
|
||||
- "This is a new requirement. Add it to specs?"
|
||||
- "This changes scope. Update the proposal?"
|
||||
|
||||
4. **The user decides** - Offer and move on. Don't pressure. Don't auto-capture.
|
||||
|
||||
---
|
||||
|
||||
## What You Don't Have To Do
|
||||
|
||||
- Follow a script
|
||||
- Ask the same questions every time
|
||||
- Produce a specific artifact
|
||||
- Reach a conclusion
|
||||
- Stay on topic if a tangent is valuable
|
||||
- Be brief (this is thinking time)
|
||||
|
||||
---
|
||||
|
||||
## Handling Different Entry Points
|
||||
|
||||
**User brings a vague idea:**
|
||||
```
|
||||
User: I'm thinking about adding real-time collaboration
|
||||
|
||||
You: Real-time collab is a big space. Let me think about this...
|
||||
|
||||
COLLABORATION SPECTRUM
|
||||
════════════════════════════════════════════
|
||||
|
||||
Awareness Coordination Sync
|
||||
│ │ │
|
||||
▼ ▼ ▼
|
||||
┌────────┐ ┌────────┐ ┌────────┐
|
||||
│Presence│ │Cursors │ │ CRDT │
|
||||
│ "3 │ │ Multi │ │Conflict│
|
||||
│online" │ │ select │ │ free │
|
||||
└────────┘ └────────┘ └────────┘
|
||||
│ │ │
|
||||
trivial moderate complex
|
||||
|
||||
Where's your head at?
|
||||
```
|
||||
|
||||
**User brings a specific problem:**
|
||||
```
|
||||
User: The auth system is a mess
|
||||
|
||||
You: [reads codebase]
|
||||
|
||||
┌─────────────────────────────────────────────┐
|
||||
│ CURRENT AUTH FLOW │
|
||||
└─────────────────────────────────────────────┘
|
||||
│
|
||||
┌─────────────┼─────────────┐
|
||||
▼ ▼ ▼
|
||||
┌─────────┐ ┌─────────┐ ┌─────────┐
|
||||
│ Google │ │ GitHub │ │ Email │
|
||||
│ OAuth │ │ OAuth │ │ Magic │
|
||||
└────┬────┘ └────┬────┘ └────┬────┘
|
||||
│ │ │
|
||||
└─────────────┼─────────────┘
|
||||
▼
|
||||
┌───────────┐
|
||||
│ Session │
|
||||
└─────┬─────┘
|
||||
│
|
||||
▼
|
||||
┌───────────┐
|
||||
│ Perms │
|
||||
└───────────┘
|
||||
|
||||
I see three tangles. Which one's burning?
|
||||
```
|
||||
|
||||
**User is stuck mid-implementation:**
|
||||
```
|
||||
User: /opsx:explore add-auth-system
|
||||
The OAuth integration is more complex than expected
|
||||
|
||||
You: [reads change artifacts]
|
||||
|
||||
You're on task 4: "Implement OAuth flow"
|
||||
|
||||
Let me trace what's involved...
|
||||
|
||||
[draws diagram, explores options, suggests paths]
|
||||
|
||||
Want to update the design to reflect this?
|
||||
Or add a spike task to investigate?
|
||||
```
|
||||
|
||||
**User wants to compare options:**
|
||||
```
|
||||
User: Should we use Postgres or SQLite?
|
||||
|
||||
You: Generic answer is boring. What's the context?
|
||||
|
||||
User: A CLI tool that tracks local dev environments
|
||||
|
||||
You: That changes everything.
|
||||
|
||||
┌─────────────────────────────────────────────────┐
|
||||
│ CLI TOOL DATA STORAGE │
|
||||
└─────────────────────────────────────────────────┘
|
||||
|
||||
Key constraints:
|
||||
• No daemon running
|
||||
• Must work offline
|
||||
• Single user
|
||||
|
||||
SQLite Postgres
|
||||
Deployment embedded ✓ needs server ✗
|
||||
Offline yes ✓ no ✗
|
||||
Single file yes ✓ no ✗
|
||||
|
||||
SQLite. Not even close.
|
||||
|
||||
Unless... is there a sync component?
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Ending Discovery
|
||||
|
||||
There's no required ending. Discovery might:
|
||||
|
||||
- **Flow into a proposal**: "Ready to start? I can create a change proposal."
|
||||
- **Result in artifact updates**: "Updated design.md with these decisions"
|
||||
- **Just provide clarity**: User has what they need, moves on
|
||||
- **Continue later**: "We can pick this up anytime"
|
||||
|
||||
When it feels like things are crystallizing, you might summarize:
|
||||
|
||||
```
|
||||
## What We Figured Out
|
||||
|
||||
**The problem**: [crystallized understanding]
|
||||
|
||||
**The approach**: [if one emerged]
|
||||
|
||||
**Open questions**: [if any remain]
|
||||
|
||||
**Next steps** (if ready):
|
||||
- Create a change proposal
|
||||
- Keep exploring: just keep talking
|
||||
```
|
||||
|
||||
But this summary is optional. Sometimes the thinking IS the value.
|
||||
|
||||
---
|
||||
|
||||
## Guardrails
|
||||
|
||||
- **Don't implement** - Never write code or implement features. Creating OpenSpec artifacts is fine, writing application code is not.
|
||||
- **Don't fake understanding** - If something is unclear, dig deeper
|
||||
- **Don't rush** - Discovery is thinking time, not task time
|
||||
- **Don't force structure** - Let patterns emerge naturally
|
||||
- **Don't auto-capture** - Offer to save insights, don't just do it
|
||||
- **Do visualize** - A good diagram is worth many paragraphs
|
||||
- **Do explore the codebase** - Ground discussions in reality
|
||||
- **Do question assumptions** - Including the user's and your own
|
||||
@@ -0,0 +1,110 @@
|
||||
---
|
||||
name: openspec-propose
|
||||
description: Propose a new change with all artifacts generated in one step. Use when the user wants to quickly describe what they want to build and get a complete proposal with design, specs, and tasks ready for implementation.
|
||||
license: MIT
|
||||
compatibility: Requires openspec CLI.
|
||||
metadata:
|
||||
author: openspec
|
||||
version: "1.0"
|
||||
generatedBy: "1.3.1"
|
||||
---
|
||||
|
||||
Propose a new change - create the change and generate all artifacts in one step.
|
||||
|
||||
I'll create a change with artifacts:
|
||||
- proposal.md (what & why)
|
||||
- design.md (how)
|
||||
- tasks.md (implementation steps)
|
||||
|
||||
When ready to implement, run /opsx:apply
|
||||
|
||||
---
|
||||
|
||||
**Input**: The user's request should include a change name (kebab-case) OR a description of what they want to build.
|
||||
|
||||
**Steps**
|
||||
|
||||
1. **If no clear input provided, ask what they want to build**
|
||||
|
||||
Use the **AskUserQuestion tool** (open-ended, no preset options) to ask:
|
||||
> "What change do you want to work on? Describe what you want to build or fix."
|
||||
|
||||
From their description, derive a kebab-case name (e.g., "add user authentication" → `add-user-auth`).
|
||||
|
||||
**IMPORTANT**: Do NOT proceed without understanding what the user wants to build.
|
||||
|
||||
2. **Create the change directory**
|
||||
```bash
|
||||
openspec new change "<name>"
|
||||
```
|
||||
This creates a scaffolded change at `openspec/changes/<name>/` with `.openspec.yaml`.
|
||||
|
||||
3. **Get the artifact build order**
|
||||
```bash
|
||||
openspec status --change "<name>" --json
|
||||
```
|
||||
Parse the JSON to get:
|
||||
- `applyRequires`: array of artifact IDs needed before implementation (e.g., `["tasks"]`)
|
||||
- `artifacts`: list of all artifacts with their status and dependencies
|
||||
|
||||
4. **Create artifacts in sequence until apply-ready**
|
||||
|
||||
Use the **TodoWrite tool** to track progress through the artifacts.
|
||||
|
||||
Loop through artifacts in dependency order (artifacts with no pending dependencies first):
|
||||
|
||||
a. **For each artifact that is `ready` (dependencies satisfied)**:
|
||||
- Get instructions:
|
||||
```bash
|
||||
openspec instructions <artifact-id> --change "<name>" --json
|
||||
```
|
||||
- The instructions JSON includes:
|
||||
- `context`: Project background (constraints for you - do NOT include in output)
|
||||
- `rules`: Artifact-specific rules (constraints for you - do NOT include in output)
|
||||
- `template`: The structure to use for your output file
|
||||
- `instruction`: Schema-specific guidance for this artifact type
|
||||
- `outputPath`: Where to write the artifact
|
||||
- `dependencies`: Completed artifacts to read for context
|
||||
- Read any completed dependency files for context
|
||||
- Create the artifact file using `template` as the structure
|
||||
- Apply `context` and `rules` as constraints - but do NOT copy them into the file
|
||||
- Show brief progress: "Created <artifact-id>"
|
||||
|
||||
b. **Continue until all `applyRequires` artifacts are complete**
|
||||
- After creating each artifact, re-run `openspec status --change "<name>" --json`
|
||||
- Check if every artifact ID in `applyRequires` has `status: "done"` in the artifacts array
|
||||
- Stop when all `applyRequires` artifacts are done
|
||||
|
||||
c. **If an artifact requires user input** (unclear context):
|
||||
- Use **AskUserQuestion tool** to clarify
|
||||
- Then continue with creation
|
||||
|
||||
5. **Show final status**
|
||||
```bash
|
||||
openspec status --change "<name>"
|
||||
```
|
||||
|
||||
**Output**
|
||||
|
||||
After completing all artifacts, summarize:
|
||||
- Change name and location
|
||||
- List of artifacts created with brief descriptions
|
||||
- What's ready: "All artifacts created! Ready for implementation."
|
||||
- Prompt: "Run `/opsx:apply` or ask me to implement to start working on the tasks."
|
||||
|
||||
**Artifact Creation Guidelines**
|
||||
|
||||
- Follow the `instruction` field from `openspec instructions` for each artifact type
|
||||
- The schema defines what each artifact should contain - follow it
|
||||
- Read dependency artifacts for context before creating new ones
|
||||
- Use `template` as the structure for your output file - fill in its sections
|
||||
- **IMPORTANT**: `context` and `rules` are constraints for YOU, not content for the file
|
||||
- Do NOT copy `<context>`, `<rules>`, `<project_context>` blocks into the artifact
|
||||
- These guide what you write, but should never appear in the output
|
||||
|
||||
**Guardrails**
|
||||
- Create ALL artifacts needed for implementation (as defined by schema's `apply.requires`)
|
||||
- Always read dependency artifacts before creating a new one
|
||||
- If context is critically unclear, ask the user - but prefer making reasonable decisions to keep momentum
|
||||
- If a change with that name already exists, ask if user wants to continue it or create a new one
|
||||
- Verify each artifact file exists after writing before proceeding to next
|
||||
@@ -55,3 +55,8 @@ uploads/
|
||||
/volumes
|
||||
/server.pid
|
||||
.claude/settings.local.json
|
||||
.opencode/plugins/emdash-notifications.js
|
||||
|
||||
### Windows / Runtime Artifacts
|
||||
*.stackdump
|
||||
NUL
|
||||
|
||||
@@ -1,29 +0,0 @@
|
||||
Stack trace:
|
||||
Frame Function Args
|
||||
0007FFFFB920 00021005FE8E (000210285F68, 00021026AB6E, 000000000000, 0007FFFFA820) msys-2.0.dll+0x1FE8E
|
||||
0007FFFFB920 0002100467F9 (000000000000, 000000000000, 000000000000, 0007FFFFBBF8) msys-2.0.dll+0x67F9
|
||||
0007FFFFB920 000210046832 (000210286019, 0007FFFFB7D8, 000000000000, 000000000000) msys-2.0.dll+0x6832
|
||||
0007FFFFB920 000210068CF6 (000000000000, 000000000000, 000000000000, 000000000000) msys-2.0.dll+0x28CF6
|
||||
0007FFFFB920 000210068E24 (0007FFFFB930, 000000000000, 000000000000, 000000000000) msys-2.0.dll+0x28E24
|
||||
0007FFFFBC00 00021006A225 (0007FFFFB930, 000000000000, 000000000000, 000000000000) msys-2.0.dll+0x2A225
|
||||
End of stack trace
|
||||
Loaded modules:
|
||||
000100400000 bash.exe
|
||||
7FF9B93D0000 ntdll.dll
|
||||
7FF9B79A0000 KERNEL32.DLL
|
||||
7FF9B6860000 KERNELBASE.dll
|
||||
7FF9B8740000 USER32.dll
|
||||
7FF9B6830000 win32u.dll
|
||||
7FF9B84F0000 GDI32.dll
|
||||
7FF9B6CD0000 gdi32full.dll
|
||||
7FF9B6790000 msvcp_win.dll
|
||||
000210040000 msys-2.0.dll
|
||||
7FF9B7000000 ucrtbase.dll
|
||||
7FF9B7370000 advapi32.dll
|
||||
7FF9B8E40000 msvcrt.dll
|
||||
7FF9B85B0000 sechost.dll
|
||||
7FF9B6FD0000 bcrypt.dll
|
||||
7FF9B90F0000 RPCRT4.dll
|
||||
7FF9B5F20000 CRYPTBASE.DLL
|
||||
7FF9B6710000 bcryptPrimitives.dll
|
||||
7FF9B86E0000 IMM32.DLL
|
||||
@@ -4,8 +4,13 @@
|
||||
|
||||
| 日期 | slug | 领域 | 关键词 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| 2026-07-03 | mvp-demo-trace-acceptance | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
|
||||
| 2026-05-29 | chatmodel-abstraction | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | archived |
|
||||
| 2026-06-23 | phase1-infrastructure | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | archived |
|
||||
| 2026-06-24 | lookup-knowledge-integration | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | archived |
|
||||
| 2026-06-25 | doc-management-ui | 前端开发/文档管理 | 文档管理页面, CRUD, 状态监控, 纯静态页面, API集成 | archived |
|
||||
| 2026-06-26 | session-storage | 会话存储/可观测 | diagnosis_session, agent_step, tool_invocation, token追踪, 多Agent路由 | openspec/changes/session-storage | archived |
|
||||
| 2026-06-29 | confidence-feedback | 质量评估/反馈机制 | evidence_score, selfEvaluation, feedback, useful, not_useful, case_library, BAD_CASE, tool_invocation规则引擎, 反馈按钮, sessionId回传 | openspec/changes/confidence-feedback | archived |
|
||||
| 2026-06-30 | session-dedup-knowledge-map | 去重/知识图谱 | RetrievedDocTracker, KnowledgeDomainService, knowledge_domain, covers, whenToRetrieve, Planner注入, ISS-001 | openspec/changes/archive/2026-06-30-session-dedup-knowledge-map | archived |
|
||||
| 2026-07-01 | executor-action-memory-relevance | 检索质量/行动记忆 | relevanceLevel, completenessHint, Min-Max归一化, RetrievedDocTracker域级记录, Executor检索约束, ISS-002 | openspec/changes/archive/2026-07-01-executor-action-memory-relevance | archived |
|
||||
| 2026-07-02 | chat-verifier-agent | Chat质量门禁/可追溯验证 | Verifier, groundedness_score, facts_checked, evidence_refs, tool_trace_summary, self_evaluation | openspec/changes/archive/2026-07-03-chat-verifier-agent | archived |
|
||||
|
||||
@@ -0,0 +1,64 @@
|
||||
# acceptance.md — confidence-feedback
|
||||
|
||||
## 实现清单
|
||||
|
||||
| 任务 | 文件 | 状态 |
|
||||
|---|---|---|
|
||||
| T0:Flyway V008 + answer 字段 | `V008__add_answer_to_diagnosis_session.sql`、`DiagnosisSession.java` | 完成 |
|
||||
| T1:EvaluationService(规则引擎) | `EvaluationService.java` | 完成 |
|
||||
| T2:ChatService 后置调用 | `ChatService.java` | 完成 |
|
||||
| T3:FeedbackController + FeedbackService | `FeedbackController.java`、`FeedbackService.java`、`FeedbackRequest.java`、`FeedbackResponse.java` | 完成 |
|
||||
| T4:CaseLibraryService | `CaseLibraryService.java` | 完成 |
|
||||
| T5:AsyncConfig | `AsyncConfig.java` | 完成 |
|
||||
|
||||
## 验证记录
|
||||
|
||||
### 静态验证(已通过)
|
||||
|
||||
- `mvn compile` BUILD SUCCESS(2026-06-30)
|
||||
- 无新增 ERROR,存量 WARNING 与本次改动无关
|
||||
- import 完整性人工检查通过
|
||||
|
||||
### 脚本验证(已通过,2026-06-30)
|
||||
|
||||
验证工具:`scripts/query_mysql.py`(本次新建)
|
||||
|
||||
| 步骤 | 操作 | 结果 |
|
||||
|---|---|---|
|
||||
| 1 | POST /api/chat 发送问题 | 200,answer 有值 |
|
||||
| 2 | 等 5 秒查 diagnosis_session | self_evaluation 写入规则引擎结果,answer 写入完整回答 |
|
||||
| 3 | POST /api/feedback useful | 200,返回 caseId;case_library 新增一行,feedback=useful,status=SUCCESS |
|
||||
| 4 | POST /api/feedback not_useful | 200,feedback=not_useful,status 仍为 SUCCESS(未被改写) |
|
||||
| 5(边界)| 重复提交 useful | 返回同一 caseId,case_library 无重复插入 |
|
||||
| 6(边界)| 非法 feedback 值 | HTTP 400 |
|
||||
|
||||
### Flyway V008 迁移
|
||||
|
||||
- 服务启动后 diagnosis_session 表存在 answer 列,验证通过(步骤 2 能写入 answer)
|
||||
|
||||
### 浏览器/人工验证(已通过,2026-06-30)
|
||||
|
||||
| 步骤 | 操作 | 结果 |
|
||||
|---|---|---|
|
||||
| 1 | 发送"今天天气怎么样" | AI 回复下方出现"有用/无用"按钮 |
|
||||
| 2 | 点击"有用" | 按钮区域替换为"已标记为有用" |
|
||||
| 3 | 网络请求确认 | POST /api/feedback 返回 HTTP 200,`success: true` |
|
||||
|
||||
### 前端反馈按钮(追加,2026-06-30)
|
||||
|
||||
**改动文件**:`app.js`、`styles.css`
|
||||
|
||||
关键设计:
|
||||
- `ChatResult` record 新增(`ChatService`),`ChatResponse` 增加 `sessionId` 字段(`ChatController`)
|
||||
- `sendQuickMessage` 读取 `chatResponse.sessionId` 存为 `this.lastSessionId`
|
||||
- `createFeedbackBar(sessionId)` 闭包绑定 sessionId,避免多轮对话时 sessionId 错位
|
||||
- `submitFeedback(feedback, barElement, sessionId)` 直接用传入参数,不依赖全局状态
|
||||
- 流式模式(`/api/chat_stream`)反馈按钮会渲染,但 sessionId 为空,点击不生效(已知限制)
|
||||
|
||||
## 已知限制
|
||||
|
||||
- 非检索工具(DateTimeTools 等)不写 tool_invocation,evidence_score = 0(已接受,符合"证据充分度"定义)
|
||||
- `@Async` 失败时 selfEvaluation 为 null,前端需处理 null(已接受)
|
||||
- CaseLibrary 的 faultCategory 固定为 GENERAL,需人工补充(已接受,Phase 2 优化)
|
||||
- LLM 观点层未实现,selfEvaluation JSON 预留 llm_opinion 扩展位(Phase 2)
|
||||
- 流式模式反馈按钮 sessionId 缺失,暂不处理(已知,后续处理流式接口时一并解决)
|
||||
@@ -0,0 +1,37 @@
|
||||
# brief.md — confidence-feedback
|
||||
|
||||
## 背景
|
||||
|
||||
DiagnosisSession 已预留 `selfEvaluation`(JSON)和 `feedback`(VARCHAR 16)两个字段,但完全为空。Agent 完成对话后不计算证据评分,也没有接收用户反馈的 API,无法支撑报告质量评估和 BadCase 追踪。
|
||||
|
||||
## 目标
|
||||
|
||||
1. 给每次对话结果自动打一个基于事实的证据充分度评分(evidence_score)
|
||||
2. 提供用户反馈 API(useful/not_useful),useful 触发案例自动沉淀,not_useful 标记 BadCase
|
||||
|
||||
## 范围
|
||||
|
||||
- `DiagnosisSession` 加 `answer` 字段(Flyway V008)
|
||||
- `EvaluationService`:基于 tool_invocation 的规则引擎,@Async 写 selfEvaluation
|
||||
- `FeedbackController` + `FeedbackService`:POST /api/feedback
|
||||
- `CaseLibraryService.createFromSession`:幂等案例沉淀
|
||||
- `AsyncConfig`:@EnableAsync
|
||||
- `ChatService`:SUCCESS 分支写 answer + 触发 evaluate;新增 `ChatResult` record 回传 sessionId
|
||||
- `ChatController.ChatResponse` 增加 `sessionId` 字段
|
||||
- 前端 `app.js`:AI 回复下方反馈按钮,点击调用 `/api/feedback`,闭包绑定 sessionId
|
||||
- 前端 `styles.css`:反馈栏样式
|
||||
|
||||
## 非目标
|
||||
|
||||
- 不实现 Verifier Agent 完整链路
|
||||
- 不实现 LLM 自评(预留扩展位,Phase 2 再做)
|
||||
- 不实现案例结构化字段自动填充(faultCategory 等暂时填 GENERAL)
|
||||
- 不实现 BadCase 自动分析或 Prompt 优化
|
||||
|
||||
## 分档
|
||||
|
||||
standard
|
||||
|
||||
## 关联 OpenSpec
|
||||
|
||||
`openspec/changes/confidence-feedback/`
|
||||
@@ -0,0 +1,115 @@
|
||||
# decisions.md — confidence-feedback
|
||||
|
||||
## Question Pool(grill 阶段)
|
||||
|
||||
| # | 问题 | 模式 | 状态 |
|
||||
|---|---|---|---|
|
||||
| Q1 | 置信度由谁计算 | user-interview | 已确认 |
|
||||
| Q2 | 反馈触发哪些后端操作 | user-interview | 已确认 |
|
||||
| Q3 | CaseLibrary 结构化字段从哪里填 | evidence-driven | 已确认(方案变更) |
|
||||
| Q4 | 验收口径 | user-interview | 已确认 |
|
||||
|
||||
---
|
||||
|
||||
## Evidence-Driven 结论
|
||||
|
||||
### Q3:CaseLibrary 内容来源
|
||||
|
||||
**初始结论**:从 `agent_step.thought` 提取(grill 阶段)
|
||||
|
||||
**修正(apply 阶段讨论后)**:
|
||||
- 代码证据:`agent_step.thought` 截断为 2000 字符,`modelOutput` 截断为 500 字符,均不是完整答案
|
||||
- `ChatService.executeChat` 第 269 行已有完整答案 `answer = response.getText()`,但未持久化
|
||||
- 决策:给 `DiagnosisSession` 加 `answer TEXT` 字段,Flyway V008 迁移,案例内容直接从 `session.answer` 取
|
||||
|
||||
---
|
||||
|
||||
## User-Interview 确认记录
|
||||
|
||||
### Q1 — 置信度由谁评估
|
||||
- 用户原话(grill):"两者都要:规则兜底 + Verifier 主打分"
|
||||
- **apply 后修正**:讨论后决定去掉 LLM 自评,仅用规则引擎(见"apply 阶段决策")
|
||||
- 最终实现:`EvaluationService` 纯规则,预留 `llm_opinion` 扩展位
|
||||
|
||||
### Q2 — 反馈触发操作
|
||||
- 用户原话:"写入 DiagnosisSession.feedback 字段, not_useful → 打 BAD_CASE 标记"
|
||||
- **apply 后修正**:BAD_CASE 不改 status,feedback 字段本身即为标记(见"apply 阶段决策")
|
||||
- 最终实现:`FeedbackService` 只写 feedback + 可选写 case_library,不改 status
|
||||
|
||||
### Q4 — 验收口径
|
||||
- 用户原话:"端到端可验证:发一次 chat → 查 DB 看 selfEvaluation 有值 → 提交 feedback → 查 DB 看 feedback + case_library"
|
||||
- 确认状态:已确认,未变化
|
||||
|
||||
---
|
||||
|
||||
## Apply 阶段决策(post-grill 重要变更)
|
||||
|
||||
### 决策 A:DiagnosisSession 加 answer 字段
|
||||
|
||||
- **问题**:案例沉淀需要完整答案,agent_step.thought 被截断,不可用
|
||||
- **决策**:新增 `answer LONGTEXT` 字段,ChatService SUCCESS 分支写入
|
||||
- **影响**:V008 Flyway 迁移,CaseLibraryService 直接读 session.answer
|
||||
|
||||
### 决策 B:去掉 LLM 自评,只用规则引擎
|
||||
|
||||
- **问题**:LLM 评估自己的答案系统性偏高分;多一次调用消耗 token;Verifier Agent 当前未实现
|
||||
- **决策**:MVP 阶段仅用基于 tool_invocation 的规则引擎
|
||||
- **理由**:规则可解释、可复现、不撒谎;Verifier 留待诊断全链路实现时再做
|
||||
- **预留**:`selfEvaluation` JSON 结构保留 `llm_opinion` 扩展位,代码底部注释说明接入点
|
||||
|
||||
### 决策 C:BAD_CASE 不改 status 字段
|
||||
|
||||
- **问题**:status 是执行状态语义(RUNNING/SUCCESS/FAILED),BAD_CASE 是质量标签,两个维度不同;覆盖 status 会破坏统计
|
||||
- **决策**:`not_useful` 通过 `feedback` 字段本身标识,查 BadCase 用 `WHERE feedback = 'not_useful'`
|
||||
|
||||
### 决策 D:评分字段重命名为 evidence_score
|
||||
|
||||
- **问题**:原名 confidence 容易误解为"答案准确性",实际衡量的是"证据收集充分度"
|
||||
- **决策**:重命名为 `evidence_score`,明确语义边界
|
||||
- **边界说明**:工具调用能证明 Agent 有尝试收集证据,但无法证明答案无幻觉;这个分数过滤最差情况(无工具调用就给答案),不能识别"调用了工具但结论仍错误"
|
||||
|
||||
### 决策 E:规则输入来源仅限 tool_invocation 事实
|
||||
|
||||
- **问题**:DateTimeTools、QueryMetricsTools 等非检索工具调用未写入 tool_invocation
|
||||
- **接受**:evidence_score 定义本来就是检索证据充分度,非检索工具排除在外是合理的,不是 bug
|
||||
- **已知限制**:调用了时间工具但 evidence_score = 0 的 session 存在
|
||||
|
||||
---
|
||||
|
||||
## 架构审计记录
|
||||
|
||||
- 接口影响:`POST /api/feedback` 是新接口(L2);ChatService 主流程返回值不变(L1)
|
||||
- 时序验证:tool_invocation 在工具执行时同步写入,evaluate @Async 在 Agent 完成后触发,无竞态问题
|
||||
- 已接受风险:
|
||||
- `@Async` 失败时 selfEvaluation 保持 null,前端需处理 null
|
||||
- 案例结构化字段(faultCategory 等)暂时填 GENERAL,后续可人工补充
|
||||
- LLM 自评预留但未实现,Phase 2 再迭代
|
||||
|
||||
### 决策 F:ChatResult record + ChatResponse.sessionId 回传
|
||||
|
||||
- **问题**:`ChatService` 内部生成 8 位 sessionId,但从不返回给前端;前端用自己的 sessionId 调 feedback 接口,后端查不到 session(400)
|
||||
- **决策**:新增 `ChatResult(answer, sessionId)` record,`executeChatWithStrategy` 链路全部返回 `ChatResult`;`ChatResponse` 增加 `sessionId` 字段;前端读取并闭包绑定至对应消息的反馈按钮
|
||||
- **影响**:`ChatService` 三个方法签名变更(内部链路),`ChatController` 调用方更新,前端 `app.js` 读取新字段
|
||||
|
||||
### 决策 G:反馈 sessionId 闭包绑定而非全局变量
|
||||
|
||||
- **问题**:最初实现用 `this.lastSessionId` 全局变量,多轮对话时点击早期消息的反馈按钮会提交最新 sessionId
|
||||
- **决策**:`createFeedbackBar(sessionId)` 接收 sessionId 参数,`submitFeedback(feedback, bar, sessionId)` 直接用传入值,不读全局状态
|
||||
- **效果**:每条 AI 回复绑定自己那轮的 sessionId,多轮对话下行为正确
|
||||
|
||||
### 项目技术栈清单
|
||||
|
||||
- ChatModel 注入:`@Autowired ChatModel chatModel`,通过 `ModelRoutingConfig` 路由
|
||||
- Repository:Spring Data JPA,`Optional<T>` 返回,方法命名约定
|
||||
- DTO:独立文件放 `dto/` 包
|
||||
- 异步:新建 `AsyncConfig.java` 加 `@EnableAsync`(项目原无此配置)
|
||||
- 无 MQ,无加密,工具类直接用 UUID.randomUUID()
|
||||
- 日志:SLF4J Logger,`LoggerFactory.getLogger()`
|
||||
- `ToolInvocationRepository.findBySessionId` 已有,可直接用
|
||||
|
||||
### 参考实现文件
|
||||
|
||||
- `ChatService.java`:executeChat/executeChatComplex 流程
|
||||
- `CaseLibraryRepository.findByDiagnosisId`:幂等检查用
|
||||
- `DiagnosisSessionRepository.findBySessionId`
|
||||
- `ToolInvocationRepository.findBySessionId`
|
||||
@@ -0,0 +1,52 @@
|
||||
# evidence.md — confidence-feedback
|
||||
|
||||
## 代码证据
|
||||
|
||||
### agent_step.thought 不可作为案例内容
|
||||
|
||||
- 文件:`AgentLoggingHook.java:135`
|
||||
- 证据:`thought` 在写入前截断为 2000 字符,`modelOutput` 截断为 500 字符
|
||||
- 结论:两者均不是返回给用户的完整答案,案例质量低
|
||||
|
||||
### ChatService 已有完整答案未持久化
|
||||
|
||||
- 文件:`ChatService.java:269`(executeChat)、`ChatService.java:353`(executeChatComplex)
|
||||
- 证据:`String answer = response.getText()` 只用于返回前端,未写入任何持久化存储
|
||||
- 结论:加 `DiagnosisSession.answer` 字段是最干净的方案
|
||||
|
||||
### ToolInvocationRepository 已有 findBySessionId
|
||||
|
||||
- 文件:`ToolInvocationRepository.java`
|
||||
- 证据:`findBySessionId(String sessionId)` 已实现,返回 `List<ToolInvocation>`
|
||||
- 结论:规则引擎可直接读取 tool_invocation 事实,无需新增查询方法
|
||||
|
||||
### tool_invocation 写入时序安全
|
||||
|
||||
- 文件:`LookupKnowledgeTool.java:144`
|
||||
- 证据:`saveToolInvocation` 在工具执行时同步调用,早于 ChatService 的 SUCCESS 分支
|
||||
- 结论:@Async evaluate 触发时 tool_invocation 数据已在库,无竞态
|
||||
|
||||
### 项目原无 @EnableAsync
|
||||
|
||||
- 证据:`grep -rn "EnableAsync"` 无任何命中(apply 前)
|
||||
- 结论:需要新建 `AsyncConfig.java`
|
||||
|
||||
### CaseLibraryRepository.findByDiagnosisId 已有幂等检查支持
|
||||
|
||||
- 文件:`CaseLibraryRepository.java`
|
||||
- 证据:`findByDiagnosisId(String diagnosisId)` 已实现
|
||||
- 结论:useful 重复提交时可用此方法检查,不重复插入
|
||||
|
||||
## 设计推导
|
||||
|
||||
### evidence_score vs confidence 命名
|
||||
|
||||
- 基于工具调用的分数衡量的是证据收集充分度,不是答案准确性
|
||||
- "confidence" 容易误解,改为 "evidence_score" 更准确
|
||||
- LLM 自评才适合叫 confidence,但当前未实现
|
||||
|
||||
### BAD_CASE 不应混入 status
|
||||
|
||||
- status 有明确执行状态语义(RUNNING/SUCCESS/FAILED)
|
||||
- 一个 SUCCESS 的 session 被标为 BAD_CASE 后,按 status 做的统计会失真
|
||||
- feedback 字段本身就够,`WHERE feedback = 'not_useful'` 即可查 BadCase
|
||||
@@ -0,0 +1,58 @@
|
||||
# Acceptance: session-dedup-knowledge-map
|
||||
|
||||
## 静态验证
|
||||
|
||||
| 项目 | 结果 | 说明 |
|
||||
|------|------|------|
|
||||
| 编译检查 | PASS | `mvn compile -q` exit code 0,所有 17 个变更文件无编译错误 |
|
||||
| 代码结构检查 | PASS | 6 个新文件(RetrievedDocTracker, DocumentFieldEnricher, KnowledgeDomainService, KnowledgeDomain, KnowledgeDomainRepository, V009 迁移)均存在且路径正确 |
|
||||
| Prompt 外部化 | PASS | `doc-field-enricher-prompt.md` 和 `domain-summary-prompt.md` 位于 `src/main/resources/prompts/`,Java 代码通过 `@PostConstruct` + `ClassPathResource` 加载 |
|
||||
| Flyway 迁移脚本 | PASS | `V009__add_knowledge_domain.sql` 存在,表结构完整 |
|
||||
| DTO 字段 | PASS | Frontmatter / KnowledgeEntry / LookupResult 新增字段均已添加 |
|
||||
| 解析器扩展 | PASS | FrontmatterParser 解析 `covers` 和 `when_to_retrieve` |
|
||||
| Jackson 替换 | PASS | KnowledgeIndexService 不再包含 extractJsonValue/extractJsonArray,改用 objectMapper.readValue |
|
||||
| Prompt 检索规则 | PASS | chat-planner-prompt.md 新增"知识库检索规则"区块(4 条规则) |
|
||||
|
||||
## 脚本验证
|
||||
|
||||
| 项目 | 结果 | 说明 |
|
||||
|------|------|------|
|
||||
| 单元测试 | 未运行 | 项目当前无针对本 change 的单元测试 |
|
||||
| 集成测试 | 未运行 | 需启动应用 + Milvus + MySQL 验证完整链路 |
|
||||
|
||||
## 浏览器/人工验证
|
||||
|
||||
| 项目 | 结果 | 说明 |
|
||||
|------|------|------|
|
||||
| V009 迁移 | PASS | Flyway 日志:`Successfully applied 1 migration to schema superbiz_agent, now at version v009` |
|
||||
| knowledge_domain 表数据 | PASS | 4 个域全部 LLM 生成 when_to_retrieve 成功(api/domain/infrastructure/troubleshooting),内容包含跨域边界引用 |
|
||||
| knowledge map 注入 Planner | PASS | 多 Agent 路径正常触发 `Supervisor → chat_planner → chat_executor`,Planner 能按域做检索规划 |
|
||||
| session 级去重 | PASS | 两个 session 均验证去重生效:session `7c517329` 去 4 次重拦截,session `9693b9fb` 6 次去重拦截 |
|
||||
| LLM 字段生成 | 未验证 | 需上传新文档后检查 metadata JSON 中是否包含 covers 和 whenToRetrieve |
|
||||
|
||||
## 未验证项
|
||||
|
||||
| 项目 | 风险 | 建议补验步骤 |
|
||||
|------|------|-------------|
|
||||
| LLM 字段生成 | 中 — 依赖外部 LLM 服务 | 上传新文档,检查 metadata JSON 中是否包含 covers 和 whenToRetrieve |
|
||||
|
||||
## 启动问题修复
|
||||
|
||||
| 问题 | 修复 | 状态 |
|
||||
|------|------|------|
|
||||
| `@PostConstruct` 中调用 `knowledgeDomainService.onDocumentChange()` 导致循环依赖 | 将域级生成从 `@PostConstruct` 移到 `@EventListener(ApplicationReadyEvent.class)` | 已修复,编译通过 |
|
||||
|
||||
## 任务完成状态
|
||||
|
||||
14/14 任务全部完成 (T1-1 ~ T6-2)。
|
||||
|
||||
## 遗留问题
|
||||
|
||||
ISS-002:Executor 无约束重复调用 `lookup_knowledge`(单会话 20+ 次),knowledge map 和检索约束只注入了 Planner 未注入 Executor。详见 `mvp/issues/ISS-002-executor-unconstrained-lookup.md`。
|
||||
|
||||
## 已知限制
|
||||
|
||||
1. **RetrievedDocTracker 为 JVM 内存存储**:应用重启后去重状态丢失,同一会话内重启无法继续去重(可接受,会话通常短于重启间隔)
|
||||
2. **Planner 只看域级 when_to_retrieve**:文档级细粒度筛选留 Phase 2
|
||||
3. **文档级 prompt 依赖同域其他文档**:首个上传到某域的文档无法获得同域参照(此时 prompt 输出"无同域其他文档")
|
||||
4. **域级 prompt 依赖其他域已入库**:首次启动且 DB 为空时,其他域信息从 L0 索引 category 列表兜底
|
||||
@@ -0,0 +1,33 @@
|
||||
# Brief: session-dedup-knowledge-map
|
||||
|
||||
## 背景
|
||||
|
||||
ISS-001:Executor 在单次对话中重复调用 `lookup_knowledge` 多达 20 次,同一文档被召回 13 次。原因是工具层无状态、Planner 无知识边界感知。
|
||||
|
||||
## 目标
|
||||
|
||||
1. 彻底消除 session 内重复文档召回(Part A)
|
||||
2. 给 Planner 注入知识图谱,让其在规划阶段就能判断需要检索哪个域、只检索一次(Part B)
|
||||
|
||||
## 范围
|
||||
|
||||
- `LookupKnowledgeTool`:session 级去重
|
||||
- `Frontmatter` / `KnowledgeEntry`:新增 covers + whenToRetrieve
|
||||
- `DocumentManagementService`:上传时 LLM 生成文档级字段
|
||||
- `KnowledgeDomainService`(新):域级聚合与 DB 存储
|
||||
- `knowledge_domain` 表(新)
|
||||
- `ChatService` + `chat-planner-prompt.md`:注入 knowledge map
|
||||
|
||||
## 非目标(Phase 2)
|
||||
|
||||
- Executor 文档级 when_to_retrieve 细粒度筛选
|
||||
- RRF 混合重排
|
||||
- 文档 frontmatter 自动生成(手动覆盖 LLM 优先已支持)
|
||||
|
||||
## 分档
|
||||
|
||||
standard
|
||||
|
||||
## 关联 OpenSpec
|
||||
|
||||
openspec/changes/session-dedup-knowledge-map/
|
||||
@@ -0,0 +1,60 @@
|
||||
# decisions.md — session-dedup-knowledge-map
|
||||
|
||||
## Question Pool
|
||||
|
||||
| # | 问题 | 类型 | 状态 |
|
||||
|---|---|---|---|
|
||||
| Q1 | domain.when_to_retrieve 来源(手动/自动聚合/LLM上传时生成) | user-interview | 已确认 |
|
||||
| Q2 | LLM 生成时机(同步上传 vs 异步补全) | user-interview | 已确认 |
|
||||
| Q3 | knowledge map 结构(域级平铺 vs 两层) | user-interview | 已确认 |
|
||||
| Q4 | domain.when_to_retrieve 存储(内存 vs DB) | user-interview | 已确认 |
|
||||
| Q5 | Executor 文档级细粒度筛选是否进 MVP | user-interview | 已确认 |
|
||||
| E1 | ThreadLocal 在多 Agent 路径是否安全 | evidence-driven | 已汇报 |
|
||||
| E2 | 6 个文档是否全部有 category 字段 | evidence-driven | 已汇报 |
|
||||
| E3 | 去重 key 设计 | evidence-driven | 已汇报 |
|
||||
| E4 | Planner prompt token 增量是否可接受 | evidence-driven | 已汇报 |
|
||||
| E5 | EvaluationService.tool_call_count 影响 | evidence-driven | 已汇报 |
|
||||
|
||||
## Evidence-Driven 结论
|
||||
|
||||
- **E1**:`AsyncConfig` 只启用 `@EnableAsync`,无 TaskDecorator。`SupervisorAgent.invoke()` 是同步阻塞调用,工具调用与主线程同线程,ThreadLocal 当前路径安全。异步扩展时需补 TaskDecorator。
|
||||
- **E2**:全部 6 个文档均有 `category` 字段:api(1)、domain(1)、infrastructure(3)、troubleshooting(1)。
|
||||
- **E3**:`KnowledgeEntry.filePath` 在 L0 内唯一,L1 `_source` 字段也是 filePath,统一用 filePath 作去重 key。
|
||||
- **E4**:当前 planner prompt 21 行,注入 knowledge map 约增加 200-400 字符,可接受。
|
||||
- **E5**:去重后 `agent_step.has_tool_call` 减少,`tool_call_count` 降低,这是修复效果,`EvaluationService` 评分规则无需改动。
|
||||
|
||||
## User-Interview 确认记录
|
||||
|
||||
**Q1** — doc.when_to_retrieve 来源
|
||||
用户原话:选 C(上传时 LLM 自动生成)
|
||||
确认状态:已确认
|
||||
|
||||
**Q2** — LLM 生成时机
|
||||
用户原话:选 X(同步,上传时当场生成)
|
||||
确认状态:已确认
|
||||
|
||||
**Q3** — knowledge map 结构
|
||||
用户原话:认可两层结构(domain → documents[])
|
||||
确认状态:已确认
|
||||
补充:Planner 只注入域级 when_to_retrieve,文档级 when_to_retrieve 留 Executor 筛选(Phase 2)
|
||||
|
||||
**Q4** — domain.when_to_retrieve 存储
|
||||
用户原话:存 DB,这样每次启动都不用让 LLM 再总结一次
|
||||
确认状态:已确认 → 新建 knowledge_domain 表,Flyway 迁移脚本
|
||||
|
||||
**Q5** — Executor 文档级细粒度筛选
|
||||
用户原话:留 Phase 2
|
||||
确认状态:已确认,MVP 不做
|
||||
|
||||
## Pre-apply 补充决策
|
||||
|
||||
- **P1:KnowledgeIndexService.parseDocumentToEntry 替换为 Jackson**:`extractJsonValue` / `extractJsonArray` 手写解析器遇到含逗号、引号的自然语言字段(whenToRetrieve)会截断。全量替换为 `objectMapper.readValue(metadata, Frontmatter.class)`,影响范围仅 `KnowledgeIndexService`,行为更健壮。(用户确认)
|
||||
- **P2:LookupResult 新增 message 字段**:去重命中时 `found=false` + `message="文档已在本会话中检索过:xxx"`,不复用 `primary.content`。语义清晰,LLM 能理解原因不会重试。(用户确认)
|
||||
|
||||
## 关键设计决策
|
||||
|
||||
1. **两级 when_to_retrieve**:文档级(upload 时 LLM 生成,存 metadata)+ 域级(文档变更时 LLM 聚合,存 knowledge_domain 表)
|
||||
2. **域级重算触发**:文档上传后、文档删除后,只重算受影响的域(不是全量);`loadIndex()` 时如果某域在 DB 没有记录,则触发生成
|
||||
3. **注入 Planner 只给域级**:knowledge map 只包含域级 when_to_retrieve + documents[](title + covers),不暴露文档级 when_to_retrieve
|
||||
4. **去重 key**:filePath(L0+L1 统一)
|
||||
5. **去重状态存储**:JVM 内 `ConcurrentHashMap<sessionId, Set<filePath>>`,`SessionContextHolder.clear()` 时同步清理
|
||||
@@ -0,0 +1,87 @@
|
||||
# Evidence: session-dedup-knowledge-map
|
||||
|
||||
## E1: ThreadLocal 在多 Agent 路径是否安全
|
||||
|
||||
**问题**:`SessionContextHolder` 基于 ThreadLocal,多 Agent 异步路径可能导致 sessionId 丢失。
|
||||
|
||||
**证据**:
|
||||
- `AsyncConfig` 只启用 `@EnableAsync`,无 `TaskDecorator`
|
||||
- `SupervisorAgent.invoke()` 是同步阻塞调用,工具调用与主线程同线程
|
||||
- 当前路径下 ThreadLocal 安全
|
||||
|
||||
**结论**:当前同步路径安全。未来引入异步扩展时需补 `TaskDecorator` 传递 ThreadLocal。
|
||||
|
||||
---
|
||||
|
||||
## E2: 6 个文档是否全部有 category 字段
|
||||
|
||||
**问题**:域聚合依赖 `category` 字段分组,需确认现有文档是否都有值。
|
||||
|
||||
**证据**:
|
||||
- 全部 6 个文档均有 `category` 字段:api(1)、domain(1)、infrastructure(3)、troubleshooting(1)
|
||||
|
||||
**结论**:现有文档无需修补,category 覆盖率 100%。
|
||||
|
||||
---
|
||||
|
||||
## E3: 去重 key 设计
|
||||
|
||||
**问题**:用什么字段唯一标识一个文档用于去重。
|
||||
|
||||
**证据**:
|
||||
- `KnowledgeEntry.filePath` 在 L0 索引内唯一
|
||||
- L1 向量索引的 `_source` 字段也是 filePath
|
||||
- 上传时 `saveToLocal()` 生成 `knowledge_base/{category}/{fileName}` 路径
|
||||
|
||||
**结论**:统一用 `filePath` 作去重 key,L0 和 L1 一致。
|
||||
|
||||
---
|
||||
|
||||
## E4: Planner prompt token 增量是否可接受
|
||||
|
||||
**问题**:knowledge map YAML 注入 Planner prompt 会增加固定 token 开销。
|
||||
|
||||
**证据**:
|
||||
- 当前 planner prompt 21 行
|
||||
- 注入 knowledge map 约增加 200-400 字符(6 个文档场景)
|
||||
- 相比 Planner 整体 prompt + 历史消息,增量占比 < 5%
|
||||
|
||||
**结论**:可接受,不构成性能瓶颈。
|
||||
|
||||
---
|
||||
|
||||
## E5: EvaluationService.tool_call_count 影响
|
||||
|
||||
**问题**:去重后 `tool_call_count` 降低,是否影响 `EvaluationService` 评分逻辑。
|
||||
|
||||
**证据**:
|
||||
- `EvaluationService` 使用 `tool_call_count` 作为评分因子
|
||||
- 去重导致重复调用被过滤,`tool_call_count` 下降
|
||||
- 这是修复效果(消除了无意义的重复调用),不是回归
|
||||
|
||||
**结论**:`EvaluationService` 评分规则无需改动。下降的 `tool_call_count` 反映了真实效率提升。
|
||||
|
||||
---
|
||||
|
||||
## P1: 手写 JSON 解析器脆弱性
|
||||
|
||||
**问题**:`KnowledgeIndexService.extractJsonValue` / `extractJsonArray` 在遇到含逗号、引号的自然语言字段时会截断。
|
||||
|
||||
**证据**:
|
||||
- `whenToRetrieve` 字段由 LLM 生成,内容为自然语言(含逗号、分号等标点)
|
||||
- 手写解析器以 `"` 和 `,` 作分隔符,自然语言中的标点会导致提前截断
|
||||
- Jackson `ObjectMapper.readValue(metadata, Frontmatter.class)` 是项目已有依赖
|
||||
|
||||
**结论**:全量替换为 Jackson,影响范围仅 `KnowledgeIndexService.parseDocumentToEntry()`,行为更健壮。
|
||||
|
||||
---
|
||||
|
||||
## P2: LookupResult 去重提示字段
|
||||
|
||||
**问题**:去重命中时如何向 LLM 返回"不要重试"的信号。
|
||||
|
||||
**证据**:
|
||||
- 复用 `primary.content` 语义不清,LLM 可能理解为正常检索结果
|
||||
- 独立 `message` 字段 + `found=false` 语义明确,LLM 能理解"已检索过"不再重试
|
||||
|
||||
**结论**:`LookupResult` 新增 `String message` 字段,去重时填入提示文本。
|
||||
@@ -0,0 +1,70 @@
|
||||
# Acceptance: executor-action-memory-relevance
|
||||
|
||||
## 分档
|
||||
|
||||
standard
|
||||
|
||||
## 任务完成状态
|
||||
|
||||
| 任务 | 状态 | 说明 |
|
||||
|------|------|------|
|
||||
| T1: RetrievedDocTracker 域级升级 | ✅ 完成 | 双层 Map 结构,域级+文档级记录 |
|
||||
| T2: LookupResult 新增字段 | ✅ 完成 | relevanceLevel / completenessHint / retrievedDomainsThisSession |
|
||||
| T3: 归一化计算逻辑 | ✅ 完成 | Min-Max 归一化 + 三等级判定 |
|
||||
| T4: LookupKnowledgeTool 集成 | ✅ 完成 | 归一化层 + 行动记忆注入 + 域拦截 |
|
||||
| T5: Executor Prompt 重写 | ✅ 完成 | 4 条检索约束,无 knowledge map |
|
||||
| T6: 入库可观测性 | ✅ 完成 | V010 + Entity + JSON 扩展 |
|
||||
| T7: BGE-M3 归一化验证测试 | ✅ 完成 | 范数=1.00000002,测试通过 |
|
||||
|
||||
## 静态验证
|
||||
|
||||
- [x] **语法/编译检查**: 所有 Java 文件编译通过
|
||||
- [x] **Impact Analysis**: LookupKnowledgeTool、RetrievedDocTracker 变更范围经 `gitnexus_impact` 检查,均为 L2 内部接口影响
|
||||
- [x] **Cross-artifact 对齐检查**: brief → proposal → design → specs → tasks 闭环,无 gap
|
||||
- [x] **Prompt 约束检查**: chat-executor-prompt.md 不包含 knowledge map,包含 4 条检索约束
|
||||
|
||||
## 脚本验证
|
||||
|
||||
- [x] **V010 Flyway 迁移**: 迁移成功,`relevance_level` 和 `dedup_reason` 列已添加
|
||||
```sql
|
||||
ALTER TABLE tool_invocation
|
||||
ADD COLUMN relevance_level VARCHAR(20),
|
||||
ADD COLUMN dedup_reason VARCHAR(32);
|
||||
```
|
||||
- [x] **FullPipelineSmokeTest**: BGE-M3 归一化测试通过(范数=1.00000002)
|
||||
- [x] **数据库数据校验**:
|
||||
- `relevance_level` 列已写入 HIGHLY_RELEVANT / REFERENCE
|
||||
- `dedup_reason` 列已写入 doc_retrieved / null
|
||||
- `retrieval_details` JSON 包含 l1_top_similarity、completeness_hint、retrieved_domains、dedup_reason
|
||||
|
||||
## 浏览器/人工验证
|
||||
|
||||
- [x] **应用启动验证**: Spring Boot 应用正常启动,端口 9900
|
||||
- [x] **Chat API 调用验证**: 通过 curl 测试 chat 接口,lookup_knowledge 调用链完整
|
||||
```
|
||||
curl -X POST "http://localhost:9900/api/chat/send" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"sessionId": "b66d799e", "question": "..."}'
|
||||
```
|
||||
- [x] **日志验证**: 应用日志可观察到 relevanceLevel、retrievedDomainsThisSession 输出
|
||||
- [x] **归一化数学验证**: l1_top_score=0.383 → l1_top_similarity=0.8085(`1 - 0.383/2.0 = 0.8085`)✅
|
||||
- [x] **域追踪验证**: `[infrastructure]` → `[infrastructure, api]` 域列表正常扩展
|
||||
|
||||
## 未验证
|
||||
|
||||
| 场景 | 原因 | 风险 | 补验建议 |
|
||||
|------|------|------|---------|
|
||||
| PRECISE 等级(L0 唯一精确匹配) | 测试会话无精确匹配场景 | 低 — L0 matchCount=1 的判断逻辑与 HIGHLY_RELEVANT 共用,实现确定性强 | 构造一条 L0 精确匹配的知识库文档后测试 |
|
||||
| domain_retrieved 域级去重 | 需要同一域全部文档已检索再查该域才触发 | 低 — isDomainRetrieved 逻辑简单,与 isDocRetrieved 等价 | Phase 2 启用域级硬限流时测试 |
|
||||
| DEDUPED 等级 | 当前 code path 去重时仍写 REFERENCE,DEDUPED 未被使用 | 低 — 设计预留,当前未启用 | Phase 2 若启用 DEDUPED 等级时验证 |
|
||||
| Phase 2 域级硬限流 | 非本次范围 | 中 — 当前仅有软约束(prompt),LLM 仍可能在 REFERENCE 下继续检索 | 实测观察,如果 lookup 调用仍偏高,启动 Phase 2 |
|
||||
|
||||
## 剩余风险
|
||||
|
||||
1. **Prompt 软约束局限性**:实测 10 次调用中 9 次为 REFERENCE,说明 LLM 仍倾向于继续检索。如果 prompt 约束效果不足,需启用 Phase 2 域级硬限流。
|
||||
2. **L1 Metadata 解析兼容性**:L1 domain 兜底路径解析 metadata JSON,如果知识库文档 frontmatter 格式不一致可能解析失败,已有 try-catch 兜底。
|
||||
|
||||
## 归档状态
|
||||
|
||||
- [ ] OpenSpec change 尚未归档
|
||||
- [ ] devflow/index.md 状态为 `implemented`,待改为 `archived`
|
||||
@@ -0,0 +1,35 @@
|
||||
# Brief: executor-action-memory-relevance
|
||||
|
||||
## 背景
|
||||
|
||||
ISS-002:Executor 在单次会话中调用 `lookup_knowledge` 20+ 次,大部分是同域换变体的冗余调用。前序 change `session-dedup-knowledge-map` 解决了文档级重复召回(ISS-001),但未解决 Executor 重复调用问题。
|
||||
|
||||
## 目标
|
||||
|
||||
- Executor 获得行动记忆(知道自己本次会话已检索了哪些域)
|
||||
- 检索结果提供归一化质量等级(PRECISE/HIGHLY_RELEVANT/REFERENCE)+ 兜底信号
|
||||
- Executor prompt 提供明确的检索约束和"放弃检索"的合法出口
|
||||
- 原始分数入库保留可观测性,但不暴露给 LLM
|
||||
|
||||
## 范围
|
||||
|
||||
- `RetrievedDocTracker`:域级 + 文档级双层记录
|
||||
- `LookupKnowledgeTool`:归一化层 + 行动记忆注入
|
||||
- `LookupResult`:新增 relevanceLevel / completenessHint / retrievedDomainsThisSession
|
||||
- `chat-executor-prompt.md`:检索约束重写
|
||||
- `ToolInvocation` + V010:入库可观测性
|
||||
|
||||
## 非目标
|
||||
|
||||
- 不给 Executor 注入 knowledge map(保持 Agent 边界)
|
||||
- 不修改 Planner prompt 或 Planner 逻辑
|
||||
- 不修改 PrimaryResult / SupplementResult 的字段(不暴露原始分数)
|
||||
- Phase 2 域级硬限制暂不实施
|
||||
|
||||
## 分档
|
||||
|
||||
standard
|
||||
|
||||
## 关联 OpenSpec change
|
||||
|
||||
openspec/changes/executor-action-memory-relevance
|
||||
@@ -0,0 +1,83 @@
|
||||
# Decisions: executor-action-memory-relevance
|
||||
|
||||
## 过程日志
|
||||
|
||||
### Clarify 阶段
|
||||
|
||||
**入口摘要**:ISS-002 Executor 无约束重复调用 lookup_knowledge(单会话 20+ 次),需要行动记忆 + 归一化质量等级 + prompt 约束来解决。
|
||||
|
||||
**slug**: `executor-action-memory-relevance`
|
||||
|
||||
**规模分档**: `standard`(涉及 7 个文件,跨 DTO/工具层/持久化/Prompt,有设计决策需澄清)
|
||||
|
||||
### Context 阶段
|
||||
|
||||
**devflow/index.md 使用状态**: 已命中。前序 change `session-dedup-knowledge-map`(archived)提供了 RetrievedDocTracker、KnowledgeDomainService、ISS-002 文档。
|
||||
|
||||
**相关 ADR**: 无直接 ADR,但 `session-dedup-knowledge-map` 的 decisions.md 和 evidence.md 记录了文档级去重和 knowledge map 注入的决策。
|
||||
|
||||
**不能违反的历史决策**:
|
||||
1. RetrievedDocTracker 的文档级去重必须保留
|
||||
2. knowledge map 只注入 Planner,不注入 Executor(本次讨论确认)
|
||||
3. L0/L1 原始分数不暴露给 LLM,只在归一化层内部使用(本次讨论确认)
|
||||
|
||||
**需进入 OpenSpec 的上下文点**:
|
||||
1. L1 score 是 L2 距离(值域 [0,+∞)),不是归一化分数——阈值设计需基于实际分布
|
||||
2. L0 的 category 可从 KnowledgeEntry.getCategory() 直接获取;L1 需解析 metadata JSON
|
||||
3. ReactAgent 是自主决策工具调用的 Agent,Prompt 约束是软约束
|
||||
|
||||
### Grill 阶段 — Question Pool
|
||||
|
||||
**维度:术语**
|
||||
1. [evidence-driven] `relevanceLevel` 三个等级(PRECISE/HIGHLY_RELEVANT/REFERENCE)的边界是否清晰,是否存在 LLM 误解的可能? → **已查证**:三个等级语义明确,PRECISE=唯一匹配、HIGHLY_RELEVANT=高分命中、REFERENCE=低置信度参考。LLM 理解风险低。
|
||||
|
||||
**维度:边界**
|
||||
2. [evidence-driven] L1 score 是 L2 距离(值域 [0,+∞)),当前代码无阈值判断。归一化阈值如何设计? → **已查证**:L2 距离典型范围取决于 BGE-M3 1024 维 embedding 的尺度,需从 `tool_invocation.retrieval_details` 中查询实际 `l1_scores` 分布才能定阈值。当前先以常量定义,标记为"需实测校准"。
|
||||
3. [evidence-driven] L1 结果的 category 提取需要解析 metadata JSON 字符串,当前 `SearchResult.metadata` 是 `toString()` 的结果。归一化层是否需要 L1 的 domain? → **已查证**:L1 的 domain 主要用于 RetrievedDocTracker 的域级记录。如果 L0 已命中且包含 category,可直接用 L0 的 category;如果仅 L1 命中,需解析 metadata 提取 category。当前知识库中 L0 大概率先命中,L1 domain 提取作为兜底路径。
|
||||
4. [user-interview] 归一化阈值(L1 score 分界线)在实测数据不足时,是否接受先用保守初始值 + 后续调优的策略? → **用户待确认**
|
||||
|
||||
**维度:验收**
|
||||
5. [evidence-driven] 现有 `tool_invocation` 表 `retrieval_details` JSON 中 `l1_scores` 存的是 L2 距离原始值,新增的 `relevance_level` 和 `completeness_hint` 入库后是否需要回填历史数据? → **已查证**:不需要回填历史数据,新列 nullable 即可,历史记录 relevance_level=null。
|
||||
|
||||
### Grill 结论
|
||||
|
||||
**evidence-driven 汇报**:
|
||||
- E1: relevanceLevel 三等级语义清晰,LLM 误解风险低
|
||||
- E2: L1 score 是 L2 距离,值域不固定,阈值需实测校准
|
||||
- E3: L0 category 直接可用,L1 category 需解析 metadata(兜底路径)
|
||||
- E4: 历史数据不回填,新列 nullable
|
||||
|
||||
**user-interview 已确认**:
|
||||
- Q4: 归一化阈值先用保守初始值 + 后续调优 → **用户已确认**,并建议用 Min-Max 归一化到 [0,1]
|
||||
|
||||
### Specify 阶段补充
|
||||
|
||||
**BGE-M3 L2 归一化实测验证**:
|
||||
- FullPipelineSmokeTest.embeddingBgeM3Works() 新增 L2 范数断言
|
||||
- 结果:范数=1.00000002,误差 < 0.01,测试通过
|
||||
- 结论:BGE-M3 输出为 L2 归一化单位向量,L2 距离数学硬上界 = 2.0
|
||||
- Min-Max 归一化公式:`similarity = 1 - min(l2Score, 2.0) / 2.0`
|
||||
|
||||
**Cross-artifact 对齐检查**:
|
||||
|
||||
| 对齐项 | 状态 |
|
||||
|--------|------|
|
||||
| brief 目标/范围/非目标 → proposal 覆盖 | 已对齐 |
|
||||
| proposal 范围/约束 → design 覆盖 | 已对齐 |
|
||||
| design 归一化/行动记忆/接口影响 → specs 覆盖 | 已对齐 |
|
||||
| specs 可观察行为 → tasks 覆盖 | 已对齐 |
|
||||
|
||||
**接口影响分级**:
|
||||
- RetrievedDocTracker 数据结构升级 → L2(内部接口,消费者只有 LookupKnowledgeTool)
|
||||
- LookupResult 新增 3 字段 → L2(工具返回值,无跨模块调用方)
|
||||
- tool_invocation 新增 2 列 → L2(Flyway nullable,不影响现有查询)
|
||||
- chat-executor-prompt.md 更新 → L1(Prompt 文本变更)
|
||||
|
||||
### Audit 阶段
|
||||
|
||||
**架构风险评估**(5 句以内):
|
||||
1. 归一化层嵌入 LookupKnowledgeTool 内部(静态方法),无跨模块耦合风险。
|
||||
2. RetrievedDocTracker 升级为双层结构,数据量级不变(文档数 × session 数),内存无风险。
|
||||
3. L1 metadata 解析 category 是兜底路径,如果 JSON 格式不一致可能解析失败——已有 try-catch 兜底。
|
||||
4. 归一化阈值 yml 配置化,运行时调优不需要改代码和重启——运维友好。
|
||||
5. Prompt 约束仍依赖 LLM 遵守——如果 Phase 1 效果不足,Phase 2 域级硬限制的 isDomainRetrieved 已就绪,无需额外改造。
|
||||
@@ -0,0 +1,65 @@
|
||||
# Evidence: executor-action-memory-relevance
|
||||
|
||||
## Evidence-driven 结论
|
||||
|
||||
### E1: relevanceLevel 三等级语义清晰度
|
||||
|
||||
- **来源**: Grill 阶段 Question Pool #1
|
||||
- **查证结果**: 三个等级语义明确,边界清晰:
|
||||
- PRECISE:L0 唯一精确匹配,LLM 应直接使用
|
||||
- HIGHLY_RELEVANT:归一化 similarity ≥ 0.75,高度相关
|
||||
- REFERENCE:归一化 similarity ≥ 0.5,相关参考
|
||||
- **结论**: LLM 误解风险低,语义边界足够清晰
|
||||
|
||||
### E2: L1 Score 值域与归一化阈值
|
||||
|
||||
- **来源**: Grill 阶段 Question Pool #2
|
||||
- **查证结果**:
|
||||
- L1 score 是 L2 距离,值域 [0, +∞)
|
||||
- BGE-M3 输出为 L2 归一化单位向量(实测范数=1.00000002),L2 距离数学硬上界 = 2.0
|
||||
- Min-Max 归一化公式:`similarity = 1 - min(l2Score, 2.0) / 2.0`
|
||||
- **结论**: 使用 `maxL2Distance=2.0` 作为归一化上界,阈值 yml 可配置
|
||||
|
||||
### E3: L1 Domain 提取兜底路径
|
||||
|
||||
- **来源**: Grill 阶段 Question Pool #3
|
||||
- **查证结果**:
|
||||
- L0 的 domain 可从 `KnowledgeEntry.getCategory()` 直接获取
|
||||
- L1 结果的 domain 需解析 `SearchResult.metadata` JSON 字符串
|
||||
- 当前知识库设计下 L0 大概率先命中,L1 domain 提取作为兜底
|
||||
- **结论**: 先尝试 L0 category,失败时解析 L1 metadata JSON(try-catch 兜底)
|
||||
|
||||
### E4: 历史数据不回填
|
||||
|
||||
- **来源**: Grill 阶段 Question Pool #5
|
||||
- **查证结果**: 新列 `relevance_level` 和 `dedup_reason` 均为 nullable,不影响现有查询
|
||||
- **结论**: 历史记录保持 null,不需要回填迁移
|
||||
|
||||
### E5: BGE-M3 L2 归一化实测验证
|
||||
|
||||
- **来源**: Specify 阶段 + FullPipelineSmokeTest
|
||||
- **查证结果**:
|
||||
- embeddingBgeM3Works() 测试新增 L2 范数断言
|
||||
- 实测范数 = 1.00000002,误差 < 0.01
|
||||
- 测试通过,BGE-M3 输出确认为 L2 归一化单位向量
|
||||
- **结论**: L2 距离上界 = 2.0 的数学依据成立
|
||||
|
||||
### E6: V010 迁移验证
|
||||
|
||||
- **来源**: Apply 阶段运行时验证
|
||||
- **查证结果**:
|
||||
- Flyway V010 迁移成功执行
|
||||
- `relevance_level` VARCHAR(20) 列可空,已正确写入
|
||||
- `dedup_reason` VARCHAR(32) 列可空,已正确写入
|
||||
- `retrieval_details` JSON 扩展字段(l1_top_similarity、relevance_level、completeness_hint、retrieved_domains、dedup_reason)全部写入
|
||||
- **结论**: 入库可观测性符合设计
|
||||
|
||||
### E7: 数据库数据校验
|
||||
|
||||
- **来源**: Apply 阶段运行时验证
|
||||
- **查证结果**:
|
||||
- session `b66d799e` 共 10 条 lookup_knowledge 调用
|
||||
- id=138: L2=0.383 → similarity=0.8085 → HIGHLY_RELEVANT(符合预期)
|
||||
- id=139-147: 主要为 REFERENCE,doc_retrieved 去重正常触发
|
||||
- retrieved_domains 域追踪:`[infrastructure]` → `[infrastructure, api]` 正常扩展
|
||||
- **结论**: 归一化、行动记忆、去重机制数据层面全部验证通过
|
||||
@@ -0,0 +1,58 @@
|
||||
# Acceptance: chat-verifier-agent
|
||||
|
||||
## Classification
|
||||
|
||||
standard
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Verifier prompt | Done | Strict JSON schema, verdict matrix, fact classifications, and `evidence_refs` are defined. |
|
||||
| VerifierInputHook | Done | Explicit verifier payload replaces raw conversation history. |
|
||||
| ChatService integration | Done | Planner, executor, and verifier are called explicitly with max two rounds. |
|
||||
| Verdict routing | Done | PASS, LOW_CONFID, and REJECT paths are handled in code. |
|
||||
| Trace summary | Done | Evidence summaries include `trace_ref` and `source_invocation_ids`. |
|
||||
| self_evaluation merge | Done | `rule_evaluation` and `verifier_evaluation` are preserved independently. |
|
||||
| Verifier observability | Done | `verifier_evaluation` persists facts, evidence refs, trace summary, rationale, score, and round. |
|
||||
|
||||
## Static Verification
|
||||
|
||||
- [x] OpenSpec artifacts exist: `proposal.md`, `design.md`, `specs/chat-verifier-agent/spec.md`, `tasks.md`, `.committed`.
|
||||
- [x] `change.json` exists and has `metadata.status = committed`.
|
||||
- [x] `.archive-ready` exists.
|
||||
- [x] devflow archive-prep files exist: `brief.md`, `evidence.md`, `decisions.md`, `acceptance.md`.
|
||||
- [x] `devflow/index.md` contains `chat-verifier-agent` with status `archived`.
|
||||
|
||||
## Script Verification
|
||||
|
||||
- [x] `mvn -q -DskipTests compile` passed.
|
||||
|
||||
## Runtime Verification
|
||||
|
||||
- [x] POST `/api/chat` with a complex question returned successfully.
|
||||
- [x] Runtime session `9138f064` showed planner, executor, and verifier execution in logs.
|
||||
- [x] Runtime session `9138f064` wrote `verifier_evaluation.verdict = LOW_CONFID`.
|
||||
- [x] Runtime session `9138f064` wrote `facts_checked[*].evidence_refs`.
|
||||
- [x] Runtime session `9138f064` wrote `tool_trace_summary[*].source_invocation_ids`.
|
||||
- [x] LOW_CONFID final answer included disclaimer and verifier-derived evidence gaps.
|
||||
|
||||
## Unverified
|
||||
|
||||
| Scenario | Reason | Risk | Follow-up |
|
||||
| --- | --- | --- | --- |
|
||||
| PASS runtime path | The exercised complex runtime case produced LOW_CONFID. | Low; PASS routing is simple pass-through after parsed verifier decision. | Add a fixture or deterministic verifier test if this becomes product-critical. |
|
||||
| REJECT runtime path | No forced contradiction case was run after traceability changes. | Medium; REJECT is the safety-critical degraded path. | Add a targeted test with a fabricated claim and evidence contradiction. |
|
||||
| Document-path-level evidence mapping | Current implementation records invocation ids and source document labels, not guaranteed canonical document paths for every retrieval mode. | Low for current audit need; medium for future UI drill-down. | Extend retrieval details with canonical document paths in a later change. |
|
||||
|
||||
## Remaining Risks
|
||||
|
||||
1. Verifier output still depends on model compliance with JSON schema; code falls back to LOW_CONFID on missing or invalid output.
|
||||
2. `AgentLoggingHook` is shared by several agent paths; current changes preserve compile and runtime behavior but should be watched in AiOps flows.
|
||||
3. `SupervisorAgent` construction remains as legacy residue in `ChatService`; runtime orchestration is explicit, but a later cleanup should remove unused supervisor construction.
|
||||
|
||||
## Archive State
|
||||
|
||||
- [x] OpenSpec change is archive-ready.
|
||||
- [x] OpenSpec change has been moved to `openspec/changes/archive/2026-07-03-chat-verifier-agent/`.
|
||||
- [x] Main spec exists at `openspec/specs/chat-verifier-agent/spec.md`.
|
||||
@@ -0,0 +1,36 @@
|
||||
# Brief: chat-verifier-agent
|
||||
|
||||
## Background
|
||||
|
||||
The complex Chat path previously returned Executor answers without a synchronous quality gate. Existing rule scoring was asynchronous and post-hoc, so it could not prevent unsupported answers from reaching users.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Add a Verifier Agent after Executor in the complex chat path.
|
||||
2. Require structured verifier output with `PASS`, `LOW_CONFID`, or `REJECT`.
|
||||
3. Route final user output in code based on verifier verdict.
|
||||
4. Persist verifier results under `diagnosis_session.self_evaluation.verifier_evaluation`.
|
||||
5. Preserve rule scoring under `rule_evaluation`.
|
||||
6. Make verifier decisions traceable to real tool invocations through `evidence_refs` and `source_invocation_ids`.
|
||||
|
||||
## Scope
|
||||
|
||||
- `ChatService`: explicit `planner -> executor -> verifier` orchestration, max two rounds, verdict routing, retry context, verifier persistence.
|
||||
- `VerifierInputHook`: explicit verifier input payload.
|
||||
- `ToolTraceSummaryService`: evidence summary from persisted tool calls.
|
||||
- `VerifierContextHolder`: round-local verifier context.
|
||||
- `SelfEvaluationMergeService`: safe JSON merge for evaluation channels.
|
||||
- `AgentLoggingHook`: concise verifier thought and fuller structured output retention.
|
||||
- `chat-verifier-prompt.md`: verifier contract, verdict matrix, and traceability schema.
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- Verifier does not call tools.
|
||||
- Verifier does not rewrite Executor output.
|
||||
- Single-agent chat path remains outside this change.
|
||||
- No database schema migration is included.
|
||||
- Document-path-level evidence attribution is deferred; current traceability is invocation-level with source document labels.
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/archive/2026-07-03-chat-verifier-agent/`
|
||||
@@ -0,0 +1,123 @@
|
||||
# Decisions: chat-verifier-agent
|
||||
|
||||
## 过程日志
|
||||
|
||||
### Clarify 阶段
|
||||
|
||||
**入口摘要**: 在 Chat 多 Agent 链路中新增 Verifier Agent,作为 Executor 输出后的质量门禁,做事实核查。
|
||||
|
||||
**slug**: `chat-verifier-agent`
|
||||
|
||||
**规模分档**: standard
|
||||
|
||||
### Context 阶段
|
||||
|
||||
**devflow/index.md 使用状态**: 已命中。前序 change `executor-action-memory-relevance`(archived)提供了 Chat 多 Agent 当前链路(Supervisor → Planner → Executor)。
|
||||
|
||||
**不能违反的历史决策**:
|
||||
1. Executor 已有完整的行动记忆和归一化质量等级,Verifier 不需要重复验证检索质量
|
||||
2. Chat Supervisor 的职责是调度,Verifier 作为子 Agent 加入后不改变 Supervisor 的定位
|
||||
3. 已有 evidence_score 做事后评分,Verifier 是事前门禁,两者不冲突
|
||||
|
||||
**需进入 OpenSpec 的上下文点**:
|
||||
1. Verifier 不需要工具调用,只是一个质量核查 Agent
|
||||
2. Verifier 需要访问 Executor 的输出 + 工具调用记录
|
||||
3. Supervisor prompt 需要重写以包含 Verifier 调度规则
|
||||
4. groundedness_score 的阈值需要在代码中定义
|
||||
|
||||
### Grill 阶段 — Question Pool
|
||||
|
||||
| # | 维度 | 问题 | 模式 | 状态 |
|
||||
|---|------|------|------|------|
|
||||
| Q1 | 术语 | evidence_score(事后评分)与 Verifier(事前门禁)职责是否冲突? | evidence-driven | 已解决 |
|
||||
| Q2 | 边界 | Verifier 需要的"工具调用记录"在 SupervisorAgent 中是否自动传递? | evidence-driven | 已解决 |
|
||||
| Q3 | 边界 | LOW_CONFID < 0.5 回调 Planner 后的新输出是否再次走 Verifier?循环上限多少? | user-interview | 已解决 |
|
||||
| Q4 | 验收 | Verifier 判决结果如何可观测?是否写入 agent_step 或 tool_invocation? | user-interview | 已解决 |
|
||||
| Q5 | 验收 | 当前 Supervisor 硬编码 prompt 是否支持多 Agent 路由变更? | evidence-driven | 已解决 |
|
||||
| Q6 | 技术 | Verifier 如何隔离 Executor 的中间推理过程,只看到干净的 query + tool 记录 + 最终答案? | user-interview | 已解决 |
|
||||
| Q7 | 验收 | groundedness_score 阈值(0.5)是否需要配置化? | user-interview | 已解决 |
|
||||
|
||||
### Evidence-driven 结论
|
||||
|
||||
| 结论 | 证据来源 | 是否已汇报用户 |
|
||||
|------|---------|-------------|
|
||||
| evidence_score(异步事后)与 Verifier(同步事前门禁)不冲突 | EvaluationService.java: @Async 注解 | 已汇报 |
|
||||
| SupervisorAgent 自动传递完整对话状态,Verifier 无需额外传递工具记录 | Spring AI Alibaba SupervisorAgent 实现 | 已汇报 |
|
||||
| Supervisor prompt 为字符串字面量,直接修改即可 | ChatService.java:353 .systemPrompt("...") | 已汇报 |
|
||||
|
||||
### User-interview 记录
|
||||
|
||||
| 问题 | 用户原话 | 确认状态 | OpenSpec 回写 |
|
||||
|------|---------|---------|-------------|
|
||||
| Q3: LOW_CONFID < 0.5 回调 Planner 循环上限? | "可以,回调一次" | 已确认 | 已回写 proposal |
|
||||
| Q4: Verifier 判决写入哪里做可观测? | "可以"(写入 diagnosis_session.self_evaluation JSON) | 已确认 | 已回写 proposal |
|
||||
| Q6: Verifier 如何隔离 Executor 中间推理? | "用 MessagesModelHook 过滤 messages" | 已确认 | 已回写 design |
|
||||
| Q7: groundedness_score 阈值是否需要配置化? | "需要配置化" | 已确认 | 已回写 design |
|
||||
|
||||
### Specify 阶段 — Cross-Artifact 对齐检查
|
||||
|
||||
| 上游 → 下游 | 检查内容 | 状态 |
|
||||
|---|---|---|
|
||||
| proposal → design | 范围、约束、关键承诺是否进入 design | 已对齐 |
|
||||
| design → specs | 关键决策、模块地图是否进入 specs | 已对齐 |
|
||||
| specs → tasks | 可观察行为是否被 tasks 覆盖为可执行切片 | 已对齐 |
|
||||
|
||||
**接口影响分级**:
|
||||
- buildChatVerifierAgent() 新增方法 → L1(内部方法,无外部消费者)
|
||||
- VerifierInputHook 类 → L1(内部 Hook,无外部消费者)
|
||||
- Supervisor prompt 重写 → L1(仅影响 Chat 多 Agent 内部调度)
|
||||
- subAgents 列表变更 → L1(Supervisor 内部配置)
|
||||
- verifier.low-confidence-threshold 配置 → L1(新增配置项,不改已有配置)
|
||||
|
||||
### Audit 阶段
|
||||
|
||||
**模块链路**:
|
||||
|
||||
```
|
||||
用户 → Supervisor → Planner(步骤) → Executor(答案+工具记录)
|
||||
│
|
||||
Supervisor 调用 Verifier
|
||||
│
|
||||
[VerifierInputHook BEFORE_MODEL]
|
||||
├─ 保留:system prompt + user query
|
||||
├─ 保留:tool call 记录(输入+返回)
|
||||
├─ 保留:Executor 最终答案
|
||||
└─ 去除:Executor 中间推理、Planner 规划过程
|
||||
│
|
||||
Verifier 判决
|
||||
│
|
||||
┌─── PASS ───→ 直接输出
|
||||
├─── LOW_CONFID≥0.5 → 带声明输出
|
||||
├─── LOW_CONFID<0.5 → 回调 Planner(一次)
|
||||
└─── REJECT → 降级输出
|
||||
│
|
||||
写入 self_evaluation JSON
|
||||
```
|
||||
|
||||
**架构风险评估**(5 句以内):
|
||||
1. Verifier 是轻量 Agent(无工具、无外部依赖),架构风险低。
|
||||
2. MessagesModelHook 纯过滤逻辑,不引入新数据源。
|
||||
3. LOW_CONFID 分级处理 + 回调仅一次的设计,避免无限循环风险。
|
||||
4. REJECT 降级确保编造内容不到达用户。
|
||||
5. 审计结论不影响现有 design/tasks,无需回写。
|
||||
|
||||
### 关键取舍
|
||||
|
||||
- 决策:LOW_CONFID < 0.5 回调 Planner 一次
|
||||
- 原因:给系统一次修正机会,但避免无限循环
|
||||
- 影响:Supervisor prompt 需维护"已回调"状态
|
||||
- 风险接受:用户已确认
|
||||
|
||||
- 决策:Verifier 判决写入 diagnosis_session.self_evaluation JSON
|
||||
- 原因:不改表结构,与 evidence_score 统一可观测体系
|
||||
- 影响:ChatService 后处理需追加 JSON
|
||||
- 风险接受:用户已确认
|
||||
|
||||
### Archive-Ready Update
|
||||
|
||||
- 实现调整:最终运行链路由 `ChatService` 显式调用 `planner -> executor -> verifier`,不再依赖 Supervisor prompt 保证 verifier 被调用。
|
||||
- 可追溯性补充:`tool_trace_summary` 增加 `trace_ref`、`source_invocation_ids`、查询样本、检索层级、相关性等级和来源文档标签。
|
||||
- 可追溯性补充:`facts_checked[*].evidence_refs` 被 prompt 要求、代码解析并持久化。
|
||||
- 验证记录:`mvn -q -DskipTests compile` 通过。
|
||||
- 验证记录:运行会话 `9138f064` 走通 planner、executor、verifier,并持久化 `verifier_evaluation.facts_checked[*].evidence_refs` 与 `tool_trace_summary[*].source_invocation_ids`。
|
||||
- 当前状态:OpenSpec change 已归档到 `openspec/changes/archive/2026-07-03-chat-verifier-agent/`,主规格已同步到 `openspec/specs/chat-verifier-agent/spec.md`。
|
||||
@@ -0,0 +1,52 @@
|
||||
# Evidence: chat-verifier-agent
|
||||
|
||||
## Code Evidence
|
||||
|
||||
### Complex chat path now invokes verifier deterministically
|
||||
|
||||
- File: `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- Evidence: `executeChatComplex` calls planner, executor, then verifier directly through `callAgent(...)`.
|
||||
- Conclusion: runtime no longer depends on prompt-only Supervisor behavior to call verifier.
|
||||
|
||||
### Verifier receives explicit inputs
|
||||
|
||||
- File: `src/main/java/com/superbiz/agent/hook/VerifierInputHook.java`
|
||||
- Evidence: the hook builds a JSON payload with `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
|
||||
- Conclusion: verifier input is stable and does not depend on guessing the last assistant message from raw history.
|
||||
|
||||
### Tool evidence is traceable to persisted invocations
|
||||
|
||||
- File: `src/main/java/com/superbiz/agent/service/ToolTraceSummaryService.java`
|
||||
- Evidence: summaries include `trace_ref`, `source_invocation_ids`, `query_samples`, `retrieval_layers`, `relevance_levels`, and `source_documents`.
|
||||
- Conclusion: verifier facts can be correlated with actual `tool_invocation` rows.
|
||||
|
||||
### Verifier facts preserve evidence references
|
||||
|
||||
- File: `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- Evidence: verifier parsing preserves `facts_checked[*].evidence_refs` and persists `tool_trace_summary` under `verifier_evaluation`.
|
||||
- Conclusion: `self_evaluation` now contains both verifier judgments and the evidence index used to form them.
|
||||
|
||||
### Evaluation channels no longer overwrite each other
|
||||
|
||||
- File: `src/main/java/com/superbiz/agent/service/SelfEvaluationMergeService.java`
|
||||
- Evidence: rule and verifier evaluations are merged into separate keys.
|
||||
- Conclusion: asynchronous rule scoring preserves verifier output.
|
||||
|
||||
### Verifier logging is less noisy
|
||||
|
||||
- File: `src/main/java/com/superbiz/agent/hook/AgentLoggingHook.java`
|
||||
- Evidence: verifier `thought` stores a concise verdict summary, while fuller model output remains available in structured storage.
|
||||
- Conclusion: `agent_step.thought` is no longer a misleading place for full verifier JSON.
|
||||
|
||||
## Runtime Evidence
|
||||
|
||||
- Compile verification passed: `mvn -q -DskipTests compile`.
|
||||
- Runtime session `9138f064` executed `planner -> executor -> verifier`.
|
||||
- Runtime session `9138f064` persisted `verifier_evaluation.facts_checked[*].evidence_refs`.
|
||||
- Runtime session `9138f064` persisted `verifier_evaluation.tool_trace_summary[*].source_invocation_ids`.
|
||||
|
||||
## Design Evidence
|
||||
|
||||
- `LOW_CONFID` returns a fixed disclaimer and verifier-derived gaps.
|
||||
- `REJECT` returns degraded output and does not pass through the raw Executor answer.
|
||||
- `retry_context` is derived from verifier-identified missing evidence facts.
|
||||
@@ -0,0 +1,65 @@
|
||||
# MVP Demo Trace Acceptance
|
||||
|
||||
## Result
|
||||
|
||||
Accepted for implementation scope.
|
||||
|
||||
## Verification
|
||||
|
||||
### Static Verification
|
||||
|
||||
- Command: `mvn -q -DskipTests compile`
|
||||
- Result: passed
|
||||
- Notes: New trace controller, service, DTO, profile, verifier fallback, and test sources compile with the project.
|
||||
|
||||
### Script Verification
|
||||
|
||||
- Command: `mvn -q "-Dtest=DiagnosisTraceServiceTest,ChatServiceSupervisorAgentTest" test`
|
||||
- Result: passed
|
||||
- Notes: Covers successful trace aggregation, missing-session 404 path via `SessionNotFoundException`, low-confidence no-retry behavior, method-tool injection, and verifier fallback when Supervisor skips `chat_verifier`.
|
||||
|
||||
### OpenSpec Verification
|
||||
|
||||
- Command: `openspec validate mvp-demo-trace-acceptance --strict`
|
||||
- Result: passed
|
||||
|
||||
### GitNexus Verification
|
||||
|
||||
- Result: skipped by user decision
|
||||
- Notes: User requested subsequent project flow to bypass GitNexus.
|
||||
|
||||
### Manual / Runtime Verification
|
||||
|
||||
- Steps: Follow `mvp/demo/README.md` with `--spring.profiles.active=mvp-demo`.
|
||||
- Result: passed
|
||||
- Notes:
|
||||
- Session `mvp-demo-payment-timeout-20260703-rerun2` completed as `SUCCESS`.
|
||||
- Chat request returned `code=200`, `success=true`, and the same `sessionId`.
|
||||
- Chat duration was `96316 ms`; persisted session duration was `95028 ms`.
|
||||
- Trace API returned `code=200`, `returnedSteps=13`, `returnedTools=12`, `hasVerifier=true`, and `verifierVerdict=LOW_CONFID`.
|
||||
- Trace agents included `planner,executor,verifier`.
|
||||
- Trace tools included `lookup_knowledge,query_logs,query_metrics`.
|
||||
- Feedback submission returned success, and a follow-up trace query showed `feedback=useful`.
|
||||
- MySQL verification confirmed `agent_step` count `13` with agents `executor,planner,verifier`.
|
||||
- MySQL verification confirmed `tool_invocation` count `12` with tools `lookup_knowledge,query_logs,query_metrics`.
|
||||
|
||||
## Completed Scope
|
||||
|
||||
- Added `GET /api/diagnosis/{sessionId}/trace`.
|
||||
- Added read-only trace aggregation from persisted diagnosis tables.
|
||||
- Added `mvp-demo` profile overlay.
|
||||
- Added payment-timeout demo acceptance documentation.
|
||||
- Added MVP note for interview storytelling.
|
||||
- Added verifier fallback so runtime trace remains complete when Supervisor returns without `verifier_output`.
|
||||
|
||||
## Known Limits
|
||||
|
||||
- `mvp-demo` is not a fully offline mock runtime.
|
||||
- Runtime still depends on available MySQL, Redis, Milvus/Zilliz, model, and embedding configuration.
|
||||
- Sensitive configuration cleanup remains intentionally deferred.
|
||||
- Supervisor can still make inefficient routing choices inside a single round; `ChatService` now invokes `chat_verifier` as a fallback when Supervisor returns without `verifier_output`, so trace completeness is preserved for the MVP demo.
|
||||
|
||||
## Handoff
|
||||
|
||||
- Runtime demo passed with current infrastructure.
|
||||
- OpenSpec archive confirmation: requested by user after successful rerun.
|
||||
@@ -0,0 +1,35 @@
|
||||
# MVP Demo Trace Acceptance Brief
|
||||
|
||||
## Background
|
||||
|
||||
- User goal: make the MVP runnable, observable, and explainable for an Agent Engineer interview.
|
||||
- Current problem: the system can execute diagnosis, but reviewers need a simple way to replay one session from final answer back to agent steps and tool evidence.
|
||||
- Associated OpenSpec: `openspec/changes/mvp-demo-trace-acceptance/`
|
||||
- Devflow scale: standard-light.
|
||||
|
||||
## Scope
|
||||
|
||||
- In scope:
|
||||
- `mvp-demo` Spring profile overlay.
|
||||
- `GET /api/diagnosis/{sessionId}/trace` read-only API.
|
||||
- Trace aggregation DTO/service/controller.
|
||||
- Focused service tests.
|
||||
- Demo and acceptance documentation.
|
||||
- Out of scope:
|
||||
- Sensitive configuration cleanup.
|
||||
- Full offline LLM/vector/database mock runtime.
|
||||
- Database schema migration.
|
||||
- Changes to chat execution, verifier routing, upload, or feedback behavior.
|
||||
- Impact area:
|
||||
- `src/main/java/com/superbiz/agent/controller`
|
||||
- `src/main/java/com/superbiz/agent/service`
|
||||
- `src/main/java/com/superbiz/agent/dto`
|
||||
- `src/main/resources/application-mvp-demo.yml`
|
||||
- `mvp/demo`
|
||||
- `mvp/notes`
|
||||
|
||||
## OpenSpec Alignment
|
||||
|
||||
- proposal coverage: covered
|
||||
- specs coverage: covered
|
||||
- tasks coverage: covered
|
||||
@@ -0,0 +1,87 @@
|
||||
# MVP Demo Trace Acceptance Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: continue the MVP toward a runnable and explainable demo by adding an `mvp-demo` profile, an end-to-end acceptance case, and a trace query API.
|
||||
- Slug: `mvp-demo-trace-acceptance`
|
||||
- Devflow scale: standard-light. The change adds a public read-only API and documentation, but does not alter core chat execution or persistence schemas.
|
||||
|
||||
## Context
|
||||
|
||||
- `devflow/index.md` was checked. Relevant history includes `session-storage`, `confidence-feedback`, `executor-action-memory-relevance`, and `chat-verifier-agent`.
|
||||
- `mvp/notes/agent-engineering-decisions.md` already recommends the next phase as "可复现 MVP Demo", including `mvp-demo` profile, fixed diagnosis case, one-click request, and `GET /api/diagnosis/{sessionId}/trace`.
|
||||
- `mvp/issues/ISS-003-mvp-design-implementation-review.md` identifies test stability, session traceability, verifier evidence chain, upload path, and SupervisorAgent consistency as recent MVP concerns. Security cleanup is intentionally deferred by user decision.
|
||||
|
||||
## Question Pool
|
||||
|
||||
| # | Dimension | Question | Mode | Status |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | Terminology | Should "trace" mean persisted diagnosis execution evidence instead of transient frontend chat history? | evidence-driven | Resolved |
|
||||
| Q2 | Boundary | Should this change modify chat execution or only expose existing persisted evidence? | evidence-driven | Resolved |
|
||||
| Q3 | Acceptance | What proves the MVP flow is end-to-end enough for demo/interview use? | evidence-driven | Resolved |
|
||||
| Q4 | Interface | What is the API impact level for `GET /api/diagnosis/{sessionId}/trace`? | evidence-driven | Resolved |
|
||||
|
||||
## Evidence-driven
|
||||
|
||||
| Conclusion | Evidence Source | Reported To User |
|
||||
|---|---|---|
|
||||
| Trace should aggregate persisted diagnosis evidence, not Redis-only chat history. | `DiagnosisSession`, `AgentStep`, `ToolInvocation` entities and repositories | Reported in progress update |
|
||||
| Core chat execution does not need to change for this slice. | Existing unified chat path and SupervisorAgent commits; requested scope is demo/profile/trace/acceptance | Reported in progress update |
|
||||
| End-to-end acceptance should cover start -> chat -> trace -> feedback. | `ChatController`, `FeedbackController`, traceable session id decision in MVP notes | Reported in progress update |
|
||||
| Trace API is additive L3 because it is a new HTTP API for frontend/demo consumers. | sm-flow interface impact rules | Recorded in OpenSpec design |
|
||||
|
||||
## User-interview
|
||||
|
||||
| Question | User Words | Confirmation | OpenSpec Writeback |
|
||||
|---|---|---|---|
|
||||
| Should security/sensitive config cleanup be included? | "安全问题先不考虑"; "敏感配置先不做" | Confirmed | Non-goal |
|
||||
| Should this be implemented under sm-flow? | "按照 sm-flow 的流程来实现吧" | Confirmed | This change follows sm-flow artifacts |
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Add a new trace API instead of embedding trace details in `/api/chat`.
|
||||
- Reason: Chat execution and observability should stay decoupled.
|
||||
- Impact: Demo can query trace after any successful chat request using the same session id.
|
||||
- Risk accepted: Response shape is new and should be treated as demo-facing contract.
|
||||
|
||||
- Decision: Keep `mvp-demo` profile as configuration overlay, not a fully mocked standalone runtime.
|
||||
- Reason: The current MVP still depends on real DB/Redis/Milvus/LLM for full chat execution; this change avoids inventing a fake runtime that hides integration behavior.
|
||||
- Impact: Demo profile improves repeatability for logs/metrics, while docs remain explicit about required external services.
|
||||
- Risk accepted: End-to-end acceptance may still require valid infrastructure and keys.
|
||||
|
||||
## Cross-Artifact Alignment
|
||||
|
||||
| Upstream -> Downstream | Check | Status |
|
||||
|---|---|---|
|
||||
| brief/prd -> proposal | Goal, scope, non-goals, and acceptance expectation are in proposal | Aligned |
|
||||
| proposal -> design | Scope, constraints, and API impact are in design | Aligned |
|
||||
| design -> specs/tasks | Trace DTO, controller/service, demo profile, and docs are represented | Aligned |
|
||||
| specs -> tasks | Observable behavior is covered by executable tasks | Aligned |
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
- Data path: HTTP trace request -> controller -> trace service -> repositories -> aggregate DTO -> `Result.success`.
|
||||
- The service is read-only and does not mutate diagnosis, step, tool, or feedback state.
|
||||
- No schema change is needed because all required fields already exist in `diagnosis_session`, `agent_step`, and `tool_invocation`.
|
||||
- Main risk is response size for large sessions; MVP mitigates by returning previews already persisted by tools rather than raw external logs.
|
||||
- The additive API is acceptable for MVP because old callers remain unaffected.
|
||||
|
||||
## Pre-apply Research
|
||||
|
||||
- Reference implementations read:
|
||||
- `ChatController` for `/api` controller conventions.
|
||||
- `FeedbackController` for simple API controller shape.
|
||||
- `GlobalExceptionHandler` and `SessionNotFoundException` for 404 handling.
|
||||
- `DiagnosisSessionRepository`, `AgentStepRepository`, `ToolInvocationRepository` for available queries.
|
||||
- `DiagnosisSession`, `AgentStep`, `ToolInvocation` for fields.
|
||||
- Impact analysis:
|
||||
- `DiagnosisSessionRepository`: LOW, direct imports in service/controller paths.
|
||||
- `AgentStepRepository`: HIGH because it participates in chat/AiOps flows. This change only consumes existing query methods and does not modify the repository.
|
||||
- `ToolInvocationRepository`: LOW.
|
||||
|
||||
## Commit Gate
|
||||
|
||||
- OpenSpec proposal/design/specs/tasks exist.
|
||||
- API impact: L3 additive collaboration API, documented in design and spec.
|
||||
- User-confirmed non-goal: sensitive configuration cleanup remains out of scope.
|
||||
- No unresolved user-interview questions remain for this slice.
|
||||
@@ -0,0 +1,25 @@
|
||||
# MVP Demo Trace Acceptance Evidence
|
||||
|
||||
## Evidence
|
||||
|
||||
| Source | Evidence | Conclusion | Reported |
|
||||
|---|---|---|---|
|
||||
| `DiagnosisSessionRepository` | Existing `findBySessionId(String)` query | Trace can locate the session without new repository methods | Yes |
|
||||
| `AgentStepRepository` | Existing `findBySessionIdOrderByStepIndex(String)` query | Agent steps can be returned in execution order | Yes |
|
||||
| `ToolInvocationRepository` | Existing `findBySessionIdOrderByIdAsc(String)` query | Tool evidence can be returned in persisted order | Yes |
|
||||
| `GlobalExceptionHandler` | Handles `SessionNotFoundException` as HTTP 404 with `Result.error(404, ...)` | Missing trace can reuse existing error contract | Yes |
|
||||
| `mvn -q "-Dtest=DiagnosisTraceServiceTest" test` | Command passed | Trace aggregation behavior is covered offline | Yes |
|
||||
| `mvn -q -DskipTests compile` | Command passed | New code compiles with the full project | Yes |
|
||||
| `gitnexus detect-changes --repo SuperBizAgent-java` | Command completed with `No changes detected` and line-ending warnings | Required GitNexus check ran; output likely does not capture newly added files | Yes |
|
||||
|
||||
## Evidence-driven Conclusions
|
||||
|
||||
- Conclusion: No database migration is required.
|
||||
- Evidence: All trace fields are available from existing `diagnosis_session`, `agent_step`, and `tool_invocation` entities.
|
||||
- Risk: Response shape becomes a new API contract.
|
||||
- User confirmation: Not required; additive L3 API recorded in OpenSpec.
|
||||
|
||||
- Conclusion: Trace aggregation can be tested without external infrastructure.
|
||||
- Evidence: `DiagnosisTraceServiceTest` uses mocked repositories and an `ObjectMapper`.
|
||||
- Risk: Runtime integration still depends on configured infrastructure.
|
||||
- User confirmation: Not required; limitation recorded in acceptance docs.
|
||||
@@ -13,6 +13,9 @@
|
||||
- [知识库检索使用指南](architecture/knowledge-retrieval-usage.md) - 文档编写和使用说明 ⭐新增
|
||||
- [会话管理](architecture/session-management.md) - Redis + MySQL 会话管理
|
||||
- [实施规划](architecture/implementation-plan.md) - 分阶段实施计划
|
||||
- [会话级去重与知识域地图](architecture/session-dedup-knowledge-map.md) - 文档级去重 + Planner 知识域地图注入解决 ISS-001 ⭐新增
|
||||
- [证据评分与用户反馈](architecture/confidence-feedback.md) - evidence_score 规则引擎 + feedback API ⭐新增
|
||||
- [行动记忆与检索归一化](architecture/action-memory-relevance.md) - Executor 行动记忆 + 归一化质量等级解决 ISS-002 ⭐新增
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -0,0 +1,275 @@
|
||||
# 行动记忆与检索质量归一化
|
||||
|
||||
Executor 行动记忆 + 归一化质量等级设计,解决 ISS-002 Executor 无约束重复检索问题。
|
||||
|
||||
---
|
||||
|
||||
## 一、问题背景
|
||||
|
||||
ISS-001 修复文档级去重后,Executor 在单次会话中仍调用 `lookup_knowledge` 20+ 次。根因:
|
||||
|
||||
1. **行动记忆缺失**:Executor 不知道自己已检索过哪些域
|
||||
2. **质量信号缺失**:检索结果没有给 LLM 判断"结果够不够"的信号
|
||||
3. **Prompt 缺少合法出口**:原 prompt 要求"所有外部信息都必须调用工具",LLM 不敢停止检索
|
||||
|
||||
---
|
||||
|
||||
## 二、整体架构
|
||||
|
||||
```
|
||||
lookup_knowledge(query)
|
||||
│
|
||||
├─ Step 1: L0 精确匹配(keywords 索引)
|
||||
├─ Step 2: L1 语义检索(Milvus 向量)
|
||||
├─ Step 3: computeRelevance()
|
||||
│ ├─ 归一化:L2 → similarity [0,1]
|
||||
│ └─ 判定:PRECISE / HIGHLY_RELEVANT / REFERENCE
|
||||
├─ Step 4: RetrievedDocTracker 检查
|
||||
│ ├─ 文档级去重 → isDocRetrieved(sessionId, docKey)
|
||||
│ ├─ 域级检查 → isDomainRetrieved(sessionId, domain)
|
||||
│ └─ 记录 → markRetrieved(sessionId, domain, docKey)
|
||||
└─ Step 5: 返回 LookupResult
|
||||
├─ primary / supplement(原始内容,不含分数)
|
||||
├─ relevanceLevel(PRECISE / HIGHLY_RELEVANT / REFERENCE)
|
||||
├─ completenessHint(兜底信号)
|
||||
└─ retrievedDomainsThisSession(行动记忆)
|
||||
```
|
||||
|
||||
### 设计原则
|
||||
|
||||
| 原则 | 说明 |
|
||||
|------|------|
|
||||
| **Agent 边界清晰** | 不给 Executor 注入 knowledge map,Executor 只知道做了什么,不用知道有什么 |
|
||||
| **分数封装** | L0/L1 原始分数不在 LookupResult 中返回 LLM,只在归一化层内部使用 |
|
||||
| **原始分数只入库** | 原始 L2 距离写进 `tool_invocation.retrieval_details` JSON 用于可观测 |
|
||||
| **软约束 + 硬拦截** | Prompt 约束(软)+ 工具层域级去重(硬)两层防御 |
|
||||
|
||||
---
|
||||
|
||||
## 三、归一化质量等级
|
||||
|
||||
### L2 距离归一化
|
||||
|
||||
BGE-M3 输出为 L2 归一化单位向量(实测范数=1.00000002),L2 距离数学硬上界 = 2.0。
|
||||
|
||||
```
|
||||
similarity = 1 - min(l2Score, maxL2Distance) / maxL2Distance
|
||||
```
|
||||
|
||||
| L2 距离 | similarity | 等级 |
|
||||
|---------|-----------|------|
|
||||
| 0.0 | 1.0 | PRECISE |
|
||||
| 0.383 | 0.8085 | HIGHLY_RELEVANT |
|
||||
| 0.5 | 0.75 | HIGHLY_RELEVANT |
|
||||
| 0.6031 | 0.6984 | REFERENCE |
|
||||
| 1.0 | 0.5 | REFERENCE 边界 |
|
||||
| 2.0+ | 0.0 | 不视为有效结果 |
|
||||
|
||||
### 三等级判定
|
||||
|
||||
| 等级 | 条件 | completenessHint | LLM 行为 |
|
||||
|------|------|-----------------|---------|
|
||||
| PRECISE | L0 matchCount == 1 | "知识库中不存在比上述结果更精准的文档" | 直接使用,禁止再检索 |
|
||||
| HIGHLY_RELEVANT | L0 命中 + similarity ≥ 0.75,或仅 L1 similarity ≥ 0.75 | "当前结果已高度相关,继续检索不太可能找到更精准的文档" | 可综合推理,大概率不需要继续查 |
|
||||
| REFERENCE | 其余命中(similarity ≥ 0.5) | "当前结果为相关参考,如需更精准信息请明确缺少的具体维度" | 可参考,如需更精准请指出缺少的维度后定向补充 |
|
||||
|
||||
### 阈值配置
|
||||
|
||||
```yaml
|
||||
retrieval:
|
||||
normalization:
|
||||
max-l2-distance: 2.0 # L2 距离上界
|
||||
highly-relevant-threshold: 0.75 # similarity ≥ 0.75 → HIGHLY_RELEVANT
|
||||
reference-threshold: 0.5 # similarity ≥ 0.5 → REFERENCE
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 四、行动记忆
|
||||
|
||||
### RetrievedDocTracker 数据结构
|
||||
|
||||
```java
|
||||
// 从单层升级为双层:session → domain → filePath 集合
|
||||
ConcurrentHashMap<String, Map<String, Set<String>>> sessionRetrievals;
|
||||
```
|
||||
|
||||
### API
|
||||
|
||||
| 方法 | 作用 |
|
||||
|------|------|
|
||||
| `markRetrieved(sessionId, domain, filePath)` | 记录一次检索 |
|
||||
| `isDocRetrieved(sessionId, filePath)` | 文档级去重 |
|
||||
| `isDomainRetrieved(sessionId, domain)` | 域级检查 |
|
||||
| `getRetrievedDomains(sessionId)` | 获取已检索域列表 |
|
||||
| `clearSession(sessionId)` | 清理会话记录 |
|
||||
|
||||
### LookupResult 返回
|
||||
|
||||
```java
|
||||
LookupResult.builder()
|
||||
.found(true)
|
||||
.primary(primaryResult)
|
||||
.supplement(supplementResult)
|
||||
.relevanceLevel("HIGHLY_RELEVANT") // PRECISE / HIGHLY_RELEVANT / REFERENCE
|
||||
.completenessHint("当前结果已高度相关...") // 兜底信号
|
||||
.retrievedDomainsThisSession(["infrastructure", "api"]) // 行动记忆
|
||||
.message("...")
|
||||
.build();
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 五、Executor Prompt 约束
|
||||
|
||||
### 4 条检索约束
|
||||
|
||||
1. **判断重复**:基于 `retrievedDomainsThisSession` 判断语义重叠
|
||||
2. **重复了怎么办**:禁止换关键词重查;先指缺少的维度,再定向补充
|
||||
3. **合法出口**:"不查全不会被追责,重复检索才会被惩罚"
|
||||
4. **利用质量信号**:PRECISE → 停止;HIGHLY_RELEVANT + 域已检索 → 禁止;REFERENCE → 指出缺少维度
|
||||
|
||||
### 关键变化
|
||||
|
||||
原有 prompt:"所有需要外部信息的地方,都必须调用对应的工具"
|
||||
→ 改为:"需要外部信息时调用工具,但须遵守下方的检索约束"
|
||||
|
||||
---
|
||||
|
||||
## 六、数据库变更
|
||||
|
||||
### V010
|
||||
|
||||
```sql
|
||||
ALTER TABLE tool_invocation
|
||||
ADD COLUMN relevance_level VARCHAR(20) COMMENT 'PRECISE/HIGHLY_RELEVANT/REFERENCE/DEDUPED',
|
||||
ADD COLUMN dedup_reason VARCHAR(32) COMMENT 'doc_retrieved/domain_retrieved/null';
|
||||
```
|
||||
|
||||
### retrieval_details JSON 扩展
|
||||
|
||||
```json
|
||||
{
|
||||
"l0_titles": ["MySQL 数据库连接池配置", "Redis 缓存配置指南"],
|
||||
"l1_scores": [0.383, 0.4502, 0.7011],
|
||||
"l1_top_score": 0.383,
|
||||
"l1_top_similarity": 0.8085,
|
||||
"relevance_level": "HIGHLY_RELEVANT",
|
||||
"completeness_hint": "当前结果已高度相关,继续检索不太可能找到更精准的文档",
|
||||
"retrieved_domains": ["infrastructure"]
|
||||
}
|
||||
```
|
||||
|
||||
扩展字段使用方式:
|
||||
|
||||
| 字段 | 用途 |
|
||||
|------|------|
|
||||
| `l1_top_score` | 原始 L2 距离最小值(可观测性) |
|
||||
| `l1_top_similarity` | 归一化后的相似度 [0,1] |
|
||||
| `relevance_level` | 归一化质量等级 |
|
||||
| `completeness_hint` | 兜底信号 |
|
||||
| `retrieved_domains` | 已检索域列表 |
|
||||
| `dedup_reason` | 去重原因(如有) |
|
||||
|
||||
---
|
||||
|
||||
## 七、使用场景
|
||||
|
||||
### 场景 1:正常检索
|
||||
|
||||
```
|
||||
用户:数据库连接池怎么配置?
|
||||
|
||||
Executor 内部:
|
||||
1. lookup_knowledge("数据库连接池配置")
|
||||
→ relevanceLevel=HIGHLY_RELEVANT (similarity=0.8085)
|
||||
→ completenessHint="当前结果已高度相关..."
|
||||
→ retrievedDomainsThisSession=["infrastructure"]
|
||||
2. 基于已有信息直接回答,不再检索
|
||||
```
|
||||
|
||||
### 场景 2:行动记忆阻止重复
|
||||
|
||||
```
|
||||
Executor 步骤列表:
|
||||
- 查数据库连接池配置
|
||||
- 查 HikariCP 参数
|
||||
- 查连接池耗尽排查
|
||||
|
||||
实际行为:
|
||||
1. lookup("数据库连接池") → relevance=HIGHLY_RELEVANT, domains=["infrastructure"]
|
||||
2. lookup("HikariCP 参数") → retrievedDomainsThisSession=["infrastructure"]
|
||||
LLM 判断:infrastructure 域已检索过,禁止换关键词重查
|
||||
→ 基于已有信息回答,指出缺少的具体维度
|
||||
3. lookup("连接池耗尽") → 同域,被 prompt 约束拦截或工具层去重拦截
|
||||
```
|
||||
|
||||
### 场景 3:PRECISE 精确匹配
|
||||
|
||||
```
|
||||
用户:ERR_TIMEOUT 是什么?
|
||||
|
||||
Executor 内部:
|
||||
1. lookup_knowledge("ERR_TIMEOUT")
|
||||
→ L0 matchCount=1(唯一精确匹配)
|
||||
→ relevanceLevel=PRECISE
|
||||
→ completenessHint="知识库中不存在比上述结果更精准的文档"
|
||||
2. 直接使用,不再检索
|
||||
```
|
||||
|
||||
### 场景 4:REFERENCE + 定向补充
|
||||
|
||||
```
|
||||
用户:如何排查生产故障?
|
||||
|
||||
Executor 内部:
|
||||
1. lookup_knowledge("故障排查")
|
||||
→ relevanceLevel=REFERENCE (similarity=0.6)
|
||||
→ retrievedDomainsThisSession=["troubleshooting"]
|
||||
2. LLM 判断:信息不足,缺少"日志分析"维度的具体步骤
|
||||
3. lookup_knowledge("日志分析步骤")
|
||||
→ 定向补充,不盲目换关键词
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 八、可观测性
|
||||
|
||||
### 查询质量分布
|
||||
|
||||
```sql
|
||||
SELECT relevance_level, COUNT(*) AS cnt
|
||||
FROM tool_invocation
|
||||
WHERE tool_name = 'lookup_knowledge'
|
||||
GROUP BY relevance_level;
|
||||
```
|
||||
|
||||
### 去重原因分布
|
||||
|
||||
```sql
|
||||
SELECT dedup_reason, COUNT(*) AS cnt
|
||||
FROM tool_invocation
|
||||
WHERE tool_name = 'lookup_knowledge'
|
||||
GROUP BY dedup_reason;
|
||||
```
|
||||
|
||||
### 归一化分数分布
|
||||
|
||||
```sql
|
||||
SELECT
|
||||
JSON_EXTRACT(retrieval_details, '$.l1_top_similarity') AS similarity,
|
||||
COUNT(*) AS cnt
|
||||
FROM tool_invocation
|
||||
WHERE tool_name = 'lookup_knowledge'
|
||||
AND retrieval_details IS NOT NULL
|
||||
GROUP BY similarity
|
||||
ORDER BY similarity;
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 九、扩展方向(Phase 2)
|
||||
|
||||
- **域级硬限流**:`isDomainRetrieved` 已就绪,在 LookupKnowledgeTool 入口直接拦截同域调用,不依赖 LLM 遵守 prompt
|
||||
- **DEDUPED 等级**:去重时单独标记为 DEDUPED 等级,与 REFERENCE 区分
|
||||
- **分数反馈调优**:基于 feedback 数据优化归一化阈值
|
||||
@@ -0,0 +1,157 @@
|
||||
# 证据评分与用户反馈架构
|
||||
|
||||
## 一、整体架构
|
||||
|
||||
```
|
||||
用户对话
|
||||
↓
|
||||
ChatService.executeChat / executeChatComplex
|
||||
↓ SUCCESS 后写入 answer,异步触发
|
||||
EvaluationService.evaluate(sessionId, answer)
|
||||
└─ 读取 tool_invocation 事实 → 规则引擎 → 写 selfEvaluation
|
||||
|
||||
用户提交反馈
|
||||
↓
|
||||
POST /api/feedback { sessionId, feedback: "useful" | "not_useful" }
|
||||
↓
|
||||
FeedbackService.submitFeedback
|
||||
├─ 写 DiagnosisSession.feedback
|
||||
├─ useful → CaseLibraryService.createFromSession → 写 case_library
|
||||
└─ not_useful → 仅写 feedback,status 不变
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 二、评分规则(evidence_score)
|
||||
|
||||
### 定位
|
||||
|
||||
`evidence_score` 衡量的是**证据收集充分度**,不是答案准确性。
|
||||
|
||||
- 能证明的:Agent 是否有尝试收集证据、检索是否命中
|
||||
- 不能证明的:答案是否有幻觉、推理是否正确
|
||||
|
||||
### 数据来源
|
||||
|
||||
规则引擎只消费 `tool_invocation` 表的事实记录,不依赖 LLM 判断。
|
||||
|
||||
### 规则定义
|
||||
|
||||
| 规则名 | 条件 | delta |
|
||||
|---|---|---|
|
||||
| `no_tool_call` | 无任何工具调用 | 直接 0 分,不参与加权 |
|
||||
| `execution_failed` | status = FAILED | 直接 0 分,不参与加权 |
|
||||
| `has_successful_tool_call` | 至少 1 次成功调用 | +30 |
|
||||
| `l0_exact_match` | 任意调用有 L0 精确匹配命中 | +35 |
|
||||
| `l1_semantic_match` | 无 L0 命中但有 L1 语义匹配 | +20 |
|
||||
| `retrieval_no_hit` | 有检索调用但无任何命中 | -10 |
|
||||
| `all_tool_calls_failed` | 全部调用失败 | -20 |
|
||||
|
||||
> L0 和 L1 互斥取高优先级(L0 命中时跳过 L1 分支)。
|
||||
|
||||
### selfEvaluation 字段格式
|
||||
|
||||
```json
|
||||
{
|
||||
"evidence_score": 65,
|
||||
"source": "rule",
|
||||
"factors": [
|
||||
{"name": "has_successful_tool_call", "delta": 30, "description": "有成功的工具调用(20次)"},
|
||||
{"name": "l0_exact_match", "delta": 35, "description": "L0 精确匹配命中"}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 说明 |
|
||||
|---|---|
|
||||
| `evidence_score` | 0-100 整数 |
|
||||
| `source` | 当前固定为 `"rule"`;预留 `"llm"` 供后续扩展 |
|
||||
| `factors` | 命中的规则列表,含 name / delta / description |
|
||||
| `llm_opinion` | 预留字段(未实现),LLM 观点叠加时在此扩展 |
|
||||
|
||||
### 已知边界
|
||||
|
||||
- 非检索工具(DateTimeTools、QueryMetricsTools 等)不写 `tool_invocation`,这类 session 的 evidence_score = 0,属于设计边界
|
||||
- 评分为异步写入(`@Async`),失败时 `selfEvaluation` 保持 null,前端需处理 null
|
||||
|
||||
---
|
||||
|
||||
## 三、反馈机制
|
||||
|
||||
### API
|
||||
|
||||
```
|
||||
POST /api/feedback
|
||||
Content-Type: application/json
|
||||
|
||||
{
|
||||
"sessionId": "xxx",
|
||||
"feedback": "useful" | "not_useful"
|
||||
}
|
||||
```
|
||||
|
||||
**响应**
|
||||
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"message": "反馈已记录",
|
||||
"caseId": "uuid 或 null"
|
||||
}
|
||||
```
|
||||
|
||||
### 后端行为
|
||||
|
||||
| feedback 值 | 操作 |
|
||||
|---|---|
|
||||
| `useful` | 写 `DiagnosisSession.feedback = "useful"`,生成 `CaseLibrary` 记录,返回 caseId |
|
||||
| `not_useful` | 写 `DiagnosisSession.feedback = "not_useful"`,status 不变 |
|
||||
| 其他值 | 返回 HTTP 400 |
|
||||
|
||||
### 重要设计决策
|
||||
|
||||
**BAD_CASE 不改 status 字段**
|
||||
|
||||
`status` 表示执行状态(RUNNING/SUCCESS/FAILED),是独立维度,不能被质量标签覆盖。
|
||||
查询 BadCase 使用:`WHERE feedback = 'not_useful'`
|
||||
|
||||
**useful 触发案例沉淀规则**
|
||||
|
||||
| CaseLibrary 字段 | 来源 |
|
||||
|---|---|
|
||||
| caseId | UUID |
|
||||
| diagnosisId | DiagnosisSession.sessionId |
|
||||
| sourceType | AUTO |
|
||||
| faultCategory | GENERAL(暂时,后续人工补充) |
|
||||
| title | query 前 100 字符 |
|
||||
| rootCause / solution | DiagnosisSession.answer(完整答案) |
|
||||
| createdBy | "system" |
|
||||
|
||||
**幂等性**:同一 sessionId 重复提交 useful,返回已有 caseId,不重复插入 case_library。
|
||||
|
||||
---
|
||||
|
||||
## 四、数据库变更
|
||||
|
||||
### V008(新增)
|
||||
|
||||
```sql
|
||||
ALTER TABLE diagnosis_session ADD COLUMN answer LONGTEXT COMMENT 'Agent 返回给用户的完整答案';
|
||||
```
|
||||
|
||||
### diagnosis_session 关键字段
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
|---|---|---|
|
||||
| `answer` | LONGTEXT | Agent 完整回答,useful 案例沉淀的内容来源 |
|
||||
| `self_evaluation` | JSON | 证据评分结果,格式见上 |
|
||||
| `feedback` | VARCHAR(16) | useful / not_useful / null |
|
||||
| `status` | VARCHAR(16) | 执行状态,不受 feedback 影响 |
|
||||
|
||||
---
|
||||
|
||||
## 五、扩展方向(Phase 2)
|
||||
|
||||
- **LLM 观点层**:在 `selfEvaluation` 的 `llm_opinion` 字段叠加 LLM 结构化观点(has_root_cause、has_solution 等),作为独立 factors,不改变现有规则逻辑
|
||||
- **案例结构化字段**:useful 触发时自动提取 faultCategory / errorCode,替代暂时的 GENERAL
|
||||
- **重复召回问题**:Executor Prompt 约束或工具层 session 维度去重(见 [ISS-001](../issues/ISS-001-duplicate-retrieval.md))
|
||||
@@ -0,0 +1,188 @@
|
||||
# 会话级去重与知识域地图
|
||||
|
||||
文档级去重 + 知识域地图注入 Planner,解决 ISS-001 Executor 重复召回同一文档问题。
|
||||
|
||||
---
|
||||
|
||||
## 一、整体架构
|
||||
|
||||
本 change 包含两个独立但互补的部分:
|
||||
|
||||
```
|
||||
Part A: 工具层去重
|
||||
LookupKnowledgeTool
|
||||
├── 维护 ConcurrentHashMap<sessionId, Set<filePath>>(JVM 内)
|
||||
├── 每次检索前过滤已召回文档
|
||||
└── SessionContextHolder.clear() 时同步清理
|
||||
|
||||
Part B: 知识域地图
|
||||
┌─────────────────────────────────────────────┐
|
||||
│ 文档上传 (DocumentManagementService) │
|
||||
│ → LLM 生成 doc.covers + doc.when_to_retrieve │
|
||||
│ → 存入 api_document.metadata │
|
||||
│ → 触发域级重算 (KnowledgeDomainService) │
|
||||
└─────────────────────────────────────────────┘
|
||||
↓
|
||||
┌─────────────────────────────────────────────┐
|
||||
│ 域级聚合 (KnowledgeDomainService) │
|
||||
│ → 读取同域所有文档的 when_to_retrieve │
|
||||
│ → LLM 生成 domain.when_to_retrieve │
|
||||
│ → 存入 knowledge_domain 表 │
|
||||
└─────────────────────────────────────────────┘
|
||||
↓
|
||||
┌─────────────────────────────────────────────┐
|
||||
│ 启动 (KnowledgeIndexService.loadIndex) │
|
||||
│ → 加载 knowledge_domain 表 │
|
||||
│ → 某域无记录则触发域级生成 │
|
||||
└─────────────────────────────────────────────┘
|
||||
↓
|
||||
┌─────────────────────────────────────────────┐
|
||||
│ Planner prompt (ChatService) │
|
||||
│ → 注入 knowledge map(域级) │
|
||||
│ → Planner 做粗粒度检索决策 │
|
||||
└─────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 二、Part A:工具层去重
|
||||
|
||||
### RetrievedDocTracker
|
||||
|
||||
session 级已召回文档追踪组件,将去重责任从 LLM 移交到工具层。
|
||||
|
||||
```java
|
||||
ConcurrentHashMap<String, Set<String>> retrieved
|
||||
key: sessionId
|
||||
value: Set<filePath>
|
||||
```
|
||||
|
||||
| 方法 | 作用 |
|
||||
|------|------|
|
||||
| `isAlreadyRetrieved(sessionId, filePath)` | 检查文档是否已召回 |
|
||||
| `markRetrieved(sessionId, filePath)` | 记录已召回文档 |
|
||||
| `clearSession(sessionId)` | 清理会话记录(SessionContextHolder.clear 触发) |
|
||||
|
||||
### 去重流程
|
||||
|
||||
```
|
||||
lookup_knowledge(query)
|
||||
→ L0 检索 → 命中一批文档
|
||||
→ 遍历结果,过滤 isAlreadyRetrieved=true 的文档
|
||||
→ 剩余文档作为 primary/supplement 返回
|
||||
→ 实际返回的文档调用 markRetrieved
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 三、Part B:知识域地图
|
||||
|
||||
### Frontmatter 新增字段
|
||||
|
||||
文档上传时 LLM 自动生成以下两个字段:
|
||||
|
||||
```yaml
|
||||
covers: ["支付失败排查", "扣款无回调"] # 业务场景标签
|
||||
when_to_retrieve: "用户描述支付失败、超时时" # 文档级检索时机
|
||||
```
|
||||
|
||||
### knowledge_domain 表
|
||||
|
||||
```sql
|
||||
CREATE TABLE knowledge_domain (
|
||||
id BIGINT AUTO_INCREMENT PRIMARY KEY,
|
||||
domain_id VARCHAR(64) NOT NULL UNIQUE,
|
||||
description VARCHAR(256),
|
||||
when_to_retrieve TEXT,
|
||||
document_count INT DEFAULT 0,
|
||||
updated_at DATETIME,
|
||||
created_at DATETIME
|
||||
);
|
||||
```
|
||||
|
||||
### Knowledge Map(注入 Planner 的 YAML)
|
||||
|
||||
```yaml
|
||||
available_knowledge_domains:
|
||||
- domain_id: "payment"
|
||||
description: "支付链路问题排查"
|
||||
when_to_retrieve: "用户问题涉及支付、退款、对账时检索;优先检索一次,勿重复"
|
||||
documents:
|
||||
- title: "支付失败排查手册"
|
||||
covers: ["支付超时", "扣款无回调"]
|
||||
- title: "退款处理指南"
|
||||
covers: ["退款未到账", "退款状态异常"]
|
||||
- domain_id: "infrastructure"
|
||||
...
|
||||
```
|
||||
|
||||
### 注入链路
|
||||
|
||||
```
|
||||
文档上传/删除
|
||||
→ KnowledgeDomainService.onDocumentChange(category)
|
||||
→ 读取同域所有文档的 when_to_retrieve
|
||||
→ LLM 聚合为 domain.when_to_retrieve
|
||||
→ 写入 knowledge_domain 表
|
||||
|
||||
应用启动
|
||||
→ KnowledgeIndexService.loadIndex()
|
||||
→ 加载 knowledge_domain → 无记录则触发聚合
|
||||
→ ChatService.buildChatPlannerAgent() 注入 prompt
|
||||
|
||||
Planner prompt 中包含知识域地图
|
||||
→ Planner 做粗粒度检索决策("查 payment 域")
|
||||
→ Executor 收到步骤后执行具体检索
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 四、关键设计决策
|
||||
|
||||
| 决策 | 方案 | 原因 |
|
||||
|------|------|------|
|
||||
| 域级 when_to_retrieve 存 DB | 持久化 | 避免每次重启调 LLM,文档变更时只重算受影响域 |
|
||||
| 文档级 when_to_retrieve 存 metadata JSON | 沿用现有路径 | 无需新增数据库字段 |
|
||||
| RetrievedDocTracker 独立于 SessionContextHolder | 职责分离 | SessionContextHolder 只持有 sessionId,Tracker 是业务状态 |
|
||||
| Planner 只看域级 | 分层决策 | 文档级 when_to_retrieve 留 Executor 筛选(Phase 2) |
|
||||
| LLM 调用同步执行 | 上传时即时生成 | 接受约 1-2s 延迟,保证数据库和 L0 索引立即一致 |
|
||||
|
||||
---
|
||||
|
||||
## 五、Agent 边界
|
||||
|
||||
```
|
||||
Planner 角色:知道"有什么域"
|
||||
└─ 知识域地图:选定要检索的域(一次规划)
|
||||
|
||||
Executor 角色:知道"做了什么"
|
||||
└─ 行动记忆:域级 + 文档级去重(ISS-002 升级为双层记忆)
|
||||
```
|
||||
|
||||
Part B(知识域地图)只注入 Planner prompt,**不注入 Executor prompt**。Executor 只通过 RetrievedDocTracker 知道自己已检索了哪些文档,不需要知道全局域有哪些。
|
||||
|
||||
---
|
||||
|
||||
## 六、数据库变更
|
||||
|
||||
### V009
|
||||
|
||||
```sql
|
||||
CREATE TABLE knowledge_domain (
|
||||
id BIGINT AUTO_INCREMENT PRIMARY KEY,
|
||||
domain_id VARCHAR(64) NOT NULL UNIQUE,
|
||||
description VARCHAR(256),
|
||||
when_to_retrieve TEXT,
|
||||
document_count INT DEFAULT 0,
|
||||
updated_at DATETIME,
|
||||
created_at DATETIME
|
||||
);
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 七、参考资料
|
||||
|
||||
- **架构文档**:`mvp/architecture/knowledge-retrieval-architecture.md`
|
||||
- **使用指南**:`mvp/architecture/knowledge-retrieval-usage.md`
|
||||
- **OpenSpec**:`openspec/changes/archive/2026-06-30-session-dedup-knowledge-map/`
|
||||
@@ -0,0 +1,94 @@
|
||||
# MVP Demo Runbook
|
||||
|
||||
This demo proves the MVP flow from user question to persisted diagnosis trace.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- MySQL, Redis, Milvus/Zilliz, and LLM/embedding configuration are available through the current project configuration.
|
||||
- Security and secret cleanup are intentionally out of scope for this MVP slice.
|
||||
- The `mvp-demo` profile enables mock Prometheus and CLS providers so log and metric tools can return repeatable evidence.
|
||||
|
||||
## Start
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||
```
|
||||
|
||||
The service listens on:
|
||||
|
||||
```text
|
||||
http://localhost:9900
|
||||
```
|
||||
|
||||
## 1. Run Chat Diagnosis
|
||||
|
||||
```powershell
|
||||
$sessionId = "mvp-demo-payment-timeout-001"
|
||||
$body = @{
|
||||
Id = $sessionId
|
||||
Question = "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/chat" `
|
||||
-ContentType "application/json" `
|
||||
-Body $body
|
||||
```
|
||||
|
||||
Expected result:
|
||||
|
||||
- `data.success` is `true`.
|
||||
- `data.sessionId` equals `mvp-demo-payment-timeout-001`.
|
||||
- `data.answer` contains a diagnosis answer.
|
||||
|
||||
## 2. Query Trace
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace"
|
||||
```
|
||||
|
||||
Expected result:
|
||||
|
||||
- `code` is `200`.
|
||||
- `data.session.sessionId` equals the chat session id.
|
||||
- `data.steps` contains planner/executor/verifier records for complex questions.
|
||||
- `data.toolInvocations` contains evidence tool calls such as `lookup_knowledge`, `query_logs`, or `query_metrics`.
|
||||
- `data.session.selfEvaluation` contains verifier or rule evaluation when available.
|
||||
|
||||
## 3. Submit Feedback
|
||||
|
||||
```powershell
|
||||
$feedback = @{
|
||||
sessionId = $sessionId
|
||||
feedback = "useful"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/feedback" `
|
||||
-ContentType "application/json" `
|
||||
-Body $feedback
|
||||
```
|
||||
|
||||
Expected result:
|
||||
|
||||
- `success` is `true`.
|
||||
- A later trace query shows `data.session.feedback` as `useful`.
|
||||
|
||||
## Demo Story
|
||||
|
||||
The important interview story is:
|
||||
|
||||
```text
|
||||
one session id
|
||||
-> user question
|
||||
-> multi-agent execution
|
||||
-> evidence tools
|
||||
-> verifier/self-evaluation
|
||||
-> final answer
|
||||
-> feedback
|
||||
-> trace API for replay and audit
|
||||
```
|
||||
@@ -0,0 +1,39 @@
|
||||
# Payment Timeout Acceptance Case
|
||||
|
||||
## Goal
|
||||
|
||||
Validate that the MVP can diagnose a payment timeout incident and expose the complete trace for replay.
|
||||
|
||||
## Input
|
||||
|
||||
- Session id: `mvp-demo-payment-timeout-001`
|
||||
- Question: `支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。`
|
||||
- Profile: `mvp-demo`
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
1. Chat returns a successful answer with the same session id.
|
||||
2. Trace API returns session metadata, final answer, ordered agent steps, and ordered tool invocations.
|
||||
3. Trace contains enough evidence to explain which tools were used and whether verifier/self-evaluation was persisted.
|
||||
4. Feedback can be submitted for the same session id.
|
||||
5. A follow-up trace query shows the persisted feedback value.
|
||||
|
||||
## Trace Fields To Inspect
|
||||
|
||||
- `data.session.query`
|
||||
- `data.session.answer`
|
||||
- `data.session.selfEvaluation`
|
||||
- `data.session.feedback`
|
||||
- `data.steps[*].agentName`
|
||||
- `data.steps[*].thought`
|
||||
- `data.toolInvocations[*].toolName`
|
||||
- `data.toolInvocations[*].inputParams`
|
||||
- `data.toolInvocations[*].outputPreview`
|
||||
- `data.toolInvocations[*].retrievalDetails`
|
||||
- `data.summary`
|
||||
|
||||
## Known Limits
|
||||
|
||||
- This case is not a full offline test. It still requires valid infrastructure for chat, persistence, vector search, and model calls.
|
||||
- Mock logs and metrics are enabled by the `mvp-demo` profile to make those evidence tools repeatable.
|
||||
- Sensitive configuration cleanup is deferred by current MVP priority.
|
||||
@@ -0,0 +1,82 @@
|
||||
# ISS-001 Executor 重复召回同一文档
|
||||
|
||||
**状态**:已修复(2026-06-30)
|
||||
**严重程度**:中(影响 token 消耗和上下文质量,不影响功能正确性)
|
||||
**发现时间**:2026-06-30
|
||||
**修复版本**:session-dedup-knowledge-map
|
||||
**架构文档**:[会话级去重与知识域地图](../architecture/session-dedup-knowledge-map.md)
|
||||
|
||||
---
|
||||
|
||||
## 现象
|
||||
|
||||
单次对话中 `lookup_knowledge` 被调用 20 次,其中"故障诊断流程规范"被重复召回约 13 次,多个文档被重复召回 3-6 次。
|
||||
|
||||
```
|
||||
tool_invocation 记录(db8bfa0f):
|
||||
L0 命中"故障诊断流程规范" × 13
|
||||
L0+L1 命中"MySQL 数据库连接池配置" × 5
|
||||
L1 命中性能类故障 × 2
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 根本原因
|
||||
|
||||
**两个层面同时缺失去重机制:**
|
||||
|
||||
1. **工具层无去重**:`LookupKnowledgeTool` 每次独立检索,不感知调用历史,同一查询关键词必然返回同一文档
|
||||
2. **Agent 层无记忆**:Executor Prompt 未要求跟踪已使用文档,LLM 每步倾向于"再确认一下",反复触发相同检索
|
||||
|
||||
**调用链路:**
|
||||
|
||||
```
|
||||
Planner step 0:制定排查计划
|
||||
Executor step 0:检索知识库 → 命中故障诊断流程规范
|
||||
Executor step 1:继续检索 → 又命中故障诊断流程规范(不知道已取过)
|
||||
Executor step 3:继续检索 → 又命中故障诊断流程规范
|
||||
... (重复 13 次)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 影响
|
||||
|
||||
- **Token 浪费**:同一文档内容反复塞入上下文,多 Agent 场景尤为明显
|
||||
- **上下文窗口压缩**:重复内容占用有效 token 空间,可能导致有用信息被截断
|
||||
- **evidence_score 失真**:`tool_call_count` 虚高,规则评分中"成功调用次数"被膨胀
|
||||
|
||||
---
|
||||
|
||||
## 修法方向
|
||||
|
||||
### 方案 A:Prompt 层约束(简单,优先验证)
|
||||
|
||||
在 `chat-executor-prompt.md` 中加规则:
|
||||
|
||||
```
|
||||
已检索过的文档不要重复检索。每次调用 lookup_knowledge 前,
|
||||
先检查对话历史中是否已有该文档的内容,有则直接使用,不再重复调用。
|
||||
```
|
||||
|
||||
优点:不改代码,立即可验证
|
||||
缺点:依赖 LLM 遵守指令,不保证 100% 生效
|
||||
|
||||
### 方案 B:工具层去重(可靠,推荐长期方案)
|
||||
|
||||
`LookupKnowledgeTool` 在 session 维度维护已召回文档 ID 集合,检索结果返回前过滤掉已召回的文档。
|
||||
|
||||
优点:彻底解决,不依赖 LLM
|
||||
缺点:需要改工具代码,需要 session 级状态传递
|
||||
|
||||
### 建议
|
||||
|
||||
MVP 阶段先做**方案 A**验证效果,若重复率明显下降则保留;
|
||||
若 LLM 不稳定遵守,再升级到**方案 B**。
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
- `src/main/resources/prompts/chat-executor-prompt.md`
|
||||
@@ -0,0 +1,87 @@
|
||||
# ISS-002 Executor 无约束重复调用 lookup_knowledge
|
||||
|
||||
**状态**:已修复
|
||||
**严重程度**:中(工具层去重已拦截重复文档,但调用本身仍浪费 token 和耗时)
|
||||
**发现时间**:2026-07-01
|
||||
**修复时间**:2026-07-01
|
||||
**关联**:ISS-001(Part A 已修,Part B 注入范围不足)
|
||||
|
||||
---
|
||||
|
||||
## 现象
|
||||
|
||||
ISS-001 修复后,session 级去重(RetrievedDocTracker)生效,同一文档不再重复召回内容。但 Executor 在单次会话中仍调用 `lookup_knowledge` 20+ 次,大部分被去重拦截返回"已检索过"。
|
||||
|
||||
实测日志(session `7c517329`,2026-07-01 13:53):
|
||||
|
||||
```
|
||||
Executor 调用 lookup_knowledge ~20 次
|
||||
去重拦截 11 次:
|
||||
- infrastructure/mysql-connection-pool.md × 6
|
||||
- api/payment-errors.md × 5
|
||||
有效检索仅 2-3 次(首次命中各域时)
|
||||
```
|
||||
|
||||
Executor 用不同的 query 变体反复查同一个域,因为 LLM 觉得"需要更多细节"。
|
||||
|
||||
---
|
||||
|
||||
## 根本原因
|
||||
|
||||
**knowledge map 和检索约束只注入了 Planner prompt,未注入 Executor prompt。**
|
||||
|
||||
当前注入范围:
|
||||
|
||||
| 组件 | knowledge map | 每域最多一次约束 |
|
||||
|------|:---:|:---:|
|
||||
| Planner prompt | 已注入 | 已注入 |
|
||||
| Executor prompt | **未注入** | **未注入** |
|
||||
|
||||
调用链路:
|
||||
|
||||
```
|
||||
Supervisor → Planner:规划一次,输出"查 infrastructure 域 + api 域"
|
||||
Supervisor → Executor:执行步骤(ReactAgent,自主决定调用工具)
|
||||
Executor step 1:lookup("MySQL 连接池配置") → 命中 infrastructure 域 ✓
|
||||
Executor step 2:lookup("HikariCP 参数调优") → 去重拦截 ✗
|
||||
Executor step 3:lookup("连接池耗尽排查步骤") → 去重拦截 ✗
|
||||
Executor step 4:lookup("支付超时排查") → 命中 api 域 ✓
|
||||
Executor step 5:lookup("ERR_TIMEOUT 错误码") → 去重拦截 ✗
|
||||
...(反复用不同变体查同域)
|
||||
```
|
||||
|
||||
Executor 看不到"每个域只查一次"的约束,也不知道已有哪些域被检索过。
|
||||
|
||||
---
|
||||
|
||||
## 影响
|
||||
|
||||
- **Token 浪费**:每次去重拦截仍需走完 L0+L1 检索流程,再返回"已检索过";LLM 也要处理这个返回信息
|
||||
- **耗时增加**:每次冗余调用约 400-500ms(L0+L1 检索 + 向量查询),20 次冗余调用浪费约 10s
|
||||
- **LLM 行为低效**:Executor 花大量 step 在重复检索上,而不是基于已有信息推理
|
||||
|
||||
---
|
||||
|
||||
## 修法方向
|
||||
|
||||
### 方案 A:Executor prompt 注入 knowledge map + 检索约束
|
||||
|
||||
在 `chat-executor-prompt.md` 或 `buildChatExecutorAgent()` 中:
|
||||
1. 注入 knowledge map(与 Planner 相同的 YAML)
|
||||
2. 添加规则:"每个域最多调用一次 lookup_knowledge;已检索过的域不要再用不同关键词重复检索"
|
||||
|
||||
优点:与 Planner 对齐,LLM 能理解域级边界
|
||||
缺点:仍依赖 LLM 遵守指令(但比纯 Prompt 约束强,因为有 knowledge map 做锚点)
|
||||
|
||||
### 方案 B:工具层硬限制(session + 域级计数)
|
||||
|
||||
在 `RetrievedDocTracker` 中增加域级计数:`ConcurrentHashMap<sessionId, Map<domain, count>>`。
|
||||
当某域检索次数 > 1 时,直接在 `LookupKnowledgeTool` 入口返回"该域已检索过,不允许再次调用"。
|
||||
|
||||
优点:100% 可靠,不依赖 LLM
|
||||
缺点:需改动 RetrievedDocTracker + LookupKnowledgeTool,需要从 filePath 反查 domain
|
||||
|
||||
### 建议
|
||||
|
||||
**先做方案 A**(改动小,与已有 knowledge map 注入逻辑一致),观察效果。
|
||||
如果 LLM 仍不遵守,再升级到方案 B。
|
||||
@@ -0,0 +1,170 @@
|
||||
# ISS-003 MVP 设计与实现 Review 收敛
|
||||
|
||||
**状态**:待规划
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-03
|
||||
**来源**:MVP 版本设计与实现 review
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
当前 MVP 已具备 Chat、Planner/Executor/Verifier、知识检索、诊断会话落库、反馈与 case library 等主线能力,但设计文档、运行时实现和可验证性之间仍存在明显偏差。
|
||||
|
||||
本 issue 用来收敛本次 review 的主要风险,方便后续拆 OpenSpec change 或工程任务。
|
||||
|
||||
---
|
||||
|
||||
## 核心问题
|
||||
|
||||
### P0:敏感配置直接提交到仓库
|
||||
|
||||
`src/main/resources/application.yml` 中包含真实基础设施地址、数据库密码、Redis 密码、Milvus token、LLM API key。
|
||||
|
||||
`src/test/java/com/superbiz/agent/service/SimpleMilvusTest.java` 中也硬编码了 Milvus/Zilliz token。
|
||||
|
||||
**影响**:
|
||||
|
||||
- 密钥泄漏后需要立即轮换。
|
||||
- 合并 worktree 后会扩大泄漏面。
|
||||
- `show-sql: true` 与 DEBUG 日志可能进一步暴露业务数据。
|
||||
|
||||
**建议**:
|
||||
|
||||
- 立即轮换已提交的 token/password/api-key。
|
||||
- 将敏感配置改为环境变量或本地 profile 覆盖。
|
||||
- 提交 `application-example.yml` 或 `.env.example`,不要提交真实值。
|
||||
|
||||
### P1:测试体系不能稳定离线运行
|
||||
|
||||
`mvn test` 编译阶段通过,但 surefire 阶段大量失败,主要原因是测试直接依赖外部 MySQL、Redis、Milvus、LLM/Embedding 服务。
|
||||
|
||||
典型失败:
|
||||
|
||||
- MySQL/Flyway 连接失败导致 repository、Redis、Spring context 测试失败。
|
||||
- Milvus 连接测试出现 `DEADLINE_EXCEEDED`。
|
||||
- 当前环境下 Mockito inline mock maker self-attach 失败。
|
||||
|
||||
**影响**:
|
||||
|
||||
- 无法在合并前获得可靠的回归信号。
|
||||
- 实现变更与环境故障混在一起,问题定位成本高。
|
||||
|
||||
**建议**:
|
||||
|
||||
- 将纯单测、H2/JPA slice、外部集成测试分离。
|
||||
- 用 Maven profile 或 JUnit tag 区分 `unit` / `integration`。
|
||||
- 默认 `mvn test` 只跑不依赖外部服务的测试。
|
||||
|
||||
### P1:会话管理设计与实现不一致
|
||||
|
||||
`mvp/architecture/session-management.md` 设计 Redis 作为主会话存储,带 `session:{session_id}` 和 TTL。
|
||||
|
||||
实际 `/api/chat` 在 `ChatController` 中使用 JVM 内存 `ConcurrentHashMap` 管理历史消息,`RedisSessionManager` 虽然存在但没有接入 controller。
|
||||
|
||||
**影响**:
|
||||
|
||||
- 应用重启后会话历史丢失。
|
||||
- 多实例部署时会话不一致。
|
||||
- Redis TTL 与设计中的生命周期不生效。
|
||||
- 前端 chat session id 与后端 diagnosis session id 存在分裂。
|
||||
|
||||
**建议**:
|
||||
|
||||
- 明确 MVP 阶段是否接受内存会话。
|
||||
- 如果接受,需要同步更新文档并标注限制。
|
||||
- 如果不接受,应将 `ChatController` 接入 `SessionManager`,统一 session id 与 diagnosis session id 的关系。
|
||||
|
||||
### P1:Verifier 证据链仍不完整
|
||||
|
||||
`ToolTraceSummaryService` 期望从 `tool_invocation` 汇总 `lookup_knowledge`、`query_logs`、`query_metrics`、`query_order` 等证据工具。
|
||||
|
||||
当前只有 `LookupKnowledgeTool` 主动写入 `tool_invocation`。`QueryMetricsTools` 和 `QueryLogsTools` 返回 JSON,但没有落库。
|
||||
|
||||
**影响**:
|
||||
|
||||
- verifier 无法稳定审计日志、指标、订单等非知识库工具事实。
|
||||
- `thought` 或模型输出中看起来做了很多推理,但可追溯工具调用证据不足。
|
||||
- 用户侧可观测性仍然偏低。
|
||||
|
||||
**建议**:
|
||||
|
||||
- 抽象统一的 `ToolInvocationRecorder`。
|
||||
- 所有 evidence tool 都必须记录 input、output preview、success、duration、trace id。
|
||||
- verifier 只消费结构化 trace summary,不依赖模型自由文本回忆工具调用。
|
||||
|
||||
### P1:上传文档路径存在重复拼接风险
|
||||
|
||||
`DocumentManagementService.saveToLocal()` 返回的是包含 `knowledge_base` 前缀的本地路径。
|
||||
|
||||
`KnowledgeIndexService.readDocument()` 又执行 `Paths.get(knowledgeBasePath, filePath)`。
|
||||
|
||||
**影响**:
|
||||
|
||||
- 上传文档进入 L0 索引后,命中时读取原文可能拼成 `knowledge_base/knowledge_base/...`。
|
||||
- 这会降低 L0 命中后的答案质量,并造成“命中但读不到原文”的隐性故障。
|
||||
|
||||
**建议**:
|
||||
|
||||
- 统一 `filePath` 语义:要么存相对 `knowledge.base-path` 的路径,要么存绝对路径。
|
||||
- `readDocument()` 对 absolute path、已带 base path 的 relative path 做兼容。
|
||||
- 增加上传文档后 L0 命中并读取原文的回归测试。
|
||||
|
||||
### P2:SupervisorAgent 构建后未使用
|
||||
|
||||
`ChatService.executeChatComplex()` 中创建了 `SupervisorAgent`,但实际仍通过 `callAgent(planner/executor/verifier)` 手写顺序编排。
|
||||
|
||||
**影响**:
|
||||
|
||||
- 代码与设计文档中的 multi-agent 编排表述不一致。
|
||||
- 后续维护者容易误判当前已由 Supervisor 执行调度。
|
||||
|
||||
**建议**:
|
||||
|
||||
- 删除未使用的 `SupervisorAgent` 构建,明确当前是手写编排。
|
||||
- 或真正切到 Spring AI Alibaba SupervisorAgent flow,并补充行为验证。
|
||||
|
||||
### P2:生产安全边界偏弱
|
||||
|
||||
`SessionConfiguration` 使用 `activateDefaultTyping + LaissezFaireSubTypeValidator` 配置 Redis JSON 反序列化。
|
||||
|
||||
`WebMvcConfig` 对所有路径放开 CORS。
|
||||
|
||||
**影响**:
|
||||
|
||||
- Redis 若被非可信写入,存在多态反序列化风险。
|
||||
- CORS 全放开适合本地 MVP,不适合公开环境。
|
||||
|
||||
**建议**:
|
||||
|
||||
- Redis value 使用明确 DTO 类型或受限 subtype validator。
|
||||
- CORS 改为按 profile 配置允许域名。
|
||||
|
||||
---
|
||||
|
||||
## 优先级建议
|
||||
|
||||
1. 先处理敏感配置和密钥轮换,避免合并后扩大泄漏范围。
|
||||
2. 建立可离线运行的单测基线,让默认 `mvn test` 可用于合并门禁。
|
||||
3. 统一 session id 与 session storage,解决前后端、Redis、diagnosis session 的语义分裂。
|
||||
4. 补齐所有 evidence tool 的 `tool_invocation` 落库,提升 verifier 可追溯性。
|
||||
5. 修正上传文档路径语义,并补回归测试。
|
||||
6. 清理或真正启用 `SupervisorAgent`,避免设计和实现长期漂移。
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `src/main/resources/application.yml`
|
||||
- `src/test/java/com/superbiz/agent/service/SimpleMilvusTest.java`
|
||||
- `src/main/java/com/superbiz/agent/controller/ChatController.java`
|
||||
- `src/main/java/com/superbiz/agent/service/session/impl/RedisSessionManager.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ToolTraceSummaryService.java`
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
- `src/main/java/com/superbiz/agent/agent/tool/QueryMetricsTools.java`
|
||||
- `src/main/java/com/superbiz/agent/agent/tool/QueryLogsTools.java`
|
||||
- `src/main/java/com/superbiz/agent/service/DocumentManagementService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- `src/main/java/com/superbiz/agent/config/SessionConfiguration.java`
|
||||
- `src/main/java/com/superbiz/agent/config/WebMvcConfig.java`
|
||||
@@ -0,0 +1,59 @@
|
||||
# ISS-004 Executor 域级检索水位控制(Phase 2)
|
||||
|
||||
**状态**:待规划
|
||||
**严重程度**:低
|
||||
**发现时间**:2026-07-01
|
||||
**关联**:ISS-002(Executor 无约束重复调用 lookup_knowledge)
|
||||
|
||||
---
|
||||
|
||||
## 现象
|
||||
|
||||
ISS-002 修复后,`lookup_knowledge` 调用已经从 20+ 次收敛到约 10 次,但仍存在同一批 domain 之间反复横跳的冗余调用。
|
||||
|
||||
当前文档级去重能阻止重复内容进入上下文,但不能阻止 LLM 继续发起相似检索请求。
|
||||
|
||||
---
|
||||
|
||||
## 根因
|
||||
|
||||
Prompt 软约束依赖 LLM 自觉遵守。在 ReactAgent 自主决策模式下,模型倾向于“再确认一步”,而不是信任已有信息。
|
||||
|
||||
---
|
||||
|
||||
## 影响
|
||||
|
||||
- 不影响核心答案正确性。
|
||||
- 增加每轮检索耗时和 token 消耗。
|
||||
- 长会话中冗余调用会随 session 继续累积。
|
||||
|
||||
---
|
||||
|
||||
## 建议方案
|
||||
|
||||
在代码层增加域级检索水位控制,而不是只依赖 prompt。
|
||||
|
||||
水位指标可以包括:
|
||||
|
||||
- 当前 session 内 `lookup_knowledge` 调用次数。
|
||||
- 当前 session 已检索 domain 数量。
|
||||
- 当前 session token 消耗。
|
||||
- 最近一次检索结果的 `relevanceLevel`。
|
||||
|
||||
决策矩阵示例:
|
||||
|
||||
| 水位 | PRECISE | HIGHLY_RELEVANT | REFERENCE | DEDUPED |
|
||||
|---|---|---|---|---|
|
||||
| 低 | 可继续 | 可继续 | 可定向补充 | 停止 |
|
||||
| 中 | 可继续 | 建议停止 | 可定向补充 | 停止 |
|
||||
| 高 | 停止 | 停止 | 停止 | 停止 |
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `src/main/java/com/superbiz/agent/dto/LookupResult.java`
|
||||
- `src/main/java/com/superbiz/agent/tool/RetrievedDocTracker.java`
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
- `src/main/resources/prompts/chat-executor-prompt.md`
|
||||
- `mvp/architecture/action-memory-relevance.md`
|
||||
@@ -0,0 +1,8 @@
|
||||
# 已知问题记录
|
||||
|
||||
| # | 标题 | 严重程度 | 状态 | 文件 |
|
||||
|---|---|---|---|---|
|
||||
| ISS-001 | Executor 重复召回同一文档 | 中 | 已修复 | [ISS-001-duplicate-retrieval.md](ISS-001-duplicate-retrieval.md) |
|
||||
| ISS-002 | Executor 无约束重复调用 lookup_knowledge | 中 | 已修复 | [ISS-002-executor-unconstrained-lookup.md](ISS-002-executor-unconstrained-lookup.md) |
|
||||
| ISS-003 | MVP 设计与实现 Review 收敛 | 高 | 待规划 | [ISS-003-mvp-design-implementation-review.md](ISS-003-mvp-design-implementation-review.md) |
|
||||
| ISS-004 | Executor 域级检索水位控制(Phase 2) | 低 | 待规划 | [ISS-004-executor-domain-hard-limit.md](ISS-004-executor-domain-hard-limit.md) |
|
||||
@@ -0,0 +1,268 @@
|
||||
# MVP Agent 工程决策记录
|
||||
|
||||
本文记录 MVP 实现过程中已经落地的一些关键修复、取舍和工程判断。目标不是写流水账,而是沉淀面试时可以讲清楚的 Agent 工程思路。
|
||||
|
||||
---
|
||||
|
||||
## 1. 统一流式与非流式 Chat 主链路
|
||||
|
||||
### 背景
|
||||
|
||||
早期 `/api/chat` 和 `/api/chat_stream` 是两条不同实现:
|
||||
|
||||
- 非流式接口会走复杂度判断,并可能进入 Planner / Executor / Verifier 多 Agent 流程。
|
||||
- 流式接口直接创建单个 ReactAgent,然后 `agent.stream()` 输出 token。
|
||||
|
||||
这导致两个接口表面都是 chat,实际能力不一致:流式接口不会进入 verifier、不会沉淀完整诊断链路,也不容易和 `diagnosis_session`、`tool_invocation` 对齐。
|
||||
|
||||
### 决策
|
||||
|
||||
将两个接口统一到同一条核心链路:
|
||||
|
||||
```text
|
||||
getOrCreateSession
|
||||
-> 读取会话历史
|
||||
-> ChatService.executeChatWithStrategy(...)
|
||||
-> 写回会话历史
|
||||
```
|
||||
|
||||
接口差异只保留在传输层:
|
||||
|
||||
- `/api/chat` 返回完整 JSON。
|
||||
- `/api/chat_stream` 通过 SSE 分块发送最终答案。
|
||||
|
||||
### 取舍
|
||||
|
||||
这样会牺牲原来的 token 级实时流式体验,但换来业务行为一致、诊断链路一致、Verifier 和 evidence trace 一致。
|
||||
|
||||
对 MVP 来说,优先保证“同一个问题不因接口不同而进入不同智能链路”,比 token 级流式更重要。
|
||||
|
||||
---
|
||||
|
||||
## 2. 会话 ID 与诊断链路统一
|
||||
|
||||
### 背景
|
||||
|
||||
原实现中:
|
||||
|
||||
- `ChatController` 用前端传入的 `Id` 在 JVM 内存里维护历史消息。
|
||||
- `ChatService` 每次执行又生成新的 8 位 sessionId,作为 `diagnosis_session` 和工具调用追踪 ID。
|
||||
|
||||
这会造成前端会话、后端诊断会话、工具证据链三者分裂。
|
||||
|
||||
### 决策
|
||||
|
||||
将前端 chat session id 作为后端诊断链路的主 session id:
|
||||
|
||||
- Redis `SessionContext` 保存聊天历史。
|
||||
- `diagnosis_session.session_id` 复用同一个 id。
|
||||
- `RunnableConfig.metadata.sessionId` 和 `SessionContextHolder` 也使用同一个 id。
|
||||
- `tool_invocation`、`agent_step`、verifier evaluation 都可按同一 session id 串起来。
|
||||
|
||||
### 企业级意义
|
||||
|
||||
Agent 系统最怕“答得出来但查不清”。统一 session id 后,一次用户请求可以完整追踪:
|
||||
|
||||
```text
|
||||
用户问题 -> Agent 步骤 -> 工具调用 -> Verifier 判断 -> 最终答案 -> 用户反馈
|
||||
```
|
||||
|
||||
这是可观测、可审计、可复盘的基础。
|
||||
|
||||
---
|
||||
|
||||
## 3. 引入统一 ToolInvocationRecorder
|
||||
|
||||
### 背景
|
||||
|
||||
Verifier 需要结构化证据链,但原实现只有 `lookup_knowledge` 主动写入 `tool_invocation`。
|
||||
|
||||
`query_logs`、`query_metrics` 虽然返回 JSON,但没有统一落库,导致 verifier 看不到日志、指标等 evidence tool 的稳定记录。
|
||||
|
||||
### 决策
|
||||
|
||||
新增 `ToolInvocationRecorder`,作为所有 evidence tool 的统一落库入口。
|
||||
|
||||
当前接入:
|
||||
|
||||
- `lookup_knowledge`
|
||||
- `query_logs`
|
||||
- `query_metrics`
|
||||
|
||||
记录字段包括:
|
||||
|
||||
- tool name
|
||||
- input params
|
||||
- output preview
|
||||
- output length
|
||||
- success
|
||||
- error message
|
||||
- duration
|
||||
- trace id / domain details
|
||||
|
||||
### 企业级意义
|
||||
|
||||
这一步把 Agent 从“模型说它查过”推进到“系统能证明它查过”。
|
||||
|
||||
后续 verifier 不应该依赖模型自由文本回忆工具调用,而应该消费结构化 trace summary。
|
||||
|
||||
---
|
||||
|
||||
## 4. Verifier 作为事实约束层
|
||||
|
||||
### 背景
|
||||
|
||||
普通 Agent 很容易在工具调用后直接生成答案,但企业场景更关心:
|
||||
|
||||
- 关键结论有没有证据
|
||||
- 证据是直接证据还是间接支持
|
||||
- 哪些事实缺口需要人工介入
|
||||
- 工具失败时是否诚实降级
|
||||
|
||||
### 决策
|
||||
|
||||
保留 Planner / Executor / Verifier 三角色:
|
||||
|
||||
- Planner 负责拆解问题。
|
||||
- Executor 负责执行查询与形成初稿。
|
||||
- Verifier 负责基于 `tool_trace_summary` 做事实核查。
|
||||
|
||||
Verifier 输出结构化 JSON,包括:
|
||||
|
||||
- verdict
|
||||
- groundedness_score
|
||||
- critical_fact_count
|
||||
- facts_checked
|
||||
- rationale
|
||||
|
||||
### 取舍
|
||||
|
||||
Verifier 会增加一次模型调用成本,但换来可解释性和质量约束。对企业级 Agent 来说,这是值得的。
|
||||
|
||||
---
|
||||
|
||||
## 5. 从手写编排切换到 SupervisorAgent
|
||||
|
||||
### 背景
|
||||
|
||||
之前 `ChatService.executeChatComplex()` 中构建了 `SupervisorAgent`,但实际仍然手写调用:
|
||||
|
||||
```text
|
||||
planner -> executor -> verifier
|
||||
```
|
||||
|
||||
这会造成代码与设计不一致,维护者容易误以为当前已经由 Supervisor 调度。
|
||||
|
||||
### 决策
|
||||
|
||||
复杂问题真正切换到 `SupervisorAgent.invoke(...)`。
|
||||
|
||||
Supervisor 负责路由:
|
||||
|
||||
```text
|
||||
chat_supervisor -> chat_planner
|
||||
chat_supervisor -> chat_executor
|
||||
chat_supervisor -> chat_verifier
|
||||
chat_supervisor -> FINISH
|
||||
```
|
||||
|
||||
外层仍保留:
|
||||
|
||||
- verifier 输出解析
|
||||
- PASS / LOW_CONFID / REJECT 判定
|
||||
- retry context
|
||||
- fallback
|
||||
- evaluation 入库
|
||||
|
||||
### 验证
|
||||
|
||||
新增离线专项测试 `ChatServiceSupervisorAgentTest`,使用 scripted `ChatModel` 验证真实 SupervisorAgent 路由顺序,不依赖真实 LLM、MySQL、Redis。
|
||||
|
||||
### 企业级意义
|
||||
|
||||
这让项目不只是“自己写 if/else 多 Agent”,而是使用框架原生 multi-agent orchestration,同时保留业务层的质量门控。
|
||||
|
||||
---
|
||||
|
||||
## 6. 文档上传路径语义统一
|
||||
|
||||
### 背景
|
||||
|
||||
上传文档时,`DocumentManagementService.saveToLocal()` 返回带 `knowledge_base` 前缀的路径。
|
||||
|
||||
而 `KnowledgeIndexService.readDocument()` 又执行:
|
||||
|
||||
```java
|
||||
Paths.get(knowledgeBasePath, filePath)
|
||||
```
|
||||
|
||||
这可能拼出:
|
||||
|
||||
```text
|
||||
knowledge_base/knowledge_base/...
|
||||
```
|
||||
|
||||
最终表现为 L0 命中文档,但读取原文失败。
|
||||
|
||||
### 决策
|
||||
|
||||
统一路径语义:
|
||||
|
||||
- 新上传文档存相对 `knowledge.base-path` 的路径,例如 `payment/runbook.md`。
|
||||
- `readDocument()` 兼容新旧路径:
|
||||
- 相对路径
|
||||
- 已带 base path 的旧相对路径
|
||||
- 绝对路径
|
||||
|
||||
### 企业级意义
|
||||
|
||||
知识库检索不能只看“命中”,还要保证命中后的内容可读、可引用、可追踪。
|
||||
|
||||
这是 RAG / Agent 系统里很典型的工程细节:检索质量问题不一定来自模型,也可能来自路径、元数据、索引和原文之间的语义不一致。
|
||||
|
||||
---
|
||||
|
||||
## 7. MVP 阶段的优先级取舍
|
||||
|
||||
当前主动暂缓的问题:
|
||||
|
||||
- 敏感配置外置与密钥轮换
|
||||
- CORS / Redis 反序列化安全边界
|
||||
- 默认 `mvn test` 离线化
|
||||
|
||||
原因不是这些不重要,而是当前目标是先跑通并讲清楚 MVP Agent 工程闭环。
|
||||
|
||||
短期优先目标:
|
||||
|
||||
```text
|
||||
可演示 -> 可观测 -> 可验证 -> 可复盘
|
||||
```
|
||||
|
||||
安全和完整测试体系属于企业落地必须项,但可以在 MVP 主链路稳定后作为下一阶段补齐。
|
||||
|
||||
---
|
||||
|
||||
## 8. 后续建议
|
||||
|
||||
下一阶段建议聚焦“可复现 MVP Demo”:
|
||||
|
||||
1. 增加 `local-demo` 或 `mvp-demo` profile。
|
||||
2. 准备固定诊断 case,例如“支付接口超时”。
|
||||
3. 提供一键初始化知识库样例。
|
||||
4. 提供一键触发复杂诊断请求的脚本。
|
||||
5. 增加 trace 查询接口:
|
||||
|
||||
```text
|
||||
GET /api/diagnosis/{sessionId}/trace
|
||||
```
|
||||
|
||||
该接口聚合:
|
||||
|
||||
- diagnosis_session
|
||||
- agent_step
|
||||
- tool_invocation
|
||||
- verifier evaluation
|
||||
- final answer
|
||||
- feedback
|
||||
|
||||
这样 MVP 就能从“功能实现”升级为“企业级 Agent 工程作品”。
|
||||
@@ -0,0 +1,39 @@
|
||||
# MVP Demo Profile 与 Trace 查询接口
|
||||
|
||||
## 背景
|
||||
|
||||
MVP 已经能跑多 Agent 诊断、工具调用、Verifier 和反馈,但对外展示时仍然缺少一个稳定的复盘入口。面试官或评审如果想确认一次 Agent 回答是否可信,不能只看最终答案,还需要看到用户原始问题、Agent 步骤顺序、工具调用证据、Verifier / self-evaluation、最终答案和用户反馈。
|
||||
|
||||
## 决策
|
||||
|
||||
新增 `mvp-demo` profile 和 trace 查询接口:
|
||||
|
||||
```text
|
||||
GET /api/diagnosis/{sessionId}/trace
|
||||
```
|
||||
|
||||
接口聚合:
|
||||
|
||||
- `diagnosis_session`
|
||||
- `agent_step`
|
||||
- `tool_invocation`
|
||||
- `self_evaluation`
|
||||
- `feedback`
|
||||
|
||||
同时在 `mvp/demo` 下沉淀端到端验收 case,把启动、提问、查 trace、提交 feedback 串成一条可演示路径。
|
||||
|
||||
## 取舍
|
||||
|
||||
`mvp-demo` profile 不是完整离线 mock 环境,仍然复用当前真实 DB / Redis / Milvus / LLM 配置,只显式打开日志和指标 mock。原因是当前阶段目标是展示企业级 Agent 工程闭环,不是隐藏真实集成复杂度。
|
||||
|
||||
这让 MVP 的讲述从“我实现了一个聊天接口”升级为:
|
||||
|
||||
```text
|
||||
我实现了一条可执行、可观测、可验收、可复盘的 Agent 诊断链路。
|
||||
```
|
||||
|
||||
## 面试表达
|
||||
|
||||
- 我没有把 trace 塞进 chat 返回值,而是做成独立只读观测接口,保持执行链路和观测链路解耦。
|
||||
- Trace API 复用已经沉淀的 `diagnosis_session`、`agent_step`、`tool_invocation` 三张表,没有引入新的 schema 风险。
|
||||
- Demo profile 只做最小 overlay,让日志和指标工具可重复,保留真实基础设施集成,方便说明 MVP 与生产化之间的差距。
|
||||
@@ -0,0 +1,116 @@
|
||||
# Design: 置信度评分与用户反馈机制
|
||||
|
||||
## 架构约束(来自 devflow)
|
||||
|
||||
- Spring Boot 3.2 + Spring AI Alibaba
|
||||
- JPA ddl-auto=validate,变更走 Flyway
|
||||
- 已有实体:`DiagnosisSession`(含 selfEvaluation JSON、feedback VARCHAR)、`CaseLibrary`
|
||||
- **新增字段**:`DiagnosisSession.answer TEXT`,存储返回给用户的完整答案,Flyway V008 迁移
|
||||
- 已有 Repository:`DiagnosisSessionRepository`、`CaseLibraryRepository`
|
||||
- 当前主流程入口:`ChatService.executeChat`(非流式)、`executeChatComplex`(多 Agent)
|
||||
|
||||
## 模块链路
|
||||
|
||||
```
|
||||
用户对话
|
||||
↓
|
||||
ChatService.executeChat / executeChatComplex
|
||||
↓ SUCCESS 后异步
|
||||
EvaluationService.evaluate(sessionId, answer, steps)
|
||||
├─ LLM 自评 → 写 selfEvaluation(含 confidence + reasoning)
|
||||
└─ 规则兜底(LLM 失败时)→ 写 selfEvaluation(含 source: "rule")
|
||||
|
||||
用户提交反馈
|
||||
↓
|
||||
POST /api/feedback { sessionId, feedback }
|
||||
↓
|
||||
FeedbackService.submitFeedback(sessionId, feedback)
|
||||
├─ 写 DiagnosisSession.feedback
|
||||
├─ feedback=useful → 写 CaseLibrary
|
||||
└─ feedback=not_useful → 更新 status=BAD_CASE
|
||||
```
|
||||
|
||||
## 数据结构定义
|
||||
|
||||
### DiagnosisSession.selfEvaluation(JSON 字符串)
|
||||
|
||||
```json
|
||||
{
|
||||
"evidence_score": 65,
|
||||
"source": "rule",
|
||||
"factors": [
|
||||
{"name": "has_successful_tool_call", "delta": 30, "description": "有成功的工具调用(2次)"},
|
||||
{"name": "l1_semantic_match", "delta": 20, "description": "L1 语义匹配命中"}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
字段说明:
|
||||
- `evidence_score`:0-100,衡量证据收集充分度(非答案准确性)
|
||||
- `source`:评分来源,当前固定为 `"rule"`;预留 `"llm"` 供后续 LLM 观点叠加
|
||||
- `factors`:命中的规则因子列表,每项含 name / delta / description,可直接用于分析
|
||||
- `llm_opinion`:预留字段,LLM 观点叠加时扩展此处,不改变现有规则逻辑
|
||||
|
||||
### FeedbackRequest(新 DTO)
|
||||
|
||||
```java
|
||||
public class FeedbackRequest {
|
||||
String sessionId; // 必填
|
||||
String feedback; // "useful" | "not_useful"
|
||||
}
|
||||
```
|
||||
|
||||
### FeedbackResponse(新 DTO)
|
||||
|
||||
```java
|
||||
public class FeedbackResponse {
|
||||
boolean success;
|
||||
String message;
|
||||
String caseId; // useful 时返回生成的 case_id,否则 null
|
||||
}
|
||||
```
|
||||
|
||||
### CaseLibrary 生成规则(useful 时)
|
||||
|
||||
| CaseLibrary 字段 | 来源 |
|
||||
|---|---|
|
||||
| caseId | UUID |
|
||||
| diagnosisId | DiagnosisSession.sessionId |
|
||||
| sourceType | SourceType.AUTO |
|
||||
| faultCategory | FaultCategory.GENERAL(暂时) |
|
||||
| title | DiagnosisSession.query 前 100 字符 |
|
||||
| rootCause | DiagnosisSession.answer(完整答案,不截断) |
|
||||
| solution | DiagnosisSession.answer(同上) |
|
||||
| createdBy | "system" |
|
||||
|
||||
## 关键技术决策
|
||||
|
||||
### 决策 1:LLM 自评异步执行
|
||||
|
||||
置信度计算在 Agent 主流程结束后异步进行(`@Async` + Spring 线程池),不阻塞用户响应。
|
||||
原因:LLM 自评耗时 1-3 秒,主流程不应等待。
|
||||
|
||||
### 决策 2:置信度规则兜底参数
|
||||
|
||||
```
|
||||
基础分:60
|
||||
工具调用加分:toolCallCount × 5,上限 +20
|
||||
步数少加分:stepCount <= 3 → +10
|
||||
status=FAILED → 直接 0
|
||||
```
|
||||
|
||||
### 决策 3:BAD_CASE 用 status 字段而非新字段
|
||||
|
||||
`DiagnosisSession.status` 已有 PENDING/RUNNING/SUCCESS/FAILED,扩展为允许包含 BAD_CASE。
|
||||
该字段是 VARCHAR 16,直接存字符串,无需枚举类(Java 端用常量控制)。
|
||||
|
||||
### 决策 4:案例内容提取策略
|
||||
|
||||
useful 时,`rootCause` 和 `solution` 从 `AgentStepRepository.findBySessionIdOrderByStepIndex` 的最后一步 `thought` 字段提取。
|
||||
如果 thought 为空,则用 DiagnosisSession.query + "(自动提取失败,请人工补充)" 占位。
|
||||
|
||||
## 接口影响等级
|
||||
|
||||
- `POST /api/feedback`:新增接口,L2(内部,前端新消费)
|
||||
- `ChatService.executeChat`:新增异步后置调用,不改返回值,L1
|
||||
- `DiagnosisSession.status` 增加 BAD_CASE 值:原调用方只读不写此字段,L2
|
||||
@@ -0,0 +1,61 @@
|
||||
# Proposal: 置信度评分与用户反馈机制
|
||||
|
||||
## 问题
|
||||
|
||||
DiagnosisSession 已预留 `selfEvaluation`(JSON)和 `feedback`(VARCHAR 16)两个字段,但目前完全为空——Agent 完成对话后不计算置信度,也没有接收用户反馈的 API,无法支撑报告质量评估和 BadCase 追踪。
|
||||
|
||||
## 建议方案
|
||||
|
||||
### 置信度评分(双轨)
|
||||
|
||||
**主轨:LLM 自评**
|
||||
- 在 ChatService 的 `executeChat` 流程结束后,追加一次轻量 LLM 调用(EvaluationService),
|
||||
将 Agent 的最终答案 + 步骤摘要传给模型,要求输出 `{"confidence": 0-100, "reasoning": "..."}` JSON。
|
||||
- 结果写入 `DiagnosisSession.selfEvaluation`。
|
||||
|
||||
**兜底轨:规则计算**
|
||||
- 若 LLM 自评失败(超时/解析失败),用规则计算:
|
||||
- 基础分 60
|
||||
- 工具调用数 > 0 每次 +5(上限 +20)
|
||||
- 步数 <= 3 额外 +10
|
||||
- 状态为 FAILED 直接 0
|
||||
- 兜底结果同样写入 `selfEvaluation`,并附 `"source": "rule"` 标记。
|
||||
|
||||
### 用户反馈 API
|
||||
|
||||
新增 `POST /api/feedback`,接收:
|
||||
```json
|
||||
{ "sessionId": "xxx", "feedback": "useful" | "not_useful" }
|
||||
```
|
||||
后端操作:
|
||||
1. 写入 `DiagnosisSession.feedback`。
|
||||
2. 若 `feedback = "not_useful"`,将 `status` 更新为 `BAD_CASE`(需要在 status 枚举扩展此值)。
|
||||
3. 若 `feedback = "useful"`,写入一条 `CaseLibrary` 记录(从 session 提取 query/answer)。
|
||||
|
||||
## 范围
|
||||
|
||||
- 新建 `EvaluationService`(置信度计算)
|
||||
- 新建 `FeedbackService`(反馈处理)
|
||||
- 新增 `POST /api/feedback` 接口(在 ChatController 或新 FeedbackController)
|
||||
- 改造 `ChatService.executeChat` 在 SUCCESS 后调用 EvaluationService
|
||||
- `DiagnosisSession.status` 枚举扩展 `BAD_CASE` 值
|
||||
- Flyway 迁移:`diagnosis_session.status` 列注释更新(不改类型,字段已存在)
|
||||
- 无需新建数据库表
|
||||
|
||||
## 非目标
|
||||
|
||||
- 不实现 Verifier Agent 完整链路(只做轻量自评,不是多 Agent 编排)
|
||||
- 不实现案例沉淀的复杂结构化字段(CaseLibrary 的 faultCategory/errorCode 等填 GENERAL/null)
|
||||
- 不实现 BadCase 的自动分析或 Prompt 优化流程
|
||||
|
||||
## 关键约束(来自 devflow)
|
||||
|
||||
- JPA ddl-auto = validate,表结构变更必须走 Flyway 迁移,但本次无需加新列
|
||||
- `DiagnosisSession.selfEvaluation` 已声明为 JSON 类型,直接用 String 写入
|
||||
- `CaseLibrary.faultCategory` 是枚举,默认填 GENERAL
|
||||
- `SourceType.AUTO` 表示系统自动生成
|
||||
|
||||
## 风险
|
||||
|
||||
- LLM 自评 prompt 质量影响分数可信度,需要在 reasoning 字段记录依据
|
||||
- BAD_CASE status 与现有 PENDING/RUNNING/SUCCESS/FAILED 并存,需确认 UI 是否受影响
|
||||
@@ -0,0 +1,44 @@
|
||||
# Functional Spec: 置信度评分与用户反馈机制
|
||||
|
||||
## REQ-1:置信度 LLM 自评
|
||||
|
||||
- Agent 对话(executeChat / executeChatComplex)成功后,异步调用 EvaluationService。
|
||||
- EvaluationService 构造 Prompt,调用 ChatModel,要求输出纯 JSON:`{"confidence": 0-100, "reasoning": "...", "source": "llm"}`。
|
||||
- 若 JSON 解析成功,写入 `DiagnosisSession.selfEvaluation`。
|
||||
- 若调用失败或解析失败,转入规则兜底(REQ-2)。
|
||||
- 验收:对话结束后数秒内,DB `diagnosis_session.self_evaluation` 非 null,且 `source` 字段存在。
|
||||
|
||||
## REQ-2:置信度规则兜底
|
||||
|
||||
- 触发条件:LLM 自评失败(任何异常)。
|
||||
- 规则:基础分 60 + toolCallCount×5(上限+20)+ (stepCount<=3 ? +10 : 0),status=FAILED 则直接 0。
|
||||
- 结果写入 `selfEvaluation`,含 `"source": "rule"`。
|
||||
- 验收:LLM 自评失败时,self_evaluation 仍有值(非 null),且 source=rule。
|
||||
|
||||
## REQ-3:反馈接收 API
|
||||
|
||||
- 接口:`POST /api/feedback`
|
||||
- 入参:`{ "sessionId": "xxx", "feedback": "useful" | "not_useful" }`
|
||||
- 出参:`{ "success": true/false, "message": "...", "caseId": "uuid 或 null" }`
|
||||
- 校验:sessionId 不能为空;feedback 只能是 useful 或 not_useful,否则返回 400。
|
||||
- 验收:接口返回 200,DB 对应行 feedback 字段有值。
|
||||
|
||||
## REQ-4:useful → 案例沉淀
|
||||
|
||||
- 触发条件:feedback = "useful"。
|
||||
- 操作:在 case_library 插入一条记录,caseId=UUID,sourceType=AUTO,faultCategory=GENERAL,
|
||||
title=query 前 100 字,rootCause/solution 来自最后一步 agent_step.thought。
|
||||
- 响应中返回 caseId。
|
||||
- 验收:提交 useful 后,case_library 表新增一行,diagnosis_id = sessionId。
|
||||
|
||||
## REQ-5:not_useful → BAD_CASE 标记
|
||||
|
||||
- 触发条件:feedback = "not_useful"。
|
||||
- 操作:feedback 字段本身即为标记,不修改 status 字段(status 保持执行状态语义)。
|
||||
- 查询 BadCase 使用:`WHERE feedback = 'not_useful'`。
|
||||
- 验收:提交 not_useful 后,DB diagnosis_session.feedback = "not_useful",status 不变。
|
||||
|
||||
## REQ-6:幂等性
|
||||
|
||||
- 同一 sessionId 重复提交 feedback,覆盖写入(不报错,不重复创建 CaseLibrary)。
|
||||
- 已有 case_library 记录时(diagnosisId 已存在),跳过插入并返回已有 caseId。
|
||||
@@ -0,0 +1,133 @@
|
||||
# Tasks: 置信度评分与用户反馈机制
|
||||
|
||||
## T0:Flyway 迁移 + DiagnosisSession 实体加字段
|
||||
|
||||
**文件**:
|
||||
- `src/main/resources/db/migration/V008__add_answer_to_diagnosis_session.sql`(新建)
|
||||
- `src/main/java/com/superbiz/agent/domain/entity/DiagnosisSession.java`(加字段)
|
||||
|
||||
**迁移脚本**:
|
||||
```sql
|
||||
ALTER TABLE diagnosis_session ADD COLUMN answer LONGTEXT COMMENT 'Agent 返回给用户的完整答案';
|
||||
```
|
||||
|
||||
**实体**:在 `DiagnosisSession` 加:
|
||||
```java
|
||||
@Column(name = "answer", columnDefinition = "LONGTEXT")
|
||||
private String answer;
|
||||
```
|
||||
|
||||
**验收标准**:应用启动不报 schema validation 错误;`diagnosis_session` 表有 answer 列
|
||||
|
||||
---
|
||||
|
||||
— LLM 自评 + 规则兜底
|
||||
|
||||
**文件**:`src/main/java/com/superbiz/agent/service/EvaluationService.java`
|
||||
|
||||
**实现**:
|
||||
- `@Service @Async` 标注
|
||||
- `evaluate(String sessionId, String answer)` 方法:
|
||||
1. 从 `DiagnosisSessionRepository` 加载 session(含 stepCount、toolCallCount、status)
|
||||
2. 调用 ChatModel 做 LLM 自评,Prompt 见下
|
||||
3. 解析 JSON → 写入 `selfEvaluation`
|
||||
4. 失败时走规则兜底
|
||||
- 规则兜底逻辑:`computeRuleScore(session)` → 返回 JSON 字符串
|
||||
|
||||
**LLM 自评 Prompt(系统提示)**:
|
||||
```
|
||||
你是一个 AI 回答质量评估器。
|
||||
请根据以下信息,评估这次 AI 回答的置信度(0-100分):
|
||||
- 用户原始问题:{query}
|
||||
- AI 的回答:{answer}
|
||||
- 工具调用次数:{toolCallCount}
|
||||
- 推理步数:{stepCount}
|
||||
|
||||
只返回一个 JSON,格式如下,不要输出任何其他内容:
|
||||
{"confidence": <0-100的整数>, "reasoning": "<评估依据,50字以内>"}
|
||||
```
|
||||
|
||||
**验收标准**:
|
||||
- LLM 正常时:DB selfEvaluation 包含 confidence 和 reasoning,source = "llm"
|
||||
- LLM 失败时:DB selfEvaluation 包含 confidence 和 source = "rule"
|
||||
|
||||
---
|
||||
|
||||
## T2:ChatService 后置调用 EvaluationService
|
||||
|
||||
**文件**:`src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
|
||||
**实现**:
|
||||
- 在 `executeChat` 的 `session.setStatus("SUCCESS")` 之后,追加 `session.setAnswer(answer)` 写入完整答案,再注入 EvaluationService 调用 `evaluate(sessionId, answer)`
|
||||
- 在 `executeChatComplex` 的 SUCCESS 分支同样补充 `session.setAnswer(answer)`
|
||||
- 注意:EvaluationService 是 @Async,调用方不等待返回值
|
||||
|
||||
**验收标准**:发送一次 chat 请求后,数秒内 DB self_evaluation 非 null
|
||||
|
||||
---
|
||||
|
||||
## T3:FeedbackController + FeedbackService
|
||||
|
||||
**文件**:
|
||||
- `src/main/java/com/superbiz/agent/controller/FeedbackController.java`(新建)
|
||||
- `src/main/java/com/superbiz/agent/service/FeedbackService.java`(新建)
|
||||
- `src/main/java/com/superbiz/agent/dto/FeedbackRequest.java`(新建)
|
||||
- `src/main/java/com/superbiz/agent/dto/FeedbackResponse.java`(新建)
|
||||
|
||||
**FeedbackService.submitFeedback(sessionId, feedback)**:
|
||||
1. 加载 session,sessionId 不存在抛异常
|
||||
2. 校验 feedback 值(useful/not_useful)
|
||||
3. 更新 `DiagnosisSession.feedback`
|
||||
4. if useful:调用 `CaseLibraryService.createFromSession(session)`
|
||||
5. if not_useful:更新 `DiagnosisSession.status = "BAD_CASE"`
|
||||
6. 保存 session
|
||||
7. 返回 FeedbackResponse
|
||||
|
||||
**幂等逻辑(useful 重复提交)**:
|
||||
- 调用 `CaseLibraryRepository.findByDiagnosisId(sessionId)` 检查
|
||||
- 已存在则返回已有 caseId,不重复插入
|
||||
|
||||
**FeedbackController**:
|
||||
```
|
||||
POST /api/feedback
|
||||
@RequestBody FeedbackRequest
|
||||
@ResponseBody FeedbackResponse
|
||||
```
|
||||
|
||||
**验收标准**:
|
||||
- useful:返回 200,feedback 字段有值,case_library 新增一行
|
||||
- not_useful:返回 200,status = BAD_CASE
|
||||
- 非法 feedback 值:返回 400
|
||||
|
||||
---
|
||||
|
||||
## T4:CaseLibraryService — createFromSession
|
||||
|
||||
**文件**:`src/main/java/com/superbiz/agent/service/CaseLibraryService.java`(新建)
|
||||
|
||||
**实现**:
|
||||
- `createFromSession(DiagnosisSession session)` → `CaseLibrary`
|
||||
- 直接从 `session.getAnswer()` 取完整答案
|
||||
- answer 为空时用占位文本 `query + "\n(自动提取失败,请人工补充)"`
|
||||
- 填写 CaseLibrary 各字段,save 后返回 caseId
|
||||
|
||||
**验收标准**:case_library 行的 diagnosis_id = sessionId,root_cause 非空
|
||||
|
||||
---
|
||||
|
||||
## T5:Spring @Async 配置
|
||||
|
||||
**文件**:检查项目是否已有 `@EnableAsync`,若无则在 `SessionConfiguration` 或新建 `AsyncConfig` 中添加
|
||||
|
||||
**验收标准**:EvaluationService 中 @Async 方法可被正确调度(不抛 bean 配置错误)
|
||||
|
||||
---
|
||||
|
||||
## T6:集成验证
|
||||
|
||||
验证步骤:
|
||||
1. 启动服务,POST /api/chat,发送一条问题
|
||||
2. 查 `diagnosis_session` 表,确认 self_evaluation 有值
|
||||
3. POST /api/feedback `{"sessionId": "xxx", "feedback": "useful"}`,确认 case_library 新增
|
||||
4. POST /api/feedback `{"sessionId": "yyy", "feedback": "not_useful"}`,确认 status = BAD_CASE
|
||||
5. 重复步骤 3,确认不重复创建 case_library
|
||||
@@ -0,0 +1 @@
|
||||
committed
|
||||
@@ -0,0 +1,157 @@
|
||||
# Design: session-dedup-knowledge-map
|
||||
|
||||
## 1. 整体架构
|
||||
|
||||
本 change 包含两个独立但互补的部分:
|
||||
|
||||
```
|
||||
Part A: 工具层去重
|
||||
LookupKnowledgeTool
|
||||
├── 维护 ConcurrentHashMap<sessionId, Set<filePath>>(JVM 内)
|
||||
├── 每次检索前过滤已召回文档
|
||||
└── SessionContextHolder.clear() 时同步清理
|
||||
|
||||
Part B: 知识图谱
|
||||
┌─────────────────────────────────────────────┐
|
||||
│ 文档上传 (DocumentManagementService) │
|
||||
│ → LLM 生成 doc.covers + doc.when_to_retrieve │
|
||||
│ → 存入 api_document.metadata │
|
||||
│ → 触发域级重算 (KnowledgeDomainService) │
|
||||
└─────────────────────────────────────────────┘
|
||||
↓
|
||||
┌─────────────────────────────────────────────┐
|
||||
│ 域级聚合 (KnowledgeDomainService) │
|
||||
│ → 读取同域所有文档的 when_to_retrieve │
|
||||
│ → LLM 生成 domain.when_to_retrieve │
|
||||
│ → 存入 knowledge_domain 表 │
|
||||
└─────────────────────────────────────────────┘
|
||||
↓
|
||||
┌─────────────────────────────────────────────┐
|
||||
│ 启动 (KnowledgeIndexService.loadIndex) │
|
||||
│ → 加载 knowledge_domain 表 │
|
||||
│ → 某域无记录则触发域级生成 │
|
||||
└─────────────────────────────────────────────┘
|
||||
↓
|
||||
┌─────────────────────────────────────────────┐
|
||||
│ Planner prompt (ChatService) │
|
||||
│ → 注入 knowledge map(域级) │
|
||||
│ → Planner 做粗粒度检索决策 │
|
||||
└─────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. 数据结构定义
|
||||
|
||||
### 2.1 Frontmatter 新增字段
|
||||
|
||||
```yaml
|
||||
# 新增两个字段,其余不变
|
||||
covers: ["支付失败排查", "扣款无回调"] # List<String>:业务场景标签,Planner 决策用
|
||||
when_to_retrieve: "用户描述支付失败、超时时" # String:文档级检索时机,LLM 上传时生成
|
||||
```
|
||||
|
||||
对应 `Frontmatter.java` 新增两个字段:
|
||||
- `List<String> covers`
|
||||
- `String whenToRetrieve`
|
||||
|
||||
对应 `KnowledgeEntry.java` 新增两个字段(同上)。
|
||||
|
||||
### 2.2 knowledge_domain 表(新表)
|
||||
|
||||
```sql
|
||||
CREATE TABLE knowledge_domain (
|
||||
id BIGINT AUTO_INCREMENT PRIMARY KEY,
|
||||
domain_id VARCHAR(64) NOT NULL UNIQUE, -- category 值,如 "payment"
|
||||
description VARCHAR(256), -- 域描述(聚合自文档 summary)
|
||||
when_to_retrieve TEXT, -- 域级检索时机(LLM 生成)
|
||||
document_count INT DEFAULT 0, -- 该域当前文档数
|
||||
updated_at DATETIME,
|
||||
created_at DATETIME
|
||||
);
|
||||
```
|
||||
|
||||
### 2.3 knowledge map 结构(注入 Planner 的 YAML 文本)
|
||||
|
||||
```yaml
|
||||
available_knowledge_domains:
|
||||
- domain_id: "payment"
|
||||
description: "支付链路问题排查"
|
||||
when_to_retrieve: "用户问题涉及支付、退款、对账时检索;优先检索一次,勿重复"
|
||||
documents:
|
||||
- title: "支付失败排查手册"
|
||||
covers: ["支付超时", "扣款无回调"]
|
||||
- title: "退款处理指南"
|
||||
covers: ["退款未到账", "退款状态异常"]
|
||||
- domain_id: "infrastructure"
|
||||
...
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. 新增组件
|
||||
|
||||
### 3.1 KnowledgeDomainService(新类)
|
||||
|
||||
职责:域级聚合与存储
|
||||
|
||||
```
|
||||
buildDomainSummary(category)
|
||||
→ 读取同域所有 KnowledgeEntry(含 when_to_retrieve)
|
||||
→ 拼装 prompt,调用 LLM
|
||||
→ 写入 knowledge_domain 表
|
||||
|
||||
buildKnowledgeMap()
|
||||
→ 读取所有 knowledge_domain 记录
|
||||
→ 拼装 YAML 文本(含 documents 列表)
|
||||
→ 返回 String(供 Planner prompt 注入)
|
||||
|
||||
onDocumentChange(category)
|
||||
→ 调用 buildDomainSummary(category)(只重算受影响域)
|
||||
```
|
||||
|
||||
### 3.2 RetrievedDocTracker(新类,或内联入 LookupKnowledgeTool)
|
||||
|
||||
职责:session 级已召回文档追踪
|
||||
|
||||
```
|
||||
ConcurrentHashMap<String, Set<String>> retrieved
|
||||
key: sessionId
|
||||
value: Set<filePath>
|
||||
|
||||
isAlreadyRetrieved(sessionId, filePath) → boolean
|
||||
markRetrieved(sessionId, filePath)
|
||||
clearSession(sessionId) ← 由 SessionContextHolder.clear() 触发
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. 改动文件清单
|
||||
|
||||
| 文件 | 改动类型 | 说明 |
|
||||
|---|---|---|
|
||||
| `LookupResult.java` | 修改 | 新增 `message` 字段(去重提示文本) |
|
||||
| `Frontmatter.java` | 修改 | 新增 `covers`、`whenToRetrieve` |
|
||||
| `KnowledgeEntry.java` | 修改 | 新增 `covers`、`whenToRetrieve` |
|
||||
| `FrontmatterParser.java` | 修改 | 解析新字段 |
|
||||
| `DocumentManagementService.java` | 修改 | upload 时调 LLM 生成文档级字段;upload/delete 后触发域级重算 |
|
||||
| `KnowledgeIndexService.java` | 修改 | loadIndex 时加载域级数据;若域无记录则触发生成 |
|
||||
| `KnowledgeDomainService.java` | 新增 | 域聚合、LLM 调用、DB 读写、buildKnowledgeMap |
|
||||
| `KnowledgeDomain.java`(entity) | 新增 | knowledge_domain 表映射 |
|
||||
| `KnowledgeDomainRepository.java` | 新增 | JPA Repository |
|
||||
| `LookupKnowledgeTool.java` | 修改 | 集成 RetrievedDocTracker,检索前过滤,检索后标记 |
|
||||
| `SessionContextHolder.java` | 修改 | clear() 时通知 RetrievedDocTracker |
|
||||
| `RetrievedDocTracker.java` | 新增 | session 级去重状态管理 |
|
||||
| `ChatService.java` | 修改 | buildChatPlannerAgent 注入 knowledge map |
|
||||
| `chat-planner-prompt.md` | 修改 | 添加 knowledge map 使用规则 |
|
||||
| `V009__add_knowledge_domain.sql` | 新增 | Flyway 建表脚本 |
|
||||
|
||||
---
|
||||
|
||||
## 5. 关键决策记录
|
||||
|
||||
1. **域级 when_to_retrieve 存 DB**:避免每次重启调 LLM;文档变更时只重算受影响域
|
||||
2. **文档级 when_to_retrieve 存 metadata JSON**:沿用现有 frontmatter 存储路径,无需新字段
|
||||
3. **RetrievedDocTracker 独立于 SessionContextHolder**:SessionContextHolder 只持有 sessionId,Tracker 是业务状态,职责分离;clear() 时通过 Tracker.clearSession() 联动
|
||||
4. **Planner 只看域级**:文档级 when_to_retrieve 留 Executor 筛选(Phase 2),MVP 不暴露给 Planner
|
||||
5. **LLM 调用同步执行**:上传时同步生成,接受约 1-2s 延迟,保证数据库和 L0 索引立即一致
|
||||
@@ -0,0 +1,61 @@
|
||||
# Proposal: session-dedup-knowledge-map
|
||||
|
||||
## 问题
|
||||
|
||||
1. **ISS-001 重复召回**:`LookupKnowledgeTool` 每次调用完全无状态,同一 session 中同一文档可被重复召回 13+ 次,浪费 token、压缩上下文窗口、导致 `tool_call_count` 虚高。
|
||||
|
||||
2. **Planner 缺少全局视野**:Planner 不知道知识库里有哪些域,只能靠 Executor 反复试探,导致低效的"盲目检索"模式。
|
||||
|
||||
## 建议方案
|
||||
|
||||
### Part A:工具层去重(彻底修复 ISS-001)
|
||||
|
||||
在 `LookupKnowledgeTool` 的 session 维度维护已召回文档 ID 集合。
|
||||
每次检索时,过滤掉已召回的文档;相同 query 命中相同文档则直接跳过(返回"已在上下文中"提示)。
|
||||
|
||||
状态存储:`ConcurrentHashMap<sessionId, Set<docKey>>`,生命周期随 session(`SessionContextHolder.clear()` 时清理)。
|
||||
|
||||
### Part B:知识图谱注入 Planner
|
||||
|
||||
启动时(`KnowledgeIndexService.loadIndex()` 完成后),将 L0 索引中的所有 `KnowledgeEntry` 聚合为域级摘要(knowledge map)。
|
||||
每次构建 Planner prompt 时(`buildChatPlannerAgent()`),将 knowledge map 注入 system prompt,让 Planner 有"知识边界"。
|
||||
|
||||
聚合策略:按 `category` 字段分组,生成结构:
|
||||
```
|
||||
available_knowledge_domains:
|
||||
- domain_id: "payment"
|
||||
description: "..."
|
||||
covers: [...]
|
||||
document_count: N
|
||||
when_to_retrieve: "..."
|
||||
```
|
||||
|
||||
知识图谱的 `description` / `when_to_retrieve` 字段来源于:
|
||||
- 选项 1:直接聚合 KnowledgeEntry 的 title/summary
|
||||
- 选项 2:文档 frontmatter 中新增 `domain_description` / `when_to_retrieve` 字段
|
||||
- 选项 3:上传时 LLM 自动生成这两个字段
|
||||
|
||||
## 范围
|
||||
|
||||
**In scope**:
|
||||
- `LookupKnowledgeTool`:添加 session 级去重状态管理
|
||||
- `KnowledgeIndexService`:添加 `buildKnowledgeMap()` 方法
|
||||
- `ChatService.buildChatPlannerAgent()`:注入 knowledge map 到 prompt
|
||||
- `chat-planner-prompt.md`:添加如何使用 knowledge map 的指令
|
||||
|
||||
**Out of scope**(本次不做):
|
||||
- `EvaluationService.tool_call_count` 的统计口径调整(去重后虚高问题自然消失,但评分规则不改)
|
||||
- RRF 混合重排
|
||||
- 文档 frontmatter 自动生成(上传时 LLM 生成,留 Phase 2)
|
||||
|
||||
## 风险
|
||||
|
||||
- Part A 引入 JVM 内存 Map,高并发时多 session 并发需线程安全
|
||||
- Part B knowledge map 注入 Planner prompt 会增加每次请求的 token 消耗(固定开销)
|
||||
- 文档 `category` 字段缺失或不规范时,聚合结果可能混乱
|
||||
|
||||
## 上下文约束
|
||||
|
||||
- `SessionContextHolder` 是 ThreadLocal,异步路径不安全(已知限制,Part A 需确认同步路径)
|
||||
- `EvaluationService` 依赖 `tool_call_count`,去重会降低此值(是修复,不是回归)
|
||||
- `KnowledgeEntry` 已有 `category` 字段,但当前数据库中的文档是否都有 `category` 需确认
|
||||
+87
@@ -0,0 +1,87 @@
|
||||
# Functional Spec: session-dedup-knowledge-map
|
||||
|
||||
## REQ-01:工具层去重(Part A)
|
||||
|
||||
**触发**:`LookupKnowledgeTool.lookupKnowledge(query)` 被调用
|
||||
|
||||
**行为**:
|
||||
1. 从 `SessionContextHolder.getSessionId()` 获取当前 sessionId;若为 null(非会话上下文)跳过去重逻辑,正常检索
|
||||
2. L0+L1 检索完成后,将结果中已在 `RetrievedDocTracker` 中标记的 filePath 过滤掉
|
||||
3. 若过滤后 L0 结果为空、L1 结果也为空(全部已召回),返回 `LookupResult.found=false`,并在结果中附带提示文本:"以下文档已在本会话中检索过:[列表],无需重复召回"
|
||||
4. 未被过滤的文档正常返回后,将其 filePath 写入 `RetrievedDocTracker`
|
||||
5. `SessionContextHolder.clear()` 调用时,`RetrievedDocTracker.clearSession(sessionId)` 同步清理
|
||||
|
||||
**验收**:
|
||||
- 同一 session 内同一文档第二次命中时,返回去重提示而非完整文档内容
|
||||
- 不同 session 之间互不影响
|
||||
- sessionId 为 null 时不影响正常检索流程
|
||||
|
||||
---
|
||||
|
||||
## REQ-02:文档级 LLM 字段生成(Part B - 文档级)
|
||||
|
||||
**触发**:`DocumentManagementService.uploadDocument()` 完成 frontmatter 解析后
|
||||
|
||||
**行为**:
|
||||
1. 若文档 frontmatter 中已包含 `covers` 和 `whenToRetrieve`,跳过 LLM 生成(作者手动填写优先)
|
||||
2. 否则,调用 LLM,输入为文档 title + summary + 正文前 1000 字符
|
||||
3. Prompt 要求 LLM 返回 JSON:`{"covers": [...], "whenToRetrieve": "..."}`
|
||||
4. 解析结果,回填到 `Frontmatter` 对象
|
||||
5. 序列化存入 `api_document.metadata`;同步更新 `KnowledgeEntry` 写入 L0 索引
|
||||
6. LLM 调用失败时,`covers` 置为空列表,`whenToRetrieve` 置为 summary(降级),不阻断上传流程
|
||||
|
||||
**验收**:
|
||||
- 上传后 `api_document.metadata` 中包含 `covers` 和 `whenToRetrieve` 字段
|
||||
- frontmatter 已有这两个字段时不覆盖
|
||||
- LLM 调用异常时文档仍上传成功,字段降级填充
|
||||
|
||||
---
|
||||
|
||||
## REQ-03:域级聚合与存储(Part B - 域级)
|
||||
|
||||
**触发**:文档上传成功后;文档删除后;`loadIndex()` 时发现某域在 `knowledge_domain` 表无记录
|
||||
|
||||
**行为**:
|
||||
1. `KnowledgeDomainService.onDocumentChange(category)` 读取该 category 下所有 `KnowledgeEntry` 的 title + covers + whenToRetrieve
|
||||
2. 调用 LLM,生成域级 `when_to_retrieve`(要求 LLM 识别域内文档边界,输出含区分语义的路由描述)
|
||||
3. 写入 `knowledge_domain` 表(upsert by domain_id),同时更新 `document_count`
|
||||
4. LLM 调用失败时,`domain.when_to_retrieve` 保留上次 DB 记录;若无历史记录则置为空字符串
|
||||
|
||||
**验收**:
|
||||
- 上传文档后,对应 category 的 `knowledge_domain` 记录被更新
|
||||
- 删除文档后,对应 category 的 `document_count` 减少,`when_to_retrieve` 重新生成
|
||||
- `loadIndex()` 时无 DB 记录的域自动触发生成
|
||||
|
||||
---
|
||||
|
||||
## REQ-04:knowledge map 注入 Planner(Part B - 注入)
|
||||
|
||||
**触发**:`ChatService.buildChatPlannerAgent()` 调用时
|
||||
|
||||
**行为**:
|
||||
1. 调用 `KnowledgeDomainService.buildKnowledgeMap()` 生成 YAML 文本
|
||||
2. YAML 结构:域列表,每个域包含 domain_id、description、when_to_retrieve、documents(title + covers)
|
||||
3. 若 knowledge_domain 表为空(无任何域记录),跳过注入,不修改 prompt
|
||||
4. 注入位置:Planner system prompt 末尾,独立区块
|
||||
|
||||
**Planner prompt 附加规则**:
|
||||
- 制定步骤时,先查看 `available_knowledge_domains`,按 `when_to_retrieve` 判断是否需要检索该域
|
||||
- 每个域最多指示 Executor 检索一次;已检索过的域不再安排检索步骤
|
||||
|
||||
**验收**:
|
||||
- Planner prompt 包含 `available_knowledge_domains` 区块
|
||||
- 无域记录时 prompt 不包含该区块(不注入空结构)
|
||||
- knowledge map 文本长度 < 1000 字符(6 个文档场景下)
|
||||
|
||||
---
|
||||
|
||||
## REQ-05:Frontmatter 字段扩展
|
||||
|
||||
**行为**:
|
||||
- `Frontmatter.java` 新增 `List<String> covers` 和 `String whenToRetrieve`
|
||||
- `KnowledgeEntry.java` 新增同名字段
|
||||
- `FrontmatterParser.java` 解析 `covers`(YAML 数组)和 `when_to_retrieve`(YAML 字符串)
|
||||
|
||||
**验收**:
|
||||
- 现有文档(无新字段)上传/解析不报错,字段为 null 或空列表
|
||||
- 含新字段的文档正确解析
|
||||
@@ -0,0 +1,74 @@
|
||||
# Tasks: session-dedup-knowledge-map
|
||||
|
||||
## T1:数据层基础
|
||||
|
||||
**T1-1:新增 knowledge_domain 表** ✅
|
||||
- 创建 `src/main/resources/db/migration/V009__add_knowledge_domain.sql`
|
||||
- 字段:id、domain_id(unique)、description、when_to_retrieve(TEXT)、document_count、created_at、updated_at
|
||||
|
||||
**T1-2:新增 KnowledgeDomain 实体和 Repository** ✅
|
||||
- `KnowledgeDomain.java`:JPA 实体,对应 knowledge_domain 表
|
||||
- `KnowledgeDomainRepository.java`:`findByDomainId(String)` + save
|
||||
|
||||
---
|
||||
|
||||
## T2:Frontmatter 扩展
|
||||
|
||||
**T2-1:Frontmatter.java / KnowledgeEntry.java 新增字段** ✅
|
||||
- `Frontmatter`:新增 `List<String> covers`、`String whenToRetrieve`
|
||||
- `KnowledgeEntry`:新增 `List<String> covers`、`String whenToRetrieve`
|
||||
|
||||
**T2-2:FrontmatterParser 解析新字段** ✅
|
||||
- 解析 YAML 中的 `covers`(List)和 `when_to_retrieve`(String)
|
||||
|
||||
**T2-3:KnowledgeIndexService 替换为 Jackson 解析** ✅
|
||||
- 全量替换手写 extractJsonValue/extractJsonArray 为 `objectMapper.readValue(metadata, Frontmatter.class)`
|
||||
|
||||
**T2-4:LookupResult 新增 message 字段** ✅
|
||||
- 新增 `String message` 字段,去重时填入提示
|
||||
|
||||
---
|
||||
|
||||
## T3:工具层去重(Part A)
|
||||
|
||||
**T3-1:新增 RetrievedDocTracker** ✅
|
||||
- `RetrievedDocTracker.java`:Spring `@Component`,`ConcurrentHashMap<String, Set<String>>`
|
||||
- 方法:`isAlreadyRetrieved`、`markRetrieved`、`clearSession`
|
||||
|
||||
**T3-2:ChatService.finally 联动 Tracker** ✅
|
||||
- `executeChat` / `executeChatComplex` 的 finally 块显式调用 `retrievedDocTracker.clearSession(sessionId)`
|
||||
|
||||
**T3-3:LookupKnowledgeTool 集成去重** ✅
|
||||
- 注入 `RetrievedDocTracker`,检索后过滤已召回文档,全部已召回时返回去重提示
|
||||
|
||||
---
|
||||
|
||||
## T4:文档级 LLM 生成(Part B 文档级)
|
||||
|
||||
**T4-1:新增 DocumentFieldEnricher + 上传时调用** ✅
|
||||
- `DocumentFieldEnricher.java`:调用 LLM 生成 covers / whenToRetrieve
|
||||
- `DocumentManagementService.uploadDocument()` frontmatter 解析后调用 enrich
|
||||
- 失败时降级(covers=空列表,whenToRetrieve=summary),不阻断上传
|
||||
|
||||
---
|
||||
|
||||
## T5:域级聚合(Part B 域级)
|
||||
|
||||
**T5-1:新增 KnowledgeDomainService** ✅
|
||||
- `buildDomainSummary`:同域文档聚合 → LLM → upsert knowledge_domain
|
||||
- `buildKnowledgeMap`:全量 knowledge_domain → YAML 字符串
|
||||
- `onDocumentChange`:触发 buildDomainSummary
|
||||
|
||||
**T5-2:loadIndex 触发域生成 + 文档变更触发域重算** ✅
|
||||
- `KnowledgeIndexService.loadIndex` 末尾:无 DB 记录的域自动触发生成
|
||||
- `DocumentManagementService.uploadDocument` / `deleteDocument` 末尾:调用 `onDocumentChange`
|
||||
|
||||
---
|
||||
|
||||
## T6:Planner 注入(Part B 注入)
|
||||
|
||||
**T6-1:ChatService 注入 knowledge map** ✅
|
||||
- `buildChatPlannerAgent()` 注入 `KnowledgeDomainService.buildKnowledgeMap()` 到 prompt
|
||||
|
||||
**T6-2:chat-planner-prompt.md 新增规则** ✅
|
||||
- 新增知识库检索规则区块,要求 Planner 按 when_to_retrieve 决策、每域最多一次
|
||||
@@ -0,0 +1,6 @@
|
||||
Archive-ready for executor-action-memory-relevance
|
||||
|
||||
Created: 2026-07-01
|
||||
Tasks complete: 7/7
|
||||
Verification: static + script + manual passed
|
||||
Unverified: PRECISE scenario, domain_retrieved scenario (low risk)
|
||||
@@ -0,0 +1,189 @@
|
||||
# Design: executor-action-memory-relevance
|
||||
|
||||
## 架构设计
|
||||
|
||||
### 整体数据流
|
||||
|
||||
```
|
||||
用户问题
|
||||
→ Supervisor → Planner(规划查哪些域)
|
||||
→ Supervisor → Executor(自主调用 lookup_knowledge)
|
||||
↓
|
||||
LookupKnowledgeTool
|
||||
├─ L0 精确匹配 → l0Matches (含 category)
|
||||
├─ L1 语义检索 → l1Results (含 L2 score)
|
||||
├─ 归一化层 → computeRelevanceLevel(l0Count, l1TopScore)
|
||||
│ L2 距离 → similarity = 1 - min(score, 2.0) / 2.0
|
||||
│ L0 唯一匹配 → PRECISE
|
||||
│ L0 命中 + L1 similarity ≥ 0.75 → HIGHLY_RELEVANT
|
||||
│ 仅 L1 similarity ≥ 0.75 → HIGHLY_RELEVANT
|
||||
│ L0 多匹配 + L1 similarity [0.5, 0.75) → REFERENCE
|
||||
│ 仅 L1 similarity [0.5, 0.75) → REFERENCE
|
||||
├─ 域级行动记忆 → RetrievedDocTracker.markRetrieved(sessionId, domain, filePath)
|
||||
│ getRetrievedDomains(sessionId) → retrievedDomainsThisSession
|
||||
├─ 文档级去重 → 保留现有逻辑
|
||||
└─ 组装 LookupResult(含 relevanceLevel, completenessHint, retrievedDomainsThisSession)
|
||||
↓
|
||||
LLM 看到:
|
||||
relevanceLevel: PRECISE
|
||||
completenessHint: "知识库中不存在比上述结果更精准的文档"
|
||||
retrievedDomainsThisSession: ["infrastructure", "api"]
|
||||
```
|
||||
|
||||
### Agent 边界(保持清晰)
|
||||
|
||||
| Agent | 知道什么 | 不知道什么 |
|
||||
|-------|---------|-----------|
|
||||
| Planner | 全域知识边界(knowledge map) | 执行细节、检索结果 |
|
||||
| Executor | 自己的行动记忆(已检索域列表) | 全域知识边界(不注入 knowledge map) |
|
||||
|
||||
行动记忆通过**工具返回值**传递,不通过 prompt 注入。
|
||||
|
||||
### 数据结构设计
|
||||
|
||||
#### 1. RetrievedDocTracker 升级
|
||||
|
||||
```java
|
||||
// 现有:sessionId → Set<filePath>(文档级)
|
||||
ConcurrentHashMap<String, Set<String>> retrieved
|
||||
|
||||
// 新增:sessionId → { domain → Set<filePath> }(域级 + 文档级)
|
||||
ConcurrentHashMap<String, Map<String, Set<String>>> sessionRetrievals
|
||||
```
|
||||
|
||||
方法列表:
|
||||
- `markRetrieved(sessionId, domain, filePath)` — 一次记录两层
|
||||
- `isDocRetrieved(sessionId, filePath)` → boolean — 文档级去重(替代现有 isAlreadyRetrieved)
|
||||
- `isDomainRetrieved(sessionId, domain)` → boolean — 域级检查(Phase 2 硬限制用)
|
||||
- `getRetrievedDomains(sessionId)` → List<String> — 行动记忆(返回给 LLM)
|
||||
- `clearSession(sessionId)` — 清理(不变)
|
||||
|
||||
#### 2. LookupResult 扩展
|
||||
|
||||
```java
|
||||
@Data @Builder
|
||||
public class LookupResult {
|
||||
boolean found;
|
||||
PrimaryResult primary; // 不变,不暴露原始分数
|
||||
SupplementResult supplement; // 不变,不暴露原始分数
|
||||
// ---- 新增 ----
|
||||
String relevanceLevel; // PRECISE / HIGHLY_RELEVANT / REFERENCE
|
||||
String completenessHint; // 兜底信号
|
||||
List<String> retrievedDomainsThisSession; // 行动记忆
|
||||
String message; // 不变
|
||||
}
|
||||
```
|
||||
|
||||
**PrimaryResult 和 SupplementResult 不加任何分数字段**。原始分数在归一化层内部消化。
|
||||
|
||||
#### 3. 归一化计算
|
||||
|
||||
`RelevanceNormalizer`(LookupKnowledgeTool 内部静态方法):
|
||||
|
||||
```
|
||||
输入:l0MatchCount, l1TopScore (L2 距离)
|
||||
输出:RelevanceAssessment { relevanceLevel, completenessHint }
|
||||
|
||||
归一化公式(BGE-M3 输出 L2 归一化单位向量,已实测验证):
|
||||
similarity = 1 - min(l2Score, maxL2Distance) / maxL2Distance
|
||||
maxL2Distance 默认 2.0,yml 可覆盖
|
||||
|
||||
判定逻辑:
|
||||
if l0MatchCount == 1 → PRECISE
|
||||
if l0MatchCount > 1 && l1Similarity >= highlyRelevantThreshold → HIGHLY_RELEVANT
|
||||
if l0MatchCount == 0 && l1Similarity >= highlyRelevantThreshold → HIGHLY_RELEVANT
|
||||
if l0MatchCount > 1 && l1Similarity >= referenceThreshold → REFERENCE
|
||||
if l0MatchCount == 0 && l1Similarity >= referenceThreshold → REFERENCE
|
||||
else → 无结果
|
||||
|
||||
completenessHint 映射:
|
||||
PRECISE → "知识库中不存在比上述结果更精准的文档"
|
||||
HIGHLY_RELEVANT → "当前结果已高度相关,继续检索不太可能找到更精准的文档"
|
||||
REFERENCE → "当前结果为相关参考,如需更精准信息请明确缺少的具体维度"
|
||||
```
|
||||
|
||||
配置项(application.yml):
|
||||
```yaml
|
||||
retrieval:
|
||||
normalization:
|
||||
max-l2-distance: 2.0 # L2 距离上界(单位向量 = 2.0)
|
||||
highly-relevant-threshold: 0.75 # similarity ≥ 0.75 → HIGHLY_RELEVANT
|
||||
reference-threshold: 0.5 # similarity ≥ 0.5 → REFERENCE
|
||||
```
|
||||
|
||||
#### 4. 入库记录扩展
|
||||
|
||||
`tool_invocation` 表新增列:
|
||||
|
||||
| 列名 | 类型 | 说明 |
|
||||
|------|------|------|
|
||||
| `relevance_level` | VARCHAR(20) | PRECISE / HIGHLY_RELEVANT / REFERENCE / DEDUPED |
|
||||
| `dedup_reason` | VARCHAR(32) | doc_retrieved / domain_retrieved / null |
|
||||
|
||||
`retrieval_details` JSON 扩展:
|
||||
```json
|
||||
{
|
||||
"l0_match_count": 2,
|
||||
"l0_titles": ["MySQL连接池配置", "HikariCP参数调优"],
|
||||
"l1_top_score": 0.52,
|
||||
"l1_top_similarity": 0.74,
|
||||
"l1_match_count": 3,
|
||||
"l1_scores": [0.52, 0.68, 0.91],
|
||||
"relevance_level": "HIGHLY_RELEVANT",
|
||||
"completeness_hint": "当前结果已高度相关...",
|
||||
"retrieved_domains": ["infrastructure"],
|
||||
"dedup_reason": null
|
||||
}
|
||||
```
|
||||
|
||||
原始 L2 score 和归一化后的 similarity 都入库,保留可观测性。
|
||||
|
||||
### Executor Prompt 设计
|
||||
|
||||
不加 knowledge map,只加基于行动记忆的行为规则:
|
||||
|
||||
```markdown
|
||||
## 检索约束
|
||||
|
||||
### 1. 判断重复:基于已检索上下文
|
||||
每次 lookup_knowledge 返回值中包含 retrievedDomainsThisSession,
|
||||
表示本次会话已检索过的知识域。如果当前问题与已检索域语义重叠,
|
||||
**禁止再次调用 lookup_knowledge**。
|
||||
|
||||
### 2. 重复了该怎么办
|
||||
如果当前想检索的内容与【已检索上下文】语义相似:
|
||||
- 禁止换关键词重新检索
|
||||
- 直接基于已有事实回答
|
||||
- 如果信息不足,先明确指出缺少什么具体维度
|
||||
(如:"缺少 HikariCP 具体配置参数"、"缺少连接池耗尽的日志样例"),
|
||||
再针对该维度进行一次定向补充检索——而非盲目换词重查
|
||||
|
||||
### 3. 合法出口:允许信息不全时给出结论
|
||||
如果你认为已有信息足以回答核心问题,即使细节不全,
|
||||
也请直接给出结论并说明局限性(如:"基于已有信息,连接池配置建议如下,
|
||||
但具体参数值需结合实际负载调整")。
|
||||
**不查全不会被追责,重复检索才会被惩罚。**
|
||||
|
||||
### 4. 利用质量信号判断
|
||||
- relevanceLevel=PRECISE → 信息精准,直接使用,不再检索
|
||||
- relevanceLevel=HIGHLY_RELEVANT + 域已在 retrievedDomainsThisSession → 禁止再次调用
|
||||
- relevanceLevel=REFERENCE → 先指出缺什么维度,再定向补充一次
|
||||
- completenessHint 是知识库给你的天花板信号,信任它
|
||||
```
|
||||
|
||||
### 关键决策
|
||||
|
||||
1. **L0/L1 原始分数不暴露给 LLM** — 在归一化层内部消化,避免 LLM 混淆尺度
|
||||
2. **BGE-M3 L2 归一化已实测验证** — 范数 1.00000002,maxL2Distance=2.0 是数学硬上界
|
||||
3. **行动记忆通过工具返回值传递** — 不通过 prompt 注入,不修改 ReactAgent prompt 构建方式
|
||||
4. **不给 Executor knowledge map** — 保持 Agent 边界:Planner 知道全域,Executor 只知道自己做了什么
|
||||
5. **Phase 2 域级硬限制暂不实施** — 先观察 prompt 约束 + 归一化信号的效果
|
||||
|
||||
### 接口影响分级
|
||||
|
||||
| 变更 | 级别 | 说明 |
|
||||
|------|------|------|
|
||||
| RetrievedDocTracker 数据结构升级 | L2 内部接口 | 消费者只有 LookupKnowledgeTool,在同一实现范围内 |
|
||||
| LookupResult 新增 3 个字段 | L2 内部接口 | 消费者是 LLM(工具返回值),无跨模块调用方 |
|
||||
| tool_invocation 表新增 2 列 | L2 内部接口 | Flyway 迁移,nullable,不影响现有查询 |
|
||||
| chat-executor-prompt.md 更新 | L1 内部实现 | Prompt 文本变更,不改变接口 |
|
||||
@@ -0,0 +1,89 @@
|
||||
# Proposal: executor-action-memory-relevance
|
||||
|
||||
## 问题
|
||||
|
||||
ISS-002:Executor 在单次会话中调用 `lookup_knowledge` 20+ 次,大部分是同域换变体的冗余调用。
|
||||
|
||||
根因:
|
||||
1. **行动记忆缺失**:Executor 不知道自己已经检索过哪些域,反复用不同关键词查同一个域
|
||||
2. **质量信号缺失**:检索结果没有归一化质量等级,LLM 无法判断"结果够不够"
|
||||
3. **Prompt 约束缺失**:现有 executor prompt 要求"所有需要外部信息的地方都必须调用工具",没有"放弃检索"的合法出口
|
||||
|
||||
## 建议方案
|
||||
|
||||
### 1. 行动记忆(通过工具返回值传递)
|
||||
|
||||
`RetrievedDocTracker` 数据结构升级:`Map<sessionId, Map<domain, Set<filePath>>>`。
|
||||
|
||||
每次 `lookup_knowledge` 返回值附带 `retrievedDomainsThisSession`,让 Executor 知道自己本次会话已检索过哪些域。
|
||||
|
||||
**不给 Executor knowledge map**——保持 Agent 边界清晰:Planner 知道全域(规划查哪个域),Executor 只知道自己做了什么(执行检索 + 基于结果推理)。
|
||||
|
||||
### 2. 归一化质量等级(封装 L0/L1 分数差异)
|
||||
|
||||
在 `LookupKnowledgeTool` 内部新增归一化层,将 L0 匹配数和 L1 score 统一为三个等级:
|
||||
|
||||
| 等级 | 含义 | LLM 应做什么 |
|
||||
|------|------|-------------|
|
||||
| `PRECISE` | 精准命中 | 直接使用,不再检索 |
|
||||
| `HIGHLY_RELEVANT` | 高度相关 | 综合推理,大概率不需要继续查 |
|
||||
| `REFERENCE` | 相关参考 | 可参考,如需更精准请明确缺什么维度 |
|
||||
|
||||
归一化逻辑:
|
||||
- L0 唯一匹配 → PRECISE
|
||||
- L0 命中 + L1 高分 → HIGHLY_RELEVANT
|
||||
- L0 多匹配 + L1 中分 → HIGHLY_RELEVANT
|
||||
- L0 多匹配 + 无 L1 → REFERENCE
|
||||
- 仅 L1 命中 → 按 score 分 HIGHLY_RELEVANT / REFERENCE
|
||||
|
||||
**L0/L1 原始分数不返回给 LLM**,只在归一化层内部使用。原始分数入库(`tool_invocation.retrieval_details`)保留可观测性。
|
||||
|
||||
### 3. 兜底信号(completenessHint)
|
||||
|
||||
每次返回附带 `completenessHint`,给 LLM "天花板"信号:
|
||||
|
||||
| relevanceLevel | completenessHint |
|
||||
|----------------|-----------------|
|
||||
| PRECISE | "知识库中不存在比上述结果更精准的文档" |
|
||||
| HIGHLY_RELEVANT | "当前结果已高度相关,继续检索不太可能找到更精准的文档" |
|
||||
| REFERENCE | "当前结果为相关参考,如需更精准信息请明确缺少的具体维度" |
|
||||
|
||||
### 4. Executor prompt 重写检索约束
|
||||
|
||||
- 基于 `retrievedDomainsThisSession` 判断重复(不是"不要重复",而是"重复了该怎么办")
|
||||
- 给 LLM 合法出口:"不查全不会被追责,重复检索才会被惩罚"
|
||||
- 利用 `relevanceLevel` + `completenessHint` 判断质量
|
||||
|
||||
### 5. 入库可观测性
|
||||
|
||||
`tool_invocation` 表新增 `relevance_level` 和 `dedup_reason` 列。
|
||||
`retrieval_details` JSON 扩展:加入归一化等级、兜底信号、已检索域、去重原因、L1 top score。
|
||||
|
||||
## 范围
|
||||
|
||||
- `LookupKnowledgeTool`:归一化层 + 行动记忆注入 + 域级拦截
|
||||
- `RetrievedDocTracker`:数据结构升级(域级记录)
|
||||
- `LookupResult`:新增 `relevanceLevel`、`completenessHint`、`retrievedDomainsThisSession`
|
||||
- `chat-executor-prompt.md`:检索约束重写
|
||||
- `ToolInvocation` 实体 + V010 迁移:新增列
|
||||
- `LookupKnowledgeTool.saveToolInvocation()`:扩展入库字段
|
||||
|
||||
## 非目标
|
||||
|
||||
- 不给 Executor 注入 knowledge map(保持 Agent 边界)
|
||||
- 不修改 Planner prompt 或 Planner 逻辑
|
||||
- 不修改 `PrimaryResult`/`SupplementResult` 的字段(不暴露原始分数给 LLM)
|
||||
- Phase 2 域级硬限制暂不实施,先观察 prompt 约束效果
|
||||
|
||||
## 风险
|
||||
|
||||
1. L1 score 阈值(0.3/0.7)需要根据实际 embedding 分布调优,当前为初始值
|
||||
2. 归一化等级可能让 LLM 过早停止检索——需实测观察 REFERENCE 场景下的行为
|
||||
3. Prompt 约束仍依赖 LLM 遵守——如果效果不足,需启用 Phase 2 域级硬限制
|
||||
|
||||
## 来自 devflow 的上下文约束
|
||||
|
||||
- 前序 change `session-dedup-knowledge-map`:已实现文档级去重(RetrievedDocTracker + filePath)和 Planner knowledge map 注入
|
||||
- ISS-001:文档级重复召回已修复
|
||||
- glossary:ReactAgent 是自主决策工具调用的 Agent,不受外部流程控制
|
||||
- JPA ddl-auto 使用 validate 模式,表结构修改必须通过 Flyway 迁移
|
||||
+110
@@ -0,0 +1,110 @@
|
||||
# Functional Spec: executor-action-memory-relevance
|
||||
|
||||
## FS-1: L2 距离归一化
|
||||
|
||||
### 需求
|
||||
LookupKnowledgeTool 内部将 L1 的 L2 距离归一化为 [0,1] 区间的 similarity 值,基于 BGE-M3 输出为 L2 归一化单位向量(已实测验证,范数=1.00000002)。
|
||||
|
||||
### 可观察行为
|
||||
- 归一化公式:`similarity = 1 - min(l2Score, maxL2Distance) / maxL2Distance`
|
||||
- `maxL2Distance` 默认 2.0,可通过 `retrieval.normalization.max-l2-distance` 覆盖
|
||||
- 归一化阈值可通过 `retrieval.normalization.highly-relevant-threshold` 和 `retrieval.normalization.reference-threshold` 配置
|
||||
- 归一化计算在 LookupKnowledgeTool 内部完成,不暴露原始分数给 LLM
|
||||
|
||||
### 验收标准
|
||||
- [ ] L2 score=0 → similarity=1.0
|
||||
- [ ] L2 score=1.0 → similarity=0.5
|
||||
- [ ] L2 score=2.0 → similarity=0.0
|
||||
- [ ] L2 score=3.0(超出上界)→ similarity=0.0(min 函数截断)
|
||||
- [ ] 配置项可通过 yml 覆盖默认值
|
||||
|
||||
## FS-2: 归一化质量等级判定
|
||||
|
||||
### 需求
|
||||
基于 L0 匹配数和归一化后的 L1 similarity,输出三等级 relevanceLevel + completenessHint。
|
||||
|
||||
### 可观察行为
|
||||
- L0 唯一匹配 → PRECISE + "知识库中不存在比上述结果更精准的文档"
|
||||
- L0 命中 + L1 similarity ≥ 0.75 → HIGHLY_RELEVANT + "当前结果已高度相关,继续检索不太可能找到更精准的文档"
|
||||
- 仅 L1 similarity ≥ 0.75 → HIGHLY_RELEVANT + 对应 hint
|
||||
- L0 多匹配 + L1 similarity [0.5, 0.75) → REFERENCE + "当前结果为相关参考,如需更精准信息请明确缺少的具体维度"
|
||||
- 仅 L1 similarity [0.5, 0.75) → REFERENCE + 对应 hint
|
||||
- L1 similarity < 0.5 → 不视为有效结果
|
||||
- 无 L0 且无 L1 → found=false
|
||||
|
||||
### 验收标准
|
||||
- [ ] L0 matchCount=1 → relevanceLevel=PRECISE
|
||||
- [ ] L0 matchCount=2, L1 similarity=0.8 → relevanceLevel=HIGHLY_RELEVANT
|
||||
- [ ] L0 matchCount=0, L1 similarity=0.8 → relevanceLevel=HIGHLY_RELEVANT
|
||||
- [ ] L0 matchCount=3, L1 similarity=0.6 → relevanceLevel=REFERENCE
|
||||
- [ ] L0 matchCount=0, L1 similarity=0.4 → found=false 或 supplement 被过滤
|
||||
- [ ] 每个 relevanceLevel 对应正确的 completenessHint
|
||||
|
||||
## FS-3: 域级行动记忆
|
||||
|
||||
### 需求
|
||||
RetrievedDocTracker 升级为域级 + 文档级双层记录,支持查询当前会话已检索的域列表。
|
||||
|
||||
### 可观察行为
|
||||
- `markRetrieved(sessionId, domain, filePath)` 一次记录两层
|
||||
- `isDocRetrieved(sessionId, filePath)` 返回文档级去重结果
|
||||
- `isDomainRetrieved(sessionId, domain)` 返回域级检查结果
|
||||
- `getRetrievedDomains(sessionId)` 返回已检索域列表
|
||||
- `clearSession(sessionId)` 清理所有记录
|
||||
- 现有 `isAlreadyRetrieved(sessionId, filePath)` 语义不变(内部委托给 isDocRetrieved)
|
||||
|
||||
### 验收标准
|
||||
- [ ] markRetrieved("s1", "infrastructure", "a.md") 后,isDocRetrieved("s1", "a.md")=true
|
||||
- [ ] markRetrieved("s1", "infrastructure", "a.md") 后,isDomainRetrieved("s1", "infrastructure")=true
|
||||
- [ ] markRetrieved("s1", "infrastructure", "a.md") 后,getRetrievedDomains("s1")=["infrastructure"]
|
||||
- [ ] markRetrieved("s1", "api", "b.md") 后,getRetrievedDomains("s1")=["infrastructure","api"]
|
||||
- [ ] clearSession("s1") 后,所有方法返回空/false
|
||||
- [ ] 线程安全:ConcurrentHashMap + ConcurrentHashMap 内层
|
||||
|
||||
## FS-4: LookupResult 返回值扩展
|
||||
|
||||
### 需求
|
||||
LookupResult 新增 relevanceLevel、completenessHint、retrievedDomainsThisSession 三个字段,让 LLM 获得行动记忆和质量信号。
|
||||
|
||||
### 可观察行为
|
||||
- 每次 lookup_knowledge 返回值包含这三个新字段
|
||||
- PrimaryResult 和 SupplementResult 不变,不暴露原始分数
|
||||
- 去重拦截时,返回值仍包含 retrievedDomainsThisSession(让 LLM 知道已检索了哪些域)
|
||||
|
||||
### 验收标准
|
||||
- [ ] 正常检索返回时,LookupResult 包含 relevanceLevel + completenessHint + retrievedDomainsThisSession
|
||||
- [ ] 文档级去重拦截时,LookupResult.message 包含去重提示,retrievedDomainsThisSession 不为 null
|
||||
- [ ] PrimaryResult 和 SupplementResult 无新增分数字段
|
||||
|
||||
## FS-5: Executor Prompt 检索约束
|
||||
|
||||
### 需求
|
||||
重写 chat-executor-prompt.md 的检索规则,从"必须调用工具"改为"基于行动记忆和质量信号判断是否需要检索"。
|
||||
|
||||
### 可观察行为
|
||||
- Prompt 不包含 knowledge map
|
||||
- Prompt 包含 4 条检索约束(判断重复、重复了该怎么办、合法出口、利用质量信号)
|
||||
- 原有规则"所有需要外部信息的地方,都必须调用对应的工具"被替换
|
||||
|
||||
### 验收标准
|
||||
- [ ] Executor prompt 不包含 knowledge map 内容
|
||||
- [ ] Executor prompt 包含"禁止换关键词重新检索"约束
|
||||
- [ ] Executor prompt 包含"不查全不会被追责"合法出口
|
||||
- [ ] Executor prompt 包含 relevanceLevel 行为指导
|
||||
|
||||
## FS-6: 入库可观测性
|
||||
|
||||
### 需求
|
||||
tool_invocation 表新增 relevance_level 和 dedup_reason 列,retrieval_details JSON 扩展。
|
||||
|
||||
### 可观察行为
|
||||
- 每次 lookup_knowledge 调用后,tool_invocation 记录包含 relevance_level 和 dedup_reason
|
||||
- retrieval_details JSON 包含 l1_top_similarity(归一化后值)、relevance_level、completeness_hint、retrieved_domains、dedup_reason
|
||||
- 历史数据新列为 null,不影响现有查询
|
||||
|
||||
### 验收标准
|
||||
- [ ] V010 迁移脚本成功执行
|
||||
- [ ] 新增 relevance_level 列 VARCHAR(20) nullable
|
||||
- [ ] 新增 dedup_reason 列 VARCHAR(32) nullable
|
||||
- [ ] saveToolInvocation() 写入新字段
|
||||
- [ ] SQL 可查询归一化等级分布:`SELECT relevance_level, COUNT(*) FROM tool_invocation WHERE tool_name='lookup_knowledge' GROUP BY relevance_level`
|
||||
@@ -0,0 +1,89 @@
|
||||
# Tasks: executor-action-memory-relevance
|
||||
|
||||
## T1: RetrievedDocTracker 域级升级
|
||||
|
||||
**文件**: `src/main/java/com/superbiz/agent/tool/RetrievedDocTracker.java`
|
||||
|
||||
**改动**:
|
||||
- 数据结构从 `ConcurrentHashMap<sessionId, Set<filePath>>` 升级为 `ConcurrentHashMap<sessionId, Map<domain, Set<filePath>>>`
|
||||
- 新增 `markRetrieved(sessionId, domain, filePath)`
|
||||
- 新增 `isDocRetrieved(sessionId, filePath)` — 从内层 Map 的 values 中查找 filePath
|
||||
- 新增 `isDomainRetrieved(sessionId, domain)` — 检查 domain key 存在
|
||||
- 新增 `getRetrievedDomains(sessionId)` → `List<String>`
|
||||
- `isAlreadyRetrieved(sessionId, filePath)` 保留(委托给 isDocRetrieved,向后兼容)
|
||||
- `clearSession(sessionId)` 清理外层 key
|
||||
|
||||
**验收**: FS-3 所有验收标准通过
|
||||
|
||||
## T2: LookupResult 新增字段
|
||||
|
||||
**文件**: `src/main/java/com/superbiz/agent/dto/LookupResult.java`
|
||||
|
||||
**改动**:
|
||||
- 新增 `String relevanceLevel`
|
||||
- 新增 `String completenessHint`
|
||||
- 新增 `List<String> retrievedDomainsThisSession`
|
||||
|
||||
**验收**: 编译通过,字段存在且类型正确
|
||||
|
||||
## T3: 归一化计算逻辑
|
||||
|
||||
**文件**: `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
|
||||
**改动**:
|
||||
- 新增配置类或字段读取 `retrieval.normalization.max-l2-distance`(默认 2.0)、`highly-relevant-threshold`(默认 0.75)、`reference-threshold`(默认 0.5)
|
||||
- 新增私有方法 `computeRelevance(int l0MatchCount, float l1TopScore)` → 返回包含 `relevanceLevel` + `completenessHint` 的 record/内部类
|
||||
- L2 距离归一化:`similarity = 1 - min(l1TopScore, maxL2Distance) / maxL2Distance`
|
||||
- 判定逻辑按 design.md 中的优先级实现
|
||||
|
||||
**验收**: FS-1 + FS-2 所有验收标准通过
|
||||
|
||||
## T4: LookupKnowledgeTool 集成归一化 + 行动记忆
|
||||
|
||||
**文件**: `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
|
||||
**改动**:
|
||||
- `lookupKnowledge()` 方法中,在 Step 4(组装结果)后、Step 5(去重过滤)前,调用 `computeRelevance()` 计算 relevanceLevel 和 completenessHint
|
||||
- 从 l0Matches 提取 domain(`l0Matches.get(0).getCategory()`),L1 结果尝试从 metadata JSON 解析 category(兜底)
|
||||
- markRetrieved 调用从 `markRetrieved(sessionId, docKey)` 改为 `markRetrieved(sessionId, domain, docKey)`
|
||||
- 去重拦截时(文档级),LookupResult 也附带 retrievedDomainsThisSession
|
||||
- LookupResult.builder() 中设置三个新字段
|
||||
|
||||
**验收**: FS-4 所有验收标准通过;日志中可看到 relevanceLevel 和 completenessHint 输出
|
||||
|
||||
## T5: Executor Prompt 重写
|
||||
|
||||
**文件**: `src/main/resources/prompts/chat-executor-prompt.md`
|
||||
|
||||
**改动**:
|
||||
- 将"所有需要外部信息的地方,都必须调用对应的工具"替换为"需要外部信息时调用工具,但须遵守下方的检索约束"
|
||||
- 新增"## 检索约束"区块,包含 4 条规则(判断重复、重复了该怎么办、合法出口、利用质量信号)
|
||||
- 不注入 knowledge map
|
||||
|
||||
**验收**: FS-5 所有验收标准通过
|
||||
|
||||
## T6: 入库可观测性
|
||||
|
||||
**文件**:
|
||||
- `src/main/resources/db/migration/V010__add_relevance_level_to_tool_invocation.sql`
|
||||
- `src/main/java/com/superbiz/agent/domain/entity/ToolInvocation.java`
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`(saveToolInvocation 方法)
|
||||
|
||||
**改动**:
|
||||
- V010: ALTER TABLE tool_invocation ADD relevance_level VARCHAR(20), ADD dedup_reason VARCHAR(32)
|
||||
- ToolInvocation 实体新增 `relevanceLevel` 和 `dedupReason` 字段
|
||||
- saveToolInvocation() 中:
|
||||
- 设置 `inv.setRelevanceLevel(...)` 和 `inv.setDedupReason(...)`
|
||||
- retrieval_details JSON 扩展:新增 l1_top_similarity、relevance_level、completeness_hint、retrieved_domains、dedup_reason 字段
|
||||
- 去重拦截时,dedupReason 设为 "doc_retrieved";域级拦截时设为 "domain_retrieved"
|
||||
|
||||
**验收**: FS-6 所有验收标准通过
|
||||
|
||||
## T7: BGE-M3 归一化验证测试
|
||||
|
||||
**文件**: `src/test/java/com/superbiz/agent/service/FullPipelineSmokeTest.java`
|
||||
|
||||
**改动**:
|
||||
- 已完成:embeddingBgeM3Works() 中新增 L2 范数断言(范数=1.00000002,测试已通过)
|
||||
|
||||
**验收**: 测试通过,范数断言 |norm - 1.0| < 0.01
|
||||
@@ -0,0 +1 @@
|
||||
ready
|
||||
@@ -0,0 +1,42 @@
|
||||
{
|
||||
"id": "chat-verifier-agent",
|
||||
"metadata": {
|
||||
"status": "archived",
|
||||
"created_at": "2026-07-02",
|
||||
"updated_at": "2026-07-03",
|
||||
"archive_readiness": "archived",
|
||||
"implementation_status": "archived"
|
||||
},
|
||||
"summary": "Add a verifier agent to the complex chat path and persist auditable verifier decisions with evidence traceability.",
|
||||
"artifacts": {
|
||||
"proposal": "proposal.md",
|
||||
"design": "design.md",
|
||||
"tasks": "tasks.md",
|
||||
"specs": [
|
||||
"specs/chat-verifier-agent/spec.md"
|
||||
],
|
||||
"devflow": "devflow/projects/2026-07-02-chat-verifier-agent"
|
||||
},
|
||||
"tasks": [
|
||||
"Verifier prompt",
|
||||
"VerifierInputHook explicit payload",
|
||||
"ChatService planner-executor-verifier orchestration",
|
||||
"Verdict routing and fixed user output templates",
|
||||
"Tool trace summary and evidence_refs traceability",
|
||||
"self_evaluation merge semantics",
|
||||
"Compile and runtime verification"
|
||||
],
|
||||
"verification": [
|
||||
{
|
||||
"type": "script",
|
||||
"command": "mvn -q -DskipTests compile",
|
||||
"result": "passed"
|
||||
},
|
||||
{
|
||||
"type": "runtime",
|
||||
"command": "POST /api/chat",
|
||||
"session_id": "9138f064",
|
||||
"result": "planner, executor, and verifier executed; verifier_evaluation contains evidence_refs and tool_trace_summary source_invocation_ids"
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,349 @@
|
||||
## Context
|
||||
|
||||
Chat 多 Agent 链路当前由 ChatService 驱动 Planner → Executor,答案输出前无质量门禁。Verifier Agent 作为 Executor 后置质量门禁,在 Executor 输出后做事实核查。
|
||||
|
||||
前序 change `executor-action-memory-relevance` 已在 Executor 侧构建了行动记忆和检索质量归一化,Verifier 不需要重复验证检索质量。
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
- Verifier 作为无工具 ReactAgent,由 ChatService 显式调用
|
||||
- Verifier 输出 verdict (PASS/LOW_CONFID/REJECT) + groundedness_score + facts_checked
|
||||
- ChatService 负责单轮显式编排:Planner → Executor → Verifier
|
||||
- ChatService 外层根据 Verifier 判决做轮次路由:PASS→输出,LOW_CONFID≥0.5→带声明输出,LOW_CONFID<0.5→补充一轮,REJECT→降级
|
||||
- Verifier 判决写入 diagnosis_session.self_evaluation JSON 容器中的 `verifier_evaluation` 槽位做可观测
|
||||
|
||||
**Non-Goals:**
|
||||
- Verifier 不调用工具
|
||||
- 不改动单 Agent 链路
|
||||
- 不修改 Executor 的输出内容
|
||||
- 不涉及数据库表结构变更
|
||||
- Verifier 不继承 Executor 的中间推理过程(通过 MessagesModelHook 过滤)
|
||||
|
||||
## Decisions
|
||||
|
||||
| 决策 | 选择 | 放弃方案 | 原因 |
|
||||
|------|------|---------|------|
|
||||
| Verifier 是否有工具 | 无工具 ReactAgent | 有工具的 Agent | 职责单一,只核查不检索 |
|
||||
| 判决分类 | PASS / LOW_CONFID / REJECT | PASS / FAIL 二分类 | LOW_CONFID 提供了弹性输出路径 |
|
||||
| 回调机制 | ChatService 外层控制最多两轮 | 全交给 Supervisor / 不回调 | 轮次上限需要硬控制,不能只靠 prompt 记忆 |
|
||||
| 可观测方案 | 写入 self_evaluation JSON 容器 | agent_step / tool_invocation / 新表 | 不改表结构,同时避免与 evidence_score 覆盖冲突 |
|
||||
| 输入隔离 | 显式状态输入 + MessagesModelHook 裁剪噪音 | 仅靠原始消息过滤 / 数据库注入 | Verifier 需要稳定读取 query、工具摘要、最终答案,不能依赖消息格式猜测 |
|
||||
| 阈值配置 | yml 配置化 | 硬编码 | 方便运维调整,不需改代码 |
|
||||
|
||||
## Verifier 输入契约
|
||||
|
||||
Verifier 的业务输入由 `ChatService` 显式组装,不依赖原始 conversation messages 的隐式结构。
|
||||
|
||||
### 必选输入
|
||||
|
||||
- `original_query`:用户原始问题
|
||||
- `executor_final_answer`:本轮 Executor 最终答案
|
||||
- `tool_trace_summary`:由工具调用事实整理出的半结构化摘要
|
||||
|
||||
### 条件输入
|
||||
|
||||
- `retry_context`:仅第二轮注入,描述上一轮 verifier 发现的证据缺口和补充约束
|
||||
|
||||
### tool_trace_summary 最小结构
|
||||
|
||||
```json
|
||||
[
|
||||
{
|
||||
"tool_name": "lookup_knowledge",
|
||||
"success": true,
|
||||
"input_summary": "查询 ERR_TIMEOUT",
|
||||
"output_summary": "命中 payment/errors.md,返回错误码定义",
|
||||
"evidence_level": "direct"
|
||||
},
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": false,
|
||||
"input_summary": "按 traceId 查询日志",
|
||||
"output_summary": "日志服务超时",
|
||||
"evidence_level": "none"
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
约束:
|
||||
|
||||
- `tool_trace_summary` 只纳入证据型工具调用,不纳入纯辅助或无业务事实意义的工具
|
||||
- `tool_trace_summary` 来源于工具调用事实,不直接透传原始日志全文
|
||||
- Verifier 基于摘要做事实核查,不直接读取数据库
|
||||
- 若某工具调用失败,仍需记录在摘要中,供 Verifier 判断证据缺口
|
||||
|
||||
### 证据型工具边界
|
||||
|
||||
默认纳入 `tool_trace_summary` 的工具:
|
||||
|
||||
- `lookup_knowledge`
|
||||
- `query_logs`
|
||||
- `query_metrics`
|
||||
- `query_order` 或其他业务事实查询类工具
|
||||
- 其他只读、能提供客观事实的工具
|
||||
|
||||
默认不纳入:
|
||||
|
||||
- `getCurrentDateTime`
|
||||
- 纯格式化、转换、控制类工具
|
||||
- 与事实核查无关的辅助工具
|
||||
|
||||
### 摘要压缩规则
|
||||
|
||||
- 每次调用只保留“最小证据摘要”,不透传原始返回全文
|
||||
- `output_summary` 控制为 1-3 句,重点描述“这次调用证明了什么 / 没能证明什么”
|
||||
- 失败调用必须保留,但统一标记:
|
||||
- `success=false`
|
||||
- `evidence_level=none`
|
||||
- 同一工具、同一主题域、同一轮次的重复调用可以折叠为一条合并摘要
|
||||
- 合并摘要至少保留:
|
||||
- 首次有效命中结果
|
||||
- 额外重复次数 / 未命中次数 / 失败次数
|
||||
|
||||
### 截断优先级
|
||||
|
||||
若 `tool_trace_summary` 过长,优先保留:
|
||||
|
||||
1. 被 `executor_final_answer` 直接引用的证据
|
||||
2. 支撑根因结论的证据
|
||||
3. 支撑修复结论的证据
|
||||
4. 与上一轮 `retry_context` 缺口直接相关的证据
|
||||
|
||||
低优先级、与最终答案无关的辅助性工具摘要可被截断。
|
||||
|
||||
### retry_context 最小结构
|
||||
|
||||
```json
|
||||
{
|
||||
"round": 1,
|
||||
"missing_evidence_facts": [
|
||||
"“根因是连接池耗尽”缺少直接证据",
|
||||
"“错误码 ERR_TIMEOUT 来自支付网关”只有间接支持"
|
||||
],
|
||||
"instruction": "仅补充以上断言相关证据,不要重复已完成检索"
|
||||
}
|
||||
```
|
||||
|
||||
### MessagesModelHook 职责边界
|
||||
|
||||
- 可以:移除 Planner/Executor 中间推理、无关闲聊和冗余 message
|
||||
- 不可以:作为 Verifier 核心业务输入的唯一来源
|
||||
- 目标:降噪,而非拼装业务事实
|
||||
|
||||
## Verifier 判决矩阵
|
||||
|
||||
Verifier 先提取并校验 `facts_checked`,再依据矩阵生成 verdict,避免只靠模型主观判断。
|
||||
|
||||
### facts_checked 分类
|
||||
|
||||
每条事实仅允许以下四类之一:
|
||||
|
||||
- `direct_evidence`:工具结果中有明确直接证据
|
||||
- `indirect_support`:可由工具结果合理推导,但不是直接陈述
|
||||
- `no_evidence`:工具结果中没有足够信息支撑
|
||||
- `contradicted`:工具结果与该事实冲突,或该事实编造了不存在的关键实体/错误码/结论
|
||||
|
||||
### 关键事实范围
|
||||
|
||||
Verifier 优先校验关键事实,至少包括:
|
||||
|
||||
- 根因结论(root cause)
|
||||
- 错误码 / 接口 / 组件归属
|
||||
- 证据来源陈述(如“日志显示”“文档说明”)
|
||||
- 明确修复结论
|
||||
|
||||
一般性建议、风险提示、非事实性表述默认不纳入关键事实,除非答案明确声称“已被证据证明”。
|
||||
|
||||
### verdict 规则
|
||||
|
||||
- `REJECT`
|
||||
- 任意关键事实为 `contradicted`
|
||||
- 或答案编造了工具/日志/文档中不存在的关键实体、错误码、结论
|
||||
|
||||
- `PASS`
|
||||
- 所有关键事实均为 `direct_evidence` 或 `indirect_support`
|
||||
- 且至少一条关键事实为 `direct_evidence`
|
||||
- 且不存在 `contradicted`
|
||||
|
||||
- `LOW_CONFID`
|
||||
- 不存在 `contradicted`
|
||||
- 但存在关键事实为 `no_evidence`
|
||||
- 或所有关键事实都只有 `indirect_support`,缺少直接锚点
|
||||
|
||||
一句话归纳:
|
||||
|
||||
- `REJECT` = 有冲突
|
||||
- `LOW_CONFID` = 无冲突但缺关键证据
|
||||
- `PASS` = 无冲突且关键事实均有支撑
|
||||
|
||||
### groundedness_score 计算
|
||||
|
||||
`groundedness_score` 不由模型自由打分,而由关键事实分类映射得到:
|
||||
|
||||
```text
|
||||
direct_evidence = 1.0
|
||||
indirect_support = 0.6
|
||||
no_evidence = 0.0
|
||||
contradicted = 0.0
|
||||
```
|
||||
|
||||
规则:
|
||||
|
||||
- 仅对关键事实计分
|
||||
- 取平均值后截断到 `[0.0, 1.0]`
|
||||
- 若存在任意关键事实为 `contradicted`,直接 verdict=`REJECT`,且 `groundedness_score=0.0`
|
||||
|
||||
### 第二轮补证据范围
|
||||
|
||||
第二轮 `retry_context` 仅回灌以下关键缺口:
|
||||
|
||||
- 关键事实为 `no_evidence`
|
||||
- 关键事实为 `indirect_support`,但仍缺直接证据锚点
|
||||
|
||||
`REJECT` 不进入第二轮补证据,直接降级输出。
|
||||
|
||||
## 用户侧输出协议
|
||||
|
||||
Verifier 的内部判决与用户侧最终输出类型分离:
|
||||
|
||||
- `PASS` → `NORMAL`
|
||||
- `LOW_CONFID` → `LOW_CONFID_WITH_DISCLAIMER`
|
||||
- `REJECT` → `DEGRADED`
|
||||
|
||||
### LOW_CONFID_WITH_DISCLAIMER
|
||||
|
||||
适用场景:
|
||||
|
||||
- 第一轮 `LOW_CONFID` 且 `groundedness_score >= threshold`
|
||||
- 第二轮后仍为 `LOW_CONFID`
|
||||
|
||||
输出规则:
|
||||
|
||||
- 使用固定免责声明前缀
|
||||
- 免责声明后拼接 `executor_final_answer`
|
||||
- 可选附加“当前证据缺口”列表,但来源必须是 verifier 的关键缺口,不得自由扩写
|
||||
|
||||
建议模板:
|
||||
|
||||
```text
|
||||
以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。
|
||||
|
||||
{executor_final_answer}
|
||||
|
||||
当前缺口:
|
||||
- ...
|
||||
- ...
|
||||
```
|
||||
|
||||
### DEGRADED
|
||||
|
||||
适用场景:
|
||||
|
||||
- 任意一轮 `REJECT`
|
||||
- 系统无法基于现有证据形成可靠结论
|
||||
|
||||
输出规则:
|
||||
|
||||
- 不透传原始 `executor_final_answer`
|
||||
- 使用固定降级模板
|
||||
- 仅允许包含:
|
||||
- 已确认信息
|
||||
- 证据缺口
|
||||
- 下一步建议
|
||||
|
||||
建议模板:
|
||||
|
||||
```text
|
||||
当前无法基于已获取证据生成可靠结论,建议人工介入。
|
||||
|
||||
已确认信息:
|
||||
- ...
|
||||
|
||||
证据缺口:
|
||||
- ...
|
||||
|
||||
建议下一步:
|
||||
- ...
|
||||
```
|
||||
|
||||
### 输出边界
|
||||
|
||||
- `LOW_CONFID_WITH_DISCLAIMER` 可以带出原始答案,但必须加固定免责声明
|
||||
- `DEGRADED` 不得透传未经验证的原始答案
|
||||
- 用户侧输出模板由代码层拼装,不依赖 Verifier 自由生成
|
||||
|
||||
## self_evaluation 存储约定
|
||||
|
||||
`diagnosis_session.self_evaluation` 统一定义为 JSON 容器对象,而不是单一评估结果:
|
||||
|
||||
```json
|
||||
{
|
||||
"rule_evaluation": {
|
||||
"evidence_score": 65,
|
||||
"source": "rule",
|
||||
"factors": []
|
||||
},
|
||||
"verifier_evaluation": {
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.42,
|
||||
"facts_checked": [],
|
||||
"rationale": "...",
|
||||
"round": 1
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
写入约束:
|
||||
|
||||
- `EvaluationService` 只负责写 `rule_evaluation`
|
||||
- `ChatService` 只负责写 `verifier_evaluation`
|
||||
- 两侧都必须使用 read-modify-write,保留另一侧已有内容
|
||||
- 禁止整段覆盖 `self_evaluation`,除非初始化为空对象
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Verifier 误判导致好答案被降级 → Mitigation: REJECT 仅用于明显编造场景,LOW_CONFID 为主要输出路径
|
||||
- [Risk] callback Planner 后新答案质量不一定提升 → Mitigation: 仅回调一次,Token 成本可控
|
||||
- [Risk] 第二轮仍可能产出 REJECT → Mitigation: 第二轮 REJECT 仍降级,不透传
|
||||
- [Risk] `self_evaluation` 被异步 evidence_score 覆盖 → Mitigation: 定义 JSON 容器槽位,统一 read-modify-write
|
||||
- [Risk] Verifier 增加 Token 消耗 → Mitigation: 单次轻量 LLM 调用,估算 <500 token
|
||||
- [Risk] 消息过滤可能导致输入契约漂移 → Mitigation: 主输入由显式状态输入提供,Hook 仅用于剔除中间推理和无关噪音
|
||||
- [Risk] 判决边界主观化,导致不同模型输出不稳定 → Mitigation: 用 facts_checked 分类 + verdict 矩阵 + 映射分数约束输出
|
||||
- [Risk] 最终用户文案随模型漂移,导致产品行为不稳定 → Mitigation: LOW_CONFID/DEGRADED 使用固定输出协议和模板
|
||||
|
||||
## Implementation Plan
|
||||
|
||||
1. 创建 `chat-verifier-prompt.md`
|
||||
2. 新建 `VerifierInputHook.java`(MessagesModelHook 实现,BEFORE_MODEL 时裁剪 messages,只保留必要上下文)
|
||||
3. `ChatService.java` 新增 `buildChatVerifierAgent()` 方法(ReactAgent,无工具,带 hook)
|
||||
4. 添加 `verifier.low-confidence-threshold: 0.5` 到 application.yml
|
||||
5. 在 `ChatService.executeChatComplex()` 中显式调用 `Planner → Executor → Verifier`
|
||||
6. 保留 `SupervisorAgent` 构造作为 legacy residue,不再依赖 prompt-only supervisor sequencing 保证 Verifier 执行
|
||||
7. 在 `ChatService.executeChatComplex()` 外层实现最多两轮调用控制
|
||||
8. 组装 Verifier 显式状态输入:`original_query` / `executor_final_answer` / `tool_trace_summary` / `retry_context`
|
||||
9. 将 `self_evaluation` 升级为 JSON 容器读写:`rule_evaluation` / `verifier_evaluation`
|
||||
10. 读取 Verifier 判决写入 `verifier_evaluation`
|
||||
|
||||
## Implementation Notes
|
||||
|
||||
### Explicit orchestration
|
||||
|
||||
The final implementation uses `ChatService` to call `planner -> executor -> verifier` directly in each outer round. This replaces the earlier prompt-only dependency on `SupervisorAgent` for verifier execution. The supervisor construction remains in the code as legacy residue, but runtime correctness is driven by explicit `callAgent(...)` ordering.
|
||||
|
||||
### Traceability model
|
||||
|
||||
The implemented verifier input and persisted evaluation include an evidence index:
|
||||
|
||||
- `tool_trace_summary[*].trace_ref`
|
||||
- `tool_trace_summary[*].source_invocation_ids`
|
||||
- `tool_trace_summary[*].query_samples`
|
||||
- `tool_trace_summary[*].retrieval_layers`
|
||||
- `tool_trace_summary[*].relevance_levels`
|
||||
- `tool_trace_summary[*].source_documents`
|
||||
|
||||
Each verifier fact may carry `facts_checked[*].evidence_refs`, which points back to `trace_ref` and the underlying `tool_invocation` ids. This closes the audit gap where verifier could list many checked facts but the reviewer could not tell which facts related to which tool calls.
|
||||
|
||||
### Observability adjustment
|
||||
|
||||
`agent_step.thought` is now intentionally concise for verifier steps. Full verifier judgment belongs in `diagnosis_session.self_evaluation.verifier_evaluation`, with `model_output` retaining the model output snapshot.
|
||||
@@ -0,0 +1,31 @@
|
||||
# Proposal: chat-verifier-agent
|
||||
|
||||
## Why
|
||||
|
||||
Chat 多 Agent 链路缺少出口质量门禁。Executor 输出答案后会直接返回给用户,无法在返回前拦截缺证据、低置信或明显编造的结论。
|
||||
|
||||
## What Changes
|
||||
|
||||
- 新增无工具 Verifier Agent,在 Executor 输出后读取答案和工具调用证据摘要,产出 `PASS` / `LOW_CONFID` / `REJECT` 判决。
|
||||
- `ChatService` 显式编排 `Planner -> Executor -> Verifier`,并根据 Verifier 判决控制最终输出或最多一次补充轮次。
|
||||
- Verifier 输入使用显式状态块:`original_query`、`executor_final_answer`、`tool_trace_summary`、第二轮可选 `retry_context`。
|
||||
- `diagnosis_session.self_evaluation` 作为 JSON 容器保存 `rule_evaluation` 与 `verifier_evaluation`,避免异步评分覆盖 Verifier 结果。
|
||||
- Verifier 结果增加可追溯证据引用:`tool_trace_summary[*].trace_ref`、`source_invocation_ids` 与 `facts_checked[*].evidence_refs`。
|
||||
- `LOW_CONFID` 和 `REJECT` 用户侧输出使用固定协议,`REJECT` 不透传未经验证的原始答案。
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `chat-verifier-agent`: Chat 多 Agent 出口事实核查、判决路由、观测存储和证据可追溯能力。
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- None.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected code: `ChatService`, chat verifier prompt, verifier input assembly, self-evaluation persistence, multi-agent runtime orchestration.
|
||||
- Affected runtime behavior: complex chat path now runs a Verifier gate after Executor and may perform one bounded retry for low-confidence evidence gaps.
|
||||
- No database schema change is required; `self_evaluation` remains the persistence container.
|
||||
- Non-goals: Verifier 不调用工具、不改写 Executor 答案、不影响单 Agent 链路、不支持超过两轮的补充编排。
|
||||
+189
@@ -0,0 +1,189 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Verifier SHALL fact-check Executor answers
|
||||
The system SHALL have a Verifier Agent that reads the Executor's answer and the tool call history, then produces a structured verdict.
|
||||
|
||||
#### Scenario: PASS verdict when all claims have evidence
|
||||
- **WHEN** all critical facts in the Executor's answer have direct or indirect support in tool call results
|
||||
- **AND** at least one critical fact has direct evidence
|
||||
- **AND** no critical fact is contradicted
|
||||
- **THEN** the Verifier SHALL output verdict="PASS" with groundedness_score ≥ 0.5
|
||||
|
||||
#### Scenario: LOW_CONFID verdict with partial evidence
|
||||
- **WHEN** no critical fact contradicts the tool results
|
||||
- **AND** some critical facts have no supporting evidence
|
||||
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
|
||||
|
||||
#### Scenario: LOW_CONFID verdict with only indirect support
|
||||
- **WHEN** no critical fact contradicts the tool results
|
||||
- **AND** all critical facts are only indirectly supported
|
||||
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
|
||||
|
||||
#### Scenario: REJECT verdict when claims contradict evidence
|
||||
- **WHEN** any critical fact in the Executor's answer contradicts tool call results
|
||||
- **OR** the answer fabricates a key entity, error code, or conclusion that does not exist in the tool evidence
|
||||
- **THEN** the Verifier SHALL output verdict="REJECT"
|
||||
|
||||
### Requirement: Verifier SHALL output structured JSON
|
||||
The Verifier SHALL output a JSON object with verdict, groundedness_score, facts_checked array, and rationale.
|
||||
|
||||
#### Scenario: Output format validation
|
||||
- **WHEN** the Verifier completes its analysis
|
||||
- **THEN** the output SHALL contain "verdict", "groundedness_score", "facts_checked", and "rationale" fields
|
||||
- **AND** groundedness_score SHALL be a float between 0.0 and 1.0
|
||||
- **AND** verdict SHALL be one of "PASS", "LOW_CONFID", or "REJECT"
|
||||
|
||||
#### Scenario: strict schema output
|
||||
- **WHEN** the Verifier returns its result
|
||||
- **THEN** it SHALL output exactly one JSON object
|
||||
- **AND** it SHALL NOT output Markdown, code fences, or explanatory text outside the JSON object
|
||||
- **AND** the JSON object SHALL include `critical_fact_count`
|
||||
- **AND** each `facts_checked` item SHALL include `fact`, `is_critical`, `verification`, and `detail`
|
||||
|
||||
### Requirement: facts_checked SHALL use a fixed classification set
|
||||
Each checked fact SHALL be labeled using a fixed evidence classification.
|
||||
|
||||
#### Scenario: fact classification values
|
||||
- **WHEN** the Verifier emits `facts_checked`
|
||||
- **THEN** each fact SHALL use one of `direct_evidence`, `indirect_support`, `no_evidence`, or `contradicted`
|
||||
|
||||
### Requirement: groundedness_score SHALL be derived from fact classifications
|
||||
The groundedness score SHALL be computed from critical fact classifications instead of being freely chosen by the model.
|
||||
|
||||
#### Scenario: contradicted fact forces reject
|
||||
- **WHEN** any critical fact is labeled `contradicted`
|
||||
- **THEN** the Verifier SHALL output verdict="REJECT"
|
||||
- **AND** groundedness_score SHALL be `0.0`
|
||||
|
||||
#### Scenario: score derived from supported facts
|
||||
- **WHEN** no critical fact is contradicted
|
||||
- **THEN** groundedness_score SHALL be computed from the mapped values of critical facts
|
||||
- **AND** the implementation SHALL use the fixed mapping `direct_evidence=1.0`, `indirect_support=0.6`, `no_evidence=0.0`
|
||||
- **AND** the result SHALL be clamped into `[0.0, 1.0]`
|
||||
|
||||
### Requirement: ChatService SHALL route based on Verifier verdict
|
||||
The system SHALL use ChatService for explicit single-round `Planner → Executor → Verifier` orchestration and SHALL use ChatService to control whether an additional round is allowed.
|
||||
|
||||
#### Scenario: PASS → direct output
|
||||
- **WHEN** Verifier outputs verdict="PASS"
|
||||
- **THEN** the system SHALL output the Executor's answer directly
|
||||
|
||||
#### Scenario: LOW_CONFID score≥0.5 → output with disclaimer
|
||||
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score ≥ 0.5
|
||||
- **THEN** the system SHALL output the Executor's answer prefixed with a fixed confidence disclaimer
|
||||
|
||||
#### Scenario: LOW_CONFID score<0.5 → trigger one additional round
|
||||
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score < 0.5 and this is the first callback
|
||||
- **THEN** the ChatService SHALL invoke one additional `Planner → Executor → Verifier` round to supplement evidence
|
||||
- **AND** after the second Verifier run, verdict="LOW_CONFID" SHALL be output with a confidence disclaimer
|
||||
- **AND** after the second Verifier run, verdict="REJECT" SHALL still produce a degraded output
|
||||
|
||||
#### Scenario: REJECT does not enter retry round
|
||||
- **WHEN** Verifier outputs verdict="REJECT"
|
||||
- **THEN** the system SHALL NOT start a retry round for evidence补充
|
||||
- **AND** it SHALL produce a degraded output directly
|
||||
|
||||
#### Scenario: REJECT → degraded output
|
||||
- **WHEN** Verifier outputs verdict="REJECT"
|
||||
- **THEN** the system SHALL output a degraded result indicating the answer cannot be reliably generated
|
||||
- **AND** it SHALL NOT pass through the raw Executor answer
|
||||
|
||||
### Requirement: User-facing verifier outputs SHALL follow fixed templates
|
||||
The system SHALL use fixed output protocols for LOW_CONFID and REJECT user-facing responses.
|
||||
|
||||
#### Scenario: LOW_CONFID uses disclaimer template
|
||||
- **WHEN** the final verdict is `LOW_CONFID`
|
||||
- **THEN** the user-facing response SHALL prepend a fixed disclaimer before the Executor answer
|
||||
- **AND** optional evidence gaps, if present, SHALL come only from verifier-identified critical gaps
|
||||
|
||||
#### Scenario: REJECT uses degraded template
|
||||
- **WHEN** the final verdict is `REJECT`
|
||||
- **THEN** the user-facing response SHALL use a degraded template
|
||||
- **AND** it SHALL include only confirmed facts, evidence gaps, and next-step suggestions
|
||||
- **AND** it SHALL NOT include unverified raw answer content
|
||||
|
||||
### Requirement: Verifier SHALL be observable
|
||||
The Verifier's verdict SHALL be persisted for observability.
|
||||
|
||||
#### Scenario: verdict written to self_evaluation
|
||||
- **WHEN** the Verifier produces a verdict
|
||||
- **THEN** the ChatService SHALL write the verdict data under `diagnosis_session.self_evaluation.verifier_evaluation`
|
||||
- **AND** existing `rule_evaluation` data SHALL be preserved
|
||||
|
||||
### Requirement: self_evaluation SHALL be a container object
|
||||
The `diagnosis_session.self_evaluation` field SHALL store multiple evaluation channels in one JSON object.
|
||||
|
||||
#### Scenario: rule evaluation stored separately
|
||||
- **WHEN** the rule-based evidence scoring completes
|
||||
- **THEN** the EvaluationService SHALL write the result under `rule_evaluation`
|
||||
- **AND** existing `verifier_evaluation` data SHALL be preserved
|
||||
|
||||
#### Scenario: verifier evaluation stored separately
|
||||
- **WHEN** the Verifier completes
|
||||
- **THEN** the ChatService SHALL write the result under `verifier_evaluation`
|
||||
- **AND** existing `rule_evaluation` data SHALL be preserved
|
||||
|
||||
#### Scenario: no whole-object overwrite after initialization
|
||||
- **WHEN** either evaluation channel updates `self_evaluation`
|
||||
- **THEN** the implementation SHALL use read-modify-write semantics
|
||||
- **AND** it SHALL NOT replace the whole JSON object except when initializing from null
|
||||
|
||||
### Requirement: Verifier SHALL consume explicit verification inputs
|
||||
The Verifier SHALL receive explicit verification inputs rather than inferring them only from raw conversation history.
|
||||
|
||||
#### Scenario: explicit input blocks available to Verifier
|
||||
- **WHEN** the Verifier starts
|
||||
- **THEN** the system SHALL provide `original_query`, `executor_final_answer`, and `tool_trace_summary` as explicit inputs
|
||||
- **AND** `retry_context` SHALL be provided on the second round only
|
||||
- **AND** message filtering MAY be used only to remove intermediate reasoning or unrelated noise
|
||||
|
||||
#### Scenario: tool trace summary derived from tool facts
|
||||
- **WHEN** the system prepares verifier inputs
|
||||
- **THEN** `tool_trace_summary` SHALL be generated from tool invocation facts
|
||||
- **AND** each summary item SHALL include tool name, success state, input summary, output summary, and evidence level
|
||||
- **AND** raw conversation history SHALL NOT be the only source of verifier evidence context
|
||||
|
||||
#### Scenario: tool trace summary preserves invocation references
|
||||
- **WHEN** the system prepares verifier inputs
|
||||
- **THEN** each summary item SHALL include a stable `trace_ref`
|
||||
- **AND** each summary item SHALL preserve `source_invocation_ids` for the tool invocation rows that contributed to the summary
|
||||
- **AND** each summary item SHOULD include query samples, retrieval layers, relevance levels, and source document labels when available
|
||||
|
||||
#### Scenario: only evidence-bearing tools included
|
||||
- **WHEN** the system generates `tool_trace_summary`
|
||||
- **THEN** it SHALL include only evidence-bearing tool invocations
|
||||
- **AND** non-evidence helper tools such as time or formatting tools SHALL be excluded by default
|
||||
|
||||
#### Scenario: failed evidence calls preserved as evidence gaps
|
||||
- **WHEN** an evidence-bearing tool invocation fails or returns no usable evidence
|
||||
- **THEN** the summary SHALL still include that invocation
|
||||
- **AND** it SHALL mark the entry as unsuccessful with an evidence level representing no evidence
|
||||
|
||||
#### Scenario: repeated tool calls may be compacted
|
||||
- **WHEN** repeated tool invocations concern the same tool, topic domain, and round
|
||||
- **THEN** the system MAY compact them into a merged summary entry
|
||||
- **AND** the merged entry SHALL preserve the first effective hit and the count of repeated, failed, or no-hit calls
|
||||
|
||||
#### Scenario: raw outputs not passed through in full
|
||||
- **WHEN** a tool invocation returns large raw content
|
||||
- **THEN** `tool_trace_summary` SHALL keep only a minimal evidence summary
|
||||
- **AND** the raw output SHALL NOT be passed through in full to the Verifier
|
||||
|
||||
#### Scenario: MessagesModelHook used only for noise reduction
|
||||
- **WHEN** a MessagesModelHook is used for the Verifier
|
||||
- **THEN** it MAY remove intermediate reasoning or irrelevant messages
|
||||
- **AND** it SHALL NOT be the primary source for assembling verifier business inputs
|
||||
|
||||
### Requirement: Verifier facts SHALL be auditable
|
||||
Verifier facts SHALL be linkable to the evidence summaries used during verification.
|
||||
|
||||
#### Scenario: facts_checked contains evidence refs
|
||||
- **WHEN** the Verifier emits `facts_checked`
|
||||
- **THEN** each fact SHALL include `evidence_refs`
|
||||
- **AND** each evidence ref SHALL point to an existing `tool_trace_summary.trace_ref`
|
||||
- **AND** each evidence ref SHALL preserve the relevant `source_invocation_ids` when available
|
||||
|
||||
#### Scenario: verifier evaluation persists traceability snapshot
|
||||
- **WHEN** the ChatService persists `verifier_evaluation`
|
||||
- **THEN** it SHALL include `traceability_version`
|
||||
- **AND** it SHALL include the `tool_trace_summary` snapshot used by the Verifier
|
||||
@@ -0,0 +1,58 @@
|
||||
# Tasks: chat-verifier-agent
|
||||
|
||||
## 1. Verifier Prompt
|
||||
|
||||
- [x] 1.1 Create `src/main/resources/prompts/chat-verifier-prompt.md`.
|
||||
- [x] 1.2 Define fixed fact classifications: `direct_evidence`, `indirect_support`, `no_evidence`, `contradicted`.
|
||||
- [x] 1.3 Define critical fact scope, verdict matrix, and `groundedness_score` mapping.
|
||||
- [x] 1.4 Define strict JSON output schema: `verdict`, `groundedness_score`, `critical_fact_count`, `facts_checked`, `rationale`.
|
||||
- [x] 1.5 Forbid Markdown, code fences, schema-extra fields, and text outside the JSON object.
|
||||
- [x] 1.6 Require `facts_checked[*].evidence_refs` for traceability to tool evidence.
|
||||
|
||||
## 2. Verifier Input Hook
|
||||
|
||||
- [x] 2.1 Add `VerifierInputHook.java` as a `MessagesModelHook` running at `BEFORE_MODEL`.
|
||||
- [x] 2.2 Replace raw verifier history with explicit payload fields: `original_query`, `executor_final_answer`, `tool_trace_summary`, `retry_context`.
|
||||
- [x] 2.3 Persist the current round `tool_trace_summary` in `VerifierContextHolder` for later verifier evaluation storage.
|
||||
|
||||
## 3. ChatService Integration
|
||||
|
||||
- [x] 3.1 Load `chatVerifierPrompt` and add `buildChatVerifierAgent()`.
|
||||
- [x] 3.2 Add configurable `verifier.low-confidence-threshold`.
|
||||
- [x] 3.3 Implement explicit per-round orchestration in `ChatService`: planner call, executor call, verifier call.
|
||||
- [x] 3.4 Keep max two outer rounds and inject `retry_context` only for the second round.
|
||||
- [x] 3.5 Parse verifier JSON directly and fall back to `LOW_CONFID` when verifier output is missing or invalid.
|
||||
- [x] 3.6 Keep `SupervisorAgent` construction as legacy residue only; runtime orchestration no longer depends on prompt-only supervisor sequencing.
|
||||
|
||||
## 4. Verdict Routing And User Output
|
||||
|
||||
- [x] 4.1 Route `PASS` to the executor answer.
|
||||
- [x] 4.2 Route `LOW_CONFID` to a fixed disclaimer plus executor answer.
|
||||
- [x] 4.3 Route `REJECT` to degraded output without passing through the raw unverified answer.
|
||||
- [x] 4.4 Build LOW_CONFID gap lists only from verifier-identified gaps.
|
||||
- [x] 4.5 Build DEGRADED confirmed facts, gaps, and next-step suggestions from verifier facts and trace summary.
|
||||
|
||||
## 5. Trace Summary And Observability
|
||||
|
||||
- [x] 5.1 Add `ToolTraceSummaryService` to build verifier evidence summaries from `tool_invocation`.
|
||||
- [x] 5.2 Include only evidence-bearing tools by default.
|
||||
- [x] 5.3 Compact repeated calls by tool and topic domain.
|
||||
- [x] 5.4 Preserve `source_invocation_ids`, `trace_ref`, query samples, retrieval layers, relevance levels, and source document labels.
|
||||
- [x] 5.5 Parse and persist `facts_checked[*].evidence_refs`.
|
||||
- [x] 5.6 Persist `verifier_evaluation.tool_trace_summary` and `traceability_version`.
|
||||
- [x] 5.7 Store concise verifier summaries in `agent_step.thought` while preserving fuller verifier output in `model_output` / `self_evaluation`.
|
||||
|
||||
## 6. self_evaluation Merge Semantics
|
||||
|
||||
- [x] 6.1 Add `SelfEvaluationMergeService`.
|
||||
- [x] 6.2 Write verifier results under `verifier_evaluation`.
|
||||
- [x] 6.3 Write rule scoring under `rule_evaluation`.
|
||||
- [x] 6.4 Preserve the other channel with read-modify-write semantics.
|
||||
|
||||
## 7. Verification
|
||||
|
||||
- [x] 7.1 Compile verification: `mvn -q -DskipTests compile`.
|
||||
- [x] 7.2 Runtime verification: `/api/chat` complex request reached `planner -> executor -> verifier`.
|
||||
- [x] 7.3 Runtime verification: session `9138f064` persisted `verifier_evaluation.facts_checked[*].evidence_refs`.
|
||||
- [x] 7.4 Runtime verification: session `9138f064` persisted `tool_trace_summary[*].source_invocation_ids`.
|
||||
- [x] 7.5 Runtime verification: LOW_CONFID user output included disclaimer and verifier-derived gaps.
|
||||
@@ -0,0 +1 @@
|
||||
mvp-demo-trace-acceptance committed on 2026-07-03
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-03
|
||||
@@ -0,0 +1,29 @@
|
||||
{
|
||||
"id": "mvp-demo-trace-acceptance",
|
||||
"metadata": {
|
||||
"status": "committed",
|
||||
"created_at": "2026-07-03",
|
||||
"updated_at": "2026-07-03",
|
||||
"implementation_status": "implemented"
|
||||
},
|
||||
"summary": "Add an MVP demo profile, a read-only diagnosis trace API, and an end-to-end acceptance case.",
|
||||
"artifacts": {
|
||||
"proposal": "proposal.md",
|
||||
"design": "design.md",
|
||||
"tasks": "tasks.md",
|
||||
"specs": [
|
||||
"specs/mvp-demo-trace-acceptance/spec.md"
|
||||
],
|
||||
"devflow": "devflow/projects/2026-07-03-mvp-demo-trace-acceptance"
|
||||
},
|
||||
"tasks": [
|
||||
"Add DiagnosisTraceResponse DTO",
|
||||
"Add DiagnosisTraceService aggregation",
|
||||
"Add DiagnosisTraceController endpoint",
|
||||
"Add mvp-demo profile",
|
||||
"Add MVP demo acceptance documentation",
|
||||
"Add focused trace service tests",
|
||||
"Run targeted verification and GitNexus change detection",
|
||||
"Update MVP notes and devflow acceptance"
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,70 @@
|
||||
## Context
|
||||
|
||||
The MVP already persists diagnosis execution data across three tables:
|
||||
|
||||
- `diagnosis_session`: query, status, answer, counts, feedback, and `self_evaluation`.
|
||||
- `agent_step`: ordered agent execution records.
|
||||
- `tool_invocation`: evidence tool calls and retrieval metadata.
|
||||
|
||||
Recent work unified chat session ids and persisted tool invocations, so a single session id can now connect user input, agent steps, evidence tools, verifier evaluation, final answer, and feedback. The missing piece is a read-only aggregation API and a documented demo profile/workflow that a reviewer can run without reading database tables manually.
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
|
||||
- Add a trace API that returns one aggregated view for a diagnosis session.
|
||||
- Keep the trace API read-only and based on existing persistence tables.
|
||||
- Add an `mvp-demo` profile that makes the demo intent explicit and keeps mock log/metric tools enabled.
|
||||
- Add a documented end-to-end acceptance case for start, chat, trace query, and feedback.
|
||||
- Add focused tests for trace aggregation.
|
||||
|
||||
**Non-Goals:**
|
||||
|
||||
- Do not clean up committed sensitive configuration in this change.
|
||||
- Do not add database migrations.
|
||||
- Do not alter `/api/chat`, `/api/chat_stream`, verifier routing, feedback, or document upload behavior.
|
||||
- Do not create a fully offline fake LLM runtime.
|
||||
|
||||
## Decisions
|
||||
|
||||
| Decision | Choice | Alternative Considered | Rationale |
|
||||
|---|---|---|---|
|
||||
| Trace API shape | Add `GET /api/diagnosis/{sessionId}/trace` | Extend `/api/chat` response | Trace is an observability concern and should not make chat responses larger or change chat clients. |
|
||||
| Aggregation ownership | New `DiagnosisTraceService` | Put aggregation in controller | Keeps controller thin and allows focused unit tests with mocked repositories. |
|
||||
| Response DTO | Dedicated nested DTO | Return raw entities or maps | DTO avoids leaking JPA entity details and gives a stable demo-facing contract. |
|
||||
| Missing session handling | Throw `SessionNotFoundException` and use existing global 404 handler | Return empty success payload | A missing trace is a real lookup miss and should be visible to callers. |
|
||||
| `self_evaluation` handling | Return raw JSON string and best-effort parsed JSON | Parse only, or ignore parse failures | Raw value preserves evidence even if JSON shape evolves; parsed value improves frontend/demo readability. |
|
||||
| Demo profile | Add `application-mvp-demo.yml` overlay | Change default `application.yml` | Overlay avoids disturbing current runtime and keeps demo choices explicit. |
|
||||
|
||||
## Interface Impact
|
||||
|
||||
- Level: L3 collaboration API.
|
||||
- Reason: This adds a new HTTP endpoint and response contract intended for frontend/demo/reviewer consumption.
|
||||
- Compatibility: Additive only. Existing callers do not need to change.
|
||||
- Documentation: The endpoint is documented in the MVP demo acceptance case.
|
||||
|
||||
## Data Structures
|
||||
|
||||
The trace response contains:
|
||||
|
||||
- `session`: session id, query, status, flow, counts, timing, created/updated time, final answer, raw self-evaluation JSON, parsed self-evaluation object, and feedback.
|
||||
- `steps`: ordered agent steps with step index, agent name, model input/output, thought, tool flag, duration, token count, and created time.
|
||||
- `toolInvocations`: ordered tool records with id, step id, tool name, input params, output preview, retrieval metadata, duration, success, error, and created time.
|
||||
- `summary`: counts derived from the returned collections and session fields.
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Trace responses may become large for long sessions. -> Mitigation: the MVP returns persisted previews and structured metadata, not raw full external logs.
|
||||
- [Risk] `self_evaluation` JSON shape may evolve. -> Mitigation: return both raw and best-effort parsed forms.
|
||||
- [Risk] Demo profile still depends on real DB/Redis/Milvus/LLM. -> Mitigation: document prerequisites and keep mock logs/metrics enabled for repeatable tool evidence.
|
||||
- [Risk] New endpoint becomes a de facto frontend contract. -> Mitigation: use a dedicated DTO and document L3 additive API impact.
|
||||
|
||||
## Migration Plan
|
||||
|
||||
- Deploying this change requires only application restart with the new code.
|
||||
- No database migration is required.
|
||||
- Rollback is deleting the new endpoint/profile/docs; persisted data remains unchanged.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- None for this slice. Security and full offline test profile remain deferred by explicit user decision.
|
||||
@@ -0,0 +1,29 @@
|
||||
## Why
|
||||
|
||||
The MVP can already execute multi-agent diagnosis, persist session traces, and collect feedback, but it is still hard to demonstrate as a complete enterprise-style workflow. A demo profile, a trace query API, and an explicit end-to-end acceptance case make the project runnable, observable, and explainable for interview and portfolio review.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Add an `mvp-demo` Spring profile that keeps the existing external infrastructure contract but turns on mock log and metric providers for repeatable demonstrations.
|
||||
- Add a read-only trace query API: `GET /api/diagnosis/{sessionId}/trace`.
|
||||
- Aggregate `diagnosis_session`, `agent_step`, `tool_invocation`, verifier/self-evaluation, final answer, and feedback into one trace response.
|
||||
- Add an end-to-end MVP acceptance case that documents startup, chat request, trace query, and feedback submission.
|
||||
- Add focused service tests for trace aggregation without requiring MySQL, Redis, Milvus, or a real LLM.
|
||||
- Record the design decision in MVP notes for interview storytelling.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `mvp-demo-trace-acceptance`: Covers the MVP demo profile, trace query API, and end-to-end acceptance workflow for a reproducible agent diagnosis demo.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- None.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected code: new trace controller/service/DTOs, `application-mvp-demo.yml`, unit tests, MVP demo documentation.
|
||||
- Affected API: adds `GET /api/diagnosis/{sessionId}/trace`. This is an additive L3 collaboration API because it is intended for frontend, demo, and external reviewer consumption.
|
||||
- Affected runtime behavior: no change to chat execution, verifier, feedback, document upload, or persistence semantics.
|
||||
- Non-goals: no sensitive configuration cleanup, no database schema migration, no replacement of existing chat endpoints, no full offline mock LLM implementation.
|
||||
+33
@@ -0,0 +1,33 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Diagnosis trace can be queried by session id
|
||||
The system SHALL expose a read-only HTTP endpoint `GET /api/diagnosis/{sessionId}/trace` that returns the persisted diagnosis trace for the requested session id.
|
||||
|
||||
#### Scenario: Existing session trace is returned
|
||||
- **WHEN** a caller requests trace data for a session id that exists in `diagnosis_session`
|
||||
- **THEN** the system returns a success response containing the session summary, ordered agent steps, ordered tool invocations, self-evaluation data, final answer, and feedback
|
||||
|
||||
#### Scenario: Missing session returns not found
|
||||
- **WHEN** a caller requests trace data for a session id that does not exist in `diagnosis_session`
|
||||
- **THEN** the system returns a 404 response using the existing session-not-found error contract
|
||||
|
||||
### Requirement: Trace aggregation is read-only
|
||||
The system MUST build trace output from existing persisted diagnosis tables and MUST NOT mutate diagnosis sessions, agent steps, tool invocations, feedback, or chat session state while serving the trace request.
|
||||
|
||||
#### Scenario: Trace query does not change persisted state
|
||||
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace`
|
||||
- **THEN** the system reads `diagnosis_session`, `agent_step`, and `tool_invocation` records and returns an aggregate without saving any of those records
|
||||
|
||||
### Requirement: MVP demo profile is available
|
||||
The system SHALL provide an `mvp-demo` Spring profile that documents the demo runtime intent and keeps mock log and metric providers enabled for repeatable diagnosis demonstrations.
|
||||
|
||||
#### Scenario: Demo profile loads mock evidence providers
|
||||
- **WHEN** the application starts with `--spring.profiles.active=mvp-demo`
|
||||
- **THEN** `prometheus.mock-enabled` and `cls.mock-enabled` are enabled by profile configuration
|
||||
|
||||
### Requirement: End-to-end MVP acceptance case is documented
|
||||
The project SHALL include an end-to-end acceptance case that demonstrates start-up, chat diagnosis, trace query, and feedback submission using the same session id.
|
||||
|
||||
#### Scenario: Reviewer follows the acceptance case
|
||||
- **WHEN** a reviewer follows the documented MVP demo acceptance steps
|
||||
- **THEN** they can run the application, submit a diagnosis question, query the trace endpoint, and submit feedback for the same session id
|
||||
@@ -0,0 +1,25 @@
|
||||
## 1. Trace Query API
|
||||
|
||||
- [x] 1.1 Add a `DiagnosisTraceResponse` DTO that represents session summary, ordered agent steps, ordered tool invocations, and derived summary counts.
|
||||
- [x] 1.2 Add `DiagnosisTraceService` that loads `DiagnosisSession`, `AgentStep`, and `ToolInvocation` records by session id and builds the response.
|
||||
- [x] 1.3 Add `DiagnosisTraceController` with `GET /api/diagnosis/{sessionId}/trace`.
|
||||
- [x] 1.4 Return 404 through `SessionNotFoundException` when the requested diagnosis session does not exist.
|
||||
|
||||
## 2. Demo Profile And Acceptance Case
|
||||
|
||||
- [x] 2.1 Add `src/main/resources/application-mvp-demo.yml` with MVP demo profile overlays and mock logs/metrics enabled.
|
||||
- [x] 2.2 Add `mvp/demo/README.md` documenting prerequisites, startup, chat request, trace query, and feedback submission.
|
||||
- [x] 2.3 Add a concrete payment-timeout acceptance case with request/response expectations.
|
||||
|
||||
## 3. Tests And Verification
|
||||
|
||||
- [x] 3.1 Add focused unit tests for `DiagnosisTraceService` success and missing-session behavior.
|
||||
- [x] 3.2 Run targeted tests for the new trace service.
|
||||
- [x] 3.3 Run compile verification.
|
||||
- [x] 3.4 Run GitNexus change detection before commit or handoff.
|
||||
|
||||
## 4. Notes And Flow Records
|
||||
|
||||
- [x] 4.1 Update MVP engineering notes with the demo/trace decision.
|
||||
- [x] 4.2 Update OpenSpec tasks as work completes.
|
||||
- [x] 4.3 Record verification results in devflow acceptance notes.
|
||||
@@ -0,0 +1,192 @@
|
||||
# chat-verifier-agent Specification
|
||||
|
||||
## Purpose
|
||||
TBD - created by archiving change chat-verifier-agent. Update Purpose after archive.
|
||||
## Requirements
|
||||
### Requirement: Verifier SHALL fact-check Executor answers
|
||||
The system SHALL have a Verifier Agent that reads the Executor's answer and the tool call history, then produces a structured verdict.
|
||||
|
||||
#### Scenario: PASS verdict when all claims have evidence
|
||||
- **WHEN** all critical facts in the Executor's answer have direct or indirect support in tool call results
|
||||
- **AND** at least one critical fact has direct evidence
|
||||
- **AND** no critical fact is contradicted
|
||||
- **THEN** the Verifier SHALL output verdict="PASS" with groundedness_score ≥ 0.5
|
||||
|
||||
#### Scenario: LOW_CONFID verdict with partial evidence
|
||||
- **WHEN** no critical fact contradicts the tool results
|
||||
- **AND** some critical facts have no supporting evidence
|
||||
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
|
||||
|
||||
#### Scenario: LOW_CONFID verdict with only indirect support
|
||||
- **WHEN** no critical fact contradicts the tool results
|
||||
- **AND** all critical facts are only indirectly supported
|
||||
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
|
||||
|
||||
#### Scenario: REJECT verdict when claims contradict evidence
|
||||
- **WHEN** any critical fact in the Executor's answer contradicts tool call results
|
||||
- **OR** the answer fabricates a key entity, error code, or conclusion that does not exist in the tool evidence
|
||||
- **THEN** the Verifier SHALL output verdict="REJECT"
|
||||
|
||||
### Requirement: Verifier SHALL output structured JSON
|
||||
The Verifier SHALL output a JSON object with verdict, groundedness_score, facts_checked array, and rationale.
|
||||
|
||||
#### Scenario: Output format validation
|
||||
- **WHEN** the Verifier completes its analysis
|
||||
- **THEN** the output SHALL contain "verdict", "groundedness_score", "facts_checked", and "rationale" fields
|
||||
- **AND** groundedness_score SHALL be a float between 0.0 and 1.0
|
||||
- **AND** verdict SHALL be one of "PASS", "LOW_CONFID", or "REJECT"
|
||||
|
||||
#### Scenario: strict schema output
|
||||
- **WHEN** the Verifier returns its result
|
||||
- **THEN** it SHALL output exactly one JSON object
|
||||
- **AND** it SHALL NOT output Markdown, code fences, or explanatory text outside the JSON object
|
||||
- **AND** the JSON object SHALL include `critical_fact_count`
|
||||
- **AND** each `facts_checked` item SHALL include `fact`, `is_critical`, `verification`, and `detail`
|
||||
|
||||
### Requirement: facts_checked SHALL use a fixed classification set
|
||||
Each checked fact SHALL be labeled using a fixed evidence classification.
|
||||
|
||||
#### Scenario: fact classification values
|
||||
- **WHEN** the Verifier emits `facts_checked`
|
||||
- **THEN** each fact SHALL use one of `direct_evidence`, `indirect_support`, `no_evidence`, or `contradicted`
|
||||
|
||||
### Requirement: groundedness_score SHALL be derived from fact classifications
|
||||
The groundedness score SHALL be computed from critical fact classifications instead of being freely chosen by the model.
|
||||
|
||||
#### Scenario: contradicted fact forces reject
|
||||
- **WHEN** any critical fact is labeled `contradicted`
|
||||
- **THEN** the Verifier SHALL output verdict="REJECT"
|
||||
- **AND** groundedness_score SHALL be `0.0`
|
||||
|
||||
#### Scenario: score derived from supported facts
|
||||
- **WHEN** no critical fact is contradicted
|
||||
- **THEN** groundedness_score SHALL be computed from the mapped values of critical facts
|
||||
- **AND** the implementation SHALL use the fixed mapping `direct_evidence=1.0`, `indirect_support=0.6`, `no_evidence=0.0`
|
||||
- **AND** the result SHALL be clamped into `[0.0, 1.0]`
|
||||
|
||||
### Requirement: ChatService SHALL route based on Verifier verdict
|
||||
The system SHALL use ChatService for explicit single-round `Planner → Executor → Verifier` orchestration and SHALL use ChatService to control whether an additional round is allowed.
|
||||
|
||||
#### Scenario: PASS → direct output
|
||||
- **WHEN** Verifier outputs verdict="PASS"
|
||||
- **THEN** the system SHALL output the Executor's answer directly
|
||||
|
||||
#### Scenario: LOW_CONFID score≥0.5 → output with disclaimer
|
||||
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score ≥ 0.5
|
||||
- **THEN** the system SHALL output the Executor's answer prefixed with a fixed confidence disclaimer
|
||||
|
||||
#### Scenario: LOW_CONFID score<0.5 → trigger one additional round
|
||||
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score < 0.5 and this is the first callback
|
||||
- **THEN** the ChatService SHALL invoke one additional `Planner → Executor → Verifier` round to supplement evidence
|
||||
- **AND** after the second Verifier run, verdict="LOW_CONFID" SHALL be output with a confidence disclaimer
|
||||
- **AND** after the second Verifier run, verdict="REJECT" SHALL still produce a degraded output
|
||||
|
||||
#### Scenario: REJECT does not enter retry round
|
||||
- **WHEN** Verifier outputs verdict="REJECT"
|
||||
- **THEN** the system SHALL NOT start a retry round for evidence补充
|
||||
- **AND** it SHALL produce a degraded output directly
|
||||
|
||||
#### Scenario: REJECT → degraded output
|
||||
- **WHEN** Verifier outputs verdict="REJECT"
|
||||
- **THEN** the system SHALL output a degraded result indicating the answer cannot be reliably generated
|
||||
- **AND** it SHALL NOT pass through the raw Executor answer
|
||||
|
||||
### Requirement: User-facing verifier outputs SHALL follow fixed templates
|
||||
The system SHALL use fixed output protocols for LOW_CONFID and REJECT user-facing responses.
|
||||
|
||||
#### Scenario: LOW_CONFID uses disclaimer template
|
||||
- **WHEN** the final verdict is `LOW_CONFID`
|
||||
- **THEN** the user-facing response SHALL prepend a fixed disclaimer before the Executor answer
|
||||
- **AND** optional evidence gaps, if present, SHALL come only from verifier-identified critical gaps
|
||||
|
||||
#### Scenario: REJECT uses degraded template
|
||||
- **WHEN** the final verdict is `REJECT`
|
||||
- **THEN** the user-facing response SHALL use a degraded template
|
||||
- **AND** it SHALL include only confirmed facts, evidence gaps, and next-step suggestions
|
||||
- **AND** it SHALL NOT include unverified raw answer content
|
||||
|
||||
### Requirement: Verifier SHALL be observable
|
||||
The Verifier's verdict SHALL be persisted for observability.
|
||||
|
||||
#### Scenario: verdict written to self_evaluation
|
||||
- **WHEN** the Verifier produces a verdict
|
||||
- **THEN** the ChatService SHALL write the verdict data under `diagnosis_session.self_evaluation.verifier_evaluation`
|
||||
- **AND** existing `rule_evaluation` data SHALL be preserved
|
||||
|
||||
### Requirement: self_evaluation SHALL be a container object
|
||||
The `diagnosis_session.self_evaluation` field SHALL store multiple evaluation channels in one JSON object.
|
||||
|
||||
#### Scenario: rule evaluation stored separately
|
||||
- **WHEN** the rule-based evidence scoring completes
|
||||
- **THEN** the EvaluationService SHALL write the result under `rule_evaluation`
|
||||
- **AND** existing `verifier_evaluation` data SHALL be preserved
|
||||
|
||||
#### Scenario: verifier evaluation stored separately
|
||||
- **WHEN** the Verifier completes
|
||||
- **THEN** the ChatService SHALL write the result under `verifier_evaluation`
|
||||
- **AND** existing `rule_evaluation` data SHALL be preserved
|
||||
|
||||
#### Scenario: no whole-object overwrite after initialization
|
||||
- **WHEN** either evaluation channel updates `self_evaluation`
|
||||
- **THEN** the implementation SHALL use read-modify-write semantics
|
||||
- **AND** it SHALL NOT replace the whole JSON object except when initializing from null
|
||||
|
||||
### Requirement: Verifier SHALL consume explicit verification inputs
|
||||
The Verifier SHALL receive explicit verification inputs rather than inferring them only from raw conversation history.
|
||||
|
||||
#### Scenario: explicit input blocks available to Verifier
|
||||
- **WHEN** the Verifier starts
|
||||
- **THEN** the system SHALL provide `original_query`, `executor_final_answer`, and `tool_trace_summary` as explicit inputs
|
||||
- **AND** `retry_context` SHALL be provided on the second round only
|
||||
- **AND** message filtering MAY be used only to remove intermediate reasoning or unrelated noise
|
||||
|
||||
#### Scenario: tool trace summary derived from tool facts
|
||||
- **WHEN** the system prepares verifier inputs
|
||||
- **THEN** `tool_trace_summary` SHALL be generated from tool invocation facts
|
||||
- **AND** each summary item SHALL include tool name, success state, input summary, output summary, and evidence level
|
||||
- **AND** raw conversation history SHALL NOT be the only source of verifier evidence context
|
||||
|
||||
#### Scenario: tool trace summary preserves invocation references
|
||||
- **WHEN** the system prepares verifier inputs
|
||||
- **THEN** each summary item SHALL include a stable `trace_ref`
|
||||
- **AND** each summary item SHALL preserve `source_invocation_ids` for the tool invocation rows that contributed to the summary
|
||||
- **AND** each summary item SHOULD include query samples, retrieval layers, relevance levels, and source document labels when available
|
||||
|
||||
#### Scenario: only evidence-bearing tools included
|
||||
- **WHEN** the system generates `tool_trace_summary`
|
||||
- **THEN** it SHALL include only evidence-bearing tool invocations
|
||||
- **AND** non-evidence helper tools such as time or formatting tools SHALL be excluded by default
|
||||
|
||||
#### Scenario: failed evidence calls preserved as evidence gaps
|
||||
- **WHEN** an evidence-bearing tool invocation fails or returns no usable evidence
|
||||
- **THEN** the summary SHALL still include that invocation
|
||||
- **AND** it SHALL mark the entry as unsuccessful with an evidence level representing no evidence
|
||||
|
||||
#### Scenario: repeated tool calls may be compacted
|
||||
- **WHEN** repeated tool invocations concern the same tool, topic domain, and round
|
||||
- **THEN** the system MAY compact them into a merged summary entry
|
||||
- **AND** the merged entry SHALL preserve the first effective hit and the count of repeated, failed, or no-hit calls
|
||||
|
||||
#### Scenario: raw outputs not passed through in full
|
||||
- **WHEN** a tool invocation returns large raw content
|
||||
- **THEN** `tool_trace_summary` SHALL keep only a minimal evidence summary
|
||||
- **AND** the raw output SHALL NOT be passed through in full to the Verifier
|
||||
|
||||
#### Scenario: MessagesModelHook used only for noise reduction
|
||||
- **WHEN** a MessagesModelHook is used for the Verifier
|
||||
- **THEN** it MAY remove intermediate reasoning or irrelevant messages
|
||||
- **AND** it SHALL NOT be the primary source for assembling verifier business inputs
|
||||
|
||||
### Requirement: Verifier facts SHALL be auditable
|
||||
Verifier facts SHALL be linkable to the evidence summaries used during verification.
|
||||
|
||||
#### Scenario: facts_checked contains evidence refs
|
||||
- **WHEN** the Verifier emits `facts_checked`
|
||||
- **THEN** each fact SHALL include `evidence_refs`
|
||||
- **AND** each evidence ref SHALL point to an existing `tool_trace_summary.trace_ref`
|
||||
- **AND** each evidence ref SHALL preserve the relevant `source_invocation_ids` when available
|
||||
|
||||
#### Scenario: verifier evaluation persists traceability snapshot
|
||||
- **WHEN** the ChatService persists `verifier_evaluation`
|
||||
- **THEN** it SHALL include `traceability_version`
|
||||
- **AND** it SHALL include the `tool_trace_summary` snapshot used by the Verifier
|
||||
@@ -0,0 +1,37 @@
|
||||
## Purpose
|
||||
|
||||
Provide a repeatable MVP demo flow that can run a chat diagnosis, expose its persisted execution trace, and submit feedback for the same session id.
|
||||
|
||||
## Requirements
|
||||
|
||||
### Requirement: Diagnosis trace can be queried by session id
|
||||
The system SHALL expose a read-only HTTP endpoint `GET /api/diagnosis/{sessionId}/trace` that returns the persisted diagnosis trace for the requested session id.
|
||||
|
||||
#### Scenario: Existing session trace is returned
|
||||
- **WHEN** a caller requests trace data for a session id that exists in `diagnosis_session`
|
||||
- **THEN** the system returns a success response containing the session summary, ordered agent steps, ordered tool invocations, self-evaluation data, final answer, and feedback
|
||||
|
||||
#### Scenario: Missing session returns not found
|
||||
- **WHEN** a caller requests trace data for a session id that does not exist in `diagnosis_session`
|
||||
- **THEN** the system returns a 404 response using the existing session-not-found error contract
|
||||
|
||||
### Requirement: Trace aggregation is read-only
|
||||
The system MUST build trace output from existing persisted diagnosis tables and MUST NOT mutate diagnosis sessions, agent steps, tool invocations, feedback, or chat session state while serving the trace request.
|
||||
|
||||
#### Scenario: Trace query does not change persisted state
|
||||
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace`
|
||||
- **THEN** the system reads `diagnosis_session`, `agent_step`, and `tool_invocation` records and returns an aggregate without saving any of those records
|
||||
|
||||
### Requirement: MVP demo profile is available
|
||||
The system SHALL provide an `mvp-demo` Spring profile that documents the demo runtime intent and keeps mock log and metric providers enabled for repeatable diagnosis demonstrations.
|
||||
|
||||
#### Scenario: Demo profile loads mock evidence providers
|
||||
- **WHEN** the application starts with `--spring.profiles.active=mvp-demo`
|
||||
- **THEN** `prometheus.mock-enabled` and `cls.mock-enabled` are enabled by profile configuration
|
||||
|
||||
### Requirement: End-to-end MVP acceptance case is documented
|
||||
The project SHALL include an end-to-end acceptance case that demonstrates start-up, chat diagnosis, trace query, and feedback submission using the same session id.
|
||||
|
||||
#### Scenario: Reviewer follows the acceptance case
|
||||
- **WHEN** a reviewer follows the documented MVP demo acceptance steps
|
||||
- **THEN** they can run the application, submit a diagnosis question, query the trace endpoint, and submit feedback for the same session id
|
||||
@@ -0,0 +1,82 @@
|
||||
#!/usr/bin/env python3
|
||||
# -*- coding: utf-8 -*-
|
||||
"""
|
||||
通用 MySQL 查询脚本
|
||||
用法:
|
||||
python scripts/query_mysql.py "SELECT * FROM diagnosis_session ORDER BY created_at DESC LIMIT 5"
|
||||
python scripts/query_mysql.py # 交互模式
|
||||
依赖:pip install pymysql
|
||||
"""
|
||||
|
||||
import sys
|
||||
import os
|
||||
|
||||
# Windows 控制台 UTF-8 输出
|
||||
if sys.stdout.encoding and sys.stdout.encoding.lower() != 'utf-8':
|
||||
sys.stdout.reconfigure(encoding='utf-8', errors='replace')
|
||||
|
||||
try:
|
||||
import pymysql
|
||||
import pymysql.cursors
|
||||
except ImportError:
|
||||
print("缺少依赖,请先执行: pip install pymysql")
|
||||
sys.exit(1)
|
||||
|
||||
# 从 application.yml 读取的连接信息
|
||||
DB_CONFIG = {
|
||||
"host": "119.29.78.52",
|
||||
"port": 33306,
|
||||
"user": "root",
|
||||
"password": "!Fucker123..",
|
||||
"database": "superbiz_agent",
|
||||
"charset": "utf8mb4",
|
||||
"cursorclass": pymysql.cursors.DictCursor,
|
||||
}
|
||||
|
||||
|
||||
def run_query(sql: str):
|
||||
conn = pymysql.connect(**DB_CONFIG)
|
||||
try:
|
||||
with conn.cursor() as cur:
|
||||
cur.execute(sql)
|
||||
if sql.strip().upper().startswith("SELECT") or sql.strip().upper().startswith("SHOW"):
|
||||
rows = cur.fetchall()
|
||||
if not rows:
|
||||
print("(空结果)")
|
||||
return
|
||||
# 打印列头
|
||||
cols = list(rows[0].keys())
|
||||
col_widths = {c: max(len(c), max(len(str(r[c])) for r in rows)) for c in cols}
|
||||
header = " | ".join(c.ljust(col_widths[c]) for c in cols)
|
||||
print(header)
|
||||
print("-" * len(header))
|
||||
for row in rows:
|
||||
print(" | ".join(str(row[c]).ljust(col_widths[c]) for c in cols))
|
||||
print(f"\n({len(rows)} 行)")
|
||||
else:
|
||||
conn.commit()
|
||||
print(f"OK,影响行数: {cur.rowcount}")
|
||||
finally:
|
||||
conn.close()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
if len(sys.argv) > 1:
|
||||
sql = " ".join(sys.argv[1:])
|
||||
run_query(sql)
|
||||
else:
|
||||
print("MySQL 交互模式(输入 exit 退出)")
|
||||
print(f"连接:{DB_CONFIG['user']}@{DB_CONFIG['host']}:{DB_CONFIG['port']}/{DB_CONFIG['database']}")
|
||||
print("-" * 50)
|
||||
while True:
|
||||
try:
|
||||
sql = input("sql> ").strip()
|
||||
if sql.lower() in ("exit", "quit", "q"):
|
||||
break
|
||||
if not sql:
|
||||
continue
|
||||
run_query(sql)
|
||||
except KeyboardInterrupt:
|
||||
break
|
||||
except Exception as e:
|
||||
print(f"错误: {e}")
|
||||
@@ -2,6 +2,7 @@ package com.superbiz.agent.agent.tool;
|
||||
|
||||
import com.fasterxml.jackson.annotation.JsonProperty;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.service.ToolInvocationRecorder;
|
||||
import lombok.Data;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
@@ -34,6 +35,11 @@ public class QueryLogsTools {
|
||||
public static final String TOOL_GET_AVAILABLE_LOG_TOPICS = "getAvailableLogTopics";
|
||||
|
||||
private final ObjectMapper objectMapper = new ObjectMapper();
|
||||
private final ToolInvocationRecorder toolInvocationRecorder;
|
||||
|
||||
public QueryLogsTools(ToolInvocationRecorder toolInvocationRecorder) {
|
||||
this.toolInvocationRecorder = toolInvocationRecorder;
|
||||
}
|
||||
|
||||
@Value("${cls.mock-enabled:false}")
|
||||
private boolean mockEnabled;
|
||||
@@ -55,6 +61,7 @@ public class QueryLogsTools {
|
||||
"Call this tool first before querying logs to understand what log topics are available. " +
|
||||
"Returns a list of log topics with their names, descriptions, and example queries.")
|
||||
public String getAvailableLogTopics() {
|
||||
long startTime = System.currentTimeMillis();
|
||||
logger.info("获取可用的日志主题列表");
|
||||
|
||||
try {
|
||||
@@ -123,11 +130,15 @@ public class QueryLogsTools {
|
||||
|
||||
output.setMessage(String.format("共有 %d 个可用的日志主题。建议使用默认地域 'ap-guangzhou' 或省略 region 参数", topics.size()));
|
||||
|
||||
return objectMapper.writerWithDefaultPrettyPrinter().writeValueAsString(output);
|
||||
String response = objectMapper.writerWithDefaultPrettyPrinter().writeValueAsString(output);
|
||||
recordInvocation(startTime, "get_available_log_topics", null, null, null, response, true, null, "logs");
|
||||
return response;
|
||||
|
||||
} catch (Exception e) {
|
||||
logger.error("获取日志主题列表失败", e);
|
||||
return "{\"success\":false,\"message\":\"获取日志主题列表失败: " + e.getMessage() + "\"}";
|
||||
String response = "{\"success\":false,\"message\":\"获取日志主题列表失败: " + e.getMessage() + "\"}";
|
||||
recordInvocation(startTime, "get_available_log_topics", null, null, null, response, false, e.getMessage(), "logs");
|
||||
return response;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -164,6 +175,7 @@ public class QueryLogsTools {
|
||||
@ToolParam(description = "查询条件,支持 Lucene 语法,如 level:ERROR OR cpu_usage:>80;为空时返回该主题近 5 条核心日志") String query,
|
||||
@ToolParam(description = "返回日志条数,默认20,最大100") Integer limit) {
|
||||
|
||||
long startTime = System.currentTimeMillis();
|
||||
int actualLimit = (limit == null || limit <= 0) ? 20 : Math.min(limit, 100);
|
||||
|
||||
String safeQuery = query == null ? "" : query;
|
||||
@@ -178,7 +190,10 @@ public class QueryLogsTools {
|
||||
logger.info("使用 Mock 数据,返回 {} 条日志", logEntries.size());
|
||||
} else {
|
||||
// 真实模式:调用 CLS API(这里预留接口,后续实现)
|
||||
return buildErrorResponse("CLS 真实查询尚未实现,请启用 mock 模式进行测试");
|
||||
String response = buildErrorResponse("CLS 真实查询尚未实现,请启用 mock 模式进行测试");
|
||||
recordInvocation(startTime, safeQuery, region, logTopic, actualLimit, response, false,
|
||||
"CLS 真实查询尚未实现,请启用 mock 模式进行测试", normalizeTopicDomain(logTopic));
|
||||
return response;
|
||||
}
|
||||
|
||||
// 构建成功响应
|
||||
@@ -193,15 +208,51 @@ public class QueryLogsTools {
|
||||
|
||||
String jsonResult = objectMapper.writerWithDefaultPrettyPrinter().writeValueAsString(output);
|
||||
logger.info("日志查询完成: 找到 {} 条日志", logEntries.size());
|
||||
recordInvocation(startTime, safeQuery, region, logTopic, actualLimit, jsonResult,
|
||||
!logEntries.isEmpty(), logEntries.isEmpty() ? "未找到匹配的日志" : null,
|
||||
normalizeTopicDomain(logTopic));
|
||||
|
||||
return jsonResult;
|
||||
|
||||
} catch (Exception e) {
|
||||
logger.error("查询日志失败", e);
|
||||
return buildErrorResponse("查询失败: " + e.getMessage());
|
||||
String response = buildErrorResponse("查询失败: " + e.getMessage());
|
||||
recordInvocation(startTime, safeQuery, region, logTopic, actualLimit, response, false,
|
||||
e.getMessage(), normalizeTopicDomain(logTopic));
|
||||
return response;
|
||||
}
|
||||
}
|
||||
|
||||
private void recordInvocation(long startTime, String query, String region, String logTopic, Integer limit,
|
||||
String output, boolean success, String errorMessage, String topicDomain) {
|
||||
Map<String, Object> input = new HashMap<>();
|
||||
input.put("query", query == null || query.isBlank() ? "DEFAULT_QUERY" : query);
|
||||
if (region != null) {
|
||||
input.put("region", region);
|
||||
}
|
||||
if (logTopic != null) {
|
||||
input.put("log_topic", logTopic);
|
||||
}
|
||||
if (limit != null) {
|
||||
input.put("limit", limit);
|
||||
}
|
||||
input.put("mock_enabled", mockEnabled);
|
||||
|
||||
toolInvocationRecorder.recordEvidenceTool(
|
||||
"query_logs",
|
||||
input,
|
||||
output,
|
||||
success,
|
||||
startTime,
|
||||
errorMessage,
|
||||
topicDomain
|
||||
);
|
||||
}
|
||||
|
||||
private String normalizeTopicDomain(String logTopic) {
|
||||
return logTopic == null || logTopic.isBlank() ? "logs" : logTopic;
|
||||
}
|
||||
|
||||
/**
|
||||
* 构建 Mock 日志数据
|
||||
|
||||
|
||||
@@ -2,6 +2,7 @@ package com.superbiz.agent.agent.tool;
|
||||
|
||||
import com.fasterxml.jackson.annotation.JsonProperty;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.service.ToolInvocationRecorder;
|
||||
import lombok.Data;
|
||||
import okhttp3.OkHttpClient;
|
||||
import okhttp3.Request;
|
||||
@@ -30,6 +31,11 @@ public class QueryMetricsTools {
|
||||
public static final String TOOL_QUERY_PROMETHEUS_ALERTS = "queryPrometheusAlerts";
|
||||
|
||||
private final ObjectMapper objectMapper = new ObjectMapper();
|
||||
private final ToolInvocationRecorder toolInvocationRecorder;
|
||||
|
||||
public QueryMetricsTools(ToolInvocationRecorder toolInvocationRecorder) {
|
||||
this.toolInvocationRecorder = toolInvocationRecorder;
|
||||
}
|
||||
|
||||
@Value("${prometheus.base-url}")
|
||||
private String prometheusBaseUrl;
|
||||
@@ -59,6 +65,7 @@ public class QueryMetricsTools {
|
||||
"This tool retrieves all currently active/firing alerts including their labels, annotations, state, and values. " +
|
||||
"Use this tool when you need to check what alerts are currently firing, investigate alert conditions, or monitor alert status.")
|
||||
public String queryPrometheusAlerts() {
|
||||
long startTime = System.currentTimeMillis();
|
||||
logger.info("开始查询 Prometheus 活动告警, Mock模式: {}", mockEnabled);
|
||||
|
||||
try {
|
||||
@@ -73,7 +80,9 @@ public class QueryMetricsTools {
|
||||
PrometheusAlertsResult result = fetchPrometheusAlerts();
|
||||
|
||||
if (!"success".equals(result.getStatus())) {
|
||||
return buildErrorResponse("Prometheus API 返回非成功状态: " + result.getStatus(), result.getError());
|
||||
String response = buildErrorResponse("Prometheus API 返回非成功状态: " + result.getStatus(), result.getError());
|
||||
recordInvocation(startTime, response, false, result.getError());
|
||||
return response;
|
||||
}
|
||||
|
||||
// 转换为简化格式,对于相同的 alertname,只保留第一个
|
||||
@@ -110,14 +119,29 @@ public class QueryMetricsTools {
|
||||
|
||||
String jsonResult = objectMapper.writerWithDefaultPrettyPrinter().writeValueAsString(output);
|
||||
logger.info("Prometheus 告警查询完成: 找到 {} 个告警", simplifiedAlerts.size());
|
||||
recordInvocation(startTime, jsonResult, true, null);
|
||||
|
||||
return jsonResult;
|
||||
|
||||
} catch (Exception e) {
|
||||
logger.error("查询 Prometheus 告警失败", e);
|
||||
return buildErrorResponse("查询失败", e.getMessage());
|
||||
String response = buildErrorResponse("查询失败", e.getMessage());
|
||||
recordInvocation(startTime, response, false, e.getMessage());
|
||||
return response;
|
||||
}
|
||||
}
|
||||
|
||||
private void recordInvocation(long startTime, String output, boolean success, String errorMessage) {
|
||||
toolInvocationRecorder.recordEvidenceTool(
|
||||
"query_metrics",
|
||||
Map.of("query", "active_prometheus_alerts", "mock_enabled", mockEnabled),
|
||||
output,
|
||||
success,
|
||||
startTime,
|
||||
errorMessage,
|
||||
"prometheus_alerts"
|
||||
);
|
||||
}
|
||||
|
||||
/**
|
||||
* 构建 Mock 告警数据
|
||||
|
||||
@@ -0,0 +1,9 @@
|
||||
package com.superbiz.agent.config;
|
||||
|
||||
import org.springframework.context.annotation.Configuration;
|
||||
import org.springframework.scheduling.annotation.EnableAsync;
|
||||
|
||||
@Configuration
|
||||
@EnableAsync
|
||||
public class AsyncConfig {
|
||||
}
|
||||
@@ -1,32 +1,30 @@
|
||||
package com.superbiz.agent.controller;
|
||||
|
||||
import com.alibaba.cloud.ai.graph.NodeOutput;
|
||||
import com.alibaba.cloud.ai.graph.OverAllState;
|
||||
import com.alibaba.cloud.ai.graph.agent.ReactAgent;
|
||||
import com.alibaba.cloud.ai.graph.streaming.OutputType;
|
||||
import com.alibaba.cloud.ai.graph.streaming.StreamingOutput;
|
||||
import lombok.Getter;
|
||||
import lombok.Setter;
|
||||
import com.superbiz.agent.domain.model.SessionContext;
|
||||
import com.superbiz.agent.service.AiOpsService;
|
||||
import com.superbiz.agent.service.ChatService;
|
||||
import com.superbiz.agent.service.session.SessionManager;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
import org.springframework.ai.chat.model.ChatModel;
|
||||
import org.springframework.ai.tool.ToolCallback;
|
||||
import org.springframework.ai.tool.ToolCallbackProvider;
|
||||
import org.springframework.beans.factory.annotation.Autowired;
|
||||
import org.springframework.beans.factory.annotation.Value;
|
||||
import org.springframework.http.MediaType;
|
||||
import org.springframework.http.ResponseEntity;
|
||||
import org.springframework.web.bind.annotation.*;
|
||||
import org.springframework.web.servlet.mvc.method.annotation.SseEmitter;
|
||||
import reactor.core.publisher.Flux;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.time.LocalDateTime;
|
||||
import java.time.ZoneId;
|
||||
import java.util.*;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
import java.util.concurrent.ExecutorService;
|
||||
import java.util.concurrent.Executors;
|
||||
import java.util.concurrent.locks.ReentrantLock;
|
||||
|
||||
/**
|
||||
* 统一 API 控制器
|
||||
@@ -44,17 +42,20 @@ public class ChatController {
|
||||
@Autowired
|
||||
private ChatService chatService;
|
||||
|
||||
@Autowired
|
||||
private SessionManager sessionManager;
|
||||
|
||||
@Autowired(required = false)
|
||||
private ToolCallbackProvider tools;
|
||||
|
||||
private final ExecutorService executor = Executors.newCachedThreadPool();
|
||||
|
||||
// 存储会话信息
|
||||
private final Map<String, SessionInfo> sessions = new ConcurrentHashMap<>();
|
||||
|
||||
// 最大历史消息窗口大小(成对计算:用户消息+AI回复=1对)
|
||||
private static final int MAX_WINDOW_SIZE = 6;
|
||||
|
||||
@Value("${session.ttl-seconds:3600}")
|
||||
private long sessionTtlSeconds;
|
||||
|
||||
/**
|
||||
* 普通对话接口(支持工具调用)
|
||||
* 与 /chat_react 逻辑一致,但直接返回完整结果而非流式输出
|
||||
@@ -71,10 +72,10 @@ public class ChatController {
|
||||
}
|
||||
|
||||
// 获取或创建会话
|
||||
SessionInfo session = getOrCreateSession(request.getId());
|
||||
SessionContext session = getOrCreateSession(request.getId());
|
||||
|
||||
// 获取历史消息
|
||||
List<Map<String, String>> history = session.getHistory();
|
||||
List<Map<String, String>> history = session.getMessageHistorySnapshot();
|
||||
logger.info("会话历史消息对数: {}", history.size() / 2);
|
||||
|
||||
// 获取注入的 ChatModel
|
||||
@@ -87,15 +88,17 @@ public class ChatController {
|
||||
|
||||
// 根据问题复杂度自动选择单 Agent 或多 Agent
|
||||
logger.info("开始 ReactAgent 对话(支持自动工具调用)");
|
||||
String fullAnswer = chatService.executeChatWithStrategy(chatModel, toolCallbacks,
|
||||
request.getQuestion(), history);
|
||||
|
||||
ChatService.ChatResult result = chatService.executeChatWithStrategy(chatModel, toolCallbacks,
|
||||
request.getQuestion(), history, session.getSessionId());
|
||||
String fullAnswer = result.answer();
|
||||
|
||||
// 更新会话历史
|
||||
session.addMessage(request.getQuestion(), fullAnswer);
|
||||
logger.info("已更新会话历史 - SessionId: {}, 当前消息对数: {}",
|
||||
request.getId(), session.getMessagePairCount());
|
||||
|
||||
return ResponseEntity.ok(ApiResponse.success(ChatResponse.success(fullAnswer)));
|
||||
session.addChatMessagePair(request.getQuestion(), fullAnswer, MAX_WINDOW_SIZE);
|
||||
sessionManager.updateSession(session);
|
||||
logger.info("已更新会话历史 - SessionId: {}, 当前消息对数: {}",
|
||||
session.getSessionId(), session.getMessagePairCount());
|
||||
|
||||
return ResponseEntity.ok(ApiResponse.success(ChatResponse.success(fullAnswer, result.sessionId())));
|
||||
|
||||
} catch (Exception e) {
|
||||
logger.error("对话失败", e);
|
||||
@@ -115,9 +118,11 @@ public class ChatController {
|
||||
return ResponseEntity.ok(ApiResponse.error("会话ID不能为空"));
|
||||
}
|
||||
|
||||
SessionInfo session = sessions.get(request.getId());
|
||||
if (session != null) {
|
||||
session.clearHistory();
|
||||
Optional<SessionContext> session = sessionManager.getSession(request.getId());
|
||||
if (session.isPresent()) {
|
||||
SessionContext context = session.get();
|
||||
context.clearMessageHistory();
|
||||
sessionManager.updateSession(context);
|
||||
return ResponseEntity.ok(ApiResponse.success("会话历史已清空"));
|
||||
} else {
|
||||
return ResponseEntity.ok(ApiResponse.error("会话不存在"));
|
||||
@@ -130,8 +135,8 @@ public class ChatController {
|
||||
}
|
||||
|
||||
/**
|
||||
* ReactAgent 对话接口(SSE 流式模式,支持多轮对话,支持自动工具调用,例如获取当前时间,查询日志,告警等)
|
||||
* 支持 session 管理,保留对话历史
|
||||
* 对话接口(SSE 流式模式)
|
||||
* 与 /chat 使用同一条 ChatService 策略链路,区别仅在于通过 SSE 分块返回最终答案。
|
||||
*/
|
||||
@PostMapping(value = "/chat_stream", produces = "text/event-stream;charset=UTF-8")
|
||||
public SseEmitter chatStream(@RequestBody ChatRequest request) {
|
||||
@@ -154,10 +159,10 @@ public class ChatController {
|
||||
logger.info("收到 ReactAgent 对话请求 - SessionId: {}, Question: {}", request.getId(), request.getQuestion());
|
||||
|
||||
// 获取或创建会话
|
||||
SessionInfo session = getOrCreateSession(request.getId());
|
||||
SessionContext session = getOrCreateSession(request.getId());
|
||||
|
||||
// 获取历史消息
|
||||
List<Map<String, String>> history = session.getHistory();
|
||||
List<Map<String, String>> history = session.getMessageHistorySnapshot();
|
||||
logger.info("ReactAgent 会话历史消息对数: {}", history.size() / 2);
|
||||
|
||||
// 获取注入的 ChatModel
|
||||
@@ -166,92 +171,25 @@ public class ChatController {
|
||||
// 记录可用工具
|
||||
chatService.logAvailableTools();
|
||||
|
||||
logger.info("开始 ReactAgent 流式对话(支持自动工具调用)");
|
||||
|
||||
// 构建系统提示词(包含历史消息)
|
||||
String systemPrompt = chatService.buildSystemPrompt(history);
|
||||
|
||||
// 创建 ReactAgent
|
||||
ReactAgent agent = chatService.createReactAgent(chatModel, systemPrompt);
|
||||
|
||||
// 用于累积完整答案
|
||||
StringBuilder fullAnswerBuilder = new StringBuilder();
|
||||
|
||||
// 使用 agent.stream() 进行流式对话
|
||||
Flux<NodeOutput> stream = agent.stream(request.getQuestion());
|
||||
|
||||
stream.subscribe(
|
||||
output -> {
|
||||
try {
|
||||
// 检查是否为 StreamingOutput 类型
|
||||
if (output instanceof StreamingOutput streamingOutput) {
|
||||
OutputType type = streamingOutput.getOutputType();
|
||||
|
||||
// 处理模型推理的流式输出
|
||||
if (type == OutputType.AGENT_MODEL_STREAMING) {
|
||||
// 流式增量内容,逐步显示
|
||||
String chunk = streamingOutput.message().getText();
|
||||
if (chunk != null && !chunk.isEmpty()) {
|
||||
fullAnswerBuilder.append(chunk);
|
||||
|
||||
// 实时发送到前端
|
||||
emitter.send(SseEmitter.event()
|
||||
.name("message")
|
||||
.data(SseMessage.content(chunk), MediaType.APPLICATION_JSON));
|
||||
|
||||
logger.info("发送流式内容: {}", chunk);
|
||||
}
|
||||
} else if (type == OutputType.AGENT_MODEL_FINISHED) {
|
||||
// 模型推理完成
|
||||
logger.info("模型输出完成");
|
||||
} else if (type == OutputType.AGENT_TOOL_FINISHED) {
|
||||
// 工具调用完成
|
||||
logger.info("工具调用完成: {}", output.node());
|
||||
} else if (type == OutputType.AGENT_HOOK_FINISHED) {
|
||||
// Hook 执行完成
|
||||
logger.debug("Hook 执行完成: {}", output.node());
|
||||
}
|
||||
}
|
||||
} catch (IOException e) {
|
||||
logger.error("发送流式消息失败", e);
|
||||
throw new RuntimeException(e);
|
||||
}
|
||||
},
|
||||
error -> {
|
||||
// 错误处理
|
||||
logger.error("ReactAgent 流式对话失败", error);
|
||||
try {
|
||||
emitter.send(SseEmitter.event()
|
||||
.name("message")
|
||||
.data(SseMessage.error(error.getMessage()), MediaType.APPLICATION_JSON));
|
||||
} catch (IOException ex) {
|
||||
logger.error("发送错误消息失败", ex);
|
||||
}
|
||||
emitter.completeWithError(error);
|
||||
},
|
||||
() -> {
|
||||
// 完成处理
|
||||
try {
|
||||
String fullAnswer = fullAnswerBuilder.toString();
|
||||
logger.info("ReactAgent 流式对话完成 - SessionId: {}, 答案长度: {}",
|
||||
request.getId(), fullAnswer.length());
|
||||
|
||||
// 更新会话历史
|
||||
session.addMessage(request.getQuestion(), fullAnswer);
|
||||
logger.info("已更新会话历史 - SessionId: {}, 当前消息对数: {}",
|
||||
request.getId(), session.getMessagePairCount());
|
||||
|
||||
// 发送完成标记
|
||||
emitter.send(SseEmitter.event()
|
||||
.name("message")
|
||||
.data(SseMessage.done(), MediaType.APPLICATION_JSON));
|
||||
emitter.complete();
|
||||
} catch (IOException e) {
|
||||
logger.error("发送完成消息失败", e);
|
||||
emitter.completeWithError(e);
|
||||
}
|
||||
}
|
||||
);
|
||||
ToolCallback[] toolCallbacks = tools != null ? tools.getToolCallbacks() : new ToolCallback[0];
|
||||
|
||||
logger.info("开始统一 ChatService 对话(SSE 分块返回)");
|
||||
ChatService.ChatResult result = chatService.executeChatWithStrategy(chatModel, toolCallbacks,
|
||||
request.getQuestion(), history, session.getSessionId());
|
||||
String fullAnswer = result.answer() == null ? "" : result.answer();
|
||||
logger.info("统一 ChatService 对话完成 - SessionId: {}, 答案长度: {}",
|
||||
result.sessionId(), fullAnswer.length());
|
||||
|
||||
session.addChatMessagePair(request.getQuestion(), fullAnswer, MAX_WINDOW_SIZE);
|
||||
sessionManager.updateSession(session);
|
||||
logger.info("已更新会话历史 - SessionId: {}, 当前消息对数: {}",
|
||||
session.getSessionId(), session.getMessagePairCount());
|
||||
|
||||
sendContentChunks(emitter, fullAnswer);
|
||||
emitter.send(SseEmitter.event()
|
||||
.name("message")
|
||||
.data(SseMessage.done(), MediaType.APPLICATION_JSON));
|
||||
emitter.complete();
|
||||
|
||||
} catch (Exception e) {
|
||||
logger.error("ReactAgent 对话初始化失败", e);
|
||||
@@ -364,12 +302,13 @@ public class ChatController {
|
||||
try {
|
||||
logger.info("收到获取会话信息请求 - SessionId: {}", sessionId);
|
||||
|
||||
SessionInfo session = sessions.get(sessionId);
|
||||
if (session != null) {
|
||||
Optional<SessionContext> session = sessionManager.getSession(sessionId);
|
||||
if (session.isPresent()) {
|
||||
SessionContext context = session.get();
|
||||
SessionInfoResponse response = new SessionInfoResponse();
|
||||
response.setSessionId(sessionId);
|
||||
response.setMessagePairCount(session.getMessagePairCount());
|
||||
response.setCreateTime(session.createTime);
|
||||
response.setMessagePairCount(context.getMessagePairCount());
|
||||
response.setCreateTime(toEpochMillis(context.getCreatedAt()));
|
||||
return ResponseEntity.ok(ApiResponse.success(response));
|
||||
} else {
|
||||
return ResponseEntity.ok(ApiResponse.error("会话不存在"));
|
||||
@@ -383,107 +322,39 @@ public class ChatController {
|
||||
|
||||
// ==================== 辅助方法 ====================
|
||||
|
||||
private SessionInfo getOrCreateSession(String sessionId) {
|
||||
if (sessionId == null || sessionId.isEmpty()) {
|
||||
sessionId = UUID.randomUUID().toString();
|
||||
}
|
||||
return sessions.computeIfAbsent(sessionId, SessionInfo::new);
|
||||
private SessionContext getOrCreateSession(String sessionId) {
|
||||
String resolvedSessionId = (sessionId == null || sessionId.isEmpty())
|
||||
? UUID.randomUUID().toString()
|
||||
: sessionId;
|
||||
return sessionManager.getSession(resolvedSessionId)
|
||||
.orElseGet(() -> {
|
||||
SessionContext context = SessionContext.builder()
|
||||
.sessionId(resolvedSessionId)
|
||||
.status("ACTIVE")
|
||||
.ttl(sessionTtlSeconds)
|
||||
.build();
|
||||
sessionManager.createSession(context, sessionTtlSeconds);
|
||||
return context;
|
||||
});
|
||||
}
|
||||
|
||||
// ==================== 内部类 ====================
|
||||
|
||||
/**
|
||||
* 会话信息
|
||||
* 管理单个会话的历史消息,支持自动清理和线程安全
|
||||
*/
|
||||
private static class SessionInfo {
|
||||
private final String sessionId;
|
||||
// 存储历史消息对:[{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]
|
||||
private final List<Map<String, String>> messageHistory;
|
||||
private final long createTime;
|
||||
private final ReentrantLock lock;
|
||||
|
||||
public SessionInfo(String sessionId) {
|
||||
this.sessionId = sessionId;
|
||||
this.messageHistory = new ArrayList<>();
|
||||
this.createTime = System.currentTimeMillis();
|
||||
this.lock = new ReentrantLock();
|
||||
private long toEpochMillis(LocalDateTime time) {
|
||||
if (time == null) {
|
||||
return 0L;
|
||||
}
|
||||
return time.atZone(ZoneId.systemDefault()).toInstant().toEpochMilli();
|
||||
}
|
||||
|
||||
/**
|
||||
* 添加一对消息(用户问题 + AI回复)
|
||||
* 自动管理历史消息窗口大小
|
||||
*/
|
||||
public void addMessage(String userQuestion, String aiAnswer) {
|
||||
lock.lock();
|
||||
try {
|
||||
// 添加用户消息
|
||||
Map<String, String> userMsg = new HashMap<>();
|
||||
userMsg.put("role", "user");
|
||||
userMsg.put("content", userQuestion);
|
||||
messageHistory.add(userMsg);
|
||||
|
||||
// 添加AI回复
|
||||
Map<String, String> assistantMsg = new HashMap<>();
|
||||
assistantMsg.put("role", "assistant");
|
||||
assistantMsg.put("content", aiAnswer);
|
||||
messageHistory.add(assistantMsg);
|
||||
|
||||
// 自动清理:保持最多 MAX_WINDOW_SIZE 对消息
|
||||
// 每对消息包含2条记录(user + assistant)
|
||||
int maxMessages = MAX_WINDOW_SIZE * 2;
|
||||
while (messageHistory.size() > maxMessages) {
|
||||
// 成对删除最旧的消息(删除前2条)
|
||||
messageHistory.remove(0); // 删除最旧的用户消息
|
||||
if (!messageHistory.isEmpty()) {
|
||||
messageHistory.remove(0); // 删除对应的AI回复
|
||||
}
|
||||
}
|
||||
|
||||
logger.debug("会话 {} 更新历史消息,当前消息对数: {}",
|
||||
sessionId, messageHistory.size() / 2);
|
||||
|
||||
} finally {
|
||||
lock.unlock();
|
||||
}
|
||||
private void sendContentChunks(SseEmitter emitter, String content) throws IOException {
|
||||
if (content == null || content.isEmpty()) {
|
||||
return;
|
||||
}
|
||||
|
||||
/**
|
||||
* 获取历史消息(线程安全)
|
||||
* 返回副本以避免并发修改
|
||||
*/
|
||||
public List<Map<String, String>> getHistory() {
|
||||
lock.lock();
|
||||
try {
|
||||
return new ArrayList<>(messageHistory);
|
||||
} finally {
|
||||
lock.unlock();
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* 清空历史消息
|
||||
*/
|
||||
public void clearHistory() {
|
||||
lock.lock();
|
||||
try {
|
||||
messageHistory.clear();
|
||||
logger.info("会话 {} 历史消息已清空", sessionId);
|
||||
} finally {
|
||||
lock.unlock();
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* 获取当前消息对数
|
||||
*/
|
||||
public int getMessagePairCount() {
|
||||
lock.lock();
|
||||
try {
|
||||
return messageHistory.size() / 2;
|
||||
} finally {
|
||||
lock.unlock();
|
||||
}
|
||||
int chunkSize = 80;
|
||||
for (int i = 0; i < content.length(); i += chunkSize) {
|
||||
int end = Math.min(i + chunkSize, content.length());
|
||||
emitter.send(SseEmitter.event()
|
||||
.name("message")
|
||||
.data(SseMessage.content(content.substring(i, end)), MediaType.APPLICATION_JSON));
|
||||
}
|
||||
}
|
||||
|
||||
@@ -537,11 +408,13 @@ public class ChatController {
|
||||
private boolean success;
|
||||
private String answer;
|
||||
private String errorMessage;
|
||||
private String sessionId;
|
||||
|
||||
public static ChatResponse success(String answer) {
|
||||
public static ChatResponse success(String answer, String sessionId) {
|
||||
ChatResponse response = new ChatResponse();
|
||||
response.setSuccess(true);
|
||||
response.setAnswer(answer);
|
||||
response.setSessionId(sessionId);
|
||||
return response;
|
||||
}
|
||||
|
||||
|
||||
@@ -0,0 +1,24 @@
|
||||
package com.superbiz.agent.controller;
|
||||
|
||||
import com.superbiz.agent.dto.DiagnosisTraceResponse;
|
||||
import com.superbiz.agent.dto.Result;
|
||||
import com.superbiz.agent.service.DiagnosisTraceService;
|
||||
import lombok.RequiredArgsConstructor;
|
||||
import org.springframework.http.ResponseEntity;
|
||||
import org.springframework.web.bind.annotation.GetMapping;
|
||||
import org.springframework.web.bind.annotation.PathVariable;
|
||||
import org.springframework.web.bind.annotation.RequestMapping;
|
||||
import org.springframework.web.bind.annotation.RestController;
|
||||
|
||||
@RestController
|
||||
@RequestMapping("/api/diagnosis")
|
||||
@RequiredArgsConstructor
|
||||
public class DiagnosisTraceController {
|
||||
|
||||
private final DiagnosisTraceService diagnosisTraceService;
|
||||
|
||||
@GetMapping("/{sessionId}/trace")
|
||||
public ResponseEntity<Result<DiagnosisTraceResponse>> getTrace(@PathVariable String sessionId) {
|
||||
return ResponseEntity.ok(Result.success(diagnosisTraceService.getTrace(sessionId)));
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,25 @@
|
||||
package com.superbiz.agent.controller;
|
||||
|
||||
import com.superbiz.agent.dto.FeedbackRequest;
|
||||
import com.superbiz.agent.dto.FeedbackResponse;
|
||||
import com.superbiz.agent.service.FeedbackService;
|
||||
import org.springframework.beans.factory.annotation.Autowired;
|
||||
import org.springframework.http.ResponseEntity;
|
||||
import org.springframework.web.bind.annotation.*;
|
||||
|
||||
@RestController
|
||||
@RequestMapping("/api")
|
||||
public class FeedbackController {
|
||||
|
||||
@Autowired
|
||||
private FeedbackService feedbackService;
|
||||
|
||||
@PostMapping("/feedback")
|
||||
public ResponseEntity<FeedbackResponse> submitFeedback(@RequestBody FeedbackRequest request) {
|
||||
FeedbackResponse response = feedbackService.submitFeedback(request.getSessionId(), request.getFeedback());
|
||||
if (!response.isSuccess()) {
|
||||
return ResponseEntity.badRequest().body(response);
|
||||
}
|
||||
return ResponseEntity.ok(response);
|
||||
}
|
||||
}
|
||||
@@ -54,6 +54,9 @@ public class DiagnosisSession {
|
||||
@Column(name = "tool_call_count")
|
||||
private Integer toolCallCount;
|
||||
|
||||
@Column(name = "answer", columnDefinition = "LONGTEXT")
|
||||
private String answer;
|
||||
|
||||
@JdbcTypeCode(SqlTypes.JSON)
|
||||
@Column(name = "self_evaluation", columnDefinition = "JSON")
|
||||
private String selfEvaluation;
|
||||
|
||||
@@ -0,0 +1,51 @@
|
||||
package com.superbiz.agent.domain.entity;
|
||||
|
||||
import jakarta.persistence.*;
|
||||
import lombok.AllArgsConstructor;
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
import lombok.NoArgsConstructor;
|
||||
|
||||
import java.time.LocalDateTime;
|
||||
|
||||
@Entity
|
||||
@Table(name = "knowledge_domain")
|
||||
@Data
|
||||
@Builder
|
||||
@NoArgsConstructor
|
||||
@AllArgsConstructor
|
||||
public class KnowledgeDomain {
|
||||
|
||||
@Id
|
||||
@GeneratedValue(strategy = GenerationType.IDENTITY)
|
||||
private Long id;
|
||||
|
||||
@Column(name = "domain_id", unique = true, nullable = false, length = 64)
|
||||
private String domainId;
|
||||
|
||||
@Column(name = "description", length = 256)
|
||||
private String description;
|
||||
|
||||
@Column(name = "when_to_retrieve", columnDefinition = "TEXT")
|
||||
private String whenToRetrieve;
|
||||
|
||||
@Column(name = "document_count", nullable = false)
|
||||
private int documentCount;
|
||||
|
||||
@Column(name = "created_at", nullable = false, updatable = false)
|
||||
private LocalDateTime createdAt;
|
||||
|
||||
@Column(name = "updated_at", nullable = false)
|
||||
private LocalDateTime updatedAt;
|
||||
|
||||
@PrePersist
|
||||
protected void onCreate() {
|
||||
createdAt = LocalDateTime.now();
|
||||
updatedAt = LocalDateTime.now();
|
||||
}
|
||||
|
||||
@PreUpdate
|
||||
protected void onUpdate() {
|
||||
updatedAt = LocalDateTime.now();
|
||||
}
|
||||
}
|
||||
@@ -61,6 +61,12 @@ public class ToolInvocation {
|
||||
@Column(name = "is_truncated")
|
||||
private Boolean isTruncated;
|
||||
|
||||
@Column(name = "relevance_level", length = 20)
|
||||
private String relevanceLevel;
|
||||
|
||||
@Column(name = "dedup_reason", length = 32)
|
||||
private String dedupReason;
|
||||
|
||||
@JdbcTypeCode(SqlTypes.JSON)
|
||||
@Column(name = "retrieval_details", columnDefinition = "JSON")
|
||||
private String retrievalDetails;
|
||||
|
||||
@@ -8,7 +8,9 @@ import lombok.NoArgsConstructor;
|
||||
import java.io.Serializable;
|
||||
import java.time.LocalDateTime;
|
||||
import java.util.ArrayList;
|
||||
import java.util.HashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
/**
|
||||
* 会话上下文数据类
|
||||
@@ -53,6 +55,12 @@ public class SessionContext implements Serializable {
|
||||
@Builder.Default
|
||||
private List<ToolCall> toolCalls = new ArrayList<>();
|
||||
|
||||
/**
|
||||
* 聊天消息历史:[{"role":"user","content":"..."}, {"role":"assistant","content":"..."}]
|
||||
*/
|
||||
@Builder.Default
|
||||
private List<Map<String, String>> messageHistory = new ArrayList<>();
|
||||
|
||||
/**
|
||||
* 会话创建时间
|
||||
*/
|
||||
@@ -79,6 +87,64 @@ public class SessionContext implements Serializable {
|
||||
this.lastActiveAt = LocalDateTime.now();
|
||||
}
|
||||
|
||||
/**
|
||||
* 添加一对聊天消息,并按消息对数裁剪窗口。
|
||||
*/
|
||||
public void addChatMessagePair(String userQuestion, String assistantAnswer, int maxPairCount) {
|
||||
if (this.messageHistory == null) {
|
||||
this.messageHistory = new ArrayList<>();
|
||||
}
|
||||
|
||||
Map<String, String> userMessage = new HashMap<>();
|
||||
userMessage.put("role", "user");
|
||||
userMessage.put("content", userQuestion);
|
||||
this.messageHistory.add(userMessage);
|
||||
|
||||
Map<String, String> assistantMessage = new HashMap<>();
|
||||
assistantMessage.put("role", "assistant");
|
||||
assistantMessage.put("content", assistantAnswer);
|
||||
this.messageHistory.add(assistantMessage);
|
||||
|
||||
int maxMessages = Math.max(maxPairCount, 0) * 2;
|
||||
while (maxMessages > 0 && this.messageHistory.size() > maxMessages) {
|
||||
this.messageHistory.remove(0);
|
||||
if (!this.messageHistory.isEmpty()) {
|
||||
this.messageHistory.remove(0);
|
||||
}
|
||||
}
|
||||
this.lastActiveAt = LocalDateTime.now();
|
||||
}
|
||||
|
||||
/**
|
||||
* 获取聊天历史副本,避免调用方直接修改内部列表。
|
||||
*/
|
||||
public List<Map<String, String>> getMessageHistorySnapshot() {
|
||||
if (this.messageHistory == null || this.messageHistory.isEmpty()) {
|
||||
return new ArrayList<>();
|
||||
}
|
||||
List<Map<String, String>> snapshot = new ArrayList<>();
|
||||
for (Map<String, String> message : this.messageHistory) {
|
||||
snapshot.add(new HashMap<>(message));
|
||||
}
|
||||
return snapshot;
|
||||
}
|
||||
|
||||
/**
|
||||
* 清空聊天历史。
|
||||
*/
|
||||
public void clearMessageHistory() {
|
||||
if (this.messageHistory == null) {
|
||||
this.messageHistory = new ArrayList<>();
|
||||
} else {
|
||||
this.messageHistory.clear();
|
||||
}
|
||||
this.lastActiveAt = LocalDateTime.now();
|
||||
}
|
||||
|
||||
public int getMessagePairCount() {
|
||||
return this.messageHistory == null ? 0 : this.messageHistory.size() / 2;
|
||||
}
|
||||
|
||||
/**
|
||||
* 更新最后活跃时间
|
||||
*/
|
||||
|
||||
@@ -0,0 +1,102 @@
|
||||
package com.superbiz.agent.dto;
|
||||
|
||||
import lombok.AllArgsConstructor;
|
||||
import lombok.Builder;
|
||||
import lombok.Data;
|
||||
import lombok.NoArgsConstructor;
|
||||
|
||||
import java.time.LocalDateTime;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
@Data
|
||||
@Builder
|
||||
@NoArgsConstructor
|
||||
@AllArgsConstructor
|
||||
public class DiagnosisTraceResponse {
|
||||
|
||||
private SessionTrace session;
|
||||
private List<AgentStepTrace> steps;
|
||||
private List<ToolInvocationTrace> toolInvocations;
|
||||
private TraceSummary summary;
|
||||
|
||||
@Data
|
||||
@Builder
|
||||
@NoArgsConstructor
|
||||
@AllArgsConstructor
|
||||
public static class SessionTrace {
|
||||
private Long id;
|
||||
private String sessionId;
|
||||
private String query;
|
||||
private String status;
|
||||
private String agentFlow;
|
||||
private Integer totalDurationMs;
|
||||
private Integer totalTokenCount;
|
||||
private Integer stepCount;
|
||||
private Integer toolCallCount;
|
||||
private String answer;
|
||||
private String selfEvaluationRaw;
|
||||
private Map<String, Object> selfEvaluation;
|
||||
private String feedback;
|
||||
private LocalDateTime createdAt;
|
||||
private LocalDateTime updatedAt;
|
||||
}
|
||||
|
||||
@Data
|
||||
@Builder
|
||||
@NoArgsConstructor
|
||||
@AllArgsConstructor
|
||||
public static class AgentStepTrace {
|
||||
private Long id;
|
||||
private String sessionId;
|
||||
private Integer stepIndex;
|
||||
private String agentName;
|
||||
private String modelInput;
|
||||
private String modelOutput;
|
||||
private String thought;
|
||||
private Boolean hasToolCall;
|
||||
private Integer durationMs;
|
||||
private Integer tokenCount;
|
||||
private LocalDateTime createdAt;
|
||||
}
|
||||
|
||||
@Data
|
||||
@Builder
|
||||
@NoArgsConstructor
|
||||
@AllArgsConstructor
|
||||
public static class ToolInvocationTrace {
|
||||
private Long id;
|
||||
private String sessionId;
|
||||
private Long stepId;
|
||||
private String toolName;
|
||||
private String inputParamsRaw;
|
||||
private Map<String, Object> inputParams;
|
||||
private String outputPreview;
|
||||
private Integer outputLength;
|
||||
private String retrievalLayer;
|
||||
private Integer l0MatchCount;
|
||||
private Integer l1MatchCount;
|
||||
private Boolean truncated;
|
||||
private String relevanceLevel;
|
||||
private String dedupReason;
|
||||
private String retrievalDetailsRaw;
|
||||
private Map<String, Object> retrievalDetails;
|
||||
private Integer durationMs;
|
||||
private Boolean success;
|
||||
private String errorMessage;
|
||||
private LocalDateTime createdAt;
|
||||
}
|
||||
|
||||
@Data
|
||||
@Builder
|
||||
@NoArgsConstructor
|
||||
@AllArgsConstructor
|
||||
public static class TraceSummary {
|
||||
private int persistedStepCount;
|
||||
private int returnedStepCount;
|
||||
private int persistedToolCallCount;
|
||||
private int returnedToolCallCount;
|
||||
private boolean hasVerifierEvaluation;
|
||||
private boolean hasFeedback;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,11 @@
|
||||
package com.superbiz.agent.dto;
|
||||
|
||||
import lombok.Getter;
|
||||
import lombok.Setter;
|
||||
|
||||
@Getter
|
||||
@Setter
|
||||
public class FeedbackRequest {
|
||||
private String sessionId;
|
||||
private String feedback;
|
||||
}
|
||||
@@ -0,0 +1,12 @@
|
||||
package com.superbiz.agent.dto;
|
||||
|
||||
import lombok.Builder;
|
||||
import lombok.Getter;
|
||||
|
||||
@Getter
|
||||
@Builder
|
||||
public class FeedbackResponse {
|
||||
private boolean success;
|
||||
private String message;
|
||||
private String caseId;
|
||||
}
|
||||
@@ -59,4 +59,14 @@ public class Frontmatter {
|
||||
* 最后更新日期(预留字段)
|
||||
*/
|
||||
private LocalDate lastUpdated;
|
||||
|
||||
/**
|
||||
* 业务场景标签,供 Planner 决策用(LLM 上传时自动生成)
|
||||
*/
|
||||
private List<String> covers;
|
||||
|
||||
/**
|
||||
* 文档级检索时机(LLM 上传时自动生成)
|
||||
*/
|
||||
private String whenToRetrieve;
|
||||
}
|
||||
|
||||
@@ -43,4 +43,14 @@ public class KnowledgeEntry {
|
||||
* 章节锚点(预留字段,MVP 不使用)
|
||||
*/
|
||||
private Map<String, String> sections;
|
||||
|
||||
/**
|
||||
* 业务场景标签,供 Planner 决策用
|
||||
*/
|
||||
private List<String> covers;
|
||||
|
||||
/**
|
||||
* 文档级检索时机
|
||||
*/
|
||||
private String whenToRetrieve;
|
||||
}
|
||||
|
||||
@@ -26,4 +26,24 @@ public class LookupResult {
|
||||
* 补充结果(L1 语义检索)
|
||||
*/
|
||||
private SupplementResult supplement;
|
||||
|
||||
/**
|
||||
* 归一化质量等级:PRECISE / HIGHLY_RELEVANT / REFERENCE
|
||||
*/
|
||||
private String relevanceLevel;
|
||||
|
||||
/**
|
||||
* 兜底信号:告诉 LLM 知识库的"天花板"
|
||||
*/
|
||||
private String completenessHint;
|
||||
|
||||
/**
|
||||
* 本次会话已检索过的域列表(行动记忆)
|
||||
*/
|
||||
private List<String> retrievedDomainsThisSession;
|
||||
|
||||
/**
|
||||
* 系统消息(如去重提示)
|
||||
*/
|
||||
private String message;
|
||||
}
|
||||
|
||||
@@ -1,25 +1,27 @@
|
||||
package com.superbiz.agent.hook;
|
||||
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.messages.MessagesModelHook;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.messages.AgentCommand;
|
||||
import com.alibaba.cloud.ai.graph.RunnableConfig;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.HookPosition;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.HookPositions;
|
||||
import com.alibaba.cloud.ai.graph.RunnableConfig;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.messages.AgentCommand;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.messages.MessagesModelHook;
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.domain.entity.AgentStep;
|
||||
import com.superbiz.agent.repository.AgentStepRepository;
|
||||
import com.superbiz.agent.util.SessionContextHolder;
|
||||
import lombok.extern.slf4j.Slf4j;
|
||||
import org.springframework.ai.chat.messages.Message;
|
||||
import org.springframework.ai.chat.messages.AssistantMessage;
|
||||
import org.springframework.ai.chat.messages.UserMessage;
|
||||
import org.springframework.ai.chat.messages.Message;
|
||||
import org.springframework.ai.chat.messages.ToolResponseMessage;
|
||||
import org.springframework.ai.chat.messages.UserMessage;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
|
||||
/**
|
||||
* Agent 日志 Hook
|
||||
* 记录 Agent 的思考过程、消息流转 + 持久化 agent_step 到 DB
|
||||
* Persists per-agent model input/output snapshots into agent_step.
|
||||
*/
|
||||
@Slf4j
|
||||
@HookPositions({HookPosition.BEFORE_MODEL, HookPosition.AFTER_MODEL})
|
||||
@@ -27,11 +29,9 @@ public class AgentLoggingHook extends MessagesModelHook {
|
||||
|
||||
private final AgentStepRepository agentStepRepository;
|
||||
private final String agentName;
|
||||
private final ObjectMapper objectMapper = new ObjectMapper();
|
||||
|
||||
/** 每个 session 的步数计数器:sessionId → stepIndex */
|
||||
private final ConcurrentHashMap<String, Integer> stepCounters = new ConcurrentHashMap<>();
|
||||
|
||||
/** beforeModel → afterModel 中间状态:sessionId_stepIndex → {stepId, startTime} */
|
||||
private final ConcurrentHashMap<String, Map<String, Object>> pendingSteps = new ConcurrentHashMap<>();
|
||||
|
||||
public AgentLoggingHook(AgentStepRepository agentStepRepository, String agentName) {
|
||||
@@ -46,60 +46,44 @@ public class AgentLoggingHook extends MessagesModelHook {
|
||||
|
||||
@Override
|
||||
public AgentCommand beforeModel(List<Message> previousMessages, RunnableConfig config) {
|
||||
// 优先从 config.metadata 取 sessionId(线程安全),兜底 ThreadLocal
|
||||
String sessionId = config.metadata("sessionId")
|
||||
.map(Object::toString)
|
||||
.orElseGet(SessionContextHolder::getSessionId);
|
||||
|
||||
boolean hasSession = (sessionId != null);
|
||||
String sessionId = resolveSessionId(config);
|
||||
boolean hasSession = sessionId != null;
|
||||
|
||||
int stepIndex = 0;
|
||||
if (hasSession) {
|
||||
stepIndex = stepCounters.merge(sessionId, 0, (old, one) -> old + 1);
|
||||
stepIndex = stepCounters.merge(sessionId, 0, (oldValue, ignored) -> oldValue + 1);
|
||||
}
|
||||
|
||||
log.info("========================================");
|
||||
log.info("*** [Agent 思考] 第 {} 轮思考开始", (hasSession ? stepCounters.get(sessionId) : 0) + 1);
|
||||
log.info("*** [Agent 思考] 当前消息数量: {}", previousMessages.size());
|
||||
log.info("*** [AgentTrace] agent={}, phase=before_model, stepIndex={}", agentName, stepIndex);
|
||||
log.info("*** [AgentTrace] messageCount={}", previousMessages.size());
|
||||
|
||||
// 打印最后几条消息
|
||||
int lastN = Math.min(3, previousMessages.size());
|
||||
if (lastN > 0) {
|
||||
log.info("*** [Agent 思考] 最近 {} 条消息:", lastN);
|
||||
log.info("*** [AgentTrace] recentMessages={}", lastN);
|
||||
List<Message> recentMessages = previousMessages.subList(previousMessages.size() - lastN, previousMessages.size());
|
||||
for (int i = 0; i < recentMessages.size(); i++) {
|
||||
Message msg = recentMessages.get(i);
|
||||
String role = getMessageRole(msg);
|
||||
log.info(" [{}] 角色: {}, 类型: {}", i + 1, role, msg.getClass().getSimpleName());
|
||||
log.info(" [{}] role={}, type={}", i + 1, getMessageRole(msg), msg.getClass().getSimpleName());
|
||||
}
|
||||
}
|
||||
|
||||
log.info("*** [Agent 思考] 准备调用模型...");
|
||||
log.info("========================================");
|
||||
|
||||
// 持久化 agent_step(beforeModel:先创建,先记 model_input 摘要)
|
||||
if (sessionId != null) {
|
||||
try {
|
||||
String modelInputSummary = buildModelInputSummary(previousMessages);
|
||||
|
||||
AgentStep step = AgentStep.builder()
|
||||
.sessionId(sessionId)
|
||||
.stepIndex(stepIndex)
|
||||
.agentName(agentName)
|
||||
.modelInput(modelInputSummary)
|
||||
.modelInput(buildModelInputSummary(previousMessages))
|
||||
.build();
|
||||
AgentStep saved = agentStepRepository.save(step);
|
||||
|
||||
// 记录中间状态供 afterModel 使用
|
||||
pendingSteps.put(sessionId + "_" + stepIndex, Map.of(
|
||||
"stepId", saved.getId(),
|
||||
"startTime", System.currentTimeMillis()
|
||||
));
|
||||
|
||||
log.debug("agent_step 已创建: sessionId={}, stepIndex={}, id={}", sessionId, stepIndex, saved.getId());
|
||||
} catch (Exception e) {
|
||||
log.error("保存 agent_step 失败", e);
|
||||
// 不中断 Agent 执行
|
||||
log.error("Failed to persist agent_step before model", e);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -108,58 +92,38 @@ public class AgentLoggingHook extends MessagesModelHook {
|
||||
|
||||
@Override
|
||||
public AgentCommand afterModel(List<Message> previousMessages, RunnableConfig config) {
|
||||
String sessionId = SessionContextHolder.getSessionId();
|
||||
boolean hasSession = (sessionId != null);
|
||||
String sessionId = resolveSessionId(config);
|
||||
int stepIndex = sessionId == null ? 0 : stepCounters.getOrDefault(sessionId, 0);
|
||||
|
||||
log.info("========================================");
|
||||
log.info("*** [Agent 思考] 第 {} 轮思考完成", (hasSession ? stepCounters.getOrDefault(sessionId, 0) : 0));
|
||||
|
||||
// 查找最后一条 AssistantMessage(模型的回复)
|
||||
AssistantMessage lastAssistant = null;
|
||||
for (int i = previousMessages.size() - 1; i >= 0; i--) {
|
||||
if (previousMessages.get(i) instanceof AssistantMessage) {
|
||||
lastAssistant = (AssistantMessage) previousMessages.get(i);
|
||||
break;
|
||||
}
|
||||
}
|
||||
log.info("*** [AgentTrace] agent={}, phase=after_model, stepIndex={}", agentName, stepIndex);
|
||||
|
||||
AssistantMessage lastAssistant = findLastAssistant(previousMessages);
|
||||
boolean hasToolCall = false;
|
||||
|
||||
if (lastAssistant != null) {
|
||||
// 打印模型返回的文本内容
|
||||
String textContent = extractTextContent(lastAssistant);
|
||||
if (textContent != null && !textContent.isEmpty()) {
|
||||
log.info("*** [Agent 思考] 模型返回文本: {}",
|
||||
textContent.length() > 500
|
||||
? textContent.substring(0, 500) + "... (已截断,总长度: " + textContent.length() + ")"
|
||||
: textContent);
|
||||
log.info("*** [AgentTrace] text={}",
|
||||
textContent.length() > 500
|
||||
? textContent.substring(0, 500) + "... (len=" + textContent.length() + ")"
|
||||
: textContent);
|
||||
}
|
||||
|
||||
// 检查是否有工具调用
|
||||
if (lastAssistant.getToolCalls() != null && !lastAssistant.getToolCalls().isEmpty()) {
|
||||
hasToolCall = true;
|
||||
log.info("*** [Agent 思考] 模型决定调用 {} 个工具:",
|
||||
lastAssistant.getToolCalls().size());
|
||||
lastAssistant.getToolCalls().forEach(toolCall -> {
|
||||
log.info(" - 工具: {}, 参数: {}",
|
||||
toolCall.name(),
|
||||
toolCall.arguments());
|
||||
});
|
||||
log.info("*** [Agent 思考] 等待工具执行结果...");
|
||||
log.info("*** [AgentTrace] toolCalls={}", lastAssistant.getToolCalls().size());
|
||||
lastAssistant.getToolCalls().forEach(toolCall ->
|
||||
log.info(" - tool={}, arguments={}", toolCall.name(), toolCall.arguments()));
|
||||
} else {
|
||||
log.info("*** [Agent 思考] 模型决定不调用工具");
|
||||
log.info("*** [Agent 思考] 这是最终答案,准备返回给用户");
|
||||
log.info("*** [AgentTrace] no tool call");
|
||||
}
|
||||
}
|
||||
|
||||
log.info("========================================");
|
||||
|
||||
// 更新 agent_step(afterModel:补全 model_output、耗时等)
|
||||
if (sessionId != null) {
|
||||
int stepIndex = stepCounters.getOrDefault(sessionId, 0);
|
||||
String stepKey = sessionId + "_" + stepIndex;
|
||||
Map<String, Object> pending = pendingSteps.remove(stepKey);
|
||||
|
||||
if (pending != null) {
|
||||
try {
|
||||
Long stepId = (Long) pending.get("stepId");
|
||||
@@ -168,20 +132,12 @@ public class AgentLoggingHook extends MessagesModelHook {
|
||||
|
||||
AgentStep step = agentStepRepository.findById(stepId).orElse(null);
|
||||
if (step != null) {
|
||||
String thought = extractTextContent(lastAssistant);
|
||||
if (thought != null && thought.length() > 2000) {
|
||||
thought = thought.substring(0, 2000);
|
||||
}
|
||||
|
||||
step.setThought(thought);
|
||||
step.setThought(buildStoredThought(lastAssistant));
|
||||
step.setHasToolCall(hasToolCall);
|
||||
step.setDurationMs(durationMs);
|
||||
|
||||
if (lastAssistant != null) {
|
||||
String outputSummary = buildModelOutputSummary(lastAssistant);
|
||||
step.setModelOutput(outputSummary);
|
||||
|
||||
// 读取实际 token 用量(由 TokenTrackingChatModel 写入)
|
||||
step.setModelOutput(buildModelOutputSummary(lastAssistant));
|
||||
Integer tokenCount = TokenUsageHolder.get();
|
||||
if (tokenCount != null) {
|
||||
step.setTokenCount(tokenCount);
|
||||
@@ -189,24 +145,32 @@ public class AgentLoggingHook extends MessagesModelHook {
|
||||
}
|
||||
|
||||
agentStepRepository.save(step);
|
||||
log.debug("agent_step 已更新: sessionId={}, stepIndex={}, duration={}ms",
|
||||
sessionId, stepIndex, durationMs);
|
||||
}
|
||||
} catch (Exception e) {
|
||||
log.error("更新 agent_step 失败", e);
|
||||
log.error("Failed to update agent_step after model", e);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// 清理 token 上下文
|
||||
TokenUsageHolder.clear();
|
||||
|
||||
return new AgentCommand(previousMessages);
|
||||
}
|
||||
|
||||
/**
|
||||
* 构建模型输入摘要(前 N 条消息的 role + 截断内容)
|
||||
*/
|
||||
private String resolveSessionId(RunnableConfig config) {
|
||||
return config.metadata("sessionId")
|
||||
.map(Object::toString)
|
||||
.orElseGet(SessionContextHolder::getSessionId);
|
||||
}
|
||||
|
||||
private AssistantMessage findLastAssistant(List<Message> previousMessages) {
|
||||
for (int i = previousMessages.size() - 1; i >= 0; i--) {
|
||||
if (previousMessages.get(i) instanceof AssistantMessage assistantMessage) {
|
||||
return assistantMessage;
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
private String buildModelInputSummary(List<Message> messages) {
|
||||
StringBuilder sb = new StringBuilder();
|
||||
int maxMessages = Math.min(messages.size(), 5);
|
||||
@@ -226,25 +190,59 @@ public class AgentLoggingHook extends MessagesModelHook {
|
||||
return result;
|
||||
}
|
||||
|
||||
/**
|
||||
* 构建模型输出摘要
|
||||
*/
|
||||
private String buildStoredThought(AssistantMessage message) {
|
||||
String text = extractTextContent(message);
|
||||
if (text == null || text.isBlank()) {
|
||||
return text;
|
||||
}
|
||||
if (!"verifier".equals(agentName)) {
|
||||
return truncate(text, 2000);
|
||||
}
|
||||
return summarizeVerifierThought(text);
|
||||
}
|
||||
|
||||
private String summarizeVerifierThought(String verifierOutput) {
|
||||
try {
|
||||
JsonNode root = objectMapper.readTree(verifierOutput);
|
||||
int factCount = root.path("facts_checked").isArray() ? root.path("facts_checked").size() : 0;
|
||||
int tracedFactCount = 0;
|
||||
if (root.path("facts_checked").isArray()) {
|
||||
for (JsonNode factNode : root.path("facts_checked")) {
|
||||
if (factNode.path("evidence_refs").isArray() && factNode.path("evidence_refs").size() > 0) {
|
||||
tracedFactCount++;
|
||||
}
|
||||
}
|
||||
}
|
||||
return "verdict=%s, score=%s, critical_fact_count=%s, facts_checked=%d, traced_facts=%d".formatted(
|
||||
root.path("verdict").asText("UNKNOWN"),
|
||||
root.path("groundedness_score").asText("0.0"),
|
||||
root.path("critical_fact_count").asText("0"),
|
||||
factCount,
|
||||
tracedFactCount
|
||||
);
|
||||
} catch (Exception e) {
|
||||
return truncate(verifierOutput, 300);
|
||||
}
|
||||
}
|
||||
|
||||
private String buildModelOutputSummary(AssistantMessage message) {
|
||||
String text = extractTextContent(message);
|
||||
if (text == null) {
|
||||
text = "";
|
||||
}
|
||||
if (text.length() > 500) {
|
||||
text = text.substring(0, 500) + "...";
|
||||
}
|
||||
int maxTextLength = "verifier".equals(agentName) ? 4000 : 500;
|
||||
text = truncate(text, maxTextLength);
|
||||
|
||||
StringBuilder sb = new StringBuilder();
|
||||
sb.append("{\"text\":\"").append(escapeJson(text)).append("\"");
|
||||
if (message.getToolCalls() != null && !message.getToolCalls().isEmpty()) {
|
||||
sb.append(",\"toolCalls\":[");
|
||||
for (int i = 0; i < message.getToolCalls().size(); i++) {
|
||||
if (i > 0) sb.append(",");
|
||||
if (i > 0) {
|
||||
sb.append(",");
|
||||
}
|
||||
sb.append("{\"name\":\"").append(escapeJson(message.getToolCalls().get(i).name()))
|
||||
.append("\",\"arguments\":").append(message.getToolCalls().get(i).arguments()).append("}");
|
||||
.append("\",\"arguments\":").append(message.getToolCalls().get(i).arguments()).append("}");
|
||||
}
|
||||
sb.append("]");
|
||||
}
|
||||
@@ -253,7 +251,9 @@ public class AgentLoggingHook extends MessagesModelHook {
|
||||
}
|
||||
|
||||
private String escapeJson(String s) {
|
||||
if (s == null) return "";
|
||||
if (s == null) {
|
||||
return "";
|
||||
}
|
||||
return s.replace("\\", "\\\\")
|
||||
.replace("\"", "\\\"")
|
||||
.replace("\n", "\\n")
|
||||
@@ -261,96 +261,70 @@ public class AgentLoggingHook extends MessagesModelHook {
|
||||
.replace("\t", "\\t");
|
||||
}
|
||||
|
||||
/**
|
||||
* 提取 AssistantMessage 的文本内容
|
||||
*/
|
||||
private String truncate(String text, int maxLength) {
|
||||
if (text == null || text.length() <= maxLength) {
|
||||
return text;
|
||||
}
|
||||
return text.substring(0, maxLength) + "...";
|
||||
}
|
||||
|
||||
private String extractTextContent(AssistantMessage message) {
|
||||
if (message == null) return null;
|
||||
if (message == null) {
|
||||
return null;
|
||||
}
|
||||
try {
|
||||
// 方法 1: 反射获取 text 字段
|
||||
try {
|
||||
java.lang.reflect.Field textField = message.getClass().getDeclaredField("text");
|
||||
textField.setAccessible(true);
|
||||
Object value = textField.get(message);
|
||||
if (value != null) {
|
||||
log.debug("通过 text 字段提取成功");
|
||||
return value.toString();
|
||||
return message.getText();
|
||||
} catch (Exception ignore) {
|
||||
// Fallback below.
|
||||
}
|
||||
|
||||
for (String fieldName : List.of("text", "content")) {
|
||||
try {
|
||||
java.lang.reflect.Field field = message.getClass().getDeclaredField(fieldName);
|
||||
field.setAccessible(true);
|
||||
Object value = field.get(message);
|
||||
if (value != null) {
|
||||
return value.toString();
|
||||
}
|
||||
} catch (NoSuchFieldException ignore) {
|
||||
// continue
|
||||
}
|
||||
} catch (NoSuchFieldException e) {
|
||||
// 尝试下一种方法
|
||||
}
|
||||
|
||||
// 方法 2: 反射获取 content 字段
|
||||
try {
|
||||
java.lang.reflect.Field contentField = message.getClass().getDeclaredField("content");
|
||||
contentField.setAccessible(true);
|
||||
Object value = contentField.get(message);
|
||||
if (value != null) {
|
||||
log.debug("通过 content 字段提取成功");
|
||||
return value.toString();
|
||||
for (String methodName : List.of("getText", "getContent")) {
|
||||
try {
|
||||
java.lang.reflect.Method method = message.getClass().getMethod(methodName);
|
||||
Object value = method.invoke(message);
|
||||
if (value != null) {
|
||||
return value.toString();
|
||||
}
|
||||
} catch (NoSuchMethodException ignore) {
|
||||
// continue
|
||||
}
|
||||
} catch (NoSuchFieldException e) {
|
||||
// 尝试下一种方法
|
||||
}
|
||||
|
||||
// 方法 3: 调用 getText() 方法
|
||||
try {
|
||||
java.lang.reflect.Method getTextMethod = message.getClass().getMethod("getText");
|
||||
Object value = getTextMethod.invoke(message);
|
||||
if (value != null) {
|
||||
log.debug("通过 getText() 方法提取成功");
|
||||
return value.toString();
|
||||
}
|
||||
} catch (NoSuchMethodException e) {
|
||||
// 尝试下一种方法
|
||||
String fallback = message.toString();
|
||||
if (fallback != null && !fallback.startsWith("AssistantMessage@")) {
|
||||
return fallback;
|
||||
}
|
||||
|
||||
// 方法 4: 调用 getContent() 方法
|
||||
try {
|
||||
java.lang.reflect.Method getContentMethod = message.getClass().getMethod("getContent");
|
||||
Object value = getContentMethod.invoke(message);
|
||||
if (value != null) {
|
||||
log.debug("通过 getContent() 方法提取成功");
|
||||
return value.toString();
|
||||
}
|
||||
} catch (NoSuchMethodException e) {
|
||||
// 方法不存在
|
||||
}
|
||||
|
||||
// 方法 5: 打印类结构信息
|
||||
log.warn("无法提取 AssistantMessage 文本内容,打印类信息:");
|
||||
log.warn("类名: {}", message.getClass().getName());
|
||||
log.warn("字段列表:");
|
||||
for (java.lang.reflect.Field field : message.getClass().getDeclaredFields()) {
|
||||
log.warn(" - {}: {}", field.getName(), field.getType().getSimpleName());
|
||||
}
|
||||
|
||||
// 方法 6: toString() 兜底
|
||||
String toString = message.toString();
|
||||
if (toString != null && !toString.startsWith("AssistantMessage@")) {
|
||||
log.debug("通过 toString() 提取");
|
||||
return toString;
|
||||
}
|
||||
|
||||
return null;
|
||||
} catch (Exception e) {
|
||||
log.error("提取 AssistantMessage 文本内容时出错", e);
|
||||
log.error("Failed to extract AssistantMessage text", e);
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* 获取消息角色
|
||||
*/
|
||||
private String getMessageRole(Message message) {
|
||||
if (message instanceof UserMessage) {
|
||||
return "User(用户)";
|
||||
} else if (message instanceof AssistantMessage) {
|
||||
return "Assistant(模型)";
|
||||
} else if (message instanceof ToolResponseMessage) {
|
||||
return "Tool(工具返回)";
|
||||
} else {
|
||||
return message.getClass().getSimpleName();
|
||||
return "user";
|
||||
}
|
||||
if (message instanceof AssistantMessage) {
|
||||
return "assistant";
|
||||
}
|
||||
if (message instanceof ToolResponseMessage) {
|
||||
return "tool";
|
||||
}
|
||||
return message.getClass().getSimpleName();
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,105 @@
|
||||
package com.superbiz.agent.hook;
|
||||
|
||||
import com.alibaba.cloud.ai.graph.RunnableConfig;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.HookPosition;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.HookPositions;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.messages.AgentCommand;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.messages.MessagesModelHook;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.service.ToolTraceSummaryService;
|
||||
import com.superbiz.agent.util.SessionContextHolder;
|
||||
import com.superbiz.agent.util.VerifierContextHolder;
|
||||
import lombok.extern.slf4j.Slf4j;
|
||||
import org.springframework.ai.chat.messages.AssistantMessage;
|
||||
import org.springframework.ai.chat.messages.Message;
|
||||
import org.springframework.ai.chat.messages.UserMessage;
|
||||
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
/**
|
||||
* Replaces verifier history with an explicit structured payload.
|
||||
*/
|
||||
@Slf4j
|
||||
@HookPositions(HookPosition.BEFORE_MODEL)
|
||||
public class VerifierInputHook extends MessagesModelHook {
|
||||
|
||||
private final ToolTraceSummaryService toolTraceSummaryService;
|
||||
private final ObjectMapper objectMapper = new ObjectMapper();
|
||||
|
||||
public VerifierInputHook(ToolTraceSummaryService toolTraceSummaryService) {
|
||||
this.toolTraceSummaryService = toolTraceSummaryService;
|
||||
}
|
||||
|
||||
@Override
|
||||
public String getName() {
|
||||
return "verifier_input_hook";
|
||||
}
|
||||
|
||||
@Override
|
||||
public AgentCommand beforeModel(List<Message> previousMessages, RunnableConfig config) {
|
||||
try {
|
||||
String sessionId = config.metadata("sessionId")
|
||||
.map(Object::toString)
|
||||
.orElseGet(SessionContextHolder::getSessionId);
|
||||
String executorFinalAnswer = VerifierContextHolder.getExecutorFinalAnswer();
|
||||
if (executorFinalAnswer == null || executorFinalAnswer.isBlank()) {
|
||||
executorFinalAnswer = extractLastAssistantText(previousMessages);
|
||||
}
|
||||
|
||||
List<Map<String, Object>> toolTraceSummary =
|
||||
toolTraceSummaryService.buildVerifierTraceSummary(sessionId, executorFinalAnswer);
|
||||
VerifierContextHolder.setToolTraceSummary(toolTraceSummary);
|
||||
|
||||
Map<String, Object> verifierInput = new LinkedHashMap<>();
|
||||
verifierInput.put("original_query", VerifierContextHolder.getOriginalQuery());
|
||||
verifierInput.put("executor_final_answer", executorFinalAnswer);
|
||||
verifierInput.put("tool_trace_summary", toolTraceSummary);
|
||||
verifierInput.put("retry_context", VerifierContextHolder.getRetryContext());
|
||||
|
||||
String payload = objectMapper.writerWithDefaultPrettyPrinter().writeValueAsString(verifierInput);
|
||||
return new AgentCommand(List.of(new UserMessage(payload)));
|
||||
} catch (Exception e) {
|
||||
log.error("Failed to build verifier input, fallback to original messages", e);
|
||||
return new AgentCommand(previousMessages);
|
||||
}
|
||||
}
|
||||
|
||||
private String extractLastAssistantText(List<Message> previousMessages) {
|
||||
for (int i = previousMessages.size() - 1; i >= 0; i--) {
|
||||
if (previousMessages.get(i) instanceof AssistantMessage assistantMessage) {
|
||||
String text = extractTextContent(assistantMessage);
|
||||
if (text != null && !text.isBlank()) {
|
||||
return text;
|
||||
}
|
||||
}
|
||||
}
|
||||
return "";
|
||||
}
|
||||
|
||||
private String extractTextContent(AssistantMessage message) {
|
||||
try {
|
||||
try {
|
||||
return message.getText();
|
||||
} catch (Exception ignore) {
|
||||
// Fallback for older implementations.
|
||||
}
|
||||
|
||||
for (String methodName : List.of("getText", "getContent")) {
|
||||
try {
|
||||
var method = message.getClass().getMethod(methodName);
|
||||
Object value = method.invoke(message);
|
||||
if (value != null) {
|
||||
return value.toString();
|
||||
}
|
||||
} catch (NoSuchMethodException ignore) {
|
||||
// continue
|
||||
}
|
||||
}
|
||||
} catch (Exception e) {
|
||||
log.debug("Failed to extract verifier assistant text", e);
|
||||
}
|
||||
return message.toString();
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,13 @@
|
||||
package com.superbiz.agent.repository;
|
||||
|
||||
import com.superbiz.agent.domain.entity.KnowledgeDomain;
|
||||
import org.springframework.data.jpa.repository.JpaRepository;
|
||||
import org.springframework.stereotype.Repository;
|
||||
|
||||
import java.util.Optional;
|
||||
|
||||
@Repository
|
||||
public interface KnowledgeDomainRepository extends JpaRepository<KnowledgeDomain, Long> {
|
||||
|
||||
Optional<KnowledgeDomain> findByDomainId(String domainId);
|
||||
}
|
||||
@@ -17,6 +17,11 @@ public interface ToolInvocationRepository extends JpaRepository<ToolInvocation,
|
||||
*/
|
||||
List<ToolInvocation> findBySessionId(String sessionId);
|
||||
|
||||
/**
|
||||
* 根据会话ID按创建顺序查询所有工具调用
|
||||
*/
|
||||
List<ToolInvocation> findBySessionIdOrderByIdAsc(String sessionId);
|
||||
|
||||
/**
|
||||
* 根据工具名查询所有调用
|
||||
*/
|
||||
|
||||
@@ -0,0 +1,53 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.superbiz.agent.domain.entity.CaseLibrary;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||
import com.superbiz.agent.domain.enums.FaultCategory;
|
||||
import com.superbiz.agent.domain.enums.SourceType;
|
||||
import com.superbiz.agent.repository.CaseLibraryRepository;
|
||||
import org.slf4j.Logger;
|
||||
import org.slf4j.LoggerFactory;
|
||||
import org.springframework.beans.factory.annotation.Autowired;
|
||||
import org.springframework.stereotype.Service;
|
||||
|
||||
import java.util.UUID;
|
||||
|
||||
@Service
|
||||
public class CaseLibraryService {
|
||||
|
||||
private static final Logger logger = LoggerFactory.getLogger(CaseLibraryService.class);
|
||||
|
||||
@Autowired
|
||||
private CaseLibraryRepository caseLibraryRepository;
|
||||
|
||||
public CaseLibrary createFromSession(DiagnosisSession session) {
|
||||
return caseLibraryRepository.findByDiagnosisId(session.getSessionId())
|
||||
.orElseGet(() -> {
|
||||
String content = session.getAnswer();
|
||||
if (content == null || content.isBlank()) {
|
||||
content = session.getQuery() + "\n(自动提取失败,请人工补充)";
|
||||
}
|
||||
|
||||
String title = session.getQuery();
|
||||
if (title.length() > 100) {
|
||||
title = title.substring(0, 100);
|
||||
}
|
||||
|
||||
CaseLibrary caseLibrary = CaseLibrary.builder()
|
||||
.caseId(UUID.randomUUID().toString())
|
||||
.diagnosisId(session.getSessionId())
|
||||
.sourceType(SourceType.AUTO)
|
||||
.faultCategory(FaultCategory.GENERAL)
|
||||
.title(title)
|
||||
.rootCause(content)
|
||||
.solution(content)
|
||||
.createdBy("system")
|
||||
.referenceCount(0)
|
||||
.build();
|
||||
|
||||
CaseLibrary saved = caseLibraryRepository.save(caseLibrary);
|
||||
logger.info("案例已沉淀: caseId={}, sessionId={}", saved.getCaseId(), session.getSessionId());
|
||||
return saved;
|
||||
});
|
||||
}
|
||||
}
|
||||
@@ -3,8 +3,10 @@ package com.superbiz.agent.service;
|
||||
import com.alibaba.cloud.ai.graph.OverAllState;
|
||||
import com.alibaba.cloud.ai.graph.RunnableConfig;
|
||||
import com.alibaba.cloud.ai.graph.agent.ReactAgent;
|
||||
import com.alibaba.cloud.ai.graph.agent.flow.agent.SupervisorAgent;
|
||||
import com.alibaba.cloud.ai.graph.agent.flow.agent.SequentialAgent;
|
||||
import com.alibaba.cloud.ai.graph.exception.GraphRunnerException;
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.agent.tool.DateTimeTools;
|
||||
import com.superbiz.agent.agent.tool.InternalDocsTools;
|
||||
import com.superbiz.agent.agent.tool.QueryLogsTools;
|
||||
@@ -13,11 +15,14 @@ import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||
import com.superbiz.agent.hook.AgentLoggingHook;
|
||||
import com.superbiz.agent.hook.TokenTrackingChatModel;
|
||||
import com.superbiz.agent.hook.TokenUsageHolder;
|
||||
import com.superbiz.agent.hook.VerifierInputHook;
|
||||
import com.superbiz.agent.repository.AgentStepRepository;
|
||||
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
||||
import com.superbiz.agent.tool.LookupKnowledgeTool;
|
||||
import com.superbiz.agent.tool.RetrievedDocTracker;
|
||||
import com.superbiz.agent.util.QuestionComplexity;
|
||||
import com.superbiz.agent.util.SessionContextHolder;
|
||||
import com.superbiz.agent.util.VerifierContextHolder;
|
||||
|
||||
import jakarta.annotation.PostConstruct;
|
||||
import org.slf4j.Logger;
|
||||
@@ -27,11 +32,14 @@ import org.springframework.ai.chat.model.ChatModel;
|
||||
import org.springframework.ai.tool.ToolCallback;
|
||||
import org.springframework.ai.tool.ToolCallbackProvider;
|
||||
import org.springframework.beans.factory.annotation.Autowired;
|
||||
import org.springframework.beans.factory.annotation.Value;
|
||||
import org.springframework.core.io.ClassPathResource;
|
||||
import org.springframework.stereotype.Service;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.nio.charset.StandardCharsets;
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Optional;
|
||||
@@ -45,6 +53,11 @@ import java.util.UUID;
|
||||
public class ChatService {
|
||||
|
||||
private static final Logger logger = LoggerFactory.getLogger(ChatService.class);
|
||||
private static final String LOW_CONFID_DISCLAIMER = "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。";
|
||||
private static final String DEGRADED_PREFIX = "当前无法基于已获取证据生成可靠结论,建议人工介入。";
|
||||
|
||||
/** 封装 answer + 后端生成的 sessionId,用于 feedback 关联 */
|
||||
public record ChatResult(String answer, String sessionId) {}
|
||||
|
||||
@Autowired
|
||||
private InternalDocsTools internalDocsTools;
|
||||
@@ -73,9 +86,32 @@ public class ChatService {
|
||||
@Autowired
|
||||
private AgentStepRepository agentStepRepository;
|
||||
|
||||
@Autowired
|
||||
private EvaluationService evaluationService;
|
||||
|
||||
@Autowired
|
||||
private RetrievedDocTracker retrievedDocTracker;
|
||||
|
||||
@Autowired
|
||||
private KnowledgeDomainService knowledgeDomainService;
|
||||
|
||||
@Autowired
|
||||
private ToolTraceSummaryService toolTraceSummaryService;
|
||||
|
||||
@Autowired
|
||||
private SelfEvaluationMergeService selfEvaluationMergeService;
|
||||
|
||||
@Value("${verifier.low-confidence-threshold:0.5}")
|
||||
private double verifierLowConfidenceThreshold;
|
||||
|
||||
@Value("${chat.complex.retry-on-low-confidence:false}")
|
||||
private boolean retryOnLowConfidence;
|
||||
|
||||
/** 多 Agent Chat 的 Prompt */
|
||||
private String chatPlannerPrompt;
|
||||
private String chatExecutorPrompt;
|
||||
private String chatVerifierPrompt;
|
||||
private final ObjectMapper objectMapper = new ObjectMapper();
|
||||
|
||||
@PostConstruct
|
||||
public void init() {
|
||||
@@ -87,6 +123,9 @@ public class ChatService {
|
||||
chatExecutorPrompt = new String(
|
||||
new ClassPathResource("prompts/chat-executor-prompt.md").getInputStream().readAllBytes(),
|
||||
StandardCharsets.UTF_8);
|
||||
chatVerifierPrompt = new String(
|
||||
new ClassPathResource("prompts/chat-verifier-prompt.md").getInputStream().readAllBytes(),
|
||||
StandardCharsets.UTF_8);
|
||||
logger.info("Chat 多 Agent Prompts 加载成功");
|
||||
} catch (IOException e) {
|
||||
logger.error("加载 Chat Prompt 文件失败", e);
|
||||
@@ -175,16 +214,19 @@ public class ChatService {
|
||||
|
||||
/**
|
||||
* 动态构建方法工具数组
|
||||
* 根据 cls.mock-enabled 决定是否包含 QueryLogsTools
|
||||
* 根据已注入的 Bean 暴露本地工具,避免 mock/真实模式下漏注入。
|
||||
*/
|
||||
public Object[] buildMethodToolsArray() {
|
||||
List<Object> methodTools = new ArrayList<>();
|
||||
methodTools.add(dateTimeTools);
|
||||
methodTools.add(lookupKnowledgeTool);
|
||||
if (queryLogsTools != null) {
|
||||
// Mock 模式:包含 QueryLogsTools
|
||||
return new Object[]{dateTimeTools, lookupKnowledgeTool};
|
||||
} else {
|
||||
// 真实模式:不包含 QueryLogsTools(由 MCP 提供日志查询功能)
|
||||
return new Object[]{dateTimeTools, lookupKnowledgeTool, queryMetricsTools};
|
||||
methodTools.add(queryLogsTools);
|
||||
}
|
||||
if (queryMetricsTools != null) {
|
||||
methodTools.add(queryMetricsTools);
|
||||
}
|
||||
return methodTools.toArray();
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -233,22 +275,21 @@ public class ChatService {
|
||||
* 执行 ReactAgent 对话(非流式)
|
||||
* @param agent ReactAgent 实例
|
||||
* @param question 用户问题
|
||||
* @return AI 回复
|
||||
* @return ChatResult(answer + sessionId)
|
||||
*/
|
||||
public String executeChat(ReactAgent agent, String question) throws GraphRunnerException {
|
||||
public ChatResult executeChat(ReactAgent agent, String question) throws GraphRunnerException {
|
||||
return executeChat(agent, question, null);
|
||||
}
|
||||
|
||||
public ChatResult executeChat(ReactAgent agent, String question, String requestedSessionId) throws GraphRunnerException {
|
||||
logger.info("========================================");
|
||||
logger.info("📝 用户问题: {}", question);
|
||||
|
||||
String sessionId = UUID.randomUUID().toString().substring(0, 8);
|
||||
String sessionId = resolveSessionId(requestedSessionId);
|
||||
long startTime = System.currentTimeMillis();
|
||||
|
||||
// 创建诊断会话
|
||||
DiagnosisSession session = DiagnosisSession.builder()
|
||||
.sessionId(sessionId)
|
||||
.query(question)
|
||||
.status("RUNNING")
|
||||
.agentFlow("CHAT")
|
||||
.build();
|
||||
// 创建或更新诊断会话
|
||||
DiagnosisSession session = startDiagnosisSession(sessionId, question);
|
||||
diagnosisSessionRepository.save(session);
|
||||
|
||||
// 设置 ThreadLocal 上下文(LookupKnowledgeTool 通过此获取 sessionId)
|
||||
@@ -267,20 +308,24 @@ public class ChatService {
|
||||
|
||||
// 更新诊断会话
|
||||
session.setStatus("SUCCESS");
|
||||
session.setAnswer(answer);
|
||||
session.setTotalDurationMs((int) duration);
|
||||
backfillSessionMetrics(session);
|
||||
diagnosisSessionRepository.save(session);
|
||||
|
||||
evaluationService.evaluate(sessionId, answer);
|
||||
|
||||
logger.info("⏱️ 总耗时: {} ms", duration);
|
||||
logger.info("📏 输出长度: {} 字符", answer.length());
|
||||
logger.info("========================================");
|
||||
|
||||
return answer;
|
||||
return new ChatResult(answer, sessionId);
|
||||
} catch (Exception e) {
|
||||
session.setStatus("FAILED");
|
||||
diagnosisSessionRepository.save(session);
|
||||
throw e;
|
||||
} finally {
|
||||
retrievedDocTracker.clearSession(sessionId);
|
||||
SessionContextHolder.clear();
|
||||
}
|
||||
}
|
||||
@@ -291,93 +336,167 @@ public class ChatService {
|
||||
* @param toolCallbacks 工具回调
|
||||
* @param question 用户问题
|
||||
* @param history 历史消息
|
||||
* @return AI 回复
|
||||
* @return ChatResult(answer + sessionId)
|
||||
*/
|
||||
public String executeChatWithStrategy(ChatModel chatModel, ToolCallback[] toolCallbacks,
|
||||
public ChatResult executeChatWithStrategy(ChatModel chatModel, ToolCallback[] toolCallbacks,
|
||||
String question, List<Map<String, String>> history) throws GraphRunnerException {
|
||||
return executeChatWithStrategy(chatModel, toolCallbacks, question, history, null);
|
||||
}
|
||||
|
||||
public ChatResult executeChatWithStrategy(ChatModel chatModel, ToolCallback[] toolCallbacks,
|
||||
String question, List<Map<String, String>> history,
|
||||
String requestedSessionId) throws GraphRunnerException {
|
||||
if (QuestionComplexity.isComplex(question)) {
|
||||
logger.info("📊 问题判定为复杂,使用多 Agent(Planner + Executor)执行");
|
||||
return executeChatComplex(chatModel, toolCallbacks, question, history);
|
||||
return executeChatComplex(chatModel, toolCallbacks, question, history, requestedSessionId);
|
||||
} else {
|
||||
logger.info("📊 问题判定为简单,使用单 Agent 执行");
|
||||
String systemPrompt = buildSystemPrompt(history);
|
||||
ReactAgent agent = createReactAgent(chatModel, systemPrompt);
|
||||
return executeChat(agent, question);
|
||||
return executeChat(agent, question, requestedSessionId);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* 多 Agent 复杂对话执行(Planner + Executor + Supervisor)
|
||||
* 多 Agent 复杂对话执行(Planner -> Executor -> Verifier)
|
||||
*/
|
||||
public String executeChatComplex(ChatModel chatModel, ToolCallback[] toolCallbacks,
|
||||
public ChatResult executeChatComplex(ChatModel chatModel, ToolCallback[] toolCallbacks,
|
||||
String question, List<Map<String, String>> history) throws GraphRunnerException {
|
||||
String sessionId = UUID.randomUUID().toString().substring(0, 8);
|
||||
return executeChatComplex(chatModel, toolCallbacks, question, history, null);
|
||||
}
|
||||
|
||||
public ChatResult executeChatComplex(ChatModel chatModel, ToolCallback[] toolCallbacks,
|
||||
String question, List<Map<String, String>> history,
|
||||
String requestedSessionId) throws GraphRunnerException {
|
||||
String sessionId = resolveSessionId(requestedSessionId);
|
||||
long startTime = System.currentTimeMillis();
|
||||
|
||||
DiagnosisSession session = DiagnosisSession.builder()
|
||||
.sessionId(sessionId)
|
||||
.query(question)
|
||||
.status("RUNNING")
|
||||
.agentFlow("CHAT")
|
||||
.build();
|
||||
DiagnosisSession session = startDiagnosisSession(sessionId, question);
|
||||
diagnosisSessionRepository.save(session);
|
||||
|
||||
SessionContextHolder.setSessionId(sessionId);
|
||||
VerifierContextHolder.setOriginalQuery(question);
|
||||
VerifierContextHolder.setRetryContext(null);
|
||||
VerifierContextHolder.setExecutorFinalAnswer(null);
|
||||
|
||||
try {
|
||||
ReactAgent planner = buildChatPlannerAgent(chatModel, toolCallbacks, history);
|
||||
ReactAgent executor = buildChatExecutorAgent(chatModel, toolCallbacks, history);
|
||||
|
||||
SupervisorAgent supervisor = SupervisorAgent.builder()
|
||||
.name("chat_supervisor")
|
||||
.description("负责调度 Planner 与 Executor 的多 Agent 控制器")
|
||||
.model(chatModel)
|
||||
.systemPrompt("你是一个智能任务调度器。分析用户问题,调用 Planner 拆解步骤,调用 Executor 执行各步骤。")
|
||||
.subAgents(List.of(planner, executor))
|
||||
VerifierDecision finalDecision = null;
|
||||
String retryContext = null;
|
||||
String answer = null;
|
||||
RunnableConfig config = RunnableConfig.builder()
|
||||
.addMetadata("sessionId", sessionId)
|
||||
.build();
|
||||
|
||||
Optional<OverAllState> stateOptional = supervisor.invoke(question);
|
||||
long duration = System.currentTimeMillis() - startTime;
|
||||
for (int round = 1; round <= 2; round++) {
|
||||
VerifierContextHolder.setRetryContext(retryContext);
|
||||
VerifierContextHolder.setToolTraceSummary(null);
|
||||
|
||||
String answer = null;
|
||||
if (stateOptional.isPresent()) {
|
||||
// 从 state 中提取 Executor 的最终输出
|
||||
OverAllState state = stateOptional.get();
|
||||
Optional<AssistantMessage> executorOutput = state.value("executor_feedback")
|
||||
.filter(AssistantMessage.class::isInstance)
|
||||
.map(AssistantMessage.class::cast);
|
||||
if (executorOutput.isPresent()) {
|
||||
answer = executorOutput.get().getText();
|
||||
ReactAgent planner = buildChatPlannerAgent(chatModel, history, retryContext);
|
||||
ReactAgent executor = buildChatExecutorAgent(chatModel, toolCallbacks, history, retryContext);
|
||||
ReactAgent verifier = buildChatVerifierAgent(chatModel);
|
||||
|
||||
SequentialAgent workflow = SequentialAgent.builder()
|
||||
.name("chat_workflow")
|
||||
.description("按固定顺序执行 Planner、Executor、Verifier 的多 Agent 工作流")
|
||||
.subAgents(List.of(planner, executor, verifier))
|
||||
.build();
|
||||
|
||||
String workflowInput = buildWorkflowInput(question, retryContext);
|
||||
Optional<OverAllState> stateOptional = workflow.invoke(workflowInput, config);
|
||||
if (stateOptional.isEmpty()) {
|
||||
finalDecision = buildVerifierFallbackDecision(round, "workflow 未返回有效状态");
|
||||
answer = buildLowConfidenceOutput(answer, finalDecision);
|
||||
persistVerifierEvaluation(session, finalDecision, round);
|
||||
break;
|
||||
}
|
||||
|
||||
String plannerPlan = extractStateText(stateOptional, "planner_plan");
|
||||
answer = extractStateText(stateOptional, "executor_feedback");
|
||||
VerifierContextHolder.setExecutorFinalAnswer(answer);
|
||||
String verifierOutput = extractStateText(stateOptional, "verifier_output");
|
||||
if ((verifierOutput == null || verifierOutput.isBlank()) && answer != null && !answer.isBlank()) {
|
||||
verifierOutput = invokeVerifierFallback(verifier, question, round, config);
|
||||
}
|
||||
finalDecision = parseVerifierDecision(verifierOutput, round);
|
||||
logger.debug("Sequential workflow round {} finished: plannerPlanLength={}, answerLength={}, verifierOutputLength={}",
|
||||
round,
|
||||
plannerPlan != null ? plannerPlan.length() : 0,
|
||||
answer != null ? answer.length() : 0,
|
||||
verifierOutput != null ? verifierOutput.length() : 0);
|
||||
|
||||
if (finalDecision == null) {
|
||||
finalDecision = buildVerifierFallbackDecision(round, "verifier_output 缺失或无法解析");
|
||||
answer = buildLowConfidenceOutput(answer, finalDecision);
|
||||
persistVerifierEvaluation(session, finalDecision, round);
|
||||
break;
|
||||
}
|
||||
|
||||
if ("PASS".equals(finalDecision.verdict())) {
|
||||
answer = answer == null || answer.isBlank() ? "抱歉,多 Agent 分析未能生成有效结论。" : answer;
|
||||
persistVerifierEvaluation(session, finalDecision, round);
|
||||
break;
|
||||
}
|
||||
|
||||
if ("REJECT".equals(finalDecision.verdict())) {
|
||||
answer = buildDegradedOutput(finalDecision);
|
||||
persistVerifierEvaluation(session, finalDecision, round);
|
||||
break;
|
||||
}
|
||||
|
||||
boolean shouldRetry = retryOnLowConfidence
|
||||
&& finalDecision.groundednessScore() < verifierLowConfidenceThreshold
|
||||
&& round < 2;
|
||||
if (!shouldRetry) {
|
||||
answer = buildLowConfidenceOutput(answer, finalDecision);
|
||||
persistVerifierEvaluation(session, finalDecision, round);
|
||||
break;
|
||||
}
|
||||
|
||||
retryContext = buildRetryContext(finalDecision);
|
||||
persistVerifierEvaluation(session, finalDecision, round);
|
||||
}
|
||||
|
||||
long duration = System.currentTimeMillis() - startTime;
|
||||
|
||||
if (answer == null || answer.isBlank()) {
|
||||
answer = "抱歉,多 Agent 分析未能生成有效结论。";
|
||||
}
|
||||
|
||||
session.setStatus("SUCCESS");
|
||||
session.setAnswer(answer);
|
||||
session.setTotalDurationMs((int) duration);
|
||||
backfillSessionMetrics(session);
|
||||
diagnosisSessionRepository.save(session);
|
||||
|
||||
evaluationService.evaluate(sessionId, answer);
|
||||
|
||||
logger.info("⏱️ 多 Agent 总耗时: {} ms", duration);
|
||||
logger.info("📏 输出长度: {} 字符", answer.length());
|
||||
|
||||
return answer;
|
||||
return new ChatResult(answer, sessionId);
|
||||
|
||||
} catch (Exception e) {
|
||||
session.setStatus("FAILED");
|
||||
diagnosisSessionRepository.save(session);
|
||||
logger.error("多 Agent 执行失败", e);
|
||||
return "执行失败: " + e.getMessage();
|
||||
return new ChatResult("执行失败: " + e.getMessage(), sessionId);
|
||||
} finally {
|
||||
retrievedDocTracker.clearSession(sessionId);
|
||||
SessionContextHolder.clear();
|
||||
VerifierContextHolder.clear();
|
||||
}
|
||||
}
|
||||
|
||||
private ReactAgent buildChatPlannerAgent(ChatModel chatModel, ToolCallback[] toolCallbacks,
|
||||
List<Map<String, String>> history) {
|
||||
private ReactAgent buildChatPlannerAgent(ChatModel chatModel, List<Map<String, String>> history,
|
||||
String retryContext) {
|
||||
StringBuilder prompt = new StringBuilder(chatPlannerPrompt);
|
||||
|
||||
// 注入 knowledge map
|
||||
String knowledgeMap = knowledgeDomainService.buildKnowledgeMap();
|
||||
if (!knowledgeMap.isBlank()) {
|
||||
prompt.append("\n\n## 可用知识库\n\n").append(knowledgeMap);
|
||||
}
|
||||
|
||||
if (!history.isEmpty()) {
|
||||
prompt.append("\n\n--- 对话历史 ---\n");
|
||||
for (Map<String, String> msg : history) {
|
||||
@@ -385,19 +504,33 @@ public class ChatService {
|
||||
}
|
||||
prompt.append("--- 对话历史结束 ---\n");
|
||||
}
|
||||
if (retryContext != null && !retryContext.isBlank()) {
|
||||
prompt.append("\n\n--- 本轮补证据约束 ---\n").append(retryContext).append("\n");
|
||||
}
|
||||
return ReactAgent.builder()
|
||||
.name("chat_planner")
|
||||
.description("负责拆解问题、规划步骤")
|
||||
.model(chatModel)
|
||||
.systemPrompt(prompt.toString())
|
||||
// Planner 不注入工具,只能规划不能执行
|
||||
.hooks(new AgentLoggingHook(agentStepRepository, "planner"))
|
||||
.outputKey("planner_plan")
|
||||
.build();
|
||||
}
|
||||
|
||||
private ReactAgent buildChatVerifierAgent(ChatModel chatModel) {
|
||||
return ReactAgent.builder()
|
||||
.name("chat_verifier")
|
||||
.description("负责验证 Executor 答案的事实准确性")
|
||||
.model(chatModel)
|
||||
.systemPrompt(chatVerifierPrompt)
|
||||
.hooks(new AgentLoggingHook(agentStepRepository, "verifier"),
|
||||
new VerifierInputHook(toolTraceSummaryService))
|
||||
.outputKey("verifier_output")
|
||||
.build();
|
||||
}
|
||||
|
||||
private ReactAgent buildChatExecutorAgent(ChatModel chatModel, ToolCallback[] toolCallbacks,
|
||||
List<Map<String, String>> history) {
|
||||
List<Map<String, String>> history, String retryContext) {
|
||||
StringBuilder prompt = new StringBuilder(chatExecutorPrompt);
|
||||
if (!history.isEmpty()) {
|
||||
prompt.append("\n\n--- 对话历史 ---\n");
|
||||
@@ -406,6 +539,9 @@ public class ChatService {
|
||||
}
|
||||
prompt.append("--- 对话历史结束 ---\n");
|
||||
}
|
||||
if (retryContext != null && !retryContext.isBlank()) {
|
||||
prompt.append("\n\n--- 本轮补证据约束 ---\n").append(retryContext).append("\n");
|
||||
}
|
||||
return ReactAgent.builder()
|
||||
.name("chat_executor")
|
||||
.description("负责执行具体步骤并及时反馈")
|
||||
@@ -418,6 +554,290 @@ public class ChatService {
|
||||
.build();
|
||||
}
|
||||
|
||||
private String resolveSessionId(String requestedSessionId) {
|
||||
if (requestedSessionId != null && !requestedSessionId.isBlank()) {
|
||||
return requestedSessionId;
|
||||
}
|
||||
return UUID.randomUUID().toString().substring(0, 8);
|
||||
}
|
||||
|
||||
private DiagnosisSession startDiagnosisSession(String sessionId, String question) {
|
||||
DiagnosisSession session = diagnosisSessionRepository.findBySessionId(sessionId)
|
||||
.orElseGet(() -> DiagnosisSession.builder()
|
||||
.sessionId(sessionId)
|
||||
.agentFlow("CHAT")
|
||||
.build());
|
||||
session.setQuery(question);
|
||||
session.setStatus("RUNNING");
|
||||
session.setAgentFlow("CHAT");
|
||||
session.setAnswer(null);
|
||||
session.setTotalDurationMs(null);
|
||||
session.setTotalTokenCount(null);
|
||||
session.setStepCount(null);
|
||||
session.setToolCallCount(null);
|
||||
return session;
|
||||
}
|
||||
|
||||
private String buildWorkflowInput(String question, String retryContext) {
|
||||
StringBuilder input = new StringBuilder();
|
||||
input.append("请按固定工作流完成本轮 Planner -> Executor -> Verifier。\n\n");
|
||||
input.append("--- 用户问题 ---\n").append(question);
|
||||
if (retryContext != null && !retryContext.isBlank()) {
|
||||
input.append("\n\n--- retry_context ---\n").append(retryContext);
|
||||
}
|
||||
input.append("\n\nVerifier 完成后由外层代码读取 verifier_output 并决定最终用户输出。");
|
||||
return input.toString();
|
||||
}
|
||||
|
||||
private String invokeVerifierFallback(ReactAgent verifier, String question, int round, RunnableConfig config) {
|
||||
try {
|
||||
logger.warn("Sequential workflow round {} finished without verifier_output, invoking chat_verifier fallback", round);
|
||||
return verifier.call("请基于 executor_final_answer 和 tool_trace_summary 输出 verifier JSON。原始问题:" + question, config)
|
||||
.getText();
|
||||
} catch (Exception e) {
|
||||
logger.error("chat_verifier fallback 执行失败", e);
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
private VerifierDecision parseVerifierDecision(String verifierOutput, int round) {
|
||||
if (verifierOutput == null || verifierOutput.isBlank()) {
|
||||
return null;
|
||||
}
|
||||
|
||||
try {
|
||||
JsonNode root = objectMapper.readTree(sanitizeJsonPayload(verifierOutput));
|
||||
List<Map<String, Object>> factsChecked = parseFactsChecked(root.path("facts_checked"));
|
||||
|
||||
return new VerifierDecision(
|
||||
root.path("verdict").asText("LOW_CONFID"),
|
||||
root.path("groundedness_score").asDouble(0.0),
|
||||
root.path("critical_fact_count").asInt(0),
|
||||
factsChecked,
|
||||
root.path("rationale").asText(""),
|
||||
round
|
||||
);
|
||||
} catch (Exception e) {
|
||||
logger.error("解析 verifier_output 失败: {}", verifierOutput, e);
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
private String sanitizeJsonPayload(String raw) {
|
||||
String trimmed = raw.trim();
|
||||
if (trimmed.startsWith("```")) {
|
||||
int firstNewline = trimmed.indexOf('\n');
|
||||
int lastFence = trimmed.lastIndexOf("```");
|
||||
if (firstNewline >= 0 && lastFence > firstNewline) {
|
||||
return trimmed.substring(firstNewline + 1, lastFence).trim();
|
||||
}
|
||||
}
|
||||
return trimmed;
|
||||
}
|
||||
|
||||
private List<Map<String, Object>> parseFactsChecked(JsonNode factsNode) {
|
||||
List<Map<String, Object>> factsChecked = new ArrayList<>();
|
||||
if (!factsNode.isArray()) {
|
||||
return factsChecked;
|
||||
}
|
||||
for (JsonNode factNode : factsNode) {
|
||||
Map<String, Object> fact = new LinkedHashMap<>();
|
||||
fact.put("fact", factNode.path("fact").asText(""));
|
||||
fact.put("is_critical", factNode.path("is_critical").asBoolean(false));
|
||||
fact.put("verification", factNode.path("verification").asText(""));
|
||||
fact.put("detail", factNode.path("detail").asText(""));
|
||||
fact.put("evidence_refs", parseEvidenceRefs(factNode.path("evidence_refs")));
|
||||
factsChecked.add(fact);
|
||||
}
|
||||
return factsChecked;
|
||||
}
|
||||
|
||||
private List<Map<String, Object>> parseEvidenceRefs(JsonNode evidenceRefsNode) {
|
||||
List<Map<String, Object>> evidenceRefs = new ArrayList<>();
|
||||
if (!evidenceRefsNode.isArray()) {
|
||||
return evidenceRefs;
|
||||
}
|
||||
for (JsonNode refNode : evidenceRefsNode) {
|
||||
Map<String, Object> evidenceRef = new LinkedHashMap<>();
|
||||
evidenceRef.put("trace_ref", refNode.path("trace_ref").asText(""));
|
||||
evidenceRef.put("tool_name", refNode.path("tool_name").asText(""));
|
||||
evidenceRef.put("topic_domain", refNode.path("topic_domain").asText(""));
|
||||
evidenceRef.put("note", refNode.path("note").asText(""));
|
||||
|
||||
List<Long> sourceInvocationIds = new ArrayList<>();
|
||||
JsonNode idsNode = refNode.path("source_invocation_ids");
|
||||
if (idsNode.isArray()) {
|
||||
for (JsonNode idNode : idsNode) {
|
||||
if (idNode.canConvertToLong()) {
|
||||
sourceInvocationIds.add(idNode.asLong());
|
||||
}
|
||||
}
|
||||
}
|
||||
evidenceRef.put("source_invocation_ids", sourceInvocationIds);
|
||||
evidenceRefs.add(evidenceRef);
|
||||
}
|
||||
return evidenceRefs;
|
||||
}
|
||||
|
||||
private VerifierDecision buildVerifierFallbackDecision(int round, String rationale) {
|
||||
return new VerifierDecision("LOW_CONFID", 0.0, 0, List.of(), rationale, round);
|
||||
}
|
||||
|
||||
private String extractStateText(Optional<OverAllState> stateOptional, String key) {
|
||||
if (stateOptional.isEmpty()) {
|
||||
return null;
|
||||
}
|
||||
return stateOptional.get().value(key)
|
||||
.map(value -> {
|
||||
if (value instanceof AssistantMessage assistantMessage) {
|
||||
return assistantMessage.getText();
|
||||
}
|
||||
return String.valueOf(value);
|
||||
})
|
||||
.orElse(null);
|
||||
}
|
||||
|
||||
private void persistVerifierEvaluation(DiagnosisSession session, VerifierDecision decision, int round) {
|
||||
if (decision == null) {
|
||||
return;
|
||||
}
|
||||
Map<String, Object> verifierEvaluation = new LinkedHashMap<>();
|
||||
verifierEvaluation.put("verdict", decision.verdict());
|
||||
verifierEvaluation.put("groundedness_score", decision.groundednessScore());
|
||||
verifierEvaluation.put("critical_fact_count", decision.criticalFactCount());
|
||||
verifierEvaluation.put("facts_checked", decision.factsChecked());
|
||||
verifierEvaluation.put("rationale", decision.rationale());
|
||||
verifierEvaluation.put("round", round);
|
||||
verifierEvaluation.put("traceability_version", "v1");
|
||||
verifierEvaluation.put("tool_trace_summary",
|
||||
Optional.ofNullable(VerifierContextHolder.getToolTraceSummary()).orElse(List.of()));
|
||||
|
||||
String merged = selfEvaluationMergeService.mergeVerifierEvaluation(session.getSelfEvaluation(), verifierEvaluation);
|
||||
session.setSelfEvaluation(merged);
|
||||
diagnosisSessionRepository.save(session);
|
||||
}
|
||||
|
||||
private String buildRetryContext(VerifierDecision decision) {
|
||||
try {
|
||||
List<String> missingFacts = extractEvidenceGaps(decision);
|
||||
Map<String, Object> retryContext = new LinkedHashMap<>();
|
||||
retryContext.put("round", decision.round());
|
||||
retryContext.put("missing_evidence_facts", missingFacts);
|
||||
retryContext.put("instruction", "仅补充以上断言相关证据,不要重复已完成检索");
|
||||
return objectMapper.writeValueAsString(retryContext);
|
||||
} catch (Exception e) {
|
||||
logger.error("构造 retry_context 失败", e);
|
||||
return "{\"round\":1,\"missing_evidence_facts\":[],\"instruction\":\"仅补充缺失证据\"}";
|
||||
}
|
||||
}
|
||||
|
||||
private String buildLowConfidenceOutput(String executorAnswer, VerifierDecision decision) {
|
||||
StringBuilder output = new StringBuilder(LOW_CONFID_DISCLAIMER);
|
||||
output.append("\n\n").append(executorAnswer == null ? "" : executorAnswer);
|
||||
|
||||
List<String> gaps = extractEvidenceGaps(decision);
|
||||
if (!gaps.isEmpty()) {
|
||||
output.append("\n\n当前缺口:");
|
||||
for (String gap : gaps) {
|
||||
output.append("\n- ").append(gap);
|
||||
}
|
||||
}
|
||||
return output.toString();
|
||||
}
|
||||
|
||||
private String buildDegradedOutput(VerifierDecision decision) {
|
||||
StringBuilder output = new StringBuilder(DEGRADED_PREFIX);
|
||||
|
||||
List<String> confirmedFacts = extractConfirmedFacts(decision);
|
||||
List<String> gaps = extractEvidenceGaps(decision);
|
||||
List<String> suggestions = buildNextStepSuggestions(decision);
|
||||
|
||||
output.append("\n\n已确认信息:");
|
||||
if (confirmedFacts.isEmpty()) {
|
||||
output.append("\n- 暂无可稳定确认的信息");
|
||||
} else {
|
||||
for (String fact : confirmedFacts) {
|
||||
output.append("\n- ").append(fact);
|
||||
}
|
||||
}
|
||||
|
||||
output.append("\n\n证据缺口:");
|
||||
if (gaps.isEmpty()) {
|
||||
output.append("\n- 当前缺少足够的直接证据支撑核心结论");
|
||||
} else {
|
||||
for (String gap : gaps) {
|
||||
output.append("\n- ").append(gap);
|
||||
}
|
||||
}
|
||||
|
||||
output.append("\n\n建议下一步:");
|
||||
for (String suggestion : suggestions) {
|
||||
output.append("\n- ").append(suggestion);
|
||||
}
|
||||
return output.toString();
|
||||
}
|
||||
|
||||
private List<String> extractConfirmedFacts(VerifierDecision decision) {
|
||||
List<String> confirmedFacts = new ArrayList<>();
|
||||
for (Map<String, Object> fact : decision.factsChecked()) {
|
||||
String verification = String.valueOf(fact.get("verification"));
|
||||
boolean critical = Boolean.TRUE.equals(fact.get("is_critical"));
|
||||
if (critical && ("direct_evidence".equals(verification) || "indirect_support".equals(verification))) {
|
||||
confirmedFacts.add(String.valueOf(fact.get("fact")));
|
||||
}
|
||||
}
|
||||
return confirmedFacts;
|
||||
}
|
||||
|
||||
private List<String> extractEvidenceGaps(VerifierDecision decision) {
|
||||
List<String> gaps = new ArrayList<>();
|
||||
for (Map<String, Object> fact : decision.factsChecked()) {
|
||||
String verification = String.valueOf(fact.get("verification"));
|
||||
boolean critical = Boolean.TRUE.equals(fact.get("is_critical"));
|
||||
if (critical && ("no_evidence".equals(verification) || "contradicted".equals(verification))) {
|
||||
gaps.add(String.valueOf(fact.get("fact")) + ":" + String.valueOf(fact.get("detail")));
|
||||
}
|
||||
}
|
||||
if (gaps.isEmpty() && "LOW_CONFID".equals(decision.verdict())) {
|
||||
for (Map<String, Object> fact : decision.factsChecked()) {
|
||||
String verification = String.valueOf(fact.get("verification"));
|
||||
boolean critical = Boolean.TRUE.equals(fact.get("is_critical"));
|
||||
if (critical && "indirect_support".equals(verification)) {
|
||||
gaps.add(String.valueOf(fact.get("fact")) + ":缺少直接证据锚点");
|
||||
}
|
||||
}
|
||||
}
|
||||
return gaps;
|
||||
}
|
||||
|
||||
private List<String> buildNextStepSuggestions(VerifierDecision decision) {
|
||||
List<String> suggestions = new ArrayList<>();
|
||||
List<Map<String, Object>> toolSummary = toolTraceSummaryService.buildVerifierTraceSummary(SessionContextHolder.getSessionId(), null);
|
||||
boolean hasKnowledgeTool = toolSummary.stream().anyMatch(item -> "lookup_knowledge".equals(item.get("tool_name")));
|
||||
boolean hasFailedEvidence = toolSummary.stream().anyMatch(item -> !Boolean.TRUE.equals(item.get("success")));
|
||||
|
||||
if (!hasKnowledgeTool) {
|
||||
suggestions.add("补充知识库或业务文档检索结果,建立可引用的证据锚点");
|
||||
}
|
||||
if (hasFailedEvidence) {
|
||||
suggestions.add("优先重试失败的证据型查询,补齐日志、指标或知识库侧证据");
|
||||
}
|
||||
if (suggestions.isEmpty()) {
|
||||
suggestions.add("围绕上述证据缺口补充只读查询,再由人工复核最终结论");
|
||||
}
|
||||
return suggestions;
|
||||
}
|
||||
|
||||
private record VerifierDecision(
|
||||
String verdict,
|
||||
double groundednessScore,
|
||||
int criticalFactCount,
|
||||
List<Map<String, Object>> factsChecked,
|
||||
String rationale,
|
||||
int round
|
||||
) {
|
||||
}
|
||||
|
||||
/** 从 agent_step 汇总 token、步数等指标回填 diagnosis_session */
|
||||
private void backfillSessionMetrics(DiagnosisSession session) {
|
||||
try {
|
||||
|
||||
@@ -0,0 +1,136 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.fasterxml.jackson.core.type.TypeReference;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.domain.entity.AgentStep;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||
import com.superbiz.agent.domain.entity.ToolInvocation;
|
||||
import com.superbiz.agent.dto.DiagnosisTraceResponse;
|
||||
import com.superbiz.agent.exception.SessionNotFoundException;
|
||||
import com.superbiz.agent.repository.AgentStepRepository;
|
||||
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
||||
import com.superbiz.agent.repository.ToolInvocationRepository;
|
||||
import lombok.RequiredArgsConstructor;
|
||||
import org.springframework.stereotype.Service;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
@Service
|
||||
@RequiredArgsConstructor
|
||||
public class DiagnosisTraceService {
|
||||
|
||||
private static final TypeReference<Map<String, Object>> JSON_MAP_TYPE = new TypeReference<>() {
|
||||
};
|
||||
|
||||
private final DiagnosisSessionRepository diagnosisSessionRepository;
|
||||
private final AgentStepRepository agentStepRepository;
|
||||
private final ToolInvocationRepository toolInvocationRepository;
|
||||
private final ObjectMapper objectMapper;
|
||||
|
||||
public DiagnosisTraceResponse getTrace(String sessionId) {
|
||||
DiagnosisSession session = diagnosisSessionRepository.findBySessionId(sessionId)
|
||||
.orElseThrow(() -> new SessionNotFoundException(sessionId));
|
||||
List<AgentStep> steps = agentStepRepository.findBySessionIdOrderByStepIndex(sessionId);
|
||||
List<ToolInvocation> toolInvocations = toolInvocationRepository.findBySessionIdOrderByIdAsc(sessionId);
|
||||
|
||||
return DiagnosisTraceResponse.builder()
|
||||
.session(toSessionTrace(session))
|
||||
.steps(steps.stream().map(this::toAgentStepTrace).toList())
|
||||
.toolInvocations(toolInvocations.stream().map(this::toToolInvocationTrace).toList())
|
||||
.summary(toSummary(session, steps, toolInvocations))
|
||||
.build();
|
||||
}
|
||||
|
||||
private DiagnosisTraceResponse.SessionTrace toSessionTrace(DiagnosisSession session) {
|
||||
return DiagnosisTraceResponse.SessionTrace.builder()
|
||||
.id(session.getId())
|
||||
.sessionId(session.getSessionId())
|
||||
.query(session.getQuery())
|
||||
.status(session.getStatus())
|
||||
.agentFlow(session.getAgentFlow())
|
||||
.totalDurationMs(session.getTotalDurationMs())
|
||||
.totalTokenCount(session.getTotalTokenCount())
|
||||
.stepCount(session.getStepCount())
|
||||
.toolCallCount(session.getToolCallCount())
|
||||
.answer(session.getAnswer())
|
||||
.selfEvaluationRaw(session.getSelfEvaluation())
|
||||
.selfEvaluation(parseJsonObject(session.getSelfEvaluation()))
|
||||
.feedback(session.getFeedback())
|
||||
.createdAt(session.getCreatedAt())
|
||||
.updatedAt(session.getUpdatedAt())
|
||||
.build();
|
||||
}
|
||||
|
||||
private DiagnosisTraceResponse.AgentStepTrace toAgentStepTrace(AgentStep step) {
|
||||
return DiagnosisTraceResponse.AgentStepTrace.builder()
|
||||
.id(step.getId())
|
||||
.sessionId(step.getSessionId())
|
||||
.stepIndex(step.getStepIndex())
|
||||
.agentName(step.getAgentName())
|
||||
.modelInput(step.getModelInput())
|
||||
.modelOutput(step.getModelOutput())
|
||||
.thought(step.getThought())
|
||||
.hasToolCall(step.getHasToolCall())
|
||||
.durationMs(step.getDurationMs())
|
||||
.tokenCount(step.getTokenCount())
|
||||
.createdAt(step.getCreatedAt())
|
||||
.build();
|
||||
}
|
||||
|
||||
private DiagnosisTraceResponse.ToolInvocationTrace toToolInvocationTrace(ToolInvocation invocation) {
|
||||
return DiagnosisTraceResponse.ToolInvocationTrace.builder()
|
||||
.id(invocation.getId())
|
||||
.sessionId(invocation.getSessionId())
|
||||
.stepId(invocation.getStepId())
|
||||
.toolName(invocation.getToolName())
|
||||
.inputParamsRaw(invocation.getInputParams())
|
||||
.inputParams(parseJsonObject(invocation.getInputParams()))
|
||||
.outputPreview(invocation.getOutputPreview())
|
||||
.outputLength(invocation.getOutputLength())
|
||||
.retrievalLayer(invocation.getRetrievalLayer())
|
||||
.l0MatchCount(invocation.getL0MatchCount())
|
||||
.l1MatchCount(invocation.getL1MatchCount())
|
||||
.truncated(invocation.getIsTruncated())
|
||||
.relevanceLevel(invocation.getRelevanceLevel())
|
||||
.dedupReason(invocation.getDedupReason())
|
||||
.retrievalDetailsRaw(invocation.getRetrievalDetails())
|
||||
.retrievalDetails(parseJsonObject(invocation.getRetrievalDetails()))
|
||||
.durationMs(invocation.getDurationMs())
|
||||
.success(invocation.getSuccess())
|
||||
.errorMessage(invocation.getErrorMessage())
|
||||
.createdAt(invocation.getCreatedAt())
|
||||
.build();
|
||||
}
|
||||
|
||||
private DiagnosisTraceResponse.TraceSummary toSummary(
|
||||
DiagnosisSession session,
|
||||
List<AgentStep> steps,
|
||||
List<ToolInvocation> toolInvocations
|
||||
) {
|
||||
Map<String, Object> selfEvaluation = parseJsonObject(session.getSelfEvaluation());
|
||||
return DiagnosisTraceResponse.TraceSummary.builder()
|
||||
.persistedStepCount(defaultInt(session.getStepCount()))
|
||||
.returnedStepCount(steps.size())
|
||||
.persistedToolCallCount(defaultInt(session.getToolCallCount()))
|
||||
.returnedToolCallCount(toolInvocations.size())
|
||||
.hasVerifierEvaluation(selfEvaluation != null && selfEvaluation.containsKey("verifier_evaluation"))
|
||||
.hasFeedback(session.getFeedback() != null && !session.getFeedback().isBlank())
|
||||
.build();
|
||||
}
|
||||
|
||||
private Map<String, Object> parseJsonObject(String json) {
|
||||
if (json == null || json.isBlank()) {
|
||||
return null;
|
||||
}
|
||||
try {
|
||||
return objectMapper.readValue(json, JSON_MAP_TYPE);
|
||||
} catch (Exception ignored) {
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
private int defaultInt(Integer value) {
|
||||
return value == null ? 0 : value;
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,135 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.dto.Frontmatter;
|
||||
import com.superbiz.agent.dto.KnowledgeEntry;
|
||||
import lombok.extern.slf4j.Slf4j;
|
||||
import org.springframework.ai.chat.model.ChatModel;
|
||||
import org.springframework.ai.chat.prompt.Prompt;
|
||||
import org.springframework.beans.factory.annotation.Autowired;
|
||||
import org.springframework.core.io.ClassPathResource;
|
||||
import org.springframework.stereotype.Service;
|
||||
|
||||
import jakarta.annotation.PostConstruct;
|
||||
import java.io.IOException;
|
||||
import java.nio.charset.StandardCharsets;
|
||||
import java.util.ArrayList;
|
||||
import java.util.List;
|
||||
import java.util.stream.Collectors;
|
||||
|
||||
/**
|
||||
* 文档字段补全服务
|
||||
* 上传时调用 LLM 生成 covers 和 whenToRetrieve
|
||||
*/
|
||||
@Slf4j
|
||||
@Service
|
||||
public class DocumentFieldEnricher {
|
||||
|
||||
@Autowired
|
||||
private ChatModel chatModel;
|
||||
|
||||
@Autowired
|
||||
private ObjectMapper objectMapper;
|
||||
|
||||
@Autowired
|
||||
private KnowledgeIndexService knowledgeIndexService;
|
||||
|
||||
private String promptTemplate;
|
||||
|
||||
@PostConstruct
|
||||
public void init() {
|
||||
try {
|
||||
promptTemplate = new String(
|
||||
new ClassPathResource("prompts/doc-field-enricher-prompt.md").getInputStream().readAllBytes(),
|
||||
StandardCharsets.UTF_8);
|
||||
log.info("DocumentFieldEnricher prompt 加载成功");
|
||||
} catch (IOException e) {
|
||||
log.error("加载 doc-field-enricher-prompt.md 失败", e);
|
||||
throw new RuntimeException("Failed to load doc-field-enricher prompt", e);
|
||||
}
|
||||
}
|
||||
|
||||
public void enrich(Frontmatter frontmatter, String bodyText) {
|
||||
enrich(frontmatter, bodyText, null);
|
||||
}
|
||||
|
||||
/**
|
||||
* 为 Frontmatter 补全 covers 和 whenToRetrieve
|
||||
* 若已有值则跳过;LLM 失败时降级,不阻断主流程
|
||||
*
|
||||
* @param frontmatter 待补全的 frontmatter
|
||||
* @param bodyText 文档正文
|
||||
* @param category 文档所属域(用于查找同域其他文档)
|
||||
*/
|
||||
public void enrich(Frontmatter frontmatter, String bodyText, String category) {
|
||||
if (frontmatter == null) return;
|
||||
|
||||
boolean needsCovers = frontmatter.getCovers() == null || frontmatter.getCovers().isEmpty();
|
||||
boolean needsWhen = frontmatter.getWhenToRetrieve() == null || frontmatter.getWhenToRetrieve().isBlank();
|
||||
|
||||
if (!needsCovers && !needsWhen) {
|
||||
log.debug("covers 和 whenToRetrieve 已存在,跳过 LLM 生成");
|
||||
return;
|
||||
}
|
||||
|
||||
try {
|
||||
String snippet = bodyText != null && bodyText.length() > 1000
|
||||
? bodyText.substring(0, 1000) : (bodyText != null ? bodyText : "");
|
||||
|
||||
String sameDomainDocs = buildSameDomainDocs(frontmatter.getTitle(), category);
|
||||
|
||||
String promptText = String.format(promptTemplate,
|
||||
frontmatter.getTitle(),
|
||||
frontmatter.getSummary(),
|
||||
sameDomainDocs,
|
||||
snippet);
|
||||
|
||||
String response = chatModel.call(new Prompt(promptText))
|
||||
.getResult().getOutput().getText();
|
||||
|
||||
// 提取 JSON 部分(防止模型输出多余文本)
|
||||
String json = extractJson(response);
|
||||
JsonNode node = objectMapper.readTree(json);
|
||||
|
||||
if (needsCovers && node.has("covers")) {
|
||||
List<String> covers = new ArrayList<>();
|
||||
node.get("covers").forEach(n -> covers.add(n.asText()));
|
||||
frontmatter.setCovers(covers);
|
||||
log.debug("LLM 生成 covers: {}", covers);
|
||||
}
|
||||
|
||||
if (needsWhen && node.has("whenToRetrieve")) {
|
||||
frontmatter.setWhenToRetrieve(node.get("whenToRetrieve").asText());
|
||||
log.debug("LLM 生成 whenToRetrieve: {}", frontmatter.getWhenToRetrieve());
|
||||
}
|
||||
|
||||
} catch (Exception e) {
|
||||
log.warn("LLM 生成文档字段失败,降级处理: title={}", frontmatter.getTitle(), e);
|
||||
if (needsCovers) frontmatter.setCovers(List.of());
|
||||
if (needsWhen) frontmatter.setWhenToRetrieve(frontmatter.getSummary());
|
||||
}
|
||||
}
|
||||
|
||||
private String extractJson(String text) {
|
||||
if (text == null) return "{}";
|
||||
int start = text.indexOf('{');
|
||||
int end = text.lastIndexOf('}');
|
||||
if (start == -1 || end == -1 || end <= start) return "{}";
|
||||
return text.substring(start, end + 1);
|
||||
}
|
||||
|
||||
/**
|
||||
* 构建同域其他文档标题列表(供 LLM 做排除判断)
|
||||
*/
|
||||
private String buildSameDomainDocs(String currentTitle, String category) {
|
||||
if (category == null || category.isBlank()) return "(无同域文档信息)";
|
||||
List<String> otherTitles = knowledgeIndexService.getAllEntries().stream()
|
||||
.filter(e -> category.equals(e.getCategory()))
|
||||
.map(KnowledgeEntry::getTitle)
|
||||
.filter(t -> t != null && !t.equals(currentTitle))
|
||||
.collect(Collectors.toList());
|
||||
if (otherTitles.isEmpty()) return "(无同域其他文档)";
|
||||
return String.join("、", otherTitles);
|
||||
}
|
||||
}
|
||||
@@ -58,6 +58,12 @@ public class DocumentManagementService {
|
||||
@Autowired
|
||||
private KnowledgeIndexService knowledgeIndexService;
|
||||
|
||||
@Autowired
|
||||
private DocumentFieldEnricher documentFieldEnricher;
|
||||
|
||||
@Autowired
|
||||
private KnowledgeDomainService knowledgeDomainService;
|
||||
|
||||
@Autowired
|
||||
private ObjectMapper objectMapper;
|
||||
|
||||
@@ -123,6 +129,8 @@ public class DocumentManagementService {
|
||||
if (frontmatterParser.hasFrontmatter(text)) {
|
||||
frontmatter = frontmatterParser.parse(text);
|
||||
if (frontmatter != null) {
|
||||
// LLM 补全 covers / whenToRetrieve(已有值则跳过)
|
||||
documentFieldEnricher.enrich(frontmatter, text, category);
|
||||
log.info("解析到frontmatter: title={}, keywords={}, time={}ms",
|
||||
frontmatter.getTitle(), frontmatter.getKeywords(), System.currentTimeMillis() - frontmatterStart);
|
||||
} else {
|
||||
@@ -196,12 +204,17 @@ public class DocumentManagementService {
|
||||
.summary(frontmatter.getSummary())
|
||||
.category(category)
|
||||
.sections(frontmatter.getSections())
|
||||
.covers(frontmatter.getCovers())
|
||||
.whenToRetrieve(frontmatter.getWhenToRetrieve())
|
||||
.build();
|
||||
|
||||
knowledgeIndexService.addToIndex(entry);
|
||||
log.info("文档已加入L0索引: docId={}, title={}", docId, frontmatter.getTitle());
|
||||
}
|
||||
|
||||
// 触发域级聚合重算
|
||||
knowledgeDomainService.onDocumentChange(category);
|
||||
|
||||
long totalTime = System.currentTimeMillis() - startTime;
|
||||
log.info("文档上传完成: docId={}, fileName={}, hasFrontmatter={}, totalTime={}ms",
|
||||
docId, fileName, frontmatter != null, totalTime);
|
||||
@@ -247,7 +260,8 @@ public class DocumentManagementService {
|
||||
private String saveToLocal(MultipartFile file, String fileName, String category) {
|
||||
try {
|
||||
// 1. 构建目标路径
|
||||
Path categoryDir = Paths.get(knowledgeBasePath, category);
|
||||
Path baseDir = Paths.get(knowledgeBasePath).normalize();
|
||||
Path categoryDir = baseDir.resolve(category).normalize();
|
||||
Files.createDirectories(categoryDir);
|
||||
|
||||
Path targetPath = categoryDir.resolve(fileName);
|
||||
@@ -255,8 +269,9 @@ public class DocumentManagementService {
|
||||
// 2. 保存文件
|
||||
file.transferTo(targetPath.toFile());
|
||||
|
||||
log.info("文件已保存到本地: {}", targetPath);
|
||||
return targetPath.toString();
|
||||
String relativePath = baseDir.relativize(targetPath.normalize()).toString().replace("\\", "/");
|
||||
log.info("文件已保存到本地: {}, storedPath={}", targetPath, relativePath);
|
||||
return relativePath;
|
||||
|
||||
} catch (IOException e) {
|
||||
throw new DocumentProcessException(
|
||||
@@ -274,7 +289,7 @@ public class DocumentManagementService {
|
||||
private void cleanupLocalFile(String localPath) {
|
||||
if (localPath != null) {
|
||||
try {
|
||||
Files.deleteIfExists(Paths.get(localPath));
|
||||
Files.deleteIfExists(resolveLocalPath(localPath));
|
||||
log.info("已清理本地文件: {}", localPath);
|
||||
} catch (IOException e) {
|
||||
log.warn("清理本地文件失败: {}", localPath, e);
|
||||
@@ -344,7 +359,7 @@ public class DocumentManagementService {
|
||||
// 删除本地文件
|
||||
if (doc.getFilePath() != null) {
|
||||
try {
|
||||
Files.deleteIfExists(Paths.get(doc.getFilePath()));
|
||||
Files.deleteIfExists(resolveLocalPath(doc.getFilePath()));
|
||||
log.info("本地文件已删除: {}", doc.getFilePath());
|
||||
} catch (IOException e) {
|
||||
log.warn("删除本地文件失败: {}", doc.getFilePath(), e);
|
||||
@@ -367,6 +382,51 @@ public class DocumentManagementService {
|
||||
// 删除元数据
|
||||
apiDocumentRepository.delete(doc);
|
||||
log.info("文档已删除,docId: {}", docId);
|
||||
|
||||
// 触发域级聚合重算
|
||||
String category = doc.getFilePath() != null
|
||||
? resolveCategory(doc.getFilePath()) : null;
|
||||
if (category != null) {
|
||||
knowledgeDomainService.onDocumentChange(category);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* 转换为响应 DTO
|
||||
*/
|
||||
/**
|
||||
* 从 filePath 解析 category(取 knowledge_base/{category}/... 中的 category 段)
|
||||
*/
|
||||
private String resolveCategory(String filePath) {
|
||||
try {
|
||||
java.nio.file.Path p = java.nio.file.Paths.get(filePath);
|
||||
// filePath 形如 knowledge_base/payment/xxx.md,取倒数第二段
|
||||
int nameCount = p.getNameCount();
|
||||
if (nameCount >= 2) {
|
||||
return p.getName(nameCount - 2).toString();
|
||||
}
|
||||
} catch (Exception ignored) {}
|
||||
return null;
|
||||
}
|
||||
|
||||
private Path resolveLocalPath(String filePath) {
|
||||
Path path = Paths.get(filePath).normalize();
|
||||
if (path.isAbsolute()) {
|
||||
return path;
|
||||
}
|
||||
|
||||
Path basePath = Paths.get(knowledgeBasePath).toAbsolutePath().normalize();
|
||||
Path baseName = basePath.getFileName();
|
||||
if (baseName != null && path.startsWith(baseName) && basePath.getParent() != null) {
|
||||
return basePath.getParent().resolve(path).normalize();
|
||||
}
|
||||
|
||||
Path pathFromWorkingDir = path.toAbsolutePath().normalize();
|
||||
if (pathFromWorkingDir.startsWith(basePath)) {
|
||||
return pathFromWorkingDir;
|
||||
}
|
||||
|
||||
return basePath.resolve(path).normalize();
|
||||
}
|
||||
|
||||
/**
|
||||
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user