feat: add chat verifier agent
This commit is contained in:
@@ -0,0 +1,156 @@
|
||||
---
|
||||
name: openspec-apply-change
|
||||
description: Implement tasks from an OpenSpec change. Use when the user wants to start implementing, continue implementation, or work through tasks.
|
||||
license: MIT
|
||||
compatibility: Requires openspec CLI.
|
||||
metadata:
|
||||
author: openspec
|
||||
version: "1.0"
|
||||
generatedBy: "1.3.1"
|
||||
---
|
||||
|
||||
Implement tasks from an OpenSpec change.
|
||||
|
||||
**Input**: Optionally specify a change name. If omitted, check if it can be inferred from conversation context. If vague or ambiguous you MUST prompt for available changes.
|
||||
|
||||
**Steps**
|
||||
|
||||
1. **Select the change**
|
||||
|
||||
If a name is provided, use it. Otherwise:
|
||||
- Infer from conversation context if the user mentioned a change
|
||||
- Auto-select if only one active change exists
|
||||
- If ambiguous, run `openspec list --json` to get available changes and use the **AskUserQuestion tool** to let the user select
|
||||
|
||||
Always announce: "Using change: <name>" and how to override (e.g., `/opsx:apply <other>`).
|
||||
|
||||
2. **Check status to understand the schema**
|
||||
```bash
|
||||
openspec status --change "<name>" --json
|
||||
```
|
||||
Parse the JSON to understand:
|
||||
- `schemaName`: The workflow being used (e.g., "spec-driven")
|
||||
- Which artifact contains the tasks (typically "tasks" for spec-driven, check status for others)
|
||||
|
||||
3. **Get apply instructions**
|
||||
|
||||
```bash
|
||||
openspec instructions apply --change "<name>" --json
|
||||
```
|
||||
|
||||
This returns:
|
||||
- `contextFiles`: artifact ID -> array of concrete file paths (varies by schema - could be proposal/specs/design/tasks or spec/tests/implementation/docs)
|
||||
- Progress (total, complete, remaining)
|
||||
- Task list with status
|
||||
- Dynamic instruction based on current state
|
||||
|
||||
**Handle states:**
|
||||
- If `state: "blocked"` (missing artifacts): show message, suggest using openspec-continue-change
|
||||
- If `state: "all_done"`: congratulate, suggest archive
|
||||
- Otherwise: proceed to implementation
|
||||
|
||||
4. **Read context files**
|
||||
|
||||
Read every file path listed under `contextFiles` from the apply instructions output.
|
||||
The files depend on the schema being used:
|
||||
- **spec-driven**: proposal, specs, design, tasks
|
||||
- Other schemas: follow the contextFiles from CLI output
|
||||
|
||||
5. **Show current progress**
|
||||
|
||||
Display:
|
||||
- Schema being used
|
||||
- Progress: "N/M tasks complete"
|
||||
- Remaining tasks overview
|
||||
- Dynamic instruction from CLI
|
||||
|
||||
6. **Implement tasks (loop until done or blocked)**
|
||||
|
||||
For each pending task:
|
||||
- Show which task is being worked on
|
||||
- Make the code changes required
|
||||
- Keep changes minimal and focused
|
||||
- Mark task complete in the tasks file: `- [ ]` → `- [x]`
|
||||
- Continue to next task
|
||||
|
||||
**Pause if:**
|
||||
- Task is unclear → ask for clarification
|
||||
- Implementation reveals a design issue → suggest updating artifacts
|
||||
- Error or blocker encountered → report and wait for guidance
|
||||
- User interrupts
|
||||
|
||||
7. **On completion or pause, show status**
|
||||
|
||||
Display:
|
||||
- Tasks completed this session
|
||||
- Overall progress: "N/M tasks complete"
|
||||
- If all done: suggest archive
|
||||
- If paused: explain why and wait for guidance
|
||||
|
||||
**Output During Implementation**
|
||||
|
||||
```
|
||||
## Implementing: <change-name> (schema: <schema-name>)
|
||||
|
||||
Working on task 3/7: <task description>
|
||||
[...implementation happening...]
|
||||
✓ Task complete
|
||||
|
||||
Working on task 4/7: <task description>
|
||||
[...implementation happening...]
|
||||
✓ Task complete
|
||||
```
|
||||
|
||||
**Output On Completion**
|
||||
|
||||
```
|
||||
## Implementation Complete
|
||||
|
||||
**Change:** <change-name>
|
||||
**Schema:** <schema-name>
|
||||
**Progress:** 7/7 tasks complete ✓
|
||||
|
||||
### Completed This Session
|
||||
- [x] Task 1
|
||||
- [x] Task 2
|
||||
...
|
||||
|
||||
All tasks complete! Ready to archive this change.
|
||||
```
|
||||
|
||||
**Output On Pause (Issue Encountered)**
|
||||
|
||||
```
|
||||
## Implementation Paused
|
||||
|
||||
**Change:** <change-name>
|
||||
**Schema:** <schema-name>
|
||||
**Progress:** 4/7 tasks complete
|
||||
|
||||
### Issue Encountered
|
||||
<description of the issue>
|
||||
|
||||
**Options:**
|
||||
1. <option 1>
|
||||
2. <option 2>
|
||||
3. Other approach
|
||||
|
||||
What would you like to do?
|
||||
```
|
||||
|
||||
**Guardrails**
|
||||
- Keep going through tasks until done or blocked
|
||||
- Always read context files before starting (from the apply instructions output)
|
||||
- If task is ambiguous, pause and ask before implementing
|
||||
- If implementation reveals issues, pause and suggest artifact updates
|
||||
- Keep code changes minimal and scoped to each task
|
||||
- Update task checkbox immediately after completing each task
|
||||
- Pause on errors, blockers, or unclear requirements - don't guess
|
||||
- Use contextFiles from CLI output, don't assume specific file names
|
||||
|
||||
**Fluid Workflow Integration**
|
||||
|
||||
This skill supports the "actions on a change" model:
|
||||
|
||||
- **Can be invoked anytime**: Before all artifacts are done (if tasks exist), after partial implementation, interleaved with other actions
|
||||
- **Allows artifact updates**: If implementation reveals design issues, suggest updating artifacts - not phase-locked, work fluidly
|
||||
@@ -0,0 +1,114 @@
|
||||
---
|
||||
name: openspec-archive-change
|
||||
description: Archive a completed change in the experimental workflow. Use when the user wants to finalize and archive a change after implementation is complete.
|
||||
license: MIT
|
||||
compatibility: Requires openspec CLI.
|
||||
metadata:
|
||||
author: openspec
|
||||
version: "1.0"
|
||||
generatedBy: "1.3.1"
|
||||
---
|
||||
|
||||
Archive a completed change in the experimental workflow.
|
||||
|
||||
**Input**: Optionally specify a change name. If omitted, check if it can be inferred from conversation context. If vague or ambiguous you MUST prompt for available changes.
|
||||
|
||||
**Steps**
|
||||
|
||||
1. **If no change name provided, prompt for selection**
|
||||
|
||||
Run `openspec list --json` to get available changes. Use the **AskUserQuestion tool** to let the user select.
|
||||
|
||||
Show only active changes (not already archived).
|
||||
Include the schema used for each change if available.
|
||||
|
||||
**IMPORTANT**: Do NOT guess or auto-select a change. Always let the user choose.
|
||||
|
||||
2. **Check artifact completion status**
|
||||
|
||||
Run `openspec status --change "<name>" --json` to check artifact completion.
|
||||
|
||||
Parse the JSON to understand:
|
||||
- `schemaName`: The workflow being used
|
||||
- `artifacts`: List of artifacts with their status (`done` or other)
|
||||
|
||||
**If any artifacts are not `done`:**
|
||||
- Display warning listing incomplete artifacts
|
||||
- Use **AskUserQuestion tool** to confirm user wants to proceed
|
||||
- Proceed if user confirms
|
||||
|
||||
3. **Check task completion status**
|
||||
|
||||
Read the tasks file (typically `tasks.md`) to check for incomplete tasks.
|
||||
|
||||
Count tasks marked with `- [ ]` (incomplete) vs `- [x]` (complete).
|
||||
|
||||
**If incomplete tasks found:**
|
||||
- Display warning showing count of incomplete tasks
|
||||
- Use **AskUserQuestion tool** to confirm user wants to proceed
|
||||
- Proceed if user confirms
|
||||
|
||||
**If no tasks file exists:** Proceed without task-related warning.
|
||||
|
||||
4. **Assess delta spec sync state**
|
||||
|
||||
Check for delta specs at `openspec/changes/<name>/specs/`. If none exist, proceed without sync prompt.
|
||||
|
||||
**If delta specs exist:**
|
||||
- Compare each delta spec with its corresponding main spec at `openspec/specs/<capability>/spec.md`
|
||||
- Determine what changes would be applied (adds, modifications, removals, renames)
|
||||
- Show a combined summary before prompting
|
||||
|
||||
**Prompt options:**
|
||||
- If changes needed: "Sync now (recommended)", "Archive without syncing"
|
||||
- If already synced: "Archive now", "Sync anyway", "Cancel"
|
||||
|
||||
If user chooses sync, use Task tool (subagent_type: "general-purpose", prompt: "Use Skill tool to invoke openspec-sync-specs for change '<name>'. Delta spec analysis: <include the analyzed delta spec summary>"). Proceed to archive regardless of choice.
|
||||
|
||||
5. **Perform the archive**
|
||||
|
||||
Create the archive directory if it doesn't exist:
|
||||
```bash
|
||||
mkdir -p openspec/changes/archive
|
||||
```
|
||||
|
||||
Generate target name using current date: `YYYY-MM-DD-<change-name>`
|
||||
|
||||
**Check if target already exists:**
|
||||
- If yes: Fail with error, suggest renaming existing archive or using different date
|
||||
- If no: Move the change directory to archive
|
||||
|
||||
```bash
|
||||
mv openspec/changes/<name> openspec/changes/archive/YYYY-MM-DD-<name>
|
||||
```
|
||||
|
||||
6. **Display summary**
|
||||
|
||||
Show archive completion summary including:
|
||||
- Change name
|
||||
- Schema that was used
|
||||
- Archive location
|
||||
- Whether specs were synced (if applicable)
|
||||
- Note about any warnings (incomplete artifacts/tasks)
|
||||
|
||||
**Output On Success**
|
||||
|
||||
```
|
||||
## Archive Complete
|
||||
|
||||
**Change:** <change-name>
|
||||
**Schema:** <schema-name>
|
||||
**Archived to:** openspec/changes/archive/YYYY-MM-DD-<name>/
|
||||
**Specs:** ✓ Synced to main specs (or "No delta specs" or "Sync skipped")
|
||||
|
||||
All artifacts complete. All tasks complete.
|
||||
```
|
||||
|
||||
**Guardrails**
|
||||
- Always prompt for change selection if not provided
|
||||
- Use artifact graph (openspec status --json) for completion checking
|
||||
- Don't block archive on warnings - just inform and confirm
|
||||
- Preserve .openspec.yaml when moving to archive (it moves with the directory)
|
||||
- Show clear summary of what happened
|
||||
- If sync is requested, use openspec-sync-specs approach (agent-driven)
|
||||
- If delta specs exist, always run the sync assessment and show the combined summary before prompting
|
||||
@@ -0,0 +1,288 @@
|
||||
---
|
||||
name: openspec-explore
|
||||
description: Enter explore mode - a thinking partner for exploring ideas, investigating problems, and clarifying requirements. Use when the user wants to think through something before or during a change.
|
||||
license: MIT
|
||||
compatibility: Requires openspec CLI.
|
||||
metadata:
|
||||
author: openspec
|
||||
version: "1.0"
|
||||
generatedBy: "1.3.1"
|
||||
---
|
||||
|
||||
Enter explore mode. Think deeply. Visualize freely. Follow the conversation wherever it goes.
|
||||
|
||||
**IMPORTANT: Explore mode is for thinking, not implementing.** You may read files, search code, and investigate the codebase, but you must NEVER write code or implement features. If the user asks you to implement something, remind them to exit explore mode first and create a change proposal. You MAY create OpenSpec artifacts (proposals, designs, specs) if the user asks—that's capturing thinking, not implementing.
|
||||
|
||||
**This is a stance, not a workflow.** There are no fixed steps, no required sequence, no mandatory outputs. You're a thinking partner helping the user explore.
|
||||
|
||||
---
|
||||
|
||||
## The Stance
|
||||
|
||||
- **Curious, not prescriptive** - Ask questions that emerge naturally, don't follow a script
|
||||
- **Open threads, not interrogations** - Surface multiple interesting directions and let the user follow what resonates. Don't funnel them through a single path of questions.
|
||||
- **Visual** - Use ASCII diagrams liberally when they'd help clarify thinking
|
||||
- **Adaptive** - Follow interesting threads, pivot when new information emerges
|
||||
- **Patient** - Don't rush to conclusions, let the shape of the problem emerge
|
||||
- **Grounded** - Explore the actual codebase when relevant, don't just theorize
|
||||
|
||||
---
|
||||
|
||||
## What You Might Do
|
||||
|
||||
Depending on what the user brings, you might:
|
||||
|
||||
**Explore the problem space**
|
||||
- Ask clarifying questions that emerge from what they said
|
||||
- Challenge assumptions
|
||||
- Reframe the problem
|
||||
- Find analogies
|
||||
|
||||
**Investigate the codebase**
|
||||
- Map existing architecture relevant to the discussion
|
||||
- Find integration points
|
||||
- Identify patterns already in use
|
||||
- Surface hidden complexity
|
||||
|
||||
**Compare options**
|
||||
- Brainstorm multiple approaches
|
||||
- Build comparison tables
|
||||
- Sketch tradeoffs
|
||||
- Recommend a path (if asked)
|
||||
|
||||
**Visualize**
|
||||
```
|
||||
┌─────────────────────────────────────────┐
|
||||
│ Use ASCII diagrams liberally │
|
||||
├─────────────────────────────────────────┤
|
||||
│ │
|
||||
│ ┌────────┐ ┌────────┐ │
|
||||
│ │ State │────────▶│ State │ │
|
||||
│ │ A │ │ B │ │
|
||||
│ └────────┘ └────────┘ │
|
||||
│ │
|
||||
│ System diagrams, state machines, │
|
||||
│ data flows, architecture sketches, │
|
||||
│ dependency graphs, comparison tables │
|
||||
│ │
|
||||
└─────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
**Surface risks and unknowns**
|
||||
- Identify what could go wrong
|
||||
- Find gaps in understanding
|
||||
- Suggest spikes or investigations
|
||||
|
||||
---
|
||||
|
||||
## OpenSpec Awareness
|
||||
|
||||
You have full context of the OpenSpec system. Use it naturally, don't force it.
|
||||
|
||||
### Check for context
|
||||
|
||||
At the start, quickly check what exists:
|
||||
```bash
|
||||
openspec list --json
|
||||
```
|
||||
|
||||
This tells you:
|
||||
- If there are active changes
|
||||
- Their names, schemas, and status
|
||||
- What the user might be working on
|
||||
|
||||
### When no change exists
|
||||
|
||||
Think freely. When insights crystallize, you might offer:
|
||||
|
||||
- "This feels solid enough to start a change. Want me to create a proposal?"
|
||||
- Or keep exploring - no pressure to formalize
|
||||
|
||||
### When a change exists
|
||||
|
||||
If the user mentions a change or you detect one is relevant:
|
||||
|
||||
1. **Read existing artifacts for context**
|
||||
- `openspec/changes/<name>/proposal.md`
|
||||
- `openspec/changes/<name>/design.md`
|
||||
- `openspec/changes/<name>/tasks.md`
|
||||
- etc.
|
||||
|
||||
2. **Reference them naturally in conversation**
|
||||
- "Your design mentions using Redis, but we just realized SQLite fits better..."
|
||||
- "The proposal scopes this to premium users, but we're now thinking everyone..."
|
||||
|
||||
3. **Offer to capture when decisions are made**
|
||||
|
||||
| Insight Type | Where to Capture |
|
||||
|----------------------------|--------------------------------|
|
||||
| New requirement discovered | `specs/<capability>/spec.md` |
|
||||
| Requirement changed | `specs/<capability>/spec.md` |
|
||||
| Design decision made | `design.md` |
|
||||
| Scope changed | `proposal.md` |
|
||||
| New work identified | `tasks.md` |
|
||||
| Assumption invalidated | Relevant artifact |
|
||||
|
||||
Example offers:
|
||||
- "That's a design decision. Capture it in design.md?"
|
||||
- "This is a new requirement. Add it to specs?"
|
||||
- "This changes scope. Update the proposal?"
|
||||
|
||||
4. **The user decides** - Offer and move on. Don't pressure. Don't auto-capture.
|
||||
|
||||
---
|
||||
|
||||
## What You Don't Have To Do
|
||||
|
||||
- Follow a script
|
||||
- Ask the same questions every time
|
||||
- Produce a specific artifact
|
||||
- Reach a conclusion
|
||||
- Stay on topic if a tangent is valuable
|
||||
- Be brief (this is thinking time)
|
||||
|
||||
---
|
||||
|
||||
## Handling Different Entry Points
|
||||
|
||||
**User brings a vague idea:**
|
||||
```
|
||||
User: I'm thinking about adding real-time collaboration
|
||||
|
||||
You: Real-time collab is a big space. Let me think about this...
|
||||
|
||||
COLLABORATION SPECTRUM
|
||||
════════════════════════════════════════════
|
||||
|
||||
Awareness Coordination Sync
|
||||
│ │ │
|
||||
▼ ▼ ▼
|
||||
┌────────┐ ┌────────┐ ┌────────┐
|
||||
│Presence│ │Cursors │ │ CRDT │
|
||||
│ "3 │ │ Multi │ │Conflict│
|
||||
│online" │ │ select │ │ free │
|
||||
└────────┘ └────────┘ └────────┘
|
||||
│ │ │
|
||||
trivial moderate complex
|
||||
|
||||
Where's your head at?
|
||||
```
|
||||
|
||||
**User brings a specific problem:**
|
||||
```
|
||||
User: The auth system is a mess
|
||||
|
||||
You: [reads codebase]
|
||||
|
||||
┌─────────────────────────────────────────────┐
|
||||
│ CURRENT AUTH FLOW │
|
||||
└─────────────────────────────────────────────┘
|
||||
│
|
||||
┌─────────────┼─────────────┐
|
||||
▼ ▼ ▼
|
||||
┌─────────┐ ┌─────────┐ ┌─────────┐
|
||||
│ Google │ │ GitHub │ │ Email │
|
||||
│ OAuth │ │ OAuth │ │ Magic │
|
||||
└────┬────┘ └────┬────┘ └────┬────┘
|
||||
│ │ │
|
||||
└─────────────┼─────────────┘
|
||||
▼
|
||||
┌───────────┐
|
||||
│ Session │
|
||||
└─────┬─────┘
|
||||
│
|
||||
▼
|
||||
┌───────────┐
|
||||
│ Perms │
|
||||
└───────────┘
|
||||
|
||||
I see three tangles. Which one's burning?
|
||||
```
|
||||
|
||||
**User is stuck mid-implementation:**
|
||||
```
|
||||
User: /opsx:explore add-auth-system
|
||||
The OAuth integration is more complex than expected
|
||||
|
||||
You: [reads change artifacts]
|
||||
|
||||
You're on task 4: "Implement OAuth flow"
|
||||
|
||||
Let me trace what's involved...
|
||||
|
||||
[draws diagram, explores options, suggests paths]
|
||||
|
||||
Want to update the design to reflect this?
|
||||
Or add a spike task to investigate?
|
||||
```
|
||||
|
||||
**User wants to compare options:**
|
||||
```
|
||||
User: Should we use Postgres or SQLite?
|
||||
|
||||
You: Generic answer is boring. What's the context?
|
||||
|
||||
User: A CLI tool that tracks local dev environments
|
||||
|
||||
You: That changes everything.
|
||||
|
||||
┌─────────────────────────────────────────────────┐
|
||||
│ CLI TOOL DATA STORAGE │
|
||||
└─────────────────────────────────────────────────┘
|
||||
|
||||
Key constraints:
|
||||
• No daemon running
|
||||
• Must work offline
|
||||
• Single user
|
||||
|
||||
SQLite Postgres
|
||||
Deployment embedded ✓ needs server ✗
|
||||
Offline yes ✓ no ✗
|
||||
Single file yes ✓ no ✗
|
||||
|
||||
SQLite. Not even close.
|
||||
|
||||
Unless... is there a sync component?
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Ending Discovery
|
||||
|
||||
There's no required ending. Discovery might:
|
||||
|
||||
- **Flow into a proposal**: "Ready to start? I can create a change proposal."
|
||||
- **Result in artifact updates**: "Updated design.md with these decisions"
|
||||
- **Just provide clarity**: User has what they need, moves on
|
||||
- **Continue later**: "We can pick this up anytime"
|
||||
|
||||
When it feels like things are crystallizing, you might summarize:
|
||||
|
||||
```
|
||||
## What We Figured Out
|
||||
|
||||
**The problem**: [crystallized understanding]
|
||||
|
||||
**The approach**: [if one emerged]
|
||||
|
||||
**Open questions**: [if any remain]
|
||||
|
||||
**Next steps** (if ready):
|
||||
- Create a change proposal
|
||||
- Keep exploring: just keep talking
|
||||
```
|
||||
|
||||
But this summary is optional. Sometimes the thinking IS the value.
|
||||
|
||||
---
|
||||
|
||||
## Guardrails
|
||||
|
||||
- **Don't implement** - Never write code or implement features. Creating OpenSpec artifacts is fine, writing application code is not.
|
||||
- **Don't fake understanding** - If something is unclear, dig deeper
|
||||
- **Don't rush** - Discovery is thinking time, not task time
|
||||
- **Don't force structure** - Let patterns emerge naturally
|
||||
- **Don't auto-capture** - Offer to save insights, don't just do it
|
||||
- **Do visualize** - A good diagram is worth many paragraphs
|
||||
- **Do explore the codebase** - Ground discussions in reality
|
||||
- **Do question assumptions** - Including the user's and your own
|
||||
@@ -0,0 +1,110 @@
|
||||
---
|
||||
name: openspec-propose
|
||||
description: Propose a new change with all artifacts generated in one step. Use when the user wants to quickly describe what they want to build and get a complete proposal with design, specs, and tasks ready for implementation.
|
||||
license: MIT
|
||||
compatibility: Requires openspec CLI.
|
||||
metadata:
|
||||
author: openspec
|
||||
version: "1.0"
|
||||
generatedBy: "1.3.1"
|
||||
---
|
||||
|
||||
Propose a new change - create the change and generate all artifacts in one step.
|
||||
|
||||
I'll create a change with artifacts:
|
||||
- proposal.md (what & why)
|
||||
- design.md (how)
|
||||
- tasks.md (implementation steps)
|
||||
|
||||
When ready to implement, run /opsx:apply
|
||||
|
||||
---
|
||||
|
||||
**Input**: The user's request should include a change name (kebab-case) OR a description of what they want to build.
|
||||
|
||||
**Steps**
|
||||
|
||||
1. **If no clear input provided, ask what they want to build**
|
||||
|
||||
Use the **AskUserQuestion tool** (open-ended, no preset options) to ask:
|
||||
> "What change do you want to work on? Describe what you want to build or fix."
|
||||
|
||||
From their description, derive a kebab-case name (e.g., "add user authentication" → `add-user-auth`).
|
||||
|
||||
**IMPORTANT**: Do NOT proceed without understanding what the user wants to build.
|
||||
|
||||
2. **Create the change directory**
|
||||
```bash
|
||||
openspec new change "<name>"
|
||||
```
|
||||
This creates a scaffolded change at `openspec/changes/<name>/` with `.openspec.yaml`.
|
||||
|
||||
3. **Get the artifact build order**
|
||||
```bash
|
||||
openspec status --change "<name>" --json
|
||||
```
|
||||
Parse the JSON to get:
|
||||
- `applyRequires`: array of artifact IDs needed before implementation (e.g., `["tasks"]`)
|
||||
- `artifacts`: list of all artifacts with their status and dependencies
|
||||
|
||||
4. **Create artifacts in sequence until apply-ready**
|
||||
|
||||
Use the **TodoWrite tool** to track progress through the artifacts.
|
||||
|
||||
Loop through artifacts in dependency order (artifacts with no pending dependencies first):
|
||||
|
||||
a. **For each artifact that is `ready` (dependencies satisfied)**:
|
||||
- Get instructions:
|
||||
```bash
|
||||
openspec instructions <artifact-id> --change "<name>" --json
|
||||
```
|
||||
- The instructions JSON includes:
|
||||
- `context`: Project background (constraints for you - do NOT include in output)
|
||||
- `rules`: Artifact-specific rules (constraints for you - do NOT include in output)
|
||||
- `template`: The structure to use for your output file
|
||||
- `instruction`: Schema-specific guidance for this artifact type
|
||||
- `outputPath`: Where to write the artifact
|
||||
- `dependencies`: Completed artifacts to read for context
|
||||
- Read any completed dependency files for context
|
||||
- Create the artifact file using `template` as the structure
|
||||
- Apply `context` and `rules` as constraints - but do NOT copy them into the file
|
||||
- Show brief progress: "Created <artifact-id>"
|
||||
|
||||
b. **Continue until all `applyRequires` artifacts are complete**
|
||||
- After creating each artifact, re-run `openspec status --change "<name>" --json`
|
||||
- Check if every artifact ID in `applyRequires` has `status: "done"` in the artifacts array
|
||||
- Stop when all `applyRequires` artifacts are done
|
||||
|
||||
c. **If an artifact requires user input** (unclear context):
|
||||
- Use **AskUserQuestion tool** to clarify
|
||||
- Then continue with creation
|
||||
|
||||
5. **Show final status**
|
||||
```bash
|
||||
openspec status --change "<name>"
|
||||
```
|
||||
|
||||
**Output**
|
||||
|
||||
After completing all artifacts, summarize:
|
||||
- Change name and location
|
||||
- List of artifacts created with brief descriptions
|
||||
- What's ready: "All artifacts created! Ready for implementation."
|
||||
- Prompt: "Run `/opsx:apply` or ask me to implement to start working on the tasks."
|
||||
|
||||
**Artifact Creation Guidelines**
|
||||
|
||||
- Follow the `instruction` field from `openspec instructions` for each artifact type
|
||||
- The schema defines what each artifact should contain - follow it
|
||||
- Read dependency artifacts for context before creating new ones
|
||||
- Use `template` as the structure for your output file - fill in its sections
|
||||
- **IMPORTANT**: `context` and `rules` are constraints for YOU, not content for the file
|
||||
- Do NOT copy `<context>`, `<rules>`, `<project_context>` blocks into the artifact
|
||||
- These guide what you write, but should never appear in the output
|
||||
|
||||
**Guardrails**
|
||||
- Create ALL artifacts needed for implementation (as defined by schema's `apply.requires`)
|
||||
- Always read dependency artifacts before creating a new one
|
||||
- If context is critically unclear, ask the user - but prefer making reasonable decisions to keep momentum
|
||||
- If a change with that name already exists, ask if user wants to continue it or create a new one
|
||||
- Verify each artifact file exists after writing before proceeding to next
|
||||
@@ -12,3 +12,4 @@
|
||||
| 2026-06-29 | confidence-feedback | 质量评估/反馈机制 | evidence_score, selfEvaluation, feedback, useful, not_useful, case_library, BAD_CASE, tool_invocation规则引擎, 反馈按钮, sessionId回传 | openspec/changes/confidence-feedback | archived |
|
||||
| 2026-06-30 | session-dedup-knowledge-map | 去重/知识图谱 | RetrievedDocTracker, KnowledgeDomainService, knowledge_domain, covers, whenToRetrieve, Planner注入, ISS-001 | openspec/changes/archive/2026-06-30-session-dedup-knowledge-map | archived |
|
||||
| 2026-07-01 | executor-action-memory-relevance | 检索质量/行动记忆 | relevanceLevel, completenessHint, Min-Max归一化, RetrievedDocTracker域级记录, Executor检索约束, ISS-002 | openspec/changes/archive/2026-07-01-executor-action-memory-relevance | archived |
|
||||
| 2026-07-02 | chat-verifier-agent | Chat质量门禁/可追溯验证 | Verifier, groundedness_score, facts_checked, evidence_refs, tool_trace_summary, self_evaluation | openspec/changes/archive/2026-07-03-chat-verifier-agent | archived |
|
||||
|
||||
@@ -0,0 +1,58 @@
|
||||
# Acceptance: chat-verifier-agent
|
||||
|
||||
## Classification
|
||||
|
||||
standard
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Verifier prompt | Done | Strict JSON schema, verdict matrix, fact classifications, and `evidence_refs` are defined. |
|
||||
| VerifierInputHook | Done | Explicit verifier payload replaces raw conversation history. |
|
||||
| ChatService integration | Done | Planner, executor, and verifier are called explicitly with max two rounds. |
|
||||
| Verdict routing | Done | PASS, LOW_CONFID, and REJECT paths are handled in code. |
|
||||
| Trace summary | Done | Evidence summaries include `trace_ref` and `source_invocation_ids`. |
|
||||
| self_evaluation merge | Done | `rule_evaluation` and `verifier_evaluation` are preserved independently. |
|
||||
| Verifier observability | Done | `verifier_evaluation` persists facts, evidence refs, trace summary, rationale, score, and round. |
|
||||
|
||||
## Static Verification
|
||||
|
||||
- [x] OpenSpec artifacts exist: `proposal.md`, `design.md`, `specs/chat-verifier-agent/spec.md`, `tasks.md`, `.committed`.
|
||||
- [x] `change.json` exists and has `metadata.status = committed`.
|
||||
- [x] `.archive-ready` exists.
|
||||
- [x] devflow archive-prep files exist: `brief.md`, `evidence.md`, `decisions.md`, `acceptance.md`.
|
||||
- [x] `devflow/index.md` contains `chat-verifier-agent` with status `archived`.
|
||||
|
||||
## Script Verification
|
||||
|
||||
- [x] `mvn -q -DskipTests compile` passed.
|
||||
|
||||
## Runtime Verification
|
||||
|
||||
- [x] POST `/api/chat` with a complex question returned successfully.
|
||||
- [x] Runtime session `9138f064` showed planner, executor, and verifier execution in logs.
|
||||
- [x] Runtime session `9138f064` wrote `verifier_evaluation.verdict = LOW_CONFID`.
|
||||
- [x] Runtime session `9138f064` wrote `facts_checked[*].evidence_refs`.
|
||||
- [x] Runtime session `9138f064` wrote `tool_trace_summary[*].source_invocation_ids`.
|
||||
- [x] LOW_CONFID final answer included disclaimer and verifier-derived evidence gaps.
|
||||
|
||||
## Unverified
|
||||
|
||||
| Scenario | Reason | Risk | Follow-up |
|
||||
| --- | --- | --- | --- |
|
||||
| PASS runtime path | The exercised complex runtime case produced LOW_CONFID. | Low; PASS routing is simple pass-through after parsed verifier decision. | Add a fixture or deterministic verifier test if this becomes product-critical. |
|
||||
| REJECT runtime path | No forced contradiction case was run after traceability changes. | Medium; REJECT is the safety-critical degraded path. | Add a targeted test with a fabricated claim and evidence contradiction. |
|
||||
| Document-path-level evidence mapping | Current implementation records invocation ids and source document labels, not guaranteed canonical document paths for every retrieval mode. | Low for current audit need; medium for future UI drill-down. | Extend retrieval details with canonical document paths in a later change. |
|
||||
|
||||
## Remaining Risks
|
||||
|
||||
1. Verifier output still depends on model compliance with JSON schema; code falls back to LOW_CONFID on missing or invalid output.
|
||||
2. `AgentLoggingHook` is shared by several agent paths; current changes preserve compile and runtime behavior but should be watched in AiOps flows.
|
||||
3. `SupervisorAgent` construction remains as legacy residue in `ChatService`; runtime orchestration is explicit, but a later cleanup should remove unused supervisor construction.
|
||||
|
||||
## Archive State
|
||||
|
||||
- [x] OpenSpec change is archive-ready.
|
||||
- [x] OpenSpec change has been moved to `openspec/changes/archive/2026-07-03-chat-verifier-agent/`.
|
||||
- [x] Main spec exists at `openspec/specs/chat-verifier-agent/spec.md`.
|
||||
@@ -0,0 +1,36 @@
|
||||
# Brief: chat-verifier-agent
|
||||
|
||||
## Background
|
||||
|
||||
The complex Chat path previously returned Executor answers without a synchronous quality gate. Existing rule scoring was asynchronous and post-hoc, so it could not prevent unsupported answers from reaching users.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Add a Verifier Agent after Executor in the complex chat path.
|
||||
2. Require structured verifier output with `PASS`, `LOW_CONFID`, or `REJECT`.
|
||||
3. Route final user output in code based on verifier verdict.
|
||||
4. Persist verifier results under `diagnosis_session.self_evaluation.verifier_evaluation`.
|
||||
5. Preserve rule scoring under `rule_evaluation`.
|
||||
6. Make verifier decisions traceable to real tool invocations through `evidence_refs` and `source_invocation_ids`.
|
||||
|
||||
## Scope
|
||||
|
||||
- `ChatService`: explicit `planner -> executor -> verifier` orchestration, max two rounds, verdict routing, retry context, verifier persistence.
|
||||
- `VerifierInputHook`: explicit verifier input payload.
|
||||
- `ToolTraceSummaryService`: evidence summary from persisted tool calls.
|
||||
- `VerifierContextHolder`: round-local verifier context.
|
||||
- `SelfEvaluationMergeService`: safe JSON merge for evaluation channels.
|
||||
- `AgentLoggingHook`: concise verifier thought and fuller structured output retention.
|
||||
- `chat-verifier-prompt.md`: verifier contract, verdict matrix, and traceability schema.
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- Verifier does not call tools.
|
||||
- Verifier does not rewrite Executor output.
|
||||
- Single-agent chat path remains outside this change.
|
||||
- No database schema migration is included.
|
||||
- Document-path-level evidence attribution is deferred; current traceability is invocation-level with source document labels.
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/archive/2026-07-03-chat-verifier-agent/`
|
||||
@@ -0,0 +1,123 @@
|
||||
# Decisions: chat-verifier-agent
|
||||
|
||||
## 过程日志
|
||||
|
||||
### Clarify 阶段
|
||||
|
||||
**入口摘要**: 在 Chat 多 Agent 链路中新增 Verifier Agent,作为 Executor 输出后的质量门禁,做事实核查。
|
||||
|
||||
**slug**: `chat-verifier-agent`
|
||||
|
||||
**规模分档**: standard
|
||||
|
||||
### Context 阶段
|
||||
|
||||
**devflow/index.md 使用状态**: 已命中。前序 change `executor-action-memory-relevance`(archived)提供了 Chat 多 Agent 当前链路(Supervisor → Planner → Executor)。
|
||||
|
||||
**不能违反的历史决策**:
|
||||
1. Executor 已有完整的行动记忆和归一化质量等级,Verifier 不需要重复验证检索质量
|
||||
2. Chat Supervisor 的职责是调度,Verifier 作为子 Agent 加入后不改变 Supervisor 的定位
|
||||
3. 已有 evidence_score 做事后评分,Verifier 是事前门禁,两者不冲突
|
||||
|
||||
**需进入 OpenSpec 的上下文点**:
|
||||
1. Verifier 不需要工具调用,只是一个质量核查 Agent
|
||||
2. Verifier 需要访问 Executor 的输出 + 工具调用记录
|
||||
3. Supervisor prompt 需要重写以包含 Verifier 调度规则
|
||||
4. groundedness_score 的阈值需要在代码中定义
|
||||
|
||||
### Grill 阶段 — Question Pool
|
||||
|
||||
| # | 维度 | 问题 | 模式 | 状态 |
|
||||
|---|------|------|------|------|
|
||||
| Q1 | 术语 | evidence_score(事后评分)与 Verifier(事前门禁)职责是否冲突? | evidence-driven | 已解决 |
|
||||
| Q2 | 边界 | Verifier 需要的"工具调用记录"在 SupervisorAgent 中是否自动传递? | evidence-driven | 已解决 |
|
||||
| Q3 | 边界 | LOW_CONFID < 0.5 回调 Planner 后的新输出是否再次走 Verifier?循环上限多少? | user-interview | 已解决 |
|
||||
| Q4 | 验收 | Verifier 判决结果如何可观测?是否写入 agent_step 或 tool_invocation? | user-interview | 已解决 |
|
||||
| Q5 | 验收 | 当前 Supervisor 硬编码 prompt 是否支持多 Agent 路由变更? | evidence-driven | 已解决 |
|
||||
| Q6 | 技术 | Verifier 如何隔离 Executor 的中间推理过程,只看到干净的 query + tool 记录 + 最终答案? | user-interview | 已解决 |
|
||||
| Q7 | 验收 | groundedness_score 阈值(0.5)是否需要配置化? | user-interview | 已解决 |
|
||||
|
||||
### Evidence-driven 结论
|
||||
|
||||
| 结论 | 证据来源 | 是否已汇报用户 |
|
||||
|------|---------|-------------|
|
||||
| evidence_score(异步事后)与 Verifier(同步事前门禁)不冲突 | EvaluationService.java: @Async 注解 | 已汇报 |
|
||||
| SupervisorAgent 自动传递完整对话状态,Verifier 无需额外传递工具记录 | Spring AI Alibaba SupervisorAgent 实现 | 已汇报 |
|
||||
| Supervisor prompt 为字符串字面量,直接修改即可 | ChatService.java:353 .systemPrompt("...") | 已汇报 |
|
||||
|
||||
### User-interview 记录
|
||||
|
||||
| 问题 | 用户原话 | 确认状态 | OpenSpec 回写 |
|
||||
|------|---------|---------|-------------|
|
||||
| Q3: LOW_CONFID < 0.5 回调 Planner 循环上限? | "可以,回调一次" | 已确认 | 已回写 proposal |
|
||||
| Q4: Verifier 判决写入哪里做可观测? | "可以"(写入 diagnosis_session.self_evaluation JSON) | 已确认 | 已回写 proposal |
|
||||
| Q6: Verifier 如何隔离 Executor 中间推理? | "用 MessagesModelHook 过滤 messages" | 已确认 | 已回写 design |
|
||||
| Q7: groundedness_score 阈值是否需要配置化? | "需要配置化" | 已确认 | 已回写 design |
|
||||
|
||||
### Specify 阶段 — Cross-Artifact 对齐检查
|
||||
|
||||
| 上游 → 下游 | 检查内容 | 状态 |
|
||||
|---|---|---|
|
||||
| proposal → design | 范围、约束、关键承诺是否进入 design | 已对齐 |
|
||||
| design → specs | 关键决策、模块地图是否进入 specs | 已对齐 |
|
||||
| specs → tasks | 可观察行为是否被 tasks 覆盖为可执行切片 | 已对齐 |
|
||||
|
||||
**接口影响分级**:
|
||||
- buildChatVerifierAgent() 新增方法 → L1(内部方法,无外部消费者)
|
||||
- VerifierInputHook 类 → L1(内部 Hook,无外部消费者)
|
||||
- Supervisor prompt 重写 → L1(仅影响 Chat 多 Agent 内部调度)
|
||||
- subAgents 列表变更 → L1(Supervisor 内部配置)
|
||||
- verifier.low-confidence-threshold 配置 → L1(新增配置项,不改已有配置)
|
||||
|
||||
### Audit 阶段
|
||||
|
||||
**模块链路**:
|
||||
|
||||
```
|
||||
用户 → Supervisor → Planner(步骤) → Executor(答案+工具记录)
|
||||
│
|
||||
Supervisor 调用 Verifier
|
||||
│
|
||||
[VerifierInputHook BEFORE_MODEL]
|
||||
├─ 保留:system prompt + user query
|
||||
├─ 保留:tool call 记录(输入+返回)
|
||||
├─ 保留:Executor 最终答案
|
||||
└─ 去除:Executor 中间推理、Planner 规划过程
|
||||
│
|
||||
Verifier 判决
|
||||
│
|
||||
┌─── PASS ───→ 直接输出
|
||||
├─── LOW_CONFID≥0.5 → 带声明输出
|
||||
├─── LOW_CONFID<0.5 → 回调 Planner(一次)
|
||||
└─── REJECT → 降级输出
|
||||
│
|
||||
写入 self_evaluation JSON
|
||||
```
|
||||
|
||||
**架构风险评估**(5 句以内):
|
||||
1. Verifier 是轻量 Agent(无工具、无外部依赖),架构风险低。
|
||||
2. MessagesModelHook 纯过滤逻辑,不引入新数据源。
|
||||
3. LOW_CONFID 分级处理 + 回调仅一次的设计,避免无限循环风险。
|
||||
4. REJECT 降级确保编造内容不到达用户。
|
||||
5. 审计结论不影响现有 design/tasks,无需回写。
|
||||
|
||||
### 关键取舍
|
||||
|
||||
- 决策:LOW_CONFID < 0.5 回调 Planner 一次
|
||||
- 原因:给系统一次修正机会,但避免无限循环
|
||||
- 影响:Supervisor prompt 需维护"已回调"状态
|
||||
- 风险接受:用户已确认
|
||||
|
||||
- 决策:Verifier 判决写入 diagnosis_session.self_evaluation JSON
|
||||
- 原因:不改表结构,与 evidence_score 统一可观测体系
|
||||
- 影响:ChatService 后处理需追加 JSON
|
||||
- 风险接受:用户已确认
|
||||
|
||||
### Archive-Ready Update
|
||||
|
||||
- 实现调整:最终运行链路由 `ChatService` 显式调用 `planner -> executor -> verifier`,不再依赖 Supervisor prompt 保证 verifier 被调用。
|
||||
- 可追溯性补充:`tool_trace_summary` 增加 `trace_ref`、`source_invocation_ids`、查询样本、检索层级、相关性等级和来源文档标签。
|
||||
- 可追溯性补充:`facts_checked[*].evidence_refs` 被 prompt 要求、代码解析并持久化。
|
||||
- 验证记录:`mvn -q -DskipTests compile` 通过。
|
||||
- 验证记录:运行会话 `9138f064` 走通 planner、executor、verifier,并持久化 `verifier_evaluation.facts_checked[*].evidence_refs` 与 `tool_trace_summary[*].source_invocation_ids`。
|
||||
- 当前状态:OpenSpec change 已归档到 `openspec/changes/archive/2026-07-03-chat-verifier-agent/`,主规格已同步到 `openspec/specs/chat-verifier-agent/spec.md`。
|
||||
@@ -0,0 +1,52 @@
|
||||
# Evidence: chat-verifier-agent
|
||||
|
||||
## Code Evidence
|
||||
|
||||
### Complex chat path now invokes verifier deterministically
|
||||
|
||||
- File: `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- Evidence: `executeChatComplex` calls planner, executor, then verifier directly through `callAgent(...)`.
|
||||
- Conclusion: runtime no longer depends on prompt-only Supervisor behavior to call verifier.
|
||||
|
||||
### Verifier receives explicit inputs
|
||||
|
||||
- File: `src/main/java/com/superbiz/agent/hook/VerifierInputHook.java`
|
||||
- Evidence: the hook builds a JSON payload with `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
|
||||
- Conclusion: verifier input is stable and does not depend on guessing the last assistant message from raw history.
|
||||
|
||||
### Tool evidence is traceable to persisted invocations
|
||||
|
||||
- File: `src/main/java/com/superbiz/agent/service/ToolTraceSummaryService.java`
|
||||
- Evidence: summaries include `trace_ref`, `source_invocation_ids`, `query_samples`, `retrieval_layers`, `relevance_levels`, and `source_documents`.
|
||||
- Conclusion: verifier facts can be correlated with actual `tool_invocation` rows.
|
||||
|
||||
### Verifier facts preserve evidence references
|
||||
|
||||
- File: `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- Evidence: verifier parsing preserves `facts_checked[*].evidence_refs` and persists `tool_trace_summary` under `verifier_evaluation`.
|
||||
- Conclusion: `self_evaluation` now contains both verifier judgments and the evidence index used to form them.
|
||||
|
||||
### Evaluation channels no longer overwrite each other
|
||||
|
||||
- File: `src/main/java/com/superbiz/agent/service/SelfEvaluationMergeService.java`
|
||||
- Evidence: rule and verifier evaluations are merged into separate keys.
|
||||
- Conclusion: asynchronous rule scoring preserves verifier output.
|
||||
|
||||
### Verifier logging is less noisy
|
||||
|
||||
- File: `src/main/java/com/superbiz/agent/hook/AgentLoggingHook.java`
|
||||
- Evidence: verifier `thought` stores a concise verdict summary, while fuller model output remains available in structured storage.
|
||||
- Conclusion: `agent_step.thought` is no longer a misleading place for full verifier JSON.
|
||||
|
||||
## Runtime Evidence
|
||||
|
||||
- Compile verification passed: `mvn -q -DskipTests compile`.
|
||||
- Runtime session `9138f064` executed `planner -> executor -> verifier`.
|
||||
- Runtime session `9138f064` persisted `verifier_evaluation.facts_checked[*].evidence_refs`.
|
||||
- Runtime session `9138f064` persisted `verifier_evaluation.tool_trace_summary[*].source_invocation_ids`.
|
||||
|
||||
## Design Evidence
|
||||
|
||||
- `LOW_CONFID` returns a fixed disclaimer and verifier-derived gaps.
|
||||
- `REJECT` returns degraded output and does not pass through the raw Executor answer.
|
||||
- `retry_context` is derived from verifier-identified missing evidence facts.
|
||||
@@ -0,0 +1,146 @@
|
||||
# ISS-003 Executor 域级检索水位控制(Phase 2)
|
||||
|
||||
**状态**:待规划
|
||||
**严重程度**:低(当前软约束已从 20+ 次收敛到 10 次,文档级去重兜住无效调用)
|
||||
**发现时间**:2026-07-01
|
||||
**关联**:ISS-002(Part A 软约束已修,LLM 遵守度不够)
|
||||
|
||||
---
|
||||
|
||||
## 现象
|
||||
|
||||
ISS-002 修复后,lookup_knowledge 调用从 20+ 次降到 10 次,但仍有 8 次冗余调用:
|
||||
|
||||
```
|
||||
id=138 → infrastructure 首次检索 HIGHLY_RELEVANT
|
||||
id=139 → infrastructure 变体 → doc_retrieved
|
||||
id=140 → infrastructure 变体 → doc_retrieved
|
||||
id=141 → api 首次检索 REFERENCE
|
||||
id=142 → api 变体 → doc_retrieved
|
||||
id=143 → infrastructure 变体 → doc_retrieved(又跳回)
|
||||
id=144-146 → infrastructure 变体 × 3 → doc_retrieved
|
||||
id=147 → api 变体 → doc_retrieved
|
||||
```
|
||||
|
||||
LLM 在两个域之间反复横跳,prompt 第 2 条"禁止换关键词重新检索"被忽略。
|
||||
|
||||
---
|
||||
|
||||
## 根本原因
|
||||
|
||||
Prompt 软约束依赖 LLM 遵守,但 LLM 在自主决策(ReactAgent)模式下倾向于"多确认一步"而不是"相信已有信息"。
|
||||
|
||||
---
|
||||
|
||||
## 影响
|
||||
|
||||
- **可接受**:文档级去重已拦截重复内容,不影响回答质量
|
||||
- **可优化**:每次冗余调用浪费 400-500ms 服务端检索时间
|
||||
- **长会话风险**:如果 session 持续追问,冗余调用会线性增长
|
||||
|
||||
---
|
||||
|
||||
## 方案:质量等级×水位决策矩阵
|
||||
|
||||
### 核心思路
|
||||
|
||||
将检索决策权从 LLM 思维链移交到代码层,动态判断是否允许下一次 `lookup_knowledge`。
|
||||
|
||||
### 决策矩阵
|
||||
|
||||
```
|
||||
水位
|
||||
低 中 高
|
||||
PRECISE 可用 可用 可用/停
|
||||
HIGHLY_ 可用 可用 停
|
||||
REFERENCE 可定向补 可定向补 停
|
||||
DEDUPED 停 停 停
|
||||
```
|
||||
|
||||
### 水位指标
|
||||
|
||||
| 水位 | 指标 | 说明 |
|
||||
|------|------|------|
|
||||
| 低 | lookup_knowledge 调用 ≤ 3 次 | 检索预算充足 |
|
||||
| 中 | lookup_knowledge 调用 4-8 次 | 适当收紧定向补充 |
|
||||
| 高 | lookup_knowledge 调用 ≥ 8 次 | 熔断,禁止再查 |
|
||||
|
||||
备选水位指标:
|
||||
- Token 消耗量
|
||||
- 已检索域数(retrievedDomainsThisSession.size())
|
||||
|
||||
### 熔断 prompt 示例
|
||||
|
||||
每次工具返回后,代码根据矩阵结果动态拼装约束注入下一步 LLM:
|
||||
|
||||
| 场景 | 熔断 prompt |
|
||||
|------|------------|
|
||||
| 水位高 + HIGHLY_RELEVANT | "水位已高,请直接给出最终结论,不要再调用 lookup_knowledge" |
|
||||
| 水位高 + REFERENCE | "水位已高,lookup_knowledge 已被限制,直接基于已有信息回答" |
|
||||
| 水位低 + REFERENCE | "水位充足,可针对缺少的维度定向补充检索一次" |
|
||||
| DEDUPED + 任何水位 | "该内容已检索过,禁止重复调用 lookup_knowledge" |
|
||||
|
||||
### 架构改动
|
||||
|
||||
```
|
||||
ChatService / ChatExecutorAgent 执行循环
|
||||
│
|
||||
├─ step N: LLM 调用工具 → lookup_knowledge 返回
|
||||
├─ step N+1:
|
||||
│ ├─ 读取 LookupResult(relevanceLevel + retrievedDomainsThisSession)
|
||||
│ ├─ 算当前水位(调用计数 / token / 域数)
|
||||
│ ├─ 查决策矩阵 → 是否允许继续检索
|
||||
│ └─ 拼装熔断 prompt → 注入 LLM 下一步 SystemMessage
|
||||
├─ step N+2: LLM 收到约束后的回复
|
||||
└─ ...
|
||||
```
|
||||
|
||||
### 与现有机制的关系
|
||||
|
||||
| 机制 | 层 | ISS-002 | ISS-003 |
|
||||
|------|-----|---------|---------|
|
||||
| 静态 Prompt 约束 | Prompt | 已实现 | 保留作为基线 |
|
||||
| 行动记忆(retrievedDomainsThisSession) | 工具返回值 | 已实现 | 复用 |
|
||||
| 归一化质量等级(relevanceLevel) | 工具返回值 | 已实现 | 复用 |
|
||||
| 文档级去重(doc_retrieved) | 工具层 | 已实现 | 保留 |
|
||||
| **域级水位决策矩阵** | 代码层 | — | **新增** |
|
||||
| DEDUPED 等级启用 | 工具返回值 | 设计预留 | 启用 |
|
||||
|
||||
---
|
||||
|
||||
## 修法方向
|
||||
|
||||
### 方案 A:决策矩阵(推荐)
|
||||
|
||||
上述质量等级×水位矩阵,在 ChatService 执行循环中做决策。
|
||||
|
||||
优点:
|
||||
- 不依赖 LLM 遵守程度
|
||||
- 保留定向补充的合法通道(比硬拦截更灵活)
|
||||
- 水位指标可配置,运维友好
|
||||
|
||||
缺点:
|
||||
- 需要改动 ChatService/ExecutorAgent 执行逻辑
|
||||
- 需要定义水位阈值(需实测校准)
|
||||
|
||||
### 方案 B:域级工具层硬限流
|
||||
|
||||
在 `LookupKnowledgeTool` 入口直接判断 `isDomainRetrieved(sessionId, domain)`,同域直接返回不检索。
|
||||
|
||||
优点:实现简单,100% 可靠
|
||||
缺点:没有"定向补充"的灵活度,误杀合法跨域查询
|
||||
|
||||
### 建议
|
||||
|
||||
**方案 A**。利用 ISS-002 已有数据结构和归一化等级,新增决策引擎层,改动量不大且灵活度高。
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `src/main/java/com/superbiz/agent/dto/LookupResult.java`
|
||||
- `src/main/java/com/superbiz/agent/tool/RetrievedDocTracker.java`
|
||||
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
|
||||
- `src/main/resources/prompts/chat-executor-prompt.md`
|
||||
- 架构设计:`mvp/architecture/action-memory-relevance.md`
|
||||
- OpenSpec:`openspec/changes/archive/2026-07-01-executor-action-memory-relevance/`
|
||||
@@ -4,3 +4,4 @@
|
||||
|---|---|---|---|---|
|
||||
| ISS-001 | Executor 重复召回同一文档 | 中 | 已修复 | [ISS-001-duplicate-retrieval.md](ISS-001-duplicate-retrieval.md) |
|
||||
| ISS-002 | Executor 无约束重复调用 lookup_knowledge | 中 | 已修复 | [ISS-002-executor-unconstrained-lookup.md](ISS-002-executor-unconstrained-lookup.md) |
|
||||
| ISS-003 | Executor 域级检索水位控制(Phase 2) | 低 | 待规划 | [ISS-003-executor-domain-hard-limit.md](ISS-003-executor-domain-hard-limit.md) |
|
||||
|
||||
@@ -0,0 +1 @@
|
||||
ready
|
||||
@@ -0,0 +1,42 @@
|
||||
{
|
||||
"id": "chat-verifier-agent",
|
||||
"metadata": {
|
||||
"status": "archived",
|
||||
"created_at": "2026-07-02",
|
||||
"updated_at": "2026-07-03",
|
||||
"archive_readiness": "archived",
|
||||
"implementation_status": "archived"
|
||||
},
|
||||
"summary": "Add a verifier agent to the complex chat path and persist auditable verifier decisions with evidence traceability.",
|
||||
"artifacts": {
|
||||
"proposal": "proposal.md",
|
||||
"design": "design.md",
|
||||
"tasks": "tasks.md",
|
||||
"specs": [
|
||||
"specs/chat-verifier-agent/spec.md"
|
||||
],
|
||||
"devflow": "devflow/projects/2026-07-02-chat-verifier-agent"
|
||||
},
|
||||
"tasks": [
|
||||
"Verifier prompt",
|
||||
"VerifierInputHook explicit payload",
|
||||
"ChatService planner-executor-verifier orchestration",
|
||||
"Verdict routing and fixed user output templates",
|
||||
"Tool trace summary and evidence_refs traceability",
|
||||
"self_evaluation merge semantics",
|
||||
"Compile and runtime verification"
|
||||
],
|
||||
"verification": [
|
||||
{
|
||||
"type": "script",
|
||||
"command": "mvn -q -DskipTests compile",
|
||||
"result": "passed"
|
||||
},
|
||||
{
|
||||
"type": "runtime",
|
||||
"command": "POST /api/chat",
|
||||
"session_id": "9138f064",
|
||||
"result": "planner, executor, and verifier executed; verifier_evaluation contains evidence_refs and tool_trace_summary source_invocation_ids"
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,349 @@
|
||||
## Context
|
||||
|
||||
Chat 多 Agent 链路当前由 ChatService 驱动 Planner → Executor,答案输出前无质量门禁。Verifier Agent 作为 Executor 后置质量门禁,在 Executor 输出后做事实核查。
|
||||
|
||||
前序 change `executor-action-memory-relevance` 已在 Executor 侧构建了行动记忆和检索质量归一化,Verifier 不需要重复验证检索质量。
|
||||
|
||||
## Goals / Non-Goals
|
||||
|
||||
**Goals:**
|
||||
- Verifier 作为无工具 ReactAgent,由 ChatService 显式调用
|
||||
- Verifier 输出 verdict (PASS/LOW_CONFID/REJECT) + groundedness_score + facts_checked
|
||||
- ChatService 负责单轮显式编排:Planner → Executor → Verifier
|
||||
- ChatService 外层根据 Verifier 判决做轮次路由:PASS→输出,LOW_CONFID≥0.5→带声明输出,LOW_CONFID<0.5→补充一轮,REJECT→降级
|
||||
- Verifier 判决写入 diagnosis_session.self_evaluation JSON 容器中的 `verifier_evaluation` 槽位做可观测
|
||||
|
||||
**Non-Goals:**
|
||||
- Verifier 不调用工具
|
||||
- 不改动单 Agent 链路
|
||||
- 不修改 Executor 的输出内容
|
||||
- 不涉及数据库表结构变更
|
||||
- Verifier 不继承 Executor 的中间推理过程(通过 MessagesModelHook 过滤)
|
||||
|
||||
## Decisions
|
||||
|
||||
| 决策 | 选择 | 放弃方案 | 原因 |
|
||||
|------|------|---------|------|
|
||||
| Verifier 是否有工具 | 无工具 ReactAgent | 有工具的 Agent | 职责单一,只核查不检索 |
|
||||
| 判决分类 | PASS / LOW_CONFID / REJECT | PASS / FAIL 二分类 | LOW_CONFID 提供了弹性输出路径 |
|
||||
| 回调机制 | ChatService 外层控制最多两轮 | 全交给 Supervisor / 不回调 | 轮次上限需要硬控制,不能只靠 prompt 记忆 |
|
||||
| 可观测方案 | 写入 self_evaluation JSON 容器 | agent_step / tool_invocation / 新表 | 不改表结构,同时避免与 evidence_score 覆盖冲突 |
|
||||
| 输入隔离 | 显式状态输入 + MessagesModelHook 裁剪噪音 | 仅靠原始消息过滤 / 数据库注入 | Verifier 需要稳定读取 query、工具摘要、最终答案,不能依赖消息格式猜测 |
|
||||
| 阈值配置 | yml 配置化 | 硬编码 | 方便运维调整,不需改代码 |
|
||||
|
||||
## Verifier 输入契约
|
||||
|
||||
Verifier 的业务输入由 `ChatService` 显式组装,不依赖原始 conversation messages 的隐式结构。
|
||||
|
||||
### 必选输入
|
||||
|
||||
- `original_query`:用户原始问题
|
||||
- `executor_final_answer`:本轮 Executor 最终答案
|
||||
- `tool_trace_summary`:由工具调用事实整理出的半结构化摘要
|
||||
|
||||
### 条件输入
|
||||
|
||||
- `retry_context`:仅第二轮注入,描述上一轮 verifier 发现的证据缺口和补充约束
|
||||
|
||||
### tool_trace_summary 最小结构
|
||||
|
||||
```json
|
||||
[
|
||||
{
|
||||
"tool_name": "lookup_knowledge",
|
||||
"success": true,
|
||||
"input_summary": "查询 ERR_TIMEOUT",
|
||||
"output_summary": "命中 payment/errors.md,返回错误码定义",
|
||||
"evidence_level": "direct"
|
||||
},
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": false,
|
||||
"input_summary": "按 traceId 查询日志",
|
||||
"output_summary": "日志服务超时",
|
||||
"evidence_level": "none"
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
约束:
|
||||
|
||||
- `tool_trace_summary` 只纳入证据型工具调用,不纳入纯辅助或无业务事实意义的工具
|
||||
- `tool_trace_summary` 来源于工具调用事实,不直接透传原始日志全文
|
||||
- Verifier 基于摘要做事实核查,不直接读取数据库
|
||||
- 若某工具调用失败,仍需记录在摘要中,供 Verifier 判断证据缺口
|
||||
|
||||
### 证据型工具边界
|
||||
|
||||
默认纳入 `tool_trace_summary` 的工具:
|
||||
|
||||
- `lookup_knowledge`
|
||||
- `query_logs`
|
||||
- `query_metrics`
|
||||
- `query_order` 或其他业务事实查询类工具
|
||||
- 其他只读、能提供客观事实的工具
|
||||
|
||||
默认不纳入:
|
||||
|
||||
- `getCurrentDateTime`
|
||||
- 纯格式化、转换、控制类工具
|
||||
- 与事实核查无关的辅助工具
|
||||
|
||||
### 摘要压缩规则
|
||||
|
||||
- 每次调用只保留“最小证据摘要”,不透传原始返回全文
|
||||
- `output_summary` 控制为 1-3 句,重点描述“这次调用证明了什么 / 没能证明什么”
|
||||
- 失败调用必须保留,但统一标记:
|
||||
- `success=false`
|
||||
- `evidence_level=none`
|
||||
- 同一工具、同一主题域、同一轮次的重复调用可以折叠为一条合并摘要
|
||||
- 合并摘要至少保留:
|
||||
- 首次有效命中结果
|
||||
- 额外重复次数 / 未命中次数 / 失败次数
|
||||
|
||||
### 截断优先级
|
||||
|
||||
若 `tool_trace_summary` 过长,优先保留:
|
||||
|
||||
1. 被 `executor_final_answer` 直接引用的证据
|
||||
2. 支撑根因结论的证据
|
||||
3. 支撑修复结论的证据
|
||||
4. 与上一轮 `retry_context` 缺口直接相关的证据
|
||||
|
||||
低优先级、与最终答案无关的辅助性工具摘要可被截断。
|
||||
|
||||
### retry_context 最小结构
|
||||
|
||||
```json
|
||||
{
|
||||
"round": 1,
|
||||
"missing_evidence_facts": [
|
||||
"“根因是连接池耗尽”缺少直接证据",
|
||||
"“错误码 ERR_TIMEOUT 来自支付网关”只有间接支持"
|
||||
],
|
||||
"instruction": "仅补充以上断言相关证据,不要重复已完成检索"
|
||||
}
|
||||
```
|
||||
|
||||
### MessagesModelHook 职责边界
|
||||
|
||||
- 可以:移除 Planner/Executor 中间推理、无关闲聊和冗余 message
|
||||
- 不可以:作为 Verifier 核心业务输入的唯一来源
|
||||
- 目标:降噪,而非拼装业务事实
|
||||
|
||||
## Verifier 判决矩阵
|
||||
|
||||
Verifier 先提取并校验 `facts_checked`,再依据矩阵生成 verdict,避免只靠模型主观判断。
|
||||
|
||||
### facts_checked 分类
|
||||
|
||||
每条事实仅允许以下四类之一:
|
||||
|
||||
- `direct_evidence`:工具结果中有明确直接证据
|
||||
- `indirect_support`:可由工具结果合理推导,但不是直接陈述
|
||||
- `no_evidence`:工具结果中没有足够信息支撑
|
||||
- `contradicted`:工具结果与该事实冲突,或该事实编造了不存在的关键实体/错误码/结论
|
||||
|
||||
### 关键事实范围
|
||||
|
||||
Verifier 优先校验关键事实,至少包括:
|
||||
|
||||
- 根因结论(root cause)
|
||||
- 错误码 / 接口 / 组件归属
|
||||
- 证据来源陈述(如“日志显示”“文档说明”)
|
||||
- 明确修复结论
|
||||
|
||||
一般性建议、风险提示、非事实性表述默认不纳入关键事实,除非答案明确声称“已被证据证明”。
|
||||
|
||||
### verdict 规则
|
||||
|
||||
- `REJECT`
|
||||
- 任意关键事实为 `contradicted`
|
||||
- 或答案编造了工具/日志/文档中不存在的关键实体、错误码、结论
|
||||
|
||||
- `PASS`
|
||||
- 所有关键事实均为 `direct_evidence` 或 `indirect_support`
|
||||
- 且至少一条关键事实为 `direct_evidence`
|
||||
- 且不存在 `contradicted`
|
||||
|
||||
- `LOW_CONFID`
|
||||
- 不存在 `contradicted`
|
||||
- 但存在关键事实为 `no_evidence`
|
||||
- 或所有关键事实都只有 `indirect_support`,缺少直接锚点
|
||||
|
||||
一句话归纳:
|
||||
|
||||
- `REJECT` = 有冲突
|
||||
- `LOW_CONFID` = 无冲突但缺关键证据
|
||||
- `PASS` = 无冲突且关键事实均有支撑
|
||||
|
||||
### groundedness_score 计算
|
||||
|
||||
`groundedness_score` 不由模型自由打分,而由关键事实分类映射得到:
|
||||
|
||||
```text
|
||||
direct_evidence = 1.0
|
||||
indirect_support = 0.6
|
||||
no_evidence = 0.0
|
||||
contradicted = 0.0
|
||||
```
|
||||
|
||||
规则:
|
||||
|
||||
- 仅对关键事实计分
|
||||
- 取平均值后截断到 `[0.0, 1.0]`
|
||||
- 若存在任意关键事实为 `contradicted`,直接 verdict=`REJECT`,且 `groundedness_score=0.0`
|
||||
|
||||
### 第二轮补证据范围
|
||||
|
||||
第二轮 `retry_context` 仅回灌以下关键缺口:
|
||||
|
||||
- 关键事实为 `no_evidence`
|
||||
- 关键事实为 `indirect_support`,但仍缺直接证据锚点
|
||||
|
||||
`REJECT` 不进入第二轮补证据,直接降级输出。
|
||||
|
||||
## 用户侧输出协议
|
||||
|
||||
Verifier 的内部判决与用户侧最终输出类型分离:
|
||||
|
||||
- `PASS` → `NORMAL`
|
||||
- `LOW_CONFID` → `LOW_CONFID_WITH_DISCLAIMER`
|
||||
- `REJECT` → `DEGRADED`
|
||||
|
||||
### LOW_CONFID_WITH_DISCLAIMER
|
||||
|
||||
适用场景:
|
||||
|
||||
- 第一轮 `LOW_CONFID` 且 `groundedness_score >= threshold`
|
||||
- 第二轮后仍为 `LOW_CONFID`
|
||||
|
||||
输出规则:
|
||||
|
||||
- 使用固定免责声明前缀
|
||||
- 免责声明后拼接 `executor_final_answer`
|
||||
- 可选附加“当前证据缺口”列表,但来源必须是 verifier 的关键缺口,不得自由扩写
|
||||
|
||||
建议模板:
|
||||
|
||||
```text
|
||||
以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。
|
||||
|
||||
{executor_final_answer}
|
||||
|
||||
当前缺口:
|
||||
- ...
|
||||
- ...
|
||||
```
|
||||
|
||||
### DEGRADED
|
||||
|
||||
适用场景:
|
||||
|
||||
- 任意一轮 `REJECT`
|
||||
- 系统无法基于现有证据形成可靠结论
|
||||
|
||||
输出规则:
|
||||
|
||||
- 不透传原始 `executor_final_answer`
|
||||
- 使用固定降级模板
|
||||
- 仅允许包含:
|
||||
- 已确认信息
|
||||
- 证据缺口
|
||||
- 下一步建议
|
||||
|
||||
建议模板:
|
||||
|
||||
```text
|
||||
当前无法基于已获取证据生成可靠结论,建议人工介入。
|
||||
|
||||
已确认信息:
|
||||
- ...
|
||||
|
||||
证据缺口:
|
||||
- ...
|
||||
|
||||
建议下一步:
|
||||
- ...
|
||||
```
|
||||
|
||||
### 输出边界
|
||||
|
||||
- `LOW_CONFID_WITH_DISCLAIMER` 可以带出原始答案,但必须加固定免责声明
|
||||
- `DEGRADED` 不得透传未经验证的原始答案
|
||||
- 用户侧输出模板由代码层拼装,不依赖 Verifier 自由生成
|
||||
|
||||
## self_evaluation 存储约定
|
||||
|
||||
`diagnosis_session.self_evaluation` 统一定义为 JSON 容器对象,而不是单一评估结果:
|
||||
|
||||
```json
|
||||
{
|
||||
"rule_evaluation": {
|
||||
"evidence_score": 65,
|
||||
"source": "rule",
|
||||
"factors": []
|
||||
},
|
||||
"verifier_evaluation": {
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.42,
|
||||
"facts_checked": [],
|
||||
"rationale": "...",
|
||||
"round": 1
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
写入约束:
|
||||
|
||||
- `EvaluationService` 只负责写 `rule_evaluation`
|
||||
- `ChatService` 只负责写 `verifier_evaluation`
|
||||
- 两侧都必须使用 read-modify-write,保留另一侧已有内容
|
||||
- 禁止整段覆盖 `self_evaluation`,除非初始化为空对象
|
||||
|
||||
## Risks / Trade-offs
|
||||
|
||||
- [Risk] Verifier 误判导致好答案被降级 → Mitigation: REJECT 仅用于明显编造场景,LOW_CONFID 为主要输出路径
|
||||
- [Risk] callback Planner 后新答案质量不一定提升 → Mitigation: 仅回调一次,Token 成本可控
|
||||
- [Risk] 第二轮仍可能产出 REJECT → Mitigation: 第二轮 REJECT 仍降级,不透传
|
||||
- [Risk] `self_evaluation` 被异步 evidence_score 覆盖 → Mitigation: 定义 JSON 容器槽位,统一 read-modify-write
|
||||
- [Risk] Verifier 增加 Token 消耗 → Mitigation: 单次轻量 LLM 调用,估算 <500 token
|
||||
- [Risk] 消息过滤可能导致输入契约漂移 → Mitigation: 主输入由显式状态输入提供,Hook 仅用于剔除中间推理和无关噪音
|
||||
- [Risk] 判决边界主观化,导致不同模型输出不稳定 → Mitigation: 用 facts_checked 分类 + verdict 矩阵 + 映射分数约束输出
|
||||
- [Risk] 最终用户文案随模型漂移,导致产品行为不稳定 → Mitigation: LOW_CONFID/DEGRADED 使用固定输出协议和模板
|
||||
|
||||
## Implementation Plan
|
||||
|
||||
1. 创建 `chat-verifier-prompt.md`
|
||||
2. 新建 `VerifierInputHook.java`(MessagesModelHook 实现,BEFORE_MODEL 时裁剪 messages,只保留必要上下文)
|
||||
3. `ChatService.java` 新增 `buildChatVerifierAgent()` 方法(ReactAgent,无工具,带 hook)
|
||||
4. 添加 `verifier.low-confidence-threshold: 0.5` 到 application.yml
|
||||
5. 在 `ChatService.executeChatComplex()` 中显式调用 `Planner → Executor → Verifier`
|
||||
6. 保留 `SupervisorAgent` 构造作为 legacy residue,不再依赖 prompt-only supervisor sequencing 保证 Verifier 执行
|
||||
7. 在 `ChatService.executeChatComplex()` 外层实现最多两轮调用控制
|
||||
8. 组装 Verifier 显式状态输入:`original_query` / `executor_final_answer` / `tool_trace_summary` / `retry_context`
|
||||
9. 将 `self_evaluation` 升级为 JSON 容器读写:`rule_evaluation` / `verifier_evaluation`
|
||||
10. 读取 Verifier 判决写入 `verifier_evaluation`
|
||||
|
||||
## Implementation Notes
|
||||
|
||||
### Explicit orchestration
|
||||
|
||||
The final implementation uses `ChatService` to call `planner -> executor -> verifier` directly in each outer round. This replaces the earlier prompt-only dependency on `SupervisorAgent` for verifier execution. The supervisor construction remains in the code as legacy residue, but runtime correctness is driven by explicit `callAgent(...)` ordering.
|
||||
|
||||
### Traceability model
|
||||
|
||||
The implemented verifier input and persisted evaluation include an evidence index:
|
||||
|
||||
- `tool_trace_summary[*].trace_ref`
|
||||
- `tool_trace_summary[*].source_invocation_ids`
|
||||
- `tool_trace_summary[*].query_samples`
|
||||
- `tool_trace_summary[*].retrieval_layers`
|
||||
- `tool_trace_summary[*].relevance_levels`
|
||||
- `tool_trace_summary[*].source_documents`
|
||||
|
||||
Each verifier fact may carry `facts_checked[*].evidence_refs`, which points back to `trace_ref` and the underlying `tool_invocation` ids. This closes the audit gap where verifier could list many checked facts but the reviewer could not tell which facts related to which tool calls.
|
||||
|
||||
### Observability adjustment
|
||||
|
||||
`agent_step.thought` is now intentionally concise for verifier steps. Full verifier judgment belongs in `diagnosis_session.self_evaluation.verifier_evaluation`, with `model_output` retaining the model output snapshot.
|
||||
@@ -0,0 +1,31 @@
|
||||
# Proposal: chat-verifier-agent
|
||||
|
||||
## Why
|
||||
|
||||
Chat 多 Agent 链路缺少出口质量门禁。Executor 输出答案后会直接返回给用户,无法在返回前拦截缺证据、低置信或明显编造的结论。
|
||||
|
||||
## What Changes
|
||||
|
||||
- 新增无工具 Verifier Agent,在 Executor 输出后读取答案和工具调用证据摘要,产出 `PASS` / `LOW_CONFID` / `REJECT` 判决。
|
||||
- `ChatService` 显式编排 `Planner -> Executor -> Verifier`,并根据 Verifier 判决控制最终输出或最多一次补充轮次。
|
||||
- Verifier 输入使用显式状态块:`original_query`、`executor_final_answer`、`tool_trace_summary`、第二轮可选 `retry_context`。
|
||||
- `diagnosis_session.self_evaluation` 作为 JSON 容器保存 `rule_evaluation` 与 `verifier_evaluation`,避免异步评分覆盖 Verifier 结果。
|
||||
- Verifier 结果增加可追溯证据引用:`tool_trace_summary[*].trace_ref`、`source_invocation_ids` 与 `facts_checked[*].evidence_refs`。
|
||||
- `LOW_CONFID` 和 `REJECT` 用户侧输出使用固定协议,`REJECT` 不透传未经验证的原始答案。
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
- `chat-verifier-agent`: Chat 多 Agent 出口事实核查、判决路由、观测存储和证据可追溯能力。
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- None.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected code: `ChatService`, chat verifier prompt, verifier input assembly, self-evaluation persistence, multi-agent runtime orchestration.
|
||||
- Affected runtime behavior: complex chat path now runs a Verifier gate after Executor and may perform one bounded retry for low-confidence evidence gaps.
|
||||
- No database schema change is required; `self_evaluation` remains the persistence container.
|
||||
- Non-goals: Verifier 不调用工具、不改写 Executor 答案、不影响单 Agent 链路、不支持超过两轮的补充编排。
|
||||
+189
@@ -0,0 +1,189 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Verifier SHALL fact-check Executor answers
|
||||
The system SHALL have a Verifier Agent that reads the Executor's answer and the tool call history, then produces a structured verdict.
|
||||
|
||||
#### Scenario: PASS verdict when all claims have evidence
|
||||
- **WHEN** all critical facts in the Executor's answer have direct or indirect support in tool call results
|
||||
- **AND** at least one critical fact has direct evidence
|
||||
- **AND** no critical fact is contradicted
|
||||
- **THEN** the Verifier SHALL output verdict="PASS" with groundedness_score ≥ 0.5
|
||||
|
||||
#### Scenario: LOW_CONFID verdict with partial evidence
|
||||
- **WHEN** no critical fact contradicts the tool results
|
||||
- **AND** some critical facts have no supporting evidence
|
||||
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
|
||||
|
||||
#### Scenario: LOW_CONFID verdict with only indirect support
|
||||
- **WHEN** no critical fact contradicts the tool results
|
||||
- **AND** all critical facts are only indirectly supported
|
||||
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
|
||||
|
||||
#### Scenario: REJECT verdict when claims contradict evidence
|
||||
- **WHEN** any critical fact in the Executor's answer contradicts tool call results
|
||||
- **OR** the answer fabricates a key entity, error code, or conclusion that does not exist in the tool evidence
|
||||
- **THEN** the Verifier SHALL output verdict="REJECT"
|
||||
|
||||
### Requirement: Verifier SHALL output structured JSON
|
||||
The Verifier SHALL output a JSON object with verdict, groundedness_score, facts_checked array, and rationale.
|
||||
|
||||
#### Scenario: Output format validation
|
||||
- **WHEN** the Verifier completes its analysis
|
||||
- **THEN** the output SHALL contain "verdict", "groundedness_score", "facts_checked", and "rationale" fields
|
||||
- **AND** groundedness_score SHALL be a float between 0.0 and 1.0
|
||||
- **AND** verdict SHALL be one of "PASS", "LOW_CONFID", or "REJECT"
|
||||
|
||||
#### Scenario: strict schema output
|
||||
- **WHEN** the Verifier returns its result
|
||||
- **THEN** it SHALL output exactly one JSON object
|
||||
- **AND** it SHALL NOT output Markdown, code fences, or explanatory text outside the JSON object
|
||||
- **AND** the JSON object SHALL include `critical_fact_count`
|
||||
- **AND** each `facts_checked` item SHALL include `fact`, `is_critical`, `verification`, and `detail`
|
||||
|
||||
### Requirement: facts_checked SHALL use a fixed classification set
|
||||
Each checked fact SHALL be labeled using a fixed evidence classification.
|
||||
|
||||
#### Scenario: fact classification values
|
||||
- **WHEN** the Verifier emits `facts_checked`
|
||||
- **THEN** each fact SHALL use one of `direct_evidence`, `indirect_support`, `no_evidence`, or `contradicted`
|
||||
|
||||
### Requirement: groundedness_score SHALL be derived from fact classifications
|
||||
The groundedness score SHALL be computed from critical fact classifications instead of being freely chosen by the model.
|
||||
|
||||
#### Scenario: contradicted fact forces reject
|
||||
- **WHEN** any critical fact is labeled `contradicted`
|
||||
- **THEN** the Verifier SHALL output verdict="REJECT"
|
||||
- **AND** groundedness_score SHALL be `0.0`
|
||||
|
||||
#### Scenario: score derived from supported facts
|
||||
- **WHEN** no critical fact is contradicted
|
||||
- **THEN** groundedness_score SHALL be computed from the mapped values of critical facts
|
||||
- **AND** the implementation SHALL use the fixed mapping `direct_evidence=1.0`, `indirect_support=0.6`, `no_evidence=0.0`
|
||||
- **AND** the result SHALL be clamped into `[0.0, 1.0]`
|
||||
|
||||
### Requirement: ChatService SHALL route based on Verifier verdict
|
||||
The system SHALL use ChatService for explicit single-round `Planner → Executor → Verifier` orchestration and SHALL use ChatService to control whether an additional round is allowed.
|
||||
|
||||
#### Scenario: PASS → direct output
|
||||
- **WHEN** Verifier outputs verdict="PASS"
|
||||
- **THEN** the system SHALL output the Executor's answer directly
|
||||
|
||||
#### Scenario: LOW_CONFID score≥0.5 → output with disclaimer
|
||||
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score ≥ 0.5
|
||||
- **THEN** the system SHALL output the Executor's answer prefixed with a fixed confidence disclaimer
|
||||
|
||||
#### Scenario: LOW_CONFID score<0.5 → trigger one additional round
|
||||
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score < 0.5 and this is the first callback
|
||||
- **THEN** the ChatService SHALL invoke one additional `Planner → Executor → Verifier` round to supplement evidence
|
||||
- **AND** after the second Verifier run, verdict="LOW_CONFID" SHALL be output with a confidence disclaimer
|
||||
- **AND** after the second Verifier run, verdict="REJECT" SHALL still produce a degraded output
|
||||
|
||||
#### Scenario: REJECT does not enter retry round
|
||||
- **WHEN** Verifier outputs verdict="REJECT"
|
||||
- **THEN** the system SHALL NOT start a retry round for evidence补充
|
||||
- **AND** it SHALL produce a degraded output directly
|
||||
|
||||
#### Scenario: REJECT → degraded output
|
||||
- **WHEN** Verifier outputs verdict="REJECT"
|
||||
- **THEN** the system SHALL output a degraded result indicating the answer cannot be reliably generated
|
||||
- **AND** it SHALL NOT pass through the raw Executor answer
|
||||
|
||||
### Requirement: User-facing verifier outputs SHALL follow fixed templates
|
||||
The system SHALL use fixed output protocols for LOW_CONFID and REJECT user-facing responses.
|
||||
|
||||
#### Scenario: LOW_CONFID uses disclaimer template
|
||||
- **WHEN** the final verdict is `LOW_CONFID`
|
||||
- **THEN** the user-facing response SHALL prepend a fixed disclaimer before the Executor answer
|
||||
- **AND** optional evidence gaps, if present, SHALL come only from verifier-identified critical gaps
|
||||
|
||||
#### Scenario: REJECT uses degraded template
|
||||
- **WHEN** the final verdict is `REJECT`
|
||||
- **THEN** the user-facing response SHALL use a degraded template
|
||||
- **AND** it SHALL include only confirmed facts, evidence gaps, and next-step suggestions
|
||||
- **AND** it SHALL NOT include unverified raw answer content
|
||||
|
||||
### Requirement: Verifier SHALL be observable
|
||||
The Verifier's verdict SHALL be persisted for observability.
|
||||
|
||||
#### Scenario: verdict written to self_evaluation
|
||||
- **WHEN** the Verifier produces a verdict
|
||||
- **THEN** the ChatService SHALL write the verdict data under `diagnosis_session.self_evaluation.verifier_evaluation`
|
||||
- **AND** existing `rule_evaluation` data SHALL be preserved
|
||||
|
||||
### Requirement: self_evaluation SHALL be a container object
|
||||
The `diagnosis_session.self_evaluation` field SHALL store multiple evaluation channels in one JSON object.
|
||||
|
||||
#### Scenario: rule evaluation stored separately
|
||||
- **WHEN** the rule-based evidence scoring completes
|
||||
- **THEN** the EvaluationService SHALL write the result under `rule_evaluation`
|
||||
- **AND** existing `verifier_evaluation` data SHALL be preserved
|
||||
|
||||
#### Scenario: verifier evaluation stored separately
|
||||
- **WHEN** the Verifier completes
|
||||
- **THEN** the ChatService SHALL write the result under `verifier_evaluation`
|
||||
- **AND** existing `rule_evaluation` data SHALL be preserved
|
||||
|
||||
#### Scenario: no whole-object overwrite after initialization
|
||||
- **WHEN** either evaluation channel updates `self_evaluation`
|
||||
- **THEN** the implementation SHALL use read-modify-write semantics
|
||||
- **AND** it SHALL NOT replace the whole JSON object except when initializing from null
|
||||
|
||||
### Requirement: Verifier SHALL consume explicit verification inputs
|
||||
The Verifier SHALL receive explicit verification inputs rather than inferring them only from raw conversation history.
|
||||
|
||||
#### Scenario: explicit input blocks available to Verifier
|
||||
- **WHEN** the Verifier starts
|
||||
- **THEN** the system SHALL provide `original_query`, `executor_final_answer`, and `tool_trace_summary` as explicit inputs
|
||||
- **AND** `retry_context` SHALL be provided on the second round only
|
||||
- **AND** message filtering MAY be used only to remove intermediate reasoning or unrelated noise
|
||||
|
||||
#### Scenario: tool trace summary derived from tool facts
|
||||
- **WHEN** the system prepares verifier inputs
|
||||
- **THEN** `tool_trace_summary` SHALL be generated from tool invocation facts
|
||||
- **AND** each summary item SHALL include tool name, success state, input summary, output summary, and evidence level
|
||||
- **AND** raw conversation history SHALL NOT be the only source of verifier evidence context
|
||||
|
||||
#### Scenario: tool trace summary preserves invocation references
|
||||
- **WHEN** the system prepares verifier inputs
|
||||
- **THEN** each summary item SHALL include a stable `trace_ref`
|
||||
- **AND** each summary item SHALL preserve `source_invocation_ids` for the tool invocation rows that contributed to the summary
|
||||
- **AND** each summary item SHOULD include query samples, retrieval layers, relevance levels, and source document labels when available
|
||||
|
||||
#### Scenario: only evidence-bearing tools included
|
||||
- **WHEN** the system generates `tool_trace_summary`
|
||||
- **THEN** it SHALL include only evidence-bearing tool invocations
|
||||
- **AND** non-evidence helper tools such as time or formatting tools SHALL be excluded by default
|
||||
|
||||
#### Scenario: failed evidence calls preserved as evidence gaps
|
||||
- **WHEN** an evidence-bearing tool invocation fails or returns no usable evidence
|
||||
- **THEN** the summary SHALL still include that invocation
|
||||
- **AND** it SHALL mark the entry as unsuccessful with an evidence level representing no evidence
|
||||
|
||||
#### Scenario: repeated tool calls may be compacted
|
||||
- **WHEN** repeated tool invocations concern the same tool, topic domain, and round
|
||||
- **THEN** the system MAY compact them into a merged summary entry
|
||||
- **AND** the merged entry SHALL preserve the first effective hit and the count of repeated, failed, or no-hit calls
|
||||
|
||||
#### Scenario: raw outputs not passed through in full
|
||||
- **WHEN** a tool invocation returns large raw content
|
||||
- **THEN** `tool_trace_summary` SHALL keep only a minimal evidence summary
|
||||
- **AND** the raw output SHALL NOT be passed through in full to the Verifier
|
||||
|
||||
#### Scenario: MessagesModelHook used only for noise reduction
|
||||
- **WHEN** a MessagesModelHook is used for the Verifier
|
||||
- **THEN** it MAY remove intermediate reasoning or irrelevant messages
|
||||
- **AND** it SHALL NOT be the primary source for assembling verifier business inputs
|
||||
|
||||
### Requirement: Verifier facts SHALL be auditable
|
||||
Verifier facts SHALL be linkable to the evidence summaries used during verification.
|
||||
|
||||
#### Scenario: facts_checked contains evidence refs
|
||||
- **WHEN** the Verifier emits `facts_checked`
|
||||
- **THEN** each fact SHALL include `evidence_refs`
|
||||
- **AND** each evidence ref SHALL point to an existing `tool_trace_summary.trace_ref`
|
||||
- **AND** each evidence ref SHALL preserve the relevant `source_invocation_ids` when available
|
||||
|
||||
#### Scenario: verifier evaluation persists traceability snapshot
|
||||
- **WHEN** the ChatService persists `verifier_evaluation`
|
||||
- **THEN** it SHALL include `traceability_version`
|
||||
- **AND** it SHALL include the `tool_trace_summary` snapshot used by the Verifier
|
||||
@@ -0,0 +1,58 @@
|
||||
# Tasks: chat-verifier-agent
|
||||
|
||||
## 1. Verifier Prompt
|
||||
|
||||
- [x] 1.1 Create `src/main/resources/prompts/chat-verifier-prompt.md`.
|
||||
- [x] 1.2 Define fixed fact classifications: `direct_evidence`, `indirect_support`, `no_evidence`, `contradicted`.
|
||||
- [x] 1.3 Define critical fact scope, verdict matrix, and `groundedness_score` mapping.
|
||||
- [x] 1.4 Define strict JSON output schema: `verdict`, `groundedness_score`, `critical_fact_count`, `facts_checked`, `rationale`.
|
||||
- [x] 1.5 Forbid Markdown, code fences, schema-extra fields, and text outside the JSON object.
|
||||
- [x] 1.6 Require `facts_checked[*].evidence_refs` for traceability to tool evidence.
|
||||
|
||||
## 2. Verifier Input Hook
|
||||
|
||||
- [x] 2.1 Add `VerifierInputHook.java` as a `MessagesModelHook` running at `BEFORE_MODEL`.
|
||||
- [x] 2.2 Replace raw verifier history with explicit payload fields: `original_query`, `executor_final_answer`, `tool_trace_summary`, `retry_context`.
|
||||
- [x] 2.3 Persist the current round `tool_trace_summary` in `VerifierContextHolder` for later verifier evaluation storage.
|
||||
|
||||
## 3. ChatService Integration
|
||||
|
||||
- [x] 3.1 Load `chatVerifierPrompt` and add `buildChatVerifierAgent()`.
|
||||
- [x] 3.2 Add configurable `verifier.low-confidence-threshold`.
|
||||
- [x] 3.3 Implement explicit per-round orchestration in `ChatService`: planner call, executor call, verifier call.
|
||||
- [x] 3.4 Keep max two outer rounds and inject `retry_context` only for the second round.
|
||||
- [x] 3.5 Parse verifier JSON directly and fall back to `LOW_CONFID` when verifier output is missing or invalid.
|
||||
- [x] 3.6 Keep `SupervisorAgent` construction as legacy residue only; runtime orchestration no longer depends on prompt-only supervisor sequencing.
|
||||
|
||||
## 4. Verdict Routing And User Output
|
||||
|
||||
- [x] 4.1 Route `PASS` to the executor answer.
|
||||
- [x] 4.2 Route `LOW_CONFID` to a fixed disclaimer plus executor answer.
|
||||
- [x] 4.3 Route `REJECT` to degraded output without passing through the raw unverified answer.
|
||||
- [x] 4.4 Build LOW_CONFID gap lists only from verifier-identified gaps.
|
||||
- [x] 4.5 Build DEGRADED confirmed facts, gaps, and next-step suggestions from verifier facts and trace summary.
|
||||
|
||||
## 5. Trace Summary And Observability
|
||||
|
||||
- [x] 5.1 Add `ToolTraceSummaryService` to build verifier evidence summaries from `tool_invocation`.
|
||||
- [x] 5.2 Include only evidence-bearing tools by default.
|
||||
- [x] 5.3 Compact repeated calls by tool and topic domain.
|
||||
- [x] 5.4 Preserve `source_invocation_ids`, `trace_ref`, query samples, retrieval layers, relevance levels, and source document labels.
|
||||
- [x] 5.5 Parse and persist `facts_checked[*].evidence_refs`.
|
||||
- [x] 5.6 Persist `verifier_evaluation.tool_trace_summary` and `traceability_version`.
|
||||
- [x] 5.7 Store concise verifier summaries in `agent_step.thought` while preserving fuller verifier output in `model_output` / `self_evaluation`.
|
||||
|
||||
## 6. self_evaluation Merge Semantics
|
||||
|
||||
- [x] 6.1 Add `SelfEvaluationMergeService`.
|
||||
- [x] 6.2 Write verifier results under `verifier_evaluation`.
|
||||
- [x] 6.3 Write rule scoring under `rule_evaluation`.
|
||||
- [x] 6.4 Preserve the other channel with read-modify-write semantics.
|
||||
|
||||
## 7. Verification
|
||||
|
||||
- [x] 7.1 Compile verification: `mvn -q -DskipTests compile`.
|
||||
- [x] 7.2 Runtime verification: `/api/chat` complex request reached `planner -> executor -> verifier`.
|
||||
- [x] 7.3 Runtime verification: session `9138f064` persisted `verifier_evaluation.facts_checked[*].evidence_refs`.
|
||||
- [x] 7.4 Runtime verification: session `9138f064` persisted `tool_trace_summary[*].source_invocation_ids`.
|
||||
- [x] 7.5 Runtime verification: LOW_CONFID user output included disclaimer and verifier-derived gaps.
|
||||
@@ -0,0 +1,192 @@
|
||||
# chat-verifier-agent Specification
|
||||
|
||||
## Purpose
|
||||
TBD - created by archiving change chat-verifier-agent. Update Purpose after archive.
|
||||
## Requirements
|
||||
### Requirement: Verifier SHALL fact-check Executor answers
|
||||
The system SHALL have a Verifier Agent that reads the Executor's answer and the tool call history, then produces a structured verdict.
|
||||
|
||||
#### Scenario: PASS verdict when all claims have evidence
|
||||
- **WHEN** all critical facts in the Executor's answer have direct or indirect support in tool call results
|
||||
- **AND** at least one critical fact has direct evidence
|
||||
- **AND** no critical fact is contradicted
|
||||
- **THEN** the Verifier SHALL output verdict="PASS" with groundedness_score ≥ 0.5
|
||||
|
||||
#### Scenario: LOW_CONFID verdict with partial evidence
|
||||
- **WHEN** no critical fact contradicts the tool results
|
||||
- **AND** some critical facts have no supporting evidence
|
||||
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
|
||||
|
||||
#### Scenario: LOW_CONFID verdict with only indirect support
|
||||
- **WHEN** no critical fact contradicts the tool results
|
||||
- **AND** all critical facts are only indirectly supported
|
||||
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
|
||||
|
||||
#### Scenario: REJECT verdict when claims contradict evidence
|
||||
- **WHEN** any critical fact in the Executor's answer contradicts tool call results
|
||||
- **OR** the answer fabricates a key entity, error code, or conclusion that does not exist in the tool evidence
|
||||
- **THEN** the Verifier SHALL output verdict="REJECT"
|
||||
|
||||
### Requirement: Verifier SHALL output structured JSON
|
||||
The Verifier SHALL output a JSON object with verdict, groundedness_score, facts_checked array, and rationale.
|
||||
|
||||
#### Scenario: Output format validation
|
||||
- **WHEN** the Verifier completes its analysis
|
||||
- **THEN** the output SHALL contain "verdict", "groundedness_score", "facts_checked", and "rationale" fields
|
||||
- **AND** groundedness_score SHALL be a float between 0.0 and 1.0
|
||||
- **AND** verdict SHALL be one of "PASS", "LOW_CONFID", or "REJECT"
|
||||
|
||||
#### Scenario: strict schema output
|
||||
- **WHEN** the Verifier returns its result
|
||||
- **THEN** it SHALL output exactly one JSON object
|
||||
- **AND** it SHALL NOT output Markdown, code fences, or explanatory text outside the JSON object
|
||||
- **AND** the JSON object SHALL include `critical_fact_count`
|
||||
- **AND** each `facts_checked` item SHALL include `fact`, `is_critical`, `verification`, and `detail`
|
||||
|
||||
### Requirement: facts_checked SHALL use a fixed classification set
|
||||
Each checked fact SHALL be labeled using a fixed evidence classification.
|
||||
|
||||
#### Scenario: fact classification values
|
||||
- **WHEN** the Verifier emits `facts_checked`
|
||||
- **THEN** each fact SHALL use one of `direct_evidence`, `indirect_support`, `no_evidence`, or `contradicted`
|
||||
|
||||
### Requirement: groundedness_score SHALL be derived from fact classifications
|
||||
The groundedness score SHALL be computed from critical fact classifications instead of being freely chosen by the model.
|
||||
|
||||
#### Scenario: contradicted fact forces reject
|
||||
- **WHEN** any critical fact is labeled `contradicted`
|
||||
- **THEN** the Verifier SHALL output verdict="REJECT"
|
||||
- **AND** groundedness_score SHALL be `0.0`
|
||||
|
||||
#### Scenario: score derived from supported facts
|
||||
- **WHEN** no critical fact is contradicted
|
||||
- **THEN** groundedness_score SHALL be computed from the mapped values of critical facts
|
||||
- **AND** the implementation SHALL use the fixed mapping `direct_evidence=1.0`, `indirect_support=0.6`, `no_evidence=0.0`
|
||||
- **AND** the result SHALL be clamped into `[0.0, 1.0]`
|
||||
|
||||
### Requirement: ChatService SHALL route based on Verifier verdict
|
||||
The system SHALL use ChatService for explicit single-round `Planner → Executor → Verifier` orchestration and SHALL use ChatService to control whether an additional round is allowed.
|
||||
|
||||
#### Scenario: PASS → direct output
|
||||
- **WHEN** Verifier outputs verdict="PASS"
|
||||
- **THEN** the system SHALL output the Executor's answer directly
|
||||
|
||||
#### Scenario: LOW_CONFID score≥0.5 → output with disclaimer
|
||||
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score ≥ 0.5
|
||||
- **THEN** the system SHALL output the Executor's answer prefixed with a fixed confidence disclaimer
|
||||
|
||||
#### Scenario: LOW_CONFID score<0.5 → trigger one additional round
|
||||
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score < 0.5 and this is the first callback
|
||||
- **THEN** the ChatService SHALL invoke one additional `Planner → Executor → Verifier` round to supplement evidence
|
||||
- **AND** after the second Verifier run, verdict="LOW_CONFID" SHALL be output with a confidence disclaimer
|
||||
- **AND** after the second Verifier run, verdict="REJECT" SHALL still produce a degraded output
|
||||
|
||||
#### Scenario: REJECT does not enter retry round
|
||||
- **WHEN** Verifier outputs verdict="REJECT"
|
||||
- **THEN** the system SHALL NOT start a retry round for evidence补充
|
||||
- **AND** it SHALL produce a degraded output directly
|
||||
|
||||
#### Scenario: REJECT → degraded output
|
||||
- **WHEN** Verifier outputs verdict="REJECT"
|
||||
- **THEN** the system SHALL output a degraded result indicating the answer cannot be reliably generated
|
||||
- **AND** it SHALL NOT pass through the raw Executor answer
|
||||
|
||||
### Requirement: User-facing verifier outputs SHALL follow fixed templates
|
||||
The system SHALL use fixed output protocols for LOW_CONFID and REJECT user-facing responses.
|
||||
|
||||
#### Scenario: LOW_CONFID uses disclaimer template
|
||||
- **WHEN** the final verdict is `LOW_CONFID`
|
||||
- **THEN** the user-facing response SHALL prepend a fixed disclaimer before the Executor answer
|
||||
- **AND** optional evidence gaps, if present, SHALL come only from verifier-identified critical gaps
|
||||
|
||||
#### Scenario: REJECT uses degraded template
|
||||
- **WHEN** the final verdict is `REJECT`
|
||||
- **THEN** the user-facing response SHALL use a degraded template
|
||||
- **AND** it SHALL include only confirmed facts, evidence gaps, and next-step suggestions
|
||||
- **AND** it SHALL NOT include unverified raw answer content
|
||||
|
||||
### Requirement: Verifier SHALL be observable
|
||||
The Verifier's verdict SHALL be persisted for observability.
|
||||
|
||||
#### Scenario: verdict written to self_evaluation
|
||||
- **WHEN** the Verifier produces a verdict
|
||||
- **THEN** the ChatService SHALL write the verdict data under `diagnosis_session.self_evaluation.verifier_evaluation`
|
||||
- **AND** existing `rule_evaluation` data SHALL be preserved
|
||||
|
||||
### Requirement: self_evaluation SHALL be a container object
|
||||
The `diagnosis_session.self_evaluation` field SHALL store multiple evaluation channels in one JSON object.
|
||||
|
||||
#### Scenario: rule evaluation stored separately
|
||||
- **WHEN** the rule-based evidence scoring completes
|
||||
- **THEN** the EvaluationService SHALL write the result under `rule_evaluation`
|
||||
- **AND** existing `verifier_evaluation` data SHALL be preserved
|
||||
|
||||
#### Scenario: verifier evaluation stored separately
|
||||
- **WHEN** the Verifier completes
|
||||
- **THEN** the ChatService SHALL write the result under `verifier_evaluation`
|
||||
- **AND** existing `rule_evaluation` data SHALL be preserved
|
||||
|
||||
#### Scenario: no whole-object overwrite after initialization
|
||||
- **WHEN** either evaluation channel updates `self_evaluation`
|
||||
- **THEN** the implementation SHALL use read-modify-write semantics
|
||||
- **AND** it SHALL NOT replace the whole JSON object except when initializing from null
|
||||
|
||||
### Requirement: Verifier SHALL consume explicit verification inputs
|
||||
The Verifier SHALL receive explicit verification inputs rather than inferring them only from raw conversation history.
|
||||
|
||||
#### Scenario: explicit input blocks available to Verifier
|
||||
- **WHEN** the Verifier starts
|
||||
- **THEN** the system SHALL provide `original_query`, `executor_final_answer`, and `tool_trace_summary` as explicit inputs
|
||||
- **AND** `retry_context` SHALL be provided on the second round only
|
||||
- **AND** message filtering MAY be used only to remove intermediate reasoning or unrelated noise
|
||||
|
||||
#### Scenario: tool trace summary derived from tool facts
|
||||
- **WHEN** the system prepares verifier inputs
|
||||
- **THEN** `tool_trace_summary` SHALL be generated from tool invocation facts
|
||||
- **AND** each summary item SHALL include tool name, success state, input summary, output summary, and evidence level
|
||||
- **AND** raw conversation history SHALL NOT be the only source of verifier evidence context
|
||||
|
||||
#### Scenario: tool trace summary preserves invocation references
|
||||
- **WHEN** the system prepares verifier inputs
|
||||
- **THEN** each summary item SHALL include a stable `trace_ref`
|
||||
- **AND** each summary item SHALL preserve `source_invocation_ids` for the tool invocation rows that contributed to the summary
|
||||
- **AND** each summary item SHOULD include query samples, retrieval layers, relevance levels, and source document labels when available
|
||||
|
||||
#### Scenario: only evidence-bearing tools included
|
||||
- **WHEN** the system generates `tool_trace_summary`
|
||||
- **THEN** it SHALL include only evidence-bearing tool invocations
|
||||
- **AND** non-evidence helper tools such as time or formatting tools SHALL be excluded by default
|
||||
|
||||
#### Scenario: failed evidence calls preserved as evidence gaps
|
||||
- **WHEN** an evidence-bearing tool invocation fails or returns no usable evidence
|
||||
- **THEN** the summary SHALL still include that invocation
|
||||
- **AND** it SHALL mark the entry as unsuccessful with an evidence level representing no evidence
|
||||
|
||||
#### Scenario: repeated tool calls may be compacted
|
||||
- **WHEN** repeated tool invocations concern the same tool, topic domain, and round
|
||||
- **THEN** the system MAY compact them into a merged summary entry
|
||||
- **AND** the merged entry SHALL preserve the first effective hit and the count of repeated, failed, or no-hit calls
|
||||
|
||||
#### Scenario: raw outputs not passed through in full
|
||||
- **WHEN** a tool invocation returns large raw content
|
||||
- **THEN** `tool_trace_summary` SHALL keep only a minimal evidence summary
|
||||
- **AND** the raw output SHALL NOT be passed through in full to the Verifier
|
||||
|
||||
#### Scenario: MessagesModelHook used only for noise reduction
|
||||
- **WHEN** a MessagesModelHook is used for the Verifier
|
||||
- **THEN** it MAY remove intermediate reasoning or irrelevant messages
|
||||
- **AND** it SHALL NOT be the primary source for assembling verifier business inputs
|
||||
|
||||
### Requirement: Verifier facts SHALL be auditable
|
||||
Verifier facts SHALL be linkable to the evidence summaries used during verification.
|
||||
|
||||
#### Scenario: facts_checked contains evidence refs
|
||||
- **WHEN** the Verifier emits `facts_checked`
|
||||
- **THEN** each fact SHALL include `evidence_refs`
|
||||
- **AND** each evidence ref SHALL point to an existing `tool_trace_summary.trace_ref`
|
||||
- **AND** each evidence ref SHALL preserve the relevant `source_invocation_ids` when available
|
||||
|
||||
#### Scenario: verifier evaluation persists traceability snapshot
|
||||
- **WHEN** the ChatService persists `verifier_evaluation`
|
||||
- **THEN** it SHALL include `traceability_version`
|
||||
- **AND** it SHALL include the `tool_trace_summary` snapshot used by the Verifier
|
||||
@@ -1,25 +1,27 @@
|
||||
package com.superbiz.agent.hook;
|
||||
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.messages.MessagesModelHook;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.messages.AgentCommand;
|
||||
import com.alibaba.cloud.ai.graph.RunnableConfig;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.HookPosition;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.HookPositions;
|
||||
import com.alibaba.cloud.ai.graph.RunnableConfig;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.messages.AgentCommand;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.messages.MessagesModelHook;
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.domain.entity.AgentStep;
|
||||
import com.superbiz.agent.repository.AgentStepRepository;
|
||||
import com.superbiz.agent.util.SessionContextHolder;
|
||||
import lombok.extern.slf4j.Slf4j;
|
||||
import org.springframework.ai.chat.messages.Message;
|
||||
import org.springframework.ai.chat.messages.AssistantMessage;
|
||||
import org.springframework.ai.chat.messages.UserMessage;
|
||||
import org.springframework.ai.chat.messages.Message;
|
||||
import org.springframework.ai.chat.messages.ToolResponseMessage;
|
||||
import org.springframework.ai.chat.messages.UserMessage;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.concurrent.ConcurrentHashMap;
|
||||
|
||||
/**
|
||||
* Agent 日志 Hook
|
||||
* 记录 Agent 的思考过程、消息流转 + 持久化 agent_step 到 DB
|
||||
* Persists per-agent model input/output snapshots into agent_step.
|
||||
*/
|
||||
@Slf4j
|
||||
@HookPositions({HookPosition.BEFORE_MODEL, HookPosition.AFTER_MODEL})
|
||||
@@ -27,11 +29,9 @@ public class AgentLoggingHook extends MessagesModelHook {
|
||||
|
||||
private final AgentStepRepository agentStepRepository;
|
||||
private final String agentName;
|
||||
private final ObjectMapper objectMapper = new ObjectMapper();
|
||||
|
||||
/** 每个 session 的步数计数器:sessionId → stepIndex */
|
||||
private final ConcurrentHashMap<String, Integer> stepCounters = new ConcurrentHashMap<>();
|
||||
|
||||
/** beforeModel → afterModel 中间状态:sessionId_stepIndex → {stepId, startTime} */
|
||||
private final ConcurrentHashMap<String, Map<String, Object>> pendingSteps = new ConcurrentHashMap<>();
|
||||
|
||||
public AgentLoggingHook(AgentStepRepository agentStepRepository, String agentName) {
|
||||
@@ -46,60 +46,44 @@ public class AgentLoggingHook extends MessagesModelHook {
|
||||
|
||||
@Override
|
||||
public AgentCommand beforeModel(List<Message> previousMessages, RunnableConfig config) {
|
||||
// 优先从 config.metadata 取 sessionId(线程安全),兜底 ThreadLocal
|
||||
String sessionId = config.metadata("sessionId")
|
||||
.map(Object::toString)
|
||||
.orElseGet(SessionContextHolder::getSessionId);
|
||||
|
||||
boolean hasSession = (sessionId != null);
|
||||
String sessionId = resolveSessionId(config);
|
||||
boolean hasSession = sessionId != null;
|
||||
|
||||
int stepIndex = 0;
|
||||
if (hasSession) {
|
||||
stepIndex = stepCounters.merge(sessionId, 0, (old, one) -> old + 1);
|
||||
stepIndex = stepCounters.merge(sessionId, 0, (oldValue, ignored) -> oldValue + 1);
|
||||
}
|
||||
|
||||
log.info("========================================");
|
||||
log.info("*** [Agent 思考] 第 {} 轮思考开始", (hasSession ? stepCounters.get(sessionId) : 0) + 1);
|
||||
log.info("*** [Agent 思考] 当前消息数量: {}", previousMessages.size());
|
||||
log.info("*** [AgentTrace] agent={}, phase=before_model, stepIndex={}", agentName, stepIndex);
|
||||
log.info("*** [AgentTrace] messageCount={}", previousMessages.size());
|
||||
|
||||
// 打印最后几条消息
|
||||
int lastN = Math.min(3, previousMessages.size());
|
||||
if (lastN > 0) {
|
||||
log.info("*** [Agent 思考] 最近 {} 条消息:", lastN);
|
||||
log.info("*** [AgentTrace] recentMessages={}", lastN);
|
||||
List<Message> recentMessages = previousMessages.subList(previousMessages.size() - lastN, previousMessages.size());
|
||||
for (int i = 0; i < recentMessages.size(); i++) {
|
||||
Message msg = recentMessages.get(i);
|
||||
String role = getMessageRole(msg);
|
||||
log.info(" [{}] 角色: {}, 类型: {}", i + 1, role, msg.getClass().getSimpleName());
|
||||
log.info(" [{}] role={}, type={}", i + 1, getMessageRole(msg), msg.getClass().getSimpleName());
|
||||
}
|
||||
}
|
||||
|
||||
log.info("*** [Agent 思考] 准备调用模型...");
|
||||
log.info("========================================");
|
||||
|
||||
// 持久化 agent_step(beforeModel:先创建,先记 model_input 摘要)
|
||||
if (sessionId != null) {
|
||||
try {
|
||||
String modelInputSummary = buildModelInputSummary(previousMessages);
|
||||
|
||||
AgentStep step = AgentStep.builder()
|
||||
.sessionId(sessionId)
|
||||
.stepIndex(stepIndex)
|
||||
.agentName(agentName)
|
||||
.modelInput(modelInputSummary)
|
||||
.modelInput(buildModelInputSummary(previousMessages))
|
||||
.build();
|
||||
AgentStep saved = agentStepRepository.save(step);
|
||||
|
||||
// 记录中间状态供 afterModel 使用
|
||||
pendingSteps.put(sessionId + "_" + stepIndex, Map.of(
|
||||
"stepId", saved.getId(),
|
||||
"startTime", System.currentTimeMillis()
|
||||
));
|
||||
|
||||
log.debug("agent_step 已创建: sessionId={}, stepIndex={}, id={}", sessionId, stepIndex, saved.getId());
|
||||
} catch (Exception e) {
|
||||
log.error("保存 agent_step 失败", e);
|
||||
// 不中断 Agent 执行
|
||||
log.error("Failed to persist agent_step before model", e);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -108,58 +92,38 @@ public class AgentLoggingHook extends MessagesModelHook {
|
||||
|
||||
@Override
|
||||
public AgentCommand afterModel(List<Message> previousMessages, RunnableConfig config) {
|
||||
String sessionId = SessionContextHolder.getSessionId();
|
||||
boolean hasSession = (sessionId != null);
|
||||
String sessionId = resolveSessionId(config);
|
||||
int stepIndex = sessionId == null ? 0 : stepCounters.getOrDefault(sessionId, 0);
|
||||
|
||||
log.info("========================================");
|
||||
log.info("*** [Agent 思考] 第 {} 轮思考完成", (hasSession ? stepCounters.getOrDefault(sessionId, 0) : 0));
|
||||
|
||||
// 查找最后一条 AssistantMessage(模型的回复)
|
||||
AssistantMessage lastAssistant = null;
|
||||
for (int i = previousMessages.size() - 1; i >= 0; i--) {
|
||||
if (previousMessages.get(i) instanceof AssistantMessage) {
|
||||
lastAssistant = (AssistantMessage) previousMessages.get(i);
|
||||
break;
|
||||
}
|
||||
}
|
||||
log.info("*** [AgentTrace] agent={}, phase=after_model, stepIndex={}", agentName, stepIndex);
|
||||
|
||||
AssistantMessage lastAssistant = findLastAssistant(previousMessages);
|
||||
boolean hasToolCall = false;
|
||||
|
||||
if (lastAssistant != null) {
|
||||
// 打印模型返回的文本内容
|
||||
String textContent = extractTextContent(lastAssistant);
|
||||
if (textContent != null && !textContent.isEmpty()) {
|
||||
log.info("*** [Agent 思考] 模型返回文本: {}",
|
||||
textContent.length() > 500
|
||||
? textContent.substring(0, 500) + "... (已截断,总长度: " + textContent.length() + ")"
|
||||
: textContent);
|
||||
log.info("*** [AgentTrace] text={}",
|
||||
textContent.length() > 500
|
||||
? textContent.substring(0, 500) + "... (len=" + textContent.length() + ")"
|
||||
: textContent);
|
||||
}
|
||||
|
||||
// 检查是否有工具调用
|
||||
if (lastAssistant.getToolCalls() != null && !lastAssistant.getToolCalls().isEmpty()) {
|
||||
hasToolCall = true;
|
||||
log.info("*** [Agent 思考] 模型决定调用 {} 个工具:",
|
||||
lastAssistant.getToolCalls().size());
|
||||
lastAssistant.getToolCalls().forEach(toolCall -> {
|
||||
log.info(" - 工具: {}, 参数: {}",
|
||||
toolCall.name(),
|
||||
toolCall.arguments());
|
||||
});
|
||||
log.info("*** [Agent 思考] 等待工具执行结果...");
|
||||
log.info("*** [AgentTrace] toolCalls={}", lastAssistant.getToolCalls().size());
|
||||
lastAssistant.getToolCalls().forEach(toolCall ->
|
||||
log.info(" - tool={}, arguments={}", toolCall.name(), toolCall.arguments()));
|
||||
} else {
|
||||
log.info("*** [Agent 思考] 模型决定不调用工具");
|
||||
log.info("*** [Agent 思考] 这是最终答案,准备返回给用户");
|
||||
log.info("*** [AgentTrace] no tool call");
|
||||
}
|
||||
}
|
||||
|
||||
log.info("========================================");
|
||||
|
||||
// 更新 agent_step(afterModel:补全 model_output、耗时等)
|
||||
if (sessionId != null) {
|
||||
int stepIndex = stepCounters.getOrDefault(sessionId, 0);
|
||||
String stepKey = sessionId + "_" + stepIndex;
|
||||
Map<String, Object> pending = pendingSteps.remove(stepKey);
|
||||
|
||||
if (pending != null) {
|
||||
try {
|
||||
Long stepId = (Long) pending.get("stepId");
|
||||
@@ -168,20 +132,12 @@ public class AgentLoggingHook extends MessagesModelHook {
|
||||
|
||||
AgentStep step = agentStepRepository.findById(stepId).orElse(null);
|
||||
if (step != null) {
|
||||
String thought = extractTextContent(lastAssistant);
|
||||
if (thought != null && thought.length() > 2000) {
|
||||
thought = thought.substring(0, 2000);
|
||||
}
|
||||
|
||||
step.setThought(thought);
|
||||
step.setThought(buildStoredThought(lastAssistant));
|
||||
step.setHasToolCall(hasToolCall);
|
||||
step.setDurationMs(durationMs);
|
||||
|
||||
if (lastAssistant != null) {
|
||||
String outputSummary = buildModelOutputSummary(lastAssistant);
|
||||
step.setModelOutput(outputSummary);
|
||||
|
||||
// 读取实际 token 用量(由 TokenTrackingChatModel 写入)
|
||||
step.setModelOutput(buildModelOutputSummary(lastAssistant));
|
||||
Integer tokenCount = TokenUsageHolder.get();
|
||||
if (tokenCount != null) {
|
||||
step.setTokenCount(tokenCount);
|
||||
@@ -189,24 +145,32 @@ public class AgentLoggingHook extends MessagesModelHook {
|
||||
}
|
||||
|
||||
agentStepRepository.save(step);
|
||||
log.debug("agent_step 已更新: sessionId={}, stepIndex={}, duration={}ms",
|
||||
sessionId, stepIndex, durationMs);
|
||||
}
|
||||
} catch (Exception e) {
|
||||
log.error("更新 agent_step 失败", e);
|
||||
log.error("Failed to update agent_step after model", e);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// 清理 token 上下文
|
||||
TokenUsageHolder.clear();
|
||||
|
||||
return new AgentCommand(previousMessages);
|
||||
}
|
||||
|
||||
/**
|
||||
* 构建模型输入摘要(前 N 条消息的 role + 截断内容)
|
||||
*/
|
||||
private String resolveSessionId(RunnableConfig config) {
|
||||
return config.metadata("sessionId")
|
||||
.map(Object::toString)
|
||||
.orElseGet(SessionContextHolder::getSessionId);
|
||||
}
|
||||
|
||||
private AssistantMessage findLastAssistant(List<Message> previousMessages) {
|
||||
for (int i = previousMessages.size() - 1; i >= 0; i--) {
|
||||
if (previousMessages.get(i) instanceof AssistantMessage assistantMessage) {
|
||||
return assistantMessage;
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
private String buildModelInputSummary(List<Message> messages) {
|
||||
StringBuilder sb = new StringBuilder();
|
||||
int maxMessages = Math.min(messages.size(), 5);
|
||||
@@ -226,25 +190,59 @@ public class AgentLoggingHook extends MessagesModelHook {
|
||||
return result;
|
||||
}
|
||||
|
||||
/**
|
||||
* 构建模型输出摘要
|
||||
*/
|
||||
private String buildStoredThought(AssistantMessage message) {
|
||||
String text = extractTextContent(message);
|
||||
if (text == null || text.isBlank()) {
|
||||
return text;
|
||||
}
|
||||
if (!"verifier".equals(agentName)) {
|
||||
return truncate(text, 2000);
|
||||
}
|
||||
return summarizeVerifierThought(text);
|
||||
}
|
||||
|
||||
private String summarizeVerifierThought(String verifierOutput) {
|
||||
try {
|
||||
JsonNode root = objectMapper.readTree(verifierOutput);
|
||||
int factCount = root.path("facts_checked").isArray() ? root.path("facts_checked").size() : 0;
|
||||
int tracedFactCount = 0;
|
||||
if (root.path("facts_checked").isArray()) {
|
||||
for (JsonNode factNode : root.path("facts_checked")) {
|
||||
if (factNode.path("evidence_refs").isArray() && factNode.path("evidence_refs").size() > 0) {
|
||||
tracedFactCount++;
|
||||
}
|
||||
}
|
||||
}
|
||||
return "verdict=%s, score=%s, critical_fact_count=%s, facts_checked=%d, traced_facts=%d".formatted(
|
||||
root.path("verdict").asText("UNKNOWN"),
|
||||
root.path("groundedness_score").asText("0.0"),
|
||||
root.path("critical_fact_count").asText("0"),
|
||||
factCount,
|
||||
tracedFactCount
|
||||
);
|
||||
} catch (Exception e) {
|
||||
return truncate(verifierOutput, 300);
|
||||
}
|
||||
}
|
||||
|
||||
private String buildModelOutputSummary(AssistantMessage message) {
|
||||
String text = extractTextContent(message);
|
||||
if (text == null) {
|
||||
text = "";
|
||||
}
|
||||
if (text.length() > 500) {
|
||||
text = text.substring(0, 500) + "...";
|
||||
}
|
||||
int maxTextLength = "verifier".equals(agentName) ? 4000 : 500;
|
||||
text = truncate(text, maxTextLength);
|
||||
|
||||
StringBuilder sb = new StringBuilder();
|
||||
sb.append("{\"text\":\"").append(escapeJson(text)).append("\"");
|
||||
if (message.getToolCalls() != null && !message.getToolCalls().isEmpty()) {
|
||||
sb.append(",\"toolCalls\":[");
|
||||
for (int i = 0; i < message.getToolCalls().size(); i++) {
|
||||
if (i > 0) sb.append(",");
|
||||
if (i > 0) {
|
||||
sb.append(",");
|
||||
}
|
||||
sb.append("{\"name\":\"").append(escapeJson(message.getToolCalls().get(i).name()))
|
||||
.append("\",\"arguments\":").append(message.getToolCalls().get(i).arguments()).append("}");
|
||||
.append("\",\"arguments\":").append(message.getToolCalls().get(i).arguments()).append("}");
|
||||
}
|
||||
sb.append("]");
|
||||
}
|
||||
@@ -253,7 +251,9 @@ public class AgentLoggingHook extends MessagesModelHook {
|
||||
}
|
||||
|
||||
private String escapeJson(String s) {
|
||||
if (s == null) return "";
|
||||
if (s == null) {
|
||||
return "";
|
||||
}
|
||||
return s.replace("\\", "\\\\")
|
||||
.replace("\"", "\\\"")
|
||||
.replace("\n", "\\n")
|
||||
@@ -261,96 +261,70 @@ public class AgentLoggingHook extends MessagesModelHook {
|
||||
.replace("\t", "\\t");
|
||||
}
|
||||
|
||||
/**
|
||||
* 提取 AssistantMessage 的文本内容
|
||||
*/
|
||||
private String truncate(String text, int maxLength) {
|
||||
if (text == null || text.length() <= maxLength) {
|
||||
return text;
|
||||
}
|
||||
return text.substring(0, maxLength) + "...";
|
||||
}
|
||||
|
||||
private String extractTextContent(AssistantMessage message) {
|
||||
if (message == null) return null;
|
||||
if (message == null) {
|
||||
return null;
|
||||
}
|
||||
try {
|
||||
// 方法 1: 反射获取 text 字段
|
||||
try {
|
||||
java.lang.reflect.Field textField = message.getClass().getDeclaredField("text");
|
||||
textField.setAccessible(true);
|
||||
Object value = textField.get(message);
|
||||
if (value != null) {
|
||||
log.debug("通过 text 字段提取成功");
|
||||
return value.toString();
|
||||
return message.getText();
|
||||
} catch (Exception ignore) {
|
||||
// Fallback below.
|
||||
}
|
||||
|
||||
for (String fieldName : List.of("text", "content")) {
|
||||
try {
|
||||
java.lang.reflect.Field field = message.getClass().getDeclaredField(fieldName);
|
||||
field.setAccessible(true);
|
||||
Object value = field.get(message);
|
||||
if (value != null) {
|
||||
return value.toString();
|
||||
}
|
||||
} catch (NoSuchFieldException ignore) {
|
||||
// continue
|
||||
}
|
||||
} catch (NoSuchFieldException e) {
|
||||
// 尝试下一种方法
|
||||
}
|
||||
|
||||
// 方法 2: 反射获取 content 字段
|
||||
try {
|
||||
java.lang.reflect.Field contentField = message.getClass().getDeclaredField("content");
|
||||
contentField.setAccessible(true);
|
||||
Object value = contentField.get(message);
|
||||
if (value != null) {
|
||||
log.debug("通过 content 字段提取成功");
|
||||
return value.toString();
|
||||
for (String methodName : List.of("getText", "getContent")) {
|
||||
try {
|
||||
java.lang.reflect.Method method = message.getClass().getMethod(methodName);
|
||||
Object value = method.invoke(message);
|
||||
if (value != null) {
|
||||
return value.toString();
|
||||
}
|
||||
} catch (NoSuchMethodException ignore) {
|
||||
// continue
|
||||
}
|
||||
} catch (NoSuchFieldException e) {
|
||||
// 尝试下一种方法
|
||||
}
|
||||
|
||||
// 方法 3: 调用 getText() 方法
|
||||
try {
|
||||
java.lang.reflect.Method getTextMethod = message.getClass().getMethod("getText");
|
||||
Object value = getTextMethod.invoke(message);
|
||||
if (value != null) {
|
||||
log.debug("通过 getText() 方法提取成功");
|
||||
return value.toString();
|
||||
}
|
||||
} catch (NoSuchMethodException e) {
|
||||
// 尝试下一种方法
|
||||
String fallback = message.toString();
|
||||
if (fallback != null && !fallback.startsWith("AssistantMessage@")) {
|
||||
return fallback;
|
||||
}
|
||||
|
||||
// 方法 4: 调用 getContent() 方法
|
||||
try {
|
||||
java.lang.reflect.Method getContentMethod = message.getClass().getMethod("getContent");
|
||||
Object value = getContentMethod.invoke(message);
|
||||
if (value != null) {
|
||||
log.debug("通过 getContent() 方法提取成功");
|
||||
return value.toString();
|
||||
}
|
||||
} catch (NoSuchMethodException e) {
|
||||
// 方法不存在
|
||||
}
|
||||
|
||||
// 方法 5: 打印类结构信息
|
||||
log.warn("无法提取 AssistantMessage 文本内容,打印类信息:");
|
||||
log.warn("类名: {}", message.getClass().getName());
|
||||
log.warn("字段列表:");
|
||||
for (java.lang.reflect.Field field : message.getClass().getDeclaredFields()) {
|
||||
log.warn(" - {}: {}", field.getName(), field.getType().getSimpleName());
|
||||
}
|
||||
|
||||
// 方法 6: toString() 兜底
|
||||
String toString = message.toString();
|
||||
if (toString != null && !toString.startsWith("AssistantMessage@")) {
|
||||
log.debug("通过 toString() 提取");
|
||||
return toString;
|
||||
}
|
||||
|
||||
return null;
|
||||
} catch (Exception e) {
|
||||
log.error("提取 AssistantMessage 文本内容时出错", e);
|
||||
log.error("Failed to extract AssistantMessage text", e);
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* 获取消息角色
|
||||
*/
|
||||
private String getMessageRole(Message message) {
|
||||
if (message instanceof UserMessage) {
|
||||
return "User(用户)";
|
||||
} else if (message instanceof AssistantMessage) {
|
||||
return "Assistant(模型)";
|
||||
} else if (message instanceof ToolResponseMessage) {
|
||||
return "Tool(工具返回)";
|
||||
} else {
|
||||
return message.getClass().getSimpleName();
|
||||
return "user";
|
||||
}
|
||||
if (message instanceof AssistantMessage) {
|
||||
return "assistant";
|
||||
}
|
||||
if (message instanceof ToolResponseMessage) {
|
||||
return "tool";
|
||||
}
|
||||
return message.getClass().getSimpleName();
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,105 @@
|
||||
package com.superbiz.agent.hook;
|
||||
|
||||
import com.alibaba.cloud.ai.graph.RunnableConfig;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.HookPosition;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.HookPositions;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.messages.AgentCommand;
|
||||
import com.alibaba.cloud.ai.graph.agent.hook.messages.MessagesModelHook;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.service.ToolTraceSummaryService;
|
||||
import com.superbiz.agent.util.SessionContextHolder;
|
||||
import com.superbiz.agent.util.VerifierContextHolder;
|
||||
import lombok.extern.slf4j.Slf4j;
|
||||
import org.springframework.ai.chat.messages.AssistantMessage;
|
||||
import org.springframework.ai.chat.messages.Message;
|
||||
import org.springframework.ai.chat.messages.UserMessage;
|
||||
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
/**
|
||||
* Replaces verifier history with an explicit structured payload.
|
||||
*/
|
||||
@Slf4j
|
||||
@HookPositions(HookPosition.BEFORE_MODEL)
|
||||
public class VerifierInputHook extends MessagesModelHook {
|
||||
|
||||
private final ToolTraceSummaryService toolTraceSummaryService;
|
||||
private final ObjectMapper objectMapper = new ObjectMapper();
|
||||
|
||||
public VerifierInputHook(ToolTraceSummaryService toolTraceSummaryService) {
|
||||
this.toolTraceSummaryService = toolTraceSummaryService;
|
||||
}
|
||||
|
||||
@Override
|
||||
public String getName() {
|
||||
return "verifier_input_hook";
|
||||
}
|
||||
|
||||
@Override
|
||||
public AgentCommand beforeModel(List<Message> previousMessages, RunnableConfig config) {
|
||||
try {
|
||||
String sessionId = config.metadata("sessionId")
|
||||
.map(Object::toString)
|
||||
.orElseGet(SessionContextHolder::getSessionId);
|
||||
String executorFinalAnswer = VerifierContextHolder.getExecutorFinalAnswer();
|
||||
if (executorFinalAnswer == null || executorFinalAnswer.isBlank()) {
|
||||
executorFinalAnswer = extractLastAssistantText(previousMessages);
|
||||
}
|
||||
|
||||
List<Map<String, Object>> toolTraceSummary =
|
||||
toolTraceSummaryService.buildVerifierTraceSummary(sessionId, executorFinalAnswer);
|
||||
VerifierContextHolder.setToolTraceSummary(toolTraceSummary);
|
||||
|
||||
Map<String, Object> verifierInput = new LinkedHashMap<>();
|
||||
verifierInput.put("original_query", VerifierContextHolder.getOriginalQuery());
|
||||
verifierInput.put("executor_final_answer", executorFinalAnswer);
|
||||
verifierInput.put("tool_trace_summary", toolTraceSummary);
|
||||
verifierInput.put("retry_context", VerifierContextHolder.getRetryContext());
|
||||
|
||||
String payload = objectMapper.writerWithDefaultPrettyPrinter().writeValueAsString(verifierInput);
|
||||
return new AgentCommand(List.of(new UserMessage(payload)));
|
||||
} catch (Exception e) {
|
||||
log.error("Failed to build verifier input, fallback to original messages", e);
|
||||
return new AgentCommand(previousMessages);
|
||||
}
|
||||
}
|
||||
|
||||
private String extractLastAssistantText(List<Message> previousMessages) {
|
||||
for (int i = previousMessages.size() - 1; i >= 0; i--) {
|
||||
if (previousMessages.get(i) instanceof AssistantMessage assistantMessage) {
|
||||
String text = extractTextContent(assistantMessage);
|
||||
if (text != null && !text.isBlank()) {
|
||||
return text;
|
||||
}
|
||||
}
|
||||
}
|
||||
return "";
|
||||
}
|
||||
|
||||
private String extractTextContent(AssistantMessage message) {
|
||||
try {
|
||||
try {
|
||||
return message.getText();
|
||||
} catch (Exception ignore) {
|
||||
// Fallback for older implementations.
|
||||
}
|
||||
|
||||
for (String methodName : List.of("getText", "getContent")) {
|
||||
try {
|
||||
var method = message.getClass().getMethod(methodName);
|
||||
Object value = method.invoke(message);
|
||||
if (value != null) {
|
||||
return value.toString();
|
||||
}
|
||||
} catch (NoSuchMethodException ignore) {
|
||||
// continue
|
||||
}
|
||||
}
|
||||
} catch (Exception e) {
|
||||
log.debug("Failed to extract verifier assistant text", e);
|
||||
}
|
||||
return message.toString();
|
||||
}
|
||||
}
|
||||
@@ -17,6 +17,11 @@ public interface ToolInvocationRepository extends JpaRepository<ToolInvocation,
|
||||
*/
|
||||
List<ToolInvocation> findBySessionId(String sessionId);
|
||||
|
||||
/**
|
||||
* 根据会话ID按创建顺序查询所有工具调用
|
||||
*/
|
||||
List<ToolInvocation> findBySessionIdOrderByIdAsc(String sessionId);
|
||||
|
||||
/**
|
||||
* 根据工具名查询所有调用
|
||||
*/
|
||||
|
||||
@@ -5,6 +5,8 @@ import com.alibaba.cloud.ai.graph.RunnableConfig;
|
||||
import com.alibaba.cloud.ai.graph.agent.ReactAgent;
|
||||
import com.alibaba.cloud.ai.graph.agent.flow.agent.SupervisorAgent;
|
||||
import com.alibaba.cloud.ai.graph.exception.GraphRunnerException;
|
||||
import com.fasterxml.jackson.databind.JsonNode;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.agent.tool.DateTimeTools;
|
||||
import com.superbiz.agent.agent.tool.InternalDocsTools;
|
||||
import com.superbiz.agent.agent.tool.QueryLogsTools;
|
||||
@@ -13,13 +15,14 @@ import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||
import com.superbiz.agent.hook.AgentLoggingHook;
|
||||
import com.superbiz.agent.hook.TokenTrackingChatModel;
|
||||
import com.superbiz.agent.hook.TokenUsageHolder;
|
||||
import com.superbiz.agent.hook.VerifierInputHook;
|
||||
import com.superbiz.agent.repository.AgentStepRepository;
|
||||
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
||||
import com.superbiz.agent.tool.LookupKnowledgeTool;
|
||||
import com.superbiz.agent.tool.RetrievedDocTracker;
|
||||
import com.superbiz.agent.util.QuestionComplexity;
|
||||
import com.superbiz.agent.util.SessionContextHolder;
|
||||
import com.superbiz.agent.service.KnowledgeDomainService;
|
||||
import com.superbiz.agent.util.VerifierContextHolder;
|
||||
|
||||
import jakarta.annotation.PostConstruct;
|
||||
import org.slf4j.Logger;
|
||||
@@ -29,11 +32,14 @@ import org.springframework.ai.chat.model.ChatModel;
|
||||
import org.springframework.ai.tool.ToolCallback;
|
||||
import org.springframework.ai.tool.ToolCallbackProvider;
|
||||
import org.springframework.beans.factory.annotation.Autowired;
|
||||
import org.springframework.beans.factory.annotation.Value;
|
||||
import org.springframework.core.io.ClassPathResource;
|
||||
import org.springframework.stereotype.Service;
|
||||
|
||||
import java.io.IOException;
|
||||
import java.nio.charset.StandardCharsets;
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
import java.util.Optional;
|
||||
@@ -47,6 +53,8 @@ import java.util.UUID;
|
||||
public class ChatService {
|
||||
|
||||
private static final Logger logger = LoggerFactory.getLogger(ChatService.class);
|
||||
private static final String LOW_CONFID_DISCLAIMER = "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。";
|
||||
private static final String DEGRADED_PREFIX = "当前无法基于已获取证据生成可靠结论,建议人工介入。";
|
||||
|
||||
/** 封装 answer + 后端生成的 sessionId,用于 feedback 关联 */
|
||||
public record ChatResult(String answer, String sessionId) {}
|
||||
@@ -87,9 +95,20 @@ public class ChatService {
|
||||
@Autowired
|
||||
private KnowledgeDomainService knowledgeDomainService;
|
||||
|
||||
@Autowired
|
||||
private ToolTraceSummaryService toolTraceSummaryService;
|
||||
|
||||
@Autowired
|
||||
private SelfEvaluationMergeService selfEvaluationMergeService;
|
||||
|
||||
@Value("${verifier.low-confidence-threshold:0.5}")
|
||||
private double verifierLowConfidenceThreshold;
|
||||
|
||||
/** 多 Agent Chat 的 Prompt */
|
||||
private String chatPlannerPrompt;
|
||||
private String chatExecutorPrompt;
|
||||
private String chatVerifierPrompt;
|
||||
private final ObjectMapper objectMapper = new ObjectMapper();
|
||||
|
||||
@PostConstruct
|
||||
public void init() {
|
||||
@@ -101,6 +120,9 @@ public class ChatService {
|
||||
chatExecutorPrompt = new String(
|
||||
new ClassPathResource("prompts/chat-executor-prompt.md").getInputStream().readAllBytes(),
|
||||
StandardCharsets.UTF_8);
|
||||
chatVerifierPrompt = new String(
|
||||
new ClassPathResource("prompts/chat-verifier-prompt.md").getInputStream().readAllBytes(),
|
||||
StandardCharsets.UTF_8);
|
||||
logger.info("Chat 多 Agent Prompts 加载成功");
|
||||
} catch (IOException e) {
|
||||
logger.error("加载 Chat Prompt 文件失败", e);
|
||||
@@ -341,34 +363,71 @@ public class ChatService {
|
||||
diagnosisSessionRepository.save(session);
|
||||
|
||||
SessionContextHolder.setSessionId(sessionId);
|
||||
VerifierContextHolder.setOriginalQuery(question);
|
||||
VerifierContextHolder.setRetryContext(null);
|
||||
VerifierContextHolder.setExecutorFinalAnswer(null);
|
||||
|
||||
try {
|
||||
ReactAgent planner = buildChatPlannerAgent(chatModel, toolCallbacks, history);
|
||||
ReactAgent executor = buildChatExecutorAgent(chatModel, toolCallbacks, history);
|
||||
|
||||
SupervisorAgent supervisor = SupervisorAgent.builder()
|
||||
.name("chat_supervisor")
|
||||
.description("负责调度 Planner 与 Executor 的多 Agent 控制器")
|
||||
.model(chatModel)
|
||||
.systemPrompt("你是一个智能任务调度器。分析用户问题,调用 Planner 拆解步骤,调用 Executor 执行各步骤。")
|
||||
.subAgents(List.of(planner, executor))
|
||||
VerifierDecision finalDecision = null;
|
||||
String retryContext = null;
|
||||
String answer = null;
|
||||
RunnableConfig config = RunnableConfig.builder()
|
||||
.addMetadata("sessionId", sessionId)
|
||||
.build();
|
||||
|
||||
Optional<OverAllState> stateOptional = supervisor.invoke(question);
|
||||
long duration = System.currentTimeMillis() - startTime;
|
||||
for (int round = 1; round <= 2; round++) {
|
||||
VerifierContextHolder.setRetryContext(retryContext);
|
||||
VerifierContextHolder.setToolTraceSummary(null);
|
||||
|
||||
String answer = null;
|
||||
if (stateOptional.isPresent()) {
|
||||
// 从 state 中提取 Executor 的最终输出
|
||||
OverAllState state = stateOptional.get();
|
||||
Optional<AssistantMessage> executorOutput = state.value("executor_feedback")
|
||||
.filter(AssistantMessage.class::isInstance)
|
||||
.map(AssistantMessage.class::cast);
|
||||
if (executorOutput.isPresent()) {
|
||||
answer = executorOutput.get().getText();
|
||||
ReactAgent planner = buildChatPlannerAgent(chatModel, history, retryContext);
|
||||
ReactAgent executor = buildChatExecutorAgent(chatModel, toolCallbacks, history, retryContext);
|
||||
ReactAgent verifier = buildChatVerifierAgent(chatModel);
|
||||
|
||||
SupervisorAgent supervisor = SupervisorAgent.builder()
|
||||
.name("chat_supervisor")
|
||||
.description("负责按单轮顺序调度 Planner、Executor、Verifier 的多 Agent 控制器")
|
||||
.model(chatModel)
|
||||
.systemPrompt(buildSupervisorPrompt(round))
|
||||
.subAgents(List.of(planner, executor, verifier))
|
||||
.build();
|
||||
|
||||
String plannerPlan = callAgent(planner, buildPlannerInput(question, retryContext), config);
|
||||
answer = callAgent(executor, buildExecutorInput(question, plannerPlan, retryContext), config);
|
||||
VerifierContextHolder.setExecutorFinalAnswer(answer);
|
||||
String verifierOutput = callAgent(verifier, "VERIFY", config);
|
||||
finalDecision = parseVerifierDecision(verifierOutput, round);
|
||||
|
||||
if (finalDecision == null) {
|
||||
finalDecision = buildVerifierFallbackDecision(round, "verifier_output 缺失或无法解析");
|
||||
answer = buildLowConfidenceOutput(answer, finalDecision);
|
||||
persistVerifierEvaluation(session, finalDecision, round);
|
||||
break;
|
||||
}
|
||||
|
||||
if ("PASS".equals(finalDecision.verdict())) {
|
||||
answer = answer == null || answer.isBlank() ? "抱歉,多 Agent 分析未能生成有效结论。" : answer;
|
||||
persistVerifierEvaluation(session, finalDecision, round);
|
||||
break;
|
||||
}
|
||||
|
||||
if ("REJECT".equals(finalDecision.verdict())) {
|
||||
answer = buildDegradedOutput(finalDecision);
|
||||
persistVerifierEvaluation(session, finalDecision, round);
|
||||
break;
|
||||
}
|
||||
|
||||
if (finalDecision.groundednessScore() >= verifierLowConfidenceThreshold || round == 2) {
|
||||
answer = buildLowConfidenceOutput(answer, finalDecision);
|
||||
persistVerifierEvaluation(session, finalDecision, round);
|
||||
break;
|
||||
}
|
||||
|
||||
retryContext = buildRetryContext(finalDecision);
|
||||
persistVerifierEvaluation(session, finalDecision, round);
|
||||
}
|
||||
|
||||
long duration = System.currentTimeMillis() - startTime;
|
||||
|
||||
if (answer == null || answer.isBlank()) {
|
||||
answer = "抱歉,多 Agent 分析未能生成有效结论。";
|
||||
}
|
||||
@@ -394,11 +453,12 @@ public class ChatService {
|
||||
} finally {
|
||||
retrievedDocTracker.clearSession(sessionId);
|
||||
SessionContextHolder.clear();
|
||||
VerifierContextHolder.clear();
|
||||
}
|
||||
}
|
||||
|
||||
private ReactAgent buildChatPlannerAgent(ChatModel chatModel, ToolCallback[] toolCallbacks,
|
||||
List<Map<String, String>> history) {
|
||||
private ReactAgent buildChatPlannerAgent(ChatModel chatModel, List<Map<String, String>> history,
|
||||
String retryContext) {
|
||||
StringBuilder prompt = new StringBuilder(chatPlannerPrompt);
|
||||
|
||||
// 注入 knowledge map
|
||||
@@ -414,6 +474,9 @@ public class ChatService {
|
||||
}
|
||||
prompt.append("--- 对话历史结束 ---\n");
|
||||
}
|
||||
if (retryContext != null && !retryContext.isBlank()) {
|
||||
prompt.append("\n\n--- 本轮补证据约束 ---\n").append(retryContext).append("\n");
|
||||
}
|
||||
return ReactAgent.builder()
|
||||
.name("chat_planner")
|
||||
.description("负责拆解问题、规划步骤")
|
||||
@@ -424,8 +487,20 @@ public class ChatService {
|
||||
.build();
|
||||
}
|
||||
|
||||
private ReactAgent buildChatVerifierAgent(ChatModel chatModel) {
|
||||
return ReactAgent.builder()
|
||||
.name("chat_verifier")
|
||||
.description("负责验证 Executor 答案的事实准确性")
|
||||
.model(chatModel)
|
||||
.systemPrompt(chatVerifierPrompt)
|
||||
.hooks(new AgentLoggingHook(agentStepRepository, "verifier"),
|
||||
new VerifierInputHook(toolTraceSummaryService))
|
||||
.outputKey("verifier_output")
|
||||
.build();
|
||||
}
|
||||
|
||||
private ReactAgent buildChatExecutorAgent(ChatModel chatModel, ToolCallback[] toolCallbacks,
|
||||
List<Map<String, String>> history) {
|
||||
List<Map<String, String>> history, String retryContext) {
|
||||
StringBuilder prompt = new StringBuilder(chatExecutorPrompt);
|
||||
if (!history.isEmpty()) {
|
||||
prompt.append("\n\n--- 对话历史 ---\n");
|
||||
@@ -434,6 +509,9 @@ public class ChatService {
|
||||
}
|
||||
prompt.append("--- 对话历史结束 ---\n");
|
||||
}
|
||||
if (retryContext != null && !retryContext.isBlank()) {
|
||||
prompt.append("\n\n--- 本轮补证据约束 ---\n").append(retryContext).append("\n");
|
||||
}
|
||||
return ReactAgent.builder()
|
||||
.name("chat_executor")
|
||||
.description("负责执行具体步骤并及时反馈")
|
||||
@@ -446,6 +524,316 @@ public class ChatService {
|
||||
.build();
|
||||
}
|
||||
|
||||
private String callAgent(ReactAgent agent, String input, RunnableConfig config) throws GraphRunnerException {
|
||||
return agent.call(input, config).getText();
|
||||
}
|
||||
|
||||
private String buildPlannerInput(String question, String retryContext) {
|
||||
if (retryContext == null || retryContext.isBlank()) {
|
||||
return question;
|
||||
}
|
||||
return question + "\n\n--- 补充约束 ---\n" + retryContext;
|
||||
}
|
||||
|
||||
private String buildExecutorInput(String question, String plannerPlan, String retryContext) {
|
||||
StringBuilder input = new StringBuilder(question);
|
||||
if (plannerPlan != null && !plannerPlan.isBlank()) {
|
||||
input.append("\n\n--- planner_plan ---\n").append(plannerPlan);
|
||||
}
|
||||
if (retryContext != null && !retryContext.isBlank()) {
|
||||
input.append("\n\n--- retry_context ---\n").append(retryContext);
|
||||
}
|
||||
return input.toString();
|
||||
}
|
||||
|
||||
private VerifierDecision parseVerifierDecision(String verifierOutput, int round) {
|
||||
if (verifierOutput == null || verifierOutput.isBlank()) {
|
||||
return null;
|
||||
}
|
||||
|
||||
try {
|
||||
JsonNode root = objectMapper.readTree(sanitizeJsonPayload(verifierOutput));
|
||||
List<Map<String, Object>> factsChecked = parseFactsChecked(root.path("facts_checked"));
|
||||
|
||||
return new VerifierDecision(
|
||||
root.path("verdict").asText("LOW_CONFID"),
|
||||
root.path("groundedness_score").asDouble(0.0),
|
||||
root.path("critical_fact_count").asInt(0),
|
||||
factsChecked,
|
||||
root.path("rationale").asText(""),
|
||||
round
|
||||
);
|
||||
} catch (Exception e) {
|
||||
logger.error("解析 verifier_output 失败: {}", verifierOutput, e);
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
private String sanitizeJsonPayload(String raw) {
|
||||
String trimmed = raw.trim();
|
||||
if (trimmed.startsWith("```")) {
|
||||
int firstNewline = trimmed.indexOf('\n');
|
||||
int lastFence = trimmed.lastIndexOf("```");
|
||||
if (firstNewline >= 0 && lastFence > firstNewline) {
|
||||
return trimmed.substring(firstNewline + 1, lastFence).trim();
|
||||
}
|
||||
}
|
||||
return trimmed;
|
||||
}
|
||||
|
||||
private List<Map<String, Object>> parseFactsChecked(JsonNode factsNode) {
|
||||
List<Map<String, Object>> factsChecked = new ArrayList<>();
|
||||
if (!factsNode.isArray()) {
|
||||
return factsChecked;
|
||||
}
|
||||
for (JsonNode factNode : factsNode) {
|
||||
Map<String, Object> fact = new LinkedHashMap<>();
|
||||
fact.put("fact", factNode.path("fact").asText(""));
|
||||
fact.put("is_critical", factNode.path("is_critical").asBoolean(false));
|
||||
fact.put("verification", factNode.path("verification").asText(""));
|
||||
fact.put("detail", factNode.path("detail").asText(""));
|
||||
fact.put("evidence_refs", parseEvidenceRefs(factNode.path("evidence_refs")));
|
||||
factsChecked.add(fact);
|
||||
}
|
||||
return factsChecked;
|
||||
}
|
||||
|
||||
private List<Map<String, Object>> parseEvidenceRefs(JsonNode evidenceRefsNode) {
|
||||
List<Map<String, Object>> evidenceRefs = new ArrayList<>();
|
||||
if (!evidenceRefsNode.isArray()) {
|
||||
return evidenceRefs;
|
||||
}
|
||||
for (JsonNode refNode : evidenceRefsNode) {
|
||||
Map<String, Object> evidenceRef = new LinkedHashMap<>();
|
||||
evidenceRef.put("trace_ref", refNode.path("trace_ref").asText(""));
|
||||
evidenceRef.put("tool_name", refNode.path("tool_name").asText(""));
|
||||
evidenceRef.put("topic_domain", refNode.path("topic_domain").asText(""));
|
||||
evidenceRef.put("note", refNode.path("note").asText(""));
|
||||
|
||||
List<Long> sourceInvocationIds = new ArrayList<>();
|
||||
JsonNode idsNode = refNode.path("source_invocation_ids");
|
||||
if (idsNode.isArray()) {
|
||||
for (JsonNode idNode : idsNode) {
|
||||
if (idNode.canConvertToLong()) {
|
||||
sourceInvocationIds.add(idNode.asLong());
|
||||
}
|
||||
}
|
||||
}
|
||||
evidenceRef.put("source_invocation_ids", sourceInvocationIds);
|
||||
evidenceRefs.add(evidenceRef);
|
||||
}
|
||||
return evidenceRefs;
|
||||
}
|
||||
|
||||
private VerifierDecision buildVerifierFallbackDecision(int round, String rationale) {
|
||||
return new VerifierDecision("LOW_CONFID", 0.0, 0, List.of(), rationale, round);
|
||||
}
|
||||
|
||||
private String buildSupervisorPrompt(int round) {
|
||||
return """
|
||||
你是一个多 Agent 调度器。每一轮必须严格按顺序完成以下动作:
|
||||
1. 先调用 chat_planner 生成执行计划
|
||||
2. 再调用 chat_executor 执行计划并形成最终答案
|
||||
3. 最后调用 chat_verifier 对 executor 最终答案做事实核查
|
||||
|
||||
规则:
|
||||
- 本轮只允许完成一次 Planner -> Executor -> Verifier 链路
|
||||
- Verifier 完成后立即停止,不要继续调用任何 Agent
|
||||
- 不要自己编造答案,最终用户输出由外层代码根据 verifier_output 决定
|
||||
- 当前是第 %d 轮,保持单轮内顺序稳定
|
||||
""".formatted(round);
|
||||
}
|
||||
|
||||
private String buildRoundInput(String question, String retryContext) {
|
||||
if (retryContext == null || retryContext.isBlank()) {
|
||||
return question;
|
||||
}
|
||||
return question + "\n\n--- 补充约束 ---\n" + retryContext;
|
||||
}
|
||||
|
||||
private String extractExecutorAnswer(Optional<OverAllState> stateOptional) {
|
||||
if (stateOptional.isEmpty()) {
|
||||
return null;
|
||||
}
|
||||
return stateOptional.get().value("executor_feedback")
|
||||
.filter(AssistantMessage.class::isInstance)
|
||||
.map(AssistantMessage.class::cast)
|
||||
.map(AssistantMessage::getText)
|
||||
.orElse(null);
|
||||
}
|
||||
|
||||
private VerifierDecision parseVerifierDecision(Optional<OverAllState> stateOptional, int round) {
|
||||
if (stateOptional.isEmpty()) {
|
||||
return null;
|
||||
}
|
||||
|
||||
Optional<AssistantMessage> verifierOutput = stateOptional.get().value("verifier_output")
|
||||
.filter(AssistantMessage.class::isInstance)
|
||||
.map(AssistantMessage.class::cast);
|
||||
|
||||
if (verifierOutput.isEmpty() || verifierOutput.get().getText() == null || verifierOutput.get().getText().isBlank()) {
|
||||
return null;
|
||||
}
|
||||
|
||||
try {
|
||||
JsonNode root = objectMapper.readTree(verifierOutput.get().getText());
|
||||
List<Map<String, Object>> factsChecked = parseFactsChecked(root.path("facts_checked"));
|
||||
|
||||
return new VerifierDecision(
|
||||
root.path("verdict").asText("LOW_CONFID"),
|
||||
root.path("groundedness_score").asDouble(0.0),
|
||||
root.path("critical_fact_count").asInt(0),
|
||||
factsChecked,
|
||||
root.path("rationale").asText(""),
|
||||
round
|
||||
);
|
||||
} catch (Exception e) {
|
||||
logger.error("解析 verifier_output 失败: {}", verifierOutput.get().getText(), e);
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
private void persistVerifierEvaluation(DiagnosisSession session, VerifierDecision decision, int round) {
|
||||
if (decision == null) {
|
||||
return;
|
||||
}
|
||||
Map<String, Object> verifierEvaluation = new LinkedHashMap<>();
|
||||
verifierEvaluation.put("verdict", decision.verdict());
|
||||
verifierEvaluation.put("groundedness_score", decision.groundednessScore());
|
||||
verifierEvaluation.put("critical_fact_count", decision.criticalFactCount());
|
||||
verifierEvaluation.put("facts_checked", decision.factsChecked());
|
||||
verifierEvaluation.put("rationale", decision.rationale());
|
||||
verifierEvaluation.put("round", round);
|
||||
verifierEvaluation.put("traceability_version", "v1");
|
||||
verifierEvaluation.put("tool_trace_summary",
|
||||
Optional.ofNullable(VerifierContextHolder.getToolTraceSummary()).orElse(List.of()));
|
||||
|
||||
String merged = selfEvaluationMergeService.mergeVerifierEvaluation(session.getSelfEvaluation(), verifierEvaluation);
|
||||
session.setSelfEvaluation(merged);
|
||||
diagnosisSessionRepository.save(session);
|
||||
}
|
||||
|
||||
private String buildRetryContext(VerifierDecision decision) {
|
||||
try {
|
||||
List<String> missingFacts = extractEvidenceGaps(decision);
|
||||
Map<String, Object> retryContext = new LinkedHashMap<>();
|
||||
retryContext.put("round", decision.round());
|
||||
retryContext.put("missing_evidence_facts", missingFacts);
|
||||
retryContext.put("instruction", "仅补充以上断言相关证据,不要重复已完成检索");
|
||||
return objectMapper.writeValueAsString(retryContext);
|
||||
} catch (Exception e) {
|
||||
logger.error("构造 retry_context 失败", e);
|
||||
return "{\"round\":1,\"missing_evidence_facts\":[],\"instruction\":\"仅补充缺失证据\"}";
|
||||
}
|
||||
}
|
||||
|
||||
private String buildLowConfidenceOutput(String executorAnswer, VerifierDecision decision) {
|
||||
StringBuilder output = new StringBuilder(LOW_CONFID_DISCLAIMER);
|
||||
output.append("\n\n").append(executorAnswer == null ? "" : executorAnswer);
|
||||
|
||||
List<String> gaps = extractEvidenceGaps(decision);
|
||||
if (!gaps.isEmpty()) {
|
||||
output.append("\n\n当前缺口:");
|
||||
for (String gap : gaps) {
|
||||
output.append("\n- ").append(gap);
|
||||
}
|
||||
}
|
||||
return output.toString();
|
||||
}
|
||||
|
||||
private String buildDegradedOutput(VerifierDecision decision) {
|
||||
StringBuilder output = new StringBuilder(DEGRADED_PREFIX);
|
||||
|
||||
List<String> confirmedFacts = extractConfirmedFacts(decision);
|
||||
List<String> gaps = extractEvidenceGaps(decision);
|
||||
List<String> suggestions = buildNextStepSuggestions(decision);
|
||||
|
||||
output.append("\n\n已确认信息:");
|
||||
if (confirmedFacts.isEmpty()) {
|
||||
output.append("\n- 暂无可稳定确认的信息");
|
||||
} else {
|
||||
for (String fact : confirmedFacts) {
|
||||
output.append("\n- ").append(fact);
|
||||
}
|
||||
}
|
||||
|
||||
output.append("\n\n证据缺口:");
|
||||
if (gaps.isEmpty()) {
|
||||
output.append("\n- 当前缺少足够的直接证据支撑核心结论");
|
||||
} else {
|
||||
for (String gap : gaps) {
|
||||
output.append("\n- ").append(gap);
|
||||
}
|
||||
}
|
||||
|
||||
output.append("\n\n建议下一步:");
|
||||
for (String suggestion : suggestions) {
|
||||
output.append("\n- ").append(suggestion);
|
||||
}
|
||||
return output.toString();
|
||||
}
|
||||
|
||||
private List<String> extractConfirmedFacts(VerifierDecision decision) {
|
||||
List<String> confirmedFacts = new ArrayList<>();
|
||||
for (Map<String, Object> fact : decision.factsChecked()) {
|
||||
String verification = String.valueOf(fact.get("verification"));
|
||||
boolean critical = Boolean.TRUE.equals(fact.get("is_critical"));
|
||||
if (critical && ("direct_evidence".equals(verification) || "indirect_support".equals(verification))) {
|
||||
confirmedFacts.add(String.valueOf(fact.get("fact")));
|
||||
}
|
||||
}
|
||||
return confirmedFacts;
|
||||
}
|
||||
|
||||
private List<String> extractEvidenceGaps(VerifierDecision decision) {
|
||||
List<String> gaps = new ArrayList<>();
|
||||
for (Map<String, Object> fact : decision.factsChecked()) {
|
||||
String verification = String.valueOf(fact.get("verification"));
|
||||
boolean critical = Boolean.TRUE.equals(fact.get("is_critical"));
|
||||
if (critical && ("no_evidence".equals(verification) || "contradicted".equals(verification))) {
|
||||
gaps.add(String.valueOf(fact.get("fact")) + ":" + String.valueOf(fact.get("detail")));
|
||||
}
|
||||
}
|
||||
if (gaps.isEmpty() && "LOW_CONFID".equals(decision.verdict())) {
|
||||
for (Map<String, Object> fact : decision.factsChecked()) {
|
||||
String verification = String.valueOf(fact.get("verification"));
|
||||
boolean critical = Boolean.TRUE.equals(fact.get("is_critical"));
|
||||
if (critical && "indirect_support".equals(verification)) {
|
||||
gaps.add(String.valueOf(fact.get("fact")) + ":缺少直接证据锚点");
|
||||
}
|
||||
}
|
||||
}
|
||||
return gaps;
|
||||
}
|
||||
|
||||
private List<String> buildNextStepSuggestions(VerifierDecision decision) {
|
||||
List<String> suggestions = new ArrayList<>();
|
||||
List<Map<String, Object>> toolSummary = toolTraceSummaryService.buildVerifierTraceSummary(SessionContextHolder.getSessionId(), null);
|
||||
boolean hasKnowledgeTool = toolSummary.stream().anyMatch(item -> "lookup_knowledge".equals(item.get("tool_name")));
|
||||
boolean hasFailedEvidence = toolSummary.stream().anyMatch(item -> !Boolean.TRUE.equals(item.get("success")));
|
||||
|
||||
if (!hasKnowledgeTool) {
|
||||
suggestions.add("补充知识库或业务文档检索结果,建立可引用的证据锚点");
|
||||
}
|
||||
if (hasFailedEvidence) {
|
||||
suggestions.add("优先重试失败的证据型查询,补齐日志、指标或知识库侧证据");
|
||||
}
|
||||
if (suggestions.isEmpty()) {
|
||||
suggestions.add("围绕上述证据缺口补充只读查询,再由人工复核最终结论");
|
||||
}
|
||||
return suggestions;
|
||||
}
|
||||
|
||||
private record VerifierDecision(
|
||||
String verdict,
|
||||
double groundednessScore,
|
||||
int criticalFactCount,
|
||||
List<Map<String, Object>> factsChecked,
|
||||
String rationale,
|
||||
int round
|
||||
) {
|
||||
}
|
||||
|
||||
/** 从 agent_step 汇总 token、步数等指标回填 diagnosis_session */
|
||||
private void backfillSessionMetrics(DiagnosisSession session) {
|
||||
try {
|
||||
|
||||
@@ -1,6 +1,5 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||
import com.superbiz.agent.domain.entity.ToolInvocation;
|
||||
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
||||
@@ -12,6 +11,7 @@ import org.springframework.scheduling.annotation.Async;
|
||||
import org.springframework.stereotype.Service;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
@@ -32,17 +32,19 @@ public class EvaluationService {
|
||||
@Autowired
|
||||
private ToolInvocationRepository toolInvocationRepository;
|
||||
|
||||
private final ObjectMapper objectMapper = new ObjectMapper();
|
||||
@Autowired
|
||||
private SelfEvaluationMergeService selfEvaluationMergeService;
|
||||
|
||||
@Async
|
||||
public void evaluate(String sessionId, String answer) {
|
||||
diagnosisSessionRepository.findBySessionId(sessionId).ifPresent(session -> {
|
||||
try {
|
||||
List<ToolInvocation> toolInvocations = toolInvocationRepository.findBySessionId(sessionId);
|
||||
String selfEvaluation = evaluateWithRules(session, toolInvocations);
|
||||
session.setSelfEvaluation(selfEvaluation);
|
||||
Map<String, Object> ruleEvaluation = evaluateWithRules(session, toolInvocations);
|
||||
String merged = selfEvaluationMergeService.mergeRuleEvaluation(session.getSelfEvaluation(), ruleEvaluation);
|
||||
session.setSelfEvaluation(merged);
|
||||
diagnosisSessionRepository.save(session);
|
||||
logger.info("证据评分已写入: sessionId={}, result={}", sessionId, selfEvaluation);
|
||||
logger.info("证据评分已写入: sessionId={}, result={}", sessionId, merged);
|
||||
} catch (Exception e) {
|
||||
logger.error("评分失败: sessionId={}", sessionId, e);
|
||||
}
|
||||
@@ -53,7 +55,7 @@ public class EvaluationService {
|
||||
// 规则引擎(事实层)
|
||||
// -------------------------------------------------------------------------
|
||||
|
||||
private String evaluateWithRules(DiagnosisSession session, List<ToolInvocation> invocations) {
|
||||
private Map<String, Object> evaluateWithRules(DiagnosisSession session, List<ToolInvocation> invocations) {
|
||||
List<Map<String, Object>> factors = new ArrayList<>();
|
||||
|
||||
if ("FAILED".equals(session.getStatus())) {
|
||||
@@ -117,19 +119,12 @@ public class EvaluationService {
|
||||
return Map.of("name", name, "delta", delta, "description", description);
|
||||
}
|
||||
|
||||
private String buildResult(int score, List<Map<String, Object>> factors) {
|
||||
try {
|
||||
Map<String, Object> result = Map.of(
|
||||
"evidence_score", score,
|
||||
"source", "rule",
|
||||
"factors", factors
|
||||
// llm_opinion: null ← 预留字段,LLM 观点叠加时在此处扩展
|
||||
);
|
||||
return objectMapper.writeValueAsString(result);
|
||||
} catch (Exception e) {
|
||||
logger.error("序列化评分结果失败", e);
|
||||
return "{\"evidence_score\":0,\"source\":\"rule\",\"factors\":[]}";
|
||||
}
|
||||
private Map<String, Object> buildResult(int score, List<Map<String, Object>> factors) {
|
||||
Map<String, Object> result = new LinkedHashMap<>();
|
||||
result.put("evidence_score", score);
|
||||
result.put("source", "rule");
|
||||
result.put("factors", factors);
|
||||
return result;
|
||||
}
|
||||
|
||||
// -------------------------------------------------------------------------
|
||||
|
||||
@@ -0,0 +1,71 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.fasterxml.jackson.core.type.TypeReference;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import lombok.extern.slf4j.Slf4j;
|
||||
import org.springframework.stereotype.Service;
|
||||
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.Map;
|
||||
|
||||
/**
|
||||
* 统一维护 diagnosis_session.self_evaluation JSON 容器。
|
||||
*/
|
||||
@Slf4j
|
||||
@Service
|
||||
public class SelfEvaluationMergeService {
|
||||
|
||||
private static final TypeReference<LinkedHashMap<String, Object>> MAP_TYPE = new TypeReference<>() {};
|
||||
|
||||
private final ObjectMapper objectMapper = new ObjectMapper();
|
||||
|
||||
public String mergeRuleEvaluation(String existingJson, Map<String, Object> ruleEvaluation) {
|
||||
return merge(existingJson, "rule_evaluation", ruleEvaluation);
|
||||
}
|
||||
|
||||
public String mergeVerifierEvaluation(String existingJson, Map<String, Object> verifierEvaluation) {
|
||||
return merge(existingJson, "verifier_evaluation", verifierEvaluation);
|
||||
}
|
||||
|
||||
private String merge(String existingJson, String key, Map<String, Object> value) {
|
||||
try {
|
||||
Map<String, Object> root = parseRoot(existingJson);
|
||||
root.put(key, value);
|
||||
return objectMapper.writeValueAsString(root);
|
||||
} catch (Exception e) {
|
||||
log.error("合并 self_evaluation 失败: key={}", key, e);
|
||||
return fallbackJson(key, value);
|
||||
}
|
||||
}
|
||||
|
||||
private Map<String, Object> parseRoot(String existingJson) throws Exception {
|
||||
if (existingJson == null || existingJson.isBlank()) {
|
||||
return new LinkedHashMap<>();
|
||||
}
|
||||
|
||||
Map<String, Object> parsed = objectMapper.readValue(existingJson, MAP_TYPE);
|
||||
if (parsed.containsKey("rule_evaluation") || parsed.containsKey("verifier_evaluation")) {
|
||||
return new LinkedHashMap<>(parsed);
|
||||
}
|
||||
|
||||
LinkedHashMap<String, Object> wrapped = new LinkedHashMap<>();
|
||||
if (parsed.containsKey("evidence_score") || parsed.containsKey("source") || parsed.containsKey("factors")) {
|
||||
wrapped.put("rule_evaluation", parsed);
|
||||
return wrapped;
|
||||
}
|
||||
if (parsed.containsKey("verdict") || parsed.containsKey("groundedness_score") || parsed.containsKey("facts_checked")) {
|
||||
wrapped.put("verifier_evaluation", parsed);
|
||||
return wrapped;
|
||||
}
|
||||
return new LinkedHashMap<>(parsed);
|
||||
}
|
||||
|
||||
private String fallbackJson(String key, Map<String, Object> value) {
|
||||
try {
|
||||
return objectMapper.writeValueAsString(Map.of(key, value));
|
||||
} catch (Exception ex) {
|
||||
log.error("兜底序列化 self_evaluation 失败", ex);
|
||||
return "{}";
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,307 @@
|
||||
package com.superbiz.agent.service;
|
||||
|
||||
import com.fasterxml.jackson.core.type.TypeReference;
|
||||
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||
import com.superbiz.agent.domain.entity.ToolInvocation;
|
||||
import com.superbiz.agent.repository.ToolInvocationRepository;
|
||||
import lombok.extern.slf4j.Slf4j;
|
||||
import org.springframework.stereotype.Service;
|
||||
|
||||
import java.util.ArrayList;
|
||||
import java.util.Comparator;
|
||||
import java.util.LinkedHashMap;
|
||||
import java.util.LinkedHashSet;
|
||||
import java.util.List;
|
||||
import java.util.Locale;
|
||||
import java.util.Map;
|
||||
import java.util.Set;
|
||||
|
||||
/**
|
||||
* Builds a verifier-facing evidence index from persisted tool invocations.
|
||||
*/
|
||||
@Slf4j
|
||||
@Service
|
||||
public class ToolTraceSummaryService {
|
||||
|
||||
private static final TypeReference<LinkedHashMap<String, Object>> MAP_TYPE = new TypeReference<>() {};
|
||||
private static final Set<String> EVIDENCE_TOOLS = Set.of("lookup_knowledge", "query_logs", "query_metrics", "query_order");
|
||||
|
||||
private final ToolInvocationRepository toolInvocationRepository;
|
||||
private final ObjectMapper objectMapper = new ObjectMapper();
|
||||
|
||||
public ToolTraceSummaryService(ToolInvocationRepository toolInvocationRepository) {
|
||||
this.toolInvocationRepository = toolInvocationRepository;
|
||||
}
|
||||
|
||||
public List<Map<String, Object>> buildVerifierTraceSummary(String sessionId, String executorFinalAnswer) {
|
||||
if (sessionId == null || sessionId.isBlank()) {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
List<ToolInvocation> invocations = toolInvocationRepository.findBySessionIdOrderByIdAsc(sessionId);
|
||||
if (invocations.isEmpty()) {
|
||||
return List.of();
|
||||
}
|
||||
|
||||
Map<String, AggregateEntry> grouped = new LinkedHashMap<>();
|
||||
for (ToolInvocation invocation : invocations) {
|
||||
if (!EVIDENCE_TOOLS.contains(invocation.getToolName())) {
|
||||
continue;
|
||||
}
|
||||
String topicDomain = extractTopicDomain(invocation);
|
||||
String key = invocation.getToolName() + "|" + topicDomain;
|
||||
AggregateEntry entry = grouped.computeIfAbsent(
|
||||
key,
|
||||
ignored -> new AggregateEntry(invocation.getToolName(), topicDomain));
|
||||
entry.absorb(invocation);
|
||||
}
|
||||
|
||||
List<AggregateEntry> rankedEntries = grouped.values().stream()
|
||||
.sorted(Comparator.comparingInt((AggregateEntry entry) -> entry.relevanceScore(executorFinalAnswer)).reversed())
|
||||
.limit(8)
|
||||
.toList();
|
||||
|
||||
List<Map<String, Object>> summaries = new ArrayList<>();
|
||||
for (int i = 0; i < rankedEntries.size(); i++) {
|
||||
summaries.add(rankedEntries.get(i).toSummary("trace-" + (i + 1)));
|
||||
}
|
||||
return summaries;
|
||||
}
|
||||
|
||||
private String extractTopicDomain(ToolInvocation invocation) {
|
||||
try {
|
||||
if (invocation.getRetrievalDetails() != null && !invocation.getRetrievalDetails().isBlank()) {
|
||||
Map<String, Object> details = objectMapper.readValue(invocation.getRetrievalDetails(), MAP_TYPE);
|
||||
Object domains = details.get("retrieved_domains");
|
||||
if (domains instanceof List<?> domainList && !domainList.isEmpty()) {
|
||||
return String.valueOf(domainList.get(0));
|
||||
}
|
||||
}
|
||||
} catch (Exception e) {
|
||||
log.debug("Failed to parse retrieved_domains, fallback to general", e);
|
||||
}
|
||||
return "general";
|
||||
}
|
||||
|
||||
private String extractInputSummary(ToolInvocation invocation) {
|
||||
String query = extractQuery(invocation);
|
||||
if (query != null && !query.isBlank()) {
|
||||
return "query=" + truncate(query, 120);
|
||||
}
|
||||
return invocation.getToolName() + " invoked";
|
||||
}
|
||||
|
||||
private String extractQuery(ToolInvocation invocation) {
|
||||
try {
|
||||
if (invocation.getInputParams() != null && !invocation.getInputParams().isBlank()) {
|
||||
Map<String, Object> params = objectMapper.readValue(invocation.getInputParams(), MAP_TYPE);
|
||||
Object query = params.get("query");
|
||||
if (query != null) {
|
||||
return String.valueOf(query);
|
||||
}
|
||||
}
|
||||
} catch (Exception e) {
|
||||
log.debug("Failed to parse invocation query", e);
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
private String extractOutputSummary(ToolInvocation invocation, String topicDomain) {
|
||||
if (!Boolean.TRUE.equals(invocation.getSuccess())) {
|
||||
if (invocation.getErrorMessage() != null && !invocation.getErrorMessage().isBlank()) {
|
||||
return "call failed: " + truncate(invocation.getErrorMessage(), 120);
|
||||
}
|
||||
return "no usable evidence returned";
|
||||
}
|
||||
|
||||
if ("lookup_knowledge".equals(invocation.getToolName())) {
|
||||
String relevance = invocation.getRelevanceLevel() != null ? invocation.getRelevanceLevel() : "UNKNOWN";
|
||||
String preview = invocation.getOutputPreview() != null && !invocation.getOutputPreview().isBlank()
|
||||
? truncate(invocation.getOutputPreview(), 160)
|
||||
: "no preview";
|
||||
return "matched domain=" + topicDomain + ", relevance=" + relevance + ", preview=" + preview;
|
||||
}
|
||||
|
||||
if (invocation.getOutputPreview() != null && !invocation.getOutputPreview().isBlank()) {
|
||||
return truncate(invocation.getOutputPreview(), 160);
|
||||
}
|
||||
return "evidence retrieved without preview";
|
||||
}
|
||||
|
||||
private String determineEvidenceLevel(ToolInvocation invocation) {
|
||||
if (!Boolean.TRUE.equals(invocation.getSuccess())) {
|
||||
return "none";
|
||||
}
|
||||
if ("PRECISE".equals(invocation.getRelevanceLevel()) || "HIGHLY_RELEVANT".equals(invocation.getRelevanceLevel())) {
|
||||
return "direct";
|
||||
}
|
||||
if ("REFERENCE".equals(invocation.getRelevanceLevel())) {
|
||||
return "indirect";
|
||||
}
|
||||
return "none";
|
||||
}
|
||||
|
||||
private List<String> extractStringList(Object value) {
|
||||
if (!(value instanceof List<?> list) || list.isEmpty()) {
|
||||
return List.of();
|
||||
}
|
||||
List<String> result = new ArrayList<>();
|
||||
for (Object item : list) {
|
||||
if (item != null) {
|
||||
result.add(String.valueOf(item));
|
||||
}
|
||||
}
|
||||
return result;
|
||||
}
|
||||
|
||||
private List<String> extractSourceDocuments(ToolInvocation invocation) {
|
||||
if (invocation.getRetrievalDetails() == null || invocation.getRetrievalDetails().isBlank()) {
|
||||
return List.of();
|
||||
}
|
||||
try {
|
||||
Map<String, Object> details = objectMapper.readValue(invocation.getRetrievalDetails(), MAP_TYPE);
|
||||
List<String> paths = extractStringList(details.get("l0_paths"));
|
||||
if (!paths.isEmpty()) {
|
||||
return paths;
|
||||
}
|
||||
List<String> titles = extractStringList(details.get("l0_titles"));
|
||||
if (!titles.isEmpty()) {
|
||||
return titles;
|
||||
}
|
||||
} catch (Exception e) {
|
||||
log.debug("Failed to parse source documents", e);
|
||||
}
|
||||
return List.of();
|
||||
}
|
||||
|
||||
private String truncate(String text, int maxLength) {
|
||||
if (text == null) {
|
||||
return "";
|
||||
}
|
||||
return text.length() <= maxLength ? text : text.substring(0, maxLength) + "...";
|
||||
}
|
||||
|
||||
private final class AggregateEntry {
|
||||
private final String toolName;
|
||||
private final String topicDomain;
|
||||
private String inputSummary;
|
||||
private String outputSummary;
|
||||
private boolean success;
|
||||
private String evidenceLevel = "none";
|
||||
private int invocationCount;
|
||||
private int failedCount;
|
||||
private int noHitCount;
|
||||
private final List<Long> sourceInvocationIds = new ArrayList<>();
|
||||
private final LinkedHashSet<String> querySamples = new LinkedHashSet<>();
|
||||
private final LinkedHashSet<String> retrievalLayers = new LinkedHashSet<>();
|
||||
private final LinkedHashSet<String> relevanceLevels = new LinkedHashSet<>();
|
||||
private final LinkedHashSet<String> sourceDocuments = new LinkedHashSet<>();
|
||||
|
||||
private AggregateEntry(String toolName, String topicDomain) {
|
||||
this.toolName = toolName;
|
||||
this.topicDomain = topicDomain;
|
||||
}
|
||||
|
||||
void absorb(ToolInvocation invocation) {
|
||||
invocationCount++;
|
||||
if (invocation.getId() != null) {
|
||||
sourceInvocationIds.add(invocation.getId());
|
||||
}
|
||||
|
||||
String query = extractQuery(invocation);
|
||||
if (query != null && !query.isBlank()) {
|
||||
querySamples.add(query);
|
||||
}
|
||||
if (invocation.getRetrievalLayer() != null && !invocation.getRetrievalLayer().isBlank()) {
|
||||
retrievalLayers.add(invocation.getRetrievalLayer());
|
||||
}
|
||||
if (invocation.getRelevanceLevel() != null && !invocation.getRelevanceLevel().isBlank()) {
|
||||
relevanceLevels.add(invocation.getRelevanceLevel());
|
||||
}
|
||||
sourceDocuments.addAll(extractSourceDocuments(invocation));
|
||||
|
||||
if (inputSummary == null || inputSummary.isBlank()) {
|
||||
inputSummary = extractInputSummary(invocation);
|
||||
}
|
||||
|
||||
boolean invocationSuccess = Boolean.TRUE.equals(invocation.getSuccess());
|
||||
if (!invocationSuccess) {
|
||||
failedCount++;
|
||||
return;
|
||||
}
|
||||
if (invocation.getDedupReason() != null) {
|
||||
noHitCount++;
|
||||
}
|
||||
|
||||
String invocationEvidenceLevel = determineEvidenceLevel(invocation);
|
||||
if (!success || evidenceRank(invocationEvidenceLevel) > evidenceRank(evidenceLevel)) {
|
||||
success = true;
|
||||
evidenceLevel = invocationEvidenceLevel;
|
||||
outputSummary = extractOutputSummary(invocation, topicDomain);
|
||||
}
|
||||
}
|
||||
|
||||
int relevanceScore(String answer) {
|
||||
int score = success ? 10 : 0;
|
||||
if ("direct".equals(evidenceLevel)) {
|
||||
score += 10;
|
||||
} else if ("indirect".equals(evidenceLevel)) {
|
||||
score += 5;
|
||||
}
|
||||
if (answer != null) {
|
||||
String normalized = answer.toLowerCase(Locale.ROOT);
|
||||
if (normalized.contains(topicDomain.toLowerCase(Locale.ROOT))) {
|
||||
score += 8;
|
||||
}
|
||||
if (normalized.contains(toolName.toLowerCase(Locale.ROOT))) {
|
||||
score += 3;
|
||||
}
|
||||
}
|
||||
return score;
|
||||
}
|
||||
|
||||
Map<String, Object> toSummary(String traceRef) {
|
||||
String mergedOutput = outputSummary == null ? "no summarized evidence" : outputSummary;
|
||||
if (invocationCount > 1) {
|
||||
StringBuilder builder = new StringBuilder(mergedOutput);
|
||||
builder.append(" (merged ").append(invocationCount).append(" invocations");
|
||||
if (failedCount > 0) {
|
||||
builder.append(", failed=").append(failedCount);
|
||||
}
|
||||
if (noHitCount > 0) {
|
||||
builder.append(", no_hit=").append(noHitCount);
|
||||
}
|
||||
builder.append(")");
|
||||
mergedOutput = builder.toString();
|
||||
}
|
||||
|
||||
Map<String, Object> summary = new LinkedHashMap<>();
|
||||
summary.put("trace_ref", traceRef);
|
||||
summary.put("tool_name", toolName);
|
||||
summary.put("success", success);
|
||||
summary.put("input_summary", inputSummary);
|
||||
summary.put("output_summary", mergedOutput);
|
||||
summary.put("evidence_level", evidenceLevel);
|
||||
summary.put("topic_domain", topicDomain);
|
||||
summary.put("source_invocation_ids", new ArrayList<>(sourceInvocationIds));
|
||||
summary.put("invocation_count", invocationCount);
|
||||
summary.put("failed_invocation_count", failedCount);
|
||||
summary.put("no_hit_invocation_count", noHitCount);
|
||||
summary.put("query_samples", new ArrayList<>(querySamples));
|
||||
summary.put("retrieval_layers", new ArrayList<>(retrievalLayers));
|
||||
summary.put("relevance_levels", new ArrayList<>(relevanceLevels));
|
||||
summary.put("source_documents", new ArrayList<>(sourceDocuments));
|
||||
return summary;
|
||||
}
|
||||
|
||||
private int evidenceRank(String level) {
|
||||
if ("direct".equals(level)) {
|
||||
return 2;
|
||||
}
|
||||
if ("indirect".equals(level)) {
|
||||
return 1;
|
||||
}
|
||||
return 0;
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,57 @@
|
||||
package com.superbiz.agent.util;
|
||||
|
||||
import java.util.List;
|
||||
import java.util.Map;
|
||||
|
||||
/**
|
||||
* Thread-local verifier context shared across one planner/executor/verifier round.
|
||||
*/
|
||||
public final class VerifierContextHolder {
|
||||
|
||||
private static final ThreadLocal<String> ORIGINAL_QUERY = new ThreadLocal<>();
|
||||
private static final ThreadLocal<String> RETRY_CONTEXT = new ThreadLocal<>();
|
||||
private static final ThreadLocal<String> EXECUTOR_FINAL_ANSWER = new ThreadLocal<>();
|
||||
private static final ThreadLocal<List<Map<String, Object>>> TOOL_TRACE_SUMMARY = new ThreadLocal<>();
|
||||
|
||||
private VerifierContextHolder() {
|
||||
}
|
||||
|
||||
public static void setOriginalQuery(String originalQuery) {
|
||||
ORIGINAL_QUERY.set(originalQuery);
|
||||
}
|
||||
|
||||
public static String getOriginalQuery() {
|
||||
return ORIGINAL_QUERY.get();
|
||||
}
|
||||
|
||||
public static void setRetryContext(String retryContext) {
|
||||
RETRY_CONTEXT.set(retryContext);
|
||||
}
|
||||
|
||||
public static String getRetryContext() {
|
||||
return RETRY_CONTEXT.get();
|
||||
}
|
||||
|
||||
public static void setExecutorFinalAnswer(String executorFinalAnswer) {
|
||||
EXECUTOR_FINAL_ANSWER.set(executorFinalAnswer);
|
||||
}
|
||||
|
||||
public static String getExecutorFinalAnswer() {
|
||||
return EXECUTOR_FINAL_ANSWER.get();
|
||||
}
|
||||
|
||||
public static void setToolTraceSummary(List<Map<String, Object>> toolTraceSummary) {
|
||||
TOOL_TRACE_SUMMARY.set(toolTraceSummary);
|
||||
}
|
||||
|
||||
public static List<Map<String, Object>> getToolTraceSummary() {
|
||||
return TOOL_TRACE_SUMMARY.get();
|
||||
}
|
||||
|
||||
public static void clear() {
|
||||
ORIGINAL_QUERY.remove();
|
||||
RETRY_CONTEXT.remove();
|
||||
EXECUTOR_FINAL_ANSWER.remove();
|
||||
TOOL_TRACE_SUMMARY.remove();
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,142 @@
|
||||
你是质量闸 verifier。你的任务是对 `executor_final_answer` 做一次基于现有证据的事实校验。
|
||||
|
||||
边界约束:
|
||||
- 不做新的检索
|
||||
- 不做超出输入证据的推理扩写
|
||||
- 不补充输入中不存在的新事实
|
||||
- 只输出一个合法 JSON 对象,不输出 Markdown,不输出代码块,不输出额外说明
|
||||
|
||||
## 输入字段
|
||||
|
||||
- `original_query`:用户原始问题
|
||||
- `executor_final_answer`:本轮 Executor 最终答案
|
||||
- `tool_trace_summary`:基于真实工具调用整理出的证据索引。每一项都带有:
|
||||
- `trace_ref`
|
||||
- `tool_name`
|
||||
- `topic_domain`
|
||||
- `source_invocation_ids`
|
||||
- `input_summary`
|
||||
- `output_summary`
|
||||
- `evidence_level`
|
||||
- `retry_context`:第二轮可选输入;若为空,按首轮处理
|
||||
|
||||
## 任务步骤
|
||||
|
||||
### 步骤一:提取关键事实
|
||||
优先提取并校验 `executor_final_answer` 里的全部实质性结论。关键事实至少包括:
|
||||
- 每一个根因结论
|
||||
- 每一个错误码、接口、组件归属或语义判断
|
||||
- 每一个明确的修复建议、参数建议、排查步骤
|
||||
- 每一个“证据来源陈述”
|
||||
|
||||
覆盖要求:
|
||||
- 不允许只抽取一个总括性事实替代整段答案
|
||||
- 如果答案给出多个根因,必须逐条拆成多个 `fact`
|
||||
- 如果答案给出多条修复建议,必须逐条拆成多个 `fact`
|
||||
- 只有寒暄、流程衔接语、与结论无关的话,才可以不纳入 `facts_checked`
|
||||
|
||||
### 步骤二:逐条校验事实
|
||||
每条事实必须输出:
|
||||
- `fact`
|
||||
- `is_critical`
|
||||
- `verification`
|
||||
- `detail`
|
||||
- `evidence_refs`
|
||||
|
||||
`verification` 只允许以下四个值:
|
||||
- `direct_evidence`
|
||||
- `indirect_support`
|
||||
- `no_evidence`
|
||||
- `contradicted`
|
||||
|
||||
### 步骤三:补齐 evidence_refs
|
||||
`evidence_refs` 必须是数组,数组元素必须引用 `tool_trace_summary` 中真实存在的证据项。每个元素包含:
|
||||
- `trace_ref`
|
||||
- `tool_name`
|
||||
- `topic_domain`
|
||||
- `source_invocation_ids`
|
||||
- `note`
|
||||
|
||||
规则:
|
||||
- 有证据支撑时,必须引用支撑该事实的证据项
|
||||
- `no_evidence` 并不等于不引用
|
||||
- 如果工具确实查过相关方向,但证据不够,仍应引用对应 trace,并在 `note` 里说明“不足以支撑”
|
||||
- 只有当确实找不到相关 trace 时,`evidence_refs` 才允许为空数组
|
||||
- 不允许编造不存在的 `trace_ref` 或 `source_invocation_ids`
|
||||
|
||||
### 步骤四:生成 verdict
|
||||
严格使用以下判定矩阵:
|
||||
1. 若任一关键事实(`is_critical=true`)为 `contradicted`
|
||||
- `verdict = "REJECT"`
|
||||
- `groundedness_score = 0.0`
|
||||
|
||||
2. 否则,若所有关键事实均为 `direct_evidence` 或 `indirect_support`
|
||||
且至少一条关键事实为 `direct_evidence`
|
||||
- `verdict = "PASS"`
|
||||
|
||||
3. 否则,若不存在 `contradicted`
|
||||
且存在关键事实为 `no_evidence`
|
||||
或所有关键事实都只有 `indirect_support`
|
||||
- `verdict = "LOW_CONFID"`
|
||||
|
||||
### 步骤五:计算 groundedness_score
|
||||
只统计 `is_critical=true` 的事实,映射如下:
|
||||
- `direct_evidence = 1.0`
|
||||
- `indirect_support = 0.6`
|
||||
- `no_evidence = 0.0`
|
||||
- `contradicted = 0.0`
|
||||
|
||||
规则:
|
||||
- 若任一关键事实为 `contradicted`,分数固定为 `0.0`
|
||||
- 否则对关键事实取平均值
|
||||
- 保留 2 位小数
|
||||
- 分数范围必须在 `[0.0, 1.0]`
|
||||
|
||||
### 步骤六:PASS 前覆盖性自检
|
||||
在输出 `PASS` 前,必须再次检查:
|
||||
- `facts_checked` 是否覆盖了 `executor_final_answer` 的全部实质性结论
|
||||
- 是否遗漏了单独出现的根因、修复建议、参数建议、排查步骤
|
||||
|
||||
如有明显遗漏,即使已校验事实都有证据,也不得输出 `PASS`。
|
||||
|
||||
### 步骤七:处理 retry_context
|
||||
若 `retry_context` 不为空:
|
||||
- 优先检查上一轮缺失证据点是否已补足
|
||||
- 不要扩展与缺口无关的新事实
|
||||
- 不要因为存在 `retry_context` 就自动降低 verdict
|
||||
|
||||
## 输出协议
|
||||
|
||||
必须输出且只能输出以下 JSON 结构:
|
||||
|
||||
{
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 0.8,
|
||||
"critical_fact_count": 2,
|
||||
"facts_checked": [
|
||||
{
|
||||
"fact": "ERR_TIMEOUT 表示请求超时",
|
||||
"is_critical": true,
|
||||
"verification": "direct_evidence",
|
||||
"detail": "知识库文档明确给出该错误码定义",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"trace_ref": "trace-1",
|
||||
"tool_name": "lookup_knowledge",
|
||||
"topic_domain": "api",
|
||||
"source_invocation_ids": [101, 104],
|
||||
"note": "trace-1 的文档摘要直接给出错误码定义"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"rationale": "所有关键事实均有支撑,且至少一条具有直接证据"
|
||||
}
|
||||
|
||||
输出要求:
|
||||
- `verdict` 只能是 `PASS` / `LOW_CONFID` / `REJECT`
|
||||
- `groundedness_score` 必须是 JSON number
|
||||
- `critical_fact_count` 必须等于 `facts_checked` 中 `is_critical=true` 的数量
|
||||
- `facts_checked` 可以为空数组,但字段不能缺失
|
||||
- 每条 `facts_checked[*]` 都必须包含 `evidence_refs`
|
||||
- 不得输出 schema 之外的字段
|
||||
Reference in New Issue
Block a user