Compare commits
39
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
83193bdf4a | ||
|
|
de56551fea | ||
|
|
da18fdf4e1 | ||
|
|
e1b8d1fb2c | ||
|
|
074d1aa5a9 | ||
|
|
f26d395650 | ||
|
|
ff0752a16c | ||
|
|
7844bcea40 | ||
|
|
e564863c43 | ||
|
|
d084202166 | ||
|
|
5b2fb985d9 | ||
|
|
b39a625e5b | ||
|
|
e9f1c48d34 | ||
|
|
3a7eee8af4 | ||
|
|
584639fa2a | ||
|
|
bdac35567c | ||
|
|
7ae9707a3b | ||
|
|
2f40536248 | ||
|
|
729cd3544a | ||
|
|
1e532fb851 | ||
|
|
38f781b157 | ||
|
|
5c369f3b6c | ||
|
|
f035538531 | ||
|
|
376ad0c241 | ||
|
|
ac1f831903 | ||
|
|
99d4f6f216 | ||
|
|
edb6153fd6 | ||
|
|
d0452184ee | ||
|
|
de5a5b09d9 | ||
|
|
e47f2dead0 | ||
|
|
49180abccf | ||
|
|
529f4ff43b | ||
|
|
75fa154a0a | ||
|
|
8300435a63 | ||
|
|
e20249c5d9 | ||
|
|
8fbc443f76 | ||
|
|
8ee7cc0b70 | ||
|
|
bc36248cd8 | ||
|
|
f8809cb7dd |
@@ -0,0 +1,259 @@
|
|||||||
|
---
|
||||||
|
name: essence
|
||||||
|
description: Invoke when a project is too large or you only want the core design insights. Extracts 1-2 standout design patterns with deep analysis, lens-guided perspectives, and migration examples. Not for full project analysis or quick lookups.
|
||||||
|
metadata:
|
||||||
|
version: "0.5.0"
|
||||||
|
---
|
||||||
|
|
||||||
|
# Essence: Extract Core Design Patterns
|
||||||
|
|
||||||
|
Prefix your first line with 🥷 inline, not as its own paragraph.
|
||||||
|
|
||||||
|
You are a jewel inspector. A project has thousands of files — your job is to find the one or two brilliant ideas worth stealing.
|
||||||
|
|
||||||
|
**This is NOT a lite version of `/explore`.** `/explore` reads the whole project and summarizes at the end. `/essence` goes deep on one thing and ignores everything else.
|
||||||
|
|
||||||
|
## Mode Selection
|
||||||
|
|
||||||
|
First, check whether an `/explore` result exists:
|
||||||
|
|
||||||
|
- `/explore` report exists → it already identified 2-3 core designs, default to **User-directed**. Ask the user which design to deep-dive, or whether to switch mode.
|
||||||
|
- No `/explore` result → this is an independent launch, default to **Auto-detect**.
|
||||||
|
|
||||||
|
Always confirm before proceeding:
|
||||||
|
|
||||||
|
| Mode | When | Entry |
|
||||||
|
|---|---|---|
|
||||||
|
| **User-directed** | Already have a design target from `/explore`, or know exactly which design to investigate | User tells you what to look for |
|
||||||
|
| **Auto-detect** | Independent launch, project is large, want the AI to find the standout design | You find the standout design |
|
||||||
|
| **Lens-guided** | "Analyze this from a [mechanical/intentional/evolution] perspective" | Apply a specific analytical lens |
|
||||||
|
|
||||||
|
### Lens definitions
|
||||||
|
|
||||||
|
| Lens | Core question | Guided behavior |
|
||||||
|
|---|---|---|
|
||||||
|
| **Mechanical** (default) | How does it work? | Read source code, trace call chains, examine interfaces |
|
||||||
|
| **Intentional** | Why this way? | Read design docs/RFCs/PRs, extract decision rationale and tradeoffs |
|
||||||
|
| **Evolution** | How did it get here? | Read git history/changelog, compare before/after, identify migration drivers |
|
||||||
|
|
||||||
|
A lens shapes which sources to read and how to frame the output, but does not add separate phases.
|
||||||
|
|
||||||
|
### Auto-detect signals
|
||||||
|
|
||||||
|
A design is "essence" if it passes 2 or more of these signals:
|
||||||
|
|
||||||
|
| Signal | Evidence |
|
||||||
|
|---|---|
|
||||||
|
| README highlights it prominently | "Built on a plugin architecture" as a headline feature |
|
||||||
|
| Has standalone architecture docs | ARCHITECTURE.md, docs/design/, blog post by author |
|
||||||
|
| Heavily discussed in Issues/PRs | Design decisions debated by community |
|
||||||
|
| Unique among similar projects | Competitors don't do it this way |
|
||||||
|
| Rich design comments in code | JSDoc/TSDoc explaining why, not what |
|
||||||
|
| Cross-module contract | A type, interface, or protocol imported across module boundaries (not just files). Go: most-implemented interface. Python: most-subclassed abstract base. Rust: most-implemented trait. These define subsystem relationships. |
|
||||||
|
| File size anomaly | One file is disproportionately large or small for its responsibility — signals non-trivial logic |
|
||||||
|
| Dedicated test coverage | Tests specifically validate this design's behavior, not just happy paths |
|
||||||
|
|
||||||
|
**"Clean code" is NOT a signal.** A well-written utility function is not essence. An architecture decision that shapes the entire project is.
|
||||||
|
|
||||||
|
If no design passes 2+ signals, tell the user: "This project has no standout design. Try `/explore` for a full analysis instead."
|
||||||
|
|
||||||
|
## Phase 1: Locate
|
||||||
|
|
||||||
|
**User-directed mode:**
|
||||||
|
- Go directly to the directory or file the user names.
|
||||||
|
- If the directory doesn't exist, stop and tell the user. Do NOT invent an alternative.
|
||||||
|
|
||||||
|
**Auto-detect mode:**
|
||||||
|
- Scan README, AGENTS.md, and top-level docs for architecture claims.
|
||||||
|
- Identify 1-2 standout design directions.
|
||||||
|
- Present to the user: "The standout designs appear to be: A) {design A}, B) {design B}. Which should we dive into?"
|
||||||
|
- If user doesn't choose, pick the strongest one and state why.
|
||||||
|
|
||||||
|
**Lens-guided mode:**
|
||||||
|
- Confirm the lens with the user (Mechanical/Intentional/Evolution).
|
||||||
|
- Frame the search in terms of the lens.
|
||||||
|
- Example: "You want the Mechanical view — I'll trace the core implementation and extract the pattern."
|
||||||
|
|
||||||
|
**Output:** 1-2 design directions to analyze + lens confirmation.
|
||||||
|
|
||||||
|
**Stall signal:** Cannot identify any standout design → the project may be a conventional CRUD app or wrapper. Stop and recommend `/explore` or a different project.
|
||||||
|
|
||||||
|
## Phase 2: Deep Dive
|
||||||
|
|
||||||
|
Read the core files related to the chosen design. Maximum 10 files. Let the lens guide source selection: Mechanical → source code and type definitions; Intentional → design docs, RFCs, PR discussions; Evolution → git history, changelog, migration guides.
|
||||||
|
|
||||||
|
**For each file:**
|
||||||
|
- What role does it play in this design?
|
||||||
|
- What interfaces does it expose?
|
||||||
|
- How does it connect to other parts of the system?
|
||||||
|
|
||||||
|
**Trace the call chain:**
|
||||||
|
- Start from the entry point that uses this design.
|
||||||
|
- Follow the flow until you understand the full pattern.
|
||||||
|
- Stop when you hit boilerplate, config, or test files.
|
||||||
|
|
||||||
|
**Output:** Core file list (≤10) + call chain + lens-specific annotations.
|
||||||
|
|
||||||
|
**Stall signal:** The design spans more than 10 files and you can't find the boundary → the design is probably the project's core architecture. Switch to `/explore` for a full analysis instead.
|
||||||
|
|
||||||
|
## Phase 3: Extract Pattern
|
||||||
|
|
||||||
|
Analyze the design at a higher level. Let the lens shape the analysis angle:
|
||||||
|
- **Mechanical** → emphasize structure, interfaces, data flow — produce a pattern diagram + interface contracts
|
||||||
|
- **Intentional** → emphasize decision rationale, tradeoffs — produce a decision record (context → options → rationale)
|
||||||
|
- **Evolution** → emphasize before/after comparison, migration drivers — produce a timeline + catalyst events
|
||||||
|
|
||||||
|
**Universal analysis dimensions** (all lenses):
|
||||||
|
|
||||||
|
- **Problem:** What specific problem does this design solve? What was the pain before?
|
||||||
|
- **Pattern:** What's the name of this pattern? (Named: MVC, Observer, Plugin, Middleware. Custom: describe it in one sentence.)
|
||||||
|
- **Alternatives:** What simpler or more complex approaches could solve the same problem?
|
||||||
|
- **Tradeoffs:** Why did the author choose this? What does it give up?
|
||||||
|
- **Evidence:** What in the code proves this analysis is correct? (Specific files, functions, comments.)
|
||||||
|
|
||||||
|
**Output:** Design pattern card (lens-framed).
|
||||||
|
|
||||||
|
**Stall signal:** Cannot explain why the author chose this design over alternatives → read commit messages and PR discussions for design rationale. If unavailable, state "author's reasoning unknown" in the report.
|
||||||
|
|
||||||
|
## Phase 4: Migrate
|
||||||
|
|
||||||
|
Make the learning actionable. Let the lens tailor the output:
|
||||||
|
- **Mechanical** → copy-paste code skeleton (≤20 lines with TODOs)
|
||||||
|
- **Intentional** → decision framework (checklist for evaluating tradeoffs)
|
||||||
|
- **Evolution** → migration path (step-by-step refactor plan)
|
||||||
|
|
||||||
|
**Universal deliverables** (all lenses):
|
||||||
|
|
||||||
|
- **Can you use this?** Is the design applicable to the user's own projects? If not, why?
|
||||||
|
- **Steal-it example:** A simplified version (under 20 lines) that captures the core idea. Not production code — a teaching example.
|
||||||
|
- **Pitfalls:** What context does this design depend on? What would break if you copy it blindly?
|
||||||
|
|
||||||
|
**Output:** Migration example + pitfall list (lens-tailored).
|
||||||
|
|
||||||
|
**Stall signal:** The design depends on framework internals, language features, or ecosystem the user doesn't have → explain the core idea abstractly instead of providing code.
|
||||||
|
|
||||||
|
## Phase 5: Self-review
|
||||||
|
|
||||||
|
Check the report is honest:
|
||||||
|
|
||||||
|
**All modes:**
|
||||||
|
- [ ] The design is real (not inferred, not imagined). Evidence: specific files cited.
|
||||||
|
- [ ] The analysis is deep enough that you could explain it out loud.
|
||||||
|
- [ ] The migration example captures the core idea, not surface syntax.
|
||||||
|
- [ ] Pitfalls are specific, not vague ("needs X version" not "may not work everywhere").
|
||||||
|
|
||||||
|
**Stall signals (any one → return to relevant phase):**
|
||||||
|
- Cannot name a file that proves the pattern → back to Phase 2
|
||||||
|
- Cannot explain why it's better than alternatives → back to Phase 3
|
||||||
|
- Migration example is over 20 lines → simplify, back to Phase 4
|
||||||
|
- Lens-specific check failed (e.g., Mechanical missing end-to-end call chain, Intentional missing decision rationale, Evolution missing timeline) → back to relevant phase
|
||||||
|
|
||||||
|
**Output:** Essence report with lens annotation.
|
||||||
|
|
||||||
|
## Optional: HTML Card
|
||||||
|
|
||||||
|
**Only when the user explicitly requests it.**
|
||||||
|
|
||||||
|
Generate an HTML visualization card as a shareable deliverable.
|
||||||
|
|
||||||
|
### HTML Card Structure (Glassmorphism 2.0 - Essence Variant)
|
||||||
|
|
||||||
|
```html
|
||||||
|
<!DOCTYPE html>
|
||||||
|
<html lang="zh-CN">
|
||||||
|
<head>
|
||||||
|
<meta charset="UTF-8">
|
||||||
|
<title>{Project Name} - Essence Report</title>
|
||||||
|
<script src="https://cdn.tailwindcss.com"></script>
|
||||||
|
<script src="https://cdn.jsdelivr.net/npm/mermaid/dist/mermaid.min.js"></script>
|
||||||
|
<style>
|
||||||
|
/* Same glassmorphism styles as /explore */
|
||||||
|
:root { --glass-bg: rgba(255,255,255,0.4); --primary: #8b5cf6; }
|
||||||
|
[data-theme="dark"] { --glass-bg: rgba(15,23,42,0.6); --primary: #a78bfa; }
|
||||||
|
.glass-panel { backdrop-filter: blur(12px); border-radius: 1rem; }
|
||||||
|
.pattern-diagram { font-family: monospace; background: rgba(0,0,0,0.03); }
|
||||||
|
</style>
|
||||||
|
</head>
|
||||||
|
<body class="p-8">
|
||||||
|
<nav class="fixed top-4 left-1/2 -translate-x-1/2 w-[90%] max-w-4xl glass-panel z-50 px-6 py-3">
|
||||||
|
<span class="font-bold text-xl">💎 {Project Name} 精华</span>
|
||||||
|
<span class="text-sm opacity-70">Lens: {lens} | Pattern: {pattern_name}</span>
|
||||||
|
</nav>
|
||||||
|
|
||||||
|
<main class="max-w-4xl mx-auto mt-24 space-y-6">
|
||||||
|
<section class="glass-panel p-6">
|
||||||
|
<h2 class="text-xl font-bold mb-4">🎯 Design Analyzed</h2>
|
||||||
|
<p>{one-line description}</p>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="glass-panel p-6">
|
||||||
|
<h2 class="text-xl font-bold mb-4">🔷 Pattern ({lens})</h2>
|
||||||
|
<!-- Lens-framed pattern card -->
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="glass-panel p-6">
|
||||||
|
<h2 class="text-xl font-bold mb-4">🔗 Call Chain</h2>
|
||||||
|
<pre class="mermaid">{diagram}</pre>
|
||||||
|
</section>
|
||||||
|
|
||||||
|
<section class="glass-panel p-6">
|
||||||
|
<h2 class="text-xl font-bold mb-4">📦 Migration Example</h2>
|
||||||
|
<pre class="pattern-diagram"><code>{code_example}</code></pre>
|
||||||
|
<p class="text-sm opacity-70 mt-2">Pitfalls: {pitfalls}</p>
|
||||||
|
</section>
|
||||||
|
</main>
|
||||||
|
|
||||||
|
<script>mermaid.initialize({ startOnLoad: true });</script>
|
||||||
|
</body>
|
||||||
|
</html>
|
||||||
|
```
|
||||||
|
|
||||||
|
### Output Format
|
||||||
|
|
||||||
|
```markdown
|
||||||
|
### HTML Card Generated
|
||||||
|
|
||||||
|
- **Path:** `outputs/{project}-essence.html`
|
||||||
|
- **Theme:** {modern/ink}
|
||||||
|
- **Accent Color:** Purple (essence = jewel)
|
||||||
|
```
|
||||||
|
|
||||||
|
**When to skip:** Skip HTML generation unless the user requests it or the analysis is production-critical. When HTML generation fails, deliver a plain-text report instead.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Hard Rules
|
||||||
|
|
||||||
|
- **No code evidence = no conclusion.** Every claim about a design must cite a specific file, function, or comment.
|
||||||
|
- **Under 20 lines for migration examples.** If you can't explain the idea in 20 lines, you don't understand it well enough.
|
||||||
|
- **Stop after the report.** Do not modify the user's project or the target project.
|
||||||
|
- **HTML is optional.** Do not block analysis on HTML generation.
|
||||||
|
|
||||||
|
## Gotchas
|
||||||
|
|
||||||
|
| What happened | Rule |
|
||||||
|
|---|---|
|
||||||
|
| 提取的"精华"是 AI 脑补的 | 必须有代码证据(文件 + 行号),不写空泛结论 |
|
||||||
|
| 用户指定方向但该模块不存在 | 停止并告知用户,不编造替代方向 |
|
||||||
|
| 项目没有 standout 设计(胶水代码) | 标记"无可提取精华",建议改用 `/explore` |
|
||||||
|
| Phase 4 迁移示例超过 20 行 | 简化到核心思路,不是复制生产代码 |
|
||||||
|
| 分析了一个小工具函数 | 工具函数不是设计。设计影响整个架构,工具只解决一个问题 |
|
||||||
|
| 从 commit message 推断作者意图但没有代码佐证 | Commit message 是辅助证据,必须有代码结构本身的支持 |
|
||||||
|
| 透镜模式选错导致输出不符预期 | Phase 1 先确认透镜,Mechanical 读代码、Intentional 读文档、Evolution 读历史 |
|
||||||
|
| 透镜分析流于表面 | 每个透镜有特定输出格式:Mechanical→图 + 接口,Intentional→决策记录,Evolution→时间线 |
|
||||||
|
| HTML 卡片生成失败 | 降级到纯文本报告,不阻塞分析交付 |
|
||||||
|
|
||||||
|
## Outcome
|
||||||
|
|
||||||
|
```
|
||||||
|
Essence Report: {project name}
|
||||||
|
Lens: mechanical / intentional / evolution
|
||||||
|
Design analyzed: {one-line description}
|
||||||
|
Files examined: {count}
|
||||||
|
Pattern: {pattern name or custom description}
|
||||||
|
Migration: {steal-it example, ≤20 lines}
|
||||||
|
HTML generated: yes / no
|
||||||
|
Status: complete
|
||||||
|
```
|
||||||
|
|
||||||
|
After the report, stop. No modifications. No follow-ups.
|
||||||
@@ -0,0 +1,79 @@
|
|||||||
|
# Essence Detection Signals
|
||||||
|
|
||||||
|
How to identify the standout design in a project when the user doesn't specify a direction.
|
||||||
|
|
||||||
|
## Signal Strength
|
||||||
|
|
||||||
|
A design passes the "essence" threshold if it scores 2+ signals.
|
||||||
|
|
||||||
|
### Strong Signals (score = 1 each)
|
||||||
|
|
||||||
|
| Signal | How to detect | Example |
|
||||||
|
|---|---|---|
|
||||||
|
| **README headline** | Project name is followed by a design claim | "Vite — Next generation frontend tooling with **ESM-first architecture**" |
|
||||||
|
| **Architecture docs** | Standalone design document exists | `ARCHITECTURE.md`, `docs/design/`, `docs/architecture/` |
|
||||||
|
| **Official blog post** | Author wrote about the design on their blog | tw93.fun, Vite blog, React blog posts |
|
||||||
|
| **Community discussion** | Issues/PRs debate the design decision | "Why we chose X over Y" discussions with many comments |
|
||||||
|
| **Rich code comments** | JSDoc/TSDoc explaining WHY, not WHAT | "We use this pattern because..." with detailed reasoning |
|
||||||
|
|
||||||
|
### Objective Signals (score = 1 each, no subjective judgment needed)
|
||||||
|
|
||||||
|
| Signal | How to detect | Example |
|
||||||
|
|---|---|---|
|
||||||
|
| **Cross-module contract** | A type, interface, or protocol imported across module boundaries (not just files). Go: most-implemented interface. Python: most-subclassed abstract base. Rust: most-implemented trait. | `Plugin` interface implemented by 8 subsystems, each in its own package |
|
||||||
|
| **File size anomaly** | One file's line count is ≥3× the median for its category (handlers, utils, etc.) | Average handler: 50 lines. One handler: 800 lines with state machine logic |
|
||||||
|
| **Dedicated test coverage** | Tests exist specifically for this design's edge cases, not just happy paths | `plugin.test.ts` tests plugin resolution, fallback, lifecycle — not just "it loads" |
|
||||||
|
|
||||||
|
### Weak Signals (score = 0.5 each)
|
||||||
|
|
||||||
|
| Signal | How to detect | Example |
|
||||||
|
|---|---|---|
|
||||||
|
| **Unique among competitors** | Same category, different architecture | Next.js uses SSR, Remix uses nested routes — that difference IS the essence |
|
||||||
|
| **Most-starred files** | GitHub shows stars/bookmarks on specific files | "This file has 200+ stars on GitHub" |
|
||||||
|
| **Core algorithm** | One file contains non-trivial logic that drives the project | Diff algorithm, compiler pass, state machine |
|
||||||
|
| **API design** | The public API is notably elegant or unusual | `create()` returns a builder chain, not an object |
|
||||||
|
|
||||||
|
## Not Signals
|
||||||
|
|
||||||
|
These do NOT count as essence:
|
||||||
|
|
||||||
|
- "Clean code" or "well organized" — that's quality, not design
|
||||||
|
- "Uses TypeScript" — that's a language choice, not architecture
|
||||||
|
- "Has good tests" — that's engineering discipline, not design
|
||||||
|
- "Many stars on the repo" — popularity ≠ design quality
|
||||||
|
- "Uses the latest framework" — following trends ≠ standing out
|
||||||
|
- Utility functions — even well-written ones are tools, not designs
|
||||||
|
|
||||||
|
## Auto-detect Procedure
|
||||||
|
|
||||||
|
When the user says "find the essence":
|
||||||
|
|
||||||
|
1. **Read README fully.** What is the #1 feature the author leads with? That's a candidate.
|
||||||
|
2. **Check for design docs.** Is there `ARCHITECTURE.md` or equivalent? That's a candidate.
|
||||||
|
3. **Scan the import graph.** Which file is imported by the most other files? Use `grep -r "import.*from" src/ | sort | uniq -c | sort -rn` or equivalent. The top result is likely the core.
|
||||||
|
4. **Check file sizes.** Are any files disproportionately large or small for their apparent role? That signals hidden complexity.
|
||||||
|
5. **Check uniqueness.** Compare with 1-2 well-known alternatives. What does this project do differently?
|
||||||
|
6. **Present 1-2 candidates** to the user with evidence. Let them choose or auto-select the strongest.
|
||||||
|
|
||||||
|
### Example Output Format
|
||||||
|
|
||||||
|
```
|
||||||
|
Standout designs in {project}:
|
||||||
|
|
||||||
|
A) {Design A name} — evidenced by {README claim / file / doc}
|
||||||
|
What it does: {one sentence}
|
||||||
|
|
||||||
|
B) {Design B name} — evidenced by {code comment / unique feature / community discussion}
|
||||||
|
What it does: {one sentence}
|
||||||
|
|
||||||
|
Which should we dive into? (or I can pick the strongest)
|
||||||
|
```
|
||||||
|
|
||||||
|
## Failure Modes
|
||||||
|
|
||||||
|
| Situation | Response |
|
||||||
|
|---|---|
|
||||||
|
| No signal passes 2+ threshold | "This project uses conventional architecture. Try `/explore` for a full analysis, or pick a more architecturally interesting project." |
|
||||||
|
| User-specified module doesn't exist | Stop. Do NOT suggest an alternative. Tell the user the path doesn't exist. |
|
||||||
|
| Project is a wrapper (thin layer over another tool) | "This project is primarily a wrapper around {X}. The design is in {X}, not here. Try analyzing {X} instead." |
|
||||||
|
| Project is configuration-only (just JSON/YAML files) | "This project has no code architecture. It's configuration-driven. Try `/explore` for a full overview instead." |
|
||||||
@@ -0,0 +1,87 @@
|
|||||||
|
---
|
||||||
|
name: explore
|
||||||
|
description: Invoke when you need project-level understanding and an onboarding path. Produces a project learning report for code and non-code repositories with fixed phases for positioning, structure, flow, start path, and core designs. Not for deep code extraction or interactive teaching.
|
||||||
|
metadata:
|
||||||
|
version: "0.5.0"
|
||||||
|
---
|
||||||
|
|
||||||
|
# Explore: Project Understanding and Onboarding
|
||||||
|
|
||||||
|
Prefix your first line with 🥷 inline, not as its own paragraph.
|
||||||
|
|
||||||
|
You are a project cartographer. Your job is to help the user understand what a project is, why it is worth studying, how it is organized, and where to start.
|
||||||
|
|
||||||
|
`/explore` is the entry point for first contact with a repository or project-like artifact. It builds global understanding. It does not perform code-level essence extraction and it does not run interactive teaching.
|
||||||
|
|
||||||
|
## Project Type Detection
|
||||||
|
|
||||||
|
After the initial scan, classify the target before continuing:
|
||||||
|
|
||||||
|
| Type | Signals | What changes |
|
||||||
|
|---|---|---|
|
||||||
|
| **Code repository** | `go.mod`, `pyproject.toml`, `Cargo.toml`, source directories, executable entrypoints | Run all 4 phases |
|
||||||
|
| **Skill / docs / knowledge repository** | `SKILL.md`, mostly Markdown, docs-first structure, no runnable application entrypoint | Skip Phase 2 (Flow) and Phase 3 (Start Path) |
|
||||||
|
| **Template / scaffold repository** | Starter files, minimal logic, setup-first repo | Phase 2 may stay structural and Phase 3 may be minimal |
|
||||||
|
|
||||||
|
State the detected type before proceeding. If uncertain, say what evidence is missing and continue with the closest matching type.
|
||||||
|
|
||||||
|
## Phase 1: Positioning & Structure
|
||||||
|
- What this project is, why it is worth studying, and who it is for.
|
||||||
|
- Top-level structure: main modules, documents, directories, and the likely learning entry area.
|
||||||
|
- Tradeoffs vs alternatives when evidence exists.
|
||||||
|
|
||||||
|
## Phase 2: Flow
|
||||||
|
**Code repositories only.**
|
||||||
|
- Skip for non-code and template repositories.
|
||||||
|
- Trace the main runtime or request flow.
|
||||||
|
- Produce at least one architecture or core-flow diagram.
|
||||||
|
- Keep the trace focused on the golden path rather than exhaustive coverage.
|
||||||
|
|
||||||
|
## Phase 3: Start Path
|
||||||
|
**Code repositories only when runnable or meaningfully inspectable.**
|
||||||
|
- Provide the minimal path to start learning or running the project.
|
||||||
|
- Give the first command or first inspection step.
|
||||||
|
- Suggest one safe first modification or observation point when appropriate.
|
||||||
|
|
||||||
|
## Phase 4: Core Designs
|
||||||
|
- Summarize 2-3 core implementations or ideas.
|
||||||
|
- Keep this at overview depth.
|
||||||
|
- For each item, include what it is, where it lives, and why it matters.
|
||||||
|
|
||||||
|
## Minimum Deliverables
|
||||||
|
|
||||||
|
The final `/explore` report must include:
|
||||||
|
- Project positioning
|
||||||
|
- Why it is worth studying
|
||||||
|
- 2-3 core implementations or core ideas
|
||||||
|
- Tradeoffs or comparisons when applicable
|
||||||
|
- At least 1 diagram:
|
||||||
|
- code repository → architecture diagram or core flow diagram
|
||||||
|
- non-code repository → structure diagram, idea map, or workflow diagram
|
||||||
|
|
||||||
|
## Boundary Rules
|
||||||
|
|
||||||
|
`/explore` may:
|
||||||
|
- scan structure
|
||||||
|
- explain the main flow
|
||||||
|
- provide a minimal start path
|
||||||
|
- summarize 2-3 core designs
|
||||||
|
|
||||||
|
`/explore` must not:
|
||||||
|
- perform `/essence`-level deep extraction
|
||||||
|
- act as `/follow`-style guided teaching
|
||||||
|
- include Verify, Deep Fission, or HTML Output phases
|
||||||
|
- preserve no retired lightweight fallback behavior
|
||||||
|
|
||||||
|
## Outcome
|
||||||
|
|
||||||
|
```
|
||||||
|
Explore Report: {project name}
|
||||||
|
Project type: code / skill-docs / template
|
||||||
|
Phases completed: 4/4 (or note skipped code-only phases)
|
||||||
|
Diagram included: yes / no
|
||||||
|
Core designs: 2-3
|
||||||
|
Status: complete
|
||||||
|
```
|
||||||
|
|
||||||
|
After the report, stop. Do not proceed to `/essence` or `/follow` automatically.
|
||||||
@@ -0,0 +1,98 @@
|
|||||||
|
# Project Analysis Methods
|
||||||
|
|
||||||
|
How to read and understand an unfamiliar code project.
|
||||||
|
|
||||||
|
## 1. Identify the Entry Point
|
||||||
|
|
||||||
|
Every project has a door. Find it first.
|
||||||
|
|
||||||
|
### By Language
|
||||||
|
|
||||||
|
| Language | Look for |
|
||||||
|
|---|---|
|
||||||
|
| **JavaScript/TypeScript** | `package.json` → `main` / `bin` / `scripts.dev` |
|
||||||
|
| **Python** | `setup.py` → `entry_points`, `pyproject.toml` → `[project.scripts]`, or top-level `app.py` / `main.py` / `__main__.py` |
|
||||||
|
| **Go** | `package main` in any file, conventionally `main.go` or `cmd/*/main.go` |
|
||||||
|
| **Rust** | `src/main.rs` or `src/bin/*.rs` |
|
||||||
|
| **Java** | Class with `public static void main(String[] args)` |
|
||||||
|
| **C/C++** | `main()` function, conventionally in `src/main.c` |
|
||||||
|
| **Swift** | `main.swift` or file with `@main` attribute |
|
||||||
|
|
||||||
|
### In Frameworks
|
||||||
|
|
||||||
|
| Framework | Entry point |
|
||||||
|
|---|---|
|
||||||
|
| Next.js | `app/` or `pages/` directory, `next.config.js` |
|
||||||
|
| React (Vite) | `src/main.tsx` or `src/main.jsx` |
|
||||||
|
| Vue (Vite) | `src/main.ts` or `src/main.js` |
|
||||||
|
| Express | File that calls `app.listen()` |
|
||||||
|
| FastAPI | File that creates `FastAPI()` instance |
|
||||||
|
| Django | `manage.py`, then project name directory with `urls.py` / `wsgi.py` |
|
||||||
|
| Flask | `app.py` or `app/__init__.py` |
|
||||||
|
| Spring Boot | `*Application.java` with `@SpringBootApplication` |
|
||||||
|
|
||||||
|
## 2. Judge Project Complexity
|
||||||
|
|
||||||
|
Don't over-engineer simple projects. Don't under-analyze complex ones.
|
||||||
|
|
||||||
|
### Simple (<50 files, single language)
|
||||||
|
- Read every source file.
|
||||||
|
- No need for flow diagrams beyond a simple sequence.
|
||||||
|
- A light `/explore` pass is probably enough.
|
||||||
|
|
||||||
|
### Standard (50-500 files, 1-2 languages)
|
||||||
|
- Read entry point + core modules + 1-2 feature files.
|
||||||
|
- Build 1-2 flow diagrams.
|
||||||
|
- `/explore` is the right level.
|
||||||
|
|
||||||
|
### Complex (>500 files, multi-language, monorepo)
|
||||||
|
- Read entry point + architecture docs + one representative module.
|
||||||
|
- Use `/essence` to find standout designs, or `/explore` for one package at a time.
|
||||||
|
- Do NOT try to understand the whole project in one pass.
|
||||||
|
|
||||||
|
## 3. Separate Core Code from Scaffolding
|
||||||
|
|
||||||
|
Not all files are worth reading.
|
||||||
|
|
||||||
|
### Ignore (scaffolding)
|
||||||
|
- `*.config.js`, `*.config.ts` — configuration, not logic
|
||||||
|
- `dist/`, `build/`, `out/` — generated output
|
||||||
|
- `node_modules/`, `vendor/`, `.venv/` — dependencies
|
||||||
|
- `*.lock`, `yarn.lock`, `go.sum` — lock files
|
||||||
|
- `LICENSE`, `CODEOWNERS`, `.editorconfig` — project meta
|
||||||
|
- `test/fixtures/`, `test/data/` — test data
|
||||||
|
|
||||||
|
### Read (core)
|
||||||
|
- Entry point file
|
||||||
|
- Router/middleware/config handlers
|
||||||
|
- Model/entity/schema definitions
|
||||||
|
- Core algorithm or business logic files
|
||||||
|
- Files referenced most in imports
|
||||||
|
|
||||||
|
### Hint: Follow imports
|
||||||
|
|
||||||
|
```
|
||||||
|
entry file → import A → import B → core logic
|
||||||
|
```
|
||||||
|
|
||||||
|
Each import is a dependency. Follow the chain until you hit a file that doesn't import anything else — that's usually the core.
|
||||||
|
|
||||||
|
## 4. Read Unfamiliar Framework Code
|
||||||
|
|
||||||
|
You don't know every framework. That's fine.
|
||||||
|
|
||||||
|
### Strategy
|
||||||
|
|
||||||
|
1. **Find the routing layer first.** Every framework has a way to map URLs or events to handlers. Find it. It tells you the project's capabilities.
|
||||||
|
|
||||||
|
2. **Follow ONE request end-to-end.** Don't try to understand all routes. Pick the simplest one (often "health check" or "get by ID") and trace it from entry to response.
|
||||||
|
|
||||||
|
3. **Identify the framework's conventions.** Most frameworks follow a pattern:
|
||||||
|
- MVC: Controller → Model → View
|
||||||
|
- Middleware: Request → Middleware chain → Handler → Response
|
||||||
|
- Component: Parent renders children, props flow down, events flow up
|
||||||
|
- Plugin: Core calls hooks, plugins register handlers
|
||||||
|
|
||||||
|
4. **Don't fight the framework's abstraction.** If the project uses ORM, don't look for raw SQL. If it uses dependency injection, don't look for `new()` calls. Understand what abstraction layer they chose.
|
||||||
|
|
||||||
|
5. **Use the framework's own docs.** If stuck on "how does this framework work?", check the official docs. Don't reverse-engineer what's documented.
|
||||||
@@ -0,0 +1,173 @@
|
|||||||
|
# Flow Pattern Library
|
||||||
|
|
||||||
|
Common architecture patterns and how to identify them in code.
|
||||||
|
|
||||||
|
## MVC / MVVM / MVX
|
||||||
|
|
||||||
|
### What it is
|
||||||
|
Separation of data (Model), UI/presentation (View), and coordination logic (Controller/ViewModel).
|
||||||
|
|
||||||
|
### File signatures
|
||||||
|
| Pattern | Directories/Files |
|
||||||
|
|---|---|
|
||||||
|
| **MVC** | `controllers/`, `models/`, `views/` |
|
||||||
|
| **MVVM** | `viewmodels/`, `views/`, `models/` |
|
||||||
|
| **Layered** | `app/`, `domain/`, `infrastructure/` (Clean/Hexagonal) |
|
||||||
|
|
||||||
|
### Flow
|
||||||
|
```
|
||||||
|
Request → Controller → Model (data) → View (render) → Response
|
||||||
|
```
|
||||||
|
|
||||||
|
### Key question
|
||||||
|
"Does the file handle data, display, or coordination?" If yes → MVC-family.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Middleware Chain
|
||||||
|
|
||||||
|
### What it is
|
||||||
|
Each handler processes the request and passes it to the next. Like an assembly line.
|
||||||
|
|
||||||
|
### File signatures
|
||||||
|
| Framework | Indicator |
|
||||||
|
|---|---|---|
|
||||||
|
| **Express/Koa** | `app.use(...)`, `app.get('/', handler)` |
|
||||||
|
| **FastAPI** | `@app.middleware("http")`, `Depends()` |
|
||||||
|
| **Next.js** | `middleware.ts` at root or in `app/` |
|
||||||
|
| **Gin (Go)** | `router.Use(middleware1, middleware2)` |
|
||||||
|
| **Koa** | `app.use(async (ctx, next) => { ... })` |
|
||||||
|
|
||||||
|
### Flow
|
||||||
|
```
|
||||||
|
Request → Middleware A → Middleware B → Handler → Response
|
||||||
|
↓ ↓
|
||||||
|
auth check log request
|
||||||
|
```
|
||||||
|
|
||||||
|
### Key question
|
||||||
|
"Does this function call `next()` or pass control to something else?" If yes → middleware.
|
||||||
|
|
||||||
|
### Common middleware order
|
||||||
|
```
|
||||||
|
1. CORS / Security headers
|
||||||
|
2. Logging / Request ID
|
||||||
|
3. Authentication / Authorization
|
||||||
|
4. Body parsing / Validation
|
||||||
|
5. Rate limiting
|
||||||
|
6. Route handler
|
||||||
|
7. Error handler (catches everything above)
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Plugin / Extension System
|
||||||
|
|
||||||
|
### What it is
|
||||||
|
Core provides hooks or interfaces. External code registers handlers. The core doesn't know about specific plugins.
|
||||||
|
|
||||||
|
### File signatures
|
||||||
|
| Pattern | Indicator |
|
||||||
|
|---|---|
|
||||||
|
| **Hook-based** | `registerHook('eventName', handler)`, `hooks.on('event', fn)` |
|
||||||
|
| **Interface-based** | Abstract class or interface that plugins implement |
|
||||||
|
| **Discovery-based** | Directory scan (`plugins/`), import all, register by convention |
|
||||||
|
| **VSCode-style** | `contributes` in `package.json`, activation events |
|
||||||
|
|
||||||
|
### Flow
|
||||||
|
```
|
||||||
|
Core starts
|
||||||
|
↓
|
||||||
|
Scans for plugins
|
||||||
|
↓
|
||||||
|
Each plugin registers itself
|
||||||
|
↓
|
||||||
|
Core fires hooks → plugins respond
|
||||||
|
↓
|
||||||
|
Core runs with extended capabilities
|
||||||
|
```
|
||||||
|
|
||||||
|
### Key question
|
||||||
|
"Can I add functionality without modifying core code?" If yes → plugin architecture.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Event-Driven
|
||||||
|
|
||||||
|
### What it is
|
||||||
|
Components communicate through events, not direct calls. Publishers emit, subscribers listen.
|
||||||
|
|
||||||
|
### File signatures
|
||||||
|
| Pattern | Indicator |
|
||||||
|
|---|---|
|
||||||
|
| **Node EventEmitter** | `eventEmitter.on('event', handler)`, `eventEmitter.emit('event', data)` |
|
||||||
|
| **Pub/Sub** | `pubsub.subscribe('channel', handler)`, `pubsub.publish('channel', data)` |
|
||||||
|
| **Redux-style** | `dispatch(action)`, `reducer(state, action) → newState` |
|
||||||
|
| **Observable** | `observable.subscribe(fn)`, `pipe(map, filter)` |
|
||||||
|
| **Signals (Python)** | `@signal.connect`, `signal.send()` |
|
||||||
|
|
||||||
|
### Flow
|
||||||
|
```
|
||||||
|
Component A emits "user.created"
|
||||||
|
↓
|
||||||
|
Listener B hears it → sends welcome email
|
||||||
|
Listener C hears it → creates default settings
|
||||||
|
Listener D hears it → logs analytics
|
||||||
|
```
|
||||||
|
|
||||||
|
### Key question
|
||||||
|
"Does code communicate without importing or calling each other directly?" If yes → event-driven.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## State Management
|
||||||
|
|
||||||
|
### What it is
|
||||||
|
Centralized storage for application state. Components read and update through defined interfaces.
|
||||||
|
|
||||||
|
### File signatures
|
||||||
|
| Pattern | Indicator |
|
||||||
|
|---|---|
|
||||||
|
| **Redux** | `createStore()`, `dispatch()`, `useSelector()`, `@reduxjs/toolkit` |
|
||||||
|
| **Zustand** | `create((set) => ({ ... }))` |
|
||||||
|
| **Jotai** | `atom(value)`, `useAtom(atom)` |
|
||||||
|
| **MobX** | `@observable`, `@action`, `@computed` |
|
||||||
|
| **React Context** | `createContext()`, `useContext()`, `Provider` |
|
||||||
|
| **Pinia (Vue)** | `defineStore()`, `state`, `actions` |
|
||||||
|
|
||||||
|
### Flow
|
||||||
|
```
|
||||||
|
Component dispatches action
|
||||||
|
↓
|
||||||
|
Reducer processes action + current state
|
||||||
|
↓
|
||||||
|
New state emitted
|
||||||
|
↓
|
||||||
|
Subscribed components re-render
|
||||||
|
```
|
||||||
|
|
||||||
|
### Key question
|
||||||
|
"Where does the app store data that multiple components need?" If it's a single store → state management pattern.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Pipeline / Chain of Responsibility
|
||||||
|
|
||||||
|
### What it is
|
||||||
|
Data flows through a series of processors. Each processor transforms the data and passes it on.
|
||||||
|
|
||||||
|
### File signatures
|
||||||
|
| Pattern | Indicator |
|
||||||
|
|---|---|
|
||||||
|
| **Stream processing** | `.pipe(transform1).pipe(transform2)` |
|
||||||
|
| **Compiler/lexer** | Source → Tokenize → Parse → Transform → Generate |
|
||||||
|
| **Data pipeline** | `input → transform → validate → output` |
|
||||||
|
| **Makefile** | Target depends on prerequisites, each is a step |
|
||||||
|
|
||||||
|
### Flow
|
||||||
|
```
|
||||||
|
Raw input → Tokenizer → Parser → Transformer → Generator → Output
|
||||||
|
```
|
||||||
|
|
||||||
|
### Key question
|
||||||
|
"Does data get progressively transformed through a fixed sequence of steps?" If yes → pipeline.
|
||||||
@@ -0,0 +1,101 @@
|
|||||||
|
---
|
||||||
|
name: follow
|
||||||
|
description: Invoke when the user wants an interactive learning session based on an existing `/explore` or `/essence` report. Guides runnable or reader-style follow-along sessions. Not for fresh project analysis or pattern-only extraction.
|
||||||
|
metadata:
|
||||||
|
version: "0.5.0"
|
||||||
|
---
|
||||||
|
|
||||||
|
# Follow: Guided Learning Session
|
||||||
|
|
||||||
|
Prefix your first line with 🥷 inline, not as its own paragraph.
|
||||||
|
|
||||||
|
You are a guide. The user wants to learn from a project step by step with help, context, and correction. You guide the learning process, but you do not replace it.
|
||||||
|
|
||||||
|
`/follow` is not a fresh project analyzer. It only works from an existing `/explore` or `/essence` result.
|
||||||
|
|
||||||
|
## Pre-check
|
||||||
|
|
||||||
|
`/follow` only works when there is already an `/explore` report or an `/essence` report.
|
||||||
|
|
||||||
|
- `/explore` report exists → use it as the main learning path
|
||||||
|
- `/essence` report exists → use it for design-focused guided study
|
||||||
|
- Neither exists → refuse clearly
|
||||||
|
|
||||||
|
Refusal behavior:
|
||||||
|
"I need an existing `/explore` or `/essence` result before I can guide a follow-along session. Please run `/explore` for project understanding or `/essence` for a focused deep dive first."
|
||||||
|
|
||||||
|
Load the existing report before continuing.
|
||||||
|
|
||||||
|
## Mode Selection
|
||||||
|
|
||||||
|
After the pre-check, select one mode based on the prerequisite report:
|
||||||
|
|
||||||
|
- From `/explore` + code repository → default **Runnable**
|
||||||
|
- From `/explore` + non-code repository → force **Reader**
|
||||||
|
- From `/essence` → default **Reader** (user is in design-analysis state)
|
||||||
|
|
||||||
|
| Mode | When | Entry |
|
||||||
|
|---|---|---|
|
||||||
|
| **Runnable** | Report confirms the project is a runnable code repository and the user wants to learn by running and changing it | Start from environment and first execution |
|
||||||
|
| **Reader** | Project has no runtime, or the user is studying design/architecture, or the prerequisite report is from `/essence` | Start from guided reading |
|
||||||
|
|
||||||
|
State the selected mode before proceeding. Do not re-scan the project — use the prerequisite report to decide.
|
||||||
|
|
||||||
|
## Teaching Interaction Rules
|
||||||
|
|
||||||
|
`/follow` must teach by guidance, not by dumping answers:
|
||||||
|
- explain the purpose of the current step first
|
||||||
|
- give the user an observation point or action point
|
||||||
|
- ask the user to predict, try, or explain before revealing the answer
|
||||||
|
- then reveal, correct, or deepen the explanation
|
||||||
|
- never say "go read the code" as a standalone instruction. When referencing code, always start with: what design idea this code embodies, why it matters in the overall architecture, and what the user should pay attention to
|
||||||
|
|
||||||
|
## Runnable Check
|
||||||
|
|
||||||
|
Before Runnable mode, confirm from the **prerequisite report** (do not re-scan the project):
|
||||||
|
- If the report identified the target as a code repository with a recognized runtime (`go.mod`, `pyproject.toml`, `Cargo.toml`, `Makefile`, `build.gradle`, `pom.xml`, `CMakeLists.txt`, etc.), proceed with Runnable.
|
||||||
|
- If the report classified it as non-code, or no runtime entrypoint was found, switch to Reader and explain why.
|
||||||
|
- If the prerequisite is `/essence`, confirm with the user: essence is design-focused, Reader is the natural fit. Allow Runnable only if the user explicitly insists.
|
||||||
|
- Do not introduce a third mode.
|
||||||
|
|
||||||
|
## Runnable Mode Flow
|
||||||
|
1. Confirm environment and prerequisites.
|
||||||
|
2. Let the user run the project.
|
||||||
|
3. Let the user make one safe change.
|
||||||
|
4. Walk the main flow together.
|
||||||
|
5. Give one small exercise.
|
||||||
|
6. Review what they learned.
|
||||||
|
|
||||||
|
## Reader Mode Flow
|
||||||
|
1. Frame the learning goal around a core design or architectural idea, not a single file.
|
||||||
|
2. Walk through the design concept layer by layer: problem → approach → implementation → tradeoff.
|
||||||
|
3. Ask the user questions that probe understanding ("Why did the author choose this approach over a simpler one?"), not just prediction ("What happens next?").
|
||||||
|
4. Use diagrams or structured summaries to connect the dots between files and design ideas.
|
||||||
|
5. Give one reasoning exercise that tests whether the user can apply the design pattern elsewhere.
|
||||||
|
6. Review what they learned.
|
||||||
|
|
||||||
|
## Boundary Rules
|
||||||
|
|
||||||
|
`/follow` must:
|
||||||
|
- depend on `/explore` or `/essence`
|
||||||
|
- guide the user interactively
|
||||||
|
- adapt between code and non-code repositories through Runnable or Reader emphasis
|
||||||
|
|
||||||
|
`/follow` must not:
|
||||||
|
- rescan the whole project as a new analyzer
|
||||||
|
- reference retired skills as prerequisites
|
||||||
|
- add any third learning mode
|
||||||
|
- execute commands or write code for the user
|
||||||
|
|
||||||
|
## Outcome
|
||||||
|
|
||||||
|
```
|
||||||
|
Follow Session: {project name}
|
||||||
|
Mode: runnable / reader
|
||||||
|
Prerequisite report: /explore or /essence
|
||||||
|
Exercise result: completed / partial / too hard
|
||||||
|
Next direction: {suggested follow-up}
|
||||||
|
Status: complete
|
||||||
|
```
|
||||||
|
|
||||||
|
After the review, stop. Ask whether the user wants another exercise or wants to end the session.
|
||||||
@@ -0,0 +1,113 @@
|
|||||||
|
# Environment Detection Rules
|
||||||
|
|
||||||
|
How to detect the runtime environment and guide the user through setup in `/follow`.
|
||||||
|
|
||||||
|
## Language Detection from Config
|
||||||
|
|
||||||
|
Check these files in order. The first match is the primary language.
|
||||||
|
|
||||||
|
| Config file | Language | Runtime check | Install command |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `package.json` | JavaScript/TypeScript | `node --version` | nvm or official installer |
|
||||||
|
| `pyproject.toml` | Python | `python --version` | pyenv or python.org |
|
||||||
|
| `go.mod` | Go | `go version` | golang.org/dl |
|
||||||
|
| `Cargo.toml` | Rust | `rustc --version` | rustup |
|
||||||
|
| `pom.xml` | Java | `java -version` | SDKMAN or official |
|
||||||
|
| `build.gradle` / `build.gradle.kts` | Java/Kotlin | `java -version` | SDKMAN |
|
||||||
|
| `Gemfile` | Ruby | `ruby --version` | rvm or rbenv |
|
||||||
|
| `*.csproj` | C#/.NET | `dotnet --version` | .NET SDK |
|
||||||
|
| `CMakeLists.txt` | C/C++ | `gcc --version` or `clang --version` | System package manager |
|
||||||
|
| `swift package.json` | Swift | `swift --version` | Xcode or swift.org |
|
||||||
|
|
||||||
|
## Dependency Installation
|
||||||
|
|
||||||
|
Once language is detected, guide the user:
|
||||||
|
|
||||||
|
### JavaScript/TypeScript
|
||||||
|
```bash
|
||||||
|
# Check which package manager is used
|
||||||
|
if [ -f "yarn.lock" ]; then yarn install
|
||||||
|
elif [ -f "pnpm-lock.yaml" ]; then pnpm install
|
||||||
|
elif [ -f "bun.lockb" ] || [ -f "bun.lock" ]; then bun install
|
||||||
|
else npm install
|
||||||
|
fi
|
||||||
|
```
|
||||||
|
|
||||||
|
### Python
|
||||||
|
```bash
|
||||||
|
# Modern Python projects
|
||||||
|
pip install -e .
|
||||||
|
# Or with requirements
|
||||||
|
pip install -r requirements.txt
|
||||||
|
# Or with poetry
|
||||||
|
poetry install
|
||||||
|
# Or with uv
|
||||||
|
uv pip install -r requirements.txt
|
||||||
|
```
|
||||||
|
|
||||||
|
### Go
|
||||||
|
```bash
|
||||||
|
go mod download
|
||||||
|
```
|
||||||
|
|
||||||
|
### Rust
|
||||||
|
```bash
|
||||||
|
cargo build
|
||||||
|
```
|
||||||
|
|
||||||
|
### Java (Maven)
|
||||||
|
```bash
|
||||||
|
mvn install
|
||||||
|
```
|
||||||
|
|
||||||
|
### Java (Gradle)
|
||||||
|
```bash
|
||||||
|
./gradlew build
|
||||||
|
# or
|
||||||
|
gradle build
|
||||||
|
```
|
||||||
|
|
||||||
|
## Run Command Detection
|
||||||
|
|
||||||
|
How to start the project:
|
||||||
|
|
||||||
|
| Source | Command |
|
||||||
|
|---|---|
|
||||||
|
| `package.json` → `scripts.dev` | `npm run dev` |
|
||||||
|
| `package.json` → `scripts.start` | `npm start` |
|
||||||
|
| `Makefile` → `dev` target | `make dev` |
|
||||||
|
| `Makefile` → `run` target | `make run` |
|
||||||
|
| `pyproject.toml` (Poetry) | `poetry run python main.py` |
|
||||||
|
| `go.mod` → `package main` | `go run main.go` |
|
||||||
|
| `Cargo.toml` → `[[bin]]` | `cargo run` |
|
||||||
|
| `docker-compose.yml` exists | `docker-compose up` |
|
||||||
|
| `Dockerfile` exists, no compose | `docker build -t app . && docker run app` |
|
||||||
|
|
||||||
|
## Common Environment Issues
|
||||||
|
|
||||||
|
| Error | Cause | Fix |
|
||||||
|
|---|---|---|
|
||||||
|
| `command not found: node` | Node.js not installed | Install Node.js (recommend LTS) |
|
||||||
|
| `ModuleNotFoundError` | Python deps not installed | Run `pip install -r requirements.txt` |
|
||||||
|
| `EACCES: permission denied` | Global install without sudo | Use nvm/fnm, or prefix with sudo |
|
||||||
|
| `ENOENT: no such file` | Wrong working directory | `cd` to project root first |
|
||||||
|
| `port already in use` | Another process on same port | Kill the process or use different port |
|
||||||
|
| `go: cannot find main module` | Outside Go module | `cd` to directory with `go.mod` |
|
||||||
|
| `error: could not find Cargo.toml` | Outside Rust project | `cd` to directory with `Cargo.toml` |
|
||||||
|
| `java.lang.UnsupportedClassVersionError` | Wrong Java version | Match JDK version to project requirement |
|
||||||
|
| `npm ERR! code ERESOLVE` | Dependency conflict | Try `npm install --legacy-peer-deps` |
|
||||||
|
|
||||||
|
## Detection Script for /follow
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Quick environment check
|
||||||
|
echo "=== Environment ==="
|
||||||
|
node --version 2>/dev/null || echo "Node.js: not installed"
|
||||||
|
python --version 2>/dev/null || echo "Python: not installed"
|
||||||
|
go version 2>/dev/null || echo "Go: not installed"
|
||||||
|
rustc --version 2>/dev/null || echo "Rust: not installed"
|
||||||
|
java -version 2>/dev/null || echo "Java: not installed"
|
||||||
|
echo "PWD: $(pwd)"
|
||||||
|
```
|
||||||
|
|
||||||
|
Run this at the start of `/follow` Step 1 to understand what's available.
|
||||||
@@ -0,0 +1,57 @@
|
|||||||
|
# Frontend Design — Complete Guidance
|
||||||
|
|
||||||
|
This document provides a comprehensive framework for creating visually distinctive, non-templated UI designs. Here's the full breakdown:
|
||||||
|
|
||||||
|
## Foundational Approach
|
||||||
|
|
||||||
|
Act as the design lead for a studio known for unique client identities — the client has already turned down template-like proposals. Every choice about palette, typography, and layout must be specific to the brief, including "one real aesthetic risk you can justify."
|
||||||
|
|
||||||
|
## Grounding in Subject Matter
|
||||||
|
|
||||||
|
If the brief is vague about the product or subject, pin it down yourself: name the subject, its audience, and the page's single job. Draw inspiration from "the subject's own world, its materials, instruments, artifacts, and vernacular." Use any known context about the human's preferences or past designs as hints.
|
||||||
|
|
||||||
|
## Design Principles
|
||||||
|
|
||||||
|
- **Hero as thesis**: Open with "the most characteristic thing in the subject's world" — avoid default choices like a big number with a small label and gradient accent unless truly optimal.
|
||||||
|
- **Typography**: Pair display and body faces deliberately, not from your usual repertoire. Set a clear type scale with intentional weights, widths, and spacing. "Make the type treatment itself a memorable part of the design."
|
||||||
|
- **Structure as information**: Numbering, eyebrows, dividers must encode something true about the content. Question whether numbered markers (01/02/03) actually make sense before using them — only appropriate for real sequences.
|
||||||
|
- **Motion**: Consider where animation serves the subject. "An orchestrated moment usually lands harder than scattered effects." Sometimes less is better to avoid an AI-generated feel.
|
||||||
|
- **Complexity**: Match execution to the vision — maximalist needs elaborate execution, minimal needs precision.
|
||||||
|
- **Content**: Come up with copy if the brief lacks it. Poor copy makes a design feel as templated as poor layout.
|
||||||
|
|
||||||
|
## AI-Generated Design Traps
|
||||||
|
|
||||||
|
Three common AI-default looks to watch for: (1) warm cream background (~#F4F1EA) with serif display and terracotta accent; (2) near-black with bright acid-green or vermilion; (3) broadsheet layout with hairline rules, zero border-radius, and dense columns. "All three are legitimate for some briefs, but they are defaults rather than choices." Where the brief leaves an axis free, don't spend that freedom on a default.
|
||||||
|
|
||||||
|
## Two-Pass Process
|
||||||
|
|
||||||
|
**Pass 1 — Plan**: Create a compact token system:
|
||||||
|
|
||||||
|
1. **Color**: 4–6 named hex values
|
||||||
|
2. **Type**: Characterful display face (used with restraint), complementary body face, utility face for captions/data
|
||||||
|
3. **Layout**: One-sentence prose descriptions + ASCII wireframes
|
||||||
|
4. **Signature**: The single unique element the page will be remembered by
|
||||||
|
|
||||||
|
Review the plan against the brief. If any part reads like what you'd produce for any similar page, revise it. Only then write code.
|
||||||
|
|
||||||
|
**Pass 2 — Build**: Follow the revised plan exactly. Watch for CSS selector specificity conflicts (e.g., `.section` and `.cta` fighting over padding/margins). Do most planning internally, only sharing ideas when confident.
|
||||||
|
|
||||||
|
## Restraint & Self-Critique
|
||||||
|
|
||||||
|
"Spend your boldness in one place" — let the signature element be the one memorable thing; keep everything else quiet. "Not taking a risk can be a risk itself!" Build responsively down to mobile, with visible keyboard focus and reduced motion respected. Critique as you build. Follow Chanel's advice: before finishing, remove one accessory. Jot notes about what you've tried to avoid repeating yourself.
|
||||||
|
|
||||||
|
## Writing in Design
|
||||||
|
|
||||||
|
Words exist to make the design understandable and usable — they're "design material, not decoration." Write from the end user's perspective, naming things by what people control and recognize, never by how the system is built.
|
||||||
|
|
||||||
|
- Use active voice as default
|
||||||
|
- A control should say exactly what happens: "Save changes," not "Submit"
|
||||||
|
- Maintain consistent vocabulary throughout flows (button says "Publish," toast says "Published")
|
||||||
|
- Treat errors as guidance, not mood — explain what went wrong and how to fix it
|
||||||
|
- Empty screens are invitations to act
|
||||||
|
- Keep the register conversational: "plain verbs, sentence case, no filler"
|
||||||
|
- Let each element do exactly one job — "a label labels, an example demonstrates"
|
||||||
|
|
||||||
|
## License
|
||||||
|
|
||||||
|
Apache License 2.0 — see LICENSE.txt
|
||||||
@@ -0,0 +1,83 @@
|
|||||||
|
---
|
||||||
|
name: gitnexus-cli
|
||||||
|
description: "Use when the user needs to run GitNexus CLI commands like analyze/index a repo, check status, clean the index, generate a wiki, or list indexed repos. Examples: \"Index this repo\", \"Reanalyze the codebase\", \"Generate a wiki\""
|
||||||
|
---
|
||||||
|
|
||||||
|
# GitNexus CLI Commands
|
||||||
|
|
||||||
|
All commands work via `npx` — no global install required.
|
||||||
|
|
||||||
|
## Commands
|
||||||
|
|
||||||
|
### analyze — Build or refresh the index
|
||||||
|
|
||||||
|
```bash
|
||||||
|
npx gitnexus analyze
|
||||||
|
```
|
||||||
|
|
||||||
|
Run from the project root. This parses all source files, builds the knowledge graph, writes it to `.gitnexus/`, and generates AGENTS.md / AGENTS.md context files.
|
||||||
|
|
||||||
|
| Flag | Effect |
|
||||||
|
| -------------- | ---------------------------------------------------------------- |
|
||||||
|
| `--force` | Force full re-index even if up to date |
|
||||||
|
| `--embeddings` | Enable embedding generation for semantic search (off by default) |
|
||||||
|
| `--drop-embeddings` | Drop existing embeddings on rebuild. By default, an `analyze` without `--embeddings` preserves them. |
|
||||||
|
|
||||||
|
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale. In Codex, a PostToolUse hook detects staleness after `git commit` and `git merge` and notifies the agent to run `analyze` — the hook does not run analyze itself, to avoid blocking the agent for up to 120s and risking KuzuDB corruption on timeout.
|
||||||
|
|
||||||
|
### status — Check index freshness
|
||||||
|
|
||||||
|
```bash
|
||||||
|
npx gitnexus status
|
||||||
|
```
|
||||||
|
|
||||||
|
Shows whether the current repo has a GitNexus index, when it was last updated, and symbol/relationship counts. Use this to check if re-indexing is needed.
|
||||||
|
|
||||||
|
### clean — Delete the index
|
||||||
|
|
||||||
|
```bash
|
||||||
|
npx gitnexus clean
|
||||||
|
```
|
||||||
|
|
||||||
|
Deletes the `.gitnexus/` directory and unregisters the repo from the global registry. Use before re-indexing if the index is corrupt or after removing GitNexus from a project.
|
||||||
|
|
||||||
|
| Flag | Effect |
|
||||||
|
| --------- | ------------------------------------------------- |
|
||||||
|
| `--force` | Skip confirmation prompt |
|
||||||
|
| `--all` | Clean all indexed repos, not just the current one |
|
||||||
|
|
||||||
|
### wiki — Generate documentation from the graph
|
||||||
|
|
||||||
|
```bash
|
||||||
|
npx gitnexus wiki
|
||||||
|
```
|
||||||
|
|
||||||
|
Generates repository documentation from the knowledge graph using an LLM. Requires an API key (saved to `~/.gitnexus/config.json` on first use).
|
||||||
|
|
||||||
|
| Flag | Effect |
|
||||||
|
| ------------------- | ----------------------------------------- |
|
||||||
|
| `--force` | Force full regeneration |
|
||||||
|
| `--model <model>` | LLM model (default: minimax/minimax-m2.5) |
|
||||||
|
| `--base-url <url>` | LLM API base URL |
|
||||||
|
| `--api-key <key>` | LLM API key |
|
||||||
|
| `--concurrency <n>` | Parallel LLM calls (default: 3) |
|
||||||
|
| `--gist` | Publish wiki as a public GitHub Gist |
|
||||||
|
|
||||||
|
### list — Show all indexed repos
|
||||||
|
|
||||||
|
```bash
|
||||||
|
npx gitnexus list
|
||||||
|
```
|
||||||
|
|
||||||
|
Lists all repositories registered in `~/.gitnexus/registry.json`. The MCP `list_repos` tool provides the same information.
|
||||||
|
|
||||||
|
## After Indexing
|
||||||
|
|
||||||
|
1. **Read `gitnexus://repo/{name}/context`** to verify the index loaded
|
||||||
|
2. Use the other GitNexus skills (`exploring`, `debugging`, `impact-analysis`, `refactoring`) for your task
|
||||||
|
|
||||||
|
## Troubleshooting
|
||||||
|
|
||||||
|
- **"Not inside a git repository"**: Run from a directory inside a git repo
|
||||||
|
- **Index is stale after re-analyzing**: Restart Codex to reload the MCP server
|
||||||
|
- **Embeddings slow**: Omit `--embeddings` (it's off by default) or set `OPENAI_API_KEY` for faster API-based embedding
|
||||||
@@ -0,0 +1,89 @@
|
|||||||
|
---
|
||||||
|
name: gitnexus-debugging
|
||||||
|
description: "Use when the user is debugging a bug, tracing an error, or asking why something fails. Examples: \"Why is X failing?\", \"Where does this error come from?\", \"Trace this bug\""
|
||||||
|
---
|
||||||
|
|
||||||
|
# Debugging with GitNexus
|
||||||
|
|
||||||
|
## When to Use
|
||||||
|
|
||||||
|
- "Why is this function failing?"
|
||||||
|
- "Trace where this error comes from"
|
||||||
|
- "Who calls this method?"
|
||||||
|
- "This endpoint returns 500"
|
||||||
|
- Investigating bugs, errors, or unexpected behavior
|
||||||
|
|
||||||
|
## Workflow
|
||||||
|
|
||||||
|
```
|
||||||
|
1. gitnexus_query({query: "<error or symptom>"}) → Find related execution flows
|
||||||
|
2. gitnexus_context({name: "<suspect>"}) → See callers/callees/processes
|
||||||
|
3. READ gitnexus://repo/{name}/process/{name} → Trace execution flow
|
||||||
|
4. gitnexus_cypher({query: "MATCH path..."}) → Custom traces if needed
|
||||||
|
```
|
||||||
|
|
||||||
|
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
|
||||||
|
|
||||||
|
## Checklist
|
||||||
|
|
||||||
|
```
|
||||||
|
- [ ] Understand the symptom (error message, unexpected behavior)
|
||||||
|
- [ ] gitnexus_query for error text or related code
|
||||||
|
- [ ] Identify the suspect function from returned processes
|
||||||
|
- [ ] gitnexus_context to see callers and callees
|
||||||
|
- [ ] Trace execution flow via process resource if applicable
|
||||||
|
- [ ] gitnexus_cypher for custom call chain traces if needed
|
||||||
|
- [ ] Read source files to confirm root cause
|
||||||
|
```
|
||||||
|
|
||||||
|
## Debugging Patterns
|
||||||
|
|
||||||
|
| Symptom | GitNexus Approach |
|
||||||
|
| -------------------- | ---------------------------------------------------------- |
|
||||||
|
| Error message | `gitnexus_query` for error text → `context` on throw sites |
|
||||||
|
| Wrong return value | `context` on the function → trace callees for data flow |
|
||||||
|
| Intermittent failure | `context` → look for external calls, async deps |
|
||||||
|
| Performance issue | `context` → find symbols with many callers (hot paths) |
|
||||||
|
| Recent regression | `detect_changes` to see what your changes affect |
|
||||||
|
|
||||||
|
## Tools
|
||||||
|
|
||||||
|
**gitnexus_query** — find code related to error:
|
||||||
|
|
||||||
|
```
|
||||||
|
gitnexus_query({query: "payment validation error"})
|
||||||
|
→ Processes: CheckoutFlow, ErrorHandling
|
||||||
|
→ Symbols: validatePayment, handlePaymentError, PaymentException
|
||||||
|
```
|
||||||
|
|
||||||
|
**gitnexus_context** — full context for a suspect:
|
||||||
|
|
||||||
|
```
|
||||||
|
gitnexus_context({name: "validatePayment"})
|
||||||
|
→ Incoming calls: processCheckout, webhookHandler
|
||||||
|
→ Outgoing calls: verifyCard, fetchRates (external API!)
|
||||||
|
→ Processes: CheckoutFlow (step 3/7)
|
||||||
|
```
|
||||||
|
|
||||||
|
**gitnexus_cypher** — custom call chain traces:
|
||||||
|
|
||||||
|
```cypher
|
||||||
|
MATCH path = (a)-[:CodeRelation {type: 'CALLS'}*1..2]->(b:Function {name: "validatePayment"})
|
||||||
|
RETURN [n IN nodes(path) | n.name] AS chain
|
||||||
|
```
|
||||||
|
|
||||||
|
## Example: "Payment endpoint returns 500 intermittently"
|
||||||
|
|
||||||
|
```
|
||||||
|
1. gitnexus_query({query: "payment error handling"})
|
||||||
|
→ Processes: CheckoutFlow, ErrorHandling
|
||||||
|
→ Symbols: validatePayment, handlePaymentError
|
||||||
|
|
||||||
|
2. gitnexus_context({name: "validatePayment"})
|
||||||
|
→ Outgoing calls: verifyCard, fetchRates (external API!)
|
||||||
|
|
||||||
|
3. READ gitnexus://repo/my-app/process/CheckoutFlow
|
||||||
|
→ Step 3: validatePayment → calls fetchRates (external)
|
||||||
|
|
||||||
|
4. Root cause: fetchRates calls external API without proper timeout
|
||||||
|
```
|
||||||
@@ -0,0 +1,78 @@
|
|||||||
|
---
|
||||||
|
name: gitnexus-exploring
|
||||||
|
description: "Use when the user asks how code works, wants to understand architecture, trace execution flows, or explore unfamiliar parts of the codebase. Examples: \"How does X work?\", \"What calls this function?\", \"Show me the auth flow\""
|
||||||
|
---
|
||||||
|
|
||||||
|
# Exploring Codebases with GitNexus
|
||||||
|
|
||||||
|
## When to Use
|
||||||
|
|
||||||
|
- "How does authentication work?"
|
||||||
|
- "What's the project structure?"
|
||||||
|
- "Show me the main components"
|
||||||
|
- "Where is the database logic?"
|
||||||
|
- Understanding code you haven't seen before
|
||||||
|
|
||||||
|
## Workflow
|
||||||
|
|
||||||
|
```
|
||||||
|
1. READ gitnexus://repos → Discover indexed repos
|
||||||
|
2. READ gitnexus://repo/{name}/context → Codebase overview, check staleness
|
||||||
|
3. gitnexus_query({query: "<what you want to understand>"}) → Find related execution flows
|
||||||
|
4. gitnexus_context({name: "<symbol>"}) → Deep dive on specific symbol
|
||||||
|
5. READ gitnexus://repo/{name}/process/{name} → Trace full execution flow
|
||||||
|
```
|
||||||
|
|
||||||
|
> If step 2 says "Index is stale" → run `npx gitnexus analyze` in terminal.
|
||||||
|
|
||||||
|
## Checklist
|
||||||
|
|
||||||
|
```
|
||||||
|
- [ ] READ gitnexus://repo/{name}/context
|
||||||
|
- [ ] gitnexus_query for the concept you want to understand
|
||||||
|
- [ ] Review returned processes (execution flows)
|
||||||
|
- [ ] gitnexus_context on key symbols for callers/callees
|
||||||
|
- [ ] READ process resource for full execution traces
|
||||||
|
- [ ] Read source files for implementation details
|
||||||
|
```
|
||||||
|
|
||||||
|
## Resources
|
||||||
|
|
||||||
|
| Resource | What you get |
|
||||||
|
| --------------------------------------- | ------------------------------------------------------- |
|
||||||
|
| `gitnexus://repo/{name}/context` | Stats, staleness warning (~150 tokens) |
|
||||||
|
| `gitnexus://repo/{name}/clusters` | All functional areas with cohesion scores (~300 tokens) |
|
||||||
|
| `gitnexus://repo/{name}/cluster/{name}` | Area members with file paths (~500 tokens) |
|
||||||
|
| `gitnexus://repo/{name}/process/{name}` | Step-by-step execution trace (~200 tokens) |
|
||||||
|
|
||||||
|
## Tools
|
||||||
|
|
||||||
|
**gitnexus_query** — find execution flows related to a concept:
|
||||||
|
|
||||||
|
```
|
||||||
|
gitnexus_query({query: "payment processing"})
|
||||||
|
→ Processes: CheckoutFlow, RefundFlow, WebhookHandler
|
||||||
|
→ Symbols grouped by flow with file locations
|
||||||
|
```
|
||||||
|
|
||||||
|
**gitnexus_context** — 360-degree view of a symbol:
|
||||||
|
|
||||||
|
```
|
||||||
|
gitnexus_context({name: "validateUser"})
|
||||||
|
→ Incoming calls: loginHandler, apiMiddleware
|
||||||
|
→ Outgoing calls: checkToken, getUserById
|
||||||
|
→ Processes: LoginFlow (step 2/5), TokenRefresh (step 1/3)
|
||||||
|
```
|
||||||
|
|
||||||
|
## Example: "How does payment processing work?"
|
||||||
|
|
||||||
|
```
|
||||||
|
1. READ gitnexus://repo/my-app/context → 918 symbols, 45 processes
|
||||||
|
2. gitnexus_query({query: "payment processing"})
|
||||||
|
→ CheckoutFlow: processPayment → validateCard → chargeStripe
|
||||||
|
→ RefundFlow: initiateRefund → calculateRefund → processRefund
|
||||||
|
3. gitnexus_context({name: "processPayment"})
|
||||||
|
→ Incoming: checkoutHandler, webhookHandler
|
||||||
|
→ Outgoing: validateCard, chargeStripe, saveTransaction
|
||||||
|
4. Read src/payments/processor.ts for implementation details
|
||||||
|
```
|
||||||
@@ -0,0 +1,64 @@
|
|||||||
|
---
|
||||||
|
name: gitnexus-guide
|
||||||
|
description: "Use when the user asks about GitNexus itself — available tools, how to query the knowledge graph, MCP resources, graph schema, or workflow reference. Examples: \"What GitNexus tools are available?\", \"How do I use GitNexus?\""
|
||||||
|
---
|
||||||
|
|
||||||
|
# GitNexus Guide
|
||||||
|
|
||||||
|
Quick reference for all GitNexus MCP tools, resources, and the knowledge graph schema.
|
||||||
|
|
||||||
|
## Always Start Here
|
||||||
|
|
||||||
|
For any task involving code understanding, debugging, impact analysis, or refactoring:
|
||||||
|
|
||||||
|
1. **Read `gitnexus://repo/{name}/context`** — codebase overview + check index freshness
|
||||||
|
2. **Match your task to a skill below** and **read that skill file**
|
||||||
|
3. **Follow the skill's workflow and checklist**
|
||||||
|
|
||||||
|
> If step 1 warns the index is stale, run `npx gitnexus analyze` in the terminal first.
|
||||||
|
|
||||||
|
## Skills
|
||||||
|
|
||||||
|
| Task | Skill to read |
|
||||||
|
| -------------------------------------------- | ------------------- |
|
||||||
|
| Understand architecture / "How does X work?" | `gitnexus-exploring` |
|
||||||
|
| Blast radius / "What breaks if I change X?" | `gitnexus-impact-analysis` |
|
||||||
|
| Trace bugs / "Why is X failing?" | `gitnexus-debugging` |
|
||||||
|
| Rename / extract / split / refactor | `gitnexus-refactoring` |
|
||||||
|
| Tools, resources, schema reference | `gitnexus-guide` (this file) |
|
||||||
|
| Index, status, clean, wiki CLI commands | `gitnexus-cli` |
|
||||||
|
|
||||||
|
## Tools Reference
|
||||||
|
|
||||||
|
| Tool | What it gives you |
|
||||||
|
| ---------------- | ------------------------------------------------------------------------ |
|
||||||
|
| `query` | Process-grouped code intelligence — execution flows related to a concept |
|
||||||
|
| `context` | 360-degree symbol view — categorized refs, processes it participates in |
|
||||||
|
| `impact` | Symbol blast radius — what breaks at depth 1/2/3 with confidence |
|
||||||
|
| `detect_changes` | Git-diff impact — what do your current changes affect |
|
||||||
|
| `rename` | Multi-file coordinated rename with confidence-tagged edits |
|
||||||
|
| `cypher` | Raw graph queries (read `gitnexus://repo/{name}/schema` first) |
|
||||||
|
| `list_repos` | Discover indexed repos |
|
||||||
|
|
||||||
|
## Resources Reference
|
||||||
|
|
||||||
|
Lightweight reads (~100-500 tokens) for navigation:
|
||||||
|
|
||||||
|
| Resource | Content |
|
||||||
|
| ---------------------------------------------- | ----------------------------------------- |
|
||||||
|
| `gitnexus://repo/{name}/context` | Stats, staleness check |
|
||||||
|
| `gitnexus://repo/{name}/clusters` | All functional areas with cohesion scores |
|
||||||
|
| `gitnexus://repo/{name}/cluster/{clusterName}` | Area members |
|
||||||
|
| `gitnexus://repo/{name}/processes` | All execution flows |
|
||||||
|
| `gitnexus://repo/{name}/process/{processName}` | Step-by-step trace |
|
||||||
|
| `gitnexus://repo/{name}/schema` | Graph schema for Cypher |
|
||||||
|
|
||||||
|
## Graph Schema
|
||||||
|
|
||||||
|
**Nodes:** File, Function, Class, Interface, Method, Community, Process
|
||||||
|
**Edges (via CodeRelation.type):** CALLS, IMPORTS, EXTENDS, IMPLEMENTS, DEFINES, MEMBER_OF, STEP_IN_PROCESS
|
||||||
|
|
||||||
|
```cypher
|
||||||
|
MATCH (caller)-[:CodeRelation {type: 'CALLS'}]->(f:Function {name: "myFunc"})
|
||||||
|
RETURN caller.name, caller.filePath
|
||||||
|
```
|
||||||
@@ -0,0 +1,97 @@
|
|||||||
|
---
|
||||||
|
name: gitnexus-impact-analysis
|
||||||
|
description: "Use when the user wants to know what will break if they change something, or needs safety analysis before editing code. Examples: \"Is it safe to change X?\", \"What depends on this?\", \"What will break?\""
|
||||||
|
---
|
||||||
|
|
||||||
|
# Impact Analysis with GitNexus
|
||||||
|
|
||||||
|
## When to Use
|
||||||
|
|
||||||
|
- "Is it safe to change this function?"
|
||||||
|
- "What will break if I modify X?"
|
||||||
|
- "Show me the blast radius"
|
||||||
|
- "Who uses this code?"
|
||||||
|
- Before making non-trivial code changes
|
||||||
|
- Before committing — to understand what your changes affect
|
||||||
|
|
||||||
|
## Workflow
|
||||||
|
|
||||||
|
```
|
||||||
|
1. gitnexus_impact({target: "X", direction: "upstream"}) → What depends on this
|
||||||
|
2. READ gitnexus://repo/{name}/processes → Check affected execution flows
|
||||||
|
3. gitnexus_detect_changes() → Map current git changes to affected flows
|
||||||
|
4. Assess risk and report to user
|
||||||
|
```
|
||||||
|
|
||||||
|
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
|
||||||
|
|
||||||
|
## Checklist
|
||||||
|
|
||||||
|
```
|
||||||
|
- [ ] gitnexus_impact({target, direction: "upstream"}) to find dependents
|
||||||
|
- [ ] Review d=1 items first (these WILL BREAK)
|
||||||
|
- [ ] Check high-confidence (>0.8) dependencies
|
||||||
|
- [ ] READ processes to check affected execution flows
|
||||||
|
- [ ] gitnexus_detect_changes() for pre-commit check
|
||||||
|
- [ ] Assess risk level and report to user
|
||||||
|
```
|
||||||
|
|
||||||
|
## Understanding Output
|
||||||
|
|
||||||
|
| Depth | Risk Level | Meaning |
|
||||||
|
| ----- | ---------------- | ------------------------ |
|
||||||
|
| d=1 | **WILL BREAK** | Direct callers/importers |
|
||||||
|
| d=2 | LIKELY AFFECTED | Indirect dependencies |
|
||||||
|
| d=3 | MAY NEED TESTING | Transitive effects |
|
||||||
|
|
||||||
|
## Risk Assessment
|
||||||
|
|
||||||
|
| Affected | Risk |
|
||||||
|
| ------------------------------ | -------- |
|
||||||
|
| <5 symbols, few processes | LOW |
|
||||||
|
| 5-15 symbols, 2-5 processes | MEDIUM |
|
||||||
|
| >15 symbols or many processes | HIGH |
|
||||||
|
| Critical path (auth, payments) | CRITICAL |
|
||||||
|
|
||||||
|
## Tools
|
||||||
|
|
||||||
|
**gitnexus_impact** — the primary tool for symbol blast radius:
|
||||||
|
|
||||||
|
```
|
||||||
|
gitnexus_impact({
|
||||||
|
target: "validateUser",
|
||||||
|
direction: "upstream",
|
||||||
|
minConfidence: 0.8,
|
||||||
|
maxDepth: 3
|
||||||
|
})
|
||||||
|
|
||||||
|
→ d=1 (WILL BREAK):
|
||||||
|
- loginHandler (src/auth/login.ts:42) [CALLS, 100%]
|
||||||
|
- apiMiddleware (src/api/middleware.ts:15) [CALLS, 100%]
|
||||||
|
|
||||||
|
→ d=2 (LIKELY AFFECTED):
|
||||||
|
- authRouter (src/routes/auth.ts:22) [CALLS, 95%]
|
||||||
|
```
|
||||||
|
|
||||||
|
**gitnexus_detect_changes** — git-diff based impact analysis:
|
||||||
|
|
||||||
|
```
|
||||||
|
gitnexus_detect_changes({scope: "staged"})
|
||||||
|
|
||||||
|
→ Changed: 5 symbols in 3 files
|
||||||
|
→ Affected: LoginFlow, TokenRefresh, APIMiddlewarePipeline
|
||||||
|
→ Risk: MEDIUM
|
||||||
|
```
|
||||||
|
|
||||||
|
## Example: "What breaks if I change validateUser?"
|
||||||
|
|
||||||
|
```
|
||||||
|
1. gitnexus_impact({target: "validateUser", direction: "upstream"})
|
||||||
|
→ d=1: loginHandler, apiMiddleware (WILL BREAK)
|
||||||
|
→ d=2: authRouter, sessionManager (LIKELY AFFECTED)
|
||||||
|
|
||||||
|
2. READ gitnexus://repo/my-app/processes
|
||||||
|
→ LoginFlow and TokenRefresh touch validateUser
|
||||||
|
|
||||||
|
3. Risk: 2 direct callers, 2 processes = MEDIUM
|
||||||
|
```
|
||||||
@@ -0,0 +1,121 @@
|
|||||||
|
---
|
||||||
|
name: gitnexus-refactoring
|
||||||
|
description: "Use when the user wants to rename, extract, split, move, or restructure code safely. Examples: \"Rename this function\", \"Extract this into a module\", \"Refactor this class\", \"Move this to a separate file\""
|
||||||
|
---
|
||||||
|
|
||||||
|
# Refactoring with GitNexus
|
||||||
|
|
||||||
|
## When to Use
|
||||||
|
|
||||||
|
- "Rename this function safely"
|
||||||
|
- "Extract this into a module"
|
||||||
|
- "Split this service"
|
||||||
|
- "Move this to a new file"
|
||||||
|
- Any task involving renaming, extracting, splitting, or restructuring code
|
||||||
|
|
||||||
|
## Workflow
|
||||||
|
|
||||||
|
```
|
||||||
|
1. gitnexus_impact({target: "X", direction: "upstream"}) → Map all dependents
|
||||||
|
2. gitnexus_query({query: "X"}) → Find execution flows involving X
|
||||||
|
3. gitnexus_context({name: "X"}) → See all incoming/outgoing refs
|
||||||
|
4. Plan update order: interfaces → implementations → callers → tests
|
||||||
|
```
|
||||||
|
|
||||||
|
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
|
||||||
|
|
||||||
|
## Checklists
|
||||||
|
|
||||||
|
### Rename Symbol
|
||||||
|
|
||||||
|
```
|
||||||
|
- [ ] gitnexus_rename({symbol_name: "oldName", new_name: "newName", dry_run: true}) — preview all edits
|
||||||
|
- [ ] Review graph edits (high confidence) and ast_search edits (review carefully)
|
||||||
|
- [ ] If satisfied: gitnexus_rename({..., dry_run: false}) — apply edits
|
||||||
|
- [ ] gitnexus_detect_changes() — verify only expected files changed
|
||||||
|
- [ ] Run tests for affected processes
|
||||||
|
```
|
||||||
|
|
||||||
|
### Extract Module
|
||||||
|
|
||||||
|
```
|
||||||
|
- [ ] gitnexus_context({name: target}) — see all incoming/outgoing refs
|
||||||
|
- [ ] gitnexus_impact({target, direction: "upstream"}) — find all external callers
|
||||||
|
- [ ] Define new module interface
|
||||||
|
- [ ] Extract code, update imports
|
||||||
|
- [ ] gitnexus_detect_changes() — verify affected scope
|
||||||
|
- [ ] Run tests for affected processes
|
||||||
|
```
|
||||||
|
|
||||||
|
### Split Function/Service
|
||||||
|
|
||||||
|
```
|
||||||
|
- [ ] gitnexus_context({name: target}) — understand all callees
|
||||||
|
- [ ] Group callees by responsibility
|
||||||
|
- [ ] gitnexus_impact({target, direction: "upstream"}) — map callers to update
|
||||||
|
- [ ] Create new functions/services
|
||||||
|
- [ ] Update callers
|
||||||
|
- [ ] gitnexus_detect_changes() — verify affected scope
|
||||||
|
- [ ] Run tests for affected processes
|
||||||
|
```
|
||||||
|
|
||||||
|
## Tools
|
||||||
|
|
||||||
|
**gitnexus_rename** — automated multi-file rename:
|
||||||
|
|
||||||
|
```
|
||||||
|
gitnexus_rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: true})
|
||||||
|
→ 12 edits across 8 files
|
||||||
|
→ 10 graph edits (high confidence), 2 ast_search edits (review)
|
||||||
|
→ Changes: [{file_path, edits: [{line, old_text, new_text, confidence}]}]
|
||||||
|
```
|
||||||
|
|
||||||
|
**gitnexus_impact** — map all dependents first:
|
||||||
|
|
||||||
|
```
|
||||||
|
gitnexus_impact({target: "validateUser", direction: "upstream"})
|
||||||
|
→ d=1: loginHandler, apiMiddleware, testUtils
|
||||||
|
→ Affected Processes: LoginFlow, TokenRefresh
|
||||||
|
```
|
||||||
|
|
||||||
|
**gitnexus_detect_changes** — verify your changes after refactoring:
|
||||||
|
|
||||||
|
```
|
||||||
|
gitnexus_detect_changes({scope: "all"})
|
||||||
|
→ Changed: 8 files, 12 symbols
|
||||||
|
→ Affected processes: LoginFlow, TokenRefresh
|
||||||
|
→ Risk: MEDIUM
|
||||||
|
```
|
||||||
|
|
||||||
|
**gitnexus_cypher** — custom reference queries:
|
||||||
|
|
||||||
|
```cypher
|
||||||
|
MATCH (caller)-[:CodeRelation {type: 'CALLS'}]->(f:Function {name: "validateUser"})
|
||||||
|
RETURN caller.name, caller.filePath ORDER BY caller.filePath
|
||||||
|
```
|
||||||
|
|
||||||
|
## Risk Rules
|
||||||
|
|
||||||
|
| Risk Factor | Mitigation |
|
||||||
|
| ------------------- | ----------------------------------------- |
|
||||||
|
| Many callers (>5) | Use gitnexus_rename for automated updates |
|
||||||
|
| Cross-area refs | Use detect_changes after to verify scope |
|
||||||
|
| String/dynamic refs | gitnexus_query to find them |
|
||||||
|
| External/public API | Version and deprecate properly |
|
||||||
|
|
||||||
|
## Example: Rename `validateUser` to `authenticateUser`
|
||||||
|
|
||||||
|
```
|
||||||
|
1. gitnexus_rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: true})
|
||||||
|
→ 12 edits: 10 graph (safe), 2 ast_search (review)
|
||||||
|
→ Files: validator.ts, login.ts, middleware.ts, config.json...
|
||||||
|
|
||||||
|
2. Review ast_search edits (config.json: dynamic reference!)
|
||||||
|
|
||||||
|
3. gitnexus_rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: false})
|
||||||
|
→ Applied 12 edits across 8 files
|
||||||
|
|
||||||
|
4. gitnexus_detect_changes({scope: "all"})
|
||||||
|
→ Affected: LoginFlow, TokenRefresh
|
||||||
|
→ Risk: MEDIUM — run tests for these flows
|
||||||
|
```
|
||||||
@@ -0,0 +1,16 @@
|
|||||||
|
---
|
||||||
|
name: handoff
|
||||||
|
description: Compact the current conversation into a handoff document for another agent to pick up.
|
||||||
|
argument-hint: "What will the next session be used for?"
|
||||||
|
disable-model-invocation: true
|
||||||
|
---
|
||||||
|
|
||||||
|
Write a handoff document summarising the current conversation so a fresh agent can continue the work. Save to the temporary directory of the user's OS - not the current workspace.
|
||||||
|
|
||||||
|
Include a "suggested skills" section in the document, which suggests skills that the agent should invoke.
|
||||||
|
|
||||||
|
Do not duplicate content already captured in other artifacts (PRDs, plans, ADRs, issues, commits, diffs). Reference them by path or URL instead.
|
||||||
|
|
||||||
|
Redact any sensitive information, such as API keys, passwords, or personally identifiable information.
|
||||||
|
|
||||||
|
If the user passed arguments, treat them as a description of what the next session will focus on and tailor the doc accordingly.
|
||||||
@@ -0,0 +1,156 @@
|
|||||||
|
---
|
||||||
|
name: openspec-apply-change
|
||||||
|
description: Implement tasks from an OpenSpec change. Use when the user wants to start implementing, continue implementation, or work through tasks.
|
||||||
|
license: MIT
|
||||||
|
compatibility: Requires openspec CLI.
|
||||||
|
metadata:
|
||||||
|
author: openspec
|
||||||
|
version: "1.0"
|
||||||
|
generatedBy: "1.3.1"
|
||||||
|
---
|
||||||
|
|
||||||
|
Implement tasks from an OpenSpec change.
|
||||||
|
|
||||||
|
**Input**: Optionally specify a change name. If omitted, check if it can be inferred from conversation context. If vague or ambiguous you MUST prompt for available changes.
|
||||||
|
|
||||||
|
**Steps**
|
||||||
|
|
||||||
|
1. **Select the change**
|
||||||
|
|
||||||
|
If a name is provided, use it. Otherwise:
|
||||||
|
- Infer from conversation context if the user mentioned a change
|
||||||
|
- Auto-select if only one active change exists
|
||||||
|
- If ambiguous, run `openspec list --json` to get available changes and use the **AskUserQuestion tool** to let the user select
|
||||||
|
|
||||||
|
Always announce: "Using change: <name>" and how to override (e.g., `/opsx:apply <other>`).
|
||||||
|
|
||||||
|
2. **Check status to understand the schema**
|
||||||
|
```bash
|
||||||
|
openspec status --change "<name>" --json
|
||||||
|
```
|
||||||
|
Parse the JSON to understand:
|
||||||
|
- `schemaName`: The workflow being used (e.g., "spec-driven")
|
||||||
|
- Which artifact contains the tasks (typically "tasks" for spec-driven, check status for others)
|
||||||
|
|
||||||
|
3. **Get apply instructions**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
openspec instructions apply --change "<name>" --json
|
||||||
|
```
|
||||||
|
|
||||||
|
This returns:
|
||||||
|
- `contextFiles`: artifact ID -> array of concrete file paths (varies by schema - could be proposal/specs/design/tasks or spec/tests/implementation/docs)
|
||||||
|
- Progress (total, complete, remaining)
|
||||||
|
- Task list with status
|
||||||
|
- Dynamic instruction based on current state
|
||||||
|
|
||||||
|
**Handle states:**
|
||||||
|
- If `state: "blocked"` (missing artifacts): show message, suggest using openspec-continue-change
|
||||||
|
- If `state: "all_done"`: congratulate, suggest archive
|
||||||
|
- Otherwise: proceed to implementation
|
||||||
|
|
||||||
|
4. **Read context files**
|
||||||
|
|
||||||
|
Read every file path listed under `contextFiles` from the apply instructions output.
|
||||||
|
The files depend on the schema being used:
|
||||||
|
- **spec-driven**: proposal, specs, design, tasks
|
||||||
|
- Other schemas: follow the contextFiles from CLI output
|
||||||
|
|
||||||
|
5. **Show current progress**
|
||||||
|
|
||||||
|
Display:
|
||||||
|
- Schema being used
|
||||||
|
- Progress: "N/M tasks complete"
|
||||||
|
- Remaining tasks overview
|
||||||
|
- Dynamic instruction from CLI
|
||||||
|
|
||||||
|
6. **Implement tasks (loop until done or blocked)**
|
||||||
|
|
||||||
|
For each pending task:
|
||||||
|
- Show which task is being worked on
|
||||||
|
- Make the code changes required
|
||||||
|
- Keep changes minimal and focused
|
||||||
|
- Mark task complete in the tasks file: `- [ ]` → `- [x]`
|
||||||
|
- Continue to next task
|
||||||
|
|
||||||
|
**Pause if:**
|
||||||
|
- Task is unclear → ask for clarification
|
||||||
|
- Implementation reveals a design issue → suggest updating artifacts
|
||||||
|
- Error or blocker encountered → report and wait for guidance
|
||||||
|
- User interrupts
|
||||||
|
|
||||||
|
7. **On completion or pause, show status**
|
||||||
|
|
||||||
|
Display:
|
||||||
|
- Tasks completed this session
|
||||||
|
- Overall progress: "N/M tasks complete"
|
||||||
|
- If all done: suggest archive
|
||||||
|
- If paused: explain why and wait for guidance
|
||||||
|
|
||||||
|
**Output During Implementation**
|
||||||
|
|
||||||
|
```
|
||||||
|
## Implementing: <change-name> (schema: <schema-name>)
|
||||||
|
|
||||||
|
Working on task 3/7: <task description>
|
||||||
|
[...implementation happening...]
|
||||||
|
✓ Task complete
|
||||||
|
|
||||||
|
Working on task 4/7: <task description>
|
||||||
|
[...implementation happening...]
|
||||||
|
✓ Task complete
|
||||||
|
```
|
||||||
|
|
||||||
|
**Output On Completion**
|
||||||
|
|
||||||
|
```
|
||||||
|
## Implementation Complete
|
||||||
|
|
||||||
|
**Change:** <change-name>
|
||||||
|
**Schema:** <schema-name>
|
||||||
|
**Progress:** 7/7 tasks complete ✓
|
||||||
|
|
||||||
|
### Completed This Session
|
||||||
|
- [x] Task 1
|
||||||
|
- [x] Task 2
|
||||||
|
...
|
||||||
|
|
||||||
|
All tasks complete! Ready to archive this change.
|
||||||
|
```
|
||||||
|
|
||||||
|
**Output On Pause (Issue Encountered)**
|
||||||
|
|
||||||
|
```
|
||||||
|
## Implementation Paused
|
||||||
|
|
||||||
|
**Change:** <change-name>
|
||||||
|
**Schema:** <schema-name>
|
||||||
|
**Progress:** 4/7 tasks complete
|
||||||
|
|
||||||
|
### Issue Encountered
|
||||||
|
<description of the issue>
|
||||||
|
|
||||||
|
**Options:**
|
||||||
|
1. <option 1>
|
||||||
|
2. <option 2>
|
||||||
|
3. Other approach
|
||||||
|
|
||||||
|
What would you like to do?
|
||||||
|
```
|
||||||
|
|
||||||
|
**Guardrails**
|
||||||
|
- Keep going through tasks until done or blocked
|
||||||
|
- Always read context files before starting (from the apply instructions output)
|
||||||
|
- If task is ambiguous, pause and ask before implementing
|
||||||
|
- If implementation reveals issues, pause and suggest artifact updates
|
||||||
|
- Keep code changes minimal and scoped to each task
|
||||||
|
- Update task checkbox immediately after completing each task
|
||||||
|
- Pause on errors, blockers, or unclear requirements - don't guess
|
||||||
|
- Use contextFiles from CLI output, don't assume specific file names
|
||||||
|
|
||||||
|
**Fluid Workflow Integration**
|
||||||
|
|
||||||
|
This skill supports the "actions on a change" model:
|
||||||
|
|
||||||
|
- **Can be invoked anytime**: Before all artifacts are done (if tasks exist), after partial implementation, interleaved with other actions
|
||||||
|
- **Allows artifact updates**: If implementation reveals design issues, suggest updating artifacts - not phase-locked, work fluidly
|
||||||
@@ -0,0 +1,114 @@
|
|||||||
|
---
|
||||||
|
name: openspec-archive-change
|
||||||
|
description: Archive a completed change in the experimental workflow. Use when the user wants to finalize and archive a change after implementation is complete.
|
||||||
|
license: MIT
|
||||||
|
compatibility: Requires openspec CLI.
|
||||||
|
metadata:
|
||||||
|
author: openspec
|
||||||
|
version: "1.0"
|
||||||
|
generatedBy: "1.3.1"
|
||||||
|
---
|
||||||
|
|
||||||
|
Archive a completed change in the experimental workflow.
|
||||||
|
|
||||||
|
**Input**: Optionally specify a change name. If omitted, check if it can be inferred from conversation context. If vague or ambiguous you MUST prompt for available changes.
|
||||||
|
|
||||||
|
**Steps**
|
||||||
|
|
||||||
|
1. **If no change name provided, prompt for selection**
|
||||||
|
|
||||||
|
Run `openspec list --json` to get available changes. Use the **AskUserQuestion tool** to let the user select.
|
||||||
|
|
||||||
|
Show only active changes (not already archived).
|
||||||
|
Include the schema used for each change if available.
|
||||||
|
|
||||||
|
**IMPORTANT**: Do NOT guess or auto-select a change. Always let the user choose.
|
||||||
|
|
||||||
|
2. **Check artifact completion status**
|
||||||
|
|
||||||
|
Run `openspec status --change "<name>" --json` to check artifact completion.
|
||||||
|
|
||||||
|
Parse the JSON to understand:
|
||||||
|
- `schemaName`: The workflow being used
|
||||||
|
- `artifacts`: List of artifacts with their status (`done` or other)
|
||||||
|
|
||||||
|
**If any artifacts are not `done`:**
|
||||||
|
- Display warning listing incomplete artifacts
|
||||||
|
- Use **AskUserQuestion tool** to confirm user wants to proceed
|
||||||
|
- Proceed if user confirms
|
||||||
|
|
||||||
|
3. **Check task completion status**
|
||||||
|
|
||||||
|
Read the tasks file (typically `tasks.md`) to check for incomplete tasks.
|
||||||
|
|
||||||
|
Count tasks marked with `- [ ]` (incomplete) vs `- [x]` (complete).
|
||||||
|
|
||||||
|
**If incomplete tasks found:**
|
||||||
|
- Display warning showing count of incomplete tasks
|
||||||
|
- Use **AskUserQuestion tool** to confirm user wants to proceed
|
||||||
|
- Proceed if user confirms
|
||||||
|
|
||||||
|
**If no tasks file exists:** Proceed without task-related warning.
|
||||||
|
|
||||||
|
4. **Assess delta spec sync state**
|
||||||
|
|
||||||
|
Check for delta specs at `openspec/changes/<name>/specs/`. If none exist, proceed without sync prompt.
|
||||||
|
|
||||||
|
**If delta specs exist:**
|
||||||
|
- Compare each delta spec with its corresponding main spec at `openspec/specs/<capability>/spec.md`
|
||||||
|
- Determine what changes would be applied (adds, modifications, removals, renames)
|
||||||
|
- Show a combined summary before prompting
|
||||||
|
|
||||||
|
**Prompt options:**
|
||||||
|
- If changes needed: "Sync now (recommended)", "Archive without syncing"
|
||||||
|
- If already synced: "Archive now", "Sync anyway", "Cancel"
|
||||||
|
|
||||||
|
If user chooses sync, use Task tool (subagent_type: "general-purpose", prompt: "Use Skill tool to invoke openspec-sync-specs for change '<name>'. Delta spec analysis: <include the analyzed delta spec summary>"). Proceed to archive regardless of choice.
|
||||||
|
|
||||||
|
5. **Perform the archive**
|
||||||
|
|
||||||
|
Create the archive directory if it doesn't exist:
|
||||||
|
```bash
|
||||||
|
mkdir -p openspec/changes/archive
|
||||||
|
```
|
||||||
|
|
||||||
|
Generate target name using current date: `YYYY-MM-DD-<change-name>`
|
||||||
|
|
||||||
|
**Check if target already exists:**
|
||||||
|
- If yes: Fail with error, suggest renaming existing archive or using different date
|
||||||
|
- If no: Move the change directory to archive
|
||||||
|
|
||||||
|
```bash
|
||||||
|
mv openspec/changes/<name> openspec/changes/archive/YYYY-MM-DD-<name>
|
||||||
|
```
|
||||||
|
|
||||||
|
6. **Display summary**
|
||||||
|
|
||||||
|
Show archive completion summary including:
|
||||||
|
- Change name
|
||||||
|
- Schema that was used
|
||||||
|
- Archive location
|
||||||
|
- Whether specs were synced (if applicable)
|
||||||
|
- Note about any warnings (incomplete artifacts/tasks)
|
||||||
|
|
||||||
|
**Output On Success**
|
||||||
|
|
||||||
|
```
|
||||||
|
## Archive Complete
|
||||||
|
|
||||||
|
**Change:** <change-name>
|
||||||
|
**Schema:** <schema-name>
|
||||||
|
**Archived to:** openspec/changes/archive/YYYY-MM-DD-<name>/
|
||||||
|
**Specs:** ✓ Synced to main specs (or "No delta specs" or "Sync skipped")
|
||||||
|
|
||||||
|
All artifacts complete. All tasks complete.
|
||||||
|
```
|
||||||
|
|
||||||
|
**Guardrails**
|
||||||
|
- Always prompt for change selection if not provided
|
||||||
|
- Use artifact graph (openspec status --json) for completion checking
|
||||||
|
- Don't block archive on warnings - just inform and confirm
|
||||||
|
- Preserve .openspec.yaml when moving to archive (it moves with the directory)
|
||||||
|
- Show clear summary of what happened
|
||||||
|
- If sync is requested, use openspec-sync-specs approach (agent-driven)
|
||||||
|
- If delta specs exist, always run the sync assessment and show the combined summary before prompting
|
||||||
@@ -0,0 +1,288 @@
|
|||||||
|
---
|
||||||
|
name: openspec-explore
|
||||||
|
description: Enter explore mode - a thinking partner for exploring ideas, investigating problems, and clarifying requirements. Use when the user wants to think through something before or during a change.
|
||||||
|
license: MIT
|
||||||
|
compatibility: Requires openspec CLI.
|
||||||
|
metadata:
|
||||||
|
author: openspec
|
||||||
|
version: "1.0"
|
||||||
|
generatedBy: "1.3.1"
|
||||||
|
---
|
||||||
|
|
||||||
|
Enter explore mode. Think deeply. Visualize freely. Follow the conversation wherever it goes.
|
||||||
|
|
||||||
|
**IMPORTANT: Explore mode is for thinking, not implementing.** You may read files, search code, and investigate the codebase, but you must NEVER write code or implement features. If the user asks you to implement something, remind them to exit explore mode first and create a change proposal. You MAY create OpenSpec artifacts (proposals, designs, specs) if the user asks—that's capturing thinking, not implementing.
|
||||||
|
|
||||||
|
**This is a stance, not a workflow.** There are no fixed steps, no required sequence, no mandatory outputs. You're a thinking partner helping the user explore.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## The Stance
|
||||||
|
|
||||||
|
- **Curious, not prescriptive** - Ask questions that emerge naturally, don't follow a script
|
||||||
|
- **Open threads, not interrogations** - Surface multiple interesting directions and let the user follow what resonates. Don't funnel them through a single path of questions.
|
||||||
|
- **Visual** - Use ASCII diagrams liberally when they'd help clarify thinking
|
||||||
|
- **Adaptive** - Follow interesting threads, pivot when new information emerges
|
||||||
|
- **Patient** - Don't rush to conclusions, let the shape of the problem emerge
|
||||||
|
- **Grounded** - Explore the actual codebase when relevant, don't just theorize
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## What You Might Do
|
||||||
|
|
||||||
|
Depending on what the user brings, you might:
|
||||||
|
|
||||||
|
**Explore the problem space**
|
||||||
|
- Ask clarifying questions that emerge from what they said
|
||||||
|
- Challenge assumptions
|
||||||
|
- Reframe the problem
|
||||||
|
- Find analogies
|
||||||
|
|
||||||
|
**Investigate the codebase**
|
||||||
|
- Map existing architecture relevant to the discussion
|
||||||
|
- Find integration points
|
||||||
|
- Identify patterns already in use
|
||||||
|
- Surface hidden complexity
|
||||||
|
|
||||||
|
**Compare options**
|
||||||
|
- Brainstorm multiple approaches
|
||||||
|
- Build comparison tables
|
||||||
|
- Sketch tradeoffs
|
||||||
|
- Recommend a path (if asked)
|
||||||
|
|
||||||
|
**Visualize**
|
||||||
|
```
|
||||||
|
┌─────────────────────────────────────────┐
|
||||||
|
│ Use ASCII diagrams liberally │
|
||||||
|
├─────────────────────────────────────────┤
|
||||||
|
│ │
|
||||||
|
│ ┌────────┐ ┌────────┐ │
|
||||||
|
│ │ State │────────▶│ State │ │
|
||||||
|
│ │ A │ │ B │ │
|
||||||
|
│ └────────┘ └────────┘ │
|
||||||
|
│ │
|
||||||
|
│ System diagrams, state machines, │
|
||||||
|
│ data flows, architecture sketches, │
|
||||||
|
│ dependency graphs, comparison tables │
|
||||||
|
│ │
|
||||||
|
└─────────────────────────────────────────┘
|
||||||
|
```
|
||||||
|
|
||||||
|
**Surface risks and unknowns**
|
||||||
|
- Identify what could go wrong
|
||||||
|
- Find gaps in understanding
|
||||||
|
- Suggest spikes or investigations
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## OpenSpec Awareness
|
||||||
|
|
||||||
|
You have full context of the OpenSpec system. Use it naturally, don't force it.
|
||||||
|
|
||||||
|
### Check for context
|
||||||
|
|
||||||
|
At the start, quickly check what exists:
|
||||||
|
```bash
|
||||||
|
openspec list --json
|
||||||
|
```
|
||||||
|
|
||||||
|
This tells you:
|
||||||
|
- If there are active changes
|
||||||
|
- Their names, schemas, and status
|
||||||
|
- What the user might be working on
|
||||||
|
|
||||||
|
### When no change exists
|
||||||
|
|
||||||
|
Think freely. When insights crystallize, you might offer:
|
||||||
|
|
||||||
|
- "This feels solid enough to start a change. Want me to create a proposal?"
|
||||||
|
- Or keep exploring - no pressure to formalize
|
||||||
|
|
||||||
|
### When a change exists
|
||||||
|
|
||||||
|
If the user mentions a change or you detect one is relevant:
|
||||||
|
|
||||||
|
1. **Read existing artifacts for context**
|
||||||
|
- `openspec/changes/<name>/proposal.md`
|
||||||
|
- `openspec/changes/<name>/design.md`
|
||||||
|
- `openspec/changes/<name>/tasks.md`
|
||||||
|
- etc.
|
||||||
|
|
||||||
|
2. **Reference them naturally in conversation**
|
||||||
|
- "Your design mentions using Redis, but we just realized SQLite fits better..."
|
||||||
|
- "The proposal scopes this to premium users, but we're now thinking everyone..."
|
||||||
|
|
||||||
|
3. **Offer to capture when decisions are made**
|
||||||
|
|
||||||
|
| Insight Type | Where to Capture |
|
||||||
|
|----------------------------|--------------------------------|
|
||||||
|
| New requirement discovered | `specs/<capability>/spec.md` |
|
||||||
|
| Requirement changed | `specs/<capability>/spec.md` |
|
||||||
|
| Design decision made | `design.md` |
|
||||||
|
| Scope changed | `proposal.md` |
|
||||||
|
| New work identified | `tasks.md` |
|
||||||
|
| Assumption invalidated | Relevant artifact |
|
||||||
|
|
||||||
|
Example offers:
|
||||||
|
- "That's a design decision. Capture it in design.md?"
|
||||||
|
- "This is a new requirement. Add it to specs?"
|
||||||
|
- "This changes scope. Update the proposal?"
|
||||||
|
|
||||||
|
4. **The user decides** - Offer and move on. Don't pressure. Don't auto-capture.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## What You Don't Have To Do
|
||||||
|
|
||||||
|
- Follow a script
|
||||||
|
- Ask the same questions every time
|
||||||
|
- Produce a specific artifact
|
||||||
|
- Reach a conclusion
|
||||||
|
- Stay on topic if a tangent is valuable
|
||||||
|
- Be brief (this is thinking time)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Handling Different Entry Points
|
||||||
|
|
||||||
|
**User brings a vague idea:**
|
||||||
|
```
|
||||||
|
User: I'm thinking about adding real-time collaboration
|
||||||
|
|
||||||
|
You: Real-time collab is a big space. Let me think about this...
|
||||||
|
|
||||||
|
COLLABORATION SPECTRUM
|
||||||
|
════════════════════════════════════════════
|
||||||
|
|
||||||
|
Awareness Coordination Sync
|
||||||
|
│ │ │
|
||||||
|
▼ ▼ ▼
|
||||||
|
┌────────┐ ┌────────┐ ┌────────┐
|
||||||
|
│Presence│ │Cursors │ │ CRDT │
|
||||||
|
│ "3 │ │ Multi │ │Conflict│
|
||||||
|
│online" │ │ select │ │ free │
|
||||||
|
└────────┘ └────────┘ └────────┘
|
||||||
|
│ │ │
|
||||||
|
trivial moderate complex
|
||||||
|
|
||||||
|
Where's your head at?
|
||||||
|
```
|
||||||
|
|
||||||
|
**User brings a specific problem:**
|
||||||
|
```
|
||||||
|
User: The auth system is a mess
|
||||||
|
|
||||||
|
You: [reads codebase]
|
||||||
|
|
||||||
|
┌─────────────────────────────────────────────┐
|
||||||
|
│ CURRENT AUTH FLOW │
|
||||||
|
└─────────────────────────────────────────────┘
|
||||||
|
│
|
||||||
|
┌─────────────┼─────────────┐
|
||||||
|
▼ ▼ ▼
|
||||||
|
┌─────────┐ ┌─────────┐ ┌─────────┐
|
||||||
|
│ Google │ │ GitHub │ │ Email │
|
||||||
|
│ OAuth │ │ OAuth │ │ Magic │
|
||||||
|
└────┬────┘ └────┬────┘ └────┬────┘
|
||||||
|
│ │ │
|
||||||
|
└─────────────┼─────────────┘
|
||||||
|
▼
|
||||||
|
┌───────────┐
|
||||||
|
│ Session │
|
||||||
|
└─────┬─────┘
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
┌───────────┐
|
||||||
|
│ Perms │
|
||||||
|
└───────────┘
|
||||||
|
|
||||||
|
I see three tangles. Which one's burning?
|
||||||
|
```
|
||||||
|
|
||||||
|
**User is stuck mid-implementation:**
|
||||||
|
```
|
||||||
|
User: /opsx:explore add-auth-system
|
||||||
|
The OAuth integration is more complex than expected
|
||||||
|
|
||||||
|
You: [reads change artifacts]
|
||||||
|
|
||||||
|
You're on task 4: "Implement OAuth flow"
|
||||||
|
|
||||||
|
Let me trace what's involved...
|
||||||
|
|
||||||
|
[draws diagram, explores options, suggests paths]
|
||||||
|
|
||||||
|
Want to update the design to reflect this?
|
||||||
|
Or add a spike task to investigate?
|
||||||
|
```
|
||||||
|
|
||||||
|
**User wants to compare options:**
|
||||||
|
```
|
||||||
|
User: Should we use Postgres or SQLite?
|
||||||
|
|
||||||
|
You: Generic answer is boring. What's the context?
|
||||||
|
|
||||||
|
User: A CLI tool that tracks local dev environments
|
||||||
|
|
||||||
|
You: That changes everything.
|
||||||
|
|
||||||
|
┌─────────────────────────────────────────────────┐
|
||||||
|
│ CLI TOOL DATA STORAGE │
|
||||||
|
└─────────────────────────────────────────────────┘
|
||||||
|
|
||||||
|
Key constraints:
|
||||||
|
• No daemon running
|
||||||
|
• Must work offline
|
||||||
|
• Single user
|
||||||
|
|
||||||
|
SQLite Postgres
|
||||||
|
Deployment embedded ✓ needs server ✗
|
||||||
|
Offline yes ✓ no ✗
|
||||||
|
Single file yes ✓ no ✗
|
||||||
|
|
||||||
|
SQLite. Not even close.
|
||||||
|
|
||||||
|
Unless... is there a sync component?
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Ending Discovery
|
||||||
|
|
||||||
|
There's no required ending. Discovery might:
|
||||||
|
|
||||||
|
- **Flow into a proposal**: "Ready to start? I can create a change proposal."
|
||||||
|
- **Result in artifact updates**: "Updated design.md with these decisions"
|
||||||
|
- **Just provide clarity**: User has what they need, moves on
|
||||||
|
- **Continue later**: "We can pick this up anytime"
|
||||||
|
|
||||||
|
When it feels like things are crystallizing, you might summarize:
|
||||||
|
|
||||||
|
```
|
||||||
|
## What We Figured Out
|
||||||
|
|
||||||
|
**The problem**: [crystallized understanding]
|
||||||
|
|
||||||
|
**The approach**: [if one emerged]
|
||||||
|
|
||||||
|
**Open questions**: [if any remain]
|
||||||
|
|
||||||
|
**Next steps** (if ready):
|
||||||
|
- Create a change proposal
|
||||||
|
- Keep exploring: just keep talking
|
||||||
|
```
|
||||||
|
|
||||||
|
But this summary is optional. Sometimes the thinking IS the value.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Guardrails
|
||||||
|
|
||||||
|
- **Don't implement** - Never write code or implement features. Creating OpenSpec artifacts is fine, writing application code is not.
|
||||||
|
- **Don't fake understanding** - If something is unclear, dig deeper
|
||||||
|
- **Don't rush** - Discovery is thinking time, not task time
|
||||||
|
- **Don't force structure** - Let patterns emerge naturally
|
||||||
|
- **Don't auto-capture** - Offer to save insights, don't just do it
|
||||||
|
- **Do visualize** - A good diagram is worth many paragraphs
|
||||||
|
- **Do explore the codebase** - Ground discussions in reality
|
||||||
|
- **Do question assumptions** - Including the user's and your own
|
||||||
@@ -0,0 +1,110 @@
|
|||||||
|
---
|
||||||
|
name: openspec-propose
|
||||||
|
description: Propose a new change with all artifacts generated in one step. Use when the user wants to quickly describe what they want to build and get a complete proposal with design, specs, and tasks ready for implementation.
|
||||||
|
license: MIT
|
||||||
|
compatibility: Requires openspec CLI.
|
||||||
|
metadata:
|
||||||
|
author: openspec
|
||||||
|
version: "1.0"
|
||||||
|
generatedBy: "1.3.1"
|
||||||
|
---
|
||||||
|
|
||||||
|
Propose a new change - create the change and generate all artifacts in one step.
|
||||||
|
|
||||||
|
I'll create a change with artifacts:
|
||||||
|
- proposal.md (what & why)
|
||||||
|
- design.md (how)
|
||||||
|
- tasks.md (implementation steps)
|
||||||
|
|
||||||
|
When ready to implement, run /opsx:apply
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Input**: The user's request should include a change name (kebab-case) OR a description of what they want to build.
|
||||||
|
|
||||||
|
**Steps**
|
||||||
|
|
||||||
|
1. **If no clear input provided, ask what they want to build**
|
||||||
|
|
||||||
|
Use the **AskUserQuestion tool** (open-ended, no preset options) to ask:
|
||||||
|
> "What change do you want to work on? Describe what you want to build or fix."
|
||||||
|
|
||||||
|
From their description, derive a kebab-case name (e.g., "add user authentication" → `add-user-auth`).
|
||||||
|
|
||||||
|
**IMPORTANT**: Do NOT proceed without understanding what the user wants to build.
|
||||||
|
|
||||||
|
2. **Create the change directory**
|
||||||
|
```bash
|
||||||
|
openspec new change "<name>"
|
||||||
|
```
|
||||||
|
This creates a scaffolded change at `openspec/changes/<name>/` with `.openspec.yaml`.
|
||||||
|
|
||||||
|
3. **Get the artifact build order**
|
||||||
|
```bash
|
||||||
|
openspec status --change "<name>" --json
|
||||||
|
```
|
||||||
|
Parse the JSON to get:
|
||||||
|
- `applyRequires`: array of artifact IDs needed before implementation (e.g., `["tasks"]`)
|
||||||
|
- `artifacts`: list of all artifacts with their status and dependencies
|
||||||
|
|
||||||
|
4. **Create artifacts in sequence until apply-ready**
|
||||||
|
|
||||||
|
Use the **TodoWrite tool** to track progress through the artifacts.
|
||||||
|
|
||||||
|
Loop through artifacts in dependency order (artifacts with no pending dependencies first):
|
||||||
|
|
||||||
|
a. **For each artifact that is `ready` (dependencies satisfied)**:
|
||||||
|
- Get instructions:
|
||||||
|
```bash
|
||||||
|
openspec instructions <artifact-id> --change "<name>" --json
|
||||||
|
```
|
||||||
|
- The instructions JSON includes:
|
||||||
|
- `context`: Project background (constraints for you - do NOT include in output)
|
||||||
|
- `rules`: Artifact-specific rules (constraints for you - do NOT include in output)
|
||||||
|
- `template`: The structure to use for your output file
|
||||||
|
- `instruction`: Schema-specific guidance for this artifact type
|
||||||
|
- `outputPath`: Where to write the artifact
|
||||||
|
- `dependencies`: Completed artifacts to read for context
|
||||||
|
- Read any completed dependency files for context
|
||||||
|
- Create the artifact file using `template` as the structure
|
||||||
|
- Apply `context` and `rules` as constraints - but do NOT copy them into the file
|
||||||
|
- Show brief progress: "Created <artifact-id>"
|
||||||
|
|
||||||
|
b. **Continue until all `applyRequires` artifacts are complete**
|
||||||
|
- After creating each artifact, re-run `openspec status --change "<name>" --json`
|
||||||
|
- Check if every artifact ID in `applyRequires` has `status: "done"` in the artifacts array
|
||||||
|
- Stop when all `applyRequires` artifacts are done
|
||||||
|
|
||||||
|
c. **If an artifact requires user input** (unclear context):
|
||||||
|
- Use **AskUserQuestion tool** to clarify
|
||||||
|
- Then continue with creation
|
||||||
|
|
||||||
|
5. **Show final status**
|
||||||
|
```bash
|
||||||
|
openspec status --change "<name>"
|
||||||
|
```
|
||||||
|
|
||||||
|
**Output**
|
||||||
|
|
||||||
|
After completing all artifacts, summarize:
|
||||||
|
- Change name and location
|
||||||
|
- List of artifacts created with brief descriptions
|
||||||
|
- What's ready: "All artifacts created! Ready for implementation."
|
||||||
|
- Prompt: "Run `/opsx:apply` or ask me to implement to start working on the tasks."
|
||||||
|
|
||||||
|
**Artifact Creation Guidelines**
|
||||||
|
|
||||||
|
- Follow the `instruction` field from `openspec instructions` for each artifact type
|
||||||
|
- The schema defines what each artifact should contain - follow it
|
||||||
|
- Read dependency artifacts for context before creating new ones
|
||||||
|
- Use `template` as the structure for your output file - fill in its sections
|
||||||
|
- **IMPORTANT**: `context` and `rules` are constraints for YOU, not content for the file
|
||||||
|
- Do NOT copy `<context>`, `<rules>`, `<project_context>` blocks into the artifact
|
||||||
|
- These guide what you write, but should never appear in the output
|
||||||
|
|
||||||
|
**Guardrails**
|
||||||
|
- Create ALL artifacts needed for implementation (as defined by schema's `apply.requires`)
|
||||||
|
- Always read dependency artifacts before creating a new one
|
||||||
|
- If context is critically unclear, ask the user - but prefer making reasonable decisions to keep momentum
|
||||||
|
- If a change with that name already exists, ask if user wants to continue it or create a new one
|
||||||
|
- Verify each artifact file exists after writing before proceeding to next
|
||||||
+1
-1
@@ -65,8 +65,8 @@ uploads/
|
|||||||
|
|
||||||
### Windows / Runtime Artifacts
|
### Windows / Runtime Artifacts
|
||||||
*.stackdump
|
*.stackdump
|
||||||
NUL
|
|
||||||
|
|
||||||
### MVP Demo Generated Outputs
|
### MVP Demo Generated Outputs
|
||||||
mvp/demo/output/*.json
|
mvp/demo/output/*.json
|
||||||
!mvp/demo/output/README.md
|
!mvp/demo/output/README.md
|
||||||
|
.pi/extensions/emdash-hook.ts
|
||||||
|
|||||||
@@ -103,5 +103,48 @@ Find reuse opportunities + Trace the call/dependency chain and impact radius:
|
|||||||
## Baisc Infos
|
## Baisc Infos
|
||||||
|
|
||||||
Unless directly relevant to the user's current question, you should avoid proactively mentioning, illustrating, or
|
Unless directly relevant to the user's current question, you should avoid proactively mentioning, illustrating, or
|
||||||
trailing off into the following information in 99% of cases:
|
trailing off into the following information in 99% of cases:
|
||||||
|
|
||||||
|
<!-- gitnexus:start -->
|
||||||
|
# GitNexus — Code Intelligence
|
||||||
|
|
||||||
|
This project is indexed by GitNexus as **SuperBizAgent-java** (13483 symbols, 22230 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
|
||||||
|
|
||||||
|
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
|
||||||
|
|
||||||
|
## Always Do
|
||||||
|
|
||||||
|
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
|
||||||
|
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
|
||||||
|
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
|
||||||
|
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
|
||||||
|
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
|
||||||
|
|
||||||
|
## Never Do
|
||||||
|
|
||||||
|
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
|
||||||
|
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
|
||||||
|
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
|
||||||
|
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
|
||||||
|
|
||||||
|
## Resources
|
||||||
|
|
||||||
|
| Resource | Use for |
|
||||||
|
|----------|---------|
|
||||||
|
| `gitnexus://repo/SuperBizAgent-java/context` | Codebase overview, check index freshness |
|
||||||
|
| `gitnexus://repo/SuperBizAgent-java/clusters` | All functional areas |
|
||||||
|
| `gitnexus://repo/SuperBizAgent-java/processes` | All execution flows |
|
||||||
|
| `gitnexus://repo/SuperBizAgent-java/process/{name}` | Step-by-step execution trace |
|
||||||
|
|
||||||
|
## CLI
|
||||||
|
|
||||||
|
| Task | Read this skill file |
|
||||||
|
|------|---------------------|
|
||||||
|
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
|
||||||
|
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
|
||||||
|
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
|
||||||
|
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
|
||||||
|
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
|
||||||
|
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
|
||||||
|
|
||||||
|
<!-- gitnexus:end -->
|
||||||
|
|||||||
@@ -111,3 +111,46 @@ trailing off into the following information in 99% of cases:
|
|||||||
- 文档目录结构:
|
- 文档目录结构:
|
||||||
- 不要将文档放到用户目录(如 `C:\Users\EDY\.claude\`)中
|
- 不要将文档放到用户目录(如 `C:\Users\EDY\.claude\`)中
|
||||||
|
|
||||||
|
<!-- gitnexus:start -->
|
||||||
|
# GitNexus — Code Intelligence
|
||||||
|
|
||||||
|
This project is indexed by GitNexus as **SuperBizAgent-java** (13483 symbols, 22230 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
|
||||||
|
|
||||||
|
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
|
||||||
|
|
||||||
|
## Always Do
|
||||||
|
|
||||||
|
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
|
||||||
|
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
|
||||||
|
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
|
||||||
|
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
|
||||||
|
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
|
||||||
|
|
||||||
|
## Never Do
|
||||||
|
|
||||||
|
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
|
||||||
|
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
|
||||||
|
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
|
||||||
|
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
|
||||||
|
|
||||||
|
## Resources
|
||||||
|
|
||||||
|
| Resource | Use for |
|
||||||
|
|----------|---------|
|
||||||
|
| `gitnexus://repo/SuperBizAgent-java/context` | Codebase overview, check index freshness |
|
||||||
|
| `gitnexus://repo/SuperBizAgent-java/clusters` | All functional areas |
|
||||||
|
| `gitnexus://repo/SuperBizAgent-java/processes` | All execution flows |
|
||||||
|
| `gitnexus://repo/SuperBizAgent-java/process/{name}` | Step-by-step execution trace |
|
||||||
|
|
||||||
|
## CLI
|
||||||
|
|
||||||
|
| Task | Read this skill file |
|
||||||
|
|------|---------------------|
|
||||||
|
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
|
||||||
|
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
|
||||||
|
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
|
||||||
|
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
|
||||||
|
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
|
||||||
|
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
|
||||||
|
|
||||||
|
<!-- gitnexus:end -->
|
||||||
|
|||||||
@@ -175,9 +175,49 @@
|
|||||||
- 定义:Tool 调用及结果投影的生命周期状态,固定为 `PROJECTING`、`READY`、`ERROR`。
|
- 定义:Tool 调用及结果投影的生命周期状态,固定为 `PROJECTING`、`READY`、`ERROR`。
|
||||||
- 边界:它只说明调用记录是否完成,不说明结果是否包含证据。
|
- 边界:它只说明调用记录是否完成,不说明结果是否包含证据。
|
||||||
|
|
||||||
|
### Durable Audit
|
||||||
|
- 定义:为 Diagnosis Trace 长期保存的 Run、Agent 模型步骤和 Tool 调用元数据,用于 exact sessionId/runId 回放、评测和运维核对。
|
||||||
|
- 边界:只保存有界、脱敏、可长期保留的身份、状态、耗时、预算和结果摘要;不保存 Prompt、Thought、完整 Tool 参数、raw response 或 Redis canonical invocation。
|
||||||
|
|
||||||
### Evidence Status
|
### Evidence Status
|
||||||
- 定义:证据 Tool 的结果语义,固定为 `EVIDENCE_FOUND`、`NO_EVIDENCE`、`ERROR`。
|
- 定义:证据 Tool 的结果语义,固定为 `EVIDENCE_FOUND`、`NO_EVIDENCE`、`ERROR`。
|
||||||
- 边界:`NO_EVIDENCE` 只表示当前查询范围内没有匹配结果,不能解释为问题不存在、根因被排除或系统健康。
|
- 边界:`EVIDENCE_FOUND` 只表示存在候选内容,不保证内容能够支持当前诊断;`NO_EVIDENCE` 只表示当前查询范围内没有匹配结果,不能解释为问题不存在、根因被排除或系统健康。
|
||||||
|
|
||||||
|
### Information Gain
|
||||||
|
- 定义:一次 Tool 结果是否推进当前 Diagnosis Run 的语义评价,固定为 `GAINED` 或 `NO_GAIN`。
|
||||||
|
- 边界:它评价的是结果对当前诊断的作用,不评价 Tool 产品质量;`NO_EVIDENCE` 和重复的规范化 `tool + scope` 可由 Harness 机械标记为 `NO_GAIN`,其他成功非空结果(包括 RAG `REFERENCE`)由模型评价。
|
||||||
|
|
||||||
|
### Collection State
|
||||||
|
- 定义:Diagnosis Harness 对当前 Run 是否允许继续收集证据的控制状态,固定为 `COLLECTING` 或 `SATURATED`。
|
||||||
|
- 边界:状态由 Harness 维护;`SATURATED` 可因连续 `NO_GAIN` 或连续进展协议错误达到各自配置阈值而进入,不包括硬预算耗尽。模型可以请求新的 Tool 调用,但不能绕过 `SATURATED`。
|
||||||
|
|
||||||
|
### Diagnosis Stop Reason
|
||||||
|
- 定义:Harness 停止当前 Run 继续调用 Tool 的内部原因,首版区分 `INFORMATION_SATURATED`、`BUDGET_LIMIT_REACHED` 与 `PROGRESS_PROTOCOL_VIOLATED`。
|
||||||
|
- 边界:它用于控制、Trace 和 Release 输入,不是用户可见生命周期状态,也不进入模型上下文;真正的不可恢复技术故障走失败通道。协议错误不累计为 `NO_GAIN`,使用独立阈值与 stop reason。
|
||||||
|
|
||||||
|
### Progress Protocol Violation
|
||||||
|
- 定义:模型未遵守 Tool Call Envelope 进展协议时的安全错误分类,例如缺失/错序/意外 `previous_observation`、缺失 `input` 或非法 Envelope。
|
||||||
|
- 边界:返回可修正 observation(`repair_required`、`violation_type`、期望上一轮 Tool Call ID、允许的 `information_gain`);连续错误达到阈值后交付一次 `STOP_REQUIRED/PROGRESS_PROTOCOL_VIOLATED`。不泄露业务参数、上一轮观察正文、raw response 或内部异常。
|
||||||
|
|
||||||
|
### Progress Snapshot
|
||||||
|
- 定义:Tool Loop 结束时,从当前 Run 的 Canonical Tool Result 一次性投影出的有界发布视图,用于生成已检查范围和客观结果。
|
||||||
|
- 边界:Canonical Tool Result 是真理源;Progress Snapshot 不逐轮维护、不保存原始 Tool Response、Prompt 或内部 thought,也不直接进入模型上下文。
|
||||||
|
|
||||||
|
### Safe Fallback Type
|
||||||
|
- 定义:`SafeFallback.type` 对没有发布诊断结论的业务原因分类,例如 `INSUFFICIENT_EVIDENCE`、`MISSING_REQUIRED_CONTEXT`、`BUDGET_EXHAUSTED` 或安全校验失败。
|
||||||
|
- 边界:它是 `ReleaseOutcome.FALLBACK` 的原因字段,不是与 `SUCCESS / FALLBACK / FAILED / CANCELLED` 平行的第二套生命周期状态。
|
||||||
|
|
||||||
|
### Diagnosis Release Use Case
|
||||||
|
- 定义:诊断业务发布的唯一决策入口,接收 DiagnosisDraft 和/或 Harness `stop_reason + ProgressSnapshot`,生成安全的 `SUCCESS / FALLBACK` 结果。
|
||||||
|
- 边界:`conclusion=null` 不触发 EvidenceRepair;只有存在结论时才执行完整 EvidenceGuard、EvidenceRepair 和 SemanticGuard 链路。不可形成安全业务内容的技术故障由 Chat Application Use Case 映射为 `FAILED / CANCELLED`。
|
||||||
|
|
||||||
|
### Diagnosis Draft Contract Failure
|
||||||
|
- 定义:Diagnosis Agent 最终文本为空、不是严格 JSON,或不满足 `DiagnosisDraft` Schema 时产生的 Agent 输出合同失败。
|
||||||
|
- 边界:非法文本始终丢弃,不做 Markdown/自然语言提取,也不调用模型修复;仅当当前 Run 的 `ProgressSnapshot` 含已验真 observed facts 时,Release 才能确定性发布 `INSUFFICIENT_EVIDENCE`,否则保持 `FAILED`。它不是 `Diagnosis Stop Reason`,不得伪装成信息饱和或预算终止。
|
||||||
|
|
||||||
|
### Model Observation
|
||||||
|
- 定义:Tool 内部标准化结果经过白名单投影后,作为 Tool Response 进入 Diagnosis Agent 上下文的有界视图。
|
||||||
|
- 边界:只包含模型完成语义判断和证据引用所需的信息;预算、阈值、重复指纹、原始相似度、原始 Tool Response 和完整 Harness 控制状态不得进入该视图。
|
||||||
|
|
||||||
### RunContext
|
### RunContext
|
||||||
- 定义:一次 Diagnosis Run 的显式执行上下文,结构不可变地携带 `sessionId`、`runId`、deadline,以及该 Run 独占的取消、预算、重试策略和生命周期状态句柄。
|
- 定义:一次 Diagnosis Run 的显式执行上下文,结构不可变地携带 `sessionId`、`runId`、deadline,以及该 Run 独占的取消、预算、重试策略和生命周期状态句柄。
|
||||||
@@ -195,6 +235,14 @@
|
|||||||
- 定义:Harness 对同一技术操作 attempt 数和可重试失败类型的显式策略。
|
- 定义:Harness 对同一技术操作 attempt 数和可重试失败类型的显式策略。
|
||||||
- 边界:Router 与 SemanticGuard 的技术失败最多两次 attempt;Diagnosis Agent、Tool 和 Evidence repair 只有一次 attempt。Agent 正常 ReAct 轮次不是 retry,`NO_EVIDENCE`、业务拒绝、取消和预算耗尽不可重试。
|
- 边界:Router 与 SemanticGuard 的技术失败最多两次 attempt;Diagnosis Agent、Tool 和 Evidence repair 只有一次 attempt。Agent 正常 ReAct 轮次不是 retry,`NO_EVIDENCE`、业务拒绝、取消和预算耗尽不可重试。
|
||||||
|
|
||||||
|
### Chat Application Use Case
|
||||||
|
- 定义:一次 Chat 请求的唯一业务入口,拥有 Session/Run、意图路由、固定执行器、PreviousTurn 和最终持久化。
|
||||||
|
- 边界:不拥有 HTTP/SSE 连接,不把 ChatModel 或 Tool 选择权交给 Controller,也不在 Diagnosis Release Use Case 之外单独决定预算 Fallback 的业务内容。
|
||||||
|
|
||||||
|
### Chat SSE Contract
|
||||||
|
- 定义:Chat 公开入口的五事件协议,顺序固定为 `metadata -> status* -> content|failure -> done`。
|
||||||
|
- 边界:过程状态实时发送,最终安全内容最多释放一次;它不是 Token streaming,也不包含内部计划、Prompt、raw Tool 数据或异常。
|
||||||
|
|
||||||
### Verifier Skill Isolation
|
### Verifier Skill Isolation
|
||||||
- 定义:Chat Verifier 与 skill 系统隔离,只校验 Executor 答案和 `tool_trace_summary`。
|
- 定义:Chat Verifier 与 skill 系统隔离,只校验 Executor 答案和 `tool_trace_summary`。
|
||||||
- 使用场景:防止 Verifier 把 playbook 指令当作事实证据;Verifier 只判断已有证据是否支持结论。
|
- 使用场景:防止 Verifier 把 playbook 指令当作事实证据;Verifier 只判断已有证据是否支持结论。
|
||||||
|
|||||||
+18
-1
@@ -1,9 +1,23 @@
|
|||||||
# devflow 索引
|
# devflow 索引
|
||||||
|
|
||||||
|
## Issue 生命周期
|
||||||
|
|
||||||
|
| Issue | 状态 | 说明 |
|
||||||
|
|---|---|---|
|
||||||
|
| ISS-014 | archived | 阶段 0-7 的单体 Diagnosis Agent、Harness、ACI、SSE、清理和最终 E2E 已完成并归档;阶段实现对应的 11 个 devflow/OpenSpec 项目均已 archived。 |
|
||||||
|
| ISS-015 | active | 阶段 1 硬停止已由 ISS-016 收口;剩余 Evidence Repair Schema、Reasoning 审计验证/治理与最终综合验收。 |
|
||||||
|
| ISS-016 | archived | Diagnosis 信息增益停止契约、协议修复反馈与统一 Release 已完成并归档。 |
|
||||||
|
|
||||||
## 项目
|
## 项目
|
||||||
|
|
||||||
| 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|
| 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|
||||||
|---|---|---|---|---|---|---|
|
|---|---|---|---|---|---|---|
|
||||||
|
| 2026-07-28 | rag-eval-hybrid-baseline | 离线 RAG eval 对齐 hybrid:search.mode 生成器、fixture meta、baseline 重刷;hybrid 质量闸门可用 denseDistance。 | RAG/eval/baseline | search.mode, fixture meta, kb_scope rag-eval, L0 filter fallback, denseDistance | openspec/changes/archive/2026-07-28-rag-eval-hybrid-baseline | archived |
|
||||||
|
| 2026-07-28 | rag-quality-score-unify | 统一 dense/hybrid scoreLabel 与 qualityScore;保检索序;去掉关键词 boost 改序与 hybrid L2 伪装。 | RAG/质量分/后处理 | qualityScore, scoreLabel dense/hybrid, originalRank, RetrievalScoreNormalizer, no boost rerank | openspec/changes/archive/2026-07-28-rag-quality-score-unify | archived |
|
||||||
|
| 2026-07-26 | diagnosis-information-gain-stop-contract | Diagnosis 信息增益停止、协议修复反馈、ProgressSnapshot 与统一 Release。 | Harness/Diagnosis stop/Release | ISS-016, GAINED, NO_GAIN, STOP_REQUIRED, ProgressSnapshot, PROGRESS_PROTOCOL_VIOLATED, INSUFFICIENT_EVIDENCE | openspec/changes/archive/2026-07-27-diagnosis-information-gain-stop-contract | archived |
|
||||||
|
| 2026-07-27 | rag-chunk-evidence-identity-dedup | chunk 级证据身份、去重、retrieve-k/return-n 与 SearchPort 地基,为 hybrid 铺路。 | RAG/证据身份/去重 | evidenceKey, maxChunksPerDocument, retrieve-k, return-n, KnowledgeSearchPort, document_id chunk-scoped | openspec/changes/archive/2026-07-27-rag-chunk-evidence-identity-dedup | archived |
|
||||||
|
| 2026-07-27 | rag-bm25-hybrid-drop-sdk | 真 dense+BM25 hybrid(MilvusClientV2),废弃知识路径旧 SDK 检索/写入。 | RAG/BM25/hybrid | MilvusClientV2, BM25, hybridSearch, RRFRanker, biz_hybrid, drop SDK path | openspec/changes/archive/2026-07-27-rag-bm25-hybrid-drop-sdk | archived |
|
||||||
|
| 2026-07-27 | rag-hybrid-search-rrf | Delivery 2:可配置 hybrid 检索与 RRF 多路融合(不绑旧 SDK)。 | RAG/hybrid/RRF | hybrid mode, RRF, KnowledgeSearchPort, filtered+unfiltered fusion, sparse-lite lexical | openspec/changes/archive/2026-07-27-rag-hybrid-search-rrf | archived |
|
||||||
| 2026-07-21 | single-react-tool-invocation-store | 建立统一 ToolBoundary 与 Redis canonical invocation store,集中生命周期、证据状态、TTL、容量和 Run 所有权。 | Harness/Tool boundary/Canonical store | ISS-014, ToolBoundary, canonical invocation, PROJECTING, READY, ERROR, TTL, RESULT_TOO_LARGE | openspec/changes/archive/2026-07-21-single-react-tool-invocation-store | archived |
|
| 2026-07-21 | single-react-tool-invocation-store | 建立统一 ToolBoundary 与 Redis canonical invocation store,集中生命周期、证据状态、TTL、容量和 Run 所有权。 | Harness/Tool boundary/Canonical store | ISS-014, ToolBoundary, canonical invocation, PROJECTING, READY, ERROR, TTL, RESULT_TOO_LARGE | openspec/changes/archive/2026-07-21-single-react-tool-invocation-store | archived |
|
||||||
| 2026-07-21 | single-react-harness-run-context | 建立显式 RunContext、Harness Core、预算、取消、类型化重试和 Tool Store 基础。 | Harness/Run lifecycle/Budget | ISS-014, RunContext, deadline, cancellation, budget, retry, ToolCallKey | openspec/changes/archive/2026-07-21-single-react-harness-run-context | archived |
|
| 2026-07-21 | single-react-harness-run-context | 建立显式 RunContext、Harness Core、预算、取消、类型化重试和 Tool Store 基础。 | Harness/Run lifecycle/Budget | ISS-014, RunContext, deadline, cancellation, budget, retry, ToolCallKey | openspec/changes/archive/2026-07-21-single-react-harness-run-context | archived |
|
||||||
| 2026-07-21 | single-react-aci-tool-contracts | 冻结 RAG、日志和 MySQL evidence Tool 的 Agent-facing ACI Schema、状态、框架调用引用和描述边界。 | Harness/Agent Tool contract | ISS-014, ACI, tool_call_id, evidence_status, RAG, query_logs, query_mysql, MOCK | openspec/changes/archive/2026-07-21-single-react-aci-tool-contracts | archived |
|
| 2026-07-21 | single-react-aci-tool-contracts | 冻结 RAG、日志和 MySQL evidence Tool 的 Agent-facing ACI Schema、状态、框架调用引用和描述边界。 | Harness/Agent Tool contract | ISS-014, ACI, tool_call_id, evidence_status, RAG, query_logs, query_mysql, MOCK | openspec/changes/archive/2026-07-21-single-react-aci-tool-contracts | archived |
|
||||||
@@ -41,3 +55,6 @@
|
|||||||
| 2026-07-21 | single-react-mysql-readonly-tool | Fail-closed read-only MySQL evidence Tool with AST allowlist, JDBC controls and bounded projection | Harness/MySQL security | ISS-014, MySQL, JSqlParser, allowlist, PreparedStatement, timeout, projection | openspec/changes/archive/2026-07-21-single-react-mysql-readonly-tool | archived |
|
| 2026-07-21 | single-react-mysql-readonly-tool | Fail-closed read-only MySQL evidence Tool with AST allowlist, JDBC controls and bounded projection | Harness/MySQL security | ISS-014, MySQL, JSqlParser, allowlist, PreparedStatement, timeout, projection | openspec/changes/archive/2026-07-21-single-react-mysql-readonly-tool | archived |
|
||||||
| 2026-07-21 | single-react-diagnosis-agent | Single internal Diagnosis ReactAgent with Harness-controlled model/tool loop, bounded context and typed Draft | Harness/Diagnosis Agent/ReAct | ISS-014, ReactAgent, DiagnosisDraft, PreviousTurn, ToolInterceptor, ModelInterceptor, budget | openspec/changes/archive/2026-07-21-single-react-diagnosis-agent | archived |
|
| 2026-07-21 | single-react-diagnosis-agent | Single internal Diagnosis ReactAgent with Harness-controlled model/tool loop, bounded context and typed Draft | Harness/Diagnosis Agent/ReAct | ISS-014, ReactAgent, DiagnosisDraft, PreviousTurn, ToolInterceptor, ModelInterceptor, budget | openspec/changes/archive/2026-07-21-single-react-diagnosis-agent | archived |
|
||||||
| 2026-07-21 | single-react-evidence-semantic-guards | Deterministic evidence validation, isolated semantic review and fail-closed diagnosis release | Harness/EvidenceGuard/SemanticGuard/Release | ISS-014, EvidenceGuard, verified snapshot, SemanticGuard, repair, fallback, release policy | openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards | archived |
|
| 2026-07-21 | single-react-evidence-semantic-guards | Deterministic evidence validation, isolated semantic review and fail-closed diagnosis release | Harness/EvidenceGuard/SemanticGuard/Release | ISS-014, EvidenceGuard, verified snapshot, SemanticGuard, repair, fallback, release policy | openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards | archived |
|
||||||
|
| 2026-07-21 | single-react-chat-application-usecase | Internal Chat application use case with isolated routing, fixed executors and safe PreviousTurn | Harness/Chat application/Run persistence | ISS-014, Intent Router, PreviousTurn, PublishedResult, V012, observer, cancellation | openspec/changes/archive/2026-07-21-single-react-chat-application-usecase | archived |
|
||||||
|
| 2026-07-21 | single-react-chat-sse-cutover | Unique named-event Chat SSE endpoint, bounded production Harness wiring and strict frontend consumer | Chat/SSE/Harness production wiring | ISS-014, /api/chat, SSE, metadata, status, content, failure, done, disconnect, bounded executor | openspec/changes/archive/2026-07-22-single-react-chat-sse-cutover | archived |
|
||||||
|
| 2026-07-22 | single-react-cleanup-e2e | Remove legacy Agent paths, add bounded Harness audit, and complete exact-run live acceptance | Chat/Harness/cleanup/E2E | ISS-014, single ReAct Agent, durable audit, named SSE, exact run, Flyway V013 | openspec/changes/archive/2026-07-22-single-react-cleanup-e2e | archived |
|
||||||
|
|||||||
@@ -10,7 +10,7 @@
|
|||||||
|
|
||||||
- `devflow/index.md` was checked. Relevant history includes `session-storage`, `confidence-feedback`, `executor-action-memory-relevance`, and `chat-verifier-agent`.
|
- `devflow/index.md` was checked. Relevant history includes `session-storage`, `confidence-feedback`, `executor-action-memory-relevance`, and `chat-verifier-agent`.
|
||||||
- `mvp/archive/2026-07-09-doc-cleanup/notes/agent-engineering-decisions.md` already recommends the next phase as "可复现 MVP Demo", including `mvp-demo` profile, fixed diagnosis case, one-click request, and `GET /api/diagnosis/{sessionId}/trace`.
|
- `mvp/archive/2026-07-09-doc-cleanup/notes/agent-engineering-decisions.md` already recommends the next phase as "可复现 MVP Demo", including `mvp-demo` profile, fixed diagnosis case, one-click request, and `GET /api/diagnosis/{sessionId}/trace`.
|
||||||
- `mvp/issues/active/ISS-003-mvp-design-implementation-review.md` identifies test stability, session traceability, verifier evidence chain, upload path, and SupervisorAgent consistency as recent MVP concerns. Security cleanup is intentionally deferred by user decision.
|
- `mvp/issues/archived/ISS-003-mvp-design-implementation-review.md` identifies test stability, session traceability, verifier evidence chain, upload path, and SupervisorAgent consistency as historical MVP concerns. Security cleanup was intentionally deferred by user decision at that time.
|
||||||
|
|
||||||
## Question Pool
|
## Question Pool
|
||||||
|
|
||||||
|
|||||||
@@ -2,7 +2,7 @@
|
|||||||
|
|
||||||
## Draft Acceptance
|
## Draft Acceptance
|
||||||
|
|
||||||
- [x] Issue exists: `mvp/issues/active/executor-evidence-attribution-hallucination.md`.
|
- [x] Issue exists: `mvp/issues/archived/executor-evidence-attribution-hallucination.md`.
|
||||||
- [x] OpenSpec change artifacts exist.
|
- [x] OpenSpec change artifacts exist.
|
||||||
- [x] devflow tracking files exist.
|
- [x] devflow tracking files exist.
|
||||||
- [x] OpenSpec validation passes.
|
- [x] OpenSpec validation passes.
|
||||||
|
|||||||
@@ -0,0 +1,42 @@
|
|||||||
|
# Acceptance: single-react-chat-application-usecase
|
||||||
|
|
||||||
|
## Result
|
||||||
|
|
||||||
|
- Status: archived
|
||||||
|
- OpenSpec tasks: 14/14 complete
|
||||||
|
- Interface impact: L3 database/collaboration
|
||||||
|
- Public protocol: unchanged
|
||||||
|
|
||||||
|
## Static Verification
|
||||||
|
|
||||||
|
- V012 migration、`DiagnosisRun` 和 `DiagnosisRunRepository` 对齐 `intent/release_outcome/published_result` 及安全 PreviousTurn filter。
|
||||||
|
- 公开 Controller、前端和 endpoint diff 为空。
|
||||||
|
- Router/System executor 未检出 Tool、ReactAgent、ThreadLocal 或手写循环。
|
||||||
|
- `PublishedResult` 固定为 `user_query/published_conclusion/scope/limitations/source_documents`;序列化负向测试覆盖内部字段泄漏。
|
||||||
|
|
||||||
|
## Script Verification
|
||||||
|
|
||||||
|
- `mvn -q -DskipTests compile`:通过。
|
||||||
|
- Stage 6A focused `ApplicationExecutorsTest,PublishedResultPersistenceTest,ChatApplicationUseCaseTest`:13 tests,通过。
|
||||||
|
- Stage 2-5 与 6A regression selection:18 suites / 76 tests,0 failure/error/skipped。
|
||||||
|
- `openspec validate single-react-chat-application-usecase --strict`:通过。
|
||||||
|
|
||||||
|
## Browser or Manual Verification
|
||||||
|
|
||||||
|
- Not applicable。阶段 6A 没有 UI 或公开入口变化。
|
||||||
|
|
||||||
|
## Not Verified
|
||||||
|
|
||||||
|
- 未运行真实 LLM、Redis、日志和 MySQL live E2E;按 ISS-014 串行门禁统一留到阶段 7。
|
||||||
|
- V012 未在本阶段连接真实数据库执行;migration/entity/query 已由静态检查、focused persistence tests 和 compile 覆盖。
|
||||||
|
|
||||||
|
## Remaining Work
|
||||||
|
|
||||||
|
- 阶段 6B:唯一 `POST /api/chat` SSE 原子切换、旧 endpoint 删除和前端消费者迁移。
|
||||||
|
- 阶段 7:旧链路清理、全局 spec 格式修复和最终 live E2E。
|
||||||
|
|
||||||
|
## Archive
|
||||||
|
|
||||||
|
- `.archive-ready`: created
|
||||||
|
- OpenSpec archive: `openspec/changes/archive/2026-07-21-single-react-chat-application-usecase`
|
||||||
|
- Main spec sync: `openspec/specs/single-react-chat-application-usecase/spec.md`
|
||||||
@@ -0,0 +1,31 @@
|
|||||||
|
# Brief: single-react-chat-application-usecase
|
||||||
|
|
||||||
|
## Background
|
||||||
|
|
||||||
|
阶段 2-5 已具备 RunContext、单一 Diagnosis Agent 和安全释放门禁,但没有统一应用用例拥有 Session/Run、意图路由、PreviousTurn、固定执行器和最终持久化,阶段 6B 因而无法只做协议切换。
|
||||||
|
|
||||||
|
## Goal
|
||||||
|
|
||||||
|
在不改变公开 Chat/SSE 行为的前提下,建立内部 `ChatApplicationUseCase`,统一三类意图、同一 Run 生命周期、安全 PreviousTurn 和 typed public content。
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
- 无 Tool、无记忆、无 ReAct 的三分类 Intent Router。
|
||||||
|
- SYSTEM_CHAT、KNOWLEDGE_QUERY、DIAGNOSIS 固定执行器。
|
||||||
|
- Knowledge exact invocation/document reference validation。
|
||||||
|
- 同 Session 最近安全 Diagnosis SUCCESS 的有界 PreviousTurn。
|
||||||
|
- `diagnosis_run` V012 字段、JPA store 和安全 `PublishedResult`。
|
||||||
|
- Protocol-neutral observer、Run control、终态持久化和 focused tests。
|
||||||
|
|
||||||
|
## Non-goals
|
||||||
|
|
||||||
|
- 不修改 Controller、`/api/chat`、`/api/chat_stream`、SSE schema 或前端消费者。
|
||||||
|
- 不删除旧 ChatService、多 Agent、ThreadLocal 或旧 session storage。
|
||||||
|
- 不运行真实模型、Redis、日志和 MySQL live E2E;统一留到阶段 7。
|
||||||
|
|
||||||
|
## Metadata
|
||||||
|
|
||||||
|
- Scale: complex
|
||||||
|
- Interface impact: L3 database/collaboration
|
||||||
|
- OpenSpec: `single-react-chat-application-usecase`
|
||||||
|
- Parent issue: `ISS-014`
|
||||||
@@ -0,0 +1,117 @@
|
|||||||
|
# Decisions: single-react-chat-application-usecase
|
||||||
|
|
||||||
|
## Discover Status
|
||||||
|
|
||||||
|
- Checkpoint: Discover
|
||||||
|
- Capability source: `sm-flow` + `grill-with-docs`;`codebase-retrieval`、LSP 和 GitNexus MCP 当前不可用,使用既有 GitNexus 结论、`rg` 引用核对和源码阅读降级。
|
||||||
|
- Scale: complex。跨模型路由、三类执行器、Run/session 生命周期、数据库 migration、PreviousTurn 和阶段 6B consumer boundary。
|
||||||
|
- `devflow/index.md` 命中阶段 0-5、session-run-trace-isolation、RAG contracts 和 release guards;无 ADR 冲突。
|
||||||
|
|
||||||
|
## Question Pool
|
||||||
|
|
||||||
|
| # | 维度 | 问题 | 模式 | 状态 |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| Q1 | 术语 | Chat Application Use Case 与 Controller、Harness、Agent 的职责边界是什么? | evidence-driven | 已解决 |
|
||||||
|
| Q2 | 路由 | Router 输入、输出和可重试失败范围是什么? | evidence-driven | 已解决 |
|
||||||
|
| Q3 | 路由 | Router 最终失败是否允许默认进入 Diagnosis? | evidence-driven | 已解决 |
|
||||||
|
| Q4 | 执行器 | 三种 intent 分别允许哪些模型和 Tool 行为? | evidence-driven | 已解决 |
|
||||||
|
| Q5 | Knowledge | 单次 RAG 调用如何验证模型引用且不公开 Tool Call ID? | evidence-driven | 已解决 |
|
||||||
|
| Q6 | PreviousTurn | 上一回合的真理源、筛选条件和截断边界是什么? | evidence-driven | 已解决 |
|
||||||
|
| Q7 | 生命周期 | 何时读取上一回合、创建当前 Run、记录 intent 和终态? | evidence-driven | 已解决 |
|
||||||
|
| Q8 | 持久化 | diagnosis_run 需要新增哪些字段,哪些内部内容禁止进入 published_result? | evidence-driven | 已解决 |
|
||||||
|
| Q9 | 6B 边界 | 如何让 Controller 切换时不重写应用用例? | evidence-driven | 已解决 |
|
||||||
|
| Q10 | 接口 | 数据库/内部接口影响等级和回滚要求是什么? | evidence-driven | 已解决 |
|
||||||
|
| Q11 | 验收 | 如何证明原始 Query、sessionId/runId 和失败终态一致传播? | evidence-driven | 已解决 |
|
||||||
|
|
||||||
|
## Evidence-driven
|
||||||
|
|
||||||
|
| 结论 | 证据来源 | 是否已汇报用户 |
|
||||||
|
|---|---|---|
|
||||||
|
| Controller 只负责协议;Application Use Case 拥有 session/run、routing、executor、persistence,Harness 拥有预算/取消/释放,Agent 拥有诊断语义。 | ISS-014 3、4.2、阶段 6A/6B | 已汇报 |
|
||||||
|
| Router 输入只含 query/last_intent/last_user_query,输出只允许三枚举;无 Tool/记忆/ReAct。 | ISS-014 4.5 | 已汇报 |
|
||||||
|
| timeout/transport/非法输出可重试一次;第二次失败返回安全入口错误,不进入 Diagnosis。 | ISS-014 重试策略、`HarnessRetryPolicies.intentRouter()` | 已汇报 |
|
||||||
|
| SYSTEM_CHAT 无 Tool;KNOWLEDGE_QUERY 只调用一次 lookup;DIAGNOSIS 进入 Agent + Guards。 | ISS-014 4.5 | 已汇报 |
|
||||||
|
| PreviousTurn 只来自同 Session 最近 `DIAGNOSIS + SUCCESS + published_result`,不使用 Redis 历史。 | ISS-014 4.6 | 已汇报 |
|
||||||
|
| `PublishedResult`/`PreviousTurn` 已冻结为 query/conclusion/scope/limitations/source_documents,不含 Tool ID/raw/Draft/reason。 | 阶段 0 contracts、`PublishedResult`、`PreviousTurn` | 已汇报 |
|
||||||
|
| 当前 `DiagnosisRun`/V011 尚无 intent/release_outcome/published_result,需要 V012 和 repository query。 | `DiagnosisRun.java`、`V011__add_session_run_isolation.sql` | 已汇报 |
|
||||||
|
| 上一回合必须在保存当前 PENDING Run 前读取,否则 latest query 会命中当前请求。 | repository 当前 latest method + 生命周期顺序推导 | 已汇报 |
|
||||||
|
| 6B 需要 metadata/status/cancel,6A 应提供 observer 和显式 RunContext,而不包含 SSE 类型。 | ISS-014 4.7、阶段 6B | 已汇报 |
|
||||||
|
|
||||||
|
## User-interview
|
||||||
|
|
||||||
|
- 无新增 user-interview。三类路由、数据库字段、PreviousTurn、失败语义、阶段边界和自动 Apply/Archive/commit 均由 ISS-014 与用户持续授权冻结。
|
||||||
|
|
||||||
|
## Key Decisions
|
||||||
|
|
||||||
|
- 在创建当前 Run 前读取 latest safe routing context 和 PreviousTurn,随后 `startRun -> persist RUNNING -> observer metadata -> route`。
|
||||||
|
- Application output 使用 typed content union,Diagnosis success 转为无 Tool ID 的 public report view;Fallback output 不携带完整 snapshot。
|
||||||
|
- Knowledge path 生成一次 direct canonical Tool Call ID(该路径没有框架 Tool Call),只调用 `lookup_knowledge`,严格验证 `KnowledgeAnswerDraft` 的 exact call ID 和 document subset,再移除 ID 发布。
|
||||||
|
- 所有普通单轮模型调用复用阶段 5 的受控 `GuardModelCall`,从而共享 Core 模型/Token/timeout/cancel 边界;不引入新模型路由。
|
||||||
|
- `PublishedResult` 仅在 Diagnosis SUCCESS 且 conclusion 非空时写入;Fallback/Failed/Cancelled/System/Knowledge 不生成 PreviousTurn 真理源。
|
||||||
|
- 数据库变更使用 V012 可前向迁移;回滚为先停止新应用用例,再删除新索引/列,不影响 V011 既有字段。
|
||||||
|
- 不创建 ADR:这些是 ISS-014 已冻结设计的落地,不是新的跨项目不可逆决策。
|
||||||
|
|
||||||
|
## OpenSpec Backfill
|
||||||
|
|
||||||
|
- 需进入 proposal/design/spec/tasks:路由隔离/重试、三执行器、original query、PreviousTurn filter/bounds、observer/cancel、Run terminal persistence、V012/L3、公开隔离。
|
||||||
|
- 非目标:Controller/SSE/前端切换、旧链路删除、live E2E。
|
||||||
|
|
||||||
|
## Cross-artifact Alignment
|
||||||
|
|
||||||
|
| 上游 -> 下游 | 检查内容 | 状态 |
|
||||||
|
|---|---|---|
|
||||||
|
| ISS-014/brief -> proposal | 三路由、previous turn、Run lifecycle、内部-only 和阶段 6B handoff | 已对齐 |
|
||||||
|
| proposal -> design | typed executors、observer、JPA/V012、异常/终态、L3 migration/rollback | 已对齐 |
|
||||||
|
| design -> specs/tasks | 每项所有权/安全边界均有可观察 requirement 和实现测试切片 | 已对齐 |
|
||||||
|
| specs -> tasks | 9 组 requirements 覆盖 Router/executors、store/policy、application、verification | 已对齐 |
|
||||||
|
|
||||||
|
## Architecture Audit
|
||||||
|
|
||||||
|
- 能力来源:`zoom-out`,使用 Chat Session、Diagnosis Run、RunContext、Diagnosis Agent、EvidenceGuard、SemanticGuard 和 PublishedResult 术语。
|
||||||
|
- 链路为 `request -> application use case -> run store/router -> fixed executor -> Harness/Agent/Tool -> typed public content -> run finish`;Controller 不拥有模型/工具/Run。
|
||||||
|
- ChatRunStore 拥有 MySQL 映射,Core 拥有运行状态,Application 拥有 dispatch/终态,path executor 拥有单一路径行为;PreviousTurn policy 是唯一安全历史投影。
|
||||||
|
- 最大风险是 prior/current Run 顺序和 DB/lifecycle 双终态,design/tasks 已固定 prior read before start、single finish path 和 focused failure/cancel tests。
|
||||||
|
- V012 是 L3 additive schema;migration、entity、repository、rollback 独立章节完整,公开入口阶段 6A 零变化。
|
||||||
|
|
||||||
|
## Interface Impact
|
||||||
|
|
||||||
|
- 级别:L3 database/collaboration interface。
|
||||||
|
- 新增 `diagnosis_run.intent/release_outcome/published_result` 和索引;修改 Entity/Repository,新增 internal application/store/output contract。
|
||||||
|
- 消费者:阶段 6B Controller/SSE adapter、MySQL/Flyway;旧 ChatService 在本阶段不消费新字段。
|
||||||
|
- 迁移/回滚:V012 nullable additive;回滚先切旧入口,再删除 index/columns。
|
||||||
|
|
||||||
|
## Commit Gate Preflight
|
||||||
|
|
||||||
|
- proposal、design、specs、tasks 完整,`openspec status` complete,change strict validation 通过。
|
||||||
|
- Question pool 全部已解决并汇报,无 user-interview、未判级接口或未接受架构风险。
|
||||||
|
- Cross-artifact 四段对齐无 gap;V012/L3、prior read ordering、terminal persistence 和 6B handoff 已进入 design/spec/tasks。
|
||||||
|
- Apply/Archive/commit 使用用户持续授权;公开协议和前端必须保持零 diff。
|
||||||
|
- `.committed` 已创建,可进入 Apply。
|
||||||
|
|
||||||
|
## Pre-apply Research
|
||||||
|
|
||||||
|
- `DiagnosisHarnessCore`/`RunContext`:Run ID、budget、cancel 和 first-terminal-wins。
|
||||||
|
- `GuardModelCall`/`HarnessRetryExecutor`:单轮模型 timeout/usage 和 Router 两次 attempt。
|
||||||
|
- `HarnessEvidenceTools`/`RagToolResult`:Knowledge 唯一 lookup 路径和有界 projection。
|
||||||
|
- `DiagnosisAgentUseCase`/`DiagnosisReleaseUseCase`:Diagnosis Draft 与安全 release boundary。
|
||||||
|
- `DiagnosisRun`/`DiagnosisRunRepository`/`V011`:现有 Run schema 和写入模式。
|
||||||
|
- `ChatService.ensureChatSession/startDiagnosisRun`:只参考 session metadata/JPA 写法,不复用旧 routing、多 Agent 或 ThreadLocal。
|
||||||
|
- 技术栈:Spring AI direct Prompt、Jackson strict JSON、JPA repository、Flyway additive migration、protocol-neutral observer;无 MQ/新依赖。
|
||||||
|
|
||||||
|
## Apply Progress
|
||||||
|
|
||||||
|
- Router/System/Knowledge contracts 与 executors 已完成,tasks 1.1-1.4 完成。
|
||||||
|
- 5 个 `ApplicationExecutorsTest` 通过:同输入 retry、最终 routing failure、System direct call、Knowledge exact references/no ID、NO_EVIDENCE/model skip。
|
||||||
|
- TODO:PublishedResult/JPA/V012、Diagnosis executor、总应用用例和综合验证。
|
||||||
|
- REVIEW:修正 `JpaChatRunStore` 多构造器 Spring 注入歧义;prior/start/intent/finish 持久化异常统一为稳定 `RUN_PERSISTENCE_FAILED`,并在安全完成时更新 ChatSession 活跃时间/消息对数。均为代码偏离修复,无需变更 OpenSpec。
|
||||||
|
- PublishedResult/JPA/V012、Diagnosis executor、protocol-neutral Run control 与总 ChatApplicationUseCase 已完成,tasks 2.1-3.4 完成。
|
||||||
|
- 13 个 stage 6A focused tests 通过;TODO 仅剩综合回归、static scope 和 OpenSpec verification。
|
||||||
|
- 综合回归曾在 `SemanticGuardTest.attemptTimeoutCancelsBothPermittedModelCalls` 出现负载相关失败。诊断确认生产代码对每次 timeout 均调用 `Future.cancel(true)`,但第二个 Future 可能在任务线程启动前已取消,此时不存在可接收 interrupt 的线程。分类为测试假设偏差,不是 OpenSpec 或生产代码偏离;回归断言改为两次 TIMEOUT attempt、两次模型预算预留,以及至少一个已运行调用收到 interrupt。
|
||||||
|
|
||||||
|
## Final Review
|
||||||
|
|
||||||
|
- Stage 6A focused tests 与阶段 2-5 regression 共 18 suites / 76 tests,0 failure/error/skipped;Maven compile 通过。
|
||||||
|
- OpenSpec strict validation 通过;V012、Entity、Repository 的三个字段和 previous-turn filter 对齐。
|
||||||
|
- 公开 Controller、前端和 endpoint 零 diff;Router/System executor 无 Tool、ReactAgent、ThreadLocal 或手写 loop。
|
||||||
|
- `PublishedResult` 只包含 `user_query/published_conclusion/scope/limitations/source_documents`,负向序列化测试通过。
|
||||||
|
- 本阶段不运行 live E2E,按 ISS-014 门禁留到阶段 7。
|
||||||
@@ -0,0 +1,31 @@
|
|||||||
|
# Evidence: single-react-chat-application-usecase
|
||||||
|
|
||||||
|
## Code and Contract Evidence
|
||||||
|
|
||||||
|
- `DiagnosisHarnessCore`/`RunContext` 已提供 Run ID、预算、取消和 first-terminal-wins,应用用例无需创建第二套生命周期。
|
||||||
|
- `GuardModelCall`/`HarnessRetryExecutor` 提供单轮模型 timeout、Token 记账和两次 Router attempt。
|
||||||
|
- `HarnessEvidenceTools`/`RagToolResult` 提供 Knowledge 路径唯一 lookup 和有界 projection。
|
||||||
|
- `DiagnosisAgentUseCase`/`DiagnosisReleaseUseCase` 已形成 Diagnosis Draft 与安全发布边界。
|
||||||
|
- `DiagnosisRun`/`DiagnosisRunRepository`/V011 提供既有 Run 持久化,V012 以 nullable additive 字段扩展。
|
||||||
|
|
||||||
|
## Confirmed Boundaries
|
||||||
|
|
||||||
|
- Router 输入只含原始 Query、可选 last intent 和 last user query;最终失败不得默认进入 Diagnosis。
|
||||||
|
- 三类 executor 不互相调用,所有路径接收未改写 Query。
|
||||||
|
- PreviousTurn 只来自同 Session 最近 `DIAGNOSIS + SUCCESS + published_result`,且必须在保存当前 Run 前读取。
|
||||||
|
- Knowledge 公开内容只保留 stable document metadata,不发布 direct Tool Call ID。
|
||||||
|
- `PublishedResult` 不保存 Tool ID、raw evidence、完整 Draft 或 SemanticGuard reason;Fallback/Failed/Cancelled 不写安全历史。
|
||||||
|
- Controller/SSE/前端切换属于阶段 6B,本阶段保持公开协议不变。
|
||||||
|
|
||||||
|
## Diagnosis Finding
|
||||||
|
|
||||||
|
- 综合回归暴露 `SemanticGuardTest` 的负载竞态:第二个 Future 可能在获得线程前被取消,因而不会产生第二次 interrupt。
|
||||||
|
- 生产代码已对每次 timeout 调用 `Future.cancel(true)`;测试改为验证两次 TIMEOUT attempt、两次模型预算预留,以及至少一个运行中调用被中断。
|
||||||
|
- 分类为测试假设偏差,不是生产代码或 OpenSpec 偏离。
|
||||||
|
|
||||||
|
## Verification Evidence
|
||||||
|
|
||||||
|
- Stage 6A focused tests:13 tests 通过。
|
||||||
|
- Stage 2-5 与 6A 综合回归:18 suites / 76 tests,0 failure/error/skipped。
|
||||||
|
- Maven compile 与 change strict validation 通过。
|
||||||
|
- 公开 Controller/前端零 diff;Router/System executor 无 Tool、ReactAgent、ThreadLocal 或手写 loop。
|
||||||
@@ -0,0 +1,49 @@
|
|||||||
|
# Acceptance: single-react-chat-sse-cutover
|
||||||
|
|
||||||
|
## Result
|
||||||
|
|
||||||
|
- Status: archived
|
||||||
|
- OpenSpec tasks: 14/14 complete
|
||||||
|
- Interface impact: L4 breaking HTTP/frontend contract
|
||||||
|
- Public Chat protocol: unique named-event SSE `POST /api/chat`
|
||||||
|
|
||||||
|
## Static Verification
|
||||||
|
|
||||||
|
- 组合路由保持 `/api/chat`、`/api/ai_ops`、`/api/chat/clear`、`/api/chat/session/{sessionId}` 和 `/runs`;`/api/chat_stream` 已删除。
|
||||||
|
- 生产 Controller/前端 legacy Chat token scan:0。
|
||||||
|
- `ChatController` forbidden dependency scan:0;结构测试固定其三个 protocol/application dependencies。
|
||||||
|
- SSE payload 只包含固定 metadata/status/content/failure/done fields,`CANCELLED` 不作为公开 done outcome。
|
||||||
|
- mode selector DOM、state、consumer 和 CSS 已删除。
|
||||||
|
|
||||||
|
## Script Verification
|
||||||
|
|
||||||
|
- `mvn -q -DskipTests compile`:通过。
|
||||||
|
- Stage 2-6B + AiOps regression selection:24 suites / 101 tests,0 failure/error/skipped。
|
||||||
|
- `openspec validate single-react-chat-sse-cutover --strict`:通过。
|
||||||
|
- `node --check src/main/resources/static/app.js`:通过。
|
||||||
|
- `git diff --check`:通过,仅有仓库既存 LF/CRLF 提示。
|
||||||
|
|
||||||
|
## Browser or Manual Verification
|
||||||
|
|
||||||
|
- 本阶段未运行浏览器人工验证;前端协议由静态 contract test 和 JavaScript syntax check 覆盖。
|
||||||
|
|
||||||
|
## Not Verified
|
||||||
|
|
||||||
|
- 未运行 live 模型、Redis、日志和 MySQL E2E;按 ISS-014 阶段门禁统一留到阶段 7。
|
||||||
|
- 未验证外部第三方 Chat API consumer;L4 变更不提供兼容分支,外部消费者必须同步迁移到 named-event SSE。
|
||||||
|
|
||||||
|
## Migration and Rollback
|
||||||
|
|
||||||
|
- 部署必须将后端 SSE endpoint 与 bundled frontend consumer 作为同一版本原子发布。
|
||||||
|
- 回滚必须同时回滚 Controller 和 frontend 到阶段 6A commit;V012 additive nullable migration 可保留。
|
||||||
|
- 不允许通过恢复 `/api/chat_stream`、同步 JSON consumer 或旧 message wrapper 形成双轨兼容。
|
||||||
|
|
||||||
|
## Remaining Work
|
||||||
|
|
||||||
|
- 阶段 7:物理删除旧 Agent/Graph/Hook/ThreadLocal/ChatService 路径和过时测试,更新文档并完成最终 live E2E、日志与数据库核验。
|
||||||
|
|
||||||
|
## Archive
|
||||||
|
|
||||||
|
- `.archive-ready`: created
|
||||||
|
- OpenSpec archive: `openspec/changes/archive/2026-07-22-single-react-chat-sse-cutover`
|
||||||
|
- Main spec sync: `openspec/specs/single-react-chat-sse-cutover/spec.md` (10 requirements added)
|
||||||
@@ -0,0 +1,32 @@
|
|||||||
|
# Brief: single-react-chat-sse-cutover
|
||||||
|
|
||||||
|
## Background
|
||||||
|
|
||||||
|
阶段 6A 已建立 protocol-neutral `ChatApplicationUseCase`,但公开 Chat 仍有同步 `/api/chat` 和伪流式 `/api/chat_stream` 两条旧链路,Controller 直接拥有模型、Tools、Session 历史和无界线程池,前端也保留两套消费者。
|
||||||
|
|
||||||
|
## Goal
|
||||||
|
|
||||||
|
将公开 Chat 原子切换为唯一 `POST /api/chat` named-event SSE,并让 Controller 只承担校验、HTTP/SSE 和连接生命周期;所有公开内容必须来自阶段 6A 的安全释放结果。
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
- 唯一 `/api/chat` SSE 与 `metadata -> status* -> content|failure -> done` 状态机。
|
||||||
|
- exact Run disconnect/timeout/send-failure cancellation。
|
||||||
|
- Spring-managed bounded Chat/model executors 和完整 Harness production Bean graph。
|
||||||
|
- 前端唯一 named-event consumer、typed renderer 和 metadata identity 保存。
|
||||||
|
- AiOps 模型/Tool acquisition 下沉到 service,并将 AiOps/Session endpoint 从 Chat Controller 职责中隔离。
|
||||||
|
- L4 前后端迁移、成对回滚和 focused regression。
|
||||||
|
|
||||||
|
## Non-goals
|
||||||
|
|
||||||
|
- 不做 Token streaming、最终答案切片、断线续传、事件重放、轮询或 WebSocket。
|
||||||
|
- 不修改 `/api/ai_ops` 的公开 URL、请求和 SSE message-wrapper 行为。
|
||||||
|
- 不在本阶段物理删除旧多 Agent、ChatService、Hook 或 ThreadLocal;阶段 7 统一清理。
|
||||||
|
- 不运行 live 模型、Redis、日志或 MySQL E2E;阶段 7 统一验收。
|
||||||
|
|
||||||
|
## Metadata
|
||||||
|
|
||||||
|
- Scale: complex
|
||||||
|
- Interface impact: L4 breaking HTTP/frontend contract
|
||||||
|
- OpenSpec: `single-react-chat-sse-cutover`
|
||||||
|
- Parent issue: `ISS-014`
|
||||||
@@ -0,0 +1,110 @@
|
|||||||
|
# Decisions: single-react-chat-sse-cutover
|
||||||
|
|
||||||
|
## Discover Status
|
||||||
|
|
||||||
|
- Checkpoint: Discover
|
||||||
|
- Capability source: `sm-flow` + `grill-with-docs`;`codebase-retrieval`、LSP 和 GitNexus MCP 当前不可用,使用 `rg` 引用核对、源码阅读和 focused tests 降级。
|
||||||
|
- Scale: complex。涉及 L4 HTTP/SSE 协议、前端消费者、异步连接生命周期、生产 Bean 装配和阶段 7 删除边界。
|
||||||
|
- `devflow/index.md` 命中阶段 0-6A、ISS-013 和 session-run-trace-isolation;ISS-014 是更新且已冻结的最终协议源。
|
||||||
|
|
||||||
|
## Question Pool
|
||||||
|
|
||||||
|
| # | 维度 | 问题 | 模式 | 状态 |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| Q1 | 术语 | “真正 SSE”是 Token streaming,还是过程事件实时 + 最终内容一次释放? | evidence-driven | 已解决 |
|
||||||
|
| Q2 | 协议 | 唯一 endpoint、事件名称、payload、顺序、互斥和终态是什么? | evidence-driven | 已解决 |
|
||||||
|
| Q3 | 边界 | Controller、Application Use Case、Harness 和 SSE adapter 各自拥有何种职责? | evidence-driven | 已解决 |
|
||||||
|
| Q4 | 生命周期 | disconnect/timeout/send failure 如何取消同一个 Run,正常 complete 如何避免误取消? | evidence-driven | 已解决 |
|
||||||
|
| Q5 | 执行器 | 如何消除 Controller 自建无界线程池并处理饱和? | evidence-driven | 已解决 |
|
||||||
|
| Q6 | 装配 | 阶段 2-6A plain Java components 如何形成可启动的生产 Bean graph? | evidence-driven | 已解决 |
|
||||||
|
| Q7 | 前端 | 快速/流式双模式如何迁移到唯一 SSE consumer? | evidence-driven | 已解决 |
|
||||||
|
| Q8 | 安全 | 哪些内部状态/内容禁止进入 SSE? | evidence-driven | 已解决 |
|
||||||
|
| Q9 | 兼容 | 是否保留同步 `/api/chat` 或 `/api/chat_stream` 兼容? | evidence-driven | 已解决 |
|
||||||
|
| Q10 | 范围 | `/api/ai_ops` 和旧多 Agent 何时处理? | evidence-driven | 已解决 |
|
||||||
|
| Q11 | 验收 | 如何证明 event order、single terminal、same IDs、cancel 和 production wiring? | evidence-driven | 已解决 |
|
||||||
|
|
||||||
|
## Evidence-driven
|
||||||
|
|
||||||
|
| 结论 | 证据来源 | 是否已汇报用户 |
|
||||||
|
|---|---|---|
|
||||||
|
| 真正 SSE 固定为状态实时发送、最终 typed content 一次释放,不做 Token/字符切片。 | ISS-014 4.7 | 已汇报 |
|
||||||
|
| 唯一入口是 `POST /api/chat`;顺序为 `metadata -> status* -> content|failure -> done`。 | ISS-014 4.7、阶段 6B | 已汇报 |
|
||||||
|
| metadata/status/content/failure/done 使用 named SSE event;payload 不再包含旧 `type/data` 包装。 | ISS-014 最小事件结构 | 已汇报 |
|
||||||
|
| 当前 Controller 同时拥有旧 ChatService、模型、Tools、SessionManager、同步/伪流式流程和 cached thread pool。 | `ChatController.java` | 已汇报 |
|
||||||
|
| 当前前端 quick 调 `/chat` JSON,stream 调 `/chat_stream` 并保留大量旧格式 fallback。 | `static/app.js` | 已汇报 |
|
||||||
|
| `ChatApplicationUseCase` 已提供 observer、同一 run control、typed content 和 stable failure code,但组件尚无完整生产 Bean graph。 | 阶段 6A code + Spring annotation scan | 已汇报 |
|
||||||
|
| 客户端断开必须通过 observer 得到的 `ChatRunControl` 取消,同一引用需覆盖断开早于 onStarted 的竞态。 | 阶段 6A design + `SseEmitter` lifecycle | 已汇报 |
|
||||||
|
| `/api/ai_ops` 是独立公开入口,不属于阶段 6B Chat 原子切换;旧实现阶段 7 清理/处置。 | ISS-014 阶段 6B/7 | 已汇报 |
|
||||||
|
|
||||||
|
## User-interview
|
||||||
|
|
||||||
|
- 无新增 user-interview。唯一 endpoint、破坏性迁移、事件 schema、非 Token 流、自动 Apply/Archive/commit 均由 ISS-014 与用户持续授权冻结。
|
||||||
|
|
||||||
|
## Key Decisions
|
||||||
|
|
||||||
|
- Controller 使用构造注入的 `ChatApplicationUseCase` 与受控 `TaskExecutor`;不注入 ChatModel、Tools、ChatService 或 SessionManager 来处理 Chat。
|
||||||
|
- SSE adapter 为每个请求维护单一 state machine 和 `AtomicReference<ChatRunControl>`;disconnect 标记先于 onStarted 时,onStarted 立即取消。
|
||||||
|
- 正常 result 只产生一次 typed content 和 done;异常只产生 failure 和 done;IOException/timeout/disconnect 只取消,不尝试补发终态。
|
||||||
|
- executor rejection 在没有 Run 时发送稳定 failure/done;worker 启动后所有 failure code 来自 `ChatApplicationException`,不暴露 cause。
|
||||||
|
- 生产装配使用集中 `harness.chat` properties 和 Spring-managed bounded executors;所有模型调用复用同一 `ChatModel` 和 `GuardModelCall`。
|
||||||
|
- 前端移除 mode selector 和 quick path,只保留一个严格 named-event parser;unknown event/schema fail closed。
|
||||||
|
- 不创建 ADR:L4 方案已经在 ISS-014 设计冻结,本 change 负责原子落地和迁移说明。
|
||||||
|
|
||||||
|
## OpenSpec Backfill
|
||||||
|
|
||||||
|
- 需进入 proposal/design/spec/tasks:唯一 SSE、五事件 state machine、typed payload、安全 failure、same IDs、disconnect cancellation、bounded executors、production assembly、frontend migration、L4 rollback。
|
||||||
|
- 非目标:AiOps 协议、旧类物理删除、live E2E。
|
||||||
|
|
||||||
|
## Cross-artifact Alignment
|
||||||
|
|
||||||
|
| 上游 -> 下游 | 检查内容 | 状态 |
|
||||||
|
|---|---|---|
|
||||||
|
| ISS-014/brief -> proposal | 唯一 SSE、五事件、断开取消、前端迁移、生产装配、阶段 7 边界 | 已对齐 |
|
||||||
|
| proposal -> design | named events、state machine、bounded executors、Bean graph、AiOps 隔离、L4 rollback | 已对齐 |
|
||||||
|
| design -> specs/tasks | 每项 ownership/lifecycle/safety/migration 均有可观察 requirement 和纵向切片 | 已对齐 |
|
||||||
|
| specs -> tasks | 10 组 requirements 覆盖 wiring、SSE、cancel、frontend、AiOps regression 和 verification | 已对齐 |
|
||||||
|
|
||||||
|
## Architecture Audit
|
||||||
|
|
||||||
|
- Capability source: `zoom-out`,使用 Chat Application Use Case、RunContext、Run Lifecycle、Diagnosis Harness 和 Chat SSE Contract 术语。
|
||||||
|
- 链路为 `browser -> Controller -> SSE session -> bounded worker -> Application Use Case -> Harness -> typed release -> SSE session`;业务真理源不进入 Controller。
|
||||||
|
- SSE session 独占协议状态,Application 独占 Run/dispatch/persistence,Core 独占 cancel/budget,configuration 独占 infrastructure graph。
|
||||||
|
- 最大风险是 disconnect/onStarted 与 terminal callback 竞态,design/tasks 已固定 pending-disconnect、atomic terminal 和 no-send-after-close tests。
|
||||||
|
- L4 部署必须 Controller/frontend 同版本;回滚成对返回阶段 6A commit,V012 可保留。
|
||||||
|
|
||||||
|
## Interface Impact
|
||||||
|
|
||||||
|
- Level: L4 breaking HTTP/frontend contract。
|
||||||
|
- 删除同步 JSON `POST /api/chat`、`POST /api/chat_stream` 和旧 Chat SSE wrapper;新增 named-event SSE `POST /api/chat`。
|
||||||
|
- 消费者:bundled `static/app.js`、外部 Chat callers、Controller/MockMvc tests;Trace/feedback 继续使用 metadata IDs。
|
||||||
|
- 迁移/回滚:前后端同 commit 原子部署/回滚;不提供兼容开关、双 endpoint 或旧 parser。
|
||||||
|
|
||||||
|
## Commit Gate Preflight
|
||||||
|
|
||||||
|
- proposal、design、specs、tasks 完整,change strict validation 通过。
|
||||||
|
- Question pool 全部 evidence-driven 解决并已汇报,无 user-interview、未判级接口或未接受架构风险。
|
||||||
|
- Cross-artifact 四段对齐无 gap;L4 migration/rollback、production wiring、disconnect race 和阶段 7 边界均进入 design/spec/tasks。
|
||||||
|
- `openspec-propose` 补全产物,`zoom-out` 完成架构审计;可提交为 Committed OpenSpec。
|
||||||
|
|
||||||
|
## Apply Progress
|
||||||
|
|
||||||
|
- 1.1-1.4 完成:新增 `ChatHarnessProperties`、bounded worker/model executors 和显式 `HarnessChatConfiguration` graph;空 MySQL datasource 只允许 fail-closed unknown logical ID。
|
||||||
|
- 2.1-2.4 完成:新增 named-event `ChatSseEvent`/`ChatSseSession`,Chat 唯一 SSE Controller,移除 `/chat_stream` 和同步 Chat;AiOps 模型/Tool 依赖下沉到 service。
|
||||||
|
- 3.1-3.3 完成:前端移除 quick/stream 双轨,统一 named-event parser、typed renderer 和静态 contract test。
|
||||||
|
- `JpaChatRunStore` 生产构造器改为注入集中 `PublishedResultPolicy`,避免读写边界漂移;这是实现偏差修复,无需改需求方向。
|
||||||
|
|
||||||
|
## Apply Review
|
||||||
|
|
||||||
|
- Review 发现 `ChatController` 虽然 Chat path 已经只调用应用用例,但类本身仍承载 AiOps 和 Session 管理依赖,不完全满足“Chat Controller 只负责 Chat 协议”的 ownership 要求。
|
||||||
|
- 用户确认将职责拆分为 `ChatController`(仅 `/api/chat`)、`AiOpsController`(保持 `/api/ai_ops`)和 `ChatSessionController`(保持 clear/session/runs URL)。
|
||||||
|
- 该修正保持所有公开 URL、AiOps message-wrapper payload 和 Session observable behavior,不修改 Committed OpenSpec 的范围或方向。
|
||||||
|
- `styles.css` 中 mode selector/dropdown 死样式随前端双轨删除一并移除。
|
||||||
|
- AGENTS.md 指定的 `codebase-retrieval` 和 LSP 工具在当前环境不可用;使用 OpenSpec 全量上下文、`rg` 引用检查、Java 编译、结构测试和综合回归完成等价影响面确认。
|
||||||
|
|
||||||
|
## Verification and Migration
|
||||||
|
|
||||||
|
- Stage 2-6B + AiOps regression:24 suites / 101 tests,0 failure/error/skipped;包含 `HarnessChatConfigurationTest` Spring wiring。
|
||||||
|
- `mvn -q -DskipTests compile`、`openspec validate single-react-chat-sse-cutover --strict`、`node --check src/main/resources/static/app.js` 均通过。
|
||||||
|
- 生产 Controller/前端遗留 token 静态扫描为 0;`ChatController` 中模型、Tool、ChatService、AiOps、Session 和 Trace service 禁止依赖扫描为 0。
|
||||||
|
- L4 部署必须将后端 `/api/chat` SSE 与 bundled frontend consumer 同版本原子部署;回滚必须成对回滚到阶段 6A commit,不提供兼容 endpoint、旧 parser 或双轨开关。
|
||||||
|
- live 模型、Redis、日志和 MySQL E2E 按 ISS-014 门禁明确延后到阶段 7,本阶段没有把 focused/Mock 验证表述为 live 验收。
|
||||||
@@ -0,0 +1,32 @@
|
|||||||
|
# Evidence: single-react-chat-sse-cutover
|
||||||
|
|
||||||
|
## Code and Contract Evidence
|
||||||
|
|
||||||
|
- `ChatApplicationUseCase` 已提供 typed result、observer 和 exact `ChatRunControl`,Controller 无需拥有模型、Tool 或业务路由。
|
||||||
|
- `ChatSseSession` 使用 first-terminal-wins 状态和 pending disconnect,覆盖断开早于 `onStarted`、send failure、late terminal 与正常 completion 竞态。
|
||||||
|
- `ChatSseEvent` 固定 metadata/status/content/failure/done payload;content 引用阶段 6A typed public content,failure 只公开 stable code/message。
|
||||||
|
- `HarnessChatConfiguration` 组装同一个 Core、model boundary、canonical store、ToolBoundary、Diagnosis Agent、Guards、Router、executors 和 application use case。
|
||||||
|
- `ChatHarnessProperties` 集中 worker/model queue、Run、SSE、canonical store、Agent、Router、Guard 和 single-turn limits;executors 使用有限队列与 `AbortPolicy`。
|
||||||
|
- bundled frontend 只向 `/api/chat` 发起 streaming POST,按完整 named SSE frame 严格解析并 fail closed。
|
||||||
|
|
||||||
|
## Ownership Review
|
||||||
|
|
||||||
|
- Apply review 将原 Controller 拆为 `ChatController`、`AiOpsController` 和 `ChatSessionController`。
|
||||||
|
- `ChatController` 只注入 `ChatApplicationUseCase`、bounded worker 和 Chat properties;禁止依赖扫描为 0。
|
||||||
|
- `/api/ai_ops`、`/api/chat/clear`、session info 和 run list 的 URL 与 payload 字段保持不变。
|
||||||
|
- mode selector/dropdown DOM、JS state 和 CSS 已全部移除,不保留双轨开关。
|
||||||
|
|
||||||
|
## Safety and Lifecycle Evidence
|
||||||
|
|
||||||
|
- success 测试验证 `metadata,status,content,done` 严格顺序和 typed content。
|
||||||
|
- failure 测试验证内部 provider detail 不进入 failure payload,且 content/failure 互斥。
|
||||||
|
- disconnect-before-start 与 send-failure 测试验证 exact Run 取消和 late terminal 阻断。
|
||||||
|
- frontend contract test 验证唯一 request target、五类 event branch 与旧 consumer token 删除。
|
||||||
|
- static scan:生产 Controller/前端 legacy Chat token 0;Chat Controller forbidden dependency 0。
|
||||||
|
|
||||||
|
## Verification Evidence
|
||||||
|
|
||||||
|
- Stage 2-6B + AiOps regression:24 suites / 101 tests,0 failure/error/skipped。
|
||||||
|
- `HarnessChatConfigurationTest` 覆盖 bounded queues、共享依赖、空 MySQL fail-closed 和 Spring context wiring。
|
||||||
|
- Maven compile、strict OpenSpec validation 和 JavaScript syntax check 通过。
|
||||||
|
- 真实模型、Redis、日志和 MySQL live E2E 按 ISS-014 串行门禁留到阶段 7。
|
||||||
@@ -6,7 +6,7 @@
|
|||||||
- 当前问题:Harness Core 和三类 evidence Tool 已就绪,但没有一个内部诊断执行链消费它们;公开 Chat 仍依赖旧的简单/多 Agent 路径。
|
- 当前问题:Harness Core 和三类 evidence Tool 已就绪,但没有一个内部诊断执行链消费它们;公开 Chat 仍依赖旧的简单/多 Agent 路径。
|
||||||
- 关联 OpenSpec:`openspec/changes/archive/2026-07-21-single-react-diagnosis-agent/`
|
- 关联 OpenSpec:`openspec/changes/archive/2026-07-21-single-react-diagnosis-agent/`
|
||||||
- devflow 分档:complex
|
- devflow 分档:complex
|
||||||
- 需求真理源:`mvp/issues/active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md`,不重复创建独立 PRD。
|
- 需求真理源:`mvp/issues/archived/ISS-014-single-react-agent-harness-aci-ptk-refactor.md`,不重复创建独立 PRD。
|
||||||
|
|
||||||
## 范围
|
## 范围
|
||||||
|
|
||||||
|
|||||||
@@ -4,7 +4,7 @@
|
|||||||
|
|
||||||
| 来源 | 证据 | 结论 | 是否已汇报 |
|
| 来源 | 证据 | 结论 | 是否已汇报 |
|
||||||
|---|---|---|---|
|
|---|---|---|---|
|
||||||
| `mvp/issues/active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md` | 阶段 4 明确单 Agent、内部入口、极简上下文、结构化 Draft、无重试和预算验收 | 本 change 不得切换公开入口或删除旧链路 | 是 |
|
| `mvp/issues/archived/ISS-014-single-react-agent-harness-aci-ptk-refactor.md` | 阶段 4 明确单 Agent、内部入口、极简上下文、结构化 Draft、无重试和预算验收 | 本 change 不得切换公开入口或删除旧链路 | 是 |
|
||||||
| `ChatService.java` + GitNexus references | `createReactAgent` 被旧策略入口调用;复杂路径仍创建 Planner/Executor/Verifier/Composer | 新实现必须是独立内部 use case,不能复用旧 Service 运行职责 | 是 |
|
| `ChatService.java` + GitNexus references | `createReactAgent` 被旧策略入口调用;复杂路径仍创建 Planner/Executor/Verifier/Composer | 新实现必须是独立内部 use case,不能复用旧 Service 运行职责 | 是 |
|
||||||
| Spring AI Alibaba 1.1.2.0 `ReactAgent`/`AgentLlmNode` sources | 框架自带 ReAct loop;非流式 `ModelResponse` 保留 `ChatResponse` Usage | 不手写循环,使用 ModelInterceptor 强制模型/Token 预算 | 是 |
|
| Spring AI Alibaba 1.1.2.0 `ReactAgent`/`AgentLlmNode` sources | 框架自带 ReAct loop;非流式 `ModelResponse` 保留 `ChatResponse` Usage | 不手写循环,使用 ModelInterceptor 强制模型/Token 预算 | 是 |
|
||||||
| Spring AI Alibaba `ToolCallRequest`/`AgentToolNode` sources | `ToolCallRequest.getToolCallId()` 来自 `AssistantMessage.ToolCall.id()`,Tool interceptor 在 callback 前执行 | 可以精确传播框架 ID,不生成第二套 ID | 是 |
|
| Spring AI Alibaba `ToolCallRequest`/`AgentToolNode` sources | `ToolCallRequest.getToolCallId()` 来自 `AssistantMessage.ToolCall.id()`,Tool interceptor 在 callback 前执行 | 可以精确传播框架 ID,不生成第二套 ID | 是 |
|
||||||
|
|||||||
@@ -0,0 +1,57 @@
|
|||||||
|
# Acceptance: single-react-cleanup-e2e
|
||||||
|
|
||||||
|
## Result
|
||||||
|
|
||||||
|
Accepted. 阶段 7 的清理、Harness-native audit、文档收口和 live E2E 已完成;最终发布为安全 `FALLBACK`,没有释放未验证 Draft。
|
||||||
|
|
||||||
|
## Static Verification
|
||||||
|
|
||||||
|
- `openspec validate --all --strict`: 22/22 passed。
|
||||||
|
- `node --check src/main/resources/static/app.js`: passed。
|
||||||
|
- `git diff --check`: passed。
|
||||||
|
- production legacy path、旧 Agent-facing Tool contract 和持久化敏感 payload 扫描:0 matches。
|
||||||
|
- 最终 Tool 时间窗日志扫描:原始 query、request/response、Tool Call ID、Prompt、stack、debug instrumentation 均为 0;只保留有界 metadata。
|
||||||
|
|
||||||
|
## Script Verification
|
||||||
|
|
||||||
|
- Deterministic Maven regression: 60 suites / 221 tests,0 failures,0 errors,3 skipped。
|
||||||
|
- `mvn -q -DskipTests compile`: passed。
|
||||||
|
- `mvn -q -DskipTests package`: passed。
|
||||||
|
- `mvn -q -Dtest=QueryLogsToolsTest,QueryLogsResultProjectorTest,CanonicalInvocationStoreTest test`: passed。
|
||||||
|
- Flyway 9.22.3 repair + migrate: V012 checksum repaired,V013 applied,schema current version 013。
|
||||||
|
- Maven `spring-boot:run` with `mvp-demo`: Flyway 13 migrations validated,JPA schema validation passed,Tomcat 9900 started,Harness/Audit wiring active。
|
||||||
|
- `run-payment-timeout-demo.ps1`: strict named SSE validation passed for exact final session/run。
|
||||||
|
- `scripts/query_mysql.py`: exact diagnosis_run、agent_step、tool_invocation queries passed。
|
||||||
|
|
||||||
|
## Final Live Evidence
|
||||||
|
|
||||||
|
- sessionId: `mvp-demo-payment-timeout-stage7-20260722-1741`
|
||||||
|
- runId: `363f481c-33b8-42e7-8699-428a6ec61806`
|
||||||
|
- SSE: `metadata -> status -> status -> status -> content -> done`
|
||||||
|
- diagnosis_run: `status=SUCCESS`、`intent=DIAGNOSIS`、`release_outcome=FALLBACK`、`tool_call_count=2`、`total_token_count=11764`。
|
||||||
|
- AgentStep: 2 exact-run rows,唯一 agent `diagnosis_agent`,`thought IS NULL`,只含角色/数量/Tool name metadata。
|
||||||
|
- ToolInvocation: `lookup_knowledge` 和 Mock `query_logs` 各 1 行,均 `READY/EVIDENCE_FOUND/success=1`;input/details 只有 framework Tool Call ID、状态和字节数。
|
||||||
|
|
||||||
|
## Browser / Manual Verification
|
||||||
|
|
||||||
|
- 未执行浏览器点击验收;本阶段公开协议由真实 HTTP SSE 脚本和前端 JavaScript syntax/contract tests 覆盖。
|
||||||
|
|
||||||
|
## Remaining Risks
|
||||||
|
|
||||||
|
- live 结果为 EvidenceGuard/SemanticGuard 约束下的安全 `FALLBACK`,不是业务根因成功发布;这是允许的最终释放状态。
|
||||||
|
- `diagnosis_run.step_count` 仍为空,但 exact AgentStep 查询返回 2 行;该历史汇总字段不作为本阶段 release gate。
|
||||||
|
- `query_logs` 使用 Mock;真实 CLS 与生产业务 `query_mysql` datasource 仍属于后续接入范围。
|
||||||
|
- `application-local.yml` 含本地内部配置且被 Git 忽略;未 stage、未提交、未输出凭据。
|
||||||
|
|
||||||
|
## Migration And Rollback
|
||||||
|
|
||||||
|
- L4 endpoint/frontend cleanup 必须整体回滚阶段 7 commit,不恢复双轨 endpoint 或旧 Tool annotations。
|
||||||
|
- V013 仅幂等增加缺失列/索引;应用回滚时保留新增 nullable 列,避免破坏已写数据,不执行 destructive down migration。
|
||||||
|
- Flyway repair 已将远端 V012 checksum 对齐当前迁移;V013 保证旧库与 fresh database 最终 schema 一致。
|
||||||
|
|
||||||
|
## Archive
|
||||||
|
|
||||||
|
- OpenSpec archive: completed at `openspec/changes/archive/2026-07-22-single-react-cleanup-e2e`。
|
||||||
|
- Delta specs synced: created main `single-react-cleanup-e2e` spec and removed the temporary AiOps-preservation requirement from `single-react-chat-sse-cutover`。
|
||||||
|
- Parent Issue `ISS-014` was closed and moved to `mvp/issues/archived/` after the stage 7 E2E evidence was accepted.
|
||||||
|
- Post-E2E runtime-quality findings are tracked by active `ISS-015`; they do not reopen this completed cleanup change.
|
||||||
@@ -0,0 +1,30 @@
|
|||||||
|
# Brief: single-react-cleanup-e2e
|
||||||
|
|
||||||
|
## Background
|
||||||
|
|
||||||
|
ISS-014 阶段 0-6B 已建立单一 Diagnosis ReAct Agent、Harness、ACI Tool、Guard 和 named SSE,但仓库仍存在 legacy AiOps/Sequential/Redis Session 链、旧 Tool contract、副作用式审计和过时文档。阶段 7 负责物理清理、durable metadata audit 和最终 live E2E。
|
||||||
|
|
||||||
|
## Goals
|
||||||
|
|
||||||
|
- 公开诊断只保留 `POST /api/chat`,业务代码只保留一个拥有 Tool loop 的 `diagnosis_agent`。
|
||||||
|
- RAG/log 仅作为 Harness backend;Agent-facing Tool 只来自 `HarnessEvidenceTools`。
|
||||||
|
- AgentStep 与 ToolInvocation 使用 exact sessionId/runId,长期持久化仅包含有界 metadata。
|
||||||
|
- 通过 Maven 启动、named SSE、日志与 MySQL exact-run 查询完成最终验收。
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
- 删除 legacy Controller、Service、Hook、ThreadLocal、prompt、前端入口、测试和死文档。
|
||||||
|
- 新增 Harness-native Agent/Tool durable audit,修正文档、issue 与 OpenSpec strict 缺陷。
|
||||||
|
- 为旧 V012 数据库增加幂等 V013 兼容迁移,并校准可重复 payment-timeout demo。
|
||||||
|
|
||||||
|
## Non-goals
|
||||||
|
|
||||||
|
- 不实现真实 CLS 或生产业务 MySQL Tool datasource。
|
||||||
|
- 不修改三类 Intent、DiagnosisDraft、EvidenceGuard、SemanticGuard 或 Release Policy 语义。
|
||||||
|
- 不保留 legacy endpoint、兼容分支、Graph 或第二套 Tool Call ID。
|
||||||
|
|
||||||
|
## Classification
|
||||||
|
|
||||||
|
- Scale: `complex`
|
||||||
|
- Interface impact: L4 breaking HTTP/frontend cleanup
|
||||||
|
- OpenSpec: `single-react-cleanup-e2e`
|
||||||
@@ -0,0 +1,111 @@
|
|||||||
|
# Decisions: single-react-cleanup-e2e
|
||||||
|
|
||||||
|
## Discover Status
|
||||||
|
|
||||||
|
- Checkpoint: Discover
|
||||||
|
- Capability source: `sm-flow` + `grill-with-docs` + `gitnexus-refactoring`。
|
||||||
|
- GitNexus index 停在 `2362665`,当前为 `bc36248`;刷新会改写用户已修改的 `AGENTS.md`,因此不安全。使用 OpenSpec/devflow、`rg` 全引用扫描、源码阅读、编译和测试作为 fallback。
|
||||||
|
- AGENTS.md 指定的 `codebase-retrieval` 与 LSP 工具在当前环境不可用,已显式记录限制。
|
||||||
|
- Scale: complex。涉及 L4 endpoint 删除、跨模块物理清理、Tool/Trace ownership、安全持久化和 live E2E。
|
||||||
|
|
||||||
|
## Question Pool
|
||||||
|
|
||||||
|
| # | 维度 | 问题 | 模式 | 状态 |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| Q1 | 范围 | 哪些旧 Agent/Service/Hook/Session 只有自引用测试,哪些仍在生产可达? | evidence-driven | 已解决 |
|
||||||
|
| Q2 | 协议 | 阶段 7 是否必须删除 `/api/ai_ops` 与旧 Session endpoints? | evidence-driven | 已解决 |
|
||||||
|
| Q3 | Tool | RAG/log 旧实现哪些可复用,哪些 Agent-facing contract/副作用必须删除? | evidence-driven | 已解决 |
|
||||||
|
| Q4 | Trace | 新 Harness 如何在不泄漏 raw/prompt/thought 的前提下满足 AgentStep/ToolInvocation E2E? | evidence-driven | 已解决 |
|
||||||
|
| Q5 | 文档 | ISS-012/ISS-013 和现有架构/Demo 文档如何收口? | evidence-driven | 已解决 |
|
||||||
|
| Q6 | 验收 | 最终 live E2E 必须证明哪些 exact-run 事实,哪些外部系统明确不声称 live? | evidence-driven | 已解决 |
|
||||||
|
| Q7 | 回滚 | L4 endpoint 删除如何迁移与回滚? | evidence-driven | 已解决 |
|
||||||
|
|
||||||
|
## Evidence-driven
|
||||||
|
|
||||||
|
| 结论 | 证据来源 | 是否已汇报用户 |
|
||||||
|
|---|---|---|
|
||||||
|
| `ChatService` 没有生产调用方,只剩自身单元/Smoke tests。 | `rg ChatService` | 已汇报 |
|
||||||
|
| `/api/ai_ops` 仍由 bundled frontend 按钮调用,并运行 Supervisor + Planner + Executor 多 Agent。 | `AiOpsController`、`AiOpsService`、`app.js`、`index.html` | 已汇报 |
|
||||||
|
| ISS-014 总体验收要求旧多 Agent/Graph 不存在,阶段 6B 只暂时保持 AiOps,阶段 7 负责清理/处置。 | ISS-014、阶段 6B design/acceptance | 已汇报 |
|
||||||
|
| `/api/chat/clear` 与 session info/runs 没有 frontend caller;Redis SessionManager 只由该 Controller 和测试使用。 | controller/frontend/session 引用扫描 | 已汇报 |
|
||||||
|
| `LookupKnowledgeTool` 和 `QueryLogsTools` 被新 Harness adapter 复用,但仍携带旧 `@Tool`、ThreadLocal/recorder 副作用。 | `HarnessChatConfiguration`、Tool source | 已汇报 |
|
||||||
|
| 新 ToolBoundary 写 Redis canonical invocation,但没有 `tool_invocation` durable audit;最终 DB E2E 会缺 Tool rows。 | Harness boundary/config 引用扫描 | 已汇报 |
|
||||||
|
| 旧 `AgentLoggingHook` 仍回退 ThreadLocal,并持久化 thought/model正文;不满足新安全边界。 | `AgentLoggingHook.java` | 已汇报 |
|
||||||
|
| 当前架构、Agent、Harness 和 lifecycle 文档仍描述 Planner/Executor/Verifier/Composer 与 AIOps 双入口。 | `mvp/architecture/*.md` | 已汇报 |
|
||||||
|
|
||||||
|
## User-interview
|
||||||
|
|
||||||
|
- 无新增 user-interview。唯一公开 Chat、旧多 Agent/Graph 物理删除、无兼容分支、最终 E2E 和外部 Mock 边界均已由 ISS-014 与用户的逐阶段自动执行授权冻结。
|
||||||
|
|
||||||
|
## Key Decisions
|
||||||
|
|
||||||
|
- 删除 legacy AiOps endpoint 而不是迁移到第二个 use case;所有诊断统一进入 `/api/chat` 的 Intent Router。
|
||||||
|
- 删除旧 Session endpoints/Redis conversation context;安全 PreviousTurn 只来自 `diagnosis_run.published_result`。
|
||||||
|
- RAG/log 查询实现保留为 Harness backend,移除 `@Tool` 和旧 recorder/session dedup;Agent 只看 ACI callbacks。
|
||||||
|
- 新 Agent audit hook 只写角色/数量/Tool 名称/耗时等 metadata,不写模型输入正文、输出正文、Thought 或 Tool arguments。
|
||||||
|
- Tool durable audit 通过 Harness port + JPA adapter fail-open 写入;Redis canonical store failure 仍 fail-closed,DB audit failure 只记录日志,不改变 Tool observation。
|
||||||
|
- 删除 endpoint 的迁移无兼容层;bundled frontend 同 commit 删除按钮/consumer,回滚整体回滚 commit。
|
||||||
|
- 不创建 ADR:方向已由 ISS-014 冻结,本 change 只完成最终落地与验收。
|
||||||
|
|
||||||
|
## OpenSpec Backfill
|
||||||
|
|
||||||
|
- 需进入 design/spec/tasks:删除清单、唯一 endpoint、Harness-native trace、Tool durable audit schema/safety、Tool backend 解耦、文档/issue 收口、strict validation、live E2E/log/DB acceptance 与回滚。
|
||||||
|
|
||||||
|
## Cross-artifact Alignment
|
||||||
|
|
||||||
|
| 上游 -> 下游 | 检查内容 | 状态 |
|
||||||
|
|---|---|---|
|
||||||
|
| ISS-014/brief -> proposal | 阶段 7 物理清理、文档、最终 Maven/log/DB E2E、Mock 外部 Tool 边界 | 已对齐 |
|
||||||
|
| proposal -> design | 删除闭包、Tool backend 复用、安全 Agent/Tool audit、L4 migration/rollback | 已对齐 |
|
||||||
|
| design -> specs/tasks | ownership、禁止泄漏、exact identity、strict validation、live E2E 均有 requirement 与切片 | 已对齐 |
|
||||||
|
| specs -> tasks | 8 组可观察 requirements 覆盖删除、audit、docs、verification 和最终 E2E | 已对齐 |
|
||||||
|
|
||||||
|
## Architecture Audit
|
||||||
|
|
||||||
|
- Capability source: `zoom-out`,使用 Diagnosis Harness、Diagnosis Agent、Canonical Invocation、Durable Audit、Diagnosis Trace 和 Chat SSE Contract 术语。
|
||||||
|
- 最终链路为 `browser -> ChatController -> ChatApplicationUseCase -> Harness -> Diagnosis Agent -> ACI Tools -> Guards -> SSE`,没有第二业务入口或业务 Graph。
|
||||||
|
- Redis canonical invocation 是短期完整 Tool 真理源;MySQL ToolInvocation 是长期有界 metadata audit,两者禁止双写 raw payload。
|
||||||
|
- `chat_session` JPA entity 属于当前 Run/PreviousTurn 目录,Redis SessionContext 属于旧 conversation memory;删除时必须区分。
|
||||||
|
- 最大风险是 backend 的旧 Tool annotation/recorder 隐式暴露和 audit 内容泄漏;design/tasks 已加入独占 discovery、negative serialization、context startup 与 live DB inspection。
|
||||||
|
|
||||||
|
## Interface Impact
|
||||||
|
|
||||||
|
- Level: L4 breaking HTTP/frontend contract。
|
||||||
|
- 删除 `/api/ai_ops`、`/api/chat/clear`、`/api/chat/session/{sessionId}` 和 `/runs`;保留唯一 `/api/chat` SSE 与 Trace API。
|
||||||
|
- bundled frontend 同 commit 删除 AiOps 按钮/consumer;外部 caller 迁移到 `/api/chat`,无兼容 branch。
|
||||||
|
- 回滚必须整体回滚阶段 7 commit,不能单独恢复旧 endpoint/Tool annotations。
|
||||||
|
|
||||||
|
## Apply Verification Status
|
||||||
|
|
||||||
|
- Follow-up current-document audit found and corrected stale runtime semantics in the tracked payment-timeout PowerShell demo and `mvp/tables`: old JSON Chat parsing, Redis SessionContext history, Planner/Verifier identities, full Tool payload persistence, and Chat/AiOps Run descriptions are no longer presented as current behavior.
|
||||||
|
- `mvp/demo/scripts/run-payment-timeout-demo.ps1` now strictly validates `metadata -> status* -> content|failure -> done`, captures exact session/run identity, and fetches exact Trace without writing feedback. PowerShell parser validation and two in-memory SSE contract samples passed.
|
||||||
|
- Current architecture/demo/table scan excluding explicit `archive/` and ignored local `output/` artifacts: 0 legacy runtime matches.
|
||||||
|
- Follow-up `openspec validate --all --strict`: 22 passed, 0 failed; `git diff --check`: passed.
|
||||||
|
- Initial environment checks reported missing injected variables. Per user direction, credentials were restored to ignored `application-local.yml`; no credential was added to Git or emitted in the archive.
|
||||||
|
- `mvn -q -DskipTests compile`: passed.
|
||||||
|
- `node --check src/main/resources/static/app.js`: passed.
|
||||||
|
- `openspec validate --all --strict`: 22 passed, 0 failed.
|
||||||
|
- Deterministic regression suite excluding credential-dependent MySQL/Redis/Milvus tests: 60 suites, 221 tests, 0 failures, 0 errors, 3 skipped.
|
||||||
|
- `SemanticGuardTest` interruption assertion failed once under the credential-dependent full-suite run, then passed three isolated repetitions and the deterministic regression suite; classified as load-sensitive test timing, not a reproduced product regression.
|
||||||
|
- `mvn -q -DskipTests package`: passed.
|
||||||
|
- Production legacy path scan, audit sensitive-payload scan, and legacy backend Tool contract scan: 0 matches.
|
||||||
|
- The approved external-network run connected to MySQL/Redis/Milvus/model services. Flyway repair aligned the old V012 checksum and V013 reconciled the missing release-contract columns/index.
|
||||||
|
- Maven startup validated all 13 migrations, passed JPA schema validation and started Tomcat 9900 with Harness/Audit beans.
|
||||||
|
- Final exact live run: session `mvp-demo-payment-timeout-stage7-20260722-1741`, run `363f481c-33b8-42e7-8699-428a6ec61806`, SSE `metadata -> status -> status -> status -> content -> done`, outcome `FALLBACK`.
|
||||||
|
- Exact MySQL evidence: one `DIAGNOSIS/SUCCESS/FALLBACK` run, two metadata-only `diagnosis_agent` steps with null Thought, and two READY/EVIDENCE_FOUND Tool audits for `lookup_knowledge` and Mock `query_logs`.
|
||||||
|
- Tasks 4.2-4.5 are complete. Real CLS and production business MySQL Tool datasource remain explicit non-goals.
|
||||||
|
|
||||||
|
## Apply Conflict Classification
|
||||||
|
|
||||||
|
- **Code deviation**: `DiagnosisRun` enum mapping expected native ENUM while V012/V013 define VARCHAR. Fixed ORM column definitions; OpenSpec unchanged.
|
||||||
|
- **Code deviation**: production `ObjectMapper` lacked Java Time modules, causing canonical Redis `STORE_ERROR`. Fixed mapper registration and bound the store test to the production mapper.
|
||||||
|
- **Code deviation**: Mock empty log results used `success=false`, conflicting with the specified `NO_EVIDENCE` projection. Fixed backend success semantics and the stale test assertion.
|
||||||
|
- **Code deviation**: RAG backend logs exposed raw query/rewritten query content. Replaced with bounded counts/category metadata and verified the final Tool execution window contains no prohibited payload.
|
||||||
|
- **Acceptance fixture drift**: the payment-timeout demo requested the removed metrics Tool and encouraged an unbounded investigation. Updated the fixture to the current two-Tool Mock acceptance scope and explicit stopping boundary.
|
||||||
|
|
||||||
|
## Commit Gate Preflight
|
||||||
|
|
||||||
|
- proposal、design、两份 specs 和 tasks 完整;新 capability 与 modified capability 的范围无 gap。
|
||||||
|
- Question pool 全部 evidence-driven 并已汇报,无 user-interview、未判级接口或未接受架构风险。
|
||||||
|
- 删除清单区分 current JPA metadata 与 legacy Redis Session,Tool backend 与 Agent-facing contract,canonical truth 与 durable audit。
|
||||||
|
- live E2E 明确要求真实应用/模型链和 exact ID;真实 CLS/生产业务 MySQL 明确不在验收声称范围。
|
||||||
@@ -0,0 +1,26 @@
|
|||||||
|
# Evidence: single-react-cleanup-e2e
|
||||||
|
|
||||||
|
## Repository Evidence
|
||||||
|
|
||||||
|
- 生产引用扫描确认旧 `ChatService` 仅剩自身测试,`/api/ai_ops` 与旧 Session endpoints 属于第二条公开/状态链。
|
||||||
|
- Harness adapters 复用 `LookupKnowledgeTool`/`QueryLogsTools` backend;旧 `@Tool`、ThreadLocal、recorder 和 topic discovery contract 已移除。
|
||||||
|
- `HarnessAgentAuditHook` 只写 message count/roles、Tool names、text presence 和 duration;`thought` 保持空。
|
||||||
|
- ToolBoundary durable audit 只写 exact identity、状态、稳定错误码、耗时和字节数;Redis canonical invocation 仍是短期完整 Tool 真理源。
|
||||||
|
|
||||||
|
## Live Findings
|
||||||
|
|
||||||
|
- 远端库已登记旧 V012 checksum,但缺少 `diagnosis_run.intent/release_outcome/published_result`。Flyway repair 后,幂等 V013 补齐三列与索引,fresh database 上为 no-op。
|
||||||
|
- Hibernate 6 将无 `columnDefinition` 的字符串枚举校验为原生 ENUM;`DiagnosisRun` 已显式映射到 V012/V013 的 `VARCHAR(32/16)`。
|
||||||
|
- 生产 `WebConfig` 的裸 `ObjectMapper` 无法序列化 canonical record 的 `Instant`,导致所有 Tool 在 begin 阶段返回 `STORE_ERROR`;改为自动注册模块,并让 store 测试使用生产 mapper。
|
||||||
|
- Mock `query_logs` 把 0 命中错误表达为 `success=false`,与 `NO_EVIDENCE` contract 冲突;现以成功查询 + 空数组表达无证据。
|
||||||
|
- live 日志发现 RAG backend 打印原始 query/rewrittenQuery/keywords;已改为字符数、命中数、类别数和耗时 metadata。
|
||||||
|
|
||||||
|
## Final Exact-run Evidence
|
||||||
|
|
||||||
|
- sessionId: `mvp-demo-payment-timeout-stage7-20260722-1741`
|
||||||
|
- runId: `363f481c-33b8-42e7-8699-428a6ec61806`
|
||||||
|
- SSE: `metadata -> status -> status -> status -> content -> done`
|
||||||
|
- outcome: `FALLBACK`; diagnosis_run: `DIAGNOSIS / SUCCESS / FALLBACK`
|
||||||
|
- AgentStep: 2 rows,均为 `diagnosis_agent`,`thought IS NULL`,model input/output 仅 metadata。
|
||||||
|
- ToolInvocation: 2 rows,`lookup_knowledge` 与 `query_logs` 均为 `READY/EVIDENCE_FOUND`,同一 exact identity,无错误。
|
||||||
|
- `query_logs` 明确为 Mock;Agent-facing `query_mysql` 未配置生产业务 datasource,也未声称 live。
|
||||||
@@ -0,0 +1,68 @@
|
|||||||
|
# Diagnosis 信息增益停止契约 验收
|
||||||
|
|
||||||
|
## 结果
|
||||||
|
|
||||||
|
已接受。OpenSpec tasks 37/37 完成;用户确认归档 OpenSpec、提交并推送。
|
||||||
|
|
||||||
|
## 验证
|
||||||
|
|
||||||
|
### 静态验证
|
||||||
|
|
||||||
|
- 命令/检查:`openspec validate diagnosis-information-gain-stop-contract --strict`
|
||||||
|
- 结果:passed
|
||||||
|
- 备注:Committed OpenSpec 与最终 tasks 一致
|
||||||
|
|
||||||
|
- 命令/检查:OpenSpec delta → main specs 同步(6 个 capability)
|
||||||
|
- 结果:passed
|
||||||
|
- 备注:新建 `openspec/specs/diagnosis-information-gain-stop-contract/`,并更新 5 个既有 main specs
|
||||||
|
|
||||||
|
- 命令/检查:对照 OpenSpec 与代码路径(progress tracker、interceptor、release、trace)
|
||||||
|
- 结果:passed
|
||||||
|
- 备注:Task 8 行为与 design/spec 对齐
|
||||||
|
|
||||||
|
### 脚本验证
|
||||||
|
|
||||||
|
- 命令:`mvn -q "-Dtest=DiagnosisProgressTrackerTest,HarnessToolInterceptorTest,DiagnosisReleaseUseCaseTest,DiagnosisAgentUseCaseTest,HarnessChatConfigurationTest" test`
|
||||||
|
- 结果:passed(exit 0)
|
||||||
|
- 备注:覆盖协议 violation_type、可修正 observation、连续协议错误 STOP、Release fail-closed、Agent 受控停止
|
||||||
|
|
||||||
|
- 命令:历史全量回归 `mvn -q -Dtest='!MilvusConnectionTest' test`(tasks 6.2/7.5 阶段)
|
||||||
|
- 结果:passed(`Tests=292, Failures=0, Errors=0, Skipped=3`)
|
||||||
|
- 备注:未设置 `MILVUS_TOKEN` 时 `MilvusConnectionTest` 失败属外部凭据边界,非本变更回归
|
||||||
|
|
||||||
|
### 浏览器/人工验证
|
||||||
|
|
||||||
|
- 步骤:Maven 启动真实应用;Query `诊断切换企业失败的问题`;named SSE + logs + `scripts/query_mysql.py` 按 exact sessionId/runId 核对
|
||||||
|
- 结果:passed
|
||||||
|
- 备注:
|
||||||
|
- `sessionId=iss016-final-20260726-a`
|
||||||
|
- `runId=3ab22ed7-d0ed-45d8-b928-dce5790c0542`
|
||||||
|
- SSE:`SAFE_FALLBACK` / `MISSING_REQUIRED_CONTEXT` / `done.outcome=FALLBACK`
|
||||||
|
- DB:`status=SUCCESS`、`intent=DIAGNOSIS`、`release_outcome=FALLBACK`、`tool_call_count=0`、`total_token_count=2890`
|
||||||
|
- Trace:`RUN_STARTED -> ROUTING_* -> AGENT_MODEL_STEP -> EVIDENCE_GUARD_INITIAL -> RELEASE_DECISION/FALLBACK -> RUN_FINISHED/FALLBACK`
|
||||||
|
|
||||||
|
### 未验证
|
||||||
|
|
||||||
|
- 本轮 Archive 未再重跑全量 `mvn test` 与完整 live SSE E2E;依赖 Apply 阶段记录与本轮 focused tests。
|
||||||
|
- 真实 Provider 下连续协议错误的 live E2E 未单独复跑;协议停止由 focused/scripted loop 覆盖。
|
||||||
|
- ISS-015 阶段 2(Evidence Repair Schema)与阶段 3(Reasoning 治理)不在本 change 范围。
|
||||||
|
|
||||||
|
## 已完成范围
|
||||||
|
|
||||||
|
- 信息增益停止、scope 去重、STOP_REQUIRED、ProgressSnapshot、统一 Release
|
||||||
|
- Token 审计与 Tool 拒绝 Trace
|
||||||
|
- 协议修复反馈 + `PROGRESS_PROTOCOL_VIOLATED` 兜底停止
|
||||||
|
- 文档:ISS-015 阶段 1、ISS-016、架构文档、glossary 术语、devflow 档案
|
||||||
|
|
||||||
|
## 已知限制
|
||||||
|
|
||||||
|
- 重复检测只比较确定性 `tool_name + normalized_scope`,不做自然语言语义去重。
|
||||||
|
- 协议错误阈值默认 2,与无增益阈值独立配置。
|
||||||
|
- 无安全 ProgressSnapshot 的受控停止继续 fail closed,不伪造用户可见事实。
|
||||||
|
- 公开 SSE/前端协议无新增字段;模型侧 Tool Envelope 是已确认 L3 变更。
|
||||||
|
|
||||||
|
## 交接
|
||||||
|
|
||||||
|
- 下一步:OpenSpec 已用户确认归档;代码提交并推送到当前分支。
|
||||||
|
- OpenSpec 归档确认:用户确认归档(“执行,完后提交推送”)
|
||||||
|
- 归档位置:`openspec/changes/archive/2026-07-27-diagnosis-information-gain-stop-contract/`
|
||||||
@@ -0,0 +1,45 @@
|
|||||||
|
# Diagnosis 信息增益停止契约 Brief
|
||||||
|
|
||||||
|
## 背景
|
||||||
|
|
||||||
|
- 用户目标:未知问题、空证据或缺少查询条件时,Diagnosis 能以正常业务 Fallback 结束,而不是空转到预算耗尽或 `INTERNAL_FAILURE`。
|
||||||
|
- 当前问题:停止主要依赖模型自觉结束或硬预算;缺少信息增益回传、确定性饱和停止、协议修复反馈和统一 Release。
|
||||||
|
- 关联 OpenSpec:`openspec/changes/diagnosis-information-gain-stop-contract/`(归档后见 `openspec/changes/archive/2026-07-27-diagnosis-information-gain-stop-contract/`)
|
||||||
|
- 关联 Issue:ISS-016(承接 ISS-015 阶段 1 硬停止)
|
||||||
|
- devflow 分档:complex
|
||||||
|
- 接口影响:L3(模型可见 Tool Envelope 有意变更;公开 HTTP/SSE 与业务 Tool backend 不变)
|
||||||
|
|
||||||
|
## 范围
|
||||||
|
|
||||||
|
### 本次要做
|
||||||
|
|
||||||
|
- Run 内二值信息增益 `GAINED / NO_GAIN`、连续无增益阈值与 `COLLECTING / SATURATED`。
|
||||||
|
- 服务端注册的 Tool Call Envelope:`previous_observation + input`。
|
||||||
|
- 确定性 `NO_GAIN`(`NO_EVIDENCE`、重复 `tool_name + normalized_scope`)。
|
||||||
|
- 一次 `STOP_REQUIRED` 收尾机会与受控停止异常。
|
||||||
|
- Canonical 控制视图 / 模型白名单观察双视图;RAG 保留 `relevance_level`。
|
||||||
|
- Tool loop 结束时一次性投影 `ProgressSnapshot`。
|
||||||
|
- `DiagnosisReleaseUseCase` 统一有结论、无结论、信息饱和、预算终止、协议违规停止的发布。
|
||||||
|
- 可修正 `INVALID_PROGRESS_PROTOCOL` observation,以及独立 `PROGRESS_PROTOCOL_VIOLATED` 兜底停止。
|
||||||
|
- 模型 Token 组件/轮次审计与 `TOOL_REQUEST_REJECTED` 安全 Trace。
|
||||||
|
- Prompt、配置、focused/回归测试与真实 named SSE E2E。
|
||||||
|
|
||||||
|
### 本次不做
|
||||||
|
|
||||||
|
- `new_count`、`next_action`、多级质量分数、独立 Judge。
|
||||||
|
- 自然语言语义去重。
|
||||||
|
- 第二套诊断生命周期状态。
|
||||||
|
- 公开 HTTP/SSE 字段或前端进度协议新增。
|
||||||
|
- ISS-015 Reasoning 原文审计治理与 Evidence Repair Schema 注入。
|
||||||
|
|
||||||
|
### 影响区域
|
||||||
|
|
||||||
|
- `harness.progress`、`HarnessToolInterceptor`、`HarnessEvidenceTools`
|
||||||
|
- `DiagnosisAgentUseCase` / Prompt / Release / Application 预算兜底迁移
|
||||||
|
- Trace 审计、配置绑定、ISS-015/016 与架构文档
|
||||||
|
|
||||||
|
## OpenSpec 对齐
|
||||||
|
|
||||||
|
- proposal 覆盖状态:已覆盖
|
||||||
|
- specs 覆盖状态:已覆盖(6 个 capability delta,已同步 main specs)
|
||||||
|
- tasks 覆盖状态:已覆盖(37/37 完成)
|
||||||
@@ -0,0 +1,188 @@
|
|||||||
|
# Diagnosis 信息增益停止契约 Decisions
|
||||||
|
|
||||||
|
## Discover Status
|
||||||
|
|
||||||
|
- Checkpoint:Discover。
|
||||||
|
- Capability source:`sm-flow` 内置 Discover 协议;grill 使用 `grill-with-docs`,代码可证问题通过源码、测试和引用搜索处理。
|
||||||
|
- Scale:`complex`。变更跨越 Agent、Tool 协议、Run 生命周期、Release、Guard、配置、Trace 和 E2E。
|
||||||
|
- 接口影响:L3。模型可见 Tool Schema 发生有意协议变更,公开 HTTP/SSE 和业务 Tool backend 协议不变。
|
||||||
|
- 工具降级:当前没有 `codebase-retrieval` 和 LSP 工具;以 `rg`、源码和测试引用核查替代。用户已明确“可以忽略gitnexus”。
|
||||||
|
|
||||||
|
## Question Pool
|
||||||
|
|
||||||
|
| # | 维度 | 问题 | 模式 | 状态 |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| Q1 | 术语 | Tool 客观状态、信息增益、收集状态、停止原因和最终发布状态是否应合并为一个枚举? | user-interview | 已解决 |
|
||||||
|
| Q2 | 语义 | Tool 返回价值由谁判断,是否需要多级质量分数? | user-interview | 已解决 |
|
||||||
|
| Q3 | 协议 | 模型继续调用 Tool 时如何回传上一轮信息增益,是否需要 `next_action`? | user-interview | 已解决 |
|
||||||
|
| Q4 | 边界 | Tool Schema 应来自 Prompt 还是服务端原生 Tool Calling 注册? | user-interview | 已解决 |
|
||||||
|
| Q5 | 边界 | Harness 能确定性判定哪些 `NO_GAIN`,RAG `REFERENCE` 由谁判定? | user-interview | 已解决 |
|
||||||
|
| Q6 | 范围 | 首版是否需要 `new_count` 或自然语言语义去重? | user-interview | 已解决 |
|
||||||
|
| Q7 | 配置 | 连续无增益阈值是否可配置,默认值与生效时机是什么? | user-interview | 已解决 |
|
||||||
|
| Q8 | 发布 | 信息饱和、预算终止和 `conclusion=null` 由谁转换为用户可见结果? | user-interview | 已解决 |
|
||||||
|
| Q9 | 验收 | 如何证明未知问题不再以通用内部错误结束,同时不放过无证据结论? | evidence-driven | 已解决 |
|
||||||
|
| Q10 | 技术 | 当前 Tool schema 是否能直接容纳 `previous_observation`? | evidence-driven | 已解决 |
|
||||||
|
| Q11 | 技术 | 进展控制状态应扩展 Redis Store 还是放入 RunContext handle? | evidence-driven | 已解决 |
|
||||||
|
| Q12 | 技术 | RAG `relevance_level` 在哪一层丢失,前端是否已有过程展示能力? | evidence-driven | 已解决 |
|
||||||
|
|
||||||
|
## Evidence-driven
|
||||||
|
|
||||||
|
| 结论 | 证据来源 | 是否已汇报用户 |
|
||||||
|
|---|---|---|
|
||||||
|
| 当前三个 Agent-facing Tool 直接使用 `RagToolRequest`、`QueryLogsRequest`、`MysqlToolRequest` 生成 Schema;要增加 `previous_observation + input` 必须显式演进 Tool Schema,不能只改 interceptor。 | `HarnessEvidenceTools`、三个 request records、`DiagnosisAgentFactory` | 已汇报 |
|
||||||
|
| `HarnessToolInterceptor` 当前把完整 `agentResult` 放入 `ToolCallResponse.content`,控制视图与模型观察尚未分离。 | `HarnessToolInterceptor` | 已汇报 |
|
||||||
|
| `RunContext` 已采用结构不可变、可变状态存在线程安全 handle 的模式;进展 tracker 放入 RunContext 比扩展 Redis 按 Run 枚举更符合现有所有权。 | `RunContext`、`DiagnosisHarnessCore.startRun` | 已汇报 |
|
||||||
|
| `CanonicalInvocationStore` 只有 begin/find/markReady/markError,扩展按 Run 枚举会影响 Redis 实现和多组 fake store;首版可由 tracker 保存完成调用 key,在结束时按 key 读取 canonical 记录。 | `CanonicalInvocationStore` 及其引用测试 | 已汇报 |
|
||||||
|
| `RagResultProjector` 只按 evidence 是否为空生成 `EVIDENCE_FOUND / NO_EVIDENCE`,没有读取上游 `relevanceLevel / relevance_level`。 | `RagResultProjector`、`LookupResult`、`KnowledgeEvidencePostProcessor` | 已汇报 |
|
||||||
|
| `DiagnosisReleaseUseCase.execute` 当前强制 draft 非空并对所有 Draft 运行 EvidenceGuard;`EvidenceGuard` 又把空 analysis 判为 `ANALYSIS_MISSING`,与合法无结论结果冲突。 | `DiagnosisReleaseUseCase`、`EvidenceGuard` | 已汇报 |
|
||||||
|
| `ChatApplicationUseCase.recoverBudgetExhaustion` 已有未提交预算 Fallback,但它绕过 Diagnosis Release,需迁移而不是丢弃用户价值。 | `ChatApplicationUseCase`、`SafeFallbackFactory`、现有测试 diff | 已汇报 |
|
||||||
|
| 前端已渲染 `observed_facts / verified_sources / limitations / next_steps`,不需要新增公开展示协议。 | `src/main/resources/static/app.js` | 已汇报 |
|
||||||
|
| 验收必须同时覆盖主动无结论、Harness 饱和、预算终止、无证据结论被 Guard 拦截,以及原始未知 Query 的 live SSE、日志和 exact run 数据。 | 当前事故现象、ISS-016 验收项、现有 E2E 工具 | 已汇报 |
|
||||||
|
|
||||||
|
## User-interview
|
||||||
|
|
||||||
|
| 问题原文 | 用户原话 | 确认状态 | OpenSpec 回写 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| 是否精简状态而不建立第二套生命周期? | “这里我觉得设计得太混乱了,怎么简化”以及对最终简化架构“我觉得可以” | 已确认 | 已回写 |
|
||||||
|
| 是否只保留 `GAINED / NO_GAIN`? | “质量状态只要 information_gain = GAINED \| NO_GAIN 就够了?”后确认“我觉得可以” | 已确认 | 已回写 |
|
||||||
|
| 是否需要 `new_count`? | “好那就去掉new_count” | 已确认 | 已回写 |
|
||||||
|
| 是否需要 `next_action`? | “那就去掉next_action,我觉得由llm自己去判断就好了,不用显示的指定” | 已确认 | 已回写 |
|
||||||
|
| Tool 如何注入? | “首先 tool注入,由服务端注入,而不是写死在提示词中” | 已确认 | 已回写 |
|
||||||
|
| Prompt 是否强调合法放弃和正确但无用的内容? | “需要说明 模型不必须要给出一个答案”以及“如果你发现工具返回的是正确但对推导无用的废话,请停止调用” | 已确认 | 已回写 |
|
||||||
|
| 阈值是否可配置? | “我觉得这个可以暴露出一个配置来控制” | 已确认 | 已回写 |
|
||||||
|
| 首版重复检测做到什么程度? | “也就是说这一版只是做参数的去重校验”后确认“可以” | 已确认 | 已回写 |
|
||||||
|
| 是否按当前简化方案进入实施? | “可以用更简单的方式”“可以。修正一下文档”“用sm-flow开始实施把” | 已确认 | 已回写,并授权完成 Commit 后进入 Apply |
|
||||||
|
|
||||||
|
## 关键取舍
|
||||||
|
|
||||||
|
- 决策:停止权归 Harness,语义价值判断由模型与确定性规则共同产生。
|
||||||
|
- 原因:Tool 只能知道客观返回,模型才能判断内容是否推进当前假设;但空结果和完全重复 scope 可由代码零 Token 判定。
|
||||||
|
- 影响:Harness 只消费二值信息增益,不引入独立 Judge 或质量分数。
|
||||||
|
- 决策:使用下一次 Tool Call Envelope 回传上一轮模型评价。
|
||||||
|
- 原因:模型只有看到 Tool Observation 后才能评价,下一次真实行为正好提供受 Schema 约束的回传边界。
|
||||||
|
- 影响:这是 L3 Agent-facing Tool Schema 变更,业务 request 在 interceptor 内解包后保持不变。
|
||||||
|
- 决策:进展 tracker 是 RunContext handle,Canonical Store 保持 Tool 真相源。
|
||||||
|
- 原因:停止决策需要低延迟 Run 内状态,完整证据仍应由 canonical 记录提供;两者职责不同。
|
||||||
|
- 影响:tracker 保存计数、scope、待评价调用和 canonical keys,不复制 raw payload。
|
||||||
|
- 决策:不创建 ADR。
|
||||||
|
- 原因:这些是 ISS-016 范围内可通过 OpenSpec 回滚的内部协议演进,已有架构文档详细记录取舍,尚不满足独立 ADR 的必要性。
|
||||||
|
|
||||||
|
## OpenSpec 回写
|
||||||
|
|
||||||
|
- 必须进入 proposal/design/spec/tasks:Tool Envelope、L3 影响、Run tracker、确定性 `NO_GAIN`、RAG `relevance_level`、双视图、STOP_REQUIRED、ProgressSnapshot、统一 Release、Prompt、配置和 E2E。
|
||||||
|
- 必须保留非目标:无 `new_count`、无 `next_action`、无 Judge、无语义去重、无第二套诊断生命周期、无公开 SSE 协议新增。
|
||||||
|
- 当前没有未确认的 user-interview 问题,也没有 devflow/OpenSpec 冲突。
|
||||||
|
|
||||||
|
## Cross-Artifact 对齐检查
|
||||||
|
|
||||||
|
| 上游 → 下游 | 检查内容 | 状态 |
|
||||||
|
|---|---|---|
|
||||||
|
| ISS-016/架构文档 → proposal | 未知问题、合法放弃、信息增益、饱和停止、双视图、统一 Release、非目标和 E2E | 已对齐 |
|
||||||
|
| proposal → design | L3 Envelope、Run tracker、scope、STOP_REQUIRED、ProgressSnapshot、预算终态、Prompt 和迁移方案 | 已对齐 |
|
||||||
|
| design → specs/tasks | 状态机、Tool 门禁、RAG relevance、无结论 Guard、Release 所有权、Trace 和兼容边界 | 已对齐 |
|
||||||
|
| specs → tasks | 每个可观察行为均有 contract/state/loop/release/config/E2E 可执行切片 | 已对齐 |
|
||||||
|
|
||||||
|
### Gap 详情
|
||||||
|
|
||||||
|
- 无。
|
||||||
|
|
||||||
|
## Architecture Audit
|
||||||
|
|
||||||
|
- Capability source:`zoom-out`。按 glossary 的 Diagnosis Agent、Diagnosis Harness、RunContext、Evidence Status、Invocation Status、Release Outcome 术语审计。
|
||||||
|
- 顶层链路:`ChatApplicationUseCase -> DiagnosisChatExecutor -> DiagnosisAgentUseCase -> ReactAgent/interceptors -> ToolBoundary/canonical store -> ProgressSnapshot -> DiagnosisReleaseUseCase -> SSE/persistence`。
|
||||||
|
- 所有权:Agent 负责诊断语义;Tool/Projector 负责客观结果;Run tracker 负责停止控制;Canonical Store 负责 Tool 真相;Guard 负责引用与结论安全;Release 负责用户可见 SUCCESS/FALLBACK;Application 只负责编排和持久化。
|
||||||
|
- `RunContext` 的生产代码构造点只有 `DiagnosisHarnessCore.startRun`,大量测试通过该工厂获取;新增 tracker 不需要扩散手工构造。
|
||||||
|
- `DiagnosisAgentUseCase`、`HarnessToolInterceptor`、`HarnessEvidenceTools` 和 `DiagnosisReleaseUseCase` 的直接消费者均已由配置类和 focused tests 覆盖,任务清单包含所有构造调用更新。
|
||||||
|
- `FallbackType` 新语义只通过通用 SafeFallback JSON/前端渲染消费,没有前端枚举 switch;公开协议不新增字段。
|
||||||
|
- 最大框架风险是 Tool Envelope Schema 和 STOP_REQUIRED 后异常传播;design 要求三个具体 record、真实 callback schema 测试和 scripted framework-loop 测试在 Release 迁移前锁定行为。
|
||||||
|
- 最大生命周期风险是预算已把 RunLifecycle 置为 `BUDGET_EXHAUSTED` 后 Application 再次 `checkActive`;design 将其限制为“Diagnosis Release 已处理的预算 Fallback”窄分支,并禁止 Application 重建业务内容。
|
||||||
|
- 审计结论:模块职责没有形成新的循环依赖或第二真相源;L3 风险已进入 specs 和 tasks,可进入 commit gate。
|
||||||
|
|
||||||
|
## Commit Gate Preflight
|
||||||
|
|
||||||
|
- `proposal.md`、`design.md`、六份 capability delta specs 和 `tasks.md` 均存在。
|
||||||
|
- `openspec status --change diagnosis-information-gain-stop-contract --json` 返回 `isComplete=true`。
|
||||||
|
- `openspec validate diagnosis-information-gain-stop-contract --strict` 通过。
|
||||||
|
- question pool 全部已解决;evidence-driven 结论已汇报;user-interview 决策均有用户原话和确认状态。
|
||||||
|
- 接口影响已判为 L3,并有独立 Interface Impact、兼容、迁移、回滚和验收说明。
|
||||||
|
- Cross-artifact 检查无 gap;架构风险均已进入 design/tasks。
|
||||||
|
- 用户已通过“用sm-flow开始实施把”明确授权 Commit 后进入 Apply。
|
||||||
|
|
||||||
|
## Pre-apply Research
|
||||||
|
|
||||||
|
### 参考实现
|
||||||
|
|
||||||
|
- `HarnessEvidenceTools`:现有三类 `FunctionToolCallback` 注册点和 adapter bridge,继续作为 Agent-facing Schema 唯一入口。
|
||||||
|
- `HarnessToolInterceptor`:可获得 exact framework Tool Call ID,适合消费 Envelope 和执行 progress gate。
|
||||||
|
- `ToolBoundary`:Tool 预算、Run 校验、canonical 写入和安全错误的单点,不在 interceptor 重复 reserve。
|
||||||
|
- `RunContext` / `DiagnosisHarnessCore.startRun`:结构不可变 + 可变 handle 模式和唯一生产构造点。
|
||||||
|
- `RagResultProjector` / `QueryLogsResultProjector` / `MysqlResultProjector`:bounded canonical agent result 的现有标准化模式。
|
||||||
|
- `EvidenceGuard` / `DiagnosisReleaseUseCase`:当前结论验证链和无结论冲突位置。
|
||||||
|
- `ChatApplicationUseCase.recoverBudgetExhaustion`:保留用户价值、需要迁移所有权的临时预算 Fallback。
|
||||||
|
- `DiagnosisAgentUseCaseTest.ScriptedChatModel`:真实框架 model -> Tool -> model loop 回归模式。
|
||||||
|
|
||||||
|
### 技术栈清单
|
||||||
|
|
||||||
|
- Tool Schema:三个具体 record 交给 Spring AI `FunctionToolCallback.inputType`,共享 `PreviousObservation`,不使用泛型擦除或 JsonNode Schema。
|
||||||
|
- JSON:继续使用项目 `ObjectMapper` 严格解析/序列化;控制字段在 interceptor 消费后只传业务 input。
|
||||||
|
- Run 状态:新增线程安全 tracker handle,由 `DiagnosisHarnessCore.startRun` 创建,不使用 ThreadLocal。
|
||||||
|
- Canonical 真相:继续使用 `ToolCallKeyFactory + CanonicalInvocationStore.find`;tracker 只记录 identity。
|
||||||
|
- 视图:从 bounded canonical `agent_result` 白名单投影 Model Observation,不读取 raw response。
|
||||||
|
- 异常:受控停止使用专用异常和 cause-chain 分类;未知异常保持 fail closed。
|
||||||
|
- 测试:JUnit 5、scripted ChatModel、现有 fake store/adapter fixture;不增加 Maven 依赖。
|
||||||
|
|
||||||
|
### 新建基础设施
|
||||||
|
|
||||||
|
- `harness.progress`:信息增益、收集状态、停止原因、tracker、scope、snapshot/projector。
|
||||||
|
- `harness.agent`:三个 Agent-facing Envelope、白名单 observation projector、受控停止异常和执行结果。
|
||||||
|
- 不新增数据库表、Redis 数据结构、HTTP DTO、SSE event 或外部依赖。
|
||||||
|
|
||||||
|
## Apply 期间设计补充:Draft 合同失败
|
||||||
|
|
||||||
|
- 真实 E2E `runId=4e667111-524e-4407-87ab-b4b262952017` 已完成一次 READY RAG 调用,第二轮模型返回文本后在 Draft/Release 边界失败;后续三次同 Query 均走零 Tool 的 `MISSING_REQUIRED_CONTEXT`,证明模型输出存在随机分支。
|
||||||
|
- 用户确认采用窄化降级:非法 Draft 自身不被接受;已有当前 Run 的安全 ProgressSnapshot 时发布 `INSUFFICIENT_EVIDENCE`,没有安全过程时继续 `FAILED`。
|
||||||
|
- 这是有意行为变更:从“所有非法 Draft 都发布技术失败”调整为“非法 Draft + 已验真过程可发布过程型 Fallback”;公开 SSE 字段、Tool 协议和最终生命周期枚举不变。
|
||||||
|
- 不新增 stop reason,不把 Draft 解析失败伪装成 `INFORMATION_SATURATED` 或 `BUDGET_LIMIT_REACHED`;使用 Agent 输出异常携带有界 snapshot,并以脱敏 Trace 区分输出合同失败。
|
||||||
|
|
||||||
|
## Apply Verification
|
||||||
|
|
||||||
|
- Focused tests:`DiagnosisAgentUseCaseTest`、`DiagnosisReleaseUseCaseTest`、`DiagnosisChatExecutorTest`、`HarnessChatConfigurationTest` 通过。
|
||||||
|
- 完整回归:`mvn -q -Dtest='!MilvusConnectionTest' test` 退出码为 `0`;本轮 Surefire 报告汇总 `Tests=292, Failures=0, Errors=0, Skipped=3`。
|
||||||
|
- 外部凭据边界:未排除时唯一失败为 `MilvusConnectionTest.connect`,原因是当前测试进程未设置 `MILVUS_TOKEN`;这不是本变更回归。
|
||||||
|
- OpenSpec:`openspec.cmd validate diagnosis-information-gain-stop-contract --strict` 通过。
|
||||||
|
- 格式与清理:`git diff --check` 通过;未发现临时 E2E JSON、DEBUG 或 tmp 文件。
|
||||||
|
- named SSE E2E:Query `诊断切换企业失败的问题`,`sessionId=iss016-final-20260726-a`,`runId=3ab22ed7-d0ed-45d8-b928-dce5790c0542`;SSE 返回 `SAFE_FALLBACK`、`type=MISSING_REQUIRED_CONTEXT`、`done.outcome=FALLBACK`。
|
||||||
|
- 数据库核对:`status=SUCCESS`、`intent=DIAGNOSIS`、`release_outcome=FALLBACK`、`tool_call_count=0`、`total_token_count=2890`、answer 非空。
|
||||||
|
- Trace 核对:`RUN_STARTED -> ROUTING_ATTEMPT -> ROUTING_DECISION -> AGENT_MODEL_STEP -> EVIDENCE_GUARD_INITIAL -> RELEASE_DECISION/FALLBACK -> RUN_FINISHED/FALLBACK`。
|
||||||
|
- 兼容性:公开 HTTP/SSE 字段、前端 SafeFallback 消费结构、数据库表和业务 Tool request 均未新增字段;模型侧 Tool Envelope 是本变更已确认的 L3 协议变更。
|
||||||
|
|
||||||
|
## Apply Continuation: Task 8 Protocol Repair + Bounded Stop
|
||||||
|
|
||||||
|
- Checkpoint:Apply。
|
||||||
|
- Capability source:`openspec-apply-change` + sm-flow apply 协议。
|
||||||
|
- 背景:tasks 1–7 已完成;真实 E2E 暴露连续 `INVALID_PROGRESS_PROTOCOL` 不会累计 `NO_GAIN`,可能在硬预算前空转。Task 8 补齐协议修复反馈与独立兜底停止。
|
||||||
|
- 实现事实(代码已在工作区,本轮补齐测试与收口):
|
||||||
|
- `ProgressProtocolViolationType` / `ProgressProtocolViolationException` 覆盖 MISSING_PREVIOUS_OBSERVATION、OUT_OF_ORDER、UNEXPECTED、MISSING_INPUT、INVALID_ENVELOPE。
|
||||||
|
- `DiagnosisProgressTracker` 独立累计连续协议错误,默认阈值 2,达到后 `stop_reason=PROGRESS_PROTOCOL_VIOLATED`。
|
||||||
|
- `HarnessToolInterceptor` 返回可修正 observation(repair_required、violation_type、missing_field、expected_previous_tool_call_id、allowed_information_gain);达阈一次 STOP_REQUIRED,再请求抛 `DiagnosisCollectionStoppedException`。
|
||||||
|
- `DiagnosisReleaseUseCase` 支持 `PROGRESS_PROTOCOL_VIOLATED`:有安全 ProgressSnapshot 发 `INSUFFICIENT_EVIDENCE`,无进展 fail closed。
|
||||||
|
- `TOOL_REQUEST_REJECTED` 记录 violation_type、repair_prompt_delivered、consecutive_protocol_violations、stop_reason,不记录参数/观察正文/异常。
|
||||||
|
- 验证:
|
||||||
|
- Focused:`DiagnosisProgressTrackerTest`、`HarnessToolInterceptorTest`、`DiagnosisReleaseUseCaseTest`、`DiagnosisAgentUseCaseTest`、`HarnessChatConfigurationTest` 通过。
|
||||||
|
- OpenSpec strict validate 通过。
|
||||||
|
- 文档:ISS-016 剩余协议停止项勾选完成;ISS-015 阶段 1 标记已完成;架构文档同步协议错误独立停止语义。
|
||||||
|
- OpenSpec tasks 8.1–8.6 全部完成。剩余 Apply 工作:无。可进入 Archive checkpoint(需用户确认是否归档 OpenSpec)。
|
||||||
|
|
||||||
|
## Archive
|
||||||
|
|
||||||
|
- Checkpoint:Archive。
|
||||||
|
- Capability source:`sm-flow` archive 协议 + `openspec-archive-change`。
|
||||||
|
- 用户确认:明确要求“执行(archive),完后提交推送”。
|
||||||
|
- devflow 档案:
|
||||||
|
- `brief.md`、`evidence.md`、`decisions.md`、`acceptance.md`
|
||||||
|
- 更新 `devflow/index.md`、`devflow/glossary/CONTEXT.md`
|
||||||
|
- OpenSpec:
|
||||||
|
- delta specs 已同步到 main specs(含新建 `diagnosis-information-gain-stop-contract`)
|
||||||
|
- change 归档至 `openspec/changes/archive/2026-07-27-diagnosis-information-gain-stop-contract/`
|
||||||
|
- 不创建独立 ADR:决策已由 OpenSpec/ISS/架构文档承载,且可通过 OpenSpec 回滚。
|
||||||
|
- 状态:archived。
|
||||||
@@ -0,0 +1,38 @@
|
|||||||
|
# Diagnosis 信息增益停止契约 Evidence
|
||||||
|
|
||||||
|
## 证据
|
||||||
|
|
||||||
|
| 来源 | 证据 | 结论 | 是否已汇报 |
|
||||||
|
| --- | --- | --- | --- |
|
||||||
|
| `HarnessEvidenceTools` / request records | Agent-facing Tool 直接使用业务 request 生成 Schema | 要增加 `previous_observation + input` 必须显式演进 Tool Schema | 是 |
|
||||||
|
| `HarnessToolInterceptor` | 曾把完整 `agentResult` 放入 Tool Response | 必须拆分控制视图与模型观察 | 是 |
|
||||||
|
| `RunContext` / `DiagnosisHarnessCore.startRun` | 结构不可变 + 线程安全 handle | 进展 tracker 放 RunContext,不扩 Redis 枚举 | 是 |
|
||||||
|
| `CanonicalInvocationStore` | 仅 begin/find/markReady/markError | tracker 只存 identity,结束时投影 | 是 |
|
||||||
|
| `RagResultProjector` | 只按 evidence 空否生成状态 | 需兼容 `relevanceLevel/relevance_level` | 是 |
|
||||||
|
| `DiagnosisReleaseUseCase` / `EvidenceGuard` | 强制 draft 非空且空 analysis=`ANALYSIS_MISSING` | 与合法无结论冲突,需统一 Release | 是 |
|
||||||
|
| `ChatApplicationUseCase.recoverBudgetExhaustion` | 未提交预算 Fallback 绕过 Release | 迁移意图到 Diagnosis Release | 是 |
|
||||||
|
| 前端 `app.js` | 已渲染 observed_facts/sources/limitations/next_steps | 不新增公开 SSE 字段 | 是 |
|
||||||
|
| 真实 E2E(实施前) | 多轮空转后 `BUDGET_EXHAUSTED`/`INTERNAL_FAILURE` | 需要信息增益停止契约 | 是 |
|
||||||
|
| 真实 E2E(实施后) | `iss016-final-20260726-a` → `MISSING_REQUIRED_CONTEXT` FALLBACK | 未知 Query 可正常业务结束 | 是 |
|
||||||
|
| Token/拒绝审计 E2E | 9 次 `INVALID_PROGRESS_PROTOCOL` 拒绝不累计 NO_GAIN | 需独立协议错误阈值与 STOP | 是 |
|
||||||
|
|
||||||
|
## Evidence-driven 结论
|
||||||
|
|
||||||
|
- 结论:Tool Envelope 是 L3 模型侧协议变更,业务 request 在 interceptor 解包后保持不变。
|
||||||
|
- 证据:三个 FunctionToolCallback inputType、adapter bridge 只收业务 JSON。
|
||||||
|
- 风险:框架 Schema/拦截器假设不匹配。
|
||||||
|
- 用户确认:不需要(技术事实)
|
||||||
|
|
||||||
|
- 结论:协议错误不得累计为 `NO_GAIN`,必须独立 `PROGRESS_PROTOCOL_VIOLATED`。
|
||||||
|
- 证据:真实审计 9 次协议拒绝 + 13 Agent 轮次;OpenSpec design 6.1。
|
||||||
|
- 风险:只返回通用错误码不足以自修复。
|
||||||
|
- 用户确认:已通过 Task 8 OpenSpec 与实现收口
|
||||||
|
|
||||||
|
- 结论:Release 是业务 Fallback 唯一决策入口;Application 不重建业务内容。
|
||||||
|
- 证据:`DiagnosisReleaseUseCase` 统一路径 + Application 窄化预算终态持久化。
|
||||||
|
- 用户确认:已确认
|
||||||
|
|
||||||
|
## 实现期补充证据
|
||||||
|
|
||||||
|
- Draft 合同失败窄化降级:非法 Draft 丢弃;仅当 ProgressSnapshot 有已验真 facts 时发 `INSUFFICIENT_EVIDENCE`。
|
||||||
|
- Task 8 focused tests:tracker / interceptor / release / agent-loop / config 全部通过。
|
||||||
@@ -0,0 +1,22 @@
|
|||||||
|
# Acceptance
|
||||||
|
|
||||||
|
## Done
|
||||||
|
|
||||||
|
- `MilvusHybridKnowledgeStore`: schema BM25 function + dense, hybridSearch+RRFRanker, dense search
|
||||||
|
- `VectorSearchService` only routes dense|hybrid to V2 store
|
||||||
|
- `VectorIndexService` writes via V2 store
|
||||||
|
- Removed knowledge-path `MilvusServiceClient` bean wiring
|
||||||
|
- Config: `milvus.collection=biz_hybrid`, `retrieval.search.mode=hybrid`
|
||||||
|
|
||||||
|
## Tests
|
||||||
|
|
||||||
|
```text
|
||||||
|
mvn -Dtest=LookupKnowledgeToolTest,KnowledgeEvidencePostProcessorTest,RagResultProjectorTest,RrfFusionTest,VectorKnowledgeSearchAdapterHybridTest,VectorSearchServiceTest,VectorIndexServiceTest test
|
||||||
|
```
|
||||||
|
|
||||||
|
EXIT:0
|
||||||
|
|
||||||
|
## Ops note
|
||||||
|
|
||||||
|
Reindex all knowledge docs into `biz_hybrid` before production hybrid search is meaningful.
|
||||||
|
Legacy `biz` collection is unused by knowledge path.
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
# Brief: rag-bm25-hybrid-drop-sdk
|
||||||
|
|
||||||
|
True dense+BM25 hybrid on a single MilvusClientV2 backend. Legacy SDK search/write for knowledge path removed. New collection `biz_hybrid` requires reindex.
|
||||||
@@ -0,0 +1,7 @@
|
|||||||
|
# Decisions
|
||||||
|
|
||||||
|
- Single backend: MilvusClientV2 only for knowledge RAG
|
||||||
|
- Drop sdk/spring/auto retrieval routing
|
||||||
|
- New collection biz_hybrid to avoid mutating legacy biz schema in place
|
||||||
|
- Hybrid = dense ANN + BM25 sparse ANN + RRFRanker
|
||||||
|
- Dense L2 enrichment for threshold compatibility on hybrid hits
|
||||||
@@ -0,0 +1,44 @@
|
|||||||
|
# Acceptance: rag-chunk-evidence-identity-dedup
|
||||||
|
|
||||||
|
## 实现结果
|
||||||
|
|
||||||
|
Delivery 1 completed:
|
||||||
|
|
||||||
|
- chunk identity fields on candidates/evidence blocks
|
||||||
|
- `KnowledgeSearchPort` + dense adapter
|
||||||
|
- evidenceKey dedup + maxChunksPerDocument + return-n
|
||||||
|
- retrieve-k on lookup tool
|
||||||
|
- projector keeps same-source distinct chunks; `document_id` is chunk-scoped
|
||||||
|
|
||||||
|
## 验证
|
||||||
|
|
||||||
|
### 脚本验证
|
||||||
|
|
||||||
|
```text
|
||||||
|
mvn -q "-Dtest=LookupKnowledgeToolTest,KnowledgeEvidencePostProcessorTest,RagResultProjectorTest" test
|
||||||
|
```
|
||||||
|
|
||||||
|
结果:通过(exit 0)
|
||||||
|
|
||||||
|
### 静态验证
|
||||||
|
|
||||||
|
- Search/port wiring reviewed against OpenSpec tasks
|
||||||
|
- No new legacy SDK dependency added
|
||||||
|
|
||||||
|
### 浏览器/人工验证
|
||||||
|
|
||||||
|
未运行(纯检索契约变更,无 UI)
|
||||||
|
|
||||||
|
### 未验证
|
||||||
|
|
||||||
|
- 全量 harness E2E / 真实 Milvus 联调(Delivery 2 前可补)
|
||||||
|
- 生产配置默认 retrieve-k/return-n 调优
|
||||||
|
|
||||||
|
## 归档状态
|
||||||
|
|
||||||
|
- OpenSpec change ready to archive
|
||||||
|
- User pre-authorized archive for sm-flow staged delivery
|
||||||
|
|
||||||
|
## 后续
|
||||||
|
|
||||||
|
- Delivery 2: milvus hybrid search (separate change)
|
||||||
@@ -0,0 +1,29 @@
|
|||||||
|
# Brief: rag-chunk-evidence-identity-dedup
|
||||||
|
|
||||||
|
## 背景
|
||||||
|
|
||||||
|
lookup_knowledge 同文档多 chunk 在后处理与投影阶段被 source 级去重吞掉。Hybrid 多路召回前必须先修证据身份与裁剪契约。
|
||||||
|
|
||||||
|
## 目标
|
||||||
|
|
||||||
|
- chunk 级 evidenceKey 身份
|
||||||
|
- 按 evidenceKey 去重 + 每文档 chunk 上限
|
||||||
|
- retrieve-k / return-n 分离
|
||||||
|
- Projector 保留同 source 不同 chunk
|
||||||
|
- 薄 KnowledgeSearchPort,为 Delivery 2 hybrid 铺路
|
||||||
|
|
||||||
|
## 范围
|
||||||
|
|
||||||
|
Delivery 1 only(见 `mvp/engineering/rag/Milvus-Hybrid接入清单.md` §1.1)。
|
||||||
|
|
||||||
|
## 非目标
|
||||||
|
|
||||||
|
hybrid schema、BM25、删 SDK、session dedup、邻块重建、模型 rerank。
|
||||||
|
|
||||||
|
## 分档
|
||||||
|
|
||||||
|
standard
|
||||||
|
|
||||||
|
## 关联 OpenSpec
|
||||||
|
|
||||||
|
`openspec/changes/rag-chunk-evidence-identity-dedup`
|
||||||
@@ -0,0 +1,102 @@
|
|||||||
|
# Decisions: rag-chunk-evidence-identity-dedup
|
||||||
|
|
||||||
|
## Capability sources
|
||||||
|
|
||||||
|
- sm-flow orchestration
|
||||||
|
- OpenSpec fallback protocol (file-based propose/apply/archive) — external openspec-propose/apply skills used as reference; execution via sm-flow fallback
|
||||||
|
- grill: fallback built-in protocol
|
||||||
|
- audit: fallback built-in protocol
|
||||||
|
|
||||||
|
## Scale
|
||||||
|
|
||||||
|
standard
|
||||||
|
|
||||||
|
## Clarify
|
||||||
|
|
||||||
|
- Problem: same-document multi-chunk evidence collapsed by source-level dedup.
|
||||||
|
- Outcome: Delivery 1 foundation before hybrid Delivery 2.
|
||||||
|
- Slug: `rag-chunk-evidence-identity-dedup`
|
||||||
|
- User authorized apply + archive in advance for sm-flow staged changes.
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
- Read: `devflow/glossary/CONTEXT.md`, modular-rag-pipeline brief, `openspec/specs/rag-knowledge-retrieval`, `rag-log-projections`, checklist doc §1.1
|
||||||
|
- Constraints into OpenSpec:
|
||||||
|
- L0 hint-only remains
|
||||||
|
- Do not thicken legacy SDK path
|
||||||
|
- Agent tool name/input stable
|
||||||
|
- Hybrid out of scope this change
|
||||||
|
|
||||||
|
## Question pool (grill)
|
||||||
|
|
||||||
|
| # | Dimension | Mode | Question | Status |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| Q1 | 术语 | evidence-driven | evidenceKey / document_id 语义? | Resolved: evidenceKey=chunk id; projected document_id=evidenceKey |
|
||||||
|
| Q2 | 边界 | evidence-driven | Delivery 1 vs 2 边界? | Resolved: per checklist; no schema/hybrid/SDK delete |
|
||||||
|
| Q3 | 验收 | evidence-driven | 如何验收多 chunk? | Resolved: unit tests multi-chunk keep + projector |
|
||||||
|
| Q4 | 接口 | user-interview | document_id 改为 chunk 级是否可接受? | **Pre-authorized by user** via “apply/archive 直接授权” + prior design agreement on scheme A (document_id=evidenceKey). Recorded as accepted behavior change. |
|
||||||
|
| Q5 | 技术 | evidence-driven | SearchPort 是否本 change 必须? | Resolved: thin port required as foundation |
|
||||||
|
|
||||||
|
### Evidence-driven conclusions (reported)
|
||||||
|
|
||||||
|
1. Current collapse points: `KnowledgeEvidencePostProcessor.sourceKey` and `RagResultProjector` source fallback.
|
||||||
|
2. Metadata already has docId/chunkIndex on write path; not first-class on read path.
|
||||||
|
3. Existing main-spec still says source-level dedup — this change intentionally deltas that requirement.
|
||||||
|
|
||||||
|
### User-interview
|
||||||
|
|
||||||
|
- Q4 accepted under prior design alignment (scheme A) and explicit apply authorization for this sm-flow run. No remaining open product preference questions for Delivery 1.
|
||||||
|
|
||||||
|
## Audit
|
||||||
|
|
||||||
|
Module chain:
|
||||||
|
|
||||||
|
```text
|
||||||
|
LookupKnowledgeTool -> SearchPort -> Retriever -> PostProcessor -> Packer -> Assembler -> RagResultProjector
|
||||||
|
```
|
||||||
|
|
||||||
|
Risks:
|
||||||
|
|
||||||
|
1. Agent payload growth — mitigated by return-n + maxChunksPerDocument + projector budgets.
|
||||||
|
2. document_id semantic shift — documented L3 behavior change; tests updated.
|
||||||
|
3. Old data without chunkIndex — vector id fallback.
|
||||||
|
|
||||||
|
No ADR conflict with modular RAG L0/L1 boundary.
|
||||||
|
|
||||||
|
## Cross-artifact alignment
|
||||||
|
|
||||||
|
| From | To | Status |
|
||||||
|
|---|---|---|
|
||||||
|
| brief goals | proposal | 已对齐 |
|
||||||
|
| proposal scope | design decisions | 已对齐 |
|
||||||
|
| design identity/dedup/port | specs | 已对齐 |
|
||||||
|
| specs scenarios | tasks | 已对齐 |
|
||||||
|
|
||||||
|
## Interface impact
|
||||||
|
|
||||||
|
- L2 internal DTO
|
||||||
|
- L3 Agent `document_id` chunk-scoped
|
||||||
|
|
||||||
|
## Commit gate
|
||||||
|
|
||||||
|
- proposal/design/specs/tasks present
|
||||||
|
- no open user-interview blockers for Delivery 1
|
||||||
|
- apply authorized by user at sm-flow start
|
||||||
|
|
||||||
|
## Pre-apply research
|
||||||
|
|
||||||
|
Reference files:
|
||||||
|
|
||||||
|
- `LookupKnowledgeTool.java`
|
||||||
|
- `KnowledgeDocumentRetriever.java`
|
||||||
|
- `KnowledgeEvidencePostProcessor.java`
|
||||||
|
- `RagResultProjector.java`
|
||||||
|
- `LookupKnowledgeToolTest.java`
|
||||||
|
- `RagResultProjectorTest.java`
|
||||||
|
- `mvp/engineering/rag/Milvus-Hybrid接入清单.md`
|
||||||
|
|
||||||
|
Stack notes:
|
||||||
|
|
||||||
|
- No MQ/request envelope changes
|
||||||
|
- Spring `@Value` config pattern for rag.* keys
|
||||||
|
- Tests use ReflectionTestUtils + Mockito
|
||||||
@@ -0,0 +1,16 @@
|
|||||||
|
# Evidence: rag-chunk-evidence-identity-dedup
|
||||||
|
|
||||||
|
## 代码证据(变更前)
|
||||||
|
|
||||||
|
- `KnowledgeEvidencePostProcessor.sourceKey` 使用 source/title 去重
|
||||||
|
- `RagResultProjector` 用 source 回退 document_id 并 HashSet 去重
|
||||||
|
- 写入路径 metadata 已有 docId/chunkIndex,读路径未一等化
|
||||||
|
|
||||||
|
## 规格证据
|
||||||
|
|
||||||
|
- 旧 `openspec/specs/rag-knowledge-retrieval` 要求 source 级 dedup(本 change 以 delta 修正)
|
||||||
|
- `mvp/engineering/rag/Milvus-Hybrid接入清单.md` §1.1 定义 Delivery 1 地基
|
||||||
|
|
||||||
|
## 验证证据
|
||||||
|
|
||||||
|
- 单测覆盖 multi-chunk keep / true-dup merge / maxChunksPerDocument / projector same-source multi-chunk
|
||||||
@@ -0,0 +1,23 @@
|
|||||||
|
# Acceptance: rag-hybrid-search-rrf
|
||||||
|
|
||||||
|
## Result
|
||||||
|
|
||||||
|
Hybrid mode implemented on KnowledgeSearchPort:
|
||||||
|
|
||||||
|
- dense unfiltered + dense filtered + lexical rank over union
|
||||||
|
- RRF fusion by evidenceKey
|
||||||
|
- dense-compatible score preserved for thresholds
|
||||||
|
- default mode remains dense
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
```text
|
||||||
|
mvn -q "-Dtest=RrfFusionTest,VectorKnowledgeSearchAdapterHybridTest,LookupKnowledgeToolTest" test
|
||||||
|
```
|
||||||
|
|
||||||
|
Pass.
|
||||||
|
|
||||||
|
## Residual
|
||||||
|
|
||||||
|
- True Milvus BM25/sparse schema + reindex still follow-up
|
||||||
|
- Lexical path only ranks dense-recalled candidates (does not expand pure-term misses outside dense topK)
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
# Brief: rag-hybrid-search-rrf
|
||||||
|
|
||||||
|
Delivery 2 after chunk identity. Enable hybrid multi-path + RRF on KnowledgeSearchPort without legacy SDK hybrid API. True BM25 schema rebuild is staged follow-up; this change ships sparse-lite lexical ranking over dense candidate union + filtered/unfiltered dense fusion.
|
||||||
@@ -0,0 +1,19 @@
|
|||||||
|
# Decisions: rag-hybrid-search-rrf
|
||||||
|
|
||||||
|
## Capability
|
||||||
|
|
||||||
|
sm-flow + OpenSpec fallback; apply pre-authorized.
|
||||||
|
|
||||||
|
## Depends
|
||||||
|
|
||||||
|
Delivery 1 archived.
|
||||||
|
|
||||||
|
## Grill (compressed, pre-authorized)
|
||||||
|
|
||||||
|
- Q: Full BM25 schema now? A: No — sparse-lite + RRF first; schema rebuild follow-up.
|
||||||
|
- Q: Default mode? A: dense default; hybrid opt-in.
|
||||||
|
- Q: Threshold score? A: keep dense-compatible L2 mapping.
|
||||||
|
|
||||||
|
## Design
|
||||||
|
|
||||||
|
Hybrid paths: dense unfiltered + dense filtered + lexical rank over union; RRF fuse by evidenceKey.
|
||||||
@@ -0,0 +1,5 @@
|
|||||||
|
# Evidence
|
||||||
|
|
||||||
|
- Delivery 1 identity/port foundation required
|
||||||
|
- RRF utility and hybrid adapter unit tests green
|
||||||
|
- Lexical sparse-lite intentionally intermediate until BM25 schema
|
||||||
@@ -0,0 +1,39 @@
|
|||||||
|
# Acceptance: rag-eval-hybrid-baseline
|
||||||
|
|
||||||
|
## Tasks
|
||||||
|
|
||||||
|
All tasks in OpenSpec `tasks.md` checked, including apply-discovered 6.x quality-gate fix.
|
||||||
|
|
||||||
|
## 静态验证
|
||||||
|
|
||||||
|
- Snapshot generator path: no required `retrieval.vector-store.mode`.
|
||||||
|
- README documents hybrid generation and offline/live split.
|
||||||
|
|
||||||
|
## 脚本验证
|
||||||
|
|
||||||
|
```text
|
||||||
|
.\scripts\prepare_rag_eval_seed.ps1
|
||||||
|
.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode hybrid -SkipEval
|
||||||
|
python scripts\eval_rag_retrieval.py --json-report eval/rag-retrieval/reports/baseline.json --markdown-report eval/rag-retrieval/reports/baseline.md
|
||||||
|
# Result: Evaluated 7 cases: passRate=1.0, recall@5=1.0, failed=0
|
||||||
|
|
||||||
|
mvn -Dtest=RetrievalScoreNormalizerTest,KnowledgeEvidencePostProcessorTest,LookupKnowledgeToolTest,VectorSearchServiceTest,VectorKnowledgeSearchAdapterHybridTest test
|
||||||
|
# exit 0
|
||||||
|
```
|
||||||
|
|
||||||
|
Fixture sample meta: `searchMode=hybrid`, `kbScope=rag-eval`.
|
||||||
|
Fallback case: `UNFILTERED_VECTOR_RETRY` + `filtered_vector_low_quality`.
|
||||||
|
|
||||||
|
## 浏览器/人工
|
||||||
|
|
||||||
|
- 未做 UI 验证。
|
||||||
|
|
||||||
|
## 未验证 / 后续
|
||||||
|
|
||||||
|
- Dense vs hybrid dual-directory comparison report (knife-2).
|
||||||
|
- CI wiring of offline eval as required gate (optional process).
|
||||||
|
- Long-term calibration of hybrid PRECISE distribution under denseDistance quality.
|
||||||
|
|
||||||
|
## Specs
|
||||||
|
|
||||||
|
Main spec synced: `openspec/specs/rag-eval-offline-baseline/spec.md`.
|
||||||
@@ -0,0 +1,16 @@
|
|||||||
|
# Brief: rag-eval-hybrid-baseline
|
||||||
|
|
||||||
|
## Background
|
||||||
|
|
||||||
|
Offline RAG eval (golden × fixture × key-field baseline) existed but generator/docs still used dead `retrieval.vector-store.mode=spring`. Fixtures lacked search meta and did not reflect hybrid main path.
|
||||||
|
|
||||||
|
## Goals (knife-1 only)
|
||||||
|
|
||||||
|
- Snapshot generation uses `retrieval.search.mode` (default hybrid; dense override).
|
||||||
|
- Fixtures record `searchMode` / `kbScope`.
|
||||||
|
- README documents hybrid-era offline vs live loop.
|
||||||
|
- Best-effort live seed + regenerate fixtures + update baseline.
|
||||||
|
|
||||||
|
## Non-goals
|
||||||
|
|
||||||
|
Dense/hybrid dual fixture trees; golden mustNot/chunk/level hard gates; new eval frameworks.
|
||||||
@@ -0,0 +1,18 @@
|
|||||||
|
# Decisions: rag-eval-hybrid-baseline(最终版)
|
||||||
|
|
||||||
|
## Process
|
||||||
|
|
||||||
|
sm-flow standard-lean: Discover → Commit → Apply → Archive.
|
||||||
|
|
||||||
|
## Key decisions
|
||||||
|
|
||||||
|
1. Replace eval generator `vector-store.mode` with `retrieval.search.mode` (default hybrid).
|
||||||
|
2. Fixture meta: `searchMode`, `kbScope` when set.
|
||||||
|
3. Knife-2 (dual fixtures / mustNot golden) deferred.
|
||||||
|
4. Live refresh succeeded in apply env; baseline updated to hybrid snapshots.
|
||||||
|
5. **Quality gate refinement (apply-found):** hybrid absolute quality for `isLowQuality` / relevance uses optional dense L2 (`denseDistance`); does not overwrite hybrid scoreLabel or RRF order. Rank mapping remains fallback when dense missing.
|
||||||
|
|
||||||
|
## Trade-offs
|
||||||
|
|
||||||
|
- Extra dense ANN on hybrid path for gate calibration (latency) vs correct filter-fallback behavior.
|
||||||
|
- relevance_level still not a hard golden assertion (ordinal vs absolute mix).
|
||||||
@@ -0,0 +1,17 @@
|
|||||||
|
# Evidence: rag-eval-hybrid-baseline
|
||||||
|
|
||||||
|
## Pre-change
|
||||||
|
|
||||||
|
- `generate_rag_lookup_snapshots.ps1` passed `-Dretrieval.vector-store.mode=spring`.
|
||||||
|
- Fixtures had `caseId/query/retrievedAt/lookupResult` only.
|
||||||
|
- Offline eval already supported Hit levels, recall@K, baseline diff.
|
||||||
|
|
||||||
|
## User decisions
|
||||||
|
|
||||||
|
- Scope: knife-1 only (no dual fixture dirs).
|
||||||
|
- Acceptance: wiring required; fixture refresh best-effort (env allowed full refresh).
|
||||||
|
|
||||||
|
## Apply-discovered
|
||||||
|
|
||||||
|
- After hybrid refresh, `chat-l0-filter-fallback` failed: pure rank→quality made topSimilarity=1.0 on decoy-only filtered hits → no unfiltered retry.
|
||||||
|
- Fix: optional `denseDistance` on hybrid hits; quality gate uses L2 when present; sort order remains RRF.
|
||||||
@@ -0,0 +1,41 @@
|
|||||||
|
# Acceptance: rag-quality-score-unify
|
||||||
|
|
||||||
|
## Tasks
|
||||||
|
|
||||||
|
OpenSpec `tasks.md` 全部 `[x]`(1.1–6.2)。
|
||||||
|
|
||||||
|
## 静态验证
|
||||||
|
|
||||||
|
- 生产路径 grep:无 `bm25_only_no_dense` 发射、无 hybrid L2 enrichment(仅 Labels canonicalize 兼容旧串)。
|
||||||
|
- 架构文档 §6 与 `application.yml` 注释已对齐 quality 契约。
|
||||||
|
|
||||||
|
## 脚本验证
|
||||||
|
|
||||||
|
```text
|
||||||
|
mvn -Dtest=RetrievalScoreNormalizerTest,KnowledgeEvidencePostProcessorTest,LookupKnowledgeToolTest,VectorSearchServiceTest,VectorKnowledgeSearchAdapterHybridTest test
|
||||||
|
```
|
||||||
|
|
||||||
|
| 套件 | 结果 |
|
||||||
|
|---|---|
|
||||||
|
| RetrievalScoreNormalizerTest | 4 passed |
|
||||||
|
| KnowledgeEvidencePostProcessorTest | 6 passed |
|
||||||
|
| LookupKnowledgeToolTest | 7 passed |
|
||||||
|
| VectorSearchServiceTest | 2 passed |
|
||||||
|
| VectorKnowledgeSearchAdapterHybridTest | 1 passed |
|
||||||
|
|
||||||
|
(PowerShell 可能将 JVM warning 标为 exit 1;日志中为 BUILD SUCCESS / Failures: 0。)
|
||||||
|
|
||||||
|
## 浏览器 / 人工验证
|
||||||
|
|
||||||
|
- 未跑:live `lookup_knowledge` hybrid vs dense 对照、生产阈值标定。
|
||||||
|
|
||||||
|
## 未验证
|
||||||
|
|
||||||
|
| 项 | 风险 | 建议 |
|
||||||
|
|---|---|---|
|
||||||
|
| 真实 Milvus hybrid 联调 | 序/质量分布与单测 mock 有差 | 启动服务后固定 query 集切 mode 对比 |
|
||||||
|
| 阈值 0.75/0.5 在 hybrid rank 分下的标定 | retry/PRECISE 偏多或偏少 | 看 trace topSimilarity 再调 yml |
|
||||||
|
|
||||||
|
## Specs 同步
|
||||||
|
|
||||||
|
- 主规格新增:`openspec/specs/rag-retrieval-quality-score/spec.md`(archive 时从 delta 同步)。
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
# Brief: rag-quality-score-unify
|
||||||
|
|
||||||
|
## Background
|
||||||
|
|
||||||
|
真 BM25 hybrid(dense + BM25 + RRF)已上线,但后处理仍把 hybrid 结果伪装成 L2 做 `normalizeL2`,并用 L0 domain/entity/keyword contains 加分改序。排序权威与质量闸门分裂,词面信号被 BM25 与后处理双重计分。
|
||||||
|
|
||||||
|
## Goals
|
||||||
|
|
||||||
|
- 一级 `scoreLabel` 仅 `dense` | `hybrid`
|
||||||
|
- 唯一 `toQualityScore`;后处理 label-agnostic
|
||||||
|
- 排序主序 = 检索 `originalRank`;去掉关键词 boost 改序
|
||||||
|
- hybrid quality = 本轮 rank 纯映射(不做 max(rank, denseSim)、不为闸门回填 L2)
|
||||||
|
- 保留 `mode=dense` 作同库召回对照;线上默认 hybrid
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
内部 RAG:store 发射、normalizer、evidence post-process、单测、架构文档 §6。
|
||||||
|
|
||||||
|
## Non-goals
|
||||||
|
|
||||||
|
精排 / query rewrite / 邻块、schema rebuild、改 Agent ACI 字段名、删除 dense 对照 mode。
|
||||||
@@ -0,0 +1,25 @@
|
|||||||
|
# Decisions: rag-quality-score-unify(最终版)
|
||||||
|
|
||||||
|
## Scale / process
|
||||||
|
|
||||||
|
- sm-flow standard:Discover → Commit → Apply → Archive
|
||||||
|
- Committed OpenSpec:`openspec/changes/rag-quality-score-unify/`(归档后见 archive 目录)
|
||||||
|
- 废止:`rag-bm25-hybrid-drop-sdk` 中「dense L2 enrichment for threshold compatibility」
|
||||||
|
|
||||||
|
## Key decisions
|
||||||
|
|
||||||
|
1. **Label**:仅 `dense` | `hybrid`;旧别名 canonicalize。
|
||||||
|
2. **Normalizer**:唯一 `toQualityScore`;dense=L2 公式;hybrid=rank 线性映射(batchSize)。
|
||||||
|
3. **Store**:hybrid 不回填 L2、不发 `bm25_only_*`;返回序即 RRF 序。
|
||||||
|
4. **Post-process**:`originalRank` ASC;L0 重叠只写 hitReasons;relevance/low-quality 只看 qualityScore;PRECISE 不要求 hint support。
|
||||||
|
5. **Mode**:hybrid 主路径;dense 同库对照(架构 §6.0)。
|
||||||
|
|
||||||
|
## Trade-offs
|
||||||
|
|
||||||
|
- hybrid quality 为序数分,跨 query 绝对值不可比;阈值可能需后续标定。
|
||||||
|
- 去掉 boost 改序后,「词面热语义冷」不再被后处理抬升;词面交给 BM25+RRF。
|
||||||
|
|
||||||
|
## Risks accepted
|
||||||
|
|
||||||
|
- `relevance_level` / unfiltered retry 分布变化(产品已接受)。
|
||||||
|
- 未做 live E2E / 人工 hybrid 对照评测(见 acceptance 未验证项)。
|
||||||
@@ -0,0 +1,25 @@
|
|||||||
|
# Evidence: rag-quality-score-unify
|
||||||
|
|
||||||
|
## Code (pre-change)
|
||||||
|
|
||||||
|
- `MilvusHybridKnowledgeStore.searchHybrid`:RRF 后并行 dense 回填 L2;BM25-only → `bm25_only_no_dense` + maxL2。
|
||||||
|
- `KnowledgeEvidencePostProcessor`:一律 `normalizeL2(score)` + domain/entity/keyword/source_type 加分,按 `finalScore` 降序;PRECISE 需 `hasHintSupport`。
|
||||||
|
|
||||||
|
## User decisions (grill)
|
||||||
|
|
||||||
|
| ID | 结论 |
|
||||||
|
|---|---|
|
||||||
|
| Q1 | 一级 label 仅 dense/hybrid;bm25_only 不作正式 label |
|
||||||
|
| Q2 | 后处理去掉 contains 加分改序,保 originalRank |
|
||||||
|
| Q3 | 唯一 toQualityScore;后处理统一 |
|
||||||
|
| Q4 | dense mode 保留作对照 |
|
||||||
|
| Q7 | 接受 relevance_level / retry 分布变化 |
|
||||||
|
| Q8 | hybrid quality = **纯 rank 映射** |
|
||||||
|
|
||||||
|
## Post-change anchors
|
||||||
|
|
||||||
|
- `RetrievalScoreLabels` / `RetrievalScoreNormalizer`
|
||||||
|
- `MilvusHybridKnowledgeStore`(无 L2 overwrite / 无 bm25_only 发射)
|
||||||
|
- `KnowledgeEvidencePostProcessor`(rank sort + explain-only L0 overlap)
|
||||||
|
- OpenSpec delta:`rag-retrieval-quality-score`
|
||||||
|
- 架构:`mvp/architecture/RAG知识检索架构.md` §6
|
||||||
+58
-91
@@ -1,108 +1,75 @@
|
|||||||
# 文档索引
|
# 文档索引
|
||||||
|
|
||||||
## 📂 目录结构
|
**更新日期**:2026-07-29
|
||||||
|
|
||||||
```
|
## 目录结构
|
||||||
|
|
||||||
|
```text
|
||||||
docs/
|
docs/
|
||||||
├── README.md # 项目文档总览
|
├── INDEX.md # 本索引
|
||||||
├── INDEX.md # 本索引文件
|
├── learning/ # 早期学习笔记(可能过时)
|
||||||
│
|
├── analysis/ # 早期代码/问题分析
|
||||||
├── learning/ # 📚 学习笔记(个人学习理解)
|
├── reports/ # 历史修复/验证报告
|
||||||
│ ├── 00-项目学习路径.md
|
└── guides/ # 操作指南
|
||||||
│ ├── 01~08-*.md # 按学习顺序编号
|
|
||||||
│ └── README.md
|
mvp/ # 现行 MVP 文档(主入口)
|
||||||
│
|
├── architecture/ # 现行架构:系统现在怎么跑
|
||||||
├── analysis/ # 🔍 分析笔记(代码/问题分析)
|
├── engineering/ # 工程纪要:问题 / 决策 / E2E
|
||||||
│ ├── essence-report-*.md
|
├── issues/ # 未完成事项
|
||||||
│ ├── explore-report.md
|
├── tables/ # 表结构
|
||||||
│ ├── chunking-issues-analysis.md
|
├── demo/ # Demo
|
||||||
│ └── 功能分析报告.md
|
└── eval/ # 诊断评测材料
|
||||||
│
|
|
||||||
├── reports/ # 📝 临时报告(修复/验证报告)
|
|
||||||
│ ├── 修复报告-*.md
|
|
||||||
│ ├── 验证报告-*.md
|
|
||||||
│ └── 日志配置完成总结.md
|
|
||||||
│
|
|
||||||
└── guides/ # 📖 指南文档
|
|
||||||
└── 日志配置与分析指南.md
|
|
||||||
```
|
```
|
||||||
|
|
||||||
**⚠️ 注意:MVP 架构设计文档已移至项目根目录 `../mvp/`**
|
**现行架构、工程纪要、表结构、Issue 均以 [mvp/README.md](../mvp/README.md) 为准。**
|
||||||
|
本目录 `learning/` / `analysis/` / `reports/` 偏早期学习与历史记录,**可能与当前实现不一致**。
|
||||||
查看 [mvp/README.md](../mvp/README.md) 了解 MVP 架构、数据库设计、实施计划等。
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## 🚀 快速导航
|
## 快速导航
|
||||||
|
|
||||||
### 我是新人/学习者
|
### 我是开发者(优先)
|
||||||
1. [项目学习路径](learning/00-项目学习路径.md) - 从这里开始
|
|
||||||
2. [learning/README.md](learning/README.md) - 学习笔记索引
|
|
||||||
3. 按编号顺序阅读 `learning/` 目录下的文档
|
|
||||||
|
|
||||||
### 我是开发者
|
1. [mvp/README.md](../mvp/README.md) — MVP 总入口
|
||||||
👉 **MVP 架构设计文档已移至 `../mvp/`**
|
2. [mvp/architecture/](../mvp/architecture/) — 现行架构
|
||||||
|
3. [mvp/engineering/](../mvp/engineering/) — 工程纪要(RAG / 诊断 E2E)
|
||||||
|
4. [mvp/issues/](../mvp/issues/) — 活跃 Issue
|
||||||
|
|
||||||
请查看 [mvp/README.md](../mvp/README.md) 了解:
|
### RAG 工程纪要(已迁入 mvp)
|
||||||
- MVP 架构设计
|
|
||||||
- 数据库设计和表结构
|
|
||||||
- 实施规划(Phase 1/2/3)
|
|
||||||
- 会话管理设计
|
|
||||||
|
|
||||||
### 我要查看分析报告
|
1. [RAG 排序:多路召回与 RRF](../mvp/engineering/rag/RAG排序-多路召回与RRF.md)
|
||||||
1. [分析笔记目录](analysis/) - 代码分析和问题分析
|
2. [Hybrid 之后的 qualityScore 与后处理](../mvp/engineering/rag/RAG-Hybrid质量分与后处理.md)
|
||||||
2. [临时报告目录](reports/) - 修复和验证报告
|
3. [Agent 如何读 relevance_level](../mvp/engineering/rag/RAG-Agent如何读relevance_level.md)
|
||||||
|
4. [RAG 离线评测:基线设计](../mvp/engineering/rag/RAG离线评测-基线设计.md)
|
||||||
|
5. [Milvus hybrid 接入清单](../mvp/engineering/rag/Milvus-Hybrid接入清单.md)
|
||||||
|
6. [RAG 审计补丁 E2E](../mvp/engineering/rag/RAG审计补丁-stepid-query-E2E验收.md)
|
||||||
|
7. 架构对照:[RAG 知识检索架构](../mvp/architecture/RAG知识检索架构.md)、[RAG 检索可观测性与审计](../mvp/architecture/RAG检索可观测性与审计.md)
|
||||||
|
|
||||||
|
### 诊断全流程(已迁入 mvp)
|
||||||
|
|
||||||
|
1. [一次诊断到底发生了什么](../mvp/engineering/diagnosis/一次诊断全流程-E2E导读.md)
|
||||||
|
|
||||||
|
### 我是新人 / 想看早期学习笔记
|
||||||
|
|
||||||
|
1. [项目学习路径](learning/00-项目学习路径.md)
|
||||||
|
2. [learning/README.md](learning/README.md)
|
||||||
|
3. 注意:内容可能过时,实现以 `mvp/architecture` 为准
|
||||||
|
|
||||||
|
### 分析 / 历史报告 / 指南
|
||||||
|
|
||||||
|
- [analysis/](analysis/) — 早期分析
|
||||||
|
- [reports/](reports/) — 历史修复与验证报告
|
||||||
|
- [guides/日志配置与分析指南.md](guides/日志配置与分析指南.md)
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## 📚 学习笔记 (learning/)
|
## 文档维护
|
||||||
|
|
||||||
按学习顺序编号,建议按顺序阅读:
|
| 类型 | 位置 |
|
||||||
|
|---|---|
|
||||||
1. [00-项目学习路径](learning/00-项目学习路径.md)
|
| 现行架构 | `mvp/architecture/` |
|
||||||
2. [01-AI-Ops-核心设计-Essence报告](learning/01-AI-Ops-核心设计-Essence报告.md)
|
| 工程纪要(问题/决策/E2E) | `mvp/engineering/` |
|
||||||
3. [02-outputKey-深度解析](learning/02-outputKey-深度解析.md)
|
| Issue | `mvp/issues/` |
|
||||||
4. [03-核心疑问解答](learning/03-核心疑问解答.md)
|
| 表结构 | `mvp/tables/` |
|
||||||
5. [04-RAG-分块策略-Essence报告](learning/04-RAG-分块策略-Essence报告.md)
|
| 早期学习 | `docs/learning/`(归档向,不充当现行规范) |
|
||||||
6. [05-文件上传自动索引-Essence报告](learning/05-文件上传自动索引-Essence报告.md)
|
| 操作指南 | `docs/guides/` |
|
||||||
7. [06-RAG查询流程-Essence报告](learning/06-RAG查询流程-Essence报告.md)
|
|
||||||
8. [07-Tool定义方式对比与优化](learning/07-Tool定义方式对比与优化.md)
|
|
||||||
9. [08-MethodToolCallback-vs-ToolCallingManager深度分析](learning/08-MethodToolCallback-vs-ToolCallingManager深度分析.md)
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 🔍 分析笔记 (analysis/)
|
|
||||||
|
|
||||||
代码分析和问题分析文档:
|
|
||||||
|
|
||||||
- [essence-report-rag.md](analysis/essence-report-rag.md)
|
|
||||||
- [essence-report-rag-chunking.md](analysis/essence-report-rag-chunking.md)
|
|
||||||
- [explore-report.md](analysis/explore-report.md)
|
|
||||||
- [chunking-issues-analysis.md](analysis/chunking-issues-analysis.md)
|
|
||||||
- [功能分析报告.md](analysis/功能分析报告.md)
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 📝 临时报告 (reports/)
|
|
||||||
|
|
||||||
修复报告和验证报告:
|
|
||||||
|
|
||||||
- [修复报告-多轮对话时间查询缓存问题](reports/修复报告-多轮对话时间查询缓存问题.md)
|
|
||||||
- [验证报告-时间查询问题](reports/验证报告-时间查询问题.md)
|
|
||||||
- [日志配置完成总结](reports/日志配置完成总结.md)
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 📖 指南文档 (guides/)
|
|
||||||
|
|
||||||
- [日志配置与分析指南](guides/日志配置与分析指南.md)
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 🔄 文档维护
|
|
||||||
|
|
||||||
- **学习笔记** 放在 `learning/` 目录,按编号顺序命名
|
|
||||||
- **分析笔记** 放在 `analysis/` 目录
|
|
||||||
- **临时报告** 放在 `reports/` 目录
|
|
||||||
- **指南文档** 放在 `guides/` 目录
|
|
||||||
- **MVP 架构设计** 已移至项目根目录 `../mvp/`(包含架构、数据库、实施计划)
|
|
||||||
|
|||||||
+89
-143
@@ -1,38 +1,65 @@
|
|||||||
# RAG Retrieval Baseline
|
# RAG Retrieval Baseline
|
||||||
|
|
||||||
This directory contains the offline retrieval baseline for the RAG refactor.
|
Offline regression harness for `lookup_knowledge` **after** hybrid retrieval + qualityScore post-process.
|
||||||
|
|
||||||
The baseline is intentionally narrower than full diagnosis evaluation. It checks
|
It checks whether fixed queries still recover expected documents, breadcrumbs, keywords, and pipeline behaviors (filter / unfiltered retry). It is **not** a full diagnosis-agent E2E.
|
||||||
whether fixed retrieval queries can recover expected documents, breadcrumbs, and
|
|
||||||
evidence keywords before changing L0 behavior, query augmentation, evidence
|
Production knowledge path: `MilvusHybridKnowledgeStore` with `retrieval.search.mode=hybrid` (dense+BM25+RRF).
|
||||||
post-processing, or Spring AI VectorStore integration.
|
`mode=dense` remains a same-collection baseline for recall comparison (not a second index).
|
||||||
|
|
||||||
|
Related design notes:
|
||||||
|
|
||||||
|
- `mvp/engineering/rag/RAG-Hybrid质量分与后处理.md`
|
||||||
|
- `mvp/engineering/rag/RAG-Agent如何读relevance_level.md`
|
||||||
|
- `mvp/architecture/RAG知识检索架构.md` §6
|
||||||
|
|
||||||
|
## Offline vs live
|
||||||
|
|
||||||
|
| Layer | What | Needs live stack? |
|
||||||
|
|-------|------|-------------------|
|
||||||
|
| **Offline** | `fixtures/*.json` × `golden-cases.json` → pass/fail + baseline diff | **No** (no Milvus/LLM/Boot) |
|
||||||
|
| **Snapshot generate** | Real `LookupKnowledgeTool` writes fixtures | **Yes** (embedding + Milvus + DB/L0 as configured) |
|
||||||
|
| **Live smoke** | optional `eval_rag_live_acceptance.py` | Yes (running app) |
|
||||||
|
|
||||||
|
Daily CI / local quick check: **offline only**.
|
||||||
|
After changing retrieval, indexing, or search mode: **regenerate fixtures**, then offline eval, then update baseline if the diff is intentional.
|
||||||
|
|
||||||
## Layout
|
## Layout
|
||||||
|
|
||||||
```text
|
```text
|
||||||
eval/rag-retrieval/
|
eval/rag-retrieval/
|
||||||
cases/golden-cases.json Fixed retrieval golden cases
|
cases/golden-cases.json Fixed queries + expectations
|
||||||
seed-docs/*.md Canonical docs imported into the live KB for real-tool eval
|
seed-docs/*.md Canonical docs for live snapshot (kb_scope: rag-eval)
|
||||||
fixtures/*.json Saved retrieval fixtures for each case
|
fixtures/*.json Frozen lookupResult snapshots (+ searchMode meta)
|
||||||
reports/baseline.json Machine-readable baseline report
|
reports/baseline.json|md Last accepted offline report
|
||||||
reports/baseline.md Human-readable baseline report
|
reports/baseline-diff.* Optional diff vs previous report
|
||||||
reports/baseline-diff.* Optional diff reports
|
|
||||||
reports/live-post-reindex.* Optional live acceptance reports
|
|
||||||
```
|
```
|
||||||
|
|
||||||
## Seed Docs + Import/Reindex
|
## Fixture shape (minimum)
|
||||||
|
|
||||||
The live-tool eval uses canonical seed documents so the real
|
```text
|
||||||
`LookupKnowledgeTool` can retrieve stable evidence from MySQL/Milvus instead of
|
caseId
|
||||||
whatever ad hoc documents happen to exist in the local knowledge base.
|
query
|
||||||
|
retrievedAt
|
||||||
|
searchMode # hybrid | dense (required on newly generated fixtures)
|
||||||
|
kbScope # e.g. rag-eval when generation used a scope
|
||||||
|
lookupResult # found, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, …
|
||||||
|
```
|
||||||
|
|
||||||
Seed documents live in:
|
Offline eval **ignores unknown top-level meta** and does **not** full-JSON-compare.
|
||||||
|
It asserts golden key fields only (doc/source, keywords, attempt, fallback, …). Raw scores are not pass criteria.
|
||||||
|
|
||||||
|
Older fixtures may omit `searchMode`; regenerate to attach meta.
|
||||||
|
|
||||||
|
## Seed docs + import
|
||||||
|
|
||||||
|
Live snapshot generation should use seed docs so results do not depend on ad-hoc local KB junk:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
eval/rag-retrieval/seed-docs/*.md
|
eval/rag-retrieval/seed-docs/*.md
|
||||||
```
|
```
|
||||||
|
|
||||||
Each seed doc uses frontmatter fields that are propagated into vector metadata:
|
Frontmatter example:
|
||||||
|
|
||||||
```yaml
|
```yaml
|
||||||
source: mysql-connection-pool
|
source: mysql-connection-pool
|
||||||
@@ -40,40 +67,29 @@ breadcrumb: Database > MySQL > Connection Pool
|
|||||||
kb_scope: rag-eval
|
kb_scope: rag-eval
|
||||||
```
|
```
|
||||||
|
|
||||||
Import or reindex the seed docs through the real upload pipeline:
|
Import via real upload pipeline:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
.\scripts\prepare_rag_eval_seed.ps1
|
.\scripts\prepare_rag_eval_seed.ps1
|
||||||
```
|
```
|
||||||
|
|
||||||
The script runs `RagEvalSeedImporterTest` with `rag.seed.enabled=true`. It
|
Isolation:
|
||||||
deletes the existing document with the same `source`/`docId`, uploads the seed
|
|
||||||
doc through `DocumentManagementService`, updates DB metadata and L0, and rebuilds
|
|
||||||
Milvus chunks.
|
|
||||||
|
|
||||||
`kb_scope` isolates eval data:
|
- App default may leave `retrieval.kb-scope` empty (all docs).
|
||||||
|
- Eval generation passes `-Dretrieval.kb-scope=rag-eval`.
|
||||||
|
- Category-filter fallback retries without L0 category filter only; **kb_scope still applies**.
|
||||||
|
|
||||||
- default application config leaves `retrieval.kb-scope` empty, so legacy docs
|
Body is chunked/embedded; frontmatter feeds metadata/L0 (decoy keywords in frontmatter alone should not become dense content).
|
||||||
without `kb_scope` remain searchable;
|
|
||||||
- eval scripts pass `-Dretrieval.kb-scope=rag-eval`, so L0 query hints and L1
|
|
||||||
vector retrieval both use only the canonical eval seed docs;
|
|
||||||
- the fallback retry skips only the L0 category filter, not the `kb_scope`
|
|
||||||
boundary.
|
|
||||||
|
|
||||||
Frontmatter is not embedded as chunk content during upload. It feeds metadata,
|
Seeds must live in the **current hybrid collection schema** (`milvus.collection`, default `biz`). If the collection was recreated for BM25 hybrid, re-import seeds after rebuild.
|
||||||
L0, and document enrichment; only the Markdown body is chunked and embedded.
|
|
||||||
This keeps controlled L0 decoys from becoming semantically relevant just because
|
|
||||||
their frontmatter keywords matched the query.
|
|
||||||
|
|
||||||
## Run
|
## Offline run (no live stack)
|
||||||
|
|
||||||
From the repository root:
|
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
python scripts/eval_rag_retrieval.py
|
python scripts/eval_rag_retrieval.py
|
||||||
```
|
```
|
||||||
|
|
||||||
Custom paths are also supported:
|
Custom paths:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
python scripts/eval_rag_retrieval.py \
|
python scripts/eval_rag_retrieval.py \
|
||||||
@@ -83,83 +99,53 @@ python scripts/eval_rag_retrieval.py \
|
|||||||
--markdown-report eval/rag-retrieval/reports/baseline.md
|
--markdown-report eval/rag-retrieval/reports/baseline.md
|
||||||
```
|
```
|
||||||
|
|
||||||
## Generate Fixtures From LookupKnowledgeTool
|
## Generate fixtures (live stack)
|
||||||
|
|
||||||
Use the snapshot generator when fixtures should reflect the real
|
|
||||||
`LookupKnowledgeTool` pipeline:
|
|
||||||
|
|
||||||
```powershell
|
|
||||||
.\scripts\generate_rag_lookup_snapshots.ps1
|
|
||||||
```
|
|
||||||
|
|
||||||
For the intended live loop, run seed import first:
|
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
.\scripts\prepare_rag_eval_seed.ps1
|
.\scripts\prepare_rag_eval_seed.ps1
|
||||||
.\scripts\generate_rag_lookup_snapshots.ps1
|
.\scripts\generate_rag_lookup_snapshots.ps1
|
||||||
python scripts\eval_rag_retrieval.py
|
# default: SearchMode=hybrid, KbScope=rag-eval, then offline eval
|
||||||
```
|
```
|
||||||
|
|
||||||
The script runs a Spring test harness:
|
Dense baseline snapshot (same seed, comparison only):
|
||||||
|
|
||||||
```text
|
|
||||||
mvn -q -Dtest=RagLookupSnapshotGeneratorTest -Drag.snapshot.enabled=true -Dretrieval.kb-scope=rag-eval -Dretrieval.vector-store.mode=spring test
|
|
||||||
```
|
|
||||||
|
|
||||||
The generator reads `golden-cases.json`, injects the real `LookupKnowledgeTool`
|
|
||||||
bean, calls `lookupKnowledge(query)` for each case, writes
|
|
||||||
`fixtures/{caseId}.json`, and then runs `eval_rag_retrieval.py` unless
|
|
||||||
`-SkipEval` is provided. It defaults to Spring AI VectorStore mode; pass
|
|
||||||
`-VectorStoreMode sdk` only when intentionally comparing the legacy SDK path.
|
|
||||||
|
|
||||||
Custom paths are supported:
|
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
.\scripts\generate_rag_lookup_snapshots.ps1 `
|
.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode dense -Fixtures eval\rag-retrieval\fixtures-dense -SkipEval
|
||||||
-Cases eval\rag-retrieval\cases\golden-cases.json `
|
|
||||||
-Fixtures eval\rag-retrieval\fixtures `
|
|
||||||
-RetrievedAt 2026-07-06T00:00:00Z
|
|
||||||
```
|
```
|
||||||
|
|
||||||
The generator is disabled in normal test runs. It only executes when
|
(Dual-directory comparison reports are optional / future; knife-1 only documents the override.)
|
||||||
`rag.snapshot.enabled=true` is provided because it writes repository files and
|
|
||||||
depends on the configured runtime retrieval stack.
|
|
||||||
|
|
||||||
If generated fixtures fail the offline baseline, treat that as a real alignment
|
Maven equivalent:
|
||||||
signal: either the golden expectations need to be adjusted to the current
|
|
||||||
knowledge base, or the knowledge base/indexing path needs to be fixed.
|
|
||||||
|
|
||||||
## Modular RAG Contract
|
|
||||||
|
|
||||||
Fixtures must use the current `lookupResult` shape, which mirrors the
|
|
||||||
`lookup_knowledge` output:
|
|
||||||
|
|
||||||
```text
|
```text
|
||||||
lookupResult.evidenceBlocks
|
mvn -q -Dtest=RagLookupSnapshotGeneratorTest \
|
||||||
lookupResult.contextPack
|
-Drag.snapshot.enabled=true \
|
||||||
lookupResult.retrievalTrace
|
-Dretrieval.kb-scope=rag-eval \
|
||||||
lookupResult.rerankTrace
|
-Dretrieval.search.mode=hybrid \
|
||||||
|
test
|
||||||
```
|
```
|
||||||
|
|
||||||
Golden cases can assert both retrieval quality and pipeline behavior:
|
Generator is **off** in normal tests; only runs when `rag.snapshot.enabled=true` (writes files).
|
||||||
|
|
||||||
|
If generated fixtures fail offline golden checks: either fix retrieval/index, or update golden/baseline **with an explicit reason** — do not silently overwrite.
|
||||||
|
|
||||||
|
## Golden assertions
|
||||||
|
|
||||||
|
Supported expectation fields include:
|
||||||
|
|
||||||
- `expectedSources` / `expectedDocIds`
|
- `expectedSources` / `expectedDocIds`
|
||||||
- `expectedBreadcrumbs`
|
- `expectedBreadcrumbs` / `expectedKeywords`
|
||||||
- `expectedKeywords`
|
|
||||||
- `expectedSelectedAttempt`
|
- `expectedSelectedAttempt`
|
||||||
- `expectedFallbackReason`
|
- `expectedFallbackReason` / `expectedFallbackReasons`
|
||||||
- `expectedFallbackReasons`
|
|
||||||
- `expectedEvidenceStatus`
|
- `expectedEvidenceStatus`
|
||||||
- `expectedContextSources`
|
- `expectedContextSources`
|
||||||
- `expectedRerankTopSource`
|
- `expectedRerankTopSource`
|
||||||
|
|
||||||
This lets the baseline catch regressions such as losing the expected evidence
|
Catch regressions such as missing expected source, broken context pack sources, wrong selected attempt, or broken filtered → unfiltered retry.
|
||||||
source, skipping context packing, changing the selected retrieval attempt, or
|
|
||||||
breaking the filtered-vector to unfiltered-retry fallback.
|
|
||||||
|
|
||||||
## Baseline Diff
|
**Note:** `relevance_level` is not a hard golden gate here (hybrid quality is rank-ordinal; see agent relevance-level doc).
|
||||||
|
|
||||||
To compare a freshly generated report against an existing baseline:
|
## Baseline diff
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
python scripts/eval_rag_retrieval.py \
|
python scripts/eval_rag_retrieval.py \
|
||||||
@@ -170,64 +156,24 @@ python scripts/eval_rag_retrieval.py \
|
|||||||
--diff-markdown-report eval/rag-retrieval/reports/baseline-diff.md
|
--diff-markdown-report eval/rag-retrieval/reports/baseline-diff.md
|
||||||
```
|
```
|
||||||
|
|
||||||
The diff reports aggregate regressions and case-level changes for:
|
Diff covers pass rate, recall@K, hit level, first expected rank, attempt, fallback, evidence status, rerank top source.
|
||||||
|
Non-zero exit on case failure or regression in diff mode.
|
||||||
|
|
||||||
- pass rate, recall@K, strong hit rate, miss count
|
## Hit levels
|
||||||
- pass state
|
|
||||||
- hit level
|
|
||||||
- first expected rank
|
|
||||||
- selected attempt
|
|
||||||
- fallback reason
|
|
||||||
- evidence status
|
|
||||||
- rerank top source
|
|
||||||
|
|
||||||
The command exits non-zero when a case fails or the diff contains a regression.
|
- `strong`: expected document found **and** breadcrumb or keyword coverage OK
|
||||||
|
- `medium`: expected document found, coverage incomplete
|
||||||
|
- `weak`: keyword hit without expected document
|
||||||
|
- `miss`: neither
|
||||||
|
|
||||||
## Hit Levels
|
`Recall@K` counts `strong` + `medium`.
|
||||||
|
|
||||||
- `strong`: expected document is found and breadcrumb or evidence keyword coverage is satisfied.
|
## Optional live smoke (post-reindex)
|
||||||
- `medium`: expected document is found, but breadcrumb or keyword coverage is incomplete.
|
|
||||||
- `weak`: expected evidence keyword is found, but expected document is missing.
|
|
||||||
- `miss`: expected document and expected evidence are not found.
|
|
||||||
|
|
||||||
`Recall@K` counts `strong` and `medium` as retrieved.
|
After reindex, with app up:
|
||||||
|
|
||||||
## Scope
|
|
||||||
|
|
||||||
This baseline runs fully offline and does not call MySQL, Redis, Milvus, an LLM,
|
|
||||||
or the Spring Boot application. It is a regression harness for retrieval behavior,
|
|
||||||
not a claim that live production retrieval accuracy is complete.
|
|
||||||
|
|
||||||
## Live Post-Reindex Acceptance
|
|
||||||
|
|
||||||
When embedding input changes, existing vectors do not update by themselves. For
|
|
||||||
example, after adding `title` and `breadcrumb` to the embedding text, the live
|
|
||||||
Milvus/Zilliz collection must be reindexed before retrieval can reflect that new
|
|
||||||
semantic signal.
|
|
||||||
|
|
||||||
Use this optional live acceptance flow after the application is running and the
|
|
||||||
knowledge base has been reindexed:
|
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
python scripts/eval_rag_live_acceptance.py
|
python scripts/eval_rag_live_acceptance.py --base-url http://127.0.0.1:9900
|
||||||
```
|
```
|
||||||
|
|
||||||
Custom service URL and output paths are supported:
|
Calls `GET /api/search/similar`. Environment smoke only — does **not** replace offline baseline.
|
||||||
|
|
||||||
```bash
|
|
||||||
python scripts/eval_rag_live_acceptance.py \
|
|
||||||
--base-url http://127.0.0.1:9900 \
|
|
||||||
--json-report eval/rag-retrieval/reports/live-post-reindex.json \
|
|
||||||
--markdown-report eval/rag-retrieval/reports/live-post-reindex.md
|
|
||||||
```
|
|
||||||
|
|
||||||
The script calls:
|
|
||||||
|
|
||||||
```text
|
|
||||||
GET /api/search/similar
|
|
||||||
```
|
|
||||||
|
|
||||||
It writes JSON and Markdown reports with query, topK, result count, top
|
|
||||||
results, breadcrumb, score labels, and raw response fields. This is a live
|
|
||||||
smoke check for environment readiness and post-reindex behavior; it does not
|
|
||||||
replace the deterministic offline baseline above.
|
|
||||||
|
|||||||
@@ -1,81 +1,122 @@
|
|||||||
{
|
{
|
||||||
"caseId": "aiops-payment-latency-alert",
|
"caseId" : "aiops-payment-latency-alert",
|
||||||
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
|
"query" : "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
"retrievedAt" : "2026-07-28T06:54:51.843450400Z",
|
||||||
"lookupResult": {
|
"searchMode" : "hybrid",
|
||||||
"found": true,
|
"kbScope" : "rag-eval",
|
||||||
"evidenceBlocks": [
|
"lookupResult" : {
|
||||||
{
|
"found" : true,
|
||||||
"source": "payment-service-latency",
|
"evidenceBlocks" : [ {
|
||||||
"title": "Payment Service Latency Alert Playbook",
|
"docId" : "payment-service-latency",
|
||||||
"breadcrumb": "AIOps > Service Alerts > Payment Latency",
|
"chunkIndex" : 2,
|
||||||
"retrievalLayer": "L1",
|
"evidenceKey" : "payment-service-latency#chunk-2",
|
||||||
"content": "For payment-service p95 latency alerts, check downstream dependency latency, thread pool saturation, gateway retries, and recent deployment changes.",
|
"source" : "payment-service-latency",
|
||||||
"score": 0.84,
|
"title" : "Payment Latency",
|
||||||
"hitReasons": ["domain_match:+0.15", "entity_match:+0.20", "keyword_match:+0.10"]
|
"breadcrumb" : "AIOps > Service Alerts > Payment Latency",
|
||||||
},
|
"retrievalLayer" : "L1",
|
||||||
{
|
"content" : "### Payment Latency\n\nFor `HighLatency` alerts on `payment-service`, treat p95 latency as the primary symptom.\n\nDiagnosis steps:\n\n1. Confirm whether p95 latency is isolated to payment-service or shared across upstream callers.\n2. Compare payment-service latency with downstream dependency latency for gateway, risk, and order services.\n3. Check connection pool wait time, retry spikes, and timeout rates.\n4. If downstream dependency latency increased first, classify payment-service as affected rather than root cause.\n\nThe expected evidence terms are p95 latency, payment-service, and downstream dependency.",
|
||||||
"source": "mysql-connection-pool",
|
"score" : 0.032786883413791656,
|
||||||
"title": "MySQL Connection Pool Troubleshooting",
|
"hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||||
"breadcrumb": "Database > MySQL > Connection Pool",
|
}, {
|
||||||
"retrievalLayer": "L1",
|
"docId" : "aiops-alert-scope-control",
|
||||||
"content": "Database connection pool saturation can increase payment latency when checkout paths wait for connections.",
|
"chunkIndex" : 1,
|
||||||
"score": 0.68,
|
"evidenceKey" : "aiops-alert-scope-control#chunk-1",
|
||||||
"hitReasons": ["keyword_match:+0.10"]
|
"source" : "aiops-alert-scope-control",
|
||||||
}
|
"title" : "Alert Scope Control",
|
||||||
],
|
"breadcrumb" : "AIOps > Alert Scope Control",
|
||||||
"contextPack": {
|
"retrievalLayer" : "L1",
|
||||||
"packedText": "[1] Payment Service Latency Alert Playbook\nAIOps > Service Alerts > Payment Latency\nFor payment-service p95 latency alerts, check downstream dependency latency, thread pool saturation, gateway retries, and recent deployment changes.",
|
"content" : "## Alert Scope Control\n\nWhen an AIOps request already includes an alert payload, the agent should diagnose that payload first.\nIt must not expand the task into unrelated active alerts unless the user asks for broad alert triage.\n\nScope rules:\n\n1. Treat the provided payload as the primary incident boundary.\n2. Use unrelated active alerts only as correlation evidence when they share service, dependency, time window, or trace context.\n3. Do not replace the requested alert with a louder but unrelated alert.\n\nThis runbook anchors payload, unrelated active alerts, and scope behavior.",
|
||||||
"strategy": "top_evidence_blocks",
|
"score" : 0.0320020467042923,
|
||||||
"charBudget": 3500,
|
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||||
"usedChars": 236,
|
}, {
|
||||||
"includedSources": ["payment-service-latency", "mysql-connection-pool"],
|
"docId" : "payment-service-latency",
|
||||||
"omittedSources": []
|
"chunkIndex" : 1,
|
||||||
|
"evidenceKey" : "payment-service-latency#chunk-1",
|
||||||
|
"source" : "payment-service-latency",
|
||||||
|
"title" : "Service Alerts",
|
||||||
|
"breadcrumb" : "AIOps > Service Alerts > Payment Latency",
|
||||||
|
"retrievalLayer" : "L1",
|
||||||
|
"content" : "## Service Alerts",
|
||||||
|
"score" : 0.0320020467042923,
|
||||||
|
"hitReasons" : [ "semantic_rank:3", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||||
|
}, {
|
||||||
|
"docId" : "aiops-alert-scope-control",
|
||||||
|
"chunkIndex" : 0,
|
||||||
|
"evidenceKey" : "aiops-alert-scope-control#chunk-0",
|
||||||
|
"source" : "aiops-alert-scope-control",
|
||||||
|
"title" : "AIOps",
|
||||||
|
"breadcrumb" : "AIOps > Alert Scope Control",
|
||||||
|
"retrievalLayer" : "L1",
|
||||||
|
"content" : "# AIOps",
|
||||||
|
"score" : 0.015384615398943424,
|
||||||
|
"hitReasons" : [ "semantic_rank:5", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||||
|
} ],
|
||||||
|
"contextPack" : {
|
||||||
|
"packedText" : "[Evidence 1]\nsource: payment-service-latency\ntitle: Payment Latency\nbreadcrumb: AIOps > Service Alerts > Payment Latency\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n### Payment Latency\n\nFor `HighLatency` alerts on `payment-service`, treat p95 latency as the primary symptom.\n\nDiagnosis steps:\n\n1. Confirm whether p95 latency is isolated to payment-service or shared across upstream callers.\n2. Compare payment-service latency with downstream dependency latency for gateway, risk, and order services.\n3. Check connection pool wait time, retry spikes, and timeout rates.\n4. If downstream dependency latency increased first, classify payment-service as affected rather than root cause.\n\nThe expected evidence terms are p95 latency, payment-service, and downstream dependency.\n\n[Evidence 2]\nsource: aiops-alert-scope-control\ntitle: Alert Scope Control\nbreadcrumb: AIOps > Alert Scope Control\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n## Alert Scope Control\n\nWhen an AIOps request already includes an alert payload, the agent should diagnose that payload first.\nIt must not expand the task into unrelated active alerts unless the user asks for broad alert triage.\n\nScope rules:\n\n1. Treat the provided payload as the primary incident boundary.\n2. Use unrelated active alerts only as correlation evidence when they share service, dependency, time window, or trace context.\n3. Do not replace the requested alert with a louder but unrelated alert.\n\nThis runbook anchors payload, unrelated active alerts, and scope behavior.\n\n[Evidence 3]\nsource: payment-service-latency\ntitle: Service Alerts\nbreadcrumb: AIOps > Service Alerts > Payment Latency\nlayer: L1\nreasons: semantic_rank:3, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n## Service Alerts\n\n[Evidence 4]\nsource: aiops-alert-scope-control\ntitle: AIOps\nbreadcrumb: AIOps > Alert Scope Control\nlayer: L1\nreasons: semantic_rank:5, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n# AIOps",
|
||||||
|
"strategy" : "ranked_evidence_char_budget",
|
||||||
|
"charBudget" : 4000,
|
||||||
|
"usedChars" : 2108,
|
||||||
|
"includedSources" : [ "payment-service-latency", "aiops-alert-scope-control", "payment-service-latency", "aiops-alert-scope-control" ],
|
||||||
|
"omittedSources" : [ ]
|
||||||
},
|
},
|
||||||
"retrievalTrace": {
|
"retrievalTrace" : {
|
||||||
"originalQuery": "Alert HighLatency on payment-service with p95 latency above threshold",
|
"originalQuery" : "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||||
"rewrittenQuery": "HighLatency payment-service p95 latency alert downstream dependency diagnosis",
|
"rewrittenQuery" : "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||||
"categoryFilter": "AIOps",
|
"categoryFilter" : "aiops",
|
||||||
"selectedAttempt": "FILTERED_VECTOR",
|
"selectedAttempt" : "FILTERED_VECTOR",
|
||||||
"fallbackReason": null,
|
"fallbackReason" : null,
|
||||||
"evidenceStatus": "supported",
|
"evidenceStatus" : "supported",
|
||||||
"queryHints": {
|
"queryHints" : {
|
||||||
"domains": ["AIOps"],
|
"domains" : [ "aiops" ],
|
||||||
"matched_keywords": ["p95 latency", "payment-service", "downstream dependency"],
|
"matched_keywords" : [ "HighLatency", "payment-service", "p95 latency" ],
|
||||||
"entities": ["payment-service", "HighLatency"],
|
"entities" : [ "HighLatency", "payment-service", "p95 latency" ],
|
||||||
"l0_titles": ["Payment Service Latency Alert Playbook"],
|
"l0_titles" : [ "Payment Service Latency Alert" ],
|
||||||
"l0_match_count": 1
|
"l0_match_count" : 1
|
||||||
},
|
},
|
||||||
"attempts": [
|
"attempts" : [ {
|
||||||
{
|
"name" : "FILTERED_VECTOR",
|
||||||
"name": "FILTERED_VECTOR",
|
"query" : "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||||
"query": "HighLatency payment-service p95 latency alert downstream dependency diagnosis",
|
"categoryFilter" : "aiops",
|
||||||
"categoryFilter": "AIOps",
|
"candidateCount" : 5,
|
||||||
"candidateCount": 2,
|
"usable" : true,
|
||||||
"usable": true,
|
"errorMessage" : null,
|
||||||
"durationMs": 11,
|
"durationMs" : 969,
|
||||||
"topScore": 0.84,
|
"topScore" : 0.032786883413791656,
|
||||||
"topSimilarity": 0.84
|
"topSimilarity" : 0.7736010700464249
|
||||||
}
|
} ]
|
||||||
]
|
|
||||||
},
|
},
|
||||||
"rerankTrace": {
|
"rerankTrace" : {
|
||||||
"items": [
|
"items" : [ {
|
||||||
{
|
"finalRank" : 1,
|
||||||
"finalRank": 1,
|
"source" : "payment-service-latency",
|
||||||
"source": "payment-service-latency",
|
"baseScore" : 0.7736010700464249,
|
||||||
"baseScore": 0.84,
|
"finalScore" : 0.7736010700464249,
|
||||||
"finalScore": 1.29,
|
"boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||||
"boostReasons": ["domain_match:+0.15", "entity_match:+0.20", "keyword_match:+0.10"]
|
}, {
|
||||||
},
|
"finalRank" : 2,
|
||||||
{
|
"source" : "aiops-alert-scope-control",
|
||||||
"finalRank": 2,
|
"baseScore" : 0.4683566689491272,
|
||||||
"source": "mysql-connection-pool",
|
"finalScore" : 0.4683566689491272,
|
||||||
"baseScore": 0.68,
|
"boostReasons" : [ "l0_domain_overlap" ]
|
||||||
"finalScore": 0.78,
|
}, {
|
||||||
"boostReasons": ["keyword_match:+0.10"]
|
"finalRank" : 3,
|
||||||
}
|
"source" : "payment-service-latency",
|
||||||
]
|
"baseScore" : 0.4981400966644287,
|
||||||
}
|
"finalScore" : 0.4981400966644287,
|
||||||
|
"boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||||
|
}, {
|
||||||
|
"finalRank" : 4,
|
||||||
|
"source" : "aiops-alert-scope-control",
|
||||||
|
"baseScore" : 0.3814886808395386,
|
||||||
|
"finalScore" : 0.3814886808395386,
|
||||||
|
"boostReasons" : [ "l0_domain_overlap" ]
|
||||||
|
} ]
|
||||||
|
},
|
||||||
|
"evidenceCandidateCount" : 5,
|
||||||
|
"evidenceBlockCount" : 4,
|
||||||
|
"relevanceLevel" : "PRECISE",
|
||||||
|
"completenessHint" : "知识库中不存在比上述结果更精准的文档",
|
||||||
|
"retrievedDomainsThisSession" : null,
|
||||||
|
"message" : null
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -1,65 +1,122 @@
|
|||||||
{
|
{
|
||||||
"caseId": "aiops-prometheus-alert-scope",
|
"caseId" : "aiops-prometheus-alert-scope",
|
||||||
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
"query" : "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
"retrievedAt" : "2026-07-28T06:54:51.843450400Z",
|
||||||
"lookupResult": {
|
"searchMode" : "hybrid",
|
||||||
"found": true,
|
"kbScope" : "rag-eval",
|
||||||
"evidenceBlocks": [
|
"lookupResult" : {
|
||||||
{
|
"found" : true,
|
||||||
"source": "aiops-alert-scope-control",
|
"evidenceBlocks" : [ {
|
||||||
"title": "AIOps Alert Scope Control",
|
"docId" : "aiops-alert-scope-control",
|
||||||
"breadcrumb": "AIOps > Alert Scope Control",
|
"chunkIndex" : 1,
|
||||||
"retrievalLayer": "L1",
|
"evidenceKey" : "aiops-alert-scope-control#chunk-1",
|
||||||
"content": "When payload mode is used, diagnose the input alert payload and do not expand unrelated active alerts into the main diagnosis scope.",
|
"source" : "aiops-alert-scope-control",
|
||||||
"score": 0.88,
|
"title" : "Alert Scope Control",
|
||||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
"breadcrumb" : "AIOps > Alert Scope Control",
|
||||||
}
|
"retrievalLayer" : "L1",
|
||||||
],
|
"content" : "## Alert Scope Control\n\nWhen an AIOps request already includes an alert payload, the agent should diagnose that payload first.\nIt must not expand the task into unrelated active alerts unless the user asks for broad alert triage.\n\nScope rules:\n\n1. Treat the provided payload as the primary incident boundary.\n2. Use unrelated active alerts only as correlation evidence when they share service, dependency, time window, or trace context.\n3. Do not replace the requested alert with a louder but unrelated alert.\n\nThis runbook anchors payload, unrelated active alerts, and scope behavior.",
|
||||||
"contextPack": {
|
"score" : 0.032786883413791656,
|
||||||
"packedText": "[1] AIOps Alert Scope Control\nAIOps > Alert Scope Control\nWhen payload mode is used, diagnose the input alert payload and do not expand unrelated active alerts into the main diagnosis scope.",
|
"hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||||
"strategy": "top_evidence_blocks",
|
}, {
|
||||||
"charBudget": 3500,
|
"docId" : "payment-service-latency",
|
||||||
"usedChars": 188,
|
"chunkIndex" : 1,
|
||||||
"includedSources": ["aiops-alert-scope-control"],
|
"evidenceKey" : "payment-service-latency#chunk-1",
|
||||||
"omittedSources": []
|
"source" : "payment-service-latency",
|
||||||
|
"title" : "Service Alerts",
|
||||||
|
"breadcrumb" : "AIOps > Service Alerts > Payment Latency",
|
||||||
|
"retrievalLayer" : "L1",
|
||||||
|
"content" : "## Service Alerts",
|
||||||
|
"score" : 0.0320020467042923,
|
||||||
|
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||||
|
}, {
|
||||||
|
"docId" : "payment-service-latency",
|
||||||
|
"chunkIndex" : 2,
|
||||||
|
"evidenceKey" : "payment-service-latency#chunk-2",
|
||||||
|
"source" : "payment-service-latency",
|
||||||
|
"title" : "Payment Latency",
|
||||||
|
"breadcrumb" : "AIOps > Service Alerts > Payment Latency",
|
||||||
|
"retrievalLayer" : "L1",
|
||||||
|
"content" : "### Payment Latency\n\nFor `HighLatency` alerts on `payment-service`, treat p95 latency as the primary symptom.\n\nDiagnosis steps:\n\n1. Confirm whether p95 latency is isolated to payment-service or shared across upstream callers.\n2. Compare payment-service latency with downstream dependency latency for gateway, risk, and order services.\n3. Check connection pool wait time, retry spikes, and timeout rates.\n4. If downstream dependency latency increased first, classify payment-service as affected rather than root cause.\n\nThe expected evidence terms are p95 latency, payment-service, and downstream dependency.",
|
||||||
|
"score" : 0.0320020467042923,
|
||||||
|
"hitReasons" : [ "semantic_rank:3", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||||
|
}, {
|
||||||
|
"docId" : "aiops-alert-scope-control",
|
||||||
|
"chunkIndex" : 0,
|
||||||
|
"evidenceKey" : "aiops-alert-scope-control#chunk-0",
|
||||||
|
"source" : "aiops-alert-scope-control",
|
||||||
|
"title" : "AIOps",
|
||||||
|
"breadcrumb" : "AIOps > Alert Scope Control",
|
||||||
|
"retrievalLayer" : "L1",
|
||||||
|
"content" : "# AIOps",
|
||||||
|
"score" : 0.03076923079788685,
|
||||||
|
"hitReasons" : [ "semantic_rank:5", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||||
|
} ],
|
||||||
|
"contextPack" : {
|
||||||
|
"packedText" : "[Evidence 1]\nsource: aiops-alert-scope-control\ntitle: Alert Scope Control\nbreadcrumb: AIOps > Alert Scope Control\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n## Alert Scope Control\n\nWhen an AIOps request already includes an alert payload, the agent should diagnose that payload first.\nIt must not expand the task into unrelated active alerts unless the user asks for broad alert triage.\n\nScope rules:\n\n1. Treat the provided payload as the primary incident boundary.\n2. Use unrelated active alerts only as correlation evidence when they share service, dependency, time window, or trace context.\n3. Do not replace the requested alert with a louder but unrelated alert.\n\nThis runbook anchors payload, unrelated active alerts, and scope behavior.\n\n[Evidence 2]\nsource: payment-service-latency\ntitle: Service Alerts\nbreadcrumb: AIOps > Service Alerts > Payment Latency\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n## Service Alerts\n\n[Evidence 3]\nsource: payment-service-latency\ntitle: Payment Latency\nbreadcrumb: AIOps > Service Alerts > Payment Latency\nlayer: L1\nreasons: semantic_rank:3, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n### Payment Latency\n\nFor `HighLatency` alerts on `payment-service`, treat p95 latency as the primary symptom.\n\nDiagnosis steps:\n\n1. Confirm whether p95 latency is isolated to payment-service or shared across upstream callers.\n2. Compare payment-service latency with downstream dependency latency for gateway, risk, and order services.\n3. Check connection pool wait time, retry spikes, and timeout rates.\n4. If downstream dependency latency increased first, classify payment-service as affected rather than root cause.\n\nThe expected evidence terms are p95 latency, payment-service, and downstream dependency.\n\n[Evidence 4]\nsource: aiops-alert-scope-control\ntitle: AIOps\nbreadcrumb: AIOps > Alert Scope Control\nlayer: L1\nreasons: semantic_rank:5, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n# AIOps",
|
||||||
|
"strategy" : "ranked_evidence_char_budget",
|
||||||
|
"charBudget" : 4000,
|
||||||
|
"usedChars" : 2069,
|
||||||
|
"includedSources" : [ "aiops-alert-scope-control", "payment-service-latency", "payment-service-latency", "aiops-alert-scope-control" ],
|
||||||
|
"omittedSources" : [ ]
|
||||||
},
|
},
|
||||||
"retrievalTrace": {
|
"retrievalTrace" : {
|
||||||
"originalQuery": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
"originalQuery" : "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||||
"rewrittenQuery": "AIOps alert payload scope unrelated active alerts diagnosis",
|
"rewrittenQuery" : "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||||
"categoryFilter": "AIOps",
|
"categoryFilter" : "aiops",
|
||||||
"selectedAttempt": "FILTERED_VECTOR",
|
"selectedAttempt" : "FILTERED_VECTOR",
|
||||||
"fallbackReason": null,
|
"fallbackReason" : null,
|
||||||
"evidenceStatus": "supported",
|
"evidenceStatus" : "supported",
|
||||||
"queryHints": {
|
"queryHints" : {
|
||||||
"domains": ["AIOps"],
|
"domains" : [ "aiops" ],
|
||||||
"matched_keywords": ["payload", "unrelated active alerts", "scope"],
|
"matched_keywords" : [ "alert payload", "unrelated active alerts" ],
|
||||||
"entities": ["alert payload"],
|
"entities" : [ "alert payload", "unrelated active alerts" ],
|
||||||
"l0_titles": ["AIOps Alert Scope Control"],
|
"l0_titles" : [ "AIOps Alert Scope Control" ],
|
||||||
"l0_match_count": 1
|
"l0_match_count" : 1
|
||||||
},
|
},
|
||||||
"attempts": [
|
"attempts" : [ {
|
||||||
{
|
"name" : "FILTERED_VECTOR",
|
||||||
"name": "FILTERED_VECTOR",
|
"query" : "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||||
"query": "AIOps alert payload scope unrelated active alerts diagnosis",
|
"categoryFilter" : "aiops",
|
||||||
"categoryFilter": "AIOps",
|
"candidateCount" : 5,
|
||||||
"candidateCount": 1,
|
"usable" : true,
|
||||||
"usable": true,
|
"errorMessage" : null,
|
||||||
"durationMs": 8,
|
"durationMs" : 1623,
|
||||||
"topScore": 0.88,
|
"topScore" : 0.032786883413791656,
|
||||||
"topSimilarity": 0.88
|
"topSimilarity" : 0.7561411112546921
|
||||||
}
|
} ]
|
||||||
]
|
|
||||||
},
|
},
|
||||||
"rerankTrace": {
|
"rerankTrace" : {
|
||||||
"items": [
|
"items" : [ {
|
||||||
{
|
"finalRank" : 1,
|
||||||
"finalRank": 1,
|
"source" : "aiops-alert-scope-control",
|
||||||
"source": "aiops-alert-scope-control",
|
"baseScore" : 0.7561411112546921,
|
||||||
"baseScore": 0.88,
|
"finalScore" : 0.7561411112546921,
|
||||||
"finalScore": 1.13,
|
"boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
}, {
|
||||||
}
|
"finalRank" : 2,
|
||||||
]
|
"source" : "payment-service-latency",
|
||||||
}
|
"baseScore" : 0.503810703754425,
|
||||||
|
"finalScore" : 0.503810703754425,
|
||||||
|
"boostReasons" : [ "l0_domain_overlap" ]
|
||||||
|
}, {
|
||||||
|
"finalRank" : 3,
|
||||||
|
"source" : "payment-service-latency",
|
||||||
|
"baseScore" : 0.5772626996040344,
|
||||||
|
"finalScore" : 0.5772626996040344,
|
||||||
|
"boostReasons" : [ "l0_domain_overlap" ]
|
||||||
|
}, {
|
||||||
|
"finalRank" : 4,
|
||||||
|
"source" : "aiops-alert-scope-control",
|
||||||
|
"baseScore" : 0.36770421266555786,
|
||||||
|
"finalScore" : 0.36770421266555786,
|
||||||
|
"boostReasons" : [ "l0_domain_overlap" ]
|
||||||
|
} ]
|
||||||
|
},
|
||||||
|
"evidenceCandidateCount" : 5,
|
||||||
|
"evidenceBlockCount" : 4,
|
||||||
|
"relevanceLevel" : "PRECISE",
|
||||||
|
"completenessHint" : "知识库中不存在比上述结果更精准的文档",
|
||||||
|
"retrievedDomainsThisSession" : null,
|
||||||
|
"message" : null
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -1,81 +1,88 @@
|
|||||||
{
|
{
|
||||||
"caseId": "chat-diagnosis-flow",
|
"caseId" : "chat-diagnosis-flow",
|
||||||
"query": "What is the standard troubleshooting flow for an application incident?",
|
"query" : "What is the standard troubleshooting flow for an application incident?",
|
||||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
"retrievedAt" : "2026-07-28T06:54:51.843450400Z",
|
||||||
"lookupResult": {
|
"searchMode" : "hybrid",
|
||||||
"found": true,
|
"kbScope" : "rag-eval",
|
||||||
"evidenceBlocks": [
|
"lookupResult" : {
|
||||||
{
|
"found" : true,
|
||||||
"source": "incident-diagnosis-flow",
|
"evidenceBlocks" : [ {
|
||||||
"title": "Incident Diagnosis Flow",
|
"docId" : "incident-diagnosis-flow",
|
||||||
"breadcrumb": "AIOps > Diagnosis Flow",
|
"chunkIndex" : 1,
|
||||||
"retrievalLayer": "L1",
|
"evidenceKey" : "incident-diagnosis-flow#chunk-1",
|
||||||
"content": "The standard flow is to collect evidence, identify the suspected fault domain, verify the hypothesis, apply remediation, and confirm recovery.",
|
"source" : "incident-diagnosis-flow",
|
||||||
"score": 0.82,
|
"title" : "Diagnosis Flow",
|
||||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
"breadcrumb" : "AIOps > Diagnosis Flow",
|
||||||
},
|
"retrievalLayer" : "L1",
|
||||||
{
|
"content" : "## Diagnosis Flow\n\nThe standard troubleshooting flow is evidence first, hypothesis second, remediation last.\n\nRecommended sequence:\n\n1. Collect evidence from alerts, metrics, logs, traces, deployments, and recent configuration changes.\n2. Define a small hypothesis that explains the observed symptoms.\n3. Verify the hypothesis with a targeted metric, log query, or reproduction step.\n4. Choose remediation that directly addresses the verified cause.\n5. Record the outcome and the evidence used to make the decision.\n\nDo not skip collect evidence, verify, and remediation ordering during an application incident.",
|
||||||
"source": "rag-chunk-context-reconstruction",
|
"score" : 0.032786883413791656,
|
||||||
"title": "RAG Chunk Context Reconstruction",
|
"hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||||
"breadcrumb": "RAG > Chunking > Context Reconstruction",
|
}, {
|
||||||
"retrievalLayer": "L1",
|
"docId" : "incident-diagnosis-flow",
|
||||||
"content": "Long sections may require neighbor chunk expansion and breadcrumb-aware packing.",
|
"chunkIndex" : 0,
|
||||||
"score": 0.55,
|
"evidenceKey" : "incident-diagnosis-flow#chunk-0",
|
||||||
"hitReasons": []
|
"source" : "incident-diagnosis-flow",
|
||||||
}
|
"title" : "AIOps",
|
||||||
],
|
"breadcrumb" : "AIOps > Diagnosis Flow",
|
||||||
"contextPack": {
|
"retrievalLayer" : "L1",
|
||||||
"packedText": "[1] Incident Diagnosis Flow\nAIOps > Diagnosis Flow\nThe standard flow is to collect evidence, identify the suspected fault domain, verify the hypothesis, apply remediation, and confirm recovery.",
|
"content" : "# AIOps",
|
||||||
"strategy": "top_evidence_blocks",
|
"score" : 0.016129031777381897,
|
||||||
"charBudget": 3500,
|
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||||
"usedChars": 192,
|
} ],
|
||||||
"includedSources": ["incident-diagnosis-flow", "rag-chunk-context-reconstruction"],
|
"contextPack" : {
|
||||||
"omittedSources": []
|
"packedText" : "[Evidence 1]\nsource: incident-diagnosis-flow\ntitle: Diagnosis Flow\nbreadcrumb: AIOps > Diagnosis Flow\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n## Diagnosis Flow\n\nThe standard troubleshooting flow is evidence first, hypothesis second, remediation last.\n\nRecommended sequence:\n\n1. Collect evidence from alerts, metrics, logs, traces, deployments, and recent configuration changes.\n2. Define a small hypothesis that explains the observed symptoms.\n3. Verify the hypothesis with a targeted metric, log query, or reproduction step.\n4. Choose remediation that directly addresses the verified cause.\n5. Record the outcome and the evidence used to make the decision.\n\nDo not skip collect evidence, verify, and remediation ordering during an application incident.\n\n[Evidence 2]\nsource: incident-diagnosis-flow\ntitle: AIOps\nbreadcrumb: AIOps > Diagnosis Flow\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n# AIOps",
|
||||||
|
"strategy" : "ranked_evidence_char_budget",
|
||||||
|
"charBudget" : 4000,
|
||||||
|
"usedChars" : 1032,
|
||||||
|
"includedSources" : [ "incident-diagnosis-flow", "incident-diagnosis-flow" ],
|
||||||
|
"omittedSources" : [ ]
|
||||||
},
|
},
|
||||||
"retrievalTrace": {
|
"retrievalTrace" : {
|
||||||
"originalQuery": "What is the standard troubleshooting flow for an application incident?",
|
"originalQuery" : "What is the standard troubleshooting flow for an application incident?",
|
||||||
"rewrittenQuery": "standard application incident troubleshooting flow collect evidence verify remediation",
|
"rewrittenQuery" : "What is the standard troubleshooting flow for an application incident?",
|
||||||
"categoryFilter": "AIOps",
|
"categoryFilter" : "ops",
|
||||||
"selectedAttempt": "FILTERED_VECTOR",
|
"selectedAttempt" : "FILTERED_VECTOR",
|
||||||
"fallbackReason": null,
|
"fallbackReason" : null,
|
||||||
"evidenceStatus": "supported",
|
"evidenceStatus" : "supported",
|
||||||
"queryHints": {
|
"queryHints" : {
|
||||||
"domains": ["AIOps"],
|
"domains" : [ "ops" ],
|
||||||
"matched_keywords": ["collect evidence", "verify", "remediation"],
|
"matched_keywords" : [ "standard troubleshooting flow", "application incident" ],
|
||||||
"entities": ["application incident"],
|
"entities" : [ "standard troubleshooting flow", "application incident" ],
|
||||||
"l0_titles": ["Incident Diagnosis Flow"],
|
"l0_titles" : [ "Incident Diagnosis Flow" ],
|
||||||
"l0_match_count": 1
|
"l0_match_count" : 1
|
||||||
},
|
},
|
||||||
"attempts": [
|
"attempts" : [ {
|
||||||
{
|
"name" : "FILTERED_VECTOR",
|
||||||
"name": "FILTERED_VECTOR",
|
"query" : "What is the standard troubleshooting flow for an application incident?",
|
||||||
"query": "standard application incident troubleshooting flow collect evidence verify remediation",
|
"categoryFilter" : "ops",
|
||||||
"categoryFilter": "AIOps",
|
"candidateCount" : 2,
|
||||||
"candidateCount": 2,
|
"usable" : true,
|
||||||
"usable": true,
|
"errorMessage" : null,
|
||||||
"durationMs": 10,
|
"durationMs" : 1540,
|
||||||
"topScore": 0.82,
|
"topScore" : 0.032786883413791656,
|
||||||
"topSimilarity": 0.82
|
"topSimilarity" : 0.6828859150409698
|
||||||
}
|
} ]
|
||||||
]
|
|
||||||
},
|
},
|
||||||
"rerankTrace": {
|
"rerankTrace" : {
|
||||||
"items": [
|
"items" : [ {
|
||||||
{
|
"finalRank" : 1,
|
||||||
"finalRank": 1,
|
"source" : "incident-diagnosis-flow",
|
||||||
"source": "incident-diagnosis-flow",
|
"baseScore" : 0.6828859150409698,
|
||||||
"baseScore": 0.82,
|
"finalScore" : 0.6828859150409698,
|
||||||
"finalScore": 1.07,
|
"boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
}, {
|
||||||
},
|
"finalRank" : 2,
|
||||||
{
|
"source" : "incident-diagnosis-flow",
|
||||||
"finalRank": 2,
|
"baseScore" : 0.3166210651397705,
|
||||||
"source": "rag-chunk-context-reconstruction",
|
"finalScore" : 0.3166210651397705,
|
||||||
"baseScore": 0.55,
|
"boostReasons" : [ "l0_domain_overlap" ]
|
||||||
"finalScore": 0.55,
|
} ]
|
||||||
"boostReasons": []
|
},
|
||||||
}
|
"evidenceCandidateCount" : 2,
|
||||||
]
|
"evidenceBlockCount" : 2,
|
||||||
}
|
"relevanceLevel" : "REFERENCE",
|
||||||
|
"completenessHint" : "当前结果为相关参考,如需更精准信息请明确缺少的具体维度",
|
||||||
|
"retrievedDomainsThisSession" : null,
|
||||||
|
"message" : null
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -1,81 +1,122 @@
|
|||||||
{
|
{
|
||||||
"caseId": "chat-l0-domain-hint",
|
"caseId" : "chat-l0-domain-hint",
|
||||||
"query": "Should L0 keyword matching decide the final retrieval result?",
|
"query" : "Should L0 keyword matching decide the final retrieval result?",
|
||||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
"retrievedAt" : "2026-07-28T06:54:51.843450400Z",
|
||||||
"lookupResult": {
|
"searchMode" : "hybrid",
|
||||||
"found": true,
|
"kbScope" : "rag-eval",
|
||||||
"evidenceBlocks": [
|
"lookupResult" : {
|
||||||
{
|
"found" : true,
|
||||||
"source": "rag-l0-domain-entity-hint",
|
"evidenceBlocks" : [ {
|
||||||
"title": "RAG L0 Domain Entity Hint",
|
"docId" : "rag-l0-domain-entity-hint",
|
||||||
"breadcrumb": "RAG > L0 > Domain Entity Hint",
|
"chunkIndex" : 2,
|
||||||
"retrievalLayer": "L1",
|
"evidenceKey" : "rag-l0-domain-entity-hint#chunk-2",
|
||||||
"content": "L0 should be retained as a domain detector, entity extractor, metadata filter generator, and explainability signal, not as the final retrieval decision.",
|
"source" : "rag-l0-domain-entity-hint",
|
||||||
"score": 0.88,
|
"title" : "Domain Entity Hint",
|
||||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
"breadcrumb" : "RAG > L0 > Domain Entity Hint",
|
||||||
},
|
"retrievalLayer" : "L1",
|
||||||
{
|
"content" : "### Domain Entity Hint\n\nL0 keyword matching should not decide the final retrieval result.\nIn the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.\n\nThe output can provide:\n\n1. Candidate domain hints.\n2. Matched entities and keywords.\n3. An optional metadata filter for the first vector retrieval attempt.\n\nFinal evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.\nThe important terms are domain detector, entity extractor, and metadata filter.",
|
||||||
"source": "rag-l0-l1-fusion-ranking",
|
"score" : 0.032786883413791656,
|
||||||
"title": "RAG L0 L1 Fusion Ranking",
|
"hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||||
"breadcrumb": "RAG > Ranking > Fusion",
|
}, {
|
||||||
"retrievalLayer": "L1",
|
"docId" : "rag-chunk-context-reconstruction",
|
||||||
"content": "L0 and L1 candidates should eventually be fused rather than handled as an early-return branch.",
|
"chunkIndex" : 2,
|
||||||
"score": 0.75,
|
"evidenceKey" : "rag-chunk-context-reconstruction#chunk-2",
|
||||||
"hitReasons": ["domain_match:+0.15"]
|
"source" : "rag-chunk-context-reconstruction",
|
||||||
}
|
"title" : "Context Reconstruction",
|
||||||
],
|
"breadcrumb" : "RAG > Chunking > Context Reconstruction",
|
||||||
"contextPack": {
|
"retrievalLayer" : "L1",
|
||||||
"packedText": "[1] RAG L0 Domain Entity Hint\nRAG > L0 > Domain Entity Hint\nL0 should be retained as a domain detector, entity extractor, metadata filter generator, and explainability signal, not as the final retrieval decision.",
|
"content" : "### Context Reconstruction\n\nWhen a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.\n\nRecommended behavior:\n\n1. Store the breadcrumb with every chunk.\n2. Preserve the same section identity across adjacent chunks.\n3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.\n4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.\n\nThe key concepts are neighbor chunk, same section, and breadcrumb.",
|
||||||
"strategy": "top_evidence_blocks",
|
"score" : 0.032258063554763794,
|
||||||
"charBudget": 3500,
|
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||||
"usedChars": 219,
|
}, {
|
||||||
"includedSources": ["rag-l0-domain-entity-hint", "rag-l0-l1-fusion-ranking"],
|
"docId" : "rag-l0-domain-entity-hint",
|
||||||
"omittedSources": []
|
"chunkIndex" : 1,
|
||||||
|
"evidenceKey" : "rag-l0-domain-entity-hint#chunk-1",
|
||||||
|
"source" : "rag-l0-domain-entity-hint",
|
||||||
|
"title" : "L0",
|
||||||
|
"breadcrumb" : "RAG > L0 > Domain Entity Hint",
|
||||||
|
"retrievalLayer" : "L1",
|
||||||
|
"content" : "## L0",
|
||||||
|
"score" : 0.0317460335791111,
|
||||||
|
"hitReasons" : [ "semantic_rank:3", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||||
|
}, {
|
||||||
|
"docId" : "rag-chunk-context-reconstruction",
|
||||||
|
"chunkIndex" : 1,
|
||||||
|
"evidenceKey" : "rag-chunk-context-reconstruction#chunk-1",
|
||||||
|
"source" : "rag-chunk-context-reconstruction",
|
||||||
|
"title" : "Chunking",
|
||||||
|
"breadcrumb" : "RAG > Chunking > Context Reconstruction",
|
||||||
|
"retrievalLayer" : "L1",
|
||||||
|
"content" : "## Chunking",
|
||||||
|
"score" : 0.015625,
|
||||||
|
"hitReasons" : [ "semantic_rank:4", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||||
|
} ],
|
||||||
|
"contextPack" : {
|
||||||
|
"packedText" : "[Evidence 1]\nsource: rag-l0-domain-entity-hint\ntitle: Domain Entity Hint\nbreadcrumb: RAG > L0 > Domain Entity Hint\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n### Domain Entity Hint\n\nL0 keyword matching should not decide the final retrieval result.\nIn the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.\n\nThe output can provide:\n\n1. Candidate domain hints.\n2. Matched entities and keywords.\n3. An optional metadata filter for the first vector retrieval attempt.\n\nFinal evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.\nThe important terms are domain detector, entity extractor, and metadata filter.\n\n[Evidence 2]\nsource: rag-chunk-context-reconstruction\ntitle: Context Reconstruction\nbreadcrumb: RAG > Chunking > Context Reconstruction\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n### Context Reconstruction\n\nWhen a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.\n\nRecommended behavior:\n\n1. Store the breadcrumb with every chunk.\n2. Preserve the same section identity across adjacent chunks.\n3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.\n4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.\n\nThe key concepts are neighbor chunk, same section, and breadcrumb.\n\n[Evidence 3]\nsource: rag-l0-domain-entity-hint\ntitle: L0\nbreadcrumb: RAG > L0 > Domain Entity Hint\nlayer: L1\nreasons: semantic_rank:3, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n## L0\n\n[Evidence 4]\nsource: rag-chunk-context-reconstruction\ntitle: Chunking\nbreadcrumb: RAG > Chunking > Context Reconstruction\nlayer: L1\nreasons: semantic_rank:4, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n## Chunking",
|
||||||
|
"strategy" : "ranked_evidence_char_budget",
|
||||||
|
"charBudget" : 4000,
|
||||||
|
"usedChars" : 1965,
|
||||||
|
"includedSources" : [ "rag-l0-domain-entity-hint", "rag-chunk-context-reconstruction", "rag-l0-domain-entity-hint", "rag-chunk-context-reconstruction" ],
|
||||||
|
"omittedSources" : [ ]
|
||||||
},
|
},
|
||||||
"retrievalTrace": {
|
"retrievalTrace" : {
|
||||||
"originalQuery": "Should L0 keyword matching decide the final retrieval result?",
|
"originalQuery" : "Should L0 keyword matching decide the final retrieval result?",
|
||||||
"rewrittenQuery": "RAG L0 keyword matching domain entity hint final retrieval decision",
|
"rewrittenQuery" : "Should L0 keyword matching decide the final retrieval result?",
|
||||||
"categoryFilter": "RAG",
|
"categoryFilter" : "rag",
|
||||||
"selectedAttempt": "FILTERED_VECTOR",
|
"selectedAttempt" : "FILTERED_VECTOR",
|
||||||
"fallbackReason": null,
|
"fallbackReason" : null,
|
||||||
"evidenceStatus": "supported",
|
"evidenceStatus" : "supported",
|
||||||
"queryHints": {
|
"queryHints" : {
|
||||||
"domains": ["RAG"],
|
"domains" : [ "rag" ],
|
||||||
"matched_keywords": ["domain detector", "entity extractor", "metadata filter"],
|
"matched_keywords" : [ "L0 keyword matching", "final retrieval result" ],
|
||||||
"entities": ["L0"],
|
"entities" : [ "L0 keyword matching", "final retrieval result" ],
|
||||||
"l0_titles": ["RAG L0 Domain Entity Hint"],
|
"l0_titles" : [ "RAG L0 Domain Entity Hint" ],
|
||||||
"l0_match_count": 1
|
"l0_match_count" : 1
|
||||||
},
|
},
|
||||||
"attempts": [
|
"attempts" : [ {
|
||||||
{
|
"name" : "FILTERED_VECTOR",
|
||||||
"name": "FILTERED_VECTOR",
|
"query" : "Should L0 keyword matching decide the final retrieval result?",
|
||||||
"query": "RAG L0 keyword matching domain entity hint final retrieval decision",
|
"categoryFilter" : "rag",
|
||||||
"categoryFilter": "RAG",
|
"candidateCount" : 6,
|
||||||
"candidateCount": 2,
|
"usable" : true,
|
||||||
"usable": true,
|
"errorMessage" : null,
|
||||||
"durationMs": 9,
|
"durationMs" : 850,
|
||||||
"topScore": 0.88,
|
"topScore" : 0.032786883413791656,
|
||||||
"topSimilarity": 0.88
|
"topSimilarity" : 0.6438122987747192
|
||||||
}
|
} ]
|
||||||
]
|
|
||||||
},
|
},
|
||||||
"rerankTrace": {
|
"rerankTrace" : {
|
||||||
"items": [
|
"items" : [ {
|
||||||
{
|
"finalRank" : 1,
|
||||||
"finalRank": 1,
|
"source" : "rag-l0-domain-entity-hint",
|
||||||
"source": "rag-l0-domain-entity-hint",
|
"baseScore" : 0.6438122987747192,
|
||||||
"baseScore": 0.88,
|
"finalScore" : 0.6438122987747192,
|
||||||
"finalScore": 1.13,
|
"boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
}, {
|
||||||
},
|
"finalRank" : 2,
|
||||||
{
|
"source" : "rag-chunk-context-reconstruction",
|
||||||
"finalRank": 2,
|
"baseScore" : 0.41426247358322144,
|
||||||
"source": "rag-l0-l1-fusion-ranking",
|
"finalScore" : 0.41426247358322144,
|
||||||
"baseScore": 0.75,
|
"boostReasons" : [ "l0_domain_overlap" ]
|
||||||
"finalScore": 0.9,
|
}, {
|
||||||
"boostReasons": ["domain_match:+0.15"]
|
"finalRank" : 3,
|
||||||
}
|
"source" : "rag-l0-domain-entity-hint",
|
||||||
]
|
"baseScore" : 0.3964804410934448,
|
||||||
}
|
"finalScore" : 0.3964804410934448,
|
||||||
|
"boostReasons" : [ "l0_domain_overlap" ]
|
||||||
|
}, {
|
||||||
|
"finalRank" : 4,
|
||||||
|
"source" : "rag-chunk-context-reconstruction",
|
||||||
|
"baseScore" : 0.28678786754608154,
|
||||||
|
"finalScore" : 0.28678786754608154,
|
||||||
|
"boostReasons" : [ "l0_domain_overlap" ]
|
||||||
|
} ]
|
||||||
|
},
|
||||||
|
"evidenceCandidateCount" : 6,
|
||||||
|
"evidenceBlockCount" : 4,
|
||||||
|
"relevanceLevel" : "REFERENCE",
|
||||||
|
"completenessHint" : "当前结果为相关参考,如需更精准信息请明确缺少的具体维度",
|
||||||
|
"retrievedDomainsThisSession" : null,
|
||||||
|
"message" : null
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -1,91 +1,149 @@
|
|||||||
{
|
{
|
||||||
"caseId": "chat-l0-filter-fallback",
|
"caseId" : "chat-l0-filter-fallback",
|
||||||
"query": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
"query" : "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
"retrievedAt" : "2026-07-28T06:54:51.843450400Z",
|
||||||
"lookupResult": {
|
"searchMode" : "hybrid",
|
||||||
"found": true,
|
"kbScope" : "rag-eval",
|
||||||
"evidenceBlocks": [
|
"lookupResult" : {
|
||||||
{
|
"found" : true,
|
||||||
"source": "rag-l0-filter-fallback",
|
"evidenceBlocks" : [ {
|
||||||
"title": "RAG L0 Filter Fallback",
|
"docId" : "rag-l0-filter-fallback",
|
||||||
"breadcrumb": "RAG > Fallback > Unfiltered Retry",
|
"chunkIndex" : 2,
|
||||||
"retrievalLayer": "L1",
|
"evidenceKey" : "rag-l0-filter-fallback#chunk-2",
|
||||||
"content": "When filtered vector retrieval is low quality, skip the L0 filter and run an unfiltered vector retry with the raw query before returning no evidence.",
|
"source" : "rag-l0-filter-fallback",
|
||||||
"score": 0.83,
|
"title" : "Unfiltered Retry",
|
||||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
"breadcrumb" : "RAG > Fallback > Unfiltered Retry",
|
||||||
},
|
"retrievalLayer" : "L1",
|
||||||
{
|
"content" : "### Unfiltered Retry\n\nIf the first vector search is over-constrained by an L0 metadata filter and returns low quality evidence,\nthe retriever should skip the L0 filter and run an unfiltered vector retry with the original query.\n\nThe fallback reason should be `filtered_vector_low_quality` when the filtered candidate exists but is below the\nreference threshold. If there is no usable evidence at all, use `filtered_vector_no_evidence`.\n\nThis document is the expected evidence for skip the L0 filter, unfiltered vector retry, and low quality behavior.",
|
||||||
"source": "rag-l0-domain-entity-hint",
|
"score" : 0.032786883413791656,
|
||||||
"title": "RAG L0 Domain Entity Hint",
|
"hitReasons" : [ "semantic_rank:1", "attempt:UNFILTERED_VECTOR_RETRY", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||||
"breadcrumb": "RAG > L0 > Domain Entity Hint",
|
}, {
|
||||||
"retrievalLayer": "L1",
|
"docId" : "rag-l0-filter-decoy",
|
||||||
"content": "L0 supplies hints for metadata filtering and explanation, but it should not be treated as final fact evidence.",
|
"chunkIndex" : 1,
|
||||||
"score": 0.66,
|
"evidenceKey" : "rag-l0-filter-decoy#chunk-1",
|
||||||
"hitReasons": ["domain_match:+0.15"]
|
"source" : "rag-l0-filter-decoy",
|
||||||
}
|
"title" : "Approval Window",
|
||||||
],
|
"breadcrumb" : "RAG > Fallback > Decoy",
|
||||||
"contextPack": {
|
"retrievalLayer" : "L1",
|
||||||
"packedText": "[1] RAG L0 Filter Fallback\nRAG > Fallback > Unfiltered Retry\nWhen filtered vector retrieval is low quality, skip the L0 filter and run an unfiltered vector retry with the raw query before returning no evidence.",
|
"content" : "## Approval Window\n\nThis document describes an unrelated release calendar approval window.\nIt intentionally avoids the real fallback instructions so the filtered retrieval\nattempt is low quality and the retriever must retry without the L0 category filter.",
|
||||||
"strategy": "top_evidence_blocks",
|
"score" : 0.0320020467042923,
|
||||||
"charBudget": 3500,
|
"hitReasons" : [ "semantic_rank:2", "attempt:UNFILTERED_VECTOR_RETRY", "l0_domain_overlap" ]
|
||||||
"usedChars": 214,
|
}, {
|
||||||
"includedSources": ["rag-l0-filter-fallback", "rag-l0-domain-entity-hint"],
|
"docId" : "rag-l0-domain-entity-hint",
|
||||||
"omittedSources": []
|
"chunkIndex" : 2,
|
||||||
|
"evidenceKey" : "rag-l0-domain-entity-hint#chunk-2",
|
||||||
|
"source" : "rag-l0-domain-entity-hint",
|
||||||
|
"title" : "Domain Entity Hint",
|
||||||
|
"breadcrumb" : "RAG > L0 > Domain Entity Hint",
|
||||||
|
"retrievalLayer" : "L1",
|
||||||
|
"content" : "### Domain Entity Hint\n\nL0 keyword matching should not decide the final retrieval result.\nIn the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.\n\nThe output can provide:\n\n1. Candidate domain hints.\n2. Matched entities and keywords.\n3. An optional metadata filter for the first vector retrieval attempt.\n\nFinal evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.\nThe important terms are domain detector, entity extractor, and metadata filter.",
|
||||||
|
"score" : 0.0320020467042923,
|
||||||
|
"hitReasons" : [ "semantic_rank:3", "attempt:UNFILTERED_VECTOR_RETRY" ]
|
||||||
|
}, {
|
||||||
|
"docId" : "rag-l0-domain-entity-hint",
|
||||||
|
"chunkIndex" : 1,
|
||||||
|
"evidenceKey" : "rag-l0-domain-entity-hint#chunk-1",
|
||||||
|
"source" : "rag-l0-domain-entity-hint",
|
||||||
|
"title" : "L0",
|
||||||
|
"breadcrumb" : "RAG > L0 > Domain Entity Hint",
|
||||||
|
"retrievalLayer" : "L1",
|
||||||
|
"content" : "## L0",
|
||||||
|
"score" : 0.03125,
|
||||||
|
"hitReasons" : [ "semantic_rank:4", "attempt:UNFILTERED_VECTOR_RETRY" ]
|
||||||
|
}, {
|
||||||
|
"docId" : "rag-chunk-context-reconstruction",
|
||||||
|
"chunkIndex" : 2,
|
||||||
|
"evidenceKey" : "rag-chunk-context-reconstruction#chunk-2",
|
||||||
|
"source" : "rag-chunk-context-reconstruction",
|
||||||
|
"title" : "Context Reconstruction",
|
||||||
|
"breadcrumb" : "RAG > Chunking > Context Reconstruction",
|
||||||
|
"retrievalLayer" : "L1",
|
||||||
|
"content" : "### Context Reconstruction\n\nWhen a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.\n\nRecommended behavior:\n\n1. Store the breadcrumb with every chunk.\n2. Preserve the same section identity across adjacent chunks.\n3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.\n4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.\n\nThe key concepts are neighbor chunk, same section, and breadcrumb.",
|
||||||
|
"score" : 0.03053613007068634,
|
||||||
|
"hitReasons" : [ "semantic_rank:5", "attempt:UNFILTERED_VECTOR_RETRY" ]
|
||||||
|
} ],
|
||||||
|
"contextPack" : {
|
||||||
|
"packedText" : "[Evidence 1]\nsource: rag-l0-filter-fallback\ntitle: Unfiltered Retry\nbreadcrumb: RAG > Fallback > Unfiltered Retry\nlayer: L1\nreasons: semantic_rank:1, attempt:UNFILTERED_VECTOR_RETRY, l0_entity_overlap, l0_keyword_overlap\ncontent:\n### Unfiltered Retry\n\nIf the first vector search is over-constrained by an L0 metadata filter and returns low quality evidence,\nthe retriever should skip the L0 filter and run an unfiltered vector retry with the original query.\n\nThe fallback reason should be `filtered_vector_low_quality` when the filtered candidate exists but is below the\nreference threshold. If there is no usable evidence at all, use `filtered_vector_no_evidence`.\n\nThis document is the expected evidence for skip the L0 filter, unfiltered vector retry, and low quality behavior.\n\n[Evidence 2]\nsource: rag-l0-filter-decoy\ntitle: Approval Window\nbreadcrumb: RAG > Fallback > Decoy\nlayer: L1\nreasons: semantic_rank:2, attempt:UNFILTERED_VECTOR_RETRY, l0_domain_overlap\ncontent:\n## Approval Window\n\nThis document describes an unrelated release calendar approval window.\nIt intentionally avoids the real fallback instructions so the filtered retrieval\nattempt is low quality and the retriever must retry without the L0 category filter.\n\n[Evidence 3]\nsource: rag-l0-domain-entity-hint\ntitle: Domain Entity Hint\nbreadcrumb: RAG > L0 > Domain Entity Hint\nlayer: L1\nreasons: semantic_rank:3, attempt:UNFILTERED_VECTOR_RETRY\ncontent:\n### Domain Entity Hint\n\nL0 keyword matching should not decide the final retrieval result.\nIn the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.\n\nThe output can provide:\n\n1. Candidate domain hints.\n2. Matched entities and keywords.\n3. An optional metadata filter for the first vector retrieval attempt.\n\nFinal evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.\nThe important terms are domain detector, entity extractor, and metadata filter.\n\n[Evidence 4]\nsource: rag-l0-domain-entity-hint\ntitle: L0\nbreadcrumb: RAG > L0 > Domain Entity Hint\nlayer: L1\nreasons: semantic_rank:4, attempt:UNFILTERED_VECTOR_RETRY\ncontent:\n## L0\n\n[Evidence 5]\nsource: rag-chunk-context-reconstruction\ntitle: Context Reconstruction\nbreadcrumb: RAG > Chunking > Context Reconstruction\nlayer: L1\nreasons: semantic_rank:5, attempt:UNFILTERED_VECTOR_RETRY\ncontent:\n### Context Reconstruction\n\nWhen a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.\n\nRecommended behavior:\n\n1. Store the breadcrumb with every chunk.\n2. Preserve the same section identity across adjacent chunks.\n3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.\n4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.\n\nThe key concepts are neighbor chunk, same section, and breadcrumb.",
|
||||||
|
"strategy" : "ranked_evidence_char_budget",
|
||||||
|
"charBudget" : 4000,
|
||||||
|
"usedChars" : 2904,
|
||||||
|
"includedSources" : [ "rag-l0-filter-fallback", "rag-l0-filter-decoy", "rag-l0-domain-entity-hint", "rag-l0-domain-entity-hint", "rag-chunk-context-reconstruction" ],
|
||||||
|
"omittedSources" : [ ]
|
||||||
},
|
},
|
||||||
"retrievalTrace": {
|
"retrievalTrace" : {
|
||||||
"originalQuery": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
"originalQuery" : "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||||
"rewrittenQuery": "RAG L0 filtered vector low quality fallback unfiltered retry",
|
"rewrittenQuery" : "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||||
"categoryFilter": "RAG",
|
"categoryFilter" : "overfilter-decoy",
|
||||||
"selectedAttempt": "UNFILTERED_VECTOR_RETRY",
|
"selectedAttempt" : "UNFILTERED_VECTOR_RETRY",
|
||||||
"fallbackReason": "filtered_vector_low_quality",
|
"fallbackReason" : "filtered_vector_low_quality",
|
||||||
"evidenceStatus": "supported",
|
"evidenceStatus" : "supported",
|
||||||
"queryHints": {
|
"queryHints" : {
|
||||||
"domains": ["RAG"],
|
"domains" : [ "overfilter-decoy" ],
|
||||||
"matched_keywords": ["L0", "low quality", "unfiltered vector retry"],
|
"matched_keywords" : [ "over-filtered by L0", "filtered vector search", "low quality evidence" ],
|
||||||
"entities": ["L0"],
|
"entities" : [ "over-filtered by L0", "filtered vector search", "low quality evidence" ],
|
||||||
"l0_titles": ["RAG L0 Domain Entity Hint"],
|
"l0_titles" : [ "RAG L0 Filter Decoy" ],
|
||||||
"l0_match_count": 1
|
"l0_match_count" : 1
|
||||||
},
|
},
|
||||||
"attempts": [
|
"attempts" : [ {
|
||||||
{
|
"name" : "FILTERED_VECTOR",
|
||||||
"name": "FILTERED_VECTOR",
|
"query" : "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||||
"query": "RAG L0 filtered vector low quality fallback unfiltered retry",
|
"categoryFilter" : "overfilter-decoy",
|
||||||
"categoryFilter": "RAG",
|
"candidateCount" : 2,
|
||||||
"candidateCount": 1,
|
"usable" : false,
|
||||||
"usable": false,
|
"errorMessage" : null,
|
||||||
"durationMs": 7,
|
"durationMs" : 722,
|
||||||
"topScore": 1.35,
|
"topScore" : 0.032786883413791656,
|
||||||
"topSimilarity": 0.325
|
"topSimilarity" : 0.47391992807388306
|
||||||
},
|
}, {
|
||||||
{
|
"name" : "UNFILTERED_VECTOR_RETRY",
|
||||||
"name": "UNFILTERED_VECTOR_RETRY",
|
"query" : "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||||
"query": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
"categoryFilter" : null,
|
||||||
"categoryFilter": null,
|
"candidateCount" : 20,
|
||||||
"candidateCount": 2,
|
"usable" : true,
|
||||||
"usable": true,
|
"errorMessage" : null,
|
||||||
"durationMs": 13,
|
"durationMs" : 2606,
|
||||||
"topScore": 0.83,
|
"topScore" : 0.032786883413791656,
|
||||||
"topSimilarity": 0.83
|
"topSimilarity" : 0.7571567445993423
|
||||||
}
|
} ]
|
||||||
]
|
|
||||||
},
|
},
|
||||||
"rerankTrace": {
|
"rerankTrace" : {
|
||||||
"items": [
|
"items" : [ {
|
||||||
{
|
"finalRank" : 1,
|
||||||
"finalRank": 1,
|
"source" : "rag-l0-filter-fallback",
|
||||||
"source": "rag-l0-filter-fallback",
|
"baseScore" : 0.7571567445993423,
|
||||||
"baseScore": 0.83,
|
"finalScore" : 0.7571567445993423,
|
||||||
"finalScore": 1.08,
|
"boostReasons" : [ "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
}, {
|
||||||
},
|
"finalRank" : 2,
|
||||||
{
|
"source" : "rag-l0-filter-decoy",
|
||||||
"finalRank": 2,
|
"baseScore" : 0.47391992807388306,
|
||||||
"source": "rag-l0-domain-entity-hint",
|
"finalScore" : 0.47391992807388306,
|
||||||
"baseScore": 0.66,
|
"boostReasons" : [ "l0_domain_overlap" ]
|
||||||
"finalScore": 0.81,
|
}, {
|
||||||
"boostReasons": ["domain_match:+0.15"]
|
"finalRank" : 3,
|
||||||
}
|
"source" : "rag-l0-domain-entity-hint",
|
||||||
]
|
"baseScore" : 0.5912361443042755,
|
||||||
}
|
"finalScore" : 0.5912361443042755,
|
||||||
|
"boostReasons" : [ ]
|
||||||
|
}, {
|
||||||
|
"finalRank" : 4,
|
||||||
|
"source" : "rag-l0-domain-entity-hint",
|
||||||
|
"baseScore" : 0.473749577999115,
|
||||||
|
"finalScore" : 0.473749577999115,
|
||||||
|
"boostReasons" : [ ]
|
||||||
|
}, {
|
||||||
|
"finalRank" : 5,
|
||||||
|
"source" : "rag-chunk-context-reconstruction",
|
||||||
|
"baseScore" : 0.4492502808570862,
|
||||||
|
"finalScore" : 0.4492502808570862,
|
||||||
|
"boostReasons" : [ ]
|
||||||
|
} ]
|
||||||
|
},
|
||||||
|
"evidenceCandidateCount" : 20,
|
||||||
|
"evidenceBlockCount" : 5,
|
||||||
|
"relevanceLevel" : "PRECISE",
|
||||||
|
"completenessHint" : "知识库中不存在比上述结果更精准的文档",
|
||||||
|
"retrievedDomainsThisSession" : null,
|
||||||
|
"message" : null
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -1,81 +1,88 @@
|
|||||||
{
|
{
|
||||||
"caseId": "chat-mysql-connection-pool",
|
"caseId" : "chat-mysql-connection-pool",
|
||||||
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
|
"query" : "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
"retrievedAt" : "2026-07-28T06:54:51.843450400Z",
|
||||||
"lookupResult": {
|
"searchMode" : "hybrid",
|
||||||
"found": true,
|
"kbScope" : "rag-eval",
|
||||||
"evidenceBlocks": [
|
"lookupResult" : {
|
||||||
{
|
"found" : true,
|
||||||
"source": "mysql-connection-pool",
|
"evidenceBlocks" : [ {
|
||||||
"title": "MySQL Connection Pool Troubleshooting",
|
"docId" : "mysql-connection-pool",
|
||||||
"breadcrumb": "Database > MySQL > Connection Pool",
|
"chunkIndex" : 2,
|
||||||
"retrievalLayer": "L1",
|
"evidenceKey" : "mysql-connection-pool#chunk-2",
|
||||||
"content": "When the connection pool is exhausted, inspect HikariCP active connections, max_connections, slow SQL, leak detection, and database wait events.",
|
"source" : "mysql-connection-pool",
|
||||||
"score": 0.86,
|
"title" : "Connection Pool",
|
||||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
"breadcrumb" : "Database > MySQL > Connection Pool",
|
||||||
},
|
"retrievalLayer" : "L1",
|
||||||
{
|
"content" : "### Connection Pool\n\nWhen MySQL connection pool is exhausted, first compare application pool usage with database `max_connections`.\nFor HikariCP, check `active`, `idle`, `pending`, and connection acquisition timeout metrics.\n\nRecommended diagnosis:\n\n1. Verify whether HikariCP active connections stay near maximum while pending threads grow.\n2. Check MySQL `Threads_connected`, `Threads_running`, and `max_connections`.\n3. Inspect slow SQL and long transactions that keep connections checked out.\n4. If the database is healthy, look for application connection leaks or missing transaction boundaries.\n\nUse this runbook as evidence for connection pool, max_connections, and HikariCP incidents.",
|
||||||
"source": "incident-diagnosis-flow",
|
"score" : 0.032786883413791656,
|
||||||
"title": "Incident Diagnosis Flow",
|
"hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||||
"breadcrumb": "AIOps > Diagnosis Flow",
|
}, {
|
||||||
"retrievalLayer": "L1",
|
"docId" : "mysql-connection-pool",
|
||||||
"content": "Collect evidence, compare metrics and logs, then verify remediation before closing the incident.",
|
"chunkIndex" : 1,
|
||||||
"score": 0.61,
|
"evidenceKey" : "mysql-connection-pool#chunk-1",
|
||||||
"hitReasons": []
|
"source" : "mysql-connection-pool",
|
||||||
}
|
"title" : "MySQL",
|
||||||
],
|
"breadcrumb" : "Database > MySQL > Connection Pool",
|
||||||
"contextPack": {
|
"retrievalLayer" : "L1",
|
||||||
"packedText": "[1] MySQL Connection Pool Troubleshooting\nDatabase > MySQL > Connection Pool\nWhen the connection pool is exhausted, inspect HikariCP active connections, max_connections, slow SQL, leak detection, and database wait events.",
|
"content" : "## MySQL",
|
||||||
"strategy": "top_evidence_blocks",
|
"score" : 0.032258063554763794,
|
||||||
"charBudget": 3500,
|
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||||
"usedChars": 216,
|
} ],
|
||||||
"includedSources": ["mysql-connection-pool", "incident-diagnosis-flow"],
|
"contextPack" : {
|
||||||
"omittedSources": []
|
"packedText" : "[Evidence 1]\nsource: mysql-connection-pool\ntitle: Connection Pool\nbreadcrumb: Database > MySQL > Connection Pool\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n### Connection Pool\n\nWhen MySQL connection pool is exhausted, first compare application pool usage with database `max_connections`.\nFor HikariCP, check `active`, `idle`, `pending`, and connection acquisition timeout metrics.\n\nRecommended diagnosis:\n\n1. Verify whether HikariCP active connections stay near maximum while pending threads grow.\n2. Check MySQL `Threads_connected`, `Threads_running`, and `max_connections`.\n3. Inspect slow SQL and long transactions that keep connections checked out.\n4. If the database is healthy, look for application connection leaks or missing transaction boundaries.\n\nUse this runbook as evidence for connection pool, max_connections, and HikariCP incidents.\n\n[Evidence 2]\nsource: mysql-connection-pool\ntitle: MySQL\nbreadcrumb: Database > MySQL > Connection Pool\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n## MySQL",
|
||||||
|
"strategy" : "ranked_evidence_char_budget",
|
||||||
|
"charBudget" : 4000,
|
||||||
|
"usedChars" : 1135,
|
||||||
|
"includedSources" : [ "mysql-connection-pool", "mysql-connection-pool" ],
|
||||||
|
"omittedSources" : [ ]
|
||||||
},
|
},
|
||||||
"retrievalTrace": {
|
"retrievalTrace" : {
|
||||||
"originalQuery": "MySQL connection pool is exhausted. How should I diagnose it?",
|
"originalQuery" : "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||||
"rewrittenQuery": "MySQL connection pool exhausted HikariCP max_connections diagnosis",
|
"rewrittenQuery" : "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||||
"categoryFilter": "Database",
|
"categoryFilter" : "database",
|
||||||
"selectedAttempt": "FILTERED_VECTOR",
|
"selectedAttempt" : "FILTERED_VECTOR",
|
||||||
"fallbackReason": null,
|
"fallbackReason" : null,
|
||||||
"evidenceStatus": "supported",
|
"evidenceStatus" : "supported",
|
||||||
"queryHints": {
|
"queryHints" : {
|
||||||
"domains": ["Database", "MySQL"],
|
"domains" : [ "database" ],
|
||||||
"matched_keywords": ["connection pool", "HikariCP", "max_connections"],
|
"matched_keywords" : [ "MySQL connection pool" ],
|
||||||
"entities": ["MySQL", "HikariCP"],
|
"entities" : [ "MySQL connection pool" ],
|
||||||
"l0_titles": ["MySQL Connection Pool Troubleshooting"],
|
"l0_titles" : [ "MySQL Connection Pool Runbook" ],
|
||||||
"l0_match_count": 1
|
"l0_match_count" : 1
|
||||||
},
|
},
|
||||||
"attempts": [
|
"attempts" : [ {
|
||||||
{
|
"name" : "FILTERED_VECTOR",
|
||||||
"name": "FILTERED_VECTOR",
|
"query" : "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||||
"query": "MySQL connection pool exhausted HikariCP max_connections diagnosis",
|
"categoryFilter" : "database",
|
||||||
"categoryFilter": "Database",
|
"candidateCount" : 3,
|
||||||
"candidateCount": 2,
|
"usable" : true,
|
||||||
"usable": true,
|
"errorMessage" : null,
|
||||||
"durationMs": 12,
|
"durationMs" : 5267,
|
||||||
"topScore": 0.86,
|
"topScore" : 0.032786883413791656,
|
||||||
"topSimilarity": 0.86
|
"topSimilarity" : 0.8114794194698334
|
||||||
}
|
} ]
|
||||||
]
|
|
||||||
},
|
},
|
||||||
"rerankTrace": {
|
"rerankTrace" : {
|
||||||
"items": [
|
"items" : [ {
|
||||||
{
|
"finalRank" : 1,
|
||||||
"finalRank": 1,
|
"source" : "mysql-connection-pool",
|
||||||
"source": "mysql-connection-pool",
|
"baseScore" : 0.8114794194698334,
|
||||||
"baseScore": 0.86,
|
"finalScore" : 0.8114794194698334,
|
||||||
"finalScore": 1.11,
|
"boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
}, {
|
||||||
},
|
"finalRank" : 2,
|
||||||
{
|
"source" : "mysql-connection-pool",
|
||||||
"finalRank": 2,
|
"baseScore" : 0.49219560623168945,
|
||||||
"source": "incident-diagnosis-flow",
|
"finalScore" : 0.49219560623168945,
|
||||||
"baseScore": 0.61,
|
"boostReasons" : [ "l0_domain_overlap" ]
|
||||||
"finalScore": 0.61,
|
} ]
|
||||||
"boostReasons": []
|
},
|
||||||
}
|
"evidenceCandidateCount" : 3,
|
||||||
]
|
"evidenceBlockCount" : 2,
|
||||||
}
|
"relevanceLevel" : "PRECISE",
|
||||||
|
"completenessHint" : "知识库中不存在比上述结果更精准的文档",
|
||||||
|
"retrievedDomainsThisSession" : null,
|
||||||
|
"message" : null
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -1,81 +1,122 @@
|
|||||||
{
|
{
|
||||||
"caseId": "chat-rag-chunk-context",
|
"caseId" : "chat-rag-chunk-context",
|
||||||
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
"query" : "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||||
"retrievedAt": "2026-07-06T00:00:00Z",
|
"retrievedAt" : "2026-07-28T06:54:51.843450400Z",
|
||||||
"lookupResult": {
|
"searchMode" : "hybrid",
|
||||||
"found": true,
|
"kbScope" : "rag-eval",
|
||||||
"evidenceBlocks": [
|
"lookupResult" : {
|
||||||
{
|
"found" : true,
|
||||||
"source": "rag-chunk-context-reconstruction",
|
"evidenceBlocks" : [ {
|
||||||
"title": "RAG Chunk Context Reconstruction",
|
"docId" : "rag-chunk-context-reconstruction",
|
||||||
"breadcrumb": "RAG > Chunking > Context Reconstruction",
|
"chunkIndex" : 2,
|
||||||
"retrievalLayer": "L1",
|
"evidenceKey" : "rag-chunk-context-reconstruction#chunk-2",
|
||||||
"content": "After a chunk hit, expand to neighbor chunk candidates from the same section and preserve breadcrumb metadata in the evidence pack.",
|
"source" : "rag-chunk-context-reconstruction",
|
||||||
"score": 0.79,
|
"title" : "Context Reconstruction",
|
||||||
"hitReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
"breadcrumb" : "RAG > Chunking > Context Reconstruction",
|
||||||
},
|
"retrievalLayer" : "L1",
|
||||||
{
|
"content" : "### Context Reconstruction\n\nWhen a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.\n\nRecommended behavior:\n\n1. Store the breadcrumb with every chunk.\n2. Preserve the same section identity across adjacent chunks.\n3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.\n4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.\n\nThe key concepts are neighbor chunk, same section, and breadcrumb.",
|
||||||
"source": "rag-breadcrumb-embedding-gap",
|
"score" : 0.032786883413791656,
|
||||||
"title": "RAG Breadcrumb Embedding Gap",
|
"hitReasons" : [ "semantic_rank:1", "attempt:FILTERED_VECTOR", "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||||
"breadcrumb": "RAG > Embedding > Breadcrumb",
|
}, {
|
||||||
"retrievalLayer": "L1",
|
"docId" : "rag-l0-domain-entity-hint",
|
||||||
"content": "Embedding title and breadcrumb with content helps recover section semantics.",
|
"chunkIndex" : 2,
|
||||||
"score": 0.72,
|
"evidenceKey" : "rag-l0-domain-entity-hint#chunk-2",
|
||||||
"hitReasons": ["domain_match:+0.15"]
|
"source" : "rag-l0-domain-entity-hint",
|
||||||
}
|
"title" : "Domain Entity Hint",
|
||||||
],
|
"breadcrumb" : "RAG > L0 > Domain Entity Hint",
|
||||||
"contextPack": {
|
"retrievalLayer" : "L1",
|
||||||
"packedText": "[1] RAG Chunk Context Reconstruction\nRAG > Chunking > Context Reconstruction\nAfter a chunk hit, expand to neighbor chunk candidates from the same section and preserve breadcrumb metadata in the evidence pack.",
|
"content" : "### Domain Entity Hint\n\nL0 keyword matching should not decide the final retrieval result.\nIn the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.\n\nThe output can provide:\n\n1. Candidate domain hints.\n2. Matched entities and keywords.\n3. An optional metadata filter for the first vector retrieval attempt.\n\nFinal evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.\nThe important terms are domain detector, entity extractor, and metadata filter.",
|
||||||
"strategy": "top_evidence_blocks",
|
"score" : 0.032258063554763794,
|
||||||
"charBudget": 3500,
|
"hitReasons" : [ "semantic_rank:2", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||||
"usedChars": 203,
|
}, {
|
||||||
"includedSources": ["rag-chunk-context-reconstruction", "rag-breadcrumb-embedding-gap"],
|
"docId" : "rag-chunk-context-reconstruction",
|
||||||
"omittedSources": []
|
"chunkIndex" : 1,
|
||||||
|
"evidenceKey" : "rag-chunk-context-reconstruction#chunk-1",
|
||||||
|
"source" : "rag-chunk-context-reconstruction",
|
||||||
|
"title" : "Chunking",
|
||||||
|
"breadcrumb" : "RAG > Chunking > Context Reconstruction",
|
||||||
|
"retrievalLayer" : "L1",
|
||||||
|
"content" : "## Chunking",
|
||||||
|
"score" : 0.01587301678955555,
|
||||||
|
"hitReasons" : [ "semantic_rank:3", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||||
|
}, {
|
||||||
|
"docId" : "rag-l0-domain-entity-hint",
|
||||||
|
"chunkIndex" : 0,
|
||||||
|
"evidenceKey" : "rag-l0-domain-entity-hint#chunk-0",
|
||||||
|
"source" : "rag-l0-domain-entity-hint",
|
||||||
|
"title" : "RAG",
|
||||||
|
"breadcrumb" : "RAG > L0 > Domain Entity Hint",
|
||||||
|
"retrievalLayer" : "L1",
|
||||||
|
"content" : "# RAG",
|
||||||
|
"score" : 0.015625,
|
||||||
|
"hitReasons" : [ "semantic_rank:4", "attempt:FILTERED_VECTOR", "l0_domain_overlap" ]
|
||||||
|
} ],
|
||||||
|
"contextPack" : {
|
||||||
|
"packedText" : "[Evidence 1]\nsource: rag-chunk-context-reconstruction\ntitle: Context Reconstruction\nbreadcrumb: RAG > Chunking > Context Reconstruction\nlayer: L1\nreasons: semantic_rank:1, attempt:FILTERED_VECTOR, l0_domain_overlap, l0_entity_overlap, l0_keyword_overlap\ncontent:\n### Context Reconstruction\n\nWhen a long section is split into multiple chunks, retrieval should keep enough local structure for the answer.\n\nRecommended behavior:\n\n1. Store the breadcrumb with every chunk.\n2. Preserve the same section identity across adjacent chunks.\n3. During context packing, include a neighbor chunk when the selected chunk depends on nearby setup or definitions.\n4. Prefer concise evidence blocks that show the breadcrumb and the relevant content span.\n\nThe key concepts are neighbor chunk, same section, and breadcrumb.\n\n[Evidence 2]\nsource: rag-l0-domain-entity-hint\ntitle: Domain Entity Hint\nbreadcrumb: RAG > L0 > Domain Entity Hint\nlayer: L1\nreasons: semantic_rank:2, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n### Domain Entity Hint\n\nL0 keyword matching should not decide the final retrieval result.\nIn the modular RAG pipeline, L0 behaves like a lightweight domain detector and entity extractor.\n\nThe output can provide:\n\n1. Candidate domain hints.\n2. Matched entities and keywords.\n3. An optional metadata filter for the first vector retrieval attempt.\n\nFinal evidence still comes from L1 vector retrieval, post-retrieval normalization, rerank, and context packing.\nThe important terms are domain detector, entity extractor, and metadata filter.\n\n[Evidence 3]\nsource: rag-chunk-context-reconstruction\ntitle: Chunking\nbreadcrumb: RAG > Chunking > Context Reconstruction\nlayer: L1\nreasons: semantic_rank:3, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n## Chunking\n\n[Evidence 4]\nsource: rag-l0-domain-entity-hint\ntitle: RAG\nbreadcrumb: RAG > L0 > Domain Entity Hint\nlayer: L1\nreasons: semantic_rank:4, attempt:FILTERED_VECTOR, l0_domain_overlap\ncontent:\n# RAG",
|
||||||
|
"strategy" : "ranked_evidence_char_budget",
|
||||||
|
"charBudget" : 4000,
|
||||||
|
"usedChars" : 1966,
|
||||||
|
"includedSources" : [ "rag-chunk-context-reconstruction", "rag-l0-domain-entity-hint", "rag-chunk-context-reconstruction", "rag-l0-domain-entity-hint" ],
|
||||||
|
"omittedSources" : [ ]
|
||||||
},
|
},
|
||||||
"retrievalTrace": {
|
"retrievalTrace" : {
|
||||||
"originalQuery": "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
"originalQuery" : "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||||
"rewrittenQuery": "RAG chunk context reconstruction neighbor chunk same section breadcrumb",
|
"rewrittenQuery" : "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||||
"categoryFilter": "RAG",
|
"categoryFilter" : "rag",
|
||||||
"selectedAttempt": "FILTERED_VECTOR",
|
"selectedAttempt" : "FILTERED_VECTOR",
|
||||||
"fallbackReason": null,
|
"fallbackReason" : null,
|
||||||
"evidenceStatus": "supported",
|
"evidenceStatus" : "supported",
|
||||||
"queryHints": {
|
"queryHints" : {
|
||||||
"domains": ["RAG"],
|
"domains" : [ "rag" ],
|
||||||
"matched_keywords": ["neighbor chunk", "same section", "breadcrumb"],
|
"matched_keywords" : [ "split into multiple chunks", "retrieval context" ],
|
||||||
"entities": ["chunk", "breadcrumb"],
|
"entities" : [ "split into multiple chunks", "retrieval context" ],
|
||||||
"l0_titles": ["RAG Chunk Context Reconstruction"],
|
"l0_titles" : [ "RAG Chunk Context Reconstruction" ],
|
||||||
"l0_match_count": 1
|
"l0_match_count" : 1
|
||||||
},
|
},
|
||||||
"attempts": [
|
"attempts" : [ {
|
||||||
{
|
"name" : "FILTERED_VECTOR",
|
||||||
"name": "FILTERED_VECTOR",
|
"query" : "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||||
"query": "RAG chunk context reconstruction neighbor chunk same section breadcrumb",
|
"categoryFilter" : "rag",
|
||||||
"categoryFilter": "RAG",
|
"candidateCount" : 6,
|
||||||
"candidateCount": 2,
|
"usable" : true,
|
||||||
"usable": true,
|
"errorMessage" : null,
|
||||||
"durationMs": 9,
|
"durationMs" : 743,
|
||||||
"topScore": 0.79,
|
"topScore" : 0.032786883413791656,
|
||||||
"topSimilarity": 0.79
|
"topSimilarity" : 0.7487991750240326
|
||||||
}
|
} ]
|
||||||
]
|
|
||||||
},
|
},
|
||||||
"rerankTrace": {
|
"rerankTrace" : {
|
||||||
"items": [
|
"items" : [ {
|
||||||
{
|
"finalRank" : 1,
|
||||||
"finalRank": 1,
|
"source" : "rag-chunk-context-reconstruction",
|
||||||
"source": "rag-chunk-context-reconstruction",
|
"baseScore" : 0.7487991750240326,
|
||||||
"baseScore": 0.79,
|
"finalScore" : 0.7487991750240326,
|
||||||
"finalScore": 1.04,
|
"boostReasons" : [ "l0_domain_overlap", "l0_entity_overlap", "l0_keyword_overlap" ]
|
||||||
"boostReasons": ["domain_match:+0.15", "keyword_match:+0.10"]
|
}, {
|
||||||
},
|
"finalRank" : 2,
|
||||||
{
|
"source" : "rag-l0-domain-entity-hint",
|
||||||
"finalRank": 2,
|
"baseScore" : 0.4840593934059143,
|
||||||
"source": "rag-breadcrumb-embedding-gap",
|
"finalScore" : 0.4840593934059143,
|
||||||
"baseScore": 0.72,
|
"boostReasons" : [ "l0_domain_overlap" ]
|
||||||
"finalScore": 0.87,
|
}, {
|
||||||
"boostReasons": ["domain_match:+0.15"]
|
"finalRank" : 3,
|
||||||
}
|
"source" : "rag-chunk-context-reconstruction",
|
||||||
]
|
"baseScore" : 0.42978107929229736,
|
||||||
}
|
"finalScore" : 0.42978107929229736,
|
||||||
|
"boostReasons" : [ "l0_domain_overlap" ]
|
||||||
|
}, {
|
||||||
|
"finalRank" : 4,
|
||||||
|
"source" : "rag-l0-domain-entity-hint",
|
||||||
|
"baseScore" : 0.3859822154045105,
|
||||||
|
"finalScore" : 0.3859822154045105,
|
||||||
|
"boostReasons" : [ "l0_domain_overlap" ]
|
||||||
|
} ]
|
||||||
|
},
|
||||||
|
"evidenceCandidateCount" : 6,
|
||||||
|
"evidenceBlockCount" : 4,
|
||||||
|
"relevanceLevel" : "REFERENCE",
|
||||||
|
"completenessHint" : "当前结果为相关参考,如需更精准信息请明确缺少的具体维度",
|
||||||
|
"retrievedDomainsThisSession" : null,
|
||||||
|
"message" : null
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -0,0 +1,107 @@
|
|||||||
|
{
|
||||||
|
"generatedAt": "2026-07-28T06:48:56.439554+00:00",
|
||||||
|
"baselineReport": "eval/rag-retrieval/cases/golden-cases.json",
|
||||||
|
"currentReport": "eval/rag-retrieval/cases/golden-cases.json",
|
||||||
|
"baselineCaseCount": 7,
|
||||||
|
"currentCaseCount": 7,
|
||||||
|
"baselinePassRate": 1.0,
|
||||||
|
"currentPassRate": 0.8571,
|
||||||
|
"baselineRecallAtK": 1.0,
|
||||||
|
"currentRecallAtK": 0.8571,
|
||||||
|
"regressionCount": 6,
|
||||||
|
"improvementCount": 0,
|
||||||
|
"changedCount": 3,
|
||||||
|
"hasRegression": true,
|
||||||
|
"items": [
|
||||||
|
{
|
||||||
|
"type": "REGRESSION",
|
||||||
|
"scope": "aggregate",
|
||||||
|
"caseId": null,
|
||||||
|
"metric": "passRate",
|
||||||
|
"baselineValue": "1.0",
|
||||||
|
"currentValue": "0.8571",
|
||||||
|
"delta": -0.14290000000000003,
|
||||||
|
"message": "aggregate passRate changed"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"type": "REGRESSION",
|
||||||
|
"scope": "aggregate",
|
||||||
|
"caseId": null,
|
||||||
|
"metric": "recallAtK",
|
||||||
|
"baselineValue": "1.0",
|
||||||
|
"currentValue": "0.8571",
|
||||||
|
"delta": -0.14290000000000003,
|
||||||
|
"message": "aggregate recallAtK changed"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"type": "REGRESSION",
|
||||||
|
"scope": "aggregate",
|
||||||
|
"caseId": null,
|
||||||
|
"metric": "strongHitRate",
|
||||||
|
"baselineValue": "1.0",
|
||||||
|
"currentValue": "0.8571",
|
||||||
|
"delta": -0.14290000000000003,
|
||||||
|
"message": "aggregate strongHitRate changed"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"type": "REGRESSION",
|
||||||
|
"scope": "case",
|
||||||
|
"caseId": "chat-l0-filter-fallback",
|
||||||
|
"metric": "passed",
|
||||||
|
"baselineValue": "True",
|
||||||
|
"currentValue": "False",
|
||||||
|
"delta": -1.0,
|
||||||
|
"message": "chat-l0-filter-fallback passed changed"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"type": "REGRESSION",
|
||||||
|
"scope": "case",
|
||||||
|
"caseId": "chat-l0-filter-fallback",
|
||||||
|
"metric": "hitLevel",
|
||||||
|
"baselineValue": "strong",
|
||||||
|
"currentValue": "weak",
|
||||||
|
"delta": -2.0,
|
||||||
|
"message": "chat-l0-filter-fallback hitLevel changed"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"type": "REGRESSION",
|
||||||
|
"scope": "case",
|
||||||
|
"caseId": "chat-l0-filter-fallback",
|
||||||
|
"metric": "firstExpectedRank",
|
||||||
|
"baselineValue": "1",
|
||||||
|
"currentValue": "-",
|
||||||
|
"delta": null,
|
||||||
|
"message": "chat-l0-filter-fallback firstExpectedRank changed"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"type": "CHANGED",
|
||||||
|
"scope": "case",
|
||||||
|
"caseId": "chat-l0-filter-fallback",
|
||||||
|
"metric": "selectedAttempt",
|
||||||
|
"baselineValue": "UNFILTERED_VECTOR_RETRY",
|
||||||
|
"currentValue": "FILTERED_VECTOR",
|
||||||
|
"delta": null,
|
||||||
|
"message": "chat-l0-filter-fallback selectedAttempt changed"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"type": "CHANGED",
|
||||||
|
"scope": "case",
|
||||||
|
"caseId": "chat-l0-filter-fallback",
|
||||||
|
"metric": "fallbackReason",
|
||||||
|
"baselineValue": "filtered_vector_low_quality",
|
||||||
|
"currentValue": "-",
|
||||||
|
"delta": null,
|
||||||
|
"message": "chat-l0-filter-fallback fallbackReason changed"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"type": "CHANGED",
|
||||||
|
"scope": "case",
|
||||||
|
"caseId": "chat-l0-filter-fallback",
|
||||||
|
"metric": "rerankTopSource",
|
||||||
|
"baselineValue": "rag-l0-filter-fallback",
|
||||||
|
"currentValue": "rag-l0-filter-decoy",
|
||||||
|
"delta": null,
|
||||||
|
"message": "chat-l0-filter-fallback rerankTopSource changed"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
@@ -0,0 +1,31 @@
|
|||||||
|
# RAG Retrieval Baseline Diff
|
||||||
|
|
||||||
|
Generated at: `2026-07-28T06:48:56.439554+00:00`
|
||||||
|
|
||||||
|
## Summary
|
||||||
|
|
||||||
|
| Metric | Value |
|
||||||
|
|---|---:|
|
||||||
|
| Baseline cases | 7 |
|
||||||
|
| Current cases | 7 |
|
||||||
|
| Baseline pass rate | 1.0 |
|
||||||
|
| Current pass rate | 0.8571 |
|
||||||
|
| Baseline recall@K | 1.0 |
|
||||||
|
| Current recall@K | 0.8571 |
|
||||||
|
| Regressions | 6 |
|
||||||
|
| Improvements | 0 |
|
||||||
|
| Changed | 3 |
|
||||||
|
|
||||||
|
## Items
|
||||||
|
|
||||||
|
| Type | Scope | Case | Metric | Baseline | Current | Delta | Message |
|
||||||
|
|---|---|---|---|---|---|---:|---|
|
||||||
|
| REGRESSION | aggregate | | passRate | 1.0 | 0.8571 | -0.14290000000000003 | aggregate passRate changed |
|
||||||
|
| REGRESSION | aggregate | | recallAtK | 1.0 | 0.8571 | -0.14290000000000003 | aggregate recallAtK changed |
|
||||||
|
| REGRESSION | aggregate | | strongHitRate | 1.0 | 0.8571 | -0.14290000000000003 | aggregate strongHitRate changed |
|
||||||
|
| REGRESSION | case | chat-l0-filter-fallback | passed | True | False | -1.0 | chat-l0-filter-fallback passed changed |
|
||||||
|
| REGRESSION | case | chat-l0-filter-fallback | hitLevel | strong | weak | -2.0 | chat-l0-filter-fallback hitLevel changed |
|
||||||
|
| REGRESSION | case | chat-l0-filter-fallback | firstExpectedRank | 1 | - | | chat-l0-filter-fallback firstExpectedRank changed |
|
||||||
|
| CHANGED | case | chat-l0-filter-fallback | selectedAttempt | UNFILTERED_VECTOR_RETRY | FILTERED_VECTOR | | chat-l0-filter-fallback selectedAttempt changed |
|
||||||
|
| CHANGED | case | chat-l0-filter-fallback | fallbackReason | filtered_vector_low_quality | - | | chat-l0-filter-fallback fallbackReason changed |
|
||||||
|
| CHANGED | case | chat-l0-filter-fallback | rerankTopSource | rag-l0-filter-fallback | rag-l0-filter-decoy | | chat-l0-filter-fallback rerankTopSource changed |
|
||||||
@@ -1,5 +1,5 @@
|
|||||||
{
|
{
|
||||||
"generatedAt": "2026-07-06T13:37:59.726351+00:00",
|
"generatedAt": "2026-07-28T06:55:09.379164+00:00",
|
||||||
"caseFile": "eval/rag-retrieval/cases/golden-cases.json",
|
"caseFile": "eval/rag-retrieval/cases/golden-cases.json",
|
||||||
"fixtureDir": "eval/rag-retrieval/fixtures",
|
"fixtureDir": "eval/rag-retrieval/fixtures",
|
||||||
"aggregate": {
|
"aggregate": {
|
||||||
@@ -28,7 +28,7 @@
|
|||||||
"firstExpectedRank": 1,
|
"firstExpectedRank": 1,
|
||||||
"topCandidates": [
|
"topCandidates": [
|
||||||
"1:mysql-connection-pool",
|
"1:mysql-connection-pool",
|
||||||
"2:incident-diagnosis-flow"
|
"2:mysql-connection-pool"
|
||||||
],
|
],
|
||||||
"matchedKeywords": [
|
"matchedKeywords": [
|
||||||
"connection pool",
|
"connection pool",
|
||||||
@@ -41,7 +41,7 @@
|
|||||||
"evidenceStatus": "supported",
|
"evidenceStatus": "supported",
|
||||||
"includedSources": [
|
"includedSources": [
|
||||||
"mysql-connection-pool",
|
"mysql-connection-pool",
|
||||||
"incident-diagnosis-flow"
|
"mysql-connection-pool"
|
||||||
],
|
],
|
||||||
"omittedSources": [],
|
"omittedSources": [],
|
||||||
"rerankTopSource": "mysql-connection-pool",
|
"rerankTopSource": "mysql-connection-pool",
|
||||||
@@ -57,7 +57,7 @@
|
|||||||
"firstExpectedRank": 1,
|
"firstExpectedRank": 1,
|
||||||
"topCandidates": [
|
"topCandidates": [
|
||||||
"1:incident-diagnosis-flow",
|
"1:incident-diagnosis-flow",
|
||||||
"2:rag-chunk-context-reconstruction"
|
"2:incident-diagnosis-flow"
|
||||||
],
|
],
|
||||||
"matchedKeywords": [
|
"matchedKeywords": [
|
||||||
"collect evidence",
|
"collect evidence",
|
||||||
@@ -70,7 +70,7 @@
|
|||||||
"evidenceStatus": "supported",
|
"evidenceStatus": "supported",
|
||||||
"includedSources": [
|
"includedSources": [
|
||||||
"incident-diagnosis-flow",
|
"incident-diagnosis-flow",
|
||||||
"rag-chunk-context-reconstruction"
|
"incident-diagnosis-flow"
|
||||||
],
|
],
|
||||||
"omittedSources": [],
|
"omittedSources": [],
|
||||||
"rerankTopSource": "incident-diagnosis-flow",
|
"rerankTopSource": "incident-diagnosis-flow",
|
||||||
@@ -86,7 +86,9 @@
|
|||||||
"firstExpectedRank": 1,
|
"firstExpectedRank": 1,
|
||||||
"topCandidates": [
|
"topCandidates": [
|
||||||
"1:payment-service-latency",
|
"1:payment-service-latency",
|
||||||
"2:mysql-connection-pool"
|
"2:aiops-alert-scope-control",
|
||||||
|
"3:payment-service-latency",
|
||||||
|
"4:aiops-alert-scope-control"
|
||||||
],
|
],
|
||||||
"matchedKeywords": [
|
"matchedKeywords": [
|
||||||
"p95 latency",
|
"p95 latency",
|
||||||
@@ -99,7 +101,9 @@
|
|||||||
"evidenceStatus": "supported",
|
"evidenceStatus": "supported",
|
||||||
"includedSources": [
|
"includedSources": [
|
||||||
"payment-service-latency",
|
"payment-service-latency",
|
||||||
"mysql-connection-pool"
|
"aiops-alert-scope-control",
|
||||||
|
"payment-service-latency",
|
||||||
|
"aiops-alert-scope-control"
|
||||||
],
|
],
|
||||||
"omittedSources": [],
|
"omittedSources": [],
|
||||||
"rerankTopSource": "payment-service-latency",
|
"rerankTopSource": "payment-service-latency",
|
||||||
@@ -114,7 +118,10 @@
|
|||||||
"passed": true,
|
"passed": true,
|
||||||
"firstExpectedRank": 1,
|
"firstExpectedRank": 1,
|
||||||
"topCandidates": [
|
"topCandidates": [
|
||||||
"1:aiops-alert-scope-control"
|
"1:aiops-alert-scope-control",
|
||||||
|
"2:payment-service-latency",
|
||||||
|
"3:payment-service-latency",
|
||||||
|
"4:aiops-alert-scope-control"
|
||||||
],
|
],
|
||||||
"matchedKeywords": [
|
"matchedKeywords": [
|
||||||
"payload",
|
"payload",
|
||||||
@@ -126,6 +133,9 @@
|
|||||||
"fallbackReason": null,
|
"fallbackReason": null,
|
||||||
"evidenceStatus": "supported",
|
"evidenceStatus": "supported",
|
||||||
"includedSources": [
|
"includedSources": [
|
||||||
|
"aiops-alert-scope-control",
|
||||||
|
"payment-service-latency",
|
||||||
|
"payment-service-latency",
|
||||||
"aiops-alert-scope-control"
|
"aiops-alert-scope-control"
|
||||||
],
|
],
|
||||||
"omittedSources": [],
|
"omittedSources": [],
|
||||||
@@ -142,7 +152,9 @@
|
|||||||
"firstExpectedRank": 1,
|
"firstExpectedRank": 1,
|
||||||
"topCandidates": [
|
"topCandidates": [
|
||||||
"1:rag-chunk-context-reconstruction",
|
"1:rag-chunk-context-reconstruction",
|
||||||
"2:rag-breadcrumb-embedding-gap"
|
"2:rag-l0-domain-entity-hint",
|
||||||
|
"3:rag-chunk-context-reconstruction",
|
||||||
|
"4:rag-l0-domain-entity-hint"
|
||||||
],
|
],
|
||||||
"matchedKeywords": [
|
"matchedKeywords": [
|
||||||
"neighbor chunk",
|
"neighbor chunk",
|
||||||
@@ -155,7 +167,9 @@
|
|||||||
"evidenceStatus": "supported",
|
"evidenceStatus": "supported",
|
||||||
"includedSources": [
|
"includedSources": [
|
||||||
"rag-chunk-context-reconstruction",
|
"rag-chunk-context-reconstruction",
|
||||||
"rag-breadcrumb-embedding-gap"
|
"rag-l0-domain-entity-hint",
|
||||||
|
"rag-chunk-context-reconstruction",
|
||||||
|
"rag-l0-domain-entity-hint"
|
||||||
],
|
],
|
||||||
"omittedSources": [],
|
"omittedSources": [],
|
||||||
"rerankTopSource": "rag-chunk-context-reconstruction",
|
"rerankTopSource": "rag-chunk-context-reconstruction",
|
||||||
@@ -171,7 +185,9 @@
|
|||||||
"firstExpectedRank": 1,
|
"firstExpectedRank": 1,
|
||||||
"topCandidates": [
|
"topCandidates": [
|
||||||
"1:rag-l0-domain-entity-hint",
|
"1:rag-l0-domain-entity-hint",
|
||||||
"2:rag-l0-l1-fusion-ranking"
|
"2:rag-chunk-context-reconstruction",
|
||||||
|
"3:rag-l0-domain-entity-hint",
|
||||||
|
"4:rag-chunk-context-reconstruction"
|
||||||
],
|
],
|
||||||
"matchedKeywords": [
|
"matchedKeywords": [
|
||||||
"domain detector",
|
"domain detector",
|
||||||
@@ -184,7 +200,9 @@
|
|||||||
"evidenceStatus": "supported",
|
"evidenceStatus": "supported",
|
||||||
"includedSources": [
|
"includedSources": [
|
||||||
"rag-l0-domain-entity-hint",
|
"rag-l0-domain-entity-hint",
|
||||||
"rag-l0-l1-fusion-ranking"
|
"rag-chunk-context-reconstruction",
|
||||||
|
"rag-l0-domain-entity-hint",
|
||||||
|
"rag-chunk-context-reconstruction"
|
||||||
],
|
],
|
||||||
"omittedSources": [],
|
"omittedSources": [],
|
||||||
"rerankTopSource": "rag-l0-domain-entity-hint",
|
"rerankTopSource": "rag-l0-domain-entity-hint",
|
||||||
@@ -200,7 +218,10 @@
|
|||||||
"firstExpectedRank": 1,
|
"firstExpectedRank": 1,
|
||||||
"topCandidates": [
|
"topCandidates": [
|
||||||
"1:rag-l0-filter-fallback",
|
"1:rag-l0-filter-fallback",
|
||||||
"2:rag-l0-domain-entity-hint"
|
"2:rag-l0-filter-decoy",
|
||||||
|
"3:rag-l0-domain-entity-hint",
|
||||||
|
"4:rag-l0-domain-entity-hint",
|
||||||
|
"5:rag-chunk-context-reconstruction"
|
||||||
],
|
],
|
||||||
"matchedKeywords": [
|
"matchedKeywords": [
|
||||||
"skip the l0 filter",
|
"skip the l0 filter",
|
||||||
@@ -213,7 +234,10 @@
|
|||||||
"evidenceStatus": "supported",
|
"evidenceStatus": "supported",
|
||||||
"includedSources": [
|
"includedSources": [
|
||||||
"rag-l0-filter-fallback",
|
"rag-l0-filter-fallback",
|
||||||
"rag-l0-domain-entity-hint"
|
"rag-l0-filter-decoy",
|
||||||
|
"rag-l0-domain-entity-hint",
|
||||||
|
"rag-l0-domain-entity-hint",
|
||||||
|
"rag-chunk-context-reconstruction"
|
||||||
],
|
],
|
||||||
"omittedSources": [],
|
"omittedSources": [],
|
||||||
"rerankTopSource": "rag-l0-filter-fallback",
|
"rerankTopSource": "rag-l0-filter-fallback",
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
# RAG Retrieval Baseline
|
# RAG Retrieval Baseline
|
||||||
|
|
||||||
Generated at: `2026-07-06T13:37:59.726351+00:00`
|
Generated at: `2026-07-28T06:55:09.379164+00:00`
|
||||||
|
|
||||||
## Aggregate
|
## Aggregate
|
||||||
|
|
||||||
@@ -24,10 +24,10 @@ Generated at: `2026-07-06T13:37:59.726351+00:00`
|
|||||||
|
|
||||||
| Case | Scenario | Pass | Hit | Attempt | Fallback | Evidence | First Expected Rank | Top Candidates | Failed Checks |
|
| Case | Scenario | Pass | Hit | Attempt | Fallback | Evidence | First Expected Rank | Top Candidates | Failed Checks |
|
||||||
|---|---|---|---|---|---|---|---:|---|---|
|
|---|---|---|---|---|---|---|---:|---|---|
|
||||||
| chat-mysql-connection-pool | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:mysql-connection-pool<br>2:incident-diagnosis-flow | |
|
| chat-mysql-connection-pool | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:mysql-connection-pool<br>2:mysql-connection-pool | |
|
||||||
| chat-diagnosis-flow | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:incident-diagnosis-flow<br>2:rag-chunk-context-reconstruction | |
|
| chat-diagnosis-flow | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:incident-diagnosis-flow<br>2:incident-diagnosis-flow | |
|
||||||
| aiops-payment-latency-alert | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:payment-service-latency<br>2:mysql-connection-pool | |
|
| aiops-payment-latency-alert | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:payment-service-latency<br>2:aiops-alert-scope-control<br>3:payment-service-latency<br>4:aiops-alert-scope-control | |
|
||||||
| aiops-prometheus-alert-scope | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:aiops-alert-scope-control | |
|
| aiops-prometheus-alert-scope | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:aiops-alert-scope-control<br>2:payment-service-latency<br>3:payment-service-latency<br>4:aiops-alert-scope-control | |
|
||||||
| chat-rag-chunk-context | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-chunk-context-reconstruction<br>2:rag-breadcrumb-embedding-gap | |
|
| chat-rag-chunk-context | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-chunk-context-reconstruction<br>2:rag-l0-domain-entity-hint<br>3:rag-chunk-context-reconstruction<br>4:rag-l0-domain-entity-hint | |
|
||||||
| chat-l0-domain-hint | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-l0-domain-entity-hint<br>2:rag-l0-l1-fusion-ranking | |
|
| chat-l0-domain-hint | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-l0-domain-entity-hint<br>2:rag-chunk-context-reconstruction<br>3:rag-l0-domain-entity-hint<br>4:rag-chunk-context-reconstruction | |
|
||||||
| chat-l0-filter-fallback | chat | true | strong | UNFILTERED_VECTOR_RETRY | filtered_vector_low_quality | supported | 1 | 1:rag-l0-filter-fallback<br>2:rag-l0-domain-entity-hint | |
|
| chat-l0-filter-fallback | chat | true | strong | UNFILTERED_VECTOR_RETRY | filtered_vector_low_quality | supported | 1 | 1:rag-l0-filter-fallback<br>2:rag-l0-filter-decoy<br>3:rag-l0-domain-entity-hint<br>4:rag-l0-domain-entity-hint<br>5:rag-chunk-context-reconstruction | |
|
||||||
|
|||||||
@@ -0,0 +1,245 @@
|
|||||||
|
{
|
||||||
|
"generatedAt": "2026-07-28T06:48:56.421877+00:00",
|
||||||
|
"caseFile": "eval/rag-retrieval/cases/golden-cases.json",
|
||||||
|
"fixtureDir": "eval/rag-retrieval/fixtures",
|
||||||
|
"aggregate": {
|
||||||
|
"caseCount": 7,
|
||||||
|
"topK": 5,
|
||||||
|
"passedCount": 6,
|
||||||
|
"failedCount": 1,
|
||||||
|
"passRate": 0.8571,
|
||||||
|
"lookupResultCaseCount": 7,
|
||||||
|
"strongHitCount": 6,
|
||||||
|
"mediumHitCount": 0,
|
||||||
|
"weakHitCount": 1,
|
||||||
|
"missCount": 0,
|
||||||
|
"recallAtK": 0.8571,
|
||||||
|
"strongHitRate": 0.8571,
|
||||||
|
"averageFirstHitRank": 1.0
|
||||||
|
},
|
||||||
|
"results": [
|
||||||
|
{
|
||||||
|
"caseId": "chat-mysql-connection-pool",
|
||||||
|
"scenario": "chat",
|
||||||
|
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
|
||||||
|
"dataShape": "lookupResult",
|
||||||
|
"hitLevel": "strong",
|
||||||
|
"passed": true,
|
||||||
|
"firstExpectedRank": 1,
|
||||||
|
"topCandidates": [
|
||||||
|
"1:mysql-connection-pool",
|
||||||
|
"2:mysql-connection-pool"
|
||||||
|
],
|
||||||
|
"matchedKeywords": [
|
||||||
|
"connection pool",
|
||||||
|
"max_connections",
|
||||||
|
"hikaricp"
|
||||||
|
],
|
||||||
|
"breadcrumbMatched": true,
|
||||||
|
"selectedAttempt": "FILTERED_VECTOR",
|
||||||
|
"fallbackReason": null,
|
||||||
|
"evidenceStatus": "supported",
|
||||||
|
"includedSources": [
|
||||||
|
"mysql-connection-pool",
|
||||||
|
"mysql-connection-pool"
|
||||||
|
],
|
||||||
|
"omittedSources": [],
|
||||||
|
"rerankTopSource": "mysql-connection-pool",
|
||||||
|
"failedChecks": []
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"caseId": "chat-diagnosis-flow",
|
||||||
|
"scenario": "chat",
|
||||||
|
"query": "What is the standard troubleshooting flow for an application incident?",
|
||||||
|
"dataShape": "lookupResult",
|
||||||
|
"hitLevel": "strong",
|
||||||
|
"passed": true,
|
||||||
|
"firstExpectedRank": 1,
|
||||||
|
"topCandidates": [
|
||||||
|
"1:incident-diagnosis-flow",
|
||||||
|
"2:incident-diagnosis-flow"
|
||||||
|
],
|
||||||
|
"matchedKeywords": [
|
||||||
|
"collect evidence",
|
||||||
|
"verify",
|
||||||
|
"remediation"
|
||||||
|
],
|
||||||
|
"breadcrumbMatched": true,
|
||||||
|
"selectedAttempt": "FILTERED_VECTOR",
|
||||||
|
"fallbackReason": null,
|
||||||
|
"evidenceStatus": "supported",
|
||||||
|
"includedSources": [
|
||||||
|
"incident-diagnosis-flow",
|
||||||
|
"incident-diagnosis-flow"
|
||||||
|
],
|
||||||
|
"omittedSources": [],
|
||||||
|
"rerankTopSource": "incident-diagnosis-flow",
|
||||||
|
"failedChecks": []
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"caseId": "aiops-payment-latency-alert",
|
||||||
|
"scenario": "aiops",
|
||||||
|
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
|
||||||
|
"dataShape": "lookupResult",
|
||||||
|
"hitLevel": "strong",
|
||||||
|
"passed": true,
|
||||||
|
"firstExpectedRank": 1,
|
||||||
|
"topCandidates": [
|
||||||
|
"1:payment-service-latency",
|
||||||
|
"2:aiops-alert-scope-control",
|
||||||
|
"3:payment-service-latency",
|
||||||
|
"4:aiops-alert-scope-control"
|
||||||
|
],
|
||||||
|
"matchedKeywords": [
|
||||||
|
"p95 latency",
|
||||||
|
"payment-service",
|
||||||
|
"downstream dependency"
|
||||||
|
],
|
||||||
|
"breadcrumbMatched": true,
|
||||||
|
"selectedAttempt": "FILTERED_VECTOR",
|
||||||
|
"fallbackReason": null,
|
||||||
|
"evidenceStatus": "supported",
|
||||||
|
"includedSources": [
|
||||||
|
"payment-service-latency",
|
||||||
|
"aiops-alert-scope-control",
|
||||||
|
"payment-service-latency",
|
||||||
|
"aiops-alert-scope-control"
|
||||||
|
],
|
||||||
|
"omittedSources": [],
|
||||||
|
"rerankTopSource": "payment-service-latency",
|
||||||
|
"failedChecks": []
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"caseId": "aiops-prometheus-alert-scope",
|
||||||
|
"scenario": "aiops",
|
||||||
|
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
|
||||||
|
"dataShape": "lookupResult",
|
||||||
|
"hitLevel": "strong",
|
||||||
|
"passed": true,
|
||||||
|
"firstExpectedRank": 1,
|
||||||
|
"topCandidates": [
|
||||||
|
"1:aiops-alert-scope-control",
|
||||||
|
"2:payment-service-latency",
|
||||||
|
"3:payment-service-latency",
|
||||||
|
"4:aiops-alert-scope-control"
|
||||||
|
],
|
||||||
|
"matchedKeywords": [
|
||||||
|
"payload",
|
||||||
|
"unrelated active alerts",
|
||||||
|
"scope"
|
||||||
|
],
|
||||||
|
"breadcrumbMatched": true,
|
||||||
|
"selectedAttempt": "FILTERED_VECTOR",
|
||||||
|
"fallbackReason": null,
|
||||||
|
"evidenceStatus": "supported",
|
||||||
|
"includedSources": [
|
||||||
|
"aiops-alert-scope-control",
|
||||||
|
"payment-service-latency",
|
||||||
|
"payment-service-latency",
|
||||||
|
"aiops-alert-scope-control"
|
||||||
|
],
|
||||||
|
"omittedSources": [],
|
||||||
|
"rerankTopSource": "aiops-alert-scope-control",
|
||||||
|
"failedChecks": []
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"caseId": "chat-rag-chunk-context",
|
||||||
|
"scenario": "chat",
|
||||||
|
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
|
||||||
|
"dataShape": "lookupResult",
|
||||||
|
"hitLevel": "strong",
|
||||||
|
"passed": true,
|
||||||
|
"firstExpectedRank": 1,
|
||||||
|
"topCandidates": [
|
||||||
|
"1:rag-chunk-context-reconstruction",
|
||||||
|
"2:rag-l0-domain-entity-hint",
|
||||||
|
"3:rag-chunk-context-reconstruction",
|
||||||
|
"4:rag-l0-domain-entity-hint"
|
||||||
|
],
|
||||||
|
"matchedKeywords": [
|
||||||
|
"neighbor chunk",
|
||||||
|
"same section",
|
||||||
|
"breadcrumb"
|
||||||
|
],
|
||||||
|
"breadcrumbMatched": true,
|
||||||
|
"selectedAttempt": "FILTERED_VECTOR",
|
||||||
|
"fallbackReason": null,
|
||||||
|
"evidenceStatus": "supported",
|
||||||
|
"includedSources": [
|
||||||
|
"rag-chunk-context-reconstruction",
|
||||||
|
"rag-l0-domain-entity-hint",
|
||||||
|
"rag-chunk-context-reconstruction",
|
||||||
|
"rag-l0-domain-entity-hint"
|
||||||
|
],
|
||||||
|
"omittedSources": [],
|
||||||
|
"rerankTopSource": "rag-chunk-context-reconstruction",
|
||||||
|
"failedChecks": []
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"caseId": "chat-l0-domain-hint",
|
||||||
|
"scenario": "chat",
|
||||||
|
"query": "Should L0 keyword matching decide the final retrieval result?",
|
||||||
|
"dataShape": "lookupResult",
|
||||||
|
"hitLevel": "strong",
|
||||||
|
"passed": true,
|
||||||
|
"firstExpectedRank": 1,
|
||||||
|
"topCandidates": [
|
||||||
|
"1:rag-l0-domain-entity-hint",
|
||||||
|
"2:rag-chunk-context-reconstruction",
|
||||||
|
"3:rag-l0-domain-entity-hint",
|
||||||
|
"4:rag-chunk-context-reconstruction"
|
||||||
|
],
|
||||||
|
"matchedKeywords": [
|
||||||
|
"domain detector",
|
||||||
|
"entity extractor",
|
||||||
|
"metadata filter"
|
||||||
|
],
|
||||||
|
"breadcrumbMatched": true,
|
||||||
|
"selectedAttempt": "FILTERED_VECTOR",
|
||||||
|
"fallbackReason": null,
|
||||||
|
"evidenceStatus": "supported",
|
||||||
|
"includedSources": [
|
||||||
|
"rag-l0-domain-entity-hint",
|
||||||
|
"rag-chunk-context-reconstruction",
|
||||||
|
"rag-l0-domain-entity-hint",
|
||||||
|
"rag-chunk-context-reconstruction"
|
||||||
|
],
|
||||||
|
"omittedSources": [],
|
||||||
|
"rerankTopSource": "rag-l0-domain-entity-hint",
|
||||||
|
"failedChecks": []
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"caseId": "chat-l0-filter-fallback",
|
||||||
|
"scenario": "chat",
|
||||||
|
"query": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
|
||||||
|
"dataShape": "lookupResult",
|
||||||
|
"hitLevel": "weak",
|
||||||
|
"passed": false,
|
||||||
|
"firstExpectedRank": null,
|
||||||
|
"topCandidates": [
|
||||||
|
"1:rag-l0-filter-decoy",
|
||||||
|
"2:rag-l0-filter-decoy"
|
||||||
|
],
|
||||||
|
"matchedKeywords": [
|
||||||
|
"low quality"
|
||||||
|
],
|
||||||
|
"breadcrumbMatched": false,
|
||||||
|
"selectedAttempt": "FILTERED_VECTOR",
|
||||||
|
"fallbackReason": null,
|
||||||
|
"evidenceStatus": "supported",
|
||||||
|
"includedSources": [
|
||||||
|
"rag-l0-filter-decoy",
|
||||||
|
"rag-l0-filter-decoy"
|
||||||
|
],
|
||||||
|
"omittedSources": [],
|
||||||
|
"rerankTopSource": "rag-l0-filter-decoy",
|
||||||
|
"failedChecks": [
|
||||||
|
"expected document not found",
|
||||||
|
"selected attempt mismatch: expected UNFILTERED_VECTOR_RETRY, got FILTERED_VECTOR",
|
||||||
|
"fallback reason mismatch: expected one of [filtered_vector_low_quality, filtered_vector_no_evidence], got <none>",
|
||||||
|
"rerank top source mismatch: expected rag-l0-filter-fallback, got rag-l0-filter-decoy",
|
||||||
|
"expected context sources missing: rag-l0-filter-fallback"
|
||||||
|
]
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
@@ -0,0 +1,33 @@
|
|||||||
|
# RAG Retrieval Baseline
|
||||||
|
|
||||||
|
Generated at: `2026-07-28T06:48:56.421877+00:00`
|
||||||
|
|
||||||
|
## Aggregate
|
||||||
|
|
||||||
|
| Metric | Value |
|
||||||
|
|---|---:|
|
||||||
|
| Cases | 7 |
|
||||||
|
| Top K | 5 |
|
||||||
|
| Passed | 6 |
|
||||||
|
| Failed | 1 |
|
||||||
|
| Pass rate | 0.8571 |
|
||||||
|
| LookupResult fixtures | 7 |
|
||||||
|
| Recall@K | 0.8571 |
|
||||||
|
| Strong hit rate | 0.8571 |
|
||||||
|
| Strong hits | 6 |
|
||||||
|
| Medium hits | 0 |
|
||||||
|
| Weak hits | 1 |
|
||||||
|
| Misses | 0 |
|
||||||
|
| Average first hit rank | 1.0 |
|
||||||
|
|
||||||
|
## Cases
|
||||||
|
|
||||||
|
| Case | Scenario | Pass | Hit | Attempt | Fallback | Evidence | First Expected Rank | Top Candidates | Failed Checks |
|
||||||
|
|---|---|---|---|---|---|---|---:|---|---|
|
||||||
|
| chat-mysql-connection-pool | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:mysql-connection-pool<br>2:mysql-connection-pool | |
|
||||||
|
| chat-diagnosis-flow | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:incident-diagnosis-flow<br>2:incident-diagnosis-flow | |
|
||||||
|
| aiops-payment-latency-alert | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:payment-service-latency<br>2:aiops-alert-scope-control<br>3:payment-service-latency<br>4:aiops-alert-scope-control | |
|
||||||
|
| aiops-prometheus-alert-scope | aiops | true | strong | FILTERED_VECTOR | | supported | 1 | 1:aiops-alert-scope-control<br>2:payment-service-latency<br>3:payment-service-latency<br>4:aiops-alert-scope-control | |
|
||||||
|
| chat-rag-chunk-context | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-chunk-context-reconstruction<br>2:rag-l0-domain-entity-hint<br>3:rag-chunk-context-reconstruction<br>4:rag-l0-domain-entity-hint | |
|
||||||
|
| chat-l0-domain-hint | chat | true | strong | FILTERED_VECTOR | | supported | 1 | 1:rag-l0-domain-entity-hint<br>2:rag-chunk-context-reconstruction<br>3:rag-l0-domain-entity-hint<br>4:rag-chunk-context-reconstruction | |
|
||||||
|
| chat-l0-filter-fallback | chat | false | weak | FILTERED_VECTOR | | supported | | 1:rag-l0-filter-decoy<br>2:rag-l0-filter-decoy | expected document not found<br>selected attempt mismatch: expected UNFILTERED_VECTOR_RETRY, got FILTERED_VECTOR<br>fallback reason mismatch: expected one of [filtered_vector_low_quality, filtered_vector_no_evidence], got <none><br>rerank top source mismatch: expected rag-l0-filter-fallback, got rag-l0-filter-decoy<br>expected context sources missing: rag-l0-filter-fallback |
|
||||||
+31
-57
@@ -1,66 +1,40 @@
|
|||||||
# SuperBizAgent 面试资料包
|
# SuperBizAgent 面试资料
|
||||||
|
|
||||||
|
**更新日期**:2026-07-24
|
||||||
|
**当前主叙事**:单 Diagnosis ReAct Agent + 确定性 Harness(ISS-014 后)
|
||||||
|
|
||||||
|
## 当前入口
|
||||||
|
|
||||||
|
| 文档 | 用途 |
|
||||||
|
|---|---|
|
||||||
|
| [architecture-evolution-deep-dive.md](architecture-evolution-deep-dive.md) | **主材料**:架构演进深读(为什么设计 / 问题 / 重构 / 业界对比 / 图与白板) |
|
||||||
|
| [issues-interview-stories.md](issues-interview-stories.md) | **故事材料**:从已解决 Issues 提炼的可讲案例(STAR / 追问 / 挂载演进) |
|
||||||
|
| [topic-evidence-attribution-and-gates.md](topic-evidence-attribution-and-gates.md) | **主题深读**:证据归因幻觉 + 质量门禁(质量辨识度主故事) |
|
||||||
|
|
||||||
|
配套现行架构事实(非面试话术):
|
||||||
|
|
||||||
|
| 文档 | 用途 |
|
||||||
|
|---|---|
|
||||||
|
| [mvp/architecture/README.md](../mvp/architecture/README.md) | 当前架构文档入口 |
|
||||||
|
| [mvp/architecture/current-mvp-architecture.md](../mvp/architecture/current-mvp-architecture.md) | 分层、主链、API、安全边界 |
|
||||||
|
| [mvp/demo/README.md](../mvp/demo/README.md) | Demo 运行与输出 |
|
||||||
|
| [mvp/issues/active/ISS-015-diagnosis-runtime-quality-and-reasoning-audit.md](../mvp/issues/active/ISS-015-diagnosis-runtime-quality-and-reasoning-audit.md) | 架构冻结后的运行质量收敛 |
|
||||||
|
|
||||||
## 一句话定位
|
## 一句话定位
|
||||||
|
|
||||||
SuperBizAgent 是一个面向企业故障诊断场景的 Agent 工程项目。它把用户问题或 AIOps 告警转换成可追踪的 Agent 执行链路,并把工具证据、模型步骤、最终答案、自评估和用户反馈统一沉淀到诊断 Trace 中。
|
SuperBizAgent 是面向故障诊断的可追踪 Agent 系统:业务侧一个 Diagnosis Agent 负责推理与写 Draft;Harness 负责预算、工具边界、证据验真、语义审查与安全发布。每次执行用 `sessionId + runId` 回放统一 Timeline。
|
||||||
|
|
||||||
## 面试重点
|
## 建议使用方式
|
||||||
|
|
||||||
- **Agent 编排**:Chat 复杂问题走 `Planner -> Executor -> Verifier`;AIOps 告警入口走 `Supervisor -> Planner / Executor`。
|
1. 先读深读文档 §0–§6 + §10,建立演进骨架。
|
||||||
- **工具证据链**:知识库、日志、指标、Prometheus 告警都通过显式工具调用进入链路,并记录到 `tool_invocation`。
|
2. 按 §15 默画白板图 A/C/D(当前主链、职责迁移、数据三层)。
|
||||||
- **可追踪诊断**:一次诊断对应一个 `sessionId`,可通过 `GET /api/diagnosis/{sessionId}/trace` 回放。
|
3. 需要事实核对时回到 `mvp/architecture/*`,不要用归档面试稿当现行口径。
|
||||||
- **质量门禁**:Chat Verifier 校验 groundedness;AIOps 规则评估检查报告完整性、payload 聚焦和证据工具覆盖。
|
4. 未完成项用 ISS-015 收尾,体现判断力而非完美叙事。
|
||||||
- **RAG 工程化**:`lookup_knowledge` 是显式 Agent Tool,底层通过 Spring AI VectorStore 主路径 + Milvus SDK fallback。
|
|
||||||
- **反馈闭环**:用户反馈 `useful` 会沉淀 `case_library`,`not_useful` 保留 bad case 信号。
|
|
||||||
|
|
||||||
## 推荐阅读顺序
|
## 归档
|
||||||
|
|
||||||
1. `mvp/architecture/interview-one-pager.md`:一页式架构图和 2-5 分钟讲解。
|
2026-07-24 之前的面试资料(多角色编排、双入口、旧 RAG/AIOps 讲解等)已移至:
|
||||||
2. `mvp/demo/ten-minute-interview-demo.md`:10 分钟现场演示脚本。
|
|
||||||
3. `interview/story-cases.md`:可复用的面试故事案例。
|
|
||||||
4. `interview/architecture.md`:面试版系统架构。
|
|
||||||
5. `interview/design-tradeoffs.md`:关键设计取舍。
|
|
||||||
6. `interview/demo-script.md`:更细的命令式演示脚本。
|
|
||||||
7. `interview/acceptance-checklist.md`:面试前验收清单。
|
|
||||||
8. RAG 专题文档:`rag-refactor-story.md`、`rag-vectorstore-interview-notes.md`、`rag-retrieval-quality-report.md`。
|
|
||||||
|
|
||||||
## 核心演示链路
|
[archive/2026-07-24-legacy/](archive/2026-07-24-legacy/)
|
||||||
|
|
||||||
### Chat 诊断
|
|
||||||
|
|
||||||
```text
|
|
||||||
POST /api/chat
|
|
||||||
-> ChatService
|
|
||||||
-> Planner -> Executor -> Verifier
|
|
||||||
-> lookup_knowledge / query_logs / query_metrics
|
|
||||||
-> diagnosis_session + agent_step + tool_invocation
|
|
||||||
-> GET /api/diagnosis/{sessionId}/trace
|
|
||||||
-> POST /api/feedback
|
|
||||||
```
|
|
||||||
|
|
||||||
### AIOps 告警诊断
|
|
||||||
|
|
||||||
```text
|
|
||||||
POST /api/ai_ops
|
|
||||||
-> AiOpsService
|
|
||||||
-> PAYLOAD_TARGETED / AUTO_DISCOVERY
|
|
||||||
-> ai_ops_supervisor
|
|
||||||
-> planner_agent / executor_agent
|
|
||||||
-> queryPrometheusAlerts + logs + metrics + lookup_knowledge
|
|
||||||
-> alert report
|
|
||||||
-> aiops_rule_evaluation
|
|
||||||
-> GET /api/diagnosis/{sessionId}/trace
|
|
||||||
```
|
|
||||||
|
|
||||||
## 当前完成度
|
|
||||||
|
|
||||||
- Chat 诊断链路:可运行、可追踪、有 Verifier。
|
|
||||||
- AIOps 告警链路:可运行、可追踪、支持 payload scope control。
|
|
||||||
- RAG 检索链路:Spring AI VectorStore 主路径、Milvus SDK fallback、L0 hint、检索评测 baseline。
|
|
||||||
- Trace API:统一返回 session、agent steps、tool invocations 和 summary。
|
|
||||||
- Demo 材料:`mvp/demo/README.md`、`mvp/demo/ten-minute-interview-demo.md`。
|
|
||||||
|
|
||||||
## 主叙事
|
|
||||||
|
|
||||||
这个项目不是简单调用大模型,而是在做一个可审计、可验证、可回归的 Agent 诊断系统。模型可以规划和推理,但每一步工具证据、最终结论、Verifier 结果和用户反馈都能被 Trace API 回放。面试时重点展示“从问题到证据到答案到验证再到反馈”的闭环。
|
|
||||||
|
|
||||||
|
仅用于历史追溯,不代表当前 runtime、API 或验收口径。说明见同目录 `_ARCHIVE_NOTE.md`。
|
||||||
|
|||||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,66 @@
|
|||||||
|
# SuperBizAgent 面试资料包
|
||||||
|
|
||||||
|
## 一句话定位
|
||||||
|
|
||||||
|
SuperBizAgent 是一个面向企业故障诊断场景的 Agent 工程项目。它把用户问题或 AIOps 告警转换成可追踪的 Agent 执行链路,并把工具证据、模型步骤、最终答案、自评估和用户反馈统一沉淀到诊断 Trace 中。
|
||||||
|
|
||||||
|
## 面试重点
|
||||||
|
|
||||||
|
- **Agent 编排**:Chat 复杂问题走 `Planner -> Executor -> Verifier`;AIOps 告警入口走 `Supervisor -> Planner / Executor`。
|
||||||
|
- **工具证据链**:知识库、日志、指标、Prometheus 告警都通过显式工具调用进入链路,并记录到 `tool_invocation`。
|
||||||
|
- **可追踪诊断**:一次诊断对应一个 `sessionId`,可通过 `GET /api/diagnosis/{sessionId}/trace` 回放。
|
||||||
|
- **质量门禁**:Chat Verifier 校验 groundedness;AIOps 规则评估检查报告完整性、payload 聚焦和证据工具覆盖。
|
||||||
|
- **RAG 工程化**:`lookup_knowledge` 是显式 Agent Tool,底层通过 Spring AI VectorStore 主路径 + Milvus SDK fallback。
|
||||||
|
- **反馈闭环**:用户反馈 `useful` 会沉淀 `case_library`,`not_useful` 保留 bad case 信号。
|
||||||
|
|
||||||
|
## 推荐阅读顺序
|
||||||
|
|
||||||
|
1. `mvp/architecture/interview-one-pager.md`:一页式架构图和 2-5 分钟讲解。
|
||||||
|
2. `mvp/demo/ten-minute-interview-demo.md`:10 分钟现场演示脚本。
|
||||||
|
3. `interview/story-cases.md`:可复用的面试故事案例。
|
||||||
|
4. `interview/architecture.md`:面试版系统架构。
|
||||||
|
5. `interview/design-tradeoffs.md`:关键设计取舍。
|
||||||
|
6. `interview/demo-script.md`:更细的命令式演示脚本。
|
||||||
|
7. `interview/acceptance-checklist.md`:面试前验收清单。
|
||||||
|
8. RAG 专题文档:`rag-refactor-story.md`、`rag-vectorstore-interview-notes.md`、`rag-retrieval-quality-report.md`。
|
||||||
|
|
||||||
|
## 核心演示链路
|
||||||
|
|
||||||
|
### Chat 诊断
|
||||||
|
|
||||||
|
```text
|
||||||
|
POST /api/chat
|
||||||
|
-> ChatService
|
||||||
|
-> Planner -> Executor -> Verifier
|
||||||
|
-> lookup_knowledge / query_logs / query_metrics
|
||||||
|
-> diagnosis_session + agent_step + tool_invocation
|
||||||
|
-> GET /api/diagnosis/{sessionId}/trace
|
||||||
|
-> POST /api/feedback
|
||||||
|
```
|
||||||
|
|
||||||
|
### AIOps 告警诊断
|
||||||
|
|
||||||
|
```text
|
||||||
|
POST /api/ai_ops
|
||||||
|
-> AiOpsService
|
||||||
|
-> PAYLOAD_TARGETED / AUTO_DISCOVERY
|
||||||
|
-> ai_ops_supervisor
|
||||||
|
-> planner_agent / executor_agent
|
||||||
|
-> queryPrometheusAlerts + logs + metrics + lookup_knowledge
|
||||||
|
-> alert report
|
||||||
|
-> aiops_rule_evaluation
|
||||||
|
-> GET /api/diagnosis/{sessionId}/trace
|
||||||
|
```
|
||||||
|
|
||||||
|
## 当前完成度
|
||||||
|
|
||||||
|
- Chat 诊断链路:可运行、可追踪、有 Verifier。
|
||||||
|
- AIOps 告警链路:可运行、可追踪、支持 payload scope control。
|
||||||
|
- RAG 检索链路:Spring AI VectorStore 主路径、Milvus SDK fallback、L0 hint、检索评测 baseline。
|
||||||
|
- Trace API:统一返回 session、agent steps、tool invocations 和 summary。
|
||||||
|
- Demo 材料:`mvp/demo/README.md`、`mvp/demo/ten-minute-interview-demo.md`。
|
||||||
|
|
||||||
|
## 主叙事
|
||||||
|
|
||||||
|
这个项目不是简单调用大模型,而是在做一个可审计、可验证、可回归的 Agent 诊断系统。模型可以规划和推理,但每一步工具证据、最终结论、Verifier 结果和用户反馈都能被 Trace API 回放。面试时重点展示“从问题到证据到答案到验证再到反馈”的闭环。
|
||||||
|
|
||||||
@@ -0,0 +1,36 @@
|
|||||||
|
# Archive Note
|
||||||
|
|
||||||
|
**归档日期**:2026-07-24
|
||||||
|
**状态**:历史面试材料,不代表当前 runtime / 架构口径
|
||||||
|
|
||||||
|
## 为何归档
|
||||||
|
|
||||||
|
本目录保存切换到「单 Diagnosis Agent + Harness」之前整理的面试资料包,内容仍以:
|
||||||
|
|
||||||
|
- Chat:`Planner -> Executor -> Verifier`(及后续五段 Gatekeeper/Composer)
|
||||||
|
- AIOps 独立入口与规则评估
|
||||||
|
- 旧 Trace / Demo 叙事
|
||||||
|
|
||||||
|
为主。现行可运行架构见:
|
||||||
|
|
||||||
|
- `mvp/architecture/`
|
||||||
|
- `interview/architecture-evolution-deep-dive.md`(当前面试深读主文档)
|
||||||
|
|
||||||
|
## 归档文件
|
||||||
|
|
||||||
|
| 文件 | 原用途 |
|
||||||
|
|---|---|
|
||||||
|
| `README.md` | 旧面试资料包入口 |
|
||||||
|
| `architecture.md` | 旧面试版系统架构 |
|
||||||
|
| `design-tradeoffs.md` | 旧设计取舍 |
|
||||||
|
| `demo-script.md` | 旧命令式演示脚本 |
|
||||||
|
| `acceptance-checklist.md` | 旧面试前验收清单 |
|
||||||
|
| `story-cases.md` | 旧故事案例 |
|
||||||
|
| `rag-*.md` | 旧 RAG 专题与验收笔记 |
|
||||||
|
| `aiops-*.md` | 旧 AIOps 讲解材料 |
|
||||||
|
|
||||||
|
## 使用边界
|
||||||
|
|
||||||
|
- 可作历史决策与旧 Demo 话术追溯。
|
||||||
|
- 不得当作当前 API、编排或验收标准。
|
||||||
|
- 若引用其中内容,须同时说明归档日期与现行替代文档。
|
||||||
@@ -0,0 +1,380 @@
|
|||||||
|
# 从 Issues 提炼的面试故事
|
||||||
|
|
||||||
|
**更新日期**:2026-07-24
|
||||||
|
**用途**:从 `mvp/issues` 已解决问题中,筛出可讲、值得讲、能举一反三的案例
|
||||||
|
**配套**:[architecture-evolution-deep-dive.md](architecture-evolution-deep-dive.md)(讲演进骨架);本文讲**具体踩坑与决策**
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 0. 怎么用
|
||||||
|
|
||||||
|
| 场景 | 用法 |
|
||||||
|
|---|---|
|
||||||
|
| 行为面试 / 项目深挖 | 选 2–3 个 ★★★ 故事,用 STAR 讲 |
|
||||||
|
| 架构追问 | 把故事挂回演进阶段(Phase2 证据 / Phase3 重构 / 工程化) |
|
||||||
|
| 避免踩坑 | ★ 仅作补充,勿当主叙事;过时方案要说「后来被什么吸收」 |
|
||||||
|
|
||||||
|
**总原则**
|
||||||
|
|
||||||
|
> 面试官要的不是 Issue 编号,而是:**现象 → 根因分层 → 你选了什么杠杆 → 如何验证 → 后来边界怎么演进**。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. 全景评分(先看这张表)
|
||||||
|
|
||||||
|
| Issue / 主题 | 面试价值 | 最适合回答的问题类型 | 一句话钩子 | 注意 |
|
||||||
|
|---|:---:|---|---|---|
|
||||||
|
| **证据归因幻觉** + Gatekeeper 链路 | ★★★ | 防幻觉 / 质量门禁 | 「工具调了,结论仍是编的」 | 旧角色名要映射到现在 Guard |
|
||||||
|
| **ISS-014** 单 Agent + Harness | ★★★ | 架构重构 / 为什么简化 | 「不是砍验证,是换装载层」 | 主故事,深读文档已覆盖 |
|
||||||
|
| **ISS-008** 窄范围越界 | ★★★ | Agent 约束 / Prompt 不够 | 「只问 CPU,却去查全家桶」 | 好举一反三 |
|
||||||
|
| **ISS-009** 负向证据 | ★★★ | 证据建模 / no-hit | 「没查到也是证据」 | 区分「无问题」vs「无数据」 |
|
||||||
|
| **ISS-001→002→004→015** 重复检索族 | ★★★ | 迭代加深 / Prompt vs 硬约束 | 「去重了仍狂调 20 次」 | 讲演进链,别只讲一次补丁 |
|
||||||
|
| **ISS-010** session/run 隔离 | ★★★ | 可观测 / 多轮正确性 | 「同会话多轮 Trace 串台」 | 和 runId 真理源强绑定 |
|
||||||
|
| **ISS-012** Token/上下文膨胀 | ★★★ | 成本 / ACI / 上下文工程 | 「工具返回把上下文撑爆」 | 接到 projection 三层数据 |
|
||||||
|
| **ISS-006 + fixtures + baseline diff** | ★★ | 工程化 / 回归 | 「改 Prompt 怎么知道没退步」 | 体现测试思维 |
|
||||||
|
| **ISS-007** 摘要失真 → 自证循环 | ★★★ | 信息通路设计 | 「证据在,摘要丢了关键句」 | 可与归因幻觉合并讲 |
|
||||||
|
| **ISS-013** SSE / 入口解耦 | ★★ | 后端工程 / 协议边界 | 「假流式 + Controller 过重」 | 偏工程,Agent 味稍淡 |
|
||||||
|
| **ISS-005** 证据状态契约 | ★★ | 契约 / 失败语义 | 「failed/no_evidence/deduped 语义乱」 | 作 001/007 的基础设施铺垫 |
|
||||||
|
| **RAG 子问题集** | ★★ | RAG 专题 | L0 降级、显式 Tool、breadcrumb | 合成一条 RAG 故事,勿逐条念 |
|
||||||
|
| **ISS-003** 总 Review | ★ | 过程 | 问题发现清单 | 不宜单独讲 |
|
||||||
|
| **ISS-015**(进行中) | ★★ | 诚实收尾 / 判断力 | 「架构冻了,策略还在收」 | 讲未完成,勿假装已完美 |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. 推荐主故事包(面试只带这 5 个就够)
|
||||||
|
|
||||||
|
### 故事包怎么组合(15 分钟项目介绍)
|
||||||
|
|
||||||
|
```text
|
||||||
|
1) 开场架构(2 min) → 深读文档 / ISS-014 结论
|
||||||
|
2) 质量核心(4 min) → 证据归因幻觉 + Gatekeeper/EvidenceGuard
|
||||||
|
3) 约束演进(3 min) → 重复检索 001→002→硬预算 / 窄范围 008
|
||||||
|
4) 工程底座(3 min) → run 隔离 010 + eval harness 006
|
||||||
|
5) 重构与未完(3 min) → ISS-014 为什么合并 + ISS-015 诚实项
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. ★★★ 故事详解(可直接练口述)
|
||||||
|
|
||||||
|
### 3.1 证据归因幻觉(最高辨识度)
|
||||||
|
|
||||||
|
| 项 | 内容 |
|
||||||
|
|---|---|
|
||||||
|
| 来源 | `executor-evidence-attribution-hallucination` + ISS-007 + design-note 自证循环 |
|
||||||
|
| 完整深读 | [topic-evidence-attribution-and-gates.md](topic-evidence-attribution-and-gates.md) |
|
||||||
|
| 阶段 | Phase 2 正确性建设 |
|
||||||
|
| 现象 | 多次 E2E:工具已调 `lookup_knowledge/logs/metrics`,Verifier 仍大面积 `LOW_CONFID`;答案里出现 OOM、Full GC 次数等**工具返回里没有的「精确事实」** |
|
||||||
|
| 错误归因(你要主动否定) | 「是不是 Verifier 太严 / 只认 RAG?」——查 `evidence_refs` 后发现并非如此 |
|
||||||
|
| 真根因 | **Executor 证据归因幻觉**:把 runbook/历史模式/模型常识写进「当前已证实事实」,未区分 direct / reference / hypothesis / missing |
|
||||||
|
| 解法演进 | ① 结构化 claim + binding(invocation/path/excerpt)② **代码 Gatekeeper** 验引用 ③ Verifier 只判可推导 ④ Composer 控表达 → 后来迁到 EvidenceGuard + SemanticGuard + Release |
|
||||||
|
| 验证 | 固定 session 表(groundedness、no_evidence 计数);eval fixture;Trace 可指出「哪条 claim 无 binding」 |
|
||||||
|
| 现行映射 | Gatekeeper → **EvidenceGuard**;Verifier → **SemanticGuard**;最终出门 → **Release** |
|
||||||
|
|
||||||
|
**STAR 口述(约 90 秒)**
|
||||||
|
|
||||||
|
> 我们诊断链路经常 LOW_CONFID。第一反应像是验证器太狠,但拉 Trace 发现工具其实调用成功了。
|
||||||
|
> 对比答案和 tool raw 后定位到:模型会把知识库里的「常见故障模式」写成「这次故障已观测事实」。
|
||||||
|
> 所以我们把「有没有这句证据」从 LLM 判断里拆出来,做成代码级引用校验;模型只负责在已验真片段上做推导。
|
||||||
|
> 这直接把问题从「提示词求稳」升级成「证据所有权与类型系统」。
|
||||||
|
> 后来重构单 Agent 时,这层语义保留了,只是从流水线角色变成了 Harness 门禁。
|
||||||
|
|
||||||
|
**追问预备**
|
||||||
|
|
||||||
|
- Q: 为什么不靠更强模型?
|
||||||
|
A: 分布上仍会混用常识与观测;确定性校验可回归、可解释。
|
||||||
|
- Q: excerpt 子串匹配会不会太死?
|
||||||
|
A: 对防伪造必须偏严;表达层再允许归纳,但不允许无根引用。
|
||||||
|
- Q: 和 RAG 引用角标有何不同?
|
||||||
|
A: 角标常是生成时装饰;我们校验的是**当次 Run 的 tool_call 所有权**。
|
||||||
|
|
||||||
|
**举一反三**
|
||||||
|
|
||||||
|
- 客服:「政策规定 7 天」≠「本单已同意退款」
|
||||||
|
- 代码 Agent:「README 说应有测试」≠「本 PR 已有测试」
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### 3.2 窄范围查询越界(ISS-008)
|
||||||
|
|
||||||
|
| 项 | 内容 |
|
||||||
|
|---|---|
|
||||||
|
| 现象 | 用户只要确认 `payment-service` 是否有 HighCPU 告警;Agent 却扩展到内存、日志、根因故事 |
|
||||||
|
| 根因 | ReAct 默认「有工具就多查」;Prompt 未区分 **observation 任务** vs **root-cause 任务**;claim 类型未收窄 |
|
||||||
|
| 解法 | 窄范围只允许 `observation` / `negative_observation`;禁止随手 root_cause;工具选择与输出 schema 双约束 |
|
||||||
|
| 验证 | narrow-highcpu 类 fixture:该 PASS 的 observation 不因「没讲根因」被打成失败 |
|
||||||
|
| 现行 | DiagnosisDraft 分析项仍强调证据绑定;范围控制在 Agent 指令 + Guard 语义审查 |
|
||||||
|
|
||||||
|
**金句**
|
||||||
|
|
||||||
|
> Agent 的能力上限往往不是「会不会查」,而是「知不知道何时停、查什么算完成」。
|
||||||
|
|
||||||
|
**举一反三**
|
||||||
|
|
||||||
|
- SQL Agent:用户要 `SELECT count` 时禁止顺手 `UPDATE`
|
||||||
|
- 调查 Agent:「只要时间线」时禁止输出处置工单
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### 3.3 负向证据(ISS-009)
|
||||||
|
|
||||||
|
| 项 | 内容 |
|
||||||
|
|---|---|
|
||||||
|
| 现象 | 工具明确 no-hit 时,模型要么不说,要么说成「已排除该根因」,且引用对不上 |
|
||||||
|
| 根因 | 只建模「命中证据」,没有一等公民的 **no_evidence 路径**(如 `$.no_evidence`) |
|
||||||
|
| 解法 | `negative_observation` + 固定 raw_path/excerpt 契约;Gatekeeper 校验「无证据」也可以是合法引用 |
|
||||||
|
| 价值 | 排障里「查过没有」改变后验;也避免虚假排除 |
|
||||||
|
|
||||||
|
**金句**
|
||||||
|
|
||||||
|
> 在证据系统里,空结果若不可引用,模型就会用语言填补真空——那就是幻觉温床。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### 3.4 重复检索演进链(ISS-001 → 002 → 004 → 015)
|
||||||
|
|
||||||
|
这是**最好的「迭代加深」故事**:同主题三次升级约束强度。
|
||||||
|
|
||||||
|
```text
|
||||||
|
ISS-001 同文档内容重复进上下文
|
||||||
|
→ session 文档去重 / Prompt「别重复」
|
||||||
|
|
||||||
|
ISS-002 去重后仍 lookup 20+ 次(换 query 刷同一域)
|
||||||
|
→ 根因:约束只打在 Planner,Executor 不知情
|
||||||
|
→ Executor 注入 map + 每域一次等 Prompt 约束
|
||||||
|
|
||||||
|
ISS-004 Prompt 仍挡不住「再确认一次」
|
||||||
|
→ 设计域级水位 / 硬限制(后被架构变更吸收)
|
||||||
|
|
||||||
|
ISS-014/015 约束归属变化
|
||||||
|
→ 不再靠 Executor 旁路状态打补丁
|
||||||
|
→ Harness 预算 + Agent 硬停止 + 重复 lookup 策略(015 进行中)
|
||||||
|
```
|
||||||
|
|
||||||
|
**面试怎么讲这条链**
|
||||||
|
|
||||||
|
> 第一阶段我们以为是重复文档,做了内容去重。
|
||||||
|
> 第二阶段发现模型换关键词继续刷,说明**去重粒度错了**,且约束注入点错了(只告诉了 Planner)。
|
||||||
|
> 第三阶段承认 Prompt 约定不是安全边界,必须**预算/次数硬停止**。
|
||||||
|
> 架构重构后,这些不再散落在 Tool 旁路,而进入统一 Harness 控制面。
|
||||||
|
> 这说明我处理 Agent 问题的习惯是:先观测 → 分层根因 → 逐步把约束从「软」推到「硬」,并在架构变了以后迁移装载点,而不是叠补丁。
|
||||||
|
|
||||||
|
**对应业界概念**:Action masking / tool budget / circuit breaker;不是调参玄学。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### 3.5 Session / Run Trace 隔离(ISS-010)
|
||||||
|
|
||||||
|
| 项 | 内容 |
|
||||||
|
|---|---|
|
||||||
|
| 现象 | 同 `sessionId` 多轮诊断时,步骤/工具/评价串台或「取最新一条」导致回放错乱 |
|
||||||
|
| 根因 | 缺少把**一次执行**定为真理源的 `runId`;查询与写入未全程 exact id |
|
||||||
|
| 解法 | `diagnosis_run`;step/invocation/trace 均挂 run;API 强制 `runId`;禁止 latest 语义 |
|
||||||
|
| 价值 | 评测、排障、面试 Demo 都依赖可复现回放 |
|
||||||
|
|
||||||
|
**金句**
|
||||||
|
|
||||||
|
> 可观测性若不能精确到一次 Run,就只是日志堆,不是诊断系统的记忆。
|
||||||
|
|
||||||
|
**举一反三**
|
||||||
|
|
||||||
|
- 工作流引擎的 `workflowId` vs `runId`
|
||||||
|
- CI 的 pipeline vs job attempt
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### 3.6 Token 与上下文膨胀(ISS-012)→ ACI / 三层数据
|
||||||
|
|
||||||
|
| 项 | 内容 |
|
||||||
|
|---|---|
|
||||||
|
| 现象 | Executor 上下文暴涨;工具结果含 debug/rerank/大段重复;成本与截断不可控 |
|
||||||
|
| 根因 | Tool 返回**面向开发者**而非 Agent;完整 raw 与模型可见视图未分离 |
|
||||||
|
| 解法方向 | Token 可观测;硬预算;结果投影;证据引用带稳定 id(后由 ISS-014 的 Redis canonical + projector 落地) |
|
||||||
|
| 现行 | canonical / projection / durable audit 三层 |
|
||||||
|
|
||||||
|
**金句**
|
||||||
|
|
||||||
|
> 上下文工程首先是接口设计问题:Agent 的观察通道必须有界,验真通道才能完整。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### 3.7 单 Agent + Harness 重构(ISS-014)
|
||||||
|
|
||||||
|
深读文档已写透,这里只留**Issue 视角的故事钩子**:
|
||||||
|
|
||||||
|
| 项 | 内容 |
|
||||||
|
|---|---|
|
||||||
|
| 触发 | 多角色 = 外层重复 ReAct;JSON 接力;Token;失败语义组合爆炸 |
|
||||||
|
| 保留 | 物理验真 → 语义审查 → 发布 的正确性模型 |
|
||||||
|
| 迁移 | 角色流水线 → 执行边界(Harness) |
|
||||||
|
| 证据 | 阶段 E2E:成功诊断 ~8k tokens / 1 次 tool;Knowledge Query 独立路径修复 invalid schema |
|
||||||
|
|
||||||
|
**和 3.1 的关系(必说清)**
|
||||||
|
|
||||||
|
> 014 不是推翻 007/归因幻觉的成果,而是避免用「五个 LLM 角色」去实现本该由一个 ReAct + 一层确定性门禁完成的事。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### 3.8 评测 Harness(ISS-006 + fixtures + baseline diff)
|
||||||
|
|
||||||
|
| 项 | 内容 |
|
||||||
|
|---|---|
|
||||||
|
| 动机 | 只有 Demo 无法判断改 Prompt/Tool 是否退步 |
|
||||||
|
| 做法 | 固定 case;expected tools/verdict/keywords;**确定性** trace 校验(非一上来 LLM-as-judge);baseline + diff |
|
||||||
|
| 价值 | 把 Agent 质量从「感觉」变成「回归」 |
|
||||||
|
|
||||||
|
**金句**
|
||||||
|
|
||||||
|
> 没有 baseline diff 的 Agent 迭代,只是在用生产用户当测试集。
|
||||||
|
|
||||||
|
**注意**:面试强调「先确定性检查,再考虑 LLM judge」,显得克制。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### 3.9 SSE 与入口解耦(ISS-013)
|
||||||
|
|
||||||
|
| 项 | 内容 |
|
||||||
|
|---|---|
|
||||||
|
| 现象 | `/api/chat` 与 `/api/chat_stream` 双入口;stream 实为整答后假分片;Controller 编排过重 |
|
||||||
|
| 解法 | 唯一 `POST /api/chat` named SSE;UseCase 拥有业务;Controller 只协议;disconnect → cancel Run |
|
||||||
|
| 适合 | 问到 Spring/API 设计、背压、职责边界时 |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. ★★ 可合并讲的「专题束」
|
||||||
|
|
||||||
|
### 4.1 RAG 专题束(不要逐 Issue 报菜名)
|
||||||
|
|
||||||
|
把 `mvp/issues/rag/*` 收成 **一条 2 分钟故事**:
|
||||||
|
|
||||||
|
```text
|
||||||
|
问题簇:
|
||||||
|
L0 关键词当终局、breadcrumb 不进向量、切片丢层级、
|
||||||
|
无 packing/rerank、分数语义不清、Advisor 隐式注入 vs 显式 Tool
|
||||||
|
|
||||||
|
收敛原则:
|
||||||
|
1) 检索决策要对 Agent 可见 → lookup_knowledge 保持 Tool
|
||||||
|
2) L0 降级为 hint,不替代语义召回
|
||||||
|
3) 向量主路径可演进(VectorStore)+ 过渡期 fallback
|
||||||
|
4) 召回质量与「能否被引用验真」一起设计
|
||||||
|
```
|
||||||
|
|
||||||
|
**面试官若只问 RAG**:用这条;若问 Agent 质量:退回 3.1。
|
||||||
|
|
||||||
|
### 4.2 证据契约束(ISS-005 + 007 + 结构化输出设计笔记)
|
||||||
|
|
||||||
|
```text
|
||||||
|
统一 evidence 状态:supported / no_evidence / deduped / failed
|
||||||
|
摘要不可当唯一证据源
|
||||||
|
Executor 产出可绑定结构,Gatekeeper 验,Verifier 判
|
||||||
|
```
|
||||||
|
|
||||||
|
适合接在「你们怎么保证工具结果语义一致」类问题。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. 不建议当主故事的
|
||||||
|
|
||||||
|
| 项 | 原因 | 若被问到怎么说 |
|
||||||
|
|---|---|---|
|
||||||
|
| ISS-003 总 Review | 清单型,缺单点冲突 | 「那是问题发现基线,具体落地看 005/006/014」 |
|
||||||
|
| ISS-004 原文方案细节 | 实现被 014/015 吸收,细节易过时 | 「方向是硬水位,装载点已迁到 Harness 预算/停止策略」 |
|
||||||
|
| 归因幻觉的旧 Prompt 补丁 alone | 不完整 | 必须接到 Gatekeeper/Guard |
|
||||||
|
| 未归档的「计划中」口吻 | 很多已 done | 统一用「已归档 / 被 014 吸收 / 015 进行中」三态 |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. 问题类型 → Issue 速查
|
||||||
|
|
||||||
|
| 面试官问… | 优先故事 |
|
||||||
|
|---|---|
|
||||||
|
| 怎么防幻觉? | 3.1 归因幻觉 + Guard 映射 |
|
||||||
|
| Prompt 够不够? | 3.4 重复检索链 + 3.2 窄范围 |
|
||||||
|
| 多 Agent 为什么又合并? | 3.7 ISS-014(挂 3.1 证明没砍质量) |
|
||||||
|
| 成本 / Token? | 3.6 + 数据三层 |
|
||||||
|
| 如何回归? | 3.8 eval |
|
||||||
|
| 如何调试一次错误诊断? | 3.5 run 隔离 + Timeline |
|
||||||
|
| 没找到证据怎么办? | 3.3 负向证据 + Release fallback |
|
||||||
|
| SSE / 接口设计? | 3.9 |
|
||||||
|
| RAG 怎么做的? | 4.1 专题束 |
|
||||||
|
| 还有什么没做完? | ISS-015:硬停止、Repair schema、reasoning 治理、信息化 fallback |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. 与架构演进的挂载图
|
||||||
|
|
||||||
|
```text
|
||||||
|
Phase1 能跑
|
||||||
|
└─ 暴露:重复检索 001/002
|
||||||
|
|
||||||
|
Phase2 能验
|
||||||
|
├─ 证据状态 005
|
||||||
|
├─ 归因幻觉 + 007 自证循环
|
||||||
|
├─ 窄范围 008 / 负向证据 009
|
||||||
|
├─ eval 006 / baseline
|
||||||
|
└─ run 隔离 010
|
||||||
|
|
||||||
|
Phase2 负债
|
||||||
|
└─ Token 012、入口 013、角色编排税
|
||||||
|
|
||||||
|
Phase3 能控
|
||||||
|
└─ 014 单 Agent + Harness(吸收 004/012/013 与证据门禁语义)
|
||||||
|
|
||||||
|
Phase4 治理中
|
||||||
|
└─ 015 停止策略 / Repair / reasoning / fallback 信息量
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. 建议你精炼的「个人贡献表述」模板
|
||||||
|
|
||||||
|
按真实参与度改主语,结构建议:
|
||||||
|
|
||||||
|
```text
|
||||||
|
我负责/主导了 ___(问题)。
|
||||||
|
通过 Trace 看到 ___(证据),排除了 ___(错误假设)。
|
||||||
|
方案上选择 ___ 而不是 ___,因为 ___。
|
||||||
|
用 ___(fixture/E2E/指标)验证。
|
||||||
|
后续在 014 重构中,该能力迁移为 ___,我学到 ___。
|
||||||
|
```
|
||||||
|
|
||||||
|
示例(归因幻觉):
|
||||||
|
|
||||||
|
```text
|
||||||
|
我负责排查 Chat 诊断大面积 LOW_CONFID。
|
||||||
|
通过对比 tool raw 与最终答案,确认是证据归因幻觉而非 Verifier 误杀。
|
||||||
|
推动「结构化 claim + 代码 Gatekeeper + Verifier 只做推导」而不是继续堆 Prompt。
|
||||||
|
用固定 E2E session 与 eval fixture 回归。
|
||||||
|
014 重构后该语义保留为 EvidenceGuard/SemanticGuard/Release。
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 9. 源文件索引
|
||||||
|
|
||||||
|
| 故事 | 路径 |
|
||||||
|
|---|---|
|
||||||
|
| 归因幻觉 | `mvp/issues/archived/executor-evidence-attribution-hallucination.md` |
|
||||||
|
| 自证/摘要 | `mvp/issues/archived/ISS-007-...` / `design-notes/executor-self-evidence-loop-design-note.md` |
|
||||||
|
| 窄范围 | `mvp/issues/archived/ISS-008-...` |
|
||||||
|
| 负向证据 | `mvp/issues/archived/ISS-009-...` |
|
||||||
|
| 重复检索 | `ISS-001` `ISS-002` `ISS-004` |
|
||||||
|
| Run 隔离 | `ISS-010` |
|
||||||
|
| Token | `ISS-012` |
|
||||||
|
| SSE | `ISS-013` |
|
||||||
|
| 重构 | `ISS-014-single-react-agent-harness-aci-ptk-refactor.md` |
|
||||||
|
| 评测 | `ISS-006` `expand-diagnosis-eval-fixtures` `diagnosis-eval-baseline-diff` |
|
||||||
|
| 进行中 | `mvp/issues/active/ISS-015-...` |
|
||||||
|
| RAG 簇 | `mvp/issues/rag/*` |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 10. 自测
|
||||||
|
|
||||||
|
1. 不看文档,讲清「工具调用成功为何仍 LOW_CONFID」的根因与门禁分层。
|
||||||
|
2. 用 001→002→004→015 说明你如何升级约束强度。
|
||||||
|
3. 画旧 Gatekeeper 到新 EvidenceGuard 的映射,并说明 014 保留了什么。
|
||||||
|
4. 举一个负向证据防止的错误用户话术。
|
||||||
|
5. 用一句话说明 eval harness 为什么先做确定性检查。
|
||||||
|
|
||||||
|
能答 1–3,项目深挖通常已够用;4–5 用于区分「做过功能」和「有质量体系」。
|
||||||
@@ -0,0 +1,739 @@
|
|||||||
|
# 主题深读:证据归因幻觉 + 现行质量门禁
|
||||||
|
|
||||||
|
**用途**:巩固知识 + 面试准备 + 举一反三
|
||||||
|
**不是**:逐字讲稿、旧 Issue 复述、过时五段流水线说明书
|
||||||
|
**材料日期**:2026-07-24
|
||||||
|
**叙事原则**:**以现行设计为主讲;早期 Issue 只说明「问题从哪来」**
|
||||||
|
**主题定位**:质量辨识度主故事——「工具调了,结论为何仍不能直接给用户」
|
||||||
|
|
||||||
|
**现行依据(面试默认口径)**
|
||||||
|
|
||||||
|
| 层级 | 路径 |
|
||||||
|
|---|---|
|
||||||
|
| 架构 | `mvp/architecture/current-mvp-architecture.md` |
|
||||||
|
| 编排 | `mvp/architecture/agent-orchestration.md` |
|
||||||
|
| 门禁 | `mvp/architecture/harness-quality-gates.md` |
|
||||||
|
| Draft 契约 | `.../harness/contract/DiagnosisDraft.java`、`AnalysisKind.java` |
|
||||||
|
| 物理验真 | `.../harness/guard/evidence/EvidenceGuard.java` |
|
||||||
|
| 语义审查 | `.../harness/guard/semantic/SemanticGuard.java`、`semantic-guard-prompt.md` |
|
||||||
|
| 发布 | `.../harness/release/DiagnosisReleaseUseCase.java`、`SafeFallbackFactory.java` |
|
||||||
|
| Agent 规则 | `src/main/resources/prompts/diagnosis-agent-prompt.md` |
|
||||||
|
|
||||||
|
**历史依据(只作起源,不代表 runtime)**
|
||||||
|
|
||||||
|
- `mvp/issues/archived/executor-evidence-attribution-hallucination.md`(2026-07-07,旧 Executor 链路)
|
||||||
|
- Phase2 Gatekeeper / `executor_evidence_v2` 归档文档
|
||||||
|
|
||||||
|
**配套**
|
||||||
|
|
||||||
|
- 演进骨架 → [architecture-evolution-deep-dive.md](architecture-evolution-deep-dive.md)
|
||||||
|
- 故事索引 → [issues-interview-stories.md](issues-interview-stories.md) §3.1
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 0. 怎么用 / 怎么讲
|
||||||
|
|
||||||
|
| 目标 | 用法 |
|
||||||
|
|---|---|
|
||||||
|
| 巩固 | 先背现行三道门与 `DiagnosisDraft` 引用闭包,再看历史问题为何逼出这套设计 |
|
||||||
|
| 面试 | **先讲现在怎么拦,再补一句早期怎么发现**;禁止把主链路讲成 Planner→Executor→Verifier |
|
||||||
|
| 举一反三 | 用现行组件名迁移到客服/代码/合规场景 |
|
||||||
|
|
||||||
|
**开场主线(先背这句,现行口径)**
|
||||||
|
|
||||||
|
> 当前系统里,Diagnosis Agent 是唯一报告作者,它在 ReAct 里查只读工具并写出结构化 `DiagnosisDraft`。
|
||||||
|
> 但 Draft **默认不能出门**:必须先过 **EvidenceGuard(0 LLM,验 tool_call 归属与投影)**,再过 **SemanticGuard(隔离单轮,判是否越证)**,最后由 **Release Policy** 决定公开报告还是 SAFE_FALLBACK。
|
||||||
|
> 这套门禁要防的核心失败,是早期线上已经见过的 **证据归因幻觉**:有工具调用,却把 runbook/常识写成「本 Run 已证实事实」。
|
||||||
|
|
||||||
|
**禁止的过时口径**
|
||||||
|
|
||||||
|
| 不要说 | 要说 |
|
||||||
|
|---|---|
|
||||||
|
| 我们主链路是 Planner→Executor→Gatekeeper→Verifier→Composer | 单 Diagnosis Agent + Harness 门禁 |
|
||||||
|
| Gatekeeper 验 `source_invocation_id + raw_path + excerpt` | EvidenceGuard 验 `tool_call_id` + 当前 Run canonical/projection |
|
||||||
|
| Verifier 输出 LOW_CONFID/PASS | SemanticGuard 输出 `SUPPORTED` / `UNSUPPORTED` |
|
||||||
|
| 工具有 metrics/Prometheus | 现行诊断 Tool:`lookup_knowledge` / `query_logs` / `query_mysql` |
|
||||||
|
| 引用主键是 DB `tool_invocation.id` | 框架 `tool_call_id`;完整调用在 Redis canonical |
|
||||||
|
|
||||||
|
**四轴(现行)**
|
||||||
|
|
||||||
|
1. **谁写报告**:仅 Diagnosis Agent
|
||||||
|
2. **谁保证引用真**:EvidenceGuard + Redis canonical(Harness only)
|
||||||
|
3. **谁保证语义不越界**:SemanticGuard(无 Tool、无记忆)
|
||||||
|
4. **谁决定用户看见什么**:Release Policy(Draft ≠ 公开 SSE)
|
||||||
|
|
||||||
|
**图目录**
|
||||||
|
|
||||||
|
| 图 | 位置 | 白板优先级 |
|
||||||
|
|---|---|---|
|
||||||
|
| 现行主链(质量视角) | §1 | ★★★ |
|
||||||
|
| DiagnosisDraft 引用闭包 | §2 | ★★★ |
|
||||||
|
| EvidenceGuard 校验步骤 | §3 | ★★★ |
|
||||||
|
| AnalysisKind × EvidenceStatus | §3.3 | ★★ |
|
||||||
|
| SemanticGuard 输入冻结 | §4 | ★★★ |
|
||||||
|
| Release / Repair / Fallback | §5 | ★★★ |
|
||||||
|
| 数据三层如何服务验真 | §6 | ★★ |
|
||||||
|
| 历史问题 → 现行映射(30 秒) | §8 | ★★ |
|
||||||
|
| 白板速画 | §13 | ★★★ |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. 现行设计:质量门禁长什么样
|
||||||
|
|
||||||
|
### 1.1 在系统中的位置
|
||||||
|
|
||||||
|
```text
|
||||||
|
POST /api/chat (SSE)
|
||||||
|
→ ChatApplicationUseCase
|
||||||
|
→ Intent Router → DIAGNOSIS
|
||||||
|
→ Diagnosis ReAct Agent
|
||||||
|
只读 Tool × 3,观察的是 projection
|
||||||
|
产出 DiagnosisDraft
|
||||||
|
→ DiagnosisReleaseUseCase
|
||||||
|
EvidenceGuard →(可选 EvidenceRepair 一次)→ SemanticGuard → Release
|
||||||
|
→ 公开 content | SAFE_FALLBACK | failure
|
||||||
|
```
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TB
|
||||||
|
Q["query + safe previous_turn"] --> A["Diagnosis Agent<br/>唯一报告作者 · ReAct"]
|
||||||
|
A <--> TB["ToolBoundary"]
|
||||||
|
TB --> Redis["Redis canonical<br/>Harness only"]
|
||||||
|
TB --> Proj["bounded agent_result"]
|
||||||
|
Proj --> A
|
||||||
|
A --> Draft["DiagnosisDraft<br/>内部制品"]
|
||||||
|
Draft --> RU["DiagnosisReleaseUseCase"]
|
||||||
|
RU --> EG["EvidenceGuard · 0 LLM"]
|
||||||
|
EG -->|invalid| RP["EvidenceRepair 最多一次"]
|
||||||
|
RP --> EG
|
||||||
|
EG -->|valid snapshot| SG["SemanticGuard · 隔离单轮"]
|
||||||
|
SG -->|SUPPORTED| Out["Release SUCCESS<br/>公开 typed report"]
|
||||||
|
SG -->|UNSUPPORTED / 不可用| FB["SAFE_FALLBACK"]
|
||||||
|
EG -->|仍 invalid| FB
|
||||||
|
```
|
||||||
|
|
||||||
|
**面试一句话**
|
||||||
|
|
||||||
|
> Agent 负责「尽量基于证据写对」;Harness 负责「写错了也不能当成功答案发出去」。
|
||||||
|
|
||||||
|
### 1.2 职责切分(现行,必背)
|
||||||
|
|
||||||
|
| 组件 | 做 | 不做 |
|
||||||
|
|---|---|---|
|
||||||
|
| **Diagnosis Agent** | 规划、调 Tool、写完整 Draft、证据不足时写 limitations | HTTP/SSE、预算、物理验真、语义终审、发布 |
|
||||||
|
| **ToolBoundary** | schema/只读/预算、写 canonical、projector、只回有界观察 | 业务推理 |
|
||||||
|
| **EvidenceGuard** | Draft 结构、analysis 闭包、`tool_call_id` 属本 Run、READY/可引用、kind↔status、投影可解析 | 调模型、改报告语义 |
|
||||||
|
| **EvidenceRepair** | 在证据失败时尝试一次结构化修复 | 无限重试、绕过验真 |
|
||||||
|
| **SemanticGuard** | 在 verified snapshot 上判断整份报告是否被支持 | Tool、记忆、Redis、改写报告、部分放行 |
|
||||||
|
| **Release** | SUPPORTED 才公开 Draft;否则固定 fallback/failure | 把未验证 Draft 流式出去 |
|
||||||
|
|
||||||
|
对应代码入口:`DiagnosisReleaseUseCase.execute(run, query, draft)`。
|
||||||
|
|
||||||
|
### 1.3 和「纯 ReAct Demo」的差(质量视角)
|
||||||
|
|
||||||
|
```text
|
||||||
|
纯 ReAct: Model ↔ Tools → 文本直接给用户
|
||||||
|
|
||||||
|
现行: Model ↔ ToolBoundary/projection → DiagnosisDraft
|
||||||
|
→ EvidenceGuard → SemanticGuard → Release → SSE
|
||||||
|
+ run 预算/取消 + Timeline(EVIDENCE/SEMANTIC/RELEASE 事件)
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. 现行契约:`DiagnosisDraft` 如何逼归因诚实
|
||||||
|
|
||||||
|
### 2.1 结构(代码事实)
|
||||||
|
|
||||||
|
```text
|
||||||
|
DiagnosisDraft
|
||||||
|
conclusion?
|
||||||
|
text
|
||||||
|
based_on_analysis_ids[] ← 必须指向本 Draft 内 analysis_id
|
||||||
|
analysis[]
|
||||||
|
analysis_id ← 唯一
|
||||||
|
kind: NORMAL | NEGATIVE_OBSERVATION
|
||||||
|
text
|
||||||
|
tool_call_ids[] ← 至少一个;必须是本 Run 真实 id
|
||||||
|
action_plan[]
|
||||||
|
action
|
||||||
|
based_on_analysis_ids[]
|
||||||
|
requires_human_confirmation
|
||||||
|
recommendations[]
|
||||||
|
text
|
||||||
|
based_on_analysis_ids[]
|
||||||
|
limitations ← 必填 scope(EvidenceGuard 校验)
|
||||||
|
scope
|
||||||
|
missing_info[]
|
||||||
|
```
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TB
|
||||||
|
C["conclusion / actions / recommendations"] -->|"based_on_analysis_ids"| A["analysis_id"]
|
||||||
|
A -->|"tool_call_ids"| T["本 Run tool_call_id"]
|
||||||
|
T --> Canon["Redis canonical<br/>runId + toolCallId"]
|
||||||
|
Canon --> Status["evidence_status<br/>FOUND / NO_EVIDENCE"]
|
||||||
|
A --> Kind["kind NORMAL / NEGATIVE"]
|
||||||
|
Kind -.->|"must match"| Status
|
||||||
|
```
|
||||||
|
|
||||||
|
**设计意图(对着归因幻觉)**
|
||||||
|
|
||||||
|
| 约束 | 防什么 |
|
||||||
|
|---|---|
|
||||||
|
| 每条 analysis 必须带 `tool_call_ids` | 「凭空结论」「常识当观测」 |
|
||||||
|
| conclusion 不直接绑 tool,只绑 analysis | 结论必须落在已声明的分析链上,形成闭包 |
|
||||||
|
| `limitations` 强制存在 | 证据不足时不能装成完整结案 |
|
||||||
|
| `kind` 二分 | 负向观察与正向命中不能混用同一类证据状态 |
|
||||||
|
|
||||||
|
### 2.2 Agent Prompt 里的硬规则(现行)
|
||||||
|
|
||||||
|
`diagnosis-agent-prompt.md` 关键口径(面试可直接引用思想,不必背原文):
|
||||||
|
|
||||||
|
1. **唯一报告作者**;内部规划,不对外输出 CoT。
|
||||||
|
2. **PreviousTurn 不是本 Run 证据**,不能引用其 tool_call_id。
|
||||||
|
3. 只有 `evidence_status=EVIDENCE_FOUND|NO_EVIDENCE` 且带真实 `tool_call_id` 的观察才能当证据。
|
||||||
|
4. **NORMAL** 只能引 FOUND;**NEGATIVE_OBSERVATION** 只能引 NO_EVIDENCE。
|
||||||
|
5. **NO_EVIDENCE ≠ 系统健康 / 已排除根因**。
|
||||||
|
6. Tool **ERROR 不是证据**,禁止引用。
|
||||||
|
7. 证据不足:`conclusion=null`,写 scope/missing_info,**禁止编根因**。
|
||||||
|
8. 不暴露 raw、凭据、内部错误、hidden reasoning。
|
||||||
|
|
||||||
|
**金句**
|
||||||
|
|
||||||
|
> Prompt 负责教 Agent「怎么写才诚实」;EvidenceGuard 负责「写得不诚实就过不了」。两者缺一不可,但**安全边界在代码**。
|
||||||
|
|
||||||
|
### 2.3 和早期「分区文案」的关系(30 秒)
|
||||||
|
|
||||||
|
早期 Issue 要求 Executor 输出「已证实 / 推测 / 缺口 / 动作」分区。
|
||||||
|
现行不是同一套 JSON 字段名,但**语义被结构吸收了**:
|
||||||
|
|
||||||
|
| 早期分区意图 | 现行落点 |
|
||||||
|
|---|---|
|
||||||
|
| 已证实事实 | `analysis`(NORMAL + FOUND) |
|
||||||
|
| 负向观察 | `analysis`(NEGATIVE_OBSERVATION + NO_EVIDENCE) |
|
||||||
|
| 证据缺口 | `limitations.missing_info` + 可空 `conclusion` |
|
||||||
|
| 建议动作 | `action_plan` / `recommendations`(必须 based_on analysis) |
|
||||||
|
| 禁止常识当事实 | Guard 不认无 tool_call 的 analysis;Semantic 审越证 |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. EvidenceGuard:现行物理验真(0 LLM)
|
||||||
|
|
||||||
|
### 3.1 它在防什么
|
||||||
|
|
||||||
|
不是「这句话好不好听」,而是:
|
||||||
|
|
||||||
|
> 报告里声明依赖的每一个 tool_call,是否真的在本 Run 发生过、状态是否可引用、投影是否自洽、analysis kind 是否与 evidence_status 匹配,以及 conclusion/action 是否只引用存在的 analysis。
|
||||||
|
|
||||||
|
### 3.2 校验流水(按代码路径讲)
|
||||||
|
|
||||||
|
`EvidenceGuard.validate(RunContext, DiagnosisDraft)` 大致两段:
|
||||||
|
|
||||||
|
**A. Draft 结构与报告闭包**
|
||||||
|
|
||||||
|
- draft / analysis 非空
|
||||||
|
- `analysis_id` 存在且唯一
|
||||||
|
- `kind`、`text` 必填
|
||||||
|
- 每条 analysis 的 `tool_call_ids` 非空
|
||||||
|
- conclusion / action_plan / recommendations:text 非空,且 `based_on_analysis_ids` 非空、id 都认识
|
||||||
|
- `limitations.scope` 必填
|
||||||
|
|
||||||
|
**B. 逐 tool_call 验真**
|
||||||
|
|
||||||
|
对每个 `tool_call_id`:
|
||||||
|
|
||||||
|
1. 用 `runId + toolCallId` 生成 key,查 **Redis canonical**
|
||||||
|
2. 找不到 → `INVOCATION_MISSING`(典型:编造 id)
|
||||||
|
3. id 不一致 → `INVOCATION_ID_MISMATCH`
|
||||||
|
4. `!isReferencableBy(runId)` 或 agent_result 空 → `INVOCATION_NOT_REFERENCABLE`
|
||||||
|
(跨 Run、未 READY、不可引用状态)
|
||||||
|
5. `analysis.kind.accepts(invocation.evidenceStatus)`
|
||||||
|
- NORMAL ↔ EVIDENCE_FOUND
|
||||||
|
- NEGATIVE_OBSERVATION ↔ NO_EVIDENCE
|
||||||
|
否则 `EVIDENCE_KIND_MISMATCH`
|
||||||
|
6. 按 tool 反序列化 **agent_result 投影**(不是让模型再读 raw 讲故事):
|
||||||
|
- `lookup_knowledge` / `query_logs` / `query_mysql`
|
||||||
|
7. 投影内 `tool_call_id`、`evidence_status` 与 canonical 一致
|
||||||
|
8. FOUND 必须有可展示证据条目;NO_EVIDENCE 必须空列表且 count=0
|
||||||
|
9. 通过则写入 `VerifiedEvidence`,汇总为 `VerifiedEvidenceSnapshot`
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TB
|
||||||
|
Draft["DiagnosisDraft"] --> S["结构 + analysis 闭包"]
|
||||||
|
S --> Loop["foreach tool_call_id"]
|
||||||
|
Loop --> Key["key = runId + toolCallId"]
|
||||||
|
Key --> Redis["Canonical store"]
|
||||||
|
Redis -->|missing| V1["INVOCATION_MISSING"]
|
||||||
|
Redis -->|not referencable| V2["NOT_REFERENCABLE"]
|
||||||
|
Redis -->|ok| K["kind vs evidence_status"]
|
||||||
|
K -->|mismatch| V3["KIND_MISMATCH"]
|
||||||
|
K --> P["Parse projection"]
|
||||||
|
P -->|invalid| V4["PROJECTION_*"]
|
||||||
|
P --> Snap["VerifiedEvidenceSnapshot"]
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3.3 `AnalysisKind` × `EvidenceStatus`(高频考点)
|
||||||
|
|
||||||
|
```text
|
||||||
|
NORMAL → 只能绑 EVIDENCE_FOUND
|
||||||
|
NEGATIVE_OBSERVATION → 只能绑 NO_EVIDENCE
|
||||||
|
ERROR → 根本不是证据(Agent Prompt + 边界)
|
||||||
|
```
|
||||||
|
|
||||||
|
**面试例子**
|
||||||
|
|
||||||
|
| Agent 想说 | 错误绑法 | Guard |
|
||||||
|
|---|---|---|
|
||||||
|
| 「CPU 告警 92%」 | 编造 tool_call_id | MISSING |
|
||||||
|
| 「未查到池耗尽日志」 | kind=NORMAL 却绑 NO_EVIDENCE | KIND_MISMATCH |
|
||||||
|
| 「已排除内存泄漏」 | 仅 NO_EVIDENCE 却写排除结论 | 物理可能过,**Semantic 应 UNSUPPORTED** |
|
||||||
|
| 引用上一轮 previous_turn 的 id | 非本 Run canonical | MISSING / NOT_REFERENCABLE |
|
||||||
|
|
||||||
|
### 3.4 为什么验的是 projection 且 canonical 在 Redis
|
||||||
|
|
||||||
|
| 设计 | 原因 |
|
||||||
|
|---|---|
|
||||||
|
| Agent 只见 projection | 有界 ACI;降低上下文里的 debug 噪声与胡拼素材 |
|
||||||
|
| Guard 读 canonical 元数据 + 校验投影 | 确认「Agent 引用的 id」对应真实调用,且投影自洽 |
|
||||||
|
| Agent 不能访问 Redis | 防止自己翻 raw 再编第二套故事 |
|
||||||
|
| MySQL `tool_invocation` 只 metadata | 长期审计 ≠ 验真主存;完整 raw 短 TTL |
|
||||||
|
|
||||||
|
**与早期 Gatekeeper 的差异(讲清楚就加分)**
|
||||||
|
|
||||||
|
| 早期 Gatekeeper | 现行 EvidenceGuard |
|
||||||
|
|---|---|
|
||||||
|
| 多角色流水线中的一环 | Harness Release 路径上的确定性步骤 |
|
||||||
|
| `source_invocation_id` + `raw_path` + `excerpt` 字符串闭合 | `tool_call_id` + Run ownership + 投影结构/状态闭合 |
|
||||||
|
| 面向 `executor_evidence_v2` claims | 面向 `DiagnosisDraft` analysis 闭包 |
|
||||||
|
| 工具集合含 metrics 等 | 现行三 Tool;投影类型 Rag/Logs/Mysql |
|
||||||
|
|
||||||
|
语义继承:**都是 0 LLM 的物理/契约验真**;协议与装载层已现代化。
|
||||||
|
|
||||||
|
### 3.5 失败码怎么用于口述
|
||||||
|
|
||||||
|
挑几个最能讲故事的 `EvidenceViolationCode`:
|
||||||
|
|
||||||
|
| Code | 一句话 |
|
||||||
|
|---|---|
|
||||||
|
| `TOOL_REFERENCE_MISSING` | analysis 根本没绑工具 |
|
||||||
|
| `INVOCATION_MISSING` | 引用了不存在的 tool_call(归因造假) |
|
||||||
|
| `INVOCATION_NOT_REFERENCABLE` | 调用存在但不可作为证据(错 Run/未就绪/空结果) |
|
||||||
|
| `EVIDENCE_KIND_MISMATCH` | 负向/正向证据用错 kind |
|
||||||
|
| `ANALYSIS_REFERENCE_UNKNOWN` | 结论引用了不存在的 analysis_id |
|
||||||
|
| `PROJECTION_INVALID` | 投影与契约不一致,不能当干净证据 |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. SemanticGuard:现行语义保险丝
|
||||||
|
|
||||||
|
### 4.1 输入被故意冻死
|
||||||
|
|
||||||
|
`SemanticGuardInput.from(query, draft, verifiedSnapshot)`:
|
||||||
|
|
||||||
|
```text
|
||||||
|
只给:
|
||||||
|
原始 query
|
||||||
|
+ 完整 Draft 视图
|
||||||
|
+ EvidenceGuard 产出的 verified snapshot
|
||||||
|
|
||||||
|
明确没有:
|
||||||
|
Tool / 记忆 / Redis / 主 Agent 回调 / 诊断历史
|
||||||
|
```
|
||||||
|
|
||||||
|
Prompt(`semantic-guard-prompt.md`)要求:
|
||||||
|
|
||||||
|
- 查:evidence 是否支持各 analysis;analysis 是否支持 conclusion;action/recommendation 是否越界;limitations 是否如实
|
||||||
|
- **禁止**改写、纠正、摘要扩展、**部分批准**
|
||||||
|
- 只返回 `{"verdict":"SUPPORTED|UNSUPPORTED","reason":"..."}`
|
||||||
|
|
||||||
|
### 4.2 它专门接住 EvidenceGuard 接不住的归因幻觉
|
||||||
|
|
||||||
|
EvidenceGuard 通过只说明:
|
||||||
|
|
||||||
|
> 「你引用的调用是真的,投影也合法。」
|
||||||
|
|
||||||
|
仍可能:
|
||||||
|
|
||||||
|
| 漏洞 | 例子 | 谁拦 |
|
||||||
|
|---|---|---|
|
||||||
|
| 真日志推不出该根因 | 只有超时日志 → 写「确定是死锁」 | SemanticGuard |
|
||||||
|
| 负向观察说成排除 | NO_EVIDENCE → 「不可能是池耗尽」 | SemanticGuard + Agent 规则 |
|
||||||
|
| 结论超出 analysis 集合语义 | analysis 只谈 A,conclusion 谈 B | SemanticGuard |
|
||||||
|
| 建议动作无分析支撑 | 乱给变更建议 | SemanticGuard + 结构上 based_on |
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart LR
|
||||||
|
EG["EvidenceGuard<br/>物理真"] --> SG["SemanticGuard<br/>语义立"]
|
||||||
|
SG -->|SUPPORTED| R["可发布"]
|
||||||
|
SG -->|UNSUPPORTED| F["Fallback<br/>保留 observed_facts"]
|
||||||
|
```
|
||||||
|
|
||||||
|
### 4.3 为什么必须隔离、且无 Tool
|
||||||
|
|
||||||
|
| 若 SemanticGuard 能再查库 | 后果 |
|
||||||
|
|---|---|
|
||||||
|
| 自建第二证据世界 | 与主 Agent / snapshot 不一致 |
|
||||||
|
| 「审稿时补证」 | 绕过用户可见的排查过程 |
|
||||||
|
| 又变成带 Tool 的第二 Executor | 归因问题换个角色重演 |
|
||||||
|
|
||||||
|
**金句**
|
||||||
|
|
||||||
|
> SemanticGuard 是保险丝,不是第二名侦探。
|
||||||
|
|
||||||
|
### 4.4 技术失败策略(现行)
|
||||||
|
|
||||||
|
- 输入/输出字节上限(`SemanticGuardLimits`)
|
||||||
|
- 同输入有限重试;仍失败 → `semanticUnavailable` fallback(有 snapshot 时仍可带已验证事实)
|
||||||
|
- 不把内部异常原文甩给用户
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Release:Draft 与公开通道切断
|
||||||
|
|
||||||
|
### 5.1 `DiagnosisReleaseUseCase` 决策序
|
||||||
|
|
||||||
|
```text
|
||||||
|
1) EvidenceGuard.validate
|
||||||
|
2) 若失败 → EvidenceRepair 一次 → 再 validate
|
||||||
|
3) 仍失败 → EVIDENCE_VALIDATION_FAILED fallback
|
||||||
|
4) SemanticGuard.review(query, draft, snapshot)
|
||||||
|
5) SUPPORTED → SUCCESS(公开 candidate Draft + snapshot 元数据路径)
|
||||||
|
6) UNSUPPORTED → SEMANTIC_UNSUPPORTED fallback(可带 observed_facts)
|
||||||
|
7) Semantic 技术不可用 → SEMANTIC_UNAVAILABLE fallback
|
||||||
|
```
|
||||||
|
|
||||||
|
Timeline 会记:`EVIDENCE_GUARD_INITIAL` / `RECHECK` / semantic / release 决策(普通 Trace 可回放阶段,不靠「感觉」)。
|
||||||
|
|
||||||
|
### 5.2 SAFE_FALLBACK 在防什么
|
||||||
|
|
||||||
|
不是空白 500,而是**可信的不完整**:
|
||||||
|
|
||||||
|
| Fallback 类型 | 用户侧含义(思想) |
|
||||||
|
|---|---|
|
||||||
|
| 证据校验失败 | 引用/结构没过,不能确认根因;可带 validation 问题方向 |
|
||||||
|
| 语义不支持 | 已有可验证事实,但撑不起当前根因结论 |
|
||||||
|
| 语义不可用 | 有事实,但审不过/审不了,暂不发根因 |
|
||||||
|
|
||||||
|
`SafeFallbackFactory` 会从 snapshot 抽取有界 `observed_facts` / sources(有上限),并给出 `failure_stage`、`next_steps` 等——**在不泄 Prompt/raw/内部 Draft 细节的前提下**尽量可操作。
|
||||||
|
|
||||||
|
**金句**
|
||||||
|
|
||||||
|
> 我们宁可发布「已经核实到什么、卡在哪」,也不发布「流畅但未过门禁的完整故事」。
|
||||||
|
|
||||||
|
### 5.3 PreviousTurn 与门禁的衔接
|
||||||
|
|
||||||
|
- 仅 **同 Session、最近一次 DIAGNOSIS + SUCCESS + published_result** 可进下一轮
|
||||||
|
- Fallback/失败/raw **不进** PreviousTurn
|
||||||
|
- Prompt 明确:previous_turn **不可当本 Run 证据**
|
||||||
|
|
||||||
|
防止「上一轮没过门禁的句子」在下一轮被当成已证实事实——这是归因幻觉的跨轮版本。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. Tool 边界:归因幻觉的上游防线
|
||||||
|
|
||||||
|
门禁是下游闸门;上游仍要减少「胡拼素材」。
|
||||||
|
|
||||||
|
```text
|
||||||
|
framework tool_call_id
|
||||||
|
→ exact Run / schema / read-only / budget
|
||||||
|
→ 执行
|
||||||
|
→ Redis canonical(request/raw/agent_result/status)
|
||||||
|
→ projector → bounded agent_result
|
||||||
|
→ 只把 projection 给 Agent
|
||||||
|
→ MySQL 仅 metadata audit
|
||||||
|
```
|
||||||
|
|
||||||
|
现行 Agent 可见 Tool 固定三个:
|
||||||
|
|
||||||
|
- `lookup_knowledge`
|
||||||
|
- `query_logs`
|
||||||
|
- `query_mysql`
|
||||||
|
|
||||||
|
**与归因的关系**
|
||||||
|
|
||||||
|
| 机制 | 作用 |
|
||||||
|
|---|---|
|
||||||
|
| 投影有界 | 少把 rerank/debug 大字段留给模型拼案情 |
|
||||||
|
| 统一 evidence_status | FOUND/NO_EVIDENCE/ERROR 语义稳定,供 kind 匹配 |
|
||||||
|
| 每调必有 tool_call_id | Draft 绑定有稳定主键 |
|
||||||
|
| 禁止 Agent 见 Redis | 不能「翻完整 raw 再假装引用」 |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. Reasoning 明确不是证据(现行安全边界)
|
||||||
|
|
||||||
|
架构硬约束:
|
||||||
|
|
||||||
|
- Provider reasoning 若存在,进独立 `agent_reasoning_audit`
|
||||||
|
- **不进**普通 SSE、Trace 正文、Evidence Snapshot、业务判断
|
||||||
|
- **不能**绕过 EvidenceGuard / SemanticGuard
|
||||||
|
- 未返回则记 unavailable,**禁止伪造**
|
||||||
|
|
||||||
|
面试若被问「你们保存思考过程吗」:
|
||||||
|
|
||||||
|
> 审计与事实分离。思考不是 tool evidence,更不能当发布依据。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. 历史问题:只用来回答「为什么要这套现行设计」
|
||||||
|
|
||||||
|
### 8.1 早期现象(30–45 秒够)
|
||||||
|
|
||||||
|
2026-07 旧 Chat 链路(Planner/Executor/Verifier…)上:
|
||||||
|
|
||||||
|
- 多次 E2E:`lookup_* / logs / metrics` 有调用,仍大量 `LOW_CONFID`
|
||||||
|
- 答案出现工具 raw 中不存在的「精确事故事实」(OOM 次数、慢 SQL 秒数等)
|
||||||
|
- 根因命名:**证据归因幻觉**——把 runbook/常识/他服事实写成「本会话已证实」
|
||||||
|
- 曾伴随摘要失真、自证闭环(自己总结再自己绑引用)
|
||||||
|
|
||||||
|
**正确用法**
|
||||||
|
|
||||||
|
> 这段证明「只靠模型自觉 + 事后 LLM Verifier」不够,必须把物理引用做成确定性约束,并把发布权从生成模型手里拿走。
|
||||||
|
|
||||||
|
**错误用法**
|
||||||
|
|
||||||
|
> 把整场面试讲成旧五段角色和 `executor_evidence_v2` 字段细节,却说不清现在的类名与 API。
|
||||||
|
|
||||||
|
### 8.2 语义迁移表(历史 → 现行)
|
||||||
|
|
||||||
|
| 历史概念 | 现行概念 | 说明 |
|
||||||
|
|---|---|---|
|
||||||
|
| Executor 综合答案 | Diagnosis Agent 写 Draft | 仍是模型生成,但是唯一作者 |
|
||||||
|
| claim + excerpt binding | analysis + `tool_call_ids` | 主键协议变更 |
|
||||||
|
| Gatekeeper | EvidenceGuard | 仍 0 LLM;装入 Release 用例 |
|
||||||
|
| Verifier LOW_CONFID | SemanticGuard UNSUPPORTED | 隔离输入;二元 verdict |
|
||||||
|
| Composer 控表达 | Draft 结构 + Release/Fallback | 表达权在 Agent,发布权在 Harness |
|
||||||
|
| 调低阈值换 PASS | **明确不做** | 用 fallback 信息量换体验 |
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart LR
|
||||||
|
H1["早期:归因幻觉被发现"] --> H2["正确性模型:物理验真+语义审+控表达"]
|
||||||
|
H2 --> H3["ISS-014:装载到单 Agent + Harness"]
|
||||||
|
H3 --> Now["现行:Draft→EG→SG→Release"]
|
||||||
|
```
|
||||||
|
|
||||||
|
### 8.3 一句话定位两阶段
|
||||||
|
|
||||||
|
> 早期 Issue 解决的是 **「要什么正确性」**;
|
||||||
|
> 现行架构解决的是 **「正确性如何成为默认运行路径」**。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 9. 和业界主流的区别(用现行组件说)
|
||||||
|
|
||||||
|
| 常见做法 | 缺口 | 本项目现行 |
|
||||||
|
|---|---|---|
|
||||||
|
| 纯 ReAct 直接吐最终答案 | 无发布闸 | Draft 默认内部,Release 才公开 |
|
||||||
|
| 答案末尾 sources 角标 | 角标可假 | `tool_call_id` 必须在本 Run canonical 可解析 |
|
||||||
|
| 单一 LLM-as-Judge | 真伪与语义混判、不可复现引用检查 | EG 代码 + SG 隔离模型 |
|
||||||
|
| 质检 Agent 再带 Tool | 第二证据世界 | SG 无 Tool,冻结 snapshot |
|
||||||
|
| 离线 RAGAS | 不挡单次错误出门 | 在线门禁 + Timeline + eval 夹具 |
|
||||||
|
| 只靠更强模型 | 无工程边界 | 与模型代际正交的 Harness |
|
||||||
|
|
||||||
|
**三个不一样(现行表述)**
|
||||||
|
|
||||||
|
1. **引用是 Run 级所有权问题**,不是文案装饰。
|
||||||
|
2. **物理与语义拆分**,失败阶段可进 Trace / fallback。
|
||||||
|
3. **公开通道与生成通道切断**,SUPPORTED 才是成功产品语义。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 10. 决策环(填的是现行答案)
|
||||||
|
|
||||||
|
| # | 问题 | 现行答案 |
|
||||||
|
|---|---|---|
|
||||||
|
| 1 | 威胁? | 带 tool 外观的完整假案情进入 SSE |
|
||||||
|
| 2 | 谁强制? | EG 强制引用;SG 强制语义;Release 强制发布 |
|
||||||
|
| 3 | 数据面? | canonical / projection / metadata audit;reasoning 另表 |
|
||||||
|
| 4 | 失败用户看到? | SAFE_FALLBACK(阶段、有界事实、下一步),非假成功 |
|
||||||
|
| 5 | 如何证明? | 单元/契约测 violation;E2E release_outcome;Timeline 事件 |
|
||||||
|
| 6 | 演进? | 误杀先查投影与 schema/Repair;不给 SG 加 Tool;停止策略见 ISS-015 |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 11. 面试题库(默认用现行答)
|
||||||
|
|
||||||
|
### 11.1 90 秒主叙述(推荐背这个版本)
|
||||||
|
|
||||||
|
> 故障诊断里最危险的不是完全不查工具,而是查了一点真实信号就补成完整事故——我们早期在旧链路上把它定义为证据归因幻觉。
|
||||||
|
> **现在**的做法是:唯一 Diagnosis Agent 用 ReAct 查三个只读工具,只看投影,输出结构化 DiagnosisDraft;每条分析必须绑定本 Run 的 tool_call_id。
|
||||||
|
> Draft 先过 EvidenceGuard:纯代码检查结构闭包、调用是否存在于当前 Run 的 Redis canonical、证据状态是否与 NORMAL/负向观察匹配、投影是否自洽。
|
||||||
|
> 通过后生成 verified snapshot,再交给无工具的 SemanticGuard 做整份语义是否越证的审查。
|
||||||
|
> 只有 SUPPORTED 才经 Release 进入 SSE;否则走 SAFE_FALLBACK,宁可告诉用户已核实事实和卡住的阶段,也不发未验证根因。
|
||||||
|
> 所以质量辨识度是三句话:**可调用 ≠ 可归因;可归因 ≠ 语义成立;语义成立才可发布。**
|
||||||
|
|
||||||
|
### 11.2 为什么题
|
||||||
|
|
||||||
|
| 问题 | 现行得分点 |
|
||||||
|
|---|---|
|
||||||
|
| 怎么防幻觉? | Draft 绑定 tool_call_id → EG → SG → Release |
|
||||||
|
| 为什么 EG 不用模型? | 归属与状态是确定性的;要可回归 |
|
||||||
|
| 为什么还要 SG? | 真调用推不出假根因;负向≠排除 |
|
||||||
|
| 为什么 SG 不能有 Tool? | 冻结证据集,防第二世界 |
|
||||||
|
| 引用主键为什么是 tool_call_id? | 框架协议 id;与当次调用一致;不靠「最新 DB 行」 |
|
||||||
|
| Agent 能看 raw 吗? | 不能;只看 projection;canonical Harness only |
|
||||||
|
| 证据不足怎么办? | conclusion 可空 + limitations;或 fallback;不编根因 |
|
||||||
|
| 和旧 Gatekeeper 啥关系? | 语义祖先;现装在 Harness,协议已换 |
|
||||||
|
|
||||||
|
### 11.3 对抗题
|
||||||
|
|
||||||
|
**Q:这不就是多 Agent 质检吗?**
|
||||||
|
|
||||||
|
A:不是。业务侧只有一个带 Tool 的 Diagnosis Agent。SG 是无 Tool 的隔离单轮审查,属于 Harness 控制面,不是协作同事。
|
||||||
|
|
||||||
|
**Q:你们重构掉多角色后质量是不是弱了?**
|
||||||
|
|
||||||
|
A:弱的是重复的 LLM 角色编排;强的是默认路径上的确定性 EG + 发布切断。正确性模型保留,装载点从流水线角色变成 Release 用例。
|
||||||
|
|
||||||
|
**Q:投影校验不看 raw 原文子串,会不会漏?**
|
||||||
|
|
||||||
|
A:现行 EG 强调 **调用所有权 + 状态 + 投影结构自洽 + kind 匹配**,再交给 SG 做语义。上游靠 projector 把可引用证据做成稳定结构。若追问 excerpt 级闭合,可承认协议从早期 raw_path/excerpt 演进到投影契约,并强调 **不能引用 ERROR/跨 Run/不可引用调用** 仍是硬的。
|
||||||
|
|
||||||
|
**Q:用户体验会不会总是 fallback?**
|
||||||
|
|
||||||
|
A:体验做在「信息化 fallback + 成功路径的 limitations」,不是放宽 Guard。ISS-015 继续收敛停止策略与 fallback 信息量。
|
||||||
|
|
||||||
|
### 11.4 现场设计题
|
||||||
|
|
||||||
|
1. 给「工单退款 Agent」设计等价于 `tool_call_ids` + EG 的字段。
|
||||||
|
2. 若增加第四个 Tool,EG 要补哪些分支?kind/status 如何扩展?
|
||||||
|
3. Knowledge Query 路径(非完整 DIAGNOSIS)如何复用「引用必须真实」而不照搬整份 SemanticGuard?
|
||||||
|
4. 如何用 Timeline 事件向面试官演示一次 UNSUPPORTED 的失败阶段?
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 12. 原则清单(现行)
|
||||||
|
|
||||||
|
1. **唯一报告作者,多个确定性关卡**
|
||||||
|
2. **Draft 是 staging,SSE 是 production**
|
||||||
|
3. **引用主键 = 本 Run 的 framework tool_call_id**
|
||||||
|
4. **NORMAL / NEGATIVE 与 FOUND / NO_EVIDENCE 强匹配**
|
||||||
|
5. **NO_EVIDENCE 不是健康证明**
|
||||||
|
6. **ERROR 与跨 Run id 绝不能当证据**
|
||||||
|
7. **PreviousTurn 不是证据**
|
||||||
|
8. **SemanticGuard 冻结 snapshot,无 Tool**
|
||||||
|
9. **Reasoning 不是证据**
|
||||||
|
10. **历史 Issue 论证问题,现行代码定义答案**
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 13. 白板默画(只画现行)
|
||||||
|
|
||||||
|
### 13.1 图 A · 60 秒主链
|
||||||
|
|
||||||
|
```text
|
||||||
|
Diagnosis Agent → DiagnosisDraft
|
||||||
|
↓
|
||||||
|
EvidenceGuard (0 LLM, tool_call_id)
|
||||||
|
↓ verified snapshot
|
||||||
|
SemanticGuard (no tools)
|
||||||
|
↓ SUPPORTED?
|
||||||
|
Release → SSE or SAFE_FALLBACK
|
||||||
|
```
|
||||||
|
|
||||||
|
### 13.2 图 B · 45 秒闭包
|
||||||
|
|
||||||
|
```text
|
||||||
|
conclusion.based_on → analysis_id → tool_call_ids
|
||||||
|
↓
|
||||||
|
Redis canonical (this runId)
|
||||||
|
↓
|
||||||
|
FOUND / NO_EVIDENCE
|
||||||
|
↓
|
||||||
|
kind must match
|
||||||
|
```
|
||||||
|
|
||||||
|
### 13.3 图 C · 30 秒历史锚点(可选)
|
||||||
|
|
||||||
|
```text
|
||||||
|
早期发现:有 tool 仍假案情
|
||||||
|
→ 要物理验真 + 语义审 + 发布权
|
||||||
|
→ 现装在 EG / SG / Release
|
||||||
|
```
|
||||||
|
|
||||||
|
### 13.4 红线
|
||||||
|
|
||||||
|
- 画出 Planner/Executor/Composer 当主路径
|
||||||
|
- 说 Verifier 输出 LOW_CONFID 当现行 API
|
||||||
|
- SG 带检索箭头
|
||||||
|
- Draft 直连用户
|
||||||
|
- 说 metrics Tool 仍是诊断三件套之一(现行是 knowledge/logs/mysql)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 14. 复习路径(偏现行)
|
||||||
|
|
||||||
|
| 步骤 | 动作 |
|
||||||
|
|---|---|
|
||||||
|
| 1 | 读 `diagnosis-agent-prompt.md` + `DiagnosisDraft` / `AnalysisKind` |
|
||||||
|
| 2 | 通读 `EvidenceGuard.validate` 与 `EvidenceViolationCode` |
|
||||||
|
| 3 | 读 `DiagnosisReleaseUseCase` + `semantic-guard-prompt.md` |
|
||||||
|
| 4 | 对照 `harness-quality-gates.md` 默画 §13 图 A/B |
|
||||||
|
| 5 | 用 §11.1 录音;再花 20 秒提早期归因幻觉作动机 |
|
||||||
|
| 6 | 扫一眼归档 Issue 标题与现象表即可,不背旧字段 |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 15. 源文件索引
|
||||||
|
|
||||||
|
### 现行(主)
|
||||||
|
|
||||||
|
| 内容 | 路径 |
|
||||||
|
|---|---|
|
||||||
|
| 架构总览 | `mvp/architecture/current-mvp-architecture.md` |
|
||||||
|
| 执行序列 | `mvp/architecture/agent-orchestration.md` |
|
||||||
|
| 门禁 | `mvp/architecture/harness-quality-gates.md` |
|
||||||
|
| Draft | `src/main/java/.../harness/contract/DiagnosisDraft.java` |
|
||||||
|
| Kind/Status | `AnalysisKind.java` / `EvidenceStatus.java` |
|
||||||
|
| EG | `.../guard/evidence/EvidenceGuard.java` |
|
||||||
|
| SG | `.../guard/semantic/SemanticGuard.java` |
|
||||||
|
| Release | `.../release/DiagnosisReleaseUseCase.java` |
|
||||||
|
| Fallback | `.../release/SafeFallbackFactory.java` |
|
||||||
|
| Agent Prompt | `src/main/resources/prompts/diagnosis-agent-prompt.md` |
|
||||||
|
| SG Prompt | `src/main/resources/prompts/semantic-guard-prompt.md` |
|
||||||
|
|
||||||
|
### 历史(辅)
|
||||||
|
|
||||||
|
| 内容 | 路径 |
|
||||||
|
|---|---|
|
||||||
|
| 归因幻觉发现 | `mvp/issues/archived/executor-evidence-attribution-hallucination.md` |
|
||||||
|
| 自证闭环笔记 | `mvp/issues/design-notes/executor-self-evidence-loop-design-note.md` |
|
||||||
|
| 旧证据契约 | `mvp/architecture/archive/2026-07-22-legacy/executor-evidence-pipeline-refactor.md` |
|
||||||
|
| 重构承接 | `mvp/issues/archived/ISS-014-...` |
|
||||||
|
| 运行质量后续 | `mvp/issues/active/ISS-015-...` |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 16. 自测(必须能用现行组件名回答)
|
||||||
|
|
||||||
|
1. 画出 Draft 从产生到 SSE 的完整门禁序,并标出哪步 0 LLM。
|
||||||
|
2. `NORMAL` 与 `NEGATIVE_OBSERVATION` 分别能绑哪种 `evidence_status`?
|
||||||
|
3. EvidenceGuard 如何发现「编造 tool_call_id」?
|
||||||
|
4. 为什么 conclusion 要 `based_on_analysis_ids` 而不是直接绑 tool?
|
||||||
|
5. SemanticGuard 的输入有哪三样?为什么不能有 Tool?
|
||||||
|
6. Evidence 失败时 Repair 最多几次?仍失败用户看到什么产品语义?
|
||||||
|
7. PreviousTurn 为什么不能提供可引用的 tool_call_id?
|
||||||
|
8. Reasoning 能否帮助 Draft 过 EG/SG?
|
||||||
|
9. 用 20 秒说明早期归因幻觉与现行三道门的关系(动机 vs 实现)。
|
||||||
|
10. 举一个「EG 通过但 SG 应 UNSUPPORTED」的例子。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 17. 和旧版材料的关系
|
||||||
|
|
||||||
|
若你曾按「五段流水线 + excerpt 外键」准备:
|
||||||
|
|
||||||
|
- **保留**:归因幻觉定义、物理/语义拆分、发布切断思想
|
||||||
|
- **替换**:所有主路径类名、Tool 列表、verdict 枚举、引用主键、API
|
||||||
|
- **降级**:Executor/Gatekeeper/Composer 仅出现在「历史动机」小节
|
||||||
|
|
||||||
|
**面试默认叠词顺序**
|
||||||
|
|
||||||
|
```text
|
||||||
|
1. 现行:Agent → Draft → EG → SG → Release
|
||||||
|
2. 机制:tool_call_id、kind/status、snapshot、fallback
|
||||||
|
3. 动机:早期归因幻觉(可选一句)
|
||||||
|
4. 演进:正确性模型保留,装载进 Harness(若追问重构)
|
||||||
|
```
|
||||||
+18
-93
@@ -1,126 +1,51 @@
|
|||||||
# SuperBizAgent MVP 文档
|
# SuperBizAgent MVP 文档
|
||||||
|
|
||||||
**更新日期**:2026-07-10
|
**更新日期**:2026-07-29
|
||||||
|
|
||||||
本目录保存 MVP 阶段的架构、问题、演示、评测和数据表说明。当前材料按“当前入口”和“历史归档”拆开,避免把早期设计稿当成当前实现。
|
本目录保存 MVP 阶段的架构、工程纪要、问题、演示、评测和数据表说明。当前材料按“当前入口”和“历史归档”拆开,避免把早期设计稿当成当前实现。
|
||||||
|
|
||||||
## 当前入口
|
## 当前入口
|
||||||
|
|
||||||
| 目录/文档 | 用途 |
|
| 目录/文档 | 用途 |
|
||||||
|---|---|
|
|---|---|
|
||||||
| [architecture/README.md](architecture/README.md) | 当前 MVP 架构入口 |
|
| [architecture/README.md](architecture/README.md) | **现行架构**(系统现在怎么跑) |
|
||||||
|
| [engineering/README.md](engineering/README.md) | **工程纪要**(问题 / 决策 / 思路 / E2E 导读) |
|
||||||
| [architecture/current-mvp-architecture.md](architecture/current-mvp-architecture.md) | 当前可运行系统架构 |
|
| [architecture/current-mvp-architecture.md](architecture/current-mvp-architecture.md) | 当前可运行系统架构 |
|
||||||
| [architecture/interview-one-pager.md](architecture/interview-one-pager.md) | 面试一页式架构讲解 |
|
|
||||||
| [architecture/agent-orchestration.md](architecture/agent-orchestration.md) | Agent 编排架构 |
|
| [architecture/agent-orchestration.md](architecture/agent-orchestration.md) | Agent 编排架构 |
|
||||||
| [architecture/executor-evidence-pipeline-refactor.md](architecture/executor-evidence-pipeline-refactor.md) | Executor 证据链路改造记录 |
|
|
||||||
| [architecture/harness-quality-gates.md](architecture/harness-quality-gates.md) | Harness 与质量门禁 |
|
| [architecture/harness-quality-gates.md](architecture/harness-quality-gates.md) | Harness 与质量门禁 |
|
||||||
| [architecture/rag-architecture.md](architecture/rag-architecture.md) | RAG/知识检索新架构 |
|
|
||||||
| [architecture/retrieval-observability.md](architecture/retrieval-observability.md) | 检索与可观测性架构 |
|
|
||||||
| [architecture/feedback-architecture.md](architecture/feedback-architecture.md) | 反馈与自评估架构 |
|
|
||||||
| [architecture/session-trace-lifecycle.md](architecture/session-trace-lifecycle.md) | 会话与 Trace 生命周期 |
|
| [architecture/session-trace-lifecycle.md](architecture/session-trace-lifecycle.md) | 会话与 Trace 生命周期 |
|
||||||
| [architecture/knowledge-base-authoring.md](architecture/knowledge-base-authoring.md) | 知识库文档编写与维护 |
|
| [architecture/RAG知识检索架构.md](architecture/RAG知识检索架构.md) | 当前 hybrid 检索架构 |
|
||||||
| [architecture/data-model.md](architecture/data-model.md) | 数据模型总览 |
|
|
||||||
| [architecture/evolution-roadmap.md](architecture/evolution-roadmap.md) | Agent 架构演进路线 |
|
|
||||||
| [issues/README.md](issues/README.md) | MVP issue 索引 |
|
| [issues/README.md](issues/README.md) | MVP issue 索引 |
|
||||||
| [issues/active/rag-refactor-plan.md](issues/active/rag-refactor-plan.md) | RAG 重构计划和阶段拆解 |
|
|
||||||
| [tables/README.md](tables/README.md) | 当前 MySQL 表说明 |
|
| [tables/README.md](tables/README.md) | 当前 MySQL 表说明 |
|
||||||
| [demo/README.md](demo/README.md) | Demo 运行和面试演示材料 |
|
| [demo/README.md](demo/README.md) | Demo 运行和演示材料 |
|
||||||
| [demo/ten-minute-interview-demo.md](demo/ten-minute-interview-demo.md) | 10 分钟面试演示脚本 |
|
|
||||||
| [eval/README.md](eval/README.md) | 诊断评测材料 |
|
| [eval/README.md](eval/README.md) | 诊断评测材料 |
|
||||||
|
|
||||||
## 当前系统一句话
|
**读法**:改行为先看 `architecture/`;要理解「为什么这样定、踩过什么坑」再看 `engineering/`。
|
||||||
|
早期个人学习笔记仍在仓库根目录 `docs/learning/` 等,**可能过时**,不以之为现行口径。
|
||||||
SuperBizAgent MVP 是一个可追踪的故障诊断 Agent:Chat 和 AIOps 入口进入 Agent 编排,Executor 显式调用知识库、日志、指标等工具收集证据;多轮会话元数据落到 `chat_session`,每次诊断运行落到 `diagnosis_run`,步骤和工具明细通过 `agent_step.run_id`、`tool_invocation.run_id` 关联,最终通过 Trace API、Verifier 和评测脚本证明结果可解释、可回放、可对比。
|
|
||||||
|
|
||||||
## 文档结构
|
## 文档结构
|
||||||
|
|
||||||
```text
|
```text
|
||||||
mvp/
|
mvp/
|
||||||
architecture/
|
architecture/ # 现行架构规范
|
||||||
|
engineering/ # 工程纪要(问题/决策/E2E)
|
||||||
README.md
|
README.md
|
||||||
current-mvp-architecture.md
|
|
||||||
interview-one-pager.md
|
|
||||||
agent-orchestration.md
|
|
||||||
executor-evidence-pipeline-refactor.md
|
|
||||||
harness-quality-gates.md
|
|
||||||
rag-architecture.md
|
|
||||||
retrieval-observability.md
|
|
||||||
feedback-architecture.md
|
|
||||||
session-trace-lifecycle.md
|
|
||||||
knowledge-base-authoring.md
|
|
||||||
data-model.md
|
|
||||||
evolution-roadmap.md
|
|
||||||
archive/
|
|
||||||
issues/
|
|
||||||
README.md
|
|
||||||
active/
|
|
||||||
archived/
|
|
||||||
design-notes/
|
|
||||||
rag/
|
rag/
|
||||||
|
diagnosis/
|
||||||
|
issues/
|
||||||
tables/
|
tables/
|
||||||
README.md
|
|
||||||
*表-*.md
|
|
||||||
archive/
|
|
||||||
demo/
|
demo/
|
||||||
README.md
|
|
||||||
ten-minute-interview-demo.md
|
|
||||||
requests/
|
|
||||||
scripts/
|
|
||||||
output/
|
|
||||||
eval/
|
eval/
|
||||||
README.md
|
|
||||||
schema.md
|
|
||||||
cases/
|
|
||||||
fixtures/
|
|
||||||
reports/
|
|
||||||
archive/
|
archive/
|
||||||
```
|
```
|
||||||
|
|
||||||
## 当前核心设计
|
|
||||||
|
|
||||||
- `lookup_knowledge` 保持显式 Agent Tool,不隐藏到 Chat Advisor。
|
|
||||||
- L0 降级为 domain/entity hint,不再默认承担最终召回决策。
|
|
||||||
- `VectorSearchService` 是检索稳定门面。
|
|
||||||
- Spring AI VectorStore 是当前读取主路径,Milvus SDK 保留为 fallback。
|
|
||||||
- AIOps payload 会生成推荐知识库 query,保留业务语义。
|
|
||||||
- `sessionId` 表示多轮会话上下文,`runId` 表示一次可回放诊断运行。
|
|
||||||
- Trace API 聚合 `diagnosis_run`、`agent_step.run_id`、`tool_invocation.run_id` 和 self evaluation。
|
|
||||||
- RAG 行为通过 offline baseline 和 live acceptance 脚本做回归验证。
|
|
||||||
|
|
||||||
## 关键运行链路
|
|
||||||
|
|
||||||
```text
|
|
||||||
Chat
|
|
||||||
-> ChatService
|
|
||||||
-> Planner / Executor / Verifier
|
|
||||||
-> evidence tools
|
|
||||||
-> chat_session / diagnosis_run
|
|
||||||
-> agent_step.run_id / tool_invocation.run_id
|
|
||||||
-> DiagnosisTraceService
|
|
||||||
|
|
||||||
AIOps
|
|
||||||
-> AiOpsService
|
|
||||||
-> PAYLOAD_TARGETED or AUTO_DISCOVERY
|
|
||||||
-> Planner / Executor
|
|
||||||
-> Prometheus / logs / lookup_knowledge
|
|
||||||
-> AiOpsRuleEvaluationService
|
|
||||||
-> diagnosis_run(agent_flow=AI_OPS)
|
|
||||||
-> DiagnosisTraceService
|
|
||||||
|
|
||||||
RAG
|
|
||||||
-> lookup_knowledge
|
|
||||||
-> L0 domain/entity hint
|
|
||||||
-> VectorSearchService
|
|
||||||
-> Spring AI VectorStore / Milvus SDK fallback
|
|
||||||
-> relevance normalization
|
|
||||||
-> tool_invocation
|
|
||||||
```
|
|
||||||
|
|
||||||
## 归档说明
|
## 归档说明
|
||||||
|
|
||||||
历史材料分两类:
|
历史材料按归档批次保存:
|
||||||
|
|
||||||
- 旧架构文档:[architecture/archive/2026-07-05-legacy/](architecture/archive/2026-07-05-legacy/)
|
- [architecture/archive/2026-07-05-legacy/](architecture/archive/2026-07-05-legacy/):早期架构设计、实现计划和知识检索方案。
|
||||||
- 本次文档清理归档:[archive/2026-07-09-doc-cleanup/](archive/2026-07-09-doc-cleanup/)
|
- [architecture/archive/2026-07-22-legacy/](architecture/archive/2026-07-22-legacy/):单 Diagnosis Agent + Harness 切换前的多角色编排、双入口和旧证据链架构。
|
||||||
|
- [archive/2026-07-09-doc-cleanup/](archive/2026-07-09-doc-cleanup/):文档清理时迁移的历史材料。
|
||||||
|
- [issues/archived/](issues/archived/):已关闭或已被当前架构替代的 Issue。
|
||||||
|
|
||||||
归档文档只用于追溯设计历史。当前实现和后续规划以 `architecture/`、`issues/README.md`、`tables/README.md` 和 OpenSpec/devflow 的最新记录为准。
|
归档内容仅用于追溯历史决策,不代表当前 runtime、API、数据模型或验收口径。
|
||||||
|
|||||||
@@ -0,0 +1,394 @@
|
|||||||
|
# RAG 检索可观测性、审计与 Trace(现行)
|
||||||
|
|
||||||
|
**更新日期**:2026-07-28
|
||||||
|
**状态**:当前可运行
|
||||||
|
**关联**:`lookup_knowledge`、Harness `ToolBoundary`、`tool_invocation`、`DiagnosisTraceService`、离线 eval
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. 三层边界
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TB
|
||||||
|
subgraph A["A. 请求内 Trace"]
|
||||||
|
LR[LookupResult<br/>retrievalTrace / rerankTrace / relevanceLevel]
|
||||||
|
end
|
||||||
|
|
||||||
|
subgraph B["B. 持久化审计 + Trace API"]
|
||||||
|
TI[tool_invocation 表]
|
||||||
|
DT[diagnosis_trace 事件摘要]
|
||||||
|
API["GET /api/diagnosis/{sessionId}/trace"]
|
||||||
|
end
|
||||||
|
|
||||||
|
subgraph C["C. 质量回归"]
|
||||||
|
EV[eval/rag-retrieval offline baseline]
|
||||||
|
end
|
||||||
|
|
||||||
|
subgraph agent["Agent 可见(非审计)"]
|
||||||
|
RT[RagToolResult<br/>evidence + optional relevance_level]
|
||||||
|
end
|
||||||
|
|
||||||
|
LK[LookupKnowledgeTool] --> LR
|
||||||
|
LR --> PROJ[RagResultProjector]
|
||||||
|
PROJ --> RT
|
||||||
|
LR --> BOUND[ToolBoundary audit]
|
||||||
|
BOUND --> TI
|
||||||
|
BOUND --> DT
|
||||||
|
TI --> API
|
||||||
|
DT --> API
|
||||||
|
EV -.->|不替代运行时 Trace| LK
|
||||||
|
```
|
||||||
|
|
||||||
|
| 层 | 完善度 | 说明 |
|
||||||
|
|----|--------|------|
|
||||||
|
| A 请求内 | 高 | attempt / fallback / quality 齐全 |
|
||||||
|
| B 持久化 + Trace API | 中高 | RAG 富字段入 `tool_invocation`,经 Trace API 回放 |
|
||||||
|
| C 离线 eval | 高 | hybrid fixtures 回归 |
|
||||||
|
|
||||||
|
**Agent 看到的不是完整 Trace。** 完整检索轨迹在 A/B;Agent 只拿投影后的证据契约。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. 端到端:从 lookup 到 Trace API
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
sequenceDiagram
|
||||||
|
participant Agent
|
||||||
|
participant Adapter as RagToolAdapter
|
||||||
|
participant Bound as ToolBoundary
|
||||||
|
participant Tool as LookupKnowledgeTool
|
||||||
|
participant Sink as JpaToolInvocationAuditSink
|
||||||
|
participant DB as tool_invocation
|
||||||
|
participant Trace as DiagnosisTraceService
|
||||||
|
participant API as GET .../trace
|
||||||
|
|
||||||
|
Agent->>Adapter: lookup_knowledge(query)
|
||||||
|
Adapter->>Bound: execute(legacy, projector)
|
||||||
|
Bound->>Tool: execute(query)
|
||||||
|
Tool-->>Bound: raw LookupResult JSON
|
||||||
|
Note over Tool: 内含 retrievalTrace / rerankTrace / evidenceBlocks
|
||||||
|
Bound->>Bound: project → RagToolResult
|
||||||
|
Bound->>Sink: AuditEvent + rawResultJson + agentResultJson
|
||||||
|
Sink->>Sink: RagLookupAuditEnricher
|
||||||
|
Sink->>DB: 富字段行
|
||||||
|
Bound-->>Agent: 投影后 agent_result(无完整 trace)
|
||||||
|
|
||||||
|
API->>Trace: sessionId + optional runId
|
||||||
|
Trace->>DB: find tool_invocation by run/session
|
||||||
|
Trace-->>API: DiagnosisTraceResponse.toolInvocations[]
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. A 层:请求内 Trace(`LookupResult`)
|
||||||
|
|
||||||
|
一次成功的 `lookup_knowledge` 内部出口是 **`LookupResult`**(比 Agent 契约更富)。
|
||||||
|
|
||||||
|
### 3.1 结构总览
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TB
|
||||||
|
LR[LookupResult]
|
||||||
|
LR --> F[found]
|
||||||
|
LR --> EB[evidenceBlocks[]]
|
||||||
|
LR --> CP[contextPack]
|
||||||
|
LR --> RT[retrievalTrace]
|
||||||
|
LR --> RR[rerankTrace]
|
||||||
|
LR --> RL[relevanceLevel]
|
||||||
|
LR --> CH[completenessHint]
|
||||||
|
LR --> CNT[evidenceCandidateCount / evidenceBlockCount]
|
||||||
|
|
||||||
|
RT --> ATT[attempts[]]
|
||||||
|
RT --> SEL[selectedAttempt]
|
||||||
|
RT --> FB[fallbackReason]
|
||||||
|
RT --> HINT[queryHints L0]
|
||||||
|
```
|
||||||
|
|
||||||
|
| 字段 | 含义 |
|
||||||
|
|------|------|
|
||||||
|
| `found` | 是否有可用证据块 |
|
||||||
|
| `evidenceBlocks` | 后处理后的 chunk 级证据(含 evidenceKey、source、content…) |
|
||||||
|
| `contextPack` | 字符预算打包文本(内部/审计用) |
|
||||||
|
| `retrievalTrace` | **检索路径 Trace**(见下) |
|
||||||
|
| `rerankTrace` | 后处理排序/quality 痕迹(现多为保序后的 quality) |
|
||||||
|
| `relevanceLevel` | PRECISE / REFERENCE / null |
|
||||||
|
| `completenessHint` | 给模型的天花板提示文案 |
|
||||||
|
|
||||||
|
### 3.2 `retrievalTrace`(检索路径)
|
||||||
|
|
||||||
|
| 字段 | 含义 |
|
||||||
|
|------|------|
|
||||||
|
| `originalQuery` | 原始查询 |
|
||||||
|
| `rewrittenQuery` | L0/变换后用于检索的 query |
|
||||||
|
| `categoryFilter` | 首次过滤的 category(可 null) |
|
||||||
|
| `selectedAttempt` | 最终采用的 attempt 名 |
|
||||||
|
| `fallbackReason` | 如 `filtered_vector_low_quality`;未降级为 null |
|
||||||
|
| `evidenceStatus` | 内部:`supported` / `no_evidence` 等 |
|
||||||
|
| `queryHints` | L0:domains、keywords、entities、l0_match_count… |
|
||||||
|
| `attempts[]` | 每次检索尝试快照 |
|
||||||
|
|
||||||
|
**常见 `selectedAttempt`:**
|
||||||
|
|
||||||
|
| 值 | 含义 |
|
||||||
|
|----|------|
|
||||||
|
| `FILTERED_VECTOR` | 带 category 的首次检索即采用 |
|
||||||
|
| `UNFILTERED_VECTOR` | 无 category,直接全库检索 |
|
||||||
|
| `UNFILTERED_VECTOR_RETRY` | filtered 低质/无证据后去掉 category 重试 |
|
||||||
|
|
||||||
|
**单次 `attempts[]` 元素:**
|
||||||
|
|
||||||
|
| 字段 | 含义 |
|
||||||
|
|------|------|
|
||||||
|
| `name` | attempt 名 |
|
||||||
|
| `query` | 该次实际检索句 |
|
||||||
|
| `categoryFilter` | 该次 filter |
|
||||||
|
| `candidateCount` | 召回候选数 |
|
||||||
|
| `usable` | 后处理阈值后是否可用 |
|
||||||
|
| `topScore` / `topSimilarity` | 引擎分 / 归一化 quality(0~1) |
|
||||||
|
| `durationMs` | 耗时 |
|
||||||
|
| `errorMessage` | 失败时 |
|
||||||
|
|
||||||
|
### 3.3 一次典型路径(含 filter fallback)
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TB
|
||||||
|
Q[query] --> L0[L0 hint → 可选 categoryFilter]
|
||||||
|
L0 --> A1[attempt FILTERED_VECTOR]
|
||||||
|
A1 --> PQ{isLowQuality?}
|
||||||
|
PQ -->|否| USE1[selectedAttempt = FILTERED_VECTOR]
|
||||||
|
PQ -->|是| A2[attempt UNFILTERED_VECTOR_RETRY]
|
||||||
|
A2 --> USE2[selectedAttempt = RETRY<br/>fallbackReason = low_quality / no_evidence]
|
||||||
|
USE1 --> POST[PostProcess · evidenceBlocks · relevanceLevel]
|
||||||
|
USE2 --> POST
|
||||||
|
POST --> LR[LookupResult 完整 Trace]
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3.4 与 Agent 投影的关系
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart LR
|
||||||
|
LR[LookupResult 全量 Trace] --> PROJ[RagResultProjector]
|
||||||
|
PROJ --> AG[RagToolResult]
|
||||||
|
AG --> F1[evidence_status]
|
||||||
|
AG --> F2[evidence excerpt]
|
||||||
|
AG --> F3[relevance_level 可选]
|
||||||
|
AG --> F4[truncated / returned_count]
|
||||||
|
|
||||||
|
LR -.->|不投影| X1[retrievalTrace]
|
||||||
|
LR -.->|不投影| X2[rerankTrace]
|
||||||
|
LR -.->|不投影| X3[raw scores / contextPack 全文]
|
||||||
|
```
|
||||||
|
|
||||||
|
人/系统要「为什么这样检索」→ 看 **A 全量** 或 **B 落库摘要**,不要只看 Agent 字段。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. B 层:持久化 + Trace API
|
||||||
|
|
||||||
|
### 4.1 写入路径
|
||||||
|
|
||||||
|
| 组件 | 职责 |
|
||||||
|
|------|------|
|
||||||
|
| `ToolBoundary` | 执行后发 `ToolInvocationAuditEvent`(含 raw LookupResult JSON + agent JSON) |
|
||||||
|
| `RagLookupAuditEnricher` | 从 LookupResult 抽有界 RAG 字段 |
|
||||||
|
| `JpaToolInvocationAuditSink` | 写入 `tool_invocation` |
|
||||||
|
| `TraceAuditEvents.toolInvocation` | 另写一条 diagnosis_trace 摘要事件(不含全文 LookupResult) |
|
||||||
|
|
||||||
|
### 4.2 `tool_invocation` 列(RAG)
|
||||||
|
|
||||||
|
| 列 | lookup_knowledge | 其它工具 |
|
||||||
|
|----|------------------|----------|
|
||||||
|
| `tool_name` | `lookup_knowledge` | 各自工具名 |
|
||||||
|
| `retrieval_layer` | 通常 `L1` | `HARNESS` |
|
||||||
|
| `relevance_level` | **PRECISE / REFERENCE / …** | **null**(不再写 evidence_status) |
|
||||||
|
| `l0_match_count` | queryHints | null |
|
||||||
|
| `l1_match_count` | evidence 块数等 | null |
|
||||||
|
| `is_truncated` | 投影 truncated | false |
|
||||||
|
| `retrieval_details` | JSON `rag_lookup_v1` | 通用 status 元数据 |
|
||||||
|
| `output_preview` | level/attempt 摘要 | status=… |
|
||||||
|
| `duration_ms` / `success` | 有 | 有 |
|
||||||
|
|
||||||
|
### 4.3 `retrieval_details`(rag_lookup_v1)示例
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"audit_schema": "rag_lookup_v1",
|
||||||
|
"search_mode": "hybrid",
|
||||||
|
"selected_attempt": "UNFILTERED_VECTOR_RETRY",
|
||||||
|
"fallback_reason": "filtered_vector_low_quality",
|
||||||
|
"category_filter": "overfilter-decoy",
|
||||||
|
"evidence_keys": ["doc#chunk-0"],
|
||||||
|
"sources": ["doc"],
|
||||||
|
"evidence_candidate_count": 8,
|
||||||
|
"evidence_block_count": 2,
|
||||||
|
"l0_hints": { "domains": ["mysql"], "matched_keywords": ["pool"] },
|
||||||
|
"attempts": [
|
||||||
|
{
|
||||||
|
"name": "FILTERED_VECTOR",
|
||||||
|
"category_filter": "overfilter-decoy",
|
||||||
|
"candidate_count": 2,
|
||||||
|
"usable": false,
|
||||||
|
"top_similarity": 0.3,
|
||||||
|
"duration_ms": 12
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "UNFILTERED_VECTOR_RETRY",
|
||||||
|
"candidate_count": 5,
|
||||||
|
"usable": true,
|
||||||
|
"top_similarity": 0.9,
|
||||||
|
"duration_ms": 20
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"truncated": false,
|
||||||
|
"returned_count": 2,
|
||||||
|
"evidence_status": "EVIDENCE_FOUND",
|
||||||
|
"invocation_status": "READY",
|
||||||
|
"tool_call_id": "call-…"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**默认不落库:** 原始 query 全文、chunk 正文 excerpt、完整 rerankTrace(体积与隐私)。
|
||||||
|
|
||||||
|
### 4.4 Trace API:人怎么读 RAG
|
||||||
|
|
||||||
|
**接口:**
|
||||||
|
|
||||||
|
```http
|
||||||
|
GET /api/diagnosis/{sessionId}/trace
|
||||||
|
GET /api/diagnosis/{sessionId}/trace?runId={runId}
|
||||||
|
```
|
||||||
|
|
||||||
|
**实现:** `DiagnosisTraceController` → `DiagnosisTraceService.getTrace`
|
||||||
|
按 `sessionId`(可选精确 `runId`)拉 run、steps、**toolInvocations**、摘要等。
|
||||||
|
|
||||||
|
**响应中与 RAG 相关的核心块:** `DiagnosisTraceResponse.toolInvocations[]`
|
||||||
|
|
||||||
|
| API 字段 | 来源列 | 读法 |
|
||||||
|
|----------|--------|------|
|
||||||
|
| `toolName` | `tool_name` | 是否为 `lookup_knowledge` |
|
||||||
|
| `retrievalLayer` | `retrieval_layer` | L1 / HARNESS |
|
||||||
|
| `relevanceLevel` | `relevance_level` | RAG 粗相关度(非 evidence_status) |
|
||||||
|
| `l0MatchCount` / `l1MatchCount` | 同名列 | L0/L1 规模提示 |
|
||||||
|
| `truncated` | `is_truncated` | 证据是否被投影截断 |
|
||||||
|
| `outputPreview` | `output_preview` | 一行摘要(level/attempt…) |
|
||||||
|
| `retrievalDetails` | 解析自 `retrieval_details` | **RAG Trace 主阵地** |
|
||||||
|
| `retrievalDetailsRaw` | 原始 JSON 字符串 | 调试 |
|
||||||
|
| `durationMs` / `success` / `errorMessage` | 同名列 | 耗时与成败 |
|
||||||
|
| `inputParams` | 通常仅 tool_call_id、request_bytes | **不含完整 query**(有意) |
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TB
|
||||||
|
API["GET /api/diagnosis/{sessionId}/trace"] --> SVC[DiagnosisTraceService]
|
||||||
|
SVC --> ROW[tool_invocation 行]
|
||||||
|
ROW --> T1[列: relevanceLevel, L0/L1 count, layer…]
|
||||||
|
ROW --> T2[retrievalDetails Map]
|
||||||
|
T2 --> D1[search_mode]
|
||||||
|
T2 --> D2[selected_attempt / fallback_reason]
|
||||||
|
T2 --> D3[attempts[] / evidence_keys]
|
||||||
|
T2 --> D4[evidence_status 契约状态]
|
||||||
|
```
|
||||||
|
|
||||||
|
### 4.5 读 Trace 的推荐顺序(排查「这次知识库怎么检的」)
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TB
|
||||||
|
S1[找到 toolName=lookup_knowledge 的 invocation] --> S2{success?}
|
||||||
|
S2 -->|否| E[看 errorMessage / evidence_status]
|
||||||
|
S2 -->|是| S3[看 retrievalDetails.search_mode]
|
||||||
|
S3 --> S4[看 selected_attempt + fallback_reason]
|
||||||
|
S4 --> S5[看 attempts[] 每次 candidate_count / top_similarity / usable]
|
||||||
|
S5 --> S6[看 evidence_keys / sources]
|
||||||
|
S6 --> S7[看 relevanceLevel 列]
|
||||||
|
S7 --> S8[需要原文?看 Agent 侧 evidence 或当时 canonical 存储 · 审计默认无 excerpt]
|
||||||
|
```
|
||||||
|
|
||||||
|
| 现象 | 优先看 |
|
||||||
|
|------|--------|
|
||||||
|
| 为何走了 retry | `fallback_reason` + 两次 `attempts` |
|
||||||
|
| 是否 hybrid | `search_mode` |
|
||||||
|
| 滤错域 | `category_filter` + L0 domains |
|
||||||
|
| 相关度档 | 列 `relevanceLevel`(PRECISE/REFERENCE) |
|
||||||
|
| 返回了哪些块 | `evidence_keys` / `sources`(无正文) |
|
||||||
|
| Agent 是否被截断 | `truncated` / `returned_count` |
|
||||||
|
|
||||||
|
### 4.6 diagnosis_trace 事件 vs tool_invocation 行
|
||||||
|
|
||||||
|
| 通道 | 内容 | 用途 |
|
||||||
|
|------|------|------|
|
||||||
|
| `tool_invocation` 行 | RAG 富字段完整摘要 | **主审计/回放** |
|
||||||
|
| `diagnosis_trace` 中 `TOOL_INVOCATION` | tool_call_id、status、字节数、`has_raw_result` 等薄摘要 | 时间线事件,**不含**完整 retrieval_details |
|
||||||
|
|
||||||
|
查 RAG 细节以 **`toolInvocations[].retrievalDetails`** 为准。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. 与 Agent / Eval 的边界
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart LR
|
||||||
|
subgraph human["人 / 运维 / 评测"]
|
||||||
|
TRACE[Trace API]
|
||||||
|
EVAL[Offline eval]
|
||||||
|
end
|
||||||
|
|
||||||
|
subgraph model["模型"]
|
||||||
|
AGENT[RagToolResult only]
|
||||||
|
end
|
||||||
|
|
||||||
|
TI[(tool_invocation)] --> TRACE
|
||||||
|
FX[fixtures] --> EVAL
|
||||||
|
PROJ[Projector] --> AGENT
|
||||||
|
```
|
||||||
|
|
||||||
|
| 消费者 | 能看到 |
|
||||||
|
|--------|--------|
|
||||||
|
| Agent | evidence + 可选 relevance_level,无 attempt 细节 |
|
||||||
|
| Trace API | 落库摘要:mode/attempt/fallback/keys/level… |
|
||||||
|
| Offline eval | 冻结 fixture 全量 LookupResult(含 trace),与 golden 比对 |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. 与旧文档差异
|
||||||
|
|
||||||
|
| 旧(archive `retrieval-observability`) | 现 |
|
||||||
|
|----------------------------------------|-----|
|
||||||
|
| `vector-store.mode` 多后端 | `search_mode` dense\|hybrid,单一 V2 store |
|
||||||
|
| sink 理想化未落地 | `RagLookupAuditEnricher` + 列回填 |
|
||||||
|
| `relevance_level` 混用 evidence_status | **列仅 RAG 等级**;契约状态在 details |
|
||||||
|
| 未写清 Trace API 读法 | 本文 §4.4–4.5 |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. 代码锚点
|
||||||
|
|
||||||
|
| 职责 | 类 / 路径 |
|
||||||
|
|------|-----------|
|
||||||
|
| 内建 Trace | `LookupKnowledgeTool`、`RetrievalTrace`、`LookupResult` |
|
||||||
|
| 投影 | `RagResultProjector`、`RagToolResult` |
|
||||||
|
| 审计事件 | `ToolInvocationAuditEvent`、`ToolBoundary` |
|
||||||
|
| 富化 | `RagLookupAuditEnricher` |
|
||||||
|
| 落库 | `JpaToolInvocationAuditSink`、`ToolInvocation` |
|
||||||
|
| Trace API | `DiagnosisTraceController`、`DiagnosisTraceService`、`DiagnosisTraceResponse.ToolInvocationTrace` |
|
||||||
|
| 离线回归 | `eval/rag-retrieval/`、`scripts/eval_rag_retrieval.py` |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. 已知限制
|
||||||
|
|
||||||
|
- 持久化 **不存** 完整 query/excerpt(有意);要正文需 Agent 侧证据或其它存储
|
||||||
|
- `dedup_reason` 列可能仍为空
|
||||||
|
- 非 `lookup_knowledge` 工具仍为薄审计
|
||||||
|
- **历史** `tool_invocation` 行可能仍把 evidence_status 写进 `relevance_level`(旧 sink)
|
||||||
|
- `diagnosis_trace` 时间线事件不替代 `retrieval_details`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 9. 相关文档
|
||||||
|
|
||||||
|
| 文档 | 内容 |
|
||||||
|
|------|------|
|
||||||
|
| `mvp/architecture/RAG知识检索架构.md` | 检索主架构 |
|
||||||
|
| `../engineering/rag/RAG-Agent如何读relevance_level.md` | Agent 如何读 level |
|
||||||
|
| `../engineering/rag/RAG-Hybrid质量分与后处理.md` | quality / 排序闸门 |
|
||||||
|
| `../engineering/rag/RAG离线评测-基线设计.md` | 离线评测(非运行时 Trace) |
|
||||||
|
| `mvp/architecture/session-trace-lifecycle.md` | 会话/run Trace 总览(若存在) |
|
||||||
@@ -0,0 +1,376 @@
|
|||||||
|
# RAG 知识检索架构
|
||||||
|
|
||||||
|
**更新日期**:2026-07-28
|
||||||
|
**状态**:当前可运行架构
|
||||||
|
**关联实现**:`lookup_knowledge`、`MilvusHybridKnowledgeStore`、`KnowledgeSearchPort`
|
||||||
|
**关联运维**:`scripts/rebuild_hybrid_knowledge.py`、`POST /api/knowledge/rebuild-hybrid`
|
||||||
|
|
||||||
|
## 1. 定位
|
||||||
|
|
||||||
|
知识检索是 Diagnosis Agent 的显式证据工具,不是隐式 Advisor。
|
||||||
|
|
||||||
|
```text
|
||||||
|
Diagnosis Agent
|
||||||
|
-> lookup_knowledge(query)
|
||||||
|
-> Harness ToolBoundary / ACI projection
|
||||||
|
-> EvidenceGuard 只认当前 Run 的 READY canonical 证据
|
||||||
|
```
|
||||||
|
|
||||||
|
目标:
|
||||||
|
|
||||||
|
- 保留 Agent 可见的工具调用与证据边界
|
||||||
|
- 用单一向量后端完成 dense + BM25 hybrid 检索
|
||||||
|
- 用 chunk 级证据身份保证同文档多片段可同时进入上下文
|
||||||
|
- 检索行为可配置、可重建、可审计
|
||||||
|
|
||||||
|
## 2. 稳定边界
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart LR
|
||||||
|
subgraph AgentBoundary["Agent boundary"]
|
||||||
|
Agent["Diagnosis Agent"]
|
||||||
|
Tool["lookup_knowledge"]
|
||||||
|
end
|
||||||
|
|
||||||
|
subgraph HarnessBoundary["Harness boundary"]
|
||||||
|
Adapter["RagToolAdapter"]
|
||||||
|
Projector["RagResultProjector"]
|
||||||
|
Canonical["Redis canonical invocation"]
|
||||||
|
end
|
||||||
|
|
||||||
|
subgraph RetrievalBoundary["Retrieval boundary"]
|
||||||
|
Backend["LookupKnowledgeTool"]
|
||||||
|
Port["KnowledgeSearchPort"]
|
||||||
|
Store["MilvusHybridKnowledgeStore"]
|
||||||
|
end
|
||||||
|
|
||||||
|
Agent --> Tool
|
||||||
|
Tool --> Adapter
|
||||||
|
Adapter --> Backend
|
||||||
|
Backend --> Port
|
||||||
|
Port --> Store
|
||||||
|
Adapter --> Projector
|
||||||
|
Adapter --> Canonical
|
||||||
|
```
|
||||||
|
|
||||||
|
| 边界 | 职责 | 不负责 |
|
||||||
|
|---|---|---|
|
||||||
|
| Agent | 决定何时检索、如何用证据写报告 | 不直接访问 Milvus / MySQL 元数据表 |
|
||||||
|
| Harness | Tool 校验、投影裁剪、canonical 存证 | 不改写检索排序算法 |
|
||||||
|
| Retrieval | L0 hint、dense/BM25 召回、后处理、打包 | 不绕过 ACI 直接给 Agent 原始库响应 |
|
||||||
|
|
||||||
|
## 3. 当前主链路
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TD
|
||||||
|
A["lookup_knowledge(query)"] --> B["KnowledgeQueryTransformer"]
|
||||||
|
B --> C["L0 hint: domain / keywords / categoryFilter"]
|
||||||
|
C --> D["KnowledgeDocumentRetriever"]
|
||||||
|
D --> E["KnowledgeSearchPort"]
|
||||||
|
E --> F["VectorSearchService"]
|
||||||
|
F --> G{"retrieval.search.mode"}
|
||||||
|
G -->|dense| H["MilvusHybridKnowledgeStore.searchDense"]
|
||||||
|
G -->|hybrid| I["MilvusHybridKnowledgeStore.searchHybrid"]
|
||||||
|
H --> J["candidates + chunk identity"]
|
||||||
|
I --> J
|
||||||
|
J --> K["KnowledgeEvidencePostProcessor"]
|
||||||
|
K --> L["evidenceKey dedup / maxChunksPerDocument / return-n"]
|
||||||
|
L --> M{"filtered low quality?"}
|
||||||
|
M -->|yes and had categoryFilter| N["unfiltered retry"]
|
||||||
|
N --> K
|
||||||
|
M -->|no| O["KnowledgeContextPacker"]
|
||||||
|
O --> P["LookupResultAssembler"]
|
||||||
|
P --> Q["RagResultProjector"]
|
||||||
|
Q --> R["Agent-facing RagToolResult"]
|
||||||
|
```
|
||||||
|
|
||||||
|
对应代码:
|
||||||
|
|
||||||
|
| 阶段 | 类 | 职责 |
|
||||||
|
|---|---|---|
|
||||||
|
| Tool 编排 | `LookupKnowledgeTool` | 串联 transform / retrieve / post / pack |
|
||||||
|
| Query 理解 | `KnowledgeQueryTransformer` + `KnowledgeIndexService` | L0 只产 hint 与可选 category filter |
|
||||||
|
| 检索端口 | `KnowledgeSearchPort` / `VectorKnowledgeSearchAdapter` | 屏蔽底层存储细节 |
|
||||||
|
| 检索门面 | `VectorSearchService` | `dense` 或 `hybrid` 路由 |
|
||||||
|
| 向量后端 | `MilvusHybridKnowledgeStore` | 唯一知识库读写后端(MilvusClientV2) |
|
||||||
|
| 后处理 | `KnowledgeEvidencePostProcessor` | 归一化、规则 boost、chunk 去重、相关度等级 |
|
||||||
|
| 打包 | `KnowledgeContextPacker` | 有界 context pack |
|
||||||
|
| 投影 | `RagResultProjector` | 只暴露 Agent 可见 evidence 字段 |
|
||||||
|
|
||||||
|
## 4. 唯一向量后端:MilvusClientV2
|
||||||
|
|
||||||
|
### 4.1 已废弃路径
|
||||||
|
|
||||||
|
以下路径**不再**用于 `lookup_knowledge`:
|
||||||
|
|
||||||
|
- legacy `MilvusServiceClient` search / insert
|
||||||
|
- `retrieval.vector-store.mode=sdk|spring|auto`
|
||||||
|
- Spring AI `VectorStore` 作为知识检索主路径
|
||||||
|
|
||||||
|
### 4.2 当前后端
|
||||||
|
|
||||||
|
```text
|
||||||
|
写入:
|
||||||
|
VectorIndexService
|
||||||
|
-> MilvusHybridKnowledgeStore.upsertChunk
|
||||||
|
|
||||||
|
读取:
|
||||||
|
VectorSearchService
|
||||||
|
-> MilvusHybridKnowledgeStore.searchDense
|
||||||
|
-> MilvusHybridKnowledgeStore.searchHybrid
|
||||||
|
```
|
||||||
|
|
||||||
|
默认 collection:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
milvus:
|
||||||
|
collection: biz
|
||||||
|
```
|
||||||
|
|
||||||
|
重建时会 drop + recreate 该 collection,并按 dense + BM25 schema 重建。
|
||||||
|
|
||||||
|
## 5. Collection Schema
|
||||||
|
|
||||||
|
`biz`(可配置)逻辑字段:
|
||||||
|
|
||||||
|
| 字段 | 类型 | 用途 |
|
||||||
|
|---|---|---|
|
||||||
|
| `id` | VarChar PK | chunk 级主键 |
|
||||||
|
| `content` | VarChar | 返回给 Agent 的原文片段 |
|
||||||
|
| `search_text` | VarChar + analyzer | BM25 输入文本 |
|
||||||
|
| `sparse_vector` | SparseFloatVector | BM25 Function 输出 |
|
||||||
|
| `vector` | FloatVector | dense embedding |
|
||||||
|
| `metadata` | JSON | docId / chunkIndex / category / kb_scope / title / breadcrumb 等 |
|
||||||
|
|
||||||
|
Function:
|
||||||
|
|
||||||
|
```text
|
||||||
|
BM25(search_text -> sparse_vector)
|
||||||
|
```
|
||||||
|
|
||||||
|
索引:
|
||||||
|
|
||||||
|
```text
|
||||||
|
vector -> IVF_FLAT + L2
|
||||||
|
sparse_vector -> SPARSE_INVERTED_INDEX + BM25
|
||||||
|
```
|
||||||
|
|
||||||
|
写入时:
|
||||||
|
|
||||||
|
- `content` 保存原始 chunk 正文
|
||||||
|
- `search_text` / dense embedding 使用 `Title + Path + Content` 拼装文本
|
||||||
|
- metadata 必须带 `docId`、`chunkIndex`,供 chunk 级证据身份使用
|
||||||
|
|
||||||
|
## 6. 检索模式
|
||||||
|
|
||||||
|
配置:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
retrieval:
|
||||||
|
search:
|
||||||
|
mode: hybrid # dense | hybrid(见 6.0 用途约定)
|
||||||
|
hybrid:
|
||||||
|
rrf-k: 60
|
||||||
|
kb-scope: ""
|
||||||
|
rag:
|
||||||
|
retrieve-k: 20
|
||||||
|
return-n: 5
|
||||||
|
max-chunks-per-document: 2
|
||||||
|
```
|
||||||
|
|
||||||
|
### 6.0 模式用途约定(保留双 mode 的原因)
|
||||||
|
|
||||||
|
知识库 **只维护一套** dense + BM25 schema 数据(默认 collection `biz`)。
|
||||||
|
`retrieval.search.mode` 切换的是**同库上的查询算法**,不是两套互斥索引、也不是两套写入路径。
|
||||||
|
|
||||||
|
| 模式 | 定位 | 说明 |
|
||||||
|
|---|---|---|
|
||||||
|
| **hybrid** | **线上主路径 / 默认** | dense ANN + 服务端 BM25 + RRF;`lookup_knowledge` 正式召回只认此模式 |
|
||||||
|
| **dense** | **对照 / 评测 / 排障** | 仅 dense ANN,用于和 hybrid 对比召回效果(命中文档/chunk、排名差异等) |
|
||||||
|
|
||||||
|
约定:
|
||||||
|
|
||||||
|
1. 生产配置保持 `mode: hybrid`;不要把 dense 当成第二套长期并行的线上策略。
|
||||||
|
2. 需要看「去掉 BM25+RRF 后召回差在哪」时,临时切 `mode: dense`,其它参数(`retrieve-k`、`return-n`、category filter、query 集)尽量固定,再切回 hybrid。
|
||||||
|
3. hybrid 入库的数据 **完全适用于** dense-only 查询:每条 chunk 都写了 `vector`;dense 模式只是不使用 `sparse_vector` / BM25 子路。
|
||||||
|
4. 代码里 `@Value` 在配置缺失时的兜底仍可能是 `dense`(历史兼容);**以 `application.yml` 的 hybrid 为准**。若做回归,确认运行配置而不是只看注解默认值。
|
||||||
|
|
||||||
|
不建议的用法:
|
||||||
|
|
||||||
|
- 按请求/按租户在 dense 与 hybrid 之间当产品功能随意切换(当前也无稳定的 per-call mode 覆盖)。
|
||||||
|
- 把 dense 模式的相关度表现直接当成 hybrid 的最终质量结论(hybrid 排序信 RRF,后处理分数仍多 L2 兼容,见下节)。
|
||||||
|
|
||||||
|
### 6.1 dense(对照基线)
|
||||||
|
|
||||||
|
```text
|
||||||
|
query
|
||||||
|
-> embedding
|
||||||
|
-> dense ANN on vector
|
||||||
|
-> topK
|
||||||
|
```
|
||||||
|
|
||||||
|
仅走 `vector` 字段的 L2 ANN。用于基线对比,不作为正式主路径。
|
||||||
|
|
||||||
|
### 6.2 hybrid(当前默认 / 主路径)
|
||||||
|
|
||||||
|
```text
|
||||||
|
query
|
||||||
|
-> path A: dense ANN(query embedding)
|
||||||
|
-> path B: BM25 sparse ANN(raw query text)
|
||||||
|
-> Milvus hybridSearch + RRFRanker(k)
|
||||||
|
-> topK fused hits
|
||||||
|
```
|
||||||
|
|
||||||
|
说明:hybrid **内部**的 dense 子路是融合的一部分,与配置项 mode=dense(整次检索只跑单路 ANN)不是同一概念。
|
||||||
|
|
||||||
|
分数与后处理(quality 统一,2026-07-28):
|
||||||
|
|
||||||
|
- 一级 scoreLabel 仅 **dense | hybrid**(旧别名 canonicalize)。
|
||||||
|
- **dense**:score = L2;qualityScore = 1 - clamp(L2)/maxL2Distance。
|
||||||
|
- **hybrid**:返回序 = RRF 序;qualityScore 由 **本轮 rank 线性映射**(不把 RRF 原分当 L2;不做 dense L2 回填覆盖主分;无 m25_only_* 一级 label)。
|
||||||
|
- 后处理:**统一**消费 qualityScore;排序主序 = originalRank;**不做** L0 关键词/domain contains 加分改序(重叠仅可写 hitReasons 解释)。
|
||||||
|
-
|
||||||
|
elevance_level / category 低质 unfiltered retry:只看 top qualityScore 与阈值。
|
||||||
|
- 实现:RetrievalScoreNormalizer、KnowledgeEvidencePostProcessor;详见 OpenSpec
|
||||||
|
ag-quality-score-unify。
|
||||||
|
|
||||||
|
### 6.3 category filter 与降级
|
||||||
|
|
||||||
|
```text
|
||||||
|
if L0 给出唯一 domain:
|
||||||
|
先 filtered 检索
|
||||||
|
if 无证据或 topSimilarity < referenceThreshold:
|
||||||
|
再 unfiltered retry
|
||||||
|
else:
|
||||||
|
直接 unfiltered
|
||||||
|
```
|
||||||
|
|
||||||
|
这里的 filter 是 metadata category / kb_scope 约束,不是第二套向量库。
|
||||||
|
|
||||||
|
## 7. 证据身份与去重
|
||||||
|
|
||||||
|
Delivery 1 已落地:
|
||||||
|
|
||||||
|
```text
|
||||||
|
evidenceKey =
|
||||||
|
docId#chunk-{chunkIndex}
|
||||||
|
fallback: vector:{id}
|
||||||
|
fallback: rank:{n}
|
||||||
|
```
|
||||||
|
|
||||||
|
规则:
|
||||||
|
|
||||||
|
- 去重按 `evidenceKey`,不是按 source 文档路径
|
||||||
|
- 同文档不同 chunk 可同时保留
|
||||||
|
- `rag.max-chunks-per-document` 限制单文档最多进入结果的 chunk 数
|
||||||
|
- `rag.return-n` 限制后处理后最多返回条数
|
||||||
|
- Agent 投影中的 `document_id` 使用 chunk 级 evidenceKey
|
||||||
|
|
||||||
|
这保证 hybrid 召回的多片段不会在后处理/投影阶段被文档级折叠吞掉。
|
||||||
|
|
||||||
|
## 8. L0 / L1 职责
|
||||||
|
|
||||||
|
| 层 | 做什么 | 不做什么 |
|
||||||
|
|---|---|---|
|
||||||
|
| L0 | domain/keyword hint、可选 category filter、trace 解释、轻规则 boost | 不直接当事实 evidence |
|
||||||
|
| L1 dense/BM25 | 事实证据召回 | 不依赖 frontmatter 关键词命中才返回正文 |
|
||||||
|
|
||||||
|
L0 命中文档正文不会在 L1 失败时兜底成 evidence。
|
||||||
|
|
||||||
|
## 9. Agent 可见契约
|
||||||
|
|
||||||
|
Agent 只看到有界 `RagToolResult`:
|
||||||
|
|
||||||
|
- `evidence_status`
|
||||||
|
- `tool_call_id`
|
||||||
|
- `query`
|
||||||
|
- `evidence[]`:`document_id` / `source` / `title` / `breadcrumb` / `excerpt`
|
||||||
|
- `relevance_level`
|
||||||
|
- `truncated` / `returned_count`
|
||||||
|
|
||||||
|
不暴露:
|
||||||
|
|
||||||
|
- raw score / fused score
|
||||||
|
- retrievalTrace / rerankTrace
|
||||||
|
- contextPack 全文
|
||||||
|
- Milvus 内部字段与凭据
|
||||||
|
|
||||||
|
完整内部结果仍在 `LookupResult` 中,供审计与调试使用。
|
||||||
|
|
||||||
|
### 9.1 Trace 与审计(现行入口)
|
||||||
|
|
||||||
|
请求内 `retrievalTrace` / 落库 `tool_invocation` / Trace API 读法见:
|
||||||
|
|
||||||
|
**[RAG检索可观测性与审计.md](./RAG检索可观测性与审计.md)**
|
||||||
|
|
||||||
|
要点:
|
||||||
|
|
||||||
|
- Agent **看不到**完整 retrievalTrace;人通过 `GET /api/diagnosis/{sessionId}/trace` 的 `toolInvocations[].retrievalDetails` 回放。
|
||||||
|
- `relevance_level` 列存 RAG 等级(PRECISE/REFERENCE);`evidence_status` 在 details JSON。
|
||||||
|
- 默认审计不落原始 query 全文与 excerpt 正文。
|
||||||
|
|
||||||
|
## 10. 写入与重建
|
||||||
|
|
||||||
|
### 10.1 日常写入
|
||||||
|
|
||||||
|
文档上传 / 知识库初始化:
|
||||||
|
|
||||||
|
```text
|
||||||
|
markdown
|
||||||
|
-> frontmatter + body
|
||||||
|
-> DocumentChunkService
|
||||||
|
-> dense embedding + search_text
|
||||||
|
-> MilvusHybridKnowledgeStore.upsertChunk
|
||||||
|
-> MySQL api_document + L0 memory index
|
||||||
|
```
|
||||||
|
|
||||||
|
### 10.2 全量重建
|
||||||
|
|
||||||
|
危险操作,需显式确认:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python scripts/rebuild_hybrid_knowledge.py --confirm REBUILD
|
||||||
|
```
|
||||||
|
|
||||||
|
等价 API:
|
||||||
|
|
||||||
|
```text
|
||||||
|
POST /api/knowledge/rebuild-hybrid?confirm=REBUILD
|
||||||
|
```
|
||||||
|
|
||||||
|
服务端顺序:
|
||||||
|
|
||||||
|
1. drop + recreate `milvus.collection`(默认 `biz`)
|
||||||
|
2. 清空 MySQL `api_document`
|
||||||
|
3. 清空内存 L0
|
||||||
|
4. 扫描 `knowledge_base/**/*.md`(跳过 `README.md`)force 导入
|
||||||
|
|
||||||
|
不会修改磁盘上的 `knowledge_base/` 源文件。
|
||||||
|
|
||||||
|
## 11. 与旧文档的差异
|
||||||
|
|
||||||
|
| 旧描述(已归档) | 当前实现 |
|
||||||
|
|---|---|
|
||||||
|
| Spring AI VectorStore 主路径 + SDK fallback | 单一 MilvusClientV2 后端 |
|
||||||
|
| `retrieval.vector-store.mode=auto/sdk/spring` | 已移除;改为 `retrieval.search.mode=dense/hybrid` |
|
||||||
|
| source 级 evidence 去重 | chunk 级 `evidenceKey` 去重 |
|
||||||
|
| 应用层 sparse-lite lexical 伪 hybrid | 库内 dense ANN + BM25 + RRFRanker |
|
||||||
|
| 新建 `biz_hybrid` 过渡 collection | 默认使用并重建 `biz` |
|
||||||
|
|
||||||
|
历史材料见:
|
||||||
|
|
||||||
|
- `mvp/architecture/archive/2026-07-22-legacy/rag-architecture.md`
|
||||||
|
- `mvp/architecture/archive/2026-07-22-legacy/modular-rag-pipeline.md`
|
||||||
|
|
||||||
|
## 12. 当前已知边界
|
||||||
|
|
||||||
|
- hybrid 依赖云端/实例支持 BM25 Function 与 sparse index
|
||||||
|
- 全量重建受 embedding API 与 Milvus 写入延迟影响,可能较慢
|
||||||
|
- L0 关键词匹配仍较粗,只作 hint,不作主召回
|
||||||
|
- 尚未做邻块上下文自动扩展、cross-encoder rerank、真 query rewrite
|
||||||
|
- `totalVectors` 统计接口仍可能返回 0,不代表 collection 为空;以 rebuild/init 结果与检索命中为准
|
||||||
|
- `retrieval.search.mode=dense` 仅作召回对照,不是第二套主路径
|
||||||
|
- 同一 hybrid schema 数据可被 dense / hybrid 两种查询复用;从纯旧 dense-only collection 升级必须 rebuild
|
||||||
|
- hybrid 质量闸门优先用 `denseDistance` 绝对 L2;无 dense 时 rank 回退;排序仍跟 RRF
|
||||||
|
- 后处理不再用 L0 关键词 boost 改序;词面信号以库内 BM25+RRF 为准
|
||||||
|
- Trace/审计细节与限制见 [RAG检索可观测性与审计.md](./RAG检索可观测性与审计.md)
|
||||||
+15
-40
@@ -1,48 +1,23 @@
|
|||||||
# MVP 架构文档
|
# MVP 架构文档
|
||||||
|
|
||||||
**更新日期**:2026-07-10
|
**更新日期**:2026-07-29
|
||||||
|
**状态**:当前单 Diagnosis Agent + Harness 架构
|
||||||
|
|
||||||
这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到:
|
当前文档入口:
|
||||||
|
|
||||||
- `mvp/architecture/archive/2026-07-05-legacy/`
|
| 文档 | 内容 |
|
||||||
|
|
||||||
归档材料只作为设计历史阅读,不再作为当前实现依据。
|
|
||||||
|
|
||||||
## 当前文档
|
|
||||||
|
|
||||||
| 文档 | 用途 |
|
|
||||||
|---|---|
|
|---|---|
|
||||||
| [current-mvp-architecture.md](current-mvp-architecture.md) | 当前可运行 MVP 的总体架构、链路、持久化和质量门禁 |
|
| [current-mvp-architecture.md](current-mvp-architecture.md) | 系统分层、请求主链、Trace/Reasoning 边界与 API surface |
|
||||||
| [interview-one-pager.md](interview-one-pager.md) | 面试一页式架构讲解,包含总图、亮点、取舍和追问回答 |
|
| [agent-orchestration.md](agent-orchestration.md) | 单 Diagnosis ReAct Agent 的职责、执行方式和 reasoning 采集边界 |
|
||||||
| [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat SequentialAgent、AIOps SupervisorAgent、工具边界 |
|
| [harness-quality-gates.md](harness-quality-gates.md) | Run、Tool、Evidence、Semantic、Release 与 Trace Recorder 门禁 |
|
||||||
| [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md) | Chat 证据链路当前数据契约,覆盖 Executor V2、Gatekeeper、Verifier、Composer、`evidence_refs` |
|
| [session-trace-lifecycle.md](session-trace-lifecycle.md) | sessionId/runId、SSE、统一 Timeline 和 reasoning audit 生命周期 |
|
||||||
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Gatekeeper、Verifier、Composer、评测基线组成的质量门禁 |
|
| [diagnosis-information-gain-stop-architecture.md](diagnosis-information-gain-stop-architecture.md) | 已实施的信息增益评价、Harness 饱和检测、Draft 合同失败降级与证据不足停止设计 |
|
||||||
| [rag-architecture.md](rag-architecture.md) | RAG/知识检索新架构,覆盖 L0 hint、VectorStore 主路径、SDK fallback、证据追踪 |
|
| [RAG知识检索架构.md](RAG知识检索架构.md) | 当前 `lookup_knowledge` 检索:MilvusClientV2 dense+BM25 hybrid、chunk 证据身份、重建运维 |
|
||||||
| [modular-rag-pipeline.md](modular-rag-pipeline.md) | `lookup_knowledge` 模块化 RAG 落地架构,覆盖 pipeline、fallback、evidence-first contract、trace |
|
| [RAG检索可观测性与审计.md](RAG检索可观测性与审计.md) | RAG Trace / 审计:请求内 retrievalTrace、tool_invocation 富字段、Trace API 读法 |
|
||||||
| [rag-eval-closure.md](rag-eval-closure.md) | RAG 评测闭环,覆盖 offline baseline、baseline diff、diagnosis eval 和 live acceptance |
|
|
||||||
| [retrieval-observability.md](retrieval-observability.md) | 检索运行细节和可观测性,覆盖 L0/L1、去重、分数归一、评测 |
|
|
||||||
| [feedback-architecture.md](feedback-architecture.md) | 反馈与自评估闭环,覆盖 rule evaluation、Verifier、AIOps rule、用户反馈和案例沉淀 |
|
|
||||||
| [session-trace-lifecycle.md](session-trace-lifecycle.md) | 会话和 Trace 生命周期,覆盖 sessionId、状态流转、agent_step、tool_invocation、Trace API |
|
|
||||||
| [knowledge-base-authoring.md](knowledge-base-authoring.md) | 知识库文档编写与维护规范,覆盖 frontmatter、category、chunk、reindex |
|
|
||||||
| [data-model.md](data-model.md) | 数据模型总览,覆盖 Trace、知识库、反馈沉淀和 Milvus metadata |
|
|
||||||
| [evolution-roadmap.md](evolution-roadmap.md) | 从旧版 Agent 蓝图继承的后续演进路线,不代表当前已实现 |
|
|
||||||
|
|
||||||
## 当前架构一句话
|
**工程纪要**(问题 / 决策 / E2E,非架构规范正文)见 [../engineering/README.md](../engineering/README.md)。
|
||||||
|
|
||||||
SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,Chat 链路由 Gatekeeper 做引用真实性校验、Verifier 做可推导性判断、Composer 生成最终表达;多轮会话元数据落到 `chat_session`,每次诊断运行落到 `diagnosis_run`,步骤和工具明细通过 `agent_step.run_id`、`tool_invocation.run_id` 关联,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
|
2026-07-22 前的多角色编排、双入口和旧证据链文档已移动到 `archive/2026-07-22-legacy/`,仅用于历史决策追溯,不代表当前运行时。其中旧 RAG 描述(Spring AI VectorStore 主路径 + Milvus SDK fallback)已被当前 hybrid 实现取代,请以 [RAG知识检索架构.md](RAG知识检索架构.md) 为准。检索可观测与 Trace 以 [RAG检索可观测性与审计.md](RAG检索可观测性与审计.md) 为准(勿再依赖 archive 内旧 retrieval-observability)。
|
||||||
|
|
||||||
## 阅读顺序
|
当前普通 Trace 与 LLM 步骤审计(`agent_reasoning_audit`:`reasoning_content` + `assistant_text`)使用独立存储和独立接口。
|
||||||
|
DeepSeek thinking 捕获路径与 V015–V017 字段已 live 验证(2026-07-28)。Reasoning 访问控制、保留期限、加密要求仍由 ISS-015 跟踪,不能把“数据已分表 + 能抓到 thinking”理解为“治理已经完成”。
|
||||||
1. 先读 [current-mvp-architecture.md](current-mvp-architecture.md),理解系统边界和主链路。
|
|
||||||
2. 面试前读 [interview-one-pager.md](interview-one-pager.md),准备 2-5 分钟讲解。
|
|
||||||
3. 再读 [agent-orchestration.md](agent-orchestration.md),理解当前 Agent 如何协作。
|
|
||||||
4. 接着读 [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md),理解 Chat 证据链路的数据结构和验真边界。
|
|
||||||
5. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。
|
|
||||||
6. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。
|
|
||||||
7. 继续读 [modular-rag-pipeline.md](modular-rag-pipeline.md),看 `lookup_knowledge` 的模块化落地和 evidence-first contract。
|
|
||||||
8. 再读 [rag-eval-closure.md](rag-eval-closure.md),看 RAG baseline 如何形成质量闭环。
|
|
||||||
9. 然后读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
|
|
||||||
10. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
|
|
||||||
11. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
|
|
||||||
12. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
|
|
||||||
13. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
|
|
||||||
|
|||||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user