Compare commits

...
Author SHA1 Message Date
zhuyongxin 7ae9707a3b feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit
(DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG
tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools,
and align MVP docs after live E2E verification.
2026-07-28 19:43:13 +08:00
zhuyongxin 2f40536248 docs(mvp): document dense+BM25 hybrid knowledge retrieval
Add current RAG architecture covering MilvusClientV2 hybrid search,
chunk evidence identity, rebuild ops, and update the MVP architecture
index and system overview links.
2026-07-28 09:56:08 +08:00
zhuyongxin 729cd3544a chore(rag): remove obsolete PowerShell rebuild script 2026-07-28 09:50:15 +08:00
zhuyongxin 1e532fb851 chore(rag): rebuild script in Python and use biz collection
Switch default knowledge collection back to biz (drop+recreate on rebuild),
replace PowerShell rebuild runner with Python, skip README.md imports, and
merge duplicate rag keys in application.yml.
2026-07-28 09:49:43 +08:00
zhuyongxin 38f781b157 feat(harness): complete protocol repair stop and archive ISS-016
Add repairable INVALID_PROGRESS_PROTOCOL observations, independent
PROGRESS_PROTOCOL_VIOLATED saturation, and controlled release paths.
Archive the OpenSpec change after syncing main specs and devflow.
2026-07-27 19:10:07 +08:00
zhuyongxin 5c369f3b6c feat(rag): add hybrid knowledge rebuild API and script
Add confirm-gated rebuild-hybrid endpoint that drops biz_hybrid, clears
api_document and L0, then force-imports knowledge_base markdown into the
dense+BM25 store. Include PowerShell runner and ops README.
2026-07-27 18:55:00 +08:00
zhuyongxin f035538531 feat(rag): dense+BM25 hybrid on MilvusClientV2, drop SDK path
Replace legacy MilvusServiceClient knowledge search/write with a single
MilvusClientV2 hybrid store (BM25 function + dense ANN + RRFRanker).
Use collection biz_hybrid and require knowledge reindex.
2026-07-27 18:49:20 +08:00
zhuyongxin 376ad0c241 feat(rag): hybrid multi-path search with RRF fusion
Add configurable hybrid mode on KnowledgeSearchPort that fuses dense
unfiltered, dense filtered, and lexical ranks via RRF while preserving
dense-compatible threshold scores. Archives Delivery 2 OpenSpec change.
2026-07-27 18:33:31 +08:00
zhuyongxin ac1f831903 feat(rag): chunk evidence identity, dedup, and search port
Preserve same-document multi-chunk evidence with evidenceKey identity,
per-document caps, retrieve-k/return-n split, and a dense KnowledgeSearchPort.
Archives Delivery 1 OpenSpec change as the foundation for hybrid retrieval.
2026-07-27 18:26:15 +08:00
zhuyongxin 99d4f6f216 docs(rag): add comments on knowledge retrieval pipeline
Document the lookup_knowledge flow from L0 hints through L1 retrieval,
post-processing, packing, and Agent projection so the boundaries and
current limitations are easier to follow.
2026-07-27 16:26:06 +08:00
aruo edb6153fd6 docs(issue): mark iss-015 partially implemented 2026-07-27 10:22:09 +08:00
aruo d0452184ee feat(harness): add information gain stop and audit 2026-07-27 01:03:34 +08:00
zhuyongxin de5a5b09d9 docs(interview): refresh materials for single-agent harness narrative
Archive pre-refactor interview notes and add current deep-dives on
architecture evolution, issue-derived stories, and evidence gates.
2026-07-24 18:14:49 +08:00
zhuyongxin e47f2dead0 docs(mvp): align architecture and audit table design 2026-07-23 20:16:29 +08:00
zhuyongxin 49180abccf docs(issue): archive legacy issues and add iss-015 2026-07-23 19:54:01 +08:00
zhuyongxin 529f4ff43b docs(demo): add successful diagnosis request 2026-07-23 19:06:37 +08:00
zhuyongxin 75fa154a0a chore(agent): add shared skills and GitNexus guidance 2026-07-23 19:05:22 +08:00
zhuyongxin 8300435a63 docs(issue): add iss-014 handoff 2026-07-23 18:03:14 +08:00
zhuyongxin e20249c5d9 feat(harness): improve trace fallback and reasoning audit 2026-07-23 17:52:01 +08:00
zhuyongxin 8fbc443f76 docs(issue): close iss-014 stage status 2026-07-22 18:08:29 +08:00
zhuyongxin 8ee7cc0b70 refactor(harness): remove legacy agent architecture 2026-07-22 18:02:01 +08:00
zhuyongxin bc36248cd8 feat(chat): cut over to single SSE endpoint 2026-07-22 10:01:12 +08:00
zhuyongxin f8809cb7dd feat(harness): add chat application use case 2026-07-22 00:57:39 +08:00
zhuyongxin ee0949d464 feat(harness): add evidence and semantic guards 2026-07-21 23:50:51 +08:00
zhuyongxin 2362665519 feat(harness): add single diagnosis react agent 2026-07-21 22:29:49 +08:00
zhuyongxin 85029d96a7 feat(harness): add readonly mysql tool 2026-07-21 21:24:42 +08:00
zhuyongxin 3e602781d6 feat(harness): add rag and log projections 2026-07-21 20:12:18 +08:00
zhuyongxin 0dbdd7d8d3 feat(harness): add canonical tool invocation boundary 2026-07-21 19:36:25 +08:00
zhuyongxin 6b74990f86 feat(harness): add run context and retry core 2026-07-21 18:36:19 +08:00
zhuyongxin 4274f3350b feat(harness): freeze aci tool contracts 2026-07-21 18:01:18 +08:00
zhuyongxin 58c39107c5 refactor(harness): freeze single-agent contracts 2026-07-21 17:33:25 +08:00
724 changed files with 48078 additions and 16343 deletions
+259
View File
@@ -0,0 +1,259 @@
---
name: essence
description: Invoke when a project is too large or you only want the core design insights. Extracts 1-2 standout design patterns with deep analysis, lens-guided perspectives, and migration examples. Not for full project analysis or quick lookups.
metadata:
version: "0.5.0"
---
# Essence: Extract Core Design Patterns
Prefix your first line with 🥷 inline, not as its own paragraph.
You are a jewel inspector. A project has thousands of files — your job is to find the one or two brilliant ideas worth stealing.
**This is NOT a lite version of `/explore`.** `/explore` reads the whole project and summarizes at the end. `/essence` goes deep on one thing and ignores everything else.
## Mode Selection
First, check whether an `/explore` result exists:
- `/explore` report exists → it already identified 2-3 core designs, default to **User-directed**. Ask the user which design to deep-dive, or whether to switch mode.
- No `/explore` result → this is an independent launch, default to **Auto-detect**.
Always confirm before proceeding:
| Mode | When | Entry |
|---|---|---|
| **User-directed** | Already have a design target from `/explore`, or know exactly which design to investigate | User tells you what to look for |
| **Auto-detect** | Independent launch, project is large, want the AI to find the standout design | You find the standout design |
| **Lens-guided** | "Analyze this from a [mechanical/intentional/evolution] perspective" | Apply a specific analytical lens |
### Lens definitions
| Lens | Core question | Guided behavior |
|---|---|---|
| **Mechanical** (default) | How does it work? | Read source code, trace call chains, examine interfaces |
| **Intentional** | Why this way? | Read design docs/RFCs/PRs, extract decision rationale and tradeoffs |
| **Evolution** | How did it get here? | Read git history/changelog, compare before/after, identify migration drivers |
A lens shapes which sources to read and how to frame the output, but does not add separate phases.
### Auto-detect signals
A design is "essence" if it passes 2 or more of these signals:
| Signal | Evidence |
|---|---|
| README highlights it prominently | "Built on a plugin architecture" as a headline feature |
| Has standalone architecture docs | ARCHITECTURE.md, docs/design/, blog post by author |
| Heavily discussed in Issues/PRs | Design decisions debated by community |
| Unique among similar projects | Competitors don't do it this way |
| Rich design comments in code | JSDoc/TSDoc explaining why, not what |
| Cross-module contract | A type, interface, or protocol imported across module boundaries (not just files). Go: most-implemented interface. Python: most-subclassed abstract base. Rust: most-implemented trait. These define subsystem relationships. |
| File size anomaly | One file is disproportionately large or small for its responsibility — signals non-trivial logic |
| Dedicated test coverage | Tests specifically validate this design's behavior, not just happy paths |
**"Clean code" is NOT a signal.** A well-written utility function is not essence. An architecture decision that shapes the entire project is.
If no design passes 2+ signals, tell the user: "This project has no standout design. Try `/explore` for a full analysis instead."
## Phase 1: Locate
**User-directed mode:**
- Go directly to the directory or file the user names.
- If the directory doesn't exist, stop and tell the user. Do NOT invent an alternative.
**Auto-detect mode:**
- Scan README, AGENTS.md, and top-level docs for architecture claims.
- Identify 1-2 standout design directions.
- Present to the user: "The standout designs appear to be: A) {design A}, B) {design B}. Which should we dive into?"
- If user doesn't choose, pick the strongest one and state why.
**Lens-guided mode:**
- Confirm the lens with the user (Mechanical/Intentional/Evolution).
- Frame the search in terms of the lens.
- Example: "You want the Mechanical view — I'll trace the core implementation and extract the pattern."
**Output:** 1-2 design directions to analyze + lens confirmation.
**Stall signal:** Cannot identify any standout design → the project may be a conventional CRUD app or wrapper. Stop and recommend `/explore` or a different project.
## Phase 2: Deep Dive
Read the core files related to the chosen design. Maximum 10 files. Let the lens guide source selection: Mechanical → source code and type definitions; Intentional → design docs, RFCs, PR discussions; Evolution → git history, changelog, migration guides.
**For each file:**
- What role does it play in this design?
- What interfaces does it expose?
- How does it connect to other parts of the system?
**Trace the call chain:**
- Start from the entry point that uses this design.
- Follow the flow until you understand the full pattern.
- Stop when you hit boilerplate, config, or test files.
**Output:** Core file list (≤10) + call chain + lens-specific annotations.
**Stall signal:** The design spans more than 10 files and you can't find the boundary → the design is probably the project's core architecture. Switch to `/explore` for a full analysis instead.
## Phase 3: Extract Pattern
Analyze the design at a higher level. Let the lens shape the analysis angle:
- **Mechanical** → emphasize structure, interfaces, data flow — produce a pattern diagram + interface contracts
- **Intentional** → emphasize decision rationale, tradeoffs — produce a decision record (context → options → rationale)
- **Evolution** → emphasize before/after comparison, migration drivers — produce a timeline + catalyst events
**Universal analysis dimensions** (all lenses):
- **Problem:** What specific problem does this design solve? What was the pain before?
- **Pattern:** What's the name of this pattern? (Named: MVC, Observer, Plugin, Middleware. Custom: describe it in one sentence.)
- **Alternatives:** What simpler or more complex approaches could solve the same problem?
- **Tradeoffs:** Why did the author choose this? What does it give up?
- **Evidence:** What in the code proves this analysis is correct? (Specific files, functions, comments.)
**Output:** Design pattern card (lens-framed).
**Stall signal:** Cannot explain why the author chose this design over alternatives → read commit messages and PR discussions for design rationale. If unavailable, state "author's reasoning unknown" in the report.
## Phase 4: Migrate
Make the learning actionable. Let the lens tailor the output:
- **Mechanical** → copy-paste code skeleton (≤20 lines with TODOs)
- **Intentional** → decision framework (checklist for evaluating tradeoffs)
- **Evolution** → migration path (step-by-step refactor plan)
**Universal deliverables** (all lenses):
- **Can you use this?** Is the design applicable to the user's own projects? If not, why?
- **Steal-it example:** A simplified version (under 20 lines) that captures the core idea. Not production code — a teaching example.
- **Pitfalls:** What context does this design depend on? What would break if you copy it blindly?
**Output:** Migration example + pitfall list (lens-tailored).
**Stall signal:** The design depends on framework internals, language features, or ecosystem the user doesn't have → explain the core idea abstractly instead of providing code.
## Phase 5: Self-review
Check the report is honest:
**All modes:**
- [ ] The design is real (not inferred, not imagined). Evidence: specific files cited.
- [ ] The analysis is deep enough that you could explain it out loud.
- [ ] The migration example captures the core idea, not surface syntax.
- [ ] Pitfalls are specific, not vague ("needs X version" not "may not work everywhere").
**Stall signals (any one → return to relevant phase):**
- Cannot name a file that proves the pattern → back to Phase 2
- Cannot explain why it's better than alternatives → back to Phase 3
- Migration example is over 20 lines → simplify, back to Phase 4
- Lens-specific check failed (e.g., Mechanical missing end-to-end call chain, Intentional missing decision rationale, Evolution missing timeline) → back to relevant phase
**Output:** Essence report with lens annotation.
## Optional: HTML Card
**Only when the user explicitly requests it.**
Generate an HTML visualization card as a shareable deliverable.
### HTML Card Structure (Glassmorphism 2.0 - Essence Variant)
```html
<!DOCTYPE html>
<html lang="zh-CN">
<head>
<meta charset="UTF-8">
<title>{Project Name} - Essence Report</title>
<script src="https://cdn.tailwindcss.com"></script>
<script src="https://cdn.jsdelivr.net/npm/mermaid/dist/mermaid.min.js"></script>
<style>
/* Same glassmorphism styles as /explore */
:root { --glass-bg: rgba(255,255,255,0.4); --primary: #8b5cf6; }
[data-theme="dark"] { --glass-bg: rgba(15,23,42,0.6); --primary: #a78bfa; }
.glass-panel { backdrop-filter: blur(12px); border-radius: 1rem; }
.pattern-diagram { font-family: monospace; background: rgba(0,0,0,0.03); }
</style>
</head>
<body class="p-8">
<nav class="fixed top-4 left-1/2 -translate-x-1/2 w-[90%] max-w-4xl glass-panel z-50 px-6 py-3">
<span class="font-bold text-xl">💎 {Project Name} 精华</span>
<span class="text-sm opacity-70">Lens: {lens} | Pattern: {pattern_name}</span>
</nav>
<main class="max-w-4xl mx-auto mt-24 space-y-6">
<section class="glass-panel p-6">
<h2 class="text-xl font-bold mb-4">🎯 Design Analyzed</h2>
<p>{one-line description}</p>
</section>
<section class="glass-panel p-6">
<h2 class="text-xl font-bold mb-4">🔷 Pattern ({lens})</h2>
<!-- Lens-framed pattern card -->
</section>
<section class="glass-panel p-6">
<h2 class="text-xl font-bold mb-4">🔗 Call Chain</h2>
<pre class="mermaid">{diagram}</pre>
</section>
<section class="glass-panel p-6">
<h2 class="text-xl font-bold mb-4">📦 Migration Example</h2>
<pre class="pattern-diagram"><code>{code_example}</code></pre>
<p class="text-sm opacity-70 mt-2">Pitfalls: {pitfalls}</p>
</section>
</main>
<script>mermaid.initialize({ startOnLoad: true });</script>
</body>
</html>
```
### Output Format
```markdown
### HTML Card Generated
- **Path:** `outputs/{project}-essence.html`
- **Theme:** {modern/ink}
- **Accent Color:** Purple (essence = jewel)
```
**When to skip:** Skip HTML generation unless the user requests it or the analysis is production-critical. When HTML generation fails, deliver a plain-text report instead.
---
## Hard Rules
- **No code evidence = no conclusion.** Every claim about a design must cite a specific file, function, or comment.
- **Under 20 lines for migration examples.** If you can't explain the idea in 20 lines, you don't understand it well enough.
- **Stop after the report.** Do not modify the user's project or the target project.
- **HTML is optional.** Do not block analysis on HTML generation.
## Gotchas
| What happened | Rule |
|---|---|
| 提取的"精华"是 AI 脑补的 | 必须有代码证据(文件 + 行号),不写空泛结论 |
| 用户指定方向但该模块不存在 | 停止并告知用户,不编造替代方向 |
| 项目没有 standout 设计(胶水代码) | 标记"无可提取精华",建议改用 `/explore` |
| Phase 4 迁移示例超过 20 行 | 简化到核心思路,不是复制生产代码 |
| 分析了一个小工具函数 | 工具函数不是设计。设计影响整个架构,工具只解决一个问题 |
| 从 commit message 推断作者意图但没有代码佐证 | Commit message 是辅助证据,必须有代码结构本身的支持 |
| 透镜模式选错导致输出不符预期 | Phase 1 先确认透镜,Mechanical 读代码、Intentional 读文档、Evolution 读历史 |
| 透镜分析流于表面 | 每个透镜有特定输出格式:Mechanical→图 + 接口,Intentional→决策记录,Evolution→时间线 |
| HTML 卡片生成失败 | 降级到纯文本报告,不阻塞分析交付 |
## Outcome
```
Essence Report: {project name}
Lens: mechanical / intentional / evolution
Design analyzed: {one-line description}
Files examined: {count}
Pattern: {pattern name or custom description}
Migration: {steal-it example, ≤20 lines}
HTML generated: yes / no
Status: complete
```
After the report, stop. No modifications. No follow-ups.
@@ -0,0 +1,79 @@
# Essence Detection Signals
How to identify the standout design in a project when the user doesn't specify a direction.
## Signal Strength
A design passes the "essence" threshold if it scores 2+ signals.
### Strong Signals (score = 1 each)
| Signal | How to detect | Example |
|---|---|---|
| **README headline** | Project name is followed by a design claim | "Vite — Next generation frontend tooling with **ESM-first architecture**" |
| **Architecture docs** | Standalone design document exists | `ARCHITECTURE.md`, `docs/design/`, `docs/architecture/` |
| **Official blog post** | Author wrote about the design on their blog | tw93.fun, Vite blog, React blog posts |
| **Community discussion** | Issues/PRs debate the design decision | "Why we chose X over Y" discussions with many comments |
| **Rich code comments** | JSDoc/TSDoc explaining WHY, not WHAT | "We use this pattern because..." with detailed reasoning |
### Objective Signals (score = 1 each, no subjective judgment needed)
| Signal | How to detect | Example |
|---|---|---|
| **Cross-module contract** | A type, interface, or protocol imported across module boundaries (not just files). Go: most-implemented interface. Python: most-subclassed abstract base. Rust: most-implemented trait. | `Plugin` interface implemented by 8 subsystems, each in its own package |
| **File size anomaly** | One file's line count is ≥3× the median for its category (handlers, utils, etc.) | Average handler: 50 lines. One handler: 800 lines with state machine logic |
| **Dedicated test coverage** | Tests exist specifically for this design's edge cases, not just happy paths | `plugin.test.ts` tests plugin resolution, fallback, lifecycle — not just "it loads" |
### Weak Signals (score = 0.5 each)
| Signal | How to detect | Example |
|---|---|---|
| **Unique among competitors** | Same category, different architecture | Next.js uses SSR, Remix uses nested routes — that difference IS the essence |
| **Most-starred files** | GitHub shows stars/bookmarks on specific files | "This file has 200+ stars on GitHub" |
| **Core algorithm** | One file contains non-trivial logic that drives the project | Diff algorithm, compiler pass, state machine |
| **API design** | The public API is notably elegant or unusual | `create()` returns a builder chain, not an object |
## Not Signals
These do NOT count as essence:
- "Clean code" or "well organized" — that's quality, not design
- "Uses TypeScript" — that's a language choice, not architecture
- "Has good tests" — that's engineering discipline, not design
- "Many stars on the repo" — popularity ≠ design quality
- "Uses the latest framework" — following trends ≠ standing out
- Utility functions — even well-written ones are tools, not designs
## Auto-detect Procedure
When the user says "find the essence":
1. **Read README fully.** What is the #1 feature the author leads with? That's a candidate.
2. **Check for design docs.** Is there `ARCHITECTURE.md` or equivalent? That's a candidate.
3. **Scan the import graph.** Which file is imported by the most other files? Use `grep -r "import.*from" src/ | sort | uniq -c | sort -rn` or equivalent. The top result is likely the core.
4. **Check file sizes.** Are any files disproportionately large or small for their apparent role? That signals hidden complexity.
5. **Check uniqueness.** Compare with 1-2 well-known alternatives. What does this project do differently?
6. **Present 1-2 candidates** to the user with evidence. Let them choose or auto-select the strongest.
### Example Output Format
```
Standout designs in {project}:
A) {Design A name} — evidenced by {README claim / file / doc}
What it does: {one sentence}
B) {Design B name} — evidenced by {code comment / unique feature / community discussion}
What it does: {one sentence}
Which should we dive into? (or I can pick the strongest)
```
## Failure Modes
| Situation | Response |
|---|---|
| No signal passes 2+ threshold | "This project uses conventional architecture. Try `/explore` for a full analysis, or pick a more architecturally interesting project." |
| User-specified module doesn't exist | Stop. Do NOT suggest an alternative. Tell the user the path doesn't exist. |
| Project is a wrapper (thin layer over another tool) | "This project is primarily a wrapper around {X}. The design is in {X}, not here. Try analyzing {X} instead." |
| Project is configuration-only (just JSON/YAML files) | "This project has no code architecture. It's configuration-driven. Try `/explore` for a full overview instead." |
+87
View File
@@ -0,0 +1,87 @@
---
name: explore
description: Invoke when you need project-level understanding and an onboarding path. Produces a project learning report for code and non-code repositories with fixed phases for positioning, structure, flow, start path, and core designs. Not for deep code extraction or interactive teaching.
metadata:
version: "0.5.0"
---
# Explore: Project Understanding and Onboarding
Prefix your first line with 🥷 inline, not as its own paragraph.
You are a project cartographer. Your job is to help the user understand what a project is, why it is worth studying, how it is organized, and where to start.
`/explore` is the entry point for first contact with a repository or project-like artifact. It builds global understanding. It does not perform code-level essence extraction and it does not run interactive teaching.
## Project Type Detection
After the initial scan, classify the target before continuing:
| Type | Signals | What changes |
|---|---|---|
| **Code repository** | `go.mod`, `pyproject.toml`, `Cargo.toml`, source directories, executable entrypoints | Run all 4 phases |
| **Skill / docs / knowledge repository** | `SKILL.md`, mostly Markdown, docs-first structure, no runnable application entrypoint | Skip Phase 2 (Flow) and Phase 3 (Start Path) |
| **Template / scaffold repository** | Starter files, minimal logic, setup-first repo | Phase 2 may stay structural and Phase 3 may be minimal |
State the detected type before proceeding. If uncertain, say what evidence is missing and continue with the closest matching type.
## Phase 1: Positioning & Structure
- What this project is, why it is worth studying, and who it is for.
- Top-level structure: main modules, documents, directories, and the likely learning entry area.
- Tradeoffs vs alternatives when evidence exists.
## Phase 2: Flow
**Code repositories only.**
- Skip for non-code and template repositories.
- Trace the main runtime or request flow.
- Produce at least one architecture or core-flow diagram.
- Keep the trace focused on the golden path rather than exhaustive coverage.
## Phase 3: Start Path
**Code repositories only when runnable or meaningfully inspectable.**
- Provide the minimal path to start learning or running the project.
- Give the first command or first inspection step.
- Suggest one safe first modification or observation point when appropriate.
## Phase 4: Core Designs
- Summarize 2-3 core implementations or ideas.
- Keep this at overview depth.
- For each item, include what it is, where it lives, and why it matters.
## Minimum Deliverables
The final `/explore` report must include:
- Project positioning
- Why it is worth studying
- 2-3 core implementations or core ideas
- Tradeoffs or comparisons when applicable
- At least 1 diagram:
- code repository → architecture diagram or core flow diagram
- non-code repository → structure diagram, idea map, or workflow diagram
## Boundary Rules
`/explore` may:
- scan structure
- explain the main flow
- provide a minimal start path
- summarize 2-3 core designs
`/explore` must not:
- perform `/essence`-level deep extraction
- act as `/follow`-style guided teaching
- include Verify, Deep Fission, or HTML Output phases
- preserve no retired lightweight fallback behavior
## Outcome
```
Explore Report: {project name}
Project type: code / skill-docs / template
Phases completed: 4/4 (or note skipped code-only phases)
Diagram included: yes / no
Core designs: 2-3
Status: complete
```
After the report, stop. Do not proceed to `/essence` or `/follow` automatically.
@@ -0,0 +1,98 @@
# Project Analysis Methods
How to read and understand an unfamiliar code project.
## 1. Identify the Entry Point
Every project has a door. Find it first.
### By Language
| Language | Look for |
|---|---|
| **JavaScript/TypeScript** | `package.json` → `main` / `bin` / `scripts.dev` |
| **Python** | `setup.py` → `entry_points`, `pyproject.toml` → `[project.scripts]`, or top-level `app.py` / `main.py` / `__main__.py` |
| **Go** | `package main` in any file, conventionally `main.go` or `cmd/*/main.go` |
| **Rust** | `src/main.rs` or `src/bin/*.rs` |
| **Java** | Class with `public static void main(String[] args)` |
| **C/C++** | `main()` function, conventionally in `src/main.c` |
| **Swift** | `main.swift` or file with `@main` attribute |
### In Frameworks
| Framework | Entry point |
|---|---|
| Next.js | `app/` or `pages/` directory, `next.config.js` |
| React (Vite) | `src/main.tsx` or `src/main.jsx` |
| Vue (Vite) | `src/main.ts` or `src/main.js` |
| Express | File that calls `app.listen()` |
| FastAPI | File that creates `FastAPI()` instance |
| Django | `manage.py`, then project name directory with `urls.py` / `wsgi.py` |
| Flask | `app.py` or `app/__init__.py` |
| Spring Boot | `*Application.java` with `@SpringBootApplication` |
## 2. Judge Project Complexity
Don't over-engineer simple projects. Don't under-analyze complex ones.
### Simple (<50 files, single language)
- Read every source file.
- No need for flow diagrams beyond a simple sequence.
- A light `/explore` pass is probably enough.
### Standard (50-500 files, 1-2 languages)
- Read entry point + core modules + 1-2 feature files.
- Build 1-2 flow diagrams.
- `/explore` is the right level.
### Complex (>500 files, multi-language, monorepo)
- Read entry point + architecture docs + one representative module.
- Use `/essence` to find standout designs, or `/explore` for one package at a time.
- Do NOT try to understand the whole project in one pass.
## 3. Separate Core Code from Scaffolding
Not all files are worth reading.
### Ignore (scaffolding)
- `*.config.js`, `*.config.ts` — configuration, not logic
- `dist/`, `build/`, `out/` — generated output
- `node_modules/`, `vendor/`, `.venv/` — dependencies
- `*.lock`, `yarn.lock`, `go.sum` — lock files
- `LICENSE`, `CODEOWNERS`, `.editorconfig` — project meta
- `test/fixtures/`, `test/data/` — test data
### Read (core)
- Entry point file
- Router/middleware/config handlers
- Model/entity/schema definitions
- Core algorithm or business logic files
- Files referenced most in imports
### Hint: Follow imports
```
entry file → import A → import B → core logic
```
Each import is a dependency. Follow the chain until you hit a file that doesn't import anything else — that's usually the core.
## 4. Read Unfamiliar Framework Code
You don't know every framework. That's fine.
### Strategy
1. **Find the routing layer first.** Every framework has a way to map URLs or events to handlers. Find it. It tells you the project's capabilities.
2. **Follow ONE request end-to-end.** Don't try to understand all routes. Pick the simplest one (often "health check" or "get by ID") and trace it from entry to response.
3. **Identify the framework's conventions.** Most frameworks follow a pattern:
- MVC: Controller → Model → View
- Middleware: Request → Middleware chain → Handler → Response
- Component: Parent renders children, props flow down, events flow up
- Plugin: Core calls hooks, plugins register handlers
4. **Don't fight the framework's abstraction.** If the project uses ORM, don't look for raw SQL. If it uses dependency injection, don't look for `new()` calls. Understand what abstraction layer they chose.
5. **Use the framework's own docs.** If stuck on "how does this framework work?", check the official docs. Don't reverse-engineer what's documented.
@@ -0,0 +1,173 @@
# Flow Pattern Library
Common architecture patterns and how to identify them in code.
## MVC / MVVM / MVX
### What it is
Separation of data (Model), UI/presentation (View), and coordination logic (Controller/ViewModel).
### File signatures
| Pattern | Directories/Files |
|---|---|
| **MVC** | `controllers/`, `models/`, `views/` |
| **MVVM** | `viewmodels/`, `views/`, `models/` |
| **Layered** | `app/`, `domain/`, `infrastructure/` (Clean/Hexagonal) |
### Flow
```
Request → Controller → Model (data) → View (render) → Response
```
### Key question
"Does the file handle data, display, or coordination?" If yes → MVC-family.
---
## Middleware Chain
### What it is
Each handler processes the request and passes it to the next. Like an assembly line.
### File signatures
| Framework | Indicator |
|---|---|---|
| **Express/Koa** | `app.use(...)`, `app.get('/', handler)` |
| **FastAPI** | `@app.middleware("http")`, `Depends()` |
| **Next.js** | `middleware.ts` at root or in `app/` |
| **Gin (Go)** | `router.Use(middleware1, middleware2)` |
| **Koa** | `app.use(async (ctx, next) => { ... })` |
### Flow
```
Request → Middleware A → Middleware B → Handler → Response
↓ ↓
auth check log request
```
### Key question
"Does this function call `next()` or pass control to something else?" If yes → middleware.
### Common middleware order
```
1. CORS / Security headers
2. Logging / Request ID
3. Authentication / Authorization
4. Body parsing / Validation
5. Rate limiting
6. Route handler
7. Error handler (catches everything above)
```
---
## Plugin / Extension System
### What it is
Core provides hooks or interfaces. External code registers handlers. The core doesn't know about specific plugins.
### File signatures
| Pattern | Indicator |
|---|---|
| **Hook-based** | `registerHook('eventName', handler)`, `hooks.on('event', fn)` |
| **Interface-based** | Abstract class or interface that plugins implement |
| **Discovery-based** | Directory scan (`plugins/`), import all, register by convention |
| **VSCode-style** | `contributes` in `package.json`, activation events |
### Flow
```
Core starts
↓
Scans for plugins
↓
Each plugin registers itself
↓
Core fires hooks → plugins respond
↓
Core runs with extended capabilities
```
### Key question
"Can I add functionality without modifying core code?" If yes → plugin architecture.
---
## Event-Driven
### What it is
Components communicate through events, not direct calls. Publishers emit, subscribers listen.
### File signatures
| Pattern | Indicator |
|---|---|
| **Node EventEmitter** | `eventEmitter.on('event', handler)`, `eventEmitter.emit('event', data)` |
| **Pub/Sub** | `pubsub.subscribe('channel', handler)`, `pubsub.publish('channel', data)` |
| **Redux-style** | `dispatch(action)`, `reducer(state, action) → newState` |
| **Observable** | `observable.subscribe(fn)`, `pipe(map, filter)` |
| **Signals (Python)** | `@signal.connect`, `signal.send()` |
### Flow
```
Component A emits "user.created"
↓
Listener B hears it → sends welcome email
Listener C hears it → creates default settings
Listener D hears it → logs analytics
```
### Key question
"Does code communicate without importing or calling each other directly?" If yes → event-driven.
---
## State Management
### What it is
Centralized storage for application state. Components read and update through defined interfaces.
### File signatures
| Pattern | Indicator |
|---|---|
| **Redux** | `createStore()`, `dispatch()`, `useSelector()`, `@reduxjs/toolkit` |
| **Zustand** | `create((set) => ({ ... }))` |
| **Jotai** | `atom(value)`, `useAtom(atom)` |
| **MobX** | `@observable`, `@action`, `@computed` |
| **React Context** | `createContext()`, `useContext()`, `Provider` |
| **Pinia (Vue)** | `defineStore()`, `state`, `actions` |
### Flow
```
Component dispatches action
↓
Reducer processes action + current state
↓
New state emitted
↓
Subscribed components re-render
```
### Key question
"Where does the app store data that multiple components need?" If it's a single store → state management pattern.
---
## Pipeline / Chain of Responsibility
### What it is
Data flows through a series of processors. Each processor transforms the data and passes it on.
### File signatures
| Pattern | Indicator |
|---|---|
| **Stream processing** | `.pipe(transform1).pipe(transform2)` |
| **Compiler/lexer** | Source → Tokenize → Parse → Transform → Generate |
| **Data pipeline** | `input → transform → validate → output` |
| **Makefile** | Target depends on prerequisites, each is a step |
### Flow
```
Raw input → Tokenizer → Parser → Transformer → Generator → Output
```
### Key question
"Does data get progressively transformed through a fixed sequence of steps?" If yes → pipeline.
+101
View File
@@ -0,0 +1,101 @@
---
name: follow
description: Invoke when the user wants an interactive learning session based on an existing `/explore` or `/essence` report. Guides runnable or reader-style follow-along sessions. Not for fresh project analysis or pattern-only extraction.
metadata:
version: "0.5.0"
---
# Follow: Guided Learning Session
Prefix your first line with 🥷 inline, not as its own paragraph.
You are a guide. The user wants to learn from a project step by step with help, context, and correction. You guide the learning process, but you do not replace it.
`/follow` is not a fresh project analyzer. It only works from an existing `/explore` or `/essence` result.
## Pre-check
`/follow` only works when there is already an `/explore` report or an `/essence` report.
- `/explore` report exists → use it as the main learning path
- `/essence` report exists → use it for design-focused guided study
- Neither exists → refuse clearly
Refusal behavior:
"I need an existing `/explore` or `/essence` result before I can guide a follow-along session. Please run `/explore` for project understanding or `/essence` for a focused deep dive first."
Load the existing report before continuing.
## Mode Selection
After the pre-check, select one mode based on the prerequisite report:
- From `/explore` + code repository → default **Runnable**
- From `/explore` + non-code repository → force **Reader**
- From `/essence` → default **Reader** (user is in design-analysis state)
| Mode | When | Entry |
|---|---|---|
| **Runnable** | Report confirms the project is a runnable code repository and the user wants to learn by running and changing it | Start from environment and first execution |
| **Reader** | Project has no runtime, or the user is studying design/architecture, or the prerequisite report is from `/essence` | Start from guided reading |
State the selected mode before proceeding. Do not re-scan the project — use the prerequisite report to decide.
## Teaching Interaction Rules
`/follow` must teach by guidance, not by dumping answers:
- explain the purpose of the current step first
- give the user an observation point or action point
- ask the user to predict, try, or explain before revealing the answer
- then reveal, correct, or deepen the explanation
- never say "go read the code" as a standalone instruction. When referencing code, always start with: what design idea this code embodies, why it matters in the overall architecture, and what the user should pay attention to
## Runnable Check
Before Runnable mode, confirm from the **prerequisite report** (do not re-scan the project):
- If the report identified the target as a code repository with a recognized runtime (`go.mod`, `pyproject.toml`, `Cargo.toml`, `Makefile`, `build.gradle`, `pom.xml`, `CMakeLists.txt`, etc.), proceed with Runnable.
- If the report classified it as non-code, or no runtime entrypoint was found, switch to Reader and explain why.
- If the prerequisite is `/essence`, confirm with the user: essence is design-focused, Reader is the natural fit. Allow Runnable only if the user explicitly insists.
- Do not introduce a third mode.
## Runnable Mode Flow
1. Confirm environment and prerequisites.
2. Let the user run the project.
3. Let the user make one safe change.
4. Walk the main flow together.
5. Give one small exercise.
6. Review what they learned.
## Reader Mode Flow
1. Frame the learning goal around a core design or architectural idea, not a single file.
2. Walk through the design concept layer by layer: problem → approach → implementation → tradeoff.
3. Ask the user questions that probe understanding ("Why did the author choose this approach over a simpler one?"), not just prediction ("What happens next?").
4. Use diagrams or structured summaries to connect the dots between files and design ideas.
5. Give one reasoning exercise that tests whether the user can apply the design pattern elsewhere.
6. Review what they learned.
## Boundary Rules
`/follow` must:
- depend on `/explore` or `/essence`
- guide the user interactively
- adapt between code and non-code repositories through Runnable or Reader emphasis
`/follow` must not:
- rescan the whole project as a new analyzer
- reference retired skills as prerequisites
- add any third learning mode
- execute commands or write code for the user
## Outcome
```
Follow Session: {project name}
Mode: runnable / reader
Prerequisite report: /explore or /essence
Exercise result: completed / partial / too hard
Next direction: {suggested follow-up}
Status: complete
```
After the review, stop. Ask whether the user wants another exercise or wants to end the session.
@@ -0,0 +1,113 @@
# Environment Detection Rules
How to detect the runtime environment and guide the user through setup in `/follow`.
## Language Detection from Config
Check these files in order. The first match is the primary language.
| Config file | Language | Runtime check | Install command |
|---|---|---|---|
| `package.json` | JavaScript/TypeScript | `node --version` | nvm or official installer |
| `pyproject.toml` | Python | `python --version` | pyenv or python.org |
| `go.mod` | Go | `go version` | golang.org/dl |
| `Cargo.toml` | Rust | `rustc --version` | rustup |
| `pom.xml` | Java | `java -version` | SDKMAN or official |
| `build.gradle` / `build.gradle.kts` | Java/Kotlin | `java -version` | SDKMAN |
| `Gemfile` | Ruby | `ruby --version` | rvm or rbenv |
| `*.csproj` | C#/.NET | `dotnet --version` | .NET SDK |
| `CMakeLists.txt` | C/C++ | `gcc --version` or `clang --version` | System package manager |
| `swift package.json` | Swift | `swift --version` | Xcode or swift.org |
## Dependency Installation
Once language is detected, guide the user:
### JavaScript/TypeScript
```bash
# Check which package manager is used
if [ -f "yarn.lock" ]; then yarn install
elif [ -f "pnpm-lock.yaml" ]; then pnpm install
elif [ -f "bun.lockb" ] || [ -f "bun.lock" ]; then bun install
else npm install
fi
```
### Python
```bash
# Modern Python projects
pip install -e .
# Or with requirements
pip install -r requirements.txt
# Or with poetry
poetry install
# Or with uv
uv pip install -r requirements.txt
```
### Go
```bash
go mod download
```
### Rust
```bash
cargo build
```
### Java (Maven)
```bash
mvn install
```
### Java (Gradle)
```bash
./gradlew build
# or
gradle build
```
## Run Command Detection
How to start the project:
| Source | Command |
|---|---|
| `package.json` → `scripts.dev` | `npm run dev` |
| `package.json` → `scripts.start` | `npm start` |
| `Makefile` → `dev` target | `make dev` |
| `Makefile` → `run` target | `make run` |
| `pyproject.toml` (Poetry) | `poetry run python main.py` |
| `go.mod` → `package main` | `go run main.go` |
| `Cargo.toml` → `[[bin]]` | `cargo run` |
| `docker-compose.yml` exists | `docker-compose up` |
| `Dockerfile` exists, no compose | `docker build -t app . && docker run app` |
## Common Environment Issues
| Error | Cause | Fix |
|---|---|---|
| `command not found: node` | Node.js not installed | Install Node.js (recommend LTS) |
| `ModuleNotFoundError` | Python deps not installed | Run `pip install -r requirements.txt` |
| `EACCES: permission denied` | Global install without sudo | Use nvm/fnm, or prefix with sudo |
| `ENOENT: no such file` | Wrong working directory | `cd` to project root first |
| `port already in use` | Another process on same port | Kill the process or use different port |
| `go: cannot find main module` | Outside Go module | `cd` to directory with `go.mod` |
| `error: could not find Cargo.toml` | Outside Rust project | `cd` to directory with `Cargo.toml` |
| `java.lang.UnsupportedClassVersionError` | Wrong Java version | Match JDK version to project requirement |
| `npm ERR! code ERESOLVE` | Dependency conflict | Try `npm install --legacy-peer-deps` |
## Detection Script for /follow
```bash
# Quick environment check
echo "=== Environment ==="
node --version 2>/dev/null || echo "Node.js: not installed"
python --version 2>/dev/null || echo "Python: not installed"
go version 2>/dev/null || echo "Go: not installed"
rustc --version 2>/dev/null || echo "Rust: not installed"
java -version 2>/dev/null || echo "Java: not installed"
echo "PWD: $(pwd)"
```
Run this at the start of `/follow` Step 1 to understand what's available.
+57
View File
@@ -0,0 +1,57 @@
# Frontend Design — Complete Guidance
This document provides a comprehensive framework for creating visually distinctive, non-templated UI designs. Here's the full breakdown:
## Foundational Approach
Act as the design lead for a studio known for unique client identities — the client has already turned down template-like proposals. Every choice about palette, typography, and layout must be specific to the brief, including "one real aesthetic risk you can justify."
## Grounding in Subject Matter
If the brief is vague about the product or subject, pin it down yourself: name the subject, its audience, and the page's single job. Draw inspiration from "the subject's own world, its materials, instruments, artifacts, and vernacular." Use any known context about the human's preferences or past designs as hints.
## Design Principles
- **Hero as thesis**: Open with "the most characteristic thing in the subject's world" — avoid default choices like a big number with a small label and gradient accent unless truly optimal.
- **Typography**: Pair display and body faces deliberately, not from your usual repertoire. Set a clear type scale with intentional weights, widths, and spacing. "Make the type treatment itself a memorable part of the design."
- **Structure as information**: Numbering, eyebrows, dividers must encode something true about the content. Question whether numbered markers (01/02/03) actually make sense before using them — only appropriate for real sequences.
- **Motion**: Consider where animation serves the subject. "An orchestrated moment usually lands harder than scattered effects." Sometimes less is better to avoid an AI-generated feel.
- **Complexity**: Match execution to the vision — maximalist needs elaborate execution, minimal needs precision.
- **Content**: Come up with copy if the brief lacks it. Poor copy makes a design feel as templated as poor layout.
## AI-Generated Design Traps
Three common AI-default looks to watch for: (1) warm cream background (~#F4F1EA) with serif display and terracotta accent; (2) near-black with bright acid-green or vermilion; (3) broadsheet layout with hairline rules, zero border-radius, and dense columns. "All three are legitimate for some briefs, but they are defaults rather than choices." Where the brief leaves an axis free, don't spend that freedom on a default.
## Two-Pass Process
**Pass 1 — Plan**: Create a compact token system:
1. **Color**: 4–6 named hex values
2. **Type**: Characterful display face (used with restraint), complementary body face, utility face for captions/data
3. **Layout**: One-sentence prose descriptions + ASCII wireframes
4. **Signature**: The single unique element the page will be remembered by
Review the plan against the brief. If any part reads like what you'd produce for any similar page, revise it. Only then write code.
**Pass 2 — Build**: Follow the revised plan exactly. Watch for CSS selector specificity conflicts (e.g., `.section` and `.cta` fighting over padding/margins). Do most planning internally, only sharing ideas when confident.
## Restraint & Self-Critique
"Spend your boldness in one place" — let the signature element be the one memorable thing; keep everything else quiet. "Not taking a risk can be a risk itself!" Build responsively down to mobile, with visible keyboard focus and reduced motion respected. Critique as you build. Follow Chanel's advice: before finishing, remove one accessory. Jot notes about what you've tried to avoid repeating yourself.
## Writing in Design
Words exist to make the design understandable and usable — they're "design material, not decoration." Write from the end user's perspective, naming things by what people control and recognize, never by how the system is built.
- Use active voice as default
- A control should say exactly what happens: "Save changes," not "Submit"
- Maintain consistent vocabulary throughout flows (button says "Publish," toast says "Published")
- Treat errors as guidance, not mood — explain what went wrong and how to fix it
- Empty screens are invitations to act
- Keep the register conversational: "plain verbs, sentence case, no filler"
- Let each element do exactly one job — "a label labels, an example demonstrates"
## License
Apache License 2.0 — see LICENSE.txt
@@ -0,0 +1,83 @@
---
name: gitnexus-cli
description: "Use when the user needs to run GitNexus CLI commands like analyze/index a repo, check status, clean the index, generate a wiki, or list indexed repos. Examples: \"Index this repo\", \"Reanalyze the codebase\", \"Generate a wiki\""
---
# GitNexus CLI Commands
All commands work via `npx` — no global install required.
## Commands
### analyze — Build or refresh the index
```bash
npx gitnexus analyze
```
Run from the project root. This parses all source files, builds the knowledge graph, writes it to `.gitnexus/`, and generates AGENTS.md / AGENTS.md context files.
| Flag | Effect |
| -------------- | ---------------------------------------------------------------- |
| `--force` | Force full re-index even if up to date |
| `--embeddings` | Enable embedding generation for semantic search (off by default) |
| `--drop-embeddings` | Drop existing embeddings on rebuild. By default, an `analyze` without `--embeddings` preserves them. |
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale. In Codex, a PostToolUse hook detects staleness after `git commit` and `git merge` and notifies the agent to run `analyze` — the hook does not run analyze itself, to avoid blocking the agent for up to 120s and risking KuzuDB corruption on timeout.
### status — Check index freshness
```bash
npx gitnexus status
```
Shows whether the current repo has a GitNexus index, when it was last updated, and symbol/relationship counts. Use this to check if re-indexing is needed.
### clean — Delete the index
```bash
npx gitnexus clean
```
Deletes the `.gitnexus/` directory and unregisters the repo from the global registry. Use before re-indexing if the index is corrupt or after removing GitNexus from a project.
| Flag | Effect |
| --------- | ------------------------------------------------- |
| `--force` | Skip confirmation prompt |
| `--all` | Clean all indexed repos, not just the current one |
### wiki — Generate documentation from the graph
```bash
npx gitnexus wiki
```
Generates repository documentation from the knowledge graph using an LLM. Requires an API key (saved to `~/.gitnexus/config.json` on first use).
| Flag | Effect |
| ------------------- | ----------------------------------------- |
| `--force` | Force full regeneration |
| `--model <model>` | LLM model (default: minimax/minimax-m2.5) |
| `--base-url <url>` | LLM API base URL |
| `--api-key <key>` | LLM API key |
| `--concurrency <n>` | Parallel LLM calls (default: 3) |
| `--gist` | Publish wiki as a public GitHub Gist |
### list — Show all indexed repos
```bash
npx gitnexus list
```
Lists all repositories registered in `~/.gitnexus/registry.json`. The MCP `list_repos` tool provides the same information.
## After Indexing
1. **Read `gitnexus://repo/{name}/context`** to verify the index loaded
2. Use the other GitNexus skills (`exploring`, `debugging`, `impact-analysis`, `refactoring`) for your task
## Troubleshooting
- **"Not inside a git repository"**: Run from a directory inside a git repo
- **Index is stale after re-analyzing**: Restart Codex to reload the MCP server
- **Embeddings slow**: Omit `--embeddings` (it's off by default) or set `OPENAI_API_KEY` for faster API-based embedding
@@ -0,0 +1,89 @@
---
name: gitnexus-debugging
description: "Use when the user is debugging a bug, tracing an error, or asking why something fails. Examples: \"Why is X failing?\", \"Where does this error come from?\", \"Trace this bug\""
---
# Debugging with GitNexus
## When to Use
- "Why is this function failing?"
- "Trace where this error comes from"
- "Who calls this method?"
- "This endpoint returns 500"
- Investigating bugs, errors, or unexpected behavior
## Workflow
```
1. gitnexus_query({query: "<error or symptom>"}) → Find related execution flows
2. gitnexus_context({name: "<suspect>"}) → See callers/callees/processes
3. READ gitnexus://repo/{name}/process/{name} → Trace execution flow
4. gitnexus_cypher({query: "MATCH path..."}) → Custom traces if needed
```
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
## Checklist
```
- [ ] Understand the symptom (error message, unexpected behavior)
- [ ] gitnexus_query for error text or related code
- [ ] Identify the suspect function from returned processes
- [ ] gitnexus_context to see callers and callees
- [ ] Trace execution flow via process resource if applicable
- [ ] gitnexus_cypher for custom call chain traces if needed
- [ ] Read source files to confirm root cause
```
## Debugging Patterns
| Symptom | GitNexus Approach |
| -------------------- | ---------------------------------------------------------- |
| Error message | `gitnexus_query` for error text → `context` on throw sites |
| Wrong return value | `context` on the function → trace callees for data flow |
| Intermittent failure | `context` → look for external calls, async deps |
| Performance issue | `context` → find symbols with many callers (hot paths) |
| Recent regression | `detect_changes` to see what your changes affect |
## Tools
**gitnexus_query** — find code related to error:
```
gitnexus_query({query: "payment validation error"})
→ Processes: CheckoutFlow, ErrorHandling
→ Symbols: validatePayment, handlePaymentError, PaymentException
```
**gitnexus_context** — full context for a suspect:
```
gitnexus_context({name: "validatePayment"})
→ Incoming calls: processCheckout, webhookHandler
→ Outgoing calls: verifyCard, fetchRates (external API!)
→ Processes: CheckoutFlow (step 3/7)
```
**gitnexus_cypher** — custom call chain traces:
```cypher
MATCH path = (a)-[:CodeRelation {type: 'CALLS'}*1..2]->(b:Function {name: "validatePayment"})
RETURN [n IN nodes(path) | n.name] AS chain
```
## Example: "Payment endpoint returns 500 intermittently"
```
1. gitnexus_query({query: "payment error handling"})
→ Processes: CheckoutFlow, ErrorHandling
→ Symbols: validatePayment, handlePaymentError
2. gitnexus_context({name: "validatePayment"})
→ Outgoing calls: verifyCard, fetchRates (external API!)
3. READ gitnexus://repo/my-app/process/CheckoutFlow
→ Step 3: validatePayment → calls fetchRates (external)
4. Root cause: fetchRates calls external API without proper timeout
```
@@ -0,0 +1,78 @@
---
name: gitnexus-exploring
description: "Use when the user asks how code works, wants to understand architecture, trace execution flows, or explore unfamiliar parts of the codebase. Examples: \"How does X work?\", \"What calls this function?\", \"Show me the auth flow\""
---
# Exploring Codebases with GitNexus
## When to Use
- "How does authentication work?"
- "What's the project structure?"
- "Show me the main components"
- "Where is the database logic?"
- Understanding code you haven't seen before
## Workflow
```
1. READ gitnexus://repos → Discover indexed repos
2. READ gitnexus://repo/{name}/context → Codebase overview, check staleness
3. gitnexus_query({query: "<what you want to understand>"}) → Find related execution flows
4. gitnexus_context({name: "<symbol>"}) → Deep dive on specific symbol
5. READ gitnexus://repo/{name}/process/{name} → Trace full execution flow
```
> If step 2 says "Index is stale" → run `npx gitnexus analyze` in terminal.
## Checklist
```
- [ ] READ gitnexus://repo/{name}/context
- [ ] gitnexus_query for the concept you want to understand
- [ ] Review returned processes (execution flows)
- [ ] gitnexus_context on key symbols for callers/callees
- [ ] READ process resource for full execution traces
- [ ] Read source files for implementation details
```
## Resources
| Resource | What you get |
| --------------------------------------- | ------------------------------------------------------- |
| `gitnexus://repo/{name}/context` | Stats, staleness warning (~150 tokens) |
| `gitnexus://repo/{name}/clusters` | All functional areas with cohesion scores (~300 tokens) |
| `gitnexus://repo/{name}/cluster/{name}` | Area members with file paths (~500 tokens) |
| `gitnexus://repo/{name}/process/{name}` | Step-by-step execution trace (~200 tokens) |
## Tools
**gitnexus_query** — find execution flows related to a concept:
```
gitnexus_query({query: "payment processing"})
→ Processes: CheckoutFlow, RefundFlow, WebhookHandler
→ Symbols grouped by flow with file locations
```
**gitnexus_context** — 360-degree view of a symbol:
```
gitnexus_context({name: "validateUser"})
→ Incoming calls: loginHandler, apiMiddleware
→ Outgoing calls: checkToken, getUserById
→ Processes: LoginFlow (step 2/5), TokenRefresh (step 1/3)
```
## Example: "How does payment processing work?"
```
1. READ gitnexus://repo/my-app/context → 918 symbols, 45 processes
2. gitnexus_query({query: "payment processing"})
→ CheckoutFlow: processPayment → validateCard → chargeStripe
→ RefundFlow: initiateRefund → calculateRefund → processRefund
3. gitnexus_context({name: "processPayment"})
→ Incoming: checkoutHandler, webhookHandler
→ Outgoing: validateCard, chargeStripe, saveTransaction
4. Read src/payments/processor.ts for implementation details
```
@@ -0,0 +1,64 @@
---
name: gitnexus-guide
description: "Use when the user asks about GitNexus itself — available tools, how to query the knowledge graph, MCP resources, graph schema, or workflow reference. Examples: \"What GitNexus tools are available?\", \"How do I use GitNexus?\""
---
# GitNexus Guide
Quick reference for all GitNexus MCP tools, resources, and the knowledge graph schema.
## Always Start Here
For any task involving code understanding, debugging, impact analysis, or refactoring:
1. **Read `gitnexus://repo/{name}/context`** — codebase overview + check index freshness
2. **Match your task to a skill below** and **read that skill file**
3. **Follow the skill's workflow and checklist**
> If step 1 warns the index is stale, run `npx gitnexus analyze` in the terminal first.
## Skills
| Task | Skill to read |
| -------------------------------------------- | ------------------- |
| Understand architecture / "How does X work?" | `gitnexus-exploring` |
| Blast radius / "What breaks if I change X?" | `gitnexus-impact-analysis` |
| Trace bugs / "Why is X failing?" | `gitnexus-debugging` |
| Rename / extract / split / refactor | `gitnexus-refactoring` |
| Tools, resources, schema reference | `gitnexus-guide` (this file) |
| Index, status, clean, wiki CLI commands | `gitnexus-cli` |
## Tools Reference
| Tool | What it gives you |
| ---------------- | ------------------------------------------------------------------------ |
| `query` | Process-grouped code intelligence — execution flows related to a concept |
| `context` | 360-degree symbol view — categorized refs, processes it participates in |
| `impact` | Symbol blast radius — what breaks at depth 1/2/3 with confidence |
| `detect_changes` | Git-diff impact — what do your current changes affect |
| `rename` | Multi-file coordinated rename with confidence-tagged edits |
| `cypher` | Raw graph queries (read `gitnexus://repo/{name}/schema` first) |
| `list_repos` | Discover indexed repos |
## Resources Reference
Lightweight reads (~100-500 tokens) for navigation:
| Resource | Content |
| ---------------------------------------------- | ----------------------------------------- |
| `gitnexus://repo/{name}/context` | Stats, staleness check |
| `gitnexus://repo/{name}/clusters` | All functional areas with cohesion scores |
| `gitnexus://repo/{name}/cluster/{clusterName}` | Area members |
| `gitnexus://repo/{name}/processes` | All execution flows |
| `gitnexus://repo/{name}/process/{processName}` | Step-by-step trace |
| `gitnexus://repo/{name}/schema` | Graph schema for Cypher |
## Graph Schema
**Nodes:** File, Function, Class, Interface, Method, Community, Process
**Edges (via CodeRelation.type):** CALLS, IMPORTS, EXTENDS, IMPLEMENTS, DEFINES, MEMBER_OF, STEP_IN_PROCESS
```cypher
MATCH (caller)-[:CodeRelation {type: 'CALLS'}]->(f:Function {name: "myFunc"})
RETURN caller.name, caller.filePath
```
@@ -0,0 +1,97 @@
---
name: gitnexus-impact-analysis
description: "Use when the user wants to know what will break if they change something, or needs safety analysis before editing code. Examples: \"Is it safe to change X?\", \"What depends on this?\", \"What will break?\""
---
# Impact Analysis with GitNexus
## When to Use
- "Is it safe to change this function?"
- "What will break if I modify X?"
- "Show me the blast radius"
- "Who uses this code?"
- Before making non-trivial code changes
- Before committing — to understand what your changes affect
## Workflow
```
1. gitnexus_impact({target: "X", direction: "upstream"}) → What depends on this
2. READ gitnexus://repo/{name}/processes → Check affected execution flows
3. gitnexus_detect_changes() → Map current git changes to affected flows
4. Assess risk and report to user
```
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
## Checklist
```
- [ ] gitnexus_impact({target, direction: "upstream"}) to find dependents
- [ ] Review d=1 items first (these WILL BREAK)
- [ ] Check high-confidence (>0.8) dependencies
- [ ] READ processes to check affected execution flows
- [ ] gitnexus_detect_changes() for pre-commit check
- [ ] Assess risk level and report to user
```
## Understanding Output
| Depth | Risk Level | Meaning |
| ----- | ---------------- | ------------------------ |
| d=1 | **WILL BREAK** | Direct callers/importers |
| d=2 | LIKELY AFFECTED | Indirect dependencies |
| d=3 | MAY NEED TESTING | Transitive effects |
## Risk Assessment
| Affected | Risk |
| ------------------------------ | -------- |
| <5 symbols, few processes | LOW |
| 5-15 symbols, 2-5 processes | MEDIUM |
| >15 symbols or many processes | HIGH |
| Critical path (auth, payments) | CRITICAL |
## Tools
**gitnexus_impact** — the primary tool for symbol blast radius:
```
gitnexus_impact({
target: "validateUser",
direction: "upstream",
minConfidence: 0.8,
maxDepth: 3
})
→ d=1 (WILL BREAK):
- loginHandler (src/auth/login.ts:42) [CALLS, 100%]
- apiMiddleware (src/api/middleware.ts:15) [CALLS, 100%]
→ d=2 (LIKELY AFFECTED):
- authRouter (src/routes/auth.ts:22) [CALLS, 95%]
```
**gitnexus_detect_changes** — git-diff based impact analysis:
```
gitnexus_detect_changes({scope: "staged"})
→ Changed: 5 symbols in 3 files
→ Affected: LoginFlow, TokenRefresh, APIMiddlewarePipeline
→ Risk: MEDIUM
```
## Example: "What breaks if I change validateUser?"
```
1. gitnexus_impact({target: "validateUser", direction: "upstream"})
→ d=1: loginHandler, apiMiddleware (WILL BREAK)
→ d=2: authRouter, sessionManager (LIKELY AFFECTED)
2. READ gitnexus://repo/my-app/processes
→ LoginFlow and TokenRefresh touch validateUser
3. Risk: 2 direct callers, 2 processes = MEDIUM
```
@@ -0,0 +1,121 @@
---
name: gitnexus-refactoring
description: "Use when the user wants to rename, extract, split, move, or restructure code safely. Examples: \"Rename this function\", \"Extract this into a module\", \"Refactor this class\", \"Move this to a separate file\""
---
# Refactoring with GitNexus
## When to Use
- "Rename this function safely"
- "Extract this into a module"
- "Split this service"
- "Move this to a new file"
- Any task involving renaming, extracting, splitting, or restructuring code
## Workflow
```
1. gitnexus_impact({target: "X", direction: "upstream"}) → Map all dependents
2. gitnexus_query({query: "X"}) → Find execution flows involving X
3. gitnexus_context({name: "X"}) → See all incoming/outgoing refs
4. Plan update order: interfaces → implementations → callers → tests
```
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
## Checklists
### Rename Symbol
```
- [ ] gitnexus_rename({symbol_name: "oldName", new_name: "newName", dry_run: true}) — preview all edits
- [ ] Review graph edits (high confidence) and ast_search edits (review carefully)
- [ ] If satisfied: gitnexus_rename({..., dry_run: false}) — apply edits
- [ ] gitnexus_detect_changes() — verify only expected files changed
- [ ] Run tests for affected processes
```
### Extract Module
```
- [ ] gitnexus_context({name: target}) — see all incoming/outgoing refs
- [ ] gitnexus_impact({target, direction: "upstream"}) — find all external callers
- [ ] Define new module interface
- [ ] Extract code, update imports
- [ ] gitnexus_detect_changes() — verify affected scope
- [ ] Run tests for affected processes
```
### Split Function/Service
```
- [ ] gitnexus_context({name: target}) — understand all callees
- [ ] Group callees by responsibility
- [ ] gitnexus_impact({target, direction: "upstream"}) — map callers to update
- [ ] Create new functions/services
- [ ] Update callers
- [ ] gitnexus_detect_changes() — verify affected scope
- [ ] Run tests for affected processes
```
## Tools
**gitnexus_rename** — automated multi-file rename:
```
gitnexus_rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: true})
→ 12 edits across 8 files
→ 10 graph edits (high confidence), 2 ast_search edits (review)
→ Changes: [{file_path, edits: [{line, old_text, new_text, confidence}]}]
```
**gitnexus_impact** — map all dependents first:
```
gitnexus_impact({target: "validateUser", direction: "upstream"})
→ d=1: loginHandler, apiMiddleware, testUtils
→ Affected Processes: LoginFlow, TokenRefresh
```
**gitnexus_detect_changes** — verify your changes after refactoring:
```
gitnexus_detect_changes({scope: "all"})
→ Changed: 8 files, 12 symbols
→ Affected processes: LoginFlow, TokenRefresh
→ Risk: MEDIUM
```
**gitnexus_cypher** — custom reference queries:
```cypher
MATCH (caller)-[:CodeRelation {type: 'CALLS'}]->(f:Function {name: "validateUser"})
RETURN caller.name, caller.filePath ORDER BY caller.filePath
```
## Risk Rules
| Risk Factor | Mitigation |
| ------------------- | ----------------------------------------- |
| Many callers (>5) | Use gitnexus_rename for automated updates |
| Cross-area refs | Use detect_changes after to verify scope |
| String/dynamic refs | gitnexus_query to find them |
| External/public API | Version and deprecate properly |
## Example: Rename `validateUser` to `authenticateUser`
```
1. gitnexus_rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: true})
→ 12 edits: 10 graph (safe), 2 ast_search (review)
→ Files: validator.ts, login.ts, middleware.ts, config.json...
2. Review ast_search edits (config.json: dynamic reference!)
3. gitnexus_rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: false})
→ Applied 12 edits across 8 files
4. gitnexus_detect_changes({scope: "all"})
→ Affected: LoginFlow, TokenRefresh
→ Risk: MEDIUM — run tests for these flows
```
+16
View File
@@ -0,0 +1,16 @@
---
name: handoff
description: Compact the current conversation into a handoff document for another agent to pick up.
argument-hint: "What will the next session be used for?"
disable-model-invocation: true
---
Write a handoff document summarising the current conversation so a fresh agent can continue the work. Save to the temporary directory of the user's OS - not the current workspace.
Include a "suggested skills" section in the document, which suggests skills that the agent should invoke.
Do not duplicate content already captured in other artifacts (PRDs, plans, ADRs, issues, commits, diffs). Reference them by path or URL instead.
Redact any sensitive information, such as API keys, passwords, or personally identifiable information.
If the user passed arguments, treat them as a description of what the next session will focus on and tailor the doc accordingly.
@@ -0,0 +1,156 @@
---
name: openspec-apply-change
description: Implement tasks from an OpenSpec change. Use when the user wants to start implementing, continue implementation, or work through tasks.
license: MIT
compatibility: Requires openspec CLI.
metadata:
author: openspec
version: "1.0"
generatedBy: "1.3.1"
---
Implement tasks from an OpenSpec change.
**Input**: Optionally specify a change name. If omitted, check if it can be inferred from conversation context. If vague or ambiguous you MUST prompt for available changes.
**Steps**
1. **Select the change**
If a name is provided, use it. Otherwise:
- Infer from conversation context if the user mentioned a change
- Auto-select if only one active change exists
- If ambiguous, run `openspec list --json` to get available changes and use the **AskUserQuestion tool** to let the user select
Always announce: "Using change: <name>" and how to override (e.g., `/opsx:apply <other>`).
2. **Check status to understand the schema**
```bash
openspec status --change "<name>" --json
```
Parse the JSON to understand:
- `schemaName`: The workflow being used (e.g., "spec-driven")
- Which artifact contains the tasks (typically "tasks" for spec-driven, check status for others)
3. **Get apply instructions**
```bash
openspec instructions apply --change "<name>" --json
```
This returns:
- `contextFiles`: artifact ID -> array of concrete file paths (varies by schema - could be proposal/specs/design/tasks or spec/tests/implementation/docs)
- Progress (total, complete, remaining)
- Task list with status
- Dynamic instruction based on current state
**Handle states:**
- If `state: "blocked"` (missing artifacts): show message, suggest using openspec-continue-change
- If `state: "all_done"`: congratulate, suggest archive
- Otherwise: proceed to implementation
4. **Read context files**
Read every file path listed under `contextFiles` from the apply instructions output.
The files depend on the schema being used:
- **spec-driven**: proposal, specs, design, tasks
- Other schemas: follow the contextFiles from CLI output
5. **Show current progress**
Display:
- Schema being used
- Progress: "N/M tasks complete"
- Remaining tasks overview
- Dynamic instruction from CLI
6. **Implement tasks (loop until done or blocked)**
For each pending task:
- Show which task is being worked on
- Make the code changes required
- Keep changes minimal and focused
- Mark task complete in the tasks file: `- [ ]` → `- [x]`
- Continue to next task
**Pause if:**
- Task is unclear → ask for clarification
- Implementation reveals a design issue → suggest updating artifacts
- Error or blocker encountered → report and wait for guidance
- User interrupts
7. **On completion or pause, show status**
Display:
- Tasks completed this session
- Overall progress: "N/M tasks complete"
- If all done: suggest archive
- If paused: explain why and wait for guidance
**Output During Implementation**
```
## Implementing: <change-name> (schema: <schema-name>)
Working on task 3/7: <task description>
[...implementation happening...]
✓ Task complete
Working on task 4/7: <task description>
[...implementation happening...]
✓ Task complete
```
**Output On Completion**
```
## Implementation Complete
**Change:** <change-name>
**Schema:** <schema-name>
**Progress:** 7/7 tasks complete ✓
### Completed This Session
- [x] Task 1
- [x] Task 2
...
All tasks complete! Ready to archive this change.
```
**Output On Pause (Issue Encountered)**
```
## Implementation Paused
**Change:** <change-name>
**Schema:** <schema-name>
**Progress:** 4/7 tasks complete
### Issue Encountered
<description of the issue>
**Options:**
1. <option 1>
2. <option 2>
3. Other approach
What would you like to do?
```
**Guardrails**
- Keep going through tasks until done or blocked
- Always read context files before starting (from the apply instructions output)
- If task is ambiguous, pause and ask before implementing
- If implementation reveals issues, pause and suggest artifact updates
- Keep code changes minimal and scoped to each task
- Update task checkbox immediately after completing each task
- Pause on errors, blockers, or unclear requirements - don't guess
- Use contextFiles from CLI output, don't assume specific file names
**Fluid Workflow Integration**
This skill supports the "actions on a change" model:
- **Can be invoked anytime**: Before all artifacts are done (if tasks exist), after partial implementation, interleaved with other actions
- **Allows artifact updates**: If implementation reveals design issues, suggest updating artifacts - not phase-locked, work fluidly
@@ -0,0 +1,114 @@
---
name: openspec-archive-change
description: Archive a completed change in the experimental workflow. Use when the user wants to finalize and archive a change after implementation is complete.
license: MIT
compatibility: Requires openspec CLI.
metadata:
author: openspec
version: "1.0"
generatedBy: "1.3.1"
---
Archive a completed change in the experimental workflow.
**Input**: Optionally specify a change name. If omitted, check if it can be inferred from conversation context. If vague or ambiguous you MUST prompt for available changes.
**Steps**
1. **If no change name provided, prompt for selection**
Run `openspec list --json` to get available changes. Use the **AskUserQuestion tool** to let the user select.
Show only active changes (not already archived).
Include the schema used for each change if available.
**IMPORTANT**: Do NOT guess or auto-select a change. Always let the user choose.
2. **Check artifact completion status**
Run `openspec status --change "<name>" --json` to check artifact completion.
Parse the JSON to understand:
- `schemaName`: The workflow being used
- `artifacts`: List of artifacts with their status (`done` or other)
**If any artifacts are not `done`:**
- Display warning listing incomplete artifacts
- Use **AskUserQuestion tool** to confirm user wants to proceed
- Proceed if user confirms
3. **Check task completion status**
Read the tasks file (typically `tasks.md`) to check for incomplete tasks.
Count tasks marked with `- [ ]` (incomplete) vs `- [x]` (complete).
**If incomplete tasks found:**
- Display warning showing count of incomplete tasks
- Use **AskUserQuestion tool** to confirm user wants to proceed
- Proceed if user confirms
**If no tasks file exists:** Proceed without task-related warning.
4. **Assess delta spec sync state**
Check for delta specs at `openspec/changes/<name>/specs/`. If none exist, proceed without sync prompt.
**If delta specs exist:**
- Compare each delta spec with its corresponding main spec at `openspec/specs/<capability>/spec.md`
- Determine what changes would be applied (adds, modifications, removals, renames)
- Show a combined summary before prompting
**Prompt options:**
- If changes needed: "Sync now (recommended)", "Archive without syncing"
- If already synced: "Archive now", "Sync anyway", "Cancel"
If user chooses sync, use Task tool (subagent_type: "general-purpose", prompt: "Use Skill tool to invoke openspec-sync-specs for change '<name>'. Delta spec analysis: <include the analyzed delta spec summary>"). Proceed to archive regardless of choice.
5. **Perform the archive**
Create the archive directory if it doesn't exist:
```bash
mkdir -p openspec/changes/archive
```
Generate target name using current date: `YYYY-MM-DD-<change-name>`
**Check if target already exists:**
- If yes: Fail with error, suggest renaming existing archive or using different date
- If no: Move the change directory to archive
```bash
mv openspec/changes/<name> openspec/changes/archive/YYYY-MM-DD-<name>
```
6. **Display summary**
Show archive completion summary including:
- Change name
- Schema that was used
- Archive location
- Whether specs were synced (if applicable)
- Note about any warnings (incomplete artifacts/tasks)
**Output On Success**
```
## Archive Complete
**Change:** <change-name>
**Schema:** <schema-name>
**Archived to:** openspec/changes/archive/YYYY-MM-DD-<name>/
**Specs:** ✓ Synced to main specs (or "No delta specs" or "Sync skipped")
All artifacts complete. All tasks complete.
```
**Guardrails**
- Always prompt for change selection if not provided
- Use artifact graph (openspec status --json) for completion checking
- Don't block archive on warnings - just inform and confirm
- Preserve .openspec.yaml when moving to archive (it moves with the directory)
- Show clear summary of what happened
- If sync is requested, use openspec-sync-specs approach (agent-driven)
- If delta specs exist, always run the sync assessment and show the combined summary before prompting
+288
View File
@@ -0,0 +1,288 @@
---
name: openspec-explore
description: Enter explore mode - a thinking partner for exploring ideas, investigating problems, and clarifying requirements. Use when the user wants to think through something before or during a change.
license: MIT
compatibility: Requires openspec CLI.
metadata:
author: openspec
version: "1.0"
generatedBy: "1.3.1"
---
Enter explore mode. Think deeply. Visualize freely. Follow the conversation wherever it goes.
**IMPORTANT: Explore mode is for thinking, not implementing.** You may read files, search code, and investigate the codebase, but you must NEVER write code or implement features. If the user asks you to implement something, remind them to exit explore mode first and create a change proposal. You MAY create OpenSpec artifacts (proposals, designs, specs) if the user asks—that's capturing thinking, not implementing.
**This is a stance, not a workflow.** There are no fixed steps, no required sequence, no mandatory outputs. You're a thinking partner helping the user explore.
---
## The Stance
- **Curious, not prescriptive** - Ask questions that emerge naturally, don't follow a script
- **Open threads, not interrogations** - Surface multiple interesting directions and let the user follow what resonates. Don't funnel them through a single path of questions.
- **Visual** - Use ASCII diagrams liberally when they'd help clarify thinking
- **Adaptive** - Follow interesting threads, pivot when new information emerges
- **Patient** - Don't rush to conclusions, let the shape of the problem emerge
- **Grounded** - Explore the actual codebase when relevant, don't just theorize
---
## What You Might Do
Depending on what the user brings, you might:
**Explore the problem space**
- Ask clarifying questions that emerge from what they said
- Challenge assumptions
- Reframe the problem
- Find analogies
**Investigate the codebase**
- Map existing architecture relevant to the discussion
- Find integration points
- Identify patterns already in use
- Surface hidden complexity
**Compare options**
- Brainstorm multiple approaches
- Build comparison tables
- Sketch tradeoffs
- Recommend a path (if asked)
**Visualize**
```
┌─────────────────────────────────────────┐
│ Use ASCII diagrams liberally │
├─────────────────────────────────────────┤
│ │
│ ┌────────┐ ┌────────┐ │
│ │ State │────────▶│ State │ │
│ │ A │ │ B │ │
│ └────────┘ └────────┘ │
│ │
│ System diagrams, state machines, │
│ data flows, architecture sketches, │
│ dependency graphs, comparison tables │
│ │
└─────────────────────────────────────────┘
```
**Surface risks and unknowns**
- Identify what could go wrong
- Find gaps in understanding
- Suggest spikes or investigations
---
## OpenSpec Awareness
You have full context of the OpenSpec system. Use it naturally, don't force it.
### Check for context
At the start, quickly check what exists:
```bash
openspec list --json
```
This tells you:
- If there are active changes
- Their names, schemas, and status
- What the user might be working on
### When no change exists
Think freely. When insights crystallize, you might offer:
- "This feels solid enough to start a change. Want me to create a proposal?"
- Or keep exploring - no pressure to formalize
### When a change exists
If the user mentions a change or you detect one is relevant:
1. **Read existing artifacts for context**
- `openspec/changes/<name>/proposal.md`
- `openspec/changes/<name>/design.md`
- `openspec/changes/<name>/tasks.md`
- etc.
2. **Reference them naturally in conversation**
- "Your design mentions using Redis, but we just realized SQLite fits better..."
- "The proposal scopes this to premium users, but we're now thinking everyone..."
3. **Offer to capture when decisions are made**
| Insight Type | Where to Capture |
|----------------------------|--------------------------------|
| New requirement discovered | `specs/<capability>/spec.md` |
| Requirement changed | `specs/<capability>/spec.md` |
| Design decision made | `design.md` |
| Scope changed | `proposal.md` |
| New work identified | `tasks.md` |
| Assumption invalidated | Relevant artifact |
Example offers:
- "That's a design decision. Capture it in design.md?"
- "This is a new requirement. Add it to specs?"
- "This changes scope. Update the proposal?"
4. **The user decides** - Offer and move on. Don't pressure. Don't auto-capture.
---
## What You Don't Have To Do
- Follow a script
- Ask the same questions every time
- Produce a specific artifact
- Reach a conclusion
- Stay on topic if a tangent is valuable
- Be brief (this is thinking time)
---
## Handling Different Entry Points
**User brings a vague idea:**
```
User: I'm thinking about adding real-time collaboration
You: Real-time collab is a big space. Let me think about this...
COLLABORATION SPECTRUM
════════════════════════════════════════════
Awareness Coordination Sync
│ │ │
▼ ▼ ▼
┌────────┐ ┌────────┐ ┌────────┐
│Presence│ │Cursors │ │ CRDT │
│ "3 │ │ Multi │ │Conflict│
│online" │ │ select │ │ free │
└────────┘ └────────┘ └────────┘
│ │ │
trivial moderate complex
Where's your head at?
```
**User brings a specific problem:**
```
User: The auth system is a mess
You: [reads codebase]
┌─────────────────────────────────────────────┐
│ CURRENT AUTH FLOW │
└─────────────────────────────────────────────┘
│
┌─────────────┼─────────────┐
▼ ▼ ▼
┌─────────┐ ┌─────────┐ ┌─────────┐
│ Google │ │ GitHub │ │ Email │
│ OAuth │ │ OAuth │ │ Magic │
└────┬────┘ └────┬────┘ └────┬────┘
│ │ │
└─────────────┼─────────────┘
▼
┌───────────┐
│ Session │
└─────┬─────┘
│
▼
┌───────────┐
│ Perms │
└───────────┘
I see three tangles. Which one's burning?
```
**User is stuck mid-implementation:**
```
User: /opsx:explore add-auth-system
The OAuth integration is more complex than expected
You: [reads change artifacts]
You're on task 4: "Implement OAuth flow"
Let me trace what's involved...
[draws diagram, explores options, suggests paths]
Want to update the design to reflect this?
Or add a spike task to investigate?
```
**User wants to compare options:**
```
User: Should we use Postgres or SQLite?
You: Generic answer is boring. What's the context?
User: A CLI tool that tracks local dev environments
You: That changes everything.
┌─────────────────────────────────────────────────┐
│ CLI TOOL DATA STORAGE │
└─────────────────────────────────────────────────┘
Key constraints:
• No daemon running
• Must work offline
• Single user
SQLite Postgres
Deployment embedded ✓ needs server ✗
Offline yes ✓ no ✗
Single file yes ✓ no ✗
SQLite. Not even close.
Unless... is there a sync component?
```
---
## Ending Discovery
There's no required ending. Discovery might:
- **Flow into a proposal**: "Ready to start? I can create a change proposal."
- **Result in artifact updates**: "Updated design.md with these decisions"
- **Just provide clarity**: User has what they need, moves on
- **Continue later**: "We can pick this up anytime"
When it feels like things are crystallizing, you might summarize:
```
## What We Figured Out
**The problem**: [crystallized understanding]
**The approach**: [if one emerged]
**Open questions**: [if any remain]
**Next steps** (if ready):
- Create a change proposal
- Keep exploring: just keep talking
```
But this summary is optional. Sometimes the thinking IS the value.
---
## Guardrails
- **Don't implement** - Never write code or implement features. Creating OpenSpec artifacts is fine, writing application code is not.
- **Don't fake understanding** - If something is unclear, dig deeper
- **Don't rush** - Discovery is thinking time, not task time
- **Don't force structure** - Let patterns emerge naturally
- **Don't auto-capture** - Offer to save insights, don't just do it
- **Do visualize** - A good diagram is worth many paragraphs
- **Do explore the codebase** - Ground discussions in reality
- **Do question assumptions** - Including the user's and your own
+110
View File
@@ -0,0 +1,110 @@
---
name: openspec-propose
description: Propose a new change with all artifacts generated in one step. Use when the user wants to quickly describe what they want to build and get a complete proposal with design, specs, and tasks ready for implementation.
license: MIT
compatibility: Requires openspec CLI.
metadata:
author: openspec
version: "1.0"
generatedBy: "1.3.1"
---
Propose a new change - create the change and generate all artifacts in one step.
I'll create a change with artifacts:
- proposal.md (what & why)
- design.md (how)
- tasks.md (implementation steps)
When ready to implement, run /opsx:apply
---
**Input**: The user's request should include a change name (kebab-case) OR a description of what they want to build.
**Steps**
1. **If no clear input provided, ask what they want to build**
Use the **AskUserQuestion tool** (open-ended, no preset options) to ask:
> "What change do you want to work on? Describe what you want to build or fix."
From their description, derive a kebab-case name (e.g., "add user authentication" → `add-user-auth`).
**IMPORTANT**: Do NOT proceed without understanding what the user wants to build.
2. **Create the change directory**
```bash
openspec new change "<name>"
```
This creates a scaffolded change at `openspec/changes/<name>/` with `.openspec.yaml`.
3. **Get the artifact build order**
```bash
openspec status --change "<name>" --json
```
Parse the JSON to get:
- `applyRequires`: array of artifact IDs needed before implementation (e.g., `["tasks"]`)
- `artifacts`: list of all artifacts with their status and dependencies
4. **Create artifacts in sequence until apply-ready**
Use the **TodoWrite tool** to track progress through the artifacts.
Loop through artifacts in dependency order (artifacts with no pending dependencies first):
a. **For each artifact that is `ready` (dependencies satisfied)**:
- Get instructions:
```bash
openspec instructions <artifact-id> --change "<name>" --json
```
- The instructions JSON includes:
- `context`: Project background (constraints for you - do NOT include in output)
- `rules`: Artifact-specific rules (constraints for you - do NOT include in output)
- `template`: The structure to use for your output file
- `instruction`: Schema-specific guidance for this artifact type
- `outputPath`: Where to write the artifact
- `dependencies`: Completed artifacts to read for context
- Read any completed dependency files for context
- Create the artifact file using `template` as the structure
- Apply `context` and `rules` as constraints - but do NOT copy them into the file
- Show brief progress: "Created <artifact-id>"
b. **Continue until all `applyRequires` artifacts are complete**
- After creating each artifact, re-run `openspec status --change "<name>" --json`
- Check if every artifact ID in `applyRequires` has `status: "done"` in the artifacts array
- Stop when all `applyRequires` artifacts are done
c. **If an artifact requires user input** (unclear context):
- Use **AskUserQuestion tool** to clarify
- Then continue with creation
5. **Show final status**
```bash
openspec status --change "<name>"
```
**Output**
After completing all artifacts, summarize:
- Change name and location
- List of artifacts created with brief descriptions
- What's ready: "All artifacts created! Ready for implementation."
- Prompt: "Run `/opsx:apply` or ask me to implement to start working on the tasks."
**Artifact Creation Guidelines**
- Follow the `instruction` field from `openspec instructions` for each artifact type
- The schema defines what each artifact should contain - follow it
- Read dependency artifacts for context before creating new ones
- Use `template` as the structure for your output file - fill in its sections
- **IMPORTANT**: `context` and `rules` are constraints for YOU, not content for the file
- Do NOT copy `<context>`, `<rules>`, `<project_context>` blocks into the artifact
- These guide what you write, but should never appear in the output
**Guardrails**
- Create ALL artifacts needed for implementation (as defined by schema's `apply.requires`)
- Always read dependency artifacts before creating a new one
- If context is critically unclear, ask the user - but prefer making reasonable decisions to keep momentum
- If a change with that name already exists, ask if user wants to continue it or create a new one
- Verify each artifact file exists after writing before proceeding to next
@@ -148,7 +148,7 @@ openspec/changes/phase-1-infrastructure/
## 敏感信息(已编辑) ## 敏感信息(已编辑)
- MySQL 密码:已配置在 application.yml(`!Fucker123..`) - MySQL 密码:已从仓库移除,使用环境变量注入
- Redis:无密码 - Redis:无密码
--- ---
+1 -1
View File
@@ -97,7 +97,7 @@ spring:
redis: redis:
host: 119.29.78.52 host: 119.29.78.52
port: 6379 port: 6379
password: '!Fucker123..' password: ${SUPERBIZ_REDIS_PASSWORD}
database: 0 database: 0
timeout: 3000 timeout: 3000
``` ```
+2 -2
View File
@@ -140,7 +140,7 @@ Error Code: 1049
datasource: datasource:
url: jdbc:mysql://119.29.78.52:33306/superbiz_agent?... url: jdbc:mysql://119.29.78.52:33306/superbiz_agent?...
username: root username: root
password: '!Fucker123..' password: ${SUPERBIZ_MYSQL_PASSWORD}
``` ```
**Redis 配置**: **Redis 配置**:
@@ -149,7 +149,7 @@ data:
redis: redis:
host: 119.29.78.52 host: 119.29.78.52
port: 6379 port: 6379
password: '!Fucker123..' password: ${SUPERBIZ_REDIS_PASSWORD}
``` ```
**Flyway 配置**: **Flyway 配置**:
+8 -1
View File
@@ -44,6 +44,13 @@ build/
app.log app.log
logs/ logs/
### Local Secrets ###
.env
.env.*
!.env.example
application-local.yml
application-*.local.yml
### Upload Files ### ### Upload Files ###
uploads/ uploads/
@@ -58,8 +65,8 @@ uploads/
### Windows / Runtime Artifacts ### Windows / Runtime Artifacts
*.stackdump *.stackdump
NUL
### MVP Demo Generated Outputs ### MVP Demo Generated Outputs
mvp/demo/output/*.json mvp/demo/output/*.json
!mvp/demo/output/README.md !mvp/demo/output/README.md
.pi/extensions/emdash-hook.ts
+44 -1
View File
@@ -103,5 +103,48 @@ Find reuse opportunities + Trace the call/dependency chain and impact radius:
## Baisc Infos ## Baisc Infos
Unless directly relevant to the user's current question, you should avoid proactively mentioning, illustrating, or Unless directly relevant to the user's current question, you should avoid proactively mentioning, illustrating, or
trailing off into the following information in 99% of cases: trailing off into the following information in 99% of cases:
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **SuperBizAgent-java** (13483 symbols, 22230 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
## Always Do
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
## Never Do
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
## Resources
| Resource | Use for |
|----------|---------|
| `gitnexus://repo/SuperBizAgent-java/context` | Codebase overview, check index freshness |
| `gitnexus://repo/SuperBizAgent-java/clusters` | All functional areas |
| `gitnexus://repo/SuperBizAgent-java/processes` | All execution flows |
| `gitnexus://repo/SuperBizAgent-java/process/{name}` | Step-by-step execution trace |
## CLI
| Task | Read this skill file |
|------|---------------------|
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
<!-- gitnexus:end -->
+43
View File
@@ -111,3 +111,46 @@ trailing off into the following information in 99% of cases:
- 文档目录结构: - 文档目录结构:
- 不要将文档放到用户目录(如 `C:\Users\EDY\.claude\`)中 - 不要将文档放到用户目录(如 `C:\Users\EDY\.claude\`)中
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **SuperBizAgent-java** (13483 symbols, 22230 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
## Always Do
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
## Never Do
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
## Resources
| Resource | Use for |
|----------|---------|
| `gitnexus://repo/SuperBizAgent-java/context` | Codebase overview, check index freshness |
| `gitnexus://repo/SuperBizAgent-java/clusters` | All functional areas |
| `gitnexus://repo/SuperBizAgent-java/processes` | All execution flows |
| `gitnexus://repo/SuperBizAgent-java/process/{name}` | Step-by-step execution trace |
## CLI
| Task | Read this skill file |
|------|---------------------|
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
<!-- gitnexus:end -->
+88
View File
@@ -155,6 +155,94 @@
- 使用场景:Executor 按 skill workflow 调用 evidence tools 收集事实,`tool_invocation` 记录这些事实证据。 - 使用场景:Executor 按 skill workflow 调用 evidence tools 收集事实,`tool_invocation` 记录这些事实证据。
- 边界:最终诊断结论必须被 evidence tools 支撑,不能仅由 skill 正文支撑。 - 边界:最终诊断结论必须被 evidence tools 支撑,不能仅由 skill 正文支撑。
### Diagnosis Harness
- 定义:围绕 Diagnosis Agent 提供确定性运行控制的边界,负责 Run、预算、取消、重试装配、Tool 调用记录、证据验真和最终释放,不承担业务诊断推理。
- 边界:Harness 不是工作流引擎,不实现 Planner/Executor/Composer 节点或自行编写 ReAct 循环。
### Diagnosis Agent
- 定义:诊断链路中唯一拥有 ReAct 工具循环并生成 `DiagnosisDraft` 的 Agent,负责规划证据查询、判断证据充分性和撰写完整诊断草稿。
- 边界:不负责意图路由、Run/Session 生命周期、证据物理验真、独立语义审查或最终发布;证据不足时必须明确停止并保留限制。
### EvidenceGuard
- 定义:Harness 内部的确定性证据验真能力,校验 Draft 引用、当前 Run 所有权、Tool 调用状态和有界 Agent 投影。
- 边界:EvidenceGuard 不调用 LLM,也不判断证据是否足以推出业务结论。
### SemanticGuard
- 定义:使用隔离上下文对完整诊断 Draft 与已验真证据做报告级语义审查的单轮 Agent。
- 边界:无工具、无记忆、无 ReAct 循环,不访问 Redis,不生成或改写用户报告。
### Invocation Status
- 定义:Tool 调用及结果投影的生命周期状态,固定为 `PROJECTING`、`READY`、`ERROR`。
- 边界:它只说明调用记录是否完成,不说明结果是否包含证据。
### Durable Audit
- 定义:为 Diagnosis Trace 长期保存的 Run、Agent 模型步骤和 Tool 调用元数据,用于 exact sessionId/runId 回放、评测和运维核对。
- 边界:只保存有界、脱敏、可长期保留的身份、状态、耗时、预算和结果摘要;不保存 Prompt、Thought、完整 Tool 参数、raw response 或 Redis canonical invocation。
### Evidence Status
- 定义:证据 Tool 的结果语义,固定为 `EVIDENCE_FOUND`、`NO_EVIDENCE`、`ERROR`。
- 边界:`EVIDENCE_FOUND` 只表示存在候选内容,不保证内容能够支持当前诊断;`NO_EVIDENCE` 只表示当前查询范围内没有匹配结果,不能解释为问题不存在、根因被排除或系统健康。
### Information Gain
- 定义:一次 Tool 结果是否推进当前 Diagnosis Run 的语义评价,固定为 `GAINED` 或 `NO_GAIN`。
- 边界:它评价的是结果对当前诊断的作用,不评价 Tool 产品质量;`NO_EVIDENCE` 和重复的规范化 `tool + scope` 可由 Harness 机械标记为 `NO_GAIN`,其他成功非空结果(包括 RAG `REFERENCE`)由模型评价。
### Collection State
- 定义:Diagnosis Harness 对当前 Run 是否允许继续收集证据的控制状态,固定为 `COLLECTING` 或 `SATURATED`。
- 边界:状态由 Harness 维护;`SATURATED` 可因连续 `NO_GAIN` 或连续进展协议错误达到各自配置阈值而进入,不包括硬预算耗尽。模型可以请求新的 Tool 调用,但不能绕过 `SATURATED`。
### Diagnosis Stop Reason
- 定义:Harness 停止当前 Run 继续调用 Tool 的内部原因,首版区分 `INFORMATION_SATURATED`、`BUDGET_LIMIT_REACHED` 与 `PROGRESS_PROTOCOL_VIOLATED`。
- 边界:它用于控制、Trace 和 Release 输入,不是用户可见生命周期状态,也不进入模型上下文;真正的不可恢复技术故障走失败通道。协议错误不累计为 `NO_GAIN`,使用独立阈值与 stop reason。
### Progress Protocol Violation
- 定义:模型未遵守 Tool Call Envelope 进展协议时的安全错误分类,例如缺失/错序/意外 `previous_observation`、缺失 `input` 或非法 Envelope。
- 边界:返回可修正 observation(`repair_required`、`violation_type`、期望上一轮 Tool Call ID、允许的 `information_gain`);连续错误达到阈值后交付一次 `STOP_REQUIRED/PROGRESS_PROTOCOL_VIOLATED`。不泄露业务参数、上一轮观察正文、raw response 或内部异常。
### Progress Snapshot
- 定义:Tool Loop 结束时,从当前 Run 的 Canonical Tool Result 一次性投影出的有界发布视图,用于生成已检查范围和客观结果。
- 边界:Canonical Tool Result 是真理源;Progress Snapshot 不逐轮维护、不保存原始 Tool Response、Prompt 或内部 thought,也不直接进入模型上下文。
### Safe Fallback Type
- 定义:`SafeFallback.type` 对没有发布诊断结论的业务原因分类,例如 `INSUFFICIENT_EVIDENCE`、`MISSING_REQUIRED_CONTEXT`、`BUDGET_EXHAUSTED` 或安全校验失败。
- 边界:它是 `ReleaseOutcome.FALLBACK` 的原因字段,不是与 `SUCCESS / FALLBACK / FAILED / CANCELLED` 平行的第二套生命周期状态。
### Diagnosis Release Use Case
- 定义:诊断业务发布的唯一决策入口,接收 DiagnosisDraft 和/或 Harness `stop_reason + ProgressSnapshot`,生成安全的 `SUCCESS / FALLBACK` 结果。
- 边界:`conclusion=null` 不触发 EvidenceRepair;只有存在结论时才执行完整 EvidenceGuard、EvidenceRepair 和 SemanticGuard 链路。不可形成安全业务内容的技术故障由 Chat Application Use Case 映射为 `FAILED / CANCELLED`。
### Diagnosis Draft Contract Failure
- 定义:Diagnosis Agent 最终文本为空、不是严格 JSON,或不满足 `DiagnosisDraft` Schema 时产生的 Agent 输出合同失败。
- 边界:非法文本始终丢弃,不做 Markdown/自然语言提取,也不调用模型修复;仅当当前 Run 的 `ProgressSnapshot` 含已验真 observed facts 时,Release 才能确定性发布 `INSUFFICIENT_EVIDENCE`,否则保持 `FAILED`。它不是 `Diagnosis Stop Reason`,不得伪装成信息饱和或预算终止。
### Model Observation
- 定义:Tool 内部标准化结果经过白名单投影后,作为 Tool Response 进入 Diagnosis Agent 上下文的有界视图。
- 边界:只包含模型完成语义判断和证据引用所需的信息;预算、阈值、重复指纹、原始相似度、原始 Tool Response 和完整 Harness 控制状态不得进入该视图。
### RunContext
- 定义:一次 Diagnosis Run 的显式执行上下文,结构不可变地携带 `sessionId`、`runId`、deadline,以及该 Run 独占的取消、预算、重试策略和生命周期状态句柄。
- 边界:RunContext 通过方法参数或框架受控 context 显式传播,不依赖 ThreadLocal;结构不可变不等于内部计数和取消状态不能变化,这些变化由线程安全句柄管理。
### Run Lifecycle
- 定义:Diagnosis Harness 对单次 Run 执行状态的内存控制,采用 first-terminal-wins 规则保证成功、失败、取消、超时和预算耗尽只能产生一个最终终态。
- 边界:Run Lifecycle 不直接等同于数据库实体写入;应用用例负责把最终状态映射到 `diagnosis_run` 持久化。
### Run Budget
- 定义:单次 Run 的模型调用、Tool 调用、单 Tool 调用、输入/输出/总 Token 和 canonical invocation 字节容量的线程安全消耗计数与门禁。
- 边界:预算上限由 Harness 配置显式提供;实际 Token 在模型响应后记录,超限后保留真实消耗并阻止后续执行。
### Harness Retry Policy
- 定义:Harness 对同一技术操作 attempt 数和可重试失败类型的显式策略。
- 边界:Router 与 SemanticGuard 的技术失败最多两次 attempt;Diagnosis Agent、Tool 和 Evidence repair 只有一次 attempt。Agent 正常 ReAct 轮次不是 retry,`NO_EVIDENCE`、业务拒绝、取消和预算耗尽不可重试。
### Chat Application Use Case
- 定义:一次 Chat 请求的唯一业务入口,拥有 Session/Run、意图路由、固定执行器、PreviousTurn 和最终持久化。
- 边界:不拥有 HTTP/SSE 连接,不把 ChatModel 或 Tool 选择权交给 Controller,也不在 Diagnosis Release Use Case 之外单独决定预算 Fallback 的业务内容。
### Chat SSE Contract
- 定义:Chat 公开入口的五事件协议,顺序固定为 `metadata -> status* -> content|failure -> done`。
- 边界:过程状态实时发送,最终安全内容最多释放一次;它不是 Token streaming,也不包含内部计划、Prompt、raw Tool 数据或异常。
### Verifier Skill Isolation ### Verifier Skill Isolation
- 定义:Chat Verifier 与 skill 系统隔离,只校验 Executor 答案和 `tool_trace_summary`。 - 定义:Chat Verifier 与 skill 系统隔离,只校验 Executor 答案和 `tool_trace_summary`。
- 使用场景:防止 Verifier 把 playbook 指令当作事实证据;Verifier 只判断已有证据是否支持结论。 - 使用场景:防止 Verifier 把 playbook 指令当作事实证据;Verifier 只判断已有证据是否支持结论。
+26 -1
View File
@@ -1,9 +1,27 @@
# devflow 索引 # devflow 索引
## Issue 生命周期
| Issue | 状态 | 说明 |
|---|---|---|
| ISS-014 | archived | 阶段 0-7 的单体 Diagnosis Agent、Harness、ACI、SSE、清理和最终 E2E 已完成并归档;阶段实现对应的 11 个 devflow/OpenSpec 项目均已 archived。 |
| ISS-015 | active | 阶段 1 硬停止已由 ISS-016 收口;剩余 Evidence Repair Schema、Reasoning 审计验证/治理与最终综合验收。 |
| ISS-016 | archived | Diagnosis 信息增益停止契约、协议修复反馈与统一 Release 已完成并归档。 |
## 项目 ## 项目
| 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 | | 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|---|---|---|---|---|---|---| |---|---|---|---|---|---|---|
| 2026-07-28 | rag-eval-hybrid-baseline | 离线 RAG eval 对齐 hybrid:search.mode 生成器、fixture meta、baseline 重刷;hybrid 质量闸门可用 denseDistance。 | RAG/eval/baseline | search.mode, fixture meta, kb_scope rag-eval, L0 filter fallback, denseDistance | openspec/changes/archive/2026-07-28-rag-eval-hybrid-baseline | archived |
| 2026-07-28 | rag-quality-score-unify | 统一 dense/hybrid scoreLabel 与 qualityScore;保检索序;去掉关键词 boost 改序与 hybrid L2 伪装。 | RAG/质量分/后处理 | qualityScore, scoreLabel dense/hybrid, originalRank, RetrievalScoreNormalizer, no boost rerank | openspec/changes/archive/2026-07-28-rag-quality-score-unify | archived |
| 2026-07-26 | diagnosis-information-gain-stop-contract | Diagnosis 信息增益停止、协议修复反馈、ProgressSnapshot 与统一 Release。 | Harness/Diagnosis stop/Release | ISS-016, GAINED, NO_GAIN, STOP_REQUIRED, ProgressSnapshot, PROGRESS_PROTOCOL_VIOLATED, INSUFFICIENT_EVIDENCE | openspec/changes/archive/2026-07-27-diagnosis-information-gain-stop-contract | archived |
| 2026-07-27 | rag-chunk-evidence-identity-dedup | chunk 级证据身份、去重、retrieve-k/return-n 与 SearchPort 地基,为 hybrid 铺路。 | RAG/证据身份/去重 | evidenceKey, maxChunksPerDocument, retrieve-k, return-n, KnowledgeSearchPort, document_id chunk-scoped | openspec/changes/archive/2026-07-27-rag-chunk-evidence-identity-dedup | archived |
| 2026-07-27 | rag-bm25-hybrid-drop-sdk | 真 dense+BM25 hybrid(MilvusClientV2),废弃知识路径旧 SDK 检索/写入。 | RAG/BM25/hybrid | MilvusClientV2, BM25, hybridSearch, RRFRanker, biz_hybrid, drop SDK path | openspec/changes/archive/2026-07-27-rag-bm25-hybrid-drop-sdk | archived |
| 2026-07-27 | rag-hybrid-search-rrf | Delivery 2:可配置 hybrid 检索与 RRF 多路融合(不绑旧 SDK)。 | RAG/hybrid/RRF | hybrid mode, RRF, KnowledgeSearchPort, filtered+unfiltered fusion, sparse-lite lexical | openspec/changes/archive/2026-07-27-rag-hybrid-search-rrf | archived |
| 2026-07-21 | single-react-tool-invocation-store | 建立统一 ToolBoundary 与 Redis canonical invocation store,集中生命周期、证据状态、TTL、容量和 Run 所有权。 | Harness/Tool boundary/Canonical store | ISS-014, ToolBoundary, canonical invocation, PROJECTING, READY, ERROR, TTL, RESULT_TOO_LARGE | openspec/changes/archive/2026-07-21-single-react-tool-invocation-store | archived |
| 2026-07-21 | single-react-harness-run-context | 建立显式 RunContext、Harness Core、预算、取消、类型化重试和 Tool Store 基础。 | Harness/Run lifecycle/Budget | ISS-014, RunContext, deadline, cancellation, budget, retry, ToolCallKey | openspec/changes/archive/2026-07-21-single-react-harness-run-context | archived |
| 2026-07-21 | single-react-aci-tool-contracts | 冻结 RAG、日志和 MySQL evidence Tool 的 Agent-facing ACI Schema、状态、框架调用引用和描述边界。 | Harness/Agent Tool contract | ISS-014, ACI, tool_call_id, evidence_status, RAG, query_logs, query_mysql, MOCK | openspec/changes/archive/2026-07-21-single-react-aci-tool-contracts | archived |
| 2026-07-21 | single-react-design-freeze | 冻结单体 Diagnosis Agent、Harness、Guard、工具证据与阶段门禁契约。 | Chat/Harness/Agent contract | ISS-014, single ReactAgent, Harness, EvidenceGuard, SemanticGuard, tool_call_id, evidence_status | openspec/changes/archive/2026-07-21-single-react-design-freeze | archived |
| 2026-07-10 | session-run-trace-isolation | 拆分会话态和运行态,引入 runId 隔离 Trace、Feedback、AIOps 和 demo 链路。 | Trace/session/run isolation | chat_session, diagnosis_run, runId, trace exact run, feedback fallback, AIOps SSE metadata, baseline drift | openspec/changes/archive/2026-07-10-session-run-trace-isolation | archived | | 2026-07-10 | session-run-trace-isolation | 拆分会话态和运行态,引入 runId 隔离 Trace、Feedback、AIOps 和 demo 链路。 | Trace/session/run isolation | chat_session, diagnosis_run, runId, trace exact run, feedback fallback, AIOps SSE metadata, baseline drift | openspec/changes/archive/2026-07-10-session-run-trace-isolation | archived |
| 2026-07-09 | interview-demo-quality-audit | 增加面试演示前置质量审计,覆盖 prompt、Gatekeeper 和评测基线。 | Agent eval/demo/Prompt audit | interview demo preflight, prompt_audit, gatekeeper rules, diagnosis baseline, 12 fixtures | openspec/changes/archive/2026-07-09-interview-demo-quality-audit | archived | | 2026-07-09 | interview-demo-quality-audit | 增加面试演示前置质量审计,覆盖 prompt、Gatekeeper 和评测基线。 | Agent eval/demo/Prompt audit | interview demo preflight, prompt_audit, gatekeeper rules, diagnosis baseline, 12 fixtures | openspec/changes/archive/2026-07-09-interview-demo-quality-audit | archived |
| 2026-07-08 | executor-composer-final-answer | 引入 Composer 生成最终回答,只使用 Verifier 允许的结论材料。 | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived | | 2026-07-08 | executor-composer-final-answer | 引入 Composer 生成最终回答,只使用 Verifier 允许的结论材料。 | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
@@ -33,3 +51,10 @@
| 2026-06-24 | lookup-knowledge-integration | 接入知识库检索,支持 L0 精确匹配和 L1 语义检索。 | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | - | archived | | 2026-06-24 | lookup-knowledge-integration | 接入知识库检索,支持 L0 精确匹配和 L1 语义检索。 | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | - | archived |
| 2026-06-23 | phase1-infrastructure | 搭建第一阶段基础设施,包括 MySQL、Redis、Milvus、Flyway 和 JPA。 | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | - | archived | | 2026-06-23 | phase1-infrastructure | 搭建第一阶段基础设施,包括 MySQL、Redis、Milvus、Flyway 和 JPA。 | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | - | archived |
| 2026-05-29 | chatmodel-abstraction | 抽象 ChatModel 和 EmbeddingModel,支持多模型路由。 | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | - | archived | | 2026-05-29 | chatmodel-abstraction | 抽象 ChatModel 和 EmbeddingModel,支持多模型路由。 | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | - | archived |
| 2026-07-21 | single-react-rag-log-projections | RAG/log projection adapters through ToolBoundary | Harness/Tool projection | ISS-014, RAG, query_logs, projection, scope, redaction, MOCK, NO_EVIDENCE | openspec/changes/archive/2026-07-21-single-react-rag-log-projections | archived |
| 2026-07-21 | single-react-mysql-readonly-tool | Fail-closed read-only MySQL evidence Tool with AST allowlist, JDBC controls and bounded projection | Harness/MySQL security | ISS-014, MySQL, JSqlParser, allowlist, PreparedStatement, timeout, projection | openspec/changes/archive/2026-07-21-single-react-mysql-readonly-tool | archived |
| 2026-07-21 | single-react-diagnosis-agent | Single internal Diagnosis ReactAgent with Harness-controlled model/tool loop, bounded context and typed Draft | Harness/Diagnosis Agent/ReAct | ISS-014, ReactAgent, DiagnosisDraft, PreviousTurn, ToolInterceptor, ModelInterceptor, budget | openspec/changes/archive/2026-07-21-single-react-diagnosis-agent | archived |
| 2026-07-21 | single-react-evidence-semantic-guards | Deterministic evidence validation, isolated semantic review and fail-closed diagnosis release | Harness/EvidenceGuard/SemanticGuard/Release | ISS-014, EvidenceGuard, verified snapshot, SemanticGuard, repair, fallback, release policy | openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards | archived |
| 2026-07-21 | single-react-chat-application-usecase | Internal Chat application use case with isolated routing, fixed executors and safe PreviousTurn | Harness/Chat application/Run persistence | ISS-014, Intent Router, PreviousTurn, PublishedResult, V012, observer, cancellation | openspec/changes/archive/2026-07-21-single-react-chat-application-usecase | archived |
| 2026-07-21 | single-react-chat-sse-cutover | Unique named-event Chat SSE endpoint, bounded production Harness wiring and strict frontend consumer | Chat/SSE/Harness production wiring | ISS-014, /api/chat, SSE, metadata, status, content, failure, done, disconnect, bounded executor | openspec/changes/archive/2026-07-22-single-react-chat-sse-cutover | archived |
| 2026-07-22 | single-react-cleanup-e2e | Remove legacy Agent paths, add bounded Harness audit, and complete exact-run live acceptance | Chat/Harness/cleanup/E2E | ISS-014, single ReAct Agent, durable audit, named SSE, exact run, Flyway V013 | openspec/changes/archive/2026-07-22-single-react-cleanup-e2e | archived |
@@ -10,7 +10,7 @@
- `devflow/index.md` was checked. Relevant history includes `session-storage`, `confidence-feedback`, `executor-action-memory-relevance`, and `chat-verifier-agent`. - `devflow/index.md` was checked. Relevant history includes `session-storage`, `confidence-feedback`, `executor-action-memory-relevance`, and `chat-verifier-agent`.
- `mvp/archive/2026-07-09-doc-cleanup/notes/agent-engineering-decisions.md` already recommends the next phase as "可复现 MVP Demo", including `mvp-demo` profile, fixed diagnosis case, one-click request, and `GET /api/diagnosis/{sessionId}/trace`. - `mvp/archive/2026-07-09-doc-cleanup/notes/agent-engineering-decisions.md` already recommends the next phase as "可复现 MVP Demo", including `mvp-demo` profile, fixed diagnosis case, one-click request, and `GET /api/diagnosis/{sessionId}/trace`.
- `mvp/issues/active/ISS-003-mvp-design-implementation-review.md` identifies test stability, session traceability, verifier evidence chain, upload path, and SupervisorAgent consistency as recent MVP concerns. Security cleanup is intentionally deferred by user decision. - `mvp/issues/archived/ISS-003-mvp-design-implementation-review.md` identifies test stability, session traceability, verifier evidence chain, upload path, and SupervisorAgent consistency as historical MVP concerns. Security cleanup was intentionally deferred by user decision at that time.
## Question Pool ## Question Pool
@@ -2,7 +2,7 @@
## Draft Acceptance ## Draft Acceptance
- [x] Issue exists: `mvp/issues/active/executor-evidence-attribution-hallucination.md`. - [x] Issue exists: `mvp/issues/archived/executor-evidence-attribution-hallucination.md`.
- [x] OpenSpec change artifacts exist. - [x] OpenSpec change artifacts exist.
- [x] devflow tracking files exist. - [x] devflow tracking files exist.
- [x] OpenSpec validation passes. - [x] OpenSpec validation passes.
@@ -0,0 +1,44 @@
# Acceptance: single-react-aci-tool-contracts
## 实现结果
- RAG Contract:最小 query Request、bounded document evidence Result。
- Log Contract:逻辑 Topic/Lookback Request、Mock provenance、Scope/Pattern/Event Result。
- MySQL Contract:逻辑 data source、参数化 SQL Request、bounded structured rows Result。
- 共享 Contract:snake_case Tool 名称、ACI 描述、不可变集合 helper,复用阶段 0 两套状态枚举。
- 旧 Tool、Chat/AIOps、Controller、持久化、数据源和公开协议未修改。
## 静态验证
- `openspec validate single-react-aci-tool-contracts --strict`:通过。
- `openspec instructions apply --change single-react-aci-tool-contracts --json`:12/12 tasks complete。
- `rg` 引用检查:新 contract 生产包未接入旧运行链路。
- 受保护文件 diff scope:旧 Tool、Chat/AIOps、Controller、Repository、resources 均为空。
## 脚本验证
- `mvn -q -DskipTests compile`:通过。
- `mvn -q '-Dtest=RagToolContractTest,QueryLogsToolContractTest,MysqlToolContractTest' test`:通过。
- `mvn -q '-Dtest=HarnessContractTest,LookupKnowledgeToolTest,QueryLogsToolsTest' test`:通过。
## 浏览器/人工验证
- 不适用。本阶段无 UI、Controller、SSE 或公开运行行为变化。
## 未验证
- 未执行 live LLM/Redis/CLS/MySQL E2E;本阶段没有接入这些运行路径,最终 live E2E 按 ISS-014 门禁留到阶段 7。
- 未验证 ToolInterceptor 的真实 ID 传播;阶段 2/3A 必须以 `ToolCallRequest.getToolCallId()` 添加集成测试。
- 未验证真实日志 adapter 的 `SourceKind` 扩展;本 Issue 首版明确只使用 Mock。
- Provider 侧旧凭据轮换仍需凭据所有者完成,仓库只能证明明文已移除。
## 剩余风险与后续门禁
- 新旧 Contract 短期并存,阶段 3B/3C 接入前不得声称旧 Tool 已符合新 ACI 输出。
- 下一阶段只能在本 change OpenSpec Archive 和 Git commit 完成后开始。
## 状态
- Stage acceptance: accepted
- OpenSpec archive: archived at `openspec/changes/archive/2026-07-21-single-react-aci-tool-contracts`
- Main spec sync: `openspec/specs/aci-evidence-tool-contracts/spec.md`(7 added requirements)
@@ -0,0 +1,32 @@
# Brief: single-react-aci-tool-contracts
## 背景
旧 RAG 和日志 Tool 暴露检索/基础设施细节与不一致状态,MySQL Tool 尚无 Agent-facing 类型。Harness 实现前需要先冻结三类最小 ACI 契约。
## 目标
- 冻结 RAG、日志、MySQL 的 Request/Result JSON Schema。
- 统一 `evidence_status`,并保持其与 invocation lifecycle 独立。
- 统一使用框架 `tool_call_id`,禁止模型传入或 Harness 生成第二套 ID。
- 冻结简短 Tool 名称/描述和 Mock 日志来源边界。
- 用三个独立契约测试锁定行为。
## 范围
- 新增 `com.superbiz.agent.harness.tool.contract` 值对象、枚举和描述常量。
- 复用阶段 0 的 `InvocationStatus` 与 `EvidenceStatus`。
- 验证 JSON、不可变集合、描述泄漏和新旧边界。
## 非目标
- 不切换旧 RAG/日志运行方法或 Chat/AIOps 注册。
- 不实现投影、Redis store、真实日志适配器或 MySQL 执行。
- 不修改 Controller/SSE 或公开协议。
## 元数据
- 分档:standard
- 接口影响:L2 前置内部契约;当前运行行为无变化
- 关联 Issue:ISS-014 阶段 1
- 关联 OpenSpec:`openspec/changes/single-react-aci-tool-contracts`
@@ -0,0 +1,116 @@
# Decisions: single-react-aci-tool-contracts
## 规模与入口
- 分档:standard。
- 入口:ISS-014 阶段 1,前置 `single-react-design-freeze` 已 Archive 并由 Git commit `58c3910` 固化。
- 目标:冻结三类 evidence Tool 的 Agent-facing ACI 契约,不接入新运行链路。
## Context
- `devflow/index.md` 命中 `single-react-design-freeze`、`modular-rag-pipeline` 和 `evidence-trace-hardening`。
- `devflow/glossary/CONTEXT.md` 已定义 Diagnosis Harness、Invocation Status、Evidence Status 与 Evidence Tools。
- 阶段 0 已确认 `tool_call_id` 使用框架 ID、两套状态语义分离、阶段串行门禁和阶段 6B 才公开切换。
- 未发现根目录旧 `CONTEXT.md` 与 glossary 冲突。
## Question Pool
| 维度 | 问题 | 模式 | 证据与结论 | 状态 |
|---|---|---|---|---|
| 术语 | invocation lifecycle 与 evidence result 是否使用同一状态? | evidence-driven | ISS-014 5.2、阶段 0 contract 明确分离;分别复用 `InvocationStatus` 与 `EvidenceStatus`。 | 已解决并汇报 |
| 术语 | `tool_call_id` 由谁生成、从哪里取得? | evidence-driven | Spring AI `AssistantMessage.ToolCall.id()` 与 Alibaba `ToolCallRequest.getToolCallId()` 提供框架 ID;普通 `ToolContext` 不自动加入该 ID。Harness 不生成第二套 ID。 | 已解决并汇报 |
| 边界 | 阶段 1 是否直接改旧 RAG/日志执行签名与返回值? | evidence-driven | ISS-014 阶段 1 只冻结 Contract,投影在阶段 3B、公开切换在阶段 6B;本阶段只新增契约代码和测试。 | 已解决并汇报 |
| 边界 | 日志阶段是否实现真实 CLS/MCP 或保留 Topic discovery? | evidence-driven | ISS-014 8.1 明确继续 Mock、删除 Agent 侧 discovery、真实适配器不在本 Issue 提前设计。 | 已解决并汇报 |
| 验收 | 如何证明契约已冻结且有界? | evidence-driven | 对三类独立 DTO 做精确 JSON、不可变集合、状态和描述泄漏测试;不以旧 Tool 集成测试代替。 | 已解决并汇报 |
| 技术 | 框架 ID 能否在后续 Harness 边界取得? | evidence-driven | 本地依赖 Spring AI Alibaba 1.1.2.0 暴露 `ToolInterceptor.interceptToolCall(ToolCallRequest, ToolCallHandler)`,request 含 `toolCallId`。 | 已解决并汇报 |
## Grill 结论
- 术语、边界、验收三类问题均已由代码、依赖 API、阶段 0 档案和 ISS-014 证明。
- 没有需要新增用户偏好或风险取舍的 `user-interview` 问题;不代理确认任何新方向。
- `grill-with-docs` 要求的代码可证问题已先查证;结论已向用户汇报。
- Proposal 已回写框架 ID 接入点、旧运行链路不切换、Mock 边界和 L2 接口影响。
## 已确认决策
- DTO 放在 Harness 的 Tool Contract 边界,复用阶段 0 的共享状态枚举,不在旧 `dto` 包继续堆叠协议。
- 三类结果只携带 `evidence_status`,canonical invocation 的 `status` 保持独立;阶段 1 通过测试冻结枚举,不提前定义存储实现。
- Tool description 使用代码常量冻结,后续 Tool adapter 注册时复用;旧 `@Tool` 注解本阶段不改,避免提前改变运行行为。
- RAG 输入仅保留 `query`;日志输入仅保留逻辑 `topic/query/lookback_minutes`;MySQL 输入仅保留逻辑 `data_source/sql/params`。
- 日志 `source_kind` 首版固定支持 `MOCK` 契约值,但保留 enum 扩展位置给后续真实适配器 change 审查。
## 能力与工具限制
- Discover 能力来源:`sm-flow` + `grill-with-docs`。
- 仓库要求的 `codebase-retrieval` 和 LSP 工具在当前工具集中不可用;已用 `rg` 引用搜索、源码阅读和本地依赖 `javap` 补足事实核对。该限制不改变契约方向,但后续 Apply 仍需通过编译和引用测试验证。
## Cross-artifact 对齐
| 链路 | 状态 | 结论 |
|---|---|---|
| brief 目标/范围/非目标 -> proposal | 已对齐 | 三类 DTO、状态、框架 ID、短描述和不切旧运行链路均有对应。 |
| proposal 范围/约束/承诺 -> design | 已对齐 | 包边界、ID 来源、record/defensive copy、三类 Schema 和迁移顺序均已设计。 |
| design 决策/接口影响/风险 -> specs/tasks | 已对齐 | L2 边界、字段、状态、描述、Mock provenance 和运行不切换均有可验证 requirement 与任务。 |
| specs 可观察行为 -> tasks | 已对齐 | 每类 Contract 都有实现与独立测试,另有状态、回归和 diff scope 验证。 |
## Architecture Audit
- 能力来源:`zoom-out`,以项目 glossary 的 Diagnosis Harness、Evidence Tools、Invocation Status 和 Evidence Status 术语审计。
- 当前输入到输出链路仍为 `ChatController/ChatService` 或 `AiOpsService -> ReactAgent -> LookupKnowledgeTool/QueryLogsTools -> 旧结果`,新契约没有运行消费者。
- 未来链路为 `Diagnosis Agent -> Alibaba ToolInterceptor/Harness -> typed Request -> adapter/store/projector -> bounded Result -> Agent observation`,阶段 1 只占有 typed contract 边界。
- 数据所有权保持明确:框架拥有 `tool_call_id`,canonical invocation 拥有生命周期,Tool-specific result 拥有证据语义与有界内容。
- 主要耦合风险是新旧契约短期并存被误当成已迁移;通过独立包、无旧调用方修改和后续阶段门禁控制,无 ADR 冲突。
## Commit Gate Preflight
- proposal、design、specs、tasks 文件完整,OpenSpec CLI 状态为 complete。
- strict validation:`openspec validate single-react-aci-tool-contracts --strict` 通过。
- question pool 中没有未汇报的 evidence-driven 结论或未确认的 user-interview 问题。
- 接口影响为 L2,当前运行消费者零变更;阶段 6B 的 L4 切换保持独立。
- cross-artifact 四段对齐无 gap,架构审计未发现需要回写的新实现约束。
- Apply 已由用户对 ISS-014 全阶段的持续授权覆盖;仍严格限制在本 Committed OpenSpec tasks 内。
## Pre-apply Research
### 参考实现
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`:确认旧 RAG 暴露 `LookupResult`、ContextPack 和检索 Trace,本阶段不改。
- `src/main/java/com/superbiz/agent/agent/tool/QueryLogsTools.java`:确认旧日志输入包含 region/logTopic/limit、存在 Topic discovery 和 mutable nested DTO,本阶段不改。
- `src/test/java/com/superbiz/agent/tool/LookupKnowledgeToolTest.java`:现有 RAG 行为回归基线。
- `src/test/java/com/superbiz/agent/agent/tool/QueryLogsToolsTest.java`:现有 Mock 日志行为回归基线。
- `src/main/java/com/superbiz/agent/harness/contract/*.java`:Java 17 record、Jackson snake_case 和共享状态枚举风格。
### 技术栈清单
- 请求/响应标准:Java 17 record + Jackson `@JsonProperty`,集合构造时 defensive copy。
- Tool 定义:当前使用 Spring AI `@Tool`,新名称/描述先以常量冻结,后续 adapter 注册复用。
- 框架调用 ID:Spring AI Alibaba `ToolCallRequest.getToolCallId()`;普通 `ToolContext` 不作为 ID 来源。
- 异常与校验:本阶段只冻结结构;Schema、ID、权限、范围和状态组合错误由后续 Pre-Tool/Projector 显式返回安全 `ERROR`。
- MQ/Consumer/加密验签:本 change 不涉及。
### 新建类型
- `AgentToolContracts` 与 contract defensive-copy helper。
- RAG Request/Result/Evidence。
- Log Topic/SourceKind/Request/Result/Scope/Pattern/Event。
- MySQL Request/Result。
- 三个独立 contract test classes。
### 影响半径
- 新增包当前应无生产调用方;旧 Chat/AIOps、Tool、Controller、Repository 和配置文件均不修改。
- 通过 focused compile/tests 和 `rg`/diff scope 证明边界。
## Apply 结果
- 冲突分类:未发现 OpenSpec 遗漏、代码偏离或方向不确定项。
- 新增共享 ACI Tool 名称/描述、defensive-copy helper 和三类 typed Request/Result records。
- 新增三个独立契约测试,覆盖精确 JSON、状态分离、框架 ID 原样保留、不可变集合、Mock provenance 和基础设施字段排除。
- 首模块对齐:共享/RAG/日志/MySQL contract 与 design/tasks 全部完成;旧 runtime 接入保持 TODO,归属后续 3B/3C/6B changes。
## Apply 验证
- 编译:`mvn -q -DskipTests compile` 通过。
- 新契约:`mvn -q '-Dtest=RagToolContractTest,QueryLogsToolContractTest,MysqlToolContractTest' test` 通过。
- 旧行为回归:`mvn -q '-Dtest=HarnessContractTest,LookupKnowledgeToolTest,QueryLogsToolsTest' test` 通过。
- 静态 scope:新 contract 生产类型当前无旧运行消费者;受保护的 Tool、Chat/AIOps、Controller、Repository 和配置文件 diff 为空。
@@ -0,0 +1,27 @@
# Evidence: single-react-aci-tool-contracts
## 文档证据
- ISS-014 5.2/5.3 冻结 `evidence_status` 与框架 `tool_call_id`;6-9 节冻结 RAG、日志和 MySQL Agent-facing Schema。
- ISS-014 阶段 1 明确只冻结 ACI Contract,不实现 Agent 架构切换、真实 CLS/MCP 或 MySQL 执行。
- 阶段 0 OpenSpec 与 devflow 已冻结 `InvocationStatus`、`EvidenceStatus`、框架 ID 真理源和阶段 6B 才公开切换。
## 代码证据
- `LookupKnowledgeTool` 仍返回包含 ContextPack/Trace 的旧 `LookupResult`,证明需要新的 bounded RAG Contract,也证明本阶段未提前切换。
- `QueryLogsTools` 仍暴露 region/logTopic/limit、Topic discovery 和旧 mutable DTO,证明逻辑 Topic/Scope/Mock provenance 契约的必要性。
- `ChatService` 与 `AiOpsService` 仍引用旧 `LookupKnowledgeTool`/`QueryLogsTools`,新 contract 生产包当前没有旧运行消费者。
- 本地 Spring AI 1.1.7 `AssistantMessage.ToolCall` 提供 `id()`;Spring AI Alibaba 1.1.2.0 `ToolCallRequest` 提供 `getToolCallId()` 与 `ToolInterceptor` 边界。
- Spring AI 1.1.7 普通 `ToolContext` 只传递调用方 context/history,不自动提供当前 Tool Call ID,因此后续必须从 Alibaba interceptor request 接入。
## Evidence-driven 结论
- lifecycle 与 evidence result 必须保持两套正交状态;已汇报并进入 OpenSpec/代码测试。
- `tool_call_id` 可以从当前框架 API 取得,Harness 无需也不得生成第二套 ID;已汇报并进入 OpenSpec。
- 阶段 1 不改旧运行签名/返回;已通过引用和 diff scope 验证。
- 日志首版保留 Mock 数据源但必须显式 `source_kind=MOCK`,不实现真实适配器;已进入 Log Contract。
- 三类 Contract 的可验证口径是精确 JSON、不可变结果、短描述和基础设施字段排除;三个独立测试均通过。
## 工具限制
- 当前会话未提供 `codebase-retrieval` 或 LSP;使用 `rg` 引用搜索、源码阅读、本地依赖 `javap`、Maven 编译和 focused tests 完成等价核对。
@@ -0,0 +1,42 @@
# Acceptance: single-react-chat-application-usecase
## Result
- Status: archived
- OpenSpec tasks: 14/14 complete
- Interface impact: L3 database/collaboration
- Public protocol: unchanged
## Static Verification
- V012 migration、`DiagnosisRun` 和 `DiagnosisRunRepository` 对齐 `intent/release_outcome/published_result` 及安全 PreviousTurn filter。
- 公开 Controller、前端和 endpoint diff 为空。
- Router/System executor 未检出 Tool、ReactAgent、ThreadLocal 或手写循环。
- `PublishedResult` 固定为 `user_query/published_conclusion/scope/limitations/source_documents`;序列化负向测试覆盖内部字段泄漏。
## Script Verification
- `mvn -q -DskipTests compile`:通过。
- Stage 6A focused `ApplicationExecutorsTest,PublishedResultPersistenceTest,ChatApplicationUseCaseTest`:13 tests,通过。
- Stage 2-5 与 6A regression selection:18 suites / 76 tests,0 failure/error/skipped。
- `openspec validate single-react-chat-application-usecase --strict`:通过。
## Browser or Manual Verification
- Not applicable。阶段 6A 没有 UI 或公开入口变化。
## Not Verified
- 未运行真实 LLM、Redis、日志和 MySQL live E2E;按 ISS-014 串行门禁统一留到阶段 7。
- V012 未在本阶段连接真实数据库执行;migration/entity/query 已由静态检查、focused persistence tests 和 compile 覆盖。
## Remaining Work
- 阶段 6B:唯一 `POST /api/chat` SSE 原子切换、旧 endpoint 删除和前端消费者迁移。
- 阶段 7:旧链路清理、全局 spec 格式修复和最终 live E2E。
## Archive
- `.archive-ready`: created
- OpenSpec archive: `openspec/changes/archive/2026-07-21-single-react-chat-application-usecase`
- Main spec sync: `openspec/specs/single-react-chat-application-usecase/spec.md`
@@ -0,0 +1,31 @@
# Brief: single-react-chat-application-usecase
## Background
阶段 2-5 已具备 RunContext、单一 Diagnosis Agent 和安全释放门禁,但没有统一应用用例拥有 Session/Run、意图路由、PreviousTurn、固定执行器和最终持久化,阶段 6B 因而无法只做协议切换。
## Goal
在不改变公开 Chat/SSE 行为的前提下,建立内部 `ChatApplicationUseCase`,统一三类意图、同一 Run 生命周期、安全 PreviousTurn 和 typed public content。
## Scope
- 无 Tool、无记忆、无 ReAct 的三分类 Intent Router。
- SYSTEM_CHAT、KNOWLEDGE_QUERY、DIAGNOSIS 固定执行器。
- Knowledge exact invocation/document reference validation。
- 同 Session 最近安全 Diagnosis SUCCESS 的有界 PreviousTurn。
- `diagnosis_run` V012 字段、JPA store 和安全 `PublishedResult`。
- Protocol-neutral observer、Run control、终态持久化和 focused tests。
## Non-goals
- 不修改 Controller、`/api/chat`、`/api/chat_stream`、SSE schema 或前端消费者。
- 不删除旧 ChatService、多 Agent、ThreadLocal 或旧 session storage。
- 不运行真实模型、Redis、日志和 MySQL live E2E;统一留到阶段 7。
## Metadata
- Scale: complex
- Interface impact: L3 database/collaboration
- OpenSpec: `single-react-chat-application-usecase`
- Parent issue: `ISS-014`
@@ -0,0 +1,117 @@
# Decisions: single-react-chat-application-usecase
## Discover Status
- Checkpoint: Discover
- Capability source: `sm-flow` + `grill-with-docs`;`codebase-retrieval`、LSP 和 GitNexus MCP 当前不可用,使用既有 GitNexus 结论、`rg` 引用核对和源码阅读降级。
- Scale: complex。跨模型路由、三类执行器、Run/session 生命周期、数据库 migration、PreviousTurn 和阶段 6B consumer boundary。
- `devflow/index.md` 命中阶段 0-5、session-run-trace-isolation、RAG contracts 和 release guards;无 ADR 冲突。
## Question Pool
| # | 维度 | 问题 | 模式 | 状态 |
|---|---|---|---|---|
| Q1 | 术语 | Chat Application Use Case 与 Controller、Harness、Agent 的职责边界是什么? | evidence-driven | 已解决 |
| Q2 | 路由 | Router 输入、输出和可重试失败范围是什么? | evidence-driven | 已解决 |
| Q3 | 路由 | Router 最终失败是否允许默认进入 Diagnosis? | evidence-driven | 已解决 |
| Q4 | 执行器 | 三种 intent 分别允许哪些模型和 Tool 行为? | evidence-driven | 已解决 |
| Q5 | Knowledge | 单次 RAG 调用如何验证模型引用且不公开 Tool Call ID? | evidence-driven | 已解决 |
| Q6 | PreviousTurn | 上一回合的真理源、筛选条件和截断边界是什么? | evidence-driven | 已解决 |
| Q7 | 生命周期 | 何时读取上一回合、创建当前 Run、记录 intent 和终态? | evidence-driven | 已解决 |
| Q8 | 持久化 | diagnosis_run 需要新增哪些字段,哪些内部内容禁止进入 published_result? | evidence-driven | 已解决 |
| Q9 | 6B 边界 | 如何让 Controller 切换时不重写应用用例? | evidence-driven | 已解决 |
| Q10 | 接口 | 数据库/内部接口影响等级和回滚要求是什么? | evidence-driven | 已解决 |
| Q11 | 验收 | 如何证明原始 Query、sessionId/runId 和失败终态一致传播? | evidence-driven | 已解决 |
## Evidence-driven
| 结论 | 证据来源 | 是否已汇报用户 |
|---|---|---|
| Controller 只负责协议;Application Use Case 拥有 session/run、routing、executor、persistence,Harness 拥有预算/取消/释放,Agent 拥有诊断语义。 | ISS-014 3、4.2、阶段 6A/6B | 已汇报 |
| Router 输入只含 query/last_intent/last_user_query,输出只允许三枚举;无 Tool/记忆/ReAct。 | ISS-014 4.5 | 已汇报 |
| timeout/transport/非法输出可重试一次;第二次失败返回安全入口错误,不进入 Diagnosis。 | ISS-014 重试策略、`HarnessRetryPolicies.intentRouter()` | 已汇报 |
| SYSTEM_CHAT 无 Tool;KNOWLEDGE_QUERY 只调用一次 lookup;DIAGNOSIS 进入 Agent + Guards。 | ISS-014 4.5 | 已汇报 |
| PreviousTurn 只来自同 Session 最近 `DIAGNOSIS + SUCCESS + published_result`,不使用 Redis 历史。 | ISS-014 4.6 | 已汇报 |
| `PublishedResult`/`PreviousTurn` 已冻结为 query/conclusion/scope/limitations/source_documents,不含 Tool ID/raw/Draft/reason。 | 阶段 0 contracts、`PublishedResult`、`PreviousTurn` | 已汇报 |
| 当前 `DiagnosisRun`/V011 尚无 intent/release_outcome/published_result,需要 V012 和 repository query。 | `DiagnosisRun.java`、`V011__add_session_run_isolation.sql` | 已汇报 |
| 上一回合必须在保存当前 PENDING Run 前读取,否则 latest query 会命中当前请求。 | repository 当前 latest method + 生命周期顺序推导 | 已汇报 |
| 6B 需要 metadata/status/cancel,6A 应提供 observer 和显式 RunContext,而不包含 SSE 类型。 | ISS-014 4.7、阶段 6B | 已汇报 |
## User-interview
- 无新增 user-interview。三类路由、数据库字段、PreviousTurn、失败语义、阶段边界和自动 Apply/Archive/commit 均由 ISS-014 与用户持续授权冻结。
## Key Decisions
- 在创建当前 Run 前读取 latest safe routing context 和 PreviousTurn,随后 `startRun -> persist RUNNING -> observer metadata -> route`。
- Application output 使用 typed content union,Diagnosis success 转为无 Tool ID 的 public report view;Fallback output 不携带完整 snapshot。
- Knowledge path 生成一次 direct canonical Tool Call ID(该路径没有框架 Tool Call),只调用 `lookup_knowledge`,严格验证 `KnowledgeAnswerDraft` 的 exact call ID 和 document subset,再移除 ID 发布。
- 所有普通单轮模型调用复用阶段 5 的受控 `GuardModelCall`,从而共享 Core 模型/Token/timeout/cancel 边界;不引入新模型路由。
- `PublishedResult` 仅在 Diagnosis SUCCESS 且 conclusion 非空时写入;Fallback/Failed/Cancelled/System/Knowledge 不生成 PreviousTurn 真理源。
- 数据库变更使用 V012 可前向迁移;回滚为先停止新应用用例,再删除新索引/列,不影响 V011 既有字段。
- 不创建 ADR:这些是 ISS-014 已冻结设计的落地,不是新的跨项目不可逆决策。
## OpenSpec Backfill
- 需进入 proposal/design/spec/tasks:路由隔离/重试、三执行器、original query、PreviousTurn filter/bounds、observer/cancel、Run terminal persistence、V012/L3、公开隔离。
- 非目标:Controller/SSE/前端切换、旧链路删除、live E2E。
## Cross-artifact Alignment
| 上游 -> 下游 | 检查内容 | 状态 |
|---|---|---|
| ISS-014/brief -> proposal | 三路由、previous turn、Run lifecycle、内部-only 和阶段 6B handoff | 已对齐 |
| proposal -> design | typed executors、observer、JPA/V012、异常/终态、L3 migration/rollback | 已对齐 |
| design -> specs/tasks | 每项所有权/安全边界均有可观察 requirement 和实现测试切片 | 已对齐 |
| specs -> tasks | 9 组 requirements 覆盖 Router/executors、store/policy、application、verification | 已对齐 |
## Architecture Audit
- 能力来源:`zoom-out`,使用 Chat Session、Diagnosis Run、RunContext、Diagnosis Agent、EvidenceGuard、SemanticGuard 和 PublishedResult 术语。
- 链路为 `request -> application use case -> run store/router -> fixed executor -> Harness/Agent/Tool -> typed public content -> run finish`;Controller 不拥有模型/工具/Run。
- ChatRunStore 拥有 MySQL 映射,Core 拥有运行状态,Application 拥有 dispatch/终态,path executor 拥有单一路径行为;PreviousTurn policy 是唯一安全历史投影。
- 最大风险是 prior/current Run 顺序和 DB/lifecycle 双终态,design/tasks 已固定 prior read before start、single finish path 和 focused failure/cancel tests。
- V012 是 L3 additive schema;migration、entity、repository、rollback 独立章节完整,公开入口阶段 6A 零变化。
## Interface Impact
- 级别:L3 database/collaboration interface。
- 新增 `diagnosis_run.intent/release_outcome/published_result` 和索引;修改 Entity/Repository,新增 internal application/store/output contract。
- 消费者:阶段 6B Controller/SSE adapter、MySQL/Flyway;旧 ChatService 在本阶段不消费新字段。
- 迁移/回滚:V012 nullable additive;回滚先切旧入口,再删除 index/columns。
## Commit Gate Preflight
- proposal、design、specs、tasks 完整,`openspec status` complete,change strict validation 通过。
- Question pool 全部已解决并汇报,无 user-interview、未判级接口或未接受架构风险。
- Cross-artifact 四段对齐无 gap;V012/L3、prior read ordering、terminal persistence 和 6B handoff 已进入 design/spec/tasks。
- Apply/Archive/commit 使用用户持续授权;公开协议和前端必须保持零 diff。
- `.committed` 已创建,可进入 Apply。
## Pre-apply Research
- `DiagnosisHarnessCore`/`RunContext`:Run ID、budget、cancel 和 first-terminal-wins。
- `GuardModelCall`/`HarnessRetryExecutor`:单轮模型 timeout/usage 和 Router 两次 attempt。
- `HarnessEvidenceTools`/`RagToolResult`:Knowledge 唯一 lookup 路径和有界 projection。
- `DiagnosisAgentUseCase`/`DiagnosisReleaseUseCase`:Diagnosis Draft 与安全 release boundary。
- `DiagnosisRun`/`DiagnosisRunRepository`/`V011`:现有 Run schema 和写入模式。
- `ChatService.ensureChatSession/startDiagnosisRun`:只参考 session metadata/JPA 写法,不复用旧 routing、多 Agent 或 ThreadLocal。
- 技术栈:Spring AI direct Prompt、Jackson strict JSON、JPA repository、Flyway additive migration、protocol-neutral observer;无 MQ/新依赖。
## Apply Progress
- Router/System/Knowledge contracts 与 executors 已完成,tasks 1.1-1.4 完成。
- 5 个 `ApplicationExecutorsTest` 通过:同输入 retry、最终 routing failure、System direct call、Knowledge exact references/no ID、NO_EVIDENCE/model skip。
- TODO:PublishedResult/JPA/V012、Diagnosis executor、总应用用例和综合验证。
- REVIEW:修正 `JpaChatRunStore` 多构造器 Spring 注入歧义;prior/start/intent/finish 持久化异常统一为稳定 `RUN_PERSISTENCE_FAILED`,并在安全完成时更新 ChatSession 活跃时间/消息对数。均为代码偏离修复,无需变更 OpenSpec。
- PublishedResult/JPA/V012、Diagnosis executor、protocol-neutral Run control 与总 ChatApplicationUseCase 已完成,tasks 2.1-3.4 完成。
- 13 个 stage 6A focused tests 通过;TODO 仅剩综合回归、static scope 和 OpenSpec verification。
- 综合回归曾在 `SemanticGuardTest.attemptTimeoutCancelsBothPermittedModelCalls` 出现负载相关失败。诊断确认生产代码对每次 timeout 均调用 `Future.cancel(true)`,但第二个 Future 可能在任务线程启动前已取消,此时不存在可接收 interrupt 的线程。分类为测试假设偏差,不是 OpenSpec 或生产代码偏离;回归断言改为两次 TIMEOUT attempt、两次模型预算预留,以及至少一个已运行调用收到 interrupt。
## Final Review
- Stage 6A focused tests 与阶段 2-5 regression 共 18 suites / 76 tests,0 failure/error/skipped;Maven compile 通过。
- OpenSpec strict validation 通过;V012、Entity、Repository 的三个字段和 previous-turn filter 对齐。
- 公开 Controller、前端和 endpoint 零 diff;Router/System executor 无 Tool、ReactAgent、ThreadLocal 或手写 loop。
- `PublishedResult` 只包含 `user_query/published_conclusion/scope/limitations/source_documents`,负向序列化测试通过。
- 本阶段不运行 live E2E,按 ISS-014 门禁留到阶段 7。
@@ -0,0 +1,31 @@
# Evidence: single-react-chat-application-usecase
## Code and Contract Evidence
- `DiagnosisHarnessCore`/`RunContext` 已提供 Run ID、预算、取消和 first-terminal-wins,应用用例无需创建第二套生命周期。
- `GuardModelCall`/`HarnessRetryExecutor` 提供单轮模型 timeout、Token 记账和两次 Router attempt。
- `HarnessEvidenceTools`/`RagToolResult` 提供 Knowledge 路径唯一 lookup 和有界 projection。
- `DiagnosisAgentUseCase`/`DiagnosisReleaseUseCase` 已形成 Diagnosis Draft 与安全发布边界。
- `DiagnosisRun`/`DiagnosisRunRepository`/V011 提供既有 Run 持久化,V012 以 nullable additive 字段扩展。
## Confirmed Boundaries
- Router 输入只含原始 Query、可选 last intent 和 last user query;最终失败不得默认进入 Diagnosis。
- 三类 executor 不互相调用,所有路径接收未改写 Query。
- PreviousTurn 只来自同 Session 最近 `DIAGNOSIS + SUCCESS + published_result`,且必须在保存当前 Run 前读取。
- Knowledge 公开内容只保留 stable document metadata,不发布 direct Tool Call ID。
- `PublishedResult` 不保存 Tool ID、raw evidence、完整 Draft 或 SemanticGuard reason;Fallback/Failed/Cancelled 不写安全历史。
- Controller/SSE/前端切换属于阶段 6B,本阶段保持公开协议不变。
## Diagnosis Finding
- 综合回归暴露 `SemanticGuardTest` 的负载竞态:第二个 Future 可能在获得线程前被取消,因而不会产生第二次 interrupt。
- 生产代码已对每次 timeout 调用 `Future.cancel(true)`;测试改为验证两次 TIMEOUT attempt、两次模型预算预留,以及至少一个运行中调用被中断。
- 分类为测试假设偏差,不是生产代码或 OpenSpec 偏离。
## Verification Evidence
- Stage 6A focused tests:13 tests 通过。
- Stage 2-5 与 6A 综合回归:18 suites / 76 tests,0 failure/error/skipped。
- Maven compile 与 change strict validation 通过。
- 公开 Controller/前端零 diff;Router/System executor 无 Tool、ReactAgent、ThreadLocal 或手写 loop。
@@ -0,0 +1,49 @@
# Acceptance: single-react-chat-sse-cutover
## Result
- Status: archived
- OpenSpec tasks: 14/14 complete
- Interface impact: L4 breaking HTTP/frontend contract
- Public Chat protocol: unique named-event SSE `POST /api/chat`
## Static Verification
- 组合路由保持 `/api/chat`、`/api/ai_ops`、`/api/chat/clear`、`/api/chat/session/{sessionId}` 和 `/runs`;`/api/chat_stream` 已删除。
- 生产 Controller/前端 legacy Chat token scan:0。
- `ChatController` forbidden dependency scan:0;结构测试固定其三个 protocol/application dependencies。
- SSE payload 只包含固定 metadata/status/content/failure/done fields,`CANCELLED` 不作为公开 done outcome。
- mode selector DOM、state、consumer 和 CSS 已删除。
## Script Verification
- `mvn -q -DskipTests compile`:通过。
- Stage 2-6B + AiOps regression selection:24 suites / 101 tests,0 failure/error/skipped。
- `openspec validate single-react-chat-sse-cutover --strict`:通过。
- `node --check src/main/resources/static/app.js`:通过。
- `git diff --check`:通过,仅有仓库既存 LF/CRLF 提示。
## Browser or Manual Verification
- 本阶段未运行浏览器人工验证;前端协议由静态 contract test 和 JavaScript syntax check 覆盖。
## Not Verified
- 未运行 live 模型、Redis、日志和 MySQL E2E;按 ISS-014 阶段门禁统一留到阶段 7。
- 未验证外部第三方 Chat API consumer;L4 变更不提供兼容分支,外部消费者必须同步迁移到 named-event SSE。
## Migration and Rollback
- 部署必须将后端 SSE endpoint 与 bundled frontend consumer 作为同一版本原子发布。
- 回滚必须同时回滚 Controller 和 frontend 到阶段 6A commit;V012 additive nullable migration 可保留。
- 不允许通过恢复 `/api/chat_stream`、同步 JSON consumer 或旧 message wrapper 形成双轨兼容。
## Remaining Work
- 阶段 7:物理删除旧 Agent/Graph/Hook/ThreadLocal/ChatService 路径和过时测试,更新文档并完成最终 live E2E、日志与数据库核验。
## Archive
- `.archive-ready`: created
- OpenSpec archive: `openspec/changes/archive/2026-07-22-single-react-chat-sse-cutover`
- Main spec sync: `openspec/specs/single-react-chat-sse-cutover/spec.md` (10 requirements added)
@@ -0,0 +1,32 @@
# Brief: single-react-chat-sse-cutover
## Background
阶段 6A 已建立 protocol-neutral `ChatApplicationUseCase`,但公开 Chat 仍有同步 `/api/chat` 和伪流式 `/api/chat_stream` 两条旧链路,Controller 直接拥有模型、Tools、Session 历史和无界线程池,前端也保留两套消费者。
## Goal
将公开 Chat 原子切换为唯一 `POST /api/chat` named-event SSE,并让 Controller 只承担校验、HTTP/SSE 和连接生命周期;所有公开内容必须来自阶段 6A 的安全释放结果。
## Scope
- 唯一 `/api/chat` SSE 与 `metadata -> status* -> content|failure -> done` 状态机。
- exact Run disconnect/timeout/send-failure cancellation。
- Spring-managed bounded Chat/model executors 和完整 Harness production Bean graph。
- 前端唯一 named-event consumer、typed renderer 和 metadata identity 保存。
- AiOps 模型/Tool acquisition 下沉到 service,并将 AiOps/Session endpoint 从 Chat Controller 职责中隔离。
- L4 前后端迁移、成对回滚和 focused regression。
## Non-goals
- 不做 Token streaming、最终答案切片、断线续传、事件重放、轮询或 WebSocket。
- 不修改 `/api/ai_ops` 的公开 URL、请求和 SSE message-wrapper 行为。
- 不在本阶段物理删除旧多 Agent、ChatService、Hook 或 ThreadLocal;阶段 7 统一清理。
- 不运行 live 模型、Redis、日志或 MySQL E2E;阶段 7 统一验收。
## Metadata
- Scale: complex
- Interface impact: L4 breaking HTTP/frontend contract
- OpenSpec: `single-react-chat-sse-cutover`
- Parent issue: `ISS-014`
@@ -0,0 +1,110 @@
# Decisions: single-react-chat-sse-cutover
## Discover Status
- Checkpoint: Discover
- Capability source: `sm-flow` + `grill-with-docs`;`codebase-retrieval`、LSP 和 GitNexus MCP 当前不可用,使用 `rg` 引用核对、源码阅读和 focused tests 降级。
- Scale: complex。涉及 L4 HTTP/SSE 协议、前端消费者、异步连接生命周期、生产 Bean 装配和阶段 7 删除边界。
- `devflow/index.md` 命中阶段 0-6A、ISS-013 和 session-run-trace-isolation;ISS-014 是更新且已冻结的最终协议源。
## Question Pool
| # | 维度 | 问题 | 模式 | 状态 |
|---|---|---|---|---|
| Q1 | 术语 | “真正 SSE”是 Token streaming,还是过程事件实时 + 最终内容一次释放? | evidence-driven | 已解决 |
| Q2 | 协议 | 唯一 endpoint、事件名称、payload、顺序、互斥和终态是什么? | evidence-driven | 已解决 |
| Q3 | 边界 | Controller、Application Use Case、Harness 和 SSE adapter 各自拥有何种职责? | evidence-driven | 已解决 |
| Q4 | 生命周期 | disconnect/timeout/send failure 如何取消同一个 Run,正常 complete 如何避免误取消? | evidence-driven | 已解决 |
| Q5 | 执行器 | 如何消除 Controller 自建无界线程池并处理饱和? | evidence-driven | 已解决 |
| Q6 | 装配 | 阶段 2-6A plain Java components 如何形成可启动的生产 Bean graph? | evidence-driven | 已解决 |
| Q7 | 前端 | 快速/流式双模式如何迁移到唯一 SSE consumer? | evidence-driven | 已解决 |
| Q8 | 安全 | 哪些内部状态/内容禁止进入 SSE? | evidence-driven | 已解决 |
| Q9 | 兼容 | 是否保留同步 `/api/chat` 或 `/api/chat_stream` 兼容? | evidence-driven | 已解决 |
| Q10 | 范围 | `/api/ai_ops` 和旧多 Agent 何时处理? | evidence-driven | 已解决 |
| Q11 | 验收 | 如何证明 event order、single terminal、same IDs、cancel 和 production wiring? | evidence-driven | 已解决 |
## Evidence-driven
| 结论 | 证据来源 | 是否已汇报用户 |
|---|---|---|
| 真正 SSE 固定为状态实时发送、最终 typed content 一次释放,不做 Token/字符切片。 | ISS-014 4.7 | 已汇报 |
| 唯一入口是 `POST /api/chat`;顺序为 `metadata -> status* -> content|failure -> done`。 | ISS-014 4.7、阶段 6B | 已汇报 |
| metadata/status/content/failure/done 使用 named SSE event;payload 不再包含旧 `type/data` 包装。 | ISS-014 最小事件结构 | 已汇报 |
| 当前 Controller 同时拥有旧 ChatService、模型、Tools、SessionManager、同步/伪流式流程和 cached thread pool。 | `ChatController.java` | 已汇报 |
| 当前前端 quick 调 `/chat` JSON,stream 调 `/chat_stream` 并保留大量旧格式 fallback。 | `static/app.js` | 已汇报 |
| `ChatApplicationUseCase` 已提供 observer、同一 run control、typed content 和 stable failure code,但组件尚无完整生产 Bean graph。 | 阶段 6A code + Spring annotation scan | 已汇报 |
| 客户端断开必须通过 observer 得到的 `ChatRunControl` 取消,同一引用需覆盖断开早于 onStarted 的竞态。 | 阶段 6A design + `SseEmitter` lifecycle | 已汇报 |
| `/api/ai_ops` 是独立公开入口,不属于阶段 6B Chat 原子切换;旧实现阶段 7 清理/处置。 | ISS-014 阶段 6B/7 | 已汇报 |
## User-interview
- 无新增 user-interview。唯一 endpoint、破坏性迁移、事件 schema、非 Token 流、自动 Apply/Archive/commit 均由 ISS-014 与用户持续授权冻结。
## Key Decisions
- Controller 使用构造注入的 `ChatApplicationUseCase` 与受控 `TaskExecutor`;不注入 ChatModel、Tools、ChatService 或 SessionManager 来处理 Chat。
- SSE adapter 为每个请求维护单一 state machine 和 `AtomicReference<ChatRunControl>`;disconnect 标记先于 onStarted 时,onStarted 立即取消。
- 正常 result 只产生一次 typed content 和 done;异常只产生 failure 和 done;IOException/timeout/disconnect 只取消,不尝试补发终态。
- executor rejection 在没有 Run 时发送稳定 failure/done;worker 启动后所有 failure code 来自 `ChatApplicationException`,不暴露 cause。
- 生产装配使用集中 `harness.chat` properties 和 Spring-managed bounded executors;所有模型调用复用同一 `ChatModel` 和 `GuardModelCall`。
- 前端移除 mode selector 和 quick path,只保留一个严格 named-event parser;unknown event/schema fail closed。
- 不创建 ADR:L4 方案已经在 ISS-014 设计冻结,本 change 负责原子落地和迁移说明。
## OpenSpec Backfill
- 需进入 proposal/design/spec/tasks:唯一 SSE、五事件 state machine、typed payload、安全 failure、same IDs、disconnect cancellation、bounded executors、production assembly、frontend migration、L4 rollback。
- 非目标:AiOps 协议、旧类物理删除、live E2E。
## Cross-artifact Alignment
| 上游 -> 下游 | 检查内容 | 状态 |
|---|---|---|
| ISS-014/brief -> proposal | 唯一 SSE、五事件、断开取消、前端迁移、生产装配、阶段 7 边界 | 已对齐 |
| proposal -> design | named events、state machine、bounded executors、Bean graph、AiOps 隔离、L4 rollback | 已对齐 |
| design -> specs/tasks | 每项 ownership/lifecycle/safety/migration 均有可观察 requirement 和纵向切片 | 已对齐 |
| specs -> tasks | 10 组 requirements 覆盖 wiring、SSE、cancel、frontend、AiOps regression 和 verification | 已对齐 |
## Architecture Audit
- Capability source: `zoom-out`,使用 Chat Application Use Case、RunContext、Run Lifecycle、Diagnosis Harness 和 Chat SSE Contract 术语。
- 链路为 `browser -> Controller -> SSE session -> bounded worker -> Application Use Case -> Harness -> typed release -> SSE session`;业务真理源不进入 Controller。
- SSE session 独占协议状态,Application 独占 Run/dispatch/persistence,Core 独占 cancel/budget,configuration 独占 infrastructure graph。
- 最大风险是 disconnect/onStarted 与 terminal callback 竞态,design/tasks 已固定 pending-disconnect、atomic terminal 和 no-send-after-close tests。
- L4 部署必须 Controller/frontend 同版本;回滚成对返回阶段 6A commit,V012 可保留。
## Interface Impact
- Level: L4 breaking HTTP/frontend contract。
- 删除同步 JSON `POST /api/chat`、`POST /api/chat_stream` 和旧 Chat SSE wrapper;新增 named-event SSE `POST /api/chat`。
- 消费者:bundled `static/app.js`、外部 Chat callers、Controller/MockMvc tests;Trace/feedback 继续使用 metadata IDs。
- 迁移/回滚:前后端同 commit 原子部署/回滚;不提供兼容开关、双 endpoint 或旧 parser。
## Commit Gate Preflight
- proposal、design、specs、tasks 完整,change strict validation 通过。
- Question pool 全部 evidence-driven 解决并已汇报,无 user-interview、未判级接口或未接受架构风险。
- Cross-artifact 四段对齐无 gap;L4 migration/rollback、production wiring、disconnect race 和阶段 7 边界均进入 design/spec/tasks。
- `openspec-propose` 补全产物,`zoom-out` 完成架构审计;可提交为 Committed OpenSpec。
## Apply Progress
- 1.1-1.4 完成:新增 `ChatHarnessProperties`、bounded worker/model executors 和显式 `HarnessChatConfiguration` graph;空 MySQL datasource 只允许 fail-closed unknown logical ID。
- 2.1-2.4 完成:新增 named-event `ChatSseEvent`/`ChatSseSession`,Chat 唯一 SSE Controller,移除 `/chat_stream` 和同步 Chat;AiOps 模型/Tool 依赖下沉到 service。
- 3.1-3.3 完成:前端移除 quick/stream 双轨,统一 named-event parser、typed renderer 和静态 contract test。
- `JpaChatRunStore` 生产构造器改为注入集中 `PublishedResultPolicy`,避免读写边界漂移;这是实现偏差修复,无需改需求方向。
## Apply Review
- Review 发现 `ChatController` 虽然 Chat path 已经只调用应用用例,但类本身仍承载 AiOps 和 Session 管理依赖,不完全满足“Chat Controller 只负责 Chat 协议”的 ownership 要求。
- 用户确认将职责拆分为 `ChatController`(仅 `/api/chat`)、`AiOpsController`(保持 `/api/ai_ops`)和 `ChatSessionController`(保持 clear/session/runs URL)。
- 该修正保持所有公开 URL、AiOps message-wrapper payload 和 Session observable behavior,不修改 Committed OpenSpec 的范围或方向。
- `styles.css` 中 mode selector/dropdown 死样式随前端双轨删除一并移除。
- AGENTS.md 指定的 `codebase-retrieval` 和 LSP 工具在当前环境不可用;使用 OpenSpec 全量上下文、`rg` 引用检查、Java 编译、结构测试和综合回归完成等价影响面确认。
## Verification and Migration
- Stage 2-6B + AiOps regression:24 suites / 101 tests,0 failure/error/skipped;包含 `HarnessChatConfigurationTest` Spring wiring。
- `mvn -q -DskipTests compile`、`openspec validate single-react-chat-sse-cutover --strict`、`node --check src/main/resources/static/app.js` 均通过。
- 生产 Controller/前端遗留 token 静态扫描为 0;`ChatController` 中模型、Tool、ChatService、AiOps、Session 和 Trace service 禁止依赖扫描为 0。
- L4 部署必须将后端 `/api/chat` SSE 与 bundled frontend consumer 同版本原子部署;回滚必须成对回滚到阶段 6A commit,不提供兼容 endpoint、旧 parser 或双轨开关。
- live 模型、Redis、日志和 MySQL E2E 按 ISS-014 门禁明确延后到阶段 7,本阶段没有把 focused/Mock 验证表述为 live 验收。
@@ -0,0 +1,32 @@
# Evidence: single-react-chat-sse-cutover
## Code and Contract Evidence
- `ChatApplicationUseCase` 已提供 typed result、observer 和 exact `ChatRunControl`,Controller 无需拥有模型、Tool 或业务路由。
- `ChatSseSession` 使用 first-terminal-wins 状态和 pending disconnect,覆盖断开早于 `onStarted`、send failure、late terminal 与正常 completion 竞态。
- `ChatSseEvent` 固定 metadata/status/content/failure/done payload;content 引用阶段 6A typed public content,failure 只公开 stable code/message。
- `HarnessChatConfiguration` 组装同一个 Core、model boundary、canonical store、ToolBoundary、Diagnosis Agent、Guards、Router、executors 和 application use case。
- `ChatHarnessProperties` 集中 worker/model queue、Run、SSE、canonical store、Agent、Router、Guard 和 single-turn limits;executors 使用有限队列与 `AbortPolicy`。
- bundled frontend 只向 `/api/chat` 发起 streaming POST,按完整 named SSE frame 严格解析并 fail closed。
## Ownership Review
- Apply review 将原 Controller 拆为 `ChatController`、`AiOpsController` 和 `ChatSessionController`。
- `ChatController` 只注入 `ChatApplicationUseCase`、bounded worker 和 Chat properties;禁止依赖扫描为 0。
- `/api/ai_ops`、`/api/chat/clear`、session info 和 run list 的 URL 与 payload 字段保持不变。
- mode selector/dropdown DOM、JS state 和 CSS 已全部移除,不保留双轨开关。
## Safety and Lifecycle Evidence
- success 测试验证 `metadata,status,content,done` 严格顺序和 typed content。
- failure 测试验证内部 provider detail 不进入 failure payload,且 content/failure 互斥。
- disconnect-before-start 与 send-failure 测试验证 exact Run 取消和 late terminal 阻断。
- frontend contract test 验证唯一 request target、五类 event branch 与旧 consumer token 删除。
- static scan:生产 Controller/前端 legacy Chat token 0;Chat Controller forbidden dependency 0。
## Verification Evidence
- Stage 2-6B + AiOps regression:24 suites / 101 tests,0 failure/error/skipped。
- `HarnessChatConfigurationTest` 覆盖 bounded queues、共享依赖、空 MySQL fail-closed 和 Spring context wiring。
- Maven compile、strict OpenSpec validation 和 JavaScript syntax check 通过。
- 真实模型、Redis、日志和 MySQL live E2E 按 ISS-014 串行门禁留到阶段 7。
@@ -0,0 +1,36 @@
# Acceptance: single-react-design-freeze
## 实现结果
- 新增类型化 Harness contracts 和序列化契约测试。
- ISS-014 已调整为 11 个串行 changes,并补齐 Tool ID、双状态、previous turn 和安全切换决策。
- Python 查询脚本和 Spring 主配置改为环境变量 Secret。
- 集成测试和历史 handoff 中的旧 Secret 副本已删除或脱敏。
- devflow glossary 新增 Harness、EvidenceGuard、SemanticGuard、Invocation Status 和 Evidence Status。
## 静态验证
- OpenSpec strict validation:通过。
- Secret 扫描:通过,已知真实 Secret 模式零匹配。
- Diff whitespace 检查:通过。
## 脚本验证
- 4 个 contract tests:通过。
- ToolInvocationRecorder、ExecutorGatekeeperService、ChatController focused baseline:通过。
- 查询脚本缺失密码环境变量时 fail fast:通过。
## 浏览器/人工验证
- 不适用;阶段 0 未切换任何公开 UI/API。
## 未验证与后续门禁
- Provider 侧旧凭据轮换待凭据所有者完成。
- Live E2E 留到阶段 7。
- 阶段 1 只有在本 change Archive 和 Git commit 后才能开始。
## 状态
- Stage acceptance: accepted
- OpenSpec archive: authorized by standing user instruction
@@ -0,0 +1,31 @@
# Brief: single-react-design-freeze
## 背景
ISS-014 将 Chat 诊断从多 Agent、Hook、ThreadLocal 和外层重试编排收敛为单体 Diagnosis ReAct Agent、确定性 Harness 和隔离 SemanticGuard。阶段 0 先冻结后续实施共同依赖的契约和安全边界。
## 目标
- 提供可复用的 Diagnosis Draft、Knowledge Answer、Fallback、published result 和 previous turn 类型。
- 固定 framework Tool Call ID、调用生命周期和证据结果双状态。
- 固定取消、重试、no-evidence、阶段门禁和安全凭据策略。
- 保持当前公开 Chat 运行行为不变。
## 范围
- `com.superbiz.agent.harness.contract` 类型层和 focused tests。
- ISS-014 的 11 个串行 OpenSpec changes 台账。
- 主配置、查询脚本和集成测试中的明文 Secret 清理。
- OpenSpec、devflow glossary 和阶段验收基线。
## 非目标
- 不接入新 Harness、Agent 或 Guards。
- 不切换 `/api/chat`,不删除旧运行链。
- 不执行 live E2E。
## 元数据
- Scale: complex
- Parent issue: `ISS-014`
- OpenSpec: `openspec/changes/single-react-design-freeze`
@@ -0,0 +1,20 @@
# Decisions: single-react-design-freeze
## 已确认决策
- ISS-014 保留为总 Issue,实施拆成 11 个独立、串行 sm-flow changes。
- 每个 change 必须完成 Apply、阶段验收、OpenSpec Archive 和 Git commit 后才能进入下一项。
- Apply、Archive 和阶段 Git commit 已获得用户对整个目标的持续授权;仅方向性决策需要暂停。
- `tool_call_id` 使用框架协议 ID,Harness 不生成第二套 ID。
- `status=PROJECTING/READY/ERROR` 表示 invocation lifecycle;`evidence_status=EVIDENCE_FOUND/NO_EVIDENCE/ERROR` 表示结果语义。
- `NO_EVIDENCE` 只能支持限定范围的 `NEGATIVE_OBSERVATION`。
- KNOWLEDGE_QUERY 使用 answer items 绑定 `tool_call_id + document_id`,首版不进入 SemanticGuard。
- `diagnosis_run` 是 previous turn 的持久化真理源;只有最近的 `DIAGNOSIS + SUCCESS` 安全发布结果可用。
- 取消采用分层、可观测语义;同步模型调用不承诺无法证明的立即硬取消。
- 底层重试压为一次 attempt,Router/SemanticGuard 的允许重试只由 Harness 执行。
- 阶段 4 和 6A 不接管公开入口,阶段 6B 才执行原子 SSE 切换。
## 权衡
- contract types 提前落地会增加少量文件,但后续 output type、validator、持久化和 SSE 可以复用,避免多阶段字符串协议漂移。
- Provider 侧凭据轮换是外部前置,阶段 0 只证明工作树已清理,不伪造外部完成状态。
@@ -0,0 +1,25 @@
# Evidence: single-react-design-freeze
## 代码与框架证据
- `ChatController` 当前直接管理会话、模型、工具、执行和伪流式 SSE,证明公开切换必须延后到阶段 6B。
- `ChatService` 当前构建 Planner/Executor/Verifier/Composer 并依赖 ThreadLocal,证明阶段 0 只能创建无运行依赖的 contract types。
- Spring AI Alibaba `Builder` 提供 ToolInterceptor、output type、tool timeout 和调用限制扩展点;`ReactAgent` 提供 interrupt。
- Spring AI 1.1.7 `SpringAiRetryProperties` 构造器默认 `maxAttempts=10`,与 Harness 单一重试所有权冲突,阶段 2 必须压为一次底层 attempt。
- `DiagnosisRun` 当前只有通用 status 和文本 answer,不足以确定性恢复安全 previous turn,因此确认后续保存 `intent/release_outcome/published_result`。
- 历史 ISS-009 和 verifier evidence 归档证明 `NO_EVIDENCE` 只能表达限定范围内无匹配结果。
## 验证证据
- `mvn -q '-Dtest=HarnessContractTest' test`:通过。
- `mvn -q '-Dtest=ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,ChatControllerTest' test`:通过。
- `mvn -q '-Dtest=HarnessContractTest,ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,ChatControllerTest' test`:通过。
- `openspec validate single-react-design-freeze --strict`:通过。
- `git diff --check`:通过,仅有 Windows line-ending warning。
- 已知旧 Secret 模式扫描:零匹配。
- 缺少 `SUPERBIZ_MYSQL_PASSWORD` 时运行查询脚本:在连接前 fail fast,符合预期。
## 未验证
- 未运行模型、Redis、MySQL 或 Milvus live E2E;按阶段门禁留到阶段 7。
- Provider 侧旧凭据是否已经轮换无法由仓库证明,必须由凭据所有者完成,并在阶段 3C/7 前核验。
@@ -0,0 +1,65 @@
# Single React Diagnosis Agent Acceptance
## 结果
已接受。阶段 4 Committed OpenSpec 的 11 项任务全部完成,新链路仅通过内部 Java/test 入口运行,公开 Chat 未切换。
## 验证
### 静态验证
- 检查:新 `harness.agent` 包与 Prompt 搜索 `ThreadLocal|SessionContextHolder|while|SequentialAgent|SupervisorAgent|StateGraph|Planner|Composer|Verifier|raw_response`。
- 结果:通过,无匹配。
- 检查:`git diff --name-only` 限定 Controller、ChatService、AiOpsService。
- 结果:通过,无 diff。
- 检查:OpenSpec artifact status、cross-artifact、L2 interface impact、`.committed`、`.archive-ready`。
- 结果:通过。
### 脚本验证
- 命令:`mvn -q -DskipTests compile`
- 结果:通过。
- 命令:`mvn -q '-Dtest=HarnessToolInterceptorTest,DiagnosisAgentUseCaseTest' test`
- 结果:通过,14 tests。
- 命令:`mvn -q '-Dtest=DiagnosisAgentUseCaseTest,HarnessToolInterceptorTest,DiagnosisHarnessCoreTest,ToolBoundaryTest,ToolAdapterTest,MysqlToolAdapterTest,CanonicalInvocationStoreTest,HarnessContractTest,RagToolContractTest,QueryLogsToolContractTest,MysqlToolContractTest' test`
- 结果:通过。
- 命令:`openspec validate single-react-diagnosis-agent --strict`
- 结果:通过。
- 命令:归档后 `openspec validate single-react-diagnosis-agent --type spec --strict`
- 结果:通过;仓库级 `openspec validate --specs --strict` 另发现前序 `mysql-readonly-tool`、`rag-log-projections` 主规格仍使用旧 scenario 格式,本阶段不跨归档边界修改。
### 浏览器/人工验证
- 不适用。本阶段无 UI、HTTP 或 SSE 行为变化。
### 未验证
- 未启动 Maven 应用、外部模型、MySQL、Redis、Milvus 或日志服务;按 ISS-014 门禁统一留到阶段 7 live E2E。
- 未验证 provider 生产响应始终携带 Token Usage;缺失 Usage 时模型调用次数仍强制,实际 Token 只能在 provider 返回时记录。
## 已完成范围
- 单一 Diagnosis ReactAgent、单 Prompt 和固定 Query/PreviousTurn 输入。
- 框架原生 tool loop、exact Tool Call ID 和三类 Harness adapter bridge。
- 模型/Token/Tool/上下文/Draft 预算与无自动 retry。
- 严格 DiagnosisDraft 输出和无证据停止规则。
- 显式 sessionId/runId audit metadata、可注入 AgentStep Hook、canonical Tool invocation。
- 公开 Chat/AiOps 与旧多 Agent 路径保持不变。
## 已知限制
- Draft 尚未经过 EvidenceGuard/SemanticGuard,不能公开发布;阶段 5 负责释放门禁。
- previous turn 由调用方传入,阶段 6A 才实现同 Session 安全 Run 选择和确定性组装。
- 真实业务 adapter/Spring Bean 装配和公开入口切换分别留给阶段 6A/6B。
- 仓库级全规格 strict validation 仍受阶段 3B/3C 主规格 scenario 标题格式阻塞,建议阶段 7 文档清理时统一修正并复验。
## Bug 修复和诊断
- 编译发现框架 `Interceptor.getName()` 必须实现,已补充稳定名称并回归。
- 首次测试修正了两个测试假设:callback 异常包装类型,以及 ToolResponseMessage 在 Prompt instructions 中的实际位置;生产规格与实现方向未变。
- 提交前自审发现默认 BeanOutputConverter schema 不允许 `conclusion=null`,与冻结的无证据契约冲突;已增加 schema post-process 和 focused regression,未放宽其他 Draft 字段。
## 交接
- 下一步:提交独立阶段 4 commit,再进入阶段 5。
- OpenSpec 归档确认:用户已对 ISS-014 每阶段 Apply、Archive 和 Git commit 提供持续授权;已归档至 `openspec/changes/archive/2026-07-21-single-react-diagnosis-agent`,主规格已同步至 `openspec/specs/single-react-diagnosis-agent/spec.md`。
@@ -0,0 +1,21 @@
# Single React Diagnosis Agent Brief
## 背景
- 用户目标:按 ISS-014 阶段 4 建立唯一的 Diagnosis ReAct Agent,并完成独立 sm-flow 归档与提交。
- 当前问题:Harness Core 和三类 evidence Tool 已就绪,但没有一个内部诊断执行链消费它们;公开 Chat 仍依赖旧的简单/多 Agent 路径。
- 关联 OpenSpec:`openspec/changes/archive/2026-07-21-single-react-diagnosis-agent/`
- devflow 分档:complex
- 需求真理源:`mvp/issues/archived/ISS-014-single-react-agent-harness-aci-ptk-refactor.md`,不重复创建独立 PRD。
## 范围
- 本次要做:单 Prompt、单 Diagnosis ReactAgent、固定 Query/PreviousTurn 输入、Harness Model/Tool interceptors、三类 Tool 注册、严格 DiagnosisDraft 解析、上下文/模型/Tool/Token 预算和内部审计注入。
- 本次不做:公开 Chat 切换、旧链路删除、EvidenceGuard、SemanticGuard、previous turn 选择、DiagnosisRun 持久化和 live E2E。
- 影响区域:`com.superbiz.agent.harness.agent`、`src/main/resources/prompts`、focused tests、OpenSpec/devflow/ISS-014 状态。
## OpenSpec 对齐
- proposal 覆盖状态:已覆盖
- specs 覆盖状态:已覆盖
- tasks 覆盖状态:已覆盖
@@ -0,0 +1,156 @@
# Decisions: single-react-diagnosis-agent
## Discover status
- Checkpoint: Discover
- Capability source: `sm-flow`,使用 `grill-with-docs` 进行代码可证问题澄清,并使用 GitNexus、源码、依赖 sources jar 与 `javap` 核对调用链和框架 API。
- Scale: complex。变更新增内部 Agent/use case,跨越模型、Tool、预算、结构化输出和审计边界,但不切换公开协议。
- `devflow/index.md` 命中 design freeze、RunContext、ACI contracts、canonical store、RAG/log projection 和 MySQL Tool;没有与 OpenSpec 冲突的 ADR。
## Question Pool
| # | 维度 | 问题 | 模式 | 状态 |
|---|---|---|---|---|
| Q1 | 术语 | 单体 Diagnosis Agent 是否包含外层 Graph 或多个报告作者? | evidence-driven | 已解决 |
| Q2 | 边界 | 阶段 4 是否切换公开 Chat 或删除旧多 Agent 链路? | evidence-driven | 已解决 |
| Q3 | 验收 | 如何证明 ReAct 正常轮次不是 Harness retry? | evidence-driven | 已解决 |
| Q4 | 技术 | 如何取得并保留框架原始 tool_call_id? | evidence-driven | 已解决 |
| Q5 | 技术 | 如何在每个模型和 Tool 边界强制 Run 预算? | evidence-driven | 已解决 |
| Q6 | 输出 | 框架生成的 DiagnosisDraft schema 是否与可空 conclusion 契约一致,并会直接返回 DiagnosisDraft? | evidence-driven | 已解决 |
| Q7 | 上下文 | previous_turn 如何进入模型且保持固定 Schema 和有界? | evidence-driven | 已解决 |
| Q8 | 审计 | 如何保留 Run/AgentStep/ToolInvocation 而不引入 ThreadLocal? | evidence-driven | 已解决 |
| Q9 | 验收 | 无证据结果如何停止而不补造根因? | evidence-driven | 已解决 |
## Evidence-driven
| 结论 | 证据来源 | 是否已汇报用户 |
|---|---|---|
| 只新增一个拥有 tool loop 的 `ReactAgent`,无 SequentialAgent、SupervisorAgent 或业务 StateGraph。 | ISS-014 阶段 4;`ChatService` 旧链路反例;Spring AI Alibaba `ReactAgent` API | 已汇报 |
| 阶段 4 只提供内部用例,公开入口保持旧实现,接口影响 L2。 | ISS-014 1174-1197、阶段 0 design freeze、GitNexus `createReactAgent` 引用 | 已汇报 |
| `ToolCallRequest.getToolCallId()` 精确暴露 `AssistantMessage.ToolCall.id()`;ToolInterceptor 可在回调执行前取得该 ID。 | framework sources `ToolCallRequest`、`AgentToolNode` | 已汇报 |
| `ModelInterceptor` 包围每次真实模型调用,适合调用 `beforeModelCall` 并记录响应 Usage;ToolBoundary 已负责 Tool 预算,不能重复计数。 | framework `AgentLlmNode`/`InterceptorChain`;`ToolBoundary` | 已汇报 |
| BeanOutputConverter 可从 DiagnosisDraft 生成格式说明,但默认 schema 不允许冻结契约的 `conclusion=null`;必须 post-process nullable conclusion,且 `ReactAgent.call` 仍返回 AssistantMessage,需要 Harness 严格解析 JSON。 | framework `DefaultBuilder`、`ReactAgent`、本地生成 schema | 已汇报 |
| RunContext 可放入 `RunnableConfig` metadata,由框架控制地传播到 ToolInterceptor;不需要 ThreadLocal。 | `RunnableConfig.addMetadata(String,Object)`、`AgentToolNode` | 已汇报 |
| 现有 `AgentLoggingHook` 优先读取 config metadata 的 sessionId/runId,可作为可选审计 Hook 复用;ToolBoundary 保持 canonical Tool invocation。 | `AgentLoggingHook`、`ToolBoundary` | 已汇报 |
| 原始 Query 不截断;超限 fail closed。PreviousTurn 先按固定 record 序列化并受独立/总上下文字节限制。 | ISS-014 极简上下文与预算规则 | 已汇报 |
| `NO_EVIDENCE` 不是错误,只能形成限定范围的 NEGATIVE_OBSERVATION;Prompt 必须要求 conclusion=null、记录 missing_info 并停止。 | glossary、ACI spec、DiagnosisDraft contract | 已汇报 |
## User-interview
- 本阶段没有新增 user-interview 问题。方向、范围、阶段串行规则、框架 ID、状态语义以及 routine Apply/Archive/Commit 持续授权均已由用户在 ISS-014 评审与前序阶段确认。
## 关键取舍
- 决策:使用框架 `ToolInterceptor` 直接桥接已注册的 Tool 定义与 Harness adapter。
- 原因:Spring AI `ToolCallback` 的调用参数不包含 Tool Call ID,而 Alibaba interceptor 明确提供原始 ID 和运行 metadata。
- 影响:ToolCallback 负责模型可见定义,interceptor 负责受控执行;未知 Tool 仍交给框架 handler 并最终失败,不伪造结果。
- 决策:模型预算由 `ModelInterceptor` 执行,Tool 预算继续由 `ToolBoundary` 执行。
- 原因:避免在 Agent 层和 Tool boundary 双重 reserve。
- 决策:阶段 4 只严格反序列化 `DiagnosisDraft`,不提前实现 EvidenceGuard。
- 原因:字段引用真实性、唯一性和语义支持属于阶段 5;本阶段只验证 Agent 能生成冻结结构并携带框架 IDs。
- 决策:复用现有 Agent Hook 注入点,不复制 AgentStep 持久化实现。
- 原因:阶段 4 不接公开运行态,阶段 6A 再装配真实 Repository 和 Run 持久化。
## OpenSpec 回写
- 需进入 proposal/design/spec/tasks:单 Agent、内部入口、ToolInterceptor 原始 ID、ModelInterceptor 模型预算、ToolBoundary 单点 Tool 预算、严格 JSON 解析、上下文字节限制、审计 Hook 注入、无证据停止、不切换公开入口。
- 不创建 ADR:这些是 ISS-014 已冻结方向和当前阶段可逆的内部装配,不满足新的难逆转架构决策条件。
## Cross-artifact 对齐
| 上游 -> 下游 | 检查内容 | 状态 |
|---|---|---|
| brief/ISS-014 -> proposal | 单 Agent、内部入口、极简上下文、预算、审计、非目标和验收预期 | 已对齐 |
| proposal -> design | Tool/Model interceptor、严格 Draft、无重试、审计注入和公开隔离 | 已对齐 |
| design -> specs/tasks | ID 传播、预算单点、输入输出限制、生命周期所有权和风险缓解 | 已对齐 |
| specs -> tasks | 7 组可观察要求均有输入/Prompt、Tool bridge、Agent use case 和 focused test 切片 | 已对齐 |
## Architecture Audit
- 能力来源:`zoom-out`,使用 glossary 的 Diagnosis Agent、Diagnosis Harness、RunContext、Evidence Status 和 Invocation Status 术语。
- 链路为 `DiagnosisAgentInput + RunContext -> internal use case -> one ReactAgent -> model/tool interceptors -> ChatModel/ToolBoundary -> strict DiagnosisDraft`,没有外层业务 Graph。
- RunContext 拥有预算/取消/生命周期,ToolBoundary/store 拥有 canonical invocation,use case 拥有输入和 Draft 解析,Agent 只拥有诊断语义;数据所有权不重叠。
- 最大耦合风险是 Spring AI Alibaba interceptor/output schema API;设计通过本地 1.1.2.0 sources jar 和实际 schema 生成结果证实,并以 focused framework-loop/schema tests 固定。
- 最大阶段风险是 Draft 在 Guards 前被误用;类名、文档和零 Controller 消费者共同保持 internal/draft 边界,阶段 5 前不得发布。
## 接口影响
- 级别:L2 内部接口。
- 变更对象:新增内部 Java use case/factory/input/limits/interceptors/registry,不修改既有方法签名。
- 消费者:本阶段只有 focused tests;阶段 6A 将成为首个生产消费者。
- 兼容性:公开 HTTP/SSE、Controller DTO、数据库和旧 Chat/AiOps 路径不变,无迁移或回滚要求。
## Pre-apply Research
### 参考实现与框架源码
- `src/main/java/com/superbiz/agent/service/ChatService.java`:`createReactAgent`、`buildChatExecutorAgent` 和旧多 Agent 反例;只复用 Builder 形态,不复用运行职责。
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`:scripted `ChatModel` 测试模式。
- `src/main/java/com/superbiz/agent/harness/core/DiagnosisHarnessCore.java`:模型/Tool/Token/容量边界。
- `src/main/java/com/superbiz/agent/harness/tool/boundary/ToolBoundary.java`:Tool budget 和 canonical record 单点。
- `src/main/java/com/superbiz/agent/harness/tool/adapter/RagToolAdapter.java`、`QueryLogsToolAdapter.java`、`MysqlToolAdapter.java`:阶段 4 唯一允许的 evidence Tool 执行入口。
- `src/main/java/com/superbiz/agent/hook/AgentLoggingHook.java`:可注入 AgentStep audit,优先读取 RunnableConfig metadata。
- Spring AI Alibaba 1.1.2.0 sources:`ReactAgent`、`DefaultBuilder`、`AgentLlmNode`、`AgentToolNode`、`ToolCallRequest`、`InterceptorChain`。
### 技术栈清单
- `ReactAgent.builder()` + 从 `DiagnosisDraft` 生成并修正 nullable conclusion 的 `.outputSchema(...)`,不新建 Graph/SequentialAgent。
- `RunnableConfig.addMetadata(String,Object)` 显式携带 sessionId/runId/RunContext,并设置 `_stream_=false`。
- `ModelInterceptor` 对每次模型调用执行 Core budget;读取 `ChatResponseMetadata.Usage`。
- `ToolInterceptor` 读取 exact framework ID,调用 adapter bridge;`ToolBoundary` 保留唯一 Tool reserve/store 边界。
- Spring `FunctionToolCallback` 仅定义模型可见 Tool Schema/description;直接 callback 执行 fail closed。
- Jackson `ObjectMapper` 负责固定输入 JSON 和严格 DiagnosisDraft 反序列化;UTF-8 字节按 `StandardCharsets.UTF_8` 计算。
- 框架 Hook 列表作为 AgentStep audit 扩展点;阶段 4 不新增 JPA/Redis/Flyway。
### 新建基础设施
- `harness.agent`:Input、Limits、Tool registry/interceptor、Model interceptor、Factory、UseCase 和异常类型。
- `prompts/diagnosis-agent-prompt.md`:唯一 Diagnosis Agent Prompt。
- focused scripted model/tool-loop tests;无需新 Maven 依赖。
### ISS-014 PRD 复用
- 阶段 4 不新建独立 `prd.md`:ISS-014 已完整覆盖问题、用户价值、数据流、Draft Schema、预算、重试、阶段边界和验收;`brief.md` 只索引本切片,不复制总 Issue。
## Commit Gate Preflight
- proposal、design、specs、tasks 完整,`openspec status` 为 complete,`openspec validate single-react-diagnosis-agent --strict` 通过。
- question pool 全部为已汇报的 evidence-driven 结论;无未确认 user-interview、接口等级或风险接受问题。
- Cross-artifact 四段对齐无 gap;架构审计的数据所有权、阶段边界和框架耦合缓解已进入 design/spec/tasks。
- 接口影响为 L2,仅新增内部 Java API;公开 Controller/SSE/JPA/旧 ChatService 保持不变。
- Apply、Archive 和阶段 Git commit 使用用户对 ISS-014 各阶段的持续授权;实现必须严格限制为 Committed OpenSpec。
- `.committed` 已创建,Committed OpenSpec 可进入 Apply。
## Apply Progress
- 首模块对齐:Input/Limits/Prompt/Tool registry/Tool interceptor 已落地,任务 1.2、2.1、2.2 完成;1.1 等待 focused test 后完成。
- 编译首次发现 `ToolInterceptor` 必须实现 `Interceptor.getName()`;分类为代码偏离,已补充稳定名称并复编译通过,无需修改 OpenSpec。
- 单 Agent use case、Model interceptor、strict Draft parser 和 audit Hook 注入已完成;scripted ChatModel 真实执行框架 `model -> Tool -> model` loop。
- 首次测试的两处失败均为测试假设偏差:Spring 会将 callback 异常包装为 `ToolExecutionException`,Tool observation 位于 `Prompt.instructions` 的 `ToolResponseMessage` 而非 `Prompt.getContents()`;已按框架真实 API 修正测试,生产设计未变。
- 阶段 4 focused tests 共 14 个通过,任务 1.1-4.2 完成;剩余任务 4.3 为综合回归、静态范围和 OpenSpec 验证。
## Apply Result
- 新增一个且仅一个 `diagnosis_agent` Factory 和内部 `DiagnosisAgentUseCase`;没有外层业务 Graph、SequentialAgent、SupervisorAgent 或手写 ReAct loop。
- 新增单一 Prompt、固定 `DiagnosisAgentInput(query, previous_turn)`、可配置 UTF-8 限制和严格 `DiagnosisDraft` 解析;不接受完整历史,不执行结构修复或 Agent retry。
- 新增 `HarnessModelInterceptor`,对每个非流式模型轮次执行 Core model/Token budget,并在 late result 返回后复查 Run active 状态。
- 新增 `HarnessEvidenceTools` 和 `HarnessToolInterceptor`,注册三类冻结 Tool,精确传播框架 Tool Call ID,并通过阶段 3B/3C adapter/ToolBoundary 返回有界结果。
- AgentStep 使用可注入框架 Hook 保留,RunnableConfig 显式传播 sessionId/runId;Run 和 Tool canonical invocation 继续由既有 Harness 边界拥有。
- 公开 Controller、ChatService、AiOpsService 无 diff;旧多 Agent 主链路继续保留。
## Apply Verification
- 编译:`mvn -q -DskipTests compile` 通过。
- 阶段 4 focused:`mvn -q '-Dtest=HarnessToolInterceptorTest,DiagnosisAgentUseCaseTest' test`,14 tests 通过。
- 综合回归:`mvn -q '-Dtest=DiagnosisAgentUseCaseTest,HarnessToolInterceptorTest,DiagnosisHarnessCoreTest,ToolBoundaryTest,ToolAdapterTest,MysqlToolAdapterTest,CanonicalInvocationStoreTest,HarnessContractTest,RagToolContractTest,QueryLogsToolContractTest,MysqlToolContractTest' test` 通过。
- OpenSpec:`openspec validate single-react-diagnosis-agent --strict` 通过。
- 静态范围:新 Agent 包无 ThreadLocal、手写 while、外层 Graph、多 Agent 类型或 raw response;Controller/ChatService/AiOpsService diff 为空。
- 未执行 live E2E:按 ISS-014 阶段门禁统一留到阶段 7。
## Archive Result
- 11/11 OpenSpec tasks 完成,`.archive-ready` 已创建。
- 主规格已同步至 `openspec/specs/single-react-diagnosis-agent/spec.md`。
- Change 已归档至 `openspec/changes/archive/2026-07-21-single-react-diagnosis-agent`。
- `devflow/index.md` 和 ISS-014 阶段表已更新为阶段 0-4 archived,下一阶段为 5。
- 提交前 schema 自审发现框架默认 BeanOutputConverter 将 `conclusion` 限为 object,与冻结的无证据 `null` 语义冲突;分类为实现设计遗漏,已回写 archived design,新增 nullable schema post-process 与回归测试,规格行为未改变。
@@ -0,0 +1,29 @@
# Single React Diagnosis Agent Evidence
## 代码与文档证据
| 来源 | 证据 | 结论 | 是否已汇报 |
|---|---|---|---|
| `mvp/issues/archived/ISS-014-single-react-agent-harness-aci-ptk-refactor.md` | 阶段 4 明确单 Agent、内部入口、极简上下文、结构化 Draft、无重试和预算验收 | 本 change 不得切换公开入口或删除旧链路 | 是 |
| `ChatService.java` + GitNexus references | `createReactAgent` 被旧策略入口调用;复杂路径仍创建 Planner/Executor/Verifier/Composer | 新实现必须是独立内部 use case,不能复用旧 Service 运行职责 | 是 |
| Spring AI Alibaba 1.1.2.0 `ReactAgent`/`AgentLlmNode` sources | 框架自带 ReAct loop;非流式 `ModelResponse` 保留 `ChatResponse` Usage | 不手写循环,使用 ModelInterceptor 强制模型/Token 预算 | 是 |
| Spring AI Alibaba `ToolCallRequest`/`AgentToolNode` sources | `ToolCallRequest.getToolCallId()` 来自 `AssistantMessage.ToolCall.id()`,Tool interceptor 在 callback 前执行 | 可以精确传播框架 ID,不生成第二套 ID | 是 |
| Spring AI Alibaba `DefaultBuilder` + 本地 BeanOutputConverter 输出 | 默认 schema 只允许 object conclusion,但冻结契约允许无证据时 conclusion=null | 生成 schema 必须 post-process nullable conclusion,UseCase 仍严格反序列化 DiagnosisDraft | 是 |
| `DiagnosisHarnessCore.java` / `ToolBoundary.java` | Core 提供 model/Token/capacity;ToolBoundary 已负责 Tool reserve/store | Agent 层只 reserve model/context,禁止双扣 Tool budget | 是 |
| 阶段 3B/3C adapters | RAG/log/MySQL 都以 `RunContext + ToolCallRequestEnvelope` 进入 ToolBoundary | Tool registry 可直接桥接既有 adapter,不复制投影或安全策略 | 是 |
| `AgentLoggingHook.java` | 优先从 RunnableConfig metadata 读取 sessionId/runId | Factory 可复用现有 Hook 注入点,不依赖 ThreadLocal | 是 |
## Evidence-driven 结论
- 单 Agent 的可验证边界是一个 `ReactAgent.call`,内部允许正常模型/Tool轮次,但外部没有 Agent retry 或第二个报告作者。
- ToolCallback 不直接提供 Tool Call ID,必须使用 Alibaba `ToolInterceptor`;已由真实 scripted framework loop 证明 ID 进入 ToolResponseMessage。
- `PreviousTurn` 只作为固定上下文,当前 Draft 只允许引用当前 Run Tool result 中的 ID。
- `NO_EVIDENCE` 必须以 `conclusion=null`、限定范围的负向观察和 missing info 结束,不能推导健康或排除根因。
- 阶段 4 的 L2 内部 API 不影响公开消费者;阶段 5/6A/6B 前 Draft 不可发布。
## 实现证据
- 新包 `com.superbiz.agent.harness.agent` 包含 Input/Limits、Prompt loader、Tool registry/interceptor、Model interceptor、Factory 和 UseCase。
- `diagnosis-agent-prompt.md` 是唯一新 Diagnosis Agent Prompt,包含 Tool、证据绑定、停止和精确 JSON 规则。
- 14 个阶段 4 tests 覆盖框架 model -> Tool -> model、ID、错误、未知 Tool、PreviousTurn、审计 metadata、nullable conclusion schema、no-evidence、解析和预算。
- 综合回归同时覆盖 Harness Core、ToolBoundary、3B/3C adapter、canonical store 和冻结 contracts。
@@ -0,0 +1,44 @@
# Acceptance: single-react-evidence-semantic-guards
## Result
- Status: archived
- OpenSpec tasks: 14/14 complete
- Interface impact: L2 internal
- Public protocol: unchanged
## Static Verification
- `git diff --check`:阶段文件无 whitespace error;工作区用户已有 `AGENTS.md`/`CLAUDE.md` 仅报告 line-ending warning,未纳入阶段范围。
- 公开 `controller`、`ChatService.java`、`AiOpsService.java` diff 为空。
- SemanticGuard 包和 prompt 不含 Tool Call ID、raw response、ReactAgent、StateGraph、ThreadLocal 或手写 while。
- 新 guard/release 包不包含业务 Agent loop。
## Script Verification
- `mvn -q -DskipTests compile`:通过。
- Stage-focused `EvidenceGuardTest,SemanticGuardTest,DiagnosisReleaseUseCaseTest`:19 tests,0 failure/error。
- Harness/Tool/Agent regression selection:18 suites / 70 tests,0 failure/error/skipped。
- `openspec validate single-react-evidence-semantic-guards --strict`:通过。
- `openspec validate --specs --strict`:16 passed / 2 failed;失败是前序 `mysql-readonly-tool`、`rag-log-projections` scenario 标题格式,不阻塞本 change,计划阶段 7 统一修复。
## Browser or Manual Verification
- Not applicable。阶段 5 没有 UI 或公开入口变化。
## Not Verified
- 未运行真实 LLM、Redis、日志和 MySQL live E2E;按 ISS-014 串行门禁统一留到阶段 7。
- Provider 是否立即响应 Java thread interrupt 取决于 SDK;Harness 已保证 Future cancel、迟到 Token 记账和迟到结果不释放,阶段 7 需用真实模型观察取消延迟。
## Remaining Work
- 阶段 6A:Chat Application Use Case、Intent Router、PreviousTurn 和 Run persistence。
- 阶段 6B:公开 SSE 原子切换。
- 阶段 7:旧链路清理、全局 spec 格式修复和最终 live E2E。
## Archive
- `.archive-ready`: created
- OpenSpec archive: `openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards`
- Main spec: `openspec/specs/single-react-evidence-semantic-guards/spec.md`
@@ -0,0 +1,32 @@
# Brief: single-react-evidence-semantic-guards
## Background
阶段 4 已能生成带框架 Tool Call ID 的 `DiagnosisDraft`,但草稿在发布前还缺少当前 Run 证据验真、独立语义审查和 fail-closed 释放边界。
## Goal
在不恢复业务 Graph、不切换公开入口的前提下,实现确定性 EvidenceGuard、一次语义不变的结构修复、无 Tool/无记忆的单轮 SemanticGuard,以及只发布原 Draft 或固定 SafeFallback 的内部 release use case。
## Scope
- Draft 结构、Analysis/报告引用和 canonical Tool invocation 验真。
- RAG/log/MySQL verified evidence snapshot。
- Evidence repair 单次调用与语义不变约束。
- SemanticGuard 模型/Token/字节预算、超时、取消、技术重试和严格二元输出。
- SUPPORTED/UNSUPPORTED/unavailable/evidence-failed release policy。
- focused tests、Harness/Tool/Agent 回归和静态范围验证。
## Non-goals
- 不切换 Chat/AiOps/SSE 公开协议。
- 不实现 Intent Router、PreviousTurn、Run persistence 或最终报告渲染。
- 不删除旧 Gatekeeper/Verifier/Composer/多 Agent 链路。
- 不运行 live model/Redis/MySQL E2E;统一留到阶段 7。
## Metadata
- Scale: complex
- Interface impact: L2 internal
- OpenSpec: `single-react-evidence-semantic-guards`
- Parent issue: `ISS-014`
@@ -0,0 +1,144 @@
# Decisions: single-react-evidence-semantic-guards
## Discover Status
- Checkpoint: Discover
- Capability source: `sm-flow`,使用 `grill-with-docs` 做 evidence-driven 澄清;`codebase-retrieval`、LSP 和 GitNexus MCP 在当前会话不可用,降级为既有 GitNexus 研究结论、`rg` 调用点核对和逐文件源码阅读。
- Scale: complex。变更跨 canonical store、三类 Tool projection、模型预算、超时/取消、结构修复、语义审查和释放边界,但不切换公开协议。
- `devflow/index.md` 命中阶段 0-4 的设计冻结、RunContext、canonical store、Tool projection 和 Diagnosis Agent;没有与当前 OpenSpec 冲突的 ADR。
## Question Pool
| # | 维度 | 问题 | 模式 | 状态 |
|---|---|---|---|---|
| Q1 | 术语 | EvidenceGuard 是否判断证据足以推出结论? | evidence-driven | 已解决 |
| Q2 | 边界 | verified snapshot 能读取哪些 canonical 字段,并向 SemanticGuard 暴露哪些内容? | evidence-driven | 已解决 |
| Q3 | 边界 | `NO_EVIDENCE` 能支持哪类 Analysis? | evidence-driven | 已解决 |
| Q4 | 修复 | 首次 EvidenceGuard 失败允许如何修复,是否可重跑 Agent 或 Tool? | evidence-driven | 已解决 |
| Q5 | 语义 | SemanticGuard 是否拥有 Tool、记忆、ReAct loop 或报告写作能力? | evidence-driven | 已解决 |
| Q6 | 重试 | 哪些 SemanticGuard 结果允许第二次 attempt? | evidence-driven | 已解决 |
| Q7 | 超时 | 单轮模型超时和 Run 取消如何终止后台调用? | evidence-driven | 已解决 |
| Q8 | 释放 | 三种 Fallback 能否包含 Draft、reason 或未验真来源? | evidence-driven | 已解决 |
| Q9 | 接口 | 阶段 5 是否切换公开入口或修改 SSE/持久化协议? | evidence-driven | 已解决 |
| Q10 | 验收 | 如何证明不存在隐式 Agent/repair/retry loop? | evidence-driven | 已解决 |
## Evidence-driven
| 结论 | 证据来源 | 是否已汇报用户 |
|---|---|---|
| EvidenceGuard 只验证结构、引用和 canonical invocation 真实性,不判断 Analysis/Conclusion 的语义充分性。 | ISS-014 4.3、阶段 5;glossary `EvidenceGuard` | 已汇报 |
| Store 只能按 `ToolCallKeyFactory.create(runId, toolCallId)` 查找;可引用记录必须是当前 Run、`READY`、含 `agent_result` 且状态为 `EVIDENCE_FOUND|NO_EVIDENCE`。 | `CanonicalInvocationStore`、`CanonicalToolInvocation.isReferencableBy` | 已汇报 |
| Snapshot 解析 `agent_result` 和必要的有界 request 字段,不读取 `raw_response`;输出按 `analysis_id` 分组且不包含 Tool Call ID。 | ISS-014 4.4、10.2、已冻结 Tool contracts | 已汇报 |
| `NORMAL` 只绑定 `EVIDENCE_FOUND`;`NEGATIVE_OBSERVATION` 只绑定 `NO_EVIDENCE`,且零结果范围来自投影/请求。 | ISS-014 4.3、10.2;`EvidenceStatus` glossary | 已汇报 |
| Evidence repair 只在首次物理验真失败后执行一次无 Tool单轮模型调用,不重跑 Diagnosis Agent 或 Tool loop。 | ISS-014 10.3、阶段 5、重试边界决策 | 已汇报 |
| SemanticGuard 复用同一 `ChatModel`,使用全新 `Prompt` 单轮调用,无 Tool/记忆/ReAct;只返回二元 verdict 和审计 reason。 | ISS-014 4.4、阶段 5;Spring AI `ChatModel.call(Prompt)` | 已汇报 |
| 仅超时、传输、解析或 Schema 技术失败可重试一次;`UNSUPPORTED` 是有效业务结果,不重试。 | `HarnessRetryPolicies.strict().semanticGuard()`、ISS-014 重试策略 | 已汇报 |
| `ChatResponseMetadata.Usage` 可复用 Core 的 model/Token 预算;受控 `Future.get(timeout)` 可在超时或 Run 取消时取消任务。 | `HarnessModelInterceptor`、`DiagnosisHarnessCore`、`RunCancellation` | 已汇报 |
| `EVIDENCE_VALIDATION_FAILED` 的来源必须为空;其余 Fallback 只能包含已验真来源,不包含 Draft 或 SemanticGuard reason。 | ISS-014 10.3、阶段 5;`SafeFallback`/`FallbackType` | 已汇报 |
| 阶段 5 只新增内部用例,公开入口切换留到阶段 6B。 | ISS-014 阶段顺序和 6B 门禁 | 已汇报 |
## User-interview
- 本阶段没有新增 user-interview 问题。上述方向、范围、修复次数、模型复用、超时/重试、Fallback 和阶段串行规则均已在 ISS-014 评审及前序对话中由用户确认。
## Key Decisions
- Snapshot 使用 Tool-specific adapter 严格解析冻结 projection;未知 Tool、ID 不一致、projection 反序列化失败或 projection 的 EvidenceStatus 与 canonical record 不一致均 fail closed。
- MySQL 的逻辑数据源和查询范围来自 canonical `request`,行/列和值来自 `agent_result`;RAG/log 以 `agent_result` 自带的稳定来源和范围为准。
- SemanticGuard 使用直接 `ChatModel.call(Prompt)`,不创建第二个 `ReactAgent`;模型调用由受控 Executor 执行,超时/取消时调用 `Future.cancel(true)`。
- Evidence repair 与 SemanticGuard 共用同一模型抽象,但各自拥有独立 prompt、严格输出类型和预算;repair 策略固定一次 attempt,不能嵌套 retry executor。
- Release use case 只在 EvidenceGuard 成功后调用 SemanticGuard;SUPPORTED 返回原 Draft 对象,所有其他路径只返回固定 `SafeFallback`。
- 不创建 ADR:这些是 ISS-014 已冻结架构的阶段实现,不产生新的难逆转跨项目决策。
## OpenSpec Backfill
- 需进入 proposal/design/spec/tasks:EvidenceGuard 规则、Tool-specific snapshot、无 raw/ID 泄漏、repair 单次边界、SemanticGuard 隔离/预算/超时/重试、原样发布与三类 Fallback、L2/公开隔离。
- 不进入本阶段:公开 Chat/SSE 装配、Run 持久化、旧多 Agent 删除和 live E2E。
## Cross-artifact Alignment
| 上游 -> 下游 | 检查内容 | 状态 |
|---|---|---|
| ISS-014/brief -> proposal | 物理验真、结构修复、隔离语义审查、固定 Fallback、阶段边界 | 已对齐 |
| proposal -> design | Store ownership、Tool-specific snapshot、timeout/cancel/retry、唯一报告作者和 L2 影响 | 已对齐 |
| design -> specs/tasks | 每个关键决策均有可观察 requirement 和对应实现/测试切片 | 已对齐 |
| specs -> tasks | 9 组 requirements 覆盖为 Evidence、model boundary、repair/release 和回归验证任务 | 已对齐 |
## Architecture Audit
- 能力来源:`zoom-out`,使用 glossary 的 Diagnosis Harness、EvidenceGuard、SemanticGuard、RunContext、Invocation Status 和 Evidence Status 术语。
- 链路为 `query + Draft + RunContext -> EvidenceGuard/store -> optional repair/shared model boundary -> SemanticGuard/shared model boundary -> original Draft or fixed Fallback`,没有外层 Graph 或第二个 Tool loop。
- Store 拥有物理调用记录,EvidenceGuard 只读并生成 snapshot;Diagnosis Agent 拥有报告语义,repair 只修 ID,SemanticGuard 只审查,release 只做二选一,数据所有权没有重叠。
- 最大运行风险是 provider 忽略 interrupt;design 要求 Future cancel、迟到 Token 记录和 `core.checkActive` 丢弃迟到结果,阶段 6A 再负责 Run 终态。
- 架构审计未发现与阶段 0-4 或 ADR 冲突;所有风险缓解已回写 design/spec/tasks。
## Interface Impact
- 级别:L2 内部接口。
- 变更对象:新增 guard/release records、interfaces、use cases、prompts 和 tests;现有方法签名不变。
- 消费者:本阶段只有 focused tests,阶段 6A 才接入 production application use case。
- 兼容性:公开 HTTP/SSE、Controller DTO、数据库、Redis key schema 和旧 Chat/AiOps 路径不变。
## Commit Gate Preflight
- proposal、design、specs、tasks 完整,`openspec status` 为 complete,`openspec validate single-react-evidence-semantic-guards --strict` 通过。
- Question pool 全部为已汇报的 evidence-driven 结论;无未确认 user-interview、接口等级或风险接受问题。
- Cross-artifact 四段对齐无 gap;架构审计的数据所有权、模型取消和唯一报告作者约束已进入 design/spec/tasks。
- 接口影响为 L2,仅新增内部 Java API;公开 Controller/SSE/JPA/Redis schema/旧 ChatService 保持不变。
- Apply、Archive 和阶段 Git commit 使用用户对 ISS-014 各阶段的持续授权;实现必须严格限制为 Committed OpenSpec。
- `.committed` 已创建,Committed OpenSpec 可进入 Apply。
## Pre-apply Research
### Reference Implementations
- `harness/tool/store/CanonicalInvocationStore.java`、`CanonicalToolInvocation.java`、`ToolCallKeyFactory.java`:当前 Run 物理验真和 lifecycle 真理源。
- `harness/tool/projection/RagResultProjector.java`、`QueryLogsResultProjector.java`、`harness/tool/mysql/MysqlResultProjector.java`:三类 `agent_result` 的唯一生产者和有界字段来源。
- `harness/agent/HarnessModelInterceptor.java`:模型调用前预算、响应 Usage 记账和 late-result active check 模式。
- `harness/retry/HarnessRetryExecutor.java`、`HarnessRetryPolicies.java`:SemanticGuard 两次技术 attempt 与 repair 一次 attempt 的装配边界。
- `harness/core/RunCancellation.java`:取消 callback 注册和 first-cancel 语义。
- `service/ExecutorGatekeeperService.java`:仅参考报告内部引用检查思想,不复用旧 Map/JPA/session DTO 或规则目录。
### Technology Stack
- Jackson strict `ObjectReader` 解析 frozen Draft/Tool projections;Tool-specific adapter 生成 typed snapshot,不透传 JsonNode/raw response。
- Spring AI `ChatModel.call(Prompt)` 执行无 Tool 单轮 guard;`ChatResponseMetadata.Usage` 接入现有 Core Token 预算。
- Java 17 `ExecutorService/Future.get(timeout)` 提供 per-attempt timeout,`Future.cancel(true)` 接入 Run cancellation;Semantic 总时限由 monotonic elapsed time 控制。
- JUnit 5 scripted `ChatModel` 和 in-memory canonical store 作为外部边界 fake;测试只通过 EvidenceGuard/SemanticGuard/Release use case 公共接口断言行为。
- 无 Controller、MQ、JPA、Flyway 或新 Maven dependency;不需要新共享基础设施。
## Apply Progress
- 首模块对齐:EvidenceGuard typed contracts、Draft/reference checks、current-Run canonical validation 和 RAG/log/MySQL snapshot adapters 已完成,tasks 1.1-1.4 完成。
- 8 个 `EvidenceGuardTest` 行为测试通过;快照序列化不含 Tool Call ID 或 `raw_response`,`NO_EVIDENCE` 仅在 `NEGATIVE_OBSERVATION` 下进入带 scope/zero-match 的快照。
- TODO:tasks 2.1-4.2,尚未实现模型边界、SemanticGuard、repair/release 和综合回归。
- REVIEW 发现 corrupted Store key 下 record ID 可能与 Draft reference 不同;分类为代码偏离,已增加 exact canonical ID 校验和回归测试,无需改变规格方向。
- REVIEW 发现 fallback `DiagnosisReleaseResult` 携带完整 snapshot 会通过 `analysis_text` 间接泄漏 Draft;分类为规格安全边界细化,已回写 design/spec,并让 fallback result 强制使用空 snapshot,安全来源只保留在 `SafeFallback.verified_sources`。
## Apply Result
- 新增确定性 EvidenceGuard:校验 typed Draft、唯一 Analysis ID、报告引用、exact current-Run Tool ID、READY lifecycle、agent result、Evidence Status 与 Analysis Kind,并严格展开 RAG/log/MySQL projection。
- verified snapshot 按 Analysis ID 分组,只含稳定 source/scope/timestamp/excerpt/values;SemanticGuard 输入不含 Tool Call ID、Redis key 或 raw response。
- 新增共享 `GuardModelCall`:复用系统 ChatModel,执行 Core model/Token/Run byte budget、per-attempt timeout、total timeout、Future cancellation 和 late-result active check。
- 新增无 Tool/无记忆/无 ReAct 的单轮 SemanticGuard,严格解析 `SUPPORTED|UNSUPPORTED + reason`;仅技术失败使用既有两次 attempt,业务 UNSUPPORTED 不重试。
- 新增一次 EvidenceRepair,只允许 ID/reference 修复;任何可见文本、kind、顺序、human-confirmation 或 limitations 变化都直接 Evidence fallback。
- 新增 DiagnosisReleaseUseCase 和固定 SafeFallbackFactory:SUPPORTED 原样返回 Draft;evidence failed、semantic unsupported/unavailable 均不返回 Draft、完整 snapshot 或审计 reason。
- 公开 Controller、ChatService、AiOpsService、SSE、JPA/Flyway 和旧多 Agent 链路无修改。
## Apply Verification
- 编译:`mvn -q -DskipTests compile` 通过。
- 阶段 5 focused:`EvidenceGuardTest` 9、`SemanticGuardTest` 4、`DiagnosisReleaseUseCaseTest` 6,共 19 tests,0 failure/error。
- 综合回归:阶段 5 + Harness Core/Retry/Tool boundary/store/projections/contracts + Diagnosis Agent,共 18 suites / 70 tests,0 failure/error/skipped。
- OpenSpec:`openspec validate single-react-evidence-semantic-guards --strict` 通过。
- 静态范围:公开 Controller/ChatService/AiOpsService diff 为空;SemanticGuard 无 Tool ID/raw response/ReactAgent/StateGraph/ThreadLocal/手写 while;新 guard/release 包无业务 Agent loop。
- 仓库级 `openspec validate --specs --strict` 为 16 passed / 2 failed;失败仍是前序 `mysql-readonly-tool`、`rag-log-projections` 的 scenario 标题格式,按计划阶段 7 统一修复。
- 未执行 live E2E:按 ISS-014 阶段门禁统一留到阶段 7。
## Archive Result
- 14/14 OpenSpec tasks 完成,`.archive-ready` 已创建。
- 主规格已同步至 `openspec/specs/single-react-evidence-semantic-guards/spec.md`,新增 9 个 requirements。
- Change 已归档至 `openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards`。
- `devflow/index.md` 和 ISS-014 阶段表已更新为阶段 0-5 archived,下一阶段为 6A。
- 未创建 ADR/compound knowledge:本阶段落实 ISS-014 已冻结边界,没有新的跨项目难逆转决策。
@@ -0,0 +1,29 @@
# Evidence: single-react-evidence-semantic-guards
## Code and Contract Evidence
- `CanonicalInvocationStore` 只提供 key lookup;`ToolCallKeyFactory` 冻结 `runId + toolCallId` 隔离,`CanonicalToolInvocation` 冻结 READY/agent_result/evidence status 可引用条件。
- `RagToolResult`、`QueryLogsToolResult`、`MysqlToolResult` 是 Agent-facing 有界 projection,足以构造 snapshot;只有 MySQL 的逻辑数据源/SQL/params 需要从 canonical request 补足。
- `AnalysisKind.accepts` 已冻结 `NORMAL -> EVIDENCE_FOUND`、`NEGATIVE_OBSERVATION -> NO_EVIDENCE`。
- `HarnessRetryPolicies.strict()` 已冻结 SemanticGuard 两次技术 attempt 和 Evidence repair 一次 attempt。
- `ChatModel.call(Prompt)`、`ChatResponseMetadata.Usage` 和 `RunCancellation.onCancel` 支持直接单轮模型调用、Token 记账与 Future cancellation,无需 ReactAgent。
## Confirmed Boundaries
- EvidenceGuard 不判断证据是否支持结论,只验证结构、引用和物理真实性。
- SemanticGuard 只接收原始 Query、移除 Tool ID 的完整 Draft view 和 verified snapshot,不访问 Redis/raw response。
- Evidence repair 只修改 ID/reference;用户可见语义发生任何变化即失败。
- `UNSUPPORTED` 是有效业务结果,不重试;只有 timeout/transport/parse/schema 技术失败可进行第二次 attempt。
- Fallback 不含 Draft、完整 snapshot 或 SemanticGuard reason;Evidence failure 的 verified sources 为空。
## Review Findings
- 增加 canonical record ID 与 Draft reference 的 exact match,防止 corrupted key 映射被误信任。
- Fallback release result 强制使用空 snapshot,防止 `analysis_text` 通过误序列化泄漏 Draft;安全来源只保留在 `SafeFallback.verified_sources`。
## Verification Evidence
- Stage-focused: 19 tests,覆盖 EvidenceGuard 9、SemanticGuard 4、DiagnosisReleaseUseCase 6。
- Regression: 18 suites / 70 tests,0 failure/error/skipped。
- Maven compile 和 change strict validation 通过。
- 公开 Controller/ChatService/AiOpsService 零 diff;SemanticGuard 无 Tool ID/raw/ReactAgent/Graph/ThreadLocal/手写 loop。
@@ -0,0 +1,45 @@
# Acceptance: single-react-harness-run-context
## 实现结果
- 新增显式 `RunContext`、取消信号、deadline 检查、Run Lifecycle 和 first-terminal-wins。
- 新增 caller-supplied limits、线程安全 Model/Tool/Token/Run bytes 预算和容量 CAS。
- 新增 strict typed retry policies/executor,记录每个实际 attempt 并阻止取消/预算异常误重试。
- 新增无 Redis 依赖的 ToolCallKeyFactory,精确保留框架 Tool Call ID。
- Spring AI `spring.ai.retry.max-attempts` 设置为 1;未修改 provider/model routing。
- 未接入旧 Chat/AIOps/Controller、ThreadLocal、JPA、Redis、Agent 或公开协议。
## 静态验证
- `openspec validate single-react-harness-run-context --strict`:通过。
- 新 Harness 包 `rg`:无 ThreadLocal/current-holder/Redis 引用。
- 旧 Chat/AIOps/Controller/JPA 调用链 diff:为空。
- 受保护配置检查:仅新增 `spring.ai.retry.max-attempts: 1`,model routing/provider 保持不变。
## 脚本验证
- `mvn -q -DskipTests compile`:通过。
- Core focused suite:通过。
- 综合回归 suite(阶段 0/1 契约 + 阶段 2 Core + ChatController):通过。
## 浏览器/人工验证
- 不适用。本阶段没有 UI、Controller、SSE 或公开协议变化。
## 未验证
- 未执行真实模型调用、Redis、CLS、MySQL 或客户端断开 live E2E;这些属于后续 Tool store/adapter/最终 E2E 门禁。
- 未将 RunContext 接入旧 ChatService;这是本阶段明确非目标,阶段 6A 才接入。
- `RunContext` 取消对已进入的同步第三方调用仍是协作式;实际 HTTP/JDBC future 取消留给后续 adapter。
- Provider 侧凭据轮换状态不由仓库证明。
## 剩余风险与后续门禁
- 关闭 SDK 隐式 retry 会使旧路径瞬时错误不再自动重试,直到后续 Router/SemanticGuard 接入 Harness;该变化已记录并可通过恢复配置回滚。
- 下一阶段 3A 必须在 ToolInterceptor 接收 RunContext 和框架 Tool Call ID,直接复用 Key/Capacity/Cancellation 门禁。
## 状态
- Stage acceptance: accepted
- OpenSpec archive: archived at `openspec/changes/archive/2026-07-21-single-react-harness-run-context`
- Main spec sync: `openspec/specs/diagnosis-harness-run-context/spec.md`(8 added requirements)
@@ -0,0 +1,32 @@
# Brief: single-react-harness-run-context
## 背景
旧 Chat/AIOps 通过多个 ThreadLocal 和业务方法内状态机传播 session/run/token/retry,无法成为后续 Tool、Agent、Guard 和新入口的稳定共同边界。Spring AI 默认 10 attempts 还会制造未被 Harness 记录的隐藏重试。
## 目标
- 建立显式、结构不可变、可异步传播的 RunContext。
- 集中实现 deadline、协作式取消、线程安全预算和唯一 Run 终态。
- 实现类型化、最多两次 attempt、逐 attempt 记录的 Harness retry。
- 提供阶段 3A 可直接使用的 Tool Call Key Factory 和单 Run 容量计数器。
- 将 Spring AI 底层 retry 压为一次。
## 范围
- 新增 Harness Core/Retry/Tool Store foundation 类型和 focused Fake tests。
- 更新 `application.yml` 的 Spring AI retry 配置。
- 更新 glossary 和 OpenSpec/devflow 档案。
## 非目标
- 不接入旧 ChatService/AiOpsService/Controller/SSE。
- 不读写现有 ThreadLocal,不删除旧实现。
- 不写 `diagnosis_run` 或 Redis,不实现 Tool projection。
## 元数据
- 分档:complex
- 接口影响:L2;另有关闭旧 SDK 隐式 retry 的有意内部行为变化
- 关联 Issue:ISS-014 阶段 2
- 关联 OpenSpec:`openspec/changes/single-react-harness-run-context`
@@ -0,0 +1,110 @@
# Decisions: single-react-harness-run-context
## 规模与入口
- 分档:complex。
- 入口:ISS-014 阶段 2;阶段 0/1 已分别由 Git commit `58c3910`、`4274f33` 固化并 Archive。
- 目标:建立后续 Tool、Agent、Guard 和入口共同依赖的显式 Harness 运行边界,不接旧 ChatService。
## Context
- `devflow/index.md` 命中 design freeze、ACI contracts 和 session/run trace isolation。
- 当前 `SessionContextHolder`、`VerifierContextHolder`、`TokenUsageHolder` 使用 ThreadLocal;新 Harness 禁止复用。
- 当前 ChatService 自行生成 runId、写 `diagnosis_run`、执行两轮 retry loop 并在 finally 清理 ThreadLocal,职责与新 Harness 边界冲突。
- `DiagnosisRun.status` 当前是字符串 `PENDING/RUNNING/SUCCESS/FAILED`;本阶段不修改实体或迁移,由后续应用用例映射。
- Spring AI 1.1.7 `SpringAiRetryProperties` 的配置前缀是 `spring.ai.retry`,默认 `maxAttempts=10`。
## Question Pool
| 维度 | 问题 | 模式 | 证据与结论 | 状态 |
|---|---|---|---|---|
| 术语 | “不可变 RunContext”是否意味着预算/取消也不能变化? | evidence-driven | ISS 要求上下文不可变同时要求计数/取消/终态;采用结构不可变 record + 线程安全单 Run 状态句柄。 | 已解决并汇报 |
| 术语 | Run 与 DiagnosisRun 是否在阶段 2 直接持久化绑定? | evidence-driven | 阶段 2 只建 Core,阶段 6A 才接应用用例;本阶段生命周期为内存执行真理,不改 JPA。 | 已解决并汇报 |
| 边界 | 是否迁移旧 ThreadLocal/ChatService 调用? | evidence-driven | ISS 1109-1111 明确禁止 ThreadLocal 和旧 ChatService 临时适配;只新增零消费者 Core。 | 已解决并汇报 |
| 边界 | 取消是否承诺立即中断同步模型调用? | evidence-driven | 阶段 0 已确认分层取消;阻止后续边界并执行资源回调,不承诺不可证明的硬中断。 | 已解决并汇报 |
| 验收 | 唯一终态如何证明? | evidence-driven | atomic first-terminal-wins lifecycle,并用取消/预算/异常/成功竞态测试证明后续终态不能覆盖。 | 已解决并汇报 |
| 验收 | 异步传播如何证明不依赖 ThreadLocal? | evidence-driven | Fake Tool 在 `CompletableFuture` 线程只接收显式 RunContext,并验证相同 session/run/state handles。 | 已解决并汇报 |
| 技术 | 隐藏重试如何关闭? | evidence-driven | 本地依赖 `SpringAiRetryProperties` 证实 `spring.ai.retry.max-attempts` 默认 10;配置改为 1 并加配置测试。 | 已解决并汇报 |
| 技术 | 尚未校准的预算默认值如何处理? | evidence-driven | ISS 要求集中配置且不伪造数值;Core 接受显式 limits,不内置默认预算,后续 Spring wiring 决定配置值。 | 已解决并汇报 |
| 技术 | Tool Call Key 是否生成 Tool ID? | evidence-driven | 阶段 0/1 冻结框架 ID;Factory 仅验证安全 segment 并拼接,不生成或改写。 | 已解决并汇报 |
## Grill 结论
- 术语、边界、验收和技术问题均可由已确认 ISS、现有代码和本地依赖 API 证明。
- 没有新的产品偏好、公开协议或风险接受度问题需要 `user-interview`;短期关闭旧 SDK retry 的影响已在 proposal 明示。
- `grill-with-docs` 的代码可证问题已先查证并向用户汇报;所有结论均已回写 proposal。
## 能力与工具限制
- Discover 能力来源:`sm-flow` + `grill-with-docs`。
- 当前工具集没有 `codebase-retrieval` 和 LSP;使用 `rg`、源码阅读、本地依赖 `jar/javap`、编译和 focused tests 补足调用链与 API 核对。
## Cross-artifact 对齐
| 链路 | 状态 | 结论 |
|---|---|---|
| brief 目标/范围/非目标 -> proposal | 已对齐 | 显式 context、Core、预算/取消/终态、retry、key/capacity、隐藏 retry 和不接旧 runtime 全部覆盖。 |
| proposal 范围/约束/承诺 -> design | 已对齐 | 类职责、并发语义、first-wins、配置变化、迁移和回滚均有明确设计。 |
| design 决策/接口影响/风险 -> specs/tasks | 已对齐 | 每个状态边界都有 scenario,配置行为变化和 zero-consumer 边界有独立任务与验证。 |
| specs 可观察行为 -> tasks | 已对齐 | 8 条 requirements 分解为 primitives、budget、Core、retry、key、配置和三层验证。 |
## Architecture Audit
- 能力来源:`zoom-out`,使用 glossary 的 RunContext、Run Lifecycle、Run Budget、Harness Retry Policy 和 Diagnosis Run 术语。
- 输入链路为未来 application use case 创建 RunContext,处理链路由 Core/Retry/Key/Capacity 通过显式参数消费,输出为唯一 RunTermination;当前 runtime 不接入。
- RunContext 拥有单 Run 内存状态,应用用例拥有数据库映射,Tool store 拥有 Redis invocation;数据所有权没有重叠。
- Core 不保存全局 Run map、不调用 Agent/数据库/Redis、不实现循环编排,因此不会演变为工作流引擎。
- 最大风险是关闭 SDK retry 对旧路径的短期行为影响;已作为显式配置变更进入 proposal/spec/test 和回滚说明,无 ADR 冲突。
## Commit Gate Preflight
- proposal、design、specs、tasks 完整,OpenSpec status complete,strict validation 通过。
- question pool 无未汇报 evidence-driven 结论、无未确认 user-interview 问题。
- 接口影响 L2;`spring.ai.retry.max-attempts=1` 的有意内部行为变化已明确影响和回滚边界。
- cross-artifact 四段对齐无 gap,架构审计约束已进入 design/spec/tasks。
- Apply 持续授权已存在;执行范围严格限制为新 Harness foundation、配置 override 和 focused tests。
## Pre-apply Research
### 参考实现与反例
- `SessionContextHolder`:ThreadLocal session/run fallback,新 Harness 明确禁止复用。
- `TokenUsageHolder` / `TokenTrackingChatModel`:当前只记录 total token 且依赖 ThreadLocal,后续 model boundary 应改为显式 RunBudget;本阶段不改旧类。
- `ChatService.executeChatComplex`:当前业务方法内创建 Run、两轮 retry、写终态和清理 ThreadLocal,是后续替换对象,不是 Core 参考实现。
- `DiagnosisRun`:现有持久化字段和字符串状态;本阶段只确认映射边界,不修改实体或 Repository。
- Spring AI 1.1.7 `SpringAiRetryProperties`:`spring.ai.retry` 前缀、默认 `maxAttempts=10`,支持精确配置 override。
### 技术栈清单
- Java 17 record 表达结构不可变 context/limits/snapshot/policy/attempt。
- `AtomicReference` 实现 first-reason/first-terminal-wins;`AtomicLong` 实现 capacity CAS;同步临界区维护复合预算一致性。
- `Clock` 和 `Supplier<String>` 注入保证 deadline/ID 可测试,不引入 scheduler 或全局 registry。
- SLF4J 只记录取消 callback 异常,不记录用户输入、Tool payload 或凭据。
- SnakeYAML 直接解析 classpath `application.yml` 验证 retry override,不启动外部 MySQL/Redis/Milvus/模型。
### 新建基础设施
- `harness.core`:RunContext、Cancellation、Lifecycle、Budget、Capacity、Core 和类型化异常/状态。
- `harness.retry`:RetryFailure/Policy/Policies/Attempt/Executor/Exception 与函数接口。
- `harness.tool.store.ToolCallKeyFactory`:纯 Key 构造,不访问 Redis。
- focused unit tests 与 Fake Model/Tool;无需新 Maven 依赖。
### 影响半径
- 新生产包在本阶段保持零消费者。
- 唯一现有运行配置变化为 `spring.ai.retry.max-attempts=1`;Model 路由、provider、Controller、JPA 和 Redis 配置保持不变。
## Apply 结果
- 冲突分类:未发现 OpenSpec 遗漏、代码偏离或方向不确定项;一次自审发现 RetryExecutor 需要无条件拦截预算/取消异常,已回写代码并通过回归测试。
- 新增 `RunContext`、Cancellation、Lifecycle、Budget、Capacity、DiagnosisHarnessCore、typed Retry 和 ToolCallKeyFactory;未接旧 Chat/AIOps/Controller/Redis/JPA。
- Spring AI 全局 retry 已由默认 10 压为 1;Harness strict policies 只允许 Router/SemanticGuard 技术失败一次显式重试。
- 首模块对齐:Run state/budget/Core/retry/key/config 与 design/tasks 全部完成;Tool interceptor/store/Agent/应用用例仍留给后续阶段。
## Apply 验证
- 编译:`mvn -q -DskipTests compile`:通过。
- Core focused:`mvn -q '-Dtest=RunContextTest,RunBudgetTest,DiagnosisHarnessCoreTest,HarnessRetryExecutorTest,ToolCallKeyFactoryTest,SpringAiRetryConfigurationTest' test`:通过(加固后复跑通过)。
- 综合回归:`mvn -q '-Dtest=HarnessContractTest,RagToolContractTest,QueryLogsToolContractTest,MysqlToolContractTest,RunContextTest,RunBudgetTest,DiagnosisHarnessCoreTest,HarnessRetryExecutorTest,ToolCallKeyFactoryTest,SpringAiRetryConfigurationTest,ChatControllerTest' test`:通过。
- 静态 scope:新 Harness 包无 ThreadLocal/current-holder/Redis 引用;旧 Chat/AIOps/Controller/JPA 调用链 diff 为空;模型路由/provider 未改。
- OpenSpec:`openspec validate single-react-harness-run-context --strict`:通过。
@@ -0,0 +1,25 @@
# Evidence: single-react-harness-run-context
## 文档与依赖证据
- ISS-014 4.2/阶段 2 要求显式 `RunContext`、deadline、取消、预算、retry、Key Factory 和 no-ThreadLocal 边界。
- 阶段 0/1 OpenSpec 已冻结框架 `tool_call_id`、两套状态语义和后续阶段串行门禁。
- 本地 Spring AI 1.1.7 `SpringAiRetryProperties` 的 `@ConfigurationProperties("spring.ai.retry")` 默认 `maxAttempts=10`;配置已覆盖为 1。
## 代码证据
- `SessionContextHolder`、`TokenUsageHolder`、`VerifierContextHolder` 当前是旧链路 ThreadLocal;新 `com.superbiz.agent.harness` 包无任何 holder/ThreadLocal/Redis 引用。
- `ChatService.executeChatComplex` 当前自行创建 run、执行两轮 retry、写 `diagnosis_run` 和清理 ThreadLocal;新 Core 不接入该方法,后续应用用例负责迁移。
- `DiagnosisRun` 仍保留现有字符串状态和 JPA Schema;阶段 2 未修改实体、Repository 或数据库。
- `DiagnosisHarnessCore` 不保存全局 Run map;RunContext 结构不可变,Cancellation/Budget/Lifecycle 为同一 Run 的线程安全句柄。
## Evidence-driven 结论
- first-reason-wins cancellation + first-terminal-wins lifecycle 可以用 AtomicReference 实现并跨异步边界共享。
- 复合 Tool/Token 预算需要同步一致性;Run bytes 用 CAS 预留避免并发超限或部分增长。
- SDK 隐藏 retry 压为一次后,Harness 才能记录 Router/SemanticGuard 的显式 attempt;Tool/Diagnosis/Evidence repair 固定一次。
- Key Factory 可直接复用阶段 3A,但不生成或改写框架 Tool Call ID,也不访问 Redis。
## 工具限制
- `codebase-retrieval` 和 LSP 不在当前工具集中;使用 `rg`、源码阅读、`jar/javap`、Maven 编译、YAML 解析和 focused tests 补足核对。
@@ -0,0 +1,26 @@
# Acceptance: single-react-mysql-readonly-tool
## Commit preflight
- OpenSpec strict validation: passed.
- Scope: fail-closed SQL validator, exact allowlist, JDBC read-only executor, bounded MySQL projection, ToolBoundary adapter and query-script safety cleanup.
- Non-goals: public Agent/Chat cutover, metadata discovery, Agent persistence database access, dynamic/tenant authorization and live production datasource provisioning.
- Security prerequisite: script write branch/default external connection values are explicitly included in this change.
## Apply acceptance
- Implemented and verified.
- Static, Maven and script evidence is recorded in `evidence.md`.
- No browser/manual verification applies; this stage adds no UI or public protocol change.
- Residual risk is limited to later live datasource provisioning/driver behavior and stage 4 integration.
## Archive acceptance
- All OpenSpec tasks are complete.
- `.archive-ready` marker is created after focused verification.
- OpenSpec is ready to move to the dated archive directory.
- Archive completed at `openspec/changes/archive/2026-07-21-single-react-mysql-readonly-tool`.
## Remaining work
- Stage 4 Diagnosis Agent integration and later live datasource/driver E2E remain outside this archive.
@@ -0,0 +1,26 @@
# Brief: single-react-mysql-readonly-tool
## Background
阶段 3A 提供了统一 ToolBoundary 和 canonical invocation store,阶段 3B 提供了 RAG/日志投影。3C 需要独立实现安全敏感的只读 MySQL Tool,避免 Agent 生成的 SQL 直接进入数据库。
## Goals
- 使用 JSqlParser 对保守 SELECT 子集进行 fail-closed AST 校验。
- 使用逻辑数据源和 schema/table/column 精确 allowlist 授权。
- 使用参数绑定、只读 JDBC、超时、取消和结果预算。
- 通过阶段 3A boundary 投影为冻结的 `MysqlToolResult`。
- 清理查询脚本的写入分支和默认连接风险。
## Non-goals
- 不接入公开 Agent/Chat 入口。
- 不提供元数据发现、动态授权、租户/行级权限或生产 datasource provisioning。
- 不查询 Agent 自身持久化数据库。
## Classification
- Scale: complex
- Interface impact: L2 internal Harness tool/adapter, plus build dependency and script safety behavior
- Issue: ISS-014 stage 3C
- Change slug: `single-react-mysql-readonly-tool`
@@ -0,0 +1,108 @@
# Decisions: single-react-mysql-readonly-tool
## Discover status
- Checkpoint: Discover
- Capability source: `sm-flow` with local ISS-014, OpenSpec contracts, existing JDBC dependency/configuration and JSqlParser 4.6 already present in the local Maven cache.
- Scale: complex, because this stage combines AST policy, authorization, JDBC resource limits, projection, and security cleanup.
## Evidence-driven findings
1. `MysqlToolRequest` and `MysqlToolResult` are already frozen under `harness.tool.contract`; no public DTO change is needed.
2. The project already has MySQL JDBC/JPA dependencies, but no Agent-facing external read-only executor or SQL policy.
3. JSqlParser 4.6 is available in the local Maven cache and exposes `CCJSqlParserUtil`, `Select`, `PlainSelect`, `Table`, `Column`, `Function`, `JdbcParameter` and visitor adapters compatible with Java 17.
4. The current application datasource points to the Agent persistence database; the new Tool must use an independently configured logical datasource map and must not reuse that datasource implicitly.
5. `scripts/query_mysql.py` currently defaults host/port/user values and contains a non-SELECT commit branch. This violates the ISS-014 security prerequisite and will be changed to read-only, environment-only behavior.
## Question pool
| Dimension | Question | Mode | Conclusion | Status |
|---|---|---|---|---|
| SQL language | Which SQL subset is executable? | evidence-driven | One SELECT, explicit columns, INNER/LEFT JOIN, predicates/group/order, parameter placeholders and allowlisted aggregates. | resolved |
| Security | How is authorization decided? | evidence-driven | Independent exact schema/table/column allowlist; parser acceptance alone is insufficient. | resolved |
| Data source | Can the Agent pass JDBC coordinates? | evidence-driven | No. Only logical data_source IDs are accepted; connection properties remain configuration/Secret data. | resolved |
| Execution | Which JDBC controls are mandatory? | evidence-driven | PreparedStatement, readOnly connection, setMaxRows, query timeout and Run cancellation. | resolved |
| Metadata | Can the Tool discover tables/columns? | evidence-driven | No. SHOW/DESCRIBE/information_schema are rejected. | resolved |
| Compatibility | Does this cut over public runtime now? | evidence-driven | No. Add internal adapter/executor; Diagnosis Agent integration is stage 4. | resolved |
## User-confirmed direction
- Use the frozen `MysqlToolRequest`/`MysqlToolResult` contract.
- Reuse the existing ToolBoundary and canonical invocation store.
- Keep stage boundaries serial: archive and commit 3C before stage 4.
- Do not pause for routine apply/archive/commit confirmation.
## Pre-apply research
### Existing implementations and dependencies
- `src/main/java/com/superbiz/agent/harness/tool/contract/MysqlToolRequest.java`
- `src/main/java/com/superbiz/agent/harness/tool/contract/MysqlToolResult.java`
- `src/main/java/com/superbiz/agent/harness/tool/boundary/ToolBoundary.java`
- `src/main/java/com/superbiz/agent/harness/tool/adapter/QueryLogsToolAdapter.java`
- `src/main/resources/application.yml`
- `pom.xml` (`mysql-connector-j` already present; add JSqlParser 4.6)
- `scripts/query_mysql.py`
### New classes
- `MysqlToolLimits`
- `MysqlDataSourceDefinition` / allowlist value objects
- `MysqlQueryPlan`
- `MysqlSqlValidator`
- `MysqlReadOnlyExecutor` and JDBC implementation
- `MysqlResultProjector`
- `MysqlToolAdapter`
### Risk controls
- Do not create a generic plugin/DSL layer.
- Do not use regex or `startsWith` as SQL authorization.
- Fail closed on parser/visitor uncertainty.
- Keep raw result canonical-only and expose only bounded projection.
## Commit checkpoint preparation
- Proposal scope, design choices, frozen SQL subset, security script cleanup and acceptance scenarios are ready for Commit artifact generation.
## Commit audit
- Capability source: `sm-flow` and local OpenSpec CLI.
- OpenSpec strict validation: passed for `single-react-mysql-readonly-tool`.
- Cross-artifact alignment:
- brief goals/non-goals -> proposal scope: aligned.
- proposal SQL/security boundaries -> design architecture: aligned.
- design validator/executor/projector decisions -> spec requirements: aligned.
- spec scenarios -> tasks for dependency, policy, JDBC, projection, adapter and verification: aligned.
- Interface impact: L2 internal Harness tool/adapter plus JSqlParser dependency and query-script behavior; no public protocol changes.
- Preflight risk accepted: parser ambiguity, datasource isolation, driver cancellation behavior and sensitive result values all fail closed or remain Harness-only.
## Commit gate
- [x] proposal, design, specs and tasks exist.
- [x] strict OpenSpec validation passes.
- [x] all evidence-driven questions are resolved.
- [x] no unresolved interface decision remains.
- [x] `.committed` marker created for Apply.
## Apply result
- Added JSqlParser 4.6 and fail-closed `MysqlSqlValidator`.
- Added immutable logical datasource/allowlist/limit/query-plan/raw-result models and independent `MysqlToolProperties` binding.
- Added `JdbcMysqlReadOnlyExecutor` with read-only connection, PreparedStatement binding, query timeout, max rows, cell/result limits and Run cancellation callback.
- Added `MysqlResultProjector` with sensitive-column redaction, row/cell/total UTF-8 bounds and `NO_EVIDENCE`.
- Added `MysqlToolAdapter` through the existing ToolBoundary; invalid SQL is rejected before database execution.
- Replaced `scripts/query_mysql.py` with environment-only, read-only transaction behavior and pre-connect write/metadata rejection.
## Apply conflicts and corrections
- `COUNT(*)` is represented by JSqlParser as an `AllColumns` parameter in this version; the visitor was corrected to permit only the explicit `COUNT(*)` exception.
- JDBC metadata access cannot be used as a checked-exception stream method reference; the implementation uses an explicit column loop.
- No OpenSpec/design conflict was found; both corrections were implementation details.
## Archive result
- Apply tasks complete and `.archive-ready` created.
- OpenSpec archived at `openspec/changes/archive/2026-07-21-single-react-mysql-readonly-tool`.
- Main capability specification added at `openspec/specs/mysql-readonly-tool/spec.md`.
- Stage 4 may consume the internal adapter only after this stage is committed.
@@ -0,0 +1,29 @@
# Evidence: single-react-mysql-readonly-tool
## Static verification
- `git diff --check`: passed before archive.
- Security scan confirms the query helper has no default external host/port/root user and no `commit()` write path.
- New MySQL Harness code receives only injected logical DataSources and does not reference `spring.datasource` or the application persistence datasource.
- OpenSpec strict validation passed for `single-react-mysql-readonly-tool`.
## Script/build verification
- `mvn -q -DskipTests compile`: passed.
- Focused suite passed: `MysqlSqlValidatorTest`, `MysqlResultProjectorTest`, `JdbcMysqlReadOnlyExecutorTest`, `MysqlToolAdapterTest`, `MysqlToolContractTest`, `ToolBoundaryTest`, `CanonicalInvocationStoreTest`.
- Python syntax compilation passed for `scripts/query_mysql.py`.
- Missing connection environment variables exit before connection with code 2.
- A write SQL invocation is rejected before connection with code 3.
## Security coverage
- Allowed: explicit allowlisted SELECT, parameter placeholders, qualified INNER JOIN and `COUNT(*)`.
- Rejected: write, WITH, subquery, UNION, wildcard projection, unknown table/column, ambiguous column, dangerous function, inline literal, CASE, FOR UPDATE, multi-statement and placeholder mismatch.
- JDBC controls verified: `setReadOnly(true)`, `PreparedStatement`, `setQueryTimeout`, `setMaxRows`, ordered parameter binding and cancellation-before-execution.
- Projection controls verified: max rows, max cell chars, total UTF-8 bytes, sensitive-column redaction, valid bounded JSON and `NO_EVIDENCE`.
## Not verified in this stage
- No live production business datasource was provisioned or queried; ISS-014 explicitly assigns live E2E to a later issue/stage.
- No public Diagnosis Agent/Chat integration was performed; stage 4 will consume the adapter internally.
- JDBC driver timeout/cancel behavior against a real remote MySQL server remains an operational integration risk.
@@ -0,0 +1,25 @@
# Acceptance: single-react-rag-log-projections
## Commit preflight
- OpenSpec strict validation: passed.
- Scope: RAG and Mock query-log projectors/adapters through the existing ToolBoundary.
- Explicit non-goals: real CLS/MCP, MySQL, Diagnosis Agent cutover, public Chat/AIOps/SSE changes, legacy recorder cleanup.
- Main residual risk: legacy log payloads contain sensitive values; projector tests must prove redaction before Agent serialization.
## Apply acceptance
- Implemented and verified.
- Static verification and focused/regression Maven tests are recorded in `evidence.md`.
- No browser/manual verification applies; this stage adds no UI or public protocol change.
- Residual risks: live Redis/CLS integration and Agent cutover remain later stages.
## Archive acceptance
- `.archive-ready` marker is created after all tasks and checks pass.
- OpenSpec is ready to move to the dated archive directory.
- Archive completed at `openspec/changes/archive/2026-07-21-single-react-rag-log-projections`.
## Remaining work
- Real Redis/CLS integration, MySQL projection, Diagnosis Agent cutover and final E2E remain later stages.
@@ -0,0 +1,24 @@
# Brief: single-react-rag-log-projections
## Background
阶段 3A 已经统一 ToolBoundary 和 canonical invocation store。下一阶段需要把现有 RAG 与 Mock 日志工具接入该边界,并把旧工具输出投影为冻结的 Agent-facing ACI 结果。
## Goal
- 提供 bounded RAG evidence projection。
- 提供带完整 logical scope 和 Mock provenance 的 query-log projection。
- 复用阶段 3A 生命周期、Run ownership、tool_call_id、预算、错误和 canonical record。
## Non-goals
- 不接入真实 CLS/MCP。
- 不实现 MySQL projection。
- 不切换 Diagnosis Agent、Chat/AIOps、SSE 或旧 recorder。
## Classification
- Scale: complex
- Interface impact: L2 internal Harness adapter/projector
- Issue: ISS-014 stage 3B
- Change slug: `single-react-rag-log-projections`
@@ -0,0 +1,72 @@
# Decisions: single-react-rag-log-projections
## Discover status
- Checkpoint: Discover
- Capability source: `sm-flow` with `grill-with-docs` codebase evidence; no external service integration required.
- Scale: complex, because two tool adapters share a lifecycle boundary and define bounded Agent-facing output semantics.
## Evidence-driven findings
1. `ToolBoundary` currently accepts `project(String rawResponse)` and already owns Run/ID/authorization/read-only/JSON/budget/lifecycle enforcement.
2. `LookupKnowledgeTool` emits `evidenceBlocks`, `contextPack`, `retrievalTrace`, `rerankTrace`, session domains, and message; these are internal retrieval/audit fields and must not be projected.
3. `QueryLogsTools` emits region, physical log topic, result limit, instance and metrics; the frozen contract requires logical topic/query/lookback scope and `source_kind=MOCK` instead.
4. Existing ACI records already define the required snake_case fields and immutable collections.
5. The request scope must be passed to the log projector through a typed adapter method rather than inferred from raw output.
## User-confirmed direction
- Use the framework-provided `tool_call_id` only.
- Keep lifecycle status and evidence status separate.
- Implement the stage in phases and complete sm-flow archive plus Git commit before the next stage.
- Adopt the request-aware projector adapter for log scope preservation.
## Question pool
| Dimension | Question | Mode | Conclusion | Status |
|---|---|---|---|---|
| Terminology | Are RAG traces and context packs Agent evidence? | evidence-driven | No. They are internal retrieval/audit details and are excluded from projection. | resolved |
| Boundary | How is log scope preserved when the generic projector has no request? | evidence-driven | Typed adapter carries `QueryLogsRequest` into a request-aware projector method. | resolved |
| Provenance | Which log source is implemented now? | evidence-driven | Existing Mock source only; result always records `source_kind=MOCK`. | resolved |
| Negative result | What does an empty query mean? | evidence-driven | `NO_EVIDENCE` for the recorded scope, with no health/problem inference. | resolved |
| Compatibility | Should legacy tools and public paths be changed now? | evidence-driven | No. Add adapters/projectors only; cutover is later. | resolved |
## Risks
- Existing mock messages contain hostnames, pod IDs, SQL literals, and stack-like text; sanitization must happen before projection.
- Collection limits and excerpt limits can make the Agent result incomplete; `truncated` must be explicit.
- The generic boundary API should remain reusable for stage 3C, so request-aware behavior belongs in an adapter or specialized projector interface.
## Discover checkpoint
- Proposal created: `openspec/changes/single-react-rag-log-projections/proposal.md`
- Context and issue evidence recorded.
- No unresolved user-interview question remains for this bounded stage; implementation direction was explicitly accepted in the conversation.
## Commit audit
- Capability source: `sm-flow` and local OpenSpec CLI.
- OpenSpec strict validation: passed for `single-react-rag-log-projections`.
- Cross-artifact alignment:
- brief goals/non-goals -> proposal scope: aligned.
- proposal boundaries and request-aware projector decision -> design: aligned.
- design projection bounds, redaction, scope and adapter ownership -> spec requirements: aligned.
- spec scenarios -> tasks for limits, RAG, logs, boundary integration and verification: aligned.
- Interface impact: L2 internal Harness adapter/projector only; no public protocol or legacy runtime cutover.
- Preflight risks accepted: legacy payload drift fails closed; sensitive log fields are redacted; total projection budget is explicit.
## Commit gate
- [x] proposal, design, specs and tasks exist.
- [x] strict OpenSpec validation passes.
- [x] all evidence-driven questions are resolved.
- [x] no unresolved interface decision remains.
- [x] `.committed` marker created for Apply.
## Archive result
- Apply tasks complete.
- `.archive-ready` marker created.
- OpenSpec archived at `openspec/changes/archive/2026-07-21-single-react-rag-log-projections`.
- Main capability specification added at `openspec/specs/rag-log-projections/spec.md`.
- Next stage remains 3C MySQL projection; no Agent cutover is implied by this archive.
@@ -0,0 +1,25 @@
# Evidence: single-react-rag-log-projections
## Static verification
- `git diff --check`: passed.
- Existing legacy files were checked and have no diff: `LookupKnowledgeTool`, `QueryLogsTools`, `ChatService`, `AiOpsService`, and `ToolInvocationRecorder`.
- OpenSpec strict validation: passed for `single-react-rag-log-projections`.
## Script/build verification
- `mvn -q -DskipTests compile`: passed.
- Focused projection tests: passed (`RagResultProjectorTest`, `QueryLogsResultProjectorTest`).
- Adapter tests: passed (`ToolAdapterTest`).
- Stage regression suite: passed (`CanonicalInvocationStoreTest`, `ToolBoundaryTest`, ACI/Core/retry/key/config/ChatController tests plus the new projection tests).
## Coverage
- RAG: bounded exact excerpts, duplicate document IDs, internal field exclusion, `NO_EVIDENCE`, truncation and framework ID.
- Logs: logical scope, Mock provenance, pattern aggregation, timeline sampling, sensitive value redaction, empty result and bounded output.
- Boundary: adapters use the existing canonical lifecycle and do not return raw payloads.
## Not verified in this stage
- Real Redis connectivity and live CLS/MCP integration.
- Diagnosis Agent/application cutover and end-to-end Maven runtime flow.
@@ -0,0 +1,43 @@
# Acceptance: single-react-tool-invocation-store
## 实现结果
- 新增 canonical record/limits/exceptions/store interface/Redis JSON adapter。
- 新增 ToolBoundary envelope/result、executor/projector interfaces 和稳定错误码。
- 实现 preflight、Run/Tool budget、Run capacity、PROJECTING、READY/ERROR、evidence semantics、TTL 和 UTF-8 limits。
- 旧 ToolInvocationRecorder、JPA、Chat/AIOps、Controller、Repository 和公开协议未修改。
## 静态验证
- `openspec validate single-react-tool-invocation-store --strict`:通过。
- 新 Harness 包 Redis 引用仅为 `RedisCanonicalInvocationStore`。
- legacy recorder/JPA/Chat/AIOps/Controller/repository/resources diff:为空。
- staged diff check 与 Secret scan:提交前执行并通过。
## 脚本验证
- `mvn -q -DskipTests compile`:通过。
- `mvn -q '-Dtest=CanonicalInvocationStoreTest,ToolBoundaryTest' test`:通过。
- 综合阶段 0/1/2/3A suite(含 Harness/ACI/Core/retry/key/config/ChatController):通过。
## 浏览器/人工验证
- 不适用。本阶段没有 UI、Controller、SSE 或公开协议变化。
## 未验证
- 未连接真实 Redis,未做 ACL/network/TTL live 验证;最终 E2E 阶段执行。
- 未接入 Alibaba ToolInterceptor,真实框架 ID 传播留给后续 Agent/application stage。
- 未实现 RAG/log/MySQL projector,留给 3B/3C。
- canonical update 的 read-TTL-write 并发窗口已记录为风险,尚未 Lua/CAS 化。
## 剩余风险与后续门禁
- 新旧 JPA audit 与 canonical store 短期并存,后续 projector 必须只以新 boundary 的 READY record 作为引用来源。
- 下一阶段 3B/3C 必须复用本 ToolBoundary,不复制 Redis 状态机。
## 状态
- Stage acceptance: accepted
- OpenSpec archive: archived at `openspec/changes/archive/2026-07-21-single-react-tool-invocation-store`
- Main spec sync: `openspec/specs/canonical-tool-invocation-store/spec.md`(7 added requirements)
@@ -0,0 +1,29 @@
# Brief: single-react-tool-invocation-store
## 背景
旧 `ToolInvocationRecorder` 依赖 ThreadLocal 和 JPA preview,不能证明完整 Tool 结果、生命周期和当前 Run 所有权。阶段 2 已提供 RunContext/Key/Capacity,需要统一 canonical ToolBoundary 和 Redis store 供后续 projector 复用。
## 目标
- 统一 Pre-Tool 门禁、PROJECTING/READY/ERROR 状态和 evidence semantics。
- 在同一 Redis record 保存 request/raw_response/agent_result、框架 ID、Run、时间和错误。
- 固定 TTL 不续期、容量/结果大小 fail-closed、raw 不静默截断。
- 用 Fake Tool/Projector/Store 覆盖 duplicate、cross-run、no-evidence、error、TTL 和 oversize。
## 范围
- Canonical invocation model/store、Redis JSON adapter、ToolBoundary 和 focused tests。
- 复用阶段 2 Core、budget、capacity、ToolCallKeyFactory。
## 非目标
- 不实现 RAG/log/MySQL projector。
- 不修改旧 recorder/JPA、Chat/AIOps、Controller/SSE 或公开协议。
## 元数据
- 分档:complex
- 接口影响:L2 内部 Harness boundary/store
- 关联 Issue:ISS-014 阶段 3A
- 关联 OpenSpec:`openspec/changes/single-react-tool-invocation-store`
@@ -0,0 +1,106 @@
# Decisions: single-react-tool-invocation-store
## 规模与入口
- 分档:complex。
- 入口:ISS-014 阶段 3A;阶段 0/1/2 已 Archive 并由 `58c3910`、`4274f33`、`6b74990` 提交。
- 目标:统一 ToolBoundary 和 Redis canonical invocation store,不实现 Tool-specific projection。
## Context
- 阶段 1 主规格已冻结 RAG/log/MySQL Agent-facing Request/Result 和 `EvidenceStatus`。
- 阶段 2 主规格已冻结 `RunContext`、预算、取消、Key Factory 和 strict retry。
- 旧 `ToolInvocationRecorder` 依赖 `SessionContextHolder`、JPA `ToolInvocation` 和 500 字符 preview,属于 durable audit 兼容路径,不是 canonical store。
- 既有 `SessionConfiguration` 提供 `RedisTemplate<String,Object>` + JSON serializer;新 store 复用 bean,不新增连接配置。
## Question Pool
| 维度 | 问题 | 模式 | 证据与结论 | 状态 |
|---|---|---|---|---|
| 术语 | canonical invocation 与旧 JPA ToolInvocation 是否同一记录? | evidence-driven | ISS-014 数据分层明确 Redis canonical 保存完整 request/raw/agent,JPA 只做 durable audit;两者分离。 | 已解决并汇报 |
| 术语 | `PROJECTING/READY/ERROR` 与 evidence status 如何组合? | evidence-driven | 阶段 0/1 规格:只有 READY 可为 FOUND/NO_EVIDENCE,ERROR 不可引用;PROJECTING 是内部暂态。 | 已解决并汇报 |
| 边界 | 阶段 3A 是否实现 RAG/log projector? | evidence-driven | ISS-014 3A 明确不实现 Tool-specific projection,3B/3C 单独接入。 | 已解决并汇报 |
| 边界 | Redis 读取是否刷新 TTL? | evidence-driven | ISS-014 固定“创建设置、读取/更新不续期”;更新采用当前剩余 TTL,不恢复初始 TTL。 | 已解决并汇报 |
| 验收 | raw 超限是否静默截断? | evidence-driven | ISS-014 明确 `RESULT_TOO_LARGE` ERROR,raw 不静默截断;agent projection 才可按 projector 预算截断并标记。 | 已解决并汇报 |
| 验收 | 缺失/重复/cross-run ID 如何处理? | evidence-driven | 阶段 3A 任务明确覆盖;Key Factory 保留框架 ID,ToolBoundary 在 store 创建前校验 Run/ID 和 duplicate。 | 已解决并汇报 |
| 技术 | Redis 如何避免 update 重置 TTL? | evidence-driven | 既有 RedisTemplate;begin 使用 setIfAbsent + TTL,update 先读取剩余 TTL 再写回相同/更短 TTL,get 不调用 expire。 | 已解决并汇报 |
| 技术 | 是否修改旧 recorder 以复用新 store? | evidence-driven | 旧链路大量测试依赖 JPA preview/evidence_refs;本阶段零消费者,保持旧 recorder 不变,避免行为回归。 | 已解决并汇报 |
## Grill 结论
- 所有术语、边界、验收和技术问题均由 ISS、阶段规格、旧代码和 Redis 配置事实证明。
- 没有新增产品偏好或兼容性取舍需要 user-interview;并发更新窗口作为已接受风险记录。
- `grill-with-docs` 的代码可证结论已回写 proposal;没有未确认问题。
## 能力与工具限制
- Discover 能力来源:`sm-flow` + `grill-with-docs`。
- 当前无 `codebase-retrieval`/LSP;使用 `rg`、源码、既有测试、本地 Redis 配置和 focused fake tests 进行等价核对。
## Cross-artifact 对齐
| 链路 | 状态 | 结论 |
|---|---|---|
| brief 目标/范围/非目标 -> proposal | 已对齐 | Boundary、canonical record、TTL/size、Fake tests 和 legacy isolation 全部覆盖。 |
| proposal 范围/约束/承诺 -> design | 已对齐 | Redis value、状态机、preflight 顺序、raw/projection 和失败处理均已设计。 |
| design 决策/接口影响/风险 -> specs/tasks | 已对齐 | L2 boundary、TTL 更新窗口、oversize/error、ID/Run 所有权有对应要求和任务。 |
| specs 可观察行为 -> tasks | 已对齐 | 7 条 requirements 分解为 store、boundary、limits/failure 和隔离验证纵向切片。 |
## Architecture Audit
- 能力来源:`zoom-out`,按 RunContext、Invocation Status、Evidence Status、canonical invocation 和 durable audit 术语审计。
- ToolBoundary 只编排一次调用;CanonicalInvocationStore 独占 Redis 状态转换;DiagnosisHarnessCore 独占 Run budget/cancellation;旧 recorder 只写 JPA audit。
- request/raw/agent result 归同一 canonical record,Agent 只获得 ToolBoundaryResult,不存在 raw 旁路。
- Redis adapter 是唯一 Redis 访问点,接口/Fake 不依赖 Redis;后续 3B/3C 可直接复用而不复制状态机。
- 风险集中在 read-TTL-write 并发窗口和 canonical raw 敏感性,已进入 design/spec/limits,无架构或 ADR 冲突。
## Commit Gate Preflight
- proposal、design、specs、tasks 完整,OpenSpec status complete,strict validation 通过。
- question pool 无未汇报 evidence-driven 或未确认 user-interview 项。
- L2 内部接口影响已记录;旧 JPA/Chat/Controller/协议不改。
- cross-artifact 无 gap,所有错误、TTL、size、ID/Run 所有权和隔离要求可由 Fake tests 验证。
- Apply 已获持续授权,范围只包括新 store/boundary 与 focused tests。
## Pre-apply Research
### 参考实现与复用
- 复用 `DiagnosisHarnessCore` 的 active/deadline/Tool budget/Run bytes 门禁。
- 复用 `ToolCallKeyFactory` 精确保留框架 Tool Call ID 并隔离 Run key。
- 复用阶段 0 `InvocationStatus` / `EvidenceStatus`,不创建字符串状态副本。
- 复用 `SessionConfiguration` 提供的 `RedisTemplate<String,Object>` 和 Spring Boot ObjectMapper。
- 旧 `ToolInvocationRecorder`/JPA preview 仅作为 durable audit 反例,本阶段不修改或调用。
### 技术栈清单
- canonical record:Java 17 record + Jackson JSON String,所有状态组合在 record transition 方法中校验。
- Redis create:`ValueOperations.setIfAbsent` + TTL;update:读取当前 remaining TTL 后写回;get 不调用 expire。
- 大小:UTF-8 bytes;record/agent result/store limits 与 RunContext 累计 capacity 双重门禁。
- 测试:Mockito RedisTemplate/ValueOperations 验证 TTL API;In-memory fake store + Fake Tool/Projector 验证 boundary,不连接外部 Redis。
### 新建类型
- canonical model/limits/store/exceptions/Redis adapter。
- ToolCallRequestEnvelope、ToolBoundaryResult、ProjectedToolResult、ToolExecutor、ToolResultProjector、ToolBoundary 和稳定 error codes。
- Redis adapter tests 与 boundary fake tests。
### 影响半径
- 新 package 在阶段 3A 保持零现有消费者。
- Redis 访问只允许出现在 `RedisCanonicalInvocationStore`;旧 Chat/AIOps/Controller/JPA/Recorder 不修改。
## Apply 结果
- 冲突分类:一次 focused test 断言错误(duplicate 场景应调用 `setIfAbsent` 两次)已修正并复跑通过;无规格偏离。
- 新增 canonical record/limits/store exceptions、Redis JSON adapter、ToolBoundary envelope/result/interfaces 和 stable error codes。
- ToolBoundary 已实现 preflight、Core budget/capacity、PROJECTING、raw/projection size、READY/ERROR 和安全返回边界;未接具体 projector。
- 首模块对齐:store/boundary/limits/failure 与 design/tasks 全部完成;RAG/log/MySQL adapters 仍留给后续阶段。
## Apply 验证
- 编译:`mvn -q -DskipTests compile`:通过。
- Store/Boundary focused:`mvn -q '-Dtest=CanonicalInvocationStoreTest,ToolBoundaryTest' test`:通过。
- 综合回归:`mvn -q '-Dtest=CanonicalInvocationStoreTest,ToolBoundaryTest,HarnessContractTest,RagToolContractTest,QueryLogsToolContractTest,MysqlToolContractTest,RunContextTest,RunBudgetTest,DiagnosisHarnessCoreTest,HarnessRetryExecutorTest,ToolCallKeyFactoryTest,SpringAiRetryConfigurationTest,ChatControllerTest' test`:通过。
- 静态隔离:新 Harness 包仅 `RedisCanonicalInvocationStore` 引用 Redis;legacy recorder/JPA/Chat/AIOps/Controller/repository/resources diff 为空。
- OpenSpec:`openspec validate single-react-tool-invocation-store --strict`:通过。
@@ -0,0 +1,20 @@
# Evidence: single-react-tool-invocation-store
## 文档与代码证据
- ISS-014 阶段 3A 明确要求统一 ToolBoundary、canonical invocation、PROJECTING/READY/ERROR、TTL/容量、ID/Run 所有权,且不实现 3B/3C projector。
- 阶段 2 已提供 `RunContext`、Run bytes capacity 和 `ToolCallKeyFactory`,本阶段直接复用。
- 旧 `ToolInvocationRecorder` 使用 JPA preview 和 ThreadLocal fallback;新 canonical record 独立保存完整 request/raw/agent,不修改旧 recorder/JPA。
- 既有 `SessionConfiguration` 提供 `RedisTemplate<String,Object>` JSON bean;`RedisCanonicalInvocationStore` 是新 Harness 包唯一 Redis 引用。
## Evidence-driven 结论
- `setIfAbsent` 确保同一 `runId+toolCallId` 不覆盖;读取不调用 expire;更新使用剩余 TTL。
- canonical record transition 只允许 PROJECTING -> READY/ERROR;READY 只接受 FOUND/NO_EVIDENCE,ERROR 不可引用。
- ToolBoundary 在执行前校验 Run、ID、JSON、授权、只读和预算;raw 只在可信 store 保存,不返回 Agent。
- UTF-8 record/Agent limits 与 Run capacity 双门禁;raw oversize 跳过 projector,Agent oversize 不返回,均产生 RESULT_TOO_LARGE。
- Fake store/Redis mock tests 已覆盖 duplicate、cross-run、unauthorized、writable、execution/projection error、NO_EVIDENCE、TTL、raw/agent oversize。
## 工具限制
- 当前无 `codebase-retrieval`/LSP;使用 `rg`、源码、Maven 编译、Mockito Redis API 和 in-memory fake 完成等价验证。
@@ -0,0 +1,57 @@
# Acceptance: single-react-cleanup-e2e
## Result
Accepted. 阶段 7 的清理、Harness-native audit、文档收口和 live E2E 已完成;最终发布为安全 `FALLBACK`,没有释放未验证 Draft。
## Static Verification
- `openspec validate --all --strict`: 22/22 passed。
- `node --check src/main/resources/static/app.js`: passed。
- `git diff --check`: passed。
- production legacy path、旧 Agent-facing Tool contract 和持久化敏感 payload 扫描:0 matches。
- 最终 Tool 时间窗日志扫描:原始 query、request/response、Tool Call ID、Prompt、stack、debug instrumentation 均为 0;只保留有界 metadata。
## Script Verification
- Deterministic Maven regression: 60 suites / 221 tests,0 failures,0 errors,3 skipped。
- `mvn -q -DskipTests compile`: passed。
- `mvn -q -DskipTests package`: passed。
- `mvn -q -Dtest=QueryLogsToolsTest,QueryLogsResultProjectorTest,CanonicalInvocationStoreTest test`: passed。
- Flyway 9.22.3 repair + migrate: V012 checksum repaired,V013 applied,schema current version 013。
- Maven `spring-boot:run` with `mvp-demo`: Flyway 13 migrations validated,JPA schema validation passed,Tomcat 9900 started,Harness/Audit wiring active。
- `run-payment-timeout-demo.ps1`: strict named SSE validation passed for exact final session/run。
- `scripts/query_mysql.py`: exact diagnosis_run、agent_step、tool_invocation queries passed。
## Final Live Evidence
- sessionId: `mvp-demo-payment-timeout-stage7-20260722-1741`
- runId: `363f481c-33b8-42e7-8699-428a6ec61806`
- SSE: `metadata -> status -> status -> status -> content -> done`
- diagnosis_run: `status=SUCCESS`、`intent=DIAGNOSIS`、`release_outcome=FALLBACK`、`tool_call_count=2`、`total_token_count=11764`。
- AgentStep: 2 exact-run rows,唯一 agent `diagnosis_agent`,`thought IS NULL`,只含角色/数量/Tool name metadata。
- ToolInvocation: `lookup_knowledge` 和 Mock `query_logs` 各 1 行,均 `READY/EVIDENCE_FOUND/success=1`;input/details 只有 framework Tool Call ID、状态和字节数。
## Browser / Manual Verification
- 未执行浏览器点击验收;本阶段公开协议由真实 HTTP SSE 脚本和前端 JavaScript syntax/contract tests 覆盖。
## Remaining Risks
- live 结果为 EvidenceGuard/SemanticGuard 约束下的安全 `FALLBACK`,不是业务根因成功发布;这是允许的最终释放状态。
- `diagnosis_run.step_count` 仍为空,但 exact AgentStep 查询返回 2 行;该历史汇总字段不作为本阶段 release gate。
- `query_logs` 使用 Mock;真实 CLS 与生产业务 `query_mysql` datasource 仍属于后续接入范围。
- `application-local.yml` 含本地内部配置且被 Git 忽略;未 stage、未提交、未输出凭据。
## Migration And Rollback
- L4 endpoint/frontend cleanup 必须整体回滚阶段 7 commit,不恢复双轨 endpoint 或旧 Tool annotations。
- V013 仅幂等增加缺失列/索引;应用回滚时保留新增 nullable 列,避免破坏已写数据,不执行 destructive down migration。
- Flyway repair 已将远端 V012 checksum 对齐当前迁移;V013 保证旧库与 fresh database 最终 schema 一致。
## Archive
- OpenSpec archive: completed at `openspec/changes/archive/2026-07-22-single-react-cleanup-e2e`。
- Delta specs synced: created main `single-react-cleanup-e2e` spec and removed the temporary AiOps-preservation requirement from `single-react-chat-sse-cutover`。
- Parent Issue `ISS-014` was closed and moved to `mvp/issues/archived/` after the stage 7 E2E evidence was accepted.
- Post-E2E runtime-quality findings are tracked by active `ISS-015`; they do not reopen this completed cleanup change.
@@ -0,0 +1,30 @@
# Brief: single-react-cleanup-e2e
## Background
ISS-014 阶段 0-6B 已建立单一 Diagnosis ReAct Agent、Harness、ACI Tool、Guard 和 named SSE,但仓库仍存在 legacy AiOps/Sequential/Redis Session 链、旧 Tool contract、副作用式审计和过时文档。阶段 7 负责物理清理、durable metadata audit 和最终 live E2E。
## Goals
- 公开诊断只保留 `POST /api/chat`,业务代码只保留一个拥有 Tool loop 的 `diagnosis_agent`。
- RAG/log 仅作为 Harness backend;Agent-facing Tool 只来自 `HarnessEvidenceTools`。
- AgentStep 与 ToolInvocation 使用 exact sessionId/runId,长期持久化仅包含有界 metadata。
- 通过 Maven 启动、named SSE、日志与 MySQL exact-run 查询完成最终验收。
## Scope
- 删除 legacy Controller、Service、Hook、ThreadLocal、prompt、前端入口、测试和死文档。
- 新增 Harness-native Agent/Tool durable audit,修正文档、issue 与 OpenSpec strict 缺陷。
- 为旧 V012 数据库增加幂等 V013 兼容迁移,并校准可重复 payment-timeout demo。
## Non-goals
- 不实现真实 CLS 或生产业务 MySQL Tool datasource。
- 不修改三类 Intent、DiagnosisDraft、EvidenceGuard、SemanticGuard 或 Release Policy 语义。
- 不保留 legacy endpoint、兼容分支、Graph 或第二套 Tool Call ID。
## Classification
- Scale: `complex`
- Interface impact: L4 breaking HTTP/frontend cleanup
- OpenSpec: `single-react-cleanup-e2e`
@@ -0,0 +1,111 @@
# Decisions: single-react-cleanup-e2e
## Discover Status
- Checkpoint: Discover
- Capability source: `sm-flow` + `grill-with-docs` + `gitnexus-refactoring`。
- GitNexus index 停在 `2362665`,当前为 `bc36248`;刷新会改写用户已修改的 `AGENTS.md`,因此不安全。使用 OpenSpec/devflow、`rg` 全引用扫描、源码阅读、编译和测试作为 fallback。
- AGENTS.md 指定的 `codebase-retrieval` 与 LSP 工具在当前环境不可用,已显式记录限制。
- Scale: complex。涉及 L4 endpoint 删除、跨模块物理清理、Tool/Trace ownership、安全持久化和 live E2E。
## Question Pool
| # | 维度 | 问题 | 模式 | 状态 |
|---|---|---|---|---|
| Q1 | 范围 | 哪些旧 Agent/Service/Hook/Session 只有自引用测试,哪些仍在生产可达? | evidence-driven | 已解决 |
| Q2 | 协议 | 阶段 7 是否必须删除 `/api/ai_ops` 与旧 Session endpoints? | evidence-driven | 已解决 |
| Q3 | Tool | RAG/log 旧实现哪些可复用,哪些 Agent-facing contract/副作用必须删除? | evidence-driven | 已解决 |
| Q4 | Trace | 新 Harness 如何在不泄漏 raw/prompt/thought 的前提下满足 AgentStep/ToolInvocation E2E? | evidence-driven | 已解决 |
| Q5 | 文档 | ISS-012/ISS-013 和现有架构/Demo 文档如何收口? | evidence-driven | 已解决 |
| Q6 | 验收 | 最终 live E2E 必须证明哪些 exact-run 事实,哪些外部系统明确不声称 live? | evidence-driven | 已解决 |
| Q7 | 回滚 | L4 endpoint 删除如何迁移与回滚? | evidence-driven | 已解决 |
## Evidence-driven
| 结论 | 证据来源 | 是否已汇报用户 |
|---|---|---|
| `ChatService` 没有生产调用方,只剩自身单元/Smoke tests。 | `rg ChatService` | 已汇报 |
| `/api/ai_ops` 仍由 bundled frontend 按钮调用,并运行 Supervisor + Planner + Executor 多 Agent。 | `AiOpsController`、`AiOpsService`、`app.js`、`index.html` | 已汇报 |
| ISS-014 总体验收要求旧多 Agent/Graph 不存在,阶段 6B 只暂时保持 AiOps,阶段 7 负责清理/处置。 | ISS-014、阶段 6B design/acceptance | 已汇报 |
| `/api/chat/clear` 与 session info/runs 没有 frontend caller;Redis SessionManager 只由该 Controller 和测试使用。 | controller/frontend/session 引用扫描 | 已汇报 |
| `LookupKnowledgeTool` 和 `QueryLogsTools` 被新 Harness adapter 复用,但仍携带旧 `@Tool`、ThreadLocal/recorder 副作用。 | `HarnessChatConfiguration`、Tool source | 已汇报 |
| 新 ToolBoundary 写 Redis canonical invocation,但没有 `tool_invocation` durable audit;最终 DB E2E 会缺 Tool rows。 | Harness boundary/config 引用扫描 | 已汇报 |
| 旧 `AgentLoggingHook` 仍回退 ThreadLocal,并持久化 thought/model正文;不满足新安全边界。 | `AgentLoggingHook.java` | 已汇报 |
| 当前架构、Agent、Harness 和 lifecycle 文档仍描述 Planner/Executor/Verifier/Composer 与 AIOps 双入口。 | `mvp/architecture/*.md` | 已汇报 |
## User-interview
- 无新增 user-interview。唯一公开 Chat、旧多 Agent/Graph 物理删除、无兼容分支、最终 E2E 和外部 Mock 边界均已由 ISS-014 与用户的逐阶段自动执行授权冻结。
## Key Decisions
- 删除 legacy AiOps endpoint 而不是迁移到第二个 use case;所有诊断统一进入 `/api/chat` 的 Intent Router。
- 删除旧 Session endpoints/Redis conversation context;安全 PreviousTurn 只来自 `diagnosis_run.published_result`。
- RAG/log 查询实现保留为 Harness backend,移除 `@Tool` 和旧 recorder/session dedup;Agent 只看 ACI callbacks。
- 新 Agent audit hook 只写角色/数量/Tool 名称/耗时等 metadata,不写模型输入正文、输出正文、Thought 或 Tool arguments。
- Tool durable audit 通过 Harness port + JPA adapter fail-open 写入;Redis canonical store failure 仍 fail-closed,DB audit failure 只记录日志,不改变 Tool observation。
- 删除 endpoint 的迁移无兼容层;bundled frontend 同 commit 删除按钮/consumer,回滚整体回滚 commit。
- 不创建 ADR:方向已由 ISS-014 冻结,本 change 只完成最终落地与验收。
## OpenSpec Backfill
- 需进入 design/spec/tasks:删除清单、唯一 endpoint、Harness-native trace、Tool durable audit schema/safety、Tool backend 解耦、文档/issue 收口、strict validation、live E2E/log/DB acceptance 与回滚。
## Cross-artifact Alignment
| 上游 -> 下游 | 检查内容 | 状态 |
|---|---|---|
| ISS-014/brief -> proposal | 阶段 7 物理清理、文档、最终 Maven/log/DB E2E、Mock 外部 Tool 边界 | 已对齐 |
| proposal -> design | 删除闭包、Tool backend 复用、安全 Agent/Tool audit、L4 migration/rollback | 已对齐 |
| design -> specs/tasks | ownership、禁止泄漏、exact identity、strict validation、live E2E 均有 requirement 与切片 | 已对齐 |
| specs -> tasks | 8 组可观察 requirements 覆盖删除、audit、docs、verification 和最终 E2E | 已对齐 |
## Architecture Audit
- Capability source: `zoom-out`,使用 Diagnosis Harness、Diagnosis Agent、Canonical Invocation、Durable Audit、Diagnosis Trace 和 Chat SSE Contract 术语。
- 最终链路为 `browser -> ChatController -> ChatApplicationUseCase -> Harness -> Diagnosis Agent -> ACI Tools -> Guards -> SSE`,没有第二业务入口或业务 Graph。
- Redis canonical invocation 是短期完整 Tool 真理源;MySQL ToolInvocation 是长期有界 metadata audit,两者禁止双写 raw payload。
- `chat_session` JPA entity 属于当前 Run/PreviousTurn 目录,Redis SessionContext 属于旧 conversation memory;删除时必须区分。
- 最大风险是 backend 的旧 Tool annotation/recorder 隐式暴露和 audit 内容泄漏;design/tasks 已加入独占 discovery、negative serialization、context startup 与 live DB inspection。
## Interface Impact
- Level: L4 breaking HTTP/frontend contract。
- 删除 `/api/ai_ops`、`/api/chat/clear`、`/api/chat/session/{sessionId}` 和 `/runs`;保留唯一 `/api/chat` SSE 与 Trace API。
- bundled frontend 同 commit 删除 AiOps 按钮/consumer;外部 caller 迁移到 `/api/chat`,无兼容 branch。
- 回滚必须整体回滚阶段 7 commit,不能单独恢复旧 endpoint/Tool annotations。
## Apply Verification Status
- Follow-up current-document audit found and corrected stale runtime semantics in the tracked payment-timeout PowerShell demo and `mvp/tables`: old JSON Chat parsing, Redis SessionContext history, Planner/Verifier identities, full Tool payload persistence, and Chat/AiOps Run descriptions are no longer presented as current behavior.
- `mvp/demo/scripts/run-payment-timeout-demo.ps1` now strictly validates `metadata -> status* -> content|failure -> done`, captures exact session/run identity, and fetches exact Trace without writing feedback. PowerShell parser validation and two in-memory SSE contract samples passed.
- Current architecture/demo/table scan excluding explicit `archive/` and ignored local `output/` artifacts: 0 legacy runtime matches.
- Follow-up `openspec validate --all --strict`: 22 passed, 0 failed; `git diff --check`: passed.
- Initial environment checks reported missing injected variables. Per user direction, credentials were restored to ignored `application-local.yml`; no credential was added to Git or emitted in the archive.
- `mvn -q -DskipTests compile`: passed.
- `node --check src/main/resources/static/app.js`: passed.
- `openspec validate --all --strict`: 22 passed, 0 failed.
- Deterministic regression suite excluding credential-dependent MySQL/Redis/Milvus tests: 60 suites, 221 tests, 0 failures, 0 errors, 3 skipped.
- `SemanticGuardTest` interruption assertion failed once under the credential-dependent full-suite run, then passed three isolated repetitions and the deterministic regression suite; classified as load-sensitive test timing, not a reproduced product regression.
- `mvn -q -DskipTests package`: passed.
- Production legacy path scan, audit sensitive-payload scan, and legacy backend Tool contract scan: 0 matches.
- The approved external-network run connected to MySQL/Redis/Milvus/model services. Flyway repair aligned the old V012 checksum and V013 reconciled the missing release-contract columns/index.
- Maven startup validated all 13 migrations, passed JPA schema validation and started Tomcat 9900 with Harness/Audit beans.
- Final exact live run: session `mvp-demo-payment-timeout-stage7-20260722-1741`, run `363f481c-33b8-42e7-8699-428a6ec61806`, SSE `metadata -> status -> status -> status -> content -> done`, outcome `FALLBACK`.
- Exact MySQL evidence: one `DIAGNOSIS/SUCCESS/FALLBACK` run, two metadata-only `diagnosis_agent` steps with null Thought, and two READY/EVIDENCE_FOUND Tool audits for `lookup_knowledge` and Mock `query_logs`.
- Tasks 4.2-4.5 are complete. Real CLS and production business MySQL Tool datasource remain explicit non-goals.
## Apply Conflict Classification
- **Code deviation**: `DiagnosisRun` enum mapping expected native ENUM while V012/V013 define VARCHAR. Fixed ORM column definitions; OpenSpec unchanged.
- **Code deviation**: production `ObjectMapper` lacked Java Time modules, causing canonical Redis `STORE_ERROR`. Fixed mapper registration and bound the store test to the production mapper.
- **Code deviation**: Mock empty log results used `success=false`, conflicting with the specified `NO_EVIDENCE` projection. Fixed backend success semantics and the stale test assertion.
- **Code deviation**: RAG backend logs exposed raw query/rewritten query content. Replaced with bounded counts/category metadata and verified the final Tool execution window contains no prohibited payload.
- **Acceptance fixture drift**: the payment-timeout demo requested the removed metrics Tool and encouraged an unbounded investigation. Updated the fixture to the current two-Tool Mock acceptance scope and explicit stopping boundary.
## Commit Gate Preflight
- proposal、design、两份 specs 和 tasks 完整;新 capability 与 modified capability 的范围无 gap。
- Question pool 全部 evidence-driven 并已汇报,无 user-interview、未判级接口或未接受架构风险。
- 删除清单区分 current JPA metadata 与 legacy Redis Session,Tool backend 与 Agent-facing contract,canonical truth 与 durable audit。
- live E2E 明确要求真实应用/模型链和 exact ID;真实 CLS/生产业务 MySQL 明确不在验收声称范围。
@@ -0,0 +1,26 @@
# Evidence: single-react-cleanup-e2e
## Repository Evidence
- 生产引用扫描确认旧 `ChatService` 仅剩自身测试,`/api/ai_ops` 与旧 Session endpoints 属于第二条公开/状态链。
- Harness adapters 复用 `LookupKnowledgeTool`/`QueryLogsTools` backend;旧 `@Tool`、ThreadLocal、recorder 和 topic discovery contract 已移除。
- `HarnessAgentAuditHook` 只写 message count/roles、Tool names、text presence 和 duration;`thought` 保持空。
- ToolBoundary durable audit 只写 exact identity、状态、稳定错误码、耗时和字节数;Redis canonical invocation 仍是短期完整 Tool 真理源。
## Live Findings
- 远端库已登记旧 V012 checksum,但缺少 `diagnosis_run.intent/release_outcome/published_result`。Flyway repair 后,幂等 V013 补齐三列与索引,fresh database 上为 no-op。
- Hibernate 6 将无 `columnDefinition` 的字符串枚举校验为原生 ENUM;`DiagnosisRun` 已显式映射到 V012/V013 的 `VARCHAR(32/16)`。
- 生产 `WebConfig` 的裸 `ObjectMapper` 无法序列化 canonical record 的 `Instant`,导致所有 Tool 在 begin 阶段返回 `STORE_ERROR`;改为自动注册模块,并让 store 测试使用生产 mapper。
- Mock `query_logs` 把 0 命中错误表达为 `success=false`,与 `NO_EVIDENCE` contract 冲突;现以成功查询 + 空数组表达无证据。
- live 日志发现 RAG backend 打印原始 query/rewrittenQuery/keywords;已改为字符数、命中数、类别数和耗时 metadata。
## Final Exact-run Evidence
- sessionId: `mvp-demo-payment-timeout-stage7-20260722-1741`
- runId: `363f481c-33b8-42e7-8699-428a6ec61806`
- SSE: `metadata -> status -> status -> status -> content -> done`
- outcome: `FALLBACK`; diagnosis_run: `DIAGNOSIS / SUCCESS / FALLBACK`
- AgentStep: 2 rows,均为 `diagnosis_agent`,`thought IS NULL`,model input/output 仅 metadata。
- ToolInvocation: 2 rows,`lookup_knowledge` 与 `query_logs` 均为 `READY/EVIDENCE_FOUND`,同一 exact identity,无错误。
- `query_logs` 明确为 Mock;Agent-facing `query_mysql` 未配置生产业务 datasource,也未声称 live。
@@ -0,0 +1,68 @@
# Diagnosis 信息增益停止契约 验收
## 结果
已接受。OpenSpec tasks 37/37 完成;用户确认归档 OpenSpec、提交并推送。
## 验证
### 静态验证
- 命令/检查:`openspec validate diagnosis-information-gain-stop-contract --strict`
- 结果:passed
- 备注:Committed OpenSpec 与最终 tasks 一致
- 命令/检查:OpenSpec delta → main specs 同步(6 个 capability)
- 结果:passed
- 备注:新建 `openspec/specs/diagnosis-information-gain-stop-contract/`,并更新 5 个既有 main specs
- 命令/检查:对照 OpenSpec 与代码路径(progress tracker、interceptor、release、trace)
- 结果:passed
- 备注:Task 8 行为与 design/spec 对齐
### 脚本验证
- 命令:`mvn -q "-Dtest=DiagnosisProgressTrackerTest,HarnessToolInterceptorTest,DiagnosisReleaseUseCaseTest,DiagnosisAgentUseCaseTest,HarnessChatConfigurationTest" test`
- 结果:passed(exit 0)
- 备注:覆盖协议 violation_type、可修正 observation、连续协议错误 STOP、Release fail-closed、Agent 受控停止
- 命令:历史全量回归 `mvn -q -Dtest='!MilvusConnectionTest' test`(tasks 6.2/7.5 阶段)
- 结果:passed(`Tests=292, Failures=0, Errors=0, Skipped=3`)
- 备注:未设置 `MILVUS_TOKEN` 时 `MilvusConnectionTest` 失败属外部凭据边界,非本变更回归
### 浏览器/人工验证
- 步骤:Maven 启动真实应用;Query `诊断切换企业失败的问题`;named SSE + logs + `scripts/query_mysql.py` 按 exact sessionId/runId 核对
- 结果:passed
- 备注:
- `sessionId=iss016-final-20260726-a`
- `runId=3ab22ed7-d0ed-45d8-b928-dce5790c0542`
- SSE:`SAFE_FALLBACK` / `MISSING_REQUIRED_CONTEXT` / `done.outcome=FALLBACK`
- DB:`status=SUCCESS`、`intent=DIAGNOSIS`、`release_outcome=FALLBACK`、`tool_call_count=0`、`total_token_count=2890`
- Trace:`RUN_STARTED -> ROUTING_* -> AGENT_MODEL_STEP -> EVIDENCE_GUARD_INITIAL -> RELEASE_DECISION/FALLBACK -> RUN_FINISHED/FALLBACK`
### 未验证
- 本轮 Archive 未再重跑全量 `mvn test` 与完整 live SSE E2E;依赖 Apply 阶段记录与本轮 focused tests。
- 真实 Provider 下连续协议错误的 live E2E 未单独复跑;协议停止由 focused/scripted loop 覆盖。
- ISS-015 阶段 2(Evidence Repair Schema)与阶段 3(Reasoning 治理)不在本 change 范围。
## 已完成范围
- 信息增益停止、scope 去重、STOP_REQUIRED、ProgressSnapshot、统一 Release
- Token 审计与 Tool 拒绝 Trace
- 协议修复反馈 + `PROGRESS_PROTOCOL_VIOLATED` 兜底停止
- 文档:ISS-015 阶段 1、ISS-016、架构文档、glossary 术语、devflow 档案
## 已知限制
- 重复检测只比较确定性 `tool_name + normalized_scope`,不做自然语言语义去重。
- 协议错误阈值默认 2,与无增益阈值独立配置。
- 无安全 ProgressSnapshot 的受控停止继续 fail closed,不伪造用户可见事实。
- 公开 SSE/前端协议无新增字段;模型侧 Tool Envelope 是已确认 L3 变更。
## 交接
- 下一步:OpenSpec 已用户确认归档;代码提交并推送到当前分支。
- OpenSpec 归档确认:用户确认归档(“执行,完后提交推送”)
- 归档位置:`openspec/changes/archive/2026-07-27-diagnosis-information-gain-stop-contract/`
@@ -0,0 +1,45 @@
# Diagnosis 信息增益停止契约 Brief
## 背景
- 用户目标:未知问题、空证据或缺少查询条件时,Diagnosis 能以正常业务 Fallback 结束,而不是空转到预算耗尽或 `INTERNAL_FAILURE`。
- 当前问题:停止主要依赖模型自觉结束或硬预算;缺少信息增益回传、确定性饱和停止、协议修复反馈和统一 Release。
- 关联 OpenSpec:`openspec/changes/diagnosis-information-gain-stop-contract/`(归档后见 `openspec/changes/archive/2026-07-27-diagnosis-information-gain-stop-contract/`)
- 关联 Issue:ISS-016(承接 ISS-015 阶段 1 硬停止)
- devflow 分档:complex
- 接口影响:L3(模型可见 Tool Envelope 有意变更;公开 HTTP/SSE 与业务 Tool backend 不变)
## 范围
### 本次要做
- Run 内二值信息增益 `GAINED / NO_GAIN`、连续无增益阈值与 `COLLECTING / SATURATED`。
- 服务端注册的 Tool Call Envelope:`previous_observation + input`。
- 确定性 `NO_GAIN`(`NO_EVIDENCE`、重复 `tool_name + normalized_scope`)。
- 一次 `STOP_REQUIRED` 收尾机会与受控停止异常。
- Canonical 控制视图 / 模型白名单观察双视图;RAG 保留 `relevance_level`。
- Tool loop 结束时一次性投影 `ProgressSnapshot`。
- `DiagnosisReleaseUseCase` 统一有结论、无结论、信息饱和、预算终止、协议违规停止的发布。
- 可修正 `INVALID_PROGRESS_PROTOCOL` observation,以及独立 `PROGRESS_PROTOCOL_VIOLATED` 兜底停止。
- 模型 Token 组件/轮次审计与 `TOOL_REQUEST_REJECTED` 安全 Trace。
- Prompt、配置、focused/回归测试与真实 named SSE E2E。
### 本次不做
- `new_count`、`next_action`、多级质量分数、独立 Judge。
- 自然语言语义去重。
- 第二套诊断生命周期状态。
- 公开 HTTP/SSE 字段或前端进度协议新增。
- ISS-015 Reasoning 原文审计治理与 Evidence Repair Schema 注入。
### 影响区域
- `harness.progress`、`HarnessToolInterceptor`、`HarnessEvidenceTools`
- `DiagnosisAgentUseCase` / Prompt / Release / Application 预算兜底迁移
- Trace 审计、配置绑定、ISS-015/016 与架构文档
## OpenSpec 对齐
- proposal 覆盖状态:已覆盖
- specs 覆盖状态:已覆盖(6 个 capability delta,已同步 main specs)
- tasks 覆盖状态:已覆盖(37/37 完成)
@@ -0,0 +1,188 @@
# Diagnosis 信息增益停止契约 Decisions
## Discover Status
- Checkpoint:Discover。
- Capability source:`sm-flow` 内置 Discover 协议;grill 使用 `grill-with-docs`,代码可证问题通过源码、测试和引用搜索处理。
- Scale:`complex`。变更跨越 Agent、Tool 协议、Run 生命周期、Release、Guard、配置、Trace 和 E2E。
- 接口影响:L3。模型可见 Tool Schema 发生有意协议变更,公开 HTTP/SSE 和业务 Tool backend 协议不变。
- 工具降级:当前没有 `codebase-retrieval` 和 LSP 工具;以 `rg`、源码和测试引用核查替代。用户已明确“可以忽略gitnexus”。
## Question Pool
| # | 维度 | 问题 | 模式 | 状态 |
|---|---|---|---|---|
| Q1 | 术语 | Tool 客观状态、信息增益、收集状态、停止原因和最终发布状态是否应合并为一个枚举? | user-interview | 已解决 |
| Q2 | 语义 | Tool 返回价值由谁判断,是否需要多级质量分数? | user-interview | 已解决 |
| Q3 | 协议 | 模型继续调用 Tool 时如何回传上一轮信息增益,是否需要 `next_action`? | user-interview | 已解决 |
| Q4 | 边界 | Tool Schema 应来自 Prompt 还是服务端原生 Tool Calling 注册? | user-interview | 已解决 |
| Q5 | 边界 | Harness 能确定性判定哪些 `NO_GAIN`,RAG `REFERENCE` 由谁判定? | user-interview | 已解决 |
| Q6 | 范围 | 首版是否需要 `new_count` 或自然语言语义去重? | user-interview | 已解决 |
| Q7 | 配置 | 连续无增益阈值是否可配置,默认值与生效时机是什么? | user-interview | 已解决 |
| Q8 | 发布 | 信息饱和、预算终止和 `conclusion=null` 由谁转换为用户可见结果? | user-interview | 已解决 |
| Q9 | 验收 | 如何证明未知问题不再以通用内部错误结束,同时不放过无证据结论? | evidence-driven | 已解决 |
| Q10 | 技术 | 当前 Tool schema 是否能直接容纳 `previous_observation`? | evidence-driven | 已解决 |
| Q11 | 技术 | 进展控制状态应扩展 Redis Store 还是放入 RunContext handle? | evidence-driven | 已解决 |
| Q12 | 技术 | RAG `relevance_level` 在哪一层丢失,前端是否已有过程展示能力? | evidence-driven | 已解决 |
## Evidence-driven
| 结论 | 证据来源 | 是否已汇报用户 |
|---|---|---|
| 当前三个 Agent-facing Tool 直接使用 `RagToolRequest`、`QueryLogsRequest`、`MysqlToolRequest` 生成 Schema;要增加 `previous_observation + input` 必须显式演进 Tool Schema,不能只改 interceptor。 | `HarnessEvidenceTools`、三个 request records、`DiagnosisAgentFactory` | 已汇报 |
| `HarnessToolInterceptor` 当前把完整 `agentResult` 放入 `ToolCallResponse.content`,控制视图与模型观察尚未分离。 | `HarnessToolInterceptor` | 已汇报 |
| `RunContext` 已采用结构不可变、可变状态存在线程安全 handle 的模式;进展 tracker 放入 RunContext 比扩展 Redis 按 Run 枚举更符合现有所有权。 | `RunContext`、`DiagnosisHarnessCore.startRun` | 已汇报 |
| `CanonicalInvocationStore` 只有 begin/find/markReady/markError,扩展按 Run 枚举会影响 Redis 实现和多组 fake store;首版可由 tracker 保存完成调用 key,在结束时按 key 读取 canonical 记录。 | `CanonicalInvocationStore` 及其引用测试 | 已汇报 |
| `RagResultProjector` 只按 evidence 是否为空生成 `EVIDENCE_FOUND / NO_EVIDENCE`,没有读取上游 `relevanceLevel / relevance_level`。 | `RagResultProjector`、`LookupResult`、`KnowledgeEvidencePostProcessor` | 已汇报 |
| `DiagnosisReleaseUseCase.execute` 当前强制 draft 非空并对所有 Draft 运行 EvidenceGuard;`EvidenceGuard` 又把空 analysis 判为 `ANALYSIS_MISSING`,与合法无结论结果冲突。 | `DiagnosisReleaseUseCase`、`EvidenceGuard` | 已汇报 |
| `ChatApplicationUseCase.recoverBudgetExhaustion` 已有未提交预算 Fallback,但它绕过 Diagnosis Release,需迁移而不是丢弃用户价值。 | `ChatApplicationUseCase`、`SafeFallbackFactory`、现有测试 diff | 已汇报 |
| 前端已渲染 `observed_facts / verified_sources / limitations / next_steps`,不需要新增公开展示协议。 | `src/main/resources/static/app.js` | 已汇报 |
| 验收必须同时覆盖主动无结论、Harness 饱和、预算终止、无证据结论被 Guard 拦截,以及原始未知 Query 的 live SSE、日志和 exact run 数据。 | 当前事故现象、ISS-016 验收项、现有 E2E 工具 | 已汇报 |
## User-interview
| 问题原文 | 用户原话 | 确认状态 | OpenSpec 回写 |
|---|---|---|---|
| 是否精简状态而不建立第二套生命周期? | “这里我觉得设计得太混乱了,怎么简化”以及对最终简化架构“我觉得可以” | 已确认 | 已回写 |
| 是否只保留 `GAINED / NO_GAIN`? | “质量状态只要 information_gain = GAINED \| NO_GAIN 就够了?”后确认“我觉得可以” | 已确认 | 已回写 |
| 是否需要 `new_count`? | “好那就去掉new_count” | 已确认 | 已回写 |
| 是否需要 `next_action`? | “那就去掉next_action,我觉得由llm自己去判断就好了,不用显示的指定” | 已确认 | 已回写 |
| Tool 如何注入? | “首先 tool注入,由服务端注入,而不是写死在提示词中” | 已确认 | 已回写 |
| Prompt 是否强调合法放弃和正确但无用的内容? | “需要说明 模型不必须要给出一个答案”以及“如果你发现工具返回的是正确但对推导无用的废话,请停止调用” | 已确认 | 已回写 |
| 阈值是否可配置? | “我觉得这个可以暴露出一个配置来控制” | 已确认 | 已回写 |
| 首版重复检测做到什么程度? | “也就是说这一版只是做参数的去重校验”后确认“可以” | 已确认 | 已回写 |
| 是否按当前简化方案进入实施? | “可以用更简单的方式”“可以。修正一下文档”“用sm-flow开始实施把” | 已确认 | 已回写,并授权完成 Commit 后进入 Apply |
## 关键取舍
- 决策:停止权归 Harness,语义价值判断由模型与确定性规则共同产生。
- 原因:Tool 只能知道客观返回,模型才能判断内容是否推进当前假设;但空结果和完全重复 scope 可由代码零 Token 判定。
- 影响:Harness 只消费二值信息增益,不引入独立 Judge 或质量分数。
- 决策:使用下一次 Tool Call Envelope 回传上一轮模型评价。
- 原因:模型只有看到 Tool Observation 后才能评价,下一次真实行为正好提供受 Schema 约束的回传边界。
- 影响:这是 L3 Agent-facing Tool Schema 变更,业务 request 在 interceptor 内解包后保持不变。
- 决策:进展 tracker 是 RunContext handle,Canonical Store 保持 Tool 真相源。
- 原因:停止决策需要低延迟 Run 内状态,完整证据仍应由 canonical 记录提供;两者职责不同。
- 影响:tracker 保存计数、scope、待评价调用和 canonical keys,不复制 raw payload。
- 决策:不创建 ADR。
- 原因:这些是 ISS-016 范围内可通过 OpenSpec 回滚的内部协议演进,已有架构文档详细记录取舍,尚不满足独立 ADR 的必要性。
## OpenSpec 回写
- 必须进入 proposal/design/spec/tasks:Tool Envelope、L3 影响、Run tracker、确定性 `NO_GAIN`、RAG `relevance_level`、双视图、STOP_REQUIRED、ProgressSnapshot、统一 Release、Prompt、配置和 E2E。
- 必须保留非目标:无 `new_count`、无 `next_action`、无 Judge、无语义去重、无第二套诊断生命周期、无公开 SSE 协议新增。
- 当前没有未确认的 user-interview 问题,也没有 devflow/OpenSpec 冲突。
## Cross-Artifact 对齐检查
| 上游 → 下游 | 检查内容 | 状态 |
|---|---|---|
| ISS-016/架构文档 → proposal | 未知问题、合法放弃、信息增益、饱和停止、双视图、统一 Release、非目标和 E2E | 已对齐 |
| proposal → design | L3 Envelope、Run tracker、scope、STOP_REQUIRED、ProgressSnapshot、预算终态、Prompt 和迁移方案 | 已对齐 |
| design → specs/tasks | 状态机、Tool 门禁、RAG relevance、无结论 Guard、Release 所有权、Trace 和兼容边界 | 已对齐 |
| specs → tasks | 每个可观察行为均有 contract/state/loop/release/config/E2E 可执行切片 | 已对齐 |
### Gap 详情
- 无。
## Architecture Audit
- Capability source:`zoom-out`。按 glossary 的 Diagnosis Agent、Diagnosis Harness、RunContext、Evidence Status、Invocation Status、Release Outcome 术语审计。
- 顶层链路:`ChatApplicationUseCase -> DiagnosisChatExecutor -> DiagnosisAgentUseCase -> ReactAgent/interceptors -> ToolBoundary/canonical store -> ProgressSnapshot -> DiagnosisReleaseUseCase -> SSE/persistence`。
- 所有权:Agent 负责诊断语义;Tool/Projector 负责客观结果;Run tracker 负责停止控制;Canonical Store 负责 Tool 真相;Guard 负责引用与结论安全;Release 负责用户可见 SUCCESS/FALLBACK;Application 只负责编排和持久化。
- `RunContext` 的生产代码构造点只有 `DiagnosisHarnessCore.startRun`,大量测试通过该工厂获取;新增 tracker 不需要扩散手工构造。
- `DiagnosisAgentUseCase`、`HarnessToolInterceptor`、`HarnessEvidenceTools` 和 `DiagnosisReleaseUseCase` 的直接消费者均已由配置类和 focused tests 覆盖,任务清单包含所有构造调用更新。
- `FallbackType` 新语义只通过通用 SafeFallback JSON/前端渲染消费,没有前端枚举 switch;公开协议不新增字段。
- 最大框架风险是 Tool Envelope Schema 和 STOP_REQUIRED 后异常传播;design 要求三个具体 record、真实 callback schema 测试和 scripted framework-loop 测试在 Release 迁移前锁定行为。
- 最大生命周期风险是预算已把 RunLifecycle 置为 `BUDGET_EXHAUSTED` 后 Application 再次 `checkActive`;design 将其限制为“Diagnosis Release 已处理的预算 Fallback”窄分支,并禁止 Application 重建业务内容。
- 审计结论:模块职责没有形成新的循环依赖或第二真相源;L3 风险已进入 specs 和 tasks,可进入 commit gate。
## Commit Gate Preflight
- `proposal.md`、`design.md`、六份 capability delta specs 和 `tasks.md` 均存在。
- `openspec status --change diagnosis-information-gain-stop-contract --json` 返回 `isComplete=true`。
- `openspec validate diagnosis-information-gain-stop-contract --strict` 通过。
- question pool 全部已解决;evidence-driven 结论已汇报;user-interview 决策均有用户原话和确认状态。
- 接口影响已判为 L3,并有独立 Interface Impact、兼容、迁移、回滚和验收说明。
- Cross-artifact 检查无 gap;架构风险均已进入 design/tasks。
- 用户已通过“用sm-flow开始实施把”明确授权 Commit 后进入 Apply。
## Pre-apply Research
### 参考实现
- `HarnessEvidenceTools`:现有三类 `FunctionToolCallback` 注册点和 adapter bridge,继续作为 Agent-facing Schema 唯一入口。
- `HarnessToolInterceptor`:可获得 exact framework Tool Call ID,适合消费 Envelope 和执行 progress gate。
- `ToolBoundary`:Tool 预算、Run 校验、canonical 写入和安全错误的单点,不在 interceptor 重复 reserve。
- `RunContext` / `DiagnosisHarnessCore.startRun`:结构不可变 + 可变 handle 模式和唯一生产构造点。
- `RagResultProjector` / `QueryLogsResultProjector` / `MysqlResultProjector`:bounded canonical agent result 的现有标准化模式。
- `EvidenceGuard` / `DiagnosisReleaseUseCase`:当前结论验证链和无结论冲突位置。
- `ChatApplicationUseCase.recoverBudgetExhaustion`:保留用户价值、需要迁移所有权的临时预算 Fallback。
- `DiagnosisAgentUseCaseTest.ScriptedChatModel`:真实框架 model -> Tool -> model loop 回归模式。
### 技术栈清单
- Tool Schema:三个具体 record 交给 Spring AI `FunctionToolCallback.inputType`,共享 `PreviousObservation`,不使用泛型擦除或 JsonNode Schema。
- JSON:继续使用项目 `ObjectMapper` 严格解析/序列化;控制字段在 interceptor 消费后只传业务 input。
- Run 状态:新增线程安全 tracker handle,由 `DiagnosisHarnessCore.startRun` 创建,不使用 ThreadLocal。
- Canonical 真相:继续使用 `ToolCallKeyFactory + CanonicalInvocationStore.find`;tracker 只记录 identity。
- 视图:从 bounded canonical `agent_result` 白名单投影 Model Observation,不读取 raw response。
- 异常:受控停止使用专用异常和 cause-chain 分类;未知异常保持 fail closed。
- 测试:JUnit 5、scripted ChatModel、现有 fake store/adapter fixture;不增加 Maven 依赖。
### 新建基础设施
- `harness.progress`:信息增益、收集状态、停止原因、tracker、scope、snapshot/projector。
- `harness.agent`:三个 Agent-facing Envelope、白名单 observation projector、受控停止异常和执行结果。
- 不新增数据库表、Redis 数据结构、HTTP DTO、SSE event 或外部依赖。
## Apply 期间设计补充:Draft 合同失败
- 真实 E2E `runId=4e667111-524e-4407-87ab-b4b262952017` 已完成一次 READY RAG 调用,第二轮模型返回文本后在 Draft/Release 边界失败;后续三次同 Query 均走零 Tool 的 `MISSING_REQUIRED_CONTEXT`,证明模型输出存在随机分支。
- 用户确认采用窄化降级:非法 Draft 自身不被接受;已有当前 Run 的安全 ProgressSnapshot 时发布 `INSUFFICIENT_EVIDENCE`,没有安全过程时继续 `FAILED`。
- 这是有意行为变更:从“所有非法 Draft 都发布技术失败”调整为“非法 Draft + 已验真过程可发布过程型 Fallback”;公开 SSE 字段、Tool 协议和最终生命周期枚举不变。
- 不新增 stop reason,不把 Draft 解析失败伪装成 `INFORMATION_SATURATED` 或 `BUDGET_LIMIT_REACHED`;使用 Agent 输出异常携带有界 snapshot,并以脱敏 Trace 区分输出合同失败。
## Apply Verification
- Focused tests:`DiagnosisAgentUseCaseTest`、`DiagnosisReleaseUseCaseTest`、`DiagnosisChatExecutorTest`、`HarnessChatConfigurationTest` 通过。
- 完整回归:`mvn -q -Dtest='!MilvusConnectionTest' test` 退出码为 `0`;本轮 Surefire 报告汇总 `Tests=292, Failures=0, Errors=0, Skipped=3`。
- 外部凭据边界:未排除时唯一失败为 `MilvusConnectionTest.connect`,原因是当前测试进程未设置 `MILVUS_TOKEN`;这不是本变更回归。
- OpenSpec:`openspec.cmd validate diagnosis-information-gain-stop-contract --strict` 通过。
- 格式与清理:`git diff --check` 通过;未发现临时 E2E JSON、DEBUG 或 tmp 文件。
- named SSE E2E:Query `诊断切换企业失败的问题`,`sessionId=iss016-final-20260726-a`,`runId=3ab22ed7-d0ed-45d8-b928-dce5790c0542`;SSE 返回 `SAFE_FALLBACK`、`type=MISSING_REQUIRED_CONTEXT`、`done.outcome=FALLBACK`。
- 数据库核对:`status=SUCCESS`、`intent=DIAGNOSIS`、`release_outcome=FALLBACK`、`tool_call_count=0`、`total_token_count=2890`、answer 非空。
- Trace 核对:`RUN_STARTED -> ROUTING_ATTEMPT -> ROUTING_DECISION -> AGENT_MODEL_STEP -> EVIDENCE_GUARD_INITIAL -> RELEASE_DECISION/FALLBACK -> RUN_FINISHED/FALLBACK`。
- 兼容性:公开 HTTP/SSE 字段、前端 SafeFallback 消费结构、数据库表和业务 Tool request 均未新增字段;模型侧 Tool Envelope 是本变更已确认的 L3 协议变更。
## Apply Continuation: Task 8 Protocol Repair + Bounded Stop
- Checkpoint:Apply。
- Capability source:`openspec-apply-change` + sm-flow apply 协议。
- 背景:tasks 1–7 已完成;真实 E2E 暴露连续 `INVALID_PROGRESS_PROTOCOL` 不会累计 `NO_GAIN`,可能在硬预算前空转。Task 8 补齐协议修复反馈与独立兜底停止。
- 实现事实(代码已在工作区,本轮补齐测试与收口):
- `ProgressProtocolViolationType` / `ProgressProtocolViolationException` 覆盖 MISSING_PREVIOUS_OBSERVATION、OUT_OF_ORDER、UNEXPECTED、MISSING_INPUT、INVALID_ENVELOPE。
- `DiagnosisProgressTracker` 独立累计连续协议错误,默认阈值 2,达到后 `stop_reason=PROGRESS_PROTOCOL_VIOLATED`。
- `HarnessToolInterceptor` 返回可修正 observation(repair_required、violation_type、missing_field、expected_previous_tool_call_id、allowed_information_gain);达阈一次 STOP_REQUIRED,再请求抛 `DiagnosisCollectionStoppedException`。
- `DiagnosisReleaseUseCase` 支持 `PROGRESS_PROTOCOL_VIOLATED`:有安全 ProgressSnapshot 发 `INSUFFICIENT_EVIDENCE`,无进展 fail closed。
- `TOOL_REQUEST_REJECTED` 记录 violation_type、repair_prompt_delivered、consecutive_protocol_violations、stop_reason,不记录参数/观察正文/异常。
- 验证:
- Focused:`DiagnosisProgressTrackerTest`、`HarnessToolInterceptorTest`、`DiagnosisReleaseUseCaseTest`、`DiagnosisAgentUseCaseTest`、`HarnessChatConfigurationTest` 通过。
- OpenSpec strict validate 通过。
- 文档:ISS-016 剩余协议停止项勾选完成;ISS-015 阶段 1 标记已完成;架构文档同步协议错误独立停止语义。
- OpenSpec tasks 8.1–8.6 全部完成。剩余 Apply 工作:无。可进入 Archive checkpoint(需用户确认是否归档 OpenSpec)。
## Archive
- Checkpoint:Archive。
- Capability source:`sm-flow` archive 协议 + `openspec-archive-change`。
- 用户确认:明确要求“执行(archive),完后提交推送”。
- devflow 档案:
- `brief.md`、`evidence.md`、`decisions.md`、`acceptance.md`
- 更新 `devflow/index.md`、`devflow/glossary/CONTEXT.md`
- OpenSpec:
- delta specs 已同步到 main specs(含新建 `diagnosis-information-gain-stop-contract`)
- change 归档至 `openspec/changes/archive/2026-07-27-diagnosis-information-gain-stop-contract/`
- 不创建独立 ADR:决策已由 OpenSpec/ISS/架构文档承载,且可通过 OpenSpec 回滚。
- 状态:archived。
@@ -0,0 +1,38 @@
# Diagnosis 信息增益停止契约 Evidence
## 证据
| 来源 | 证据 | 结论 | 是否已汇报 |
| --- | --- | --- | --- |
| `HarnessEvidenceTools` / request records | Agent-facing Tool 直接使用业务 request 生成 Schema | 要增加 `previous_observation + input` 必须显式演进 Tool Schema | 是 |
| `HarnessToolInterceptor` | 曾把完整 `agentResult` 放入 Tool Response | 必须拆分控制视图与模型观察 | 是 |
| `RunContext` / `DiagnosisHarnessCore.startRun` | 结构不可变 + 线程安全 handle | 进展 tracker 放 RunContext,不扩 Redis 枚举 | 是 |
| `CanonicalInvocationStore` | 仅 begin/find/markReady/markError | tracker 只存 identity,结束时投影 | 是 |
| `RagResultProjector` | 只按 evidence 空否生成状态 | 需兼容 `relevanceLevel/relevance_level` | 是 |
| `DiagnosisReleaseUseCase` / `EvidenceGuard` | 强制 draft 非空且空 analysis=`ANALYSIS_MISSING` | 与合法无结论冲突,需统一 Release | 是 |
| `ChatApplicationUseCase.recoverBudgetExhaustion` | 未提交预算 Fallback 绕过 Release | 迁移意图到 Diagnosis Release | 是 |
| 前端 `app.js` | 已渲染 observed_facts/sources/limitations/next_steps | 不新增公开 SSE 字段 | 是 |
| 真实 E2E(实施前) | 多轮空转后 `BUDGET_EXHAUSTED`/`INTERNAL_FAILURE` | 需要信息增益停止契约 | 是 |
| 真实 E2E(实施后) | `iss016-final-20260726-a` → `MISSING_REQUIRED_CONTEXT` FALLBACK | 未知 Query 可正常业务结束 | 是 |
| Token/拒绝审计 E2E | 9 次 `INVALID_PROGRESS_PROTOCOL` 拒绝不累计 NO_GAIN | 需独立协议错误阈值与 STOP | 是 |
## Evidence-driven 结论
- 结论:Tool Envelope 是 L3 模型侧协议变更,业务 request 在 interceptor 解包后保持不变。
- 证据:三个 FunctionToolCallback inputType、adapter bridge 只收业务 JSON。
- 风险:框架 Schema/拦截器假设不匹配。
- 用户确认:不需要(技术事实)
- 结论:协议错误不得累计为 `NO_GAIN`,必须独立 `PROGRESS_PROTOCOL_VIOLATED`。
- 证据:真实审计 9 次协议拒绝 + 13 Agent 轮次;OpenSpec design 6.1。
- 风险:只返回通用错误码不足以自修复。
- 用户确认:已通过 Task 8 OpenSpec 与实现收口
- 结论:Release 是业务 Fallback 唯一决策入口;Application 不重建业务内容。
- 证据:`DiagnosisReleaseUseCase` 统一路径 + Application 窄化预算终态持久化。
- 用户确认:已确认
## 实现期补充证据
- Draft 合同失败窄化降级:非法 Draft 丢弃;仅当 ProgressSnapshot 有已验真 facts 时发 `INSUFFICIENT_EVIDENCE`。
- Task 8 focused tests:tracker / interceptor / release / agent-loop / config 全部通过。
@@ -0,0 +1,22 @@
# Acceptance
## Done
- `MilvusHybridKnowledgeStore`: schema BM25 function + dense, hybridSearch+RRFRanker, dense search
- `VectorSearchService` only routes dense|hybrid to V2 store
- `VectorIndexService` writes via V2 store
- Removed knowledge-path `MilvusServiceClient` bean wiring
- Config: `milvus.collection=biz_hybrid`, `retrieval.search.mode=hybrid`
## Tests
```text
mvn -Dtest=LookupKnowledgeToolTest,KnowledgeEvidencePostProcessorTest,RagResultProjectorTest,RrfFusionTest,VectorKnowledgeSearchAdapterHybridTest,VectorSearchServiceTest,VectorIndexServiceTest test
```
EXIT:0
## Ops note
Reindex all knowledge docs into `biz_hybrid` before production hybrid search is meaningful.
Legacy `biz` collection is unused by knowledge path.
@@ -0,0 +1,3 @@
# Brief: rag-bm25-hybrid-drop-sdk
True dense+BM25 hybrid on a single MilvusClientV2 backend. Legacy SDK search/write for knowledge path removed. New collection `biz_hybrid` requires reindex.
@@ -0,0 +1,7 @@
# Decisions
- Single backend: MilvusClientV2 only for knowledge RAG
- Drop sdk/spring/auto retrieval routing
- New collection biz_hybrid to avoid mutating legacy biz schema in place
- Hybrid = dense ANN + BM25 sparse ANN + RRFRanker
- Dense L2 enrichment for threshold compatibility on hybrid hits
@@ -0,0 +1,44 @@
# Acceptance: rag-chunk-evidence-identity-dedup
## 实现结果
Delivery 1 completed:
- chunk identity fields on candidates/evidence blocks
- `KnowledgeSearchPort` + dense adapter
- evidenceKey dedup + maxChunksPerDocument + return-n
- retrieve-k on lookup tool
- projector keeps same-source distinct chunks; `document_id` is chunk-scoped
## 验证
### 脚本验证
```text
mvn -q "-Dtest=LookupKnowledgeToolTest,KnowledgeEvidencePostProcessorTest,RagResultProjectorTest" test
```
结果:通过(exit 0)
### 静态验证
- Search/port wiring reviewed against OpenSpec tasks
- No new legacy SDK dependency added
### 浏览器/人工验证
未运行(纯检索契约变更,无 UI)
### 未验证
- 全量 harness E2E / 真实 Milvus 联调(Delivery 2 前可补)
- 生产配置默认 retrieve-k/return-n 调优
## 归档状态
- OpenSpec change ready to archive
- User pre-authorized archive for sm-flow staged delivery
## 后续
- Delivery 2: milvus hybrid search (separate change)
@@ -0,0 +1,29 @@
# Brief: rag-chunk-evidence-identity-dedup
## 背景
lookup_knowledge 同文档多 chunk 在后处理与投影阶段被 source 级去重吞掉。Hybrid 多路召回前必须先修证据身份与裁剪契约。
## 目标
- chunk 级 evidenceKey 身份
- 按 evidenceKey 去重 + 每文档 chunk 上限
- retrieve-k / return-n 分离
- Projector 保留同 source 不同 chunk
- 薄 KnowledgeSearchPort,为 Delivery 2 hybrid 铺路
## 范围
Delivery 1 only(见 `docs/Milvus-Hybrid接入清单.md` §1.1)。
## 非目标
hybrid schema、BM25、删 SDK、session dedup、邻块重建、模型 rerank。
## 分档
standard
## 关联 OpenSpec
`openspec/changes/rag-chunk-evidence-identity-dedup`
@@ -0,0 +1,102 @@
# Decisions: rag-chunk-evidence-identity-dedup
## Capability sources
- sm-flow orchestration
- OpenSpec fallback protocol (file-based propose/apply/archive) — external openspec-propose/apply skills used as reference; execution via sm-flow fallback
- grill: fallback built-in protocol
- audit: fallback built-in protocol
## Scale
standard
## Clarify
- Problem: same-document multi-chunk evidence collapsed by source-level dedup.
- Outcome: Delivery 1 foundation before hybrid Delivery 2.
- Slug: `rag-chunk-evidence-identity-dedup`
- User authorized apply + archive in advance for sm-flow staged changes.
## Context
- Read: `devflow/glossary/CONTEXT.md`, modular-rag-pipeline brief, `openspec/specs/rag-knowledge-retrieval`, `rag-log-projections`, checklist doc §1.1
- Constraints into OpenSpec:
- L0 hint-only remains
- Do not thicken legacy SDK path
- Agent tool name/input stable
- Hybrid out of scope this change
## Question pool (grill)
| # | Dimension | Mode | Question | Status |
|---|---|---|---|---|
| Q1 | 术语 | evidence-driven | evidenceKey / document_id 语义? | Resolved: evidenceKey=chunk id; projected document_id=evidenceKey |
| Q2 | 边界 | evidence-driven | Delivery 1 vs 2 边界? | Resolved: per checklist; no schema/hybrid/SDK delete |
| Q3 | 验收 | evidence-driven | 如何验收多 chunk? | Resolved: unit tests multi-chunk keep + projector |
| Q4 | 接口 | user-interview | document_id 改为 chunk 级是否可接受? | **Pre-authorized by user** via “apply/archive 直接授权” + prior design agreement on scheme A (document_id=evidenceKey). Recorded as accepted behavior change. |
| Q5 | 技术 | evidence-driven | SearchPort 是否本 change 必须? | Resolved: thin port required as foundation |
### Evidence-driven conclusions (reported)
1. Current collapse points: `KnowledgeEvidencePostProcessor.sourceKey` and `RagResultProjector` source fallback.
2. Metadata already has docId/chunkIndex on write path; not first-class on read path.
3. Existing main-spec still says source-level dedup — this change intentionally deltas that requirement.
### User-interview
- Q4 accepted under prior design alignment (scheme A) and explicit apply authorization for this sm-flow run. No remaining open product preference questions for Delivery 1.
## Audit
Module chain:
```text
LookupKnowledgeTool -> SearchPort -> Retriever -> PostProcessor -> Packer -> Assembler -> RagResultProjector
```
Risks:
1. Agent payload growth — mitigated by return-n + maxChunksPerDocument + projector budgets.
2. document_id semantic shift — documented L3 behavior change; tests updated.
3. Old data without chunkIndex — vector id fallback.
No ADR conflict with modular RAG L0/L1 boundary.
## Cross-artifact alignment
| From | To | Status |
|---|---|---|
| brief goals | proposal | 已对齐 |
| proposal scope | design decisions | 已对齐 |
| design identity/dedup/port | specs | 已对齐 |
| specs scenarios | tasks | 已对齐 |
## Interface impact
- L2 internal DTO
- L3 Agent `document_id` chunk-scoped
## Commit gate
- proposal/design/specs/tasks present
- no open user-interview blockers for Delivery 1
- apply authorized by user at sm-flow start
## Pre-apply research
Reference files:
- `LookupKnowledgeTool.java`
- `KnowledgeDocumentRetriever.java`
- `KnowledgeEvidencePostProcessor.java`
- `RagResultProjector.java`
- `LookupKnowledgeToolTest.java`
- `RagResultProjectorTest.java`
- `docs/Milvus-Hybrid接入清单.md`
Stack notes:
- No MQ/request envelope changes
- Spring `@Value` config pattern for rag.* keys
- Tests use ReflectionTestUtils + Mockito
@@ -0,0 +1,16 @@
# Evidence: rag-chunk-evidence-identity-dedup
## 代码证据(变更前)
- `KnowledgeEvidencePostProcessor.sourceKey` 使用 source/title 去重
- `RagResultProjector` 用 source 回退 document_id 并 HashSet 去重
- 写入路径 metadata 已有 docId/chunkIndex,读路径未一等化
## 规格证据
- 旧 `openspec/specs/rag-knowledge-retrieval` 要求 source 级 dedup(本 change 以 delta 修正)
- `docs/Milvus-Hybrid接入清单.md` §1.1 定义 Delivery 1 地基
## 验证证据
- 单测覆盖 multi-chunk keep / true-dup merge / maxChunksPerDocument / projector same-source multi-chunk
@@ -0,0 +1,23 @@
# Acceptance: rag-hybrid-search-rrf
## Result
Hybrid mode implemented on KnowledgeSearchPort:
- dense unfiltered + dense filtered + lexical rank over union
- RRF fusion by evidenceKey
- dense-compatible score preserved for thresholds
- default mode remains dense
## Verification
```text
mvn -q "-Dtest=RrfFusionTest,VectorKnowledgeSearchAdapterHybridTest,LookupKnowledgeToolTest" test
```
Pass.
## Residual
- True Milvus BM25/sparse schema + reindex still follow-up
- Lexical path only ranks dense-recalled candidates (does not expand pure-term misses outside dense topK)
@@ -0,0 +1,3 @@
# Brief: rag-hybrid-search-rrf
Delivery 2 after chunk identity. Enable hybrid multi-path + RRF on KnowledgeSearchPort without legacy SDK hybrid API. True BM25 schema rebuild is staged follow-up; this change ships sparse-lite lexical ranking over dense candidate union + filtered/unfiltered dense fusion.
@@ -0,0 +1,19 @@
# Decisions: rag-hybrid-search-rrf
## Capability
sm-flow + OpenSpec fallback; apply pre-authorized.
## Depends
Delivery 1 archived.
## Grill (compressed, pre-authorized)
- Q: Full BM25 schema now? A: No — sparse-lite + RRF first; schema rebuild follow-up.
- Q: Default mode? A: dense default; hybrid opt-in.
- Q: Threshold score? A: keep dense-compatible L2 mapping.
## Design
Hybrid paths: dense unfiltered + dense filtered + lexical rank over union; RRF fuse by evidenceKey.
@@ -0,0 +1,5 @@
# Evidence
- Delivery 1 identity/port foundation required
- RRF utility and hybrid adapter unit tests green
- Lexical sparse-lite intentionally intermediate until BM25 schema
@@ -0,0 +1,39 @@
# Acceptance: rag-eval-hybrid-baseline
## Tasks
All tasks in OpenSpec `tasks.md` checked, including apply-discovered 6.x quality-gate fix.
## 静态验证
- Snapshot generator path: no required `retrieval.vector-store.mode`.
- README documents hybrid generation and offline/live split.
## 脚本验证
```text
.\scripts\prepare_rag_eval_seed.ps1
.\scripts\generate_rag_lookup_snapshots.ps1 -SearchMode hybrid -SkipEval
python scripts\eval_rag_retrieval.py --json-report eval/rag-retrieval/reports/baseline.json --markdown-report eval/rag-retrieval/reports/baseline.md
# Result: Evaluated 7 cases: passRate=1.0, recall@5=1.0, failed=0
mvn -Dtest=RetrievalScoreNormalizerTest,KnowledgeEvidencePostProcessorTest,LookupKnowledgeToolTest,VectorSearchServiceTest,VectorKnowledgeSearchAdapterHybridTest test
# exit 0
```
Fixture sample meta: `searchMode=hybrid`, `kbScope=rag-eval`.
Fallback case: `UNFILTERED_VECTOR_RETRY` + `filtered_vector_low_quality`.
## 浏览器/人工
- 未做 UI 验证。
## 未验证 / 后续
- Dense vs hybrid dual-directory comparison report (knife-2).
- CI wiring of offline eval as required gate (optional process).
- Long-term calibration of hybrid PRECISE distribution under denseDistance quality.
## Specs
Main spec synced: `openspec/specs/rag-eval-offline-baseline/spec.md`.
@@ -0,0 +1,16 @@
# Brief: rag-eval-hybrid-baseline
## Background
Offline RAG eval (golden × fixture × key-field baseline) existed but generator/docs still used dead `retrieval.vector-store.mode=spring`. Fixtures lacked search meta and did not reflect hybrid main path.
## Goals (knife-1 only)
- Snapshot generation uses `retrieval.search.mode` (default hybrid; dense override).
- Fixtures record `searchMode` / `kbScope`.
- README documents hybrid-era offline vs live loop.
- Best-effort live seed + regenerate fixtures + update baseline.
## Non-goals
Dense/hybrid dual fixture trees; golden mustNot/chunk/level hard gates; new eval frameworks.
@@ -0,0 +1,18 @@
# Decisions: rag-eval-hybrid-baseline(最终版)
## Process
sm-flow standard-lean: Discover → Commit → Apply → Archive.
## Key decisions
1. Replace eval generator `vector-store.mode` with `retrieval.search.mode` (default hybrid).
2. Fixture meta: `searchMode`, `kbScope` when set.
3. Knife-2 (dual fixtures / mustNot golden) deferred.
4. Live refresh succeeded in apply env; baseline updated to hybrid snapshots.
5. **Quality gate refinement (apply-found):** hybrid absolute quality for `isLowQuality` / relevance uses optional dense L2 (`denseDistance`); does not overwrite hybrid scoreLabel or RRF order. Rank mapping remains fallback when dense missing.
## Trade-offs
- Extra dense ANN on hybrid path for gate calibration (latency) vs correct filter-fallback behavior.
- relevance_level still not a hard golden assertion (ordinal vs absolute mix).
@@ -0,0 +1,17 @@
# Evidence: rag-eval-hybrid-baseline
## Pre-change
- `generate_rag_lookup_snapshots.ps1` passed `-Dretrieval.vector-store.mode=spring`.
- Fixtures had `caseId/query/retrievedAt/lookupResult` only.
- Offline eval already supported Hit levels, recall@K, baseline diff.
## User decisions
- Scope: knife-1 only (no dual fixture dirs).
- Acceptance: wiring required; fixture refresh best-effort (env allowed full refresh).
## Apply-discovered
- After hybrid refresh, `chat-l0-filter-fallback` failed: pure rank→quality made topSimilarity=1.0 on decoy-only filtered hits → no unfiltered retry.
- Fix: optional `denseDistance` on hybrid hits; quality gate uses L2 when present; sort order remains RRF.
@@ -0,0 +1,41 @@
# Acceptance: rag-quality-score-unify
## Tasks
OpenSpec `tasks.md` 全部 `[x]`(1.1–6.2)。
## 静态验证
- 生产路径 grep:无 `bm25_only_no_dense` 发射、无 hybrid L2 enrichment(仅 Labels canonicalize 兼容旧串)。
- 架构文档 §6 与 `application.yml` 注释已对齐 quality 契约。
## 脚本验证
```text
mvn -Dtest=RetrievalScoreNormalizerTest,KnowledgeEvidencePostProcessorTest,LookupKnowledgeToolTest,VectorSearchServiceTest,VectorKnowledgeSearchAdapterHybridTest test
```
| 套件 | 结果 |
|---|---|
| RetrievalScoreNormalizerTest | 4 passed |
| KnowledgeEvidencePostProcessorTest | 6 passed |
| LookupKnowledgeToolTest | 7 passed |
| VectorSearchServiceTest | 2 passed |
| VectorKnowledgeSearchAdapterHybridTest | 1 passed |
(PowerShell 可能将 JVM warning 标为 exit 1;日志中为 BUILD SUCCESS / Failures: 0。)
## 浏览器 / 人工验证
- 未跑:live `lookup_knowledge` hybrid vs dense 对照、生产阈值标定。
## 未验证
| 项 | 风险 | 建议 |
|---|---|---|
| 真实 Milvus hybrid 联调 | 序/质量分布与单测 mock 有差 | 启动服务后固定 query 集切 mode 对比 |
| 阈值 0.75/0.5 在 hybrid rank 分下的标定 | retry/PRECISE 偏多或偏少 | 看 trace topSimilarity 再调 yml |
## Specs 同步
- 主规格新增:`openspec/specs/rag-retrieval-quality-score/spec.md`(archive 时从 delta 同步)。
@@ -0,0 +1,21 @@
# Brief: rag-quality-score-unify
## Background
真 BM25 hybrid(dense + BM25 + RRF)已上线,但后处理仍把 hybrid 结果伪装成 L2 做 `normalizeL2`,并用 L0 domain/entity/keyword contains 加分改序。排序权威与质量闸门分裂,词面信号被 BM25 与后处理双重计分。
## Goals
- 一级 `scoreLabel` 仅 `dense` | `hybrid`
- 唯一 `toQualityScore`;后处理 label-agnostic
- 排序主序 = 检索 `originalRank`;去掉关键词 boost 改序
- hybrid quality = 本轮 rank 纯映射(不做 max(rank, denseSim)、不为闸门回填 L2)
- 保留 `mode=dense` 作同库召回对照;线上默认 hybrid
## Scope
内部 RAG:store 发射、normalizer、evidence post-process、单测、架构文档 §6。
## Non-goals
精排 / query rewrite / 邻块、schema rebuild、改 Agent ACI 字段名、删除 dense 对照 mode。
@@ -0,0 +1,25 @@
# Decisions: rag-quality-score-unify(最终版)
## Scale / process
- sm-flow standard:Discover → Commit → Apply → Archive
- Committed OpenSpec:`openspec/changes/rag-quality-score-unify/`(归档后见 archive 目录)
- 废止:`rag-bm25-hybrid-drop-sdk` 中「dense L2 enrichment for threshold compatibility」
## Key decisions
1. **Label**:仅 `dense` | `hybrid`;旧别名 canonicalize。
2. **Normalizer**:唯一 `toQualityScore`;dense=L2 公式;hybrid=rank 线性映射(batchSize)。
3. **Store**:hybrid 不回填 L2、不发 `bm25_only_*`;返回序即 RRF 序。
4. **Post-process**:`originalRank` ASC;L0 重叠只写 hitReasons;relevance/low-quality 只看 qualityScore;PRECISE 不要求 hint support。
5. **Mode**:hybrid 主路径;dense 同库对照(架构 §6.0)。
## Trade-offs
- hybrid quality 为序数分,跨 query 绝对值不可比;阈值可能需后续标定。
- 去掉 boost 改序后,「词面热语义冷」不再被后处理抬升;词面交给 BM25+RRF。
## Risks accepted
- `relevance_level` / unfiltered retry 分布变化(产品已接受)。
- 未做 live E2E / 人工 hybrid 对照评测(见 acceptance 未验证项)。
@@ -0,0 +1,25 @@
# Evidence: rag-quality-score-unify
## Code (pre-change)
- `MilvusHybridKnowledgeStore.searchHybrid`:RRF 后并行 dense 回填 L2;BM25-only → `bm25_only_no_dense` + maxL2。
- `KnowledgeEvidencePostProcessor`:一律 `normalizeL2(score)` + domain/entity/keyword/source_type 加分,按 `finalScore` 降序;PRECISE 需 `hasHintSupport`。
## User decisions (grill)
| ID | 结论 |
|---|---|
| Q1 | 一级 label 仅 dense/hybrid;bm25_only 不作正式 label |
| Q2 | 后处理去掉 contains 加分改序,保 originalRank |
| Q3 | 唯一 toQualityScore;后处理统一 |
| Q4 | dense mode 保留作对照 |
| Q7 | 接受 relevance_level / retry 分布变化 |
| Q8 | hybrid quality = **纯 rank 映射** |
## Post-change anchors
- `RetrievalScoreLabels` / `RetrievalScoreNormalizer`
- `MilvusHybridKnowledgeStore`(无 L2 overwrite / 无 bm25_only 发射)
- `KnowledgeEvidencePostProcessor`(rank sort + explain-only L0 overlap)
- OpenSpec delta:`rag-retrieval-quality-score`
- 架构:`mvp/architecture/RAG知识检索架构.md` §6
+12
View File
@@ -53,6 +53,18 @@ docs/
1. [分析笔记目录](analysis/) - 代码分析和问题分析 1. [分析笔记目录](analysis/) - 代码分析和问题分析
2. [临时报告目录](reports/) - 修复和验证报告 2. [临时报告目录](reports/) - 修复和验证报告
### RAG 设计讨论(docs 根目录)
1. [RAG 排序:多路召回与 RRF](RAG排序-多路召回与RRF.md) - K、融合、L0 边界
2. [Hybrid 之后的 qualityScore 与后处理](RAG-Hybrid质量分与后处理.md) - 上一代问题、L2 伪装、统一归一化
3. [Agent 如何读 relevance_level](RAG-Agent如何读relevance_level.md) - 粗相关度标签的含义与误读
4. [RAG 离线评测:讨论、设计与落地](RAG离线评测-基线设计.md) - Golden/Fixture、流程图、hybrid 对齐与闸门修复
5. [Milvus hybrid 接入清单](Milvus-Hybrid接入清单.md)
6. [RAG Trace / 审计(架构)](../mvp/architecture/RAG检索可观测性与审计.md) - 请求内 trace、tool_invocation、Trace API
### 诊断全流程(E2E 导读)
1. [一次诊断到底发生了什么](一次诊断全流程-E2E导读.md) - SUCCESS 全流程:阶段拆解、token、timeline、字段词典
2. [RAG 审计补丁 E2E:step_id + query](RAG审计补丁-stepid-query-E2E验收.md) - 审计字段 live 验收(含业务 FALLBACK 样本)
--- ---
## 📚 学习笔记 (learning/) ## 📚 学习笔记 (learning/)
+931
View File
@@ -0,0 +1,931 @@
# Milvus Hybrid Search 接入对照清单
**日期**:2026-07-27
**前提**:旧 Milvus SDK 直连检索路径后续废弃,不作为长期实现基础
**目标**:在现有 `lookup_knowledge` pipeline 上接入 dense + sparse/BM25 混合检索,融合优先走服务端 RRF
**关联文档**:
- `docs/RAG排序-多路召回与RRF.md`(排序与多路召回判断框架)
- 本文后续实现讨论以本节 **「交付拆分:分块去重 + Hybrid 同规划」** 为基线
---
## 1. 结论先说
可以接,而且和前面讨论的多路召回 / RRF 高度一致。
但(写作当时)项目 **还不具备 hybrid 运行条件**,缺的不是“再调一次 search”,而是:
```text
1. schema 只有 dense,没有 sparse/BM25 字段
2. 写入只产 dense embedding
3. 检索抽象仍以单路 similaritySearch 为中心
4. 后处理仍承担了过多“伪融合”职责
5. 证据去重粒度偏文档/source 级,同文档多 chunk 会被吞掉
```
**接入原则:**
```text
- 不继续加厚旧 SDK search 分支
- 以“检索端口 + 写入端口”抽象为准
- hybrid 融合尽量下沉到向量库(RRFRanker)
- 应用层保留:filter 策略、chunk 去重、return-n、阈值、投影
- 分块去重与 hybrid 同规划、分里程碑交付(先共用地基,再开 hybrid)
```
```mermaid
flowchart TB
subgraph gaps["写作时的缺口"]
G1[无 sparse/BM25 schema]
G2[写入只有 dense]
G3[单路 similaritySearch]
G4[后处理伪融合]
G5[source 级去重吞 chunk]
end
subgraph principles["接入原则"]
P1[端口抽象 · 不堆旧 SDK]
P2[融合下沉向量库 RRF]
P3[应用层:filter/dedup/return-n/投影]
P4[先地基后 hybrid 分里程碑]
end
gaps --> principles
```
## 实现状态(2026-07-27,后续已完成)
| 里程碑 | 状态 | 说明 |
|---|---|---|
| 交付 1 chunk 身份/去重/SearchPort | **已完成并归档** | `2026-07-27-rag-chunk-evidence-identity-dedup` |
| 交付 2a 应用层 multi-path+RRF | **已完成并归档** | `2026-07-27-rag-hybrid-search-rrf`(已被 2b 取代为生产路径) |
| 交付 2b 真 BM25 hybrid + 废弃 SDK | **已完成并归档** | `2026-07-27-rag-bm25-hybrid-drop-sdk` |
```mermaid
flowchart LR
D1[交付1<br/>chunk 身份/去重] --> D2a[交付2a<br/>app RRF]
D2a --> D2b[交付2b<br/>真 BM25 hybrid]
D2b --> NOW[生产:V2 store + hybrid mode]
```
**当前生产知识路径:**
```text
VectorIndexService / VectorSearchService
-> MilvusHybridKnowledgeStore (MilvusClientV2 only)
collection: milvus.collection (default biz)
mode: retrieval.search.mode = dense | hybrid
hybrid: dense ANN + BM25 sparse ANN + RRFRanker
```
```mermaid
flowchart TB
subgraph write["写入"]
UP[upload / init / rebuild] --> VIS[VectorIndexService]
VIS --> STORE[MilvusHybridKnowledgeStore]
end
subgraph read["检索"]
LK[lookup_knowledge] --> VSS[VectorSearchService]
VSS -->|dense| SD[searchDense]
VSS -->|hybrid| SH[searchHybrid + RRF]
SD --> STORE
SH --> STORE
end
STORE --> COL[(Milvus collection biz<br/>dense + BM25 schema)]
```
**运维必做:** 全量重灌知识库到 hybrid schema collection(配置名以 `milvus.collection` 为准,常见 `biz`);旧纯 dense collection 不能直接当 hybrid 用。
---
## 1.1 交付拆分:分块去重 + Hybrid 同规划
> 实现讨论基线。后续排期、拆 PR、评审范围,默认按本节两个交付理解。
### 判断
**可以一起做,而且应该绑在同一条改造主线上**;
但不要理解成「一个 PR 把 hybrid 全做完」。
更准确的表述:
```text
同一条演进线,两层交付:
交付 1:共用地基(证据身份 + chunk 去重 + 检索裁剪 + 端口雏形)
交付 2:hybrid(schema/写入/查询/RRF + 阈值校准)
```
### 为什么必须同规划
两边改的是同一条链上的相邻环节:
```text
检索命中
-> 候选身份(docId / chunkIndex / evidenceKey) ← hybrid 要,去重也要
-> 去重 / 单文档 chunk 上限 ← 分块去重
-> 排序融合(现在规则 / 以后库内 RRF) ← hybrid
-> return-n / Agent 投影
```
```mermaid
flowchart TB
HIT[检索命中] --> ID[候选身份<br/>docId / chunkIndex / evidenceKey]
ID --> DEDUP[chunk 去重 · 每文档上限]
DEDUP --> FUSE[排序融合 · 库内 RRF]
FUSE --> RET[return-n]
RET --> PROJ[Agent 投影]
ID -.->|交付1 地基| D1[chunk identity]
DEDUP -.-> D1
FUSE -.->|交付2| D2[hybrid]
```
若拆开且顺序错误:
| 只做一项 | 后果 |
|---|---|
| 只做 hybrid,不做 chunk 去重 | 多路召回更多同文档片段,仍被 source 级去重吞掉,**hybrid 收益被吃掉** |
| 只做去重,完全不管候选/端口结构 | 能立刻改善,但接 hybrid 时往往还要再改一遍 DTO 与映射 |
因此:
> **分块去重不是 hybrid 的可选项,而是 hybrid 生效的前提。**
> 设计上当一件事;代码上分两个可独立验证的里程碑。
### 必须放进同一批(交付 1 公共地基)
这些强烈建议同一波完成,作为后续实现讨论的最小必选范围:
| 项 | 原因 |
|---|---|
| 候选补 `docId` / `chunkIndex` / `evidenceKey` | 去重 key 与 hybrid hit 身份统一 |
| 后处理按 `docId#chunkIndex`(或 vector id fallback)去重 | 修复「同文档多 chunk 被吞」 |
| `maxChunksPerDocument` | 放开多 chunk 后防止单文档刷屏 |
| `retrieve-k` / `return-n` 分离 | hybrid 扩召回时必需;现在 K=3 也不该三者混用 |
| Projector 去重语义对齐 | 后处理放出的多 chunk,不能在投影阶段再按 `source` 砍成 1 条 |
| `SearchHit` / `RetrievedEvidenceCandidate` 字段对齐 | 避免 hybrid 再引入第三套结果结构 |
| (建议)`KnowledgeSearchPort` 雏形 | 检索调用面先稳定,后续只换实现 |
可称为:
```text
「检索结果身份与裁剪契约」
```
**不上 hybrid 也有独立价值**,并且为交付 2 铺路。
### 不要硬塞进交付 1 的同一 PR
可同规划、建议第二波(交付 2):
| 项 | 原因 |
|---|---|
| 新 collection + sparse/BM25 schema | 数据迁移/重灌,风险独立 |
| 全量重索引 | 耗时长,需单独验证 |
| 打开 `search.mode=hybrid` | 依赖 sparse 数据已就绪 |
| fused score 阈值重标定 | 要 hybrid 跑起来后有样本 |
| 删除旧 SDK 读路径 | 最后做,降低回滚成本 |
否则单个交付会同时碰:业务排序逻辑 + 数据迁移 + 基础设施,评审、回滚、评测都困难。
### 交付 1:chunk 级证据身份 + 去重 + 检索裁剪
**主题:** 让同一次检索内,同文档多个相关 chunk 能作为独立证据存活,并为 hybrid 统一 hit 模型。
**范围(实现讨论默认包含):**
```text
1. RetrievedEvidenceCandidate / EvidenceBlock
- 补 docId、chunkIndex、evidenceKey
2. KnowledgeDocumentRetriever
- 从 metadata 抽取 docId/chunkIndex
- evidenceKey 规则:
docId + "#chunk-" + chunkIndex
fallback: "vector:" + id
fallback: "rank:" + originalRank
3. KnowledgeEvidencePostProcessor
- 去重 key = evidenceKey(不再 source/title 优先)
- maxChunksPerDocument(建议默认 2)
- 同 key 才 merge;merge 不覆盖更高分 content
4. RagResultProjector
- 按 evidence 身份去重(chunk 级 document_id 或显式 chunk 身份)
- 禁止再仅用 source 当“每文档一条”的唯一键
5. 配置
- rag.retrieve-k
- rag.return-n
- rag.max-chunks-per-document
- 逐步弱化/废弃单一 rag.top-k 身兼多职
6. (建议同批)KnowledgeSearchPort / SearchRequest / SearchHit 雏形
- 即使底层暂时仍是 dense-only,调用面先稳定
7. 单测
- 同 doc 两 chunk 都保留
- 同 doc+chunk 真重复只留一条
- 超 maxChunksPerDocument 裁掉低分
- projector 不再误杀同 source 不同 chunk
```
**明确不包含:**
```text
- sparse/BM25 schema
- 全量重灌
- hybridSearch 开关
- 旧 SDK 删除
```
**完成定义(交付 1 Done):**
```text
[ ] 同文档多相关 chunk 可同时出现在内部 evidenceBlocks
[ ] Agent 投影后仍能看到多于 1 条同文档片段(未超预算时)
[ ] retrieve-k / return-n 可配置且行为可测
[ ] 候选身份字段稳定,足够支撑后续 hybrid hit 映射
[ ] 不依赖旧 SDK 新增逻辑
```
### 交付 2:Hybrid 写入 + 查询
**主题:** 在交付 1 的身份/裁剪契约稳定后,打开 dense + sparse/BM25 与库内 RRF。
**范围:**
```text
1. 新 collection schema(dense + sparse/BM25 + 必要标量字段)
2. 写入 dense + sparse,doc_id 级删除与重灌
3. KnowledgeSearchPort 实现 hybrid 模式
4. ranker = RRF(默认)/ Weighted(可配路权)
5. category filter 策略:
- 唯一 domain 时 filtered hybrid
- 低质时 unfiltered 兜底(或双路径轻量合并)
6. 分数语义区分 fused/dense/sparse,重标定 found/relevance
7. 回归评测与延迟对比
8. 冻结并最终删除旧 SDK 读路径
```
**完成定义(交付 2 Done):**
```text
[ ] 可配置 dense | hybrid 切换
[ ] hybrid 默认 RRF,术语类与语义类回归不回退
[ ] filter 误杀有兜底
[ ] 应用层仍按 chunk 身份去重,hybrid 多命中不会被 source 级逻辑误伤
[ ] 新逻辑不再依赖旧 SDK search
```
### 节奏与评审方式
```text
规划:一件事(检索质量主线)
设计评审:按交付 1 + 交付 2 两章看
开发:
先合并交付 1(可独立上线/验证)
再做交付 2(数据迁移 + hybrid 开关)
验收:
交付 1 用“同文档多 chunk”用例
交付 2 用“术语/语义/filter 兜底/延迟”用例
```
### 反模式(实现讨论时直接否决)
```text
❌ 一个大 PR:去重 + schema 重灌 + hybrid 开关 + 删 SDK
❌ 先上 hybrid、后补 chunk 去重
❌ 交付 1 仍按 source 去重,只把 hybrid 分数接进来
❌ 为 hybrid 新建第三套与 candidate/EvidenceBlock 并行的结果模型长期共存
❌ 在旧 SDK search 实现里继续堆 hybrid 细节作为长期方案
```
---
## 2. 现状对照
| 层级 | 当前实现 | Hybrid 需要 |
|---|---|---|
| Collection | `id / vector / content / metadata` | 至少再有 sparse/BM25 文本检索能力 |
| 写入 | `VectorIndexService` 只写 dense | 同步维护 dense + sparse/BM25 |
| 检索门面 | `VectorSearchService`:`sdk \| spring \| auto` | 单端口:`search(query, options)`,内部可 hybrid |
| 旧 SDK 路径 | `MilvusServiceClient.search` | **废弃,不再作为主实现** |
| Spring AI 路径 | `VectorStore.similaritySearch` | 可作过渡 dense 读路径,但 hybrid 能力要单独确认/扩展 |
| 后处理 | 规则 boost + source 去重 | 融合交给库;后处理做裁剪/等级/打包 |
| Agent 投影 | `RagResultProjector` | 基本不动 |
当前关键文件:
```text
写入:
VectorIndexService
DocumentChunkService
VectorEmbeddingService
MilvusClientFactory # schema/index 创建(旧)
读取:
VectorSearchService # 门面,含 sdk/spring 路由
KnowledgeDocumentRetriever
KnowledgeEvidencePostProcessor
LookupKnowledgeTool
配置:
retrieval.vector-store.mode
retrieval.kb-scope
rag.top-k
```
---
## 3. 目标架构(不绑旧 SDK)
```text
┌─────────────────────────┐
upload/init │ KnowledgeWritePort │
chunk + embed -> │ - upsertChunks() │
│ - deleteByDocId() │
└───────────┬─────────────┘
│
▼
Vector DB
dense + sparse/BM25
metadata filters
▲
┌───────────┴─────────────┐
lookup_knowledge │ KnowledgeSearchPort │
query + options-> │ - search() │
│ - mode: DENSE/HYBRID │
└───────────┬─────────────┘
│
▼
KnowledgeDocumentRetriever
│
▼
PostProcess(dedup/chunk cap/threshold/pack)
│
▼
LookupResult / Projector
```
```mermaid
flowchart TB
subgraph write_port["写入边界"]
W[KnowledgeWritePort<br/>upsert / deleteByDocId]
end
subgraph search_port["检索边界"]
S[KnowledgeSearchPort<br/>mode DENSE / HYBRID]
end
W --> VDB[(Vector DB<br/>dense + BM25 sparse<br/>metadata filter)]
S --> VDB
UP[upload/init] --> W
LK[lookup_knowledge] --> S
S --> RET[DocumentRetriever]
RET --> POST[PostProcess<br/>dedup / cap / threshold / pack]
POST --> PROJ[LookupResult / Projector]
PROJ --> AG[Agent]
```
说明:
- **Port** 是应用边界,实现可换成 Spring AI、Milvus 新客户端、或其他封装。
- 旧 `MilvusServiceClient` 检索实现可以暂时留着,但 **新功能不要往里堆**。
- Hybrid 是 `KnowledgeSearchPort` 的一种 mode,不是再开一套平行 tool。
---
## 4. Schema 改造清单
### 4.1 建议逻辑模型
```text
id string PK # chunk 级唯一 id
doc_id string # 文档 id(从 metadata 提升为一等字段更稳)
chunk_index int
content text/varchar # 原始 chunk 正文(给 BM25 / 返回)
title string nullable
breadcrumb string nullable
category string nullable
kb_scope string nullable
dense_vector float vector # embedding(title/path/content)
sparse_vector sparse vector # BM25 或 sparse embedding
metadata json # 兼容扩展字段
```
### 4.2 和现状差异
| 字段 | 现状 | 建议 |
|---|---|---|
| `vector` | 有 | 可改名 `dense_vector`,或保留别名兼容 |
| `content` | 有,仅存储/返回 | 同时作为 BM25 输入文本 |
| `sparse_vector` | 无 | **新增,hybrid 必需** |
| `docId/chunkIndex` | 塞在 JSON metadata | 建议提升为可过滤/可排序字段 |
| `category/kb_scope` | metadata JSON | 建议提升,filter 更稳 |
### 4.3 索引
```text
dense_vector -> 向量索引(COSINE/IP/L2,与 embedding 一致)
sparse_vector -> 稀疏倒排 / BM25 索引
category/kb_scope/doc_id -> 标量过滤索引(如需要)
```
### 4.4 迁移策略
不要幻想“只改 search 方法”:
1. **新建 collection 或新版本 collection**(推荐)
2. 全量重灌知识库(dense + sparse)
3. 双写一段时间(可选)
4. 切换读路径到 hybrid
5. 下线旧 collection / 旧 SDK 读路径
就地改老 collection 风险高:已有数据无 sparse,历史 metadata 形态也不统一。
---
## 5. 写入路径改造清单
### 5.1 需要动的职责
| 类/模块 | 现在 | 改造 |
|---|---|---|
| `DocumentChunkService` | 产出 chunk 正文/title/breadcrumb | 基本可复用 |
| `VectorEmbeddingService` | 只做 dense embed | 保留;sparse/BM25 另算或交给库 |
| `VectorIndexService` | 组装 metadata + insert dense | 升级为 write port 实现:dense+sparse 一并 upsert |
| 删除逻辑 | 按 `metadata.docId` / `_source` 删 | 统一按 `doc_id` 删,避免路径不一致 |
### 5.2 写入时每条 chunk 必须具备
```text
- dense_vector: embed(buildEmbeddingText(chunk))
- sparse 输入: 建议用“可检索文本”
title + breadcrumb + content
而不是只丢 raw content
- doc_id / chunk_index / category / kb_scope
- 稳定 chunk id(doc_id + chunk_index 派生)
```
```mermaid
flowchart LR
CHUNK[DocumentChunk] --> EMB[dense embed]
CHUNK --> ST[search_text<br/>title+path+content]
CHUNK --> META[docId/chunkIndex<br/>category/kb_scope]
EMB --> ROW[upsert row]
ST --> ROW
META --> ROW
ROW --> FN[BM25 Function<br/>search_text → sparse]
ROW --> COL[(collection)]
FN --> COL
```
### 5.3 注意
- embedding 文本可以继续拼 `Title/Path/Content`
- **返回给 Agent 的 content 仍应是原文 chunk**,不要返回 embedding 拼接串
- BM25 文本建议包含 title/breadcrumb,否则专有名词在标题里时字面路会弱
---
## 6. 检索路径改造清单
### 6.1 新检索端口(建议)
不要继续扩:
```text
searchSimilarDocuments(query, topK, category)
```
建议收敛成:
```text
SearchRequest {
query: string
retrieveK: int # 例如 20
returnN: int # 例如 5,可在后处理裁
mode: DENSE | HYBRID
categoryFilter?: string
kbScope?: string
ranker: RRF | WEIGHTED
rrfK: int # 默认 60
weights?: {dense, sparse}
}
SearchHit {
id, docId, chunkIndex
content, title, breadcrumb
source, category
scores: {
fused?, denseRank?, sparseRank?, raw?...
}
metadata
}
```
```mermaid
flowchart TB
REQ[SearchRequest<br/>query · retrieveK · mode<br/>filter · rrfK] --> PORT[KnowledgeSearchPort]
PORT -->|DENSE| D[dense ANN only]
PORT -->|HYBRID| H[dense + BM25 + RRF]
D --> HIT[SearchHit 列表<br/>id/docId/chunk · content · ranks]
H --> HIT
HIT --> APP[后处理 / 投影]
```
### 6.2 `VectorSearchService` 怎么演进
短期:
```text
保留门面类名也可
但内部:
- 不再把 sdk 当长期分支
- 增加 hybridSearch(...) 能力
- mode 配置改为:
dense | hybrid
(spring 仅作 dense 兼容实现)
```
中期:
```text
VectorSearchService 实现 KnowledgeSearchPort
旧 sdk 分支删除或仅 test/fallback 开关默认关
```
### 6.3 Hybrid 查询语义
```text
路 A: dense(query_embedding) limit=retrieveK
路 B: bm25/sparse(query_text) limit=retrieveK
可选过滤: category / kb_scope
融合: RRFRanker(k=60) 或 WeightedRanker
输出: top retrieveK/returnN
```
对应我们之前的公式:
```text
RRF_w(d) = Σ w_i / (k + rank_i(d))
```
- 用 RRF:先不调权重
- 用 Weighted:调的是 **dense/sparse 路权**,不是 keyword contains 加分
### 6.4 filtered + unfiltered 还要不要?
还要,但定位变了:
| 能力 | 放哪 |
|---|---|
| dense + bm25 融合 | **库内 hybrid** |
| category filter 开/关 | 应用策略层,可变成两次 hybrid 或 filter 参数 |
| chunk 去重 / 每文档上限 | 应用后处理 |
| found / relevanceLevel | 应用后处理 |
推荐策略:
```text
if 唯一 domain:
hybrid(query, filter=category) # 主路
若低质量:
hybrid(query, filter=null) # 兜底
或并行两条 hybrid 再做一次轻量合并
else:
hybrid(query, filter=null)
```
注意:这里的“两条”是 **filter 策略双路径**,不是再手写一套 dense/bm25 融合。
---
## 7. 和现有 pipeline 的衔接(按类)
### 7.1 基本不动
| 类 | 原因 |
|---|---|
| `LookupKnowledgeTool` | 继续编排 transform → retrieve → post → pack |
| `KnowledgeQueryTransformer` | 仍产 categoryFilter / hints |
| `KnowledgeContextPacker` | 仍做字符预算 |
| `LookupResultAssembler` | 仍组装内部结果 |
| `RagToolAdapter` / `RagResultProjector` | Agent 契约保持稳定 |
### 7.2 要改
| 类 | 改什么 |
|---|---|
| `KnowledgeDocumentRetriever` | 调新 search port;透传 retrieveK/mode;把 docId/chunkIndex 提成候选一等字段 |
| `KnowledgeEvidencePostProcessor` | 弱化“跨路融合”职责;保留 dedup、chunk cap、阈值、轻精排 |
| `VectorSearchService` | 成为 hybrid 入口,去掉对旧 sdk 的依赖增长 |
| `VectorIndexService` | 写入 dense+sparse,统一 doc_id 删除 |
| schema 工厂/初始化 | 新 collection 定义与索引 |
### 7.3 后处理职责重新划分
**交给 Milvus hybrid:**
- dense/sparse 多路召回
- RRF / weighted 融合
- 基础 topK
**留给应用层:**
```text
1. chunk 级去重(docId#chunkIndex)
2. maxChunksPerDocument
3. return-n 裁剪
4. relevanceLevel / isLowQuality
5. 可选轻规则精排(title/breadcrumb 命中)
6. context pack 与 Agent 投影
```
**应降级或删除的:**
```text
把 keyword contains 大额加分当主排序器
在应用层重复实现一套 dense+bm25 分数硬加
```
L0 仍只做:
```text
- 是否启用 category filter
- 轻量精排特征
- trace 解释
```
---
## 8. 配置建议(示意)
```properties
# 检索模式:dense | hybrid
retrieval.search.mode=hybrid
# 旧 sdk 读路径默认关闭(后续删除)
retrieval.legacy-sdk.enabled=false
# 召回/返回分离
rag.retrieve-k=20
rag.return-n=5
rag.max-chunks-per-document=2
# 融合
retrieval.hybrid.ranker=rrf
retrieval.hybrid.rrf-k=60
# 若用 weighted:
# retrieval.hybrid.ranker=weighted
# retrieval.hybrid.weight.dense=1.0
# retrieval.hybrid.weight.sparse=0.8
# filter 策略
retrieval.filter.retry-unfiltered-on-low-quality=true
retrieval.normalization.reference-threshold=0.5
```
说明:
- `rag.top-k` 应逐步废弃,避免“召回/返回/展示”一个参数打天下
- 阈值字段若 hybrid 后分数语义变化,需要重新校准,不能照搬旧 L2 经验值
---
## 9. 分数与阈值:hybrid 后要重标定
当前后处理默认假设:
```text
score ≈ 兼容 L2 距离
normalizeL2 后得到 0~1
```
hybrid 后常见变化:
| 来源 | 语义 |
|---|---|
| dense raw | L2 / cosine |
| sparse/BM25 raw | 另一套 |
| fused RRF | 名次融合分,不是相似度概率 |
因此:
1. `SearchHit` 要区分 `fusedScore` / `denseScore` / `sparseScore`
2. `isLowQuality` 不要直接拿 RRF 分当旧 L2 用
3. 过渡期可:
- 用“是否有命中 + 规则完整性”判断 found
- 或只对 dense 分做阈值,RRF 只负责排序
4. 重新用 15~30 条回归 query 标定
---
## 10. 分阶段落地(推荐)
> 与 **§1.1 交付 1 / 交付 2** 对齐。Phase 编号用于执行拆解;对外沟通优先用两个交付里程碑。
### 交付 1 对应 Phase
#### Phase 0:应用层前提(交付 1 核心)
```text
[ ] chunk 级去重(不要 source 级吞 chunk)
[ ] 候选暴露 docId / chunkIndex / evidenceKey
[ ] retrieve-k / return-n 分离
[ ] maxChunksPerDocument
[ ] Projector 按证据身份去重(不再 source 唯一)
[ ] 明确 legacy-sdk 读路径仅兼容、默认关或冻结
[ ] 单测:同文档多 chunk / 真重复 / 单文档上限
```
#### Phase 1:检索端口收敛(交付 1 建议同批或紧随)
```text
[ ] 定义 KnowledgeSearchPort / SearchRequest / SearchHit
[ ] VectorSearchService 适配该端口(先 dense-only 也可)
[ ] KnowledgeDocumentRetriever 只依赖端口
[ ] 单测用 fake search port,不再绑 SDK 细节
```
**交付 1 出口:** 不上 hybrid 也可合并;hybrid 所需 hit 身份与裁剪契约已稳定。
### 交付 2 对应 Phase
#### Phase 2:写入与 schema 支持 sparse/BM25
```text
[ ] 新 collection schema
[ ] 写入 dense + sparse/BM25 文本
[ ] doc_id 级删除与重灌
[ ] 知识库全量重建脚本/任务
```
#### Phase 3:打开 hybrid 读路径
```text
[ ] search.mode=hybrid
[ ] ranker=rrf
[ ] category filter 策略接入
[ ] low-quality 时 unfiltered 兜底
[ ] trace 记录 dense/sparse/fused 信息(内部)
[ ] 确认 hybrid 多命中仍走 chunk 级去重,不被 source 误伤
```
#### Phase 4:瘦身后处理 + 下线旧路径
```text
[ ] 规则 boost 降为轻精排或可关
[ ] 删除/隔离旧 SDK search 实现
[ ] 校准 found/relevance 阈值
[ ] 回归评测与延迟对比
```
**交付 2 出口:** dense|hybrid 可切换;RRF 默认可用;旧 SDK 检索不再被新逻辑依赖。
---
## 11. 类级改造对照表
| 类 | 优先级 | 动作 | 是否依赖旧 SDK |
|---|---|---|---|
| `KnowledgeDocumentRetriever` | P0 | 接新端口,透传 hybrid 选项,补 chunk 身份 | 否 |
| `KnowledgeEvidencePostProcessor` | P0 | 去重改 chunk 级;融合职责外移 | 否 |
| `LookupKnowledgeTool` | P1 | 使用 retrieve-k/return-n;保留 filter 降级策略 | 否 |
| `VectorSearchService` | P0 | 增加 hybrid;冻结/移除 sdk 增长 | 实现可无 SDK |
| `VectorIndexService` | P0 | dense+sparse 写入,doc_id 删除 | 实现可无 SDK |
| `MilvusClientFactory` | P2 | 仅迁移期维护;新 schema 建议新模块 | 旧 |
| `VectorEmbeddingService` | P1 | 继续 dense;不塞 hybrid 逻辑 | 否 |
| `RagResultProjector` | P2 | 若 document_id 变 chunk 级,同步语义 | 否 |
| Spring AI `VectorStore` | P2 | 可继续承载 dense;hybrid 需单独能力层 | 否 |
---
## 12. 测试清单
### 单元
```text
[ ] RRF 融合结果顺序(可用 fixture,不连库)
[ ] chunk 去重:同 doc 不同 chunk 都保留
[ ] 同 doc 超过 maxChunksPerDocument 被裁
[ ] category filter 低质时走 unfiltered
[ ] SearchHit 字段映射:docId/chunkIndex/content
```
### 集成 / 回归
```text
[ ] 专有名词/错误码 query:hybrid 应优于 pure dense
[ ] 换说法语义 query:hybrid 不低于 pure dense
[ ] 错误 domain filter:unfiltered 兜底仍能找回
[ ] 长文档多 chunk:返回不少于 2 个相关片段(若存在)
[ ] 延迟:hybrid P95 可接受
[ ] 重灌后旧 doc 删除干净,无幽灵 chunk
```
### 兼容
```text
[ ] mode=dense 仍可用(回滚开关)
[ ] Agent 契约字段不破(evidence/document_id/excerpt)
[ ] tool_invocation / trace 仍有 selectedAttempt 与基本检索信息
```
---
## 13. 明确不做的事
1. **继续在旧 SDK `search` 上叠 hybrid 细节当长期方案**
2. **应用层把 dense raw 分和 BM25 raw 分直接相加**
3. **用 L0 contains 大额加分替代库内 RRF**
4. **只改查询、不重灌 sparse 数据**
5. **hybrid 后仍拿旧 L2 阈值硬套 fused score**
6. **让 Agent 直接依赖内部 fused/raw score 字段**(除非契约明确升级)
---
## 14. 和前序讨论的对齐
| 讨论结论 | 在本清单中的落点 |
|---|---|
| K=3 不必先上复杂 rerank | `retrieve-k=20, return-n=5` |
| filtered + unfiltered 有价值 | hybrid 之上的 filter 策略双路径 |
| BM25 是跨维度召回 | schema sparse/BM25 + hybrid 路 |
| 跨路优先 RRF | 库内 `RRFRanker` |
| `w_i` 是路权 | `WeightedRanker` / 配置 weight.dense/sparse |
| L0 只做导航 | 仅影响 filter 与轻精排,不负责主融合 |
| 旧 SDK 后续废弃 | 新开发只走 search/write port,不绑 SDK |
---
## 15. 最小可交付定义(MVP)
MVP 拆成两个可独立验收的里程碑(与 §1.1 一致)。
### MVP-1:分块去重与身份契约(交付 1)
```text
1. 候选/证据具备 docId、chunkIndex、evidenceKey
2. 后处理与投影均按 chunk 身份去重
3. maxChunksPerDocument 生效
4. retrieve-k / return-n 分离且可测
5. 同文档多相关 chunk 在未超预算时可同时到达 Agent
6. 不新增对旧 SDK 的依赖
```
### MVP-2:Hybrid 接入(交付 2)
```text
1. 新 collection 可写入 dense + BM25/sparse
2. 知识库可全量重建
3. lookup_knowledge 可通过配置切换 dense/hybrid
4. hybrid 默认 RRF 融合
5. 应用层 chunk 去重与 return-n 在 hybrid 下仍正确
6. 旧 SDK 检索不再被新逻辑依赖
7. 至少一套回归 query 证明:
- 术语类 query 不回退
- 语义类 query 不回退
- filter 误杀有兜底
```
只有 MVP-1 完成,才建议开始 MVP-2 的数据迁移与开关切换。
---
## 16. 建议的下一步实现顺序(动手时)
实现讨论与排期默认按此顺序:
```text
1. 交付 1 / MVP-1
- chunk 身份
- chunk 去重 + maxChunksPerDocument
- retrieve-k / return-n
- Projector 对齐
- SearchPort 雏形(建议)
2. 交付 2 / MVP-2
- schema + 重灌
- hybrid RRF 读路径
- filter 兜底
- 阈值校准
- 下线旧 SDK 读路径
```
**不要跳过交付 1 直接做 hybrid。**
交付 1 不依赖旧 SDK,不阻塞后续 hybrid,且单独合并就有质量收益。
---
## 17. 后续实现讨论检查清单
开会或开 PR 前,用下面问题对齐范围:
```text
[ ] 本次是交付 1、交付 2,还是仅其中子项?
[ ] 是否改动了证据身份字段(docId/chunkIndex/evidenceKey)?
[ ] 去重 key 是否仍存在 source 级路径?
[ ] retrieve-k 与 return-n 是否仍混用 top-k?
[ ] 是否把 schema 重灌/hybrid 开关误塞进交付 1?
[ ] 是否有新增旧 SDK 依赖?
[ ] 单测是否覆盖“同文档多 chunk”?
[ ] 若已 hybrid:fused score 是否被误当成旧 L2 阈值?
```
+396
View File
@@ -0,0 +1,396 @@
# Agent 如何读 `relevance_level`:它是什么、不是什么
**日期**:2026-07-28
**范围**:`lookup_knowledge` 投影给 Agent 的粗粒度相关度标签
**读者**:要在 Diagnosis Agent / 工具契约里正确使用知识库结果的工程与提示词同学
**关联**:
- ACI 契约:`RagToolResult.relevance_level` / `RagRelevanceLevel`
- 计算:`KnowledgeEvidencePostProcessor` → `qualityScore` 阈值
- 质量统一:`docs/RAG-Hybrid质量分与后处理.md`
- 多路与 RRF:`docs/RAG排序-多路召回与RRF.md`
- **运行时 Trace / 审计**:`mvp/architecture/RAG检索可观测性与审计.md`
---
## 1. 一句话定义
**`relevance_level` 是对「这一次 `lookup_knowledge` 调用整体有多相关」的粗档标签,不是某一条 evidence 的分数,也不是 0~1 的相似度。**
Agent 真正写诊断、做引用时,仍应以 `evidence[]` 里的 excerpt 为准;`relevance_level` 只帮助判断:**这批评据大概有多硬、还要不要再查。**
```mermaid
flowchart LR
subgraph tool["lookup_knowledge 结果"]
ES[evidence_status]
EV[evidence excerpt]
RL[relevance_level]
TR[truncated]
end
ES --> DEC{Agent 决策}
EV --> DEC
RL --> DEC
TR --> DEC
DEC --> A1[写结论 / 引用]
DEC --> A2[再查 / 换工具]
DEC --> A3[证据不足降级]
```
---
## 2. 它出现在哪里
成功(或有结果)的知识库工具投影里,典型形状:
```json
{
"evidence_status": "EVIDENCE_FOUND",
"tool_call_id": "call-…",
"query": "用户/Agent 的检索句",
"evidence": [
{
"document_id": "doc#chunk-0",
"source": "…",
"title": "…",
"breadcrumb": "…",
"excerpt": "…"
}
],
"returned_count": 3,
"relevance_level": "PRECISE",
"truncated": false
}
```
要点:
- JSON 字段名是 **`relevance_level`**(snake_case)
- Java 枚举:`RagRelevanceLevel`(`PRECISE` / `HIGHLY_RELEVANT` / `REFERENCE`)
- 无可用证据或质量不够时,字段常为 **null / 省略**(`NON_NULL`)
- Agent **看不到** raw L2、RRF 分、`retrievalTrace`、`rerankTrace`(ACI 有意裁掉)
```mermaid
flowchart TB
subgraph internal["内部 LookupResult(审计/调试)"]
QS[qualityScore / topSimilarity]
RT[retrievalTrace / rerankTrace]
CP[contextPack 全文]
SC[raw score / scoreLabel]
end
subgraph project["RagResultProjector"]
P[裁剪与规范化]
end
subgraph agent["Agent 可见 RagToolResult"]
A1[evidence_status]
A2[tool_call_id / query]
A3[evidence excerpt 列表]
A4[relevance_level 可选]
A5[returned_count / truncated]
end
internal --> P --> agent
```
---
## 3. 枚举值怎么理解
| 值 | 产品语义 | Agent 侧更合理的用法 |
|----|----------|----------------------|
| **PRECISE** | 整体很贴:当前 top 证据质量高,继续同题检索不太可能更准 | 优先依据 `evidence[]` 组织结论;避免无意义的重复 `lookup_knowledge` |
| **HIGHLY_RELEVANT** | 高度相关(枚举保留) | 与 PRECISE 类似,略保守表述即可 |
| **REFERENCE** | 可作参考,但不到「已精准命中」 | 可引用,但结论留余地;缺维度时换 query 再查或叠日志/指标 |
| **null / 不出现** | 没有可报的粗相关档(无证据或 top 质量偏低) | **不要**当成知识库已证实;按证据不足处理 |
### 和 `evidence_status` 的分工
| 字段 | 回答的问题 |
|------|------------|
| `evidence_status` | 这次有没有合法、可引用的证据(如 `EVIDENCE_FOUND` / `NO_EVIDENCE`) |
| `relevance_level` | **有证据时**,整体有多贴(粗档) |
| `evidence[]` | 具体可以引用哪些片段 |
没有证据时,不应指望靠 `relevance_level`「升级」出结论;契约上也不会用 level 把空结果扮成有证据。
```mermaid
flowchart TB
ES{evidence_status}
ES -->|NO_EVIDENCE| N1[不要当知识库已证实]
ES -->|EVIDENCE_FOUND| RL{relevance_level}
RL -->|PRECISE| U1[优先引用 excerpt · 少重复检索]
RL -->|REFERENCE| U2[可引用 · 结论留余地]
RL -->|null / 缺省| U3[有块但质量偏低 · 慎用强结论]
RL -->|HIGHLY_RELEVANT| U4[与 PRECISE 类似 · 略保守]
```
---
## 4. 它是怎么算出来的(实现口径)
### 4.1 只看「本轮第一名」的 qualityScore
后处理在完成排序、去重、截断之后:
```text
取 originalRank 最优(排序后第一条)的 qualityScore
≥ highly-relevant-threshold (默认 0.75) → PRECISE
≥ reference-threshold (默认 0.5) → REFERENCE
否则 / 无可用证据 → null
```
配置(`application.yml`):
```yaml
retrieval:
normalization:
max-l2-distance: 2.0
highly-relevant-threshold: 0.75
reference-threshold: 0.5
```
因此:
- level 描述的是 **整次调用的 top 质量**,不是每条 evidence 各打一档
- 列表里第 2、第 3 条即使偏弱,只要 top1 够高,整次仍可能是 `PRECISE`
```mermaid
flowchart TB
RET[检索有序候选] --> POST[后处理:保 originalRank · 去重 · 截断]
POST --> TOP[取排序后第一条 qualityScore]
TOP --> T1{≥ 0.75?}
T1 -->|是| PRECISE[PRECISE]
T1 -->|否| T2{≥ 0.5?}
T2 -->|是| REF[REFERENCE]
T2 -->|否| NULL[null / 不报档]
```
### 4.2 qualityScore 从哪来(和检索 mode 绑定)
统一经 `RetrievalScoreNormalizer.toQualityScore`:
| `retrieval.search.mode` | top qualityScore 含义 |
|-------------------------|------------------------|
| **dense** | top1 的 L2 归一化:约 `1 - L2 / maxL2Distance` |
| **hybrid** | **优先**同 id 的 `denseDistance`(绝对 L2 质量,供闸门/level);无 dense 邻域时 **rank 回退** |
```mermaid
flowchart LR
subgraph dense_mode["mode=dense"]
L2[score = L2] --> QD[quality = 1 - L2/max]
end
subgraph hybrid_mode["mode=hybrid"]
RRF[RRF 序 = originalRank] --> SORT[列表顺序]
DD[denseDistance 可选] --> QH{有 dense?}
QH -->|是| QL[quality = L2 归一化]
QH -->|否| QR[quality = rank 映射]
end
QD --> LV[relevance_level]
QL --> LV
QR --> LV
```
读 level 时注意:
> **dense 下的 PRECISE ≈「向量足够近」**
> **hybrid 下的 PRECISE ≈「top1 的绝对/回退 quality 跨过了 0.75」**;排序仍跟 RRF,不是「又变回只信 L2 排序」。
不要把 hybrid 的 level 读成与 dense **完全同一把尺子**,但也不要再假设「rank1 永远 PRECISE」(在附带 denseDistance 后,远邻 decoy 可以很低分并触发 filter fallback)。
### 4.3 关于 HIGHLY_RELEVANT
枚举和旧文档里仍有三档。历史上大致是:
```text
高分 + L0 hint 支撑 → PRECISE
高分但无 hint → HIGHLY_RELEVANT
中等分 → REFERENCE
```
质量分统一之后,当前实现是:**≥ 0.75 直接 PRECISE**,不再要求 L0 contains 才能精准。
因此运行时 **很少再单独产出 HIGHLY_RELEVANT**;读旧 trace / 旧快照时仍可能见到。
---
## 5. Agent 应该怎么读(建议协议)
### 5.1 推荐读法
```text
1. 先看 evidence_status
2. 再读 evidence[] 的 excerpt(唯一可引用正文)
3. 用 relevance_level 调节「敢多敢少」与「要不要再查」
```
```mermaid
flowchart TB
START[收到 RagToolResult] --> S1{evidence_status}
S1 -->|无证据| FAIL[不编造 · 换工具或安全降级]
S1 -->|有证据| S2[精读 evidence excerpt]
S2 --> S3{relevance_level}
S3 -->|PRECISE| C1[结论可较硬 · 少重复同 query 检索]
S3 -->|REFERENCE| C2[结论留余地 · 可换问法或叠日志指标]
S3 -->|缺省| C3[慎用强结论 · 优先补查]
C1 --> CITE[引用必须落在 excerpt]
C2 --> CITE
C3 --> CITE
```
| 组合 | 建议行为 |
|------|----------|
| FOUND + PRECISE | 以 excerpt 为主写结论;少重复同 query 检索 |
| FOUND + REFERENCE | 可引用,表述保守;缺关键事实则改写 query 或换工具 |
| FOUND 但 level 空 | 有块但质量闸门偏低:慎用强结论,优先补查 |
| NO_EVIDENCE | 不编造知识库依据;走其他证据工具或安全降级 |
### 5.2 明确不要这样读
1. **不要当逐条相关度**
没有 `evidence[i].relevance_level`;不能说「第 2 条是 REFERENCE」。
2. **不要当连续分数**
没有 0.83;只有粗档。不要在推理里假装有精确分。
3. **不要在 hybrid 下当成绝对语义相似度**
序数 quality 下 PRECISE 很常见,表示「本轮第一够格」,不等于「全局语义必近」。
4. **不要代替 excerpt 引用**
level 不能当证据正文;Gatekeeper / EvidenceGuard 认的是可核对片段与引用约束。
5. **不要和 Harness 验真混为一谈**
level 是检索侧粗标;工具生命周期、`evidence_status`、守卫校验是另一层。
6. **不要用它驱动跨 mode 对比**
同一 query 切 dense/hybrid 时,比命中集合与排名;别只比「是不是都 PRECISE」。
---
## 6. Agent 看不见、但会影响 level 的内部量
便于排查「为什么突然全是 PRECISE / 总是 null」:
| 内部量 | 作用 | Agent 是否可见 |
|--------|------|----------------|
| `qualityScore` / `topSimilarity` | 定 level、低质 retry | 否 |
| `originalRank` | 排序权威;hybrid quality 输入 | 否 |
| `score` + `scoreLabel` | dense=L2 / hybrid=融合侧 | 否 |
| L0 domain/keyword | 现仅 hitReasons 解释,不改序、不抬 level | 否(reasons 也可能被投影裁掉) |
| `completenessHint` | 内部完整度文案 | 通常否 |
| category filter + unfiltered retry | 低质时可能换一批 evidence 再定 level | 过程 trace 否 |
投影原则(ACI):模型只要能理解与引用结果;**不给 raw score、阈值、trace。**
```mermaid
flowchart TB
subgraph pipe["检索管道内部"]
STORE[Milvus hybrid/dense]
NORM[toQualityScore]
POST[PostProcessor]
STORE --> NORM --> POST
end
POST --> LV[relevance_level]
POST --> EB[evidenceBlocks]
POST --> TR[traces · 通常不投影]
EB --> PROJ[RagResultProjector]
LV --> PROJ
PROJ --> AGENT[Agent Observation]
```
---
## 7. 和相邻概念的边界
```text
evidence_status 有没有证据
relevance_level 有的话有多贴(粗)
evidence[].excerpt 贴在哪一段文字上(细、可引用)
truncated 列表是否被预算截断(可能还有更好的没展示)
information_gain 等 诊断环路里「这轮工具对任务有没有增益」(另一契约)
```
```mermaid
flowchart TB
subgraph fields["同一次工具结果里的分工"]
ES[evidence_status<br/>有没有]
RL[relevance_level<br/>有多贴·粗]
EX[excerpt<br/>说什么·细]
TC[truncated<br/>是否被截断]
end
ES --> USE[Agent 使用]
RL --> USE
EX --> USE
TC --> USE
USE --> NOTE[结论锚在 excerpt<br/>level 只调力度]
```
`truncated=true` 时:即使 `PRECISE`,也只说明 **已返回子集里的 top 很强**,不保证库内没有更相关却被截掉的块。
---
## 8. 提示词 / 产品文案可用的短说明
可直接给模型或文档的精简版:
```text
relevance_level 是本次知识库检索的整体相关度粗标:
- PRECISE:当前证据整体很贴,优先引用 evidence 写结论,避免无意义重复检索
- REFERENCE:仅供参考,结论需留余地,必要时换问法或改用其他工具
- 缺省:不要把本次结果当作高置信知识库证实
务必以 evidence 中的 excerpt 为唯一引用依据;不要编造未出现的文档内容。
在 hybrid 检索下,PRECISE 更多表示「本轮排序第一档」,不是精确相似度分数。
```
---
## 9. 常见误读示例
| 误读 | 更正 |
|------|------|
| 「PRECISE 所以三条 evidence 都精准」 | 只保证 top 质量跨线;其余条只是同批返回 |
| 「没有 relevance_level 就是工具失败」 | 更可能是无证据或质量偏低;看 `evidence_status` |
| 「hybrid 全是 PRECISE 说明召回完美」 | 可能只是 rank1→quality=1.0 的档位特性 |
| 「REFERENCE 的 excerpt 不能引用」 | 可以引用,但结论强度应下调 |
| 「level 高就可以跳过 excerpt」 | 不可;引用与验真仍看正文 |
---
## 10. 结语
`relevance_level` 是检索链路送给 Agent 的 **粗粒度驾驶辅助**:
- 告诉模型这批评据大概硬不硬
- **不**替代 excerpt,**不**暴露打分细节,**不**等于逐条标注
在 dense 模式下,它更接近「向量有多近」;
在 hybrid 主路径下,它更接近「本轮融合第一名是否跨过质量门槛」。
读的时候记住三句即可:
```text
1. 先 status,再 excerpt,最后才看 level
2. level 管「敢多敢少」,excerpt 管「说了什么」
3. hybrid 的 PRECISE ≠ 绝对语义满分
```
---
## 附录:代码与规格锚点
| 项 | 位置 |
|----|------|
| Agent 结果契约 | `RagToolResult` / `RagRelevanceLevel` |
| 投影 | `RagResultProjector` |
| 等级计算 | `KnowledgeEvidencePostProcessor.computeRelevance` |
| 质量分 | `RetrievalScoreNormalizer` |
| 规格 | `openspec/specs/aci-evidence-tool-contracts`、`rag-retrieval-quality-score` |
| 架构 | `mvp/architecture/RAG知识检索架构.md` §9 |
+592
View File
@@ -0,0 +1,592 @@
# 混合检索上线之后:为什么还要统一 qualityScore,以及上一代后处理错在哪
**日期**:2026-07-28
**范围**:hybrid 检索后的分数语义、后处理排序、相关度闸门、scoreLabel 约定
**读者**:已经(或准备)上 dense+BM25+RRF,却发现「召回变了、质量判断还拧着」的工程同学
**关联实现**:
- `lookup_knowledge` 模块化链路
- `MilvusHybridKnowledgeStore`(dense / hybrid)
- `RetrievalScoreNormalizer` / `KnowledgeEvidencePostProcessor`
- OpenSpec / devflow:`rag-quality-score-unify`
- 前置讨论:`docs/RAG排序-多路召回与RRF.md`
- 架构:`mvp/architecture/RAG知识检索架构.md` §6
---
## 1. 引言:融合排好了序,不等于质量链路闭环了
上一篇文章(《诊断 Agent 场景下的 RAG 排序:从 K=3 规则加分,到多路召回与 RRF》)回答的是:
> 候选太少时别急着上精排;跨路不要硬加原始分;优先 RRF;L0 只做导航。
那一轮讨论之后,工程上陆续落地了:
1. **chunk 级证据身份与去重**(`docId#chunkIndex`,同文档多片段可并存)
2. **真 hybrid**:Milvus 服务端 dense ANN + BM25 sparse + `hybridSearch` + `RRFRanker`
3. **单一知识后端**(`MilvusClientV2`),去掉 sdk/spring 多路由主路径
4. **mode 开关**:`hybrid` 线上主路径,`dense` 同库对照评测
主缺口从「假 hybrid / 粗去重」变成了另一件事:
```text
库内:RRF 已经决定谁先谁后
应用:后处理仍假装每条 score 都是 L2
再用 L0 关键词 contains 加分改序
```
于是出现一种很拧的现象:
- 检索层已经是 **混合检索的世界**
- 质量层还活在 **单路 dense + 规则 boost 的世界**
本文记录的,就是这次对「拧」的拆解、拍板与落地口径:
**统一 qualityScore,废止 hybrid 的 L2 伪装,去掉关键词 boost 改序。**
目标不是再推一套更复杂的模型,而是回答:
> hybrid 上线之后,排序权威和质量闸门到底听谁的?后处理还该不该拿关键词打分?
```mermaid
flowchart TB
subgraph retrieval["检索层 · 已 hybrid"]
Q[query] --> D[dense ANN]
Q --> B[BM25 sparse]
D --> RRF[hybridSearch + RRF]
B --> RRF
RRF --> ORD[RRF 序]
end
subgraph post_old["后处理 · 仍 L2 世界"]
ORD --> FAKE[伪装 / 回填 L2]
FAKE --> BOOST[关键词 boost 改序]
BOOST --> GATE[阈值 / level / retry]
end
post_old --> PAIN[排序与闸门拧巴]
```
---
## 2. 上一代(hybrid 刚落地时)到底长什么样
### 2.1 检索侧:已经是真混合
`mode=hybrid` 时大致是:
```text
query
├─ dense ANN(query embedding) → vector / L2
└─ BM25 sparse(EmbeddedText) → sparse_vector / BM25
│
▼
Milvus hybridSearch + RRFRanker(k)
│
▼
融合后的 hit 列表(RRF 序)
```
```mermaid
flowchart LR
Q[query] --> EMB[embedding]
Q --> TXT[raw text]
EMB --> DA[dense ANN<br/>vector / L2]
TXT --> BA[BM25 ANN<br/>sparse]
DA --> HS[Milvus hybridSearch]
BA --> HS
HS --> RR[RRFRanker]
RR --> HITS[有序 hits]
```
这比应用层 sparse-lite / 伪 hybrid 前进了一大步:词面与语义在**库内**融合,chunk 身份也不会在后处理被文档级折叠吞掉。
### 2.2 分数侧:仍在「骗」后处理
后处理历史契约默认:
```text
score ≈ L2 距离(越小越好)
baseScore = 1 - clamp(L2) / maxL2Distance # 越大越好
再 + domain/entity/keyword boost
按 finalScore 重排
用 baseScore 定 relevance_level / 是否低质 retry
```
为了迁就这套契约,hybrid 路径做了补丁:
```text
hybrid 融合结果
-> 再跑一路 dense
-> 按 id 把 L2 回填到 score,label 改成 l2_distance
-> 仅 BM25 命中、dense 没命中:
score = maxL2Distance
label = bm25_only_no_dense
```
```mermaid
flowchart TB
H[hybrid RRF 结果] --> P[并行 dense 探测]
P --> M{同 id 有 L2?}
M -->|是| O1[score=L2 · label=l2_distance]
M -->|否| O2[score=maxL2 · label=bm25_only]
O1 --> N[normalizeL2 + boost 重排]
O2 --> N
N --> BAD[BM25-only 好证据被当成最差]
```
意图是好的:让 `normalizeL2` 和 0.75/0.5 阈值「还能用」。
副作用也很清楚:
| 现象 | 后果 |
|------|------|
| RRF 决定顺序,L2 决定「好不好」 | 两套真理,互相打架 |
| BM25-only 好证据被标成最远 L2 | quality≈0,像低质,甚至触发 unfiltered retry |
| `bm25_only_no_dense` 像第三种 label | 概念膨胀:mode 其实只有 dense/hybrid |
| 多打一路 dense 只为回填 | 延迟与复杂度,换来的是语义自洽的假象 |
一句话:
> **不是拿 RRF 分错误地套了 L2 公式,而是排序信 RRF,打分/闸门仍假装大家都是 L2。**
### 2.3 后处理侧:关键词 boost 改主序
典型逻辑:
```text
baseScore = normalizeL2(score)
finalScore = baseScore
+ domain_match (+0.15)
+ entity_match (+0.20)
+ keyword_match (+0.10)
+ source_type (+0.05)
按 finalScore 降序
```
在 **还没有库内 BM25** 时,这套东西多少能补一点词面。
在 **已经 hybrid** 之后,问题变成:
1. **双重计分**
BM25 已经在 RRF 里投过票;后处理再用 contains 加分,等于词面再抬一次。
2. **contains 比 BM25 更糙**
无 IDF、无文档长度、短词子串误命中——正好制造「词频/词面高、相关度低却排前面」。
3. **冲掉 RRF 序**
花了 hybrid 买到的融合序,被 L0 词表二次改写。
4. **PRECISE 还绑 hint**
高质量还要 `hasHintSupport`(同样是 contains),把导航层信号抬成等级门槛。
结合上一篇文章的判断——**L0 / 关键词适合做提示,不适合当最终裁判**——hybrid 上线后,后处理 boost 改序已经从「可接受的轻启发式」滑向「明确的设计债」。
---
## 3. 问题清单:chunk 去重 + hybrid 之后,还剩什么
可以分成四层(本次主要收口前两层):
### 3.1 正确性 / 契约(本次主战场)
1. hybrid **没有**独立的融合分归一化,只有 L2 兼容补丁
2. 后处理关键词打分不合理,会抬升词面热、语义冷的片段
3. `scoreLabel` 语义混乱:`l2_distance` / `rrf_fused` / `bm25_only_*` 混用
4. `mode=dense` 与 hybrid 内部 dense 子路概念易混(mode 是整次查询算法,不是「第三套库」)
### 3.2 质量上限(未在本 change 做完)
- 无固定 RAG 评测报表驱动阈值标定
- 无邻块扩展、真 query rewrite、cross-encoder 精排
- 中文 analyzer / 分词策略未产品化钉死
### 3.3 工程债(部分清理、部分保留)
- 应用层 `LexicalRanker` / 自研 `RrfFusion` 可能仍像「还有 app-layer hybrid」
- Spring AI starter 仍可作 sidecar,但不是知识主路径(starter 至 2.0.0 仍无 BM25 hybrid)
- 写入先删后插,非强 upsert;Milvus / MySQL / L0 三方一致性靠流程
### 3.4 运维边界
- hybrid 依赖 BM25 Function + sparse index
- 全量 rebuild 受 embedding 与写入延迟约束
- `totalVectors` 一类统计可能仍不可信
本次 change(`rag-quality-score-unify`)**有意只收口 3.1**:
让 hybrid 的排序权威与质量闸门重新对齐,而不是同时上精排模型。
---
## 4. 关键澄清:L2 是什么,它是不是「后处理」本身
讨论中容易把「L2」和「后处理打分流程」混成一个词。需要拆开:
### 4.1 L2 是度量
**L2 = 欧氏距离**,dense ANN 常用 metric:
- 越小越相似
- 单位向量场景下可用 `maxL2Distance≈2` 做上界
- 归一化相似度:`1 - clamp(L2) / maxL2`
### 4.2 后处理是流水线
后处理消费的是「约定好的 score」,历史上**假定**它是 L2,于是:
```text
score(L2) → normalizeL2 → baseScore → (+boost) → finalScore → 排序/等级
```
```mermaid
flowchart LR
subgraph metric["度量层"]
L2[L2 距离]
end
subgraph pipe["后处理流水线"]
N[normalize]
B[可选 boost]
S[排序 / 截断]
G[等级 / 闸门]
N --> B --> S --> G
end
L2 -.->|历史上假定输入是 L2| N
RRF[RRF 融合分] -.->|量纲不同 · 不能直接套| N
```
所以:
- L2 ≠ 后处理
- L2 = dense 路径的自然距离
- 后处理 = 把某种 score 变成 quality / 等级 / 截断结果的流程
hybrid 的问题是:**流程还在,输入契约已经不再总是 L2。**
### 4.3 category 降级也不是 L2 存在的唯一理由
filtered → unfiltered retry 用的是:
```text
isLowQuality = 无可用证据 或 topSimilarity < referenceThreshold
```
`topSimilarity` 来自归一化后的质量分。
unfiltered 只是**再检一次**,尺子本来就该是统一 quality,而不是「专为降级准备的 L2」。
```mermaid
flowchart TB
F[带 category 的检索] --> Q{isLowQuality?}
Q -->|是| U[unfiltered retry]
Q -->|否| K[采用本次结果]
U --> M[合并/替换为 retry 结果]
K --> OUT[后处理出口]
M --> OUT
```
---
## 5. 设计拍板:统一成什么
### 5.1 两层概念,不要混
| 层 | 只有什么 | 不是什么 |
|----|----------|----------|
| **检索 mode** | `dense` \| `hybrid` | 不是三套库 |
| **一级 scoreLabel** | `dense` \| `hybrid` | 不是 `bm25_only` 第三种模式 |
- `mode`:整次查询怎么跑(配置 `retrieval.search.mode`)
- `scoreLabel`:这条 hit 的 `score` 怎么解释
```mermaid
flowchart TB
CFG[retrieval.search.mode] --> M1[dense 整次只跑 ANN]
CFG --> M2[hybrid 整次 dense+BM25+RRF]
M1 --> L1[scoreLabel=dense]
M2 --> L2[scoreLabel=hybrid]
L2 -.->|不是| L3[bm25_only 第三种 mode]
```
`bm25_only_no_dense` **不是第三种检索**,只是旧链路里「这条 hybrid 命中没有 dense L2 可回填」的补丁标签。统一后应降级为历史别名(canonicalize → `hybrid`),不再一级发射。
### 5.2 检索回来带什么
建议最小契约:
```text
rank (originalRank) // 1 最好;hybrid = RRF 序;dense = ANN 序
score // 引擎主分;量纲由 label 解释
scoreLabel // dense | hybrid
rawScore? // 可选调试
denseDistance? // hybrid 可选:同 id 的 L2,仅供闸门
```
| label | score 含义 |
|-------|------------|
| `dense` | L2 距离(越小越好) |
| `hybrid` | 引擎融合分可放 raw/score;**排序不看其量纲** |
```mermaid
flowchart LR
subgraph emit["Store 发射"]
LAB[scoreLabel<br/>dense | hybrid]
SCR[score / rawScore]
RNK[列表序 → originalRank]
DD[denseDistance? 仅 hybrid]
end
LAB --> NORM
SCR --> NORM
RNK --> SORT
DD --> NORM
NORM[toQualityScore] --> QS[qualityScore]
SORT[按 rank 排序] --> LIST[evidence 顺序]
QS --> GATE[level / isLowQuality]
```
### 5.3 唯一归一化点
```text
qualityScore = toQualityScore(label, score, rank, batchSize, maxL2, denseDistance?)
// 输出统一:[0,1],越大越好
```
分支只允许出现在这里:
```text
dense → 1 - clamp(L2)/maxL2
hybrid → 优先 denseDistance 的 L2 归一化(绝对质量 / 闸门)
无 dense 时 rank 线性回退
```
```mermaid
flowchart TB
IN[label + score + rank + denseDistance?] --> C{canonicalize label}
C -->|dense| L2[l2ToQuality score]
C -->|hybrid| H{denseDistance?}
H -->|有| L2H[l2ToQuality denseDistance]
H -->|无| RK[rankToQuality]
L2 --> OUT[qualityScore 0..1]
L2H --> OUT
RK --> OUT
```
**演进说明:** 切片 1 曾用纯 rank 做 hybrid quality;eval 发现 rank1 恒高会杀死 L0 filter fallback。
现行约定:**排序仍纯 RRF;闸门可用 denseDistance 绝对质量**,且 **不得** 再把主分/label 伪装成 L2。
### 5.4 后处理:统一流程,不要按 label 再分叉业务
```text
candidates
→ 每条 toQualityScore(...) ← 唯一认 label 的地方
→ qualityScore + originalRank
→ 统一:按 rank 排序 / 去重 / 每文档 chunk 上限 / return-n
→ 统一:relevance_level、isLowQuality(只看 qualityScore)
```
```mermaid
flowchart TB
CAND[candidates] --> QS[toQualityScore 每条]
QS --> SORT[sort by originalRank ASC]
SORT --> DEDUP[evidenceKey 去重]
DEDUP --> CAP[max-chunks / return-n]
CAP --> REL[relevance_level]
CAP --> LOW[isLowQuality → filter retry]
CAP --> OUT[EvidenceBlocks]
```
可以记成:
> **Label 只活在进后处理之前的适配器里;后处理是 label-agnostic 的。**
> **排序听 rank;闸门听 qualityScore(hybrid 可含 denseDistance)。**
### 5.5 后处理还改不改?——要改,而且和归一化同一刀
后处理合理职责是 **裁剪与装配**,不是第二套检索:
| 保留 | 去掉或降级 |
|------|------------|
| evidenceKey 去重 | domain/entity/keyword **加分改序** |
| max-chunks-per-document | contains 当相关度代理 |
| return-n | PRECISE 强制 hint support |
| excerpt 截断、EvidenceBlock | |
| 统一 qualityScore 闸门 | |
L0 仍可:
- 导航:category filter(失败 unfiltered retry)
- 解释:`hitReasons` 记 `l0_keyword_overlap` 等(**零分值**)
词面该不该高:交给 **BM25 子路 + RRF**。
语义该不该近:交给 **dense 子路**(融合时已参与;dense-only mode 对照时单独看)。
### 5.6 mode=dense 还要不要
要,但定位清楚:
| 模式 | 定位 |
|------|------|
| **hybrid** | 线上主路径 / 默认 |
| **dense** | 同库对照、评测、排障——看「去掉 BM25+RRF 后差在哪」 |
注意:
- hybrid **入库**数据完全适用于 dense 查询(每条都写了 `vector`)
- hybrid **内部**仍有 dense 子路——那是融合的一部分,≠ `mode=dense`
- 对照时固定 `retrieve-k` / `return-n` / filter / query 集,只切 mode
- 优先比命中集合与排名;`relevance_level` 在 hybrid 下是序数 quality,慎作跨 mode 绝对值对比
---
## 6. 目标数据流(落地后)
```text
VectorSearchService (mode=dense|hybrid)
-> hits{ originalRank, score, scoreLabel=dense|hybrid, rawScore?, denseDistance? }
-> KnowledgeDocumentRetriever / SearchPort
-> KnowledgeEvidencePostProcessor
qualityScore = RetrievalScoreNormalizer.toQualityScore(...)
sort by originalRank ASC
evidenceKey dedup / max-chunks / return-n
relevance_level & topSimilarity from qualityScore
L0 overlap → hitReasons only
-> ContextPack / Assembler / Projector
```
```mermaid
flowchart TB
Q[query] --> VSS[VectorSearchService]
VSS -->|mode=dense| SD[searchDense]
VSS -->|mode=hybrid| SH[searchHybrid + 可选 denseDistance]
SD --> PORT[KnowledgeSearchPort / Adapter]
SH --> PORT
PORT --> POST[KnowledgeEvidencePostProcessor]
POST --> PACK[ContextPacker]
POST --> ASM[LookupResultAssembler]
ASM --> PROJ[RagResultProjector]
PROJ --> AGENT[Agent 可见契约]
```
与上一代对比:
| 环节 | 上一代 | 现在 |
|------|--------|------|
| hybrid score | 常被 L2 覆盖 | 保持融合侧;label=`hybrid` |
| BM25-only | maxL2 + `bm25_only_*` | 普通 hybrid hit,quality 看 rank |
| 归一化 | 一律当 L2 | 按 label 唯一转换 |
| 排序 | finalScore(含 boost) | originalRank |
| L0 关键词 | +分改序 | 仅解释 |
| 质量闸门 | baseScore(L2 兼容) | qualityScore |
---
## 7. 行为变化:必须说清楚的协议调整
这是**有意的行为变化**(对内检索质量语义;Agent ACI 字段名可不变):
1. hybrid 下证据顺序更贴近 **RRF**,不再被 contains 抬到前面
2. 「词面很准、dense 略远」的命中,不再被默认打成低质占位
3. `relevance_level` / category unfiltered retry 的触发分布可能变化
4. hybrid 的 quality 是**本轮序数分**,跨 query 绝对值不可比;阈值可能需后续标定
5. dense 对照模式:质量仍走 L2 归一化,行为更接近旧 dense 主路径
未改:
- Agent 可见字段结构(evidence 列表、relevance 枚举名等)
- Milvus hybrid schema / 不必为本次 rebuild
- `mode=dense` 开关本身
---
## 8. 和上一篇文章的衔接:阶段进度
对照 `RAG排序-多路召回与RRF.md` 的推进顺序:
| 阶段 | 内容 | 状态(截至 2026-07-28) |
|------|------|-------------------------|
| Phase 0 | chunk 去重、retrieve-k/return-n、身份 | **已落地** |
| Phase 1~2 | 多路 + RRF;真 BM25 hybrid | **已落地**(库内 hybrid,非 app-layer 伪融合) |
| 分数职责分离 | 排序 vs 可用性/质量闸门 | **本次收口**(qualityScore 统一) |
| 去掉 L0 当裁判 | 关键词不改主序 | **本次收口** |
| Phase 3 | 可插拔模型 Rerank | **未做**(候选池与评测闭环仍优先) |
| 邻块 / query rewrite | 上下文与问句改写 | **未做** |
因此,本次文章不是推翻上一篇,而是补上上一篇写到「融合之后」却还没写完的半截:
> 融合解决「谁进来、谁先排」;
> 归一化与后处理决定「算不算够好、会不会被规则再次打乱」。
---
## 9. 实现锚点(便于对照代码)
| 组件 | 职责 |
|------|------|
| `RetrievalScoreLabels` | `dense` / `hybrid` + 旧别名 canonicalize |
| `RetrievalScoreNormalizer` | 唯一 `toQualityScore` |
| `MilvusHybridKnowledgeStore` | 发射 label;hybrid 不 L2 覆盖 |
| `VectorSearchService` | mode 路由;SearchResult 契约注释 |
| `KnowledgeEvidencePostProcessor` | rank 保序、去 boost 改序、quality 闸门 |
| 架构 §6.0 | mode 用途:hybrid 主路径 / dense 对照 |
验证(单测,非 live E2E):
- normalizer:L2 边界、rank 单调、别名
- post-process:rank 不被 keyword 打乱;caps/return-n
- lookup tool:保序;context pack 元数据仍在
已知未验证:真实 Milvus 联调对照、阈值标定。
---
## 10. 实践清单:以后别再踩的坑
1. **不要**为了复用旧 `normalizeL2`,把 hybrid 结果伪装成 L2。
2. **不要**在已经 BM25 hybrid 之后,再用 L0 contains 大额加分改主序。
3. **不要**把 `bm25_only` 当成第三种检索模式。
4. **不要**把 hybrid 内部的 dense 子路,和 `mode=dense` 整次查询混为一谈。
5. **要**让 label 差异停在适配器;后处理只认 qualityScore + rank。
6. **要**用 dense mode 做召回对照,而不是第二套长期线上策略。
7. **要**接受:hybrid 序数 quality 与绝对阈值之间,需要观测后再调,而不是再发明一层伪装。
8. **下一步再考虑**精排模型——在契约掰直、有固定 query 回归集之后。
---
## 11. 结语
混合检索落地,解决的是「漏」和「跨路硬加分」里很大一块。
但若后处理仍活在 L2 + 关键词 boost 的旧世界,hybrid 买到的 RRF 序和质量信号会被悄悄改写,甚至惩罚「只在 BM25 路很强」的好证据。
这次收口的核心就三句:
```text
1. 一级 label 只有 dense / hybrid
2. 唯一 toQualityScore;后处理统一、保 rank
3. L0 关键词可以解释,不可以再当排序裁判
```
它不是 RAG 的终点,而是 hybrid 从「能跑」变成「分数语义自洽」的必要一步。
在此之后,评测闭环、阈值标定、邻块与精排,才有干净的基线可谈。
---
## 附录 A:术语
| 术语 | 含义 |
|------|------|
| L2 | 欧氏距离;dense ANN 常用;越小越相似 |
| RRF | Reciprocal Rank Fusion;用名次融合多路,不融合原始分 |
| scoreLabel | 一级分数语义:`dense` \| `hybrid` |
| qualityScore | 归一化后的 0~1 质量分(越大越好),供等级与低质闸门 |
| originalRank | 检索返回名次;后处理排序权威 |
| mode=dense | 整次只跑 dense ANN(对照) |
| mode=hybrid | dense+BM25+RRF(主路径) |
| L0 | query hint / 可选 category filter;不作事实证据、不改主序 |
## 附录 B:相关材料
| 材料 | 路径 |
|------|------|
| 多路与 RRF 讨论 | `docs/RAG排序-多路召回与RRF.md` |
| 当前架构 | `mvp/architecture/RAG知识检索架构.md` |
| OpenSpec 归档 | `openspec/changes/archive/2026-07-28-rag-quality-score-unify/` |
| 主规格 | `openspec/specs/rag-retrieval-quality-score/spec.md` |
| devflow | `devflow/projects/2026-07-28-rag-quality-score-unify/` |

Some files were not shown because too many files have changed in this diff Show More