Compare commits
10
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
e47f2dead0 | ||
|
|
49180abccf | ||
|
|
529f4ff43b | ||
|
|
75fa154a0a | ||
|
|
8300435a63 | ||
|
|
e20249c5d9 | ||
|
|
8fbc443f76 | ||
|
|
8ee7cc0b70 | ||
|
|
bc36248cd8 | ||
|
|
f8809cb7dd |
@@ -0,0 +1,259 @@
|
||||
---
|
||||
name: essence
|
||||
description: Invoke when a project is too large or you only want the core design insights. Extracts 1-2 standout design patterns with deep analysis, lens-guided perspectives, and migration examples. Not for full project analysis or quick lookups.
|
||||
metadata:
|
||||
version: "0.5.0"
|
||||
---
|
||||
|
||||
# Essence: Extract Core Design Patterns
|
||||
|
||||
Prefix your first line with 🥷 inline, not as its own paragraph.
|
||||
|
||||
You are a jewel inspector. A project has thousands of files — your job is to find the one or two brilliant ideas worth stealing.
|
||||
|
||||
**This is NOT a lite version of `/explore`.** `/explore` reads the whole project and summarizes at the end. `/essence` goes deep on one thing and ignores everything else.
|
||||
|
||||
## Mode Selection
|
||||
|
||||
First, check whether an `/explore` result exists:
|
||||
|
||||
- `/explore` report exists → it already identified 2-3 core designs, default to **User-directed**. Ask the user which design to deep-dive, or whether to switch mode.
|
||||
- No `/explore` result → this is an independent launch, default to **Auto-detect**.
|
||||
|
||||
Always confirm before proceeding:
|
||||
|
||||
| Mode | When | Entry |
|
||||
|---|---|---|
|
||||
| **User-directed** | Already have a design target from `/explore`, or know exactly which design to investigate | User tells you what to look for |
|
||||
| **Auto-detect** | Independent launch, project is large, want the AI to find the standout design | You find the standout design |
|
||||
| **Lens-guided** | "Analyze this from a [mechanical/intentional/evolution] perspective" | Apply a specific analytical lens |
|
||||
|
||||
### Lens definitions
|
||||
|
||||
| Lens | Core question | Guided behavior |
|
||||
|---|---|---|
|
||||
| **Mechanical** (default) | How does it work? | Read source code, trace call chains, examine interfaces |
|
||||
| **Intentional** | Why this way? | Read design docs/RFCs/PRs, extract decision rationale and tradeoffs |
|
||||
| **Evolution** | How did it get here? | Read git history/changelog, compare before/after, identify migration drivers |
|
||||
|
||||
A lens shapes which sources to read and how to frame the output, but does not add separate phases.
|
||||
|
||||
### Auto-detect signals
|
||||
|
||||
A design is "essence" if it passes 2 or more of these signals:
|
||||
|
||||
| Signal | Evidence |
|
||||
|---|---|
|
||||
| README highlights it prominently | "Built on a plugin architecture" as a headline feature |
|
||||
| Has standalone architecture docs | ARCHITECTURE.md, docs/design/, blog post by author |
|
||||
| Heavily discussed in Issues/PRs | Design decisions debated by community |
|
||||
| Unique among similar projects | Competitors don't do it this way |
|
||||
| Rich design comments in code | JSDoc/TSDoc explaining why, not what |
|
||||
| Cross-module contract | A type, interface, or protocol imported across module boundaries (not just files). Go: most-implemented interface. Python: most-subclassed abstract base. Rust: most-implemented trait. These define subsystem relationships. |
|
||||
| File size anomaly | One file is disproportionately large or small for its responsibility — signals non-trivial logic |
|
||||
| Dedicated test coverage | Tests specifically validate this design's behavior, not just happy paths |
|
||||
|
||||
**"Clean code" is NOT a signal.** A well-written utility function is not essence. An architecture decision that shapes the entire project is.
|
||||
|
||||
If no design passes 2+ signals, tell the user: "This project has no standout design. Try `/explore` for a full analysis instead."
|
||||
|
||||
## Phase 1: Locate
|
||||
|
||||
**User-directed mode:**
|
||||
- Go directly to the directory or file the user names.
|
||||
- If the directory doesn't exist, stop and tell the user. Do NOT invent an alternative.
|
||||
|
||||
**Auto-detect mode:**
|
||||
- Scan README, AGENTS.md, and top-level docs for architecture claims.
|
||||
- Identify 1-2 standout design directions.
|
||||
- Present to the user: "The standout designs appear to be: A) {design A}, B) {design B}. Which should we dive into?"
|
||||
- If user doesn't choose, pick the strongest one and state why.
|
||||
|
||||
**Lens-guided mode:**
|
||||
- Confirm the lens with the user (Mechanical/Intentional/Evolution).
|
||||
- Frame the search in terms of the lens.
|
||||
- Example: "You want the Mechanical view — I'll trace the core implementation and extract the pattern."
|
||||
|
||||
**Output:** 1-2 design directions to analyze + lens confirmation.
|
||||
|
||||
**Stall signal:** Cannot identify any standout design → the project may be a conventional CRUD app or wrapper. Stop and recommend `/explore` or a different project.
|
||||
|
||||
## Phase 2: Deep Dive
|
||||
|
||||
Read the core files related to the chosen design. Maximum 10 files. Let the lens guide source selection: Mechanical → source code and type definitions; Intentional → design docs, RFCs, PR discussions; Evolution → git history, changelog, migration guides.
|
||||
|
||||
**For each file:**
|
||||
- What role does it play in this design?
|
||||
- What interfaces does it expose?
|
||||
- How does it connect to other parts of the system?
|
||||
|
||||
**Trace the call chain:**
|
||||
- Start from the entry point that uses this design.
|
||||
- Follow the flow until you understand the full pattern.
|
||||
- Stop when you hit boilerplate, config, or test files.
|
||||
|
||||
**Output:** Core file list (≤10) + call chain + lens-specific annotations.
|
||||
|
||||
**Stall signal:** The design spans more than 10 files and you can't find the boundary → the design is probably the project's core architecture. Switch to `/explore` for a full analysis instead.
|
||||
|
||||
## Phase 3: Extract Pattern
|
||||
|
||||
Analyze the design at a higher level. Let the lens shape the analysis angle:
|
||||
- **Mechanical** → emphasize structure, interfaces, data flow — produce a pattern diagram + interface contracts
|
||||
- **Intentional** → emphasize decision rationale, tradeoffs — produce a decision record (context → options → rationale)
|
||||
- **Evolution** → emphasize before/after comparison, migration drivers — produce a timeline + catalyst events
|
||||
|
||||
**Universal analysis dimensions** (all lenses):
|
||||
|
||||
- **Problem:** What specific problem does this design solve? What was the pain before?
|
||||
- **Pattern:** What's the name of this pattern? (Named: MVC, Observer, Plugin, Middleware. Custom: describe it in one sentence.)
|
||||
- **Alternatives:** What simpler or more complex approaches could solve the same problem?
|
||||
- **Tradeoffs:** Why did the author choose this? What does it give up?
|
||||
- **Evidence:** What in the code proves this analysis is correct? (Specific files, functions, comments.)
|
||||
|
||||
**Output:** Design pattern card (lens-framed).
|
||||
|
||||
**Stall signal:** Cannot explain why the author chose this design over alternatives → read commit messages and PR discussions for design rationale. If unavailable, state "author's reasoning unknown" in the report.
|
||||
|
||||
## Phase 4: Migrate
|
||||
|
||||
Make the learning actionable. Let the lens tailor the output:
|
||||
- **Mechanical** → copy-paste code skeleton (≤20 lines with TODOs)
|
||||
- **Intentional** → decision framework (checklist for evaluating tradeoffs)
|
||||
- **Evolution** → migration path (step-by-step refactor plan)
|
||||
|
||||
**Universal deliverables** (all lenses):
|
||||
|
||||
- **Can you use this?** Is the design applicable to the user's own projects? If not, why?
|
||||
- **Steal-it example:** A simplified version (under 20 lines) that captures the core idea. Not production code — a teaching example.
|
||||
- **Pitfalls:** What context does this design depend on? What would break if you copy it blindly?
|
||||
|
||||
**Output:** Migration example + pitfall list (lens-tailored).
|
||||
|
||||
**Stall signal:** The design depends on framework internals, language features, or ecosystem the user doesn't have → explain the core idea abstractly instead of providing code.
|
||||
|
||||
## Phase 5: Self-review
|
||||
|
||||
Check the report is honest:
|
||||
|
||||
**All modes:**
|
||||
- [ ] The design is real (not inferred, not imagined). Evidence: specific files cited.
|
||||
- [ ] The analysis is deep enough that you could explain it out loud.
|
||||
- [ ] The migration example captures the core idea, not surface syntax.
|
||||
- [ ] Pitfalls are specific, not vague ("needs X version" not "may not work everywhere").
|
||||
|
||||
**Stall signals (any one → return to relevant phase):**
|
||||
- Cannot name a file that proves the pattern → back to Phase 2
|
||||
- Cannot explain why it's better than alternatives → back to Phase 3
|
||||
- Migration example is over 20 lines → simplify, back to Phase 4
|
||||
- Lens-specific check failed (e.g., Mechanical missing end-to-end call chain, Intentional missing decision rationale, Evolution missing timeline) → back to relevant phase
|
||||
|
||||
**Output:** Essence report with lens annotation.
|
||||
|
||||
## Optional: HTML Card
|
||||
|
||||
**Only when the user explicitly requests it.**
|
||||
|
||||
Generate an HTML visualization card as a shareable deliverable.
|
||||
|
||||
### HTML Card Structure (Glassmorphism 2.0 - Essence Variant)
|
||||
|
||||
```html
|
||||
<!DOCTYPE html>
|
||||
<html lang="zh-CN">
|
||||
<head>
|
||||
<meta charset="UTF-8">
|
||||
<title>{Project Name} - Essence Report</title>
|
||||
<script src="https://cdn.tailwindcss.com"></script>
|
||||
<script src="https://cdn.jsdelivr.net/npm/mermaid/dist/mermaid.min.js"></script>
|
||||
<style>
|
||||
/* Same glassmorphism styles as /explore */
|
||||
:root { --glass-bg: rgba(255,255,255,0.4); --primary: #8b5cf6; }
|
||||
[data-theme="dark"] { --glass-bg: rgba(15,23,42,0.6); --primary: #a78bfa; }
|
||||
.glass-panel { backdrop-filter: blur(12px); border-radius: 1rem; }
|
||||
.pattern-diagram { font-family: monospace; background: rgba(0,0,0,0.03); }
|
||||
</style>
|
||||
</head>
|
||||
<body class="p-8">
|
||||
<nav class="fixed top-4 left-1/2 -translate-x-1/2 w-[90%] max-w-4xl glass-panel z-50 px-6 py-3">
|
||||
<span class="font-bold text-xl">💎 {Project Name} 精华</span>
|
||||
<span class="text-sm opacity-70">Lens: {lens} | Pattern: {pattern_name}</span>
|
||||
</nav>
|
||||
|
||||
<main class="max-w-4xl mx-auto mt-24 space-y-6">
|
||||
<section class="glass-panel p-6">
|
||||
<h2 class="text-xl font-bold mb-4">🎯 Design Analyzed</h2>
|
||||
<p>{one-line description}</p>
|
||||
</section>
|
||||
|
||||
<section class="glass-panel p-6">
|
||||
<h2 class="text-xl font-bold mb-4">🔷 Pattern ({lens})</h2>
|
||||
<!-- Lens-framed pattern card -->
|
||||
</section>
|
||||
|
||||
<section class="glass-panel p-6">
|
||||
<h2 class="text-xl font-bold mb-4">🔗 Call Chain</h2>
|
||||
<pre class="mermaid">{diagram}</pre>
|
||||
</section>
|
||||
|
||||
<section class="glass-panel p-6">
|
||||
<h2 class="text-xl font-bold mb-4">📦 Migration Example</h2>
|
||||
<pre class="pattern-diagram"><code>{code_example}</code></pre>
|
||||
<p class="text-sm opacity-70 mt-2">Pitfalls: {pitfalls}</p>
|
||||
</section>
|
||||
</main>
|
||||
|
||||
<script>mermaid.initialize({ startOnLoad: true });</script>
|
||||
</body>
|
||||
</html>
|
||||
```
|
||||
|
||||
### Output Format
|
||||
|
||||
```markdown
|
||||
### HTML Card Generated
|
||||
|
||||
- **Path:** `outputs/{project}-essence.html`
|
||||
- **Theme:** {modern/ink}
|
||||
- **Accent Color:** Purple (essence = jewel)
|
||||
```
|
||||
|
||||
**When to skip:** Skip HTML generation unless the user requests it or the analysis is production-critical. When HTML generation fails, deliver a plain-text report instead.
|
||||
|
||||
---
|
||||
|
||||
## Hard Rules
|
||||
|
||||
- **No code evidence = no conclusion.** Every claim about a design must cite a specific file, function, or comment.
|
||||
- **Under 20 lines for migration examples.** If you can't explain the idea in 20 lines, you don't understand it well enough.
|
||||
- **Stop after the report.** Do not modify the user's project or the target project.
|
||||
- **HTML is optional.** Do not block analysis on HTML generation.
|
||||
|
||||
## Gotchas
|
||||
|
||||
| What happened | Rule |
|
||||
|---|---|
|
||||
| 提取的"精华"是 AI 脑补的 | 必须有代码证据(文件 + 行号),不写空泛结论 |
|
||||
| 用户指定方向但该模块不存在 | 停止并告知用户,不编造替代方向 |
|
||||
| 项目没有 standout 设计(胶水代码) | 标记"无可提取精华",建议改用 `/explore` |
|
||||
| Phase 4 迁移示例超过 20 行 | 简化到核心思路,不是复制生产代码 |
|
||||
| 分析了一个小工具函数 | 工具函数不是设计。设计影响整个架构,工具只解决一个问题 |
|
||||
| 从 commit message 推断作者意图但没有代码佐证 | Commit message 是辅助证据,必须有代码结构本身的支持 |
|
||||
| 透镜模式选错导致输出不符预期 | Phase 1 先确认透镜,Mechanical 读代码、Intentional 读文档、Evolution 读历史 |
|
||||
| 透镜分析流于表面 | 每个透镜有特定输出格式:Mechanical→图 + 接口,Intentional→决策记录,Evolution→时间线 |
|
||||
| HTML 卡片生成失败 | 降级到纯文本报告,不阻塞分析交付 |
|
||||
|
||||
## Outcome
|
||||
|
||||
```
|
||||
Essence Report: {project name}
|
||||
Lens: mechanical / intentional / evolution
|
||||
Design analyzed: {one-line description}
|
||||
Files examined: {count}
|
||||
Pattern: {pattern name or custom description}
|
||||
Migration: {steal-it example, ≤20 lines}
|
||||
HTML generated: yes / no
|
||||
Status: complete
|
||||
```
|
||||
|
||||
After the report, stop. No modifications. No follow-ups.
|
||||
@@ -0,0 +1,79 @@
|
||||
# Essence Detection Signals
|
||||
|
||||
How to identify the standout design in a project when the user doesn't specify a direction.
|
||||
|
||||
## Signal Strength
|
||||
|
||||
A design passes the "essence" threshold if it scores 2+ signals.
|
||||
|
||||
### Strong Signals (score = 1 each)
|
||||
|
||||
| Signal | How to detect | Example |
|
||||
|---|---|---|
|
||||
| **README headline** | Project name is followed by a design claim | "Vite — Next generation frontend tooling with **ESM-first architecture**" |
|
||||
| **Architecture docs** | Standalone design document exists | `ARCHITECTURE.md`, `docs/design/`, `docs/architecture/` |
|
||||
| **Official blog post** | Author wrote about the design on their blog | tw93.fun, Vite blog, React blog posts |
|
||||
| **Community discussion** | Issues/PRs debate the design decision | "Why we chose X over Y" discussions with many comments |
|
||||
| **Rich code comments** | JSDoc/TSDoc explaining WHY, not WHAT | "We use this pattern because..." with detailed reasoning |
|
||||
|
||||
### Objective Signals (score = 1 each, no subjective judgment needed)
|
||||
|
||||
| Signal | How to detect | Example |
|
||||
|---|---|---|
|
||||
| **Cross-module contract** | A type, interface, or protocol imported across module boundaries (not just files). Go: most-implemented interface. Python: most-subclassed abstract base. Rust: most-implemented trait. | `Plugin` interface implemented by 8 subsystems, each in its own package |
|
||||
| **File size anomaly** | One file's line count is ≥3× the median for its category (handlers, utils, etc.) | Average handler: 50 lines. One handler: 800 lines with state machine logic |
|
||||
| **Dedicated test coverage** | Tests exist specifically for this design's edge cases, not just happy paths | `plugin.test.ts` tests plugin resolution, fallback, lifecycle — not just "it loads" |
|
||||
|
||||
### Weak Signals (score = 0.5 each)
|
||||
|
||||
| Signal | How to detect | Example |
|
||||
|---|---|---|
|
||||
| **Unique among competitors** | Same category, different architecture | Next.js uses SSR, Remix uses nested routes — that difference IS the essence |
|
||||
| **Most-starred files** | GitHub shows stars/bookmarks on specific files | "This file has 200+ stars on GitHub" |
|
||||
| **Core algorithm** | One file contains non-trivial logic that drives the project | Diff algorithm, compiler pass, state machine |
|
||||
| **API design** | The public API is notably elegant or unusual | `create()` returns a builder chain, not an object |
|
||||
|
||||
## Not Signals
|
||||
|
||||
These do NOT count as essence:
|
||||
|
||||
- "Clean code" or "well organized" — that's quality, not design
|
||||
- "Uses TypeScript" — that's a language choice, not architecture
|
||||
- "Has good tests" — that's engineering discipline, not design
|
||||
- "Many stars on the repo" — popularity ≠ design quality
|
||||
- "Uses the latest framework" — following trends ≠ standing out
|
||||
- Utility functions — even well-written ones are tools, not designs
|
||||
|
||||
## Auto-detect Procedure
|
||||
|
||||
When the user says "find the essence":
|
||||
|
||||
1. **Read README fully.** What is the #1 feature the author leads with? That's a candidate.
|
||||
2. **Check for design docs.** Is there `ARCHITECTURE.md` or equivalent? That's a candidate.
|
||||
3. **Scan the import graph.** Which file is imported by the most other files? Use `grep -r "import.*from" src/ | sort | uniq -c | sort -rn` or equivalent. The top result is likely the core.
|
||||
4. **Check file sizes.** Are any files disproportionately large or small for their apparent role? That signals hidden complexity.
|
||||
5. **Check uniqueness.** Compare with 1-2 well-known alternatives. What does this project do differently?
|
||||
6. **Present 1-2 candidates** to the user with evidence. Let them choose or auto-select the strongest.
|
||||
|
||||
### Example Output Format
|
||||
|
||||
```
|
||||
Standout designs in {project}:
|
||||
|
||||
A) {Design A name} — evidenced by {README claim / file / doc}
|
||||
What it does: {one sentence}
|
||||
|
||||
B) {Design B name} — evidenced by {code comment / unique feature / community discussion}
|
||||
What it does: {one sentence}
|
||||
|
||||
Which should we dive into? (or I can pick the strongest)
|
||||
```
|
||||
|
||||
## Failure Modes
|
||||
|
||||
| Situation | Response |
|
||||
|---|---|
|
||||
| No signal passes 2+ threshold | "This project uses conventional architecture. Try `/explore` for a full analysis, or pick a more architecturally interesting project." |
|
||||
| User-specified module doesn't exist | Stop. Do NOT suggest an alternative. Tell the user the path doesn't exist. |
|
||||
| Project is a wrapper (thin layer over another tool) | "This project is primarily a wrapper around {X}. The design is in {X}, not here. Try analyzing {X} instead." |
|
||||
| Project is configuration-only (just JSON/YAML files) | "This project has no code architecture. It's configuration-driven. Try `/explore` for a full overview instead." |
|
||||
@@ -0,0 +1,87 @@
|
||||
---
|
||||
name: explore
|
||||
description: Invoke when you need project-level understanding and an onboarding path. Produces a project learning report for code and non-code repositories with fixed phases for positioning, structure, flow, start path, and core designs. Not for deep code extraction or interactive teaching.
|
||||
metadata:
|
||||
version: "0.5.0"
|
||||
---
|
||||
|
||||
# Explore: Project Understanding and Onboarding
|
||||
|
||||
Prefix your first line with 🥷 inline, not as its own paragraph.
|
||||
|
||||
You are a project cartographer. Your job is to help the user understand what a project is, why it is worth studying, how it is organized, and where to start.
|
||||
|
||||
`/explore` is the entry point for first contact with a repository or project-like artifact. It builds global understanding. It does not perform code-level essence extraction and it does not run interactive teaching.
|
||||
|
||||
## Project Type Detection
|
||||
|
||||
After the initial scan, classify the target before continuing:
|
||||
|
||||
| Type | Signals | What changes |
|
||||
|---|---|---|
|
||||
| **Code repository** | `go.mod`, `pyproject.toml`, `Cargo.toml`, source directories, executable entrypoints | Run all 4 phases |
|
||||
| **Skill / docs / knowledge repository** | `SKILL.md`, mostly Markdown, docs-first structure, no runnable application entrypoint | Skip Phase 2 (Flow) and Phase 3 (Start Path) |
|
||||
| **Template / scaffold repository** | Starter files, minimal logic, setup-first repo | Phase 2 may stay structural and Phase 3 may be minimal |
|
||||
|
||||
State the detected type before proceeding. If uncertain, say what evidence is missing and continue with the closest matching type.
|
||||
|
||||
## Phase 1: Positioning & Structure
|
||||
- What this project is, why it is worth studying, and who it is for.
|
||||
- Top-level structure: main modules, documents, directories, and the likely learning entry area.
|
||||
- Tradeoffs vs alternatives when evidence exists.
|
||||
|
||||
## Phase 2: Flow
|
||||
**Code repositories only.**
|
||||
- Skip for non-code and template repositories.
|
||||
- Trace the main runtime or request flow.
|
||||
- Produce at least one architecture or core-flow diagram.
|
||||
- Keep the trace focused on the golden path rather than exhaustive coverage.
|
||||
|
||||
## Phase 3: Start Path
|
||||
**Code repositories only when runnable or meaningfully inspectable.**
|
||||
- Provide the minimal path to start learning or running the project.
|
||||
- Give the first command or first inspection step.
|
||||
- Suggest one safe first modification or observation point when appropriate.
|
||||
|
||||
## Phase 4: Core Designs
|
||||
- Summarize 2-3 core implementations or ideas.
|
||||
- Keep this at overview depth.
|
||||
- For each item, include what it is, where it lives, and why it matters.
|
||||
|
||||
## Minimum Deliverables
|
||||
|
||||
The final `/explore` report must include:
|
||||
- Project positioning
|
||||
- Why it is worth studying
|
||||
- 2-3 core implementations or core ideas
|
||||
- Tradeoffs or comparisons when applicable
|
||||
- At least 1 diagram:
|
||||
- code repository → architecture diagram or core flow diagram
|
||||
- non-code repository → structure diagram, idea map, or workflow diagram
|
||||
|
||||
## Boundary Rules
|
||||
|
||||
`/explore` may:
|
||||
- scan structure
|
||||
- explain the main flow
|
||||
- provide a minimal start path
|
||||
- summarize 2-3 core designs
|
||||
|
||||
`/explore` must not:
|
||||
- perform `/essence`-level deep extraction
|
||||
- act as `/follow`-style guided teaching
|
||||
- include Verify, Deep Fission, or HTML Output phases
|
||||
- preserve no retired lightweight fallback behavior
|
||||
|
||||
## Outcome
|
||||
|
||||
```
|
||||
Explore Report: {project name}
|
||||
Project type: code / skill-docs / template
|
||||
Phases completed: 4/4 (or note skipped code-only phases)
|
||||
Diagram included: yes / no
|
||||
Core designs: 2-3
|
||||
Status: complete
|
||||
```
|
||||
|
||||
After the report, stop. Do not proceed to `/essence` or `/follow` automatically.
|
||||
@@ -0,0 +1,98 @@
|
||||
# Project Analysis Methods
|
||||
|
||||
How to read and understand an unfamiliar code project.
|
||||
|
||||
## 1. Identify the Entry Point
|
||||
|
||||
Every project has a door. Find it first.
|
||||
|
||||
### By Language
|
||||
|
||||
| Language | Look for |
|
||||
|---|---|
|
||||
| **JavaScript/TypeScript** | `package.json` → `main` / `bin` / `scripts.dev` |
|
||||
| **Python** | `setup.py` → `entry_points`, `pyproject.toml` → `[project.scripts]`, or top-level `app.py` / `main.py` / `__main__.py` |
|
||||
| **Go** | `package main` in any file, conventionally `main.go` or `cmd/*/main.go` |
|
||||
| **Rust** | `src/main.rs` or `src/bin/*.rs` |
|
||||
| **Java** | Class with `public static void main(String[] args)` |
|
||||
| **C/C++** | `main()` function, conventionally in `src/main.c` |
|
||||
| **Swift** | `main.swift` or file with `@main` attribute |
|
||||
|
||||
### In Frameworks
|
||||
|
||||
| Framework | Entry point |
|
||||
|---|---|
|
||||
| Next.js | `app/` or `pages/` directory, `next.config.js` |
|
||||
| React (Vite) | `src/main.tsx` or `src/main.jsx` |
|
||||
| Vue (Vite) | `src/main.ts` or `src/main.js` |
|
||||
| Express | File that calls `app.listen()` |
|
||||
| FastAPI | File that creates `FastAPI()` instance |
|
||||
| Django | `manage.py`, then project name directory with `urls.py` / `wsgi.py` |
|
||||
| Flask | `app.py` or `app/__init__.py` |
|
||||
| Spring Boot | `*Application.java` with `@SpringBootApplication` |
|
||||
|
||||
## 2. Judge Project Complexity
|
||||
|
||||
Don't over-engineer simple projects. Don't under-analyze complex ones.
|
||||
|
||||
### Simple (<50 files, single language)
|
||||
- Read every source file.
|
||||
- No need for flow diagrams beyond a simple sequence.
|
||||
- A light `/explore` pass is probably enough.
|
||||
|
||||
### Standard (50-500 files, 1-2 languages)
|
||||
- Read entry point + core modules + 1-2 feature files.
|
||||
- Build 1-2 flow diagrams.
|
||||
- `/explore` is the right level.
|
||||
|
||||
### Complex (>500 files, multi-language, monorepo)
|
||||
- Read entry point + architecture docs + one representative module.
|
||||
- Use `/essence` to find standout designs, or `/explore` for one package at a time.
|
||||
- Do NOT try to understand the whole project in one pass.
|
||||
|
||||
## 3. Separate Core Code from Scaffolding
|
||||
|
||||
Not all files are worth reading.
|
||||
|
||||
### Ignore (scaffolding)
|
||||
- `*.config.js`, `*.config.ts` — configuration, not logic
|
||||
- `dist/`, `build/`, `out/` — generated output
|
||||
- `node_modules/`, `vendor/`, `.venv/` — dependencies
|
||||
- `*.lock`, `yarn.lock`, `go.sum` — lock files
|
||||
- `LICENSE`, `CODEOWNERS`, `.editorconfig` — project meta
|
||||
- `test/fixtures/`, `test/data/` — test data
|
||||
|
||||
### Read (core)
|
||||
- Entry point file
|
||||
- Router/middleware/config handlers
|
||||
- Model/entity/schema definitions
|
||||
- Core algorithm or business logic files
|
||||
- Files referenced most in imports
|
||||
|
||||
### Hint: Follow imports
|
||||
|
||||
```
|
||||
entry file → import A → import B → core logic
|
||||
```
|
||||
|
||||
Each import is a dependency. Follow the chain until you hit a file that doesn't import anything else — that's usually the core.
|
||||
|
||||
## 4. Read Unfamiliar Framework Code
|
||||
|
||||
You don't know every framework. That's fine.
|
||||
|
||||
### Strategy
|
||||
|
||||
1. **Find the routing layer first.** Every framework has a way to map URLs or events to handlers. Find it. It tells you the project's capabilities.
|
||||
|
||||
2. **Follow ONE request end-to-end.** Don't try to understand all routes. Pick the simplest one (often "health check" or "get by ID") and trace it from entry to response.
|
||||
|
||||
3. **Identify the framework's conventions.** Most frameworks follow a pattern:
|
||||
- MVC: Controller → Model → View
|
||||
- Middleware: Request → Middleware chain → Handler → Response
|
||||
- Component: Parent renders children, props flow down, events flow up
|
||||
- Plugin: Core calls hooks, plugins register handlers
|
||||
|
||||
4. **Don't fight the framework's abstraction.** If the project uses ORM, don't look for raw SQL. If it uses dependency injection, don't look for `new()` calls. Understand what abstraction layer they chose.
|
||||
|
||||
5. **Use the framework's own docs.** If stuck on "how does this framework work?", check the official docs. Don't reverse-engineer what's documented.
|
||||
@@ -0,0 +1,173 @@
|
||||
# Flow Pattern Library
|
||||
|
||||
Common architecture patterns and how to identify them in code.
|
||||
|
||||
## MVC / MVVM / MVX
|
||||
|
||||
### What it is
|
||||
Separation of data (Model), UI/presentation (View), and coordination logic (Controller/ViewModel).
|
||||
|
||||
### File signatures
|
||||
| Pattern | Directories/Files |
|
||||
|---|---|
|
||||
| **MVC** | `controllers/`, `models/`, `views/` |
|
||||
| **MVVM** | `viewmodels/`, `views/`, `models/` |
|
||||
| **Layered** | `app/`, `domain/`, `infrastructure/` (Clean/Hexagonal) |
|
||||
|
||||
### Flow
|
||||
```
|
||||
Request → Controller → Model (data) → View (render) → Response
|
||||
```
|
||||
|
||||
### Key question
|
||||
"Does the file handle data, display, or coordination?" If yes → MVC-family.
|
||||
|
||||
---
|
||||
|
||||
## Middleware Chain
|
||||
|
||||
### What it is
|
||||
Each handler processes the request and passes it to the next. Like an assembly line.
|
||||
|
||||
### File signatures
|
||||
| Framework | Indicator |
|
||||
|---|---|---|
|
||||
| **Express/Koa** | `app.use(...)`, `app.get('/', handler)` |
|
||||
| **FastAPI** | `@app.middleware("http")`, `Depends()` |
|
||||
| **Next.js** | `middleware.ts` at root or in `app/` |
|
||||
| **Gin (Go)** | `router.Use(middleware1, middleware2)` |
|
||||
| **Koa** | `app.use(async (ctx, next) => { ... })` |
|
||||
|
||||
### Flow
|
||||
```
|
||||
Request → Middleware A → Middleware B → Handler → Response
|
||||
↓ ↓
|
||||
auth check log request
|
||||
```
|
||||
|
||||
### Key question
|
||||
"Does this function call `next()` or pass control to something else?" If yes → middleware.
|
||||
|
||||
### Common middleware order
|
||||
```
|
||||
1. CORS / Security headers
|
||||
2. Logging / Request ID
|
||||
3. Authentication / Authorization
|
||||
4. Body parsing / Validation
|
||||
5. Rate limiting
|
||||
6. Route handler
|
||||
7. Error handler (catches everything above)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Plugin / Extension System
|
||||
|
||||
### What it is
|
||||
Core provides hooks or interfaces. External code registers handlers. The core doesn't know about specific plugins.
|
||||
|
||||
### File signatures
|
||||
| Pattern | Indicator |
|
||||
|---|---|
|
||||
| **Hook-based** | `registerHook('eventName', handler)`, `hooks.on('event', fn)` |
|
||||
| **Interface-based** | Abstract class or interface that plugins implement |
|
||||
| **Discovery-based** | Directory scan (`plugins/`), import all, register by convention |
|
||||
| **VSCode-style** | `contributes` in `package.json`, activation events |
|
||||
|
||||
### Flow
|
||||
```
|
||||
Core starts
|
||||
↓
|
||||
Scans for plugins
|
||||
↓
|
||||
Each plugin registers itself
|
||||
↓
|
||||
Core fires hooks → plugins respond
|
||||
↓
|
||||
Core runs with extended capabilities
|
||||
```
|
||||
|
||||
### Key question
|
||||
"Can I add functionality without modifying core code?" If yes → plugin architecture.
|
||||
|
||||
---
|
||||
|
||||
## Event-Driven
|
||||
|
||||
### What it is
|
||||
Components communicate through events, not direct calls. Publishers emit, subscribers listen.
|
||||
|
||||
### File signatures
|
||||
| Pattern | Indicator |
|
||||
|---|---|
|
||||
| **Node EventEmitter** | `eventEmitter.on('event', handler)`, `eventEmitter.emit('event', data)` |
|
||||
| **Pub/Sub** | `pubsub.subscribe('channel', handler)`, `pubsub.publish('channel', data)` |
|
||||
| **Redux-style** | `dispatch(action)`, `reducer(state, action) → newState` |
|
||||
| **Observable** | `observable.subscribe(fn)`, `pipe(map, filter)` |
|
||||
| **Signals (Python)** | `@signal.connect`, `signal.send()` |
|
||||
|
||||
### Flow
|
||||
```
|
||||
Component A emits "user.created"
|
||||
↓
|
||||
Listener B hears it → sends welcome email
|
||||
Listener C hears it → creates default settings
|
||||
Listener D hears it → logs analytics
|
||||
```
|
||||
|
||||
### Key question
|
||||
"Does code communicate without importing or calling each other directly?" If yes → event-driven.
|
||||
|
||||
---
|
||||
|
||||
## State Management
|
||||
|
||||
### What it is
|
||||
Centralized storage for application state. Components read and update through defined interfaces.
|
||||
|
||||
### File signatures
|
||||
| Pattern | Indicator |
|
||||
|---|---|
|
||||
| **Redux** | `createStore()`, `dispatch()`, `useSelector()`, `@reduxjs/toolkit` |
|
||||
| **Zustand** | `create((set) => ({ ... }))` |
|
||||
| **Jotai** | `atom(value)`, `useAtom(atom)` |
|
||||
| **MobX** | `@observable`, `@action`, `@computed` |
|
||||
| **React Context** | `createContext()`, `useContext()`, `Provider` |
|
||||
| **Pinia (Vue)** | `defineStore()`, `state`, `actions` |
|
||||
|
||||
### Flow
|
||||
```
|
||||
Component dispatches action
|
||||
↓
|
||||
Reducer processes action + current state
|
||||
↓
|
||||
New state emitted
|
||||
↓
|
||||
Subscribed components re-render
|
||||
```
|
||||
|
||||
### Key question
|
||||
"Where does the app store data that multiple components need?" If it's a single store → state management pattern.
|
||||
|
||||
---
|
||||
|
||||
## Pipeline / Chain of Responsibility
|
||||
|
||||
### What it is
|
||||
Data flows through a series of processors. Each processor transforms the data and passes it on.
|
||||
|
||||
### File signatures
|
||||
| Pattern | Indicator |
|
||||
|---|---|
|
||||
| **Stream processing** | `.pipe(transform1).pipe(transform2)` |
|
||||
| **Compiler/lexer** | Source → Tokenize → Parse → Transform → Generate |
|
||||
| **Data pipeline** | `input → transform → validate → output` |
|
||||
| **Makefile** | Target depends on prerequisites, each is a step |
|
||||
|
||||
### Flow
|
||||
```
|
||||
Raw input → Tokenizer → Parser → Transformer → Generator → Output
|
||||
```
|
||||
|
||||
### Key question
|
||||
"Does data get progressively transformed through a fixed sequence of steps?" If yes → pipeline.
|
||||
@@ -0,0 +1,101 @@
|
||||
---
|
||||
name: follow
|
||||
description: Invoke when the user wants an interactive learning session based on an existing `/explore` or `/essence` report. Guides runnable or reader-style follow-along sessions. Not for fresh project analysis or pattern-only extraction.
|
||||
metadata:
|
||||
version: "0.5.0"
|
||||
---
|
||||
|
||||
# Follow: Guided Learning Session
|
||||
|
||||
Prefix your first line with 🥷 inline, not as its own paragraph.
|
||||
|
||||
You are a guide. The user wants to learn from a project step by step with help, context, and correction. You guide the learning process, but you do not replace it.
|
||||
|
||||
`/follow` is not a fresh project analyzer. It only works from an existing `/explore` or `/essence` result.
|
||||
|
||||
## Pre-check
|
||||
|
||||
`/follow` only works when there is already an `/explore` report or an `/essence` report.
|
||||
|
||||
- `/explore` report exists → use it as the main learning path
|
||||
- `/essence` report exists → use it for design-focused guided study
|
||||
- Neither exists → refuse clearly
|
||||
|
||||
Refusal behavior:
|
||||
"I need an existing `/explore` or `/essence` result before I can guide a follow-along session. Please run `/explore` for project understanding or `/essence` for a focused deep dive first."
|
||||
|
||||
Load the existing report before continuing.
|
||||
|
||||
## Mode Selection
|
||||
|
||||
After the pre-check, select one mode based on the prerequisite report:
|
||||
|
||||
- From `/explore` + code repository → default **Runnable**
|
||||
- From `/explore` + non-code repository → force **Reader**
|
||||
- From `/essence` → default **Reader** (user is in design-analysis state)
|
||||
|
||||
| Mode | When | Entry |
|
||||
|---|---|---|
|
||||
| **Runnable** | Report confirms the project is a runnable code repository and the user wants to learn by running and changing it | Start from environment and first execution |
|
||||
| **Reader** | Project has no runtime, or the user is studying design/architecture, or the prerequisite report is from `/essence` | Start from guided reading |
|
||||
|
||||
State the selected mode before proceeding. Do not re-scan the project — use the prerequisite report to decide.
|
||||
|
||||
## Teaching Interaction Rules
|
||||
|
||||
`/follow` must teach by guidance, not by dumping answers:
|
||||
- explain the purpose of the current step first
|
||||
- give the user an observation point or action point
|
||||
- ask the user to predict, try, or explain before revealing the answer
|
||||
- then reveal, correct, or deepen the explanation
|
||||
- never say "go read the code" as a standalone instruction. When referencing code, always start with: what design idea this code embodies, why it matters in the overall architecture, and what the user should pay attention to
|
||||
|
||||
## Runnable Check
|
||||
|
||||
Before Runnable mode, confirm from the **prerequisite report** (do not re-scan the project):
|
||||
- If the report identified the target as a code repository with a recognized runtime (`go.mod`, `pyproject.toml`, `Cargo.toml`, `Makefile`, `build.gradle`, `pom.xml`, `CMakeLists.txt`, etc.), proceed with Runnable.
|
||||
- If the report classified it as non-code, or no runtime entrypoint was found, switch to Reader and explain why.
|
||||
- If the prerequisite is `/essence`, confirm with the user: essence is design-focused, Reader is the natural fit. Allow Runnable only if the user explicitly insists.
|
||||
- Do not introduce a third mode.
|
||||
|
||||
## Runnable Mode Flow
|
||||
1. Confirm environment and prerequisites.
|
||||
2. Let the user run the project.
|
||||
3. Let the user make one safe change.
|
||||
4. Walk the main flow together.
|
||||
5. Give one small exercise.
|
||||
6. Review what they learned.
|
||||
|
||||
## Reader Mode Flow
|
||||
1. Frame the learning goal around a core design or architectural idea, not a single file.
|
||||
2. Walk through the design concept layer by layer: problem → approach → implementation → tradeoff.
|
||||
3. Ask the user questions that probe understanding ("Why did the author choose this approach over a simpler one?"), not just prediction ("What happens next?").
|
||||
4. Use diagrams or structured summaries to connect the dots between files and design ideas.
|
||||
5. Give one reasoning exercise that tests whether the user can apply the design pattern elsewhere.
|
||||
6. Review what they learned.
|
||||
|
||||
## Boundary Rules
|
||||
|
||||
`/follow` must:
|
||||
- depend on `/explore` or `/essence`
|
||||
- guide the user interactively
|
||||
- adapt between code and non-code repositories through Runnable or Reader emphasis
|
||||
|
||||
`/follow` must not:
|
||||
- rescan the whole project as a new analyzer
|
||||
- reference retired skills as prerequisites
|
||||
- add any third learning mode
|
||||
- execute commands or write code for the user
|
||||
|
||||
## Outcome
|
||||
|
||||
```
|
||||
Follow Session: {project name}
|
||||
Mode: runnable / reader
|
||||
Prerequisite report: /explore or /essence
|
||||
Exercise result: completed / partial / too hard
|
||||
Next direction: {suggested follow-up}
|
||||
Status: complete
|
||||
```
|
||||
|
||||
After the review, stop. Ask whether the user wants another exercise or wants to end the session.
|
||||
@@ -0,0 +1,113 @@
|
||||
# Environment Detection Rules
|
||||
|
||||
How to detect the runtime environment and guide the user through setup in `/follow`.
|
||||
|
||||
## Language Detection from Config
|
||||
|
||||
Check these files in order. The first match is the primary language.
|
||||
|
||||
| Config file | Language | Runtime check | Install command |
|
||||
|---|---|---|---|
|
||||
| `package.json` | JavaScript/TypeScript | `node --version` | nvm or official installer |
|
||||
| `pyproject.toml` | Python | `python --version` | pyenv or python.org |
|
||||
| `go.mod` | Go | `go version` | golang.org/dl |
|
||||
| `Cargo.toml` | Rust | `rustc --version` | rustup |
|
||||
| `pom.xml` | Java | `java -version` | SDKMAN or official |
|
||||
| `build.gradle` / `build.gradle.kts` | Java/Kotlin | `java -version` | SDKMAN |
|
||||
| `Gemfile` | Ruby | `ruby --version` | rvm or rbenv |
|
||||
| `*.csproj` | C#/.NET | `dotnet --version` | .NET SDK |
|
||||
| `CMakeLists.txt` | C/C++ | `gcc --version` or `clang --version` | System package manager |
|
||||
| `swift package.json` | Swift | `swift --version` | Xcode or swift.org |
|
||||
|
||||
## Dependency Installation
|
||||
|
||||
Once language is detected, guide the user:
|
||||
|
||||
### JavaScript/TypeScript
|
||||
```bash
|
||||
# Check which package manager is used
|
||||
if [ -f "yarn.lock" ]; then yarn install
|
||||
elif [ -f "pnpm-lock.yaml" ]; then pnpm install
|
||||
elif [ -f "bun.lockb" ] || [ -f "bun.lock" ]; then bun install
|
||||
else npm install
|
||||
fi
|
||||
```
|
||||
|
||||
### Python
|
||||
```bash
|
||||
# Modern Python projects
|
||||
pip install -e .
|
||||
# Or with requirements
|
||||
pip install -r requirements.txt
|
||||
# Or with poetry
|
||||
poetry install
|
||||
# Or with uv
|
||||
uv pip install -r requirements.txt
|
||||
```
|
||||
|
||||
### Go
|
||||
```bash
|
||||
go mod download
|
||||
```
|
||||
|
||||
### Rust
|
||||
```bash
|
||||
cargo build
|
||||
```
|
||||
|
||||
### Java (Maven)
|
||||
```bash
|
||||
mvn install
|
||||
```
|
||||
|
||||
### Java (Gradle)
|
||||
```bash
|
||||
./gradlew build
|
||||
# or
|
||||
gradle build
|
||||
```
|
||||
|
||||
## Run Command Detection
|
||||
|
||||
How to start the project:
|
||||
|
||||
| Source | Command |
|
||||
|---|---|
|
||||
| `package.json` → `scripts.dev` | `npm run dev` |
|
||||
| `package.json` → `scripts.start` | `npm start` |
|
||||
| `Makefile` → `dev` target | `make dev` |
|
||||
| `Makefile` → `run` target | `make run` |
|
||||
| `pyproject.toml` (Poetry) | `poetry run python main.py` |
|
||||
| `go.mod` → `package main` | `go run main.go` |
|
||||
| `Cargo.toml` → `[[bin]]` | `cargo run` |
|
||||
| `docker-compose.yml` exists | `docker-compose up` |
|
||||
| `Dockerfile` exists, no compose | `docker build -t app . && docker run app` |
|
||||
|
||||
## Common Environment Issues
|
||||
|
||||
| Error | Cause | Fix |
|
||||
|---|---|---|
|
||||
| `command not found: node` | Node.js not installed | Install Node.js (recommend LTS) |
|
||||
| `ModuleNotFoundError` | Python deps not installed | Run `pip install -r requirements.txt` |
|
||||
| `EACCES: permission denied` | Global install without sudo | Use nvm/fnm, or prefix with sudo |
|
||||
| `ENOENT: no such file` | Wrong working directory | `cd` to project root first |
|
||||
| `port already in use` | Another process on same port | Kill the process or use different port |
|
||||
| `go: cannot find main module` | Outside Go module | `cd` to directory with `go.mod` |
|
||||
| `error: could not find Cargo.toml` | Outside Rust project | `cd` to directory with `Cargo.toml` |
|
||||
| `java.lang.UnsupportedClassVersionError` | Wrong Java version | Match JDK version to project requirement |
|
||||
| `npm ERR! code ERESOLVE` | Dependency conflict | Try `npm install --legacy-peer-deps` |
|
||||
|
||||
## Detection Script for /follow
|
||||
|
||||
```bash
|
||||
# Quick environment check
|
||||
echo "=== Environment ==="
|
||||
node --version 2>/dev/null || echo "Node.js: not installed"
|
||||
python --version 2>/dev/null || echo "Python: not installed"
|
||||
go version 2>/dev/null || echo "Go: not installed"
|
||||
rustc --version 2>/dev/null || echo "Rust: not installed"
|
||||
java -version 2>/dev/null || echo "Java: not installed"
|
||||
echo "PWD: $(pwd)"
|
||||
```
|
||||
|
||||
Run this at the start of `/follow` Step 1 to understand what's available.
|
||||
@@ -0,0 +1,57 @@
|
||||
# Frontend Design — Complete Guidance
|
||||
|
||||
This document provides a comprehensive framework for creating visually distinctive, non-templated UI designs. Here's the full breakdown:
|
||||
|
||||
## Foundational Approach
|
||||
|
||||
Act as the design lead for a studio known for unique client identities — the client has already turned down template-like proposals. Every choice about palette, typography, and layout must be specific to the brief, including "one real aesthetic risk you can justify."
|
||||
|
||||
## Grounding in Subject Matter
|
||||
|
||||
If the brief is vague about the product or subject, pin it down yourself: name the subject, its audience, and the page's single job. Draw inspiration from "the subject's own world, its materials, instruments, artifacts, and vernacular." Use any known context about the human's preferences or past designs as hints.
|
||||
|
||||
## Design Principles
|
||||
|
||||
- **Hero as thesis**: Open with "the most characteristic thing in the subject's world" — avoid default choices like a big number with a small label and gradient accent unless truly optimal.
|
||||
- **Typography**: Pair display and body faces deliberately, not from your usual repertoire. Set a clear type scale with intentional weights, widths, and spacing. "Make the type treatment itself a memorable part of the design."
|
||||
- **Structure as information**: Numbering, eyebrows, dividers must encode something true about the content. Question whether numbered markers (01/02/03) actually make sense before using them — only appropriate for real sequences.
|
||||
- **Motion**: Consider where animation serves the subject. "An orchestrated moment usually lands harder than scattered effects." Sometimes less is better to avoid an AI-generated feel.
|
||||
- **Complexity**: Match execution to the vision — maximalist needs elaborate execution, minimal needs precision.
|
||||
- **Content**: Come up with copy if the brief lacks it. Poor copy makes a design feel as templated as poor layout.
|
||||
|
||||
## AI-Generated Design Traps
|
||||
|
||||
Three common AI-default looks to watch for: (1) warm cream background (~#F4F1EA) with serif display and terracotta accent; (2) near-black with bright acid-green or vermilion; (3) broadsheet layout with hairline rules, zero border-radius, and dense columns. "All three are legitimate for some briefs, but they are defaults rather than choices." Where the brief leaves an axis free, don't spend that freedom on a default.
|
||||
|
||||
## Two-Pass Process
|
||||
|
||||
**Pass 1 — Plan**: Create a compact token system:
|
||||
|
||||
1. **Color**: 4–6 named hex values
|
||||
2. **Type**: Characterful display face (used with restraint), complementary body face, utility face for captions/data
|
||||
3. **Layout**: One-sentence prose descriptions + ASCII wireframes
|
||||
4. **Signature**: The single unique element the page will be remembered by
|
||||
|
||||
Review the plan against the brief. If any part reads like what you'd produce for any similar page, revise it. Only then write code.
|
||||
|
||||
**Pass 2 — Build**: Follow the revised plan exactly. Watch for CSS selector specificity conflicts (e.g., `.section` and `.cta` fighting over padding/margins). Do most planning internally, only sharing ideas when confident.
|
||||
|
||||
## Restraint & Self-Critique
|
||||
|
||||
"Spend your boldness in one place" — let the signature element be the one memorable thing; keep everything else quiet. "Not taking a risk can be a risk itself!" Build responsively down to mobile, with visible keyboard focus and reduced motion respected. Critique as you build. Follow Chanel's advice: before finishing, remove one accessory. Jot notes about what you've tried to avoid repeating yourself.
|
||||
|
||||
## Writing in Design
|
||||
|
||||
Words exist to make the design understandable and usable — they're "design material, not decoration." Write from the end user's perspective, naming things by what people control and recognize, never by how the system is built.
|
||||
|
||||
- Use active voice as default
|
||||
- A control should say exactly what happens: "Save changes," not "Submit"
|
||||
- Maintain consistent vocabulary throughout flows (button says "Publish," toast says "Published")
|
||||
- Treat errors as guidance, not mood — explain what went wrong and how to fix it
|
||||
- Empty screens are invitations to act
|
||||
- Keep the register conversational: "plain verbs, sentence case, no filler"
|
||||
- Let each element do exactly one job — "a label labels, an example demonstrates"
|
||||
|
||||
## License
|
||||
|
||||
Apache License 2.0 — see LICENSE.txt
|
||||
@@ -0,0 +1,83 @@
|
||||
---
|
||||
name: gitnexus-cli
|
||||
description: "Use when the user needs to run GitNexus CLI commands like analyze/index a repo, check status, clean the index, generate a wiki, or list indexed repos. Examples: \"Index this repo\", \"Reanalyze the codebase\", \"Generate a wiki\""
|
||||
---
|
||||
|
||||
# GitNexus CLI Commands
|
||||
|
||||
All commands work via `npx` — no global install required.
|
||||
|
||||
## Commands
|
||||
|
||||
### analyze — Build or refresh the index
|
||||
|
||||
```bash
|
||||
npx gitnexus analyze
|
||||
```
|
||||
|
||||
Run from the project root. This parses all source files, builds the knowledge graph, writes it to `.gitnexus/`, and generates AGENTS.md / AGENTS.md context files.
|
||||
|
||||
| Flag | Effect |
|
||||
| -------------- | ---------------------------------------------------------------- |
|
||||
| `--force` | Force full re-index even if up to date |
|
||||
| `--embeddings` | Enable embedding generation for semantic search (off by default) |
|
||||
| `--drop-embeddings` | Drop existing embeddings on rebuild. By default, an `analyze` without `--embeddings` preserves them. |
|
||||
|
||||
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale. In Codex, a PostToolUse hook detects staleness after `git commit` and `git merge` and notifies the agent to run `analyze` — the hook does not run analyze itself, to avoid blocking the agent for up to 120s and risking KuzuDB corruption on timeout.
|
||||
|
||||
### status — Check index freshness
|
||||
|
||||
```bash
|
||||
npx gitnexus status
|
||||
```
|
||||
|
||||
Shows whether the current repo has a GitNexus index, when it was last updated, and symbol/relationship counts. Use this to check if re-indexing is needed.
|
||||
|
||||
### clean — Delete the index
|
||||
|
||||
```bash
|
||||
npx gitnexus clean
|
||||
```
|
||||
|
||||
Deletes the `.gitnexus/` directory and unregisters the repo from the global registry. Use before re-indexing if the index is corrupt or after removing GitNexus from a project.
|
||||
|
||||
| Flag | Effect |
|
||||
| --------- | ------------------------------------------------- |
|
||||
| `--force` | Skip confirmation prompt |
|
||||
| `--all` | Clean all indexed repos, not just the current one |
|
||||
|
||||
### wiki — Generate documentation from the graph
|
||||
|
||||
```bash
|
||||
npx gitnexus wiki
|
||||
```
|
||||
|
||||
Generates repository documentation from the knowledge graph using an LLM. Requires an API key (saved to `~/.gitnexus/config.json` on first use).
|
||||
|
||||
| Flag | Effect |
|
||||
| ------------------- | ----------------------------------------- |
|
||||
| `--force` | Force full regeneration |
|
||||
| `--model <model>` | LLM model (default: minimax/minimax-m2.5) |
|
||||
| `--base-url <url>` | LLM API base URL |
|
||||
| `--api-key <key>` | LLM API key |
|
||||
| `--concurrency <n>` | Parallel LLM calls (default: 3) |
|
||||
| `--gist` | Publish wiki as a public GitHub Gist |
|
||||
|
||||
### list — Show all indexed repos
|
||||
|
||||
```bash
|
||||
npx gitnexus list
|
||||
```
|
||||
|
||||
Lists all repositories registered in `~/.gitnexus/registry.json`. The MCP `list_repos` tool provides the same information.
|
||||
|
||||
## After Indexing
|
||||
|
||||
1. **Read `gitnexus://repo/{name}/context`** to verify the index loaded
|
||||
2. Use the other GitNexus skills (`exploring`, `debugging`, `impact-analysis`, `refactoring`) for your task
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
- **"Not inside a git repository"**: Run from a directory inside a git repo
|
||||
- **Index is stale after re-analyzing**: Restart Codex to reload the MCP server
|
||||
- **Embeddings slow**: Omit `--embeddings` (it's off by default) or set `OPENAI_API_KEY` for faster API-based embedding
|
||||
@@ -0,0 +1,89 @@
|
||||
---
|
||||
name: gitnexus-debugging
|
||||
description: "Use when the user is debugging a bug, tracing an error, or asking why something fails. Examples: \"Why is X failing?\", \"Where does this error come from?\", \"Trace this bug\""
|
||||
---
|
||||
|
||||
# Debugging with GitNexus
|
||||
|
||||
## When to Use
|
||||
|
||||
- "Why is this function failing?"
|
||||
- "Trace where this error comes from"
|
||||
- "Who calls this method?"
|
||||
- "This endpoint returns 500"
|
||||
- Investigating bugs, errors, or unexpected behavior
|
||||
|
||||
## Workflow
|
||||
|
||||
```
|
||||
1. gitnexus_query({query: "<error or symptom>"}) → Find related execution flows
|
||||
2. gitnexus_context({name: "<suspect>"}) → See callers/callees/processes
|
||||
3. READ gitnexus://repo/{name}/process/{name} → Trace execution flow
|
||||
4. gitnexus_cypher({query: "MATCH path..."}) → Custom traces if needed
|
||||
```
|
||||
|
||||
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
|
||||
|
||||
## Checklist
|
||||
|
||||
```
|
||||
- [ ] Understand the symptom (error message, unexpected behavior)
|
||||
- [ ] gitnexus_query for error text or related code
|
||||
- [ ] Identify the suspect function from returned processes
|
||||
- [ ] gitnexus_context to see callers and callees
|
||||
- [ ] Trace execution flow via process resource if applicable
|
||||
- [ ] gitnexus_cypher for custom call chain traces if needed
|
||||
- [ ] Read source files to confirm root cause
|
||||
```
|
||||
|
||||
## Debugging Patterns
|
||||
|
||||
| Symptom | GitNexus Approach |
|
||||
| -------------------- | ---------------------------------------------------------- |
|
||||
| Error message | `gitnexus_query` for error text → `context` on throw sites |
|
||||
| Wrong return value | `context` on the function → trace callees for data flow |
|
||||
| Intermittent failure | `context` → look for external calls, async deps |
|
||||
| Performance issue | `context` → find symbols with many callers (hot paths) |
|
||||
| Recent regression | `detect_changes` to see what your changes affect |
|
||||
|
||||
## Tools
|
||||
|
||||
**gitnexus_query** — find code related to error:
|
||||
|
||||
```
|
||||
gitnexus_query({query: "payment validation error"})
|
||||
→ Processes: CheckoutFlow, ErrorHandling
|
||||
→ Symbols: validatePayment, handlePaymentError, PaymentException
|
||||
```
|
||||
|
||||
**gitnexus_context** — full context for a suspect:
|
||||
|
||||
```
|
||||
gitnexus_context({name: "validatePayment"})
|
||||
→ Incoming calls: processCheckout, webhookHandler
|
||||
→ Outgoing calls: verifyCard, fetchRates (external API!)
|
||||
→ Processes: CheckoutFlow (step 3/7)
|
||||
```
|
||||
|
||||
**gitnexus_cypher** — custom call chain traces:
|
||||
|
||||
```cypher
|
||||
MATCH path = (a)-[:CodeRelation {type: 'CALLS'}*1..2]->(b:Function {name: "validatePayment"})
|
||||
RETURN [n IN nodes(path) | n.name] AS chain
|
||||
```
|
||||
|
||||
## Example: "Payment endpoint returns 500 intermittently"
|
||||
|
||||
```
|
||||
1. gitnexus_query({query: "payment error handling"})
|
||||
→ Processes: CheckoutFlow, ErrorHandling
|
||||
→ Symbols: validatePayment, handlePaymentError
|
||||
|
||||
2. gitnexus_context({name: "validatePayment"})
|
||||
→ Outgoing calls: verifyCard, fetchRates (external API!)
|
||||
|
||||
3. READ gitnexus://repo/my-app/process/CheckoutFlow
|
||||
→ Step 3: validatePayment → calls fetchRates (external)
|
||||
|
||||
4. Root cause: fetchRates calls external API without proper timeout
|
||||
```
|
||||
@@ -0,0 +1,78 @@
|
||||
---
|
||||
name: gitnexus-exploring
|
||||
description: "Use when the user asks how code works, wants to understand architecture, trace execution flows, or explore unfamiliar parts of the codebase. Examples: \"How does X work?\", \"What calls this function?\", \"Show me the auth flow\""
|
||||
---
|
||||
|
||||
# Exploring Codebases with GitNexus
|
||||
|
||||
## When to Use
|
||||
|
||||
- "How does authentication work?"
|
||||
- "What's the project structure?"
|
||||
- "Show me the main components"
|
||||
- "Where is the database logic?"
|
||||
- Understanding code you haven't seen before
|
||||
|
||||
## Workflow
|
||||
|
||||
```
|
||||
1. READ gitnexus://repos → Discover indexed repos
|
||||
2. READ gitnexus://repo/{name}/context → Codebase overview, check staleness
|
||||
3. gitnexus_query({query: "<what you want to understand>"}) → Find related execution flows
|
||||
4. gitnexus_context({name: "<symbol>"}) → Deep dive on specific symbol
|
||||
5. READ gitnexus://repo/{name}/process/{name} → Trace full execution flow
|
||||
```
|
||||
|
||||
> If step 2 says "Index is stale" → run `npx gitnexus analyze` in terminal.
|
||||
|
||||
## Checklist
|
||||
|
||||
```
|
||||
- [ ] READ gitnexus://repo/{name}/context
|
||||
- [ ] gitnexus_query for the concept you want to understand
|
||||
- [ ] Review returned processes (execution flows)
|
||||
- [ ] gitnexus_context on key symbols for callers/callees
|
||||
- [ ] READ process resource for full execution traces
|
||||
- [ ] Read source files for implementation details
|
||||
```
|
||||
|
||||
## Resources
|
||||
|
||||
| Resource | What you get |
|
||||
| --------------------------------------- | ------------------------------------------------------- |
|
||||
| `gitnexus://repo/{name}/context` | Stats, staleness warning (~150 tokens) |
|
||||
| `gitnexus://repo/{name}/clusters` | All functional areas with cohesion scores (~300 tokens) |
|
||||
| `gitnexus://repo/{name}/cluster/{name}` | Area members with file paths (~500 tokens) |
|
||||
| `gitnexus://repo/{name}/process/{name}` | Step-by-step execution trace (~200 tokens) |
|
||||
|
||||
## Tools
|
||||
|
||||
**gitnexus_query** — find execution flows related to a concept:
|
||||
|
||||
```
|
||||
gitnexus_query({query: "payment processing"})
|
||||
→ Processes: CheckoutFlow, RefundFlow, WebhookHandler
|
||||
→ Symbols grouped by flow with file locations
|
||||
```
|
||||
|
||||
**gitnexus_context** — 360-degree view of a symbol:
|
||||
|
||||
```
|
||||
gitnexus_context({name: "validateUser"})
|
||||
→ Incoming calls: loginHandler, apiMiddleware
|
||||
→ Outgoing calls: checkToken, getUserById
|
||||
→ Processes: LoginFlow (step 2/5), TokenRefresh (step 1/3)
|
||||
```
|
||||
|
||||
## Example: "How does payment processing work?"
|
||||
|
||||
```
|
||||
1. READ gitnexus://repo/my-app/context → 918 symbols, 45 processes
|
||||
2. gitnexus_query({query: "payment processing"})
|
||||
→ CheckoutFlow: processPayment → validateCard → chargeStripe
|
||||
→ RefundFlow: initiateRefund → calculateRefund → processRefund
|
||||
3. gitnexus_context({name: "processPayment"})
|
||||
→ Incoming: checkoutHandler, webhookHandler
|
||||
→ Outgoing: validateCard, chargeStripe, saveTransaction
|
||||
4. Read src/payments/processor.ts for implementation details
|
||||
```
|
||||
@@ -0,0 +1,64 @@
|
||||
---
|
||||
name: gitnexus-guide
|
||||
description: "Use when the user asks about GitNexus itself — available tools, how to query the knowledge graph, MCP resources, graph schema, or workflow reference. Examples: \"What GitNexus tools are available?\", \"How do I use GitNexus?\""
|
||||
---
|
||||
|
||||
# GitNexus Guide
|
||||
|
||||
Quick reference for all GitNexus MCP tools, resources, and the knowledge graph schema.
|
||||
|
||||
## Always Start Here
|
||||
|
||||
For any task involving code understanding, debugging, impact analysis, or refactoring:
|
||||
|
||||
1. **Read `gitnexus://repo/{name}/context`** — codebase overview + check index freshness
|
||||
2. **Match your task to a skill below** and **read that skill file**
|
||||
3. **Follow the skill's workflow and checklist**
|
||||
|
||||
> If step 1 warns the index is stale, run `npx gitnexus analyze` in the terminal first.
|
||||
|
||||
## Skills
|
||||
|
||||
| Task | Skill to read |
|
||||
| -------------------------------------------- | ------------------- |
|
||||
| Understand architecture / "How does X work?" | `gitnexus-exploring` |
|
||||
| Blast radius / "What breaks if I change X?" | `gitnexus-impact-analysis` |
|
||||
| Trace bugs / "Why is X failing?" | `gitnexus-debugging` |
|
||||
| Rename / extract / split / refactor | `gitnexus-refactoring` |
|
||||
| Tools, resources, schema reference | `gitnexus-guide` (this file) |
|
||||
| Index, status, clean, wiki CLI commands | `gitnexus-cli` |
|
||||
|
||||
## Tools Reference
|
||||
|
||||
| Tool | What it gives you |
|
||||
| ---------------- | ------------------------------------------------------------------------ |
|
||||
| `query` | Process-grouped code intelligence — execution flows related to a concept |
|
||||
| `context` | 360-degree symbol view — categorized refs, processes it participates in |
|
||||
| `impact` | Symbol blast radius — what breaks at depth 1/2/3 with confidence |
|
||||
| `detect_changes` | Git-diff impact — what do your current changes affect |
|
||||
| `rename` | Multi-file coordinated rename with confidence-tagged edits |
|
||||
| `cypher` | Raw graph queries (read `gitnexus://repo/{name}/schema` first) |
|
||||
| `list_repos` | Discover indexed repos |
|
||||
|
||||
## Resources Reference
|
||||
|
||||
Lightweight reads (~100-500 tokens) for navigation:
|
||||
|
||||
| Resource | Content |
|
||||
| ---------------------------------------------- | ----------------------------------------- |
|
||||
| `gitnexus://repo/{name}/context` | Stats, staleness check |
|
||||
| `gitnexus://repo/{name}/clusters` | All functional areas with cohesion scores |
|
||||
| `gitnexus://repo/{name}/cluster/{clusterName}` | Area members |
|
||||
| `gitnexus://repo/{name}/processes` | All execution flows |
|
||||
| `gitnexus://repo/{name}/process/{processName}` | Step-by-step trace |
|
||||
| `gitnexus://repo/{name}/schema` | Graph schema for Cypher |
|
||||
|
||||
## Graph Schema
|
||||
|
||||
**Nodes:** File, Function, Class, Interface, Method, Community, Process
|
||||
**Edges (via CodeRelation.type):** CALLS, IMPORTS, EXTENDS, IMPLEMENTS, DEFINES, MEMBER_OF, STEP_IN_PROCESS
|
||||
|
||||
```cypher
|
||||
MATCH (caller)-[:CodeRelation {type: 'CALLS'}]->(f:Function {name: "myFunc"})
|
||||
RETURN caller.name, caller.filePath
|
||||
```
|
||||
@@ -0,0 +1,97 @@
|
||||
---
|
||||
name: gitnexus-impact-analysis
|
||||
description: "Use when the user wants to know what will break if they change something, or needs safety analysis before editing code. Examples: \"Is it safe to change X?\", \"What depends on this?\", \"What will break?\""
|
||||
---
|
||||
|
||||
# Impact Analysis with GitNexus
|
||||
|
||||
## When to Use
|
||||
|
||||
- "Is it safe to change this function?"
|
||||
- "What will break if I modify X?"
|
||||
- "Show me the blast radius"
|
||||
- "Who uses this code?"
|
||||
- Before making non-trivial code changes
|
||||
- Before committing — to understand what your changes affect
|
||||
|
||||
## Workflow
|
||||
|
||||
```
|
||||
1. gitnexus_impact({target: "X", direction: "upstream"}) → What depends on this
|
||||
2. READ gitnexus://repo/{name}/processes → Check affected execution flows
|
||||
3. gitnexus_detect_changes() → Map current git changes to affected flows
|
||||
4. Assess risk and report to user
|
||||
```
|
||||
|
||||
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
|
||||
|
||||
## Checklist
|
||||
|
||||
```
|
||||
- [ ] gitnexus_impact({target, direction: "upstream"}) to find dependents
|
||||
- [ ] Review d=1 items first (these WILL BREAK)
|
||||
- [ ] Check high-confidence (>0.8) dependencies
|
||||
- [ ] READ processes to check affected execution flows
|
||||
- [ ] gitnexus_detect_changes() for pre-commit check
|
||||
- [ ] Assess risk level and report to user
|
||||
```
|
||||
|
||||
## Understanding Output
|
||||
|
||||
| Depth | Risk Level | Meaning |
|
||||
| ----- | ---------------- | ------------------------ |
|
||||
| d=1 | **WILL BREAK** | Direct callers/importers |
|
||||
| d=2 | LIKELY AFFECTED | Indirect dependencies |
|
||||
| d=3 | MAY NEED TESTING | Transitive effects |
|
||||
|
||||
## Risk Assessment
|
||||
|
||||
| Affected | Risk |
|
||||
| ------------------------------ | -------- |
|
||||
| <5 symbols, few processes | LOW |
|
||||
| 5-15 symbols, 2-5 processes | MEDIUM |
|
||||
| >15 symbols or many processes | HIGH |
|
||||
| Critical path (auth, payments) | CRITICAL |
|
||||
|
||||
## Tools
|
||||
|
||||
**gitnexus_impact** — the primary tool for symbol blast radius:
|
||||
|
||||
```
|
||||
gitnexus_impact({
|
||||
target: "validateUser",
|
||||
direction: "upstream",
|
||||
minConfidence: 0.8,
|
||||
maxDepth: 3
|
||||
})
|
||||
|
||||
→ d=1 (WILL BREAK):
|
||||
- loginHandler (src/auth/login.ts:42) [CALLS, 100%]
|
||||
- apiMiddleware (src/api/middleware.ts:15) [CALLS, 100%]
|
||||
|
||||
→ d=2 (LIKELY AFFECTED):
|
||||
- authRouter (src/routes/auth.ts:22) [CALLS, 95%]
|
||||
```
|
||||
|
||||
**gitnexus_detect_changes** — git-diff based impact analysis:
|
||||
|
||||
```
|
||||
gitnexus_detect_changes({scope: "staged"})
|
||||
|
||||
→ Changed: 5 symbols in 3 files
|
||||
→ Affected: LoginFlow, TokenRefresh, APIMiddlewarePipeline
|
||||
→ Risk: MEDIUM
|
||||
```
|
||||
|
||||
## Example: "What breaks if I change validateUser?"
|
||||
|
||||
```
|
||||
1. gitnexus_impact({target: "validateUser", direction: "upstream"})
|
||||
→ d=1: loginHandler, apiMiddleware (WILL BREAK)
|
||||
→ d=2: authRouter, sessionManager (LIKELY AFFECTED)
|
||||
|
||||
2. READ gitnexus://repo/my-app/processes
|
||||
→ LoginFlow and TokenRefresh touch validateUser
|
||||
|
||||
3. Risk: 2 direct callers, 2 processes = MEDIUM
|
||||
```
|
||||
@@ -0,0 +1,121 @@
|
||||
---
|
||||
name: gitnexus-refactoring
|
||||
description: "Use when the user wants to rename, extract, split, move, or restructure code safely. Examples: \"Rename this function\", \"Extract this into a module\", \"Refactor this class\", \"Move this to a separate file\""
|
||||
---
|
||||
|
||||
# Refactoring with GitNexus
|
||||
|
||||
## When to Use
|
||||
|
||||
- "Rename this function safely"
|
||||
- "Extract this into a module"
|
||||
- "Split this service"
|
||||
- "Move this to a new file"
|
||||
- Any task involving renaming, extracting, splitting, or restructuring code
|
||||
|
||||
## Workflow
|
||||
|
||||
```
|
||||
1. gitnexus_impact({target: "X", direction: "upstream"}) → Map all dependents
|
||||
2. gitnexus_query({query: "X"}) → Find execution flows involving X
|
||||
3. gitnexus_context({name: "X"}) → See all incoming/outgoing refs
|
||||
4. Plan update order: interfaces → implementations → callers → tests
|
||||
```
|
||||
|
||||
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
|
||||
|
||||
## Checklists
|
||||
|
||||
### Rename Symbol
|
||||
|
||||
```
|
||||
- [ ] gitnexus_rename({symbol_name: "oldName", new_name: "newName", dry_run: true}) — preview all edits
|
||||
- [ ] Review graph edits (high confidence) and ast_search edits (review carefully)
|
||||
- [ ] If satisfied: gitnexus_rename({..., dry_run: false}) — apply edits
|
||||
- [ ] gitnexus_detect_changes() — verify only expected files changed
|
||||
- [ ] Run tests for affected processes
|
||||
```
|
||||
|
||||
### Extract Module
|
||||
|
||||
```
|
||||
- [ ] gitnexus_context({name: target}) — see all incoming/outgoing refs
|
||||
- [ ] gitnexus_impact({target, direction: "upstream"}) — find all external callers
|
||||
- [ ] Define new module interface
|
||||
- [ ] Extract code, update imports
|
||||
- [ ] gitnexus_detect_changes() — verify affected scope
|
||||
- [ ] Run tests for affected processes
|
||||
```
|
||||
|
||||
### Split Function/Service
|
||||
|
||||
```
|
||||
- [ ] gitnexus_context({name: target}) — understand all callees
|
||||
- [ ] Group callees by responsibility
|
||||
- [ ] gitnexus_impact({target, direction: "upstream"}) — map callers to update
|
||||
- [ ] Create new functions/services
|
||||
- [ ] Update callers
|
||||
- [ ] gitnexus_detect_changes() — verify affected scope
|
||||
- [ ] Run tests for affected processes
|
||||
```
|
||||
|
||||
## Tools
|
||||
|
||||
**gitnexus_rename** — automated multi-file rename:
|
||||
|
||||
```
|
||||
gitnexus_rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: true})
|
||||
→ 12 edits across 8 files
|
||||
→ 10 graph edits (high confidence), 2 ast_search edits (review)
|
||||
→ Changes: [{file_path, edits: [{line, old_text, new_text, confidence}]}]
|
||||
```
|
||||
|
||||
**gitnexus_impact** — map all dependents first:
|
||||
|
||||
```
|
||||
gitnexus_impact({target: "validateUser", direction: "upstream"})
|
||||
→ d=1: loginHandler, apiMiddleware, testUtils
|
||||
→ Affected Processes: LoginFlow, TokenRefresh
|
||||
```
|
||||
|
||||
**gitnexus_detect_changes** — verify your changes after refactoring:
|
||||
|
||||
```
|
||||
gitnexus_detect_changes({scope: "all"})
|
||||
→ Changed: 8 files, 12 symbols
|
||||
→ Affected processes: LoginFlow, TokenRefresh
|
||||
→ Risk: MEDIUM
|
||||
```
|
||||
|
||||
**gitnexus_cypher** — custom reference queries:
|
||||
|
||||
```cypher
|
||||
MATCH (caller)-[:CodeRelation {type: 'CALLS'}]->(f:Function {name: "validateUser"})
|
||||
RETURN caller.name, caller.filePath ORDER BY caller.filePath
|
||||
```
|
||||
|
||||
## Risk Rules
|
||||
|
||||
| Risk Factor | Mitigation |
|
||||
| ------------------- | ----------------------------------------- |
|
||||
| Many callers (>5) | Use gitnexus_rename for automated updates |
|
||||
| Cross-area refs | Use detect_changes after to verify scope |
|
||||
| String/dynamic refs | gitnexus_query to find them |
|
||||
| External/public API | Version and deprecate properly |
|
||||
|
||||
## Example: Rename `validateUser` to `authenticateUser`
|
||||
|
||||
```
|
||||
1. gitnexus_rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: true})
|
||||
→ 12 edits: 10 graph (safe), 2 ast_search (review)
|
||||
→ Files: validator.ts, login.ts, middleware.ts, config.json...
|
||||
|
||||
2. Review ast_search edits (config.json: dynamic reference!)
|
||||
|
||||
3. gitnexus_rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: false})
|
||||
→ Applied 12 edits across 8 files
|
||||
|
||||
4. gitnexus_detect_changes({scope: "all"})
|
||||
→ Affected: LoginFlow, TokenRefresh
|
||||
→ Risk: MEDIUM — run tests for these flows
|
||||
```
|
||||
@@ -0,0 +1,16 @@
|
||||
---
|
||||
name: handoff
|
||||
description: Compact the current conversation into a handoff document for another agent to pick up.
|
||||
argument-hint: "What will the next session be used for?"
|
||||
disable-model-invocation: true
|
||||
---
|
||||
|
||||
Write a handoff document summarising the current conversation so a fresh agent can continue the work. Save to the temporary directory of the user's OS - not the current workspace.
|
||||
|
||||
Include a "suggested skills" section in the document, which suggests skills that the agent should invoke.
|
||||
|
||||
Do not duplicate content already captured in other artifacts (PRDs, plans, ADRs, issues, commits, diffs). Reference them by path or URL instead.
|
||||
|
||||
Redact any sensitive information, such as API keys, passwords, or personally identifiable information.
|
||||
|
||||
If the user passed arguments, treat them as a description of what the next session will focus on and tailor the doc accordingly.
|
||||
@@ -0,0 +1,156 @@
|
||||
---
|
||||
name: openspec-apply-change
|
||||
description: Implement tasks from an OpenSpec change. Use when the user wants to start implementing, continue implementation, or work through tasks.
|
||||
license: MIT
|
||||
compatibility: Requires openspec CLI.
|
||||
metadata:
|
||||
author: openspec
|
||||
version: "1.0"
|
||||
generatedBy: "1.3.1"
|
||||
---
|
||||
|
||||
Implement tasks from an OpenSpec change.
|
||||
|
||||
**Input**: Optionally specify a change name. If omitted, check if it can be inferred from conversation context. If vague or ambiguous you MUST prompt for available changes.
|
||||
|
||||
**Steps**
|
||||
|
||||
1. **Select the change**
|
||||
|
||||
If a name is provided, use it. Otherwise:
|
||||
- Infer from conversation context if the user mentioned a change
|
||||
- Auto-select if only one active change exists
|
||||
- If ambiguous, run `openspec list --json` to get available changes and use the **AskUserQuestion tool** to let the user select
|
||||
|
||||
Always announce: "Using change: <name>" and how to override (e.g., `/opsx:apply <other>`).
|
||||
|
||||
2. **Check status to understand the schema**
|
||||
```bash
|
||||
openspec status --change "<name>" --json
|
||||
```
|
||||
Parse the JSON to understand:
|
||||
- `schemaName`: The workflow being used (e.g., "spec-driven")
|
||||
- Which artifact contains the tasks (typically "tasks" for spec-driven, check status for others)
|
||||
|
||||
3. **Get apply instructions**
|
||||
|
||||
```bash
|
||||
openspec instructions apply --change "<name>" --json
|
||||
```
|
||||
|
||||
This returns:
|
||||
- `contextFiles`: artifact ID -> array of concrete file paths (varies by schema - could be proposal/specs/design/tasks or spec/tests/implementation/docs)
|
||||
- Progress (total, complete, remaining)
|
||||
- Task list with status
|
||||
- Dynamic instruction based on current state
|
||||
|
||||
**Handle states:**
|
||||
- If `state: "blocked"` (missing artifacts): show message, suggest using openspec-continue-change
|
||||
- If `state: "all_done"`: congratulate, suggest archive
|
||||
- Otherwise: proceed to implementation
|
||||
|
||||
4. **Read context files**
|
||||
|
||||
Read every file path listed under `contextFiles` from the apply instructions output.
|
||||
The files depend on the schema being used:
|
||||
- **spec-driven**: proposal, specs, design, tasks
|
||||
- Other schemas: follow the contextFiles from CLI output
|
||||
|
||||
5. **Show current progress**
|
||||
|
||||
Display:
|
||||
- Schema being used
|
||||
- Progress: "N/M tasks complete"
|
||||
- Remaining tasks overview
|
||||
- Dynamic instruction from CLI
|
||||
|
||||
6. **Implement tasks (loop until done or blocked)**
|
||||
|
||||
For each pending task:
|
||||
- Show which task is being worked on
|
||||
- Make the code changes required
|
||||
- Keep changes minimal and focused
|
||||
- Mark task complete in the tasks file: `- [ ]` → `- [x]`
|
||||
- Continue to next task
|
||||
|
||||
**Pause if:**
|
||||
- Task is unclear → ask for clarification
|
||||
- Implementation reveals a design issue → suggest updating artifacts
|
||||
- Error or blocker encountered → report and wait for guidance
|
||||
- User interrupts
|
||||
|
||||
7. **On completion or pause, show status**
|
||||
|
||||
Display:
|
||||
- Tasks completed this session
|
||||
- Overall progress: "N/M tasks complete"
|
||||
- If all done: suggest archive
|
||||
- If paused: explain why and wait for guidance
|
||||
|
||||
**Output During Implementation**
|
||||
|
||||
```
|
||||
## Implementing: <change-name> (schema: <schema-name>)
|
||||
|
||||
Working on task 3/7: <task description>
|
||||
[...implementation happening...]
|
||||
✓ Task complete
|
||||
|
||||
Working on task 4/7: <task description>
|
||||
[...implementation happening...]
|
||||
✓ Task complete
|
||||
```
|
||||
|
||||
**Output On Completion**
|
||||
|
||||
```
|
||||
## Implementation Complete
|
||||
|
||||
**Change:** <change-name>
|
||||
**Schema:** <schema-name>
|
||||
**Progress:** 7/7 tasks complete ✓
|
||||
|
||||
### Completed This Session
|
||||
- [x] Task 1
|
||||
- [x] Task 2
|
||||
...
|
||||
|
||||
All tasks complete! Ready to archive this change.
|
||||
```
|
||||
|
||||
**Output On Pause (Issue Encountered)**
|
||||
|
||||
```
|
||||
## Implementation Paused
|
||||
|
||||
**Change:** <change-name>
|
||||
**Schema:** <schema-name>
|
||||
**Progress:** 4/7 tasks complete
|
||||
|
||||
### Issue Encountered
|
||||
<description of the issue>
|
||||
|
||||
**Options:**
|
||||
1. <option 1>
|
||||
2. <option 2>
|
||||
3. Other approach
|
||||
|
||||
What would you like to do?
|
||||
```
|
||||
|
||||
**Guardrails**
|
||||
- Keep going through tasks until done or blocked
|
||||
- Always read context files before starting (from the apply instructions output)
|
||||
- If task is ambiguous, pause and ask before implementing
|
||||
- If implementation reveals issues, pause and suggest artifact updates
|
||||
- Keep code changes minimal and scoped to each task
|
||||
- Update task checkbox immediately after completing each task
|
||||
- Pause on errors, blockers, or unclear requirements - don't guess
|
||||
- Use contextFiles from CLI output, don't assume specific file names
|
||||
|
||||
**Fluid Workflow Integration**
|
||||
|
||||
This skill supports the "actions on a change" model:
|
||||
|
||||
- **Can be invoked anytime**: Before all artifacts are done (if tasks exist), after partial implementation, interleaved with other actions
|
||||
- **Allows artifact updates**: If implementation reveals design issues, suggest updating artifacts - not phase-locked, work fluidly
|
||||
@@ -0,0 +1,114 @@
|
||||
---
|
||||
name: openspec-archive-change
|
||||
description: Archive a completed change in the experimental workflow. Use when the user wants to finalize and archive a change after implementation is complete.
|
||||
license: MIT
|
||||
compatibility: Requires openspec CLI.
|
||||
metadata:
|
||||
author: openspec
|
||||
version: "1.0"
|
||||
generatedBy: "1.3.1"
|
||||
---
|
||||
|
||||
Archive a completed change in the experimental workflow.
|
||||
|
||||
**Input**: Optionally specify a change name. If omitted, check if it can be inferred from conversation context. If vague or ambiguous you MUST prompt for available changes.
|
||||
|
||||
**Steps**
|
||||
|
||||
1. **If no change name provided, prompt for selection**
|
||||
|
||||
Run `openspec list --json` to get available changes. Use the **AskUserQuestion tool** to let the user select.
|
||||
|
||||
Show only active changes (not already archived).
|
||||
Include the schema used for each change if available.
|
||||
|
||||
**IMPORTANT**: Do NOT guess or auto-select a change. Always let the user choose.
|
||||
|
||||
2. **Check artifact completion status**
|
||||
|
||||
Run `openspec status --change "<name>" --json` to check artifact completion.
|
||||
|
||||
Parse the JSON to understand:
|
||||
- `schemaName`: The workflow being used
|
||||
- `artifacts`: List of artifacts with their status (`done` or other)
|
||||
|
||||
**If any artifacts are not `done`:**
|
||||
- Display warning listing incomplete artifacts
|
||||
- Use **AskUserQuestion tool** to confirm user wants to proceed
|
||||
- Proceed if user confirms
|
||||
|
||||
3. **Check task completion status**
|
||||
|
||||
Read the tasks file (typically `tasks.md`) to check for incomplete tasks.
|
||||
|
||||
Count tasks marked with `- [ ]` (incomplete) vs `- [x]` (complete).
|
||||
|
||||
**If incomplete tasks found:**
|
||||
- Display warning showing count of incomplete tasks
|
||||
- Use **AskUserQuestion tool** to confirm user wants to proceed
|
||||
- Proceed if user confirms
|
||||
|
||||
**If no tasks file exists:** Proceed without task-related warning.
|
||||
|
||||
4. **Assess delta spec sync state**
|
||||
|
||||
Check for delta specs at `openspec/changes/<name>/specs/`. If none exist, proceed without sync prompt.
|
||||
|
||||
**If delta specs exist:**
|
||||
- Compare each delta spec with its corresponding main spec at `openspec/specs/<capability>/spec.md`
|
||||
- Determine what changes would be applied (adds, modifications, removals, renames)
|
||||
- Show a combined summary before prompting
|
||||
|
||||
**Prompt options:**
|
||||
- If changes needed: "Sync now (recommended)", "Archive without syncing"
|
||||
- If already synced: "Archive now", "Sync anyway", "Cancel"
|
||||
|
||||
If user chooses sync, use Task tool (subagent_type: "general-purpose", prompt: "Use Skill tool to invoke openspec-sync-specs for change '<name>'. Delta spec analysis: <include the analyzed delta spec summary>"). Proceed to archive regardless of choice.
|
||||
|
||||
5. **Perform the archive**
|
||||
|
||||
Create the archive directory if it doesn't exist:
|
||||
```bash
|
||||
mkdir -p openspec/changes/archive
|
||||
```
|
||||
|
||||
Generate target name using current date: `YYYY-MM-DD-<change-name>`
|
||||
|
||||
**Check if target already exists:**
|
||||
- If yes: Fail with error, suggest renaming existing archive or using different date
|
||||
- If no: Move the change directory to archive
|
||||
|
||||
```bash
|
||||
mv openspec/changes/<name> openspec/changes/archive/YYYY-MM-DD-<name>
|
||||
```
|
||||
|
||||
6. **Display summary**
|
||||
|
||||
Show archive completion summary including:
|
||||
- Change name
|
||||
- Schema that was used
|
||||
- Archive location
|
||||
- Whether specs were synced (if applicable)
|
||||
- Note about any warnings (incomplete artifacts/tasks)
|
||||
|
||||
**Output On Success**
|
||||
|
||||
```
|
||||
## Archive Complete
|
||||
|
||||
**Change:** <change-name>
|
||||
**Schema:** <schema-name>
|
||||
**Archived to:** openspec/changes/archive/YYYY-MM-DD-<name>/
|
||||
**Specs:** ✓ Synced to main specs (or "No delta specs" or "Sync skipped")
|
||||
|
||||
All artifacts complete. All tasks complete.
|
||||
```
|
||||
|
||||
**Guardrails**
|
||||
- Always prompt for change selection if not provided
|
||||
- Use artifact graph (openspec status --json) for completion checking
|
||||
- Don't block archive on warnings - just inform and confirm
|
||||
- Preserve .openspec.yaml when moving to archive (it moves with the directory)
|
||||
- Show clear summary of what happened
|
||||
- If sync is requested, use openspec-sync-specs approach (agent-driven)
|
||||
- If delta specs exist, always run the sync assessment and show the combined summary before prompting
|
||||
@@ -0,0 +1,288 @@
|
||||
---
|
||||
name: openspec-explore
|
||||
description: Enter explore mode - a thinking partner for exploring ideas, investigating problems, and clarifying requirements. Use when the user wants to think through something before or during a change.
|
||||
license: MIT
|
||||
compatibility: Requires openspec CLI.
|
||||
metadata:
|
||||
author: openspec
|
||||
version: "1.0"
|
||||
generatedBy: "1.3.1"
|
||||
---
|
||||
|
||||
Enter explore mode. Think deeply. Visualize freely. Follow the conversation wherever it goes.
|
||||
|
||||
**IMPORTANT: Explore mode is for thinking, not implementing.** You may read files, search code, and investigate the codebase, but you must NEVER write code or implement features. If the user asks you to implement something, remind them to exit explore mode first and create a change proposal. You MAY create OpenSpec artifacts (proposals, designs, specs) if the user asks—that's capturing thinking, not implementing.
|
||||
|
||||
**This is a stance, not a workflow.** There are no fixed steps, no required sequence, no mandatory outputs. You're a thinking partner helping the user explore.
|
||||
|
||||
---
|
||||
|
||||
## The Stance
|
||||
|
||||
- **Curious, not prescriptive** - Ask questions that emerge naturally, don't follow a script
|
||||
- **Open threads, not interrogations** - Surface multiple interesting directions and let the user follow what resonates. Don't funnel them through a single path of questions.
|
||||
- **Visual** - Use ASCII diagrams liberally when they'd help clarify thinking
|
||||
- **Adaptive** - Follow interesting threads, pivot when new information emerges
|
||||
- **Patient** - Don't rush to conclusions, let the shape of the problem emerge
|
||||
- **Grounded** - Explore the actual codebase when relevant, don't just theorize
|
||||
|
||||
---
|
||||
|
||||
## What You Might Do
|
||||
|
||||
Depending on what the user brings, you might:
|
||||
|
||||
**Explore the problem space**
|
||||
- Ask clarifying questions that emerge from what they said
|
||||
- Challenge assumptions
|
||||
- Reframe the problem
|
||||
- Find analogies
|
||||
|
||||
**Investigate the codebase**
|
||||
- Map existing architecture relevant to the discussion
|
||||
- Find integration points
|
||||
- Identify patterns already in use
|
||||
- Surface hidden complexity
|
||||
|
||||
**Compare options**
|
||||
- Brainstorm multiple approaches
|
||||
- Build comparison tables
|
||||
- Sketch tradeoffs
|
||||
- Recommend a path (if asked)
|
||||
|
||||
**Visualize**
|
||||
```
|
||||
┌─────────────────────────────────────────┐
|
||||
│ Use ASCII diagrams liberally │
|
||||
├─────────────────────────────────────────┤
|
||||
│ │
|
||||
│ ┌────────┐ ┌────────┐ │
|
||||
│ │ State │────────▶│ State │ │
|
||||
│ │ A │ │ B │ │
|
||||
│ └────────┘ └────────┘ │
|
||||
│ │
|
||||
│ System diagrams, state machines, │
|
||||
│ data flows, architecture sketches, │
|
||||
│ dependency graphs, comparison tables │
|
||||
│ │
|
||||
└─────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
**Surface risks and unknowns**
|
||||
- Identify what could go wrong
|
||||
- Find gaps in understanding
|
||||
- Suggest spikes or investigations
|
||||
|
||||
---
|
||||
|
||||
## OpenSpec Awareness
|
||||
|
||||
You have full context of the OpenSpec system. Use it naturally, don't force it.
|
||||
|
||||
### Check for context
|
||||
|
||||
At the start, quickly check what exists:
|
||||
```bash
|
||||
openspec list --json
|
||||
```
|
||||
|
||||
This tells you:
|
||||
- If there are active changes
|
||||
- Their names, schemas, and status
|
||||
- What the user might be working on
|
||||
|
||||
### When no change exists
|
||||
|
||||
Think freely. When insights crystallize, you might offer:
|
||||
|
||||
- "This feels solid enough to start a change. Want me to create a proposal?"
|
||||
- Or keep exploring - no pressure to formalize
|
||||
|
||||
### When a change exists
|
||||
|
||||
If the user mentions a change or you detect one is relevant:
|
||||
|
||||
1. **Read existing artifacts for context**
|
||||
- `openspec/changes/<name>/proposal.md`
|
||||
- `openspec/changes/<name>/design.md`
|
||||
- `openspec/changes/<name>/tasks.md`
|
||||
- etc.
|
||||
|
||||
2. **Reference them naturally in conversation**
|
||||
- "Your design mentions using Redis, but we just realized SQLite fits better..."
|
||||
- "The proposal scopes this to premium users, but we're now thinking everyone..."
|
||||
|
||||
3. **Offer to capture when decisions are made**
|
||||
|
||||
| Insight Type | Where to Capture |
|
||||
|----------------------------|--------------------------------|
|
||||
| New requirement discovered | `specs/<capability>/spec.md` |
|
||||
| Requirement changed | `specs/<capability>/spec.md` |
|
||||
| Design decision made | `design.md` |
|
||||
| Scope changed | `proposal.md` |
|
||||
| New work identified | `tasks.md` |
|
||||
| Assumption invalidated | Relevant artifact |
|
||||
|
||||
Example offers:
|
||||
- "That's a design decision. Capture it in design.md?"
|
||||
- "This is a new requirement. Add it to specs?"
|
||||
- "This changes scope. Update the proposal?"
|
||||
|
||||
4. **The user decides** - Offer and move on. Don't pressure. Don't auto-capture.
|
||||
|
||||
---
|
||||
|
||||
## What You Don't Have To Do
|
||||
|
||||
- Follow a script
|
||||
- Ask the same questions every time
|
||||
- Produce a specific artifact
|
||||
- Reach a conclusion
|
||||
- Stay on topic if a tangent is valuable
|
||||
- Be brief (this is thinking time)
|
||||
|
||||
---
|
||||
|
||||
## Handling Different Entry Points
|
||||
|
||||
**User brings a vague idea:**
|
||||
```
|
||||
User: I'm thinking about adding real-time collaboration
|
||||
|
||||
You: Real-time collab is a big space. Let me think about this...
|
||||
|
||||
COLLABORATION SPECTRUM
|
||||
════════════════════════════════════════════
|
||||
|
||||
Awareness Coordination Sync
|
||||
│ │ │
|
||||
▼ ▼ ▼
|
||||
┌────────┐ ┌────────┐ ┌────────┐
|
||||
│Presence│ │Cursors │ │ CRDT │
|
||||
│ "3 │ │ Multi │ │Conflict│
|
||||
│online" │ │ select │ │ free │
|
||||
└────────┘ └────────┘ └────────┘
|
||||
│ │ │
|
||||
trivial moderate complex
|
||||
|
||||
Where's your head at?
|
||||
```
|
||||
|
||||
**User brings a specific problem:**
|
||||
```
|
||||
User: The auth system is a mess
|
||||
|
||||
You: [reads codebase]
|
||||
|
||||
┌─────────────────────────────────────────────┐
|
||||
│ CURRENT AUTH FLOW │
|
||||
└─────────────────────────────────────────────┘
|
||||
│
|
||||
┌─────────────┼─────────────┐
|
||||
▼ ▼ ▼
|
||||
┌─────────┐ ┌─────────┐ ┌─────────┐
|
||||
│ Google │ │ GitHub │ │ Email │
|
||||
│ OAuth │ │ OAuth │ │ Magic │
|
||||
└────┬────┘ └────┬────┘ └────┬────┘
|
||||
│ │ │
|
||||
└─────────────┼─────────────┘
|
||||
▼
|
||||
┌───────────┐
|
||||
│ Session │
|
||||
└─────┬─────┘
|
||||
│
|
||||
▼
|
||||
┌───────────┐
|
||||
│ Perms │
|
||||
└───────────┘
|
||||
|
||||
I see three tangles. Which one's burning?
|
||||
```
|
||||
|
||||
**User is stuck mid-implementation:**
|
||||
```
|
||||
User: /opsx:explore add-auth-system
|
||||
The OAuth integration is more complex than expected
|
||||
|
||||
You: [reads change artifacts]
|
||||
|
||||
You're on task 4: "Implement OAuth flow"
|
||||
|
||||
Let me trace what's involved...
|
||||
|
||||
[draws diagram, explores options, suggests paths]
|
||||
|
||||
Want to update the design to reflect this?
|
||||
Or add a spike task to investigate?
|
||||
```
|
||||
|
||||
**User wants to compare options:**
|
||||
```
|
||||
User: Should we use Postgres or SQLite?
|
||||
|
||||
You: Generic answer is boring. What's the context?
|
||||
|
||||
User: A CLI tool that tracks local dev environments
|
||||
|
||||
You: That changes everything.
|
||||
|
||||
┌─────────────────────────────────────────────────┐
|
||||
│ CLI TOOL DATA STORAGE │
|
||||
└─────────────────────────────────────────────────┘
|
||||
|
||||
Key constraints:
|
||||
• No daemon running
|
||||
• Must work offline
|
||||
• Single user
|
||||
|
||||
SQLite Postgres
|
||||
Deployment embedded ✓ needs server ✗
|
||||
Offline yes ✓ no ✗
|
||||
Single file yes ✓ no ✗
|
||||
|
||||
SQLite. Not even close.
|
||||
|
||||
Unless... is there a sync component?
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Ending Discovery
|
||||
|
||||
There's no required ending. Discovery might:
|
||||
|
||||
- **Flow into a proposal**: "Ready to start? I can create a change proposal."
|
||||
- **Result in artifact updates**: "Updated design.md with these decisions"
|
||||
- **Just provide clarity**: User has what they need, moves on
|
||||
- **Continue later**: "We can pick this up anytime"
|
||||
|
||||
When it feels like things are crystallizing, you might summarize:
|
||||
|
||||
```
|
||||
## What We Figured Out
|
||||
|
||||
**The problem**: [crystallized understanding]
|
||||
|
||||
**The approach**: [if one emerged]
|
||||
|
||||
**Open questions**: [if any remain]
|
||||
|
||||
**Next steps** (if ready):
|
||||
- Create a change proposal
|
||||
- Keep exploring: just keep talking
|
||||
```
|
||||
|
||||
But this summary is optional. Sometimes the thinking IS the value.
|
||||
|
||||
---
|
||||
|
||||
## Guardrails
|
||||
|
||||
- **Don't implement** - Never write code or implement features. Creating OpenSpec artifacts is fine, writing application code is not.
|
||||
- **Don't fake understanding** - If something is unclear, dig deeper
|
||||
- **Don't rush** - Discovery is thinking time, not task time
|
||||
- **Don't force structure** - Let patterns emerge naturally
|
||||
- **Don't auto-capture** - Offer to save insights, don't just do it
|
||||
- **Do visualize** - A good diagram is worth many paragraphs
|
||||
- **Do explore the codebase** - Ground discussions in reality
|
||||
- **Do question assumptions** - Including the user's and your own
|
||||
@@ -0,0 +1,110 @@
|
||||
---
|
||||
name: openspec-propose
|
||||
description: Propose a new change with all artifacts generated in one step. Use when the user wants to quickly describe what they want to build and get a complete proposal with design, specs, and tasks ready for implementation.
|
||||
license: MIT
|
||||
compatibility: Requires openspec CLI.
|
||||
metadata:
|
||||
author: openspec
|
||||
version: "1.0"
|
||||
generatedBy: "1.3.1"
|
||||
---
|
||||
|
||||
Propose a new change - create the change and generate all artifacts in one step.
|
||||
|
||||
I'll create a change with artifacts:
|
||||
- proposal.md (what & why)
|
||||
- design.md (how)
|
||||
- tasks.md (implementation steps)
|
||||
|
||||
When ready to implement, run /opsx:apply
|
||||
|
||||
---
|
||||
|
||||
**Input**: The user's request should include a change name (kebab-case) OR a description of what they want to build.
|
||||
|
||||
**Steps**
|
||||
|
||||
1. **If no clear input provided, ask what they want to build**
|
||||
|
||||
Use the **AskUserQuestion tool** (open-ended, no preset options) to ask:
|
||||
> "What change do you want to work on? Describe what you want to build or fix."
|
||||
|
||||
From their description, derive a kebab-case name (e.g., "add user authentication" → `add-user-auth`).
|
||||
|
||||
**IMPORTANT**: Do NOT proceed without understanding what the user wants to build.
|
||||
|
||||
2. **Create the change directory**
|
||||
```bash
|
||||
openspec new change "<name>"
|
||||
```
|
||||
This creates a scaffolded change at `openspec/changes/<name>/` with `.openspec.yaml`.
|
||||
|
||||
3. **Get the artifact build order**
|
||||
```bash
|
||||
openspec status --change "<name>" --json
|
||||
```
|
||||
Parse the JSON to get:
|
||||
- `applyRequires`: array of artifact IDs needed before implementation (e.g., `["tasks"]`)
|
||||
- `artifacts`: list of all artifacts with their status and dependencies
|
||||
|
||||
4. **Create artifacts in sequence until apply-ready**
|
||||
|
||||
Use the **TodoWrite tool** to track progress through the artifacts.
|
||||
|
||||
Loop through artifacts in dependency order (artifacts with no pending dependencies first):
|
||||
|
||||
a. **For each artifact that is `ready` (dependencies satisfied)**:
|
||||
- Get instructions:
|
||||
```bash
|
||||
openspec instructions <artifact-id> --change "<name>" --json
|
||||
```
|
||||
- The instructions JSON includes:
|
||||
- `context`: Project background (constraints for you - do NOT include in output)
|
||||
- `rules`: Artifact-specific rules (constraints for you - do NOT include in output)
|
||||
- `template`: The structure to use for your output file
|
||||
- `instruction`: Schema-specific guidance for this artifact type
|
||||
- `outputPath`: Where to write the artifact
|
||||
- `dependencies`: Completed artifacts to read for context
|
||||
- Read any completed dependency files for context
|
||||
- Create the artifact file using `template` as the structure
|
||||
- Apply `context` and `rules` as constraints - but do NOT copy them into the file
|
||||
- Show brief progress: "Created <artifact-id>"
|
||||
|
||||
b. **Continue until all `applyRequires` artifacts are complete**
|
||||
- After creating each artifact, re-run `openspec status --change "<name>" --json`
|
||||
- Check if every artifact ID in `applyRequires` has `status: "done"` in the artifacts array
|
||||
- Stop when all `applyRequires` artifacts are done
|
||||
|
||||
c. **If an artifact requires user input** (unclear context):
|
||||
- Use **AskUserQuestion tool** to clarify
|
||||
- Then continue with creation
|
||||
|
||||
5. **Show final status**
|
||||
```bash
|
||||
openspec status --change "<name>"
|
||||
```
|
||||
|
||||
**Output**
|
||||
|
||||
After completing all artifacts, summarize:
|
||||
- Change name and location
|
||||
- List of artifacts created with brief descriptions
|
||||
- What's ready: "All artifacts created! Ready for implementation."
|
||||
- Prompt: "Run `/opsx:apply` or ask me to implement to start working on the tasks."
|
||||
|
||||
**Artifact Creation Guidelines**
|
||||
|
||||
- Follow the `instruction` field from `openspec instructions` for each artifact type
|
||||
- The schema defines what each artifact should contain - follow it
|
||||
- Read dependency artifacts for context before creating new ones
|
||||
- Use `template` as the structure for your output file - fill in its sections
|
||||
- **IMPORTANT**: `context` and `rules` are constraints for YOU, not content for the file
|
||||
- Do NOT copy `<context>`, `<rules>`, `<project_context>` blocks into the artifact
|
||||
- These guide what you write, but should never appear in the output
|
||||
|
||||
**Guardrails**
|
||||
- Create ALL artifacts needed for implementation (as defined by schema's `apply.requires`)
|
||||
- Always read dependency artifacts before creating a new one
|
||||
- If context is critically unclear, ask the user - but prefer making reasonable decisions to keep momentum
|
||||
- If a change with that name already exists, ask if user wants to continue it or create a new one
|
||||
- Verify each artifact file exists after writing before proceeding to next
|
||||
@@ -105,3 +105,46 @@ Find reuse opportunities + Trace the call/dependency chain and impact radius:
|
||||
Unless directly relevant to the user's current question, you should avoid proactively mentioning, illustrating, or
|
||||
trailing off into the following information in 99% of cases:
|
||||
|
||||
<!-- gitnexus:start -->
|
||||
# GitNexus — Code Intelligence
|
||||
|
||||
This project is indexed by GitNexus as **SuperBizAgent-java** (13483 symbols, 22230 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
|
||||
|
||||
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
|
||||
|
||||
## Always Do
|
||||
|
||||
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
|
||||
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
|
||||
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
|
||||
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
|
||||
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
|
||||
|
||||
## Never Do
|
||||
|
||||
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
|
||||
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
|
||||
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
|
||||
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
|
||||
|
||||
## Resources
|
||||
|
||||
| Resource | Use for |
|
||||
|----------|---------|
|
||||
| `gitnexus://repo/SuperBizAgent-java/context` | Codebase overview, check index freshness |
|
||||
| `gitnexus://repo/SuperBizAgent-java/clusters` | All functional areas |
|
||||
| `gitnexus://repo/SuperBizAgent-java/processes` | All execution flows |
|
||||
| `gitnexus://repo/SuperBizAgent-java/process/{name}` | Step-by-step execution trace |
|
||||
|
||||
## CLI
|
||||
|
||||
| Task | Read this skill file |
|
||||
|------|---------------------|
|
||||
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
|
||||
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
|
||||
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
|
||||
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
|
||||
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
|
||||
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
|
||||
|
||||
<!-- gitnexus:end -->
|
||||
|
||||
@@ -111,3 +111,46 @@ trailing off into the following information in 99% of cases:
|
||||
- 文档目录结构:
|
||||
- 不要将文档放到用户目录(如 `C:\Users\EDY\.claude\`)中
|
||||
|
||||
<!-- gitnexus:start -->
|
||||
# GitNexus — Code Intelligence
|
||||
|
||||
This project is indexed by GitNexus as **SuperBizAgent-java** (13483 symbols, 22230 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
|
||||
|
||||
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
|
||||
|
||||
## Always Do
|
||||
|
||||
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
|
||||
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
|
||||
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
|
||||
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
|
||||
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
|
||||
|
||||
## Never Do
|
||||
|
||||
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
|
||||
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
|
||||
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
|
||||
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
|
||||
|
||||
## Resources
|
||||
|
||||
| Resource | Use for |
|
||||
|----------|---------|
|
||||
| `gitnexus://repo/SuperBizAgent-java/context` | Codebase overview, check index freshness |
|
||||
| `gitnexus://repo/SuperBizAgent-java/clusters` | All functional areas |
|
||||
| `gitnexus://repo/SuperBizAgent-java/processes` | All execution flows |
|
||||
| `gitnexus://repo/SuperBizAgent-java/process/{name}` | Step-by-step execution trace |
|
||||
|
||||
## CLI
|
||||
|
||||
| Task | Read this skill file |
|
||||
|------|---------------------|
|
||||
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
|
||||
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
|
||||
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
|
||||
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
|
||||
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
|
||||
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
|
||||
|
||||
<!-- gitnexus:end -->
|
||||
|
||||
@@ -175,6 +175,10 @@
|
||||
- 定义:Tool 调用及结果投影的生命周期状态,固定为 `PROJECTING`、`READY`、`ERROR`。
|
||||
- 边界:它只说明调用记录是否完成,不说明结果是否包含证据。
|
||||
|
||||
### Durable Audit
|
||||
- 定义:为 Diagnosis Trace 长期保存的 Run、Agent 模型步骤和 Tool 调用元数据,用于 exact sessionId/runId 回放、评测和运维核对。
|
||||
- 边界:只保存有界、脱敏、可长期保留的身份、状态、耗时、预算和结果摘要;不保存 Prompt、Thought、完整 Tool 参数、raw response 或 Redis canonical invocation。
|
||||
|
||||
### Evidence Status
|
||||
- 定义:证据 Tool 的结果语义,固定为 `EVIDENCE_FOUND`、`NO_EVIDENCE`、`ERROR`。
|
||||
- 边界:`NO_EVIDENCE` 只表示当前查询范围内没有匹配结果,不能解释为问题不存在、根因被排除或系统健康。
|
||||
@@ -195,6 +199,14 @@
|
||||
- 定义:Harness 对同一技术操作 attempt 数和可重试失败类型的显式策略。
|
||||
- 边界:Router 与 SemanticGuard 的技术失败最多两次 attempt;Diagnosis Agent、Tool 和 Evidence repair 只有一次 attempt。Agent 正常 ReAct 轮次不是 retry,`NO_EVIDENCE`、业务拒绝、取消和预算耗尽不可重试。
|
||||
|
||||
### Chat Application Use Case
|
||||
- 定义:一次 Chat 请求的唯一业务入口,拥有 Session/Run、意图路由、固定执行器、PreviousTurn 和最终持久化。
|
||||
- 边界:不拥有 HTTP/SSE 连接,也不把 ChatModel 或 Tool 选择权交给 Controller。
|
||||
|
||||
### Chat SSE Contract
|
||||
- 定义:Chat 公开入口的五事件协议,顺序固定为 `metadata -> status* -> content|failure -> done`。
|
||||
- 边界:过程状态实时发送,最终安全内容最多释放一次;它不是 Token streaming,也不包含内部计划、Prompt、raw Tool 数据或异常。
|
||||
|
||||
### Verifier Skill Isolation
|
||||
- 定义:Chat Verifier 与 skill 系统隔离,只校验 Executor 答案和 `tool_trace_summary`。
|
||||
- 使用场景:防止 Verifier 把 playbook 指令当作事实证据;Verifier 只判断已有证据是否支持结论。
|
||||
|
||||
@@ -1,5 +1,12 @@
|
||||
# devflow 索引
|
||||
|
||||
## Issue 生命周期
|
||||
|
||||
| Issue | 状态 | 说明 |
|
||||
|---|---|---|
|
||||
| ISS-014 | archived | 阶段 0-7 的单体 Diagnosis Agent、Harness、ACI、SSE、清理和最终 E2E 已完成并归档;阶段实现对应的 11 个 devflow/OpenSpec 项目均已 archived。 |
|
||||
| ISS-015 | active | 承接 ISS-014 E2E 后发现的 Agent 硬停止、Evidence Repair Schema、Reasoning 审计验证/治理和 Fallback 信息质量问题。 |
|
||||
|
||||
## 项目
|
||||
|
||||
| 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|
||||
@@ -41,3 +48,6 @@
|
||||
| 2026-07-21 | single-react-mysql-readonly-tool | Fail-closed read-only MySQL evidence Tool with AST allowlist, JDBC controls and bounded projection | Harness/MySQL security | ISS-014, MySQL, JSqlParser, allowlist, PreparedStatement, timeout, projection | openspec/changes/archive/2026-07-21-single-react-mysql-readonly-tool | archived |
|
||||
| 2026-07-21 | single-react-diagnosis-agent | Single internal Diagnosis ReactAgent with Harness-controlled model/tool loop, bounded context and typed Draft | Harness/Diagnosis Agent/ReAct | ISS-014, ReactAgent, DiagnosisDraft, PreviousTurn, ToolInterceptor, ModelInterceptor, budget | openspec/changes/archive/2026-07-21-single-react-diagnosis-agent | archived |
|
||||
| 2026-07-21 | single-react-evidence-semantic-guards | Deterministic evidence validation, isolated semantic review and fail-closed diagnosis release | Harness/EvidenceGuard/SemanticGuard/Release | ISS-014, EvidenceGuard, verified snapshot, SemanticGuard, repair, fallback, release policy | openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards | archived |
|
||||
| 2026-07-21 | single-react-chat-application-usecase | Internal Chat application use case with isolated routing, fixed executors and safe PreviousTurn | Harness/Chat application/Run persistence | ISS-014, Intent Router, PreviousTurn, PublishedResult, V012, observer, cancellation | openspec/changes/archive/2026-07-21-single-react-chat-application-usecase | archived |
|
||||
| 2026-07-21 | single-react-chat-sse-cutover | Unique named-event Chat SSE endpoint, bounded production Harness wiring and strict frontend consumer | Chat/SSE/Harness production wiring | ISS-014, /api/chat, SSE, metadata, status, content, failure, done, disconnect, bounded executor | openspec/changes/archive/2026-07-22-single-react-chat-sse-cutover | archived |
|
||||
| 2026-07-22 | single-react-cleanup-e2e | Remove legacy Agent paths, add bounded Harness audit, and complete exact-run live acceptance | Chat/Harness/cleanup/E2E | ISS-014, single ReAct Agent, durable audit, named SSE, exact run, Flyway V013 | openspec/changes/archive/2026-07-22-single-react-cleanup-e2e | archived |
|
||||
|
||||
@@ -10,7 +10,7 @@
|
||||
|
||||
- `devflow/index.md` was checked. Relevant history includes `session-storage`, `confidence-feedback`, `executor-action-memory-relevance`, and `chat-verifier-agent`.
|
||||
- `mvp/archive/2026-07-09-doc-cleanup/notes/agent-engineering-decisions.md` already recommends the next phase as "可复现 MVP Demo", including `mvp-demo` profile, fixed diagnosis case, one-click request, and `GET /api/diagnosis/{sessionId}/trace`.
|
||||
- `mvp/issues/active/ISS-003-mvp-design-implementation-review.md` identifies test stability, session traceability, verifier evidence chain, upload path, and SupervisorAgent consistency as recent MVP concerns. Security cleanup is intentionally deferred by user decision.
|
||||
- `mvp/issues/archived/ISS-003-mvp-design-implementation-review.md` identifies test stability, session traceability, verifier evidence chain, upload path, and SupervisorAgent consistency as historical MVP concerns. Security cleanup was intentionally deferred by user decision at that time.
|
||||
|
||||
## Question Pool
|
||||
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
## Draft Acceptance
|
||||
|
||||
- [x] Issue exists: `mvp/issues/active/executor-evidence-attribution-hallucination.md`.
|
||||
- [x] Issue exists: `mvp/issues/archived/executor-evidence-attribution-hallucination.md`.
|
||||
- [x] OpenSpec change artifacts exist.
|
||||
- [x] devflow tracking files exist.
|
||||
- [x] OpenSpec validation passes.
|
||||
|
||||
@@ -0,0 +1,42 @@
|
||||
# Acceptance: single-react-chat-application-usecase
|
||||
|
||||
## Result
|
||||
|
||||
- Status: archived
|
||||
- OpenSpec tasks: 14/14 complete
|
||||
- Interface impact: L3 database/collaboration
|
||||
- Public protocol: unchanged
|
||||
|
||||
## Static Verification
|
||||
|
||||
- V012 migration、`DiagnosisRun` 和 `DiagnosisRunRepository` 对齐 `intent/release_outcome/published_result` 及安全 PreviousTurn filter。
|
||||
- 公开 Controller、前端和 endpoint diff 为空。
|
||||
- Router/System executor 未检出 Tool、ReactAgent、ThreadLocal 或手写循环。
|
||||
- `PublishedResult` 固定为 `user_query/published_conclusion/scope/limitations/source_documents`;序列化负向测试覆盖内部字段泄漏。
|
||||
|
||||
## Script Verification
|
||||
|
||||
- `mvn -q -DskipTests compile`:通过。
|
||||
- Stage 6A focused `ApplicationExecutorsTest,PublishedResultPersistenceTest,ChatApplicationUseCaseTest`:13 tests,通过。
|
||||
- Stage 2-5 与 6A regression selection:18 suites / 76 tests,0 failure/error/skipped。
|
||||
- `openspec validate single-react-chat-application-usecase --strict`:通过。
|
||||
|
||||
## Browser or Manual Verification
|
||||
|
||||
- Not applicable。阶段 6A 没有 UI 或公开入口变化。
|
||||
|
||||
## Not Verified
|
||||
|
||||
- 未运行真实 LLM、Redis、日志和 MySQL live E2E;按 ISS-014 串行门禁统一留到阶段 7。
|
||||
- V012 未在本阶段连接真实数据库执行;migration/entity/query 已由静态检查、focused persistence tests 和 compile 覆盖。
|
||||
|
||||
## Remaining Work
|
||||
|
||||
- 阶段 6B:唯一 `POST /api/chat` SSE 原子切换、旧 endpoint 删除和前端消费者迁移。
|
||||
- 阶段 7:旧链路清理、全局 spec 格式修复和最终 live E2E。
|
||||
|
||||
## Archive
|
||||
|
||||
- `.archive-ready`: created
|
||||
- OpenSpec archive: `openspec/changes/archive/2026-07-21-single-react-chat-application-usecase`
|
||||
- Main spec sync: `openspec/specs/single-react-chat-application-usecase/spec.md`
|
||||
@@ -0,0 +1,31 @@
|
||||
# Brief: single-react-chat-application-usecase
|
||||
|
||||
## Background
|
||||
|
||||
阶段 2-5 已具备 RunContext、单一 Diagnosis Agent 和安全释放门禁,但没有统一应用用例拥有 Session/Run、意图路由、PreviousTurn、固定执行器和最终持久化,阶段 6B 因而无法只做协议切换。
|
||||
|
||||
## Goal
|
||||
|
||||
在不改变公开 Chat/SSE 行为的前提下,建立内部 `ChatApplicationUseCase`,统一三类意图、同一 Run 生命周期、安全 PreviousTurn 和 typed public content。
|
||||
|
||||
## Scope
|
||||
|
||||
- 无 Tool、无记忆、无 ReAct 的三分类 Intent Router。
|
||||
- SYSTEM_CHAT、KNOWLEDGE_QUERY、DIAGNOSIS 固定执行器。
|
||||
- Knowledge exact invocation/document reference validation。
|
||||
- 同 Session 最近安全 Diagnosis SUCCESS 的有界 PreviousTurn。
|
||||
- `diagnosis_run` V012 字段、JPA store 和安全 `PublishedResult`。
|
||||
- Protocol-neutral observer、Run control、终态持久化和 focused tests。
|
||||
|
||||
## Non-goals
|
||||
|
||||
- 不修改 Controller、`/api/chat`、`/api/chat_stream`、SSE schema 或前端消费者。
|
||||
- 不删除旧 ChatService、多 Agent、ThreadLocal 或旧 session storage。
|
||||
- 不运行真实模型、Redis、日志和 MySQL live E2E;统一留到阶段 7。
|
||||
|
||||
## Metadata
|
||||
|
||||
- Scale: complex
|
||||
- Interface impact: L3 database/collaboration
|
||||
- OpenSpec: `single-react-chat-application-usecase`
|
||||
- Parent issue: `ISS-014`
|
||||
@@ -0,0 +1,117 @@
|
||||
# Decisions: single-react-chat-application-usecase
|
||||
|
||||
## Discover Status
|
||||
|
||||
- Checkpoint: Discover
|
||||
- Capability source: `sm-flow` + `grill-with-docs`;`codebase-retrieval`、LSP 和 GitNexus MCP 当前不可用,使用既有 GitNexus 结论、`rg` 引用核对和源码阅读降级。
|
||||
- Scale: complex。跨模型路由、三类执行器、Run/session 生命周期、数据库 migration、PreviousTurn 和阶段 6B consumer boundary。
|
||||
- `devflow/index.md` 命中阶段 0-5、session-run-trace-isolation、RAG contracts 和 release guards;无 ADR 冲突。
|
||||
|
||||
## Question Pool
|
||||
|
||||
| # | 维度 | 问题 | 模式 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | 术语 | Chat Application Use Case 与 Controller、Harness、Agent 的职责边界是什么? | evidence-driven | 已解决 |
|
||||
| Q2 | 路由 | Router 输入、输出和可重试失败范围是什么? | evidence-driven | 已解决 |
|
||||
| Q3 | 路由 | Router 最终失败是否允许默认进入 Diagnosis? | evidence-driven | 已解决 |
|
||||
| Q4 | 执行器 | 三种 intent 分别允许哪些模型和 Tool 行为? | evidence-driven | 已解决 |
|
||||
| Q5 | Knowledge | 单次 RAG 调用如何验证模型引用且不公开 Tool Call ID? | evidence-driven | 已解决 |
|
||||
| Q6 | PreviousTurn | 上一回合的真理源、筛选条件和截断边界是什么? | evidence-driven | 已解决 |
|
||||
| Q7 | 生命周期 | 何时读取上一回合、创建当前 Run、记录 intent 和终态? | evidence-driven | 已解决 |
|
||||
| Q8 | 持久化 | diagnosis_run 需要新增哪些字段,哪些内部内容禁止进入 published_result? | evidence-driven | 已解决 |
|
||||
| Q9 | 6B 边界 | 如何让 Controller 切换时不重写应用用例? | evidence-driven | 已解决 |
|
||||
| Q10 | 接口 | 数据库/内部接口影响等级和回滚要求是什么? | evidence-driven | 已解决 |
|
||||
| Q11 | 验收 | 如何证明原始 Query、sessionId/runId 和失败终态一致传播? | evidence-driven | 已解决 |
|
||||
|
||||
## Evidence-driven
|
||||
|
||||
| 结论 | 证据来源 | 是否已汇报用户 |
|
||||
|---|---|---|
|
||||
| Controller 只负责协议;Application Use Case 拥有 session/run、routing、executor、persistence,Harness 拥有预算/取消/释放,Agent 拥有诊断语义。 | ISS-014 3、4.2、阶段 6A/6B | 已汇报 |
|
||||
| Router 输入只含 query/last_intent/last_user_query,输出只允许三枚举;无 Tool/记忆/ReAct。 | ISS-014 4.5 | 已汇报 |
|
||||
| timeout/transport/非法输出可重试一次;第二次失败返回安全入口错误,不进入 Diagnosis。 | ISS-014 重试策略、`HarnessRetryPolicies.intentRouter()` | 已汇报 |
|
||||
| SYSTEM_CHAT 无 Tool;KNOWLEDGE_QUERY 只调用一次 lookup;DIAGNOSIS 进入 Agent + Guards。 | ISS-014 4.5 | 已汇报 |
|
||||
| PreviousTurn 只来自同 Session 最近 `DIAGNOSIS + SUCCESS + published_result`,不使用 Redis 历史。 | ISS-014 4.6 | 已汇报 |
|
||||
| `PublishedResult`/`PreviousTurn` 已冻结为 query/conclusion/scope/limitations/source_documents,不含 Tool ID/raw/Draft/reason。 | 阶段 0 contracts、`PublishedResult`、`PreviousTurn` | 已汇报 |
|
||||
| 当前 `DiagnosisRun`/V011 尚无 intent/release_outcome/published_result,需要 V012 和 repository query。 | `DiagnosisRun.java`、`V011__add_session_run_isolation.sql` | 已汇报 |
|
||||
| 上一回合必须在保存当前 PENDING Run 前读取,否则 latest query 会命中当前请求。 | repository 当前 latest method + 生命周期顺序推导 | 已汇报 |
|
||||
| 6B 需要 metadata/status/cancel,6A 应提供 observer 和显式 RunContext,而不包含 SSE 类型。 | ISS-014 4.7、阶段 6B | 已汇报 |
|
||||
|
||||
## User-interview
|
||||
|
||||
- 无新增 user-interview。三类路由、数据库字段、PreviousTurn、失败语义、阶段边界和自动 Apply/Archive/commit 均由 ISS-014 与用户持续授权冻结。
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- 在创建当前 Run 前读取 latest safe routing context 和 PreviousTurn,随后 `startRun -> persist RUNNING -> observer metadata -> route`。
|
||||
- Application output 使用 typed content union,Diagnosis success 转为无 Tool ID 的 public report view;Fallback output 不携带完整 snapshot。
|
||||
- Knowledge path 生成一次 direct canonical Tool Call ID(该路径没有框架 Tool Call),只调用 `lookup_knowledge`,严格验证 `KnowledgeAnswerDraft` 的 exact call ID 和 document subset,再移除 ID 发布。
|
||||
- 所有普通单轮模型调用复用阶段 5 的受控 `GuardModelCall`,从而共享 Core 模型/Token/timeout/cancel 边界;不引入新模型路由。
|
||||
- `PublishedResult` 仅在 Diagnosis SUCCESS 且 conclusion 非空时写入;Fallback/Failed/Cancelled/System/Knowledge 不生成 PreviousTurn 真理源。
|
||||
- 数据库变更使用 V012 可前向迁移;回滚为先停止新应用用例,再删除新索引/列,不影响 V011 既有字段。
|
||||
- 不创建 ADR:这些是 ISS-014 已冻结设计的落地,不是新的跨项目不可逆决策。
|
||||
|
||||
## OpenSpec Backfill
|
||||
|
||||
- 需进入 proposal/design/spec/tasks:路由隔离/重试、三执行器、original query、PreviousTurn filter/bounds、observer/cancel、Run terminal persistence、V012/L3、公开隔离。
|
||||
- 非目标:Controller/SSE/前端切换、旧链路删除、live E2E。
|
||||
|
||||
## Cross-artifact Alignment
|
||||
|
||||
| 上游 -> 下游 | 检查内容 | 状态 |
|
||||
|---|---|---|
|
||||
| ISS-014/brief -> proposal | 三路由、previous turn、Run lifecycle、内部-only 和阶段 6B handoff | 已对齐 |
|
||||
| proposal -> design | typed executors、observer、JPA/V012、异常/终态、L3 migration/rollback | 已对齐 |
|
||||
| design -> specs/tasks | 每项所有权/安全边界均有可观察 requirement 和实现测试切片 | 已对齐 |
|
||||
| specs -> tasks | 9 组 requirements 覆盖 Router/executors、store/policy、application、verification | 已对齐 |
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
- 能力来源:`zoom-out`,使用 Chat Session、Diagnosis Run、RunContext、Diagnosis Agent、EvidenceGuard、SemanticGuard 和 PublishedResult 术语。
|
||||
- 链路为 `request -> application use case -> run store/router -> fixed executor -> Harness/Agent/Tool -> typed public content -> run finish`;Controller 不拥有模型/工具/Run。
|
||||
- ChatRunStore 拥有 MySQL 映射,Core 拥有运行状态,Application 拥有 dispatch/终态,path executor 拥有单一路径行为;PreviousTurn policy 是唯一安全历史投影。
|
||||
- 最大风险是 prior/current Run 顺序和 DB/lifecycle 双终态,design/tasks 已固定 prior read before start、single finish path 和 focused failure/cancel tests。
|
||||
- V012 是 L3 additive schema;migration、entity、repository、rollback 独立章节完整,公开入口阶段 6A 零变化。
|
||||
|
||||
## Interface Impact
|
||||
|
||||
- 级别:L3 database/collaboration interface。
|
||||
- 新增 `diagnosis_run.intent/release_outcome/published_result` 和索引;修改 Entity/Repository,新增 internal application/store/output contract。
|
||||
- 消费者:阶段 6B Controller/SSE adapter、MySQL/Flyway;旧 ChatService 在本阶段不消费新字段。
|
||||
- 迁移/回滚:V012 nullable additive;回滚先切旧入口,再删除 index/columns。
|
||||
|
||||
## Commit Gate Preflight
|
||||
|
||||
- proposal、design、specs、tasks 完整,`openspec status` complete,change strict validation 通过。
|
||||
- Question pool 全部已解决并汇报,无 user-interview、未判级接口或未接受架构风险。
|
||||
- Cross-artifact 四段对齐无 gap;V012/L3、prior read ordering、terminal persistence 和 6B handoff 已进入 design/spec/tasks。
|
||||
- Apply/Archive/commit 使用用户持续授权;公开协议和前端必须保持零 diff。
|
||||
- `.committed` 已创建,可进入 Apply。
|
||||
|
||||
## Pre-apply Research
|
||||
|
||||
- `DiagnosisHarnessCore`/`RunContext`:Run ID、budget、cancel 和 first-terminal-wins。
|
||||
- `GuardModelCall`/`HarnessRetryExecutor`:单轮模型 timeout/usage 和 Router 两次 attempt。
|
||||
- `HarnessEvidenceTools`/`RagToolResult`:Knowledge 唯一 lookup 路径和有界 projection。
|
||||
- `DiagnosisAgentUseCase`/`DiagnosisReleaseUseCase`:Diagnosis Draft 与安全 release boundary。
|
||||
- `DiagnosisRun`/`DiagnosisRunRepository`/`V011`:现有 Run schema 和写入模式。
|
||||
- `ChatService.ensureChatSession/startDiagnosisRun`:只参考 session metadata/JPA 写法,不复用旧 routing、多 Agent 或 ThreadLocal。
|
||||
- 技术栈:Spring AI direct Prompt、Jackson strict JSON、JPA repository、Flyway additive migration、protocol-neutral observer;无 MQ/新依赖。
|
||||
|
||||
## Apply Progress
|
||||
|
||||
- Router/System/Knowledge contracts 与 executors 已完成,tasks 1.1-1.4 完成。
|
||||
- 5 个 `ApplicationExecutorsTest` 通过:同输入 retry、最终 routing failure、System direct call、Knowledge exact references/no ID、NO_EVIDENCE/model skip。
|
||||
- TODO:PublishedResult/JPA/V012、Diagnosis executor、总应用用例和综合验证。
|
||||
- REVIEW:修正 `JpaChatRunStore` 多构造器 Spring 注入歧义;prior/start/intent/finish 持久化异常统一为稳定 `RUN_PERSISTENCE_FAILED`,并在安全完成时更新 ChatSession 活跃时间/消息对数。均为代码偏离修复,无需变更 OpenSpec。
|
||||
- PublishedResult/JPA/V012、Diagnosis executor、protocol-neutral Run control 与总 ChatApplicationUseCase 已完成,tasks 2.1-3.4 完成。
|
||||
- 13 个 stage 6A focused tests 通过;TODO 仅剩综合回归、static scope 和 OpenSpec verification。
|
||||
- 综合回归曾在 `SemanticGuardTest.attemptTimeoutCancelsBothPermittedModelCalls` 出现负载相关失败。诊断确认生产代码对每次 timeout 均调用 `Future.cancel(true)`,但第二个 Future 可能在任务线程启动前已取消,此时不存在可接收 interrupt 的线程。分类为测试假设偏差,不是 OpenSpec 或生产代码偏离;回归断言改为两次 TIMEOUT attempt、两次模型预算预留,以及至少一个已运行调用收到 interrupt。
|
||||
|
||||
## Final Review
|
||||
|
||||
- Stage 6A focused tests 与阶段 2-5 regression 共 18 suites / 76 tests,0 failure/error/skipped;Maven compile 通过。
|
||||
- OpenSpec strict validation 通过;V012、Entity、Repository 的三个字段和 previous-turn filter 对齐。
|
||||
- 公开 Controller、前端和 endpoint 零 diff;Router/System executor 无 Tool、ReactAgent、ThreadLocal 或手写 loop。
|
||||
- `PublishedResult` 只包含 `user_query/published_conclusion/scope/limitations/source_documents`,负向序列化测试通过。
|
||||
- 本阶段不运行 live E2E,按 ISS-014 门禁留到阶段 7。
|
||||
@@ -0,0 +1,31 @@
|
||||
# Evidence: single-react-chat-application-usecase
|
||||
|
||||
## Code and Contract Evidence
|
||||
|
||||
- `DiagnosisHarnessCore`/`RunContext` 已提供 Run ID、预算、取消和 first-terminal-wins,应用用例无需创建第二套生命周期。
|
||||
- `GuardModelCall`/`HarnessRetryExecutor` 提供单轮模型 timeout、Token 记账和两次 Router attempt。
|
||||
- `HarnessEvidenceTools`/`RagToolResult` 提供 Knowledge 路径唯一 lookup 和有界 projection。
|
||||
- `DiagnosisAgentUseCase`/`DiagnosisReleaseUseCase` 已形成 Diagnosis Draft 与安全发布边界。
|
||||
- `DiagnosisRun`/`DiagnosisRunRepository`/V011 提供既有 Run 持久化,V012 以 nullable additive 字段扩展。
|
||||
|
||||
## Confirmed Boundaries
|
||||
|
||||
- Router 输入只含原始 Query、可选 last intent 和 last user query;最终失败不得默认进入 Diagnosis。
|
||||
- 三类 executor 不互相调用,所有路径接收未改写 Query。
|
||||
- PreviousTurn 只来自同 Session 最近 `DIAGNOSIS + SUCCESS + published_result`,且必须在保存当前 Run 前读取。
|
||||
- Knowledge 公开内容只保留 stable document metadata,不发布 direct Tool Call ID。
|
||||
- `PublishedResult` 不保存 Tool ID、raw evidence、完整 Draft 或 SemanticGuard reason;Fallback/Failed/Cancelled 不写安全历史。
|
||||
- Controller/SSE/前端切换属于阶段 6B,本阶段保持公开协议不变。
|
||||
|
||||
## Diagnosis Finding
|
||||
|
||||
- 综合回归暴露 `SemanticGuardTest` 的负载竞态:第二个 Future 可能在获得线程前被取消,因而不会产生第二次 interrupt。
|
||||
- 生产代码已对每次 timeout 调用 `Future.cancel(true)`;测试改为验证两次 TIMEOUT attempt、两次模型预算预留,以及至少一个运行中调用被中断。
|
||||
- 分类为测试假设偏差,不是生产代码或 OpenSpec 偏离。
|
||||
|
||||
## Verification Evidence
|
||||
|
||||
- Stage 6A focused tests:13 tests 通过。
|
||||
- Stage 2-5 与 6A 综合回归:18 suites / 76 tests,0 failure/error/skipped。
|
||||
- Maven compile 与 change strict validation 通过。
|
||||
- 公开 Controller/前端零 diff;Router/System executor 无 Tool、ReactAgent、ThreadLocal 或手写 loop。
|
||||
@@ -0,0 +1,49 @@
|
||||
# Acceptance: single-react-chat-sse-cutover
|
||||
|
||||
## Result
|
||||
|
||||
- Status: archived
|
||||
- OpenSpec tasks: 14/14 complete
|
||||
- Interface impact: L4 breaking HTTP/frontend contract
|
||||
- Public Chat protocol: unique named-event SSE `POST /api/chat`
|
||||
|
||||
## Static Verification
|
||||
|
||||
- 组合路由保持 `/api/chat`、`/api/ai_ops`、`/api/chat/clear`、`/api/chat/session/{sessionId}` 和 `/runs`;`/api/chat_stream` 已删除。
|
||||
- 生产 Controller/前端 legacy Chat token scan:0。
|
||||
- `ChatController` forbidden dependency scan:0;结构测试固定其三个 protocol/application dependencies。
|
||||
- SSE payload 只包含固定 metadata/status/content/failure/done fields,`CANCELLED` 不作为公开 done outcome。
|
||||
- mode selector DOM、state、consumer 和 CSS 已删除。
|
||||
|
||||
## Script Verification
|
||||
|
||||
- `mvn -q -DskipTests compile`:通过。
|
||||
- Stage 2-6B + AiOps regression selection:24 suites / 101 tests,0 failure/error/skipped。
|
||||
- `openspec validate single-react-chat-sse-cutover --strict`:通过。
|
||||
- `node --check src/main/resources/static/app.js`:通过。
|
||||
- `git diff --check`:通过,仅有仓库既存 LF/CRLF 提示。
|
||||
|
||||
## Browser or Manual Verification
|
||||
|
||||
- 本阶段未运行浏览器人工验证;前端协议由静态 contract test 和 JavaScript syntax check 覆盖。
|
||||
|
||||
## Not Verified
|
||||
|
||||
- 未运行 live 模型、Redis、日志和 MySQL E2E;按 ISS-014 阶段门禁统一留到阶段 7。
|
||||
- 未验证外部第三方 Chat API consumer;L4 变更不提供兼容分支,外部消费者必须同步迁移到 named-event SSE。
|
||||
|
||||
## Migration and Rollback
|
||||
|
||||
- 部署必须将后端 SSE endpoint 与 bundled frontend consumer 作为同一版本原子发布。
|
||||
- 回滚必须同时回滚 Controller 和 frontend 到阶段 6A commit;V012 additive nullable migration 可保留。
|
||||
- 不允许通过恢复 `/api/chat_stream`、同步 JSON consumer 或旧 message wrapper 形成双轨兼容。
|
||||
|
||||
## Remaining Work
|
||||
|
||||
- 阶段 7:物理删除旧 Agent/Graph/Hook/ThreadLocal/ChatService 路径和过时测试,更新文档并完成最终 live E2E、日志与数据库核验。
|
||||
|
||||
## Archive
|
||||
|
||||
- `.archive-ready`: created
|
||||
- OpenSpec archive: `openspec/changes/archive/2026-07-22-single-react-chat-sse-cutover`
|
||||
- Main spec sync: `openspec/specs/single-react-chat-sse-cutover/spec.md` (10 requirements added)
|
||||
@@ -0,0 +1,32 @@
|
||||
# Brief: single-react-chat-sse-cutover
|
||||
|
||||
## Background
|
||||
|
||||
阶段 6A 已建立 protocol-neutral `ChatApplicationUseCase`,但公开 Chat 仍有同步 `/api/chat` 和伪流式 `/api/chat_stream` 两条旧链路,Controller 直接拥有模型、Tools、Session 历史和无界线程池,前端也保留两套消费者。
|
||||
|
||||
## Goal
|
||||
|
||||
将公开 Chat 原子切换为唯一 `POST /api/chat` named-event SSE,并让 Controller 只承担校验、HTTP/SSE 和连接生命周期;所有公开内容必须来自阶段 6A 的安全释放结果。
|
||||
|
||||
## Scope
|
||||
|
||||
- 唯一 `/api/chat` SSE 与 `metadata -> status* -> content|failure -> done` 状态机。
|
||||
- exact Run disconnect/timeout/send-failure cancellation。
|
||||
- Spring-managed bounded Chat/model executors 和完整 Harness production Bean graph。
|
||||
- 前端唯一 named-event consumer、typed renderer 和 metadata identity 保存。
|
||||
- AiOps 模型/Tool acquisition 下沉到 service,并将 AiOps/Session endpoint 从 Chat Controller 职责中隔离。
|
||||
- L4 前后端迁移、成对回滚和 focused regression。
|
||||
|
||||
## Non-goals
|
||||
|
||||
- 不做 Token streaming、最终答案切片、断线续传、事件重放、轮询或 WebSocket。
|
||||
- 不修改 `/api/ai_ops` 的公开 URL、请求和 SSE message-wrapper 行为。
|
||||
- 不在本阶段物理删除旧多 Agent、ChatService、Hook 或 ThreadLocal;阶段 7 统一清理。
|
||||
- 不运行 live 模型、Redis、日志或 MySQL E2E;阶段 7 统一验收。
|
||||
|
||||
## Metadata
|
||||
|
||||
- Scale: complex
|
||||
- Interface impact: L4 breaking HTTP/frontend contract
|
||||
- OpenSpec: `single-react-chat-sse-cutover`
|
||||
- Parent issue: `ISS-014`
|
||||
@@ -0,0 +1,110 @@
|
||||
# Decisions: single-react-chat-sse-cutover
|
||||
|
||||
## Discover Status
|
||||
|
||||
- Checkpoint: Discover
|
||||
- Capability source: `sm-flow` + `grill-with-docs`;`codebase-retrieval`、LSP 和 GitNexus MCP 当前不可用,使用 `rg` 引用核对、源码阅读和 focused tests 降级。
|
||||
- Scale: complex。涉及 L4 HTTP/SSE 协议、前端消费者、异步连接生命周期、生产 Bean 装配和阶段 7 删除边界。
|
||||
- `devflow/index.md` 命中阶段 0-6A、ISS-013 和 session-run-trace-isolation;ISS-014 是更新且已冻结的最终协议源。
|
||||
|
||||
## Question Pool
|
||||
|
||||
| # | 维度 | 问题 | 模式 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | 术语 | “真正 SSE”是 Token streaming,还是过程事件实时 + 最终内容一次释放? | evidence-driven | 已解决 |
|
||||
| Q2 | 协议 | 唯一 endpoint、事件名称、payload、顺序、互斥和终态是什么? | evidence-driven | 已解决 |
|
||||
| Q3 | 边界 | Controller、Application Use Case、Harness 和 SSE adapter 各自拥有何种职责? | evidence-driven | 已解决 |
|
||||
| Q4 | 生命周期 | disconnect/timeout/send failure 如何取消同一个 Run,正常 complete 如何避免误取消? | evidence-driven | 已解决 |
|
||||
| Q5 | 执行器 | 如何消除 Controller 自建无界线程池并处理饱和? | evidence-driven | 已解决 |
|
||||
| Q6 | 装配 | 阶段 2-6A plain Java components 如何形成可启动的生产 Bean graph? | evidence-driven | 已解决 |
|
||||
| Q7 | 前端 | 快速/流式双模式如何迁移到唯一 SSE consumer? | evidence-driven | 已解决 |
|
||||
| Q8 | 安全 | 哪些内部状态/内容禁止进入 SSE? | evidence-driven | 已解决 |
|
||||
| Q9 | 兼容 | 是否保留同步 `/api/chat` 或 `/api/chat_stream` 兼容? | evidence-driven | 已解决 |
|
||||
| Q10 | 范围 | `/api/ai_ops` 和旧多 Agent 何时处理? | evidence-driven | 已解决 |
|
||||
| Q11 | 验收 | 如何证明 event order、single terminal、same IDs、cancel 和 production wiring? | evidence-driven | 已解决 |
|
||||
|
||||
## Evidence-driven
|
||||
|
||||
| 结论 | 证据来源 | 是否已汇报用户 |
|
||||
|---|---|---|
|
||||
| 真正 SSE 固定为状态实时发送、最终 typed content 一次释放,不做 Token/字符切片。 | ISS-014 4.7 | 已汇报 |
|
||||
| 唯一入口是 `POST /api/chat`;顺序为 `metadata -> status* -> content|failure -> done`。 | ISS-014 4.7、阶段 6B | 已汇报 |
|
||||
| metadata/status/content/failure/done 使用 named SSE event;payload 不再包含旧 `type/data` 包装。 | ISS-014 最小事件结构 | 已汇报 |
|
||||
| 当前 Controller 同时拥有旧 ChatService、模型、Tools、SessionManager、同步/伪流式流程和 cached thread pool。 | `ChatController.java` | 已汇报 |
|
||||
| 当前前端 quick 调 `/chat` JSON,stream 调 `/chat_stream` 并保留大量旧格式 fallback。 | `static/app.js` | 已汇报 |
|
||||
| `ChatApplicationUseCase` 已提供 observer、同一 run control、typed content 和 stable failure code,但组件尚无完整生产 Bean graph。 | 阶段 6A code + Spring annotation scan | 已汇报 |
|
||||
| 客户端断开必须通过 observer 得到的 `ChatRunControl` 取消,同一引用需覆盖断开早于 onStarted 的竞态。 | 阶段 6A design + `SseEmitter` lifecycle | 已汇报 |
|
||||
| `/api/ai_ops` 是独立公开入口,不属于阶段 6B Chat 原子切换;旧实现阶段 7 清理/处置。 | ISS-014 阶段 6B/7 | 已汇报 |
|
||||
|
||||
## User-interview
|
||||
|
||||
- 无新增 user-interview。唯一 endpoint、破坏性迁移、事件 schema、非 Token 流、自动 Apply/Archive/commit 均由 ISS-014 与用户持续授权冻结。
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Controller 使用构造注入的 `ChatApplicationUseCase` 与受控 `TaskExecutor`;不注入 ChatModel、Tools、ChatService 或 SessionManager 来处理 Chat。
|
||||
- SSE adapter 为每个请求维护单一 state machine 和 `AtomicReference<ChatRunControl>`;disconnect 标记先于 onStarted 时,onStarted 立即取消。
|
||||
- 正常 result 只产生一次 typed content 和 done;异常只产生 failure 和 done;IOException/timeout/disconnect 只取消,不尝试补发终态。
|
||||
- executor rejection 在没有 Run 时发送稳定 failure/done;worker 启动后所有 failure code 来自 `ChatApplicationException`,不暴露 cause。
|
||||
- 生产装配使用集中 `harness.chat` properties 和 Spring-managed bounded executors;所有模型调用复用同一 `ChatModel` 和 `GuardModelCall`。
|
||||
- 前端移除 mode selector 和 quick path,只保留一个严格 named-event parser;unknown event/schema fail closed。
|
||||
- 不创建 ADR:L4 方案已经在 ISS-014 设计冻结,本 change 负责原子落地和迁移说明。
|
||||
|
||||
## OpenSpec Backfill
|
||||
|
||||
- 需进入 proposal/design/spec/tasks:唯一 SSE、五事件 state machine、typed payload、安全 failure、same IDs、disconnect cancellation、bounded executors、production assembly、frontend migration、L4 rollback。
|
||||
- 非目标:AiOps 协议、旧类物理删除、live E2E。
|
||||
|
||||
## Cross-artifact Alignment
|
||||
|
||||
| 上游 -> 下游 | 检查内容 | 状态 |
|
||||
|---|---|---|
|
||||
| ISS-014/brief -> proposal | 唯一 SSE、五事件、断开取消、前端迁移、生产装配、阶段 7 边界 | 已对齐 |
|
||||
| proposal -> design | named events、state machine、bounded executors、Bean graph、AiOps 隔离、L4 rollback | 已对齐 |
|
||||
| design -> specs/tasks | 每项 ownership/lifecycle/safety/migration 均有可观察 requirement 和纵向切片 | 已对齐 |
|
||||
| specs -> tasks | 10 组 requirements 覆盖 wiring、SSE、cancel、frontend、AiOps regression 和 verification | 已对齐 |
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
- Capability source: `zoom-out`,使用 Chat Application Use Case、RunContext、Run Lifecycle、Diagnosis Harness 和 Chat SSE Contract 术语。
|
||||
- 链路为 `browser -> Controller -> SSE session -> bounded worker -> Application Use Case -> Harness -> typed release -> SSE session`;业务真理源不进入 Controller。
|
||||
- SSE session 独占协议状态,Application 独占 Run/dispatch/persistence,Core 独占 cancel/budget,configuration 独占 infrastructure graph。
|
||||
- 最大风险是 disconnect/onStarted 与 terminal callback 竞态,design/tasks 已固定 pending-disconnect、atomic terminal 和 no-send-after-close tests。
|
||||
- L4 部署必须 Controller/frontend 同版本;回滚成对返回阶段 6A commit,V012 可保留。
|
||||
|
||||
## Interface Impact
|
||||
|
||||
- Level: L4 breaking HTTP/frontend contract。
|
||||
- 删除同步 JSON `POST /api/chat`、`POST /api/chat_stream` 和旧 Chat SSE wrapper;新增 named-event SSE `POST /api/chat`。
|
||||
- 消费者:bundled `static/app.js`、外部 Chat callers、Controller/MockMvc tests;Trace/feedback 继续使用 metadata IDs。
|
||||
- 迁移/回滚:前后端同 commit 原子部署/回滚;不提供兼容开关、双 endpoint 或旧 parser。
|
||||
|
||||
## Commit Gate Preflight
|
||||
|
||||
- proposal、design、specs、tasks 完整,change strict validation 通过。
|
||||
- Question pool 全部 evidence-driven 解决并已汇报,无 user-interview、未判级接口或未接受架构风险。
|
||||
- Cross-artifact 四段对齐无 gap;L4 migration/rollback、production wiring、disconnect race 和阶段 7 边界均进入 design/spec/tasks。
|
||||
- `openspec-propose` 补全产物,`zoom-out` 完成架构审计;可提交为 Committed OpenSpec。
|
||||
|
||||
## Apply Progress
|
||||
|
||||
- 1.1-1.4 完成:新增 `ChatHarnessProperties`、bounded worker/model executors 和显式 `HarnessChatConfiguration` graph;空 MySQL datasource 只允许 fail-closed unknown logical ID。
|
||||
- 2.1-2.4 完成:新增 named-event `ChatSseEvent`/`ChatSseSession`,Chat 唯一 SSE Controller,移除 `/chat_stream` 和同步 Chat;AiOps 模型/Tool 依赖下沉到 service。
|
||||
- 3.1-3.3 完成:前端移除 quick/stream 双轨,统一 named-event parser、typed renderer 和静态 contract test。
|
||||
- `JpaChatRunStore` 生产构造器改为注入集中 `PublishedResultPolicy`,避免读写边界漂移;这是实现偏差修复,无需改需求方向。
|
||||
|
||||
## Apply Review
|
||||
|
||||
- Review 发现 `ChatController` 虽然 Chat path 已经只调用应用用例,但类本身仍承载 AiOps 和 Session 管理依赖,不完全满足“Chat Controller 只负责 Chat 协议”的 ownership 要求。
|
||||
- 用户确认将职责拆分为 `ChatController`(仅 `/api/chat`)、`AiOpsController`(保持 `/api/ai_ops`)和 `ChatSessionController`(保持 clear/session/runs URL)。
|
||||
- 该修正保持所有公开 URL、AiOps message-wrapper payload 和 Session observable behavior,不修改 Committed OpenSpec 的范围或方向。
|
||||
- `styles.css` 中 mode selector/dropdown 死样式随前端双轨删除一并移除。
|
||||
- AGENTS.md 指定的 `codebase-retrieval` 和 LSP 工具在当前环境不可用;使用 OpenSpec 全量上下文、`rg` 引用检查、Java 编译、结构测试和综合回归完成等价影响面确认。
|
||||
|
||||
## Verification and Migration
|
||||
|
||||
- Stage 2-6B + AiOps regression:24 suites / 101 tests,0 failure/error/skipped;包含 `HarnessChatConfigurationTest` Spring wiring。
|
||||
- `mvn -q -DskipTests compile`、`openspec validate single-react-chat-sse-cutover --strict`、`node --check src/main/resources/static/app.js` 均通过。
|
||||
- 生产 Controller/前端遗留 token 静态扫描为 0;`ChatController` 中模型、Tool、ChatService、AiOps、Session 和 Trace service 禁止依赖扫描为 0。
|
||||
- L4 部署必须将后端 `/api/chat` SSE 与 bundled frontend consumer 同版本原子部署;回滚必须成对回滚到阶段 6A commit,不提供兼容 endpoint、旧 parser 或双轨开关。
|
||||
- live 模型、Redis、日志和 MySQL E2E 按 ISS-014 门禁明确延后到阶段 7,本阶段没有把 focused/Mock 验证表述为 live 验收。
|
||||
@@ -0,0 +1,32 @@
|
||||
# Evidence: single-react-chat-sse-cutover
|
||||
|
||||
## Code and Contract Evidence
|
||||
|
||||
- `ChatApplicationUseCase` 已提供 typed result、observer 和 exact `ChatRunControl`,Controller 无需拥有模型、Tool 或业务路由。
|
||||
- `ChatSseSession` 使用 first-terminal-wins 状态和 pending disconnect,覆盖断开早于 `onStarted`、send failure、late terminal 与正常 completion 竞态。
|
||||
- `ChatSseEvent` 固定 metadata/status/content/failure/done payload;content 引用阶段 6A typed public content,failure 只公开 stable code/message。
|
||||
- `HarnessChatConfiguration` 组装同一个 Core、model boundary、canonical store、ToolBoundary、Diagnosis Agent、Guards、Router、executors 和 application use case。
|
||||
- `ChatHarnessProperties` 集中 worker/model queue、Run、SSE、canonical store、Agent、Router、Guard 和 single-turn limits;executors 使用有限队列与 `AbortPolicy`。
|
||||
- bundled frontend 只向 `/api/chat` 发起 streaming POST,按完整 named SSE frame 严格解析并 fail closed。
|
||||
|
||||
## Ownership Review
|
||||
|
||||
- Apply review 将原 Controller 拆为 `ChatController`、`AiOpsController` 和 `ChatSessionController`。
|
||||
- `ChatController` 只注入 `ChatApplicationUseCase`、bounded worker 和 Chat properties;禁止依赖扫描为 0。
|
||||
- `/api/ai_ops`、`/api/chat/clear`、session info 和 run list 的 URL 与 payload 字段保持不变。
|
||||
- mode selector/dropdown DOM、JS state 和 CSS 已全部移除,不保留双轨开关。
|
||||
|
||||
## Safety and Lifecycle Evidence
|
||||
|
||||
- success 测试验证 `metadata,status,content,done` 严格顺序和 typed content。
|
||||
- failure 测试验证内部 provider detail 不进入 failure payload,且 content/failure 互斥。
|
||||
- disconnect-before-start 与 send-failure 测试验证 exact Run 取消和 late terminal 阻断。
|
||||
- frontend contract test 验证唯一 request target、五类 event branch 与旧 consumer token 删除。
|
||||
- static scan:生产 Controller/前端 legacy Chat token 0;Chat Controller forbidden dependency 0。
|
||||
|
||||
## Verification Evidence
|
||||
|
||||
- Stage 2-6B + AiOps regression:24 suites / 101 tests,0 failure/error/skipped。
|
||||
- `HarnessChatConfigurationTest` 覆盖 bounded queues、共享依赖、空 MySQL fail-closed 和 Spring context wiring。
|
||||
- Maven compile、strict OpenSpec validation 和 JavaScript syntax check 通过。
|
||||
- 真实模型、Redis、日志和 MySQL live E2E 按 ISS-014 串行门禁留到阶段 7。
|
||||
@@ -6,7 +6,7 @@
|
||||
- 当前问题:Harness Core 和三类 evidence Tool 已就绪,但没有一个内部诊断执行链消费它们;公开 Chat 仍依赖旧的简单/多 Agent 路径。
|
||||
- 关联 OpenSpec:`openspec/changes/archive/2026-07-21-single-react-diagnosis-agent/`
|
||||
- devflow 分档:complex
|
||||
- 需求真理源:`mvp/issues/active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md`,不重复创建独立 PRD。
|
||||
- 需求真理源:`mvp/issues/archived/ISS-014-single-react-agent-harness-aci-ptk-refactor.md`,不重复创建独立 PRD。
|
||||
|
||||
## 范围
|
||||
|
||||
|
||||
@@ -4,7 +4,7 @@
|
||||
|
||||
| 来源 | 证据 | 结论 | 是否已汇报 |
|
||||
|---|---|---|---|
|
||||
| `mvp/issues/active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md` | 阶段 4 明确单 Agent、内部入口、极简上下文、结构化 Draft、无重试和预算验收 | 本 change 不得切换公开入口或删除旧链路 | 是 |
|
||||
| `mvp/issues/archived/ISS-014-single-react-agent-harness-aci-ptk-refactor.md` | 阶段 4 明确单 Agent、内部入口、极简上下文、结构化 Draft、无重试和预算验收 | 本 change 不得切换公开入口或删除旧链路 | 是 |
|
||||
| `ChatService.java` + GitNexus references | `createReactAgent` 被旧策略入口调用;复杂路径仍创建 Planner/Executor/Verifier/Composer | 新实现必须是独立内部 use case,不能复用旧 Service 运行职责 | 是 |
|
||||
| Spring AI Alibaba 1.1.2.0 `ReactAgent`/`AgentLlmNode` sources | 框架自带 ReAct loop;非流式 `ModelResponse` 保留 `ChatResponse` Usage | 不手写循环,使用 ModelInterceptor 强制模型/Token 预算 | 是 |
|
||||
| Spring AI Alibaba `ToolCallRequest`/`AgentToolNode` sources | `ToolCallRequest.getToolCallId()` 来自 `AssistantMessage.ToolCall.id()`,Tool interceptor 在 callback 前执行 | 可以精确传播框架 ID,不生成第二套 ID | 是 |
|
||||
|
||||
@@ -0,0 +1,57 @@
|
||||
# Acceptance: single-react-cleanup-e2e
|
||||
|
||||
## Result
|
||||
|
||||
Accepted. 阶段 7 的清理、Harness-native audit、文档收口和 live E2E 已完成;最终发布为安全 `FALLBACK`,没有释放未验证 Draft。
|
||||
|
||||
## Static Verification
|
||||
|
||||
- `openspec validate --all --strict`: 22/22 passed。
|
||||
- `node --check src/main/resources/static/app.js`: passed。
|
||||
- `git diff --check`: passed。
|
||||
- production legacy path、旧 Agent-facing Tool contract 和持久化敏感 payload 扫描:0 matches。
|
||||
- 最终 Tool 时间窗日志扫描:原始 query、request/response、Tool Call ID、Prompt、stack、debug instrumentation 均为 0;只保留有界 metadata。
|
||||
|
||||
## Script Verification
|
||||
|
||||
- Deterministic Maven regression: 60 suites / 221 tests,0 failures,0 errors,3 skipped。
|
||||
- `mvn -q -DskipTests compile`: passed。
|
||||
- `mvn -q -DskipTests package`: passed。
|
||||
- `mvn -q -Dtest=QueryLogsToolsTest,QueryLogsResultProjectorTest,CanonicalInvocationStoreTest test`: passed。
|
||||
- Flyway 9.22.3 repair + migrate: V012 checksum repaired,V013 applied,schema current version 013。
|
||||
- Maven `spring-boot:run` with `mvp-demo`: Flyway 13 migrations validated,JPA schema validation passed,Tomcat 9900 started,Harness/Audit wiring active。
|
||||
- `run-payment-timeout-demo.ps1`: strict named SSE validation passed for exact final session/run。
|
||||
- `scripts/query_mysql.py`: exact diagnosis_run、agent_step、tool_invocation queries passed。
|
||||
|
||||
## Final Live Evidence
|
||||
|
||||
- sessionId: `mvp-demo-payment-timeout-stage7-20260722-1741`
|
||||
- runId: `363f481c-33b8-42e7-8699-428a6ec61806`
|
||||
- SSE: `metadata -> status -> status -> status -> content -> done`
|
||||
- diagnosis_run: `status=SUCCESS`、`intent=DIAGNOSIS`、`release_outcome=FALLBACK`、`tool_call_count=2`、`total_token_count=11764`。
|
||||
- AgentStep: 2 exact-run rows,唯一 agent `diagnosis_agent`,`thought IS NULL`,只含角色/数量/Tool name metadata。
|
||||
- ToolInvocation: `lookup_knowledge` 和 Mock `query_logs` 各 1 行,均 `READY/EVIDENCE_FOUND/success=1`;input/details 只有 framework Tool Call ID、状态和字节数。
|
||||
|
||||
## Browser / Manual Verification
|
||||
|
||||
- 未执行浏览器点击验收;本阶段公开协议由真实 HTTP SSE 脚本和前端 JavaScript syntax/contract tests 覆盖。
|
||||
|
||||
## Remaining Risks
|
||||
|
||||
- live 结果为 EvidenceGuard/SemanticGuard 约束下的安全 `FALLBACK`,不是业务根因成功发布;这是允许的最终释放状态。
|
||||
- `diagnosis_run.step_count` 仍为空,但 exact AgentStep 查询返回 2 行;该历史汇总字段不作为本阶段 release gate。
|
||||
- `query_logs` 使用 Mock;真实 CLS 与生产业务 `query_mysql` datasource 仍属于后续接入范围。
|
||||
- `application-local.yml` 含本地内部配置且被 Git 忽略;未 stage、未提交、未输出凭据。
|
||||
|
||||
## Migration And Rollback
|
||||
|
||||
- L4 endpoint/frontend cleanup 必须整体回滚阶段 7 commit,不恢复双轨 endpoint 或旧 Tool annotations。
|
||||
- V013 仅幂等增加缺失列/索引;应用回滚时保留新增 nullable 列,避免破坏已写数据,不执行 destructive down migration。
|
||||
- Flyway repair 已将远端 V012 checksum 对齐当前迁移;V013 保证旧库与 fresh database 最终 schema 一致。
|
||||
|
||||
## Archive
|
||||
|
||||
- OpenSpec archive: completed at `openspec/changes/archive/2026-07-22-single-react-cleanup-e2e`。
|
||||
- Delta specs synced: created main `single-react-cleanup-e2e` spec and removed the temporary AiOps-preservation requirement from `single-react-chat-sse-cutover`。
|
||||
- Parent Issue `ISS-014` was closed and moved to `mvp/issues/archived/` after the stage 7 E2E evidence was accepted.
|
||||
- Post-E2E runtime-quality findings are tracked by active `ISS-015`; they do not reopen this completed cleanup change.
|
||||
@@ -0,0 +1,30 @@
|
||||
# Brief: single-react-cleanup-e2e
|
||||
|
||||
## Background
|
||||
|
||||
ISS-014 阶段 0-6B 已建立单一 Diagnosis ReAct Agent、Harness、ACI Tool、Guard 和 named SSE,但仓库仍存在 legacy AiOps/Sequential/Redis Session 链、旧 Tool contract、副作用式审计和过时文档。阶段 7 负责物理清理、durable metadata audit 和最终 live E2E。
|
||||
|
||||
## Goals
|
||||
|
||||
- 公开诊断只保留 `POST /api/chat`,业务代码只保留一个拥有 Tool loop 的 `diagnosis_agent`。
|
||||
- RAG/log 仅作为 Harness backend;Agent-facing Tool 只来自 `HarnessEvidenceTools`。
|
||||
- AgentStep 与 ToolInvocation 使用 exact sessionId/runId,长期持久化仅包含有界 metadata。
|
||||
- 通过 Maven 启动、named SSE、日志与 MySQL exact-run 查询完成最终验收。
|
||||
|
||||
## Scope
|
||||
|
||||
- 删除 legacy Controller、Service、Hook、ThreadLocal、prompt、前端入口、测试和死文档。
|
||||
- 新增 Harness-native Agent/Tool durable audit,修正文档、issue 与 OpenSpec strict 缺陷。
|
||||
- 为旧 V012 数据库增加幂等 V013 兼容迁移,并校准可重复 payment-timeout demo。
|
||||
|
||||
## Non-goals
|
||||
|
||||
- 不实现真实 CLS 或生产业务 MySQL Tool datasource。
|
||||
- 不修改三类 Intent、DiagnosisDraft、EvidenceGuard、SemanticGuard 或 Release Policy 语义。
|
||||
- 不保留 legacy endpoint、兼容分支、Graph 或第二套 Tool Call ID。
|
||||
|
||||
## Classification
|
||||
|
||||
- Scale: `complex`
|
||||
- Interface impact: L4 breaking HTTP/frontend cleanup
|
||||
- OpenSpec: `single-react-cleanup-e2e`
|
||||
@@ -0,0 +1,111 @@
|
||||
# Decisions: single-react-cleanup-e2e
|
||||
|
||||
## Discover Status
|
||||
|
||||
- Checkpoint: Discover
|
||||
- Capability source: `sm-flow` + `grill-with-docs` + `gitnexus-refactoring`。
|
||||
- GitNexus index 停在 `2362665`,当前为 `bc36248`;刷新会改写用户已修改的 `AGENTS.md`,因此不安全。使用 OpenSpec/devflow、`rg` 全引用扫描、源码阅读、编译和测试作为 fallback。
|
||||
- AGENTS.md 指定的 `codebase-retrieval` 与 LSP 工具在当前环境不可用,已显式记录限制。
|
||||
- Scale: complex。涉及 L4 endpoint 删除、跨模块物理清理、Tool/Trace ownership、安全持久化和 live E2E。
|
||||
|
||||
## Question Pool
|
||||
|
||||
| # | 维度 | 问题 | 模式 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | 范围 | 哪些旧 Agent/Service/Hook/Session 只有自引用测试,哪些仍在生产可达? | evidence-driven | 已解决 |
|
||||
| Q2 | 协议 | 阶段 7 是否必须删除 `/api/ai_ops` 与旧 Session endpoints? | evidence-driven | 已解决 |
|
||||
| Q3 | Tool | RAG/log 旧实现哪些可复用,哪些 Agent-facing contract/副作用必须删除? | evidence-driven | 已解决 |
|
||||
| Q4 | Trace | 新 Harness 如何在不泄漏 raw/prompt/thought 的前提下满足 AgentStep/ToolInvocation E2E? | evidence-driven | 已解决 |
|
||||
| Q5 | 文档 | ISS-012/ISS-013 和现有架构/Demo 文档如何收口? | evidence-driven | 已解决 |
|
||||
| Q6 | 验收 | 最终 live E2E 必须证明哪些 exact-run 事实,哪些外部系统明确不声称 live? | evidence-driven | 已解决 |
|
||||
| Q7 | 回滚 | L4 endpoint 删除如何迁移与回滚? | evidence-driven | 已解决 |
|
||||
|
||||
## Evidence-driven
|
||||
|
||||
| 结论 | 证据来源 | 是否已汇报用户 |
|
||||
|---|---|---|
|
||||
| `ChatService` 没有生产调用方,只剩自身单元/Smoke tests。 | `rg ChatService` | 已汇报 |
|
||||
| `/api/ai_ops` 仍由 bundled frontend 按钮调用,并运行 Supervisor + Planner + Executor 多 Agent。 | `AiOpsController`、`AiOpsService`、`app.js`、`index.html` | 已汇报 |
|
||||
| ISS-014 总体验收要求旧多 Agent/Graph 不存在,阶段 6B 只暂时保持 AiOps,阶段 7 负责清理/处置。 | ISS-014、阶段 6B design/acceptance | 已汇报 |
|
||||
| `/api/chat/clear` 与 session info/runs 没有 frontend caller;Redis SessionManager 只由该 Controller 和测试使用。 | controller/frontend/session 引用扫描 | 已汇报 |
|
||||
| `LookupKnowledgeTool` 和 `QueryLogsTools` 被新 Harness adapter 复用,但仍携带旧 `@Tool`、ThreadLocal/recorder 副作用。 | `HarnessChatConfiguration`、Tool source | 已汇报 |
|
||||
| 新 ToolBoundary 写 Redis canonical invocation,但没有 `tool_invocation` durable audit;最终 DB E2E 会缺 Tool rows。 | Harness boundary/config 引用扫描 | 已汇报 |
|
||||
| 旧 `AgentLoggingHook` 仍回退 ThreadLocal,并持久化 thought/model正文;不满足新安全边界。 | `AgentLoggingHook.java` | 已汇报 |
|
||||
| 当前架构、Agent、Harness 和 lifecycle 文档仍描述 Planner/Executor/Verifier/Composer 与 AIOps 双入口。 | `mvp/architecture/*.md` | 已汇报 |
|
||||
|
||||
## User-interview
|
||||
|
||||
- 无新增 user-interview。唯一公开 Chat、旧多 Agent/Graph 物理删除、无兼容分支、最终 E2E 和外部 Mock 边界均已由 ISS-014 与用户的逐阶段自动执行授权冻结。
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- 删除 legacy AiOps endpoint 而不是迁移到第二个 use case;所有诊断统一进入 `/api/chat` 的 Intent Router。
|
||||
- 删除旧 Session endpoints/Redis conversation context;安全 PreviousTurn 只来自 `diagnosis_run.published_result`。
|
||||
- RAG/log 查询实现保留为 Harness backend,移除 `@Tool` 和旧 recorder/session dedup;Agent 只看 ACI callbacks。
|
||||
- 新 Agent audit hook 只写角色/数量/Tool 名称/耗时等 metadata,不写模型输入正文、输出正文、Thought 或 Tool arguments。
|
||||
- Tool durable audit 通过 Harness port + JPA adapter fail-open 写入;Redis canonical store failure 仍 fail-closed,DB audit failure 只记录日志,不改变 Tool observation。
|
||||
- 删除 endpoint 的迁移无兼容层;bundled frontend 同 commit 删除按钮/consumer,回滚整体回滚 commit。
|
||||
- 不创建 ADR:方向已由 ISS-014 冻结,本 change 只完成最终落地与验收。
|
||||
|
||||
## OpenSpec Backfill
|
||||
|
||||
- 需进入 design/spec/tasks:删除清单、唯一 endpoint、Harness-native trace、Tool durable audit schema/safety、Tool backend 解耦、文档/issue 收口、strict validation、live E2E/log/DB acceptance 与回滚。
|
||||
|
||||
## Cross-artifact Alignment
|
||||
|
||||
| 上游 -> 下游 | 检查内容 | 状态 |
|
||||
|---|---|---|
|
||||
| ISS-014/brief -> proposal | 阶段 7 物理清理、文档、最终 Maven/log/DB E2E、Mock 外部 Tool 边界 | 已对齐 |
|
||||
| proposal -> design | 删除闭包、Tool backend 复用、安全 Agent/Tool audit、L4 migration/rollback | 已对齐 |
|
||||
| design -> specs/tasks | ownership、禁止泄漏、exact identity、strict validation、live E2E 均有 requirement 与切片 | 已对齐 |
|
||||
| specs -> tasks | 8 组可观察 requirements 覆盖删除、audit、docs、verification 和最终 E2E | 已对齐 |
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
- Capability source: `zoom-out`,使用 Diagnosis Harness、Diagnosis Agent、Canonical Invocation、Durable Audit、Diagnosis Trace 和 Chat SSE Contract 术语。
|
||||
- 最终链路为 `browser -> ChatController -> ChatApplicationUseCase -> Harness -> Diagnosis Agent -> ACI Tools -> Guards -> SSE`,没有第二业务入口或业务 Graph。
|
||||
- Redis canonical invocation 是短期完整 Tool 真理源;MySQL ToolInvocation 是长期有界 metadata audit,两者禁止双写 raw payload。
|
||||
- `chat_session` JPA entity 属于当前 Run/PreviousTurn 目录,Redis SessionContext 属于旧 conversation memory;删除时必须区分。
|
||||
- 最大风险是 backend 的旧 Tool annotation/recorder 隐式暴露和 audit 内容泄漏;design/tasks 已加入独占 discovery、negative serialization、context startup 与 live DB inspection。
|
||||
|
||||
## Interface Impact
|
||||
|
||||
- Level: L4 breaking HTTP/frontend contract。
|
||||
- 删除 `/api/ai_ops`、`/api/chat/clear`、`/api/chat/session/{sessionId}` 和 `/runs`;保留唯一 `/api/chat` SSE 与 Trace API。
|
||||
- bundled frontend 同 commit 删除 AiOps 按钮/consumer;外部 caller 迁移到 `/api/chat`,无兼容 branch。
|
||||
- 回滚必须整体回滚阶段 7 commit,不能单独恢复旧 endpoint/Tool annotations。
|
||||
|
||||
## Apply Verification Status
|
||||
|
||||
- Follow-up current-document audit found and corrected stale runtime semantics in the tracked payment-timeout PowerShell demo and `mvp/tables`: old JSON Chat parsing, Redis SessionContext history, Planner/Verifier identities, full Tool payload persistence, and Chat/AiOps Run descriptions are no longer presented as current behavior.
|
||||
- `mvp/demo/scripts/run-payment-timeout-demo.ps1` now strictly validates `metadata -> status* -> content|failure -> done`, captures exact session/run identity, and fetches exact Trace without writing feedback. PowerShell parser validation and two in-memory SSE contract samples passed.
|
||||
- Current architecture/demo/table scan excluding explicit `archive/` and ignored local `output/` artifacts: 0 legacy runtime matches.
|
||||
- Follow-up `openspec validate --all --strict`: 22 passed, 0 failed; `git diff --check`: passed.
|
||||
- Initial environment checks reported missing injected variables. Per user direction, credentials were restored to ignored `application-local.yml`; no credential was added to Git or emitted in the archive.
|
||||
- `mvn -q -DskipTests compile`: passed.
|
||||
- `node --check src/main/resources/static/app.js`: passed.
|
||||
- `openspec validate --all --strict`: 22 passed, 0 failed.
|
||||
- Deterministic regression suite excluding credential-dependent MySQL/Redis/Milvus tests: 60 suites, 221 tests, 0 failures, 0 errors, 3 skipped.
|
||||
- `SemanticGuardTest` interruption assertion failed once under the credential-dependent full-suite run, then passed three isolated repetitions and the deterministic regression suite; classified as load-sensitive test timing, not a reproduced product regression.
|
||||
- `mvn -q -DskipTests package`: passed.
|
||||
- Production legacy path scan, audit sensitive-payload scan, and legacy backend Tool contract scan: 0 matches.
|
||||
- The approved external-network run connected to MySQL/Redis/Milvus/model services. Flyway repair aligned the old V012 checksum and V013 reconciled the missing release-contract columns/index.
|
||||
- Maven startup validated all 13 migrations, passed JPA schema validation and started Tomcat 9900 with Harness/Audit beans.
|
||||
- Final exact live run: session `mvp-demo-payment-timeout-stage7-20260722-1741`, run `363f481c-33b8-42e7-8699-428a6ec61806`, SSE `metadata -> status -> status -> status -> content -> done`, outcome `FALLBACK`.
|
||||
- Exact MySQL evidence: one `DIAGNOSIS/SUCCESS/FALLBACK` run, two metadata-only `diagnosis_agent` steps with null Thought, and two READY/EVIDENCE_FOUND Tool audits for `lookup_knowledge` and Mock `query_logs`.
|
||||
- Tasks 4.2-4.5 are complete. Real CLS and production business MySQL Tool datasource remain explicit non-goals.
|
||||
|
||||
## Apply Conflict Classification
|
||||
|
||||
- **Code deviation**: `DiagnosisRun` enum mapping expected native ENUM while V012/V013 define VARCHAR. Fixed ORM column definitions; OpenSpec unchanged.
|
||||
- **Code deviation**: production `ObjectMapper` lacked Java Time modules, causing canonical Redis `STORE_ERROR`. Fixed mapper registration and bound the store test to the production mapper.
|
||||
- **Code deviation**: Mock empty log results used `success=false`, conflicting with the specified `NO_EVIDENCE` projection. Fixed backend success semantics and the stale test assertion.
|
||||
- **Code deviation**: RAG backend logs exposed raw query/rewritten query content. Replaced with bounded counts/category metadata and verified the final Tool execution window contains no prohibited payload.
|
||||
- **Acceptance fixture drift**: the payment-timeout demo requested the removed metrics Tool and encouraged an unbounded investigation. Updated the fixture to the current two-Tool Mock acceptance scope and explicit stopping boundary.
|
||||
|
||||
## Commit Gate Preflight
|
||||
|
||||
- proposal、design、两份 specs 和 tasks 完整;新 capability 与 modified capability 的范围无 gap。
|
||||
- Question pool 全部 evidence-driven 并已汇报,无 user-interview、未判级接口或未接受架构风险。
|
||||
- 删除清单区分 current JPA metadata 与 legacy Redis Session,Tool backend 与 Agent-facing contract,canonical truth 与 durable audit。
|
||||
- live E2E 明确要求真实应用/模型链和 exact ID;真实 CLS/生产业务 MySQL 明确不在验收声称范围。
|
||||
@@ -0,0 +1,26 @@
|
||||
# Evidence: single-react-cleanup-e2e
|
||||
|
||||
## Repository Evidence
|
||||
|
||||
- 生产引用扫描确认旧 `ChatService` 仅剩自身测试,`/api/ai_ops` 与旧 Session endpoints 属于第二条公开/状态链。
|
||||
- Harness adapters 复用 `LookupKnowledgeTool`/`QueryLogsTools` backend;旧 `@Tool`、ThreadLocal、recorder 和 topic discovery contract 已移除。
|
||||
- `HarnessAgentAuditHook` 只写 message count/roles、Tool names、text presence 和 duration;`thought` 保持空。
|
||||
- ToolBoundary durable audit 只写 exact identity、状态、稳定错误码、耗时和字节数;Redis canonical invocation 仍是短期完整 Tool 真理源。
|
||||
|
||||
## Live Findings
|
||||
|
||||
- 远端库已登记旧 V012 checksum,但缺少 `diagnosis_run.intent/release_outcome/published_result`。Flyway repair 后,幂等 V013 补齐三列与索引,fresh database 上为 no-op。
|
||||
- Hibernate 6 将无 `columnDefinition` 的字符串枚举校验为原生 ENUM;`DiagnosisRun` 已显式映射到 V012/V013 的 `VARCHAR(32/16)`。
|
||||
- 生产 `WebConfig` 的裸 `ObjectMapper` 无法序列化 canonical record 的 `Instant`,导致所有 Tool 在 begin 阶段返回 `STORE_ERROR`;改为自动注册模块,并让 store 测试使用生产 mapper。
|
||||
- Mock `query_logs` 把 0 命中错误表达为 `success=false`,与 `NO_EVIDENCE` contract 冲突;现以成功查询 + 空数组表达无证据。
|
||||
- live 日志发现 RAG backend 打印原始 query/rewrittenQuery/keywords;已改为字符数、命中数、类别数和耗时 metadata。
|
||||
|
||||
## Final Exact-run Evidence
|
||||
|
||||
- sessionId: `mvp-demo-payment-timeout-stage7-20260722-1741`
|
||||
- runId: `363f481c-33b8-42e7-8699-428a6ec61806`
|
||||
- SSE: `metadata -> status -> status -> status -> content -> done`
|
||||
- outcome: `FALLBACK`; diagnosis_run: `DIAGNOSIS / SUCCESS / FALLBACK`
|
||||
- AgentStep: 2 rows,均为 `diagnosis_agent`,`thought IS NULL`,model input/output 仅 metadata。
|
||||
- ToolInvocation: 2 rows,`lookup_knowledge` 与 `query_logs` 均为 `READY/EVIDENCE_FOUND`,同一 exact identity,无错误。
|
||||
- `query_logs` 明确为 Mock;Agent-facing `query_mysql` 未配置生产业务 datasource,也未声称 live。
|
||||
+11
-67
@@ -1,6 +1,6 @@
|
||||
# SuperBizAgent MVP 文档
|
||||
|
||||
**更新日期**:2026-07-10
|
||||
**更新日期**:2026-07-23
|
||||
|
||||
本目录保存 MVP 阶段的架构、问题、演示、评测和数据表说明。当前材料按“当前入口”和“历史归档”拆开,避免把早期设计稿当成当前实现。
|
||||
|
||||
@@ -10,27 +10,15 @@
|
||||
|---|---|
|
||||
| [architecture/README.md](architecture/README.md) | 当前 MVP 架构入口 |
|
||||
| [architecture/current-mvp-architecture.md](architecture/current-mvp-architecture.md) | 当前可运行系统架构 |
|
||||
| [architecture/interview-one-pager.md](architecture/interview-one-pager.md) | 面试一页式架构讲解 |
|
||||
| [architecture/agent-orchestration.md](architecture/agent-orchestration.md) | Agent 编排架构 |
|
||||
| [architecture/executor-evidence-pipeline-refactor.md](architecture/executor-evidence-pipeline-refactor.md) | Executor 证据链路改造记录 |
|
||||
| [architecture/harness-quality-gates.md](architecture/harness-quality-gates.md) | Harness 与质量门禁 |
|
||||
| [architecture/rag-architecture.md](architecture/rag-architecture.md) | RAG/知识检索新架构 |
|
||||
| [architecture/retrieval-observability.md](architecture/retrieval-observability.md) | 检索与可观测性架构 |
|
||||
| [architecture/feedback-architecture.md](architecture/feedback-architecture.md) | 反馈与自评估架构 |
|
||||
| [architecture/session-trace-lifecycle.md](architecture/session-trace-lifecycle.md) | 会话与 Trace 生命周期 |
|
||||
| [architecture/knowledge-base-authoring.md](architecture/knowledge-base-authoring.md) | 知识库文档编写与维护 |
|
||||
| [architecture/data-model.md](architecture/data-model.md) | 数据模型总览 |
|
||||
| [architecture/evolution-roadmap.md](architecture/evolution-roadmap.md) | Agent 架构演进路线 |
|
||||
| [issues/README.md](issues/README.md) | MVP issue 索引 |
|
||||
| [issues/active/rag-refactor-plan.md](issues/active/rag-refactor-plan.md) | RAG 重构计划和阶段拆解 |
|
||||
| [tables/README.md](tables/README.md) | 当前 MySQL 表说明 |
|
||||
| [demo/README.md](demo/README.md) | Demo 运行和面试演示材料 |
|
||||
| [demo/ten-minute-interview-demo.md](demo/ten-minute-interview-demo.md) | 10 分钟面试演示脚本 |
|
||||
| [demo/README.md](demo/README.md) | Demo 运行和演示材料 |
|
||||
| [eval/README.md](eval/README.md) | 诊断评测材料 |
|
||||
|
||||
## 当前系统一句话
|
||||
|
||||
SuperBizAgent MVP 是一个可追踪的故障诊断 Agent:Chat 和 AIOps 入口进入 Agent 编排,Executor 显式调用知识库、日志、指标等工具收集证据;多轮会话元数据落到 `chat_session`,每次诊断运行落到 `diagnosis_run`,步骤和工具明细通过 `agent_step.run_id`、`tool_invocation.run_id` 关联,最终通过 Trace API、Verifier 和评测脚本证明结果可解释、可回放、可对比。
|
||||
当前架构、运行链路和后续规划分别以 `architecture/`、`issues/README.md`、`tables/README.md` 以及 OpenSpec/devflow 的最新记录为准。
|
||||
|
||||
## 文档结构
|
||||
|
||||
@@ -39,18 +27,12 @@ mvp/
|
||||
architecture/
|
||||
README.md
|
||||
current-mvp-architecture.md
|
||||
interview-one-pager.md
|
||||
agent-orchestration.md
|
||||
executor-evidence-pipeline-refactor.md
|
||||
harness-quality-gates.md
|
||||
rag-architecture.md
|
||||
retrieval-observability.md
|
||||
feedback-architecture.md
|
||||
session-trace-lifecycle.md
|
||||
knowledge-base-authoring.md
|
||||
data-model.md
|
||||
evolution-roadmap.md
|
||||
archive/
|
||||
2026-07-05-legacy/
|
||||
2026-07-22-legacy/
|
||||
issues/
|
||||
README.md
|
||||
active/
|
||||
@@ -76,51 +58,13 @@ mvp/
|
||||
archive/
|
||||
```
|
||||
|
||||
## 当前核心设计
|
||||
|
||||
- `lookup_knowledge` 保持显式 Agent Tool,不隐藏到 Chat Advisor。
|
||||
- L0 降级为 domain/entity hint,不再默认承担最终召回决策。
|
||||
- `VectorSearchService` 是检索稳定门面。
|
||||
- Spring AI VectorStore 是当前读取主路径,Milvus SDK 保留为 fallback。
|
||||
- AIOps payload 会生成推荐知识库 query,保留业务语义。
|
||||
- `sessionId` 表示多轮会话上下文,`runId` 表示一次可回放诊断运行。
|
||||
- Trace API 聚合 `diagnosis_run`、`agent_step.run_id`、`tool_invocation.run_id` 和 self evaluation。
|
||||
- RAG 行为通过 offline baseline 和 live acceptance 脚本做回归验证。
|
||||
|
||||
## 关键运行链路
|
||||
|
||||
```text
|
||||
Chat
|
||||
-> ChatService
|
||||
-> Planner / Executor / Verifier
|
||||
-> evidence tools
|
||||
-> chat_session / diagnosis_run
|
||||
-> agent_step.run_id / tool_invocation.run_id
|
||||
-> DiagnosisTraceService
|
||||
|
||||
AIOps
|
||||
-> AiOpsService
|
||||
-> PAYLOAD_TARGETED or AUTO_DISCOVERY
|
||||
-> Planner / Executor
|
||||
-> Prometheus / logs / lookup_knowledge
|
||||
-> AiOpsRuleEvaluationService
|
||||
-> diagnosis_run(agent_flow=AI_OPS)
|
||||
-> DiagnosisTraceService
|
||||
|
||||
RAG
|
||||
-> lookup_knowledge
|
||||
-> L0 domain/entity hint
|
||||
-> VectorSearchService
|
||||
-> Spring AI VectorStore / Milvus SDK fallback
|
||||
-> relevance normalization
|
||||
-> tool_invocation
|
||||
```
|
||||
|
||||
## 归档说明
|
||||
|
||||
历史材料分两类:
|
||||
历史材料按归档批次保存:
|
||||
|
||||
- 旧架构文档:[architecture/archive/2026-07-05-legacy/](architecture/archive/2026-07-05-legacy/)
|
||||
- 本次文档清理归档:[archive/2026-07-09-doc-cleanup/](archive/2026-07-09-doc-cleanup/)
|
||||
- [architecture/archive/2026-07-05-legacy/](architecture/archive/2026-07-05-legacy/):早期架构设计、实现计划和知识检索方案。
|
||||
- [architecture/archive/2026-07-22-legacy/](architecture/archive/2026-07-22-legacy/):单 Diagnosis Agent + Harness 切换前的多角色编排、双入口和旧证据链架构。
|
||||
- [archive/2026-07-09-doc-cleanup/](archive/2026-07-09-doc-cleanup/):文档清理时迁移的历史材料。
|
||||
- [issues/archived/](issues/archived/):已关闭或已被当前架构替代的 Issue。
|
||||
|
||||
归档文档只用于追溯设计历史。当前实现和后续规划以 `architecture/`、`issues/README.md`、`tables/README.md` 和 OpenSpec/devflow 的最新记录为准。
|
||||
归档内容仅用于追溯历史决策,不代表当前 runtime、API、数据模型或验收口径。
|
||||
|
||||
+10
-41
@@ -1,48 +1,17 @@
|
||||
# MVP 架构文档
|
||||
|
||||
**更新日期**:2026-07-10
|
||||
**更新日期**:2026-07-23
|
||||
**状态**:当前单 Diagnosis Agent + Harness 架构
|
||||
|
||||
这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到:
|
||||
当前文档入口:
|
||||
|
||||
- `mvp/architecture/archive/2026-07-05-legacy/`
|
||||
|
||||
归档材料只作为设计历史阅读,不再作为当前实现依据。
|
||||
|
||||
## 当前文档
|
||||
|
||||
| 文档 | 用途 |
|
||||
| 文档 | 内容 |
|
||||
|---|---|
|
||||
| [current-mvp-architecture.md](current-mvp-architecture.md) | 当前可运行 MVP 的总体架构、链路、持久化和质量门禁 |
|
||||
| [interview-one-pager.md](interview-one-pager.md) | 面试一页式架构讲解,包含总图、亮点、取舍和追问回答 |
|
||||
| [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat SequentialAgent、AIOps SupervisorAgent、工具边界 |
|
||||
| [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md) | Chat 证据链路当前数据契约,覆盖 Executor V2、Gatekeeper、Verifier、Composer、`evidence_refs` |
|
||||
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Gatekeeper、Verifier、Composer、评测基线组成的质量门禁 |
|
||||
| [rag-architecture.md](rag-architecture.md) | RAG/知识检索新架构,覆盖 L0 hint、VectorStore 主路径、SDK fallback、证据追踪 |
|
||||
| [modular-rag-pipeline.md](modular-rag-pipeline.md) | `lookup_knowledge` 模块化 RAG 落地架构,覆盖 pipeline、fallback、evidence-first contract、trace |
|
||||
| [rag-eval-closure.md](rag-eval-closure.md) | RAG 评测闭环,覆盖 offline baseline、baseline diff、diagnosis eval 和 live acceptance |
|
||||
| [retrieval-observability.md](retrieval-observability.md) | 检索运行细节和可观测性,覆盖 L0/L1、去重、分数归一、评测 |
|
||||
| [feedback-architecture.md](feedback-architecture.md) | 反馈与自评估闭环,覆盖 rule evaluation、Verifier、AIOps rule、用户反馈和案例沉淀 |
|
||||
| [session-trace-lifecycle.md](session-trace-lifecycle.md) | 会话和 Trace 生命周期,覆盖 sessionId、状态流转、agent_step、tool_invocation、Trace API |
|
||||
| [knowledge-base-authoring.md](knowledge-base-authoring.md) | 知识库文档编写与维护规范,覆盖 frontmatter、category、chunk、reindex |
|
||||
| [data-model.md](data-model.md) | 数据模型总览,覆盖 Trace、知识库、反馈沉淀和 Milvus metadata |
|
||||
| [evolution-roadmap.md](evolution-roadmap.md) | 从旧版 Agent 蓝图继承的后续演进路线,不代表当前已实现 |
|
||||
| [current-mvp-architecture.md](current-mvp-architecture.md) | 系统分层、请求主链、Trace/Reasoning 边界与 API surface |
|
||||
| [agent-orchestration.md](agent-orchestration.md) | 单 Diagnosis ReAct Agent 的职责、执行方式和 reasoning 采集边界 |
|
||||
| [harness-quality-gates.md](harness-quality-gates.md) | Run、Tool、Evidence、Semantic、Release 与 Trace Recorder 门禁 |
|
||||
| [session-trace-lifecycle.md](session-trace-lifecycle.md) | sessionId/runId、SSE、统一 Timeline 和 reasoning audit 生命周期 |
|
||||
|
||||
## 当前架构一句话
|
||||
2026-07-22 前的多角色编排、双入口和旧证据链文档已移动到 `archive/2026-07-22-legacy/`,仅用于历史决策追溯,不代表当前运行时。
|
||||
|
||||
SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,Chat 链路由 Gatekeeper 做引用真实性校验、Verifier 做可推导性判断、Composer 生成最终表达;多轮会话元数据落到 `chat_session`,每次诊断运行落到 `diagnosis_run`,步骤和工具明细通过 `agent_step.run_id`、`tool_invocation.run_id` 关联,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
|
||||
|
||||
## 阅读顺序
|
||||
|
||||
1. 先读 [current-mvp-architecture.md](current-mvp-architecture.md),理解系统边界和主链路。
|
||||
2. 面试前读 [interview-one-pager.md](interview-one-pager.md),准备 2-5 分钟讲解。
|
||||
3. 再读 [agent-orchestration.md](agent-orchestration.md),理解当前 Agent 如何协作。
|
||||
4. 接着读 [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md),理解 Chat 证据链路的数据结构和验真边界。
|
||||
5. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。
|
||||
6. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。
|
||||
7. 继续读 [modular-rag-pipeline.md](modular-rag-pipeline.md),看 `lookup_knowledge` 的模块化落地和 evidence-first contract。
|
||||
8. 再读 [rag-eval-closure.md](rag-eval-closure.md),看 RAG baseline 如何形成质量闭环。
|
||||
9. 然后读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
|
||||
10. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
|
||||
11. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
|
||||
12. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
|
||||
13. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
|
||||
当前普通 Trace 与 Provider reasoning 审计使用独立存储和独立接口。Reasoning 访问控制、保留期限、加密要求以及真实 Provider/V015 验证仍由 ISS-015 跟踪,不能把“数据已分表”理解为“治理已经完成”。
|
||||
|
||||
@@ -1,237 +1,69 @@
|
||||
# Agent 编排架构
|
||||
# Diagnosis Agent 执行架构
|
||||
|
||||
**更新日期**:2026-07-08
|
||||
**更新日期**:2026-07-23
|
||||
**状态**:当前可运行架构
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
||||
|
||||
## 1. 设计定位
|
||||
## 1. 单 Agent 原则
|
||||
|
||||
旧版 Agent 架构把系统描述为 Supervisor、Planner、SubAgent、Verifier 的团队协作。当前 MVP 保留这个核心思想,但实现更收敛:
|
||||
当前业务诊断只有一个 `Diagnosis Agent`。它使用框架 `ReactAgent` 完成规划、行动、观察和最终 Draft,但项目不在外层复制 ReAct 状态机,也不使用业务 Graph 或多角色协作链。
|
||||
|
||||
- Chat 链路使用固定顺序工作流:`Planner -> Executor -> Gatekeeper -> Verifier -> Composer`。
|
||||
- AIOps 链路使用 `SupervisorAgent` 调度 `Planner + Executor`,最终由规则评估器做轻量验证。
|
||||
- 当前没有拆分 ExternalApiSubAgent、InternalErrorSubAgent、DatabaseSubAgent;这些作为后续演进方向保留。
|
||||
- 证据工具不直接散落在各个 Agent 里,而是通过 Spring AI ToolCallback / `@Tool` 统一暴露。
|
||||
## 2. 职责
|
||||
|
||||
## 2. 当前 Agent 全景
|
||||
Diagnosis Agent:
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph Chat["Chat diagnosis"]
|
||||
ChatIn["POST /api/chat"] --> ChatService["ChatService"]
|
||||
ChatService --> ChatPlanner["chat_planner"]
|
||||
ChatPlanner --> ChatExecutor["chat_executor"]
|
||||
ChatExecutor --> ChatTools["evidence tools"]
|
||||
ChatTools --> ChatExecutor
|
||||
ChatExecutor --> ChatGatekeeper["ExecutorGatekeeperService"]
|
||||
ChatGatekeeper --> ChatVerifier["chat_verifier"]
|
||||
ChatVerifier --> ChatDecision{"PASS / LOW_CONFID / REJECT"}
|
||||
ChatDecision --> ChatComposer["chat_composer"]
|
||||
ChatComposer --> ChatAnswer["final answer"]
|
||||
end
|
||||
- 接收当前 query 与可选、受限的安全 PreviousTurn。
|
||||
- 自主选择只读 evidence Tool。
|
||||
- 根据 Agent projection 判断是否需要继续查询。
|
||||
- 输出结构化 `DiagnosisDraft`,每条 analysis 绑定 framework `tool_call_id`。
|
||||
- 证据不足时明确限制,不补造事实。
|
||||
|
||||
subgraph AiOps["AIOps diagnosis"]
|
||||
AiOpsIn["POST /api/ai_ops"] --> AiOpsService["AiOpsService"]
|
||||
AiOpsService --> Supervisor["ai_ops_supervisor"]
|
||||
Supervisor --> AiOpsPlanner["planner_agent"]
|
||||
Supervisor --> AiOpsExecutor["executor_agent"]
|
||||
AiOpsPlanner --> AiOpsExecutor
|
||||
AiOpsExecutor --> AiOpsTools["Prometheus / logs / lookup_knowledge"]
|
||||
AiOpsTools --> AiOpsReport["alert report"]
|
||||
AiOpsReport --> AiOpsRule["AiOpsRuleEvaluationService"]
|
||||
end
|
||||
Diagnosis Agent 不负责:
|
||||
|
||||
subgraph Trace["Trace persistence"]
|
||||
ChatSession["chat_session"]
|
||||
Run["diagnosis_run"]
|
||||
Step["agent_step"]
|
||||
Invocation["tool_invocation"]
|
||||
SelfEval["self_evaluation"]
|
||||
end
|
||||
- HTTP/SSE、Session/Run 生命周期和持久化。
|
||||
- 模型/Tool/Token/timeout/cancel 预算。
|
||||
- Tool 参数授权、raw response 投影或证据物理验真。
|
||||
- SemanticGuard 与最终发布决定。
|
||||
|
||||
ChatService --> ChatSession
|
||||
ChatService --> Run
|
||||
ChatPlanner --> Step
|
||||
ChatExecutor --> Step
|
||||
ChatGatekeeper --> SelfEval
|
||||
ChatVerifier --> Step
|
||||
ChatTools --> Invocation
|
||||
ChatDecision --> SelfEval
|
||||
ChatComposer --> Step
|
||||
|
||||
AiOpsService --> ChatSession
|
||||
AiOpsService --> Run
|
||||
AiOpsPlanner --> Step
|
||||
AiOpsExecutor --> Step
|
||||
AiOpsTools --> Invocation
|
||||
AiOpsRule --> SelfEval
|
||||
```
|
||||
|
||||
## 3. Chat 编排
|
||||
|
||||
Chat 复杂诊断采用 `SequentialAgent`,顺序固定:
|
||||
|
||||
```text
|
||||
chat_planner
|
||||
-> chat_executor
|
||||
-> lookup_knowledge / query_logs / query_metrics / date_time
|
||||
-> outputs executor_evidence_v2
|
||||
-> VerifierInputHook / ExecutorGatekeeperService
|
||||
-> validates source_invocation_id / raw_path / evidence_excerpt
|
||||
-> chat_verifier
|
||||
-> judges whether verified evidence can derive claims
|
||||
-> chat_composer
|
||||
-> writes final user-facing answer
|
||||
```
|
||||
|
||||
关键行为:
|
||||
|
||||
| 角色 | 当前职责 | 输出 |
|
||||
|---|---|---|
|
||||
| `chat_planner` | 拆解问题,注入知识域地图和对话历史,给出排查方向 | `planner_plan` |
|
||||
| `chat_executor` | 按计划调用证据工具,抽取带 `source_invocation_id + raw_path + evidence_excerpt` 的微观事实 | `executor_evidence_v2` |
|
||||
| `ExecutorGatekeeperService` | 在 Verifier 前做代码级引用验真,拒绝伪造 ID、错配 raw_path、错配 excerpt | `gatekeeper_result` |
|
||||
| `chat_verifier` | 只判断已验真 evidence excerpt 是否能推出 claim,不做新检索 | `verifier_output` |
|
||||
| `chat_composer` | 只表达 Verifier 允许输出的 claims、缺口和建议,生成最终用户答复 | `composer_output` |
|
||||
|
||||
Chat 链路最多支持两轮验证:
|
||||
## 3. 执行序列
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant C as ChatService
|
||||
participant P as chat_planner
|
||||
participant E as chat_executor
|
||||
participant T as tools
|
||||
participant G as gatekeeper
|
||||
participant V as chat_verifier
|
||||
participant M as chat_composer
|
||||
participant R as diagnosis_run
|
||||
participant App as Chat Application
|
||||
participant Core as Harness Core
|
||||
participant Agent as Diagnosis Agent
|
||||
participant Tool as ACI Tool Boundary
|
||||
participant Audit as Audit Hook / Trace Recorder
|
||||
participant EG as EvidenceGuard
|
||||
participant SG as SemanticGuard
|
||||
participant Release as Release Policy
|
||||
|
||||
C->>P: 原始问题 + history + retry_context
|
||||
P-->>C: planner_plan
|
||||
C->>E: planner_plan + 上下文
|
||||
E->>T: 调用证据工具
|
||||
T-->>E: 证据结果
|
||||
E-->>C: executor_evidence_v2
|
||||
C->>G: executor_structured_output + tool_invocation.evidence_refs
|
||||
G-->>C: gatekeeper_result
|
||||
C->>V: executor_structured_output + gatekeeper_result + tool_trace_summary
|
||||
V-->>C: PASS / LOW_CONFID / REJECT
|
||||
C->>R: 写入 verifier_evaluation
|
||||
alt LOW_CONFID 且允许补证据
|
||||
C->>P: retry_context: 仅补缺失证据
|
||||
else PASS 或 REJECT
|
||||
C->>M: allowed_claims + missing_info + recommended_actions
|
||||
M-->>C: composer_output
|
||||
C->>R: 保存 Composer 最终 answer
|
||||
end
|
||||
App->>Core: start RunContext
|
||||
App->>Agent: query + safe previous_turn
|
||||
Agent->>Tool: tool name + framework tool_call_id + typed args
|
||||
Tool-->>Agent: bounded agent_result
|
||||
Tool->>Audit: bounded Tool lifecycle metadata
|
||||
Agent->>Audit: step metadata + Provider reasoning availability
|
||||
Agent-->>App: DiagnosisDraft
|
||||
App->>EG: Draft + current Run canonical invocations
|
||||
EG-->>App: verified snapshot or deterministic failure
|
||||
App->>SG: query + full Draft + verified snapshot
|
||||
SG-->>App: SUPPORTED / UNSUPPORTED
|
||||
App->>Release: decide public content
|
||||
Release-->>App: report or fixed fallback
|
||||
```
|
||||
|
||||
决策语义:
|
||||
## 4. PreviousTurn
|
||||
|
||||
| Verdict | 行为 |
|
||||
|---|---|
|
||||
| `PASS` | 把 Verifier 允许表达的 claims 交给 Composer 输出 |
|
||||
| `LOW_CONFID` | 如果分数低于阈值且仍有轮次,构造 `retry_context` 补证据;否则输出低置信提示 |
|
||||
| `REJECT` | 输出降级答复,只保留已确认信息和下一步建议 |
|
||||
PreviousTurn 只来自同一 Session 最近一个 `DIAGNOSIS + SUCCESS + published_result`。Fallback、失败、取消、raw evidence 和完整历史都不能进入下一轮;字段与字节上限由 Harness 配置控制。
|
||||
|
||||
## 4. AIOps 编排
|
||||
## 5. Provider Reasoning 审计
|
||||
|
||||
AIOps 使用 `SupervisorAgent` 调度两个子 Agent:
|
||||
`HarnessAgentAuditHook` 在每次模型步骤结束后检查 `AssistantMessage` metadata。当前识别 `reasoning_content`、`reasoningContent`、`reasoning` 和 `thinking`,但只接受 Provider 实际返回的非空文本:
|
||||
|
||||
```text
|
||||
ai_ops_supervisor
|
||||
-> planner_agent
|
||||
-> executor_agent
|
||||
-> final report
|
||||
-> AiOpsRuleEvaluationService
|
||||
```
|
||||
- 有内容时写入 `agent_reasoning_audit`,单条最多保留 32000 个字符,并记录 UTF-8 `content_bytes`。
|
||||
- 无内容时写入 `reasoning_available=false`、`reasoning_content=NULL`、`content_bytes=0`,不得根据最终回答反推或生成 reasoning。
|
||||
- `agent_step.thought` 始终为空;步骤表只记录 message count、roles、是否有文本、Tool names、reasoning availability 和字节数等 metadata。
|
||||
- 普通 `diagnosis_trace_event` 的 `AGENT_MODEL_STEP` 只记录 reasoning availability/bytes,不保存 reasoning 原文。
|
||||
- Reasoning 只用于受限审计,不进入 Agent 后续上下文,不参与 EvidenceGuard、SemanticGuard 或 Release Policy 的事实判断。
|
||||
|
||||
与 Chat 的差异:
|
||||
|
||||
- AIOps 的输入可能是结构化告警 payload。
|
||||
- payload 模式会进入 `PAYLOAD_TARGETED`,最终报告必须聚焦输入告警。
|
||||
- 无 payload 时进入 `AUTO_DISCOVERY`,先通过告警工具发现活跃告警。
|
||||
- 当前 AIOps 不使用 LLM Verifier,而使用轻量规则评估器写入 `self_evaluation.aiops_rule_evaluation`。
|
||||
|
||||
## 5. 工具边界
|
||||
|
||||
当前 Executor 可用工具来自两类:
|
||||
|
||||
```text
|
||||
methodTools
|
||||
-> dateTimeTools
|
||||
-> lookupKnowledgeTool
|
||||
-> queryMetricsTools
|
||||
-> queryLogsTools when mock enabled
|
||||
|
||||
ToolCallbackProvider
|
||||
-> framework-discovered tools
|
||||
```
|
||||
|
||||
工具调用必须写入 `tool_invocation`。其中 `lookup_knowledge` 额外记录:
|
||||
|
||||
- L0/L1 命中数量。
|
||||
- 检索层。
|
||||
- relevance level。
|
||||
- retrieved domains。
|
||||
- dedup reason。
|
||||
|
||||
## 6. Skill / Playbook 流程
|
||||
|
||||
当前 Skill 是诊断流程编排提示,不是事实证据来源。Planner 只能看到 `SkillRegistry.listAll()` 暴露的 name/description 元数据;Executor 才能通过 Spring AI Alibaba 官方 `SkillsAgentHook` 使用 `read_skill` 读取完整 `SKILL.md`。
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Registry["SkillRegistry<br/>active skill metadata"] --> PlannerHook["PlannerSkillMetadataHook"]
|
||||
PlannerHook --> Planner["Planner<br/>metadata only"]
|
||||
Planner --> Plan["planner_plan<br/>selected_skill + steps"]
|
||||
|
||||
Registry --> ExecutorHook["SkillsAgentHook"]
|
||||
ExecutorHook --> ReadSkill["read_skill"]
|
||||
Plan --> Executor["Executor"]
|
||||
Executor --> ReadSkill
|
||||
ReadSkill --> SkillBody["SKILL.md workflow"]
|
||||
SkillBody --> Executor
|
||||
Executor --> EvidenceTools["lookup_knowledge / logs / metrics"]
|
||||
EvidenceTools --> ToolTrace["tool_invocation evidence"]
|
||||
Executor --> Gatekeeper["Gatekeeper"]
|
||||
Gatekeeper --> Verifier["Verifier"]
|
||||
ToolTrace --> Verifier
|
||||
Verifier --> Composer["Composer"]
|
||||
```
|
||||
|
||||
| 角色 | Skill 可见性 | 工具权限 |
|
||||
|---|---|---|
|
||||
| Planner | 只看 skill name / description,并输出 `selected_skill` | 不暴露 `read_skill` |
|
||||
| Executor | 读取 Planner 选中的 skill 正文 | 暴露官方 `read_skill` 和证据工具 |
|
||||
| Gatekeeper | 不看 skill catalog,也不读 skill 正文 | 只读取 Executor 输出和 `tool_invocation.retrieval_details.evidence_refs` |
|
||||
| Verifier | 不看 skill catalog,也不读 skill 正文 | 只读取 Gatekeeper 结果、结构化 claims 和 trace summary |
|
||||
| Composer | 不看 skill catalog,也不读 skill 正文 | 只读取 Verifier 允许表达的内容 |
|
||||
|
||||
## 7. 与旧版设计的差异
|
||||
|
||||
| 旧版设想 | 当前实现 |
|
||||
|---|---|
|
||||
| Supervisor + Planner + 多个专科 SubAgent + Verifier | Chat: Planner + Executor + Gatekeeper + Verifier + Composer;AIOps: Supervisor + Planner + Executor |
|
||||
| ExternalApiSubAgent / InternalErrorSubAgent / DatabaseSubAgent | 暂未拆分,能力通过通用 Executor + 工具 + Prompt 约束实现 |
|
||||
| 每个 SubAgent 专属工具集 | 当前 Executor 持有统一证据工具集合 |
|
||||
| Verifier 支持 PASS / REVISE / REJECT | 当前 Chat Verifier 输出 PASS / LOW_CONFID / REJECT |
|
||||
| Skill 驱动不同诊断流程 | 当前以 Planner 元数据选择 + Executor 读取 playbook 的方式接入 |
|
||||
|
||||
## 8. 后续演进
|
||||
|
||||
当诊断场景和工具复杂度继续上升时,再考虑拆分:
|
||||
|
||||
- `ExternalApiSubAgent`:接口文档、错误码、请求参数、第三方日志。
|
||||
- `DatabaseSubAgent`:连接池、慢 SQL、死锁、索引建议。
|
||||
- `CacheSubAgent`:Redis 超时、连接、热点 key、内存风险。
|
||||
- `GenericDiagnosisSubAgent`:专项 Agent 失败后的兜底。
|
||||
|
||||
拆分前提:
|
||||
|
||||
- 当前 Executor prompt 已难以维护。
|
||||
- 不同故障类型的工具权限明显不同。
|
||||
- Trace 能证明某类问题需要独立的推理策略。
|
||||
- 评测集能覆盖拆分前后的行为差异。
|
||||
当前查询隔离已经实现,完整访问治理和真实 Provider 行为验证仍属于 ISS-015。
|
||||
|
||||
@@ -209,7 +209,7 @@ Deferred future enhancements:
|
||||
|
||||
## 10. Supporting Materials
|
||||
|
||||
- `mvp/issues/rag-refactor-plan.md`
|
||||
- `mvp/issues/archived/rag-refactor-plan.md`
|
||||
- `eval/rag-retrieval/README.md`
|
||||
- `scripts/eval_rag_live_acceptance.py`
|
||||
- `interview/rag-refactor-story.md`
|
||||
|
||||
@@ -0,0 +1,32 @@
|
||||
# MVP 架构归档(2026-07-22)
|
||||
|
||||
**状态**:历史快照,禁止作为当前实现依据
|
||||
|
||||
本目录保存切换到单 Diagnosis Agent + Harness 之前的架构文档,包括多角色编排、Chat/AIOps 双入口、Executor/Verifier 证据链、旧 RAG 设计和旧数据模型。
|
||||
|
||||
这些文档保留原有内容和相对链接,用于追溯当时的设计背景。当前 runtime、API、数据模型和验收口径请从 [../../../README.md](../../../README.md) 进入。
|
||||
|
||||
## 归档内容
|
||||
|
||||
| 文档 | 用途 |
|
||||
|---|---|
|
||||
| [current-mvp-architecture.md](current-mvp-architecture.md) | 归档时的整体架构快照 |
|
||||
| [interview-one-pager.md](interview-one-pager.md) | 旧架构讲解材料 |
|
||||
| [agent-orchestration.md](agent-orchestration.md) | Chat SequentialAgent 与 AIOps SupervisorAgent 编排 |
|
||||
| [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md) | Executor、Gatekeeper、Verifier 与 Composer 证据链 |
|
||||
| [harness-quality-gates.md](harness-quality-gates.md) | 旧多角色 Harness 质量门禁 |
|
||||
| [rag-architecture.md](rag-architecture.md) | 旧 RAG 总体架构 |
|
||||
| [modular-rag-pipeline.md](modular-rag-pipeline.md) | 旧模块化 RAG 方案 |
|
||||
| [rag-eval-closure.md](rag-eval-closure.md) | 旧 RAG 评测闭环 |
|
||||
| [retrieval-observability.md](retrieval-observability.md) | 旧检索可观测性设计 |
|
||||
| [feedback-architecture.md](feedback-architecture.md) | 旧反馈与自评估设计 |
|
||||
| [session-trace-lifecycle.md](session-trace-lifecycle.md) | 旧会话和 Trace 生命周期 |
|
||||
| [knowledge-base-authoring.md](knowledge-base-authoring.md) | 知识库编写规范快照 |
|
||||
| [data-model.md](data-model.md) | 旧数据模型总览 |
|
||||
| [evolution-roadmap.md](evolution-roadmap.md) | 旧架构演进路线 |
|
||||
|
||||
## 归档边界
|
||||
|
||||
- 不根据本目录新增或修改运行时代码。
|
||||
- 不将本目录描述的接口和表结构视为当前契约。
|
||||
- 如需引用历史决策,应同时说明其归档日期和当前替代方案。
|
||||
@@ -0,0 +1,3 @@
|
||||
# Archive Note
|
||||
|
||||
本目录保存 2026-07-22 单 Diagnosis Agent + Harness 切换前的当前架构文档。内容用于历史决策追溯,不代表现行 runtime、API 或验收口径。
|
||||
@@ -0,0 +1,237 @@
|
||||
# Agent 编排架构
|
||||
|
||||
**更新日期**:2026-07-08
|
||||
**状态**:当前可运行架构
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
||||
|
||||
## 1. 设计定位
|
||||
|
||||
旧版 Agent 架构把系统描述为 Supervisor、Planner、SubAgent、Verifier 的团队协作。当前 MVP 保留这个核心思想,但实现更收敛:
|
||||
|
||||
- Chat 链路使用固定顺序工作流:`Planner -> Executor -> Gatekeeper -> Verifier -> Composer`。
|
||||
- AIOps 链路使用 `SupervisorAgent` 调度 `Planner + Executor`,最终由规则评估器做轻量验证。
|
||||
- 当前没有拆分 ExternalApiSubAgent、InternalErrorSubAgent、DatabaseSubAgent;这些作为后续演进方向保留。
|
||||
- 证据工具不直接散落在各个 Agent 里,而是通过 Spring AI ToolCallback / `@Tool` 统一暴露。
|
||||
|
||||
## 2. 当前 Agent 全景
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph Chat["Chat diagnosis"]
|
||||
ChatIn["POST /api/chat"] --> ChatService["ChatService"]
|
||||
ChatService --> ChatPlanner["chat_planner"]
|
||||
ChatPlanner --> ChatExecutor["chat_executor"]
|
||||
ChatExecutor --> ChatTools["evidence tools"]
|
||||
ChatTools --> ChatExecutor
|
||||
ChatExecutor --> ChatGatekeeper["ExecutorGatekeeperService"]
|
||||
ChatGatekeeper --> ChatVerifier["chat_verifier"]
|
||||
ChatVerifier --> ChatDecision{"PASS / LOW_CONFID / REJECT"}
|
||||
ChatDecision --> ChatComposer["chat_composer"]
|
||||
ChatComposer --> ChatAnswer["final answer"]
|
||||
end
|
||||
|
||||
subgraph AiOps["AIOps diagnosis"]
|
||||
AiOpsIn["POST /api/ai_ops"] --> AiOpsService["AiOpsService"]
|
||||
AiOpsService --> Supervisor["ai_ops_supervisor"]
|
||||
Supervisor --> AiOpsPlanner["planner_agent"]
|
||||
Supervisor --> AiOpsExecutor["executor_agent"]
|
||||
AiOpsPlanner --> AiOpsExecutor
|
||||
AiOpsExecutor --> AiOpsTools["Prometheus / logs / lookup_knowledge"]
|
||||
AiOpsTools --> AiOpsReport["alert report"]
|
||||
AiOpsReport --> AiOpsRule["AiOpsRuleEvaluationService"]
|
||||
end
|
||||
|
||||
subgraph Trace["Trace persistence"]
|
||||
ChatSession["chat_session"]
|
||||
Run["diagnosis_run"]
|
||||
Step["agent_step"]
|
||||
Invocation["tool_invocation"]
|
||||
SelfEval["self_evaluation"]
|
||||
end
|
||||
|
||||
ChatService --> ChatSession
|
||||
ChatService --> Run
|
||||
ChatPlanner --> Step
|
||||
ChatExecutor --> Step
|
||||
ChatGatekeeper --> SelfEval
|
||||
ChatVerifier --> Step
|
||||
ChatTools --> Invocation
|
||||
ChatDecision --> SelfEval
|
||||
ChatComposer --> Step
|
||||
|
||||
AiOpsService --> ChatSession
|
||||
AiOpsService --> Run
|
||||
AiOpsPlanner --> Step
|
||||
AiOpsExecutor --> Step
|
||||
AiOpsTools --> Invocation
|
||||
AiOpsRule --> SelfEval
|
||||
```
|
||||
|
||||
## 3. Chat 编排
|
||||
|
||||
Chat 复杂诊断采用 `SequentialAgent`,顺序固定:
|
||||
|
||||
```text
|
||||
chat_planner
|
||||
-> chat_executor
|
||||
-> lookup_knowledge / query_logs / query_metrics / date_time
|
||||
-> outputs executor_evidence_v2
|
||||
-> VerifierInputHook / ExecutorGatekeeperService
|
||||
-> validates source_invocation_id / raw_path / evidence_excerpt
|
||||
-> chat_verifier
|
||||
-> judges whether verified evidence can derive claims
|
||||
-> chat_composer
|
||||
-> writes final user-facing answer
|
||||
```
|
||||
|
||||
关键行为:
|
||||
|
||||
| 角色 | 当前职责 | 输出 |
|
||||
|---|---|---|
|
||||
| `chat_planner` | 拆解问题,注入知识域地图和对话历史,给出排查方向 | `planner_plan` |
|
||||
| `chat_executor` | 按计划调用证据工具,抽取带 `source_invocation_id + raw_path + evidence_excerpt` 的微观事实 | `executor_evidence_v2` |
|
||||
| `ExecutorGatekeeperService` | 在 Verifier 前做代码级引用验真,拒绝伪造 ID、错配 raw_path、错配 excerpt | `gatekeeper_result` |
|
||||
| `chat_verifier` | 只判断已验真 evidence excerpt 是否能推出 claim,不做新检索 | `verifier_output` |
|
||||
| `chat_composer` | 只表达 Verifier 允许输出的 claims、缺口和建议,生成最终用户答复 | `composer_output` |
|
||||
|
||||
Chat 链路最多支持两轮验证:
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant C as ChatService
|
||||
participant P as chat_planner
|
||||
participant E as chat_executor
|
||||
participant T as tools
|
||||
participant G as gatekeeper
|
||||
participant V as chat_verifier
|
||||
participant M as chat_composer
|
||||
participant R as diagnosis_run
|
||||
|
||||
C->>P: 原始问题 + history + retry_context
|
||||
P-->>C: planner_plan
|
||||
C->>E: planner_plan + 上下文
|
||||
E->>T: 调用证据工具
|
||||
T-->>E: 证据结果
|
||||
E-->>C: executor_evidence_v2
|
||||
C->>G: executor_structured_output + tool_invocation.evidence_refs
|
||||
G-->>C: gatekeeper_result
|
||||
C->>V: executor_structured_output + gatekeeper_result + tool_trace_summary
|
||||
V-->>C: PASS / LOW_CONFID / REJECT
|
||||
C->>R: 写入 verifier_evaluation
|
||||
alt LOW_CONFID 且允许补证据
|
||||
C->>P: retry_context: 仅补缺失证据
|
||||
else PASS 或 REJECT
|
||||
C->>M: allowed_claims + missing_info + recommended_actions
|
||||
M-->>C: composer_output
|
||||
C->>R: 保存 Composer 最终 answer
|
||||
end
|
||||
```
|
||||
|
||||
决策语义:
|
||||
|
||||
| Verdict | 行为 |
|
||||
|---|---|
|
||||
| `PASS` | 把 Verifier 允许表达的 claims 交给 Composer 输出 |
|
||||
| `LOW_CONFID` | 如果分数低于阈值且仍有轮次,构造 `retry_context` 补证据;否则输出低置信提示 |
|
||||
| `REJECT` | 输出降级答复,只保留已确认信息和下一步建议 |
|
||||
|
||||
## 4. AIOps 编排
|
||||
|
||||
AIOps 使用 `SupervisorAgent` 调度两个子 Agent:
|
||||
|
||||
```text
|
||||
ai_ops_supervisor
|
||||
-> planner_agent
|
||||
-> executor_agent
|
||||
-> final report
|
||||
-> AiOpsRuleEvaluationService
|
||||
```
|
||||
|
||||
与 Chat 的差异:
|
||||
|
||||
- AIOps 的输入可能是结构化告警 payload。
|
||||
- payload 模式会进入 `PAYLOAD_TARGETED`,最终报告必须聚焦输入告警。
|
||||
- 无 payload 时进入 `AUTO_DISCOVERY`,先通过告警工具发现活跃告警。
|
||||
- 当前 AIOps 不使用 LLM Verifier,而使用轻量规则评估器写入 `self_evaluation.aiops_rule_evaluation`。
|
||||
|
||||
## 5. 工具边界
|
||||
|
||||
当前 Executor 可用工具来自两类:
|
||||
|
||||
```text
|
||||
methodTools
|
||||
-> dateTimeTools
|
||||
-> lookupKnowledgeTool
|
||||
-> queryMetricsTools
|
||||
-> queryLogsTools when mock enabled
|
||||
|
||||
ToolCallbackProvider
|
||||
-> framework-discovered tools
|
||||
```
|
||||
|
||||
工具调用必须写入 `tool_invocation`。其中 `lookup_knowledge` 额外记录:
|
||||
|
||||
- L0/L1 命中数量。
|
||||
- 检索层。
|
||||
- relevance level。
|
||||
- retrieved domains。
|
||||
- dedup reason。
|
||||
|
||||
## 6. Skill / Playbook 流程
|
||||
|
||||
当前 Skill 是诊断流程编排提示,不是事实证据来源。Planner 只能看到 `SkillRegistry.listAll()` 暴露的 name/description 元数据;Executor 才能通过 Spring AI Alibaba 官方 `SkillsAgentHook` 使用 `read_skill` 读取完整 `SKILL.md`。
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Registry["SkillRegistry<br/>active skill metadata"] --> PlannerHook["PlannerSkillMetadataHook"]
|
||||
PlannerHook --> Planner["Planner<br/>metadata only"]
|
||||
Planner --> Plan["planner_plan<br/>selected_skill + steps"]
|
||||
|
||||
Registry --> ExecutorHook["SkillsAgentHook"]
|
||||
ExecutorHook --> ReadSkill["read_skill"]
|
||||
Plan --> Executor["Executor"]
|
||||
Executor --> ReadSkill
|
||||
ReadSkill --> SkillBody["SKILL.md workflow"]
|
||||
SkillBody --> Executor
|
||||
Executor --> EvidenceTools["lookup_knowledge / logs / metrics"]
|
||||
EvidenceTools --> ToolTrace["tool_invocation evidence"]
|
||||
Executor --> Gatekeeper["Gatekeeper"]
|
||||
Gatekeeper --> Verifier["Verifier"]
|
||||
ToolTrace --> Verifier
|
||||
Verifier --> Composer["Composer"]
|
||||
```
|
||||
|
||||
| 角色 | Skill 可见性 | 工具权限 |
|
||||
|---|---|---|
|
||||
| Planner | 只看 skill name / description,并输出 `selected_skill` | 不暴露 `read_skill` |
|
||||
| Executor | 读取 Planner 选中的 skill 正文 | 暴露官方 `read_skill` 和证据工具 |
|
||||
| Gatekeeper | 不看 skill catalog,也不读 skill 正文 | 只读取 Executor 输出和 `tool_invocation.retrieval_details.evidence_refs` |
|
||||
| Verifier | 不看 skill catalog,也不读 skill 正文 | 只读取 Gatekeeper 结果、结构化 claims 和 trace summary |
|
||||
| Composer | 不看 skill catalog,也不读 skill 正文 | 只读取 Verifier 允许表达的内容 |
|
||||
|
||||
## 7. 与旧版设计的差异
|
||||
|
||||
| 旧版设想 | 当前实现 |
|
||||
|---|---|
|
||||
| Supervisor + Planner + 多个专科 SubAgent + Verifier | Chat: Planner + Executor + Gatekeeper + Verifier + Composer;AIOps: Supervisor + Planner + Executor |
|
||||
| ExternalApiSubAgent / InternalErrorSubAgent / DatabaseSubAgent | 暂未拆分,能力通过通用 Executor + 工具 + Prompt 约束实现 |
|
||||
| 每个 SubAgent 专属工具集 | 当前 Executor 持有统一证据工具集合 |
|
||||
| Verifier 支持 PASS / REVISE / REJECT | 当前 Chat Verifier 输出 PASS / LOW_CONFID / REJECT |
|
||||
| Skill 驱动不同诊断流程 | 当前以 Planner 元数据选择 + Executor 读取 playbook 的方式接入 |
|
||||
|
||||
## 8. 后续演进
|
||||
|
||||
当诊断场景和工具复杂度继续上升时,再考虑拆分:
|
||||
|
||||
- `ExternalApiSubAgent`:接口文档、错误码、请求参数、第三方日志。
|
||||
- `DatabaseSubAgent`:连接池、慢 SQL、死锁、索引建议。
|
||||
- `CacheSubAgent`:Redis 超时、连接、热点 key、内存风险。
|
||||
- `GenericDiagnosisSubAgent`:专项 Agent 失败后的兜底。
|
||||
|
||||
拆分前提:
|
||||
|
||||
- 当前 Executor prompt 已难以维护。
|
||||
- 不同故障类型的工具权限明显不同。
|
||||
- Trace 能证明某类问题需要独立的推理策略。
|
||||
- 评测集能覆盖拆分前后的行为差异。
|
||||
@@ -0,0 +1,444 @@
|
||||
# 当前 MVP 架构
|
||||
|
||||
**更新日期**:2026-07-08
|
||||
**状态**:当前可运行架构
|
||||
**适用范围**:Demo、面试讲解、后续迭代规划
|
||||
|
||||
## 1. 系统定位
|
||||
|
||||
SuperBizAgent MVP 不是通用 Chatbot,而是面向故障诊断的 Agent 工程项目。
|
||||
|
||||
核心目标:
|
||||
|
||||
- 支持用户主动发起的 Chat 诊断。
|
||||
- 支持 AIOps 告警触发的自动诊断。
|
||||
- 保留 Agent 的规划、执行、验证过程。
|
||||
- 工具调用必须显式、可追踪、可回放。
|
||||
- RAG 检索必须通过 `lookup_knowledge` 暴露证据链。
|
||||
- 每次诊断都沉淀 session、step、tool invocation 和 self evaluation。
|
||||
|
||||
## 2. 总体分层
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph API["API Layer"]
|
||||
ChatController["ChatController"]
|
||||
TraceController["DiagnosisTraceController"]
|
||||
SearchController["SearchController"]
|
||||
DocumentController["DocumentController"]
|
||||
end
|
||||
|
||||
subgraph App["Application Service"]
|
||||
ChatService["ChatService"]
|
||||
AiOpsService["AiOpsService"]
|
||||
TraceService["DiagnosisTraceService"]
|
||||
end
|
||||
|
||||
subgraph Agent["Agent Orchestration"]
|
||||
Supervisor["Supervisor"]
|
||||
Planner["Planner"]
|
||||
Executor["Executor"]
|
||||
Gatekeeper["Gatekeeper"]
|
||||
Verifier["Verifier"]
|
||||
Composer["Composer"]
|
||||
end
|
||||
|
||||
subgraph Tools["Evidence Tools"]
|
||||
KnowledgeTool["lookup_knowledge"]
|
||||
LogsTool["query_logs"]
|
||||
MetricsTool["query_metrics"]
|
||||
AlertsTool["queryPrometheusAlerts"]
|
||||
end
|
||||
|
||||
subgraph Skills["Skill / Playbook"]
|
||||
SkillRegistry["SkillRegistry"]
|
||||
PlannerSkillHook["PlannerSkillMetadataHook"]
|
||||
SkillsHook["SkillsAgentHook"]
|
||||
ReadSkill["read_skill"]
|
||||
end
|
||||
|
||||
subgraph RAG["RAG Retrieval"]
|
||||
L0["KnowledgeIndexService"]
|
||||
VectorSearch["VectorSearchService"]
|
||||
VectorStore["Spring AI VectorStore"]
|
||||
SdkFallback["Milvus SDK fallback"]
|
||||
end
|
||||
|
||||
subgraph Store["Persistence and Trace"]
|
||||
ChatSession["chat_session"]
|
||||
Run["diagnosis_run"]
|
||||
Step["agent_step"]
|
||||
Invocation["tool_invocation"]
|
||||
ApiDoc["api_document"]
|
||||
Milvus["Milvus/Zilliz"]
|
||||
end
|
||||
|
||||
API --> App
|
||||
ChatService --> Agent
|
||||
AiOpsService --> Agent
|
||||
SkillRegistry --> PlannerSkillHook
|
||||
PlannerSkillHook --> Planner
|
||||
SkillRegistry --> SkillsHook
|
||||
SkillsHook --> Executor
|
||||
Executor --> ReadSkill
|
||||
Agent --> Tools
|
||||
KnowledgeTool --> RAG
|
||||
RAG --> Store
|
||||
Tools --> Invocation
|
||||
Agent --> Step
|
||||
App --> Session
|
||||
TraceService --> Session
|
||||
TraceService --> Step
|
||||
TraceService --> Invocation
|
||||
```
|
||||
|
||||
```text
|
||||
API Layer
|
||||
-> ChatController
|
||||
-> DiagnosisTraceController
|
||||
-> SearchController
|
||||
-> DocumentController
|
||||
|
||||
Application Service
|
||||
-> ChatService
|
||||
-> AiOpsService
|
||||
-> DiagnosisTraceService
|
||||
|
||||
Agent Orchestration
|
||||
-> Supervisor
|
||||
-> Planner
|
||||
-> Executor
|
||||
-> Gatekeeper
|
||||
-> Verifier
|
||||
-> Composer
|
||||
|
||||
Evidence Tools
|
||||
-> lookup_knowledge
|
||||
-> query_logs
|
||||
-> query_metrics
|
||||
-> queryPrometheusAlerts
|
||||
|
||||
Skill / Playbook
|
||||
-> SkillRegistry
|
||||
-> PlannerSkillMetadataHook gives Planner name/description only
|
||||
-> SkillsAgentHook gives Executor read_skill
|
||||
-> Verifier is isolated from skills
|
||||
|
||||
RAG Retrieval
|
||||
-> KnowledgeIndexService
|
||||
-> VectorSearchService
|
||||
-> Spring AI VectorStore
|
||||
-> Milvus SDK fallback
|
||||
|
||||
Persistence
|
||||
-> chat_session
|
||||
-> diagnosis_run
|
||||
-> agent_step.run_id
|
||||
-> tool_invocation.run_id
|
||||
-> api_document
|
||||
-> Milvus/Zilliz collection
|
||||
|
||||
Quality Gates
|
||||
-> executor gatekeeper
|
||||
-> chat verifier
|
||||
-> AIOps rule evaluation
|
||||
-> diagnosis eval baseline
|
||||
-> RAG retrieval baseline
|
||||
```
|
||||
|
||||
## 3. Chat 诊断链路
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
actor User as 用户
|
||||
participant API as POST /api/chat
|
||||
participant Chat as ChatService
|
||||
participant Planner as Planner Agent
|
||||
participant Executor as Executor Agent
|
||||
participant Tool as Evidence Tools
|
||||
participant Gatekeeper as Gatekeeper Hook
|
||||
participant Verifier as Verifier Agent
|
||||
participant Composer as Composer Agent
|
||||
participant DB as Trace Tables
|
||||
participant Trace as Trace API
|
||||
|
||||
User->>API: 提交诊断问题
|
||||
API->>Chat: execute chat strategy
|
||||
Chat->>DB: 创建 chat_session metadata + diagnosis_run(runId)
|
||||
Chat->>Planner: 复杂问题进入规划
|
||||
Planner->>DB: 写入 agent_step.run_id
|
||||
Planner->>Executor: 下发排查方向
|
||||
Executor->>Tool: lookup_knowledge / logs / metrics
|
||||
Tool->>DB: 写入 tool_invocation.run_id
|
||||
Tool-->>Executor: 返回证据
|
||||
Executor->>Gatekeeper: 输出 executor_evidence_v2
|
||||
Gatekeeper->>DB: 读取 tool_invocation.evidence_refs 并校验引用
|
||||
Gatekeeper->>Verifier: 传入已验真的 claims / excerpts
|
||||
Verifier->>DB: 合并 diagnosis_run.self_evaluation.verifier_evaluation
|
||||
Verifier->>Composer: 传入 allowed_claims / missing_info / actions
|
||||
Composer->>Chat: 生成最终用户答复
|
||||
Chat->>DB: 保存 diagnosis_run.answer
|
||||
User->>Trace: GET /api/diagnosis/{sessionId}/trace?runId=...
|
||||
Trace->>DB: 聚合 run / step / tool
|
||||
Trace-->>User: 返回可回放诊断链路
|
||||
```
|
||||
|
||||
```text
|
||||
POST /api/chat
|
||||
-> ChatService
|
||||
-> 简单问题:轻量回答
|
||||
-> 复杂诊断:Agent 编排
|
||||
-> Planner 制定排查方向
|
||||
-> Executor 调用证据工具
|
||||
-> lookup_knowledge
|
||||
-> query_logs
|
||||
-> query_metrics
|
||||
-> Gatekeeper 校验 Executor 证据引用真实性
|
||||
-> Verifier 判断 claim 是否能由已核验证据推出
|
||||
-> Composer 生成最终用户答复
|
||||
-> 保存 chat_session metadata
|
||||
-> 保存 diagnosis_run
|
||||
-> 保存 agent_step.run_id
|
||||
-> 保存 tool_invocation.run_id
|
||||
-> 合并 diagnosis_run.self_evaluation.verifier_evaluation
|
||||
```
|
||||
|
||||
Chat 链路的质量门禁由三段组成:Gatekeeper 先做代码级引用验真,Verifier 再做 LLM 可推导性判断,Composer 最后控制对用户的表达边界。Gatekeeper、Verifier、Composer 的输出合并到当前 `diagnosis_run.self_evaluation.verifier_evaluation`,Trace API 会展示该验证结果。
|
||||
|
||||
Agent 编排细节见 [agent-orchestration.md](agent-orchestration.md)。
|
||||
|
||||
关键代码:
|
||||
|
||||
- `src/main/java/com/superbiz/agent/controller/ChatController.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/SelfEvaluationMergeService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
|
||||
|
||||
## 4. AIOps 诊断链路
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Request["POST /api/ai_ops"] --> Payload{"包含告警 payload?"}
|
||||
Payload -->|是| Targeted["PAYLOAD_TARGETED"]
|
||||
Payload -->|否| Discovery["AUTO_DISCOVERY"]
|
||||
|
||||
Targeted --> BuildPrompt["构造聚焦 payload 的诊断 prompt"]
|
||||
Targeted --> QueryAug["生成 recommended lookup_knowledge query"]
|
||||
Discovery --> DiscoverAlert["通过 queryPrometheusAlerts 发现活跃告警"]
|
||||
|
||||
BuildPrompt --> Plan["Planner 规划排查"]
|
||||
QueryAug --> Plan
|
||||
DiscoverAlert --> Plan
|
||||
|
||||
Plan --> Execute["Executor 收集证据"]
|
||||
Execute --> Knowledge["lookup_knowledge"]
|
||||
Execute --> Metrics["query_metrics / Prometheus"]
|
||||
Execute --> Logs["query_logs"]
|
||||
|
||||
Knowledge --> Report["告警分析报告"]
|
||||
Metrics --> Report
|
||||
Logs --> Report
|
||||
|
||||
Report --> RuleEval["AiOpsRuleEvaluationService"]
|
||||
RuleEval --> SelfEval["self_evaluation.aiops_rule_evaluation"]
|
||||
Report --> Trace["DiagnosisTraceService"]
|
||||
SelfEval --> Trace
|
||||
```
|
||||
|
||||
```text
|
||||
POST /api/ai_ops
|
||||
-> AiOpsService
|
||||
-> 判断是否有告警 payload
|
||||
-> PAYLOAD_TARGETED
|
||||
-> AUTO_DISCOVERY
|
||||
-> 构造 AIOps 诊断 prompt
|
||||
-> payload 模式补充 recommended lookup_knowledge query
|
||||
-> Agent 编排
|
||||
-> Planner / Executor
|
||||
-> Prometheus / logs / knowledge tools
|
||||
-> 生成告警分析报告
|
||||
-> AiOpsRuleEvaluationService
|
||||
-> 合并 diagnosis_run.self_evaluation.aiops_rule_evaluation
|
||||
-> Trace API 可查看全链路
|
||||
```
|
||||
|
||||
AIOps 保留两种模式:
|
||||
|
||||
| 模式 | 触发条件 | 行为 |
|
||||
|---|---|---|
|
||||
| `PAYLOAD_TARGETED` | 请求包含 alertName、service、severity、description、timeRange 等字段 | 以 payload 为唯一主诊断对象,并生成推荐知识库 query |
|
||||
| `AUTO_DISCOVERY` | 请求没有明确告警 payload | 先查询当前活跃告警,再选择目标排查 |
|
||||
|
||||
AIOps 当前使用轻量规则验证器,重点检查:
|
||||
|
||||
- 最终报告是否存在。
|
||||
- payload 模式是否聚焦输入告警。
|
||||
- 是否使用关键证据工具,例如 `lookup_knowledge`、日志、指标。
|
||||
|
||||
关键代码:
|
||||
|
||||
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/AiOpsRuleEvaluationService.java`
|
||||
|
||||
## 5. RAG 位置
|
||||
|
||||
RAG 不是隐藏在 Chat Advisor 里的隐式能力,而是 Executor 可以显式调用的工具:
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Executor["Executor Agent"] --> Tool["lookup_knowledge Tool"]
|
||||
Tool --> L0["L0 domain/entity hint"]
|
||||
Tool --> Search["VectorSearchService"]
|
||||
L0 --> Search
|
||||
Search --> VectorStore["Spring AI VectorStore"]
|
||||
Search --> Fallback["Milvus SDK fallback"]
|
||||
VectorStore --> Normalize["score/rawScore/scoreLabel"]
|
||||
Fallback --> Normalize
|
||||
Normalize --> Evidence["evidence output"]
|
||||
Evidence --> Invocation["tool_invocation"]
|
||||
Evidence --> Executor
|
||||
```
|
||||
|
||||
```text
|
||||
Executor
|
||||
-> lookup_knowledge(query)
|
||||
-> L0 domain/entity hint
|
||||
-> VectorSearchService
|
||||
-> Spring AI VectorStore
|
||||
-> Milvus SDK fallback
|
||||
-> evidence shaping
|
||||
-> tool_invocation
|
||||
```
|
||||
|
||||
保留显式工具的原因:
|
||||
|
||||
- Agent 何时检索、检索什么、证据是什么,必须能在 trace 中解释。
|
||||
- AIOps payload 到 query 的业务映射需要项目内控制。
|
||||
- `tool_invocation` 是后续评测、回放和面试讲解的核心材料。
|
||||
|
||||
RAG 总体设计见 [rag-architecture.md](rag-architecture.md),检索运行细节见 [retrieval-observability.md](retrieval-observability.md)。
|
||||
|
||||
## 6. 持久化模型
|
||||
|
||||
当前诊断持久化以 session/run/trace 明细为核心:
|
||||
|
||||
```text
|
||||
chat_session
|
||||
-> 多轮会话目录和元数据
|
||||
-> session_id / status / message_pair_count
|
||||
|
||||
diagnosis_run
|
||||
-> 一次诊断运行的主记录
|
||||
-> run_id / session_id
|
||||
-> query / status / agent_flow / answer
|
||||
-> self_evaluation
|
||||
-> step_count / tool_call_count / duration
|
||||
|
||||
agent_step
|
||||
-> Agent 模型调用步骤
|
||||
-> session_id / run_id
|
||||
-> step_index / agent_name
|
||||
-> model_input / model_output / thought
|
||||
-> duration / token_count
|
||||
|
||||
tool_invocation
|
||||
-> 工具调用事实
|
||||
-> session_id / run_id
|
||||
-> tool_name / input_params / output_preview
|
||||
-> retrieval_layer / retrieval_details
|
||||
-> retrieval_details.evidence_refs
|
||||
-> relevance_level / dedup_reason
|
||||
-> duration / success
|
||||
```
|
||||
|
||||
说明:
|
||||
|
||||
- 旧的 `diagnosis_record` 已不是当前主模型。
|
||||
- `diagnosis_session` 已降级为历史兼容和回滚表,新执行写入 `chat_session + diagnosis_run`。
|
||||
- `api_document` 仍用于文档元数据管理。
|
||||
- 文档向量内容存放在 Milvus/Zilliz collection 中。
|
||||
|
||||
会话和 Trace 生命周期见 [session-trace-lifecycle.md](session-trace-lifecycle.md),完整数据关系见 [data-model.md](data-model.md)。
|
||||
|
||||
## 7. Trace API
|
||||
|
||||
```text
|
||||
GET /api/diagnosis/{sessionId}/trace
|
||||
GET /api/diagnosis/{sessionId}/trace?runId=run-...
|
||||
```
|
||||
|
||||
Trace API 聚合:
|
||||
|
||||
- 会话元数据、运行状态和最终报告。
|
||||
- Agent step 序列。
|
||||
- 工具调用和检索细节。
|
||||
- Chat Gatekeeper / Verifier / Composer 结果。
|
||||
- AIOps rule evaluation 结果。
|
||||
|
||||
Trace 是本项目区别于普通问答系统的关键:答案不是孤立文本,而是可以追溯到 Agent 决策、工具调用和证据来源。
|
||||
|
||||
Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明见 [harness-quality-gates.md](harness-quality-gates.md),用户反馈与 `self_evaluation` 闭环见 [feedback-architecture.md](feedback-architecture.md)。
|
||||
|
||||
## 8. 质量门禁
|
||||
|
||||
当前质量门禁分层如下:
|
||||
|
||||
| 门禁 | 位置 | 作用 |
|
||||
|---|---|---|
|
||||
| Executor Gatekeeper | `VerifierInputHook` / `ExecutorGatekeeperService` | 校验 Executor 引用的 invocation、`raw_path`、`evidence_excerpt` 是否真实 |
|
||||
| Chat Verifier | `ChatService` | 判断已验真证据是否能推出 Executor claims |
|
||||
| Chat Composer | `ChatService` | 只表达 Verifier 允许输出的内容,避免把 no-evidence 说成已排除 |
|
||||
| AIOps Rule Evaluation | `AiOpsRuleEvaluationService` | 校验告警诊断是否聚焦 payload 并使用证据 |
|
||||
| Diagnosis Eval Baseline | `mvp/eval/` | 固化诊断 trace 和报告行为 |
|
||||
| RAG Retrieval Baseline | `eval/rag-retrieval/` | 固化检索召回行为,避免 RAG 重构回退 |
|
||||
| Live RAG Acceptance | `scripts/eval_rag_live_acceptance.py` | 在运行环境中验证重建索引后的真实检索 |
|
||||
|
||||
## 9. 当前完成状态
|
||||
|
||||
已经完成:
|
||||
|
||||
- Chat 和 AIOps 两条入口链路。
|
||||
- 显式 `lookup_knowledge` Agent Tool。
|
||||
- L0 从最终决策降级为 domain/entity hint。
|
||||
- `VectorSearchService` 作为稳定检索门面。
|
||||
- Spring AI VectorStore 读取路径。
|
||||
- Milvus SDK fallback。
|
||||
- `score` / `rawScore` / `scoreLabel` 分数语义拆分。
|
||||
- `title`、`breadcrumb`、`content` 参与 embedding 文本。
|
||||
- `tool_invocation` 记录检索层、relevance level、dedup reason。
|
||||
- Chat verifier 和 AIOps rule evaluation 合并进 `self_evaluation`。
|
||||
- Chat Executor 结构化输出 `executor_evidence_v2`,不再直接承担最终用户答复。
|
||||
- `tool_invocation.retrieval_details.evidence_refs` 支持 `raw_path` 精确引用和 `$.no_evidence` 负向证据。
|
||||
- Gatekeeper 对 Executor 引用做代码级验真,并在审计中记录 `rule_set_version` 和规则元数据摘要。
|
||||
- Verifier 只判断可推导性。
|
||||
- Composer 在 Verifier 之后生成最终用户表达,并限制 negative observation 过度表述。
|
||||
- RAG offline baseline 和 live acceptance 脚本。
|
||||
|
||||
暂不作为当前已完成能力声明:
|
||||
|
||||
- 完整 QueryTransformer / MultiQuery。
|
||||
- BM25、RRF、cross-encoder rerank。
|
||||
- 完整邻居 chunk / section context expansion。
|
||||
- VectorStore 写入路径全面迁移。
|
||||
- 完整 LLM-based AIOps verifier。
|
||||
|
||||
后续 Agent 拆分、Skill/Playbook、MCP 工具协议化和进程隔离等方向见 [evolution-roadmap.md](evolution-roadmap.md)。
|
||||
|
||||
## 10. 关键代码索引
|
||||
|
||||
| 能力 | 代码 |
|
||||
|---|---|
|
||||
| Chat 入口与编排 | `ChatController`, `ChatService` |
|
||||
| AIOps 入口与编排 | `ChatController.aiOps`, `AiOpsService` |
|
||||
| AIOps 规则验证 | `AiOpsRuleEvaluationService` |
|
||||
| 知识库工具 | `LookupKnowledgeTool` |
|
||||
| L0 hint | `KnowledgeIndexService` |
|
||||
| 向量检索门面 | `VectorSearchService` |
|
||||
| 文档切片 | `DocumentChunkService` |
|
||||
| 向量写入 | `VectorIndexService` |
|
||||
| Spring AI VectorStore 配置辅助 | `SpringAiVectorStoreSidecarService` |
|
||||
| Trace 聚合 | `DiagnosisTraceService` |
|
||||
| 工具调用记录 | `ToolInvocationRecorder` |
|
||||
| Executor 引用验真 | `ExecutorGatekeeperService`, `VerifierInputHook` |
|
||||
| self_evaluation 合并 | `SelfEvaluationMergeService` |
|
||||
@@ -0,0 +1,258 @@
|
||||
# Harness 与质量门禁架构
|
||||
|
||||
**更新日期**:2026-07-08
|
||||
**状态**:当前可运行架构 + 后续门禁规划
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
||||
|
||||
## 1. 设计目标
|
||||
|
||||
Agent 系统的核心风险不是“没有答案”,而是:
|
||||
|
||||
- 答案引用了不存在的证据。
|
||||
- 工具调用失败后仍然编造结论。
|
||||
- 检索结果相关性不足但被当作强证据。
|
||||
- 多轮诊断重复检索同一文档,浪费上下文。
|
||||
- 最终报告无法回放执行过程。
|
||||
|
||||
因此当前 MVP 的 Harness 不是单个组件,而是一组约束:
|
||||
|
||||
```text
|
||||
Prompt contract
|
||||
+ Tool boundary
|
||||
+ Agent hooks
|
||||
+ Trace persistence
|
||||
+ Gatekeeper deterministic validation
|
||||
+ Verifier / rule evaluation
|
||||
+ Eval baseline
|
||||
```
|
||||
|
||||
## 2. Harness 总图
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
Input["User / AIOps input"] --> Prompt["Prompt contract"]
|
||||
Prompt --> Agent["Planner / Executor / Verifier / Composer"]
|
||||
Agent --> Tools["Evidence tools"]
|
||||
Tools --> Invocation["tool_invocation"]
|
||||
Agent --> StepHook["AgentLoggingHook"]
|
||||
StepHook --> Step["agent_step"]
|
||||
Agent --> Run["diagnosis_run"]
|
||||
|
||||
Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
|
||||
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
Agent --> Gatekeeper
|
||||
Invocation --> TraceSummary["ToolTraceSummaryService"]
|
||||
Gatekeeper --> Verifier["chat_verifier"]
|
||||
TraceSummary --> Verifier
|
||||
Verifier --> SelfEval["self_evaluation.verifier_evaluation"]
|
||||
|
||||
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
|
||||
AiOpsRule --> AiOpsEval["self_evaluation.aiops_rule_evaluation"]
|
||||
|
||||
Run --> TraceAPI["DiagnosisTraceService"]
|
||||
Step --> TraceAPI
|
||||
Invocation --> TraceAPI
|
||||
SelfEval --> TraceAPI
|
||||
AiOpsEval --> TraceAPI
|
||||
|
||||
TraceAPI --> Eval["diagnosis eval / RAG eval"]
|
||||
```
|
||||
|
||||
## 3. Prompt Contract
|
||||
|
||||
当前 Prompt 按角色拆分:
|
||||
|
||||
| Prompt | 用途 |
|
||||
|---|---|
|
||||
| `supervisor-prompt.md` | AIOps Supervisor 调度 Planner / Executor |
|
||||
| `planner-prompt.md` | AIOps Planner 规划、再规划、输出告警报告 |
|
||||
| `executor-prompt.md` | AIOps Executor 按步骤调用工具 |
|
||||
| `chat-planner-prompt.md` | Chat 复杂问题规划 |
|
||||
| `chat-executor-prompt.md` | Chat 执行工具并输出 `executor_evidence_v2` 微观事实 |
|
||||
| `chat-verifier-prompt.md` | 基于 Gatekeeper 已验真的证据判断 claims 是否可推出 |
|
||||
| `chat-composer-prompt.md` | 基于 Verifier 允许表达的内容生成最终用户答复 |
|
||||
|
||||
Prompt 层当前承担的门禁:
|
||||
|
||||
- 禁止凭记忆回答错误码、接口定义、排障步骤。
|
||||
- 需要外部信息时必须调用工具。
|
||||
- 工具连续失败或返回空结果时,最终报告必须诚实说明。
|
||||
- Chat Executor 不允许在窄范围问题中扩展根因、风险或修复建议。
|
||||
- Chat Verifier 不允许做新检索,只能判断已验真证据是否可推出 claims。
|
||||
- Chat Composer 不允许补事实,尤其不能把 `$.no_evidence` 表达为“已排除/确认没有”。
|
||||
- AIOps payload 模式必须聚焦输入告警。
|
||||
|
||||
Chat 链路还会在 `verifier_evaluation.prompt_audit` 中持久化紧凑 Prompt 审计快照:
|
||||
|
||||
```json
|
||||
{
|
||||
"version": "chat-prompts-v1",
|
||||
"prompts": [
|
||||
{
|
||||
"name": "chat_executor",
|
||||
"version": "chat-executor-v2",
|
||||
"resource": "prompts/chat-executor-prompt.md"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
该快照只保存版本和资源路径,不保存完整 Prompt 文本。它用于面试演示、trace 回放和离线 baseline 解释“本次诊断使用了哪套 Prompt 契约”。
|
||||
|
||||
## 4. Trace Hooks
|
||||
|
||||
`AgentLoggingHook` 是当前 Agent step 可观测性的核心。
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant A as Agent
|
||||
participant H as AgentLoggingHook
|
||||
participant DB as agent_step
|
||||
|
||||
A->>H: before_model(messages, sessionId, runId)
|
||||
H->>DB: 写入 session_id / run_id / model_input / step_index / agent_name
|
||||
A-->>A: LLM 推理
|
||||
A->>H: after_model(messages, sessionId, runId)
|
||||
H->>DB: 回填 model_output / thought / has_tool_call / duration / token_count
|
||||
```
|
||||
|
||||
记录内容:
|
||||
|
||||
- 最近输入消息摘要。
|
||||
- Agent 输出摘要。
|
||||
- 是否包含 tool call。
|
||||
- duration。
|
||||
- token count。
|
||||
- Verifier 的 JSON 输出摘要。
|
||||
|
||||
新写入必须带 `run_id`;`session_id` 仍保留用于粗粒度排查和历史兼容。
|
||||
|
||||
## 5. Tool Invocation 门禁
|
||||
|
||||
工具调用记录由 `ToolInvocationRecorder` 和具体工具共同完成。
|
||||
|
||||
核心记录:
|
||||
|
||||
```text
|
||||
tool_name
|
||||
input_params
|
||||
output_preview
|
||||
retrieval_layer
|
||||
l0_match_count
|
||||
l1_match_count
|
||||
retrieval_details
|
||||
-> evidence_refs
|
||||
relevance_level
|
||||
dedup_reason
|
||||
duration_ms
|
||||
success
|
||||
error_message
|
||||
```
|
||||
|
||||
对 `lookup_knowledge` 的质量约束:
|
||||
|
||||
- L0 只作为 hint,不绕过 L1。
|
||||
- 检索结果归一化为 `PRECISE`、`HIGHLY_RELEVANT`、`REFERENCE`。
|
||||
- 同 session 内重复文档会被 `RetrievedDocTracker` 去重。
|
||||
- dedup、no evidence、failed 等状态进入 `retrieval_details.evidence_status`。
|
||||
- `retrieval_details.evidence_refs` 记录可被 Executor 引用的最小证据文本,格式为 `raw_path + text`。
|
||||
- no-hit / no-evidence 工具结果会生成 `raw_path=$.no_evidence` 的负向证据引用,语义仅限“本次查询未检索到匹配证据”。
|
||||
|
||||
## 6. Gatekeeper 与 Verifier 门禁
|
||||
|
||||
Chat Verifier 前置一层 Gatekeeper。Gatekeeper 不调用 LLM,只用代码检查 Executor 输出的证据引用是否真实存在。
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Invocation["tool_invocation"] --> EvidenceRefs["retrieval_details.evidence_refs"]
|
||||
ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
EvidenceRefs --> Gatekeeper
|
||||
Gatekeeper --> GateResult["gatekeeper_result"]
|
||||
Invocation --> Summary["ToolTraceSummaryService"]
|
||||
Summary --> EvidenceIndex["tool_trace_summary"]
|
||||
GateResult --> Verifier["chat_verifier"]
|
||||
ExecutorOutput --> Verifier
|
||||
EvidenceIndex --> Verifier
|
||||
Verifier --> Verdict{"verdict"}
|
||||
Verdict -->|PASS| Composer["chat_composer"]
|
||||
Composer --> Pass["输出最终答复"]
|
||||
Verdict -->|LOW_CONFID| Low["补证据或低置信输出"]
|
||||
Verdict -->|REJECT| Reject["降级输出"]
|
||||
```
|
||||
|
||||
Gatekeeper 检查:
|
||||
|
||||
| 检查 | 失败语义 |
|
||||
|---|---|
|
||||
| `answer_version=executor_evidence_v2` | 非结构化或旧结构输出降为低置信 |
|
||||
| `source_invocation_id` 真实存在 | 伪造 ID 直接拒绝 |
|
||||
| `tool_name` 与 invocation 对齐 | 张冠李戴直接拒绝 |
|
||||
| `raw_path` 存在于 `evidence_refs` | 无中生有直接拒绝 |
|
||||
| `evidence_excerpt` 由 `evidence_refs[].text` 支撑 | excerpt 编造或错配直接拒绝 |
|
||||
| `negative_observation` 只能引用 `$.no_evidence` | 用正向日志证明“没查到”直接拒绝 |
|
||||
|
||||
Gatekeeper 审计还会记录 `rule_set_version` 和已启用规则元数据摘要。当前规则元数据来自本地 `gatekeeper-rules.json`,规则执行仍是确定性 Java 代码。
|
||||
|
||||
Verifier 输出:
|
||||
|
||||
```json
|
||||
{
|
||||
"verdict": "PASS|LOW_CONFID|REJECT",
|
||||
"groundedness_score": 0.8,
|
||||
"critical_fact_count": 2,
|
||||
"claim_checks": [],
|
||||
"facts_checked": [],
|
||||
"rationale": "..."
|
||||
}
|
||||
```
|
||||
|
||||
Verifier 不再逐字核验 excerpt 真伪;这由 Gatekeeper 完成。Verifier 只回答一个问题:`claim_text` 是否能由已经验真的 `evidence_excerpt` 推导出来。
|
||||
|
||||
结果写入:
|
||||
|
||||
```text
|
||||
diagnosis_run.self_evaluation.verifier_evaluation
|
||||
```
|
||||
|
||||
其中同时持久化 `executor_structured_output`、`gatekeeper_result`、`tool_trace_summary`、`prompt_audit` 和 `composer_output`,用于 Trace 回放。
|
||||
|
||||
## 7. AIOps 规则门禁
|
||||
|
||||
AIOps 当前不走 Chat Verifier,而是用 `AiOpsRuleEvaluationService` 做轻量检查。
|
||||
|
||||
检查重点:
|
||||
|
||||
- 最终报告是否存在。
|
||||
- payload 模式是否围绕输入告警展开。
|
||||
- 是否调用证据工具,尤其是 `lookup_knowledge`、日志、指标。
|
||||
- 是否把无关活跃告警扩展成主诊断对象。
|
||||
|
||||
结果写入:
|
||||
|
||||
```text
|
||||
diagnosis_run.self_evaluation.aiops_rule_evaluation
|
||||
```
|
||||
|
||||
## 8. Eval Baseline
|
||||
|
||||
当前质量门禁还包括离线评测资产:
|
||||
|
||||
| 评测 | 位置 | 作用 |
|
||||
|---|---|---|
|
||||
| Diagnosis eval | `mvp/eval/` | 检查诊断 trace、报告和证据行为 |
|
||||
| RAG retrieval eval | `eval/rag-retrieval/` | 检查固定检索 query 的召回稳定性 |
|
||||
| Live RAG acceptance | `scripts/eval_rag_live_acceptance.py` | 检查运行环境中真实 `/api/search/similar` 行为 |
|
||||
|
||||
## 9. 后续门禁规划
|
||||
|
||||
从旧版设计继承但尚未完整实现的门禁:
|
||||
|
||||
- 工具参数 schema 校验。
|
||||
- 同一工具调用次数上限。
|
||||
- 工具超时的统一熔断。
|
||||
- Gatekeeper 规则远程化或三层分离:索引层、元数据层、规则实现层。
|
||||
- Prompt 版本回滚和更细粒度变更审计。
|
||||
- Verifier 对 AIOps 报告的 LLM 级事实校验。
|
||||
|
||||
这些应在评测集扩大后逐步加入,避免一次性把诊断流程卡得过死。
|
||||
+1
-1
@@ -2,7 +2,7 @@
|
||||
|
||||
**更新日期**:2026-07-06
|
||||
**状态**:当前主架构 + 后续演进边界
|
||||
**关联计划**:[`mvp/issues/active/rag-refactor-plan.md`](../issues/active/rag-refactor-plan.md)
|
||||
**关联计划**:[`mvp/issues/archived/rag-refactor-plan.md`](../../../issues/archived/rag-refactor-plan.md)
|
||||
|
||||
## 1. 架构目标
|
||||
|
||||
@@ -0,0 +1,160 @@
|
||||
# 会话与 Trace 生命周期
|
||||
|
||||
**更新日期**:2026-07-10
|
||||
**状态**:当前可运行架构
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/session-management.md`
|
||||
|
||||
## 1. 定位
|
||||
|
||||
当前 MVP 把“会话态”和“运行态”拆开:
|
||||
|
||||
```text
|
||||
chat_session(sessionId)
|
||||
-> diagnosis_run(runId)
|
||||
-> agent_step(runId)
|
||||
-> tool_invocation(runId)
|
||||
```
|
||||
|
||||
- `sessionId` 表示多轮会话目录和 Redis 上下文。
|
||||
- `runId` 表示一次可回放诊断执行。
|
||||
- `DiagnosisTraceService` 聚合一个 run 的主记录、步骤和工具调用,形成可回放 Trace。
|
||||
- `diagnosis_session` 只保留为历史兼容和回滚表。
|
||||
|
||||
## 2. 生命周期总图
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Start["request: chat / ai_ops"] --> Resolve["resolve sessionId"]
|
||||
Resolve --> Session["ensure chat_session metadata"]
|
||||
Session --> Run["create diagnosis_run(runId)"]
|
||||
Run --> Running["run.status = RUNNING"]
|
||||
|
||||
Running --> Agent["Agent workflow"]
|
||||
Agent --> Context["execution context(sessionId, runId)"]
|
||||
Context --> StepHook["AgentLoggingHook"]
|
||||
StepHook --> Step["agent_step(session_id, run_id)"]
|
||||
Context --> Tool["Evidence tools"]
|
||||
Tool --> Invocation["tool_invocation(session_id, run_id)"]
|
||||
Invocation --> Gatekeeper["Gatekeeper evidence validation"]
|
||||
|
||||
Agent --> Final{"workflow result"}
|
||||
Final -->|success| Success["run.status = SUCCESS, answer saved"]
|
||||
Final -->|failed| Failed["run.status = FAILED"]
|
||||
|
||||
Success --> Evaluation["diagnosis_run.self_evaluation merge"]
|
||||
Failed --> Evaluation
|
||||
Evaluation --> Trace["GET /api/diagnosis/{sessionId}/trace?runId=..."]
|
||||
Success --> Feedback["POST /api/feedback(sessionId, runId)"]
|
||||
Feedback --> Case["useful -> case_library(run_id)"]
|
||||
```
|
||||
|
||||
## 3. ID 规则
|
||||
|
||||
| ID | 来源 | 含义 |
|
||||
|---|---|---|
|
||||
| `sessionId` | Chat request `Id`、AIOps payload `sessionId`,缺失时由服务生成 | 多轮会话目录和 Redis 上下文 |
|
||||
| `runId` | 每次有效 Chat/AIOps 执行创建 | 一次诊断运行和 Trace 回放边界 |
|
||||
|
||||
设计含义:
|
||||
|
||||
- 同一个 `sessionId` 可以贯穿多轮 Chat。
|
||||
- 每次有效 Chat/AIOps 执行都会创建新的 `runId`。
|
||||
- Trace 和 Feedback 新客户端应传 `runId`;只传 `sessionId` 时兼容解析 latest run。
|
||||
- latest run 排序使用 `diagnosis_run.created_at DESC, id DESC`,不使用 `updated_at`。
|
||||
|
||||
## 4. 运行状态流转
|
||||
|
||||
```mermaid
|
||||
stateDiagram-v2
|
||||
[*] --> PENDING
|
||||
PENDING --> RUNNING: start diagnosis
|
||||
RUNNING --> SUCCESS: workflow completed
|
||||
RUNNING --> FAILED: exception / empty state
|
||||
SUCCESS --> SUCCESS: feedback submitted
|
||||
FAILED --> FAILED: feedback submitted
|
||||
```
|
||||
|
||||
字段边界:
|
||||
|
||||
| 字段 | 所属表 | 含义 |
|
||||
|---|---|---|
|
||||
| `status` | `diagnosis_run` | 单次运行执行状态 |
|
||||
| `answer` | `diagnosis_run` | 本次运行最终报告或答复 |
|
||||
| `self_evaluation` | `diagnosis_run` | 本次运行系统自评估 JSON |
|
||||
| `feedback` | `diagnosis_run` | 本次运行用户反馈 |
|
||||
|
||||
`feedback` 不修改 `status`。一个执行成功但用户标记 `not_useful` 的 run,仍然应该是 `SUCCESS + feedback=not_useful`。
|
||||
|
||||
## 5. agent_step 写入
|
||||
|
||||
`AgentLoggingHook` 在模型调用前后写入和回填 `agent_step`。
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant Agent as Agent
|
||||
participant Hook as AgentLoggingHook
|
||||
participant DB as agent_step
|
||||
|
||||
Agent->>Hook: before_model(messages, sessionId, runId)
|
||||
Hook->>DB: insert step(session_id, run_id, model_input, step_index)
|
||||
Agent->>Hook: after_model(output, sessionId, runId)
|
||||
Hook->>DB: update model_output, duration, token_count, has_tool_call
|
||||
```
|
||||
|
||||
新写入必须带 `run_id`,同时保留 `session_id` 便于粗粒度排查。
|
||||
|
||||
## 6. tool_invocation 写入
|
||||
|
||||
工具调用记录同样通过执行上下文拿到 `sessionId + runId`:
|
||||
|
||||
```text
|
||||
ToolInvocationRecorder
|
||||
-> tool_invocation.session_id
|
||||
-> tool_invocation.run_id
|
||||
-> retrieval_details / evidence_refs
|
||||
```
|
||||
|
||||
Verifier、Gatekeeper 和 EvaluationService 应按 `run_id` 读取工具调用,避免同一 session 的其他 run 参与评分或证据校验。
|
||||
|
||||
## 7. Trace API 聚合
|
||||
|
||||
```text
|
||||
GET /api/diagnosis/{sessionId}/trace
|
||||
GET /api/diagnosis/{sessionId}/trace?runId=run-...
|
||||
```
|
||||
|
||||
聚合逻辑:
|
||||
|
||||
```text
|
||||
diagnosis_run by sessionId + runId
|
||||
+ chat_session metadata when available
|
||||
+ agent_step where run_id = runId, ordered by the Trace API
|
||||
+ tool_invocation where run_id = runId order by id
|
||||
-> DiagnosisTraceResponse
|
||||
```
|
||||
|
||||
当 `runId` 缺失时,Trace API 为兼容旧客户端解析最新 run,并在响应中返回 resolved `runId`。当 `runId` 属于其他 `sessionId` 时,API 必须拒绝,不能泄漏其他会话的 Trace。
|
||||
|
||||
## 8. Chat 与 AIOps 差异
|
||||
|
||||
| 维度 | Chat | AIOps |
|
||||
|---|---|---|
|
||||
| `agent_flow` | `CHAT` | `AI_OPS` |
|
||||
| 编排方式 | `SequentialAgent`: Planner -> Executor -> Gatekeeper -> Verifier -> Composer | `SupervisorAgent`: Planner + Executor |
|
||||
| 自评估 | `rule_evaluation` + `verifier_evaluation` | `aiops_rule_evaluation` |
|
||||
| 答案字段 | Chat 最终答复 | 告警分析报告 |
|
||||
| runId 暴露 | `/api/chat` JSON response | `/api/ai_ops` SSE metadata message |
|
||||
|
||||
## 9. 清理与边界
|
||||
|
||||
- Redis 会话历史用于多轮上下文,不是长期审计记录。
|
||||
- MySQL `diagnosis_run + agent_step + tool_invocation` 是主要可回放来源。
|
||||
- `chat_session.expires_at` 只是目录元数据;Redis 消息历史可独立过期。
|
||||
- `RetrievedDocTracker` 仍是 session 级运行时去重状态,诊断结束后清理。
|
||||
|
||||
## 10. 后续增强
|
||||
|
||||
1. Trace API 增加更结构化的 `self_evaluation` 展示。
|
||||
2. `agent_step` 与 `tool_invocation.step_id` 建立更严格关联。
|
||||
3. 旧 `diagnosis_session` 只读观察期结束后,再评估数据库层面的约束收紧或归档策略。
|
||||
@@ -1,444 +1,101 @@
|
||||
# 当前 MVP 架构
|
||||
|
||||
**更新日期**:2026-07-08
|
||||
**更新日期**:2026-07-23
|
||||
**状态**:当前可运行架构
|
||||
**适用范围**:Demo、面试讲解、后续迭代规划
|
||||
|
||||
## 1. 系统定位
|
||||
|
||||
SuperBizAgent MVP 不是通用 Chatbot,而是面向故障诊断的 Agent 工程项目。
|
||||
SuperBizAgent 是面向故障诊断的可追踪 Agent 应用。当前系统只保留一个拥有 Tool loop 的 `Diagnosis Agent`;Harness 负责确定性的预算、取消、工具边界、证据验真、语义审查和安全发布。
|
||||
|
||||
核心目标:
|
||||
|
||||
- 支持用户主动发起的 Chat 诊断。
|
||||
- 支持 AIOps 告警触发的自动诊断。
|
||||
- 保留 Agent 的规划、执行、验证过程。
|
||||
- 工具调用必须显式、可追踪、可回放。
|
||||
- RAG 检索必须通过 `lookup_knowledge` 暴露证据链。
|
||||
- 每次诊断都沉淀 session、step、tool invocation 和 self evaluation。
|
||||
|
||||
## 2. 总体分层
|
||||
## 2. 分层
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph API["API Layer"]
|
||||
ChatController["ChatController"]
|
||||
TraceController["DiagnosisTraceController"]
|
||||
SearchController["SearchController"]
|
||||
DocumentController["DocumentController"]
|
||||
end
|
||||
|
||||
subgraph App["Application Service"]
|
||||
ChatService["ChatService"]
|
||||
AiOpsService["AiOpsService"]
|
||||
TraceService["DiagnosisTraceService"]
|
||||
end
|
||||
|
||||
subgraph Agent["Agent Orchestration"]
|
||||
Supervisor["Supervisor"]
|
||||
Planner["Planner"]
|
||||
Executor["Executor"]
|
||||
Gatekeeper["Gatekeeper"]
|
||||
Verifier["Verifier"]
|
||||
Composer["Composer"]
|
||||
end
|
||||
|
||||
subgraph Tools["Evidence Tools"]
|
||||
KnowledgeTool["lookup_knowledge"]
|
||||
LogsTool["query_logs"]
|
||||
MetricsTool["query_metrics"]
|
||||
AlertsTool["queryPrometheusAlerts"]
|
||||
end
|
||||
|
||||
subgraph Skills["Skill / Playbook"]
|
||||
SkillRegistry["SkillRegistry"]
|
||||
PlannerSkillHook["PlannerSkillMetadataHook"]
|
||||
SkillsHook["SkillsAgentHook"]
|
||||
ReadSkill["read_skill"]
|
||||
end
|
||||
|
||||
subgraph RAG["RAG Retrieval"]
|
||||
L0["KnowledgeIndexService"]
|
||||
VectorSearch["VectorSearchService"]
|
||||
VectorStore["Spring AI VectorStore"]
|
||||
SdkFallback["Milvus SDK fallback"]
|
||||
end
|
||||
|
||||
subgraph Store["Persistence and Trace"]
|
||||
ChatSession["chat_session"]
|
||||
Run["diagnosis_run"]
|
||||
Step["agent_step"]
|
||||
Invocation["tool_invocation"]
|
||||
ApiDoc["api_document"]
|
||||
Milvus["Milvus/Zilliz"]
|
||||
end
|
||||
|
||||
API --> App
|
||||
ChatService --> Agent
|
||||
AiOpsService --> Agent
|
||||
SkillRegistry --> PlannerSkillHook
|
||||
PlannerSkillHook --> Planner
|
||||
SkillRegistry --> SkillsHook
|
||||
SkillsHook --> Executor
|
||||
Executor --> ReadSkill
|
||||
Agent --> Tools
|
||||
KnowledgeTool --> RAG
|
||||
RAG --> Store
|
||||
Tools --> Invocation
|
||||
Agent --> Step
|
||||
App --> Session
|
||||
TraceService --> Session
|
||||
TraceService --> Step
|
||||
TraceService --> Invocation
|
||||
Browser["Browser / API client"] --> Chat["POST /api/chat named SSE"]
|
||||
Chat --> App["ChatApplicationUseCase"]
|
||||
App --> Router["Intent Router"]
|
||||
Router --> System["System Chat"]
|
||||
Router --> Knowledge["Knowledge Query"]
|
||||
Router --> Diagnosis["Diagnosis Agent"]
|
||||
Diagnosis --> Tools["Harness ACI Tools"]
|
||||
Tools --> Canonical["Redis canonical invocation"]
|
||||
Diagnosis --> Evidence["EvidenceGuard"]
|
||||
Evidence --> Semantic["SemanticGuard"]
|
||||
Semantic --> Release["Release Policy"]
|
||||
Release --> Chat
|
||||
App --> Run["diagnosis_run"]
|
||||
Diagnosis --> Step["agent_step metadata audit"]
|
||||
App --> Timeline["diagnosis_trace_event"]
|
||||
Diagnosis --> Reasoning["agent_reasoning_audit restricted"]
|
||||
Tools --> Invocation["tool_invocation metadata audit"]
|
||||
Run --> Trace["Diagnosis Trace API"]
|
||||
Step --> Trace
|
||||
Invocation --> Trace
|
||||
Timeline --> Trace
|
||||
Reasoning --> ReasoningAPI["Reasoning Audit API"]
|
||||
```
|
||||
|
||||
```text
|
||||
API Layer
|
||||
-> ChatController
|
||||
-> DiagnosisTraceController
|
||||
-> SearchController
|
||||
-> DocumentController
|
||||
|
||||
Application Service
|
||||
-> ChatService
|
||||
-> AiOpsService
|
||||
-> DiagnosisTraceService
|
||||
|
||||
Agent Orchestration
|
||||
-> Supervisor
|
||||
-> Planner
|
||||
-> Executor
|
||||
-> Gatekeeper
|
||||
-> Verifier
|
||||
-> Composer
|
||||
|
||||
Evidence Tools
|
||||
-> lookup_knowledge
|
||||
-> query_logs
|
||||
-> query_metrics
|
||||
-> queryPrometheusAlerts
|
||||
|
||||
Skill / Playbook
|
||||
-> SkillRegistry
|
||||
-> PlannerSkillMetadataHook gives Planner name/description only
|
||||
-> SkillsAgentHook gives Executor read_skill
|
||||
-> Verifier is isolated from skills
|
||||
|
||||
RAG Retrieval
|
||||
-> KnowledgeIndexService
|
||||
-> VectorSearchService
|
||||
-> Spring AI VectorStore
|
||||
-> Milvus SDK fallback
|
||||
|
||||
Persistence
|
||||
-> chat_session
|
||||
-> diagnosis_run
|
||||
-> agent_step.run_id
|
||||
-> tool_invocation.run_id
|
||||
-> api_document
|
||||
-> Milvus/Zilliz collection
|
||||
|
||||
Quality Gates
|
||||
-> executor gatekeeper
|
||||
-> chat verifier
|
||||
-> AIOps rule evaluation
|
||||
-> diagnosis eval baseline
|
||||
-> RAG retrieval baseline
|
||||
```
|
||||
|
||||
## 3. Chat 诊断链路
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
actor User as 用户
|
||||
participant API as POST /api/chat
|
||||
participant Chat as ChatService
|
||||
participant Planner as Planner Agent
|
||||
participant Executor as Executor Agent
|
||||
participant Tool as Evidence Tools
|
||||
participant Gatekeeper as Gatekeeper Hook
|
||||
participant Verifier as Verifier Agent
|
||||
participant Composer as Composer Agent
|
||||
participant DB as Trace Tables
|
||||
participant Trace as Trace API
|
||||
|
||||
User->>API: 提交诊断问题
|
||||
API->>Chat: execute chat strategy
|
||||
Chat->>DB: 创建 chat_session metadata + diagnosis_run(runId)
|
||||
Chat->>Planner: 复杂问题进入规划
|
||||
Planner->>DB: 写入 agent_step.run_id
|
||||
Planner->>Executor: 下发排查方向
|
||||
Executor->>Tool: lookup_knowledge / logs / metrics
|
||||
Tool->>DB: 写入 tool_invocation.run_id
|
||||
Tool-->>Executor: 返回证据
|
||||
Executor->>Gatekeeper: 输出 executor_evidence_v2
|
||||
Gatekeeper->>DB: 读取 tool_invocation.evidence_refs 并校验引用
|
||||
Gatekeeper->>Verifier: 传入已验真的 claims / excerpts
|
||||
Verifier->>DB: 合并 diagnosis_run.self_evaluation.verifier_evaluation
|
||||
Verifier->>Composer: 传入 allowed_claims / missing_info / actions
|
||||
Composer->>Chat: 生成最终用户答复
|
||||
Chat->>DB: 保存 diagnosis_run.answer
|
||||
User->>Trace: GET /api/diagnosis/{sessionId}/trace?runId=...
|
||||
Trace->>DB: 聚合 run / step / tool
|
||||
Trace-->>User: 返回可回放诊断链路
|
||||
```
|
||||
## 3. 唯一 Chat 主链
|
||||
|
||||
```text
|
||||
POST /api/chat
|
||||
-> ChatService
|
||||
-> 简单问题:轻量回答
|
||||
-> 复杂诊断:Agent 编排
|
||||
-> Planner 制定排查方向
|
||||
-> Executor 调用证据工具
|
||||
-> lookup_knowledge
|
||||
-> query_logs
|
||||
-> query_metrics
|
||||
-> Gatekeeper 校验 Executor 证据引用真实性
|
||||
-> Verifier 判断 claim 是否能由已核验证据推出
|
||||
-> Composer 生成最终用户答复
|
||||
-> 保存 chat_session metadata
|
||||
-> 保存 diagnosis_run
|
||||
-> 保存 agent_step.run_id
|
||||
-> 保存 tool_invocation.run_id
|
||||
-> 合并 diagnosis_run.self_evaluation.verifier_evaluation
|
||||
-> metadata(session_id, run_id)
|
||||
-> status*
|
||||
-> ChatApplicationUseCase
|
||||
-> SYSTEM_CHAT | KNOWLEDGE_QUERY | DIAGNOSIS
|
||||
-> content | failure
|
||||
-> done(SUCCESS | FALLBACK | FAILED)
|
||||
```
|
||||
|
||||
Chat 链路的质量门禁由三段组成:Gatekeeper 先做代码级引用验真,Verifier 再做 LLM 可推导性判断,Composer 最后控制对用户的表达边界。Gatekeeper、Verifier、Composer 的输出合并到当前 `diagnosis_run.self_evaluation.verifier_evaluation`,Trace API 会展示该验证结果。
|
||||
- Controller 只处理请求校验、bounded worker、SSE 和 disconnect。
|
||||
- Application Use Case 拥有 Session/Run、路由、PreviousTurn 和终态持久化。
|
||||
- Diagnosis Agent 是唯一报告作者和唯一拥有 evidence Tool loop 的业务 Agent。
|
||||
- EvidenceGuard 只做确定性结构/引用验真;SemanticGuard 在隔离上下文做整份报告语义审查。
|
||||
- 未通过 Release Policy 的 Draft 永不进入公开 SSE。
|
||||
|
||||
Agent 编排细节见 [agent-orchestration.md](agent-orchestration.md)。
|
||||
## 4. Tool 与数据边界
|
||||
|
||||
关键代码:
|
||||
Agent 只看到三个固定 Tool:
|
||||
|
||||
- `src/main/java/com/superbiz/agent/controller/ChatController.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/SelfEvaluationMergeService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
|
||||
- `lookup_knowledge`
|
||||
- `query_logs`
|
||||
- `query_mysql`
|
||||
|
||||
## 4. AIOps 诊断链路
|
||||
每次调用由框架提供 `tool_call_id`,Harness 校验 exact run、只读、Schema、预算和容量。Redis 保存 TTL 内完整 canonical invocation;MySQL `tool_invocation` 只保存长期有界 metadata,不保存完整参数、SQL/日志正文、raw response 或 Agent projection。
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Request["POST /api/ai_ops"] --> Payload{"包含告警 payload?"}
|
||||
Payload -->|是| Targeted["PAYLOAD_TARGETED"]
|
||||
Payload -->|否| Discovery["AUTO_DISCOVERY"]
|
||||
|
||||
Targeted --> BuildPrompt["构造聚焦 payload 的诊断 prompt"]
|
||||
Targeted --> QueryAug["生成 recommended lookup_knowledge query"]
|
||||
Discovery --> DiscoverAlert["通过 queryPrometheusAlerts 发现活跃告警"]
|
||||
|
||||
BuildPrompt --> Plan["Planner 规划排查"]
|
||||
QueryAug --> Plan
|
||||
DiscoverAlert --> Plan
|
||||
|
||||
Plan --> Execute["Executor 收集证据"]
|
||||
Execute --> Knowledge["lookup_knowledge"]
|
||||
Execute --> Metrics["query_metrics / Prometheus"]
|
||||
Execute --> Logs["query_logs"]
|
||||
|
||||
Knowledge --> Report["告警分析报告"]
|
||||
Metrics --> Report
|
||||
Logs --> Report
|
||||
|
||||
Report --> RuleEval["AiOpsRuleEvaluationService"]
|
||||
RuleEval --> SelfEval["self_evaluation.aiops_rule_evaluation"]
|
||||
Report --> Trace["DiagnosisTraceService"]
|
||||
SelfEval --> Trace
|
||||
```
|
||||
## 5. Trace 与持久化
|
||||
|
||||
```text
|
||||
POST /api/ai_ops
|
||||
-> AiOpsService
|
||||
-> 判断是否有告警 payload
|
||||
-> PAYLOAD_TARGETED
|
||||
-> AUTO_DISCOVERY
|
||||
-> 构造 AIOps 诊断 prompt
|
||||
-> payload 模式补充 recommended lookup_knowledge query
|
||||
-> Agent 编排
|
||||
-> Planner / Executor
|
||||
-> Prometheus / logs / knowledge tools
|
||||
-> 生成告警分析报告
|
||||
-> AiOpsRuleEvaluationService
|
||||
-> 合并 diagnosis_run.self_evaluation.aiops_rule_evaluation
|
||||
-> Trace API 可查看全链路
|
||||
chat_session(sessionId)
|
||||
-> diagnosis_run(runId)
|
||||
-> agent_step(runId)
|
||||
-> tool_invocation(runId)
|
||||
-> diagnosis_trace_event(runId)
|
||||
-> agent_reasoning_audit(runId, restricted)
|
||||
```
|
||||
|
||||
AIOps 保留两种模式:
|
||||
- `chat_session` 是 JPA Run 目录与多轮 metadata,不保存完整对话历史。
|
||||
- `diagnosis_run` 是 Run 状态、intent、release outcome、安全发布结果和预算汇总真理源。
|
||||
- `agent_step` 只保存模型步骤 metadata,不保存 Prompt、消息正文、模型正文或 Thought。
|
||||
- `tool_invocation` 只保存 Tool durable audit metadata;完整调用由 Redis canonical store 短期保存。
|
||||
- `diagnosis_trace_event` 是追加式统一 Timeline,记录 Run、Routing、Agent、Tool、Evidence、Semantic 和 Release 生命周期事件;`details` 只能保存有界安全 metadata。
|
||||
- `agent_reasoning_audit` 与普通 Trace 分表,只保存 Provider 实际返回的 reasoning 或明确的 unavailable 记录;reasoning 不属于事实证据。
|
||||
|
||||
| 模式 | 触发条件 | 行为 |
|
||||
|---|---|---|
|
||||
| `PAYLOAD_TARGETED` | 请求包含 alertName、service、severity、description、timeRange 等字段 | 以 payload 为唯一主诊断对象,并生成推荐知识库 query |
|
||||
| `AUTO_DISCOVERY` | 请求没有明确告警 payload | 先查询当前活跃告警,再选择目标排查 |
|
||||
## 6. 公开 API
|
||||
|
||||
AIOps 当前使用轻量规则验证器,重点检查:
|
||||
当前诊断执行入口只有 `POST /api/chat`。诊断审计读取分为:
|
||||
|
||||
- 最终报告是否存在。
|
||||
- payload 模式是否聚焦输入告警。
|
||||
- 是否使用关键证据工具,例如 `lookup_knowledge`、日志、指标。
|
||||
- `GET /api/diagnosis/{sessionId}/trace?runId={runId}`:普通 Trace,返回 Run、步骤、Tool metadata 和统一 Timeline,不返回 reasoning 原文。
|
||||
- `GET /api/diagnosis/{sessionId}/trace/reasoning?runId={runId}`:独立 reasoning 审计读取,`runId` 必填并校验其属于 path `sessionId`。
|
||||
|
||||
关键代码:
|
||||
Reasoning endpoint 是敏感审计面,不属于普通业务 API。当前已完成数据和查询隔离;认证授权、保留期限、加密要求及真实 Provider 验证仍由 ISS-015 收敛。Feedback、文档与检索 API 保持独立;已删除的旧诊断和 Redis conversation Session endpoint 不提供兼容分支。
|
||||
|
||||
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/AiOpsRuleEvaluationService.java`
|
||||
## 7. 安全边界
|
||||
|
||||
## 5. RAG 位置
|
||||
|
||||
RAG 不是隐藏在 Chat Advisor 里的隐式能力,而是 Executor 可以显式调用的工具:
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Executor["Executor Agent"] --> Tool["lookup_knowledge Tool"]
|
||||
Tool --> L0["L0 domain/entity hint"]
|
||||
Tool --> Search["VectorSearchService"]
|
||||
L0 --> Search
|
||||
Search --> VectorStore["Spring AI VectorStore"]
|
||||
Search --> Fallback["Milvus SDK fallback"]
|
||||
VectorStore --> Normalize["score/rawScore/scoreLabel"]
|
||||
Fallback --> Normalize
|
||||
Normalize --> Evidence["evidence output"]
|
||||
Evidence --> Invocation["tool_invocation"]
|
||||
Evidence --> Executor
|
||||
```
|
||||
|
||||
```text
|
||||
Executor
|
||||
-> lookup_knowledge(query)
|
||||
-> L0 domain/entity hint
|
||||
-> VectorSearchService
|
||||
-> Spring AI VectorStore
|
||||
-> Milvus SDK fallback
|
||||
-> evidence shaping
|
||||
-> tool_invocation
|
||||
```
|
||||
|
||||
保留显式工具的原因:
|
||||
|
||||
- Agent 何时检索、检索什么、证据是什么,必须能在 trace 中解释。
|
||||
- AIOps payload 到 query 的业务映射需要项目内控制。
|
||||
- `tool_invocation` 是后续评测、回放和面试讲解的核心材料。
|
||||
|
||||
RAG 总体设计见 [rag-architecture.md](rag-architecture.md),检索运行细节见 [retrieval-observability.md](retrieval-observability.md)。
|
||||
|
||||
## 6. 持久化模型
|
||||
|
||||
当前诊断持久化以 session/run/trace 明细为核心:
|
||||
|
||||
```text
|
||||
chat_session
|
||||
-> 多轮会话目录和元数据
|
||||
-> session_id / status / message_pair_count
|
||||
|
||||
diagnosis_run
|
||||
-> 一次诊断运行的主记录
|
||||
-> run_id / session_id
|
||||
-> query / status / agent_flow / answer
|
||||
-> self_evaluation
|
||||
-> step_count / tool_call_count / duration
|
||||
|
||||
agent_step
|
||||
-> Agent 模型调用步骤
|
||||
-> session_id / run_id
|
||||
-> step_index / agent_name
|
||||
-> model_input / model_output / thought
|
||||
-> duration / token_count
|
||||
|
||||
tool_invocation
|
||||
-> 工具调用事实
|
||||
-> session_id / run_id
|
||||
-> tool_name / input_params / output_preview
|
||||
-> retrieval_layer / retrieval_details
|
||||
-> retrieval_details.evidence_refs
|
||||
-> relevance_level / dedup_reason
|
||||
-> duration / success
|
||||
```
|
||||
|
||||
说明:
|
||||
|
||||
- 旧的 `diagnosis_record` 已不是当前主模型。
|
||||
- `diagnosis_session` 已降级为历史兼容和回滚表,新执行写入 `chat_session + diagnosis_run`。
|
||||
- `api_document` 仍用于文档元数据管理。
|
||||
- 文档向量内容存放在 Milvus/Zilliz collection 中。
|
||||
|
||||
会话和 Trace 生命周期见 [session-trace-lifecycle.md](session-trace-lifecycle.md),完整数据关系见 [data-model.md](data-model.md)。
|
||||
|
||||
## 7. Trace API
|
||||
|
||||
```text
|
||||
GET /api/diagnosis/{sessionId}/trace
|
||||
GET /api/diagnosis/{sessionId}/trace?runId=run-...
|
||||
```
|
||||
|
||||
Trace API 聚合:
|
||||
|
||||
- 会话元数据、运行状态和最终报告。
|
||||
- Agent step 序列。
|
||||
- 工具调用和检索细节。
|
||||
- Chat Gatekeeper / Verifier / Composer 结果。
|
||||
- AIOps rule evaluation 结果。
|
||||
|
||||
Trace 是本项目区别于普通问答系统的关键:答案不是孤立文本,而是可以追溯到 Agent 决策、工具调用和证据来源。
|
||||
|
||||
Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明见 [harness-quality-gates.md](harness-quality-gates.md),用户反馈与 `self_evaluation` 闭环见 [feedback-architecture.md](feedback-architecture.md)。
|
||||
|
||||
## 8. 质量门禁
|
||||
|
||||
当前质量门禁分层如下:
|
||||
|
||||
| 门禁 | 位置 | 作用 |
|
||||
|---|---|---|
|
||||
| Executor Gatekeeper | `VerifierInputHook` / `ExecutorGatekeeperService` | 校验 Executor 引用的 invocation、`raw_path`、`evidence_excerpt` 是否真实 |
|
||||
| Chat Verifier | `ChatService` | 判断已验真证据是否能推出 Executor claims |
|
||||
| Chat Composer | `ChatService` | 只表达 Verifier 允许输出的内容,避免把 no-evidence 说成已排除 |
|
||||
| AIOps Rule Evaluation | `AiOpsRuleEvaluationService` | 校验告警诊断是否聚焦 payload 并使用证据 |
|
||||
| Diagnosis Eval Baseline | `mvp/eval/` | 固化诊断 trace 和报告行为 |
|
||||
| RAG Retrieval Baseline | `eval/rag-retrieval/` | 固化检索召回行为,避免 RAG 重构回退 |
|
||||
| Live RAG Acceptance | `scripts/eval_rag_live_acceptance.py` | 在运行环境中验证重建索引后的真实检索 |
|
||||
|
||||
## 9. 当前完成状态
|
||||
|
||||
已经完成:
|
||||
|
||||
- Chat 和 AIOps 两条入口链路。
|
||||
- 显式 `lookup_knowledge` Agent Tool。
|
||||
- L0 从最终决策降级为 domain/entity hint。
|
||||
- `VectorSearchService` 作为稳定检索门面。
|
||||
- Spring AI VectorStore 读取路径。
|
||||
- Milvus SDK fallback。
|
||||
- `score` / `rawScore` / `scoreLabel` 分数语义拆分。
|
||||
- `title`、`breadcrumb`、`content` 参与 embedding 文本。
|
||||
- `tool_invocation` 记录检索层、relevance level、dedup reason。
|
||||
- Chat verifier 和 AIOps rule evaluation 合并进 `self_evaluation`。
|
||||
- Chat Executor 结构化输出 `executor_evidence_v2`,不再直接承担最终用户答复。
|
||||
- `tool_invocation.retrieval_details.evidence_refs` 支持 `raw_path` 精确引用和 `$.no_evidence` 负向证据。
|
||||
- Gatekeeper 对 Executor 引用做代码级验真,并在审计中记录 `rule_set_version` 和规则元数据摘要。
|
||||
- Verifier 只判断可推导性。
|
||||
- Composer 在 Verifier 之后生成最终用户表达,并限制 negative observation 过度表述。
|
||||
- RAG offline baseline 和 live acceptance 脚本。
|
||||
|
||||
暂不作为当前已完成能力声明:
|
||||
|
||||
- 完整 QueryTransformer / MultiQuery。
|
||||
- BM25、RRF、cross-encoder rerank。
|
||||
- 完整邻居 chunk / section context expansion。
|
||||
- VectorStore 写入路径全面迁移。
|
||||
- 完整 LLM-based AIOps verifier。
|
||||
|
||||
后续 Agent 拆分、Skill/Playbook、MCP 工具协议化和进程隔离等方向见 [evolution-roadmap.md](evolution-roadmap.md)。
|
||||
|
||||
## 10. 关键代码索引
|
||||
|
||||
| 能力 | 代码 |
|
||||
|---|---|
|
||||
| Chat 入口与编排 | `ChatController`, `ChatService` |
|
||||
| AIOps 入口与编排 | `ChatController.aiOps`, `AiOpsService` |
|
||||
| AIOps 规则验证 | `AiOpsRuleEvaluationService` |
|
||||
| 知识库工具 | `LookupKnowledgeTool` |
|
||||
| L0 hint | `KnowledgeIndexService` |
|
||||
| 向量检索门面 | `VectorSearchService` |
|
||||
| 文档切片 | `DocumentChunkService` |
|
||||
| 向量写入 | `VectorIndexService` |
|
||||
| Spring AI VectorStore 配置辅助 | `SpringAiVectorStoreSidecarService` |
|
||||
| Trace 聚合 | `DiagnosisTraceService` |
|
||||
| 工具调用记录 | `ToolInvocationRecorder` |
|
||||
| Executor 引用验真 | `ExecutorGatekeeperService`, `VerifierInputHook` |
|
||||
| self_evaluation 合并 | `SelfEvaluationMergeService` |
|
||||
- 普通 SSE、Trace、Evidence Snapshot、业务结果和应用日志不输出或保存 reasoning 原文。
|
||||
- 仅当 Provider 在模型 metadata 中实际返回 reasoning 时,审计 Hook 才将其截断后写入独立表;Provider 未返回时不得伪造。
|
||||
- Reasoning 不能作为事实证据,也不能绕过 EvidenceGuard 或 SemanticGuard。
|
||||
- 不向 Agent 暴露 Redis、canonical key、完整 Tool 请求/响应或数据库凭据。
|
||||
- EvidenceGuard 只接受当前 Run 的 READY canonical invocation。
|
||||
- SemanticGuard 无 Tool、无记忆、无回调主 Agent 能力。
|
||||
- technical failure 与 guard rejection 只能产生 stable failure 或固定 safe fallback。
|
||||
|
||||
@@ -1,258 +1,64 @@
|
||||
# Harness 与质量门禁架构
|
||||
# Harness 与质量门禁
|
||||
|
||||
**更新日期**:2026-07-08
|
||||
**状态**:当前可运行架构 + 后续门禁规划
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
||||
**更新日期**:2026-07-23
|
||||
**状态**:当前可运行架构
|
||||
|
||||
## 1. 设计目标
|
||||
## 1. Harness 定位
|
||||
|
||||
Agent 系统的核心风险不是“没有答案”,而是:
|
||||
Harness 是确定性执行边界,不承担业务推理。它统一管理:
|
||||
|
||||
- 答案引用了不存在的证据。
|
||||
- 工具调用失败后仍然编造结论。
|
||||
- 检索结果相关性不足但被当作强证据。
|
||||
- 多轮诊断重复检索同一文档,浪费上下文。
|
||||
- 最终报告无法回放执行过程。
|
||||
- RunContext、deadline、first-terminal-wins lifecycle 与客户端取消。
|
||||
- 模型调用、Tool 调用、Token、字节数和单 Tool 次数预算。
|
||||
- 类型化 retry policy;Diagnosis Agent 和 Tool 调用不自动重试。
|
||||
- ToolBoundary、canonical invocation 与 Agent projection。
|
||||
- EvidenceGuard、Evidence repair、SemanticGuard 与 Release Policy。
|
||||
- metadata-only durable audit、统一 Trace Timeline 与独立 reasoning 审计。
|
||||
|
||||
因此当前 MVP 的 Harness 不是单个组件,而是一组约束:
|
||||
## 2. ToolBoundary
|
||||
|
||||
```text
|
||||
Prompt contract
|
||||
+ Tool boundary
|
||||
+ Agent hooks
|
||||
+ Trace persistence
|
||||
+ Gatekeeper deterministic validation
|
||||
+ Verifier / rule evaluation
|
||||
+ Eval baseline
|
||||
framework tool_call_id
|
||||
-> exact Run / schema / authorization / read-only / budget
|
||||
-> backend execution
|
||||
-> raw response -> Redis canonical invocation
|
||||
-> projector -> bounded agent_result
|
||||
-> ToolInvocation durable metadata audit
|
||||
-> Agent observation
|
||||
```
|
||||
|
||||
## 2. Harness 总图
|
||||
Redis canonical invocation 可在 TTL 内保存完整 request/raw_response/agent_result,受独立前缀、容量和 Harness-only 访问保护。Durable audit 只保存 identity、Tool 名、状态、耗时和字节数;audit 写入失败可观测但不改变 canonical Tool 结果。
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
Input["User / AIOps input"] --> Prompt["Prompt contract"]
|
||||
Prompt --> Agent["Planner / Executor / Verifier / Composer"]
|
||||
Agent --> Tools["Evidence tools"]
|
||||
Tools --> Invocation["tool_invocation"]
|
||||
Agent --> StepHook["AgentLoggingHook"]
|
||||
StepHook --> Step["agent_step"]
|
||||
Agent --> Run["diagnosis_run"]
|
||||
## 3. EvidenceGuard
|
||||
|
||||
Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
|
||||
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
Agent --> Gatekeeper
|
||||
Invocation --> TraceSummary["ToolTraceSummaryService"]
|
||||
Gatekeeper --> Verifier["chat_verifier"]
|
||||
TraceSummary --> Verifier
|
||||
Verifier --> SelfEval["self_evaluation.verifier_evaluation"]
|
||||
EvidenceGuard 不调用模型。它校验 Draft schema、analysis ID、当前 Run Tool ownership、READY 状态、evidence status 和每条结论的引用闭包,并生成只包含 Agent projection 的 verified snapshot。
|
||||
|
||||
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
|
||||
AiOpsRule --> AiOpsEval["self_evaluation.aiops_rule_evaluation"]
|
||||
## 4. SemanticGuard
|
||||
|
||||
Run --> TraceAPI["DiagnosisTraceService"]
|
||||
Step --> TraceAPI
|
||||
Invocation --> TraceAPI
|
||||
SelfEval --> TraceAPI
|
||||
AiOpsEval --> TraceAPI
|
||||
SemanticGuard 使用隔离的单轮模型调用,只接收原始 query、完整 Draft 和 verified snapshot。它无 Tool、无记忆、不访问 Redis、不改写报告;技术失败最多按相同输入重试一次,仍失败则安全降级。
|
||||
|
||||
TraceAPI --> Eval["diagnosis eval / RAG eval"]
|
||||
```
|
||||
## 5. Release Policy
|
||||
|
||||
## 3. Prompt Contract
|
||||
- `SUPPORTED`:发布 Diagnosis Agent 原始安全 Draft 的 typed report。
|
||||
- `UNSUPPORTED` 或 evidence failure:发布有界 SAFE_FALLBACK,说明 `failure_stage`、已验证 `observed_facts`、`validation_issues`、限制和 `next_steps`;不得泄露 Prompt、原始 Draft、原始 Tool 载荷或内部异常。
|
||||
- technical failure:发布 stable failure,不泄漏内部异常。
|
||||
- cancel/timeout:结束 exact Run,禁止 late content。
|
||||
|
||||
当前 Prompt 按角色拆分:
|
||||
## 6. Trace Recorder
|
||||
|
||||
| Prompt | 用途 |
|
||||
|---|---|
|
||||
| `supervisor-prompt.md` | AIOps Supervisor 调度 Planner / Executor |
|
||||
| `planner-prompt.md` | AIOps Planner 规划、再规划、输出告警报告 |
|
||||
| `executor-prompt.md` | AIOps Executor 按步骤调用工具 |
|
||||
| `chat-planner-prompt.md` | Chat 复杂问题规划 |
|
||||
| `chat-executor-prompt.md` | Chat 执行工具并输出 `executor_evidence_v2` 微观事实 |
|
||||
| `chat-verifier-prompt.md` | 基于 Gatekeeper 已验真的证据判断 claims 是否可推出 |
|
||||
| `chat-composer-prompt.md` | 基于 Verifier 允许表达的内容生成最终用户答复 |
|
||||
`diagnosis_trace_event` 是追加式统一 Timeline。Recorder 按 exact `runId` 分配递增 `sequence_no`,覆盖:
|
||||
|
||||
Prompt 层当前承担的门禁:
|
||||
- `RUN`:Run 开始和终态。
|
||||
- `ROUTING`:路由尝试与最终 intent。
|
||||
- `AGENT`:模型步骤及有界 metadata。
|
||||
- `TOOL`:Tool 调用状态和耗时。
|
||||
- `EVIDENCE`:首次校验、Repair 尝试和复检。
|
||||
- `SEMANTIC`:语义校验尝试与判定。
|
||||
- `RELEASE`:最终发布决策。
|
||||
|
||||
- 禁止凭记忆回答错误码、接口定义、排障步骤。
|
||||
- 需要外部信息时必须调用工具。
|
||||
- 工具连续失败或返回空结果时,最终报告必须诚实说明。
|
||||
- Chat Executor 不允许在窄范围问题中扩展根因、风险或修复建议。
|
||||
- Chat Verifier 不允许做新检索,只能判断已验真证据是否可推出 claims。
|
||||
- Chat Composer 不允许补事实,尤其不能把 `$.no_evidence` 表达为“已排除/确认没有”。
|
||||
- AIOps payload 模式必须聚焦输入告警。
|
||||
Trace 写入失败只记录警告,不应改变业务执行结果;`details` 禁止包含 Prompt、Thought、Draft 正文或 raw Tool payload。普通 Trace API 按 `sequence_no, id` 返回 Timeline。
|
||||
|
||||
Chat 链路还会在 `verifier_evaluation.prompt_audit` 中持久化紧凑 Prompt 审计快照:
|
||||
## 7. Audit 安全
|
||||
|
||||
```json
|
||||
{
|
||||
"version": "chat-prompts-v1",
|
||||
"prompts": [
|
||||
{
|
||||
"name": "chat_executor",
|
||||
"version": "chat-executor-v2",
|
||||
"resource": "prompts/chat-executor-prompt.md"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
AgentStep 不保存 Prompt、消息正文、模型正文、Tool arguments 或 Thought。ToolInvocation 不保存完整 request、SQL/日志 query、raw response 或 Agent projection。Provider reasoning 仅写入独立 `agent_reasoning_audit`,不进入普通 Trace 或发布结果;无 Provider 内容时必须记录 unavailable,不能伪造。应用日志不得打印这些字段。
|
||||
|
||||
该快照只保存版本和资源路径,不保存完整 Prompt 文本。它用于面试演示、trace 回放和离线 baseline 解释“本次诊断使用了哪套 Prompt 契约”。
|
||||
|
||||
## 4. Trace Hooks
|
||||
|
||||
`AgentLoggingHook` 是当前 Agent step 可观测性的核心。
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant A as Agent
|
||||
participant H as AgentLoggingHook
|
||||
participant DB as agent_step
|
||||
|
||||
A->>H: before_model(messages, sessionId, runId)
|
||||
H->>DB: 写入 session_id / run_id / model_input / step_index / agent_name
|
||||
A-->>A: LLM 推理
|
||||
A->>H: after_model(messages, sessionId, runId)
|
||||
H->>DB: 回填 model_output / thought / has_tool_call / duration / token_count
|
||||
```
|
||||
|
||||
记录内容:
|
||||
|
||||
- 最近输入消息摘要。
|
||||
- Agent 输出摘要。
|
||||
- 是否包含 tool call。
|
||||
- duration。
|
||||
- token count。
|
||||
- Verifier 的 JSON 输出摘要。
|
||||
|
||||
新写入必须带 `run_id`;`session_id` 仍保留用于粗粒度排查和历史兼容。
|
||||
|
||||
## 5. Tool Invocation 门禁
|
||||
|
||||
工具调用记录由 `ToolInvocationRecorder` 和具体工具共同完成。
|
||||
|
||||
核心记录:
|
||||
|
||||
```text
|
||||
tool_name
|
||||
input_params
|
||||
output_preview
|
||||
retrieval_layer
|
||||
l0_match_count
|
||||
l1_match_count
|
||||
retrieval_details
|
||||
-> evidence_refs
|
||||
relevance_level
|
||||
dedup_reason
|
||||
duration_ms
|
||||
success
|
||||
error_message
|
||||
```
|
||||
|
||||
对 `lookup_knowledge` 的质量约束:
|
||||
|
||||
- L0 只作为 hint,不绕过 L1。
|
||||
- 检索结果归一化为 `PRECISE`、`HIGHLY_RELEVANT`、`REFERENCE`。
|
||||
- 同 session 内重复文档会被 `RetrievedDocTracker` 去重。
|
||||
- dedup、no evidence、failed 等状态进入 `retrieval_details.evidence_status`。
|
||||
- `retrieval_details.evidence_refs` 记录可被 Executor 引用的最小证据文本,格式为 `raw_path + text`。
|
||||
- no-hit / no-evidence 工具结果会生成 `raw_path=$.no_evidence` 的负向证据引用,语义仅限“本次查询未检索到匹配证据”。
|
||||
|
||||
## 6. Gatekeeper 与 Verifier 门禁
|
||||
|
||||
Chat Verifier 前置一层 Gatekeeper。Gatekeeper 不调用 LLM,只用代码检查 Executor 输出的证据引用是否真实存在。
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Invocation["tool_invocation"] --> EvidenceRefs["retrieval_details.evidence_refs"]
|
||||
ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
EvidenceRefs --> Gatekeeper
|
||||
Gatekeeper --> GateResult["gatekeeper_result"]
|
||||
Invocation --> Summary["ToolTraceSummaryService"]
|
||||
Summary --> EvidenceIndex["tool_trace_summary"]
|
||||
GateResult --> Verifier["chat_verifier"]
|
||||
ExecutorOutput --> Verifier
|
||||
EvidenceIndex --> Verifier
|
||||
Verifier --> Verdict{"verdict"}
|
||||
Verdict -->|PASS| Composer["chat_composer"]
|
||||
Composer --> Pass["输出最终答复"]
|
||||
Verdict -->|LOW_CONFID| Low["补证据或低置信输出"]
|
||||
Verdict -->|REJECT| Reject["降级输出"]
|
||||
```
|
||||
|
||||
Gatekeeper 检查:
|
||||
|
||||
| 检查 | 失败语义 |
|
||||
|---|---|
|
||||
| `answer_version=executor_evidence_v2` | 非结构化或旧结构输出降为低置信 |
|
||||
| `source_invocation_id` 真实存在 | 伪造 ID 直接拒绝 |
|
||||
| `tool_name` 与 invocation 对齐 | 张冠李戴直接拒绝 |
|
||||
| `raw_path` 存在于 `evidence_refs` | 无中生有直接拒绝 |
|
||||
| `evidence_excerpt` 由 `evidence_refs[].text` 支撑 | excerpt 编造或错配直接拒绝 |
|
||||
| `negative_observation` 只能引用 `$.no_evidence` | 用正向日志证明“没查到”直接拒绝 |
|
||||
|
||||
Gatekeeper 审计还会记录 `rule_set_version` 和已启用规则元数据摘要。当前规则元数据来自本地 `gatekeeper-rules.json`,规则执行仍是确定性 Java 代码。
|
||||
|
||||
Verifier 输出:
|
||||
|
||||
```json
|
||||
{
|
||||
"verdict": "PASS|LOW_CONFID|REJECT",
|
||||
"groundedness_score": 0.8,
|
||||
"critical_fact_count": 2,
|
||||
"claim_checks": [],
|
||||
"facts_checked": [],
|
||||
"rationale": "..."
|
||||
}
|
||||
```
|
||||
|
||||
Verifier 不再逐字核验 excerpt 真伪;这由 Gatekeeper 完成。Verifier 只回答一个问题:`claim_text` 是否能由已经验真的 `evidence_excerpt` 推导出来。
|
||||
|
||||
结果写入:
|
||||
|
||||
```text
|
||||
diagnosis_run.self_evaluation.verifier_evaluation
|
||||
```
|
||||
|
||||
其中同时持久化 `executor_structured_output`、`gatekeeper_result`、`tool_trace_summary`、`prompt_audit` 和 `composer_output`,用于 Trace 回放。
|
||||
|
||||
## 7. AIOps 规则门禁
|
||||
|
||||
AIOps 当前不走 Chat Verifier,而是用 `AiOpsRuleEvaluationService` 做轻量检查。
|
||||
|
||||
检查重点:
|
||||
|
||||
- 最终报告是否存在。
|
||||
- payload 模式是否围绕输入告警展开。
|
||||
- 是否调用证据工具,尤其是 `lookup_knowledge`、日志、指标。
|
||||
- 是否把无关活跃告警扩展成主诊断对象。
|
||||
|
||||
结果写入:
|
||||
|
||||
```text
|
||||
diagnosis_run.self_evaluation.aiops_rule_evaluation
|
||||
```
|
||||
|
||||
## 8. Eval Baseline
|
||||
|
||||
当前质量门禁还包括离线评测资产:
|
||||
|
||||
| 评测 | 位置 | 作用 |
|
||||
|---|---|---|
|
||||
| Diagnosis eval | `mvp/eval/` | 检查诊断 trace、报告和证据行为 |
|
||||
| RAG retrieval eval | `eval/rag-retrieval/` | 检查固定检索 query 的召回稳定性 |
|
||||
| Live RAG acceptance | `scripts/eval_rag_live_acceptance.py` | 检查运行环境中真实 `/api/search/similar` 行为 |
|
||||
|
||||
## 9. 后续门禁规划
|
||||
|
||||
从旧版设计继承但尚未完整实现的门禁:
|
||||
|
||||
- 工具参数 schema 校验。
|
||||
- 同一工具调用次数上限。
|
||||
- 工具超时的统一熔断。
|
||||
- Gatekeeper 规则远程化或三层分离:索引层、元数据层、规则实现层。
|
||||
- Prompt 版本回滚和更细粒度变更审计。
|
||||
- Verifier 对 AIOps 报告的 LLM 级事实校验。
|
||||
|
||||
这些应在评测集扩大后逐步加入,避免一次性把诊断流程卡得过死。
|
||||
Reasoning endpoint 当前已与普通 Trace 分离并执行 `sessionId + runId` 归属校验,但访问控制、保留期限和加密要求尚未完成,继续由 ISS-015 跟踪。
|
||||
|
||||
@@ -1,160 +1,55 @@
|
||||
# 会话与 Trace 生命周期
|
||||
# Session、Run 与 Trace 生命周期
|
||||
|
||||
**更新日期**:2026-07-10
|
||||
**更新日期**:2026-07-23
|
||||
**状态**:当前可运行架构
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/session-management.md`
|
||||
|
||||
## 1. 定位
|
||||
## 1. Identity
|
||||
|
||||
当前 MVP 把“会话态”和“运行态”拆开:
|
||||
- `sessionId`:多轮对话目录,由客户端传入或应用生成。
|
||||
- `runId`:一次 Chat 执行,由 Harness 生成并在 SSE metadata 首事件返回。
|
||||
- 所有 Run、AgentStep、ToolInvocation 和 Trace 查询必须使用同一个 exact ID;禁止用“最新一条”替代。
|
||||
|
||||
## 2. 生命周期
|
||||
|
||||
```text
|
||||
chat_session(sessionId)
|
||||
-> diagnosis_run(runId)
|
||||
-> agent_step(runId)
|
||||
-> tool_invocation(runId)
|
||||
request accepted
|
||||
-> start RunContext
|
||||
-> persist diagnosis_run RUNNING
|
||||
-> append diagnosis_trace_event RUN_STARTED
|
||||
-> metadata(session_id, run_id)
|
||||
-> route / execute / guard / release
|
||||
-> SUCCESS | FALLBACK | FAILED | CANCELLED
|
||||
-> append diagnosis_trace_event RUN_FINISHED
|
||||
-> persist terminal state and budget usage
|
||||
```
|
||||
|
||||
- `sessionId` 表示多轮会话目录和 Redis 上下文。
|
||||
- `runId` 表示一次可回放诊断执行。
|
||||
- `DiagnosisTraceService` 聚合一个 run 的主记录、步骤和工具调用,形成可回放 Trace。
|
||||
- `diagnosis_session` 只保留为历史兼容和回滚表。
|
||||
disconnect、timeout 与 send failure 通过同一个 `ChatRunControl` 请求取消。正常 SSE complete 在 callback 前标记 terminal,避免误取消;late content 被 state machine 拒绝。
|
||||
|
||||
## 2. 生命周期总图
|
||||
## 3. Trace 聚合
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Start["request: chat / ai_ops"] --> Resolve["resolve sessionId"]
|
||||
Resolve --> Session["ensure chat_session metadata"]
|
||||
Session --> Run["create diagnosis_run(runId)"]
|
||||
Run --> Running["run.status = RUNNING"]
|
||||
`GET /api/diagnosis/{sessionId}/trace?runId={runId}` 聚合:
|
||||
|
||||
Running --> Agent["Agent workflow"]
|
||||
Agent --> Context["execution context(sessionId, runId)"]
|
||||
Context --> StepHook["AgentLoggingHook"]
|
||||
StepHook --> Step["agent_step(session_id, run_id)"]
|
||||
Context --> Tool["Evidence tools"]
|
||||
Tool --> Invocation["tool_invocation(session_id, run_id)"]
|
||||
Invocation --> Gatekeeper["Gatekeeper evidence validation"]
|
||||
- `chat_session` metadata。
|
||||
- exact `diagnosis_run` 状态、intent、release outcome、安全 answer 与预算。
|
||||
- `agent_step` metadata-only 模型步骤。
|
||||
- `tool_invocation` metadata-only Tool durable audit。
|
||||
- `diagnosis_trace_event` 按 `sequence_no, id` 排序的统一生命周期 Timeline。
|
||||
|
||||
Agent --> Final{"workflow result"}
|
||||
Final -->|success| Success["run.status = SUCCESS, answer saved"]
|
||||
Final -->|failed| Failed["run.status = FAILED"]
|
||||
Redis canonical invocation 不是 Trace API 的长期响应内容;它只供当前 Run EvidenceGuard 验真。
|
||||
|
||||
Success --> Evaluation["diagnosis_run.self_evaluation merge"]
|
||||
Failed --> Evaluation
|
||||
Evaluation --> Trace["GET /api/diagnosis/{sessionId}/trace?runId=..."]
|
||||
Success --> Feedback["POST /api/feedback(sessionId, runId)"]
|
||||
Feedback --> Case["useful -> case_library(run_id)"]
|
||||
```
|
||||
普通 Trace 不读取 `agent_reasoning_audit`,也不返回 reasoning 原文。Agent 模型步骤和 Timeline 只暴露 `reasoning_available`、`reasoning_bytes` 等有界 metadata。
|
||||
|
||||
## 3. ID 规则
|
||||
## 4. PreviousTurn
|
||||
|
||||
| ID | 来源 | 含义 |
|
||||
|---|---|---|
|
||||
| `sessionId` | Chat request `Id`、AIOps payload `sessionId`,缺失时由服务生成 | 多轮会话目录和 Redis 上下文 |
|
||||
| `runId` | 每次有效 Chat/AIOps 执行创建 | 一次诊断运行和 Trace 回放边界 |
|
||||
应用在创建当前 Run 前读取同 Session 最近安全发布结果。只允许结构化 PublishedResult 的固定字段进入 PreviousTurn,且执行字节上限;完整历史、失败、Fallback、Tool raw data 和 guard reason 均排除。
|
||||
|
||||
设计含义:
|
||||
## 5. Reasoning 审计查询
|
||||
|
||||
- 同一个 `sessionId` 可以贯穿多轮 Chat。
|
||||
- 每次有效 Chat/AIOps 执行都会创建新的 `runId`。
|
||||
- Trace 和 Feedback 新客户端应传 `runId`;只传 `sessionId` 时兼容解析 latest run。
|
||||
- latest run 排序使用 `diagnosis_run.created_at DESC, id DESC`,不使用 `updated_at`。
|
||||
`GET /api/diagnosis/{sessionId}/trace/reasoning?runId={runId}` 提供独立 reasoning 审计读取:
|
||||
|
||||
## 4. 运行状态流转
|
||||
- `runId` 必填,服务端先验证 Run 存在且属于 path `sessionId`,禁止跨 Session 串读。
|
||||
- 结果按 `step_index` 返回 Agent、reasoning availability、受限原文、字节数和创建时间。
|
||||
- Provider 未返回 reasoning 时仍保留 unavailable 记录,以区分“没有返回”与“审计遗漏”。
|
||||
- Reasoning 数据不回流到 PreviousTurn,不进入普通 Trace、SSE、Evidence Snapshot 或发布结果。
|
||||
|
||||
```mermaid
|
||||
stateDiagram-v2
|
||||
[*] --> PENDING
|
||||
PENDING --> RUNNING: start diagnosis
|
||||
RUNNING --> SUCCESS: workflow completed
|
||||
RUNNING --> FAILED: exception / empty state
|
||||
SUCCESS --> SUCCESS: feedback submitted
|
||||
FAILED --> FAILED: feedback submitted
|
||||
```
|
||||
|
||||
字段边界:
|
||||
|
||||
| 字段 | 所属表 | 含义 |
|
||||
|---|---|---|
|
||||
| `status` | `diagnosis_run` | 单次运行执行状态 |
|
||||
| `answer` | `diagnosis_run` | 本次运行最终报告或答复 |
|
||||
| `self_evaluation` | `diagnosis_run` | 本次运行系统自评估 JSON |
|
||||
| `feedback` | `diagnosis_run` | 本次运行用户反馈 |
|
||||
|
||||
`feedback` 不修改 `status`。一个执行成功但用户标记 `not_useful` 的 run,仍然应该是 `SUCCESS + feedback=not_useful`。
|
||||
|
||||
## 5. agent_step 写入
|
||||
|
||||
`AgentLoggingHook` 在模型调用前后写入和回填 `agent_step`。
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant Agent as Agent
|
||||
participant Hook as AgentLoggingHook
|
||||
participant DB as agent_step
|
||||
|
||||
Agent->>Hook: before_model(messages, sessionId, runId)
|
||||
Hook->>DB: insert step(session_id, run_id, model_input, step_index)
|
||||
Agent->>Hook: after_model(output, sessionId, runId)
|
||||
Hook->>DB: update model_output, duration, token_count, has_tool_call
|
||||
```
|
||||
|
||||
新写入必须带 `run_id`,同时保留 `session_id` 便于粗粒度排查。
|
||||
|
||||
## 6. tool_invocation 写入
|
||||
|
||||
工具调用记录同样通过执行上下文拿到 `sessionId + runId`:
|
||||
|
||||
```text
|
||||
ToolInvocationRecorder
|
||||
-> tool_invocation.session_id
|
||||
-> tool_invocation.run_id
|
||||
-> retrieval_details / evidence_refs
|
||||
```
|
||||
|
||||
Verifier、Gatekeeper 和 EvaluationService 应按 `run_id` 读取工具调用,避免同一 session 的其他 run 参与评分或证据校验。
|
||||
|
||||
## 7. Trace API 聚合
|
||||
|
||||
```text
|
||||
GET /api/diagnosis/{sessionId}/trace
|
||||
GET /api/diagnosis/{sessionId}/trace?runId=run-...
|
||||
```
|
||||
|
||||
聚合逻辑:
|
||||
|
||||
```text
|
||||
diagnosis_run by sessionId + runId
|
||||
+ chat_session metadata when available
|
||||
+ agent_step where run_id = runId, ordered by the Trace API
|
||||
+ tool_invocation where run_id = runId order by id
|
||||
-> DiagnosisTraceResponse
|
||||
```
|
||||
|
||||
当 `runId` 缺失时,Trace API 为兼容旧客户端解析最新 run,并在响应中返回 resolved `runId`。当 `runId` 属于其他 `sessionId` 时,API 必须拒绝,不能泄漏其他会话的 Trace。
|
||||
|
||||
## 8. Chat 与 AIOps 差异
|
||||
|
||||
| 维度 | Chat | AIOps |
|
||||
|---|---|---|
|
||||
| `agent_flow` | `CHAT` | `AI_OPS` |
|
||||
| 编排方式 | `SequentialAgent`: Planner -> Executor -> Gatekeeper -> Verifier -> Composer | `SupervisorAgent`: Planner + Executor |
|
||||
| 自评估 | `rule_evaluation` + `verifier_evaluation` | `aiops_rule_evaluation` |
|
||||
| 答案字段 | Chat 最终答复 | 告警分析报告 |
|
||||
| runId 暴露 | `/api/chat` JSON response | `/api/ai_ops` SSE metadata message |
|
||||
|
||||
## 9. 清理与边界
|
||||
|
||||
- Redis 会话历史用于多轮上下文,不是长期审计记录。
|
||||
- MySQL `diagnosis_run + agent_step + tool_invocation` 是主要可回放来源。
|
||||
- `chat_session.expires_at` 只是目录元数据;Redis 消息历史可独立过期。
|
||||
- `RetrievedDocTracker` 仍是 session 级运行时去重状态,诊断结束后清理。
|
||||
|
||||
## 10. 后续增强
|
||||
|
||||
1. Trace API 增加更结构化的 `self_evaluation` 展示。
|
||||
2. `agent_step` 与 `tool_invocation.step_id` 建立更严格关联。
|
||||
3. 旧 `diagnosis_session` 只读观察期结束后,再评估数据库层面的约束收紧或归档策略。
|
||||
该端点属于敏感审计面。当前完成了分表、独立查询和归属校验;身份认证、权限模型、保留期限、加密及真实 Provider/V015 验证尚未完成,由 ISS-015 阶段 3 收敛。
|
||||
|
||||
+21
-197
@@ -1,205 +1,29 @@
|
||||
# MVP 演示手册
|
||||
# 单 Diagnosis Agent Demo
|
||||
|
||||
本目录用于演示 MVP 从用户问题到诊断 Trace 的完整闭环。
|
||||
**更新日期**:2026-07-22
|
||||
|
||||
面试时建议先读:
|
||||
## 运行前提
|
||||
|
||||
- `ten-minute-interview-demo.md`:10 分钟现场演示脚本。
|
||||
- `interview-walkthrough.md`:面试讲解话术。
|
||||
- `evidence-pipeline-scenarios.md`:PASS / LOW_CONFID / REJECT / no-evidence 场景矩阵。
|
||||
- `trace-inspection-checklist.md`:Trace 字段检查清单。
|
||||
- `scripts/run-interview-demo-check.ps1`:面试预检脚本,包含服务可达性、Chat、Trace、反馈和 summary 输出。
|
||||
- `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。
|
||||
- `interview-q-and-a.md`:面试追问回答,覆盖 Agent 工程取舍、审计和评测。
|
||||
- `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。
|
||||
- `requests/narrow-highcpu-chat.json`:窄范围正向观察请求。
|
||||
- `requests/hikari-no-evidence-chat.json`:no-evidence 负向观察请求。
|
||||
- `requests/safety-unsupported-claim-chat.json`:安全降级讨论请求。
|
||||
- 应用、MySQL、Redis、Milvus 与模型配置可用。
|
||||
- `cls.mock-enabled=true` 用于 query_logs Mock 证据。
|
||||
- query_mysql 只使用 `mysql-tool.datasources` 配置的隔离只读数据源;不查询应用数据库。
|
||||
|
||||
## 1. 前置条件
|
||||
## 主流程
|
||||
|
||||
- MySQL、Redis、Milvus/Zilliz、LLM 和 embedding 配置可用。
|
||||
- 安全和密钥清理不属于当前 MVP 演示范围。
|
||||
- `mvp-demo` profile 会启用 mock Prometheus 和 mock CLS,让日志和指标工具返回可复现证据。
|
||||
1. 启动应用:`mvn spring-boot:run`。
|
||||
2. 向 `POST /api/chat` 提交 `requests/payment-timeout-chat.json`。
|
||||
3. 验证 SSE:`metadata -> status* -> content|failure -> done`。
|
||||
4. 保存 metadata 的 exact `session_id` 与 `run_id`。
|
||||
5. 检查 `logs/application.log` 的 Run/Tool/Guard/Release 状态,确认无 Prompt、Thought 或 raw Tool payload。
|
||||
6. 使用 `scripts/query_mysql.py` 按 exact runId 查询 `diagnosis_run`、`agent_step`、`tool_invocation`。
|
||||
7. 打开 `/trace.html?sessionId=...&runId=...` 检查聚合 Trace。
|
||||
|
||||
## 2. 启动服务
|
||||
## 验收重点
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||
```
|
||||
- 只有 `diagnosis_agent` 具有 Tool loop。
|
||||
- Tool 只包含 `lookup_knowledge`、`query_logs`、`query_mysql`。
|
||||
- EvidenceGuard/SemanticGuard 完成前没有 content。
|
||||
- SUCCESS 发布 typed diagnosis report;证据或语义不支持时发布固定 safe fallback。
|
||||
- query_logs 为 Mock;不把它表述为真实 CLS live 结果。
|
||||
|
||||
服务地址:
|
||||
|
||||
```text
|
||||
http://localhost:9900
|
||||
```
|
||||
|
||||
## 3. Chat 诊断 Demo
|
||||
|
||||
最快方式:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1
|
||||
```
|
||||
|
||||
脚本会生成:
|
||||
|
||||
```text
|
||||
mvp/demo/output/chat-response.json
|
||||
mvp/demo/output/trace-response.json
|
||||
mvp/demo/output/feedback-response.json
|
||||
mvp/demo/output/interview-demo-summary.json
|
||||
```
|
||||
|
||||
手动请求:
|
||||
|
||||
```powershell
|
||||
$sessionId = "mvp-demo-payment-timeout-001"
|
||||
$body = @{
|
||||
Id = $sessionId
|
||||
Question = "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/chat" `
|
||||
-ContentType "application/json" `
|
||||
-Body $body
|
||||
```
|
||||
|
||||
如果要继续手动查询同一次诊断运行,先保留响应中的 run id:
|
||||
|
||||
```powershell
|
||||
$chat = Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/chat" `
|
||||
-ContentType "application/json" `
|
||||
-Body $body
|
||||
|
||||
$runId = $chat.data.runId
|
||||
```
|
||||
|
||||
期望结果:
|
||||
|
||||
- `data.success = true`
|
||||
- `data.sessionId = mvp-demo-payment-timeout-001`
|
||||
- `data.runId` 为本次诊断运行的唯一 ID
|
||||
- `data.answer` 包含诊断答复
|
||||
|
||||
## 4. 查询 Trace
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace?runId=$runId"
|
||||
```
|
||||
|
||||
期望结果:
|
||||
|
||||
- `code = 200`
|
||||
- `data.runId` 等于 `$runId`
|
||||
- `data.session.sessionId` 等于 Chat session id
|
||||
- `data.run.runId` 等于 `$runId`
|
||||
- `data.steps` 包含 planner / executor / verifier 等步骤
|
||||
- `data.toolInvocations` 包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具
|
||||
- `data.session.selfEvaluation` 包含 verifier 或 rule evaluation
|
||||
- Chat V2 链路中,`data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` 记录 Chat Prompt 审计版本
|
||||
- Chat V2 链路中,`data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` 记录 Gatekeeper 规则集版本
|
||||
|
||||
## 5. 提交反馈
|
||||
|
||||
```powershell
|
||||
$feedback = @{
|
||||
sessionId = $sessionId
|
||||
runId = $runId
|
||||
feedback = "useful"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/feedback" `
|
||||
-ContentType "application/json" `
|
||||
-Body $feedback
|
||||
```
|
||||
|
||||
期望结果:
|
||||
|
||||
- `success = true`
|
||||
- `runId = $runId`
|
||||
- 后续精确 Trace 中 `data.session.feedback = useful`
|
||||
- useful 反馈会尝试沉淀 `case_library`
|
||||
|
||||
## 6. AIOps 告警诊断 Demo
|
||||
|
||||
```powershell
|
||||
$aiopsSessionId = "mvp-demo-aiops-payment-cpu-001"
|
||||
$aiopsBody = @{
|
||||
sessionId = $aiopsSessionId
|
||||
alertName = "HighCPUUsage"
|
||||
service = "payment-service"
|
||||
severity = "P1"
|
||||
description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
|
||||
timeRange = "last_15m"
|
||||
userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-WebRequest `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/ai_ops" `
|
||||
-ContentType "application/json" `
|
||||
-Body $aiopsBody
|
||||
```
|
||||
|
||||
期望结果:
|
||||
|
||||
- SSE 首条是 `type=metadata` 的 `message` 事件,包含 sessionId `mvp-demo-aiops-payment-cpu-001` 和本次 AIOps `runId`
|
||||
- 后续流式输出包含 AIOps 告警分析报告
|
||||
- 报告聚焦输入的 `HighCPUUsage/payment-service`
|
||||
- 精确 Trace 中 `data.session.agentFlow = AI_OPS`
|
||||
- `data.session.answer` 包含最终告警报告
|
||||
- `data.toolInvocations` 包含证据工具调用
|
||||
|
||||
查询 AIOps Trace 时优先使用 SSE metadata 中的 runId:
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace?runId=$aiopsRunId"
|
||||
```
|
||||
|
||||
## 7. Demo 主线
|
||||
|
||||
Chat 主线:
|
||||
|
||||
```text
|
||||
一个 session id + 一个 run id
|
||||
-> 用户问题
|
||||
-> 多 Agent 执行
|
||||
-> 证据工具
|
||||
-> Verifier / self_evaluation
|
||||
-> 最终答案
|
||||
-> 用户反馈
|
||||
-> Trace API 回放
|
||||
```
|
||||
|
||||
AIOps 主线:
|
||||
|
||||
```text
|
||||
一个 session id + 一个 run id
|
||||
-> 告警 payload
|
||||
-> AIOps Planner / Executor
|
||||
-> 证据工具
|
||||
-> 告警分析报告
|
||||
-> AIOps rule evaluation
|
||||
-> Trace API 回放
|
||||
```
|
||||
|
||||
## 8. Evidence Pipeline 场景矩阵
|
||||
|
||||
面试时不要把所有安全场景都压到 live LLM 现场表现上。建议使用:
|
||||
|
||||
- `scripts/run-interview-demo-check.ps1` 跑主路径和预检 summary。
|
||||
- `evidence-pipeline-scenarios.md` 讲解 PASS / LOW_CONFID / REJECT / no-evidence 矩阵。
|
||||
- `mvp/eval/reports/baseline-report.md` 证明固定 fixture 12/12 通过。
|
||||
|
||||
这样可以同时展示真实链路和确定性回归能力。
|
||||
旧多角色与第二诊断入口的 demo 已保存在 `archive/2026-07-22-legacy/`,不代表当前运行时。
|
||||
|
||||
@@ -0,0 +1,205 @@
|
||||
# MVP 演示手册
|
||||
|
||||
本目录用于演示 MVP 从用户问题到诊断 Trace 的完整闭环。
|
||||
|
||||
面试时建议先读:
|
||||
|
||||
- `ten-minute-interview-demo.md`:10 分钟现场演示脚本。
|
||||
- `interview-walkthrough.md`:面试讲解话术。
|
||||
- `evidence-pipeline-scenarios.md`:PASS / LOW_CONFID / REJECT / no-evidence 场景矩阵。
|
||||
- `trace-inspection-checklist.md`:Trace 字段检查清单。
|
||||
- `scripts/run-interview-demo-check.ps1`:面试预检脚本,包含服务可达性、Chat、Trace、反馈和 summary 输出。
|
||||
- `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。
|
||||
- `interview-q-and-a.md`:面试追问回答,覆盖 Agent 工程取舍、审计和评测。
|
||||
- `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。
|
||||
- `requests/narrow-highcpu-chat.json`:窄范围正向观察请求。
|
||||
- `requests/hikari-no-evidence-chat.json`:no-evidence 负向观察请求。
|
||||
- `requests/safety-unsupported-claim-chat.json`:安全降级讨论请求。
|
||||
|
||||
## 1. 前置条件
|
||||
|
||||
- MySQL、Redis、Milvus/Zilliz、LLM 和 embedding 配置可用。
|
||||
- 安全和密钥清理不属于当前 MVP 演示范围。
|
||||
- `mvp-demo` profile 会启用 mock Prometheus 和 mock CLS,让日志和指标工具返回可复现证据。
|
||||
|
||||
## 2. 启动服务
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||
```
|
||||
|
||||
服务地址:
|
||||
|
||||
```text
|
||||
http://localhost:9900
|
||||
```
|
||||
|
||||
## 3. Chat 诊断 Demo
|
||||
|
||||
最快方式:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1
|
||||
```
|
||||
|
||||
脚本会生成:
|
||||
|
||||
```text
|
||||
mvp/demo/output/chat-response.json
|
||||
mvp/demo/output/trace-response.json
|
||||
mvp/demo/output/feedback-response.json
|
||||
mvp/demo/output/interview-demo-summary.json
|
||||
```
|
||||
|
||||
手动请求:
|
||||
|
||||
```powershell
|
||||
$sessionId = "mvp-demo-payment-timeout-001"
|
||||
$body = @{
|
||||
Id = $sessionId
|
||||
Question = "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/chat" `
|
||||
-ContentType "application/json" `
|
||||
-Body $body
|
||||
```
|
||||
|
||||
如果要继续手动查询同一次诊断运行,先保留响应中的 run id:
|
||||
|
||||
```powershell
|
||||
$chat = Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/chat" `
|
||||
-ContentType "application/json" `
|
||||
-Body $body
|
||||
|
||||
$runId = $chat.data.runId
|
||||
```
|
||||
|
||||
期望结果:
|
||||
|
||||
- `data.success = true`
|
||||
- `data.sessionId = mvp-demo-payment-timeout-001`
|
||||
- `data.runId` 为本次诊断运行的唯一 ID
|
||||
- `data.answer` 包含诊断答复
|
||||
|
||||
## 4. 查询 Trace
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace?runId=$runId"
|
||||
```
|
||||
|
||||
期望结果:
|
||||
|
||||
- `code = 200`
|
||||
- `data.runId` 等于 `$runId`
|
||||
- `data.session.sessionId` 等于 Chat session id
|
||||
- `data.run.runId` 等于 `$runId`
|
||||
- `data.steps` 包含 planner / executor / verifier 等步骤
|
||||
- `data.toolInvocations` 包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具
|
||||
- `data.session.selfEvaluation` 包含 verifier 或 rule evaluation
|
||||
- Chat V2 链路中,`data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` 记录 Chat Prompt 审计版本
|
||||
- Chat V2 链路中,`data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` 记录 Gatekeeper 规则集版本
|
||||
|
||||
## 5. 提交反馈
|
||||
|
||||
```powershell
|
||||
$feedback = @{
|
||||
sessionId = $sessionId
|
||||
runId = $runId
|
||||
feedback = "useful"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/feedback" `
|
||||
-ContentType "application/json" `
|
||||
-Body $feedback
|
||||
```
|
||||
|
||||
期望结果:
|
||||
|
||||
- `success = true`
|
||||
- `runId = $runId`
|
||||
- 后续精确 Trace 中 `data.session.feedback = useful`
|
||||
- useful 反馈会尝试沉淀 `case_library`
|
||||
|
||||
## 6. AIOps 告警诊断 Demo
|
||||
|
||||
```powershell
|
||||
$aiopsSessionId = "mvp-demo-aiops-payment-cpu-001"
|
||||
$aiopsBody = @{
|
||||
sessionId = $aiopsSessionId
|
||||
alertName = "HighCPUUsage"
|
||||
service = "payment-service"
|
||||
severity = "P1"
|
||||
description = "服务 payment-service 的 CPU 使用率持续超过 80%,当前值为 92%。实例: pod-payment-service-7d8f9c6b5-x2k4m。"
|
||||
timeRange = "last_15m"
|
||||
userRequest = "请结合 Prometheus 活动告警、system-metrics 日志和知识库生成告警分析报告。"
|
||||
} | ConvertTo-Json
|
||||
|
||||
Invoke-WebRequest `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/ai_ops" `
|
||||
-ContentType "application/json" `
|
||||
-Body $aiopsBody
|
||||
```
|
||||
|
||||
期望结果:
|
||||
|
||||
- SSE 首条是 `type=metadata` 的 `message` 事件,包含 sessionId `mvp-demo-aiops-payment-cpu-001` 和本次 AIOps `runId`
|
||||
- 后续流式输出包含 AIOps 告警分析报告
|
||||
- 报告聚焦输入的 `HighCPUUsage/payment-service`
|
||||
- 精确 Trace 中 `data.session.agentFlow = AI_OPS`
|
||||
- `data.session.answer` 包含最终告警报告
|
||||
- `data.toolInvocations` 包含证据工具调用
|
||||
|
||||
查询 AIOps Trace 时优先使用 SSE metadata 中的 runId:
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace?runId=$aiopsRunId"
|
||||
```
|
||||
|
||||
## 7. Demo 主线
|
||||
|
||||
Chat 主线:
|
||||
|
||||
```text
|
||||
一个 session id + 一个 run id
|
||||
-> 用户问题
|
||||
-> 多 Agent 执行
|
||||
-> 证据工具
|
||||
-> Verifier / self_evaluation
|
||||
-> 最终答案
|
||||
-> 用户反馈
|
||||
-> Trace API 回放
|
||||
```
|
||||
|
||||
AIOps 主线:
|
||||
|
||||
```text
|
||||
一个 session id + 一个 run id
|
||||
-> 告警 payload
|
||||
-> AIOps Planner / Executor
|
||||
-> 证据工具
|
||||
-> 告警分析报告
|
||||
-> AIOps rule evaluation
|
||||
-> Trace API 回放
|
||||
```
|
||||
|
||||
## 8. Evidence Pipeline 场景矩阵
|
||||
|
||||
面试时不要把所有安全场景都压到 live LLM 现场表现上。建议使用:
|
||||
|
||||
- `scripts/run-interview-demo-check.ps1` 跑主路径和预检 summary。
|
||||
- `evidence-pipeline-scenarios.md` 讲解 PASS / LOW_CONFID / REJECT / no-evidence 矩阵。
|
||||
- `mvp/eval/reports/baseline-report.md` 证明固定 fixture 12/12 通过。
|
||||
|
||||
这样可以同时展示真实链路和确定性回归能力。
|
||||
@@ -0,0 +1,3 @@
|
||||
# Archive Note
|
||||
|
||||
本目录保存旧多角色与旧诊断入口 demo。当前 demo 以 `mvp/demo/README.md` 和唯一 `/api/chat` named SSE 为准。
|
||||
@@ -0,0 +1,58 @@
|
||||
# Trace 检查清单
|
||||
|
||||
运行 `scripts/run-interview-demo-check.ps1` 后,用这份清单检查 `trace-response.json` 和 `interview-demo-summary.json`。
|
||||
|
||||
## 1. Session
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
| `data.runId` / `data.run.runId` | 是否等于 demo 响应中的 `runId` | `runId` 精确绑定这一次诊断运行 |
|
||||
| `data.session.sessionId` | 是否等于 `mvp-demo-payment-timeout-001` | `sessionId` 保留多轮上下文,Trace 精确回放依赖 `runId` |
|
||||
| `data.session.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
|
||||
| `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
|
||||
| `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | 如果是 Chat V2 链路,是否记录 Gatekeeper 规则版本 | 安全规则可审计、可回归 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` | 如果是 Chat V2 链路,是否记录 Prompt 审计版本 | Prompt 变更可解释、可回归 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.prompts[*].version` | 是否记录 planner / executor / verifier / composer 版本 | 便于定位 Prompt 变更影响 |
|
||||
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在当前 diagnosis run 上 |
|
||||
|
||||
## 2. Agent 步骤
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
| `data.steps[*].agentName` | 是否有 Planner / Executor / Verifier 或等价步骤 | 流程被拆成可检查的 Agent 步骤 |
|
||||
| `data.steps[*].thought` | 是否有高层步骤摘要 | 内部过程可审计,不只看最终文本 |
|
||||
| `data.steps[*].durationMs` | 是否有步骤耗时 | Trace 可用于耗时分析 |
|
||||
| `data.steps[*].tokenCount` | 如可用,是否记录 token | Trace 可用于模型成本分析 |
|
||||
|
||||
## 3. 工具证据
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
| `data.toolInvocations[*].toolName` | 是否包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具 | Agent 通过工具收集证据,而不是无依据猜测 |
|
||||
| `data.toolInvocations[*].inputParams` | 是否能看到每个工具的入参 | 工具输入可审计、可调试 |
|
||||
| `data.toolInvocations[*].outputPreview` | 是否有受控长度的证据预览 | 保留证据但不倾倒巨大 payload |
|
||||
| `data.toolInvocations[*].success` | 是否区分成功和失败 | 工具失败对 Verifier 和 reviewer 可见 |
|
||||
| `data.toolInvocations[*].retrievalDetails` | 是否包含检索 metadata | 检索质量可事后检查 |
|
||||
| `data.toolInvocations[*].retrievalDetails.evidence_refs` | 是否包含 `raw_path + text` | Gatekeeper 可以用代码核对 Executor 引用 |
|
||||
| `data.toolInvocations[*].relevanceLevel` | 是否有相关性等级 | 可解释检索结果强弱 |
|
||||
|
||||
## 4. Summary
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
| `data.summary.persistedStepCount` | step 行是否持久化 | Trace 来自存储,不是响应内存 |
|
||||
| `data.summary.persistedToolCallCount` | tool 行是否持久化 | 工具证据在请求结束后仍可回放 |
|
||||
| `data.summary.hasVerifierEvaluation` | 是否存在 Verifier 结果 | 最终答案经过质量门 |
|
||||
| `data.summary.hasFeedback` | 提交反馈后是否为 true | 人类反馈闭环完成 |
|
||||
|
||||
## 5. 好的结果长什么样
|
||||
|
||||
```text
|
||||
同一个 session id + run id
|
||||
-> 最终答案
|
||||
-> 持久化 agent steps
|
||||
-> 持久化 evidence tool calls
|
||||
-> verifier / self-evaluation
|
||||
-> feedback attached to the same run
|
||||
```
|
||||
@@ -2,11 +2,13 @@
|
||||
|
||||
本目录是本地 Demo 响应的默认输出位置。
|
||||
|
||||
生成文件会被 Git 忽略:
|
||||
当前 named SSE Demo 生成以下文件,均被 Git 忽略:
|
||||
|
||||
- `chat-response.json`
|
||||
- `chat-sse.txt`
|
||||
- `chat-events.json`
|
||||
- `trace-response.json`
|
||||
- `feedback-response.json`
|
||||
|
||||
目录中可能存在旧版 Demo 生成的 `chat-response.json`、`feedback-response.json` 或历史 Trace;它们不是当前架构的验收证据。阶段验收必须使用本次 SSE metadata 返回的 exact `session_id + run_id` 重新生成结果。
|
||||
|
||||
保留此 README 是为了让目录存在于仓库中。
|
||||
|
||||
|
||||
@@ -2,41 +2,34 @@
|
||||
|
||||
## 1. 目标
|
||||
|
||||
验证 MVP 能诊断支付超时问题,并暴露完整 Trace 供回放。
|
||||
验证当前单 Diagnosis Agent 能诊断支付超时问题,并为一次精确 Run 暴露安全、可核对的 Trace。
|
||||
|
||||
## 2. 输入
|
||||
|
||||
- Session id:`mvp-demo-payment-timeout-001`
|
||||
- 问题:`支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。`
|
||||
- `session_id`:`mvp-demo-payment-timeout-001`
|
||||
- 问题:结合知识库和日志证据诊断支付超时,并给出有边界的修复建议
|
||||
- Profile:`mvp-demo`
|
||||
|
||||
## 3. 验收标准
|
||||
|
||||
1. Chat 返回成功答复,且 session id 与请求一致,并返回本次诊断的 run id。
|
||||
2. Trace API 使用 `sessionId + runId` 返回会话元数据、运行摘要、最终答案、按顺序排列的 agent steps 和 tool invocations。
|
||||
3. Trace 中有足够证据说明用了哪些工具,以及 verifier / self-evaluation 是否已持久化。
|
||||
4. 可以使用同一个 session id 和本次 run id 提交反馈。
|
||||
5. 后续精确 Trace 查询能看到已持久化的 feedback 值。
|
||||
1. `POST /api/chat` 按 `metadata -> status* -> content|failure -> done` 顺序发送 named SSE;`content` 与 `failure` 必须互斥且只出现一次。
|
||||
2. `metadata.session_id` 和 `metadata.run_id` 非空;该精确 ID 对能唯一定位持久化 Run 与 Trace。
|
||||
3. Run 的 `intent=DIAGNOSIS`,终态、`release_outcome` 与 SSE `done` outcome 一致。
|
||||
4. `agent_step.agent_name` 只出现 `diagnosis_agent`。AgentStep 只保存消息数量/角色、输出是否存在、Tool 名称等 metadata,不保存 Prompt、消息正文、模型正文或 Thought。
|
||||
5. 每条 `tool_invocation` 都属于精确 `run_id`,Tool 名称属于 ACI allowlist,只保存有界的身份、状态、错误码、耗时和字节数 metadata;不得包含 SQL、日志查询正文、raw response、凭据或 evidence body。
|
||||
6. `content` 只在 Harness guards 与 Release Policy 完成后发布;guard 或技术失败只能发布固定安全 fallback,不能泄漏 Agent 或 Tool 原始 JSON。
|
||||
7. `query_logs` 明确标记为 Mock。`query_mysql` 只针对配置的隔离只读数据源验证;本验收不声称接入真实 CLS 或生产业务 MySQL。
|
||||
|
||||
## 4. 需要检查的 Trace 字段
|
||||
|
||||
- `data.runId`
|
||||
- `data.run.runId`
|
||||
- `data.session.query`
|
||||
- `data.session.answer`
|
||||
- `data.session.selfEvaluation`
|
||||
- `data.session.feedback`
|
||||
- `data.steps[*].agentName`
|
||||
- `data.steps[*].thought`
|
||||
- `data.toolInvocations[*].toolName`
|
||||
- `data.toolInvocations[*].inputParams`
|
||||
- `data.toolInvocations[*].outputPreview`
|
||||
- `data.toolInvocations[*].retrievalDetails`
|
||||
- `data.runId` 与 `data.run.runId`
|
||||
- `data.session.sessionId`、`data.session.intent`、`data.session.releaseOutcome`
|
||||
- `data.steps[*].agentName`、`data.steps[*].thought` 与 metadata 字段
|
||||
- `data.toolInvocations[*].toolName`、identity、status、error code、duration 与 size 字段
|
||||
- `data.summary`
|
||||
|
||||
## 5. 已知边界
|
||||
|
||||
- 这不是完整离线测试,仍需要有效的 chat、持久化、向量检索和模型调用环境。
|
||||
- `mvp-demo` profile 启用 mock 日志和指标,让证据工具返回更稳定。
|
||||
- 敏感配置清理不属于当前 MVP 优先级。
|
||||
|
||||
- 这是一次真实应用 E2E 验收,不能替代确定性的单元与契约测试。
|
||||
- 外部模型、Redis、Milvus 和隔离数据源的可用性可能影响 live Run;失败时必须记录精确 session/run identity。
|
||||
- 敏感配置治理和生产 CLS/业务 MySQL 接入不属于本次 MVP 验收范围。
|
||||
|
||||
@@ -0,0 +1,4 @@
|
||||
{
|
||||
"Id": "mvp-demo-payment-redis-timeout-success-001",
|
||||
"Question": "诊断 payment-service 的 Redis 连接超时。只调用 query_logs 一次,日志范围使用 APPLICATION,查询词使用 payment-service redis,不调用其他工具。若返回 EVIDENCE_FOUND,只陈述该工具结果中明确出现的 payment-service Redis 超时事实;不要把同批结果里的其他服务事件写入分析或结论。结论必须严格由这条日志事实支持,action_plan 和 recommendations 保持为空;limitations.scope 写明本次 Mock 日志查询范围,limitations.missing_info 写明尚未用生产日志确认。"
|
||||
}
|
||||
@@ -1,4 +1,4 @@
|
||||
{
|
||||
"Id": "mvp-demo-payment-timeout-001",
|
||||
"Question": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
|
||||
"Question": "支付接口最近出现超时。在给出结论前必须实际调用 lookup_knowledge 和 query_logs 各一次,日志范围使用 APPLICATION 并查询 payment-service error slow database;随后结束诊断,证据不足时明确说明缺口。"
|
||||
}
|
||||
|
||||
@@ -7,55 +7,122 @@ param(
|
||||
|
||||
$ErrorActionPreference = "Stop"
|
||||
|
||||
function ConvertFrom-NamedSse {
|
||||
param([Parameter(Mandatory = $true)][string]$Content)
|
||||
|
||||
$events = @()
|
||||
foreach ($frame in ($Content -split "(?:\r?\n){2,}")) {
|
||||
if ([string]::IsNullOrWhiteSpace($frame)) {
|
||||
continue
|
||||
}
|
||||
|
||||
$name = $null
|
||||
$dataLines = @()
|
||||
foreach ($line in ($frame -split "\r?\n")) {
|
||||
if ($line.StartsWith("event:")) {
|
||||
$name = $line.Substring(6).Trim()
|
||||
} elseif ($line.StartsWith("data:")) {
|
||||
$dataLines += $line.Substring(5).TrimStart()
|
||||
}
|
||||
}
|
||||
|
||||
if ([string]::IsNullOrWhiteSpace($name) -or $dataLines.Count -eq 0) {
|
||||
throw "Invalid named SSE frame: $frame"
|
||||
}
|
||||
$rawData = $dataLines -join "`n"
|
||||
$events += [pscustomobject]@{
|
||||
name = $name
|
||||
payload = $rawData | ConvertFrom-Json
|
||||
}
|
||||
}
|
||||
return @($events)
|
||||
}
|
||||
|
||||
function Assert-ChatSseContract {
|
||||
param(
|
||||
[Parameter(Mandatory = $true)][array]$Events,
|
||||
[Parameter(Mandatory = $true)][string]$ExpectedSessionId
|
||||
)
|
||||
|
||||
if ($Events.Count -lt 3) {
|
||||
throw "Chat SSE must contain metadata, a terminal event, and done"
|
||||
}
|
||||
if ($Events[0].name -ne "metadata") {
|
||||
throw "First Chat SSE event must be metadata"
|
||||
}
|
||||
if ($Events[-1].name -ne "done") {
|
||||
throw "Last Chat SSE event must be done"
|
||||
}
|
||||
|
||||
$terminalEvents = @($Events | Where-Object { $_.name -in @("content", "failure") })
|
||||
if ($terminalEvents.Count -ne 1 -or $Events[-2].name -ne $terminalEvents[0].name) {
|
||||
throw "Chat SSE must contain exactly one content or failure immediately before done"
|
||||
}
|
||||
|
||||
$allowed = @("metadata", "status", "content", "failure", "done")
|
||||
$unknown = @($Events | Where-Object { $_.name -notin $allowed })
|
||||
if ($unknown.Count -gt 0) {
|
||||
throw "Chat SSE contains unknown events: $($unknown.name -join ', ')"
|
||||
}
|
||||
$invalidMiddle = @()
|
||||
if ($Events.Count -gt 3) {
|
||||
$invalidMiddle = @($Events[1..($Events.Count - 3)] |
|
||||
Where-Object { $_.name -ne "status" })
|
||||
}
|
||||
if ($invalidMiddle.Count -gt 0) {
|
||||
throw "Only status events are allowed between metadata and the terminal event"
|
||||
}
|
||||
|
||||
$metadata = $Events[0].payload
|
||||
if ($metadata.session_id -ne $ExpectedSessionId) {
|
||||
throw "SSE session_id does not match the requested SessionId"
|
||||
}
|
||||
if ([string]::IsNullOrWhiteSpace([string]$metadata.run_id)) {
|
||||
throw "SSE metadata is missing run_id"
|
||||
}
|
||||
|
||||
$outcome = [string]$Events[-1].payload.outcome
|
||||
if ($terminalEvents[0].name -eq "failure" -and $outcome -ne "FAILED") {
|
||||
throw "A failure event must end with outcome FAILED"
|
||||
}
|
||||
if ($terminalEvents[0].name -eq "content" -and $outcome -notin @("SUCCESS", "FALLBACK")) {
|
||||
throw "A content event must end with outcome SUCCESS or FALLBACK"
|
||||
}
|
||||
}
|
||||
|
||||
New-Item -ItemType Directory -Force -Path $OutputDir | Out-Null
|
||||
|
||||
$request = Get-Content -Raw -Encoding UTF8 -Path $RequestFile | ConvertFrom-Json
|
||||
$request.Id = $SessionId
|
||||
$body = $request | ConvertTo-Json -Depth 8
|
||||
|
||||
Write-Host "正在运行支付超时 Chat 诊断 Demo..."
|
||||
Write-Host "Running payment-timeout Chat E2E"
|
||||
Write-Host "BaseUrl: $BaseUrl"
|
||||
Write-Host "SessionId: $SessionId"
|
||||
|
||||
$chat = Invoke-RestMethod `
|
||||
$response = Invoke-WebRequest `
|
||||
-UseBasicParsing `
|
||||
-Method Post `
|
||||
-Uri "$BaseUrl/api/chat" `
|
||||
-Headers @{ Accept = "text/event-stream" } `
|
||||
-ContentType "application/json; charset=utf-8" `
|
||||
-Body $body
|
||||
|
||||
$chat | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-response.json"
|
||||
Write-Host "已保存 Chat 响应: $OutputDir/chat-response.json"
|
||||
$response.Content | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-sse.txt"
|
||||
$events = @(ConvertFrom-NamedSse -Content $response.Content)
|
||||
Assert-ChatSseContract -Events $events -ExpectedSessionId $SessionId
|
||||
$events | ConvertTo-Json -Depth 50 | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-events.json"
|
||||
|
||||
$runId = $chat.data.runId
|
||||
if (-not $runId) {
|
||||
throw "Chat 响应缺少 runId,无法查询精确 Trace。"
|
||||
}
|
||||
$runId = [string]$events[0].payload.run_id
|
||||
Write-Host "RunId: $runId"
|
||||
Write-Host "SSE sequence: $($events.name -join ' -> ')"
|
||||
|
||||
$trace = Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "$BaseUrl/api/diagnosis/$SessionId/trace?runId=$([System.Uri]::EscapeDataString($runId))"
|
||||
|
||||
$trace | ConvertTo-Json -Depth 50 | Set-Content -Encoding UTF8 -Path "$OutputDir/trace-response.json"
|
||||
Write-Host "已保存 Trace 响应: $OutputDir/trace-response.json"
|
||||
|
||||
$feedbackBody = @{
|
||||
sessionId = $SessionId
|
||||
runId = $runId
|
||||
feedback = "useful"
|
||||
} | ConvertTo-Json
|
||||
|
||||
$feedback = Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "$BaseUrl/api/feedback" `
|
||||
-ContentType "application/json; charset=utf-8" `
|
||||
-Body $feedbackBody
|
||||
|
||||
$feedback | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/feedback-response.json"
|
||||
Write-Host "已保存反馈响应: $OutputDir/feedback-response.json"
|
||||
|
||||
Write-Host ""
|
||||
Write-Host "Demo 已完成,请检查:"
|
||||
Write-Host "- mvp/demo/output/chat-response.json"
|
||||
Write-Host "- mvp/demo/output/trace-response.json"
|
||||
Write-Host "- mvp/demo/output/feedback-response.json"
|
||||
Write-Host "E2E artifacts:"
|
||||
Write-Host "- $OutputDir/chat-sse.txt"
|
||||
Write-Host "- $OutputDir/chat-events.json"
|
||||
Write-Host "- $OutputDir/trace-response.json"
|
||||
|
||||
@@ -1,58 +1,12 @@
|
||||
# Trace 检查清单
|
||||
|
||||
运行 `scripts/run-interview-demo-check.ps1` 后,用这份清单检查 `trace-response.json` 和 `interview-demo-summary.json`。
|
||||
|
||||
## 1. Session
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
| `data.runId` / `data.run.runId` | 是否等于 demo 响应中的 `runId` | `runId` 精确绑定这一次诊断运行 |
|
||||
| `data.session.sessionId` | 是否等于 `mvp-demo-payment-timeout-001` | `sessionId` 保留多轮上下文,Trace 精确回放依赖 `runId` |
|
||||
| `data.session.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
|
||||
| `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
|
||||
| `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | 如果是 Chat V2 链路,是否记录 Gatekeeper 规则版本 | 安全规则可审计、可回归 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` | 如果是 Chat V2 链路,是否记录 Prompt 审计版本 | Prompt 变更可解释、可回归 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.prompts[*].version` | 是否记录 planner / executor / verifier / composer 版本 | 便于定位 Prompt 变更影响 |
|
||||
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在当前 diagnosis run 上 |
|
||||
|
||||
## 2. Agent 步骤
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
| `data.steps[*].agentName` | 是否有 Planner / Executor / Verifier 或等价步骤 | 流程被拆成可检查的 Agent 步骤 |
|
||||
| `data.steps[*].thought` | 是否有高层步骤摘要 | 内部过程可审计,不只看最终文本 |
|
||||
| `data.steps[*].durationMs` | 是否有步骤耗时 | Trace 可用于耗时分析 |
|
||||
| `data.steps[*].tokenCount` | 如可用,是否记录 token | Trace 可用于模型成本分析 |
|
||||
|
||||
## 3. 工具证据
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
| `data.toolInvocations[*].toolName` | 是否包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具 | Agent 通过工具收集证据,而不是无依据猜测 |
|
||||
| `data.toolInvocations[*].inputParams` | 是否能看到每个工具的入参 | 工具输入可审计、可调试 |
|
||||
| `data.toolInvocations[*].outputPreview` | 是否有受控长度的证据预览 | 保留证据但不倾倒巨大 payload |
|
||||
| `data.toolInvocations[*].success` | 是否区分成功和失败 | 工具失败对 Verifier 和 reviewer 可见 |
|
||||
| `data.toolInvocations[*].retrievalDetails` | 是否包含检索 metadata | 检索质量可事后检查 |
|
||||
| `data.toolInvocations[*].retrievalDetails.evidence_refs` | 是否包含 `raw_path + text` | Gatekeeper 可以用代码核对 Executor 引用 |
|
||||
| `data.toolInvocations[*].relevanceLevel` | 是否有相关性等级 | 可解释检索结果强弱 |
|
||||
|
||||
## 4. Summary
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
| `data.summary.persistedStepCount` | step 行是否持久化 | Trace 来自存储,不是响应内存 |
|
||||
| `data.summary.persistedToolCallCount` | tool 行是否持久化 | 工具证据在请求结束后仍可回放 |
|
||||
| `data.summary.hasVerifierEvaluation` | 是否存在 Verifier 结果 | 最终答案经过质量门 |
|
||||
| `data.summary.hasFeedback` | 提交反馈后是否为 true | 人类反馈闭环完成 |
|
||||
|
||||
## 5. 好的结果长什么样
|
||||
|
||||
```text
|
||||
同一个 session id + run id
|
||||
-> 最终答案
|
||||
-> 持久化 agent steps
|
||||
-> 持久化 evidence tool calls
|
||||
-> verifier / self-evaluation
|
||||
-> feedback attached to the same run
|
||||
```
|
||||
- [ ] SSE metadata 的 `session_id`、`run_id` 非空且与数据库完全一致。
|
||||
- [ ] `diagnosis_run.intent=DIAGNOSIS`,status/release_outcome 与 done outcome 一致。
|
||||
- [ ] `agent_step.agent_name` 只出现 `diagnosis_agent`。
|
||||
- [ ] AgentStep model_input/model_output 只含 metadata,thought 为空。
|
||||
- [ ] ToolInvocation 全部属于 exact runId,Tool 名在 ACI allowlist 内。
|
||||
- [ ] ToolInvocation input/output/retrieval details 不含 SQL、日志 query、raw response 或 evidence body。
|
||||
- [ ] Run 的模型/Tool/Token/字节预算均未超过集中配置。
|
||||
- [ ] content 只出现一次并来自 Release Policy;failure 与 content 互斥。
|
||||
- [ ] `logs/application.log` 不含 Prompt、Thought、完整 Tool 参数、raw response、vendor exception 或 stack 泄漏。
|
||||
- [ ] query_logs 标记为 Mock;query_mysql 只使用隔离只读 datasource contract。
|
||||
|
||||
+11
-10
@@ -1,6 +1,6 @@
|
||||
# MVP Issues 索引
|
||||
|
||||
**更新日期**:2026-07-20
|
||||
**更新日期**:2026-07-23
|
||||
**状态**:按活跃问题、设计笔记、RAG 问题集和已归档问题整理
|
||||
|
||||
## 目录约定
|
||||
@@ -9,20 +9,14 @@
|
||||
|---|---|
|
||||
| [active/](active/) | 仍需要规划或实现的问题 |
|
||||
| [design-notes/](design-notes/) | 已形成方向、用于指导后续实现的设计记录 |
|
||||
| [rag/](rag/) | RAG 子问题集合;多数已合并到 RAG 重构计划 |
|
||||
| [rag/](rag/) | RAG 历史子问题集合;不代表当前实施计划 |
|
||||
| [archived/](archived/) | 已修复、已实施或已归档的问题 |
|
||||
|
||||
## 活跃问题
|
||||
|
||||
| 名称 | 标题 | 严重程度 | 状态 | 文件 |
|
||||
|---|---|---|---|---|
|
||||
| ISS-003 | MVP 设计与实现 Review 收敛 | 高 | 待规划 | [active/ISS-003-mvp-design-implementation-review.md](active/ISS-003-mvp-design-implementation-review.md) |
|
||||
| ISS-004 | Executor 域级检索水位控制 | 低 | 待规划 | [active/ISS-004-executor-domain-hard-limit.md](active/ISS-004-executor-domain-hard-limit.md) |
|
||||
| ISS-012 | Executor Token 预算与上下文膨胀 | 高 | 待规划 | [active/ISS-012-executor-token-budget-and-context-growth.md](active/ISS-012-executor-token-budget-and-context-growth.md) |
|
||||
| ISS-013 | Chat 入口解耦与真正 SSE 收敛 | 高 | 待规划 | [active/ISS-013-chat-entry-decoupling-and-sse.md](active/ISS-013-chat-entry-decoupling-and-sse.md) |
|
||||
| ISS-014 | 单体 ReAct Agent、Harness 与 ACI 工具瘦身 | 高 | 待阶段 0 冻结 | [active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md](active/ISS-014-single-react-agent-harness-aci-ptk-refactor.md) |
|
||||
| executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [active/executor-evidence-attribution-hallucination.md](active/executor-evidence-attribution-hallucination.md) |
|
||||
| rag-refactor-plan | RAG 检索重构计划 | 高 | 待规划 | [active/rag-refactor-plan.md](active/rag-refactor-plan.md) |
|
||||
| ISS-015 | 诊断运行质量与 Reasoning 审计收敛 | 高 | 待实施 | [active/ISS-015-diagnosis-runtime-quality-and-reasoning-audit.md](active/ISS-015-diagnosis-runtime-quality-and-reasoning-audit.md) |
|
||||
|
||||
## 设计笔记
|
||||
|
||||
@@ -33,7 +27,7 @@
|
||||
|
||||
## RAG 问题集
|
||||
|
||||
这些问题已经收敛到 [active/rag-refactor-plan.md](active/rag-refactor-plan.md),单个文件保留用于追溯原始问题和设计背景。
|
||||
这些问题曾经收敛到 RAG 重构计划,单个文件保留用于追溯原始问题和设计背景。原计划已因架构变化归档,不作为当前实施依据。
|
||||
|
||||
| 名称 | 标题 | 状态 | 文件 |
|
||||
|---|---|---|---|
|
||||
@@ -57,12 +51,19 @@
|
||||
|---|---|---|---|
|
||||
| ISS-001 | Executor 重复召回同一文档 | 已修复 | [archived/ISS-001-duplicate-retrieval.md](archived/ISS-001-duplicate-retrieval.md) |
|
||||
| ISS-002 | Executor 无约束重复调用 lookup_knowledge | 已修复 | [archived/ISS-002-executor-unconstrained-lookup.md](archived/ISS-002-executor-unconstrained-lookup.md) |
|
||||
| ISS-003 | MVP 设计与实现 Review 收敛 | Review 基线过时 | [archived/ISS-003-mvp-design-implementation-review.md](archived/ISS-003-mvp-design-implementation-review.md) |
|
||||
| ISS-004 | Executor 域级检索水位控制 | 已被 ISS-014 吸收 | [archived/ISS-004-executor-domain-hard-limit.md](archived/ISS-004-executor-domain-hard-limit.md) |
|
||||
| ISS-005 | 证据链补齐与降级契约收敛 | 已归档 | [archived/ISS-005-evidence-trace-hardening.md](archived/ISS-005-evidence-trace-hardening.md) |
|
||||
| ISS-006 | 固定诊断评测集与回归 Harness | 已归档 | [archived/ISS-006-diagnosis-eval-harness.md](archived/ISS-006-diagnosis-eval-harness.md) |
|
||||
| ISS-007 | Verifier 证据摘要保真与工具命中质量问题 | 已实施 | [archived/ISS-007-verifier-evidence-summary-fidelity.md](archived/ISS-007-verifier-evidence-summary-fidelity.md) |
|
||||
| ISS-008 | Executor 窄范围查询越界 | 已修复 | [archived/ISS-008-executor-narrow-scope-overreach.md](archived/ISS-008-executor-narrow-scope-overreach.md) |
|
||||
| ISS-009 | negative_observation 精确引用 no-evidence 结果 | 已修复 | [archived/ISS-009-negative-observation-no-evidence-reference.md](archived/ISS-009-negative-observation-no-evidence-reference.md) |
|
||||
| ISS-010 | 同 session 多轮诊断 Trace 隔离 | 已归档 | [archived/ISS-010-session-run-trace-isolation.md](archived/ISS-010-session-run-trace-isolation.md) |
|
||||
| ISS-012 | Executor Token 预算与上下文膨胀 | 已归档 | [archived/ISS-012-executor-token-budget-and-context-growth.md](archived/ISS-012-executor-token-budget-and-context-growth.md) |
|
||||
| ISS-013 | Chat 入口解耦与真正 SSE 收敛 | 已归档 | [archived/ISS-013-chat-entry-decoupling-and-sse.md](archived/ISS-013-chat-entry-decoupling-and-sse.md) |
|
||||
| ISS-014 | 单体 ReAct Agent、Harness 与 ACI 工具瘦身 | 已完成并归档 | [archived/ISS-014-single-react-agent-harness-aci-ptk-refactor.md](archived/ISS-014-single-react-agent-harness-aci-ptk-refactor.md) |
|
||||
| executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 旧架构问题被 ISS-014 替代 | [archived/executor-evidence-attribution-hallucination.md](archived/executor-evidence-attribution-hallucination.md) |
|
||||
| rag-refactor-plan | RAG 检索重构计划 | 设计过时 | [archived/rag-refactor-plan.md](archived/rag-refactor-plan.md) |
|
||||
| diagnosis-eval-baseline-diff | 诊断评测 baseline diff 与回归判断 | 已归档 | [archived/diagnosis-eval-baseline-diff.md](archived/diagnosis-eval-baseline-diff.md) |
|
||||
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 已归档 | [archived/expand-diagnosis-eval-fixtures.md](archived/expand-diagnosis-eval-fixtures.md) |
|
||||
| mvp-demo-interview-runbook | Plan C 面试可复现 Demo 包 | 已归档 | [archived/mvp-demo-interview-runbook.md](archived/mvp-demo-interview-runbook.md) |
|
||||
|
||||
@@ -0,0 +1,103 @@
|
||||
# ISS-015 诊断运行质量与 Reasoning 审计收敛
|
||||
|
||||
**状态**:待实施
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-23
|
||||
**来源**:ISS-014 阶段 7 及后续真实 E2E 验证
|
||||
**关联**:ISS-014、ISS-004、executor-evidence-attribution-hallucination
|
||||
|
||||
---
|
||||
|
||||
## 1. 背景
|
||||
|
||||
ISS-014 已完成单体 Diagnosis ReAct Agent、Harness、ACI Tool、EvidenceGuard、SemanticGuard、SSE 和旧架构清理,并通过阶段 7 E2E 以及后续 Diagnosis/Knowledge Query 成功用例验证主链路。真实 E2E 同时暴露了新的运行质量问题:Agent 可能重复检索直至预算耗尽,Evidence Repair 可能因输出 Schema 不完整而解析失败,Reasoning 审计能力尚未完成真实 Provider 和数据库验证,安全 Fallback 在部分失败路径下仍缺少足够信息。
|
||||
|
||||
这些问题发生在 ISS-014 架构切换和验收之后,不重新打开 ISS-014,也不改变已经冻结的单体 Agent/Harness/ACI 架构。本 Issue 作为后续唯一追踪入口。
|
||||
|
||||
## 2. 用户体验问题
|
||||
|
||||
### 2.1 审计需要保留 Agent 思考记录
|
||||
|
||||
用户需要在后续审计中查看 Agent 的思考过程,确认它为什么选择某个 Tool、如何从观察结果继续行动以及在哪个阶段停止。该记录必须是独立的审计数据,不能混入业务事实、Evidence Snapshot、普通 Trace 或 SSE,也不能把未由 Provider 返回的内容伪造成思考记录。
|
||||
|
||||
### 2.2 失败结果不能只给空泛结论
|
||||
|
||||
当前失败路径可能直接返回“当前证据无法完成真实性校验,无法确认根因”。这虽然避免了无依据结论,但没有告诉用户已经观察到什么、校验在哪个阶段失败、缺少哪些信息以及下一步可以怎么查,导致用户无法判断系统是否真的执行过有效排查。
|
||||
|
||||
Fallback 必须在不暴露 Prompt、原始 Draft、原始 Tool 载荷和内部异常的前提下,提供有界的 `failure_stage`、`observed_facts`、`validation_issues`、查询范围、已完成的 Tool/证据阶段和可执行的 `next_steps`。
|
||||
|
||||
## 3. 已有验证基线
|
||||
|
||||
- Diagnosis 成功:`sessionId=mvp-demo-success-e2e-20260722-2032`,`runId=e72d9c7b-2f78-44c8-9086-f2e9d6987f24`,`release_outcome=SUCCESS`,总 Token 8455,Tool 调用 1 次。
|
||||
- Knowledge Query 成功:`sessionId=mvp-demo-order-timeout-fixed-20260723`,`runId=2b704d1b-8f11-40ee-a4c6-5e5e04d3847c`,`release_outcome=SUCCESS`,总 Token 1998。
|
||||
- 失败会话 `session_ud9pzde7r_1784785531138` 暴露两类问题:Evidence Repair `PARSE_ERROR`;Diagnosis Agent 重复调用 `lookup_knowledge`,最终在 12 次 Tool 调用、45087 Token 后进入 `BUDGET_EXHAUSTED`。
|
||||
- `diagnosis_trace_event`、信息化 Fallback、`agent_reasoning_audit` 和 reasoning 查询接口已经实现并通过 focused tests;真实 Provider/V015 验证仍属于本 Issue。
|
||||
|
||||
## 4. 目标
|
||||
|
||||
1. Agent 在证据轮次或单 Tool 预算接近上限时停止继续检索,并生成当前证据允许的最终 Draft。
|
||||
2. Evidence Repair 使用真实 `DiagnosisDraft` JSON Schema,消除结构修复阶段的 `PARSE_ERROR`。
|
||||
3. Reasoning 审计在真实 Provider 返回和不返回 reasoning 的两种情况下都有明确、可查询的审计结果。
|
||||
4. Reasoning 原文保持独立受限存储,不进入普通 SSE、Trace、证据快照或业务结果。
|
||||
5. Fallback 在不泄露 Prompt、Draft 和原始 Tool 数据的前提下,说明失败阶段、已观察事实和校验问题。
|
||||
6. 用户在失败时至少能知道系统执行到哪个阶段、看到了哪些有界事实、哪些校验未通过以及下一步应补充什么信息。
|
||||
|
||||
## 5. 实施阶段
|
||||
|
||||
### 阶段 1:Diagnosis Agent 硬停止策略
|
||||
|
||||
- 限制证据收集轮次和重复 `lookup_knowledge`。
|
||||
- 在预算耗尽前向 Agent 注入停止信号并要求输出最终 Draft。
|
||||
- 区分正常 ReAct 轮次、新 Tool Action 和真正的预算耗尽。
|
||||
- 增加重复检索、临界预算和无足够证据时的回归测试。
|
||||
|
||||
### 阶段 2:Evidence Repair Schema
|
||||
|
||||
- 使用框架 `BeanOutputConverter<DiagnosisDraft>` 或等价结构化转换器注入实际 JSON Schema。
|
||||
- 保持 Repair 无 Tool、最多一次、失败即安全 Fallback 的现有边界。
|
||||
- 覆盖合法修复、Schema 非法、解析失败和二次 EvidenceGuard 失败。
|
||||
|
||||
### 阶段 3:Reasoning 审计验证与治理
|
||||
|
||||
- 使用真实 Provider 验证 reasoning metadata 的键和返回行为。
|
||||
- Provider 不返回 reasoning 时写入 `reasoning_available=false`,不得伪造内容。
|
||||
- 验证 V015 数据库迁移和 `sessionId + runId` 精确 reasoning 查询。
|
||||
- 明确 reasoning 审计接口的访问控制、保留期限和加密要求。
|
||||
|
||||
### 阶段 4:Fallback 信息质量与最终 E2E
|
||||
|
||||
- 工具成功但无法构造 verified snapshot 时,返回有界 `observed_facts` 和 `validation_issues`。
|
||||
- 确保普通 Trace 只记录 `reasoning_available/reasoning_bytes`,不返回 reasoning 原文。
|
||||
- 运行至少一个 Diagnosis SUCCESS、一个信息化 FALLBACK 和一个 reasoning unavailable 的真实 E2E。
|
||||
- 按 exact `sessionId + runId` 核对 SSE、Trace、Reasoning Audit、ToolInvocation 和 Run 终态。
|
||||
|
||||
## 6. 验收标准
|
||||
|
||||
- 重复检索场景不会因第 9 次同 Tool 调用才被动触发 `BUDGET_EXHAUSTED`。
|
||||
- Agent 在有证据时可形成合法 Draft;证据不足时形成无强结论的合法 Draft 或信息化 Fallback。
|
||||
- Evidence Repair 的真实模型输出不再出现因缺失 `DiagnosisDraft` Schema 导致的 `PARSE_ERROR`。
|
||||
- V015 在真实数据库中迁移成功,reasoning 查询严格校验 path `sessionId` 与 query `runId` 的归属关系。
|
||||
- Provider 无 reasoning 时仍有审计记录;Provider 有 reasoning 时原文只存在于受限审计接口。
|
||||
- 普通 SSE、Trace、日志、Evidence Snapshot 和发布结果均不包含 reasoning 原文、Prompt 或原始 Tool 载荷。
|
||||
- 审计查询能按精确 `sessionId + runId` 返回每个 Agent 模型步骤的 reasoning 可用性、步骤顺序、字节数和受限原文;不存在跨 Session/Run 串读。
|
||||
- 失败 Fallback 不再只返回固定的“无法确认根因”文本;至少包含失败阶段、非敏感观察事实、具体校验问题、证据范围/缺口和下一步建议。
|
||||
- `observed_facts` 和 `validation_issues` 有长度和字段边界,不能泄露 Prompt、完整上下文、原始 Tool 响应、凭据或内部堆栈。
|
||||
- 在无证据、EvidenceGuard 失败、Evidence Repair 失败、SemanticGuard UNSUPPORTED 和预算耗尽等路径下,用户都能区分失败原因,而不是收到同一种空泛结论。
|
||||
- 三类最终 E2E 证据和精确数据库核验完成归档。
|
||||
|
||||
## 7. 流程门禁
|
||||
|
||||
- 每个阶段使用独立 OpenSpec change 和完整串行 sm-flow。
|
||||
- 每阶段完成 Apply、focused tests、验收证据、Archive 和独立 Git commit 后再进入下一阶段。
|
||||
- 阶段 1-3 不运行完整 live E2E;阶段 4 统一完成最终真实验证。
|
||||
- 不重新引入 Planner/Executor/Verifier/Composer、业务 Graph、第二 Chat 入口或通用工作流 DSL。
|
||||
|
||||
## 8. 相关文件
|
||||
|
||||
- `mvp/issues/archived/ISS-014-single-react-agent-harness-aci-ptk-refactor.md`
|
||||
- `src/main/java/com/superbiz/agent/harness/agent/DiagnosisAgentFactory.java`
|
||||
- `src/main/java/com/superbiz/agent/harness/release/EvidenceRepair.java`
|
||||
- `src/main/java/com/superbiz/agent/harness/release/SafeFallbackFactory.java`
|
||||
- `src/main/java/com/superbiz/agent/harness/audit/HarnessAgentAuditHook.java`
|
||||
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
|
||||
- `src/main/resources/db/migration/V015__create_agent_reasoning_audit.sql`
|
||||
+10
-1
@@ -1,9 +1,18 @@
|
||||
# ISS-003 MVP 设计与实现 Review 收敛
|
||||
|
||||
**状态**:待规划
|
||||
**状态**:已归档(Review 基线过时)
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-03
|
||||
**来源**:MVP 版本设计与实现 review
|
||||
**归档时间**:2026-07-23
|
||||
|
||||
---
|
||||
|
||||
## 归档说明
|
||||
|
||||
本 Issue 基于早期 Planner/Executor/Verifier、Redis Session 和旧 Tool Trace 架构进行总体 Review。其主要问题后来分别由凭据清理、Session/Run 隔离、证据链加固、文档路径修复和 ISS-014 单体 Diagnosis Agent/Harness 重构处理或替代;原文件中的类名、入口、优先级和建议方案已经不能准确描述当前系统。
|
||||
|
||||
仍可能存在的测试隔离、CORS 或 Redis 序列化风险不继续挂在本旧 Review 下。后续需要处理时,应基于当前代码和 ISS-014 架构重新建立范围明确的 Issue。本 Issue 以“Review 基线过时”归档,不作为当前缺陷清单或实施依据。
|
||||
|
||||
---
|
||||
|
||||
+10
-1
@@ -1,9 +1,18 @@
|
||||
# ISS-004 Executor 域级检索水位控制(Phase 2)
|
||||
|
||||
**状态**:待规划
|
||||
**状态**:已归档(范围被 ISS-014 吸收)
|
||||
**严重程度**:低
|
||||
**发现时间**:2026-07-01
|
||||
**关联**:ISS-002(Executor 无约束重复调用 lookup_knowledge)
|
||||
**归档时间**:2026-07-23
|
||||
|
||||
---
|
||||
|
||||
## 归档说明
|
||||
|
||||
本问题不再按旧 Executor 的 Session/Domain 水位矩阵独立实施。当前架构已经由 ISS-014 收敛为单体 Diagnosis ReAct Agent 与 Harness,重复 `lookup_knowledge`、证据轮次上限和预算耗尽前强制生成最终 Draft 统一由 ISS-014 的 Agent 硬停止策略处理。
|
||||
|
||||
原问题仍然有效,但实现入口、生命周期边界和验收方式已经改变;剩余工作统一追踪到 [ISS-015 Diagnosis Agent 硬停止策略](../active/ISS-015-diagnosis-runtime-quality-and-reasoning-audit.md#阶段-1diagnosis-agent-硬停止策略),本 Issue 以“范围被吸收”归档,不表示硬停止策略已经完成。
|
||||
|
||||
---
|
||||
|
||||
+2
-1
@@ -1,6 +1,7 @@
|
||||
# ISS-012 Executor Token 预算与上下文膨胀
|
||||
|
||||
**状态**:待规划
|
||||
**状态**:已被 ISS-014 吸收并归档(2026-07-22)
|
||||
**吸收结果**:单 Diagnosis Agent、Harness 集中预算、ACI Tool projection、exact Run audit 与安全 Fallback 已替代本 Issue 的旧 Executor 方案;最终 live 数据见 ISS-014 阶段 7 验收。
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-20
|
||||
**关联**:ISS-002、ISS-004、ISS-011
|
||||
+2
-1
@@ -1,6 +1,7 @@
|
||||
# ISS-013 Chat 入口解耦与真正 SSE 收敛
|
||||
|
||||
**状态**:待规划
|
||||
**状态**:已被 ISS-014 吸收并归档(2026-07-22)
|
||||
**吸收结果**:唯一 `/api/chat` named SSE、Chat Application Use Case、bounded executor、exact Run cancel 和前端单 consumer 已由 ISS-014 阶段 6A/6B 完成,最终 E2E 归入阶段 7。
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-20
|
||||
**关联**:ISS-011、ISS-012
|
||||
@@ -0,0 +1,55 @@
|
||||
# SuperBizAgent ISS-014 归档与 ISS-015 交接文档
|
||||
|
||||
## 当前上下文
|
||||
|
||||
- 分支:`refactor/chat-single-react-harness`
|
||||
- 任务:停止端到端验证,提交当前代码,更新 ISS-014,并交接后续开发。
|
||||
- 本轮没有继续启动服务或发送真实请求;9900 端口确认没有监听进程。
|
||||
- 工作区仍包含用户已有的文档、技能和演示输出改动;提交时只纳入 ISS-014 实现相关文件。
|
||||
|
||||
## 本次已完成
|
||||
|
||||
1. 新增独立追加式 `diagnosis_trace_event` 审计表和记录器,覆盖 Run、Routing、Agent、Tool、Evidence、Semantic、Release 生命周期事件。
|
||||
2. Knowledge Query 使用 `BeanOutputConverter<KnowledgeAnswerDraft>` 注入实际 JSON Schema,修复缺少 Schema 导致的 `Knowledge answer is invalid`。
|
||||
3. Fallback 增加 `failure_stage`、`observed_facts`、`validation_issues`,为失败提供可操作的阶段、事实和校验信息,同时不暴露 Prompt、原始 Draft 或 Tool 载荷。
|
||||
4. 新增 `agent_reasoning_audit` 表、Repository、Hook 持久化和专用查询接口。只保存模型供应商实际返回的 reasoning 元数据,长度上限 32000 字符;普通 Trace 只展示是否存在及字节数,不返回原文。
|
||||
5. ISS-014 已完成并归档;E2E 后新发现的问题已转移到 ISS-015。
|
||||
|
||||
## 已知验证结果
|
||||
|
||||
- Diagnosis 成功:`mvp-demo-success-e2e-20260722-2032` / `e72d9c7b-2f78-44c8-9086-f2e9d6987f24`,SUCCESS,8455 tokens,1 次 Tool 调用。
|
||||
- Knowledge Query 成功:`mvp-demo-order-timeout-fixed-20260723` / `2b704d1b-8f11-40ee-a4c6-5e5e04d3847c`,SUCCESS,1998 tokens。
|
||||
- 已通过的重点测试:`HarnessAgentAuditHookTest`、`DiagnosisReleaseUseCaseTest`、`DiagnosisTraceServiceTest`、`HarnessContractTest`、`HarnessChatConfigurationTest`、`ApplicationExecutorsTest`、`ChatApplicationUseCaseTest`、`JpaToolInvocationAuditSinkTest`、`JpaDiagnosisTraceRecorderTest`,以及编译和 `git diff --check`。
|
||||
- V015 真实数据库迁移、reasoning endpoint 的真实数据库/E2E 尚未验证。
|
||||
|
||||
## ISS-015 后续实现顺序
|
||||
|
||||
1. 为 Diagnosis Agent 增加硬停止策略:限制证据轮次、阻止重复 `lookup_knowledge`,在预算耗尽前强制输出最终 Draft。
|
||||
2. 给 Evidence Repair 注入真实 `DiagnosisDraft` Schema,消除运行时 `PARSE_ERROR`。
|
||||
3. 使用真实 Provider 验证 reasoning metadata;无 reasoning 时也要写入 `reasoning_available=false` 的审计记录。
|
||||
4. 在受控环境验证 V015 迁移和 reasoning 查询接口,并用精确 `sessionId + runId` 对齐数据。
|
||||
5. 评估 reasoning 审计表的访问控制、保留期限和加密策略。
|
||||
6. 继续细化工具成功但无法构造 verified snapshot 时的 fallback 事实和校验问题。
|
||||
|
||||
## 重要约束
|
||||
|
||||
- 不恢复完整端到端验证,除非用户明确要求。
|
||||
- 不把 reasoning 审计内容当作事实证据,也不加入普通 SSE、Trace 或业务结果。
|
||||
- 不提交无关用户文件、演示输出、环境配置和凭据。
|
||||
- ISS-015 的后续阶段仍按独立 sm-flow/OpenSpec、归档、提交后再进入下一阶段。
|
||||
|
||||
## 建议技能
|
||||
|
||||
- `diagnose`:继续处理 Agent 重复工具调用、预算耗尽和 Evidence Repair 解析失败。
|
||||
- `sm-flow`:开始下一阶段实现时按 OpenSpec-first 流程执行和归档。
|
||||
- `gitnexus-debugging`:定位 Agent 停止条件和跨 Harness 调用链问题。
|
||||
- `gitnexus-impact-analysis`:编辑共享 Harness/Agent 方法前评估调用影响。
|
||||
|
||||
## 关键文件
|
||||
|
||||
- 已归档 Issue:`mvp/issues/archived/ISS-014-single-react-agent-harness-aci-ptk-refactor.md`
|
||||
- 后续 Issue:`mvp/issues/active/ISS-015-diagnosis-runtime-quality-and-reasoning-audit.md`
|
||||
- Trace:`src/main/java/com/superbiz/agent/harness/audit/`、`src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
|
||||
- Reasoning:`src/main/java/com/superbiz/agent/domain/entity/AgentReasoningAudit.java`、`src/main/java/com/superbiz/agent/harness/audit/HarnessAgentAuditHook.java`、`src/main/resources/db/migration/V015__create_agent_reasoning_audit.sql`
|
||||
- Fallback:`src/main/java/com/superbiz/agent/harness/release/SafeFallbackFactory.java`
|
||||
- Knowledge:`src/main/java/com/superbiz/agent/harness/application/executor/KnowledgeQueryExecutor.java`
|
||||
+53
-31
@@ -1,11 +1,12 @@
|
||||
# ISS-014 单体 ReAct Agent、Harness 与 ACI 工具瘦身
|
||||
|
||||
**状态**:实施中(阶段 0-5 已归档,下一阶段 6A)
|
||||
**状态**:已完成并归档(阶段 0-7 已验收,后续问题转移至 ISS-015)
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-20
|
||||
**目标分支**:`refactor/chat-single-react-harness`
|
||||
**基线分支**:`refactor/mvp1.0`
|
||||
**关联**:ISS-003、ISS-012、ISS-013、executor-evidence-attribution-hallucination
|
||||
**归档时间**:2026-07-23
|
||||
|
||||
---
|
||||
|
||||
@@ -137,7 +138,7 @@ DIAGNOSIS
|
||||
- 调用其他 Agent;
|
||||
- 外层重试状态机。
|
||||
|
||||
不要求 Agent 输出 Thought 或 Chain of Thought。内部规划不形成对外协议,也不持久化为业务事实。
|
||||
不要求 Agent 输出 Thought 或 Chain of Thought。内部规划不形成对外协议,也不持久化为业务事实。若模型供应商实际返回 reasoning 元数据,可按独立审计数据保留受限副本;该副本不进入普通 Trace、证据快照或业务结果。
|
||||
|
||||
### 4.2 Harness
|
||||
|
||||
@@ -509,7 +510,7 @@ EvidenceGuard 或 SemanticGuard 降级使用 `content + done(FALLBACK)`;Router
|
||||
|
||||
禁止输出:
|
||||
|
||||
- Chain of Thought 和内部规划;
|
||||
- Chain of Thought 和内部规划(普通 SSE/Trace 不输出;独立审计存储仅限供应商实际返回的 reasoning 元数据);
|
||||
- System Prompt、模型 Prompt 和完整上下文;
|
||||
- 未脱敏 Tool 参数;
|
||||
- 原始日志、完整数据库行和完整检索 Trace;
|
||||
@@ -1046,9 +1047,9 @@ ISS-014 是总设计 Issue,不创建跨阶段共享的 OpenSpec change。以
|
||||
| 3C | `single-react-mysql-readonly-tool` | Completed;已归档 |
|
||||
| 4 | `single-react-diagnosis-agent` | Completed;已归档 |
|
||||
| 5 | `single-react-evidence-semantic-guards` | Completed;已归档 |
|
||||
| 6A | `single-react-chat-application-usecase` | Pending |
|
||||
| 6B | `single-react-chat-sse-cutover` | Pending |
|
||||
| 7 | `single-react-cleanup-e2e` | Pending |
|
||||
| 6A | `single-react-chat-application-usecase` | Completed;已归档 |
|
||||
| 6B | `single-react-chat-sse-cutover` | Completed;已归档 |
|
||||
| 7 | `single-react-cleanup-e2e` | Completed;已归档 |
|
||||
|
||||
每个 change 独立完成 Discover、Commit、Apply、阶段验收、Archive 和 Git commit;前一阶段 Archive 且提交后才允许启动下一阶段,不并行实施相邻阶段。
|
||||
|
||||
@@ -1307,27 +1308,27 @@ ISS-014 是总设计 Issue,不创建跨阶段共享的 OpenSpec change。以
|
||||
|
||||
## 14. 总体验收标准
|
||||
|
||||
- [ ] 复杂诊断只存在一个拥有工具循环的 Diagnosis ReAct Agent。
|
||||
- [ ] 不存在业务 StateGraph 或 Planner/Executor/Composer 多 Agent 主链路。
|
||||
- [ ] Harness 不承担业务推理,不演变为工作流引擎。
|
||||
- [ ] EvidenceGuard 是 Harness 内的确定性能力,不是独立编排节点。
|
||||
- [ ] SemanticGuard 使用完全隔离上下文,无工具、无记忆、无回调循环。
|
||||
- [ ] SemanticGuard 不使用未校准数值置信度控制在线释放。
|
||||
- [ ] 所有 Agent-facing evidence Tool 符合 ACI 状态和 Tool Call ID 契约。
|
||||
- [ ] RAG 不再向 Agent 返回 ContextPack/Trace/Rerank 等审计数据。
|
||||
- [ ] query_logs 不要求 Agent 先调用 Topic discovery,且返回聚合、抽样、脱敏结果。
|
||||
- [ ] MySQL Tool 只读、安全解析、参数绑定、allowlist、超时和结果上限全部生效。
|
||||
- [ ] Tool 原始结果不会未经有界投影进入 Agent 上下文;Redis canonical evidence 当前可暂不脱敏,但不得被 Agent 直接读取。
|
||||
- [ ] Redis 每次 Tool Call 单 Key 保存,状态、TTL、容量、ACL 和日志禁泄漏规则均有测试。
|
||||
- [ ] 每个 Tool Call ID 可以按 exact runId 和调用记录状态验真,并取得对应 `agent_result`。
|
||||
- [ ] 只保留一个 `/api/chat` SSE 接口。
|
||||
- [ ] SSE 不输出 Thought、Prompt、原始 Tool 载荷和未验证结论。
|
||||
- [ ] 客户端断开、模型/Tool/SemanticGuard 超时均有明确取消和 Run 终态。
|
||||
- [ ] 所有重试由 Harness 按类型化策略装配和记录,无 SDK/HTTP/数据库隐藏重试或整个 Diagnosis Agent 重跑。
|
||||
- [ ] 最终 E2E 能按 sessionId/runId 对齐 SSE、日志、AgentStep、ToolInvocation 和最终答案。
|
||||
- [ ] Token、工具调用、Tool 投影、Redis TTL 和总延迟预算均来自集中配置,并有可验证的强制上限和耗尽原因。
|
||||
- [ ] 11 个 OpenSpec changes 均已独立 Archive,并分别对应一个范围清晰的 Git commit。
|
||||
- [ ] 不保留旧兼容分支、注释代码、本地 refs 卸载和硬编码凭据。
|
||||
- [x] 复杂诊断只存在一个拥有工具循环的 Diagnosis ReAct Agent。
|
||||
- [x] 不存在业务 StateGraph 或 Planner/Executor/Composer 多 Agent 主链路。
|
||||
- [x] Harness 不承担业务推理,不演变为工作流引擎。
|
||||
- [x] EvidenceGuard 是 Harness 内的确定性能力,不是独立编排节点。
|
||||
- [x] SemanticGuard 使用完全隔离上下文,无工具、无记忆、无回调循环。
|
||||
- [x] SemanticGuard 不使用未校准数值置信度控制在线释放。
|
||||
- [x] 所有 Agent-facing evidence Tool 符合 ACI 状态和 Tool Call ID 契约。
|
||||
- [x] RAG 不再向 Agent 返回 ContextPack/Trace/Rerank 等审计数据。
|
||||
- [x] query_logs 不要求 Agent 先调用 Topic discovery,且返回聚合、抽样、脱敏结果。
|
||||
- [x] MySQL Tool 只读、安全解析、参数绑定、allowlist、超时和结果上限全部生效。
|
||||
- [x] Tool 原始结果不会未经有界投影进入 Agent 上下文;Redis canonical evidence 当前可暂不脱敏,但不得被 Agent 直接读取。
|
||||
- [x] Redis 每次 Tool Call 单 Key 保存,状态、TTL、容量、ACL 和日志禁泄漏规则均有测试。
|
||||
- [x] 每个 Tool Call ID 可以按 exact runId 和调用记录状态验真,并取得对应 `agent_result`。
|
||||
- [x] 只保留一个 `/api/chat` SSE 接口。
|
||||
- [x] SSE 不输出 Thought、Prompt、原始 Tool 载荷和未验证结论。
|
||||
- [x] 客户端断开、模型/Tool/SemanticGuard 超时均有明确取消和 Run 终态。
|
||||
- [x] 所有重试由 Harness 按类型化策略装配和记录,无 SDK/HTTP/数据库隐藏重试或整个 Diagnosis Agent 重跑。
|
||||
- [x] 最终 E2E 能按 sessionId/runId 对齐 SSE、日志、AgentStep、ToolInvocation 和最终答案。
|
||||
- [x] Token、工具调用、Tool 投影、Redis TTL 和总延迟预算均来自集中配置,并有可验证的强制上限和耗尽原因。
|
||||
- [x] 11 个 OpenSpec changes 均已独立 Archive,并分别对应一个范围清晰的 Git commit。
|
||||
- [x] 不保留旧兼容分支、注释代码、本地 refs 卸载和硬编码凭据。
|
||||
|
||||
## 15. 非目标
|
||||
|
||||
@@ -1335,7 +1336,7 @@ ISS-014 是总设计 Issue,不创建跨阶段共享的 OpenSpec change。以
|
||||
- 不引入业务 StateGraph;
|
||||
- 不实现 Harness 插件市场、通用 Pipeline DSL 或策略语言;
|
||||
- 不保留旧同步 Chat、`/chat_stream` 或旧 Tool Contract;
|
||||
- 不输出或持久化 Chain of Thought;
|
||||
- 不向普通 SSE、Trace、证据快照或业务结果输出 Chain of Thought;若供应商实际返回 reasoning 元数据,只允许写入独立的受限审计表,不将其作为业务事实或证据。
|
||||
- 不使用本地 `refs/*.md` 作为 Tool 原始结果主存储;
|
||||
- 不向 Agent 暴露 Redis Client、Redis Tool、key、连接信息或完整调用记录;
|
||||
- 不提供历史来源召回 Tool、通用 Memory Tool 或完整上下文恢复能力;
|
||||
@@ -1386,10 +1387,31 @@ ISS-014 是总设计 Issue,不创建跨阶段共享的 OpenSpec change。以
|
||||
|
||||
实现中不得为这些风险创建 HarnessPlugin、Graph、通用 Tool DSL 或大量预留接口;按各阶段最小范围实现并通过门禁验证。
|
||||
|
||||
## 17. 相关文件
|
||||
## 17. 阶段 7 验证与后续问题转移
|
||||
|
||||
- `mvp/issues/active/ISS-012-executor-token-budget-and-context-growth.md`
|
||||
- `mvp/issues/active/ISS-013-chat-entry-decoupling-and-sse.md`
|
||||
阶段 7 已完成真实 E2E、日志和数据库验收;后续 Diagnosis 与 Knowledge Query 成功用例进一步验证了新主链路。真实验证中新发现的运行质量问题不重新打开本 Issue,统一转移到 [ISS-015 诊断运行质量与 Reasoning 审计收敛](../active/ISS-015-diagnosis-runtime-quality-and-reasoning-audit.md)。
|
||||
|
||||
### 已完成并纳入本次提交
|
||||
|
||||
- 新增独立追加式 `diagnosis_trace_event` 表,按 `sessionId + runId` 记录 Run、Routing、Agent、Tool、Evidence、Semantic 和 Release 生命周期事件;普通 Trace 查询不再依赖其他业务表拼接完整时间线。
|
||||
- Knowledge Query 的 `KnowledgeAnswerDraft` 使用框架 `BeanOutputConverter` 注入实际 JSON Schema,避免模型因缺少结构约束返回 `Knowledge answer is invalid`。
|
||||
- Fallback 增加 `failure_stage`、`observed_facts` 和 `validation_issues`,在不暴露 Prompt、原始 Draft 或 Tool 载荷的前提下,为用户提供可审计的失败阶段、观察事实和校验问题。
|
||||
- 新增独立 `agent_reasoning_audit` 表和 reasoning 审计查询接口。Agent Hook 只保存供应商实际返回的 reasoning 元数据并限制长度;普通 Trace 仅记录是否存在及字节数,不返回原文。
|
||||
|
||||
### 已验证证据
|
||||
|
||||
- Diagnosis 成功用例:`sessionId=mvp-demo-success-e2e-20260722-2032`,`runId=e72d9c7b-2f78-44c8-9086-f2e9d6987f24`,`release_outcome=SUCCESS`,总 Token 8455,Tool 调用 1 次。
|
||||
- Knowledge Query 成功用例:`sessionId=mvp-demo-order-timeout-fixed-20260723`,`runId=2b704d1b-8f11-40ee-a4c6-5e5e04d3847c`,`release_outcome=SUCCESS`,总 Token 1998。
|
||||
- 相关 focused tests 和编译已通过。V015、reasoning endpoint、Agent 硬停止和 Evidence Repair 的后续真实验证属于 ISS-015,不再作为 ISS-014 归档门禁。
|
||||
|
||||
### 已转移范围
|
||||
|
||||
Diagnosis Agent 硬停止、Evidence Repair Schema、Reasoning 审计真实验证与治理、Fallback 信息质量已完整迁移到 ISS-015。ISS-014 的单体 Agent/Harness/ACI 架构、阶段实现和 E2E 验收至此关闭。
|
||||
|
||||
## 18. 相关文件
|
||||
|
||||
- `mvp/issues/archived/ISS-012-executor-token-budget-and-context-growth.md`
|
||||
- `mvp/issues/archived/ISS-013-chat-entry-decoupling-and-sse.md`
|
||||
- `mvp/disscus/plan.md`
|
||||
- `src/main/java/com/superbiz/agent/controller/ChatController.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
+10
-1
@@ -1,10 +1,19 @@
|
||||
# Executor 证据归因幻觉
|
||||
|
||||
**状态**:待规划
|
||||
**状态**:已归档(旧架构问题被 ISS-014 替代)
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-07
|
||||
**来源**:Trace Workbench 复核 / 最近 Chat 诊断会话低置信分析
|
||||
**关联**:ISS-005(证据链补齐与降级契约收敛)、ISS-006(固定诊断评测集与回归 Harness)、`chat-verifier-agent`、`diagnosis-playbook-skills`
|
||||
**归档时间**:2026-07-23
|
||||
|
||||
---
|
||||
|
||||
## 归档说明
|
||||
|
||||
本问题针对旧 Planner/Executor/Verifier/Composer 链路提出,相关 Executor 与 Verifier 主链路现已被 ISS-014 的单体 Diagnosis Agent、EvidenceGuard、SemanticGuard 和确定性 Release/Fallback 机制替代。旧架构下的输出分区、Verifier `LOW_CONFID` 和 Prompt 修补不再是当前实现入口。
|
||||
|
||||
证据引用真实性、语义支持性和信息化 Fallback 已由 ISS-014 统一承接;E2E 后新发现的 Agent 停止策略、Evidence Repair Schema 和审计验证继续在 [ISS-015 诊断运行质量与 Reasoning 审计收敛](../active/ISS-015-diagnosis-runtime-quality-and-reasoning-audit.md) 中追踪。本 Issue 以“旧架构问题被替代”归档,不表示所有诊断质量问题已经关闭。
|
||||
|
||||
---
|
||||
|
||||
@@ -1,9 +1,18 @@
|
||||
# RAG 检索重构计划
|
||||
|
||||
**状态**:待规划
|
||||
**状态**:已归档(设计过时)
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-05
|
||||
**范围**:RAG、检索、知识库、Agent Tool、AIOps 诊断证据链
|
||||
**归档时间**:2026-07-23
|
||||
|
||||
---
|
||||
|
||||
## 归档说明
|
||||
|
||||
本计划基于旧的 Milvus SDK、L0/L1 分层和 `VectorSearchService` 架构,包含 Spring AI VectorStore 旁路迁移、Query Transformer 和多阶段检索替换路线;当前实现已经收敛到 ISS-014 定义的单体 Diagnosis Agent、Harness 和显式 `lookup_knowledge` Tool Contract,原计划中的内部边界、验收口径和阶段拆分均已过时。
|
||||
|
||||
本文仅保留作为历史设计背景和 RAG 子问题来源,不再作为当前实现依据。后续若需要改造检索基础设施,应基于当前代码和 ISS-014 的 ACI/证据投影约束新建独立 Issue,不直接沿用本文的阶段计划。
|
||||
|
||||
---
|
||||
|
||||
@@ -0,0 +1,50 @@
|
||||
# Agent 推理审计表:agent_reasoning_audit
|
||||
|
||||
**状态**:当前独立敏感审计表;治理待 ISS-015 收敛
|
||||
**来源**:`V015__create_agent_reasoning_audit.sql`、`AgentReasoningAudit`
|
||||
|
||||
## 定位
|
||||
|
||||
`agent_reasoning_audit` 保存模型 Provider 在 Agent 步骤 metadata 中实际返回的 reasoning 内容。它与普通 Trace、AgentStep、Evidence Snapshot 和业务发布结果物理分离,不能作为事实证据或诊断结论来源。
|
||||
|
||||
Provider 未返回 reasoning 时仍写入 unavailable 记录,防止把“没有返回”误判为“审计链路漏写”。系统不得从最终回答、Tool 调用或其他字段生成伪 reasoning。
|
||||
|
||||
## 字段
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
|---|---|---|---|
|
||||
| `id` | BIGINT | 是 | 自增主键 |
|
||||
| `session_id` | VARCHAR(64) | 是 | 所属 Chat Session ID |
|
||||
| `run_id` | VARCHAR(64) | 是 | 所属 Diagnosis Run ID |
|
||||
| `step_index` | INT | 是 | Agent 模型步骤序号 |
|
||||
| `agent_name` | VARCHAR(64) | 是 | Agent 身份,当前为 `diagnosis_agent` |
|
||||
| `reasoning_available` | BOOLEAN | 是 | Provider 是否实际返回非空 reasoning |
|
||||
| `reasoning_content` | LONGTEXT | 否 | Provider reasoning 原文;Hook 当前最多保留 32000 个字符 |
|
||||
| `content_bytes` | INT | 是 | 截断后 reasoning 的 UTF-8 字节数;unavailable 时为 0 |
|
||||
| `created_at` | DATETIME | 是 | 创建时间,默认当前时间 |
|
||||
|
||||
## 索引
|
||||
|
||||
| 索引 | 字段 | 用途 |
|
||||
|---|---|---|
|
||||
| `idx_reasoning_run_step` | `run_id, step_index` | exact Run 下按模型步骤查询 |
|
||||
| `idx_reasoning_session_created` | `session_id, created_at` | Session 范围审计排查 |
|
||||
|
||||
## 关系
|
||||
|
||||
- `run_id` 逻辑关联 `diagnosis_run.run_id`。
|
||||
- `session_id` 逻辑关联 `chat_session.session_id`。
|
||||
- 通过 `session_id + run_id + step_index` 与 `agent_step` 逻辑对应,不建立数据库外键。
|
||||
|
||||
## 写入规则
|
||||
|
||||
- Provider 返回非空 reasoning:`reasoning_available=true`,保存截断后的原文及实际 UTF-8 字节数。
|
||||
- Provider 未返回 reasoning:`reasoning_available=false`,`reasoning_content=NULL`,`content_bytes=0`。
|
||||
- `agent_step.thought` 继续保持为空;普通 Trace 仅保留 availability/bytes metadata。
|
||||
- Reasoning 不进入 SSE、PreviousTurn、Evidence Snapshot、发布结果或应用日志。
|
||||
|
||||
## 查询与治理
|
||||
|
||||
当前独立接口为 `GET /api/diagnosis/{sessionId}/trace/reasoning?runId={runId}`。`runId` 必填,服务端校验其属于 path `sessionId`,并按 `step_index` 返回记录。
|
||||
|
||||
分表和归属校验已经实现,但不能等同于完整安全治理。身份认证、角色授权、保留/删除期限、静态与传输加密以及真实 Provider/V015 验证仍由 ISS-015 阶段 3 跟踪;治理完成前不应将该接口暴露给普通业务用户。
|
||||
@@ -5,7 +5,7 @@
|
||||
|
||||
## 定位
|
||||
|
||||
`agent_step` 记录一次诊断运行中每个 Agent 步骤的模型输入、输出、耗时和 Token 消耗。`run_id` 是执行隔离边界;Trace 页面展示顺序以 Trace API 返回顺序为准。
|
||||
`agent_step` 记录 Diagnosis Agent 模型步骤的有界审计 metadata。`run_id` 是执行隔离边界;当前写入不得保存 Prompt、消息正文、模型正文、Tool arguments 或 Thought。Provider reasoning 使用独立 `agent_reasoning_audit` 表,不复用历史 `thought` 字段。
|
||||
|
||||
## 字段
|
||||
|
||||
@@ -15,10 +15,10 @@
|
||||
| `session_id` | VARCHAR(64) | 是 | 所属会话目录 ID,保留用于粗粒度过滤和兼容 |
|
||||
| `run_id` | VARCHAR(64) | 否 | 所属 `diagnosis_run.run_id`;新执行应写入 |
|
||||
| `step_index` | INT | 是 | 步骤序号,从 0 开始 |
|
||||
| `agent_name` | VARCHAR(32) | 是 | Agent 名称,例如 planner、executor、verifier、composer |
|
||||
| `model_input` | TEXT | 否 | 模型输入摘要;`V006` 已从 JSON 改为 TEXT |
|
||||
| `model_output` | TEXT | 否 | 模型输出摘要;`V006` 已从 JSON 改为 TEXT |
|
||||
| `thought` | TEXT | 否 | Agent 思考过程或调试摘要 |
|
||||
| `agent_name` | VARCHAR(32) | 是 | 当前 Harness 写入固定为 `diagnosis_agent` |
|
||||
| `model_input` | TEXT | 否 | JSON metadata,仅包含 message count 与 roles |
|
||||
| `model_output` | TEXT | 否 | JSON metadata,仅包含 text presence 与 Tool names |
|
||||
| `thought` | TEXT | 否 | 当前 Harness 必须写空;字段仅保留历史兼容 |
|
||||
| `has_tool_call` | BOOLEAN | 否 | 本步骤是否触发工具调用 |
|
||||
| `duration_ms` | INT | 否 | 本步骤耗时 |
|
||||
| `token_count` | INT | 否 | 本步骤 Token 消耗 |
|
||||
@@ -37,9 +37,11 @@
|
||||
- `agent_step.run_id` 逻辑关联 `diagnosis_run.run_id`。
|
||||
- `agent_step.session_id` 保留为 `chat_session.session_id` 的冗余关联,便于粗粒度过滤和兼容查询。
|
||||
- `tool_invocation.step_id` 可关联 `agent_step.id`,但当前允许为空且不强制外键。
|
||||
- `agent_reasoning_audit` 通过相同的 `session_id + run_id + step_index` 逻辑定位模型步骤,不建立数据库外键。
|
||||
|
||||
## 注意点
|
||||
|
||||
- 前端展示步骤时应使用 Trace API 返回顺序;服务端会在同一 `run_id` 范围内整理步骤顺序。
|
||||
- 新 Trace、Verifier 和评测读路径应按 `run_id` 取数,避免同一 `sessionId` 多轮诊断混入。
|
||||
- Verifier 应在 Executor 循环完成后出现;如果 `step_index` 中 Verifier 提前,通常意味着编排或记录顺序有问题。
|
||||
- 新 Trace 和验收读路径必须按 exact `run_id` 取数,避免同一 `sessionId` 多次运行混入。
|
||||
- 当前 Run 若出现 `diagnosis_agent` 之外的新写入,或 `thought` 非空,视为审计边界违规。
|
||||
- `model_output` 可保存 `has_text`、`tool_names`、`reasoning_available` 和 `reasoning_bytes` 等有界 metadata,不得保存 reasoning 原文。
|
||||
|
||||
+10
-6
@@ -1,6 +1,6 @@
|
||||
# MVP 数据表索引
|
||||
|
||||
**更新日期**:2026-07-10
|
||||
**更新日期**:2026-07-23
|
||||
**状态**:当前表文档入口
|
||||
|
||||
本目录保存当前 MVP 使用的数据表说明。详细结构以 Flyway migration 和实体类为准;本目录用于面试讲解、排查索引和快速理解数据流。
|
||||
@@ -9,10 +9,12 @@
|
||||
|
||||
| 表 | 用途 | 文档 |
|
||||
|---|---|---|
|
||||
| `chat_session` | 会话目录元数据,保存同一个 `sessionId` 的多轮会话状态快照 | [聊天会话表-chat_session.md](聊天会话表-chat_session.md) |
|
||||
| `diagnosis_run` | 运行级主记录,保存一次 Chat/AIOps 诊断的 query、状态、答案、自评估和反馈 | [诊断运行表-diagnosis_run.md](诊断运行表-diagnosis_run.md) |
|
||||
| `agent_step` | Agent 步骤记录,按 `run_id` 隔离回放执行链路 | [Agent步骤表-agent_step.md](Agent步骤表-agent_step.md) |
|
||||
| `tool_invocation` | 工具调用记录,按 `run_id` 支撑 Trace、Verifier 和评测 | [工具调用表-tool_invocation.md](工具调用表-tool_invocation.md) |
|
||||
| `chat_session` | Chat 会话目录 metadata;不保存完整消息历史 | [聊天会话表-chat_session.md](聊天会话表-chat_session.md) |
|
||||
| `diagnosis_run` | `/api/chat` 运行主记录,保存 intent、终态与安全发布结果 | [诊断运行表-diagnosis_run.md](诊断运行表-diagnosis_run.md) |
|
||||
| `agent_step` | Diagnosis Agent metadata-only 模型步骤审计 | [Agent步骤表-agent_step.md](Agent步骤表-agent_step.md) |
|
||||
| `tool_invocation` | Harness ToolBoundary metadata-only 长期审计 | [工具调用表-tool_invocation.md](工具调用表-tool_invocation.md) |
|
||||
| `diagnosis_trace_event` | 按 Run 追加的统一诊断生命周期 Timeline | [诊断Trace事件表-diagnosis_trace_event.md](诊断Trace事件表-diagnosis_trace_event.md) |
|
||||
| `agent_reasoning_audit` | Provider reasoning 独立敏感审计;不属于普通 Trace | [Agent推理审计表-agent_reasoning_audit.md](Agent推理审计表-agent_reasoning_audit.md) |
|
||||
| `api_document` | 知识库文档元数据,和向量库 chunk 通过 `doc_id` 关联 | [文档元数据表-api_document.md](文档元数据表-api_document.md) |
|
||||
| `knowledge_domain` | 知识域元数据,支撑 RAG domain hint 和检索策略 | [知识域表-knowledge_domain.md](知识域表-knowledge_domain.md) |
|
||||
| `case_library` | 用户反馈沉淀出的高质量诊断案例 | [案例库表-case_library.md](案例库表-case_library.md) |
|
||||
@@ -31,6 +33,8 @@ chat_session.session_id
|
||||
-> diagnosis_run.session_id
|
||||
-> agent_step.run_id
|
||||
-> tool_invocation.run_id
|
||||
-> diagnosis_trace_event.run_id
|
||||
-> agent_reasoning_audit.run_id (restricted)
|
||||
-> case_library.diagnosis_id (new AUTO cases use run_id)
|
||||
|
||||
diagnosis_session.session_id
|
||||
@@ -43,4 +47,4 @@ knowledge_domain.domain_id
|
||||
-> api_document metadata.category / vector chunk metadata.category
|
||||
```
|
||||
|
||||
当前实现主要使用逻辑关联,不依赖数据库外键。
|
||||
当前实现主要使用 `session_id + run_id` 逻辑关联,不依赖数据库外键。`agent_reasoning_audit` 是敏感审计数据,不与普通 Trace、Evidence Snapshot 或业务结果合并读取。
|
||||
|
||||
@@ -5,7 +5,7 @@
|
||||
|
||||
## 定位
|
||||
|
||||
`tool_invocation` 记录 Agent 在一次诊断运行中显式调用工具的事实,包括工具名、入参、输出摘要、检索层级、证据引用和失败信息。它是 Trace、Verifier、评测和人工排查的共同数据源。
|
||||
`tool_invocation` 是 Harness ToolBoundary 的长期 metadata-only 审计表。它记录 exact Run/Tool identity、状态、稳定错误码、耗时与字节数;完整请求、raw response 和 Agent projection 只短期存在于 Redis canonical invocation,不写入本表。
|
||||
|
||||
## 字段
|
||||
|
||||
@@ -15,20 +15,20 @@
|
||||
| `session_id` | VARCHAR(64) | 是 | 所属会话目录 ID,保留用于粗粒度过滤和兼容 |
|
||||
| `run_id` | VARCHAR(64) | 否 | 所属 `diagnosis_run.run_id`;新执行应写入 |
|
||||
| `step_id` | BIGINT | 否 | 可关联 `agent_step.id` |
|
||||
| `tool_name` | VARCHAR(64) | 是 | 工具名称,例如 `lookup_knowledge`、日志查询、指标查询 |
|
||||
| `input_params` | JSON | 是 | 工具入参 |
|
||||
| `output_preview` | TEXT | 否 | 工具输出摘要或前缀 |
|
||||
| `output_length` | INT | 否 | 工具输出字符数 |
|
||||
| `retrieval_layer` | VARCHAR(8) | 否 | 检索层级,例如 `L0`、`L1`、`L0+L1` |
|
||||
| `tool_name` | VARCHAR(64) | 是 | ACI Tool 名:`lookup_knowledge`、`query_logs` 或 `query_mysql` |
|
||||
| `input_params` | JSON | 是 | 仅 `tool_call_id` 与 `request_bytes` metadata,不含 Tool 参数正文 |
|
||||
| `output_preview` | TEXT | 否 | 仅 invocation/evidence status metadata |
|
||||
| `output_length` | INT | 否 | Agent projection UTF-8 字节数 |
|
||||
| `retrieval_layer` | VARCHAR(8) | 否 | 当前 Harness 审计固定为 `HARNESS` |
|
||||
| `l0_match_count` | INT | 否 | L0 命中数量 |
|
||||
| `l1_match_count` | INT | 否 | L1 命中数量 |
|
||||
| `is_truncated` | BOOLEAN | 否 | 输出是否被截断 |
|
||||
| `relevance_level` | VARCHAR(20) | 否 | 归一化质量等级:`PRECISE`、`HIGHLY_RELEVANT`、`REFERENCE`、`DEDUPED` |
|
||||
| `dedup_reason` | VARCHAR(32) | 否 | 去重原因,例如 `doc_retrieved`、`domain_retrieved` |
|
||||
| `retrieval_details` | JSON | 否 | 检索明细、证据引用、Gatekeeper 可用导航信息 |
|
||||
| `relevance_level` | VARCHAR(20) | 否 | 当前 Harness 复用该字段保存 evidence status |
|
||||
| `dedup_reason` | VARCHAR(32) | 否 | 历史字段;当前 Harness 不写入 |
|
||||
| `retrieval_details` | JSON | 否 | `tool_call_id`、status、evidence status、result bytes 与可选稳定错误码 |
|
||||
| `duration_ms` | INT | 否 | 工具耗时 |
|
||||
| `success` | BOOLEAN | 否 | 工具是否成功 |
|
||||
| `error_message` | TEXT | 否 | 失败原因 |
|
||||
| `error_message` | TEXT | 否 | 仅稳定错误码,不保存内部异常或 vendor message |
|
||||
| `created_at` | DATETIME | 是 | 创建时间 |
|
||||
|
||||
## 索引
|
||||
@@ -36,7 +36,7 @@
|
||||
| 索引 | 字段 | 用途 |
|
||||
|---|---|---|
|
||||
| `idx_session_id` | `session_id` | 历史兼容和粗粒度排查 |
|
||||
| `idx_tool_invocation_run_id` | `run_id, id` | Trace、Verifier、评测按运行查询工具调用 |
|
||||
| `idx_tool_invocation_run_id` | `run_id, id` | Trace 与验收按 exact Run 查询 Tool 审计 |
|
||||
| `idx_tool_name` | `tool_name` | 按工具类型排查 |
|
||||
| `idx_retrieval_layer` | `retrieval_layer` | 观察 RAG L0/L1 行为 |
|
||||
|
||||
@@ -48,22 +48,19 @@
|
||||
|
||||
## 关键 JSON
|
||||
|
||||
`retrieval_details` 是扩展字段。当前重要结构包括:
|
||||
当前 Harness 写入的 `retrieval_details` 结构为:
|
||||
|
||||
```json
|
||||
{
|
||||
"evidence_status": "supported",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.logs[0]",
|
||||
"text": "工具返回中可核对的最小证据文本"
|
||||
}
|
||||
]
|
||||
"tool_call_id": "framework-call-id",
|
||||
"status": "READY",
|
||||
"evidence_status": "EVIDENCE_FOUND",
|
||||
"agent_result_bytes": 512
|
||||
}
|
||||
```
|
||||
|
||||
## 注意点
|
||||
|
||||
- Verifier 不应只信任 RAG 证据;所有工具只要能提供 `evidence_refs`,都应该进入可校验证据链。
|
||||
- `output_preview` 只适合展示和排查,不应被当成完整原始输出。
|
||||
- `$.no_evidence` 只代表“本次工具未命中证据”,不能推导为“故障不存在”。
|
||||
- 本表不是完整证据真理源,EvidenceGuard 只读取当前 Run 的 Redis canonical invocation。
|
||||
- `input_params`、`output_preview` 和 `retrieval_details` 均不得出现 SQL、日志 query、evidence body、凭据或 raw response。
|
||||
- audit 写入失败应记录安全 warning,但不能改变已确定的 canonical Tool 结果。
|
||||
|
||||
@@ -5,7 +5,7 @@
|
||||
|
||||
## 定位
|
||||
|
||||
`knowledge_domain` 保存知识库领域级元数据,用来帮助 Planner/Executor 判断什么时候检索某一类知识,并为 RAG 的 domain hint、去重和可观测性提供基础信息。
|
||||
`knowledge_domain` 保存知识库领域级元数据,为 RAG backend 的 domain hint、检索选择和可观测性提供基础信息;它不是 Agent-facing Tool contract。
|
||||
|
||||
## 字段
|
||||
|
||||
@@ -33,4 +33,4 @@
|
||||
## 注意点
|
||||
|
||||
- `when_to_retrieve` 是检索策略提示,不是事实证据。
|
||||
- Executor / Verifier 不能把领域描述当作诊断结论依据;事实仍应来自工具返回的证据块或证据引用。
|
||||
- Diagnosis Agent 与 SemanticGuard 不能把领域描述当作诊断结论依据;事实仍应来自当前 Run 经 EvidenceGuard 验真的 Tool evidence。
|
||||
|
||||
@@ -5,7 +5,7 @@
|
||||
|
||||
## 定位
|
||||
|
||||
`chat_session` 保存多轮 Chat 会话的元数据,用于把同一个 `sessionId` 下的多次诊断运行组织在一起。它不保存完整对话历史;正文消息仍由 Redis `SessionContext.messageHistory` 管理。
|
||||
`chat_session` 保存 Chat 会话目录元数据,用于把同一个 `sessionId` 下的多次运行组织在一起。它不保存完整对话历史;当前多轮只从最近一次安全发布的 `diagnosis_run.published_result` 构造有界 `PreviousTurn`,不再使用 Redis `SessionContext`。
|
||||
|
||||
## 字段
|
||||
|
||||
@@ -14,10 +14,10 @@
|
||||
| `id` | BIGINT | 是 | 自增主键 |
|
||||
| `session_id` | VARCHAR(64) | 是 | 会话目录 ID,外部 API 仍通过它定位会话 |
|
||||
| `status` | VARCHAR(16) | 否 | `ACTIVE`、`EXPIRED`、`CLOSED` |
|
||||
| `message_pair_count` | INT | 否 | Redis 会话中问答轮次数的快照 |
|
||||
| `message_pair_count` | INT | 否 | 历史兼容计数;当前运行不依赖它恢复消息正文 |
|
||||
| `created_at` | DATETIME | 是 | 创建时间 |
|
||||
| `last_active_at` | DATETIME | 否 | 最近活跃时间 |
|
||||
| `expires_at` | DATETIME | 否 | 目录元数据,可为空;Redis 消息历史可独立过期 |
|
||||
| `expires_at` | DATETIME | 否 | 会话目录过期元数据,可为空 |
|
||||
|
||||
## 索引
|
||||
|
||||
@@ -38,3 +38,4 @@
|
||||
|
||||
- `chat_session` 是会话元数据,不是诊断执行记录。
|
||||
- 不要把 query、answer、self_evaluation、feedback 写入该表;这些属于 `diagnosis_run`。
|
||||
- 不要从该表或 Redis 恢复完整对话正文;安全追问上下文只来自成功发布的结构化结果。
|
||||
|
||||
@@ -0,0 +1,52 @@
|
||||
# 诊断 Trace 事件表:diagnosis_trace_event
|
||||
|
||||
**状态**:当前统一 Trace Timeline
|
||||
**来源**:`V014__create_diagnosis_trace_event.sql`、`DiagnosisTraceEvent`
|
||||
|
||||
## 定位
|
||||
|
||||
`diagnosis_trace_event` 是按 Run 追加的统一生命周期事件表。它记录诊断执行经过哪些阶段、每个阶段的稳定事件类型和结果,使普通 Trace 不再依赖从多个业务表推测完整时序。
|
||||
|
||||
该表只保存有界安全 metadata,不保存 Prompt、Thought、DiagnosisDraft 正文或 raw Tool payload。事件写入失败可观测,但不改变业务执行结果。
|
||||
|
||||
## 字段
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
|---|---|---|---|
|
||||
| `id` | BIGINT | 是 | 自增主键;同一 `sequence_no` 下作为稳定次序补充 |
|
||||
| `session_id` | VARCHAR(64) | 是 | 所属 Chat Session ID |
|
||||
| `run_id` | VARCHAR(64) | 是 | 所属 Diagnosis Run ID |
|
||||
| `sequence_no` | INT | 是 | 同一 Run 内递增事件序号 |
|
||||
| `phase` | VARCHAR(32) | 是 | `RUN`、`ROUTING`、`AGENT`、`TOOL`、`EVIDENCE`、`SEMANTIC` 或 `RELEASE` |
|
||||
| `event_type` | VARCHAR(48) | 是 | 稳定事件类型,如 `RUN_STARTED`、`AGENT_MODEL_STEP`、`RELEASE_DECISION` |
|
||||
| `status` | VARCHAR(32) | 是 | 稳定事件结果,如 `SUCCEEDED`、`FAILED`、`PASSED`、`REJECTED`、`FALLBACK` |
|
||||
| `attempt_no` | INT | 否 | Retry 或模型尝试序号;不适用时为空 |
|
||||
| `duration_ms` | INT | 否 | 事件耗时;无法计算时为空 |
|
||||
| `details` | JSON | 是 | 有界安全 metadata;不得包含敏感正文和原始载荷 |
|
||||
| `created_at` | DATETIME | 是 | 创建时间,默认当前时间 |
|
||||
|
||||
## 索引
|
||||
|
||||
| 索引 | 字段 | 用途 |
|
||||
|---|---|---|
|
||||
| `idx_trace_event_run_sequence` | `run_id, sequence_no, id` | exact Run Timeline 顺序查询 |
|
||||
| `idx_trace_event_session_created` | `session_id, created_at, id` | Session 范围历史排查 |
|
||||
| `idx_trace_event_phase` | `phase, event_type` | 按生命周期阶段和事件类型统计 |
|
||||
|
||||
## 关系
|
||||
|
||||
- `run_id` 逻辑关联 `diagnosis_run.run_id`。
|
||||
- `session_id` 逻辑关联 `chat_session.session_id`,同时用于 Run 所有权排查。
|
||||
- 不建立数据库外键;应用必须使用 exact `sessionId + runId` 做归属校验。
|
||||
|
||||
## 事件范围
|
||||
|
||||
当前事件类型覆盖 Run 开始/结束、路由尝试/决策、Agent 模型步骤、Tool 调用、Evidence 初检/Repair/复检、Semantic 尝试/决策和 Release 决策。
|
||||
|
||||
普通 Trace API 按 `sequence_no ASC, id ASC` 返回 Timeline。Agent 模型事件最多包含 reasoning availability 和字节数,不包含 reasoning 原文。
|
||||
|
||||
## 注意点
|
||||
|
||||
- 该表是追加式审计时间线,不是 `diagnosis_run` 终态的替代品。
|
||||
- `details` 必须保持结构化、有界和非敏感;禁止将其他表的正文复制进来。
|
||||
- 不应仅依赖 `created_at` 排序,同一 Run 必须使用 `sequence_no, id`。
|
||||
@@ -5,7 +5,7 @@
|
||||
|
||||
## 定位
|
||||
|
||||
`diagnosis_session` 是旧版 session 级诊断主记录。`V011` 之后,新 Chat/AIOps 执行的运行态写入已经切到 `chat_session + diagnosis_run`;本表保留用于历史兼容、迁移回填和回滚比较。
|
||||
`diagnosis_session` 是旧版 session 级诊断主记录。`V011` 之后,当前 `/api/chat` 执行的运行态写入已经切到 `chat_session + diagnosis_run`;本表保留用于历史兼容、迁移回填和回滚比较。
|
||||
|
||||
## 字段
|
||||
|
||||
|
||||
@@ -5,7 +5,7 @@
|
||||
|
||||
## 定位
|
||||
|
||||
`diagnosis_run` 表示一次可回放的 Chat 或 AIOps 诊断执行。`run_id` 是运行级边界,Trace、反馈、自评估、案例沉淀和统计都应优先按 `run_id` 绑定。
|
||||
`diagnosis_run` 表示一次 `/api/chat` 应用执行。`run_id` 是运行级边界,SSE、Trace、AgentStep 与 ToolInvocation 必须按 metadata 返回的 exact `run_id` 绑定。
|
||||
|
||||
## 字段
|
||||
|
||||
@@ -14,11 +14,14 @@
|
||||
| `id` | BIGINT | 是 | 自增主键 |
|
||||
| `run_id` | VARCHAR(64) | 是 | 运行唯一 ID,格式为 `run-` + UUID |
|
||||
| `session_id` | VARCHAR(64) | 是 | 所属 `chat_session.session_id` |
|
||||
| `query` | TEXT | 是 | 本次 Chat 问题或 AIOps 告警摘要 |
|
||||
| `status` | VARCHAR(16) | 否 | `PENDING`、`RUNNING`、`SUCCESS`、`FAILED` |
|
||||
| `agent_flow` | VARCHAR(32) | 否 | `CHAT` 或 `AI_OPS` |
|
||||
| `answer` | LONGTEXT | 否 | 本次运行的最终答复或告警报告 |
|
||||
| `self_evaluation` | JSON | 否 | 本次运行的 rule、verifier、aiops 自评估容器 |
|
||||
| `query` | TEXT | 是 | 本次 Chat 用户问题 |
|
||||
| `status` | VARCHAR(16) | 否 | `PENDING`、`RUNNING`、`SUCCESS`、`FALLBACK`、`FAILED` 或 `CANCELLED` |
|
||||
| `agent_flow` | VARCHAR(32) | 否 | 历史兼容字段;当前公开执行统一来自 Chat Harness |
|
||||
| `answer` | LONGTEXT | 否 | 安全发布的最终文本兼容字段 |
|
||||
| `intent` | VARCHAR(32) | 否 | `SYSTEM_CHAT`、`KNOWLEDGE_QUERY` 或 `DIAGNOSIS` |
|
||||
| `release_outcome` | VARCHAR(16) | 否 | `SUCCESS`、`FALLBACK`、`FAILED` 或 `CANCELLED` |
|
||||
| `published_result` | JSON | 否 | Release Policy 允许发布的结构化安全结果 |
|
||||
| `self_evaluation` | JSON | 否 | 历史兼容字段;当前 Harness 不写入旧 verifier/AiOps 结构 |
|
||||
| `feedback` | VARCHAR(16) | 否 | 本次运行的用户反馈 |
|
||||
| `total_duration_ms` | INT | 否 | 本次运行总耗时 |
|
||||
| `total_token_count` | INT | 否 | 本次运行 Token 消耗 |
|
||||
@@ -35,13 +38,15 @@
|
||||
| `idx_diagnosis_run_session_created` | `session_id, created_at, id` | session 下最新运行解析和运行列表 |
|
||||
| `idx_diagnosis_run_session_run` | `session_id, run_id` | exact trace / feedback ownership 校验 |
|
||||
| `idx_diagnosis_run_status` | `status` | 状态筛选 |
|
||||
| `idx_diagnosis_run_agent_flow` | `agent_flow` | 区分 Chat / AIOps |
|
||||
| `idx_diagnosis_run_agent_flow` | `agent_flow` | 历史兼容筛选 |
|
||||
|
||||
## 关系
|
||||
|
||||
- `diagnosis_run.session_id` 逻辑关联 `chat_session.session_id`。
|
||||
- `agent_step.run_id` 逻辑关联 `diagnosis_run.run_id`。
|
||||
- `tool_invocation.run_id` 逻辑关联 `diagnosis_run.run_id`。
|
||||
- `diagnosis_trace_event.run_id` 逻辑关联 `diagnosis_run.run_id`,形成统一生命周期 Timeline。
|
||||
- `agent_reasoning_audit.run_id` 逻辑关联 `diagnosis_run.run_id`,但只能通过独立敏感审计读路径查询。
|
||||
- 新的自动案例沉淀使用 `case_library.diagnosis_id = diagnosis_run.run_id`。
|
||||
|
||||
## 注意点
|
||||
@@ -49,3 +54,5 @@
|
||||
- `GET /api/diagnosis/{sessionId}/trace` 未带 `runId` 时只为兼容解析 latest run;新 demo 和新客户端应传 `runId`。
|
||||
- latest run 排序使用 `created_at DESC, id DESC`,避免 feedback 或自评估更新 `updated_at` 后改变回放目标。
|
||||
- 历史 `diagnosis_session` 会被迁移成兼容 run,但旧混合数据不能被还原成真实多轮边界。
|
||||
- 当前诊断发布结果以 `release_outcome + published_result` 为准,不得从旧 self-evaluation 推断 Release Policy 结果。
|
||||
- `diagnosis_run` 不保存 reasoning 原文;普通 Run/Trace 查询也不得通过聚合将其带出。
|
||||
|
||||
+1
@@ -0,0 +1 @@
|
||||
ready
|
||||
@@ -0,0 +1 @@
|
||||
committed
|
||||
+2
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-21
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user