Compare commits

..
Author SHA1 Message Date
zhuyongxin 83193bdf4a docs(harness): add overall architecture learning note
- Add architecture note: config assembly (three-layer organization),
  HTTP entry (thin controller + SSE state machine + disconnect cancel),
  session & memory system (PreviousTurn injection, terminology calibration,
  skill as procedural long-term memory), knowledge base write pipeline
2026-08-10 18:35:23 +08:00
zhuyongxin de56551fea docs(harness): add LLM judge design note (interview Q&A) 2026-08-10 11:34:37 +08:00
zhuyongxin da18fdf4e1 docs(harness): add agent domain learning note and mark all nine domains complete
- Add agent note: factory assembly, dual interceptors (model budget/token audit,
  tool five gates), use case loop shell, controlled stop recovery, dual-view projection
- Mark agent domain complete in roadmap (all nine domains done)
2026-08-10 10:12:04 +08:00
aruo e1b8d1fb2c docs(harness): add evidence-chain, application-audit, contract-state and interview review notes; annotate guard/release core classes 2026-08-08 00:12:25 +08:00
zhuyongxin 074d1aa5a9 docs(harness): add MySQL sandbox learning note and mark tool domain complete
- Add MySQL sandbox note: three defense layers (semantic/connection/output),
  fail-closed validation, allowlist, parameterization, cancellation, redaction
- Mark tool domain complete in roadmap; next: guard
2026-08-06 17:46:53 +08:00
zhuyongxin f26d395650 docs(harness): annotate RAG backend core classes and add retrieval learning note
- Annotate LookupKnowledgeTool, KnowledgeEvidencePostProcessor, RrfFusion, KnowledgeDocumentRetriever
- Add RAG retrieval learning note: L0 navigation, multi-recall + RRF, qualityScore,
  degradation, contract semantics, validation (audit + offline eval), discussion insights
2026-08-05 18:37:50 +08:00
zhuyongxin ff0752a16c docs(harness): annotate tool domain classes and add tool chain learning notes
- Annotate 43 tool domain classes (contract/projection/boundary/store/adapter/mysql)
- Add tool registration and execution chain learning note
- Add tool call chain runtime journey note (model decision to observation)
2026-08-04 18:37:17 +08:00
zhuyongxin 7844bcea40 docs(harness): annotate progress core classes and add code-level learning notes
- Annotate DiagnosisProgressTracker, HarnessToolInterceptor, DiagnosisProgressProjector,
  DiagnosisReleaseUseCase, ToolBoundary, ToolBoundaryResult, CanonicalToolInvocation
- Add progress code learning note: interceptor gates, tracker state machine,
  canonical lifecycle, execution gate, projection/release pipeline
- Update learning roadmap: progress marked as deeply learned, next is tool domain
2026-08-03 18:37:23 +08:00
wdm1802 e564863c43 docs(harness): add retry guide and learning roadmap; annotate retry core 2026-08-03 01:29:36 +08:00
wdm1802 d084202166 docs(harness): add budget flow sequence and execution control notes 2026-08-02 20:58:13 +08:00
zhuyongxin 5b2fb985d9 docs(harness): annotate entry orchestration, enums, and exception recovery paths 2026-07-30 19:08:03 +08:00
zhuyongxin b39a625e5b docs(mvp): add remaining harness audit guides and engineering index 2026-07-30 19:04:59 +08:00
zhuyongxin e9f1c48d34 docs(harness): add loop inner/outer exception handling guide 2026-07-30 18:55:43 +08:00
zhuyongxin 3a7eee8af4 docs(mvp): add harness design and progressive guides 2026-07-29 19:04:44 +08:00
zhuyongxin 584639fa2a docs(mvp): move engineering notes under mvp/engineering
Relocate RAG and diagnosis decision/E2E writeups from docs/ root into
mvp/engineering so architecture, issues, and engineering narrative stay
together. Update indexes and cross-links; leave docs/learning as legacy.
2026-07-29 10:49:45 +08:00
zhuyongxin bdac35567c docs(issues): add ISS-017 L0 filter fallback hardening
Track L0 hard-filter + unfiltered retry brittleness. Keep current
behavior; prefer A+C later and defer soft-constraint redesign (D).
2026-07-29 10:48:47 +08:00
zhuyongxin 7ae9707a3b feat(harness,rag): dual LLM audit fields, run conclusion, and hybrid quality
Persist provider reasoning and assistant text separately on agent_reasoning_audit
(DeepSeekAssistantMessage path), extract diagnosis_run.conclusion, enrich RAG
tool audit (step_id/query/qualityScore), gate empty mysql tools, drop devtools,
and align MVP docs after live E2E verification.
2026-07-28 19:43:13 +08:00
zhuyongxin 2f40536248 docs(mvp): document dense+BM25 hybrid knowledge retrieval
Add current RAG architecture covering MilvusClientV2 hybrid search,
chunk evidence identity, rebuild ops, and update the MVP architecture
index and system overview links.
2026-07-28 09:56:08 +08:00
zhuyongxin 729cd3544a chore(rag): remove obsolete PowerShell rebuild script 2026-07-28 09:50:15 +08:00
zhuyongxin 1e532fb851 chore(rag): rebuild script in Python and use biz collection
Switch default knowledge collection back to biz (drop+recreate on rebuild),
replace PowerShell rebuild runner with Python, skip README.md imports, and
merge duplicate rag keys in application.yml.
2026-07-28 09:49:43 +08:00
zhuyongxin 38f781b157 feat(harness): complete protocol repair stop and archive ISS-016
Add repairable INVALID_PROGRESS_PROTOCOL observations, independent
PROGRESS_PROTOCOL_VIOLATED saturation, and controlled release paths.
Archive the OpenSpec change after syncing main specs and devflow.
2026-07-27 19:10:07 +08:00
zhuyongxin 5c369f3b6c feat(rag): add hybrid knowledge rebuild API and script
Add confirm-gated rebuild-hybrid endpoint that drops biz_hybrid, clears
api_document and L0, then force-imports knowledge_base markdown into the
dense+BM25 store. Include PowerShell runner and ops README.
2026-07-27 18:55:00 +08:00
zhuyongxin f035538531 feat(rag): dense+BM25 hybrid on MilvusClientV2, drop SDK path
Replace legacy MilvusServiceClient knowledge search/write with a single
MilvusClientV2 hybrid store (BM25 function + dense ANN + RRFRanker).
Use collection biz_hybrid and require knowledge reindex.
2026-07-27 18:49:20 +08:00
zhuyongxin 376ad0c241 feat(rag): hybrid multi-path search with RRF fusion
Add configurable hybrid mode on KnowledgeSearchPort that fuses dense
unfiltered, dense filtered, and lexical ranks via RRF while preserving
dense-compatible threshold scores. Archives Delivery 2 OpenSpec change.
2026-07-27 18:33:31 +08:00
zhuyongxin ac1f831903 feat(rag): chunk evidence identity, dedup, and search port
Preserve same-document multi-chunk evidence with evidenceKey identity,
per-document caps, retrieve-k/return-n split, and a dense KnowledgeSearchPort.
Archives Delivery 1 OpenSpec change as the foundation for hybrid retrieval.
2026-07-27 18:26:15 +08:00
zhuyongxin 99d4f6f216 docs(rag): add comments on knowledge retrieval pipeline
Document the lookup_knowledge flow from L0 hints through L1 retrieval,
post-processing, packing, and Agent projection so the boundaries and
current limitations are easier to follow.
2026-07-27 16:26:06 +08:00
aruo edb6153fd6 docs(issue): mark iss-015 partially implemented 2026-07-27 10:22:09 +08:00
aruo d0452184ee feat(harness): add information gain stop and audit 2026-07-27 01:03:34 +08:00
zhuyongxin de5a5b09d9 docs(interview): refresh materials for single-agent harness narrative
Archive pre-refactor interview notes and add current deep-dives on
architecture evolution, issue-derived stories, and evidence gates.
2026-07-24 18:14:49 +08:00
zhuyongxin e47f2dead0 docs(mvp): align architecture and audit table design 2026-07-23 20:16:29 +08:00
zhuyongxin 49180abccf docs(issue): archive legacy issues and add iss-015 2026-07-23 19:54:01 +08:00
zhuyongxin 529f4ff43b docs(demo): add successful diagnosis request 2026-07-23 19:06:37 +08:00
zhuyongxin 75fa154a0a chore(agent): add shared skills and GitNexus guidance 2026-07-23 19:05:22 +08:00
zhuyongxin 8300435a63 docs(issue): add iss-014 handoff 2026-07-23 18:03:14 +08:00
zhuyongxin e20249c5d9 feat(harness): improve trace fallback and reasoning audit 2026-07-23 17:52:01 +08:00
zhuyongxin 8fbc443f76 docs(issue): close iss-014 stage status 2026-07-22 18:08:29 +08:00
zhuyongxin 8ee7cc0b70 refactor(harness): remove legacy agent architecture 2026-07-22 18:02:01 +08:00
zhuyongxin bc36248cd8 feat(chat): cut over to single SSE endpoint 2026-07-22 10:01:12 +08:00
zhuyongxin f8809cb7dd feat(harness): add chat application use case 2026-07-22 00:57:39 +08:00
zhuyongxin ee0949d464 feat(harness): add evidence and semantic guards 2026-07-21 23:50:51 +08:00
zhuyongxin 2362665519 feat(harness): add single diagnosis react agent 2026-07-21 22:29:49 +08:00
zhuyongxin 85029d96a7 feat(harness): add readonly mysql tool 2026-07-21 21:24:42 +08:00
zhuyongxin 3e602781d6 feat(harness): add rag and log projections 2026-07-21 20:12:18 +08:00
zhuyongxin 0dbdd7d8d3 feat(harness): add canonical tool invocation boundary 2026-07-21 19:36:25 +08:00
zhuyongxin 6b74990f86 feat(harness): add run context and retry core 2026-07-21 18:36:19 +08:00
zhuyongxin 4274f3350b feat(harness): freeze aci tool contracts 2026-07-21 18:01:18 +08:00
zhuyongxin 58c39107c5 refactor(harness): freeze single-agent contracts 2026-07-21 17:33:25 +08:00
zhuyongxin 30d3296043 docs(mvp): archive session run trace issue 2026-07-11 16:46:20 +08:00
zhuyongxin 3578709896 docs(openspec): archive session run isolation 2026-07-10 22:52:39 +08:00
zhuyongxin f9df94377b feat(trace): finish run-aware demo verification 2026-07-10 21:37:43 +08:00
zhuyongxin 78c1477198 feat(trace): isolate aiops runs 2026-07-10 20:57:46 +08:00
zhuyongxin d928a1968a feat(trace): bind feedback to runs 2026-07-10 20:29:56 +08:00
zhuyongxin 027aed1eeb feat(trace): add run-scoped trace reads 2026-07-10 20:07:23 +08:00
zhuyongxin 26d5529280 feat(trace): isolate chat runs 2026-07-10 19:02:04 +08:00
zhuyongxin 6fdbd34bab docs(openspec): tighten run isolation contract 2026-07-10 17:56:51 +08:00
zhuyongxin 52bf0302c6 feat(trace): add session run isolation schema 2026-07-10 17:47:56 +08:00
zhuyongxin 841437fa06 docs(mvp): organize mvp documentation 2026-07-09 13:40:08 +08:00
zhuyongxin 9c9a0024d4 feat(demo): add interview quality audit 2026-07-09 11:18:49 +08:00
zhuyongxin a6c2d4459c docs(openspec): propose interview demo quality audit 2026-07-09 10:34:33 +08:00
zhuyongxin da45fa3fb0 docs(devflow): sort index by date 2026-07-09 10:24:35 +08:00
aruo db0f229285 feat(eval): add evidence pipeline acceptance closure 2026-07-09 00:47:48 +08:00
aruo a77c947cd4 docs(architecture): align evidence pipeline design 2026-07-08 23:53:21 +08:00
aruo 9a84b3de34 feat(agent): support no-evidence references 2026-07-08 23:30:00 +08:00
aruo 7b8c75e571 feat(agent): harden verifier evidence references 2026-07-08 16:12:56 +08:00
aruo a08672b31e feat(eval): add executor audit closure checks 2026-07-08 10:21:39 +08:00
aruo 6015bcbf6f feat(agent): add composer final answer 2026-07-08 09:51:07 +08:00
aruo a5b4502c72 docs(openspec): propose executor composer final answer 2026-07-08 02:49:46 +08:00
aruo 39c0c5f8be docs(mvp): clarify executor v2 implementation issue 2026-07-08 02:43:12 +08:00
aruo 1b31e78be5 feat(agent): add verifier claim checks 2026-07-08 02:33:02 +08:00
aruo c5e496e715 feat(agent): add executor gatekeeper hook 2026-07-08 02:01:49 +08:00
aruo 050cbc8fee feat(agent): add executor evidence v2 contract 2026-07-08 01:37:15 +08:00
zhuyongxin a6afbfaa9d chore: add editorconfig 2026-07-07 21:08:51 +08:00
zhuyongxin 0ee27eb523 feat(trace): improve session workbench review 2026-07-07 19:06:18 +08:00
zhuyongxin 3b62a8940c chore(openspec): archive executor evidence output contract 2026-07-07 19:04:24 +08:00
zhuyongxin 04eb50e2b4 feat(agent): add executor evidence output contract 2026-07-07 19:02:02 +08:00
zhuyongxin 7f2e47ca38 docs(mvp): record executor evidence loop design 2026-07-07 18:51:47 +08:00
aruo aa035b828c feat(trace): add diagnosis trace workbench 2026-07-07 01:24:56 +08:00
aruo b3315ead52 fix(agent): harden live diagnosis skill observability 2026-07-07 00:20:34 +08:00
zhuyongxin 64adb998cf chore(rag): add eval knowledge base mirror 2026-07-06 21:48:21 +08:00
zhuyongxin ed7efc58b7 feat(rag): close eval pipeline with live snapshots 2026-07-06 21:39:27 +08:00
zhuyongxin cf3333d607 feat(rag): modularize knowledge retrieval pipeline 2026-07-06 17:06:05 +08:00
zhuyongxin a375daead7 fix: clean up aiops mojibake text 2026-07-06 11:38:01 +08:00
zhuyongxin a5a0e0c6be Merge remote-tracking branch 'origin/refactor/mvp1.0' into refactor/mvp1.0 2026-07-06 10:54:45 +08:00
zhuyongxin 9e8e20b3b5 chore: update agent instructions 2026-07-06 10:47:54 +08:00
aruo 3dfe3dbe53 Update devflow glossary for skills 2026-07-06 10:43:21 +08:00
zhuyongxin 37083fc92a chore: add agent skills 2026-07-06 10:18:10 +08:00
1041 changed files with 85803 additions and 12875 deletions
+117
View File
@@ -0,0 +1,117 @@
---
name: diagnose
description: Disciplined diagnosis loop for hard bugs and performance regressions. Reproduce → minimise → hypothesise → instrument → fix → regression-test. Use when user says "diagnose this" / "debug this", reports a bug, says something is broken/throwing/failing, or describes a performance regression.
---
# Diagnose
A discipline for hard bugs. Skip phases only when explicitly justified.
When exploring the codebase, use the project's domain glossary to get a clear mental model of the relevant modules, and check ADRs in the area you're touching.
## Phase 1 — Build a feedback loop
**This is the skill.** Everything else is mechanical. If you have a fast, deterministic, agent-runnable pass/fail signal for the bug, you will find the cause — bisection, hypothesis-testing, and instrumentation all just consume that signal. If you don't have one, no amount of staring at code will save you.
Spend disproportionate effort here. **Be aggressive. Be creative. Refuse to give up.**
### Ways to construct one — try them in roughly this order
1. **Failing test** at whatever seam reaches the bug — unit, integration, e2e.
2. **Curl / HTTP script** against a running dev server.
3. **CLI invocation** with a fixture input, diffing stdout against a known-good snapshot.
4. **Headless browser script** (Playwright / Puppeteer) — drives the UI, asserts on DOM/console/network.
5. **Replay a captured trace.** Save a real network request / payload / event log to disk; replay it through the code path in isolation.
6. **Throwaway harness.** Spin up a minimal subset of the system (one service, mocked deps) that exercises the bug code path with a single function call.
7. **Property / fuzz loop.** If the bug is "sometimes wrong output", run 1000 random inputs and look for the failure mode.
8. **Bisection harness.** If the bug appeared between two known states (commit, dataset, version), automate "boot at state X, check, repeat" so you can `git bisect run` it.
9. **Differential loop.** Run the same input through old-version vs new-version (or two configs) and diff outputs.
10. **HITL bash script.** Last resort. If a human must click, drive _them_ with `scripts/hitl-loop.template.sh` so the loop is still structured. Captured output feeds back to you.
Build the right feedback loop, and the bug is 90% fixed.
### Iterate on the loop itself
Treat the loop as a product. Once you have _a_ loop, ask:
- Can I make it faster? (Cache setup, skip unrelated init, narrow the test scope.)
- Can I make the signal sharper? (Assert on the specific symptom, not "didn't crash".)
- Can I make it more deterministic? (Pin time, seed RNG, isolate filesystem, freeze network.)
A 30-second flaky loop is barely better than no loop. A 2-second deterministic loop is a debugging superpower.
### Non-deterministic bugs
The goal is not a clean repro but a **higher reproduction rate**. Loop the trigger 100×, parallelise, add stress, narrow timing windows, inject sleeps. A 50%-flake bug is debuggable; 1% is not — keep raising the rate until it's debuggable.
### When you genuinely cannot build a loop
Stop and say so explicitly. List what you tried. Ask the user for: (a) access to whatever environment reproduces it, (b) a captured artifact (HAR file, log dump, core dump, screen recording with timestamps), or (c) permission to add temporary production instrumentation. Do **not** proceed to hypothesise without a loop.
Do not proceed to Phase 2 until you have a loop you believe in.
## Phase 2 — Reproduce
Run the loop. Watch the bug appear.
Confirm:
- [ ] The loop produces the failure mode the **user** described — not a different failure that happens to be nearby. Wrong bug = wrong fix.
- [ ] The failure is reproducible across multiple runs (or, for non-deterministic bugs, reproducible at a high enough rate to debug against).
- [ ] You have captured the exact symptom (error message, wrong output, slow timing) so later phases can verify the fix actually addresses it.
Do not proceed until you reproduce the bug.
## Phase 3 — Hypothesise
Generate **3–5 ranked hypotheses** before testing any of them. Single-hypothesis generation anchors on the first plausible idea.
Each hypothesis must be **falsifiable**: state the prediction it makes.
> Format: "If <X> is the cause, then <changing Y> will make the bug disappear / <changing Z> will make it worse."
If you cannot state the prediction, the hypothesis is a vibe — discard or sharpen it.
**Show the ranked list to the user before testing.** They often have domain knowledge that re-ranks instantly ("we just deployed a change to #3"), or know hypotheses they've already ruled out. Cheap checkpoint, big time saver. Don't block on it — proceed with your ranking if the user is AFK.
## Phase 4 — Instrument
Each probe must map to a specific prediction from Phase 3. **Change one variable at a time.**
Tool preference:
1. **Debugger / REPL inspection** if the env supports it. One breakpoint beats ten logs.
2. **Targeted logs** at the boundaries that distinguish hypotheses.
3. Never "log everything and grep".
**Tag every debug log** with a unique prefix, e.g. `[DEBUG-a4f2]`. Cleanup at the end becomes a single grep. Untagged logs survive; tagged logs die.
**Perf branch.** For performance regressions, logs are usually wrong. Instead: establish a baseline measurement (timing harness, `performance.now()`, profiler, query plan), then bisect. Measure first, fix second.
## Phase 5 — Fix + regression test
Write the regression test **before the fix** — but only if there is a **correct seam** for it.
A correct seam is one where the test exercises the **real bug pattern** as it occurs at the call site. If the only available seam is too shallow (single-caller test when the bug needs multiple callers, unit test that can't replicate the chain that triggered the bug), a regression test there gives false confidence.
**If no correct seam exists, that itself is the finding.** Note it. The codebase architecture is preventing the bug from being locked down. Flag this for the next phase.
If a correct seam exists:
1. Turn the minimised repro into a failing test at that seam.
2. Watch it fail.
3. Apply the fix.
4. Watch it pass.
5. Re-run the Phase 1 feedback loop against the original (un-minimised) scenario.
## Phase 6 — Cleanup + post-mortem
Required before declaring done:
- [ ] Original repro no longer reproduces (re-run the Phase 1 loop)
- [ ] Regression test passes (or absence of seam is documented)
- [ ] All `[DEBUG-...]` instrumentation removed (`grep` the prefix)
- [ ] Throwaway prototypes deleted (or moved to a clearly-marked debug location)
- [ ] The hypothesis that turned out correct is stated in the commit / PR message — so the next debugger learns
**Then ask: what would have prevented this bug?** If the answer involves architectural change (no good test seam, tangled callers, hidden coupling) hand off to the `/improve-codebase-architecture` skill with the specifics. Make the recommendation **after** the fix is in, not before — you have more information now than when you started.
+259
View File
@@ -0,0 +1,259 @@
---
name: essence
description: Invoke when a project is too large or you only want the core design insights. Extracts 1-2 standout design patterns with deep analysis, lens-guided perspectives, and migration examples. Not for full project analysis or quick lookups.
metadata:
version: "0.5.0"
---
# Essence: Extract Core Design Patterns
Prefix your first line with 🥷 inline, not as its own paragraph.
You are a jewel inspector. A project has thousands of files — your job is to find the one or two brilliant ideas worth stealing.
**This is NOT a lite version of `/explore`.** `/explore` reads the whole project and summarizes at the end. `/essence` goes deep on one thing and ignores everything else.
## Mode Selection
First, check whether an `/explore` result exists:
- `/explore` report exists → it already identified 2-3 core designs, default to **User-directed**. Ask the user which design to deep-dive, or whether to switch mode.
- No `/explore` result → this is an independent launch, default to **Auto-detect**.
Always confirm before proceeding:
| Mode | When | Entry |
|---|---|---|
| **User-directed** | Already have a design target from `/explore`, or know exactly which design to investigate | User tells you what to look for |
| **Auto-detect** | Independent launch, project is large, want the AI to find the standout design | You find the standout design |
| **Lens-guided** | "Analyze this from a [mechanical/intentional/evolution] perspective" | Apply a specific analytical lens |
### Lens definitions
| Lens | Core question | Guided behavior |
|---|---|---|
| **Mechanical** (default) | How does it work? | Read source code, trace call chains, examine interfaces |
| **Intentional** | Why this way? | Read design docs/RFCs/PRs, extract decision rationale and tradeoffs |
| **Evolution** | How did it get here? | Read git history/changelog, compare before/after, identify migration drivers |
A lens shapes which sources to read and how to frame the output, but does not add separate phases.
### Auto-detect signals
A design is "essence" if it passes 2 or more of these signals:
| Signal | Evidence |
|---|---|
| README highlights it prominently | "Built on a plugin architecture" as a headline feature |
| Has standalone architecture docs | ARCHITECTURE.md, docs/design/, blog post by author |
| Heavily discussed in Issues/PRs | Design decisions debated by community |
| Unique among similar projects | Competitors don't do it this way |
| Rich design comments in code | JSDoc/TSDoc explaining why, not what |
| Cross-module contract | A type, interface, or protocol imported across module boundaries (not just files). Go: most-implemented interface. Python: most-subclassed abstract base. Rust: most-implemented trait. These define subsystem relationships. |
| File size anomaly | One file is disproportionately large or small for its responsibility — signals non-trivial logic |
| Dedicated test coverage | Tests specifically validate this design's behavior, not just happy paths |
**"Clean code" is NOT a signal.** A well-written utility function is not essence. An architecture decision that shapes the entire project is.
If no design passes 2+ signals, tell the user: "This project has no standout design. Try `/explore` for a full analysis instead."
## Phase 1: Locate
**User-directed mode:**
- Go directly to the directory or file the user names.
- If the directory doesn't exist, stop and tell the user. Do NOT invent an alternative.
**Auto-detect mode:**
- Scan README, AGENTS.md, and top-level docs for architecture claims.
- Identify 1-2 standout design directions.
- Present to the user: "The standout designs appear to be: A) {design A}, B) {design B}. Which should we dive into?"
- If user doesn't choose, pick the strongest one and state why.
**Lens-guided mode:**
- Confirm the lens with the user (Mechanical/Intentional/Evolution).
- Frame the search in terms of the lens.
- Example: "You want the Mechanical view — I'll trace the core implementation and extract the pattern."
**Output:** 1-2 design directions to analyze + lens confirmation.
**Stall signal:** Cannot identify any standout design → the project may be a conventional CRUD app or wrapper. Stop and recommend `/explore` or a different project.
## Phase 2: Deep Dive
Read the core files related to the chosen design. Maximum 10 files. Let the lens guide source selection: Mechanical → source code and type definitions; Intentional → design docs, RFCs, PR discussions; Evolution → git history, changelog, migration guides.
**For each file:**
- What role does it play in this design?
- What interfaces does it expose?
- How does it connect to other parts of the system?
**Trace the call chain:**
- Start from the entry point that uses this design.
- Follow the flow until you understand the full pattern.
- Stop when you hit boilerplate, config, or test files.
**Output:** Core file list (≤10) + call chain + lens-specific annotations.
**Stall signal:** The design spans more than 10 files and you can't find the boundary → the design is probably the project's core architecture. Switch to `/explore` for a full analysis instead.
## Phase 3: Extract Pattern
Analyze the design at a higher level. Let the lens shape the analysis angle:
- **Mechanical** → emphasize structure, interfaces, data flow — produce a pattern diagram + interface contracts
- **Intentional** → emphasize decision rationale, tradeoffs — produce a decision record (context → options → rationale)
- **Evolution** → emphasize before/after comparison, migration drivers — produce a timeline + catalyst events
**Universal analysis dimensions** (all lenses):
- **Problem:** What specific problem does this design solve? What was the pain before?
- **Pattern:** What's the name of this pattern? (Named: MVC, Observer, Plugin, Middleware. Custom: describe it in one sentence.)
- **Alternatives:** What simpler or more complex approaches could solve the same problem?
- **Tradeoffs:** Why did the author choose this? What does it give up?
- **Evidence:** What in the code proves this analysis is correct? (Specific files, functions, comments.)
**Output:** Design pattern card (lens-framed).
**Stall signal:** Cannot explain why the author chose this design over alternatives → read commit messages and PR discussions for design rationale. If unavailable, state "author's reasoning unknown" in the report.
## Phase 4: Migrate
Make the learning actionable. Let the lens tailor the output:
- **Mechanical** → copy-paste code skeleton (≤20 lines with TODOs)
- **Intentional** → decision framework (checklist for evaluating tradeoffs)
- **Evolution** → migration path (step-by-step refactor plan)
**Universal deliverables** (all lenses):
- **Can you use this?** Is the design applicable to the user's own projects? If not, why?
- **Steal-it example:** A simplified version (under 20 lines) that captures the core idea. Not production code — a teaching example.
- **Pitfalls:** What context does this design depend on? What would break if you copy it blindly?
**Output:** Migration example + pitfall list (lens-tailored).
**Stall signal:** The design depends on framework internals, language features, or ecosystem the user doesn't have → explain the core idea abstractly instead of providing code.
## Phase 5: Self-review
Check the report is honest:
**All modes:**
- [ ] The design is real (not inferred, not imagined). Evidence: specific files cited.
- [ ] The analysis is deep enough that you could explain it out loud.
- [ ] The migration example captures the core idea, not surface syntax.
- [ ] Pitfalls are specific, not vague ("needs X version" not "may not work everywhere").
**Stall signals (any one → return to relevant phase):**
- Cannot name a file that proves the pattern → back to Phase 2
- Cannot explain why it's better than alternatives → back to Phase 3
- Migration example is over 20 lines → simplify, back to Phase 4
- Lens-specific check failed (e.g., Mechanical missing end-to-end call chain, Intentional missing decision rationale, Evolution missing timeline) → back to relevant phase
**Output:** Essence report with lens annotation.
## Optional: HTML Card
**Only when the user explicitly requests it.**
Generate an HTML visualization card as a shareable deliverable.
### HTML Card Structure (Glassmorphism 2.0 - Essence Variant)
```html
<!DOCTYPE html>
<html lang="zh-CN">
<head>
<meta charset="UTF-8">
<title>{Project Name} - Essence Report</title>
<script src="https://cdn.tailwindcss.com"></script>
<script src="https://cdn.jsdelivr.net/npm/mermaid/dist/mermaid.min.js"></script>
<style>
/* Same glassmorphism styles as /explore */
:root { --glass-bg: rgba(255,255,255,0.4); --primary: #8b5cf6; }
[data-theme="dark"] { --glass-bg: rgba(15,23,42,0.6); --primary: #a78bfa; }
.glass-panel { backdrop-filter: blur(12px); border-radius: 1rem; }
.pattern-diagram { font-family: monospace; background: rgba(0,0,0,0.03); }
</style>
</head>
<body class="p-8">
<nav class="fixed top-4 left-1/2 -translate-x-1/2 w-[90%] max-w-4xl glass-panel z-50 px-6 py-3">
<span class="font-bold text-xl">💎 {Project Name} 精华</span>
<span class="text-sm opacity-70">Lens: {lens} | Pattern: {pattern_name}</span>
</nav>
<main class="max-w-4xl mx-auto mt-24 space-y-6">
<section class="glass-panel p-6">
<h2 class="text-xl font-bold mb-4">🎯 Design Analyzed</h2>
<p>{one-line description}</p>
</section>
<section class="glass-panel p-6">
<h2 class="text-xl font-bold mb-4">🔷 Pattern ({lens})</h2>
<!-- Lens-framed pattern card -->
</section>
<section class="glass-panel p-6">
<h2 class="text-xl font-bold mb-4">🔗 Call Chain</h2>
<pre class="mermaid">{diagram}</pre>
</section>
<section class="glass-panel p-6">
<h2 class="text-xl font-bold mb-4">📦 Migration Example</h2>
<pre class="pattern-diagram"><code>{code_example}</code></pre>
<p class="text-sm opacity-70 mt-2">Pitfalls: {pitfalls}</p>
</section>
</main>
<script>mermaid.initialize({ startOnLoad: true });</script>
</body>
</html>
```
### Output Format
```markdown
### HTML Card Generated
- **Path:** `outputs/{project}-essence.html`
- **Theme:** {modern/ink}
- **Accent Color:** Purple (essence = jewel)
```
**When to skip:** Skip HTML generation unless the user requests it or the analysis is production-critical. When HTML generation fails, deliver a plain-text report instead.
---
## Hard Rules
- **No code evidence = no conclusion.** Every claim about a design must cite a specific file, function, or comment.
- **Under 20 lines for migration examples.** If you can't explain the idea in 20 lines, you don't understand it well enough.
- **Stop after the report.** Do not modify the user's project or the target project.
- **HTML is optional.** Do not block analysis on HTML generation.
## Gotchas
| What happened | Rule |
|---|---|
| 提取的"精华"是 AI 脑补的 | 必须有代码证据(文件 + 行号),不写空泛结论 |
| 用户指定方向但该模块不存在 | 停止并告知用户,不编造替代方向 |
| 项目没有 standout 设计(胶水代码) | 标记"无可提取精华",建议改用 `/explore` |
| Phase 4 迁移示例超过 20 行 | 简化到核心思路,不是复制生产代码 |
| 分析了一个小工具函数 | 工具函数不是设计。设计影响整个架构,工具只解决一个问题 |
| 从 commit message 推断作者意图但没有代码佐证 | Commit message 是辅助证据,必须有代码结构本身的支持 |
| 透镜模式选错导致输出不符预期 | Phase 1 先确认透镜,Mechanical 读代码、Intentional 读文档、Evolution 读历史 |
| 透镜分析流于表面 | 每个透镜有特定输出格式:Mechanical→图 + 接口,Intentional→决策记录,Evolution→时间线 |
| HTML 卡片生成失败 | 降级到纯文本报告,不阻塞分析交付 |
## Outcome
```
Essence Report: {project name}
Lens: mechanical / intentional / evolution
Design analyzed: {one-line description}
Files examined: {count}
Pattern: {pattern name or custom description}
Migration: {steal-it example, ≤20 lines}
HTML generated: yes / no
Status: complete
```
After the report, stop. No modifications. No follow-ups.
@@ -0,0 +1,79 @@
# Essence Detection Signals
How to identify the standout design in a project when the user doesn't specify a direction.
## Signal Strength
A design passes the "essence" threshold if it scores 2+ signals.
### Strong Signals (score = 1 each)
| Signal | How to detect | Example |
|---|---|---|
| **README headline** | Project name is followed by a design claim | "Vite — Next generation frontend tooling with **ESM-first architecture**" |
| **Architecture docs** | Standalone design document exists | `ARCHITECTURE.md`, `docs/design/`, `docs/architecture/` |
| **Official blog post** | Author wrote about the design on their blog | tw93.fun, Vite blog, React blog posts |
| **Community discussion** | Issues/PRs debate the design decision | "Why we chose X over Y" discussions with many comments |
| **Rich code comments** | JSDoc/TSDoc explaining WHY, not WHAT | "We use this pattern because..." with detailed reasoning |
### Objective Signals (score = 1 each, no subjective judgment needed)
| Signal | How to detect | Example |
|---|---|---|
| **Cross-module contract** | A type, interface, or protocol imported across module boundaries (not just files). Go: most-implemented interface. Python: most-subclassed abstract base. Rust: most-implemented trait. | `Plugin` interface implemented by 8 subsystems, each in its own package |
| **File size anomaly** | One file's line count is ≥3× the median for its category (handlers, utils, etc.) | Average handler: 50 lines. One handler: 800 lines with state machine logic |
| **Dedicated test coverage** | Tests exist specifically for this design's edge cases, not just happy paths | `plugin.test.ts` tests plugin resolution, fallback, lifecycle — not just "it loads" |
### Weak Signals (score = 0.5 each)
| Signal | How to detect | Example |
|---|---|---|
| **Unique among competitors** | Same category, different architecture | Next.js uses SSR, Remix uses nested routes — that difference IS the essence |
| **Most-starred files** | GitHub shows stars/bookmarks on specific files | "This file has 200+ stars on GitHub" |
| **Core algorithm** | One file contains non-trivial logic that drives the project | Diff algorithm, compiler pass, state machine |
| **API design** | The public API is notably elegant or unusual | `create()` returns a builder chain, not an object |
## Not Signals
These do NOT count as essence:
- "Clean code" or "well organized" — that's quality, not design
- "Uses TypeScript" — that's a language choice, not architecture
- "Has good tests" — that's engineering discipline, not design
- "Many stars on the repo" — popularity ≠ design quality
- "Uses the latest framework" — following trends ≠ standing out
- Utility functions — even well-written ones are tools, not designs
## Auto-detect Procedure
When the user says "find the essence":
1. **Read README fully.** What is the #1 feature the author leads with? That's a candidate.
2. **Check for design docs.** Is there `ARCHITECTURE.md` or equivalent? That's a candidate.
3. **Scan the import graph.** Which file is imported by the most other files? Use `grep -r "import.*from" src/ | sort | uniq -c | sort -rn` or equivalent. The top result is likely the core.
4. **Check file sizes.** Are any files disproportionately large or small for their apparent role? That signals hidden complexity.
5. **Check uniqueness.** Compare with 1-2 well-known alternatives. What does this project do differently?
6. **Present 1-2 candidates** to the user with evidence. Let them choose or auto-select the strongest.
### Example Output Format
```
Standout designs in {project}:
A) {Design A name} — evidenced by {README claim / file / doc}
What it does: {one sentence}
B) {Design B name} — evidenced by {code comment / unique feature / community discussion}
What it does: {one sentence}
Which should we dive into? (or I can pick the strongest)
```
## Failure Modes
| Situation | Response |
|---|---|
| No signal passes 2+ threshold | "This project uses conventional architecture. Try `/explore` for a full analysis, or pick a more architecturally interesting project." |
| User-specified module doesn't exist | Stop. Do NOT suggest an alternative. Tell the user the path doesn't exist. |
| Project is a wrapper (thin layer over another tool) | "This project is primarily a wrapper around {X}. The design is in {X}, not here. Try analyzing {X} instead." |
| Project is configuration-only (just JSON/YAML files) | "This project has no code architecture. It's configuration-driven. Try `/explore` for a full overview instead." |
+87
View File
@@ -0,0 +1,87 @@
---
name: explore
description: Invoke when you need project-level understanding and an onboarding path. Produces a project learning report for code and non-code repositories with fixed phases for positioning, structure, flow, start path, and core designs. Not for deep code extraction or interactive teaching.
metadata:
version: "0.5.0"
---
# Explore: Project Understanding and Onboarding
Prefix your first line with 🥷 inline, not as its own paragraph.
You are a project cartographer. Your job is to help the user understand what a project is, why it is worth studying, how it is organized, and where to start.
`/explore` is the entry point for first contact with a repository or project-like artifact. It builds global understanding. It does not perform code-level essence extraction and it does not run interactive teaching.
## Project Type Detection
After the initial scan, classify the target before continuing:
| Type | Signals | What changes |
|---|---|---|
| **Code repository** | `go.mod`, `pyproject.toml`, `Cargo.toml`, source directories, executable entrypoints | Run all 4 phases |
| **Skill / docs / knowledge repository** | `SKILL.md`, mostly Markdown, docs-first structure, no runnable application entrypoint | Skip Phase 2 (Flow) and Phase 3 (Start Path) |
| **Template / scaffold repository** | Starter files, minimal logic, setup-first repo | Phase 2 may stay structural and Phase 3 may be minimal |
State the detected type before proceeding. If uncertain, say what evidence is missing and continue with the closest matching type.
## Phase 1: Positioning & Structure
- What this project is, why it is worth studying, and who it is for.
- Top-level structure: main modules, documents, directories, and the likely learning entry area.
- Tradeoffs vs alternatives when evidence exists.
## Phase 2: Flow
**Code repositories only.**
- Skip for non-code and template repositories.
- Trace the main runtime or request flow.
- Produce at least one architecture or core-flow diagram.
- Keep the trace focused on the golden path rather than exhaustive coverage.
## Phase 3: Start Path
**Code repositories only when runnable or meaningfully inspectable.**
- Provide the minimal path to start learning or running the project.
- Give the first command or first inspection step.
- Suggest one safe first modification or observation point when appropriate.
## Phase 4: Core Designs
- Summarize 2-3 core implementations or ideas.
- Keep this at overview depth.
- For each item, include what it is, where it lives, and why it matters.
## Minimum Deliverables
The final `/explore` report must include:
- Project positioning
- Why it is worth studying
- 2-3 core implementations or core ideas
- Tradeoffs or comparisons when applicable
- At least 1 diagram:
- code repository → architecture diagram or core flow diagram
- non-code repository → structure diagram, idea map, or workflow diagram
## Boundary Rules
`/explore` may:
- scan structure
- explain the main flow
- provide a minimal start path
- summarize 2-3 core designs
`/explore` must not:
- perform `/essence`-level deep extraction
- act as `/follow`-style guided teaching
- include Verify, Deep Fission, or HTML Output phases
- preserve no retired lightweight fallback behavior
## Outcome
```
Explore Report: {project name}
Project type: code / skill-docs / template
Phases completed: 4/4 (or note skipped code-only phases)
Diagram included: yes / no
Core designs: 2-3
Status: complete
```
After the report, stop. Do not proceed to `/essence` or `/follow` automatically.
@@ -0,0 +1,98 @@
# Project Analysis Methods
How to read and understand an unfamiliar code project.
## 1. Identify the Entry Point
Every project has a door. Find it first.
### By Language
| Language | Look for |
|---|---|
| **JavaScript/TypeScript** | `package.json` → `main` / `bin` / `scripts.dev` |
| **Python** | `setup.py` → `entry_points`, `pyproject.toml` → `[project.scripts]`, or top-level `app.py` / `main.py` / `__main__.py` |
| **Go** | `package main` in any file, conventionally `main.go` or `cmd/*/main.go` |
| **Rust** | `src/main.rs` or `src/bin/*.rs` |
| **Java** | Class with `public static void main(String[] args)` |
| **C/C++** | `main()` function, conventionally in `src/main.c` |
| **Swift** | `main.swift` or file with `@main` attribute |
### In Frameworks
| Framework | Entry point |
|---|---|
| Next.js | `app/` or `pages/` directory, `next.config.js` |
| React (Vite) | `src/main.tsx` or `src/main.jsx` |
| Vue (Vite) | `src/main.ts` or `src/main.js` |
| Express | File that calls `app.listen()` |
| FastAPI | File that creates `FastAPI()` instance |
| Django | `manage.py`, then project name directory with `urls.py` / `wsgi.py` |
| Flask | `app.py` or `app/__init__.py` |
| Spring Boot | `*Application.java` with `@SpringBootApplication` |
## 2. Judge Project Complexity
Don't over-engineer simple projects. Don't under-analyze complex ones.
### Simple (<50 files, single language)
- Read every source file.
- No need for flow diagrams beyond a simple sequence.
- A light `/explore` pass is probably enough.
### Standard (50-500 files, 1-2 languages)
- Read entry point + core modules + 1-2 feature files.
- Build 1-2 flow diagrams.
- `/explore` is the right level.
### Complex (>500 files, multi-language, monorepo)
- Read entry point + architecture docs + one representative module.
- Use `/essence` to find standout designs, or `/explore` for one package at a time.
- Do NOT try to understand the whole project in one pass.
## 3. Separate Core Code from Scaffolding
Not all files are worth reading.
### Ignore (scaffolding)
- `*.config.js`, `*.config.ts` — configuration, not logic
- `dist/`, `build/`, `out/` — generated output
- `node_modules/`, `vendor/`, `.venv/` — dependencies
- `*.lock`, `yarn.lock`, `go.sum` — lock files
- `LICENSE`, `CODEOWNERS`, `.editorconfig` — project meta
- `test/fixtures/`, `test/data/` — test data
### Read (core)
- Entry point file
- Router/middleware/config handlers
- Model/entity/schema definitions
- Core algorithm or business logic files
- Files referenced most in imports
### Hint: Follow imports
```
entry file → import A → import B → core logic
```
Each import is a dependency. Follow the chain until you hit a file that doesn't import anything else — that's usually the core.
## 4. Read Unfamiliar Framework Code
You don't know every framework. That's fine.
### Strategy
1. **Find the routing layer first.** Every framework has a way to map URLs or events to handlers. Find it. It tells you the project's capabilities.
2. **Follow ONE request end-to-end.** Don't try to understand all routes. Pick the simplest one (often "health check" or "get by ID") and trace it from entry to response.
3. **Identify the framework's conventions.** Most frameworks follow a pattern:
- MVC: Controller → Model → View
- Middleware: Request → Middleware chain → Handler → Response
- Component: Parent renders children, props flow down, events flow up
- Plugin: Core calls hooks, plugins register handlers
4. **Don't fight the framework's abstraction.** If the project uses ORM, don't look for raw SQL. If it uses dependency injection, don't look for `new()` calls. Understand what abstraction layer they chose.
5. **Use the framework's own docs.** If stuck on "how does this framework work?", check the official docs. Don't reverse-engineer what's documented.
@@ -0,0 +1,173 @@
# Flow Pattern Library
Common architecture patterns and how to identify them in code.
## MVC / MVVM / MVX
### What it is
Separation of data (Model), UI/presentation (View), and coordination logic (Controller/ViewModel).
### File signatures
| Pattern | Directories/Files |
|---|---|
| **MVC** | `controllers/`, `models/`, `views/` |
| **MVVM** | `viewmodels/`, `views/`, `models/` |
| **Layered** | `app/`, `domain/`, `infrastructure/` (Clean/Hexagonal) |
### Flow
```
Request → Controller → Model (data) → View (render) → Response
```
### Key question
"Does the file handle data, display, or coordination?" If yes → MVC-family.
---
## Middleware Chain
### What it is
Each handler processes the request and passes it to the next. Like an assembly line.
### File signatures
| Framework | Indicator |
|---|---|---|
| **Express/Koa** | `app.use(...)`, `app.get('/', handler)` |
| **FastAPI** | `@app.middleware("http")`, `Depends()` |
| **Next.js** | `middleware.ts` at root or in `app/` |
| **Gin (Go)** | `router.Use(middleware1, middleware2)` |
| **Koa** | `app.use(async (ctx, next) => { ... })` |
### Flow
```
Request → Middleware A → Middleware B → Handler → Response
↓ ↓
auth check log request
```
### Key question
"Does this function call `next()` or pass control to something else?" If yes → middleware.
### Common middleware order
```
1. CORS / Security headers
2. Logging / Request ID
3. Authentication / Authorization
4. Body parsing / Validation
5. Rate limiting
6. Route handler
7. Error handler (catches everything above)
```
---
## Plugin / Extension System
### What it is
Core provides hooks or interfaces. External code registers handlers. The core doesn't know about specific plugins.
### File signatures
| Pattern | Indicator |
|---|---|
| **Hook-based** | `registerHook('eventName', handler)`, `hooks.on('event', fn)` |
| **Interface-based** | Abstract class or interface that plugins implement |
| **Discovery-based** | Directory scan (`plugins/`), import all, register by convention |
| **VSCode-style** | `contributes` in `package.json`, activation events |
### Flow
```
Core starts
↓
Scans for plugins
↓
Each plugin registers itself
↓
Core fires hooks → plugins respond
↓
Core runs with extended capabilities
```
### Key question
"Can I add functionality without modifying core code?" If yes → plugin architecture.
---
## Event-Driven
### What it is
Components communicate through events, not direct calls. Publishers emit, subscribers listen.
### File signatures
| Pattern | Indicator |
|---|---|
| **Node EventEmitter** | `eventEmitter.on('event', handler)`, `eventEmitter.emit('event', data)` |
| **Pub/Sub** | `pubsub.subscribe('channel', handler)`, `pubsub.publish('channel', data)` |
| **Redux-style** | `dispatch(action)`, `reducer(state, action) → newState` |
| **Observable** | `observable.subscribe(fn)`, `pipe(map, filter)` |
| **Signals (Python)** | `@signal.connect`, `signal.send()` |
### Flow
```
Component A emits "user.created"
↓
Listener B hears it → sends welcome email
Listener C hears it → creates default settings
Listener D hears it → logs analytics
```
### Key question
"Does code communicate without importing or calling each other directly?" If yes → event-driven.
---
## State Management
### What it is
Centralized storage for application state. Components read and update through defined interfaces.
### File signatures
| Pattern | Indicator |
|---|---|
| **Redux** | `createStore()`, `dispatch()`, `useSelector()`, `@reduxjs/toolkit` |
| **Zustand** | `create((set) => ({ ... }))` |
| **Jotai** | `atom(value)`, `useAtom(atom)` |
| **MobX** | `@observable`, `@action`, `@computed` |
| **React Context** | `createContext()`, `useContext()`, `Provider` |
| **Pinia (Vue)** | `defineStore()`, `state`, `actions` |
### Flow
```
Component dispatches action
↓
Reducer processes action + current state
↓
New state emitted
↓
Subscribed components re-render
```
### Key question
"Where does the app store data that multiple components need?" If it's a single store → state management pattern.
---
## Pipeline / Chain of Responsibility
### What it is
Data flows through a series of processors. Each processor transforms the data and passes it on.
### File signatures
| Pattern | Indicator |
|---|---|
| **Stream processing** | `.pipe(transform1).pipe(transform2)` |
| **Compiler/lexer** | Source → Tokenize → Parse → Transform → Generate |
| **Data pipeline** | `input → transform → validate → output` |
| **Makefile** | Target depends on prerequisites, each is a step |
### Flow
```
Raw input → Tokenizer → Parser → Transformer → Generator → Output
```
### Key question
"Does data get progressively transformed through a fixed sequence of steps?" If yes → pipeline.
+101
View File
@@ -0,0 +1,101 @@
---
name: follow
description: Invoke when the user wants an interactive learning session based on an existing `/explore` or `/essence` report. Guides runnable or reader-style follow-along sessions. Not for fresh project analysis or pattern-only extraction.
metadata:
version: "0.5.0"
---
# Follow: Guided Learning Session
Prefix your first line with 🥷 inline, not as its own paragraph.
You are a guide. The user wants to learn from a project step by step with help, context, and correction. You guide the learning process, but you do not replace it.
`/follow` is not a fresh project analyzer. It only works from an existing `/explore` or `/essence` result.
## Pre-check
`/follow` only works when there is already an `/explore` report or an `/essence` report.
- `/explore` report exists → use it as the main learning path
- `/essence` report exists → use it for design-focused guided study
- Neither exists → refuse clearly
Refusal behavior:
"I need an existing `/explore` or `/essence` result before I can guide a follow-along session. Please run `/explore` for project understanding or `/essence` for a focused deep dive first."
Load the existing report before continuing.
## Mode Selection
After the pre-check, select one mode based on the prerequisite report:
- From `/explore` + code repository → default **Runnable**
- From `/explore` + non-code repository → force **Reader**
- From `/essence` → default **Reader** (user is in design-analysis state)
| Mode | When | Entry |
|---|---|---|
| **Runnable** | Report confirms the project is a runnable code repository and the user wants to learn by running and changing it | Start from environment and first execution |
| **Reader** | Project has no runtime, or the user is studying design/architecture, or the prerequisite report is from `/essence` | Start from guided reading |
State the selected mode before proceeding. Do not re-scan the project — use the prerequisite report to decide.
## Teaching Interaction Rules
`/follow` must teach by guidance, not by dumping answers:
- explain the purpose of the current step first
- give the user an observation point or action point
- ask the user to predict, try, or explain before revealing the answer
- then reveal, correct, or deepen the explanation
- never say "go read the code" as a standalone instruction. When referencing code, always start with: what design idea this code embodies, why it matters in the overall architecture, and what the user should pay attention to
## Runnable Check
Before Runnable mode, confirm from the **prerequisite report** (do not re-scan the project):
- If the report identified the target as a code repository with a recognized runtime (`go.mod`, `pyproject.toml`, `Cargo.toml`, `Makefile`, `build.gradle`, `pom.xml`, `CMakeLists.txt`, etc.), proceed with Runnable.
- If the report classified it as non-code, or no runtime entrypoint was found, switch to Reader and explain why.
- If the prerequisite is `/essence`, confirm with the user: essence is design-focused, Reader is the natural fit. Allow Runnable only if the user explicitly insists.
- Do not introduce a third mode.
## Runnable Mode Flow
1. Confirm environment and prerequisites.
2. Let the user run the project.
3. Let the user make one safe change.
4. Walk the main flow together.
5. Give one small exercise.
6. Review what they learned.
## Reader Mode Flow
1. Frame the learning goal around a core design or architectural idea, not a single file.
2. Walk through the design concept layer by layer: problem → approach → implementation → tradeoff.
3. Ask the user questions that probe understanding ("Why did the author choose this approach over a simpler one?"), not just prediction ("What happens next?").
4. Use diagrams or structured summaries to connect the dots between files and design ideas.
5. Give one reasoning exercise that tests whether the user can apply the design pattern elsewhere.
6. Review what they learned.
## Boundary Rules
`/follow` must:
- depend on `/explore` or `/essence`
- guide the user interactively
- adapt between code and non-code repositories through Runnable or Reader emphasis
`/follow` must not:
- rescan the whole project as a new analyzer
- reference retired skills as prerequisites
- add any third learning mode
- execute commands or write code for the user
## Outcome
```
Follow Session: {project name}
Mode: runnable / reader
Prerequisite report: /explore or /essence
Exercise result: completed / partial / too hard
Next direction: {suggested follow-up}
Status: complete
```
After the review, stop. Ask whether the user wants another exercise or wants to end the session.
@@ -0,0 +1,113 @@
# Environment Detection Rules
How to detect the runtime environment and guide the user through setup in `/follow`.
## Language Detection from Config
Check these files in order. The first match is the primary language.
| Config file | Language | Runtime check | Install command |
|---|---|---|---|
| `package.json` | JavaScript/TypeScript | `node --version` | nvm or official installer |
| `pyproject.toml` | Python | `python --version` | pyenv or python.org |
| `go.mod` | Go | `go version` | golang.org/dl |
| `Cargo.toml` | Rust | `rustc --version` | rustup |
| `pom.xml` | Java | `java -version` | SDKMAN or official |
| `build.gradle` / `build.gradle.kts` | Java/Kotlin | `java -version` | SDKMAN |
| `Gemfile` | Ruby | `ruby --version` | rvm or rbenv |
| `*.csproj` | C#/.NET | `dotnet --version` | .NET SDK |
| `CMakeLists.txt` | C/C++ | `gcc --version` or `clang --version` | System package manager |
| `swift package.json` | Swift | `swift --version` | Xcode or swift.org |
## Dependency Installation
Once language is detected, guide the user:
### JavaScript/TypeScript
```bash
# Check which package manager is used
if [ -f "yarn.lock" ]; then yarn install
elif [ -f "pnpm-lock.yaml" ]; then pnpm install
elif [ -f "bun.lockb" ] || [ -f "bun.lock" ]; then bun install
else npm install
fi
```
### Python
```bash
# Modern Python projects
pip install -e .
# Or with requirements
pip install -r requirements.txt
# Or with poetry
poetry install
# Or with uv
uv pip install -r requirements.txt
```
### Go
```bash
go mod download
```
### Rust
```bash
cargo build
```
### Java (Maven)
```bash
mvn install
```
### Java (Gradle)
```bash
./gradlew build
# or
gradle build
```
## Run Command Detection
How to start the project:
| Source | Command |
|---|---|
| `package.json` → `scripts.dev` | `npm run dev` |
| `package.json` → `scripts.start` | `npm start` |
| `Makefile` → `dev` target | `make dev` |
| `Makefile` → `run` target | `make run` |
| `pyproject.toml` (Poetry) | `poetry run python main.py` |
| `go.mod` → `package main` | `go run main.go` |
| `Cargo.toml` → `[[bin]]` | `cargo run` |
| `docker-compose.yml` exists | `docker-compose up` |
| `Dockerfile` exists, no compose | `docker build -t app . && docker run app` |
## Common Environment Issues
| Error | Cause | Fix |
|---|---|---|
| `command not found: node` | Node.js not installed | Install Node.js (recommend LTS) |
| `ModuleNotFoundError` | Python deps not installed | Run `pip install -r requirements.txt` |
| `EACCES: permission denied` | Global install without sudo | Use nvm/fnm, or prefix with sudo |
| `ENOENT: no such file` | Wrong working directory | `cd` to project root first |
| `port already in use` | Another process on same port | Kill the process or use different port |
| `go: cannot find main module` | Outside Go module | `cd` to directory with `go.mod` |
| `error: could not find Cargo.toml` | Outside Rust project | `cd` to directory with `Cargo.toml` |
| `java.lang.UnsupportedClassVersionError` | Wrong Java version | Match JDK version to project requirement |
| `npm ERR! code ERESOLVE` | Dependency conflict | Try `npm install --legacy-peer-deps` |
## Detection Script for /follow
```bash
# Quick environment check
echo "=== Environment ==="
node --version 2>/dev/null || echo "Node.js: not installed"
python --version 2>/dev/null || echo "Python: not installed"
go version 2>/dev/null || echo "Go: not installed"
rustc --version 2>/dev/null || echo "Rust: not installed"
java -version 2>/dev/null || echo "Java: not installed"
echo "PWD: $(pwd)"
```
Run this at the start of `/follow` Step 1 to understand what's available.
+57
View File
@@ -0,0 +1,57 @@
# Frontend Design — Complete Guidance
This document provides a comprehensive framework for creating visually distinctive, non-templated UI designs. Here's the full breakdown:
## Foundational Approach
Act as the design lead for a studio known for unique client identities — the client has already turned down template-like proposals. Every choice about palette, typography, and layout must be specific to the brief, including "one real aesthetic risk you can justify."
## Grounding in Subject Matter
If the brief is vague about the product or subject, pin it down yourself: name the subject, its audience, and the page's single job. Draw inspiration from "the subject's own world, its materials, instruments, artifacts, and vernacular." Use any known context about the human's preferences or past designs as hints.
## Design Principles
- **Hero as thesis**: Open with "the most characteristic thing in the subject's world" — avoid default choices like a big number with a small label and gradient accent unless truly optimal.
- **Typography**: Pair display and body faces deliberately, not from your usual repertoire. Set a clear type scale with intentional weights, widths, and spacing. "Make the type treatment itself a memorable part of the design."
- **Structure as information**: Numbering, eyebrows, dividers must encode something true about the content. Question whether numbered markers (01/02/03) actually make sense before using them — only appropriate for real sequences.
- **Motion**: Consider where animation serves the subject. "An orchestrated moment usually lands harder than scattered effects." Sometimes less is better to avoid an AI-generated feel.
- **Complexity**: Match execution to the vision — maximalist needs elaborate execution, minimal needs precision.
- **Content**: Come up with copy if the brief lacks it. Poor copy makes a design feel as templated as poor layout.
## AI-Generated Design Traps
Three common AI-default looks to watch for: (1) warm cream background (~#F4F1EA) with serif display and terracotta accent; (2) near-black with bright acid-green or vermilion; (3) broadsheet layout with hairline rules, zero border-radius, and dense columns. "All three are legitimate for some briefs, but they are defaults rather than choices." Where the brief leaves an axis free, don't spend that freedom on a default.
## Two-Pass Process
**Pass 1 — Plan**: Create a compact token system:
1. **Color**: 4–6 named hex values
2. **Type**: Characterful display face (used with restraint), complementary body face, utility face for captions/data
3. **Layout**: One-sentence prose descriptions + ASCII wireframes
4. **Signature**: The single unique element the page will be remembered by
Review the plan against the brief. If any part reads like what you'd produce for any similar page, revise it. Only then write code.
**Pass 2 — Build**: Follow the revised plan exactly. Watch for CSS selector specificity conflicts (e.g., `.section` and `.cta` fighting over padding/margins). Do most planning internally, only sharing ideas when confident.
## Restraint & Self-Critique
"Spend your boldness in one place" — let the signature element be the one memorable thing; keep everything else quiet. "Not taking a risk can be a risk itself!" Build responsively down to mobile, with visible keyboard focus and reduced motion respected. Critique as you build. Follow Chanel's advice: before finishing, remove one accessory. Jot notes about what you've tried to avoid repeating yourself.
## Writing in Design
Words exist to make the design understandable and usable — they're "design material, not decoration." Write from the end user's perspective, naming things by what people control and recognize, never by how the system is built.
- Use active voice as default
- A control should say exactly what happens: "Save changes," not "Submit"
- Maintain consistent vocabulary throughout flows (button says "Publish," toast says "Published")
- Treat errors as guidance, not mood — explain what went wrong and how to fix it
- Empty screens are invitations to act
- Keep the register conversational: "plain verbs, sentence case, no filler"
- Let each element do exactly one job — "a label labels, an example demonstrates"
## License
Apache License 2.0 — see LICENSE.txt
@@ -0,0 +1,83 @@
---
name: gitnexus-cli
description: "Use when the user needs to run GitNexus CLI commands like analyze/index a repo, check status, clean the index, generate a wiki, or list indexed repos. Examples: \"Index this repo\", \"Reanalyze the codebase\", \"Generate a wiki\""
---
# GitNexus CLI Commands
All commands work via `npx` — no global install required.
## Commands
### analyze — Build or refresh the index
```bash
npx gitnexus analyze
```
Run from the project root. This parses all source files, builds the knowledge graph, writes it to `.gitnexus/`, and generates AGENTS.md / AGENTS.md context files.
| Flag | Effect |
| -------------- | ---------------------------------------------------------------- |
| `--force` | Force full re-index even if up to date |
| `--embeddings` | Enable embedding generation for semantic search (off by default) |
| `--drop-embeddings` | Drop existing embeddings on rebuild. By default, an `analyze` without `--embeddings` preserves them. |
**When to run:** First time in a project, after major code changes, or when `gitnexus://repo/{name}/context` reports the index is stale. In Codex, a PostToolUse hook detects staleness after `git commit` and `git merge` and notifies the agent to run `analyze` — the hook does not run analyze itself, to avoid blocking the agent for up to 120s and risking KuzuDB corruption on timeout.
### status — Check index freshness
```bash
npx gitnexus status
```
Shows whether the current repo has a GitNexus index, when it was last updated, and symbol/relationship counts. Use this to check if re-indexing is needed.
### clean — Delete the index
```bash
npx gitnexus clean
```
Deletes the `.gitnexus/` directory and unregisters the repo from the global registry. Use before re-indexing if the index is corrupt or after removing GitNexus from a project.
| Flag | Effect |
| --------- | ------------------------------------------------- |
| `--force` | Skip confirmation prompt |
| `--all` | Clean all indexed repos, not just the current one |
### wiki — Generate documentation from the graph
```bash
npx gitnexus wiki
```
Generates repository documentation from the knowledge graph using an LLM. Requires an API key (saved to `~/.gitnexus/config.json` on first use).
| Flag | Effect |
| ------------------- | ----------------------------------------- |
| `--force` | Force full regeneration |
| `--model <model>` | LLM model (default: minimax/minimax-m2.5) |
| `--base-url <url>` | LLM API base URL |
| `--api-key <key>` | LLM API key |
| `--concurrency <n>` | Parallel LLM calls (default: 3) |
| `--gist` | Publish wiki as a public GitHub Gist |
### list — Show all indexed repos
```bash
npx gitnexus list
```
Lists all repositories registered in `~/.gitnexus/registry.json`. The MCP `list_repos` tool provides the same information.
## After Indexing
1. **Read `gitnexus://repo/{name}/context`** to verify the index loaded
2. Use the other GitNexus skills (`exploring`, `debugging`, `impact-analysis`, `refactoring`) for your task
## Troubleshooting
- **"Not inside a git repository"**: Run from a directory inside a git repo
- **Index is stale after re-analyzing**: Restart Codex to reload the MCP server
- **Embeddings slow**: Omit `--embeddings` (it's off by default) or set `OPENAI_API_KEY` for faster API-based embedding
@@ -0,0 +1,89 @@
---
name: gitnexus-debugging
description: "Use when the user is debugging a bug, tracing an error, or asking why something fails. Examples: \"Why is X failing?\", \"Where does this error come from?\", \"Trace this bug\""
---
# Debugging with GitNexus
## When to Use
- "Why is this function failing?"
- "Trace where this error comes from"
- "Who calls this method?"
- "This endpoint returns 500"
- Investigating bugs, errors, or unexpected behavior
## Workflow
```
1. gitnexus_query({query: "<error or symptom>"}) → Find related execution flows
2. gitnexus_context({name: "<suspect>"}) → See callers/callees/processes
3. READ gitnexus://repo/{name}/process/{name} → Trace execution flow
4. gitnexus_cypher({query: "MATCH path..."}) → Custom traces if needed
```
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
## Checklist
```
- [ ] Understand the symptom (error message, unexpected behavior)
- [ ] gitnexus_query for error text or related code
- [ ] Identify the suspect function from returned processes
- [ ] gitnexus_context to see callers and callees
- [ ] Trace execution flow via process resource if applicable
- [ ] gitnexus_cypher for custom call chain traces if needed
- [ ] Read source files to confirm root cause
```
## Debugging Patterns
| Symptom | GitNexus Approach |
| -------------------- | ---------------------------------------------------------- |
| Error message | `gitnexus_query` for error text → `context` on throw sites |
| Wrong return value | `context` on the function → trace callees for data flow |
| Intermittent failure | `context` → look for external calls, async deps |
| Performance issue | `context` → find symbols with many callers (hot paths) |
| Recent regression | `detect_changes` to see what your changes affect |
## Tools
**gitnexus_query** — find code related to error:
```
gitnexus_query({query: "payment validation error"})
→ Processes: CheckoutFlow, ErrorHandling
→ Symbols: validatePayment, handlePaymentError, PaymentException
```
**gitnexus_context** — full context for a suspect:
```
gitnexus_context({name: "validatePayment"})
→ Incoming calls: processCheckout, webhookHandler
→ Outgoing calls: verifyCard, fetchRates (external API!)
→ Processes: CheckoutFlow (step 3/7)
```
**gitnexus_cypher** — custom call chain traces:
```cypher
MATCH path = (a)-[:CodeRelation {type: 'CALLS'}*1..2]->(b:Function {name: "validatePayment"})
RETURN [n IN nodes(path) | n.name] AS chain
```
## Example: "Payment endpoint returns 500 intermittently"
```
1. gitnexus_query({query: "payment error handling"})
→ Processes: CheckoutFlow, ErrorHandling
→ Symbols: validatePayment, handlePaymentError
2. gitnexus_context({name: "validatePayment"})
→ Outgoing calls: verifyCard, fetchRates (external API!)
3. READ gitnexus://repo/my-app/process/CheckoutFlow
→ Step 3: validatePayment → calls fetchRates (external)
4. Root cause: fetchRates calls external API without proper timeout
```
@@ -0,0 +1,78 @@
---
name: gitnexus-exploring
description: "Use when the user asks how code works, wants to understand architecture, trace execution flows, or explore unfamiliar parts of the codebase. Examples: \"How does X work?\", \"What calls this function?\", \"Show me the auth flow\""
---
# Exploring Codebases with GitNexus
## When to Use
- "How does authentication work?"
- "What's the project structure?"
- "Show me the main components"
- "Where is the database logic?"
- Understanding code you haven't seen before
## Workflow
```
1. READ gitnexus://repos → Discover indexed repos
2. READ gitnexus://repo/{name}/context → Codebase overview, check staleness
3. gitnexus_query({query: "<what you want to understand>"}) → Find related execution flows
4. gitnexus_context({name: "<symbol>"}) → Deep dive on specific symbol
5. READ gitnexus://repo/{name}/process/{name} → Trace full execution flow
```
> If step 2 says "Index is stale" → run `npx gitnexus analyze` in terminal.
## Checklist
```
- [ ] READ gitnexus://repo/{name}/context
- [ ] gitnexus_query for the concept you want to understand
- [ ] Review returned processes (execution flows)
- [ ] gitnexus_context on key symbols for callers/callees
- [ ] READ process resource for full execution traces
- [ ] Read source files for implementation details
```
## Resources
| Resource | What you get |
| --------------------------------------- | ------------------------------------------------------- |
| `gitnexus://repo/{name}/context` | Stats, staleness warning (~150 tokens) |
| `gitnexus://repo/{name}/clusters` | All functional areas with cohesion scores (~300 tokens) |
| `gitnexus://repo/{name}/cluster/{name}` | Area members with file paths (~500 tokens) |
| `gitnexus://repo/{name}/process/{name}` | Step-by-step execution trace (~200 tokens) |
## Tools
**gitnexus_query** — find execution flows related to a concept:
```
gitnexus_query({query: "payment processing"})
→ Processes: CheckoutFlow, RefundFlow, WebhookHandler
→ Symbols grouped by flow with file locations
```
**gitnexus_context** — 360-degree view of a symbol:
```
gitnexus_context({name: "validateUser"})
→ Incoming calls: loginHandler, apiMiddleware
→ Outgoing calls: checkToken, getUserById
→ Processes: LoginFlow (step 2/5), TokenRefresh (step 1/3)
```
## Example: "How does payment processing work?"
```
1. READ gitnexus://repo/my-app/context → 918 symbols, 45 processes
2. gitnexus_query({query: "payment processing"})
→ CheckoutFlow: processPayment → validateCard → chargeStripe
→ RefundFlow: initiateRefund → calculateRefund → processRefund
3. gitnexus_context({name: "processPayment"})
→ Incoming: checkoutHandler, webhookHandler
→ Outgoing: validateCard, chargeStripe, saveTransaction
4. Read src/payments/processor.ts for implementation details
```
@@ -0,0 +1,64 @@
---
name: gitnexus-guide
description: "Use when the user asks about GitNexus itself — available tools, how to query the knowledge graph, MCP resources, graph schema, or workflow reference. Examples: \"What GitNexus tools are available?\", \"How do I use GitNexus?\""
---
# GitNexus Guide
Quick reference for all GitNexus MCP tools, resources, and the knowledge graph schema.
## Always Start Here
For any task involving code understanding, debugging, impact analysis, or refactoring:
1. **Read `gitnexus://repo/{name}/context`** — codebase overview + check index freshness
2. **Match your task to a skill below** and **read that skill file**
3. **Follow the skill's workflow and checklist**
> If step 1 warns the index is stale, run `npx gitnexus analyze` in the terminal first.
## Skills
| Task | Skill to read |
| -------------------------------------------- | ------------------- |
| Understand architecture / "How does X work?" | `gitnexus-exploring` |
| Blast radius / "What breaks if I change X?" | `gitnexus-impact-analysis` |
| Trace bugs / "Why is X failing?" | `gitnexus-debugging` |
| Rename / extract / split / refactor | `gitnexus-refactoring` |
| Tools, resources, schema reference | `gitnexus-guide` (this file) |
| Index, status, clean, wiki CLI commands | `gitnexus-cli` |
## Tools Reference
| Tool | What it gives you |
| ---------------- | ------------------------------------------------------------------------ |
| `query` | Process-grouped code intelligence — execution flows related to a concept |
| `context` | 360-degree symbol view — categorized refs, processes it participates in |
| `impact` | Symbol blast radius — what breaks at depth 1/2/3 with confidence |
| `detect_changes` | Git-diff impact — what do your current changes affect |
| `rename` | Multi-file coordinated rename with confidence-tagged edits |
| `cypher` | Raw graph queries (read `gitnexus://repo/{name}/schema` first) |
| `list_repos` | Discover indexed repos |
## Resources Reference
Lightweight reads (~100-500 tokens) for navigation:
| Resource | Content |
| ---------------------------------------------- | ----------------------------------------- |
| `gitnexus://repo/{name}/context` | Stats, staleness check |
| `gitnexus://repo/{name}/clusters` | All functional areas with cohesion scores |
| `gitnexus://repo/{name}/cluster/{clusterName}` | Area members |
| `gitnexus://repo/{name}/processes` | All execution flows |
| `gitnexus://repo/{name}/process/{processName}` | Step-by-step trace |
| `gitnexus://repo/{name}/schema` | Graph schema for Cypher |
## Graph Schema
**Nodes:** File, Function, Class, Interface, Method, Community, Process
**Edges (via CodeRelation.type):** CALLS, IMPORTS, EXTENDS, IMPLEMENTS, DEFINES, MEMBER_OF, STEP_IN_PROCESS
```cypher
MATCH (caller)-[:CodeRelation {type: 'CALLS'}]->(f:Function {name: "myFunc"})
RETURN caller.name, caller.filePath
```
@@ -0,0 +1,97 @@
---
name: gitnexus-impact-analysis
description: "Use when the user wants to know what will break if they change something, or needs safety analysis before editing code. Examples: \"Is it safe to change X?\", \"What depends on this?\", \"What will break?\""
---
# Impact Analysis with GitNexus
## When to Use
- "Is it safe to change this function?"
- "What will break if I modify X?"
- "Show me the blast radius"
- "Who uses this code?"
- Before making non-trivial code changes
- Before committing — to understand what your changes affect
## Workflow
```
1. gitnexus_impact({target: "X", direction: "upstream"}) → What depends on this
2. READ gitnexus://repo/{name}/processes → Check affected execution flows
3. gitnexus_detect_changes() → Map current git changes to affected flows
4. Assess risk and report to user
```
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
## Checklist
```
- [ ] gitnexus_impact({target, direction: "upstream"}) to find dependents
- [ ] Review d=1 items first (these WILL BREAK)
- [ ] Check high-confidence (>0.8) dependencies
- [ ] READ processes to check affected execution flows
- [ ] gitnexus_detect_changes() for pre-commit check
- [ ] Assess risk level and report to user
```
## Understanding Output
| Depth | Risk Level | Meaning |
| ----- | ---------------- | ------------------------ |
| d=1 | **WILL BREAK** | Direct callers/importers |
| d=2 | LIKELY AFFECTED | Indirect dependencies |
| d=3 | MAY NEED TESTING | Transitive effects |
## Risk Assessment
| Affected | Risk |
| ------------------------------ | -------- |
| <5 symbols, few processes | LOW |
| 5-15 symbols, 2-5 processes | MEDIUM |
| >15 symbols or many processes | HIGH |
| Critical path (auth, payments) | CRITICAL |
## Tools
**gitnexus_impact** — the primary tool for symbol blast radius:
```
gitnexus_impact({
target: "validateUser",
direction: "upstream",
minConfidence: 0.8,
maxDepth: 3
})
→ d=1 (WILL BREAK):
- loginHandler (src/auth/login.ts:42) [CALLS, 100%]
- apiMiddleware (src/api/middleware.ts:15) [CALLS, 100%]
→ d=2 (LIKELY AFFECTED):
- authRouter (src/routes/auth.ts:22) [CALLS, 95%]
```
**gitnexus_detect_changes** — git-diff based impact analysis:
```
gitnexus_detect_changes({scope: "staged"})
→ Changed: 5 symbols in 3 files
→ Affected: LoginFlow, TokenRefresh, APIMiddlewarePipeline
→ Risk: MEDIUM
```
## Example: "What breaks if I change validateUser?"
```
1. gitnexus_impact({target: "validateUser", direction: "upstream"})
→ d=1: loginHandler, apiMiddleware (WILL BREAK)
→ d=2: authRouter, sessionManager (LIKELY AFFECTED)
2. READ gitnexus://repo/my-app/processes
→ LoginFlow and TokenRefresh touch validateUser
3. Risk: 2 direct callers, 2 processes = MEDIUM
```
@@ -0,0 +1,121 @@
---
name: gitnexus-refactoring
description: "Use when the user wants to rename, extract, split, move, or restructure code safely. Examples: \"Rename this function\", \"Extract this into a module\", \"Refactor this class\", \"Move this to a separate file\""
---
# Refactoring with GitNexus
## When to Use
- "Rename this function safely"
- "Extract this into a module"
- "Split this service"
- "Move this to a new file"
- Any task involving renaming, extracting, splitting, or restructuring code
## Workflow
```
1. gitnexus_impact({target: "X", direction: "upstream"}) → Map all dependents
2. gitnexus_query({query: "X"}) → Find execution flows involving X
3. gitnexus_context({name: "X"}) → See all incoming/outgoing refs
4. Plan update order: interfaces → implementations → callers → tests
```
> If "Index is stale" → run `npx gitnexus analyze` in terminal.
## Checklists
### Rename Symbol
```
- [ ] gitnexus_rename({symbol_name: "oldName", new_name: "newName", dry_run: true}) — preview all edits
- [ ] Review graph edits (high confidence) and ast_search edits (review carefully)
- [ ] If satisfied: gitnexus_rename({..., dry_run: false}) — apply edits
- [ ] gitnexus_detect_changes() — verify only expected files changed
- [ ] Run tests for affected processes
```
### Extract Module
```
- [ ] gitnexus_context({name: target}) — see all incoming/outgoing refs
- [ ] gitnexus_impact({target, direction: "upstream"}) — find all external callers
- [ ] Define new module interface
- [ ] Extract code, update imports
- [ ] gitnexus_detect_changes() — verify affected scope
- [ ] Run tests for affected processes
```
### Split Function/Service
```
- [ ] gitnexus_context({name: target}) — understand all callees
- [ ] Group callees by responsibility
- [ ] gitnexus_impact({target, direction: "upstream"}) — map callers to update
- [ ] Create new functions/services
- [ ] Update callers
- [ ] gitnexus_detect_changes() — verify affected scope
- [ ] Run tests for affected processes
```
## Tools
**gitnexus_rename** — automated multi-file rename:
```
gitnexus_rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: true})
→ 12 edits across 8 files
→ 10 graph edits (high confidence), 2 ast_search edits (review)
→ Changes: [{file_path, edits: [{line, old_text, new_text, confidence}]}]
```
**gitnexus_impact** — map all dependents first:
```
gitnexus_impact({target: "validateUser", direction: "upstream"})
→ d=1: loginHandler, apiMiddleware, testUtils
→ Affected Processes: LoginFlow, TokenRefresh
```
**gitnexus_detect_changes** — verify your changes after refactoring:
```
gitnexus_detect_changes({scope: "all"})
→ Changed: 8 files, 12 symbols
→ Affected processes: LoginFlow, TokenRefresh
→ Risk: MEDIUM
```
**gitnexus_cypher** — custom reference queries:
```cypher
MATCH (caller)-[:CodeRelation {type: 'CALLS'}]->(f:Function {name: "validateUser"})
RETURN caller.name, caller.filePath ORDER BY caller.filePath
```
## Risk Rules
| Risk Factor | Mitigation |
| ------------------- | ----------------------------------------- |
| Many callers (>5) | Use gitnexus_rename for automated updates |
| Cross-area refs | Use detect_changes after to verify scope |
| String/dynamic refs | gitnexus_query to find them |
| External/public API | Version and deprecate properly |
## Example: Rename `validateUser` to `authenticateUser`
```
1. gitnexus_rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: true})
→ 12 edits: 10 graph (safe), 2 ast_search (review)
→ Files: validator.ts, login.ts, middleware.ts, config.json...
2. Review ast_search edits (config.json: dynamic reference!)
3. gitnexus_rename({symbol_name: "validateUser", new_name: "authenticateUser", dry_run: false})
→ Applied 12 edits across 8 files
4. gitnexus_detect_changes({scope: "all"})
→ Affected: LoginFlow, TokenRefresh
→ Risk: MEDIUM — run tests for these flows
```
@@ -0,0 +1,47 @@
# ADR Format
ADRs live in `docs/adr/` and use sequential numbering: `0001-slug.md`, `0002-slug.md`, etc.
Create the `docs/adr/` directory lazily — only when the first ADR is needed.
## Template
```md
# {Short title of the decision}
{1-3 sentences: what's the context, what did we decide, and why.}
```
That's it. An ADR can be a single paragraph. The value is in recording *that* a decision was made and *why* — not in filling out sections.
## Optional sections
Only include these when they add genuine value. Most ADRs won't need them.
- **Status** frontmatter (`proposed | accepted | deprecated | superseded by ADR-NNNN`) — useful when decisions are revisited
- **Considered Options** — only when the rejected alternatives are worth remembering
- **Consequences** — only when non-obvious downstream effects need to be called out
## Numbering
Scan `docs/adr/` for the highest existing number and increment by one.
## When to offer an ADR
All three of these must be true:
1. **Hard to reverse** — the cost of changing your mind later is meaningful
2. **Surprising without context** — a future reader will look at the code and wonder "why on earth did they do it this way?"
3. **The result of a real trade-off** — there were genuine alternatives and you picked one for specific reasons
If a decision is easy to reverse, skip it — you'll just reverse it. If it's not surprising, nobody will wonder why. If there was no real alternative, there's nothing to record beyond "we did the obvious thing."
### What qualifies
- **Architectural shape.** "We're using a monorepo." "The write model is event-sourced, the read model is projected into Postgres."
- **Integration patterns between contexts.** "Ordering and Billing communicate via domain events, not synchronous HTTP."
- **Technology choices that carry lock-in.** Database, message bus, auth provider, deployment target. Not every library — just the ones that would take a quarter to swap out.
- **Boundary and scope decisions.** "Customer data is owned by the Customer context; other contexts reference it by ID only." The explicit no-s are as valuable as the yes-s.
- **Deliberate deviations from the obvious path.** "We're using manual SQL instead of an ORM because X." Anything where a reasonable reader would assume the opposite. These stop the next engineer from "fixing" something that was deliberate.
- **Constraints not visible in the code.** "We can't use AWS because of compliance requirements." "Response times must be under 200ms because of the partner API contract."
- **Rejected alternatives when the rejection is non-obvious.** If you considered GraphQL and picked REST for subtle reasons, record it — otherwise someone will suggest GraphQL again in six months.
@@ -0,0 +1,77 @@
# CONTEXT.md Format
## Structure
```md
# {Context Name}
{One or two sentence description of what this context is and why it exists.}
## Language
**Order**:
{A concise description of the term}
_Avoid_: Purchase, transaction
**Invoice**:
A request for payment sent to a customer after delivery.
_Avoid_: Bill, payment request
**Customer**:
A person or organization that places orders.
_Avoid_: Client, buyer, account
## Relationships
- An **Order** produces one or more **Invoices**
- An **Invoice** belongs to exactly one **Customer**
## Example dialogue
> **Dev:** "When a **Customer** places an **Order**, do we create the **Invoice** immediately?"
> **Domain expert:** "No — an **Invoice** is only generated once a **Fulfillment** is confirmed."
## Flagged ambiguities
- "account" was used to mean both **Customer** and **User** — resolved: these are distinct concepts.
```
## Rules
- **Be opinionated.** When multiple words exist for the same concept, pick the best one and list the others as aliases to avoid.
- **Flag conflicts explicitly.** If a term is used ambiguously, call it out in "Flagged ambiguities" with a clear resolution.
- **Keep definitions tight.** One sentence max. Define what it IS, not what it does.
- **Show relationships.** Use bold term names and express cardinality where obvious.
- **Only include terms specific to this project's context.** General programming concepts (timeouts, error types, utility patterns) don't belong even if the project uses them extensively. Before adding a term, ask: is this a concept unique to this context, or a general programming concept? Only the former belongs.
- **Group terms under subheadings** when natural clusters emerge. If all terms belong to a single cohesive area, a flat list is fine.
- **Write an example dialogue.** A conversation between a dev and a domain expert that demonstrates how the terms interact naturally and clarifies boundaries between related concepts.
## Single vs multi-context repos
**Single context (most repos):** One `CONTEXT.md` at the repo root.
**Multiple contexts:** A `CONTEXT-MAP.md` at the repo root lists the contexts, where they live, and how they relate to each other:
```md
# Context Map
## Contexts
- [Ordering](./src/ordering/CONTEXT.md) — receives and tracks customer orders
- [Billing](./src/billing/CONTEXT.md) — generates invoices and processes payments
- [Fulfillment](./src/fulfillment/CONTEXT.md) — manages warehouse picking and shipping
## Relationships
- **Ordering → Fulfillment**: Ordering emits `OrderPlaced` events; Fulfillment consumes them to start picking
- **Fulfillment → Billing**: Fulfillment emits `ShipmentDispatched` events; Billing consumes them to generate invoices
- **Ordering ↔ Billing**: Shared types for `CustomerId` and `Money`
```
The skill infers which structure applies:
- If `CONTEXT-MAP.md` exists, read it to find contexts
- If only a root `CONTEXT.md` exists, single context
- If neither exists, create a root `CONTEXT.md` lazily when the first term is resolved
When multiple contexts exist, infer which one the current topic relates to. If unclear, ask.
+88
View File
@@ -0,0 +1,88 @@
---
name: grill-with-docs
description: Grilling session that challenges your plan against the existing domain model, sharpens terminology, and updates documentation (CONTEXT.md, ADRs) inline as decisions crystallise. Use when user wants to stress-test a plan against their project's language and documented decisions.
---
<what-to-do>
Interview me relentlessly about every aspect of this plan until we reach a shared understanding. Walk down each branch of the design tree, resolving dependencies between decisions one-by-one. For each question, provide your recommended answer.
Ask the questions one at a time, waiting for feedback on each question before continuing.
If a question can be answered by exploring the codebase, explore the codebase instead.
</what-to-do>
<supporting-info>
## Domain awareness
During codebase exploration, also look for existing documentation:
### File structure
Most repos have a single context:
```
/
├── CONTEXT.md
├── docs/
│ └── adr/
│ ├── 0001-event-sourced-orders.md
│ └── 0002-postgres-for-write-model.md
└── src/
```
If a `CONTEXT-MAP.md` exists at the root, the repo has multiple contexts. The map points to where each one lives:
```
/
├── CONTEXT-MAP.md
├── docs/
│ └── adr/ ← system-wide decisions
├── src/
│ ├── ordering/
│ │ ├── CONTEXT.md
│ │ └── docs/adr/ ← context-specific decisions
│ └── billing/
│ ├── CONTEXT.md
│ └── docs/adr/
```
Create files lazily — only when you have something to write. If no `CONTEXT.md` exists, create one when the first term is resolved. If no `docs/adr/` exists, create it when the first ADR is needed.
## During the session
### Challenge against the glossary
When the user uses a term that conflicts with the existing language in `CONTEXT.md`, call it out immediately. "Your glossary defines 'cancellation' as X, but you seem to mean Y — which is it?"
### Sharpen fuzzy language
When the user uses vague or overloaded terms, propose a precise canonical term. "You're saying 'account' — do you mean the Customer or the User? Those are different things."
### Discuss concrete scenarios
When domain relationships are being discussed, stress-test them with specific scenarios. Invent scenarios that probe edge cases and force the user to be precise about the boundaries between concepts.
### Cross-reference with code
When the user states how something works, check whether the code agrees. If you find a contradiction, surface it: "Your code cancels entire Orders, but you just said partial cancellation is possible — which is right?"
### Update CONTEXT.md inline
When a term is resolved, update `CONTEXT.md` right there. Don't batch these up — capture them as they happen. Use the format in [CONTEXT-FORMAT.md](./CONTEXT-FORMAT.md).
`CONTEXT.md` should be totally devoid of implementation details. Do not treat `CONTEXT.md` as a spec, a scratch pad, or a repository for implementation decisions. It is a glossary and nothing else.
### Offer ADRs sparingly
Only offer to create an ADR when all three are true:
1. **Hard to reverse** — the cost of changing your mind later is meaningful
2. **Surprising without context** — a future reader will wonder "why did they do it this way?"
3. **The result of a real trade-off** — there were genuine alternatives and you picked one for specific reasons
If any of the three is missing, skip the ADR. Use the format in [ADR-FORMAT.md](./ADR-FORMAT.md).
</supporting-info>
+16
View File
@@ -0,0 +1,16 @@
---
name: handoff
description: Compact the current conversation into a handoff document for another agent to pick up.
argument-hint: "What will the next session be used for?"
disable-model-invocation: true
---
Write a handoff document summarising the current conversation so a fresh agent can continue the work. Save to the temporary directory of the user's OS - not the current workspace.
Include a "suggested skills" section in the document, which suggests skills that the agent should invoke.
Do not duplicate content already captured in other artifacts (PRDs, plans, ADRs, issues, commits, diffs). Reference them by path or URL instead.
Redact any sensitive information, such as API keys, passwords, or personally identifiable information.
If the user passed arguments, treat them as a description of what the next session will focus on and tailor the doc accordingly.
@@ -0,0 +1,156 @@
---
name: openspec-apply-change
description: Implement tasks from an OpenSpec change. Use when the user wants to start implementing, continue implementation, or work through tasks.
license: MIT
compatibility: Requires openspec CLI.
metadata:
author: openspec
version: "1.0"
generatedBy: "1.3.1"
---
Implement tasks from an OpenSpec change.
**Input**: Optionally specify a change name. If omitted, check if it can be inferred from conversation context. If vague or ambiguous you MUST prompt for available changes.
**Steps**
1. **Select the change**
If a name is provided, use it. Otherwise:
- Infer from conversation context if the user mentioned a change
- Auto-select if only one active change exists
- If ambiguous, run `openspec list --json` to get available changes and use the **AskUserQuestion tool** to let the user select
Always announce: "Using change: <name>" and how to override (e.g., `/opsx:apply <other>`).
2. **Check status to understand the schema**
```bash
openspec status --change "<name>" --json
```
Parse the JSON to understand:
- `schemaName`: The workflow being used (e.g., "spec-driven")
- Which artifact contains the tasks (typically "tasks" for spec-driven, check status for others)
3. **Get apply instructions**
```bash
openspec instructions apply --change "<name>" --json
```
This returns:
- `contextFiles`: artifact ID -> array of concrete file paths (varies by schema - could be proposal/specs/design/tasks or spec/tests/implementation/docs)
- Progress (total, complete, remaining)
- Task list with status
- Dynamic instruction based on current state
**Handle states:**
- If `state: "blocked"` (missing artifacts): show message, suggest using openspec-continue-change
- If `state: "all_done"`: congratulate, suggest archive
- Otherwise: proceed to implementation
4. **Read context files**
Read every file path listed under `contextFiles` from the apply instructions output.
The files depend on the schema being used:
- **spec-driven**: proposal, specs, design, tasks
- Other schemas: follow the contextFiles from CLI output
5. **Show current progress**
Display:
- Schema being used
- Progress: "N/M tasks complete"
- Remaining tasks overview
- Dynamic instruction from CLI
6. **Implement tasks (loop until done or blocked)**
For each pending task:
- Show which task is being worked on
- Make the code changes required
- Keep changes minimal and focused
- Mark task complete in the tasks file: `- [ ]` → `- [x]`
- Continue to next task
**Pause if:**
- Task is unclear → ask for clarification
- Implementation reveals a design issue → suggest updating artifacts
- Error or blocker encountered → report and wait for guidance
- User interrupts
7. **On completion or pause, show status**
Display:
- Tasks completed this session
- Overall progress: "N/M tasks complete"
- If all done: suggest archive
- If paused: explain why and wait for guidance
**Output During Implementation**
```
## Implementing: <change-name> (schema: <schema-name>)
Working on task 3/7: <task description>
[...implementation happening...]
✓ Task complete
Working on task 4/7: <task description>
[...implementation happening...]
✓ Task complete
```
**Output On Completion**
```
## Implementation Complete
**Change:** <change-name>
**Schema:** <schema-name>
**Progress:** 7/7 tasks complete ✓
### Completed This Session
- [x] Task 1
- [x] Task 2
...
All tasks complete! Ready to archive this change.
```
**Output On Pause (Issue Encountered)**
```
## Implementation Paused
**Change:** <change-name>
**Schema:** <schema-name>
**Progress:** 4/7 tasks complete
### Issue Encountered
<description of the issue>
**Options:**
1. <option 1>
2. <option 2>
3. Other approach
What would you like to do?
```
**Guardrails**
- Keep going through tasks until done or blocked
- Always read context files before starting (from the apply instructions output)
- If task is ambiguous, pause and ask before implementing
- If implementation reveals issues, pause and suggest artifact updates
- Keep code changes minimal and scoped to each task
- Update task checkbox immediately after completing each task
- Pause on errors, blockers, or unclear requirements - don't guess
- Use contextFiles from CLI output, don't assume specific file names
**Fluid Workflow Integration**
This skill supports the "actions on a change" model:
- **Can be invoked anytime**: Before all artifacts are done (if tasks exist), after partial implementation, interleaved with other actions
- **Allows artifact updates**: If implementation reveals design issues, suggest updating artifacts - not phase-locked, work fluidly
@@ -0,0 +1,114 @@
---
name: openspec-archive-change
description: Archive a completed change in the experimental workflow. Use when the user wants to finalize and archive a change after implementation is complete.
license: MIT
compatibility: Requires openspec CLI.
metadata:
author: openspec
version: "1.0"
generatedBy: "1.3.1"
---
Archive a completed change in the experimental workflow.
**Input**: Optionally specify a change name. If omitted, check if it can be inferred from conversation context. If vague or ambiguous you MUST prompt for available changes.
**Steps**
1. **If no change name provided, prompt for selection**
Run `openspec list --json` to get available changes. Use the **AskUserQuestion tool** to let the user select.
Show only active changes (not already archived).
Include the schema used for each change if available.
**IMPORTANT**: Do NOT guess or auto-select a change. Always let the user choose.
2. **Check artifact completion status**
Run `openspec status --change "<name>" --json` to check artifact completion.
Parse the JSON to understand:
- `schemaName`: The workflow being used
- `artifacts`: List of artifacts with their status (`done` or other)
**If any artifacts are not `done`:**
- Display warning listing incomplete artifacts
- Use **AskUserQuestion tool** to confirm user wants to proceed
- Proceed if user confirms
3. **Check task completion status**
Read the tasks file (typically `tasks.md`) to check for incomplete tasks.
Count tasks marked with `- [ ]` (incomplete) vs `- [x]` (complete).
**If incomplete tasks found:**
- Display warning showing count of incomplete tasks
- Use **AskUserQuestion tool** to confirm user wants to proceed
- Proceed if user confirms
**If no tasks file exists:** Proceed without task-related warning.
4. **Assess delta spec sync state**
Check for delta specs at `openspec/changes/<name>/specs/`. If none exist, proceed without sync prompt.
**If delta specs exist:**
- Compare each delta spec with its corresponding main spec at `openspec/specs/<capability>/spec.md`
- Determine what changes would be applied (adds, modifications, removals, renames)
- Show a combined summary before prompting
**Prompt options:**
- If changes needed: "Sync now (recommended)", "Archive without syncing"
- If already synced: "Archive now", "Sync anyway", "Cancel"
If user chooses sync, use Task tool (subagent_type: "general-purpose", prompt: "Use Skill tool to invoke openspec-sync-specs for change '<name>'. Delta spec analysis: <include the analyzed delta spec summary>"). Proceed to archive regardless of choice.
5. **Perform the archive**
Create the archive directory if it doesn't exist:
```bash
mkdir -p openspec/changes/archive
```
Generate target name using current date: `YYYY-MM-DD-<change-name>`
**Check if target already exists:**
- If yes: Fail with error, suggest renaming existing archive or using different date
- If no: Move the change directory to archive
```bash
mv openspec/changes/<name> openspec/changes/archive/YYYY-MM-DD-<name>
```
6. **Display summary**
Show archive completion summary including:
- Change name
- Schema that was used
- Archive location
- Whether specs were synced (if applicable)
- Note about any warnings (incomplete artifacts/tasks)
**Output On Success**
```
## Archive Complete
**Change:** <change-name>
**Schema:** <schema-name>
**Archived to:** openspec/changes/archive/YYYY-MM-DD-<name>/
**Specs:** ✓ Synced to main specs (or "No delta specs" or "Sync skipped")
All artifacts complete. All tasks complete.
```
**Guardrails**
- Always prompt for change selection if not provided
- Use artifact graph (openspec status --json) for completion checking
- Don't block archive on warnings - just inform and confirm
- Preserve .openspec.yaml when moving to archive (it moves with the directory)
- Show clear summary of what happened
- If sync is requested, use openspec-sync-specs approach (agent-driven)
- If delta specs exist, always run the sync assessment and show the combined summary before prompting
+288
View File
@@ -0,0 +1,288 @@
---
name: openspec-explore
description: Enter explore mode - a thinking partner for exploring ideas, investigating problems, and clarifying requirements. Use when the user wants to think through something before or during a change.
license: MIT
compatibility: Requires openspec CLI.
metadata:
author: openspec
version: "1.0"
generatedBy: "1.3.1"
---
Enter explore mode. Think deeply. Visualize freely. Follow the conversation wherever it goes.
**IMPORTANT: Explore mode is for thinking, not implementing.** You may read files, search code, and investigate the codebase, but you must NEVER write code or implement features. If the user asks you to implement something, remind them to exit explore mode first and create a change proposal. You MAY create OpenSpec artifacts (proposals, designs, specs) if the user asks—that's capturing thinking, not implementing.
**This is a stance, not a workflow.** There are no fixed steps, no required sequence, no mandatory outputs. You're a thinking partner helping the user explore.
---
## The Stance
- **Curious, not prescriptive** - Ask questions that emerge naturally, don't follow a script
- **Open threads, not interrogations** - Surface multiple interesting directions and let the user follow what resonates. Don't funnel them through a single path of questions.
- **Visual** - Use ASCII diagrams liberally when they'd help clarify thinking
- **Adaptive** - Follow interesting threads, pivot when new information emerges
- **Patient** - Don't rush to conclusions, let the shape of the problem emerge
- **Grounded** - Explore the actual codebase when relevant, don't just theorize
---
## What You Might Do
Depending on what the user brings, you might:
**Explore the problem space**
- Ask clarifying questions that emerge from what they said
- Challenge assumptions
- Reframe the problem
- Find analogies
**Investigate the codebase**
- Map existing architecture relevant to the discussion
- Find integration points
- Identify patterns already in use
- Surface hidden complexity
**Compare options**
- Brainstorm multiple approaches
- Build comparison tables
- Sketch tradeoffs
- Recommend a path (if asked)
**Visualize**
```
┌─────────────────────────────────────────┐
│ Use ASCII diagrams liberally │
├─────────────────────────────────────────┤
│ │
│ ┌────────┐ ┌────────┐ │
│ │ State │────────▶│ State │ │
│ │ A │ │ B │ │
│ └────────┘ └────────┘ │
│ │
│ System diagrams, state machines, │
│ data flows, architecture sketches, │
│ dependency graphs, comparison tables │
│ │
└─────────────────────────────────────────┘
```
**Surface risks and unknowns**
- Identify what could go wrong
- Find gaps in understanding
- Suggest spikes or investigations
---
## OpenSpec Awareness
You have full context of the OpenSpec system. Use it naturally, don't force it.
### Check for context
At the start, quickly check what exists:
```bash
openspec list --json
```
This tells you:
- If there are active changes
- Their names, schemas, and status
- What the user might be working on
### When no change exists
Think freely. When insights crystallize, you might offer:
- "This feels solid enough to start a change. Want me to create a proposal?"
- Or keep exploring - no pressure to formalize
### When a change exists
If the user mentions a change or you detect one is relevant:
1. **Read existing artifacts for context**
- `openspec/changes/<name>/proposal.md`
- `openspec/changes/<name>/design.md`
- `openspec/changes/<name>/tasks.md`
- etc.
2. **Reference them naturally in conversation**
- "Your design mentions using Redis, but we just realized SQLite fits better..."
- "The proposal scopes this to premium users, but we're now thinking everyone..."
3. **Offer to capture when decisions are made**
| Insight Type | Where to Capture |
|----------------------------|--------------------------------|
| New requirement discovered | `specs/<capability>/spec.md` |
| Requirement changed | `specs/<capability>/spec.md` |
| Design decision made | `design.md` |
| Scope changed | `proposal.md` |
| New work identified | `tasks.md` |
| Assumption invalidated | Relevant artifact |
Example offers:
- "That's a design decision. Capture it in design.md?"
- "This is a new requirement. Add it to specs?"
- "This changes scope. Update the proposal?"
4. **The user decides** - Offer and move on. Don't pressure. Don't auto-capture.
---
## What You Don't Have To Do
- Follow a script
- Ask the same questions every time
- Produce a specific artifact
- Reach a conclusion
- Stay on topic if a tangent is valuable
- Be brief (this is thinking time)
---
## Handling Different Entry Points
**User brings a vague idea:**
```
User: I'm thinking about adding real-time collaboration
You: Real-time collab is a big space. Let me think about this...
COLLABORATION SPECTRUM
════════════════════════════════════════════
Awareness Coordination Sync
│ │ │
▼ ▼ ▼
┌────────┐ ┌────────┐ ┌────────┐
│Presence│ │Cursors │ │ CRDT │
│ "3 │ │ Multi │ │Conflict│
│online" │ │ select │ │ free │
└────────┘ └────────┘ └────────┘
│ │ │
trivial moderate complex
Where's your head at?
```
**User brings a specific problem:**
```
User: The auth system is a mess
You: [reads codebase]
┌─────────────────────────────────────────────┐
│ CURRENT AUTH FLOW │
└─────────────────────────────────────────────┘
│
┌─────────────┼─────────────┐
▼ ▼ ▼
┌─────────┐ ┌─────────┐ ┌─────────┐
│ Google │ │ GitHub │ │ Email │
│ OAuth │ │ OAuth │ │ Magic │
└────┬────┘ └────┬────┘ └────┬────┘
│ │ │
└─────────────┼─────────────┘
▼
┌───────────┐
│ Session │
└─────┬─────┘
│
▼
┌───────────┐
│ Perms │
└───────────┘
I see three tangles. Which one's burning?
```
**User is stuck mid-implementation:**
```
User: /opsx:explore add-auth-system
The OAuth integration is more complex than expected
You: [reads change artifacts]
You're on task 4: "Implement OAuth flow"
Let me trace what's involved...
[draws diagram, explores options, suggests paths]
Want to update the design to reflect this?
Or add a spike task to investigate?
```
**User wants to compare options:**
```
User: Should we use Postgres or SQLite?
You: Generic answer is boring. What's the context?
User: A CLI tool that tracks local dev environments
You: That changes everything.
┌─────────────────────────────────────────────────┐
│ CLI TOOL DATA STORAGE │
└─────────────────────────────────────────────────┘
Key constraints:
• No daemon running
• Must work offline
• Single user
SQLite Postgres
Deployment embedded ✓ needs server ✗
Offline yes ✓ no ✗
Single file yes ✓ no ✗
SQLite. Not even close.
Unless... is there a sync component?
```
---
## Ending Discovery
There's no required ending. Discovery might:
- **Flow into a proposal**: "Ready to start? I can create a change proposal."
- **Result in artifact updates**: "Updated design.md with these decisions"
- **Just provide clarity**: User has what they need, moves on
- **Continue later**: "We can pick this up anytime"
When it feels like things are crystallizing, you might summarize:
```
## What We Figured Out
**The problem**: [crystallized understanding]
**The approach**: [if one emerged]
**Open questions**: [if any remain]
**Next steps** (if ready):
- Create a change proposal
- Keep exploring: just keep talking
```
But this summary is optional. Sometimes the thinking IS the value.
---
## Guardrails
- **Don't implement** - Never write code or implement features. Creating OpenSpec artifacts is fine, writing application code is not.
- **Don't fake understanding** - If something is unclear, dig deeper
- **Don't rush** - Discovery is thinking time, not task time
- **Don't force structure** - Let patterns emerge naturally
- **Don't auto-capture** - Offer to save insights, don't just do it
- **Do visualize** - A good diagram is worth many paragraphs
- **Do explore the codebase** - Ground discussions in reality
- **Do question assumptions** - Including the user's and your own
+110
View File
@@ -0,0 +1,110 @@
---
name: openspec-propose
description: Propose a new change with all artifacts generated in one step. Use when the user wants to quickly describe what they want to build and get a complete proposal with design, specs, and tasks ready for implementation.
license: MIT
compatibility: Requires openspec CLI.
metadata:
author: openspec
version: "1.0"
generatedBy: "1.3.1"
---
Propose a new change - create the change and generate all artifacts in one step.
I'll create a change with artifacts:
- proposal.md (what & why)
- design.md (how)
- tasks.md (implementation steps)
When ready to implement, run /opsx:apply
---
**Input**: The user's request should include a change name (kebab-case) OR a description of what they want to build.
**Steps**
1. **If no clear input provided, ask what they want to build**
Use the **AskUserQuestion tool** (open-ended, no preset options) to ask:
> "What change do you want to work on? Describe what you want to build or fix."
From their description, derive a kebab-case name (e.g., "add user authentication" → `add-user-auth`).
**IMPORTANT**: Do NOT proceed without understanding what the user wants to build.
2. **Create the change directory**
```bash
openspec new change "<name>"
```
This creates a scaffolded change at `openspec/changes/<name>/` with `.openspec.yaml`.
3. **Get the artifact build order**
```bash
openspec status --change "<name>" --json
```
Parse the JSON to get:
- `applyRequires`: array of artifact IDs needed before implementation (e.g., `["tasks"]`)
- `artifacts`: list of all artifacts with their status and dependencies
4. **Create artifacts in sequence until apply-ready**
Use the **TodoWrite tool** to track progress through the artifacts.
Loop through artifacts in dependency order (artifacts with no pending dependencies first):
a. **For each artifact that is `ready` (dependencies satisfied)**:
- Get instructions:
```bash
openspec instructions <artifact-id> --change "<name>" --json
```
- The instructions JSON includes:
- `context`: Project background (constraints for you - do NOT include in output)
- `rules`: Artifact-specific rules (constraints for you - do NOT include in output)
- `template`: The structure to use for your output file
- `instruction`: Schema-specific guidance for this artifact type
- `outputPath`: Where to write the artifact
- `dependencies`: Completed artifacts to read for context
- Read any completed dependency files for context
- Create the artifact file using `template` as the structure
- Apply `context` and `rules` as constraints - but do NOT copy them into the file
- Show brief progress: "Created <artifact-id>"
b. **Continue until all `applyRequires` artifacts are complete**
- After creating each artifact, re-run `openspec status --change "<name>" --json`
- Check if every artifact ID in `applyRequires` has `status: "done"` in the artifacts array
- Stop when all `applyRequires` artifacts are done
c. **If an artifact requires user input** (unclear context):
- Use **AskUserQuestion tool** to clarify
- Then continue with creation
5. **Show final status**
```bash
openspec status --change "<name>"
```
**Output**
After completing all artifacts, summarize:
- Change name and location
- List of artifacts created with brief descriptions
- What's ready: "All artifacts created! Ready for implementation."
- Prompt: "Run `/opsx:apply` or ask me to implement to start working on the tasks."
**Artifact Creation Guidelines**
- Follow the `instruction` field from `openspec instructions` for each artifact type
- The schema defines what each artifact should contain - follow it
- Read dependency artifacts for context before creating new ones
- Use `template` as the structure for your output file - fill in its sections
- **IMPORTANT**: `context` and `rules` are constraints for YOU, not content for the file
- Do NOT copy `<context>`, `<rules>`, `<project_context>` blocks into the artifact
- These guide what you write, but should never appear in the output
**Guardrails**
- Create ALL artifacts needed for implementation (as defined by schema's `apply.requires`)
- Always read dependency artifacts before creating a new one
- If context is critically unclear, ask the user - but prefer making reasonable decisions to keep momentum
- If a change with that name already exists, ask if user wants to continue it or create a new one
- Verify each artifact file exists after writing before proceeding to next
+109
View File
@@ -0,0 +1,109 @@
---
name: tdd
description: Test-driven development with red-green-refactor loop. Use when user wants to build features or fix bugs using TDD, mentions "red-green-refactor", wants integration tests, or asks for test-first development.
---
# Test-Driven Development
## Philosophy
**Core principle**: Tests should verify behavior through public interfaces, not implementation details. Code can change entirely; tests shouldn't.
**Good tests** are integration-style: they exercise real code paths through public APIs. They describe _what_ the system does, not _how_ it does it. A good test reads like a specification - "user can checkout with valid cart" tells you exactly what capability exists. These tests survive refactors because they don't care about internal structure.
**Bad tests** are coupled to implementation. They mock internal collaborators, test private methods, or verify through external means (like querying a database directly instead of using the interface). The warning sign: your test breaks when you refactor, but behavior hasn't changed. If you rename an internal function and tests fail, those tests were testing implementation, not behavior.
See [tests.md](tests.md) for examples and [mocking.md](mocking.md) for mocking guidelines.
## Anti-Pattern: Horizontal Slices
**DO NOT write all tests first, then all implementation.** This is "horizontal slicing" - treating RED as "write all tests" and GREEN as "write all code."
This produces **crap tests**:
- Tests written in bulk test _imagined_ behavior, not _actual_ behavior
- You end up testing the _shape_ of things (data structures, function signatures) rather than user-facing behavior
- Tests become insensitive to real changes - they pass when behavior breaks, fail when behavior is fine
- You outrun your headlights, committing to test structure before understanding the implementation
**Correct approach**: Vertical slices via tracer bullets. One test → one implementation → repeat. Each test responds to what you learned from the previous cycle. Because you just wrote the code, you know exactly what behavior matters and how to verify it.
```
WRONG (horizontal):
RED: test1, test2, test3, test4, test5
GREEN: impl1, impl2, impl3, impl4, impl5
RIGHT (vertical):
RED→GREEN: test1→impl1
RED→GREEN: test2→impl2
RED→GREEN: test3→impl3
...
```
## Workflow
### 1. Planning
When exploring the codebase, use the project's domain glossary so that test names and interface vocabulary match the project's language, and respect ADRs in the area you're touching.
Before writing any code:
- [ ] Confirm with user what interface changes are needed
- [ ] Confirm with user which behaviors to test (prioritize)
- [ ] Identify opportunities for [deep modules](deep-modules.md) (small interface, deep implementation)
- [ ] Design interfaces for [testability](interface-design.md)
- [ ] List the behaviors to test (not implementation steps)
- [ ] Get user approval on the plan
Ask: "What should the public interface look like? Which behaviors are most important to test?"
**You can't test everything.** Confirm with the user exactly which behaviors matter most. Focus testing effort on critical paths and complex logic, not every possible edge case.
### 2. Tracer Bullet
Write ONE test that confirms ONE thing about the system:
```
RED: Write test for first behavior → test fails
GREEN: Write minimal code to pass → test passes
```
This is your tracer bullet - proves the path works end-to-end.
### 3. Incremental Loop
For each remaining behavior:
```
RED: Write next test → fails
GREEN: Minimal code to pass → passes
```
Rules:
- One test at a time
- Only enough code to pass current test
- Don't anticipate future tests
- Keep tests focused on observable behavior
### 4. Refactor
After all tests pass, look for [refactor candidates](refactoring.md):
- [ ] Extract duplication
- [ ] Deepen modules (move complexity behind simple interfaces)
- [ ] Apply SOLID principles where natural
- [ ] Consider what new code reveals about existing code
- [ ] Run tests after each refactor step
**Never refactor while RED.** Get to GREEN first.
## Checklist Per Cycle
```
[ ] Test describes behavior, not implementation
[ ] Test uses public interface only
[ ] Test would survive internal refactor
[ ] Code is minimal for this test
[ ] No speculative features added
```
+33
View File
@@ -0,0 +1,33 @@
# Deep Modules
From "A Philosophy of Software Design":
**Deep module** = small interface + lots of implementation
```
┌─────────────────────┐
│ Small Interface │ ← Few methods, simple params
├─────────────────────┤
│ │
│ │
│ Deep Implementation│ ← Complex logic hidden
│ │
│ │
└─────────────────────┘
```
**Shallow module** = large interface + little implementation (avoid)
```
┌─────────────────────────────────┐
│ Large Interface │ ← Many methods, complex params
├─────────────────────────────────┤
│ Thin Implementation │ ← Just passes through
└─────────────────────────────────┘
```
When designing interfaces, ask:
- Can I reduce the number of methods?
- Can I simplify the parameters?
- Can I hide more complexity inside?
+31
View File
@@ -0,0 +1,31 @@
# Interface Design for Testability
Good interfaces make testing natural:
1. **Accept dependencies, don't create them**
```typescript
// Testable
function processOrder(order, paymentGateway) {}
// Hard to test
function processOrder(order) {
const gateway = new StripeGateway();
}
```
2. **Return results, don't produce side effects**
```typescript
// Testable
function calculateDiscount(cart): Discount {}
// Hard to test
function applyDiscount(cart): void {
cart.total -= discount;
}
```
3. **Small surface area**
- Fewer methods = fewer tests needed
- Fewer params = simpler test setup
+59
View File
@@ -0,0 +1,59 @@
# When to Mock
Mock at **system boundaries** only:
- External APIs (payment, email, etc.)
- Databases (sometimes - prefer test DB)
- Time/randomness
- File system (sometimes)
Don't mock:
- Your own classes/modules
- Internal collaborators
- Anything you control
## Designing for Mockability
At system boundaries, design interfaces that are easy to mock:
**1. Use dependency injection**
Pass external dependencies in rather than creating them internally:
```typescript
// Easy to mock
function processPayment(order, paymentClient) {
return paymentClient.charge(order.total);
}
// Hard to mock
function processPayment(order) {
const client = new StripeClient(process.env.STRIPE_KEY);
return client.charge(order.total);
}
```
**2. Prefer SDK-style interfaces over generic fetchers**
Create specific functions for each external operation instead of one generic function with conditional logic:
```typescript
// GOOD: Each function is independently mockable
const api = {
getUser: (id) => fetch(`/users/${id}`),
getOrders: (userId) => fetch(`/users/${userId}/orders`),
createOrder: (data) => fetch('/orders', { method: 'POST', body: data }),
};
// BAD: Mocking requires conditional logic inside the mock
const api = {
fetch: (endpoint, options) => fetch(endpoint, options),
};
```
The SDK approach means:
- Each mock returns one specific shape
- No conditional logic in test setup
- Easier to see which endpoints a test exercises
- Type safety per endpoint
+10
View File
@@ -0,0 +1,10 @@
# Refactor Candidates
After TDD cycle, look for:
- **Duplication** → Extract function/class
- **Long methods** → Break into private helpers (keep tests on public interface)
- **Shallow modules** → Combine or deepen
- **Feature envy** → Move logic to where data lives
- **Primitive obsession** → Introduce value objects
- **Existing code** the new code reveals as problematic
+61
View File
@@ -0,0 +1,61 @@
# Good and Bad Tests
## Good Tests
**Integration-style**: Test through real interfaces, not mocks of internal parts.
```typescript
// GOOD: Tests observable behavior
test("user can checkout with valid cart", async () => {
const cart = createCart();
cart.add(product);
const result = await checkout(cart, paymentMethod);
expect(result.status).toBe("confirmed");
});
```
Characteristics:
- Tests behavior users/callers care about
- Uses public API only
- Survives internal refactors
- Describes WHAT, not HOW
- One logical assertion per test
## Bad Tests
**Implementation-detail tests**: Coupled to internal structure.
```typescript
// BAD: Tests implementation details
test("checkout calls paymentService.process", async () => {
const mockPayment = jest.mock(paymentService);
await checkout(cart, payment);
expect(mockPayment.process).toHaveBeenCalledWith(cart.total);
});
```
Red flags:
- Mocking internal collaborators
- Testing private methods
- Asserting on call counts/order
- Test breaks when refactoring without behavior change
- Test name describes HOW not WHAT
- Verifying through external means instead of interface
```typescript
// BAD: Bypasses interface to verify
test("createUser saves to database", async () => {
await createUser({ name: "Alice" });
const row = await db.query("SELECT * FROM users WHERE name = ?", ["Alice"]);
expect(row).toBeDefined();
});
// GOOD: Verifies through interface
test("createUser makes user retrievable", async () => {
const user = await createUser({ name: "Alice" });
const retrieved = await getUser(user.id);
expect(retrieved.name).toBe("Alice");
});
```
+76
View File
@@ -0,0 +1,76 @@
---
name: to-prd
description: Turn the current conversation context into a PRD and publish it to the project issue tracker. Use when user wants to create a PRD from the current context.
---
This skill takes the current conversation context and codebase understanding and produces a PRD. Do NOT interview the user — just synthesize what you already know.
The issue tracker and triage label vocabulary should have been provided to you — run `/setup-matt-pocock-skills` if not.
## Process
1. Explore the repo to understand the current state of the codebase, if you haven't already. Use the project's domain glossary vocabulary throughout the PRD, and respect any ADRs in the area you're touching.
2. Sketch out the major modules you will need to build or modify to complete the implementation. Actively look for opportunities to extract deep modules that can be tested in isolation.
A deep module (as opposed to a shallow module) is one which encapsulates a lot of functionality in a simple, testable interface which rarely changes.
Check with the user that these modules match their expectations. Check with the user which modules they want tests written for.
3. Write the PRD using the template below, then publish it to the project issue tracker. Apply the `ready-for-agent` triage label - no need for additional triage.
<prd-template>
## Problem Statement
The problem that the user is facing, from the user's perspective.
## Solution
The solution to the problem, from the user's perspective.
## User Stories
A LONG, numbered list of user stories. Each user story should be in the format of:
1. As an <actor>, I want a <feature>, so that <benefit>
<user-story-example>
1. As a mobile bank customer, I want to see balance on my accounts, so that I can make better informed decisions about my spending
</user-story-example>
This list of user stories should be extremely extensive and cover all aspects of the feature.
## Implementation Decisions
A list of implementation decisions that were made. This can include:
- The modules that will be built/modified
- The interfaces of those modules that will be modified
- Technical clarifications from the developer
- Architectural decisions
- Schema changes
- API contracts
- Specific interactions
Do NOT include specific file paths or code snippets. They may end up being outdated very quickly.
Exception: if a prototype produced a snippet that encodes a decision more precisely than prose can (state machine, reducer, schema, type shape), inline it within the relevant decision and note briefly that it came from a prototype. Trim to the decision-rich parts — not a working demo, just the important bits.
## Testing Decisions
A list of testing decisions that were made. Include:
- A description of what makes a good test (only test external behavior, not implementation details)
- Which modules will be tested
- Prior art for the tests (i.e. similar types of tests in the codebase)
## Out of Scope
A description of the things that are out of scope for this PRD.
## Further Notes
Any further notes about the feature.
</prd-template>
+7
View File
@@ -0,0 +1,7 @@
---
name: zoom-out
description: Tell the agent to zoom out and give broader context or a higher-level perspective. Use when you're unfamiliar with a section of code or need to understand how it fits into the bigger picture.
disable-model-invocation: true
---
I don't know this area of code well. Go up a layer of abstraction. Give me a map of all the relevant modules and callers, using the project's domain glossary vocabulary.
@@ -148,7 +148,7 @@ openspec/changes/phase-1-infrastructure/
## 敏感信息(已编辑)
- MySQL 密码:已配置在 application.yml(`!Fucker123..`)
- MySQL 密码:已从仓库移除,使用环境变量注入
- Redis:无密码
---
+1 -1
View File
@@ -97,7 +97,7 @@ spring:
redis:
host: 119.29.78.52
port: 6379
password: '!Fucker123..'
password: ${SUPERBIZ_REDIS_PASSWORD}
database: 0
timeout: 3000
```
+2 -2
View File
@@ -140,7 +140,7 @@ Error Code: 1049
datasource:
url: jdbc:mysql://119.29.78.52:33306/superbiz_agent?...
username: root
password: '!Fucker123..'
password: ${SUPERBIZ_MYSQL_PASSWORD}
```
**Redis 配置**:
@@ -149,7 +149,7 @@ data:
redis:
host: 119.29.78.52
port: 6379
password: '!Fucker123..'
password: ${SUPERBIZ_REDIS_PASSWORD}
```
**Flyway 配置**:
+13
View File
@@ -0,0 +1,13 @@
root = true
[*]
charset = utf-8
end_of_line = crlf
insert_final_newline = true
trim_trailing_whitespace = true
[*.md]
trim_trailing_whitespace = false
[*.{java,xml,yml,yaml,properties,json,sql,txt,ps1}]
charset = utf-8
+8 -2
View File
@@ -44,12 +44,18 @@ build/
app.log
logs/
### Local Secrets ###
.env
.env.*
!.env.example
application-local.yml
application-*.local.yml
### Upload Files ###
uploads/
### Temp Scripts ###
*.sh
*.py
### docker
/volumes
@@ -59,8 +65,8 @@ uploads/
### Windows / Runtime Artifacts
*.stackdump
NUL
### MVP Demo Generated Outputs
mvp/demo/output/*.json
!mvp/demo/output/README.md
.pi/extensions/emdash-hook.ts
+108 -1
View File
@@ -1,7 +1,114 @@
# CLAUDE.md
## Defaults
- Reply in **Chinese** unless I explicitly ask for English.
- No emojis.
- Do not truncate important outputs (logs, diffs, stack traces, commands, or critical reasoning that affects
safety/correctness).
## Refactor policy (legacy code)
- When existing code is a "big ball of mud" (hard to maintain, clearly bad design,
full of hacks), prefer a **clean, full refactor** over stacking more patches
on top of it.
- A refactor may completely replace internal structure
(functions, modules, classes, data flow).
- By default, try to preserve externally observable behaviour.
If you intentionally change behaviour or protocols, you MUST:
- Call out clearly that this is a **behaviour/protocol change**.
- Explain why the change is necessary and which code paths/consumers are affected.
- Update or add tests to cover the new behaviour.
## Before touching code (mandatory)
Find reuse opportunities + Trace the call/dependency chain and impact radius:
- Use semantic code search first via `codebase-retrieval` tool.
- Confirm understanding with LSP: `goToDefinition`, `findReferences`.
- Use Grep/Glob for verifying and understanding additional code snippets.
## Red lines
- No copy-paste duplication.
- Do not break existing externally observable behaviour **unless**:
- It is part of a deliberate refactor as described in the refactor policy, and
- You clearly document the behavioural change and its impact.
- Do not proceed with a known-wrong approach.
- Critical paths must have explicit error handling.
- Never implement "blindly": always confirm understanding via code reading + references.
## Task sizing
- **Simple**
- Criteria — single file, clear requirement, < 20 lines changed,
clearly local impact.
- Handling — after doing the "Before touching code" steps
(research + impact analysis + internal three-question checklist),
you may execute directly with minimal explanation.
- A very short context line is enough;
a full breakdown of the checklist is not required.
- **Medium**
- Criteria — 2–5 files, or requires some research, or impact is not obviously local.
- Handling — write a short plan (bullet points) → then implement.
- Briefly surface the checklist result in the reply
(1–3 short lines describing real issue, key reuse, and main impact).
- **Complex**
- Criteria — architecture changes, multiple modules, high uncertainty or risk.
- Handling — follow this workflow:
1. **RESEARCH**: inspect code and facts only (no proposals yet).
2. **PLAN**: present options + tradeoffs + recommendation;
use `AskUserQuestion` actively to align with the user;
wait for user's confirmation.
3. **EXECUTE**: implement exactly the approved plan.
4. **REVIEW**: self-check (tests, edge cases, cleanup).
## Git
- Do not commit unless I explicitly ask.
- Do not push unless I explicitly ask.
- Before writing a commit message, glance at a few recent commits and match the repo's style:
- `git log -n 5 --oneline`
- If there is no obvious existing style, use this default format:
- `<type>(<scope>): <description>`
- Before any commit: run `git diff` and confirm the exact scope of changes.
- Never force-push to `main` / `master` unless the user approves.
- Do not add attribution lines in commit messages.
## Security
- Never hardcode secrets (keys/passwords/tokens).
- Never commit `.env` files or any credentials.
- Validate user input at trust boundaries (APIs, CLIs, external data sources).
## Quality & cleanup
- Prefer clarity and simplicity first (KISS); apply DRY to remove obvious
copy-paste duplication when it does not hurt readability.
- If you change a function signature, update **all** call sites.
- After changes:
- Remove temporary files.
- Remove dead/commented-out code.
- Remove unused imports.
- Remove debug logging that is no longer needed.
- Run the smallest meaningful verification (lint/test/build) for the parts you touched.
## Windows / PowerShell (if used)
- PowerShell does not support `&&`; use `;` to chain commands.
- Quote paths that contain spaces or non-ASCII characters.
## Baisc Infos
Unless directly relevant to the user's current question, you should avoid proactively mentioning, illustrating, or
trailing off into the following information in 99% of cases:
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **SuperBizAgent-java** (7988 symbols, 12713 relationships, 297 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **SuperBizAgent-java** (13483 symbols, 22230 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
+1 -2
View File
@@ -111,11 +111,10 @@ trailing off into the following information in 99% of cases:
- 文档目录结构:
- 不要将文档放到用户目录(如 `C:\Users\EDY\.claude\`)中
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **SuperBizAgent-java** (7988 symbols, 12713 relationships, 297 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
This project is indexed by GitNexus as **SuperBizAgent-java** (13483 symbols, 22230 relationships, 300 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
+151 -2
View File
@@ -72,9 +72,10 @@
### SessionContext
- 定义:会话上下文数据类,存储在 Redis 中的会话数据
- 包含字段:sessionId、userId、businessId、traceId、status、toolCalls、TTL
- 包含字段:sessionId、userId、businessId、traceId、status、toolCalls、messageHistory、TTL
- 序列化方式:JSON(GenericJackson2JsonRedisSerializer)
- 使用场景:多轮对话上下文管理、工具调用历史追踪
- 边界:messageHistory 是热路径对话历史缓存,用于下一轮 prompt 上下文;长期审计的问题和答案应落到 Diagnosis Run,而不是依赖 Redis TTL 内的上下文正文。
### ToolCall
- 定义:工具调用记录数据类,追踪 Agent 使用的工具及其结果
@@ -87,6 +88,21 @@
- 核心方法:createSession、getSession、updateSession、deleteSession、refreshSession、addToolCall
- 使用场景:分布式会话管理、Agent 状态维护
### Chat Session
- 定义:一次多轮对话上下文,由 `sessionId` 唯一标识。
- 使用场景:保存用户连续对话的上下文窗口、会话状态和最近活跃时间。
- 边界:Chat Session 不代表一次诊断执行;同一个 Chat Session 可以包含多次 Diagnosis Run。
### Diagnosis Run
- 定义:一次独立诊断执行,由 `runId` 唯一标识,属于一个 Chat Session。
- 使用场景:保存某一轮诊断的 query、answer、status、耗时、token、反馈和自评估结果。
- 边界:Diagnosis Run 是 Trace、Feedback 和 Evidence score 的绑定对象;多轮对话中的每次 `/api/chat` 或 `/api/ai_ops` 执行都应创建新的 Diagnosis Run。
### Diagnosis Trace
- 定义:一次 Diagnosis Run 的可回放执行轨迹,由 run 主记录、AgentStep 和 ToolInvocation 聚合形成。
- 使用场景:Trace API、Trace UI、Verifier 审计、评测 fixture 和人工排查。
- 边界:Diagnosis Trace 是聚合视图,不要求单独的 trace 主表;当前 trace 明细由 `agent_step` 和 `tool_invocation` 表承载。
### Flyway
- 定义:数据库版本迁移工具,管理 SQL 脚本的版本化执行
- 配置:spring.flyway.enabled=true, baseline-on-migrate=true
@@ -105,4 +121,137 @@
- 枚举类型在数据库中存储为 VARCHAR,JPA 使用 `@Enumerated(EnumType.STRING)` + `columnDefinition = "VARCHAR"`
- JPA ddl-auto 使用 `validate` 模式,表结构修改必须通过 Flyway 迁移脚本
- Redis 会话 TTL 由调用方指定,不同场景使用不同过期时间(短诊断 5 分钟,长会话 1 小时)
- Repository 查询方法遵循 Spring Data JPA 命名约定,复杂查询使用 `@Query`
- Repository 查询方法遵循 Spring Data JPA 命名约定,复杂查询使用 `@Query`
## Diagnosis Playbook Skills
### Diagnosis Playbook Skill
- 定义:项目内可版本化的诊断流程包,存放在 `src/main/resources/skills/{skill-name}/SKILL.md`。
- 使用场景:把高频故障诊断流程从大 prompt / 知识库文档中抽出,形成可审查、可复用、可按需加载的 playbook。
- 边界:skill 只定义排查 workflow、证据顺序、停止条件、低置信度行为和报告规则;事实性知识仍放在 `knowledge_base/`,事实证据仍来自 evidence tools。
### SkillRegistry
- 定义:Spring AI Alibaba Agent Framework 的 skill 元数据和正文读取入口。本项目使用 `ClasspathSkillRegistry` 从 classpath `skills/` 加载 skill。
- 使用场景:统一提供 skill `name` / `description` 元数据,并支撑 Executor 通过官方 `read_skill` 读取完整 `SKILL.md`。
- 当前约束:`SkillConfig.SingleSkillRegistry` 临时只暴露 active skill `diagnose-mysql-connection-pool`,用于验证单 skill 流程和避免一次性注入全部 skill。
### PlannerSkillMetadataHook
- 定义:项目本地 hook,只向 Planner 注入结构化 `skill_catalog` 元数据。
- 使用场景:Planner 根据 skill `name` / `description` 选择 `selected_skill`,输出 `selection_reason` 和执行计划。
- 边界:Planner 不暴露官方 `read_skill` 工具,不读取完整 `SKILL.md`;Planner 只能选择 skill,不能执行 skill。
### SkillsAgentHook
- 定义:Spring AI Alibaba 官方 skill hook,会同时注入官方 Skills System prompt,并暴露 `read_skill` 工具。
- 使用场景:只挂到 Executor 和 single-agent Chat;Executor 根据 `planner_plan.selected_skill` 读取完整 playbook 后再调用证据工具。
- 边界:不要挂到 Planner,否则 Planner 会获得 `read_skill` 工具并可能读取完整 skill;Verifier 也不能挂该 hook。
### read_skill
- 定义:官方 skill 读取工具,参数为 `skill_name`,返回对应 `SKILL.md` 正文。
- 使用场景:Executor 在执行场景化诊断前读取 Planner 选中的 playbook。
- 边界:`read_skill` 是流程指导工具,不是事实证据工具;不应作为诊断事实写入 `tool_invocation` 证据链。
### Evidence Tools
- 定义:产生可验证诊断事实的工具集合,包括 `lookup_knowledge`、`query_logs`、`query_metrics`、告警/Prometheus 工具等。
- 使用场景:Executor 按 skill workflow 调用 evidence tools 收集事实,`tool_invocation` 记录这些事实证据。
- 边界:最终诊断结论必须被 evidence tools 支撑,不能仅由 skill 正文支撑。
### Diagnosis Harness
- 定义:围绕 Diagnosis Agent 提供确定性运行控制的边界,负责 Run、预算、取消、重试装配、Tool 调用记录、证据验真和最终释放,不承担业务诊断推理。
- 边界:Harness 不是工作流引擎,不实现 Planner/Executor/Composer 节点或自行编写 ReAct 循环。
### Diagnosis Agent
- 定义:诊断链路中唯一拥有 ReAct 工具循环并生成 `DiagnosisDraft` 的 Agent,负责规划证据查询、判断证据充分性和撰写完整诊断草稿。
- 边界:不负责意图路由、Run/Session 生命周期、证据物理验真、独立语义审查或最终发布;证据不足时必须明确停止并保留限制。
### EvidenceGuard
- 定义:Harness 内部的确定性证据验真能力,校验 Draft 引用、当前 Run 所有权、Tool 调用状态和有界 Agent 投影。
- 边界:EvidenceGuard 不调用 LLM,也不判断证据是否足以推出业务结论。
### SemanticGuard
- 定义:使用隔离上下文对完整诊断 Draft 与已验真证据做报告级语义审查的单轮 Agent。
- 边界:无工具、无记忆、无 ReAct 循环,不访问 Redis,不生成或改写用户报告。
### Invocation Status
- 定义:Tool 调用及结果投影的生命周期状态,固定为 `PROJECTING`、`READY`、`ERROR`。
- 边界:它只说明调用记录是否完成,不说明结果是否包含证据。
### Durable Audit
- 定义:为 Diagnosis Trace 长期保存的 Run、Agent 模型步骤和 Tool 调用元数据,用于 exact sessionId/runId 回放、评测和运维核对。
- 边界:只保存有界、脱敏、可长期保留的身份、状态、耗时、预算和结果摘要;不保存 Prompt、Thought、完整 Tool 参数、raw response 或 Redis canonical invocation。
### Evidence Status
- 定义:证据 Tool 的结果语义,固定为 `EVIDENCE_FOUND`、`NO_EVIDENCE`、`ERROR`。
- 边界:`EVIDENCE_FOUND` 只表示存在候选内容,不保证内容能够支持当前诊断;`NO_EVIDENCE` 只表示当前查询范围内没有匹配结果,不能解释为问题不存在、根因被排除或系统健康。
### Information Gain
- 定义:一次 Tool 结果是否推进当前 Diagnosis Run 的语义评价,固定为 `GAINED` 或 `NO_GAIN`。
- 边界:它评价的是结果对当前诊断的作用,不评价 Tool 产品质量;`NO_EVIDENCE` 和重复的规范化 `tool + scope` 可由 Harness 机械标记为 `NO_GAIN`,其他成功非空结果(包括 RAG `REFERENCE`)由模型评价。
### Collection State
- 定义:Diagnosis Harness 对当前 Run 是否允许继续收集证据的控制状态,固定为 `COLLECTING` 或 `SATURATED`。
- 边界:状态由 Harness 维护;`SATURATED` 可因连续 `NO_GAIN` 或连续进展协议错误达到各自配置阈值而进入,不包括硬预算耗尽。模型可以请求新的 Tool 调用,但不能绕过 `SATURATED`。
### Diagnosis Stop Reason
- 定义:Harness 停止当前 Run 继续调用 Tool 的内部原因,首版区分 `INFORMATION_SATURATED`、`BUDGET_LIMIT_REACHED` 与 `PROGRESS_PROTOCOL_VIOLATED`。
- 边界:它用于控制、Trace 和 Release 输入,不是用户可见生命周期状态,也不进入模型上下文;真正的不可恢复技术故障走失败通道。协议错误不累计为 `NO_GAIN`,使用独立阈值与 stop reason。
### Progress Protocol Violation
- 定义:模型未遵守 Tool Call Envelope 进展协议时的安全错误分类,例如缺失/错序/意外 `previous_observation`、缺失 `input` 或非法 Envelope。
- 边界:返回可修正 observation(`repair_required`、`violation_type`、期望上一轮 Tool Call ID、允许的 `information_gain`);连续错误达到阈值后交付一次 `STOP_REQUIRED/PROGRESS_PROTOCOL_VIOLATED`。不泄露业务参数、上一轮观察正文、raw response 或内部异常。
### Progress Snapshot
- 定义:Tool Loop 结束时,从当前 Run 的 Canonical Tool Result 一次性投影出的有界发布视图,用于生成已检查范围和客观结果。
- 边界:Canonical Tool Result 是真理源;Progress Snapshot 不逐轮维护、不保存原始 Tool Response、Prompt 或内部 thought,也不直接进入模型上下文。
### Safe Fallback Type
- 定义:`SafeFallback.type` 对没有发布诊断结论的业务原因分类,例如 `INSUFFICIENT_EVIDENCE`、`MISSING_REQUIRED_CONTEXT`、`BUDGET_EXHAUSTED` 或安全校验失败。
- 边界:它是 `ReleaseOutcome.FALLBACK` 的原因字段,不是与 `SUCCESS / FALLBACK / FAILED / CANCELLED` 平行的第二套生命周期状态。
### Diagnosis Release Use Case
- 定义:诊断业务发布的唯一决策入口,接收 DiagnosisDraft 和/或 Harness `stop_reason + ProgressSnapshot`,生成安全的 `SUCCESS / FALLBACK` 结果。
- 边界:`conclusion=null` 不触发 EvidenceRepair;只有存在结论时才执行完整 EvidenceGuard、EvidenceRepair 和 SemanticGuard 链路。不可形成安全业务内容的技术故障由 Chat Application Use Case 映射为 `FAILED / CANCELLED`。
### Diagnosis Draft Contract Failure
- 定义:Diagnosis Agent 最终文本为空、不是严格 JSON,或不满足 `DiagnosisDraft` Schema 时产生的 Agent 输出合同失败。
- 边界:非法文本始终丢弃,不做 Markdown/自然语言提取,也不调用模型修复;仅当当前 Run 的 `ProgressSnapshot` 含已验真 observed facts 时,Release 才能确定性发布 `INSUFFICIENT_EVIDENCE`,否则保持 `FAILED`。它不是 `Diagnosis Stop Reason`,不得伪装成信息饱和或预算终止。
### Model Observation
- 定义:Tool 内部标准化结果经过白名单投影后,作为 Tool Response 进入 Diagnosis Agent 上下文的有界视图。
- 边界:只包含模型完成语义判断和证据引用所需的信息;预算、阈值、重复指纹、原始相似度、原始 Tool Response 和完整 Harness 控制状态不得进入该视图。
### RunContext
- 定义:一次 Diagnosis Run 的显式执行上下文,结构不可变地携带 `sessionId`、`runId`、deadline,以及该 Run 独占的取消、预算、重试策略和生命周期状态句柄。
- 边界:RunContext 通过方法参数或框架受控 context 显式传播,不依赖 ThreadLocal;结构不可变不等于内部计数和取消状态不能变化,这些变化由线程安全句柄管理。
### Run Lifecycle
- 定义:Diagnosis Harness 对单次 Run 执行状态的内存控制,采用 first-terminal-wins 规则保证成功、失败、取消、超时和预算耗尽只能产生一个最终终态。
- 边界:Run Lifecycle 不直接等同于数据库实体写入;应用用例负责把最终状态映射到 `diagnosis_run` 持久化。
### Run Budget
- 定义:单次 Run 的模型调用、Tool 调用、单 Tool 调用、输入/输出/总 Token 和 canonical invocation 字节容量的线程安全消耗计数与门禁。
- 边界:预算上限由 Harness 配置显式提供;实际 Token 在模型响应后记录,超限后保留真实消耗并阻止后续执行。
### Harness Retry Policy
- 定义:Harness 对同一技术操作 attempt 数和可重试失败类型的显式策略。
- 边界:Router 与 SemanticGuard 的技术失败最多两次 attempt;Diagnosis Agent、Tool 和 Evidence repair 只有一次 attempt。Agent 正常 ReAct 轮次不是 retry,`NO_EVIDENCE`、业务拒绝、取消和预算耗尽不可重试。
### Chat Application Use Case
- 定义:一次 Chat 请求的唯一业务入口,拥有 Session/Run、意图路由、固定执行器、PreviousTurn 和最终持久化。
- 边界:不拥有 HTTP/SSE 连接,不把 ChatModel 或 Tool 选择权交给 Controller,也不在 Diagnosis Release Use Case 之外单独决定预算 Fallback 的业务内容。
### Chat SSE Contract
- 定义:Chat 公开入口的五事件协议,顺序固定为 `metadata -> status* -> content|failure -> done`。
- 边界:过程状态实时发送,最终安全内容最多释放一次;它不是 Token streaming,也不包含内部计划、Prompt、raw Tool 数据或异常。
### Verifier Skill Isolation
- 定义:Chat Verifier 与 skill 系统隔离,只校验 Executor 答案和 `tool_trace_summary`。
- 使用场景:防止 Verifier 把 playbook 指令当作事实证据;Verifier 只判断已有证据是否支持结论。
- 边界:Verifier 不接收 `skill_catalog`,不暴露 `read_skill`,不读取 `SKILL.md`。
## Diagnosis Playbook Business Rules
- Planner 只看 skill metadata,输出 `selected_skill`、`selection_reason` 和 plan。
- Executor 才能调用 `read_skill(selected_skill)`,并且读取 skill 后仍必须调用 evidence tools。
- Skill 正文不得替代 `lookup_knowledge`、日志、指标或告警数据。
- Verifier 只基于 `tool_trace_summary` 校验事实,不基于 skill 正文校验事实。
- 当前阶段保留单 active skill 白名单:`diagnose-mysql-connection-pool`。
+57 -21
View File
@@ -1,24 +1,60 @@
# devflow 索引
# devflow 索引
## Issue 生命周期
| Issue | 状态 | 说明 |
|---|---|---|
| ISS-014 | archived | 阶段 0-7 的单体 Diagnosis Agent、Harness、ACI、SSE、清理和最终 E2E 已完成并归档;阶段实现对应的 11 个 devflow/OpenSpec 项目均已 archived。 |
| ISS-015 | active | 阶段 1 硬停止已由 ISS-016 收口;剩余 Evidence Repair Schema、Reasoning 审计验证/治理与最终综合验收。 |
| ISS-016 | archived | Diagnosis 信息增益停止契约、协议修复反馈与统一 Release 已完成并归档。 |
## 项目
| 日期 | slug | 领域 | 关键词 | 状态 |
|---|---|---|---|---|
| 2026-07-05 | diagnosis-playbook-skills | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented |
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
| 2026-07-04 | diagnosis-eval-harness | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
| 2026-07-04 | evidence-trace-hardening | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
| 2026-07-03 | mvp-demo-trace-acceptance | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
| 2026-07-04 | aiops-traceable-diagnosis-entry | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
| 2026-07-04 | aiops-alert-scope-control | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
| 2026-05-29 | chatmodel-abstraction | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | archived |
| 2026-06-23 | phase1-infrastructure | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | archived |
| 2026-06-24 | lookup-knowledge-integration | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | archived |
| 2026-06-25 | doc-management-ui | 前端开发/文档管理 | 文档管理页面, CRUD, 状态监控, 纯静态页面, API集成 | archived |
| 2026-06-26 | session-storage | 会话存储/可观测 | diagnosis_session, agent_step, tool_invocation, token追踪, 多Agent路由 | openspec/changes/session-storage | archived |
| 2026-06-29 | confidence-feedback | 质量评估/反馈机制 | evidence_score, selfEvaluation, feedback, useful, not_useful, case_library, BAD_CASE, tool_invocation规则引擎, 反馈按钮, sessionId回传 | openspec/changes/confidence-feedback | archived |
| 2026-06-30 | session-dedup-knowledge-map | 去重/知识图谱 | RetrievedDocTracker, KnowledgeDomainService, knowledge_domain, covers, whenToRetrieve, Planner注入, ISS-001 | openspec/changes/archive/2026-06-30-session-dedup-knowledge-map | archived |
| 2026-07-01 | executor-action-memory-relevance | 检索质量/行动记忆 | relevanceLevel, completenessHint, Min-Max归一化, RetrievedDocTracker域级记录, Executor检索约束, ISS-002 | openspec/changes/archive/2026-07-01-executor-action-memory-relevance | archived |
| 2026-07-02 | chat-verifier-agent | Chat质量门禁/可追溯验证 | Verifier, groundedness_score, facts_checked, evidence_refs, tool_trace_summary, self_evaluation | openspec/changes/archive/2026-07-03-chat-verifier-agent | archived |
| 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|---|---|---|---|---|---|---|
| 2026-07-28 | rag-eval-hybrid-baseline | 离线 RAG eval 对齐 hybrid:search.mode 生成器、fixture meta、baseline 重刷;hybrid 质量闸门可用 denseDistance。 | RAG/eval/baseline | search.mode, fixture meta, kb_scope rag-eval, L0 filter fallback, denseDistance | openspec/changes/archive/2026-07-28-rag-eval-hybrid-baseline | archived |
| 2026-07-28 | rag-quality-score-unify | 统一 dense/hybrid scoreLabel 与 qualityScore;保检索序;去掉关键词 boost 改序与 hybrid L2 伪装。 | RAG/质量分/后处理 | qualityScore, scoreLabel dense/hybrid, originalRank, RetrievalScoreNormalizer, no boost rerank | openspec/changes/archive/2026-07-28-rag-quality-score-unify | archived |
| 2026-07-26 | diagnosis-information-gain-stop-contract | Diagnosis 信息增益停止、协议修复反馈、ProgressSnapshot 与统一 Release。 | Harness/Diagnosis stop/Release | ISS-016, GAINED, NO_GAIN, STOP_REQUIRED, ProgressSnapshot, PROGRESS_PROTOCOL_VIOLATED, INSUFFICIENT_EVIDENCE | openspec/changes/archive/2026-07-27-diagnosis-information-gain-stop-contract | archived |
| 2026-07-27 | rag-chunk-evidence-identity-dedup | chunk 级证据身份、去重、retrieve-k/return-n 与 SearchPort 地基,为 hybrid 铺路。 | RAG/证据身份/去重 | evidenceKey, maxChunksPerDocument, retrieve-k, return-n, KnowledgeSearchPort, document_id chunk-scoped | openspec/changes/archive/2026-07-27-rag-chunk-evidence-identity-dedup | archived |
| 2026-07-27 | rag-bm25-hybrid-drop-sdk | 真 dense+BM25 hybrid(MilvusClientV2),废弃知识路径旧 SDK 检索/写入。 | RAG/BM25/hybrid | MilvusClientV2, BM25, hybridSearch, RRFRanker, biz_hybrid, drop SDK path | openspec/changes/archive/2026-07-27-rag-bm25-hybrid-drop-sdk | archived |
| 2026-07-27 | rag-hybrid-search-rrf | Delivery 2:可配置 hybrid 检索与 RRF 多路融合(不绑旧 SDK)。 | RAG/hybrid/RRF | hybrid mode, RRF, KnowledgeSearchPort, filtered+unfiltered fusion, sparse-lite lexical | openspec/changes/archive/2026-07-27-rag-hybrid-search-rrf | archived |
| 2026-07-21 | single-react-tool-invocation-store | 建立统一 ToolBoundary 与 Redis canonical invocation store,集中生命周期、证据状态、TTL、容量和 Run 所有权。 | Harness/Tool boundary/Canonical store | ISS-014, ToolBoundary, canonical invocation, PROJECTING, READY, ERROR, TTL, RESULT_TOO_LARGE | openspec/changes/archive/2026-07-21-single-react-tool-invocation-store | archived |
| 2026-07-21 | single-react-harness-run-context | 建立显式 RunContext、Harness Core、预算、取消、类型化重试和 Tool Store 基础。 | Harness/Run lifecycle/Budget | ISS-014, RunContext, deadline, cancellation, budget, retry, ToolCallKey | openspec/changes/archive/2026-07-21-single-react-harness-run-context | archived |
| 2026-07-21 | single-react-aci-tool-contracts | 冻结 RAG、日志和 MySQL evidence Tool 的 Agent-facing ACI Schema、状态、框架调用引用和描述边界。 | Harness/Agent Tool contract | ISS-014, ACI, tool_call_id, evidence_status, RAG, query_logs, query_mysql, MOCK | openspec/changes/archive/2026-07-21-single-react-aci-tool-contracts | archived |
| 2026-07-21 | single-react-design-freeze | 冻结单体 Diagnosis Agent、Harness、Guard、工具证据与阶段门禁契约。 | Chat/Harness/Agent contract | ISS-014, single ReactAgent, Harness, EvidenceGuard, SemanticGuard, tool_call_id, evidence_status | openspec/changes/archive/2026-07-21-single-react-design-freeze | archived |
| 2026-07-10 | session-run-trace-isolation | 拆分会话态和运行态,引入 runId 隔离 Trace、Feedback、AIOps 和 demo 链路。 | Trace/session/run isolation | chat_session, diagnosis_run, runId, trace exact run, feedback fallback, AIOps SSE metadata, baseline drift | openspec/changes/archive/2026-07-10-session-run-trace-isolation | archived |
| 2026-07-09 | interview-demo-quality-audit | 增加面试演示前置质量审计,覆盖 prompt、Gatekeeper 和评测基线。 | Agent eval/demo/Prompt audit | interview demo preflight, prompt_audit, gatekeeper rules, diagnosis baseline, 12 fixtures | openspec/changes/archive/2026-07-09-interview-demo-quality-audit | archived |
| 2026-07-08 | executor-composer-final-answer | 引入 Composer 生成最终回答,只使用 Verifier 允许的结论材料。 | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
| 2026-07-08 | diagnosis-eval-demo-gatekeeper-closure | 收敛诊断评测、稳定 demo 场景和 Gatekeeper 审计元数据。 | Agent eval/demo/Gatekeeper | diagnosis eval matrix, stable demo scenarios, Gatekeeper rule set version, audit metadata | openspec/changes/archive/2026-07-08-diagnosis-eval-demo-gatekeeper-closure | archived |
| 2026-07-08 | verifier-evidence-reference-fidelity | 强化 Verifier 对 evidence_refs、raw_path 和 no_evidence 的保真校验。 | Chat质量门禁/证据归因 | evidence_refs, raw_path, Gatekeeper severity, verifier evidence excerpt, HikariCP mock, no_evidence | openspec/changes/archive/2026-07-08-verifier-evidence-reference-fidelity | archived |
| 2026-07-07 | executor-evidence-output-contract | 设计 Executor 结构化证据输出,解决证据归因幻觉和 LOW_CONFID 问题。 | Chat质量门禁/证据归因 | Executor structured output, evidence bindings, Verifier structured claims, LOW_CONFID, hallucination | openspec/changes/archive/2026-07-07-executor-evidence-output-contract | archived |
| 2026-07-07 | executor-v2-output-contract | 将 Executor 输出升级为 V2 契约,移除面向用户的最终回答字段。 | Chat质量门禁/证据归因 | executor_evidence_v2, user_facing_answer removal, diagnosis_summary removal, structured renderer | openspec/changes/archive/2026-07-07-executor-v2-output-contract | archived |
| 2026-07-07 | executor-gatekeeper-hook | 在 Executor 与 Verifier 之间接入 Gatekeeper,校验证据绑定来源。 | Chat质量门禁/证据归因 | Gatekeeper, verifier payload, source_invocation_ids, tool_name match, self_evaluation | openspec/changes/archive/2026-07-07-executor-gatekeeper-hook | archived |
| 2026-07-07 | executor-verifier-claim-checks | 增加 Verifier claim_checks 和事实校验兼容逻辑。 | Chat质量门禁/证据归因 | Verifier claim_checks, facts_checked compatibility, effective verdict guardrail, malformed output downgrade | openspec/changes/archive/2026-07-07-executor-verifier-claim-checks | archived |
| 2026-07-06 | rag-eval-pipeline-closure | 建立 RAG 评测闭环,加入 fixture、快照和 baseline diff。 | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
| 2026-07-06 | modular-rag-pipeline | 将 lookup_knowledge 改造成模块化 RAG 管线,补齐证据块和检索追踪。 | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
| 2026-07-05 | diagnosis-playbook-skills | 增加诊断 Playbook Skill,沉淀支付超时、MySQL 池、Redis 超时等套路。 | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented |
| 2026-07-05 | mvp-demo-interview-runbook | 准备可复现的 MVP 面试演示包、运行手册和 Trace 检查清单。 | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
| 2026-07-05 | diagnosis-eval-baseline-diff | 增加诊断评测 baseline diff,用于判断回归和证据覆盖变化。 | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
| 2026-07-04 | expand-diagnosis-eval-fixtures | 扩充诊断评测 fixture,覆盖 Redis、慢响应和 JVM 内存风险。 | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
| 2026-07-04 | diagnosis-eval-harness | 建立固定诊断评测 Harness,输出 trace、证据覆盖和 verdict 分布。 | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
| 2026-07-04 | evidence-trace-hardening | 强化工具调用证据链、降级契约和离线验证能力。 | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
| 2026-07-04 | aiops-traceable-diagnosis-entry | 增加可追踪的 AIOps 告警诊断入口,打通 sessionId 和 Trace API。 | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
| 2026-07-04 | aiops-alert-scope-control | 收敛 AIOps 告警诊断范围,区分 payload 定向和自动发现模式。 | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
| 2026-07-03 | mvp-demo-trace-acceptance | 增加 MVP demo 的 Trace 验收,覆盖会话、步骤、工具和反馈链路。 | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
| 2026-07-02 | chat-verifier-agent | 增加 Chat Verifier Agent,用 groundedness 和 evidence_refs 校验回答。 | Chat质量门禁/可追溯验证 | Verifier, groundedness_score, facts_checked, evidence_refs, tool_trace_summary, self_evaluation | openspec/changes/archive/2026-07-03-chat-verifier-agent | archived |
| 2026-07-01 | executor-action-memory-relevance | 增加行动记忆和相关性信号,约束 Executor 重复检索。 | 检索质量/行动记忆 | relevanceLevel, completenessHint, Min-Max归一化, RetrievedDocTracker域级记录, Executor检索约束, ISS-002 | openspec/changes/archive/2026-07-01-executor-action-memory-relevance | archived |
| 2026-06-30 | session-dedup-knowledge-map | 引入会话级去重和知识域地图,减少重复召回。 | 去重/知识图谱 | RetrievedDocTracker, KnowledgeDomainService, knowledge_domain, covers, whenToRetrieve, Planner注入, ISS-001 | openspec/changes/archive/2026-06-30-session-dedup-knowledge-map | archived |
| 2026-06-29 | confidence-feedback | 建立质量评估和用户反馈机制,并把有用反馈沉淀为案例。 | 质量评估/反馈机制 | evidence_score, selfEvaluation, feedback, useful, not_useful, case_library, BAD_CASE, tool_invocation规则引擎, 反馈按钮, sessionId回传 | openspec/changes/confidence-feedback | archived |
| 2026-06-26 | session-storage | 建立通用会话存储,记录 session、agent step 和 tool invocation。 | 会话存储/可观测 | diagnosis_session, agent_step, tool_invocation, token追踪, 多Agent路由 | openspec/changes/session-storage | archived |
| 2026-06-25 | doc-management-ui | 实现文档管理页面,支持文档 CRUD、状态监控和 API 集成。 | 前端开发/文档管理 | 文档管理页面, CRUD, 状态监控, 纯静态页面, API集成 | - | archived |
| 2026-06-24 | lookup-knowledge-integration | 接入知识库检索,支持 L0 精确匹配和 L1 语义检索。 | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | - | archived |
| 2026-06-23 | phase1-infrastructure | 搭建第一阶段基础设施,包括 MySQL、Redis、Milvus、Flyway 和 JPA。 | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | - | archived |
| 2026-05-29 | chatmodel-abstraction | 抽象 ChatModel 和 EmbeddingModel,支持多模型路由。 | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | - | archived |
| 2026-07-21 | single-react-rag-log-projections | RAG/log projection adapters through ToolBoundary | Harness/Tool projection | ISS-014, RAG, query_logs, projection, scope, redaction, MOCK, NO_EVIDENCE | openspec/changes/archive/2026-07-21-single-react-rag-log-projections | archived |
| 2026-07-21 | single-react-mysql-readonly-tool | Fail-closed read-only MySQL evidence Tool with AST allowlist, JDBC controls and bounded projection | Harness/MySQL security | ISS-014, MySQL, JSqlParser, allowlist, PreparedStatement, timeout, projection | openspec/changes/archive/2026-07-21-single-react-mysql-readonly-tool | archived |
| 2026-07-21 | single-react-diagnosis-agent | Single internal Diagnosis ReactAgent with Harness-controlled model/tool loop, bounded context and typed Draft | Harness/Diagnosis Agent/ReAct | ISS-014, ReactAgent, DiagnosisDraft, PreviousTurn, ToolInterceptor, ModelInterceptor, budget | openspec/changes/archive/2026-07-21-single-react-diagnosis-agent | archived |
| 2026-07-21 | single-react-evidence-semantic-guards | Deterministic evidence validation, isolated semantic review and fail-closed diagnosis release | Harness/EvidenceGuard/SemanticGuard/Release | ISS-014, EvidenceGuard, verified snapshot, SemanticGuard, repair, fallback, release policy | openspec/changes/archive/2026-07-21-single-react-evidence-semantic-guards | archived |
| 2026-07-21 | single-react-chat-application-usecase | Internal Chat application use case with isolated routing, fixed executors and safe PreviousTurn | Harness/Chat application/Run persistence | ISS-014, Intent Router, PreviousTurn, PublishedResult, V012, observer, cancellation | openspec/changes/archive/2026-07-21-single-react-chat-application-usecase | archived |
| 2026-07-21 | single-react-chat-sse-cutover | Unique named-event Chat SSE endpoint, bounded production Harness wiring and strict frontend consumer | Chat/SSE/Harness production wiring | ISS-014, /api/chat, SSE, metadata, status, content, failure, done, disconnect, bounded executor | openspec/changes/archive/2026-07-22-single-react-chat-sse-cutover | archived |
| 2026-07-22 | single-react-cleanup-e2e | Remove legacy Agent paths, add bounded Harness audit, and complete exact-run live acceptance | Chat/Harness/cleanup/E2E | ISS-014, single ReAct Agent, durable audit, named SSE, exact run, Flyway V013 | openspec/changes/archive/2026-07-22-single-react-cleanup-e2e | archived |
@@ -48,7 +48,7 @@
## 遗留问题
ISS-002:Executor 无约束重复调用 `lookup_knowledge`(单会话 20+ 次),knowledge map 和检索约束只注入了 Planner 未注入 Executor。详见 `mvp/issues/ISS-002-executor-unconstrained-lookup.md`。
ISS-002:Executor 无约束重复调用 `lookup_knowledge`(单会话 20+ 次),knowledge map 和检索约束只注入了 Planner 未注入 Executor。详见 `mvp/issues/archived/ISS-002-executor-unconstrained-lookup.md`。
## 已知限制
@@ -9,8 +9,8 @@
## Context
- `devflow/index.md` was checked. Relevant history includes `session-storage`, `confidence-feedback`, `executor-action-memory-relevance`, and `chat-verifier-agent`.
- `mvp/notes/agent-engineering-decisions.md` already recommends the next phase as "可复现 MVP Demo", including `mvp-demo` profile, fixed diagnosis case, one-click request, and `GET /api/diagnosis/{sessionId}/trace`.
- `mvp/issues/ISS-003-mvp-design-implementation-review.md` identifies test stability, session traceability, verifier evidence chain, upload path, and SupervisorAgent consistency as recent MVP concerns. Security cleanup is intentionally deferred by user decision.
- `mvp/archive/2026-07-09-doc-cleanup/notes/agent-engineering-decisions.md` already recommends the next phase as "可复现 MVP Demo", including `mvp-demo` profile, fixed diagnosis case, one-click request, and `GET /api/diagnosis/{sessionId}/trace`.
- `mvp/issues/archived/ISS-003-mvp-design-implementation-review.md` identifies test stability, session traceability, verifier evidence chain, upload path, and SupervisorAgent consistency as historical MVP concerns. Security cleanup was intentionally deferred by user decision at that time.
## Question Pool
@@ -0,0 +1,101 @@
# Modular RAG Pipeline — Acceptance
## 验收状态
状态:通过,OpenSpec 已归档。
任务完成:
- OpenSpec tasks:31/31 完成。
- Review 后新增去重边界修复和回归测试。
- OpenSpec archive:`openspec/changes/archive/2026-07-06-modular-rag-pipeline`。
## 静态验证
```powershell
openspec validate modular-rag-pipeline --strict
```
结果:
```text
Change 'modular-rag-pipeline' is valid
```
```powershell
git diff --check
```
结果:
```text
PASS
```
说明:仅出现 Windows LF/CRLF warning,无 whitespace error。
## 脚本验证
```powershell
mvn -q -DskipTests compile
```
结果:PASS。
```powershell
mvn -q "-Dtest=LookupKnowledgeToolTest,ToolInvocationRecorderTest" test
```
结果:PASS。
覆盖:
- filtered L1 成功不 retry。
- filtered L1 低质量触发 raw unfiltered retry。
- filtered L1 无 evidence 触发 raw unfiltered retry。
- L0 hint 不作为 standalone fact evidence。
- 无 L0 hint 时直接 unfiltered vector search。
- rerank 使用 hint match 并记录 trace。
- context pack 保留 source/title/breadcrumb/hit reasons。
- evidence blocks 按 source 去重。
- session dedup 不再返回可消费 evidence/context。
- recorder 记录 retrieval trace、rerank trace、context pack summary 和 evidence summaries。
```powershell
$env:MILVUS_TOKEN = <application.yml 中的 milvus.token>; mvn -q test
```
结果:PASS。
说明:
- `MilvusConnectionTest` 需要 `MILVUS_TOKEN` 环境变量,直接读 `System.getenv`,不会自动读 `application.yml`。
- 注入该环境变量后完整测试通过。
## 实现验收
已验证行为:
- `LookupKnowledgeTool` 已变为 pipeline orchestrator。
- `LookupResult` 新契约包含 `evidenceBlocks`、`contextPack`、`retrievalTrace`、`rerankTrace`。
- 旧 `primary/supplement` 字段和 DTO 已删除。
- `ToolInvocationRecorder` 不再依赖 `result.getPrimary()`。
- filtered retrieval 失败时会记录 `filtered_vector_no_evidence` 或 `filtered_vector_low_quality`。
- no-evidence 情况返回 `found=false` 且保留 retrieval trace。
- session dedup 情况返回 `found=false` 且 evidence/context 为空。
## 未验证项
人工 Demo 未执行:
- 还没有通过真实 Chat/AIOps 会话观察 Agent 是否稳定按 `contextPack.packedText` 和 `evidenceBlocks` 引用证据。
风险:
- 工具 JSON 契约是 L4 breaking change,prompt 已更新,但真实对话行为仍建议做一次端到端 demo。
## 后续建议
- 增加一组 RAG eval cases,固定 query、期望 evidence source、期望 fallback path。
- 将 `MilvusConnectionTest` 改成 Spring 配置驱动或 integration profile,避免配置源混用。
- 后续可在评测数据足够后再考虑 model-based rerank 或 hybrid retrieval。
@@ -0,0 +1,54 @@
# Modular RAG Pipeline — Brief
## 背景
`lookup_knowledge` 已经能返回知识库证据,但实现集中在 `LookupKnowledgeTool` 内部:L0 查询分析、L1 向量召回、相关性归一化、证据组装、会话去重和 trace 入库耦合在一起。
旧返回契约 `primary/supplement` 也延续了“L0 是主结果、L1 是补充”的语义,和当前设计目标不一致。新的目标是让 L0 只作为 query understanding / filter / rerank / trace 信号,让 L1 向量检索成为事实证据来源。
## 目标
- 将 `lookup_knowledge` 改造成模块化 RAG pipeline。
- 保留显式 Agent tool 边界,不改工具名和 query 参数。
- L0 只提供领域、关键词、实体、category filter 和 trace hint。
- L1 filtered vector retrieval 失败或低质量时,降级为 raw query unfiltered L1 retry。
- 输出 evidence-first contract:`evidenceBlocks`、`contextPack`、`retrievalTrace`、`rerankTrace`。
- 保持 `tool_invocation` 表结构稳定,把新 trace 写入 `retrieval_details` JSON。
## 范围
已完成:
- 新增 pipeline DTO:`KnowledgeQuery`、`RetrievedEvidenceCandidate`、`ContextPack`、`RetrievalTrace`、`RerankTrace`、`EvidencePostprocessResult`。
- 新增 pipeline service:`KnowledgeQueryTransformer`、`KnowledgeDocumentRetriever`、`KnowledgeEvidencePostProcessor`、`KnowledgeContextPacker`、`LookupResultAssembler`。
- 重构 `LookupKnowledgeTool` 为薄 orchestration 层。
- 迁移 `LookupResult`,删除 `primary/supplement` 字段和 `PrimaryResult` / `SupplementResult` 类。
- 更新 `ToolInvocationRecorder`,记录 query transform、retrieval trace、context pack summary、rerank trace、fallback reason 和 evidence summaries。
- 更新 executor prompt 和 RAG 架构文档。
- 补充 lookup、recorder、fallback、rerank、context pack、session dedup 测试。
非目标:
- 不引入 implicit Advisor。
- 不引入 cross-encoder、BM25、RRF、Elasticsearch、OpenSearch。
- 不改文档上传、chunk、embedding 写入、Milvus schema。
- 不改变 Agent 何时调用 `lookup_knowledge`。
## 关联 OpenSpec
- `openspec/changes/archive/2026-07-06-modular-rag-pipeline`
## 接口影响
级别:L4 breaking interface。
原因:
- 删除旧 `LookupResult.primary` / `LookupResult.supplement`。
- `lookup_knowledge` tool JSON 输出形状变化。
缓解:
- 工具名和输入参数保持不变。
- in-repo 消费方、测试和 prompt 同步迁移。
- `tool_invocation` 表结构不变。
@@ -0,0 +1,72 @@
# Modular RAG Pipeline — Decisions
## D1: `lookup_knowledge` 保持显式工具
不把知识检索做成隐式 Advisor。Agent 仍显式调用 `lookup_knowledge(query)`,这样 trace、Verifier、Eval 都能看到工具调用边界。
## D2: L0 只做 query understanding
L0 产出:
- `domainHints`
- `matchedKeywords`
- `entities`
- `categoryFilter`
- `l0Titles`
- `l0MatchCount`
L0 不再直接转成 fact evidence。L0 hint 可以影响 filter、rerank、trace,但不能在 L1 失败时冒充知识证据。
## D3: MVP 降级策略采用 unfiltered L1 retry
流程:
```text
filtered L1 with L0 category filter
-> empty / no final evidence / below reference threshold
-> raw query unfiltered L1 retry
-> still no evidence => no_evidence
```
取舍:
- 简单、可解释、适合 MVP。
- 避免引入 BM25/RRF/multi-query/cross-encoder 的复杂度。
- 代价是低质量场景多一次向量查询,已通过 trace 记录 attempt duration。
## D4: 删除 `primary/supplement`
这是一次 L4 breaking interface change。
删除原因:
- `primary/supplement` 绑定旧语义:L0 primary、L1 supplement。
- 新设计中事实证据来自 `evidenceBlocks/contextPack`。
迁移结果:
- `LookupResult` 暴露 evidence-first 字段。
- `PrimaryResult` / `SupplementResult` 已删除。
- 生产代码和测试不再引用 `getPrimary()` / `getSupplement()`。
## D5: Trace 表结构保持稳定
`tool_invocation` 表不新增列。新增信息写入 `retrieval_details` JSON:
- `query_transform`
- `retrieval_trace`
- `context_pack_summary`
- `rerank_trace`
- `fallback_reason`
- `evidence_blocks`
原因:当前 trace、Verifier、Eval 已经以 `tool_invocation` 为证据入口,JSON details 足够承载 RAG 细节,避免 schema churn。
## D6: 会话去重不返回可消费证据
Review 后修正:
- dedup result 的 `found=false` 必须和 evidence/context 语义一致。
- 返回消息说明文档已检索过。
- 不再返回 `evidenceBlocks/contextPack`,避免 Agent 重复使用同一证据。
- 保留 `retrievalTrace` 和 `retrievedDomainsThisSession` 便于可观测。
@@ -0,0 +1,56 @@
# Modular RAG Pipeline — Evidence
## 代码证据
关键入口:
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
- `src/main/java/com/superbiz/agent/dto/LookupResult.java`
新增模块:
- `KnowledgeQueryTransformer`:复用 `KnowledgeIndexService.analyzeQuery`,把 L0 转成 query hints 和可选 `categoryFilter`。
- `KnowledgeDocumentRetriever`:封装 `VectorSearchService.searchSimilarDocuments(query, topK, category)`,统一 filtered / unfiltered attempt。
- `KnowledgeEvidencePostProcessor`:归一化 L2、创建 evidence blocks、source dedup、规则 rerank、输出 `RerankTrace`。
- `KnowledgeContextPacker`:按字符预算打包 evidence,保留 source/title/breadcrumb/hit reasons。
- `LookupResultAssembler`:统一组装 evidence-first result、no-evidence result、session dedup result。
## 设计证据
已有文档约束:
- `mvp/architecture/rag-architecture.md`:RAG 应表达为可解释 pipeline,而不是一坨工具逻辑。
- `mvp/architecture/retrieval-observability.md`:L0 是 hint/explainability 层,L1 是语义检索主路径。
- `devflow/glossary/CONTEXT.md`:`lookup_knowledge` 是显式 Agent evidence tool,`tool_invocation` 是 trace / verifier / eval 的证据来源。
OpenSpec 对齐:
- `openspec/changes/modular-rag-pipeline/proposal.md`
- `openspec/changes/modular-rag-pipeline/design.md`
- `openspec/changes/modular-rag-pipeline/specs/rag-knowledge-retrieval/spec.md`
- `openspec/changes/modular-rag-pipeline/tasks.md`
## 用户确认
- 一次到位做模块化 RAG,而不是只做小补丁。
- L0 不再作为事实证据兜底。
- filtered L1 不准时,MVP 降级为 raw query unfiltered L1 retry。
- 可以新增字段,并删除旧字段以换取后续流程清晰。
## Review 发现
Review 中发现一个非阻塞但应修复的问题:
- 会话去重命中时,返回 `found=false` 但仍带 `evidenceBlocks/contextPack`,可能导致 Agent 重复消费同一份证据。
修复:
- `LookupResultAssembler.deduped` 清空可消费 evidence/context,只保留 message、trace、relevance hint 和 session domain memory。
- 新增 `LookupKnowledgeToolTest.sessionDedupDoesNotReturnConsumableEvidenceAgain`。
## 非阻塞观察
- `MilvusConnectionTest` 仍直接依赖 `MILVUS_TOKEN` 环境变量;主配置中已有 token,但测试不读 Spring 配置。
- 测试日志仍有 ANTLR 版本 warning,不影响测试通过。
- 控制台在部分命令输出中仍会出现中文编码显示问题,但源码按 UTF-8 读取时关键用户提示文本正常。
@@ -0,0 +1,78 @@
# Acceptance: rag-eval-pipeline-closure
## Status
Archived.
## Acceptance Criteria
| Item | Status | Notes |
|---|---|---|
| Modular fixture support | Done | Evaluator reads `lookupResult.evidenceBlocks/contextPack/retrievalTrace/rerankTrace`. |
| LookupResult-only contract | Done | Evaluator fails fixtures that do not expose `lookupResult`. |
| Golden modular assertions | Done | Cases assert selected attempt, accepted fallback reason, evidence status, context sources, and rerank top source. |
| Fallback coverage | Done | Added `chat-l0-filter-fallback` for filtered low-quality/no-evidence to unfiltered retry. |
| Real tool snapshot generation | Done | Added `RagLookupSnapshotGeneratorTest` and `generate_rag_lookup_snapshots.ps1`, defaulting to Spring AI VectorStore mode. |
| Seed docs import/reindex | Done | Added canonical seed docs, `RagEvalSeedImporterTest`, and `prepare_rag_eval_seed.ps1`. |
| Eval metadata isolation | Done | Added `kb_scope` metadata and `retrieval.kb-scope` filtering for L0 and L1. |
| Frontmatter body split | Done | Upload chunking embeds Markdown body, while frontmatter feeds metadata and L0. |
| Baseline diff | Done | `--compare-to` writes JSON/Markdown diff and exits non-zero on regression. |
| Documentation | Done | Updated RAG eval README and added `mvp/architecture/rag-eval-closure.md`. |
## Verification
```powershell
python scripts\eval_rag_retrieval.py
```
Result: passed. 7 cases, passRate=1.0, recall@5=1.0.
```powershell
mvn -q "-Dtest=RagLookupSnapshotGeneratorTest" test
```
Result: passed. The snapshot generator stays disabled unless `rag.snapshot.enabled=true` is provided.
```powershell
mvn -q "-Dtest=RagLookupSnapshotGeneratorTest" "-Drag.snapshot.enabled=true" "-Drag.snapshot.fixtures=<temp-fixtures>" "-Drag.snapshot.retrievedAt=2026-07-06T00:00:00Z" "-Dretrieval.kb-scope=rag-eval" "-Dretrieval.vector-store.mode=spring" test
python scripts\eval_rag_retrieval.py --fixtures <temp-fixtures> --json-report <temp-current.json> --markdown-report <temp-current.md>
```
Result: passed. Spring AI VectorStore live snapshot produced 7 cases, passRate=1.0, recall@5=1.0. The fallback case used `selectedAttempt=UNFILTERED_VECTOR_RETRY` and `fallbackReason=filtered_vector_no_evidence`; the expected source `rag-l0-filter-fallback` remained rank 1.
```powershell
python scripts\eval_rag_retrieval.py --json-report <temp-current.json> --markdown-report <temp-current.md> --compare-to eval\rag-retrieval\reports\baseline.json --diff-json-report <temp-diff.json> --diff-markdown-report <temp-diff.md>
```
Result: passed. regressions=0.
```powershell
mvn -q "-Dtest=FrontmatterParserTest,VectorIndexServiceTest,VectorSearchServiceTest,DocumentManagementServiceTest,RagLookupSnapshotGeneratorTest,RagEvalSeedImporterTest" test
```
Result: passed. The seed importer and snapshot generator remain disabled unless their system properties are explicitly enabled.
```powershell
$null = [scriptblock]::Create((Get-Content -Raw scripts\prepare_rag_eval_seed.ps1))
$null = [scriptblock]::Create((Get-Content -Raw scripts\generate_rag_lookup_snapshots.ps1))
```
Result: PowerShell syntax OK.
```powershell
.\scripts\prepare_rag_eval_seed.ps1
```
Result: passed. Seed docs were imported through `DocumentManagementService` and reindexed into the configured runtime DB/vector stack.
```powershell
.\scripts\generate_rag_lookup_snapshots.ps1 -Fixtures <temp-fixtures> -RetrievedAt 2026-07-06T00:00:00Z -SkipEval
```
Result: passed after defaulting the script to `retrieval.vector-store.mode=spring`. The script generated live fixtures through the real `LookupKnowledgeTool` and then the offline evaluator reported 7 cases, passRate=1.0, recall@5=1.0.
```powershell
git diff --check
```
Result: no whitespace errors. Git reported only LF/CRLF conversion warnings.
@@ -0,0 +1,21 @@
# Brief: rag-eval-pipeline-closure
## Background
The modular RAG pipeline now returns `LookupResult` with `evidenceBlocks`, `contextPack`, `retrievalTrace`, and `rerankTrace`. The RAG retrieval baseline must validate that full contract, so it can detect regressions in fallback behavior, context packing, or rerank trace.
## Goals
1. Reuse the existing offline RAG retrieval baseline.
2. Extend it to support modular `LookupResult` fixtures.
3. Add golden assertions for selected attempt, fallback reason, evidence status, context sources, and rerank top source.
4. Add a RAG baseline diff path for regression detection.
5. Add a snapshot generator that calls the real `LookupKnowledgeTool`.
6. Document how RAG baseline and diagnosis baseline form a quality loop.
## Non-Goals
- No new production API.
- No LLM-as-judge scoring.
- No production API behavior changes.
- No replacement for diagnosis eval.
@@ -0,0 +1,57 @@
# Decisions: rag-eval-pipeline-closure
## D1. Reuse the existing evaluator
Decision: extend `scripts/eval_rag_retrieval.py` instead of creating a second evaluator.
Reason: the old evaluator already owns golden cases, fixtures, hit-level classification, and Markdown/JSON reports. Extending it keeps one RAG baseline path.
## D2. Use LookupResult as the only fixture contract
Decision: support `lookupResult` only.
Reason: the MVP has moved to evidence-first RAG. Keeping an older fixture contract would weaken the baseline and let incomplete fixtures bypass context packing, retrieval trace, and rerank checks.
## D3. Make modular assertions opt-in per case
Decision: use fields such as `expectedSelectedAttempt`, `expectedFallbackReason`/`expectedFallbackReasons`, `expectedEvidenceStatus`, `expectedContextSources`, and `expectedRerankTopSource`.
Reason: golden cases can be strict where the pipeline path matters without forcing every historical case to assert every new field.
## D4. Diff remains deterministic
Decision: RAG diff compares report fields only and does not call live services or models.
Reason: this keeps it suitable for local regression checks and CI-style gates.
## D5. Isolate live eval docs with kb_scope
Decision: add `kb_scope` metadata and use `rag-eval` for canonical eval seed documents.
Reason: local production documents are not stable enough for golden retrieval expectations. Scope isolation lets real `LookupKnowledgeTool` snapshots use the same MySQL/Milvus stack while avoiding accidental matches from unrelated local data.
Default runtime keeps `retrieval.kb-scope` empty so legacy documents without `kb_scope` remain searchable. Eval scripts pass `-Dretrieval.kb-scope=rag-eval`. The same scope applies to L0 query hints and L1 vector retrieval.
## D6. Import seed docs through the real upload pipeline
Decision: seed docs are imported by `RagEvalSeedImporterTest` through `DocumentManagementService.uploadDocument`.
Reason: this updates DB metadata, L0 index state, local knowledge files, and Milvus chunks in the same way as normal document ingestion. A direct Milvus-only seed would make the live eval less representative.
## D7. Strip frontmatter before chunk embedding
Decision: uploaded Markdown frontmatter feeds metadata/L0 but is stripped before chunking and embedding.
Reason: frontmatter is a control plane, not evidence text. Keeping it in chunks lets L0-only keywords artificially improve vector similarity, especially for fallback decoy cases.
## D8. Treat retry behavior as the stable fallback contract
Decision: the fallback golden case accepts both `filtered_vector_low_quality` and `filtered_vector_no_evidence`, while still requiring `selectedAttempt=UNFILTERED_VECTOR_RETRY`, expected evidence source, context packing, and rerank top source.
Reason: Spring AI VectorStore and the Milvus SDK can differ on whether an over-filtered first pass returns a weak candidate or no candidate. The MVP contract is that the retriever skips only the L0 category filter, keeps `kb_scope`, retries the original query, and returns the correct evidence.
## D9. Default live snapshots to Spring AI VectorStore
Decision: `generate_rag_lookup_snapshots.ps1` defaults to `retrieval.vector-store.mode=spring`.
Reason: Spring AI VectorStore is the current framework path for the project and should be the default live verification route. SDK mode remains available through `-VectorStoreMode sdk` for comparison.
@@ -0,0 +1,87 @@
# Evidence: rag-eval-pipeline-closure
## Changed Assets
- `scripts/eval_rag_retrieval.py`
- `eval/rag-retrieval/cases/golden-cases.json`
- `eval/rag-retrieval/fixtures/*.json`
- `eval/rag-retrieval/reports/baseline.json`
- `eval/rag-retrieval/reports/baseline.md`
- `eval/rag-retrieval/README.md`
- `mvp/architecture/rag-eval-closure.md`
- `src/test/java/com/superbiz/agent/eval/RagLookupSnapshotGeneratorTest.java`
- `src/test/java/com/superbiz/agent/eval/RagEvalSeedImporterTest.java`
- `scripts/generate_rag_lookup_snapshots.ps1`
- `scripts/prepare_rag_eval_seed.ps1`
- `eval/rag-retrieval/seed-docs/*.md`
- `src/main/java/com/superbiz/agent/dto/Frontmatter.java`
- `src/main/java/com/superbiz/agent/dto/KnowledgeEntry.java`
- `src/main/java/com/superbiz/agent/service/FrontmatterParser.java`
- `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java`
- `src/main/java/com/superbiz/agent/service/DocumentManagementService.java`
- `src/main/java/com/superbiz/agent/service/VectorIndexService.java`
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
- `src/main/java/com/superbiz/agent/service/SpringAiVectorStoreSidecarService.java`
- `src/main/resources/application.yml`
## Baseline Result
```text
Evaluated 7 cases: passRate=1.0, recall@5=1.0, failed=0
```
## Regression Signals
The evaluator now fails on:
- missing expected source
- missing breadcrumb or evidence keyword
- non-`lookupResult` fixture
- selected attempt mismatch
- fallback reason mismatch
- fallback reason outside accepted values
- evidence status mismatch
- missing context source
- rerank top source mismatch
The snapshot generator now provides:
- real `LookupKnowledgeTool` invocation
- one fixture per golden case
- explicit opt-in through `rag.snapshot.enabled=true`
- optional post-generation baseline evaluation
- scoped retrieval through `retrieval.kb-scope=rag-eval`
- Spring AI VectorStore by default through `retrieval.vector-store.mode=spring`
- scoped L0 hints through the same `retrieval.kb-scope`
The seed importer now provides:
- canonical eval docs under `eval/rag-retrieval/seed-docs`
- real `DocumentManagementService` import/reindex
- stable `source`/`docId` metadata
- `kb_scope=rag-eval` isolation from local non-eval documents
- frontmatter stripping before chunk embedding
- an over-filter decoy seed doc for fallback-path evaluation
The diff now detects:
- aggregate pass/recall regression
- case pass regression
- hit-level regression
- first-rank regression
- selected attempt/fallback/evidence/rerank changes
## Final Spring Live Snapshot
```text
.\scripts\generate_rag_lookup_snapshots.ps1 -Fixtures <temp-fixtures> -RetrievedAt 2026-07-06T00:00:00Z
Evaluated 7 cases: passRate=1.0, recall@5=1.0, failed=0
```
Key fallback trace:
```text
selectedAttempt=UNFILTERED_VECTOR_RETRY
fallbackReason=filtered_vector_no_evidence
rerankTopSource=rag-l0-filter-fallback
```
@@ -0,0 +1,16 @@
# Acceptance: executor-evidence-output-contract
## Draft Acceptance
- [x] Issue exists: `mvp/issues/archived/executor-evidence-attribution-hallucination.md`.
- [x] OpenSpec change artifacts exist.
- [x] devflow tracking files exist.
- [x] OpenSpec validation passes.
- [x] Implementation updates Chat Executor prompt.
- [x] Implementation passes structured Executor output to Verifier.
- [x] Verifier prefers structured claims and still falls back safely.
- [x] Focused tests cover parsing, payload assembly, and unsupported confirmed claims.
## Notes
This project is currently in proposal/design stage. Runtime code is intentionally not changed yet.
@@ -0,0 +1,23 @@
# Brief: executor-evidence-output-contract
## Summary
Executor currently returns natural-language diagnosis answers that may mix confirmed evidence, runbook guidance, historical patterns, and unsupported inference. Verifier catches many unsupported facts, but only after extracting claims from prose.
This project defines a structured Executor evidence-attribution contract and updates the Verifier input/verification path to consume it.
## Goal
Make Chat Executor output machine-checkable so confirmed claims are explicitly bound to current-session evidence, while hypotheses and evidence gaps remain visibly separate.
## Scope
- Chat Executor prompt contract.
- Executor structured output parsing.
- Verifier payload extension.
- Chat Verifier prompt behavior.
- Focused tests/eval fixtures.
## Related OpenSpec
`openspec/changes/executor-evidence-output-contract/`
@@ -0,0 +1,71 @@
# Decisions: executor-evidence-output-contract
## sm-flow Progress
### Clarify
Entry summary: recent Chat diagnosis sessions are `LOW_CONFID` because Executor presents unsupported or weakly supported details as confirmed facts after successful tool calls.
Slug: `executor-evidence-output-contract`
Scale: standard. This affects prompts, verifier input assembly, parsing behavior, and tests, but does not require a database schema change.
### Context
Relevant history:
- `executor-action-memory-relevance`: Executor already has retrieval quality constraints and should avoid repeated `lookup_knowledge`.
- `chat-verifier-agent`: Verifier should not see intermediate reasoning; it receives explicit `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
- `evidence-trace-hardening`: evidence-bearing tools persist stable traces and no-evidence semantics.
- `modular-rag-pipeline`: `lookup_knowledge` exposes evidence blocks and context packs; L0 hints are not fact evidence.
Current code shape:
- `src/main/resources/prompts/chat-executor-prompt.md` is the Chat Executor prompt.
- `src/main/resources/prompts/executor-prompt.md` is for the AiOps flow and is not the target of this Chat change.
- `VerifierInputHook` currently builds a payload with `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
- Verifier prompt currently extracts facts from `executor_final_answer`.
### Grill
Question: Should Executor output only JSON or JSON plus readable answer?
Decision: use one JSON object containing both machine fields and `user_facing_answer`. This avoids losing a readable Chinese answer while giving Verifier structured claims.
Question: Should evidence binding use `chunk_id`?
Decision: no. Use generic binding fields because `query_logs` and `query_metrics` do not naturally expose RAG chunks.
Question: Should Verifier trust Executor-provided claims completely?
Decision: no. Verifier should verify structured claims first, then scan `user_facing_answer` for extra confirmed-sounding facts omitted from `claims`.
Question: What happens when Executor JSON is malformed?
Decision: preserve raw final answer, mark parse failure, and fall back to existing natural-language verification.
### Specify
OpenSpec artifacts:
- `proposal.md`: why and scope
- `design.md`: contract, verifier behavior, risks
- `specs/chat-verifier-agent/spec.md`: modified and added requirements
- `tasks.md`: implementation checklist
### Audit
Cross-artifact alignment:
- Issue describes evidence attribution hallucination.
- Proposal scopes the fix to Executor output and Verifier consumption.
- Design preserves existing verifier isolation.
- Spec adds observable behavior without changing database schema.
- Tasks remain implementation-oriented and unchecked.
Interface impact:
- Prompt/output contract: L2 internal Agent contract change.
- Verifier payload: L2 internal structured input extension.
- Database schema: no change.
- External HTTP API: no intended change.
@@ -0,0 +1,17 @@
# Evidence: executor-evidence-output-contract
## Repository Evidence
- `chat-executor-prompt.md` currently requires using real tool data but does not require a structured evidence-attribution output.
- `chat-verifier-prompt.md` currently extracts facts from `executor_final_answer` prose and compares them with `tool_trace_summary`.
- `VerifierInputHook` currently provides `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
- `openspec/specs/chat-verifier-agent/spec.md` already requires explicit verifier inputs, auditable evidence refs, fixed verdicts, and low-confidence handling.
- `openspec/specs/evidence-trace-hardening/spec.md` already distinguishes failed, no-hit, deduped, and successful evidence-tool traces.
## Runtime Evidence From Recent Sessions
Recent MySQL inspection showed repeated `LOW_CONFID` verifier results with many `no_evidence` facts. Typical unsupported claims included OOM, Full GC frequency, specific slow SQL timings, lock waits, and service-specific timeout details that were not supported by current-session tool traces.
## Design Evidence
This change preserves the previous design that Verifier should not inspect intermediate reasoning. The new structured output is still final Executor output, not hidden chain-of-thought.
@@ -0,0 +1,48 @@
# Acceptance: executor-gatekeeper-hook
## Implementation Result
Completed stage two of Executor Structured Output V2.
- Added `ExecutorGatekeeperService`.
- Added initial `schema.executor_v2` and `evidence.invocation_ref` rules.
- Added `gatekeeper_result` to Verifier payload.
- Stored `gatekeeper_result` in `VerifierContextHolder`.
- Persisted `gatekeeper_result` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- Updated Verifier prompt so Gatekeeper fail must not produce PASS.
- Added focused tests for schema failure, valid pass, fabricated invocation ids, tool name mismatch, hook payload, and persistence.
## Static Verification
- `cmd /c openspec validate executor-gatekeeper-hook`
- Result: passed.
- Coverage: OpenSpec change validity.
## Script Verification
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test`
- Result: passed.
- Coverage: 23 focused tests for Gatekeeper service, hook integration, and ChatService persistence.
## Browser / Manual Verification
Not run. This stage changes backend validation and audit behavior only.
## Unverified
- Full live application run with a real LLM.
- MySQL trace inspection after a real chat session.
- Excerpt similarity or phrase/utilization rules.
Reason: This phase intentionally covers deterministic schema and invocation-reference validation. Full live verification is better after Verifier V2 and Composer are implemented.
## Remaining Work
- Phase three: Verifier V2 `claim_checks`.
- Phase four: Composer final-answer generation.
- Phase five: eval fixtures and full audit closure.
## Archive Status
Devflow archive files created for stage two. OpenSpec archive is expected before moving to stage three.
@@ -0,0 +1,44 @@
# Brief: executor-gatekeeper-hook
## Background
Stage one of Executor Structured Output V2 changed Chat Executor output to `executor_evidence_v2`, removing final-expression fields from Executor. That made the output structured, but it did not yet prevent deterministic evidence attribution failures such as fabricated invocation ids, removed fields, empty evidence bindings, or mismatched tool names.
## Goal
Add a deterministic Gatekeeper between Executor output parsing and Verifier model execution.
The Gatekeeper should:
- Validate the initial Executor V2 schema.
- Validate `claims[].evidence_bindings[].source_invocation_ids` against current-session `tool_invocation` rows.
- Validate evidence binding `tool_name` against the persisted invocation tool name.
- Expose a small `gatekeeper_result` to Verifier and audit persistence.
## Scope
Included:
- New Gatekeeper validation service.
- `schema.executor_v2` initial rule.
- `evidence.invocation_ref` initial rule.
- `VerifierInputHook` payload integration.
- `VerifierContextHolder` storage.
- `ChatService` verifier evaluation persistence.
- Minimal verifier prompt update.
- Focused tests for Gatekeeper, hook payload, fabricated invocation ids, tool name mismatch, and persistence.
Excluded:
- No Executor retry on Gatekeeper failure.
- No excerpt similarity rule in this phase.
- No hallucination phrase or evidence utilization rule in this phase.
- No Verifier V2 `claim_checks`.
- No Composer.
- No database schema changes.
## OpenSpec
- Change: `openspec/changes/executor-gatekeeper-hook`
- Parent stage: `openspec/changes/archive/2026-07-07-executor-v2-output-contract`
@@ -0,0 +1,51 @@
# Decisions: executor-gatekeeper-hook
## Key Decisions
### Gatekeeper stays in VerifierInputHook
Decision: Gatekeeper is integrated inside `VerifierInputHook`, after Executor output parsing and before Verifier model execution.
Reason: The user explicitly chose to keep this version in the Verifier hook and not move validation into Executor hook. This preserves the current workflow orchestration.
### No retry in this phase
Decision: Gatekeeper failure does not trigger automatic Executor retry.
Reason: Retry behavior is intentionally deferred. This phase only validates, exposes, and audits deterministic failures.
### Initial rule set is intentionally small
Decision: Stage two implements only `schema.executor_v2` and `evidence.invocation_ref` as hard checks.
Reason: These rules catch the highest-confidence physical failures with low implementation risk. Excerpt similarity, hallucination phrases, and evidence utilization remain later enhancements.
### No new database schema
Decision: Persist `gatekeeper_result` in existing `diagnosis_session.self_evaluation.verifier_evaluation`.
Reason: The user asked to keep database fields minimal. Existing JSON audit storage is enough for this phase.
### Internal interface impact
Decision: This is an L2 internal interface extension.
Impact:
- Verifier payload gains `gatekeeper_result`.
- `VerifierContextHolder` gains Gatekeeper result storage.
- `self_evaluation.verifier_evaluation` gains `gatekeeper_result`.
- No external API, DTO, database table, or schema migration changes.
## Deferred Decisions
- Whether Gatekeeper should later trigger Executor retry.
- Whether `evidence.excerpt_similarity` should be hard fail or warn-only.
- Whether hallucination phrase and evidence utilization rules should be config-driven from metadata files.
- How Verifier V2 `claim_checks` should enforce Gatekeeper failures in code, beyond prompt instruction.
## Remaining Risks
- Verifier prompt compliance is not a deterministic guarantee; stage three should make Gatekeeper fail incompatible with PASS in Verifier V2 behavior.
- Excerpt authenticity is not checked in this phase, so real invocation ids can still be paired with misleading excerpt text until a later rule is implemented.
@@ -0,0 +1,54 @@
# Evidence: executor-gatekeeper-hook
## Context Evidence
- `executor-v2-output-contract` established `executor_evidence_v2` and removed Executor final-expression fields.
- `VerifierInputHook` is the existing integration point for explicit Verifier payload construction.
- `ChatService.persistVerifierEvaluation(...)` is the existing persistence path for verifier audit snapshots.
- `ToolInvocationRepository.findBySessionIdOrderByIdAsc(...)` provides the current-session invocation pool used by Gatekeeper.
## Implementation Evidence
- `src/main/java/com/superbiz/agent/service/ExecutorGatekeeperService.java`
- Implements `schema.executor_v2`.
- Implements `evidence.invocation_ref`.
- Returns `status`, `failed_rules`, `warnings`, and `errors`.
- `src/main/java/com/superbiz/agent/hook/VerifierInputHook.java`
- Runs Gatekeeper after parsing Executor output and building trace summary.
- Adds `gatekeeper_result` to Verifier payload.
- Stores `gatekeeper_result` in `VerifierContextHolder`.
- `src/main/java/com/superbiz/agent/util/VerifierContextHolder.java`
- Stores per-request Gatekeeper result for later persistence.
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- Wires `ExecutorGatekeeperService` into verifier hook construction.
- Persists `gatekeeper_result` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- `src/main/resources/prompts/chat-verifier-prompt.md`
- Documents `gatekeeper_result` as an input.
- States Gatekeeper fail must not produce PASS.
## Test Evidence
- `src/test/java/com/superbiz/agent/service/ExecutorGatekeeperServiceTest.java`
- Covers schema failure and valid pass behavior.
- Covers fabricated invocation ids and tool name mismatch.
- `src/test/java/com/superbiz/agent/hook/VerifierInputHookTest.java`
- Covers Verifier payload containing `gatekeeper_result`.
- Covers hook behavior for fabricated invocation ids.
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`
- Covers persistence of `gatekeeper_result` into verifier evaluation.
## Validation Evidence
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test`
- Result: passed.
- Coverage: 23 focused tests.
- `cmd /c openspec validate executor-gatekeeper-hook`
- Result: passed.
@@ -0,0 +1,41 @@
# Acceptance: executor-v2-output-contract
## Implementation Result
Completed stage one of Executor Structured Output V2.
- Executor prompt now emits `executor_evidence_v2`.
- Executor output no longer includes `diagnosis_summary` or `user_facing_answer`.
- ChatService PASS path renders V2 structured output into readable Chinese.
- VerifierInputHook remains parse-only and accepts V2 output without final-expression fields.
## Static Verification
- `cmd /c openspec validate executor-v2-output-contract`
- Result: passed.
- Coverage: OpenSpec syntax and change validity.
## Script Verification
- `mvn "-Dtest=VerifierInputHookTest,ChatServiceSequentialAgentTest" test`
- Result: passed.
- Coverage: VerifierInputHook V2 parsing; ChatService PASS rendering for V2; existing sequential workflow tests.
## Browser / Manual Verification
Not run. This stage changes backend prompt/runtime contract and unit-level behavior only.
## Unverified
- Full live application run with a real LLM.
- MySQL trace inspection after a real chat session.
Reason: Stage one is covered by focused unit tests; live verification is more useful after Gatekeeper and Composer phases.
## Remaining Work
- Phase two: Gatekeeper in `VerifierInputHook`.
- Phase three: Verifier V2 `claim_checks`.
- Phase four: Composer.
- Phase five: eval fixtures and full audit closure.
@@ -0,0 +1,32 @@
# Brief: executor-v2-output-contract
## Background
`executor_evidence_v1` still made Chat Executor produce both evidence attribution and final user-facing prose through `diagnosis_summary` and `user_facing_answer`.
This kept Executor in a "diagnose and narrate" role and left room for unsupported conclusions to appear before later verification and composition stages.
## Goal
Narrow Chat Executor output to `executor_evidence_v2`: structured diagnostic material only, with final expression removed from Executor.
## Scope
- Update Chat Executor prompt to emit `executor_evidence_v2`.
- Remove `diagnosis_summary` and `user_facing_answer` from Executor output.
- Keep `claims`, `hypotheses`, `recommended_actions`, `missing_info`, and evidence bindings.
- Add temporary ChatService rendering for PASS + V2 output so normal users do not see raw JSON.
- Preserve V1 `user_facing_answer` extraction for compatibility.
## Non-Goals
- No Gatekeeper implementation.
- No Verifier V2 `claim_checks`.
- No Composer.
- No Planner changes.
- No database schema changes.
- No evidence tool signature changes.
## Related OpenSpec
`openspec/changes/archive/2026-07-07-executor-v2-output-contract/`
@@ -0,0 +1,39 @@
# Decisions: executor-v2-output-contract
## Key Decisions
### Executor V2 removes final-expression fields
Decision: Chat Executor final output now uses `executor_evidence_v2` and must not include `diagnosis_summary` or `user_facing_answer`.
Reason: Executor should collect evidence and produce structured diagnostic material, not write final user-facing conclusions.
### Temporary renderer bridges the gap before Composer
Decision: `ChatService` renders V2 structured fields into readable Chinese only when Verifier returns `PASS`.
Reason: Composer is a later phase, but external users must not receive raw JSON during this intermediate stage.
### V1 compatibility remains
Decision: Existing V1 `user_facing_answer` extraction remains.
Reason: It keeps old tests and any lingering V1 output compatible while the staged migration continues.
### Gatekeeper and Verifier V2 are deferred
Decision: This phase does not add Gatekeeper or `claim_checks`.
Reason: The user requested phase-by-phase implementation with archive and commit after each phase. Gatekeeper is phase two.
## Interface Impact
- Internal Agent output contract: L4, because fields are removed.
- Verifier payload: L2, because raw `executor_final_answer` and parsed `executor_structured_output` remain.
- External Chat answer: compatible intent; users still get readable Chinese.
## Risks
- The temporary renderer is not a full Composer and should be replaced in the Composer phase.
- Verifier prompt still uses V1 `facts_checked`; Verifier V2 is a later phase.
@@ -0,0 +1,21 @@
# Evidence: executor-v2-output-contract
## Context Used
- `devflow/projects/2026-07-07-executor-evidence-output-contract`: V1 evidence-attribution contract kept `user_facing_answer`.
- `devflow/projects/2026-07-02-chat-verifier-agent`: Verifier consumes explicit inputs and should not see intermediate reasoning.
- `devflow/projects/2026-07-04-evidence-trace-hardening`: evidence summaries and tool invocation references are the evidence foundation.
- `mvp/issues/design-notes/executor-structured-output-v2.md`: staged implementation design; stage one is Executor V2 output contract.
## Code Evidence
- `src/main/resources/prompts/chat-executor-prompt.md`: V2 contract now uses `answer_version="executor_evidence_v2"` and removes final-expression fields.
- `src/main/java/com/superbiz/agent/service/ChatService.java`: PASS path now tries V1 `user_facing_answer`, then renders V2 structured output to readable Chinese.
- `src/main/resources/prompts/chat-verifier-prompt.md`: `user_facing_answer` is now described as compatibility-only.
- `src/test/java/com/superbiz/agent/hook/VerifierInputHookTest.java`: V2 structured output without `user_facing_answer` parses successfully.
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`: PASS + V2 output renders Chinese and does not expose raw JSON.
## Key Finding
The previous V1 contract intentionally kept `user_facing_answer`, but the V2 staged design intentionally removes it. This is an internal Agent contract break, mitigated by a temporary renderer until Composer is implemented.
@@ -0,0 +1,58 @@
# Acceptance
## Implementation Result
Implemented stage three of Executor Structured Output V2:
- Verifier prompt now validates claim derivability rather than scanning final natural-language output.
- Verifier output supports `claim_checks`.
- `ChatService` derives compatibility `facts_checked` from `claim_checks`.
- `ChatService` persists both `claim_checks` and `facts_checked`.
- Effective verdict guardrails prevent Gatekeeper failures and malformed structured output from remaining `PASS`.
- Verifier logging summarizes `claim_checks`.
## Static Verification
- Reviewed `git diff --stat` and changed files are scoped to stage three implementation, tests, OpenSpec/devflow, and the issue handoff document.
- `cmd /c openspec validate executor-verifier-claim-checks` passed.
- After OpenSpec archive, `cmd /c openspec validate --specs` passed.
## Script Verification
Passed:
```powershell
mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
```
Coverage:
- Verifier payload and Gatekeeper hook behavior.
- Gatekeeper schema/invocation validation.
- `claim_checks` parsing and persistence.
- `claim_checks` to `facts_checked` compatibility mapping.
- all claim verification mapping classes.
- Gatekeeper failure downgrade from model `PASS`.
- malformed Executor output downgrade from model `PASS`.
The targeted Maven test command was re-run after OpenSpec archive and passed.
## Browser / Manual Verification
Not run. This stage changes backend prompt, parser, audit, and tests only.
## OpenSpec Archive Status
Archived:
```text
openspec/changes/archive/2026-07-07-executor-verifier-claim-checks
```
Archive follow-up: the generated canonical `chat-verifier-agent` spec was reviewed and amended to preserve pre-existing verifier input and Gatekeeper schema scenarios while adding the new claim-check scenarios.
## Remaining Risks
- Composer is not implemented in this stage; final PASS rendering still uses the temporary V2 renderer until stage four.
- Full eval fixture expansion is deferred to stage five.
- Existing Maven warnings about duplicate test dependency and Lombok builder defaults remain outside this stage.
@@ -0,0 +1,41 @@
# Executor Verifier Claim Checks
## Background
Stage one moved Chat Executor to `executor_evidence_v2`, and stage two added deterministic Gatekeeper checks before Verifier. After those stages, Verifier still primarily used the legacy `facts_checked` contract and could still treat `executor_final_answer` as a fact source.
That left two risks:
- Verifier could still extract extra confirmed facts from natural-language Executor output.
- Downstream audit and retry consumers could not distinguish V2 claim-level verification from legacy fact checks.
## Goal
Make Verifier V2 claim-oriented:
- verify `executor_structured_output.claims` as the primary target;
- emit `claim_checks` as the authoritative V2 result;
- keep `facts_checked` only as a compatibility projection;
- enforce code-side guardrails so Gatekeeper failures or malformed structured output cannot remain effective `PASS`.
## Scope
- Updated `chat-verifier-prompt.md` to frame verification as claim derivability.
- Extended `ChatService` to parse, normalize, map, and persist `claim_checks`.
- Added effective verdict guardrails for Gatekeeper failure and malformed/missing Executor structured output.
- Updated verifier logging summaries to count `claim_checks`.
- Updated sequential workflow tests to cover claim mapping and downgrade behavior.
## Non-Goals
- No Composer integration in this phase.
- No final-answer material filtering beyond existing templates and temporary V2 renderer.
- No Executor retry behavior change.
- No Gatekeeper rule expansion.
- No database schema migration.
## OpenSpec
- Active change before archive: `openspec/changes/executor-verifier-claim-checks`
- Capability: `chat-verifier-agent`
- Scale: standard
@@ -0,0 +1,57 @@
# Decisions
## Scope Decision
Stage three is limited to Verifier V2 claim checks. Composer is explicitly deferred to stage four.
Reason: Composer requires stable verifier output and allowed-material filtering; mixing it into this stage would make rollback and acceptance unclear.
## Contract Decision
`claim_checks` is the authoritative V2 verifier output.
`facts_checked` remains as a compatibility projection generated from `claim_checks` when present.
Reason: existing low-confidence rendering, retry context, trace output, and evaluation code still depend on `facts_checked`.
## Mapping Decision
Claim verification maps to legacy facts as follows:
| claim verification | legacy facts_checked verification |
|---|---|
| `direct_observation` | `direct_evidence` |
| `reasonable_inference` | `indirect_support` |
| `overstated` | `indirect_support` |
| `unsupported` | `no_evidence` |
| `external_unknown` | `no_evidence` |
| `contradicted` | `contradicted` |
## Guardrail Decision
Effective verdict is enforced in code:
- `gatekeeper_result.status=fail` cannot remain `PASS`.
- `evidence.invocation_ref` failure downgrades to `REJECT`.
- Other Gatekeeper failures downgrade at least to `LOW_CONFID`.
- missing/malformed Executor structured output cannot remain `PASS`.
Reason: prompt compliance is not deterministic enough for safety-critical evidence attribution.
## Apply Fix Record
Initial targeted Maven verification failed because older tests expected PASS to return Executor natural-language output or V1 `user_facing_answer`.
Classification: test drift from the committed OpenSpec, not a design blocker.
Resolution: update tests to use valid Executor V2 output for PASS paths and assert downgrade behavior for malformed or Gatekeeper-failed outputs.
## Interface Impact
L2 internal contract extension:
- Verifier output gains `claim_checks`.
- Existing `facts_checked` remains available.
- Persistence JSON gains `claim_checks` under existing `self_evaluation`.
No external API or database schema changes.
@@ -0,0 +1,35 @@
# Evidence
## Relevant History
- `executor-v2-output-contract`: Executor emits `executor_evidence_v2` and no longer emits final-expression fields.
- `executor-gatekeeper-hook`: Gatekeeper validates schema and invocation references before Verifier and persists `gatekeeper_result`.
- `chat-verifier-agent`: Existing Verifier used `facts_checked`, low-confidence rendering, retry context, and verifier audit.
## Code Evidence
- `src/main/resources/prompts/chat-verifier-prompt.md`: Verifier prompt now makes `executor_structured_output.claims` primary and treats `executor_final_answer` as debug/fallback only.
- `src/main/java/com/superbiz/agent/service/ChatService.java`: parses `claim_checks`, maps them to compatibility `facts_checked`, persists both, and applies effective verdict guardrails.
- `src/main/java/com/superbiz/agent/hook/AgentLoggingHook.java`: verifier thought summaries now include `claim_checks`.
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`: covers claim-check mapping, Gatekeeper downgrade, malformed output downgrade, and V2 PASS paths.
## Evidence-Driven Conclusions
- `facts_checked` cannot be removed yet because existing low-confidence templates, retry context, trace tooling, and eval paths still consume it.
- Prompt-only prevention is insufficient for Gatekeeper failures; `ChatService` must enforce effective verdict downgrades in code.
- No database schema migration is needed because `claim_checks` is persisted inside existing `diagnosis_session.self_evaluation`.
- Composer remains stage four and must not be mixed into this stage.
## Verification Evidence
Script verification passed:
```powershell
mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
cmd /c openspec validate executor-verifier-claim-checks
```
Known existing warnings:
- Maven reports duplicate `spring-boot-starter-test` dependency in `pom.xml`.
- Existing Lombok `@Builder` default warnings remain.
@@ -0,0 +1,48 @@
# diagnosis-eval-demo-gatekeeper-closure Acceptance
## Static / Structure Verification
- `cmd /c openspec validate diagnosis-eval-demo-gatekeeper-closure --strict`
- Result: passed.
- `cmd /c openspec validate --specs`
- Result: passed, 10 specs passed.
## Script Verification
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,VerifierInputHookTest" test`
- Result: 36 tests, 0 failures, 0 errors.
- `mvn "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest,ToolInvocationRecorderTest,QueryLogsToolsTest" test`
- Result: 61 tests, 0 failures, 0 errors.
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`
- Result after E2E startup fix: 23 tests, 0 failures, 0 errors.
## Live E2E Verification
- Start command: `mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"`.
- Demo command: `powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1`.
- Result: chat, trace, and feedback requests completed successfully.
- Output files:
- `mvp/demo/output/chat-response.json`
- `mvp/demo/output/trace-response.json`
- `mvp/demo/output/feedback-response.json`
- Trace observations:
- `hasVerifierEvaluation=true`
- `gatekeeper_result.rule_set_version=gatekeeper-rules-v1`
## Fixed During Verification
- E2E startup initially failed because Spring could not instantiate `ExecutorGatekeeperService`.
- Root cause: two public constructors and no explicit `@Autowired` constructor.
- Fix: annotate the production constructor with `@Autowired`.
## Residual Risk
- The live payment-timeout path can still produce `LOW_CONFID` because model-generated evidence bindings may omit some explicit `source_invocation_id` values.
- This is not a blocker for this change because deterministic matrix behavior is covered by saved fixtures and baseline evaluation.
- Existing Maven warnings remain: duplicate `spring-boot-starter-test` declaration and Lombok `@Builder` default warnings.
## Archive Status
- Devflow archive artifacts created.
- OpenSpec change archived to `openspec/changes/archive/2026-07-08-diagnosis-eval-demo-gatekeeper-closure`.
- Main specs synced by `cmd /c openspec archive diagnosis-eval-demo-gatekeeper-closure --yes`.
@@ -0,0 +1,33 @@
# diagnosis-eval-demo-gatekeeper-closure Brief
## Background
The Chat evidence pipeline already had Executor V2 structured output, deterministic Gatekeeper validation, Verifier claim checks, and Composer final rendering. The missing piece was an interview-ready acceptance story that made the anti-hallucination behavior easy to demonstrate and regress.
## Goal
Close the next three interview-readiness gaps together:
- diagnosis eval fixture matrix
- stable demo data set
- Gatekeeper rule configuration and audit version
## Scope
- Expand `mvp/eval` with matrix-oriented cases, fixtures, and baseline reports.
- Add stable demo request payloads and scenario documentation.
- Add a lightweight local Gatekeeper rule catalog with `rule_set_version` and rule metadata in `gatekeeper_result`.
- Update architecture, demo, and eval docs to describe the current implementation.
## Non-goals
- No new public HTTP endpoint.
- No new database table.
- No Planner `scope_contract`.
- No Gatekeeper retry loop.
- No remote or dynamic rule execution engine.
## OpenSpec
- Change: `openspec/changes/diagnosis-eval-demo-gatekeeper-closure`
- Interface impact: L2 internal contract change.
@@ -0,0 +1,144 @@
# diagnosis-eval-demo-gatekeeper-closure Decisions
## Clarify
- Entry summary: implement the next three interview-readiness items together: diagnosis eval fixture matrix, stable demo data set, and Gatekeeper rule configuration/audit version.
- Slug: `diagnosis-eval-demo-gatekeeper-closure`
- Devflow scale: `standard`
- Interface impact: expected L2 internal contract change because `gatekeeper_result` audit JSON will gain rule metadata/version fields.
## Context
- `devflow/index.md` used: related entries found for diagnosis eval harness, fixture expansion, MVP demo runbook, Gatekeeper hook, and verifier evidence reference fidelity.
- Relevant glossary:
- Evidence Tools produce incident facts and must be recorded in `tool_invocation`.
- Verifier should not use skills/runbooks as incident evidence.
- `tool_invocation.retrieval_details` is the structured evidence/audit home for tool-specific details.
- Historical constraints that must enter OpenSpec:
- Diagnosis eval is offline and deterministic; no LLM-as-judge.
- Demo assets should be runnable, but fixed regression should use saved fixtures.
- Gatekeeper remains in the Verifier hook path.
- No new database table for Gatekeeper audit; use `self_evaluation.verifier_evaluation.gatekeeper_result`.
- `$.no_evidence` is a query no-hit signal, not proof that a problem is impossible.
## Question Pool
| ID | Dimension | Mode | Question | Status |
|---|---|---|---|---|
| Q1 | Terminology | evidence-driven | What names should this change use for the matrix, demo set, and Gatekeeper rule metadata? | Resolved |
| Q2 | Boundary | evidence-driven | Should this change alter public APIs, database schema, Planner output, or retry behavior? | Resolved |
| Q3 | Acceptance | evidence-driven | Which existing tests and baseline assets define the current acceptance style? | Resolved |
| Q4 | Technical | evidence-driven | Where should Gatekeeper rule metadata live with minimal implementation risk? | Pending code research |
| Q5 | Scope | user-interview | Should the stable demo set be documentation/payloads only, or should it include live E2E scripts for all scenarios? | Confirmed |
## Evidence-driven Conclusions
- Q1 conclusion: use `diagnosis eval matrix`, `stable demo scenarios`, and `Gatekeeper rule set version` as terms.
- Q2 conclusion: keep this as an internal contract change. Do not add public endpoints, tables, Planner `scope_contract`, or Gatekeeper retry.
- Q3 conclusion: existing `DiagnosisTraceEvaluatorTest`, `ExecutorGatekeeperServiceTest`, `VerifierInputHookTest`, `ToolInvocationRecorderTest`, and `mvp/eval/reports` define the current acceptance style.
- Q4 conclusion: Gatekeeper metadata should live behind a small rule catalog loaded by `ExecutorGatekeeperService`; the audit output should include a rule set version and enabled rule metadata summary, without adding tables or remote registry.
## User-interview Confirmations
- Q5 confirmed by resumed objective: complete items 1/2/3 with sm-flow, archive, submit, and run end-to-end if necessary.
- Implementation interpretation: stable demo scenarios will be fixed request payloads and runbook docs plus deterministic fixture-backed eval. Live E2E remains necessary only for at least one main path or where unit/fixture evidence is insufficient.
## OpenSpec Backfill
- Created Draft proposal at `openspec/changes/diagnosis-eval-demo-gatekeeper-closure/proposal.md`.
- Context constraints from historical devflow entries were written into the proposal.
- Scope confirmation and Gatekeeper catalog placement were written into the proposal/design.
## Current Checkpoint
- Discover completed.
- No implementation files changed yet.
## Specify / Alignment
### Cross-artifact Alignment
| Check | Status | Notes |
|---|---|---|
| brief/proposal goals -> proposal | Aligned | Proposal covers eval matrix, stable demo scenarios, and Gatekeeper rule catalog/audit version. |
| proposal scope/constraints -> design | Aligned | Design records offline deterministic eval, fixture-backed demo distinction, local rule catalog, and no new table/API. |
| design decisions -> specs/tasks | Aligned | Specs cover eval matrix, rule set version validation, demo scenarios, and Gatekeeper rule metadata; tasks cover matching implementation slices. |
| specs observable behavior -> tasks | Aligned | Each requirement has an executable task and acceptance check. |
### Interface Impact
- Level: L2 internal contract change.
- Reason: `gatekeeper_result` internal audit JSON gains `rule_set_version` and rule metadata summary. Eval case/result fields may gain optional rule set checks. No public HTTP API, database schema, or external DTO contract changes.
## Audit
Input -> processing -> output chain:
```text
mvp/demo request docs + mvp/eval fixtures
-> DiagnosisTraceEvaluator
-> baseline reports
-> interview/demo evidence
Gatekeeper rule catalog
-> ExecutorGatekeeperService
-> VerifierInputHook / ChatService persisted self_evaluation
-> Trace and eval audit
```
Architecture risk assessment:
1. The change is intentionally internal and should not add new public consumers.
2. Gatekeeper catalog must stay metadata-only; dynamic rule execution would be a different, riskier architecture.
3. Fixture-backed demo scenarios should be documented as deterministic regression artifacts, not live LLM guarantees.
4. Baseline report churn is expected and must be committed with case/fixture changes.
5. No devflow/OpenSpec conflict found.
## Commit Gate
- `cmd /c openspec validate diagnosis-eval-demo-gatekeeper-closure --strict`: passed.
- `cmd /c openspec validate --specs`: passed, 10 specs passed.
- File completeness:
- proposal.md: present.
- design.md: present.
- specs: present for `diagnosis-eval-harness`, `mvp-demo-trace-acceptance`, `chat-verifier-agent`.
- tasks.md: present.
- Consistency:
- Proposal concepts have corresponding design sections.
- Design decisions are reflected in specs/tasks.
- Task acceptance checks are verifiable.
## Current Checkpoint
- Commit completed.
- `.committed` marker created.
## Apply Verification
- Focused verification passed:
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,VerifierInputHookTest" test`
- Result: 36 tests, 0 failures, 0 errors.
- Broader relevant regression passed:
- `mvn "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest,ToolInvocationRecorderTest,QueryLogsToolsTest" test`
- Result: 61 tests, 0 failures, 0 errors.
- E2E startup repro found a Spring bean construction issue:
- Command: `mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"`
- Failure: `ExecutorGatekeeperService` had two public constructors and no annotated constructor, so Spring attempted a no-arg constructor and failed with `No default constructor found`.
- Classification: code deviation from OpenSpec implementation intent, not a spec gap.
- Fix: annotate the production constructor with `@Autowired`.
- Post-fix focused regression passed:
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`
- Result: 23 tests, 0 failures, 0 errors.
- Live E2E passed for demo compatibility:
- Start: `mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"`
- Run: `powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1`
- Result: `/api/chat`, `/api/diagnosis/{sessionId}/trace`, and `/api/feedback` completed successfully.
- Trace summary included `hasVerifierEvaluation=true`.
- Persisted Gatekeeper audit included `rule_set_version=gatekeeper-rules-v1`.
- Residual quality note: the live payment-timeout response remained `LOW_CONFID` because some model-produced evidence bindings still lacked explicit `source_invocation_id`; deterministic PASS/LOW_CONFID/REJECT claims are covered by fixture-backed eval.
## Archive Readiness
- OpenSpec tasks 1-4 completed.
- Verification is recorded in devflow acceptance artifacts.
- Remaining known risk: live LLM output is not deterministic and may still produce LOW_CONFID on the payment-timeout path; this is intentionally documented as demo compatibility, not a fixed PASS guarantee.
@@ -0,0 +1,22 @@
# diagnosis-eval-demo-gatekeeper-closure Evidence
## Code And Artifact Evidence
- Gatekeeper rule metadata lives in `src/main/resources/gatekeeper/gatekeeper-rules.json`.
- `ExecutorGatekeeperService` loads the local catalog, uses configured threshold parameters, and emits `rule_set_version` plus enabled rule metadata.
- `VerifierInputHook` and `ChatService` preserve Gatekeeper audit metadata in fallback/default paths.
- `DiagnosisTraceEvaluator` can optionally validate expected Gatekeeper rule set version.
- `mvp/eval/cases/diagnosis-cases.json` now includes narrow-scope and no-evidence matrix cases.
- `mvp/eval/reports/baseline-report.json` and `.md` were regenerated for the expanded fixed matrix.
- `mvp/demo/evidence-pipeline-scenarios.md` documents live vs fixture-backed demo scenarios.
## Decisions
- Keep this phase internal: no public API, no DB schema, no Planner output change.
- Keep Gatekeeper deterministic Java validation; the catalog is metadata/config only.
- Treat live demo as compatibility evidence and fixture-backed eval as deterministic regression evidence.
- Persist audit under the existing `self_evaluation.verifier_evaluation.gatekeeper_result` structure.
## Runtime Finding
The first Maven E2E startup found a real integration issue: `ExecutorGatekeeperService` had multiple public constructors without an annotated constructor, so Spring could not instantiate the service. The fix was to annotate the production constructor with `@Autowired`.
@@ -0,0 +1,57 @@
# Acceptance
## Implementation Result
Implemented stage four of Executor Structured Output V2:
- Added `chat_composer` prompt and Agent.
- Final answers for PASS, LOW_CONFID, and REJECT now use Composer when Verifier decision is valid.
- Composer input is filtered from Verifier decision and Executor structured output.
- Unsupported, external-unknown, and contradicted claims are excluded from confirmed final-answer material.
- REJECT Composer input has `allowed_hypotheses=[]`.
- Malformed Composer output uses deterministic safe fallback.
- Fallback does not expose raw Composer JSON, raw Executor JSON, or Executor `user_facing_answer`.
- `composer_output` is persisted in verifier audit.
## Static Verification
- Reviewed implementation diff for stage-four scope.
- `cmd /c openspec validate executor-composer-final-answer` passed.
- `cmd /c openspec validate --specs` passed before archive.
## Script Verification
Passed:
```powershell
mvn "-Dtest=ChatServiceSequentialAgentTest" test
mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
```
Coverage:
- Composer prompt loading and invocation.
- valid Composer output as final answer source.
- malformed Composer output fallback.
- PASS no raw Executor JSON leakage.
- LOW_CONFID separation of confirmed material, possible directions, and gaps.
- REJECT safe output without raw Executor answer.
- Gatekeeper and Verifier stage compatibility.
## Browser / Manual Verification
Not run. This stage changes backend prompt, routing, parser, audit, and tests only.
## OpenSpec Archive Status
Archived:
```text
openspec/changes/archive/2026-07-08-executor-composer-final-answer
```
## Remaining Risks
- Stage five still needs broader eval fixture coverage for full evidence-attribution regressions.
- Composer prompt quality can be improved after real run traces are collected.
- Existing Maven warnings about duplicate test dependency and Lombok builder defaults remain outside this stage.
@@ -0,0 +1,54 @@
# Executor Composer Final Answer
## Background
Stages one to three moved the Chat diagnosis chain to structured Executor output, deterministic Gatekeeper validation, and Verifier `claim_checks`.
Before this stage, `ChatService` still owned final answer rendering. PASS paths could use a temporary V2 renderer, while LOW_CONFID and REJECT paths used templates. That left final user-facing expression too close to Executor material and made it harder to prove that only Verifier-allowed claims reached the user.
## Goal
Add a Composer expression layer after Verifier:
```text
chat_planner
-> chat_executor
-> VerifierInputHook + Gatekeeper
-> chat_verifier
-> chat_composer
-> final answer
```
Composer produces user-facing answers from filtered material only:
- `allowed_claims`
- `allowed_hypotheses`
- `missing_info`
- `recommended_actions`
- `rationale`
## Scope
- Added `chat-composer-prompt.md`.
- Added `chat_composer` Agent construction in `ChatService`.
- Added Composer input filtering from Verifier decision and Executor structured output.
- Replaced PASS temporary V2 renderer usage with Composer-or-safe-fallback rendering.
- Routed LOW_CONFID and REJECT final answers through Composer when Verifier output is valid.
- Added deterministic fallback for malformed Composer output.
- Persisted `composer_output` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- Updated sequential workflow tests.
## Non-Goals
- No Planner changes.
- No Executor retry changes.
- No Gatekeeper rule expansion.
- No Verifier classification expansion.
- No database schema migration.
- No stage-five eval fixture expansion.
## OpenSpec
- Active change before archive: `openspec/changes/executor-composer-final-answer`
- Capabilities: `chat-composer-agent`, `chat-verifier-agent`
- Scale: standard
@@ -0,0 +1,70 @@
# Decisions
## Scope Decision
Stage four is limited to Composer final-answer generation and routing.
Reason: Gatekeeper and Verifier contracts were stabilized in earlier stages; this phase should only close the final-expression path.
## Composer Responsibility
Composer is an expression layer, not a diagnosis layer.
It may rephrase and organize only filtered material. It must not call tools, introduce new facts, rejudge root cause, or read raw Executor/tool output.
## Filtering Decision
`ChatService` owns Composer input filtering:
| Verifier classification | Composer handling |
|---|---|
| `direct_observation` | `allowed_claims` |
| `reasonable_inference` | `allowed_claims`, with bounded wording |
| `overstated` | `allowed_hypotheses` or `missing_info` |
| `unsupported` | `missing_info` |
| `external_unknown` | `missing_info` |
| `contradicted` | `missing_info` / REJECT-safe output |
For REJECT, `allowed_hypotheses` is always empty.
## Fallback Decision
Malformed Composer output falls back to deterministic rendering from filtered Composer input.
Fallback must never expose:
- raw Composer JSON;
- raw Executor JSON;
- Executor `user_facing_answer`;
- full unscreened tool output.
## Audit Decision
No new table is added. Composer output is persisted under:
```text
diagnosis_session.self_evaluation.verifier_evaluation.composer_output
```
The audit snapshot is intentionally compact and stores status plus parsed user-facing fields.
## Apply Fix Record
Initial targeted verification exposed test drift:
- test file had a UTF-8 BOM and failed Java compilation;
- scripted chat model did not recognize `COMPOSER_TEST_PROMPT`;
- older tests expected three-Agent execution and temporary V2 renderer behavior;
- LOW_CONFID assertions required indirect support to disappear instead of appearing as a possible direction.
Resolution: remove BOM, add Composer script branch, and update assertions to match the committed Composer contract.
## Interface Impact
L2 internal behavior change:
- external Chat API still returns a final answer string;
- internal final-answer source changes from Executor/temporary renderer to Composer or safe fallback;
- audit JSON gains `composer_output` under existing `self_evaluation`.
No database schema change.
@@ -0,0 +1,36 @@
# Evidence
## Relevant History
- `executor-v2-output-contract`: Executor emits structured diagnostic material and no final-expression fields.
- `executor-gatekeeper-hook`: Gatekeeper validates deterministic evidence failures before Verifier.
- `executor-verifier-claim-checks`: Verifier emits `claim_checks` and effective verdict guardrails.
## Code Evidence
- `src/main/resources/prompts/chat-composer-prompt.md`: defines Composer as an expression layer with strict JSON output.
- `src/main/java/com/superbiz/agent/service/ChatService.java`: loads Composer prompt, invokes `chat_composer`, filters Composer input, parses Composer output, falls back safely, and persists Composer audit.
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`: covers Composer invocation, fallback, REJECT/LOW_CONFID behavior, and no raw JSON leakage.
## Evidence-Driven Conclusions
- Composer must be after Verifier because Verifier `claim_checks` are the authority for allowed final-answer material.
- Composer must not receive raw tool output or full unscreened Executor output because that would re-open the evidence attribution problem.
- Verifier malformed/missing output should not invoke Composer because there is no trustworthy decision to filter with.
- Fixed fallback remains necessary because Composer is an LLM call with a strict JSON contract and can produce malformed output.
## Verification Evidence
Passed:
```powershell
mvn "-Dtest=ChatServiceSequentialAgentTest" test
mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
cmd /c openspec validate executor-composer-final-answer
cmd /c openspec validate --specs
```
Known existing warnings:
- Maven reports duplicate `spring-boot-starter-test` dependency in `pom.xml`.
- Existing Lombok `@Builder` default warnings remain.
@@ -0,0 +1,52 @@
# Acceptance
## Static Verification
- `openspec validate verifier-evidence-reference-fidelity --strict`: passed.
- `openspec validate --specs`: passed.
## Script Verification
- `mvn "-Dtest=ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,QueryLogsToolsTest,ChatServiceSequentialAgentTest" test`
- Passed: 38 tests.
- `mvn "-Dtest=ExecutorGatekeeperServiceTest" test`
- Passed: 9 tests.
- `mvn test`
- Failed on unrelated environment-gated `MilvusConnectionTest.connect`: `MILVUS_TOKEN` environment variable was not set.
- Other executed tests in the run progressed until that single failure; focused tests for this change passed.
## End-to-End Verification
The Java service was restarted with `mvn spring-boot:run`; logs were written under `logs/`.
| Case | Session | Result | Gatekeeper Audit |
|---|---|---|---|
| HikariCP positive | `iss007-hikari-positive-20260708-1553` | PASS; confirmed order-service HikariCP timeout and pool saturation logs | `pass / none`, checked_bindings=2 |
| HikariCP negative | `iss007-hikari-negative-20260708-1555` | LOW_CONFID; no `generic-service`; no false positive for inventory-service | `fail / low_confid` |
| HighMemoryUsage positive | `iss007-memory-positive-20260708-1558` | PASS; confirmed HighMemoryUsage 91%, did not confirm memory leak | `pass / none`, checked_bindings=1 |
| SlowResponse positive | `iss007-slow-positive-20260708-1600` | PASS; confirmed SlowResponse and slow request logs, no DB pool root cause | `pass / none`, checked_bindings=7 |
| Narrow HighCPUUsage | `iss007-narrow-highcpu-20260708-1602` | PASS; only covered payment-service HighCPUUsage | `pass / none`, checked_bindings=1 |
## Database Audit
`scripts/query_mysql.py` was used to verify:
- `diagnosis_session.self_evaluation.verifier_evaluation.verdict`
- `gatekeeper_result.status`
- `gatekeeper_result.severity`
- `gatekeeper_result.checked_bindings`
- no-hit HikariCP query rows persist `evidence_status=no_evidence`
## Remaining Risk
- Negative no-hit claims still have incomplete precise references when Executor uses `$.logs` for empty arrays. Gatekeeper correctly downgrades to `LOW_CONFID`.
- Prompt-only scope control is improved but not a hard contract. A future `scope_contract` may still be needed.
- Full test suite requires `MILVUS_TOKEN` to pass `MilvusConnectionTest`.
## OpenSpec Archive
- `openspec archive verifier-evidence-reference-fidelity --yes`: succeeded.
- Main specs updated:
- `openspec/specs/chat-verifier-agent/spec.md`
- `openspec/specs/evidence-trace-hardening/spec.md`
- Non-blocking warning: proposal did not use OpenSpec's preferred `## Why` / `## What Changes` headers, but archive completed.
@@ -0,0 +1,44 @@
# Verifier Evidence Reference Fidelity
## Background
ISS-007 came from end-to-end diagnosis cases where raw tool output and Executor `evidence_excerpt` contained enough facts, but Verifier still returned `LOW_CONFID` because the verifier-facing summary compressed away key details.
The affected flow is:
```text
chat_planner
-> chat_executor
-> VerifierInputHook / Gatekeeper
-> chat_verifier
-> chat_composer
```
The change hardens the evidence handoff between Executor, Gatekeeper, and Verifier.
## Goal
Make Executor cite concrete tool evidence, make Gatekeeper validate that citation with code, and make Verifier judge whether verified evidence can derive the claim.
## Scope
- Persist `tool_invocation.retrieval_details.evidence_refs`.
- Use `source_invocation_id + raw_path + evidence_excerpt` as the precise evidence binding.
- Add Gatekeeper `severity` and checked binding audit.
- Keep Gatekeeper in the Verifier hook path.
- Keep `tool_trace_summary` as navigation/audit context, not the only evidence source.
- Fix HikariCP mock positive/no-hit behavior.
- Tighten Executor/Verifier prompts for narrow-scope evidence handling.
## Non-goals
- No Planner `scope_contract`.
- No new database table.
- No full JSONPath engine.
- No change to external HTTP API.
- No retry rollback from Gatekeeper to Executor in this phase.
## OpenSpec
- Change: `openspec/changes/verifier-evidence-reference-fidelity`
- Source issue: `mvp/issues/archived/ISS-007-verifier-evidence-summary-fidelity.md`
@@ -0,0 +1,54 @@
# Decisions
## Evidence Reference
Use `source_invocation_id + raw_path + evidence_excerpt` as the precise evidence reference for Executor claim bindings.
Reason:
- Invocation ID alone only identifies a tool call, not the evidence inside it.
- `raw_path` is enough for the first version when paired with `retrieval_details.evidence_refs`.
- `evidence_excerpt` remains the text Verifier reads, but only after Gatekeeper validates it.
## Raw Path
Only support stable locators in the first version:
- `$.alerts[i]`
- `$.logs[i]`
- `$.evidence_blocks[i]`
No full JSONPath engine is introduced.
## Gatekeeper Severity
Gatekeeper output includes:
- `status`
- `severity`
- `checked_bindings`
- `failed_rules`
- `warnings`
- `errors`
Severity meaning:
- `none`: precise references passed.
- `low_confid`: evidence is missing or incomplete, but not fabricated.
- `reject`: fabricated ID, wrong tool, unknown raw path, or mismatched excerpt.
## Verifier Boundary
Verifier uses verified claim-local excerpts as primary derivability evidence. `tool_trace_summary` remains available for navigation and audit, but no longer needs to carry every concrete fact.
## Hook Placement
Gatekeeper remains in the Verifier input hook path. This version does not retry Executor on Gatekeeper failure.
## Planner
Planner is not changed. `scope_contract` remains a later-stage idea. This phase uses prompt constraints to reduce narrow-scope over-expansion.
## Database
No new tables. Evidence refs are stored in `tool_invocation.retrieval_details.evidence_refs`; audit is stored in `diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result`.
@@ -0,0 +1,33 @@
# Evidence
## Existing Context
- Existing `chat-verifier-agent` spec still used `source_invocation_ids` and `tool_trace_summary` as the main verifier evidence context.
- Existing `evidence-trace-hardening` spec already established `tool_invocation.retrieval_details` as the right place for structured tool-specific facts.
- Prior devflow projects established that runbook/skill content is guidance, not incident evidence.
## Code Findings
- `VerifierInputHook` previously backfilled plural `source_invocation_ids` from `tool_trace_summary` by tool name.
- `ExecutorGatekeeperService` previously validated invocation existence and tool name, but not `raw_path` or excerpt authenticity.
- `ToolInvocationRecorder` persisted retrieval details but did not generate claim-addressable `evidence_refs`.
- `QueryLogsTools` could fall back to `generic-service` placeholder logs on no-hit.
## Implementation Evidence
- `ToolInvocationRecorder` now extracts:
- `$.alerts[i]` for `query_metrics`
- `$.logs[i]` for `query_logs`
- `$.evidence_blocks[i]` for `lookup_knowledge`
- `ExecutorGatekeeperService` now validates:
- invocation existence
- tool name
- raw path presence
- `retrieval_details.evidence_refs`
- excerpt similarity/support
- `VerifierInputHook` only auto-fills a singular `source_invocation_id` when exactly one candidate exists and never invents `raw_path`.
- `QueryLogsTools` returns HikariCP mock logs for `order-service` and returns empty no-hit results for unrelated services.
## Residual Finding
The HikariCP negative E2E no longer has generic-service pollution, but the model still issued an extra broad HikariCP query without the service filter and used order-service as context. This is a remaining narrow-scope behavior issue, not a mock evidence pollution issue.
@@ -0,0 +1,46 @@
# Acceptance
## Static Verification
- `openspec validate interview-demo-quality-audit --strict`
- Result: passed.
- Coverage: OpenSpec proposal/design/spec/tasks consistency.
- PowerShell parser/runtime readiness check:
- Command: `powershell -NoProfile -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1 -BaseUrl http://127.0.0.1:1 -OutputDir target/demo-check-syntax`
- Result: expected failure with actionable readiness message.
- Coverage: script parses under Windows PowerShell and fails before issuing diagnosis requests when service is unreachable.
## Script Verification
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
- Result: passed.
- Coverage: 12/12 fixed eval fixtures, Prompt audit evaluator checks, Gatekeeper rule metadata checks, regenerated baseline reports.
- `mvn -q "-Dtest=ChatServiceSequentialAgentTest" test`
- Result: passed.
- Coverage: Chat verifier evaluation persists `prompt_audit`.
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ChatServiceSequentialAgentTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`
- Result: passed.
- Coverage: broader eval, baseline diff, Chat sequential flow, Gatekeeper, and Verifier input hook regression set.
- `mvn -q -DskipTests compile`
- Result: passed.
- Coverage: main source compilation.
## E2E Verification
- Started service with:
- `mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo`
- Ran:
- `powershell -NoProfile -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1 -BaseUrl http://localhost:9900 -SessionId mvp-demo-interview-quality-audit-001`
- Result: passed.
- Summary:
- `chatSuccess=true`
- `verdict=LOW_CONFID`
- `gatekeeperStatus=fail`
- `gatekeeperRuleSetVersion=gatekeeper-rules-v1`
- `promptAuditVersion=chat-prompts-v1`
- tools included `lookup_knowledge`, `query_logs`, `query_metrics`, and `get_available_log_topics`
- Note: live E2E remains a compatibility check, not the deterministic PASS oracle. The fixed fixture baseline is the regression source of truth.
## Not Verified
- Browser UI inspection was not required for this change because the scope is backend trace/eval/demo script documentation, not frontend behavior.
@@ -0,0 +1,23 @@
# Interview Demo Quality Audit Brief
## Background
The MVP already demonstrates traceable Agent diagnosis with Planner, Executor, Gatekeeper, Verifier, Composer, evidence tools, trace persistence, and deterministic eval fixtures. The remaining interview-readiness gap is not a new Agent architecture; it is making the demo easier to run and making prompt/rule changes easier to audit.
## Goal
Stabilize the interview demo path, expand fixture-backed evaluation, and persist prompt/Gatekeeper audit metadata so the project can explain and verify Agent behavior during interviews.
## Scope
- Add prompt audit metadata to Chat verifier evaluation.
- Extend deterministic eval cases and baseline reports.
- Add an interview demo preflight/check script.
- Update MVP demo and architecture documentation.
## Non-goals
- No public API or database schema changes.
- No new SubAgent split, MCP migration, process isolation, or AIOps LLM Verifier.
- No guarantee that every live LLM run returns PASS.
@@ -0,0 +1,111 @@
# interview-demo-quality-audit Decisions
## Clarify
- Entry summary: stabilize the interview demo, expand deterministic eval coverage, and add Prompt/Gatekeeper version audit.
- Slug: `interview-demo-quality-audit`.
- Devflow scale: `standard`.
- Interface impact: L2 internal contract change because `verifier_evaluation` gains `prompt_audit`; no public HTTP API or database schema change.
## Context
- `devflow/index.md` used: related entries found for `diagnosis-eval-demo-gatekeeper-closure`, `executor-composer-final-answer`, `verifier-evidence-reference-fidelity`, `mvp-demo-interview-runbook`, and `diagnosis-eval-baseline-diff`.
- Relevant glossary:
- Evidence Tools produce incident facts and must be recorded in `tool_invocation`.
- Verifier should not use skills/runbooks as incident evidence.
- Diagnosis Playbook Skill is workflow guidance, not a fact source.
- Historical constraints that must enter OpenSpec:
- Diagnosis eval is deterministic and fixture-backed; no LLM-as-judge.
- Stable demo scenarios are documentation/payloads plus deterministic fixtures; live E2E is a compatibility check, not a guaranteed PASS oracle.
- Gatekeeper rule metadata is already metadata-only and should not become dynamic rule execution.
- Composer is the final expression layer and must not leak raw Executor JSON.
## Question Pool
| ID | Dimension | Mode | Question | Status |
|---|---|---|---|---|
| Q1 | Terminology | evidence-driven | What should the new audit metadata be called? | Resolved |
| Q2 | Boundary | evidence-driven | Does this require public API or schema changes? | Resolved |
| Q3 | Acceptance | evidence-driven | Which current assets define deterministic acceptance? | Resolved |
| Q4 | Technical | evidence-driven | Where should prompt version metadata live with minimal implementation risk? | Resolved |
| Q5 | Scope | user-interview | Should live E2E be mandatory for all scenarios? | Confirmed by objective as conditional |
## Evidence-driven Conclusions
- Q1 conclusion: use `prompt_audit` for prompt version metadata and keep existing `gatekeeper_result.rule_set_version`.
- Q2 conclusion: keep this as an internal trace/self-evaluation contract change. Do not add endpoints, tables, or new Agent roles.
- Q3 conclusion: `DiagnosisTraceEvaluatorTest`, baseline reports, fixed fixtures, and demo scripts define current acceptance style.
- Q4 conclusion: add a small Chat prompt audit catalog near `ChatService` prompt loading and persist a compact snapshot with verifier evaluation.
- Q5 conclusion: run live E2E with `mvp-demo` profile if dependencies are available; otherwise record the blocker and rely on deterministic eval/unit evidence.
## Specify / Alignment
| Check | Status | Notes |
|---|---|---|
| proposal goals -> proposal | Aligned | Proposal covers demo preflight, eval expansion, prompt audit, and docs. |
| proposal scope/constraints -> design | Aligned | Design records no public API/schema changes, prompt audit shape, eval fields, and demo script behavior. |
| design decisions -> specs/tasks | Aligned | Specs cover persisted prompt audit, evaluator checks, baseline, and demo script outputs. |
| specs observable behavior -> tasks | Aligned | Each requirement has implementation and verification tasks. |
## Audit
Input -> processing -> output chain:
```text
prompt resource metadata
-> ChatService / PromptAudit snapshot
-> verifier_evaluation.prompt_audit
-> Trace API / eval fixtures
-> DiagnosisTraceEvaluator baseline
run-interview-demo-check.ps1
-> service readiness
-> chat / trace / feedback
-> mvp/demo/output summary
```
Architecture risk assessment:
1. The audit shape is intentionally compact and internal; storing full prompt text would create noisy traces and possible sensitive-content risk.
2. Eval should assert versions by explicit metadata, not by prompt content hashes that churn during local prompt edits.
3. Live demo checks may still be LOW_CONFID because LLM output is not deterministic; deterministic fixtures remain the regression source of truth.
4. No devflow/OpenSpec conflict found.
## Commit Gate
- `openspec validate interview-demo-quality-audit --strict`: passed.
- File completeness:
- `proposal.md`: present.
- `design.md`: present.
- `specs/`: present for `chat-verifier-agent`, `diagnosis-eval-harness`, and `mvp-demo-trace-acceptance`.
- `tasks.md`: present.
- Consistency:
- Proposal goals map to design sections.
- Design decisions map to spec requirements and executable tasks.
- Task acceptance checks are verifiable.
- `.committed` marker created.
## Current Checkpoint
- Commit completed.
- Apply is authorized by the original objective: "完成后归档提交".
## Pre-apply Research
- Capability source: sm-flow built-in apply protocol. `openspec-apply-change` was not invoked directly in this session.
- Repository semantic search/LSP note: the requested `codebase-retrieval` and LSP tools were not available in the exposed toolset, so impact analysis used `rg`, direct file reads, OpenSpec/devflow artifacts, and targeted tests.
- Reference implementation and reuse:
- `ChatService.persistVerifierEvaluation(...)` is the single persistence point for Chat verifier/composer audit data; prompt audit was added there to cover normal, fallback, and degraded Composer paths.
- `DiagnosisTraceEvaluator` and `DiagnosisEvalReportWriter` are the deterministic eval extension points; no LLM judge was introduced.
- `mvp/demo/scripts/run-payment-timeout-demo.ps1` provided the request/trace/feedback flow reused by the new interview preflight script.
- Interface impact remains L2 internal trace contract: `verifier_evaluation.prompt_audit` and eval report fields are added; no public endpoint, table, or request DTO changed.
## Apply Notes
- Added compact Chat prompt audit metadata: `chat-prompts-v1`, with planner/executor/verifier/composer prompt versions and resource paths.
- Extended diagnosis eval schema, result reporting, baseline fixtures, JSON report, and Markdown report for Prompt audit and Gatekeeper rule metadata.
- Added two fixture-backed audit cases:
- `prompt-gatekeeper-audit-closure`
- `audit-metadata-low-confid`
- Added `mvp/demo/scripts/run-interview-demo-check.ps1` to run service readiness, Chat, Trace, feedback, and summary output.
- Updated MVP demo/eval/architecture docs to explain `prompt_audit.version`, `gatekeeper_result.rule_set_version`, and deterministic fixture baseline.
@@ -0,0 +1,58 @@
# Evidence
## Context Files Read
- `devflow/index.md`
- `devflow/glossary/CONTEXT.md`
- `devflow/projects/2026-07-08-diagnosis-eval-demo-gatekeeper-closure/decisions.md`
- `devflow/projects/2026-07-08-executor-composer-final-answer/decisions.md`
- `mvp/architecture/current-mvp-architecture.md`
- `mvp/architecture/agent-orchestration.md`
- `mvp/architecture/executor-evidence-pipeline-refactor.md`
- `mvp/architecture/harness-quality-gates.md`
- `mvp/demo/README.md`
- `mvp/demo/ten-minute-interview-demo.md`
- `mvp/eval/README.md`
- `mvp/eval/cases/diagnosis-cases.json`
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- `src/main/java/com/superbiz/agent/service/ExecutorGatekeeperService.java`
- `src/main/java/com/superbiz/agent/eval/DiagnosisTraceEvaluator.java`
- `src/main/resources/gatekeeper/gatekeeper-rules.json`
## Tooling Note
The required `codebase-retrieval` and LSP tools were not exposed in this session. Impact analysis used `rg`, direct file reads, existing OpenSpec/devflow artifacts, and targeted tests instead.
## Implementation Evidence
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- Adds `prompt_audit` under `verifier_evaluation` through the shared `persistVerifierEvaluation(...)` path.
- Uses compact metadata only: audit version, prompt names, prompt versions, and resource paths.
- `src/main/java/com/superbiz/agent/eval/DiagnosisTraceEvaluator.java`
- Adds deterministic checks for `requirePromptAudit`, `expectedPromptAuditVersion`, `expectedPromptVersions`, and `requireGatekeeperRules`.
- `src/main/java/com/superbiz/agent/eval/DiagnosisEvalReportWriter.java`
- Adds Prompt Audit and Gatekeeper rule count columns to Markdown reports.
- `mvp/eval/cases/diagnosis-cases.json`
- Expands fixed baseline to 12 fixture-backed cases.
- `mvp/eval/fixtures/prompt-gatekeeper-audit-closure-pass.json`
- Positive PASS fixture proving Prompt audit and Gatekeeper rule metadata closure.
- `mvp/eval/fixtures/audit-metadata-low-confid.json`
- LOW_CONFID fixture proving safe answer behavior while audit metadata remains present.
- `mvp/demo/scripts/run-interview-demo-check.ps1`
- Adds service readiness, Chat, Trace, feedback, and summary output for interview preflight.
## Verification Evidence
- OpenSpec:
- `openspec validate interview-demo-quality-audit --strict`: passed before archive.
- `openspec validate --specs --strict`: 10 specs passed after merging deltas into main specs.
- Unit/eval:
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`: passed.
- `mvn -q "-Dtest=ChatServiceSequentialAgentTest" test`: passed.
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ChatServiceSequentialAgentTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`: passed.
- Compile:
- `mvn -q -DskipTests compile`: passed.
- E2E:
- Started `mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo`.
- Ran `mvp/demo/scripts/run-interview-demo-check.ps1` against `http://localhost:9900`.
- Summary recorded `chatSuccess=true`, `verdict=LOW_CONFID`, `gatekeeperRuleSetVersion=gatekeeper-rules-v1`, and `promptAuditVersion=chat-prompts-v1`.
@@ -0,0 +1,103 @@
# Acceptance
## 实现结果
- OpenSpec tasks: `42/42` complete。
- Phase commits:
- `52bf030 feat(trace): add session run isolation schema`
- `6fdbd34 docs(openspec): tighten run isolation contract`
- `26d5529 feat(trace): isolate chat runs`
- `027aed1 feat(trace): add run-scoped trace reads`
- `d928a19 feat(trace): bind feedback to runs`
- `78c1477 feat(trace): isolate aiops runs`
- `f9df943 feat(trace): finish run-aware demo verification`
- OpenSpec archive: `openspec/changes/archive/2026-07-10-session-run-trace-isolation`
## 静态验证
```powershell
node --check src\main\resources\static\app.js
node --check src\main\resources\static\trace.js
openspec validate session-run-trace-isolation --strict
git diff --check -- . ':!devflow/index.md'
```
结果:通过。
## 脚本验证
PowerShell demo 脚本解析:
```powershell
$scripts = @(
'mvp\demo\scripts\run-payment-timeout-demo.ps1',
'mvp\demo\scripts\run-interview-demo-check.ps1'
)
foreach ($script in $scripts) {
[scriptblock]::Create((Get-Content -Raw -Encoding UTF8 $script)) | Out-Null
}
```
结果:通过。
Focused tests:
```powershell
mvn -q "-Dtest=ChatControllerTest,DiagnosisTraceServiceTest,FeedbackControllerTest,FeedbackServiceTest,AiOpsServiceTest" test
```
结果:通过。
Baseline / regression:
```powershell
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest" test
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
```
结果:通过,无 baseline drift。
## E2E 验证
使用 Maven 启动:
```powershell
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
```
E2E 使用同一 `sessionId` 连续两轮 Chat:
- `sessionId`: `e2e-phase6-chat-codex-20260710-2120`
- `run1`: `run-e2a97696-4398-4abc-90e4-28f45c838f92`
- `run2`: `run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172`
验证结果:
- Chat1 / Chat2 均成功。
- run1 exact trace 返回 run1。
- run2 exact trace 返回 run2。
- session-only latest trace 返回 run2。
- DB 中同一 session 有两条 `diagnosis_run`。
- step/tool rows 按 `run_id` 隔离,mixed row check 为 0。
- `chat_session.message_pair_count = 2`。
## 日志验证
检查:
- `target/e2e/phase6-mvn-20260710-211831.out.log`
- `logs/application.log`
- `logs/chat.log`
结果:能找到 E2E `sessionId`、两个 `runId`、Chat execution、run persistence 和 trace lookup 相关日志。
## 浏览器/人工验证
未单独进行浏览器点击验证。Trace UI 的本次验收通过静态语法检查、URL/runId 参数代码审查和后端 exact trace E2E 共同覆盖。建议后续手动打开 `trace.html?sessionId=...&runId=...` 做展示层冒烟。
## 剩余风险 / 后续事项
- 缺少 `runId` 的 Feedback fallback 是短期兼容路径,客户端全部迁移后可收紧。
- `diagnosis_session` 仍保留为历史兼容和回滚表,后续需要观察窗口后再评估约束收紧或归档策略。
- `case_library.diagnosis_id` 仍是过渡字段,旧值可能为 `session_id`,新自动值为 `run_id`。
- 历史 mixed trace 不能恢复真实多轮边界,只能按 compatibility run 查询。
@@ -0,0 +1,41 @@
# Session / Run / Trace Isolation
## 背景
同一个 `sessionId` 以前同时代表多轮 Chat 上下文和一次持久化诊断 Trace。端到端验证发现,同一 `sessionId` 连续两轮 Chat 时,Redis 多轮上下文是正确的,但 MySQL 中 `diagnosis_session` 会被后一轮覆盖,`agent_step` 和 `tool_invocation` 会按同一个 `session_id` 混在一起。
这会导致 Trace 回放、Verifier/Evaluation 读数、Feedback 绑定和 `case_library` 来源都可能跨轮污染。
## 目标
- 将会话态和运行态拆开:`chat_session` 保存会话元数据,`diagnosis_run` 保存一次诊断运行。
- 引入正式 API 字段 `runId`,作为一次可回放诊断执行的边界。
- `agent_step` 和 `tool_invocation` 保留原 Trace 明细角色,新增 `run_id` 并按 run 隔离读写。
- Trace、Feedback、CaseLibrary、AIOps、demo 脚本和 Trace UI 都支持 run-aware 流程。
- 保留旧 `diagnosis_session` 作为历史兼容和回滚表。
- 完成 Maven E2E、DB 检查、日志检查和 baseline drift 验证。
## 范围
- Flyway/JPA 增加 `chat_session`、`diagnosis_run`,并给 `agent_step`、`tool_invocation` 增加 `run_id`。
- Chat 每次有效执行创建一个新的 `diagnosis_run`,响应返回 `sessionId + runId`。
- Trace API 支持 latest-run fallback 和 exact-run 查询:`GET /api/diagnosis/{sessionId}/trace?runId=...`。
- 新增 run list API:`GET /api/chat/session/{sessionId}/runs`。
- Feedback 优先绑定 `runId`,缺省时短期 fallback 到 latest run 并返回 `fallbackToLatestRun=true`。
- AIOps 每次有效执行创建并透出 `runId`,SSE 保持 `message` event name 并发送 `type=metadata`。
- MVP demo、Trace UI、表文档和架构文档统一为 `chat_session -> diagnosis_run -> trace detail(run_id)`。
## 非目标
- 不新增 `diagnosis_trace` 或 `trace_event` 主表。
- 不实现完整 run-list UI。
- 不删除旧 `diagnosis_session`。
- 不改变 Redis 对话历史窗口策略。
- 不把完整多轮正文历史持久化到 MySQL。
- 不尝试把历史混合 trace 还原成真实多轮边界。
## 关联
- OpenSpec: `openspec/changes/archive/2026-07-10-session-run-trace-isolation`
- Change slug: `session-run-trace-isolation`
- 分档: complex
@@ -0,0 +1,48 @@
# Decisions
## 核心决策
| 决策 | 选择 | 理由 |
|---|---|---|
| 领域拆分 | 新增 `chat_session` 和 `diagnosis_run` | 会话元数据和一次诊断执行的生命周期不同,继续塞在一张表会导致上下文膨胀和边界混淆 |
| Trace 明细 | 复用 `agent_step` / `tool_invocation`,增加 `run_id` | 现有明细表已经能表达 Trace,隔离需要 run key,不需要新事件模型 |
| API 身份 | `runId = "run-" + UUID` | 外部 ID 不依赖数据库自增 ID,碰撞风险低 |
| Trace 兼容 | 缺少 `runId` 时按 `created_at DESC, id DESC` 解析 latest run | 保留旧客户端兼容性,避免 feedback/eval 更新 `updated_at` 后改变 latest 判定 |
| 历史迁移 | 每条旧 `diagnosis_session` 生成一条 compatibility run | 旧混合数据没有真实轮次边界,不能伪造多 run 历史 |
| Feedback fallback | 缺少 `runId` 时短期绑定 latest run 并返回 `fallbackToLatestRun=true` | 老客户端可继续工作,同时让歧义可观测 |
| Case provenance | 新自动案例写 `case_library.diagnosis_id = run_id` | 保留旧列,文档声明过渡语义 |
| AIOps 范围 | 同一个 change 内完成 AIOps run isolation | AIOps 是一等 Trace 入口,不能留下同类混合 trace bug |
| 所有权校验 | 服务层校验 run/session ownership,暂不加 DB 外键 | 兼容历史 orphan rows 和回滚窗口 |
## 用户确认
- 选择拆 `chat_session` 和 `diagnosis_run`,不只是在旧表加字段。
- `chat_session` 第一阶段只保存元数据,不保存完整对话正文。
- 完整多轮对话历史继续放在 Redis `SessionContext.messageHistory`。
- `runId` 是正式 API 字段。
- Trace 缺少 `runId` 时短期默认查 latest run。
- Feedback 缺少 `runId` 时短期 fallback,长期可再收紧。
- 每次有效 Chat/AIOps 都创建 run。
- 不新增 `diagnosis_trace` / `trace_event` 主表。
- 旧 `diagnosis_session` 保留用于历史和回滚,新代码不再写新执行态。
- demo 脚本和 Trace UI 做最小 `runId` 支持。
## 接口影响
级别:L4。
- 新 API 响应字段:`runId`。
- Trace API 新 query 参数:`runId`。
- 新 API:`GET /api/chat/session/{sessionId}/runs`。
- Feedback request 新增 optional/preferred `runId`。
- Feedback response 新增 bound `runId` 和 `fallbackToLatestRun`。
- `/api/ai_ops` SSE 保持 event name `message`,新增 `type=metadata` 消息。
- DB contract 新增两张表和两个 `run_id` 列。
- 旧 `sessionId` only 调用仍兼容,但 fallback 必须可观测。
## 风险接受
- 历史混合 trace 无法真实拆分,只能作为 compatibility run。
- 上下文传播同时依赖 `RunnableConfig.metadata` 和 `SessionContextHolder`,后续改动必须注意 `sessionId/runId` 同步。
- `case_library.diagnosis_id` 在过渡期存在 `session_id` 和 `run_id` 两种语义。
- 缺少 `runId` 的 Feedback 仍有歧义,后续客户端迁移完成后可收紧为参数错误。
@@ -0,0 +1,76 @@
# Evidence
## 上下文证据
- `SessionContext.messageHistory` 和 `getMessagePairCount()` 证明 Redis 承载热对话历史;MySQL 只需要长期审计的会话目录和运行记录。
- `CaseLibraryService.createFromSession` 原先按 `DiagnosisSession.sessionId` 去重并映射 query/answer,因此 run 隔离后需要新增 `createFromRun`。
- 旧 `mvp/architecture/data-model.md` 把 `case_library.diagnosis_id` 解释为 `diagnosis_session.session_id`,本次改为过渡语义:旧数据可能是 `session_id`,新自动案例是 `run_id`。
- 既有 Trace OpenSpec 要求 `GET /api/diagnosis/{sessionId}/trace` 是只读端点;latest-run 和 exact-run 查询都必须保持只读。
- ISS-010 的 E2E 事实显示同一 `sessionId` 两轮 Chat 会产生 MySQL Trace 混合,是本 change 的直接触发证据。
## 实现证据
- Phase 1 增加 `V011__add_session_run_isolation.sql`,创建 `chat_session`、`diagnosis_run`,并为 `agent_step` / `tool_invocation` 增加 nullable `run_id`。
- Phase 2 将 Chat 写路径切到 `chat_session + diagnosis_run`,并让 Hook/Tool/Evaluation/Gatekeeper 使用 run-scoped 数据。
- Phase 3 将 Trace API 改为 latest-run / exact-run 双模式,并加入 lightweight run summaries。
- Phase 4 将 Feedback 和 CaseLibrary 绑定到 run,保留没有 run-backed 数据时的 legacy fallback。
- Phase 5 将 AIOps 接入 run isolation,SSE metadata 暴露 `sessionId + runId`。
- Phase 6 更新 demo 脚本、Trace UI、MVP 架构文档和表文档,并修正 review 后发现的 session-only 文档残留。
## E2E 证据
Maven 启动命令:
```powershell
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
```
日志:
- `target/e2e/phase6-mvn-20260710-211831.out.log`
- `target/e2e/phase6-mvn-20260710-211831.err.log`
- `logs/application.log`
- `logs/chat.log`
E2E session:
- `sessionId`: `e2e-phase6-chat-codex-20260710-2120`
- `run1`: `run-e2a97696-4398-4abc-90e4-28f45c838f92`
- `run2`: `run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172`
结果:
- 两轮 Chat 都成功,并复用同一个 `sessionId`。
- 两轮返回不同 `runId`。
- run1 exact trace 只返回 run1。
- run2 exact trace 只返回 run2。
- session-only Trace latest fallback 返回 run2。
- `chat_session.message_pair_count = 2`,证明多轮上下文连续。
## DB 证据
通过 `scripts/query_mysql.py` 检查:
- `diagnosis_run` 中该 E2E session 有 2 条 `SUCCESS / CHAT` 运行。
- `agent_step` 按 run 分组:run1 `10` 行,run2 `9` 行。
- `tool_invocation` 按 run 分组:run1 `14` 行,run2 `8` 行。
- mixed row check 为 `0`,没有 NULL 或 unexpected `run_id` 混入该 E2E session。
## Baseline 证据
运行:
```powershell
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest" test
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
```
结果:
- 两组 baseline / regression 命令通过。
- baseline harness 使用离线 fixture,不依赖 live DB/session tables。
- 未观察到 baseline drift。
## 工具限制
AGENTS 要求的 `codebase-retrieval` 和 LSP 工具在本会话不可用。替代验证使用 OpenSpec、`rg`、定向阅读、 focused tests、E2E、DB 查询和日志检查。
@@ -0,0 +1,44 @@
# Acceptance: single-react-aci-tool-contracts
## 实现结果
- RAG Contract:最小 query Request、bounded document evidence Result。
- Log Contract:逻辑 Topic/Lookback Request、Mock provenance、Scope/Pattern/Event Result。
- MySQL Contract:逻辑 data source、参数化 SQL Request、bounded structured rows Result。
- 共享 Contract:snake_case Tool 名称、ACI 描述、不可变集合 helper,复用阶段 0 两套状态枚举。
- 旧 Tool、Chat/AIOps、Controller、持久化、数据源和公开协议未修改。
## 静态验证
- `openspec validate single-react-aci-tool-contracts --strict`:通过。
- `openspec instructions apply --change single-react-aci-tool-contracts --json`:12/12 tasks complete。
- `rg` 引用检查:新 contract 生产包未接入旧运行链路。
- 受保护文件 diff scope:旧 Tool、Chat/AIOps、Controller、Repository、resources 均为空。
## 脚本验证
- `mvn -q -DskipTests compile`:通过。
- `mvn -q '-Dtest=RagToolContractTest,QueryLogsToolContractTest,MysqlToolContractTest' test`:通过。
- `mvn -q '-Dtest=HarnessContractTest,LookupKnowledgeToolTest,QueryLogsToolsTest' test`:通过。
## 浏览器/人工验证
- 不适用。本阶段无 UI、Controller、SSE 或公开运行行为变化。
## 未验证
- 未执行 live LLM/Redis/CLS/MySQL E2E;本阶段没有接入这些运行路径,最终 live E2E 按 ISS-014 门禁留到阶段 7。
- 未验证 ToolInterceptor 的真实 ID 传播;阶段 2/3A 必须以 `ToolCallRequest.getToolCallId()` 添加集成测试。
- 未验证真实日志 adapter 的 `SourceKind` 扩展;本 Issue 首版明确只使用 Mock。
- Provider 侧旧凭据轮换仍需凭据所有者完成,仓库只能证明明文已移除。
## 剩余风险与后续门禁
- 新旧 Contract 短期并存,阶段 3B/3C 接入前不得声称旧 Tool 已符合新 ACI 输出。
- 下一阶段只能在本 change OpenSpec Archive 和 Git commit 完成后开始。
## 状态
- Stage acceptance: accepted
- OpenSpec archive: archived at `openspec/changes/archive/2026-07-21-single-react-aci-tool-contracts`
- Main spec sync: `openspec/specs/aci-evidence-tool-contracts/spec.md`(7 added requirements)
@@ -0,0 +1,32 @@
# Brief: single-react-aci-tool-contracts
## 背景
旧 RAG 和日志 Tool 暴露检索/基础设施细节与不一致状态,MySQL Tool 尚无 Agent-facing 类型。Harness 实现前需要先冻结三类最小 ACI 契约。
## 目标
- 冻结 RAG、日志、MySQL 的 Request/Result JSON Schema。
- 统一 `evidence_status`,并保持其与 invocation lifecycle 独立。
- 统一使用框架 `tool_call_id`,禁止模型传入或 Harness 生成第二套 ID。
- 冻结简短 Tool 名称/描述和 Mock 日志来源边界。
- 用三个独立契约测试锁定行为。
## 范围
- 新增 `com.superbiz.agent.harness.tool.contract` 值对象、枚举和描述常量。
- 复用阶段 0 的 `InvocationStatus` 与 `EvidenceStatus`。
- 验证 JSON、不可变集合、描述泄漏和新旧边界。
## 非目标
- 不切换旧 RAG/日志运行方法或 Chat/AIOps 注册。
- 不实现投影、Redis store、真实日志适配器或 MySQL 执行。
- 不修改 Controller/SSE 或公开协议。
## 元数据
- 分档:standard
- 接口影响:L2 前置内部契约;当前运行行为无变化
- 关联 Issue:ISS-014 阶段 1
- 关联 OpenSpec:`openspec/changes/single-react-aci-tool-contracts`
@@ -0,0 +1,116 @@
# Decisions: single-react-aci-tool-contracts
## 规模与入口
- 分档:standard。
- 入口:ISS-014 阶段 1,前置 `single-react-design-freeze` 已 Archive 并由 Git commit `58c3910` 固化。
- 目标:冻结三类 evidence Tool 的 Agent-facing ACI 契约,不接入新运行链路。
## Context
- `devflow/index.md` 命中 `single-react-design-freeze`、`modular-rag-pipeline` 和 `evidence-trace-hardening`。
- `devflow/glossary/CONTEXT.md` 已定义 Diagnosis Harness、Invocation Status、Evidence Status 与 Evidence Tools。
- 阶段 0 已确认 `tool_call_id` 使用框架 ID、两套状态语义分离、阶段串行门禁和阶段 6B 才公开切换。
- 未发现根目录旧 `CONTEXT.md` 与 glossary 冲突。
## Question Pool
| 维度 | 问题 | 模式 | 证据与结论 | 状态 |
|---|---|---|---|---|
| 术语 | invocation lifecycle 与 evidence result 是否使用同一状态? | evidence-driven | ISS-014 5.2、阶段 0 contract 明确分离;分别复用 `InvocationStatus` 与 `EvidenceStatus`。 | 已解决并汇报 |
| 术语 | `tool_call_id` 由谁生成、从哪里取得? | evidence-driven | Spring AI `AssistantMessage.ToolCall.id()` 与 Alibaba `ToolCallRequest.getToolCallId()` 提供框架 ID;普通 `ToolContext` 不自动加入该 ID。Harness 不生成第二套 ID。 | 已解决并汇报 |
| 边界 | 阶段 1 是否直接改旧 RAG/日志执行签名与返回值? | evidence-driven | ISS-014 阶段 1 只冻结 Contract,投影在阶段 3B、公开切换在阶段 6B;本阶段只新增契约代码和测试。 | 已解决并汇报 |
| 边界 | 日志阶段是否实现真实 CLS/MCP 或保留 Topic discovery? | evidence-driven | ISS-014 8.1 明确继续 Mock、删除 Agent 侧 discovery、真实适配器不在本 Issue 提前设计。 | 已解决并汇报 |
| 验收 | 如何证明契约已冻结且有界? | evidence-driven | 对三类独立 DTO 做精确 JSON、不可变集合、状态和描述泄漏测试;不以旧 Tool 集成测试代替。 | 已解决并汇报 |
| 技术 | 框架 ID 能否在后续 Harness 边界取得? | evidence-driven | 本地依赖 Spring AI Alibaba 1.1.2.0 暴露 `ToolInterceptor.interceptToolCall(ToolCallRequest, ToolCallHandler)`,request 含 `toolCallId`。 | 已解决并汇报 |
## Grill 结论
- 术语、边界、验收三类问题均已由代码、依赖 API、阶段 0 档案和 ISS-014 证明。
- 没有需要新增用户偏好或风险取舍的 `user-interview` 问题;不代理确认任何新方向。
- `grill-with-docs` 要求的代码可证问题已先查证;结论已向用户汇报。
- Proposal 已回写框架 ID 接入点、旧运行链路不切换、Mock 边界和 L2 接口影响。
## 已确认决策
- DTO 放在 Harness 的 Tool Contract 边界,复用阶段 0 的共享状态枚举,不在旧 `dto` 包继续堆叠协议。
- 三类结果只携带 `evidence_status`,canonical invocation 的 `status` 保持独立;阶段 1 通过测试冻结枚举,不提前定义存储实现。
- Tool description 使用代码常量冻结,后续 Tool adapter 注册时复用;旧 `@Tool` 注解本阶段不改,避免提前改变运行行为。
- RAG 输入仅保留 `query`;日志输入仅保留逻辑 `topic/query/lookback_minutes`;MySQL 输入仅保留逻辑 `data_source/sql/params`。
- 日志 `source_kind` 首版固定支持 `MOCK` 契约值,但保留 enum 扩展位置给后续真实适配器 change 审查。
## 能力与工具限制
- Discover 能力来源:`sm-flow` + `grill-with-docs`。
- 仓库要求的 `codebase-retrieval` 和 LSP 工具在当前工具集中不可用;已用 `rg` 引用搜索、源码阅读和本地依赖 `javap` 补足事实核对。该限制不改变契约方向,但后续 Apply 仍需通过编译和引用测试验证。
## Cross-artifact 对齐
| 链路 | 状态 | 结论 |
|---|---|---|
| brief 目标/范围/非目标 -> proposal | 已对齐 | 三类 DTO、状态、框架 ID、短描述和不切旧运行链路均有对应。 |
| proposal 范围/约束/承诺 -> design | 已对齐 | 包边界、ID 来源、record/defensive copy、三类 Schema 和迁移顺序均已设计。 |
| design 决策/接口影响/风险 -> specs/tasks | 已对齐 | L2 边界、字段、状态、描述、Mock provenance 和运行不切换均有可验证 requirement 与任务。 |
| specs 可观察行为 -> tasks | 已对齐 | 每类 Contract 都有实现与独立测试,另有状态、回归和 diff scope 验证。 |
## Architecture Audit
- 能力来源:`zoom-out`,以项目 glossary 的 Diagnosis Harness、Evidence Tools、Invocation Status 和 Evidence Status 术语审计。
- 当前输入到输出链路仍为 `ChatController/ChatService` 或 `AiOpsService -> ReactAgent -> LookupKnowledgeTool/QueryLogsTools -> 旧结果`,新契约没有运行消费者。
- 未来链路为 `Diagnosis Agent -> Alibaba ToolInterceptor/Harness -> typed Request -> adapter/store/projector -> bounded Result -> Agent observation`,阶段 1 只占有 typed contract 边界。
- 数据所有权保持明确:框架拥有 `tool_call_id`,canonical invocation 拥有生命周期,Tool-specific result 拥有证据语义与有界内容。
- 主要耦合风险是新旧契约短期并存被误当成已迁移;通过独立包、无旧调用方修改和后续阶段门禁控制,无 ADR 冲突。
## Commit Gate Preflight
- proposal、design、specs、tasks 文件完整,OpenSpec CLI 状态为 complete。
- strict validation:`openspec validate single-react-aci-tool-contracts --strict` 通过。
- question pool 中没有未汇报的 evidence-driven 结论或未确认的 user-interview 问题。
- 接口影响为 L2,当前运行消费者零变更;阶段 6B 的 L4 切换保持独立。
- cross-artifact 四段对齐无 gap,架构审计未发现需要回写的新实现约束。
- Apply 已由用户对 ISS-014 全阶段的持续授权覆盖;仍严格限制在本 Committed OpenSpec tasks 内。
## Pre-apply Research
### 参考实现
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`:确认旧 RAG 暴露 `LookupResult`、ContextPack 和检索 Trace,本阶段不改。
- `src/main/java/com/superbiz/agent/agent/tool/QueryLogsTools.java`:确认旧日志输入包含 region/logTopic/limit、存在 Topic discovery 和 mutable nested DTO,本阶段不改。
- `src/test/java/com/superbiz/agent/tool/LookupKnowledgeToolTest.java`:现有 RAG 行为回归基线。
- `src/test/java/com/superbiz/agent/agent/tool/QueryLogsToolsTest.java`:现有 Mock 日志行为回归基线。
- `src/main/java/com/superbiz/agent/harness/contract/*.java`:Java 17 record、Jackson snake_case 和共享状态枚举风格。
### 技术栈清单
- 请求/响应标准:Java 17 record + Jackson `@JsonProperty`,集合构造时 defensive copy。
- Tool 定义:当前使用 Spring AI `@Tool`,新名称/描述先以常量冻结,后续 adapter 注册复用。
- 框架调用 ID:Spring AI Alibaba `ToolCallRequest.getToolCallId()`;普通 `ToolContext` 不作为 ID 来源。
- 异常与校验:本阶段只冻结结构;Schema、ID、权限、范围和状态组合错误由后续 Pre-Tool/Projector 显式返回安全 `ERROR`。
- MQ/Consumer/加密验签:本 change 不涉及。
### 新建类型
- `AgentToolContracts` 与 contract defensive-copy helper。
- RAG Request/Result/Evidence。
- Log Topic/SourceKind/Request/Result/Scope/Pattern/Event。
- MySQL Request/Result。
- 三个独立 contract test classes。
### 影响半径
- 新增包当前应无生产调用方;旧 Chat/AIOps、Tool、Controller、Repository 和配置文件均不修改。
- 通过 focused compile/tests 和 `rg`/diff scope 证明边界。
## Apply 结果
- 冲突分类:未发现 OpenSpec 遗漏、代码偏离或方向不确定项。
- 新增共享 ACI Tool 名称/描述、defensive-copy helper 和三类 typed Request/Result records。
- 新增三个独立契约测试,覆盖精确 JSON、状态分离、框架 ID 原样保留、不可变集合、Mock provenance 和基础设施字段排除。
- 首模块对齐:共享/RAG/日志/MySQL contract 与 design/tasks 全部完成;旧 runtime 接入保持 TODO,归属后续 3B/3C/6B changes。
## Apply 验证
- 编译:`mvn -q -DskipTests compile` 通过。
- 新契约:`mvn -q '-Dtest=RagToolContractTest,QueryLogsToolContractTest,MysqlToolContractTest' test` 通过。
- 旧行为回归:`mvn -q '-Dtest=HarnessContractTest,LookupKnowledgeToolTest,QueryLogsToolsTest' test` 通过。
- 静态 scope:新 contract 生产类型当前无旧运行消费者;受保护的 Tool、Chat/AIOps、Controller、Repository 和配置文件 diff 为空。
@@ -0,0 +1,27 @@
# Evidence: single-react-aci-tool-contracts
## 文档证据
- ISS-014 5.2/5.3 冻结 `evidence_status` 与框架 `tool_call_id`;6-9 节冻结 RAG、日志和 MySQL Agent-facing Schema。
- ISS-014 阶段 1 明确只冻结 ACI Contract,不实现 Agent 架构切换、真实 CLS/MCP 或 MySQL 执行。
- 阶段 0 OpenSpec 与 devflow 已冻结 `InvocationStatus`、`EvidenceStatus`、框架 ID 真理源和阶段 6B 才公开切换。
## 代码证据
- `LookupKnowledgeTool` 仍返回包含 ContextPack/Trace 的旧 `LookupResult`,证明需要新的 bounded RAG Contract,也证明本阶段未提前切换。
- `QueryLogsTools` 仍暴露 region/logTopic/limit、Topic discovery 和旧 mutable DTO,证明逻辑 Topic/Scope/Mock provenance 契约的必要性。
- `ChatService` 与 `AiOpsService` 仍引用旧 `LookupKnowledgeTool`/`QueryLogsTools`,新 contract 生产包当前没有旧运行消费者。
- 本地 Spring AI 1.1.7 `AssistantMessage.ToolCall` 提供 `id()`;Spring AI Alibaba 1.1.2.0 `ToolCallRequest` 提供 `getToolCallId()` 与 `ToolInterceptor` 边界。
- Spring AI 1.1.7 普通 `ToolContext` 只传递调用方 context/history,不自动提供当前 Tool Call ID,因此后续必须从 Alibaba interceptor request 接入。
## Evidence-driven 结论
- lifecycle 与 evidence result 必须保持两套正交状态;已汇报并进入 OpenSpec/代码测试。
- `tool_call_id` 可以从当前框架 API 取得,Harness 无需也不得生成第二套 ID;已汇报并进入 OpenSpec。
- 阶段 1 不改旧运行签名/返回;已通过引用和 diff scope 验证。
- 日志首版保留 Mock 数据源但必须显式 `source_kind=MOCK`,不实现真实适配器;已进入 Log Contract。
- 三类 Contract 的可验证口径是精确 JSON、不可变结果、短描述和基础设施字段排除;三个独立测试均通过。
## 工具限制
- 当前会话未提供 `codebase-retrieval` 或 LSP;使用 `rg` 引用搜索、源码阅读、本地依赖 `javap`、Maven 编译和 focused tests 完成等价核对。
@@ -0,0 +1,42 @@
# Acceptance: single-react-chat-application-usecase
## Result
- Status: archived
- OpenSpec tasks: 14/14 complete
- Interface impact: L3 database/collaboration
- Public protocol: unchanged
## Static Verification
- V012 migration、`DiagnosisRun` 和 `DiagnosisRunRepository` 对齐 `intent/release_outcome/published_result` 及安全 PreviousTurn filter。
- 公开 Controller、前端和 endpoint diff 为空。
- Router/System executor 未检出 Tool、ReactAgent、ThreadLocal 或手写循环。
- `PublishedResult` 固定为 `user_query/published_conclusion/scope/limitations/source_documents`;序列化负向测试覆盖内部字段泄漏。
## Script Verification
- `mvn -q -DskipTests compile`:通过。
- Stage 6A focused `ApplicationExecutorsTest,PublishedResultPersistenceTest,ChatApplicationUseCaseTest`:13 tests,通过。
- Stage 2-5 与 6A regression selection:18 suites / 76 tests,0 failure/error/skipped。
- `openspec validate single-react-chat-application-usecase --strict`:通过。
## Browser or Manual Verification
- Not applicable。阶段 6A 没有 UI 或公开入口变化。
## Not Verified
- 未运行真实 LLM、Redis、日志和 MySQL live E2E;按 ISS-014 串行门禁统一留到阶段 7。
- V012 未在本阶段连接真实数据库执行;migration/entity/query 已由静态检查、focused persistence tests 和 compile 覆盖。
## Remaining Work
- 阶段 6B:唯一 `POST /api/chat` SSE 原子切换、旧 endpoint 删除和前端消费者迁移。
- 阶段 7:旧链路清理、全局 spec 格式修复和最终 live E2E。
## Archive
- `.archive-ready`: created
- OpenSpec archive: `openspec/changes/archive/2026-07-21-single-react-chat-application-usecase`
- Main spec sync: `openspec/specs/single-react-chat-application-usecase/spec.md`
@@ -0,0 +1,31 @@
# Brief: single-react-chat-application-usecase
## Background
阶段 2-5 已具备 RunContext、单一 Diagnosis Agent 和安全释放门禁,但没有统一应用用例拥有 Session/Run、意图路由、PreviousTurn、固定执行器和最终持久化,阶段 6B 因而无法只做协议切换。
## Goal
在不改变公开 Chat/SSE 行为的前提下,建立内部 `ChatApplicationUseCase`,统一三类意图、同一 Run 生命周期、安全 PreviousTurn 和 typed public content。
## Scope
- 无 Tool、无记忆、无 ReAct 的三分类 Intent Router。
- SYSTEM_CHAT、KNOWLEDGE_QUERY、DIAGNOSIS 固定执行器。
- Knowledge exact invocation/document reference validation。
- 同 Session 最近安全 Diagnosis SUCCESS 的有界 PreviousTurn。
- `diagnosis_run` V012 字段、JPA store 和安全 `PublishedResult`。
- Protocol-neutral observer、Run control、终态持久化和 focused tests。
## Non-goals
- 不修改 Controller、`/api/chat`、`/api/chat_stream`、SSE schema 或前端消费者。
- 不删除旧 ChatService、多 Agent、ThreadLocal 或旧 session storage。
- 不运行真实模型、Redis、日志和 MySQL live E2E;统一留到阶段 7。
## Metadata
- Scale: complex
- Interface impact: L3 database/collaboration
- OpenSpec: `single-react-chat-application-usecase`
- Parent issue: `ISS-014`

Some files were not shown because too many files have changed in this diff Show More