Compare commits

..
Author SHA1 Message Date
zhuyongxin 4b9cf7c5cc refactor(trace): enforce run-only diagnosis model 2026-07-20 14:11:34 +08:00
zhuyongxin 190013c901 feat(graph): complete stategraph cleanup and acceptance 2026-07-20 10:23:27 +08:00
zhuyongxin 208a231113 test(graph): establish diagnosis stategraph suite 2026-07-17 18:53:29 +08:00
zhuyongxin 99e490f227 feat(graph): cut over chat diagnosis stategraph 2026-07-17 18:30:08 +08:00
zhuyongxin 1460dd1e99 feat(graph): add diagnosis real nodes 2026-07-17 16:44:05 +08:00
zhuyongxin 42ba204532 feat(graph): add diagnosis routing skeleton 2026-07-17 11:17:13 +08:00
zhuyongxin 581daffdad docs(openspec): archive stategraph design freeze 2026-07-17 10:24:13 +08:00
zhuyongxin a36fe72639 docs(mvp): plan chat diagnosis stategraph refactor 2026-07-16 18:59:35 +08:00
zhuyongxin 30d3296043 docs(mvp): archive session run trace issue 2026-07-11 16:46:20 +08:00
zhuyongxin 3578709896 docs(openspec): archive session run isolation 2026-07-10 22:52:39 +08:00
zhuyongxin f9df94377b feat(trace): finish run-aware demo verification 2026-07-10 21:37:43 +08:00
zhuyongxin 78c1477198 feat(trace): isolate aiops runs 2026-07-10 20:57:46 +08:00
zhuyongxin d928a1968a feat(trace): bind feedback to runs 2026-07-10 20:29:56 +08:00
zhuyongxin 027aed1eeb feat(trace): add run-scoped trace reads 2026-07-10 20:07:23 +08:00
zhuyongxin 26d5529280 feat(trace): isolate chat runs 2026-07-10 19:02:04 +08:00
zhuyongxin 6fdbd34bab docs(openspec): tighten run isolation contract 2026-07-10 17:56:51 +08:00
zhuyongxin 52bf0302c6 feat(trace): add session run isolation schema 2026-07-10 17:47:56 +08:00
zhuyongxin 841437fa06 docs(mvp): organize mvp documentation 2026-07-09 13:40:08 +08:00
zhuyongxin 9c9a0024d4 feat(demo): add interview quality audit 2026-07-09 11:18:49 +08:00
zhuyongxin a6c2d4459c docs(openspec): propose interview demo quality audit 2026-07-09 10:34:33 +08:00
zhuyongxin da45fa3fb0 docs(devflow): sort index by date 2026-07-09 10:24:35 +08:00
aruo db0f229285 feat(eval): add evidence pipeline acceptance closure 2026-07-09 00:47:48 +08:00
aruo a77c947cd4 docs(architecture): align evidence pipeline design 2026-07-08 23:53:21 +08:00
aruo 9a84b3de34 feat(agent): support no-evidence references 2026-07-08 23:30:00 +08:00
aruo 7b8c75e571 feat(agent): harden verifier evidence references 2026-07-08 16:12:56 +08:00
aruo a08672b31e feat(eval): add executor audit closure checks 2026-07-08 10:21:39 +08:00
aruo 6015bcbf6f feat(agent): add composer final answer 2026-07-08 09:51:07 +08:00
aruo a5b4502c72 docs(openspec): propose executor composer final answer 2026-07-08 02:49:46 +08:00
aruo 39c0c5f8be docs(mvp): clarify executor v2 implementation issue 2026-07-08 02:43:12 +08:00
aruo 1b31e78be5 feat(agent): add verifier claim checks 2026-07-08 02:33:02 +08:00
aruo c5e496e715 feat(agent): add executor gatekeeper hook 2026-07-08 02:01:49 +08:00
aruo 050cbc8fee feat(agent): add executor evidence v2 contract 2026-07-08 01:37:15 +08:00
zhuyongxin a6afbfaa9d chore: add editorconfig 2026-07-07 21:08:51 +08:00
zhuyongxin 0ee27eb523 feat(trace): improve session workbench review 2026-07-07 19:06:18 +08:00
zhuyongxin 3b62a8940c chore(openspec): archive executor evidence output contract 2026-07-07 19:04:24 +08:00
zhuyongxin 04eb50e2b4 feat(agent): add executor evidence output contract 2026-07-07 19:02:02 +08:00
zhuyongxin 7f2e47ca38 docs(mvp): record executor evidence loop design 2026-07-07 18:51:47 +08:00
aruo aa035b828c feat(trace): add diagnosis trace workbench 2026-07-07 01:24:56 +08:00
aruo b3315ead52 fix(agent): harden live diagnosis skill observability 2026-07-07 00:20:34 +08:00
zhuyongxin 64adb998cf chore(rag): add eval knowledge base mirror 2026-07-06 21:48:21 +08:00
zhuyongxin ed7efc58b7 feat(rag): close eval pipeline with live snapshots 2026-07-06 21:39:27 +08:00
zhuyongxin cf3333d607 feat(rag): modularize knowledge retrieval pipeline 2026-07-06 17:06:05 +08:00
zhuyongxin a375daead7 fix: clean up aiops mojibake text 2026-07-06 11:38:01 +08:00
zhuyongxin a5a0e0c6be Merge remote-tracking branch 'origin/refactor/mvp1.0' into refactor/mvp1.0 2026-07-06 10:54:45 +08:00
zhuyongxin 9e8e20b3b5 chore: update agent instructions 2026-07-06 10:47:54 +08:00
aruo 3dfe3dbe53 Update devflow glossary for skills 2026-07-06 10:43:21 +08:00
zhuyongxin 37083fc92a chore: add agent skills 2026-07-06 10:18:10 +08:00
546 changed files with 43213 additions and 4700 deletions
+117
View File
@@ -0,0 +1,117 @@
---
name: diagnose
description: Disciplined diagnosis loop for hard bugs and performance regressions. Reproduce → minimise → hypothesise → instrument → fix → regression-test. Use when user says "diagnose this" / "debug this", reports a bug, says something is broken/throwing/failing, or describes a performance regression.
---
# Diagnose
A discipline for hard bugs. Skip phases only when explicitly justified.
When exploring the codebase, use the project's domain glossary to get a clear mental model of the relevant modules, and check ADRs in the area you're touching.
## Phase 1 — Build a feedback loop
**This is the skill.** Everything else is mechanical. If you have a fast, deterministic, agent-runnable pass/fail signal for the bug, you will find the cause — bisection, hypothesis-testing, and instrumentation all just consume that signal. If you don't have one, no amount of staring at code will save you.
Spend disproportionate effort here. **Be aggressive. Be creative. Refuse to give up.**
### Ways to construct one — try them in roughly this order
1. **Failing test** at whatever seam reaches the bug — unit, integration, e2e.
2. **Curl / HTTP script** against a running dev server.
3. **CLI invocation** with a fixture input, diffing stdout against a known-good snapshot.
4. **Headless browser script** (Playwright / Puppeteer) — drives the UI, asserts on DOM/console/network.
5. **Replay a captured trace.** Save a real network request / payload / event log to disk; replay it through the code path in isolation.
6. **Throwaway harness.** Spin up a minimal subset of the system (one service, mocked deps) that exercises the bug code path with a single function call.
7. **Property / fuzz loop.** If the bug is "sometimes wrong output", run 1000 random inputs and look for the failure mode.
8. **Bisection harness.** If the bug appeared between two known states (commit, dataset, version), automate "boot at state X, check, repeat" so you can `git bisect run` it.
9. **Differential loop.** Run the same input through old-version vs new-version (or two configs) and diff outputs.
10. **HITL bash script.** Last resort. If a human must click, drive _them_ with `scripts/hitl-loop.template.sh` so the loop is still structured. Captured output feeds back to you.
Build the right feedback loop, and the bug is 90% fixed.
### Iterate on the loop itself
Treat the loop as a product. Once you have _a_ loop, ask:
- Can I make it faster? (Cache setup, skip unrelated init, narrow the test scope.)
- Can I make the signal sharper? (Assert on the specific symptom, not "didn't crash".)
- Can I make it more deterministic? (Pin time, seed RNG, isolate filesystem, freeze network.)
A 30-second flaky loop is barely better than no loop. A 2-second deterministic loop is a debugging superpower.
### Non-deterministic bugs
The goal is not a clean repro but a **higher reproduction rate**. Loop the trigger 100×, parallelise, add stress, narrow timing windows, inject sleeps. A 50%-flake bug is debuggable; 1% is not — keep raising the rate until it's debuggable.
### When you genuinely cannot build a loop
Stop and say so explicitly. List what you tried. Ask the user for: (a) access to whatever environment reproduces it, (b) a captured artifact (HAR file, log dump, core dump, screen recording with timestamps), or (c) permission to add temporary production instrumentation. Do **not** proceed to hypothesise without a loop.
Do not proceed to Phase 2 until you have a loop you believe in.
## Phase 2 — Reproduce
Run the loop. Watch the bug appear.
Confirm:
- [ ] The loop produces the failure mode the **user** described — not a different failure that happens to be nearby. Wrong bug = wrong fix.
- [ ] The failure is reproducible across multiple runs (or, for non-deterministic bugs, reproducible at a high enough rate to debug against).
- [ ] You have captured the exact symptom (error message, wrong output, slow timing) so later phases can verify the fix actually addresses it.
Do not proceed until you reproduce the bug.
## Phase 3 — Hypothesise
Generate **3–5 ranked hypotheses** before testing any of them. Single-hypothesis generation anchors on the first plausible idea.
Each hypothesis must be **falsifiable**: state the prediction it makes.
> Format: "If <X> is the cause, then <changing Y> will make the bug disappear / <changing Z> will make it worse."
If you cannot state the prediction, the hypothesis is a vibe — discard or sharpen it.
**Show the ranked list to the user before testing.** They often have domain knowledge that re-ranks instantly ("we just deployed a change to #3"), or know hypotheses they've already ruled out. Cheap checkpoint, big time saver. Don't block on it — proceed with your ranking if the user is AFK.
## Phase 4 — Instrument
Each probe must map to a specific prediction from Phase 3. **Change one variable at a time.**
Tool preference:
1. **Debugger / REPL inspection** if the env supports it. One breakpoint beats ten logs.
2. **Targeted logs** at the boundaries that distinguish hypotheses.
3. Never "log everything and grep".
**Tag every debug log** with a unique prefix, e.g. `[DEBUG-a4f2]`. Cleanup at the end becomes a single grep. Untagged logs survive; tagged logs die.
**Perf branch.** For performance regressions, logs are usually wrong. Instead: establish a baseline measurement (timing harness, `performance.now()`, profiler, query plan), then bisect. Measure first, fix second.
## Phase 5 — Fix + regression test
Write the regression test **before the fix** — but only if there is a **correct seam** for it.
A correct seam is one where the test exercises the **real bug pattern** as it occurs at the call site. If the only available seam is too shallow (single-caller test when the bug needs multiple callers, unit test that can't replicate the chain that triggered the bug), a regression test there gives false confidence.
**If no correct seam exists, that itself is the finding.** Note it. The codebase architecture is preventing the bug from being locked down. Flag this for the next phase.
If a correct seam exists:
1. Turn the minimised repro into a failing test at that seam.
2. Watch it fail.
3. Apply the fix.
4. Watch it pass.
5. Re-run the Phase 1 feedback loop against the original (un-minimised) scenario.
## Phase 6 — Cleanup + post-mortem
Required before declaring done:
- [ ] Original repro no longer reproduces (re-run the Phase 1 loop)
- [ ] Regression test passes (or absence of seam is documented)
- [ ] All `[DEBUG-...]` instrumentation removed (`grep` the prefix)
- [ ] Throwaway prototypes deleted (or moved to a clearly-marked debug location)
- [ ] The hypothesis that turned out correct is stated in the commit / PR message — so the next debugger learns
**Then ask: what would have prevented this bug?** If the answer involves architectural change (no good test seam, tangled callers, hidden coupling) hand off to the `/improve-codebase-architecture` skill with the specifics. Make the recommendation **after** the fix is in, not before — you have more information now than when you started.
@@ -0,0 +1,47 @@
# ADR Format
ADRs live in `docs/adr/` and use sequential numbering: `0001-slug.md`, `0002-slug.md`, etc.
Create the `docs/adr/` directory lazily — only when the first ADR is needed.
## Template
```md
# {Short title of the decision}
{1-3 sentences: what's the context, what did we decide, and why.}
```
That's it. An ADR can be a single paragraph. The value is in recording *that* a decision was made and *why* — not in filling out sections.
## Optional sections
Only include these when they add genuine value. Most ADRs won't need them.
- **Status** frontmatter (`proposed | accepted | deprecated | superseded by ADR-NNNN`) — useful when decisions are revisited
- **Considered Options** — only when the rejected alternatives are worth remembering
- **Consequences** — only when non-obvious downstream effects need to be called out
## Numbering
Scan `docs/adr/` for the highest existing number and increment by one.
## When to offer an ADR
All three of these must be true:
1. **Hard to reverse** — the cost of changing your mind later is meaningful
2. **Surprising without context** — a future reader will look at the code and wonder "why on earth did they do it this way?"
3. **The result of a real trade-off** — there were genuine alternatives and you picked one for specific reasons
If a decision is easy to reverse, skip it — you'll just reverse it. If it's not surprising, nobody will wonder why. If there was no real alternative, there's nothing to record beyond "we did the obvious thing."
### What qualifies
- **Architectural shape.** "We're using a monorepo." "The write model is event-sourced, the read model is projected into Postgres."
- **Integration patterns between contexts.** "Ordering and Billing communicate via domain events, not synchronous HTTP."
- **Technology choices that carry lock-in.** Database, message bus, auth provider, deployment target. Not every library — just the ones that would take a quarter to swap out.
- **Boundary and scope decisions.** "Customer data is owned by the Customer context; other contexts reference it by ID only." The explicit no-s are as valuable as the yes-s.
- **Deliberate deviations from the obvious path.** "We're using manual SQL instead of an ORM because X." Anything where a reasonable reader would assume the opposite. These stop the next engineer from "fixing" something that was deliberate.
- **Constraints not visible in the code.** "We can't use AWS because of compliance requirements." "Response times must be under 200ms because of the partner API contract."
- **Rejected alternatives when the rejection is non-obvious.** If you considered GraphQL and picked REST for subtle reasons, record it — otherwise someone will suggest GraphQL again in six months.
@@ -0,0 +1,77 @@
# CONTEXT.md Format
## Structure
```md
# {Context Name}
{One or two sentence description of what this context is and why it exists.}
## Language
**Order**:
{A concise description of the term}
_Avoid_: Purchase, transaction
**Invoice**:
A request for payment sent to a customer after delivery.
_Avoid_: Bill, payment request
**Customer**:
A person or organization that places orders.
_Avoid_: Client, buyer, account
## Relationships
- An **Order** produces one or more **Invoices**
- An **Invoice** belongs to exactly one **Customer**
## Example dialogue
> **Dev:** "When a **Customer** places an **Order**, do we create the **Invoice** immediately?"
> **Domain expert:** "No — an **Invoice** is only generated once a **Fulfillment** is confirmed."
## Flagged ambiguities
- "account" was used to mean both **Customer** and **User** — resolved: these are distinct concepts.
```
## Rules
- **Be opinionated.** When multiple words exist for the same concept, pick the best one and list the others as aliases to avoid.
- **Flag conflicts explicitly.** If a term is used ambiguously, call it out in "Flagged ambiguities" with a clear resolution.
- **Keep definitions tight.** One sentence max. Define what it IS, not what it does.
- **Show relationships.** Use bold term names and express cardinality where obvious.
- **Only include terms specific to this project's context.** General programming concepts (timeouts, error types, utility patterns) don't belong even if the project uses them extensively. Before adding a term, ask: is this a concept unique to this context, or a general programming concept? Only the former belongs.
- **Group terms under subheadings** when natural clusters emerge. If all terms belong to a single cohesive area, a flat list is fine.
- **Write an example dialogue.** A conversation between a dev and a domain expert that demonstrates how the terms interact naturally and clarifies boundaries between related concepts.
## Single vs multi-context repos
**Single context (most repos):** One `CONTEXT.md` at the repo root.
**Multiple contexts:** A `CONTEXT-MAP.md` at the repo root lists the contexts, where they live, and how they relate to each other:
```md
# Context Map
## Contexts
- [Ordering](./src/ordering/CONTEXT.md) — receives and tracks customer orders
- [Billing](./src/billing/CONTEXT.md) — generates invoices and processes payments
- [Fulfillment](./src/fulfillment/CONTEXT.md) — manages warehouse picking and shipping
## Relationships
- **Ordering → Fulfillment**: Ordering emits `OrderPlaced` events; Fulfillment consumes them to start picking
- **Fulfillment → Billing**: Fulfillment emits `ShipmentDispatched` events; Billing consumes them to generate invoices
- **Ordering ↔ Billing**: Shared types for `CustomerId` and `Money`
```
The skill infers which structure applies:
- If `CONTEXT-MAP.md` exists, read it to find contexts
- If only a root `CONTEXT.md` exists, single context
- If neither exists, create a root `CONTEXT.md` lazily when the first term is resolved
When multiple contexts exist, infer which one the current topic relates to. If unclear, ask.
+88
View File
@@ -0,0 +1,88 @@
---
name: grill-with-docs
description: Grilling session that challenges your plan against the existing domain model, sharpens terminology, and updates documentation (CONTEXT.md, ADRs) inline as decisions crystallise. Use when user wants to stress-test a plan against their project's language and documented decisions.
---
<what-to-do>
Interview me relentlessly about every aspect of this plan until we reach a shared understanding. Walk down each branch of the design tree, resolving dependencies between decisions one-by-one. For each question, provide your recommended answer.
Ask the questions one at a time, waiting for feedback on each question before continuing.
If a question can be answered by exploring the codebase, explore the codebase instead.
</what-to-do>
<supporting-info>
## Domain awareness
During codebase exploration, also look for existing documentation:
### File structure
Most repos have a single context:
```
/
├── CONTEXT.md
├── docs/
│ └── adr/
│ ├── 0001-event-sourced-orders.md
│ └── 0002-postgres-for-write-model.md
└── src/
```
If a `CONTEXT-MAP.md` exists at the root, the repo has multiple contexts. The map points to where each one lives:
```
/
├── CONTEXT-MAP.md
├── docs/
│ └── adr/ ← system-wide decisions
├── src/
│ ├── ordering/
│ │ ├── CONTEXT.md
│ │ └── docs/adr/ ← context-specific decisions
│ └── billing/
│ ├── CONTEXT.md
│ └── docs/adr/
```
Create files lazily — only when you have something to write. If no `CONTEXT.md` exists, create one when the first term is resolved. If no `docs/adr/` exists, create it when the first ADR is needed.
## During the session
### Challenge against the glossary
When the user uses a term that conflicts with the existing language in `CONTEXT.md`, call it out immediately. "Your glossary defines 'cancellation' as X, but you seem to mean Y — which is it?"
### Sharpen fuzzy language
When the user uses vague or overloaded terms, propose a precise canonical term. "You're saying 'account' — do you mean the Customer or the User? Those are different things."
### Discuss concrete scenarios
When domain relationships are being discussed, stress-test them with specific scenarios. Invent scenarios that probe edge cases and force the user to be precise about the boundaries between concepts.
### Cross-reference with code
When the user states how something works, check whether the code agrees. If you find a contradiction, surface it: "Your code cancels entire Orders, but you just said partial cancellation is possible — which is right?"
### Update CONTEXT.md inline
When a term is resolved, update `CONTEXT.md` right there. Don't batch these up — capture them as they happen. Use the format in [CONTEXT-FORMAT.md](./CONTEXT-FORMAT.md).
`CONTEXT.md` should be totally devoid of implementation details. Do not treat `CONTEXT.md` as a spec, a scratch pad, or a repository for implementation decisions. It is a glossary and nothing else.
### Offer ADRs sparingly
Only offer to create an ADR when all three are true:
1. **Hard to reverse** — the cost of changing your mind later is meaningful
2. **Surprising without context** — a future reader will wonder "why did they do it this way?"
3. **The result of a real trade-off** — there were genuine alternatives and you picked one for specific reasons
If any of the three is missing, skip the ADR. Use the format in [ADR-FORMAT.md](./ADR-FORMAT.md).
</supporting-info>
+109
View File
@@ -0,0 +1,109 @@
---
name: tdd
description: Test-driven development with red-green-refactor loop. Use when user wants to build features or fix bugs using TDD, mentions "red-green-refactor", wants integration tests, or asks for test-first development.
---
# Test-Driven Development
## Philosophy
**Core principle**: Tests should verify behavior through public interfaces, not implementation details. Code can change entirely; tests shouldn't.
**Good tests** are integration-style: they exercise real code paths through public APIs. They describe _what_ the system does, not _how_ it does it. A good test reads like a specification - "user can checkout with valid cart" tells you exactly what capability exists. These tests survive refactors because they don't care about internal structure.
**Bad tests** are coupled to implementation. They mock internal collaborators, test private methods, or verify through external means (like querying a database directly instead of using the interface). The warning sign: your test breaks when you refactor, but behavior hasn't changed. If you rename an internal function and tests fail, those tests were testing implementation, not behavior.
See [tests.md](tests.md) for examples and [mocking.md](mocking.md) for mocking guidelines.
## Anti-Pattern: Horizontal Slices
**DO NOT write all tests first, then all implementation.** This is "horizontal slicing" - treating RED as "write all tests" and GREEN as "write all code."
This produces **crap tests**:
- Tests written in bulk test _imagined_ behavior, not _actual_ behavior
- You end up testing the _shape_ of things (data structures, function signatures) rather than user-facing behavior
- Tests become insensitive to real changes - they pass when behavior breaks, fail when behavior is fine
- You outrun your headlights, committing to test structure before understanding the implementation
**Correct approach**: Vertical slices via tracer bullets. One test → one implementation → repeat. Each test responds to what you learned from the previous cycle. Because you just wrote the code, you know exactly what behavior matters and how to verify it.
```
WRONG (horizontal):
RED: test1, test2, test3, test4, test5
GREEN: impl1, impl2, impl3, impl4, impl5
RIGHT (vertical):
RED→GREEN: test1→impl1
RED→GREEN: test2→impl2
RED→GREEN: test3→impl3
...
```
## Workflow
### 1. Planning
When exploring the codebase, use the project's domain glossary so that test names and interface vocabulary match the project's language, and respect ADRs in the area you're touching.
Before writing any code:
- [ ] Confirm with user what interface changes are needed
- [ ] Confirm with user which behaviors to test (prioritize)
- [ ] Identify opportunities for [deep modules](deep-modules.md) (small interface, deep implementation)
- [ ] Design interfaces for [testability](interface-design.md)
- [ ] List the behaviors to test (not implementation steps)
- [ ] Get user approval on the plan
Ask: "What should the public interface look like? Which behaviors are most important to test?"
**You can't test everything.** Confirm with the user exactly which behaviors matter most. Focus testing effort on critical paths and complex logic, not every possible edge case.
### 2. Tracer Bullet
Write ONE test that confirms ONE thing about the system:
```
RED: Write test for first behavior → test fails
GREEN: Write minimal code to pass → test passes
```
This is your tracer bullet - proves the path works end-to-end.
### 3. Incremental Loop
For each remaining behavior:
```
RED: Write next test → fails
GREEN: Minimal code to pass → passes
```
Rules:
- One test at a time
- Only enough code to pass current test
- Don't anticipate future tests
- Keep tests focused on observable behavior
### 4. Refactor
After all tests pass, look for [refactor candidates](refactoring.md):
- [ ] Extract duplication
- [ ] Deepen modules (move complexity behind simple interfaces)
- [ ] Apply SOLID principles where natural
- [ ] Consider what new code reveals about existing code
- [ ] Run tests after each refactor step
**Never refactor while RED.** Get to GREEN first.
## Checklist Per Cycle
```
[ ] Test describes behavior, not implementation
[ ] Test uses public interface only
[ ] Test would survive internal refactor
[ ] Code is minimal for this test
[ ] No speculative features added
```
+33
View File
@@ -0,0 +1,33 @@
# Deep Modules
From "A Philosophy of Software Design":
**Deep module** = small interface + lots of implementation
```
┌─────────────────────┐
│ Small Interface │ ← Few methods, simple params
├─────────────────────┤
│ │
│ │
│ Deep Implementation│ ← Complex logic hidden
│ │
│ │
└─────────────────────┘
```
**Shallow module** = large interface + little implementation (avoid)
```
┌─────────────────────────────────┐
│ Large Interface │ ← Many methods, complex params
├─────────────────────────────────┤
│ Thin Implementation │ ← Just passes through
└─────────────────────────────────┘
```
When designing interfaces, ask:
- Can I reduce the number of methods?
- Can I simplify the parameters?
- Can I hide more complexity inside?
+31
View File
@@ -0,0 +1,31 @@
# Interface Design for Testability
Good interfaces make testing natural:
1. **Accept dependencies, don't create them**
```typescript
// Testable
function processOrder(order, paymentGateway) {}
// Hard to test
function processOrder(order) {
const gateway = new StripeGateway();
}
```
2. **Return results, don't produce side effects**
```typescript
// Testable
function calculateDiscount(cart): Discount {}
// Hard to test
function applyDiscount(cart): void {
cart.total -= discount;
}
```
3. **Small surface area**
- Fewer methods = fewer tests needed
- Fewer params = simpler test setup
+59
View File
@@ -0,0 +1,59 @@
# When to Mock
Mock at **system boundaries** only:
- External APIs (payment, email, etc.)
- Databases (sometimes - prefer test DB)
- Time/randomness
- File system (sometimes)
Don't mock:
- Your own classes/modules
- Internal collaborators
- Anything you control
## Designing for Mockability
At system boundaries, design interfaces that are easy to mock:
**1. Use dependency injection**
Pass external dependencies in rather than creating them internally:
```typescript
// Easy to mock
function processPayment(order, paymentClient) {
return paymentClient.charge(order.total);
}
// Hard to mock
function processPayment(order) {
const client = new StripeClient(process.env.STRIPE_KEY);
return client.charge(order.total);
}
```
**2. Prefer SDK-style interfaces over generic fetchers**
Create specific functions for each external operation instead of one generic function with conditional logic:
```typescript
// GOOD: Each function is independently mockable
const api = {
getUser: (id) => fetch(`/users/${id}`),
getOrders: (userId) => fetch(`/users/${userId}/orders`),
createOrder: (data) => fetch('/orders', { method: 'POST', body: data }),
};
// BAD: Mocking requires conditional logic inside the mock
const api = {
fetch: (endpoint, options) => fetch(endpoint, options),
};
```
The SDK approach means:
- Each mock returns one specific shape
- No conditional logic in test setup
- Easier to see which endpoints a test exercises
- Type safety per endpoint
+10
View File
@@ -0,0 +1,10 @@
# Refactor Candidates
After TDD cycle, look for:
- **Duplication** → Extract function/class
- **Long methods** → Break into private helpers (keep tests on public interface)
- **Shallow modules** → Combine or deepen
- **Feature envy** → Move logic to where data lives
- **Primitive obsession** → Introduce value objects
- **Existing code** the new code reveals as problematic
+61
View File
@@ -0,0 +1,61 @@
# Good and Bad Tests
## Good Tests
**Integration-style**: Test through real interfaces, not mocks of internal parts.
```typescript
// GOOD: Tests observable behavior
test("user can checkout with valid cart", async () => {
const cart = createCart();
cart.add(product);
const result = await checkout(cart, paymentMethod);
expect(result.status).toBe("confirmed");
});
```
Characteristics:
- Tests behavior users/callers care about
- Uses public API only
- Survives internal refactors
- Describes WHAT, not HOW
- One logical assertion per test
## Bad Tests
**Implementation-detail tests**: Coupled to internal structure.
```typescript
// BAD: Tests implementation details
test("checkout calls paymentService.process", async () => {
const mockPayment = jest.mock(paymentService);
await checkout(cart, payment);
expect(mockPayment.process).toHaveBeenCalledWith(cart.total);
});
```
Red flags:
- Mocking internal collaborators
- Testing private methods
- Asserting on call counts/order
- Test breaks when refactoring without behavior change
- Test name describes HOW not WHAT
- Verifying through external means instead of interface
```typescript
// BAD: Bypasses interface to verify
test("createUser saves to database", async () => {
await createUser({ name: "Alice" });
const row = await db.query("SELECT * FROM users WHERE name = ?", ["Alice"]);
expect(row).toBeDefined();
});
// GOOD: Verifies through interface
test("createUser makes user retrievable", async () => {
const user = await createUser({ name: "Alice" });
const retrieved = await getUser(user.id);
expect(retrieved.name).toBe("Alice");
});
```
+76
View File
@@ -0,0 +1,76 @@
---
name: to-prd
description: Turn the current conversation context into a PRD and publish it to the project issue tracker. Use when user wants to create a PRD from the current context.
---
This skill takes the current conversation context and codebase understanding and produces a PRD. Do NOT interview the user — just synthesize what you already know.
The issue tracker and triage label vocabulary should have been provided to you — run `/setup-matt-pocock-skills` if not.
## Process
1. Explore the repo to understand the current state of the codebase, if you haven't already. Use the project's domain glossary vocabulary throughout the PRD, and respect any ADRs in the area you're touching.
2. Sketch out the major modules you will need to build or modify to complete the implementation. Actively look for opportunities to extract deep modules that can be tested in isolation.
A deep module (as opposed to a shallow module) is one which encapsulates a lot of functionality in a simple, testable interface which rarely changes.
Check with the user that these modules match their expectations. Check with the user which modules they want tests written for.
3. Write the PRD using the template below, then publish it to the project issue tracker. Apply the `ready-for-agent` triage label - no need for additional triage.
<prd-template>
## Problem Statement
The problem that the user is facing, from the user's perspective.
## Solution
The solution to the problem, from the user's perspective.
## User Stories
A LONG, numbered list of user stories. Each user story should be in the format of:
1. As an <actor>, I want a <feature>, so that <benefit>
<user-story-example>
1. As a mobile bank customer, I want to see balance on my accounts, so that I can make better informed decisions about my spending
</user-story-example>
This list of user stories should be extremely extensive and cover all aspects of the feature.
## Implementation Decisions
A list of implementation decisions that were made. This can include:
- The modules that will be built/modified
- The interfaces of those modules that will be modified
- Technical clarifications from the developer
- Architectural decisions
- Schema changes
- API contracts
- Specific interactions
Do NOT include specific file paths or code snippets. They may end up being outdated very quickly.
Exception: if a prototype produced a snippet that encodes a decision more precisely than prose can (state machine, reducer, schema, type shape), inline it within the relevant decision and note briefly that it came from a prototype. Trim to the decision-rich parts — not a working demo, just the important bits.
## Testing Decisions
A list of testing decisions that were made. Include:
- A description of what makes a good test (only test external behavior, not implementation details)
- Which modules will be tested
- Prior art for the tests (i.e. similar types of tests in the codebase)
## Out of Scope
A description of the things that are out of scope for this PRD.
## Further Notes
Any further notes about the feature.
</prd-template>
+7
View File
@@ -0,0 +1,7 @@
---
name: zoom-out
description: Tell the agent to zoom out and give broader context or a higher-level perspective. Use when you're unfamiliar with a section of code or need to understand how it fits into the bigger picture.
disable-model-invocation: true
---
I don't know this area of code well. Go up a layer of abstraction. Give me a map of all the relevant modules and callers, using the project's domain glossary vocabulary.
+13
View File
@@ -0,0 +1,13 @@
root = true
[*]
charset = utf-8
end_of_line = crlf
insert_final_newline = true
trim_trailing_whitespace = true
[*.md]
trim_trailing_whitespace = false
[*.{java,xml,yml,yaml,properties,json,sql,txt,ps1}]
charset = utf-8
-1
View File
@@ -49,7 +49,6 @@ uploads/
### Temp Scripts ###
*.sh
*.py
### docker
/volumes
+106 -42
View File
@@ -1,43 +1,107 @@
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
# CLAUDE.md
## Defaults
- Reply in **Chinese** unless I explicitly ask for English.
- No emojis.
- Do not truncate important outputs (logs, diffs, stack traces, commands, or critical reasoning that affects
safety/correctness).
## Refactor policy (legacy code)
- When existing code is a "big ball of mud" (hard to maintain, clearly bad design,
full of hacks), prefer a **clean, full refactor** over stacking more patches
on top of it.
- A refactor may completely replace internal structure
(functions, modules, classes, data flow).
- By default, try to preserve externally observable behaviour.
If you intentionally change behaviour or protocols, you MUST:
- Call out clearly that this is a **behaviour/protocol change**.
- Explain why the change is necessary and which code paths/consumers are affected.
- Update or add tests to cover the new behaviour.
## Before touching code (mandatory)
Find reuse opportunities + Trace the call/dependency chain and impact radius:
- Use semantic code search first via `codebase-retrieval` tool.
- Confirm understanding with LSP: `goToDefinition`, `findReferences`.
- Use Grep/Glob for verifying and understanding additional code snippets.
## Red lines
- No copy-paste duplication.
- Do not break existing externally observable behaviour **unless**:
- It is part of a deliberate refactor as described in the refactor policy, and
- You clearly document the behavioural change and its impact.
- Do not proceed with a known-wrong approach.
- Critical paths must have explicit error handling.
- Never implement "blindly": always confirm understanding via code reading + references.
## Task sizing
- **Simple**
- Criteria — single file, clear requirement, < 20 lines changed,
clearly local impact.
- Handling — after doing the "Before touching code" steps
(research + impact analysis + internal three-question checklist),
you may execute directly with minimal explanation.
- A very short context line is enough;
a full breakdown of the checklist is not required.
- **Medium**
- Criteria — 2–5 files, or requires some research, or impact is not obviously local.
- Handling — write a short plan (bullet points) → then implement.
- Briefly surface the checklist result in the reply
(1–3 short lines describing real issue, key reuse, and main impact).
- **Complex**
- Criteria — architecture changes, multiple modules, high uncertainty or risk.
- Handling — follow this workflow:
1. **RESEARCH**: inspect code and facts only (no proposals yet).
2. **PLAN**: present options + tradeoffs + recommendation;
use `AskUserQuestion` actively to align with the user;
wait for user's confirmation.
3. **EXECUTE**: implement exactly the approved plan.
4. **REVIEW**: self-check (tests, edge cases, cleanup).
## Git
- Do not commit unless I explicitly ask.
- Do not push unless I explicitly ask.
- Before writing a commit message, glance at a few recent commits and match the repo's style:
- `git log -n 5 --oneline`
- If there is no obvious existing style, use this default format:
- `<type>(<scope>): <description>`
- Before any commit: run `git diff` and confirm the exact scope of changes.
- Never force-push to `main` / `master` unless the user approves.
- Do not add attribution lines in commit messages.
## Security
- Never hardcode secrets (keys/passwords/tokens).
- Never commit `.env` files or any credentials.
- Validate user input at trust boundaries (APIs, CLIs, external data sources).
## Quality & cleanup
- Prefer clarity and simplicity first (KISS); apply DRY to remove obvious
copy-paste duplication when it does not hurt readability.
- If you change a function signature, update **all** call sites.
- After changes:
- Remove temporary files.
- Remove dead/commented-out code.
- Remove unused imports.
- Remove debug logging that is no longer needed.
- Run the smallest meaningful verification (lint/test/build) for the parts you touched.
## Windows / PowerShell (if used)
- PowerShell does not support `&&`; use `;` to chain commands.
- Quote paths that contain spaces or non-ASCII characters.
## Baisc Infos
Unless directly relevant to the user's current question, you should avoid proactively mentioning, illustrating, or
trailing off into the following information in 99% of cases:
This project is indexed by GitNexus as **SuperBizAgent-java** (7988 symbols, 12713 relationships, 297 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
## Always Do
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
## Never Do
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
## Resources
| Resource | Use for |
|----------|---------|
| `gitnexus://repo/SuperBizAgent-java/context` | Codebase overview, check index freshness |
| `gitnexus://repo/SuperBizAgent-java/clusters` | All functional areas |
| `gitnexus://repo/SuperBizAgent-java/processes` | All execution flows |
| `gitnexus://repo/SuperBizAgent-java/process/{name}` | Step-by-step execution trace |
## CLI
| Task | Read this skill file |
|------|---------------------|
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
<!-- gitnexus:end -->
-44
View File
@@ -111,47 +111,3 @@ trailing off into the following information in 99% of cases:
- 文档目录结构:
- 不要将文档放到用户目录(如 `C:\Users\EDY\.claude\`)中
<!-- gitnexus:start -->
# GitNexus — Code Intelligence
This project is indexed by GitNexus as **SuperBizAgent-java** (7988 symbols, 12713 relationships, 297 execution flows). Use the GitNexus MCP tools to understand code, assess impact, and navigate safely.
> If any GitNexus tool warns the index is stale, run `npx gitnexus analyze` in terminal first.
## Always Do
- **MUST run impact analysis before editing any symbol.** Before modifying a function, class, or method, run `gitnexus_impact({target: "symbolName", direction: "upstream"})` and report the blast radius (direct callers, affected processes, risk level) to the user.
- **MUST run `gitnexus_detect_changes()` before committing** to verify your changes only affect expected symbols and execution flows.
- **MUST warn the user** if impact analysis returns HIGH or CRITICAL risk before proceeding with edits.
- When exploring unfamiliar code, use `gitnexus_query({query: "concept"})` to find execution flows instead of grepping. It returns process-grouped results ranked by relevance.
- When you need full context on a specific symbol — callers, callees, which execution flows it participates in — use `gitnexus_context({name: "symbolName"})`.
## Never Do
- NEVER edit a function, class, or method without first running `gitnexus_impact` on it.
- NEVER ignore HIGH or CRITICAL risk warnings from impact analysis.
- NEVER rename symbols with find-and-replace — use `gitnexus_rename` which understands the call graph.
- NEVER commit changes without running `gitnexus_detect_changes()` to check affected scope.
## Resources
| Resource | Use for |
|----------|---------|
| `gitnexus://repo/SuperBizAgent-java/context` | Codebase overview, check index freshness |
| `gitnexus://repo/SuperBizAgent-java/clusters` | All functional areas |
| `gitnexus://repo/SuperBizAgent-java/processes` | All execution flows |
| `gitnexus://repo/SuperBizAgent-java/process/{name}` | Step-by-step execution trace |
## CLI
| Task | Read this skill file |
|------|---------------------|
| Understand architecture / "How does X work?" | `.claude/skills/gitnexus/gitnexus-exploring/SKILL.md` |
| Blast radius / "What breaks if I change X?" | `.claude/skills/gitnexus/gitnexus-impact-analysis/SKILL.md` |
| Trace bugs / "Why is X failing?" | `.claude/skills/gitnexus/gitnexus-debugging/SKILL.md` |
| Rename / extract / split / refactor | `.claude/skills/gitnexus/gitnexus-refactoring/SKILL.md` |
| Tools, resources, schema reference | `.claude/skills/gitnexus/gitnexus-guide/SKILL.md` |
| Index, status, clean, wiki CLI commands | `.claude/skills/gitnexus/gitnexus-cli/SKILL.md` |
<!-- gitnexus:end -->
+68 -2
View File
@@ -72,9 +72,10 @@
### SessionContext
- 定义:会话上下文数据类,存储在 Redis 中的会话数据
- 包含字段:sessionId、userId、businessId、traceId、status、toolCalls、TTL
- 包含字段:sessionId、userId、businessId、traceId、status、toolCalls、messageHistory、TTL
- 序列化方式:JSON(GenericJackson2JsonRedisSerializer)
- 使用场景:多轮对话上下文管理、工具调用历史追踪
- 边界:messageHistory 是热路径对话历史缓存,用于下一轮 prompt 上下文;长期审计的问题和答案应落到 Diagnosis Run,而不是依赖 Redis TTL 内的上下文正文。
### ToolCall
- 定义:工具调用记录数据类,追踪 Agent 使用的工具及其结果
@@ -87,6 +88,26 @@
- 核心方法:createSession、getSession、updateSession、deleteSession、refreshSession、addToolCall
- 使用场景:分布式会话管理、Agent 状态维护
### Chat Session
- 定义:一次多轮对话上下文,由 `sessionId` 唯一标识。
- 使用场景:保存用户连续对话的上下文窗口、会话状态和最近活跃时间。
- 边界:Chat Session 不代表一次诊断执行;同一个 Chat Session 可以包含多次 Diagnosis Run。
### Diagnosis Run
- 定义:一次独立诊断执行,由 `runId` 唯一标识,属于一个 Chat Session。
- 使用场景:保存某一轮诊断的 query、answer、status、耗时、token、反馈和自评估结果。
- 边界:Diagnosis Run 是 Trace、Feedback 和 Evidence score 的绑定对象;多轮对话中的每次 `/api/chat` 或 `/api/ai_ops` 执行都应创建新的 Diagnosis Run。
### Diagnosis Trace
- 定义:一次 Diagnosis Run 的可回放执行轨迹,由 run 主记录、AgentStep 和 ToolInvocation 聚合形成。
- 使用场景:Trace API、Trace UI、Verifier 审计、评测 fixture 和人工排查。
- 边界:Diagnosis Trace 是聚合视图,不要求单独的 trace 主表;当前 trace 明细由 `agent_step` 和 `tool_invocation` 表承载。
### Diagnosis Orchestration Trace
- 定义:一次 Diagnosis Run 的紧凑编排审计摘要,记录实际节点路径、条件边原因、技术重试、降级和终止原因。
- 使用场景:解释诊断编排为何进入某个节点、为何重试或为何提前终止,并支撑路由验收和人工审计。
- 边界:它是 Diagnosis Trace 的编排维度,不是完整事件日志、自评估结果或持久恢复检查点;不保存 Prompt、模型思考、工具原文和完整编排上下文快照。
### Flyway
- 定义:数据库版本迁移工具,管理 SQL 脚本的版本化执行
- 配置:spring.flyway.enabled=true, baseline-on-migrate=true
@@ -105,4 +126,49 @@
- 枚举类型在数据库中存储为 VARCHAR,JPA 使用 `@Enumerated(EnumType.STRING)` + `columnDefinition = "VARCHAR"`
- JPA ddl-auto 使用 `validate` 模式,表结构修改必须通过 Flyway 迁移脚本
- Redis 会话 TTL 由调用方指定,不同场景使用不同过期时间(短诊断 5 分钟,长会话 1 小时)
- Repository 查询方法遵循 Spring Data JPA 命名约定,复杂查询使用 `@Query`
- Repository 查询方法遵循 Spring Data JPA 命名约定,复杂查询使用 `@Query`
## Diagnosis Playbook Skills
### Diagnosis Playbook Skill
- 定义:项目内可版本化的诊断流程包,存放在 `src/main/resources/skills/{skill-name}/SKILL.md`。
- 使用场景:把高频故障诊断流程从大 prompt / 知识库文档中抽出,形成可审查、可复用、可按需加载的 playbook。
- 边界:skill 只定义排查 workflow、证据顺序、停止条件、低置信度行为和报告规则;事实性知识仍放在 `knowledge_base/`,事实证据仍来自 evidence tools。
### SkillRegistry
- 定义:Spring AI Alibaba Agent Framework 的 skill 元数据和正文读取入口。本项目使用 `ClasspathSkillRegistry` 从 classpath `skills/` 加载 skill。
- 使用场景:统一提供 skill `name` / `description` 元数据,并支撑 Executor 通过官方 `read_skill` 读取完整 `SKILL.md`。
- 当前约束:`SkillConfig.SingleSkillRegistry` 临时只暴露 active skill `diagnose-mysql-connection-pool`,用于验证单 skill 流程和避免一次性注入全部 skill。
### PlannerSkillMetadataHook
- 定义:项目本地 hook,只向 Planner 注入结构化 `skill_catalog` 元数据。
- 使用场景:Planner 根据 skill `name` / `description` 选择 `selected_skill`,输出 `selection_reason` 和执行计划。
- 边界:Planner 不暴露官方 `read_skill` 工具,不读取完整 `SKILL.md`;Planner 只能选择 skill,不能执行 skill。
### SkillsAgentHook
- 定义:Spring AI Alibaba 官方 skill hook,会同时注入官方 Skills System prompt,并暴露 `read_skill` 工具。
- 使用场景:只挂到 Executor 和 single-agent Chat;Executor 根据 `planner_plan.selected_skill` 读取完整 playbook 后再调用证据工具。
- 边界:不要挂到 Planner,否则 Planner 会获得 `read_skill` 工具并可能读取完整 skill;Verifier 也不能挂该 hook。
### read_skill
- 定义:官方 skill 读取工具,参数为 `skill_name`,返回对应 `SKILL.md` 正文。
- 使用场景:Executor 在执行场景化诊断前读取 Planner 选中的 playbook。
- 边界:`read_skill` 是流程指导工具,不是事实证据工具;不应作为诊断事实写入 `tool_invocation` 证据链。
### Evidence Tools
- 定义:产生可验证诊断事实的工具集合,包括 `lookup_knowledge`、`query_logs`、`query_metrics`、告警/Prometheus 工具等。
- 使用场景:Executor 按 skill workflow 调用 evidence tools 收集事实,`tool_invocation` 记录这些事实证据。
- 边界:最终诊断结论必须被 evidence tools 支撑,不能仅由 skill 正文支撑。
### Verifier Skill Isolation
- 定义:Chat Verifier 与 skill 系统隔离,只校验 Gatekeeper 投影后的 `verified_executor_output` 和 `verified_evidence`。
- 使用场景:防止 Verifier 把 playbook 指令当作事实证据;Verifier 只判断已有证据是否支持结论。
- 边界:Verifier 不接收 `skill_catalog`,不暴露 `read_skill`,不读取 `SKILL.md`、完整 `tool_trace_summary` 或未经验真的 Executor 自由文本。
## Diagnosis Playbook Business Rules
- Planner 只看 skill metadata,输出 `selected_skill`、`selection_reason` 和 plan。
- Executor 才能调用 `read_skill(selected_skill)`,并且读取 skill 后仍必须调用 evidence tools。
- Skill 正文不得替代 `lookup_knowledge`、日志、指标或告警数据。
- Verifier 只基于 Gatekeeper 通过的 `verified_executor_output` 和 `verified_evidence` 校验事实,不基于 skill 正文、完整工具 Trace 或未验真输出校验事实。
- 当前阶段保留单 active skill 白名单:`diagnose-mysql-connection-pool`。
+37 -20
View File
@@ -2,23 +2,40 @@
## 项目
| 日期 | slug | 领域 | 关键词 | 状态 |
|---|---|---|---|---|
| 2026-07-05 | diagnosis-playbook-skills | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented |
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
| 2026-07-04 | diagnosis-eval-harness | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
| 2026-07-04 | evidence-trace-hardening | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
| 2026-07-03 | mvp-demo-trace-acceptance | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
| 2026-07-04 | aiops-traceable-diagnosis-entry | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
| 2026-07-04 | aiops-alert-scope-control | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
| 2026-05-29 | chatmodel-abstraction | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | archived |
| 2026-06-23 | phase1-infrastructure | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | archived |
| 2026-06-24 | lookup-knowledge-integration | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | archived |
| 2026-06-25 | doc-management-ui | 前端开发/文档管理 | 文档管理页面, CRUD, 状态监控, 纯静态页面, API集成 | archived |
| 2026-06-26 | session-storage | 会话存储/可观测 | diagnosis_session, agent_step, tool_invocation, token追踪, 多Agent路由 | openspec/changes/session-storage | archived |
| 2026-06-29 | confidence-feedback | 质量评估/反馈机制 | evidence_score, selfEvaluation, feedback, useful, not_useful, case_library, BAD_CASE, tool_invocation规则引擎, 反馈按钮, sessionId回传 | openspec/changes/confidence-feedback | archived |
| 2026-06-30 | session-dedup-knowledge-map | 去重/知识图谱 | RetrievedDocTracker, KnowledgeDomainService, knowledge_domain, covers, whenToRetrieve, Planner注入, ISS-001 | openspec/changes/archive/2026-06-30-session-dedup-knowledge-map | archived |
| 2026-07-01 | executor-action-memory-relevance | 检索质量/行动记忆 | relevanceLevel, completenessHint, Min-Max归一化, RetrievedDocTracker域级记录, Executor检索约束, ISS-002 | openspec/changes/archive/2026-07-01-executor-action-memory-relevance | archived |
| 2026-07-02 | chat-verifier-agent | Chat质量门禁/可追溯验证 | Verifier, groundedness_score, facts_checked, evidence_refs, tool_trace_summary, self_evaluation | openspec/changes/archive/2026-07-03-chat-verifier-agent | archived |
| 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|---|---|---|---|---|---|---|
| 2026-07-17 | chat-diagnosis-stategraph-cleanup-docs | 清理旧诊断编排闭包,对齐当前文档与 demo contract,并完成 ISS-011 最终 live、日志和数据库验收。 | Chat diagnosis orchestration/cleanup | legacy closure, current docs, orchestration trace, Maven E2E, MySQL ownership | openspec/changes/archive/2026-07-20-chat-diagnosis-stategraph-cleanup-docs | archived |
| 2026-07-17 | chat-diagnosis-stategraph-test-suite | 建立 Workflow、Node Contract、Chat Integration 三层权威测试体系并退役旧 Hook implementation tests。 | Chat diagnosis orchestration/testing | workflow test, node contract, Chat integration, coverage matrix, Hook test retirement | openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-test-suite | archived |
| 2026-07-17 | chat-diagnosis-stategraph-chatservice-cutover | 将复杂 Chat 单轨切换到 Diagnosis StateGraph,并增加 Run 级 orchestration trace 和 verified-only Verifier 输入。 | Chat diagnosis orchestration/production cutover | ChatService, CompiledGraph stream, runId metadata, orchestration trace, verified-only prompt, V012 | openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-chatservice-cutover | archived |
| 2026-07-17 | chat-diagnosis-stategraph-real-nodes | 接入真实 Agent/Java Nodes、显式 Gatekeeper、可信输入投影、关键证据补查与安全 Fallback,暂不切换生产入口。 | Chat diagnosis orchestration/nodes | ReactAgent adapter, Gatekeeper node, verified input, evidence retry, safe fallback, CompiledGraph | openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-real-nodes | archived |
| 2026-07-17 | chat-diagnosis-stategraph-routing-skeleton | 实现未接生产入口的 Diagnosis StateGraph 骨架、有限路由和 Fake Node 测试。 | Chat diagnosis orchestration/graph | StateGraph, fake node, conditional edge, retry counter, orchestration events, trace builder | openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-routing-skeleton | archived |
| 2026-07-17 | chat-diagnosis-stategraph-design-freeze | 冻结 ISS-011 的 Graph State、条件边、有限重试、安全降级、审计和测试迁移边界。 | Chat diagnosis orchestration/design | StateGraph, runId, Gatekeeper, verified evidence, fallback, orchestration trace, test migration | openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-design-freeze | archived |
| 2026-07-10 | session-run-trace-isolation | 拆分会话态和运行态,引入 runId 隔离 Trace、Feedback、AIOps 和 demo 链路。 | Trace/session/run isolation | chat_session, diagnosis_run, runId, trace exact run, feedback fallback, AIOps SSE metadata, baseline drift | openspec/changes/archive/2026-07-10-session-run-trace-isolation | archived |
| 2026-07-09 | interview-demo-quality-audit | 增加面试演示前置质量审计,覆盖 prompt、Gatekeeper 和评测基线。 | Agent eval/demo/Prompt audit | interview demo preflight, prompt_audit, gatekeeper rules, diagnosis baseline, 12 fixtures | openspec/changes/archive/2026-07-09-interview-demo-quality-audit | archived |
| 2026-07-08 | executor-composer-final-answer | 引入 Composer 生成最终回答,只使用 Verifier 允许的结论材料。 | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
| 2026-07-08 | diagnosis-eval-demo-gatekeeper-closure | 收敛诊断评测、稳定 demo 场景和 Gatekeeper 审计元数据。 | Agent eval/demo/Gatekeeper | diagnosis eval matrix, stable demo scenarios, Gatekeeper rule set version, audit metadata | openspec/changes/archive/2026-07-08-diagnosis-eval-demo-gatekeeper-closure | archived |
| 2026-07-08 | verifier-evidence-reference-fidelity | 强化 Verifier 对 evidence_refs、raw_path 和 no_evidence 的保真校验。 | Chat质量门禁/证据归因 | evidence_refs, raw_path, Gatekeeper severity, verifier evidence excerpt, HikariCP mock, no_evidence | openspec/changes/archive/2026-07-08-verifier-evidence-reference-fidelity | archived |
| 2026-07-07 | executor-evidence-output-contract | 设计 Executor 结构化证据输出,解决证据归因幻觉和 LOW_CONFID 问题。 | Chat质量门禁/证据归因 | Executor structured output, evidence bindings, Verifier structured claims, LOW_CONFID, hallucination | openspec/changes/archive/2026-07-07-executor-evidence-output-contract | archived |
| 2026-07-07 | executor-v2-output-contract | 将 Executor 输出升级为 V2 契约,移除面向用户的最终回答字段。 | Chat质量门禁/证据归因 | executor_evidence_v2, user_facing_answer removal, diagnosis_summary removal, structured renderer | openspec/changes/archive/2026-07-07-executor-v2-output-contract | archived |
| 2026-07-07 | executor-gatekeeper-hook | 在 Executor 与 Verifier 之间接入 Gatekeeper,校验证据绑定来源。 | Chat质量门禁/证据归因 | Gatekeeper, verifier payload, source_invocation_ids, tool_name match, self_evaluation | openspec/changes/archive/2026-07-07-executor-gatekeeper-hook | archived |
| 2026-07-07 | executor-verifier-claim-checks | 增加 Verifier claim_checks 和事实校验兼容逻辑。 | Chat质量门禁/证据归因 | Verifier claim_checks, facts_checked compatibility, effective verdict guardrail, malformed output downgrade | openspec/changes/archive/2026-07-07-executor-verifier-claim-checks | archived |
| 2026-07-06 | rag-eval-pipeline-closure | 建立 RAG 评测闭环,加入 fixture、快照和 baseline diff。 | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
| 2026-07-06 | modular-rag-pipeline | 将 lookup_knowledge 改造成模块化 RAG 管线,补齐证据块和检索追踪。 | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
| 2026-07-05 | diagnosis-playbook-skills | 增加诊断 Playbook Skill,沉淀支付超时、MySQL 池、Redis 超时等套路。 | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented |
| 2026-07-05 | mvp-demo-interview-runbook | 准备可复现的 MVP 面试演示包、运行手册和 Trace 检查清单。 | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
| 2026-07-05 | diagnosis-eval-baseline-diff | 增加诊断评测 baseline diff,用于判断回归和证据覆盖变化。 | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
| 2026-07-04 | expand-diagnosis-eval-fixtures | 扩充诊断评测 fixture,覆盖 Redis、慢响应和 JVM 内存风险。 | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
| 2026-07-04 | diagnosis-eval-harness | 建立固定诊断评测 Harness,输出 trace、证据覆盖和 verdict 分布。 | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
| 2026-07-04 | evidence-trace-hardening | 强化工具调用证据链、降级契约和离线验证能力。 | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
| 2026-07-04 | aiops-traceable-diagnosis-entry | 增加可追踪的 AIOps 告警诊断入口,打通 sessionId 和 Trace API。 | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
| 2026-07-04 | aiops-alert-scope-control | 收敛 AIOps 告警诊断范围,区分 payload 定向和自动发现模式。 | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
| 2026-07-03 | mvp-demo-trace-acceptance | 增加 MVP demo 的 Trace 验收,覆盖会话、步骤、工具和反馈链路。 | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
| 2026-07-02 | chat-verifier-agent | 增加 Chat Verifier Agent,用 groundedness 和 evidence_refs 校验回答。 | Chat质量门禁/可追溯验证 | Verifier, groundedness_score, facts_checked, evidence_refs, tool_trace_summary, self_evaluation | openspec/changes/archive/2026-07-03-chat-verifier-agent | archived |
| 2026-07-01 | executor-action-memory-relevance | 增加行动记忆和相关性信号,约束 Executor 重复检索。 | 检索质量/行动记忆 | relevanceLevel, completenessHint, Min-Max归一化, RetrievedDocTracker域级记录, Executor检索约束, ISS-002 | openspec/changes/archive/2026-07-01-executor-action-memory-relevance | archived |
| 2026-06-30 | session-dedup-knowledge-map | 引入会话级去重和知识域地图,减少重复召回。 | 去重/知识图谱 | RetrievedDocTracker, KnowledgeDomainService, knowledge_domain, covers, whenToRetrieve, Planner注入, ISS-001 | openspec/changes/archive/2026-06-30-session-dedup-knowledge-map | archived |
| 2026-06-29 | confidence-feedback | 建立质量评估和用户反馈机制,并把有用反馈沉淀为案例。 | 质量评估/反馈机制 | evidence_score, selfEvaluation, feedback, useful, not_useful, case_library, BAD_CASE, tool_invocation规则引擎, 反馈按钮, sessionId回传 | openspec/changes/confidence-feedback | archived |
| 2026-06-26 | session-storage | 建立通用会话存储,记录 session、agent step 和 tool invocation。 | 会话存储/可观测 | diagnosis_session, agent_step, tool_invocation, token追踪, 多Agent路由 | openspec/changes/session-storage | archived |
| 2026-06-25 | doc-management-ui | 实现文档管理页面,支持文档 CRUD、状态监控和 API 集成。 | 前端开发/文档管理 | 文档管理页面, CRUD, 状态监控, 纯静态页面, API集成 | - | archived |
| 2026-06-24 | lookup-knowledge-integration | 接入知识库检索,支持 L0 精确匹配和 L1 语义检索。 | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | - | archived |
| 2026-06-23 | phase1-infrastructure | 搭建第一阶段基础设施,包括 MySQL、Redis、Milvus、Flyway 和 JPA。 | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | - | archived |
| 2026-05-29 | chatmodel-abstraction | 抽象 ChatModel 和 EmbeddingModel,支持多模型路由。 | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | - | archived |
@@ -48,7 +48,7 @@
## 遗留问题
ISS-002:Executor 无约束重复调用 `lookup_knowledge`(单会话 20+ 次),knowledge map 和检索约束只注入了 Planner 未注入 Executor。详见 `mvp/issues/ISS-002-executor-unconstrained-lookup.md`。
ISS-002:Executor 无约束重复调用 `lookup_knowledge`(单会话 20+ 次),knowledge map 和检索约束只注入了 Planner 未注入 Executor。详见 `mvp/issues/archived/ISS-002-executor-unconstrained-lookup.md`。
## 已知限制
@@ -9,8 +9,8 @@
## Context
- `devflow/index.md` was checked. Relevant history includes `session-storage`, `confidence-feedback`, `executor-action-memory-relevance`, and `chat-verifier-agent`.
- `mvp/notes/agent-engineering-decisions.md` already recommends the next phase as "可复现 MVP Demo", including `mvp-demo` profile, fixed diagnosis case, one-click request, and `GET /api/diagnosis/{sessionId}/trace`.
- `mvp/issues/ISS-003-mvp-design-implementation-review.md` identifies test stability, session traceability, verifier evidence chain, upload path, and SupervisorAgent consistency as recent MVP concerns. Security cleanup is intentionally deferred by user decision.
- `mvp/archive/2026-07-09-doc-cleanup/notes/agent-engineering-decisions.md` already recommends the next phase as "可复现 MVP Demo", including `mvp-demo` profile, fixed diagnosis case, one-click request, and `GET /api/diagnosis/{sessionId}/trace`.
- `mvp/issues/active/ISS-003-mvp-design-implementation-review.md` identifies test stability, session traceability, verifier evidence chain, upload path, and SupervisorAgent consistency as recent MVP concerns. Security cleanup is intentionally deferred by user decision.
## Question Pool
@@ -0,0 +1,101 @@
# Modular RAG Pipeline — Acceptance
## 验收状态
状态:通过,OpenSpec 已归档。
任务完成:
- OpenSpec tasks:31/31 完成。
- Review 后新增去重边界修复和回归测试。
- OpenSpec archive:`openspec/changes/archive/2026-07-06-modular-rag-pipeline`。
## 静态验证
```powershell
openspec validate modular-rag-pipeline --strict
```
结果:
```text
Change 'modular-rag-pipeline' is valid
```
```powershell
git diff --check
```
结果:
```text
PASS
```
说明:仅出现 Windows LF/CRLF warning,无 whitespace error。
## 脚本验证
```powershell
mvn -q -DskipTests compile
```
结果:PASS。
```powershell
mvn -q "-Dtest=LookupKnowledgeToolTest,ToolInvocationRecorderTest" test
```
结果:PASS。
覆盖:
- filtered L1 成功不 retry。
- filtered L1 低质量触发 raw unfiltered retry。
- filtered L1 无 evidence 触发 raw unfiltered retry。
- L0 hint 不作为 standalone fact evidence。
- 无 L0 hint 时直接 unfiltered vector search。
- rerank 使用 hint match 并记录 trace。
- context pack 保留 source/title/breadcrumb/hit reasons。
- evidence blocks 按 source 去重。
- session dedup 不再返回可消费 evidence/context。
- recorder 记录 retrieval trace、rerank trace、context pack summary 和 evidence summaries。
```powershell
$env:MILVUS_TOKEN = <application.yml 中的 milvus.token>; mvn -q test
```
结果:PASS。
说明:
- `MilvusConnectionTest` 需要 `MILVUS_TOKEN` 环境变量,直接读 `System.getenv`,不会自动读 `application.yml`。
- 注入该环境变量后完整测试通过。
## 实现验收
已验证行为:
- `LookupKnowledgeTool` 已变为 pipeline orchestrator。
- `LookupResult` 新契约包含 `evidenceBlocks`、`contextPack`、`retrievalTrace`、`rerankTrace`。
- 旧 `primary/supplement` 字段和 DTO 已删除。
- `ToolInvocationRecorder` 不再依赖 `result.getPrimary()`。
- filtered retrieval 失败时会记录 `filtered_vector_no_evidence` 或 `filtered_vector_low_quality`。
- no-evidence 情况返回 `found=false` 且保留 retrieval trace。
- session dedup 情况返回 `found=false` 且 evidence/context 为空。
## 未验证项
人工 Demo 未执行:
- 还没有通过真实 Chat/AIOps 会话观察 Agent 是否稳定按 `contextPack.packedText` 和 `evidenceBlocks` 引用证据。
风险:
- 工具 JSON 契约是 L4 breaking change,prompt 已更新,但真实对话行为仍建议做一次端到端 demo。
## 后续建议
- 增加一组 RAG eval cases,固定 query、期望 evidence source、期望 fallback path。
- 将 `MilvusConnectionTest` 改成 Spring 配置驱动或 integration profile,避免配置源混用。
- 后续可在评测数据足够后再考虑 model-based rerank 或 hybrid retrieval。
@@ -0,0 +1,54 @@
# Modular RAG Pipeline — Brief
## 背景
`lookup_knowledge` 已经能返回知识库证据,但实现集中在 `LookupKnowledgeTool` 内部:L0 查询分析、L1 向量召回、相关性归一化、证据组装、会话去重和 trace 入库耦合在一起。
旧返回契约 `primary/supplement` 也延续了“L0 是主结果、L1 是补充”的语义,和当前设计目标不一致。新的目标是让 L0 只作为 query understanding / filter / rerank / trace 信号,让 L1 向量检索成为事实证据来源。
## 目标
- 将 `lookup_knowledge` 改造成模块化 RAG pipeline。
- 保留显式 Agent tool 边界,不改工具名和 query 参数。
- L0 只提供领域、关键词、实体、category filter 和 trace hint。
- L1 filtered vector retrieval 失败或低质量时,降级为 raw query unfiltered L1 retry。
- 输出 evidence-first contract:`evidenceBlocks`、`contextPack`、`retrievalTrace`、`rerankTrace`。
- 保持 `tool_invocation` 表结构稳定,把新 trace 写入 `retrieval_details` JSON。
## 范围
已完成:
- 新增 pipeline DTO:`KnowledgeQuery`、`RetrievedEvidenceCandidate`、`ContextPack`、`RetrievalTrace`、`RerankTrace`、`EvidencePostprocessResult`。
- 新增 pipeline service:`KnowledgeQueryTransformer`、`KnowledgeDocumentRetriever`、`KnowledgeEvidencePostProcessor`、`KnowledgeContextPacker`、`LookupResultAssembler`。
- 重构 `LookupKnowledgeTool` 为薄 orchestration 层。
- 迁移 `LookupResult`,删除 `primary/supplement` 字段和 `PrimaryResult` / `SupplementResult` 类。
- 更新 `ToolInvocationRecorder`,记录 query transform、retrieval trace、context pack summary、rerank trace、fallback reason 和 evidence summaries。
- 更新 executor prompt 和 RAG 架构文档。
- 补充 lookup、recorder、fallback、rerank、context pack、session dedup 测试。
非目标:
- 不引入 implicit Advisor。
- 不引入 cross-encoder、BM25、RRF、Elasticsearch、OpenSearch。
- 不改文档上传、chunk、embedding 写入、Milvus schema。
- 不改变 Agent 何时调用 `lookup_knowledge`。
## 关联 OpenSpec
- `openspec/changes/archive/2026-07-06-modular-rag-pipeline`
## 接口影响
级别:L4 breaking interface。
原因:
- 删除旧 `LookupResult.primary` / `LookupResult.supplement`。
- `lookup_knowledge` tool JSON 输出形状变化。
缓解:
- 工具名和输入参数保持不变。
- in-repo 消费方、测试和 prompt 同步迁移。
- `tool_invocation` 表结构不变。
@@ -0,0 +1,72 @@
# Modular RAG Pipeline — Decisions
## D1: `lookup_knowledge` 保持显式工具
不把知识检索做成隐式 Advisor。Agent 仍显式调用 `lookup_knowledge(query)`,这样 trace、Verifier、Eval 都能看到工具调用边界。
## D2: L0 只做 query understanding
L0 产出:
- `domainHints`
- `matchedKeywords`
- `entities`
- `categoryFilter`
- `l0Titles`
- `l0MatchCount`
L0 不再直接转成 fact evidence。L0 hint 可以影响 filter、rerank、trace,但不能在 L1 失败时冒充知识证据。
## D3: MVP 降级策略采用 unfiltered L1 retry
流程:
```text
filtered L1 with L0 category filter
-> empty / no final evidence / below reference threshold
-> raw query unfiltered L1 retry
-> still no evidence => no_evidence
```
取舍:
- 简单、可解释、适合 MVP。
- 避免引入 BM25/RRF/multi-query/cross-encoder 的复杂度。
- 代价是低质量场景多一次向量查询,已通过 trace 记录 attempt duration。
## D4: 删除 `primary/supplement`
这是一次 L4 breaking interface change。
删除原因:
- `primary/supplement` 绑定旧语义:L0 primary、L1 supplement。
- 新设计中事实证据来自 `evidenceBlocks/contextPack`。
迁移结果:
- `LookupResult` 暴露 evidence-first 字段。
- `PrimaryResult` / `SupplementResult` 已删除。
- 生产代码和测试不再引用 `getPrimary()` / `getSupplement()`。
## D5: Trace 表结构保持稳定
`tool_invocation` 表不新增列。新增信息写入 `retrieval_details` JSON:
- `query_transform`
- `retrieval_trace`
- `context_pack_summary`
- `rerank_trace`
- `fallback_reason`
- `evidence_blocks`
原因:当前 trace、Verifier、Eval 已经以 `tool_invocation` 为证据入口,JSON details 足够承载 RAG 细节,避免 schema churn。
## D6: 会话去重不返回可消费证据
Review 后修正:
- dedup result 的 `found=false` 必须和 evidence/context 语义一致。
- 返回消息说明文档已检索过。
- 不再返回 `evidenceBlocks/contextPack`,避免 Agent 重复使用同一证据。
- 保留 `retrievalTrace` 和 `retrievedDomainsThisSession` 便于可观测。
@@ -0,0 +1,56 @@
# Modular RAG Pipeline — Evidence
## 代码证据
关键入口:
- `src/main/java/com/superbiz/agent/tool/LookupKnowledgeTool.java`
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
- `src/main/java/com/superbiz/agent/dto/LookupResult.java`
新增模块:
- `KnowledgeQueryTransformer`:复用 `KnowledgeIndexService.analyzeQuery`,把 L0 转成 query hints 和可选 `categoryFilter`。
- `KnowledgeDocumentRetriever`:封装 `VectorSearchService.searchSimilarDocuments(query, topK, category)`,统一 filtered / unfiltered attempt。
- `KnowledgeEvidencePostProcessor`:归一化 L2、创建 evidence blocks、source dedup、规则 rerank、输出 `RerankTrace`。
- `KnowledgeContextPacker`:按字符预算打包 evidence,保留 source/title/breadcrumb/hit reasons。
- `LookupResultAssembler`:统一组装 evidence-first result、no-evidence result、session dedup result。
## 设计证据
已有文档约束:
- `mvp/architecture/rag-architecture.md`:RAG 应表达为可解释 pipeline,而不是一坨工具逻辑。
- `mvp/architecture/retrieval-observability.md`:L0 是 hint/explainability 层,L1 是语义检索主路径。
- `devflow/glossary/CONTEXT.md`:`lookup_knowledge` 是显式 Agent evidence tool,`tool_invocation` 是 trace / verifier / eval 的证据来源。
OpenSpec 对齐:
- `openspec/changes/modular-rag-pipeline/proposal.md`
- `openspec/changes/modular-rag-pipeline/design.md`
- `openspec/changes/modular-rag-pipeline/specs/rag-knowledge-retrieval/spec.md`
- `openspec/changes/modular-rag-pipeline/tasks.md`
## 用户确认
- 一次到位做模块化 RAG,而不是只做小补丁。
- L0 不再作为事实证据兜底。
- filtered L1 不准时,MVP 降级为 raw query unfiltered L1 retry。
- 可以新增字段,并删除旧字段以换取后续流程清晰。
## Review 发现
Review 中发现一个非阻塞但应修复的问题:
- 会话去重命中时,返回 `found=false` 但仍带 `evidenceBlocks/contextPack`,可能导致 Agent 重复消费同一份证据。
修复:
- `LookupResultAssembler.deduped` 清空可消费 evidence/context,只保留 message、trace、relevance hint 和 session domain memory。
- 新增 `LookupKnowledgeToolTest.sessionDedupDoesNotReturnConsumableEvidenceAgain`。
## 非阻塞观察
- `MilvusConnectionTest` 仍直接依赖 `MILVUS_TOKEN` 环境变量;主配置中已有 token,但测试不读 Spring 配置。
- 测试日志仍有 ANTLR 版本 warning,不影响测试通过。
- 控制台在部分命令输出中仍会出现中文编码显示问题,但源码按 UTF-8 读取时关键用户提示文本正常。
@@ -0,0 +1,78 @@
# Acceptance: rag-eval-pipeline-closure
## Status
Archived.
## Acceptance Criteria
| Item | Status | Notes |
|---|---|---|
| Modular fixture support | Done | Evaluator reads `lookupResult.evidenceBlocks/contextPack/retrievalTrace/rerankTrace`. |
| LookupResult-only contract | Done | Evaluator fails fixtures that do not expose `lookupResult`. |
| Golden modular assertions | Done | Cases assert selected attempt, accepted fallback reason, evidence status, context sources, and rerank top source. |
| Fallback coverage | Done | Added `chat-l0-filter-fallback` for filtered low-quality/no-evidence to unfiltered retry. |
| Real tool snapshot generation | Done | Added `RagLookupSnapshotGeneratorTest` and `generate_rag_lookup_snapshots.ps1`, defaulting to Spring AI VectorStore mode. |
| Seed docs import/reindex | Done | Added canonical seed docs, `RagEvalSeedImporterTest`, and `prepare_rag_eval_seed.ps1`. |
| Eval metadata isolation | Done | Added `kb_scope` metadata and `retrieval.kb-scope` filtering for L0 and L1. |
| Frontmatter body split | Done | Upload chunking embeds Markdown body, while frontmatter feeds metadata and L0. |
| Baseline diff | Done | `--compare-to` writes JSON/Markdown diff and exits non-zero on regression. |
| Documentation | Done | Updated RAG eval README and added `mvp/architecture/rag-eval-closure.md`. |
## Verification
```powershell
python scripts\eval_rag_retrieval.py
```
Result: passed. 7 cases, passRate=1.0, recall@5=1.0.
```powershell
mvn -q "-Dtest=RagLookupSnapshotGeneratorTest" test
```
Result: passed. The snapshot generator stays disabled unless `rag.snapshot.enabled=true` is provided.
```powershell
mvn -q "-Dtest=RagLookupSnapshotGeneratorTest" "-Drag.snapshot.enabled=true" "-Drag.snapshot.fixtures=<temp-fixtures>" "-Drag.snapshot.retrievedAt=2026-07-06T00:00:00Z" "-Dretrieval.kb-scope=rag-eval" "-Dretrieval.vector-store.mode=spring" test
python scripts\eval_rag_retrieval.py --fixtures <temp-fixtures> --json-report <temp-current.json> --markdown-report <temp-current.md>
```
Result: passed. Spring AI VectorStore live snapshot produced 7 cases, passRate=1.0, recall@5=1.0. The fallback case used `selectedAttempt=UNFILTERED_VECTOR_RETRY` and `fallbackReason=filtered_vector_no_evidence`; the expected source `rag-l0-filter-fallback` remained rank 1.
```powershell
python scripts\eval_rag_retrieval.py --json-report <temp-current.json> --markdown-report <temp-current.md> --compare-to eval\rag-retrieval\reports\baseline.json --diff-json-report <temp-diff.json> --diff-markdown-report <temp-diff.md>
```
Result: passed. regressions=0.
```powershell
mvn -q "-Dtest=FrontmatterParserTest,VectorIndexServiceTest,VectorSearchServiceTest,DocumentManagementServiceTest,RagLookupSnapshotGeneratorTest,RagEvalSeedImporterTest" test
```
Result: passed. The seed importer and snapshot generator remain disabled unless their system properties are explicitly enabled.
```powershell
$null = [scriptblock]::Create((Get-Content -Raw scripts\prepare_rag_eval_seed.ps1))
$null = [scriptblock]::Create((Get-Content -Raw scripts\generate_rag_lookup_snapshots.ps1))
```
Result: PowerShell syntax OK.
```powershell
.\scripts\prepare_rag_eval_seed.ps1
```
Result: passed. Seed docs were imported through `DocumentManagementService` and reindexed into the configured runtime DB/vector stack.
```powershell
.\scripts\generate_rag_lookup_snapshots.ps1 -Fixtures <temp-fixtures> -RetrievedAt 2026-07-06T00:00:00Z -SkipEval
```
Result: passed after defaulting the script to `retrieval.vector-store.mode=spring`. The script generated live fixtures through the real `LookupKnowledgeTool` and then the offline evaluator reported 7 cases, passRate=1.0, recall@5=1.0.
```powershell
git diff --check
```
Result: no whitespace errors. Git reported only LF/CRLF conversion warnings.
@@ -0,0 +1,21 @@
# Brief: rag-eval-pipeline-closure
## Background
The modular RAG pipeline now returns `LookupResult` with `evidenceBlocks`, `contextPack`, `retrievalTrace`, and `rerankTrace`. The RAG retrieval baseline must validate that full contract, so it can detect regressions in fallback behavior, context packing, or rerank trace.
## Goals
1. Reuse the existing offline RAG retrieval baseline.
2. Extend it to support modular `LookupResult` fixtures.
3. Add golden assertions for selected attempt, fallback reason, evidence status, context sources, and rerank top source.
4. Add a RAG baseline diff path for regression detection.
5. Add a snapshot generator that calls the real `LookupKnowledgeTool`.
6. Document how RAG baseline and diagnosis baseline form a quality loop.
## Non-Goals
- No new production API.
- No LLM-as-judge scoring.
- No production API behavior changes.
- No replacement for diagnosis eval.
@@ -0,0 +1,57 @@
# Decisions: rag-eval-pipeline-closure
## D1. Reuse the existing evaluator
Decision: extend `scripts/eval_rag_retrieval.py` instead of creating a second evaluator.
Reason: the old evaluator already owns golden cases, fixtures, hit-level classification, and Markdown/JSON reports. Extending it keeps one RAG baseline path.
## D2. Use LookupResult as the only fixture contract
Decision: support `lookupResult` only.
Reason: the MVP has moved to evidence-first RAG. Keeping an older fixture contract would weaken the baseline and let incomplete fixtures bypass context packing, retrieval trace, and rerank checks.
## D3. Make modular assertions opt-in per case
Decision: use fields such as `expectedSelectedAttempt`, `expectedFallbackReason`/`expectedFallbackReasons`, `expectedEvidenceStatus`, `expectedContextSources`, and `expectedRerankTopSource`.
Reason: golden cases can be strict where the pipeline path matters without forcing every historical case to assert every new field.
## D4. Diff remains deterministic
Decision: RAG diff compares report fields only and does not call live services or models.
Reason: this keeps it suitable for local regression checks and CI-style gates.
## D5. Isolate live eval docs with kb_scope
Decision: add `kb_scope` metadata and use `rag-eval` for canonical eval seed documents.
Reason: local production documents are not stable enough for golden retrieval expectations. Scope isolation lets real `LookupKnowledgeTool` snapshots use the same MySQL/Milvus stack while avoiding accidental matches from unrelated local data.
Default runtime keeps `retrieval.kb-scope` empty so legacy documents without `kb_scope` remain searchable. Eval scripts pass `-Dretrieval.kb-scope=rag-eval`. The same scope applies to L0 query hints and L1 vector retrieval.
## D6. Import seed docs through the real upload pipeline
Decision: seed docs are imported by `RagEvalSeedImporterTest` through `DocumentManagementService.uploadDocument`.
Reason: this updates DB metadata, L0 index state, local knowledge files, and Milvus chunks in the same way as normal document ingestion. A direct Milvus-only seed would make the live eval less representative.
## D7. Strip frontmatter before chunk embedding
Decision: uploaded Markdown frontmatter feeds metadata/L0 but is stripped before chunking and embedding.
Reason: frontmatter is a control plane, not evidence text. Keeping it in chunks lets L0-only keywords artificially improve vector similarity, especially for fallback decoy cases.
## D8. Treat retry behavior as the stable fallback contract
Decision: the fallback golden case accepts both `filtered_vector_low_quality` and `filtered_vector_no_evidence`, while still requiring `selectedAttempt=UNFILTERED_VECTOR_RETRY`, expected evidence source, context packing, and rerank top source.
Reason: Spring AI VectorStore and the Milvus SDK can differ on whether an over-filtered first pass returns a weak candidate or no candidate. The MVP contract is that the retriever skips only the L0 category filter, keeps `kb_scope`, retries the original query, and returns the correct evidence.
## D9. Default live snapshots to Spring AI VectorStore
Decision: `generate_rag_lookup_snapshots.ps1` defaults to `retrieval.vector-store.mode=spring`.
Reason: Spring AI VectorStore is the current framework path for the project and should be the default live verification route. SDK mode remains available through `-VectorStoreMode sdk` for comparison.
@@ -0,0 +1,87 @@
# Evidence: rag-eval-pipeline-closure
## Changed Assets
- `scripts/eval_rag_retrieval.py`
- `eval/rag-retrieval/cases/golden-cases.json`
- `eval/rag-retrieval/fixtures/*.json`
- `eval/rag-retrieval/reports/baseline.json`
- `eval/rag-retrieval/reports/baseline.md`
- `eval/rag-retrieval/README.md`
- `mvp/architecture/rag-eval-closure.md`
- `src/test/java/com/superbiz/agent/eval/RagLookupSnapshotGeneratorTest.java`
- `src/test/java/com/superbiz/agent/eval/RagEvalSeedImporterTest.java`
- `scripts/generate_rag_lookup_snapshots.ps1`
- `scripts/prepare_rag_eval_seed.ps1`
- `eval/rag-retrieval/seed-docs/*.md`
- `src/main/java/com/superbiz/agent/dto/Frontmatter.java`
- `src/main/java/com/superbiz/agent/dto/KnowledgeEntry.java`
- `src/main/java/com/superbiz/agent/service/FrontmatterParser.java`
- `src/main/java/com/superbiz/agent/service/KnowledgeIndexService.java`
- `src/main/java/com/superbiz/agent/service/DocumentManagementService.java`
- `src/main/java/com/superbiz/agent/service/VectorIndexService.java`
- `src/main/java/com/superbiz/agent/service/VectorSearchService.java`
- `src/main/java/com/superbiz/agent/service/SpringAiVectorStoreSidecarService.java`
- `src/main/resources/application.yml`
## Baseline Result
```text
Evaluated 7 cases: passRate=1.0, recall@5=1.0, failed=0
```
## Regression Signals
The evaluator now fails on:
- missing expected source
- missing breadcrumb or evidence keyword
- non-`lookupResult` fixture
- selected attempt mismatch
- fallback reason mismatch
- fallback reason outside accepted values
- evidence status mismatch
- missing context source
- rerank top source mismatch
The snapshot generator now provides:
- real `LookupKnowledgeTool` invocation
- one fixture per golden case
- explicit opt-in through `rag.snapshot.enabled=true`
- optional post-generation baseline evaluation
- scoped retrieval through `retrieval.kb-scope=rag-eval`
- Spring AI VectorStore by default through `retrieval.vector-store.mode=spring`
- scoped L0 hints through the same `retrieval.kb-scope`
The seed importer now provides:
- canonical eval docs under `eval/rag-retrieval/seed-docs`
- real `DocumentManagementService` import/reindex
- stable `source`/`docId` metadata
- `kb_scope=rag-eval` isolation from local non-eval documents
- frontmatter stripping before chunk embedding
- an over-filter decoy seed doc for fallback-path evaluation
The diff now detects:
- aggregate pass/recall regression
- case pass regression
- hit-level regression
- first-rank regression
- selected attempt/fallback/evidence/rerank changes
## Final Spring Live Snapshot
```text
.\scripts\generate_rag_lookup_snapshots.ps1 -Fixtures <temp-fixtures> -RetrievedAt 2026-07-06T00:00:00Z
Evaluated 7 cases: passRate=1.0, recall@5=1.0, failed=0
```
Key fallback trace:
```text
selectedAttempt=UNFILTERED_VECTOR_RETRY
fallbackReason=filtered_vector_no_evidence
rerankTopSource=rag-l0-filter-fallback
```
@@ -0,0 +1,16 @@
# Acceptance: executor-evidence-output-contract
## Draft Acceptance
- [x] Issue exists: `mvp/issues/active/executor-evidence-attribution-hallucination.md`.
- [x] OpenSpec change artifacts exist.
- [x] devflow tracking files exist.
- [x] OpenSpec validation passes.
- [x] Implementation updates Chat Executor prompt.
- [x] Implementation passes structured Executor output to Verifier.
- [x] Verifier prefers structured claims and still falls back safely.
- [x] Focused tests cover parsing, payload assembly, and unsupported confirmed claims.
## Notes
This project is currently in proposal/design stage. Runtime code is intentionally not changed yet.
@@ -0,0 +1,23 @@
# Brief: executor-evidence-output-contract
## Summary
Executor currently returns natural-language diagnosis answers that may mix confirmed evidence, runbook guidance, historical patterns, and unsupported inference. Verifier catches many unsupported facts, but only after extracting claims from prose.
This project defines a structured Executor evidence-attribution contract and updates the Verifier input/verification path to consume it.
## Goal
Make Chat Executor output machine-checkable so confirmed claims are explicitly bound to current-session evidence, while hypotheses and evidence gaps remain visibly separate.
## Scope
- Chat Executor prompt contract.
- Executor structured output parsing.
- Verifier payload extension.
- Chat Verifier prompt behavior.
- Focused tests/eval fixtures.
## Related OpenSpec
`openspec/changes/executor-evidence-output-contract/`
@@ -0,0 +1,71 @@
# Decisions: executor-evidence-output-contract
## sm-flow Progress
### Clarify
Entry summary: recent Chat diagnosis sessions are `LOW_CONFID` because Executor presents unsupported or weakly supported details as confirmed facts after successful tool calls.
Slug: `executor-evidence-output-contract`
Scale: standard. This affects prompts, verifier input assembly, parsing behavior, and tests, but does not require a database schema change.
### Context
Relevant history:
- `executor-action-memory-relevance`: Executor already has retrieval quality constraints and should avoid repeated `lookup_knowledge`.
- `chat-verifier-agent`: Verifier should not see intermediate reasoning; it receives explicit `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
- `evidence-trace-hardening`: evidence-bearing tools persist stable traces and no-evidence semantics.
- `modular-rag-pipeline`: `lookup_knowledge` exposes evidence blocks and context packs; L0 hints are not fact evidence.
Current code shape:
- `src/main/resources/prompts/chat-executor-prompt.md` is the Chat Executor prompt.
- `src/main/resources/prompts/executor-prompt.md` is for the AiOps flow and is not the target of this Chat change.
- `VerifierInputHook` currently builds a payload with `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
- Verifier prompt currently extracts facts from `executor_final_answer`.
### Grill
Question: Should Executor output only JSON or JSON plus readable answer?
Decision: use one JSON object containing both machine fields and `user_facing_answer`. This avoids losing a readable Chinese answer while giving Verifier structured claims.
Question: Should evidence binding use `chunk_id`?
Decision: no. Use generic binding fields because `query_logs` and `query_metrics` do not naturally expose RAG chunks.
Question: Should Verifier trust Executor-provided claims completely?
Decision: no. Verifier should verify structured claims first, then scan `user_facing_answer` for extra confirmed-sounding facts omitted from `claims`.
Question: What happens when Executor JSON is malformed?
Decision: preserve raw final answer, mark parse failure, and fall back to existing natural-language verification.
### Specify
OpenSpec artifacts:
- `proposal.md`: why and scope
- `design.md`: contract, verifier behavior, risks
- `specs/chat-verifier-agent/spec.md`: modified and added requirements
- `tasks.md`: implementation checklist
### Audit
Cross-artifact alignment:
- Issue describes evidence attribution hallucination.
- Proposal scopes the fix to Executor output and Verifier consumption.
- Design preserves existing verifier isolation.
- Spec adds observable behavior without changing database schema.
- Tasks remain implementation-oriented and unchecked.
Interface impact:
- Prompt/output contract: L2 internal Agent contract change.
- Verifier payload: L2 internal structured input extension.
- Database schema: no change.
- External HTTP API: no intended change.
@@ -0,0 +1,17 @@
# Evidence: executor-evidence-output-contract
## Repository Evidence
- `chat-executor-prompt.md` currently requires using real tool data but does not require a structured evidence-attribution output.
- `chat-verifier-prompt.md` currently extracts facts from `executor_final_answer` prose and compares them with `tool_trace_summary`.
- `VerifierInputHook` currently provides `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
- `openspec/specs/chat-verifier-agent/spec.md` already requires explicit verifier inputs, auditable evidence refs, fixed verdicts, and low-confidence handling.
- `openspec/specs/evidence-trace-hardening/spec.md` already distinguishes failed, no-hit, deduped, and successful evidence-tool traces.
## Runtime Evidence From Recent Sessions
Recent MySQL inspection showed repeated `LOW_CONFID` verifier results with many `no_evidence` facts. Typical unsupported claims included OOM, Full GC frequency, specific slow SQL timings, lock waits, and service-specific timeout details that were not supported by current-session tool traces.
## Design Evidence
This change preserves the previous design that Verifier should not inspect intermediate reasoning. The new structured output is still final Executor output, not hidden chain-of-thought.
@@ -0,0 +1,48 @@
# Acceptance: executor-gatekeeper-hook
## Implementation Result
Completed stage two of Executor Structured Output V2.
- Added `ExecutorGatekeeperService`.
- Added initial `schema.executor_v2` and `evidence.invocation_ref` rules.
- Added `gatekeeper_result` to Verifier payload.
- Stored `gatekeeper_result` in `VerifierContextHolder`.
- Persisted `gatekeeper_result` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- Updated Verifier prompt so Gatekeeper fail must not produce PASS.
- Added focused tests for schema failure, valid pass, fabricated invocation ids, tool name mismatch, hook payload, and persistence.
## Static Verification
- `cmd /c openspec validate executor-gatekeeper-hook`
- Result: passed.
- Coverage: OpenSpec change validity.
## Script Verification
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test`
- Result: passed.
- Coverage: 23 focused tests for Gatekeeper service, hook integration, and ChatService persistence.
## Browser / Manual Verification
Not run. This stage changes backend validation and audit behavior only.
## Unverified
- Full live application run with a real LLM.
- MySQL trace inspection after a real chat session.
- Excerpt similarity or phrase/utilization rules.
Reason: This phase intentionally covers deterministic schema and invocation-reference validation. Full live verification is better after Verifier V2 and Composer are implemented.
## Remaining Work
- Phase three: Verifier V2 `claim_checks`.
- Phase four: Composer final-answer generation.
- Phase five: eval fixtures and full audit closure.
## Archive Status
Devflow archive files created for stage two. OpenSpec archive is expected before moving to stage three.
@@ -0,0 +1,44 @@
# Brief: executor-gatekeeper-hook
## Background
Stage one of Executor Structured Output V2 changed Chat Executor output to `executor_evidence_v2`, removing final-expression fields from Executor. That made the output structured, but it did not yet prevent deterministic evidence attribution failures such as fabricated invocation ids, removed fields, empty evidence bindings, or mismatched tool names.
## Goal
Add a deterministic Gatekeeper between Executor output parsing and Verifier model execution.
The Gatekeeper should:
- Validate the initial Executor V2 schema.
- Validate `claims[].evidence_bindings[].source_invocation_ids` against current-session `tool_invocation` rows.
- Validate evidence binding `tool_name` against the persisted invocation tool name.
- Expose a small `gatekeeper_result` to Verifier and audit persistence.
## Scope
Included:
- New Gatekeeper validation service.
- `schema.executor_v2` initial rule.
- `evidence.invocation_ref` initial rule.
- `VerifierInputHook` payload integration.
- `VerifierContextHolder` storage.
- `ChatService` verifier evaluation persistence.
- Minimal verifier prompt update.
- Focused tests for Gatekeeper, hook payload, fabricated invocation ids, tool name mismatch, and persistence.
Excluded:
- No Executor retry on Gatekeeper failure.
- No excerpt similarity rule in this phase.
- No hallucination phrase or evidence utilization rule in this phase.
- No Verifier V2 `claim_checks`.
- No Composer.
- No database schema changes.
## OpenSpec
- Change: `openspec/changes/executor-gatekeeper-hook`
- Parent stage: `openspec/changes/archive/2026-07-07-executor-v2-output-contract`
@@ -0,0 +1,51 @@
# Decisions: executor-gatekeeper-hook
## Key Decisions
### Gatekeeper stays in VerifierInputHook
Decision: Gatekeeper is integrated inside `VerifierInputHook`, after Executor output parsing and before Verifier model execution.
Reason: The user explicitly chose to keep this version in the Verifier hook and not move validation into Executor hook. This preserves the current workflow orchestration.
### No retry in this phase
Decision: Gatekeeper failure does not trigger automatic Executor retry.
Reason: Retry behavior is intentionally deferred. This phase only validates, exposes, and audits deterministic failures.
### Initial rule set is intentionally small
Decision: Stage two implements only `schema.executor_v2` and `evidence.invocation_ref` as hard checks.
Reason: These rules catch the highest-confidence physical failures with low implementation risk. Excerpt similarity, hallucination phrases, and evidence utilization remain later enhancements.
### No new database schema
Decision: Persist `gatekeeper_result` in existing `diagnosis_session.self_evaluation.verifier_evaluation`.
Reason: The user asked to keep database fields minimal. Existing JSON audit storage is enough for this phase.
### Internal interface impact
Decision: This is an L2 internal interface extension.
Impact:
- Verifier payload gains `gatekeeper_result`.
- `VerifierContextHolder` gains Gatekeeper result storage.
- `self_evaluation.verifier_evaluation` gains `gatekeeper_result`.
- No external API, DTO, database table, or schema migration changes.
## Deferred Decisions
- Whether Gatekeeper should later trigger Executor retry.
- Whether `evidence.excerpt_similarity` should be hard fail or warn-only.
- Whether hallucination phrase and evidence utilization rules should be config-driven from metadata files.
- How Verifier V2 `claim_checks` should enforce Gatekeeper failures in code, beyond prompt instruction.
## Remaining Risks
- Verifier prompt compliance is not a deterministic guarantee; stage three should make Gatekeeper fail incompatible with PASS in Verifier V2 behavior.
- Excerpt authenticity is not checked in this phase, so real invocation ids can still be paired with misleading excerpt text until a later rule is implemented.
@@ -0,0 +1,54 @@
# Evidence: executor-gatekeeper-hook
## Context Evidence
- `executor-v2-output-contract` established `executor_evidence_v2` and removed Executor final-expression fields.
- `VerifierInputHook` is the existing integration point for explicit Verifier payload construction.
- `ChatService.persistVerifierEvaluation(...)` is the existing persistence path for verifier audit snapshots.
- `ToolInvocationRepository.findBySessionIdOrderByIdAsc(...)` provides the current-session invocation pool used by Gatekeeper.
## Implementation Evidence
- `src/main/java/com/superbiz/agent/service/ExecutorGatekeeperService.java`
- Implements `schema.executor_v2`.
- Implements `evidence.invocation_ref`.
- Returns `status`, `failed_rules`, `warnings`, and `errors`.
- `src/main/java/com/superbiz/agent/hook/VerifierInputHook.java`
- Runs Gatekeeper after parsing Executor output and building trace summary.
- Adds `gatekeeper_result` to Verifier payload.
- Stores `gatekeeper_result` in `VerifierContextHolder`.
- `src/main/java/com/superbiz/agent/util/VerifierContextHolder.java`
- Stores per-request Gatekeeper result for later persistence.
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- Wires `ExecutorGatekeeperService` into verifier hook construction.
- Persists `gatekeeper_result` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- `src/main/resources/prompts/chat-verifier-prompt.md`
- Documents `gatekeeper_result` as an input.
- States Gatekeeper fail must not produce PASS.
## Test Evidence
- `src/test/java/com/superbiz/agent/service/ExecutorGatekeeperServiceTest.java`
- Covers schema failure and valid pass behavior.
- Covers fabricated invocation ids and tool name mismatch.
- `src/test/java/com/superbiz/agent/hook/VerifierInputHookTest.java`
- Covers Verifier payload containing `gatekeeper_result`.
- Covers hook behavior for fabricated invocation ids.
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`
- Covers persistence of `gatekeeper_result` into verifier evaluation.
## Validation Evidence
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test`
- Result: passed.
- Coverage: 23 focused tests.
- `cmd /c openspec validate executor-gatekeeper-hook`
- Result: passed.
@@ -0,0 +1,41 @@
# Acceptance: executor-v2-output-contract
## Implementation Result
Completed stage one of Executor Structured Output V2.
- Executor prompt now emits `executor_evidence_v2`.
- Executor output no longer includes `diagnosis_summary` or `user_facing_answer`.
- ChatService PASS path renders V2 structured output into readable Chinese.
- VerifierInputHook remains parse-only and accepts V2 output without final-expression fields.
## Static Verification
- `cmd /c openspec validate executor-v2-output-contract`
- Result: passed.
- Coverage: OpenSpec syntax and change validity.
## Script Verification
- `mvn "-Dtest=VerifierInputHookTest,ChatServiceSequentialAgentTest" test`
- Result: passed.
- Coverage: VerifierInputHook V2 parsing; ChatService PASS rendering for V2; existing sequential workflow tests.
## Browser / Manual Verification
Not run. This stage changes backend prompt/runtime contract and unit-level behavior only.
## Unverified
- Full live application run with a real LLM.
- MySQL trace inspection after a real chat session.
Reason: Stage one is covered by focused unit tests; live verification is more useful after Gatekeeper and Composer phases.
## Remaining Work
- Phase two: Gatekeeper in `VerifierInputHook`.
- Phase three: Verifier V2 `claim_checks`.
- Phase four: Composer.
- Phase five: eval fixtures and full audit closure.
@@ -0,0 +1,32 @@
# Brief: executor-v2-output-contract
## Background
`executor_evidence_v1` still made Chat Executor produce both evidence attribution and final user-facing prose through `diagnosis_summary` and `user_facing_answer`.
This kept Executor in a "diagnose and narrate" role and left room for unsupported conclusions to appear before later verification and composition stages.
## Goal
Narrow Chat Executor output to `executor_evidence_v2`: structured diagnostic material only, with final expression removed from Executor.
## Scope
- Update Chat Executor prompt to emit `executor_evidence_v2`.
- Remove `diagnosis_summary` and `user_facing_answer` from Executor output.
- Keep `claims`, `hypotheses`, `recommended_actions`, `missing_info`, and evidence bindings.
- Add temporary ChatService rendering for PASS + V2 output so normal users do not see raw JSON.
- Preserve V1 `user_facing_answer` extraction for compatibility.
## Non-Goals
- No Gatekeeper implementation.
- No Verifier V2 `claim_checks`.
- No Composer.
- No Planner changes.
- No database schema changes.
- No evidence tool signature changes.
## Related OpenSpec
`openspec/changes/archive/2026-07-07-executor-v2-output-contract/`
@@ -0,0 +1,39 @@
# Decisions: executor-v2-output-contract
## Key Decisions
### Executor V2 removes final-expression fields
Decision: Chat Executor final output now uses `executor_evidence_v2` and must not include `diagnosis_summary` or `user_facing_answer`.
Reason: Executor should collect evidence and produce structured diagnostic material, not write final user-facing conclusions.
### Temporary renderer bridges the gap before Composer
Decision: `ChatService` renders V2 structured fields into readable Chinese only when Verifier returns `PASS`.
Reason: Composer is a later phase, but external users must not receive raw JSON during this intermediate stage.
### V1 compatibility remains
Decision: Existing V1 `user_facing_answer` extraction remains.
Reason: It keeps old tests and any lingering V1 output compatible while the staged migration continues.
### Gatekeeper and Verifier V2 are deferred
Decision: This phase does not add Gatekeeper or `claim_checks`.
Reason: The user requested phase-by-phase implementation with archive and commit after each phase. Gatekeeper is phase two.
## Interface Impact
- Internal Agent output contract: L4, because fields are removed.
- Verifier payload: L2, because raw `executor_final_answer` and parsed `executor_structured_output` remain.
- External Chat answer: compatible intent; users still get readable Chinese.
## Risks
- The temporary renderer is not a full Composer and should be replaced in the Composer phase.
- Verifier prompt still uses V1 `facts_checked`; Verifier V2 is a later phase.
@@ -0,0 +1,21 @@
# Evidence: executor-v2-output-contract
## Context Used
- `devflow/projects/2026-07-07-executor-evidence-output-contract`: V1 evidence-attribution contract kept `user_facing_answer`.
- `devflow/projects/2026-07-02-chat-verifier-agent`: Verifier consumes explicit inputs and should not see intermediate reasoning.
- `devflow/projects/2026-07-04-evidence-trace-hardening`: evidence summaries and tool invocation references are the evidence foundation.
- `mvp/issues/design-notes/executor-structured-output-v2.md`: staged implementation design; stage one is Executor V2 output contract.
## Code Evidence
- `src/main/resources/prompts/chat-executor-prompt.md`: V2 contract now uses `answer_version="executor_evidence_v2"` and removes final-expression fields.
- `src/main/java/com/superbiz/agent/service/ChatService.java`: PASS path now tries V1 `user_facing_answer`, then renders V2 structured output to readable Chinese.
- `src/main/resources/prompts/chat-verifier-prompt.md`: `user_facing_answer` is now described as compatibility-only.
- `src/test/java/com/superbiz/agent/hook/VerifierInputHookTest.java`: V2 structured output without `user_facing_answer` parses successfully.
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`: PASS + V2 output renders Chinese and does not expose raw JSON.
## Key Finding
The previous V1 contract intentionally kept `user_facing_answer`, but the V2 staged design intentionally removes it. This is an internal Agent contract break, mitigated by a temporary renderer until Composer is implemented.
@@ -0,0 +1,58 @@
# Acceptance
## Implementation Result
Implemented stage three of Executor Structured Output V2:
- Verifier prompt now validates claim derivability rather than scanning final natural-language output.
- Verifier output supports `claim_checks`.
- `ChatService` derives compatibility `facts_checked` from `claim_checks`.
- `ChatService` persists both `claim_checks` and `facts_checked`.
- Effective verdict guardrails prevent Gatekeeper failures and malformed structured output from remaining `PASS`.
- Verifier logging summarizes `claim_checks`.
## Static Verification
- Reviewed `git diff --stat` and changed files are scoped to stage three implementation, tests, OpenSpec/devflow, and the issue handoff document.
- `cmd /c openspec validate executor-verifier-claim-checks` passed.
- After OpenSpec archive, `cmd /c openspec validate --specs` passed.
## Script Verification
Passed:
```powershell
mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
```
Coverage:
- Verifier payload and Gatekeeper hook behavior.
- Gatekeeper schema/invocation validation.
- `claim_checks` parsing and persistence.
- `claim_checks` to `facts_checked` compatibility mapping.
- all claim verification mapping classes.
- Gatekeeper failure downgrade from model `PASS`.
- malformed Executor output downgrade from model `PASS`.
The targeted Maven test command was re-run after OpenSpec archive and passed.
## Browser / Manual Verification
Not run. This stage changes backend prompt, parser, audit, and tests only.
## OpenSpec Archive Status
Archived:
```text
openspec/changes/archive/2026-07-07-executor-verifier-claim-checks
```
Archive follow-up: the generated canonical `chat-verifier-agent` spec was reviewed and amended to preserve pre-existing verifier input and Gatekeeper schema scenarios while adding the new claim-check scenarios.
## Remaining Risks
- Composer is not implemented in this stage; final PASS rendering still uses the temporary V2 renderer until stage four.
- Full eval fixture expansion is deferred to stage five.
- Existing Maven warnings about duplicate test dependency and Lombok builder defaults remain outside this stage.
@@ -0,0 +1,41 @@
# Executor Verifier Claim Checks
## Background
Stage one moved Chat Executor to `executor_evidence_v2`, and stage two added deterministic Gatekeeper checks before Verifier. After those stages, Verifier still primarily used the legacy `facts_checked` contract and could still treat `executor_final_answer` as a fact source.
That left two risks:
- Verifier could still extract extra confirmed facts from natural-language Executor output.
- Downstream audit and retry consumers could not distinguish V2 claim-level verification from legacy fact checks.
## Goal
Make Verifier V2 claim-oriented:
- verify `executor_structured_output.claims` as the primary target;
- emit `claim_checks` as the authoritative V2 result;
- keep `facts_checked` only as a compatibility projection;
- enforce code-side guardrails so Gatekeeper failures or malformed structured output cannot remain effective `PASS`.
## Scope
- Updated `chat-verifier-prompt.md` to frame verification as claim derivability.
- Extended `ChatService` to parse, normalize, map, and persist `claim_checks`.
- Added effective verdict guardrails for Gatekeeper failure and malformed/missing Executor structured output.
- Updated verifier logging summaries to count `claim_checks`.
- Updated sequential workflow tests to cover claim mapping and downgrade behavior.
## Non-Goals
- No Composer integration in this phase.
- No final-answer material filtering beyond existing templates and temporary V2 renderer.
- No Executor retry behavior change.
- No Gatekeeper rule expansion.
- No database schema migration.
## OpenSpec
- Active change before archive: `openspec/changes/executor-verifier-claim-checks`
- Capability: `chat-verifier-agent`
- Scale: standard
@@ -0,0 +1,57 @@
# Decisions
## Scope Decision
Stage three is limited to Verifier V2 claim checks. Composer is explicitly deferred to stage four.
Reason: Composer requires stable verifier output and allowed-material filtering; mixing it into this stage would make rollback and acceptance unclear.
## Contract Decision
`claim_checks` is the authoritative V2 verifier output.
`facts_checked` remains as a compatibility projection generated from `claim_checks` when present.
Reason: existing low-confidence rendering, retry context, trace output, and evaluation code still depend on `facts_checked`.
## Mapping Decision
Claim verification maps to legacy facts as follows:
| claim verification | legacy facts_checked verification |
|---|---|
| `direct_observation` | `direct_evidence` |
| `reasonable_inference` | `indirect_support` |
| `overstated` | `indirect_support` |
| `unsupported` | `no_evidence` |
| `external_unknown` | `no_evidence` |
| `contradicted` | `contradicted` |
## Guardrail Decision
Effective verdict is enforced in code:
- `gatekeeper_result.status=fail` cannot remain `PASS`.
- `evidence.invocation_ref` failure downgrades to `REJECT`.
- Other Gatekeeper failures downgrade at least to `LOW_CONFID`.
- missing/malformed Executor structured output cannot remain `PASS`.
Reason: prompt compliance is not deterministic enough for safety-critical evidence attribution.
## Apply Fix Record
Initial targeted Maven verification failed because older tests expected PASS to return Executor natural-language output or V1 `user_facing_answer`.
Classification: test drift from the committed OpenSpec, not a design blocker.
Resolution: update tests to use valid Executor V2 output for PASS paths and assert downgrade behavior for malformed or Gatekeeper-failed outputs.
## Interface Impact
L2 internal contract extension:
- Verifier output gains `claim_checks`.
- Existing `facts_checked` remains available.
- Persistence JSON gains `claim_checks` under existing `self_evaluation`.
No external API or database schema changes.
@@ -0,0 +1,35 @@
# Evidence
## Relevant History
- `executor-v2-output-contract`: Executor emits `executor_evidence_v2` and no longer emits final-expression fields.
- `executor-gatekeeper-hook`: Gatekeeper validates schema and invocation references before Verifier and persists `gatekeeper_result`.
- `chat-verifier-agent`: Existing Verifier used `facts_checked`, low-confidence rendering, retry context, and verifier audit.
## Code Evidence
- `src/main/resources/prompts/chat-verifier-prompt.md`: Verifier prompt now makes `executor_structured_output.claims` primary and treats `executor_final_answer` as debug/fallback only.
- `src/main/java/com/superbiz/agent/service/ChatService.java`: parses `claim_checks`, maps them to compatibility `facts_checked`, persists both, and applies effective verdict guardrails.
- `src/main/java/com/superbiz/agent/hook/AgentLoggingHook.java`: verifier thought summaries now include `claim_checks`.
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`: covers claim-check mapping, Gatekeeper downgrade, malformed output downgrade, and V2 PASS paths.
## Evidence-Driven Conclusions
- `facts_checked` cannot be removed yet because existing low-confidence templates, retry context, trace tooling, and eval paths still consume it.
- Prompt-only prevention is insufficient for Gatekeeper failures; `ChatService` must enforce effective verdict downgrades in code.
- No database schema migration is needed because `claim_checks` is persisted inside existing `diagnosis_session.self_evaluation`.
- Composer remains stage four and must not be mixed into this stage.
## Verification Evidence
Script verification passed:
```powershell
mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
cmd /c openspec validate executor-verifier-claim-checks
```
Known existing warnings:
- Maven reports duplicate `spring-boot-starter-test` dependency in `pom.xml`.
- Existing Lombok `@Builder` default warnings remain.
@@ -0,0 +1,48 @@
# diagnosis-eval-demo-gatekeeper-closure Acceptance
## Static / Structure Verification
- `cmd /c openspec validate diagnosis-eval-demo-gatekeeper-closure --strict`
- Result: passed.
- `cmd /c openspec validate --specs`
- Result: passed, 10 specs passed.
## Script Verification
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,VerifierInputHookTest" test`
- Result: 36 tests, 0 failures, 0 errors.
- `mvn "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest,ToolInvocationRecorderTest,QueryLogsToolsTest" test`
- Result: 61 tests, 0 failures, 0 errors.
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`
- Result after E2E startup fix: 23 tests, 0 failures, 0 errors.
## Live E2E Verification
- Start command: `mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"`.
- Demo command: `powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1`.
- Result: chat, trace, and feedback requests completed successfully.
- Output files:
- `mvp/demo/output/chat-response.json`
- `mvp/demo/output/trace-response.json`
- `mvp/demo/output/feedback-response.json`
- Trace observations:
- `hasVerifierEvaluation=true`
- `gatekeeper_result.rule_set_version=gatekeeper-rules-v1`
## Fixed During Verification
- E2E startup initially failed because Spring could not instantiate `ExecutorGatekeeperService`.
- Root cause: two public constructors and no explicit `@Autowired` constructor.
- Fix: annotate the production constructor with `@Autowired`.
## Residual Risk
- The live payment-timeout path can still produce `LOW_CONFID` because model-generated evidence bindings may omit some explicit `source_invocation_id` values.
- This is not a blocker for this change because deterministic matrix behavior is covered by saved fixtures and baseline evaluation.
- Existing Maven warnings remain: duplicate `spring-boot-starter-test` declaration and Lombok `@Builder` default warnings.
## Archive Status
- Devflow archive artifacts created.
- OpenSpec change archived to `openspec/changes/archive/2026-07-08-diagnosis-eval-demo-gatekeeper-closure`.
- Main specs synced by `cmd /c openspec archive diagnosis-eval-demo-gatekeeper-closure --yes`.
@@ -0,0 +1,33 @@
# diagnosis-eval-demo-gatekeeper-closure Brief
## Background
The Chat evidence pipeline already had Executor V2 structured output, deterministic Gatekeeper validation, Verifier claim checks, and Composer final rendering. The missing piece was an interview-ready acceptance story that made the anti-hallucination behavior easy to demonstrate and regress.
## Goal
Close the next three interview-readiness gaps together:
- diagnosis eval fixture matrix
- stable demo data set
- Gatekeeper rule configuration and audit version
## Scope
- Expand `mvp/eval` with matrix-oriented cases, fixtures, and baseline reports.
- Add stable demo request payloads and scenario documentation.
- Add a lightweight local Gatekeeper rule catalog with `rule_set_version` and rule metadata in `gatekeeper_result`.
- Update architecture, demo, and eval docs to describe the current implementation.
## Non-goals
- No new public HTTP endpoint.
- No new database table.
- No Planner `scope_contract`.
- No Gatekeeper retry loop.
- No remote or dynamic rule execution engine.
## OpenSpec
- Change: `openspec/changes/diagnosis-eval-demo-gatekeeper-closure`
- Interface impact: L2 internal contract change.
@@ -0,0 +1,144 @@
# diagnosis-eval-demo-gatekeeper-closure Decisions
## Clarify
- Entry summary: implement the next three interview-readiness items together: diagnosis eval fixture matrix, stable demo data set, and Gatekeeper rule configuration/audit version.
- Slug: `diagnosis-eval-demo-gatekeeper-closure`
- Devflow scale: `standard`
- Interface impact: expected L2 internal contract change because `gatekeeper_result` audit JSON will gain rule metadata/version fields.
## Context
- `devflow/index.md` used: related entries found for diagnosis eval harness, fixture expansion, MVP demo runbook, Gatekeeper hook, and verifier evidence reference fidelity.
- Relevant glossary:
- Evidence Tools produce incident facts and must be recorded in `tool_invocation`.
- Verifier should not use skills/runbooks as incident evidence.
- `tool_invocation.retrieval_details` is the structured evidence/audit home for tool-specific details.
- Historical constraints that must enter OpenSpec:
- Diagnosis eval is offline and deterministic; no LLM-as-judge.
- Demo assets should be runnable, but fixed regression should use saved fixtures.
- Gatekeeper remains in the Verifier hook path.
- No new database table for Gatekeeper audit; use `self_evaluation.verifier_evaluation.gatekeeper_result`.
- `$.no_evidence` is a query no-hit signal, not proof that a problem is impossible.
## Question Pool
| ID | Dimension | Mode | Question | Status |
|---|---|---|---|---|
| Q1 | Terminology | evidence-driven | What names should this change use for the matrix, demo set, and Gatekeeper rule metadata? | Resolved |
| Q2 | Boundary | evidence-driven | Should this change alter public APIs, database schema, Planner output, or retry behavior? | Resolved |
| Q3 | Acceptance | evidence-driven | Which existing tests and baseline assets define the current acceptance style? | Resolved |
| Q4 | Technical | evidence-driven | Where should Gatekeeper rule metadata live with minimal implementation risk? | Pending code research |
| Q5 | Scope | user-interview | Should the stable demo set be documentation/payloads only, or should it include live E2E scripts for all scenarios? | Confirmed |
## Evidence-driven Conclusions
- Q1 conclusion: use `diagnosis eval matrix`, `stable demo scenarios`, and `Gatekeeper rule set version` as terms.
- Q2 conclusion: keep this as an internal contract change. Do not add public endpoints, tables, Planner `scope_contract`, or Gatekeeper retry.
- Q3 conclusion: existing `DiagnosisTraceEvaluatorTest`, `ExecutorGatekeeperServiceTest`, `VerifierInputHookTest`, `ToolInvocationRecorderTest`, and `mvp/eval/reports` define the current acceptance style.
- Q4 conclusion: Gatekeeper metadata should live behind a small rule catalog loaded by `ExecutorGatekeeperService`; the audit output should include a rule set version and enabled rule metadata summary, without adding tables or remote registry.
## User-interview Confirmations
- Q5 confirmed by resumed objective: complete items 1/2/3 with sm-flow, archive, submit, and run end-to-end if necessary.
- Implementation interpretation: stable demo scenarios will be fixed request payloads and runbook docs plus deterministic fixture-backed eval. Live E2E remains necessary only for at least one main path or where unit/fixture evidence is insufficient.
## OpenSpec Backfill
- Created Draft proposal at `openspec/changes/diagnosis-eval-demo-gatekeeper-closure/proposal.md`.
- Context constraints from historical devflow entries were written into the proposal.
- Scope confirmation and Gatekeeper catalog placement were written into the proposal/design.
## Current Checkpoint
- Discover completed.
- No implementation files changed yet.
## Specify / Alignment
### Cross-artifact Alignment
| Check | Status | Notes |
|---|---|---|
| brief/proposal goals -> proposal | Aligned | Proposal covers eval matrix, stable demo scenarios, and Gatekeeper rule catalog/audit version. |
| proposal scope/constraints -> design | Aligned | Design records offline deterministic eval, fixture-backed demo distinction, local rule catalog, and no new table/API. |
| design decisions -> specs/tasks | Aligned | Specs cover eval matrix, rule set version validation, demo scenarios, and Gatekeeper rule metadata; tasks cover matching implementation slices. |
| specs observable behavior -> tasks | Aligned | Each requirement has an executable task and acceptance check. |
### Interface Impact
- Level: L2 internal contract change.
- Reason: `gatekeeper_result` internal audit JSON gains `rule_set_version` and rule metadata summary. Eval case/result fields may gain optional rule set checks. No public HTTP API, database schema, or external DTO contract changes.
## Audit
Input -> processing -> output chain:
```text
mvp/demo request docs + mvp/eval fixtures
-> DiagnosisTraceEvaluator
-> baseline reports
-> interview/demo evidence
Gatekeeper rule catalog
-> ExecutorGatekeeperService
-> VerifierInputHook / ChatService persisted self_evaluation
-> Trace and eval audit
```
Architecture risk assessment:
1. The change is intentionally internal and should not add new public consumers.
2. Gatekeeper catalog must stay metadata-only; dynamic rule execution would be a different, riskier architecture.
3. Fixture-backed demo scenarios should be documented as deterministic regression artifacts, not live LLM guarantees.
4. Baseline report churn is expected and must be committed with case/fixture changes.
5. No devflow/OpenSpec conflict found.
## Commit Gate
- `cmd /c openspec validate diagnosis-eval-demo-gatekeeper-closure --strict`: passed.
- `cmd /c openspec validate --specs`: passed, 10 specs passed.
- File completeness:
- proposal.md: present.
- design.md: present.
- specs: present for `diagnosis-eval-harness`, `mvp-demo-trace-acceptance`, `chat-verifier-agent`.
- tasks.md: present.
- Consistency:
- Proposal concepts have corresponding design sections.
- Design decisions are reflected in specs/tasks.
- Task acceptance checks are verifiable.
## Current Checkpoint
- Commit completed.
- `.committed` marker created.
## Apply Verification
- Focused verification passed:
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,VerifierInputHookTest" test`
- Result: 36 tests, 0 failures, 0 errors.
- Broader relevant regression passed:
- `mvn "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest,ToolInvocationRecorderTest,QueryLogsToolsTest" test`
- Result: 61 tests, 0 failures, 0 errors.
- E2E startup repro found a Spring bean construction issue:
- Command: `mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"`
- Failure: `ExecutorGatekeeperService` had two public constructors and no annotated constructor, so Spring attempted a no-arg constructor and failed with `No default constructor found`.
- Classification: code deviation from OpenSpec implementation intent, not a spec gap.
- Fix: annotate the production constructor with `@Autowired`.
- Post-fix focused regression passed:
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`
- Result: 23 tests, 0 failures, 0 errors.
- Live E2E passed for demo compatibility:
- Start: `mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"`
- Run: `powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1`
- Result: `/api/chat`, `/api/diagnosis/{sessionId}/trace`, and `/api/feedback` completed successfully.
- Trace summary included `hasVerifierEvaluation=true`.
- Persisted Gatekeeper audit included `rule_set_version=gatekeeper-rules-v1`.
- Residual quality note: the live payment-timeout response remained `LOW_CONFID` because some model-produced evidence bindings still lacked explicit `source_invocation_id`; deterministic PASS/LOW_CONFID/REJECT claims are covered by fixture-backed eval.
## Archive Readiness
- OpenSpec tasks 1-4 completed.
- Verification is recorded in devflow acceptance artifacts.
- Remaining known risk: live LLM output is not deterministic and may still produce LOW_CONFID on the payment-timeout path; this is intentionally documented as demo compatibility, not a fixed PASS guarantee.
@@ -0,0 +1,22 @@
# diagnosis-eval-demo-gatekeeper-closure Evidence
## Code And Artifact Evidence
- Gatekeeper rule metadata lives in `src/main/resources/gatekeeper/gatekeeper-rules.json`.
- `ExecutorGatekeeperService` loads the local catalog, uses configured threshold parameters, and emits `rule_set_version` plus enabled rule metadata.
- `VerifierInputHook` and `ChatService` preserve Gatekeeper audit metadata in fallback/default paths.
- `DiagnosisTraceEvaluator` can optionally validate expected Gatekeeper rule set version.
- `mvp/eval/cases/diagnosis-cases.json` now includes narrow-scope and no-evidence matrix cases.
- `mvp/eval/reports/baseline-report.json` and `.md` were regenerated for the expanded fixed matrix.
- `mvp/demo/evidence-pipeline-scenarios.md` documents live vs fixture-backed demo scenarios.
## Decisions
- Keep this phase internal: no public API, no DB schema, no Planner output change.
- Keep Gatekeeper deterministic Java validation; the catalog is metadata/config only.
- Treat live demo as compatibility evidence and fixture-backed eval as deterministic regression evidence.
- Persist audit under the existing `self_evaluation.verifier_evaluation.gatekeeper_result` structure.
## Runtime Finding
The first Maven E2E startup found a real integration issue: `ExecutorGatekeeperService` had multiple public constructors without an annotated constructor, so Spring could not instantiate the service. The fix was to annotate the production constructor with `@Autowired`.
@@ -0,0 +1,57 @@
# Acceptance
## Implementation Result
Implemented stage four of Executor Structured Output V2:
- Added `chat_composer` prompt and Agent.
- Final answers for PASS, LOW_CONFID, and REJECT now use Composer when Verifier decision is valid.
- Composer input is filtered from Verifier decision and Executor structured output.
- Unsupported, external-unknown, and contradicted claims are excluded from confirmed final-answer material.
- REJECT Composer input has `allowed_hypotheses=[]`.
- Malformed Composer output uses deterministic safe fallback.
- Fallback does not expose raw Composer JSON, raw Executor JSON, or Executor `user_facing_answer`.
- `composer_output` is persisted in verifier audit.
## Static Verification
- Reviewed implementation diff for stage-four scope.
- `cmd /c openspec validate executor-composer-final-answer` passed.
- `cmd /c openspec validate --specs` passed before archive.
## Script Verification
Passed:
```powershell
mvn "-Dtest=ChatServiceSequentialAgentTest" test
mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
```
Coverage:
- Composer prompt loading and invocation.
- valid Composer output as final answer source.
- malformed Composer output fallback.
- PASS no raw Executor JSON leakage.
- LOW_CONFID separation of confirmed material, possible directions, and gaps.
- REJECT safe output without raw Executor answer.
- Gatekeeper and Verifier stage compatibility.
## Browser / Manual Verification
Not run. This stage changes backend prompt, routing, parser, audit, and tests only.
## OpenSpec Archive Status
Archived:
```text
openspec/changes/archive/2026-07-08-executor-composer-final-answer
```
## Remaining Risks
- Stage five still needs broader eval fixture coverage for full evidence-attribution regressions.
- Composer prompt quality can be improved after real run traces are collected.
- Existing Maven warnings about duplicate test dependency and Lombok builder defaults remain outside this stage.
@@ -0,0 +1,54 @@
# Executor Composer Final Answer
## Background
Stages one to three moved the Chat diagnosis chain to structured Executor output, deterministic Gatekeeper validation, and Verifier `claim_checks`.
Before this stage, `ChatService` still owned final answer rendering. PASS paths could use a temporary V2 renderer, while LOW_CONFID and REJECT paths used templates. That left final user-facing expression too close to Executor material and made it harder to prove that only Verifier-allowed claims reached the user.
## Goal
Add a Composer expression layer after Verifier:
```text
chat_planner
-> chat_executor
-> VerifierInputHook + Gatekeeper
-> chat_verifier
-> chat_composer
-> final answer
```
Composer produces user-facing answers from filtered material only:
- `allowed_claims`
- `allowed_hypotheses`
- `missing_info`
- `recommended_actions`
- `rationale`
## Scope
- Added `chat-composer-prompt.md`.
- Added `chat_composer` Agent construction in `ChatService`.
- Added Composer input filtering from Verifier decision and Executor structured output.
- Replaced PASS temporary V2 renderer usage with Composer-or-safe-fallback rendering.
- Routed LOW_CONFID and REJECT final answers through Composer when Verifier output is valid.
- Added deterministic fallback for malformed Composer output.
- Persisted `composer_output` under `diagnosis_session.self_evaluation.verifier_evaluation`.
- Updated sequential workflow tests.
## Non-Goals
- No Planner changes.
- No Executor retry changes.
- No Gatekeeper rule expansion.
- No Verifier classification expansion.
- No database schema migration.
- No stage-five eval fixture expansion.
## OpenSpec
- Active change before archive: `openspec/changes/executor-composer-final-answer`
- Capabilities: `chat-composer-agent`, `chat-verifier-agent`
- Scale: standard
@@ -0,0 +1,70 @@
# Decisions
## Scope Decision
Stage four is limited to Composer final-answer generation and routing.
Reason: Gatekeeper and Verifier contracts were stabilized in earlier stages; this phase should only close the final-expression path.
## Composer Responsibility
Composer is an expression layer, not a diagnosis layer.
It may rephrase and organize only filtered material. It must not call tools, introduce new facts, rejudge root cause, or read raw Executor/tool output.
## Filtering Decision
`ChatService` owns Composer input filtering:
| Verifier classification | Composer handling |
|---|---|
| `direct_observation` | `allowed_claims` |
| `reasonable_inference` | `allowed_claims`, with bounded wording |
| `overstated` | `allowed_hypotheses` or `missing_info` |
| `unsupported` | `missing_info` |
| `external_unknown` | `missing_info` |
| `contradicted` | `missing_info` / REJECT-safe output |
For REJECT, `allowed_hypotheses` is always empty.
## Fallback Decision
Malformed Composer output falls back to deterministic rendering from filtered Composer input.
Fallback must never expose:
- raw Composer JSON;
- raw Executor JSON;
- Executor `user_facing_answer`;
- full unscreened tool output.
## Audit Decision
No new table is added. Composer output is persisted under:
```text
diagnosis_session.self_evaluation.verifier_evaluation.composer_output
```
The audit snapshot is intentionally compact and stores status plus parsed user-facing fields.
## Apply Fix Record
Initial targeted verification exposed test drift:
- test file had a UTF-8 BOM and failed Java compilation;
- scripted chat model did not recognize `COMPOSER_TEST_PROMPT`;
- older tests expected three-Agent execution and temporary V2 renderer behavior;
- LOW_CONFID assertions required indirect support to disappear instead of appearing as a possible direction.
Resolution: remove BOM, add Composer script branch, and update assertions to match the committed Composer contract.
## Interface Impact
L2 internal behavior change:
- external Chat API still returns a final answer string;
- internal final-answer source changes from Executor/temporary renderer to Composer or safe fallback;
- audit JSON gains `composer_output` under existing `self_evaluation`.
No database schema change.
@@ -0,0 +1,36 @@
# Evidence
## Relevant History
- `executor-v2-output-contract`: Executor emits structured diagnostic material and no final-expression fields.
- `executor-gatekeeper-hook`: Gatekeeper validates deterministic evidence failures before Verifier.
- `executor-verifier-claim-checks`: Verifier emits `claim_checks` and effective verdict guardrails.
## Code Evidence
- `src/main/resources/prompts/chat-composer-prompt.md`: defines Composer as an expression layer with strict JSON output.
- `src/main/java/com/superbiz/agent/service/ChatService.java`: loads Composer prompt, invokes `chat_composer`, filters Composer input, parses Composer output, falls back safely, and persists Composer audit.
- `src/test/java/com/superbiz/agent/service/ChatServiceSequentialAgentTest.java`: covers Composer invocation, fallback, REJECT/LOW_CONFID behavior, and no raw JSON leakage.
## Evidence-Driven Conclusions
- Composer must be after Verifier because Verifier `claim_checks` are the authority for allowed final-answer material.
- Composer must not receive raw tool output or full unscreened Executor output because that would re-open the evidence attribution problem.
- Verifier malformed/missing output should not invoke Composer because there is no trustworthy decision to filter with.
- Fixed fallback remains necessary because Composer is an LLM call with a strict JSON contract and can produce malformed output.
## Verification Evidence
Passed:
```powershell
mvn "-Dtest=ChatServiceSequentialAgentTest" test
mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
cmd /c openspec validate executor-composer-final-answer
cmd /c openspec validate --specs
```
Known existing warnings:
- Maven reports duplicate `spring-boot-starter-test` dependency in `pom.xml`.
- Existing Lombok `@Builder` default warnings remain.
@@ -0,0 +1,52 @@
# Acceptance
## Static Verification
- `openspec validate verifier-evidence-reference-fidelity --strict`: passed.
- `openspec validate --specs`: passed.
## Script Verification
- `mvn "-Dtest=ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,QueryLogsToolsTest,ChatServiceSequentialAgentTest" test`
- Passed: 38 tests.
- `mvn "-Dtest=ExecutorGatekeeperServiceTest" test`
- Passed: 9 tests.
- `mvn test`
- Failed on unrelated environment-gated `MilvusConnectionTest.connect`: `MILVUS_TOKEN` environment variable was not set.
- Other executed tests in the run progressed until that single failure; focused tests for this change passed.
## End-to-End Verification
The Java service was restarted with `mvn spring-boot:run`; logs were written under `logs/`.
| Case | Session | Result | Gatekeeper Audit |
|---|---|---|---|
| HikariCP positive | `iss007-hikari-positive-20260708-1553` | PASS; confirmed order-service HikariCP timeout and pool saturation logs | `pass / none`, checked_bindings=2 |
| HikariCP negative | `iss007-hikari-negative-20260708-1555` | LOW_CONFID; no `generic-service`; no false positive for inventory-service | `fail / low_confid` |
| HighMemoryUsage positive | `iss007-memory-positive-20260708-1558` | PASS; confirmed HighMemoryUsage 91%, did not confirm memory leak | `pass / none`, checked_bindings=1 |
| SlowResponse positive | `iss007-slow-positive-20260708-1600` | PASS; confirmed SlowResponse and slow request logs, no DB pool root cause | `pass / none`, checked_bindings=7 |
| Narrow HighCPUUsage | `iss007-narrow-highcpu-20260708-1602` | PASS; only covered payment-service HighCPUUsage | `pass / none`, checked_bindings=1 |
## Database Audit
`scripts/query_mysql.py` was used to verify:
- `diagnosis_session.self_evaluation.verifier_evaluation.verdict`
- `gatekeeper_result.status`
- `gatekeeper_result.severity`
- `gatekeeper_result.checked_bindings`
- no-hit HikariCP query rows persist `evidence_status=no_evidence`
## Remaining Risk
- Negative no-hit claims still have incomplete precise references when Executor uses `$.logs` for empty arrays. Gatekeeper correctly downgrades to `LOW_CONFID`.
- Prompt-only scope control is improved but not a hard contract. A future `scope_contract` may still be needed.
- Full test suite requires `MILVUS_TOKEN` to pass `MilvusConnectionTest`.
## OpenSpec Archive
- `openspec archive verifier-evidence-reference-fidelity --yes`: succeeded.
- Main specs updated:
- `openspec/specs/chat-verifier-agent/spec.md`
- `openspec/specs/evidence-trace-hardening/spec.md`
- Non-blocking warning: proposal did not use OpenSpec's preferred `## Why` / `## What Changes` headers, but archive completed.
@@ -0,0 +1,44 @@
# Verifier Evidence Reference Fidelity
## Background
ISS-007 came from end-to-end diagnosis cases where raw tool output and Executor `evidence_excerpt` contained enough facts, but Verifier still returned `LOW_CONFID` because the verifier-facing summary compressed away key details.
The affected flow is:
```text
chat_planner
-> chat_executor
-> VerifierInputHook / Gatekeeper
-> chat_verifier
-> chat_composer
```
The change hardens the evidence handoff between Executor, Gatekeeper, and Verifier.
## Goal
Make Executor cite concrete tool evidence, make Gatekeeper validate that citation with code, and make Verifier judge whether verified evidence can derive the claim.
## Scope
- Persist `tool_invocation.retrieval_details.evidence_refs`.
- Use `source_invocation_id + raw_path + evidence_excerpt` as the precise evidence binding.
- Add Gatekeeper `severity` and checked binding audit.
- Keep Gatekeeper in the Verifier hook path.
- Keep `tool_trace_summary` as navigation/audit context, not the only evidence source.
- Fix HikariCP mock positive/no-hit behavior.
- Tighten Executor/Verifier prompts for narrow-scope evidence handling.
## Non-goals
- No Planner `scope_contract`.
- No new database table.
- No full JSONPath engine.
- No change to external HTTP API.
- No retry rollback from Gatekeeper to Executor in this phase.
## OpenSpec
- Change: `openspec/changes/verifier-evidence-reference-fidelity`
- Source issue: `mvp/issues/archived/ISS-007-verifier-evidence-summary-fidelity.md`
@@ -0,0 +1,54 @@
# Decisions
## Evidence Reference
Use `source_invocation_id + raw_path + evidence_excerpt` as the precise evidence reference for Executor claim bindings.
Reason:
- Invocation ID alone only identifies a tool call, not the evidence inside it.
- `raw_path` is enough for the first version when paired with `retrieval_details.evidence_refs`.
- `evidence_excerpt` remains the text Verifier reads, but only after Gatekeeper validates it.
## Raw Path
Only support stable locators in the first version:
- `$.alerts[i]`
- `$.logs[i]`
- `$.evidence_blocks[i]`
No full JSONPath engine is introduced.
## Gatekeeper Severity
Gatekeeper output includes:
- `status`
- `severity`
- `checked_bindings`
- `failed_rules`
- `warnings`
- `errors`
Severity meaning:
- `none`: precise references passed.
- `low_confid`: evidence is missing or incomplete, but not fabricated.
- `reject`: fabricated ID, wrong tool, unknown raw path, or mismatched excerpt.
## Verifier Boundary
Verifier uses verified claim-local excerpts as primary derivability evidence. `tool_trace_summary` remains available for navigation and audit, but no longer needs to carry every concrete fact.
## Hook Placement
Gatekeeper remains in the Verifier input hook path. This version does not retry Executor on Gatekeeper failure.
## Planner
Planner is not changed. `scope_contract` remains a later-stage idea. This phase uses prompt constraints to reduce narrow-scope over-expansion.
## Database
No new tables. Evidence refs are stored in `tool_invocation.retrieval_details.evidence_refs`; audit is stored in `diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result`.
@@ -0,0 +1,33 @@
# Evidence
## Existing Context
- Existing `chat-verifier-agent` spec still used `source_invocation_ids` and `tool_trace_summary` as the main verifier evidence context.
- Existing `evidence-trace-hardening` spec already established `tool_invocation.retrieval_details` as the right place for structured tool-specific facts.
- Prior devflow projects established that runbook/skill content is guidance, not incident evidence.
## Code Findings
- `VerifierInputHook` previously backfilled plural `source_invocation_ids` from `tool_trace_summary` by tool name.
- `ExecutorGatekeeperService` previously validated invocation existence and tool name, but not `raw_path` or excerpt authenticity.
- `ToolInvocationRecorder` persisted retrieval details but did not generate claim-addressable `evidence_refs`.
- `QueryLogsTools` could fall back to `generic-service` placeholder logs on no-hit.
## Implementation Evidence
- `ToolInvocationRecorder` now extracts:
- `$.alerts[i]` for `query_metrics`
- `$.logs[i]` for `query_logs`
- `$.evidence_blocks[i]` for `lookup_knowledge`
- `ExecutorGatekeeperService` now validates:
- invocation existence
- tool name
- raw path presence
- `retrieval_details.evidence_refs`
- excerpt similarity/support
- `VerifierInputHook` only auto-fills a singular `source_invocation_id` when exactly one candidate exists and never invents `raw_path`.
- `QueryLogsTools` returns HikariCP mock logs for `order-service` and returns empty no-hit results for unrelated services.
## Residual Finding
The HikariCP negative E2E no longer has generic-service pollution, but the model still issued an extra broad HikariCP query without the service filter and used order-service as context. This is a remaining narrow-scope behavior issue, not a mock evidence pollution issue.
@@ -0,0 +1,46 @@
# Acceptance
## Static Verification
- `openspec validate interview-demo-quality-audit --strict`
- Result: passed.
- Coverage: OpenSpec proposal/design/spec/tasks consistency.
- PowerShell parser/runtime readiness check:
- Command: `powershell -NoProfile -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1 -BaseUrl http://127.0.0.1:1 -OutputDir target/demo-check-syntax`
- Result: expected failure with actionable readiness message.
- Coverage: script parses under Windows PowerShell and fails before issuing diagnosis requests when service is unreachable.
## Script Verification
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
- Result: passed.
- Coverage: 12/12 fixed eval fixtures, Prompt audit evaluator checks, Gatekeeper rule metadata checks, regenerated baseline reports.
- `mvn -q "-Dtest=ChatServiceSequentialAgentTest" test`
- Result: passed.
- Coverage: Chat verifier evaluation persists `prompt_audit`.
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ChatServiceSequentialAgentTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`
- Result: passed.
- Coverage: broader eval, baseline diff, Chat sequential flow, Gatekeeper, and Verifier input hook regression set.
- `mvn -q -DskipTests compile`
- Result: passed.
- Coverage: main source compilation.
## E2E Verification
- Started service with:
- `mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo`
- Ran:
- `powershell -NoProfile -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1 -BaseUrl http://localhost:9900 -SessionId mvp-demo-interview-quality-audit-001`
- Result: passed.
- Summary:
- `chatSuccess=true`
- `verdict=LOW_CONFID`
- `gatekeeperStatus=fail`
- `gatekeeperRuleSetVersion=gatekeeper-rules-v1`
- `promptAuditVersion=chat-prompts-v1`
- tools included `lookup_knowledge`, `query_logs`, `query_metrics`, and `get_available_log_topics`
- Note: live E2E remains a compatibility check, not the deterministic PASS oracle. The fixed fixture baseline is the regression source of truth.
## Not Verified
- Browser UI inspection was not required for this change because the scope is backend trace/eval/demo script documentation, not frontend behavior.
@@ -0,0 +1,23 @@
# Interview Demo Quality Audit Brief
## Background
The MVP already demonstrates traceable Agent diagnosis with Planner, Executor, Gatekeeper, Verifier, Composer, evidence tools, trace persistence, and deterministic eval fixtures. The remaining interview-readiness gap is not a new Agent architecture; it is making the demo easier to run and making prompt/rule changes easier to audit.
## Goal
Stabilize the interview demo path, expand fixture-backed evaluation, and persist prompt/Gatekeeper audit metadata so the project can explain and verify Agent behavior during interviews.
## Scope
- Add prompt audit metadata to Chat verifier evaluation.
- Extend deterministic eval cases and baseline reports.
- Add an interview demo preflight/check script.
- Update MVP demo and architecture documentation.
## Non-goals
- No public API or database schema changes.
- No new SubAgent split, MCP migration, process isolation, or AIOps LLM Verifier.
- No guarantee that every live LLM run returns PASS.
@@ -0,0 +1,111 @@
# interview-demo-quality-audit Decisions
## Clarify
- Entry summary: stabilize the interview demo, expand deterministic eval coverage, and add Prompt/Gatekeeper version audit.
- Slug: `interview-demo-quality-audit`.
- Devflow scale: `standard`.
- Interface impact: L2 internal contract change because `verifier_evaluation` gains `prompt_audit`; no public HTTP API or database schema change.
## Context
- `devflow/index.md` used: related entries found for `diagnosis-eval-demo-gatekeeper-closure`, `executor-composer-final-answer`, `verifier-evidence-reference-fidelity`, `mvp-demo-interview-runbook`, and `diagnosis-eval-baseline-diff`.
- Relevant glossary:
- Evidence Tools produce incident facts and must be recorded in `tool_invocation`.
- Verifier should not use skills/runbooks as incident evidence.
- Diagnosis Playbook Skill is workflow guidance, not a fact source.
- Historical constraints that must enter OpenSpec:
- Diagnosis eval is deterministic and fixture-backed; no LLM-as-judge.
- Stable demo scenarios are documentation/payloads plus deterministic fixtures; live E2E is a compatibility check, not a guaranteed PASS oracle.
- Gatekeeper rule metadata is already metadata-only and should not become dynamic rule execution.
- Composer is the final expression layer and must not leak raw Executor JSON.
## Question Pool
| ID | Dimension | Mode | Question | Status |
|---|---|---|---|---|
| Q1 | Terminology | evidence-driven | What should the new audit metadata be called? | Resolved |
| Q2 | Boundary | evidence-driven | Does this require public API or schema changes? | Resolved |
| Q3 | Acceptance | evidence-driven | Which current assets define deterministic acceptance? | Resolved |
| Q4 | Technical | evidence-driven | Where should prompt version metadata live with minimal implementation risk? | Resolved |
| Q5 | Scope | user-interview | Should live E2E be mandatory for all scenarios? | Confirmed by objective as conditional |
## Evidence-driven Conclusions
- Q1 conclusion: use `prompt_audit` for prompt version metadata and keep existing `gatekeeper_result.rule_set_version`.
- Q2 conclusion: keep this as an internal trace/self-evaluation contract change. Do not add endpoints, tables, or new Agent roles.
- Q3 conclusion: `DiagnosisTraceEvaluatorTest`, baseline reports, fixed fixtures, and demo scripts define current acceptance style.
- Q4 conclusion: add a small Chat prompt audit catalog near `ChatService` prompt loading and persist a compact snapshot with verifier evaluation.
- Q5 conclusion: run live E2E with `mvp-demo` profile if dependencies are available; otherwise record the blocker and rely on deterministic eval/unit evidence.
## Specify / Alignment
| Check | Status | Notes |
|---|---|---|
| proposal goals -> proposal | Aligned | Proposal covers demo preflight, eval expansion, prompt audit, and docs. |
| proposal scope/constraints -> design | Aligned | Design records no public API/schema changes, prompt audit shape, eval fields, and demo script behavior. |
| design decisions -> specs/tasks | Aligned | Specs cover persisted prompt audit, evaluator checks, baseline, and demo script outputs. |
| specs observable behavior -> tasks | Aligned | Each requirement has implementation and verification tasks. |
## Audit
Input -> processing -> output chain:
```text
prompt resource metadata
-> ChatService / PromptAudit snapshot
-> verifier_evaluation.prompt_audit
-> Trace API / eval fixtures
-> DiagnosisTraceEvaluator baseline
run-interview-demo-check.ps1
-> service readiness
-> chat / trace / feedback
-> mvp/demo/output summary
```
Architecture risk assessment:
1. The audit shape is intentionally compact and internal; storing full prompt text would create noisy traces and possible sensitive-content risk.
2. Eval should assert versions by explicit metadata, not by prompt content hashes that churn during local prompt edits.
3. Live demo checks may still be LOW_CONFID because LLM output is not deterministic; deterministic fixtures remain the regression source of truth.
4. No devflow/OpenSpec conflict found.
## Commit Gate
- `openspec validate interview-demo-quality-audit --strict`: passed.
- File completeness:
- `proposal.md`: present.
- `design.md`: present.
- `specs/`: present for `chat-verifier-agent`, `diagnosis-eval-harness`, and `mvp-demo-trace-acceptance`.
- `tasks.md`: present.
- Consistency:
- Proposal goals map to design sections.
- Design decisions map to spec requirements and executable tasks.
- Task acceptance checks are verifiable.
- `.committed` marker created.
## Current Checkpoint
- Commit completed.
- Apply is authorized by the original objective: "完成后归档提交".
## Pre-apply Research
- Capability source: sm-flow built-in apply protocol. `openspec-apply-change` was not invoked directly in this session.
- Repository semantic search/LSP note: the requested `codebase-retrieval` and LSP tools were not available in the exposed toolset, so impact analysis used `rg`, direct file reads, OpenSpec/devflow artifacts, and targeted tests.
- Reference implementation and reuse:
- `ChatService.persistVerifierEvaluation(...)` is the single persistence point for Chat verifier/composer audit data; prompt audit was added there to cover normal, fallback, and degraded Composer paths.
- `DiagnosisTraceEvaluator` and `DiagnosisEvalReportWriter` are the deterministic eval extension points; no LLM judge was introduced.
- `mvp/demo/scripts/run-payment-timeout-demo.ps1` provided the request/trace/feedback flow reused by the new interview preflight script.
- Interface impact remains L2 internal trace contract: `verifier_evaluation.prompt_audit` and eval report fields are added; no public endpoint, table, or request DTO changed.
## Apply Notes
- Added compact Chat prompt audit metadata: `chat-prompts-v1`, with planner/executor/verifier/composer prompt versions and resource paths.
- Extended diagnosis eval schema, result reporting, baseline fixtures, JSON report, and Markdown report for Prompt audit and Gatekeeper rule metadata.
- Added two fixture-backed audit cases:
- `prompt-gatekeeper-audit-closure`
- `audit-metadata-low-confid`
- Added `mvp/demo/scripts/run-interview-demo-check.ps1` to run service readiness, Chat, Trace, feedback, and summary output.
- Updated MVP demo/eval/architecture docs to explain `prompt_audit.version`, `gatekeeper_result.rule_set_version`, and deterministic fixture baseline.
@@ -0,0 +1,58 @@
# Evidence
## Context Files Read
- `devflow/index.md`
- `devflow/glossary/CONTEXT.md`
- `devflow/projects/2026-07-08-diagnosis-eval-demo-gatekeeper-closure/decisions.md`
- `devflow/projects/2026-07-08-executor-composer-final-answer/decisions.md`
- `mvp/architecture/current-mvp-architecture.md`
- `mvp/architecture/agent-orchestration.md`
- `mvp/architecture/executor-evidence-pipeline-refactor.md`
- `mvp/architecture/harness-quality-gates.md`
- `mvp/demo/README.md`
- `mvp/demo/ten-minute-interview-demo.md`
- `mvp/eval/README.md`
- `mvp/eval/cases/diagnosis-cases.json`
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- `src/main/java/com/superbiz/agent/service/ExecutorGatekeeperService.java`
- `src/main/java/com/superbiz/agent/eval/DiagnosisTraceEvaluator.java`
- `src/main/resources/gatekeeper/gatekeeper-rules.json`
## Tooling Note
The required `codebase-retrieval` and LSP tools were not exposed in this session. Impact analysis used `rg`, direct file reads, existing OpenSpec/devflow artifacts, and targeted tests instead.
## Implementation Evidence
- `src/main/java/com/superbiz/agent/service/ChatService.java`
- Adds `prompt_audit` under `verifier_evaluation` through the shared `persistVerifierEvaluation(...)` path.
- Uses compact metadata only: audit version, prompt names, prompt versions, and resource paths.
- `src/main/java/com/superbiz/agent/eval/DiagnosisTraceEvaluator.java`
- Adds deterministic checks for `requirePromptAudit`, `expectedPromptAuditVersion`, `expectedPromptVersions`, and `requireGatekeeperRules`.
- `src/main/java/com/superbiz/agent/eval/DiagnosisEvalReportWriter.java`
- Adds Prompt Audit and Gatekeeper rule count columns to Markdown reports.
- `mvp/eval/cases/diagnosis-cases.json`
- Expands fixed baseline to 12 fixture-backed cases.
- `mvp/eval/fixtures/prompt-gatekeeper-audit-closure-pass.json`
- Positive PASS fixture proving Prompt audit and Gatekeeper rule metadata closure.
- `mvp/eval/fixtures/audit-metadata-low-confid.json`
- LOW_CONFID fixture proving safe answer behavior while audit metadata remains present.
- `mvp/demo/scripts/run-interview-demo-check.ps1`
- Adds service readiness, Chat, Trace, feedback, and summary output for interview preflight.
## Verification Evidence
- OpenSpec:
- `openspec validate interview-demo-quality-audit --strict`: passed before archive.
- `openspec validate --specs --strict`: 10 specs passed after merging deltas into main specs.
- Unit/eval:
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`: passed.
- `mvn -q "-Dtest=ChatServiceSequentialAgentTest" test`: passed.
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ChatServiceSequentialAgentTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`: passed.
- Compile:
- `mvn -q -DskipTests compile`: passed.
- E2E:
- Started `mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo`.
- Ran `mvp/demo/scripts/run-interview-demo-check.ps1` against `http://localhost:9900`.
- Summary recorded `chatSuccess=true`, `verdict=LOW_CONFID`, `gatekeeperRuleSetVersion=gatekeeper-rules-v1`, and `promptAuditVersion=chat-prompts-v1`.
@@ -0,0 +1,103 @@
# Acceptance
## 实现结果
- OpenSpec tasks: `42/42` complete。
- Phase commits:
- `52bf030 feat(trace): add session run isolation schema`
- `6fdbd34 docs(openspec): tighten run isolation contract`
- `26d5529 feat(trace): isolate chat runs`
- `027aed1 feat(trace): add run-scoped trace reads`
- `d928a19 feat(trace): bind feedback to runs`
- `78c1477 feat(trace): isolate aiops runs`
- `f9df943 feat(trace): finish run-aware demo verification`
- OpenSpec archive: `openspec/changes/archive/2026-07-10-session-run-trace-isolation`
## 静态验证
```powershell
node --check src\main\resources\static\app.js
node --check src\main\resources\static\trace.js
openspec validate session-run-trace-isolation --strict
git diff --check -- . ':!devflow/index.md'
```
结果:通过。
## 脚本验证
PowerShell demo 脚本解析:
```powershell
$scripts = @(
'mvp\demo\scripts\run-payment-timeout-demo.ps1',
'mvp\demo\scripts\run-interview-demo-check.ps1'
)
foreach ($script in $scripts) {
[scriptblock]::Create((Get-Content -Raw -Encoding UTF8 $script)) | Out-Null
}
```
结果:通过。
Focused tests:
```powershell
mvn -q "-Dtest=ChatControllerTest,DiagnosisTraceServiceTest,FeedbackControllerTest,FeedbackServiceTest,AiOpsServiceTest" test
```
结果:通过。
Baseline / regression:
```powershell
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest" test
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
```
结果:通过,无 baseline drift。
## E2E 验证
使用 Maven 启动:
```powershell
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
```
E2E 使用同一 `sessionId` 连续两轮 Chat:
- `sessionId`: `e2e-phase6-chat-codex-20260710-2120`
- `run1`: `run-e2a97696-4398-4abc-90e4-28f45c838f92`
- `run2`: `run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172`
验证结果:
- Chat1 / Chat2 均成功。
- run1 exact trace 返回 run1。
- run2 exact trace 返回 run2。
- session-only latest trace 返回 run2。
- DB 中同一 session 有两条 `diagnosis_run`。
- step/tool rows 按 `run_id` 隔离,mixed row check 为 0。
- `chat_session.message_pair_count = 2`。
## 日志验证
检查:
- `target/e2e/phase6-mvn-20260710-211831.out.log`
- `logs/application.log`
- `logs/chat.log`
结果:能找到 E2E `sessionId`、两个 `runId`、Chat execution、run persistence 和 trace lookup 相关日志。
## 浏览器/人工验证
未单独进行浏览器点击验证。Trace UI 的本次验收通过静态语法检查、URL/runId 参数代码审查和后端 exact trace E2E 共同覆盖。建议后续手动打开 `trace.html?sessionId=...&runId=...` 做展示层冒烟。
## 剩余风险 / 后续事项
- 缺少 `runId` 的 Feedback fallback 是短期兼容路径,客户端全部迁移后可收紧。
- `diagnosis_session` 仍保留为历史兼容和回滚表,后续需要观察窗口后再评估约束收紧或归档策略。
- `case_library.diagnosis_id` 仍是过渡字段,旧值可能为 `session_id`,新自动值为 `run_id`。
- 历史 mixed trace 不能恢复真实多轮边界,只能按 compatibility run 查询。
@@ -0,0 +1,41 @@
# Session / Run / Trace Isolation
## 背景
同一个 `sessionId` 以前同时代表多轮 Chat 上下文和一次持久化诊断 Trace。端到端验证发现,同一 `sessionId` 连续两轮 Chat 时,Redis 多轮上下文是正确的,但 MySQL 中 `diagnosis_session` 会被后一轮覆盖,`agent_step` 和 `tool_invocation` 会按同一个 `session_id` 混在一起。
这会导致 Trace 回放、Verifier/Evaluation 读数、Feedback 绑定和 `case_library` 来源都可能跨轮污染。
## 目标
- 将会话态和运行态拆开:`chat_session` 保存会话元数据,`diagnosis_run` 保存一次诊断运行。
- 引入正式 API 字段 `runId`,作为一次可回放诊断执行的边界。
- `agent_step` 和 `tool_invocation` 保留原 Trace 明细角色,新增 `run_id` 并按 run 隔离读写。
- Trace、Feedback、CaseLibrary、AIOps、demo 脚本和 Trace UI 都支持 run-aware 流程。
- 保留旧 `diagnosis_session` 作为历史兼容和回滚表。
- 完成 Maven E2E、DB 检查、日志检查和 baseline drift 验证。
## 范围
- Flyway/JPA 增加 `chat_session`、`diagnosis_run`,并给 `agent_step`、`tool_invocation` 增加 `run_id`。
- Chat 每次有效执行创建一个新的 `diagnosis_run`,响应返回 `sessionId + runId`。
- Trace API 支持 latest-run fallback 和 exact-run 查询:`GET /api/diagnosis/{sessionId}/trace?runId=...`。
- 新增 run list API:`GET /api/chat/session/{sessionId}/runs`。
- Feedback 优先绑定 `runId`,缺省时短期 fallback 到 latest run 并返回 `fallbackToLatestRun=true`。
- AIOps 每次有效执行创建并透出 `runId`,SSE 保持 `message` event name 并发送 `type=metadata`。
- MVP demo、Trace UI、表文档和架构文档统一为 `chat_session -> diagnosis_run -> trace detail(run_id)`。
## 非目标
- 不新增 `diagnosis_trace` 或 `trace_event` 主表。
- 不实现完整 run-list UI。
- 不删除旧 `diagnosis_session`。
- 不改变 Redis 对话历史窗口策略。
- 不把完整多轮正文历史持久化到 MySQL。
- 不尝试把历史混合 trace 还原成真实多轮边界。
## 关联
- OpenSpec: `openspec/changes/archive/2026-07-10-session-run-trace-isolation`
- Change slug: `session-run-trace-isolation`
- 分档: complex
@@ -0,0 +1,48 @@
# Decisions
## 核心决策
| 决策 | 选择 | 理由 |
|---|---|---|
| 领域拆分 | 新增 `chat_session` 和 `diagnosis_run` | 会话元数据和一次诊断执行的生命周期不同,继续塞在一张表会导致上下文膨胀和边界混淆 |
| Trace 明细 | 复用 `agent_step` / `tool_invocation`,增加 `run_id` | 现有明细表已经能表达 Trace,隔离需要 run key,不需要新事件模型 |
| API 身份 | `runId = "run-" + UUID` | 外部 ID 不依赖数据库自增 ID,碰撞风险低 |
| Trace 兼容 | 缺少 `runId` 时按 `created_at DESC, id DESC` 解析 latest run | 保留旧客户端兼容性,避免 feedback/eval 更新 `updated_at` 后改变 latest 判定 |
| 历史迁移 | 每条旧 `diagnosis_session` 生成一条 compatibility run | 旧混合数据没有真实轮次边界,不能伪造多 run 历史 |
| Feedback fallback | 缺少 `runId` 时短期绑定 latest run 并返回 `fallbackToLatestRun=true` | 老客户端可继续工作,同时让歧义可观测 |
| Case provenance | 新自动案例写 `case_library.diagnosis_id = run_id` | 保留旧列,文档声明过渡语义 |
| AIOps 范围 | 同一个 change 内完成 AIOps run isolation | AIOps 是一等 Trace 入口,不能留下同类混合 trace bug |
| 所有权校验 | 服务层校验 run/session ownership,暂不加 DB 外键 | 兼容历史 orphan rows 和回滚窗口 |
## 用户确认
- 选择拆 `chat_session` 和 `diagnosis_run`,不只是在旧表加字段。
- `chat_session` 第一阶段只保存元数据,不保存完整对话正文。
- 完整多轮对话历史继续放在 Redis `SessionContext.messageHistory`。
- `runId` 是正式 API 字段。
- Trace 缺少 `runId` 时短期默认查 latest run。
- Feedback 缺少 `runId` 时短期 fallback,长期可再收紧。
- 每次有效 Chat/AIOps 都创建 run。
- 不新增 `diagnosis_trace` / `trace_event` 主表。
- 旧 `diagnosis_session` 保留用于历史和回滚,新代码不再写新执行态。
- demo 脚本和 Trace UI 做最小 `runId` 支持。
## 接口影响
级别:L4。
- 新 API 响应字段:`runId`。
- Trace API 新 query 参数:`runId`。
- 新 API:`GET /api/chat/session/{sessionId}/runs`。
- Feedback request 新增 optional/preferred `runId`。
- Feedback response 新增 bound `runId` 和 `fallbackToLatestRun`。
- `/api/ai_ops` SSE 保持 event name `message`,新增 `type=metadata` 消息。
- DB contract 新增两张表和两个 `run_id` 列。
- 旧 `sessionId` only 调用仍兼容,但 fallback 必须可观测。
## 风险接受
- 历史混合 trace 无法真实拆分,只能作为 compatibility run。
- 上下文传播同时依赖 `RunnableConfig.metadata` 和 `SessionContextHolder`,后续改动必须注意 `sessionId/runId` 同步。
- `case_library.diagnosis_id` 在过渡期存在 `session_id` 和 `run_id` 两种语义。
- 缺少 `runId` 的 Feedback 仍有歧义,后续客户端迁移完成后可收紧为参数错误。
@@ -0,0 +1,76 @@
# Evidence
## 上下文证据
- `SessionContext.messageHistory` 和 `getMessagePairCount()` 证明 Redis 承载热对话历史;MySQL 只需要长期审计的会话目录和运行记录。
- `CaseLibraryService.createFromSession` 原先按 `DiagnosisSession.sessionId` 去重并映射 query/answer,因此 run 隔离后需要新增 `createFromRun`。
- 旧 `mvp/architecture/data-model.md` 把 `case_library.diagnosis_id` 解释为 `diagnosis_session.session_id`,本次改为过渡语义:旧数据可能是 `session_id`,新自动案例是 `run_id`。
- 既有 Trace OpenSpec 要求 `GET /api/diagnosis/{sessionId}/trace` 是只读端点;latest-run 和 exact-run 查询都必须保持只读。
- ISS-010 的 E2E 事实显示同一 `sessionId` 两轮 Chat 会产生 MySQL Trace 混合,是本 change 的直接触发证据。
## 实现证据
- Phase 1 增加 `V011__add_session_run_isolation.sql`,创建 `chat_session`、`diagnosis_run`,并为 `agent_step` / `tool_invocation` 增加 nullable `run_id`。
- Phase 2 将 Chat 写路径切到 `chat_session + diagnosis_run`,并让 Hook/Tool/Evaluation/Gatekeeper 使用 run-scoped 数据。
- Phase 3 将 Trace API 改为 latest-run / exact-run 双模式,并加入 lightweight run summaries。
- Phase 4 将 Feedback 和 CaseLibrary 绑定到 run,保留没有 run-backed 数据时的 legacy fallback。
- Phase 5 将 AIOps 接入 run isolation,SSE metadata 暴露 `sessionId + runId`。
- Phase 6 更新 demo 脚本、Trace UI、MVP 架构文档和表文档,并修正 review 后发现的 session-only 文档残留。
## E2E 证据
Maven 启动命令:
```powershell
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
```
日志:
- `target/e2e/phase6-mvn-20260710-211831.out.log`
- `target/e2e/phase6-mvn-20260710-211831.err.log`
- `logs/application.log`
- `logs/chat.log`
E2E session:
- `sessionId`: `e2e-phase6-chat-codex-20260710-2120`
- `run1`: `run-e2a97696-4398-4abc-90e4-28f45c838f92`
- `run2`: `run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172`
结果:
- 两轮 Chat 都成功,并复用同一个 `sessionId`。
- 两轮返回不同 `runId`。
- run1 exact trace 只返回 run1。
- run2 exact trace 只返回 run2。
- session-only Trace latest fallback 返回 run2。
- `chat_session.message_pair_count = 2`,证明多轮上下文连续。
## DB 证据
通过 `scripts/query_mysql.py` 检查:
- `diagnosis_run` 中该 E2E session 有 2 条 `SUCCESS / CHAT` 运行。
- `agent_step` 按 run 分组:run1 `10` 行,run2 `9` 行。
- `tool_invocation` 按 run 分组:run1 `14` 行,run2 `8` 行。
- mixed row check 为 `0`,没有 NULL 或 unexpected `run_id` 混入该 E2E session。
## Baseline 证据
运行:
```powershell
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest" test
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
```
结果:
- 两组 baseline / regression 命令通过。
- baseline harness 使用离线 fixture,不依赖 live DB/session tables。
- 未观察到 baseline drift。
## 工具限制
AGENTS 要求的 `codebase-retrieval` 和 LSP 工具在本会话不可用。替代验证使用 OpenSpec、`rg`、定向阅读、 focused tests、E2E、DB 查询和日志检查。
@@ -0,0 +1,52 @@
# Chat Diagnosis StateGraph ChatService Cutover 验收
## 结果
已接受。OpenSpec tasks 26/26 完成,阶段 3 可归档。
## 验证
### 静态验证
- `git diff --check`:通过。
- source/reference isolation:cutover 核心文件中的 `SequentialAgent`、`VerifierContextHolder`、`VerifierInputHook`、`tool_trace_summary` 命中 0;核心 TODO/FIXME/placeholder 命中 0。
- schema whitelist:V012 为 1 个 ALTER TABLE、1 个 ADD COLUMN、0 CREATE、0 DROP,目标仅 `diagnosis_run.orchestration_trace JSON NULL`。
- `openspec validate chat-diagnosis-stategraph-chatservice-cutover --strict`:通过。
- `openspec validate --specs --strict`:14 passed,0 failed。
### 脚本验证
- focused Maven regression:Graph runtime/result mapper/real nodes、ChatService cutover、Trace、Controller、Repository、Gatekeeper、Composer、Eval,共 27 suites / 119 tests;0 failures、0 errors、0 skipped。
- `mvn -q -DskipTests test-compile`:通过。
- `ChatServiceGraphIntegrationTest` 覆盖 SUCCESS、handled Fallback、unhandled failure、blank answer/partial trace 和同 session 多 Run 隔离。
### 浏览器/人工验证
- 未运行。阶段 3 的公开协议由 Controller/service tests 覆盖,完整人工/live 验收按用户规则保留到阶段 5。
### 未验证
- 未使用 Maven 启动应用做 live E2E。
- 未检查 `logs/` 运行日志。
- 未执行 `scripts/query_mysql.py` 查询真实数据库。
- 原因:用户明确要求只有阶段 5 全部完成后统一执行端到端、日志和数据库验收;阶段 3 只做风险相关自动化验证。
- 剩余风险:真实模型/工具调用下的 Prompt 行为、Flyway 在真实 MySQL 的应用结果和最终 Trace 数据须由阶段 5 E2E 证明。
## 已完成范围
- 复杂 Chat 唯一生产编排切换到 Diagnosis StateGraph,公开 ChatResult 保持兼容。
- Run 生命周期、metrics、Eval、finally cleanup 与 runId 隔离保持;handled Fallback=SUCCESS,未处理/空答案=FAILED。
- 新增 Run 级 compact orchestration trace,并只在 Trace run 对象暴露解析结果。
- Verifier Prompt/self-evaluation 使用 verified-only 数据,不伪造未发生 verdict/event。
- 移除冲突的 Sequential 实现测试并以 Graph/public contract tests 替代。
## Bug 修复和诊断
- 替代测试错误引用 `com.superbiz.agent.tool.ToolInvocationRecorder`:通过定义/引用搜索确认类型位于同一 `service` 包,删除错误 import。
- no-answer 新测试错误期待 inner cause:分类为测试断言偏差,改为公开 wrapper error,同时保留 FAILED 与真实 partial trace 核心断言。
## 交接
- OpenSpec archive:已同步 5 份 delta specs,并归档到 `openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-chatservice-cutover/`。
- 下一步:审查精确 Git diff 并完成阶段 3 独立提交,之后才启动阶段 4。
- OpenSpec 归档确认:用户已明确授权后续阶段直接实现/归档;归档已完成。
@@ -0,0 +1,21 @@
# Chat Diagnosis StateGraph ChatService Cutover Brief
## 背景
- 用户目标:将 ISS-011 阶段 3 作为独立 sm-flow,正式切换复杂 Chat 生产编排并增加 Run 级 orchestration trace。
- 当前问题:真实 Diagnosis Graph Nodes 已存在,但生产入口仍依赖 SequentialAgent、外层重试和 ThreadLocal/Hook 隐式状态。
- 关联 OpenSpec:`openspec/changes/chat-diagnosis-stategraph-chatservice-cutover/`
- devflow 分档:complex。
## 范围
- 本次要做:复杂 Chat 单轨 Graph cutover;复用四类 Agent builder;Graph result/self-evaluation 映射;V012 Run JSON 字段;Trace run-only 投影;verified-only Verifier Prompt;必要回归测试。
- 本次不做:不改 `/api/chat` 请求/响应,不做历史 trace 回填,不删除仍供历史代码/测试使用的 Hook/ThreadLocal 类型,不完成阶段 4 测试体系全面收敛,不运行 live E2E/log/DB 验收。
- 影响区域:ChatService、Diagnosis Graph runtime/result mapper、DiagnosisRun/Flyway、Trace DTO/service、Verifier Prompt、Graph/Service/Trace tests。
## OpenSpec 对齐
- proposal 覆盖状态:已覆盖。
- design 覆盖状态:已覆盖。
- specs 覆盖状态:已覆盖,5 个 delta capabilities。
- tasks 覆盖状态:26/26 已完成。
@@ -0,0 +1,207 @@
# Chat Diagnosis StateGraph ChatService Cutover Decisions
## Entry Summary
- 问题:复杂 Chat 仍使用 SequentialAgent + ThreadLocal/Hook 隐式状态机,真实 Graph 尚未成为生产入口,也没有 Run 级 orchestration trace 持久化和 API 投影。
- 期望:阶段 3 独立完成生产 cutover、Run/Trace 映射和必要测试,归档并提交后才进入阶段 4。
- 分档:complex。
- Change:`chat-diagnosis-stategraph-chatservice-cutover`。
- 授权:用户已要求后续阶段直接实现,不再逐 checkpoint 等待;阶段门禁、独立 archive/commit 与阶段 5 才 E2E 约束不变。
## Context Sources
- `mvp/issues/active/ISS-011-chat-diagnosis-stategraph-orchestration.md` 阶段 3、协议影响、Run 状态和验收章节。
- 阶段 0–2 OpenSpec archives、devflow acceptance/decisions 与当前四份 Graph 主 specs。
- `devflow/glossary/CONTEXT.md` 中 Chat Session、Diagnosis Run、Diagnosis Trace、Diagnosis Orchestration Trace 的边界。
- `ChatService` 生产调用链、四个 Agent builder、Run persistence/self-evaluation/metrics 逻辑。
- `DiagnosisRun`、V011 migration、`DiagnosisTraceResponse`、`DiagnosisTraceService` 与相关 tests。
- `DiagnosisGraphFactory`、真实 action factory、Node adapters、final trace builder 和本地 CompiledGraph API。
## Question Pool
| # | 维度 | 问题 | 模式 | 状态 |
|---|---|---|---|---|
| Q1 | 术语 | orchestration trace 是否属于 self-evaluation 或完整 Trace event log? | evidence-driven | 已解决 |
| Q2 | 边界 | 阶段 3 是否改变 `/api/chat`,以及 Trace 字段出现在哪一层? | user-interview(既有冻结决策) | 已确认 |
| Q3 | 生命周期 | LOW_CONFID/REJECT/Fallback 是否应使用 FAILED 或新增 DEGRADED? | user-interview(既有冻结决策) | 已确认 |
| Q4 | 错误处理 | Graph 异常时如何 best-effort 保存已有编排事实而不伪造 event? | evidence-driven | 已解决 |
| Q5 | 技术实现 | 是否复用现有 Agent factory/Hook/ToolCallback 与真实 Node assembly? | evidence-driven | 已解决 |
| Q6 | 安全 | Graph Verifier Prompt/self-evaluation 是否可继续使用 raw Executor 和完整 tool trace? | evidence-driven | 已解决 |
| Q7 | 验收 | 阶段 3 是否需要新增测试,是否现在运行 live E2E? | user-interview(用户最新规则) | 已确认 |
## Evidence-driven Findings
- Q1:glossary 与阶段 0 spec 已定义 orchestration trace 为 Run 级紧凑编排摘要,与 self-evaluation、AgentStep/ToolInvocation 和 checkpoint 分离;无需新增术语或 ADR。
- Q4:所有已定义 Agent/Java 失败应由 Graph 路由到 handled Fallback 并返回 final state;只有真实可取得的 final/partial state 才可构造 trace。无法取得 state 的未处理异常标记 Run FAILED,不得生成虚假 transition。
- Q5:`DiagnosisRealGraphActionsFactory` 已提供真实 Node assembly;ChatService 现有四个 builder、AgentLoggingHook、Skills hooks、method tools 和 ToolCallbacks 可直接构造 ReactAgent,再通过 `ReactAgentDiagnosisInvoker` 注入,不得复制 Prompt/Agent factory。
- Q6:阶段 2 spec 已禁止 Graph Verifier 使用 raw Executor/full tool trace;当前 `chat-verifier-prompt.md` 仍描述旧 Hook payload,是阶段 3 必须同步的明确 gap。self-evaluation 可保存 Graph 的 verified output、Gatekeeper audit 和 Composer audit,但不能重新引入 raw 输入。
## User-interview Confirmations
| 问题 | 用户原话/既有确认 | 确认状态 | OpenSpec 回写 |
|---|---|---|---|
| Q2 外部协议与 Trace 层级 | ISS-011 已冻结 `/api/chat` 不变,Trace 只在 `run.orchestrationTrace` 增解析对象 | 已确认 | proposal |
| Q3 Run 生命周期 | ISS-011 已冻结安全 Fallback 为 SUCCESS,不新增 DEGRADED;只有无法生成安全响应的未处理失败为 FAILED | 已确认 | proposal |
| Q7 验收节奏 | “端到端只在最后阶段全部完成后才验证;每个阶段如果有必要添加单元测试验收的话,就加” | 已确认 | proposal |
## Interface Impact
- 等级:L4。
- 原因:复杂 Chat 内部状态机正式切换;DiagnosisRun 增加数据库 JSON 契约;Trace run 对象新增字段;旧固定顺序消费者/测试不再成立。
- 保持兼容:`/api/chat` request/response、sessionId/runId、Executor/Verifier/Composer 输出契约不变。
- 加法变化:只新增 `diagnosis_run.orchestration_trace` 和 `run.orchestrationTrace`。
- 回滚:revert 阶段 3 代码/Prompt/spec,数据库列可保留 nullable;不保留运行时双轨开关。
## Discover Status
- `devflow/index.md`:命中阶段 0–2 archives 和 Run/Trace 历史项目。
- Glossary:相关术语已存在且无冲突,不需更新。
- ADR:阶段 0 已记录难以逆转的状态机/Trace 决策,本阶段没有新的三条件 ADR。
- 未解决问题:0。
- Draft 产物:当前只创建 proposal + decisions;design/specs/tasks 留到 Commit checkpoint。
## Grill-with-docs Review
### Domain model stress test
- 同一 session 连续两个复杂 Chat run:每次编译/执行使用自己的 runId threadId 和 metadata;orchestration trace 只写各自 DiagnosisRun,session 投影不复制,符合 Session/Run/Trace 领域边界。
- Gatekeeper REJECT 或 Planner/Executor 技术失败:Graph 进入确定性 Fallback,返回安全非空答案;Run 为 SUCCESS,trace `degraded=true`,不会把诊断质量塞进 Run status。
- Composer 成功但 verdict=LOW_CONFID/REJECT:Run 仍为 SUCCESS;self-evaluation 保存 effective verdict,orchestration trace 保存实际路径,两者职责不混合。
- Graph 未处理异常:使用流式 NodeOutput 捕获最后一个真实 state,只有其中已有 events 时才 best-effort 构造 partial trace;没有 event 时不伪造 transition,Run 标记 FAILED。
- 历史/AI_OPS run:新增列 nullable,Trace DTO 对旧 run 可返回 null;不增加历史回填或跨 flow 假数据。
### Design tree conclusions
- 生产切换采用单轨,不增加 feature flag 或保留 Sequential/Graph 双运行;回滚依靠 Git revert,nullable 列可保留。
- ChatService 负责 Run 生命周期和调用一个专用 Graph runtime/orchestrator;复杂 Agent 构造从旧 Sequential 私有流程中搬移或封装复用,不复制第二套 builder。
- Graph 执行优先使用可观察的 `CompiledGraph.stream(..., config)` 收集最后真实 state,以支持异常时 best-effort trace;正常终止仍以 END state 的 `final_answer` 为唯一答案来源。
- Graph result mapping 形成独立组件:final answer、Verifier evaluation、Composer audit、Gatekeeper audit、trace JSON 均从显式 final state读取,不再访问 VerifierContextHolder。
- Verifier Prompt 必须更新为 `diagnosis_context`、`verified_executor_output`、`verified_evidence`、`gatekeeper_audit`、`verdict_ceiling` 和可选 retry context;删除 raw Executor/full tool trace/Hook Gatekeeper 说明。
- 阶段 3 必须有 production cutover 与 Trace DTO/Service tests;阶段 4 再做测试体系全面改名、夹具收敛和旧测试删除。
### Documentation result
- 术语与 `devflow/glossary/CONTEXT.md` 完全一致,无需修改 glossary。
- 状态机、Run status 和 orchestration trace 隔离均来自阶段 0 已归档 ADR/规格,不创建重复 ADR。
- proposal 已反映单轨切换、L4 接口影响、流式 partial-state 处理、Prompt 安全边界和阶段 5 E2E 延期。
- Grill question pool 全部关闭;无需要再次询问用户的产品取舍。
## Architecture Audit
### Module and caller map
`ChatController` 的 normal/SSE 两个入口都调用 `ChatService.executeChatWithStrategy`,复杂分支进入 `executeChatComplex`;对外仍只消费 `ChatResult(answer, sessionId, runId)`。阶段 3 将内部链路变为 `ChatService Run lifecycle -> complex Chat Graph runtime/Agent assembly -> real Diagnosis Nodes -> Graph result mapper -> DiagnosisRun persistence -> DiagnosisTraceService`。Trace UI/Eval 当前读取兼容 `session.selfEvaluation`,它仍由 Run self-evaluation 投影;新增 orchestration trace 只属于 `run`。AIOps、Feedback、CaseLibrary 和 Run list 继续使用 DiagnosisRun 既有字段,不消费新增 orchestration trace。
| 模块 | 所有权 | 允许的依赖/影响 |
|---|---|---|
| ChatController | `/api/chat` normal/SSE 协议 | 继续只依赖 ChatResult;无字段变化 |
| ChatService | Chat Session/Diagnosis Run 生命周期、成功/失败保存、metrics、Eval、finally cleanup | 调用一个 Graph runtime/result mapper;不再拥有条件边/重试/Gatekeeper/Composer 路由 |
| Complex Chat Graph runtime | 每请求 Agent assembly、initial state、RunnableConfig、CompiledGraph stream | 复用现有 Prompt/tool/skill/logging;不持久化跨 Run 状态 |
| Diagnosis Graph | Node status、verified material、events、final answer | 保持阶段 1/2 路由和 counter 所有权 |
| Graph result mapper | final/partial state 到安全 evaluation/trace DTO | 不读 ThreadLocal,不查询其他 run,不持久化 raw material |
| DiagnosisRun/Flyway | 当前 Run 的 orchestration JSON | 仅一个 nullable JSON 列;历史/AIOps 可为 null |
| DiagnosisTraceService/DTO | exact/latest Run 查询和 API 投影 | 只在 RunTrace 加 parsed map;session/top-level/run-list 不重复 |
### Lifecycle and coupling audit
ReactAgent 与 CompiledGraph 都按请求构造,因为 Planner/Executor system Prompt 含本次 history/knowledge map;Graph State 和 runtime last-state holder 也是 invocation-scoped,不进入 singleton 可变字段。SessionContextHolder 仍只包围当前请求并在 finally 清理;Graph Node config 以 runId threadId + metadata 绑定 AgentStep、ToolInvocation 和 Gatekeeper。success persistence 在 Eval 之前完成,EvaluationService 再按 runId 合并 rule channel;handled Fallback 与 Composer 使用同一 SUCCESS 路径。Trace JSON 与 self-evaluation 分栏,避免把路线、质量和完整 Agent/tool 明细耦合到一个容器。
### Consumer impact audit
- ChatController normal/SSE:ChatResult 协议不变,L4 内部切换不要求调用方迁移。
- Trace UI/demo:继续读取兼容 `session.selfEvaluation`;新的 `run.orchestrationTrace` 是加法字段,最终脚本断言留到阶段 5。
- DiagnosisTraceEvaluator:现有 fixtures 不变;工具覆盖仍可从 `toolInvocations` 读取,兼容 `executor_structured_output` 保存 verified projection。pre-verification Fallback 质量未来应以 trace degraded 而非伪 verdict 判断,属于阶段 4 测试体系/阶段 5文档收尾。
- AIOps/Feedback/CaseLibrary/Run list:实体新增 nullable 字段不改变 builder call sites或查询语义。
- 数据库:Hibernate validate 要求 V012 与 entity 同批;回滚代码可忽略保留列。
### Cross-artifact alignment
| 对齐链 | 结果 | 证据 |
|---|---|---|
| brief/proposal 目标、范围、非目标 → proposal | 已对齐 | 单轨 cutover、Run/Trace、Prompt、阶段边界与 E2E 延期均明确 |
| proposal 承诺与约束 → design | 已对齐 | 10 项决策覆盖 runtime、state、config、partial state、result、persistence、DB/API/Prompt/tests |
| design 架构/接口结论 → specs/tasks | 已对齐 | L4、单轨、Run status、verified-only、Trace 唯一投影和 schema whitelist 均有 requirement/task |
| specs 可观察行为 → tasks 可执行切片 | 已对齐 | 6 组 26 个切片覆盖数据、runtime、Prompt、cutover、tests、handoff |
### Audit result
未发现与阶段 0–2、Run/Trace ADR 或 glossary 冲突。审计确认不能把 Agent factory 复制进 Node,也不能把 pre-verification Fallback 伪装为 Verifier verdict;两项已在 design/specs/tasks 固定。唯一跨阶段依赖是阶段 4 让测试/Eval 以 orchestration degraded 理解无 Verifier verdict 的安全 Fallback,已记录但不阻塞阶段 3 production correctness。架构风险可接受,cross-artifact gap=0,无未解决接口消费者。
## Commit Gate
- schema:spec-driven;proposal/design/specs/tasks 全部 done,applyRequires=`tasks` 已满足。
- OpenSpec:当前 change strict validation 通过;14 个主 specs 全部 strict pass。
- 规格结构:5 个 delta capabilities、23 条 requirements、72 个 scenarios;tasks 26 个可执行 checkbox。
- Cross-artifact:4/4 已对齐,gap=0。
- Interface impact:L4;design 已独立记录消费者、加法 DB/Trace 协议、单轨迁移和 Git revert 回滚。
- Question pool:所有 evidence-driven 已查证;所有 user-interview 已由 ISS-011/用户原话确认;无未决项。
- Preflight:`git diff --check` 通过;Commit checkpoint 尚未修改 Java、SQL、Prompt 或测试。
- 结论:Draft OpenSpec 已达到可执行状态,创建 `.committed`,阶段 3 Apply 只能以这些产物为依据。
## Apply Authorization
- 用户原话:“直接实现吧,不用找我授权了”。
- 本阶段在 Commit gate 后直接进入 Apply;不扩大到阶段 4 全面测试迁移或阶段 5 live E2E。
## Pre-apply Research
### Reference implementations read
- `ChatService.executeChatComplex`、四个 complex Agent builders、Run start/save、metrics、Prompt audit、self-evaluation merge:迁移源实现和外部兼容基线。
- `DiagnosisGraphFactory`、`DiagnosisRealGraphActionsFactory`、四个 Agent adapters、Gatekeeper/VerifiedInput/Fallback、`DiagnosisOrchestrationTraceBuilder`:Graph 路由/计数/安全材料唯一真理源。
- `DiagnosisRun`、V011 migration、`DiagnosisTraceResponse`、`DiagnosisTraceService`、`DiagnosisTraceServiceTest`:Run JSON 字段和 exact/latest Trace 映射标准。
- `ChatController` normal/SSE、`DiagnosisTraceController`:公开协议调用者与响应包装边界。
- `SelfEvaluationMergeService`、`EvaluationService`、`DiagnosisTraceEvaluator`:evaluation container、异步 rule merge 和兼容消费者。
- 本地 graph-core 1.1.2.0 `javap`:CompiledGraph `stream/invoke/state`、NodeOutput `state()`、RunnableConfig `threadId/metadata` 的真实 API。
- `chat-*-prompt.md`:现有 Prompt 组装与 Verifier 旧 Hook payload gap。
### Technology inventory
| 类别 | 项目标准 / 本阶段使用 |
|---|---|
| Graph execution | `CompiledGraph.stream(initial, config)` + NodeOutput.state,当前请求线程阻塞消费 |
| Run context | SessionContextHolder + RunnableConfig threadId/metadata;runId 是唯一执行边界 |
| Agent assembly | ReactAgent builder、AgentLoggingHook、PlannerSkillMetadataHook/SkillsAgentHook、现有 tools/callbacks |
| JSON persistence | Jackson map serialization;实体 String + `@JdbcTypeCode(SqlTypes.JSON)`;Flyway JSON column |
| Trace API | Lombok DTO builder + DiagnosisTraceService parsed Map;exact/latest Run 查询 |
| Evaluation | SelfEvaluationMergeService container;EvaluationService 按 runId 异步合并 rule channel |
| Tests | JUnit 5,通过 public service/runtime/Trace API;只 mock repository/model/tool 外部边界 |
| MQ/Consumer | 不涉及 |
### New infrastructure and reuse
- 新增一个深接口的 complex Chat Graph runtime/result mapper;复用既有 Graph/Node/Agent builders,不增加新依赖或第二套路由。
- 新增 V012、DiagnosisRun 字段和 RunTrace parsed map;不新增表、repository method 或历史 backfill。
- TDD tracer bullet 从 Trace DTO/Service 的唯一 Run 投影开始,再进入 runtime final-state mapping,最后切换 ChatService。
- 研究未发现 devflow/OpenSpec 冲突,技术清单足以开始实现。
## Apply Progress
- TDD Trace slice RED:DiagnosisTraceServiceTest 明确缺少 DiagnosisRun builder 字段和 RunTrace getter。
- GREEN:V012、DiagnosisRun JSON 字段、RunTrace parsed map、DiagnosisTraceService mapping 完成;10 个 Trace service tests 通过。
- 失败分类:首次 GREEN 运行仅测试夹具用裸 ObjectMapper 无法序列化 LocalDateTime,属于 test harness 偏差;改为项目可用的 `findAndRegisterModules()` 后原始 focused loop 通过,未改业务协议。
- Trace API JSON 断言证明顶层/session 无重复字段、run 为解析对象、无 raw 字段;null/invalid JSON fail closed。
## Apply Completion
- 复杂 Chat 已单轨切换到 `ChatDiagnosisGraphRuntime`;生产 `ChatService` 不再创建 `SequentialAgent`,不再读取 `VerifierContextHolder`,Graph Verifier 只保留 `AgentLoggingHook`。
- runtime 使用 query-only initial state、`threadId=runId` 和 sessionId/runId metadata,流式保留最后真实 state;异常只保存真实 partial state/trace。
- `DiagnosisGraphResultMapper` 只从 verified projection 构造兼容 self-evaluation;Verifier 未完成时不伪造 verdict,不持久化 full tool trace/raw Executor。
- Verifier Prompt 与 runtime payload 已统一为 verified-only,Prompt audit 更新为 `chat-prompts-v2` / `chat-verifier-v3`。
- 旧 `ChatServiceSequentialAgentTest` 因绑定已删除的固定顺序/ThreadLocal/score retry 实现而移除,由 public service/runtime/route tests 替代。
- V012 schema 白名单测试确认只新增 `diagnosis_run.orchestration_trace JSON NULL`,无其他 schema object。
## Verification Summary
- focused regression:27 suites / 119 tests,0 failures、0 errors、0 skipped。
- Maven test compilation:通过。
- OpenSpec:当前 change strict pass;主 specs 14/14 strict pass。
- 静态门禁:`git diff --check` 通过;cutover 禁用引用 0;核心 TODO/placeholder 0;schema whitelist 为 1 ALTER / 1 ADD COLUMN / 0 CREATE / 0 DROP。
- 实现期唯一失败分类:no-answer service test 曾错误期待 inner cause 文本,属于测试断言偏差;按公开 wrapper error + FAILED/partial trace 契约修正后通过,OpenSpec 与生产代码无需变更。
- 按阶段门禁未运行 Maven live E2E、未检查 `logs/`、未执行 `scripts/query_mysql.py`;统一保留到阶段 5。
## Archive Result
- 5 份 delta specs 已同步:新增 9、修改 13、删除 1 条 requirements。
- OpenSpec 已归档到 `openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-chatservice-cutover/`。
- CLI 对 proposal 结构给出非阻塞建议(大 change/delta 拆分与 SHALL/scenario 启发式);tasks 26/26、delta specs 和 strict validation 均通过,未形成验收阻塞。
@@ -0,0 +1,21 @@
# Chat Diagnosis StateGraph ChatService Cutover Evidence
## 证据
| 来源 | 证据 | 结论 | 是否已汇报 |
|---|---|---|---|
| `ChatService` + source isolation search | 复杂路径只有一次 Graph runtime 调用,无 SequentialAgent、VerifierContextHolder、VerifierInputHook 或 score retry | 单轨 cutover 完成,简单 Chat 路径未改 | 是 |
| `ChatDiagnosisGraphRuntimeTest` / `DiagnosisRealGraphIntegrationTest` | query-only state、runId thread/metadata、Composer/Fallback、empty stream、blank answer、partial trace、Gatekeeper once/twice | Graph identity、路由和真实 partial-state 边界可观察 | 是 |
| `DiagnosisGraphResultMapperTest` / `ChatVerifierPromptContractTest` | verified projection、effective verdict 兼容、pre-verification 无伪 verdict、禁用 raw/full trace 字段 | self-evaluation 与 Prompt 安全边界一致 | 是 |
| `ChatServiceGraphIntegrationTest` | ChatResult、agent_flow、SUCCESS/Fallback/FAILED、metrics/Eval、partial trace、多 Run 隔离 | 生产生命周期与兼容返回满足阶段 3 规格 | 是 |
| `DiagnosisTraceServiceTest` | `run.orchestrationTrace` 解析对象、top/session 无重复、invalid/null fail closed | Trace 字段所有权保持 Run 隔离 | 是 |
| `DiagnosisRunSchemaContractTest` | V012 executable SQL 与唯一允许语句精确相等 | schema 仅增加一个 nullable JSON 列 | 是 |
| focused Maven suite | 27 suites / 119 tests,0 failures/errors/skipped | Graph、Controller、Repository、Gatekeeper、Composer、Eval 回归通过 | 是 |
| OpenSpec/static gates | change strict、14/14 主 specs、test compile、diff check、禁用引用/占位/schema whitelist 全通过 | 规格、编译和源级隔离闭环 | 是 |
## Evidence-driven 结论
- orchestration trace 必须独立于 self-evaluation 和详细 Agent/tool Trace;V012 与 RunTrace 的实现保持这一边界。
- handled Fallback 表示安全降级答案,Run 仍为 SUCCESS;未处理、空答案或 invariant 失败才是 FAILED。
- Verifier/Gatekeeper 只以显式 Graph state 传递可信材料;继续使用 Hook/ThreadLocal 会形成双 Gatekeeper 和不安全输入,因此生产路径已彻底移除该依赖。
- 旧 Sequential 测试验证的是已废止内部机制,替换为 public service + runtime + route contract tests 才能保持真实回归价值。
@@ -0,0 +1,49 @@
# Acceptance
## 静态验证
- 旧闭包路径不存在,executable legacy refs=0。
- current-doc stale architecture refs=0;历史兼容引用有明确限定。
- `git diff --check` 通过;无新增 migration/schema 变更;临时调试标记=0。
- 当前 change strict 和 16 个 main specs strict 全部通过。
## 脚本验证
- `mvn -q -DskipTests test-compile`:通过。
- 39-suite authoritative/focused Maven command:157 tests,0 failure/error/skipped。
- 归档前 `mvn clean` + test compilation + 43-suite deterministic command:189 tests,0 failure/error/skipped。
- `DiagnosisTraceEvaluatorTest` + `DiagnosisEvalBaselineDiffTest`:12/12 baseline,same diff=0。
- PowerShell parser + `InterviewDemoScriptContractTest`:通过。
- `run-interview-demo-check.ps1 -SessionId iss-011-stage5-20260720015557 -OutputDir target/iss-011-stage5-output-current`:exit 0。
- `scripts/query_mysql.py` exact queries:V012、Run JSON、AgentStep/ToolInvocation ownership 全部通过。
- `openspec validate --all --strict --no-interactive`:17/17 passed;`git diff --check`、current-doc、schema、debug source、port/temp scope checks 通过。
## Live E2E
| 项目 | 结果 |
|---|---|
| Maven profile | `mvp-demo` |
| sessionId | `iss-011-stage5-20260720015557` |
| runId | `run-808ac38f-3ad0-4462-a6d0-ed50d8686473` |
| Run | `CHAT/SUCCESS` |
| answer / metrics | 109 chars / 75964ms / 111802 tokens / 8 steps / 12 tools |
| Graph | `stategraph-v1`, `fallback`, `fallback_completed`, degraded=true, 3 transitions, retry=0 |
| evaluation / feedback | non-empty / useful |
| new ERROR | 0 |
| DB ownership | wrong owner=0,wrong-session rows=0 |
| process cleanup | owned PIDs stopped,9900 released |
## 浏览器/人工验证
- 不适用。本阶段验收入口是 API/PowerShell executable contract,无 UI 改动。
## 未验证
- 无 OpenSpec 必需项未验证。
## 归档状态
- ISS-011 已归档至 `mvp/issues/archived/ISS-011-chat-diagnosis-stategraph-orchestration.md`。
- OpenSpec 归档路径:`openspec/changes/archive/2026-07-20-chat-diagnosis-stategraph-cleanup-docs`。
- OpenSpec CLI 已同步主 specs:新增 `chat-diagnosis-stategraph-cleanup-docs`,更新 `mvp-demo-trace-acceptance` 3 项 requirement。
- 不 push。
@@ -0,0 +1,32 @@
# Chat Diagnosis StateGraph Cleanup And Final Acceptance
## 背景
ISS-011 阶段 0-4 已完成 StateGraph 设计冻结、路由骨架、真实 Nodes、ChatService 单轨切换和三层权威测试。阶段 5 负责删除旧 Hook/ThreadLocal/full-trace service 闭包、对齐当前文档与 demo executable contract,并以唯一 Run 完成最终 Maven、日志和数据库验收。
## 目标
- 只保留 bounded StateGraph 复杂 Chat 编排和 verified-only Verifier 输入。
- 让 current architecture/eval/demo 文档与 Run-owned orchestration trace 一致。
- 先通过确定性回归,再用 Maven `mvp-demo` 证明 exact Chat/Trace/feedback、日志和数据库 ownership。
- 所有门禁通过后关闭 ISS-011,并归档阶段 5 OpenSpec。
## 范围
- 删除 `VerifierInputHook`、`VerifierContextHolder`、`ToolTraceSummaryService` 及其 focused test。
- 更新 current architecture/eval/demo 文档和 interview demo check。
- 修复 live 暴露的 Graph event classloader 边界与 nested ReactAgent resume config 问题。
- 完成 Graph/Chat/Trace/Eval 回归、Maven live E2E、日志/MySQL 核验、进程清理和 Issue 生命周期收口。
## 非目标
- 不修改公开 Chat/feedback API、Executor/Verifier/Composer 业务协议或数据库 schema。
- 不重写 archived issues、历史 design notes 和 legacy fixtures。
- 不删除旧 Trace/fixture 对 `tool_trace_summary` 的只读兼容。
- 不 push,不删除失败尝试的审计数据。
## 元数据
- 分档:complex
- OpenSpec:`chat-diagnosis-stategraph-cleanup-docs`
- 接口影响:L2 内部类型/状态表示修复;外部 API/DTO/schema 不变
@@ -0,0 +1,159 @@
# Chat Diagnosis StateGraph Cleanup, Final Acceptance And Documentation Decisions
## Entry Summary
- 问题:ISS-011 运行时已切换且测试体系已收敛,但旧 Hook/ThreadLocal/service死代码、当前架构文档和最终 live 证据尚未闭环。
- 期望:阶段 5完成清理、文档、自动化回归、eval、Maven E2E、日志/DB 验收、Issue 归档和独立提交。
- 分档:complex;接口影响 L2 内部删除 + 文档/demo 验收增强,外部 API/DB 协议不变。
- Change:`chat-diagnosis-stategraph-cleanup-docs`。
- 授权:用户已明确要求直接实现;本阶段按此前规则执行唯一最终 E2E。
## Context Sources
- ISS-011 阶段 5、测试策略、协议影响、验收标准和冻结决策。
- 阶段 0–4 OpenSpec archives、devflow acceptance 与提交 `581daff`、`42ba204`、`1460dd1`、`99e490f`、`208a231`。
- 全仓 `VerifierInputHook`/`VerifierContextHolder`/`ToolTraceSummaryService` 定义与引用搜索。
- `mvp/architecture/README.md` 列出的 current docs、`mvp/eval/README.md`、`mvp/demo/` scripts/checklist。
- `mvp-demo` profile、payment-timeout request、`scripts/query_mysql.py` 和 logs/ 现有布局。
## Question Pool
| # | 维度 | 问题 | 模式 | 状态 |
|---|---|---|---|---|
| Q1 | 清理 | 旧 Hook/ThreadLocal/trace summary service 是否还有生产消费者? | evidence-driven | 已解决 |
| Q2 | 文档 | 哪些旧引用应更新,哪些历史材料应保留? | evidence-driven | 已解决 |
| Q3 | E2E | 最终 live 场景如何绑定唯一 session/run 并证明 Graph 路径? | evidence-driven | 已解决 |
| Q4 | 日志/DB | 如何避免用旧日志/latest DB 记录冒充当前证据? | evidence-driven | 已解决 |
| Q5 | 验收 | 何时允许启动 Maven、是否需要日志和 DB 查询? | user-interview(用户最新规则) | 已确认 |
| Q6 | 关闭 | 何时把 ISS-011 从 active 移到 archived? | evidence-driven | 已解决 |
## Evidence-driven Findings
- Q1:旧闭包只有 `VerifierInputHook -> VerifierContextHolder + ToolTraceSummaryService`,以及 `ToolTraceSummaryServiceTest`;ChatService/Graph/Trace/Eval 均无引用,可整体删除。
- 实现前规格校正:`ChatVerifierPromptContractTest` 和 `DiagnosisGraphTestSuiteStructureTest` 必须保留旧类型名称的负向字符串断言;这不构成 executable reference。OpenSpec 已收紧为无定义/import/实例化/type-use,允许负向 guard literal。
- Q2:current architecture index 仍列 `agent-orchestration.md` 等为当前真理源,因此必须更新;`mvp/issues/design-notes`、archived issues、历史 eval fixtures 保留时间点/兼容语义,不做大规模重写。
- Q3:`run-interview-demo-check.ps1` 已用 Chat response runId 查询 exact Trace/feedback,最适合扩展 `run.orchestrationTrace` fail-fast 和 summary,不另建重复脚本。
- Q4:E2E 使用唯一 timestamp sessionId;日志记录启动前 byte/time 边界并按 session/run 搜索;DB 所有核心查询带 exact sessionId/runId,另查询错误 ownership count。
- Q6:只有实现、回归、eval、live E2E、日志、DB 和 OpenSpec门禁全部通过后,Issue checkbox 才可完成并移动到 archived。
## User-interview Confirmation
| 问题 | 用户原话 | 状态 | OpenSpec 回写 |
|---|---|---|---|
| Q5 最终验收节奏 | “端到端只在最后阶段全部完成后才验证……日志在log文件夹,项目库有查询数据库的py工具” | 已确认 | proposal |
## Grill-with-docs Result
- Session/Run/Trace 术语保持不变;新增强调 `orchestration_trace` 是 Run 路由摘要,不属于 self-evaluation 或日志。
- StateGraph、Workflow/Node Contract/Chat Integration 属于实现/测试架构术语,不修改业务 glossary。
- 当前文档必须使用 explicit Gatekeeper Node、verified-only Verifier 和 bounded evidence retry;历史设计笔记仍可描述当时 Hook 架构。
- 删除旧闭包是阶段 0 已冻结单轨迁移的自然收尾,不形成新的难逆转权衡,无需 ADR。
## Discover Status
- `devflow/index.md`:命中阶段 0–4 archives。
- 接口影响:L2 内部类型删除;外部 API/DTO/DB/Prompt/状态语义无变化。
- E2E 入口/脚本/日志/DB 工具已定位;真实执行留到 Apply 最后。
- 未解决问题:0。
- Draft 产物:proposal + decisions;尚未生成 design/spec/tasks,尚未删除代码或启动应用。
## Architecture Audit
### Module and evidence map
`ChatController -> ChatService -> ChatDiagnosisGraphRuntime -> DiagnosisRealGraphActionsFactory -> explicit Nodes -> DiagnosisGraphResultMapper -> DiagnosisRun/Trace` 是唯一当前 Chat链。旧 `VerifierInputHook -> ToolTraceSummaryService/VerifierContextHolder` 已从主链断开,删除不改变输入/输出或持久化。阶段 5新增的 demo script assertion只消费 exact Trace `run.orchestrationTrace`,DB/log检查是验收消费者,不成为运行时业务依赖。
| 模块 | 所有权 | 阶段 5动作 |
|---|---|---|
| Graph/Chat runtime | 路由、Node、Run 生命周期 | 不改行为,仅回归 |
| Legacy Hook closure | 旧 Sequential Verifier payload | 整体删除 |
| Current architecture docs | 当前实现真理源 | 更新 StateGraph/verified-only/trace |
| Historical docs/fixtures | 时间点/兼容记录 | 保留,不冒充当前实现 |
| Demo check | live Chat/Trace/feedback executable contract | 增加 exact orchestration trace fail-fast |
| logs/MySQL | live运行证据 | 只读本次 session/run |
| ISS/OpenSpec/devflow | 生命周期与交接 | 所有门禁通过后归档 |
### Lifecycle and failure ownership
- 自动化门禁失败:不启动 live Maven,修复代码/测试/规格后重跑。
- live startup失败:应用未 ready,不执行 demo/DB成功声明,先读启动输出和新日志诊断。
- Chat/Trace/feedback失败:保留 exact response/run证据,ISS保持 active。
- log ERROR:逐条分类;未解释 ERROR阻塞验收。
- DB不一致:以 exact run为准,不能用 API成功掩盖 persistence偏差。
- finally:无论成功失败都停止本轮进程并确认端口,不扩大到未知已有进程。
### Consumer and compatibility audit
- 外部 API/DTO/DB consumer无迁移;demo summary仅加字段。
- Trace UI/eval 对历史 `tool_trace_summary` 的读取保留,旧 fixture不批量迁移。
- current docs消费者将看到新 StateGraph架构;历史链接仍可追溯 old Hook设计。
- Issue move只改变文档位置/index,代码/运行时不依赖该路径。
### Cross-artifact alignment
| 上游 → 下游 | 检查内容 | 状态 |
|---|---|---|
| brief/proposal → proposal | cleanup、current docs、demo、regression、live/log/DB、Issue closure | 已对齐 |
| proposal → design | 删除闭包、current/history边界、顺序、identity、日志/DB、cleanup | 已对齐 |
| design → specs/tasks | 负向 literal例外、E2E字段、exact evidence、进程清理、Issue gate | 已对齐 |
| specs → tasks | 每条 requirement有可执行 cleanup/docs/test/live/log/DB/closure slice | 已对齐 |
### Audit result
审计确认阶段 5不需要新运行时抽象或 DB migration;主要风险来自外部 live状态和证据归属,已通过 unique session/run、log boundary、exact DB queries和process ownership缓解。规格误把负向名称 literal 当 executable reference 的 gap 已修正。接口影响 L2,cross-artifact gap=0,无新 ADR。
## Commit Gate
- schema:spec-driven;proposal/design/2 delta specs/tasks 全部 done,applyRequires=`tasks` 已满足。
- OpenSpec:当前 change strict pass;16 个主 specs strict pass。
- Cross-artifact:4/4 已对齐,gap=0;负向 guard literal例外已写入 proposal/design/spec/tasks。
- Question pool:5 个 evidence-driven 已解决,1 个 user-interview 已由用户原话确认,无未决项。
- Interface impact:L2 internal type removal + demo/docs enhancement;外部协议/DB无变化。
- Preflight:`git diff --check` 通过;尚未删除代码、修改 current docs/script或启动应用。
- 结论:Draft OpenSpec 达到可执行状态,创建 `.committed` 后进入 Apply。
## Apply Progress
### Legacy closure removal
- 已删除 `VerifierInputHook`、`VerifierContextHolder`、`ToolTraceSummaryService` 和 `ToolTraceSummaryServiceTest`,四个路径均不存在。
- `rg` 对 `src/main`、`src/test` 的旧类型扫描仅命中 `ChatVerifierPromptContractTest` 和 `DiagnosisGraphTestSuiteStructureTest` 中的负向守卫字符串;无定义、import、实例化、继承或类型依赖。
- 删除后 focused 回归覆盖 Executor parser、Gatekeeper service/node、VerifiedInput、Verifier、Composer、Fallback、Workflow、Node Contract、Chat integration、Trace、result mapper 和结构契约:14 suites / 82 tests,0 failure、0 error、0 skipped。
- `mvn -q -DskipTests test-compile` 通过;Graph/shared protocol 真理源保留,Spring 当前链路所需类型可完整编译。
### Current docs and demo contract
- architecture index、编排、session/trace、current MVP、evidence pipeline、quality gates、feedback、retrieval 和 eval 文档已切换为 bounded StateGraph、显式 Gatekeeper/Verified Input、verified-only Verifier、有限重试/Fallback 和 Run-owned `orchestration_trace`。
- current-doc scan 对 `SequentialAgent`、旧 Hook/ThreadLocal/service 及旧测试类名为 0 命中;`tool_trace_summary` 仅剩 4 处,均明确标注为旧 Run/fixture 只读兼容,不是当前 Verifier 输入。
- interview demo check 绑定 Chat 返回的 exact runId,校验 Chat/Trace ownership、Run CHAT/SUCCESS、Agent/tool/self-evaluation、Graph trace 六个字段和 feedback success;summary 新增 orchestration version、final node、termination reason、degraded、transition count 和 evidence retry count。
- PowerShell parser 语法检查通过;`InterviewDemoScriptContractTest` 2 tests 通过,覆盖 exact runId URL/response、orchestration fail-fast 和 summary 字段。
### Final deterministic gates
- authoritative/focused regression:39 suites / 157 tests,0 failure、0 error、0 skipped;覆盖三层 Graph、全部 Graph Node/router/trace builder、Chat/Trace/Gatekeeper/Composer、Controller、Repository、schema、feedback/tool recorder 和 demo contract。
- fixed diagnosis eval:12/12 passed,verdict distribution 为 PASS=5、LOW_CONFID=6、REJECT=1;same-baseline diff 无 regression、0 items。
- `mvn -q -DskipTests test-compile` 通过;当前 change strict 通过,16 个主 specs strict 全部通过。
- `git diff --check`、legacy executable refs、current-doc stale refs 和 unexpected schema change 检查全部通过。
- focused 回归日志中的 Graph ERROR/exception stack trace 来自 `ChatServiceGraphIntegrationTest` 对 FAILED/no-answer/unhandled failure 的显式契约用例,Maven exit 0,不是未解释的 live ERROR。
### Live failure diagnosis and correction
- 首次 live identity:sessionId=`iss-011-stage5-20260717140450`,runId=`run-6db680f8-f764-49d1-995f-0e55a4b05a06`。demo contract 在 exact Trace `run.orchestrationTrace=null` 处 fail-fast,未提交 feedback;Run 为 `CHAT/FAILED`,步骤/工具均为 0。
- 新日志根因:`DiagnosisOrchestrationTraceBuilder` 收到类名相同但 classloader identity 不同的 `OrchestrationEvent`,`instanceof` 失败并抛出 `orchestration events contain unsupported value`。这是 Spring Boot DevTools live classloader 才暴露的 Graph state 表示缺陷,单元 JVM 未复现。
- 冲突分类:代码偏离/运行时兼容 bug,OpenSpec 对 non-empty orchestration trace 和 live Maven startup 的要求正确,不修改验收口径。
- RED:新增 portable event map builder 回归,修复前 1 test error;GREEN:Graph state 的 production/test actions 改存 classloader-neutral Map,builder兼容 local record/Map,Node/Workflow assertions 改读 Map。
- 修复后先运行 5-suite Graph/Node/Runtime/Chat integration focused gate,再运行完整 39 suites / 157 tests,全部 0 failure/error/skipped;首次 Maven 进程链已按 ownership 停止,9900 已释放。
- 第二个 live blocker 为外层 Graph `RunnableConfig` 的 resume metadata 被原样传给内层 ReactAgent,触发 `Resume request without a configured checkpoint saver`。回归先证明 nested config 与 outer config 同一且含 `HUMAN_FEEDBACK`,再改为保留 sessionId/runId、剔除 resume/state-update/checkpoint 控制信息的独立配置;5 suites / 50 tests 和随后完整 39 suites / 157 tests 通过,临时 `[DEBUG-ISS011-NODE]` 探针已删除且源码扫描为 0。
- 2026-07-17 的一次长请求在工具执行期间遭遇外部 MySQL 瞬时 `Connection is closed`,留下精确 `RUNNING` 失败尝试;仓库查询工具随后证明数据库恢复且 server `wait_timeout=28800`。该失败未被当作验收通过,失败 Run 保留为真实审计记录。
### Accepted live E2E
- 启动:`mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"`;启动前 9900 空闲,隐藏进程链为 cmd `25080` -> Maven Java `17860` -> app Java `10732`,readiness 后执行固定 payment-timeout demo。
- identity:sessionId=`iss-011-stage5-20260720015557`,runId=`run-808ac38f-3ad0-4462-a6d0-ed50d8686473`;Chat/Trace exact identity 一致,answer 长度 109,feedback request success。
- Trace:Run=`CHAT/SUCCESS`,8 AgentSteps、12 ToolInvocations,self-evaluation 非空;`stategraph-v1`,final node=`fallback`,termination=`fallback_completed`,degraded=`true`,3 transitions,evidence retry count=0。
- Fallback 原因是 Gatekeeper LOW_CONFID 且本轮模型输出缺少 `source_invocation_id`;这是按冻结契约执行的安全降级,answer 非空且未绕过 Gatekeeper,日志/DB 均可审计。
- 最终日志发生跨日 rollover:7 月 20 日 active `application.log`/`chat.log` 全部属于本轮;`application-error.log` 最后写入仍为 7 月 17 日。本轮 `rg " ERROR "` 对 application/chat 为 0,新增 error-file bytes 为 0;session/run、Graph 75964ms、evaluation、exact Trace 和 useful feedback 均有关联日志。
- MySQL:V012 `orchestration_trace` 为 nullable JSON;exact Run answer=109、duration=75964、token=111802、steps=8、tools=12、evaluation len=8822、trace len=438、feedback=useful;JSON 路由与 Trace 完全一致。
- ownership:AgentStep 8、ToolInvocation 12,各自 distinct session/run=1、wrong owner=0;唯一 session 下 wrong run/step/tool 均为 0。
- cleanup:只停止 PID `10732/17860/25080`,最终 9900 已释放,无剩余 owned process。
@@ -0,0 +1,38 @@
# Evidence
## Source And Dependency Evidence
- 旧闭包四个文件已删除;`src/main`/`src/test` 旧类型扫描仅剩两个测试中的负向名称守卫,无 definition/import/instantiation/type dependency。
- 当前真理源保留 `ExecutorEvidenceParser`、`ExecutorGatekeeperService`、`GatekeeperNode`、`VerifiedInputNode`、`VerifierNodeAdapter` 和 `DiagnosisGraphResultMapper`。
- current docs 不再描述 SequentialAgent、Hook Gatekeeper 或 full-trace Verifier;4 处 `tool_trace_summary` 均明确为历史只读兼容。
## Deterministic Evidence
- 删除后 focused:14 suites / 82 tests,0 failure/error/skipped。
- 最终 authoritative/focused:39 suites / 157 tests,0 failure/error/skipped。
- 归档前 `mvn clean` 后重建验证:43 suites / 189 tests,0 failure/error/skipped;额外覆盖 4 个无需外部服务的现存测试类。
- diagnosis eval:12/12 passed;PASS=5、LOW_CONFID=6、REJECT=1;same-baseline diff=0。
- `mvn -q -DskipTests test-compile`、PowerShell parser、OpenSpec current strict、16 main specs strict、`git diff --check`、legacy/current-doc/schema scans 均通过。
## Diagnose Evidence
- DevTools live classloader 使 record `instanceof` 边界失效;portable event map regression 先 RED,Node state 改用 Map 且 builder 兼容 record/Map 后 GREEN。
- outer Graph resume metadata 污染 nested ReactAgent;nested config isolation regression 先 RED,保留 session/run metadata并剔除 resume/state-update/checkpoint 控制信息后 GREEN。
- 两次修复后均重跑 focused 和完整 deterministic gate;临时调试探针为 0。
## Accepted Live Evidence
- sessionId:`iss-011-stage5-20260720015557`
- runId:`run-808ac38f-3ad0-4462-a6d0-ed50d8686473`
- Maven profile:`mvp-demo`;Chat/Trace/feedback script exit 0。
- Run:CHAT/SUCCESS;answer=109 chars;8 steps;12 tools;self-evaluation 非空;feedback=useful。
- Graph:stategraph-v1;planner -> executor -> gatekeeper -> fallback;termination=fallback_completed;degraded=true;evidence retries=0。
- Logs:本轮 active application/chat 中 ERROR=0,error appender 无新写入;session/run、evaluation、Trace、feedback 可关联。
- DB:V012 JSON column 存在;Run fields/JSON 与 API 一致;step/tool wrong owner=0;unique-session wrong rows=0。
- Cleanup:owned process chain 已停止,9900 已释放。
- Archive preflight:临时 `target/iss-011-stage5*` 目录为 0,OpenSpec strict 17/17,current-doc stale=0,schema diff=0,`git diff --check` 通过。
## Known Limits
- 本次真实模型遗漏 `source_invocation_id`,Gatekeeper 按契约降为 LOW_CONFID 并进入安全 Fallback;这是成功且可审计的 degraded Run,不是完整 Composer 正常路径。
- 2026-07-17 外部 MySQL 瞬时断连留下一个 RUNNING 失败尝试;它未计入验收且保留审计,不影响 2026-07-20 exact accepted Run。
@@ -0,0 +1,70 @@
# Chat Diagnosis StateGraph Design Freeze Acceptance
## 结果
已接受。阶段 0 完成设计冻结,没有修改运行时实现。
## 验证
### 静态验证
- 命令:`git diff --check`
- 结果:passed
- 备注:仅有现有 LF/CRLF 提示,无 whitespace error。
- 检查:Git changed/untracked 路径运行时拒绝列表。
- 结果:passed,12 个路径,`src/`、Maven、运行配置、脚本、数据库迁移命中 0。
- 备注:阶段 0 只包含 Issue、glossary、OpenSpec/devflow 和执行记录。
- 检查:proposal → design → specs → tasks 四向对齐。
- 结果:passed,gap=0。
- 备注:状态、路由、Fallback、审计、测试迁移和六阶段边界均闭环。
### 脚本验证
- 命令:`openspec validate chat-diagnosis-stategraph-design-freeze --type change --strict --json`
- 结果:passed,1/1。
- 备注:当前 change 无结构或场景格式问题。
- 命令:`openspec validate --specs --strict --json`
- 结果:passed,11/11。
- 备注:归档前主规格基线未回归。
### 浏览器/人工验证
- 结果:not run。
- 原因:阶段 0 无 UI 或运行行为。
### 未验证
- 单元测试:not run。阶段 0 无代码行为,新增或运行单元测试没有新的验收价值。
- Maven E2E:not run。用户明确要求仅阶段 5 在全部实现完成后统一执行。
- `logs/`:not inspected for stage acceptance;保留到阶段 5。
- 数据库:not queried for stage acceptance;保留到阶段 5。
- 备注:此前改造前 live baseline 仅为 research,不计入本阶段验收。
## 已完成范围
- 冻结最小 Graph State、所有条件边、四类独立计数和终止路径。
- 冻结 verified-input、两类 Fallback、Run 状态与 orchestration trace 边界。
- 冻结 L4 消费者、兼容、迁移、回滚和测试替换策略。
- 创建 ADR-001,并规定阶段 1–5 引用本阶段 archive。
- OpenSpec apply tasks 7/7 完成。
## 已知限制
- 运行时仍为旧 Sequential 编排,这是阶段 0 的有意状态。
- 新设计尚未经过 Graph 编译、Node 单测、Chat 集成或最终 E2E;后续阶段逐项证明。
- 主运行时 specs 暂时仍描述旧行为,只新增设计基线 capability。
## Bug 修复和诊断
- 更正此前错误流程模型:删除未提交的总 change,改为六个独立 sm-flow/change。
- 更正此前错误 E2E 门禁:阶段 0–4 不做 E2E,阶段 5 统一验证。
## 交接
- 下一步:阶段 0 Git commit;提交完成后才能创建阶段 1 change。
- Delta sync:新增 `chat-diagnosis-stategraph-design-freeze` 主 spec,共 6 个 requirements,无现有 spec 修改/删除。
- OpenSpec 归档确认:用户已在目标中明确要求每阶段 archive,并在后续澄清中再次确认,视为已授权。
- OpenSpec 归档结果:已同步主 spec,并归档到 `openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-design-freeze/`。
@@ -0,0 +1,27 @@
# ADR-001: StateGraph owns run-scoped diagnosis control
**状态**:已接受
**日期**:2026-07-17
复杂 Chat 使用 Spring AI Alibaba StateGraph 管理一次 Diagnosis Run 内的顺序、条件边、有限重试、终止和降级;ReactAgent 只执行 Planner、Executor、Verifier、Composer 的语义任务,Gatekeeper、verified-input builder、evidence retry prepare 和固定 Fallback 使用确定性 Java Node。这样可以精确恢复失败位置并形成可审计路径,同时继续复用现有 Agent、工具和证据协议。
Gatekeeper 从 `VerifierInputHook` 的隐式执行迁为显式且唯一的 Graph Node。Verifier 只能消费 Gatekeeper passed checked bindings 投影出的 `verified_executor_output` 和 `verified_evidence`;前置验证失败的 Fallback 不得输出 Executor claim。
编排审计只属于当前 `runId`:有界 `orchestration_events` 压缩后写入 `diagnosis_run.orchestration_trace`,Trace API 只在 `run.orchestrationTrace` 暴露解析对象。它不替代 self evaluation、AgentStep、ToolInvocation 或持久 Graph checkpoint,也不新增 trace 明细表。
## Considered Options
- 继续在 `ChatService` 外层叠加 `for/if`:拒绝,失败恢复位置、循环上限和路由原因仍然隐式。
- 使用 SupervisorAgent:拒绝,当前是固定诊断 Pipeline,不需要动态选择专科 Agent。
- 直接将父 Graph State 交给 `ReactAgent.asNode(...)`:首版拒绝,无法证明 messages、outputKey 和私有执行上下文隔离。
- 长期保留 Sequential/StateGraph feature flag 双轨:拒绝,会形成两个编排真理源并增加安全规则漂移。
## Compatibility And Migration
`/api/chat`、`executor_evidence_v2`、Verifier 和 Composer 输出协议保持不变。Trace API 只增加 run-scoped 字段;数据库只增加 nullable JSON 列,历史 Run 不回填。
实施必须按六个独立 sm-flow/change 依次完成:设计冻结、路由骨架、真实节点、ChatService/Trace 切换、测试体系、清理与最终验收。阶段 1–5 必须读取阶段 0 archive,偏离本 ADR 时通过当阶段 OpenSpec 显式修正。
## Rollback
每阶段使用独立 Git commit,可整体 revert 当前阶段。生产切换后回滚代码时允许保留 nullable `orchestration_trace` 列;不得用配置重新形成长期双轨。
@@ -0,0 +1,20 @@
# Chat Diagnosis StateGraph Design Freeze Brief
## 背景
- 用户目标:将 ISS-011 阶段 0–5 分别作为独立 sm-flow,前一阶段 archive 并 Git commit 后才进入下一阶段。
- 当前问题:复杂 Chat 的跨 Agent 状态机分散在 ChatService、SequentialAgent、VerifierInputHook 和 ThreadLocal 中;实现前需先冻结统一设计。
- 关联 OpenSpec:`openspec/changes/chat-diagnosis-stategraph-design-freeze/`
- devflow 分档:complex
## 范围
- 本次要做:冻结最小 Graph State、完整路由、有限重试、安全 Fallback、run-scoped 审计、L4 接口影响和测试替换边界。
- 本次不做:任何 Java、SQL、Prompt、配置、运行时 spec 或运行行为修改;不运行 Maven E2E。
- 影响区域:ISS-011、glossary、OpenSpec 设计基线、阶段 0 ADR 和后续五阶段交接契约。
## OpenSpec 对齐
- proposal 覆盖状态:已覆盖阶段 0 目标、范围、非目标、验收和风险。
- specs 覆盖状态:新增设计基线 capability,不提前修改运行时 capabilities。
- tasks 覆盖状态:7/7 完成,且仅包含文档、ADR 和静态验证。
@@ -0,0 +1,120 @@
# Chat Diagnosis StateGraph Design Freeze Decisions
## Question Pool
| # | 维度 | 问题 | 模式 | 状态 |
|---|---|---|---|---|
| Q1 | 术语 | Diagnosis Orchestration Trace 与 Diagnosis Trace、self evaluation、Graph checkpoint 的边界是什么? | evidence-driven | 已解决 |
| Q2 | 边界 | ISS-011 是一个总 sm-flow,还是阶段 0–5 各自独立 sm-flow? | user-interview | 已解决 |
| Q3 | 验收 | Maven E2E 在每阶段执行还是只在最终阶段执行? | user-interview | 已解决 |
| Q4 | 验收 | 每阶段是否强制新增并运行单元测试? | user-interview | 已解决 |
| Q5 | 技术 | 锁定的 StateGraph 版本是否提供条件边、config-aware Node/Edge、recursion limit 和 threadId? | evidence-driven | 已解决 |
| Q6 | 架构 | Gatekeeper、Verifier 输入与 Composer 的可信材料边界应复用什么现有契约? | evidence-driven | 已解决 |
| Q7 | 接口 | 阶段 0 本身和最终设计分别属于什么接口影响等级? | evidence-driven | 已解决 |
| Q8 | 归档 | 每阶段是否应在进入下一阶段前独立 archive 和 Git commit? | user-interview | 已解决 |
## Evidence-driven
| 结论 | 证据来源 | 是否已汇报用户 |
|---|---|---|
| Orchestration Trace 是 run 级紧凑编排摘要,不是事件日志、自评估或 Graph checkpoint | `devflow/glossary/CONTEXT.md`、ISS-011 §7.5、session-run-trace-isolation decisions | 已汇报 |
| Gatekeeper 从 Hook 迁为显式 Node 是有意替换旧内部架构,不允许双入口 | executor-gatekeeper-hook decisions、ISS-011 §2.3/§7.2/§14 | 已汇报 |
| Verifier verified evidence 必须来自 Gatekeeper 通过的 checked bindings;Composer 不得读取 raw Executor/tool output | verifier-evidence-reference-fidelity、executor-composer-final-answer 历史档案、ISS-011 §7.4 | 已汇报 |
| 本地 1.1.2.0 API 支持条件边、config-aware Node/Edge、recursion limit、RunnableConfig.threadId 和 ReactAgent.call(input, config) | Maven dependency tree 与本地 JAR `javap` 研究记录 | 已汇报 |
| 阶段 0 是文档/规格交付,无运行时接口变化;其冻结的最终目标涉及状态机、DB 和 Trace API,属于 L4 | sm-flow operating-rules、ISS-011 §10/§14 | 已汇报 |
| 主 OpenSpec 的外层 groundedness round、Hook Gatekeeper 和完整 tool_trace_summary 输入与冻结设计冲突 | `openspec/specs/chat-verifier-agent/spec.md` 等主规格 | 已汇报 |
## User-interview
| 问题原文 | 用户原话 | 确认状态 | OpenSpec 回写 |
|---|---|---|---|
| ISS-011 的阶段关系如何映射 sm-flow? | “iss-011里每个阶段,都是一个sm-flow,而不是将一整个iss-011打包进一个sm-flow中” | 已确认 | 已回写 proposal 范围与非目标 |
| Maven E2E 何时执行? | “端到端只在最后阶段全部完成后才验证” | 已确认 | 已回写 acceptance 与 out of scope |
| 阶段单元测试是否强制? | “每个阶段如果有必要添加单元测试验收的话,就加,没有必要的话就跳过单元测试” | 已确认 | 已回写 acceptance |
| 每阶段如何进入下一阶段? | “每个阶段需要归档完并提交才能进入下一个阶段” | 已确认 | 已回写阶段门禁 |
## OpenSpec Input Context
### devflow index
已命中并读取:
- `session-run-trace-isolation`
- `executor-gatekeeper-hook`
- `verifier-evidence-reference-fidelity`
- `executor-evidence-output-contract`
- `executor-v2-output-contract`
- `executor-verifier-claim-checks`
- `executor-composer-final-answer`
- `mvp-demo-trace-acceptance`
### Historical constraints entering OpenSpec
- `runId` 是运行态和 Trace 的所有权边界。
- AgentStep、ToolInvocation 与 self evaluation 的既有 run 绑定必须保留。
- checked binding 是 verified evidence 的复用来源。
- Composer 的 allowed-material 安全边界不得削弱。
- 旧 Gatekeeper-in-Hook 决策在本设计中被显式废止,不能保留并行入口。
- 主 OpenSpec 的旧运行时要求只能在对应实现阶段修改,阶段 0 不提前宣称代码已切换。
## Key Decisions
### 六个独立交付单元
阶段 0–5 分别使用独立 slug、OpenSpec change、devflow 项目、Archive 和 Git commit。前一 change 归档并提交后才创建下一 change。
### 阶段 0 不承载后续实现
阶段 0 只冻结设计并归档长期上下文。阶段 1–5 的代码、数据库、测试和清理任务不进入本 change 的 tasks。
### 验收分层
阶段 0 没有运行时行为变化,不新增或运行单元测试,也不运行 Maven E2E。验收使用 OpenSpec strict validation、结构检查和文档一致性检查。Maven E2E、`logs/` 与数据库只在阶段 5 统一执行。
### 接口影响
- 当前 change:L1 文档/设计交付,不改变调用方可观察行为。
- 冻结目标:L4,涉及内部状态机语义、Trace API 加字段、数据库契约、迁移和回滚。
- 兼容:保持 `/api/chat` 和 Agent 输出协议;Trace API 为加法式 run 字段。
- 回滚:后续运行时代码可按阶段 Git revert;nullable DB 列可保留,不引入双轨配置。
- 消费者:ChatController/ChatService、Trace DTO/Service/UI/demo、数据库迁移、评测与运维审计。
## Risks Accepted
- 阶段 0 archive 后,冻结设计已完成但运行时仍为旧 Sequential;这是有意的阶段性状态。
- 六个 changes 的一致性由后续每次 context/commit gate 读取并审计本档案保证。
- 不为阶段 0 的纯设计变更新增无行为价值的单元测试。
## Cross-Artifact 对齐检查
| 上游 → 下游 | 检查内容 | 状态 |
|---|---|---|
| brief/prd → proposal | Issue 的阶段 0 目标、范围、非目标和验收已进入 proposal | 已对齐 |
| proposal → 设计产物 | 状态、路由、重试、Fallback、审计、测试迁移和六阶段边界均进入 design | 已对齐 |
| 设计产物 → specs/tasks | L4 影响、安全边界、终止规则和阶段 0 文档工作均进入 spec 或 tasks | 已对齐 |
| specs → tasks | 设计基线、源文档、ADR、严格验证和无运行时变更检查均有可执行任务 | 已对齐 |
Gap:无。
## Architecture Audit
输入链为 `ChatController -> ChatService`,目标处理链为 Run 生命周期 → Graph orchestrator → 白名单 Agent/Java Nodes,输出仍为 `ChatResult + DiagnosisRun + exact RunTrace`。Graph State 只属于当前 run,verified evidence 只来自当前 run 的 passed checked bindings,编排摘要只写当前 Diagnosis Run。旧“Gatekeeper stays in VerifierInputHook”决策被 ISS-011 的显式 Node 有意替代,但旧 Gatekeeper 规则本身继续复用;旧“不新增 trace 主表”和 run ownership 决策保持成立。主要耦合风险是 Hook/Node 双执行、阶段性主 spec 漂移和 Trace 投影重复,design 已分别通过单入口迁移、按阶段修改运行时 specs、唯一 `run.orchestrationTrace` 投影缓解。架构风险已由用户确认的 ISS-011 冻结决策和六阶段交付口径接受,无需返回 grill。
## Commit Preflight
- proposal、design、specs、tasks 完整:通过。
- 所有 user-interview 已确认:通过。
- 未汇报 evidence-driven 结论:无。
- L4 接口影响、消费者、兼容、迁移和回滚:已在 design 独立章节记录。
- OpenSpec strict validation:change 1/1 通过,主 specs 11/11 通过。
- Apply 授权:用户启动目标时要求分阶段完整执行、archive 和 Git commit,后续又确认六个独立 sm-flow;视为当前阶段连续执行授权。
## Apply Evidence
- Task 1.1:ISS-011 §3–14 与 committed design 的 State 字段、条件边、四类计数、Fallback、安全审计和测试迁移语义一致;未发现需回写 OpenSpec 的 drift,未修改运行时代码。
- Task 1.2:将 glossary 中新 Orchestration Trace 的具体 StateGraph/Graph State 表述收敛为“诊断编排/编排上下文快照”;Verifier 的 verified-output/evidence 名称作为跨节点协议保留。
- Task 2.1:创建 ADR-001,记录 StateGraph 控制边界、显式 Gatekeeper、run-scoped trace、替代方案、兼容、迁移和回滚。
- Task 2.2:六个独立 change 的唯一边界已写入 design、design-baseline spec、ADR 和根计划;阶段 1–5 必须引用阶段 0 archive。
- Task 3.1:最终 strict validation 为 change 1/1、主 specs 11/11;proposal/design/spec/tasks 对状态、Fallback、审计、运行时非目标和测试迁移均形成闭环,cross-artifact gap 为 0。
- Task 3.2:Git 枚举 12 个 changed/untracked 路径,运行时拒绝列表命中 0;无 `src/`、Maven、运行配置、脚本或数据库迁移变更。
- Task 3.3:单元测试未运行,因为阶段 0 无代码行为;Maven E2E、`logs/` 和数据库核验未运行,因为用户明确要求只在阶段 5 全部实现后统一执行。改造前 research baseline 不计入阶段验收。
@@ -0,0 +1,30 @@
# Chat Diagnosis StateGraph Design Freeze Evidence
## 证据
| 来源 | 证据 | 结论 | 是否已汇报 |
|---|---|---|---|
| ISS-011 §3–14 | 已定义目标流程、26 个 State 字段、有限回边、Fallback、Trace 和测试策略 | 可形成无开放分支的阶段 0 设计基线 | 是 |
| `ChatService` / Controller 引用核查 | 生产链为 Controller → ChatService → SequentialAgent,细粒度状态散布在 ChatService | StateGraph 应只接管 Run 内控制,ChatService 保留生命周期 | 是 |
| `VerifierInputHook` / `VerifierContextHolder` 引用核查 | Gatekeeper、工具摘要和解析结果通过 Hook/ThreadLocal 隐式传播 | Gatekeeper 与 verified-input 必须迁为显式 Node,不能保留双入口 | 是 |
| executor-gatekeeper-hook 历史 decisions | 旧阶段曾决定 Gatekeeper 留在 Hook 且不重试 | 本设计有意替换入口,但继续复用确定性规则 | 是 |
| session-run-trace-isolation 历史 decisions | runId 是执行/Trace 所有权边界,AgentStep/ToolInvocation 已承载明细 | orchestration trace 只附着当前 Run,不新增明细主表 | 是 |
| Verifier/Composer 历史档案 | checked bindings 和 allowed material 是可信边界 | Builder 不重复验真,Fallback 不泄漏 raw claim | 是 |
| 本地 Graph Core 1.1.2.0 JAR | 已验证条件边、config-aware Node/Edge、recursion limit、threadId 与 Agent call API | 阶段 1 可基于锁定签名实现,不采用网上漂移示例 | 是 |
| 主 OpenSpec 核查 | 旧 spec 仍描述 Hook Gatekeeper、完整 tool trace 和外层 groundedness round | 阶段 0 不提前改运行时 spec,后续对应实现阶段再修改 | 是 |
| 用户原话 | 六个独立 sm-flow;仅阶段 5 E2E;单测按必要性 | 阶段门禁与测试口径已固定 | 是 |
## Evidence-driven 结论
- 结论:阶段 0 运行时影响为 L1,但冻结目标为 L4。
- 证据:本阶段 changed paths 无 `src/`/SQL/配置;最终目标改变状态机、Trace API 和 DB。
- 风险:设计完成不等于运行时完成。
- 用户确认:已确认分阶段交付。
- 结论:旧 Gatekeeper-in-Hook 决策应被显式 Node 有意替代。
- 证据:当前 Hook 引用与 ISS-011 条件边要求冲突。
- 风险:迁移期双执行。
- 用户确认:ISS-011 冻结决策已确认。
- 结论:阶段 0 不需要单元测试或 E2E。
- 证据:Git 运行时拒绝列表命中 0,OpenSpec strict 和文档一致性已覆盖本阶段可观察产物。
- 风险:运行时正确性仍未证明,将由阶段 1–5 测试和最终 E2E 证明。
- 用户确认:已确认。
@@ -0,0 +1,42 @@
# Chat Diagnosis StateGraph Real Nodes 验收
## 结果
已接受。OpenSpec tasks 27/27 完成;真实 Nodes 可构造和测试,旧生产 Sequential 路径保持可用,阶段 2 未切换生产入口。
## 静态验证
- `openspec validate chat-diagnosis-stategraph-real-nodes --strict`:通过。
- `openspec validate --specs --strict`:13 passed,0 failed。
- `git diff --check`:通过;只有 LF/CRLF 转换提示,无 whitespace error。
- 源码/引用检查:`ChatService` Graph 引用 0;新增源码 forbidden refs 0;protocol 反向依赖 0;DB/Trace/Prompt diff 0。
- 装配对齐:无核心 TODO/FIXME/placeholder;Graph factory 继续拥有所有 retry counter 与 Planner reset。
## 脚本验证
- 新 protocol/Node/CompiledGraph/Router/Trace focused suite:通过。
- `mvn -q "-Dtest=ChatServiceSequentialAgentTest,VerifierInputHookTest,ExecutorGatekeeperServiceTest" test`:通过。
- 两组测试合计:100 tests,0 failures,0 errors,0 skipped。
- `mvn -q -DskipTests test-compile`:通过。
## 浏览器/人工验证
- 未运行。阶段 2 不改变用户入口或 UI,没有独立人工验收价值。
## 未验证
- Maven live E2E、`logs/` 日志和 `scripts/query_mysql.py` 数据库核验未运行。
- 原因:用户明确要求只在阶段 5 全部实现后统一做最终 E2E;阶段 2 尚未切换生产入口。
- 风险:当前验收只证明组件/Graph 契约与旧路径回归,不证明真实生产装配、模型、DB 和 Trace 全链路。
## 已完成范围
- 中立共享 protocol 与旧路径行为保持型委托。
- 四个 Agent adapters、显式 Gatekeeper、可信投影、关键证据补查、Composer 与两类 Fallback。
- 真实 CompiledGraph 装配、critical gap 路由修复和完整 focused test 证据。
## 交接
- 下一步:归档本 change、独立提交阶段 2,然后启动阶段 3 `chat-diagnosis-stategraph-chatservice-cutover`。
- OpenSpec 归档:已同步主 specs,并归档到 `openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-real-nodes/`。
- 归档授权:用户已明确“直接实现吧,不用找我授权了”,授权后续阶段在门禁通过后直接归档和提交。
@@ -0,0 +1,21 @@
# Chat Diagnosis StateGraph Real Nodes Brief
## 背景
- 用户目标:将 ISS-011 阶段 2 作为独立 sm-flow,接入真实 Agent/Java Nodes,并在归档、验收和独立 Git 提交后才进入阶段 3。
- 当前问题:阶段 1 只有 Fake Node 路由骨架;Executor/Verifier/Composer 协议逻辑分散在 Hook 与 ChatService,真实 Graph 尚不能安全调用 Agent、Gatekeeper 或投影可信材料。
- 关联 OpenSpec:`openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-real-nodes/`
- devflow 分档:complex
## 范围
- 本次要做:共享无状态 protocol 组件;Planner/Executor/Verifier/Composer adapters;显式 Gatekeeper、Verified Input、Evidence Retry、Fallback Nodes;真实 CompiledGraph 装配和 focused tests。
- 本次不做:不切换 ChatService 生产入口,不修改 DB、Trace API、Prompt 契约,不删除旧 Sequential/Hook,不运行 live E2E。
- 影响区域:`diagnosis.protocol`、`graph.diagnosis`、`VerifierInputHook`、`ChatService` 共享逻辑委托及对应测试。
## OpenSpec 对齐
- proposal 覆盖状态:已覆盖。
- design 覆盖状态:已覆盖;接口影响为 L2,生产切换明确延期到阶段 3。
- specs 覆盖状态:已覆盖真实 Nodes 安全边界,并修正 critical evidence gap 路由条件。
- tasks 覆盖状态:27/27 完成。
@@ -0,0 +1,218 @@
# Chat Diagnosis StateGraph Real Nodes Decisions
## Question Pool
| # | 维度 | 问题 | 模式 | 状态 |
|---|---|---|---|---|
| Q1 | 边界 | 阶段 2 是否切换 ChatService/DB/Trace? | user-interview(六阶段已确认) | 已解决 |
| Q2 | 复用 | 新 Nodes 如何避免复制 Hook/ChatService 解析与安全渲染? | evidence-driven | 已解决 |
| Q3 | Agent | Adapter 如何调用真实 ReactAgent 并保留 Hook/ToolCallback/run config? | evidence-driven | 已解决 |
| Q4 | Gatekeeper | 显式 Node 如何保证当前 run、单次调用和 fail-closed 标准化? | evidence-driven | 已解决 |
| Q5 | 安全 | passed checked bindings 如何投影为 verified output/evidence? | evidence-driven | 已解决 |
| Q6 | 补证据 | 哪些 facts 可触发 evidence retry,completed queries 如何表达? | evidence-driven | 已解决 |
| Q7 | Fallback | 前置验证失败与 Composer 后置失败可分别使用哪些材料? | evidence-driven | 已解决 |
| Q8 | Prompt | 阶段 2 是否立即修改共享 Verifier Prompt? | evidence-driven | 已解决 |
| Q9 | 验收 | 阶段 2 是否需要单元/集成测试与 E2E? | evidence-driven + user rule | 已解决 |
## Evidence-driven
| 结论 | 证据来源 | 是否已汇报用户 |
|---|---|---|
| ReactAgent.call(input, config) 返回 AssistantMessage,现有 AgentLoggingHook 从 config metadata 读取 sessionId/runId | 本地 1.1.2.0 javap、AgentLoggingHook | 已汇报 |
| Executor parser 当前在 VerifierInputHook,Verifier/Composer parser 与 renderer 当前在 ChatService | 定向源码阅读和 references | 已汇报 |
| checked_bindings 提供 claim_id/tool_name/source_invocation_id/raw_path/matched_text/status | ExecutorGatekeeperService | 已汇报 |
| Gatekeeper.validateRun 使用 run-scoped ToolInvocation repository | ExecutorGatekeeperService tests/code | 已汇报 |
| Graph path不能注册旧 VerifierInputHook,否则 Gatekeeper 会双执行且输入包含 full tool trace | 阶段 0 ADR、Hook 源码 | 已汇报 |
| evidence retry 只允许 critical no_evidence/indirect_support | 阶段 0 design/ISS-011 | 已汇报 |
| 阶段 1 router 未检查 is_critical,属于实现偏离 | Router 源码与阶段 0 baseline 对照 | 已汇报 |
| 共享 Prompt 暂时仍服务旧 Sequential path,本阶段直接修改会提前破坏生产 payload | ChatService agent builder + VerifierInputHook + prompt | 已汇报 |
| Node 输入/安全边界新增且旧解析会重构,单元和 focused regression 必要;E2E 不必要 | 阶段边界和用户规则 | 已汇报 |
## User-interview
| 问题原文 | 用户原话 | 确认状态 | OpenSpec 回写 |
|---|---|---|---|
| 阶段 2 是否独立 sm-flow? | “iss-011里每个阶段,都是一个sm-flow” | 已确认 | 独立 change |
| 是否可提前切生产入口? | “每个阶段需要归档完并提交才能进入下一个阶段” | 已确认不可提前阶段 3 | Out of Scope |
| 阶段 2 是否执行 E2E? | “端到端只在最后阶段全部完成后才验证” | 已确认不执行 | Acceptance |
| 是否添加单元测试? | “如果有必要添加单元测试验收的话,就加” | 已确认规则;本阶段判定必要 | Acceptance |
## Context And Handoff
- 阶段 0 archive/commit:design baseline / `581daff`
- 阶段 1 archive/commit:routing skeleton / `42ba204`
- 当前 change:`chat-diagnosis-stategraph-real-nodes`
- 后续 change:`chat-diagnosis-stategraph-chatservice-cutover`,只能在本阶段 archive + commit 后创建。
## Technical Decisions
### Shared protocol components
- 抽取 Executor、Verifier、Composer 解析器,旧 Hook/ChatService 委托新组件。
- 抽取 Composer safe input builder / renderer,旧 ChatService 保持相同输出。
- JSON sanitization 只存在一份共享实现,不在每个 Node copy。
- 抽取过程是行为保持 refactor;现有 focused tests 是回归门禁。
### Agent adapter boundary
- `DiagnosisAgentInvoker` 是最小 port:`invoke(String, RunnableConfig) -> String`。
- `ReactAgentDiagnosisInvoker` 只包装 `ReactAgent.call(...).getText()`。
- Planner/Executor/Verifier/Composer adapters 各自拥有白名单 input projector、parser 和 status event。
- tests 使用 fake invoker,另有 ReactAgent wrapper test。
### Legacy Hook coexistence
- 旧 Sequential path 在阶段 3 前仍通过 VerifierInputHook 执行 Gatekeeper。
- Graph Verifier Agent 不注册该 Hook;显式 Gatekeeper Node 是 Graph 中唯一 validation 入口。
- Parser/enricher 可共享,但 Hook 的 legacy payload/prompt 暂不改变。
- 阶段 3 切换生产入口并同步 Verifier Prompt;阶段 5 删除旧隐式结构。
### Gatekeeper normalization
- raw pass → PASS。
- raw fail + severity low_confid → LOW_CONFID。
- raw fail + severity reject → REJECT。
- 缺失、unknown、异常或不一致 → REJECT。
- verified_binding_count 只统计 checked_bindings.status=pass。
### Verified projection
- 通过项按 claim_id + source_invocation_id + tool_name + raw_path 与原 binding 精确匹配。
- verified_executor_output 只包含 answer_version 与至少一条 passed binding 的 filtered claims;合法零 claim 保持空列表。
- verified_evidence 只含 claim_id/source_invocation_id/tool_name/raw_path/matched_text。
- hypotheses、失败 binding、未引用 ToolInvocation、raw Executor output 不投影。
### Evidence retry
- extractor 要求 `is_critical=true` 且 verification 为 no_evidence/indirect_support。
- gap 字段:claim_id(从 `claim-id: text` 提取或空)、fact、verification、reason。
- completed queries 由 verified evidence 的 tool/invocation/path 去重生成。
- prior verified output/evidence 原样只读进入 retry_context。
- 约束固定:max_retry=1、do_not_repeat_successful_queries、only_execute_incremental_queries、preserve_prior_verified_claims。
- Executor 输入声明“增量执行、完整输出”,Java 不合并 claims。
### Interface impact
- 等级:L2 internal interface;共享 parser 委托保持旧行为。
- 新消费者:阶段 3 Graph orchestrator/Agent factory。
- 当前外部 API/DB/生产路由:无变化。
- 回滚:revert 本阶段提交;旧 Sequential path仍完整。
- 有意修正:Router 只允许 critical gap,属于对冻结基线的代码修复。
## Risks Accepted
- 真实模型尚未执行;Node contract 通过 fake invoker/real service tests证明,阶段 5 才 E2E。
- Prompt 输入说明与 Graph input 的最终同步推迟到阶段 3,避免当前旧生产 Hook 提前不兼容。
- 旧 Hook 暂时仍存在,但不进入 Graph action graph;阶段 5 必须删除。
## Architecture Audit
### Module and ownership map
`Diagnosis Context + RunnableConfig` → Agent adapters / deterministic Nodes → owned Graph State fields → `DiagnosisGraphRouter` → next Node or END → `final_answer + orchestration_events`。阶段 2 只提供这条可构造链,阶段 3 才由 `ChatService` 创建 Run、构造 Agent 实例并调用 Graph。
| 模块 | 数据所有权 | 允许依赖 |
|---|---|---|
| `diagnosis.protocol` | JSON contract、纯解析结果、安全 input/rendering | Jackson 与纯 DTO;不依赖 Graph/Hook/ChatService/ThreadLocal |
| `graph.diagnosis` Agent adapters | 白名单输入、Agent attempt status/event | protocol、注入的 invoker、Graph API |
| Gatekeeper Node | raw Gatekeeper result、normalized status/count | ExecutorGatekeeperService、当前 run config |
| Verified Input / Retry Prepare | verified projection、critical gaps、retry context | raw result 的只读投影与 protocol DTO |
| Router / Factory | 条件边、有限计数、调用次序 | Graph State;不解析 Agent raw output |
| legacy Hook / ChatService | 阶段 3 前的生产 Sequential 流程 | 只委托 protocol;不得消费真实 Graph action set |
### Lifecycle and coupling audit
Graph State 和 events 都是 invocation-scoped,runId 只从当前 RunnableConfig 获取,Nodes 不持有跨 Run 可变状态。旧路径与 Graph 路径阶段性共享的只有无状态 protocol 组件和 Gatekeeper service,不共享 ThreadLocal 或 Agent output。ReactAgent/invoker 必须构造注入,阶段 2 不复制 ChatService Prompt/Agent factory。唯一有意的阶段性耦合是 legacy consumers 改为委托 protocol,这由旧 focused tests 和完整 revert 保护。阶段 3 前生产入口、DB、Trace、Prompt 均保持隔离。
### Cross-artifact alignment
| 对齐链 | 结果 | 证据 |
|---|---|---|
| brief/proposal 目标、范围、非目标 → proposal | 已对齐 | 独立阶段边界、真实 Nodes、无生产切换均明确 |
| proposal 承诺与约束 → design | 已对齐 | invoker、共享组件、Gatekeeper、投影、retry、Fallback、L2 均有决策 |
| design 架构/接口结论 → specs/tasks | 已对齐 | 中立 protocol 包、构造注入、fail-closed 和生产隔离均有任务/行为 |
| specs 可观察行为 → tasks 可执行切片 | 已对齐 | 每个 requirement 至少由一个实现任务和一个测试/验收任务覆盖 |
### Audit result
审计发现共享协议组件包所有权与 Agent 实例装配边界需要显式化,已回写 design/tasks。未发现状态字段、路由计数、run 生命周期或阶段边界的新冲突。接口影响维持 L2,消费者都在本 change 与下一阶段明确范围内。剩余风险是共享抽取的旧行为漂移和 binding 投影泄漏,均有 focused regression 与 mixed-binding tests。架构风险可接受,无未解决问题。
## Commit Gate
- proposal/design/specs/tasks:全部存在,OpenSpec status `isComplete=true`。
- 当前 change strict validation:通过。
- 主 specs strict validation:13 passed,0 failed。
- 规格结构:9 requirements、22 scenarios;tasks:27 个 checkbox 切片。
- Cross-artifact:4/4 已对齐,gap=0。
- 接口影响:L2,已在 design 独立章节记录消费者、兼容和回滚。
- Question pool:所有 evidence-driven 已汇报;所有 user-interview 已确认;无未决项。
- Preflight scope:`git diff --check` 通过;Commit checkpoint 未修改业务代码。
- 结论:Draft OpenSpec 达到可执行状态,创建 `.committed`。
## Apply Authorization
- 用户原话:“直接实现吧,不用找我授权了”。
- 解释:阶段 2–5 后续 checkpoint 可在前置门禁通过后直接继续,不再因 Apply 或 Archive 授权暂停。
- 不扩大范围:每阶段仍须独立 OpenSpec、验收归档、Git commit;阶段 0–4 不做 E2E,阶段 5 才统一执行。
## Pre-apply Research
### Reference implementations read
- `src/main/java/com/superbiz/agent/hook/VerifierInputHook.java`:Executor JSON sanitization、parse status、tool-name normalization、唯一 invocation 回填与 legacy Gatekeeper payload。
- `src/main/java/com/superbiz/agent/service/ChatService.java`:Verifier claim/fact parser、Gatekeeper ceiling、Composer allowed-material builder、Composer parser/safe renderer、旧 retry context。
- `src/main/java/com/superbiz/agent/service/ExecutorGatekeeperService.java`:`validateRun`、severity normalization source、checked binding 与 matched_text 结构。
- `src/main/java/com/superbiz/agent/graph/diagnosis/DiagnosisGraphFactory.java`:config-aware action ports、technical retry/evidence retry counter 所有权和 recursion limit。
- `src/main/java/com/superbiz/agent/graph/diagnosis/OrchestrationEvent.java` 与 `src/test/java/com/superbiz/agent/graph/diagnosis/ScriptedDiagnosisGraphActions.java`:每 attempt 单 event 形态。
- `src/main/java/com/superbiz/agent/hook/AgentLoggingHook.java`:RunnableConfig metadata 中 sessionId/runId 的审计读取方式。
- `src/test/java/com/superbiz/agent/hook/VerifierInputHookTest.java`、`ChatServiceSequentialAgentTest.java`、`DiagnosisGraphRoutingTest.java`:JUnit 5、Mockito 边界替身和 CompiledGraph observable behavior 测试风格。
- 本地 `1.1.2.0` JAR `javap`:`ReactAgent.call(String,RunnableConfig)` 与 `RunnableConfig.threadId/metadata` 公共 API。
### Technology inventory
| 类别 | 项目标准 / 本阶段使用 |
|---|---|
| JSON contract | Jackson `ObjectMapper`、`LinkedHashMap` 保持稳定字段顺序;共享 sanitization 只存在一份 |
| Graph Node | `AsyncNodeActionWithConfig` 返回 `CompletableFuture<Map<String,Object>>`;state 默认 Replace、events Append |
| Agent boundary | 构造注入的 `DiagnosisAgentInvoker`;ReactAgent wrapper 透传同一个 RunnableConfig |
| Run scope | `config.metadata("runId")` 为 Gatekeeper 唯一运行边界;缺失即 fail closed |
| Error handling | 合法 contract/invalid output/temporary/permanent 分离;unknown exception 不推断为可重试 |
| Tests | JUnit 5;只在 ReactAgent、repository 等系统边界使用 fake/mock;Node/CompiledGraph 走公开 action/graph 接口 |
| Request/response | 本阶段不改 Controller/DTO/API,不适用 |
| MQ/Consumer | 本阶段不涉及,不适用 |
| DB/Trace/Prompt | 本阶段禁止修改,阶段 3 处理 |
### Reuse and new infrastructure
- 新建中立 `com.superbiz.agent.diagnosis.protocol`:`JsonPayloadSupport`、Executor/Verifier/Composer protocol、safe input/rendering、`EvidenceGapExtractor`;不得依赖 Graph/Hook/ChatService/ThreadLocal。
- 新建 `graph.diagnosis` Node 层:invoker wrapper、failure classifier、四个 Agent adapters、Gatekeeper、Verified Input、Retry Prepare、Fallback 和 action assembly。
- 不新建 Prompt/Agent factory;阶段 3 通过构造注入已有 ReactAgent 实例。
- 不新增 Maven dependency、数据库迁移、配置项或生产 consumer。
### Pre-apply conclusion
参考实现、API 签名、异常和测试标准已足以指导实现;未发现 devflow/OpenSpec 冲突。进入 TDD tracer bullet,先锁定共享 Executor parser 的合法 no-evidence 与 malformed 行为。
## Apply Completion And Verification
### Assembly alignment
- 首个共享 protocol 模块与最终真实 Graph action assembly 均逐项对照 design/specs:invoker 显式透传 RunnableConfig,Gatekeeper 单次 fail-closed,Verified Input 精确投影,Verifier/Composer 固定输入重试,critical evidence retry 与两类 Fallback 边界全部落地。
- `DiagnosisGraphFactory` 继续独占技术重试计数、evidence retry 计数和 Planner mode/reset;Node 不重复拥有编排计数。
- 新增源码中无 `ThreadLocal`、`VerifierInputHook`、`tool_trace_summary`、`raw_executor`、TODO 或 FIXME;protocol 包无 Graph/Hook/ChatService 反向依赖。
- `ChatService` 无 Diagnosis Graph/CompiledGraph 引用,生产切换保持在阶段 3;DB migration、Trace DTO/entity/repository 和 prompts 均无 diff。
### Automated verification
- 新 protocol/Node/真实 CompiledGraph/Router/Trace focused suite:通过。
- 旧 `ChatServiceSequentialAgentTest`、`VerifierInputHookTest`、`ExecutorGatekeeperServiceTest` 回归:通过。
- 合计 100 tests,0 failures,0 errors,0 skipped。
- `mvn -q -DskipTests test-compile`:通过。
- `openspec validate chat-diagnosis-stategraph-real-nodes --strict`:通过。
- `openspec validate --specs --strict`:13 passed,0 failed。
- `git diff --check`:通过;仅报告 Git 既有 LF/CRLF 转换提示,无 whitespace error。
### Deferred final verification
- 阶段 2 按用户确认不运行 Maven live E2E,不启动应用,不检查 `logs/`,不查询数据库。
- 上述端到端、日志和 `scripts/query_mysql.py` 数据库核验统一保留到阶段 5 全部实现完成后执行。
@@ -0,0 +1,25 @@
# Chat Diagnosis StateGraph Real Nodes Evidence
## 证据
| 来源 | 证据 | 结论 | 是否已汇报 |
|---|---|---|---|
| `VerifierInputHook.java`、`ChatService.java` | Executor/Verifier/Composer 解析与安全渲染原实现 | 抽取到中立 protocol 并让旧路径委托,避免双真理源 | 是 |
| `ExecutorGatekeeperService.java` | `validateRun` 与 checked binding/matched_text 契约 | Graph Gatekeeper 必须按当前 runId 单次调用并 fail closed | 是 |
| 本地 Graph/ReactAgent 1.1.2.0 API | `ReactAgent.call(String,RunnableConfig)` 与 config-aware Graph action | 最小 invoker 可精确透传输入和当前 Run metadata | 是 |
| 阶段 0/1 OpenSpec archives | 路由、计数、状态与安全边界冻结 | 阶段 2 不改变 Graph counter 所有权或生产入口 | 是 |
| 新 protocol/Node/CompiledGraph tests | PASS、REJECT、LOW_CONFID、critical retry、固定输入重试和安全 fallback | 真实 Nodes 的可观察路径和材料边界已覆盖 | 是 |
| 旧 Sequential/Hook/Gatekeeper tests | 共享抽取后的旧路径回归 | 阶段 2 未破坏当前生产控制流 | 是 |
## Evidence-driven 结论
- Graph Verifier 不得注册旧 `VerifierInputHook`,否则会双执行 Gatekeeper 并泄漏完整 tool trace。
- Verified Input 必须用 claim/invocation/tool/path 精确关联 passed binding;不能复刻 Gatekeeper 判断或读取未引用工具结果。
- Router 与 Retry Prepare 必须共用 critical-gap 提取规则:仅 `is_critical=true` 的 `no_evidence`/`indirect_support`。
- 技术重试输入必须字节一致,且不得重跑前序 Node;计数仍由 Graph factory 统一拥有。
- 当前 ChatService 无 Graph 引用,DB/Trace/Prompt 无 diff,满足阶段 2 的生产隔离要求。
## 风险与后续证据
- 本阶段使用 fake invoker 和 focused tests,不证明真实模型/外部基础设施联通;阶段 5 最终 E2E 统一补证。
- Prompt 输入说明、生产装配、Run/Trace 持久化属于阶段 3,不能提前从阶段 2 证据推断已完成。
@@ -0,0 +1,69 @@
# Chat Diagnosis StateGraph Routing Skeleton Acceptance
## 结果
已接受。阶段 1 完成未接生产入口的 Graph 骨架和 Fake Node 路由体系。
## 验证
### 静态验证
- 命令:`git diff --check`
- 结果:passed。
- 检查:production refs、forbidden deps、23 changed paths 白名单。
- 结果:passed,outside refs=0,forbidden refs=0,out-of-scope=0。
### 脚本验证
- 命令:`mvn -q "-Dtest=DiagnosisGraphRoutingTest,DiagnosisOrchestrationTraceBuilderTest" test`
- 结果:passed,35 tests(29 routing + 6 trace),0 failure/error。
- 覆盖:正常、三类技术 retry、Executor 不重试、Gatekeeper、evidence retry、Composer、unknown fail-closed、threadId、Append events 和 trace。
- 命令:`mvn -q "-DskipTests" test`
- 结果:passed。
- 覆盖:main/test compilation。
- 命令:`openspec validate chat-diagnosis-stategraph-routing-skeleton --type change --strict --json`
- 结果:passed,1/1。
- 命令:`openspec validate --specs --strict --json`
- 结果:passed,12/12(归档前)。
### 浏览器/人工验证
- 结果:not run。
- 原因:无 UI 或生产入口变化。
### 未验证
- Maven E2E:not run,用户要求仅阶段 5 全部实现后统一执行。
- `logs/`:not inspected,保留到阶段 5。
- 数据库:not queried,保留到阶段 5。
- 真实模型/Agent:not invoked,属于阶段 2。
## 已完成范围
- 显式 Graph Core direct dependency。
- 26 state keys、typed status、topology/route/reason constants。
- 八个 config-aware action ports 和真实 CompiledGraph factory。
- Planner/Verifier/Composer 技术计数、一次 evidence retry、recursion limit 32。
- fail-closed router、events Append 和 trace builder。
- 35 个 Fake Node/trace tests。
## 已知限制
- Skeleton 没有生产消费者,阶段 3 才切换 ChatService。
- Fake Node 只证明控制流,不证明真实 Agent JSON、Gatekeeper 或安全输入映射。
- orchestration trace 尚未持久化或通过 API 暴露。
## Bug 修复和诊断
- 架构审计发现并修正 Graph Core 传递依赖所有权,改为 BOM 管理的直接依赖。
- 自查补齐全部正常节点 threadId、Verifier REJECT 和缺失 effective verdict 场景。
## 交接
- 下一步:阶段 1 Git commit;完成后才能创建阶段 2 change。
- Delta sync:新增 `chat-diagnosis-stategraph-routing-skeleton` 主 spec,共 6 requirements。
- OpenSpec 归档确认:用户已要求每阶段 archive,授权已存在。
- OpenSpec 归档结果:已同步主 spec,并归档到 `openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-routing-skeleton/`。
@@ -0,0 +1,21 @@
# Chat Diagnosis StateGraph Routing Skeleton Brief
## 背景
- 用户目标:ISS-011 阶段 1 独立完成 Graph 骨架和 Fake Node 路由测试,archive/commit 后才能进入真实 Node 阶段。
- 当前问题:阶段 0 只有设计基线,仓库此前没有可编译 StateGraph 或条件边验证。
- 关联 OpenSpec:`openspec/changes/chat-diagnosis-stategraph-routing-skeleton/`
- devflow 分档:complex
- 前置基线:`openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-design-freeze/`
## 范围
- 本次要做:Graph Core 直接依赖、状态/枚举、config-aware action ports、deterministic router、Graph factory、有限计数、events/trace builder、Fake Node tests。
- 本次不做:真实 Agent/Gatekeeper、ChatService、Hook、DB、Trace API、旧实现清理、Maven E2E。
- 影响区域:`pom.xml`、`com.superbiz.agent.graph.diagnosis`、对应 test 包。
## OpenSpec 对齐
- proposal 覆盖状态:已覆盖 skeleton-only 目标、L2 边界、测试与 E2E 非目标。
- specs 覆盖状态:6 个 requirements 覆盖编译、Planner、Executor/Gatekeeper、Verifier、Composer、trace。
- tasks 覆盖状态:12/12 完成。
@@ -0,0 +1,112 @@
# Chat Diagnosis StateGraph Routing Skeleton Decisions
## Question Pool
| # | 维度 | 问题 | 模式 | 状态 |
|---|---|---|---|---|
| Q1 | 术语 | 阶段 1 的 Graph State、event、trace 含义从哪里继承? | evidence-driven | 已解决 |
| Q2 | 边界 | 阶段 1 是否接入 ChatService 或真实 Agent/Service? | evidence-driven | 已解决 |
| Q3 | 技术 | 锁定 1.1.2.0 是否支持所需 Node/Edge、策略和 recursion limit? | evidence-driven | 已解决 |
| Q4 | 技术 | Node action port 是否需要 RunnableConfig? | evidence-driven | 已解决 |
| Q5 | 验收 | 阶段 1 是否有必要添加单元测试? | evidence-driven + user rule | 已解决 |
| Q6 | 接口 | 新骨架的接口影响等级与消费者是什么? | evidence-driven | 已解决 |
| Q7 | 循环 | recursion limit 应取多少,是否替代业务计数? | evidence-driven | 已解决 |
| Q8 | 阶段 | 是否在本 change 实现真实 nodes、DB、Trace API 或旧代码清理? | user-interview(已由六阶段口径确认) | 已解决 |
## Evidence-driven
| 结论 | 证据来源 | 是否已汇报用户 |
|---|---|---|
| 阶段 1 必须完整继承阶段 0 baseline | archived design/spec/ADR | 已汇报 |
| 本阶段只做 Fake Node 骨架,不接真实模型 | ISS-011 阶段 1/2 边界 | 已汇报 |
| 1.1.2.0 支持 AsyncNodeActionWithConfig、AsyncEdgeActionWithConfig、conditional edges、Replace/Append、recursionLimit 和 threadId | 本地 JAR `javap` / `javap -c` | 已汇报 |
| AppendStrategy 将 list/collection 追加为有序列表,可用于 node terminal events | 本地 AppendStrategy bytecode | 已汇报 |
| 路由是新增行为且分支多,单元测试有必要 | 阶段 1 完成标准与用户“必要则加”规则 | 已汇报 |
| 最坏合法路径少于 20 次 Node 执行,limit 32 有安全余量 | 冻结路由矩阵的路径计数 | 已汇报 |
| 当前仓库没有 Graph 包或实现 | `rg` / package 目录核查 | 已汇报 |
| 生产代码将直接 import Graph Core,必须从传递依赖提升为 BOM 管理的直接依赖 | pom 与 dependency tree / 架构审计 | 已汇报 |
## User-interview
| 问题原文 | 用户原话 | 确认状态 | OpenSpec 回写 |
|---|---|---|---|
| 阶段 1 是否应独立执行 sm-flow? | “iss-011里每个阶段,都是一个sm-flow” | 已确认 | 本 change 独立边界 |
| 阶段 1 是否执行 E2E? | “端到端只在最后阶段全部完成后才验证” | 已确认 | Out of Scope / Acceptance |
| 阶段 1 是否添加单元测试? | “如果有必要添加单元测试验收的话,就加” | 已确认规则;本阶段判定必要 | Acceptance |
| 是否可提前实现阶段 2–5? | “每个阶段需要归档完并提交才能进入下一个阶段” | 已确认不可提前 | Out of Scope |
## Context And Handoff
- 前置 archive:`openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-design-freeze/`
- 前置 commit:`581daff`
- 当前 change:`chat-diagnosis-stategraph-routing-skeleton`
- 后续 change:`chat-diagnosis-stategraph-real-nodes`,只能在本阶段 archive + commit 后创建。
## Technical Decisions
### Package and API
- 新包:`com.superbiz.agent.graph.diagnosis`。
- `pom.xml` 显式声明 Graph Core,版本继续由现有 BOM 管理。
- Graph 工厂接收 config-aware Node action ports,不依赖 Spring Bean 或真实 Agent。
- Fake actions 只放 `src/test`。
- 状态读取集中在 typed helper/router,避免各 edge 复制字符串解析。
### Retry ownership
- Planner/Verifier/Composer Node wrapper 根据进入节点前的上次技术失败状态增加各自 retry count。
- 第一次失败时 count=0,允许 self-loop;重入后 count=1,第二次失败直接 Fallback。
- Evidence Retry Node 将 run-level evidence count 增加一次并重置 planner retry count。
- Edge 只选择 route,不隐式修改状态。
### Event and trace
- 每个 Fake/后续真实 Node 返回一个 event list,AppendStrategy 负责累积。
- event 包含 node/outcome/reasonCode/attempt,不包含 payload。
- trace builder 由相邻 events 生成 transition;final node/reason 来自最后 event。
- events 为空时拒绝构造 trace,避免伪造路径。
### Interface impact
- 等级:L2 internal interface。
- 新消费者:阶段 2 Node adapters、阶段 3 ChatService orchestrator、阶段 4 tests。
- 外部 API/DB/运行路径:无变化。
- 回滚:revert 本阶段提交即可;因为未接生产入口,没有数据迁移。
- 兼容:后续 action adapters 必须实现已冻结 ports,不得传递父 State 全量数据。
## Risks Accepted
- Stage 1 skeleton 会暂时存在但未被生产调用,这是阶段边界要求,不是死代码最终状态。
- 真实 Agent 状态映射尚未验证,由阶段 2 独立 change 负责。
- Maven E2E 不运行,生产路径完全未改变。
## Apply Evidence
- Task 1.1:显式声明 BOM 管理的 Graph Core;新增 26 个 state keys、状态枚举、拓扑/route/reason 常量、默认 Replace + events Append 策略和安全 typed reads。
- Task 1.2:新增不可变 event/transition/trace records,构造期拒绝空 routing metadata 和非法计数,map 输出只有冻结字段。
- Task 2.1:新增八个 non-null config-aware action ports,Fake/真实 Node 共用同一 Graph 接入面。
- Task 2.2:新增纯 router;所有未知/缺失状态 fail closed,LOW_CONFID guard 只读 facts_checked,不构造 retry context。
- Task 2.3:Graph Factory 注册八节点和全部条件边;wrapper 独占 retry/evidence control state,compile recursion limit=32。
- 首模块 Maven compile:passed(31s),锁定 1.1.2.0 API 假设成立。
- Task 3.1:trace builder 从 event 单向派生 transitions/final reason/degraded/evidence count,空或异类 events 显式失败。
- Task 4.1:新增严格 FIFO Fake Node fixture,记录 sequence/calls/threadId,每 attempt 只追加一个 terminal event,意外调用立即失败。
- Task 4.2:真实 CompiledGraph normal/Planner/Executor/Gatekeeper focused tests 首轮通过(Maven exit 0,约 67s)。
- Task 4.3:Verifier/evidence/Composer 路由与独立计数测试通过(Maven exit 0,约 12s)。
- Task 4.4:routing + trace focused suite 通过(Maven exit 0,约 28s),覆盖事件顺序、degraded、map 白名单、不可变性和非法输入。
- Task 5.1:35 tests(29 routing + 6 trace)全通过;Maven test compilation、change strict 1/1、主 specs 12/12、diff check 均通过。
- Task 5.2:现有 production refs=0,新 Graph 对真实 Service/DB/Trace refs=0,23 个 changed paths 全部命中阶段白名单;Maven E2E/log/DB 按用户口径保留到阶段 5。
- Review:补齐所有正常节点 threadId 传播、Verifier REJECT→Composer 和 completed-without-verdict fail-closed 用例;增强后 focused suite exit 0。
## Cross-Artifact 对齐检查
| 上游 → 下游 | 检查内容 | 状态 |
|---|---|---|
| brief/prd → proposal | 阶段 1 目标、Fake Node 边界、测试口径和 E2E 非目标 | 已对齐 |
| proposal → 设计产物 | direct dependency、状态、ports、router、counter、trace、测试架构 | 已对齐 |
| 设计产物 → specs/tasks | 所有可观察路由、终止、安全默认和实现模块 | 已对齐 |
| specs → tasks | 编译、全路由、trace、focused tests、生产隔离检查 | 已对齐 |
Gap:无。
## Architecture Audit
输入是 run-scoped 初始 state 与 `RunnableConfig`,处理链是 CompiledGraph → config-aware action ports → deterministic router/counters,输出是最终 state 与纯路由 trace;阶段 1 没有 Controller/Service/DB 消费者。Graph State 由单次 invoke 所有,events 只由 Node append,transitions 只由 builder 派生,避免双写。阶段 2 只实现 action ports,阶段 3 才将 orchestrator 交给 ChatService,因此当前未接生产入口是刻意生命周期边界。审计发现的唯一缺口是 Graph Core 直接依赖所有权,已回写 proposal/design/tasks。与阶段 0 archive 无冲突,L2 风险可由 Fake Node CompiledGraph tests 和 revert 单提交控制。
@@ -0,0 +1,31 @@
# Chat Diagnosis StateGraph Routing Skeleton Evidence
## 证据
| 来源 | 证据 | 结论 | 是否已汇报 |
|---|---|---|---|
| 阶段 0 archive/ADR | 冻结 26 个 state keys、完整路由、四类计数、run audit | 阶段 1 实现未偏离基线 | 是 |
| 本地 Graph Core 1.1.2.0 `javap` | config-aware Node/Edge、conditional edge、recursion limit、threadId 存在 | 使用公共锁定 API 可编译 | 是 |
| AppendStrategy bytecode | list/collection 按顺序追加 | 每 Node 返回单 event list 可形成实际路径 | 是 |
| Maven compile | `mvn -q "-DskipTests" compile` exit 0 | direct dependency 和 Factory API 编译成立 | 是 |
| Fake Node CompiledGraph tests | 29 routing tests,0 failure/error | 全条件边、retry、threadId、unknown fail-closed 成立 | 是 |
| Trace tests | 6 tests,0 failure/error | transitions、degraded、map 白名单、不可变和非法输入成立 | 是 |
| Test compilation | `mvn -q "-DskipTests" test` exit 0 | 全测试源可编译 | 是 |
| OpenSpec validation | change 1/1、主 specs 12/12 strict | artifacts 与既有规格无回归 | 是 |
| 生产隔离检查 | outside Graph refs=0,Graph 对真实 Service/DB/Trace refs=0 | 阶段 1 未接生产入口 | 是 |
| Git 路径白名单 | 23 changed paths,out-of-scope=0 | 无跨阶段文件混入 | 是 |
## Evidence-driven 结论
- 结论:直接声明 Graph Core 是正确依赖所有权。
- 证据:生产代码直接 import Graph Core;依赖原先仅由 Agent Framework 传递。
- 风险:BOM 升级仍需重新跑真实 Graph tests。
- 用户确认:不需要,属于构建稳健性。
- 结论:recursion limit 32 足够且没有替代业务计数。
- 证据:最坏合法路径低于 20;第二次技术失败和第二次 LOW_CONFID tests 均终止。
- 风险:未来新增循环必须重算。
- 用户确认:不需要,冻结业务上限未变。
- 结论:单元测试必要,E2E 不必要。
- 证据:阶段新增条件边行为但未接生产入口;35 tests 直接验证 Graph。
- 风险:真实 Agent 映射仍留给阶段 2。
- 用户确认:符合用户按必要性和最终阶段 E2E 规则。
@@ -0,0 +1,50 @@
# Chat Diagnosis StateGraph Test Suite 验收
## 结果
已接受。OpenSpec tasks 21/21 完成,阶段 4为 test-only,无生产行为变更。
## 验证
### 静态验证
- `git diff --check`:通过。
- `git diff --name-only HEAD -- src/main`:0 个文件。
- test inventory:三个权威类均存在;`ChatServiceSequentialAgentTest`/`VerifierInputHookTest` 均不存在;`ScriptedDiagnosisGraphActions` 定义 1 处。
- `openspec validate chat-diagnosis-stategraph-test-suite --strict`:通过。
- `openspec validate --specs --strict`:15 passed,0 failed。
### 脚本验证
- 三层权威 suite:4 suites / 43 tests,0 failures/errors/skipped。
- 完整保留安全回归:31 suites / 126 tests,0 failures/errors/skipped。
- `mvn -q -DskipTests test-compile`:通过。
### 浏览器/人工验证
- 未运行;阶段 4仅重构自动化测试,且用户要求最终人工/live 验收到阶段 5统一执行。
### 未验证
- 未使用 Maven 启动应用做 live E2E。
- 未检查 `logs/`。
- 未执行 `scripts/query_mysql.py`。
- 剩余风险:真实模型、工具、Flyway/MySQL 和最终 Trace 内容仍需阶段 5 E2E/log/DB 证据。
## 已完成范围
- 建立 Workflow、Node Contract、Chat Integration 三层权威测试体系。
- 补齐 ceiling/second LOW_CONFID、tool-failure legal snapshot、partial-pass REJECT 和 Verifier invalid status 边界。
- 删除旧 Hook implementation test,并保留 parser/Gatekeeper/projection/Composer 等安全回归。
- 用结构测试和 source gate 防止旧 Sequential/Hook 测试回归。
## 已知限制
- 生产 `VerifierInputHook`/`VerifierContextHolder` 类型仍存在但已无生产/权威测试消费者;阶段 5清理。
- component tests 与权威层存在有意的分层 overlap,详见 `evidence.md`。
## 交接
- OpenSpec archive:已同步 2 份 delta specs,并归档到 `openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-test-suite/`。
- 下一步:审查精确 Git diff并完成阶段 4独立提交,之后启动阶段 5。
- OpenSpec 归档确认:用户已授权直接执行后续归档;归档已完成。
@@ -0,0 +1,20 @@
# Chat Diagnosis StateGraph Test Suite Brief
## 背景
- 用户目标:将 ISS-011 阶段 4作为独立 sm-flow,建立以 Graph 路径和外部行为为中心的新测试体系。
- 当前问题:阶段 1–3 覆盖充分但组织分散,旧 Hook payload test 仍会约束已退出生产的隐式状态机。
- 关联 OpenSpec:`openspec/changes/chat-diagnosis-stategraph-test-suite/`
- devflow 分档:complex;接口影响 L1 test-only。
## 范围
- 本次要做:Workflow/Node Contract/Chat Integration 三层权威入口、Issue 路径矩阵补齐、旧 Hook test 退役、保留安全回归。
- 本次不做:不改生产源码,不删除生产 Hook/ThreadLocal 类型,不运行 live E2E/log/DB 验收。
- 影响区域:Diagnosis Graph tests、ChatService integration test、OpenSpec/devflow。
## OpenSpec 对齐
- proposal/design/specs:已覆盖,2 个 delta capabilities。
- tasks:21/21 已完成。
- 生产行为:无变化,`src/main` diff=0。
@@ -0,0 +1,121 @@
# Chat Diagnosis StateGraph Test Suite Decisions
## Entry Summary
- 问题:阶段 1–3 的安全覆盖已充足,但测试命名/组织仍是逐步实现产物,尚未形成 ISS-011 指定的 Workflow、Node Contract、Chat Integration 三层权威体系。
- 期望:阶段 4只重构测试结构并补齐矩阵,不改变生产行为;归档并提交后才进入阶段 5。
- 分档:complex(路径矩阵广、涉及旧安全测试退役,但接口影响为 L1 test-only)。
- Change:`chat-diagnosis-stategraph-test-suite`。
- 授权:用户已要求直接实现,阶段门禁与阶段 5 才 live E2E 的约束不变。
## Context Sources
- ISS-011 阶段 4、测试策略和验收标准。
- `chat-diagnosis-stategraph-design-freeze` 的 test migration requirement。
- 阶段 1–3 archive/acceptance 和当前 119-test focused baseline。
- `DiagnosisGraphRoutingTest`、各 Node tests、`DiagnosisRealGraphIntegrationTest`、`ChatDiagnosisGraphRuntimeTest`、`ChatServiceGraphIntegrationTest`。
- `VerifierInputHookTest`、protocol parser tests、`ExecutorGatekeeperServiceTest` 与 Trace/Controller/Repository/Eval tests。
## Question Pool
| # | 维度 | 问题 | 模式 | 状态 |
|---|---|---|---|---|
| Q1 | 术语 | Workflow、Node Contract、Chat Integration 的职责边界是什么? | evidence-driven | 已解决 |
| Q2 | 边界 | 是否把所有现有 Graph unit tests 合并成三个巨型类? | evidence-driven | 已解决 |
| Q3 | 旧测试 | `VerifierInputHookTest` 在显式 Graph Gatekeeper 后应保留、改写还是删除? | evidence-driven | 已解决 |
| Q4 | 验收 | 如何证明阶段 4矩阵完整且没有恢复固定 Sequential 顺序? | evidence-driven | 已解决 |
| Q5 | 阶段 | 是否允许为测试可测性改生产代码或执行 live E2E? | user-interview(既有冻结规则) | 已确认 |
## Evidence-driven Findings
- Q1:Issue 已明确三类测试;现有 `DiagnosisGraphRoutingTest` 对应 Workflow,分散 Node/Protocol tests 对应 Node Contract,`ChatServiceGraphIntegrationTest` 对应外部生命周期。
- Q2:现有细粒度测试失败定位清晰,全部合并会制造大文件;应保留专用 unit tests,同时新增/重命名三层权威入口并共享夹具。
- Q3:生产 Graph Verifier 已不注册 Hook,Hook test 仍验证 raw/full-trace/ThreadLocal payload,与 verified-only 生产协议冲突;parser/Gatekeeper/投影行为已有独立 tests,故阶段 4删除 Hook test,生产类型留到阶段 5。
- Q4:以 Issue 必需路径清单建立 requirement-to-test matrix;source check 禁止 `SequentialAgent`/`VerifierInputHookTest` 成为新 suite 依赖,并运行保留安全回归。
## User-interview Confirmation
| 问题 | 用户原话/既有确认 | 状态 | OpenSpec 回写 |
|---|---|---|---|
| Q5 阶段边界 | “端到端只在最后阶段全部完成后才验证;每个阶段如果有必要添加单元测试验收的话,就加” | 已确认 | proposal |
## Grill-with-docs Result
- 术语不进入业务 glossary:Workflow/Node Contract/Integration 是测试架构术语,不改变 Session、Run、Trace、Gatekeeper 或 evidence gap 领域定义。
- 具体场景压力测试:同 session 多 run 属于 Chat Integration;Gatekeeper REJECT/LOW_CONFID 与 retry exhaustion 属于 Workflow;failed binding 过滤与 verified-only payload 属于 Node Contract。
- 旧 Hook test 的有效行为已分别迁移到 `ExecutorEvidenceParserTest`、`ExecutorGatekeeperServiceTest`、`GatekeeperNodeTest` 和 `VerifiedInputNodeTest`;删除不会丢失安全真理源。
- 没有难以逆转的新架构取舍,不创建 ADR;测试组织可以在保持行为矩阵的前提下继续演进。
## Discover Status
- `devflow/index.md`:命中阶段 0–3 archive。
- 生产调用链:阶段 3 已冻结且测试阶段默认不修改。
- 接口影响:L1 test-only;无 API/DTO/DB/Prompt/运行时消费者变化。
- 未解决问题:0。
- Draft 产物:proposal + decisions;尚未生成 design/spec/tasks,尚未修改测试代码。
## Architecture Audit
### Test ownership map
`ISS-011 path matrix -> DiagnosisGraphWorkflowTest -> ScriptedDiagnosisGraphActions -> real DiagnosisGraphFactory` 负责控制流;`Node/protocol contracts -> DiagnosisGraphNodeContractTest + focused component tests -> real Node actions/parsers/Gatekeeper` 负责安全投影;`public lifecycle -> ChatServiceGraphIntegrationTest -> Run repositories/Eval/Trace mapping` 负责外部行为。Controller、Repository、Trace、Eval 和 protocol tests 是三层体系的下游安全消费者,不应被重写成 Graph 内部顺序断言。
### Data and lifecycle ownership
- Workflow fixtures 只拥有脚本状态、调用计数和 events,不创建 Session/Run。
- Node Contract fixtures 只拥有 invoker input/output 和 mock Gatekeeper current-run result,不持久化生产实体。
- Chat Integration fixtures 只观察 ChatService public result 和 current Run persistence,不推断内部 Node 次序。
- `VerifierInputHookTest` 删除后不产生数据契约缺口:Executor parser、Gatekeeper、passed-binding projection 各自已有单一测试所有者。
### Coupling risks
- 最大风险是 Workflow 与 Node Contract 都断言完整路径而重复;设计将 Fake route matrix 与 real-node input/security matrix分开。
- `ScriptedDiagnosisGraphActions` 是唯一 Fake topology fixture;不新增第二套 Graph builder。
- 测试-only阶段禁止 `src/main` diff,避免为测试便利扩大 production API。
- 类重命名使用 Git rename,旧名称仅允许出现在 OpenSpec/devflow迁移说明中。
### Cross-artifact alignment
| 上游 → 下游 | 检查内容 | 状态 |
|---|---|---|
| brief/proposal → proposal | 三层体系、旧 Hook test退役、保留安全回归、阶段 5 E2E 延期 | 已对齐 |
| proposal → design | rename-not-copy、职责归属、production diff=0、rollback | 已对齐 |
| design → specs/tasks | Workflow/Node/Integration矩阵、Hook test删除、source inventory和验证门禁 | 已对齐 |
| specs → tasks | 每类 scenario 均有 rename、补缺、回归或静态验收任务 | 已对齐 |
### Audit result
架构审计未发现业务 glossary、阶段 0–3 specs 或生产行为冲突。新 capability 只描述测试验证系统,modified design-freeze requirement 完整保留并细化旧测试替换边界。接口影响保持 L1,cross-artifact gap=0,无需回写生产 spec 或创建 ADR。
## Commit Gate
- schema:spec-driven;proposal/design/2 delta specs/tasks 全部 done,applyRequires=`tasks` 已满足。
- OpenSpec:当前 change strict pass;15 个主 specs strict pass。
- Cross-artifact:4/4 已对齐,gap=0。
- Question pool:4 个 evidence-driven 已查证,1 个 user-interview 由用户既有原话确认,无未决项。
- Interface impact:L1 test-only;design 有独立影响/回滚章节,默认 `src/main` diff=0。
- Preflight:`git diff --check` 通过,当前仅 Draft OpenSpec/devflow 与根目录计划文件,无测试/生产代码修改。
- 结论:Draft OpenSpec 达到可执行状态,创建 `.committed` 后进入 Apply。
## Apply Completion
- `DiagnosisGraphRoutingTest` 已以 Git rename 演进为 `DiagnosisGraphWorkflowTest`;所有 scripted run 统一断言 events 与真实 sequence 完全一致。
- `DiagnosisRealGraphIntegrationTest` 已演进为 `DiagnosisGraphNodeContractTest`;补齐合法工具失败限制、partial-pass REJECT 和 Verifier invalid 无伪 verdict。
- Workflow 显式断言 Gatekeeper ceiling 将模型 PASS 限制为 LOW_CONFID 且不触发 evidence retry,以及第二次 LOW_CONFID 不再补证据。
- `ChatServiceGraphIntegrationTest` 增加 SessionContextHolder finally cleanup 断言;public Run/Trace/Eval/multi-run 覆盖保持。
- `VerifierInputHookTest` 已删除;生产 Hook/ThreadLocal 类型未修改,留待阶段 5清理。
- 新增 `DiagnosisGraphTestSuiteStructureTest`,保证三个权威类存在、Sequential/Hook实现测试不存在且权威测试不引用旧实现。
## Verification Summary
- 新权威层:4 suites / 43 tests,0 failures/errors/skipped。
- 完整保留安全回归:31 suites / 126 tests,0 failures/errors/skipped。
- Maven test compilation:通过。
- OpenSpec:当前 change strict pass;主 specs 15/15 strict pass。
- 静态门禁:`git diff --check` 通过;required authoritative classes=3;legacy tests=0;scripted fixture definitions=1;`src/main` diff=0。
- 按阶段门禁未运行 Maven live E2E、未检查 `logs/`、未执行 `scripts/query_mysql.py`;统一保留到阶段 5。
## Archive Result
- 2 份 delta specs 已同步:新增 test-suite capability 6 条 requirements,修改 design-freeze test migration requirement 1 条。
- OpenSpec 已归档到 `openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-test-suite/`。
@@ -0,0 +1,42 @@
# Chat Diagnosis StateGraph Test Suite Evidence
## Requirement-to-test Matrix
| ISS-011 路径/边界 | 权威测试 | 辅助证据 |
|---|---|---|
| PASS 与精确 event/transition | `DiagnosisGraphWorkflowTest.normalPassPathUsesCompiledGraphAndPreservesEventOrder` | `DiagnosisGraphNodeContractTest.compiledRealNodeGraphCompletesPassPathInExactOrder` |
| Planner 一次技术重试/耗尽/非重试失败 | `DiagnosisGraphWorkflowTest.planner*` | Planner adapter focused test |
| Executor FAILED/TOOL_BLOCKED/INVALID/no-evidence | `DiagnosisGraphWorkflowTest.executor*` | `DiagnosisGraphNodeContractTest.legalSnapshotAfterToolFailureRemainsCompletedAndReachesGatekeeper` |
| Gatekeeper PASS/LOW/REJECT/unknown | `DiagnosisGraphWorkflowTest.gatekeeper*` | Gatekeeper Node/service tests |
| partial-pass REJECT 不泄漏 claim | Workflow unsafe route | `DiagnosisGraphNodeContractTest.gatekeeperRejectSkipsVerifierAndUsesPreVerificationFallback` |
| verified-only passed-binding projection | Workflow VerifiedInput route | `VerifiedInputNodeTest` + Verifier adapter test |
| Verifier 技术重试/耗尽/非重试失败 | `DiagnosisGraphWorkflowTest.verifier*` | `DiagnosisGraphNodeContractTest.verifierInvalidOutputSetsExecutionStatusWithoutFabricatedVerdict` |
| critical evidence retry、无 gap、non-critical、ceiling、第二次 LOW | `DiagnosisGraphWorkflowTest.*LowConfidence*` / `secondLowConfidence*` | EvidenceRetryPrepareNodeTest + real full-snapshot integration |
| 第二轮完整 snapshot 再验真 | Workflow evidence retry | `DiagnosisGraphNodeContractTest.criticalGapPerformsOneIncrementalRoundAndRevalidatesCompleteSnapshot` |
| Verifier REJECT 安全表达 | `DiagnosisGraphWorkflowTest.verifierRejectStillRunsComposerWithSafeMaterial` | ComposerSafeInputBuilderTest |
| Composer retry/耗尽/非重试/安全 Fallback | `DiagnosisGraphWorkflowTest.composer*` | ComposerNodeAdapterTest + FallbackNodeTest |
| ChatResult/Run/Trace/evaluation/cleanup | `ChatServiceGraphIntegrationTest` | Controller/Trace/Repository/Eval tests |
| 同 session 多 run 隔离 | `ChatServiceGraphIntegrationTest.sameSessionCreatesDistinctRunIdsAndKeepsTracePerRun` | Run repository/Trace exact-run tests |
| 三层结构与旧实现测试退役 | `DiagnosisGraphTestSuiteStructureTest` | source inventory command |
## 旧 Hook test 安全映射
| 旧行为 | 新真理源 |
|---|---|
| fenced/prefixed/malformed Executor JSON | `ExecutorEvidenceParserTest` |
| invocation/raw_path/evidence excerpt真实性 | `ExecutorGatekeeperServiceTest` |
| Gatekeeper status/severity/audit | `GatekeeperNodeTest` |
| passed binding 精确投影 | `VerifiedInputNodeTest` |
| Verifier verified-only payload | `VerifierNodeAdapterTest` + `ChatVerifierPromptContractTest` |
## 验证证据
- 31 suites / 126 tests:0 failures、0 errors、0 skipped。
- test compilation、当前 change strict、主 specs 15/15、diff check 全通过。
- `src/main` diff=0;三个权威类存在;旧 Sequential/Hook tests不存在;Scripted fixture 定义唯一。
## Intentional overlap and limits
- Workflow 与 Node Contract 都覆盖 PASS/REJECT,但前者验证 route/event,后者验证真实输入/安全材料;这是分层证据,不是复制 fixture。
- 细粒度 component tests 保留以定位失败,不要求全部搬进三个权威类。
- 真实模型/工具、日志和数据库行为未在阶段 4验证,统一由阶段 5 live E2E承担。
+149 -2
View File
@@ -12,12 +12,59 @@ post-processing, or Spring AI VectorStore integration.
```text
eval/rag-retrieval/
cases/golden-cases.json Fixed retrieval golden cases
fixtures/*.json Saved retrieval candidates for each case
seed-docs/*.md Canonical docs imported into the live KB for real-tool eval
fixtures/*.json Saved retrieval fixtures for each case
reports/baseline.json Machine-readable baseline report
reports/baseline.md Human-readable baseline report
reports/baseline-diff.* Optional diff reports
reports/live-post-reindex.* Optional live acceptance reports
```
## Seed Docs + Import/Reindex
The live-tool eval uses canonical seed documents so the real
`LookupKnowledgeTool` can retrieve stable evidence from MySQL/Milvus instead of
whatever ad hoc documents happen to exist in the local knowledge base.
Seed documents live in:
```text
eval/rag-retrieval/seed-docs/*.md
```
Each seed doc uses frontmatter fields that are propagated into vector metadata:
```yaml
source: mysql-connection-pool
breadcrumb: Database > MySQL > Connection Pool
kb_scope: rag-eval
```
Import or reindex the seed docs through the real upload pipeline:
```powershell
.\scripts\prepare_rag_eval_seed.ps1
```
The script runs `RagEvalSeedImporterTest` with `rag.seed.enabled=true`. It
deletes the existing document with the same `source`/`docId`, uploads the seed
doc through `DocumentManagementService`, updates DB metadata and L0, and rebuilds
Milvus chunks.
`kb_scope` isolates eval data:
- default application config leaves `retrieval.kb-scope` empty, so legacy docs
without `kb_scope` remain searchable;
- eval scripts pass `-Dretrieval.kb-scope=rag-eval`, so L0 query hints and L1
vector retrieval both use only the canonical eval seed docs;
- the fallback retry skips only the L0 category filter, not the `kb_scope`
boundary.
Frontmatter is not embedded as chunk content during upload. It feeds metadata,
L0, and document enrichment; only the Markdown body is chunked and embedded.
This keeps controlled L0 decoys from becoming semantically relevant just because
their frontmatter keywords matched the query.
## Run
From the repository root:
@@ -36,6 +83,106 @@ python scripts/eval_rag_retrieval.py \
--markdown-report eval/rag-retrieval/reports/baseline.md
```
## Generate Fixtures From LookupKnowledgeTool
Use the snapshot generator when fixtures should reflect the real
`LookupKnowledgeTool` pipeline:
```powershell
.\scripts\generate_rag_lookup_snapshots.ps1
```
For the intended live loop, run seed import first:
```powershell
.\scripts\prepare_rag_eval_seed.ps1
.\scripts\generate_rag_lookup_snapshots.ps1
python scripts\eval_rag_retrieval.py
```
The script runs a Spring test harness:
```text
mvn -q -Dtest=RagLookupSnapshotGeneratorTest -Drag.snapshot.enabled=true -Dretrieval.kb-scope=rag-eval -Dretrieval.vector-store.mode=spring test
```
The generator reads `golden-cases.json`, injects the real `LookupKnowledgeTool`
bean, calls `lookupKnowledge(query)` for each case, writes
`fixtures/{caseId}.json`, and then runs `eval_rag_retrieval.py` unless
`-SkipEval` is provided. It defaults to Spring AI VectorStore mode; pass
`-VectorStoreMode sdk` only when intentionally comparing the legacy SDK path.
Custom paths are supported:
```powershell
.\scripts\generate_rag_lookup_snapshots.ps1 `
-Cases eval\rag-retrieval\cases\golden-cases.json `
-Fixtures eval\rag-retrieval\fixtures `
-RetrievedAt 2026-07-06T00:00:00Z
```
The generator is disabled in normal test runs. It only executes when
`rag.snapshot.enabled=true` is provided because it writes repository files and
depends on the configured runtime retrieval stack.
If generated fixtures fail the offline baseline, treat that as a real alignment
signal: either the golden expectations need to be adjusted to the current
knowledge base, or the knowledge base/indexing path needs to be fixed.
## Modular RAG Contract
Fixtures must use the current `lookupResult` shape, which mirrors the
`lookup_knowledge` output:
```text
lookupResult.evidenceBlocks
lookupResult.contextPack
lookupResult.retrievalTrace
lookupResult.rerankTrace
```
Golden cases can assert both retrieval quality and pipeline behavior:
- `expectedSources` / `expectedDocIds`
- `expectedBreadcrumbs`
- `expectedKeywords`
- `expectedSelectedAttempt`
- `expectedFallbackReason`
- `expectedFallbackReasons`
- `expectedEvidenceStatus`
- `expectedContextSources`
- `expectedRerankTopSource`
This lets the baseline catch regressions such as losing the expected evidence
source, skipping context packing, changing the selected retrieval attempt, or
breaking the filtered-vector to unfiltered-retry fallback.
## Baseline Diff
To compare a freshly generated report against an existing baseline:
```bash
python scripts/eval_rag_retrieval.py \
--json-report eval/rag-retrieval/reports/current.json \
--markdown-report eval/rag-retrieval/reports/current.md \
--compare-to eval/rag-retrieval/reports/baseline.json \
--diff-json-report eval/rag-retrieval/reports/baseline-diff.json \
--diff-markdown-report eval/rag-retrieval/reports/baseline-diff.md
```
The diff reports aggregate regressions and case-level changes for:
- pass rate, recall@K, strong hit rate, miss count
- pass state
- hit level
- first expected rank
- selected attempt
- fallback reason
- evidence status
- rerank top source
The command exits non-zero when a case fails or the diff contains a regression.
## Hit Levels
- `strong`: expected document is found and breadcrumb or evidence keyword coverage is satisfied.
@@ -81,6 +228,6 @@ GET /api/search/similar
```
It writes JSON and Markdown reports with query, topK, result count, top
candidates, breadcrumb, score labels, and raw response fields. This is a live
results, breadcrumb, score labels, and raw response fields. This is a live
smoke check for environment readiness and post-reindex behavior; it does not
replace the deterministic offline baseline above.
@@ -8,8 +8,14 @@
"scenario": "chat",
"query": "MySQL connection pool is exhausted. How should I diagnose it?",
"expectedDocIds": ["mysql-connection-pool"],
"expectedSources": ["mysql-connection-pool"],
"expectedBreadcrumbs": ["Database > MySQL > Connection Pool"],
"expectedKeywords": ["connection pool", "max_connections", "HikariCP"],
"expectedSelectedAttempt": "FILTERED_VECTOR",
"expectedFallbackReason": null,
"expectedEvidenceStatus": "supported",
"expectedContextSources": ["mysql-connection-pool"],
"expectedRerankTopSource": "mysql-connection-pool",
"notes": "Covers precise database troubleshooting retrieval."
},
{
@@ -17,8 +23,14 @@
"scenario": "chat",
"query": "What is the standard troubleshooting flow for an application incident?",
"expectedDocIds": ["incident-diagnosis-flow"],
"expectedSources": ["incident-diagnosis-flow"],
"expectedBreadcrumbs": ["AIOps > Diagnosis Flow"],
"expectedKeywords": ["collect evidence", "verify", "remediation"],
"expectedSelectedAttempt": "FILTERED_VECTOR",
"expectedFallbackReason": null,
"expectedEvidenceStatus": "supported",
"expectedContextSources": ["incident-diagnosis-flow"],
"expectedRerankTopSource": "incident-diagnosis-flow",
"notes": "Covers process-style knowledge where breadcrumb matters."
},
{
@@ -26,8 +38,14 @@
"scenario": "aiops",
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
"expectedDocIds": ["payment-service-latency"],
"expectedSources": ["payment-service-latency"],
"expectedBreadcrumbs": ["AIOps > Service Alerts > Payment Latency"],
"expectedKeywords": ["p95 latency", "payment-service", "downstream dependency"],
"expectedSelectedAttempt": "FILTERED_VECTOR",
"expectedFallbackReason": null,
"expectedEvidenceStatus": "supported",
"expectedContextSources": ["payment-service-latency"],
"expectedRerankTopSource": "payment-service-latency",
"notes": "Covers alert payload terms that should become retrieval hints."
},
{
@@ -35,8 +53,14 @@
"scenario": "aiops",
"query": "When an AIOps request already includes alert payload, should the agent diagnose unrelated active alerts?",
"expectedDocIds": ["aiops-alert-scope-control"],
"expectedSources": ["aiops-alert-scope-control"],
"expectedBreadcrumbs": ["AIOps > Alert Scope Control"],
"expectedKeywords": ["payload", "unrelated active alerts", "scope"],
"expectedSelectedAttempt": "FILTERED_VECTOR",
"expectedFallbackReason": null,
"expectedEvidenceStatus": "supported",
"expectedContextSources": ["aiops-alert-scope-control"],
"expectedRerankTopSource": "aiops-alert-scope-control",
"notes": "Covers scoped alert diagnosis behavior."
},
{
@@ -44,8 +68,14 @@
"scenario": "chat",
"query": "If a long section is split into multiple chunks, how do we keep retrieval context?",
"expectedDocIds": ["rag-chunk-context-reconstruction"],
"expectedSources": ["rag-chunk-context-reconstruction"],
"expectedBreadcrumbs": ["RAG > Chunking > Context Reconstruction"],
"expectedKeywords": ["neighbor chunk", "same section", "breadcrumb"],
"expectedSelectedAttempt": "FILTERED_VECTOR",
"expectedFallbackReason": null,
"expectedEvidenceStatus": "supported",
"expectedContextSources": ["rag-chunk-context-reconstruction"],
"expectedRerankTopSource": "rag-chunk-context-reconstruction",
"notes": "Covers the known RAG refactor issue around context reconstruction."
},
{
@@ -53,9 +83,30 @@
"scenario": "chat",
"query": "Should L0 keyword matching decide the final retrieval result?",
"expectedDocIds": ["rag-l0-domain-entity-hint"],
"expectedSources": ["rag-l0-domain-entity-hint"],
"expectedBreadcrumbs": ["RAG > L0 > Domain Entity Hint"],
"expectedKeywords": ["domain detector", "entity extractor", "metadata filter"],
"expectedSelectedAttempt": "FILTERED_VECTOR",
"expectedFallbackReason": null,
"expectedEvidenceStatus": "supported",
"expectedContextSources": ["rag-l0-domain-entity-hint"],
"expectedRerankTopSource": "rag-l0-domain-entity-hint",
"notes": "Covers the target L0 role after refactor."
},
{
"caseId": "chat-l0-filter-fallback",
"scenario": "chat",
"query": "RAG query was over-filtered by L0 and filtered vector search returned low quality evidence. What should happen?",
"expectedDocIds": ["rag-l0-filter-fallback"],
"expectedSources": ["rag-l0-filter-fallback"],
"expectedBreadcrumbs": ["RAG > Fallback > Unfiltered Retry"],
"expectedKeywords": ["skip the L0 filter", "unfiltered vector retry", "low quality"],
"expectedSelectedAttempt": "UNFILTERED_VECTOR_RETRY",
"expectedFallbackReasons": ["filtered_vector_low_quality", "filtered_vector_no_evidence"],
"expectedEvidenceStatus": "supported",
"expectedContextSources": ["rag-l0-filter-fallback"],
"expectedRerankTopSource": "rag-l0-filter-fallback",
"notes": "Covers the MVP fallback rule: if filtered L1 is low quality, retry raw query without L0 filter."
}
]
}
@@ -1,25 +1,81 @@
{
"caseId": "aiops-payment-latency-alert",
"query": "Alert HighLatency on payment-service with p95 latency above threshold",
"retrievedAt": "2026-07-05T00:00:00Z",
"candidates": [
{
"rank": 1,
"docId": "payment-service-latency",
"title": "Payment Service Latency Alert Playbook",
"breadcrumb": "AIOps > Service Alerts > Payment Latency",
"content": "For payment-service p95 latency alerts, check downstream dependency latency, thread pool saturation, gateway retries, and recent deployment changes.",
"score": 0.84,
"retrievalLayer": "L1"
"retrievedAt": "2026-07-06T00:00:00Z",
"lookupResult": {
"found": true,
"evidenceBlocks": [
{
"source": "payment-service-latency",
"title": "Payment Service Latency Alert Playbook",
"breadcrumb": "AIOps > Service Alerts > Payment Latency",
"retrievalLayer": "L1",
"content": "For payment-service p95 latency alerts, check downstream dependency latency, thread pool saturation, gateway retries, and recent deployment changes.",
"score": 0.84,
"hitReasons": ["domain_match:+0.15", "entity_match:+0.20", "keyword_match:+0.10"]
},
{
"source": "mysql-connection-pool",
"title": "MySQL Connection Pool Troubleshooting",
"breadcrumb": "Database > MySQL > Connection Pool",
"retrievalLayer": "L1",
"content": "Database connection pool saturation can increase payment latency when checkout paths wait for connections.",
"score": 0.68,
"hitReasons": ["keyword_match:+0.10"]
}
],
"contextPack": {
"packedText": "[1] Payment Service Latency Alert Playbook\nAIOps > Service Alerts > Payment Latency\nFor payment-service p95 latency alerts, check downstream dependency latency, thread pool saturation, gateway retries, and recent deployment changes.",
"strategy": "top_evidence_blocks",
"charBudget": 3500,
"usedChars": 236,
"includedSources": ["payment-service-latency", "mysql-connection-pool"],
"omittedSources": []
},
{
"rank": 2,
"docId": "mysql-connection-pool",
"title": "MySQL Connection Pool Troubleshooting",
"breadcrumb": "Database > MySQL > Connection Pool",
"content": "Database connection pool saturation can increase payment latency when checkout paths wait for connections.",
"score": 0.68,
"retrievalLayer": "L1"
"retrievalTrace": {
"originalQuery": "Alert HighLatency on payment-service with p95 latency above threshold",
"rewrittenQuery": "HighLatency payment-service p95 latency alert downstream dependency diagnosis",
"categoryFilter": "AIOps",
"selectedAttempt": "FILTERED_VECTOR",
"fallbackReason": null,
"evidenceStatus": "supported",
"queryHints": {
"domains": ["AIOps"],
"matched_keywords": ["p95 latency", "payment-service", "downstream dependency"],
"entities": ["payment-service", "HighLatency"],
"l0_titles": ["Payment Service Latency Alert Playbook"],
"l0_match_count": 1
},
"attempts": [
{
"name": "FILTERED_VECTOR",
"query": "HighLatency payment-service p95 latency alert downstream dependency diagnosis",
"categoryFilter": "AIOps",
"candidateCount": 2,
"usable": true,
"durationMs": 11,
"topScore": 0.84,
"topSimilarity": 0.84
}
]
},
"rerankTrace": {
"items": [
{
"finalRank": 1,
"source": "payment-service-latency",
"baseScore": 0.84,
"finalScore": 1.29,
"boostReasons": ["domain_match:+0.15", "entity_match:+0.20", "keyword_match:+0.10"]
},
{
"finalRank": 2,
"source": "mysql-connection-pool",
"baseScore": 0.68,
"finalScore": 0.78,
"boostReasons": ["keyword_match:+0.10"]
}
]
}
]
}
}

Some files were not shown because too many files have changed in this diff Show More