feat(graph): complete stategraph cleanup and acceptance

This commit is contained in:
zhuyongxin
2026-07-20 10:23:27 +08:00
parent 208a231113
commit 190013c901
46 changed files with 1223 additions and 1211 deletions
@@ -0,0 +1,91 @@
# chat-diagnosis-stategraph-cleanup-docs Specification
## Purpose
TBD - created by archiving change chat-diagnosis-stategraph-cleanup-docs. Update Purpose after archive.
## Requirements
### Requirement: Legacy implicit verifier orchestration SHALL be removed
The final StateGraph implementation SHALL have no production or test definition, import, instantiation, or executable type dependency for the legacy Verifier Hook, its ThreadLocal context, or its dedicated full-trace summary producer. Negative contract tests MAY retain identifier strings solely to prevent regression.
#### Scenario: Legacy source inventory is inspected
- **WHEN** stage 5 cleanup completes
- **THEN** `VerifierInputHook`, `VerifierContextHolder`, and `ToolTraceSummaryService` source files SHALL NOT exist
- **AND** no production or test source SHALL import, instantiate, extend, or type-reference those types
- **AND** identifier string literals SHALL only be allowed in negative source/Prompt contract assertions
#### Scenario: Graph safety contracts are rerun
- **WHEN** the legacy closure is removed
- **THEN** Executor parser, Gatekeeper service/node, Verified Input, Verifier, Composer, Fallback, Trace, and Chat integration tests SHALL pass
- **AND** Maven test compilation and live Spring startup SHALL succeed without the removed beans
### Requirement: Current documentation SHALL describe the real StateGraph architecture
Current architecture, eval, and demo documentation SHALL describe complex Chat as explicit bounded Diagnosis StateGraph orchestration with run-owned events and verified-only Verifier input.
#### Scenario: Current docs are inspected
- **WHEN** a maintainer follows the architecture index and active eval/demo guides
- **THEN** Chat orchestration SHALL be described as conditional StateGraph Nodes rather than SequentialAgent or Hook Gatekeeper
- **AND** Verifier input SHALL use verified Executor projection/evidence rather than full tool trace
- **AND** `run.orchestrationTrace` SHALL be documented separately from self-evaluation and detailed Agent/tool Trace
#### Scenario: Historical docs are inspected
- **WHEN** a maintainer reads archived issues, design notes, or legacy fixtures
- **THEN** those materials MAY retain their original Hook/full-trace terminology
- **AND** they SHALL NOT be indexed as the current implementation truth
### Requirement: Final deterministic regression SHALL remain green
The project SHALL run the authoritative Graph suite, retained public/security contracts, Maven test compilation, and fixed diagnosis eval baseline before live acceptance.
#### Scenario: Final automated gates run
- **WHEN** stage 5 implementation and documentation are complete
- **THEN** Workflow, Node Contract, Chat Integration, Trace, Controller, Repository, Gatekeeper, Composer, parser/projection, and Eval tests SHALL pass
- **AND** the fixed diagnosis eval report/diff SHALL show no unexpected regression
- **AND** OpenSpec strict and source/whitespace checks SHALL pass
### Requirement: Final live E2E SHALL prove exact StateGraph Run ownership
The final acceptance SHALL start the application through Maven with the `mvp-demo` profile and SHALL execute Chat, exact Trace, and feedback requests using one unique sessionId and the returned runId.
#### Scenario: Live Chat Graph completes
- **WHEN** the unique payment-timeout demo request completes
- **THEN** the response SHALL be successful with a non-empty answer/sessionId/runId
- **AND** exact Trace SHALL return the same runId with a non-empty `run.orchestrationTrace`
- **AND** orchestration trace SHALL include version, final node, termination reason, transitions, degraded, and evidence retry count
- **AND** feedback SHALL attach to the same runId
#### Scenario: Live application is cleaned up
- **WHEN** live verification finishes or fails
- **THEN** the Maven/Java processes started by this stage SHALL be stopped
- **AND** port 9900 SHALL no longer be owned by the stage 5 process
- **AND** temporary target output/log files SHALL be removed after evidence extraction
### Requirement: Final logs and database SHALL corroborate the E2E Run
Stage 5 SHALL inspect only the current live run's new log segment and SHALL query exact database rows with the repository MySQL tool.
#### Scenario: New logs are inspected
- **WHEN** E2E returns sessionId and runId
- **THEN** new application/chat logs SHALL contain evidence for the current request/run lifecycle
- **AND** no unexplained ERROR in the new stage 5 segment SHALL invalidate acceptance
#### Scenario: Exact database run is queried
- **WHEN** `scripts/query_mysql.py` queries the E2E sessionId/runId
- **THEN** V012 orchestration column SHALL exist
- **AND** DiagnosisRun SHALL be CHAT/SUCCESS with non-empty answer, metrics, self-evaluation, orchestration trace, and useful feedback
- **AND** orchestration JSON fields SHALL match the exact Trace response
- **AND** AgentStep/ToolInvocation ownership checks SHALL contain no row from another run/session
### Requirement: ISS-011 SHALL close only after all final gates pass
The Issue SHALL remain active until cleanup, docs, deterministic regression, live E2E, logs, database, process cleanup, and OpenSpec validation are complete.
#### Scenario: Any final gate fails
- **WHEN** a required stage 5 gate is incomplete or failed
- **THEN** ISS-011 SHALL remain active
- **AND** acceptance SHALL record the blocker without marking the OpenSpec complete
#### Scenario: All final gates pass
- **WHEN** all stage 5 acceptance evidence is archived
- **THEN** ISS-011 checkboxes/status SHALL be completed
- **AND** the Issue SHALL move from active to archived with its index entry updated
@@ -37,11 +37,12 @@ The system SHALL provide an `mvp-demo` Spring profile that documents the demo ru
- **THEN** `prometheus.mock-enabled` and `cls.mock-enabled` are enabled by profile configuration
### Requirement: End-to-end MVP acceptance case is documented
The project SHALL include an end-to-end acceptance case that demonstrates start-up, chat diagnosis, trace query, and feedback submission using the same `sessionId + runId`.
The project SHALL include an end-to-end acceptance case that demonstrates Maven start-up, chat diagnosis, exact trace query, orchestration trace inspection, feedback submission, log inspection, and exact database verification using the same `sessionId + runId`.
#### Scenario: Reviewer follows the acceptance case
- **WHEN** a reviewer follows the documented MVP demo acceptance steps
- **THEN** they can run the application, submit a diagnosis question, query the exact trace endpoint, and submit feedback for the same run
- **THEN** they can run the application, submit a diagnosis question, query the exact trace endpoint, inspect `run.orchestrationTrace`, and submit feedback for the same run
- **AND** they can correlate that run with new logs and exact DiagnosisRun/AgentStep/ToolInvocation database records
### Requirement: MVP demo SHALL provide an interview runbook
The MVP demo SHALL include a concise interview runbook that explains how to demonstrate the Agent flow and how to narrate the engineering value.
@@ -67,14 +68,16 @@ The MVP demo SHALL provide scripts and request payloads for running the payment-
- **THEN** it SHALL write chat, exact trace, and feedback responses under a demo output directory
### Requirement: MVP demo SHALL be reproducible for interviews
The MVP demo SHALL provide a repeatable way to show a diagnosis answer, trace, verifier evaluation, and feedback.
The MVP demo SHALL provide a repeatable way to show a diagnosis answer, exact run trace, StateGraph orchestration summary, verifier evaluation, and feedback.
#### Scenario: interview demo check script records an evidence bundle
- **WHEN** the user runs the interview demo check script against a running `mvp-demo` service
- **THEN** the script SHALL submit a fixed Chat diagnosis request
- **AND** it SHALL fetch the trace for the same `sessionId + runId`
- **AND** it SHALL fail if `run.orchestrationTrace` or its version/final node/termination reason is missing
- **AND** it SHALL submit useful feedback for that run
- **AND** it SHALL write chat, trace, feedback, and summary outputs under `mvp/demo/output/`
- **AND** it SHALL write chat, trace, feedback, and summary outputs under the configured output directory
- **AND** summary SHALL include final node, termination reason, degraded, transition count, and evidence retry count
#### Scenario: interview demo check fails with actionable readiness output
- **WHEN** the target service is not reachable
@@ -82,16 +85,17 @@ The MVP demo SHALL provide a repeatable way to show a diagnosis answer, trace, v
- **AND** the failure message SHALL name the base URL and the expected startup profile
#### Scenario: interview documentation explains audit fields
- **WHEN** an interviewer asks how prompt or Gatekeeper changes are audited
- **THEN** the demo documentation SHALL point to `prompt_audit.version` and `gatekeeper_result.rule_set_version`
- **AND** it SHALL explain that deterministic eval fixtures are the regression source of truth
- **WHEN** an interviewer asks how Prompt, Gatekeeper, or Graph routing changes are audited
- **THEN** the demo documentation SHALL point to `prompt_audit.version`, `gatekeeper_result.rule_set_version`, and `run.orchestrationTrace`
- **AND** it SHALL explain that deterministic eval fixtures and the Graph test suite are the regression source of truth
### Requirement: MVP demo SHALL provide a trace inspection checklist
The MVP demo SHALL document which trace fields to inspect for evidence, verifier behavior, and run-level auditability.
The MVP demo SHALL document which trace fields to inspect for StateGraph routing, evidence, verifier behavior, and run-level auditability.
#### Scenario: Checklist maps fields to interview claims
- **WHEN** a developer reviews a trace response
- **THEN** the checklist SHALL map concrete JSON paths to the claims made in the interview walkthrough
- **WHEN** a developer reviews an exact trace response
- **THEN** the checklist SHALL map `run.orchestrationTrace` version/transitions/final node/termination reason/degraded/evidence retry count to Graph routing claims
- **AND** it SHALL map AgentStep, ToolInvocation, self-evaluation, Prompt audit, Gatekeeper audit, answer, and feedback paths to their separate responsibilities
### Requirement: MVP demo SHALL provide a browser trace workbench
The MVP demo SHALL provide a browser-accessible static page for inspecting one