feat(graph): cut over chat diagnosis stategraph

This commit is contained in:
zhuyongxin
2026-07-17 18:30:08 +08:00
parent 1460dd1e99
commit 99e490f227
36 changed files with 2640 additions and 1623 deletions
@@ -0,0 +1,110 @@
# chat-diagnosis-stategraph-chatservice-cutover Specification
## Purpose
TBD - created by archiving change chat-diagnosis-stategraph-chatservice-cutover. Update Purpose after archive.
## Requirements
### Requirement: Complex Chat SHALL use the real Diagnosis StateGraph as its only production orchestrator
The system SHALL execute each complex Chat request through one real Diagnosis CompiledGraph and SHALL NOT create, invoke, or fall back to a SequentialAgent workflow.
#### Scenario: Complex Chat starts
- **WHEN** `executeChatWithStrategy` classifies a valid question as complex
- **THEN** ChatService SHALL create the current Diagnosis Run and invoke one Diagnosis CompiledGraph
- **AND** the production path SHALL NOT maintain an outer score-based retry loop
#### Scenario: Graph dependencies are assembled
- **WHEN** the complex Chat Graph is constructed
- **THEN** Planner, Executor, Verifier, and Composer SHALL use the existing project Prompt, tool, skill, and AgentLoggingHook assembly rules
- **AND** the Graph Verifier SHALL NOT register VerifierInputHook
- **AND** Gatekeeper SHALL run only as the explicit Graph Node
### Requirement: Graph invocation SHALL preserve current Run ownership
The system SHALL use the current `runId` as Graph `threadId` and SHALL pass both current `sessionId` and `runId` in RunnableConfig metadata.
#### Scenario: Agent Node runs
- **WHEN** any Agent Node is invoked for a complex Chat run
- **THEN** its RunnableConfig threadId SHALL equal the current runId
- **AND** AgentStep and ToolInvocation writes SHALL retain the current sessionId and runId
#### Scenario: Gatekeeper validates Executor output
- **WHEN** the Gatekeeper Node runs
- **THEN** it SHALL validate only tool invocations belonging to the runId in RunnableConfig metadata
- **AND** data from another run in the same session SHALL NOT be considered
### Requirement: Initial Graph State SHALL contain only bounded diagnosis control data
ChatService SHALL initialize Diagnosis Graph State with the current query context, NORMAL Planner mode, zero independent retry counters, and an empty orchestration event list.
#### Scenario: Initial state is projected
- **WHEN** a complex Chat run enters Planner for the first time
- **THEN** `diagnosis_context` SHALL contain the current query/original query
- **AND** complete conversation history SHALL NOT be stored in parent Graph State
- **AND** history MAY remain in the Planner and Executor system Prompt assembled for this request
#### Scenario: Counters are initialized
- **WHEN** the Graph starts
- **THEN** planner, verifier, composer, and evidence retry counters SHALL each be zero
- **AND** Planner mode SHALL be NORMAL
### Requirement: Graph final state SHALL map to the existing Chat result and Run lifecycle
The system SHALL use non-empty `final_answer` from a handled Graph terminal state as the existing ChatResult answer. Run status SHALL express execution lifecycle rather than diagnosis quality.
#### Scenario: Composer completes
- **WHEN** Graph reaches Composer and produces a safe non-empty final answer
- **THEN** ChatResult SHALL preserve the current answer/sessionId/runId protocol
- **AND** DiagnosisRun SHALL be saved as SUCCESS with answer, duration, token count, step count, and tool count
- **AND** EvaluationService SHALL evaluate the current runId
#### Scenario: Handled Fallback completes
- **WHEN** Graph reaches deterministic Fallback and produces a safe non-empty answer
- **THEN** DiagnosisRun SHALL be SUCCESS
- **AND** diagnosis quality SHALL be expressed by verifier fields when available or `orchestrationTrace.degraded=true`
- **AND** no DEGRADED Run status SHALL be introduced
#### Scenario: Graph cannot produce a safe response
- **WHEN** Graph has an unhandled failure, final state is unavailable, final answer is blank, or required successful-result persistence fails
- **THEN** DiagnosisRun SHALL be marked FAILED when it can still be saved
- **AND** the system SHALL NOT report a successful Graph result
### Requirement: Graph state SHALL be the only source for verifier evaluation persistence
The system SHALL build `diagnosis_run.self_evaluation.verifier_evaluation` from explicit Graph State and SHALL NOT read VerifierContextHolder on the complex Chat path.
#### Scenario: Verifier completed
- **WHEN** final Graph State contains a completed Verifier result
- **THEN** verifier evaluation SHALL include execution status, model verdict, effective verdict, groundedness, claim/fact checks, rationale, round, Gatekeeper audit, verified Executor output/evidence, Prompt audit, and Composer audit when available
- **AND** the compatibility `verdict` field SHALL equal effective verdict
#### Scenario: Pre-verification Fallback completed
- **WHEN** Graph reaches Fallback before Verifier completes
- **THEN** verifier evaluation SHALL preserve available execution statuses, Gatekeeper audit, Prompt audit, and fallback context
- **AND** it SHALL NOT fabricate a model or effective verdict
#### Scenario: Evaluation payload is inspected
- **WHEN** verifier evaluation is persisted
- **THEN** it SHALL NOT contain raw Executor text or complete tool trace summary
- **AND** `executor_structured_output` SHALL contain at most the verified projection retained for compatibility
### Requirement: Stage 3 verification SHALL not run the final live E2E
The change SHALL use focused automated tests for production cutover and contracts while reserving Maven live startup, log inspection, and database querying for stage 5.
#### Scenario: Stage 3 is accepted
- **WHEN** stage 3 verification completes
- **THEN** Graph cutover, Run/Trace, Prompt, Controller/Eval regressions, Maven test compilation, and OpenSpec strict validation SHALL have passed
- **AND** live E2E, `logs/`, and `scripts/query_mysql.py` SHALL be recorded as intentionally deferred to stage 5