feat(graph): cut over chat diagnosis stategraph
This commit is contained in:
@@ -0,0 +1,110 @@
|
||||
# chat-diagnosis-stategraph-chatservice-cutover Specification
|
||||
|
||||
## Purpose
|
||||
TBD - created by archiving change chat-diagnosis-stategraph-chatservice-cutover. Update Purpose after archive.
|
||||
## Requirements
|
||||
### Requirement: Complex Chat SHALL use the real Diagnosis StateGraph as its only production orchestrator
|
||||
|
||||
The system SHALL execute each complex Chat request through one real Diagnosis CompiledGraph and SHALL NOT create, invoke, or fall back to a SequentialAgent workflow.
|
||||
|
||||
#### Scenario: Complex Chat starts
|
||||
|
||||
- **WHEN** `executeChatWithStrategy` classifies a valid question as complex
|
||||
- **THEN** ChatService SHALL create the current Diagnosis Run and invoke one Diagnosis CompiledGraph
|
||||
- **AND** the production path SHALL NOT maintain an outer score-based retry loop
|
||||
|
||||
#### Scenario: Graph dependencies are assembled
|
||||
|
||||
- **WHEN** the complex Chat Graph is constructed
|
||||
- **THEN** Planner, Executor, Verifier, and Composer SHALL use the existing project Prompt, tool, skill, and AgentLoggingHook assembly rules
|
||||
- **AND** the Graph Verifier SHALL NOT register VerifierInputHook
|
||||
- **AND** Gatekeeper SHALL run only as the explicit Graph Node
|
||||
|
||||
### Requirement: Graph invocation SHALL preserve current Run ownership
|
||||
|
||||
The system SHALL use the current `runId` as Graph `threadId` and SHALL pass both current `sessionId` and `runId` in RunnableConfig metadata.
|
||||
|
||||
#### Scenario: Agent Node runs
|
||||
|
||||
- **WHEN** any Agent Node is invoked for a complex Chat run
|
||||
- **THEN** its RunnableConfig threadId SHALL equal the current runId
|
||||
- **AND** AgentStep and ToolInvocation writes SHALL retain the current sessionId and runId
|
||||
|
||||
#### Scenario: Gatekeeper validates Executor output
|
||||
|
||||
- **WHEN** the Gatekeeper Node runs
|
||||
- **THEN** it SHALL validate only tool invocations belonging to the runId in RunnableConfig metadata
|
||||
- **AND** data from another run in the same session SHALL NOT be considered
|
||||
|
||||
### Requirement: Initial Graph State SHALL contain only bounded diagnosis control data
|
||||
|
||||
ChatService SHALL initialize Diagnosis Graph State with the current query context, NORMAL Planner mode, zero independent retry counters, and an empty orchestration event list.
|
||||
|
||||
#### Scenario: Initial state is projected
|
||||
|
||||
- **WHEN** a complex Chat run enters Planner for the first time
|
||||
- **THEN** `diagnosis_context` SHALL contain the current query/original query
|
||||
- **AND** complete conversation history SHALL NOT be stored in parent Graph State
|
||||
- **AND** history MAY remain in the Planner and Executor system Prompt assembled for this request
|
||||
|
||||
#### Scenario: Counters are initialized
|
||||
|
||||
- **WHEN** the Graph starts
|
||||
- **THEN** planner, verifier, composer, and evidence retry counters SHALL each be zero
|
||||
- **AND** Planner mode SHALL be NORMAL
|
||||
|
||||
### Requirement: Graph final state SHALL map to the existing Chat result and Run lifecycle
|
||||
|
||||
The system SHALL use non-empty `final_answer` from a handled Graph terminal state as the existing ChatResult answer. Run status SHALL express execution lifecycle rather than diagnosis quality.
|
||||
|
||||
#### Scenario: Composer completes
|
||||
|
||||
- **WHEN** Graph reaches Composer and produces a safe non-empty final answer
|
||||
- **THEN** ChatResult SHALL preserve the current answer/sessionId/runId protocol
|
||||
- **AND** DiagnosisRun SHALL be saved as SUCCESS with answer, duration, token count, step count, and tool count
|
||||
- **AND** EvaluationService SHALL evaluate the current runId
|
||||
|
||||
#### Scenario: Handled Fallback completes
|
||||
|
||||
- **WHEN** Graph reaches deterministic Fallback and produces a safe non-empty answer
|
||||
- **THEN** DiagnosisRun SHALL be SUCCESS
|
||||
- **AND** diagnosis quality SHALL be expressed by verifier fields when available or `orchestrationTrace.degraded=true`
|
||||
- **AND** no DEGRADED Run status SHALL be introduced
|
||||
|
||||
#### Scenario: Graph cannot produce a safe response
|
||||
|
||||
- **WHEN** Graph has an unhandled failure, final state is unavailable, final answer is blank, or required successful-result persistence fails
|
||||
- **THEN** DiagnosisRun SHALL be marked FAILED when it can still be saved
|
||||
- **AND** the system SHALL NOT report a successful Graph result
|
||||
|
||||
### Requirement: Graph state SHALL be the only source for verifier evaluation persistence
|
||||
|
||||
The system SHALL build `diagnosis_run.self_evaluation.verifier_evaluation` from explicit Graph State and SHALL NOT read VerifierContextHolder on the complex Chat path.
|
||||
|
||||
#### Scenario: Verifier completed
|
||||
|
||||
- **WHEN** final Graph State contains a completed Verifier result
|
||||
- **THEN** verifier evaluation SHALL include execution status, model verdict, effective verdict, groundedness, claim/fact checks, rationale, round, Gatekeeper audit, verified Executor output/evidence, Prompt audit, and Composer audit when available
|
||||
- **AND** the compatibility `verdict` field SHALL equal effective verdict
|
||||
|
||||
#### Scenario: Pre-verification Fallback completed
|
||||
|
||||
- **WHEN** Graph reaches Fallback before Verifier completes
|
||||
- **THEN** verifier evaluation SHALL preserve available execution statuses, Gatekeeper audit, Prompt audit, and fallback context
|
||||
- **AND** it SHALL NOT fabricate a model or effective verdict
|
||||
|
||||
#### Scenario: Evaluation payload is inspected
|
||||
|
||||
- **WHEN** verifier evaluation is persisted
|
||||
- **THEN** it SHALL NOT contain raw Executor text or complete tool trace summary
|
||||
- **AND** `executor_structured_output` SHALL contain at most the verified projection retained for compatibility
|
||||
|
||||
### Requirement: Stage 3 verification SHALL not run the final live E2E
|
||||
|
||||
The change SHALL use focused automated tests for production cutover and contracts while reserving Maven live startup, log inspection, and database querying for stage 5.
|
||||
|
||||
#### Scenario: Stage 3 is accepted
|
||||
|
||||
- **WHEN** stage 3 verification completes
|
||||
- **THEN** Graph cutover, Run/Trace, Prompt, Controller/Eval regressions, Maven test compilation, and OpenSpec strict validation SHALL have passed
|
||||
- **AND** live E2E, `logs/`, and `scripts/query_mysql.py` SHALL be recorded as intentionally deferred to stage 5
|
||||
Reference in New Issue
Block a user