Files
SuperBizAgent-java/openspec/specs/chat-diagnosis-stategraph-chatservice-cutover/spec.md
T

5.7 KiB

chat-diagnosis-stategraph-chatservice-cutover Specification

Purpose

TBD - created by archiving change chat-diagnosis-stategraph-chatservice-cutover. Update Purpose after archive.

Requirements

Requirement: Complex Chat SHALL use the real Diagnosis StateGraph as its only production orchestrator

The system SHALL execute each complex Chat request through one real Diagnosis CompiledGraph and SHALL NOT create, invoke, or fall back to a SequentialAgent workflow.

Scenario: Complex Chat starts

  • WHEN executeChatWithStrategy classifies a valid question as complex
  • THEN ChatService SHALL create the current Diagnosis Run and invoke one Diagnosis CompiledGraph
  • AND the production path SHALL NOT maintain an outer score-based retry loop

Scenario: Graph dependencies are assembled

  • WHEN the complex Chat Graph is constructed
  • THEN Planner, Executor, Verifier, and Composer SHALL use the existing project Prompt, tool, skill, and AgentLoggingHook assembly rules
  • AND the Graph Verifier SHALL NOT register VerifierInputHook
  • AND Gatekeeper SHALL run only as the explicit Graph Node

Requirement: Graph invocation SHALL preserve current Run ownership

The system SHALL use the current runId as Graph threadId and SHALL pass both current sessionId and runId in RunnableConfig metadata.

Scenario: Agent Node runs

  • WHEN any Agent Node is invoked for a complex Chat run
  • THEN its RunnableConfig threadId SHALL equal the current runId
  • AND AgentStep and ToolInvocation writes SHALL retain the current sessionId and runId

Scenario: Gatekeeper validates Executor output

  • WHEN the Gatekeeper Node runs
  • THEN it SHALL validate only tool invocations belonging to the runId in RunnableConfig metadata
  • AND data from another run in the same session SHALL NOT be considered

Requirement: Initial Graph State SHALL contain only bounded diagnosis control data

ChatService SHALL initialize Diagnosis Graph State with the current query context, NORMAL Planner mode, zero independent retry counters, and an empty orchestration event list.

Scenario: Initial state is projected

  • WHEN a complex Chat run enters Planner for the first time
  • THEN diagnosis_context SHALL contain the current query/original query
  • AND complete conversation history SHALL NOT be stored in parent Graph State
  • AND history MAY remain in the Planner and Executor system Prompt assembled for this request

Scenario: Counters are initialized

  • WHEN the Graph starts
  • THEN planner, verifier, composer, and evidence retry counters SHALL each be zero
  • AND Planner mode SHALL be NORMAL

Requirement: Graph final state SHALL map to the existing Chat result and Run lifecycle

The system SHALL use non-empty final_answer from a handled Graph terminal state as the existing ChatResult answer. Run status SHALL express execution lifecycle rather than diagnosis quality.

Scenario: Composer completes

  • WHEN Graph reaches Composer and produces a safe non-empty final answer
  • THEN ChatResult SHALL preserve the current answer/sessionId/runId protocol
  • AND DiagnosisRun SHALL be saved as SUCCESS with answer, duration, token count, step count, and tool count
  • AND EvaluationService SHALL evaluate the current runId

Scenario: Handled Fallback completes

  • WHEN Graph reaches deterministic Fallback and produces a safe non-empty answer
  • THEN DiagnosisRun SHALL be SUCCESS
  • AND diagnosis quality SHALL be expressed by verifier fields when available or orchestrationTrace.degraded=true
  • AND no DEGRADED Run status SHALL be introduced

Scenario: Graph cannot produce a safe response

  • WHEN Graph has an unhandled failure, final state is unavailable, final answer is blank, or required successful-result persistence fails
  • THEN DiagnosisRun SHALL be marked FAILED when it can still be saved
  • AND the system SHALL NOT report a successful Graph result

Requirement: Graph state SHALL be the only source for verifier evaluation persistence

The system SHALL build diagnosis_run.self_evaluation.verifier_evaluation from explicit Graph State and SHALL NOT read VerifierContextHolder on the complex Chat path.

Scenario: Verifier completed

  • WHEN final Graph State contains a completed Verifier result
  • THEN verifier evaluation SHALL include execution status, model verdict, effective verdict, groundedness, claim/fact checks, rationale, round, Gatekeeper audit, verified Executor output/evidence, Prompt audit, and Composer audit when available
  • AND the compatibility verdict field SHALL equal effective verdict

Scenario: Pre-verification Fallback completed

  • WHEN Graph reaches Fallback before Verifier completes
  • THEN verifier evaluation SHALL preserve available execution statuses, Gatekeeper audit, Prompt audit, and fallback context
  • AND it SHALL NOT fabricate a model or effective verdict

Scenario: Evaluation payload is inspected

  • WHEN verifier evaluation is persisted
  • THEN it SHALL NOT contain raw Executor text or complete tool trace summary
  • AND executor_structured_output SHALL contain at most the verified projection retained for compatibility

Requirement: Stage 3 verification SHALL not run the final live E2E

The change SHALL use focused automated tests for production cutover and contracts while reserving Maven live startup, log inspection, and database querying for stage 5.

Scenario: Stage 3 is accepted

  • WHEN stage 3 verification completes
  • THEN Graph cutover, Run/Trace, Prompt, Controller/Eval regressions, Maven test compilation, and OpenSpec strict validation SHALL have passed
  • AND live E2E, logs/, and scripts/query_mysql.py SHALL be recorded as intentionally deferred to stage 5