8.1 KiB
single-react-chat-sse-cutover Specification
Purpose
TBD - created by archiving change single-react-chat-sse-cutover. Update Purpose after archive.
Requirements
Requirement: Chat SHALL expose one SSE endpoint
The public Chat API SHALL expose only POST /api/chat with text/event-stream. The synchronous JSON Chat path and POST /api/chat_stream MUST be removed without a compatibility branch.
Scenario: Accepted Chat request
- WHEN a client posts a non-blank Chat query
- THEN
/api/chatreturns an SSE emitter backed only byChatApplicationUseCase
Scenario: Removed legacy endpoint
- WHEN a client posts to
/api/chat_stream - THEN no Chat handler is mapped to that endpoint
Scenario: Invalid request
- WHEN a client posts a blank Chat query
- THEN the server rejects it before creating a Run or starting an SSE event sequence
Requirement: SSE events SHALL follow one strict state machine
Every accepted Chat stream SHALL emit named events in the order metadata -> status* -> content|failure -> done. Metadata and done MUST occur exactly once, status MAY occur zero or more times, content and failure MUST be mutually exclusive and each MUST occur at most once.
Scenario: Successful diagnosis
- WHEN a Diagnosis Run safely releases a report
- THEN the stream emits one metadata event, zero or more safe status events, one content event, and one SUCCESS done event in order
Scenario: Safe fallback
- WHEN the release boundary returns a Fallback
- THEN the stream emits one content event containing only the fixed Fallback followed by one FALLBACK done event
Scenario: Technical failure
- WHEN routing or Harness execution cannot produce safe content
- THEN the stream emits one stable failure event followed by one FAILED done event and emits no content
Requirement: SSE payloads SHALL be typed and safe
Metadata SHALL contain only session_id and run_id; status SHALL contain only a fixed code and safe message; content SHALL contain content_type plus the stage 6A typed public payload; failure SHALL contain a stable code and safe message; done SHALL contain only SUCCESS, FALLBACK, or FAILED outcome. The stream MUST NOT expose prompts, thoughts, raw Tool data, complete evidence, internal exceptions, vendor errors, stack traces, unverified Drafts, or internal guard reasons.
Scenario: Metadata identity
- WHEN the use case starts one Run
- THEN the metadata IDs exactly match the IDs persisted and returned by that same Run
Scenario: Content release boundary
- WHEN status events are emitted before SemanticGuard completes
- THEN no diagnosis content is emitted until
ChatApplicationUseCasereturns released typed content
Scenario: Failure sanitization
- WHEN an internal exception contains provider or persistence details
- THEN the failure event contains only the stable application code and public message
Requirement: Final content SHALL not be disguised token streaming
The Chat stream SHALL send process status while work is running and SHALL release final safe content in one content event. It MUST NOT split a completed answer by character or fixed chunk size.
Scenario: Long final report
- WHEN a safe final report is larger than the legacy chunk size
- THEN it is sent as one typed content event rather than multiple content chunks
Requirement: Client disconnect SHALL cancel the exact Run
SSE completion before terminal release, timeout, error, or send failure SHALL request CLIENT_DISCONNECT cancellation through the ChatRunControl belonging to the same metadata IDs. A disconnect that occurs before onStarted MUST be remembered and applied when the control becomes available. No event SHALL be sent after disconnection.
Scenario: Disconnect during model work
- WHEN the client disconnects after metadata while a model call is pending
- THEN the same Run is cancelled, late content is discarded, and no content/failure/done is attempted
Scenario: Disconnect before Run control publication
- WHEN a disconnect is observed before the application observer receives
onStarted - THEN the observer cancels that Run immediately when its control is published
Scenario: Normal completion callback
- WHEN content or failure and done have completed normally
- THEN the emitter completion callback does not change the already terminal Run outcome
Requirement: Controller SHALL remain a bounded protocol adapter
The Controller SHALL validate request shape, admit work to a Spring-managed bounded executor, map application observer/results to SSE, and manage connection lifecycle. It MUST NOT select or invoke ChatModel, ToolCallbackProvider, ChatService routing, Session history, Diagnosis Agent, Guards, or Tool logic, and MUST NOT construct an executor.
Scenario: Worker saturation
- WHEN the bounded Chat worker rejects a request before a Run starts
- THEN the endpoint returns a stable unavailable HTTP response and does not run work on the request thread
Scenario: Dependency inspection
- WHEN Chat Controller dependencies and source are inspected
- THEN only protocol/application collaborators are present for Chat and no model, Tool, old strategy, history, or executor construction remains
Requirement: Production SHALL assemble one Harness graph
Spring production configuration SHALL assemble one shared Core, ChatModel boundary, canonical store, Tool boundary, three evidence adapters, Diagnosis Agent, EvidenceGuard, repair, SemanticGuard, release use case, three fixed executors, JPA Run store, Router, and Chat Application Use Case. All Run paths SHALL share the same Core and model configuration.
Scenario: Application context wiring
- WHEN a focused Spring context loads the Harness Chat configuration with controlled dependencies
- THEN exactly one
ChatApplicationUseCasegraph is created without missing or ambiguous beans
Scenario: MySQL Tool is not configured
- WHEN no logical MySQL datasource is configured
- THEN the query_mysql Tool remains bounded and fails closed for unknown data sources without using the application database
Requirement: Runtime limits and executors SHALL be centralized and bounded
Run budgets, timeouts, byte limits, canonical TTL/capacity, worker sizing and model executor sizing SHALL come from harness.chat configuration with positive validated defaults. Worker and model executors SHALL use bounded queues and rejection policies and SHALL be shut down by Spring.
Scenario: Configuration defaults
- WHEN the production properties load without overrides
- THEN every Harness limit is positive, canonical limits preserve their invariants, and both executor queues have finite capacity
Scenario: Model executor saturation
- WHEN the model executor cannot accept another submitted model call
- THEN the operation fails as a controlled technical failure and does not create an unbounded queue or caller-runs execution
Requirement: Bundled frontend SHALL consume only the new Chat SSE contract
The bundled frontend SHALL send every Chat message to /api/chat, parse complete named SSE frames, store metadata IDs for Trace/feedback, render safe typed content once, surface stable failure, and finish only after a valid done event. It MUST NOT retain quick/stream mode selection, synchronous Chat JSON parsing, /chat_stream, legacy message-wrapper fallback, or raw non-JSON content fallback.
Scenario: Frontend success
- WHEN the browser receives metadata, statuses, one typed content and SUCCESS done
- THEN it records the exact IDs, updates progress, renders the payload once, and marks the message complete
Scenario: Frontend protocol violation
- WHEN the browser receives an unknown, duplicate, out-of-order, mutually conflicting, or malformed event
- THEN it fails closed and does not render the payload as trusted assistant content
Scenario: Frontend request target
- WHEN static Chat consumer source is inspected
- THEN it contains one
/chatstreaming request and no/chat_streamor synchronous Chat consumer