Files
SuperBizAgent-java/openspec/specs/single-react-chat-sse-cutover/spec.md
T

127 lines
8.5 KiB
Markdown

# single-react-chat-sse-cutover Specification
## Purpose
TBD - created by archiving change single-react-chat-sse-cutover. Update Purpose after archive.
## Requirements
### Requirement: Chat SHALL expose one SSE endpoint
The public Chat API SHALL expose only `POST /api/chat` with `text/event-stream`. The synchronous JSON Chat path and `POST /api/chat_stream` MUST be removed without a compatibility branch.
#### Scenario: Accepted Chat request
- **WHEN** a client posts a non-blank Chat query
- **THEN** `/api/chat` returns an SSE emitter backed only by `ChatApplicationUseCase`
#### Scenario: Removed legacy endpoint
- **WHEN** a client posts to `/api/chat_stream`
- **THEN** no Chat handler is mapped to that endpoint
#### Scenario: Invalid request
- **WHEN** a client posts a blank Chat query
- **THEN** the server rejects it before creating a Run or starting an SSE event sequence
### Requirement: SSE events SHALL follow one strict state machine
Every accepted Chat stream SHALL emit named events in the order `metadata -> status* -> content|failure -> done`. Metadata and done MUST occur exactly once, status MAY occur zero or more times, content and failure MUST be mutually exclusive and each MUST occur at most once.
#### Scenario: Successful diagnosis
- **WHEN** a Diagnosis Run safely releases a report
- **THEN** the stream emits one metadata event, zero or more safe status events, one content event, and one SUCCESS done event in order
#### Scenario: Safe fallback
- **WHEN** the release boundary returns a Fallback
- **THEN** the stream emits one content event containing only the fixed Fallback followed by one FALLBACK done event
#### Scenario: Technical failure
- **WHEN** routing or Harness execution cannot produce safe content
- **THEN** the stream emits one stable failure event followed by one FAILED done event and emits no content
### Requirement: SSE payloads SHALL be typed and safe
Metadata SHALL contain only `session_id` and `run_id`; status SHALL contain only a fixed code and safe message; content SHALL contain `content_type` plus the stage 6A typed public payload; failure SHALL contain a stable code and safe message; done SHALL contain only `SUCCESS`, `FALLBACK`, or `FAILED` outcome. The stream MUST NOT expose prompts, thoughts, raw Tool data, complete evidence, internal exceptions, vendor errors, stack traces, unverified Drafts, or internal guard reasons.
#### Scenario: Metadata identity
- **WHEN** the use case starts one Run
- **THEN** the metadata IDs exactly match the IDs persisted and returned by that same Run
#### Scenario: Content release boundary
- **WHEN** status events are emitted before SemanticGuard completes
- **THEN** no diagnosis content is emitted until `ChatApplicationUseCase` returns released typed content
#### Scenario: Failure sanitization
- **WHEN** an internal exception contains provider or persistence details
- **THEN** the failure event contains only the stable application code and public message
### Requirement: Final content SHALL not be disguised token streaming
The Chat stream SHALL send process status while work is running and SHALL release final safe content in one content event. It MUST NOT split a completed answer by character or fixed chunk size.
#### Scenario: Long final report
- **WHEN** a safe final report is larger than the legacy chunk size
- **THEN** it is sent as one typed content event rather than multiple content chunks
### Requirement: Client disconnect SHALL cancel the exact Run
SSE completion before terminal release, timeout, error, or send failure SHALL request CLIENT_DISCONNECT cancellation through the `ChatRunControl` belonging to the same metadata IDs. A disconnect that occurs before `onStarted` MUST be remembered and applied when the control becomes available. No event SHALL be sent after disconnection.
#### Scenario: Disconnect during model work
- **WHEN** the client disconnects after metadata while a model call is pending
- **THEN** the same Run is cancelled, late content is discarded, and no content/failure/done is attempted
#### Scenario: Disconnect before Run control publication
- **WHEN** a disconnect is observed before the application observer receives `onStarted`
- **THEN** the observer cancels that Run immediately when its control is published
#### Scenario: Normal completion callback
- **WHEN** content or failure and done have completed normally
- **THEN** the emitter completion callback does not change the already terminal Run outcome
### Requirement: Controller SHALL remain a bounded protocol adapter
The Controller SHALL validate request shape, admit work to a Spring-managed bounded executor, map application observer/results to SSE, and manage connection lifecycle. It MUST NOT select or invoke ChatModel, ToolCallbackProvider, ChatService routing, Session history, Diagnosis Agent, Guards, or Tool logic, and MUST NOT construct an executor.
#### Scenario: Worker saturation
- **WHEN** the bounded Chat worker rejects a request before a Run starts
- **THEN** the endpoint returns a stable unavailable HTTP response and does not run work on the request thread
#### Scenario: Dependency inspection
- **WHEN** Chat Controller dependencies and source are inspected
- **THEN** only protocol/application collaborators are present for Chat and no model, Tool, old strategy, history, or executor construction remains
### Requirement: Production SHALL assemble one Harness graph
Spring production configuration SHALL assemble one shared Core, ChatModel boundary, canonical store, Tool boundary, three evidence adapters, Diagnosis Agent, EvidenceGuard, repair, SemanticGuard, release use case, three fixed executors, JPA Run store, Router, and Chat Application Use Case. All Run paths SHALL share the same Core and model configuration.
#### Scenario: Application context wiring
- **WHEN** a focused Spring context loads the Harness Chat configuration with controlled dependencies
- **THEN** exactly one `ChatApplicationUseCase` graph is created without missing or ambiguous beans
#### Scenario: MySQL Tool is not configured
- **WHEN** no logical MySQL datasource is configured
- **THEN** the query_mysql Tool remains bounded and fails closed for unknown data sources without using the application database
### Requirement: Runtime limits and executors SHALL be centralized and bounded
Run budgets, timeouts, byte limits, canonical TTL/capacity, worker sizing and model executor sizing SHALL come from `harness.chat` configuration with positive validated defaults. Worker and model executors SHALL use bounded queues and rejection policies and SHALL be shut down by Spring.
#### Scenario: Configuration defaults
- **WHEN** the production properties load without overrides
- **THEN** every Harness limit is positive, canonical limits preserve their invariants, and both executor queues have finite capacity
#### Scenario: Model executor saturation
- **WHEN** the model executor cannot accept another submitted model call
- **THEN** the operation fails as a controlled technical failure and does not create an unbounded queue or caller-runs execution
### Requirement: Bundled frontend SHALL consume only the new Chat SSE contract
The bundled frontend SHALL send every Chat message to `/api/chat`, parse complete named SSE frames, store metadata IDs for Trace/feedback, render safe typed content once, surface stable failure, and finish only after a valid done event. It MUST NOT retain quick/stream mode selection, synchronous Chat JSON parsing, `/chat_stream`, legacy message-wrapper fallback, or raw non-JSON content fallback.
#### Scenario: Frontend success
- **WHEN** the browser receives metadata, statuses, one typed content and SUCCESS done
- **THEN** it records the exact IDs, updates progress, renders the payload once, and marks the message complete
#### Scenario: Frontend protocol violation
- **WHEN** the browser receives an unknown, duplicate, out-of-order, mutually conflicting, or malformed event
- **THEN** it fails closed and does not render the payload as trusted assistant content
#### Scenario: Frontend request target
- **WHEN** static Chat consumer source is inspected
- **THEN** it contains one `/chat` streaming request and no `/chat_stream` or synchronous Chat consumer
### Requirement: AiOps public behavior SHALL remain isolated
The `/api/ai_ops` URL, request schema and event behavior SHALL remain unchanged in stage 6B. Any internal model/Tool dependency movement needed to keep Chat Controller protocol-only MUST preserve existing AiOps observable behavior.
#### Scenario: AiOps regression
- **WHEN** existing AiOps Controller and service tests run after Chat cutover
- **THEN** existing metadata and analysis behavior remains compatible