Files
SuperBizAgent-java/openspec/specs/single-react-cleanup-e2e/spec.md
T

6.8 KiB

single-react-cleanup-e2e Specification

Purpose

TBD - created by archiving change single-react-cleanup-e2e. Update Purpose after archive.

Requirements

Requirement: Legacy diagnosis orchestration SHALL be physically absent

Production and test source SHALL contain no legacy ChatService, AiOps Supervisor/Planner/Executor, Sequential Planner/Executor/Verifier/Composer, legacy Gatekeeper/Verifier helpers, ThreadLocal execution context, or their dedicated prompts and obsolete tests. The only business Agent with a Tool loop SHALL be the Diagnosis Agent constructed inside the Harness.

Scenario: Source inventory after cleanup

  • WHEN production source, resources and tests are scanned
  • THEN legacy orchestration classes, prompts, ThreadLocals and self-only tests are absent rather than disabled or commented out

Scenario: Agent construction inventory

  • WHEN business Agent builders and framework orchestration types are inspected
  • THEN only the Diagnosis Agent has evidence Tool callbacks and no Supervisor/Sequential business workflow remains

Requirement: Public diagnosis SHALL use one endpoint

The application SHALL expose POST /api/chat as the only public diagnosis execution endpoint. /api/ai_ops, legacy Chat Session management endpoints and bundled frontend callers for those endpoints MUST be removed without a compatibility branch.

Scenario: Bundled frontend diagnosis

  • WHEN a user submits a diagnosis from the bundled frontend
  • THEN it uses the named-event SSE /api/chat consumer and exposes no AiOps mode or button

Scenario: Removed endpoint scan

  • WHEN Controller mappings and frontend request targets are inspected
  • THEN no /api/ai_ops, /api/chat/clear or /api/chat/session execution/management mapping remains

Requirement: Evidence backends SHALL not be Agent contracts

RAG and log query implementations MAY be reused behind Harness adapters, but MUST NOT expose @Tool, ToolCallbackProvider, topic-discovery-first behavior, legacy Tool descriptions, Session ThreadLocal, session dedup or per-backend durable recorder side effects. Agent-facing Tool names, schemas and descriptions SHALL come only from HarnessEvidenceTools and AgentToolContracts.

Scenario: Tool discovery inspection

  • WHEN Spring Tool annotations and Agent callback registration are inspected
  • THEN only lookup_knowledge, query_logs and query_mysql Harness ACI callbacks are available to the Diagnosis Agent

Scenario: Backend invocation

  • WHEN a Harness adapter calls RAG or Mock logs
  • THEN the backend returns raw adapter input without reading ThreadLocal or writing a second ToolInvocation record

Requirement: Agent durable audit SHALL be metadata-only

Every persisted Diagnosis AgentStep SHALL use the exact RunnableConfig sessionId/runId and MAY contain only bounded metadata such as message count/roles, Tool names, text presence, duration and budget counters. It MUST NOT persist or log Prompt text, message content, model output text, Tool arguments, raw evidence or Thought, and MUST NOT fall back to ThreadLocal identity.

Scenario: Model step persistence

  • WHEN the Diagnosis Agent performs model calls
  • THEN AgentStep rows use the SSE metadata identity and contain no user query, evidence body, Tool argument or chain-of-thought text

Scenario: Missing metadata

  • WHEN an audit hook is invoked without exact sessionId/runId metadata
  • THEN it skips persistence with a safe warning rather than inventing or reading implicit identity

Requirement: ToolBoundary SHALL write safe durable audit

For each accepted Harness evidence Tool call, ToolBoundary SHALL attempt to write one durable audit row with exact sessionId/runId/toolCallId/toolName, invocation/evidence status, stable error code, duration and byte counts. Durable audit MUST NOT contain the complete request, SQL/log query body, raw response or Agent projection content. Audit persistence failure SHALL be observable but MUST NOT change the canonical Tool result.

Scenario: Successful Tool call

  • WHEN ToolBoundary completes a READY invocation
  • THEN Redis retains the canonical record and MySQL receives one metadata-only ToolInvocation row for the same Run and Tool call

Scenario: Failed Tool call

  • WHEN ToolBoundary returns a stable error after an accepted request
  • THEN the durable row records only stable status/error metadata and no internal exception or raw payload

Scenario: Audit database failure

  • WHEN durable audit persistence throws after canonical result determination
  • THEN ToolBoundary logs a safe warning and returns the unchanged canonical Tool result

Requirement: Current documentation SHALL describe the single Harness architecture

Current MVP architecture, Agent/Harness, Trace, API and demo documents SHALL describe the single Diagnosis Agent, explicit Harness ownership, ACI Tools, named SSE and safe durable audit. ISS-012 and ISS-013 SHALL be recorded as absorbed by ISS-014; historical archived documents MAY retain historical descriptions.

Scenario: Current documentation scan

  • WHEN non-archived current architecture and demo documents are inspected
  • THEN they do not present Planner/Executor/Verifier/Composer or /api/ai_ops as current runtime behavior

Requirement: Repository verification SHALL be clean

The final implementation SHALL compile, pass focused and relevant full regressions, pass JavaScript syntax and static legacy scans, and pass strict OpenSpec validation for all current specs. No dead imports, temporary files, hardcoded credentials, compatibility flags or unexplained legacy references may be introduced.

Scenario: Automated verification

  • WHEN the stage 7 verification suite runs
  • THEN all selected tests/build/static/OpenSpec checks pass with no legacy runtime token in production source

Requirement: Final live E2E SHALL correlate exact Run evidence

The final acceptance SHALL start the Spring Boot application, send a real diagnosis request through /api/chat, strictly parse the named SSE sequence, inspect logs/, and query MySQL with scripts/query_mysql.py using the exact metadata sessionId/runId. It SHALL verify Run intent/outcome/status, one Diagnosis Agent identity, bounded model/Tool counts, ToolInvocation identity, safe final content/fallback and no prohibited content leakage.

Scenario: Live successful or safe fallback diagnosis

  • WHEN the configured model and infrastructure process the fixed diagnosis request
  • THEN SSE, logs, diagnosis_run, agent_step and tool_invocation evidence agree on the exact identity and final release is SUCCESS or documented safe FALLBACK

Scenario: External Tool boundary statement

  • WHEN final acceptance is archived
  • THEN it distinguishes Mock query_logs and isolated query_mysql contract evidence from unverified real CLS or production business MySQL integration