refactor(harness): remove legacy agent architecture

This commit is contained in:
zhuyongxin
2026-07-22 18:02:01 +08:00
parent bc36248cd8
commit 8ee7cc0b70
148 changed files with 3091 additions and 13889 deletions
@@ -10,22 +10,52 @@ Define a fail-closed, parameterized, read-only MySQL evidence Tool that reuses T
The Tool SHALL parse exactly one SQL statement with JSqlParser and SHALL accept only a single `SELECT` with explicit projection columns, supported predicates/grouping/ordering, `INNER JOIN` or `LEFT JOIN`, parameter placeholders and allowlisted functions. It SHALL reject writes, CTEs, subqueries, set operations, wildcard projections except `COUNT(*)`, metadata discovery, unsupported joins/functions, `FOR UPDATE`, multiple statements and unknown/ambiguous AST structures.
#### Scenario: Reject unsupported SQL before execution
- **WHEN** an agent submits a write statement, a second statement, a subquery, or an unallowlisted function
- **THEN** validation fails closed and no JDBC execution is attempted
### Requirement: Data-source and identifier authorization SHALL use exact independent allowlists
The Tool SHALL accept only a logical `data_source` ID and SHALL authorize every schema, table and column used in projection, join, predicate, grouping and ordering against the configured exact allowlist. It SHALL reject unknown data sources, schemas, tables, columns, aliases and ambiguous unqualified columns. Agent input SHALL NOT provide JDBC coordinates or authorization controls.
#### Scenario: Reject an identifier outside the configured allowlist
- **WHEN** a valid-looking query references an unknown data source, table, column, alias, or ambiguous unqualified column
- **THEN** authorization fails before connection creation and agent-provided JDBC coordinates are ignored
### Requirement: Parameter binding and JDBC execution SHALL be read-only and bounded
The executor SHALL use a configured logical datasource, a read-only JDBC connection, `PreparedStatement` parameter binding, query timeout, max rows and Run cancellation/deadline checks. Placeholder count SHALL exactly match `params`. The executor SHALL not expose connection details or raw JDBC failures to the Agent.
#### Scenario: Execute an authorized query within read-only bounds
- **WHEN** an authorized query has exactly matching parameters and an active Run
- **THEN** it executes through a read-only prepared statement with timeout and row limits, while cancellation or a deadline stops execution and raw JDBC details remain hidden
### Requirement: MySQL projection SHALL be bounded and evidence-aware
The projector SHALL expose only ordered columns, bounded JSON-safe rows, returned count and truncation. It SHALL enforce row, cell and total UTF-8 limits, redact sensitive column values, return `NO_EVIDENCE` for a successful empty result, and never expose raw JDBC metadata or credentials.
#### Scenario: Project bounded rows and empty evidence safely
- **WHEN** JDBC returns rows containing oversized or sensitive values, or returns a successful empty result
- **THEN** the projection redacts and bounds values, reports truncation when data is removed, and returns `NO_EVIDENCE` for the empty result without exposing metadata or credentials
### Requirement: MySQL Tool SHALL reuse canonical Harness ownership
The adapter SHALL pass the exact framework `tool_call_id` and RunContext through the existing ToolBoundary and canonical invocation store. It SHALL not create a second ID, use a parallel store, return raw SQL results, or modify legacy audit/public runtime paths.
#### Scenario: Preserve framework invocation identity
- **WHEN** the adapter invokes an authorized MySQL query
- **THEN** ToolBoundary and the canonical invocation store receive the exact framework `tool_call_id` and RunContext, with no second identifier or parallel raw-result path
### Requirement: The query helper script SHALL be read-only and secret-free by default
The repository query helper SHALL require connection values from environment variables, reject non-SELECT and metadata discovery SQL before connection, and SHALL NOT commit writes or expose hardcoded external connection defaults.
#### Scenario: Reject unsafe helper SQL without connecting
- **WHEN** the helper receives a non-`SELECT` or metadata-discovery statement, or missing environment connection values
- **THEN** it exits before opening a connection and does not reveal or invent credentials or external connection defaults
@@ -10,14 +10,34 @@ Define bounded RAG and Mock query-log projections that execute through the stage
The RAG adapter SHALL accept the logical `query` request, execute the existing knowledge tool through ToolBoundary, and project only `RagToolResult` fields. Context packs, retrieval traces, rerank traces, scores, hit reasons, domains, messages and full document bodies SHALL NOT appear in the Agent result.
#### Scenario: Project only bounded RAG evidence
- **WHEN** the adapter receives a logical query and the knowledge backend returns a result
- **THEN** it invokes through ToolBoundary and exposes only the bounded `RagToolResult` evidence fields, excluding retrieval internals and full document bodies
### Requirement: Query-log projection SHALL preserve logical scope and Mock provenance
The query-log adapter SHALL accept only logical topic, query and optional lookback minutes, execute the existing Mock source through ToolBoundary, and project `source_kind=MOCK`, complete scope, match count, returned count, bounded patterns, bounded timeline events and truncation.
#### Scenario: Return a bounded Mock query-log projection
- **WHEN** the adapter receives a logical topic, query, and optional lookback
- **THEN** it invokes the Mock source through ToolBoundary and returns `source_kind=MOCK`, the complete logical scope, bounded matches and timeline events, counts, and truncation
### Requirement: Projection SHALL redact and bound sensitive log content
The log projector SHALL exclude instance and metrics fields and redact credentials, token-like values, host/pod identifiers, PIDs, IP addresses, SQL literals and stack-like suffixes from Agent-facing messages. It SHALL enforce per-item, collection and total UTF-8 bounds and set `truncated=true` when any bound removes data.
#### Scenario: Redact sensitive log fields and mark truncation
- **WHEN** a log result contains credentials, identifiers, SQL literals, stack suffixes, or content beyond configured UTF-8 bounds
- **THEN** those values are omitted or redacted and `truncated=true` indicates removed content
### Requirement: Adapters SHALL reuse canonical boundary ownership
Both adapters SHALL pass the framework `tool_call_id` and RunContext to the existing ToolBoundary and SHALL NOT create a second ID, write a parallel store, return raw responses, or modify legacy audit paths.
#### Scenario: Preserve canonical identity for both adapters
- **WHEN** either the RAG or query-log adapter executes
- **THEN** it passes the exact framework `tool_call_id` and RunContext through ToolBoundary without creating a second identifier, parallel store, raw response path, or legacy audit side effect
@@ -117,10 +117,3 @@ The bundled frontend SHALL send every Chat message to `/api/chat`, parse complet
#### Scenario: Frontend request target
- **WHEN** static Chat consumer source is inspected
- **THEN** it contains one `/chat` streaming request and no `/chat_stream` or synchronous Chat consumer
### Requirement: AiOps public behavior SHALL remain isolated
The `/api/ai_ops` URL, request schema and event behavior SHALL remain unchanged in stage 6B. Any internal model/Tool dependency movement needed to keep Chat Controller protocol-only MUST preserve existing AiOps observable behavior.
#### Scenario: AiOps regression
- **WHEN** existing AiOps Controller and service tests run after Chat cutover
- **THEN** existing metadata and analysis behavior remains compatible
@@ -0,0 +1,88 @@
# single-react-cleanup-e2e Specification
## Purpose
TBD - created by archiving change single-react-cleanup-e2e. Update Purpose after archive.
## Requirements
### Requirement: Legacy diagnosis orchestration SHALL be physically absent
Production and test source SHALL contain no legacy ChatService, AiOps Supervisor/Planner/Executor, Sequential Planner/Executor/Verifier/Composer, legacy Gatekeeper/Verifier helpers, ThreadLocal execution context, or their dedicated prompts and obsolete tests. The only business Agent with a Tool loop SHALL be the Diagnosis Agent constructed inside the Harness.
#### Scenario: Source inventory after cleanup
- **WHEN** production source, resources and tests are scanned
- **THEN** legacy orchestration classes, prompts, ThreadLocals and self-only tests are absent rather than disabled or commented out
#### Scenario: Agent construction inventory
- **WHEN** business Agent builders and framework orchestration types are inspected
- **THEN** only the Diagnosis Agent has evidence Tool callbacks and no Supervisor/Sequential business workflow remains
### Requirement: Public diagnosis SHALL use one endpoint
The application SHALL expose `POST /api/chat` as the only public diagnosis execution endpoint. `/api/ai_ops`, legacy Chat Session management endpoints and bundled frontend callers for those endpoints MUST be removed without a compatibility branch.
#### Scenario: Bundled frontend diagnosis
- **WHEN** a user submits a diagnosis from the bundled frontend
- **THEN** it uses the named-event SSE `/api/chat` consumer and exposes no AiOps mode or button
#### Scenario: Removed endpoint scan
- **WHEN** Controller mappings and frontend request targets are inspected
- **THEN** no `/api/ai_ops`, `/api/chat/clear` or `/api/chat/session` execution/management mapping remains
### Requirement: Evidence backends SHALL not be Agent contracts
RAG and log query implementations MAY be reused behind Harness adapters, but MUST NOT expose `@Tool`, `ToolCallbackProvider`, topic-discovery-first behavior, legacy Tool descriptions, Session ThreadLocal, session dedup or per-backend durable recorder side effects. Agent-facing Tool names, schemas and descriptions SHALL come only from `HarnessEvidenceTools` and `AgentToolContracts`.
#### Scenario: Tool discovery inspection
- **WHEN** Spring Tool annotations and Agent callback registration are inspected
- **THEN** only `lookup_knowledge`, `query_logs` and `query_mysql` Harness ACI callbacks are available to the Diagnosis Agent
#### Scenario: Backend invocation
- **WHEN** a Harness adapter calls RAG or Mock logs
- **THEN** the backend returns raw adapter input without reading ThreadLocal or writing a second ToolInvocation record
### Requirement: Agent durable audit SHALL be metadata-only
Every persisted Diagnosis AgentStep SHALL use the exact RunnableConfig sessionId/runId and MAY contain only bounded metadata such as message count/roles, Tool names, text presence, duration and budget counters. It MUST NOT persist or log Prompt text, message content, model output text, Tool arguments, raw evidence or Thought, and MUST NOT fall back to ThreadLocal identity.
#### Scenario: Model step persistence
- **WHEN** the Diagnosis Agent performs model calls
- **THEN** AgentStep rows use the SSE metadata identity and contain no user query, evidence body, Tool argument or chain-of-thought text
#### Scenario: Missing metadata
- **WHEN** an audit hook is invoked without exact sessionId/runId metadata
- **THEN** it skips persistence with a safe warning rather than inventing or reading implicit identity
### Requirement: ToolBoundary SHALL write safe durable audit
For each accepted Harness evidence Tool call, ToolBoundary SHALL attempt to write one durable audit row with exact sessionId/runId/toolCallId/toolName, invocation/evidence status, stable error code, duration and byte counts. Durable audit MUST NOT contain the complete request, SQL/log query body, raw response or Agent projection content. Audit persistence failure SHALL be observable but MUST NOT change the canonical Tool result.
#### Scenario: Successful Tool call
- **WHEN** ToolBoundary completes a READY invocation
- **THEN** Redis retains the canonical record and MySQL receives one metadata-only ToolInvocation row for the same Run and Tool call
#### Scenario: Failed Tool call
- **WHEN** ToolBoundary returns a stable error after an accepted request
- **THEN** the durable row records only stable status/error metadata and no internal exception or raw payload
#### Scenario: Audit database failure
- **WHEN** durable audit persistence throws after canonical result determination
- **THEN** ToolBoundary logs a safe warning and returns the unchanged canonical Tool result
### Requirement: Current documentation SHALL describe the single Harness architecture
Current MVP architecture, Agent/Harness, Trace, API and demo documents SHALL describe the single Diagnosis Agent, explicit Harness ownership, ACI Tools, named SSE and safe durable audit. ISS-012 and ISS-013 SHALL be recorded as absorbed by ISS-014; historical archived documents MAY retain historical descriptions.
#### Scenario: Current documentation scan
- **WHEN** non-archived current architecture and demo documents are inspected
- **THEN** they do not present Planner/Executor/Verifier/Composer or `/api/ai_ops` as current runtime behavior
### Requirement: Repository verification SHALL be clean
The final implementation SHALL compile, pass focused and relevant full regressions, pass JavaScript syntax and static legacy scans, and pass strict OpenSpec validation for all current specs. No dead imports, temporary files, hardcoded credentials, compatibility flags or unexplained legacy references may be introduced.
#### Scenario: Automated verification
- **WHEN** the stage 7 verification suite runs
- **THEN** all selected tests/build/static/OpenSpec checks pass with no legacy runtime token in production source
### Requirement: Final live E2E SHALL correlate exact Run evidence
The final acceptance SHALL start the Spring Boot application, send a real diagnosis request through `/api/chat`, strictly parse the named SSE sequence, inspect `logs/`, and query MySQL with `scripts/query_mysql.py` using the exact metadata sessionId/runId. It SHALL verify Run intent/outcome/status, one Diagnosis Agent identity, bounded model/Tool counts, ToolInvocation identity, safe final content/fallback and no prohibited content leakage.
#### Scenario: Live successful or safe fallback diagnosis
- **WHEN** the configured model and infrastructure process the fixed diagnosis request
- **THEN** SSE, logs, diagnosis_run, agent_step and tool_invocation evidence agree on the exact identity and final release is SUCCESS or documented safe FALLBACK
#### Scenario: External Tool boundary statement
- **WHEN** final acceptance is archived
- **THEN** it distinguishes Mock query_logs and isolated query_mysql contract evidence from unverified real CLS or production business MySQL integration