# single-react-diagnosis-agent Specification ## Purpose 定义内部单体 Diagnosis ReactAgent 的极简输入、框架 ReAct/tool loop、Harness 模型与 Tool 边界、结构化 DiagnosisDraft、无证据停止、确定性预算和审计隔离要求。 ## Requirements ### Requirement: Diagnosis execution SHALL use one framework ReactAgent The internal diagnosis path SHALL create exactly one `diagnosis_agent` with the framework-provided ReAct/tool loop. It SHALL NOT create Planner, Executor, Verifier, Composer, SequentialAgent, SupervisorAgent, an outer business Graph, or a hand-written model/Tool loop. One use-case invocation SHALL call the Diagnosis Agent once and SHALL NOT wrap the Agent or any individual model call in Harness retry. #### Scenario: Diagnosis needs Tool evidence - **WHEN** the model emits one or more Tool actions before its final response - **THEN** the same Diagnosis ReactAgent executes those actions through its framework loop and returns one final Draft without invoking another Agent #### Scenario: Agent execution fails - **WHEN** the single Diagnosis Agent invocation throws or returns invalid structured output - **THEN** the internal use case fails closed without automatically invoking the Agent or model again ### Requirement: Diagnosis input SHALL contain only current Query and optional PreviousTurn The internal use case SHALL accept a non-blank current Query and an optional frozen `PreviousTurn`, serialize them as the fixed `query` and `previous_turn` input fields, and SHALL NOT load or accept complete Session history, Redis memory, prior raw Tool results, or model-generated history summaries. The current Query SHALL retain its original text and SHALL NOT be rewritten or silently truncated. #### Scenario: Follow-up diagnosis has a safe previous turn - **WHEN** a caller supplies a `PreviousTurn` - **THEN** the Agent receives exactly the frozen previous-turn fields plus the current Query and no complete conversation history #### Scenario: Query exceeds configured context budget - **WHEN** the current Query exceeds its UTF-8 byte limit - **THEN** the use case rejects it before any model call instead of truncating or rewriting it ### Requirement: Every model round SHALL be controlled by RunContext The Diagnosis Agent SHALL receive `RunContext` explicitly and SHALL use a model interceptor to call `DiagnosisHarnessCore.beforeModelCall` for every framework ReAct model round. Non-streaming model response Usage SHALL be recorded into the same Run budget when available. Cancellation, deadline, model-call exhaustion, or Token exhaustion SHALL prevent subsequent controlled work. #### Scenario: ReAct performs two model rounds - **WHEN** one model round requests a Tool and the next produces the Draft - **THEN** the same Run budget records two model calls and no Harness retry attempt #### Scenario: Model-call budget is exhausted - **WHEN** the framework attempts a model round beyond the configured maximum - **THEN** the call is rejected before reaching ChatModel and the Run records budget exhaustion ### Requirement: Evidence Tools SHALL execute through the Harness boundary The Diagnosis Agent SHALL expose only the frozen `lookup_knowledge`, `query_logs`, and `query_mysql` definitions. A Tool interceptor SHALL propagate the exact framework Tool Call ID, Run ID, Tool name, and raw JSON arguments into the corresponding stage 3B/3C adapter and `ToolBoundary`. Successful observations SHALL contain only the bounded `agent_result`; failed observations SHALL contain only stable error semantics and SHALL NOT contain raw responses, internal exceptions, credentials, or invocation lifecycle internals. #### Scenario: Framework requests RAG evidence - **WHEN** the model calls `lookup_knowledge` with framework ID `call-1` - **THEN** the RAG adapter and Agent observation use exactly `call-1`, and the canonical invocation is owned by the current Run #### Scenario: Tool execution fails - **WHEN** a registered adapter returns an error result - **THEN** the Agent receives `evidence_status=ERROR`, the framework Tool Call ID and a stable error code without automatic Tool retry or raw failure detail #### Scenario: Unknown Tool is requested - **WHEN** a model requests a Tool outside the three registered definitions - **THEN** the Harness does not authorize or emulate it and does not create a canonical evidence record ### Requirement: Diagnosis output SHALL be a bounded DiagnosisDraft The Agent SHALL receive the generated schema for `DiagnosisDraft` and SHALL return JSON that the internal use case strictly parses into the frozen record. The use case SHALL enforce configured UTF-8 limits for query, previous turn, total input and Draft output and account accepted input/output bytes against the Run capacity. It SHALL reject blank, fenced, prefixed, malformed, oversized or schema-incompatible output without repair or retry. #### Scenario: Supported Draft is returned - **WHEN** the Agent returns valid JSON containing Conclusion, Analysis items, Action Plan, Recommendations and Limitations - **THEN** the use case returns a `DiagnosisDraft` whose Analysis Tool Call IDs are the exact strings emitted by the Agent #### Scenario: Model returns prose around JSON - **WHEN** the final response contains a Markdown fence or explanatory prefix around an otherwise valid object - **THEN** strict parsing fails and the Diagnosis Agent is not invoked a second time ### Requirement: Insufficient evidence SHALL terminate without a fabricated conclusion The single Prompt SHALL require every normal Analysis item to cite current-Run evidence Tool Call IDs and SHALL restrict `NO_EVIDENCE` to scoped `NEGATIVE_OBSERVATION`. If current evidence cannot support a diagnosis, the Agent SHALL stop the current ReAct execution with `conclusion=null`, describe the actual scope and missing information in `limitations`, and SHALL NOT infer that the problem does not exist or fabricate a root cause. #### Scenario: Tool finds no evidence - **WHEN** the only completed Tool observation has `evidence_status=NO_EVIDENCE` - **THEN** the final Draft has no confirmed Conclusion, records the bounded negative observation and missing information, and makes no additional automatic retry ### Requirement: Stage-four execution SHALL remain internal and auditable The new use case SHALL be callable only as an internal Java/test entry in this stage and SHALL NOT be wired into public Chat or AIOps controllers. It SHALL propagate the same sessionId/runId through RunnableConfig metadata, allow the existing AgentStep Hook to be injected, retain ToolBoundary canonical invocation audit, and leave final Run success/persistence ownership to later Guard/application stages. #### Scenario: Internal audited execution completes - **WHEN** an injected audit Hook observes a model round and a Tool executes - **THEN** both observe the same RunContext sessionId/runId and the public Chat call path remains unchanged #### Scenario: Stage four is archived - **WHEN** focused Agent tests pass and the change is archived - **THEN** `ChatController`, public SSE behavior and the old multi-Agent implementation remain available for stages 5, 6A and 6B