## MODIFIED Requirements ### Requirement: Evidence Tools SHALL execute through the Harness boundary The Diagnosis Agent SHALL expose only the frozen `lookup_knowledge`, `query_logs`, and `query_mysql` definitions through native Tool Calling. Each Agent-facing Tool input SHALL be a typed Envelope containing optional `previous_observation` and required business `input`. A Tool interceptor SHALL validate and consume the previous observation, propagate the exact framework Tool Call ID, Run ID, Tool name, and unwrapped business JSON into the corresponding adapter and `ToolBoundary`, and SHALL enforce Run progress before execution. Successful observations SHALL use a bounded per-Tool whitelist projection; failed or control observations SHALL contain only stable safe semantics and SHALL NOT contain raw responses, internal exceptions, credentials, invocation lifecycle internals, counters, thresholds, or remaining budget. #### Scenario: Framework requests first RAG evidence - **WHEN** the model calls `lookup_knowledge` with framework ID `call-1`, no pending evaluation and a typed business input - **THEN** the RAG adapter receives only the unwrapped business request, uses exactly `call-1`, and the canonical invocation is owned by the current Run #### Scenario: Model continues after a non-empty result - **WHEN** the last successful Tool result is pending semantic evaluation and the model requests another Tool - **THEN** the Envelope must identify that exact prior Tool Call and contain `GAINED` or `NO_GAIN` before the new business Tool can execute #### Scenario: Tool execution fails - **WHEN** a registered adapter returns an error result - **THEN** the Agent receives `evidence_status=ERROR`, the framework Tool Call ID and a stable error code without automatic Tool retry or raw failure detail #### Scenario: Unknown Tool is requested - **WHEN** a model requests a Tool outside the three registered definitions - **THEN** the Harness does not authorize or emulate it and does not create a canonical evidence record ### Requirement: Insufficient evidence SHALL terminate without a fabricated conclusion The single Chinese Prompt SHALL require every normal Analysis item to cite current-Run evidence Tool Call IDs and SHALL restrict `NO_EVIDENCE` to scoped `NEGATIVE_OBSERVATION`. The Agent SHALL NOT be required to find a root cause. If current evidence cannot support a diagnosis, no required context exists for a bounded Tool call, or available results are correct but do not advance any diagnosis hypothesis, the Agent SHALL stop with `conclusion=null`, describe actual scope and missing information in `limitations`, and SHALL NOT infer that the problem does not exist, fabricate a root cause, or make equivalent Tool calls merely to show activity. #### Scenario: Tool finds no evidence - **WHEN** completed Tool observations have no diagnostic information gain - **THEN** the final Draft has no confirmed Conclusion, records bounded checked scope and missing information, and does not make an equivalent retry #### Scenario: No bounded Tool query is possible - **WHEN** the Query lacks required enterprise, time, service or error context - **THEN** the Agent may perform zero Tool calls and returns a no-conclusion Draft whose `limitations.missing_info` identifies the required context #### Scenario: Harness requires stop - **WHEN** the Agent receives `STOP_REQUIRED` - **THEN** it emits a bounded final Draft without another Tool call ## ADDED Requirements ### Requirement: Diagnosis execution SHALL preserve controlled stop outcomes The internal Agent use case SHALL distinguish a valid Draft, controlled information saturation, and budget termination from an unclassified Agent failure. It SHALL return a bounded internal execution result containing optional Draft, ProgressSnapshot and stop reason, and SHALL NOT convert a recognized controlled stop into `DiagnosisAgentOutputException`. #### Scenario: Model ignores STOP_REQUIRED - **WHEN** the framework surfaces the typed collection-stopped signal after the final completion opportunity - **THEN** the Agent use case returns no Draft with `INFORMATION_SATURATED` and the current ProgressSnapshot #### Scenario: Unclassified framework failure - **WHEN** Agent execution throws an exception unrelated to controlled stop, cancellation, or budget termination - **THEN** execution still fails closed and no safe progress is fabricated #### Scenario: Invalid final Draft after verified checks - **WHEN** the final model text is empty or violates the strict DiagnosisDraft contract after the current Run has completed READY canonical Tool checks - **THEN** the invalid text is discarded, the output failure carries only the bounded ProgressSnapshot and safe failure metadata, and no model repair or loose JSON extraction occurs #### Scenario: Invalid final Draft without verified checks - **WHEN** the final model text violates the strict DiagnosisDraft contract before any publishable ProgressSnapshot exists - **THEN** execution remains failed and MUST NOT fabricate missing context, observed facts or a no-conclusion Draft