Files

5.0 KiB

MODIFIED Requirements

Requirement: Evidence Tools SHALL execute through the Harness boundary

The Diagnosis Agent SHALL expose only the frozen lookup_knowledge, query_logs, and query_mysql definitions through native Tool Calling. Each Agent-facing Tool input SHALL be a typed Envelope containing optional previous_observation and required business input. A Tool interceptor SHALL validate and consume the previous observation, propagate the exact framework Tool Call ID, Run ID, Tool name, and unwrapped business JSON into the corresponding adapter and ToolBoundary, and SHALL enforce Run progress before execution. Successful observations SHALL use a bounded per-Tool whitelist projection; failed or control observations SHALL contain only stable safe semantics and SHALL NOT contain raw responses, internal exceptions, credentials, invocation lifecycle internals, counters, thresholds, or remaining budget.

Scenario: Framework requests first RAG evidence

  • WHEN the model calls lookup_knowledge with framework ID call-1, no pending evaluation and a typed business input
  • THEN the RAG adapter receives only the unwrapped business request, uses exactly call-1, and the canonical invocation is owned by the current Run

Scenario: Model continues after a non-empty result

  • WHEN the last successful Tool result is pending semantic evaluation and the model requests another Tool
  • THEN the Envelope must identify that exact prior Tool Call and contain GAINED or NO_GAIN before the new business Tool can execute

Scenario: Tool execution fails

  • WHEN a registered adapter returns an error result
  • THEN the Agent receives evidence_status=ERROR, the framework Tool Call ID and a stable error code without automatic Tool retry or raw failure detail

Scenario: Unknown Tool is requested

  • WHEN a model requests a Tool outside the three registered definitions
  • THEN the Harness does not authorize or emulate it and does not create a canonical evidence record

Requirement: Insufficient evidence SHALL terminate without a fabricated conclusion

The single Chinese Prompt SHALL require every normal Analysis item to cite current-Run evidence Tool Call IDs and SHALL restrict NO_EVIDENCE to scoped NEGATIVE_OBSERVATION. The Agent SHALL NOT be required to find a root cause. If current evidence cannot support a diagnosis, no required context exists for a bounded Tool call, or available results are correct but do not advance any diagnosis hypothesis, the Agent SHALL stop with conclusion=null, describe actual scope and missing information in limitations, and SHALL NOT infer that the problem does not exist, fabricate a root cause, or make equivalent Tool calls merely to show activity.

Scenario: Tool finds no evidence

  • WHEN completed Tool observations have no diagnostic information gain
  • THEN the final Draft has no confirmed Conclusion, records bounded checked scope and missing information, and does not make an equivalent retry

Scenario: No bounded Tool query is possible

  • WHEN the Query lacks required enterprise, time, service or error context
  • THEN the Agent may perform zero Tool calls and returns a no-conclusion Draft whose limitations.missing_info identifies the required context

Scenario: Harness requires stop

  • WHEN the Agent receives STOP_REQUIRED
  • THEN it emits a bounded final Draft without another Tool call

ADDED Requirements

Requirement: Diagnosis execution SHALL preserve controlled stop outcomes

The internal Agent use case SHALL distinguish a valid Draft, controlled information saturation, and budget termination from an unclassified Agent failure. It SHALL return a bounded internal execution result containing optional Draft, ProgressSnapshot and stop reason, and SHALL NOT convert a recognized controlled stop into DiagnosisAgentOutputException.

Scenario: Model ignores STOP_REQUIRED

  • WHEN the framework surfaces the typed collection-stopped signal after the final completion opportunity
  • THEN the Agent use case returns no Draft with INFORMATION_SATURATED and the current ProgressSnapshot

Scenario: Unclassified framework failure

  • WHEN Agent execution throws an exception unrelated to controlled stop, cancellation, or budget termination
  • THEN execution still fails closed and no safe progress is fabricated

Scenario: Invalid final Draft after verified checks

  • WHEN the final model text is empty or violates the strict DiagnosisDraft contract after the current Run has completed READY canonical Tool checks
  • THEN the invalid text is discarded, the output failure carries only the bounded ProgressSnapshot and safe failure metadata, and no model repair or loose JSON extraction occurs

Scenario: Invalid final Draft without verified checks

  • WHEN the final model text violates the strict DiagnosisDraft contract before any publishable ProgressSnapshot exists
  • THEN execution remains failed and MUST NOT fabricate missing context, observed facts or a no-conclusion Draft