5.0 KiB
MODIFIED Requirements
Requirement: Evidence Tools SHALL execute through the Harness boundary
The Diagnosis Agent SHALL expose only the frozen lookup_knowledge, query_logs, and query_mysql definitions through native Tool Calling. Each Agent-facing Tool input SHALL be a typed Envelope containing optional previous_observation and required business input. A Tool interceptor SHALL validate and consume the previous observation, propagate the exact framework Tool Call ID, Run ID, Tool name, and unwrapped business JSON into the corresponding adapter and ToolBoundary, and SHALL enforce Run progress before execution. Successful observations SHALL use a bounded per-Tool whitelist projection; failed or control observations SHALL contain only stable safe semantics and SHALL NOT contain raw responses, internal exceptions, credentials, invocation lifecycle internals, counters, thresholds, or remaining budget.
Scenario: Framework requests first RAG evidence
- WHEN the model calls
lookup_knowledgewith framework IDcall-1, no pending evaluation and a typed business input - THEN the RAG adapter receives only the unwrapped business request, uses exactly
call-1, and the canonical invocation is owned by the current Run
Scenario: Model continues after a non-empty result
- WHEN the last successful Tool result is pending semantic evaluation and the model requests another Tool
- THEN the Envelope must identify that exact prior Tool Call and contain
GAINEDorNO_GAINbefore the new business Tool can execute
Scenario: Tool execution fails
- WHEN a registered adapter returns an error result
- THEN the Agent receives
evidence_status=ERROR, the framework Tool Call ID and a stable error code without automatic Tool retry or raw failure detail
Scenario: Unknown Tool is requested
- WHEN a model requests a Tool outside the three registered definitions
- THEN the Harness does not authorize or emulate it and does not create a canonical evidence record
Requirement: Insufficient evidence SHALL terminate without a fabricated conclusion
The single Chinese Prompt SHALL require every normal Analysis item to cite current-Run evidence Tool Call IDs and SHALL restrict NO_EVIDENCE to scoped NEGATIVE_OBSERVATION. The Agent SHALL NOT be required to find a root cause. If current evidence cannot support a diagnosis, no required context exists for a bounded Tool call, or available results are correct but do not advance any diagnosis hypothesis, the Agent SHALL stop with conclusion=null, describe actual scope and missing information in limitations, and SHALL NOT infer that the problem does not exist, fabricate a root cause, or make equivalent Tool calls merely to show activity.
Scenario: Tool finds no evidence
- WHEN completed Tool observations have no diagnostic information gain
- THEN the final Draft has no confirmed Conclusion, records bounded checked scope and missing information, and does not make an equivalent retry
Scenario: No bounded Tool query is possible
- WHEN the Query lacks required enterprise, time, service or error context
- THEN the Agent may perform zero Tool calls and returns a no-conclusion Draft whose
limitations.missing_infoidentifies the required context
Scenario: Harness requires stop
- WHEN the Agent receives
STOP_REQUIRED - THEN it emits a bounded final Draft without another Tool call
ADDED Requirements
Requirement: Diagnosis execution SHALL preserve controlled stop outcomes
The internal Agent use case SHALL distinguish a valid Draft, controlled information saturation, and budget termination from an unclassified Agent failure. It SHALL return a bounded internal execution result containing optional Draft, ProgressSnapshot and stop reason, and SHALL NOT convert a recognized controlled stop into DiagnosisAgentOutputException.
Scenario: Model ignores STOP_REQUIRED
- WHEN the framework surfaces the typed collection-stopped signal after the final completion opportunity
- THEN the Agent use case returns no Draft with
INFORMATION_SATURATEDand the current ProgressSnapshot
Scenario: Unclassified framework failure
- WHEN Agent execution throws an exception unrelated to controlled stop, cancellation, or budget termination
- THEN execution still fails closed and no safe progress is fabricated
Scenario: Invalid final Draft after verified checks
- WHEN the final model text is empty or violates the strict DiagnosisDraft contract after the current Run has completed READY canonical Tool checks
- THEN the invalid text is discarded, the output failure carries only the bounded ProgressSnapshot and safe failure metadata, and no model repair or loose JSON extraction occurs
Scenario: Invalid final Draft without verified checks
- WHEN the final model text violates the strict DiagnosisDraft contract before any publishable ProgressSnapshot exists
- THEN execution remains failed and MUST NOT fabricate missing context, observed facts or a no-conclusion Draft