Files
SuperBizAgent-java/openspec/changes/executor-evidence-output-contract/proposal.md
T

2.9 KiB

Proposal: executor-evidence-output-contract

Why

Recent Chat diagnosis sessions frequently end as LOW_CONFID even though evidence tools were called successfully. The main failure mode is not missing tool execution; it is that Executor blends tool evidence, runbook/reference patterns, and model inference into one natural-language answer, then presents unsupported details as confirmed incident facts.

Verifier currently receives original_query, executor_final_answer, tool_trace_summary, and optional retry_context. It can catch unsupported claims, but it must first infer facts from unstructured prose. That makes the quality gate reactive and noisy: unsupported claims are detected after the answer has already been shaped as a confident story.

What Changes

  • Update the Chat Executor prompt so its final output follows a strict evidence-attribution JSON contract.
  • Require confirmed claims to carry explicit evidence bindings and require unsupported items to be placed in hypotheses, missing_info, or recommended_actions instead of confirmed conclusions.
  • Extend the verifier input contract with a parsed or raw executor_structured_output field while preserving the existing executor_final_answer field for fallback compatibility.
  • Update the Chat Verifier prompt so it prefers Executor-provided structured claims over natural-language fact extraction.
  • Keep Verifier isolated from Planner/Executor intermediate reasoning; it still receives only the original query, Executor final output, structured claim contract, tool trace summary, and retry context.
  • Do not change evidence tool signatures, database schema, or the Planner -> Executor -> Verifier topology.

Capabilities

New Capabilities

None.

Modified Capabilities

  • chat-verifier-agent: Chat verification now consumes an Executor evidence-attribution contract when available.

Impact

  • Affected prompts: chat-executor-prompt.md, chat-verifier-prompt.md.
  • Affected integration: VerifierInputHook or equivalent verifier payload assembly must include executor_structured_output when Executor returns valid structured JSON.
  • Affected parsing: ChatService may parse Executor output to separate machine contract from user-facing text, with fallback when parsing fails.
  • Affected tests/eval: prompt contract tests, verifier input assembly tests, and unsupported-claim regression fixtures.
  • No database schema change is required. Structured Executor output can be persisted in existing step output fields and verifier snapshots.

Out Of Scope

  • Loosening Verifier scoring or PASS criteria.
  • Raising verifier.low-confidence-threshold to hide unsupported claims.
  • Treating runbook, skill, or historical case text as current incident evidence unless it was returned as evidence for the current query and bound explicitly.
  • Adding a new verifier tool or allowing Verifier to perform retrieval.
  • Replacing the existing tool trace summary contract.