Files
SuperBizAgent-java/openspec/changes/executor-evidence-output-contract/design.md
T

7.8 KiB

Context

Current Chat flow explicitly runs Planner -> Executor -> Verifier. VerifierInputHook builds a structured verifier payload containing:

  • original_query
  • executor_final_answer
  • tool_trace_summary
  • retry_context

This satisfies the earlier design goal that Verifier should not see Planner/Executor intermediate reasoning. However, Executor's answer is still natural language. Verifier must infer claims from prose, then compare those claims to tool traces. Recent low-confidence sessions show that Executor often converts weak or unrelated evidence into confirmed conclusions before Verifier sees it.

The fix should move evidence attribution earlier: Executor must state which claims are confirmed, which evidence supports them, and which points remain hypotheses or gaps.

Goals / Non-Goals

Goals:

  • Make Executor final output machine-checkable.
  • Separate confirmed claims, hypotheses, recommended actions, and missing information.
  • Make every confirmed claim bind to concrete tool evidence identifiers or excerpts.
  • Let Verifier consume structured claims directly instead of reconstructing all facts from prose.
  • Preserve Verifier input isolation from intermediate reasoning.
  • Preserve existing behavior when Executor output is malformed by falling back to natural-language verification and low-confidence handling.

Non-Goals:

  • Do not add database tables or columns.
  • Do not change evidence tool method signatures.
  • Do not let Verifier call tools.
  • Do not require full raw tool outputs in verifier input.
  • Do not make user-facing answers become raw JSON-only unless the product layer explicitly chooses that rendering.

Decisions

Decision: Executor emits an evidence-attribution contract

Executor final output SHALL be a single JSON object. The recommended contract is:

{
  "answer_version": "executor_evidence_v1",
  "diagnosis_summary": "1-2 sentence summary using only supported facts",
  "claims": [
    {
      "claim_id": "claim-1",
      "claim_type": "root_cause",
      "claim_text": "The payment-service connection pool is saturated",
      "support_level": "direct",
      "evidence_bindings": [
        {
          "source_type": "tool_trace",
          "source_id": "trace-1",
          "tool_name": "query_metrics",
          "source_invocation_ids": [101],
          "evidence_excerpt": "active connections reached max pool size"
        }
      ]
    }
  ],
  "hypotheses": [
    {
      "hypothesis_text": "A connection leak may be contributing",
      "basis": "metrics show saturation, but no leak evidence was returned",
      "needed_evidence": ["connection lifetime metrics", "leak detection logs"]
    }
  ],
  "recommended_actions": [
    {
      "action_text": "Check HikariCP active/idle/pending connection metrics",
      "reason": "Needed to confirm pool saturation scope",
      "evidence_bindings": []
    }
  ],
  "missing_info": [
    "No log evidence confirming a connection leak was returned"
  ],
  "user_facing_answer": "Chinese answer rendered from the same confirmed claims, hypotheses, actions, and gaps"
}

Rationale: this keeps the machine contract explicit while still allowing the product to return a readable Chinese answer.

Decision: Evidence binding uses generic source fields

chunk_id alone is too specific to lookup_knowledge. The binding shape SHALL support all evidence-bearing tools:

  • source_type: tool_trace, lookup_evidence_block, or another stable source family
  • source_id: tool_trace_summary.trace_ref, evidence block id, or equivalent stable id
  • tool_name: evidence tool name when available
  • source_invocation_ids: persisted tool_invocation ids when available
  • evidence_excerpt: bounded excerpt copied or summarized from tool evidence

Rationale: query_logs, query_metrics, and lookup_knowledge expose evidence differently. A generic binding prevents the prompt contract from overfitting to RAG chunks.

Decision: Confirmed claims are stricter than hypotheses and actions

Confirmed claims SHALL contain only facts supported by current-session evidence. Runbook instructions, skill guidance, and historical case patterns SHALL NOT appear as confirmed claims unless the current tool trace supports that exact fact.

Unsupported but useful diagnostic ideas SHALL be placed under hypotheses or recommended_actions.

Rationale: this directly addresses the observed hallucination: turning plausible patterns into current incident facts.

Decision: Verifier prefers structured claims but keeps fallback

Verifier prompt SHALL use this order:

  1. If executor_structured_output.claims is valid, verify each claim directly.
  2. Also scan user_facing_answer for extra confirmed-sounding facts not present in claims; mark them as facts to check.
  3. If the structured output is missing or invalid, fall back to the existing natural-language extraction from executor_final_answer.

Invalid structured output SHALL NOT crash the flow. It SHOULD produce a fallback verifier decision and make malformed structure visible in verifier_evaluation.

Rationale: the new contract should improve precision without making runtime brittle.

Decision: Verifier still does not see intermediate reasoning

The verifier payload MAY add:

  • executor_structured_output
  • executor_output_parse_status

It SHALL continue to exclude Planner reasoning, Executor intermediate reasoning, and raw conversation noise.

Rationale: this preserves the original verifier design: verify the final answer and tool traces, not hidden reasoning.

Decision: User-facing output remains Chinese

For Chat user-facing pages and answers, the rendered answer SHALL be Chinese. If the Executor emits JSON, either:

  • user_facing_answer is returned to the user after verifier routing, or
  • ChatService renders a Chinese answer from the structured contract.

The raw machine contract may remain visible in trace/debug views, but the normal user answer should not become an English/JSON-only artifact.

Rationale: this aligns with current UI/product language requirements while keeping machine-checkable evidence.

Risks / Trade-offs

  • [Risk] LLM may emit malformed JSON. Mitigation: parse-status fallback and verifier natural-language fallback.
  • [Risk] Executor may put unsupported facts in user_facing_answer but omit them from claims. Mitigation: Verifier scans user_facing_answer for extra confirmed-sounding facts.
  • [Risk] Evidence excerpts may be fabricated. Mitigation: Verifier checks binding ids against tool_trace_summary and treats missing refs as no evidence.
  • [Risk] More verbose output increases tokens. Mitigation: keep excerpts bounded and put detailed raw evidence only in tool traces.
  • [Risk] Prompt-only enforcement is soft. Mitigation: add focused tests/eval fixtures and later consider code-level schema validation.

Migration Plan

  1. Update chat-executor-prompt.md with the evidence-attribution JSON contract.
  2. Add parsing in Chat runtime for Executor output:
    • valid JSON -> preserve parsed executor_structured_output
    • invalid JSON -> mark parse status and keep raw executor_final_answer
  3. Update verifier payload assembly to include executor_structured_output and parse status.
  4. Update chat-verifier-prompt.md to prefer structured claims and verify extra facts in user_facing_answer.
  5. Persist parsed or raw structured output in existing trace/evaluation snapshots without schema changes.
  6. Add focused tests and regression fixtures for unsupported confirmed claims.
  7. Validate OpenSpec and run targeted tests.

Open Questions

  • Should user_facing_answer be mandatory in Executor JSON, or should ChatService render it from structured fields?
  • Should malformed Executor JSON force LOW_CONFID, or should Verifier decide based on natural-language fallback alone?
  • Should support_level allow only direct, indirect, and none, or should it also include contradicted for Executor self-reporting?