## Context Current state after stage two: ```text chat_planner -> chat_executor -> VerifierInputHook + Gatekeeper -> chat_verifier -> ChatService final rendering ``` Verifier receives explicit inputs: - `original_query` - `executor_final_answer` - `executor_structured_output` - `executor_output_parse_status` - `tool_trace_summary` - `gatekeeper_result` - `retry_context` However, Verifier output is still primarily: ```json { "verdict": "PASS", "groundedness_score": 0.8, "critical_fact_count": 1, "facts_checked": [], "rationale": "..." } ``` Stage three introduces V2 output while preserving the old compatibility field. ## Verifier V2 Output Verifier should output: ```json { "verdict": "LOW_CONFID", "groundedness_score": 0.62, "critical_fact_count": 1, "claim_checks": [ { "claim_id": "claim-1", "claim_text": "payment-service CPU usage is high", "claim_type": "symptom", "verification": "direct_observation", "detail": "query_metrics shows CPU=92%", "evidence_refs": [ { "trace_ref": "trace-1", "tool_name": "query_metrics", "source_invocation_ids": [394], "note": "metrics summary contains CPU=92%" } ] } ], "hypothesis_checks": [], "facts_checked": [], "rationale": "..." } ``` `facts_checked` remains for compatibility. If Verifier does not emit it, `ChatService` must derive it from `claim_checks`. ## Claim Verification Set `claim_checks[].verification` is limited to: | Value | Meaning | Legacy mapping | |---|---|---| | `direct_observation` | Evidence directly observes the claim | `direct_evidence` | | `reasonable_inference` | Evidence can reasonably support the claim, but not as direct observation | `indirect_support` | | `overstated` | Evidence partially supports the claim, but the claim says too much | `indirect_support` | | `unsupported` | Evidence is insufficient | `no_evidence` | | `external_unknown` | Claim introduces evidence-external entity/value/root cause | `no_evidence` | | `contradicted` | Claim conflicts with evidence | `contradicted` | ## Compatibility Mapping `ChatService` must keep old downstream behavior alive by producing `facts_checked`. Suggested mapping: ```text facts_checked[].fact = "{claim_id}: {claim_text}" facts_checked[].is_critical = claim_type in ["root_cause", "symptom", "impact", "risk"] facts_checked[].verification = mapped legacy verification facts_checked[].detail = claim_checks[].detail facts_checked[].evidence_refs = claim_checks[].evidence_refs ``` If Verifier emits both `claim_checks` and `facts_checked`, `claim_checks` is authoritative. `facts_checked` may be replaced by the deterministic compatibility projection to avoid inconsistent audit data. If Verifier emits only old `facts_checked`, ChatService keeps the old path. ## Verdict Guardrails Gatekeeper fail: - If `gatekeeper_result.status = fail`, effective verdict must not be `PASS`. - If the model returns `PASS`, ChatService should downgrade the effective verdict to `LOW_CONFID` or `REJECT`. - For this phase, `evidence.invocation_ref` failure should downgrade to `REJECT`; other Gatekeeper failures should downgrade to `LOW_CONFID`. Malformed or missing structured output: - If `executor_output_parse_status.status` is `missing` or `malformed`, Verifier should not use natural-language extraction to produce PASS. - Effective verdict should be `LOW_CONFID`. ## Prompt Boundary The prompt should say: - Primary target is `executor_structured_output.claims`. - Do not extract additional confirmed facts from `executor_final_answer` when structured output is valid. - `executor_final_answer` is debug/fallback only. - `claim_checks` is the primary output. - `facts_checked` is compatibility output. ## Interface Impact - L2 internal contract extension. - No external API change. - No database schema change. - Audit JSON gains `claim_checks`. ## Risks / Mitigations - Risk: old low-confidence templates rely on `facts_checked`. - Mitigation: derive `facts_checked` from `claim_checks`. - Risk: prompt-only Gatekeeper PASS prevention is insufficient. - Mitigation: add code-side effective verdict guard. - Risk: Agent logging only summarizes `facts_checked`. - Mitigation: update logging to understand `claim_checks` while keeping old summary compatibility.