feat(agent): add verifier claim checks
This commit is contained in:
@@ -0,0 +1,140 @@
|
||||
## Context
|
||||
|
||||
Current state after stage two:
|
||||
|
||||
```text
|
||||
chat_planner
|
||||
-> chat_executor
|
||||
-> VerifierInputHook + Gatekeeper
|
||||
-> chat_verifier
|
||||
-> ChatService final rendering
|
||||
```
|
||||
|
||||
Verifier receives explicit inputs:
|
||||
|
||||
- `original_query`
|
||||
- `executor_final_answer`
|
||||
- `executor_structured_output`
|
||||
- `executor_output_parse_status`
|
||||
- `tool_trace_summary`
|
||||
- `gatekeeper_result`
|
||||
- `retry_context`
|
||||
|
||||
However, Verifier output is still primarily:
|
||||
|
||||
```json
|
||||
{
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 0.8,
|
||||
"critical_fact_count": 1,
|
||||
"facts_checked": [],
|
||||
"rationale": "..."
|
||||
}
|
||||
```
|
||||
|
||||
Stage three introduces V2 output while preserving the old compatibility field.
|
||||
|
||||
## Verifier V2 Output
|
||||
|
||||
Verifier should output:
|
||||
|
||||
```json
|
||||
{
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.62,
|
||||
"critical_fact_count": 1,
|
||||
"claim_checks": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_text": "payment-service CPU usage is high",
|
||||
"claim_type": "symptom",
|
||||
"verification": "direct_observation",
|
||||
"detail": "query_metrics shows CPU=92%",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"trace_ref": "trace-1",
|
||||
"tool_name": "query_metrics",
|
||||
"source_invocation_ids": [394],
|
||||
"note": "metrics summary contains CPU=92%"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypothesis_checks": [],
|
||||
"facts_checked": [],
|
||||
"rationale": "..."
|
||||
}
|
||||
```
|
||||
|
||||
`facts_checked` remains for compatibility. If Verifier does not emit it, `ChatService` must derive it from `claim_checks`.
|
||||
|
||||
## Claim Verification Set
|
||||
|
||||
`claim_checks[].verification` is limited to:
|
||||
|
||||
| Value | Meaning | Legacy mapping |
|
||||
|---|---|---|
|
||||
| `direct_observation` | Evidence directly observes the claim | `direct_evidence` |
|
||||
| `reasonable_inference` | Evidence can reasonably support the claim, but not as direct observation | `indirect_support` |
|
||||
| `overstated` | Evidence partially supports the claim, but the claim says too much | `indirect_support` |
|
||||
| `unsupported` | Evidence is insufficient | `no_evidence` |
|
||||
| `external_unknown` | Claim introduces evidence-external entity/value/root cause | `no_evidence` |
|
||||
| `contradicted` | Claim conflicts with evidence | `contradicted` |
|
||||
|
||||
## Compatibility Mapping
|
||||
|
||||
`ChatService` must keep old downstream behavior alive by producing `facts_checked`.
|
||||
|
||||
Suggested mapping:
|
||||
|
||||
```text
|
||||
facts_checked[].fact = "{claim_id}: {claim_text}"
|
||||
facts_checked[].is_critical = claim_type in ["root_cause", "symptom", "impact", "risk"]
|
||||
facts_checked[].verification = mapped legacy verification
|
||||
facts_checked[].detail = claim_checks[].detail
|
||||
facts_checked[].evidence_refs = claim_checks[].evidence_refs
|
||||
```
|
||||
|
||||
If Verifier emits both `claim_checks` and `facts_checked`, `claim_checks` is authoritative. `facts_checked` may be replaced by the deterministic compatibility projection to avoid inconsistent audit data.
|
||||
|
||||
If Verifier emits only old `facts_checked`, ChatService keeps the old path.
|
||||
|
||||
## Verdict Guardrails
|
||||
|
||||
Gatekeeper fail:
|
||||
|
||||
- If `gatekeeper_result.status = fail`, effective verdict must not be `PASS`.
|
||||
- If the model returns `PASS`, ChatService should downgrade the effective verdict to `LOW_CONFID` or `REJECT`.
|
||||
- For this phase, `evidence.invocation_ref` failure should downgrade to `REJECT`; other Gatekeeper failures should downgrade to `LOW_CONFID`.
|
||||
|
||||
Malformed or missing structured output:
|
||||
|
||||
- If `executor_output_parse_status.status` is `missing` or `malformed`, Verifier should not use natural-language extraction to produce PASS.
|
||||
- Effective verdict should be `LOW_CONFID`.
|
||||
|
||||
## Prompt Boundary
|
||||
|
||||
The prompt should say:
|
||||
|
||||
- Primary target is `executor_structured_output.claims`.
|
||||
- Do not extract additional confirmed facts from `executor_final_answer` when structured output is valid.
|
||||
- `executor_final_answer` is debug/fallback only.
|
||||
- `claim_checks` is the primary output.
|
||||
- `facts_checked` is compatibility output.
|
||||
|
||||
## Interface Impact
|
||||
|
||||
- L2 internal contract extension.
|
||||
- No external API change.
|
||||
- No database schema change.
|
||||
- Audit JSON gains `claim_checks`.
|
||||
|
||||
## Risks / Mitigations
|
||||
|
||||
- Risk: old low-confidence templates rely on `facts_checked`.
|
||||
- Mitigation: derive `facts_checked` from `claim_checks`.
|
||||
- Risk: prompt-only Gatekeeper PASS prevention is insufficient.
|
||||
- Mitigation: add code-side effective verdict guard.
|
||||
- Risk: Agent logging only summarizes `facts_checked`.
|
||||
- Mitigation: update logging to understand `claim_checks` while keeping old summary compatibility.
|
||||
|
||||
Reference in New Issue
Block a user