feat(agent): add verifier claim checks
This commit is contained in:
@@ -0,0 +1 @@
|
||||
ready
|
||||
@@ -0,0 +1 @@
|
||||
committed
|
||||
@@ -0,0 +1,2 @@
|
||||
schema: spec-driven
|
||||
created: 2026-07-07
|
||||
@@ -0,0 +1,110 @@
|
||||
# Decisions: executor-verifier-claim-checks
|
||||
|
||||
## sm-flow Progress
|
||||
|
||||
### Clarify
|
||||
|
||||
Entry summary: implement stage three of Executor Structured Output V2 by making Verifier V2 claim-oriented.
|
||||
|
||||
Slug: `executor-verifier-claim-checks`
|
||||
|
||||
Scale: standard. This changes internal verifier output/audit contracts and parsing logic, but does not change external APIs or database schema.
|
||||
|
||||
### Context
|
||||
|
||||
Relevant history:
|
||||
|
||||
- `executor-v2-output-contract`: Executor emits `executor_evidence_v2` and no longer emits final-expression fields.
|
||||
- `executor-gatekeeper-hook`: Gatekeeper validates schema and invocation references before Verifier and persists `gatekeeper_result`.
|
||||
- `chat-verifier-agent`: Current Verifier still primarily uses `facts_checked`.
|
||||
|
||||
Current code shape:
|
||||
|
||||
- `chat-verifier-prompt.md` still frames the task around `executor_final_answer` and `facts_checked`.
|
||||
- `ChatService.parseVerifierDecision(...)` only parses `facts_checked`.
|
||||
- `buildLowConfidenceOutput(...)` and `buildRetryContext(...)` consume `VerifierDecision.factsChecked()`.
|
||||
- `persistVerifierEvaluation(...)` already persists parse status, structured output, trace summary, and gatekeeper result.
|
||||
- `AgentLoggingHook` summarizes verifier output using `facts_checked`.
|
||||
|
||||
### Grill
|
||||
|
||||
Question pool:
|
||||
|
||||
| Question | Mode | Resolution |
|
||||
|---|---|---|
|
||||
| Should Verifier still scan `executor_final_answer` when structured output is valid? | evidence-driven | No. Stage three explicitly makes claims the primary target and raw output debug/fallback only. |
|
||||
| Should `facts_checked` be removed now? | evidence-driven | No. It remains a compatibility projection for low-confidence templates, retry context, eval, and trace tooling. |
|
||||
| Should Gatekeeper fail prevention be prompt-only? | evidence-driven | No. Stage two noted prompt compliance is not deterministic; stage three adds code-side effective verdict guard. |
|
||||
| Does this require database migration? | evidence-driven | No. `claim_checks` is stored under existing JSON self_evaluation. |
|
||||
| Does this introduce Composer? | evidence-driven | No. Composer is stage four. |
|
||||
|
||||
No user-interview questions are open for this stage.
|
||||
|
||||
### Specify
|
||||
|
||||
OpenSpec artifacts:
|
||||
|
||||
- `proposal.md`: why and scope.
|
||||
- `design.md`: Verifier V2 output, compatibility mapping, verdict guardrails, risks.
|
||||
- `specs/chat-verifier-agent/spec.md`: observable requirements for claim checks, compatibility facts, persistence, and guardrails.
|
||||
- `tasks.md`: executable implementation and verification checklist.
|
||||
|
||||
### Audit
|
||||
|
||||
Architecture risk summary:
|
||||
|
||||
- This is an L2 internal verifier contract extension.
|
||||
- Existing consumers continue using `facts_checked`, which is now generated from `claim_checks` when present.
|
||||
- Gatekeeper PASS prevention becomes deterministic in code, reducing reliance on prompt compliance.
|
||||
- No external API or database schema changes are introduced.
|
||||
|
||||
Cross-artifact alignment:
|
||||
|
||||
| Source | Target | Status |
|
||||
|---|---|---|
|
||||
| issue stage three | proposal | aligned |
|
||||
| proposal scope / non-goals | design | aligned |
|
||||
| design output contract and guardrails | specs | aligned |
|
||||
| specs observable behavior | tasks | aligned |
|
||||
|
||||
Interface impact:
|
||||
|
||||
- Verifier output: L2 internal extension with `claim_checks`.
|
||||
- Persistence JSON: L2 internal audit extension in existing `self_evaluation`.
|
||||
- External HTTP/API behavior: unchanged.
|
||||
|
||||
### Commit
|
||||
|
||||
Commit gate result: passed.
|
||||
|
||||
- `proposal.md`, `design.md`, `specs/chat-verifier-agent/spec.md`, and `tasks.md` exist.
|
||||
- `cmd /c openspec validate executor-verifier-claim-checks` passed.
|
||||
- `cmd /c openspec status --change executor-verifier-claim-checks` reports 4/4 artifacts complete.
|
||||
- No unresolved user-interview questions remain for this stage.
|
||||
|
||||
### Apply
|
||||
|
||||
Capability source: OpenSpec CLI + `openspec-apply-change` protocol, executed through the local shell tool. No separate semantic/LSP tools are available in this session, so implementation evidence used OpenSpec artifacts, `rg`/diff inspection, and targeted tests.
|
||||
|
||||
Implemented changes:
|
||||
|
||||
- Updated `chat-verifier-prompt.md` so `executor_structured_output.claims` is the primary verification target.
|
||||
- Added `claim_checks` parsing, normalization, persistence, and compatibility mapping to legacy `facts_checked` in `ChatService`.
|
||||
- Added code-side effective verdict guardrails:
|
||||
- missing/malformed structured output cannot remain `PASS`;
|
||||
- `gatekeeper_result.status=fail` cannot remain `PASS`;
|
||||
- `evidence.invocation_ref` failures downgrade to `REJECT`;
|
||||
- other Gatekeeper failures downgrade at least to `LOW_CONFID`.
|
||||
- Updated verifier logging summaries to account for `claim_checks`.
|
||||
- Updated sequential workflow tests to use valid Executor V2 output for PASS paths and to verify downgrade paths for Gatekeeper failure and malformed output.
|
||||
|
||||
Conflict / fix record:
|
||||
|
||||
- Initial targeted Maven verification failed because legacy tests still expected PASS to return Executor natural-language output or V1 `user_facing_answer`.
|
||||
- Classification: code/test drift from the committed OpenSpec, not a design blocker.
|
||||
- Resolution: updated tests to assert the stage-three contract: valid V2 structured output may PASS through the temporary renderer, while missing/malformed/V1-style output cannot produce effective PASS through natural-language fallback.
|
||||
|
||||
Verification:
|
||||
|
||||
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test` passed.
|
||||
- `cmd /c openspec validate executor-verifier-claim-checks` passed.
|
||||
@@ -0,0 +1,140 @@
|
||||
## Context
|
||||
|
||||
Current state after stage two:
|
||||
|
||||
```text
|
||||
chat_planner
|
||||
-> chat_executor
|
||||
-> VerifierInputHook + Gatekeeper
|
||||
-> chat_verifier
|
||||
-> ChatService final rendering
|
||||
```
|
||||
|
||||
Verifier receives explicit inputs:
|
||||
|
||||
- `original_query`
|
||||
- `executor_final_answer`
|
||||
- `executor_structured_output`
|
||||
- `executor_output_parse_status`
|
||||
- `tool_trace_summary`
|
||||
- `gatekeeper_result`
|
||||
- `retry_context`
|
||||
|
||||
However, Verifier output is still primarily:
|
||||
|
||||
```json
|
||||
{
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 0.8,
|
||||
"critical_fact_count": 1,
|
||||
"facts_checked": [],
|
||||
"rationale": "..."
|
||||
}
|
||||
```
|
||||
|
||||
Stage three introduces V2 output while preserving the old compatibility field.
|
||||
|
||||
## Verifier V2 Output
|
||||
|
||||
Verifier should output:
|
||||
|
||||
```json
|
||||
{
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.62,
|
||||
"critical_fact_count": 1,
|
||||
"claim_checks": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_text": "payment-service CPU usage is high",
|
||||
"claim_type": "symptom",
|
||||
"verification": "direct_observation",
|
||||
"detail": "query_metrics shows CPU=92%",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"trace_ref": "trace-1",
|
||||
"tool_name": "query_metrics",
|
||||
"source_invocation_ids": [394],
|
||||
"note": "metrics summary contains CPU=92%"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypothesis_checks": [],
|
||||
"facts_checked": [],
|
||||
"rationale": "..."
|
||||
}
|
||||
```
|
||||
|
||||
`facts_checked` remains for compatibility. If Verifier does not emit it, `ChatService` must derive it from `claim_checks`.
|
||||
|
||||
## Claim Verification Set
|
||||
|
||||
`claim_checks[].verification` is limited to:
|
||||
|
||||
| Value | Meaning | Legacy mapping |
|
||||
|---|---|---|
|
||||
| `direct_observation` | Evidence directly observes the claim | `direct_evidence` |
|
||||
| `reasonable_inference` | Evidence can reasonably support the claim, but not as direct observation | `indirect_support` |
|
||||
| `overstated` | Evidence partially supports the claim, but the claim says too much | `indirect_support` |
|
||||
| `unsupported` | Evidence is insufficient | `no_evidence` |
|
||||
| `external_unknown` | Claim introduces evidence-external entity/value/root cause | `no_evidence` |
|
||||
| `contradicted` | Claim conflicts with evidence | `contradicted` |
|
||||
|
||||
## Compatibility Mapping
|
||||
|
||||
`ChatService` must keep old downstream behavior alive by producing `facts_checked`.
|
||||
|
||||
Suggested mapping:
|
||||
|
||||
```text
|
||||
facts_checked[].fact = "{claim_id}: {claim_text}"
|
||||
facts_checked[].is_critical = claim_type in ["root_cause", "symptom", "impact", "risk"]
|
||||
facts_checked[].verification = mapped legacy verification
|
||||
facts_checked[].detail = claim_checks[].detail
|
||||
facts_checked[].evidence_refs = claim_checks[].evidence_refs
|
||||
```
|
||||
|
||||
If Verifier emits both `claim_checks` and `facts_checked`, `claim_checks` is authoritative. `facts_checked` may be replaced by the deterministic compatibility projection to avoid inconsistent audit data.
|
||||
|
||||
If Verifier emits only old `facts_checked`, ChatService keeps the old path.
|
||||
|
||||
## Verdict Guardrails
|
||||
|
||||
Gatekeeper fail:
|
||||
|
||||
- If `gatekeeper_result.status = fail`, effective verdict must not be `PASS`.
|
||||
- If the model returns `PASS`, ChatService should downgrade the effective verdict to `LOW_CONFID` or `REJECT`.
|
||||
- For this phase, `evidence.invocation_ref` failure should downgrade to `REJECT`; other Gatekeeper failures should downgrade to `LOW_CONFID`.
|
||||
|
||||
Malformed or missing structured output:
|
||||
|
||||
- If `executor_output_parse_status.status` is `missing` or `malformed`, Verifier should not use natural-language extraction to produce PASS.
|
||||
- Effective verdict should be `LOW_CONFID`.
|
||||
|
||||
## Prompt Boundary
|
||||
|
||||
The prompt should say:
|
||||
|
||||
- Primary target is `executor_structured_output.claims`.
|
||||
- Do not extract additional confirmed facts from `executor_final_answer` when structured output is valid.
|
||||
- `executor_final_answer` is debug/fallback only.
|
||||
- `claim_checks` is the primary output.
|
||||
- `facts_checked` is compatibility output.
|
||||
|
||||
## Interface Impact
|
||||
|
||||
- L2 internal contract extension.
|
||||
- No external API change.
|
||||
- No database schema change.
|
||||
- Audit JSON gains `claim_checks`.
|
||||
|
||||
## Risks / Mitigations
|
||||
|
||||
- Risk: old low-confidence templates rely on `facts_checked`.
|
||||
- Mitigation: derive `facts_checked` from `claim_checks`.
|
||||
- Risk: prompt-only Gatekeeper PASS prevention is insufficient.
|
||||
- Mitigation: add code-side effective verdict guard.
|
||||
- Risk: Agent logging only summarizes `facts_checked`.
|
||||
- Mitigation: update logging to understand `claim_checks` while keeping old summary compatibility.
|
||||
|
||||
@@ -0,0 +1,47 @@
|
||||
## Why
|
||||
|
||||
Stage one moved Executor to `executor_evidence_v2`, and stage two added deterministic Gatekeeper checks before Verifier. The Verifier still mainly operates through the legacy `facts_checked` contract and the prompt still allows fallback extraction from `executor_final_answer`.
|
||||
|
||||
That keeps two problems alive:
|
||||
|
||||
- Verifier can still treat natural-language Executor output as a fact source.
|
||||
- Downstream code cannot distinguish claim-level verification results from legacy natural-language fact checks.
|
||||
|
||||
This phase makes Verifier V2 claim-oriented: Verifier evaluates `executor_structured_output.claims` for whether each claim can be reasonably derived from evidence, emits `claim_checks`, and keeps `facts_checked` only as a compatibility projection.
|
||||
|
||||
## What Changes
|
||||
|
||||
- Update `chat-verifier-prompt.md` so the primary verification target is `executor_structured_output.claims`.
|
||||
- Add Verifier V2 output field `claim_checks`.
|
||||
- Preserve compatibility by mapping `claim_checks` into legacy `facts_checked`.
|
||||
- Parse and persist `claim_checks` in `ChatService`.
|
||||
- Ensure `gatekeeper_result.status=fail` cannot result in an effective `PASS`.
|
||||
- Ensure missing or malformed Executor structured output does not fall back to natural-language fact extraction for PASS.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### New Capabilities
|
||||
|
||||
None.
|
||||
|
||||
### Modified Capabilities
|
||||
|
||||
- `chat-verifier-agent`: Verifier output now includes claim-level checks and uses claim derivability as the primary groundedness contract.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affected prompt: `src/main/resources/prompts/chat-verifier-prompt.md`.
|
||||
- Affected service: `ChatService.parseVerifierDecision(...)`, retry context generation, verifier persistence.
|
||||
- Affected audit: `diagnosis_session.self_evaluation.verifier_evaluation` gains `claim_checks` and keeps `facts_checked`.
|
||||
- Affected logging: verifier thought summaries may count `claim_checks`.
|
||||
- Affected tests: `ChatServiceSequentialAgentTest` and focused verifier parsing tests.
|
||||
- Database schema: no table or column change.
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No Composer in this phase.
|
||||
- No final-answer material filtering in this phase beyond existing LOW_CONFID/REJECT templates and temporary V2 renderer.
|
||||
- No Executor retry behavior change.
|
||||
- No Gatekeeper rule expansion.
|
||||
- No database schema migration.
|
||||
|
||||
+109
@@ -0,0 +1,109 @@
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Verifier SHALL fact-check Executor answers
|
||||
The system SHALL have a Verifier Agent that reads structured Executor claims and the tool call history, then produces a structured verdict based on claim derivability.
|
||||
|
||||
#### Scenario: PASS verdict when all claims have evidence
|
||||
- **WHEN** all critical claims in `executor_structured_output.claims` have direct observation or reasonable inference support in tool call results
|
||||
- **AND** at least one critical claim has direct observation
|
||||
- **AND** no critical claim is contradicted, unsupported, external unknown, or overstated
|
||||
- **AND** `gatekeeper_result.status` is not `fail`
|
||||
- **THEN** the Verifier MAY output verdict="PASS" with groundedness_score ≥ 0.5
|
||||
|
||||
#### Scenario: LOW_CONFID verdict with partial evidence
|
||||
- **WHEN** no critical claim contradicts the tool results
|
||||
- **AND** some critical claims are `unsupported`, `external_unknown`, or `overstated`
|
||||
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
|
||||
|
||||
#### Scenario: LOW_CONFID verdict with only inference support
|
||||
- **WHEN** no critical claim contradicts the tool results
|
||||
- **AND** all critical claims are only `reasonable_inference`
|
||||
- **THEN** the Verifier SHALL output verdict="LOW_CONFID"
|
||||
|
||||
#### Scenario: REJECT verdict when claims contradict evidence
|
||||
- **WHEN** any critical claim in `executor_structured_output.claims` contradicts tool call results
|
||||
- **OR** the claim fabricates a key entity, error code, or conclusion that does not exist in the tool evidence
|
||||
- **THEN** the Verifier SHALL output verdict="REJECT"
|
||||
|
||||
#### Scenario: Structured Executor claims are verified first
|
||||
- **WHEN** `executor_structured_output.claims` is present and valid
|
||||
- **THEN** Verifier SHALL verify each structured claim against `tool_trace_summary` through `claim_checks`
|
||||
- **AND** each claim's evidence bindings SHALL reference existing trace or invocation identifiers when those identifiers are available
|
||||
- **AND** a claim with fabricated or missing evidence references SHALL NOT be classified as `direct_observation`
|
||||
- **AND** Verifier SHALL NOT add extra confirmed facts from `executor_final_answer` that are absent from `executor_structured_output.claims`
|
||||
|
||||
#### Scenario: Malformed structured output cannot pass through natural language fallback
|
||||
- **WHEN** Executor does not return parseable structured output
|
||||
- **THEN** Verifier SHALL NOT produce an effective `PASS` by extracting facts from `executor_final_answer`
|
||||
- **AND** the effective verdict SHALL be `LOW_CONFID`
|
||||
|
||||
### Requirement: Verifier SHALL output structured JSON
|
||||
The Verifier SHALL output a JSON object with verdict, groundedness_score, claim_checks array, compatibility facts_checked array, and rationale.
|
||||
|
||||
#### Scenario: claim-level verifier output is accepted
|
||||
- **WHEN** the Verifier checks Executor V2 structured output
|
||||
- **THEN** the output SHALL contain `verdict`, `groundedness_score`, `critical_fact_count`, `claim_checks`, `facts_checked`, and `rationale`
|
||||
- **AND** `claim_checks` SHALL be the primary V2 verification result
|
||||
- **AND** `facts_checked` SHALL remain available for compatibility
|
||||
|
||||
### Requirement: facts_checked SHALL use a fixed classification set
|
||||
The system SHALL continue to expose legacy `facts_checked` using its fixed verification classification set.
|
||||
|
||||
#### Scenario: claim checks are mapped to legacy facts
|
||||
- **WHEN** Verifier output contains `claim_checks`
|
||||
- **THEN** ChatService SHALL derive compatibility `facts_checked`
|
||||
- **AND** `direct_observation` SHALL map to `direct_evidence`
|
||||
- **AND** `reasonable_inference` and `overstated` SHALL map to `indirect_support`
|
||||
- **AND** `unsupported` and `external_unknown` SHALL map to `no_evidence`
|
||||
- **AND** `contradicted` SHALL map to `contradicted`
|
||||
|
||||
### Requirement: Verifier SHALL be observable
|
||||
The Verifier's verdict SHALL be persisted for observability.
|
||||
|
||||
#### Scenario: claim checks written to self_evaluation
|
||||
- **WHEN** the Verifier evaluation is persisted
|
||||
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `claim_checks`
|
||||
- **AND** it SHALL continue to include compatibility `facts_checked`
|
||||
- **AND** existing fields such as `verdict`, `groundedness_score`, `rationale`, `executor_output_parse_status`, `tool_trace_summary`, and `gatekeeper_result` SHALL be preserved
|
||||
|
||||
### Requirement: Verifier SHALL consume explicit verification inputs
|
||||
The Verifier SHALL receive explicit verification inputs rather than inferring them only from raw conversation history.
|
||||
|
||||
#### Scenario: structured claims are the primary verification target
|
||||
- **WHEN** `executor_output_parse_status.status` is `valid`
|
||||
- **AND** `executor_structured_output.claims` is available
|
||||
- **THEN** Verifier SHALL verify each claim through `claim_checks`
|
||||
- **AND** Verifier SHALL NOT add extra confirmed facts from `executor_final_answer` that are absent from `executor_structured_output.claims`
|
||||
|
||||
#### Scenario: malformed structured output cannot pass through natural language fallback
|
||||
- **WHEN** `executor_output_parse_status.status` is `missing` or `malformed`
|
||||
- **THEN** Verifier SHALL NOT produce an effective `PASS` by extracting facts from `executor_final_answer`
|
||||
- **AND** the effective verdict SHALL be `LOW_CONFID`
|
||||
|
||||
### Requirement: Executor Gatekeeper SHALL validate deterministic structured-output failures
|
||||
The system SHALL run deterministic Gatekeeper checks after Executor output parsing and before Verifier model execution.
|
||||
|
||||
#### Scenario: gatekeeper fail prevents PASS
|
||||
- **WHEN** `gatekeeper_result.status` is `fail`
|
||||
- **AND** the Verifier model returns `verdict = "PASS"`
|
||||
- **THEN** ChatService SHALL downgrade the effective verdict
|
||||
- **AND** the effective verdict SHALL NOT be `PASS`
|
||||
|
||||
#### Scenario: invocation reference failure downgrades to reject
|
||||
- **WHEN** `gatekeeper_result.failed_rules` contains `evidence.invocation_ref`
|
||||
- **AND** the Verifier model returns `verdict = "PASS"`
|
||||
- **THEN** ChatService SHALL set the effective verdict to `REJECT`
|
||||
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Verifier claim checks SHALL use a fixed derivability classification set
|
||||
The Verifier SHALL classify each structured claim using a fixed derivability classification set.
|
||||
|
||||
#### Scenario: claim check verification values are constrained
|
||||
- **WHEN** Verifier emits `claim_checks`
|
||||
- **THEN** each item SHALL use one of `direct_observation`, `reasonable_inference`, `overstated`, `unsupported`, `external_unknown`, or `contradicted`
|
||||
|
||||
#### Scenario: claim check evidence references remain auditable
|
||||
- **WHEN** Verifier emits `claim_checks`
|
||||
- **THEN** each claim check SHALL include `claim_id`, `verification`, `detail`, and `evidence_refs`
|
||||
- **AND** every evidence ref SHALL preserve available `trace_ref`, `tool_name`, and `source_invocation_ids`
|
||||
@@ -0,0 +1,39 @@
|
||||
## 1. Prompt Contract
|
||||
|
||||
- [x] 1.1 Update `chat-verifier-prompt.md` so `executor_structured_output.claims` is the primary verification target.
|
||||
- [x] 1.2 Remove the valid-structured-output path that scans `executor_final_answer` for extra confirmed facts.
|
||||
- [x] 1.3 Add `claim_checks` and optional `hypothesis_checks` to the output contract.
|
||||
- [x] 1.4 Keep `facts_checked` as compatibility output.
|
||||
- [x] 1.5 State that missing/malformed structured output cannot produce PASS through natural-language fallback.
|
||||
|
||||
## 2. Parser And Compatibility Mapping
|
||||
|
||||
- [x] 2.1 Extend `VerifierDecision` to store `claim_checks`.
|
||||
- [x] 2.2 Parse `claim_checks` from verifier output.
|
||||
- [x] 2.3 Derive compatibility `facts_checked` from `claim_checks` when present.
|
||||
- [x] 2.4 Preserve old `facts_checked` parsing when `claim_checks` is absent.
|
||||
- [x] 2.5 Persist `claim_checks` in verifier evaluation.
|
||||
|
||||
## 3. Effective Verdict Guardrails
|
||||
|
||||
- [x] 3.1 Add code-side guard so `gatekeeper_result.status=fail` cannot result in effective PASS.
|
||||
- [x] 3.2 Downgrade `evidence.invocation_ref` failures to REJECT.
|
||||
- [x] 3.3 Downgrade other Gatekeeper failures to at least LOW_CONFID.
|
||||
- [x] 3.4 Ensure missing/malformed Executor structured output cannot produce effective PASS.
|
||||
|
||||
## 4. Compatibility Consumers
|
||||
|
||||
- [x] 4.1 Ensure low-confidence rendering still uses compatibility `facts_checked`.
|
||||
- [x] 4.2 Ensure `buildRetryContext(...)` still receives evidence gaps from compatibility `facts_checked`.
|
||||
- [x] 4.3 Update verifier logging summary to account for `claim_checks`.
|
||||
- [x] 4.4 Preserve existing traceability and Gatekeeper audit fields.
|
||||
|
||||
## 5. Tests And Verification
|
||||
|
||||
- [x] 5.1 Add or update tests for parsing `claim_checks`.
|
||||
- [x] 5.2 Add tests for `claim_checks` to `facts_checked` compatibility mapping.
|
||||
- [x] 5.3 Add tests for each verification mapping class: `direct_observation`, `reasonable_inference`, `overstated`, `unsupported`, `external_unknown`, `contradicted`.
|
||||
- [x] 5.4 Add tests proving Gatekeeper fail cannot remain PASS.
|
||||
- [x] 5.5 Add tests proving malformed/missing structured output cannot remain PASS.
|
||||
- [x] 5.6 Run targeted tests.
|
||||
- [x] 5.7 Validate this OpenSpec change.
|
||||
Reference in New Issue
Block a user