52 lines
2.4 KiB
Markdown
52 lines
2.4 KiB
Markdown
# Tasks
|
|
|
|
## 1. Diagnosis eval matrix
|
|
|
|
- [x] Add matrix-oriented eval cases for narrow-scope supported evidence and no-evidence negative observation.
|
|
- [x] Add matching offline trace fixtures with V2 audit closure fields.
|
|
- [x] Extend evaluator case/result model to optionally check `gatekeeper_result.rule_set_version`.
|
|
- [x] Regenerate baseline JSON and Markdown reports.
|
|
- [x] Update eval README/schema to describe the matrix and rule set version check.
|
|
|
|
Acceptance:
|
|
|
|
- `DiagnosisTraceEvaluatorTest` passes with the new total case count and verdict distribution.
|
|
- Every fixed case resolves to an existing fixture.
|
|
- V2 matrix cases include Gatekeeper, claim checks, Composer output, and final-answer leakage checks where relevant.
|
|
|
|
## 2. Stable demo scenarios
|
|
|
|
- [x] Add stable demo request payloads for supported positive, no-evidence, and safety/reject discussion scenarios.
|
|
- [x] Add a demo scenario guide that maps each payload or fixture to interview claims and expected trace fields.
|
|
- [x] Update existing demo README/runbook references so reviewers know which path is live and which paths are fixture-backed.
|
|
|
|
Acceptance:
|
|
|
|
- Demo docs clearly distinguish live script path from deterministic fixture-backed scenarios.
|
|
- Each scenario has a stable session id or fixture id.
|
|
- No demo doc claims unsupported production behavior.
|
|
|
|
## 3. Gatekeeper rule catalog and audit version
|
|
|
|
- [x] Add a lightweight Gatekeeper rule catalog with version and rule metadata.
|
|
- [x] Include `rule_set_version` and enabled rule metadata summary in every Gatekeeper result, including pass/fallback/internal-error results.
|
|
- [x] Add tests for catalog loading and audit fields.
|
|
- [x] Keep validation logic deterministic; do not add dynamic script execution or remote config.
|
|
|
|
Acceptance:
|
|
|
|
- `ExecutorGatekeeperServiceTest` proves Gatekeeper output contains the rule set version and rule metadata.
|
|
- Existing Gatekeeper pass/low_confid/reject behavior remains unchanged.
|
|
|
|
## 4. Verification
|
|
|
|
- [x] Run focused tests for eval and Gatekeeper changes.
|
|
- [x] Run broader relevant regression tests if focused changes touch shared code.
|
|
- [x] Run at least one live end-to-end check if unit/fixture evidence is insufficient to prove demo path compatibility.
|
|
- [x] Record verification commands and results in devflow acceptance.
|
|
|
|
Acceptance:
|
|
|
|
- All required tests pass, or failures are classified and fixed before archive.
|
|
- If live E2E is skipped, the reason is documented and fixture coverage must prove the requested behavior.
|