Files

67 lines
3.7 KiB
Markdown

# diagnosis-eval-demo-gatekeeper-closure
## Problem
The Chat evidence pipeline now has Executor V2 structured output, deterministic Gatekeeper validation, Verifier claim checks, and Composer final rendering. The runtime path is stronger than the project-level acceptance story around it.
Current gaps:
- Diagnosis eval fixtures cover several V2 audit behaviors, but they do not yet form an explicit interview-ready matrix for positive evidence, no-evidence, reject, low-confidence, narrow-scope, and raw-output leakage.
- Demo assets are still centered on the original payment-timeout path. They do not clearly package the stable PASS / LOW_CONFID / REJECT examples needed for an Agent engineering interview.
- Gatekeeper rules are hard-coded constants and thresholds. Audit output identifies failed rules, but it does not expose a rule set version or rule metadata that can be discussed, tested, and evolved.
## Proposed Change
Create an acceptance closure layer around the existing evidence pipeline:
```text
stable demo scenarios
-> saved offline diagnosis fixtures
-> deterministic diagnosis eval matrix
-> Gatekeeper rule metadata/version
-> persisted audit result that records the rule set version
```
This change keeps the current runtime architecture. It does not introduce new Agents or new database tables. It strengthens the project as an interview-ready Agent engineering artifact by making the anti-hallucination behavior demonstrable and regressable.
## Scope
- Expand `mvp/eval` case definitions and trace fixtures into an explicit evidence-pipeline matrix.
- Add or update baseline reports so the fixed matrix remains deterministic.
- Add stable demo request payloads and runbook docs for PASS, LOW_CONFID / no-evidence, and REJECT / fabricated-reference style scenarios.
- Add Gatekeeper rule metadata/configuration with a small rule set version.
- Include Gatekeeper rule set version and loaded rule metadata summary in `gatekeeper_result`.
- Add focused tests for eval matrix behavior and Gatekeeper rule metadata/audit version.
- Update architecture/demo/eval docs as needed.
## Non-goals
- No new public HTTP endpoint.
- No new database table.
- No Planner `scope_contract`.
- No retry loop from Gatekeeper back to Executor.
- No full JSONPath engine.
- No LLM-as-judge.
- No production-grade remote rule registry.
- No automatic prompt optimization.
## Context Constraints
- Diagnosis eval should remain offline and deterministic.
- Stable demo data should reuse existing mock tools and `knowledge_base` where possible.
- Gatekeeper audit should remain under `diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result`.
- Gatekeeper rules should stay lightweight: rule id, description, severity, enabled flag, and config parameters are enough for this phase.
- `$.no_evidence` means only "this tool query returned no matching evidence"; it must not become a confirmed absence of the underlying problem.
## Assumptions
- This is a `standard` sm-flow change because it touches eval assets, demo assets, Gatekeeper internals, tests, and docs.
- Interface impact is expected to be L2 internal contract change: internal JSON audit fields expand, but no external HTTP contract or database schema changes.
- Existing V2 audit fixtures and ISS-008 / ISS-009 validation records are acceptable seeds for the matrix.
## Risks
- If demo scenarios depend on live LLM output, they may still be nondeterministic. The fixed eval matrix must use saved fixtures.
- If Gatekeeper config becomes too flexible, it could obscure deterministic safety rules. This phase should only expose metadata and simple parameters.
- Baseline reports must be updated together with case/fixture changes, or the evaluator tests will become noisy.