# diagnosis-eval-demo-gatekeeper-closure ## Problem The Chat evidence pipeline now has Executor V2 structured output, deterministic Gatekeeper validation, Verifier claim checks, and Composer final rendering. The runtime path is stronger than the project-level acceptance story around it. Current gaps: - Diagnosis eval fixtures cover several V2 audit behaviors, but they do not yet form an explicit interview-ready matrix for positive evidence, no-evidence, reject, low-confidence, narrow-scope, and raw-output leakage. - Demo assets are still centered on the original payment-timeout path. They do not clearly package the stable PASS / LOW_CONFID / REJECT examples needed for an Agent engineering interview. - Gatekeeper rules are hard-coded constants and thresholds. Audit output identifies failed rules, but it does not expose a rule set version or rule metadata that can be discussed, tested, and evolved. ## Proposed Change Create an acceptance closure layer around the existing evidence pipeline: ```text stable demo scenarios -> saved offline diagnosis fixtures -> deterministic diagnosis eval matrix -> Gatekeeper rule metadata/version -> persisted audit result that records the rule set version ``` This change keeps the current runtime architecture. It does not introduce new Agents or new database tables. It strengthens the project as an interview-ready Agent engineering artifact by making the anti-hallucination behavior demonstrable and regressable. ## Scope - Expand `mvp/eval` case definitions and trace fixtures into an explicit evidence-pipeline matrix. - Add or update baseline reports so the fixed matrix remains deterministic. - Add stable demo request payloads and runbook docs for PASS, LOW_CONFID / no-evidence, and REJECT / fabricated-reference style scenarios. - Add Gatekeeper rule metadata/configuration with a small rule set version. - Include Gatekeeper rule set version and loaded rule metadata summary in `gatekeeper_result`. - Add focused tests for eval matrix behavior and Gatekeeper rule metadata/audit version. - Update architecture/demo/eval docs as needed. ## Non-goals - No new public HTTP endpoint. - No new database table. - No Planner `scope_contract`. - No retry loop from Gatekeeper back to Executor. - No full JSONPath engine. - No LLM-as-judge. - No production-grade remote rule registry. - No automatic prompt optimization. ## Context Constraints - Diagnosis eval should remain offline and deterministic. - Stable demo data should reuse existing mock tools and `knowledge_base` where possible. - Gatekeeper audit should remain under `diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result`. - Gatekeeper rules should stay lightweight: rule id, description, severity, enabled flag, and config parameters are enough for this phase. - `$.no_evidence` means only "this tool query returned no matching evidence"; it must not become a confirmed absence of the underlying problem. ## Assumptions - This is a `standard` sm-flow change because it touches eval assets, demo assets, Gatekeeper internals, tests, and docs. - Interface impact is expected to be L2 internal contract change: internal JSON audit fields expand, but no external HTTP contract or database schema changes. - Existing V2 audit fixtures and ISS-008 / ISS-009 validation records are acceptable seeds for the matrix. ## Risks - If demo scenarios depend on live LLM output, they may still be nondeterministic. The fixed eval matrix must use saved fixtures. - If Gatekeeper config becomes too flexible, it could obscure deterministic safety rules. This phase should only expose metadata and simple parameters. - Baseline reports must be updated together with case/fixture changes, or the evaluator tests will become noisy.