Files

3.7 KiB

diagnosis-eval-demo-gatekeeper-closure

Problem

The Chat evidence pipeline now has Executor V2 structured output, deterministic Gatekeeper validation, Verifier claim checks, and Composer final rendering. The runtime path is stronger than the project-level acceptance story around it.

Current gaps:

  • Diagnosis eval fixtures cover several V2 audit behaviors, but they do not yet form an explicit interview-ready matrix for positive evidence, no-evidence, reject, low-confidence, narrow-scope, and raw-output leakage.
  • Demo assets are still centered on the original payment-timeout path. They do not clearly package the stable PASS / LOW_CONFID / REJECT examples needed for an Agent engineering interview.
  • Gatekeeper rules are hard-coded constants and thresholds. Audit output identifies failed rules, but it does not expose a rule set version or rule metadata that can be discussed, tested, and evolved.

Proposed Change

Create an acceptance closure layer around the existing evidence pipeline:

stable demo scenarios
  -> saved offline diagnosis fixtures
  -> deterministic diagnosis eval matrix
  -> Gatekeeper rule metadata/version
  -> persisted audit result that records the rule set version

This change keeps the current runtime architecture. It does not introduce new Agents or new database tables. It strengthens the project as an interview-ready Agent engineering artifact by making the anti-hallucination behavior demonstrable and regressable.

Scope

  • Expand mvp/eval case definitions and trace fixtures into an explicit evidence-pipeline matrix.
  • Add or update baseline reports so the fixed matrix remains deterministic.
  • Add stable demo request payloads and runbook docs for PASS, LOW_CONFID / no-evidence, and REJECT / fabricated-reference style scenarios.
  • Add Gatekeeper rule metadata/configuration with a small rule set version.
  • Include Gatekeeper rule set version and loaded rule metadata summary in gatekeeper_result.
  • Add focused tests for eval matrix behavior and Gatekeeper rule metadata/audit version.
  • Update architecture/demo/eval docs as needed.

Non-goals

  • No new public HTTP endpoint.
  • No new database table.
  • No Planner scope_contract.
  • No retry loop from Gatekeeper back to Executor.
  • No full JSONPath engine.
  • No LLM-as-judge.
  • No production-grade remote rule registry.
  • No automatic prompt optimization.

Context Constraints

  • Diagnosis eval should remain offline and deterministic.
  • Stable demo data should reuse existing mock tools and knowledge_base where possible.
  • Gatekeeper audit should remain under diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result.
  • Gatekeeper rules should stay lightweight: rule id, description, severity, enabled flag, and config parameters are enough for this phase.
  • $.no_evidence means only "this tool query returned no matching evidence"; it must not become a confirmed absence of the underlying problem.

Assumptions

  • This is a standard sm-flow change because it touches eval assets, demo assets, Gatekeeper internals, tests, and docs.
  • Interface impact is expected to be L2 internal contract change: internal JSON audit fields expand, but no external HTTP contract or database schema changes.
  • Existing V2 audit fixtures and ISS-008 / ISS-009 validation records are acceptable seeds for the matrix.

Risks

  • If demo scenarios depend on live LLM output, they may still be nondeterministic. The fixed eval matrix must use saved fixtures.
  • If Gatekeeper config becomes too flexible, it could obscure deterministic safety rules. This phase should only expose metadata and simple parameters.
  • Baseline reports must be updated together with case/fixture changes, or the evaluator tests will become noisy.