67 lines
3.7 KiB
Markdown
67 lines
3.7 KiB
Markdown
# diagnosis-eval-demo-gatekeeper-closure
|
|
|
|
## Problem
|
|
|
|
The Chat evidence pipeline now has Executor V2 structured output, deterministic Gatekeeper validation, Verifier claim checks, and Composer final rendering. The runtime path is stronger than the project-level acceptance story around it.
|
|
|
|
Current gaps:
|
|
|
|
- Diagnosis eval fixtures cover several V2 audit behaviors, but they do not yet form an explicit interview-ready matrix for positive evidence, no-evidence, reject, low-confidence, narrow-scope, and raw-output leakage.
|
|
- Demo assets are still centered on the original payment-timeout path. They do not clearly package the stable PASS / LOW_CONFID / REJECT examples needed for an Agent engineering interview.
|
|
- Gatekeeper rules are hard-coded constants and thresholds. Audit output identifies failed rules, but it does not expose a rule set version or rule metadata that can be discussed, tested, and evolved.
|
|
|
|
## Proposed Change
|
|
|
|
Create an acceptance closure layer around the existing evidence pipeline:
|
|
|
|
```text
|
|
stable demo scenarios
|
|
-> saved offline diagnosis fixtures
|
|
-> deterministic diagnosis eval matrix
|
|
-> Gatekeeper rule metadata/version
|
|
-> persisted audit result that records the rule set version
|
|
```
|
|
|
|
This change keeps the current runtime architecture. It does not introduce new Agents or new database tables. It strengthens the project as an interview-ready Agent engineering artifact by making the anti-hallucination behavior demonstrable and regressable.
|
|
|
|
## Scope
|
|
|
|
- Expand `mvp/eval` case definitions and trace fixtures into an explicit evidence-pipeline matrix.
|
|
- Add or update baseline reports so the fixed matrix remains deterministic.
|
|
- Add stable demo request payloads and runbook docs for PASS, LOW_CONFID / no-evidence, and REJECT / fabricated-reference style scenarios.
|
|
- Add Gatekeeper rule metadata/configuration with a small rule set version.
|
|
- Include Gatekeeper rule set version and loaded rule metadata summary in `gatekeeper_result`.
|
|
- Add focused tests for eval matrix behavior and Gatekeeper rule metadata/audit version.
|
|
- Update architecture/demo/eval docs as needed.
|
|
|
|
## Non-goals
|
|
|
|
- No new public HTTP endpoint.
|
|
- No new database table.
|
|
- No Planner `scope_contract`.
|
|
- No retry loop from Gatekeeper back to Executor.
|
|
- No full JSONPath engine.
|
|
- No LLM-as-judge.
|
|
- No production-grade remote rule registry.
|
|
- No automatic prompt optimization.
|
|
|
|
## Context Constraints
|
|
|
|
- Diagnosis eval should remain offline and deterministic.
|
|
- Stable demo data should reuse existing mock tools and `knowledge_base` where possible.
|
|
- Gatekeeper audit should remain under `diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result`.
|
|
- Gatekeeper rules should stay lightweight: rule id, description, severity, enabled flag, and config parameters are enough for this phase.
|
|
- `$.no_evidence` means only "this tool query returned no matching evidence"; it must not become a confirmed absence of the underlying problem.
|
|
|
|
## Assumptions
|
|
|
|
- This is a `standard` sm-flow change because it touches eval assets, demo assets, Gatekeeper internals, tests, and docs.
|
|
- Interface impact is expected to be L2 internal contract change: internal JSON audit fields expand, but no external HTTP contract or database schema changes.
|
|
- Existing V2 audit fixtures and ISS-008 / ISS-009 validation records are acceptable seeds for the matrix.
|
|
|
|
## Risks
|
|
|
|
- If demo scenarios depend on live LLM output, they may still be nondeterministic. The fixed eval matrix must use saved fixtures.
|
|
- If Gatekeeper config becomes too flexible, it could obscure deterministic safety rules. This phase should only expose metadata and simple parameters.
|
|
- Baseline reports must be updated together with case/fixture changes, or the evaluator tests will become noisy.
|