feat(eval): add evidence pipeline acceptance closure
This commit is contained in:
+66
@@ -0,0 +1,66 @@
|
||||
# diagnosis-eval-demo-gatekeeper-closure
|
||||
|
||||
## Problem
|
||||
|
||||
The Chat evidence pipeline now has Executor V2 structured output, deterministic Gatekeeper validation, Verifier claim checks, and Composer final rendering. The runtime path is stronger than the project-level acceptance story around it.
|
||||
|
||||
Current gaps:
|
||||
|
||||
- Diagnosis eval fixtures cover several V2 audit behaviors, but they do not yet form an explicit interview-ready matrix for positive evidence, no-evidence, reject, low-confidence, narrow-scope, and raw-output leakage.
|
||||
- Demo assets are still centered on the original payment-timeout path. They do not clearly package the stable PASS / LOW_CONFID / REJECT examples needed for an Agent engineering interview.
|
||||
- Gatekeeper rules are hard-coded constants and thresholds. Audit output identifies failed rules, but it does not expose a rule set version or rule metadata that can be discussed, tested, and evolved.
|
||||
|
||||
## Proposed Change
|
||||
|
||||
Create an acceptance closure layer around the existing evidence pipeline:
|
||||
|
||||
```text
|
||||
stable demo scenarios
|
||||
-> saved offline diagnosis fixtures
|
||||
-> deterministic diagnosis eval matrix
|
||||
-> Gatekeeper rule metadata/version
|
||||
-> persisted audit result that records the rule set version
|
||||
```
|
||||
|
||||
This change keeps the current runtime architecture. It does not introduce new Agents or new database tables. It strengthens the project as an interview-ready Agent engineering artifact by making the anti-hallucination behavior demonstrable and regressable.
|
||||
|
||||
## Scope
|
||||
|
||||
- Expand `mvp/eval` case definitions and trace fixtures into an explicit evidence-pipeline matrix.
|
||||
- Add or update baseline reports so the fixed matrix remains deterministic.
|
||||
- Add stable demo request payloads and runbook docs for PASS, LOW_CONFID / no-evidence, and REJECT / fabricated-reference style scenarios.
|
||||
- Add Gatekeeper rule metadata/configuration with a small rule set version.
|
||||
- Include Gatekeeper rule set version and loaded rule metadata summary in `gatekeeper_result`.
|
||||
- Add focused tests for eval matrix behavior and Gatekeeper rule metadata/audit version.
|
||||
- Update architecture/demo/eval docs as needed.
|
||||
|
||||
## Non-goals
|
||||
|
||||
- No new public HTTP endpoint.
|
||||
- No new database table.
|
||||
- No Planner `scope_contract`.
|
||||
- No retry loop from Gatekeeper back to Executor.
|
||||
- No full JSONPath engine.
|
||||
- No LLM-as-judge.
|
||||
- No production-grade remote rule registry.
|
||||
- No automatic prompt optimization.
|
||||
|
||||
## Context Constraints
|
||||
|
||||
- Diagnosis eval should remain offline and deterministic.
|
||||
- Stable demo data should reuse existing mock tools and `knowledge_base` where possible.
|
||||
- Gatekeeper audit should remain under `diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result`.
|
||||
- Gatekeeper rules should stay lightweight: rule id, description, severity, enabled flag, and config parameters are enough for this phase.
|
||||
- `$.no_evidence` means only "this tool query returned no matching evidence"; it must not become a confirmed absence of the underlying problem.
|
||||
|
||||
## Assumptions
|
||||
|
||||
- This is a `standard` sm-flow change because it touches eval assets, demo assets, Gatekeeper internals, tests, and docs.
|
||||
- Interface impact is expected to be L2 internal contract change: internal JSON audit fields expand, but no external HTTP contract or database schema changes.
|
||||
- Existing V2 audit fixtures and ISS-008 / ISS-009 validation records are acceptable seeds for the matrix.
|
||||
|
||||
## Risks
|
||||
|
||||
- If demo scenarios depend on live LLM output, they may still be nondeterministic. The fixed eval matrix must use saved fixtures.
|
||||
- If Gatekeeper config becomes too flexible, it could obscure deterministic safety rules. This phase should only expose metadata and simple parameters.
|
||||
- Baseline reports must be updated together with case/fixture changes, or the evaluator tests will become noisy.
|
||||
Reference in New Issue
Block a user