feat(eval): add evidence pipeline acceptance closure

This commit is contained in:
aruo
2026-07-09 00:47:48 +08:00
parent a77c947cd4
commit db0f229285
46 changed files with 1434 additions and 53 deletions
@@ -0,0 +1,110 @@
# Design
## Data Flow
```text
Gatekeeper rule catalog
-> ExecutorGatekeeperService.validate(...)
-> gatekeeper_result.rule_set_version + rule metadata summary
-> DiagnosisTraceEvaluator fixture checks
-> baseline report / demo runbook
```
```text
Stable demo scenarios
-> request payloads and docs
-> optional live run for main path
-> saved trace fixtures for deterministic matrix
-> mvp/eval baseline report
```
## Eval Matrix
The diagnosis eval matrix remains offline and deterministic. It should cover these rows:
| Matrix row | Expected signal |
|---|---|
| Positive supported evidence | `PASS`, Gatekeeper `pass`, Composer `valid` |
| Narrow-scope observation | `PASS`, one or minimal claims, no forbidden over-expansion |
| No-evidence negative observation | `PASS` or allowed non-reject verdict, `$.no_evidence`, no overstatement |
| Unsupported claim filtering | `LOW_CONFID`, unsupported claim not in final answer |
| Gatekeeper fabricated reference | `REJECT` or `LOW_CONFID`, Gatekeeper `fail` |
| Composer fallback | no raw Executor protocol leakage |
The evaluator should validate rule set version only for cases that opt into the new check. This keeps old fixtures readable while allowing the new matrix to prove Gatekeeper metadata persistence.
## Stable Demo Scenarios
Stable demo scenarios are source-controlled payloads and documentation, not a live-only test harness. The demo set should include:
- A supported positive path.
- A no-evidence / negative-observation path.
- A safety/reject path explained through fixed fixture or evaluator output.
Only the main path needs a live script in this phase. Other scenarios may be represented by payloads, fixture names, and expected trace fields.
## Gatekeeper Rule Catalog
Gatekeeper rules stay local and lightweight:
```json
{
"version": "gatekeeper-rules-v1",
"rules": [
{
"id": "evidence.raw_path",
"description": "raw_path must exist in retrieval_details.evidence_refs",
"default_severity": "reject",
"enabled": true
}
]
}
```
The first implementation may use an in-memory default catalog or a classpath JSON resource. It must expose:
- rule set version
- enabled rule ids
- rule descriptions
- severity defaults or threshold parameters when present
Gatekeeper validation logic remains deterministic Java code. The catalog is metadata/config, not a dynamic scripting engine.
## Audit Persistence
`gatekeeper_result` should include:
```json
{
"rule_set_version": "gatekeeper-rules-v1",
"rules": [
{
"id": "evidence.raw_path",
"description": "raw_path must exist in retrieval_details.evidence_refs",
"enabled": true,
"default_severity": "reject"
}
]
}
```
Existing fields remain:
- `status`
- `severity`
- `checked_bindings`
- `failed_rules`
- `warnings`
- `errors`
## Interface Impact
Impact level: L2 internal contract change.
This expands internal audit JSON and eval case/result fields. It does not change external HTTP APIs, database schema, Planner output, or public DTO contracts.
## Risks
- Too much Gatekeeper flexibility could weaken safety. This phase only adds metadata/configuration and keeps rule implementations fixed in code.
- Demo scenarios should not promise deterministic LLM behavior. Deterministic claims should point to fixture-backed eval results.
- Baseline updates must be made together with new fixtures and tests.
@@ -0,0 +1,66 @@
# diagnosis-eval-demo-gatekeeper-closure
## Problem
The Chat evidence pipeline now has Executor V2 structured output, deterministic Gatekeeper validation, Verifier claim checks, and Composer final rendering. The runtime path is stronger than the project-level acceptance story around it.
Current gaps:
- Diagnosis eval fixtures cover several V2 audit behaviors, but they do not yet form an explicit interview-ready matrix for positive evidence, no-evidence, reject, low-confidence, narrow-scope, and raw-output leakage.
- Demo assets are still centered on the original payment-timeout path. They do not clearly package the stable PASS / LOW_CONFID / REJECT examples needed for an Agent engineering interview.
- Gatekeeper rules are hard-coded constants and thresholds. Audit output identifies failed rules, but it does not expose a rule set version or rule metadata that can be discussed, tested, and evolved.
## Proposed Change
Create an acceptance closure layer around the existing evidence pipeline:
```text
stable demo scenarios
-> saved offline diagnosis fixtures
-> deterministic diagnosis eval matrix
-> Gatekeeper rule metadata/version
-> persisted audit result that records the rule set version
```
This change keeps the current runtime architecture. It does not introduce new Agents or new database tables. It strengthens the project as an interview-ready Agent engineering artifact by making the anti-hallucination behavior demonstrable and regressable.
## Scope
- Expand `mvp/eval` case definitions and trace fixtures into an explicit evidence-pipeline matrix.
- Add or update baseline reports so the fixed matrix remains deterministic.
- Add stable demo request payloads and runbook docs for PASS, LOW_CONFID / no-evidence, and REJECT / fabricated-reference style scenarios.
- Add Gatekeeper rule metadata/configuration with a small rule set version.
- Include Gatekeeper rule set version and loaded rule metadata summary in `gatekeeper_result`.
- Add focused tests for eval matrix behavior and Gatekeeper rule metadata/audit version.
- Update architecture/demo/eval docs as needed.
## Non-goals
- No new public HTTP endpoint.
- No new database table.
- No Planner `scope_contract`.
- No retry loop from Gatekeeper back to Executor.
- No full JSONPath engine.
- No LLM-as-judge.
- No production-grade remote rule registry.
- No automatic prompt optimization.
## Context Constraints
- Diagnosis eval should remain offline and deterministic.
- Stable demo data should reuse existing mock tools and `knowledge_base` where possible.
- Gatekeeper audit should remain under `diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result`.
- Gatekeeper rules should stay lightweight: rule id, description, severity, enabled flag, and config parameters are enough for this phase.
- `$.no_evidence` means only "this tool query returned no matching evidence"; it must not become a confirmed absence of the underlying problem.
## Assumptions
- This is a `standard` sm-flow change because it touches eval assets, demo assets, Gatekeeper internals, tests, and docs.
- Interface impact is expected to be L2 internal contract change: internal JSON audit fields expand, but no external HTTP contract or database schema changes.
- Existing V2 audit fixtures and ISS-008 / ISS-009 validation records are acceptable seeds for the matrix.
## Risks
- If demo scenarios depend on live LLM output, they may still be nondeterministic. The fixed eval matrix must use saved fixtures.
- If Gatekeeper config becomes too flexible, it could obscure deterministic safety rules. This phase should only expose metadata and simple parameters.
- Baseline reports must be updated together with case/fixture changes, or the evaluator tests will become noisy.
@@ -0,0 +1,26 @@
## ADDED Requirements
### Requirement: Gatekeeper SHALL expose rule catalog metadata in audit output
Gatekeeper SHALL include the rule catalog version and enabled rule metadata in its validation result.
#### Scenario: Gatekeeper pass includes rule metadata
- **WHEN** Gatekeeper returns `status=pass`
- **THEN** the result SHALL include `rule_set_version`
- **AND** it SHALL include a rule metadata summary
#### Scenario: Gatekeeper fail includes rule metadata
- **WHEN** Gatekeeper returns `status=fail`
- **THEN** the result SHALL include `rule_set_version`
- **AND** it SHALL include a rule metadata summary
#### Scenario: Gatekeeper fallback pass includes rule metadata
- **WHEN** the Verifier input hook returns a fallback Gatekeeper pass because no Gatekeeper service is available
- **THEN** the result SHOULD still include the default rule set version and an empty or default rule metadata summary
### Requirement: Gatekeeper rule catalog SHALL remain deterministic
The Gatekeeper rule catalog SHALL configure metadata and simple parameters only; validation behavior SHALL remain deterministic Java code.
#### Scenario: Rule metadata is lightweight
- **WHEN** rule metadata is loaded
- **THEN** each enabled rule SHOULD expose an id, description, enabled flag, and default severity or relevant parameter
- **AND** rule metadata SHALL NOT execute dynamic scripts
@@ -0,0 +1,28 @@
## ADDED Requirements
### Requirement: Evaluation harness SHALL expose an evidence-pipeline matrix
The evaluation harness SHALL include fixed cases that demonstrate the current Chat evidence pipeline across positive evidence, narrow-scope observation, no-evidence negative observation, low-confidence filtering, reject safety, and composer fallback behavior.
#### Scenario: Matrix cases are fixture backed
- **WHEN** the fixed diagnosis case file is evaluated
- **THEN** each matrix case SHALL resolve to an offline trace fixture
- **AND** evaluation SHALL not require a live LLM or running application
#### Scenario: Matrix cases preserve V2 audit closure
- **WHEN** a matrix case requires V2 audit closure
- **THEN** its fixture SHALL include `gatekeeper_result`
- **AND** it SHALL include `claim_checks`
- **AND** it SHALL include `composer_output`
### Requirement: Evaluation harness SHALL validate Gatekeeper rule set version when requested
The evaluation harness SHALL be able to assert the Gatekeeper rule set version recorded in a fixture.
#### Scenario: Expected rule set version matches
- **WHEN** an evaluation case declares `expectedGatekeeperRuleSetVersion`
- **AND** the fixture has the same `selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version`
- **THEN** the rule set version check SHALL pass
#### Scenario: Expected rule set version mismatches
- **WHEN** an evaluation case declares `expectedGatekeeperRuleSetVersion`
- **AND** the fixture has a different or missing rule set version
- **THEN** the case SHALL fail with a clear failed check
@@ -0,0 +1,15 @@
## ADDED Requirements
### Requirement: MVP demo SHALL provide stable evidence-pipeline scenarios
The MVP demo SHALL provide stable scenarios that explain how to demonstrate positive evidence, no-evidence, and safety/reject behavior for an Agent engineering interview.
#### Scenario: Scenario guide maps demo inputs to evidence claims
- **WHEN** a reviewer opens the demo scenario guide
- **THEN** it SHALL list the supported positive, no-evidence, and safety/reject scenarios
- **AND** it SHALL map each scenario to a request payload or fixture id
- **AND** it SHALL describe the expected Gatekeeper, Verifier, Composer, and trace fields to inspect
#### Scenario: Live and fixture-backed scenarios are distinguished
- **WHEN** a demo scenario is fixture-backed rather than live-scripted
- **THEN** the documentation SHALL say so explicitly
- **AND** it SHALL avoid promising deterministic live LLM output for that scenario
@@ -0,0 +1,51 @@
# Tasks
## 1. Diagnosis eval matrix
- [x] Add matrix-oriented eval cases for narrow-scope supported evidence and no-evidence negative observation.
- [x] Add matching offline trace fixtures with V2 audit closure fields.
- [x] Extend evaluator case/result model to optionally check `gatekeeper_result.rule_set_version`.
- [x] Regenerate baseline JSON and Markdown reports.
- [x] Update eval README/schema to describe the matrix and rule set version check.
Acceptance:
- `DiagnosisTraceEvaluatorTest` passes with the new total case count and verdict distribution.
- Every fixed case resolves to an existing fixture.
- V2 matrix cases include Gatekeeper, claim checks, Composer output, and final-answer leakage checks where relevant.
## 2. Stable demo scenarios
- [x] Add stable demo request payloads for supported positive, no-evidence, and safety/reject discussion scenarios.
- [x] Add a demo scenario guide that maps each payload or fixture to interview claims and expected trace fields.
- [x] Update existing demo README/runbook references so reviewers know which path is live and which paths are fixture-backed.
Acceptance:
- Demo docs clearly distinguish live script path from deterministic fixture-backed scenarios.
- Each scenario has a stable session id or fixture id.
- No demo doc claims unsupported production behavior.
## 3. Gatekeeper rule catalog and audit version
- [x] Add a lightweight Gatekeeper rule catalog with version and rule metadata.
- [x] Include `rule_set_version` and enabled rule metadata summary in every Gatekeeper result, including pass/fallback/internal-error results.
- [x] Add tests for catalog loading and audit fields.
- [x] Keep validation logic deterministic; do not add dynamic script execution or remote config.
Acceptance:
- `ExecutorGatekeeperServiceTest` proves Gatekeeper output contains the rule set version and rule metadata.
- Existing Gatekeeper pass/low_confid/reject behavior remains unchanged.
## 4. Verification
- [x] Run focused tests for eval and Gatekeeper changes.
- [x] Run broader relevant regression tests if focused changes touch shared code.
- [x] Run at least one live end-to-end check if unit/fixture evidence is insufficient to prove demo path compatibility.
- [x] Record verification commands and results in devflow acceptance.
Acceptance:
- All required tests pass, or failures are classified and fixed before archive.
- If live E2E is skipped, the reason is documented and fixture coverage must prove the requested behavior.
+26 -1
View File
@@ -1,4 +1,4 @@
# chat-verifier-agent Specification
# chat-verifier-agent Specification
## Purpose
TBD - created by archiving change chat-verifier-agent. Update Purpose after archive.
@@ -531,3 +531,28 @@ The Executor prompt SHALL instruct Executor to keep narrow confirmation question
- **THEN** Executor SHOULD output the minimum necessary claims, normally one and at most two
- **AND** those claims SHALL be `observation` or `negative_observation` unless current-session evidence proves more
- **AND** Executor SHALL NOT emit unrelated root-cause, remediation, or excluded-topic claims as confirmed facts
### Requirement: Gatekeeper SHALL expose rule catalog metadata in audit output
Gatekeeper SHALL include the rule catalog version and enabled rule metadata in its validation result.
#### Scenario: Gatekeeper pass includes rule metadata
- **WHEN** Gatekeeper returns `status=pass`
- **THEN** the result SHALL include `rule_set_version`
- **AND** it SHALL include a rule metadata summary
#### Scenario: Gatekeeper fail includes rule metadata
- **WHEN** Gatekeeper returns `status=fail`
- **THEN** the result SHALL include `rule_set_version`
- **AND** it SHALL include a rule metadata summary
#### Scenario: Gatekeeper fallback pass includes rule metadata
- **WHEN** the Verifier input hook returns a fallback Gatekeeper pass because no Gatekeeper service is available
- **THEN** the result SHOULD still include the default rule set version and an empty or default rule metadata summary
### Requirement: Gatekeeper rule catalog SHALL remain deterministic
The Gatekeeper rule catalog SHALL configure metadata and simple parameters only; validation behavior SHALL remain deterministic Java code.
#### Scenario: Rule metadata is lightweight
- **WHEN** rule metadata is loaded
- **THEN** each enabled rule SHOULD expose an id, description, enabled flag, and default severity or relevant parameter
- **AND** rule metadata SHALL NOT execute dynamic scripts
+27 -2
View File
@@ -2,9 +2,7 @@
## Purpose
Provide a repeatable offline evaluation harness for MVP diagnosis Agent behavior, so prompt, tool, retrieval, and verifier changes can be checked against fixed trace-based regression cases.
## Requirements
### Requirement: Evaluation harness SHALL define fixed diagnosis cases
The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria.
@@ -170,3 +168,30 @@ The evaluation harness SHALL detect configured unsafe or unsupported claim text
#### Scenario: Composer fallback still avoids raw JSON leakage
- **WHEN** a trace records Composer fallback rendering
- **THEN** the evaluator SHALL still enforce final-answer raw JSON leakage checks
### Requirement: Evaluation harness SHALL expose an evidence-pipeline matrix
The evaluation harness SHALL include fixed cases that demonstrate the current Chat evidence pipeline across positive evidence, narrow-scope observation, no-evidence negative observation, low-confidence filtering, reject safety, and composer fallback behavior.
#### Scenario: Matrix cases are fixture backed
- **WHEN** the fixed diagnosis case file is evaluated
- **THEN** each matrix case SHALL resolve to an offline trace fixture
- **AND** evaluation SHALL not require a live LLM or running application
#### Scenario: Matrix cases preserve V2 audit closure
- **WHEN** a matrix case requires V2 audit closure
- **THEN** its fixture SHALL include `gatekeeper_result`
- **AND** it SHALL include `claim_checks`
- **AND** it SHALL include `composer_output`
### Requirement: Evaluation harness SHALL validate Gatekeeper rule set version when requested
The evaluation harness SHALL be able to assert the Gatekeeper rule set version recorded in a fixture.
#### Scenario: Expected rule set version matches
- **WHEN** an evaluation case declares `expectedGatekeeperRuleSetVersion`
- **AND** the fixture has the same `selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version`
- **THEN** the rule set version check SHALL pass
#### Scenario: Expected rule set version mismatches
- **WHEN** an evaluation case declares `expectedGatekeeperRuleSetVersion`
- **AND** the fixture has a different or missing rule set version
- **THEN** the case SHALL fail with a clear failed check
@@ -98,3 +98,17 @@ without requiring raw JSON inspection first.
planner `read_skill` text mentions, executor `read_skill` text mentions, and
verifier `read_skill` text mentions
### Requirement: MVP demo SHALL provide stable evidence-pipeline scenarios
The MVP demo SHALL provide stable scenarios that explain how to demonstrate positive evidence, no-evidence, and safety/reject behavior for an Agent engineering interview.
#### Scenario: Scenario guide maps demo inputs to evidence claims
- **WHEN** a reviewer opens the demo scenario guide
- **THEN** it SHALL list the supported positive, no-evidence, and safety/reject scenarios
- **AND** it SHALL map each scenario to a request payload or fixture id
- **AND** it SHALL describe the expected Gatekeeper, Verifier, Composer, and trace fields to inspect
#### Scenario: Live and fixture-backed scenarios are distinguished
- **WHEN** a demo scenario is fixture-backed rather than live-scripted
- **THEN** the documentation SHALL say so explicitly
- **AND** it SHALL avoid promising deterministic live LLM output for that scenario