feat(eval): add evidence pipeline acceptance closure

This commit is contained in:
aruo
2026-07-09 00:47:48 +08:00
parent a77c947cd4
commit db0f229285
46 changed files with 1434 additions and 53 deletions
+26 -1
View File
@@ -1,4 +1,4 @@
# chat-verifier-agent Specification
# chat-verifier-agent Specification
## Purpose
TBD - created by archiving change chat-verifier-agent. Update Purpose after archive.
@@ -531,3 +531,28 @@ The Executor prompt SHALL instruct Executor to keep narrow confirmation question
- **THEN** Executor SHOULD output the minimum necessary claims, normally one and at most two
- **AND** those claims SHALL be `observation` or `negative_observation` unless current-session evidence proves more
- **AND** Executor SHALL NOT emit unrelated root-cause, remediation, or excluded-topic claims as confirmed facts
### Requirement: Gatekeeper SHALL expose rule catalog metadata in audit output
Gatekeeper SHALL include the rule catalog version and enabled rule metadata in its validation result.
#### Scenario: Gatekeeper pass includes rule metadata
- **WHEN** Gatekeeper returns `status=pass`
- **THEN** the result SHALL include `rule_set_version`
- **AND** it SHALL include a rule metadata summary
#### Scenario: Gatekeeper fail includes rule metadata
- **WHEN** Gatekeeper returns `status=fail`
- **THEN** the result SHALL include `rule_set_version`
- **AND** it SHALL include a rule metadata summary
#### Scenario: Gatekeeper fallback pass includes rule metadata
- **WHEN** the Verifier input hook returns a fallback Gatekeeper pass because no Gatekeeper service is available
- **THEN** the result SHOULD still include the default rule set version and an empty or default rule metadata summary
### Requirement: Gatekeeper rule catalog SHALL remain deterministic
The Gatekeeper rule catalog SHALL configure metadata and simple parameters only; validation behavior SHALL remain deterministic Java code.
#### Scenario: Rule metadata is lightweight
- **WHEN** rule metadata is loaded
- **THEN** each enabled rule SHOULD expose an id, description, enabled flag, and default severity or relevant parameter
- **AND** rule metadata SHALL NOT execute dynamic scripts
+27 -2
View File
@@ -2,9 +2,7 @@
## Purpose
Provide a repeatable offline evaluation harness for MVP diagnosis Agent behavior, so prompt, tool, retrieval, and verifier changes can be checked against fixed trace-based regression cases.
## Requirements
### Requirement: Evaluation harness SHALL define fixed diagnosis cases
The system SHALL provide a small fixed set of MVP diagnosis evaluation cases with explicit expected trace and answer criteria.
@@ -170,3 +168,30 @@ The evaluation harness SHALL detect configured unsafe or unsupported claim text
#### Scenario: Composer fallback still avoids raw JSON leakage
- **WHEN** a trace records Composer fallback rendering
- **THEN** the evaluator SHALL still enforce final-answer raw JSON leakage checks
### Requirement: Evaluation harness SHALL expose an evidence-pipeline matrix
The evaluation harness SHALL include fixed cases that demonstrate the current Chat evidence pipeline across positive evidence, narrow-scope observation, no-evidence negative observation, low-confidence filtering, reject safety, and composer fallback behavior.
#### Scenario: Matrix cases are fixture backed
- **WHEN** the fixed diagnosis case file is evaluated
- **THEN** each matrix case SHALL resolve to an offline trace fixture
- **AND** evaluation SHALL not require a live LLM or running application
#### Scenario: Matrix cases preserve V2 audit closure
- **WHEN** a matrix case requires V2 audit closure
- **THEN** its fixture SHALL include `gatekeeper_result`
- **AND** it SHALL include `claim_checks`
- **AND** it SHALL include `composer_output`
### Requirement: Evaluation harness SHALL validate Gatekeeper rule set version when requested
The evaluation harness SHALL be able to assert the Gatekeeper rule set version recorded in a fixture.
#### Scenario: Expected rule set version matches
- **WHEN** an evaluation case declares `expectedGatekeeperRuleSetVersion`
- **AND** the fixture has the same `selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version`
- **THEN** the rule set version check SHALL pass
#### Scenario: Expected rule set version mismatches
- **WHEN** an evaluation case declares `expectedGatekeeperRuleSetVersion`
- **AND** the fixture has a different or missing rule set version
- **THEN** the case SHALL fail with a clear failed check
@@ -98,3 +98,17 @@ without requiring raw JSON inspection first.
planner `read_skill` text mentions, executor `read_skill` text mentions, and
verifier `read_skill` text mentions
### Requirement: MVP demo SHALL provide stable evidence-pipeline scenarios
The MVP demo SHALL provide stable scenarios that explain how to demonstrate positive evidence, no-evidence, and safety/reject behavior for an Agent engineering interview.
#### Scenario: Scenario guide maps demo inputs to evidence claims
- **WHEN** a reviewer opens the demo scenario guide
- **THEN** it SHALL list the supported positive, no-evidence, and safety/reject scenarios
- **AND** it SHALL map each scenario to a request payload or fixture id
- **AND** it SHALL describe the expected Gatekeeper, Verifier, Composer, and trace fields to inspect
#### Scenario: Live and fixture-backed scenarios are distinguished
- **WHEN** a demo scenario is fixture-backed rather than live-scripted
- **THEN** the documentation SHALL say so explicitly
- **AND** it SHALL avoid promising deterministic live LLM output for that scenario