50 lines
3.2 KiB
Markdown
50 lines
3.2 KiB
Markdown
## Context
|
|
|
|
Executor Structured Output V2 introduced three audit layers after the original fixed-case eval harness was created:
|
|
|
|
- Gatekeeper result in `selfEvaluation.verifier_evaluation.gatekeeper_result`
|
|
- Verifier claim-level checks in `selfEvaluation.verifier_evaluation.claim_checks`
|
|
- Composer result in `selfEvaluation.verifier_evaluation.composer_output`
|
|
|
|
The existing harness proves that a saved trace has expected answer keywords, evidence tools, verdicts, and basic structured claim bindings. It does not yet prove that the V2 audit chain is internally consistent or that Composer filtered unsupported claims before writing the final user answer.
|
|
|
|
## Goals / Non-Goals
|
|
|
|
**Goals:**
|
|
|
|
- Make V2 audit fields part of offline regression checks.
|
|
- Catch physical/protocol regressions: missing Gatekeeper audit, PASS after Gatekeeper fail, missing claim checks, missing Composer audit, raw Executor JSON leakage, and unsupported claims leaking into the final answer.
|
|
- Add fixture cases that exercise the new checks without starting the application.
|
|
- Keep the evaluator deterministic and easy to explain to another implementation agent.
|
|
|
|
**Non-Goals:**
|
|
|
|
- Do not redesign Planner, Executor, Gatekeeper, Verifier, or Composer.
|
|
- Do not introduce an LLM-based evaluator.
|
|
- Do not require live MySQL, Redis, Milvus, or application startup for the baseline fixture tests.
|
|
- Do not add a broad new database schema; trace audit remains read from existing JSON fields.
|
|
|
|
## Decisions
|
|
|
|
1. Extend eval case expectations instead of hardcoding every V2 rule globally.
|
|
|
|
Some legacy or intentionally partial fixtures may not contain all V2 audit fields. Case-level expectations let the fixed baseline explicitly say which trace must prove Gatekeeper, claim checks, Composer, or leakage prevention. The default can remain backward-compatible while new V2 fixtures opt in to stricter checks.
|
|
|
|
2. Keep consistency checks lexical and deterministic.
|
|
|
|
The evaluator will not determine semantic truth from scratch. It will compare configured keywords against the final answer and audit fields. This is enough to catch the intended regression class: unsupported or contradicted claims being rendered as confirmed final answers.
|
|
|
|
3. Treat Gatekeeper failure as a hard regression if paired with `PASS`.
|
|
|
|
The production chain already guards this. The eval harness should independently fail any trace where `gatekeeper_result.status=fail` and Verifier verdict remains `PASS`, because that means the audit layer can no longer be trusted.
|
|
|
|
4. Record audit signals in eval results.
|
|
|
|
Per-case results should expose the observed Gatekeeper status, Composer status, and claim-check count so baseline JSON/Markdown reports remain useful during review.
|
|
|
|
## Risks / Trade-offs
|
|
|
|
- [Risk] Keyword-based final-answer checks can miss paraphrases. -> Mitigation: use them only for regression-sensitive fixtures where the unsafe claim keyword is deliberately fixed.
|
|
- [Risk] Adding too many case fields makes fixtures harder to maintain. -> Mitigation: keep expectation fields small and optional.
|
|
- [Risk] Baseline report changes may look like a product behavior change. -> Mitigation: document that this phase changes only eval fixtures/rules unless a production bug is discovered and fixed.
|