feat(agent): add executor evidence output contract
This commit is contained in:
@@ -5,6 +5,7 @@
|
||||
| 日期 | slug | 领域 | 关键词 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| 2026-07-05 | diagnosis-playbook-skills | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented |
|
||||
| 2026-07-07 | executor-evidence-output-contract | Chat质量门禁/证据归因 | Executor structured output, evidence bindings, Verifier structured claims, LOW_CONFID, hallucination | openspec/changes/executor-evidence-output-contract | proposed |
|
||||
| 2026-07-06 | rag-eval-pipeline-closure | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
|
||||
| 2026-07-06 | modular-rag-pipeline | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
|
||||
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
|
||||
|
||||
@@ -0,0 +1,16 @@
|
||||
# Acceptance: executor-evidence-output-contract
|
||||
|
||||
## Draft Acceptance
|
||||
|
||||
- [x] Issue exists: `mvp/issues/executor-evidence-attribution-hallucination.md`.
|
||||
- [x] OpenSpec change artifacts exist.
|
||||
- [x] devflow tracking files exist.
|
||||
- [x] OpenSpec validation passes.
|
||||
- [x] Implementation updates Chat Executor prompt.
|
||||
- [x] Implementation passes structured Executor output to Verifier.
|
||||
- [x] Verifier prefers structured claims and still falls back safely.
|
||||
- [x] Focused tests cover parsing, payload assembly, and unsupported confirmed claims.
|
||||
|
||||
## Notes
|
||||
|
||||
This project is currently in proposal/design stage. Runtime code is intentionally not changed yet.
|
||||
@@ -0,0 +1,23 @@
|
||||
# Brief: executor-evidence-output-contract
|
||||
|
||||
## Summary
|
||||
|
||||
Executor currently returns natural-language diagnosis answers that may mix confirmed evidence, runbook guidance, historical patterns, and unsupported inference. Verifier catches many unsupported facts, but only after extracting claims from prose.
|
||||
|
||||
This project defines a structured Executor evidence-attribution contract and updates the Verifier input/verification path to consume it.
|
||||
|
||||
## Goal
|
||||
|
||||
Make Chat Executor output machine-checkable so confirmed claims are explicitly bound to current-session evidence, while hypotheses and evidence gaps remain visibly separate.
|
||||
|
||||
## Scope
|
||||
|
||||
- Chat Executor prompt contract.
|
||||
- Executor structured output parsing.
|
||||
- Verifier payload extension.
|
||||
- Chat Verifier prompt behavior.
|
||||
- Focused tests/eval fixtures.
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/executor-evidence-output-contract/`
|
||||
@@ -0,0 +1,71 @@
|
||||
# Decisions: executor-evidence-output-contract
|
||||
|
||||
## sm-flow Progress
|
||||
|
||||
### Clarify
|
||||
|
||||
Entry summary: recent Chat diagnosis sessions are `LOW_CONFID` because Executor presents unsupported or weakly supported details as confirmed facts after successful tool calls.
|
||||
|
||||
Slug: `executor-evidence-output-contract`
|
||||
|
||||
Scale: standard. This affects prompts, verifier input assembly, parsing behavior, and tests, but does not require a database schema change.
|
||||
|
||||
### Context
|
||||
|
||||
Relevant history:
|
||||
|
||||
- `executor-action-memory-relevance`: Executor already has retrieval quality constraints and should avoid repeated `lookup_knowledge`.
|
||||
- `chat-verifier-agent`: Verifier should not see intermediate reasoning; it receives explicit `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
|
||||
- `evidence-trace-hardening`: evidence-bearing tools persist stable traces and no-evidence semantics.
|
||||
- `modular-rag-pipeline`: `lookup_knowledge` exposes evidence blocks and context packs; L0 hints are not fact evidence.
|
||||
|
||||
Current code shape:
|
||||
|
||||
- `src/main/resources/prompts/chat-executor-prompt.md` is the Chat Executor prompt.
|
||||
- `src/main/resources/prompts/executor-prompt.md` is for the AiOps flow and is not the target of this Chat change.
|
||||
- `VerifierInputHook` currently builds a payload with `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
|
||||
- Verifier prompt currently extracts facts from `executor_final_answer`.
|
||||
|
||||
### Grill
|
||||
|
||||
Question: Should Executor output only JSON or JSON plus readable answer?
|
||||
|
||||
Decision: use one JSON object containing both machine fields and `user_facing_answer`. This avoids losing a readable Chinese answer while giving Verifier structured claims.
|
||||
|
||||
Question: Should evidence binding use `chunk_id`?
|
||||
|
||||
Decision: no. Use generic binding fields because `query_logs` and `query_metrics` do not naturally expose RAG chunks.
|
||||
|
||||
Question: Should Verifier trust Executor-provided claims completely?
|
||||
|
||||
Decision: no. Verifier should verify structured claims first, then scan `user_facing_answer` for extra confirmed-sounding facts omitted from `claims`.
|
||||
|
||||
Question: What happens when Executor JSON is malformed?
|
||||
|
||||
Decision: preserve raw final answer, mark parse failure, and fall back to existing natural-language verification.
|
||||
|
||||
### Specify
|
||||
|
||||
OpenSpec artifacts:
|
||||
|
||||
- `proposal.md`: why and scope
|
||||
- `design.md`: contract, verifier behavior, risks
|
||||
- `specs/chat-verifier-agent/spec.md`: modified and added requirements
|
||||
- `tasks.md`: implementation checklist
|
||||
|
||||
### Audit
|
||||
|
||||
Cross-artifact alignment:
|
||||
|
||||
- Issue describes evidence attribution hallucination.
|
||||
- Proposal scopes the fix to Executor output and Verifier consumption.
|
||||
- Design preserves existing verifier isolation.
|
||||
- Spec adds observable behavior without changing database schema.
|
||||
- Tasks remain implementation-oriented and unchecked.
|
||||
|
||||
Interface impact:
|
||||
|
||||
- Prompt/output contract: L2 internal Agent contract change.
|
||||
- Verifier payload: L2 internal structured input extension.
|
||||
- Database schema: no change.
|
||||
- External HTTP API: no intended change.
|
||||
@@ -0,0 +1,17 @@
|
||||
# Evidence: executor-evidence-output-contract
|
||||
|
||||
## Repository Evidence
|
||||
|
||||
- `chat-executor-prompt.md` currently requires using real tool data but does not require a structured evidence-attribution output.
|
||||
- `chat-verifier-prompt.md` currently extracts facts from `executor_final_answer` prose and compares them with `tool_trace_summary`.
|
||||
- `VerifierInputHook` currently provides `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
|
||||
- `openspec/specs/chat-verifier-agent/spec.md` already requires explicit verifier inputs, auditable evidence refs, fixed verdicts, and low-confidence handling.
|
||||
- `openspec/specs/evidence-trace-hardening/spec.md` already distinguishes failed, no-hit, deduped, and successful evidence-tool traces.
|
||||
|
||||
## Runtime Evidence From Recent Sessions
|
||||
|
||||
Recent MySQL inspection showed repeated `LOW_CONFID` verifier results with many `no_evidence` facts. Typical unsupported claims included OOM, Full GC frequency, specific slow SQL timings, lock waits, and service-specific timeout details that were not supported by current-session tool traces.
|
||||
|
||||
## Design Evidence
|
||||
|
||||
This change preserves the previous design that Verifier should not inspect intermediate reasoning. The new structured output is still final Executor output, not hidden chain-of-thought.
|
||||
Reference in New Issue
Block a user