feat(agent): add executor evidence output contract
This commit is contained in:
@@ -0,0 +1,71 @@
|
||||
# Decisions: executor-evidence-output-contract
|
||||
|
||||
## sm-flow Progress
|
||||
|
||||
### Clarify
|
||||
|
||||
Entry summary: recent Chat diagnosis sessions are `LOW_CONFID` because Executor presents unsupported or weakly supported details as confirmed facts after successful tool calls.
|
||||
|
||||
Slug: `executor-evidence-output-contract`
|
||||
|
||||
Scale: standard. This affects prompts, verifier input assembly, parsing behavior, and tests, but does not require a database schema change.
|
||||
|
||||
### Context
|
||||
|
||||
Relevant history:
|
||||
|
||||
- `executor-action-memory-relevance`: Executor already has retrieval quality constraints and should avoid repeated `lookup_knowledge`.
|
||||
- `chat-verifier-agent`: Verifier should not see intermediate reasoning; it receives explicit `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
|
||||
- `evidence-trace-hardening`: evidence-bearing tools persist stable traces and no-evidence semantics.
|
||||
- `modular-rag-pipeline`: `lookup_knowledge` exposes evidence blocks and context packs; L0 hints are not fact evidence.
|
||||
|
||||
Current code shape:
|
||||
|
||||
- `src/main/resources/prompts/chat-executor-prompt.md` is the Chat Executor prompt.
|
||||
- `src/main/resources/prompts/executor-prompt.md` is for the AiOps flow and is not the target of this Chat change.
|
||||
- `VerifierInputHook` currently builds a payload with `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
|
||||
- Verifier prompt currently extracts facts from `executor_final_answer`.
|
||||
|
||||
### Grill
|
||||
|
||||
Question: Should Executor output only JSON or JSON plus readable answer?
|
||||
|
||||
Decision: use one JSON object containing both machine fields and `user_facing_answer`. This avoids losing a readable Chinese answer while giving Verifier structured claims.
|
||||
|
||||
Question: Should evidence binding use `chunk_id`?
|
||||
|
||||
Decision: no. Use generic binding fields because `query_logs` and `query_metrics` do not naturally expose RAG chunks.
|
||||
|
||||
Question: Should Verifier trust Executor-provided claims completely?
|
||||
|
||||
Decision: no. Verifier should verify structured claims first, then scan `user_facing_answer` for extra confirmed-sounding facts omitted from `claims`.
|
||||
|
||||
Question: What happens when Executor JSON is malformed?
|
||||
|
||||
Decision: preserve raw final answer, mark parse failure, and fall back to existing natural-language verification.
|
||||
|
||||
### Specify
|
||||
|
||||
OpenSpec artifacts:
|
||||
|
||||
- `proposal.md`: why and scope
|
||||
- `design.md`: contract, verifier behavior, risks
|
||||
- `specs/chat-verifier-agent/spec.md`: modified and added requirements
|
||||
- `tasks.md`: implementation checklist
|
||||
|
||||
### Audit
|
||||
|
||||
Cross-artifact alignment:
|
||||
|
||||
- Issue describes evidence attribution hallucination.
|
||||
- Proposal scopes the fix to Executor output and Verifier consumption.
|
||||
- Design preserves existing verifier isolation.
|
||||
- Spec adds observable behavior without changing database schema.
|
||||
- Tasks remain implementation-oriented and unchecked.
|
||||
|
||||
Interface impact:
|
||||
|
||||
- Prompt/output contract: L2 internal Agent contract change.
|
||||
- Verifier payload: L2 internal structured input extension.
|
||||
- Database schema: no change.
|
||||
- External HTTP API: no intended change.
|
||||
Reference in New Issue
Block a user