72 lines
3.0 KiB
Markdown
72 lines
3.0 KiB
Markdown
# Decisions: executor-evidence-output-contract
|
|
|
|
## sm-flow Progress
|
|
|
|
### Clarify
|
|
|
|
Entry summary: recent Chat diagnosis sessions are `LOW_CONFID` because Executor presents unsupported or weakly supported details as confirmed facts after successful tool calls.
|
|
|
|
Slug: `executor-evidence-output-contract`
|
|
|
|
Scale: standard. This affects prompts, verifier input assembly, parsing behavior, and tests, but does not require a database schema change.
|
|
|
|
### Context
|
|
|
|
Relevant history:
|
|
|
|
- `executor-action-memory-relevance`: Executor already has retrieval quality constraints and should avoid repeated `lookup_knowledge`.
|
|
- `chat-verifier-agent`: Verifier should not see intermediate reasoning; it receives explicit `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
|
|
- `evidence-trace-hardening`: evidence-bearing tools persist stable traces and no-evidence semantics.
|
|
- `modular-rag-pipeline`: `lookup_knowledge` exposes evidence blocks and context packs; L0 hints are not fact evidence.
|
|
|
|
Current code shape:
|
|
|
|
- `src/main/resources/prompts/chat-executor-prompt.md` is the Chat Executor prompt.
|
|
- `src/main/resources/prompts/executor-prompt.md` is for the AiOps flow and is not the target of this Chat change.
|
|
- `VerifierInputHook` currently builds a payload with `original_query`, `executor_final_answer`, `tool_trace_summary`, and `retry_context`.
|
|
- Verifier prompt currently extracts facts from `executor_final_answer`.
|
|
|
|
### Grill
|
|
|
|
Question: Should Executor output only JSON or JSON plus readable answer?
|
|
|
|
Decision: use one JSON object containing both machine fields and `user_facing_answer`. This avoids losing a readable Chinese answer while giving Verifier structured claims.
|
|
|
|
Question: Should evidence binding use `chunk_id`?
|
|
|
|
Decision: no. Use generic binding fields because `query_logs` and `query_metrics` do not naturally expose RAG chunks.
|
|
|
|
Question: Should Verifier trust Executor-provided claims completely?
|
|
|
|
Decision: no. Verifier should verify structured claims first, then scan `user_facing_answer` for extra confirmed-sounding facts omitted from `claims`.
|
|
|
|
Question: What happens when Executor JSON is malformed?
|
|
|
|
Decision: preserve raw final answer, mark parse failure, and fall back to existing natural-language verification.
|
|
|
|
### Specify
|
|
|
|
OpenSpec artifacts:
|
|
|
|
- `proposal.md`: why and scope
|
|
- `design.md`: contract, verifier behavior, risks
|
|
- `specs/chat-verifier-agent/spec.md`: modified and added requirements
|
|
- `tasks.md`: implementation checklist
|
|
|
|
### Audit
|
|
|
|
Cross-artifact alignment:
|
|
|
|
- Issue describes evidence attribution hallucination.
|
|
- Proposal scopes the fix to Executor output and Verifier consumption.
|
|
- Design preserves existing verifier isolation.
|
|
- Spec adds observable behavior without changing database schema.
|
|
- Tasks remain implementation-oriented and unchecked.
|
|
|
|
Interface impact:
|
|
|
|
- Prompt/output contract: L2 internal Agent contract change.
|
|
- Verifier payload: L2 internal structured input extension.
|
|
- Database schema: no change.
|
|
- External HTTP API: no intended change.
|