Compare commits
26
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
4b9cf7c5cc | ||
|
|
190013c901 | ||
|
|
208a231113 | ||
|
|
99e490f227 | ||
|
|
1460dd1e99 | ||
|
|
42ba204532 | ||
|
|
581daffdad | ||
|
|
a36fe72639 | ||
|
|
30d3296043 | ||
|
|
3578709896 | ||
|
|
f9df94377b | ||
|
|
78c1477198 | ||
|
|
d928a1968a | ||
|
|
027aed1eeb | ||
|
|
26d5529280 | ||
|
|
6fdbd34bab | ||
|
|
52bf0302c6 | ||
|
|
841437fa06 | ||
|
|
9c9a0024d4 | ||
|
|
a6c2d4459c | ||
|
|
da45fa3fb0 | ||
|
|
db0f229285 | ||
|
|
a77c947cd4 | ||
|
|
9a84b3de34 | ||
|
|
7b8c75e571 | ||
|
|
a08672b31e |
@@ -72,9 +72,10 @@
|
||||
|
||||
### SessionContext
|
||||
- 定义:会话上下文数据类,存储在 Redis 中的会话数据
|
||||
- 包含字段:sessionId、userId、businessId、traceId、status、toolCalls、TTL
|
||||
- 包含字段:sessionId、userId、businessId、traceId、status、toolCalls、messageHistory、TTL
|
||||
- 序列化方式:JSON(GenericJackson2JsonRedisSerializer)
|
||||
- 使用场景:多轮对话上下文管理、工具调用历史追踪
|
||||
- 边界:messageHistory 是热路径对话历史缓存,用于下一轮 prompt 上下文;长期审计的问题和答案应落到 Diagnosis Run,而不是依赖 Redis TTL 内的上下文正文。
|
||||
|
||||
### ToolCall
|
||||
- 定义:工具调用记录数据类,追踪 Agent 使用的工具及其结果
|
||||
@@ -87,6 +88,26 @@
|
||||
- 核心方法:createSession、getSession、updateSession、deleteSession、refreshSession、addToolCall
|
||||
- 使用场景:分布式会话管理、Agent 状态维护
|
||||
|
||||
### Chat Session
|
||||
- 定义:一次多轮对话上下文,由 `sessionId` 唯一标识。
|
||||
- 使用场景:保存用户连续对话的上下文窗口、会话状态和最近活跃时间。
|
||||
- 边界:Chat Session 不代表一次诊断执行;同一个 Chat Session 可以包含多次 Diagnosis Run。
|
||||
|
||||
### Diagnosis Run
|
||||
- 定义:一次独立诊断执行,由 `runId` 唯一标识,属于一个 Chat Session。
|
||||
- 使用场景:保存某一轮诊断的 query、answer、status、耗时、token、反馈和自评估结果。
|
||||
- 边界:Diagnosis Run 是 Trace、Feedback 和 Evidence score 的绑定对象;多轮对话中的每次 `/api/chat` 或 `/api/ai_ops` 执行都应创建新的 Diagnosis Run。
|
||||
|
||||
### Diagnosis Trace
|
||||
- 定义:一次 Diagnosis Run 的可回放执行轨迹,由 run 主记录、AgentStep 和 ToolInvocation 聚合形成。
|
||||
- 使用场景:Trace API、Trace UI、Verifier 审计、评测 fixture 和人工排查。
|
||||
- 边界:Diagnosis Trace 是聚合视图,不要求单独的 trace 主表;当前 trace 明细由 `agent_step` 和 `tool_invocation` 表承载。
|
||||
|
||||
### Diagnosis Orchestration Trace
|
||||
- 定义:一次 Diagnosis Run 的紧凑编排审计摘要,记录实际节点路径、条件边原因、技术重试、降级和终止原因。
|
||||
- 使用场景:解释诊断编排为何进入某个节点、为何重试或为何提前终止,并支撑路由验收和人工审计。
|
||||
- 边界:它是 Diagnosis Trace 的编排维度,不是完整事件日志、自评估结果或持久恢复检查点;不保存 Prompt、模型思考、工具原文和完整编排上下文快照。
|
||||
|
||||
### Flyway
|
||||
- 定义:数据库版本迁移工具,管理 SQL 脚本的版本化执行
|
||||
- 配置:spring.flyway.enabled=true, baseline-on-migrate=true
|
||||
@@ -140,14 +161,14 @@
|
||||
- 边界:最终诊断结论必须被 evidence tools 支撑,不能仅由 skill 正文支撑。
|
||||
|
||||
### Verifier Skill Isolation
|
||||
- 定义:Chat Verifier 与 skill 系统隔离,只校验 Executor 答案和 `tool_trace_summary`。
|
||||
- 定义:Chat Verifier 与 skill 系统隔离,只校验 Gatekeeper 投影后的 `verified_executor_output` 和 `verified_evidence`。
|
||||
- 使用场景:防止 Verifier 把 playbook 指令当作事实证据;Verifier 只判断已有证据是否支持结论。
|
||||
- 边界:Verifier 不接收 `skill_catalog`,不暴露 `read_skill`,不读取 `SKILL.md`。
|
||||
- 边界:Verifier 不接收 `skill_catalog`,不暴露 `read_skill`,不读取 `SKILL.md`、完整 `tool_trace_summary` 或未经验真的 Executor 自由文本。
|
||||
|
||||
## Diagnosis Playbook Business Rules
|
||||
|
||||
- Planner 只看 skill metadata,输出 `selected_skill`、`selection_reason` 和 plan。
|
||||
- Executor 才能调用 `read_skill(selected_skill)`,并且读取 skill 后仍必须调用 evidence tools。
|
||||
- Skill 正文不得替代 `lookup_knowledge`、日志、指标或告警数据。
|
||||
- Verifier 只基于 `tool_trace_summary` 校验事实,不基于 skill 正文校验事实。
|
||||
- Verifier 只基于 Gatekeeper 通过的 `verified_executor_output` 和 `verified_evidence` 校验事实,不基于 skill 正文、完整工具 Trace 或未验真输出校验事实。
|
||||
- 当前阶段保留单 active skill 白名单:`diagnose-mysql-connection-pool`。
|
||||
|
||||
+37
-27
@@ -2,30 +2,40 @@
|
||||
|
||||
## 项目
|
||||
|
||||
| 日期 | slug | 领域 | 关键词 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| 2026-07-05 | diagnosis-playbook-skills | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented |
|
||||
| 2026-07-07 | executor-evidence-output-contract | Chat质量门禁/证据归因 | Executor structured output, evidence bindings, Verifier structured claims, LOW_CONFID, hallucination | openspec/changes/archive/2026-07-07-executor-evidence-output-contract | archived |
|
||||
| 2026-07-07 | executor-v2-output-contract | Chat质量门禁/证据归因 | executor_evidence_v2, user_facing_answer removal, diagnosis_summary removal, structured renderer | openspec/changes/archive/2026-07-07-executor-v2-output-contract | archived |
|
||||
| 2026-07-07 | executor-gatekeeper-hook | Chat质量门禁/证据归因 | Gatekeeper, verifier payload, source_invocation_ids, tool_name match, self_evaluation | openspec/changes/archive/2026-07-07-executor-gatekeeper-hook | archived |
|
||||
| 2026-07-07 | executor-verifier-claim-checks | Chat质量门禁/证据归因 | Verifier claim_checks, facts_checked compatibility, effective verdict guardrail, malformed output downgrade | openspec/changes/archive/2026-07-07-executor-verifier-claim-checks | archived |
|
||||
| 2026-07-08 | executor-composer-final-answer | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
|
||||
| 2026-07-06 | rag-eval-pipeline-closure | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
|
||||
| 2026-07-06 | modular-rag-pipeline | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
|
||||
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
|
||||
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
|
||||
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
|
||||
| 2026-07-04 | diagnosis-eval-harness | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
|
||||
| 2026-07-04 | evidence-trace-hardening | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
|
||||
| 2026-07-03 | mvp-demo-trace-acceptance | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
|
||||
| 2026-07-04 | aiops-traceable-diagnosis-entry | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
|
||||
| 2026-07-04 | aiops-alert-scope-control | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
|
||||
| 2026-05-29 | chatmodel-abstraction | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | archived |
|
||||
| 2026-06-23 | phase1-infrastructure | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | archived |
|
||||
| 2026-06-24 | lookup-knowledge-integration | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | archived |
|
||||
| 2026-06-25 | doc-management-ui | 前端开发/文档管理 | 文档管理页面, CRUD, 状态监控, 纯静态页面, API集成 | archived |
|
||||
| 2026-06-26 | session-storage | 会话存储/可观测 | diagnosis_session, agent_step, tool_invocation, token追踪, 多Agent路由 | openspec/changes/session-storage | archived |
|
||||
| 2026-06-29 | confidence-feedback | 质量评估/反馈机制 | evidence_score, selfEvaluation, feedback, useful, not_useful, case_library, BAD_CASE, tool_invocation规则引擎, 反馈按钮, sessionId回传 | openspec/changes/confidence-feedback | archived |
|
||||
| 2026-06-30 | session-dedup-knowledge-map | 去重/知识图谱 | RetrievedDocTracker, KnowledgeDomainService, knowledge_domain, covers, whenToRetrieve, Planner注入, ISS-001 | openspec/changes/archive/2026-06-30-session-dedup-knowledge-map | archived |
|
||||
| 2026-07-01 | executor-action-memory-relevance | 检索质量/行动记忆 | relevanceLevel, completenessHint, Min-Max归一化, RetrievedDocTracker域级记录, Executor检索约束, ISS-002 | openspec/changes/archive/2026-07-01-executor-action-memory-relevance | archived |
|
||||
| 2026-07-02 | chat-verifier-agent | Chat质量门禁/可追溯验证 | Verifier, groundedness_score, facts_checked, evidence_refs, tool_trace_summary, self_evaluation | openspec/changes/archive/2026-07-03-chat-verifier-agent | archived |
|
||||
| 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|
||||
|---|---|---|---|---|---|---|
|
||||
| 2026-07-17 | chat-diagnosis-stategraph-cleanup-docs | 清理旧诊断编排闭包,对齐当前文档与 demo contract,并完成 ISS-011 最终 live、日志和数据库验收。 | Chat diagnosis orchestration/cleanup | legacy closure, current docs, orchestration trace, Maven E2E, MySQL ownership | openspec/changes/archive/2026-07-20-chat-diagnosis-stategraph-cleanup-docs | archived |
|
||||
| 2026-07-17 | chat-diagnosis-stategraph-test-suite | 建立 Workflow、Node Contract、Chat Integration 三层权威测试体系并退役旧 Hook implementation tests。 | Chat diagnosis orchestration/testing | workflow test, node contract, Chat integration, coverage matrix, Hook test retirement | openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-test-suite | archived |
|
||||
| 2026-07-17 | chat-diagnosis-stategraph-chatservice-cutover | 将复杂 Chat 单轨切换到 Diagnosis StateGraph,并增加 Run 级 orchestration trace 和 verified-only Verifier 输入。 | Chat diagnosis orchestration/production cutover | ChatService, CompiledGraph stream, runId metadata, orchestration trace, verified-only prompt, V012 | openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-chatservice-cutover | archived |
|
||||
| 2026-07-17 | chat-diagnosis-stategraph-real-nodes | 接入真实 Agent/Java Nodes、显式 Gatekeeper、可信输入投影、关键证据补查与安全 Fallback,暂不切换生产入口。 | Chat diagnosis orchestration/nodes | ReactAgent adapter, Gatekeeper node, verified input, evidence retry, safe fallback, CompiledGraph | openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-real-nodes | archived |
|
||||
| 2026-07-17 | chat-diagnosis-stategraph-routing-skeleton | 实现未接生产入口的 Diagnosis StateGraph 骨架、有限路由和 Fake Node 测试。 | Chat diagnosis orchestration/graph | StateGraph, fake node, conditional edge, retry counter, orchestration events, trace builder | openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-routing-skeleton | archived |
|
||||
| 2026-07-17 | chat-diagnosis-stategraph-design-freeze | 冻结 ISS-011 的 Graph State、条件边、有限重试、安全降级、审计和测试迁移边界。 | Chat diagnosis orchestration/design | StateGraph, runId, Gatekeeper, verified evidence, fallback, orchestration trace, test migration | openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-design-freeze | archived |
|
||||
| 2026-07-10 | session-run-trace-isolation | 拆分会话态和运行态,引入 runId 隔离 Trace、Feedback、AIOps 和 demo 链路。 | Trace/session/run isolation | chat_session, diagnosis_run, runId, trace exact run, feedback fallback, AIOps SSE metadata, baseline drift | openspec/changes/archive/2026-07-10-session-run-trace-isolation | archived |
|
||||
| 2026-07-09 | interview-demo-quality-audit | 增加面试演示前置质量审计,覆盖 prompt、Gatekeeper 和评测基线。 | Agent eval/demo/Prompt audit | interview demo preflight, prompt_audit, gatekeeper rules, diagnosis baseline, 12 fixtures | openspec/changes/archive/2026-07-09-interview-demo-quality-audit | archived |
|
||||
| 2026-07-08 | executor-composer-final-answer | 引入 Composer 生成最终回答,只使用 Verifier 允许的结论材料。 | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
|
||||
| 2026-07-08 | diagnosis-eval-demo-gatekeeper-closure | 收敛诊断评测、稳定 demo 场景和 Gatekeeper 审计元数据。 | Agent eval/demo/Gatekeeper | diagnosis eval matrix, stable demo scenarios, Gatekeeper rule set version, audit metadata | openspec/changes/archive/2026-07-08-diagnosis-eval-demo-gatekeeper-closure | archived |
|
||||
| 2026-07-08 | verifier-evidence-reference-fidelity | 强化 Verifier 对 evidence_refs、raw_path 和 no_evidence 的保真校验。 | Chat质量门禁/证据归因 | evidence_refs, raw_path, Gatekeeper severity, verifier evidence excerpt, HikariCP mock, no_evidence | openspec/changes/archive/2026-07-08-verifier-evidence-reference-fidelity | archived |
|
||||
| 2026-07-07 | executor-evidence-output-contract | 设计 Executor 结构化证据输出,解决证据归因幻觉和 LOW_CONFID 问题。 | Chat质量门禁/证据归因 | Executor structured output, evidence bindings, Verifier structured claims, LOW_CONFID, hallucination | openspec/changes/archive/2026-07-07-executor-evidence-output-contract | archived |
|
||||
| 2026-07-07 | executor-v2-output-contract | 将 Executor 输出升级为 V2 契约,移除面向用户的最终回答字段。 | Chat质量门禁/证据归因 | executor_evidence_v2, user_facing_answer removal, diagnosis_summary removal, structured renderer | openspec/changes/archive/2026-07-07-executor-v2-output-contract | archived |
|
||||
| 2026-07-07 | executor-gatekeeper-hook | 在 Executor 与 Verifier 之间接入 Gatekeeper,校验证据绑定来源。 | Chat质量门禁/证据归因 | Gatekeeper, verifier payload, source_invocation_ids, tool_name match, self_evaluation | openspec/changes/archive/2026-07-07-executor-gatekeeper-hook | archived |
|
||||
| 2026-07-07 | executor-verifier-claim-checks | 增加 Verifier claim_checks 和事实校验兼容逻辑。 | Chat质量门禁/证据归因 | Verifier claim_checks, facts_checked compatibility, effective verdict guardrail, malformed output downgrade | openspec/changes/archive/2026-07-07-executor-verifier-claim-checks | archived |
|
||||
| 2026-07-06 | rag-eval-pipeline-closure | 建立 RAG 评测闭环,加入 fixture、快照和 baseline diff。 | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
|
||||
| 2026-07-06 | modular-rag-pipeline | 将 lookup_knowledge 改造成模块化 RAG 管线,补齐证据块和检索追踪。 | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
|
||||
| 2026-07-05 | diagnosis-playbook-skills | 增加诊断 Playbook Skill,沉淀支付超时、MySQL 池、Redis 超时等套路。 | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented |
|
||||
| 2026-07-05 | mvp-demo-interview-runbook | 准备可复现的 MVP 面试演示包、运行手册和 Trace 检查清单。 | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
|
||||
| 2026-07-05 | diagnosis-eval-baseline-diff | 增加诊断评测 baseline diff,用于判断回归和证据覆盖变化。 | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
|
||||
| 2026-07-04 | expand-diagnosis-eval-fixtures | 扩充诊断评测 fixture,覆盖 Redis、慢响应和 JVM 内存风险。 | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
|
||||
| 2026-07-04 | diagnosis-eval-harness | 建立固定诊断评测 Harness,输出 trace、证据覆盖和 verdict 分布。 | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
|
||||
| 2026-07-04 | evidence-trace-hardening | 强化工具调用证据链、降级契约和离线验证能力。 | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
|
||||
| 2026-07-04 | aiops-traceable-diagnosis-entry | 增加可追踪的 AIOps 告警诊断入口,打通 sessionId 和 Trace API。 | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
|
||||
| 2026-07-04 | aiops-alert-scope-control | 收敛 AIOps 告警诊断范围,区分 payload 定向和自动发现模式。 | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
|
||||
| 2026-07-03 | mvp-demo-trace-acceptance | 增加 MVP demo 的 Trace 验收,覆盖会话、步骤、工具和反馈链路。 | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
|
||||
| 2026-07-02 | chat-verifier-agent | 增加 Chat Verifier Agent,用 groundedness 和 evidence_refs 校验回答。 | Chat质量门禁/可追溯验证 | Verifier, groundedness_score, facts_checked, evidence_refs, tool_trace_summary, self_evaluation | openspec/changes/archive/2026-07-03-chat-verifier-agent | archived |
|
||||
| 2026-07-01 | executor-action-memory-relevance | 增加行动记忆和相关性信号,约束 Executor 重复检索。 | 检索质量/行动记忆 | relevanceLevel, completenessHint, Min-Max归一化, RetrievedDocTracker域级记录, Executor检索约束, ISS-002 | openspec/changes/archive/2026-07-01-executor-action-memory-relevance | archived |
|
||||
| 2026-06-30 | session-dedup-knowledge-map | 引入会话级去重和知识域地图,减少重复召回。 | 去重/知识图谱 | RetrievedDocTracker, KnowledgeDomainService, knowledge_domain, covers, whenToRetrieve, Planner注入, ISS-001 | openspec/changes/archive/2026-06-30-session-dedup-knowledge-map | archived |
|
||||
| 2026-06-29 | confidence-feedback | 建立质量评估和用户反馈机制,并把有用反馈沉淀为案例。 | 质量评估/反馈机制 | evidence_score, selfEvaluation, feedback, useful, not_useful, case_library, BAD_CASE, tool_invocation规则引擎, 反馈按钮, sessionId回传 | openspec/changes/confidence-feedback | archived |
|
||||
| 2026-06-26 | session-storage | 建立通用会话存储,记录 session、agent step 和 tool invocation。 | 会话存储/可观测 | diagnosis_session, agent_step, tool_invocation, token追踪, 多Agent路由 | openspec/changes/session-storage | archived |
|
||||
| 2026-06-25 | doc-management-ui | 实现文档管理页面,支持文档 CRUD、状态监控和 API 集成。 | 前端开发/文档管理 | 文档管理页面, CRUD, 状态监控, 纯静态页面, API集成 | - | archived |
|
||||
| 2026-06-24 | lookup-knowledge-integration | 接入知识库检索,支持 L0 精确匹配和 L1 语义检索。 | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | - | archived |
|
||||
| 2026-06-23 | phase1-infrastructure | 搭建第一阶段基础设施,包括 MySQL、Redis、Milvus、Flyway 和 JPA。 | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | - | archived |
|
||||
| 2026-05-29 | chatmodel-abstraction | 抽象 ChatModel 和 EmbeddingModel,支持多模型路由。 | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | - | archived |
|
||||
|
||||
@@ -48,7 +48,7 @@
|
||||
|
||||
## 遗留问题
|
||||
|
||||
ISS-002:Executor 无约束重复调用 `lookup_knowledge`(单会话 20+ 次),knowledge map 和检索约束只注入了 Planner 未注入 Executor。详见 `mvp/issues/ISS-002-executor-unconstrained-lookup.md`。
|
||||
ISS-002:Executor 无约束重复调用 `lookup_knowledge`(单会话 20+ 次),knowledge map 和检索约束只注入了 Planner 未注入 Executor。详见 `mvp/issues/archived/ISS-002-executor-unconstrained-lookup.md`。
|
||||
|
||||
## 已知限制
|
||||
|
||||
|
||||
@@ -9,8 +9,8 @@
|
||||
## Context
|
||||
|
||||
- `devflow/index.md` was checked. Relevant history includes `session-storage`, `confidence-feedback`, `executor-action-memory-relevance`, and `chat-verifier-agent`.
|
||||
- `mvp/notes/agent-engineering-decisions.md` already recommends the next phase as "可复现 MVP Demo", including `mvp-demo` profile, fixed diagnosis case, one-click request, and `GET /api/diagnosis/{sessionId}/trace`.
|
||||
- `mvp/issues/ISS-003-mvp-design-implementation-review.md` identifies test stability, session traceability, verifier evidence chain, upload path, and SupervisorAgent consistency as recent MVP concerns. Security cleanup is intentionally deferred by user decision.
|
||||
- `mvp/archive/2026-07-09-doc-cleanup/notes/agent-engineering-decisions.md` already recommends the next phase as "可复现 MVP Demo", including `mvp-demo` profile, fixed diagnosis case, one-click request, and `GET /api/diagnosis/{sessionId}/trace`.
|
||||
- `mvp/issues/active/ISS-003-mvp-design-implementation-review.md` identifies test stability, session traceability, verifier evidence chain, upload path, and SupervisorAgent consistency as recent MVP concerns. Security cleanup is intentionally deferred by user decision.
|
||||
|
||||
## Question Pool
|
||||
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
## Draft Acceptance
|
||||
|
||||
- [x] Issue exists: `mvp/issues/executor-evidence-attribution-hallucination.md`.
|
||||
- [x] Issue exists: `mvp/issues/active/executor-evidence-attribution-hallucination.md`.
|
||||
- [x] OpenSpec change artifacts exist.
|
||||
- [x] devflow tracking files exist.
|
||||
- [x] OpenSpec validation passes.
|
||||
|
||||
@@ -5,7 +5,7 @@
|
||||
- `devflow/projects/2026-07-07-executor-evidence-output-contract`: V1 evidence-attribution contract kept `user_facing_answer`.
|
||||
- `devflow/projects/2026-07-02-chat-verifier-agent`: Verifier consumes explicit inputs and should not see intermediate reasoning.
|
||||
- `devflow/projects/2026-07-04-evidence-trace-hardening`: evidence summaries and tool invocation references are the evidence foundation.
|
||||
- `mvp/issues/executor-structured-output-v2.md`: staged implementation design; stage one is Executor V2 output contract.
|
||||
- `mvp/issues/design-notes/executor-structured-output-v2.md`: staged implementation design; stage one is Executor V2 output contract.
|
||||
|
||||
## Code Evidence
|
||||
|
||||
|
||||
@@ -0,0 +1,48 @@
|
||||
# diagnosis-eval-demo-gatekeeper-closure Acceptance
|
||||
|
||||
## Static / Structure Verification
|
||||
|
||||
- `cmd /c openspec validate diagnosis-eval-demo-gatekeeper-closure --strict`
|
||||
- Result: passed.
|
||||
- `cmd /c openspec validate --specs`
|
||||
- Result: passed, 10 specs passed.
|
||||
|
||||
## Script Verification
|
||||
|
||||
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,VerifierInputHookTest" test`
|
||||
- Result: 36 tests, 0 failures, 0 errors.
|
||||
- `mvn "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest,ToolInvocationRecorderTest,QueryLogsToolsTest" test`
|
||||
- Result: 61 tests, 0 failures, 0 errors.
|
||||
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`
|
||||
- Result after E2E startup fix: 23 tests, 0 failures, 0 errors.
|
||||
|
||||
## Live E2E Verification
|
||||
|
||||
- Start command: `mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"`.
|
||||
- Demo command: `powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1`.
|
||||
- Result: chat, trace, and feedback requests completed successfully.
|
||||
- Output files:
|
||||
- `mvp/demo/output/chat-response.json`
|
||||
- `mvp/demo/output/trace-response.json`
|
||||
- `mvp/demo/output/feedback-response.json`
|
||||
- Trace observations:
|
||||
- `hasVerifierEvaluation=true`
|
||||
- `gatekeeper_result.rule_set_version=gatekeeper-rules-v1`
|
||||
|
||||
## Fixed During Verification
|
||||
|
||||
- E2E startup initially failed because Spring could not instantiate `ExecutorGatekeeperService`.
|
||||
- Root cause: two public constructors and no explicit `@Autowired` constructor.
|
||||
- Fix: annotate the production constructor with `@Autowired`.
|
||||
|
||||
## Residual Risk
|
||||
|
||||
- The live payment-timeout path can still produce `LOW_CONFID` because model-generated evidence bindings may omit some explicit `source_invocation_id` values.
|
||||
- This is not a blocker for this change because deterministic matrix behavior is covered by saved fixtures and baseline evaluation.
|
||||
- Existing Maven warnings remain: duplicate `spring-boot-starter-test` declaration and Lombok `@Builder` default warnings.
|
||||
|
||||
## Archive Status
|
||||
|
||||
- Devflow archive artifacts created.
|
||||
- OpenSpec change archived to `openspec/changes/archive/2026-07-08-diagnosis-eval-demo-gatekeeper-closure`.
|
||||
- Main specs synced by `cmd /c openspec archive diagnosis-eval-demo-gatekeeper-closure --yes`.
|
||||
@@ -0,0 +1,33 @@
|
||||
# diagnosis-eval-demo-gatekeeper-closure Brief
|
||||
|
||||
## Background
|
||||
|
||||
The Chat evidence pipeline already had Executor V2 structured output, deterministic Gatekeeper validation, Verifier claim checks, and Composer final rendering. The missing piece was an interview-ready acceptance story that made the anti-hallucination behavior easy to demonstrate and regress.
|
||||
|
||||
## Goal
|
||||
|
||||
Close the next three interview-readiness gaps together:
|
||||
|
||||
- diagnosis eval fixture matrix
|
||||
- stable demo data set
|
||||
- Gatekeeper rule configuration and audit version
|
||||
|
||||
## Scope
|
||||
|
||||
- Expand `mvp/eval` with matrix-oriented cases, fixtures, and baseline reports.
|
||||
- Add stable demo request payloads and scenario documentation.
|
||||
- Add a lightweight local Gatekeeper rule catalog with `rule_set_version` and rule metadata in `gatekeeper_result`.
|
||||
- Update architecture, demo, and eval docs to describe the current implementation.
|
||||
|
||||
## Non-goals
|
||||
|
||||
- No new public HTTP endpoint.
|
||||
- No new database table.
|
||||
- No Planner `scope_contract`.
|
||||
- No Gatekeeper retry loop.
|
||||
- No remote or dynamic rule execution engine.
|
||||
|
||||
## OpenSpec
|
||||
|
||||
- Change: `openspec/changes/diagnosis-eval-demo-gatekeeper-closure`
|
||||
- Interface impact: L2 internal contract change.
|
||||
@@ -0,0 +1,144 @@
|
||||
# diagnosis-eval-demo-gatekeeper-closure Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: implement the next three interview-readiness items together: diagnosis eval fixture matrix, stable demo data set, and Gatekeeper rule configuration/audit version.
|
||||
- Slug: `diagnosis-eval-demo-gatekeeper-closure`
|
||||
- Devflow scale: `standard`
|
||||
- Interface impact: expected L2 internal contract change because `gatekeeper_result` audit JSON will gain rule metadata/version fields.
|
||||
|
||||
## Context
|
||||
|
||||
- `devflow/index.md` used: related entries found for diagnosis eval harness, fixture expansion, MVP demo runbook, Gatekeeper hook, and verifier evidence reference fidelity.
|
||||
- Relevant glossary:
|
||||
- Evidence Tools produce incident facts and must be recorded in `tool_invocation`.
|
||||
- Verifier should not use skills/runbooks as incident evidence.
|
||||
- `tool_invocation.retrieval_details` is the structured evidence/audit home for tool-specific details.
|
||||
- Historical constraints that must enter OpenSpec:
|
||||
- Diagnosis eval is offline and deterministic; no LLM-as-judge.
|
||||
- Demo assets should be runnable, but fixed regression should use saved fixtures.
|
||||
- Gatekeeper remains in the Verifier hook path.
|
||||
- No new database table for Gatekeeper audit; use `self_evaluation.verifier_evaluation.gatekeeper_result`.
|
||||
- `$.no_evidence` is a query no-hit signal, not proof that a problem is impossible.
|
||||
|
||||
## Question Pool
|
||||
|
||||
| ID | Dimension | Mode | Question | Status |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | Terminology | evidence-driven | What names should this change use for the matrix, demo set, and Gatekeeper rule metadata? | Resolved |
|
||||
| Q2 | Boundary | evidence-driven | Should this change alter public APIs, database schema, Planner output, or retry behavior? | Resolved |
|
||||
| Q3 | Acceptance | evidence-driven | Which existing tests and baseline assets define the current acceptance style? | Resolved |
|
||||
| Q4 | Technical | evidence-driven | Where should Gatekeeper rule metadata live with minimal implementation risk? | Pending code research |
|
||||
| Q5 | Scope | user-interview | Should the stable demo set be documentation/payloads only, or should it include live E2E scripts for all scenarios? | Confirmed |
|
||||
|
||||
## Evidence-driven Conclusions
|
||||
|
||||
- Q1 conclusion: use `diagnosis eval matrix`, `stable demo scenarios`, and `Gatekeeper rule set version` as terms.
|
||||
- Q2 conclusion: keep this as an internal contract change. Do not add public endpoints, tables, Planner `scope_contract`, or Gatekeeper retry.
|
||||
- Q3 conclusion: existing `DiagnosisTraceEvaluatorTest`, `ExecutorGatekeeperServiceTest`, `VerifierInputHookTest`, `ToolInvocationRecorderTest`, and `mvp/eval/reports` define the current acceptance style.
|
||||
- Q4 conclusion: Gatekeeper metadata should live behind a small rule catalog loaded by `ExecutorGatekeeperService`; the audit output should include a rule set version and enabled rule metadata summary, without adding tables or remote registry.
|
||||
|
||||
## User-interview Confirmations
|
||||
|
||||
- Q5 confirmed by resumed objective: complete items 1/2/3 with sm-flow, archive, submit, and run end-to-end if necessary.
|
||||
- Implementation interpretation: stable demo scenarios will be fixed request payloads and runbook docs plus deterministic fixture-backed eval. Live E2E remains necessary only for at least one main path or where unit/fixture evidence is insufficient.
|
||||
|
||||
## OpenSpec Backfill
|
||||
|
||||
- Created Draft proposal at `openspec/changes/diagnosis-eval-demo-gatekeeper-closure/proposal.md`.
|
||||
- Context constraints from historical devflow entries were written into the proposal.
|
||||
- Scope confirmation and Gatekeeper catalog placement were written into the proposal/design.
|
||||
|
||||
## Current Checkpoint
|
||||
|
||||
- Discover completed.
|
||||
- No implementation files changed yet.
|
||||
|
||||
## Specify / Alignment
|
||||
|
||||
### Cross-artifact Alignment
|
||||
|
||||
| Check | Status | Notes |
|
||||
|---|---|---|
|
||||
| brief/proposal goals -> proposal | Aligned | Proposal covers eval matrix, stable demo scenarios, and Gatekeeper rule catalog/audit version. |
|
||||
| proposal scope/constraints -> design | Aligned | Design records offline deterministic eval, fixture-backed demo distinction, local rule catalog, and no new table/API. |
|
||||
| design decisions -> specs/tasks | Aligned | Specs cover eval matrix, rule set version validation, demo scenarios, and Gatekeeper rule metadata; tasks cover matching implementation slices. |
|
||||
| specs observable behavior -> tasks | Aligned | Each requirement has an executable task and acceptance check. |
|
||||
|
||||
### Interface Impact
|
||||
|
||||
- Level: L2 internal contract change.
|
||||
- Reason: `gatekeeper_result` internal audit JSON gains `rule_set_version` and rule metadata summary. Eval case/result fields may gain optional rule set checks. No public HTTP API, database schema, or external DTO contract changes.
|
||||
|
||||
## Audit
|
||||
|
||||
Input -> processing -> output chain:
|
||||
|
||||
```text
|
||||
mvp/demo request docs + mvp/eval fixtures
|
||||
-> DiagnosisTraceEvaluator
|
||||
-> baseline reports
|
||||
-> interview/demo evidence
|
||||
|
||||
Gatekeeper rule catalog
|
||||
-> ExecutorGatekeeperService
|
||||
-> VerifierInputHook / ChatService persisted self_evaluation
|
||||
-> Trace and eval audit
|
||||
```
|
||||
|
||||
Architecture risk assessment:
|
||||
|
||||
1. The change is intentionally internal and should not add new public consumers.
|
||||
2. Gatekeeper catalog must stay metadata-only; dynamic rule execution would be a different, riskier architecture.
|
||||
3. Fixture-backed demo scenarios should be documented as deterministic regression artifacts, not live LLM guarantees.
|
||||
4. Baseline report churn is expected and must be committed with case/fixture changes.
|
||||
5. No devflow/OpenSpec conflict found.
|
||||
|
||||
## Commit Gate
|
||||
|
||||
- `cmd /c openspec validate diagnosis-eval-demo-gatekeeper-closure --strict`: passed.
|
||||
- `cmd /c openspec validate --specs`: passed, 10 specs passed.
|
||||
- File completeness:
|
||||
- proposal.md: present.
|
||||
- design.md: present.
|
||||
- specs: present for `diagnosis-eval-harness`, `mvp-demo-trace-acceptance`, `chat-verifier-agent`.
|
||||
- tasks.md: present.
|
||||
- Consistency:
|
||||
- Proposal concepts have corresponding design sections.
|
||||
- Design decisions are reflected in specs/tasks.
|
||||
- Task acceptance checks are verifiable.
|
||||
|
||||
## Current Checkpoint
|
||||
|
||||
- Commit completed.
|
||||
- `.committed` marker created.
|
||||
|
||||
## Apply Verification
|
||||
|
||||
- Focused verification passed:
|
||||
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,VerifierInputHookTest" test`
|
||||
- Result: 36 tests, 0 failures, 0 errors.
|
||||
- Broader relevant regression passed:
|
||||
- `mvn "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest,ToolInvocationRecorderTest,QueryLogsToolsTest" test`
|
||||
- Result: 61 tests, 0 failures, 0 errors.
|
||||
- E2E startup repro found a Spring bean construction issue:
|
||||
- Command: `mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"`
|
||||
- Failure: `ExecutorGatekeeperService` had two public constructors and no annotated constructor, so Spring attempted a no-arg constructor and failed with `No default constructor found`.
|
||||
- Classification: code deviation from OpenSpec implementation intent, not a spec gap.
|
||||
- Fix: annotate the production constructor with `@Autowired`.
|
||||
- Post-fix focused regression passed:
|
||||
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`
|
||||
- Result: 23 tests, 0 failures, 0 errors.
|
||||
- Live E2E passed for demo compatibility:
|
||||
- Start: `mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"`
|
||||
- Run: `powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1`
|
||||
- Result: `/api/chat`, `/api/diagnosis/{sessionId}/trace`, and `/api/feedback` completed successfully.
|
||||
- Trace summary included `hasVerifierEvaluation=true`.
|
||||
- Persisted Gatekeeper audit included `rule_set_version=gatekeeper-rules-v1`.
|
||||
- Residual quality note: the live payment-timeout response remained `LOW_CONFID` because some model-produced evidence bindings still lacked explicit `source_invocation_id`; deterministic PASS/LOW_CONFID/REJECT claims are covered by fixture-backed eval.
|
||||
|
||||
## Archive Readiness
|
||||
|
||||
- OpenSpec tasks 1-4 completed.
|
||||
- Verification is recorded in devflow acceptance artifacts.
|
||||
- Remaining known risk: live LLM output is not deterministic and may still produce LOW_CONFID on the payment-timeout path; this is intentionally documented as demo compatibility, not a fixed PASS guarantee.
|
||||
@@ -0,0 +1,22 @@
|
||||
# diagnosis-eval-demo-gatekeeper-closure Evidence
|
||||
|
||||
## Code And Artifact Evidence
|
||||
|
||||
- Gatekeeper rule metadata lives in `src/main/resources/gatekeeper/gatekeeper-rules.json`.
|
||||
- `ExecutorGatekeeperService` loads the local catalog, uses configured threshold parameters, and emits `rule_set_version` plus enabled rule metadata.
|
||||
- `VerifierInputHook` and `ChatService` preserve Gatekeeper audit metadata in fallback/default paths.
|
||||
- `DiagnosisTraceEvaluator` can optionally validate expected Gatekeeper rule set version.
|
||||
- `mvp/eval/cases/diagnosis-cases.json` now includes narrow-scope and no-evidence matrix cases.
|
||||
- `mvp/eval/reports/baseline-report.json` and `.md` were regenerated for the expanded fixed matrix.
|
||||
- `mvp/demo/evidence-pipeline-scenarios.md` documents live vs fixture-backed demo scenarios.
|
||||
|
||||
## Decisions
|
||||
|
||||
- Keep this phase internal: no public API, no DB schema, no Planner output change.
|
||||
- Keep Gatekeeper deterministic Java validation; the catalog is metadata/config only.
|
||||
- Treat live demo as compatibility evidence and fixture-backed eval as deterministic regression evidence.
|
||||
- Persist audit under the existing `self_evaluation.verifier_evaluation.gatekeeper_result` structure.
|
||||
|
||||
## Runtime Finding
|
||||
|
||||
The first Maven E2E startup found a real integration issue: `ExecutorGatekeeperService` had multiple public constructors without an annotated constructor, so Spring could not instantiate the service. The fix was to annotate the production constructor with `@Autowired`.
|
||||
@@ -0,0 +1,52 @@
|
||||
# Acceptance
|
||||
|
||||
## Static Verification
|
||||
|
||||
- `openspec validate verifier-evidence-reference-fidelity --strict`: passed.
|
||||
- `openspec validate --specs`: passed.
|
||||
|
||||
## Script Verification
|
||||
|
||||
- `mvn "-Dtest=ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,QueryLogsToolsTest,ChatServiceSequentialAgentTest" test`
|
||||
- Passed: 38 tests.
|
||||
- `mvn "-Dtest=ExecutorGatekeeperServiceTest" test`
|
||||
- Passed: 9 tests.
|
||||
- `mvn test`
|
||||
- Failed on unrelated environment-gated `MilvusConnectionTest.connect`: `MILVUS_TOKEN` environment variable was not set.
|
||||
- Other executed tests in the run progressed until that single failure; focused tests for this change passed.
|
||||
|
||||
## End-to-End Verification
|
||||
|
||||
The Java service was restarted with `mvn spring-boot:run`; logs were written under `logs/`.
|
||||
|
||||
| Case | Session | Result | Gatekeeper Audit |
|
||||
|---|---|---|---|
|
||||
| HikariCP positive | `iss007-hikari-positive-20260708-1553` | PASS; confirmed order-service HikariCP timeout and pool saturation logs | `pass / none`, checked_bindings=2 |
|
||||
| HikariCP negative | `iss007-hikari-negative-20260708-1555` | LOW_CONFID; no `generic-service`; no false positive for inventory-service | `fail / low_confid` |
|
||||
| HighMemoryUsage positive | `iss007-memory-positive-20260708-1558` | PASS; confirmed HighMemoryUsage 91%, did not confirm memory leak | `pass / none`, checked_bindings=1 |
|
||||
| SlowResponse positive | `iss007-slow-positive-20260708-1600` | PASS; confirmed SlowResponse and slow request logs, no DB pool root cause | `pass / none`, checked_bindings=7 |
|
||||
| Narrow HighCPUUsage | `iss007-narrow-highcpu-20260708-1602` | PASS; only covered payment-service HighCPUUsage | `pass / none`, checked_bindings=1 |
|
||||
|
||||
## Database Audit
|
||||
|
||||
`scripts/query_mysql.py` was used to verify:
|
||||
|
||||
- `diagnosis_session.self_evaluation.verifier_evaluation.verdict`
|
||||
- `gatekeeper_result.status`
|
||||
- `gatekeeper_result.severity`
|
||||
- `gatekeeper_result.checked_bindings`
|
||||
- no-hit HikariCP query rows persist `evidence_status=no_evidence`
|
||||
|
||||
## Remaining Risk
|
||||
|
||||
- Negative no-hit claims still have incomplete precise references when Executor uses `$.logs` for empty arrays. Gatekeeper correctly downgrades to `LOW_CONFID`.
|
||||
- Prompt-only scope control is improved but not a hard contract. A future `scope_contract` may still be needed.
|
||||
- Full test suite requires `MILVUS_TOKEN` to pass `MilvusConnectionTest`.
|
||||
|
||||
## OpenSpec Archive
|
||||
|
||||
- `openspec archive verifier-evidence-reference-fidelity --yes`: succeeded.
|
||||
- Main specs updated:
|
||||
- `openspec/specs/chat-verifier-agent/spec.md`
|
||||
- `openspec/specs/evidence-trace-hardening/spec.md`
|
||||
- Non-blocking warning: proposal did not use OpenSpec's preferred `## Why` / `## What Changes` headers, but archive completed.
|
||||
@@ -0,0 +1,44 @@
|
||||
# Verifier Evidence Reference Fidelity
|
||||
|
||||
## Background
|
||||
|
||||
ISS-007 came from end-to-end diagnosis cases where raw tool output and Executor `evidence_excerpt` contained enough facts, but Verifier still returned `LOW_CONFID` because the verifier-facing summary compressed away key details.
|
||||
|
||||
The affected flow is:
|
||||
|
||||
```text
|
||||
chat_planner
|
||||
-> chat_executor
|
||||
-> VerifierInputHook / Gatekeeper
|
||||
-> chat_verifier
|
||||
-> chat_composer
|
||||
```
|
||||
|
||||
The change hardens the evidence handoff between Executor, Gatekeeper, and Verifier.
|
||||
|
||||
## Goal
|
||||
|
||||
Make Executor cite concrete tool evidence, make Gatekeeper validate that citation with code, and make Verifier judge whether verified evidence can derive the claim.
|
||||
|
||||
## Scope
|
||||
|
||||
- Persist `tool_invocation.retrieval_details.evidence_refs`.
|
||||
- Use `source_invocation_id + raw_path + evidence_excerpt` as the precise evidence binding.
|
||||
- Add Gatekeeper `severity` and checked binding audit.
|
||||
- Keep Gatekeeper in the Verifier hook path.
|
||||
- Keep `tool_trace_summary` as navigation/audit context, not the only evidence source.
|
||||
- Fix HikariCP mock positive/no-hit behavior.
|
||||
- Tighten Executor/Verifier prompts for narrow-scope evidence handling.
|
||||
|
||||
## Non-goals
|
||||
|
||||
- No Planner `scope_contract`.
|
||||
- No new database table.
|
||||
- No full JSONPath engine.
|
||||
- No change to external HTTP API.
|
||||
- No retry rollback from Gatekeeper to Executor in this phase.
|
||||
|
||||
## OpenSpec
|
||||
|
||||
- Change: `openspec/changes/verifier-evidence-reference-fidelity`
|
||||
- Source issue: `mvp/issues/archived/ISS-007-verifier-evidence-summary-fidelity.md`
|
||||
@@ -0,0 +1,54 @@
|
||||
# Decisions
|
||||
|
||||
## Evidence Reference
|
||||
|
||||
Use `source_invocation_id + raw_path + evidence_excerpt` as the precise evidence reference for Executor claim bindings.
|
||||
|
||||
Reason:
|
||||
|
||||
- Invocation ID alone only identifies a tool call, not the evidence inside it.
|
||||
- `raw_path` is enough for the first version when paired with `retrieval_details.evidence_refs`.
|
||||
- `evidence_excerpt` remains the text Verifier reads, but only after Gatekeeper validates it.
|
||||
|
||||
## Raw Path
|
||||
|
||||
Only support stable locators in the first version:
|
||||
|
||||
- `$.alerts[i]`
|
||||
- `$.logs[i]`
|
||||
- `$.evidence_blocks[i]`
|
||||
|
||||
No full JSONPath engine is introduced.
|
||||
|
||||
## Gatekeeper Severity
|
||||
|
||||
Gatekeeper output includes:
|
||||
|
||||
- `status`
|
||||
- `severity`
|
||||
- `checked_bindings`
|
||||
- `failed_rules`
|
||||
- `warnings`
|
||||
- `errors`
|
||||
|
||||
Severity meaning:
|
||||
|
||||
- `none`: precise references passed.
|
||||
- `low_confid`: evidence is missing or incomplete, but not fabricated.
|
||||
- `reject`: fabricated ID, wrong tool, unknown raw path, or mismatched excerpt.
|
||||
|
||||
## Verifier Boundary
|
||||
|
||||
Verifier uses verified claim-local excerpts as primary derivability evidence. `tool_trace_summary` remains available for navigation and audit, but no longer needs to carry every concrete fact.
|
||||
|
||||
## Hook Placement
|
||||
|
||||
Gatekeeper remains in the Verifier input hook path. This version does not retry Executor on Gatekeeper failure.
|
||||
|
||||
## Planner
|
||||
|
||||
Planner is not changed. `scope_contract` remains a later-stage idea. This phase uses prompt constraints to reduce narrow-scope over-expansion.
|
||||
|
||||
## Database
|
||||
|
||||
No new tables. Evidence refs are stored in `tool_invocation.retrieval_details.evidence_refs`; audit is stored in `diagnosis_session.self_evaluation.verifier_evaluation.gatekeeper_result`.
|
||||
@@ -0,0 +1,33 @@
|
||||
# Evidence
|
||||
|
||||
## Existing Context
|
||||
|
||||
- Existing `chat-verifier-agent` spec still used `source_invocation_ids` and `tool_trace_summary` as the main verifier evidence context.
|
||||
- Existing `evidence-trace-hardening` spec already established `tool_invocation.retrieval_details` as the right place for structured tool-specific facts.
|
||||
- Prior devflow projects established that runbook/skill content is guidance, not incident evidence.
|
||||
|
||||
## Code Findings
|
||||
|
||||
- `VerifierInputHook` previously backfilled plural `source_invocation_ids` from `tool_trace_summary` by tool name.
|
||||
- `ExecutorGatekeeperService` previously validated invocation existence and tool name, but not `raw_path` or excerpt authenticity.
|
||||
- `ToolInvocationRecorder` persisted retrieval details but did not generate claim-addressable `evidence_refs`.
|
||||
- `QueryLogsTools` could fall back to `generic-service` placeholder logs on no-hit.
|
||||
|
||||
## Implementation Evidence
|
||||
|
||||
- `ToolInvocationRecorder` now extracts:
|
||||
- `$.alerts[i]` for `query_metrics`
|
||||
- `$.logs[i]` for `query_logs`
|
||||
- `$.evidence_blocks[i]` for `lookup_knowledge`
|
||||
- `ExecutorGatekeeperService` now validates:
|
||||
- invocation existence
|
||||
- tool name
|
||||
- raw path presence
|
||||
- `retrieval_details.evidence_refs`
|
||||
- excerpt similarity/support
|
||||
- `VerifierInputHook` only auto-fills a singular `source_invocation_id` when exactly one candidate exists and never invents `raw_path`.
|
||||
- `QueryLogsTools` returns HikariCP mock logs for `order-service` and returns empty no-hit results for unrelated services.
|
||||
|
||||
## Residual Finding
|
||||
|
||||
The HikariCP negative E2E no longer has generic-service pollution, but the model still issued an extra broad HikariCP query without the service filter and used order-service as context. This is a remaining narrow-scope behavior issue, not a mock evidence pollution issue.
|
||||
@@ -0,0 +1,46 @@
|
||||
# Acceptance
|
||||
|
||||
## Static Verification
|
||||
|
||||
- `openspec validate interview-demo-quality-audit --strict`
|
||||
- Result: passed.
|
||||
- Coverage: OpenSpec proposal/design/spec/tasks consistency.
|
||||
- PowerShell parser/runtime readiness check:
|
||||
- Command: `powershell -NoProfile -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1 -BaseUrl http://127.0.0.1:1 -OutputDir target/demo-check-syntax`
|
||||
- Result: expected failure with actionable readiness message.
|
||||
- Coverage: script parses under Windows PowerShell and fails before issuing diagnosis requests when service is unreachable.
|
||||
|
||||
## Script Verification
|
||||
|
||||
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
|
||||
- Result: passed.
|
||||
- Coverage: 12/12 fixed eval fixtures, Prompt audit evaluator checks, Gatekeeper rule metadata checks, regenerated baseline reports.
|
||||
- `mvn -q "-Dtest=ChatServiceSequentialAgentTest" test`
|
||||
- Result: passed.
|
||||
- Coverage: Chat verifier evaluation persists `prompt_audit`.
|
||||
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ChatServiceSequentialAgentTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`
|
||||
- Result: passed.
|
||||
- Coverage: broader eval, baseline diff, Chat sequential flow, Gatekeeper, and Verifier input hook regression set.
|
||||
- `mvn -q -DskipTests compile`
|
||||
- Result: passed.
|
||||
- Coverage: main source compilation.
|
||||
|
||||
## E2E Verification
|
||||
|
||||
- Started service with:
|
||||
- `mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo`
|
||||
- Ran:
|
||||
- `powershell -NoProfile -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1 -BaseUrl http://localhost:9900 -SessionId mvp-demo-interview-quality-audit-001`
|
||||
- Result: passed.
|
||||
- Summary:
|
||||
- `chatSuccess=true`
|
||||
- `verdict=LOW_CONFID`
|
||||
- `gatekeeperStatus=fail`
|
||||
- `gatekeeperRuleSetVersion=gatekeeper-rules-v1`
|
||||
- `promptAuditVersion=chat-prompts-v1`
|
||||
- tools included `lookup_knowledge`, `query_logs`, `query_metrics`, and `get_available_log_topics`
|
||||
- Note: live E2E remains a compatibility check, not the deterministic PASS oracle. The fixed fixture baseline is the regression source of truth.
|
||||
|
||||
## Not Verified
|
||||
|
||||
- Browser UI inspection was not required for this change because the scope is backend trace/eval/demo script documentation, not frontend behavior.
|
||||
@@ -0,0 +1,23 @@
|
||||
# Interview Demo Quality Audit Brief
|
||||
|
||||
## Background
|
||||
|
||||
The MVP already demonstrates traceable Agent diagnosis with Planner, Executor, Gatekeeper, Verifier, Composer, evidence tools, trace persistence, and deterministic eval fixtures. The remaining interview-readiness gap is not a new Agent architecture; it is making the demo easier to run and making prompt/rule changes easier to audit.
|
||||
|
||||
## Goal
|
||||
|
||||
Stabilize the interview demo path, expand fixture-backed evaluation, and persist prompt/Gatekeeper audit metadata so the project can explain and verify Agent behavior during interviews.
|
||||
|
||||
## Scope
|
||||
|
||||
- Add prompt audit metadata to Chat verifier evaluation.
|
||||
- Extend deterministic eval cases and baseline reports.
|
||||
- Add an interview demo preflight/check script.
|
||||
- Update MVP demo and architecture documentation.
|
||||
|
||||
## Non-goals
|
||||
|
||||
- No public API or database schema changes.
|
||||
- No new SubAgent split, MCP migration, process isolation, or AIOps LLM Verifier.
|
||||
- No guarantee that every live LLM run returns PASS.
|
||||
|
||||
@@ -0,0 +1,111 @@
|
||||
# interview-demo-quality-audit Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: stabilize the interview demo, expand deterministic eval coverage, and add Prompt/Gatekeeper version audit.
|
||||
- Slug: `interview-demo-quality-audit`.
|
||||
- Devflow scale: `standard`.
|
||||
- Interface impact: L2 internal contract change because `verifier_evaluation` gains `prompt_audit`; no public HTTP API or database schema change.
|
||||
|
||||
## Context
|
||||
|
||||
- `devflow/index.md` used: related entries found for `diagnosis-eval-demo-gatekeeper-closure`, `executor-composer-final-answer`, `verifier-evidence-reference-fidelity`, `mvp-demo-interview-runbook`, and `diagnosis-eval-baseline-diff`.
|
||||
- Relevant glossary:
|
||||
- Evidence Tools produce incident facts and must be recorded in `tool_invocation`.
|
||||
- Verifier should not use skills/runbooks as incident evidence.
|
||||
- Diagnosis Playbook Skill is workflow guidance, not a fact source.
|
||||
- Historical constraints that must enter OpenSpec:
|
||||
- Diagnosis eval is deterministic and fixture-backed; no LLM-as-judge.
|
||||
- Stable demo scenarios are documentation/payloads plus deterministic fixtures; live E2E is a compatibility check, not a guaranteed PASS oracle.
|
||||
- Gatekeeper rule metadata is already metadata-only and should not become dynamic rule execution.
|
||||
- Composer is the final expression layer and must not leak raw Executor JSON.
|
||||
|
||||
## Question Pool
|
||||
|
||||
| ID | Dimension | Mode | Question | Status |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | Terminology | evidence-driven | What should the new audit metadata be called? | Resolved |
|
||||
| Q2 | Boundary | evidence-driven | Does this require public API or schema changes? | Resolved |
|
||||
| Q3 | Acceptance | evidence-driven | Which current assets define deterministic acceptance? | Resolved |
|
||||
| Q4 | Technical | evidence-driven | Where should prompt version metadata live with minimal implementation risk? | Resolved |
|
||||
| Q5 | Scope | user-interview | Should live E2E be mandatory for all scenarios? | Confirmed by objective as conditional |
|
||||
|
||||
## Evidence-driven Conclusions
|
||||
|
||||
- Q1 conclusion: use `prompt_audit` for prompt version metadata and keep existing `gatekeeper_result.rule_set_version`.
|
||||
- Q2 conclusion: keep this as an internal trace/self-evaluation contract change. Do not add endpoints, tables, or new Agent roles.
|
||||
- Q3 conclusion: `DiagnosisTraceEvaluatorTest`, baseline reports, fixed fixtures, and demo scripts define current acceptance style.
|
||||
- Q4 conclusion: add a small Chat prompt audit catalog near `ChatService` prompt loading and persist a compact snapshot with verifier evaluation.
|
||||
- Q5 conclusion: run live E2E with `mvp-demo` profile if dependencies are available; otherwise record the blocker and rely on deterministic eval/unit evidence.
|
||||
|
||||
## Specify / Alignment
|
||||
|
||||
| Check | Status | Notes |
|
||||
|---|---|---|
|
||||
| proposal goals -> proposal | Aligned | Proposal covers demo preflight, eval expansion, prompt audit, and docs. |
|
||||
| proposal scope/constraints -> design | Aligned | Design records no public API/schema changes, prompt audit shape, eval fields, and demo script behavior. |
|
||||
| design decisions -> specs/tasks | Aligned | Specs cover persisted prompt audit, evaluator checks, baseline, and demo script outputs. |
|
||||
| specs observable behavior -> tasks | Aligned | Each requirement has implementation and verification tasks. |
|
||||
|
||||
## Audit
|
||||
|
||||
Input -> processing -> output chain:
|
||||
|
||||
```text
|
||||
prompt resource metadata
|
||||
-> ChatService / PromptAudit snapshot
|
||||
-> verifier_evaluation.prompt_audit
|
||||
-> Trace API / eval fixtures
|
||||
-> DiagnosisTraceEvaluator baseline
|
||||
|
||||
run-interview-demo-check.ps1
|
||||
-> service readiness
|
||||
-> chat / trace / feedback
|
||||
-> mvp/demo/output summary
|
||||
```
|
||||
|
||||
Architecture risk assessment:
|
||||
|
||||
1. The audit shape is intentionally compact and internal; storing full prompt text would create noisy traces and possible sensitive-content risk.
|
||||
2. Eval should assert versions by explicit metadata, not by prompt content hashes that churn during local prompt edits.
|
||||
3. Live demo checks may still be LOW_CONFID because LLM output is not deterministic; deterministic fixtures remain the regression source of truth.
|
||||
4. No devflow/OpenSpec conflict found.
|
||||
|
||||
## Commit Gate
|
||||
|
||||
- `openspec validate interview-demo-quality-audit --strict`: passed.
|
||||
- File completeness:
|
||||
- `proposal.md`: present.
|
||||
- `design.md`: present.
|
||||
- `specs/`: present for `chat-verifier-agent`, `diagnosis-eval-harness`, and `mvp-demo-trace-acceptance`.
|
||||
- `tasks.md`: present.
|
||||
- Consistency:
|
||||
- Proposal goals map to design sections.
|
||||
- Design decisions map to spec requirements and executable tasks.
|
||||
- Task acceptance checks are verifiable.
|
||||
- `.committed` marker created.
|
||||
|
||||
## Current Checkpoint
|
||||
|
||||
- Commit completed.
|
||||
- Apply is authorized by the original objective: "完成后归档提交".
|
||||
|
||||
## Pre-apply Research
|
||||
|
||||
- Capability source: sm-flow built-in apply protocol. `openspec-apply-change` was not invoked directly in this session.
|
||||
- Repository semantic search/LSP note: the requested `codebase-retrieval` and LSP tools were not available in the exposed toolset, so impact analysis used `rg`, direct file reads, OpenSpec/devflow artifacts, and targeted tests.
|
||||
- Reference implementation and reuse:
|
||||
- `ChatService.persistVerifierEvaluation(...)` is the single persistence point for Chat verifier/composer audit data; prompt audit was added there to cover normal, fallback, and degraded Composer paths.
|
||||
- `DiagnosisTraceEvaluator` and `DiagnosisEvalReportWriter` are the deterministic eval extension points; no LLM judge was introduced.
|
||||
- `mvp/demo/scripts/run-payment-timeout-demo.ps1` provided the request/trace/feedback flow reused by the new interview preflight script.
|
||||
- Interface impact remains L2 internal trace contract: `verifier_evaluation.prompt_audit` and eval report fields are added; no public endpoint, table, or request DTO changed.
|
||||
|
||||
## Apply Notes
|
||||
|
||||
- Added compact Chat prompt audit metadata: `chat-prompts-v1`, with planner/executor/verifier/composer prompt versions and resource paths.
|
||||
- Extended diagnosis eval schema, result reporting, baseline fixtures, JSON report, and Markdown report for Prompt audit and Gatekeeper rule metadata.
|
||||
- Added two fixture-backed audit cases:
|
||||
- `prompt-gatekeeper-audit-closure`
|
||||
- `audit-metadata-low-confid`
|
||||
- Added `mvp/demo/scripts/run-interview-demo-check.ps1` to run service readiness, Chat, Trace, feedback, and summary output.
|
||||
- Updated MVP demo/eval/architecture docs to explain `prompt_audit.version`, `gatekeeper_result.rule_set_version`, and deterministic fixture baseline.
|
||||
@@ -0,0 +1,58 @@
|
||||
# Evidence
|
||||
|
||||
## Context Files Read
|
||||
|
||||
- `devflow/index.md`
|
||||
- `devflow/glossary/CONTEXT.md`
|
||||
- `devflow/projects/2026-07-08-diagnosis-eval-demo-gatekeeper-closure/decisions.md`
|
||||
- `devflow/projects/2026-07-08-executor-composer-final-answer/decisions.md`
|
||||
- `mvp/architecture/current-mvp-architecture.md`
|
||||
- `mvp/architecture/agent-orchestration.md`
|
||||
- `mvp/architecture/executor-evidence-pipeline-refactor.md`
|
||||
- `mvp/architecture/harness-quality-gates.md`
|
||||
- `mvp/demo/README.md`
|
||||
- `mvp/demo/ten-minute-interview-demo.md`
|
||||
- `mvp/eval/README.md`
|
||||
- `mvp/eval/cases/diagnosis-cases.json`
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ExecutorGatekeeperService.java`
|
||||
- `src/main/java/com/superbiz/agent/eval/DiagnosisTraceEvaluator.java`
|
||||
- `src/main/resources/gatekeeper/gatekeeper-rules.json`
|
||||
|
||||
## Tooling Note
|
||||
|
||||
The required `codebase-retrieval` and LSP tools were not exposed in this session. Impact analysis used `rg`, direct file reads, existing OpenSpec/devflow artifacts, and targeted tests instead.
|
||||
|
||||
## Implementation Evidence
|
||||
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- Adds `prompt_audit` under `verifier_evaluation` through the shared `persistVerifierEvaluation(...)` path.
|
||||
- Uses compact metadata only: audit version, prompt names, prompt versions, and resource paths.
|
||||
- `src/main/java/com/superbiz/agent/eval/DiagnosisTraceEvaluator.java`
|
||||
- Adds deterministic checks for `requirePromptAudit`, `expectedPromptAuditVersion`, `expectedPromptVersions`, and `requireGatekeeperRules`.
|
||||
- `src/main/java/com/superbiz/agent/eval/DiagnosisEvalReportWriter.java`
|
||||
- Adds Prompt Audit and Gatekeeper rule count columns to Markdown reports.
|
||||
- `mvp/eval/cases/diagnosis-cases.json`
|
||||
- Expands fixed baseline to 12 fixture-backed cases.
|
||||
- `mvp/eval/fixtures/prompt-gatekeeper-audit-closure-pass.json`
|
||||
- Positive PASS fixture proving Prompt audit and Gatekeeper rule metadata closure.
|
||||
- `mvp/eval/fixtures/audit-metadata-low-confid.json`
|
||||
- LOW_CONFID fixture proving safe answer behavior while audit metadata remains present.
|
||||
- `mvp/demo/scripts/run-interview-demo-check.ps1`
|
||||
- Adds service readiness, Chat, Trace, feedback, and summary output for interview preflight.
|
||||
|
||||
## Verification Evidence
|
||||
|
||||
- OpenSpec:
|
||||
- `openspec validate interview-demo-quality-audit --strict`: passed before archive.
|
||||
- `openspec validate --specs --strict`: 10 specs passed after merging deltas into main specs.
|
||||
- Unit/eval:
|
||||
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`: passed.
|
||||
- `mvn -q "-Dtest=ChatServiceSequentialAgentTest" test`: passed.
|
||||
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ChatServiceSequentialAgentTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`: passed.
|
||||
- Compile:
|
||||
- `mvn -q -DskipTests compile`: passed.
|
||||
- E2E:
|
||||
- Started `mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo`.
|
||||
- Ran `mvp/demo/scripts/run-interview-demo-check.ps1` against `http://localhost:9900`.
|
||||
- Summary recorded `chatSuccess=true`, `verdict=LOW_CONFID`, `gatekeeperRuleSetVersion=gatekeeper-rules-v1`, and `promptAuditVersion=chat-prompts-v1`.
|
||||
@@ -0,0 +1,103 @@
|
||||
# Acceptance
|
||||
|
||||
## 实现结果
|
||||
|
||||
- OpenSpec tasks: `42/42` complete。
|
||||
- Phase commits:
|
||||
- `52bf030 feat(trace): add session run isolation schema`
|
||||
- `6fdbd34 docs(openspec): tighten run isolation contract`
|
||||
- `26d5529 feat(trace): isolate chat runs`
|
||||
- `027aed1 feat(trace): add run-scoped trace reads`
|
||||
- `d928a19 feat(trace): bind feedback to runs`
|
||||
- `78c1477 feat(trace): isolate aiops runs`
|
||||
- `f9df943 feat(trace): finish run-aware demo verification`
|
||||
- OpenSpec archive: `openspec/changes/archive/2026-07-10-session-run-trace-isolation`
|
||||
|
||||
## 静态验证
|
||||
|
||||
```powershell
|
||||
node --check src\main\resources\static\app.js
|
||||
node --check src\main\resources\static\trace.js
|
||||
openspec validate session-run-trace-isolation --strict
|
||||
git diff --check -- . ':!devflow/index.md'
|
||||
```
|
||||
|
||||
结果:通过。
|
||||
|
||||
## 脚本验证
|
||||
|
||||
PowerShell demo 脚本解析:
|
||||
|
||||
```powershell
|
||||
$scripts = @(
|
||||
'mvp\demo\scripts\run-payment-timeout-demo.ps1',
|
||||
'mvp\demo\scripts\run-interview-demo-check.ps1'
|
||||
)
|
||||
foreach ($script in $scripts) {
|
||||
[scriptblock]::Create((Get-Content -Raw -Encoding UTF8 $script)) | Out-Null
|
||||
}
|
||||
```
|
||||
|
||||
结果:通过。
|
||||
|
||||
Focused tests:
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=ChatControllerTest,DiagnosisTraceServiceTest,FeedbackControllerTest,FeedbackServiceTest,AiOpsServiceTest" test
|
||||
```
|
||||
|
||||
结果:通过。
|
||||
|
||||
Baseline / regression:
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest" test
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
|
||||
```
|
||||
|
||||
结果:通过,无 baseline drift。
|
||||
|
||||
## E2E 验证
|
||||
|
||||
使用 Maven 启动:
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
|
||||
```
|
||||
|
||||
E2E 使用同一 `sessionId` 连续两轮 Chat:
|
||||
|
||||
- `sessionId`: `e2e-phase6-chat-codex-20260710-2120`
|
||||
- `run1`: `run-e2a97696-4398-4abc-90e4-28f45c838f92`
|
||||
- `run2`: `run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172`
|
||||
|
||||
验证结果:
|
||||
|
||||
- Chat1 / Chat2 均成功。
|
||||
- run1 exact trace 返回 run1。
|
||||
- run2 exact trace 返回 run2。
|
||||
- session-only latest trace 返回 run2。
|
||||
- DB 中同一 session 有两条 `diagnosis_run`。
|
||||
- step/tool rows 按 `run_id` 隔离,mixed row check 为 0。
|
||||
- `chat_session.message_pair_count = 2`。
|
||||
|
||||
## 日志验证
|
||||
|
||||
检查:
|
||||
|
||||
- `target/e2e/phase6-mvn-20260710-211831.out.log`
|
||||
- `logs/application.log`
|
||||
- `logs/chat.log`
|
||||
|
||||
结果:能找到 E2E `sessionId`、两个 `runId`、Chat execution、run persistence 和 trace lookup 相关日志。
|
||||
|
||||
## 浏览器/人工验证
|
||||
|
||||
未单独进行浏览器点击验证。Trace UI 的本次验收通过静态语法检查、URL/runId 参数代码审查和后端 exact trace E2E 共同覆盖。建议后续手动打开 `trace.html?sessionId=...&runId=...` 做展示层冒烟。
|
||||
|
||||
## 剩余风险 / 后续事项
|
||||
|
||||
- 缺少 `runId` 的 Feedback fallback 是短期兼容路径,客户端全部迁移后可收紧。
|
||||
- `diagnosis_session` 仍保留为历史兼容和回滚表,后续需要观察窗口后再评估约束收紧或归档策略。
|
||||
- `case_library.diagnosis_id` 仍是过渡字段,旧值可能为 `session_id`,新自动值为 `run_id`。
|
||||
- 历史 mixed trace 不能恢复真实多轮边界,只能按 compatibility run 查询。
|
||||
@@ -0,0 +1,41 @@
|
||||
# Session / Run / Trace Isolation
|
||||
|
||||
## 背景
|
||||
|
||||
同一个 `sessionId` 以前同时代表多轮 Chat 上下文和一次持久化诊断 Trace。端到端验证发现,同一 `sessionId` 连续两轮 Chat 时,Redis 多轮上下文是正确的,但 MySQL 中 `diagnosis_session` 会被后一轮覆盖,`agent_step` 和 `tool_invocation` 会按同一个 `session_id` 混在一起。
|
||||
|
||||
这会导致 Trace 回放、Verifier/Evaluation 读数、Feedback 绑定和 `case_library` 来源都可能跨轮污染。
|
||||
|
||||
## 目标
|
||||
|
||||
- 将会话态和运行态拆开:`chat_session` 保存会话元数据,`diagnosis_run` 保存一次诊断运行。
|
||||
- 引入正式 API 字段 `runId`,作为一次可回放诊断执行的边界。
|
||||
- `agent_step` 和 `tool_invocation` 保留原 Trace 明细角色,新增 `run_id` 并按 run 隔离读写。
|
||||
- Trace、Feedback、CaseLibrary、AIOps、demo 脚本和 Trace UI 都支持 run-aware 流程。
|
||||
- 保留旧 `diagnosis_session` 作为历史兼容和回滚表。
|
||||
- 完成 Maven E2E、DB 检查、日志检查和 baseline drift 验证。
|
||||
|
||||
## 范围
|
||||
|
||||
- Flyway/JPA 增加 `chat_session`、`diagnosis_run`,并给 `agent_step`、`tool_invocation` 增加 `run_id`。
|
||||
- Chat 每次有效执行创建一个新的 `diagnosis_run`,响应返回 `sessionId + runId`。
|
||||
- Trace API 支持 latest-run fallback 和 exact-run 查询:`GET /api/diagnosis/{sessionId}/trace?runId=...`。
|
||||
- 新增 run list API:`GET /api/chat/session/{sessionId}/runs`。
|
||||
- Feedback 优先绑定 `runId`,缺省时短期 fallback 到 latest run 并返回 `fallbackToLatestRun=true`。
|
||||
- AIOps 每次有效执行创建并透出 `runId`,SSE 保持 `message` event name 并发送 `type=metadata`。
|
||||
- MVP demo、Trace UI、表文档和架构文档统一为 `chat_session -> diagnosis_run -> trace detail(run_id)`。
|
||||
|
||||
## 非目标
|
||||
|
||||
- 不新增 `diagnosis_trace` 或 `trace_event` 主表。
|
||||
- 不实现完整 run-list UI。
|
||||
- 不删除旧 `diagnosis_session`。
|
||||
- 不改变 Redis 对话历史窗口策略。
|
||||
- 不把完整多轮正文历史持久化到 MySQL。
|
||||
- 不尝试把历史混合 trace 还原成真实多轮边界。
|
||||
|
||||
## 关联
|
||||
|
||||
- OpenSpec: `openspec/changes/archive/2026-07-10-session-run-trace-isolation`
|
||||
- Change slug: `session-run-trace-isolation`
|
||||
- 分档: complex
|
||||
@@ -0,0 +1,48 @@
|
||||
# Decisions
|
||||
|
||||
## 核心决策
|
||||
|
||||
| 决策 | 选择 | 理由 |
|
||||
|---|---|---|
|
||||
| 领域拆分 | 新增 `chat_session` 和 `diagnosis_run` | 会话元数据和一次诊断执行的生命周期不同,继续塞在一张表会导致上下文膨胀和边界混淆 |
|
||||
| Trace 明细 | 复用 `agent_step` / `tool_invocation`,增加 `run_id` | 现有明细表已经能表达 Trace,隔离需要 run key,不需要新事件模型 |
|
||||
| API 身份 | `runId = "run-" + UUID` | 外部 ID 不依赖数据库自增 ID,碰撞风险低 |
|
||||
| Trace 兼容 | 缺少 `runId` 时按 `created_at DESC, id DESC` 解析 latest run | 保留旧客户端兼容性,避免 feedback/eval 更新 `updated_at` 后改变 latest 判定 |
|
||||
| 历史迁移 | 每条旧 `diagnosis_session` 生成一条 compatibility run | 旧混合数据没有真实轮次边界,不能伪造多 run 历史 |
|
||||
| Feedback fallback | 缺少 `runId` 时短期绑定 latest run 并返回 `fallbackToLatestRun=true` | 老客户端可继续工作,同时让歧义可观测 |
|
||||
| Case provenance | 新自动案例写 `case_library.diagnosis_id = run_id` | 保留旧列,文档声明过渡语义 |
|
||||
| AIOps 范围 | 同一个 change 内完成 AIOps run isolation | AIOps 是一等 Trace 入口,不能留下同类混合 trace bug |
|
||||
| 所有权校验 | 服务层校验 run/session ownership,暂不加 DB 外键 | 兼容历史 orphan rows 和回滚窗口 |
|
||||
|
||||
## 用户确认
|
||||
|
||||
- 选择拆 `chat_session` 和 `diagnosis_run`,不只是在旧表加字段。
|
||||
- `chat_session` 第一阶段只保存元数据,不保存完整对话正文。
|
||||
- 完整多轮对话历史继续放在 Redis `SessionContext.messageHistory`。
|
||||
- `runId` 是正式 API 字段。
|
||||
- Trace 缺少 `runId` 时短期默认查 latest run。
|
||||
- Feedback 缺少 `runId` 时短期 fallback,长期可再收紧。
|
||||
- 每次有效 Chat/AIOps 都创建 run。
|
||||
- 不新增 `diagnosis_trace` / `trace_event` 主表。
|
||||
- 旧 `diagnosis_session` 保留用于历史和回滚,新代码不再写新执行态。
|
||||
- demo 脚本和 Trace UI 做最小 `runId` 支持。
|
||||
|
||||
## 接口影响
|
||||
|
||||
级别:L4。
|
||||
|
||||
- 新 API 响应字段:`runId`。
|
||||
- Trace API 新 query 参数:`runId`。
|
||||
- 新 API:`GET /api/chat/session/{sessionId}/runs`。
|
||||
- Feedback request 新增 optional/preferred `runId`。
|
||||
- Feedback response 新增 bound `runId` 和 `fallbackToLatestRun`。
|
||||
- `/api/ai_ops` SSE 保持 event name `message`,新增 `type=metadata` 消息。
|
||||
- DB contract 新增两张表和两个 `run_id` 列。
|
||||
- 旧 `sessionId` only 调用仍兼容,但 fallback 必须可观测。
|
||||
|
||||
## 风险接受
|
||||
|
||||
- 历史混合 trace 无法真实拆分,只能作为 compatibility run。
|
||||
- 上下文传播同时依赖 `RunnableConfig.metadata` 和 `SessionContextHolder`,后续改动必须注意 `sessionId/runId` 同步。
|
||||
- `case_library.diagnosis_id` 在过渡期存在 `session_id` 和 `run_id` 两种语义。
|
||||
- 缺少 `runId` 的 Feedback 仍有歧义,后续客户端迁移完成后可收紧为参数错误。
|
||||
@@ -0,0 +1,76 @@
|
||||
# Evidence
|
||||
|
||||
## 上下文证据
|
||||
|
||||
- `SessionContext.messageHistory` 和 `getMessagePairCount()` 证明 Redis 承载热对话历史;MySQL 只需要长期审计的会话目录和运行记录。
|
||||
- `CaseLibraryService.createFromSession` 原先按 `DiagnosisSession.sessionId` 去重并映射 query/answer,因此 run 隔离后需要新增 `createFromRun`。
|
||||
- 旧 `mvp/architecture/data-model.md` 把 `case_library.diagnosis_id` 解释为 `diagnosis_session.session_id`,本次改为过渡语义:旧数据可能是 `session_id`,新自动案例是 `run_id`。
|
||||
- 既有 Trace OpenSpec 要求 `GET /api/diagnosis/{sessionId}/trace` 是只读端点;latest-run 和 exact-run 查询都必须保持只读。
|
||||
- ISS-010 的 E2E 事实显示同一 `sessionId` 两轮 Chat 会产生 MySQL Trace 混合,是本 change 的直接触发证据。
|
||||
|
||||
## 实现证据
|
||||
|
||||
- Phase 1 增加 `V011__add_session_run_isolation.sql`,创建 `chat_session`、`diagnosis_run`,并为 `agent_step` / `tool_invocation` 增加 nullable `run_id`。
|
||||
- Phase 2 将 Chat 写路径切到 `chat_session + diagnosis_run`,并让 Hook/Tool/Evaluation/Gatekeeper 使用 run-scoped 数据。
|
||||
- Phase 3 将 Trace API 改为 latest-run / exact-run 双模式,并加入 lightweight run summaries。
|
||||
- Phase 4 将 Feedback 和 CaseLibrary 绑定到 run,保留没有 run-backed 数据时的 legacy fallback。
|
||||
- Phase 5 将 AIOps 接入 run isolation,SSE metadata 暴露 `sessionId + runId`。
|
||||
- Phase 6 更新 demo 脚本、Trace UI、MVP 架构文档和表文档,并修正 review 后发现的 session-only 文档残留。
|
||||
|
||||
## E2E 证据
|
||||
|
||||
Maven 启动命令:
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
|
||||
```
|
||||
|
||||
日志:
|
||||
|
||||
- `target/e2e/phase6-mvn-20260710-211831.out.log`
|
||||
- `target/e2e/phase6-mvn-20260710-211831.err.log`
|
||||
- `logs/application.log`
|
||||
- `logs/chat.log`
|
||||
|
||||
E2E session:
|
||||
|
||||
- `sessionId`: `e2e-phase6-chat-codex-20260710-2120`
|
||||
- `run1`: `run-e2a97696-4398-4abc-90e4-28f45c838f92`
|
||||
- `run2`: `run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172`
|
||||
|
||||
结果:
|
||||
|
||||
- 两轮 Chat 都成功,并复用同一个 `sessionId`。
|
||||
- 两轮返回不同 `runId`。
|
||||
- run1 exact trace 只返回 run1。
|
||||
- run2 exact trace 只返回 run2。
|
||||
- session-only Trace latest fallback 返回 run2。
|
||||
- `chat_session.message_pair_count = 2`,证明多轮上下文连续。
|
||||
|
||||
## DB 证据
|
||||
|
||||
通过 `scripts/query_mysql.py` 检查:
|
||||
|
||||
- `diagnosis_run` 中该 E2E session 有 2 条 `SUCCESS / CHAT` 运行。
|
||||
- `agent_step` 按 run 分组:run1 `10` 行,run2 `9` 行。
|
||||
- `tool_invocation` 按 run 分组:run1 `14` 行,run2 `8` 行。
|
||||
- mixed row check 为 `0`,没有 NULL 或 unexpected `run_id` 混入该 E2E session。
|
||||
|
||||
## Baseline 证据
|
||||
|
||||
运行:
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest" test
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
|
||||
```
|
||||
|
||||
结果:
|
||||
|
||||
- 两组 baseline / regression 命令通过。
|
||||
- baseline harness 使用离线 fixture,不依赖 live DB/session tables。
|
||||
- 未观察到 baseline drift。
|
||||
|
||||
## 工具限制
|
||||
|
||||
AGENTS 要求的 `codebase-retrieval` 和 LSP 工具在本会话不可用。替代验证使用 OpenSpec、`rg`、定向阅读、 focused tests、E2E、DB 查询和日志检查。
|
||||
+52
@@ -0,0 +1,52 @@
|
||||
# Chat Diagnosis StateGraph ChatService Cutover 验收
|
||||
|
||||
## 结果
|
||||
|
||||
已接受。OpenSpec tasks 26/26 完成,阶段 3 可归档。
|
||||
|
||||
## 验证
|
||||
|
||||
### 静态验证
|
||||
|
||||
- `git diff --check`:通过。
|
||||
- source/reference isolation:cutover 核心文件中的 `SequentialAgent`、`VerifierContextHolder`、`VerifierInputHook`、`tool_trace_summary` 命中 0;核心 TODO/FIXME/placeholder 命中 0。
|
||||
- schema whitelist:V012 为 1 个 ALTER TABLE、1 个 ADD COLUMN、0 CREATE、0 DROP,目标仅 `diagnosis_run.orchestration_trace JSON NULL`。
|
||||
- `openspec validate chat-diagnosis-stategraph-chatservice-cutover --strict`:通过。
|
||||
- `openspec validate --specs --strict`:14 passed,0 failed。
|
||||
|
||||
### 脚本验证
|
||||
|
||||
- focused Maven regression:Graph runtime/result mapper/real nodes、ChatService cutover、Trace、Controller、Repository、Gatekeeper、Composer、Eval,共 27 suites / 119 tests;0 failures、0 errors、0 skipped。
|
||||
- `mvn -q -DskipTests test-compile`:通过。
|
||||
- `ChatServiceGraphIntegrationTest` 覆盖 SUCCESS、handled Fallback、unhandled failure、blank answer/partial trace 和同 session 多 Run 隔离。
|
||||
|
||||
### 浏览器/人工验证
|
||||
|
||||
- 未运行。阶段 3 的公开协议由 Controller/service tests 覆盖,完整人工/live 验收按用户规则保留到阶段 5。
|
||||
|
||||
### 未验证
|
||||
|
||||
- 未使用 Maven 启动应用做 live E2E。
|
||||
- 未检查 `logs/` 运行日志。
|
||||
- 未执行 `scripts/query_mysql.py` 查询真实数据库。
|
||||
- 原因:用户明确要求只有阶段 5 全部完成后统一执行端到端、日志和数据库验收;阶段 3 只做风险相关自动化验证。
|
||||
- 剩余风险:真实模型/工具调用下的 Prompt 行为、Flyway 在真实 MySQL 的应用结果和最终 Trace 数据须由阶段 5 E2E 证明。
|
||||
|
||||
## 已完成范围
|
||||
|
||||
- 复杂 Chat 唯一生产编排切换到 Diagnosis StateGraph,公开 ChatResult 保持兼容。
|
||||
- Run 生命周期、metrics、Eval、finally cleanup 与 runId 隔离保持;handled Fallback=SUCCESS,未处理/空答案=FAILED。
|
||||
- 新增 Run 级 compact orchestration trace,并只在 Trace run 对象暴露解析结果。
|
||||
- Verifier Prompt/self-evaluation 使用 verified-only 数据,不伪造未发生 verdict/event。
|
||||
- 移除冲突的 Sequential 实现测试并以 Graph/public contract tests 替代。
|
||||
|
||||
## Bug 修复和诊断
|
||||
|
||||
- 替代测试错误引用 `com.superbiz.agent.tool.ToolInvocationRecorder`:通过定义/引用搜索确认类型位于同一 `service` 包,删除错误 import。
|
||||
- no-answer 新测试错误期待 inner cause:分类为测试断言偏差,改为公开 wrapper error,同时保留 FAILED 与真实 partial trace 核心断言。
|
||||
|
||||
## 交接
|
||||
|
||||
- OpenSpec archive:已同步 5 份 delta specs,并归档到 `openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-chatservice-cutover/`。
|
||||
- 下一步:审查精确 Git diff 并完成阶段 3 独立提交,之后才启动阶段 4。
|
||||
- OpenSpec 归档确认:用户已明确授权后续阶段直接实现/归档;归档已完成。
|
||||
@@ -0,0 +1,21 @@
|
||||
# Chat Diagnosis StateGraph ChatService Cutover Brief
|
||||
|
||||
## 背景
|
||||
|
||||
- 用户目标:将 ISS-011 阶段 3 作为独立 sm-flow,正式切换复杂 Chat 生产编排并增加 Run 级 orchestration trace。
|
||||
- 当前问题:真实 Diagnosis Graph Nodes 已存在,但生产入口仍依赖 SequentialAgent、外层重试和 ThreadLocal/Hook 隐式状态。
|
||||
- 关联 OpenSpec:`openspec/changes/chat-diagnosis-stategraph-chatservice-cutover/`
|
||||
- devflow 分档:complex。
|
||||
|
||||
## 范围
|
||||
|
||||
- 本次要做:复杂 Chat 单轨 Graph cutover;复用四类 Agent builder;Graph result/self-evaluation 映射;V012 Run JSON 字段;Trace run-only 投影;verified-only Verifier Prompt;必要回归测试。
|
||||
- 本次不做:不改 `/api/chat` 请求/响应,不做历史 trace 回填,不删除仍供历史代码/测试使用的 Hook/ThreadLocal 类型,不完成阶段 4 测试体系全面收敛,不运行 live E2E/log/DB 验收。
|
||||
- 影响区域:ChatService、Diagnosis Graph runtime/result mapper、DiagnosisRun/Flyway、Trace DTO/service、Verifier Prompt、Graph/Service/Trace tests。
|
||||
|
||||
## OpenSpec 对齐
|
||||
|
||||
- proposal 覆盖状态:已覆盖。
|
||||
- design 覆盖状态:已覆盖。
|
||||
- specs 覆盖状态:已覆盖,5 个 delta capabilities。
|
||||
- tasks 覆盖状态:26/26 已完成。
|
||||
+207
@@ -0,0 +1,207 @@
|
||||
# Chat Diagnosis StateGraph ChatService Cutover Decisions
|
||||
|
||||
## Entry Summary
|
||||
|
||||
- 问题:复杂 Chat 仍使用 SequentialAgent + ThreadLocal/Hook 隐式状态机,真实 Graph 尚未成为生产入口,也没有 Run 级 orchestration trace 持久化和 API 投影。
|
||||
- 期望:阶段 3 独立完成生产 cutover、Run/Trace 映射和必要测试,归档并提交后才进入阶段 4。
|
||||
- 分档:complex。
|
||||
- Change:`chat-diagnosis-stategraph-chatservice-cutover`。
|
||||
- 授权:用户已要求后续阶段直接实现,不再逐 checkpoint 等待;阶段门禁、独立 archive/commit 与阶段 5 才 E2E 约束不变。
|
||||
|
||||
## Context Sources
|
||||
|
||||
- `mvp/issues/active/ISS-011-chat-diagnosis-stategraph-orchestration.md` 阶段 3、协议影响、Run 状态和验收章节。
|
||||
- 阶段 0–2 OpenSpec archives、devflow acceptance/decisions 与当前四份 Graph 主 specs。
|
||||
- `devflow/glossary/CONTEXT.md` 中 Chat Session、Diagnosis Run、Diagnosis Trace、Diagnosis Orchestration Trace 的边界。
|
||||
- `ChatService` 生产调用链、四个 Agent builder、Run persistence/self-evaluation/metrics 逻辑。
|
||||
- `DiagnosisRun`、V011 migration、`DiagnosisTraceResponse`、`DiagnosisTraceService` 与相关 tests。
|
||||
- `DiagnosisGraphFactory`、真实 action factory、Node adapters、final trace builder 和本地 CompiledGraph API。
|
||||
|
||||
## Question Pool
|
||||
|
||||
| # | 维度 | 问题 | 模式 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | 术语 | orchestration trace 是否属于 self-evaluation 或完整 Trace event log? | evidence-driven | 已解决 |
|
||||
| Q2 | 边界 | 阶段 3 是否改变 `/api/chat`,以及 Trace 字段出现在哪一层? | user-interview(既有冻结决策) | 已确认 |
|
||||
| Q3 | 生命周期 | LOW_CONFID/REJECT/Fallback 是否应使用 FAILED 或新增 DEGRADED? | user-interview(既有冻结决策) | 已确认 |
|
||||
| Q4 | 错误处理 | Graph 异常时如何 best-effort 保存已有编排事实而不伪造 event? | evidence-driven | 已解决 |
|
||||
| Q5 | 技术实现 | 是否复用现有 Agent factory/Hook/ToolCallback 与真实 Node assembly? | evidence-driven | 已解决 |
|
||||
| Q6 | 安全 | Graph Verifier Prompt/self-evaluation 是否可继续使用 raw Executor 和完整 tool trace? | evidence-driven | 已解决 |
|
||||
| Q7 | 验收 | 阶段 3 是否需要新增测试,是否现在运行 live E2E? | user-interview(用户最新规则) | 已确认 |
|
||||
|
||||
## Evidence-driven Findings
|
||||
|
||||
- Q1:glossary 与阶段 0 spec 已定义 orchestration trace 为 Run 级紧凑编排摘要,与 self-evaluation、AgentStep/ToolInvocation 和 checkpoint 分离;无需新增术语或 ADR。
|
||||
- Q4:所有已定义 Agent/Java 失败应由 Graph 路由到 handled Fallback 并返回 final state;只有真实可取得的 final/partial state 才可构造 trace。无法取得 state 的未处理异常标记 Run FAILED,不得生成虚假 transition。
|
||||
- Q5:`DiagnosisRealGraphActionsFactory` 已提供真实 Node assembly;ChatService 现有四个 builder、AgentLoggingHook、Skills hooks、method tools 和 ToolCallbacks 可直接构造 ReactAgent,再通过 `ReactAgentDiagnosisInvoker` 注入,不得复制 Prompt/Agent factory。
|
||||
- Q6:阶段 2 spec 已禁止 Graph Verifier 使用 raw Executor/full tool trace;当前 `chat-verifier-prompt.md` 仍描述旧 Hook payload,是阶段 3 必须同步的明确 gap。self-evaluation 可保存 Graph 的 verified output、Gatekeeper audit 和 Composer audit,但不能重新引入 raw 输入。
|
||||
|
||||
## User-interview Confirmations
|
||||
|
||||
| 问题 | 用户原话/既有确认 | 确认状态 | OpenSpec 回写 |
|
||||
|---|---|---|---|
|
||||
| Q2 外部协议与 Trace 层级 | ISS-011 已冻结 `/api/chat` 不变,Trace 只在 `run.orchestrationTrace` 增解析对象 | 已确认 | proposal |
|
||||
| Q3 Run 生命周期 | ISS-011 已冻结安全 Fallback 为 SUCCESS,不新增 DEGRADED;只有无法生成安全响应的未处理失败为 FAILED | 已确认 | proposal |
|
||||
| Q7 验收节奏 | “端到端只在最后阶段全部完成后才验证;每个阶段如果有必要添加单元测试验收的话,就加” | 已确认 | proposal |
|
||||
|
||||
## Interface Impact
|
||||
|
||||
- 等级:L4。
|
||||
- 原因:复杂 Chat 内部状态机正式切换;DiagnosisRun 增加数据库 JSON 契约;Trace run 对象新增字段;旧固定顺序消费者/测试不再成立。
|
||||
- 保持兼容:`/api/chat` request/response、sessionId/runId、Executor/Verifier/Composer 输出契约不变。
|
||||
- 加法变化:只新增 `diagnosis_run.orchestration_trace` 和 `run.orchestrationTrace`。
|
||||
- 回滚:revert 阶段 3 代码/Prompt/spec,数据库列可保留 nullable;不保留运行时双轨开关。
|
||||
|
||||
## Discover Status
|
||||
|
||||
- `devflow/index.md`:命中阶段 0–2 archives 和 Run/Trace 历史项目。
|
||||
- Glossary:相关术语已存在且无冲突,不需更新。
|
||||
- ADR:阶段 0 已记录难以逆转的状态机/Trace 决策,本阶段没有新的三条件 ADR。
|
||||
- 未解决问题:0。
|
||||
- Draft 产物:当前只创建 proposal + decisions;design/specs/tasks 留到 Commit checkpoint。
|
||||
|
||||
## Grill-with-docs Review
|
||||
|
||||
### Domain model stress test
|
||||
|
||||
- 同一 session 连续两个复杂 Chat run:每次编译/执行使用自己的 runId threadId 和 metadata;orchestration trace 只写各自 DiagnosisRun,session 投影不复制,符合 Session/Run/Trace 领域边界。
|
||||
- Gatekeeper REJECT 或 Planner/Executor 技术失败:Graph 进入确定性 Fallback,返回安全非空答案;Run 为 SUCCESS,trace `degraded=true`,不会把诊断质量塞进 Run status。
|
||||
- Composer 成功但 verdict=LOW_CONFID/REJECT:Run 仍为 SUCCESS;self-evaluation 保存 effective verdict,orchestration trace 保存实际路径,两者职责不混合。
|
||||
- Graph 未处理异常:使用流式 NodeOutput 捕获最后一个真实 state,只有其中已有 events 时才 best-effort 构造 partial trace;没有 event 时不伪造 transition,Run 标记 FAILED。
|
||||
- 历史/AI_OPS run:新增列 nullable,Trace DTO 对旧 run 可返回 null;不增加历史回填或跨 flow 假数据。
|
||||
|
||||
### Design tree conclusions
|
||||
|
||||
- 生产切换采用单轨,不增加 feature flag 或保留 Sequential/Graph 双运行;回滚依靠 Git revert,nullable 列可保留。
|
||||
- ChatService 负责 Run 生命周期和调用一个专用 Graph runtime/orchestrator;复杂 Agent 构造从旧 Sequential 私有流程中搬移或封装复用,不复制第二套 builder。
|
||||
- Graph 执行优先使用可观察的 `CompiledGraph.stream(..., config)` 收集最后真实 state,以支持异常时 best-effort trace;正常终止仍以 END state 的 `final_answer` 为唯一答案来源。
|
||||
- Graph result mapping 形成独立组件:final answer、Verifier evaluation、Composer audit、Gatekeeper audit、trace JSON 均从显式 final state读取,不再访问 VerifierContextHolder。
|
||||
- Verifier Prompt 必须更新为 `diagnosis_context`、`verified_executor_output`、`verified_evidence`、`gatekeeper_audit`、`verdict_ceiling` 和可选 retry context;删除 raw Executor/full tool trace/Hook Gatekeeper 说明。
|
||||
- 阶段 3 必须有 production cutover 与 Trace DTO/Service tests;阶段 4 再做测试体系全面改名、夹具收敛和旧测试删除。
|
||||
|
||||
### Documentation result
|
||||
|
||||
- 术语与 `devflow/glossary/CONTEXT.md` 完全一致,无需修改 glossary。
|
||||
- 状态机、Run status 和 orchestration trace 隔离均来自阶段 0 已归档 ADR/规格,不创建重复 ADR。
|
||||
- proposal 已反映单轨切换、L4 接口影响、流式 partial-state 处理、Prompt 安全边界和阶段 5 E2E 延期。
|
||||
- Grill question pool 全部关闭;无需要再次询问用户的产品取舍。
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
### Module and caller map
|
||||
|
||||
`ChatController` 的 normal/SSE 两个入口都调用 `ChatService.executeChatWithStrategy`,复杂分支进入 `executeChatComplex`;对外仍只消费 `ChatResult(answer, sessionId, runId)`。阶段 3 将内部链路变为 `ChatService Run lifecycle -> complex Chat Graph runtime/Agent assembly -> real Diagnosis Nodes -> Graph result mapper -> DiagnosisRun persistence -> DiagnosisTraceService`。Trace UI/Eval 当前读取兼容 `session.selfEvaluation`,它仍由 Run self-evaluation 投影;新增 orchestration trace 只属于 `run`。AIOps、Feedback、CaseLibrary 和 Run list 继续使用 DiagnosisRun 既有字段,不消费新增 orchestration trace。
|
||||
|
||||
| 模块 | 所有权 | 允许的依赖/影响 |
|
||||
|---|---|---|
|
||||
| ChatController | `/api/chat` normal/SSE 协议 | 继续只依赖 ChatResult;无字段变化 |
|
||||
| ChatService | Chat Session/Diagnosis Run 生命周期、成功/失败保存、metrics、Eval、finally cleanup | 调用一个 Graph runtime/result mapper;不再拥有条件边/重试/Gatekeeper/Composer 路由 |
|
||||
| Complex Chat Graph runtime | 每请求 Agent assembly、initial state、RunnableConfig、CompiledGraph stream | 复用现有 Prompt/tool/skill/logging;不持久化跨 Run 状态 |
|
||||
| Diagnosis Graph | Node status、verified material、events、final answer | 保持阶段 1/2 路由和 counter 所有权 |
|
||||
| Graph result mapper | final/partial state 到安全 evaluation/trace DTO | 不读 ThreadLocal,不查询其他 run,不持久化 raw material |
|
||||
| DiagnosisRun/Flyway | 当前 Run 的 orchestration JSON | 仅一个 nullable JSON 列;历史/AIOps 可为 null |
|
||||
| DiagnosisTraceService/DTO | exact/latest Run 查询和 API 投影 | 只在 RunTrace 加 parsed map;session/top-level/run-list 不重复 |
|
||||
|
||||
### Lifecycle and coupling audit
|
||||
|
||||
ReactAgent 与 CompiledGraph 都按请求构造,因为 Planner/Executor system Prompt 含本次 history/knowledge map;Graph State 和 runtime last-state holder 也是 invocation-scoped,不进入 singleton 可变字段。SessionContextHolder 仍只包围当前请求并在 finally 清理;Graph Node config 以 runId threadId + metadata 绑定 AgentStep、ToolInvocation 和 Gatekeeper。success persistence 在 Eval 之前完成,EvaluationService 再按 runId 合并 rule channel;handled Fallback 与 Composer 使用同一 SUCCESS 路径。Trace JSON 与 self-evaluation 分栏,避免把路线、质量和完整 Agent/tool 明细耦合到一个容器。
|
||||
|
||||
### Consumer impact audit
|
||||
|
||||
- ChatController normal/SSE:ChatResult 协议不变,L4 内部切换不要求调用方迁移。
|
||||
- Trace UI/demo:继续读取兼容 `session.selfEvaluation`;新的 `run.orchestrationTrace` 是加法字段,最终脚本断言留到阶段 5。
|
||||
- DiagnosisTraceEvaluator:现有 fixtures 不变;工具覆盖仍可从 `toolInvocations` 读取,兼容 `executor_structured_output` 保存 verified projection。pre-verification Fallback 质量未来应以 trace degraded 而非伪 verdict 判断,属于阶段 4 测试体系/阶段 5文档收尾。
|
||||
- AIOps/Feedback/CaseLibrary/Run list:实体新增 nullable 字段不改变 builder call sites或查询语义。
|
||||
- 数据库:Hibernate validate 要求 V012 与 entity 同批;回滚代码可忽略保留列。
|
||||
|
||||
### Cross-artifact alignment
|
||||
|
||||
| 对齐链 | 结果 | 证据 |
|
||||
|---|---|---|
|
||||
| brief/proposal 目标、范围、非目标 → proposal | 已对齐 | 单轨 cutover、Run/Trace、Prompt、阶段边界与 E2E 延期均明确 |
|
||||
| proposal 承诺与约束 → design | 已对齐 | 10 项决策覆盖 runtime、state、config、partial state、result、persistence、DB/API/Prompt/tests |
|
||||
| design 架构/接口结论 → specs/tasks | 已对齐 | L4、单轨、Run status、verified-only、Trace 唯一投影和 schema whitelist 均有 requirement/task |
|
||||
| specs 可观察行为 → tasks 可执行切片 | 已对齐 | 6 组 26 个切片覆盖数据、runtime、Prompt、cutover、tests、handoff |
|
||||
|
||||
### Audit result
|
||||
|
||||
未发现与阶段 0–2、Run/Trace ADR 或 glossary 冲突。审计确认不能把 Agent factory 复制进 Node,也不能把 pre-verification Fallback 伪装为 Verifier verdict;两项已在 design/specs/tasks 固定。唯一跨阶段依赖是阶段 4 让测试/Eval 以 orchestration degraded 理解无 Verifier verdict 的安全 Fallback,已记录但不阻塞阶段 3 production correctness。架构风险可接受,cross-artifact gap=0,无未解决接口消费者。
|
||||
|
||||
## Commit Gate
|
||||
|
||||
- schema:spec-driven;proposal/design/specs/tasks 全部 done,applyRequires=`tasks` 已满足。
|
||||
- OpenSpec:当前 change strict validation 通过;14 个主 specs 全部 strict pass。
|
||||
- 规格结构:5 个 delta capabilities、23 条 requirements、72 个 scenarios;tasks 26 个可执行 checkbox。
|
||||
- Cross-artifact:4/4 已对齐,gap=0。
|
||||
- Interface impact:L4;design 已独立记录消费者、加法 DB/Trace 协议、单轨迁移和 Git revert 回滚。
|
||||
- Question pool:所有 evidence-driven 已查证;所有 user-interview 已由 ISS-011/用户原话确认;无未决项。
|
||||
- Preflight:`git diff --check` 通过;Commit checkpoint 尚未修改 Java、SQL、Prompt 或测试。
|
||||
- 结论:Draft OpenSpec 已达到可执行状态,创建 `.committed`,阶段 3 Apply 只能以这些产物为依据。
|
||||
|
||||
## Apply Authorization
|
||||
|
||||
- 用户原话:“直接实现吧,不用找我授权了”。
|
||||
- 本阶段在 Commit gate 后直接进入 Apply;不扩大到阶段 4 全面测试迁移或阶段 5 live E2E。
|
||||
|
||||
## Pre-apply Research
|
||||
|
||||
### Reference implementations read
|
||||
|
||||
- `ChatService.executeChatComplex`、四个 complex Agent builders、Run start/save、metrics、Prompt audit、self-evaluation merge:迁移源实现和外部兼容基线。
|
||||
- `DiagnosisGraphFactory`、`DiagnosisRealGraphActionsFactory`、四个 Agent adapters、Gatekeeper/VerifiedInput/Fallback、`DiagnosisOrchestrationTraceBuilder`:Graph 路由/计数/安全材料唯一真理源。
|
||||
- `DiagnosisRun`、V011 migration、`DiagnosisTraceResponse`、`DiagnosisTraceService`、`DiagnosisTraceServiceTest`:Run JSON 字段和 exact/latest Trace 映射标准。
|
||||
- `ChatController` normal/SSE、`DiagnosisTraceController`:公开协议调用者与响应包装边界。
|
||||
- `SelfEvaluationMergeService`、`EvaluationService`、`DiagnosisTraceEvaluator`:evaluation container、异步 rule merge 和兼容消费者。
|
||||
- 本地 graph-core 1.1.2.0 `javap`:CompiledGraph `stream/invoke/state`、NodeOutput `state()`、RunnableConfig `threadId/metadata` 的真实 API。
|
||||
- `chat-*-prompt.md`:现有 Prompt 组装与 Verifier 旧 Hook payload gap。
|
||||
|
||||
### Technology inventory
|
||||
|
||||
| 类别 | 项目标准 / 本阶段使用 |
|
||||
|---|---|
|
||||
| Graph execution | `CompiledGraph.stream(initial, config)` + NodeOutput.state,当前请求线程阻塞消费 |
|
||||
| Run context | SessionContextHolder + RunnableConfig threadId/metadata;runId 是唯一执行边界 |
|
||||
| Agent assembly | ReactAgent builder、AgentLoggingHook、PlannerSkillMetadataHook/SkillsAgentHook、现有 tools/callbacks |
|
||||
| JSON persistence | Jackson map serialization;实体 String + `@JdbcTypeCode(SqlTypes.JSON)`;Flyway JSON column |
|
||||
| Trace API | Lombok DTO builder + DiagnosisTraceService parsed Map;exact/latest Run 查询 |
|
||||
| Evaluation | SelfEvaluationMergeService container;EvaluationService 按 runId 异步合并 rule channel |
|
||||
| Tests | JUnit 5,通过 public service/runtime/Trace API;只 mock repository/model/tool 外部边界 |
|
||||
| MQ/Consumer | 不涉及 |
|
||||
|
||||
### New infrastructure and reuse
|
||||
|
||||
- 新增一个深接口的 complex Chat Graph runtime/result mapper;复用既有 Graph/Node/Agent builders,不增加新依赖或第二套路由。
|
||||
- 新增 V012、DiagnosisRun 字段和 RunTrace parsed map;不新增表、repository method 或历史 backfill。
|
||||
- TDD tracer bullet 从 Trace DTO/Service 的唯一 Run 投影开始,再进入 runtime final-state mapping,最后切换 ChatService。
|
||||
- 研究未发现 devflow/OpenSpec 冲突,技术清单足以开始实现。
|
||||
|
||||
## Apply Progress
|
||||
|
||||
- TDD Trace slice RED:DiagnosisTraceServiceTest 明确缺少 DiagnosisRun builder 字段和 RunTrace getter。
|
||||
- GREEN:V012、DiagnosisRun JSON 字段、RunTrace parsed map、DiagnosisTraceService mapping 完成;10 个 Trace service tests 通过。
|
||||
- 失败分类:首次 GREEN 运行仅测试夹具用裸 ObjectMapper 无法序列化 LocalDateTime,属于 test harness 偏差;改为项目可用的 `findAndRegisterModules()` 后原始 focused loop 通过,未改业务协议。
|
||||
- Trace API JSON 断言证明顶层/session 无重复字段、run 为解析对象、无 raw 字段;null/invalid JSON fail closed。
|
||||
|
||||
## Apply Completion
|
||||
|
||||
- 复杂 Chat 已单轨切换到 `ChatDiagnosisGraphRuntime`;生产 `ChatService` 不再创建 `SequentialAgent`,不再读取 `VerifierContextHolder`,Graph Verifier 只保留 `AgentLoggingHook`。
|
||||
- runtime 使用 query-only initial state、`threadId=runId` 和 sessionId/runId metadata,流式保留最后真实 state;异常只保存真实 partial state/trace。
|
||||
- `DiagnosisGraphResultMapper` 只从 verified projection 构造兼容 self-evaluation;Verifier 未完成时不伪造 verdict,不持久化 full tool trace/raw Executor。
|
||||
- Verifier Prompt 与 runtime payload 已统一为 verified-only,Prompt audit 更新为 `chat-prompts-v2` / `chat-verifier-v3`。
|
||||
- 旧 `ChatServiceSequentialAgentTest` 因绑定已删除的固定顺序/ThreadLocal/score retry 实现而移除,由 public service/runtime/route tests 替代。
|
||||
- V012 schema 白名单测试确认只新增 `diagnosis_run.orchestration_trace JSON NULL`,无其他 schema object。
|
||||
|
||||
## Verification Summary
|
||||
|
||||
- focused regression:27 suites / 119 tests,0 failures、0 errors、0 skipped。
|
||||
- Maven test compilation:通过。
|
||||
- OpenSpec:当前 change strict pass;主 specs 14/14 strict pass。
|
||||
- 静态门禁:`git diff --check` 通过;cutover 禁用引用 0;核心 TODO/placeholder 0;schema whitelist 为 1 ALTER / 1 ADD COLUMN / 0 CREATE / 0 DROP。
|
||||
- 实现期唯一失败分类:no-answer service test 曾错误期待 inner cause 文本,属于测试断言偏差;按公开 wrapper error + FAILED/partial trace 契约修正后通过,OpenSpec 与生产代码无需变更。
|
||||
- 按阶段门禁未运行 Maven live E2E、未检查 `logs/`、未执行 `scripts/query_mysql.py`;统一保留到阶段 5。
|
||||
|
||||
## Archive Result
|
||||
|
||||
- 5 份 delta specs 已同步:新增 9、修改 13、删除 1 条 requirements。
|
||||
- OpenSpec 已归档到 `openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-chatservice-cutover/`。
|
||||
- CLI 对 proposal 结构给出非阻塞建议(大 change/delta 拆分与 SHALL/scenario 启发式);tasks 26/26、delta specs 和 strict validation 均通过,未形成验收阻塞。
|
||||
@@ -0,0 +1,21 @@
|
||||
# Chat Diagnosis StateGraph ChatService Cutover Evidence
|
||||
|
||||
## 证据
|
||||
|
||||
| 来源 | 证据 | 结论 | 是否已汇报 |
|
||||
|---|---|---|---|
|
||||
| `ChatService` + source isolation search | 复杂路径只有一次 Graph runtime 调用,无 SequentialAgent、VerifierContextHolder、VerifierInputHook 或 score retry | 单轨 cutover 完成,简单 Chat 路径未改 | 是 |
|
||||
| `ChatDiagnosisGraphRuntimeTest` / `DiagnosisRealGraphIntegrationTest` | query-only state、runId thread/metadata、Composer/Fallback、empty stream、blank answer、partial trace、Gatekeeper once/twice | Graph identity、路由和真实 partial-state 边界可观察 | 是 |
|
||||
| `DiagnosisGraphResultMapperTest` / `ChatVerifierPromptContractTest` | verified projection、effective verdict 兼容、pre-verification 无伪 verdict、禁用 raw/full trace 字段 | self-evaluation 与 Prompt 安全边界一致 | 是 |
|
||||
| `ChatServiceGraphIntegrationTest` | ChatResult、agent_flow、SUCCESS/Fallback/FAILED、metrics/Eval、partial trace、多 Run 隔离 | 生产生命周期与兼容返回满足阶段 3 规格 | 是 |
|
||||
| `DiagnosisTraceServiceTest` | `run.orchestrationTrace` 解析对象、top/session 无重复、invalid/null fail closed | Trace 字段所有权保持 Run 隔离 | 是 |
|
||||
| `DiagnosisRunSchemaContractTest` | V012 executable SQL 与唯一允许语句精确相等 | schema 仅增加一个 nullable JSON 列 | 是 |
|
||||
| focused Maven suite | 27 suites / 119 tests,0 failures/errors/skipped | Graph、Controller、Repository、Gatekeeper、Composer、Eval 回归通过 | 是 |
|
||||
| OpenSpec/static gates | change strict、14/14 主 specs、test compile、diff check、禁用引用/占位/schema whitelist 全通过 | 规格、编译和源级隔离闭环 | 是 |
|
||||
|
||||
## Evidence-driven 结论
|
||||
|
||||
- orchestration trace 必须独立于 self-evaluation 和详细 Agent/tool Trace;V012 与 RunTrace 的实现保持这一边界。
|
||||
- handled Fallback 表示安全降级答案,Run 仍为 SUCCESS;未处理、空答案或 invariant 失败才是 FAILED。
|
||||
- Verifier/Gatekeeper 只以显式 Graph state 传递可信材料;继续使用 Hook/ThreadLocal 会形成双 Gatekeeper 和不安全输入,因此生产路径已彻底移除该依赖。
|
||||
- 旧 Sequential 测试验证的是已废止内部机制,替换为 public service + runtime + route contract tests 才能保持真实回归价值。
|
||||
@@ -0,0 +1,49 @@
|
||||
# Acceptance
|
||||
|
||||
## 静态验证
|
||||
|
||||
- 旧闭包路径不存在,executable legacy refs=0。
|
||||
- current-doc stale architecture refs=0;历史兼容引用有明确限定。
|
||||
- `git diff --check` 通过;无新增 migration/schema 变更;临时调试标记=0。
|
||||
- 当前 change strict 和 16 个 main specs strict 全部通过。
|
||||
|
||||
## 脚本验证
|
||||
|
||||
- `mvn -q -DskipTests test-compile`:通过。
|
||||
- 39-suite authoritative/focused Maven command:157 tests,0 failure/error/skipped。
|
||||
- 归档前 `mvn clean` + test compilation + 43-suite deterministic command:189 tests,0 failure/error/skipped。
|
||||
- `DiagnosisTraceEvaluatorTest` + `DiagnosisEvalBaselineDiffTest`:12/12 baseline,same diff=0。
|
||||
- PowerShell parser + `InterviewDemoScriptContractTest`:通过。
|
||||
- `run-interview-demo-check.ps1 -SessionId iss-011-stage5-20260720015557 -OutputDir target/iss-011-stage5-output-current`:exit 0。
|
||||
- `scripts/query_mysql.py` exact queries:V012、Run JSON、AgentStep/ToolInvocation ownership 全部通过。
|
||||
- `openspec validate --all --strict --no-interactive`:17/17 passed;`git diff --check`、current-doc、schema、debug source、port/temp scope checks 通过。
|
||||
|
||||
## Live E2E
|
||||
|
||||
| 项目 | 结果 |
|
||||
|---|---|
|
||||
| Maven profile | `mvp-demo` |
|
||||
| sessionId | `iss-011-stage5-20260720015557` |
|
||||
| runId | `run-808ac38f-3ad0-4462-a6d0-ed50d8686473` |
|
||||
| Run | `CHAT/SUCCESS` |
|
||||
| answer / metrics | 109 chars / 75964ms / 111802 tokens / 8 steps / 12 tools |
|
||||
| Graph | `stategraph-v1`, `fallback`, `fallback_completed`, degraded=true, 3 transitions, retry=0 |
|
||||
| evaluation / feedback | non-empty / useful |
|
||||
| new ERROR | 0 |
|
||||
| DB ownership | wrong owner=0,wrong-session rows=0 |
|
||||
| process cleanup | owned PIDs stopped,9900 released |
|
||||
|
||||
## 浏览器/人工验证
|
||||
|
||||
- 不适用。本阶段验收入口是 API/PowerShell executable contract,无 UI 改动。
|
||||
|
||||
## 未验证
|
||||
|
||||
- 无 OpenSpec 必需项未验证。
|
||||
|
||||
## 归档状态
|
||||
|
||||
- ISS-011 已归档至 `mvp/issues/archived/ISS-011-chat-diagnosis-stategraph-orchestration.md`。
|
||||
- OpenSpec 归档路径:`openspec/changes/archive/2026-07-20-chat-diagnosis-stategraph-cleanup-docs`。
|
||||
- OpenSpec CLI 已同步主 specs:新增 `chat-diagnosis-stategraph-cleanup-docs`,更新 `mvp-demo-trace-acceptance` 3 项 requirement。
|
||||
- 不 push。
|
||||
@@ -0,0 +1,32 @@
|
||||
# Chat Diagnosis StateGraph Cleanup And Final Acceptance
|
||||
|
||||
## 背景
|
||||
|
||||
ISS-011 阶段 0-4 已完成 StateGraph 设计冻结、路由骨架、真实 Nodes、ChatService 单轨切换和三层权威测试。阶段 5 负责删除旧 Hook/ThreadLocal/full-trace service 闭包、对齐当前文档与 demo executable contract,并以唯一 Run 完成最终 Maven、日志和数据库验收。
|
||||
|
||||
## 目标
|
||||
|
||||
- 只保留 bounded StateGraph 复杂 Chat 编排和 verified-only Verifier 输入。
|
||||
- 让 current architecture/eval/demo 文档与 Run-owned orchestration trace 一致。
|
||||
- 先通过确定性回归,再用 Maven `mvp-demo` 证明 exact Chat/Trace/feedback、日志和数据库 ownership。
|
||||
- 所有门禁通过后关闭 ISS-011,并归档阶段 5 OpenSpec。
|
||||
|
||||
## 范围
|
||||
|
||||
- 删除 `VerifierInputHook`、`VerifierContextHolder`、`ToolTraceSummaryService` 及其 focused test。
|
||||
- 更新 current architecture/eval/demo 文档和 interview demo check。
|
||||
- 修复 live 暴露的 Graph event classloader 边界与 nested ReactAgent resume config 问题。
|
||||
- 完成 Graph/Chat/Trace/Eval 回归、Maven live E2E、日志/MySQL 核验、进程清理和 Issue 生命周期收口。
|
||||
|
||||
## 非目标
|
||||
|
||||
- 不修改公开 Chat/feedback API、Executor/Verifier/Composer 业务协议或数据库 schema。
|
||||
- 不重写 archived issues、历史 design notes 和 legacy fixtures。
|
||||
- 不删除旧 Trace/fixture 对 `tool_trace_summary` 的只读兼容。
|
||||
- 不 push,不删除失败尝试的审计数据。
|
||||
|
||||
## 元数据
|
||||
|
||||
- 分档:complex
|
||||
- OpenSpec:`chat-diagnosis-stategraph-cleanup-docs`
|
||||
- 接口影响:L2 内部类型/状态表示修复;外部 API/DTO/schema 不变
|
||||
@@ -0,0 +1,159 @@
|
||||
# Chat Diagnosis StateGraph Cleanup, Final Acceptance And Documentation Decisions
|
||||
|
||||
## Entry Summary
|
||||
|
||||
- 问题:ISS-011 运行时已切换且测试体系已收敛,但旧 Hook/ThreadLocal/service死代码、当前架构文档和最终 live 证据尚未闭环。
|
||||
- 期望:阶段 5完成清理、文档、自动化回归、eval、Maven E2E、日志/DB 验收、Issue 归档和独立提交。
|
||||
- 分档:complex;接口影响 L2 内部删除 + 文档/demo 验收增强,外部 API/DB 协议不变。
|
||||
- Change:`chat-diagnosis-stategraph-cleanup-docs`。
|
||||
- 授权:用户已明确要求直接实现;本阶段按此前规则执行唯一最终 E2E。
|
||||
|
||||
## Context Sources
|
||||
|
||||
- ISS-011 阶段 5、测试策略、协议影响、验收标准和冻结决策。
|
||||
- 阶段 0–4 OpenSpec archives、devflow acceptance 与提交 `581daff`、`42ba204`、`1460dd1`、`99e490f`、`208a231`。
|
||||
- 全仓 `VerifierInputHook`/`VerifierContextHolder`/`ToolTraceSummaryService` 定义与引用搜索。
|
||||
- `mvp/architecture/README.md` 列出的 current docs、`mvp/eval/README.md`、`mvp/demo/` scripts/checklist。
|
||||
- `mvp-demo` profile、payment-timeout request、`scripts/query_mysql.py` 和 logs/ 现有布局。
|
||||
|
||||
## Question Pool
|
||||
|
||||
| # | 维度 | 问题 | 模式 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | 清理 | 旧 Hook/ThreadLocal/trace summary service 是否还有生产消费者? | evidence-driven | 已解决 |
|
||||
| Q2 | 文档 | 哪些旧引用应更新,哪些历史材料应保留? | evidence-driven | 已解决 |
|
||||
| Q3 | E2E | 最终 live 场景如何绑定唯一 session/run 并证明 Graph 路径? | evidence-driven | 已解决 |
|
||||
| Q4 | 日志/DB | 如何避免用旧日志/latest DB 记录冒充当前证据? | evidence-driven | 已解决 |
|
||||
| Q5 | 验收 | 何时允许启动 Maven、是否需要日志和 DB 查询? | user-interview(用户最新规则) | 已确认 |
|
||||
| Q6 | 关闭 | 何时把 ISS-011 从 active 移到 archived? | evidence-driven | 已解决 |
|
||||
|
||||
## Evidence-driven Findings
|
||||
|
||||
- Q1:旧闭包只有 `VerifierInputHook -> VerifierContextHolder + ToolTraceSummaryService`,以及 `ToolTraceSummaryServiceTest`;ChatService/Graph/Trace/Eval 均无引用,可整体删除。
|
||||
- 实现前规格校正:`ChatVerifierPromptContractTest` 和 `DiagnosisGraphTestSuiteStructureTest` 必须保留旧类型名称的负向字符串断言;这不构成 executable reference。OpenSpec 已收紧为无定义/import/实例化/type-use,允许负向 guard literal。
|
||||
- Q2:current architecture index 仍列 `agent-orchestration.md` 等为当前真理源,因此必须更新;`mvp/issues/design-notes`、archived issues、历史 eval fixtures 保留时间点/兼容语义,不做大规模重写。
|
||||
- Q3:`run-interview-demo-check.ps1` 已用 Chat response runId 查询 exact Trace/feedback,最适合扩展 `run.orchestrationTrace` fail-fast 和 summary,不另建重复脚本。
|
||||
- Q4:E2E 使用唯一 timestamp sessionId;日志记录启动前 byte/time 边界并按 session/run 搜索;DB 所有核心查询带 exact sessionId/runId,另查询错误 ownership count。
|
||||
- Q6:只有实现、回归、eval、live E2E、日志、DB 和 OpenSpec门禁全部通过后,Issue checkbox 才可完成并移动到 archived。
|
||||
|
||||
## User-interview Confirmation
|
||||
|
||||
| 问题 | 用户原话 | 状态 | OpenSpec 回写 |
|
||||
|---|---|---|---|
|
||||
| Q5 最终验收节奏 | “端到端只在最后阶段全部完成后才验证……日志在log文件夹,项目库有查询数据库的py工具” | 已确认 | proposal |
|
||||
|
||||
## Grill-with-docs Result
|
||||
|
||||
- Session/Run/Trace 术语保持不变;新增强调 `orchestration_trace` 是 Run 路由摘要,不属于 self-evaluation 或日志。
|
||||
- StateGraph、Workflow/Node Contract/Chat Integration 属于实现/测试架构术语,不修改业务 glossary。
|
||||
- 当前文档必须使用 explicit Gatekeeper Node、verified-only Verifier 和 bounded evidence retry;历史设计笔记仍可描述当时 Hook 架构。
|
||||
- 删除旧闭包是阶段 0 已冻结单轨迁移的自然收尾,不形成新的难逆转权衡,无需 ADR。
|
||||
|
||||
## Discover Status
|
||||
|
||||
- `devflow/index.md`:命中阶段 0–4 archives。
|
||||
- 接口影响:L2 内部类型删除;外部 API/DTO/DB/Prompt/状态语义无变化。
|
||||
- E2E 入口/脚本/日志/DB 工具已定位;真实执行留到 Apply 最后。
|
||||
- 未解决问题:0。
|
||||
- Draft 产物:proposal + decisions;尚未生成 design/spec/tasks,尚未删除代码或启动应用。
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
### Module and evidence map
|
||||
|
||||
`ChatController -> ChatService -> ChatDiagnosisGraphRuntime -> DiagnosisRealGraphActionsFactory -> explicit Nodes -> DiagnosisGraphResultMapper -> DiagnosisRun/Trace` 是唯一当前 Chat链。旧 `VerifierInputHook -> ToolTraceSummaryService/VerifierContextHolder` 已从主链断开,删除不改变输入/输出或持久化。阶段 5新增的 demo script assertion只消费 exact Trace `run.orchestrationTrace`,DB/log检查是验收消费者,不成为运行时业务依赖。
|
||||
|
||||
| 模块 | 所有权 | 阶段 5动作 |
|
||||
|---|---|---|
|
||||
| Graph/Chat runtime | 路由、Node、Run 生命周期 | 不改行为,仅回归 |
|
||||
| Legacy Hook closure | 旧 Sequential Verifier payload | 整体删除 |
|
||||
| Current architecture docs | 当前实现真理源 | 更新 StateGraph/verified-only/trace |
|
||||
| Historical docs/fixtures | 时间点/兼容记录 | 保留,不冒充当前实现 |
|
||||
| Demo check | live Chat/Trace/feedback executable contract | 增加 exact orchestration trace fail-fast |
|
||||
| logs/MySQL | live运行证据 | 只读本次 session/run |
|
||||
| ISS/OpenSpec/devflow | 生命周期与交接 | 所有门禁通过后归档 |
|
||||
|
||||
### Lifecycle and failure ownership
|
||||
|
||||
- 自动化门禁失败:不启动 live Maven,修复代码/测试/规格后重跑。
|
||||
- live startup失败:应用未 ready,不执行 demo/DB成功声明,先读启动输出和新日志诊断。
|
||||
- Chat/Trace/feedback失败:保留 exact response/run证据,ISS保持 active。
|
||||
- log ERROR:逐条分类;未解释 ERROR阻塞验收。
|
||||
- DB不一致:以 exact run为准,不能用 API成功掩盖 persistence偏差。
|
||||
- finally:无论成功失败都停止本轮进程并确认端口,不扩大到未知已有进程。
|
||||
|
||||
### Consumer and compatibility audit
|
||||
|
||||
- 外部 API/DTO/DB consumer无迁移;demo summary仅加字段。
|
||||
- Trace UI/eval 对历史 `tool_trace_summary` 的读取保留,旧 fixture不批量迁移。
|
||||
- current docs消费者将看到新 StateGraph架构;历史链接仍可追溯 old Hook设计。
|
||||
- Issue move只改变文档位置/index,代码/运行时不依赖该路径。
|
||||
|
||||
### Cross-artifact alignment
|
||||
|
||||
| 上游 → 下游 | 检查内容 | 状态 |
|
||||
|---|---|---|
|
||||
| brief/proposal → proposal | cleanup、current docs、demo、regression、live/log/DB、Issue closure | 已对齐 |
|
||||
| proposal → design | 删除闭包、current/history边界、顺序、identity、日志/DB、cleanup | 已对齐 |
|
||||
| design → specs/tasks | 负向 literal例外、E2E字段、exact evidence、进程清理、Issue gate | 已对齐 |
|
||||
| specs → tasks | 每条 requirement有可执行 cleanup/docs/test/live/log/DB/closure slice | 已对齐 |
|
||||
|
||||
### Audit result
|
||||
|
||||
审计确认阶段 5不需要新运行时抽象或 DB migration;主要风险来自外部 live状态和证据归属,已通过 unique session/run、log boundary、exact DB queries和process ownership缓解。规格误把负向名称 literal 当 executable reference 的 gap 已修正。接口影响 L2,cross-artifact gap=0,无新 ADR。
|
||||
|
||||
## Commit Gate
|
||||
|
||||
- schema:spec-driven;proposal/design/2 delta specs/tasks 全部 done,applyRequires=`tasks` 已满足。
|
||||
- OpenSpec:当前 change strict pass;16 个主 specs strict pass。
|
||||
- Cross-artifact:4/4 已对齐,gap=0;负向 guard literal例外已写入 proposal/design/spec/tasks。
|
||||
- Question pool:5 个 evidence-driven 已解决,1 个 user-interview 已由用户原话确认,无未决项。
|
||||
- Interface impact:L2 internal type removal + demo/docs enhancement;外部协议/DB无变化。
|
||||
- Preflight:`git diff --check` 通过;尚未删除代码、修改 current docs/script或启动应用。
|
||||
- 结论:Draft OpenSpec 达到可执行状态,创建 `.committed` 后进入 Apply。
|
||||
|
||||
## Apply Progress
|
||||
|
||||
### Legacy closure removal
|
||||
|
||||
- 已删除 `VerifierInputHook`、`VerifierContextHolder`、`ToolTraceSummaryService` 和 `ToolTraceSummaryServiceTest`,四个路径均不存在。
|
||||
- `rg` 对 `src/main`、`src/test` 的旧类型扫描仅命中 `ChatVerifierPromptContractTest` 和 `DiagnosisGraphTestSuiteStructureTest` 中的负向守卫字符串;无定义、import、实例化、继承或类型依赖。
|
||||
- 删除后 focused 回归覆盖 Executor parser、Gatekeeper service/node、VerifiedInput、Verifier、Composer、Fallback、Workflow、Node Contract、Chat integration、Trace、result mapper 和结构契约:14 suites / 82 tests,0 failure、0 error、0 skipped。
|
||||
- `mvn -q -DskipTests test-compile` 通过;Graph/shared protocol 真理源保留,Spring 当前链路所需类型可完整编译。
|
||||
|
||||
### Current docs and demo contract
|
||||
|
||||
- architecture index、编排、session/trace、current MVP、evidence pipeline、quality gates、feedback、retrieval 和 eval 文档已切换为 bounded StateGraph、显式 Gatekeeper/Verified Input、verified-only Verifier、有限重试/Fallback 和 Run-owned `orchestration_trace`。
|
||||
- current-doc scan 对 `SequentialAgent`、旧 Hook/ThreadLocal/service 及旧测试类名为 0 命中;`tool_trace_summary` 仅剩 4 处,均明确标注为旧 Run/fixture 只读兼容,不是当前 Verifier 输入。
|
||||
- interview demo check 绑定 Chat 返回的 exact runId,校验 Chat/Trace ownership、Run CHAT/SUCCESS、Agent/tool/self-evaluation、Graph trace 六个字段和 feedback success;summary 新增 orchestration version、final node、termination reason、degraded、transition count 和 evidence retry count。
|
||||
- PowerShell parser 语法检查通过;`InterviewDemoScriptContractTest` 2 tests 通过,覆盖 exact runId URL/response、orchestration fail-fast 和 summary 字段。
|
||||
|
||||
### Final deterministic gates
|
||||
|
||||
- authoritative/focused regression:39 suites / 157 tests,0 failure、0 error、0 skipped;覆盖三层 Graph、全部 Graph Node/router/trace builder、Chat/Trace/Gatekeeper/Composer、Controller、Repository、schema、feedback/tool recorder 和 demo contract。
|
||||
- fixed diagnosis eval:12/12 passed,verdict distribution 为 PASS=5、LOW_CONFID=6、REJECT=1;same-baseline diff 无 regression、0 items。
|
||||
- `mvn -q -DskipTests test-compile` 通过;当前 change strict 通过,16 个主 specs strict 全部通过。
|
||||
- `git diff --check`、legacy executable refs、current-doc stale refs 和 unexpected schema change 检查全部通过。
|
||||
- focused 回归日志中的 Graph ERROR/exception stack trace 来自 `ChatServiceGraphIntegrationTest` 对 FAILED/no-answer/unhandled failure 的显式契约用例,Maven exit 0,不是未解释的 live ERROR。
|
||||
|
||||
### Live failure diagnosis and correction
|
||||
|
||||
- 首次 live identity:sessionId=`iss-011-stage5-20260717140450`,runId=`run-6db680f8-f764-49d1-995f-0e55a4b05a06`。demo contract 在 exact Trace `run.orchestrationTrace=null` 处 fail-fast,未提交 feedback;Run 为 `CHAT/FAILED`,步骤/工具均为 0。
|
||||
- 新日志根因:`DiagnosisOrchestrationTraceBuilder` 收到类名相同但 classloader identity 不同的 `OrchestrationEvent`,`instanceof` 失败并抛出 `orchestration events contain unsupported value`。这是 Spring Boot DevTools live classloader 才暴露的 Graph state 表示缺陷,单元 JVM 未复现。
|
||||
- 冲突分类:代码偏离/运行时兼容 bug,OpenSpec 对 non-empty orchestration trace 和 live Maven startup 的要求正确,不修改验收口径。
|
||||
- RED:新增 portable event map builder 回归,修复前 1 test error;GREEN:Graph state 的 production/test actions 改存 classloader-neutral Map,builder兼容 local record/Map,Node/Workflow assertions 改读 Map。
|
||||
- 修复后先运行 5-suite Graph/Node/Runtime/Chat integration focused gate,再运行完整 39 suites / 157 tests,全部 0 failure/error/skipped;首次 Maven 进程链已按 ownership 停止,9900 已释放。
|
||||
|
||||
- 第二个 live blocker 为外层 Graph `RunnableConfig` 的 resume metadata 被原样传给内层 ReactAgent,触发 `Resume request without a configured checkpoint saver`。回归先证明 nested config 与 outer config 同一且含 `HUMAN_FEEDBACK`,再改为保留 sessionId/runId、剔除 resume/state-update/checkpoint 控制信息的独立配置;5 suites / 50 tests 和随后完整 39 suites / 157 tests 通过,临时 `[DEBUG-ISS011-NODE]` 探针已删除且源码扫描为 0。
|
||||
- 2026-07-17 的一次长请求在工具执行期间遭遇外部 MySQL 瞬时 `Connection is closed`,留下精确 `RUNNING` 失败尝试;仓库查询工具随后证明数据库恢复且 server `wait_timeout=28800`。该失败未被当作验收通过,失败 Run 保留为真实审计记录。
|
||||
|
||||
### Accepted live E2E
|
||||
|
||||
- 启动:`mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"`;启动前 9900 空闲,隐藏进程链为 cmd `25080` -> Maven Java `17860` -> app Java `10732`,readiness 后执行固定 payment-timeout demo。
|
||||
- identity:sessionId=`iss-011-stage5-20260720015557`,runId=`run-808ac38f-3ad0-4462-a6d0-ed50d8686473`;Chat/Trace exact identity 一致,answer 长度 109,feedback request success。
|
||||
- Trace:Run=`CHAT/SUCCESS`,8 AgentSteps、12 ToolInvocations,self-evaluation 非空;`stategraph-v1`,final node=`fallback`,termination=`fallback_completed`,degraded=`true`,3 transitions,evidence retry count=0。
|
||||
- Fallback 原因是 Gatekeeper LOW_CONFID 且本轮模型输出缺少 `source_invocation_id`;这是按冻结契约执行的安全降级,answer 非空且未绕过 Gatekeeper,日志/DB 均可审计。
|
||||
- 最终日志发生跨日 rollover:7 月 20 日 active `application.log`/`chat.log` 全部属于本轮;`application-error.log` 最后写入仍为 7 月 17 日。本轮 `rg " ERROR "` 对 application/chat 为 0,新增 error-file bytes 为 0;session/run、Graph 75964ms、evaluation、exact Trace 和 useful feedback 均有关联日志。
|
||||
- MySQL:V012 `orchestration_trace` 为 nullable JSON;exact Run answer=109、duration=75964、token=111802、steps=8、tools=12、evaluation len=8822、trace len=438、feedback=useful;JSON 路由与 Trace 完全一致。
|
||||
- ownership:AgentStep 8、ToolInvocation 12,各自 distinct session/run=1、wrong owner=0;唯一 session 下 wrong run/step/tool 均为 0。
|
||||
- cleanup:只停止 PID `10732/17860/25080`,最终 9900 已释放,无剩余 owned process。
|
||||
@@ -0,0 +1,38 @@
|
||||
# Evidence
|
||||
|
||||
## Source And Dependency Evidence
|
||||
|
||||
- 旧闭包四个文件已删除;`src/main`/`src/test` 旧类型扫描仅剩两个测试中的负向名称守卫,无 definition/import/instantiation/type dependency。
|
||||
- 当前真理源保留 `ExecutorEvidenceParser`、`ExecutorGatekeeperService`、`GatekeeperNode`、`VerifiedInputNode`、`VerifierNodeAdapter` 和 `DiagnosisGraphResultMapper`。
|
||||
- current docs 不再描述 SequentialAgent、Hook Gatekeeper 或 full-trace Verifier;4 处 `tool_trace_summary` 均明确为历史只读兼容。
|
||||
|
||||
## Deterministic Evidence
|
||||
|
||||
- 删除后 focused:14 suites / 82 tests,0 failure/error/skipped。
|
||||
- 最终 authoritative/focused:39 suites / 157 tests,0 failure/error/skipped。
|
||||
- 归档前 `mvn clean` 后重建验证:43 suites / 189 tests,0 failure/error/skipped;额外覆盖 4 个无需外部服务的现存测试类。
|
||||
- diagnosis eval:12/12 passed;PASS=5、LOW_CONFID=6、REJECT=1;same-baseline diff=0。
|
||||
- `mvn -q -DskipTests test-compile`、PowerShell parser、OpenSpec current strict、16 main specs strict、`git diff --check`、legacy/current-doc/schema scans 均通过。
|
||||
|
||||
## Diagnose Evidence
|
||||
|
||||
- DevTools live classloader 使 record `instanceof` 边界失效;portable event map regression 先 RED,Node state 改用 Map 且 builder 兼容 record/Map 后 GREEN。
|
||||
- outer Graph resume metadata 污染 nested ReactAgent;nested config isolation regression 先 RED,保留 session/run metadata并剔除 resume/state-update/checkpoint 控制信息后 GREEN。
|
||||
- 两次修复后均重跑 focused 和完整 deterministic gate;临时调试探针为 0。
|
||||
|
||||
## Accepted Live Evidence
|
||||
|
||||
- sessionId:`iss-011-stage5-20260720015557`
|
||||
- runId:`run-808ac38f-3ad0-4462-a6d0-ed50d8686473`
|
||||
- Maven profile:`mvp-demo`;Chat/Trace/feedback script exit 0。
|
||||
- Run:CHAT/SUCCESS;answer=109 chars;8 steps;12 tools;self-evaluation 非空;feedback=useful。
|
||||
- Graph:stategraph-v1;planner -> executor -> gatekeeper -> fallback;termination=fallback_completed;degraded=true;evidence retries=0。
|
||||
- Logs:本轮 active application/chat 中 ERROR=0,error appender 无新写入;session/run、evaluation、Trace、feedback 可关联。
|
||||
- DB:V012 JSON column 存在;Run fields/JSON 与 API 一致;step/tool wrong owner=0;unique-session wrong rows=0。
|
||||
- Cleanup:owned process chain 已停止,9900 已释放。
|
||||
- Archive preflight:临时 `target/iss-011-stage5*` 目录为 0,OpenSpec strict 17/17,current-doc stale=0,schema diff=0,`git diff --check` 通过。
|
||||
|
||||
## Known Limits
|
||||
|
||||
- 本次真实模型遗漏 `source_invocation_id`,Gatekeeper 按契约降为 LOW_CONFID 并进入安全 Fallback;这是成功且可审计的 degraded Run,不是完整 Composer 正常路径。
|
||||
- 2026-07-17 外部 MySQL 瞬时断连留下一个 RUNNING 失败尝试;它未计入验收且保留审计,不影响 2026-07-20 exact accepted Run。
|
||||
@@ -0,0 +1,70 @@
|
||||
# Chat Diagnosis StateGraph Design Freeze Acceptance
|
||||
|
||||
## 结果
|
||||
|
||||
已接受。阶段 0 完成设计冻结,没有修改运行时实现。
|
||||
|
||||
## 验证
|
||||
|
||||
### 静态验证
|
||||
|
||||
- 命令:`git diff --check`
|
||||
- 结果:passed
|
||||
- 备注:仅有现有 LF/CRLF 提示,无 whitespace error。
|
||||
|
||||
- 检查:Git changed/untracked 路径运行时拒绝列表。
|
||||
- 结果:passed,12 个路径,`src/`、Maven、运行配置、脚本、数据库迁移命中 0。
|
||||
- 备注:阶段 0 只包含 Issue、glossary、OpenSpec/devflow 和执行记录。
|
||||
|
||||
- 检查:proposal → design → specs → tasks 四向对齐。
|
||||
- 结果:passed,gap=0。
|
||||
- 备注:状态、路由、Fallback、审计、测试迁移和六阶段边界均闭环。
|
||||
|
||||
### 脚本验证
|
||||
|
||||
- 命令:`openspec validate chat-diagnosis-stategraph-design-freeze --type change --strict --json`
|
||||
- 结果:passed,1/1。
|
||||
- 备注:当前 change 无结构或场景格式问题。
|
||||
|
||||
- 命令:`openspec validate --specs --strict --json`
|
||||
- 结果:passed,11/11。
|
||||
- 备注:归档前主规格基线未回归。
|
||||
|
||||
### 浏览器/人工验证
|
||||
|
||||
- 结果:not run。
|
||||
- 原因:阶段 0 无 UI 或运行行为。
|
||||
|
||||
### 未验证
|
||||
|
||||
- 单元测试:not run。阶段 0 无代码行为,新增或运行单元测试没有新的验收价值。
|
||||
- Maven E2E:not run。用户明确要求仅阶段 5 在全部实现完成后统一执行。
|
||||
- `logs/`:not inspected for stage acceptance;保留到阶段 5。
|
||||
- 数据库:not queried for stage acceptance;保留到阶段 5。
|
||||
- 备注:此前改造前 live baseline 仅为 research,不计入本阶段验收。
|
||||
|
||||
## 已完成范围
|
||||
|
||||
- 冻结最小 Graph State、所有条件边、四类独立计数和终止路径。
|
||||
- 冻结 verified-input、两类 Fallback、Run 状态与 orchestration trace 边界。
|
||||
- 冻结 L4 消费者、兼容、迁移、回滚和测试替换策略。
|
||||
- 创建 ADR-001,并规定阶段 1–5 引用本阶段 archive。
|
||||
- OpenSpec apply tasks 7/7 完成。
|
||||
|
||||
## 已知限制
|
||||
|
||||
- 运行时仍为旧 Sequential 编排,这是阶段 0 的有意状态。
|
||||
- 新设计尚未经过 Graph 编译、Node 单测、Chat 集成或最终 E2E;后续阶段逐项证明。
|
||||
- 主运行时 specs 暂时仍描述旧行为,只新增设计基线 capability。
|
||||
|
||||
## Bug 修复和诊断
|
||||
|
||||
- 更正此前错误流程模型:删除未提交的总 change,改为六个独立 sm-flow/change。
|
||||
- 更正此前错误 E2E 门禁:阶段 0–4 不做 E2E,阶段 5 统一验证。
|
||||
|
||||
## 交接
|
||||
|
||||
- 下一步:阶段 0 Git commit;提交完成后才能创建阶段 1 change。
|
||||
- Delta sync:新增 `chat-diagnosis-stategraph-design-freeze` 主 spec,共 6 个 requirements,无现有 spec 修改/删除。
|
||||
- OpenSpec 归档确认:用户已在目标中明确要求每阶段 archive,并在后续澄清中再次确认,视为已授权。
|
||||
- OpenSpec 归档结果:已同步主 spec,并归档到 `openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-design-freeze/`。
|
||||
+27
@@ -0,0 +1,27 @@
|
||||
# ADR-001: StateGraph owns run-scoped diagnosis control
|
||||
|
||||
**状态**:已接受
|
||||
**日期**:2026-07-17
|
||||
|
||||
复杂 Chat 使用 Spring AI Alibaba StateGraph 管理一次 Diagnosis Run 内的顺序、条件边、有限重试、终止和降级;ReactAgent 只执行 Planner、Executor、Verifier、Composer 的语义任务,Gatekeeper、verified-input builder、evidence retry prepare 和固定 Fallback 使用确定性 Java Node。这样可以精确恢复失败位置并形成可审计路径,同时继续复用现有 Agent、工具和证据协议。
|
||||
|
||||
Gatekeeper 从 `VerifierInputHook` 的隐式执行迁为显式且唯一的 Graph Node。Verifier 只能消费 Gatekeeper passed checked bindings 投影出的 `verified_executor_output` 和 `verified_evidence`;前置验证失败的 Fallback 不得输出 Executor claim。
|
||||
|
||||
编排审计只属于当前 `runId`:有界 `orchestration_events` 压缩后写入 `diagnosis_run.orchestration_trace`,Trace API 只在 `run.orchestrationTrace` 暴露解析对象。它不替代 self evaluation、AgentStep、ToolInvocation 或持久 Graph checkpoint,也不新增 trace 明细表。
|
||||
|
||||
## Considered Options
|
||||
|
||||
- 继续在 `ChatService` 外层叠加 `for/if`:拒绝,失败恢复位置、循环上限和路由原因仍然隐式。
|
||||
- 使用 SupervisorAgent:拒绝,当前是固定诊断 Pipeline,不需要动态选择专科 Agent。
|
||||
- 直接将父 Graph State 交给 `ReactAgent.asNode(...)`:首版拒绝,无法证明 messages、outputKey 和私有执行上下文隔离。
|
||||
- 长期保留 Sequential/StateGraph feature flag 双轨:拒绝,会形成两个编排真理源并增加安全规则漂移。
|
||||
|
||||
## Compatibility And Migration
|
||||
|
||||
`/api/chat`、`executor_evidence_v2`、Verifier 和 Composer 输出协议保持不变。Trace API 只增加 run-scoped 字段;数据库只增加 nullable JSON 列,历史 Run 不回填。
|
||||
|
||||
实施必须按六个独立 sm-flow/change 依次完成:设计冻结、路由骨架、真实节点、ChatService/Trace 切换、测试体系、清理与最终验收。阶段 1–5 必须读取阶段 0 archive,偏离本 ADR 时通过当阶段 OpenSpec 显式修正。
|
||||
|
||||
## Rollback
|
||||
|
||||
每阶段使用独立 Git commit,可整体 revert 当前阶段。生产切换后回滚代码时允许保留 nullable `orchestration_trace` 列;不得用配置重新形成长期双轨。
|
||||
@@ -0,0 +1,20 @@
|
||||
# Chat Diagnosis StateGraph Design Freeze Brief
|
||||
|
||||
## 背景
|
||||
|
||||
- 用户目标:将 ISS-011 阶段 0–5 分别作为独立 sm-flow,前一阶段 archive 并 Git commit 后才进入下一阶段。
|
||||
- 当前问题:复杂 Chat 的跨 Agent 状态机分散在 ChatService、SequentialAgent、VerifierInputHook 和 ThreadLocal 中;实现前需先冻结统一设计。
|
||||
- 关联 OpenSpec:`openspec/changes/chat-diagnosis-stategraph-design-freeze/`
|
||||
- devflow 分档:complex
|
||||
|
||||
## 范围
|
||||
|
||||
- 本次要做:冻结最小 Graph State、完整路由、有限重试、安全 Fallback、run-scoped 审计、L4 接口影响和测试替换边界。
|
||||
- 本次不做:任何 Java、SQL、Prompt、配置、运行时 spec 或运行行为修改;不运行 Maven E2E。
|
||||
- 影响区域:ISS-011、glossary、OpenSpec 设计基线、阶段 0 ADR 和后续五阶段交接契约。
|
||||
|
||||
## OpenSpec 对齐
|
||||
|
||||
- proposal 覆盖状态:已覆盖阶段 0 目标、范围、非目标、验收和风险。
|
||||
- specs 覆盖状态:新增设计基线 capability,不提前修改运行时 capabilities。
|
||||
- tasks 覆盖状态:7/7 完成,且仅包含文档、ADR 和静态验证。
|
||||
@@ -0,0 +1,120 @@
|
||||
# Chat Diagnosis StateGraph Design Freeze Decisions
|
||||
|
||||
## Question Pool
|
||||
|
||||
| # | 维度 | 问题 | 模式 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | 术语 | Diagnosis Orchestration Trace 与 Diagnosis Trace、self evaluation、Graph checkpoint 的边界是什么? | evidence-driven | 已解决 |
|
||||
| Q2 | 边界 | ISS-011 是一个总 sm-flow,还是阶段 0–5 各自独立 sm-flow? | user-interview | 已解决 |
|
||||
| Q3 | 验收 | Maven E2E 在每阶段执行还是只在最终阶段执行? | user-interview | 已解决 |
|
||||
| Q4 | 验收 | 每阶段是否强制新增并运行单元测试? | user-interview | 已解决 |
|
||||
| Q5 | 技术 | 锁定的 StateGraph 版本是否提供条件边、config-aware Node/Edge、recursion limit 和 threadId? | evidence-driven | 已解决 |
|
||||
| Q6 | 架构 | Gatekeeper、Verifier 输入与 Composer 的可信材料边界应复用什么现有契约? | evidence-driven | 已解决 |
|
||||
| Q7 | 接口 | 阶段 0 本身和最终设计分别属于什么接口影响等级? | evidence-driven | 已解决 |
|
||||
| Q8 | 归档 | 每阶段是否应在进入下一阶段前独立 archive 和 Git commit? | user-interview | 已解决 |
|
||||
|
||||
## Evidence-driven
|
||||
|
||||
| 结论 | 证据来源 | 是否已汇报用户 |
|
||||
|---|---|---|
|
||||
| Orchestration Trace 是 run 级紧凑编排摘要,不是事件日志、自评估或 Graph checkpoint | `devflow/glossary/CONTEXT.md`、ISS-011 §7.5、session-run-trace-isolation decisions | 已汇报 |
|
||||
| Gatekeeper 从 Hook 迁为显式 Node 是有意替换旧内部架构,不允许双入口 | executor-gatekeeper-hook decisions、ISS-011 §2.3/§7.2/§14 | 已汇报 |
|
||||
| Verifier verified evidence 必须来自 Gatekeeper 通过的 checked bindings;Composer 不得读取 raw Executor/tool output | verifier-evidence-reference-fidelity、executor-composer-final-answer 历史档案、ISS-011 §7.4 | 已汇报 |
|
||||
| 本地 1.1.2.0 API 支持条件边、config-aware Node/Edge、recursion limit、RunnableConfig.threadId 和 ReactAgent.call(input, config) | Maven dependency tree 与本地 JAR `javap` 研究记录 | 已汇报 |
|
||||
| 阶段 0 是文档/规格交付,无运行时接口变化;其冻结的最终目标涉及状态机、DB 和 Trace API,属于 L4 | sm-flow operating-rules、ISS-011 §10/§14 | 已汇报 |
|
||||
| 主 OpenSpec 的外层 groundedness round、Hook Gatekeeper 和完整 tool_trace_summary 输入与冻结设计冲突 | `openspec/specs/chat-verifier-agent/spec.md` 等主规格 | 已汇报 |
|
||||
|
||||
## User-interview
|
||||
|
||||
| 问题原文 | 用户原话 | 确认状态 | OpenSpec 回写 |
|
||||
|---|---|---|---|
|
||||
| ISS-011 的阶段关系如何映射 sm-flow? | “iss-011里每个阶段,都是一个sm-flow,而不是将一整个iss-011打包进一个sm-flow中” | 已确认 | 已回写 proposal 范围与非目标 |
|
||||
| Maven E2E 何时执行? | “端到端只在最后阶段全部完成后才验证” | 已确认 | 已回写 acceptance 与 out of scope |
|
||||
| 阶段单元测试是否强制? | “每个阶段如果有必要添加单元测试验收的话,就加,没有必要的话就跳过单元测试” | 已确认 | 已回写 acceptance |
|
||||
| 每阶段如何进入下一阶段? | “每个阶段需要归档完并提交才能进入下一个阶段” | 已确认 | 已回写阶段门禁 |
|
||||
|
||||
## OpenSpec Input Context
|
||||
|
||||
### devflow index
|
||||
|
||||
已命中并读取:
|
||||
|
||||
- `session-run-trace-isolation`
|
||||
- `executor-gatekeeper-hook`
|
||||
- `verifier-evidence-reference-fidelity`
|
||||
- `executor-evidence-output-contract`
|
||||
- `executor-v2-output-contract`
|
||||
- `executor-verifier-claim-checks`
|
||||
- `executor-composer-final-answer`
|
||||
- `mvp-demo-trace-acceptance`
|
||||
|
||||
### Historical constraints entering OpenSpec
|
||||
|
||||
- `runId` 是运行态和 Trace 的所有权边界。
|
||||
- AgentStep、ToolInvocation 与 self evaluation 的既有 run 绑定必须保留。
|
||||
- checked binding 是 verified evidence 的复用来源。
|
||||
- Composer 的 allowed-material 安全边界不得削弱。
|
||||
- 旧 Gatekeeper-in-Hook 决策在本设计中被显式废止,不能保留并行入口。
|
||||
- 主 OpenSpec 的旧运行时要求只能在对应实现阶段修改,阶段 0 不提前宣称代码已切换。
|
||||
|
||||
## Key Decisions
|
||||
|
||||
### 六个独立交付单元
|
||||
|
||||
阶段 0–5 分别使用独立 slug、OpenSpec change、devflow 项目、Archive 和 Git commit。前一 change 归档并提交后才创建下一 change。
|
||||
|
||||
### 阶段 0 不承载后续实现
|
||||
|
||||
阶段 0 只冻结设计并归档长期上下文。阶段 1–5 的代码、数据库、测试和清理任务不进入本 change 的 tasks。
|
||||
|
||||
### 验收分层
|
||||
|
||||
阶段 0 没有运行时行为变化,不新增或运行单元测试,也不运行 Maven E2E。验收使用 OpenSpec strict validation、结构检查和文档一致性检查。Maven E2E、`logs/` 与数据库只在阶段 5 统一执行。
|
||||
|
||||
### 接口影响
|
||||
|
||||
- 当前 change:L1 文档/设计交付,不改变调用方可观察行为。
|
||||
- 冻结目标:L4,涉及内部状态机语义、Trace API 加字段、数据库契约、迁移和回滚。
|
||||
- 兼容:保持 `/api/chat` 和 Agent 输出协议;Trace API 为加法式 run 字段。
|
||||
- 回滚:后续运行时代码可按阶段 Git revert;nullable DB 列可保留,不引入双轨配置。
|
||||
- 消费者:ChatController/ChatService、Trace DTO/Service/UI/demo、数据库迁移、评测与运维审计。
|
||||
|
||||
## Risks Accepted
|
||||
|
||||
- 阶段 0 archive 后,冻结设计已完成但运行时仍为旧 Sequential;这是有意的阶段性状态。
|
||||
- 六个 changes 的一致性由后续每次 context/commit gate 读取并审计本档案保证。
|
||||
- 不为阶段 0 的纯设计变更新增无行为价值的单元测试。
|
||||
|
||||
## Cross-Artifact 对齐检查
|
||||
|
||||
| 上游 → 下游 | 检查内容 | 状态 |
|
||||
|---|---|---|
|
||||
| brief/prd → proposal | Issue 的阶段 0 目标、范围、非目标和验收已进入 proposal | 已对齐 |
|
||||
| proposal → 设计产物 | 状态、路由、重试、Fallback、审计、测试迁移和六阶段边界均进入 design | 已对齐 |
|
||||
| 设计产物 → specs/tasks | L4 影响、安全边界、终止规则和阶段 0 文档工作均进入 spec 或 tasks | 已对齐 |
|
||||
| specs → tasks | 设计基线、源文档、ADR、严格验证和无运行时变更检查均有可执行任务 | 已对齐 |
|
||||
|
||||
Gap:无。
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
输入链为 `ChatController -> ChatService`,目标处理链为 Run 生命周期 → Graph orchestrator → 白名单 Agent/Java Nodes,输出仍为 `ChatResult + DiagnosisRun + exact RunTrace`。Graph State 只属于当前 run,verified evidence 只来自当前 run 的 passed checked bindings,编排摘要只写当前 Diagnosis Run。旧“Gatekeeper stays in VerifierInputHook”决策被 ISS-011 的显式 Node 有意替代,但旧 Gatekeeper 规则本身继续复用;旧“不新增 trace 主表”和 run ownership 决策保持成立。主要耦合风险是 Hook/Node 双执行、阶段性主 spec 漂移和 Trace 投影重复,design 已分别通过单入口迁移、按阶段修改运行时 specs、唯一 `run.orchestrationTrace` 投影缓解。架构风险已由用户确认的 ISS-011 冻结决策和六阶段交付口径接受,无需返回 grill。
|
||||
|
||||
## Commit Preflight
|
||||
|
||||
- proposal、design、specs、tasks 完整:通过。
|
||||
- 所有 user-interview 已确认:通过。
|
||||
- 未汇报 evidence-driven 结论:无。
|
||||
- L4 接口影响、消费者、兼容、迁移和回滚:已在 design 独立章节记录。
|
||||
- OpenSpec strict validation:change 1/1 通过,主 specs 11/11 通过。
|
||||
- Apply 授权:用户启动目标时要求分阶段完整执行、archive 和 Git commit,后续又确认六个独立 sm-flow;视为当前阶段连续执行授权。
|
||||
|
||||
## Apply Evidence
|
||||
|
||||
- Task 1.1:ISS-011 §3–14 与 committed design 的 State 字段、条件边、四类计数、Fallback、安全审计和测试迁移语义一致;未发现需回写 OpenSpec 的 drift,未修改运行时代码。
|
||||
- Task 1.2:将 glossary 中新 Orchestration Trace 的具体 StateGraph/Graph State 表述收敛为“诊断编排/编排上下文快照”;Verifier 的 verified-output/evidence 名称作为跨节点协议保留。
|
||||
- Task 2.1:创建 ADR-001,记录 StateGraph 控制边界、显式 Gatekeeper、run-scoped trace、替代方案、兼容、迁移和回滚。
|
||||
- Task 2.2:六个独立 change 的唯一边界已写入 design、design-baseline spec、ADR 和根计划;阶段 1–5 必须引用阶段 0 archive。
|
||||
- Task 3.1:最终 strict validation 为 change 1/1、主 specs 11/11;proposal/design/spec/tasks 对状态、Fallback、审计、运行时非目标和测试迁移均形成闭环,cross-artifact gap 为 0。
|
||||
- Task 3.2:Git 枚举 12 个 changed/untracked 路径,运行时拒绝列表命中 0;无 `src/`、Maven、运行配置、脚本或数据库迁移变更。
|
||||
- Task 3.3:单元测试未运行,因为阶段 0 无代码行为;Maven E2E、`logs/` 和数据库核验未运行,因为用户明确要求只在阶段 5 全部实现后统一执行。改造前 research baseline 不计入阶段验收。
|
||||
@@ -0,0 +1,30 @@
|
||||
# Chat Diagnosis StateGraph Design Freeze Evidence
|
||||
|
||||
## 证据
|
||||
|
||||
| 来源 | 证据 | 结论 | 是否已汇报 |
|
||||
|---|---|---|---|
|
||||
| ISS-011 §3–14 | 已定义目标流程、26 个 State 字段、有限回边、Fallback、Trace 和测试策略 | 可形成无开放分支的阶段 0 设计基线 | 是 |
|
||||
| `ChatService` / Controller 引用核查 | 生产链为 Controller → ChatService → SequentialAgent,细粒度状态散布在 ChatService | StateGraph 应只接管 Run 内控制,ChatService 保留生命周期 | 是 |
|
||||
| `VerifierInputHook` / `VerifierContextHolder` 引用核查 | Gatekeeper、工具摘要和解析结果通过 Hook/ThreadLocal 隐式传播 | Gatekeeper 与 verified-input 必须迁为显式 Node,不能保留双入口 | 是 |
|
||||
| executor-gatekeeper-hook 历史 decisions | 旧阶段曾决定 Gatekeeper 留在 Hook 且不重试 | 本设计有意替换入口,但继续复用确定性规则 | 是 |
|
||||
| session-run-trace-isolation 历史 decisions | runId 是执行/Trace 所有权边界,AgentStep/ToolInvocation 已承载明细 | orchestration trace 只附着当前 Run,不新增明细主表 | 是 |
|
||||
| Verifier/Composer 历史档案 | checked bindings 和 allowed material 是可信边界 | Builder 不重复验真,Fallback 不泄漏 raw claim | 是 |
|
||||
| 本地 Graph Core 1.1.2.0 JAR | 已验证条件边、config-aware Node/Edge、recursion limit、threadId 与 Agent call API | 阶段 1 可基于锁定签名实现,不采用网上漂移示例 | 是 |
|
||||
| 主 OpenSpec 核查 | 旧 spec 仍描述 Hook Gatekeeper、完整 tool trace 和外层 groundedness round | 阶段 0 不提前改运行时 spec,后续对应实现阶段再修改 | 是 |
|
||||
| 用户原话 | 六个独立 sm-flow;仅阶段 5 E2E;单测按必要性 | 阶段门禁与测试口径已固定 | 是 |
|
||||
|
||||
## Evidence-driven 结论
|
||||
|
||||
- 结论:阶段 0 运行时影响为 L1,但冻结目标为 L4。
|
||||
- 证据:本阶段 changed paths 无 `src/`/SQL/配置;最终目标改变状态机、Trace API 和 DB。
|
||||
- 风险:设计完成不等于运行时完成。
|
||||
- 用户确认:已确认分阶段交付。
|
||||
- 结论:旧 Gatekeeper-in-Hook 决策应被显式 Node 有意替代。
|
||||
- 证据:当前 Hook 引用与 ISS-011 条件边要求冲突。
|
||||
- 风险:迁移期双执行。
|
||||
- 用户确认:ISS-011 冻结决策已确认。
|
||||
- 结论:阶段 0 不需要单元测试或 E2E。
|
||||
- 证据:Git 运行时拒绝列表命中 0,OpenSpec strict 和文档一致性已覆盖本阶段可观察产物。
|
||||
- 风险:运行时正确性仍未证明,将由阶段 1–5 测试和最终 E2E 证明。
|
||||
- 用户确认:已确认。
|
||||
@@ -0,0 +1,42 @@
|
||||
# Chat Diagnosis StateGraph Real Nodes 验收
|
||||
|
||||
## 结果
|
||||
|
||||
已接受。OpenSpec tasks 27/27 完成;真实 Nodes 可构造和测试,旧生产 Sequential 路径保持可用,阶段 2 未切换生产入口。
|
||||
|
||||
## 静态验证
|
||||
|
||||
- `openspec validate chat-diagnosis-stategraph-real-nodes --strict`:通过。
|
||||
- `openspec validate --specs --strict`:13 passed,0 failed。
|
||||
- `git diff --check`:通过;只有 LF/CRLF 转换提示,无 whitespace error。
|
||||
- 源码/引用检查:`ChatService` Graph 引用 0;新增源码 forbidden refs 0;protocol 反向依赖 0;DB/Trace/Prompt diff 0。
|
||||
- 装配对齐:无核心 TODO/FIXME/placeholder;Graph factory 继续拥有所有 retry counter 与 Planner reset。
|
||||
|
||||
## 脚本验证
|
||||
|
||||
- 新 protocol/Node/CompiledGraph/Router/Trace focused suite:通过。
|
||||
- `mvn -q "-Dtest=ChatServiceSequentialAgentTest,VerifierInputHookTest,ExecutorGatekeeperServiceTest" test`:通过。
|
||||
- 两组测试合计:100 tests,0 failures,0 errors,0 skipped。
|
||||
- `mvn -q -DskipTests test-compile`:通过。
|
||||
|
||||
## 浏览器/人工验证
|
||||
|
||||
- 未运行。阶段 2 不改变用户入口或 UI,没有独立人工验收价值。
|
||||
|
||||
## 未验证
|
||||
|
||||
- Maven live E2E、`logs/` 日志和 `scripts/query_mysql.py` 数据库核验未运行。
|
||||
- 原因:用户明确要求只在阶段 5 全部实现后统一做最终 E2E;阶段 2 尚未切换生产入口。
|
||||
- 风险:当前验收只证明组件/Graph 契约与旧路径回归,不证明真实生产装配、模型、DB 和 Trace 全链路。
|
||||
|
||||
## 已完成范围
|
||||
|
||||
- 中立共享 protocol 与旧路径行为保持型委托。
|
||||
- 四个 Agent adapters、显式 Gatekeeper、可信投影、关键证据补查、Composer 与两类 Fallback。
|
||||
- 真实 CompiledGraph 装配、critical gap 路由修复和完整 focused test 证据。
|
||||
|
||||
## 交接
|
||||
|
||||
- 下一步:归档本 change、独立提交阶段 2,然后启动阶段 3 `chat-diagnosis-stategraph-chatservice-cutover`。
|
||||
- OpenSpec 归档:已同步主 specs,并归档到 `openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-real-nodes/`。
|
||||
- 归档授权:用户已明确“直接实现吧,不用找我授权了”,授权后续阶段在门禁通过后直接归档和提交。
|
||||
@@ -0,0 +1,21 @@
|
||||
# Chat Diagnosis StateGraph Real Nodes Brief
|
||||
|
||||
## 背景
|
||||
|
||||
- 用户目标:将 ISS-011 阶段 2 作为独立 sm-flow,接入真实 Agent/Java Nodes,并在归档、验收和独立 Git 提交后才进入阶段 3。
|
||||
- 当前问题:阶段 1 只有 Fake Node 路由骨架;Executor/Verifier/Composer 协议逻辑分散在 Hook 与 ChatService,真实 Graph 尚不能安全调用 Agent、Gatekeeper 或投影可信材料。
|
||||
- 关联 OpenSpec:`openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-real-nodes/`
|
||||
- devflow 分档:complex
|
||||
|
||||
## 范围
|
||||
|
||||
- 本次要做:共享无状态 protocol 组件;Planner/Executor/Verifier/Composer adapters;显式 Gatekeeper、Verified Input、Evidence Retry、Fallback Nodes;真实 CompiledGraph 装配和 focused tests。
|
||||
- 本次不做:不切换 ChatService 生产入口,不修改 DB、Trace API、Prompt 契约,不删除旧 Sequential/Hook,不运行 live E2E。
|
||||
- 影响区域:`diagnosis.protocol`、`graph.diagnosis`、`VerifierInputHook`、`ChatService` 共享逻辑委托及对应测试。
|
||||
|
||||
## OpenSpec 对齐
|
||||
|
||||
- proposal 覆盖状态:已覆盖。
|
||||
- design 覆盖状态:已覆盖;接口影响为 L2,生产切换明确延期到阶段 3。
|
||||
- specs 覆盖状态:已覆盖真实 Nodes 安全边界,并修正 critical evidence gap 路由条件。
|
||||
- tasks 覆盖状态:27/27 完成。
|
||||
@@ -0,0 +1,218 @@
|
||||
# Chat Diagnosis StateGraph Real Nodes Decisions
|
||||
|
||||
## Question Pool
|
||||
|
||||
| # | 维度 | 问题 | 模式 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | 边界 | 阶段 2 是否切换 ChatService/DB/Trace? | user-interview(六阶段已确认) | 已解决 |
|
||||
| Q2 | 复用 | 新 Nodes 如何避免复制 Hook/ChatService 解析与安全渲染? | evidence-driven | 已解决 |
|
||||
| Q3 | Agent | Adapter 如何调用真实 ReactAgent 并保留 Hook/ToolCallback/run config? | evidence-driven | 已解决 |
|
||||
| Q4 | Gatekeeper | 显式 Node 如何保证当前 run、单次调用和 fail-closed 标准化? | evidence-driven | 已解决 |
|
||||
| Q5 | 安全 | passed checked bindings 如何投影为 verified output/evidence? | evidence-driven | 已解决 |
|
||||
| Q6 | 补证据 | 哪些 facts 可触发 evidence retry,completed queries 如何表达? | evidence-driven | 已解决 |
|
||||
| Q7 | Fallback | 前置验证失败与 Composer 后置失败可分别使用哪些材料? | evidence-driven | 已解决 |
|
||||
| Q8 | Prompt | 阶段 2 是否立即修改共享 Verifier Prompt? | evidence-driven | 已解决 |
|
||||
| Q9 | 验收 | 阶段 2 是否需要单元/集成测试与 E2E? | evidence-driven + user rule | 已解决 |
|
||||
|
||||
## Evidence-driven
|
||||
|
||||
| 结论 | 证据来源 | 是否已汇报用户 |
|
||||
|---|---|---|
|
||||
| ReactAgent.call(input, config) 返回 AssistantMessage,现有 AgentLoggingHook 从 config metadata 读取 sessionId/runId | 本地 1.1.2.0 javap、AgentLoggingHook | 已汇报 |
|
||||
| Executor parser 当前在 VerifierInputHook,Verifier/Composer parser 与 renderer 当前在 ChatService | 定向源码阅读和 references | 已汇报 |
|
||||
| checked_bindings 提供 claim_id/tool_name/source_invocation_id/raw_path/matched_text/status | ExecutorGatekeeperService | 已汇报 |
|
||||
| Gatekeeper.validateRun 使用 run-scoped ToolInvocation repository | ExecutorGatekeeperService tests/code | 已汇报 |
|
||||
| Graph path不能注册旧 VerifierInputHook,否则 Gatekeeper 会双执行且输入包含 full tool trace | 阶段 0 ADR、Hook 源码 | 已汇报 |
|
||||
| evidence retry 只允许 critical no_evidence/indirect_support | 阶段 0 design/ISS-011 | 已汇报 |
|
||||
| 阶段 1 router 未检查 is_critical,属于实现偏离 | Router 源码与阶段 0 baseline 对照 | 已汇报 |
|
||||
| 共享 Prompt 暂时仍服务旧 Sequential path,本阶段直接修改会提前破坏生产 payload | ChatService agent builder + VerifierInputHook + prompt | 已汇报 |
|
||||
| Node 输入/安全边界新增且旧解析会重构,单元和 focused regression 必要;E2E 不必要 | 阶段边界和用户规则 | 已汇报 |
|
||||
|
||||
## User-interview
|
||||
|
||||
| 问题原文 | 用户原话 | 确认状态 | OpenSpec 回写 |
|
||||
|---|---|---|---|
|
||||
| 阶段 2 是否独立 sm-flow? | “iss-011里每个阶段,都是一个sm-flow” | 已确认 | 独立 change |
|
||||
| 是否可提前切生产入口? | “每个阶段需要归档完并提交才能进入下一个阶段” | 已确认不可提前阶段 3 | Out of Scope |
|
||||
| 阶段 2 是否执行 E2E? | “端到端只在最后阶段全部完成后才验证” | 已确认不执行 | Acceptance |
|
||||
| 是否添加单元测试? | “如果有必要添加单元测试验收的话,就加” | 已确认规则;本阶段判定必要 | Acceptance |
|
||||
|
||||
## Context And Handoff
|
||||
|
||||
- 阶段 0 archive/commit:design baseline / `581daff`
|
||||
- 阶段 1 archive/commit:routing skeleton / `42ba204`
|
||||
- 当前 change:`chat-diagnosis-stategraph-real-nodes`
|
||||
- 后续 change:`chat-diagnosis-stategraph-chatservice-cutover`,只能在本阶段 archive + commit 后创建。
|
||||
|
||||
## Technical Decisions
|
||||
|
||||
### Shared protocol components
|
||||
|
||||
- 抽取 Executor、Verifier、Composer 解析器,旧 Hook/ChatService 委托新组件。
|
||||
- 抽取 Composer safe input builder / renderer,旧 ChatService 保持相同输出。
|
||||
- JSON sanitization 只存在一份共享实现,不在每个 Node copy。
|
||||
- 抽取过程是行为保持 refactor;现有 focused tests 是回归门禁。
|
||||
|
||||
### Agent adapter boundary
|
||||
|
||||
- `DiagnosisAgentInvoker` 是最小 port:`invoke(String, RunnableConfig) -> String`。
|
||||
- `ReactAgentDiagnosisInvoker` 只包装 `ReactAgent.call(...).getText()`。
|
||||
- Planner/Executor/Verifier/Composer adapters 各自拥有白名单 input projector、parser 和 status event。
|
||||
- tests 使用 fake invoker,另有 ReactAgent wrapper test。
|
||||
|
||||
### Legacy Hook coexistence
|
||||
|
||||
- 旧 Sequential path 在阶段 3 前仍通过 VerifierInputHook 执行 Gatekeeper。
|
||||
- Graph Verifier Agent 不注册该 Hook;显式 Gatekeeper Node 是 Graph 中唯一 validation 入口。
|
||||
- Parser/enricher 可共享,但 Hook 的 legacy payload/prompt 暂不改变。
|
||||
- 阶段 3 切换生产入口并同步 Verifier Prompt;阶段 5 删除旧隐式结构。
|
||||
|
||||
### Gatekeeper normalization
|
||||
|
||||
- raw pass → PASS。
|
||||
- raw fail + severity low_confid → LOW_CONFID。
|
||||
- raw fail + severity reject → REJECT。
|
||||
- 缺失、unknown、异常或不一致 → REJECT。
|
||||
- verified_binding_count 只统计 checked_bindings.status=pass。
|
||||
|
||||
### Verified projection
|
||||
|
||||
- 通过项按 claim_id + source_invocation_id + tool_name + raw_path 与原 binding 精确匹配。
|
||||
- verified_executor_output 只包含 answer_version 与至少一条 passed binding 的 filtered claims;合法零 claim 保持空列表。
|
||||
- verified_evidence 只含 claim_id/source_invocation_id/tool_name/raw_path/matched_text。
|
||||
- hypotheses、失败 binding、未引用 ToolInvocation、raw Executor output 不投影。
|
||||
|
||||
### Evidence retry
|
||||
|
||||
- extractor 要求 `is_critical=true` 且 verification 为 no_evidence/indirect_support。
|
||||
- gap 字段:claim_id(从 `claim-id: text` 提取或空)、fact、verification、reason。
|
||||
- completed queries 由 verified evidence 的 tool/invocation/path 去重生成。
|
||||
- prior verified output/evidence 原样只读进入 retry_context。
|
||||
- 约束固定:max_retry=1、do_not_repeat_successful_queries、only_execute_incremental_queries、preserve_prior_verified_claims。
|
||||
- Executor 输入声明“增量执行、完整输出”,Java 不合并 claims。
|
||||
|
||||
### Interface impact
|
||||
|
||||
- 等级:L2 internal interface;共享 parser 委托保持旧行为。
|
||||
- 新消费者:阶段 3 Graph orchestrator/Agent factory。
|
||||
- 当前外部 API/DB/生产路由:无变化。
|
||||
- 回滚:revert 本阶段提交;旧 Sequential path仍完整。
|
||||
- 有意修正:Router 只允许 critical gap,属于对冻结基线的代码修复。
|
||||
|
||||
## Risks Accepted
|
||||
|
||||
- 真实模型尚未执行;Node contract 通过 fake invoker/real service tests证明,阶段 5 才 E2E。
|
||||
- Prompt 输入说明与 Graph input 的最终同步推迟到阶段 3,避免当前旧生产 Hook 提前不兼容。
|
||||
- 旧 Hook 暂时仍存在,但不进入 Graph action graph;阶段 5 必须删除。
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
### Module and ownership map
|
||||
|
||||
`Diagnosis Context + RunnableConfig` → Agent adapters / deterministic Nodes → owned Graph State fields → `DiagnosisGraphRouter` → next Node or END → `final_answer + orchestration_events`。阶段 2 只提供这条可构造链,阶段 3 才由 `ChatService` 创建 Run、构造 Agent 实例并调用 Graph。
|
||||
|
||||
| 模块 | 数据所有权 | 允许依赖 |
|
||||
|---|---|---|
|
||||
| `diagnosis.protocol` | JSON contract、纯解析结果、安全 input/rendering | Jackson 与纯 DTO;不依赖 Graph/Hook/ChatService/ThreadLocal |
|
||||
| `graph.diagnosis` Agent adapters | 白名单输入、Agent attempt status/event | protocol、注入的 invoker、Graph API |
|
||||
| Gatekeeper Node | raw Gatekeeper result、normalized status/count | ExecutorGatekeeperService、当前 run config |
|
||||
| Verified Input / Retry Prepare | verified projection、critical gaps、retry context | raw result 的只读投影与 protocol DTO |
|
||||
| Router / Factory | 条件边、有限计数、调用次序 | Graph State;不解析 Agent raw output |
|
||||
| legacy Hook / ChatService | 阶段 3 前的生产 Sequential 流程 | 只委托 protocol;不得消费真实 Graph action set |
|
||||
|
||||
### Lifecycle and coupling audit
|
||||
|
||||
Graph State 和 events 都是 invocation-scoped,runId 只从当前 RunnableConfig 获取,Nodes 不持有跨 Run 可变状态。旧路径与 Graph 路径阶段性共享的只有无状态 protocol 组件和 Gatekeeper service,不共享 ThreadLocal 或 Agent output。ReactAgent/invoker 必须构造注入,阶段 2 不复制 ChatService Prompt/Agent factory。唯一有意的阶段性耦合是 legacy consumers 改为委托 protocol,这由旧 focused tests 和完整 revert 保护。阶段 3 前生产入口、DB、Trace、Prompt 均保持隔离。
|
||||
|
||||
### Cross-artifact alignment
|
||||
|
||||
| 对齐链 | 结果 | 证据 |
|
||||
|---|---|---|
|
||||
| brief/proposal 目标、范围、非目标 → proposal | 已对齐 | 独立阶段边界、真实 Nodes、无生产切换均明确 |
|
||||
| proposal 承诺与约束 → design | 已对齐 | invoker、共享组件、Gatekeeper、投影、retry、Fallback、L2 均有决策 |
|
||||
| design 架构/接口结论 → specs/tasks | 已对齐 | 中立 protocol 包、构造注入、fail-closed 和生产隔离均有任务/行为 |
|
||||
| specs 可观察行为 → tasks 可执行切片 | 已对齐 | 每个 requirement 至少由一个实现任务和一个测试/验收任务覆盖 |
|
||||
|
||||
### Audit result
|
||||
|
||||
审计发现共享协议组件包所有权与 Agent 实例装配边界需要显式化,已回写 design/tasks。未发现状态字段、路由计数、run 生命周期或阶段边界的新冲突。接口影响维持 L2,消费者都在本 change 与下一阶段明确范围内。剩余风险是共享抽取的旧行为漂移和 binding 投影泄漏,均有 focused regression 与 mixed-binding tests。架构风险可接受,无未解决问题。
|
||||
|
||||
## Commit Gate
|
||||
|
||||
- proposal/design/specs/tasks:全部存在,OpenSpec status `isComplete=true`。
|
||||
- 当前 change strict validation:通过。
|
||||
- 主 specs strict validation:13 passed,0 failed。
|
||||
- 规格结构:9 requirements、22 scenarios;tasks:27 个 checkbox 切片。
|
||||
- Cross-artifact:4/4 已对齐,gap=0。
|
||||
- 接口影响:L2,已在 design 独立章节记录消费者、兼容和回滚。
|
||||
- Question pool:所有 evidence-driven 已汇报;所有 user-interview 已确认;无未决项。
|
||||
- Preflight scope:`git diff --check` 通过;Commit checkpoint 未修改业务代码。
|
||||
- 结论:Draft OpenSpec 达到可执行状态,创建 `.committed`。
|
||||
|
||||
## Apply Authorization
|
||||
|
||||
- 用户原话:“直接实现吧,不用找我授权了”。
|
||||
- 解释:阶段 2–5 后续 checkpoint 可在前置门禁通过后直接继续,不再因 Apply 或 Archive 授权暂停。
|
||||
- 不扩大范围:每阶段仍须独立 OpenSpec、验收归档、Git commit;阶段 0–4 不做 E2E,阶段 5 才统一执行。
|
||||
|
||||
## Pre-apply Research
|
||||
|
||||
### Reference implementations read
|
||||
|
||||
- `src/main/java/com/superbiz/agent/hook/VerifierInputHook.java`:Executor JSON sanitization、parse status、tool-name normalization、唯一 invocation 回填与 legacy Gatekeeper payload。
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`:Verifier claim/fact parser、Gatekeeper ceiling、Composer allowed-material builder、Composer parser/safe renderer、旧 retry context。
|
||||
- `src/main/java/com/superbiz/agent/service/ExecutorGatekeeperService.java`:`validateRun`、severity normalization source、checked binding 与 matched_text 结构。
|
||||
- `src/main/java/com/superbiz/agent/graph/diagnosis/DiagnosisGraphFactory.java`:config-aware action ports、technical retry/evidence retry counter 所有权和 recursion limit。
|
||||
- `src/main/java/com/superbiz/agent/graph/diagnosis/OrchestrationEvent.java` 与 `src/test/java/com/superbiz/agent/graph/diagnosis/ScriptedDiagnosisGraphActions.java`:每 attempt 单 event 形态。
|
||||
- `src/main/java/com/superbiz/agent/hook/AgentLoggingHook.java`:RunnableConfig metadata 中 sessionId/runId 的审计读取方式。
|
||||
- `src/test/java/com/superbiz/agent/hook/VerifierInputHookTest.java`、`ChatServiceSequentialAgentTest.java`、`DiagnosisGraphRoutingTest.java`:JUnit 5、Mockito 边界替身和 CompiledGraph observable behavior 测试风格。
|
||||
- 本地 `1.1.2.0` JAR `javap`:`ReactAgent.call(String,RunnableConfig)` 与 `RunnableConfig.threadId/metadata` 公共 API。
|
||||
|
||||
### Technology inventory
|
||||
|
||||
| 类别 | 项目标准 / 本阶段使用 |
|
||||
|---|---|
|
||||
| JSON contract | Jackson `ObjectMapper`、`LinkedHashMap` 保持稳定字段顺序;共享 sanitization 只存在一份 |
|
||||
| Graph Node | `AsyncNodeActionWithConfig` 返回 `CompletableFuture<Map<String,Object>>`;state 默认 Replace、events Append |
|
||||
| Agent boundary | 构造注入的 `DiagnosisAgentInvoker`;ReactAgent wrapper 透传同一个 RunnableConfig |
|
||||
| Run scope | `config.metadata("runId")` 为 Gatekeeper 唯一运行边界;缺失即 fail closed |
|
||||
| Error handling | 合法 contract/invalid output/temporary/permanent 分离;unknown exception 不推断为可重试 |
|
||||
| Tests | JUnit 5;只在 ReactAgent、repository 等系统边界使用 fake/mock;Node/CompiledGraph 走公开 action/graph 接口 |
|
||||
| Request/response | 本阶段不改 Controller/DTO/API,不适用 |
|
||||
| MQ/Consumer | 本阶段不涉及,不适用 |
|
||||
| DB/Trace/Prompt | 本阶段禁止修改,阶段 3 处理 |
|
||||
|
||||
### Reuse and new infrastructure
|
||||
|
||||
- 新建中立 `com.superbiz.agent.diagnosis.protocol`:`JsonPayloadSupport`、Executor/Verifier/Composer protocol、safe input/rendering、`EvidenceGapExtractor`;不得依赖 Graph/Hook/ChatService/ThreadLocal。
|
||||
- 新建 `graph.diagnosis` Node 层:invoker wrapper、failure classifier、四个 Agent adapters、Gatekeeper、Verified Input、Retry Prepare、Fallback 和 action assembly。
|
||||
- 不新建 Prompt/Agent factory;阶段 3 通过构造注入已有 ReactAgent 实例。
|
||||
- 不新增 Maven dependency、数据库迁移、配置项或生产 consumer。
|
||||
|
||||
### Pre-apply conclusion
|
||||
|
||||
参考实现、API 签名、异常和测试标准已足以指导实现;未发现 devflow/OpenSpec 冲突。进入 TDD tracer bullet,先锁定共享 Executor parser 的合法 no-evidence 与 malformed 行为。
|
||||
|
||||
## Apply Completion And Verification
|
||||
|
||||
### Assembly alignment
|
||||
|
||||
- 首个共享 protocol 模块与最终真实 Graph action assembly 均逐项对照 design/specs:invoker 显式透传 RunnableConfig,Gatekeeper 单次 fail-closed,Verified Input 精确投影,Verifier/Composer 固定输入重试,critical evidence retry 与两类 Fallback 边界全部落地。
|
||||
- `DiagnosisGraphFactory` 继续独占技术重试计数、evidence retry 计数和 Planner mode/reset;Node 不重复拥有编排计数。
|
||||
- 新增源码中无 `ThreadLocal`、`VerifierInputHook`、`tool_trace_summary`、`raw_executor`、TODO 或 FIXME;protocol 包无 Graph/Hook/ChatService 反向依赖。
|
||||
- `ChatService` 无 Diagnosis Graph/CompiledGraph 引用,生产切换保持在阶段 3;DB migration、Trace DTO/entity/repository 和 prompts 均无 diff。
|
||||
|
||||
### Automated verification
|
||||
|
||||
- 新 protocol/Node/真实 CompiledGraph/Router/Trace focused suite:通过。
|
||||
- 旧 `ChatServiceSequentialAgentTest`、`VerifierInputHookTest`、`ExecutorGatekeeperServiceTest` 回归:通过。
|
||||
- 合计 100 tests,0 failures,0 errors,0 skipped。
|
||||
- `mvn -q -DskipTests test-compile`:通过。
|
||||
- `openspec validate chat-diagnosis-stategraph-real-nodes --strict`:通过。
|
||||
- `openspec validate --specs --strict`:13 passed,0 failed。
|
||||
- `git diff --check`:通过;仅报告 Git 既有 LF/CRLF 转换提示,无 whitespace error。
|
||||
|
||||
### Deferred final verification
|
||||
|
||||
- 阶段 2 按用户确认不运行 Maven live E2E,不启动应用,不检查 `logs/`,不查询数据库。
|
||||
- 上述端到端、日志和 `scripts/query_mysql.py` 数据库核验统一保留到阶段 5 全部实现完成后执行。
|
||||
@@ -0,0 +1,25 @@
|
||||
# Chat Diagnosis StateGraph Real Nodes Evidence
|
||||
|
||||
## 证据
|
||||
|
||||
| 来源 | 证据 | 结论 | 是否已汇报 |
|
||||
|---|---|---|---|
|
||||
| `VerifierInputHook.java`、`ChatService.java` | Executor/Verifier/Composer 解析与安全渲染原实现 | 抽取到中立 protocol 并让旧路径委托,避免双真理源 | 是 |
|
||||
| `ExecutorGatekeeperService.java` | `validateRun` 与 checked binding/matched_text 契约 | Graph Gatekeeper 必须按当前 runId 单次调用并 fail closed | 是 |
|
||||
| 本地 Graph/ReactAgent 1.1.2.0 API | `ReactAgent.call(String,RunnableConfig)` 与 config-aware Graph action | 最小 invoker 可精确透传输入和当前 Run metadata | 是 |
|
||||
| 阶段 0/1 OpenSpec archives | 路由、计数、状态与安全边界冻结 | 阶段 2 不改变 Graph counter 所有权或生产入口 | 是 |
|
||||
| 新 protocol/Node/CompiledGraph tests | PASS、REJECT、LOW_CONFID、critical retry、固定输入重试和安全 fallback | 真实 Nodes 的可观察路径和材料边界已覆盖 | 是 |
|
||||
| 旧 Sequential/Hook/Gatekeeper tests | 共享抽取后的旧路径回归 | 阶段 2 未破坏当前生产控制流 | 是 |
|
||||
|
||||
## Evidence-driven 结论
|
||||
|
||||
- Graph Verifier 不得注册旧 `VerifierInputHook`,否则会双执行 Gatekeeper 并泄漏完整 tool trace。
|
||||
- Verified Input 必须用 claim/invocation/tool/path 精确关联 passed binding;不能复刻 Gatekeeper 判断或读取未引用工具结果。
|
||||
- Router 与 Retry Prepare 必须共用 critical-gap 提取规则:仅 `is_critical=true` 的 `no_evidence`/`indirect_support`。
|
||||
- 技术重试输入必须字节一致,且不得重跑前序 Node;计数仍由 Graph factory 统一拥有。
|
||||
- 当前 ChatService 无 Graph 引用,DB/Trace/Prompt 无 diff,满足阶段 2 的生产隔离要求。
|
||||
|
||||
## 风险与后续证据
|
||||
|
||||
- 本阶段使用 fake invoker 和 focused tests,不证明真实模型/外部基础设施联通;阶段 5 最终 E2E 统一补证。
|
||||
- Prompt 输入说明、生产装配、Run/Trace 持久化属于阶段 3,不能提前从阶段 2 证据推断已完成。
|
||||
@@ -0,0 +1,69 @@
|
||||
# Chat Diagnosis StateGraph Routing Skeleton Acceptance
|
||||
|
||||
## 结果
|
||||
|
||||
已接受。阶段 1 完成未接生产入口的 Graph 骨架和 Fake Node 路由体系。
|
||||
|
||||
## 验证
|
||||
|
||||
### 静态验证
|
||||
|
||||
- 命令:`git diff --check`
|
||||
- 结果:passed。
|
||||
- 检查:production refs、forbidden deps、23 changed paths 白名单。
|
||||
- 结果:passed,outside refs=0,forbidden refs=0,out-of-scope=0。
|
||||
|
||||
### 脚本验证
|
||||
|
||||
- 命令:`mvn -q "-Dtest=DiagnosisGraphRoutingTest,DiagnosisOrchestrationTraceBuilderTest" test`
|
||||
- 结果:passed,35 tests(29 routing + 6 trace),0 failure/error。
|
||||
- 覆盖:正常、三类技术 retry、Executor 不重试、Gatekeeper、evidence retry、Composer、unknown fail-closed、threadId、Append events 和 trace。
|
||||
|
||||
- 命令:`mvn -q "-DskipTests" test`
|
||||
- 结果:passed。
|
||||
- 覆盖:main/test compilation。
|
||||
|
||||
- 命令:`openspec validate chat-diagnosis-stategraph-routing-skeleton --type change --strict --json`
|
||||
- 结果:passed,1/1。
|
||||
|
||||
- 命令:`openspec validate --specs --strict --json`
|
||||
- 结果:passed,12/12(归档前)。
|
||||
|
||||
### 浏览器/人工验证
|
||||
|
||||
- 结果:not run。
|
||||
- 原因:无 UI 或生产入口变化。
|
||||
|
||||
### 未验证
|
||||
|
||||
- Maven E2E:not run,用户要求仅阶段 5 全部实现后统一执行。
|
||||
- `logs/`:not inspected,保留到阶段 5。
|
||||
- 数据库:not queried,保留到阶段 5。
|
||||
- 真实模型/Agent:not invoked,属于阶段 2。
|
||||
|
||||
## 已完成范围
|
||||
|
||||
- 显式 Graph Core direct dependency。
|
||||
- 26 state keys、typed status、topology/route/reason constants。
|
||||
- 八个 config-aware action ports 和真实 CompiledGraph factory。
|
||||
- Planner/Verifier/Composer 技术计数、一次 evidence retry、recursion limit 32。
|
||||
- fail-closed router、events Append 和 trace builder。
|
||||
- 35 个 Fake Node/trace tests。
|
||||
|
||||
## 已知限制
|
||||
|
||||
- Skeleton 没有生产消费者,阶段 3 才切换 ChatService。
|
||||
- Fake Node 只证明控制流,不证明真实 Agent JSON、Gatekeeper 或安全输入映射。
|
||||
- orchestration trace 尚未持久化或通过 API 暴露。
|
||||
|
||||
## Bug 修复和诊断
|
||||
|
||||
- 架构审计发现并修正 Graph Core 传递依赖所有权,改为 BOM 管理的直接依赖。
|
||||
- 自查补齐全部正常节点 threadId、Verifier REJECT 和缺失 effective verdict 场景。
|
||||
|
||||
## 交接
|
||||
|
||||
- 下一步:阶段 1 Git commit;完成后才能创建阶段 2 change。
|
||||
- Delta sync:新增 `chat-diagnosis-stategraph-routing-skeleton` 主 spec,共 6 requirements。
|
||||
- OpenSpec 归档确认:用户已要求每阶段 archive,授权已存在。
|
||||
- OpenSpec 归档结果:已同步主 spec,并归档到 `openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-routing-skeleton/`。
|
||||
@@ -0,0 +1,21 @@
|
||||
# Chat Diagnosis StateGraph Routing Skeleton Brief
|
||||
|
||||
## 背景
|
||||
|
||||
- 用户目标:ISS-011 阶段 1 独立完成 Graph 骨架和 Fake Node 路由测试,archive/commit 后才能进入真实 Node 阶段。
|
||||
- 当前问题:阶段 0 只有设计基线,仓库此前没有可编译 StateGraph 或条件边验证。
|
||||
- 关联 OpenSpec:`openspec/changes/chat-diagnosis-stategraph-routing-skeleton/`
|
||||
- devflow 分档:complex
|
||||
- 前置基线:`openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-design-freeze/`
|
||||
|
||||
## 范围
|
||||
|
||||
- 本次要做:Graph Core 直接依赖、状态/枚举、config-aware action ports、deterministic router、Graph factory、有限计数、events/trace builder、Fake Node tests。
|
||||
- 本次不做:真实 Agent/Gatekeeper、ChatService、Hook、DB、Trace API、旧实现清理、Maven E2E。
|
||||
- 影响区域:`pom.xml`、`com.superbiz.agent.graph.diagnosis`、对应 test 包。
|
||||
|
||||
## OpenSpec 对齐
|
||||
|
||||
- proposal 覆盖状态:已覆盖 skeleton-only 目标、L2 边界、测试与 E2E 非目标。
|
||||
- specs 覆盖状态:6 个 requirements 覆盖编译、Planner、Executor/Gatekeeper、Verifier、Composer、trace。
|
||||
- tasks 覆盖状态:12/12 完成。
|
||||
@@ -0,0 +1,112 @@
|
||||
# Chat Diagnosis StateGraph Routing Skeleton Decisions
|
||||
|
||||
## Question Pool
|
||||
|
||||
| # | 维度 | 问题 | 模式 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | 术语 | 阶段 1 的 Graph State、event、trace 含义从哪里继承? | evidence-driven | 已解决 |
|
||||
| Q2 | 边界 | 阶段 1 是否接入 ChatService 或真实 Agent/Service? | evidence-driven | 已解决 |
|
||||
| Q3 | 技术 | 锁定 1.1.2.0 是否支持所需 Node/Edge、策略和 recursion limit? | evidence-driven | 已解决 |
|
||||
| Q4 | 技术 | Node action port 是否需要 RunnableConfig? | evidence-driven | 已解决 |
|
||||
| Q5 | 验收 | 阶段 1 是否有必要添加单元测试? | evidence-driven + user rule | 已解决 |
|
||||
| Q6 | 接口 | 新骨架的接口影响等级与消费者是什么? | evidence-driven | 已解决 |
|
||||
| Q7 | 循环 | recursion limit 应取多少,是否替代业务计数? | evidence-driven | 已解决 |
|
||||
| Q8 | 阶段 | 是否在本 change 实现真实 nodes、DB、Trace API 或旧代码清理? | user-interview(已由六阶段口径确认) | 已解决 |
|
||||
|
||||
## Evidence-driven
|
||||
|
||||
| 结论 | 证据来源 | 是否已汇报用户 |
|
||||
|---|---|---|
|
||||
| 阶段 1 必须完整继承阶段 0 baseline | archived design/spec/ADR | 已汇报 |
|
||||
| 本阶段只做 Fake Node 骨架,不接真实模型 | ISS-011 阶段 1/2 边界 | 已汇报 |
|
||||
| 1.1.2.0 支持 AsyncNodeActionWithConfig、AsyncEdgeActionWithConfig、conditional edges、Replace/Append、recursionLimit 和 threadId | 本地 JAR `javap` / `javap -c` | 已汇报 |
|
||||
| AppendStrategy 将 list/collection 追加为有序列表,可用于 node terminal events | 本地 AppendStrategy bytecode | 已汇报 |
|
||||
| 路由是新增行为且分支多,单元测试有必要 | 阶段 1 完成标准与用户“必要则加”规则 | 已汇报 |
|
||||
| 最坏合法路径少于 20 次 Node 执行,limit 32 有安全余量 | 冻结路由矩阵的路径计数 | 已汇报 |
|
||||
| 当前仓库没有 Graph 包或实现 | `rg` / package 目录核查 | 已汇报 |
|
||||
| 生产代码将直接 import Graph Core,必须从传递依赖提升为 BOM 管理的直接依赖 | pom 与 dependency tree / 架构审计 | 已汇报 |
|
||||
|
||||
## User-interview
|
||||
|
||||
| 问题原文 | 用户原话 | 确认状态 | OpenSpec 回写 |
|
||||
|---|---|---|---|
|
||||
| 阶段 1 是否应独立执行 sm-flow? | “iss-011里每个阶段,都是一个sm-flow” | 已确认 | 本 change 独立边界 |
|
||||
| 阶段 1 是否执行 E2E? | “端到端只在最后阶段全部完成后才验证” | 已确认 | Out of Scope / Acceptance |
|
||||
| 阶段 1 是否添加单元测试? | “如果有必要添加单元测试验收的话,就加” | 已确认规则;本阶段判定必要 | Acceptance |
|
||||
| 是否可提前实现阶段 2–5? | “每个阶段需要归档完并提交才能进入下一个阶段” | 已确认不可提前 | Out of Scope |
|
||||
|
||||
## Context And Handoff
|
||||
|
||||
- 前置 archive:`openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-design-freeze/`
|
||||
- 前置 commit:`581daff`
|
||||
- 当前 change:`chat-diagnosis-stategraph-routing-skeleton`
|
||||
- 后续 change:`chat-diagnosis-stategraph-real-nodes`,只能在本阶段 archive + commit 后创建。
|
||||
|
||||
## Technical Decisions
|
||||
|
||||
### Package and API
|
||||
|
||||
- 新包:`com.superbiz.agent.graph.diagnosis`。
|
||||
- `pom.xml` 显式声明 Graph Core,版本继续由现有 BOM 管理。
|
||||
- Graph 工厂接收 config-aware Node action ports,不依赖 Spring Bean 或真实 Agent。
|
||||
- Fake actions 只放 `src/test`。
|
||||
- 状态读取集中在 typed helper/router,避免各 edge 复制字符串解析。
|
||||
|
||||
### Retry ownership
|
||||
|
||||
- Planner/Verifier/Composer Node wrapper 根据进入节点前的上次技术失败状态增加各自 retry count。
|
||||
- 第一次失败时 count=0,允许 self-loop;重入后 count=1,第二次失败直接 Fallback。
|
||||
- Evidence Retry Node 将 run-level evidence count 增加一次并重置 planner retry count。
|
||||
- Edge 只选择 route,不隐式修改状态。
|
||||
|
||||
### Event and trace
|
||||
|
||||
- 每个 Fake/后续真实 Node 返回一个 event list,AppendStrategy 负责累积。
|
||||
- event 包含 node/outcome/reasonCode/attempt,不包含 payload。
|
||||
- trace builder 由相邻 events 生成 transition;final node/reason 来自最后 event。
|
||||
- events 为空时拒绝构造 trace,避免伪造路径。
|
||||
|
||||
### Interface impact
|
||||
|
||||
- 等级:L2 internal interface。
|
||||
- 新消费者:阶段 2 Node adapters、阶段 3 ChatService orchestrator、阶段 4 tests。
|
||||
- 外部 API/DB/运行路径:无变化。
|
||||
- 回滚:revert 本阶段提交即可;因为未接生产入口,没有数据迁移。
|
||||
- 兼容:后续 action adapters 必须实现已冻结 ports,不得传递父 State 全量数据。
|
||||
|
||||
## Risks Accepted
|
||||
- Stage 1 skeleton 会暂时存在但未被生产调用,这是阶段边界要求,不是死代码最终状态。
|
||||
- 真实 Agent 状态映射尚未验证,由阶段 2 独立 change 负责。
|
||||
- Maven E2E 不运行,生产路径完全未改变。
|
||||
|
||||
## Apply Evidence
|
||||
|
||||
- Task 1.1:显式声明 BOM 管理的 Graph Core;新增 26 个 state keys、状态枚举、拓扑/route/reason 常量、默认 Replace + events Append 策略和安全 typed reads。
|
||||
- Task 1.2:新增不可变 event/transition/trace records,构造期拒绝空 routing metadata 和非法计数,map 输出只有冻结字段。
|
||||
- Task 2.1:新增八个 non-null config-aware action ports,Fake/真实 Node 共用同一 Graph 接入面。
|
||||
- Task 2.2:新增纯 router;所有未知/缺失状态 fail closed,LOW_CONFID guard 只读 facts_checked,不构造 retry context。
|
||||
- Task 2.3:Graph Factory 注册八节点和全部条件边;wrapper 独占 retry/evidence control state,compile recursion limit=32。
|
||||
- 首模块 Maven compile:passed(31s),锁定 1.1.2.0 API 假设成立。
|
||||
- Task 3.1:trace builder 从 event 单向派生 transitions/final reason/degraded/evidence count,空或异类 events 显式失败。
|
||||
- Task 4.1:新增严格 FIFO Fake Node fixture,记录 sequence/calls/threadId,每 attempt 只追加一个 terminal event,意外调用立即失败。
|
||||
- Task 4.2:真实 CompiledGraph normal/Planner/Executor/Gatekeeper focused tests 首轮通过(Maven exit 0,约 67s)。
|
||||
- Task 4.3:Verifier/evidence/Composer 路由与独立计数测试通过(Maven exit 0,约 12s)。
|
||||
- Task 4.4:routing + trace focused suite 通过(Maven exit 0,约 28s),覆盖事件顺序、degraded、map 白名单、不可变性和非法输入。
|
||||
- Task 5.1:35 tests(29 routing + 6 trace)全通过;Maven test compilation、change strict 1/1、主 specs 12/12、diff check 均通过。
|
||||
- Task 5.2:现有 production refs=0,新 Graph 对真实 Service/DB/Trace refs=0,23 个 changed paths 全部命中阶段白名单;Maven E2E/log/DB 按用户口径保留到阶段 5。
|
||||
- Review:补齐所有正常节点 threadId 传播、Verifier REJECT→Composer 和 completed-without-verdict fail-closed 用例;增强后 focused suite exit 0。
|
||||
|
||||
## Cross-Artifact 对齐检查
|
||||
|
||||
| 上游 → 下游 | 检查内容 | 状态 |
|
||||
|---|---|---|
|
||||
| brief/prd → proposal | 阶段 1 目标、Fake Node 边界、测试口径和 E2E 非目标 | 已对齐 |
|
||||
| proposal → 设计产物 | direct dependency、状态、ports、router、counter、trace、测试架构 | 已对齐 |
|
||||
| 设计产物 → specs/tasks | 所有可观察路由、终止、安全默认和实现模块 | 已对齐 |
|
||||
| specs → tasks | 编译、全路由、trace、focused tests、生产隔离检查 | 已对齐 |
|
||||
|
||||
Gap:无。
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
输入是 run-scoped 初始 state 与 `RunnableConfig`,处理链是 CompiledGraph → config-aware action ports → deterministic router/counters,输出是最终 state 与纯路由 trace;阶段 1 没有 Controller/Service/DB 消费者。Graph State 由单次 invoke 所有,events 只由 Node append,transitions 只由 builder 派生,避免双写。阶段 2 只实现 action ports,阶段 3 才将 orchestrator 交给 ChatService,因此当前未接生产入口是刻意生命周期边界。审计发现的唯一缺口是 Graph Core 直接依赖所有权,已回写 proposal/design/tasks。与阶段 0 archive 无冲突,L2 风险可由 Fake Node CompiledGraph tests 和 revert 单提交控制。
|
||||
@@ -0,0 +1,31 @@
|
||||
# Chat Diagnosis StateGraph Routing Skeleton Evidence
|
||||
|
||||
## 证据
|
||||
|
||||
| 来源 | 证据 | 结论 | 是否已汇报 |
|
||||
|---|---|---|---|
|
||||
| 阶段 0 archive/ADR | 冻结 26 个 state keys、完整路由、四类计数、run audit | 阶段 1 实现未偏离基线 | 是 |
|
||||
| 本地 Graph Core 1.1.2.0 `javap` | config-aware Node/Edge、conditional edge、recursion limit、threadId 存在 | 使用公共锁定 API 可编译 | 是 |
|
||||
| AppendStrategy bytecode | list/collection 按顺序追加 | 每 Node 返回单 event list 可形成实际路径 | 是 |
|
||||
| Maven compile | `mvn -q "-DskipTests" compile` exit 0 | direct dependency 和 Factory API 编译成立 | 是 |
|
||||
| Fake Node CompiledGraph tests | 29 routing tests,0 failure/error | 全条件边、retry、threadId、unknown fail-closed 成立 | 是 |
|
||||
| Trace tests | 6 tests,0 failure/error | transitions、degraded、map 白名单、不可变和非法输入成立 | 是 |
|
||||
| Test compilation | `mvn -q "-DskipTests" test` exit 0 | 全测试源可编译 | 是 |
|
||||
| OpenSpec validation | change 1/1、主 specs 12/12 strict | artifacts 与既有规格无回归 | 是 |
|
||||
| 生产隔离检查 | outside Graph refs=0,Graph 对真实 Service/DB/Trace refs=0 | 阶段 1 未接生产入口 | 是 |
|
||||
| Git 路径白名单 | 23 changed paths,out-of-scope=0 | 无跨阶段文件混入 | 是 |
|
||||
|
||||
## Evidence-driven 结论
|
||||
|
||||
- 结论:直接声明 Graph Core 是正确依赖所有权。
|
||||
- 证据:生产代码直接 import Graph Core;依赖原先仅由 Agent Framework 传递。
|
||||
- 风险:BOM 升级仍需重新跑真实 Graph tests。
|
||||
- 用户确认:不需要,属于构建稳健性。
|
||||
- 结论:recursion limit 32 足够且没有替代业务计数。
|
||||
- 证据:最坏合法路径低于 20;第二次技术失败和第二次 LOW_CONFID tests 均终止。
|
||||
- 风险:未来新增循环必须重算。
|
||||
- 用户确认:不需要,冻结业务上限未变。
|
||||
- 结论:单元测试必要,E2E 不必要。
|
||||
- 证据:阶段新增条件边行为但未接生产入口;35 tests 直接验证 Graph。
|
||||
- 风险:真实 Agent 映射仍留给阶段 2。
|
||||
- 用户确认:符合用户按必要性和最终阶段 E2E 规则。
|
||||
@@ -0,0 +1,50 @@
|
||||
# Chat Diagnosis StateGraph Test Suite 验收
|
||||
|
||||
## 结果
|
||||
|
||||
已接受。OpenSpec tasks 21/21 完成,阶段 4为 test-only,无生产行为变更。
|
||||
|
||||
## 验证
|
||||
|
||||
### 静态验证
|
||||
|
||||
- `git diff --check`:通过。
|
||||
- `git diff --name-only HEAD -- src/main`:0 个文件。
|
||||
- test inventory:三个权威类均存在;`ChatServiceSequentialAgentTest`/`VerifierInputHookTest` 均不存在;`ScriptedDiagnosisGraphActions` 定义 1 处。
|
||||
- `openspec validate chat-diagnosis-stategraph-test-suite --strict`:通过。
|
||||
- `openspec validate --specs --strict`:15 passed,0 failed。
|
||||
|
||||
### 脚本验证
|
||||
|
||||
- 三层权威 suite:4 suites / 43 tests,0 failures/errors/skipped。
|
||||
- 完整保留安全回归:31 suites / 126 tests,0 failures/errors/skipped。
|
||||
- `mvn -q -DskipTests test-compile`:通过。
|
||||
|
||||
### 浏览器/人工验证
|
||||
|
||||
- 未运行;阶段 4仅重构自动化测试,且用户要求最终人工/live 验收到阶段 5统一执行。
|
||||
|
||||
### 未验证
|
||||
|
||||
- 未使用 Maven 启动应用做 live E2E。
|
||||
- 未检查 `logs/`。
|
||||
- 未执行 `scripts/query_mysql.py`。
|
||||
- 剩余风险:真实模型、工具、Flyway/MySQL 和最终 Trace 内容仍需阶段 5 E2E/log/DB 证据。
|
||||
|
||||
## 已完成范围
|
||||
|
||||
- 建立 Workflow、Node Contract、Chat Integration 三层权威测试体系。
|
||||
- 补齐 ceiling/second LOW_CONFID、tool-failure legal snapshot、partial-pass REJECT 和 Verifier invalid status 边界。
|
||||
- 删除旧 Hook implementation test,并保留 parser/Gatekeeper/projection/Composer 等安全回归。
|
||||
- 用结构测试和 source gate 防止旧 Sequential/Hook 测试回归。
|
||||
|
||||
## 已知限制
|
||||
|
||||
- 生产 `VerifierInputHook`/`VerifierContextHolder` 类型仍存在但已无生产/权威测试消费者;阶段 5清理。
|
||||
- component tests 与权威层存在有意的分层 overlap,详见 `evidence.md`。
|
||||
|
||||
## 交接
|
||||
|
||||
- OpenSpec archive:已同步 2 份 delta specs,并归档到 `openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-test-suite/`。
|
||||
- 下一步:审查精确 Git diff并完成阶段 4独立提交,之后启动阶段 5。
|
||||
- OpenSpec 归档确认:用户已授权直接执行后续归档;归档已完成。
|
||||
@@ -0,0 +1,20 @@
|
||||
# Chat Diagnosis StateGraph Test Suite Brief
|
||||
|
||||
## 背景
|
||||
|
||||
- 用户目标:将 ISS-011 阶段 4作为独立 sm-flow,建立以 Graph 路径和外部行为为中心的新测试体系。
|
||||
- 当前问题:阶段 1–3 覆盖充分但组织分散,旧 Hook payload test 仍会约束已退出生产的隐式状态机。
|
||||
- 关联 OpenSpec:`openspec/changes/chat-diagnosis-stategraph-test-suite/`
|
||||
- devflow 分档:complex;接口影响 L1 test-only。
|
||||
|
||||
## 范围
|
||||
|
||||
- 本次要做:Workflow/Node Contract/Chat Integration 三层权威入口、Issue 路径矩阵补齐、旧 Hook test 退役、保留安全回归。
|
||||
- 本次不做:不改生产源码,不删除生产 Hook/ThreadLocal 类型,不运行 live E2E/log/DB 验收。
|
||||
- 影响区域:Diagnosis Graph tests、ChatService integration test、OpenSpec/devflow。
|
||||
|
||||
## OpenSpec 对齐
|
||||
|
||||
- proposal/design/specs:已覆盖,2 个 delta capabilities。
|
||||
- tasks:21/21 已完成。
|
||||
- 生产行为:无变化,`src/main` diff=0。
|
||||
@@ -0,0 +1,121 @@
|
||||
# Chat Diagnosis StateGraph Test Suite Decisions
|
||||
|
||||
## Entry Summary
|
||||
|
||||
- 问题:阶段 1–3 的安全覆盖已充足,但测试命名/组织仍是逐步实现产物,尚未形成 ISS-011 指定的 Workflow、Node Contract、Chat Integration 三层权威体系。
|
||||
- 期望:阶段 4只重构测试结构并补齐矩阵,不改变生产行为;归档并提交后才进入阶段 5。
|
||||
- 分档:complex(路径矩阵广、涉及旧安全测试退役,但接口影响为 L1 test-only)。
|
||||
- Change:`chat-diagnosis-stategraph-test-suite`。
|
||||
- 授权:用户已要求直接实现,阶段门禁与阶段 5 才 live E2E 的约束不变。
|
||||
|
||||
## Context Sources
|
||||
|
||||
- ISS-011 阶段 4、测试策略和验收标准。
|
||||
- `chat-diagnosis-stategraph-design-freeze` 的 test migration requirement。
|
||||
- 阶段 1–3 archive/acceptance 和当前 119-test focused baseline。
|
||||
- `DiagnosisGraphRoutingTest`、各 Node tests、`DiagnosisRealGraphIntegrationTest`、`ChatDiagnosisGraphRuntimeTest`、`ChatServiceGraphIntegrationTest`。
|
||||
- `VerifierInputHookTest`、protocol parser tests、`ExecutorGatekeeperServiceTest` 与 Trace/Controller/Repository/Eval tests。
|
||||
|
||||
## Question Pool
|
||||
|
||||
| # | 维度 | 问题 | 模式 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | 术语 | Workflow、Node Contract、Chat Integration 的职责边界是什么? | evidence-driven | 已解决 |
|
||||
| Q2 | 边界 | 是否把所有现有 Graph unit tests 合并成三个巨型类? | evidence-driven | 已解决 |
|
||||
| Q3 | 旧测试 | `VerifierInputHookTest` 在显式 Graph Gatekeeper 后应保留、改写还是删除? | evidence-driven | 已解决 |
|
||||
| Q4 | 验收 | 如何证明阶段 4矩阵完整且没有恢复固定 Sequential 顺序? | evidence-driven | 已解决 |
|
||||
| Q5 | 阶段 | 是否允许为测试可测性改生产代码或执行 live E2E? | user-interview(既有冻结规则) | 已确认 |
|
||||
|
||||
## Evidence-driven Findings
|
||||
|
||||
- Q1:Issue 已明确三类测试;现有 `DiagnosisGraphRoutingTest` 对应 Workflow,分散 Node/Protocol tests 对应 Node Contract,`ChatServiceGraphIntegrationTest` 对应外部生命周期。
|
||||
- Q2:现有细粒度测试失败定位清晰,全部合并会制造大文件;应保留专用 unit tests,同时新增/重命名三层权威入口并共享夹具。
|
||||
- Q3:生产 Graph Verifier 已不注册 Hook,Hook test 仍验证 raw/full-trace/ThreadLocal payload,与 verified-only 生产协议冲突;parser/Gatekeeper/投影行为已有独立 tests,故阶段 4删除 Hook test,生产类型留到阶段 5。
|
||||
- Q4:以 Issue 必需路径清单建立 requirement-to-test matrix;source check 禁止 `SequentialAgent`/`VerifierInputHookTest` 成为新 suite 依赖,并运行保留安全回归。
|
||||
|
||||
## User-interview Confirmation
|
||||
|
||||
| 问题 | 用户原话/既有确认 | 状态 | OpenSpec 回写 |
|
||||
|---|---|---|---|
|
||||
| Q5 阶段边界 | “端到端只在最后阶段全部完成后才验证;每个阶段如果有必要添加单元测试验收的话,就加” | 已确认 | proposal |
|
||||
|
||||
## Grill-with-docs Result
|
||||
|
||||
- 术语不进入业务 glossary:Workflow/Node Contract/Integration 是测试架构术语,不改变 Session、Run、Trace、Gatekeeper 或 evidence gap 领域定义。
|
||||
- 具体场景压力测试:同 session 多 run 属于 Chat Integration;Gatekeeper REJECT/LOW_CONFID 与 retry exhaustion 属于 Workflow;failed binding 过滤与 verified-only payload 属于 Node Contract。
|
||||
- 旧 Hook test 的有效行为已分别迁移到 `ExecutorEvidenceParserTest`、`ExecutorGatekeeperServiceTest`、`GatekeeperNodeTest` 和 `VerifiedInputNodeTest`;删除不会丢失安全真理源。
|
||||
- 没有难以逆转的新架构取舍,不创建 ADR;测试组织可以在保持行为矩阵的前提下继续演进。
|
||||
|
||||
## Discover Status
|
||||
|
||||
- `devflow/index.md`:命中阶段 0–3 archive。
|
||||
- 生产调用链:阶段 3 已冻结且测试阶段默认不修改。
|
||||
- 接口影响:L1 test-only;无 API/DTO/DB/Prompt/运行时消费者变化。
|
||||
- 未解决问题:0。
|
||||
- Draft 产物:proposal + decisions;尚未生成 design/spec/tasks,尚未修改测试代码。
|
||||
|
||||
## Architecture Audit
|
||||
|
||||
### Test ownership map
|
||||
|
||||
`ISS-011 path matrix -> DiagnosisGraphWorkflowTest -> ScriptedDiagnosisGraphActions -> real DiagnosisGraphFactory` 负责控制流;`Node/protocol contracts -> DiagnosisGraphNodeContractTest + focused component tests -> real Node actions/parsers/Gatekeeper` 负责安全投影;`public lifecycle -> ChatServiceGraphIntegrationTest -> Run repositories/Eval/Trace mapping` 负责外部行为。Controller、Repository、Trace、Eval 和 protocol tests 是三层体系的下游安全消费者,不应被重写成 Graph 内部顺序断言。
|
||||
|
||||
### Data and lifecycle ownership
|
||||
|
||||
- Workflow fixtures 只拥有脚本状态、调用计数和 events,不创建 Session/Run。
|
||||
- Node Contract fixtures 只拥有 invoker input/output 和 mock Gatekeeper current-run result,不持久化生产实体。
|
||||
- Chat Integration fixtures 只观察 ChatService public result 和 current Run persistence,不推断内部 Node 次序。
|
||||
- `VerifierInputHookTest` 删除后不产生数据契约缺口:Executor parser、Gatekeeper、passed-binding projection 各自已有单一测试所有者。
|
||||
|
||||
### Coupling risks
|
||||
|
||||
- 最大风险是 Workflow 与 Node Contract 都断言完整路径而重复;设计将 Fake route matrix 与 real-node input/security matrix分开。
|
||||
- `ScriptedDiagnosisGraphActions` 是唯一 Fake topology fixture;不新增第二套 Graph builder。
|
||||
- 测试-only阶段禁止 `src/main` diff,避免为测试便利扩大 production API。
|
||||
- 类重命名使用 Git rename,旧名称仅允许出现在 OpenSpec/devflow迁移说明中。
|
||||
|
||||
### Cross-artifact alignment
|
||||
|
||||
| 上游 → 下游 | 检查内容 | 状态 |
|
||||
|---|---|---|
|
||||
| brief/proposal → proposal | 三层体系、旧 Hook test退役、保留安全回归、阶段 5 E2E 延期 | 已对齐 |
|
||||
| proposal → design | rename-not-copy、职责归属、production diff=0、rollback | 已对齐 |
|
||||
| design → specs/tasks | Workflow/Node/Integration矩阵、Hook test删除、source inventory和验证门禁 | 已对齐 |
|
||||
| specs → tasks | 每类 scenario 均有 rename、补缺、回归或静态验收任务 | 已对齐 |
|
||||
|
||||
### Audit result
|
||||
|
||||
架构审计未发现业务 glossary、阶段 0–3 specs 或生产行为冲突。新 capability 只描述测试验证系统,modified design-freeze requirement 完整保留并细化旧测试替换边界。接口影响保持 L1,cross-artifact gap=0,无需回写生产 spec 或创建 ADR。
|
||||
|
||||
## Commit Gate
|
||||
|
||||
- schema:spec-driven;proposal/design/2 delta specs/tasks 全部 done,applyRequires=`tasks` 已满足。
|
||||
- OpenSpec:当前 change strict pass;15 个主 specs strict pass。
|
||||
- Cross-artifact:4/4 已对齐,gap=0。
|
||||
- Question pool:4 个 evidence-driven 已查证,1 个 user-interview 由用户既有原话确认,无未决项。
|
||||
- Interface impact:L1 test-only;design 有独立影响/回滚章节,默认 `src/main` diff=0。
|
||||
- Preflight:`git diff --check` 通过,当前仅 Draft OpenSpec/devflow 与根目录计划文件,无测试/生产代码修改。
|
||||
- 结论:Draft OpenSpec 达到可执行状态,创建 `.committed` 后进入 Apply。
|
||||
|
||||
## Apply Completion
|
||||
|
||||
- `DiagnosisGraphRoutingTest` 已以 Git rename 演进为 `DiagnosisGraphWorkflowTest`;所有 scripted run 统一断言 events 与真实 sequence 完全一致。
|
||||
- `DiagnosisRealGraphIntegrationTest` 已演进为 `DiagnosisGraphNodeContractTest`;补齐合法工具失败限制、partial-pass REJECT 和 Verifier invalid 无伪 verdict。
|
||||
- Workflow 显式断言 Gatekeeper ceiling 将模型 PASS 限制为 LOW_CONFID 且不触发 evidence retry,以及第二次 LOW_CONFID 不再补证据。
|
||||
- `ChatServiceGraphIntegrationTest` 增加 SessionContextHolder finally cleanup 断言;public Run/Trace/Eval/multi-run 覆盖保持。
|
||||
- `VerifierInputHookTest` 已删除;生产 Hook/ThreadLocal 类型未修改,留待阶段 5清理。
|
||||
- 新增 `DiagnosisGraphTestSuiteStructureTest`,保证三个权威类存在、Sequential/Hook实现测试不存在且权威测试不引用旧实现。
|
||||
|
||||
## Verification Summary
|
||||
|
||||
- 新权威层:4 suites / 43 tests,0 failures/errors/skipped。
|
||||
- 完整保留安全回归:31 suites / 126 tests,0 failures/errors/skipped。
|
||||
- Maven test compilation:通过。
|
||||
- OpenSpec:当前 change strict pass;主 specs 15/15 strict pass。
|
||||
- 静态门禁:`git diff --check` 通过;required authoritative classes=3;legacy tests=0;scripted fixture definitions=1;`src/main` diff=0。
|
||||
- 按阶段门禁未运行 Maven live E2E、未检查 `logs/`、未执行 `scripts/query_mysql.py`;统一保留到阶段 5。
|
||||
|
||||
## Archive Result
|
||||
|
||||
- 2 份 delta specs 已同步:新增 test-suite capability 6 条 requirements,修改 design-freeze test migration requirement 1 条。
|
||||
- OpenSpec 已归档到 `openspec/changes/archive/2026-07-17-chat-diagnosis-stategraph-test-suite/`。
|
||||
@@ -0,0 +1,42 @@
|
||||
# Chat Diagnosis StateGraph Test Suite Evidence
|
||||
|
||||
## Requirement-to-test Matrix
|
||||
|
||||
| ISS-011 路径/边界 | 权威测试 | 辅助证据 |
|
||||
|---|---|---|
|
||||
| PASS 与精确 event/transition | `DiagnosisGraphWorkflowTest.normalPassPathUsesCompiledGraphAndPreservesEventOrder` | `DiagnosisGraphNodeContractTest.compiledRealNodeGraphCompletesPassPathInExactOrder` |
|
||||
| Planner 一次技术重试/耗尽/非重试失败 | `DiagnosisGraphWorkflowTest.planner*` | Planner adapter focused test |
|
||||
| Executor FAILED/TOOL_BLOCKED/INVALID/no-evidence | `DiagnosisGraphWorkflowTest.executor*` | `DiagnosisGraphNodeContractTest.legalSnapshotAfterToolFailureRemainsCompletedAndReachesGatekeeper` |
|
||||
| Gatekeeper PASS/LOW/REJECT/unknown | `DiagnosisGraphWorkflowTest.gatekeeper*` | Gatekeeper Node/service tests |
|
||||
| partial-pass REJECT 不泄漏 claim | Workflow unsafe route | `DiagnosisGraphNodeContractTest.gatekeeperRejectSkipsVerifierAndUsesPreVerificationFallback` |
|
||||
| verified-only passed-binding projection | Workflow VerifiedInput route | `VerifiedInputNodeTest` + Verifier adapter test |
|
||||
| Verifier 技术重试/耗尽/非重试失败 | `DiagnosisGraphWorkflowTest.verifier*` | `DiagnosisGraphNodeContractTest.verifierInvalidOutputSetsExecutionStatusWithoutFabricatedVerdict` |
|
||||
| critical evidence retry、无 gap、non-critical、ceiling、第二次 LOW | `DiagnosisGraphWorkflowTest.*LowConfidence*` / `secondLowConfidence*` | EvidenceRetryPrepareNodeTest + real full-snapshot integration |
|
||||
| 第二轮完整 snapshot 再验真 | Workflow evidence retry | `DiagnosisGraphNodeContractTest.criticalGapPerformsOneIncrementalRoundAndRevalidatesCompleteSnapshot` |
|
||||
| Verifier REJECT 安全表达 | `DiagnosisGraphWorkflowTest.verifierRejectStillRunsComposerWithSafeMaterial` | ComposerSafeInputBuilderTest |
|
||||
| Composer retry/耗尽/非重试/安全 Fallback | `DiagnosisGraphWorkflowTest.composer*` | ComposerNodeAdapterTest + FallbackNodeTest |
|
||||
| ChatResult/Run/Trace/evaluation/cleanup | `ChatServiceGraphIntegrationTest` | Controller/Trace/Repository/Eval tests |
|
||||
| 同 session 多 run 隔离 | `ChatServiceGraphIntegrationTest.sameSessionCreatesDistinctRunIdsAndKeepsTracePerRun` | Run repository/Trace exact-run tests |
|
||||
| 三层结构与旧实现测试退役 | `DiagnosisGraphTestSuiteStructureTest` | source inventory command |
|
||||
|
||||
## 旧 Hook test 安全映射
|
||||
|
||||
| 旧行为 | 新真理源 |
|
||||
|---|---|
|
||||
| fenced/prefixed/malformed Executor JSON | `ExecutorEvidenceParserTest` |
|
||||
| invocation/raw_path/evidence excerpt真实性 | `ExecutorGatekeeperServiceTest` |
|
||||
| Gatekeeper status/severity/audit | `GatekeeperNodeTest` |
|
||||
| passed binding 精确投影 | `VerifiedInputNodeTest` |
|
||||
| Verifier verified-only payload | `VerifierNodeAdapterTest` + `ChatVerifierPromptContractTest` |
|
||||
|
||||
## 验证证据
|
||||
|
||||
- 31 suites / 126 tests:0 failures、0 errors、0 skipped。
|
||||
- test compilation、当前 change strict、主 specs 15/15、diff check 全通过。
|
||||
- `src/main` diff=0;三个权威类存在;旧 Sequential/Hook tests不存在;Scripted fixture 定义唯一。
|
||||
|
||||
## Intentional overlap and limits
|
||||
|
||||
- Workflow 与 Node Contract 都覆盖 PASS/REJECT,但前者验证 route/event,后者验证真实输入/安全材料;这是分层证据,不是复制 fixture。
|
||||
- 细粒度 component tests 保留以定位失败,不要求全部搬进三个权威类。
|
||||
- 真实模型/工具、日志和数据库行为未在阶段 4验证,统一由阶段 5 live E2E承担。
|
||||
+192
@@ -0,0 +1,192 @@
|
||||
<!doctype html>
|
||||
<html lang="zh-CN">
|
||||
<head>
|
||||
<meta charset="utf-8">
|
||||
<meta name="viewport" content="width=device-width, initial-scale=1">
|
||||
<title>Spring AI Alibaba Graph:诊断编排改造</title>
|
||||
<script src="https://cdn.jsdelivr.net/npm/mermaid/dist/mermaid.min.js"></script>
|
||||
<style>
|
||||
:root { --bg:#f4f7fb; --card:rgba(255,255,255,.86); --text:#172033; --muted:#5c6880; --blue:#2563eb; --line:#dbe3f0; }
|
||||
* { box-sizing:border-box; }
|
||||
body { margin:0; font-family:Inter,"PingFang SC","Microsoft YaHei",sans-serif; background:linear-gradient(135deg,#eef4ff,#f8fafc); color:var(--text); line-height:1.75; }
|
||||
.wrap { max-width:1000px; margin:0 auto; padding:42px 24px 110px; }
|
||||
header,.card { background:var(--card); border:1px solid rgba(255,255,255,.9); box-shadow:0 12px 35px rgba(35,55,90,.09); backdrop-filter:blur(14px); border-radius:20px; padding:28px; margin-bottom:22px; }
|
||||
h1 { margin:0 0 8px; font-size:34px; }
|
||||
h2 { margin-top:0; color:#173c85; }
|
||||
h3 { color:#244c92; }
|
||||
code,pre { font-family:"Cascadia Code",Consolas,monospace; }
|
||||
pre { background:#101827; color:#e5edf9; padding:18px; border-radius:14px; overflow:auto; }
|
||||
.tag { display:inline-block; padding:4px 10px; margin-right:6px; border-radius:999px; background:#e4edff; color:#2453a6; font-size:13px; }
|
||||
.mnemonic-card { background:#fff8cf; border:2px dashed #e6b800; padding:16px; border-radius:14px; }
|
||||
.fission-section { background:#fff1f2; border-left:5px solid #e11d48; padding:18px; border-radius:12px; }
|
||||
.truth { border-left:4px solid #2563eb; background:#eff6ff; padding:14px; border-radius:10px; }
|
||||
details { background:#f8fafc; border:1px solid var(--line); border-radius:12px; padding:12px 15px; margin:9px 0; }
|
||||
summary { cursor:pointer; font-weight:700; }
|
||||
.search { position:fixed; bottom:22px; left:50%; transform:translateX(-50%); width:min(720px,calc(100% - 36px)); background:rgba(15,23,42,.93); padding:12px; border-radius:16px; box-shadow:0 15px 40px rgba(0,0,0,.25); z-index:5; }
|
||||
.search input { width:100%; border:0; outline:0; border-radius:10px; padding:12px 14px; font-size:15px; }
|
||||
.hidden { display:none !important; }
|
||||
ul,ol { padding-left:24px; }
|
||||
</style>
|
||||
</head>
|
||||
<body>
|
||||
<div class="wrap" id="content-area">
|
||||
<header>
|
||||
<h1>Spring AI Alibaba Graph:诊断编排改造</h1>
|
||||
<p>作者:叫我小杨同学的小码酱</p>
|
||||
<span class="tag">StateGraph</span><span class="tag">Agent 编排</span><span class="tag">条件边</span><span class="tag">故障诊断</span>
|
||||
</header>
|
||||
|
||||
<section class="card">
|
||||
<h2>0. 核心摘要</h2>
|
||||
<p><strong>让 ReactAgent 继续负责做事,让 StateGraph 负责下一步去哪里。</strong></p>
|
||||
<p>生活类比:Planner、Executor、Verifier 是医院科室,Graph 是分诊和转诊制度。</p>
|
||||
<p class="truth">官方核心模型是 State、Nodes、Edges。本地依赖 1.1.2.0 已确认支持 StateGraph、条件边、编译配置、中断和 threadId。</p>
|
||||
</section>
|
||||
|
||||
<section class="card">
|
||||
<h2>1. 概念破冰</h2>
|
||||
<div class="mnemonic-card">状态记事实,节点做任务,边管下一步,检查点管恢复。</div>
|
||||
<p>当前 ChatService 已经是半个状态机:SequentialAgent 运行 Planner、Executor、Verifier,外层 Java 再根据 PASS、LOW_CONFID、REJECT 决定 Composer 或重试。Graph 改造的价值,是把分散的控制权显式化。</p>
|
||||
<pre>当前:SequentialAgent + 外层 if/else
|
||||
目标:StateGraph 条件边 + ReactAgent 语义节点</pre>
|
||||
</section>
|
||||
|
||||
<section class="card">
|
||||
<h2>2. 深度解析</h2>
|
||||
<p>Supervisor 适合动态选择专科 Agent;Gatekeeper、Verifier 等强制门禁应由代码边控制。普通工具失败也不需要 interrupt,只有等待人工输入或审批时才暂停。</p>
|
||||
<div class="mermaid">
|
||||
flowchart TD
|
||||
S["START"] --> P["Planner"]
|
||||
P --> E["Executor"]
|
||||
E -- "结构有效" --> G["Gatekeeper"]
|
||||
E -- "阻断或非法" --> F["Fallback"]
|
||||
G -- "允许" --> V["Verifier"]
|
||||
G -- "拒绝" --> F
|
||||
V -- "通过或拒绝" --> C["Composer"]
|
||||
V -- "低置信且有预算" --> R["Retry Guard"]
|
||||
R -- "补证据" --> P
|
||||
R -- "停止" --> C
|
||||
C --> X["END"]
|
||||
F --> X
|
||||
</div>
|
||||
|
||||
<h3>状态设计</h3>
|
||||
<p>保存统一诊断上下文、计划、Executor 结构化输出、Gatekeeper 结果、Verifier verdict、重试轮次和最终结果。不要在 State 里复制所有原始日志或完整思考过程。</p>
|
||||
|
||||
<h3>核心伪代码</h3>
|
||||
<pre>StateGraph graph = new StateGraph("diagnosis_workflow", strategies)
|
||||
.addNode("planner", plannerNode)
|
||||
.addNode("executor", executorNode)
|
||||
.addNode("gatekeeper", gatekeeperNode)
|
||||
.addNode("verifier", verifierNode)
|
||||
.addNode("retry_guard", retryGuardNode)
|
||||
.addNode("composer", composerNode)
|
||||
.addNode("fallback", fallbackNode)
|
||||
.addEdge(START, "planner")
|
||||
.addEdge("planner", "executor")
|
||||
.addConditionalEdges("executor", routeAfterExecutor,
|
||||
Map.of("gatekeeper","gatekeeper",
|
||||
"retry_guard","retry_guard",
|
||||
"fallback","fallback"))
|
||||
.addConditionalEdges("gatekeeper", routeAfterGatekeeper,
|
||||
Map.of("verifier","verifier","fallback","fallback"))
|
||||
.addConditionalEdges("verifier", routeAfterVerifier,
|
||||
Map.of("retry_guard","retry_guard",
|
||||
"composer","composer",
|
||||
"fallback","fallback"))
|
||||
.addConditionalEdges("retry_guard", routeAfterRetry,
|
||||
Map.of("planner","planner","composer","composer"))
|
||||
.addEdge("composer", END)
|
||||
.addEdge("fallback", END);</pre>
|
||||
|
||||
<h3>运行边界</h3>
|
||||
<pre>RunnableConfig config = RunnableConfig.builder()
|
||||
.threadId(runId)
|
||||
.addMetadata("sessionId", sessionId)
|
||||
.addMetadata("runId", runId)
|
||||
.build();</pre>
|
||||
<p>一个 diagnosis_run 使用一个 Graph thread,避免同一 session 下多个 run 共享检查点。MemorySaver 不是跨重启持久化。</p>
|
||||
|
||||
<h3>ReactAgent 的接入</h3>
|
||||
<p>本地 ReactAgent 提供 asNode(boolean, boolean)。当前项目第一阶段更适合用适配节点调用已有 Agent,显式控制输入、outputKey 和解析;状态契约稳定后再评估直接 asNode。</p>
|
||||
</section>
|
||||
|
||||
<section class="card fission-section">
|
||||
<h2>3. 深度裂变</h2>
|
||||
<h3>🔍 搜索内化:改造的是控制权,不是 Agent</h3>
|
||||
<p>Graph 的节点可以是 LLM,也可以是普通 Java 代码;ReactAgent 本身已经是子图。所谓“Multi-Agent 改 Graph”,实际上是把跨 Agent 状态转换交给父 Graph。</p>
|
||||
<p>官方页面示例有 OverAllStaste 拼写错误,实际类型是 OverAllState。网站主分支可能领先于本地依赖,最终必须以项目 JAR 和编译测试为准。</p>
|
||||
</section>
|
||||
|
||||
<section class="card">
|
||||
<h2>4. 实战指南</h2>
|
||||
<ol>
|
||||
<li>先定义节点结果状态,不改 Prompt。</li>
|
||||
<li>把现有 Planner、Executor、Verifier 包装为 Node。</li>
|
||||
<li>Gatekeeper 和固定降级做成 Java Node。</li>
|
||||
<li>迁移现有两轮 LOW_CONFID 控制。</li>
|
||||
<li>加入 Executor 阻断、非法结构和 Gatekeeper REJECT 条件边。</li>
|
||||
<li>保留旧 Sequential 链路作为短期回退。</li>
|
||||
<li>最后再增加 HITL、并行和专科 SubAgent。</li>
|
||||
</ol>
|
||||
<h3>避坑</h3>
|
||||
<ul>
|
||||
<li>不要让 Supervisor 决定是否跳过安全门禁。</li>
|
||||
<li>不要把 no_evidence 当成 Executor 失败。</li>
|
||||
<li>不要让多个 run 共用 sessionId 作为 Graph threadId。</li>
|
||||
<li>不要把 MemorySaver 当成生产持久化。</li>
|
||||
<li>不要未验证 messages 传播就直接大量使用 asNode(true, true)。</li>
|
||||
</ul>
|
||||
</section>
|
||||
|
||||
<section class="card">
|
||||
<h2>5. 温故知新</h2>
|
||||
<h3>FAQ</h3>
|
||||
<details><summary>1. Graph 会替代 ReactAgent 吗?</summary><p>不会,ReactAgent 可以作为子图节点继续使用。</p></details>
|
||||
<details><summary>2. 为什么不用 Supervisor 控制失败?</summary><p>失败跳转是确定性规则,不需要增加一次模型决策。</p></details>
|
||||
<details><summary>3. no_evidence 是否直接 fallback?</summary><p>不一定,合法 no-evidence 引用仍要经过 Gatekeeper 和 Verifier。</p></details>
|
||||
<details><summary>4. Composer 必须是 Agent 吗?</summary><p>正常表达可以使用轻量 Agent,系统失败要保留固定模板。</p></details>
|
||||
<details><summary>5. threadId 用什么?</summary><p>当前数据模型下优先使用 runId,sessionId 作为元数据。</p></details>
|
||||
<details><summary>6. 何时需要 Checkpointer?</summary><p>需要暂停、恢复和检查 Graph 历史状态时。</p></details>
|
||||
<details><summary>7. 能直接使用 agent.asNode 吗?</summary><p>可以,但要验证 outputKey、messages 和父子检查点。</p></details>
|
||||
<details><summary>8. AIOps 要单独 Graph 吗?</summary><p>入口和输出策略独立,诊断核心可以共用。</p></details>
|
||||
|
||||
<h3>自测题</h3>
|
||||
<ol>
|
||||
<li>为什么 Executor 工具阻断不应由 Verifier 决定重试?</li>
|
||||
<li>ReplaceStrategy 和 AppendStrategy 各适合什么状态?</li>
|
||||
<li>为什么合法 no_evidence 仍然需要 Gatekeeper?</li>
|
||||
<li>Supervisor 与条件边的决策权有什么不同?</li>
|
||||
<li>为什么 Graph threadId 更适合使用 runId?</li>
|
||||
<li>什么情况下才应该配置 interruptBefore?</li>
|
||||
</ol>
|
||||
</section>
|
||||
</div>
|
||||
|
||||
<div class="search"><input id="search-input" placeholder="搜索本文内容……"></div>
|
||||
<script>
|
||||
mermaid.initialize({ startOnLoad: true, theme: 'neutral' });
|
||||
window.onload = function() {
|
||||
const input = document.getElementById('search-input');
|
||||
if(!input) return;
|
||||
input.addEventListener('input', (e) => {
|
||||
const term = e.target.value.toLowerCase().trim();
|
||||
const contentArea = document.getElementById('content-area');
|
||||
const blocks = contentArea.querySelectorAll('p, li, blockquote, .fission-section, .mnemonic-card, details, .mermaid');
|
||||
if(term.length === 0) {
|
||||
blocks.forEach(el => el.classList.remove('hidden'));
|
||||
document.querySelectorAll('h1, h2, h3').forEach(el => el.classList.remove('hidden'));
|
||||
return;
|
||||
}
|
||||
blocks.forEach(el => el.classList.add('hidden'));
|
||||
document.querySelectorAll('h1, h2, h3').forEach(el => el.classList.add('hidden'));
|
||||
blocks.forEach(el => {
|
||||
if(el.innerText.toLowerCase().includes(term)) {
|
||||
el.classList.remove('hidden');
|
||||
}
|
||||
});
|
||||
});
|
||||
};
|
||||
</script>
|
||||
</body>
|
||||
</html>
|
||||
+383
@@ -0,0 +1,383 @@
|
||||
# Spring AI Alibaba Graph:诊断编排改造
|
||||
|
||||
作者:叫我小杨同学的小码酱
|
||||
标签:Spring AI Alibaba、StateGraph、Agent 编排、条件边、故障诊断
|
||||
|
||||
## 0. 核心摘要
|
||||
|
||||
一句话:让 ReactAgent 继续负责“做事”,让 StateGraph 负责“下一步去哪里”。
|
||||
|
||||
生活类比:Planner、Executor、Verifier 是医院里的不同科室,Graph 是分诊和转诊制度;不能让某个科室自己决定跳过检验和会诊。
|
||||
|
||||
真理锚点:官方文档将 Graph 概括为 State、Nodes、Edges,核心关系是“节点完成工作,边决定下一步做什么”。当前项目依赖的 `spring-ai-alibaba-graph-core:1.1.2.0` 本地 JAR 已确认提供 `StateGraph.addConditionalEdges(...)`、`CompileConfig.interruptBefore/After(...)` 和 `RunnableConfig.threadId(...)`。
|
||||
|
||||
## 1. 概念破冰
|
||||
|
||||
> 巧记:状态记事实,节点做任务,边管下一步,检查点管恢复。
|
||||
|
||||
当前 `ChatService` 已经像一个半成品状态机:`SequentialAgent` 固定执行 Planner、Executor、Verifier,外层 Java 循环再判断 PASS、LOW_CONFID、REJECT,并决定 Composer 或下一轮。问题不是 Agent 不够多,而是状态转换分散在 `SequentialAgent`、Hook 和外层 `if/else` 中。
|
||||
|
||||
```text
|
||||
当前
|
||||
SequentialAgent: Planner -> Executor -> Verifier
|
||||
|
|
||||
ChatService 外层: PASS / LOW_CONFID / REJECT -> Composer / retry
|
||||
|
||||
目标
|
||||
StateGraph 显式表示所有阶段和条件边
|
||||
ReactAgent 作为图中的语义节点继续复用
|
||||
```
|
||||
|
||||
## 2. 深度解析
|
||||
|
||||
### 2.1 为什么不是直接换成 SupervisorAgent
|
||||
|
||||
SupervisorAgent 适合在多个专科 Agent 之间动态选择,例如 Database、Redis、JVM。当前诊断链路中的 Gatekeeper、Verifier 是不能随意跳过的质量门禁。如果让 LLM Supervisor 决定下一步,关键流程会从代码控制变成模型决策。
|
||||
|
||||
当前真正需要的是确定性路由:
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
S["START"] --> P["Planner"]
|
||||
P --> E["Executor"]
|
||||
E -- "证据结构有效" --> G["Gatekeeper"]
|
||||
E -- "执行阻断或非法输出" --> F["Fallback"]
|
||||
G -- "PASS 或 LOW_CONFID" --> V["Verifier"]
|
||||
G -- "REJECT" --> F
|
||||
V -- "PASS" --> C["Composer"]
|
||||
V -- "LOW_CONFID 且有预算" --> R["Retry Guard"]
|
||||
V -- "REJECT 或无预算" --> C
|
||||
R -- "允许补证据" --> P
|
||||
R -- "停止" --> C
|
||||
C --> X["END"]
|
||||
F --> X
|
||||
```
|
||||
|
||||
### 2.2 State 应保存什么
|
||||
|
||||
状态应该保存跨节点需要共享的原始事实和结构化结果,不保存拼好的 Prompt,也不保存无边界增长的模型思考过程。
|
||||
|
||||
建议的语义状态:
|
||||
|
||||
- `diagnosis_context`:入口适配后的统一诊断上下文。
|
||||
- `planner_plan`:Planner 的结构化计划。
|
||||
- `executor_output`:`executor_evidence_v2`。
|
||||
- `executor_status`:成功、阻断、非法输出或失败。
|
||||
- `gatekeeper_result`:代码验真结果。
|
||||
- `verifier_output`、`verdict`:可推导性结果。
|
||||
- `retry_context`、`round`:有限补证据状态。
|
||||
- `final_answer`:最终表达。
|
||||
- `failure_reason`:确定性的失败原因。
|
||||
|
||||
现有 Trace 已由 `agent_step` 和 `tool_invocation` 持久化,Graph State 不需要复制所有原始日志。
|
||||
|
||||
### 2.3 KeyStrategy 如何选择
|
||||
|
||||
诊断状态大多使用 `ReplaceStrategy`,因为每个阶段产生当前轮的最新结果。只有确实需要累计的轻量事件列表才使用 `AppendStrategy`。
|
||||
|
||||
```java
|
||||
KeyStrategyFactory diagnosisStateStrategies() {
|
||||
return () -> {
|
||||
Map<String, KeyStrategy> strategies = new HashMap<>();
|
||||
strategies.put("diagnosis_context", new ReplaceStrategy());
|
||||
strategies.put("planner_plan", new ReplaceStrategy());
|
||||
strategies.put("executor_output", new ReplaceStrategy());
|
||||
strategies.put("executor_status", new ReplaceStrategy());
|
||||
strategies.put("gatekeeper_result", new ReplaceStrategy());
|
||||
strategies.put("verifier_output", new ReplaceStrategy());
|
||||
strategies.put("verdict", new ReplaceStrategy());
|
||||
strategies.put("retry_context", new ReplaceStrategy());
|
||||
strategies.put("round", new ReplaceStrategy());
|
||||
strategies.put("final_answer", new ReplaceStrategy());
|
||||
strategies.put("failure_reason", new ReplaceStrategy());
|
||||
return strategies;
|
||||
};
|
||||
}
|
||||
```
|
||||
|
||||
### 2.4 ReactAgent 怎样放进 Graph
|
||||
|
||||
本地 `1.1.2.0` 的 `ReactAgent` 提供 `asNode(boolean includeContents, boolean returnReasoningContents)`,可以直接作为子图节点。但当前项目 Planner、Executor、Verifier 的输入组织和输出解析已经有较多定制,第一阶段更推荐使用适配节点显式调用现有 Agent:
|
||||
|
||||
```java
|
||||
var plannerNode = node_async((state, config) -> {
|
||||
DiagnosisContext context = requireContext(state);
|
||||
String prompt = plannerInput(context, state.value("retry_context").orElse(null));
|
||||
|
||||
AssistantMessage response = plannerAgent.call(prompt, childConfig(config, "planner"));
|
||||
PlannerPlan plan = plannerParser.parse(extractText(response));
|
||||
|
||||
return Map.of(
|
||||
"planner_plan", plan,
|
||||
"failure_reason", ""
|
||||
);
|
||||
});
|
||||
```
|
||||
|
||||
适配节点的好处是不会意外把父图全部 `messages` 注入所有 Agent,也能继续复用当前解析器、Hook、Prompt 和 ToolCallback。
|
||||
|
||||
等统一状态契约稳定后,可以评估:
|
||||
|
||||
```java
|
||||
graph.addNode("planner", plannerAgent.asNode(false, false));
|
||||
```
|
||||
|
||||
但需要先验证父子图的 `messages`、outputKey 和 Checkpointer 是否符合预期。
|
||||
|
||||
### 2.5 Executor 节点只报告状态,不决定路由
|
||||
|
||||
```java
|
||||
var executorNode = node_async((state, config) -> {
|
||||
try {
|
||||
PlannerPlan plan = requirePlan(state);
|
||||
DiagnosisContext context = requireContext(state);
|
||||
|
||||
AssistantMessage response = executorAgent.call(
|
||||
executorInput(context, plan),
|
||||
childConfig(config, "executor")
|
||||
);
|
||||
|
||||
ExecutorEvidence output = executorParser.parse(extractText(response));
|
||||
|
||||
if (!output.isStructurallyValid()) {
|
||||
return Map.of(
|
||||
"executor_status", "INVALID_OUTPUT",
|
||||
"failure_reason", "executor_evidence_v2 解析失败"
|
||||
);
|
||||
}
|
||||
|
||||
// no_evidence 仍然是合法结构,需要交给 Gatekeeper 验证真实引用。
|
||||
return Map.of(
|
||||
"executor_status", "COMPLETED",
|
||||
"executor_output", output
|
||||
);
|
||||
}
|
||||
catch (ToolCapabilityBlockedException e) {
|
||||
return Map.of(
|
||||
"executor_status", "TOOL_BLOCKED",
|
||||
"failure_reason", e.getMessage()
|
||||
);
|
||||
}
|
||||
catch (Exception e) {
|
||||
return Map.of(
|
||||
"executor_status", "FAILED",
|
||||
"failure_reason", safeMessage(e)
|
||||
);
|
||||
}
|
||||
});
|
||||
```
|
||||
|
||||
注意:`no_evidence` 不是 Executor 失败。当前项目已经用 `$.no_evidence` 表达“查询成功但无匹配证据”,它仍应进入 Gatekeeper 和 Verifier,防止被过度表达为“问题不存在”。
|
||||
|
||||
### 2.6 Gatekeeper 节点保持纯代码
|
||||
|
||||
```java
|
||||
var gatekeeperNode = node_async((state, config) -> {
|
||||
ExecutorEvidence output = requireExecutorOutput(state);
|
||||
String runId = metadata(config, "runId");
|
||||
|
||||
GatekeeperResult result = executorGatekeeperService.validate(runId, output);
|
||||
|
||||
return Map.of(
|
||||
"gatekeeper_result", result,
|
||||
"gatekeeper_status", result.severity()
|
||||
);
|
||||
});
|
||||
```
|
||||
|
||||
Gatekeeper 不需要改造成 Agent。它负责确定性引用验真,是 Graph 中的普通 Java Node。
|
||||
|
||||
### 2.7 条件边是改造核心
|
||||
|
||||
```java
|
||||
StateGraph graph = new StateGraph("diagnosis_workflow", diagnosisStateStrategies())
|
||||
.addNode("planner", plannerNode)
|
||||
.addNode("executor", executorNode)
|
||||
.addNode("gatekeeper", gatekeeperNode)
|
||||
.addNode("verifier", verifierNode)
|
||||
.addNode("retry_guard", retryGuardNode)
|
||||
.addNode("composer", composerNode)
|
||||
.addNode("fallback", fallbackNode)
|
||||
|
||||
.addEdge(START, "planner")
|
||||
.addEdge("planner", "executor")
|
||||
|
||||
.addConditionalEdges(
|
||||
"executor",
|
||||
edge_async(state -> switch (stringValue(state, "executor_status")) {
|
||||
case "COMPLETED" -> "gatekeeper";
|
||||
case "INVALID_OUTPUT" -> retryAvailable(state) ? "retry_guard" : "fallback";
|
||||
case "TOOL_BLOCKED", "FAILED" -> "fallback";
|
||||
default -> "fallback";
|
||||
}),
|
||||
Map.of(
|
||||
"gatekeeper", "gatekeeper",
|
||||
"retry_guard", "retry_guard",
|
||||
"fallback", "fallback"
|
||||
)
|
||||
)
|
||||
|
||||
.addConditionalEdges(
|
||||
"gatekeeper",
|
||||
edge_async(state -> switch (stringValue(state, "gatekeeper_status")) {
|
||||
case "REJECT" -> "fallback";
|
||||
default -> "verifier";
|
||||
}),
|
||||
Map.of("verifier", "verifier", "fallback", "fallback")
|
||||
)
|
||||
|
||||
.addConditionalEdges(
|
||||
"verifier",
|
||||
edge_async(state -> switch (stringValue(state, "verdict")) {
|
||||
case "LOW_CONFID" -> retryAvailable(state) ? "retry_guard" : "composer";
|
||||
case "PASS", "REJECT" -> "composer";
|
||||
default -> "fallback";
|
||||
}),
|
||||
Map.of(
|
||||
"retry_guard", "retry_guard",
|
||||
"composer", "composer",
|
||||
"fallback", "fallback"
|
||||
)
|
||||
)
|
||||
|
||||
.addConditionalEdges(
|
||||
"retry_guard",
|
||||
edge_async(state -> shouldRetry(state) ? "planner" : "composer"),
|
||||
Map.of("planner", "planner", "composer", "composer")
|
||||
)
|
||||
|
||||
.addEdge("composer", END)
|
||||
.addEdge("fallback", END);
|
||||
```
|
||||
|
||||
### 2.8 编译和执行
|
||||
|
||||
```java
|
||||
SaverConfig saverConfig = SaverConfig.builder()
|
||||
.register(new MemorySaver())
|
||||
.build();
|
||||
|
||||
CompileConfig compileConfig = CompileConfig.builder()
|
||||
.recursionLimit(20)
|
||||
.saverConfig(saverConfig)
|
||||
.build();
|
||||
|
||||
CompiledGraph compiledGraph = graph.compile(compileConfig);
|
||||
|
||||
RunnableConfig runConfig = RunnableConfig.builder()
|
||||
// 一个 diagnosis_run 对应一个 Graph thread,避免同一 session 下多个 run 混状态。
|
||||
.threadId(runId)
|
||||
.addMetadata("sessionId", sessionId)
|
||||
.addMetadata("runId", runId)
|
||||
.build();
|
||||
|
||||
Map<String, Object> initialState = Map.of(
|
||||
"diagnosis_context", diagnosisContext,
|
||||
"round", 1
|
||||
);
|
||||
|
||||
Optional<OverAllState> finalState = compiledGraph.invoke(initialState, runConfig);
|
||||
```
|
||||
|
||||
`MemorySaver` 只适合进程内检查点,不等价于重启后可恢复的持久化。当前项目已有数据库 Trace,可以先把 Graph 用于流程控制;只有真正需要跨进程暂停恢复时,再引入持久 Checkpointer 或显式恢复模型。
|
||||
|
||||
### 2.9 Chat 与 AIOps 怎样共用
|
||||
|
||||
入口适配不同,公共 Graph 接收统一 `DiagnosisContext`:
|
||||
|
||||
```java
|
||||
DiagnosisContext chatContext = chatAdapter.from(question, history);
|
||||
DiagnosisContext aiOpsContext = aiOpsAdapter.from(alertPayload);
|
||||
|
||||
DiagnosisResult chatResult = diagnosisGraph.execute(chatContext, sessionId, runId);
|
||||
DiagnosisResult aiOpsResult = diagnosisGraph.execute(aiOpsContext, sessionId, runId);
|
||||
|
||||
return chatOutputAdapter.render(chatResult);
|
||||
return aiOpsOutputAdapter.render(aiOpsResult);
|
||||
```
|
||||
|
||||
Chat 的普通问答仍走轻量链路;复杂诊断才进入公共 Graph。AIOps 在入口阶段固定主告警范围,Graph 内部继续复用证据收集和验证。
|
||||
|
||||
### 2.10 人工中断怎么放
|
||||
|
||||
普通工具失败不需要 interrupt,条件边即可。只有确实需要等待人工输入或审批时才增加节点:
|
||||
|
||||
```java
|
||||
CompileConfig compileConfig = CompileConfig.builder()
|
||||
.saverConfig(saverConfig)
|
||||
.interruptBefore("human_review")
|
||||
.build();
|
||||
```
|
||||
|
||||
恢复时使用相同 `threadId` 和 checkpoint 信息,并通过 `RunnableConfig.builder(oldConfig).resume()` 或状态更新接口继续。具体恢复协议需要结合当前版本做集成测试,不能只凭文档假设。
|
||||
|
||||
## 3. 深度裂变
|
||||
|
||||
<div class="fission-section">
|
||||
|
||||
### 🔍 搜索内化:真正的改造对象不是 Agent,而是控制权
|
||||
|
||||
官方文档和本地 `1.1.2.0` JAR 均证明 Graph 节点既可以是 LLM,也可以是普通 Java 代码,条件边由状态决定目标节点。`ReactAgent` 本身已经是一个子图,并提供 `asNode(...)` 适配能力。
|
||||
|
||||
因此“从 Multi-Agent 改成 Graph”并不准确。更准确的是:把跨 Agent 的状态转换从高层 Flow 抽出来,交给父 Graph;ReactAgent 继续作为子图存在。
|
||||
|
||||
文档页面示例存在 `OverAllStaste` 拼写错误,实际类名是 `OverAllState`。网站主分支可能领先于本地依赖,因此最终应以项目锁定版本的 JAR 签名和编译测试为准。
|
||||
|
||||
</div>
|
||||
|
||||
## 4. 实战指南
|
||||
|
||||
### 4.1 最小迁移顺序
|
||||
|
||||
1. 定义统一的节点结果状态,不先改 Prompt。
|
||||
2. 把现有 Planner、Executor、Verifier 调用包装为 Graph Node。
|
||||
3. 将 Gatekeeper 和固定降级模板做成普通 Java Node。
|
||||
4. 先迁移当前两轮 LOW_CONFID 循环。
|
||||
5. 为 Executor 阻断、非法结构、Gatekeeper REJECT 增加条件边。
|
||||
6. 保留原 Sequential 链路作为回退,完成行为对比后再删除。
|
||||
7. 最后再考虑 HITL、并行和专科 SubAgent。
|
||||
|
||||
### 4.2 常见反模式
|
||||
|
||||
- 把每个异常都交给 LLM Supervisor 决策。
|
||||
- Graph State 存放所有原始日志和完整思考过程。
|
||||
- 把 `no_evidence` 当作 Executor 执行失败。
|
||||
- 同一 session 的多个 run 共用一个 Graph `threadId`。
|
||||
- 一开始就设计几十个节点和完整 Incident 状态机。
|
||||
- 未验证父子图消息传播就直接大量使用 `ReactAgent.asNode(true, true)`。
|
||||
- 把 `MemorySaver` 当成生产级持久化。
|
||||
|
||||
### 4.3 ROI
|
||||
|
||||
收益:条件分支可见、失败可测试、门禁不可跳过、Trace 更容易与节点对齐。
|
||||
代价:需要维护状态契约、条件边和父子图上下文,并增加 Graph 级测试。
|
||||
判断标准:如果当前只有固定顺序且失败直接结束,SequentialAgent 更简单;当局部重试、降级、HITL 和多入口策略已经出现时,StateGraph 的控制收益开始超过复杂度。
|
||||
|
||||
## 5. 温故知新
|
||||
|
||||
### FAQ
|
||||
|
||||
1. **Graph 会替代 ReactAgent 吗?** 不会,ReactAgent 可以作为 Graph 节点或由适配节点调用。
|
||||
2. **为什么不用 Supervisor 控制失败?** 失败跳转是确定性规则,不应增加一次 LLM 决策。
|
||||
3. **no_evidence 是否直接走 fallback?** 不一定。合法的 no-evidence 引用仍要经过 Gatekeeper 和 Verifier。
|
||||
4. **Composer 是否必须是 Agent?** 正常表达可以是轻量 Agent,系统失败场景应保留固定模板。
|
||||
5. **threadId 用 sessionId 还是 runId?** 当前模型下优先用 runId,避免同一会话多次运行状态串扰。
|
||||
6. **什么时候需要 Checkpointer?** 需要暂停、恢复、查看历史 Graph 状态时;普通 Trace 持久化不自动等于 Graph Checkpoint。
|
||||
7. **能否直接使用 agent.asNode?** 可以,但需要验证输入、outputKey、messages 和父子 Checkpointer 行为。
|
||||
8. **AIOps 是否要单独一张 Graph?** 可以先共用诊断核心,入口范围策略和输出报告保持独立。
|
||||
|
||||
### 自测题
|
||||
|
||||
1. 为什么 Executor 工具阻断不应该由 Verifier 判断是否重试?
|
||||
2. `ReplaceStrategy` 和 `AppendStrategy` 在诊断状态中分别适合什么数据?
|
||||
3. 为什么合法 `no_evidence` 仍然需要 Gatekeeper?
|
||||
4. SupervisorAgent 和 StateGraph 条件边的决策权有什么区别?
|
||||
5. 为什么当前 Graph `threadId` 更适合使用 runId?
|
||||
6. 什么情况下才应该增加 `interruptBefore`?
|
||||
|
||||
### 参考资源
|
||||
|
||||
- https://java2ai.com/docs/frameworks/graph-core/core/core-library
|
||||
- https://java2ai.com/docs/frameworks/graph-core/quick-start
|
||||
- Spring AI Alibaba 本地依赖:`spring-ai-alibaba-graph-core:1.1.2.0`
|
||||
- 当前项目:`ChatService`、`AiOpsService`、`ExecutorGatekeeperService`、`VerifierInputHook`
|
||||
+30
-18
@@ -1,8 +1,8 @@
|
||||
# SuperBizAgent MVP 文档
|
||||
|
||||
**更新日期**:2026-07-05
|
||||
**更新日期**:2026-07-20
|
||||
|
||||
本目录保存 MVP 阶段的架构、问题、演示、评测和数据表说明。当前架构入口已经整理到 `mvp/architecture/`,旧版架构材料已归档,避免继续把历史方案当成当前实现。
|
||||
本目录保存 MVP 阶段的架构、问题、演示、评测和数据表说明。当前材料按“当前入口”和“历史归档”拆开,避免把早期设计稿当成当前实现。
|
||||
|
||||
## 当前入口
|
||||
|
||||
@@ -10,8 +10,10 @@
|
||||
|---|---|
|
||||
| [architecture/README.md](architecture/README.md) | 当前 MVP 架构入口 |
|
||||
| [architecture/current-mvp-architecture.md](architecture/current-mvp-architecture.md) | 当前可运行系统架构 |
|
||||
| [architecture/stategraph-runtime-architecture.md](architecture/stategraph-runtime-architecture.md) | 复杂 Chat StateGraph 运行时架构 |
|
||||
| [architecture/interview-one-pager.md](architecture/interview-one-pager.md) | 面试一页式架构讲解 |
|
||||
| [architecture/agent-orchestration.md](architecture/agent-orchestration.md) | Agent 编排架构 |
|
||||
| [architecture/executor-evidence-pipeline-refactor.md](architecture/executor-evidence-pipeline-refactor.md) | Executor 证据链路改造记录 |
|
||||
| [architecture/harness-quality-gates.md](architecture/harness-quality-gates.md) | Harness 与质量门禁 |
|
||||
| [architecture/rag-architecture.md](architecture/rag-architecture.md) | RAG/知识检索新架构 |
|
||||
| [architecture/retrieval-observability.md](architecture/retrieval-observability.md) | 检索与可观测性架构 |
|
||||
@@ -20,15 +22,16 @@
|
||||
| [architecture/knowledge-base-authoring.md](architecture/knowledge-base-authoring.md) | 知识库文档编写与维护 |
|
||||
| [architecture/data-model.md](architecture/data-model.md) | 数据模型总览 |
|
||||
| [architecture/evolution-roadmap.md](architecture/evolution-roadmap.md) | Agent 架构演进路线 |
|
||||
| [issues/rag-refactor-plan.md](issues/rag-refactor-plan.md) | RAG 重构计划和阶段拆解 |
|
||||
| [issues/README.md](issues/README.md) | MVP issue 索引 |
|
||||
| [issues/active/rag-refactor-plan.md](issues/active/rag-refactor-plan.md) | RAG 重构计划和阶段拆解 |
|
||||
| [tables/README.md](tables/README.md) | 当前 MySQL 表说明 |
|
||||
| [demo/README.md](demo/README.md) | Demo 运行和面试演示材料 |
|
||||
| [demo/ten-minute-interview-demo.md](demo/ten-minute-interview-demo.md) | 10 分钟面试演示脚本 |
|
||||
| [eval/README.md](eval/README.md) | 诊断评测材料 |
|
||||
| [issues/README.md](issues/README.md) | MVP issue 索引 |
|
||||
|
||||
## 当前系统一句话
|
||||
|
||||
SuperBizAgent MVP 是一个可追踪的故障诊断 Agent:Chat 和 AIOps 入口进入 Agent 编排,Executor 显式调用知识库、日志、指标等工具收集证据,诊断过程落到 `diagnosis_session`、`agent_step`、`tool_invocation`,最终通过 Trace API、Verifier 和评测脚本证明结果可解释、可回放、可对比。
|
||||
SuperBizAgent MVP 是一个可追踪的故障诊断 Agent:Chat 和 AIOps 入口进入 Agent 编排,Executor 显式调用知识库、日志、指标等工具收集证据;多轮会话元数据落到 `chat_session`,每次诊断运行落到 `diagnosis_run`,步骤和工具明细通过 `agent_step.run_id`、`tool_invocation.run_id` 关联,最终通过 Trace API、Verifier 和评测脚本证明结果可解释、可回放、可对比。
|
||||
|
||||
## 文档结构
|
||||
|
||||
@@ -37,8 +40,10 @@ mvp/
|
||||
architecture/
|
||||
README.md
|
||||
current-mvp-architecture.md
|
||||
stategraph-runtime-architecture.md
|
||||
interview-one-pager.md
|
||||
agent-orchestration.md
|
||||
executor-evidence-pipeline-refactor.md
|
||||
harness-quality-gates.md
|
||||
rag-architecture.md
|
||||
retrieval-observability.md
|
||||
@@ -47,12 +52,16 @@ mvp/
|
||||
knowledge-base-authoring.md
|
||||
data-model.md
|
||||
evolution-roadmap.md
|
||||
archive/2026-07-05-legacy/
|
||||
archive/
|
||||
issues/
|
||||
README.md
|
||||
rag-refactor-plan.md
|
||||
ISS-*.md
|
||||
rag-*.md
|
||||
active/
|
||||
archived/
|
||||
rag/
|
||||
tables/
|
||||
README.md
|
||||
*表-*.md
|
||||
archive/
|
||||
demo/
|
||||
README.md
|
||||
ten-minute-interview-demo.md
|
||||
@@ -65,9 +74,7 @@ mvp/
|
||||
cases/
|
||||
fixtures/
|
||||
reports/
|
||||
notes/
|
||||
plan/
|
||||
tables/
|
||||
archive/
|
||||
```
|
||||
|
||||
## 当前核心设计
|
||||
@@ -77,7 +84,8 @@ mvp/
|
||||
- `VectorSearchService` 是检索稳定门面。
|
||||
- Spring AI VectorStore 是当前读取主路径,Milvus SDK 保留为 fallback。
|
||||
- AIOps payload 会生成推荐知识库 query,保留业务语义。
|
||||
- Trace API 聚合 session、step、tool invocation 和 self evaluation。
|
||||
- `sessionId` 表示多轮会话上下文,`runId` 表示一次可回放诊断运行。
|
||||
- Trace API 聚合 `diagnosis_run`、`agent_step.run_id`、`tool_invocation.run_id` 和 self evaluation。
|
||||
- RAG 行为通过 offline baseline 和 live acceptance 脚本做回归验证。
|
||||
|
||||
## 关键运行链路
|
||||
@@ -87,7 +95,8 @@ Chat
|
||||
-> ChatService
|
||||
-> Planner / Executor / Verifier
|
||||
-> evidence tools
|
||||
-> diagnosis_session / agent_step / tool_invocation
|
||||
-> chat_session / diagnosis_run
|
||||
-> agent_step.run_id / tool_invocation.run_id
|
||||
-> DiagnosisTraceService
|
||||
|
||||
AIOps
|
||||
@@ -96,6 +105,7 @@ AIOps
|
||||
-> Planner / Executor
|
||||
-> Prometheus / logs / lookup_knowledge
|
||||
-> AiOpsRuleEvaluationService
|
||||
-> diagnosis_run(agent_flow=AI_OPS)
|
||||
-> DiagnosisTraceService
|
||||
|
||||
RAG
|
||||
@@ -107,10 +117,12 @@ RAG
|
||||
-> tool_invocation
|
||||
```
|
||||
|
||||
## 旧文档说明
|
||||
## 归档说明
|
||||
|
||||
旧版架构文档已移动到:
|
||||
历史材料分三类:
|
||||
|
||||
- [architecture/archive/2026-07-05-legacy/](architecture/archive/2026-07-05-legacy/)
|
||||
- 旧架构文档:[architecture/archive/2026-07-05-legacy/](architecture/archive/2026-07-05-legacy/)
|
||||
- 本次文档清理归档:[archive/2026-07-09-doc-cleanup/](archive/2026-07-09-doc-cleanup/)
|
||||
- Run-only v2 和 StateGraph 收口后的旧材料:[archive/2026-07-20-doc-cleanup/](archive/2026-07-20-doc-cleanup/)
|
||||
|
||||
归档文档只用于追溯设计历史。当前实现和后续规划以 `architecture/current-mvp-architecture.md` 与 `architecture/rag-architecture.md` 为准。
|
||||
归档文档只用于追溯设计历史。当前实现和后续规划以 `architecture/`、`issues/README.md`、`tables/README.md` 和 OpenSpec/devflow 的最新记录为准。
|
||||
|
||||
+19
-15
@@ -1,6 +1,6 @@
|
||||
# MVP 架构文档
|
||||
|
||||
**更新日期**:2026-07-06
|
||||
**更新日期**:2026-07-20
|
||||
|
||||
这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到:
|
||||
|
||||
@@ -13,9 +13,11 @@
|
||||
| 文档 | 用途 |
|
||||
|---|---|
|
||||
| [current-mvp-architecture.md](current-mvp-architecture.md) | 当前可运行 MVP 的总体架构、链路、持久化和质量门禁 |
|
||||
| [stategraph-runtime-architecture.md](stategraph-runtime-architecture.md) | 最新复杂 Chat StateGraph 运行时权威快照,覆盖 Node 路由、状态边界、Run/Trace 持久化和验收层次 |
|
||||
| [interview-one-pager.md](interview-one-pager.md) | 面试一页式架构讲解,包含总图、亮点、取舍和追问回答 |
|
||||
| [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat SequentialAgent、AIOps SupervisorAgent、工具边界 |
|
||||
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Verifier、评测基线组成的质量门禁 |
|
||||
| [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat bounded StateGraph、AIOps SupervisorAgent、工具边界 |
|
||||
| [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md) | Chat 证据链路当前数据契约,覆盖 Executor V2、Gatekeeper、Verifier、Composer、`evidence_refs` |
|
||||
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、StateGraph、Trace、Gatekeeper、Verifier、Composer、评测基线组成的质量门禁 |
|
||||
| [rag-architecture.md](rag-architecture.md) | RAG/知识检索新架构,覆盖 L0 hint、VectorStore 主路径、SDK fallback、证据追踪 |
|
||||
| [modular-rag-pipeline.md](modular-rag-pipeline.md) | `lookup_knowledge` 模块化 RAG 落地架构,覆盖 pipeline、fallback、evidence-first contract、trace |
|
||||
| [rag-eval-closure.md](rag-eval-closure.md) | RAG 评测闭环,覆盖 offline baseline、baseline diff、diagnosis eval 和 live acceptance |
|
||||
@@ -28,19 +30,21 @@
|
||||
|
||||
## 当前架构一句话
|
||||
|
||||
SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,执行过程落到 `diagnosis_session`、`agent_step`、`tool_invocation`,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
|
||||
SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:复杂 Chat 由有界 StateGraph 显式编排 Planner、Executor、Gatekeeper、Verified Input、Verifier、Composer 与安全 Fallback,Executor 通过工具收集日志、指标和知识库证据;多轮会话元数据落到 `chat_session`,每次诊断运行落到 `diagnosis_run`,步骤和工具明细通过 `run_id` 关联,Graph 路由摘要独立保存为 `orchestration_trace`,最终通过精确 Run Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
|
||||
|
||||
## 阅读顺序
|
||||
|
||||
1. 先读 [current-mvp-architecture.md](current-mvp-architecture.md),理解系统边界和主链路。
|
||||
2. 面试前读 [interview-one-pager.md](interview-one-pager.md),准备 2-5 分钟讲解。
|
||||
3. 再读 [agent-orchestration.md](agent-orchestration.md),理解当前 Agent 如何协作。
|
||||
4. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。
|
||||
5. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。
|
||||
6. 继续读 [modular-rag-pipeline.md](modular-rag-pipeline.md),看 `lookup_knowledge` 的模块化落地和 evidence-first contract。
|
||||
7. 再读 [rag-eval-closure.md](rag-eval-closure.md),看 RAG baseline 如何形成质量闭环。
|
||||
8. 然后读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
|
||||
9. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
|
||||
10. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
|
||||
11. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
|
||||
12. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
|
||||
2. 再读 [stategraph-runtime-architecture.md](stategraph-runtime-architecture.md),理解复杂 Chat 的真实运行时和路由边界。
|
||||
3. 面试前读 [interview-one-pager.md](interview-one-pager.md),准备 2-5 分钟讲解。
|
||||
4. 再读 [agent-orchestration.md](agent-orchestration.md),理解当前 Agent 如何协作。
|
||||
5. 接着读 [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md),理解 Chat 证据链路的数据结构和验真边界。
|
||||
6. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。
|
||||
7. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。
|
||||
8. 继续读 [modular-rag-pipeline.md](modular-rag-pipeline.md),看 `lookup_knowledge` 的模块化落地和 evidence-first contract。
|
||||
9. 再读 [rag-eval-closure.md](rag-eval-closure.md),看 RAG baseline 如何形成质量闭环。
|
||||
10. 然后读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
|
||||
11. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
|
||||
12. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
|
||||
13. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
|
||||
14. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
|
||||
|
||||
@@ -1,14 +1,14 @@
|
||||
# Agent 编排架构
|
||||
|
||||
**更新日期**:2026-07-05
|
||||
**状态**:当前可运行架构
|
||||
**更新日期**:2026-07-20
|
||||
**状态**:当前可运行架构
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
||||
|
||||
## 1. 设计定位
|
||||
|
||||
旧版 Agent 架构把系统描述为 Supervisor、Planner、SubAgent、Verifier 的团队协作。当前 MVP 保留这个核心思想,但实现更收敛:
|
||||
|
||||
- Chat 链路使用固定顺序工作流:`Planner -> Executor -> Verifier`。
|
||||
- Chat 复杂诊断使用有递归上限的显式 StateGraph;正常路径是 `Planner -> Executor -> Gatekeeper -> Verified Input -> Verifier -> Composer`,条件边负责有限技术重试、一次补证据和安全 Fallback。
|
||||
- AIOps 链路使用 `SupervisorAgent` 调度 `Planner + Executor`,最终由规则评估器做轻量验证。
|
||||
- 当前没有拆分 ExternalApiSubAgent、InternalErrorSubAgent、DatabaseSubAgent;这些作为后续演进方向保留。
|
||||
- 证据工具不直接散落在各个 Agent 里,而是通过 Spring AI ToolCallback / `@Tool` 统一暴露。
|
||||
@@ -19,13 +19,19 @@
|
||||
flowchart TB
|
||||
subgraph Chat["Chat diagnosis"]
|
||||
ChatIn["POST /api/chat"] --> ChatService["ChatService"]
|
||||
ChatService --> ChatPlanner["chat_planner"]
|
||||
ChatService --> ChatGraph["ChatDiagnosisGraphRuntime / StateGraph"]
|
||||
ChatGraph --> ChatPlanner["Planner Node"]
|
||||
ChatPlanner --> ChatExecutor["chat_executor"]
|
||||
ChatExecutor --> ChatTools["evidence tools"]
|
||||
ChatTools --> ChatExecutor
|
||||
ChatExecutor --> ChatVerifier["chat_verifier"]
|
||||
ChatExecutor --> ChatGatekeeper["Gatekeeper Node"]
|
||||
ChatGatekeeper --> VerifiedInput["Verified Input Node"]
|
||||
VerifiedInput --> ChatVerifier["Verifier Node"]
|
||||
ChatVerifier --> ChatDecision{"PASS / LOW_CONFID / REJECT"}
|
||||
ChatDecision --> ChatAnswer["final answer"]
|
||||
ChatDecision --> ChatComposer["Composer Node"]
|
||||
ChatDecision --> ChatFallback["Fallback Node"]
|
||||
ChatComposer --> ChatAnswer["final answer"]
|
||||
ChatFallback --> ChatAnswer
|
||||
end
|
||||
|
||||
subgraph AiOps["AIOps diagnosis"]
|
||||
@@ -40,20 +46,26 @@ flowchart TB
|
||||
end
|
||||
|
||||
subgraph Trace["Trace persistence"]
|
||||
Session["diagnosis_session"]
|
||||
ChatSession["chat_session"]
|
||||
Run["diagnosis_run"]
|
||||
Step["agent_step"]
|
||||
Invocation["tool_invocation"]
|
||||
SelfEval["self_evaluation"]
|
||||
end
|
||||
|
||||
ChatService --> Session
|
||||
ChatService --> ChatSession
|
||||
ChatService --> Run
|
||||
ChatPlanner --> Step
|
||||
ChatExecutor --> Step
|
||||
ChatGatekeeper --> SelfEval
|
||||
ChatGraph --> Run
|
||||
ChatVerifier --> Step
|
||||
ChatTools --> Invocation
|
||||
ChatDecision --> SelfEval
|
||||
ChatComposer --> Step
|
||||
|
||||
AiOpsService --> Session
|
||||
AiOpsService --> ChatSession
|
||||
AiOpsService --> Run
|
||||
AiOpsPlanner --> Step
|
||||
AiOpsExecutor --> Step
|
||||
AiOpsTools --> Invocation
|
||||
@@ -62,15 +74,13 @@ flowchart TB
|
||||
|
||||
## 3. Chat 编排
|
||||
|
||||
Chat 复杂诊断采用 `SequentialAgent`,顺序固定:
|
||||
Chat 复杂诊断采用 `ChatDiagnosisGraphRuntime` 执行、`DiagnosisGraphFactory` 编译的 bounded StateGraph。它有一条正常路径和显式条件边,不再依赖固定顺序 Agent 或隐式前置校验:
|
||||
|
||||
```text
|
||||
chat_planner
|
||||
-> chat_executor
|
||||
-> lookup_knowledge / query_logs / query_metrics / date_time
|
||||
-> chat_verifier
|
||||
-> reads tool_trace_summary
|
||||
-> outputs verifier JSON
|
||||
START -> PLANNER -> EXECUTOR -> GATEKEEPER -> VERIFIED_INPUT -> VERIFIER -> COMPOSER -> END
|
||||
| | | | |
|
||||
+ retry + fallback + fallback + retry + retry/fallback
|
||||
+ EVIDENCE_RETRY -> PLANNER (最多一次)
|
||||
```
|
||||
|
||||
关键行为:
|
||||
@@ -78,44 +88,68 @@ chat_planner
|
||||
| 角色 | 当前职责 | 输出 |
|
||||
|---|---|---|
|
||||
| `chat_planner` | 拆解问题,注入知识域地图和对话历史,给出排查方向 | `planner_plan` |
|
||||
| `chat_executor` | 按计划调用证据工具,组合工具返回形成诊断答复 | `executor_feedback` |
|
||||
| `chat_verifier` | 只基于已有证据校验 Executor 答案,不做新检索 | `verifier_output` |
|
||||
| `chat_executor` | 按计划调用证据工具,抽取带 `source_invocation_id + raw_path + evidence_excerpt` 的微观事实 | `executor_evidence_v2` |
|
||||
| `GatekeeperNode` / `ExecutorGatekeeperService` | 按当前 `runId` 做代码级引用验真,拒绝伪造 ID、错配 raw_path、错配 excerpt | `gatekeeper_result` |
|
||||
| `VerifiedInputNode` | 只投影 Gatekeeper 通过的 claims 与 matched evidence,隔离完整工具 Trace | `verified_executor_output`、`verified_evidence` |
|
||||
| `chat_verifier` | 只判断已验真的 evidence excerpt 是否能推出 claim,不做新检索、不读取完整工具 Trace | `verifier_output` |
|
||||
| `chat_composer` | 只表达 Verifier 允许输出的 claims、缺口和建议,生成最终用户答复 | `composer_output` |
|
||||
| `FallbackNode` | 在不可恢复失败或路由上限触发时生成非空安全答复 | `final_answer`、degraded trace |
|
||||
|
||||
Chat 链路最多支持两轮验证:
|
||||
Chat Graph 支持有限技术重试,并只允许一次 evidence retry;所有分支最终进入 Composer 或 Fallback:
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant C as ChatService
|
||||
participant C as ChatService / StateGraph
|
||||
participant P as chat_planner
|
||||
participant E as chat_executor
|
||||
participant T as tools
|
||||
participant G as gatekeeper
|
||||
participant V as chat_verifier
|
||||
participant S as diagnosis_session
|
||||
participant M as chat_composer
|
||||
participant R as diagnosis_run
|
||||
|
||||
C->>P: 原始问题 + history + retry_context
|
||||
P-->>C: planner_plan
|
||||
C->>E: planner_plan + 上下文
|
||||
E->>T: 调用证据工具
|
||||
T-->>E: 证据结果
|
||||
E-->>C: executor_feedback
|
||||
C->>V: executor_final_answer + tool_trace_summary
|
||||
E-->>C: executor_evidence_v2
|
||||
C->>G: executor_output + run-owned tool_invocation.evidence_refs
|
||||
G-->>C: gatekeeper_result
|
||||
C->>C: VerifiedInputNode projects passed claims/evidence
|
||||
C->>V: verified_executor_output + verified_evidence + gatekeeper_audit
|
||||
V-->>C: PASS / LOW_CONFID / REJECT
|
||||
C->>S: 写入 verifier_evaluation
|
||||
alt LOW_CONFID 且允许补证据
|
||||
C->>R: 写入 verifier_evaluation
|
||||
alt LOW_CONFID 且允许一次补证据
|
||||
C->>P: retry_context: 仅补缺失证据
|
||||
else PASS 或 REJECT
|
||||
C-->>S: 保存最终 answer
|
||||
else PASS / LOW_CONFID 可输出
|
||||
C->>M: allowed_claims + missing_info + recommended_actions
|
||||
M-->>C: composer_output
|
||||
C->>R: 保存 Composer 最终 answer
|
||||
else 不可恢复失败
|
||||
C->>R: Fallback 安全答复
|
||||
end
|
||||
C->>R: 保存 orchestration_trace(version/transitions/final_node/termination_reason/degraded/evidence_retry_count)
|
||||
```
|
||||
|
||||
决策语义:
|
||||
|
||||
| Verdict | 行为 |
|
||||
|---|---|
|
||||
| `PASS` | 输出 Executor 答案 |
|
||||
| `PASS` | 把 Verifier 允许表达的 claims 交给 Composer 输出 |
|
||||
| `LOW_CONFID` | 如果分数低于阈值且仍有轮次,构造 `retry_context` 补证据;否则输出低置信提示 |
|
||||
| `REJECT` | 输出降级答复,只保留已确认信息和下一步建议 |
|
||||
| `REJECT` | Verifier 完成后仍进入 Composer,但 Composer 只能表达允许材料和诊断限制;Gatekeeper REJECT 才直接进入 Fallback |
|
||||
|
||||
### 3.1 运行时边界
|
||||
|
||||
- 外层 Graph 使用 `runId` 作为 `RunnableConfig.threadId`,metadata 只承载 `sessionId/runId` 等业务身份。
|
||||
- `ReactAgentDiagnosisInvoker` 为 nested Agent 创建独立 config,不传播外层 human-feedback、state-update、checkpoint/resume 控制 metadata。
|
||||
- Graph State 默认 replace,只有 `orchestration_events` append;事件使用 portable Map,避免 DevTools classloader 的 record identity 问题。
|
||||
- Planner、Verifier、Composer 各最多一次技术重试;evidence retry 独立计数且最多一次;Graph recursion limit 为 32。
|
||||
- `DiagnosisGraphResultMapper` 负责质量评估,`DiagnosisOrchestrationTraceBuilder` 负责路由摘要,两者分别写入 `self_evaluation` 和 `orchestration_trace`。
|
||||
|
||||
完整运行时拓扑、路由表和失败语义见 [stategraph-runtime-architecture.md](stategraph-runtime-architecture.md)。
|
||||
|
||||
## 4. AIOps 编排
|
||||
|
||||
@@ -177,21 +211,25 @@ flowchart LR
|
||||
SkillBody --> Executor
|
||||
Executor --> EvidenceTools["lookup_knowledge / logs / metrics"]
|
||||
EvidenceTools --> ToolTrace["tool_invocation evidence"]
|
||||
Executor --> Verifier["Verifier"]
|
||||
Executor --> Gatekeeper["Gatekeeper"]
|
||||
Gatekeeper --> Verifier["Verifier"]
|
||||
ToolTrace --> Verifier
|
||||
Verifier --> Composer["Composer"]
|
||||
```
|
||||
|
||||
| 角色 | Skill 可见性 | 工具权限 |
|
||||
|---|---|---|
|
||||
| Planner | 只看 skill name / description,并输出 `selected_skill` | 不暴露 `read_skill` |
|
||||
| Executor | 读取 Planner 选中的 skill 正文 | 暴露官方 `read_skill` 和证据工具 |
|
||||
| Verifier | 不看 skill catalog,也不读 skill 正文 | 只读取 `tool_trace_summary` |
|
||||
| Gatekeeper | 不看 skill catalog,也不读 skill 正文 | 只读取 Executor 输出和 `tool_invocation.retrieval_details.evidence_refs` |
|
||||
| Verifier | 不看 skill catalog,也不读 skill 正文 | 只读取 Verified Input 投影和 Gatekeeper audit |
|
||||
| Composer | 不看 skill catalog,也不读 skill 正文 | 只读取 Verifier 允许表达的内容 |
|
||||
|
||||
## 7. 与旧版设计的差异
|
||||
|
||||
| 旧版设想 | 当前实现 |
|
||||
|---|---|
|
||||
| Supervisor + Planner + 多个专科 SubAgent + Verifier | Chat: Planner + Executor + Verifier;AIOps: Supervisor + Planner + Executor |
|
||||
| Supervisor + Planner + 多个专科 SubAgent + Verifier | Chat: Planner + Executor + Gatekeeper + Verifier + Composer;AIOps: Supervisor + Planner + Executor |
|
||||
| ExternalApiSubAgent / InternalErrorSubAgent / DatabaseSubAgent | 暂未拆分,能力通过通用 Executor + 工具 + Prompt 约束实现 |
|
||||
| 每个 SubAgent 专属工具集 | 当前 Executor 持有统一证据工具集合 |
|
||||
| Verifier 支持 PASS / REVISE / REJECT | 当前 Chat Verifier 输出 PASS / LOW_CONFID / REJECT |
|
||||
|
||||
@@ -154,4 +154,4 @@ ALTER TABLE diagnosis_session ADD COLUMN answer LONGTEXT COMMENT 'Agent 返回
|
||||
|
||||
- **LLM 观点层**:在 `selfEvaluation` 的 `llm_opinion` 字段叠加 LLM 结构化观点(has_root_cause、has_solution 等),作为独立 factors,不改变现有规则逻辑
|
||||
- **案例结构化字段**:useful 触发时自动提取 faultCategory / errorCode,替代暂时的 GENERAL
|
||||
- **重复召回问题**:Executor Prompt 约束或工具层 session 维度去重(见 [ISS-001](../issues/ISS-001-duplicate-retrieval.md))
|
||||
- **重复召回问题**:Executor Prompt 约束或工具层 session 维度去重(见 [ISS-001](../../../issues/archived/ISS-001-duplicate-retrieval.md))
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
# 当前 MVP 架构
|
||||
|
||||
**更新日期**:2026-07-05
|
||||
**状态**:当前可运行架构
|
||||
**更新日期**:2026-07-20
|
||||
**状态**:当前可运行架构
|
||||
**适用范围**:Demo、面试讲解、后续迭代规划
|
||||
|
||||
## 1. 系统定位
|
||||
@@ -36,9 +36,14 @@ flowchart TB
|
||||
|
||||
subgraph Agent["Agent Orchestration"]
|
||||
Supervisor["Supervisor"]
|
||||
StateGraph["Chat Diagnosis StateGraph"]
|
||||
Planner["Planner"]
|
||||
Executor["Executor"]
|
||||
Gatekeeper["Gatekeeper"]
|
||||
VerifiedInput["Verified Input"]
|
||||
Verifier["Verifier"]
|
||||
Composer["Composer"]
|
||||
Fallback["Fallback"]
|
||||
end
|
||||
|
||||
subgraph Tools["Evidence Tools"]
|
||||
@@ -63,7 +68,8 @@ flowchart TB
|
||||
end
|
||||
|
||||
subgraph Store["Persistence and Trace"]
|
||||
Session["diagnosis_session"]
|
||||
ChatSession["chat_session"]
|
||||
Run["diagnosis_run"]
|
||||
Step["agent_step"]
|
||||
Invocation["tool_invocation"]
|
||||
ApiDoc["api_document"]
|
||||
@@ -71,7 +77,14 @@ flowchart TB
|
||||
end
|
||||
|
||||
API --> App
|
||||
ChatService --> Agent
|
||||
ChatService --> StateGraph
|
||||
StateGraph --> Planner
|
||||
StateGraph --> Executor
|
||||
StateGraph --> Gatekeeper
|
||||
StateGraph --> VerifiedInput
|
||||
StateGraph --> Verifier
|
||||
StateGraph --> Composer
|
||||
StateGraph --> Fallback
|
||||
AiOpsService --> Agent
|
||||
SkillRegistry --> PlannerSkillHook
|
||||
PlannerSkillHook --> Planner
|
||||
@@ -83,8 +96,8 @@ flowchart TB
|
||||
RAG --> Store
|
||||
Tools --> Invocation
|
||||
Agent --> Step
|
||||
App --> Session
|
||||
TraceService --> Session
|
||||
App --> ChatSession
|
||||
TraceService --> ChatSession
|
||||
TraceService --> Step
|
||||
TraceService --> Invocation
|
||||
```
|
||||
@@ -105,7 +118,9 @@ Agent Orchestration
|
||||
-> Supervisor
|
||||
-> Planner
|
||||
-> Executor
|
||||
-> Gatekeeper
|
||||
-> Verifier
|
||||
-> Composer
|
||||
|
||||
Evidence Tools
|
||||
-> lookup_knowledge
|
||||
@@ -126,13 +141,15 @@ RAG Retrieval
|
||||
-> Milvus SDK fallback
|
||||
|
||||
Persistence
|
||||
-> diagnosis_session
|
||||
-> agent_step
|
||||
-> tool_invocation
|
||||
-> chat_session
|
||||
-> diagnosis_run
|
||||
-> agent_step.run_id
|
||||
-> tool_invocation.run_id
|
||||
-> api_document
|
||||
-> Milvus/Zilliz collection
|
||||
|
||||
Quality Gates
|
||||
-> executor gatekeeper
|
||||
-> chat verifier
|
||||
-> AIOps rule evaluation
|
||||
-> diagnosis eval baseline
|
||||
@@ -150,23 +167,32 @@ sequenceDiagram
|
||||
participant Planner as Planner Agent
|
||||
participant Executor as Executor Agent
|
||||
participant Tool as Evidence Tools
|
||||
participant Gatekeeper as Gatekeeper Node
|
||||
participant Projection as Verified Input Node
|
||||
participant Verifier as Verifier Agent
|
||||
participant Composer as Composer Agent
|
||||
participant DB as Trace Tables
|
||||
participant Trace as Trace API
|
||||
|
||||
User->>API: 提交诊断问题
|
||||
API->>Chat: execute chat strategy
|
||||
Chat->>DB: 创建 chat_session metadata + diagnosis_run(runId)
|
||||
Chat->>Planner: 复杂问题进入规划
|
||||
Planner->>DB: 写入 agent_step
|
||||
Planner->>DB: 写入 agent_step.run_id
|
||||
Planner->>Executor: 下发排查方向
|
||||
Executor->>Tool: lookup_knowledge / logs / metrics
|
||||
Tool->>DB: 写入 tool_invocation
|
||||
Tool->>DB: 写入 tool_invocation.run_id
|
||||
Tool-->>Executor: 返回证据
|
||||
Executor->>Verifier: 生成候选诊断并校验
|
||||
Verifier->>DB: 合并 self_evaluation.verifier_evaluation
|
||||
Chat->>DB: 保存 diagnosis_session.answer
|
||||
User->>Trace: GET /api/diagnosis/{sessionId}/trace
|
||||
Trace->>DB: 聚合 session / step / tool
|
||||
Executor->>Gatekeeper: 输出 executor_evidence_v2
|
||||
Gatekeeper->>DB: 读取 tool_invocation.evidence_refs 并校验引用
|
||||
Gatekeeper->>Projection: 输出通过验真的 bindings
|
||||
Projection->>Verifier: 只传入 verified claims / evidence
|
||||
Verifier->>DB: 合并 diagnosis_run.self_evaluation.verifier_evaluation
|
||||
Verifier->>Composer: 传入 allowed_claims / missing_info / actions
|
||||
Composer->>Chat: 生成最终用户答复
|
||||
Chat->>DB: 保存 answer / self_evaluation / orchestration_trace
|
||||
User->>Trace: GET /api/diagnosis/{sessionId}/trace?runId=...
|
||||
Trace->>DB: 聚合 run / step / tool
|
||||
Trace-->>User: 返回可回放诊断链路
|
||||
```
|
||||
|
||||
@@ -180,16 +206,21 @@ POST /api/chat
|
||||
-> lookup_knowledge
|
||||
-> query_logs
|
||||
-> query_metrics
|
||||
-> Verifier 校验最终诊断
|
||||
-> 保存 diagnosis_session
|
||||
-> 保存 agent_step
|
||||
-> 保存 tool_invocation
|
||||
-> 合并 self_evaluation.verifier_evaluation
|
||||
-> Gatekeeper Node 校验 Executor 证据引用真实性
|
||||
-> Verified Input Node 只投影通过验真的 claims/evidence
|
||||
-> Verifier 判断 claim 是否能由已验真证据推出
|
||||
-> Composer 生成最终用户答复
|
||||
-> 保存 chat_session metadata
|
||||
-> 保存 diagnosis_run
|
||||
-> 保存 agent_step.run_id
|
||||
-> 保存 tool_invocation.run_id
|
||||
-> 合并 diagnosis_run.self_evaluation.verifier_evaluation
|
||||
-> 保存 diagnosis_run.orchestration_trace
|
||||
```
|
||||
|
||||
Chat 链路的质量门禁是 LLM Verifier。Verifier 输出合并到 `diagnosis_session.self_evaluation.verifier_evaluation`,Trace API 会展示该验证结果。
|
||||
Chat 链路的质量门禁由四段组成:Gatekeeper 先做代码级引用验真,Verified Input 再隔离未通过的 binding,Verifier 做 LLM 可推导性判断,Composer 最后控制对用户的表达边界。结构化质量结果合并到 `diagnosis_run.self_evaluation.verifier_evaluation`;Node 路由、重试和降级摘要独立写入 `diagnosis_run.orchestration_trace`。
|
||||
|
||||
Agent 编排细节见 [agent-orchestration.md](agent-orchestration.md)。
|
||||
最新运行时、路由和状态边界见 [stategraph-runtime-architecture.md](stategraph-runtime-architecture.md),Agent 职责见 [agent-orchestration.md](agent-orchestration.md)。
|
||||
|
||||
关键代码:
|
||||
|
||||
@@ -242,7 +273,7 @@ POST /api/ai_ops
|
||||
-> Prometheus / logs / knowledge tools
|
||||
-> 生成告警分析报告
|
||||
-> AiOpsRuleEvaluationService
|
||||
-> 合并 self_evaluation.aiops_rule_evaluation
|
||||
-> 合并 diagnosis_run.self_evaluation.aiops_rule_evaluation
|
||||
-> Trace API 可查看全链路
|
||||
```
|
||||
|
||||
@@ -304,32 +335,42 @@ RAG 总体设计见 [rag-architecture.md](rag-architecture.md),检索运行细
|
||||
|
||||
## 6. 持久化模型
|
||||
|
||||
当前诊断持久化以三张表为核心:
|
||||
当前诊断持久化以 session/run/trace 明细为核心:
|
||||
|
||||
```text
|
||||
diagnosis_session
|
||||
-> 一次诊断会话的主记录
|
||||
chat_session
|
||||
-> 多轮会话目录和元数据
|
||||
-> session_id / status / message_pair_count
|
||||
|
||||
diagnosis_run
|
||||
-> 一次诊断运行的主记录
|
||||
-> run_id / session_id
|
||||
-> query / status / agent_flow / answer
|
||||
-> self_evaluation
|
||||
-> orchestration_trace
|
||||
-> step_count / tool_call_count / duration
|
||||
|
||||
agent_step
|
||||
-> Agent 模型调用步骤
|
||||
-> session_id / run_id
|
||||
-> step_index / agent_name
|
||||
-> model_input / model_output / thought
|
||||
-> duration / token_count
|
||||
|
||||
tool_invocation
|
||||
-> 工具调用事实
|
||||
-> session_id / run_id
|
||||
-> tool_name / input_params / output_preview
|
||||
-> retrieval_layer / retrieval_details
|
||||
-> retrieval_details.evidence_refs
|
||||
-> relevance_level / dedup_reason
|
||||
-> duration / success
|
||||
```
|
||||
|
||||
说明:
|
||||
|
||||
- 旧的 `diagnosis_record` 已不是当前主模型,迁移脚本中已经由 `diagnosis_session + agent_step + tool_invocation` 取代。
|
||||
- 旧的 `diagnosis_record` 已不是当前主模型。
|
||||
- 当前运行时只使用 `chat_session + diagnosis_run`,不再映射、读取或写入 `diagnosis_session`。
|
||||
- `api_document` 仍用于文档元数据管理。
|
||||
- 文档向量内容存放在 Milvus/Zilliz collection 中。
|
||||
|
||||
@@ -339,19 +380,21 @@ tool_invocation
|
||||
|
||||
```text
|
||||
GET /api/diagnosis/{sessionId}/trace
|
||||
GET /api/diagnosis/{sessionId}/trace?runId=run-...
|
||||
```
|
||||
|
||||
Trace API 聚合:
|
||||
|
||||
- 会话状态和最终报告。
|
||||
- 会话元数据、运行状态和最终报告。
|
||||
- Agent step 序列。
|
||||
- 工具调用和检索细节。
|
||||
- Chat verifier 结果。
|
||||
- Chat Gatekeeper / Verifier / Composer 结果。
|
||||
- Chat `run.orchestrationTrace` 路由摘要,解析自 `diagnosis_run.orchestration_trace`,独立于 self-evaluation 和步骤/工具明细。
|
||||
- AIOps rule evaluation 结果。
|
||||
|
||||
Trace 是本项目区别于普通问答系统的关键:答案不是孤立文本,而是可以追溯到 Agent 决策、工具调用和证据来源。
|
||||
|
||||
Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gates.md](harness-quality-gates.md),用户反馈与 `self_evaluation` 闭环见 [feedback-architecture.md](feedback-architecture.md)。
|
||||
Prompt、StateGraph、Gatekeeper、Verifier、Composer 和评测门禁的完整说明见 [harness-quality-gates.md](harness-quality-gates.md),用户反馈与 `self_evaluation` 闭环见 [feedback-architecture.md](feedback-architecture.md)。
|
||||
|
||||
## 8. 质量门禁
|
||||
|
||||
@@ -359,7 +402,11 @@ Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gate
|
||||
|
||||
| 门禁 | 位置 | 作用 |
|
||||
|---|---|---|
|
||||
| Chat Verifier | `ChatService` | 校验普通诊断回答质量 |
|
||||
| Executor Gatekeeper | `GatekeeperNode` / `ExecutorGatekeeperService` | 校验当前 Run 的 invocation、`raw_path`、`evidence_excerpt` 是否真实 |
|
||||
| Verified Input | `VerifiedInputNode` | 仅投影 Gatekeeper 通过的 claims/evidence,阻断完整工具 Trace 进入 Verifier |
|
||||
| Chat Verifier | `VerifierNodeAdapter` | 判断已验真证据是否能推出 Executor claims |
|
||||
| Chat Composer / Fallback | `ComposerNodeAdapter` / `FallbackNode` | 输出受控答复;异常分支也必须安全终止 |
|
||||
| Graph routing | `DiagnosisGraphWorkflowTest` / `run.orchestrationTrace` | 验证条件边、有限重试、最终节点和终止原因 |
|
||||
| AIOps Rule Evaluation | `AiOpsRuleEvaluationService` | 校验告警诊断是否聚焦 payload 并使用证据 |
|
||||
| Diagnosis Eval Baseline | `mvp/eval/` | 固化诊断 trace 和报告行为 |
|
||||
| RAG Retrieval Baseline | `eval/rag-retrieval/` | 固化检索召回行为,避免 RAG 重构回退 |
|
||||
@@ -370,6 +417,7 @@ Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gate
|
||||
已经完成:
|
||||
|
||||
- Chat 和 AIOps 两条入口链路。
|
||||
- Chat 复杂诊断已单轨切换到 bounded StateGraph,并持久化 Run-owned `orchestration_trace`;Graph event 使用 portable Map,nested ReactAgent config 与外层 checkpoint/resume 控制信息隔离。
|
||||
- 显式 `lookup_knowledge` Agent Tool。
|
||||
- L0 从最终决策降级为 domain/entity hint。
|
||||
- `VectorSearchService` 作为稳定检索门面。
|
||||
@@ -379,6 +427,11 @@ Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gate
|
||||
- `title`、`breadcrumb`、`content` 参与 embedding 文本。
|
||||
- `tool_invocation` 记录检索层、relevance level、dedup reason。
|
||||
- Chat verifier 和 AIOps rule evaluation 合并进 `self_evaluation`。
|
||||
- Chat Executor 结构化输出 `executor_evidence_v2`,不再直接承担最终用户答复。
|
||||
- `tool_invocation.retrieval_details.evidence_refs` 支持 `raw_path` 精确引用和 `$.no_evidence` 负向证据。
|
||||
- Gatekeeper 对 Executor 引用做代码级验真,并在审计中记录 `rule_set_version` 和规则元数据摘要。
|
||||
- Verifier 只判断可推导性。
|
||||
- Composer 在 Verifier 之后生成最终用户表达,并限制 negative observation 过度表述。
|
||||
- RAG offline baseline 和 live acceptance 脚本。
|
||||
|
||||
暂不作为当前已完成能力声明:
|
||||
@@ -389,13 +442,16 @@ Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gate
|
||||
- VectorStore 写入路径全面迁移。
|
||||
- 完整 LLM-based AIOps verifier。
|
||||
|
||||
后续 Agent 拆分、Skill/Playbook、MCP 工具协议化和进程隔离等方向见 [evolution-roadmap.md](evolution-roadmap.md)。
|
||||
后续 Agent 拆分、Playbook 版本化增强、MCP 工具协议化和进程隔离等方向见 [evolution-roadmap.md](evolution-roadmap.md)。
|
||||
|
||||
## 10. 关键代码索引
|
||||
|
||||
| 能力 | 代码 |
|
||||
|---|---|
|
||||
| Chat 入口与编排 | `ChatController`, `ChatService` |
|
||||
| Chat StateGraph runtime | `ChatDiagnosisGraphRuntime`, `DiagnosisGraphFactory`, `DiagnosisGraphRouter`, `DiagnosisGraphState` |
|
||||
| Chat Node assembly/config isolation | `DiagnosisRealGraphActionsFactory`, `ReactAgentDiagnosisInvoker` |
|
||||
| Graph result/audit mapping | `DiagnosisGraphResultMapper`, `DiagnosisOrchestrationTraceBuilder` |
|
||||
| AIOps 入口与编排 | `ChatController.aiOps`, `AiOpsService` |
|
||||
| AIOps 规则验证 | `AiOpsRuleEvaluationService` |
|
||||
| 知识库工具 | `LookupKnowledgeTool` |
|
||||
@@ -406,4 +462,5 @@ Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gate
|
||||
| Spring AI VectorStore 配置辅助 | `SpringAiVectorStoreSidecarService` |
|
||||
| Trace 聚合 | `DiagnosisTraceService` |
|
||||
| 工具调用记录 | `ToolInvocationRecorder` |
|
||||
| Executor 引用验真与投影 | `GatekeeperNode`, `ExecutorGatekeeperService`, `VerifiedInputNode` |
|
||||
| self_evaluation 合并 | `SelfEvaluationMergeService` |
|
||||
|
||||
+96
-149
@@ -1,78 +1,73 @@
|
||||
# 数据模型总览
|
||||
|
||||
**更新日期**:2026-07-05
|
||||
**更新日期**:2026-07-20
|
||||
**状态**:当前可运行架构
|
||||
|
||||
## 1. 定位
|
||||
|
||||
本文从架构角度说明当前 MVP 的核心数据模型。详细字段仍以 Flyway migration 和 `mvp/tables/` 为准。
|
||||
本文从架构角度说明当前 MVP 的核心数据模型。详细字段以 Flyway migration、实体类和 `mvp/tables/` 为准。
|
||||
|
||||
核心数据分三组:
|
||||
|
||||
- 诊断 Trace:`diagnosis_session`、`agent_step`、`tool_invocation`
|
||||
- 会话与诊断 Trace:`chat_session`、`diagnosis_run`、`agent_step`、`tool_invocation`
|
||||
- 知识库:`api_document`、`knowledge_domain`、Milvus/Zilliz metadata
|
||||
- 反馈沉淀:`case_library`
|
||||
|
||||
当前运行时只使用 `chat_session + diagnosis_run`;旧 `diagnosis_session` 表不属于当前版本契约。
|
||||
|
||||
## 2. 总体关系
|
||||
|
||||
```mermaid
|
||||
erDiagram
|
||||
diagnosis_session ||--o{ agent_step : has
|
||||
diagnosis_session ||--o{ tool_invocation : has
|
||||
diagnosis_session ||--o| case_library : creates_when_useful
|
||||
chat_session ||--o{ diagnosis_run : owns
|
||||
diagnosis_run ||--o{ agent_step : has
|
||||
diagnosis_run ||--o{ tool_invocation : has
|
||||
diagnosis_run ||--o| case_library : creates_when_useful
|
||||
api_document ||--o{ milvus_chunk : indexed_as
|
||||
knowledge_domain ||--o{ api_document : groups
|
||||
|
||||
diagnosis_session {
|
||||
chat_session {
|
||||
bigint id
|
||||
varchar session_id
|
||||
varchar status
|
||||
int message_pair_count
|
||||
datetime last_active_at
|
||||
datetime expires_at
|
||||
}
|
||||
|
||||
diagnosis_run {
|
||||
bigint id
|
||||
varchar run_id
|
||||
varchar session_id
|
||||
text query
|
||||
varchar status
|
||||
varchar agent_flow
|
||||
longtext answer
|
||||
json self_evaluation
|
||||
json orchestration_trace
|
||||
varchar feedback
|
||||
}
|
||||
|
||||
agent_step {
|
||||
bigint id
|
||||
varchar session_id
|
||||
varchar run_id
|
||||
int step_index
|
||||
varchar agent_name
|
||||
text model_input
|
||||
text model_output
|
||||
text thought
|
||||
boolean has_tool_call
|
||||
}
|
||||
|
||||
tool_invocation {
|
||||
bigint id
|
||||
varchar session_id
|
||||
varchar run_id
|
||||
bigint step_id
|
||||
varchar tool_name
|
||||
json input_params
|
||||
text output_preview
|
||||
varchar retrieval_layer
|
||||
json retrieval_details
|
||||
varchar relevance_level
|
||||
varchar dedup_reason
|
||||
}
|
||||
|
||||
api_document {
|
||||
bigint id
|
||||
varchar doc_id
|
||||
varchar file_name
|
||||
varchar file_path
|
||||
varchar status
|
||||
int chunk_count
|
||||
text metadata
|
||||
}
|
||||
|
||||
knowledge_domain {
|
||||
bigint id
|
||||
varchar domain_id
|
||||
varchar description
|
||||
text when_to_retrieve
|
||||
int document_count
|
||||
}
|
||||
|
||||
case_library {
|
||||
@@ -84,123 +79,100 @@ erDiagram
|
||||
text root_cause
|
||||
text solution
|
||||
}
|
||||
|
||||
milvus_chunk {
|
||||
varchar id
|
||||
text content
|
||||
json metadata
|
||||
vector vector
|
||||
}
|
||||
```
|
||||
|
||||
说明:Milvus/Zilliz collection 不是 MySQL 表,图中的 `milvus_chunk` 是逻辑模型。
|
||||
|
||||
## 3. 诊断 Trace 模型
|
||||
## 3. 会话与运行模型
|
||||
|
||||
### diagnosis_session
|
||||
### chat_session
|
||||
|
||||
会话级主记录。
|
||||
|
||||
关键字段:
|
||||
`chat_session` 是会话目录表,保存 `sessionId` 的元数据:
|
||||
|
||||
| 字段 | 说明 |
|
||||
|---|---|
|
||||
| `session_id` | 外部关联键,Trace 和 Feedback 都使用它 |
|
||||
| `query` | 用户原始问题或 AIOps 输入摘要 |
|
||||
| `status` | 执行状态 |
|
||||
| `session_id` | 外部会话 ID,用于多轮上下文和 run 列表 |
|
||||
| `status` | 会话目录状态 |
|
||||
| `message_pair_count` | Redis 对话轮次数快照 |
|
||||
| `last_active_at` | 最近活跃时间 |
|
||||
| `expires_at` | 可为空的目录 TTL 元数据 |
|
||||
|
||||
它不保存完整对话历史,正文消息仍由 Redis `SessionContext.messageHistory` 管理。
|
||||
|
||||
### diagnosis_run
|
||||
|
||||
`diagnosis_run` 是一次可回放诊断执行的主记录:
|
||||
|
||||
| 字段 | 说明 |
|
||||
|---|---|
|
||||
| `run_id` | 运行 ID,格式为 `run-` + UUID |
|
||||
| `session_id` | 所属 `chat_session.session_id` |
|
||||
| `query` | 本次 Chat 问题或 AIOps 告警摘要 |
|
||||
| `status` | 本次执行状态 |
|
||||
| `agent_flow` | `CHAT` / `AI_OPS` |
|
||||
| `answer` | 最终答复或告警报告 |
|
||||
| `self_evaluation` | rule/verifier/aiops 自评估容器 |
|
||||
| `feedback` | 用户反馈 |
|
||||
| `answer` | 本次运行最终答复或告警报告 |
|
||||
| `self_evaluation` | 本次运行的 rule/verifier/aiops 自评估容器 |
|
||||
| `orchestration_trace` | nullable JSON;复杂 Chat 的 StateGraph 路由摘要,非 StateGraph Run 可为空 |
|
||||
| `feedback` | 本次运行的用户反馈 |
|
||||
|
||||
同一个 `sessionId` 可以有多个 `runId`。Trace、反馈、评测和案例沉淀都应优先使用 `runId`,避免多轮同 session 下的数据混合。
|
||||
|
||||
## 4. Trace 明细模型
|
||||
|
||||
### agent_step
|
||||
|
||||
记录模型调用步骤。
|
||||
|
||||
用途:
|
||||
|
||||
- 回放 Agent 推理过程。
|
||||
- 查看 Planner / Executor / Verifier 的输入输出摘要。
|
||||
- 统计 step count、duration、token count。
|
||||
`agent_step` 记录模型调用步骤。新写入同时保留 `session_id` 和 `run_id`,其中 `run_id` 是回放边界。Trace 页面和评测应先按 `run_id` 隔离取数,展示顺序以 Trace API 返回顺序为准。
|
||||
|
||||
### tool_invocation
|
||||
|
||||
记录工具调用事实。
|
||||
`tool_invocation` 记录显式工具调用事实。`retrieval_details.evidence_refs` 是 Chat 证据链路的关键字段:
|
||||
|
||||
用途:
|
||||
|
||||
- 给 Trace API 展示证据。
|
||||
- 给 Verifier 构造 `tool_trace_summary`。
|
||||
- 给 `EvaluationService` 计算 evidence score。
|
||||
- 给 RAG eval 和人工排查提供检索细节。
|
||||
|
||||
## 4. 知识库模型
|
||||
|
||||
### api_document
|
||||
|
||||
MySQL 中的文档元数据表。
|
||||
|
||||
职责:
|
||||
|
||||
- 管理上传文件。
|
||||
- 保存 file hash,用于去重。
|
||||
- 记录索引状态和 chunk 数量。
|
||||
- 保存 frontmatter JSON。
|
||||
|
||||
### knowledge_domain
|
||||
|
||||
领域级元数据。
|
||||
|
||||
职责:
|
||||
|
||||
- 按 category 聚合文档。
|
||||
- 存储领域描述。
|
||||
- 存储 `when_to_retrieve`,辅助 Planner/Executor 判断什么时候检索该领域。
|
||||
|
||||
### Milvus/Zilliz metadata
|
||||
|
||||
向量 collection 中每个 chunk 的 metadata 主要包括:
|
||||
|
||||
```text
|
||||
docId
|
||||
_source
|
||||
chunkIndex
|
||||
totalChunks
|
||||
title
|
||||
breadcrumb
|
||||
category
|
||||
```json
|
||||
{
|
||||
"evidence_status": "supported",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.logs[0]",
|
||||
"text": "2026-07-08 23:05:28 ERROR order-service HikariPool-1 - Connection is not available..."
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
这些字段支撑:
|
||||
`$.no_evidence` 只表示“本次工具查询未检索到匹配证据”,不能被解释为“问题不存在”或“根因已排除”。
|
||||
|
||||
- category filter。
|
||||
- source 展示。
|
||||
- breadcrumb 上下文。
|
||||
- docId 删除和重建索引。
|
||||
- evidence block 构造。
|
||||
## 5. Run 级审计分层
|
||||
|
||||
## 5. 反馈沉淀模型
|
||||
一次 Run 的可审计信息分为三层,不能互相替代:
|
||||
|
||||
### case_library
|
||||
| 层次 | 存储 | 语义 |
|
||||
|---|---|---|
|
||||
| 执行明细 | `agent_step`、`tool_invocation` | 模型步骤、工具输入输出、检索细节和证据事实 |
|
||||
| 质量评估 | `diagnosis_run.self_evaluation` | rule/verifier/aiops 判断、Gatekeeper 审计、Prompt 版本和允许输出材料 |
|
||||
| 编排摘要 | `diagnosis_run.orchestration_trace` | StateGraph transitions、final node、termination reason、degraded、evidence retry count |
|
||||
|
||||
`useful` 反馈会触发 `CaseLibraryService.createFromSession`。
|
||||
`orchestration_trace` 只通过 exact Run 的 `run.orchestrationTrace` 暴露,Trace 响应不再包含兼容 `session` 投影。Run 状态仍只表达执行生命周期:安全 Fallback 是 `SUCCESS + degraded=true`,只有无法生成安全响应或未处理失败才是 `FAILED`。
|
||||
|
||||
## 6. 反馈沉淀模型
|
||||
|
||||
`useful` 反馈会触发 `CaseLibraryService.createFromRun`。
|
||||
|
||||
当前自动映射:
|
||||
|
||||
| 字段 | 来源 |
|
||||
|---|---|
|
||||
| `case_id` | UUID |
|
||||
| `diagnosis_id` | `diagnosis_session.session_id` |
|
||||
| `diagnosis_id` | `diagnosis_run.run_id` |
|
||||
| `source_type` | `AUTO` |
|
||||
| `fault_category` | 当前默认 `GENERAL` |
|
||||
| `title` | session query 前 100 字符 |
|
||||
| `root_cause` | session answer |
|
||||
| `solution` | session answer |
|
||||
| `title` | run query 前 100 字符 |
|
||||
| `root_cause` | run answer |
|
||||
| `solution` | run answer |
|
||||
| `created_by` | `system` |
|
||||
|
||||
## 6. self_evaluation 结构
|
||||
## 7. self_evaluation 结构
|
||||
|
||||
`diagnosis_session.self_evaluation` 是 JSON 容器:
|
||||
`diagnosis_run.self_evaluation` 是运行级 JSON 容器:
|
||||
|
||||
```json
|
||||
{
|
||||
@@ -210,49 +182,24 @@ category
|
||||
}
|
||||
```
|
||||
|
||||
边界:
|
||||
Chat 通常写入 `rule_evaluation` 和 `verifier_evaluation`;AIOps 写入 `aiops_rule_evaluation`。
|
||||
|
||||
- `rule_evaluation` 评估证据收集充分度。
|
||||
- `verifier_evaluation` 评估 Chat 答案关键事实是否有证据支撑。
|
||||
- `aiops_rule_evaluation` 评估 AIOps 报告是否聚焦告警并使用证据。
|
||||
|
||||
## 7. 数据写入时序
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant API as API
|
||||
participant Svc as ChatService/AiOpsService
|
||||
participant Session as diagnosis_session
|
||||
participant Agent as Agent
|
||||
participant Step as agent_step
|
||||
participant Tool as tool_invocation
|
||||
participant Eval as self_evaluation
|
||||
participant Feedback as case_library
|
||||
|
||||
API->>Svc: request
|
||||
Svc->>Session: create/update RUNNING
|
||||
Agent->>Step: before/after model
|
||||
Agent->>Tool: tool call record
|
||||
Svc->>Session: SUCCESS/FAILED + answer
|
||||
Svc->>Eval: merge evaluation
|
||||
API->>Svc: feedback useful
|
||||
Svc->>Feedback: create case
|
||||
```
|
||||
`orchestration_trace` 不放入该 JSON,避免把答案质量和 Graph 路由混成同一审计维度。
|
||||
|
||||
## 8. 当前边界和后续
|
||||
|
||||
当前边界:
|
||||
|
||||
- `agent_step.session_id` 和 `tool_invocation.session_id` 通过 sessionId 关联,不强制外键。
|
||||
- `tool_invocation.step_id` 可为空。
|
||||
- Milvus chunk 与 `api_document` 通过 metadata.docId 逻辑关联。
|
||||
- `case_library` 与 session 通过 `diagnosis_id=session_id` 关联。
|
||||
- `chat_session` 只存会话元数据,不存完整正文历史。
|
||||
- `diagnosis_run` 存一次运行的长期审计状态。
|
||||
- `agent_step.run_id` 和 `tool_invocation.run_id` 是 Trace、Verifier、Eval 的运行边界。
|
||||
- `diagnosis_run.orchestration_trace` 是 nullable Run-owned Graph 摘要;非 StateGraph Run 可以为空。
|
||||
- 当前实现主要使用逻辑关联,不依赖数据库外键。
|
||||
- `case_library.diagnosis_id` 是过渡字段,新值按 `run_id` 解释,旧值可能按 `session_id` 解释。
|
||||
- 当前 Java 运行时不存在 `diagnosis_session` entity/repository 或 fallback。
|
||||
|
||||
后续可增强:
|
||||
|
||||
1. 增加 run id,支持同 session 多次独立诊断。
|
||||
2. 强化 `tool_invocation.step_id` 关联。
|
||||
3. 将 evidence block 结构化保存。
|
||||
4. 将 `case_library` 的 rootCause/solution 从完整 answer 中结构化抽取。
|
||||
|
||||
1. 强化 `tool_invocation.step_id` 关联。
|
||||
2. 将 Gatekeeper 规则配置化时的规则元数据保存为可审计版本。
|
||||
3. 将 `case_library` 的 rootCause/solution 从完整 answer 中结构化抽取。
|
||||
|
||||
@@ -1,12 +1,12 @@
|
||||
# Agent 架构演进路线
|
||||
|
||||
**更新日期**:2026-07-05
|
||||
**更新日期**:2026-07-20
|
||||
**状态**:后续演进设计,不代表当前已实现
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
||||
|
||||
## 1. 为什么需要演进路线
|
||||
|
||||
旧版 `agent-architecture.md` 包含很多生产级设想:专科 SubAgent、Skill 体系、进程隔离、回退路由、MCP 工具协议化、进化引擎。它们不应作为当前 MVP 事实写入主架构,但可以作为后续扩展路线。
|
||||
旧版 `agent-architecture.md` 包含很多生产级设想:专科 SubAgent、完整 Skill 治理、进程隔离、跨 Agent 回退、MCP 工具协议化、进化引擎。当前已经落地 bounded StateGraph、安全 Fallback 和基础 Skill/Playbook 接入;本文件只描述它们之上的后续增强。
|
||||
|
||||
当前原则:
|
||||
|
||||
@@ -18,12 +18,12 @@
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
MVP["Current MVP: Planner + Executor + Verifier"] --> Split{"Executor 是否过载?"}
|
||||
MVP["Current MVP: bounded StateGraph + evidence gates"] --> Split{"Executor 是否过载?"}
|
||||
Split -->|是| SubAgents["专科 SubAgent"]
|
||||
Split -->|否| Keep["继续强化通用 Executor"]
|
||||
|
||||
SubAgents --> Skills["Skill / Playbook 体系"]
|
||||
Skills --> Fallback["回退路由"]
|
||||
SubAgents --> Skills["Skill / Playbook 版本化治理"]
|
||||
Skills --> Fallback["跨 SubAgent 回退路由"]
|
||||
Fallback --> Isolation["进程或 Pod 隔离"]
|
||||
|
||||
MVP --> ToolGrowth{"工具数量和来源是否增长?"}
|
||||
@@ -59,9 +59,9 @@ flowchart TD
|
||||
- 过早拆分会增加 Prompt、评测和 trace 分析成本。
|
||||
- 没有足够分类评测前,拆分可能只是移动复杂度。
|
||||
|
||||
## 4. Skill / Playbook 体系
|
||||
## 4. Skill / Playbook 版本化治理
|
||||
|
||||
旧版设计中的 Skill 可以在当前项目中演进为可版本化的诊断 Playbook。
|
||||
当前已经通过 Planner metadata selection + Executor `read_skill` 接入诊断 Playbook。下一阶段不是重新建设 Skill 入口,而是增加版本、评测、回退和审计治理。
|
||||
|
||||
```text
|
||||
fault_category
|
||||
@@ -73,14 +73,14 @@ fault_category
|
||||
-> evaluation checks
|
||||
```
|
||||
|
||||
优先落地方向:
|
||||
当前已覆盖的方向:
|
||||
|
||||
- AIOps 告警处理 Playbook。
|
||||
- 支付超时 Playbook。
|
||||
- MySQL 连接池风险 Playbook。
|
||||
- Redis timeout Playbook。
|
||||
|
||||
落地前提:
|
||||
后续增强前提:
|
||||
|
||||
- 每个 Playbook 至少有 3-5 个 eval case。
|
||||
- Playbook 失败时可以回退到通用 Executor。
|
||||
|
||||
@@ -1,78 +1,45 @@
|
||||
# Current Chat Agent Data Contracts
|
||||
# Chat Evidence Pipeline Contracts
|
||||
|
||||
**状态**:当前实现
|
||||
**日期**:2026-07-07
|
||||
**范围**:当前 Chat 复杂诊断链路的数据结构定义
|
||||
**状态**:当前实现
|
||||
**更新日期**:2026-07-17
|
||||
**范围**:Chat 复杂诊断链路中的 Planner、Executor、Gatekeeper、Verifier、Composer 数据契约
|
||||
|
||||
当前代码实现是三 Agent 顺序链路:
|
||||
当前 Chat 复杂诊断链路是:
|
||||
|
||||
```text
|
||||
chat_planner -> chat_executor -> chat_verifier
|
||||
chat_planner
|
||||
-> chat_executor
|
||||
-> GatekeeperNode / ExecutorGatekeeperService
|
||||
-> VerifiedInputNode
|
||||
-> chat_verifier
|
||||
-> chat_composer
|
||||
-> final answer
|
||||
```
|
||||
|
||||
对应 `ChatService.executeChatComplex(...)` 中的 `SequentialAgent`。
|
||||
设计原则:
|
||||
|
||||
- Planner 暂不输出 `scope_contract`。
|
||||
- Executor 只做证据收集和微观事实提炼,不生成最终用户答案。
|
||||
- Gatekeeper 在 Verifier 前做代码级引用真实性校验。
|
||||
- Verifier 判断 claim 是否能由已核验证据推出。
|
||||
- Composer 只表达 Verifier 允许输出的内容。
|
||||
|
||||
---
|
||||
|
||||
## 1. Workflow Input
|
||||
## 1. Planner
|
||||
|
||||
由 `ChatService.buildWorkflowInput(...)` 构造,传给 `chat_workflow`。
|
||||
|
||||
```text
|
||||
请按固定工作流完成本轮 Planner -> Executor -> Verifier。
|
||||
|
||||
--- 用户问题 ---
|
||||
{question}
|
||||
|
||||
--- retry_context ---
|
||||
{retry_context}
|
||||
|
||||
Verifier 完成后由外层代码读取 verifier_output 并决定最终用户输出。
|
||||
```
|
||||
|
||||
| 字段 | 来源 | 定义 |
|
||||
|---|---|---|
|
||||
| `question` | 用户输入 | 用户本轮原始问题 |
|
||||
| `retry_context` | ChatService | 第二轮补证据约束;首轮为空 |
|
||||
|
||||
---
|
||||
|
||||
## 2. chat_planner
|
||||
|
||||
### 2.1 Input
|
||||
|
||||
`chat_planner` 的输入来自 workflow input 和 system prompt 追加上下文。
|
||||
Planner 当前保持不变,输出 `planner_plan`:
|
||||
|
||||
```json
|
||||
{
|
||||
"question": "用户原始问题",
|
||||
"history": [],
|
||||
"available_knowledge_domains": "...",
|
||||
"skill_catalog": {},
|
||||
"retry_context": null
|
||||
"selected_skill": "diagnose-mysql-connection-pool",
|
||||
"selection_reason": "选择该 skill 的原因",
|
||||
"plan": ["步骤1", "步骤2"],
|
||||
"reasoning": "规划思路"
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 来源 | 定义 |
|
||||
|---|---|---|
|
||||
| `question` | workflow input | 用户原始问题 |
|
||||
| `history` | `ChatService.buildChatPlannerAgent(...)` | 对话历史,拼接到 planner system prompt |
|
||||
| `available_knowledge_domains` | `KnowledgeDomainService.buildKnowledgeMap()` | 可用知识域地图,拼接到 planner system prompt |
|
||||
| `skill_catalog` | `PlannerSkillMetadataHook` | Planner 可见的 skill name/description 元数据 |
|
||||
| `retry_context` | `ChatService` | Verifier 低置信后构造的补证据上下文 |
|
||||
|
||||
### 2.2 Output:`planner_plan`
|
||||
|
||||
当前 prompt 要求输出 JSON:
|
||||
|
||||
```json
|
||||
{
|
||||
"selected_skill": "匹配的 skill 名称;如果没有匹配则为 null",
|
||||
"selection_reason": "选择该 skill 的原因;如果没有匹配则说明不使用 skill",
|
||||
"plan": ["步骤1描述", "步骤2描述", "步骤3描述"],
|
||||
"reasoning": "规划思路说明"
|
||||
}
|
||||
```
|
||||
字段定义:
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
@@ -81,434 +48,398 @@ Verifier 完成后由外层代码读取 verifier_output 并决定最终用户输
|
||||
| `plan` | array | 给 Executor 的执行步骤 |
|
||||
| `reasoning` | string | 规划思路说明 |
|
||||
|
||||
运行态输出 key:
|
||||
当前边界:
|
||||
|
||||
```text
|
||||
planner_plan
|
||||
```
|
||||
- 不新增 `scope_contract`。
|
||||
- 不要求 Planner 显式列出 forbidden actions。
|
||||
- 窄范围控制先由 Executor Prompt 约束,后续如仍不稳定再引入 Planner contract。
|
||||
|
||||
---
|
||||
|
||||
## 3. chat_executor
|
||||
## 2. Executor
|
||||
|
||||
### 3.1 Input
|
||||
Executor 输出 `executor_evidence_v2`。它不是最终答复,而是给 Gatekeeper、Verifier、Composer 使用的结构化诊断材料。
|
||||
|
||||
`chat_executor` 接收前序 `planner_plan`,并通过 system prompt 获得历史、skill 读取约束、retry 约束和工具权限。
|
||||
### 2.1 输出结构
|
||||
|
||||
```json
|
||||
{
|
||||
"planner_plan": {},
|
||||
"history": [],
|
||||
"retry_context": null,
|
||||
"tool_permissions": {
|
||||
"method_tools": ["dateTimeTools", "lookupKnowledgeTool", "queryMetricsTools", "queryLogsTools"],
|
||||
"tool_callbacks": []
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 来源 | 定义 |
|
||||
|---|---|---|
|
||||
| `planner_plan` | `chat_planner` | Planner 输出的计划 |
|
||||
| `history` | `ChatService.buildChatExecutorAgent(...)` | 对话历史,拼接到 executor system prompt |
|
||||
| `retry_context` | `ChatService` | 本轮补证据约束 |
|
||||
| `method_tools` | `ChatService.buildMethodToolsArray()` | Executor 可直接调用的本地工具 |
|
||||
| `tool_callbacks` | `ToolCallback[]` | 框架发现或外部注入工具 |
|
||||
| `read_skill` | `SkillsAgentHook` | 当存在 skillRegistry 时,Executor 可读取 Planner 选中的 skill |
|
||||
|
||||
### 3.2 Output:`executor_feedback`
|
||||
|
||||
当前 `chat-executor-prompt.md` 要求输出一个 JSON 对象,即 `executor_evidence_v1`。
|
||||
|
||||
```json
|
||||
{
|
||||
"answer_version": "executor_evidence_v1",
|
||||
"diagnosis_summary": "1-2句话总结,仅包含有证据支撑的事实和证据边界",
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "root_cause",
|
||||
"claim_text": "事实断言或有限结论",
|
||||
"claim_type": "observation",
|
||||
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"source_type": "tool_trace",
|
||||
"source_id": "工具返回中的 evidence block id、trace_ref 或可定位标识",
|
||||
"tool_name": "lookup_knowledge/query_logs/query_metrics/read_skill 等",
|
||||
"source_invocation_ids": [],
|
||||
"evidence_excerpt": "从工具返回中摘取的原话、指标值、日志片段或关键数据"
|
||||
"source_id": "",
|
||||
"tool_name": "query_metrics",
|
||||
"source_invocation_id": 517,
|
||||
"raw_path": "$.alerts[0]",
|
||||
"evidence_excerpt": "HighCPUUsage, service=payment-service, state=firing, current=92%, duration=25m"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [
|
||||
{
|
||||
"hypothesis_text": "未被证实但值得排查的方向",
|
||||
"basis": "它基于哪些已知证据或为什么只是推测",
|
||||
"needed_evidence": ["需要补充的证据"]
|
||||
}
|
||||
],
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "建议动作",
|
||||
"reason": "为什么建议做这个动作",
|
||||
"evidence_bindings": []
|
||||
}
|
||||
],
|
||||
"missing_info": [
|
||||
"导致无法确认完整根因的证据缺口"
|
||||
],
|
||||
"user_facing_answer": "面向用户的中文回答。必须与 claims/hypotheses/recommended_actions/missing_info 一致。"
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [],
|
||||
"missing_info": []
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
字段定义:
|
||||
|
||||
| 字段 | 类型 | 必填 | 定义 |
|
||||
|---|---|---:|---|
|
||||
| `answer_version` | string | 是 | 固定为 `executor_evidence_v2` |
|
||||
| `claims` | array | 是 | Executor 提出的待验证事实断言 |
|
||||
| `claims[].claim_id` | string | 是 | claim 标识 |
|
||||
| `claims[].claim_type` | string | 是 | `observation`、`negative_observation`、`symptom`、`root_cause` 等;窄范围任务只允许前两者 |
|
||||
| `claims[].claim_text` | string | 是 | 事实断言文本 |
|
||||
| `claims[].support_level` | string | 是 | `direct` 或 `indirect` |
|
||||
| `claims[].evidence_bindings` | array | 是 | 支撑该 claim 的证据绑定,不能为空 |
|
||||
| `evidence_bindings[].source_type` | string | 否 | 当前通常为 `tool_trace` |
|
||||
| `evidence_bindings[].source_id` | string | 否 | 兼容字段,不作为精确引用主键 |
|
||||
| `evidence_bindings[].tool_name` | string | 是 | `query_logs`、`query_metrics`、`lookup_knowledge` 等 |
|
||||
| `evidence_bindings[].source_invocation_id` | number/null | 是 | 来源 `tool_invocation.id`;缺失时 Gatekeeper 只在能唯一匹配时回填 |
|
||||
| `evidence_bindings[].raw_path` | string | 是 | 工具返回中的稳定定位路径 |
|
||||
| `evidence_bindings[].evidence_excerpt` | string | 是 | 工具返回中的原文片段或系统抽取的最小证据文本 |
|
||||
| `hypotheses` | array | 是 | 未证实但值得排查的方向,不是 confirmed fact |
|
||||
| `recommended_actions` | array | 是 | 下一步动作;本期只允许证据收集或继续排查动作 |
|
||||
| `missing_info` | array | 是 | 无法确认结论所缺少的证据 |
|
||||
|
||||
禁止字段:
|
||||
|
||||
- `diagnosis_summary`
|
||||
- `user_facing_answer`
|
||||
- `source_invocation_ids` 作为主引用字段
|
||||
|
||||
### 2.2 raw_path
|
||||
|
||||
当前支持的精确路径:
|
||||
|
||||
| 工具 | 正向证据路径 | 负向证据路径 |
|
||||
|---|---|---|
|
||||
| `answer_version` | string | 当前固定为 `executor_evidence_v1` |
|
||||
| `diagnosis_summary` | string | 有证据边界的简短诊断摘要 |
|
||||
| `claims` | array | 已证实或有明确间接支撑的事实断言 |
|
||||
| `claims[].claim_id` | string | claim 标识 |
|
||||
| `claims[].claim_type` | string | claim 类型,例如 `root_cause`、`symptom`、`impact` |
|
||||
| `claims[].claim_text` | string | 事实断言文本 |
|
||||
| `claims[].support_level` | string | `direct` 或 `indirect` |
|
||||
| `claims[].evidence_bindings` | array | 支撑 claim 的证据绑定,不能为空 |
|
||||
| `evidence_bindings[].source_type` | string | 证据来源类型,例如 `tool_trace` |
|
||||
| `evidence_bindings[].source_id` | string | evidence block id、trace_ref 或其它定位标识 |
|
||||
| `evidence_bindings[].tool_name` | string | 来源工具名 |
|
||||
| `evidence_bindings[].source_invocation_ids` | array | 来源 `tool_invocation.id` |
|
||||
| `evidence_bindings[].evidence_excerpt` | string | 工具返回中的原话、指标值、日志片段或关键数据 |
|
||||
| `hypotheses` | array | 未证实但值得排查的方向 |
|
||||
| `hypotheses[].hypothesis_text` | string | 假设文本 |
|
||||
| `hypotheses[].basis` | string | 假设依据和未证实原因 |
|
||||
| `hypotheses[].needed_evidence` | array | 确认该假设还需要的证据 |
|
||||
| `recommended_actions` | array | 建议动作 |
|
||||
| `recommended_actions[].action_text` | string | 建议动作文本 |
|
||||
| `recommended_actions[].reason` | string | 建议原因 |
|
||||
| `recommended_actions[].evidence_bindings` | array | 建议动作关联证据,可为空 |
|
||||
| `missing_info` | array | 证据缺口 |
|
||||
| `user_facing_answer` | string | 候选用户答案,PASS 时由 ChatService 提取输出 |
|
||||
| `query_metrics` | `$.alerts[i]` | `$.no_evidence` |
|
||||
| `query_logs` | `$.logs[i]` | `$.no_evidence` |
|
||||
| `lookup_knowledge` | `$.evidence_blocks[i]` | `$.no_evidence` |
|
||||
|
||||
运行态输出 key:
|
||||
约束:
|
||||
|
||||
```text
|
||||
executor_feedback
|
||||
```
|
||||
- `raw_path` 必须指向数组条目或 `$.no_evidence`。
|
||||
- 禁止字段级子路径,例如 `$.alerts[0].state`、`$.logs[0].message`。
|
||||
- 同一条工具数组项只能绑定一次;多个字段应合并进同一个 `evidence_excerpt`。
|
||||
|
||||
---
|
||||
### 2.3 negative_observation
|
||||
|
||||
## 4. chat_verifier
|
||||
|
||||
### 4.1 Input
|
||||
|
||||
`VerifierInputHook` 会在 Verifier 调用前替换消息历史,构造显式 JSON payload。
|
||||
当工具明确返回 no-hit / no-evidence 时,Executor 可以输出 `negative_observation`:
|
||||
|
||||
```json
|
||||
{
|
||||
"original_query": "用户原始问题",
|
||||
"executor_final_answer": "{...executor_feedback raw text...}",
|
||||
"executor_structured_output": {},
|
||||
"executor_output_parse_status": {
|
||||
"status": "valid",
|
||||
"detail": "parsed executor evidence contract"
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "negative_observation",
|
||||
"claim_text": "当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"source_invocation_id": 517,
|
||||
"raw_path": "$.no_evidence",
|
||||
"evidence_excerpt": "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
语义边界:
|
||||
|
||||
- `$.no_evidence` 只表示“该工具对当前查询返回无匹配证据”。
|
||||
- 不表示“问题绝对不存在”。
|
||||
- 不表示“根因被排除”。
|
||||
- 不表示“系统已经健康”。
|
||||
- `negative_observation` 的 `evidence_bindings` 只能绑定 `$.no_evidence`,不能混绑其它服务的正向日志。
|
||||
|
||||
### 2.4 窄范围任务
|
||||
|
||||
窄范围任务指用户只要求确认某个服务、告警、日志、错误、订单或时间窗口。
|
||||
|
||||
Executor 必须遵守:
|
||||
|
||||
- 只输出 `observation` / `negative_observation`。
|
||||
- claim 数量通常 1 条,最多 2 条。
|
||||
- claim 数量限制不限制 `evidence_bindings` 数量。
|
||||
- 不输出根因、风险、修复建议、经验推断。
|
||||
- 不把 Runbook / Skill / 知识库通用知识写成当前环境事实。
|
||||
- 精确查询返回 no-evidence 后,不得放宽关键词、删除服务名或扩大服务范围继续查。
|
||||
|
||||
---
|
||||
|
||||
## 3. Tool Invocation Evidence Refs
|
||||
|
||||
工具调用落库到 `tool_invocation`,其中 `retrieval_details.evidence_refs` 是 Gatekeeper 的主校验源。
|
||||
|
||||
### 3.1 正向证据
|
||||
|
||||
```json
|
||||
{
|
||||
"evidence_status": "supported",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.logs[0]",
|
||||
"text": "2026-07-08 23:05:28 ERROR order-service HikariPool-1 - Connection is not available..."
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
### 3.2 负向证据
|
||||
|
||||
```json
|
||||
{
|
||||
"evidence_status": "no_evidence",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.no_evidence",
|
||||
"text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
字段定义:
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `evidence_status` | string | `supported`、`no_evidence`、`deduped`、`failed` |
|
||||
| `evidence_refs[].raw_path` | string | 证据在工具返回中的稳定定位符 |
|
||||
| `evidence_refs[].text` | string | 系统抽取的最小证据文本,供 Gatekeeper 和 Verifier 使用 |
|
||||
|
||||
---
|
||||
|
||||
## 4. Gatekeeper
|
||||
|
||||
Gatekeeper 是 StateGraph 中的显式 Node,调用 `ExecutorGatekeeperService` 对当前 Run 的证据引用做代码级真实性校验。
|
||||
|
||||
### 4.1 输入
|
||||
|
||||
- `sessionId + runId`(来自 `RunnableConfig`,工具查询以 `runId` 为边界)
|
||||
- `executor_structured_output`
|
||||
- 当前 run 的 `tool_invocation`
|
||||
|
||||
### 4.2 输出
|
||||
|
||||
```json
|
||||
{
|
||||
"status": "pass",
|
||||
"severity": "none",
|
||||
"rule_set_version": "gatekeeper-rules-v1",
|
||||
"rules": [
|
||||
{
|
||||
"id": "evidence.raw_path",
|
||||
"description": "raw_path must exist in retrieval_details.evidence_refs",
|
||||
"enabled": true,
|
||||
"default_severity": "reject"
|
||||
}
|
||||
],
|
||||
"checked_bindings": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"tool_name": "query_logs",
|
||||
"source_invocation_id": 517,
|
||||
"raw_path": "$.no_evidence",
|
||||
"matched_text": "query_logs returned no evidence; ...",
|
||||
"status": "pass"
|
||||
}
|
||||
],
|
||||
"failed_rules": [],
|
||||
"warnings": [],
|
||||
"errors": []
|
||||
}
|
||||
```
|
||||
|
||||
字段定义:
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `status` | string | `pass` 或 `fail` |
|
||||
| `severity` | string | `none`、`low_confid`、`reject` |
|
||||
| `rule_set_version` | string | 当前加载的 Gatekeeper 规则集版本 |
|
||||
| `rules` | array | 已启用规则的轻量元数据摘要 |
|
||||
| `checked_bindings` | array | 每条证据绑定的校验结果 |
|
||||
| `failed_rules` | array | 失败规则 id |
|
||||
| `warnings` | array | 自动回填等非阻断信息 |
|
||||
| `errors` | array | 失败明细 |
|
||||
|
||||
校验规则:
|
||||
|
||||
- `answer_version` 必须是 `executor_evidence_v2`。
|
||||
- 不允许 `diagnosis_summary` / `user_facing_answer`。
|
||||
- 每个 claim 必须有非空 `evidence_bindings`。
|
||||
- `tool_name` 必须和真实 invocation 对齐。
|
||||
- `source_invocation_id` 必须存在;缺失时只在 `tool_name + raw_path + evidence_excerpt` 能唯一匹配真实 invocation 时回填。
|
||||
- `raw_path` 必须存在于 `retrieval_details.evidence_refs`。
|
||||
- `evidence_excerpt` 必须由 `evidence_refs[].text` 支撑。
|
||||
- `negative_observation` 只能绑定 `$.no_evidence`。
|
||||
|
||||
规则配置:
|
||||
|
||||
- 当前规则元数据位于 `src/main/resources/gatekeeper/gatekeeper-rules.json`。
|
||||
- 规则实现仍是确定性 Java 代码,不执行动态脚本。
|
||||
- 当前配置只承载规则 id、描述、默认 severity、启用状态和简单参数,例如 excerpt token overlap 阈值。
|
||||
|
||||
失败分级:
|
||||
|
||||
| 场景 | severity |
|
||||
|---|---|
|
||||
| 伪造 invocation id | `reject` |
|
||||
| tool_name 与 invocation 不匹配 | `reject` |
|
||||
| raw_path 不存在 | `reject` |
|
||||
| excerpt 与 matched_text 不匹配 | `reject` |
|
||||
| negative_observation 绑定正向日志 | `reject` |
|
||||
| 缺少 raw_path / invocation id 且无法唯一回填 | `low_confid` |
|
||||
| 旧 invocation 没有 `evidence_refs` | `low_confid` |
|
||||
|
||||
---
|
||||
|
||||
## 5. Verifier
|
||||
|
||||
`VerifiedInputNode` 只保留 Gatekeeper 检查通过的 claim/binding,并为 Verifier 构造最小输入:
|
||||
|
||||
```json
|
||||
{
|
||||
"diagnosis_context": {
|
||||
"query": "用户原始问题"
|
||||
},
|
||||
"tool_trace_summary": [],
|
||||
"verified_executor_output": {
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": []
|
||||
},
|
||||
"verified_evidence": [],
|
||||
"gatekeeper_audit": {},
|
||||
"verdict_ceiling": "PASS",
|
||||
"retry_context": null
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 来源 | 定义 |
|
||||
|---|---|---|
|
||||
| `original_query` | `VerifierContextHolder` | 用户原始问题 |
|
||||
| `executor_final_answer` | `VerifierContextHolder` 或上一条 AssistantMessage | Executor 原始输出文本 |
|
||||
| `executor_structured_output` | `VerifierInputHook.parseExecutorOutput(...)` | Executor 输出可解析且包含 `claims` 时的 JSON 对象;否则为 null |
|
||||
| `executor_output_parse_status.status` | `VerifierInputHook` | `valid` / `missing` / `malformed` |
|
||||
| `executor_output_parse_status.detail` | `VerifierInputHook` | 解析状态说明 |
|
||||
| `tool_trace_summary` | `ToolTraceSummaryService.buildVerifierTraceSummary(...)` | 基于真实 `tool_invocation` 构建的证据索引 |
|
||||
| `retry_context` | `VerifierContextHolder` | 当前补证据上下文 |
|
||||
Verifier 职责:
|
||||
|
||||
### 4.2 `tool_trace_summary`
|
||||
- 不调用工具。
|
||||
- 不读 skill。
|
||||
- 不逐字核验 excerpt 真伪;这由 Gatekeeper 完成。
|
||||
- 只判断 `claim_text` 是否能由已核验的 `evidence_excerpt` 推出。
|
||||
- 不读取 Executor 原始答复或完整工具 Trace,只读取 verified projection。
|
||||
- 对 `gatekeeper_result.severity=reject` 不得输出 `PASS`。
|
||||
- 对 `gatekeeper_result.severity=low_confid` 不得输出 `PASS`。
|
||||
|
||||
`ToolTraceSummaryService` 聚合 evidence tools:
|
||||
|
||||
```text
|
||||
lookup_knowledge, query_logs, query_metrics, query_order
|
||||
```
|
||||
|
||||
输出项结构:
|
||||
|
||||
```json
|
||||
{
|
||||
"trace_ref": "trace-1",
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"input_summary": "query=payment-service timeout",
|
||||
"output_summary": "log_evidence: ...",
|
||||
"evidence_level": "direct",
|
||||
"topic_domain": "general",
|
||||
"source_invocation_ids": [394],
|
||||
"invocation_count": 1,
|
||||
"failed_invocation_count": 0,
|
||||
"no_hit_invocation_count": 0,
|
||||
"query_samples": ["payment-service timeout"],
|
||||
"retrieval_layers": [],
|
||||
"relevance_levels": [],
|
||||
"source_documents": []
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `trace_ref` | string | Verifier 可引用的证据摘要编号 |
|
||||
| `tool_name` | string | 聚合后的工具名 |
|
||||
| `success` | boolean | 是否存在可用证据 |
|
||||
| `input_summary` | string | 工具输入摘要 |
|
||||
| `output_summary` | string | 工具输出摘要 |
|
||||
| `evidence_level` | string | `direct` / `indirect` / `none` |
|
||||
| `topic_domain` | string | 主题域,优先来自 `retrieval_details.retrieved_domains` |
|
||||
| `source_invocation_ids` | array | 聚合的 `tool_invocation.id` |
|
||||
| `invocation_count` | number | 聚合调用次数 |
|
||||
| `failed_invocation_count` | number | 失败调用次数 |
|
||||
| `no_hit_invocation_count` | number | 无证据或去重调用次数 |
|
||||
| `query_samples` | array | 查询样例 |
|
||||
| `retrieval_layers` | array | 检索层级 |
|
||||
| `relevance_levels` | array | 相关性等级 |
|
||||
| `source_documents` | array | 来源文档标签 |
|
||||
|
||||
### 4.3 Output:`verifier_output`
|
||||
|
||||
当前 `chat-verifier-prompt.md` 要求输出:
|
||||
输出:
|
||||
|
||||
```json
|
||||
{
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 0.8,
|
||||
"critical_fact_count": 2,
|
||||
"facts_checked": [
|
||||
"groundedness_score": 1.0,
|
||||
"critical_fact_count": 1,
|
||||
"claim_checks": [],
|
||||
"facts_checked": [],
|
||||
"rationale": "..."
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6. Composer
|
||||
|
||||
Composer 位于 Verifier 之后,输入是 ChatService 过滤后的允许表达材料。
|
||||
|
||||
输入概念:
|
||||
|
||||
| 字段 | 定义 |
|
||||
|---|---|
|
||||
| `original_query` | 用户原始问题 |
|
||||
| `verdict` | `PASS` / `LOW_CONFID` / `REJECT` |
|
||||
| `allowed_claims` | Verifier 允许表达的 claims |
|
||||
| `allowed_hypotheses` | Verifier 允许表达的假设 |
|
||||
| `missing_info` | 证据缺口 |
|
||||
| `recommended_actions` | 允许表达的建议动作 |
|
||||
| `rationale` | Verifier 判定理由 |
|
||||
|
||||
输出:
|
||||
|
||||
```json
|
||||
{
|
||||
"answer_summary": "一句话概括",
|
||||
"recommended_actions": [
|
||||
{
|
||||
"fact": "ERR_TIMEOUT 表示请求超时",
|
||||
"is_critical": true,
|
||||
"verification": "direct_evidence",
|
||||
"detail": "知识库文档明确给出该错误码定义",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"trace_ref": "trace-1",
|
||||
"tool_name": "lookup_knowledge",
|
||||
"topic_domain": "api",
|
||||
"source_invocation_ids": [101, 104],
|
||||
"note": "trace-1 的文档摘要直接给出错误码定义"
|
||||
}
|
||||
]
|
||||
"action_text": "下一步动作",
|
||||
"reason": "原因"
|
||||
}
|
||||
],
|
||||
"rationale": "所有关键事实均有支撑,且至少一条具有直接证据"
|
||||
"user_facing_answer": "最终给用户看的中文答案"
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `verdict` | string | `PASS` / `LOW_CONFID` / `REJECT` |
|
||||
| `groundedness_score` | number | 关键事实证据支撑评分 |
|
||||
| `critical_fact_count` | number | `facts_checked` 中 `is_critical=true` 的数量 |
|
||||
| `facts_checked` | array | Verifier 校验过的事实列表 |
|
||||
| `facts_checked[].fact` | string | 被校验事实 |
|
||||
| `facts_checked[].is_critical` | boolean | 是否关键事实 |
|
||||
| `facts_checked[].verification` | string | `direct_evidence` / `indirect_support` / `no_evidence` / `contradicted` |
|
||||
| `facts_checked[].detail` | string | 校验说明 |
|
||||
| `facts_checked[].evidence_refs` | array | 证据引用 |
|
||||
| `evidence_refs[].trace_ref` | string | 引用的 `tool_trace_summary.trace_ref` |
|
||||
| `evidence_refs[].tool_name` | string | 引用工具 |
|
||||
| `evidence_refs[].topic_domain` | string | 引用主题域 |
|
||||
| `evidence_refs[].source_invocation_ids` | array | 引用的 `tool_invocation.id` |
|
||||
| `evidence_refs[].note` | string | 引用说明 |
|
||||
| `rationale` | string | verdict 判定理由 |
|
||||
表达边界:
|
||||
|
||||
运行态输出 key:
|
||||
|
||||
```text
|
||||
verifier_output
|
||||
```
|
||||
- Composer 不补事实、不补根因、不调用工具。
|
||||
- 只表达 `allowed_claims`、`allowed_hypotheses`、`missing_info`、`recommended_actions`。
|
||||
- 当 claim 是 `negative_observation` 或证据来自 `$.no_evidence` 时,只能表达“当前查询未检索到 / 本次检索未发现匹配证据”。
|
||||
- 禁止表达“问题不存在”“已排除该问题”“确认没有”“日志层面已排除”等过度结论。
|
||||
|
||||
---
|
||||
|
||||
## 5. VerifierDecision
|
||||
## 7. Trace Persistence
|
||||
|
||||
`ChatService.parseVerifierDecision(...)` 将 `verifier_output` 解析为内部 record:
|
||||
|
||||
```json
|
||||
{
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundednessScore": 0.5,
|
||||
"criticalFactCount": 2,
|
||||
"factsChecked": [],
|
||||
"rationale": "证据不足",
|
||||
"round": 1
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `verdict` | string | Verifier verdict |
|
||||
| `groundednessScore` | number | groundedness score |
|
||||
| `criticalFactCount` | number | 关键事实数量 |
|
||||
| `factsChecked` | array | 解析后的 facts_checked |
|
||||
| `rationale` | string | 判定理由 |
|
||||
| `round` | number | 当前验证轮次 |
|
||||
|
||||
---
|
||||
|
||||
## 6. retry_context
|
||||
|
||||
当 `LOW_CONFID` 且满足重试条件时,`ChatService.buildRetryContext(...)` 构造:
|
||||
|
||||
```json
|
||||
{
|
||||
"round": 1,
|
||||
"missing_evidence_facts": [
|
||||
"某关键事实:缺少直接证据"
|
||||
],
|
||||
"instruction": "仅补充以上断言相关证据,不要重复已完成检索"
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `round` | number | 触发 retry 的轮次 |
|
||||
| `missing_evidence_facts` | array | 来自 Verifier 的证据缺口 |
|
||||
| `instruction` | string | 补证据约束 |
|
||||
|
||||
---
|
||||
|
||||
## 7. diagnosis_session.self_evaluation.verifier_evaluation
|
||||
|
||||
`ChatService.persistVerifierEvaluation(...)` 将 Verifier 结果合并进 `diagnosis_session.self_evaluation`。
|
||||
`diagnosis_run.self_evaluation.verifier_evaluation` 持久化:
|
||||
|
||||
```json
|
||||
{
|
||||
"verifier_evaluation": {
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.5,
|
||||
"critical_fact_count": 2,
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 1.0,
|
||||
"critical_fact_count": 1,
|
||||
"claim_checks": [],
|
||||
"facts_checked": [],
|
||||
"rationale": "证据不足",
|
||||
"rationale": "...",
|
||||
"round": 1,
|
||||
"traceability_version": "v1",
|
||||
"executor_output_parse_status": {
|
||||
"status": "valid",
|
||||
"detail": "parsed executor evidence contract"
|
||||
},
|
||||
"executor_output_parse_status": {},
|
||||
"executor_structured_output": {},
|
||||
"tool_trace_summary": []
|
||||
"gatekeeper_result": {
|
||||
"rule_set_version": "gatekeeper-rules-v1"
|
||||
},
|
||||
"composer_output": {},
|
||||
"verified_evidence": []
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `verifier_evaluation.verdict` | string | Verifier verdict |
|
||||
| `verifier_evaluation.groundedness_score` | number | groundedness score |
|
||||
| `verifier_evaluation.critical_fact_count` | number | 关键事实数量 |
|
||||
| `verifier_evaluation.facts_checked` | array | 校验事实列表 |
|
||||
| `verifier_evaluation.rationale` | string | 判定理由 |
|
||||
| `verifier_evaluation.round` | number | 验证轮次 |
|
||||
| `verifier_evaluation.traceability_version` | string | 当前固定为 `v1` |
|
||||
| `verifier_evaluation.executor_output_parse_status` | object | Executor 输出解析状态 |
|
||||
| `verifier_evaluation.executor_structured_output` | object/null | 解析后的 Executor 结构化输出 |
|
||||
| `verifier_evaluation.tool_trace_summary` | array | Verifier 使用的工具证据索引 |
|
||||
Trace API 可用于回放:
|
||||
|
||||
- Executor 输出了哪些 claim。
|
||||
- 每个 claim 引用了哪些 `source_invocation_id + raw_path + evidence_excerpt`。
|
||||
- Gatekeeper 是否通过、是否自动回填。
|
||||
- Verifier 如何判断可推导性。
|
||||
- Composer 最终如何表达给用户。
|
||||
- `run.orchestrationTrace` 如何经过条件边、有限重试并终止。
|
||||
|
||||
当前版本不生成、读取或展示 `verifier_evaluation.tool_trace_summary`;Verifier 只消费 verified projection。
|
||||
|
||||
---
|
||||
|
||||
## 8. Final Answer Rendering
|
||||
## 8. 当前已验证样例
|
||||
|
||||
ChatService 根据 Verifier verdict 决定最终 `diagnosis_session.answer`。
|
||||
|
||||
| Verdict | 当前行为 |
|
||||
|---|---|
|
||||
| `PASS` | 优先提取 `executor_feedback.user_facing_answer`;提取失败则使用 executor 原文 |
|
||||
| `LOW_CONFID` | 输出低置信模板:已确认信息、当前缺口、建议下一步 |
|
||||
| `REJECT` | 输出降级模板:已确认信息、证据缺口、建议下一步 |
|
||||
|
||||
低置信模板使用:
|
||||
|
||||
```text
|
||||
以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。
|
||||
|
||||
已确认信息:
|
||||
- ...
|
||||
|
||||
当前缺口:
|
||||
- ...
|
||||
|
||||
建议下一步:
|
||||
- ...
|
||||
```
|
||||
|
||||
拒绝模板使用:
|
||||
|
||||
```text
|
||||
当前无法基于已获取证据生成可靠结论。
|
||||
|
||||
已确认信息:
|
||||
- ...
|
||||
|
||||
证据缺口:
|
||||
- ...
|
||||
|
||||
建议下一步:
|
||||
- ...
|
||||
```
|
||||
| 场景 | sessionId | 结果 |
|
||||
|---|---|---|
|
||||
| HighCPUUsage 窄范围正向确认 | `iss008-narrow-highcpu-rerun-20260708-215510` | `PASS`,1 条 `observation`,无越界 claim |
|
||||
| HikariCP negative_observation | `iss009-hikari-negative-latest-20260708-232428` | `PASS`,`raw_path=$.no_evidence`,无过度表达 |
|
||||
|
||||
---
|
||||
|
||||
## 9. Trace Persistence Data
|
||||
## 9. 仍需记录或后续补强
|
||||
|
||||
### 9.1 diagnosis_session
|
||||
当前架构文档已记录主链路、数据契约和语义边界。后续如果继续实现,建议再补:
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `session_id` | string | 会话 id |
|
||||
| `query` | text | 用户问题 |
|
||||
| `status` | string | 会话状态 |
|
||||
| `agent_flow` | string | 当前 Chat 链路为 `CHAT` |
|
||||
| `total_duration_ms` | number | 总耗时 |
|
||||
| `total_token_count` | number | 总 token |
|
||||
| `step_count` | number | agent step 数 |
|
||||
| `tool_call_count` | number | tool invocation 数 |
|
||||
| `answer` | longtext | 最终用户答案 |
|
||||
| `self_evaluation` | json | 包含 verifier_evaluation |
|
||||
| `feedback` | string | 用户反馈 |
|
||||
|
||||
### 9.2 agent_step
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `session_id` | string | 会话 id |
|
||||
| `step_index` | number | 步骤序号 |
|
||||
| `agent_name` | string | `planner` / `executor` / `verifier` |
|
||||
| `model_input` | text | 模型输入摘要 |
|
||||
| `model_output` | text | 模型输出摘要 |
|
||||
| `thought` | text | hook 记录的摘要信息 |
|
||||
| `has_tool_call` | boolean | 是否包含工具调用 |
|
||||
| `duration_ms` | number | 模型调用耗时 |
|
||||
| `token_count` | number | token 数 |
|
||||
|
||||
### 9.3 tool_invocation
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `id` | number | 工具调用 id |
|
||||
| `session_id` | string | 会话 id |
|
||||
| `step_id` | number | 对应 agent_step id |
|
||||
| `tool_name` | string | 工具名 |
|
||||
| `input_params` | json | 工具输入参数 |
|
||||
| `output_preview` | text | 工具输出预览 |
|
||||
| `output_length` | number | 原始输出长度 |
|
||||
| `retrieval_layer` | string | 检索层 |
|
||||
| `l0_match_count` | number | L0 命中数 |
|
||||
| `l1_match_count` | number | L1 命中数 |
|
||||
| `is_truncated` | boolean | 输出是否截断 |
|
||||
| `relevance_level` | string | 相关性等级 |
|
||||
| `dedup_reason` | string | 去重原因 |
|
||||
| `retrieval_details` | json | 检索细节 |
|
||||
| `duration_ms` | number | 工具耗时 |
|
||||
| `success` | boolean | 是否成功 |
|
||||
| `error_message` | text | 错误信息 |
|
||||
1. Planner `scope_contract` 的 ADR:只有当 Prompt-first 无法稳定控制越界时再引入。
|
||||
2. 更完整的 Gatekeeper 规则配置化:当前只有本地轻量 metadata/catalog,后续如果做索引层、元数据层、远程规则层,需要单独记录加载顺序、变更审批和回滚策略。
|
||||
3. Prompt version 记录:当前 prompt 变更没有版本号,后续如果需要回滚和对比,应记录 prompt version。
|
||||
|
||||
@@ -1,15 +1,15 @@
|
||||
# 反馈与自评估架构
|
||||
|
||||
**更新日期**:2026-07-05
|
||||
**状态**:当前可运行架构
|
||||
**更新日期**:2026-07-17
|
||||
**状态**:当前可运行架构
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/confidence-feedback.md`
|
||||
|
||||
## 1. 定位
|
||||
|
||||
反馈架构包含两条闭环:
|
||||
|
||||
1. 系统自评估:基于工具调用、Verifier、AIOps 规则检查,写入 `diagnosis_session.self_evaluation`。
|
||||
2. 用户反馈:用户标记 `useful` 或 `not_useful`,写入 `diagnosis_session.feedback`,其中 `useful` 会沉淀案例。
|
||||
1. 系统自评估:基于当前 run 的工具调用、Gatekeeper、Verifier、Composer、AIOps 规则检查,写入 `diagnosis_run.self_evaluation`。
|
||||
2. 用户反馈:用户标记 `useful` 或 `not_useful`,优先写入 `diagnosis_run.feedback`,其中 `useful` 会沉淀案例。
|
||||
|
||||
当前重要边界:
|
||||
|
||||
@@ -21,13 +21,17 @@
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Answer["Chat / AIOps final answer"] --> Session["diagnosis_session.answer"]
|
||||
Answer["Chat / AIOps final answer"] --> Run["diagnosis_run.answer"]
|
||||
|
||||
subgraph SelfEval["Self evaluation"]
|
||||
Invocation["tool_invocation"] --> RuleEval["EvaluationService: rule_evaluation"]
|
||||
Invocation --> TraceSummary["ToolTraceSummaryService"]
|
||||
TraceSummary --> Verifier["chat_verifier"]
|
||||
Invocation --> EvidenceRefs["evidence_refs"]
|
||||
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
Gatekeeper --> Projection["VerifiedInputNode"]
|
||||
Projection --> Verifier["chat_verifier"]
|
||||
Verifier --> VerifierEval["verifier_evaluation"]
|
||||
Verifier --> Composer["chat_composer"]
|
||||
Composer --> VerifierEval
|
||||
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
|
||||
AiOpsRule --> AiOpsEval["aiops_rule_evaluation"]
|
||||
end
|
||||
@@ -35,24 +39,24 @@ flowchart TD
|
||||
RuleEval --> Merge["SelfEvaluationMergeService"]
|
||||
VerifierEval --> Merge
|
||||
AiOpsEval --> Merge
|
||||
Merge --> SelfJson["diagnosis_session.self_evaluation"]
|
||||
Merge --> SelfJson["diagnosis_run.self_evaluation"]
|
||||
|
||||
subgraph UserFeedback["User feedback"]
|
||||
UI["Feedback bar"] --> API["POST /api/feedback"]
|
||||
API --> FeedbackService["FeedbackService"]
|
||||
FeedbackService --> FeedbackField["diagnosis_session.feedback"]
|
||||
FeedbackService --> FeedbackField["diagnosis_run.feedback"]
|
||||
FeedbackService --> Useful{"feedback == useful?"}
|
||||
Useful -->|yes| CaseService["CaseLibraryService.createFromSession"]
|
||||
Useful -->|yes| CaseService["CaseLibraryService.createFromRun"]
|
||||
CaseService --> Case["case_library"]
|
||||
Useful -->|no| BadCase["Bad case by feedback=not_useful"]
|
||||
end
|
||||
|
||||
Session --> UI
|
||||
Run --> UI
|
||||
```
|
||||
|
||||
## 3. self_evaluation JSON
|
||||
|
||||
`SelfEvaluationMergeService` 统一维护 `diagnosis_session.self_evaluation`。
|
||||
`SelfEvaluationMergeService` 只维护当前运行的 `diagnosis_run.self_evaluation`,不再解析或写入旧 `diagnosis_session` 数据。
|
||||
|
||||
当前结构:
|
||||
|
||||
@@ -67,11 +71,16 @@ flowchart TD
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 0.8,
|
||||
"critical_fact_count": 2,
|
||||
"claim_checks": [],
|
||||
"facts_checked": [],
|
||||
"rationale": "...",
|
||||
"round": 1,
|
||||
"traceability_version": "v1",
|
||||
"tool_trace_summary": []
|
||||
"executor_output_parse_status": {},
|
||||
"executor_structured_output": {},
|
||||
"gatekeeper_result": {},
|
||||
"composer_output": {},
|
||||
"verified_evidence": []
|
||||
},
|
||||
"aiops_rule_evaluation": {
|
||||
"verdict": "...",
|
||||
@@ -80,10 +89,7 @@ flowchart TD
|
||||
}
|
||||
```
|
||||
|
||||
兼容逻辑:
|
||||
|
||||
- 如果旧 JSON 根节点包含 `evidence_score`,会被包进 `rule_evaluation`。
|
||||
- 如果旧 JSON 根节点包含 `verdict` / `groundedness_score`,会被包进 `verifier_evaluation`。
|
||||
输入必须使用当前分层 JSON:`rule_evaluation`、`verifier_evaluation`、`aiops_rule_evaluation`。旧扁平 JSON 不再自动包装。
|
||||
|
||||
## 4. 规则评分
|
||||
|
||||
@@ -117,17 +123,28 @@ flowchart TD
|
||||
|
||||
## 5. Chat Verifier 自评估
|
||||
|
||||
Chat Verifier 校验 Executor 的最终答案是否被证据支撑。
|
||||
Chat 自评估分三步:
|
||||
|
||||
1. Gatekeeper 用代码校验 Executor 的引用是否真实。
|
||||
2. Verifier 判断已验真的 `evidence_excerpt` 是否能推出 `claim_text`。
|
||||
3. Composer 只把 Verifier 允许表达的内容写成最终用户答复。
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Answer["executor_final_answer"] --> Verifier["chat_verifier"]
|
||||
Invocation["tool_invocation"] --> Summary["ToolTraceSummaryService"]
|
||||
Summary --> Evidence["tool_trace_summary"]
|
||||
Evidence --> Verifier
|
||||
ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
Invocation["tool_invocation"] --> EvidenceRefs["retrieval_details.evidence_refs"]
|
||||
EvidenceRefs --> Gatekeeper
|
||||
Gatekeeper --> GateResult["gatekeeper_result"]
|
||||
GateResult --> Projection["VerifiedInputNode"]
|
||||
ExecutorOutput --> Projection
|
||||
Projection --> VerifiedOutput["verified_executor_output + verified_evidence"]
|
||||
VerifiedOutput --> Verifier["chat_verifier"]
|
||||
Verifier --> Output["verifier_output JSON"]
|
||||
Output --> Composer["chat_composer"]
|
||||
Composer --> ComposerOutput["composer_output"]
|
||||
Output --> Merge["SelfEvaluationMergeService.mergeVerifierEvaluation"]
|
||||
Merge --> Session["diagnosis_session.self_evaluation.verifier_evaluation"]
|
||||
ComposerOutput --> Merge
|
||||
Merge --> Run["diagnosis_run.self_evaluation.verifier_evaluation"]
|
||||
```
|
||||
|
||||
Verifier 输出:
|
||||
@@ -137,16 +154,26 @@ Verifier 输出:
|
||||
| `verdict` | `PASS` / `LOW_CONFID` / `REJECT` |
|
||||
| `groundedness_score` | 关键事实证据支撑度 |
|
||||
| `critical_fact_count` | 关键事实数量 |
|
||||
| `claim_checks` | 对 Executor 结构化 claims 的逐条可推导性判断 |
|
||||
| `facts_checked` | 逐条事实校验 |
|
||||
| `rationale` | 判定原因 |
|
||||
| `tool_trace_summary` | 本次校验使用的证据索引 |
|
||||
| `executor_structured_output` | Executor 输出的结构化 claims 与证据绑定 |
|
||||
| `verified_evidence` | Gatekeeper 通过并投影给 Verifier 的最小 matched evidence |
|
||||
| `gatekeeper_result` | 引用真实性校验结果 |
|
||||
| `composer_output` | 最终表达的解析状态和摘要 |
|
||||
|
||||
ChatService 根据 verdict 决定:
|
||||
|
||||
- `PASS`:输出 Executor 答案。
|
||||
- `PASS`:把允许表达的 claims 交给 Composer 输出。
|
||||
- `LOW_CONFID`:必要时构造 `retry_context` 补证据;否则输出低置信提示。
|
||||
- `REJECT`:降级输出,只保留已确认信息。
|
||||
|
||||
边界:
|
||||
|
||||
- `executor_final_answer` 只作为 debug/fallback 上下文;结构化输出有效时,Verifier 不得从中抽取额外确认事实。
|
||||
- `$.no_evidence` 只能表达“当前查询未检索到匹配证据”,不能表达“已排除/确认没有”。
|
||||
- `run.orchestrationTrace` 是独立的 StateGraph 路由摘要,不属于 `self_evaluation`;当前版本不生成或读取 `tool_trace_summary`。
|
||||
|
||||
## 6. AIOps 规则自评估
|
||||
|
||||
AIOps 当前使用 `AiOpsRuleEvaluationService`,结果写入 `aiops_rule_evaluation`。
|
||||
@@ -167,6 +194,7 @@ POST /api/feedback
|
||||
Content-Type: application/json
|
||||
|
||||
{
|
||||
"runId": "run-xxx",
|
||||
"sessionId": "xxx",
|
||||
"feedback": "useful" | "not_useful"
|
||||
}
|
||||
@@ -178,6 +206,7 @@ Content-Type: application/json
|
||||
{
|
||||
"success": true,
|
||||
"message": "反馈已记录",
|
||||
"runId": "run-xxx",
|
||||
"caseId": "uuid 或 null"
|
||||
}
|
||||
```
|
||||
@@ -186,10 +215,12 @@ Content-Type: application/json
|
||||
|
||||
| feedback | 行为 |
|
||||
|---|---|
|
||||
| `useful` | 写入 `DiagnosisSession.feedback`,调用 `CaseLibraryService.createFromSession` |
|
||||
| `not_useful` | 写入 `DiagnosisSession.feedback`,不改变 session status |
|
||||
| `useful` | 写入 `DiagnosisRun.feedback`,调用 `CaseLibraryService.createFromRun` |
|
||||
| `not_useful` | 写入 `DiagnosisRun.feedback`,不改变 run status |
|
||||
| 其他值 | 返回 HTTP 400 |
|
||||
|
||||
新版本要求请求必须携带 `sessionId + runId`。后端验证 `runId` 属于 `sessionId`;缺少 `runId`、Run 不存在或归属错误时直接拒绝,不绑定 latest run,也不回退历史表。
|
||||
|
||||
## 8. 案例沉淀
|
||||
|
||||
`useful` 反馈会生成或复用 `case_library` 记录。
|
||||
@@ -199,7 +230,7 @@ Content-Type: application/json
|
||||
| CaseLibrary 字段 | 来源 |
|
||||
|---|---|
|
||||
| `caseId` | UUID |
|
||||
| `diagnosisId` | `DiagnosisSession.sessionId` |
|
||||
| `diagnosisId` | `DiagnosisRun.runId` |
|
||||
| `sourceType` | `AUTO` |
|
||||
| `faultCategory` | 当前固定为 `GENERAL` |
|
||||
| `title` | `query` 前 100 字符 |
|
||||
@@ -210,7 +241,7 @@ Content-Type: application/json
|
||||
幂等性:
|
||||
|
||||
```text
|
||||
case_library.diagnosisId == sessionId
|
||||
case_library.diagnosisId == runId
|
||||
-> existing case: return existing
|
||||
-> missing case: create new
|
||||
```
|
||||
@@ -229,23 +260,23 @@ Trace API 会展示:
|
||||
|
||||
| 视角 | 数据来源 |
|
||||
|---|---|
|
||||
| 执行是否成功 | `diagnosis_session.status` |
|
||||
| 执行是否成功 | `diagnosis_run.status` |
|
||||
| 证据是否充分 | `self_evaluation.rule_evaluation` / `verifier_evaluation` |
|
||||
| 用户是否认可 | `diagnosis_session.feedback` |
|
||||
| 用户是否认可 | `diagnosis_run.feedback` |
|
||||
|
||||
## 10. 后续增强
|
||||
|
||||
近期优先:
|
||||
|
||||
1. 将 `rule_evaluation` 与 `verifier_evaluation` 在 Trace API 中结构化展示。
|
||||
2. `not_useful` 反馈沉淀 bad case,而不是只写字段。
|
||||
3. useful 案例自动提取 faultCategory、errorCode、service、rootCause、solution。
|
||||
4. AIOps 引入 LLM Verifier。
|
||||
5. 把反馈和 eval baseline 打通,形成可回归的质量改进闭环。
|
||||
2. 将 ISS-008 / ISS-009 这类 E2E 通过样例固化进 diagnosis eval fixtures。
|
||||
3. `not_useful` 反馈沉淀 bad case,而不是只写字段。
|
||||
4. useful 案例自动提取 faultCategory、errorCode、service、rootCause、solution。
|
||||
5. AIOps 引入 LLM Verifier。
|
||||
6. 把反馈和 eval baseline 打通,形成可回归的质量改进闭环。
|
||||
|
||||
暂不优先:
|
||||
|
||||
- 用用户反馈直接修改 session status。
|
||||
- 仅凭 `evidence_score` 判断答案正确。
|
||||
- 在没有人工审核时自动把 bad case 反向写入 Prompt。
|
||||
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
# Harness 与质量门禁架构
|
||||
|
||||
**更新日期**:2026-07-05
|
||||
**状态**:当前可运行架构 + 后续门禁规划
|
||||
**更新日期**:2026-07-17
|
||||
**状态**:当前可运行架构 + 后续门禁规划
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
||||
|
||||
## 1. 设计目标
|
||||
@@ -20,7 +20,9 @@ Agent 系统的核心风险不是“没有答案”,而是:
|
||||
Prompt contract
|
||||
+ Tool boundary
|
||||
+ Agent hooks
|
||||
+ StateGraph routing contract
|
||||
+ Trace persistence
|
||||
+ Gatekeeper deterministic validation
|
||||
+ Verifier / rule evaluation
|
||||
+ Eval baseline
|
||||
```
|
||||
@@ -30,21 +32,26 @@ Prompt contract
|
||||
```mermaid
|
||||
flowchart TB
|
||||
Input["User / AIOps input"] --> Prompt["Prompt contract"]
|
||||
Prompt --> Agent["Planner / Executor / Verifier"]
|
||||
Prompt --> Agent["Diagnosis StateGraph Nodes"]
|
||||
Agent --> Tools["Evidence tools"]
|
||||
Tools --> Invocation["tool_invocation"]
|
||||
Agent --> StepHook["AgentLoggingHook"]
|
||||
StepHook --> Step["agent_step"]
|
||||
Agent --> Session["diagnosis_session"]
|
||||
Agent --> Run["diagnosis_run"]
|
||||
|
||||
Invocation --> TraceSummary["ToolTraceSummaryService"]
|
||||
TraceSummary --> Verifier["chat_verifier"]
|
||||
Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
|
||||
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
Agent --> Gatekeeper
|
||||
Gatekeeper --> Projection["VerifiedInputNode"]
|
||||
Projection --> Verifier["chat_verifier"]
|
||||
Verifier --> SelfEval["self_evaluation.verifier_evaluation"]
|
||||
|
||||
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
|
||||
AiOpsRule --> AiOpsEval["self_evaluation.aiops_rule_evaluation"]
|
||||
|
||||
Session --> TraceAPI["DiagnosisTraceService"]
|
||||
Run --> TraceAPI["DiagnosisTraceService"]
|
||||
Agent --> Routing["diagnosis_run.orchestration_trace"]
|
||||
Routing --> TraceAPI
|
||||
Step --> TraceAPI
|
||||
Invocation --> TraceAPI
|
||||
SelfEval --> TraceAPI
|
||||
@@ -63,17 +70,37 @@ flowchart TB
|
||||
| `planner-prompt.md` | AIOps Planner 规划、再规划、输出告警报告 |
|
||||
| `executor-prompt.md` | AIOps Executor 按步骤调用工具 |
|
||||
| `chat-planner-prompt.md` | Chat 复杂问题规划 |
|
||||
| `chat-executor-prompt.md` | Chat 执行工具并形成诊断答复 |
|
||||
| `chat-verifier-prompt.md` | 校验 Executor 答案是否被工具证据支撑 |
|
||||
| `chat-executor-prompt.md` | Chat 执行工具并输出 `executor_evidence_v2` 微观事实 |
|
||||
| `chat-verifier-prompt.md` | 基于 Gatekeeper 已验真的证据判断 claims 是否可推出 |
|
||||
| `chat-composer-prompt.md` | 基于 Verifier 允许表达的内容生成最终用户答复 |
|
||||
|
||||
Prompt 层当前承担的门禁:
|
||||
|
||||
- 禁止凭记忆回答错误码、接口定义、排障步骤。
|
||||
- 需要外部信息时必须调用工具。
|
||||
- 工具连续失败或返回空结果时,最终报告必须诚实说明。
|
||||
- Chat Verifier 不允许做新检索,只能校验已有证据。
|
||||
- Chat Executor 不允许在窄范围问题中扩展根因、风险或修复建议。
|
||||
- Chat Verifier 不允许做新检索,只能判断已验真证据是否可推出 claims。
|
||||
- Chat Composer 不允许补事实,尤其不能把 `$.no_evidence` 表达为“已排除/确认没有”。
|
||||
- AIOps payload 模式必须聚焦输入告警。
|
||||
|
||||
Chat 链路还会在 `verifier_evaluation.prompt_audit` 中持久化紧凑 Prompt 审计快照:
|
||||
|
||||
```json
|
||||
{
|
||||
"version": "chat-prompts-v1",
|
||||
"prompts": [
|
||||
{
|
||||
"name": "chat_executor",
|
||||
"version": "chat-executor-v2",
|
||||
"resource": "prompts/chat-executor-prompt.md"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
该快照只保存版本和资源路径,不保存完整 Prompt 文本。它用于面试演示、trace 回放和离线 baseline 解释“本次诊断使用了哪套 Prompt 契约”。
|
||||
|
||||
## 4. Trace Hooks
|
||||
|
||||
`AgentLoggingHook` 是当前 Agent step 可观测性的核心。
|
||||
@@ -85,10 +112,10 @@ sequenceDiagram
|
||||
participant H as AgentLoggingHook
|
||||
participant DB as agent_step
|
||||
|
||||
A->>H: before_model(messages, sessionId)
|
||||
H->>DB: 写入 model_input / step_index / agent_name
|
||||
A->>H: before_model(messages, sessionId, runId)
|
||||
H->>DB: 写入 session_id / run_id / model_input / step_index / agent_name
|
||||
A-->>A: LLM 推理
|
||||
A->>H: after_model(messages, sessionId)
|
||||
A->>H: after_model(messages, sessionId, runId)
|
||||
H->>DB: 回填 model_output / thought / has_tool_call / duration / token_count
|
||||
```
|
||||
|
||||
@@ -101,6 +128,8 @@ sequenceDiagram
|
||||
- token count。
|
||||
- Verifier 的 JSON 输出摘要。
|
||||
|
||||
新写入必须带 `run_id`;`session_id` 只用于会话归属和粗粒度排查,不能替代 Run 边界。
|
||||
|
||||
## 5. Tool Invocation 门禁
|
||||
|
||||
工具调用记录由 `ToolInvocationRecorder` 和具体工具共同完成。
|
||||
@@ -115,6 +144,7 @@ retrieval_layer
|
||||
l0_match_count
|
||||
l1_match_count
|
||||
retrieval_details
|
||||
-> evidence_refs
|
||||
relevance_level
|
||||
dedup_reason
|
||||
duration_ms
|
||||
@@ -128,23 +158,45 @@ error_message
|
||||
- 检索结果归一化为 `PRECISE`、`HIGHLY_RELEVANT`、`REFERENCE`。
|
||||
- 同 session 内重复文档会被 `RetrievedDocTracker` 去重。
|
||||
- dedup、no evidence、failed 等状态进入 `retrieval_details.evidence_status`。
|
||||
- `retrieval_details.evidence_refs` 记录可被 Executor 引用的最小证据文本,格式为 `raw_path + text`。
|
||||
- no-hit / no-evidence 工具结果会生成 `raw_path=$.no_evidence` 的负向证据引用,语义仅限“本次查询未检索到匹配证据”。
|
||||
|
||||
## 6. Verifier 门禁
|
||||
## 6. Gatekeeper 与 Verifier 门禁
|
||||
|
||||
Chat Verifier 的输入不是原始工具日志,而是 `ToolTraceSummaryService` 构造的证据索引。
|
||||
Chat StateGraph 在 Verifier 前显式执行 Gatekeeper 和 Verified Input。Gatekeeper 不调用 LLM,只用代码检查 Executor 输出的证据引用是否真实存在;Verified Input 只投影通过的 binding。
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Invocation["tool_invocation"] --> Summary["ToolTraceSummaryService"]
|
||||
Summary --> EvidenceIndex["tool_trace_summary"]
|
||||
EvidenceIndex --> Verifier["chat_verifier"]
|
||||
ExecutorAnswer["executor_final_answer"] --> Verifier
|
||||
Invocation["tool_invocation"] --> EvidenceRefs["retrieval_details.evidence_refs"]
|
||||
ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
EvidenceRefs --> Gatekeeper
|
||||
Gatekeeper --> GateResult["gatekeeper_result"]
|
||||
GateResult --> Projection["VerifiedInputNode"]
|
||||
Projection --> VerifiedClaims["verified_executor_output"]
|
||||
Projection --> VerifiedEvidence["verified_evidence"]
|
||||
GateResult --> Verifier["chat_verifier"]
|
||||
VerifiedClaims --> Verifier
|
||||
VerifiedEvidence --> Verifier
|
||||
Verifier --> Verdict{"verdict"}
|
||||
Verdict -->|PASS| Pass["输出原答案"]
|
||||
Verdict -->|PASS| Composer["chat_composer"]
|
||||
Composer --> Pass["输出最终答复"]
|
||||
Verdict -->|LOW_CONFID| Low["补证据或低置信输出"]
|
||||
Verdict -->|REJECT| Reject["降级输出"]
|
||||
```
|
||||
|
||||
Gatekeeper 检查:
|
||||
|
||||
| 检查 | 失败语义 |
|
||||
|---|---|
|
||||
| `answer_version=executor_evidence_v2` | 非结构化或旧结构输出降为低置信 |
|
||||
| `source_invocation_id` 真实存在 | 伪造 ID 直接拒绝 |
|
||||
| `tool_name` 与 invocation 对齐 | 张冠李戴直接拒绝 |
|
||||
| `raw_path` 存在于 `evidence_refs` | 无中生有直接拒绝 |
|
||||
| `evidence_excerpt` 由 `evidence_refs[].text` 支撑 | excerpt 编造或错配直接拒绝 |
|
||||
| `negative_observation` 只能引用 `$.no_evidence` | 用正向日志证明“没查到”直接拒绝 |
|
||||
|
||||
Gatekeeper 审计还会记录 `rule_set_version` 和已启用规则元数据摘要。当前规则元数据来自本地 `gatekeeper-rules.json`,规则执行仍是确定性 Java 代码。
|
||||
|
||||
Verifier 输出:
|
||||
|
||||
```json
|
||||
@@ -152,17 +204,22 @@ Verifier 输出:
|
||||
"verdict": "PASS|LOW_CONFID|REJECT",
|
||||
"groundedness_score": 0.8,
|
||||
"critical_fact_count": 2,
|
||||
"claim_checks": [],
|
||||
"facts_checked": [],
|
||||
"rationale": "..."
|
||||
}
|
||||
```
|
||||
|
||||
Verifier 不再逐字核验 excerpt 真伪;这由 Gatekeeper 完成。Verifier 只回答一个问题:`claim_text` 是否能由已经验真的 `evidence_excerpt` 推导出来。
|
||||
|
||||
结果写入:
|
||||
|
||||
```text
|
||||
diagnosis_session.self_evaluation.verifier_evaluation
|
||||
diagnosis_run.self_evaluation.verifier_evaluation
|
||||
```
|
||||
|
||||
其中持久化 verified `executor_structured_output`、`verified_evidence`、`gatekeeper_result`、`prompt_audit` 和 `composer_output`,用于 Trace 回放。Graph 路由另存 `diagnosis_run.orchestration_trace`;当前版本不生成或读取 `tool_trace_summary`。
|
||||
|
||||
## 7. AIOps 规则门禁
|
||||
|
||||
AIOps 当前不走 Chat Verifier,而是用 `AiOpsRuleEvaluationService` 做轻量检查。
|
||||
@@ -177,7 +234,7 @@ AIOps 当前不走 Chat Verifier,而是用 `AiOpsRuleEvaluationService` 做轻
|
||||
结果写入:
|
||||
|
||||
```text
|
||||
diagnosis_session.self_evaluation.aiops_rule_evaluation
|
||||
diagnosis_run.self_evaluation.aiops_rule_evaluation
|
||||
```
|
||||
|
||||
## 8. Eval Baseline
|
||||
@@ -197,9 +254,8 @@ diagnosis_session.self_evaluation.aiops_rule_evaluation
|
||||
- 工具参数 schema 校验。
|
||||
- 同一工具调用次数上限。
|
||||
- 工具超时的统一熔断。
|
||||
- 报告中的数值与工具返回值自动对齐校验。
|
||||
- Prompt 版本记录和回滚。
|
||||
- Gatekeeper 规则远程化或三层分离:索引层、元数据层、规则实现层。
|
||||
- Prompt 版本回滚和更细粒度变更审计。
|
||||
- Verifier 对 AIOps 报告的 LLM 级事实校验。
|
||||
|
||||
这些应在评测集扩大后逐步加入,避免一次性把诊断流程卡得过死。
|
||||
|
||||
|
||||
@@ -1,11 +1,12 @@
|
||||
# 面试一页式架构讲解
|
||||
|
||||
**更新日期**:2026-07-20
|
||||
**用途**:面试现场 2-5 分钟讲清项目
|
||||
**适合场景**:开场介绍、架构追问、Demo 前铺垫
|
||||
|
||||
## 1. 一句话
|
||||
|
||||
SuperBizAgent 是一个面向企业故障诊断的可追踪 Agent 系统:它把用户问题或 AIOps 告警转换成 Planner、Executor、Verifier 的诊断链路,所有工具证据、模型步骤、最终答案、自评估和用户反馈都能通过同一个 `sessionId` 回放。
|
||||
SuperBizAgent 是一个面向企业故障诊断的可追踪 Agent 系统:复杂 Chat 通过 bounded StateGraph 编排 Planner、Executor、Gatekeeper、Verified Input、Verifier、Composer 和安全 Fallback,`sessionId` 保留多轮上下文,`runId` 精确绑定一次诊断运行;工具证据、模型步骤、Graph 路由、自评估、最终答案和用户反馈都能按 `sessionId + runId` 回放。
|
||||
|
||||
## 2. 一张图
|
||||
|
||||
@@ -16,7 +17,7 @@ flowchart TB
|
||||
API --> Chat["ChatService"]
|
||||
API --> AiOps["AiOpsService"]
|
||||
|
||||
Chat --> ChatFlow["Chat: Planner -> Executor -> Verifier"]
|
||||
Chat --> ChatFlow["Chat StateGraph: Planner -> Executor -> Gatekeeper -> Verified Input -> Verifier -> Composer/Fallback"]
|
||||
AiOps --> AiOpsFlow["AIOps: Supervisor -> Planner / Executor"]
|
||||
|
||||
ChatFlow --> Tools["Evidence Tools"]
|
||||
@@ -34,14 +35,19 @@ flowchart TB
|
||||
AiOpsFlow --> Trace
|
||||
Tools --> Trace
|
||||
|
||||
Trace --> Session["diagnosis_session"]
|
||||
Trace --> ChatSession["chat_session"]
|
||||
Trace --> Run["diagnosis_run"]
|
||||
Trace --> Step["agent_step"]
|
||||
Trace --> Invocation["tool_invocation"]
|
||||
|
||||
Invocation --> Verifier["Verifier / Rule Evaluation"]
|
||||
Invocation --> Gate["Gatekeeper / Rule Evaluation"]
|
||||
Gate --> Projection["Verified Input"]
|
||||
Projection --> Verifier["Verifier"]
|
||||
Verifier --> SelfEval["self_evaluation"]
|
||||
|
||||
Session --> TraceAPI["GET /api/diagnosis/{sessionId}/trace"]
|
||||
Run --> RouteTrace["orchestration_trace"]
|
||||
Run --> TraceAPI["GET /api/diagnosis/{sessionId}/trace?runId=..."]
|
||||
RouteTrace --> TraceAPI
|
||||
Step --> TraceAPI
|
||||
Invocation --> TraceAPI
|
||||
SelfEval --> TraceAPI
|
||||
@@ -55,24 +61,24 @@ flowchart TB
|
||||
```text
|
||||
这个项目不是把问题直接丢给大模型,而是把诊断拆成可审计的执行链路。
|
||||
|
||||
Chat 复杂问题走 Planner -> Executor -> Verifier:
|
||||
Planner 负责拆解,Executor 负责调用知识库、日志和指标工具,Verifier 只基于已有工具证据校验最终答案。
|
||||
Chat 复杂问题走 bounded StateGraph:
|
||||
Planner 负责拆解,Executor 调用知识库、日志和指标工具并提炼带证据引用的微观事实;Gatekeeper 用代码核对 invocation、raw_path 和 excerpt 是否真实;Verified Input 只投影通过的 binding;Verifier 判断这些事实能否由已验真的证据推出;Composer 只把允许表达的结论写成最终答案,不可恢复分支由 Fallback 生成安全答复。
|
||||
|
||||
AIOps 告警入口走 Supervisor 调度 Planner/Executor:
|
||||
如果请求里有 alert payload,系统会进入 PAYLOAD_TARGETED 模式,报告必须聚焦这个告警,而不是被当前环境中的其他活跃告警带偏。
|
||||
|
||||
所有过程都会落到 diagnosis_session、agent_step、tool_invocation。
|
||||
所以我可以用一个 sessionId 回放:模型怎么规划、调了哪些工具、工具返回什么、Verifier 怎么判定、用户最后是否反馈有用。
|
||||
会话元数据会落到 chat_session,每次诊断运行会落到 diagnosis_run,步骤和工具明细通过 run_id 关联;证据质量写入 self_evaluation,Graph 路由独立写入 orchestration_trace。
|
||||
所以我可以用 sessionId + runId 精确回放:模型怎么规划、调了哪些工具、Gatekeeper 怎么验真、Graph 为什么重试或降级、Verifier 怎么判定、Composer/Fallback 如何结束、用户最后是否反馈有用。
|
||||
```
|
||||
|
||||
## 4. 五个亮点
|
||||
|
||||
| 亮点 | 怎么讲 |
|
||||
|---|---|
|
||||
| 可追踪 Agent | 每次诊断都有 `sessionId`,Trace API 可以回放 session、step、tool |
|
||||
| 可追踪 StateGraph | 每次诊断都有 `runId`,Trace API 可以回放 run、step、tool 和 `run.orchestrationTrace`;同一 `sessionId` 可有多次独立 run |
|
||||
| 显式工具证据链 | `lookup_knowledge`、日志、指标都记录到 `tool_invocation` |
|
||||
| RAG 工程化 | L0 降级为 hint,Spring AI VectorStore 做主检索,SDK fallback 保底 |
|
||||
| 质量门禁 | Chat Verifier 校验 groundedness,AIOps rule evaluation 控制告警聚焦 |
|
||||
| 质量门禁 | Chat Gatekeeper 验引用、Verifier 判可推导、Composer 控表达,AIOps rule evaluation 控制告警聚焦 |
|
||||
| 反馈闭环 | useful 反馈沉淀 `case_library`,not_useful 保留 bad case 信号 |
|
||||
|
||||
## 5. 三个关键取舍
|
||||
@@ -99,15 +105,14 @@ aiops_rule_evaluation -> AIOps 报告是否聚焦告警并使用证据
|
||||
|
||||
| 追问 | 回答方向 |
|
||||
|---|---|
|
||||
| 怎么防止幻觉? | Executor 必须用工具;Verifier 只基于 `tool_trace_summary` 校验;LOW_CONFID/REJECT 会降级输出 |
|
||||
| 怎么防止幻觉? | Executor 输出 `executor_evidence_v2`,每个 claim 绑定 `source_invocation_id + raw_path + evidence_excerpt`;Gatekeeper 用 `tool_invocation.retrieval_details.evidence_refs` 核验引用真实性;Verifier 只判断可推导性;Composer 防止把 no-evidence 说成已排除 |
|
||||
| RAG 质量怎么保证? | offline golden cases + live acceptance + trace inspection 三层验证 |
|
||||
| 为什么 L0 不直接返回? | L0 子串命中不等于语义相关,当前只做 domain/entity hint 和 metadata filter |
|
||||
| AIOps 如何避免跑偏? | payload 模式生成 recommended query,并用 rule evaluation 检查报告聚焦输入告警 |
|
||||
| 下一步怎么演进? | evidence block、邻居 chunk、Playbook、AIOps LLM Verifier、MCP 工具协议化 |
|
||||
| 下一步怎么演进? | 固化 E2E fixture、Prompt version、Gatekeeper 规则配置化、邻居 chunk、AIOps LLM Verifier、MCP 工具协议化 |
|
||||
|
||||
## 7. 现场演示入口
|
||||
|
||||
- Demo 脚本:`mvp/demo/ten-minute-interview-demo.md`
|
||||
- 故事案例:`interview/story-cases.md`
|
||||
- 架构细节:`mvp/architecture/README.md`
|
||||
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
**更新日期**:2026-07-06
|
||||
**状态**:当前主架构 + 后续演进边界
|
||||
**关联计划**:`mvp/issues/rag-refactor-plan.md`
|
||||
**关联计划**:[`mvp/issues/active/rag-refactor-plan.md`](../issues/active/rag-refactor-plan.md)
|
||||
|
||||
## 1. 架构目标
|
||||
|
||||
|
||||
@@ -141,7 +141,8 @@ post-retrieval 层再把检索候选归一为:
|
||||
|
||||
- 给 Agent 输出 completeness hint。
|
||||
- 写入 `tool_invocation.relevance_level`。
|
||||
- 给 Verifier 构造 `tool_trace_summary`。
|
||||
- 给 Gatekeeper 提供 `evidence_refs` 引用验真源。
|
||||
- 由 Gatekeeper 核验后,经 `VerifiedInputNode` 给 Verifier 构造最小 `verified_evidence` 投影。
|
||||
- 供 EvaluationService 计算 evidence score。
|
||||
|
||||
## 6. 文档切片和 metadata
|
||||
@@ -178,11 +179,14 @@ flowchart LR
|
||||
LookupResult --> Recorder["ToolInvocationRecorder"]
|
||||
Recorder --> Invocation["tool_invocation"]
|
||||
Invocation --> Trace["DiagnosisTraceService"]
|
||||
Invocation --> Summary["ToolTraceSummaryService"]
|
||||
Summary --> Verifier["chat_verifier"]
|
||||
Invocation --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
Gatekeeper --> Projection["VerifiedInputNode"]
|
||||
Projection --> Verifier["chat_verifier"]
|
||||
Invocation --> Eval["EvaluationService / RAG eval"]
|
||||
```
|
||||
|
||||
当前 StateGraph 不生成或读取 `tool_trace_summary`,也不会把完整工具调用摘要输入 Verifier。
|
||||
|
||||
`tool_invocation` 中与检索相关的字段:
|
||||
|
||||
```text
|
||||
@@ -205,9 +209,25 @@ success
|
||||
- evidence status。
|
||||
- dedup reason。
|
||||
- evidence block summaries。
|
||||
- evidence refs:`raw_path + text`,用于核对 Executor 的 `evidence_excerpt`。
|
||||
- context pack summary。
|
||||
- rerank trace。
|
||||
|
||||
其中 `evidence_refs` 是当前 Chat 证据链路的精确引用源:
|
||||
|
||||
```json
|
||||
{
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.evidence_blocks[0]",
|
||||
"text": "最小证据文本"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
如果检索返回 no evidence,应使用 `raw_path=$.no_evidence` 记录负向证据。它只能说明“本次检索没有匹配证据”,不能作为“问题不存在”的证明。
|
||||
|
||||
## 8. 去重与行动记忆
|
||||
|
||||
当前 session 级去重由 `RetrievedDocTracker` 负责。
|
||||
|
||||
@@ -1,68 +1,68 @@
|
||||
# 会话与 Trace 生命周期
|
||||
|
||||
**更新日期**:2026-07-05
|
||||
**状态**:当前可运行架构
|
||||
**更新日期**:2026-07-17
|
||||
**状态**:当前可运行架构
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/session-management.md`
|
||||
|
||||
## 1. 定位
|
||||
|
||||
旧版会话设计以 Redis 会话为主,MySQL 作为可选长期沉淀。当前 MVP 的可追踪诊断已经转为 MySQL Trace 三表为主:
|
||||
当前 MVP 把“会话态”和“运行态”拆开:
|
||||
|
||||
```text
|
||||
diagnosis_session
|
||||
-> agent_step
|
||||
-> tool_invocation
|
||||
chat_session(sessionId)
|
||||
-> diagnosis_run(runId)
|
||||
-> agent_step(runId)
|
||||
-> tool_invocation(runId)
|
||||
```
|
||||
|
||||
因此本文描述的是当前可运行链路:
|
||||
|
||||
- `sessionId` 是一次诊断和后续 trace/feedback 的关联键。
|
||||
- `diagnosis_session` 保存会话级状态、问题、答案、自评估和反馈。
|
||||
- `agent_step` 保存每个 Agent 模型调用。
|
||||
- `tool_invocation` 保存工具调用事实。
|
||||
- `DiagnosisTraceService` 聚合三类记录,形成可回放 trace。
|
||||
- `sessionId` 表示多轮会话目录和 Redis 上下文。
|
||||
- `runId` 表示一次可回放诊断执行。
|
||||
- `DiagnosisTraceService` 聚合一个 run 的主记录、步骤和工具调用,形成可回放 Trace。
|
||||
- 运行时不再映射、读取或写入 `diagnosis_session`;数据库中的旧表不属于当前版本契约。
|
||||
|
||||
## 2. 生命周期总图
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Start["request: chat / ai_ops"] --> Resolve["resolve sessionId"]
|
||||
Resolve --> Create["create or reset diagnosis_session"]
|
||||
Create --> Running["status = RUNNING"]
|
||||
Resolve --> Session["ensure chat_session metadata"]
|
||||
Session --> Run["create diagnosis_run(runId)"]
|
||||
Run --> Running["run.status = RUNNING"]
|
||||
|
||||
Running --> Agent["Agent workflow"]
|
||||
Agent --> StepHook["AgentLoggingHook"]
|
||||
StepHook --> Step["agent_step"]
|
||||
Agent --> Tool["Evidence tools"]
|
||||
Tool --> Invocation["tool_invocation"]
|
||||
Running --> Agent["Chat StateGraph / AIOps workflow"]
|
||||
Agent --> Context["execution context(sessionId, runId)"]
|
||||
Context --> StepHook["AgentLoggingHook"]
|
||||
StepHook --> Step["agent_step(session_id, run_id)"]
|
||||
Context --> Tool["Evidence tools"]
|
||||
Tool --> Invocation["tool_invocation(session_id, run_id)"]
|
||||
Invocation --> Gatekeeper["Gatekeeper evidence validation"]
|
||||
|
||||
Agent --> Final{"workflow result"}
|
||||
Final -->|success| Success["status = SUCCESS, answer saved"]
|
||||
Final -->|failed| Failed["status = FAILED"]
|
||||
Agent --> GraphTrace["Chat: save orchestration_trace"]
|
||||
GraphTrace --> Final{"workflow result"}
|
||||
Final -->|success| Success["run.status = SUCCESS, answer saved"]
|
||||
Final -->|failed| Failed["run.status = FAILED"]
|
||||
|
||||
Success --> Evaluation["self_evaluation merge"]
|
||||
Success --> Evaluation["diagnosis_run.self_evaluation merge"]
|
||||
Failed --> Evaluation
|
||||
Evaluation --> Trace["GET /api/diagnosis/{sessionId}/trace"]
|
||||
Success --> Feedback["POST /api/feedback"]
|
||||
Feedback --> Case["useful -> case_library"]
|
||||
Evaluation --> Trace["GET /api/diagnosis/{sessionId}/trace?runId=..."]
|
||||
Success --> Feedback["POST /api/feedback(sessionId, runId)"]
|
||||
Feedback --> Case["useful -> case_library(run_id)"]
|
||||
```
|
||||
|
||||
## 3. sessionId 规则
|
||||
## 3. ID 规则
|
||||
|
||||
| 链路 | sessionId 来源 |
|
||||
|---|---|
|
||||
| Chat | 如果请求带 sessionId,则复用;否则生成短 UUID |
|
||||
| AIOps | 如果 payload 带 sessionId,则复用;否则生成 UUID |
|
||||
| Trace | URL path 中的 `{sessionId}` |
|
||||
| Feedback | request body 中的 `sessionId` |
|
||||
| ID | 来源 | 含义 |
|
||||
|---|---|---|
|
||||
| `sessionId` | Chat request `Id`、AIOps payload `sessionId`,缺失时由服务生成 | 多轮会话目录和 Redis 上下文 |
|
||||
| `runId` | 每次有效 Chat/AIOps 执行创建 | 一次诊断运行和 Trace 回放边界 |
|
||||
|
||||
设计含义:
|
||||
|
||||
- 同一个 `sessionId` 可以贯穿诊断、trace 查询和用户反馈。
|
||||
- 当前诊断开始时会重置当前 session 的运行态字段,例如 answer、duration、step/tool count。
|
||||
- `sessionId` 是业务关联键,不依赖数据库自增 ID 暴露给外部。
|
||||
- 同一个 `sessionId` 可以贯穿多轮 Chat。
|
||||
- 每次有效 Chat/AIOps 执行都会创建新的 `runId`。
|
||||
- Trace 和 Feedback 必须同时传 `sessionId + runId`;后端不解析 latest run,也不回退旧表。
|
||||
|
||||
## 4. 状态流转
|
||||
## 4. 运行状态流转
|
||||
|
||||
```mermaid
|
||||
stateDiagram-v2
|
||||
@@ -76,14 +76,15 @@ stateDiagram-v2
|
||||
|
||||
字段边界:
|
||||
|
||||
| 字段 | 含义 |
|
||||
|---|---|
|
||||
| `status` | 执行状态:`PENDING` / `RUNNING` / `SUCCESS` / `FAILED` |
|
||||
| `answer` | Agent 最终返回给用户的报告或答复 |
|
||||
| `self_evaluation` | 系统自评估 JSON |
|
||||
| `feedback` | 用户反馈:`useful` / `not_useful` / null |
|
||||
| 字段 | 所属表 | 含义 |
|
||||
|---|---|---|
|
||||
| `status` | `diagnosis_run` | 单次运行执行状态 |
|
||||
| `answer` | `diagnosis_run` | 本次运行最终报告或答复 |
|
||||
| `self_evaluation` | `diagnosis_run` | 本次运行系统自评估 JSON |
|
||||
| `orchestration_trace` | `diagnosis_run` | Chat StateGraph 路由摘要;包含 version、transitions、final node、termination reason、degraded 和 evidence retry count |
|
||||
| `feedback` | `diagnosis_run` | 本次运行用户反馈 |
|
||||
|
||||
`feedback` 不修改 `status`。一个执行成功但用户标记 `not_useful` 的 session,仍然应该是 `SUCCESS + feedback=not_useful`。
|
||||
`feedback` 不修改 `status`。一个执行成功但用户标记 `not_useful` 的 run,仍然应该是 `SUCCESS + feedback=not_useful`。
|
||||
|
||||
## 5. agent_step 写入
|
||||
|
||||
@@ -92,105 +93,74 @@ stateDiagram-v2
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant Agent as ReactAgent
|
||||
participant Agent as Agent
|
||||
participant Hook as AgentLoggingHook
|
||||
participant DB as agent_step
|
||||
|
||||
Agent->>Hook: before_model(messages, sessionId)
|
||||
Hook->>DB: insert step_index / agent_name / model_input
|
||||
Agent-->>Agent: model call
|
||||
Agent->>Hook: after_model(messages, sessionId)
|
||||
Hook->>DB: update model_output / thought / has_tool_call / duration / token_count
|
||||
Agent->>Hook: before_model(messages, sessionId, runId)
|
||||
Hook->>DB: insert step(session_id, run_id, model_input, step_index)
|
||||
Agent->>Hook: after_model(output, sessionId, runId)
|
||||
Hook->>DB: update model_output, duration, token_count, has_tool_call
|
||||
```
|
||||
|
||||
当前记录:
|
||||
|
||||
- `session_id`
|
||||
- `step_index`
|
||||
- `agent_name`
|
||||
- `model_input`
|
||||
- `model_output`
|
||||
- `thought`
|
||||
- `has_tool_call`
|
||||
- `duration_ms`
|
||||
- `token_count`
|
||||
新写入必须带 `run_id`,同时保留 `session_id` 便于粗粒度排查。
|
||||
|
||||
## 6. tool_invocation 写入
|
||||
|
||||
工具调用记录真实工具事实,不记录模型猜测。
|
||||
|
||||
关键字段:
|
||||
工具调用记录同样通过执行上下文拿到 `sessionId + runId`:
|
||||
|
||||
```text
|
||||
session_id
|
||||
step_id
|
||||
tool_name
|
||||
input_params
|
||||
output_preview
|
||||
output_length
|
||||
retrieval_layer
|
||||
l0_match_count
|
||||
l1_match_count
|
||||
retrieval_details
|
||||
relevance_level
|
||||
dedup_reason
|
||||
duration_ms
|
||||
success
|
||||
error_message
|
||||
ToolInvocationRecorder
|
||||
-> tool_invocation.session_id
|
||||
-> tool_invocation.run_id
|
||||
-> retrieval_details / evidence_refs
|
||||
```
|
||||
|
||||
对 `lookup_knowledge`,`retrieval_details` 会承载 L0/L1、领域、证据状态、去重等检索细节。对非检索工具,检索字段可以为空。
|
||||
Gatekeeper 和 EvaluationService 按 `run_id` 读取工具调用,避免同一 session 的其他 run 参与评分或证据校验。Verifier 只读取 `VerifiedInputNode` 生成的 verified projection,不直接读取完整工具调用列表。
|
||||
|
||||
## 7. Trace API 聚合
|
||||
|
||||
```text
|
||||
GET /api/diagnosis/{sessionId}/trace
|
||||
GET /api/diagnosis/{sessionId}/trace?runId=run-...
|
||||
```
|
||||
|
||||
聚合逻辑:
|
||||
|
||||
```text
|
||||
diagnosis_session by sessionId
|
||||
+ agent_step ordered by step_index
|
||||
+ tool_invocation ordered by id
|
||||
diagnosis_run by sessionId + runId
|
||||
+ run.orchestrationTrace parsed from diagnosis_run.orchestration_trace
|
||||
+ chat_session metadata when available
|
||||
+ agent_step where run_id = runId, ordered by the Trace API
|
||||
+ tool_invocation where run_id = runId order by id
|
||||
-> DiagnosisTraceResponse
|
||||
```
|
||||
|
||||
Trace 视图回答的问题:
|
||||
当 `runId` 缺失时,Trace API 直接拒绝请求。当 `runId` 属于其他 `sessionId` 时,API 同样拒绝,不能泄漏其他会话的 Trace。
|
||||
|
||||
- 这次诊断是否成功?
|
||||
- 哪些 Agent 参与了?
|
||||
- 每一步模型输入输出是什么摘要?
|
||||
- 调用了哪些工具?
|
||||
- 工具返回了什么证据?
|
||||
- Verifier / AIOps rule 是否通过?
|
||||
- 用户是否反馈有用?
|
||||
`run.orchestrationTrace` 只属于精确 Run 投影,响应不再提供兼容 `session` 对象。它解释 Graph 路由;`selfEvaluation` 解释证据/答案质量;`steps` 和 `toolInvocations` 保存详细执行证据,三者职责互不替代。非 StateGraph Run 的该字段可以为空。
|
||||
|
||||
## 8. Chat 与 AIOps 差异
|
||||
|
||||
| 维度 | Chat | AIOps |
|
||||
|---|---|---|
|
||||
| `agent_flow` | `CHAT` | `AI_OPS` |
|
||||
| 编排方式 | `SequentialAgent`: Planner -> Executor -> Verifier | `SupervisorAgent`: Planner + Executor |
|
||||
| 编排方式 | bounded `StateGraph`: Planner / Executor / Gatekeeper / Verified Input / Verifier / Composer / Fallback | `SupervisorAgent`: Planner + Executor |
|
||||
| 自评估 | `rule_evaluation` + `verifier_evaluation` | `aiops_rule_evaluation` |
|
||||
| 答案字段 | Chat 最终答复 | 告警分析报告 |
|
||||
| payload | 用户自然语言 + history | alert payload 或 auto-discovery |
|
||||
| runId 暴露 | `/api/chat` JSON response | `/api/ai_ops` SSE metadata message |
|
||||
|
||||
Chat StateGraph 的权威自动化验收分三层:`DiagnosisGraphWorkflowTest` 验证路由,`DiagnosisGraphNodeContractTest` 验证真实 Node 输入输出,`ChatServiceGraphIntegrationTest` 验证 Run 生命周期、Trace 持久化和对外集成。
|
||||
|
||||
## 9. 清理与边界
|
||||
|
||||
当前会话持久化边界:
|
||||
|
||||
- MySQL Trace 记录是主要可回放来源。
|
||||
- Chat 历史仍可作为请求上下文传入 Agent,但不是本文档的主持久化模型。
|
||||
- Redis 主会话存储是历史设计,不作为当前架构事实。
|
||||
- `RetrievedDocTracker` 是 session 级运行时去重状态,诊断结束后清理。
|
||||
- Redis 会话历史用于多轮上下文,不是长期审计记录。
|
||||
- MySQL `diagnosis_run + agent_step + tool_invocation` 是主要可回放来源。
|
||||
- `chat_session.expires_at` 只是目录元数据;Redis 消息历史可独立过期。
|
||||
- `RetrievedDocTracker` 仍是 session 级运行时去重状态,诊断结束后清理。
|
||||
|
||||
## 10. 后续增强
|
||||
|
||||
可考虑:
|
||||
|
||||
1. Trace API 增加更结构化的 `self_evaluation` 展示。
|
||||
2. `agent_step` 与 `tool_invocation.step_id` 建立更严格关联。
|
||||
3. 对多轮同 session 诊断增加 run id,避免复用 session 时历史记录混杂。
|
||||
4. 为 Trace 增加导出能力,服务面试演示和回归分析。
|
||||
|
||||
3. 数据库中的旧 `diagnosis_session` 表按独立数据治理任务决定是否物理删除;当前应用不再依赖它。
|
||||
|
||||
@@ -0,0 +1,199 @@
|
||||
# Chat StateGraph 运行时架构
|
||||
|
||||
**更新日期**:2026-07-20
|
||||
**状态**:当前复杂 Chat 诊断的权威运行时架构
|
||||
**适用范围**:`POST /api/chat` 的复杂诊断路径;简单 Chat 和 AIOps 使用各自链路
|
||||
|
||||
## 1. 架构定位
|
||||
|
||||
复杂 Chat 已单轨切换为 Spring AI Alibaba bounded `StateGraph`。`ChatService` 负责 Run 生命周期和持久化,`ChatDiagnosisGraphRuntime` 负责执行 Graph,`DiagnosisGraphFactory` 负责声明 Node 与条件边,Agent/Java Node 负责各自的语义任务或确定性校验。
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
API["POST /api/chat"] --> Chat["ChatService.executeChatComplex"]
|
||||
Chat --> Session["chat_session"]
|
||||
Chat --> Run["diagnosis_run: RUNNING"]
|
||||
Chat --> Actions["DiagnosisRealGraphActionsFactory"]
|
||||
Actions --> Runtime["ChatDiagnosisGraphRuntime"]
|
||||
Runtime --> Factory["DiagnosisGraphFactory"]
|
||||
Factory --> Graph["Compiled StateGraph"]
|
||||
|
||||
Graph --> Planner["PlannerNodeAdapter"]
|
||||
Graph --> Executor["ExecutorNodeAdapter"]
|
||||
Graph --> Gatekeeper["GatekeeperNode"]
|
||||
Graph --> Projection["VerifiedInputNode"]
|
||||
Graph --> Verifier["VerifierNodeAdapter"]
|
||||
Graph --> Retry["EvidenceRetryPrepareNode"]
|
||||
Graph --> Composer["ComposerNodeAdapter"]
|
||||
Graph --> Fallback["FallbackNode"]
|
||||
|
||||
Executor --> Tools["lookup_knowledge / logs / metrics"]
|
||||
Tools --> Invocations["tool_invocation"]
|
||||
Planner --> Steps["agent_step"]
|
||||
Executor --> Steps
|
||||
Verifier --> Steps
|
||||
Composer --> Steps
|
||||
|
||||
Graph --> Mapper["DiagnosisGraphResultMapper"]
|
||||
Graph --> TraceBuilder["DiagnosisOrchestrationTraceBuilder"]
|
||||
Mapper --> SelfEval["diagnosis_run.self_evaluation"]
|
||||
TraceBuilder --> RouteTrace["diagnosis_run.orchestration_trace"]
|
||||
Chat --> RunDone["diagnosis_run: SUCCESS / FAILED"]
|
||||
RunDone --> TraceAPI["exact Run Trace API"]
|
||||
SelfEval --> TraceAPI
|
||||
RouteTrace --> TraceAPI
|
||||
Steps --> TraceAPI
|
||||
Invocations --> TraceAPI
|
||||
```
|
||||
|
||||
## 2. 运行生命周期
|
||||
|
||||
一次复杂 Chat 运行按以下顺序执行:
|
||||
|
||||
1. `ChatService` 解析或创建 `sessionId`,生成唯一 `runId`。
|
||||
2. 确保 `chat_session` 元数据存在,并创建 `diagnosis_run`,初始状态为 `RUNNING`、`agent_flow=CHAT`。
|
||||
3. 构建 Planner、Executor、Verifier、Composer 四个 `ReactAgent`,再由 `DiagnosisRealGraphActionsFactory` 组合 Java Nodes。
|
||||
4. `ChatDiagnosisGraphRuntime` 以 `runId` 作为 Graph `threadId`,将 `sessionId/runId` 放入 `RunnableConfig.metadata`。
|
||||
5. `DiagnosisGraphFactory` 编译 StateGraph 并执行,Graph recursion limit 固定为 32。
|
||||
6. Graph 返回非空 `final_answer` 后,`DiagnosisGraphResultMapper` 生成 verifier evaluation,`DiagnosisOrchestrationTraceBuilder` 压缩路由摘要。
|
||||
7. `ChatService` 保存答案、耗时、步骤数、工具数、自评估和编排摘要,将 Run 标记为 `SUCCESS`。
|
||||
8. 未处理异常会尽力保存 partial state/partial trace,再将 Run 标记为 `FAILED`;能够生成安全 Fallback 的路径仍是 `SUCCESS`,并通过 `degraded=true` 表达质量降级。
|
||||
9. `finally` 清理本轮检索追踪和 session/run ThreadLocal,避免跨 Run 污染。
|
||||
|
||||
## 3. Graph 拓扑
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Start([START]) --> Planner[Planner]
|
||||
Planner -->|COMPLETED| Executor[Executor]
|
||||
Planner -->|technical retry once| Planner
|
||||
Planner -->|non-retryable / exhausted| Fallback[Fallback]
|
||||
|
||||
Executor -->|COMPLETED| Gatekeeper[Gatekeeper]
|
||||
Executor -->|INVALID_OUTPUT / TOOL_BLOCKED / FAILED| Fallback
|
||||
|
||||
Gatekeeper -->|PASS| VerifiedInput[Verified Input]
|
||||
Gatekeeper -->|LOW_CONFID + verified bindings| VerifiedInput
|
||||
Gatekeeper -->|REJECT / no verified binding| Fallback
|
||||
|
||||
VerifiedInput --> Verifier[Verifier]
|
||||
Verifier -->|technical retry once| Verifier
|
||||
Verifier -->|LOW_CONFID + critical gap + retry allowed| EvidenceRetry[Evidence Retry]
|
||||
EvidenceRetry -->|EVIDENCE_GAP_ONLY| Planner
|
||||
Verifier -->|completed and no retry| Composer[Composer]
|
||||
Verifier -->|non-retryable / exhausted| Fallback
|
||||
|
||||
Composer -->|COMPLETED| End([END])
|
||||
Composer -->|technical retry once| Composer
|
||||
Composer -->|non-retryable / exhausted| Fallback
|
||||
Fallback --> End
|
||||
```
|
||||
|
||||
路由规则:
|
||||
|
||||
| 节点 | 继续条件 | 重试 | 安全终止 |
|
||||
|---|---|---|---|
|
||||
| Planner | `COMPLETED` 进入 Executor | `INVALID_OUTPUT` / `RETRYABLE_FAILED` 最多一次技术重试 | `NON_RETRYABLE_FAILED` 或重试耗尽进入 Fallback |
|
||||
| Executor | 只有 `COMPLETED` 进入 Gatekeeper | 不做 Graph 技术重试 | 非法输出、工具阻断或执行失败进入 Fallback |
|
||||
| Gatekeeper | `PASS`,或 `LOW_CONFID` 且至少一个 verified binding | 不重试 | `REJECT` 或零 verified binding 进入 Fallback |
|
||||
| Verifier | 完成后由 `effective_verdict` 决定 Composer 或补证据 | 技术失败最多一次;证据补查最多一次 | 不可重试失败或技术重试耗尽进入 Fallback |
|
||||
| Composer | `COMPLETED` 结束 | 技术失败最多一次 | 不可重试失败或重试耗尽进入 Fallback |
|
||||
| Fallback | 生成非空确定性安全答复 | 不重试 | 直接结束并标记 `degraded=true` |
|
||||
|
||||
Evidence retry 只有同时满足以下条件才发生:
|
||||
|
||||
- `evidence_retry_count < 1`。
|
||||
- Gatekeeper 给出的 verifier verdict ceiling 仍允许 `PASS`。
|
||||
- Verifier 输出包含可提取的 critical evidence gap。
|
||||
|
||||
补证据时 Planner 进入 `EVIDENCE_GAP_ONLY`,`planner_retry_count` 重置;Executor 只执行增量查询,但重新输出完整 `executor_evidence_v2` 快照。
|
||||
|
||||
## 4. 状态与执行边界
|
||||
|
||||
### Graph State
|
||||
|
||||
Graph State 只保存跨 Node 的控制信息和结构化结果:
|
||||
|
||||
- `diagnosis_context`、`planner_plan`、`executor_output`。
|
||||
- `gatekeeper_result`、`verified_executor_output`、`verified_evidence`。
|
||||
- `verifier_output`、`composer_output`、`final_answer`。
|
||||
- Planner/Verifier/Composer 技术重试计数和 `evidence_retry_count`。
|
||||
- `orchestration_events` 有界追加事件。
|
||||
|
||||
默认状态键使用 replace strategy,只有 `orchestration_events` 使用 append strategy。事件在 Graph 边界存为 classloader-neutral Map:`node/outcome/reason_code/attempt`,避免 DevTools restart classloader 造成 record 类型身份不一致。
|
||||
|
||||
### RunnableConfig
|
||||
|
||||
外层 Graph config 使用:
|
||||
|
||||
- `threadId = runId`。
|
||||
- metadata 包含 `sessionId` 和 `runId`。
|
||||
|
||||
调用 nested `ReactAgent` 时,`ReactAgentDiagnosisInvoker` 创建独立 config,只保留业务身份与 store,不向子 Agent 传播外层 Graph 的 human-feedback、state-update、checkpoint/resume 控制 metadata,避免父 Graph 恢复语义污染子 Graph。
|
||||
|
||||
## 5. 证据信任边界
|
||||
|
||||
```text
|
||||
Executor output
|
||||
-> source_invocation_id + raw_path + evidence_excerpt
|
||||
-> GatekeeperNode / ExecutorGatekeeperService
|
||||
-> 按当前 runId 读取 tool_invocation
|
||||
-> 验证 invocation ownership、raw_path、excerpt
|
||||
-> VerifiedInputNode
|
||||
-> 只投影通过的 claims/bindings/evidence
|
||||
-> VerifierNodeAdapter
|
||||
-> 只判断已验真证据是否支持 claim
|
||||
-> ComposerNodeAdapter
|
||||
-> 只表达允许输出的结论、限制和建议
|
||||
```
|
||||
|
||||
Verifier 不读取完整工具 Trace,不执行新检索,也不读取 Skill 正文。当前版本不生成或读取 `tool_trace_summary`。
|
||||
|
||||
## 6. Run 级审计模型
|
||||
|
||||
| 审计层 | 存储/API | 回答的问题 |
|
||||
|---|---|---|
|
||||
| 执行明细 | `agent_step`、`tool_invocation` | 模型和工具实际做了什么? |
|
||||
| 证据与答案质量 | `diagnosis_run.self_evaluation` | 引用是否真实、claim 是否可推导、Prompt/Gatekeeper 版本是什么? |
|
||||
| Graph 路由 | `diagnosis_run.orchestration_trace` / `run.orchestrationTrace` | 走了哪些 Node、为何重试或降级、在哪里结束? |
|
||||
|
||||
`orchestration_trace` 当前结构:
|
||||
|
||||
```json
|
||||
{
|
||||
"version": "stategraph-v1",
|
||||
"transitions": [
|
||||
{"from": "planner", "to": "executor", "reason_code": "completed", "attempt": 1}
|
||||
],
|
||||
"final_node": "composer",
|
||||
"termination_reason": "composer_completed",
|
||||
"degraded": false,
|
||||
"evidence_retry_count": 0
|
||||
}
|
||||
```
|
||||
|
||||
该字段只属于 exact Run 投影,响应不提供兼容 `session` 投影;非 StateGraph Run 可以为空。
|
||||
|
||||
## 7. 验收层次
|
||||
|
||||
| 测试层 | 权威测试 | 覆盖 |
|
||||
|---|---|---|
|
||||
| Workflow | `DiagnosisGraphWorkflowTest` | 全部分支、有限重试、补证据和 Fallback |
|
||||
| Node Contract | `DiagnosisGraphNodeContractTest` | 真实 Node 的输入投影、输出状态和证据边界 |
|
||||
| Runtime | `ChatDiagnosisGraphRuntimeTest` | config、最终状态、partial failure/trace |
|
||||
| Chat Integration | `ChatServiceGraphIntegrationTest` | Run 生命周期、答案、自评估、编排摘要和失败持久化 |
|
||||
| Trace Contract | `DiagnosisTraceServiceTest` | exact run ownership 和 `run.orchestrationTrace` 投影 |
|
||||
| Demo Contract | `InterviewDemoScriptContractTest` | exact runId、Graph 字段和 summary 输出 |
|
||||
|
||||
## 8. 关键代码
|
||||
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- `src/main/java/com/superbiz/agent/graph/diagnosis/ChatDiagnosisGraphRuntime.java`
|
||||
- `src/main/java/com/superbiz/agent/graph/diagnosis/DiagnosisGraphFactory.java`
|
||||
- `src/main/java/com/superbiz/agent/graph/diagnosis/DiagnosisGraphRouter.java`
|
||||
- `src/main/java/com/superbiz/agent/graph/diagnosis/DiagnosisGraphState.java`
|
||||
- `src/main/java/com/superbiz/agent/graph/diagnosis/DiagnosisRealGraphActionsFactory.java`
|
||||
- `src/main/java/com/superbiz/agent/graph/diagnosis/ReactAgentDiagnosisInvoker.java`
|
||||
- `src/main/java/com/superbiz/agent/graph/diagnosis/DiagnosisGraphResultMapper.java`
|
||||
- `src/main/java/com/superbiz/agent/graph/diagnosis/DiagnosisOrchestrationTraceBuilder.java`
|
||||
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
|
||||
@@ -0,0 +1,17 @@
|
||||
# 2026-07-09 MVP 文档清理归档
|
||||
|
||||
本目录保存本次清理中从当前入口移出的历史设计材料。这些文档仍有追溯价值,但不再代表当前可运行实现。
|
||||
|
||||
## 归档内容
|
||||
|
||||
| 目录 | 内容 | 归档原因 |
|
||||
|---|---|---|
|
||||
| `discuss/` | 早期 Executor Prompt、L0、RAG 讨论稿 | 已被当前 architecture、OpenSpec change 和 issue 取代 |
|
||||
| `plan/` | `session-storage-design.md` | 会话存储已实现,当前表以 Flyway 和 `mvp/tables/` 为准 |
|
||||
| `notes/` | 早期工程决策和 Demo Trace 验收笔记 | 相关内容已沉淀到 architecture、demo、eval 和 devflow |
|
||||
|
||||
## 使用原则
|
||||
|
||||
- 当前架构以 `mvp/architecture/` 为准。
|
||||
- 当前表结构以 `mvp/tables/`、Flyway migration 和实体类为准。
|
||||
- 当前问题入口以 `mvp/issues/README.md` 为准。
|
||||
@@ -0,0 +1,19 @@
|
||||
# 2026-07-20 MVP 文档清理归档
|
||||
|
||||
本目录保存已被当前 StateGraph、Executor Evidence Pipeline 和 Run-only v2 架构替代的历史材料。归档文件仅用于追溯,不代表当前运行时、接口或数据模型契约。
|
||||
|
||||
## 归档清单
|
||||
|
||||
| 原路径 | 归档文件 | 归档原因 | 当前依据 |
|
||||
|---|---|---|---|
|
||||
| `mvp/.backup/database-design-backup-20240622.md` | [database-design-backup-20240622.md](database/database-design-backup-20240622.md) | 早期数据库设计备份,模型和表关系已过时 | [data-model.md](../../architecture/data-model.md) |
|
||||
| `mvp/issues/design-notes/executor-self-evidence-loop-design-note.md` | [executor-self-evidence-loop-design-note.md](issues/design-notes/executor-self-evidence-loop-design-note.md) | 早期问题分析已被可执行证据链契约吸收 | [executor-evidence-pipeline-refactor.md](../../architecture/executor-evidence-pipeline-refactor.md) |
|
||||
| `mvp/issues/design-notes/executor-structured-output-v2.md` | [executor-structured-output-v2.md](issues/design-notes/executor-structured-output-v2.md) | 分阶段实施说明已完成并由当前契约和 OpenSpec 接管 | [executor-evidence-pipeline-refactor.md](../../architecture/executor-evidence-pipeline-refactor.md)、[OpenSpec 主规格](../../../openspec/specs/) |
|
||||
| `mvp/tables/诊断会话表-diagnosis_session.md` | [诊断会话表-diagnosis_session.md](tables/诊断会话表-diagnosis_session.md) | 运行时已不再映射、读取或写入旧会话级诊断表 | [session-trace-lifecycle.md](../../architecture/session-trace-lifecycle.md)、[诊断运行表-diagnosis_run.md](../../tables/诊断运行表-diagnosis_run.md) |
|
||||
|
||||
## 使用约束
|
||||
|
||||
- 当前架构从 `mvp/architecture/README.md` 进入。
|
||||
- 当前问题状态从 `mvp/issues/README.md` 进入。
|
||||
- 当前表模型从 `mvp/tables/README.md` 进入。
|
||||
- 归档中的兼容、回退和阶段状态描述均为历史快照,不得用于推导当前行为。
|
||||
+2
@@ -1,5 +1,7 @@
|
||||
# 数据库设计文档
|
||||
|
||||
> 归档说明:这是 2024 年数据库设计备份,已被当前 `chat_session + diagnosis_run + run-scoped trace detail` 模型替代,仅用于历史追溯。
|
||||
|
||||
## 一、设计原则
|
||||
|
||||
### 1.1 核心原则
|
||||
+2
@@ -1,5 +1,7 @@
|
||||
# Executor 自证循环与证据摘要链路设计记录
|
||||
|
||||
> 归档说明:本文是实施前的问题分析。当前实现以 `mvp/architecture/executor-evidence-pipeline-refactor.md` 和 OpenSpec 主规格为准。
|
||||
|
||||
**状态**:已形成方向,待创建 OpenSpec change
|
||||
**严重程度**:高
|
||||
**记录时间**:2026-07-07
|
||||
+2
@@ -1,5 +1,7 @@
|
||||
# Executor Structured Output V2 可执行设计与实施 Issue
|
||||
|
||||
> 归档说明:本文记录旧的分阶段实施方案,其中的兼容期和待启动状态不再适用。当前实现以 `mvp/architecture/executor-evidence-pipeline-refactor.md` 和 OpenSpec 主规格为准。
|
||||
|
||||
**状态**:阶段四待启动,前三阶段已归档并提交
|
||||
**严重程度**:高
|
||||
**创建日期**:2026-07-07
|
||||
@@ -0,0 +1,50 @@
|
||||
# 诊断会话表:diagnosis_session
|
||||
|
||||
> 归档说明:本文记录 Run-only v2 之前的旧表语义。当前 Java 运行时不再映射、读取或写入 `diagnosis_session`;数据库中物理表是否保留由独立数据治理任务决定。
|
||||
|
||||
**状态**:历史快照,不属于当前版本契约
|
||||
**历史来源**:`V005__create_session_storage.sql`、`V008__add_answer_to_diagnosis_session.sql`
|
||||
|
||||
## 定位
|
||||
|
||||
`diagnosis_session` 是旧版 session 级诊断主记录。`V011` 之后,新 Chat/AIOps 执行的运行态写入已经切到 `chat_session + diagnosis_run`;本表保留用于历史兼容、迁移回填和回滚比较。
|
||||
|
||||
## 字段
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
|---|---|---|---|
|
||||
| `id` | BIGINT | 是 | 自增主键 |
|
||||
| `session_id` | VARCHAR(64) | 是 | 旧版会话唯一 ID,也是兼容 run 回填来源 |
|
||||
| `query` | TEXT | 是 | 用户原始问题或 AIOps 输入摘要 |
|
||||
| `status` | VARCHAR(16) | 否 | `PENDING`、`RUNNING`、`SUCCESS`、`FAILED` |
|
||||
| `agent_flow` | VARCHAR(32) | 否 | `CHAT` 或 `AI_OPS` |
|
||||
| `total_duration_ms` | INT | 否 | 总耗时,单位毫秒 |
|
||||
| `total_token_count` | INT | 否 | 总 Token 消耗 |
|
||||
| `step_count` | INT | 否 | Agent 步骤数 |
|
||||
| `tool_call_count` | INT | 否 | 工具调用次数 |
|
||||
| `answer` | LONGTEXT | 否 | 返回给用户的最终答案 |
|
||||
| `self_evaluation` | JSON | 否 | rule、verifier、aiops 等自评估结果容器 |
|
||||
| `feedback` | VARCHAR(16) | 否 | 用户反馈:`useful`、`not_useful` 或空 |
|
||||
| `created_at` | DATETIME | 是 | 创建时间 |
|
||||
| `updated_at` | DATETIME | 是 | 更新时间 |
|
||||
|
||||
## 索引
|
||||
|
||||
| 索引 | 字段 | 用途 |
|
||||
|---|---|---|
|
||||
| `session_id` unique | `session_id` | 会话唯一约束 |
|
||||
| `idx_created_at` | `created_at` | 按时间查询 |
|
||||
| `idx_status` | `status` | 按状态筛选 |
|
||||
| `idx_agent_flow` | `agent_flow` | 区分 Chat / AIOps |
|
||||
|
||||
## 关系
|
||||
|
||||
- 迁移时每条 `diagnosis_session` 会生成一条兼容 `diagnosis_run`。
|
||||
- 历史 `agent_step.run_id` 和 `tool_invocation.run_id` 会尽量从兼容 `diagnosis_run` 回填。
|
||||
- 历史自动案例可能仍使用 `case_library.diagnosis_id = diagnosis_session.session_id`。
|
||||
|
||||
## 注意点
|
||||
|
||||
- 新执行不应再把 query、answer、self_evaluation、feedback、统计计数写入本表。
|
||||
- 新 Trace 聚合优先读取 `diagnosis_run + agent_step.run_id + tool_invocation.run_id`。
|
||||
- 历史 fallback 仅在没有 run-backed 数据时读取本表。
|
||||
+53
-10
@@ -6,9 +6,15 @@
|
||||
|
||||
- `ten-minute-interview-demo.md`:10 分钟现场演示脚本。
|
||||
- `interview-walkthrough.md`:面试讲解话术。
|
||||
- `evidence-pipeline-scenarios.md`:PASS / LOW_CONFID / REJECT / no-evidence 场景矩阵。
|
||||
- `trace-inspection-checklist.md`:Trace 字段检查清单。
|
||||
- `scripts/run-interview-demo-check.ps1`:面试预检脚本,绑定 exact runId,强制校验 Run orchestration trace,并输出 Chat、Trace、反馈和 summary。
|
||||
- `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。
|
||||
- `interview-q-and-a.md`:面试追问回答,覆盖 Agent 工程取舍、审计和评测。
|
||||
- `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。
|
||||
- `requests/narrow-highcpu-chat.json`:窄范围正向观察请求。
|
||||
- `requests/hikari-no-evidence-chat.json`:no-evidence 负向观察请求。
|
||||
- `requests/safety-unsupported-claim-chat.json`:安全降级讨论请求。
|
||||
|
||||
## 1. 前置条件
|
||||
|
||||
@@ -33,7 +39,7 @@ http://localhost:9900
|
||||
最快方式:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1
|
||||
```
|
||||
|
||||
脚本会生成:
|
||||
@@ -42,8 +48,11 @@ powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-de
|
||||
mvp/demo/output/chat-response.json
|
||||
mvp/demo/output/trace-response.json
|
||||
mvp/demo/output/feedback-response.json
|
||||
mvp/demo/output/interview-demo-summary.json
|
||||
```
|
||||
|
||||
自动化验收应传入唯一 `-SessionId`,并用 `-OutputDir target/...` 避免覆盖仓库样例。脚本从 Chat 响应取得 exact `runId`,缺少 `data.run.orchestrationTrace` 或 version/final node/termination reason/transitions/degraded/evidence retry count 时会立即失败。summary 额外包含 `orchestrationVersion`、`finalNode`、`terminationReason`、`degraded`、`transitionCount` 和 `evidenceRetryCount`。
|
||||
|
||||
手动请求:
|
||||
|
||||
```powershell
|
||||
@@ -60,10 +69,23 @@ Invoke-RestMethod `
|
||||
-Body $body
|
||||
```
|
||||
|
||||
如果要继续手动查询同一次诊断运行,先保留响应中的 run id:
|
||||
|
||||
```powershell
|
||||
$chat = Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/chat" `
|
||||
-ContentType "application/json" `
|
||||
-Body $body
|
||||
|
||||
$runId = $chat.data.runId
|
||||
```
|
||||
|
||||
期望结果:
|
||||
|
||||
- `data.success = true`
|
||||
- `data.sessionId = mvp-demo-payment-timeout-001`
|
||||
- `data.runId` 为本次诊断运行的唯一 ID
|
||||
- `data.answer` 包含诊断答复
|
||||
|
||||
## 4. 查询 Trace
|
||||
@@ -71,22 +93,31 @@ Invoke-RestMethod `
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace"
|
||||
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace?runId=$runId"
|
||||
```
|
||||
|
||||
期望结果:
|
||||
|
||||
- `code = 200`
|
||||
- `data.runId` 等于 `$runId`
|
||||
- `data.session.sessionId` 等于 Chat session id
|
||||
- `data.run.runId` 等于 `$runId`
|
||||
- `data.run.orchestrationTrace.version` 非空
|
||||
- `data.run.orchestrationTrace.final_node` 和 `termination_reason` 非空
|
||||
- `data.run.orchestrationTrace.transitions` 是本次 Graph 的条件边记录
|
||||
- `data.run.orchestrationTrace.degraded` 和 `evidence_retry_count` 记录安全降级与补证据次数
|
||||
- `data.steps` 包含 planner / executor / verifier 等步骤
|
||||
- `data.toolInvocations` 包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具
|
||||
- `data.session.selfEvaluation` 包含 verifier 或 rule evaluation
|
||||
- Chat V2 链路中,`data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` 记录 Chat Prompt 审计版本
|
||||
- Chat V2 链路中,`data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` 记录 Gatekeeper 规则集版本
|
||||
|
||||
## 5. 提交反馈
|
||||
|
||||
```powershell
|
||||
$feedback = @{
|
||||
sessionId = $sessionId
|
||||
runId = $runId
|
||||
feedback = "useful"
|
||||
} | ConvertTo-Json
|
||||
|
||||
@@ -100,7 +131,8 @@ Invoke-RestMethod `
|
||||
期望结果:
|
||||
|
||||
- `success = true`
|
||||
- 后续 Trace 中 `data.session.feedback = useful`
|
||||
- `runId = $runId`
|
||||
- 后续精确 Trace 中 `data.session.feedback = useful`
|
||||
- useful 反馈会尝试沉淀 `case_library`
|
||||
|
||||
## 6. AIOps 告警诊断 Demo
|
||||
@@ -126,19 +158,19 @@ Invoke-WebRequest `
|
||||
|
||||
期望结果:
|
||||
|
||||
- SSE 首条包含 `session` 消息,sessionId 为 `mvp-demo-aiops-payment-cpu-001`
|
||||
- SSE 首条是 `type=metadata` 的 `message` 事件,包含 sessionId `mvp-demo-aiops-payment-cpu-001` 和本次 AIOps `runId`
|
||||
- 后续流式输出包含 AIOps 告警分析报告
|
||||
- 报告聚焦输入的 `HighCPUUsage/payment-service`
|
||||
- 同一 session 的 Trace 中 `data.session.agentFlow = AI_OPS`
|
||||
- 精确 Trace 中 `data.session.agentFlow = AI_OPS`
|
||||
- `data.session.answer` 包含最终告警报告
|
||||
- `data.toolInvocations` 包含证据工具调用
|
||||
|
||||
查询 AIOps Trace:
|
||||
查询 AIOps Trace 时优先使用 SSE metadata 中的 runId:
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace"
|
||||
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace?runId=$aiopsRunId"
|
||||
```
|
||||
|
||||
## 7. Demo 主线
|
||||
@@ -146,11 +178,12 @@ Invoke-RestMethod `
|
||||
Chat 主线:
|
||||
|
||||
```text
|
||||
一个 session id
|
||||
一个 session id + 一个 run id
|
||||
-> 用户问题
|
||||
-> 多 Agent 执行
|
||||
-> bounded StateGraph(Planner / Executor / Gatekeeper / Verified Input / Verifier / Composer / Fallback)
|
||||
-> 证据工具
|
||||
-> Verifier / self_evaluation
|
||||
-> run.orchestrationTrace 路由摘要
|
||||
-> 最终答案
|
||||
-> 用户反馈
|
||||
-> Trace API 回放
|
||||
@@ -159,7 +192,7 @@ Chat 主线:
|
||||
AIOps 主线:
|
||||
|
||||
```text
|
||||
一个 session id
|
||||
一个 session id + 一个 run id
|
||||
-> 告警 payload
|
||||
-> AIOps Planner / Executor
|
||||
-> 证据工具
|
||||
@@ -167,3 +200,13 @@ AIOps 主线:
|
||||
-> AIOps rule evaluation
|
||||
-> Trace API 回放
|
||||
```
|
||||
|
||||
## 8. Evidence Pipeline 场景矩阵
|
||||
|
||||
面试时不要把所有安全场景都压到 live LLM 现场表现上。建议使用:
|
||||
|
||||
- `scripts/run-interview-demo-check.ps1` 跑主路径和预检 summary。
|
||||
- `evidence-pipeline-scenarios.md` 讲解 PASS / LOW_CONFID / REJECT / no-evidence 矩阵。
|
||||
- `mvp/eval/reports/baseline-report.md` 证明固定 fixture 12/12 通过。
|
||||
|
||||
这样可以同时展示真实链路和确定性回归能力。
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
## 1. 目标
|
||||
|
||||
验证旧版 `/api/ai_ops` 入口可以作为可追踪的告警触发诊断入口,并且 payload 模式下报告聚焦输入告警。
|
||||
验证 `/api/ai_ops` 入口可以作为可追踪的告警触发诊断入口,并且 payload 模式下报告聚焦输入告警。
|
||||
|
||||
## 2. 输入
|
||||
|
||||
@@ -25,14 +25,14 @@
|
||||
|
||||
## 3. 验收标准
|
||||
|
||||
1. SSE 流输出 `session` 消息,且包含请求中的 session id。
|
||||
2. AIOps 执行创建或更新 `diagnosis_session`,并写入 `agent_flow = AI_OPS`。
|
||||
3. 持久化的 session query 包含告警名、服务名、等级、时间范围和描述。
|
||||
4. 如果生成最终报告,`diagnosis_session.answer` 包含该报告。
|
||||
5. `GET /api/diagnosis/{sessionId}/trace` 返回 AIOps session、按顺序排列的 agent steps 和 tool invocations。
|
||||
1. SSE 流首条输出 `type=metadata` 的 `message` 事件,且包含请求中的 session id 和本次 AIOps run id。
|
||||
2. AIOps 执行创建 `diagnosis_run`,并写入 `agent_flow = AI_OPS`。
|
||||
3. 持久化的 run query 包含告警名、服务名、等级、时间范围和描述。
|
||||
4. 如果生成最终报告,`diagnosis_run.answer` 包含该报告。
|
||||
5. `GET /api/diagnosis/{sessionId}/trace?runId=...` 返回 AIOps run、按顺序排列的 agent steps 和 tool invocations。
|
||||
6. payload 模式下,报告主线聚焦 `HighCPUUsage/payment-service`。
|
||||
7. 其他活跃告警最多作为相关风险或上下文出现,不应展开成完整独立根因章节。
|
||||
8. `self_evaluation.aiops_rule_evaluation` 存在,并能反映报告完整性、payload 聚焦和证据工具覆盖情况。
|
||||
8. `diagnosis_run.self_evaluation.aiops_rule_evaluation` 存在,并能反映报告完整性、payload 聚焦和证据工具覆盖情况。
|
||||
|
||||
## 4. 已知边界
|
||||
|
||||
|
||||
@@ -0,0 +1,63 @@
|
||||
# Evidence Pipeline Demo Scenarios
|
||||
|
||||
这份清单用于面试时说明 Chat 证据链路如何覆盖 `PASS`、`LOW_CONFID`、`REJECT` 和 no-evidence 场景。
|
||||
|
||||
重点区别:
|
||||
|
||||
- Live demo 证明本地服务、工具、Trace、Feedback 主链路能跑通。
|
||||
- Fixture-backed eval 证明固定安全场景可以确定性回归,不依赖 LLM 当场随机输出。
|
||||
|
||||
## Scenario Matrix
|
||||
|
||||
| 场景 | 类型 | 输入/证据 | 期望讲点 |
|
||||
|---|---|---|---|
|
||||
| Payment timeout | Live 主路径 | `requests/payment-timeout-chat.json` | 完整 Chat -> Trace -> Feedback 闭环 |
|
||||
| Narrow HighCPU observation | Live 可尝试 + fixture-backed | `requests/narrow-highcpu-chat.json` / `mvp/eval/fixtures/narrow-highcpu-observation-pass.json` | Executor 只输出观察类 claim,Gatekeeper 验引用,Verifier PASS |
|
||||
| Hikari no-evidence | Live 可尝试 + fixture-backed | `requests/hikari-no-evidence-chat.json` / `mvp/eval/fixtures/hikari-no-evidence-negative-observation-pass.json` | `$.no_evidence` 只表示本次查询无匹配证据,Composer 不说“已排除” |
|
||||
| Unsupported claim filtering | Fixture-backed | `requests/safety-unsupported-claim-chat.json` / `mvp/eval/fixtures/unsupported-claim-filtering-low-confid.json` | Verifier 将 unsupported claim 降为 LOW_CONFID,最终答案不确认“主库故障” |
|
||||
| Fabricated invocation reject | Fixture-backed | `mvp/eval/fixtures/gatekeeper-fabricated-invocation-reject.json` | Gatekeeper 拦截伪造 invocation,最终 REJECT/降级 |
|
||||
| Composer fallback | Fixture-backed | `mvp/eval/fixtures/composer-fallback-no-raw-json-low-confid.json` | 即使 Composer 输出异常,也不能把 Executor JSON 泄漏给用户 |
|
||||
|
||||
## Trace Fields To Inspect
|
||||
|
||||
| 能力 | JSON path |
|
||||
|---|---|
|
||||
| Executor V2 输出 | `data.session.selfEvaluation.verifier_evaluation.executor_structured_output` |
|
||||
| Gatekeeper 结果 | `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.status` |
|
||||
| Gatekeeper 规则版本 | `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` |
|
||||
| 证据绑定校验 | `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.checked_bindings` |
|
||||
| Verifier claim checks | `data.session.selfEvaluation.verifier_evaluation.claim_checks` |
|
||||
| Composer 输出 | `data.session.selfEvaluation.verifier_evaluation.composer_output` |
|
||||
| 工具证据引用 | `data.toolInvocations[*].retrievalDetails.evidence_refs` |
|
||||
|
||||
## How To Present It
|
||||
|
||||
```text
|
||||
我把现场 demo 和固定 eval 分开。
|
||||
现场 demo 证明系统能跑通真实链路;
|
||||
fixture-backed eval 证明反幻觉安全场景可以稳定回归。
|
||||
Gatekeeper 的规则版本也进入 trace,所以后续调整阈值或规则时可以审计。
|
||||
```
|
||||
|
||||
## Optional Live Requests
|
||||
|
||||
手动发送某个请求样例:
|
||||
|
||||
```powershell
|
||||
$body = Get-Content -Raw -Encoding UTF8 "mvp/demo/requests/narrow-highcpu-chat.json"
|
||||
Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/chat" `
|
||||
-ContentType "application/json" `
|
||||
-Body $body
|
||||
```
|
||||
|
||||
然后保留响应里的 `runId`,查询同一 run 的 Trace:
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/mvp-demo-narrow-highcpu-001/trace?runId=$runId"
|
||||
```
|
||||
|
||||
注意:除 payment-timeout 主路径外,其它 live 请求是“可尝试”的演示入口;稳定验收以 `mvp/eval` fixture 和 baseline 为准。
|
||||
@@ -0,0 +1,33 @@
|
||||
# 面试追问 Q&A
|
||||
|
||||
## 为什么不用普通 Chatbot?
|
||||
|
||||
这个项目的重点不是生成一段诊断文本,而是把诊断拆成可审计链路:Planner 拆解问题,Executor 调工具拿证据,Gatekeeper 用代码核验证据引用,Verifier 判断可推导性,Composer 生成最终表达。`sessionId` 保留多轮上下文,`runId` 精确绑定一次诊断运行,Trace 和 Feedback 都可以按 `sessionId + runId` 回放和定位。
|
||||
|
||||
## 为什么 RAG 要做成显式工具?
|
||||
|
||||
`lookup_knowledge` 保持显式工具调用,才能在 `tool_invocation` 里看到 Agent 查了什么、命中了什么、相关性等级是什么,以及最终答案是否真的使用了这些证据。隐式 Advisor 更方便,但不利于审计 Agent 决策。
|
||||
|
||||
## 怎么防止 Executor 幻觉?
|
||||
|
||||
Executor 不直接负责最终用户答案,而是输出 `executor_evidence_v2` 的微观事实和证据引用。Gatekeeper 会校验 `source_invocation_id`、`raw_path`、`evidence_excerpt` 是否真实存在;Verifier 再判断 claim 是否能由已验真的证据推出;Composer 只表达 Verifier 允许输出的内容。
|
||||
|
||||
## LOW_CONFID 是失败吗?
|
||||
|
||||
不是。`LOW_CONFID` 表示当前证据不足以支撑强结论,但系统仍然可以安全表达已确认事实和缺失信息。面试时可以把它作为“没有证据就不强答”的质量门禁,而不是模型能力失败。
|
||||
|
||||
## Prompt 改了怎么审计?
|
||||
|
||||
Chat verifier evaluation 里会记录 `prompt_audit.version`,并列出 planner、executor、verifier、composer 的 Prompt 版本和资源路径。它不保存完整 Prompt 文本,只保留用于回放和回归解释的紧凑元数据。
|
||||
|
||||
## Gatekeeper 改了怎么审计?
|
||||
|
||||
Gatekeeper 结果里记录 `gatekeeper_result.rule_set_version` 和已启用规则元数据摘要。规则执行仍是确定性 Java 代码,版本和规则元数据用于解释“这次引用验真用的是哪套规则”。
|
||||
|
||||
## 为什么现在不拆 SubAgent?
|
||||
|
||||
当前 MVP 的主要风险不是 Agent 数量不够,而是证据、验证和回归是否稳定。文档里的演进路线把 SubAgent 放在 P2:等故障类型、工具权限和评测集足够明确后再拆,避免只是移动复杂度。
|
||||
|
||||
## 为什么 baseline 比 live demo 更重要?
|
||||
|
||||
live demo 证明链路在当前环境能跑通,但 LLM 和外部依赖会波动。`mvp/eval` 的固定 fixture baseline 是确定性回归来源,用来判断 Prompt、工具、Gatekeeper、Verifier 或 Composer 的改动有没有让系统退化。
|
||||
@@ -13,7 +13,7 @@
|
||||
关键主张不是“模型回答了一次”,而是:
|
||||
|
||||
```text
|
||||
系统能展示用了什么证据、答案如何被检查、如何用 sessionId 回放整次诊断。
|
||||
系统能展示用了什么证据、答案如何被检查、如何用 sessionId + runId 精确回放这次诊断。
|
||||
```
|
||||
|
||||
## 2. Demo 流程
|
||||
@@ -23,7 +23,8 @@
|
||||
3. 打开 `mvp/demo/output/chat-response.json`。
|
||||
4. 打开 `mvp/demo/output/trace-response.json`。
|
||||
5. 指出证据工具和 verifier evaluation。
|
||||
6. 提交 feedback,并展示它挂在同一个 session 上。
|
||||
6. 提交 feedback,并展示它挂在当前 run 上。
|
||||
7. 打开 `evidence-pipeline-scenarios.md`,说明 PASS / LOW_CONFID / REJECT / no-evidence 的固定回归矩阵。
|
||||
|
||||
## 3. 命令
|
||||
|
||||
@@ -58,7 +59,7 @@ mvp/demo/output/chat-response.json
|
||||
话术:
|
||||
|
||||
```text
|
||||
这是用户看到的答案。这里的 sessionId 是稳定的,所以我后面可以追踪这一次回答是怎么来的。
|
||||
这是用户看到的答案。这里的 sessionId 是稳定的,同时响应里会返回 runId,所以我后面可以精确追踪这一次回答是怎么来的。
|
||||
```
|
||||
|
||||
### 4.2 证据 Trace
|
||||
@@ -111,7 +112,7 @@ mvp/demo/output/feedback-response.json
|
||||
话术:
|
||||
|
||||
```text
|
||||
feedback 会挂在同一个 diagnosis session 上。
|
||||
feedback 会挂在当前 diagnosis run 上。
|
||||
这让后续挖掘 useful case 或 not_useful bad case 成为可能。
|
||||
```
|
||||
|
||||
@@ -125,6 +126,14 @@ Demo 证明真实链路能跑通,offline eval baseline 证明固定 case 可
|
||||
这两者分开是有意的:Demo 面向人类审阅,eval 面向自动化信号。
|
||||
```
|
||||
|
||||
如果被问到怎么防止证据归因幻觉,可以补充:
|
||||
|
||||
```text
|
||||
Executor 的 claim 必须绑定 source_invocation_id、raw_path 和 evidence_excerpt。
|
||||
Gatekeeper 用代码核验这些引用,并把 rule_set_version 写进 trace。
|
||||
Verifier 只判断已核验证据能否推出 claim,Composer 只表达允许输出的内容。
|
||||
```
|
||||
|
||||
## 5. 强面试表达
|
||||
|
||||
```text
|
||||
@@ -141,4 +150,3 @@ traceability、evidence persistence、verifier gating、feedback 和 regression
|
||||
mvp-demo profile mock 了日志和指标,但不是完整生产运行环境。
|
||||
密钥清理、默认隔离测试和生产可靠性是后续 hardening 工作。
|
||||
```
|
||||
|
||||
|
||||
@@ -12,14 +12,16 @@
|
||||
|
||||
## 3. 验收标准
|
||||
|
||||
1. Chat 返回成功答复,且 session id 与请求一致。
|
||||
2. Trace API 返回 session 元数据、最终答案、按顺序排列的 agent steps 和 tool invocations。
|
||||
1. Chat 返回成功答复,且 session id 与请求一致,并返回本次诊断的 run id。
|
||||
2. Trace API 使用 `sessionId + runId` 返回会话元数据、运行摘要、最终答案、按顺序排列的 agent steps 和 tool invocations。
|
||||
3. Trace 中有足够证据说明用了哪些工具,以及 verifier / self-evaluation 是否已持久化。
|
||||
4. 可以使用同一个 session id 提交反馈。
|
||||
5. 后续 Trace 查询能看到已持久化的 feedback 值。
|
||||
4. 可以使用同一个 session id 和本次 run id 提交反馈。
|
||||
5. 后续精确 Trace 查询能看到已持久化的 feedback 值。
|
||||
|
||||
## 4. 需要检查的 Trace 字段
|
||||
|
||||
- `data.runId`
|
||||
- `data.run.runId`
|
||||
- `data.session.query`
|
||||
- `data.session.answer`
|
||||
- `data.session.selfEvaluation`
|
||||
|
||||
@@ -0,0 +1,4 @@
|
||||
{
|
||||
"Id": "mvp-demo-hikari-no-evidence-001",
|
||||
"Question": "确认 inventory-service 当前是否有 HikariCP 连接池耗尽日志。"
|
||||
}
|
||||
@@ -0,0 +1,4 @@
|
||||
{
|
||||
"Id": "mvp-demo-narrow-highcpu-001",
|
||||
"Question": "确认 payment-service 当前是否存在 HighCPUUsage 告警。"
|
||||
}
|
||||
@@ -0,0 +1,4 @@
|
||||
{
|
||||
"Id": "mvp-demo-safety-unsupported-001",
|
||||
"Question": "订单超时是否可以确认由数据库主库故障导致?请只基于当前证据回答。"
|
||||
}
|
||||
@@ -0,0 +1,233 @@
|
||||
param(
|
||||
[string]$BaseUrl = "http://localhost:9900",
|
||||
[string]$SessionId = "mvp-demo-interview-payment-timeout-001",
|
||||
[string]$RequestFile = "$PSScriptRoot/../requests/payment-timeout-chat.json",
|
||||
[string]$OutputDir = "$PSScriptRoot/../output"
|
||||
)
|
||||
|
||||
$ErrorActionPreference = "Stop"
|
||||
|
||||
function Test-ServiceReachable {
|
||||
param([string]$Url)
|
||||
|
||||
try {
|
||||
$request = [System.Net.WebRequest]::Create($Url)
|
||||
$request.Method = "GET"
|
||||
$request.Timeout = 5000
|
||||
$response = $request.GetResponse()
|
||||
$response.Close()
|
||||
return $true
|
||||
} catch [System.Net.WebException] {
|
||||
if ($_.Exception.Response -ne $null) {
|
||||
$_.Exception.Response.Close()
|
||||
return $true
|
||||
}
|
||||
return $false
|
||||
}
|
||||
}
|
||||
|
||||
function Get-TraceData {
|
||||
param($TraceResponse)
|
||||
|
||||
if ($TraceResponse.PSObject.Properties.Name -contains "data") {
|
||||
return $TraceResponse.data
|
||||
}
|
||||
return $TraceResponse
|
||||
}
|
||||
|
||||
function Get-SelfEvaluation {
|
||||
param($TraceData)
|
||||
|
||||
if ($null -eq $TraceData -or $null -eq $TraceData.session) {
|
||||
return $null
|
||||
}
|
||||
return $TraceData.session.selfEvaluation
|
||||
}
|
||||
|
||||
function Get-ToolNames {
|
||||
param($TraceData)
|
||||
|
||||
if ($null -eq $TraceData -or $null -eq $TraceData.toolInvocations) {
|
||||
return @()
|
||||
}
|
||||
return @($TraceData.toolInvocations | ForEach-Object { $_.toolName } | Where-Object { $_ } | Sort-Object -Unique)
|
||||
}
|
||||
|
||||
New-Item -ItemType Directory -Force -Path $OutputDir | Out-Null
|
||||
|
||||
Write-Host "Running interview demo preflight..."
|
||||
Write-Host "BaseUrl: $BaseUrl"
|
||||
Write-Host "SessionId: $SessionId"
|
||||
|
||||
if (-not (Test-ServiceReachable -Url $BaseUrl)) {
|
||||
throw "Service is not reachable: $BaseUrl. Start the app with mvp-demo profile first: mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo"
|
||||
}
|
||||
|
||||
$request = Get-Content -Raw -Encoding UTF8 -Path $RequestFile | ConvertFrom-Json
|
||||
$request.Id = $SessionId
|
||||
$body = $request | ConvertTo-Json -Depth 8
|
||||
|
||||
$chatRequest = @{
|
||||
Method = "Post"
|
||||
Uri = "$BaseUrl/api/chat"
|
||||
ContentType = "application/json; charset=utf-8"
|
||||
Body = $body
|
||||
}
|
||||
$chat = Invoke-RestMethod @chatRequest
|
||||
|
||||
$chatPath = Join-Path $OutputDir "chat-response.json"
|
||||
$chat | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $chatPath
|
||||
|
||||
$runId = $chat.data.runId
|
||||
if ($chat.data.success -ne $true) {
|
||||
throw "Chat response was not successful."
|
||||
}
|
||||
if ([string]::IsNullOrWhiteSpace([string]$chat.data.answer)) {
|
||||
throw "Chat response did not include a non-empty answer."
|
||||
}
|
||||
if ($chat.data.sessionId -ne $SessionId) {
|
||||
throw "Chat response sessionId '$($chat.data.sessionId)' did not match requested sessionId '$SessionId'."
|
||||
}
|
||||
if (-not $runId) {
|
||||
throw "Chat response did not include runId; exact trace verification cannot continue."
|
||||
}
|
||||
|
||||
$traceRequest = @{
|
||||
Method = "Get"
|
||||
Uri = "$BaseUrl/api/diagnosis/$SessionId/trace?runId=$([System.Uri]::EscapeDataString($runId))"
|
||||
}
|
||||
$trace = Invoke-RestMethod @traceRequest
|
||||
|
||||
$tracePath = Join-Path $OutputDir "trace-response.json"
|
||||
$trace | ConvertTo-Json -Depth 80 | Set-Content -Encoding UTF8 -Path $tracePath
|
||||
|
||||
$traceData = Get-TraceData -TraceResponse $trace
|
||||
if ($null -eq $traceData -or $null -eq $traceData.run) {
|
||||
throw "Exact Trace response did not include data.run."
|
||||
}
|
||||
if ($traceData.runId -ne $runId -or $traceData.run.runId -ne $runId) {
|
||||
throw "Exact Trace runId did not match Chat runId '$runId'."
|
||||
}
|
||||
if ($traceData.run.sessionId -ne $SessionId) {
|
||||
throw "Exact Trace run did not belong to requested sessionId '$SessionId'."
|
||||
}
|
||||
|
||||
$orchestrationTrace = $traceData.run.orchestrationTrace
|
||||
if ($null -eq $orchestrationTrace) {
|
||||
throw "Exact Trace data.run.orchestrationTrace is missing."
|
||||
}
|
||||
foreach ($field in @("version", "final_node", "termination_reason")) {
|
||||
if (-not ($orchestrationTrace.PSObject.Properties.Name -contains $field) -or
|
||||
[string]::IsNullOrWhiteSpace([string]$orchestrationTrace.$field)) {
|
||||
throw "Exact Trace data.run.orchestrationTrace.$field is missing."
|
||||
}
|
||||
}
|
||||
foreach ($field in @("transitions", "degraded", "evidence_retry_count")) {
|
||||
if (-not ($orchestrationTrace.PSObject.Properties.Name -contains $field)) {
|
||||
throw "Exact Trace data.run.orchestrationTrace.$field is missing."
|
||||
}
|
||||
}
|
||||
if ($null -eq $orchestrationTrace.transitions) {
|
||||
throw "Exact Trace data.run.orchestrationTrace.transitions must be an array."
|
||||
}
|
||||
if ([int]$orchestrationTrace.evidence_retry_count -lt 0) {
|
||||
throw "Exact Trace data.run.orchestrationTrace.evidence_retry_count must not be negative."
|
||||
}
|
||||
if ($traceData.run.status -ne "SUCCESS" -or $traceData.run.agentFlow -ne "CHAT") {
|
||||
throw "Exact Trace run must be CHAT/SUCCESS."
|
||||
}
|
||||
if ([string]::IsNullOrWhiteSpace([string]$traceData.run.answer)) {
|
||||
throw "Exact Trace run did not include a non-empty answer."
|
||||
}
|
||||
if (@($traceData.steps).Count -eq 0 -or @($traceData.toolInvocations).Count -eq 0) {
|
||||
throw "Exact Trace did not include both Agent steps and tool invocation evidence."
|
||||
}
|
||||
if ($null -eq $traceData.run.selfEvaluation) {
|
||||
throw "Exact Trace run did not include selfEvaluation."
|
||||
}
|
||||
|
||||
$feedbackBody = @{
|
||||
sessionId = $SessionId
|
||||
runId = $runId
|
||||
feedback = "useful"
|
||||
} | ConvertTo-Json
|
||||
|
||||
$feedbackRequest = @{
|
||||
Method = "Post"
|
||||
Uri = "$BaseUrl/api/feedback"
|
||||
ContentType = "application/json; charset=utf-8"
|
||||
Body = $feedbackBody
|
||||
}
|
||||
$feedback = Invoke-RestMethod @feedbackRequest
|
||||
if ($feedback.success -ne $true) {
|
||||
throw "Feedback request was not successful for runId '$runId'."
|
||||
}
|
||||
|
||||
$feedbackPath = Join-Path $OutputDir "feedback-response.json"
|
||||
$feedback | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $feedbackPath
|
||||
|
||||
$selfEvaluation = Get-SelfEvaluation -TraceData $traceData
|
||||
$verifierEvaluation = $null
|
||||
if ($null -ne $selfEvaluation) {
|
||||
$verifierEvaluation = $selfEvaluation.verifier_evaluation
|
||||
}
|
||||
|
||||
$gatekeeperResult = $null
|
||||
$promptAudit = $null
|
||||
if ($null -ne $verifierEvaluation) {
|
||||
$gatekeeperResult = $verifierEvaluation.gatekeeper_result
|
||||
$promptAudit = $verifierEvaluation.prompt_audit
|
||||
}
|
||||
|
||||
$verdict = $null
|
||||
$gatekeeperStatus = $null
|
||||
$gatekeeperRuleSetVersion = $null
|
||||
$promptAuditVersion = $null
|
||||
if ($null -ne $verifierEvaluation) {
|
||||
$verdict = $verifierEvaluation.verdict
|
||||
}
|
||||
if ($null -ne $gatekeeperResult) {
|
||||
$gatekeeperStatus = $gatekeeperResult.status
|
||||
$gatekeeperRuleSetVersion = $gatekeeperResult.rule_set_version
|
||||
}
|
||||
if ($null -ne $promptAudit) {
|
||||
$promptAuditVersion = $promptAudit.version
|
||||
}
|
||||
$toolNames = Get-ToolNames -TraceData $traceData
|
||||
$transitionCount = @($orchestrationTrace.transitions).Count
|
||||
$summaryPath = Join-Path $OutputDir "interview-demo-summary.json"
|
||||
|
||||
$summary = [ordered]@{
|
||||
sessionId = $SessionId
|
||||
runId = $runId
|
||||
baseUrl = $BaseUrl
|
||||
chatSuccess = $chat.data.success
|
||||
verdict = $verdict
|
||||
gatekeeperStatus = $gatekeeperStatus
|
||||
gatekeeperRuleSetVersion = $gatekeeperRuleSetVersion
|
||||
promptAuditVersion = $promptAuditVersion
|
||||
orchestrationVersion = $orchestrationTrace.version
|
||||
finalNode = $orchestrationTrace.final_node
|
||||
terminationReason = $orchestrationTrace.termination_reason
|
||||
degraded = [bool]$orchestrationTrace.degraded
|
||||
transitionCount = $transitionCount
|
||||
evidenceRetryCount = [int]$orchestrationTrace.evidence_retry_count
|
||||
toolNames = $toolNames
|
||||
paths = [ordered]@{
|
||||
chat = $chatPath
|
||||
trace = $tracePath
|
||||
feedback = $feedbackPath
|
||||
summary = $summaryPath
|
||||
}
|
||||
}
|
||||
|
||||
$summary | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path $summaryPath
|
||||
|
||||
Write-Host ""
|
||||
Write-Host "Interview demo preflight completed."
|
||||
Write-Host "Verdict: $($summary.verdict)"
|
||||
Write-Host "Gatekeeper rules: $($summary.gatekeeperRuleSetVersion)"
|
||||
Write-Host "Prompt audit: $($summary.promptAuditVersion)"
|
||||
Write-Host "Graph final node: $($summary.finalNode)"
|
||||
Write-Host "Graph termination: $($summary.terminationReason)"
|
||||
Write-Host "Summary: $summaryPath"
|
||||
@@ -26,15 +26,22 @@ $chat = Invoke-RestMethod `
|
||||
$chat | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-response.json"
|
||||
Write-Host "已保存 Chat 响应: $OutputDir/chat-response.json"
|
||||
|
||||
$runId = $chat.data.runId
|
||||
if (-not $runId) {
|
||||
throw "Chat 响应缺少 runId,无法查询精确 Trace。"
|
||||
}
|
||||
Write-Host "RunId: $runId"
|
||||
|
||||
$trace = Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "$BaseUrl/api/diagnosis/$SessionId/trace"
|
||||
-Uri "$BaseUrl/api/diagnosis/$SessionId/trace?runId=$([System.Uri]::EscapeDataString($runId))"
|
||||
|
||||
$trace | ConvertTo-Json -Depth 50 | Set-Content -Encoding UTF8 -Path "$OutputDir/trace-response.json"
|
||||
Write-Host "已保存 Trace 响应: $OutputDir/trace-response.json"
|
||||
|
||||
$feedbackBody = @{
|
||||
sessionId = $SessionId
|
||||
runId = $runId
|
||||
feedback = "useful"
|
||||
} | ConvertTo-Json
|
||||
|
||||
|
||||
@@ -37,7 +37,7 @@ http://localhost:9900
|
||||
推荐使用固定脚本:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1
|
||||
```
|
||||
|
||||
脚本会写出:
|
||||
@@ -46,13 +46,14 @@ powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-de
|
||||
mvp/demo/output/chat-response.json
|
||||
mvp/demo/output/trace-response.json
|
||||
mvp/demo/output/feedback-response.json
|
||||
mvp/demo/output/interview-demo-summary.json
|
||||
```
|
||||
|
||||
现场话术:
|
||||
|
||||
```text
|
||||
这里我用固定 sessionId 跑一个支付接口超时问题。
|
||||
固定 sessionId 的好处是,后面 trace 和 feedback 都能关联到同一次诊断。
|
||||
固定 sessionId 的好处是保留多轮上下文;每次诊断还会返回 runId,后面 trace 和 feedback 都用这个 runId 精确关联到同一次运行。
|
||||
```
|
||||
|
||||
## 3. 展示用户答案
|
||||
@@ -67,6 +68,7 @@ mvp/demo/output/chat-response.json
|
||||
|
||||
```text
|
||||
data.sessionId
|
||||
data.runId
|
||||
data.answer
|
||||
```
|
||||
|
||||
@@ -75,7 +77,7 @@ data.answer
|
||||
```text
|
||||
这是用户看到的答案。
|
||||
但这个项目的重点不是这段文字,而是这段文字是否有证据链。
|
||||
接下来我用同一个 sessionId 查 trace。
|
||||
接下来我用同一个 sessionId 加 runId 查 trace。
|
||||
```
|
||||
|
||||
## 4. 展示 Trace
|
||||
@@ -98,6 +100,8 @@ data.toolInvocations[*].outputPreview
|
||||
data.toolInvocations[*].retrievalLayer
|
||||
data.toolInvocations[*].relevanceLevel
|
||||
data.summary.hasVerifierEvaluation
|
||||
data.session.selfEvaluation.verifier_evaluation.prompt_audit.version
|
||||
data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version
|
||||
```
|
||||
|
||||
现场话术:
|
||||
@@ -155,6 +159,10 @@ Verifier 不做新检索,只看工具 trace 汇总。
|
||||
如果 PASS,就输出原答案。
|
||||
如果 LOW_CONFID,可以补证据或加低置信提示。
|
||||
如果 REJECT,就降级输出,只保留已确认信息。
|
||||
|
||||
Prompt 和 Gatekeeper 的版本也会进入 trace。
|
||||
`prompt_audit.version` 用于说明本次 Chat 使用哪套 Prompt 契约,`gatekeeper_result.rule_set_version` 用于说明引用验真的规则版本。
|
||||
固定 fixture baseline 是回归判断来源,live demo 主要证明当前环境链路可跑通。
|
||||
```
|
||||
|
||||
## 7. 展示反馈闭环
|
||||
@@ -175,7 +183,7 @@ caseId
|
||||
现场话术:
|
||||
|
||||
```text
|
||||
用户反馈 useful 会写回同一个 diagnosis_session。
|
||||
用户反馈 useful 会写回当前 diagnosis_run。
|
||||
后端会把这次诊断自动沉淀到 case_library,后续可以做案例检索或 bad case 分析。
|
||||
|
||||
这里 status 和 feedback 是分开的:
|
||||
@@ -226,6 +234,7 @@ AIOps 有两个模式。
|
||||
mvp/demo/output/chat-response.json
|
||||
mvp/demo/output/trace-response.json
|
||||
mvp/demo/output/feedback-response.json
|
||||
mvp/demo/output/interview-demo-summary.json
|
||||
```
|
||||
|
||||
降级话术:
|
||||
|
||||
@@ -1,18 +1,36 @@
|
||||
# Trace 检查清单
|
||||
|
||||
运行 `scripts/run-payment-timeout-demo.ps1` 后,用这份清单检查 `trace-response.json`。
|
||||
运行 `scripts/run-interview-demo-check.ps1` 后,用这份清单检查 `trace-response.json` 和 `interview-demo-summary.json`。
|
||||
|
||||
## 1. Session
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
| `data.session.sessionId` | 是否等于 `mvp-demo-payment-timeout-001` | 一个 session id 串起 chat、工具、verifier、feedback 和 trace |
|
||||
| `data.session.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
|
||||
| `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
|
||||
| `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
|
||||
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在同一次诊断上 |
|
||||
| `data.runId` / `data.run.runId` | 是否等于 demo 响应中的 `runId` | `runId` 精确绑定这一次诊断运行 |
|
||||
| `data.run.sessionId` | 是否等于本次 Chat 请求的唯一 sessionId | `sessionId` 保留多轮上下文,Trace 精确回放依赖 `runId` |
|
||||
| `data.run.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
|
||||
| `data.run.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
|
||||
| `data.run.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
|
||||
| `data.run.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | 是否记录 Gatekeeper 规则版本 | 安全规则可审计、可回归 |
|
||||
| `data.run.selfEvaluation.verifier_evaluation.prompt_audit.version` | 是否记录 Prompt 审计版本 | Prompt 变更可解释、可回归 |
|
||||
| `data.run.selfEvaluation.verifier_evaluation.prompt_audit.prompts[*].version` | 是否记录 planner / executor / verifier / composer 版本 | 便于定位 Prompt 变更影响 |
|
||||
| `data.run.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在当前 diagnosis run 上 |
|
||||
|
||||
## 2. Agent 步骤
|
||||
## 2. StateGraph 路由
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
| `data.run.orchestrationTrace.version` | 是否存在当前 trace contract 版本 | 路由摘要可演进、可兼容 |
|
||||
| `data.run.orchestrationTrace.transitions[*]` | 是否记录实际经过的 Node 和 route | Graph 条件边不是从日志推断 |
|
||||
| `data.run.orchestrationTrace.final_node` | 最终是 Composer 还是 Fallback | 正常输出与安全降级明确区分 |
|
||||
| `data.run.orchestrationTrace.termination_reason` | 是否给出终止原因 | 每次 Run 都有可解释终点 |
|
||||
| `data.run.orchestrationTrace.degraded` | 是否发生安全降级 | fallback 是可审计行为 |
|
||||
| `data.run.orchestrationTrace.evidence_retry_count` | 是否为 0 或 1 | 补证据循环有硬上限 |
|
||||
| `interview-demo-summary.json.finalNode` 等摘要字段 | 是否与 exact Trace 一致 | summary 只消费 Run 路由真理源 |
|
||||
|
||||
`orchestrationTrace` 负责路由;`selfEvaluation` 负责证据和答案质量;AgentStep/ToolInvocation 负责详细执行与工具证据。三者不能互相替代。
|
||||
|
||||
## 3. Agent 步骤
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
@@ -21,7 +39,7 @@
|
||||
| `data.steps[*].durationMs` | 是否有步骤耗时 | Trace 可用于耗时分析 |
|
||||
| `data.steps[*].tokenCount` | 如可用,是否记录 token | Trace 可用于模型成本分析 |
|
||||
|
||||
## 3. 工具证据
|
||||
## 4. 工具证据
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
@@ -30,9 +48,10 @@
|
||||
| `data.toolInvocations[*].outputPreview` | 是否有受控长度的证据预览 | 保留证据但不倾倒巨大 payload |
|
||||
| `data.toolInvocations[*].success` | 是否区分成功和失败 | 工具失败对 Verifier 和 reviewer 可见 |
|
||||
| `data.toolInvocations[*].retrievalDetails` | 是否包含检索 metadata | 检索质量可事后检查 |
|
||||
| `data.toolInvocations[*].retrievalDetails.evidence_refs` | 是否包含 `raw_path + text` | Gatekeeper 可以用代码核对 Executor 引用 |
|
||||
| `data.toolInvocations[*].relevanceLevel` | 是否有相关性等级 | 可解释检索结果强弱 |
|
||||
|
||||
## 4. Summary
|
||||
## 5. Summary
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
@@ -41,13 +60,14 @@
|
||||
| `data.summary.hasVerifierEvaluation` | 是否存在 Verifier 结果 | 最终答案经过质量门 |
|
||||
| `data.summary.hasFeedback` | 提交反馈后是否为 true | 人类反馈闭环完成 |
|
||||
|
||||
## 5. 好的结果长什么样
|
||||
## 6. 好的结果长什么样
|
||||
|
||||
```text
|
||||
同一个 session id
|
||||
同一个 session id + run id
|
||||
-> 最终答案
|
||||
-> 持久化 agent steps
|
||||
-> 持久化 evidence tool calls
|
||||
-> verifier / self-evaluation
|
||||
-> feedback attached to the same session
|
||||
-> run.orchestrationTrace routing summary
|
||||
-> feedback attached to the same run
|
||||
```
|
||||
|
||||
+56
-14
@@ -1,6 +1,18 @@
|
||||
# Diagnosis Eval Harness
|
||||
|
||||
This folder contains the first fixed-case evaluation set for the MVP diagnosis Agent.
|
||||
This folder contains the fixed offline regression set for the MVP diagnosis Agent.
|
||||
|
||||
## Background
|
||||
|
||||
The current complex Chat diagnosis chain is a bounded StateGraph:
|
||||
|
||||
```text
|
||||
Planner -> Executor -> Gatekeeper -> Verified Input -> Verifier -> Composer -> final answer
|
||||
| |
|
||||
+ bounded evidence retry + safe Fallback
|
||||
```
|
||||
|
||||
Executor V2 structured output, deterministic Gatekeeper audit, verified-only Verifier input, `claim_checks`, Composer rendering, and StateGraph routing are covered by deterministic tests so future prompt, tool, or graph changes can be checked without relying on a one-off demo.
|
||||
|
||||
## Scope
|
||||
|
||||
@@ -14,7 +26,28 @@ This folder contains the first fixed-case evaluation set for the MVP diagnosis A
|
||||
|
||||
## Current Mode
|
||||
|
||||
The first version evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.
|
||||
The baseline evaluates saved trace fixtures. It does not start the application and does not require MySQL, Redis, Milvus, or a real LLM.
|
||||
|
||||
The committed baseline currently contains:
|
||||
|
||||
```text
|
||||
12 fixed cases
|
||||
12 passing fixture evaluations
|
||||
5 PASS verdicts
|
||||
6 LOW_CONFID verdicts
|
||||
1 REJECT verdict
|
||||
```
|
||||
|
||||
The V2 evidence-pipeline matrix covers:
|
||||
|
||||
- Positive supported evidence for a narrow HighCPU observation.
|
||||
- No-evidence `negative_observation` using `$.no_evidence`.
|
||||
- Gatekeeper failure for a fabricated tool invocation reference.
|
||||
- Unsupported claim filtering before the final answer.
|
||||
- Composer fallback rendering without raw Executor JSON leakage.
|
||||
- Gatekeeper rule set version audit for new matrix fixtures.
|
||||
- Prompt audit version checks for planner, executor, verifier, and composer prompts.
|
||||
- Gatekeeper rule metadata checks for enabled rule id and default severity.
|
||||
|
||||
## Verification
|
||||
|
||||
@@ -24,29 +57,38 @@ Run the focused evaluator test:
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test
|
||||
```
|
||||
|
||||
The committed baseline report represents the current fixed fixture set:
|
||||
Run the authoritative Graph layers plus the fixed evaluator checks:
|
||||
|
||||
```text
|
||||
5 fixed cases
|
||||
5 passing fixture evaluations
|
||||
2 PASS verdicts
|
||||
3 LOW_CONFID verdicts
|
||||
```powershell
|
||||
mvn -q "-Dtest=DiagnosisGraphWorkflowTest,DiagnosisGraphNodeContractTest,ChatServiceGraphIntegrationTest,DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest" test
|
||||
```
|
||||
|
||||
When fixtures or evaluator rules change, regenerate the report from the same case file and fixture directory, then update both JSON and Markdown outputs together.
|
||||
When fixtures or evaluator rules change, regenerate both baseline reports from the same case file and fixture directory, then update JSON and Markdown together.
|
||||
|
||||
## Interview Story
|
||||
## Regression Signal
|
||||
|
||||
The harness gives the MVP a repeatable baseline:
|
||||
The harness is deterministic code, not an LLM judge:
|
||||
|
||||
```text
|
||||
fixed diagnosis case
|
||||
-> saved or runtime trace
|
||||
-> saved trace fixture
|
||||
-> rule-based trace validation
|
||||
-> JSON / Markdown report
|
||||
-> regression signal for prompts, tools, retrieval, and verifier behavior
|
||||
-> regression signal for prompts, tools, retrieval, verifier, and composer behavior
|
||||
```
|
||||
|
||||
Stage 5 adds these V2 checks:
|
||||
|
||||
- `gatekeeper_result.status=fail` cannot coexist with Verifier `PASS`.
|
||||
- Required V2 fixtures must include `gatekeeper_result`, `claim_checks`, and `composer_output`.
|
||||
- `claim_checks` must be structurally auditable.
|
||||
- Composer output must record whether normal parsing or fallback rendering was used.
|
||||
- Gatekeeper rule set version can be asserted per fixture.
|
||||
- Prompt audit version and per-prompt versions can be asserted per fixture.
|
||||
- Gatekeeper rule metadata can be required per fixture.
|
||||
- Final answers must not leak raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`.
|
||||
- Configured unsupported claim keywords must not appear as confirmed final-answer content.
|
||||
|
||||
## Baseline Diff
|
||||
|
||||
Baseline diff compares a current report against `reports/baseline-report.json`.
|
||||
@@ -58,4 +100,4 @@ current report
|
||||
-> regressions, improvements, and changed signals
|
||||
```
|
||||
|
||||
Use it to answer: did a prompt, tool, retrieval, or verifier change make the Agent worse than the fixed baseline?
|
||||
Use it to answer: did a prompt, tool, retrieval, verifier, or composer change make the Agent worse than the fixed baseline?
|
||||
|
||||
@@ -1,4 +1,67 @@
|
||||
[
|
||||
{
|
||||
"id": "narrow-highcpu-observation",
|
||||
"title": "Narrow HighCPU observation",
|
||||
"question": "确认 payment-service 当前是否存在 HighCPUUsage 告警。",
|
||||
"traceFixture": "narrow-highcpu-observation-pass.json",
|
||||
"expectedRootCauseKeywords": ["payment-service", "HighCPUUsage", "92%"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_metrics"],
|
||||
"allowedVerdicts": ["PASS"],
|
||||
"forbiddenAnswerKeywords": ["根因", "修复建议", "通常情况下"],
|
||||
"requireV2AuditClosure": true,
|
||||
"requireClaimChecks": true,
|
||||
"requireComposerOutput": true,
|
||||
"expectedGatekeeperStatuses": ["pass"],
|
||||
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
|
||||
"expectedComposerStatuses": ["valid"],
|
||||
"forbiddenConfirmedClaimKeywords": ["数据库连接池"]
|
||||
},
|
||||
{
|
||||
"id": "prompt-gatekeeper-audit-closure",
|
||||
"title": "Prompt and Gatekeeper audit closure",
|
||||
"question": "确认 payment-service 当前是否存在 HighCPUUsage 告警,并检查审计元数据是否完整。",
|
||||
"traceFixture": "prompt-gatekeeper-audit-closure-pass.json",
|
||||
"expectedRootCauseKeywords": ["payment-service", "HighCPUUsage", "92%"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_metrics"],
|
||||
"allowedVerdicts": ["PASS"],
|
||||
"forbiddenAnswerKeywords": ["根因", "修复建议", "通常情况下"],
|
||||
"requireV2AuditClosure": true,
|
||||
"requireClaimChecks": true,
|
||||
"requireComposerOutput": true,
|
||||
"expectedGatekeeperStatuses": ["pass"],
|
||||
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
|
||||
"expectedComposerStatuses": ["valid"],
|
||||
"forbiddenConfirmedClaimKeywords": ["数据库连接池"],
|
||||
"requirePromptAudit": true,
|
||||
"expectedPromptAuditVersion": "chat-prompts-v1",
|
||||
"expectedPromptVersions": {
|
||||
"chat_planner": "chat-planner-v1",
|
||||
"chat_executor": "chat-executor-v2",
|
||||
"chat_verifier": "chat-verifier-v2",
|
||||
"chat_composer": "chat-composer-v1"
|
||||
},
|
||||
"requireGatekeeperRules": true
|
||||
},
|
||||
{
|
||||
"id": "hikari-no-evidence-negative-observation",
|
||||
"title": "Hikari no-evidence negative observation",
|
||||
"question": "确认 inventory-service 当前是否有 HikariCP 连接池耗尽日志。",
|
||||
"traceFixture": "hikari-no-evidence-negative-observation-pass.json",
|
||||
"expectedRootCauseKeywords": ["未检索到", "HikariCP", "匹配证据"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_logs"],
|
||||
"allowedVerdicts": ["PASS"],
|
||||
"forbiddenAnswerKeywords": ["已排除", "确认没有", "日志层面已排除"],
|
||||
"requireV2AuditClosure": true,
|
||||
"requireClaimChecks": true,
|
||||
"requireComposerOutput": true,
|
||||
"expectedGatekeeperStatuses": ["pass"],
|
||||
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
|
||||
"expectedComposerStatuses": ["valid"],
|
||||
"forbiddenConfirmedClaimKeywords": ["已排除 HikariCP"]
|
||||
},
|
||||
{
|
||||
"id": "payment-timeout",
|
||||
"title": "Payment API timeout",
|
||||
@@ -53,5 +116,83 @@
|
||||
"requiredEvidenceTools": ["query_metrics", "query_logs"],
|
||||
"allowedVerdicts": ["LOW_CONFID", "PASS"],
|
||||
"forbiddenAnswerKeywords": ["可以忽略"]
|
||||
},
|
||||
{
|
||||
"id": "gatekeeper-fabricated-invocation",
|
||||
"title": "Gatekeeper fabricated invocation",
|
||||
"question": "支付失败是否能确认由日志中的连接池耗尽导致?",
|
||||
"traceFixture": "gatekeeper-fabricated-invocation-reject.json",
|
||||
"expectedRootCauseKeywords": ["证据", "引用", "失败"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_logs"],
|
||||
"allowedVerdicts": ["REJECT", "LOW_CONFID"],
|
||||
"forbiddenAnswerKeywords": ["已经完全确认"],
|
||||
"requireV2AuditClosure": true,
|
||||
"requireClaimChecks": true,
|
||||
"requireComposerOutput": true,
|
||||
"expectedGatekeeperStatuses": ["fail"],
|
||||
"expectedComposerStatuses": ["valid"],
|
||||
"forbiddenConfirmedClaimKeywords": ["连接池耗尽导致支付失败"]
|
||||
},
|
||||
{
|
||||
"id": "unsupported-claim-filtering",
|
||||
"title": "Unsupported claim filtering",
|
||||
"question": "订单超时是否可以确认由数据库主库故障导致?",
|
||||
"traceFixture": "unsupported-claim-filtering-low-confid.json",
|
||||
"expectedRootCauseKeywords": ["超时", "证据"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_logs"],
|
||||
"allowedVerdicts": ["LOW_CONFID"],
|
||||
"forbiddenAnswerKeywords": ["已经确认"],
|
||||
"requireV2AuditClosure": true,
|
||||
"requireClaimChecks": true,
|
||||
"requireComposerOutput": true,
|
||||
"expectedGatekeeperStatuses": ["pass"],
|
||||
"expectedComposerStatuses": ["valid"],
|
||||
"forbiddenConfirmedClaimKeywords": ["主库故障"]
|
||||
},
|
||||
{
|
||||
"id": "audit-metadata-low-confid",
|
||||
"title": "Audit metadata low confidence",
|
||||
"question": "订单超时是否可以确认由数据库主库故障导致,并检查审计元数据是否完整?",
|
||||
"traceFixture": "audit-metadata-low-confid.json",
|
||||
"expectedRootCauseKeywords": ["超时", "证据"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_logs"],
|
||||
"allowedVerdicts": ["LOW_CONFID"],
|
||||
"forbiddenAnswerKeywords": ["已经确认"],
|
||||
"requireV2AuditClosure": true,
|
||||
"requireClaimChecks": true,
|
||||
"requireComposerOutput": true,
|
||||
"expectedGatekeeperStatuses": ["pass"],
|
||||
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
|
||||
"expectedComposerStatuses": ["valid"],
|
||||
"forbiddenConfirmedClaimKeywords": ["主库故障"],
|
||||
"requirePromptAudit": true,
|
||||
"expectedPromptAuditVersion": "chat-prompts-v1",
|
||||
"expectedPromptVersions": {
|
||||
"chat_planner": "chat-planner-v1",
|
||||
"chat_executor": "chat-executor-v2",
|
||||
"chat_verifier": "chat-verifier-v2",
|
||||
"chat_composer": "chat-composer-v1"
|
||||
},
|
||||
"requireGatekeeperRules": true
|
||||
},
|
||||
{
|
||||
"id": "composer-fallback-no-raw-json",
|
||||
"title": "Composer fallback no raw JSON",
|
||||
"question": "库存服务慢响应是否可以直接输出 Executor JSON?",
|
||||
"traceFixture": "composer-fallback-no-raw-json-low-confid.json",
|
||||
"expectedRootCauseKeywords": ["慢响应", "证据"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_metrics"],
|
||||
"allowedVerdicts": ["LOW_CONFID"],
|
||||
"forbiddenAnswerKeywords": ["executor_evidence_v2", "answer_version", "claim_id"],
|
||||
"requireV2AuditClosure": true,
|
||||
"requireClaimChecks": true,
|
||||
"requireComposerOutput": true,
|
||||
"expectedGatekeeperStatuses": ["pass"],
|
||||
"expectedComposerStatuses": ["composer_malformed"],
|
||||
"forbiddenConfirmedClaimKeywords": ["线程池已经耗尽"]
|
||||
}
|
||||
]
|
||||
|
||||
@@ -0,0 +1,192 @@
|
||||
{
|
||||
"run": {
|
||||
"sessionId": "eval-audit-metadata-low-confid",
|
||||
"query": "订单超时是否可以确认由数据库主库故障导致,并检查审计元数据是否完整?",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 45000,
|
||||
"toolCallCount": 1,
|
||||
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:日志显示订单接口出现超时。\n\n仍需补充信息:目前没有数据库故障日志或主库状态证据,不能把该方向写成确认根因。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.42,
|
||||
"critical_fact_count": 2,
|
||||
"prompt_audit": {
|
||||
"version": "chat-prompts-v1",
|
||||
"prompts": [
|
||||
{
|
||||
"name": "chat_planner",
|
||||
"version": "chat-planner-v1",
|
||||
"resource": "prompts/chat-planner-prompt.md"
|
||||
},
|
||||
{
|
||||
"name": "chat_executor",
|
||||
"version": "chat-executor-v2",
|
||||
"resource": "prompts/chat-executor-prompt.md"
|
||||
},
|
||||
{
|
||||
"name": "chat_verifier",
|
||||
"version": "chat-verifier-v2",
|
||||
"resource": "prompts/chat-verifier-prompt.md"
|
||||
},
|
||||
{
|
||||
"name": "chat_composer",
|
||||
"version": "chat-composer-v1",
|
||||
"resource": "prompts/chat-composer-prompt.md"
|
||||
}
|
||||
]
|
||||
},
|
||||
"gatekeeper_result": {
|
||||
"status": "pass",
|
||||
"severity": "none",
|
||||
"rule_set_version": "gatekeeper-rules-v1",
|
||||
"rules": [
|
||||
{
|
||||
"id": "evidence.invocation",
|
||||
"description": "source_invocation_id must reference an existing tool invocation",
|
||||
"enabled": true,
|
||||
"default_severity": "reject"
|
||||
},
|
||||
{
|
||||
"id": "evidence.excerpt",
|
||||
"description": "evidence_excerpt must be supported by recorded evidence text",
|
||||
"enabled": true,
|
||||
"default_severity": "reject"
|
||||
}
|
||||
],
|
||||
"checked_bindings": [
|
||||
{
|
||||
"claim_id": "claim-timeout",
|
||||
"tool_name": "query_logs",
|
||||
"source_invocation_id": 22,
|
||||
"raw_path": "$.logs[0]",
|
||||
"matched_text": "order api timeout",
|
||||
"status": "pass"
|
||||
}
|
||||
],
|
||||
"failed_rules": [],
|
||||
"warnings": [],
|
||||
"errors": []
|
||||
},
|
||||
"executor_structured_output": {
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-timeout",
|
||||
"claim_type": "symptom",
|
||||
"claim_text": "订单接口出现超时",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"source_type": "tool_trace",
|
||||
"tool_name": "query_logs",
|
||||
"source_invocation_id": 22,
|
||||
"raw_path": "$.logs[0]",
|
||||
"evidence_excerpt": "order api timeout"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"claim_id": "claim-db-primary",
|
||||
"claim_type": "root_cause",
|
||||
"claim_text": "数据库主库故障导致订单超时",
|
||||
"support_level": "weak",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"source_type": "tool_trace",
|
||||
"tool_name": "query_logs",
|
||||
"source_invocation_id": 22,
|
||||
"raw_path": "$.logs[0]",
|
||||
"evidence_excerpt": "order api timeout"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "补充查询数据库主库状态和错误日志",
|
||||
"reason": "当前只有订单接口超时日志"
|
||||
}
|
||||
],
|
||||
"missing_info": ["数据库主库状态", "数据库错误日志"]
|
||||
},
|
||||
"claim_checks": [
|
||||
{
|
||||
"claim_id": "claim-timeout",
|
||||
"claim_text": "订单接口出现超时",
|
||||
"claim_type": "symptom",
|
||||
"verification": "direct_observation",
|
||||
"detail": "日志直接记录 order api timeout",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"source_invocation_id": 22,
|
||||
"raw_path": "$.logs[0]"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"claim_id": "claim-db-primary",
|
||||
"claim_text": "数据库主库故障导致订单超时",
|
||||
"claim_type": "root_cause",
|
||||
"verification": "unsupported",
|
||||
"detail": "日志只能证明订单接口超时,不能证明数据库主库故障",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"source_invocation_id": 22,
|
||||
"raw_path": "$.logs[0]"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"facts_checked": [],
|
||||
"composer_output": {
|
||||
"status": "valid",
|
||||
"answer_summary": "日志显示订单接口超时,但数据库方向证据不足。",
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "补充查询数据库主库状态和错误日志",
|
||||
"reason": "当前只有订单接口超时日志"
|
||||
}
|
||||
],
|
||||
"user_facing_answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:日志显示订单接口出现超时。\n\n仍需补充信息:目前没有数据库故障日志或主库状态证据,不能把该方向写成确认根因。"
|
||||
},
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 22,
|
||||
"sessionId": "eval-audit-metadata-low-confid",
|
||||
"toolName": "query_logs",
|
||||
"outputPreview": "order api timeout",
|
||||
"retrievalDetails": {
|
||||
"evidence_status": "supported",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.logs[0]",
|
||||
"text": "order api timeout"
|
||||
}
|
||||
]
|
||||
},
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 1,
|
||||
"returnedToolCallCount": 1,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,142 @@
|
||||
{
|
||||
"run": {
|
||||
"sessionId": "eval-composer-fallback-no-raw-json",
|
||||
"query": "库存服务慢响应是否可以直接输出 Executor JSON?",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 47000,
|
||||
"toolCallCount": 1,
|
||||
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:指标显示库存服务出现慢响应。\n\n仍需补充信息:当前没有线程池队列或线程耗尽证据,不能确认线程池方向。\n\n建议动作:补充查询库存服务线程池指标。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.45,
|
||||
"critical_fact_count": 2,
|
||||
"gatekeeper_result": {
|
||||
"status": "pass",
|
||||
"failed_rules": [],
|
||||
"warnings": [],
|
||||
"errors": []
|
||||
},
|
||||
"executor_structured_output": {
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-slow-response",
|
||||
"claim_type": "symptom",
|
||||
"claim_text": "库存服务出现慢响应",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"tool_invocation_id": 1,
|
||||
"tool_name": "query_metrics",
|
||||
"evidence_excerpt": "inventory p99 latency increased"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"claim_id": "claim-thread-pool",
|
||||
"claim_type": "root_cause",
|
||||
"claim_text": "线程池已经耗尽",
|
||||
"support_level": "weak",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"tool_invocation_id": 1,
|
||||
"tool_name": "query_metrics",
|
||||
"evidence_excerpt": "inventory p99 latency increased"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "补充查询库存服务线程池指标",
|
||||
"reason": "当前只有慢响应指标"
|
||||
}
|
||||
],
|
||||
"missing_info": ["线程池队列长度", "活跃线程数"]
|
||||
},
|
||||
"claim_checks": [
|
||||
{
|
||||
"claim_id": "claim-slow-response",
|
||||
"claim_text": "库存服务出现慢响应",
|
||||
"claim_type": "symptom",
|
||||
"verification": "direct_observation",
|
||||
"detail": "指标显示 inventory p99 latency increased",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"tool_invocation_id": 1,
|
||||
"tool_name": "query_metrics"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"claim_id": "claim-thread-pool",
|
||||
"claim_text": "线程池已经耗尽",
|
||||
"claim_type": "root_cause",
|
||||
"verification": "external_unknown",
|
||||
"detail": "没有线程池队列或活跃线程指标,不能确认该结论",
|
||||
"evidence_refs": []
|
||||
}
|
||||
],
|
||||
"facts_checked": [
|
||||
{
|
||||
"fact": "库存服务出现慢响应",
|
||||
"verification": "direct_evidence",
|
||||
"detail": "指标显示 inventory p99 latency increased",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"tool_invocation_id": 1,
|
||||
"tool_name": "query_metrics"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"fact": "线程池已经耗尽",
|
||||
"verification": "external_unknown",
|
||||
"detail": "没有线程池队列或活跃线程指标,不能确认该结论",
|
||||
"evidence_refs": []
|
||||
}
|
||||
],
|
||||
"composer_output": {
|
||||
"status": "composer_malformed",
|
||||
"detail": "used safe fallback rendering",
|
||||
"answer_summary": "指标显示库存服务出现慢响应。",
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "补充查询库存服务线程池指标",
|
||||
"reason": "当前只有慢响应指标"
|
||||
}
|
||||
],
|
||||
"user_facing_answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:指标显示库存服务出现慢响应。\n\n仍需补充信息:当前没有线程池队列或线程耗尽证据,不能确认线程池方向。\n\n建议动作:补充查询库存服务线程池指标。"
|
||||
},
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_metrics",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-composer-fallback-no-raw-json",
|
||||
"toolName": "query_metrics",
|
||||
"outputPreview": "inventory p99 latency increased",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 1,
|
||||
"returnedToolCallCount": 1,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,110 @@
|
||||
{
|
||||
"run": {
|
||||
"sessionId": "eval-gatekeeper-fabricated-invocation",
|
||||
"query": "支付失败是否能确认由日志中的连接池耗尽导致?",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 39000,
|
||||
"toolCallCount": 1,
|
||||
"answer": "当前无法基于已获取证据生成可靠结论。\n\n证据引用校验失败:Executor 引用了不存在的工具调用记录,因此不能把连接池问题作为确认结论。建议重新收集日志证据后再判断。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "REJECT",
|
||||
"groundedness_score": 0.1,
|
||||
"critical_fact_count": 1,
|
||||
"gatekeeper_result": {
|
||||
"status": "fail",
|
||||
"failed_rules": ["evidence.invocation_ref"],
|
||||
"warnings": [],
|
||||
"errors": [
|
||||
{
|
||||
"rule_id": "evidence.invocation_ref",
|
||||
"field": "claims[0].evidence_bindings[0].tool_invocation_id",
|
||||
"message": "tool_invocation_id does not exist"
|
||||
}
|
||||
]
|
||||
},
|
||||
"executor_structured_output": {
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "root_cause",
|
||||
"claim_text": "连接池耗尽导致支付失败",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"tool_invocation_id": 99,
|
||||
"tool_name": "query_logs",
|
||||
"evidence_excerpt": "connection pool exhausted"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "重新查询支付服务错误日志",
|
||||
"reason": "当前 Executor 证据引用无法回溯"
|
||||
}
|
||||
],
|
||||
"missing_info": ["需要有效的日志工具调用记录"]
|
||||
},
|
||||
"claim_checks": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_text": "连接池耗尽导致支付失败",
|
||||
"claim_type": "root_cause",
|
||||
"verification": "unsupported",
|
||||
"detail": "Gatekeeper 已判定证据引用不存在,不能确认该结论",
|
||||
"evidence_refs": []
|
||||
}
|
||||
],
|
||||
"facts_checked": [
|
||||
{
|
||||
"fact": "连接池耗尽导致支付失败",
|
||||
"verification": "unsupported",
|
||||
"detail": "Gatekeeper 已判定证据引用不存在,不能确认该结论",
|
||||
"evidence_refs": []
|
||||
}
|
||||
],
|
||||
"composer_output": {
|
||||
"status": "valid",
|
||||
"answer_summary": "证据引用校验失败,不能确认根因。",
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "重新查询支付服务错误日志",
|
||||
"reason": "当前 Executor 证据引用无法回溯"
|
||||
}
|
||||
],
|
||||
"user_facing_answer": "当前无法基于已获取证据生成可靠结论。\n\n证据引用校验失败:Executor 引用了不存在的工具调用记录,因此不能把连接池问题作为确认结论。建议重新收集日志证据后再判断。"
|
||||
},
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 1,
|
||||
"sessionId": "eval-gatekeeper-fabricated-invocation",
|
||||
"toolName": "query_logs",
|
||||
"outputPreview": "payment failed without matching connection pool exhaustion entry",
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 1,
|
||||
"returnedToolCallCount": 1,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,126 @@
|
||||
{
|
||||
"run": {
|
||||
"sessionId": "eval-hikari-no-evidence-negative-observation",
|
||||
"query": "确认 inventory-service 当前是否有 HikariCP 连接池耗尽日志。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 21000,
|
||||
"toolCallCount": 1,
|
||||
"answer": "本次查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志;这只表示当前查询没有匹配证据,仍不能据此判断系统一定健康。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 1.0,
|
||||
"critical_fact_count": 1,
|
||||
"gatekeeper_result": {
|
||||
"status": "pass",
|
||||
"severity": "none",
|
||||
"rule_set_version": "gatekeeper-rules-v1",
|
||||
"rules": [
|
||||
{
|
||||
"id": "evidence.raw_path",
|
||||
"description": "raw_path must exist in retrieval_details.evidence_refs",
|
||||
"enabled": true,
|
||||
"default_severity": "reject"
|
||||
}
|
||||
],
|
||||
"checked_bindings": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"tool_name": "query_logs",
|
||||
"source_invocation_id": 12,
|
||||
"raw_path": "$.no_evidence",
|
||||
"matched_text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志",
|
||||
"status": "pass"
|
||||
}
|
||||
],
|
||||
"failed_rules": [],
|
||||
"warnings": [],
|
||||
"errors": []
|
||||
},
|
||||
"executor_structured_output": {
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "negative_observation",
|
||||
"claim_text": "当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"source_type": "tool_trace",
|
||||
"tool_name": "query_logs",
|
||||
"source_invocation_id": 12,
|
||||
"raw_path": "$.no_evidence",
|
||||
"evidence_excerpt": "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [],
|
||||
"missing_info": [
|
||||
"仅查询了 application-logs 中 inventory-service HikariCP 相关日志"
|
||||
]
|
||||
},
|
||||
"claim_checks": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_text": "当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。",
|
||||
"claim_type": "negative_observation",
|
||||
"verification": "direct_observation",
|
||||
"detail": "$.no_evidence 只支持本次查询未检索到匹配证据。",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"source_invocation_id": 12,
|
||||
"raw_path": "$.no_evidence"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"facts_checked": [],
|
||||
"composer_output": {
|
||||
"status": "valid",
|
||||
"answer_summary": "本次查询未检索到匹配日志。",
|
||||
"recommended_actions": [],
|
||||
"user_facing_answer": "本次查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志;这只表示当前查询没有匹配证据,仍不能据此判断系统一定健康。"
|
||||
},
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"source_invocation_ids": [12],
|
||||
"evidence_level": "no_evidence"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 12,
|
||||
"sessionId": "eval-hikari-no-evidence-negative-observation",
|
||||
"toolName": "query_logs",
|
||||
"outputPreview": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志",
|
||||
"retrievalDetails": {
|
||||
"evidence_status": "no_evidence",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.no_evidence",
|
||||
"text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志"
|
||||
}
|
||||
]
|
||||
},
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 1,
|
||||
"returnedToolCallCount": 1,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -1,5 +1,5 @@
|
||||
{
|
||||
"session": {
|
||||
"run": {
|
||||
"sessionId": "eval-jvm-memory-risk",
|
||||
"query": "订单服务内存使用率过高,请判断是否存在 OOM 风险。",
|
||||
"status": "SUCCESS",
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
{
|
||||
"session": {
|
||||
"run": {
|
||||
"sessionId": "eval-mysql-pool",
|
||||
"query": "订单服务大量请求超时,请判断是否和 MySQL 连接池有关。",
|
||||
"status": "SUCCESS",
|
||||
|
||||
@@ -0,0 +1,124 @@
|
||||
{
|
||||
"run": {
|
||||
"sessionId": "eval-narrow-highcpu-observation",
|
||||
"query": "确认 payment-service 当前是否存在 HighCPUUsage 告警。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 18000,
|
||||
"toolCallCount": 1,
|
||||
"answer": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 1.0,
|
||||
"critical_fact_count": 1,
|
||||
"gatekeeper_result": {
|
||||
"status": "pass",
|
||||
"severity": "none",
|
||||
"rule_set_version": "gatekeeper-rules-v1",
|
||||
"rules": [
|
||||
{
|
||||
"id": "evidence.raw_path",
|
||||
"description": "raw_path must exist in retrieval_details.evidence_refs",
|
||||
"enabled": true,
|
||||
"default_severity": "reject"
|
||||
}
|
||||
],
|
||||
"checked_bindings": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"tool_name": "query_metrics",
|
||||
"source_invocation_id": 11,
|
||||
"raw_path": "$.alerts[0]",
|
||||
"matched_text": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m",
|
||||
"status": "pass"
|
||||
}
|
||||
],
|
||||
"failed_rules": [],
|
||||
"warnings": [],
|
||||
"errors": []
|
||||
},
|
||||
"executor_structured_output": {
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "observation",
|
||||
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"source_type": "tool_trace",
|
||||
"tool_name": "query_metrics",
|
||||
"source_invocation_id": 11,
|
||||
"raw_path": "$.alerts[0]",
|
||||
"evidence_excerpt": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [],
|
||||
"missing_info": []
|
||||
},
|
||||
"claim_checks": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
|
||||
"claim_type": "observation",
|
||||
"verification": "direct_observation",
|
||||
"detail": "已核验的指标证据直接包含服务名、告警名和 CPU 当前值。",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"source_invocation_id": 11,
|
||||
"raw_path": "$.alerts[0]"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"facts_checked": [],
|
||||
"composer_output": {
|
||||
"status": "valid",
|
||||
"answer_summary": "payment-service 当前存在 HighCPUUsage 告警。",
|
||||
"recommended_actions": [],
|
||||
"user_facing_answer": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。"
|
||||
},
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_metrics",
|
||||
"success": true,
|
||||
"source_invocation_ids": [11],
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 11,
|
||||
"sessionId": "eval-narrow-highcpu-observation",
|
||||
"toolName": "query_metrics",
|
||||
"outputPreview": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m",
|
||||
"retrievalDetails": {
|
||||
"evidence_status": "supported",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.alerts[0]",
|
||||
"text": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m"
|
||||
}
|
||||
]
|
||||
},
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 1,
|
||||
"returnedToolCallCount": 1,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -1,5 +1,5 @@
|
||||
{
|
||||
"session": {
|
||||
"run": {
|
||||
"sessionId": "eval-payment-timeout",
|
||||
"query": "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。",
|
||||
"status": "SUCCESS",
|
||||
|
||||
@@ -0,0 +1,155 @@
|
||||
{
|
||||
"run": {
|
||||
"sessionId": "eval-prompt-gatekeeper-audit-closure",
|
||||
"query": "确认 payment-service 当前是否存在 HighCPUUsage 告警,并检查审计元数据是否完整。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 19000,
|
||||
"toolCallCount": 1,
|
||||
"answer": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 1.0,
|
||||
"critical_fact_count": 1,
|
||||
"prompt_audit": {
|
||||
"version": "chat-prompts-v1",
|
||||
"prompts": [
|
||||
{
|
||||
"name": "chat_planner",
|
||||
"version": "chat-planner-v1",
|
||||
"resource": "prompts/chat-planner-prompt.md"
|
||||
},
|
||||
{
|
||||
"name": "chat_executor",
|
||||
"version": "chat-executor-v2",
|
||||
"resource": "prompts/chat-executor-prompt.md"
|
||||
},
|
||||
{
|
||||
"name": "chat_verifier",
|
||||
"version": "chat-verifier-v2",
|
||||
"resource": "prompts/chat-verifier-prompt.md"
|
||||
},
|
||||
{
|
||||
"name": "chat_composer",
|
||||
"version": "chat-composer-v1",
|
||||
"resource": "prompts/chat-composer-prompt.md"
|
||||
}
|
||||
]
|
||||
},
|
||||
"gatekeeper_result": {
|
||||
"status": "pass",
|
||||
"severity": "none",
|
||||
"rule_set_version": "gatekeeper-rules-v1",
|
||||
"rules": [
|
||||
{
|
||||
"id": "evidence.invocation",
|
||||
"description": "source_invocation_id must reference an existing tool invocation",
|
||||
"enabled": true,
|
||||
"default_severity": "reject"
|
||||
},
|
||||
{
|
||||
"id": "evidence.raw_path",
|
||||
"description": "raw_path must exist in retrieval_details.evidence_refs",
|
||||
"enabled": true,
|
||||
"default_severity": "reject"
|
||||
}
|
||||
],
|
||||
"checked_bindings": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"tool_name": "query_metrics",
|
||||
"source_invocation_id": 21,
|
||||
"raw_path": "$.alerts[0]",
|
||||
"matched_text": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m",
|
||||
"status": "pass"
|
||||
}
|
||||
],
|
||||
"failed_rules": [],
|
||||
"warnings": [],
|
||||
"errors": []
|
||||
},
|
||||
"executor_structured_output": {
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "observation",
|
||||
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"source_type": "tool_trace",
|
||||
"tool_name": "query_metrics",
|
||||
"source_invocation_id": 21,
|
||||
"raw_path": "$.alerts[0]",
|
||||
"evidence_excerpt": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [],
|
||||
"missing_info": []
|
||||
},
|
||||
"claim_checks": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
|
||||
"claim_type": "observation",
|
||||
"verification": "direct_observation",
|
||||
"detail": "已核验的指标证据直接包含服务名、告警名和 CPU 当前值。",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"source_invocation_id": 21,
|
||||
"raw_path": "$.alerts[0]"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"facts_checked": [],
|
||||
"composer_output": {
|
||||
"status": "valid",
|
||||
"answer_summary": "payment-service 当前存在 HighCPUUsage 告警。",
|
||||
"recommended_actions": [],
|
||||
"user_facing_answer": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。"
|
||||
},
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_metrics",
|
||||
"success": true,
|
||||
"source_invocation_ids": [21],
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 21,
|
||||
"sessionId": "eval-prompt-gatekeeper-audit-closure",
|
||||
"toolName": "query_metrics",
|
||||
"outputPreview": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m",
|
||||
"retrievalDetails": {
|
||||
"evidence_status": "supported",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.alerts[0]",
|
||||
"text": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m"
|
||||
}
|
||||
]
|
||||
},
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 1,
|
||||
"returnedToolCallCount": 1,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user