Compare commits
13
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
30d3296043 | ||
|
|
3578709896 | ||
|
|
f9df94377b | ||
|
|
78c1477198 | ||
|
|
d928a1968a | ||
|
|
027aed1eeb | ||
|
|
26d5529280 | ||
|
|
6fdbd34bab | ||
|
|
52bf0302c6 | ||
|
|
841437fa06 | ||
|
|
9c9a0024d4 | ||
|
|
a6c2d4459c | ||
|
|
da45fa3fb0 |
@@ -72,9 +72,10 @@
|
|||||||
|
|
||||||
### SessionContext
|
### SessionContext
|
||||||
- 定义:会话上下文数据类,存储在 Redis 中的会话数据
|
- 定义:会话上下文数据类,存储在 Redis 中的会话数据
|
||||||
- 包含字段:sessionId、userId、businessId、traceId、status、toolCalls、TTL
|
- 包含字段:sessionId、userId、businessId、traceId、status、toolCalls、messageHistory、TTL
|
||||||
- 序列化方式:JSON(GenericJackson2JsonRedisSerializer)
|
- 序列化方式:JSON(GenericJackson2JsonRedisSerializer)
|
||||||
- 使用场景:多轮对话上下文管理、工具调用历史追踪
|
- 使用场景:多轮对话上下文管理、工具调用历史追踪
|
||||||
|
- 边界:messageHistory 是热路径对话历史缓存,用于下一轮 prompt 上下文;长期审计的问题和答案应落到 Diagnosis Run,而不是依赖 Redis TTL 内的上下文正文。
|
||||||
|
|
||||||
### ToolCall
|
### ToolCall
|
||||||
- 定义:工具调用记录数据类,追踪 Agent 使用的工具及其结果
|
- 定义:工具调用记录数据类,追踪 Agent 使用的工具及其结果
|
||||||
@@ -87,6 +88,21 @@
|
|||||||
- 核心方法:createSession、getSession、updateSession、deleteSession、refreshSession、addToolCall
|
- 核心方法:createSession、getSession、updateSession、deleteSession、refreshSession、addToolCall
|
||||||
- 使用场景:分布式会话管理、Agent 状态维护
|
- 使用场景:分布式会话管理、Agent 状态维护
|
||||||
|
|
||||||
|
### Chat Session
|
||||||
|
- 定义:一次多轮对话上下文,由 `sessionId` 唯一标识。
|
||||||
|
- 使用场景:保存用户连续对话的上下文窗口、会话状态和最近活跃时间。
|
||||||
|
- 边界:Chat Session 不代表一次诊断执行;同一个 Chat Session 可以包含多次 Diagnosis Run。
|
||||||
|
|
||||||
|
### Diagnosis Run
|
||||||
|
- 定义:一次独立诊断执行,由 `runId` 唯一标识,属于一个 Chat Session。
|
||||||
|
- 使用场景:保存某一轮诊断的 query、answer、status、耗时、token、反馈和自评估结果。
|
||||||
|
- 边界:Diagnosis Run 是 Trace、Feedback 和 Evidence score 的绑定对象;多轮对话中的每次 `/api/chat` 或 `/api/ai_ops` 执行都应创建新的 Diagnosis Run。
|
||||||
|
|
||||||
|
### Diagnosis Trace
|
||||||
|
- 定义:一次 Diagnosis Run 的可回放执行轨迹,由 run 主记录、AgentStep 和 ToolInvocation 聚合形成。
|
||||||
|
- 使用场景:Trace API、Trace UI、Verifier 审计、评测 fixture 和人工排查。
|
||||||
|
- 边界:Diagnosis Trace 是聚合视图,不要求单独的 trace 主表;当前 trace 明细由 `agent_step` 和 `tool_invocation` 表承载。
|
||||||
|
|
||||||
### Flyway
|
### Flyway
|
||||||
- 定义:数据库版本迁移工具,管理 SQL 脚本的版本化执行
|
- 定义:数据库版本迁移工具,管理 SQL 脚本的版本化执行
|
||||||
- 配置:spring.flyway.enabled=true, baseline-on-migrate=true
|
- 配置:spring.flyway.enabled=true, baseline-on-migrate=true
|
||||||
|
|||||||
+31
-29
@@ -2,32 +2,34 @@
|
|||||||
|
|
||||||
## 项目
|
## 项目
|
||||||
|
|
||||||
| 日期 | slug | 领域 | 关键词 | 状态 |
|
| 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|
||||||
|---|---|---|---|---|
|
|---|---|---|---|---|---|---|
|
||||||
| 2026-07-05 | diagnosis-playbook-skills | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented |
|
| 2026-07-10 | session-run-trace-isolation | 拆分会话态和运行态,引入 runId 隔离 Trace、Feedback、AIOps 和 demo 链路。 | Trace/session/run isolation | chat_session, diagnosis_run, runId, trace exact run, feedback fallback, AIOps SSE metadata, baseline drift | openspec/changes/archive/2026-07-10-session-run-trace-isolation | archived |
|
||||||
| 2026-07-07 | executor-evidence-output-contract | Chat质量门禁/证据归因 | Executor structured output, evidence bindings, Verifier structured claims, LOW_CONFID, hallucination | openspec/changes/archive/2026-07-07-executor-evidence-output-contract | archived |
|
| 2026-07-09 | interview-demo-quality-audit | 增加面试演示前置质量审计,覆盖 prompt、Gatekeeper 和评测基线。 | Agent eval/demo/Prompt audit | interview demo preflight, prompt_audit, gatekeeper rules, diagnosis baseline, 12 fixtures | openspec/changes/archive/2026-07-09-interview-demo-quality-audit | archived |
|
||||||
| 2026-07-07 | executor-v2-output-contract | Chat质量门禁/证据归因 | executor_evidence_v2, user_facing_answer removal, diagnosis_summary removal, structured renderer | openspec/changes/archive/2026-07-07-executor-v2-output-contract | archived |
|
| 2026-07-08 | executor-composer-final-answer | 引入 Composer 生成最终回答,只使用 Verifier 允许的结论材料。 | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
|
||||||
| 2026-07-07 | executor-gatekeeper-hook | Chat质量门禁/证据归因 | Gatekeeper, verifier payload, source_invocation_ids, tool_name match, self_evaluation | openspec/changes/archive/2026-07-07-executor-gatekeeper-hook | archived |
|
| 2026-07-08 | diagnosis-eval-demo-gatekeeper-closure | 收敛诊断评测、稳定 demo 场景和 Gatekeeper 审计元数据。 | Agent eval/demo/Gatekeeper | diagnosis eval matrix, stable demo scenarios, Gatekeeper rule set version, audit metadata | openspec/changes/archive/2026-07-08-diagnosis-eval-demo-gatekeeper-closure | archived |
|
||||||
| 2026-07-07 | executor-verifier-claim-checks | Chat质量门禁/证据归因 | Verifier claim_checks, facts_checked compatibility, effective verdict guardrail, malformed output downgrade | openspec/changes/archive/2026-07-07-executor-verifier-claim-checks | archived |
|
| 2026-07-08 | verifier-evidence-reference-fidelity | 强化 Verifier 对 evidence_refs、raw_path 和 no_evidence 的保真校验。 | Chat质量门禁/证据归因 | evidence_refs, raw_path, Gatekeeper severity, verifier evidence excerpt, HikariCP mock, no_evidence | openspec/changes/archive/2026-07-08-verifier-evidence-reference-fidelity | archived |
|
||||||
| 2026-07-08 | executor-composer-final-answer | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
|
| 2026-07-07 | executor-evidence-output-contract | 设计 Executor 结构化证据输出,解决证据归因幻觉和 LOW_CONFID 问题。 | Chat质量门禁/证据归因 | Executor structured output, evidence bindings, Verifier structured claims, LOW_CONFID, hallucination | openspec/changes/archive/2026-07-07-executor-evidence-output-contract | archived |
|
||||||
| 2026-07-08 | diagnosis-eval-demo-gatekeeper-closure | Agent eval/demo/Gatekeeper | diagnosis eval matrix, stable demo scenarios, Gatekeeper rule set version, audit metadata | openspec/changes/archive/2026-07-08-diagnosis-eval-demo-gatekeeper-closure | archived |
|
| 2026-07-07 | executor-v2-output-contract | 将 Executor 输出升级为 V2 契约,移除面向用户的最终回答字段。 | Chat质量门禁/证据归因 | executor_evidence_v2, user_facing_answer removal, diagnosis_summary removal, structured renderer | openspec/changes/archive/2026-07-07-executor-v2-output-contract | archived |
|
||||||
| 2026-07-08 | verifier-evidence-reference-fidelity | Chat质量门禁/证据归因 | evidence_refs, raw_path, Gatekeeper severity, verifier evidence excerpt, HikariCP mock, no_evidence | openspec/changes/archive/2026-07-08-verifier-evidence-reference-fidelity | archived |
|
| 2026-07-07 | executor-gatekeeper-hook | 在 Executor 与 Verifier 之间接入 Gatekeeper,校验证据绑定来源。 | Chat质量门禁/证据归因 | Gatekeeper, verifier payload, source_invocation_ids, tool_name match, self_evaluation | openspec/changes/archive/2026-07-07-executor-gatekeeper-hook | archived |
|
||||||
| 2026-07-06 | rag-eval-pipeline-closure | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
|
| 2026-07-07 | executor-verifier-claim-checks | 增加 Verifier claim_checks 和事实校验兼容逻辑。 | Chat质量门禁/证据归因 | Verifier claim_checks, facts_checked compatibility, effective verdict guardrail, malformed output downgrade | openspec/changes/archive/2026-07-07-executor-verifier-claim-checks | archived |
|
||||||
| 2026-07-06 | modular-rag-pipeline | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
|
| 2026-07-06 | rag-eval-pipeline-closure | 建立 RAG 评测闭环,加入 fixture、快照和 baseline diff。 | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
|
||||||
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
|
| 2026-07-06 | modular-rag-pipeline | 将 lookup_knowledge 改造成模块化 RAG 管线,补齐证据块和检索追踪。 | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
|
||||||
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
|
| 2026-07-05 | diagnosis-playbook-skills | 增加诊断 Playbook Skill,沉淀支付超时、MySQL 池、Redis 超时等套路。 | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented |
|
||||||
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
|
| 2026-07-05 | mvp-demo-interview-runbook | 准备可复现的 MVP 面试演示包、运行手册和 Trace 检查清单。 | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
|
||||||
| 2026-07-04 | diagnosis-eval-harness | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
|
| 2026-07-05 | diagnosis-eval-baseline-diff | 增加诊断评测 baseline diff,用于判断回归和证据覆盖变化。 | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
|
||||||
| 2026-07-04 | evidence-trace-hardening | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
|
| 2026-07-04 | expand-diagnosis-eval-fixtures | 扩充诊断评测 fixture,覆盖 Redis、慢响应和 JVM 内存风险。 | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
|
||||||
| 2026-07-03 | mvp-demo-trace-acceptance | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
|
| 2026-07-04 | diagnosis-eval-harness | 建立固定诊断评测 Harness,输出 trace、证据覆盖和 verdict 分布。 | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
|
||||||
| 2026-07-04 | aiops-traceable-diagnosis-entry | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
|
| 2026-07-04 | evidence-trace-hardening | 强化工具调用证据链、降级契约和离线验证能力。 | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
|
||||||
| 2026-07-04 | aiops-alert-scope-control | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
|
| 2026-07-04 | aiops-traceable-diagnosis-entry | 增加可追踪的 AIOps 告警诊断入口,打通 sessionId 和 Trace API。 | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
|
||||||
| 2026-05-29 | chatmodel-abstraction | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | archived |
|
| 2026-07-04 | aiops-alert-scope-control | 收敛 AIOps 告警诊断范围,区分 payload 定向和自动发现模式。 | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
|
||||||
| 2026-06-23 | phase1-infrastructure | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | archived |
|
| 2026-07-03 | mvp-demo-trace-acceptance | 增加 MVP demo 的 Trace 验收,覆盖会话、步骤、工具和反馈链路。 | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
|
||||||
| 2026-06-24 | lookup-knowledge-integration | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | archived |
|
| 2026-07-02 | chat-verifier-agent | 增加 Chat Verifier Agent,用 groundedness 和 evidence_refs 校验回答。 | Chat质量门禁/可追溯验证 | Verifier, groundedness_score, facts_checked, evidence_refs, tool_trace_summary, self_evaluation | openspec/changes/archive/2026-07-03-chat-verifier-agent | archived |
|
||||||
| 2026-06-25 | doc-management-ui | 前端开发/文档管理 | 文档管理页面, CRUD, 状态监控, 纯静态页面, API集成 | archived |
|
| 2026-07-01 | executor-action-memory-relevance | 增加行动记忆和相关性信号,约束 Executor 重复检索。 | 检索质量/行动记忆 | relevanceLevel, completenessHint, Min-Max归一化, RetrievedDocTracker域级记录, Executor检索约束, ISS-002 | openspec/changes/archive/2026-07-01-executor-action-memory-relevance | archived |
|
||||||
| 2026-06-26 | session-storage | 会话存储/可观测 | diagnosis_session, agent_step, tool_invocation, token追踪, 多Agent路由 | openspec/changes/session-storage | archived |
|
| 2026-06-30 | session-dedup-knowledge-map | 引入会话级去重和知识域地图,减少重复召回。 | 去重/知识图谱 | RetrievedDocTracker, KnowledgeDomainService, knowledge_domain, covers, whenToRetrieve, Planner注入, ISS-001 | openspec/changes/archive/2026-06-30-session-dedup-knowledge-map | archived |
|
||||||
| 2026-06-29 | confidence-feedback | 质量评估/反馈机制 | evidence_score, selfEvaluation, feedback, useful, not_useful, case_library, BAD_CASE, tool_invocation规则引擎, 反馈按钮, sessionId回传 | openspec/changes/confidence-feedback | archived |
|
| 2026-06-29 | confidence-feedback | 建立质量评估和用户反馈机制,并把有用反馈沉淀为案例。 | 质量评估/反馈机制 | evidence_score, selfEvaluation, feedback, useful, not_useful, case_library, BAD_CASE, tool_invocation规则引擎, 反馈按钮, sessionId回传 | openspec/changes/confidence-feedback | archived |
|
||||||
| 2026-06-30 | session-dedup-knowledge-map | 去重/知识图谱 | RetrievedDocTracker, KnowledgeDomainService, knowledge_domain, covers, whenToRetrieve, Planner注入, ISS-001 | openspec/changes/archive/2026-06-30-session-dedup-knowledge-map | archived |
|
| 2026-06-26 | session-storage | 建立通用会话存储,记录 session、agent step 和 tool invocation。 | 会话存储/可观测 | diagnosis_session, agent_step, tool_invocation, token追踪, 多Agent路由 | openspec/changes/session-storage | archived |
|
||||||
| 2026-07-01 | executor-action-memory-relevance | 检索质量/行动记忆 | relevanceLevel, completenessHint, Min-Max归一化, RetrievedDocTracker域级记录, Executor检索约束, ISS-002 | openspec/changes/archive/2026-07-01-executor-action-memory-relevance | archived |
|
| 2026-06-25 | doc-management-ui | 实现文档管理页面,支持文档 CRUD、状态监控和 API 集成。 | 前端开发/文档管理 | 文档管理页面, CRUD, 状态监控, 纯静态页面, API集成 | - | archived |
|
||||||
| 2026-07-02 | chat-verifier-agent | Chat质量门禁/可追溯验证 | Verifier, groundedness_score, facts_checked, evidence_refs, tool_trace_summary, self_evaluation | openspec/changes/archive/2026-07-03-chat-verifier-agent | archived |
|
| 2026-06-24 | lookup-knowledge-integration | 接入知识库检索,支持 L0 精确匹配和 L1 语义检索。 | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | - | archived |
|
||||||
|
| 2026-06-23 | phase1-infrastructure | 搭建第一阶段基础设施,包括 MySQL、Redis、Milvus、Flyway 和 JPA。 | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | - | archived |
|
||||||
|
| 2026-05-29 | chatmodel-abstraction | 抽象 ChatModel 和 EmbeddingModel,支持多模型路由。 | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | - | archived |
|
||||||
|
|||||||
@@ -48,7 +48,7 @@
|
|||||||
|
|
||||||
## 遗留问题
|
## 遗留问题
|
||||||
|
|
||||||
ISS-002:Executor 无约束重复调用 `lookup_knowledge`(单会话 20+ 次),knowledge map 和检索约束只注入了 Planner 未注入 Executor。详见 `mvp/issues/ISS-002-executor-unconstrained-lookup.md`。
|
ISS-002:Executor 无约束重复调用 `lookup_knowledge`(单会话 20+ 次),knowledge map 和检索约束只注入了 Planner 未注入 Executor。详见 `mvp/issues/archived/ISS-002-executor-unconstrained-lookup.md`。
|
||||||
|
|
||||||
## 已知限制
|
## 已知限制
|
||||||
|
|
||||||
|
|||||||
@@ -9,8 +9,8 @@
|
|||||||
## Context
|
## Context
|
||||||
|
|
||||||
- `devflow/index.md` was checked. Relevant history includes `session-storage`, `confidence-feedback`, `executor-action-memory-relevance`, and `chat-verifier-agent`.
|
- `devflow/index.md` was checked. Relevant history includes `session-storage`, `confidence-feedback`, `executor-action-memory-relevance`, and `chat-verifier-agent`.
|
||||||
- `mvp/notes/agent-engineering-decisions.md` already recommends the next phase as "可复现 MVP Demo", including `mvp-demo` profile, fixed diagnosis case, one-click request, and `GET /api/diagnosis/{sessionId}/trace`.
|
- `mvp/archive/2026-07-09-doc-cleanup/notes/agent-engineering-decisions.md` already recommends the next phase as "可复现 MVP Demo", including `mvp-demo` profile, fixed diagnosis case, one-click request, and `GET /api/diagnosis/{sessionId}/trace`.
|
||||||
- `mvp/issues/ISS-003-mvp-design-implementation-review.md` identifies test stability, session traceability, verifier evidence chain, upload path, and SupervisorAgent consistency as recent MVP concerns. Security cleanup is intentionally deferred by user decision.
|
- `mvp/issues/active/ISS-003-mvp-design-implementation-review.md` identifies test stability, session traceability, verifier evidence chain, upload path, and SupervisorAgent consistency as recent MVP concerns. Security cleanup is intentionally deferred by user decision.
|
||||||
|
|
||||||
## Question Pool
|
## Question Pool
|
||||||
|
|
||||||
|
|||||||
@@ -2,7 +2,7 @@
|
|||||||
|
|
||||||
## Draft Acceptance
|
## Draft Acceptance
|
||||||
|
|
||||||
- [x] Issue exists: `mvp/issues/executor-evidence-attribution-hallucination.md`.
|
- [x] Issue exists: `mvp/issues/active/executor-evidence-attribution-hallucination.md`.
|
||||||
- [x] OpenSpec change artifacts exist.
|
- [x] OpenSpec change artifacts exist.
|
||||||
- [x] devflow tracking files exist.
|
- [x] devflow tracking files exist.
|
||||||
- [x] OpenSpec validation passes.
|
- [x] OpenSpec validation passes.
|
||||||
|
|||||||
@@ -5,7 +5,7 @@
|
|||||||
- `devflow/projects/2026-07-07-executor-evidence-output-contract`: V1 evidence-attribution contract kept `user_facing_answer`.
|
- `devflow/projects/2026-07-07-executor-evidence-output-contract`: V1 evidence-attribution contract kept `user_facing_answer`.
|
||||||
- `devflow/projects/2026-07-02-chat-verifier-agent`: Verifier consumes explicit inputs and should not see intermediate reasoning.
|
- `devflow/projects/2026-07-02-chat-verifier-agent`: Verifier consumes explicit inputs and should not see intermediate reasoning.
|
||||||
- `devflow/projects/2026-07-04-evidence-trace-hardening`: evidence summaries and tool invocation references are the evidence foundation.
|
- `devflow/projects/2026-07-04-evidence-trace-hardening`: evidence summaries and tool invocation references are the evidence foundation.
|
||||||
- `mvp/issues/executor-structured-output-v2.md`: staged implementation design; stage one is Executor V2 output contract.
|
- `mvp/issues/design-notes/executor-structured-output-v2.md`: staged implementation design; stage one is Executor V2 output contract.
|
||||||
|
|
||||||
## Code Evidence
|
## Code Evidence
|
||||||
|
|
||||||
|
|||||||
@@ -41,4 +41,4 @@ Make Executor cite concrete tool evidence, make Gatekeeper validate that citatio
|
|||||||
## OpenSpec
|
## OpenSpec
|
||||||
|
|
||||||
- Change: `openspec/changes/verifier-evidence-reference-fidelity`
|
- Change: `openspec/changes/verifier-evidence-reference-fidelity`
|
||||||
- Source issue: `mvp/issues/ISS-007-verifier-evidence-summary-fidelity.md`
|
- Source issue: `mvp/issues/archived/ISS-007-verifier-evidence-summary-fidelity.md`
|
||||||
|
|||||||
@@ -0,0 +1,46 @@
|
|||||||
|
# Acceptance
|
||||||
|
|
||||||
|
## Static Verification
|
||||||
|
|
||||||
|
- `openspec validate interview-demo-quality-audit --strict`
|
||||||
|
- Result: passed.
|
||||||
|
- Coverage: OpenSpec proposal/design/spec/tasks consistency.
|
||||||
|
- PowerShell parser/runtime readiness check:
|
||||||
|
- Command: `powershell -NoProfile -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1 -BaseUrl http://127.0.0.1:1 -OutputDir target/demo-check-syntax`
|
||||||
|
- Result: expected failure with actionable readiness message.
|
||||||
|
- Coverage: script parses under Windows PowerShell and fails before issuing diagnosis requests when service is unreachable.
|
||||||
|
|
||||||
|
## Script Verification
|
||||||
|
|
||||||
|
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
|
||||||
|
- Result: passed.
|
||||||
|
- Coverage: 12/12 fixed eval fixtures, Prompt audit evaluator checks, Gatekeeper rule metadata checks, regenerated baseline reports.
|
||||||
|
- `mvn -q "-Dtest=ChatServiceSequentialAgentTest" test`
|
||||||
|
- Result: passed.
|
||||||
|
- Coverage: Chat verifier evaluation persists `prompt_audit`.
|
||||||
|
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ChatServiceSequentialAgentTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`
|
||||||
|
- Result: passed.
|
||||||
|
- Coverage: broader eval, baseline diff, Chat sequential flow, Gatekeeper, and Verifier input hook regression set.
|
||||||
|
- `mvn -q -DskipTests compile`
|
||||||
|
- Result: passed.
|
||||||
|
- Coverage: main source compilation.
|
||||||
|
|
||||||
|
## E2E Verification
|
||||||
|
|
||||||
|
- Started service with:
|
||||||
|
- `mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo`
|
||||||
|
- Ran:
|
||||||
|
- `powershell -NoProfile -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1 -BaseUrl http://localhost:9900 -SessionId mvp-demo-interview-quality-audit-001`
|
||||||
|
- Result: passed.
|
||||||
|
- Summary:
|
||||||
|
- `chatSuccess=true`
|
||||||
|
- `verdict=LOW_CONFID`
|
||||||
|
- `gatekeeperStatus=fail`
|
||||||
|
- `gatekeeperRuleSetVersion=gatekeeper-rules-v1`
|
||||||
|
- `promptAuditVersion=chat-prompts-v1`
|
||||||
|
- tools included `lookup_knowledge`, `query_logs`, `query_metrics`, and `get_available_log_topics`
|
||||||
|
- Note: live E2E remains a compatibility check, not the deterministic PASS oracle. The fixed fixture baseline is the regression source of truth.
|
||||||
|
|
||||||
|
## Not Verified
|
||||||
|
|
||||||
|
- Browser UI inspection was not required for this change because the scope is backend trace/eval/demo script documentation, not frontend behavior.
|
||||||
@@ -0,0 +1,23 @@
|
|||||||
|
# Interview Demo Quality Audit Brief
|
||||||
|
|
||||||
|
## Background
|
||||||
|
|
||||||
|
The MVP already demonstrates traceable Agent diagnosis with Planner, Executor, Gatekeeper, Verifier, Composer, evidence tools, trace persistence, and deterministic eval fixtures. The remaining interview-readiness gap is not a new Agent architecture; it is making the demo easier to run and making prompt/rule changes easier to audit.
|
||||||
|
|
||||||
|
## Goal
|
||||||
|
|
||||||
|
Stabilize the interview demo path, expand fixture-backed evaluation, and persist prompt/Gatekeeper audit metadata so the project can explain and verify Agent behavior during interviews.
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
- Add prompt audit metadata to Chat verifier evaluation.
|
||||||
|
- Extend deterministic eval cases and baseline reports.
|
||||||
|
- Add an interview demo preflight/check script.
|
||||||
|
- Update MVP demo and architecture documentation.
|
||||||
|
|
||||||
|
## Non-goals
|
||||||
|
|
||||||
|
- No public API or database schema changes.
|
||||||
|
- No new SubAgent split, MCP migration, process isolation, or AIOps LLM Verifier.
|
||||||
|
- No guarantee that every live LLM run returns PASS.
|
||||||
|
|
||||||
@@ -0,0 +1,111 @@
|
|||||||
|
# interview-demo-quality-audit Decisions
|
||||||
|
|
||||||
|
## Clarify
|
||||||
|
|
||||||
|
- Entry summary: stabilize the interview demo, expand deterministic eval coverage, and add Prompt/Gatekeeper version audit.
|
||||||
|
- Slug: `interview-demo-quality-audit`.
|
||||||
|
- Devflow scale: `standard`.
|
||||||
|
- Interface impact: L2 internal contract change because `verifier_evaluation` gains `prompt_audit`; no public HTTP API or database schema change.
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
- `devflow/index.md` used: related entries found for `diagnosis-eval-demo-gatekeeper-closure`, `executor-composer-final-answer`, `verifier-evidence-reference-fidelity`, `mvp-demo-interview-runbook`, and `diagnosis-eval-baseline-diff`.
|
||||||
|
- Relevant glossary:
|
||||||
|
- Evidence Tools produce incident facts and must be recorded in `tool_invocation`.
|
||||||
|
- Verifier should not use skills/runbooks as incident evidence.
|
||||||
|
- Diagnosis Playbook Skill is workflow guidance, not a fact source.
|
||||||
|
- Historical constraints that must enter OpenSpec:
|
||||||
|
- Diagnosis eval is deterministic and fixture-backed; no LLM-as-judge.
|
||||||
|
- Stable demo scenarios are documentation/payloads plus deterministic fixtures; live E2E is a compatibility check, not a guaranteed PASS oracle.
|
||||||
|
- Gatekeeper rule metadata is already metadata-only and should not become dynamic rule execution.
|
||||||
|
- Composer is the final expression layer and must not leak raw Executor JSON.
|
||||||
|
|
||||||
|
## Question Pool
|
||||||
|
|
||||||
|
| ID | Dimension | Mode | Question | Status |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| Q1 | Terminology | evidence-driven | What should the new audit metadata be called? | Resolved |
|
||||||
|
| Q2 | Boundary | evidence-driven | Does this require public API or schema changes? | Resolved |
|
||||||
|
| Q3 | Acceptance | evidence-driven | Which current assets define deterministic acceptance? | Resolved |
|
||||||
|
| Q4 | Technical | evidence-driven | Where should prompt version metadata live with minimal implementation risk? | Resolved |
|
||||||
|
| Q5 | Scope | user-interview | Should live E2E be mandatory for all scenarios? | Confirmed by objective as conditional |
|
||||||
|
|
||||||
|
## Evidence-driven Conclusions
|
||||||
|
|
||||||
|
- Q1 conclusion: use `prompt_audit` for prompt version metadata and keep existing `gatekeeper_result.rule_set_version`.
|
||||||
|
- Q2 conclusion: keep this as an internal trace/self-evaluation contract change. Do not add endpoints, tables, or new Agent roles.
|
||||||
|
- Q3 conclusion: `DiagnosisTraceEvaluatorTest`, baseline reports, fixed fixtures, and demo scripts define current acceptance style.
|
||||||
|
- Q4 conclusion: add a small Chat prompt audit catalog near `ChatService` prompt loading and persist a compact snapshot with verifier evaluation.
|
||||||
|
- Q5 conclusion: run live E2E with `mvp-demo` profile if dependencies are available; otherwise record the blocker and rely on deterministic eval/unit evidence.
|
||||||
|
|
||||||
|
## Specify / Alignment
|
||||||
|
|
||||||
|
| Check | Status | Notes |
|
||||||
|
|---|---|---|
|
||||||
|
| proposal goals -> proposal | Aligned | Proposal covers demo preflight, eval expansion, prompt audit, and docs. |
|
||||||
|
| proposal scope/constraints -> design | Aligned | Design records no public API/schema changes, prompt audit shape, eval fields, and demo script behavior. |
|
||||||
|
| design decisions -> specs/tasks | Aligned | Specs cover persisted prompt audit, evaluator checks, baseline, and demo script outputs. |
|
||||||
|
| specs observable behavior -> tasks | Aligned | Each requirement has implementation and verification tasks. |
|
||||||
|
|
||||||
|
## Audit
|
||||||
|
|
||||||
|
Input -> processing -> output chain:
|
||||||
|
|
||||||
|
```text
|
||||||
|
prompt resource metadata
|
||||||
|
-> ChatService / PromptAudit snapshot
|
||||||
|
-> verifier_evaluation.prompt_audit
|
||||||
|
-> Trace API / eval fixtures
|
||||||
|
-> DiagnosisTraceEvaluator baseline
|
||||||
|
|
||||||
|
run-interview-demo-check.ps1
|
||||||
|
-> service readiness
|
||||||
|
-> chat / trace / feedback
|
||||||
|
-> mvp/demo/output summary
|
||||||
|
```
|
||||||
|
|
||||||
|
Architecture risk assessment:
|
||||||
|
|
||||||
|
1. The audit shape is intentionally compact and internal; storing full prompt text would create noisy traces and possible sensitive-content risk.
|
||||||
|
2. Eval should assert versions by explicit metadata, not by prompt content hashes that churn during local prompt edits.
|
||||||
|
3. Live demo checks may still be LOW_CONFID because LLM output is not deterministic; deterministic fixtures remain the regression source of truth.
|
||||||
|
4. No devflow/OpenSpec conflict found.
|
||||||
|
|
||||||
|
## Commit Gate
|
||||||
|
|
||||||
|
- `openspec validate interview-demo-quality-audit --strict`: passed.
|
||||||
|
- File completeness:
|
||||||
|
- `proposal.md`: present.
|
||||||
|
- `design.md`: present.
|
||||||
|
- `specs/`: present for `chat-verifier-agent`, `diagnosis-eval-harness`, and `mvp-demo-trace-acceptance`.
|
||||||
|
- `tasks.md`: present.
|
||||||
|
- Consistency:
|
||||||
|
- Proposal goals map to design sections.
|
||||||
|
- Design decisions map to spec requirements and executable tasks.
|
||||||
|
- Task acceptance checks are verifiable.
|
||||||
|
- `.committed` marker created.
|
||||||
|
|
||||||
|
## Current Checkpoint
|
||||||
|
|
||||||
|
- Commit completed.
|
||||||
|
- Apply is authorized by the original objective: "完成后归档提交".
|
||||||
|
|
||||||
|
## Pre-apply Research
|
||||||
|
|
||||||
|
- Capability source: sm-flow built-in apply protocol. `openspec-apply-change` was not invoked directly in this session.
|
||||||
|
- Repository semantic search/LSP note: the requested `codebase-retrieval` and LSP tools were not available in the exposed toolset, so impact analysis used `rg`, direct file reads, OpenSpec/devflow artifacts, and targeted tests.
|
||||||
|
- Reference implementation and reuse:
|
||||||
|
- `ChatService.persistVerifierEvaluation(...)` is the single persistence point for Chat verifier/composer audit data; prompt audit was added there to cover normal, fallback, and degraded Composer paths.
|
||||||
|
- `DiagnosisTraceEvaluator` and `DiagnosisEvalReportWriter` are the deterministic eval extension points; no LLM judge was introduced.
|
||||||
|
- `mvp/demo/scripts/run-payment-timeout-demo.ps1` provided the request/trace/feedback flow reused by the new interview preflight script.
|
||||||
|
- Interface impact remains L2 internal trace contract: `verifier_evaluation.prompt_audit` and eval report fields are added; no public endpoint, table, or request DTO changed.
|
||||||
|
|
||||||
|
## Apply Notes
|
||||||
|
|
||||||
|
- Added compact Chat prompt audit metadata: `chat-prompts-v1`, with planner/executor/verifier/composer prompt versions and resource paths.
|
||||||
|
- Extended diagnosis eval schema, result reporting, baseline fixtures, JSON report, and Markdown report for Prompt audit and Gatekeeper rule metadata.
|
||||||
|
- Added two fixture-backed audit cases:
|
||||||
|
- `prompt-gatekeeper-audit-closure`
|
||||||
|
- `audit-metadata-low-confid`
|
||||||
|
- Added `mvp/demo/scripts/run-interview-demo-check.ps1` to run service readiness, Chat, Trace, feedback, and summary output.
|
||||||
|
- Updated MVP demo/eval/architecture docs to explain `prompt_audit.version`, `gatekeeper_result.rule_set_version`, and deterministic fixture baseline.
|
||||||
@@ -0,0 +1,58 @@
|
|||||||
|
# Evidence
|
||||||
|
|
||||||
|
## Context Files Read
|
||||||
|
|
||||||
|
- `devflow/index.md`
|
||||||
|
- `devflow/glossary/CONTEXT.md`
|
||||||
|
- `devflow/projects/2026-07-08-diagnosis-eval-demo-gatekeeper-closure/decisions.md`
|
||||||
|
- `devflow/projects/2026-07-08-executor-composer-final-answer/decisions.md`
|
||||||
|
- `mvp/architecture/current-mvp-architecture.md`
|
||||||
|
- `mvp/architecture/agent-orchestration.md`
|
||||||
|
- `mvp/architecture/executor-evidence-pipeline-refactor.md`
|
||||||
|
- `mvp/architecture/harness-quality-gates.md`
|
||||||
|
- `mvp/demo/README.md`
|
||||||
|
- `mvp/demo/ten-minute-interview-demo.md`
|
||||||
|
- `mvp/eval/README.md`
|
||||||
|
- `mvp/eval/cases/diagnosis-cases.json`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/ExecutorGatekeeperService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/eval/DiagnosisTraceEvaluator.java`
|
||||||
|
- `src/main/resources/gatekeeper/gatekeeper-rules.json`
|
||||||
|
|
||||||
|
## Tooling Note
|
||||||
|
|
||||||
|
The required `codebase-retrieval` and LSP tools were not exposed in this session. Impact analysis used `rg`, direct file reads, existing OpenSpec/devflow artifacts, and targeted tests instead.
|
||||||
|
|
||||||
|
## Implementation Evidence
|
||||||
|
|
||||||
|
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||||
|
- Adds `prompt_audit` under `verifier_evaluation` through the shared `persistVerifierEvaluation(...)` path.
|
||||||
|
- Uses compact metadata only: audit version, prompt names, prompt versions, and resource paths.
|
||||||
|
- `src/main/java/com/superbiz/agent/eval/DiagnosisTraceEvaluator.java`
|
||||||
|
- Adds deterministic checks for `requirePromptAudit`, `expectedPromptAuditVersion`, `expectedPromptVersions`, and `requireGatekeeperRules`.
|
||||||
|
- `src/main/java/com/superbiz/agent/eval/DiagnosisEvalReportWriter.java`
|
||||||
|
- Adds Prompt Audit and Gatekeeper rule count columns to Markdown reports.
|
||||||
|
- `mvp/eval/cases/diagnosis-cases.json`
|
||||||
|
- Expands fixed baseline to 12 fixture-backed cases.
|
||||||
|
- `mvp/eval/fixtures/prompt-gatekeeper-audit-closure-pass.json`
|
||||||
|
- Positive PASS fixture proving Prompt audit and Gatekeeper rule metadata closure.
|
||||||
|
- `mvp/eval/fixtures/audit-metadata-low-confid.json`
|
||||||
|
- LOW_CONFID fixture proving safe answer behavior while audit metadata remains present.
|
||||||
|
- `mvp/demo/scripts/run-interview-demo-check.ps1`
|
||||||
|
- Adds service readiness, Chat, Trace, feedback, and summary output for interview preflight.
|
||||||
|
|
||||||
|
## Verification Evidence
|
||||||
|
|
||||||
|
- OpenSpec:
|
||||||
|
- `openspec validate interview-demo-quality-audit --strict`: passed before archive.
|
||||||
|
- `openspec validate --specs --strict`: 10 specs passed after merging deltas into main specs.
|
||||||
|
- Unit/eval:
|
||||||
|
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`: passed.
|
||||||
|
- `mvn -q "-Dtest=ChatServiceSequentialAgentTest" test`: passed.
|
||||||
|
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ChatServiceSequentialAgentTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`: passed.
|
||||||
|
- Compile:
|
||||||
|
- `mvn -q -DskipTests compile`: passed.
|
||||||
|
- E2E:
|
||||||
|
- Started `mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo`.
|
||||||
|
- Ran `mvp/demo/scripts/run-interview-demo-check.ps1` against `http://localhost:9900`.
|
||||||
|
- Summary recorded `chatSuccess=true`, `verdict=LOW_CONFID`, `gatekeeperRuleSetVersion=gatekeeper-rules-v1`, and `promptAuditVersion=chat-prompts-v1`.
|
||||||
@@ -0,0 +1,103 @@
|
|||||||
|
# Acceptance
|
||||||
|
|
||||||
|
## 实现结果
|
||||||
|
|
||||||
|
- OpenSpec tasks: `42/42` complete。
|
||||||
|
- Phase commits:
|
||||||
|
- `52bf030 feat(trace): add session run isolation schema`
|
||||||
|
- `6fdbd34 docs(openspec): tighten run isolation contract`
|
||||||
|
- `26d5529 feat(trace): isolate chat runs`
|
||||||
|
- `027aed1 feat(trace): add run-scoped trace reads`
|
||||||
|
- `d928a19 feat(trace): bind feedback to runs`
|
||||||
|
- `78c1477 feat(trace): isolate aiops runs`
|
||||||
|
- `f9df943 feat(trace): finish run-aware demo verification`
|
||||||
|
- OpenSpec archive: `openspec/changes/archive/2026-07-10-session-run-trace-isolation`
|
||||||
|
|
||||||
|
## 静态验证
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
node --check src\main\resources\static\app.js
|
||||||
|
node --check src\main\resources\static\trace.js
|
||||||
|
openspec validate session-run-trace-isolation --strict
|
||||||
|
git diff --check -- . ':!devflow/index.md'
|
||||||
|
```
|
||||||
|
|
||||||
|
结果:通过。
|
||||||
|
|
||||||
|
## 脚本验证
|
||||||
|
|
||||||
|
PowerShell demo 脚本解析:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
$scripts = @(
|
||||||
|
'mvp\demo\scripts\run-payment-timeout-demo.ps1',
|
||||||
|
'mvp\demo\scripts\run-interview-demo-check.ps1'
|
||||||
|
)
|
||||||
|
foreach ($script in $scripts) {
|
||||||
|
[scriptblock]::Create((Get-Content -Raw -Encoding UTF8 $script)) | Out-Null
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
结果:通过。
|
||||||
|
|
||||||
|
Focused tests:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
mvn -q "-Dtest=ChatControllerTest,DiagnosisTraceServiceTest,FeedbackControllerTest,FeedbackServiceTest,AiOpsServiceTest" test
|
||||||
|
```
|
||||||
|
|
||||||
|
结果:通过。
|
||||||
|
|
||||||
|
Baseline / regression:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest" test
|
||||||
|
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
|
||||||
|
```
|
||||||
|
|
||||||
|
结果:通过,无 baseline drift。
|
||||||
|
|
||||||
|
## E2E 验证
|
||||||
|
|
||||||
|
使用 Maven 启动:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
|
||||||
|
```
|
||||||
|
|
||||||
|
E2E 使用同一 `sessionId` 连续两轮 Chat:
|
||||||
|
|
||||||
|
- `sessionId`: `e2e-phase6-chat-codex-20260710-2120`
|
||||||
|
- `run1`: `run-e2a97696-4398-4abc-90e4-28f45c838f92`
|
||||||
|
- `run2`: `run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172`
|
||||||
|
|
||||||
|
验证结果:
|
||||||
|
|
||||||
|
- Chat1 / Chat2 均成功。
|
||||||
|
- run1 exact trace 返回 run1。
|
||||||
|
- run2 exact trace 返回 run2。
|
||||||
|
- session-only latest trace 返回 run2。
|
||||||
|
- DB 中同一 session 有两条 `diagnosis_run`。
|
||||||
|
- step/tool rows 按 `run_id` 隔离,mixed row check 为 0。
|
||||||
|
- `chat_session.message_pair_count = 2`。
|
||||||
|
|
||||||
|
## 日志验证
|
||||||
|
|
||||||
|
检查:
|
||||||
|
|
||||||
|
- `target/e2e/phase6-mvn-20260710-211831.out.log`
|
||||||
|
- `logs/application.log`
|
||||||
|
- `logs/chat.log`
|
||||||
|
|
||||||
|
结果:能找到 E2E `sessionId`、两个 `runId`、Chat execution、run persistence 和 trace lookup 相关日志。
|
||||||
|
|
||||||
|
## 浏览器/人工验证
|
||||||
|
|
||||||
|
未单独进行浏览器点击验证。Trace UI 的本次验收通过静态语法检查、URL/runId 参数代码审查和后端 exact trace E2E 共同覆盖。建议后续手动打开 `trace.html?sessionId=...&runId=...` 做展示层冒烟。
|
||||||
|
|
||||||
|
## 剩余风险 / 后续事项
|
||||||
|
|
||||||
|
- 缺少 `runId` 的 Feedback fallback 是短期兼容路径,客户端全部迁移后可收紧。
|
||||||
|
- `diagnosis_session` 仍保留为历史兼容和回滚表,后续需要观察窗口后再评估约束收紧或归档策略。
|
||||||
|
- `case_library.diagnosis_id` 仍是过渡字段,旧值可能为 `session_id`,新自动值为 `run_id`。
|
||||||
|
- 历史 mixed trace 不能恢复真实多轮边界,只能按 compatibility run 查询。
|
||||||
@@ -0,0 +1,41 @@
|
|||||||
|
# Session / Run / Trace Isolation
|
||||||
|
|
||||||
|
## 背景
|
||||||
|
|
||||||
|
同一个 `sessionId` 以前同时代表多轮 Chat 上下文和一次持久化诊断 Trace。端到端验证发现,同一 `sessionId` 连续两轮 Chat 时,Redis 多轮上下文是正确的,但 MySQL 中 `diagnosis_session` 会被后一轮覆盖,`agent_step` 和 `tool_invocation` 会按同一个 `session_id` 混在一起。
|
||||||
|
|
||||||
|
这会导致 Trace 回放、Verifier/Evaluation 读数、Feedback 绑定和 `case_library` 来源都可能跨轮污染。
|
||||||
|
|
||||||
|
## 目标
|
||||||
|
|
||||||
|
- 将会话态和运行态拆开:`chat_session` 保存会话元数据,`diagnosis_run` 保存一次诊断运行。
|
||||||
|
- 引入正式 API 字段 `runId`,作为一次可回放诊断执行的边界。
|
||||||
|
- `agent_step` 和 `tool_invocation` 保留原 Trace 明细角色,新增 `run_id` 并按 run 隔离读写。
|
||||||
|
- Trace、Feedback、CaseLibrary、AIOps、demo 脚本和 Trace UI 都支持 run-aware 流程。
|
||||||
|
- 保留旧 `diagnosis_session` 作为历史兼容和回滚表。
|
||||||
|
- 完成 Maven E2E、DB 检查、日志检查和 baseline drift 验证。
|
||||||
|
|
||||||
|
## 范围
|
||||||
|
|
||||||
|
- Flyway/JPA 增加 `chat_session`、`diagnosis_run`,并给 `agent_step`、`tool_invocation` 增加 `run_id`。
|
||||||
|
- Chat 每次有效执行创建一个新的 `diagnosis_run`,响应返回 `sessionId + runId`。
|
||||||
|
- Trace API 支持 latest-run fallback 和 exact-run 查询:`GET /api/diagnosis/{sessionId}/trace?runId=...`。
|
||||||
|
- 新增 run list API:`GET /api/chat/session/{sessionId}/runs`。
|
||||||
|
- Feedback 优先绑定 `runId`,缺省时短期 fallback 到 latest run 并返回 `fallbackToLatestRun=true`。
|
||||||
|
- AIOps 每次有效执行创建并透出 `runId`,SSE 保持 `message` event name 并发送 `type=metadata`。
|
||||||
|
- MVP demo、Trace UI、表文档和架构文档统一为 `chat_session -> diagnosis_run -> trace detail(run_id)`。
|
||||||
|
|
||||||
|
## 非目标
|
||||||
|
|
||||||
|
- 不新增 `diagnosis_trace` 或 `trace_event` 主表。
|
||||||
|
- 不实现完整 run-list UI。
|
||||||
|
- 不删除旧 `diagnosis_session`。
|
||||||
|
- 不改变 Redis 对话历史窗口策略。
|
||||||
|
- 不把完整多轮正文历史持久化到 MySQL。
|
||||||
|
- 不尝试把历史混合 trace 还原成真实多轮边界。
|
||||||
|
|
||||||
|
## 关联
|
||||||
|
|
||||||
|
- OpenSpec: `openspec/changes/archive/2026-07-10-session-run-trace-isolation`
|
||||||
|
- Change slug: `session-run-trace-isolation`
|
||||||
|
- 分档: complex
|
||||||
@@ -0,0 +1,48 @@
|
|||||||
|
# Decisions
|
||||||
|
|
||||||
|
## 核心决策
|
||||||
|
|
||||||
|
| 决策 | 选择 | 理由 |
|
||||||
|
|---|---|---|
|
||||||
|
| 领域拆分 | 新增 `chat_session` 和 `diagnosis_run` | 会话元数据和一次诊断执行的生命周期不同,继续塞在一张表会导致上下文膨胀和边界混淆 |
|
||||||
|
| Trace 明细 | 复用 `agent_step` / `tool_invocation`,增加 `run_id` | 现有明细表已经能表达 Trace,隔离需要 run key,不需要新事件模型 |
|
||||||
|
| API 身份 | `runId = "run-" + UUID` | 外部 ID 不依赖数据库自增 ID,碰撞风险低 |
|
||||||
|
| Trace 兼容 | 缺少 `runId` 时按 `created_at DESC, id DESC` 解析 latest run | 保留旧客户端兼容性,避免 feedback/eval 更新 `updated_at` 后改变 latest 判定 |
|
||||||
|
| 历史迁移 | 每条旧 `diagnosis_session` 生成一条 compatibility run | 旧混合数据没有真实轮次边界,不能伪造多 run 历史 |
|
||||||
|
| Feedback fallback | 缺少 `runId` 时短期绑定 latest run 并返回 `fallbackToLatestRun=true` | 老客户端可继续工作,同时让歧义可观测 |
|
||||||
|
| Case provenance | 新自动案例写 `case_library.diagnosis_id = run_id` | 保留旧列,文档声明过渡语义 |
|
||||||
|
| AIOps 范围 | 同一个 change 内完成 AIOps run isolation | AIOps 是一等 Trace 入口,不能留下同类混合 trace bug |
|
||||||
|
| 所有权校验 | 服务层校验 run/session ownership,暂不加 DB 外键 | 兼容历史 orphan rows 和回滚窗口 |
|
||||||
|
|
||||||
|
## 用户确认
|
||||||
|
|
||||||
|
- 选择拆 `chat_session` 和 `diagnosis_run`,不只是在旧表加字段。
|
||||||
|
- `chat_session` 第一阶段只保存元数据,不保存完整对话正文。
|
||||||
|
- 完整多轮对话历史继续放在 Redis `SessionContext.messageHistory`。
|
||||||
|
- `runId` 是正式 API 字段。
|
||||||
|
- Trace 缺少 `runId` 时短期默认查 latest run。
|
||||||
|
- Feedback 缺少 `runId` 时短期 fallback,长期可再收紧。
|
||||||
|
- 每次有效 Chat/AIOps 都创建 run。
|
||||||
|
- 不新增 `diagnosis_trace` / `trace_event` 主表。
|
||||||
|
- 旧 `diagnosis_session` 保留用于历史和回滚,新代码不再写新执行态。
|
||||||
|
- demo 脚本和 Trace UI 做最小 `runId` 支持。
|
||||||
|
|
||||||
|
## 接口影响
|
||||||
|
|
||||||
|
级别:L4。
|
||||||
|
|
||||||
|
- 新 API 响应字段:`runId`。
|
||||||
|
- Trace API 新 query 参数:`runId`。
|
||||||
|
- 新 API:`GET /api/chat/session/{sessionId}/runs`。
|
||||||
|
- Feedback request 新增 optional/preferred `runId`。
|
||||||
|
- Feedback response 新增 bound `runId` 和 `fallbackToLatestRun`。
|
||||||
|
- `/api/ai_ops` SSE 保持 event name `message`,新增 `type=metadata` 消息。
|
||||||
|
- DB contract 新增两张表和两个 `run_id` 列。
|
||||||
|
- 旧 `sessionId` only 调用仍兼容,但 fallback 必须可观测。
|
||||||
|
|
||||||
|
## 风险接受
|
||||||
|
|
||||||
|
- 历史混合 trace 无法真实拆分,只能作为 compatibility run。
|
||||||
|
- 上下文传播同时依赖 `RunnableConfig.metadata` 和 `SessionContextHolder`,后续改动必须注意 `sessionId/runId` 同步。
|
||||||
|
- `case_library.diagnosis_id` 在过渡期存在 `session_id` 和 `run_id` 两种语义。
|
||||||
|
- 缺少 `runId` 的 Feedback 仍有歧义,后续客户端迁移完成后可收紧为参数错误。
|
||||||
@@ -0,0 +1,76 @@
|
|||||||
|
# Evidence
|
||||||
|
|
||||||
|
## 上下文证据
|
||||||
|
|
||||||
|
- `SessionContext.messageHistory` 和 `getMessagePairCount()` 证明 Redis 承载热对话历史;MySQL 只需要长期审计的会话目录和运行记录。
|
||||||
|
- `CaseLibraryService.createFromSession` 原先按 `DiagnosisSession.sessionId` 去重并映射 query/answer,因此 run 隔离后需要新增 `createFromRun`。
|
||||||
|
- 旧 `mvp/architecture/data-model.md` 把 `case_library.diagnosis_id` 解释为 `diagnosis_session.session_id`,本次改为过渡语义:旧数据可能是 `session_id`,新自动案例是 `run_id`。
|
||||||
|
- 既有 Trace OpenSpec 要求 `GET /api/diagnosis/{sessionId}/trace` 是只读端点;latest-run 和 exact-run 查询都必须保持只读。
|
||||||
|
- ISS-010 的 E2E 事实显示同一 `sessionId` 两轮 Chat 会产生 MySQL Trace 混合,是本 change 的直接触发证据。
|
||||||
|
|
||||||
|
## 实现证据
|
||||||
|
|
||||||
|
- Phase 1 增加 `V011__add_session_run_isolation.sql`,创建 `chat_session`、`diagnosis_run`,并为 `agent_step` / `tool_invocation` 增加 nullable `run_id`。
|
||||||
|
- Phase 2 将 Chat 写路径切到 `chat_session + diagnosis_run`,并让 Hook/Tool/Evaluation/Gatekeeper 使用 run-scoped 数据。
|
||||||
|
- Phase 3 将 Trace API 改为 latest-run / exact-run 双模式,并加入 lightweight run summaries。
|
||||||
|
- Phase 4 将 Feedback 和 CaseLibrary 绑定到 run,保留没有 run-backed 数据时的 legacy fallback。
|
||||||
|
- Phase 5 将 AIOps 接入 run isolation,SSE metadata 暴露 `sessionId + runId`。
|
||||||
|
- Phase 6 更新 demo 脚本、Trace UI、MVP 架构文档和表文档,并修正 review 后发现的 session-only 文档残留。
|
||||||
|
|
||||||
|
## E2E 证据
|
||||||
|
|
||||||
|
Maven 启动命令:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
|
||||||
|
```
|
||||||
|
|
||||||
|
日志:
|
||||||
|
|
||||||
|
- `target/e2e/phase6-mvn-20260710-211831.out.log`
|
||||||
|
- `target/e2e/phase6-mvn-20260710-211831.err.log`
|
||||||
|
- `logs/application.log`
|
||||||
|
- `logs/chat.log`
|
||||||
|
|
||||||
|
E2E session:
|
||||||
|
|
||||||
|
- `sessionId`: `e2e-phase6-chat-codex-20260710-2120`
|
||||||
|
- `run1`: `run-e2a97696-4398-4abc-90e4-28f45c838f92`
|
||||||
|
- `run2`: `run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172`
|
||||||
|
|
||||||
|
结果:
|
||||||
|
|
||||||
|
- 两轮 Chat 都成功,并复用同一个 `sessionId`。
|
||||||
|
- 两轮返回不同 `runId`。
|
||||||
|
- run1 exact trace 只返回 run1。
|
||||||
|
- run2 exact trace 只返回 run2。
|
||||||
|
- session-only Trace latest fallback 返回 run2。
|
||||||
|
- `chat_session.message_pair_count = 2`,证明多轮上下文连续。
|
||||||
|
|
||||||
|
## DB 证据
|
||||||
|
|
||||||
|
通过 `scripts/query_mysql.py` 检查:
|
||||||
|
|
||||||
|
- `diagnosis_run` 中该 E2E session 有 2 条 `SUCCESS / CHAT` 运行。
|
||||||
|
- `agent_step` 按 run 分组:run1 `10` 行,run2 `9` 行。
|
||||||
|
- `tool_invocation` 按 run 分组:run1 `14` 行,run2 `8` 行。
|
||||||
|
- mixed row check 为 `0`,没有 NULL 或 unexpected `run_id` 混入该 E2E session。
|
||||||
|
|
||||||
|
## Baseline 证据
|
||||||
|
|
||||||
|
运行:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest" test
|
||||||
|
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
|
||||||
|
```
|
||||||
|
|
||||||
|
结果:
|
||||||
|
|
||||||
|
- 两组 baseline / regression 命令通过。
|
||||||
|
- baseline harness 使用离线 fixture,不依赖 live DB/session tables。
|
||||||
|
- 未观察到 baseline drift。
|
||||||
|
|
||||||
|
## 工具限制
|
||||||
|
|
||||||
|
AGENTS 要求的 `codebase-retrieval` 和 LSP 工具在本会话不可用。替代验证使用 OpenSpec、`rg`、定向阅读、 focused tests、E2E、DB 查询和日志检查。
|
||||||
+28
-18
@@ -1,8 +1,8 @@
|
|||||||
# SuperBizAgent MVP 文档
|
# SuperBizAgent MVP 文档
|
||||||
|
|
||||||
**更新日期**:2026-07-05
|
**更新日期**:2026-07-10
|
||||||
|
|
||||||
本目录保存 MVP 阶段的架构、问题、演示、评测和数据表说明。当前架构入口已经整理到 `mvp/architecture/`,旧版架构材料已归档,避免继续把历史方案当成当前实现。
|
本目录保存 MVP 阶段的架构、问题、演示、评测和数据表说明。当前材料按“当前入口”和“历史归档”拆开,避免把早期设计稿当成当前实现。
|
||||||
|
|
||||||
## 当前入口
|
## 当前入口
|
||||||
|
|
||||||
@@ -12,6 +12,7 @@
|
|||||||
| [architecture/current-mvp-architecture.md](architecture/current-mvp-architecture.md) | 当前可运行系统架构 |
|
| [architecture/current-mvp-architecture.md](architecture/current-mvp-architecture.md) | 当前可运行系统架构 |
|
||||||
| [architecture/interview-one-pager.md](architecture/interview-one-pager.md) | 面试一页式架构讲解 |
|
| [architecture/interview-one-pager.md](architecture/interview-one-pager.md) | 面试一页式架构讲解 |
|
||||||
| [architecture/agent-orchestration.md](architecture/agent-orchestration.md) | Agent 编排架构 |
|
| [architecture/agent-orchestration.md](architecture/agent-orchestration.md) | Agent 编排架构 |
|
||||||
|
| [architecture/executor-evidence-pipeline-refactor.md](architecture/executor-evidence-pipeline-refactor.md) | Executor 证据链路改造记录 |
|
||||||
| [architecture/harness-quality-gates.md](architecture/harness-quality-gates.md) | Harness 与质量门禁 |
|
| [architecture/harness-quality-gates.md](architecture/harness-quality-gates.md) | Harness 与质量门禁 |
|
||||||
| [architecture/rag-architecture.md](architecture/rag-architecture.md) | RAG/知识检索新架构 |
|
| [architecture/rag-architecture.md](architecture/rag-architecture.md) | RAG/知识检索新架构 |
|
||||||
| [architecture/retrieval-observability.md](architecture/retrieval-observability.md) | 检索与可观测性架构 |
|
| [architecture/retrieval-observability.md](architecture/retrieval-observability.md) | 检索与可观测性架构 |
|
||||||
@@ -20,15 +21,16 @@
|
|||||||
| [architecture/knowledge-base-authoring.md](architecture/knowledge-base-authoring.md) | 知识库文档编写与维护 |
|
| [architecture/knowledge-base-authoring.md](architecture/knowledge-base-authoring.md) | 知识库文档编写与维护 |
|
||||||
| [architecture/data-model.md](architecture/data-model.md) | 数据模型总览 |
|
| [architecture/data-model.md](architecture/data-model.md) | 数据模型总览 |
|
||||||
| [architecture/evolution-roadmap.md](architecture/evolution-roadmap.md) | Agent 架构演进路线 |
|
| [architecture/evolution-roadmap.md](architecture/evolution-roadmap.md) | Agent 架构演进路线 |
|
||||||
| [issues/rag-refactor-plan.md](issues/rag-refactor-plan.md) | RAG 重构计划和阶段拆解 |
|
| [issues/README.md](issues/README.md) | MVP issue 索引 |
|
||||||
|
| [issues/active/rag-refactor-plan.md](issues/active/rag-refactor-plan.md) | RAG 重构计划和阶段拆解 |
|
||||||
|
| [tables/README.md](tables/README.md) | 当前 MySQL 表说明 |
|
||||||
| [demo/README.md](demo/README.md) | Demo 运行和面试演示材料 |
|
| [demo/README.md](demo/README.md) | Demo 运行和面试演示材料 |
|
||||||
| [demo/ten-minute-interview-demo.md](demo/ten-minute-interview-demo.md) | 10 分钟面试演示脚本 |
|
| [demo/ten-minute-interview-demo.md](demo/ten-minute-interview-demo.md) | 10 分钟面试演示脚本 |
|
||||||
| [eval/README.md](eval/README.md) | 诊断评测材料 |
|
| [eval/README.md](eval/README.md) | 诊断评测材料 |
|
||||||
| [issues/README.md](issues/README.md) | MVP issue 索引 |
|
|
||||||
|
|
||||||
## 当前系统一句话
|
## 当前系统一句话
|
||||||
|
|
||||||
SuperBizAgent MVP 是一个可追踪的故障诊断 Agent:Chat 和 AIOps 入口进入 Agent 编排,Executor 显式调用知识库、日志、指标等工具收集证据,诊断过程落到 `diagnosis_session`、`agent_step`、`tool_invocation`,最终通过 Trace API、Verifier 和评测脚本证明结果可解释、可回放、可对比。
|
SuperBizAgent MVP 是一个可追踪的故障诊断 Agent:Chat 和 AIOps 入口进入 Agent 编排,Executor 显式调用知识库、日志、指标等工具收集证据;多轮会话元数据落到 `chat_session`,每次诊断运行落到 `diagnosis_run`,步骤和工具明细通过 `agent_step.run_id`、`tool_invocation.run_id` 关联,最终通过 Trace API、Verifier 和评测脚本证明结果可解释、可回放、可对比。
|
||||||
|
|
||||||
## 文档结构
|
## 文档结构
|
||||||
|
|
||||||
@@ -39,6 +41,7 @@ mvp/
|
|||||||
current-mvp-architecture.md
|
current-mvp-architecture.md
|
||||||
interview-one-pager.md
|
interview-one-pager.md
|
||||||
agent-orchestration.md
|
agent-orchestration.md
|
||||||
|
executor-evidence-pipeline-refactor.md
|
||||||
harness-quality-gates.md
|
harness-quality-gates.md
|
||||||
rag-architecture.md
|
rag-architecture.md
|
||||||
retrieval-observability.md
|
retrieval-observability.md
|
||||||
@@ -47,12 +50,17 @@ mvp/
|
|||||||
knowledge-base-authoring.md
|
knowledge-base-authoring.md
|
||||||
data-model.md
|
data-model.md
|
||||||
evolution-roadmap.md
|
evolution-roadmap.md
|
||||||
archive/2026-07-05-legacy/
|
archive/
|
||||||
issues/
|
issues/
|
||||||
README.md
|
README.md
|
||||||
rag-refactor-plan.md
|
active/
|
||||||
ISS-*.md
|
archived/
|
||||||
rag-*.md
|
design-notes/
|
||||||
|
rag/
|
||||||
|
tables/
|
||||||
|
README.md
|
||||||
|
*表-*.md
|
||||||
|
archive/
|
||||||
demo/
|
demo/
|
||||||
README.md
|
README.md
|
||||||
ten-minute-interview-demo.md
|
ten-minute-interview-demo.md
|
||||||
@@ -65,9 +73,7 @@ mvp/
|
|||||||
cases/
|
cases/
|
||||||
fixtures/
|
fixtures/
|
||||||
reports/
|
reports/
|
||||||
notes/
|
archive/
|
||||||
plan/
|
|
||||||
tables/
|
|
||||||
```
|
```
|
||||||
|
|
||||||
## 当前核心设计
|
## 当前核心设计
|
||||||
@@ -77,7 +83,8 @@ mvp/
|
|||||||
- `VectorSearchService` 是检索稳定门面。
|
- `VectorSearchService` 是检索稳定门面。
|
||||||
- Spring AI VectorStore 是当前读取主路径,Milvus SDK 保留为 fallback。
|
- Spring AI VectorStore 是当前读取主路径,Milvus SDK 保留为 fallback。
|
||||||
- AIOps payload 会生成推荐知识库 query,保留业务语义。
|
- AIOps payload 会生成推荐知识库 query,保留业务语义。
|
||||||
- Trace API 聚合 session、step、tool invocation 和 self evaluation。
|
- `sessionId` 表示多轮会话上下文,`runId` 表示一次可回放诊断运行。
|
||||||
|
- Trace API 聚合 `diagnosis_run`、`agent_step.run_id`、`tool_invocation.run_id` 和 self evaluation。
|
||||||
- RAG 行为通过 offline baseline 和 live acceptance 脚本做回归验证。
|
- RAG 行为通过 offline baseline 和 live acceptance 脚本做回归验证。
|
||||||
|
|
||||||
## 关键运行链路
|
## 关键运行链路
|
||||||
@@ -87,7 +94,8 @@ Chat
|
|||||||
-> ChatService
|
-> ChatService
|
||||||
-> Planner / Executor / Verifier
|
-> Planner / Executor / Verifier
|
||||||
-> evidence tools
|
-> evidence tools
|
||||||
-> diagnosis_session / agent_step / tool_invocation
|
-> chat_session / diagnosis_run
|
||||||
|
-> agent_step.run_id / tool_invocation.run_id
|
||||||
-> DiagnosisTraceService
|
-> DiagnosisTraceService
|
||||||
|
|
||||||
AIOps
|
AIOps
|
||||||
@@ -96,6 +104,7 @@ AIOps
|
|||||||
-> Planner / Executor
|
-> Planner / Executor
|
||||||
-> Prometheus / logs / lookup_knowledge
|
-> Prometheus / logs / lookup_knowledge
|
||||||
-> AiOpsRuleEvaluationService
|
-> AiOpsRuleEvaluationService
|
||||||
|
-> diagnosis_run(agent_flow=AI_OPS)
|
||||||
-> DiagnosisTraceService
|
-> DiagnosisTraceService
|
||||||
|
|
||||||
RAG
|
RAG
|
||||||
@@ -107,10 +116,11 @@ RAG
|
|||||||
-> tool_invocation
|
-> tool_invocation
|
||||||
```
|
```
|
||||||
|
|
||||||
## 旧文档说明
|
## 归档说明
|
||||||
|
|
||||||
旧版架构文档已移动到:
|
历史材料分两类:
|
||||||
|
|
||||||
- [architecture/archive/2026-07-05-legacy/](architecture/archive/2026-07-05-legacy/)
|
- 旧架构文档:[architecture/archive/2026-07-05-legacy/](architecture/archive/2026-07-05-legacy/)
|
||||||
|
- 本次文档清理归档:[archive/2026-07-09-doc-cleanup/](archive/2026-07-09-doc-cleanup/)
|
||||||
|
|
||||||
归档文档只用于追溯设计历史。当前实现和后续规划以 `architecture/current-mvp-architecture.md` 与 `architecture/rag-architecture.md` 为准。
|
归档文档只用于追溯设计历史。当前实现和后续规划以 `architecture/`、`issues/README.md`、`tables/README.md` 和 OpenSpec/devflow 的最新记录为准。
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
# MVP 架构文档
|
# MVP 架构文档
|
||||||
|
|
||||||
**更新日期**:2026-07-08
|
**更新日期**:2026-07-10
|
||||||
|
|
||||||
这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到:
|
这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到:
|
||||||
|
|
||||||
@@ -29,7 +29,7 @@
|
|||||||
|
|
||||||
## 当前架构一句话
|
## 当前架构一句话
|
||||||
|
|
||||||
SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,Chat 链路由 Gatekeeper 做引用真实性校验、Verifier 做可推导性判断、Composer 生成最终表达;执行过程落到 `diagnosis_session`、`agent_step`、`tool_invocation`,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
|
SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,Chat 链路由 Gatekeeper 做引用真实性校验、Verifier 做可推导性判断、Composer 生成最终表达;多轮会话元数据落到 `chat_session`,每次诊断运行落到 `diagnosis_run`,步骤和工具明细通过 `agent_step.run_id`、`tool_invocation.run_id` 关联,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
|
||||||
|
|
||||||
## 阅读顺序
|
## 阅读顺序
|
||||||
|
|
||||||
|
|||||||
@@ -42,13 +42,15 @@ flowchart TB
|
|||||||
end
|
end
|
||||||
|
|
||||||
subgraph Trace["Trace persistence"]
|
subgraph Trace["Trace persistence"]
|
||||||
Session["diagnosis_session"]
|
ChatSession["chat_session"]
|
||||||
|
Run["diagnosis_run"]
|
||||||
Step["agent_step"]
|
Step["agent_step"]
|
||||||
Invocation["tool_invocation"]
|
Invocation["tool_invocation"]
|
||||||
SelfEval["self_evaluation"]
|
SelfEval["self_evaluation"]
|
||||||
end
|
end
|
||||||
|
|
||||||
ChatService --> Session
|
ChatService --> ChatSession
|
||||||
|
ChatService --> Run
|
||||||
ChatPlanner --> Step
|
ChatPlanner --> Step
|
||||||
ChatExecutor --> Step
|
ChatExecutor --> Step
|
||||||
ChatGatekeeper --> SelfEval
|
ChatGatekeeper --> SelfEval
|
||||||
@@ -57,7 +59,8 @@ flowchart TB
|
|||||||
ChatDecision --> SelfEval
|
ChatDecision --> SelfEval
|
||||||
ChatComposer --> Step
|
ChatComposer --> Step
|
||||||
|
|
||||||
AiOpsService --> Session
|
AiOpsService --> ChatSession
|
||||||
|
AiOpsService --> Run
|
||||||
AiOpsPlanner --> Step
|
AiOpsPlanner --> Step
|
||||||
AiOpsExecutor --> Step
|
AiOpsExecutor --> Step
|
||||||
AiOpsTools --> Invocation
|
AiOpsTools --> Invocation
|
||||||
@@ -103,7 +106,7 @@ sequenceDiagram
|
|||||||
participant G as gatekeeper
|
participant G as gatekeeper
|
||||||
participant V as chat_verifier
|
participant V as chat_verifier
|
||||||
participant M as chat_composer
|
participant M as chat_composer
|
||||||
participant S as diagnosis_session
|
participant R as diagnosis_run
|
||||||
|
|
||||||
C->>P: 原始问题 + history + retry_context
|
C->>P: 原始问题 + history + retry_context
|
||||||
P-->>C: planner_plan
|
P-->>C: planner_plan
|
||||||
@@ -115,13 +118,13 @@ sequenceDiagram
|
|||||||
G-->>C: gatekeeper_result
|
G-->>C: gatekeeper_result
|
||||||
C->>V: executor_structured_output + gatekeeper_result + tool_trace_summary
|
C->>V: executor_structured_output + gatekeeper_result + tool_trace_summary
|
||||||
V-->>C: PASS / LOW_CONFID / REJECT
|
V-->>C: PASS / LOW_CONFID / REJECT
|
||||||
C->>S: 写入 verifier_evaluation
|
C->>R: 写入 verifier_evaluation
|
||||||
alt LOW_CONFID 且允许补证据
|
alt LOW_CONFID 且允许补证据
|
||||||
C->>P: retry_context: 仅补缺失证据
|
C->>P: retry_context: 仅补缺失证据
|
||||||
else PASS 或 REJECT
|
else PASS 或 REJECT
|
||||||
C->>M: allowed_claims + missing_info + recommended_actions
|
C->>M: allowed_claims + missing_info + recommended_actions
|
||||||
M-->>C: composer_output
|
M-->>C: composer_output
|
||||||
C->>S: 保存 Composer 最终 answer
|
C->>R: 保存 Composer 最终 answer
|
||||||
end
|
end
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|||||||
@@ -154,4 +154,4 @@ ALTER TABLE diagnosis_session ADD COLUMN answer LONGTEXT COMMENT 'Agent 返回
|
|||||||
|
|
||||||
- **LLM 观点层**:在 `selfEvaluation` 的 `llm_opinion` 字段叠加 LLM 结构化观点(has_root_cause、has_solution 等),作为独立 factors,不改变现有规则逻辑
|
- **LLM 观点层**:在 `selfEvaluation` 的 `llm_opinion` 字段叠加 LLM 结构化观点(has_root_cause、has_solution 等),作为独立 factors,不改变现有规则逻辑
|
||||||
- **案例结构化字段**:useful 触发时自动提取 faultCategory / errorCode,替代暂时的 GENERAL
|
- **案例结构化字段**:useful 触发时自动提取 faultCategory / errorCode,替代暂时的 GENERAL
|
||||||
- **重复召回问题**:Executor Prompt 约束或工具层 session 维度去重(见 [ISS-001](../issues/ISS-001-duplicate-retrieval.md))
|
- **重复召回问题**:Executor Prompt 约束或工具层 session 维度去重(见 [ISS-001](../../../issues/archived/ISS-001-duplicate-retrieval.md))
|
||||||
|
|||||||
@@ -65,7 +65,8 @@ flowchart TB
|
|||||||
end
|
end
|
||||||
|
|
||||||
subgraph Store["Persistence and Trace"]
|
subgraph Store["Persistence and Trace"]
|
||||||
Session["diagnosis_session"]
|
ChatSession["chat_session"]
|
||||||
|
Run["diagnosis_run"]
|
||||||
Step["agent_step"]
|
Step["agent_step"]
|
||||||
Invocation["tool_invocation"]
|
Invocation["tool_invocation"]
|
||||||
ApiDoc["api_document"]
|
ApiDoc["api_document"]
|
||||||
@@ -130,9 +131,10 @@ RAG Retrieval
|
|||||||
-> Milvus SDK fallback
|
-> Milvus SDK fallback
|
||||||
|
|
||||||
Persistence
|
Persistence
|
||||||
-> diagnosis_session
|
-> chat_session
|
||||||
-> agent_step
|
-> diagnosis_run
|
||||||
-> tool_invocation
|
-> agent_step.run_id
|
||||||
|
-> tool_invocation.run_id
|
||||||
-> api_document
|
-> api_document
|
||||||
-> Milvus/Zilliz collection
|
-> Milvus/Zilliz collection
|
||||||
|
|
||||||
@@ -163,21 +165,22 @@ sequenceDiagram
|
|||||||
|
|
||||||
User->>API: 提交诊断问题
|
User->>API: 提交诊断问题
|
||||||
API->>Chat: execute chat strategy
|
API->>Chat: execute chat strategy
|
||||||
|
Chat->>DB: 创建 chat_session metadata + diagnosis_run(runId)
|
||||||
Chat->>Planner: 复杂问题进入规划
|
Chat->>Planner: 复杂问题进入规划
|
||||||
Planner->>DB: 写入 agent_step
|
Planner->>DB: 写入 agent_step.run_id
|
||||||
Planner->>Executor: 下发排查方向
|
Planner->>Executor: 下发排查方向
|
||||||
Executor->>Tool: lookup_knowledge / logs / metrics
|
Executor->>Tool: lookup_knowledge / logs / metrics
|
||||||
Tool->>DB: 写入 tool_invocation
|
Tool->>DB: 写入 tool_invocation.run_id
|
||||||
Tool-->>Executor: 返回证据
|
Tool-->>Executor: 返回证据
|
||||||
Executor->>Gatekeeper: 输出 executor_evidence_v2
|
Executor->>Gatekeeper: 输出 executor_evidence_v2
|
||||||
Gatekeeper->>DB: 读取 tool_invocation.evidence_refs 并校验引用
|
Gatekeeper->>DB: 读取 tool_invocation.evidence_refs 并校验引用
|
||||||
Gatekeeper->>Verifier: 传入已验真的 claims / excerpts
|
Gatekeeper->>Verifier: 传入已验真的 claims / excerpts
|
||||||
Verifier->>DB: 合并 self_evaluation.verifier_evaluation
|
Verifier->>DB: 合并 diagnosis_run.self_evaluation.verifier_evaluation
|
||||||
Verifier->>Composer: 传入 allowed_claims / missing_info / actions
|
Verifier->>Composer: 传入 allowed_claims / missing_info / actions
|
||||||
Composer->>Chat: 生成最终用户答复
|
Composer->>Chat: 生成最终用户答复
|
||||||
Chat->>DB: 保存 diagnosis_session.answer
|
Chat->>DB: 保存 diagnosis_run.answer
|
||||||
User->>Trace: GET /api/diagnosis/{sessionId}/trace
|
User->>Trace: GET /api/diagnosis/{sessionId}/trace?runId=...
|
||||||
Trace->>DB: 聚合 session / step / tool
|
Trace->>DB: 聚合 run / step / tool
|
||||||
Trace-->>User: 返回可回放诊断链路
|
Trace-->>User: 返回可回放诊断链路
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -194,13 +197,14 @@ POST /api/chat
|
|||||||
-> Gatekeeper 校验 Executor 证据引用真实性
|
-> Gatekeeper 校验 Executor 证据引用真实性
|
||||||
-> Verifier 判断 claim 是否能由已核验证据推出
|
-> Verifier 判断 claim 是否能由已核验证据推出
|
||||||
-> Composer 生成最终用户答复
|
-> Composer 生成最终用户答复
|
||||||
-> 保存 diagnosis_session
|
-> 保存 chat_session metadata
|
||||||
-> 保存 agent_step
|
-> 保存 diagnosis_run
|
||||||
-> 保存 tool_invocation
|
-> 保存 agent_step.run_id
|
||||||
-> 合并 self_evaluation.verifier_evaluation
|
-> 保存 tool_invocation.run_id
|
||||||
|
-> 合并 diagnosis_run.self_evaluation.verifier_evaluation
|
||||||
```
|
```
|
||||||
|
|
||||||
Chat 链路的质量门禁由三段组成:Gatekeeper 先做代码级引用验真,Verifier 再做 LLM 可推导性判断,Composer 最后控制对用户的表达边界。Gatekeeper、Verifier、Composer 的输出合并到 `diagnosis_session.self_evaluation.verifier_evaluation`,Trace API 会展示该验证结果。
|
Chat 链路的质量门禁由三段组成:Gatekeeper 先做代码级引用验真,Verifier 再做 LLM 可推导性判断,Composer 最后控制对用户的表达边界。Gatekeeper、Verifier、Composer 的输出合并到当前 `diagnosis_run.self_evaluation.verifier_evaluation`,Trace API 会展示该验证结果。
|
||||||
|
|
||||||
Agent 编排细节见 [agent-orchestration.md](agent-orchestration.md)。
|
Agent 编排细节见 [agent-orchestration.md](agent-orchestration.md)。
|
||||||
|
|
||||||
@@ -255,7 +259,7 @@ POST /api/ai_ops
|
|||||||
-> Prometheus / logs / knowledge tools
|
-> Prometheus / logs / knowledge tools
|
||||||
-> 生成告警分析报告
|
-> 生成告警分析报告
|
||||||
-> AiOpsRuleEvaluationService
|
-> AiOpsRuleEvaluationService
|
||||||
-> 合并 self_evaluation.aiops_rule_evaluation
|
-> 合并 diagnosis_run.self_evaluation.aiops_rule_evaluation
|
||||||
-> Trace API 可查看全链路
|
-> Trace API 可查看全链路
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -317,23 +321,30 @@ RAG 总体设计见 [rag-architecture.md](rag-architecture.md),检索运行细
|
|||||||
|
|
||||||
## 6. 持久化模型
|
## 6. 持久化模型
|
||||||
|
|
||||||
当前诊断持久化以三张表为核心:
|
当前诊断持久化以 session/run/trace 明细为核心:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
diagnosis_session
|
chat_session
|
||||||
-> 一次诊断会话的主记录
|
-> 多轮会话目录和元数据
|
||||||
|
-> session_id / status / message_pair_count
|
||||||
|
|
||||||
|
diagnosis_run
|
||||||
|
-> 一次诊断运行的主记录
|
||||||
|
-> run_id / session_id
|
||||||
-> query / status / agent_flow / answer
|
-> query / status / agent_flow / answer
|
||||||
-> self_evaluation
|
-> self_evaluation
|
||||||
-> step_count / tool_call_count / duration
|
-> step_count / tool_call_count / duration
|
||||||
|
|
||||||
agent_step
|
agent_step
|
||||||
-> Agent 模型调用步骤
|
-> Agent 模型调用步骤
|
||||||
|
-> session_id / run_id
|
||||||
-> step_index / agent_name
|
-> step_index / agent_name
|
||||||
-> model_input / model_output / thought
|
-> model_input / model_output / thought
|
||||||
-> duration / token_count
|
-> duration / token_count
|
||||||
|
|
||||||
tool_invocation
|
tool_invocation
|
||||||
-> 工具调用事实
|
-> 工具调用事实
|
||||||
|
-> session_id / run_id
|
||||||
-> tool_name / input_params / output_preview
|
-> tool_name / input_params / output_preview
|
||||||
-> retrieval_layer / retrieval_details
|
-> retrieval_layer / retrieval_details
|
||||||
-> retrieval_details.evidence_refs
|
-> retrieval_details.evidence_refs
|
||||||
@@ -343,7 +354,8 @@ tool_invocation
|
|||||||
|
|
||||||
说明:
|
说明:
|
||||||
|
|
||||||
- 旧的 `diagnosis_record` 已不是当前主模型,迁移脚本中已经由 `diagnosis_session + agent_step + tool_invocation` 取代。
|
- 旧的 `diagnosis_record` 已不是当前主模型。
|
||||||
|
- `diagnosis_session` 已降级为历史兼容和回滚表,新执行写入 `chat_session + diagnosis_run`。
|
||||||
- `api_document` 仍用于文档元数据管理。
|
- `api_document` 仍用于文档元数据管理。
|
||||||
- 文档向量内容存放在 Milvus/Zilliz collection 中。
|
- 文档向量内容存放在 Milvus/Zilliz collection 中。
|
||||||
|
|
||||||
@@ -353,11 +365,12 @@ tool_invocation
|
|||||||
|
|
||||||
```text
|
```text
|
||||||
GET /api/diagnosis/{sessionId}/trace
|
GET /api/diagnosis/{sessionId}/trace
|
||||||
|
GET /api/diagnosis/{sessionId}/trace?runId=run-...
|
||||||
```
|
```
|
||||||
|
|
||||||
Trace API 聚合:
|
Trace API 聚合:
|
||||||
|
|
||||||
- 会话状态和最终报告。
|
- 会话元数据、运行状态和最终报告。
|
||||||
- Agent step 序列。
|
- Agent step 序列。
|
||||||
- 工具调用和检索细节。
|
- 工具调用和检索细节。
|
||||||
- Chat Gatekeeper / Verifier / Composer 结果。
|
- Chat Gatekeeper / Verifier / Composer 结果。
|
||||||
|
|||||||
+70
-194
@@ -1,31 +1,44 @@
|
|||||||
# 数据模型总览
|
# 数据模型总览
|
||||||
|
|
||||||
**更新日期**:2026-07-08
|
**更新日期**:2026-07-10
|
||||||
**状态**:当前可运行架构
|
**状态**:当前可运行架构
|
||||||
|
|
||||||
## 1. 定位
|
## 1. 定位
|
||||||
|
|
||||||
本文从架构角度说明当前 MVP 的核心数据模型。详细字段仍以 Flyway migration 和 `mvp/tables/` 为准。
|
本文从架构角度说明当前 MVP 的核心数据模型。详细字段以 Flyway migration、实体类和 `mvp/tables/` 为准。
|
||||||
|
|
||||||
核心数据分三组:
|
核心数据分三组:
|
||||||
|
|
||||||
- 诊断 Trace:`diagnosis_session`、`agent_step`、`tool_invocation`
|
- 会话与诊断 Trace:`chat_session`、`diagnosis_run`、`agent_step`、`tool_invocation`
|
||||||
- 知识库:`api_document`、`knowledge_domain`、Milvus/Zilliz metadata
|
- 知识库:`api_document`、`knowledge_domain`、Milvus/Zilliz metadata
|
||||||
- 反馈沉淀:`case_library`
|
- 反馈沉淀:`case_library`
|
||||||
|
|
||||||
|
`diagnosis_session` 仍保留为历史兼容和回滚表,不再是新执行写入的主模型。
|
||||||
|
|
||||||
## 2. 总体关系
|
## 2. 总体关系
|
||||||
|
|
||||||
```mermaid
|
```mermaid
|
||||||
erDiagram
|
erDiagram
|
||||||
diagnosis_session ||--o{ agent_step : has
|
chat_session ||--o{ diagnosis_run : owns
|
||||||
diagnosis_session ||--o{ tool_invocation : has
|
diagnosis_run ||--o{ agent_step : has
|
||||||
diagnosis_session ||--o| case_library : creates_when_useful
|
diagnosis_run ||--o{ tool_invocation : has
|
||||||
|
diagnosis_run ||--o| case_library : creates_when_useful
|
||||||
api_document ||--o{ milvus_chunk : indexed_as
|
api_document ||--o{ milvus_chunk : indexed_as
|
||||||
knowledge_domain ||--o{ api_document : groups
|
knowledge_domain ||--o{ api_document : groups
|
||||||
|
|
||||||
diagnosis_session {
|
chat_session {
|
||||||
bigint id
|
bigint id
|
||||||
varchar session_id
|
varchar session_id
|
||||||
|
varchar status
|
||||||
|
int message_pair_count
|
||||||
|
datetime last_active_at
|
||||||
|
datetime expires_at
|
||||||
|
}
|
||||||
|
|
||||||
|
diagnosis_run {
|
||||||
|
bigint id
|
||||||
|
varchar run_id
|
||||||
|
varchar session_id
|
||||||
text query
|
text query
|
||||||
varchar status
|
varchar status
|
||||||
varchar agent_flow
|
varchar agent_flow
|
||||||
@@ -37,42 +50,23 @@ erDiagram
|
|||||||
agent_step {
|
agent_step {
|
||||||
bigint id
|
bigint id
|
||||||
varchar session_id
|
varchar session_id
|
||||||
|
varchar run_id
|
||||||
int step_index
|
int step_index
|
||||||
varchar agent_name
|
varchar agent_name
|
||||||
text model_input
|
text model_input
|
||||||
text model_output
|
text model_output
|
||||||
text thought
|
|
||||||
boolean has_tool_call
|
boolean has_tool_call
|
||||||
}
|
}
|
||||||
|
|
||||||
tool_invocation {
|
tool_invocation {
|
||||||
bigint id
|
bigint id
|
||||||
varchar session_id
|
varchar session_id
|
||||||
|
varchar run_id
|
||||||
|
bigint step_id
|
||||||
varchar tool_name
|
varchar tool_name
|
||||||
json input_params
|
json input_params
|
||||||
text output_preview
|
text output_preview
|
||||||
varchar retrieval_layer
|
|
||||||
json retrieval_details
|
json retrieval_details
|
||||||
varchar relevance_level
|
|
||||||
varchar dedup_reason
|
|
||||||
}
|
|
||||||
|
|
||||||
api_document {
|
|
||||||
bigint id
|
|
||||||
varchar doc_id
|
|
||||||
varchar file_name
|
|
||||||
varchar file_path
|
|
||||||
varchar status
|
|
||||||
int chunk_count
|
|
||||||
text metadata
|
|
||||||
}
|
|
||||||
|
|
||||||
knowledge_domain {
|
|
||||||
bigint id
|
|
||||||
varchar domain_id
|
|
||||||
varchar description
|
|
||||||
text when_to_retrieve
|
|
||||||
int document_count
|
|
||||||
}
|
}
|
||||||
|
|
||||||
case_library {
|
case_library {
|
||||||
@@ -84,58 +78,52 @@ erDiagram
|
|||||||
text root_cause
|
text root_cause
|
||||||
text solution
|
text solution
|
||||||
}
|
}
|
||||||
|
|
||||||
milvus_chunk {
|
|
||||||
varchar id
|
|
||||||
text content
|
|
||||||
json metadata
|
|
||||||
vector vector
|
|
||||||
}
|
|
||||||
```
|
```
|
||||||
|
|
||||||
说明:Milvus/Zilliz collection 不是 MySQL 表,图中的 `milvus_chunk` 是逻辑模型。
|
说明:Milvus/Zilliz collection 不是 MySQL 表,图中的 `milvus_chunk` 是逻辑模型。
|
||||||
|
|
||||||
## 3. 诊断 Trace 模型
|
## 3. 会话与运行模型
|
||||||
|
|
||||||
### diagnosis_session
|
### chat_session
|
||||||
|
|
||||||
会话级主记录。
|
`chat_session` 是会话目录表,保存 `sessionId` 的元数据:
|
||||||
|
|
||||||
关键字段:
|
|
||||||
|
|
||||||
| 字段 | 说明 |
|
| 字段 | 说明 |
|
||||||
|---|---|
|
|---|---|
|
||||||
| `session_id` | 外部关联键,Trace 和 Feedback 都使用它 |
|
| `session_id` | 外部会话 ID,用于多轮上下文和 run 列表 |
|
||||||
| `query` | 用户原始问题或 AIOps 输入摘要 |
|
| `status` | 会话目录状态 |
|
||||||
| `status` | 执行状态 |
|
| `message_pair_count` | Redis 对话轮次数快照 |
|
||||||
|
| `last_active_at` | 最近活跃时间 |
|
||||||
|
| `expires_at` | 可为空的目录 TTL 元数据 |
|
||||||
|
|
||||||
|
它不保存完整对话历史,正文消息仍由 Redis `SessionContext.messageHistory` 管理。
|
||||||
|
|
||||||
|
### diagnosis_run
|
||||||
|
|
||||||
|
`diagnosis_run` 是一次可回放诊断执行的主记录:
|
||||||
|
|
||||||
|
| 字段 | 说明 |
|
||||||
|
|---|---|
|
||||||
|
| `run_id` | 运行 ID,格式为 `run-` + UUID |
|
||||||
|
| `session_id` | 所属 `chat_session.session_id` |
|
||||||
|
| `query` | 本次 Chat 问题或 AIOps 告警摘要 |
|
||||||
|
| `status` | 本次执行状态 |
|
||||||
| `agent_flow` | `CHAT` / `AI_OPS` |
|
| `agent_flow` | `CHAT` / `AI_OPS` |
|
||||||
| `answer` | 最终答复或告警报告 |
|
| `answer` | 本次运行最终答复或告警报告 |
|
||||||
| `self_evaluation` | rule/verifier/aiops 自评估容器 |
|
| `self_evaluation` | 本次运行的 rule/verifier/aiops 自评估容器 |
|
||||||
| `feedback` | 用户反馈 |
|
| `feedback` | 本次运行的用户反馈 |
|
||||||
|
|
||||||
|
同一个 `sessionId` 可以有多个 `runId`。Trace、反馈、评测和案例沉淀都应优先使用 `runId`,避免多轮同 session 下的数据混合。
|
||||||
|
|
||||||
|
## 4. Trace 明细模型
|
||||||
|
|
||||||
### agent_step
|
### agent_step
|
||||||
|
|
||||||
记录模型调用步骤。
|
`agent_step` 记录模型调用步骤。新写入同时保留 `session_id` 和 `run_id`,其中 `run_id` 是回放边界。Trace 页面和评测应先按 `run_id` 隔离取数,展示顺序以 Trace API 返回顺序为准。
|
||||||
|
|
||||||
用途:
|
|
||||||
|
|
||||||
- 回放 Agent 推理过程。
|
|
||||||
- 查看 Planner / Executor / Verifier 的输入输出摘要。
|
|
||||||
- 统计 step count、duration、token count。
|
|
||||||
|
|
||||||
### tool_invocation
|
### tool_invocation
|
||||||
|
|
||||||
记录工具调用事实。
|
`tool_invocation` 记录显式工具调用事实。`retrieval_details.evidence_refs` 是 Chat 证据链路的关键字段:
|
||||||
|
|
||||||
用途:
|
|
||||||
|
|
||||||
- 给 Trace API 展示证据。
|
|
||||||
- 给 Gatekeeper 提供 `retrieval_details.evidence_refs` 引用验真源。
|
|
||||||
- 给 Verifier 构造 `tool_trace_summary` 审计导航。
|
|
||||||
- 给 `EvaluationService` 计算 evidence score。
|
|
||||||
- 给 RAG eval 和人工排查提供检索细节。
|
|
||||||
|
|
||||||
`retrieval_details.evidence_refs` 是当前 Chat 证据链路的关键字段:
|
|
||||||
|
|
||||||
```json
|
```json
|
||||||
{
|
{
|
||||||
@@ -149,83 +137,28 @@ erDiagram
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
字段边界:
|
|
||||||
|
|
||||||
| 字段 | 说明 |
|
|
||||||
|---|---|
|
|
||||||
| `evidence_status` | 工具证据状态,例如 `supported`、`no_evidence`、`deduped`、`failed` |
|
|
||||||
| `evidence_refs[].raw_path` | Executor 可引用的稳定路径,例如 `$.logs[0]`、`$.alerts[0]`、`$.evidence_blocks[0]`、`$.no_evidence` |
|
|
||||||
| `evidence_refs[].text` | 系统抽取的最小证据文本,Gatekeeper 用它核对 `evidence_excerpt` |
|
|
||||||
|
|
||||||
`$.no_evidence` 只表示“本次工具查询未检索到匹配证据”,不能被解释为“问题不存在”或“根因已排除”。
|
`$.no_evidence` 只表示“本次工具查询未检索到匹配证据”,不能被解释为“问题不存在”或“根因已排除”。
|
||||||
|
|
||||||
## 4. 知识库模型
|
|
||||||
|
|
||||||
### api_document
|
|
||||||
|
|
||||||
MySQL 中的文档元数据表。
|
|
||||||
|
|
||||||
职责:
|
|
||||||
|
|
||||||
- 管理上传文件。
|
|
||||||
- 保存 file hash,用于去重。
|
|
||||||
- 记录索引状态和 chunk 数量。
|
|
||||||
- 保存 frontmatter JSON。
|
|
||||||
|
|
||||||
### knowledge_domain
|
|
||||||
|
|
||||||
领域级元数据。
|
|
||||||
|
|
||||||
职责:
|
|
||||||
|
|
||||||
- 按 category 聚合文档。
|
|
||||||
- 存储领域描述。
|
|
||||||
- 存储 `when_to_retrieve`,辅助 Planner/Executor 判断什么时候检索该领域。
|
|
||||||
|
|
||||||
### Milvus/Zilliz metadata
|
|
||||||
|
|
||||||
向量 collection 中每个 chunk 的 metadata 主要包括:
|
|
||||||
|
|
||||||
```text
|
|
||||||
docId
|
|
||||||
_source
|
|
||||||
chunkIndex
|
|
||||||
totalChunks
|
|
||||||
title
|
|
||||||
breadcrumb
|
|
||||||
category
|
|
||||||
```
|
|
||||||
|
|
||||||
这些字段支撑:
|
|
||||||
|
|
||||||
- category filter。
|
|
||||||
- source 展示。
|
|
||||||
- breadcrumb 上下文。
|
|
||||||
- docId 删除和重建索引。
|
|
||||||
- evidence block 构造。
|
|
||||||
|
|
||||||
## 5. 反馈沉淀模型
|
## 5. 反馈沉淀模型
|
||||||
|
|
||||||
### case_library
|
`useful` 反馈会触发 `CaseLibraryService.createFromRun`。
|
||||||
|
|
||||||
`useful` 反馈会触发 `CaseLibraryService.createFromSession`。
|
|
||||||
|
|
||||||
当前自动映射:
|
当前自动映射:
|
||||||
|
|
||||||
| 字段 | 来源 |
|
| 字段 | 来源 |
|
||||||
|---|---|
|
|---|---|
|
||||||
| `case_id` | UUID |
|
| `case_id` | UUID |
|
||||||
| `diagnosis_id` | `diagnosis_session.session_id` |
|
| `diagnosis_id` | 新数据为 `diagnosis_run.run_id`;历史数据可能为 `diagnosis_session.session_id` |
|
||||||
| `source_type` | `AUTO` |
|
| `source_type` | `AUTO` |
|
||||||
| `fault_category` | 当前默认 `GENERAL` |
|
| `fault_category` | 当前默认 `GENERAL` |
|
||||||
| `title` | session query 前 100 字符 |
|
| `title` | run query 前 100 字符 |
|
||||||
| `root_cause` | session answer |
|
| `root_cause` | run answer |
|
||||||
| `solution` | session answer |
|
| `solution` | run answer |
|
||||||
| `created_by` | `system` |
|
| `created_by` | `system` |
|
||||||
|
|
||||||
## 6. self_evaluation 结构
|
## 6. self_evaluation 结构
|
||||||
|
|
||||||
`diagnosis_session.self_evaluation` 是 JSON 容器:
|
`diagnosis_run.self_evaluation` 是运行级 JSON 容器:
|
||||||
|
|
||||||
```json
|
```json
|
||||||
{
|
{
|
||||||
@@ -235,78 +168,21 @@ category
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
边界:
|
Chat 通常写入 `rule_evaluation` 和 `verifier_evaluation`;AIOps 写入 `aiops_rule_evaluation`。
|
||||||
|
|
||||||
- `rule_evaluation` 评估证据收集充分度。
|
## 7. 当前边界和后续
|
||||||
- `verifier_evaluation` 评估 Chat 结构化 claims 是否能由已验真证据推出,并保存 Gatekeeper、Verifier、Composer 的审计数据。
|
|
||||||
- `aiops_rule_evaluation` 评估 AIOps 报告是否聚焦告警并使用证据。
|
|
||||||
|
|
||||||
当前 `verifier_evaluation` 关键结构:
|
|
||||||
|
|
||||||
```json
|
|
||||||
{
|
|
||||||
"verdict": "PASS",
|
|
||||||
"groundedness_score": 1.0,
|
|
||||||
"critical_fact_count": 1,
|
|
||||||
"claim_checks": [],
|
|
||||||
"facts_checked": [],
|
|
||||||
"rationale": "...",
|
|
||||||
"round": 1,
|
|
||||||
"traceability_version": "v1",
|
|
||||||
"executor_output_parse_status": {},
|
|
||||||
"executor_structured_output": {},
|
|
||||||
"gatekeeper_result": {},
|
|
||||||
"composer_output": {},
|
|
||||||
"tool_trace_summary": []
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
必要审计字段:
|
|
||||||
|
|
||||||
| 字段 | 说明 |
|
|
||||||
|---|---|
|
|
||||||
| `executor_output_parse_status` | Executor 输出是否能解析为 `executor_evidence_v2` |
|
|
||||||
| `executor_structured_output` | Executor 结构化 claims、hypotheses、recommended_actions、missing_info |
|
|
||||||
| `gatekeeper_result` | 引用真实性校验结果,包括 rule set version、checked bindings、failed rules、warnings、errors |
|
|
||||||
| `composer_output` | Composer 最终表达及解析状态 |
|
|
||||||
| `tool_trace_summary` | Verifier 调用时使用的工具调用导航索引,不是唯一证据源 |
|
|
||||||
|
|
||||||
## 7. 数据写入时序
|
|
||||||
|
|
||||||
```mermaid
|
|
||||||
sequenceDiagram
|
|
||||||
autonumber
|
|
||||||
participant API as API
|
|
||||||
participant Svc as ChatService/AiOpsService
|
|
||||||
participant Session as diagnosis_session
|
|
||||||
participant Agent as Agent
|
|
||||||
participant Step as agent_step
|
|
||||||
participant Tool as tool_invocation
|
|
||||||
participant Eval as self_evaluation
|
|
||||||
participant Feedback as case_library
|
|
||||||
|
|
||||||
API->>Svc: request
|
|
||||||
Svc->>Session: create/update RUNNING
|
|
||||||
Agent->>Step: before/after model
|
|
||||||
Agent->>Tool: tool call record
|
|
||||||
Svc->>Session: SUCCESS/FAILED + answer
|
|
||||||
Svc->>Eval: merge evaluation
|
|
||||||
API->>Svc: feedback useful
|
|
||||||
Svc->>Feedback: create case
|
|
||||||
```
|
|
||||||
|
|
||||||
## 8. 当前边界和后续
|
|
||||||
|
|
||||||
当前边界:
|
当前边界:
|
||||||
|
|
||||||
- `agent_step.session_id` 和 `tool_invocation.session_id` 通过 sessionId 关联,不强制外键。
|
- `chat_session` 只存会话元数据,不存完整正文历史。
|
||||||
- `tool_invocation.step_id` 可为空。
|
- `diagnosis_run` 存一次运行的长期审计状态。
|
||||||
- Milvus chunk 与 `api_document` 通过 metadata.docId 逻辑关联。
|
- `agent_step.run_id` 和 `tool_invocation.run_id` 是 Trace、Verifier、Eval 的运行边界。
|
||||||
- `case_library` 与 session 通过 `diagnosis_id=session_id` 关联。
|
- 当前实现主要使用逻辑关联,不依赖数据库外键。
|
||||||
|
- `case_library.diagnosis_id` 是过渡字段,新值按 `run_id` 解释,旧值可能按 `session_id` 解释。
|
||||||
|
- `diagnosis_session` 只作为历史兼容和回滚表保留。
|
||||||
|
|
||||||
后续可增强:
|
后续可增强:
|
||||||
|
|
||||||
1. 增加 run id,支持同 session 多次独立诊断。
|
1. 强化 `tool_invocation.step_id` 关联。
|
||||||
2. 强化 `tool_invocation.step_id` 关联。
|
2. 将 Gatekeeper 规则配置化时的规则元数据保存为可审计版本。
|
||||||
3. 将 Gatekeeper 规则配置化时的规则元数据保存为可审计版本。
|
3. 将 `case_library` 的 rootCause/solution 从完整 answer 中结构化抽取。
|
||||||
4. 将 `case_library` 的 rootCause/solution 从完整 answer 中结构化抽取。
|
|
||||||
|
|||||||
@@ -391,7 +391,7 @@ Composer 位于 Verifier 之后,输入是 ChatService 过滤后的允许表达
|
|||||||
|
|
||||||
## 7. Trace Persistence
|
## 7. Trace Persistence
|
||||||
|
|
||||||
`diagnosis_session.self_evaluation.verifier_evaluation` 持久化:
|
`diagnosis_run.self_evaluation.verifier_evaluation` 持久化:
|
||||||
|
|
||||||
```json
|
```json
|
||||||
{
|
{
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
# 反馈与自评估架构
|
# 反馈与自评估架构
|
||||||
|
|
||||||
**更新日期**:2026-07-08
|
**更新日期**:2026-07-10
|
||||||
**状态**:当前可运行架构
|
**状态**:当前可运行架构
|
||||||
**参考历史文档**:`archive/2026-07-05-legacy/confidence-feedback.md`
|
**参考历史文档**:`archive/2026-07-05-legacy/confidence-feedback.md`
|
||||||
|
|
||||||
@@ -8,8 +8,8 @@
|
|||||||
|
|
||||||
反馈架构包含两条闭环:
|
反馈架构包含两条闭环:
|
||||||
|
|
||||||
1. 系统自评估:基于工具调用、Gatekeeper、Verifier、Composer、AIOps 规则检查,写入 `diagnosis_session.self_evaluation`。
|
1. 系统自评估:基于当前 run 的工具调用、Gatekeeper、Verifier、Composer、AIOps 规则检查,写入 `diagnosis_run.self_evaluation`。
|
||||||
2. 用户反馈:用户标记 `useful` 或 `not_useful`,写入 `diagnosis_session.feedback`,其中 `useful` 会沉淀案例。
|
2. 用户反馈:用户标记 `useful` 或 `not_useful`,优先写入 `diagnosis_run.feedback`,其中 `useful` 会沉淀案例。
|
||||||
|
|
||||||
当前重要边界:
|
当前重要边界:
|
||||||
|
|
||||||
@@ -21,7 +21,7 @@
|
|||||||
|
|
||||||
```mermaid
|
```mermaid
|
||||||
flowchart TD
|
flowchart TD
|
||||||
Answer["Chat / AIOps final answer"] --> Session["diagnosis_session.answer"]
|
Answer["Chat / AIOps final answer"] --> Run["diagnosis_run.answer"]
|
||||||
|
|
||||||
subgraph SelfEval["Self evaluation"]
|
subgraph SelfEval["Self evaluation"]
|
||||||
Invocation["tool_invocation"] --> RuleEval["EvaluationService: rule_evaluation"]
|
Invocation["tool_invocation"] --> RuleEval["EvaluationService: rule_evaluation"]
|
||||||
@@ -40,24 +40,24 @@ flowchart TD
|
|||||||
RuleEval --> Merge["SelfEvaluationMergeService"]
|
RuleEval --> Merge["SelfEvaluationMergeService"]
|
||||||
VerifierEval --> Merge
|
VerifierEval --> Merge
|
||||||
AiOpsEval --> Merge
|
AiOpsEval --> Merge
|
||||||
Merge --> SelfJson["diagnosis_session.self_evaluation"]
|
Merge --> SelfJson["diagnosis_run.self_evaluation"]
|
||||||
|
|
||||||
subgraph UserFeedback["User feedback"]
|
subgraph UserFeedback["User feedback"]
|
||||||
UI["Feedback bar"] --> API["POST /api/feedback"]
|
UI["Feedback bar"] --> API["POST /api/feedback"]
|
||||||
API --> FeedbackService["FeedbackService"]
|
API --> FeedbackService["FeedbackService"]
|
||||||
FeedbackService --> FeedbackField["diagnosis_session.feedback"]
|
FeedbackService --> FeedbackField["diagnosis_run.feedback"]
|
||||||
FeedbackService --> Useful{"feedback == useful?"}
|
FeedbackService --> Useful{"feedback == useful?"}
|
||||||
Useful -->|yes| CaseService["CaseLibraryService.createFromSession"]
|
Useful -->|yes| CaseService["CaseLibraryService.createFromRun"]
|
||||||
CaseService --> Case["case_library"]
|
CaseService --> Case["case_library"]
|
||||||
Useful -->|no| BadCase["Bad case by feedback=not_useful"]
|
Useful -->|no| BadCase["Bad case by feedback=not_useful"]
|
||||||
end
|
end
|
||||||
|
|
||||||
Session --> UI
|
Run --> UI
|
||||||
```
|
```
|
||||||
|
|
||||||
## 3. self_evaluation JSON
|
## 3. self_evaluation JSON
|
||||||
|
|
||||||
`SelfEvaluationMergeService` 统一维护 `diagnosis_session.self_evaluation`。
|
`SelfEvaluationMergeService` 统一维护当前运行的 `diagnosis_run.self_evaluation`。历史兼容数据可能仍存在于 `diagnosis_session.self_evaluation`,但新 Chat/AIOps 执行不再写旧表。
|
||||||
|
|
||||||
当前结构:
|
当前结构:
|
||||||
|
|
||||||
@@ -149,7 +149,7 @@ flowchart LR
|
|||||||
Composer --> ComposerOutput["composer_output"]
|
Composer --> ComposerOutput["composer_output"]
|
||||||
Output --> Merge["SelfEvaluationMergeService.mergeVerifierEvaluation"]
|
Output --> Merge["SelfEvaluationMergeService.mergeVerifierEvaluation"]
|
||||||
ComposerOutput --> Merge
|
ComposerOutput --> Merge
|
||||||
Merge --> Session["diagnosis_session.self_evaluation.verifier_evaluation"]
|
Merge --> Run["diagnosis_run.self_evaluation.verifier_evaluation"]
|
||||||
```
|
```
|
||||||
|
|
||||||
Verifier 输出:
|
Verifier 输出:
|
||||||
@@ -198,6 +198,7 @@ POST /api/feedback
|
|||||||
Content-Type: application/json
|
Content-Type: application/json
|
||||||
|
|
||||||
{
|
{
|
||||||
|
"runId": "run-xxx",
|
||||||
"sessionId": "xxx",
|
"sessionId": "xxx",
|
||||||
"feedback": "useful" | "not_useful"
|
"feedback": "useful" | "not_useful"
|
||||||
}
|
}
|
||||||
@@ -209,6 +210,8 @@ Content-Type: application/json
|
|||||||
{
|
{
|
||||||
"success": true,
|
"success": true,
|
||||||
"message": "反馈已记录",
|
"message": "反馈已记录",
|
||||||
|
"runId": "run-xxx",
|
||||||
|
"fallbackToLatestRun": false,
|
||||||
"caseId": "uuid 或 null"
|
"caseId": "uuid 或 null"
|
||||||
}
|
}
|
||||||
```
|
```
|
||||||
@@ -217,10 +220,16 @@ Content-Type: application/json
|
|||||||
|
|
||||||
| feedback | 行为 |
|
| feedback | 行为 |
|
||||||
|---|---|
|
|---|---|
|
||||||
| `useful` | 写入 `DiagnosisSession.feedback`,调用 `CaseLibraryService.createFromSession` |
|
| `useful` | 写入 `DiagnosisRun.feedback`,调用 `CaseLibraryService.createFromRun` |
|
||||||
| `not_useful` | 写入 `DiagnosisSession.feedback`,不改变 session status |
|
| `not_useful` | 写入 `DiagnosisRun.feedback`,不改变 run status |
|
||||||
| 其他值 | 返回 HTTP 400 |
|
| 其他值 | 返回 HTTP 400 |
|
||||||
|
|
||||||
|
兼容行为:
|
||||||
|
|
||||||
|
- 请求带 `runId` 时,后端验证 `runId` 属于 `sessionId`。
|
||||||
|
- 请求缺少 `runId` 且存在 run-backed 数据时,后端绑定 latest run,并返回 `fallbackToLatestRun=true` 和实际 `runId`。
|
||||||
|
- 仅当没有 `diagnosis_run` 但存在历史 `diagnosis_session` 时,才使用历史 fallback;该路径不声明 latest-run fallback。
|
||||||
|
|
||||||
## 8. 案例沉淀
|
## 8. 案例沉淀
|
||||||
|
|
||||||
`useful` 反馈会生成或复用 `case_library` 记录。
|
`useful` 反馈会生成或复用 `case_library` 记录。
|
||||||
@@ -230,7 +239,7 @@ Content-Type: application/json
|
|||||||
| CaseLibrary 字段 | 来源 |
|
| CaseLibrary 字段 | 来源 |
|
||||||
|---|---|
|
|---|---|
|
||||||
| `caseId` | UUID |
|
| `caseId` | UUID |
|
||||||
| `diagnosisId` | `DiagnosisSession.sessionId` |
|
| `diagnosisId` | 新数据为 `DiagnosisRun.runId`;历史数据可能为 `DiagnosisSession.sessionId` |
|
||||||
| `sourceType` | `AUTO` |
|
| `sourceType` | `AUTO` |
|
||||||
| `faultCategory` | 当前固定为 `GENERAL` |
|
| `faultCategory` | 当前固定为 `GENERAL` |
|
||||||
| `title` | `query` 前 100 字符 |
|
| `title` | `query` 前 100 字符 |
|
||||||
@@ -241,7 +250,7 @@ Content-Type: application/json
|
|||||||
幂等性:
|
幂等性:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
case_library.diagnosisId == sessionId
|
case_library.diagnosisId == runId
|
||||||
-> existing case: return existing
|
-> existing case: return existing
|
||||||
-> missing case: create new
|
-> missing case: create new
|
||||||
```
|
```
|
||||||
@@ -260,9 +269,9 @@ Trace API 会展示:
|
|||||||
|
|
||||||
| 视角 | 数据来源 |
|
| 视角 | 数据来源 |
|
||||||
|---|---|
|
|---|---|
|
||||||
| 执行是否成功 | `diagnosis_session.status` |
|
| 执行是否成功 | `diagnosis_run.status` |
|
||||||
| 证据是否充分 | `self_evaluation.rule_evaluation` / `verifier_evaluation` |
|
| 证据是否充分 | `self_evaluation.rule_evaluation` / `verifier_evaluation` |
|
||||||
| 用户是否认可 | `diagnosis_session.feedback` |
|
| 用户是否认可 | `diagnosis_run.feedback` |
|
||||||
|
|
||||||
## 10. 后续增强
|
## 10. 后续增强
|
||||||
|
|
||||||
|
|||||||
@@ -36,7 +36,7 @@ flowchart TB
|
|||||||
Tools --> Invocation["tool_invocation"]
|
Tools --> Invocation["tool_invocation"]
|
||||||
Agent --> StepHook["AgentLoggingHook"]
|
Agent --> StepHook["AgentLoggingHook"]
|
||||||
StepHook --> Step["agent_step"]
|
StepHook --> Step["agent_step"]
|
||||||
Agent --> Session["diagnosis_session"]
|
Agent --> Run["diagnosis_run"]
|
||||||
|
|
||||||
Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
|
Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
|
||||||
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
|
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
|
||||||
@@ -49,7 +49,7 @@ flowchart TB
|
|||||||
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
|
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
|
||||||
AiOpsRule --> AiOpsEval["self_evaluation.aiops_rule_evaluation"]
|
AiOpsRule --> AiOpsEval["self_evaluation.aiops_rule_evaluation"]
|
||||||
|
|
||||||
Session --> TraceAPI["DiagnosisTraceService"]
|
Run --> TraceAPI["DiagnosisTraceService"]
|
||||||
Step --> TraceAPI
|
Step --> TraceAPI
|
||||||
Invocation --> TraceAPI
|
Invocation --> TraceAPI
|
||||||
SelfEval --> TraceAPI
|
SelfEval --> TraceAPI
|
||||||
@@ -82,6 +82,23 @@ Prompt 层当前承担的门禁:
|
|||||||
- Chat Composer 不允许补事实,尤其不能把 `$.no_evidence` 表达为“已排除/确认没有”。
|
- Chat Composer 不允许补事实,尤其不能把 `$.no_evidence` 表达为“已排除/确认没有”。
|
||||||
- AIOps payload 模式必须聚焦输入告警。
|
- AIOps payload 模式必须聚焦输入告警。
|
||||||
|
|
||||||
|
Chat 链路还会在 `verifier_evaluation.prompt_audit` 中持久化紧凑 Prompt 审计快照:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"version": "chat-prompts-v1",
|
||||||
|
"prompts": [
|
||||||
|
{
|
||||||
|
"name": "chat_executor",
|
||||||
|
"version": "chat-executor-v2",
|
||||||
|
"resource": "prompts/chat-executor-prompt.md"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
该快照只保存版本和资源路径,不保存完整 Prompt 文本。它用于面试演示、trace 回放和离线 baseline 解释“本次诊断使用了哪套 Prompt 契约”。
|
||||||
|
|
||||||
## 4. Trace Hooks
|
## 4. Trace Hooks
|
||||||
|
|
||||||
`AgentLoggingHook` 是当前 Agent step 可观测性的核心。
|
`AgentLoggingHook` 是当前 Agent step 可观测性的核心。
|
||||||
@@ -93,10 +110,10 @@ sequenceDiagram
|
|||||||
participant H as AgentLoggingHook
|
participant H as AgentLoggingHook
|
||||||
participant DB as agent_step
|
participant DB as agent_step
|
||||||
|
|
||||||
A->>H: before_model(messages, sessionId)
|
A->>H: before_model(messages, sessionId, runId)
|
||||||
H->>DB: 写入 model_input / step_index / agent_name
|
H->>DB: 写入 session_id / run_id / model_input / step_index / agent_name
|
||||||
A-->>A: LLM 推理
|
A-->>A: LLM 推理
|
||||||
A->>H: after_model(messages, sessionId)
|
A->>H: after_model(messages, sessionId, runId)
|
||||||
H->>DB: 回填 model_output / thought / has_tool_call / duration / token_count
|
H->>DB: 回填 model_output / thought / has_tool_call / duration / token_count
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -109,6 +126,8 @@ sequenceDiagram
|
|||||||
- token count。
|
- token count。
|
||||||
- Verifier 的 JSON 输出摘要。
|
- Verifier 的 JSON 输出摘要。
|
||||||
|
|
||||||
|
新写入必须带 `run_id`;`session_id` 仍保留用于粗粒度排查和历史兼容。
|
||||||
|
|
||||||
## 5. Tool Invocation 门禁
|
## 5. Tool Invocation 门禁
|
||||||
|
|
||||||
工具调用记录由 `ToolInvocationRecorder` 和具体工具共同完成。
|
工具调用记录由 `ToolInvocationRecorder` 和具体工具共同完成。
|
||||||
@@ -193,10 +212,10 @@ Verifier 不再逐字核验 excerpt 真伪;这由 Gatekeeper 完成。Verifier
|
|||||||
结果写入:
|
结果写入:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
diagnosis_session.self_evaluation.verifier_evaluation
|
diagnosis_run.self_evaluation.verifier_evaluation
|
||||||
```
|
```
|
||||||
|
|
||||||
其中同时持久化 `executor_structured_output`、`gatekeeper_result`、`tool_trace_summary` 和 `composer_output`,用于 Trace 回放。
|
其中同时持久化 `executor_structured_output`、`gatekeeper_result`、`tool_trace_summary`、`prompt_audit` 和 `composer_output`,用于 Trace 回放。
|
||||||
|
|
||||||
## 7. AIOps 规则门禁
|
## 7. AIOps 规则门禁
|
||||||
|
|
||||||
@@ -212,7 +231,7 @@ AIOps 当前不走 Chat Verifier,而是用 `AiOpsRuleEvaluationService` 做轻
|
|||||||
结果写入:
|
结果写入:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
diagnosis_session.self_evaluation.aiops_rule_evaluation
|
diagnosis_run.self_evaluation.aiops_rule_evaluation
|
||||||
```
|
```
|
||||||
|
|
||||||
## 8. Eval Baseline
|
## 8. Eval Baseline
|
||||||
@@ -233,7 +252,7 @@ diagnosis_session.self_evaluation.aiops_rule_evaluation
|
|||||||
- 同一工具调用次数上限。
|
- 同一工具调用次数上限。
|
||||||
- 工具超时的统一熔断。
|
- 工具超时的统一熔断。
|
||||||
- Gatekeeper 规则远程化或三层分离:索引层、元数据层、规则实现层。
|
- Gatekeeper 规则远程化或三层分离:索引层、元数据层、规则实现层。
|
||||||
- Prompt 版本记录和回滚。
|
- Prompt 版本回滚和更细粒度变更审计。
|
||||||
- Verifier 对 AIOps 报告的 LLM 级事实校验。
|
- Verifier 对 AIOps 报告的 LLM 级事实校验。
|
||||||
|
|
||||||
这些应在评测集扩大后逐步加入,避免一次性把诊断流程卡得过死。
|
这些应在评测集扩大后逐步加入,避免一次性把诊断流程卡得过死。
|
||||||
|
|||||||
@@ -5,7 +5,7 @@
|
|||||||
|
|
||||||
## 1. 一句话
|
## 1. 一句话
|
||||||
|
|
||||||
SuperBizAgent 是一个面向企业故障诊断的可追踪 Agent 系统:它把用户问题或 AIOps 告警转换成 Planner、Executor、Gatekeeper、Verifier、Composer 的诊断链路,所有工具证据、模型步骤、最终答案、自评估和用户反馈都能通过同一个 `sessionId` 回放。
|
SuperBizAgent 是一个面向企业故障诊断的可追踪 Agent 系统:它把用户问题或 AIOps 告警转换成 Planner、Executor、Gatekeeper、Verifier、Composer 的诊断链路,`sessionId` 保留多轮上下文,`runId` 精确绑定一次诊断运行;所有工具证据、模型步骤、最终答案、自评估和用户反馈都能按 `sessionId + runId` 回放。
|
||||||
|
|
||||||
## 2. 一张图
|
## 2. 一张图
|
||||||
|
|
||||||
@@ -34,14 +34,15 @@ flowchart TB
|
|||||||
AiOpsFlow --> Trace
|
AiOpsFlow --> Trace
|
||||||
Tools --> Trace
|
Tools --> Trace
|
||||||
|
|
||||||
Trace --> Session["diagnosis_session"]
|
Trace --> ChatSession["chat_session"]
|
||||||
|
Trace --> Run["diagnosis_run"]
|
||||||
Trace --> Step["agent_step"]
|
Trace --> Step["agent_step"]
|
||||||
Trace --> Invocation["tool_invocation"]
|
Trace --> Invocation["tool_invocation"]
|
||||||
|
|
||||||
Invocation --> Verifier["Verifier / Rule Evaluation"]
|
Invocation --> Verifier["Verifier / Rule Evaluation"]
|
||||||
Verifier --> SelfEval["self_evaluation"]
|
Verifier --> SelfEval["self_evaluation"]
|
||||||
|
|
||||||
Session --> TraceAPI["GET /api/diagnosis/{sessionId}/trace"]
|
Run --> TraceAPI["GET /api/diagnosis/{sessionId}/trace?runId=..."]
|
||||||
Step --> TraceAPI
|
Step --> TraceAPI
|
||||||
Invocation --> TraceAPI
|
Invocation --> TraceAPI
|
||||||
SelfEval --> TraceAPI
|
SelfEval --> TraceAPI
|
||||||
@@ -61,15 +62,15 @@ Planner 负责拆解,Executor 只负责调用知识库、日志和指标工具
|
|||||||
AIOps 告警入口走 Supervisor 调度 Planner/Executor:
|
AIOps 告警入口走 Supervisor 调度 Planner/Executor:
|
||||||
如果请求里有 alert payload,系统会进入 PAYLOAD_TARGETED 模式,报告必须聚焦这个告警,而不是被当前环境中的其他活跃告警带偏。
|
如果请求里有 alert payload,系统会进入 PAYLOAD_TARGETED 模式,报告必须聚焦这个告警,而不是被当前环境中的其他活跃告警带偏。
|
||||||
|
|
||||||
所有过程都会落到 diagnosis_session、agent_step、tool_invocation。
|
会话元数据会落到 chat_session,每次诊断运行会落到 diagnosis_run,步骤和工具明细通过 run_id 关联。
|
||||||
所以我可以用一个 sessionId 回放:模型怎么规划、调了哪些工具、工具返回什么、Gatekeeper 怎么验真、Verifier 怎么判定、Composer 最后怎么表达、用户最后是否反馈有用。
|
所以我可以用 sessionId + runId 精确回放:模型怎么规划、调了哪些工具、工具返回什么、Gatekeeper 怎么验真、Verifier 怎么判定、Composer 最后怎么表达、用户最后是否反馈有用。
|
||||||
```
|
```
|
||||||
|
|
||||||
## 4. 五个亮点
|
## 4. 五个亮点
|
||||||
|
|
||||||
| 亮点 | 怎么讲 |
|
| 亮点 | 怎么讲 |
|
||||||
|---|---|
|
|---|---|
|
||||||
| 可追踪 Agent | 每次诊断都有 `sessionId`,Trace API 可以回放 session、step、tool |
|
| 可追踪 Agent | 每次诊断都有 `runId`,Trace API 可以回放 run、step、tool;同一 `sessionId` 可有多次独立 run |
|
||||||
| 显式工具证据链 | `lookup_knowledge`、日志、指标都记录到 `tool_invocation` |
|
| 显式工具证据链 | `lookup_knowledge`、日志、指标都记录到 `tool_invocation` |
|
||||||
| RAG 工程化 | L0 降级为 hint,Spring AI VectorStore 做主检索,SDK fallback 保底 |
|
| RAG 工程化 | L0 降级为 hint,Spring AI VectorStore 做主检索,SDK fallback 保底 |
|
||||||
| 质量门禁 | Chat Gatekeeper 验引用、Verifier 判可推导、Composer 控表达,AIOps rule evaluation 控制告警聚焦 |
|
| 质量门禁 | Chat Gatekeeper 验引用、Verifier 判可推导、Composer 控表达,AIOps rule evaluation 控制告警聚焦 |
|
||||||
|
|||||||
@@ -2,7 +2,7 @@
|
|||||||
|
|
||||||
**更新日期**:2026-07-06
|
**更新日期**:2026-07-06
|
||||||
**状态**:当前主架构 + 后续演进边界
|
**状态**:当前主架构 + 后续演进边界
|
||||||
**关联计划**:`mvp/issues/rag-refactor-plan.md`
|
**关联计划**:[`mvp/issues/active/rag-refactor-plan.md`](../issues/active/rag-refactor-plan.md)
|
||||||
|
|
||||||
## 1. 架构目标
|
## 1. 架构目标
|
||||||
|
|
||||||
|
|||||||
@@ -1,69 +1,68 @@
|
|||||||
# 会话与 Trace 生命周期
|
# 会话与 Trace 生命周期
|
||||||
|
|
||||||
**更新日期**:2026-07-08
|
**更新日期**:2026-07-10
|
||||||
**状态**:当前可运行架构
|
**状态**:当前可运行架构
|
||||||
**参考历史文档**:`archive/2026-07-05-legacy/session-management.md`
|
**参考历史文档**:`archive/2026-07-05-legacy/session-management.md`
|
||||||
|
|
||||||
## 1. 定位
|
## 1. 定位
|
||||||
|
|
||||||
旧版会话设计以 Redis 会话为主,MySQL 作为可选长期沉淀。当前 MVP 的可追踪诊断已经转为 MySQL Trace 三表为主:
|
当前 MVP 把“会话态”和“运行态”拆开:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
diagnosis_session
|
chat_session(sessionId)
|
||||||
-> agent_step
|
-> diagnosis_run(runId)
|
||||||
-> tool_invocation
|
-> agent_step(runId)
|
||||||
|
-> tool_invocation(runId)
|
||||||
```
|
```
|
||||||
|
|
||||||
因此本文描述的是当前可运行链路:
|
- `sessionId` 表示多轮会话目录和 Redis 上下文。
|
||||||
|
- `runId` 表示一次可回放诊断执行。
|
||||||
- `sessionId` 是一次诊断和后续 trace/feedback 的关联键。
|
- `DiagnosisTraceService` 聚合一个 run 的主记录、步骤和工具调用,形成可回放 Trace。
|
||||||
- `diagnosis_session` 保存会话级状态、问题、答案、自评估和反馈。
|
- `diagnosis_session` 只保留为历史兼容和回滚表。
|
||||||
- `agent_step` 保存每个 Agent 模型调用。
|
|
||||||
- `tool_invocation` 保存工具调用事实。
|
|
||||||
- `DiagnosisTraceService` 聚合三类记录,形成可回放 trace。
|
|
||||||
|
|
||||||
## 2. 生命周期总图
|
## 2. 生命周期总图
|
||||||
|
|
||||||
```mermaid
|
```mermaid
|
||||||
flowchart TD
|
flowchart TD
|
||||||
Start["request: chat / ai_ops"] --> Resolve["resolve sessionId"]
|
Start["request: chat / ai_ops"] --> Resolve["resolve sessionId"]
|
||||||
Resolve --> Create["create or reset diagnosis_session"]
|
Resolve --> Session["ensure chat_session metadata"]
|
||||||
Create --> Running["status = RUNNING"]
|
Session --> Run["create diagnosis_run(runId)"]
|
||||||
|
Run --> Running["run.status = RUNNING"]
|
||||||
|
|
||||||
Running --> Agent["Agent workflow"]
|
Running --> Agent["Agent workflow"]
|
||||||
Agent --> StepHook["AgentLoggingHook"]
|
Agent --> Context["execution context(sessionId, runId)"]
|
||||||
StepHook --> Step["agent_step"]
|
Context --> StepHook["AgentLoggingHook"]
|
||||||
Agent --> Tool["Evidence tools"]
|
StepHook --> Step["agent_step(session_id, run_id)"]
|
||||||
Tool --> Invocation["tool_invocation"]
|
Context --> Tool["Evidence tools"]
|
||||||
|
Tool --> Invocation["tool_invocation(session_id, run_id)"]
|
||||||
Invocation --> Gatekeeper["Gatekeeper evidence validation"]
|
Invocation --> Gatekeeper["Gatekeeper evidence validation"]
|
||||||
|
|
||||||
Agent --> Final{"workflow result"}
|
Agent --> Final{"workflow result"}
|
||||||
Final -->|success| Success["status = SUCCESS, answer saved"]
|
Final -->|success| Success["run.status = SUCCESS, answer saved"]
|
||||||
Final -->|failed| Failed["status = FAILED"]
|
Final -->|failed| Failed["run.status = FAILED"]
|
||||||
|
|
||||||
Success --> Evaluation["self_evaluation merge"]
|
Success --> Evaluation["diagnosis_run.self_evaluation merge"]
|
||||||
Failed --> Evaluation
|
Failed --> Evaluation
|
||||||
Evaluation --> Trace["GET /api/diagnosis/{sessionId}/trace"]
|
Evaluation --> Trace["GET /api/diagnosis/{sessionId}/trace?runId=..."]
|
||||||
Success --> Feedback["POST /api/feedback"]
|
Success --> Feedback["POST /api/feedback(sessionId, runId)"]
|
||||||
Feedback --> Case["useful -> case_library"]
|
Feedback --> Case["useful -> case_library(run_id)"]
|
||||||
```
|
```
|
||||||
|
|
||||||
## 3. sessionId 规则
|
## 3. ID 规则
|
||||||
|
|
||||||
| 链路 | sessionId 来源 |
|
| ID | 来源 | 含义 |
|
||||||
|---|---|
|
|---|---|---|
|
||||||
| Chat | 如果请求带 sessionId,则复用;否则生成短 UUID |
|
| `sessionId` | Chat request `Id`、AIOps payload `sessionId`,缺失时由服务生成 | 多轮会话目录和 Redis 上下文 |
|
||||||
| AIOps | 如果 payload 带 sessionId,则复用;否则生成 UUID |
|
| `runId` | 每次有效 Chat/AIOps 执行创建 | 一次诊断运行和 Trace 回放边界 |
|
||||||
| Trace | URL path 中的 `{sessionId}` |
|
|
||||||
| Feedback | request body 中的 `sessionId` |
|
|
||||||
|
|
||||||
设计含义:
|
设计含义:
|
||||||
|
|
||||||
- 同一个 `sessionId` 可以贯穿诊断、trace 查询和用户反馈。
|
- 同一个 `sessionId` 可以贯穿多轮 Chat。
|
||||||
- 当前诊断开始时会重置当前 session 的运行态字段,例如 answer、duration、step/tool count。
|
- 每次有效 Chat/AIOps 执行都会创建新的 `runId`。
|
||||||
- `sessionId` 是业务关联键,不依赖数据库自增 ID 暴露给外部。
|
- Trace 和 Feedback 新客户端应传 `runId`;只传 `sessionId` 时兼容解析 latest run。
|
||||||
|
- latest run 排序使用 `diagnosis_run.created_at DESC, id DESC`,不使用 `updated_at`。
|
||||||
|
|
||||||
## 4. 状态流转
|
## 4. 运行状态流转
|
||||||
|
|
||||||
```mermaid
|
```mermaid
|
||||||
stateDiagram-v2
|
stateDiagram-v2
|
||||||
@@ -77,14 +76,14 @@ stateDiagram-v2
|
|||||||
|
|
||||||
字段边界:
|
字段边界:
|
||||||
|
|
||||||
| 字段 | 含义 |
|
| 字段 | 所属表 | 含义 |
|
||||||
|---|---|
|
|---|---|---|
|
||||||
| `status` | 执行状态:`PENDING` / `RUNNING` / `SUCCESS` / `FAILED` |
|
| `status` | `diagnosis_run` | 单次运行执行状态 |
|
||||||
| `answer` | Agent 最终返回给用户的报告或答复 |
|
| `answer` | `diagnosis_run` | 本次运行最终报告或答复 |
|
||||||
| `self_evaluation` | 系统自评估 JSON |
|
| `self_evaluation` | `diagnosis_run` | 本次运行系统自评估 JSON |
|
||||||
| `feedback` | 用户反馈:`useful` / `not_useful` / null |
|
| `feedback` | `diagnosis_run` | 本次运行用户反馈 |
|
||||||
|
|
||||||
`feedback` 不修改 `status`。一个执行成功但用户标记 `not_useful` 的 session,仍然应该是 `SUCCESS + feedback=not_useful`。
|
`feedback` 不修改 `status`。一个执行成功但用户标记 `not_useful` 的 run,仍然应该是 `SUCCESS + feedback=not_useful`。
|
||||||
|
|
||||||
## 5. agent_step 写入
|
## 5. agent_step 写入
|
||||||
|
|
||||||
@@ -93,93 +92,49 @@ stateDiagram-v2
|
|||||||
```mermaid
|
```mermaid
|
||||||
sequenceDiagram
|
sequenceDiagram
|
||||||
autonumber
|
autonumber
|
||||||
participant Agent as ReactAgent
|
participant Agent as Agent
|
||||||
participant Hook as AgentLoggingHook
|
participant Hook as AgentLoggingHook
|
||||||
participant DB as agent_step
|
participant DB as agent_step
|
||||||
|
|
||||||
Agent->>Hook: before_model(messages, sessionId)
|
Agent->>Hook: before_model(messages, sessionId, runId)
|
||||||
Hook->>DB: insert step_index / agent_name / model_input
|
Hook->>DB: insert step(session_id, run_id, model_input, step_index)
|
||||||
Agent-->>Agent: model call
|
Agent->>Hook: after_model(output, sessionId, runId)
|
||||||
Agent->>Hook: after_model(messages, sessionId)
|
Hook->>DB: update model_output, duration, token_count, has_tool_call
|
||||||
Hook->>DB: update model_output / thought / has_tool_call / duration / token_count
|
|
||||||
```
|
```
|
||||||
|
|
||||||
当前记录:
|
新写入必须带 `run_id`,同时保留 `session_id` 便于粗粒度排查。
|
||||||
|
|
||||||
- `session_id`
|
|
||||||
- `step_index`
|
|
||||||
- `agent_name`
|
|
||||||
- `model_input`
|
|
||||||
- `model_output`
|
|
||||||
- `thought`
|
|
||||||
- `has_tool_call`
|
|
||||||
- `duration_ms`
|
|
||||||
- `token_count`
|
|
||||||
|
|
||||||
## 6. tool_invocation 写入
|
## 6. tool_invocation 写入
|
||||||
|
|
||||||
工具调用记录真实工具事实,不记录模型猜测。
|
工具调用记录同样通过执行上下文拿到 `sessionId + runId`:
|
||||||
|
|
||||||
关键字段:
|
|
||||||
|
|
||||||
```text
|
```text
|
||||||
session_id
|
ToolInvocationRecorder
|
||||||
step_id
|
-> tool_invocation.session_id
|
||||||
tool_name
|
-> tool_invocation.run_id
|
||||||
input_params
|
-> retrieval_details / evidence_refs
|
||||||
output_preview
|
|
||||||
output_length
|
|
||||||
retrieval_layer
|
|
||||||
l0_match_count
|
|
||||||
l1_match_count
|
|
||||||
retrieval_details
|
|
||||||
-> evidence_refs
|
|
||||||
relevance_level
|
|
||||||
dedup_reason
|
|
||||||
duration_ms
|
|
||||||
success
|
|
||||||
error_message
|
|
||||||
```
|
```
|
||||||
|
|
||||||
对 `lookup_knowledge`,`retrieval_details` 会承载 L0/L1、领域、证据状态、去重等检索细节。对日志、指标和知识库工具,`retrieval_details.evidence_refs` 会记录 Gatekeeper 可核验的最小证据引用:
|
Verifier、Gatekeeper 和 EvaluationService 应按 `run_id` 读取工具调用,避免同一 session 的其他 run 参与评分或证据校验。
|
||||||
|
|
||||||
```json
|
|
||||||
{
|
|
||||||
"evidence_refs": [
|
|
||||||
{
|
|
||||||
"raw_path": "$.logs[0]",
|
|
||||||
"text": "最小证据文本"
|
|
||||||
}
|
|
||||||
]
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
当工具明确没有返回匹配证据时,可以记录 `raw_path=$.no_evidence`。该路径只表示“本次工具查询未检索到匹配证据”,不表示问题被排除。
|
|
||||||
|
|
||||||
## 7. Trace API 聚合
|
## 7. Trace API 聚合
|
||||||
|
|
||||||
```text
|
```text
|
||||||
GET /api/diagnosis/{sessionId}/trace
|
GET /api/diagnosis/{sessionId}/trace
|
||||||
|
GET /api/diagnosis/{sessionId}/trace?runId=run-...
|
||||||
```
|
```
|
||||||
|
|
||||||
聚合逻辑:
|
聚合逻辑:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
diagnosis_session by sessionId
|
diagnosis_run by sessionId + runId
|
||||||
+ agent_step ordered by step_index
|
+ chat_session metadata when available
|
||||||
+ tool_invocation ordered by id
|
+ agent_step where run_id = runId, ordered by the Trace API
|
||||||
|
+ tool_invocation where run_id = runId order by id
|
||||||
-> DiagnosisTraceResponse
|
-> DiagnosisTraceResponse
|
||||||
```
|
```
|
||||||
|
|
||||||
Trace 视图回答的问题:
|
当 `runId` 缺失时,Trace API 为兼容旧客户端解析最新 run,并在响应中返回 resolved `runId`。当 `runId` 属于其他 `sessionId` 时,API 必须拒绝,不能泄漏其他会话的 Trace。
|
||||||
|
|
||||||
- 这次诊断是否成功?
|
|
||||||
- 哪些 Agent 参与了?
|
|
||||||
- 每一步模型输入输出是什么摘要?
|
|
||||||
- 调用了哪些工具?
|
|
||||||
- 工具返回了什么证据?
|
|
||||||
- Gatekeeper / Verifier / Composer / AIOps rule 是否通过?
|
|
||||||
- 用户是否反馈有用?
|
|
||||||
|
|
||||||
## 8. Chat 与 AIOps 差异
|
## 8. Chat 与 AIOps 差异
|
||||||
|
|
||||||
@@ -189,22 +144,17 @@ Trace 视图回答的问题:
|
|||||||
| 编排方式 | `SequentialAgent`: Planner -> Executor -> Gatekeeper -> Verifier -> Composer | `SupervisorAgent`: Planner + Executor |
|
| 编排方式 | `SequentialAgent`: Planner -> Executor -> Gatekeeper -> Verifier -> Composer | `SupervisorAgent`: Planner + Executor |
|
||||||
| 自评估 | `rule_evaluation` + `verifier_evaluation` | `aiops_rule_evaluation` |
|
| 自评估 | `rule_evaluation` + `verifier_evaluation` | `aiops_rule_evaluation` |
|
||||||
| 答案字段 | Chat 最终答复 | 告警分析报告 |
|
| 答案字段 | Chat 最终答复 | 告警分析报告 |
|
||||||
| payload | 用户自然语言 + history | alert payload 或 auto-discovery |
|
| runId 暴露 | `/api/chat` JSON response | `/api/ai_ops` SSE metadata message |
|
||||||
|
|
||||||
## 9. 清理与边界
|
## 9. 清理与边界
|
||||||
|
|
||||||
当前会话持久化边界:
|
- Redis 会话历史用于多轮上下文,不是长期审计记录。
|
||||||
|
- MySQL `diagnosis_run + agent_step + tool_invocation` 是主要可回放来源。
|
||||||
- MySQL Trace 记录是主要可回放来源。
|
- `chat_session.expires_at` 只是目录元数据;Redis 消息历史可独立过期。
|
||||||
- Chat 历史仍可作为请求上下文传入 Agent,但不是本文档的主持久化模型。
|
- `RetrievedDocTracker` 仍是 session 级运行时去重状态,诊断结束后清理。
|
||||||
- Redis 主会话存储是历史设计,不作为当前架构事实。
|
|
||||||
- `RetrievedDocTracker` 是 session 级运行时去重状态,诊断结束后清理。
|
|
||||||
|
|
||||||
## 10. 后续增强
|
## 10. 后续增强
|
||||||
|
|
||||||
可考虑:
|
|
||||||
|
|
||||||
1. Trace API 增加更结构化的 `self_evaluation` 展示。
|
1. Trace API 增加更结构化的 `self_evaluation` 展示。
|
||||||
2. `agent_step` 与 `tool_invocation.step_id` 建立更严格关联。
|
2. `agent_step` 与 `tool_invocation.step_id` 建立更严格关联。
|
||||||
3. 对多轮同 session 诊断增加 run id,避免复用 session 时历史记录混杂。
|
3. 旧 `diagnosis_session` 只读观察期结束后,再评估数据库层面的约束收紧或归档策略。
|
||||||
4. 为 Trace 增加导出能力,服务面试演示和回归分析。
|
|
||||||
|
|||||||
@@ -0,0 +1,17 @@
|
|||||||
|
# 2026-07-09 MVP 文档清理归档
|
||||||
|
|
||||||
|
本目录保存本次清理中从当前入口移出的历史设计材料。这些文档仍有追溯价值,但不再代表当前可运行实现。
|
||||||
|
|
||||||
|
## 归档内容
|
||||||
|
|
||||||
|
| 目录 | 内容 | 归档原因 |
|
||||||
|
|---|---|---|
|
||||||
|
| `discuss/` | 早期 Executor Prompt、L0、RAG 讨论稿 | 已被当前 architecture、OpenSpec change 和 issue 取代 |
|
||||||
|
| `plan/` | `session-storage-design.md` | 会话存储已实现,当前表以 Flyway 和 `mvp/tables/` 为准 |
|
||||||
|
| `notes/` | 早期工程决策和 Demo Trace 验收笔记 | 相关内容已沉淀到 architecture、demo、eval 和 devflow |
|
||||||
|
|
||||||
|
## 使用原则
|
||||||
|
|
||||||
|
- 当前架构以 `mvp/architecture/` 为准。
|
||||||
|
- 当前表结构以 `mvp/tables/`、Flyway migration 和实体类为准。
|
||||||
|
- 当前问题入口以 `mvp/issues/README.md` 为准。
|
||||||
+33
-11
@@ -8,7 +8,9 @@
|
|||||||
- `interview-walkthrough.md`:面试讲解话术。
|
- `interview-walkthrough.md`:面试讲解话术。
|
||||||
- `evidence-pipeline-scenarios.md`:PASS / LOW_CONFID / REJECT / no-evidence 场景矩阵。
|
- `evidence-pipeline-scenarios.md`:PASS / LOW_CONFID / REJECT / no-evidence 场景矩阵。
|
||||||
- `trace-inspection-checklist.md`:Trace 字段检查清单。
|
- `trace-inspection-checklist.md`:Trace 字段检查清单。
|
||||||
|
- `scripts/run-interview-demo-check.ps1`:面试预检脚本,包含服务可达性、Chat、Trace、反馈和 summary 输出。
|
||||||
- `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。
|
- `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。
|
||||||
|
- `interview-q-and-a.md`:面试追问回答,覆盖 Agent 工程取舍、审计和评测。
|
||||||
- `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。
|
- `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。
|
||||||
- `requests/narrow-highcpu-chat.json`:窄范围正向观察请求。
|
- `requests/narrow-highcpu-chat.json`:窄范围正向观察请求。
|
||||||
- `requests/hikari-no-evidence-chat.json`:no-evidence 负向观察请求。
|
- `requests/hikari-no-evidence-chat.json`:no-evidence 负向观察请求。
|
||||||
@@ -37,7 +39,7 @@ http://localhost:9900
|
|||||||
最快方式:
|
最快方式:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1
|
||||||
```
|
```
|
||||||
|
|
||||||
脚本会生成:
|
脚本会生成:
|
||||||
@@ -46,6 +48,7 @@ powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-de
|
|||||||
mvp/demo/output/chat-response.json
|
mvp/demo/output/chat-response.json
|
||||||
mvp/demo/output/trace-response.json
|
mvp/demo/output/trace-response.json
|
||||||
mvp/demo/output/feedback-response.json
|
mvp/demo/output/feedback-response.json
|
||||||
|
mvp/demo/output/interview-demo-summary.json
|
||||||
```
|
```
|
||||||
|
|
||||||
手动请求:
|
手动请求:
|
||||||
@@ -64,10 +67,23 @@ Invoke-RestMethod `
|
|||||||
-Body $body
|
-Body $body
|
||||||
```
|
```
|
||||||
|
|
||||||
|
如果要继续手动查询同一次诊断运行,先保留响应中的 run id:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
$chat = Invoke-RestMethod `
|
||||||
|
-Method Post `
|
||||||
|
-Uri "http://localhost:9900/api/chat" `
|
||||||
|
-ContentType "application/json" `
|
||||||
|
-Body $body
|
||||||
|
|
||||||
|
$runId = $chat.data.runId
|
||||||
|
```
|
||||||
|
|
||||||
期望结果:
|
期望结果:
|
||||||
|
|
||||||
- `data.success = true`
|
- `data.success = true`
|
||||||
- `data.sessionId = mvp-demo-payment-timeout-001`
|
- `data.sessionId = mvp-demo-payment-timeout-001`
|
||||||
|
- `data.runId` 为本次诊断运行的唯一 ID
|
||||||
- `data.answer` 包含诊断答复
|
- `data.answer` 包含诊断答复
|
||||||
|
|
||||||
## 4. 查询 Trace
|
## 4. 查询 Trace
|
||||||
@@ -75,22 +91,27 @@ Invoke-RestMethod `
|
|||||||
```powershell
|
```powershell
|
||||||
Invoke-RestMethod `
|
Invoke-RestMethod `
|
||||||
-Method Get `
|
-Method Get `
|
||||||
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace"
|
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace?runId=$runId"
|
||||||
```
|
```
|
||||||
|
|
||||||
期望结果:
|
期望结果:
|
||||||
|
|
||||||
- `code = 200`
|
- `code = 200`
|
||||||
|
- `data.runId` 等于 `$runId`
|
||||||
- `data.session.sessionId` 等于 Chat session id
|
- `data.session.sessionId` 等于 Chat session id
|
||||||
|
- `data.run.runId` 等于 `$runId`
|
||||||
- `data.steps` 包含 planner / executor / verifier 等步骤
|
- `data.steps` 包含 planner / executor / verifier 等步骤
|
||||||
- `data.toolInvocations` 包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具
|
- `data.toolInvocations` 包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具
|
||||||
- `data.session.selfEvaluation` 包含 verifier 或 rule evaluation
|
- `data.session.selfEvaluation` 包含 verifier 或 rule evaluation
|
||||||
|
- Chat V2 链路中,`data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` 记录 Chat Prompt 审计版本
|
||||||
|
- Chat V2 链路中,`data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` 记录 Gatekeeper 规则集版本
|
||||||
|
|
||||||
## 5. 提交反馈
|
## 5. 提交反馈
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
$feedback = @{
|
$feedback = @{
|
||||||
sessionId = $sessionId
|
sessionId = $sessionId
|
||||||
|
runId = $runId
|
||||||
feedback = "useful"
|
feedback = "useful"
|
||||||
} | ConvertTo-Json
|
} | ConvertTo-Json
|
||||||
|
|
||||||
@@ -104,7 +125,8 @@ Invoke-RestMethod `
|
|||||||
期望结果:
|
期望结果:
|
||||||
|
|
||||||
- `success = true`
|
- `success = true`
|
||||||
- 后续 Trace 中 `data.session.feedback = useful`
|
- `runId = $runId`
|
||||||
|
- 后续精确 Trace 中 `data.session.feedback = useful`
|
||||||
- useful 反馈会尝试沉淀 `case_library`
|
- useful 反馈会尝试沉淀 `case_library`
|
||||||
|
|
||||||
## 6. AIOps 告警诊断 Demo
|
## 6. AIOps 告警诊断 Demo
|
||||||
@@ -130,19 +152,19 @@ Invoke-WebRequest `
|
|||||||
|
|
||||||
期望结果:
|
期望结果:
|
||||||
|
|
||||||
- SSE 首条包含 `session` 消息,sessionId 为 `mvp-demo-aiops-payment-cpu-001`
|
- SSE 首条是 `type=metadata` 的 `message` 事件,包含 sessionId `mvp-demo-aiops-payment-cpu-001` 和本次 AIOps `runId`
|
||||||
- 后续流式输出包含 AIOps 告警分析报告
|
- 后续流式输出包含 AIOps 告警分析报告
|
||||||
- 报告聚焦输入的 `HighCPUUsage/payment-service`
|
- 报告聚焦输入的 `HighCPUUsage/payment-service`
|
||||||
- 同一 session 的 Trace 中 `data.session.agentFlow = AI_OPS`
|
- 精确 Trace 中 `data.session.agentFlow = AI_OPS`
|
||||||
- `data.session.answer` 包含最终告警报告
|
- `data.session.answer` 包含最终告警报告
|
||||||
- `data.toolInvocations` 包含证据工具调用
|
- `data.toolInvocations` 包含证据工具调用
|
||||||
|
|
||||||
查询 AIOps Trace:
|
查询 AIOps Trace 时优先使用 SSE metadata 中的 runId:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
Invoke-RestMethod `
|
Invoke-RestMethod `
|
||||||
-Method Get `
|
-Method Get `
|
||||||
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace"
|
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace?runId=$aiopsRunId"
|
||||||
```
|
```
|
||||||
|
|
||||||
## 7. Demo 主线
|
## 7. Demo 主线
|
||||||
@@ -150,7 +172,7 @@ Invoke-RestMethod `
|
|||||||
Chat 主线:
|
Chat 主线:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
一个 session id
|
一个 session id + 一个 run id
|
||||||
-> 用户问题
|
-> 用户问题
|
||||||
-> 多 Agent 执行
|
-> 多 Agent 执行
|
||||||
-> 证据工具
|
-> 证据工具
|
||||||
@@ -163,7 +185,7 @@ Chat 主线:
|
|||||||
AIOps 主线:
|
AIOps 主线:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
一个 session id
|
一个 session id + 一个 run id
|
||||||
-> 告警 payload
|
-> 告警 payload
|
||||||
-> AIOps Planner / Executor
|
-> AIOps Planner / Executor
|
||||||
-> 证据工具
|
-> 证据工具
|
||||||
@@ -176,8 +198,8 @@ AIOps 主线:
|
|||||||
|
|
||||||
面试时不要把所有安全场景都压到 live LLM 现场表现上。建议使用:
|
面试时不要把所有安全场景都压到 live LLM 现场表现上。建议使用:
|
||||||
|
|
||||||
- `scripts/run-payment-timeout-demo.ps1` 跑主路径。
|
- `scripts/run-interview-demo-check.ps1` 跑主路径和预检 summary。
|
||||||
- `evidence-pipeline-scenarios.md` 讲解 PASS / LOW_CONFID / REJECT / no-evidence 矩阵。
|
- `evidence-pipeline-scenarios.md` 讲解 PASS / LOW_CONFID / REJECT / no-evidence 矩阵。
|
||||||
- `mvp/eval/reports/baseline-report.md` 证明固定 fixture 10/10 通过。
|
- `mvp/eval/reports/baseline-report.md` 证明固定 fixture 12/12 通过。
|
||||||
|
|
||||||
这样可以同时展示真实链路和确定性回归能力。
|
这样可以同时展示真实链路和确定性回归能力。
|
||||||
|
|||||||
@@ -2,7 +2,7 @@
|
|||||||
|
|
||||||
## 1. 目标
|
## 1. 目标
|
||||||
|
|
||||||
验证旧版 `/api/ai_ops` 入口可以作为可追踪的告警触发诊断入口,并且 payload 模式下报告聚焦输入告警。
|
验证 `/api/ai_ops` 入口可以作为可追踪的告警触发诊断入口,并且 payload 模式下报告聚焦输入告警。
|
||||||
|
|
||||||
## 2. 输入
|
## 2. 输入
|
||||||
|
|
||||||
@@ -25,14 +25,14 @@
|
|||||||
|
|
||||||
## 3. 验收标准
|
## 3. 验收标准
|
||||||
|
|
||||||
1. SSE 流输出 `session` 消息,且包含请求中的 session id。
|
1. SSE 流首条输出 `type=metadata` 的 `message` 事件,且包含请求中的 session id 和本次 AIOps run id。
|
||||||
2. AIOps 执行创建或更新 `diagnosis_session`,并写入 `agent_flow = AI_OPS`。
|
2. AIOps 执行创建 `diagnosis_run`,并写入 `agent_flow = AI_OPS`。
|
||||||
3. 持久化的 session query 包含告警名、服务名、等级、时间范围和描述。
|
3. 持久化的 run query 包含告警名、服务名、等级、时间范围和描述。
|
||||||
4. 如果生成最终报告,`diagnosis_session.answer` 包含该报告。
|
4. 如果生成最终报告,`diagnosis_run.answer` 包含该报告。
|
||||||
5. `GET /api/diagnosis/{sessionId}/trace` 返回 AIOps session、按顺序排列的 agent steps 和 tool invocations。
|
5. `GET /api/diagnosis/{sessionId}/trace?runId=...` 返回 AIOps run、按顺序排列的 agent steps 和 tool invocations。
|
||||||
6. payload 模式下,报告主线聚焦 `HighCPUUsage/payment-service`。
|
6. payload 模式下,报告主线聚焦 `HighCPUUsage/payment-service`。
|
||||||
7. 其他活跃告警最多作为相关风险或上下文出现,不应展开成完整独立根因章节。
|
7. 其他活跃告警最多作为相关风险或上下文出现,不应展开成完整独立根因章节。
|
||||||
8. `self_evaluation.aiops_rule_evaluation` 存在,并能反映报告完整性、payload 聚焦和证据工具覆盖情况。
|
8. `diagnosis_run.self_evaluation.aiops_rule_evaluation` 存在,并能反映报告完整性、payload 聚焦和证据工具覆盖情况。
|
||||||
|
|
||||||
## 4. 已知边界
|
## 4. 已知边界
|
||||||
|
|
||||||
|
|||||||
@@ -52,12 +52,12 @@ Invoke-RestMethod `
|
|||||||
-Body $body
|
-Body $body
|
||||||
```
|
```
|
||||||
|
|
||||||
然后查询同一 session:
|
然后保留响应里的 `runId`,查询同一 run 的 Trace:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
Invoke-RestMethod `
|
Invoke-RestMethod `
|
||||||
-Method Get `
|
-Method Get `
|
||||||
-Uri "http://localhost:9900/api/diagnosis/mvp-demo-narrow-highcpu-001/trace"
|
-Uri "http://localhost:9900/api/diagnosis/mvp-demo-narrow-highcpu-001/trace?runId=$runId"
|
||||||
```
|
```
|
||||||
|
|
||||||
注意:除 payment-timeout 主路径外,其它 live 请求是“可尝试”的演示入口;稳定验收以 `mvp/eval` fixture 和 baseline 为准。
|
注意:除 payment-timeout 主路径外,其它 live 请求是“可尝试”的演示入口;稳定验收以 `mvp/eval` fixture 和 baseline 为准。
|
||||||
|
|||||||
@@ -0,0 +1,33 @@
|
|||||||
|
# 面试追问 Q&A
|
||||||
|
|
||||||
|
## 为什么不用普通 Chatbot?
|
||||||
|
|
||||||
|
这个项目的重点不是生成一段诊断文本,而是把诊断拆成可审计链路:Planner 拆解问题,Executor 调工具拿证据,Gatekeeper 用代码核验证据引用,Verifier 判断可推导性,Composer 生成最终表达。`sessionId` 保留多轮上下文,`runId` 精确绑定一次诊断运行,Trace 和 Feedback 都可以按 `sessionId + runId` 回放和定位。
|
||||||
|
|
||||||
|
## 为什么 RAG 要做成显式工具?
|
||||||
|
|
||||||
|
`lookup_knowledge` 保持显式工具调用,才能在 `tool_invocation` 里看到 Agent 查了什么、命中了什么、相关性等级是什么,以及最终答案是否真的使用了这些证据。隐式 Advisor 更方便,但不利于审计 Agent 决策。
|
||||||
|
|
||||||
|
## 怎么防止 Executor 幻觉?
|
||||||
|
|
||||||
|
Executor 不直接负责最终用户答案,而是输出 `executor_evidence_v2` 的微观事实和证据引用。Gatekeeper 会校验 `source_invocation_id`、`raw_path`、`evidence_excerpt` 是否真实存在;Verifier 再判断 claim 是否能由已验真的证据推出;Composer 只表达 Verifier 允许输出的内容。
|
||||||
|
|
||||||
|
## LOW_CONFID 是失败吗?
|
||||||
|
|
||||||
|
不是。`LOW_CONFID` 表示当前证据不足以支撑强结论,但系统仍然可以安全表达已确认事实和缺失信息。面试时可以把它作为“没有证据就不强答”的质量门禁,而不是模型能力失败。
|
||||||
|
|
||||||
|
## Prompt 改了怎么审计?
|
||||||
|
|
||||||
|
Chat verifier evaluation 里会记录 `prompt_audit.version`,并列出 planner、executor、verifier、composer 的 Prompt 版本和资源路径。它不保存完整 Prompt 文本,只保留用于回放和回归解释的紧凑元数据。
|
||||||
|
|
||||||
|
## Gatekeeper 改了怎么审计?
|
||||||
|
|
||||||
|
Gatekeeper 结果里记录 `gatekeeper_result.rule_set_version` 和已启用规则元数据摘要。规则执行仍是确定性 Java 代码,版本和规则元数据用于解释“这次引用验真用的是哪套规则”。
|
||||||
|
|
||||||
|
## 为什么现在不拆 SubAgent?
|
||||||
|
|
||||||
|
当前 MVP 的主要风险不是 Agent 数量不够,而是证据、验证和回归是否稳定。文档里的演进路线把 SubAgent 放在 P2:等故障类型、工具权限和评测集足够明确后再拆,避免只是移动复杂度。
|
||||||
|
|
||||||
|
## 为什么 baseline 比 live demo 更重要?
|
||||||
|
|
||||||
|
live demo 证明链路在当前环境能跑通,但 LLM 和外部依赖会波动。`mvp/eval` 的固定 fixture baseline 是确定性回归来源,用来判断 Prompt、工具、Gatekeeper、Verifier 或 Composer 的改动有没有让系统退化。
|
||||||
@@ -13,7 +13,7 @@
|
|||||||
关键主张不是“模型回答了一次”,而是:
|
关键主张不是“模型回答了一次”,而是:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
系统能展示用了什么证据、答案如何被检查、如何用 sessionId 回放整次诊断。
|
系统能展示用了什么证据、答案如何被检查、如何用 sessionId + runId 精确回放这次诊断。
|
||||||
```
|
```
|
||||||
|
|
||||||
## 2. Demo 流程
|
## 2. Demo 流程
|
||||||
@@ -23,7 +23,7 @@
|
|||||||
3. 打开 `mvp/demo/output/chat-response.json`。
|
3. 打开 `mvp/demo/output/chat-response.json`。
|
||||||
4. 打开 `mvp/demo/output/trace-response.json`。
|
4. 打开 `mvp/demo/output/trace-response.json`。
|
||||||
5. 指出证据工具和 verifier evaluation。
|
5. 指出证据工具和 verifier evaluation。
|
||||||
6. 提交 feedback,并展示它挂在同一个 session 上。
|
6. 提交 feedback,并展示它挂在当前 run 上。
|
||||||
7. 打开 `evidence-pipeline-scenarios.md`,说明 PASS / LOW_CONFID / REJECT / no-evidence 的固定回归矩阵。
|
7. 打开 `evidence-pipeline-scenarios.md`,说明 PASS / LOW_CONFID / REJECT / no-evidence 的固定回归矩阵。
|
||||||
|
|
||||||
## 3. 命令
|
## 3. 命令
|
||||||
@@ -59,7 +59,7 @@ mvp/demo/output/chat-response.json
|
|||||||
话术:
|
话术:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
这是用户看到的答案。这里的 sessionId 是稳定的,所以我后面可以追踪这一次回答是怎么来的。
|
这是用户看到的答案。这里的 sessionId 是稳定的,同时响应里会返回 runId,所以我后面可以精确追踪这一次回答是怎么来的。
|
||||||
```
|
```
|
||||||
|
|
||||||
### 4.2 证据 Trace
|
### 4.2 证据 Trace
|
||||||
@@ -112,7 +112,7 @@ mvp/demo/output/feedback-response.json
|
|||||||
话术:
|
话术:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
feedback 会挂在同一个 diagnosis session 上。
|
feedback 会挂在当前 diagnosis run 上。
|
||||||
这让后续挖掘 useful case 或 not_useful bad case 成为可能。
|
这让后续挖掘 useful case 或 not_useful bad case 成为可能。
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|||||||
@@ -12,14 +12,16 @@
|
|||||||
|
|
||||||
## 3. 验收标准
|
## 3. 验收标准
|
||||||
|
|
||||||
1. Chat 返回成功答复,且 session id 与请求一致。
|
1. Chat 返回成功答复,且 session id 与请求一致,并返回本次诊断的 run id。
|
||||||
2. Trace API 返回 session 元数据、最终答案、按顺序排列的 agent steps 和 tool invocations。
|
2. Trace API 使用 `sessionId + runId` 返回会话元数据、运行摘要、最终答案、按顺序排列的 agent steps 和 tool invocations。
|
||||||
3. Trace 中有足够证据说明用了哪些工具,以及 verifier / self-evaluation 是否已持久化。
|
3. Trace 中有足够证据说明用了哪些工具,以及 verifier / self-evaluation 是否已持久化。
|
||||||
4. 可以使用同一个 session id 提交反馈。
|
4. 可以使用同一个 session id 和本次 run id 提交反馈。
|
||||||
5. 后续 Trace 查询能看到已持久化的 feedback 值。
|
5. 后续精确 Trace 查询能看到已持久化的 feedback 值。
|
||||||
|
|
||||||
## 4. 需要检查的 Trace 字段
|
## 4. 需要检查的 Trace 字段
|
||||||
|
|
||||||
|
- `data.runId`
|
||||||
|
- `data.run.runId`
|
||||||
- `data.session.query`
|
- `data.session.query`
|
||||||
- `data.session.answer`
|
- `data.session.answer`
|
||||||
- `data.session.selfEvaluation`
|
- `data.session.selfEvaluation`
|
||||||
|
|||||||
@@ -0,0 +1,168 @@
|
|||||||
|
param(
|
||||||
|
[string]$BaseUrl = "http://localhost:9900",
|
||||||
|
[string]$SessionId = "mvp-demo-interview-payment-timeout-001",
|
||||||
|
[string]$RequestFile = "$PSScriptRoot/../requests/payment-timeout-chat.json",
|
||||||
|
[string]$OutputDir = "$PSScriptRoot/../output"
|
||||||
|
)
|
||||||
|
|
||||||
|
$ErrorActionPreference = "Stop"
|
||||||
|
|
||||||
|
function Test-ServiceReachable {
|
||||||
|
param([string]$Url)
|
||||||
|
|
||||||
|
try {
|
||||||
|
$request = [System.Net.WebRequest]::Create($Url)
|
||||||
|
$request.Method = "GET"
|
||||||
|
$request.Timeout = 5000
|
||||||
|
$response = $request.GetResponse()
|
||||||
|
$response.Close()
|
||||||
|
return $true
|
||||||
|
} catch [System.Net.WebException] {
|
||||||
|
if ($_.Exception.Response -ne $null) {
|
||||||
|
$_.Exception.Response.Close()
|
||||||
|
return $true
|
||||||
|
}
|
||||||
|
return $false
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function Get-TraceData {
|
||||||
|
param($TraceResponse)
|
||||||
|
|
||||||
|
if ($TraceResponse.PSObject.Properties.Name -contains "data") {
|
||||||
|
return $TraceResponse.data
|
||||||
|
}
|
||||||
|
return $TraceResponse
|
||||||
|
}
|
||||||
|
|
||||||
|
function Get-SelfEvaluation {
|
||||||
|
param($TraceData)
|
||||||
|
|
||||||
|
if ($null -eq $TraceData -or $null -eq $TraceData.session) {
|
||||||
|
return $null
|
||||||
|
}
|
||||||
|
return $TraceData.session.selfEvaluation
|
||||||
|
}
|
||||||
|
|
||||||
|
function Get-ToolNames {
|
||||||
|
param($TraceData)
|
||||||
|
|
||||||
|
if ($null -eq $TraceData -or $null -eq $TraceData.toolInvocations) {
|
||||||
|
return @()
|
||||||
|
}
|
||||||
|
return @($TraceData.toolInvocations | ForEach-Object { $_.toolName } | Where-Object { $_ } | Sort-Object -Unique)
|
||||||
|
}
|
||||||
|
|
||||||
|
New-Item -ItemType Directory -Force -Path $OutputDir | Out-Null
|
||||||
|
|
||||||
|
Write-Host "Running interview demo preflight..."
|
||||||
|
Write-Host "BaseUrl: $BaseUrl"
|
||||||
|
Write-Host "SessionId: $SessionId"
|
||||||
|
|
||||||
|
if (-not (Test-ServiceReachable -Url $BaseUrl)) {
|
||||||
|
throw "Service is not reachable: $BaseUrl. Start the app with mvp-demo profile first: mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo"
|
||||||
|
}
|
||||||
|
|
||||||
|
$request = Get-Content -Raw -Encoding UTF8 -Path $RequestFile | ConvertFrom-Json
|
||||||
|
$request.Id = $SessionId
|
||||||
|
$body = $request | ConvertTo-Json -Depth 8
|
||||||
|
|
||||||
|
$chatRequest = @{
|
||||||
|
Method = "Post"
|
||||||
|
Uri = "$BaseUrl/api/chat"
|
||||||
|
ContentType = "application/json; charset=utf-8"
|
||||||
|
Body = $body
|
||||||
|
}
|
||||||
|
$chat = Invoke-RestMethod @chatRequest
|
||||||
|
|
||||||
|
$chatPath = Join-Path $OutputDir "chat-response.json"
|
||||||
|
$chat | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $chatPath
|
||||||
|
|
||||||
|
$runId = $chat.data.runId
|
||||||
|
if (-not $runId) {
|
||||||
|
throw "Chat response did not include runId; exact trace verification cannot continue."
|
||||||
|
}
|
||||||
|
|
||||||
|
$traceRequest = @{
|
||||||
|
Method = "Get"
|
||||||
|
Uri = "$BaseUrl/api/diagnosis/$SessionId/trace?runId=$([System.Uri]::EscapeDataString($runId))"
|
||||||
|
}
|
||||||
|
$trace = Invoke-RestMethod @traceRequest
|
||||||
|
|
||||||
|
$tracePath = Join-Path $OutputDir "trace-response.json"
|
||||||
|
$trace | ConvertTo-Json -Depth 80 | Set-Content -Encoding UTF8 -Path $tracePath
|
||||||
|
|
||||||
|
$feedbackBody = @{
|
||||||
|
sessionId = $SessionId
|
||||||
|
runId = $runId
|
||||||
|
feedback = "useful"
|
||||||
|
} | ConvertTo-Json
|
||||||
|
|
||||||
|
$feedbackRequest = @{
|
||||||
|
Method = "Post"
|
||||||
|
Uri = "$BaseUrl/api/feedback"
|
||||||
|
ContentType = "application/json; charset=utf-8"
|
||||||
|
Body = $feedbackBody
|
||||||
|
}
|
||||||
|
$feedback = Invoke-RestMethod @feedbackRequest
|
||||||
|
|
||||||
|
$feedbackPath = Join-Path $OutputDir "feedback-response.json"
|
||||||
|
$feedback | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $feedbackPath
|
||||||
|
|
||||||
|
$traceData = Get-TraceData -TraceResponse $trace
|
||||||
|
$selfEvaluation = Get-SelfEvaluation -TraceData $traceData
|
||||||
|
$verifierEvaluation = $null
|
||||||
|
if ($null -ne $selfEvaluation) {
|
||||||
|
$verifierEvaluation = $selfEvaluation.verifier_evaluation
|
||||||
|
}
|
||||||
|
|
||||||
|
$gatekeeperResult = $null
|
||||||
|
$promptAudit = $null
|
||||||
|
if ($null -ne $verifierEvaluation) {
|
||||||
|
$gatekeeperResult = $verifierEvaluation.gatekeeper_result
|
||||||
|
$promptAudit = $verifierEvaluation.prompt_audit
|
||||||
|
}
|
||||||
|
|
||||||
|
$verdict = $null
|
||||||
|
$gatekeeperStatus = $null
|
||||||
|
$gatekeeperRuleSetVersion = $null
|
||||||
|
$promptAuditVersion = $null
|
||||||
|
if ($null -ne $verifierEvaluation) {
|
||||||
|
$verdict = $verifierEvaluation.verdict
|
||||||
|
}
|
||||||
|
if ($null -ne $gatekeeperResult) {
|
||||||
|
$gatekeeperStatus = $gatekeeperResult.status
|
||||||
|
$gatekeeperRuleSetVersion = $gatekeeperResult.rule_set_version
|
||||||
|
}
|
||||||
|
if ($null -ne $promptAudit) {
|
||||||
|
$promptAuditVersion = $promptAudit.version
|
||||||
|
}
|
||||||
|
$toolNames = Get-ToolNames -TraceData $traceData
|
||||||
|
$summaryPath = Join-Path $OutputDir "interview-demo-summary.json"
|
||||||
|
|
||||||
|
$summary = [ordered]@{
|
||||||
|
sessionId = $SessionId
|
||||||
|
runId = $runId
|
||||||
|
baseUrl = $BaseUrl
|
||||||
|
chatSuccess = $chat.data.success
|
||||||
|
verdict = $verdict
|
||||||
|
gatekeeperStatus = $gatekeeperStatus
|
||||||
|
gatekeeperRuleSetVersion = $gatekeeperRuleSetVersion
|
||||||
|
promptAuditVersion = $promptAuditVersion
|
||||||
|
toolNames = $toolNames
|
||||||
|
paths = [ordered]@{
|
||||||
|
chat = $chatPath
|
||||||
|
trace = $tracePath
|
||||||
|
feedback = $feedbackPath
|
||||||
|
summary = $summaryPath
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
$summary | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path $summaryPath
|
||||||
|
|
||||||
|
Write-Host ""
|
||||||
|
Write-Host "Interview demo preflight completed."
|
||||||
|
Write-Host "Verdict: $($summary.verdict)"
|
||||||
|
Write-Host "Gatekeeper rules: $($summary.gatekeeperRuleSetVersion)"
|
||||||
|
Write-Host "Prompt audit: $($summary.promptAuditVersion)"
|
||||||
|
Write-Host "Summary: $summaryPath"
|
||||||
@@ -26,15 +26,22 @@ $chat = Invoke-RestMethod `
|
|||||||
$chat | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-response.json"
|
$chat | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-response.json"
|
||||||
Write-Host "已保存 Chat 响应: $OutputDir/chat-response.json"
|
Write-Host "已保存 Chat 响应: $OutputDir/chat-response.json"
|
||||||
|
|
||||||
|
$runId = $chat.data.runId
|
||||||
|
if (-not $runId) {
|
||||||
|
throw "Chat 响应缺少 runId,无法查询精确 Trace。"
|
||||||
|
}
|
||||||
|
Write-Host "RunId: $runId"
|
||||||
|
|
||||||
$trace = Invoke-RestMethod `
|
$trace = Invoke-RestMethod `
|
||||||
-Method Get `
|
-Method Get `
|
||||||
-Uri "$BaseUrl/api/diagnosis/$SessionId/trace"
|
-Uri "$BaseUrl/api/diagnosis/$SessionId/trace?runId=$([System.Uri]::EscapeDataString($runId))"
|
||||||
|
|
||||||
$trace | ConvertTo-Json -Depth 50 | Set-Content -Encoding UTF8 -Path "$OutputDir/trace-response.json"
|
$trace | ConvertTo-Json -Depth 50 | Set-Content -Encoding UTF8 -Path "$OutputDir/trace-response.json"
|
||||||
Write-Host "已保存 Trace 响应: $OutputDir/trace-response.json"
|
Write-Host "已保存 Trace 响应: $OutputDir/trace-response.json"
|
||||||
|
|
||||||
$feedbackBody = @{
|
$feedbackBody = @{
|
||||||
sessionId = $SessionId
|
sessionId = $SessionId
|
||||||
|
runId = $runId
|
||||||
feedback = "useful"
|
feedback = "useful"
|
||||||
} | ConvertTo-Json
|
} | ConvertTo-Json
|
||||||
|
|
||||||
|
|||||||
@@ -37,7 +37,7 @@ http://localhost:9900
|
|||||||
推荐使用固定脚本:
|
推荐使用固定脚本:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1
|
||||||
```
|
```
|
||||||
|
|
||||||
脚本会写出:
|
脚本会写出:
|
||||||
@@ -46,13 +46,14 @@ powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-de
|
|||||||
mvp/demo/output/chat-response.json
|
mvp/demo/output/chat-response.json
|
||||||
mvp/demo/output/trace-response.json
|
mvp/demo/output/trace-response.json
|
||||||
mvp/demo/output/feedback-response.json
|
mvp/demo/output/feedback-response.json
|
||||||
|
mvp/demo/output/interview-demo-summary.json
|
||||||
```
|
```
|
||||||
|
|
||||||
现场话术:
|
现场话术:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
这里我用固定 sessionId 跑一个支付接口超时问题。
|
这里我用固定 sessionId 跑一个支付接口超时问题。
|
||||||
固定 sessionId 的好处是,后面 trace 和 feedback 都能关联到同一次诊断。
|
固定 sessionId 的好处是保留多轮上下文;每次诊断还会返回 runId,后面 trace 和 feedback 都用这个 runId 精确关联到同一次运行。
|
||||||
```
|
```
|
||||||
|
|
||||||
## 3. 展示用户答案
|
## 3. 展示用户答案
|
||||||
@@ -67,6 +68,7 @@ mvp/demo/output/chat-response.json
|
|||||||
|
|
||||||
```text
|
```text
|
||||||
data.sessionId
|
data.sessionId
|
||||||
|
data.runId
|
||||||
data.answer
|
data.answer
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -75,7 +77,7 @@ data.answer
|
|||||||
```text
|
```text
|
||||||
这是用户看到的答案。
|
这是用户看到的答案。
|
||||||
但这个项目的重点不是这段文字,而是这段文字是否有证据链。
|
但这个项目的重点不是这段文字,而是这段文字是否有证据链。
|
||||||
接下来我用同一个 sessionId 查 trace。
|
接下来我用同一个 sessionId 加 runId 查 trace。
|
||||||
```
|
```
|
||||||
|
|
||||||
## 4. 展示 Trace
|
## 4. 展示 Trace
|
||||||
@@ -98,6 +100,8 @@ data.toolInvocations[*].outputPreview
|
|||||||
data.toolInvocations[*].retrievalLayer
|
data.toolInvocations[*].retrievalLayer
|
||||||
data.toolInvocations[*].relevanceLevel
|
data.toolInvocations[*].relevanceLevel
|
||||||
data.summary.hasVerifierEvaluation
|
data.summary.hasVerifierEvaluation
|
||||||
|
data.session.selfEvaluation.verifier_evaluation.prompt_audit.version
|
||||||
|
data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version
|
||||||
```
|
```
|
||||||
|
|
||||||
现场话术:
|
现场话术:
|
||||||
@@ -155,6 +159,10 @@ Verifier 不做新检索,只看工具 trace 汇总。
|
|||||||
如果 PASS,就输出原答案。
|
如果 PASS,就输出原答案。
|
||||||
如果 LOW_CONFID,可以补证据或加低置信提示。
|
如果 LOW_CONFID,可以补证据或加低置信提示。
|
||||||
如果 REJECT,就降级输出,只保留已确认信息。
|
如果 REJECT,就降级输出,只保留已确认信息。
|
||||||
|
|
||||||
|
Prompt 和 Gatekeeper 的版本也会进入 trace。
|
||||||
|
`prompt_audit.version` 用于说明本次 Chat 使用哪套 Prompt 契约,`gatekeeper_result.rule_set_version` 用于说明引用验真的规则版本。
|
||||||
|
固定 fixture baseline 是回归判断来源,live demo 主要证明当前环境链路可跑通。
|
||||||
```
|
```
|
||||||
|
|
||||||
## 7. 展示反馈闭环
|
## 7. 展示反馈闭环
|
||||||
@@ -175,7 +183,7 @@ caseId
|
|||||||
现场话术:
|
现场话术:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
用户反馈 useful 会写回同一个 diagnosis_session。
|
用户反馈 useful 会写回当前 diagnosis_run。
|
||||||
后端会把这次诊断自动沉淀到 case_library,后续可以做案例检索或 bad case 分析。
|
后端会把这次诊断自动沉淀到 case_library,后续可以做案例检索或 bad case 分析。
|
||||||
|
|
||||||
这里 status 和 feedback 是分开的:
|
这里 status 和 feedback 是分开的:
|
||||||
@@ -226,6 +234,7 @@ AIOps 有两个模式。
|
|||||||
mvp/demo/output/chat-response.json
|
mvp/demo/output/chat-response.json
|
||||||
mvp/demo/output/trace-response.json
|
mvp/demo/output/trace-response.json
|
||||||
mvp/demo/output/feedback-response.json
|
mvp/demo/output/feedback-response.json
|
||||||
|
mvp/demo/output/interview-demo-summary.json
|
||||||
```
|
```
|
||||||
|
|
||||||
降级话术:
|
降级话术:
|
||||||
|
|||||||
@@ -1,17 +1,20 @@
|
|||||||
# Trace 检查清单
|
# Trace 检查清单
|
||||||
|
|
||||||
运行 `scripts/run-payment-timeout-demo.ps1` 后,用这份清单检查 `trace-response.json`。
|
运行 `scripts/run-interview-demo-check.ps1` 后,用这份清单检查 `trace-response.json` 和 `interview-demo-summary.json`。
|
||||||
|
|
||||||
## 1. Session
|
## 1. Session
|
||||||
|
|
||||||
| JSON path | 检查点 | 面试讲点 |
|
| JSON path | 检查点 | 面试讲点 |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| `data.session.sessionId` | 是否等于 `mvp-demo-payment-timeout-001` | 一个 session id 串起 chat、工具、verifier、feedback 和 trace |
|
| `data.runId` / `data.run.runId` | 是否等于 demo 响应中的 `runId` | `runId` 精确绑定这一次诊断运行 |
|
||||||
|
| `data.session.sessionId` | 是否等于 `mvp-demo-payment-timeout-001` | `sessionId` 保留多轮上下文,Trace 精确回放依赖 `runId` |
|
||||||
| `data.session.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
|
| `data.session.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
|
||||||
| `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
|
| `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
|
||||||
| `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
|
| `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
|
||||||
| `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | 如果是 Chat V2 链路,是否记录 Gatekeeper 规则版本 | 安全规则可审计、可回归 |
|
| `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | 如果是 Chat V2 链路,是否记录 Gatekeeper 规则版本 | 安全规则可审计、可回归 |
|
||||||
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在同一次诊断上 |
|
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` | 如果是 Chat V2 链路,是否记录 Prompt 审计版本 | Prompt 变更可解释、可回归 |
|
||||||
|
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.prompts[*].version` | 是否记录 planner / executor / verifier / composer 版本 | 便于定位 Prompt 变更影响 |
|
||||||
|
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在当前 diagnosis run 上 |
|
||||||
|
|
||||||
## 2. Agent 步骤
|
## 2. Agent 步骤
|
||||||
|
|
||||||
@@ -46,10 +49,10 @@
|
|||||||
## 5. 好的结果长什么样
|
## 5. 好的结果长什么样
|
||||||
|
|
||||||
```text
|
```text
|
||||||
同一个 session id
|
同一个 session id + run id
|
||||||
-> 最终答案
|
-> 最终答案
|
||||||
-> 持久化 agent steps
|
-> 持久化 agent steps
|
||||||
-> 持久化 evidence tool calls
|
-> 持久化 evidence tool calls
|
||||||
-> verifier / self-evaluation
|
-> verifier / self-evaluation
|
||||||
-> feedback attached to the same session
|
-> feedback attached to the same run
|
||||||
```
|
```
|
||||||
|
|||||||
+8
-4
@@ -29,10 +29,10 @@ The baseline evaluates saved trace fixtures. It does not start the application a
|
|||||||
The committed baseline currently contains:
|
The committed baseline currently contains:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
10 fixed cases
|
12 fixed cases
|
||||||
10 passing fixture evaluations
|
12 passing fixture evaluations
|
||||||
4 PASS verdicts
|
5 PASS verdicts
|
||||||
5 LOW_CONFID verdicts
|
6 LOW_CONFID verdicts
|
||||||
1 REJECT verdict
|
1 REJECT verdict
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -44,6 +44,8 @@ The V2 evidence-pipeline matrix covers:
|
|||||||
- Unsupported claim filtering before the final answer.
|
- Unsupported claim filtering before the final answer.
|
||||||
- Composer fallback rendering without raw Executor JSON leakage.
|
- Composer fallback rendering without raw Executor JSON leakage.
|
||||||
- Gatekeeper rule set version audit for new matrix fixtures.
|
- Gatekeeper rule set version audit for new matrix fixtures.
|
||||||
|
- Prompt audit version checks for planner, executor, verifier, and composer prompts.
|
||||||
|
- Gatekeeper rule metadata checks for enabled rule id and default severity.
|
||||||
|
|
||||||
## Verification
|
## Verification
|
||||||
|
|
||||||
@@ -80,6 +82,8 @@ Stage 5 adds these V2 checks:
|
|||||||
- `claim_checks` must be structurally auditable.
|
- `claim_checks` must be structurally auditable.
|
||||||
- Composer output must record whether normal parsing or fallback rendering was used.
|
- Composer output must record whether normal parsing or fallback rendering was used.
|
||||||
- Gatekeeper rule set version can be asserted per fixture.
|
- Gatekeeper rule set version can be asserted per fixture.
|
||||||
|
- Prompt audit version and per-prompt versions can be asserted per fixture.
|
||||||
|
- Gatekeeper rule metadata can be required per fixture.
|
||||||
- Final answers must not leak raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`.
|
- Final answers must not leak raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`.
|
||||||
- Configured unsupported claim keywords must not appear as confirmed final-answer content.
|
- Configured unsupported claim keywords must not appear as confirmed final-answer content.
|
||||||
|
|
||||||
|
|||||||
@@ -17,6 +17,33 @@
|
|||||||
"expectedComposerStatuses": ["valid"],
|
"expectedComposerStatuses": ["valid"],
|
||||||
"forbiddenConfirmedClaimKeywords": ["数据库连接池"]
|
"forbiddenConfirmedClaimKeywords": ["数据库连接池"]
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
"id": "prompt-gatekeeper-audit-closure",
|
||||||
|
"title": "Prompt and Gatekeeper audit closure",
|
||||||
|
"question": "确认 payment-service 当前是否存在 HighCPUUsage 告警,并检查审计元数据是否完整。",
|
||||||
|
"traceFixture": "prompt-gatekeeper-audit-closure-pass.json",
|
||||||
|
"expectedRootCauseKeywords": ["payment-service", "HighCPUUsage", "92%"],
|
||||||
|
"minKeywordMatches": 2,
|
||||||
|
"requiredEvidenceTools": ["query_metrics"],
|
||||||
|
"allowedVerdicts": ["PASS"],
|
||||||
|
"forbiddenAnswerKeywords": ["根因", "修复建议", "通常情况下"],
|
||||||
|
"requireV2AuditClosure": true,
|
||||||
|
"requireClaimChecks": true,
|
||||||
|
"requireComposerOutput": true,
|
||||||
|
"expectedGatekeeperStatuses": ["pass"],
|
||||||
|
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
|
||||||
|
"expectedComposerStatuses": ["valid"],
|
||||||
|
"forbiddenConfirmedClaimKeywords": ["数据库连接池"],
|
||||||
|
"requirePromptAudit": true,
|
||||||
|
"expectedPromptAuditVersion": "chat-prompts-v1",
|
||||||
|
"expectedPromptVersions": {
|
||||||
|
"chat_planner": "chat-planner-v1",
|
||||||
|
"chat_executor": "chat-executor-v2",
|
||||||
|
"chat_verifier": "chat-verifier-v2",
|
||||||
|
"chat_composer": "chat-composer-v1"
|
||||||
|
},
|
||||||
|
"requireGatekeeperRules": true
|
||||||
|
},
|
||||||
{
|
{
|
||||||
"id": "hikari-no-evidence-negative-observation",
|
"id": "hikari-no-evidence-negative-observation",
|
||||||
"title": "Hikari no-evidence negative observation",
|
"title": "Hikari no-evidence negative observation",
|
||||||
@@ -124,6 +151,33 @@
|
|||||||
"expectedComposerStatuses": ["valid"],
|
"expectedComposerStatuses": ["valid"],
|
||||||
"forbiddenConfirmedClaimKeywords": ["主库故障"]
|
"forbiddenConfirmedClaimKeywords": ["主库故障"]
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
"id": "audit-metadata-low-confid",
|
||||||
|
"title": "Audit metadata low confidence",
|
||||||
|
"question": "订单超时是否可以确认由数据库主库故障导致,并检查审计元数据是否完整?",
|
||||||
|
"traceFixture": "audit-metadata-low-confid.json",
|
||||||
|
"expectedRootCauseKeywords": ["超时", "证据"],
|
||||||
|
"minKeywordMatches": 2,
|
||||||
|
"requiredEvidenceTools": ["query_logs"],
|
||||||
|
"allowedVerdicts": ["LOW_CONFID"],
|
||||||
|
"forbiddenAnswerKeywords": ["已经确认"],
|
||||||
|
"requireV2AuditClosure": true,
|
||||||
|
"requireClaimChecks": true,
|
||||||
|
"requireComposerOutput": true,
|
||||||
|
"expectedGatekeeperStatuses": ["pass"],
|
||||||
|
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
|
||||||
|
"expectedComposerStatuses": ["valid"],
|
||||||
|
"forbiddenConfirmedClaimKeywords": ["主库故障"],
|
||||||
|
"requirePromptAudit": true,
|
||||||
|
"expectedPromptAuditVersion": "chat-prompts-v1",
|
||||||
|
"expectedPromptVersions": {
|
||||||
|
"chat_planner": "chat-planner-v1",
|
||||||
|
"chat_executor": "chat-executor-v2",
|
||||||
|
"chat_verifier": "chat-verifier-v2",
|
||||||
|
"chat_composer": "chat-composer-v1"
|
||||||
|
},
|
||||||
|
"requireGatekeeperRules": true
|
||||||
|
},
|
||||||
{
|
{
|
||||||
"id": "composer-fallback-no-raw-json",
|
"id": "composer-fallback-no-raw-json",
|
||||||
"title": "Composer fallback no raw JSON",
|
"title": "Composer fallback no raw JSON",
|
||||||
|
|||||||
@@ -0,0 +1,192 @@
|
|||||||
|
{
|
||||||
|
"session": {
|
||||||
|
"sessionId": "eval-audit-metadata-low-confid",
|
||||||
|
"query": "订单超时是否可以确认由数据库主库故障导致,并检查审计元数据是否完整?",
|
||||||
|
"status": "SUCCESS",
|
||||||
|
"agentFlow": "CHAT",
|
||||||
|
"totalDurationMs": 45000,
|
||||||
|
"toolCallCount": 1,
|
||||||
|
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:日志显示订单接口出现超时。\n\n仍需补充信息:目前没有数据库故障日志或主库状态证据,不能把该方向写成确认根因。",
|
||||||
|
"selfEvaluation": {
|
||||||
|
"verifier_evaluation": {
|
||||||
|
"verdict": "LOW_CONFID",
|
||||||
|
"groundedness_score": 0.42,
|
||||||
|
"critical_fact_count": 2,
|
||||||
|
"prompt_audit": {
|
||||||
|
"version": "chat-prompts-v1",
|
||||||
|
"prompts": [
|
||||||
|
{
|
||||||
|
"name": "chat_planner",
|
||||||
|
"version": "chat-planner-v1",
|
||||||
|
"resource": "prompts/chat-planner-prompt.md"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "chat_executor",
|
||||||
|
"version": "chat-executor-v2",
|
||||||
|
"resource": "prompts/chat-executor-prompt.md"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "chat_verifier",
|
||||||
|
"version": "chat-verifier-v2",
|
||||||
|
"resource": "prompts/chat-verifier-prompt.md"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "chat_composer",
|
||||||
|
"version": "chat-composer-v1",
|
||||||
|
"resource": "prompts/chat-composer-prompt.md"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"gatekeeper_result": {
|
||||||
|
"status": "pass",
|
||||||
|
"severity": "none",
|
||||||
|
"rule_set_version": "gatekeeper-rules-v1",
|
||||||
|
"rules": [
|
||||||
|
{
|
||||||
|
"id": "evidence.invocation",
|
||||||
|
"description": "source_invocation_id must reference an existing tool invocation",
|
||||||
|
"enabled": true,
|
||||||
|
"default_severity": "reject"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "evidence.excerpt",
|
||||||
|
"description": "evidence_excerpt must be supported by recorded evidence text",
|
||||||
|
"enabled": true,
|
||||||
|
"default_severity": "reject"
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"checked_bindings": [
|
||||||
|
{
|
||||||
|
"claim_id": "claim-timeout",
|
||||||
|
"tool_name": "query_logs",
|
||||||
|
"source_invocation_id": 22,
|
||||||
|
"raw_path": "$.logs[0]",
|
||||||
|
"matched_text": "order api timeout",
|
||||||
|
"status": "pass"
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"failed_rules": [],
|
||||||
|
"warnings": [],
|
||||||
|
"errors": []
|
||||||
|
},
|
||||||
|
"executor_structured_output": {
|
||||||
|
"answer_version": "executor_evidence_v2",
|
||||||
|
"claims": [
|
||||||
|
{
|
||||||
|
"claim_id": "claim-timeout",
|
||||||
|
"claim_type": "symptom",
|
||||||
|
"claim_text": "订单接口出现超时",
|
||||||
|
"support_level": "direct",
|
||||||
|
"evidence_bindings": [
|
||||||
|
{
|
||||||
|
"source_type": "tool_trace",
|
||||||
|
"tool_name": "query_logs",
|
||||||
|
"source_invocation_id": 22,
|
||||||
|
"raw_path": "$.logs[0]",
|
||||||
|
"evidence_excerpt": "order api timeout"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"claim_id": "claim-db-primary",
|
||||||
|
"claim_type": "root_cause",
|
||||||
|
"claim_text": "数据库主库故障导致订单超时",
|
||||||
|
"support_level": "weak",
|
||||||
|
"evidence_bindings": [
|
||||||
|
{
|
||||||
|
"source_type": "tool_trace",
|
||||||
|
"tool_name": "query_logs",
|
||||||
|
"source_invocation_id": 22,
|
||||||
|
"raw_path": "$.logs[0]",
|
||||||
|
"evidence_excerpt": "order api timeout"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"hypotheses": [],
|
||||||
|
"recommended_actions": [
|
||||||
|
{
|
||||||
|
"action_text": "补充查询数据库主库状态和错误日志",
|
||||||
|
"reason": "当前只有订单接口超时日志"
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"missing_info": ["数据库主库状态", "数据库错误日志"]
|
||||||
|
},
|
||||||
|
"claim_checks": [
|
||||||
|
{
|
||||||
|
"claim_id": "claim-timeout",
|
||||||
|
"claim_text": "订单接口出现超时",
|
||||||
|
"claim_type": "symptom",
|
||||||
|
"verification": "direct_observation",
|
||||||
|
"detail": "日志直接记录 order api timeout",
|
||||||
|
"evidence_refs": [
|
||||||
|
{
|
||||||
|
"source_invocation_id": 22,
|
||||||
|
"raw_path": "$.logs[0]"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"claim_id": "claim-db-primary",
|
||||||
|
"claim_text": "数据库主库故障导致订单超时",
|
||||||
|
"claim_type": "root_cause",
|
||||||
|
"verification": "unsupported",
|
||||||
|
"detail": "日志只能证明订单接口超时,不能证明数据库主库故障",
|
||||||
|
"evidence_refs": [
|
||||||
|
{
|
||||||
|
"source_invocation_id": 22,
|
||||||
|
"raw_path": "$.logs[0]"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"facts_checked": [],
|
||||||
|
"composer_output": {
|
||||||
|
"status": "valid",
|
||||||
|
"answer_summary": "日志显示订单接口超时,但数据库方向证据不足。",
|
||||||
|
"recommended_actions": [
|
||||||
|
{
|
||||||
|
"action_text": "补充查询数据库主库状态和错误日志",
|
||||||
|
"reason": "当前只有订单接口超时日志"
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"user_facing_answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:日志显示订单接口出现超时。\n\n仍需补充信息:目前没有数据库故障日志或主库状态证据,不能把该方向写成确认根因。"
|
||||||
|
},
|
||||||
|
"tool_trace_summary": [
|
||||||
|
{
|
||||||
|
"tool_name": "query_logs",
|
||||||
|
"success": true,
|
||||||
|
"evidence_level": "direct"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"steps": [],
|
||||||
|
"toolInvocations": [
|
||||||
|
{
|
||||||
|
"id": 22,
|
||||||
|
"sessionId": "eval-audit-metadata-low-confid",
|
||||||
|
"toolName": "query_logs",
|
||||||
|
"outputPreview": "order api timeout",
|
||||||
|
"retrievalDetails": {
|
||||||
|
"evidence_status": "supported",
|
||||||
|
"evidence_refs": [
|
||||||
|
{
|
||||||
|
"raw_path": "$.logs[0]",
|
||||||
|
"text": "order api timeout"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"success": true
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"summary": {
|
||||||
|
"persistedStepCount": 3,
|
||||||
|
"returnedStepCount": 3,
|
||||||
|
"persistedToolCallCount": 1,
|
||||||
|
"returnedToolCallCount": 1,
|
||||||
|
"hasVerifierEvaluation": true,
|
||||||
|
"hasFeedback": false
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,155 @@
|
|||||||
|
{
|
||||||
|
"session": {
|
||||||
|
"sessionId": "eval-prompt-gatekeeper-audit-closure",
|
||||||
|
"query": "确认 payment-service 当前是否存在 HighCPUUsage 告警,并检查审计元数据是否完整。",
|
||||||
|
"status": "SUCCESS",
|
||||||
|
"agentFlow": "CHAT",
|
||||||
|
"totalDurationMs": 19000,
|
||||||
|
"toolCallCount": 1,
|
||||||
|
"answer": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
|
||||||
|
"selfEvaluation": {
|
||||||
|
"verifier_evaluation": {
|
||||||
|
"verdict": "PASS",
|
||||||
|
"groundedness_score": 1.0,
|
||||||
|
"critical_fact_count": 1,
|
||||||
|
"prompt_audit": {
|
||||||
|
"version": "chat-prompts-v1",
|
||||||
|
"prompts": [
|
||||||
|
{
|
||||||
|
"name": "chat_planner",
|
||||||
|
"version": "chat-planner-v1",
|
||||||
|
"resource": "prompts/chat-planner-prompt.md"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "chat_executor",
|
||||||
|
"version": "chat-executor-v2",
|
||||||
|
"resource": "prompts/chat-executor-prompt.md"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "chat_verifier",
|
||||||
|
"version": "chat-verifier-v2",
|
||||||
|
"resource": "prompts/chat-verifier-prompt.md"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "chat_composer",
|
||||||
|
"version": "chat-composer-v1",
|
||||||
|
"resource": "prompts/chat-composer-prompt.md"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"gatekeeper_result": {
|
||||||
|
"status": "pass",
|
||||||
|
"severity": "none",
|
||||||
|
"rule_set_version": "gatekeeper-rules-v1",
|
||||||
|
"rules": [
|
||||||
|
{
|
||||||
|
"id": "evidence.invocation",
|
||||||
|
"description": "source_invocation_id must reference an existing tool invocation",
|
||||||
|
"enabled": true,
|
||||||
|
"default_severity": "reject"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "evidence.raw_path",
|
||||||
|
"description": "raw_path must exist in retrieval_details.evidence_refs",
|
||||||
|
"enabled": true,
|
||||||
|
"default_severity": "reject"
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"checked_bindings": [
|
||||||
|
{
|
||||||
|
"claim_id": "claim-1",
|
||||||
|
"tool_name": "query_metrics",
|
||||||
|
"source_invocation_id": 21,
|
||||||
|
"raw_path": "$.alerts[0]",
|
||||||
|
"matched_text": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m",
|
||||||
|
"status": "pass"
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"failed_rules": [],
|
||||||
|
"warnings": [],
|
||||||
|
"errors": []
|
||||||
|
},
|
||||||
|
"executor_structured_output": {
|
||||||
|
"answer_version": "executor_evidence_v2",
|
||||||
|
"claims": [
|
||||||
|
{
|
||||||
|
"claim_id": "claim-1",
|
||||||
|
"claim_type": "observation",
|
||||||
|
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
|
||||||
|
"support_level": "direct",
|
||||||
|
"evidence_bindings": [
|
||||||
|
{
|
||||||
|
"source_type": "tool_trace",
|
||||||
|
"tool_name": "query_metrics",
|
||||||
|
"source_invocation_id": 21,
|
||||||
|
"raw_path": "$.alerts[0]",
|
||||||
|
"evidence_excerpt": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"hypotheses": [],
|
||||||
|
"recommended_actions": [],
|
||||||
|
"missing_info": []
|
||||||
|
},
|
||||||
|
"claim_checks": [
|
||||||
|
{
|
||||||
|
"claim_id": "claim-1",
|
||||||
|
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
|
||||||
|
"claim_type": "observation",
|
||||||
|
"verification": "direct_observation",
|
||||||
|
"detail": "已核验的指标证据直接包含服务名、告警名和 CPU 当前值。",
|
||||||
|
"evidence_refs": [
|
||||||
|
{
|
||||||
|
"source_invocation_id": 21,
|
||||||
|
"raw_path": "$.alerts[0]"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"facts_checked": [],
|
||||||
|
"composer_output": {
|
||||||
|
"status": "valid",
|
||||||
|
"answer_summary": "payment-service 当前存在 HighCPUUsage 告警。",
|
||||||
|
"recommended_actions": [],
|
||||||
|
"user_facing_answer": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。"
|
||||||
|
},
|
||||||
|
"tool_trace_summary": [
|
||||||
|
{
|
||||||
|
"tool_name": "query_metrics",
|
||||||
|
"success": true,
|
||||||
|
"source_invocation_ids": [21],
|
||||||
|
"evidence_level": "direct"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"steps": [],
|
||||||
|
"toolInvocations": [
|
||||||
|
{
|
||||||
|
"id": 21,
|
||||||
|
"sessionId": "eval-prompt-gatekeeper-audit-closure",
|
||||||
|
"toolName": "query_metrics",
|
||||||
|
"outputPreview": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m",
|
||||||
|
"retrievalDetails": {
|
||||||
|
"evidence_status": "supported",
|
||||||
|
"evidence_refs": [
|
||||||
|
{
|
||||||
|
"raw_path": "$.alerts[0]",
|
||||||
|
"text": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"success": true
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"summary": {
|
||||||
|
"persistedStepCount": 3,
|
||||||
|
"returnedStepCount": 3,
|
||||||
|
"persistedToolCallCount": 1,
|
||||||
|
"returnedToolCallCount": 1,
|
||||||
|
"hasVerifierEvaluation": true,
|
||||||
|
"hasFeedback": false
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -1,14 +1,14 @@
|
|||||||
{
|
{
|
||||||
"totalCases" : 10,
|
"totalCases" : 12,
|
||||||
"passedCases" : 10,
|
"passedCases" : 12,
|
||||||
"passRate" : 1.0,
|
"passRate" : 1.0,
|
||||||
"verdictDistribution" : {
|
"verdictDistribution" : {
|
||||||
"PASS" : 4,
|
"PASS" : 5,
|
||||||
"LOW_CONFID" : 5,
|
"LOW_CONFID" : 6,
|
||||||
"REJECT" : 1
|
"REJECT" : 1
|
||||||
},
|
},
|
||||||
"averageToolCallCount" : 1.5,
|
"averageToolCallCount" : 1.4166666666666667,
|
||||||
"averageDurationMs" : 39800.0,
|
"averageDurationMs" : 38500.0,
|
||||||
"results" : [ {
|
"results" : [ {
|
||||||
"caseId" : "narrow-highcpu-observation",
|
"caseId" : "narrow-highcpu-observation",
|
||||||
"title" : "Narrow HighCPU observation",
|
"title" : "Narrow HighCPU observation",
|
||||||
@@ -22,10 +22,31 @@
|
|||||||
},
|
},
|
||||||
"gatekeeperStatus" : "pass",
|
"gatekeeperStatus" : "pass",
|
||||||
"gatekeeperRuleSetVersion" : "gatekeeper-rules-v1",
|
"gatekeeperRuleSetVersion" : "gatekeeper-rules-v1",
|
||||||
|
"promptAuditVersion" : null,
|
||||||
"composerStatus" : "valid",
|
"composerStatus" : "valid",
|
||||||
"claimCheckCount" : 1,
|
"claimCheckCount" : 1,
|
||||||
|
"gatekeeperRuleCount" : 1,
|
||||||
"toolCallCount" : 1,
|
"toolCallCount" : 1,
|
||||||
"durationMs" : 18000
|
"durationMs" : 18000
|
||||||
|
}, {
|
||||||
|
"caseId" : "prompt-gatekeeper-audit-closure",
|
||||||
|
"title" : "Prompt and Gatekeeper audit closure",
|
||||||
|
"passed" : true,
|
||||||
|
"failedChecks" : [ ],
|
||||||
|
"verdict" : "PASS",
|
||||||
|
"matchedKeywordCount" : 3,
|
||||||
|
"requiredKeywordCount" : 3,
|
||||||
|
"evidenceCoverage" : {
|
||||||
|
"query_metrics" : true
|
||||||
|
},
|
||||||
|
"gatekeeperStatus" : "pass",
|
||||||
|
"gatekeeperRuleSetVersion" : "gatekeeper-rules-v1",
|
||||||
|
"promptAuditVersion" : "chat-prompts-v1",
|
||||||
|
"composerStatus" : "valid",
|
||||||
|
"claimCheckCount" : 1,
|
||||||
|
"gatekeeperRuleCount" : 2,
|
||||||
|
"toolCallCount" : 1,
|
||||||
|
"durationMs" : 19000
|
||||||
}, {
|
}, {
|
||||||
"caseId" : "hikari-no-evidence-negative-observation",
|
"caseId" : "hikari-no-evidence-negative-observation",
|
||||||
"title" : "Hikari no-evidence negative observation",
|
"title" : "Hikari no-evidence negative observation",
|
||||||
@@ -39,8 +60,10 @@
|
|||||||
},
|
},
|
||||||
"gatekeeperStatus" : "pass",
|
"gatekeeperStatus" : "pass",
|
||||||
"gatekeeperRuleSetVersion" : "gatekeeper-rules-v1",
|
"gatekeeperRuleSetVersion" : "gatekeeper-rules-v1",
|
||||||
|
"promptAuditVersion" : null,
|
||||||
"composerStatus" : "valid",
|
"composerStatus" : "valid",
|
||||||
"claimCheckCount" : 1,
|
"claimCheckCount" : 1,
|
||||||
|
"gatekeeperRuleCount" : 1,
|
||||||
"toolCallCount" : 1,
|
"toolCallCount" : 1,
|
||||||
"durationMs" : 21000
|
"durationMs" : 21000
|
||||||
}, {
|
}, {
|
||||||
@@ -58,8 +81,10 @@
|
|||||||
},
|
},
|
||||||
"gatekeeperStatus" : null,
|
"gatekeeperStatus" : null,
|
||||||
"gatekeeperRuleSetVersion" : null,
|
"gatekeeperRuleSetVersion" : null,
|
||||||
|
"promptAuditVersion" : null,
|
||||||
"composerStatus" : null,
|
"composerStatus" : null,
|
||||||
"claimCheckCount" : null,
|
"claimCheckCount" : null,
|
||||||
|
"gatekeeperRuleCount" : null,
|
||||||
"toolCallCount" : 3,
|
"toolCallCount" : 3,
|
||||||
"durationMs" : 42000
|
"durationMs" : 42000
|
||||||
}, {
|
}, {
|
||||||
@@ -76,8 +101,10 @@
|
|||||||
},
|
},
|
||||||
"gatekeeperStatus" : null,
|
"gatekeeperStatus" : null,
|
||||||
"gatekeeperRuleSetVersion" : null,
|
"gatekeeperRuleSetVersion" : null,
|
||||||
|
"promptAuditVersion" : null,
|
||||||
"composerStatus" : null,
|
"composerStatus" : null,
|
||||||
"claimCheckCount" : null,
|
"claimCheckCount" : null,
|
||||||
|
"gatekeeperRuleCount" : null,
|
||||||
"toolCallCount" : 2,
|
"toolCallCount" : 2,
|
||||||
"durationMs" : 51000
|
"durationMs" : 51000
|
||||||
}, {
|
}, {
|
||||||
@@ -93,8 +120,10 @@
|
|||||||
},
|
},
|
||||||
"gatekeeperStatus" : null,
|
"gatekeeperStatus" : null,
|
||||||
"gatekeeperRuleSetVersion" : null,
|
"gatekeeperRuleSetVersion" : null,
|
||||||
|
"promptAuditVersion" : null,
|
||||||
"composerStatus" : null,
|
"composerStatus" : null,
|
||||||
"claimCheckCount" : null,
|
"claimCheckCount" : null,
|
||||||
|
"gatekeeperRuleCount" : null,
|
||||||
"toolCallCount" : 1,
|
"toolCallCount" : 1,
|
||||||
"durationMs" : 36000
|
"durationMs" : 36000
|
||||||
}, {
|
}, {
|
||||||
@@ -111,8 +140,10 @@
|
|||||||
},
|
},
|
||||||
"gatekeeperStatus" : null,
|
"gatekeeperStatus" : null,
|
||||||
"gatekeeperRuleSetVersion" : null,
|
"gatekeeperRuleSetVersion" : null,
|
||||||
|
"promptAuditVersion" : null,
|
||||||
"composerStatus" : null,
|
"composerStatus" : null,
|
||||||
"claimCheckCount" : null,
|
"claimCheckCount" : null,
|
||||||
|
"gatekeeperRuleCount" : null,
|
||||||
"toolCallCount" : 2,
|
"toolCallCount" : 2,
|
||||||
"durationMs" : 47000
|
"durationMs" : 47000
|
||||||
}, {
|
}, {
|
||||||
@@ -129,8 +160,10 @@
|
|||||||
},
|
},
|
||||||
"gatekeeperStatus" : null,
|
"gatekeeperStatus" : null,
|
||||||
"gatekeeperRuleSetVersion" : null,
|
"gatekeeperRuleSetVersion" : null,
|
||||||
|
"promptAuditVersion" : null,
|
||||||
"composerStatus" : null,
|
"composerStatus" : null,
|
||||||
"claimCheckCount" : null,
|
"claimCheckCount" : null,
|
||||||
|
"gatekeeperRuleCount" : null,
|
||||||
"toolCallCount" : 2,
|
"toolCallCount" : 2,
|
||||||
"durationMs" : 53000
|
"durationMs" : 53000
|
||||||
}, {
|
}, {
|
||||||
@@ -146,8 +179,10 @@
|
|||||||
},
|
},
|
||||||
"gatekeeperStatus" : "fail",
|
"gatekeeperStatus" : "fail",
|
||||||
"gatekeeperRuleSetVersion" : null,
|
"gatekeeperRuleSetVersion" : null,
|
||||||
|
"promptAuditVersion" : null,
|
||||||
"composerStatus" : "valid",
|
"composerStatus" : "valid",
|
||||||
"claimCheckCount" : 1,
|
"claimCheckCount" : 1,
|
||||||
|
"gatekeeperRuleCount" : null,
|
||||||
"toolCallCount" : 1,
|
"toolCallCount" : 1,
|
||||||
"durationMs" : 39000
|
"durationMs" : 39000
|
||||||
}, {
|
}, {
|
||||||
@@ -163,10 +198,31 @@
|
|||||||
},
|
},
|
||||||
"gatekeeperStatus" : "pass",
|
"gatekeeperStatus" : "pass",
|
||||||
"gatekeeperRuleSetVersion" : null,
|
"gatekeeperRuleSetVersion" : null,
|
||||||
|
"promptAuditVersion" : null,
|
||||||
"composerStatus" : "valid",
|
"composerStatus" : "valid",
|
||||||
"claimCheckCount" : 2,
|
"claimCheckCount" : 2,
|
||||||
|
"gatekeeperRuleCount" : null,
|
||||||
"toolCallCount" : 1,
|
"toolCallCount" : 1,
|
||||||
"durationMs" : 44000
|
"durationMs" : 44000
|
||||||
|
}, {
|
||||||
|
"caseId" : "audit-metadata-low-confid",
|
||||||
|
"title" : "Audit metadata low confidence",
|
||||||
|
"passed" : true,
|
||||||
|
"failedChecks" : [ ],
|
||||||
|
"verdict" : "LOW_CONFID",
|
||||||
|
"matchedKeywordCount" : 2,
|
||||||
|
"requiredKeywordCount" : 2,
|
||||||
|
"evidenceCoverage" : {
|
||||||
|
"query_logs" : true
|
||||||
|
},
|
||||||
|
"gatekeeperStatus" : "pass",
|
||||||
|
"gatekeeperRuleSetVersion" : "gatekeeper-rules-v1",
|
||||||
|
"promptAuditVersion" : "chat-prompts-v1",
|
||||||
|
"composerStatus" : "valid",
|
||||||
|
"claimCheckCount" : 2,
|
||||||
|
"gatekeeperRuleCount" : 2,
|
||||||
|
"toolCallCount" : 1,
|
||||||
|
"durationMs" : 45000
|
||||||
}, {
|
}, {
|
||||||
"caseId" : "composer-fallback-no-raw-json",
|
"caseId" : "composer-fallback-no-raw-json",
|
||||||
"title" : "Composer fallback no raw JSON",
|
"title" : "Composer fallback no raw JSON",
|
||||||
@@ -180,8 +236,10 @@
|
|||||||
},
|
},
|
||||||
"gatekeeperStatus" : "pass",
|
"gatekeeperStatus" : "pass",
|
||||||
"gatekeeperRuleSetVersion" : null,
|
"gatekeeperRuleSetVersion" : null,
|
||||||
|
"promptAuditVersion" : null,
|
||||||
"composerStatus" : "composer_malformed",
|
"composerStatus" : "composer_malformed",
|
||||||
"claimCheckCount" : 2,
|
"claimCheckCount" : 2,
|
||||||
|
"gatekeeperRuleCount" : null,
|
||||||
"toolCallCount" : 1,
|
"toolCallCount" : 1,
|
||||||
"durationMs" : 47000
|
"durationMs" : 47000
|
||||||
} ]
|
} ]
|
||||||
|
|||||||
@@ -1,28 +1,30 @@
|
|||||||
# Diagnosis Eval Report
|
# Diagnosis Eval Report
|
||||||
|
|
||||||
- Total cases: 10
|
- Total cases: 12
|
||||||
- Passed cases: 10
|
- Passed cases: 12
|
||||||
- Pass rate: 100.00%
|
- Pass rate: 100.00%
|
||||||
- Average tool calls: 1.50
|
- Average tool calls: 1.42
|
||||||
- Average duration ms: 39800.00
|
- Average duration ms: 38500.00
|
||||||
|
|
||||||
## Verdict Distribution
|
## Verdict Distribution
|
||||||
|
|
||||||
- PASS: 4
|
- PASS: 5
|
||||||
- LOW_CONFID: 5
|
- LOW_CONFID: 6
|
||||||
- REJECT: 1
|
- REJECT: 1
|
||||||
|
|
||||||
## Cases
|
## Cases
|
||||||
|
|
||||||
| Case | Result | Verdict | Gatekeeper | Rule Set | Composer | Claim Checks | Keywords | Tool Calls | Duration ms | Failed Checks |
|
| Case | Result | Verdict | Gatekeeper | Rule Set | Prompt Audit | Composer | Claim Checks | Rules | Keywords | Tool Calls | Duration ms | Failed Checks |
|
||||||
| --- | --- | --- | --- | --- | --- | ---: | --- | ---: | ---: | --- |
|
| --- | --- | --- | --- | --- | --- | --- | ---: | ---: | --- | ---: | ---: | --- |
|
||||||
| narrow-highcpu-observation | PASS | PASS | pass | gatekeeper-rules-v1 | valid | 1 | 3/3 | 1 | 18000 | - |
|
| narrow-highcpu-observation | PASS | PASS | pass | gatekeeper-rules-v1 | - | valid | 1 | 1 | 3/3 | 1 | 18000 | - |
|
||||||
| hikari-no-evidence-negative-observation | PASS | PASS | pass | gatekeeper-rules-v1 | valid | 1 | 3/3 | 1 | 21000 | - |
|
| prompt-gatekeeper-audit-closure | PASS | PASS | pass | gatekeeper-rules-v1 | chat-prompts-v1 | valid | 1 | 2 | 3/3 | 1 | 19000 | - |
|
||||||
| payment-timeout | PASS | PASS | - | - | - | - | 3/3 | 3 | 42000 | - |
|
| hikari-no-evidence-negative-observation | PASS | PASS | pass | gatekeeper-rules-v1 | - | valid | 1 | 1 | 3/3 | 1 | 21000 | - |
|
||||||
| mysql-pool-exhausted | PASS | LOW_CONFID | - | - | - | - | 3/3 | 2 | 51000 | - |
|
| payment-timeout | PASS | PASS | - | - | - | - | - | - | 3/3 | 3 | 42000 | - |
|
||||||
| redis-timeout | PASS | LOW_CONFID | - | - | - | - | 2/2 | 1 | 36000 | - |
|
| mysql-pool-exhausted | PASS | LOW_CONFID | - | - | - | - | - | - | 3/3 | 2 | 51000 | - |
|
||||||
| slow-response | PASS | PASS | - | - | - | - | 2/2 | 2 | 47000 | - |
|
| redis-timeout | PASS | LOW_CONFID | - | - | - | - | - | - | 2/2 | 1 | 36000 | - |
|
||||||
| jvm-memory-risk | PASS | LOW_CONFID | - | - | - | - | 3/3 | 2 | 53000 | - |
|
| slow-response | PASS | PASS | - | - | - | - | - | - | 2/2 | 2 | 47000 | - |
|
||||||
| gatekeeper-fabricated-invocation | PASS | REJECT | fail | - | valid | 1 | 3/3 | 1 | 39000 | - |
|
| jvm-memory-risk | PASS | LOW_CONFID | - | - | - | - | - | - | 3/3 | 2 | 53000 | - |
|
||||||
| unsupported-claim-filtering | PASS | LOW_CONFID | pass | - | valid | 2 | 2/2 | 1 | 44000 | - |
|
| gatekeeper-fabricated-invocation | PASS | REJECT | fail | - | - | valid | 1 | - | 3/3 | 1 | 39000 | - |
|
||||||
| composer-fallback-no-raw-json | PASS | LOW_CONFID | pass | - | composer_malformed | 2 | 2/2 | 1 | 47000 | - |
|
| unsupported-claim-filtering | PASS | LOW_CONFID | pass | - | - | valid | 2 | - | 2/2 | 1 | 44000 | - |
|
||||||
|
| audit-metadata-low-confid | PASS | LOW_CONFID | pass | gatekeeper-rules-v1 | chat-prompts-v1 | valid | 2 | 2 | 2/2 | 1 | 45000 | - |
|
||||||
|
| composer-fallback-no-raw-json | PASS | LOW_CONFID | pass | - | - | composer_malformed | 2 | - | 2/2 | 1 | 47000 | - |
|
||||||
|
|||||||
+23
-1
@@ -34,7 +34,16 @@ baseline report:整套固定集当前认可的结果
|
|||||||
"expectedGatekeeperStatuses": ["pass"],
|
"expectedGatekeeperStatuses": ["pass"],
|
||||||
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
|
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
|
||||||
"expectedComposerStatuses": ["valid"],
|
"expectedComposerStatuses": ["valid"],
|
||||||
"forbiddenConfirmedClaimKeywords": ["主库故障"]
|
"forbiddenConfirmedClaimKeywords": ["主库故障"],
|
||||||
|
"requirePromptAudit": true,
|
||||||
|
"expectedPromptAuditVersion": "chat-prompts-v1",
|
||||||
|
"expectedPromptVersions": {
|
||||||
|
"chat_planner": "chat-planner-v1",
|
||||||
|
"chat_executor": "chat-executor-v2",
|
||||||
|
"chat_verifier": "chat-verifier-v2",
|
||||||
|
"chat_composer": "chat-composer-v1"
|
||||||
|
},
|
||||||
|
"requireGatekeeperRules": true
|
||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -58,6 +67,10 @@ baseline report:整套固定集当前认可的结果
|
|||||||
| `expectedGatekeeperRuleSetVersion` | 期望的 Gatekeeper 规则集版本 | 配置后校验 `gatekeeper_result.rule_set_version` |
|
| `expectedGatekeeperRuleSetVersion` | 期望的 Gatekeeper 规则集版本 | 配置后校验 `gatekeeper_result.rule_set_version` |
|
||||||
| `expectedComposerStatuses` | 允许的 Composer 状态 | 实际 `composer_output.status` 不在列表中则失败 |
|
| `expectedComposerStatuses` | 允许的 Composer 状态 | 实际 `composer_output.status` 不在列表中则失败 |
|
||||||
| `forbiddenConfirmedClaimKeywords` | 不得进入最终答案的未支持结论关键词 | 用于证明 unsupported/external_unknown claim 被过滤 |
|
| `forbiddenConfirmedClaimKeywords` | 不得进入最终答案的未支持结论关键词 | 用于证明 unsupported/external_unknown claim 被过滤 |
|
||||||
|
| `requirePromptAudit` | 是否要求 Prompt 审计元数据 | 要求 `prompt_audit.version` 存在 |
|
||||||
|
| `expectedPromptAuditVersion` | 期望的 Prompt 审计目录版本 | 配置后校验 `prompt_audit.version` |
|
||||||
|
| `expectedPromptVersions` | 期望的各角色 Prompt 版本 | 校验 `prompt_audit.prompts[*].name/version` |
|
||||||
|
| `requireGatekeeperRules` | 是否要求 Gatekeeper 规则元数据 | 要求 `gatekeeper_result.rules` 非空,且每条规则有 `id`、`enabled`、`default_severity` |
|
||||||
|
|
||||||
## 2. Trace Fixture
|
## 2. Trace Fixture
|
||||||
|
|
||||||
@@ -72,6 +85,9 @@ fixture 是一次 Agent 运行后的 trace 快照。评测器只读取当前规
|
|||||||
| `session.selfEvaluation.verifier_evaluation.verdict` | Verifier 判定 | 必须存在并符合 case 的 `allowedVerdicts` |
|
| `session.selfEvaluation.verifier_evaluation.verdict` | Verifier 判定 | 必须存在并符合 case 的 `allowedVerdicts` |
|
||||||
| `session.selfEvaluation.verifier_evaluation.gatekeeper_result.status` | Gatekeeper 结果 | V2 case 必须存在;`fail` 不允许搭配 `PASS` |
|
| `session.selfEvaluation.verifier_evaluation.gatekeeper_result.status` | Gatekeeper 结果 | V2 case 必须存在;`fail` 不允许搭配 `PASS` |
|
||||||
| `session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | Gatekeeper 规则集版本 | 新矩阵 case 可显式断言该版本 |
|
| `session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | Gatekeeper 规则集版本 | 新矩阵 case 可显式断言该版本 |
|
||||||
|
| `session.selfEvaluation.verifier_evaluation.gatekeeper_result.rules` | Gatekeeper 规则元数据摘要 | 审计 case 可要求规则列表非空且字段完整 |
|
||||||
|
| `session.selfEvaluation.verifier_evaluation.prompt_audit.version` | Chat Prompt 审计目录版本 | 审计 case 可显式断言该版本 |
|
||||||
|
| `session.selfEvaluation.verifier_evaluation.prompt_audit.prompts` | 各 Chat Prompt 名称、版本和资源路径 | 审计 case 可断言 planner、executor、verifier、composer 版本 |
|
||||||
| `session.selfEvaluation.verifier_evaluation.claim_checks` | Verifier V2 claim 级校验 | V2 case 必须存在;每项需要 `claim_id`、`verification`、`detail` |
|
| `session.selfEvaluation.verifier_evaluation.claim_checks` | Verifier V2 claim 级校验 | V2 case 必须存在;每项需要 `claim_id`、`verification`、`detail` |
|
||||||
| `session.selfEvaluation.verifier_evaluation.composer_output.status` | Composer 渲染状态 | V2 case 必须存在;记录 `valid`、`composer_malformed` 等 |
|
| `session.selfEvaluation.verifier_evaluation.composer_output.status` | Composer 渲染状态 | V2 case 必须存在;记录 `valid`、`composer_malformed` 等 |
|
||||||
| `session.selfEvaluation.verifier_evaluation.executor_structured_output.claims[*].evidence_bindings` | Executor claim 证据绑定 | 如果结构化输出存在,每条 claim 需要证据绑定 |
|
| `session.selfEvaluation.verifier_evaluation.executor_structured_output.claims[*].evidence_bindings` | Executor claim 证据绑定 | 如果结构化输出存在,每条 claim 需要证据绑定 |
|
||||||
@@ -95,8 +111,10 @@ Java 类型:`DiagnosisEvalResult`
|
|||||||
| `evidenceCoverage` | 每个必需工具是否出现 |
|
| `evidenceCoverage` | 每个必需工具是否出现 |
|
||||||
| `gatekeeperStatus` | 读到的 `gatekeeper_result.status` |
|
| `gatekeeperStatus` | 读到的 `gatekeeper_result.status` |
|
||||||
| `gatekeeperRuleSetVersion` | 读到的 `gatekeeper_result.rule_set_version` |
|
| `gatekeeperRuleSetVersion` | 读到的 `gatekeeper_result.rule_set_version` |
|
||||||
|
| `promptAuditVersion` | 读到的 `prompt_audit.version` |
|
||||||
| `composerStatus` | 读到的 `composer_output.status` |
|
| `composerStatus` | 读到的 `composer_output.status` |
|
||||||
| `claimCheckCount` | `claim_checks` 数量 |
|
| `claimCheckCount` | `claim_checks` 数量 |
|
||||||
|
| `gatekeeperRuleCount` | `gatekeeper_result.rules` 数量 |
|
||||||
| `toolCallCount` | trace 中工具调用总数 |
|
| `toolCallCount` | trace 中工具调用总数 |
|
||||||
| `durationMs` | trace 总耗时 |
|
| `durationMs` | trace 总耗时 |
|
||||||
|
|
||||||
@@ -130,6 +148,10 @@ Executor structured output
|
|||||||
- V2 case 必须有 `gatekeeper_result`、`claim_checks`、`composer_output`。
|
- V2 case 必须有 `gatekeeper_result`、`claim_checks`、`composer_output`。
|
||||||
- `gatekeeper_result.status = fail` 时,Verifier verdict 不能是 `PASS`。
|
- `gatekeeper_result.status = fail` 时,Verifier verdict 不能是 `PASS`。
|
||||||
- 配置 `expectedGatekeeperRuleSetVersion` 的 case 必须匹配 `gatekeeper_result.rule_set_version`。
|
- 配置 `expectedGatekeeperRuleSetVersion` 的 case 必须匹配 `gatekeeper_result.rule_set_version`。
|
||||||
|
- 配置 `requirePromptAudit` 的 case 必须包含 `prompt_audit.version`。
|
||||||
|
- 配置 `expectedPromptAuditVersion` 的 case 必须匹配 `prompt_audit.version`。
|
||||||
|
- 配置 `expectedPromptVersions` 的 case 必须能在 `prompt_audit.prompts` 中找到对应角色和版本。
|
||||||
|
- 配置 `requireGatekeeperRules` 的 case 必须包含非空 `gatekeeper_result.rules`,且每条规则有 `id`、`enabled`、`default_severity`。
|
||||||
- `claim_checks[*].verification` 只能是 `direct_observation`、`reasonable_inference`、`overstated`、`unsupported`、`external_unknown`、`contradicted`。
|
- `claim_checks[*].verification` 只能是 `direct_observation`、`reasonable_inference`、`overstated`、`unsupported`、`external_unknown`、`contradicted`。
|
||||||
- Composer 输出必须记录 `status`。
|
- Composer 输出必须记录 `status`。
|
||||||
- 最终答案不能泄漏 `executor_evidence_v2`、`answer_version`、`evidence_bindings`、`claim_id`。
|
- 最终答案不能泄漏 `executor_evidence_v2`、`answer_version`、`evidence_bindings`、`claim_id`。
|
||||||
|
|||||||
+56
-38
@@ -1,47 +1,65 @@
|
|||||||
# 已知问题记录
|
# MVP Issues 索引
|
||||||
|
|
||||||
| # | 标题 | 严重程度 | 状态 | 文件 |
|
**更新日期**:2026-07-10
|
||||||
|---|---|---|---|---|
|
**状态**:按活跃问题、设计笔记、RAG 问题集和已归档问题整理
|
||||||
| ISS-001 | Executor 重复召回同一文档 | 中 | 已修复 | [ISS-001-duplicate-retrieval.md](ISS-001-duplicate-retrieval.md) |
|
|
||||||
| ISS-002 | Executor 无约束重复调用 lookup_knowledge | 中 | 已修复 | [ISS-002-executor-unconstrained-lookup.md](ISS-002-executor-unconstrained-lookup.md) |
|
|
||||||
| ISS-003 | MVP 设计与实现 Review 收敛 | 高 | 待规划 | [ISS-003-mvp-design-implementation-review.md](ISS-003-mvp-design-implementation-review.md) |
|
|
||||||
| ISS-004 | Executor 域级检索水位控制(Phase 2) | 低 | 待规划 | [ISS-004-executor-domain-hard-limit.md](ISS-004-executor-domain-hard-limit.md) |
|
|
||||||
| ISS-005 | 证据链补齐与降级契约收敛 | 高 | 已归档 | [ISS-005-evidence-trace-hardening.md](ISS-005-evidence-trace-hardening.md) |
|
|
||||||
| ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) |
|
|
||||||
| ISS-007 | Verifier 证据摘要保真与工具命中质量问题 | 高 | 已实施 | [ISS-007-verifier-evidence-summary-fidelity.md](ISS-007-verifier-evidence-summary-fidelity.md) |
|
|
||||||
| ISS-008 | Executor 窄范围查询越界 | 中 | 已修复 | [ISS-008-executor-narrow-scope-overreach.md](ISS-008-executor-narrow-scope-overreach.md) |
|
|
||||||
| ISS-009 | negative_observation 精确引用 no-evidence 结果 | 中 | 已修复 | [ISS-009-negative-observation-no-evidence-reference.md](ISS-009-negative-observation-no-evidence-reference.md) |
|
|
||||||
| executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [executor-evidence-attribution-hallucination.md](executor-evidence-attribution-hallucination.md) |
|
|
||||||
| executor-self-evidence-loop-design-note | Executor 自证循环与证据摘要链路设计记录 | 高 | 已形成方向 | [executor-self-evidence-loop-design-note.md](executor-self-evidence-loop-design-note.md) |
|
|
||||||
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) |
|
|
||||||
| diagnosis-eval-baseline-diff | 诊断评测 baseline diff 与回归判断 | 中 | 已归档 | [diagnosis-eval-baseline-diff.md](diagnosis-eval-baseline-diff.md) |
|
|
||||||
| mvp-demo-interview-runbook | Plan C 面试可复现 Demo 包 | 中 | 已归档 | [mvp-demo-interview-runbook.md](mvp-demo-interview-runbook.md) |
|
|
||||||
|
|
||||||
## RAG 重构计划
|
## 目录约定
|
||||||
|
|
||||||
|
| 目录 | 用途 |
|
||||||
|
|---|---|
|
||||||
|
| [active/](active/) | 仍需要规划或实现的问题 |
|
||||||
|
| [design-notes/](design-notes/) | 已形成方向、用于指导后续实现的设计记录 |
|
||||||
|
| [rag/](rag/) | RAG 子问题集合;多数已合并到 RAG 重构计划 |
|
||||||
|
| [archived/](archived/) | 已修复、已实施或已归档的问题 |
|
||||||
|
|
||||||
|
## 活跃问题
|
||||||
|
|
||||||
| 名称 | 标题 | 严重程度 | 状态 | 文件 |
|
| 名称 | 标题 | 严重程度 | 状态 | 文件 |
|
||||||
|---|---|---|---|---|
|
|---|---|---|---|---|
|
||||||
| rag-refactor-plan | RAG 检索重构计划 | 高 | 待规划 | [rag-refactor-plan.md](rag-refactor-plan.md) |
|
| ISS-003 | MVP 设计与实现 Review 收敛 | 高 | 待规划 | [active/ISS-003-mvp-design-implementation-review.md](active/ISS-003-mvp-design-implementation-review.md) |
|
||||||
|
| ISS-004 | Executor 域级检索水位控制 | 低 | 待规划 | [active/ISS-004-executor-domain-hard-limit.md](active/ISS-004-executor-domain-hard-limit.md) |
|
||||||
|
| executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [active/executor-evidence-attribution-hallucination.md](active/executor-evidence-attribution-hallucination.md) |
|
||||||
|
| rag-refactor-plan | RAG 检索重构计划 | 高 | 待规划 | [active/rag-refactor-plan.md](active/rag-refactor-plan.md) |
|
||||||
|
|
||||||
## RAG 检索问题
|
## 设计笔记
|
||||||
|
|
||||||
| 名称 | 标题 | 严重程度 | 状态 | 文件 |
|
| 名称 | 标题 | 状态 | 文件 |
|
||||||
|---|---|---|---|---|
|
|---|---|---|---|
|
||||||
| chunk-context-reconstruction | RAG 切片上下文重建缺失 | 高 | 已合并到重构计划 | [rag-chunk-context-reconstruction.md](rag-chunk-context-reconstruction.md) |
|
| executor-self-evidence-loop-design-note | Executor 自证循环与证据摘要链路设计记录 | 已形成方向 | [design-notes/executor-self-evidence-loop-design-note.md](design-notes/executor-self-evidence-loop-design-note.md) |
|
||||||
| breadcrumb-embedding-gap | RAG breadcrumb 未参与向量语义 | 高 | 已合并到重构计划 | [rag-breadcrumb-embedding-gap.md](rag-breadcrumb-embedding-gap.md) |
|
| executor-structured-output-v2 | Executor 结构化输出 V2 阶段设计 | 部分已实施,保留为后续改造参考 | [design-notes/executor-structured-output-v2.md](design-notes/executor-structured-output-v2.md) |
|
||||||
| l0-l1-fusion-ranking | RAG L0 和 L1 未真正融合排序 | 中 | 已合并到重构计划 | [rag-l0-l1-fusion-ranking.md](rag-l0-l1-fusion-ranking.md) |
|
|
||||||
| l0-keyword-matching-quality | RAG L0 关键词匹配质量不足 | 中 | 已合并到重构计划 | [rag-l0-keyword-matching-quality.md](rag-l0-keyword-matching-quality.md) |
|
|
||||||
| l1-score-calibration | RAG L1 分数阈值未校准 | 中 | 已合并到重构计划 | [rag-l1-score-calibration.md](rag-l1-score-calibration.md) |
|
|
||||||
| context-packing-and-reranking | RAG 缺少上下文打包和 Rerank | 中 | 已合并到重构计划 | [rag-context-packing-and-reranking.md](rag-context-packing-and-reranking.md) |
|
|
||||||
| upload-chunk-parameter-drift | RAG 上传切片参数未真正生效 | 低 | 已合并到重构计划 | [rag-upload-chunk-parameter-drift.md](rag-upload-chunk-parameter-drift.md) |
|
|
||||||
| query-rewrite-gap | RAG 查询改写能力薄弱 | 中 | 已合并到重构计划 | [rag-query-rewrite-gap.md](rag-query-rewrite-gap.md) |
|
|
||||||
|
|
||||||
## RAG 框架化改造
|
## RAG 问题集
|
||||||
|
|
||||||
| 名称 | 标题 | 严重程度 | 状态 | 文件 |
|
这些问题已经收敛到 [active/rag-refactor-plan.md](active/rag-refactor-plan.md),单个文件保留用于追溯原始问题和设计背景。
|
||||||
|---|---|---|---|---|
|
|
||||||
| spring-ai-vectorstore-migration | RAG 迁移到 Spring AI VectorStore 检索抽象 | 高 | 已合并到重构计划 | [rag-spring-ai-vectorstore-migration.md](rag-spring-ai-vectorstore-migration.md) |
|
| 名称 | 标题 | 状态 | 文件 |
|
||||||
| spring-ai-query-transformer | RAG 接入 Spring AI Query Transformer | 中 | 已合并到重构计划 | [rag-spring-ai-query-transformer.md](rag-spring-ai-query-transformer.md) |
|
|---|---|---|---|
|
||||||
| spring-ai-document-postprocessor | RAG 使用 DocumentPostProcessor 做后处理 | 中 | 已合并到重构计划 | [rag-spring-ai-document-postprocessor.md](rag-spring-ai-document-postprocessor.md) |
|
| breadcrumb-embedding-gap | RAG breadcrumb 未参与向量语义 | 已合并到重构计划 | [rag/rag-breadcrumb-embedding-gap.md](rag/rag-breadcrumb-embedding-gap.md) |
|
||||||
| l0-domain-entity-hint | RAG 将 L0 降级为领域和实体 Hint | 中 | 已合并到重构计划 | [rag-l0-domain-entity-hint.md](rag-l0-domain-entity-hint.md) |
|
| chunk-context-reconstruction | RAG 切片上下文重建缺失 | 已合并到重构计划 | [rag/rag-chunk-context-reconstruction.md](rag/rag-chunk-context-reconstruction.md) |
|
||||||
| spring-ai-advisor-boundary | RAG 明确 Spring AI Advisor 与 Agent Tool 的边界 | 中 | 已合并到重构计划 | [rag-spring-ai-advisor-boundary.md](rag-spring-ai-advisor-boundary.md) |
|
| context-packing-and-reranking | RAG 缺少上下文打包和 Rerank | 已合并到重构计划 | [rag/rag-context-packing-and-reranking.md](rag/rag-context-packing-and-reranking.md) |
|
||||||
|
| l0-domain-entity-hint | RAG 将 L0 降级为领域和实体 Hint | 已合并到重构计划 | [rag/rag-l0-domain-entity-hint.md](rag/rag-l0-domain-entity-hint.md) |
|
||||||
|
| l0-keyword-matching-quality | RAG L0 关键词匹配质量不足 | 已合并到重构计划 | [rag/rag-l0-keyword-matching-quality.md](rag/rag-l0-keyword-matching-quality.md) |
|
||||||
|
| l0-l1-fusion-ranking | RAG L0 和 L1 未真正融合排序 | 已合并到重构计划 | [rag/rag-l0-l1-fusion-ranking.md](rag/rag-l0-l1-fusion-ranking.md) |
|
||||||
|
| l1-score-calibration | RAG L1 分数阈值未校准 | 已合并到重构计划 | [rag/rag-l1-score-calibration.md](rag/rag-l1-score-calibration.md) |
|
||||||
|
| query-rewrite-gap | RAG 查询改写能力薄弱 | 已合并到重构计划 | [rag/rag-query-rewrite-gap.md](rag/rag-query-rewrite-gap.md) |
|
||||||
|
| spring-ai-advisor-boundary | Spring AI Advisor 与 Agent Tool 边界 | 已合并到重构计划 | [rag/rag-spring-ai-advisor-boundary.md](rag/rag-spring-ai-advisor-boundary.md) |
|
||||||
|
| spring-ai-document-postprocessor | 使用 DocumentPostProcessor 做后处理 | 已合并到重构计划 | [rag/rag-spring-ai-document-postprocessor.md](rag/rag-spring-ai-document-postprocessor.md) |
|
||||||
|
| spring-ai-query-transformer | 接入 Spring AI Query Transformer | 已合并到重构计划 | [rag/rag-spring-ai-query-transformer.md](rag/rag-spring-ai-query-transformer.md) |
|
||||||
|
| spring-ai-vectorstore-migration | 迁移到 Spring AI VectorStore 检索抽象 | 已合并到重构计划 | [rag/rag-spring-ai-vectorstore-migration.md](rag/rag-spring-ai-vectorstore-migration.md) |
|
||||||
|
| upload-chunk-parameter-drift | 上传切片参数未真正生效 | 已合并到重构计划 | [rag/rag-upload-chunk-parameter-drift.md](rag/rag-upload-chunk-parameter-drift.md) |
|
||||||
|
|
||||||
|
## 已归档问题
|
||||||
|
|
||||||
|
| 名称 | 标题 | 状态 | 文件 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| ISS-001 | Executor 重复召回同一文档 | 已修复 | [archived/ISS-001-duplicate-retrieval.md](archived/ISS-001-duplicate-retrieval.md) |
|
||||||
|
| ISS-002 | Executor 无约束重复调用 lookup_knowledge | 已修复 | [archived/ISS-002-executor-unconstrained-lookup.md](archived/ISS-002-executor-unconstrained-lookup.md) |
|
||||||
|
| ISS-005 | 证据链补齐与降级契约收敛 | 已归档 | [archived/ISS-005-evidence-trace-hardening.md](archived/ISS-005-evidence-trace-hardening.md) |
|
||||||
|
| ISS-006 | 固定诊断评测集与回归 Harness | 已归档 | [archived/ISS-006-diagnosis-eval-harness.md](archived/ISS-006-diagnosis-eval-harness.md) |
|
||||||
|
| ISS-007 | Verifier 证据摘要保真与工具命中质量问题 | 已实施 | [archived/ISS-007-verifier-evidence-summary-fidelity.md](archived/ISS-007-verifier-evidence-summary-fidelity.md) |
|
||||||
|
| ISS-008 | Executor 窄范围查询越界 | 已修复 | [archived/ISS-008-executor-narrow-scope-overreach.md](archived/ISS-008-executor-narrow-scope-overreach.md) |
|
||||||
|
| ISS-009 | negative_observation 精确引用 no-evidence 结果 | 已修复 | [archived/ISS-009-negative-observation-no-evidence-reference.md](archived/ISS-009-negative-observation-no-evidence-reference.md) |
|
||||||
|
| ISS-010 | 同 session 多轮诊断 Trace 隔离 | 已归档 | [archived/ISS-010-session-run-trace-isolation.md](archived/ISS-010-session-run-trace-isolation.md) |
|
||||||
|
| diagnosis-eval-baseline-diff | 诊断评测 baseline diff 与回归判断 | 已归档 | [archived/diagnosis-eval-baseline-diff.md](archived/diagnosis-eval-baseline-diff.md) |
|
||||||
|
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 已归档 | [archived/expand-diagnosis-eval-fixtures.md](archived/expand-diagnosis-eval-fixtures.md) |
|
||||||
|
| mvp-demo-interview-runbook | Plan C 面试可复现 Demo 包 | 已归档 | [archived/mvp-demo-interview-runbook.md](archived/mvp-demo-interview-runbook.md) |
|
||||||
|
|||||||
@@ -347,19 +347,19 @@ RAG、Agent、AIOps、数据库记录互相关联,必须分阶段推进,每
|
|||||||
|
|
||||||
本计划合并以下问题和改造方向:
|
本计划合并以下问题和改造方向:
|
||||||
|
|
||||||
- [rag-chunk-context-reconstruction.md](rag-chunk-context-reconstruction.md)
|
- [rag-chunk-context-reconstruction.md](../rag/rag-chunk-context-reconstruction.md)
|
||||||
- [rag-breadcrumb-embedding-gap.md](rag-breadcrumb-embedding-gap.md)
|
- [rag-breadcrumb-embedding-gap.md](../rag/rag-breadcrumb-embedding-gap.md)
|
||||||
- [rag-l0-l1-fusion-ranking.md](rag-l0-l1-fusion-ranking.md)
|
- [rag-l0-l1-fusion-ranking.md](../rag/rag-l0-l1-fusion-ranking.md)
|
||||||
- [rag-l0-keyword-matching-quality.md](rag-l0-keyword-matching-quality.md)
|
- [rag-l0-keyword-matching-quality.md](../rag/rag-l0-keyword-matching-quality.md)
|
||||||
- [rag-l1-score-calibration.md](rag-l1-score-calibration.md)
|
- [rag-l1-score-calibration.md](../rag/rag-l1-score-calibration.md)
|
||||||
- [rag-context-packing-and-reranking.md](rag-context-packing-and-reranking.md)
|
- [rag-context-packing-and-reranking.md](../rag/rag-context-packing-and-reranking.md)
|
||||||
- [rag-upload-chunk-parameter-drift.md](rag-upload-chunk-parameter-drift.md)
|
- [rag-upload-chunk-parameter-drift.md](../rag/rag-upload-chunk-parameter-drift.md)
|
||||||
- [rag-query-rewrite-gap.md](rag-query-rewrite-gap.md)
|
- [rag-query-rewrite-gap.md](../rag/rag-query-rewrite-gap.md)
|
||||||
- [rag-spring-ai-vectorstore-migration.md](rag-spring-ai-vectorstore-migration.md)
|
- [rag-spring-ai-vectorstore-migration.md](../rag/rag-spring-ai-vectorstore-migration.md)
|
||||||
- [rag-spring-ai-query-transformer.md](rag-spring-ai-query-transformer.md)
|
- [rag-spring-ai-query-transformer.md](../rag/rag-spring-ai-query-transformer.md)
|
||||||
- [rag-spring-ai-document-postprocessor.md](rag-spring-ai-document-postprocessor.md)
|
- [rag-spring-ai-document-postprocessor.md](../rag/rag-spring-ai-document-postprocessor.md)
|
||||||
- [rag-l0-domain-entity-hint.md](rag-l0-domain-entity-hint.md)
|
- [rag-l0-domain-entity-hint.md](../rag/rag-l0-domain-entity-hint.md)
|
||||||
- [rag-spring-ai-advisor-boundary.md](rag-spring-ai-advisor-boundary.md)
|
- [rag-spring-ai-advisor-boundary.md](../rag/rag-spring-ai-advisor-boundary.md)
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
+1
-1
@@ -4,7 +4,7 @@
|
|||||||
**严重程度**:中(影响 token 消耗和上下文质量,不影响功能正确性)
|
**严重程度**:中(影响 token 消耗和上下文质量,不影响功能正确性)
|
||||||
**发现时间**:2026-06-30
|
**发现时间**:2026-06-30
|
||||||
**修复版本**:session-dedup-knowledge-map
|
**修复版本**:session-dedup-knowledge-map
|
||||||
**历史架构文档**:[会话级去重与知识域地图](../architecture/archive/2026-07-05-legacy/session-dedup-knowledge-map.md)
|
**历史架构文档**:[会话级去重与知识域地图](../../architecture/archive/2026-07-05-legacy/session-dedup-knowledge-map.md)
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -0,0 +1,586 @@
|
|||||||
|
# ISS-010 同 session 多轮诊断 Trace 隔离
|
||||||
|
|
||||||
|
**状态**:已归档
|
||||||
|
**严重程度**:高
|
||||||
|
**发现时间**:2026-07-10
|
||||||
|
**来源**:同一 `sessionId` 多轮 Chat E2E 验证
|
||||||
|
|
||||||
|
**归档日期**:2026-07-10
|
||||||
|
**OpenSpec**:`openspec/changes/archive/2026-07-10-session-run-trace-isolation`
|
||||||
|
**实现提交**:`52bf030`、`26d5529`、`027aed1`、`d928a19`、`78c1477`、`f9df943`
|
||||||
|
**归档提交**:`3578709`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 背景
|
||||||
|
|
||||||
|
归档结论:当前 MVP 已将“会话态”和“运行态”拆开。`sessionId` 表示多轮会话目录和 Redis 上下文;`runId` 表示一次可回放诊断执行。Trace、Feedback、Evaluation、AIOps 和案例沉淀的新路径都按 `runId` 隔离。
|
||||||
|
|
||||||
|
原问题中 Chat 链路同时存在两类“会话”语义:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Redis SessionContext
|
||||||
|
-> 保存同一 sessionId 的多轮对话历史
|
||||||
|
-> 用于下一轮模型上下文
|
||||||
|
|
||||||
|
MySQL diagnosis_session / agent_step / tool_invocation
|
||||||
|
-> 保存诊断 Trace
|
||||||
|
-> 用于 Trace API、Verifier、Evidence score、Feedback 和评测
|
||||||
|
```
|
||||||
|
|
||||||
|
多轮对话需要继续复用 `sessionId`,否则无法保留上下文。但一次诊断 Trace 应该是可独立回放、可独立评分、可独立反馈的执行单元。
|
||||||
|
|
||||||
|
当前实现只按 `sessionId` 关联 Trace,导致同一个 `sessionId` 下多轮诊断的 step/tool 记录混在一起。
|
||||||
|
|
||||||
|
当前实现已改为:
|
||||||
|
|
||||||
|
```text
|
||||||
|
chat_session(sessionId)
|
||||||
|
-> diagnosis_run(runId)
|
||||||
|
-> agent_step.run_id
|
||||||
|
-> tool_invocation.run_id
|
||||||
|
```
|
||||||
|
|
||||||
|
旧 `diagnosis_session` 保留为历史兼容和回滚表,新 Chat/AIOps 执行不再写入新的运行态。
|
||||||
|
|
||||||
|
## 归档结果
|
||||||
|
|
||||||
|
- OpenSpec 已归档到 `openspec/changes/archive/2026-07-10-session-run-trace-isolation`。
|
||||||
|
- 主规格已同步到 `openspec/specs/session-run-trace-isolation/spec.md`。
|
||||||
|
- `mvp/architecture/` 和 `mvp/tables/` 已更新为 `chat_session -> diagnosis_run -> agent_step/tool_invocation(run_id)` 模型。
|
||||||
|
- Demo 脚本和 Trace UI 已支持 `sessionId + runId` 精确 Trace 和 Feedback。
|
||||||
|
- Maven E2E、`scripts/query_mysql.py` DB 检查、`logs/` 日志检查和 baseline drift 检查均已通过;未观察到 baseline drift。
|
||||||
|
- `devflow/projects/2026-07-10-session-run-trace-isolation/` 已保存 brief、evidence、decisions、acceptance。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## E2E 证据
|
||||||
|
|
||||||
|
本次使用 `mvp-demo` profile 通过 Maven 启动服务,并用同一个 `sessionId` 连续请求两轮 `/api/chat`:
|
||||||
|
|
||||||
|
```text
|
||||||
|
sessionId = e2e-multiturn-codex-20260710-1615
|
||||||
|
round 1 = 支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。
|
||||||
|
round 2 = 基于上一轮结论,只列出目前最缺的三类证据,以及下一步应该优先查哪个系统。
|
||||||
|
```
|
||||||
|
|
||||||
|
验证结果:
|
||||||
|
|
||||||
|
- 第一轮成功,走多 Agent:`planner -> executor -> verifier -> composer`。
|
||||||
|
- 第二轮成功,日志显示进入请求时 `会话历史消息对数: 1`,说明 Redis 历史上下文被复用。
|
||||||
|
- `/api/chat/session/{sessionId}` 返回 `messagePairCount=2`。
|
||||||
|
- `diagnosis_session` 只有一行,`query` 被第二轮问题覆盖。
|
||||||
|
- `agent_step` 返回 14 行,包含第一轮多 Agent step 和第二轮 `intelligent_assistant` step。
|
||||||
|
- `tool_invocation` 返回 19 行,包含两轮工具调用。
|
||||||
|
- `self_evaluation.verifier_evaluation` 仍保留第一轮 Verifier 结果;第二轮简单问答没有新的 Verifier,但 rule evaluation 会基于同 session 全部工具调用重新计算。
|
||||||
|
|
||||||
|
关键入库形态:
|
||||||
|
|
||||||
|
```text
|
||||||
|
diagnosis_session
|
||||||
|
session_id = e2e-multiturn-codex-20260710-1615
|
||||||
|
query = round 2 question
|
||||||
|
status = SUCCESS
|
||||||
|
step_count = 14
|
||||||
|
tool_call_count = 19
|
||||||
|
|
||||||
|
agent_step
|
||||||
|
round 1: planner, executor..., verifier, composer
|
||||||
|
round 2: intelligent_assistant...
|
||||||
|
|
||||||
|
tool_invocation
|
||||||
|
round 1 tools + round 2 tools all under same session_id
|
||||||
|
```
|
||||||
|
|
||||||
|
本次验证产物保存在:
|
||||||
|
|
||||||
|
- `target/e2e/request-round1.json`
|
||||||
|
- `target/e2e/response-round1.json`
|
||||||
|
- `target/e2e/request-round2.json`
|
||||||
|
- `target/e2e/response-round2.json`
|
||||||
|
- `target/e2e/trace-after-round2.json`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 核心问题
|
||||||
|
|
||||||
|
### P0:Trace 不是单次诊断的稳定回放
|
||||||
|
|
||||||
|
`GET /api/diagnosis/{sessionId}/trace` 会聚合同一 `sessionId` 下所有 `agent_step` 和 `tool_invocation`。
|
||||||
|
|
||||||
|
多轮之后,Trace 不再表示某一轮诊断,而是混合历史执行轨迹。
|
||||||
|
|
||||||
|
### P0:Verifier 和评分可能读取跨轮证据
|
||||||
|
|
||||||
|
Verifier、Gatekeeper、`ToolTraceSummaryService` 和 `EvaluationService` 当前主要按 `sessionId` 查询工具调用。
|
||||||
|
|
||||||
|
如果上一轮和当前轮证据混在一起,当前轮可能引用或评分到历史工具结果。
|
||||||
|
|
||||||
|
### P1:反馈语义不清晰
|
||||||
|
|
||||||
|
`feedback` 当前在 `diagnosis_session` 上按 `sessionId` 保存。
|
||||||
|
|
||||||
|
多轮之后,用户反馈的是哪一轮答案不再明确。`useful` 反馈沉淀到 `case_library` 时也可能关联到最新主表答案,而不是用户实际评价的那一轮。
|
||||||
|
|
||||||
|
### P1:`diagnosis_session` 字段被覆盖但子表追加
|
||||||
|
|
||||||
|
主表 `query/answer/status/self_evaluation/step_count/tool_call_count` 表示最新运行或混合统计,子表却保留多轮历史。
|
||||||
|
|
||||||
|
这会让 Trace summary、数据库统计和人工排查产生歧义。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 已确认决策
|
||||||
|
|
||||||
|
### D1:`runId` 是正式 API 字段
|
||||||
|
|
||||||
|
`/api/chat` 和 `/api/ai_ops` 的响应或 SSE 消息需要暴露本次执行的 `runId`。
|
||||||
|
|
||||||
|
```text
|
||||||
|
sessionId = 多轮对话上下文 ID
|
||||||
|
runId = 本轮诊断执行 ID
|
||||||
|
```
|
||||||
|
|
||||||
|
新客户端应优先用 `runId` 查询 Trace 和提交 Feedback。旧客户端只传 `sessionId` 时,服务端兼容解析该 session 的最新 run。
|
||||||
|
|
||||||
|
### D2:拆分会话态和运行态
|
||||||
|
|
||||||
|
不再把 session、run、trace 全部塞进 `diagnosis_session` 一张主表。
|
||||||
|
|
||||||
|
新增两张主表:
|
||||||
|
|
||||||
|
```text
|
||||||
|
chat_session
|
||||||
|
-> 多轮对话上下文主表
|
||||||
|
|
||||||
|
diagnosis_run
|
||||||
|
-> 单次诊断执行主表
|
||||||
|
```
|
||||||
|
|
||||||
|
Trace 继续使用现有明细表表达:
|
||||||
|
|
||||||
|
```text
|
||||||
|
agent_step
|
||||||
|
tool_invocation
|
||||||
|
```
|
||||||
|
|
||||||
|
暂不新增单独的 `diagnosis_trace` 或 `trace_event` 主表。
|
||||||
|
|
||||||
|
### D3:`runId` 格式
|
||||||
|
|
||||||
|
使用 `run-` + UUID 全量字符串。
|
||||||
|
|
||||||
|
```text
|
||||||
|
run-550e8400-e29b-41d4-a716-446655440000
|
||||||
|
```
|
||||||
|
|
||||||
|
### D4:Trace API 兼容旧路径
|
||||||
|
|
||||||
|
```text
|
||||||
|
GET /api/diagnosis/{sessionId}/trace
|
||||||
|
-> 查该 session 最新 run
|
||||||
|
|
||||||
|
GET /api/diagnosis/{sessionId}/trace?runId=run-xxx
|
||||||
|
-> 查指定 run
|
||||||
|
```
|
||||||
|
|
||||||
|
指定 `runId` 时必须校验该 run 属于 path 中的 `sessionId`。
|
||||||
|
|
||||||
|
最新 run 建议按 `diagnosis_run.created_at DESC, id DESC` 解析,避免旧 run 因反馈或异步评分更新 `updated_at` 后被误认为最新。
|
||||||
|
|
||||||
|
### D5:Feedback 优先绑定 run
|
||||||
|
|
||||||
|
Feedback request 支持 `runId`。
|
||||||
|
|
||||||
|
- 有 `runId`:绑定指定 run。
|
||||||
|
- 无 `runId`:短期兼容绑定该 `sessionId` 最新 run,并显式标记 fallback。
|
||||||
|
- `case_library.diagnosis_id` 新数据保存 `run_id`。
|
||||||
|
|
||||||
|
兼容语义:历史 `case_library.diagnosis_id` 可能保存 `diagnosis_session.session_id`;本 change 之后自动沉淀的新数据保存 `diagnosis_run.run_id`。查询、幂等和文档需要在过渡期识别两种来源,避免把旧案例误判为无效数据。
|
||||||
|
|
||||||
|
### D6:所有 `/api/chat` 执行请求都创建 run
|
||||||
|
|
||||||
|
只要请求通过参数校验并进入 `ChatService.executeChatWithStrategy`,就创建新的 diagnosis run。
|
||||||
|
|
||||||
|
- 简单问答也创建 run。
|
||||||
|
- 复杂诊断也创建 run。
|
||||||
|
- 空问题等参数校验失败不创建 run。
|
||||||
|
|
||||||
|
### D7:AIOps 同步纳入 run 隔离
|
||||||
|
|
||||||
|
每次 `/api/ai_ops` 执行也创建新的 diagnosis run。AIOps 的 step、tool invocation 和 rule evaluation 都按 `runId` 隔离。
|
||||||
|
|
||||||
|
阶段说明:AIOps 可作为独立实现切片排在 Chat 之后,但必须在本 change 整体完成前落地;Chat-only 的中间状态只能作为过渡验证状态,不能作为生产完成状态归档。
|
||||||
|
|
||||||
|
### D8:`chat_session` 第一阶段只保存会话元数据
|
||||||
|
|
||||||
|
`chat_session` 是会话目录/索引表,不保存完整对话历史正文。
|
||||||
|
|
||||||
|
建议保存:
|
||||||
|
|
||||||
|
```text
|
||||||
|
session_id
|
||||||
|
status
|
||||||
|
message_pair_count
|
||||||
|
created_at
|
||||||
|
last_active_at
|
||||||
|
expires_at
|
||||||
|
```
|
||||||
|
|
||||||
|
完整多轮对话历史继续放在 Redis `SessionContext.messageHistory`,用于下一轮 prompt 上下文。
|
||||||
|
|
||||||
|
每轮需要长期审计的用户问题和最终回答保存到 `diagnosis_run.query` / `diagnosis_run.answer`。
|
||||||
|
|
||||||
|
`chat_session.expires_at` 只表示 MySQL 会话目录的过期/清理元数据;Redis TTL 到期后,`SessionContext.messageHistory` 可能不再存在,但已经持久化的 `diagnosis_run`、`agent_step` 和 `tool_invocation` 仍作为审计记录保留。
|
||||||
|
|
||||||
|
如果未来需要长期保存完整聊天历史,再单独设计 `chat_message` 表,不在本阶段引入。
|
||||||
|
|
||||||
|
### D9:旧 `diagnosis_session` 表保留但新代码不再写入
|
||||||
|
|
||||||
|
新增 `diagnosis_run` 后,旧 `diagnosis_session` 不立即删除、不立即改造成 view、不直接重命名。
|
||||||
|
|
||||||
|
迁移策略:
|
||||||
|
|
||||||
|
1. 新增 `chat_session` / `diagnosis_run`。
|
||||||
|
2. 为 `agent_step` / `tool_invocation` 新增 nullable `run_id`。
|
||||||
|
3. 将旧 `diagnosis_session` 数据迁移/复制为 `diagnosis_run` 兼容记录。
|
||||||
|
4. 为旧 `agent_step` / `tool_invocation` 回填对应 `run_id`。
|
||||||
|
5. 增加必要索引和查询方法,先保持兼容读取。
|
||||||
|
6. 新代码切换为只写 `chat_session` 和 `diagnosis_run`,并为新 step/tool 写入 `run_id`。
|
||||||
|
7. 验证新旧数据 `run_id` 覆盖情况后,再将新写路径要求 `run_id` 非空,并补充索引/约束。
|
||||||
|
8. 旧 `diagnosis_session` 暂时保留,用于历史核对和回滚窗口。
|
||||||
|
9. 后续确认无依赖后,再单独归档或删除旧表。
|
||||||
|
|
||||||
|
### D10:提供轻量 run 列表 API
|
||||||
|
|
||||||
|
新增轻量查询接口,用于查看一个 Chat Session 下有哪些 Diagnosis Run。
|
||||||
|
|
||||||
|
```text
|
||||||
|
GET /api/chat/session/{sessionId}/runs
|
||||||
|
```
|
||||||
|
|
||||||
|
建议返回字段:
|
||||||
|
|
||||||
|
```text
|
||||||
|
runId
|
||||||
|
sessionId
|
||||||
|
query
|
||||||
|
status
|
||||||
|
agentFlow
|
||||||
|
answerPreview
|
||||||
|
stepCount
|
||||||
|
toolCallCount
|
||||||
|
createdAt
|
||||||
|
updatedAt
|
||||||
|
```
|
||||||
|
|
||||||
|
该接口只读 `diagnosis_run` 主表,不展开 `agent_step` / `tool_invocation` 大字段。
|
||||||
|
|
||||||
|
### D11:Feedback 缺少 `runId` 时短期兼容,长期收紧
|
||||||
|
|
||||||
|
Feedback 新协议优先要求 `runId`。
|
||||||
|
|
||||||
|
短期兼容策略:
|
||||||
|
|
||||||
|
- 有 `runId`:绑定指定 run。
|
||||||
|
- 无 `runId`:绑定该 `sessionId` 最新 run。
|
||||||
|
- 无 `runId` fallback 时,在响应或日志中明确标记 `fallbackToLatestRun=true`,并返回实际绑定的 `runId`。
|
||||||
|
|
||||||
|
长期收紧策略:
|
||||||
|
|
||||||
|
- 当前端、demo 脚本和外部调用方都完成 `runId` 传递后,再评估是否将缺少 `runId` 改为参数错误。
|
||||||
|
|
||||||
|
### D12:同步更新 demo 脚本和 Trace UI 的 `runId` 最小支持
|
||||||
|
|
||||||
|
本 issue 实施范围包含 demo 脚本和 Trace UI 的最小协议适配。
|
||||||
|
|
||||||
|
范围:
|
||||||
|
|
||||||
|
- Demo 脚本读取 `/api/chat` 或 `/api/ai_ops` 返回的 `runId`。
|
||||||
|
- Demo 脚本查询 Trace 时传 `?runId=...`。
|
||||||
|
- Trace UI 支持 URL 参数 `?sessionId=...&runId=...`。
|
||||||
|
- Trace UI 查询时如果有 `runId`,带上 `runId`。
|
||||||
|
- 不在本阶段实现完整 run 列表 UI。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 目标语义
|
||||||
|
|
||||||
|
引入明确的 `sessionId` / `runId` 分层:
|
||||||
|
|
||||||
|
```text
|
||||||
|
sessionId = 多轮对话上下文
|
||||||
|
runId = 单次诊断执行 / 单次可回放 Trace
|
||||||
|
```
|
||||||
|
|
||||||
|
目标关系:
|
||||||
|
|
||||||
|
```text
|
||||||
|
chat_session(sessionId)
|
||||||
|
-> one conversation context
|
||||||
|
-> conversation metadata / TTL / last active state
|
||||||
|
|
||||||
|
Redis SessionContext(sessionId)
|
||||||
|
-> hot messageHistory cache
|
||||||
|
-> supports prompt context window
|
||||||
|
|
||||||
|
diagnosis_run(runId, sessionId)
|
||||||
|
-> one diagnosis run
|
||||||
|
|
||||||
|
agent_step(runId, sessionId)
|
||||||
|
-> steps of one run
|
||||||
|
|
||||||
|
tool_invocation(runId, sessionId)
|
||||||
|
-> tool calls of one run
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 建议方案
|
||||||
|
|
||||||
|
采用“拆分主表 + 复用现有 Trace 明细表”的方案:
|
||||||
|
|
||||||
|
1. 新增 `chat_session`。
|
||||||
|
2. 新增 `diagnosis_run`。
|
||||||
|
3. 逐步迁移当前 `diagnosis_session` 语义到 `diagnosis_run`。
|
||||||
|
4. `agent_step` 新增 `run_id`,继续保留 `session_id` 作为冗余筛选和兼容字段。
|
||||||
|
5. `tool_invocation` 新增 `run_id`,继续保留 `session_id` 作为冗余筛选和兼容字段。
|
||||||
|
6. Trace API 聚合 `diagnosis_run + agent_step + tool_invocation`。
|
||||||
|
|
||||||
|
建议核心字段:
|
||||||
|
|
||||||
|
```text
|
||||||
|
chat_session
|
||||||
|
id
|
||||||
|
session_id unique
|
||||||
|
status
|
||||||
|
message_pair_count
|
||||||
|
created_at
|
||||||
|
last_active_at
|
||||||
|
expires_at
|
||||||
|
|
||||||
|
diagnosis_run
|
||||||
|
id
|
||||||
|
run_id unique
|
||||||
|
session_id
|
||||||
|
query
|
||||||
|
status
|
||||||
|
agent_flow
|
||||||
|
answer
|
||||||
|
self_evaluation
|
||||||
|
feedback
|
||||||
|
total_duration_ms
|
||||||
|
total_token_count
|
||||||
|
step_count
|
||||||
|
tool_call_count
|
||||||
|
created_at
|
||||||
|
updated_at
|
||||||
|
|
||||||
|
agent_step(run_id, step_index)
|
||||||
|
tool_invocation(run_id, id)
|
||||||
|
```
|
||||||
|
|
||||||
|
理由:
|
||||||
|
|
||||||
|
- `chat_session` 只表达会话态,避免会话上下文和诊断结果混在一起。
|
||||||
|
- `diagnosis_run` 只表达一次执行,天然隔离每轮 Trace、评分和反馈。
|
||||||
|
- `agent_step` / `tool_invocation` 已足够表达 Trace 明细,暂不需要额外 trace 主表。
|
||||||
|
- 后续如果需要统一时间线,再增加 `trace_event`,不阻塞本次隔离。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 分阶段计划
|
||||||
|
|
||||||
|
### Phase 0:协议基线和数据边界
|
||||||
|
|
||||||
|
目标:先把语义定死,避免实现中反复。
|
||||||
|
|
||||||
|
已确认基线:
|
||||||
|
|
||||||
|
1. `/api/chat` 是否返回 `runId`。
|
||||||
|
2. `GET /api/diagnosis/{sessionId}/trace` 默认查最新 run 还是要求显式传 `runId`。
|
||||||
|
3. Feedback 是否优先绑定 `runId`,只有旧请求缺失 `runId` 时才回退最新 run。
|
||||||
|
4. AIOps 是否和 Chat 同步接入 `runId`。
|
||||||
|
5. 简单问答是否也创建 diagnosis run。
|
||||||
|
|
||||||
|
建议默认:
|
||||||
|
|
||||||
|
- `/api/chat` 返回 `sessionId + runId`。
|
||||||
|
- `GET /api/diagnosis/{sessionId}/trace` 兼容查最新 run。
|
||||||
|
- `GET /api/diagnosis/{sessionId}/trace?runId=...` 查指定 run。
|
||||||
|
- Feedback 优先按 `runId` 绑定。
|
||||||
|
- Chat 简单问答也创建 run。
|
||||||
|
- AIOps 同步接入 run 隔离。
|
||||||
|
|
||||||
|
### Phase 1:Schema 迁移和历史数据兼容
|
||||||
|
|
||||||
|
目标:引入 `chat_session` / `diagnosis_run`,并保留旧数据可查询。
|
||||||
|
|
||||||
|
任务:
|
||||||
|
|
||||||
|
- Flyway 新增 `chat_session`。
|
||||||
|
- Flyway 新增 `diagnosis_run`。
|
||||||
|
- 为旧 `diagnosis_session` 生成兼容 `diagnosis_run` 记录。
|
||||||
|
- 为旧 `agent_step` / `tool_invocation` 回填对应 `run_id`。
|
||||||
|
- 增加 `find latest run by sessionId` 查询。
|
||||||
|
- 增加 `find by runId` 查询。
|
||||||
|
- 保留旧 `diagnosis_session` 一段时间,新代码不再写入。
|
||||||
|
|
||||||
|
验收:
|
||||||
|
|
||||||
|
- 旧 session 的 Trace 仍可查。
|
||||||
|
- 新索引存在。
|
||||||
|
- 不改变旧 `/api/chat` 必需字段。
|
||||||
|
- `agent_step` / `tool_invocation` 支持 nullable `run_id` 并完成旧数据回填。
|
||||||
|
- 新增 repository 查询可以按 `sessionId` 找最新 run、按 `runId` 找指定 run。
|
||||||
|
|
||||||
|
### Phase 2:Chat 写入切到 runId
|
||||||
|
|
||||||
|
目标:每轮 `/api/chat` 创建一个新的 run,step/tool 按 run 隔离。
|
||||||
|
|
||||||
|
任务:
|
||||||
|
|
||||||
|
- `ChatService` 每次执行生成新的 `runId`。
|
||||||
|
- `ChatController` 确保 `chat_session` 存在并更新会话态。
|
||||||
|
- `ChatService` 按 `runId` 创建 `diagnosis_run`。
|
||||||
|
- `AgentLoggingHook` 写入 `agent_step.run_id`。
|
||||||
|
- `ToolInvocationRecorder` 写入 `tool_invocation.run_id`。
|
||||||
|
- `SessionContextHolder` 或新的上下文 holder 同时携带 `sessionId + runId`。
|
||||||
|
- `backfillSessionMetrics` 按 `runId` 统计。
|
||||||
|
- `EvaluationService` 按 `runId` 读取工具调用。
|
||||||
|
|
||||||
|
验收:
|
||||||
|
|
||||||
|
- 同一 `sessionId` 连续两轮后,`diagnosis_run` 有两行不同 `run_id`。
|
||||||
|
- `/api/chat` 响应增加正式字段 `runId`。
|
||||||
|
- 两轮 `agent_step` / `tool_invocation` 分别按各自 `run_id` 查询。
|
||||||
|
- Redis `messagePairCount` 仍为 2,证明上下文不被破坏。
|
||||||
|
|
||||||
|
### Phase 3:Trace API 兼容和精确查询
|
||||||
|
|
||||||
|
目标:Trace API 可查最新 run,也可查指定 run,并能列出一个 session 下的 run。
|
||||||
|
|
||||||
|
任务:
|
||||||
|
|
||||||
|
- `GET /api/diagnosis/{sessionId}/trace` 从 `diagnosis_run` 默认解析最新 run。
|
||||||
|
- 增加 `runId` query 参数。
|
||||||
|
- Trace response 增加 `runId`。
|
||||||
|
- 新增 `GET /api/chat/session/{sessionId}/runs`。
|
||||||
|
|
||||||
|
验收:
|
||||||
|
|
||||||
|
- 不传 `runId` 返回最新 run。
|
||||||
|
- 传第一轮 `runId` 只返回第一轮 step/tool。
|
||||||
|
- 传第二轮 `runId` 只返回第二轮 step/tool。
|
||||||
|
- run 列表 API 只返回轻量 run 摘要,不展开 trace 明细。
|
||||||
|
- Demo 脚本和 Trace UI 的 `runId` 最小适配按 OpenSpec tasks 放到 Phase 6,避免 Phase 3 同时混入前端/脚本范围。
|
||||||
|
|
||||||
|
### Phase 4:Feedback 和 CaseLibrary 绑定 run
|
||||||
|
|
||||||
|
目标:反馈明确评价哪一轮诊断。
|
||||||
|
|
||||||
|
任务:
|
||||||
|
|
||||||
|
- Feedback request 支持 `runId`。
|
||||||
|
- 旧请求只有 `sessionId` 时短期绑定最新 run,并显式标记 fallback。
|
||||||
|
- `case_library.diagnosis_id` 新数据保存 `run_id`。
|
||||||
|
- `CaseLibraryService` 以 run 为来源生成 case,并用 `run_id` 做新数据幂等键。
|
||||||
|
- 文档说明 `diagnosis_id` 的过渡语义:旧数据可能是 `session_id`,新数据是 `run_id`。
|
||||||
|
|
||||||
|
验收:
|
||||||
|
|
||||||
|
- 同 session 多轮后,对第一轮提交 feedback 不会覆盖第二轮。
|
||||||
|
- useful 生成 case 时能定位到对应 run 的 query/answer。
|
||||||
|
|
||||||
|
### Phase 5:AIOps 同步 run 隔离
|
||||||
|
|
||||||
|
目标:AIOps 使用同样的 run 语义,避免另一条入口继续混杂。
|
||||||
|
|
||||||
|
任务:
|
||||||
|
|
||||||
|
- `AiOpsService` 生成并返回/透出 `runId`。
|
||||||
|
- AIOps `agent_step` / `tool_invocation` 按 `runId` 隔离。
|
||||||
|
- AIOps rule evaluation 写入当前 run 的 `diagnosis_run.self_evaluation.aiops_rule_evaluation`。
|
||||||
|
- AIOps Trace 查询兼容 `sessionId + runId`。
|
||||||
|
|
||||||
|
验收:
|
||||||
|
|
||||||
|
- 同一 AIOps `sessionId` 重跑不会混合 step/tool。
|
||||||
|
- AIOps rule evaluation 只读取当前 run 工具调用。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 暂不做
|
||||||
|
|
||||||
|
1. 暂不新增 `diagnosis_trace` 或 `trace_event` 主表。
|
||||||
|
2. 暂不做完整 run 列表 UI。
|
||||||
|
3. 暂不删除历史 Trace 数据。
|
||||||
|
4. 暂不改变 Redis 多轮上下文窗口策略。
|
||||||
|
5. 暂不立即物理删除旧 `diagnosis_session` 表。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 风险
|
||||||
|
|
||||||
|
### 1. 兼容风险
|
||||||
|
|
||||||
|
现有脚本、Trace 页面和反馈接口可能只知道 `sessionId`。
|
||||||
|
|
||||||
|
缓解:保留 `sessionId` 默认查最新 run 的行为。
|
||||||
|
|
||||||
|
### 2. 异步上下文风险
|
||||||
|
|
||||||
|
工具调用和 Agent hook 依赖 ThreadLocal / RunnableConfig 传递上下文。
|
||||||
|
|
||||||
|
缓解:统一上下文对象,明确 `sessionId` 和 `runId` 必须同时传递。
|
||||||
|
|
||||||
|
### 3. 历史数据回填风险
|
||||||
|
|
||||||
|
旧数据没有真实 run 边界,只能按当前 `diagnosis_session` 生成一条兼容 `diagnosis_run`。
|
||||||
|
|
||||||
|
缓解:旧数据视为单 run,不尝试拆分历史混合数据。
|
||||||
|
|
||||||
|
### 4. 评分口径变化风险
|
||||||
|
|
||||||
|
按 `runId` 隔离后,工具调用数和 evidence score 可能下降,但语义更正确。
|
||||||
|
|
||||||
|
缓解:更新 eval fixture 和 baseline,记录这是预期行为变化。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 已收敛问题
|
||||||
|
|
||||||
|
本 issue 当前已经收敛以下设计边界:
|
||||||
|
|
||||||
|
- `runId` 是正式 API 字段。
|
||||||
|
- `chat_session` 和 `diagnosis_run` 拆分为两张主表。
|
||||||
|
- Trace 明细继续由 `agent_step` / `tool_invocation` 承载。
|
||||||
|
- `chat_session` 只保存元数据,不保存完整对话历史。
|
||||||
|
- 旧 `diagnosis_session` 保留但新代码不再写入。
|
||||||
|
- 提供轻量 run 列表 API。
|
||||||
|
- Feedback 缺少 `runId` 时短期兼容、长期收紧。
|
||||||
|
- Demo 脚本和 Trace UI 做 `runId` 最小支持。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 相关文件
|
||||||
|
|
||||||
|
- `mvp/architecture/session-trace-lifecycle.md`
|
||||||
|
- `mvp/architecture/data-model.md`
|
||||||
|
- `mvp/tables/聊天会话表-chat_session.md`
|
||||||
|
- `mvp/tables/诊断运行表-diagnosis_run.md`
|
||||||
|
- `mvp/tables/诊断会话表-diagnosis_session.md`
|
||||||
|
- `mvp/tables/Agent步骤表-agent_step.md`
|
||||||
|
- `mvp/tables/工具调用表-tool_invocation.md`
|
||||||
|
- `mvp/tables/案例库表-case_library.md`
|
||||||
|
- `src/main/java/com/superbiz/agent/controller/ChatController.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/EvaluationService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/FeedbackService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/CaseLibraryService.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/hook/AgentLoggingHook.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
|
||||||
|
- `src/main/java/com/superbiz/agent/util/SessionContextHolder.java`
|
||||||
|
- `src/main/resources/db/migration/V005__create_session_storage.sql`
|
||||||
@@ -0,0 +1,45 @@
|
|||||||
|
# Agent 步骤表:agent_step
|
||||||
|
|
||||||
|
**状态**:当前表
|
||||||
|
**来源**:`V005__create_session_storage.sql`、`V006__fix_agent_step_json_to_text.sql`、`AgentStep`
|
||||||
|
|
||||||
|
## 定位
|
||||||
|
|
||||||
|
`agent_step` 记录一次诊断运行中每个 Agent 步骤的模型输入、输出、耗时和 Token 消耗。`run_id` 是执行隔离边界;Trace 页面展示顺序以 Trace API 返回顺序为准。
|
||||||
|
|
||||||
|
## 字段
|
||||||
|
|
||||||
|
| 字段 | 类型 | 必填 | 说明 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `id` | BIGINT | 是 | 自增主键 |
|
||||||
|
| `session_id` | VARCHAR(64) | 是 | 所属会话目录 ID,保留用于粗粒度过滤和兼容 |
|
||||||
|
| `run_id` | VARCHAR(64) | 否 | 所属 `diagnosis_run.run_id`;新执行应写入 |
|
||||||
|
| `step_index` | INT | 是 | 步骤序号,从 0 开始 |
|
||||||
|
| `agent_name` | VARCHAR(32) | 是 | Agent 名称,例如 planner、executor、verifier、composer |
|
||||||
|
| `model_input` | TEXT | 否 | 模型输入摘要;`V006` 已从 JSON 改为 TEXT |
|
||||||
|
| `model_output` | TEXT | 否 | 模型输出摘要;`V006` 已从 JSON 改为 TEXT |
|
||||||
|
| `thought` | TEXT | 否 | Agent 思考过程或调试摘要 |
|
||||||
|
| `has_tool_call` | BOOLEAN | 否 | 本步骤是否触发工具调用 |
|
||||||
|
| `duration_ms` | INT | 否 | 本步骤耗时 |
|
||||||
|
| `token_count` | INT | 否 | 本步骤 Token 消耗 |
|
||||||
|
| `created_at` | DATETIME | 是 | 创建时间 |
|
||||||
|
|
||||||
|
## 索引
|
||||||
|
|
||||||
|
| 索引 | 字段 | 用途 |
|
||||||
|
|---|---|---|
|
||||||
|
| `idx_session_step` | `session_id, step_index` | 历史兼容和粗粒度排查 |
|
||||||
|
| `idx_agent_step_run_step` | `run_id, step_index` | 按运行筛选步骤并辅助顺序查询 |
|
||||||
|
| `idx_agent_name` | `agent_name` | 按 Agent 类型筛选 |
|
||||||
|
|
||||||
|
## 关系
|
||||||
|
|
||||||
|
- `agent_step.run_id` 逻辑关联 `diagnosis_run.run_id`。
|
||||||
|
- `agent_step.session_id` 保留为 `chat_session.session_id` 的冗余关联,便于粗粒度过滤和兼容查询。
|
||||||
|
- `tool_invocation.step_id` 可关联 `agent_step.id`,但当前允许为空且不强制外键。
|
||||||
|
|
||||||
|
## 注意点
|
||||||
|
|
||||||
|
- 前端展示步骤时应使用 Trace API 返回顺序;服务端会在同一 `run_id` 范围内整理步骤顺序。
|
||||||
|
- 新 Trace、Verifier 和评测读路径应按 `run_id` 取数,避免同一 `sessionId` 多轮诊断混入。
|
||||||
|
- Verifier 应在 Executor 循环完成后出现;如果 `step_index` 中 Verifier 提前,通常意味着编排或记录顺序有问题。
|
||||||
@@ -0,0 +1,46 @@
|
|||||||
|
# MVP 数据表索引
|
||||||
|
|
||||||
|
**更新日期**:2026-07-10
|
||||||
|
**状态**:当前表文档入口
|
||||||
|
|
||||||
|
本目录保存当前 MVP 使用的数据表说明。详细结构以 Flyway migration 和实体类为准;本目录用于面试讲解、排查索引和快速理解数据流。
|
||||||
|
|
||||||
|
## 当前表
|
||||||
|
|
||||||
|
| 表 | 用途 | 文档 |
|
||||||
|
|---|---|---|
|
||||||
|
| `chat_session` | 会话目录元数据,保存同一个 `sessionId` 的多轮会话状态快照 | [聊天会话表-chat_session.md](聊天会话表-chat_session.md) |
|
||||||
|
| `diagnosis_run` | 运行级主记录,保存一次 Chat/AIOps 诊断的 query、状态、答案、自评估和反馈 | [诊断运行表-diagnosis_run.md](诊断运行表-diagnosis_run.md) |
|
||||||
|
| `agent_step` | Agent 步骤记录,按 `run_id` 隔离回放执行链路 | [Agent步骤表-agent_step.md](Agent步骤表-agent_step.md) |
|
||||||
|
| `tool_invocation` | 工具调用记录,按 `run_id` 支撑 Trace、Verifier 和评测 | [工具调用表-tool_invocation.md](工具调用表-tool_invocation.md) |
|
||||||
|
| `api_document` | 知识库文档元数据,和向量库 chunk 通过 `doc_id` 关联 | [文档元数据表-api_document.md](文档元数据表-api_document.md) |
|
||||||
|
| `knowledge_domain` | 知识域元数据,支撑 RAG domain hint 和检索策略 | [知识域表-knowledge_domain.md](知识域表-knowledge_domain.md) |
|
||||||
|
| `case_library` | 用户反馈沉淀出的高质量诊断案例 | [案例库表-case_library.md](案例库表-case_library.md) |
|
||||||
|
| `diagnosis_session` | 历史兼容和回滚表,新执行写入不再依赖它 | [诊断会话表-diagnosis_session.md](诊断会话表-diagnosis_session.md) |
|
||||||
|
|
||||||
|
## 已归档表
|
||||||
|
|
||||||
|
| 表 | 归档原因 | 文档 |
|
||||||
|
|---|---|---|
|
||||||
|
| `diagnosis_record` | 已由 `V007` 删除,历史上被 `diagnosis_session + agent_step + tool_invocation` 替代;当前新模型是 `chat_session + diagnosis_run + trace detail` | [archive/2026-07-09-doc-cleanup/旧诊断记录表-diagnosis_record.md](archive/2026-07-09-doc-cleanup/旧诊断记录表-diagnosis_record.md) |
|
||||||
|
|
||||||
|
## 核心关系
|
||||||
|
|
||||||
|
```text
|
||||||
|
chat_session.session_id
|
||||||
|
-> diagnosis_run.session_id
|
||||||
|
-> agent_step.run_id
|
||||||
|
-> tool_invocation.run_id
|
||||||
|
-> case_library.diagnosis_id (new AUTO cases use run_id)
|
||||||
|
|
||||||
|
diagnosis_session.session_id
|
||||||
|
-> historical compatibility / rollback only
|
||||||
|
|
||||||
|
api_document.doc_id
|
||||||
|
-> vector chunk metadata.docId / doc_id
|
||||||
|
|
||||||
|
knowledge_domain.domain_id
|
||||||
|
-> api_document metadata.category / vector chunk metadata.category
|
||||||
|
```
|
||||||
|
|
||||||
|
当前实现主要使用逻辑关联,不依赖数据库外键。
|
||||||
@@ -1,332 +0,0 @@
|
|||||||
# api_document - 文档元数据表
|
|
||||||
|
|
||||||
## 表定位
|
|
||||||
|
|
||||||
**文档管理表**:管理接口文档的元信息,不负责文档检索(检索由 Milvus 负责)
|
|
||||||
|
|
||||||
## 设计理念
|
|
||||||
|
|
||||||
### 文档管理,不是文档检索
|
|
||||||
|
|
||||||
**核心定位**:
|
|
||||||
- MySQL 负责文档元数据管理(状态、版本、去重)
|
|
||||||
- Milvus 负责文档内容存储和检索
|
|
||||||
- 通过 doc_id 关联两者
|
|
||||||
|
|
||||||
**MVP版本原则**:
|
|
||||||
- ✅ 最简字段,满足基本管理需求
|
|
||||||
- ✅ 文件去重(基于 file_hash)
|
|
||||||
- ✅ 状态追踪(索引进度)
|
|
||||||
- ✅ 硬删除(同步删除 Milvus 数据)
|
|
||||||
- ❌ 暂不支持:软删除、启用开关、版本管理(Phase 2)
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 表结构(MVP版)
|
|
||||||
|
|
||||||
```sql
|
|
||||||
CREATE TABLE api_document (
|
|
||||||
-- 主键
|
|
||||||
id BIGINT PRIMARY KEY AUTO_INCREMENT,
|
|
||||||
doc_id VARCHAR(64) UNIQUE NOT NULL COMMENT '文档唯一ID(UUID),关联Milvus',
|
|
||||||
|
|
||||||
-- 文档分类
|
|
||||||
fault_category VARCHAR(32) DEFAULT 'EXTERNAL_API' COMMENT '文档类别',
|
|
||||||
fault_source VARCHAR(128) COMMENT '文档归属(省份/服务名)',
|
|
||||||
api_name VARCHAR(128) COMMENT '接口名称',
|
|
||||||
version VARCHAR(32) DEFAULT 'v1.0' COMMENT '文档版本',
|
|
||||||
|
|
||||||
-- 文件信息
|
|
||||||
file_name VARCHAR(256) NOT NULL COMMENT '原始文件名',
|
|
||||||
file_path VARCHAR(512) COMMENT '文件存储路径',
|
|
||||||
file_hash VARCHAR(64) COMMENT '文件MD5 hash(用于去重)',
|
|
||||||
file_size BIGINT COMMENT '文件大小(字节)',
|
|
||||||
|
|
||||||
-- 索引状态
|
|
||||||
status VARCHAR(16) DEFAULT 'PENDING' COMMENT '索引状态(PENDING/PROCESSING/INDEXED/FAILED)',
|
|
||||||
chunk_count INT DEFAULT 0 COMMENT '分块数量',
|
|
||||||
error_message TEXT COMMENT '失败原因',
|
|
||||||
|
|
||||||
-- 时间字段
|
|
||||||
indexed_at DATETIME COMMENT '索引完成时间',
|
|
||||||
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
|
|
||||||
updated_at DATETIME DEFAULT CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP,
|
|
||||||
|
|
||||||
-- 索引
|
|
||||||
UNIQUE INDEX uk_file_hash (file_hash),
|
|
||||||
INDEX idx_doc_id (doc_id),
|
|
||||||
INDEX idx_fault_source (fault_source),
|
|
||||||
INDEX idx_status (status),
|
|
||||||
INDEX idx_created_at (created_at)
|
|
||||||
) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4 COMMENT='文档元数据表(MVP版)';
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 字段说明
|
|
||||||
|
|
||||||
| 字段 | 类型 | 必填 | 说明 |
|
|
||||||
|------|------|------|------|
|
|
||||||
| doc_id | VARCHAR(64) | 是 | **核心**:文档唯一ID,关联 Milvus |
|
|
||||||
| fault_category | VARCHAR(32) | 否 | 文档类别 |
|
|
||||||
| fault_source | VARCHAR(128) | 否 | 文档归属(省份/服务名)|
|
|
||||||
| api_name | VARCHAR(128) | 否 | 接口名称 |
|
|
||||||
| version | VARCHAR(32) | 否 | 文档版本 |
|
|
||||||
| file_name | VARCHAR(256) | 是 | 原始文件名 |
|
|
||||||
| file_path | VARCHAR(512) | 否 | 文件存储路径 |
|
|
||||||
| file_hash | VARCHAR(64) | 否 | **去重关键**:文件MD5 |
|
|
||||||
| file_size | BIGINT | 否 | 文件大小 |
|
|
||||||
| status | VARCHAR(16) | 是 | **状态追踪**:PENDING/PROCESSING/INDEXED/FAILED |
|
|
||||||
| chunk_count | INT | 否 | 分块数量 |
|
|
||||||
| error_message | TEXT | 否 | 失败原因 |
|
|
||||||
| indexed_at | DATETIME | 否 | 索引完成时间 |
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 核心设计决策
|
|
||||||
|
|
||||||
### 1. doc_id:MySQL 与 Milvus 的桥梁
|
|
||||||
|
|
||||||
```
|
|
||||||
作用:
|
|
||||||
- MySQL:通过 doc_id 管理文档元数据
|
|
||||||
- Milvus:每个 chunk 的 metadata 中携带 doc_id
|
|
||||||
|
|
||||||
关联关系:
|
|
||||||
api_document (MySQL)
|
|
||||||
doc_id: doc-001
|
|
||||||
↓ 1:N
|
|
||||||
Milvus chunks
|
|
||||||
chunk_1: {doc_id: 'doc-001', text: '...', vector: [...]}
|
|
||||||
chunk_2: {doc_id: 'doc-001', text: '...', vector: [...]}
|
|
||||||
|
|
||||||
管理操作:
|
|
||||||
- 删除文档:
|
|
||||||
DELETE FROM milvus_collection WHERE metadata["doc_id"] == 'doc-001';
|
|
||||||
DELETE FROM api_document WHERE doc_id = 'doc-001';
|
|
||||||
```
|
|
||||||
|
|
||||||
### 2. file_hash:文件去重
|
|
||||||
|
|
||||||
```
|
|
||||||
去重流程:
|
|
||||||
1. 用户上传文件
|
|
||||||
↓
|
|
||||||
2. 计算文件 MD5
|
|
||||||
file_hash = md5(file_content)
|
|
||||||
↓
|
|
||||||
3. 检查是否已存在
|
|
||||||
SELECT * FROM api_document WHERE file_hash = 'abc123...';
|
|
||||||
↓
|
|
||||||
4a. 如果存在 → 提示"文档已存在"
|
|
||||||
4b. 如果不存在 → 继续导入
|
|
||||||
|
|
||||||
唯一约束:UNIQUE INDEX uk_file_hash (file_hash)
|
|
||||||
```
|
|
||||||
|
|
||||||
### 3. status:状态追踪
|
|
||||||
|
|
||||||
```
|
|
||||||
状态流转:
|
|
||||||
PENDING (待处理)
|
|
||||||
↓
|
|
||||||
PROCESSING (处理中)
|
|
||||||
↓ 成功
|
|
||||||
INDEXED (已索引)
|
|
||||||
↓ 失败
|
|
||||||
FAILED (失败)
|
|
||||||
|
|
||||||
用途:
|
|
||||||
- 批量导入时监控进度
|
|
||||||
- 失败重试
|
|
||||||
- 统计索引成功率
|
|
||||||
```
|
|
||||||
|
|
||||||
### 4. 硬删除策略(MVP)
|
|
||||||
|
|
||||||
```
|
|
||||||
删除文档时:
|
|
||||||
1. 删除 Milvus 中的所有分块
|
|
||||||
2. 删除 MySQL 元数据
|
|
||||||
3. 可选:删除原始文件
|
|
||||||
|
|
||||||
特点:
|
|
||||||
- 简单直接
|
|
||||||
- 数据彻底删除
|
|
||||||
- 不可恢复(需谨慎)
|
|
||||||
|
|
||||||
Phase 2 可增强:
|
|
||||||
- 软删除(archived_at)
|
|
||||||
- 启用开关(enabled)
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 数据流
|
|
||||||
|
|
||||||
### 场景1:导入新文档
|
|
||||||
|
|
||||||
```
|
|
||||||
1. 用户上传文件
|
|
||||||
↓
|
|
||||||
2. 计算 hash
|
|
||||||
↓
|
|
||||||
3. 检查去重(MySQL)
|
|
||||||
↓
|
|
||||||
4. 插入元数据(status=PROCESSING)
|
|
||||||
↓
|
|
||||||
5. 后台处理:解析 → 分块 → 向量化 → 存入 Milvus
|
|
||||||
↓
|
|
||||||
6. 更新状态(status=INDEXED, chunk_count=15)
|
|
||||||
```
|
|
||||||
|
|
||||||
### 场景2:删除文档
|
|
||||||
|
|
||||||
```
|
|
||||||
1. 用户删除文档
|
|
||||||
↓
|
|
||||||
2. 删除 Milvus 数据(WHERE metadata["doc_id"] == 'xxx')
|
|
||||||
↓
|
|
||||||
3. 删除 MySQL 元数据
|
|
||||||
↓
|
|
||||||
4. 可选:删除原始文件
|
|
||||||
```
|
|
||||||
|
|
||||||
### 场景3:重新索引
|
|
||||||
|
|
||||||
```
|
|
||||||
1. 删除旧数据(Milvus + MySQL)
|
|
||||||
↓
|
|
||||||
2. 重新导入(同场景1)
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 典型查询
|
|
||||||
|
|
||||||
```sql
|
|
||||||
-- 查看文档列表
|
|
||||||
SELECT doc_id, file_name, version, status, chunk_count, indexed_at
|
|
||||||
FROM api_document
|
|
||||||
WHERE fault_source = '广东'
|
|
||||||
AND status = 'INDEXED'
|
|
||||||
ORDER BY indexed_at DESC;
|
|
||||||
|
|
||||||
-- 查询失败的文档
|
|
||||||
SELECT doc_id, file_name, error_message
|
|
||||||
FROM api_document
|
|
||||||
WHERE status = 'FAILED';
|
|
||||||
|
|
||||||
-- 统计各状态文档数量
|
|
||||||
SELECT status, COUNT(*) as count
|
|
||||||
FROM api_document
|
|
||||||
GROUP BY status;
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 与 Milvus 的协作
|
|
||||||
|
|
||||||
### Milvus Collection Schema
|
|
||||||
|
|
||||||
```python
|
|
||||||
{
|
|
||||||
"collection_name": "api_doc_collection",
|
|
||||||
"fields": [
|
|
||||||
{"name": "id", "type": "VARCHAR", "is_primary": true},
|
|
||||||
{"name": "content", "type": "VARCHAR"},
|
|
||||||
{"name": "vector", "type": "FLOAT_VECTOR", "dim": 1536},
|
|
||||||
{"name": "metadata", "type": "JSON"}
|
|
||||||
]
|
|
||||||
}
|
|
||||||
|
|
||||||
# metadata 结构
|
|
||||||
{
|
|
||||||
"doc_id": "doc-001", # 关联 MySQL
|
|
||||||
"_source": "/path/to/file",
|
|
||||||
"_file_name": "xxx.docx",
|
|
||||||
"chunkIndex": 0,
|
|
||||||
"totalChunks": 15
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Java 代码示例
|
|
||||||
|
|
||||||
```java
|
|
||||||
// 插入时携带 doc_id
|
|
||||||
Map<String, Object> metadata = new HashMap<>();
|
|
||||||
metadata.put("doc_id", docId); // 关联 MySQL
|
|
||||||
metadata.put("_source", filePath);
|
|
||||||
metadata.put("chunkIndex", chunkIndex);
|
|
||||||
|
|
||||||
// 删除文档的所有分块
|
|
||||||
String expr = String.format("metadata[\"doc_id\"] == \"%s\"", docId);
|
|
||||||
milvusClient.delete(DeleteParam.newBuilder()
|
|
||||||
.withCollectionName(COLLECTION_NAME)
|
|
||||||
.withExpr(expr)
|
|
||||||
.build());
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 数据示例
|
|
||||||
|
|
||||||
```sql
|
|
||||||
-- 外部接口文档
|
|
||||||
INSERT INTO api_document VALUES
|
|
||||||
(1, 'doc-001', 'EXTERNAL_API', '广东', '社保查询', 'v2.1',
|
|
||||||
'广东社保查询v2.1.docx', '/docs/guangdong/social-v2.1.docx',
|
|
||||||
'abc123...', 1048576,
|
|
||||||
'INDEXED', 15, NULL, '2024-06-15 10:30:00', NOW(), NOW());
|
|
||||||
|
|
||||||
-- 内部服务文档
|
|
||||||
INSERT INTO api_document VALUES
|
|
||||||
(2, 'doc-002', 'INTERNAL_ERROR', 'order-service', '订单服务API', 'v1.0',
|
|
||||||
'订单服务API文档.pdf', '/docs/internal/order-service-api.pdf',
|
|
||||||
'def456...', 2097152,
|
|
||||||
'INDEXED', 20, NULL, '2024-06-14 15:20:00', NOW(), NOW());
|
|
||||||
|
|
||||||
-- 处理失败的文档
|
|
||||||
INSERT INTO api_document VALUES
|
|
||||||
(3, 'doc-003', 'EXTERNAL_API', '江苏', '公积金查询', 'v1.5',
|
|
||||||
'江苏公积金查询.html', '/docs/jiangsu/fund-v1.5.html',
|
|
||||||
'ghi789...', 512000,
|
|
||||||
'FAILED', 0, '不支持HTML格式', NULL, NOW(), NOW());
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 数据量预估
|
|
||||||
|
|
||||||
```
|
|
||||||
预估:100-200 条
|
|
||||||
- 外部接口文档:50-100 条
|
|
||||||
- 内部服务文档:20-50 条
|
|
||||||
- 其他文档:30-50 条
|
|
||||||
|
|
||||||
存储:
|
|
||||||
- 单条记录:约 1KB
|
|
||||||
- 200 条:约 200KB
|
|
||||||
|
|
||||||
结论:数据量很小
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## MVP 版本的简化
|
|
||||||
|
|
||||||
```
|
|
||||||
Phase 1(当前):
|
|
||||||
✅ 基础字段和表结构
|
|
||||||
✅ 文件去重(file_hash)
|
|
||||||
✅ 状态追踪(status)
|
|
||||||
✅ 硬删除
|
|
||||||
✅ 通过 doc_id 关联 Milvus
|
|
||||||
|
|
||||||
Phase 2(未来增强):
|
|
||||||
❌ enabled(启用开关)
|
|
||||||
❌ archived_at(软删除)
|
|
||||||
❌ batch_id(批次管理)
|
|
||||||
❌ status 细化
|
|
||||||
❌ tags(标签分类)
|
|
||||||
```
|
|
||||||
+3
-1
@@ -1,4 +1,6 @@
|
|||||||
# diagnosis_record - 诊断记录表
|
# diagnosis_record - 旧诊断记录表
|
||||||
|
|
||||||
|
> 归档说明:`diagnosis_record` 已在 `V007__drop_diagnosis_record.sql` 中删除,当前主模型是 `diagnosis_session + agent_step + tool_invocation`。本文只用于追溯早期设计。
|
||||||
|
|
||||||
## 表定位
|
## 表定位
|
||||||
|
|
||||||
@@ -1,265 +0,0 @@
|
|||||||
# case_library - 案例库表
|
|
||||||
|
|
||||||
## 表定位
|
|
||||||
|
|
||||||
**知识沉淀表**:存储高质量诊断案例,支持相似案例推荐
|
|
||||||
|
|
||||||
## 设计理念
|
|
||||||
|
|
||||||
### 知识沉淀,系统越用越智能
|
|
||||||
|
|
||||||
**核心价值**:
|
|
||||||
- 质量过滤:只存储高质量案例(成功诊断 + 用户反馈有用)
|
|
||||||
- 知识沉淀:历史诊断经验可复用
|
|
||||||
- 提升准确率:相似问题提供历史参考
|
|
||||||
- 加速诊断:快速推荐相似案例
|
|
||||||
|
|
||||||
**MVP版本设计原则**:
|
|
||||||
- ✅ 能用:满足基本案例推荐功能
|
|
||||||
- ✅ 简单:字段不多,逻辑清晰
|
|
||||||
- ✅ 可扩展:后续可增加字段
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 表结构(MVP版)
|
|
||||||
|
|
||||||
```sql
|
|
||||||
CREATE TABLE case_library (
|
|
||||||
-- 主键
|
|
||||||
id BIGINT PRIMARY KEY AUTO_INCREMENT,
|
|
||||||
case_id VARCHAR(64) UNIQUE NOT NULL COMMENT '案例唯一ID(UUID)',
|
|
||||||
|
|
||||||
-- 来源关联
|
|
||||||
diagnosis_id VARCHAR(64) COMMENT '关联诊断记录(可选,人工录入时为空)',
|
|
||||||
source_type VARCHAR(16) DEFAULT 'AUTO' COMMENT '来源类型(AUTO:自动生成/MANUAL:人工录入)',
|
|
||||||
|
|
||||||
-- 案例分类
|
|
||||||
fault_category VARCHAR(32) COMMENT '故障类别(EXTERNAL_API/INTERNAL_ERROR/DATABASE...)',
|
|
||||||
fault_source VARCHAR(128) COMMENT '故障源(省份/服务名/类名...)',
|
|
||||||
fault_target VARCHAR(256) COMMENT '故障目标(接口URL/方法名/SQL...)',
|
|
||||||
error_code VARCHAR(64) COMMENT '错误码',
|
|
||||||
|
|
||||||
-- 案例内容
|
|
||||||
title VARCHAR(256) NOT NULL COMMENT '案例标题(简短描述)',
|
|
||||||
root_cause TEXT NOT NULL COMMENT '根因分析',
|
|
||||||
solution TEXT NOT NULL COMMENT '解决方案',
|
|
||||||
|
|
||||||
-- 简单统计
|
|
||||||
reference_count INT DEFAULT 0 COMMENT '引用次数(被推荐的次数)',
|
|
||||||
|
|
||||||
-- 元数据
|
|
||||||
created_by VARCHAR(64) COMMENT '创建人',
|
|
||||||
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
|
|
||||||
updated_at DATETIME DEFAULT CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP,
|
|
||||||
|
|
||||||
-- 索引
|
|
||||||
INDEX idx_fault_category (fault_category),
|
|
||||||
INDEX idx_error_code (error_code),
|
|
||||||
INDEX idx_fault_source (fault_source),
|
|
||||||
INDEX idx_fault_target (fault_target(100)),
|
|
||||||
INDEX idx_diagnosis_id (diagnosis_id),
|
|
||||||
INDEX idx_reference_count (reference_count),
|
|
||||||
INDEX idx_created_at (created_at)
|
|
||||||
) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4 COMMENT='案例库表(MVP版)';
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 字段说明
|
|
||||||
|
|
||||||
| 字段 | 类型 | 必填 | 说明 |
|
|
||||||
|------|------|------|------|
|
|
||||||
| case_id | VARCHAR(64) | 是 | 案例唯一标识(UUID)|
|
|
||||||
| diagnosis_id | VARCHAR(64) | 否 | 关联诊断记录(人工录入时为空)|
|
|
||||||
| source_type | VARCHAR(16) | 是 | 来源:AUTO(自动)/MANUAL(人工)|
|
|
||||||
| fault_category | VARCHAR(32) | 否 | 故障类别 |
|
|
||||||
| fault_source | VARCHAR(128) | 否 | 故障源 |
|
|
||||||
| fault_target | VARCHAR(256) | 否 | 故障目标(与 diagnosis_record 一致)|
|
|
||||||
| error_code | VARCHAR(64) | 否 | 错误码 |
|
|
||||||
| title | VARCHAR(256) | 是 | 案例标题 |
|
|
||||||
| root_cause | TEXT | 是 | 根因分析(核心内容)|
|
|
||||||
| solution | TEXT | 是 | 解决方案(核心内容)|
|
|
||||||
| reference_count | INT | 是 | 引用次数(用于排序)|
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 核心设计决策
|
|
||||||
|
|
||||||
### 1. 案例来源
|
|
||||||
|
|
||||||
```
|
|
||||||
来源1:自动生成(source_type=AUTO)
|
|
||||||
├─ 触发条件:诊断成功 + 用户反馈"有用"
|
|
||||||
├─ 关联诊断:diagnosis_id 不为空
|
|
||||||
└─ 质量保证:用户验证过
|
|
||||||
|
|
||||||
来源2:人工录入(source_type=MANUAL)
|
|
||||||
├─ 运维团队总结的经典案例
|
|
||||||
├─ diagnosis_id 为空
|
|
||||||
└─ 质量最高
|
|
||||||
|
|
||||||
注意:诊断失败或用户反馈"无用"的不自动生成案例
|
|
||||||
```
|
|
||||||
|
|
||||||
### 2. 简化的评分机制(MVP)
|
|
||||||
|
|
||||||
```
|
|
||||||
MVP版本:只按 reference_count 排序
|
|
||||||
- 引用次数多的排前面
|
|
||||||
- 简单有效
|
|
||||||
|
|
||||||
Phase 2 可增强:
|
|
||||||
- 增加 useful_count(用户反馈有用次数)
|
|
||||||
- 增加 score(综合评分)
|
|
||||||
- 增加 is_featured(人工标记的经典案例)
|
|
||||||
```
|
|
||||||
|
|
||||||
### 3. 与 diagnosis_record 的关系
|
|
||||||
|
|
||||||
```
|
|
||||||
关系:一对一(可选)
|
|
||||||
- 一次诊断 → 可以生成一个案例
|
|
||||||
- 通过 diagnosis_id 关联
|
|
||||||
- diagnosis_id 可为空(人工录入案例)
|
|
||||||
|
|
||||||
流程:
|
|
||||||
diagnosis_record(成功)
|
|
||||||
↓
|
|
||||||
用户反馈"有用"
|
|
||||||
↓
|
|
||||||
自动生成 case_library
|
|
||||||
↓
|
|
||||||
后续可人工修正、合并相似案例
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 数据示例
|
|
||||||
|
|
||||||
### 示例1:外部接口故障案例
|
|
||||||
```sql
|
|
||||||
INSERT INTO case_library VALUES
|
|
||||||
(1, 'case-001', 'diag-001', 'AUTO', 'EXTERNAL_API', '广东', '/api/v1/guangdong/social-security', '40003',
|
|
||||||
'广东社保查询idCard字段缺失',
|
|
||||||
'请求报文中未传入idCard字段,导致参数校验失败',
|
|
||||||
'前端表单增加idCard必填校验;后端增加参数校验提示',
|
|
||||||
15, 'system', NOW(), NOW());
|
|
||||||
```
|
|
||||||
|
|
||||||
### 示例2:内部错误案例
|
|
||||||
```sql
|
|
||||||
INSERT INTO case_library VALUES
|
|
||||||
(2, 'case-002', 'diag-045', 'AUTO', 'INTERNAL_ERROR', 'order-service', 'OrderController.createOrder()', 'NullPointerException',
|
|
||||||
'订单服务创建订单空指针异常',
|
|
||||||
'OrderController.createOrder()方法中user对象为null,未做空判断',
|
|
||||||
'在第45行添加空判断:if (user == null) throw new BizException("用户信息不存在")',
|
|
||||||
8, 'system', NOW(), NOW());
|
|
||||||
```
|
|
||||||
|
|
||||||
### 示例3:人工录入案例
|
|
||||||
```sql
|
|
||||||
INSERT INTO case_library VALUES
|
|
||||||
(3, 'case-003', NULL, 'MANUAL', 'DATABASE', 'mysql-master-01', 'UPDATE orders SET status=? WHERE order_id=?', '1213',
|
|
||||||
'订单库存更新死锁通用处理',
|
|
||||||
'两个事务互相等待对方释放锁',
|
|
||||||
'调整事务加锁顺序:统一先锁订单,再锁库存;或使用乐观锁',
|
|
||||||
3, 'admin', NOW(), NOW());
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 典型查询
|
|
||||||
|
|
||||||
### 精确匹配查询
|
|
||||||
```sql
|
|
||||||
-- 按错误码查询
|
|
||||||
SELECT * FROM case_library
|
|
||||||
WHERE error_code = '40003'
|
|
||||||
ORDER BY reference_count DESC
|
|
||||||
LIMIT 5;
|
|
||||||
|
|
||||||
-- 按故障类别 + 错误码 + 故障目标查询
|
|
||||||
SELECT * FROM case_library
|
|
||||||
WHERE fault_category = 'INTERNAL_ERROR'
|
|
||||||
AND error_code = 'NullPointerException'
|
|
||||||
AND fault_target = 'OrderController.createOrder()'
|
|
||||||
ORDER BY reference_count DESC
|
|
||||||
LIMIT 5;
|
|
||||||
```
|
|
||||||
|
|
||||||
### 统计分析
|
|
||||||
```sql
|
|
||||||
-- 统计案例分布
|
|
||||||
SELECT
|
|
||||||
fault_category,
|
|
||||||
COUNT(*) as count,
|
|
||||||
AVG(reference_count) as avg_reference
|
|
||||||
FROM case_library
|
|
||||||
GROUP BY fault_category
|
|
||||||
ORDER BY count DESC;
|
|
||||||
|
|
||||||
-- Top 引用案例
|
|
||||||
SELECT title, reference_count, created_at
|
|
||||||
FROM case_library
|
|
||||||
ORDER BY reference_count DESC
|
|
||||||
LIMIT 10;
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 与 Milvus 的配合
|
|
||||||
|
|
||||||
### 混合检索策略
|
|
||||||
|
|
||||||
```
|
|
||||||
1. 精确匹配(MySQL)
|
|
||||||
- 按 error_code 查询
|
|
||||||
- 按 fault_category + fault_source 查询
|
|
||||||
- 优点:快速、准确
|
|
||||||
|
|
||||||
2. 语义检索(Milvus)
|
|
||||||
- 将案例内容向量化
|
|
||||||
- 按语义相似度查询
|
|
||||||
- 优点:能找到相似但不同错误码的案例
|
|
||||||
|
|
||||||
3. 混合策略(推荐)
|
|
||||||
Step 1: 先精确匹配(MySQL)
|
|
||||||
Step 2: 如果结果 < 3 个,补充语义检索(Milvus)
|
|
||||||
Step 3: 合并去重,按 reference_count 排序
|
|
||||||
Step 4: 返回 Top 5
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 数据量预估
|
|
||||||
|
|
||||||
```
|
|
||||||
预估:500-1000 条
|
|
||||||
- 初期:每月新增 10-20 条
|
|
||||||
- 稳定期:每月新增 5-10 条
|
|
||||||
- 总量:1-2 年达到稳定
|
|
||||||
|
|
||||||
存储:
|
|
||||||
- 单条记录:约 2KB
|
|
||||||
- 1000 条:约 2MB
|
|
||||||
|
|
||||||
结论:数据量很小
|
|
||||||
```
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## MVP 版本的简化
|
|
||||||
|
|
||||||
```
|
|
||||||
Phase 1(当前):
|
|
||||||
✅ 基础字段和表结构
|
|
||||||
✅ 自动生成案例
|
|
||||||
✅ 人工录入案例
|
|
||||||
✅ 按 reference_count 简单排序
|
|
||||||
|
|
||||||
Phase 2(未来增强):
|
|
||||||
❌ useful_count + score(复杂评分)
|
|
||||||
❌ 版本管理
|
|
||||||
❌ 标签分类(tags)
|
|
||||||
❌ 案例合并功能
|
|
||||||
```
|
|
||||||
@@ -0,0 +1,69 @@
|
|||||||
|
# 工具调用表:tool_invocation
|
||||||
|
|
||||||
|
**状态**:当前表
|
||||||
|
**来源**:`V005__create_session_storage.sql`、`V010__add_relevance_level_to_tool_invocation.sql`、`ToolInvocation`
|
||||||
|
|
||||||
|
## 定位
|
||||||
|
|
||||||
|
`tool_invocation` 记录 Agent 在一次诊断运行中显式调用工具的事实,包括工具名、入参、输出摘要、检索层级、证据引用和失败信息。它是 Trace、Verifier、评测和人工排查的共同数据源。
|
||||||
|
|
||||||
|
## 字段
|
||||||
|
|
||||||
|
| 字段 | 类型 | 必填 | 说明 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `id` | BIGINT | 是 | 自增主键 |
|
||||||
|
| `session_id` | VARCHAR(64) | 是 | 所属会话目录 ID,保留用于粗粒度过滤和兼容 |
|
||||||
|
| `run_id` | VARCHAR(64) | 否 | 所属 `diagnosis_run.run_id`;新执行应写入 |
|
||||||
|
| `step_id` | BIGINT | 否 | 可关联 `agent_step.id` |
|
||||||
|
| `tool_name` | VARCHAR(64) | 是 | 工具名称,例如 `lookup_knowledge`、日志查询、指标查询 |
|
||||||
|
| `input_params` | JSON | 是 | 工具入参 |
|
||||||
|
| `output_preview` | TEXT | 否 | 工具输出摘要或前缀 |
|
||||||
|
| `output_length` | INT | 否 | 工具输出字符数 |
|
||||||
|
| `retrieval_layer` | VARCHAR(8) | 否 | 检索层级,例如 `L0`、`L1`、`L0+L1` |
|
||||||
|
| `l0_match_count` | INT | 否 | L0 命中数量 |
|
||||||
|
| `l1_match_count` | INT | 否 | L1 命中数量 |
|
||||||
|
| `is_truncated` | BOOLEAN | 否 | 输出是否被截断 |
|
||||||
|
| `relevance_level` | VARCHAR(20) | 否 | 归一化质量等级:`PRECISE`、`HIGHLY_RELEVANT`、`REFERENCE`、`DEDUPED` |
|
||||||
|
| `dedup_reason` | VARCHAR(32) | 否 | 去重原因,例如 `doc_retrieved`、`domain_retrieved` |
|
||||||
|
| `retrieval_details` | JSON | 否 | 检索明细、证据引用、Gatekeeper 可用导航信息 |
|
||||||
|
| `duration_ms` | INT | 否 | 工具耗时 |
|
||||||
|
| `success` | BOOLEAN | 否 | 工具是否成功 |
|
||||||
|
| `error_message` | TEXT | 否 | 失败原因 |
|
||||||
|
| `created_at` | DATETIME | 是 | 创建时间 |
|
||||||
|
|
||||||
|
## 索引
|
||||||
|
|
||||||
|
| 索引 | 字段 | 用途 |
|
||||||
|
|---|---|---|
|
||||||
|
| `idx_session_id` | `session_id` | 历史兼容和粗粒度排查 |
|
||||||
|
| `idx_tool_invocation_run_id` | `run_id, id` | Trace、Verifier、评测按运行查询工具调用 |
|
||||||
|
| `idx_tool_name` | `tool_name` | 按工具类型排查 |
|
||||||
|
| `idx_retrieval_layer` | `retrieval_layer` | 观察 RAG L0/L1 行为 |
|
||||||
|
|
||||||
|
## 关系
|
||||||
|
|
||||||
|
- `tool_invocation.run_id` 逻辑关联 `diagnosis_run.run_id`。
|
||||||
|
- `tool_invocation.session_id` 保留为 `chat_session.session_id` 的冗余关联,便于粗粒度过滤和兼容查询。
|
||||||
|
- `tool_invocation.step_id` 可关联 `agent_step.id`,但当前不强制。
|
||||||
|
|
||||||
|
## 关键 JSON
|
||||||
|
|
||||||
|
`retrieval_details` 是扩展字段。当前重要结构包括:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"evidence_status": "supported",
|
||||||
|
"evidence_refs": [
|
||||||
|
{
|
||||||
|
"raw_path": "$.logs[0]",
|
||||||
|
"text": "工具返回中可核对的最小证据文本"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## 注意点
|
||||||
|
|
||||||
|
- Verifier 不应只信任 RAG 证据;所有工具只要能提供 `evidence_refs`,都应该进入可校验证据链。
|
||||||
|
- `output_preview` 只适合展示和排查,不应被当成完整原始输出。
|
||||||
|
- `$.no_evidence` 只代表“本次工具未命中证据”,不能推导为“故障不存在”。
|
||||||
@@ -0,0 +1,51 @@
|
|||||||
|
# 文档元数据表:api_document
|
||||||
|
|
||||||
|
**状态**:当前表
|
||||||
|
**来源**:`V003__create_api_document.sql`、`V004__add_metadata_to_api_document.sql`、`ApiDocument`
|
||||||
|
|
||||||
|
## 定位
|
||||||
|
|
||||||
|
`api_document` 是知识库文档的 MySQL 元数据表。它不保存向量正文,正文切片和向量检索由 Milvus/Zilliz collection 承担;两侧通过 `doc_id` 和 chunk metadata 逻辑关联。
|
||||||
|
|
||||||
|
## 字段
|
||||||
|
|
||||||
|
| 字段 | 类型 | 必填 | 说明 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `id` | BIGINT | 是 | 自增主键 |
|
||||||
|
| `doc_id` | VARCHAR(64) | 是 | 文档唯一 ID,关联向量库 chunk metadata |
|
||||||
|
| `fault_category` | VARCHAR(32) | 否 | 文档类别,默认 `EXTERNAL_API`;实体侧使用 `FaultCategory` |
|
||||||
|
| `fault_source` | VARCHAR(128) | 否 | 文档归属,例如服务名、省份或系统来源 |
|
||||||
|
| `api_name` | VARCHAR(128) | 否 | 接口或文档主题名称 |
|
||||||
|
| `version` | VARCHAR(32) | 否 | 文档版本,默认 `v1.0` |
|
||||||
|
| `file_name` | VARCHAR(256) | 是 | 原始文件名 |
|
||||||
|
| `file_path` | VARCHAR(512) | 否 | 文件存储路径 |
|
||||||
|
| `file_hash` | VARCHAR(64) | 否 | 文件 MD5,用于去重 |
|
||||||
|
| `file_size` | BIGINT | 否 | 文件大小,单位字节 |
|
||||||
|
| `status` | VARCHAR(16) | 否 | 索引状态:`PENDING`、`PROCESSING`、`INDEXED`、`FAILED` |
|
||||||
|
| `chunk_count` | INT | 否 | 向量库切片数量 |
|
||||||
|
| `error_message` | TEXT | 否 | 索引失败原因 |
|
||||||
|
| `metadata` | TEXT | 否 | frontmatter 元数据 JSON 字符串 |
|
||||||
|
| `indexed_at` | DATETIME | 否 | 索引完成时间 |
|
||||||
|
| `created_at` | DATETIME | 是 | 创建时间 |
|
||||||
|
| `updated_at` | DATETIME | 是 | 更新时间 |
|
||||||
|
|
||||||
|
## 索引
|
||||||
|
|
||||||
|
| 索引 | 字段 | 用途 |
|
||||||
|
|---|---|---|
|
||||||
|
| `uk_file_hash` | `file_hash` | 文件去重 |
|
||||||
|
| `idx_doc_id` | `doc_id` | 按文档 ID 查询 |
|
||||||
|
| `idx_fault_source` | `fault_source` | 按来源筛选 |
|
||||||
|
| `idx_status` | `status` | 查看索引状态 |
|
||||||
|
| `idx_created_at` | `created_at` | 按上传时间排序 |
|
||||||
|
|
||||||
|
## 关系
|
||||||
|
|
||||||
|
- `api_document.doc_id` 与向量库 chunk metadata 中的 `docId` / `doc_id` 逻辑关联。
|
||||||
|
- `knowledge_domain.domain_id` 与文档 metadata 中的 `category` 形成领域聚合关系;当前没有数据库外键。
|
||||||
|
|
||||||
|
## 注意点
|
||||||
|
|
||||||
|
- 删除文档时需要同时处理 MySQL 元数据和向量库 chunk。
|
||||||
|
- `metadata` 是 JSON 字符串,不是 MySQL JSON 列。
|
||||||
|
- 表字段以 Flyway 为准;实体默认值和枚举可能与迁移脚本的 SQL 默认值存在历史差异,排查时优先看实际迁移和数据库结构。
|
||||||
@@ -0,0 +1,51 @@
|
|||||||
|
# 案例库表:case_library
|
||||||
|
|
||||||
|
**状态**:当前表
|
||||||
|
**来源**:`V002__create_case_library.sql`、`CaseLibrary`
|
||||||
|
|
||||||
|
## 定位
|
||||||
|
|
||||||
|
`case_library` 保存高质量诊断案例,用于后续相似案例推荐和知识沉淀。当前自动沉淀路径来自 `useful` 用户反馈:新数据把 `diagnosis_run` 中的 query 和 answer 映射为案例内容。
|
||||||
|
|
||||||
|
## 字段
|
||||||
|
|
||||||
|
| 字段 | 类型 | 必填 | 说明 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `id` | BIGINT | 是 | 自增主键 |
|
||||||
|
| `case_id` | VARCHAR(64) | 是 | 案例唯一 ID |
|
||||||
|
| `diagnosis_id` | VARCHAR(64) | 否 | 关联诊断来源;新自动生成时存 `diagnosis_run.run_id`,历史数据可能是 `diagnosis_session.session_id` |
|
||||||
|
| `source_type` | VARCHAR(16) | 否 | 来源类型:`AUTO` 或 `MANUAL` |
|
||||||
|
| `fault_category` | VARCHAR(32) | 否 | 故障类别,实体侧使用 `FaultCategory` |
|
||||||
|
| `fault_source` | VARCHAR(128) | 否 | 故障源,例如服务、系统或省份 |
|
||||||
|
| `fault_target` | VARCHAR(256) | 否 | 故障目标,例如接口、方法、SQL 或组件 |
|
||||||
|
| `error_code` | VARCHAR(64) | 否 | 错误码或异常类型 |
|
||||||
|
| `title` | VARCHAR(256) | 是 | 案例标题 |
|
||||||
|
| `root_cause` | TEXT | 是 | 根因分析 |
|
||||||
|
| `solution` | TEXT | 是 | 解决方案 |
|
||||||
|
| `reference_count` | INT | 否 | 被推荐次数 |
|
||||||
|
| `created_by` | VARCHAR(64) | 否 | 创建人 |
|
||||||
|
| `created_at` | DATETIME | 是 | 创建时间 |
|
||||||
|
| `updated_at` | DATETIME | 是 | 更新时间 |
|
||||||
|
|
||||||
|
## 索引
|
||||||
|
|
||||||
|
| 索引 | 字段 | 用途 |
|
||||||
|
|---|---|---|
|
||||||
|
| `idx_fault_category` | `fault_category` | 按故障类别筛选 |
|
||||||
|
| `idx_error_code` | `error_code` | 按错误码精确匹配 |
|
||||||
|
| `idx_fault_source` | `fault_source` | 按故障源筛选 |
|
||||||
|
| `idx_fault_target` | `fault_target(100)` | 按故障目标筛选 |
|
||||||
|
| `idx_diagnosis_id` | `diagnosis_id` | 追溯来源运行或历史会话 |
|
||||||
|
| `idx_reference_count` | `reference_count` | 推荐排序 |
|
||||||
|
| `idx_created_at` | `created_at` | 时间排序 |
|
||||||
|
|
||||||
|
## 关系
|
||||||
|
|
||||||
|
- `case_library.diagnosis_id` 是过渡字段:新自动案例逻辑关联 `diagnosis_run.run_id`,历史自动案例可能仍是 `diagnosis_session.session_id`。
|
||||||
|
- 人工录入案例可以不填写 `diagnosis_id`。
|
||||||
|
|
||||||
|
## 注意点
|
||||||
|
|
||||||
|
- 旧文档里提到的 `diagnosis_record` 已被 `V007` 删除,不再是当前主模型。
|
||||||
|
- 查询新自动案例时优先按 `run_id` 追溯;遇到旧值时再按历史 `session_id` 解释。
|
||||||
|
- 当前自动沉淀仍比较粗:`root_cause` 和 `solution` 都可能来自完整 answer。后续可从结构化结论中拆分根因、证据和修复建议。
|
||||||
@@ -0,0 +1,36 @@
|
|||||||
|
# 知识域表:knowledge_domain
|
||||||
|
|
||||||
|
**状态**:当前表
|
||||||
|
**来源**:`V009__add_knowledge_domain.sql`、`KnowledgeDomain`
|
||||||
|
|
||||||
|
## 定位
|
||||||
|
|
||||||
|
`knowledge_domain` 保存知识库领域级元数据,用来帮助 Planner/Executor 判断什么时候检索某一类知识,并为 RAG 的 domain hint、去重和可观测性提供基础信息。
|
||||||
|
|
||||||
|
## 字段
|
||||||
|
|
||||||
|
| 字段 | 类型 | 必填 | 说明 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `id` | BIGINT | 是 | 自增主键 |
|
||||||
|
| `domain_id` | VARCHAR(64) | 是 | 领域 ID,通常对应文档 category,例如 `payment`、`infrastructure` |
|
||||||
|
| `description` | VARCHAR(256) | 否 | 领域描述 |
|
||||||
|
| `when_to_retrieve` | TEXT | 否 | 何时检索该领域的提示说明 |
|
||||||
|
| `document_count` | INT | 是 | 当前领域文档数量 |
|
||||||
|
| `created_at` | DATETIME | 是 | 创建时间 |
|
||||||
|
| `updated_at` | DATETIME | 是 | 更新时间 |
|
||||||
|
|
||||||
|
## 索引
|
||||||
|
|
||||||
|
| 索引 | 字段 | 用途 |
|
||||||
|
|---|---|---|
|
||||||
|
| unique | `domain_id` | 保证领域 ID 唯一 |
|
||||||
|
|
||||||
|
## 关系
|
||||||
|
|
||||||
|
- `knowledge_domain.domain_id` 与 `api_document.metadata` 或向量库 chunk metadata 中的 `category` 逻辑关联。
|
||||||
|
- 当前没有数据库外键,领域文档数量由服务逻辑维护。
|
||||||
|
|
||||||
|
## 注意点
|
||||||
|
|
||||||
|
- `when_to_retrieve` 是检索策略提示,不是事实证据。
|
||||||
|
- Executor / Verifier 不能把领域描述当作诊断结论依据;事实仍应来自工具返回的证据块或证据引用。
|
||||||
@@ -0,0 +1,40 @@
|
|||||||
|
# 聊天会话表:chat_session
|
||||||
|
|
||||||
|
**状态**:当前会话目录表
|
||||||
|
**来源**:`V011__add_session_run_isolation.sql`、`ChatSession`
|
||||||
|
|
||||||
|
## 定位
|
||||||
|
|
||||||
|
`chat_session` 保存多轮 Chat 会话的元数据,用于把同一个 `sessionId` 下的多次诊断运行组织在一起。它不保存完整对话历史;正文消息仍由 Redis `SessionContext.messageHistory` 管理。
|
||||||
|
|
||||||
|
## 字段
|
||||||
|
|
||||||
|
| 字段 | 类型 | 必填 | 说明 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `id` | BIGINT | 是 | 自增主键 |
|
||||||
|
| `session_id` | VARCHAR(64) | 是 | 会话目录 ID,外部 API 仍通过它定位会话 |
|
||||||
|
| `status` | VARCHAR(16) | 否 | `ACTIVE`、`EXPIRED`、`CLOSED` |
|
||||||
|
| `message_pair_count` | INT | 否 | Redis 会话中问答轮次数的快照 |
|
||||||
|
| `created_at` | DATETIME | 是 | 创建时间 |
|
||||||
|
| `last_active_at` | DATETIME | 否 | 最近活跃时间 |
|
||||||
|
| `expires_at` | DATETIME | 否 | 目录元数据,可为空;Redis 消息历史可独立过期 |
|
||||||
|
|
||||||
|
## 索引
|
||||||
|
|
||||||
|
| 索引 | 字段 | 用途 |
|
||||||
|
|---|---|---|
|
||||||
|
| `session_id` unique | `session_id` | 会话目录唯一约束 |
|
||||||
|
| `idx_chat_session_last_active` | `last_active_at` | 最近会话列表和排查 |
|
||||||
|
| `idx_chat_session_status` | `status` | 按状态筛选 |
|
||||||
|
| `idx_chat_session_expires_at` | `expires_at` | 过期目录排查 |
|
||||||
|
|
||||||
|
## 关系
|
||||||
|
|
||||||
|
- `diagnosis_run.session_id` 逻辑关联 `chat_session.session_id`。
|
||||||
|
- 当前不强制数据库外键,服务层校验 run/session ownership。
|
||||||
|
- 一个 `chat_session` 可以拥有多个 `diagnosis_run`。
|
||||||
|
|
||||||
|
## 注意点
|
||||||
|
|
||||||
|
- `chat_session` 是会话元数据,不是诊断执行记录。
|
||||||
|
- 不要把 query、answer、self_evaluation、feedback 写入该表;这些属于 `diagnosis_run`。
|
||||||
@@ -0,0 +1,48 @@
|
|||||||
|
# 诊断会话表:diagnosis_session
|
||||||
|
|
||||||
|
**状态**:历史兼容和回滚表
|
||||||
|
**来源**:`V005__create_session_storage.sql`、`V008__add_answer_to_diagnosis_session.sql`、`DiagnosisSession`
|
||||||
|
|
||||||
|
## 定位
|
||||||
|
|
||||||
|
`diagnosis_session` 是旧版 session 级诊断主记录。`V011` 之后,新 Chat/AIOps 执行的运行态写入已经切到 `chat_session + diagnosis_run`;本表保留用于历史兼容、迁移回填和回滚比较。
|
||||||
|
|
||||||
|
## 字段
|
||||||
|
|
||||||
|
| 字段 | 类型 | 必填 | 说明 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `id` | BIGINT | 是 | 自增主键 |
|
||||||
|
| `session_id` | VARCHAR(64) | 是 | 旧版会话唯一 ID,也是兼容 run 回填来源 |
|
||||||
|
| `query` | TEXT | 是 | 用户原始问题或 AIOps 输入摘要 |
|
||||||
|
| `status` | VARCHAR(16) | 否 | `PENDING`、`RUNNING`、`SUCCESS`、`FAILED` |
|
||||||
|
| `agent_flow` | VARCHAR(32) | 否 | `CHAT` 或 `AI_OPS` |
|
||||||
|
| `total_duration_ms` | INT | 否 | 总耗时,单位毫秒 |
|
||||||
|
| `total_token_count` | INT | 否 | 总 Token 消耗 |
|
||||||
|
| `step_count` | INT | 否 | Agent 步骤数 |
|
||||||
|
| `tool_call_count` | INT | 否 | 工具调用次数 |
|
||||||
|
| `answer` | LONGTEXT | 否 | 返回给用户的最终答案 |
|
||||||
|
| `self_evaluation` | JSON | 否 | rule、verifier、aiops 等自评估结果容器 |
|
||||||
|
| `feedback` | VARCHAR(16) | 否 | 用户反馈:`useful`、`not_useful` 或空 |
|
||||||
|
| `created_at` | DATETIME | 是 | 创建时间 |
|
||||||
|
| `updated_at` | DATETIME | 是 | 更新时间 |
|
||||||
|
|
||||||
|
## 索引
|
||||||
|
|
||||||
|
| 索引 | 字段 | 用途 |
|
||||||
|
|---|---|---|
|
||||||
|
| `session_id` unique | `session_id` | 会话唯一约束 |
|
||||||
|
| `idx_created_at` | `created_at` | 按时间查询 |
|
||||||
|
| `idx_status` | `status` | 按状态筛选 |
|
||||||
|
| `idx_agent_flow` | `agent_flow` | 区分 Chat / AIOps |
|
||||||
|
|
||||||
|
## 关系
|
||||||
|
|
||||||
|
- 迁移时每条 `diagnosis_session` 会生成一条兼容 `diagnosis_run`。
|
||||||
|
- 历史 `agent_step.run_id` 和 `tool_invocation.run_id` 会尽量从兼容 `diagnosis_run` 回填。
|
||||||
|
- 历史自动案例可能仍使用 `case_library.diagnosis_id = diagnosis_session.session_id`。
|
||||||
|
|
||||||
|
## 注意点
|
||||||
|
|
||||||
|
- 新执行不应再把 query、answer、self_evaluation、feedback、统计计数写入本表。
|
||||||
|
- 新 Trace 聚合优先读取 `diagnosis_run + agent_step.run_id + tool_invocation.run_id`。
|
||||||
|
- 历史 fallback 仅在没有 run-backed 数据时读取本表。
|
||||||
@@ -0,0 +1,51 @@
|
|||||||
|
# 诊断运行表:diagnosis_run
|
||||||
|
|
||||||
|
**状态**:当前诊断运行主表
|
||||||
|
**来源**:`V011__add_session_run_isolation.sql`、`DiagnosisRun`
|
||||||
|
|
||||||
|
## 定位
|
||||||
|
|
||||||
|
`diagnosis_run` 表示一次可回放的 Chat 或 AIOps 诊断执行。`run_id` 是运行级边界,Trace、反馈、自评估、案例沉淀和统计都应优先按 `run_id` 绑定。
|
||||||
|
|
||||||
|
## 字段
|
||||||
|
|
||||||
|
| 字段 | 类型 | 必填 | 说明 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `id` | BIGINT | 是 | 自增主键 |
|
||||||
|
| `run_id` | VARCHAR(64) | 是 | 运行唯一 ID,格式为 `run-` + UUID |
|
||||||
|
| `session_id` | VARCHAR(64) | 是 | 所属 `chat_session.session_id` |
|
||||||
|
| `query` | TEXT | 是 | 本次 Chat 问题或 AIOps 告警摘要 |
|
||||||
|
| `status` | VARCHAR(16) | 否 | `PENDING`、`RUNNING`、`SUCCESS`、`FAILED` |
|
||||||
|
| `agent_flow` | VARCHAR(32) | 否 | `CHAT` 或 `AI_OPS` |
|
||||||
|
| `answer` | LONGTEXT | 否 | 本次运行的最终答复或告警报告 |
|
||||||
|
| `self_evaluation` | JSON | 否 | 本次运行的 rule、verifier、aiops 自评估容器 |
|
||||||
|
| `feedback` | VARCHAR(16) | 否 | 本次运行的用户反馈 |
|
||||||
|
| `total_duration_ms` | INT | 否 | 本次运行总耗时 |
|
||||||
|
| `total_token_count` | INT | 否 | 本次运行 Token 消耗 |
|
||||||
|
| `step_count` | INT | 否 | 本次运行的 Agent 步骤数 |
|
||||||
|
| `tool_call_count` | INT | 否 | 本次运行的工具调用数 |
|
||||||
|
| `created_at` | DATETIME | 是 | 创建时间 |
|
||||||
|
| `updated_at` | DATETIME | 是 | 更新时间 |
|
||||||
|
|
||||||
|
## 索引
|
||||||
|
|
||||||
|
| 索引 | 字段 | 用途 |
|
||||||
|
|---|---|---|
|
||||||
|
| `run_id` unique | `run_id` | 运行唯一约束 |
|
||||||
|
| `idx_diagnosis_run_session_created` | `session_id, created_at, id` | session 下最新运行解析和运行列表 |
|
||||||
|
| `idx_diagnosis_run_session_run` | `session_id, run_id` | exact trace / feedback ownership 校验 |
|
||||||
|
| `idx_diagnosis_run_status` | `status` | 状态筛选 |
|
||||||
|
| `idx_diagnosis_run_agent_flow` | `agent_flow` | 区分 Chat / AIOps |
|
||||||
|
|
||||||
|
## 关系
|
||||||
|
|
||||||
|
- `diagnosis_run.session_id` 逻辑关联 `chat_session.session_id`。
|
||||||
|
- `agent_step.run_id` 逻辑关联 `diagnosis_run.run_id`。
|
||||||
|
- `tool_invocation.run_id` 逻辑关联 `diagnosis_run.run_id`。
|
||||||
|
- 新的自动案例沉淀使用 `case_library.diagnosis_id = diagnosis_run.run_id`。
|
||||||
|
|
||||||
|
## 注意点
|
||||||
|
|
||||||
|
- `GET /api/diagnosis/{sessionId}/trace` 未带 `runId` 时只为兼容解析 latest run;新 demo 和新客户端应传 `runId`。
|
||||||
|
- latest run 排序使用 `created_at DESC, id DESC`,避免 feedback 或自评估更新 `updated_at` 后改变回放目标。
|
||||||
|
- 历史 `diagnosis_session` 会被迁移成兼容 run,但旧混合数据不能被还原成真实多轮边界。
|
||||||
@@ -30,7 +30,7 @@
|
|||||||
|
|
||||||
| Source | Alignment |
|
| Source | Alignment |
|
||||||
|---|---|
|
|---|---|
|
||||||
| Issue objective | Stage four in `mvp/issues/executor-structured-output-v2.md` requires Composer output final answer from Verifier-allowed material. Covered by proposal, design, specs, and tasks. |
|
| Issue objective | Stage four in `mvp/issues/design-notes/executor-structured-output-v2.md` requires Composer output final answer from Verifier-allowed material. Covered by proposal, design, specs, and tasks. |
|
||||||
| Proposal -> design | Proposal says Composer owns final expression; design defines input filtering, output parsing, fallback, and audit. |
|
| Proposal -> design | Proposal says Composer owns final expression; design defines input filtering, output parsing, fallback, and audit. |
|
||||||
| Design -> specs | Design decisions are reflected in `chat-composer-agent` requirements and modified `chat-verifier-agent` routing requirements. |
|
| Design -> specs | Design decisions are reflected in `chat-composer-agent` requirements and modified `chat-verifier-agent` routing requirements. |
|
||||||
| Specs -> tasks | Each required behavior has implementation and test tasks, including malformed fallback and no raw output leakage. |
|
| Specs -> tasks | Each required behavior has implementation and test tasks, including malformed fallback and no raw output leakage. |
|
||||||
|
|||||||
+1
-1
@@ -2,7 +2,7 @@
|
|||||||
|
|
||||||
## Discover Context
|
## Discover Context
|
||||||
|
|
||||||
- Source issue: `mvp/issues/ISS-007-verifier-evidence-summary-fidelity.md`.
|
- Source issue: `mvp/issues/archived/ISS-007-verifier-evidence-summary-fidelity.md`.
|
||||||
- Related existing specs: `chat-verifier-agent`, `evidence-trace-hardening`.
|
- Related existing specs: `chat-verifier-agent`, `evidence-trace-hardening`.
|
||||||
- Related devflow records: `executor-gatekeeper-hook`, `executor-verifier-claim-checks`, `executor-composer-final-answer`, `evidence-trace-hardening`.
|
- Related devflow records: `executor-gatekeeper-hook`, `executor-verifier-claim-checks`, `executor-composer-final-answer`, `evidence-trace-hardening`.
|
||||||
- Current repo instruction requested semantic code search and LSP confirmation before code changes; those tools are not exposed in this environment, so implementation will use `rg`, direct code reading, and focused tests as fallback evidence.
|
- Current repo instruction requested semantic code search and LSP confirmation before code changes; those tools are not exposed in this environment, so implementation will use `rg`, direct code reading, and focused tests as fallback evidence.
|
||||||
|
|||||||
@@ -0,0 +1 @@
|
|||||||
|
ready
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
committed
|
||||||
@@ -0,0 +1,117 @@
|
|||||||
|
# Design: Interview Demo Quality Audit
|
||||||
|
|
||||||
|
## Overview
|
||||||
|
|
||||||
|
The change adds auditability and demo readiness without changing Agent routing.
|
||||||
|
|
||||||
|
```text
|
||||||
|
ChatService prompt resources
|
||||||
|
-> PromptAuditService
|
||||||
|
-> verifier_evaluation.prompt_audit
|
||||||
|
-> Trace API
|
||||||
|
-> DiagnosisTraceEvaluator
|
||||||
|
-> baseline report
|
||||||
|
|
||||||
|
mvp/demo/scripts/run-interview-demo-check.ps1
|
||||||
|
-> health/readiness check
|
||||||
|
-> payment timeout chat
|
||||||
|
-> trace fetch
|
||||||
|
-> feedback
|
||||||
|
-> demo output bundle
|
||||||
|
```
|
||||||
|
|
||||||
|
## Prompt Audit
|
||||||
|
|
||||||
|
Add a compact `prompt_audit` object under:
|
||||||
|
|
||||||
|
```text
|
||||||
|
diagnosis_session.self_evaluation.verifier_evaluation.prompt_audit
|
||||||
|
```
|
||||||
|
|
||||||
|
Shape:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"version": "chat-prompts-v1",
|
||||||
|
"prompts": [
|
||||||
|
{
|
||||||
|
"name": "chat_planner",
|
||||||
|
"version": "chat-planner-v1",
|
||||||
|
"resource": "prompts/chat-planner-prompt.md"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Design choices:
|
||||||
|
|
||||||
|
- Use explicit local metadata, not full prompt hashes, because the interview goal is explainable version audit rather than cryptographic integrity.
|
||||||
|
- Keep this metadata in code or a small resource catalog near prompt loading.
|
||||||
|
- Persist prompt audit with every Chat verifier evaluation, including fallback/degraded paths.
|
||||||
|
- Do not include full prompt text in trace.
|
||||||
|
|
||||||
|
## Eval Expansion
|
||||||
|
|
||||||
|
Extend `DiagnosisEvalCase` with optional fields:
|
||||||
|
|
||||||
|
- `requirePromptAudit`
|
||||||
|
- `expectedPromptAuditVersion`
|
||||||
|
- `expectedPromptVersions`
|
||||||
|
- `requireGatekeeperRules`
|
||||||
|
|
||||||
|
Evaluator behavior:
|
||||||
|
|
||||||
|
- If `requirePromptAudit=true`, `verifier_evaluation.prompt_audit.version` must exist.
|
||||||
|
- If `expectedPromptAuditVersion` is set, it must match.
|
||||||
|
- If `expectedPromptVersions` is set, each listed prompt name/version pair must exist.
|
||||||
|
- If `requireGatekeeperRules=true`, `gatekeeper_result.rules` must be a non-empty list and each item must include `id`, `enabled`, and `default_severity`.
|
||||||
|
|
||||||
|
Add at least two fixture-backed cases:
|
||||||
|
|
||||||
|
- A positive audit closure case that requires prompt audit + Gatekeeper rules.
|
||||||
|
- A metadata-gap negative case represented as `LOW_CONFID`/safe final answer, used to prove the evaluator catches missing audit metadata when configured.
|
||||||
|
|
||||||
|
The baseline must remain fully passing after fixtures are updated.
|
||||||
|
|
||||||
|
## Demo Stabilization
|
||||||
|
|
||||||
|
Add a PowerShell script:
|
||||||
|
|
||||||
|
```text
|
||||||
|
mvp/demo/scripts/run-interview-demo-check.ps1
|
||||||
|
```
|
||||||
|
|
||||||
|
Responsibilities:
|
||||||
|
|
||||||
|
- Accept base URL and session id parameters.
|
||||||
|
- Check that the service is reachable.
|
||||||
|
- Run the existing payment-timeout chat request.
|
||||||
|
- Fetch trace for the same session id.
|
||||||
|
- Submit useful feedback.
|
||||||
|
- Write outputs under `mvp/demo/output/`.
|
||||||
|
- Emit a concise summary with session id, verdict, Gatekeeper rule version, prompt audit version, and output paths.
|
||||||
|
|
||||||
|
The script should fail fast with actionable messages when the service is unavailable.
|
||||||
|
|
||||||
|
## Documentation
|
||||||
|
|
||||||
|
Add/update:
|
||||||
|
|
||||||
|
- `mvp/demo/README.md`: mention the preflight script.
|
||||||
|
- `mvp/demo/ten-minute-interview-demo.md`: use the preflight script as the recommended path.
|
||||||
|
- `mvp/demo/interview-q-and-a.md`: concise interview answers for Agent engineering tradeoffs.
|
||||||
|
- `mvp/architecture/harness-quality-gates.md`: record prompt audit as part of the quality gate.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
Required:
|
||||||
|
|
||||||
|
- Targeted unit/eval tests for prompt audit persistence and evaluator checks.
|
||||||
|
- Regenerated baseline JSON/Markdown reports.
|
||||||
|
- OpenSpec validation.
|
||||||
|
|
||||||
|
E2E:
|
||||||
|
|
||||||
|
- If local dependencies are available, run Spring Boot with `mvp-demo` profile and execute the new preflight script.
|
||||||
|
- If unavailable, record the reason and rely on deterministic unit/eval evidence.
|
||||||
|
|
||||||
@@ -0,0 +1,58 @@
|
|||||||
|
# Change: Interview Demo Quality Audit
|
||||||
|
|
||||||
|
## Problem
|
||||||
|
|
||||||
|
SuperBizAgent MVP is now strong enough to demonstrate traceable Agent engineering, but the interview path still has three gaps:
|
||||||
|
|
||||||
|
- The live demo has a run script, but no preflight command that checks service readiness and produces a concise interview evidence bundle.
|
||||||
|
- The diagnosis eval baseline covers the V2 evidence pipeline, but it does not yet assert prompt-version audit data and has limited coverage for audit metadata gaps.
|
||||||
|
- Gatekeeper already exposes `rule_set_version`, but prompt versions are not persisted with the verifier evaluation, making prompt changes harder to explain, compare, and roll back in an interview.
|
||||||
|
|
||||||
|
This change stabilizes the MVP as an interview artifact rather than adding a new diagnosis architecture.
|
||||||
|
|
||||||
|
## Proposed Solution
|
||||||
|
|
||||||
|
Implement a small internal quality/audit increment:
|
||||||
|
|
||||||
|
1. Add prompt version audit metadata to Chat verifier evaluation.
|
||||||
|
2. Extend deterministic diagnosis eval cases/fixtures to assert prompt audit metadata and audit-metadata failures.
|
||||||
|
3. Add an interview demo preflight script and documentation that can be run before or during a demo to verify service readiness, execute the payment timeout path, fetch trace, and record key audit fields.
|
||||||
|
4. Add/update MVP documentation for interview Q&A and the new audit/preflight workflow.
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
In scope:
|
||||||
|
|
||||||
|
- Internal `diagnosis_session.self_evaluation.verifier_evaluation` audit JSON.
|
||||||
|
- Diagnosis eval case schema, evaluator checks, fixtures, and baseline reports.
|
||||||
|
- MVP demo scripts/docs.
|
||||||
|
- Architecture/demo documentation for prompt and Gatekeeper version audit.
|
||||||
|
|
||||||
|
Out of scope:
|
||||||
|
|
||||||
|
- Public HTTP API changes.
|
||||||
|
- Database schema changes.
|
||||||
|
- New Agent roles, MCP tool server migration, process isolation, or AIOps LLM Verifier.
|
||||||
|
- Replacing existing `Planner -> Executor -> Gatekeeper -> Verifier -> Composer` orchestration.
|
||||||
|
- Guaranteeing live LLM `PASS` for every demo run. Live demo compatibility and deterministic fixture regression are separate acceptance paths.
|
||||||
|
|
||||||
|
## Context Constraints From devflow
|
||||||
|
|
||||||
|
- Evidence Tools produce incident facts and must be recorded in `tool_invocation`.
|
||||||
|
- Chat quality gates are layered: Gatekeeper verifies evidence references, Verifier judges derivability, Composer controls expression.
|
||||||
|
- `diagnosis eval` is deterministic and fixture-backed; no LLM-as-judge.
|
||||||
|
- Demo assets should be runnable, but interview safety should not depend solely on live LLM behavior.
|
||||||
|
- Gatekeeper rule metadata is metadata-only; dynamic rule execution is out of scope.
|
||||||
|
|
||||||
|
## Interface Impact
|
||||||
|
|
||||||
|
Level: L2 internal contract change.
|
||||||
|
|
||||||
|
Reason: `verifier_evaluation` gains a compact `prompt_audit` object. Existing public API shape remains the same, and the value is exposed only through already-existing trace/self-evaluation JSON.
|
||||||
|
|
||||||
|
## Risks
|
||||||
|
|
||||||
|
- Baseline report churn is expected when adding cases; JSON and Markdown reports must be regenerated together.
|
||||||
|
- Prompt audit must be deterministic and stable enough for eval fixtures; avoid hashing full prompt text with environment-specific content.
|
||||||
|
- Demo preflight must not hardcode secrets and must tolerate local service unavailability with clear failure messages.
|
||||||
|
|
||||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user