Compare commits
16
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
30d3296043 | ||
|
|
3578709896 | ||
|
|
f9df94377b | ||
|
|
78c1477198 | ||
|
|
d928a1968a | ||
|
|
027aed1eeb | ||
|
|
26d5529280 | ||
|
|
6fdbd34bab | ||
|
|
52bf0302c6 | ||
|
|
841437fa06 | ||
|
|
9c9a0024d4 | ||
|
|
a6c2d4459c | ||
|
|
da45fa3fb0 | ||
|
|
db0f229285 | ||
|
|
a77c947cd4 | ||
|
|
9a84b3de34 |
@@ -72,9 +72,10 @@
|
||||
|
||||
### SessionContext
|
||||
- 定义:会话上下文数据类,存储在 Redis 中的会话数据
|
||||
- 包含字段:sessionId、userId、businessId、traceId、status、toolCalls、TTL
|
||||
- 包含字段:sessionId、userId、businessId、traceId、status、toolCalls、messageHistory、TTL
|
||||
- 序列化方式:JSON(GenericJackson2JsonRedisSerializer)
|
||||
- 使用场景:多轮对话上下文管理、工具调用历史追踪
|
||||
- 边界:messageHistory 是热路径对话历史缓存,用于下一轮 prompt 上下文;长期审计的问题和答案应落到 Diagnosis Run,而不是依赖 Redis TTL 内的上下文正文。
|
||||
|
||||
### ToolCall
|
||||
- 定义:工具调用记录数据类,追踪 Agent 使用的工具及其结果
|
||||
@@ -87,6 +88,21 @@
|
||||
- 核心方法:createSession、getSession、updateSession、deleteSession、refreshSession、addToolCall
|
||||
- 使用场景:分布式会话管理、Agent 状态维护
|
||||
|
||||
### Chat Session
|
||||
- 定义:一次多轮对话上下文,由 `sessionId` 唯一标识。
|
||||
- 使用场景:保存用户连续对话的上下文窗口、会话状态和最近活跃时间。
|
||||
- 边界:Chat Session 不代表一次诊断执行;同一个 Chat Session 可以包含多次 Diagnosis Run。
|
||||
|
||||
### Diagnosis Run
|
||||
- 定义:一次独立诊断执行,由 `runId` 唯一标识,属于一个 Chat Session。
|
||||
- 使用场景:保存某一轮诊断的 query、answer、status、耗时、token、反馈和自评估结果。
|
||||
- 边界:Diagnosis Run 是 Trace、Feedback 和 Evidence score 的绑定对象;多轮对话中的每次 `/api/chat` 或 `/api/ai_ops` 执行都应创建新的 Diagnosis Run。
|
||||
|
||||
### Diagnosis Trace
|
||||
- 定义:一次 Diagnosis Run 的可回放执行轨迹,由 run 主记录、AgentStep 和 ToolInvocation 聚合形成。
|
||||
- 使用场景:Trace API、Trace UI、Verifier 审计、评测 fixture 和人工排查。
|
||||
- 边界:Diagnosis Trace 是聚合视图,不要求单独的 trace 主表;当前 trace 明细由 `agent_step` 和 `tool_invocation` 表承载。
|
||||
|
||||
### Flyway
|
||||
- 定义:数据库版本迁移工具,管理 SQL 脚本的版本化执行
|
||||
- 配置:spring.flyway.enabled=true, baseline-on-migrate=true
|
||||
|
||||
+31
-28
@@ -2,31 +2,34 @@
|
||||
|
||||
## 项目
|
||||
|
||||
| 日期 | slug | 领域 | 关键词 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| 2026-07-05 | diagnosis-playbook-skills | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented |
|
||||
| 2026-07-07 | executor-evidence-output-contract | Chat质量门禁/证据归因 | Executor structured output, evidence bindings, Verifier structured claims, LOW_CONFID, hallucination | openspec/changes/archive/2026-07-07-executor-evidence-output-contract | archived |
|
||||
| 2026-07-07 | executor-v2-output-contract | Chat质量门禁/证据归因 | executor_evidence_v2, user_facing_answer removal, diagnosis_summary removal, structured renderer | openspec/changes/archive/2026-07-07-executor-v2-output-contract | archived |
|
||||
| 2026-07-07 | executor-gatekeeper-hook | Chat质量门禁/证据归因 | Gatekeeper, verifier payload, source_invocation_ids, tool_name match, self_evaluation | openspec/changes/archive/2026-07-07-executor-gatekeeper-hook | archived |
|
||||
| 2026-07-07 | executor-verifier-claim-checks | Chat质量门禁/证据归因 | Verifier claim_checks, facts_checked compatibility, effective verdict guardrail, malformed output downgrade | openspec/changes/archive/2026-07-07-executor-verifier-claim-checks | archived |
|
||||
| 2026-07-08 | executor-composer-final-answer | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
|
||||
| 2026-07-08 | verifier-evidence-reference-fidelity | Chat质量门禁/证据归因 | evidence_refs, raw_path, Gatekeeper severity, verifier evidence excerpt, HikariCP mock, no_evidence | openspec/changes/archive/2026-07-08-verifier-evidence-reference-fidelity | archived |
|
||||
| 2026-07-06 | rag-eval-pipeline-closure | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
|
||||
| 2026-07-06 | modular-rag-pipeline | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
|
||||
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
|
||||
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
|
||||
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
|
||||
| 2026-07-04 | diagnosis-eval-harness | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
|
||||
| 2026-07-04 | evidence-trace-hardening | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
|
||||
| 2026-07-03 | mvp-demo-trace-acceptance | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
|
||||
| 2026-07-04 | aiops-traceable-diagnosis-entry | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
|
||||
| 2026-07-04 | aiops-alert-scope-control | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
|
||||
| 2026-05-29 | chatmodel-abstraction | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | archived |
|
||||
| 2026-06-23 | phase1-infrastructure | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | archived |
|
||||
| 2026-06-24 | lookup-knowledge-integration | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | archived |
|
||||
| 2026-06-25 | doc-management-ui | 前端开发/文档管理 | 文档管理页面, CRUD, 状态监控, 纯静态页面, API集成 | archived |
|
||||
| 2026-06-26 | session-storage | 会话存储/可观测 | diagnosis_session, agent_step, tool_invocation, token追踪, 多Agent路由 | openspec/changes/session-storage | archived |
|
||||
| 2026-06-29 | confidence-feedback | 质量评估/反馈机制 | evidence_score, selfEvaluation, feedback, useful, not_useful, case_library, BAD_CASE, tool_invocation规则引擎, 反馈按钮, sessionId回传 | openspec/changes/confidence-feedback | archived |
|
||||
| 2026-06-30 | session-dedup-knowledge-map | 去重/知识图谱 | RetrievedDocTracker, KnowledgeDomainService, knowledge_domain, covers, whenToRetrieve, Planner注入, ISS-001 | openspec/changes/archive/2026-06-30-session-dedup-knowledge-map | archived |
|
||||
| 2026-07-01 | executor-action-memory-relevance | 检索质量/行动记忆 | relevanceLevel, completenessHint, Min-Max归一化, RetrievedDocTracker域级记录, Executor检索约束, ISS-002 | openspec/changes/archive/2026-07-01-executor-action-memory-relevance | archived |
|
||||
| 2026-07-02 | chat-verifier-agent | Chat质量门禁/可追溯验证 | Verifier, groundedness_score, facts_checked, evidence_refs, tool_trace_summary, self_evaluation | openspec/changes/archive/2026-07-03-chat-verifier-agent | archived |
|
||||
| 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|
||||
|---|---|---|---|---|---|---|
|
||||
| 2026-07-10 | session-run-trace-isolation | 拆分会话态和运行态,引入 runId 隔离 Trace、Feedback、AIOps 和 demo 链路。 | Trace/session/run isolation | chat_session, diagnosis_run, runId, trace exact run, feedback fallback, AIOps SSE metadata, baseline drift | openspec/changes/archive/2026-07-10-session-run-trace-isolation | archived |
|
||||
| 2026-07-09 | interview-demo-quality-audit | 增加面试演示前置质量审计,覆盖 prompt、Gatekeeper 和评测基线。 | Agent eval/demo/Prompt audit | interview demo preflight, prompt_audit, gatekeeper rules, diagnosis baseline, 12 fixtures | openspec/changes/archive/2026-07-09-interview-demo-quality-audit | archived |
|
||||
| 2026-07-08 | executor-composer-final-answer | 引入 Composer 生成最终回答,只使用 Verifier 允许的结论材料。 | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
|
||||
| 2026-07-08 | diagnosis-eval-demo-gatekeeper-closure | 收敛诊断评测、稳定 demo 场景和 Gatekeeper 审计元数据。 | Agent eval/demo/Gatekeeper | diagnosis eval matrix, stable demo scenarios, Gatekeeper rule set version, audit metadata | openspec/changes/archive/2026-07-08-diagnosis-eval-demo-gatekeeper-closure | archived |
|
||||
| 2026-07-08 | verifier-evidence-reference-fidelity | 强化 Verifier 对 evidence_refs、raw_path 和 no_evidence 的保真校验。 | Chat质量门禁/证据归因 | evidence_refs, raw_path, Gatekeeper severity, verifier evidence excerpt, HikariCP mock, no_evidence | openspec/changes/archive/2026-07-08-verifier-evidence-reference-fidelity | archived |
|
||||
| 2026-07-07 | executor-evidence-output-contract | 设计 Executor 结构化证据输出,解决证据归因幻觉和 LOW_CONFID 问题。 | Chat质量门禁/证据归因 | Executor structured output, evidence bindings, Verifier structured claims, LOW_CONFID, hallucination | openspec/changes/archive/2026-07-07-executor-evidence-output-contract | archived |
|
||||
| 2026-07-07 | executor-v2-output-contract | 将 Executor 输出升级为 V2 契约,移除面向用户的最终回答字段。 | Chat质量门禁/证据归因 | executor_evidence_v2, user_facing_answer removal, diagnosis_summary removal, structured renderer | openspec/changes/archive/2026-07-07-executor-v2-output-contract | archived |
|
||||
| 2026-07-07 | executor-gatekeeper-hook | 在 Executor 与 Verifier 之间接入 Gatekeeper,校验证据绑定来源。 | Chat质量门禁/证据归因 | Gatekeeper, verifier payload, source_invocation_ids, tool_name match, self_evaluation | openspec/changes/archive/2026-07-07-executor-gatekeeper-hook | archived |
|
||||
| 2026-07-07 | executor-verifier-claim-checks | 增加 Verifier claim_checks 和事实校验兼容逻辑。 | Chat质量门禁/证据归因 | Verifier claim_checks, facts_checked compatibility, effective verdict guardrail, malformed output downgrade | openspec/changes/archive/2026-07-07-executor-verifier-claim-checks | archived |
|
||||
| 2026-07-06 | rag-eval-pipeline-closure | 建立 RAG 评测闭环,加入 fixture、快照和 baseline diff。 | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
|
||||
| 2026-07-06 | modular-rag-pipeline | 将 lookup_knowledge 改造成模块化 RAG 管线,补齐证据块和检索追踪。 | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
|
||||
| 2026-07-05 | diagnosis-playbook-skills | 增加诊断 Playbook Skill,沉淀支付超时、MySQL 池、Redis 超时等套路。 | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented |
|
||||
| 2026-07-05 | mvp-demo-interview-runbook | 准备可复现的 MVP 面试演示包、运行手册和 Trace 检查清单。 | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
|
||||
| 2026-07-05 | diagnosis-eval-baseline-diff | 增加诊断评测 baseline diff,用于判断回归和证据覆盖变化。 | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
|
||||
| 2026-07-04 | expand-diagnosis-eval-fixtures | 扩充诊断评测 fixture,覆盖 Redis、慢响应和 JVM 内存风险。 | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
|
||||
| 2026-07-04 | diagnosis-eval-harness | 建立固定诊断评测 Harness,输出 trace、证据覆盖和 verdict 分布。 | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
|
||||
| 2026-07-04 | evidence-trace-hardening | 强化工具调用证据链、降级契约和离线验证能力。 | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
|
||||
| 2026-07-04 | aiops-traceable-diagnosis-entry | 增加可追踪的 AIOps 告警诊断入口,打通 sessionId 和 Trace API。 | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
|
||||
| 2026-07-04 | aiops-alert-scope-control | 收敛 AIOps 告警诊断范围,区分 payload 定向和自动发现模式。 | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
|
||||
| 2026-07-03 | mvp-demo-trace-acceptance | 增加 MVP demo 的 Trace 验收,覆盖会话、步骤、工具和反馈链路。 | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
|
||||
| 2026-07-02 | chat-verifier-agent | 增加 Chat Verifier Agent,用 groundedness 和 evidence_refs 校验回答。 | Chat质量门禁/可追溯验证 | Verifier, groundedness_score, facts_checked, evidence_refs, tool_trace_summary, self_evaluation | openspec/changes/archive/2026-07-03-chat-verifier-agent | archived |
|
||||
| 2026-07-01 | executor-action-memory-relevance | 增加行动记忆和相关性信号,约束 Executor 重复检索。 | 检索质量/行动记忆 | relevanceLevel, completenessHint, Min-Max归一化, RetrievedDocTracker域级记录, Executor检索约束, ISS-002 | openspec/changes/archive/2026-07-01-executor-action-memory-relevance | archived |
|
||||
| 2026-06-30 | session-dedup-knowledge-map | 引入会话级去重和知识域地图,减少重复召回。 | 去重/知识图谱 | RetrievedDocTracker, KnowledgeDomainService, knowledge_domain, covers, whenToRetrieve, Planner注入, ISS-001 | openspec/changes/archive/2026-06-30-session-dedup-knowledge-map | archived |
|
||||
| 2026-06-29 | confidence-feedback | 建立质量评估和用户反馈机制,并把有用反馈沉淀为案例。 | 质量评估/反馈机制 | evidence_score, selfEvaluation, feedback, useful, not_useful, case_library, BAD_CASE, tool_invocation规则引擎, 反馈按钮, sessionId回传 | openspec/changes/confidence-feedback | archived |
|
||||
| 2026-06-26 | session-storage | 建立通用会话存储,记录 session、agent step 和 tool invocation。 | 会话存储/可观测 | diagnosis_session, agent_step, tool_invocation, token追踪, 多Agent路由 | openspec/changes/session-storage | archived |
|
||||
| 2026-06-25 | doc-management-ui | 实现文档管理页面,支持文档 CRUD、状态监控和 API 集成。 | 前端开发/文档管理 | 文档管理页面, CRUD, 状态监控, 纯静态页面, API集成 | - | archived |
|
||||
| 2026-06-24 | lookup-knowledge-integration | 接入知识库检索,支持 L0 精确匹配和 L1 语义检索。 | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | - | archived |
|
||||
| 2026-06-23 | phase1-infrastructure | 搭建第一阶段基础设施,包括 MySQL、Redis、Milvus、Flyway 和 JPA。 | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | - | archived |
|
||||
| 2026-05-29 | chatmodel-abstraction | 抽象 ChatModel 和 EmbeddingModel,支持多模型路由。 | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | - | archived |
|
||||
|
||||
@@ -48,7 +48,7 @@
|
||||
|
||||
## 遗留问题
|
||||
|
||||
ISS-002:Executor 无约束重复调用 `lookup_knowledge`(单会话 20+ 次),knowledge map 和检索约束只注入了 Planner 未注入 Executor。详见 `mvp/issues/ISS-002-executor-unconstrained-lookup.md`。
|
||||
ISS-002:Executor 无约束重复调用 `lookup_knowledge`(单会话 20+ 次),knowledge map 和检索约束只注入了 Planner 未注入 Executor。详见 `mvp/issues/archived/ISS-002-executor-unconstrained-lookup.md`。
|
||||
|
||||
## 已知限制
|
||||
|
||||
|
||||
@@ -9,8 +9,8 @@
|
||||
## Context
|
||||
|
||||
- `devflow/index.md` was checked. Relevant history includes `session-storage`, `confidence-feedback`, `executor-action-memory-relevance`, and `chat-verifier-agent`.
|
||||
- `mvp/notes/agent-engineering-decisions.md` already recommends the next phase as "可复现 MVP Demo", including `mvp-demo` profile, fixed diagnosis case, one-click request, and `GET /api/diagnosis/{sessionId}/trace`.
|
||||
- `mvp/issues/ISS-003-mvp-design-implementation-review.md` identifies test stability, session traceability, verifier evidence chain, upload path, and SupervisorAgent consistency as recent MVP concerns. Security cleanup is intentionally deferred by user decision.
|
||||
- `mvp/archive/2026-07-09-doc-cleanup/notes/agent-engineering-decisions.md` already recommends the next phase as "可复现 MVP Demo", including `mvp-demo` profile, fixed diagnosis case, one-click request, and `GET /api/diagnosis/{sessionId}/trace`.
|
||||
- `mvp/issues/active/ISS-003-mvp-design-implementation-review.md` identifies test stability, session traceability, verifier evidence chain, upload path, and SupervisorAgent consistency as recent MVP concerns. Security cleanup is intentionally deferred by user decision.
|
||||
|
||||
## Question Pool
|
||||
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
## Draft Acceptance
|
||||
|
||||
- [x] Issue exists: `mvp/issues/executor-evidence-attribution-hallucination.md`.
|
||||
- [x] Issue exists: `mvp/issues/active/executor-evidence-attribution-hallucination.md`.
|
||||
- [x] OpenSpec change artifacts exist.
|
||||
- [x] devflow tracking files exist.
|
||||
- [x] OpenSpec validation passes.
|
||||
|
||||
@@ -5,7 +5,7 @@
|
||||
- `devflow/projects/2026-07-07-executor-evidence-output-contract`: V1 evidence-attribution contract kept `user_facing_answer`.
|
||||
- `devflow/projects/2026-07-02-chat-verifier-agent`: Verifier consumes explicit inputs and should not see intermediate reasoning.
|
||||
- `devflow/projects/2026-07-04-evidence-trace-hardening`: evidence summaries and tool invocation references are the evidence foundation.
|
||||
- `mvp/issues/executor-structured-output-v2.md`: staged implementation design; stage one is Executor V2 output contract.
|
||||
- `mvp/issues/design-notes/executor-structured-output-v2.md`: staged implementation design; stage one is Executor V2 output contract.
|
||||
|
||||
## Code Evidence
|
||||
|
||||
|
||||
@@ -0,0 +1,48 @@
|
||||
# diagnosis-eval-demo-gatekeeper-closure Acceptance
|
||||
|
||||
## Static / Structure Verification
|
||||
|
||||
- `cmd /c openspec validate diagnosis-eval-demo-gatekeeper-closure --strict`
|
||||
- Result: passed.
|
||||
- `cmd /c openspec validate --specs`
|
||||
- Result: passed, 10 specs passed.
|
||||
|
||||
## Script Verification
|
||||
|
||||
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,VerifierInputHookTest" test`
|
||||
- Result: 36 tests, 0 failures, 0 errors.
|
||||
- `mvn "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest,ToolInvocationRecorderTest,QueryLogsToolsTest" test`
|
||||
- Result: 61 tests, 0 failures, 0 errors.
|
||||
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`
|
||||
- Result after E2E startup fix: 23 tests, 0 failures, 0 errors.
|
||||
|
||||
## Live E2E Verification
|
||||
|
||||
- Start command: `mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"`.
|
||||
- Demo command: `powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1`.
|
||||
- Result: chat, trace, and feedback requests completed successfully.
|
||||
- Output files:
|
||||
- `mvp/demo/output/chat-response.json`
|
||||
- `mvp/demo/output/trace-response.json`
|
||||
- `mvp/demo/output/feedback-response.json`
|
||||
- Trace observations:
|
||||
- `hasVerifierEvaluation=true`
|
||||
- `gatekeeper_result.rule_set_version=gatekeeper-rules-v1`
|
||||
|
||||
## Fixed During Verification
|
||||
|
||||
- E2E startup initially failed because Spring could not instantiate `ExecutorGatekeeperService`.
|
||||
- Root cause: two public constructors and no explicit `@Autowired` constructor.
|
||||
- Fix: annotate the production constructor with `@Autowired`.
|
||||
|
||||
## Residual Risk
|
||||
|
||||
- The live payment-timeout path can still produce `LOW_CONFID` because model-generated evidence bindings may omit some explicit `source_invocation_id` values.
|
||||
- This is not a blocker for this change because deterministic matrix behavior is covered by saved fixtures and baseline evaluation.
|
||||
- Existing Maven warnings remain: duplicate `spring-boot-starter-test` declaration and Lombok `@Builder` default warnings.
|
||||
|
||||
## Archive Status
|
||||
|
||||
- Devflow archive artifacts created.
|
||||
- OpenSpec change archived to `openspec/changes/archive/2026-07-08-diagnosis-eval-demo-gatekeeper-closure`.
|
||||
- Main specs synced by `cmd /c openspec archive diagnosis-eval-demo-gatekeeper-closure --yes`.
|
||||
@@ -0,0 +1,33 @@
|
||||
# diagnosis-eval-demo-gatekeeper-closure Brief
|
||||
|
||||
## Background
|
||||
|
||||
The Chat evidence pipeline already had Executor V2 structured output, deterministic Gatekeeper validation, Verifier claim checks, and Composer final rendering. The missing piece was an interview-ready acceptance story that made the anti-hallucination behavior easy to demonstrate and regress.
|
||||
|
||||
## Goal
|
||||
|
||||
Close the next three interview-readiness gaps together:
|
||||
|
||||
- diagnosis eval fixture matrix
|
||||
- stable demo data set
|
||||
- Gatekeeper rule configuration and audit version
|
||||
|
||||
## Scope
|
||||
|
||||
- Expand `mvp/eval` with matrix-oriented cases, fixtures, and baseline reports.
|
||||
- Add stable demo request payloads and scenario documentation.
|
||||
- Add a lightweight local Gatekeeper rule catalog with `rule_set_version` and rule metadata in `gatekeeper_result`.
|
||||
- Update architecture, demo, and eval docs to describe the current implementation.
|
||||
|
||||
## Non-goals
|
||||
|
||||
- No new public HTTP endpoint.
|
||||
- No new database table.
|
||||
- No Planner `scope_contract`.
|
||||
- No Gatekeeper retry loop.
|
||||
- No remote or dynamic rule execution engine.
|
||||
|
||||
## OpenSpec
|
||||
|
||||
- Change: `openspec/changes/diagnosis-eval-demo-gatekeeper-closure`
|
||||
- Interface impact: L2 internal contract change.
|
||||
@@ -0,0 +1,144 @@
|
||||
# diagnosis-eval-demo-gatekeeper-closure Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: implement the next three interview-readiness items together: diagnosis eval fixture matrix, stable demo data set, and Gatekeeper rule configuration/audit version.
|
||||
- Slug: `diagnosis-eval-demo-gatekeeper-closure`
|
||||
- Devflow scale: `standard`
|
||||
- Interface impact: expected L2 internal contract change because `gatekeeper_result` audit JSON will gain rule metadata/version fields.
|
||||
|
||||
## Context
|
||||
|
||||
- `devflow/index.md` used: related entries found for diagnosis eval harness, fixture expansion, MVP demo runbook, Gatekeeper hook, and verifier evidence reference fidelity.
|
||||
- Relevant glossary:
|
||||
- Evidence Tools produce incident facts and must be recorded in `tool_invocation`.
|
||||
- Verifier should not use skills/runbooks as incident evidence.
|
||||
- `tool_invocation.retrieval_details` is the structured evidence/audit home for tool-specific details.
|
||||
- Historical constraints that must enter OpenSpec:
|
||||
- Diagnosis eval is offline and deterministic; no LLM-as-judge.
|
||||
- Demo assets should be runnable, but fixed regression should use saved fixtures.
|
||||
- Gatekeeper remains in the Verifier hook path.
|
||||
- No new database table for Gatekeeper audit; use `self_evaluation.verifier_evaluation.gatekeeper_result`.
|
||||
- `$.no_evidence` is a query no-hit signal, not proof that a problem is impossible.
|
||||
|
||||
## Question Pool
|
||||
|
||||
| ID | Dimension | Mode | Question | Status |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | Terminology | evidence-driven | What names should this change use for the matrix, demo set, and Gatekeeper rule metadata? | Resolved |
|
||||
| Q2 | Boundary | evidence-driven | Should this change alter public APIs, database schema, Planner output, or retry behavior? | Resolved |
|
||||
| Q3 | Acceptance | evidence-driven | Which existing tests and baseline assets define the current acceptance style? | Resolved |
|
||||
| Q4 | Technical | evidence-driven | Where should Gatekeeper rule metadata live with minimal implementation risk? | Pending code research |
|
||||
| Q5 | Scope | user-interview | Should the stable demo set be documentation/payloads only, or should it include live E2E scripts for all scenarios? | Confirmed |
|
||||
|
||||
## Evidence-driven Conclusions
|
||||
|
||||
- Q1 conclusion: use `diagnosis eval matrix`, `stable demo scenarios`, and `Gatekeeper rule set version` as terms.
|
||||
- Q2 conclusion: keep this as an internal contract change. Do not add public endpoints, tables, Planner `scope_contract`, or Gatekeeper retry.
|
||||
- Q3 conclusion: existing `DiagnosisTraceEvaluatorTest`, `ExecutorGatekeeperServiceTest`, `VerifierInputHookTest`, `ToolInvocationRecorderTest`, and `mvp/eval/reports` define the current acceptance style.
|
||||
- Q4 conclusion: Gatekeeper metadata should live behind a small rule catalog loaded by `ExecutorGatekeeperService`; the audit output should include a rule set version and enabled rule metadata summary, without adding tables or remote registry.
|
||||
|
||||
## User-interview Confirmations
|
||||
|
||||
- Q5 confirmed by resumed objective: complete items 1/2/3 with sm-flow, archive, submit, and run end-to-end if necessary.
|
||||
- Implementation interpretation: stable demo scenarios will be fixed request payloads and runbook docs plus deterministic fixture-backed eval. Live E2E remains necessary only for at least one main path or where unit/fixture evidence is insufficient.
|
||||
|
||||
## OpenSpec Backfill
|
||||
|
||||
- Created Draft proposal at `openspec/changes/diagnosis-eval-demo-gatekeeper-closure/proposal.md`.
|
||||
- Context constraints from historical devflow entries were written into the proposal.
|
||||
- Scope confirmation and Gatekeeper catalog placement were written into the proposal/design.
|
||||
|
||||
## Current Checkpoint
|
||||
|
||||
- Discover completed.
|
||||
- No implementation files changed yet.
|
||||
|
||||
## Specify / Alignment
|
||||
|
||||
### Cross-artifact Alignment
|
||||
|
||||
| Check | Status | Notes |
|
||||
|---|---|---|
|
||||
| brief/proposal goals -> proposal | Aligned | Proposal covers eval matrix, stable demo scenarios, and Gatekeeper rule catalog/audit version. |
|
||||
| proposal scope/constraints -> design | Aligned | Design records offline deterministic eval, fixture-backed demo distinction, local rule catalog, and no new table/API. |
|
||||
| design decisions -> specs/tasks | Aligned | Specs cover eval matrix, rule set version validation, demo scenarios, and Gatekeeper rule metadata; tasks cover matching implementation slices. |
|
||||
| specs observable behavior -> tasks | Aligned | Each requirement has an executable task and acceptance check. |
|
||||
|
||||
### Interface Impact
|
||||
|
||||
- Level: L2 internal contract change.
|
||||
- Reason: `gatekeeper_result` internal audit JSON gains `rule_set_version` and rule metadata summary. Eval case/result fields may gain optional rule set checks. No public HTTP API, database schema, or external DTO contract changes.
|
||||
|
||||
## Audit
|
||||
|
||||
Input -> processing -> output chain:
|
||||
|
||||
```text
|
||||
mvp/demo request docs + mvp/eval fixtures
|
||||
-> DiagnosisTraceEvaluator
|
||||
-> baseline reports
|
||||
-> interview/demo evidence
|
||||
|
||||
Gatekeeper rule catalog
|
||||
-> ExecutorGatekeeperService
|
||||
-> VerifierInputHook / ChatService persisted self_evaluation
|
||||
-> Trace and eval audit
|
||||
```
|
||||
|
||||
Architecture risk assessment:
|
||||
|
||||
1. The change is intentionally internal and should not add new public consumers.
|
||||
2. Gatekeeper catalog must stay metadata-only; dynamic rule execution would be a different, riskier architecture.
|
||||
3. Fixture-backed demo scenarios should be documented as deterministic regression artifacts, not live LLM guarantees.
|
||||
4. Baseline report churn is expected and must be committed with case/fixture changes.
|
||||
5. No devflow/OpenSpec conflict found.
|
||||
|
||||
## Commit Gate
|
||||
|
||||
- `cmd /c openspec validate diagnosis-eval-demo-gatekeeper-closure --strict`: passed.
|
||||
- `cmd /c openspec validate --specs`: passed, 10 specs passed.
|
||||
- File completeness:
|
||||
- proposal.md: present.
|
||||
- design.md: present.
|
||||
- specs: present for `diagnosis-eval-harness`, `mvp-demo-trace-acceptance`, `chat-verifier-agent`.
|
||||
- tasks.md: present.
|
||||
- Consistency:
|
||||
- Proposal concepts have corresponding design sections.
|
||||
- Design decisions are reflected in specs/tasks.
|
||||
- Task acceptance checks are verifiable.
|
||||
|
||||
## Current Checkpoint
|
||||
|
||||
- Commit completed.
|
||||
- `.committed` marker created.
|
||||
|
||||
## Apply Verification
|
||||
|
||||
- Focused verification passed:
|
||||
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,VerifierInputHookTest" test`
|
||||
- Result: 36 tests, 0 failures, 0 errors.
|
||||
- Broader relevant regression passed:
|
||||
- `mvn "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest,ToolInvocationRecorderTest,QueryLogsToolsTest" test`
|
||||
- Result: 61 tests, 0 failures, 0 errors.
|
||||
- E2E startup repro found a Spring bean construction issue:
|
||||
- Command: `mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"`
|
||||
- Failure: `ExecutorGatekeeperService` had two public constructors and no annotated constructor, so Spring attempted a no-arg constructor and failed with `No default constructor found`.
|
||||
- Classification: code deviation from OpenSpec implementation intent, not a spec gap.
|
||||
- Fix: annotate the production constructor with `@Autowired`.
|
||||
- Post-fix focused regression passed:
|
||||
- `mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`
|
||||
- Result: 23 tests, 0 failures, 0 errors.
|
||||
- Live E2E passed for demo compatibility:
|
||||
- Start: `mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"`
|
||||
- Run: `powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1`
|
||||
- Result: `/api/chat`, `/api/diagnosis/{sessionId}/trace`, and `/api/feedback` completed successfully.
|
||||
- Trace summary included `hasVerifierEvaluation=true`.
|
||||
- Persisted Gatekeeper audit included `rule_set_version=gatekeeper-rules-v1`.
|
||||
- Residual quality note: the live payment-timeout response remained `LOW_CONFID` because some model-produced evidence bindings still lacked explicit `source_invocation_id`; deterministic PASS/LOW_CONFID/REJECT claims are covered by fixture-backed eval.
|
||||
|
||||
## Archive Readiness
|
||||
|
||||
- OpenSpec tasks 1-4 completed.
|
||||
- Verification is recorded in devflow acceptance artifacts.
|
||||
- Remaining known risk: live LLM output is not deterministic and may still produce LOW_CONFID on the payment-timeout path; this is intentionally documented as demo compatibility, not a fixed PASS guarantee.
|
||||
@@ -0,0 +1,22 @@
|
||||
# diagnosis-eval-demo-gatekeeper-closure Evidence
|
||||
|
||||
## Code And Artifact Evidence
|
||||
|
||||
- Gatekeeper rule metadata lives in `src/main/resources/gatekeeper/gatekeeper-rules.json`.
|
||||
- `ExecutorGatekeeperService` loads the local catalog, uses configured threshold parameters, and emits `rule_set_version` plus enabled rule metadata.
|
||||
- `VerifierInputHook` and `ChatService` preserve Gatekeeper audit metadata in fallback/default paths.
|
||||
- `DiagnosisTraceEvaluator` can optionally validate expected Gatekeeper rule set version.
|
||||
- `mvp/eval/cases/diagnosis-cases.json` now includes narrow-scope and no-evidence matrix cases.
|
||||
- `mvp/eval/reports/baseline-report.json` and `.md` were regenerated for the expanded fixed matrix.
|
||||
- `mvp/demo/evidence-pipeline-scenarios.md` documents live vs fixture-backed demo scenarios.
|
||||
|
||||
## Decisions
|
||||
|
||||
- Keep this phase internal: no public API, no DB schema, no Planner output change.
|
||||
- Keep Gatekeeper deterministic Java validation; the catalog is metadata/config only.
|
||||
- Treat live demo as compatibility evidence and fixture-backed eval as deterministic regression evidence.
|
||||
- Persist audit under the existing `self_evaluation.verifier_evaluation.gatekeeper_result` structure.
|
||||
|
||||
## Runtime Finding
|
||||
|
||||
The first Maven E2E startup found a real integration issue: `ExecutorGatekeeperService` had multiple public constructors without an annotated constructor, so Spring could not instantiate the service. The fix was to annotate the production constructor with `@Autowired`.
|
||||
@@ -41,4 +41,4 @@ Make Executor cite concrete tool evidence, make Gatekeeper validate that citatio
|
||||
## OpenSpec
|
||||
|
||||
- Change: `openspec/changes/verifier-evidence-reference-fidelity`
|
||||
- Source issue: `mvp/issues/ISS-007-verifier-evidence-summary-fidelity.md`
|
||||
- Source issue: `mvp/issues/archived/ISS-007-verifier-evidence-summary-fidelity.md`
|
||||
|
||||
@@ -0,0 +1,46 @@
|
||||
# Acceptance
|
||||
|
||||
## Static Verification
|
||||
|
||||
- `openspec validate interview-demo-quality-audit --strict`
|
||||
- Result: passed.
|
||||
- Coverage: OpenSpec proposal/design/spec/tasks consistency.
|
||||
- PowerShell parser/runtime readiness check:
|
||||
- Command: `powershell -NoProfile -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1 -BaseUrl http://127.0.0.1:1 -OutputDir target/demo-check-syntax`
|
||||
- Result: expected failure with actionable readiness message.
|
||||
- Coverage: script parses under Windows PowerShell and fails before issuing diagnosis requests when service is unreachable.
|
||||
|
||||
## Script Verification
|
||||
|
||||
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
|
||||
- Result: passed.
|
||||
- Coverage: 12/12 fixed eval fixtures, Prompt audit evaluator checks, Gatekeeper rule metadata checks, regenerated baseline reports.
|
||||
- `mvn -q "-Dtest=ChatServiceSequentialAgentTest" test`
|
||||
- Result: passed.
|
||||
- Coverage: Chat verifier evaluation persists `prompt_audit`.
|
||||
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ChatServiceSequentialAgentTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`
|
||||
- Result: passed.
|
||||
- Coverage: broader eval, baseline diff, Chat sequential flow, Gatekeeper, and Verifier input hook regression set.
|
||||
- `mvn -q -DskipTests compile`
|
||||
- Result: passed.
|
||||
- Coverage: main source compilation.
|
||||
|
||||
## E2E Verification
|
||||
|
||||
- Started service with:
|
||||
- `mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo`
|
||||
- Ran:
|
||||
- `powershell -NoProfile -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1 -BaseUrl http://localhost:9900 -SessionId mvp-demo-interview-quality-audit-001`
|
||||
- Result: passed.
|
||||
- Summary:
|
||||
- `chatSuccess=true`
|
||||
- `verdict=LOW_CONFID`
|
||||
- `gatekeeperStatus=fail`
|
||||
- `gatekeeperRuleSetVersion=gatekeeper-rules-v1`
|
||||
- `promptAuditVersion=chat-prompts-v1`
|
||||
- tools included `lookup_knowledge`, `query_logs`, `query_metrics`, and `get_available_log_topics`
|
||||
- Note: live E2E remains a compatibility check, not the deterministic PASS oracle. The fixed fixture baseline is the regression source of truth.
|
||||
|
||||
## Not Verified
|
||||
|
||||
- Browser UI inspection was not required for this change because the scope is backend trace/eval/demo script documentation, not frontend behavior.
|
||||
@@ -0,0 +1,23 @@
|
||||
# Interview Demo Quality Audit Brief
|
||||
|
||||
## Background
|
||||
|
||||
The MVP already demonstrates traceable Agent diagnosis with Planner, Executor, Gatekeeper, Verifier, Composer, evidence tools, trace persistence, and deterministic eval fixtures. The remaining interview-readiness gap is not a new Agent architecture; it is making the demo easier to run and making prompt/rule changes easier to audit.
|
||||
|
||||
## Goal
|
||||
|
||||
Stabilize the interview demo path, expand fixture-backed evaluation, and persist prompt/Gatekeeper audit metadata so the project can explain and verify Agent behavior during interviews.
|
||||
|
||||
## Scope
|
||||
|
||||
- Add prompt audit metadata to Chat verifier evaluation.
|
||||
- Extend deterministic eval cases and baseline reports.
|
||||
- Add an interview demo preflight/check script.
|
||||
- Update MVP demo and architecture documentation.
|
||||
|
||||
## Non-goals
|
||||
|
||||
- No public API or database schema changes.
|
||||
- No new SubAgent split, MCP migration, process isolation, or AIOps LLM Verifier.
|
||||
- No guarantee that every live LLM run returns PASS.
|
||||
|
||||
@@ -0,0 +1,111 @@
|
||||
# interview-demo-quality-audit Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: stabilize the interview demo, expand deterministic eval coverage, and add Prompt/Gatekeeper version audit.
|
||||
- Slug: `interview-demo-quality-audit`.
|
||||
- Devflow scale: `standard`.
|
||||
- Interface impact: L2 internal contract change because `verifier_evaluation` gains `prompt_audit`; no public HTTP API or database schema change.
|
||||
|
||||
## Context
|
||||
|
||||
- `devflow/index.md` used: related entries found for `diagnosis-eval-demo-gatekeeper-closure`, `executor-composer-final-answer`, `verifier-evidence-reference-fidelity`, `mvp-demo-interview-runbook`, and `diagnosis-eval-baseline-diff`.
|
||||
- Relevant glossary:
|
||||
- Evidence Tools produce incident facts and must be recorded in `tool_invocation`.
|
||||
- Verifier should not use skills/runbooks as incident evidence.
|
||||
- Diagnosis Playbook Skill is workflow guidance, not a fact source.
|
||||
- Historical constraints that must enter OpenSpec:
|
||||
- Diagnosis eval is deterministic and fixture-backed; no LLM-as-judge.
|
||||
- Stable demo scenarios are documentation/payloads plus deterministic fixtures; live E2E is a compatibility check, not a guaranteed PASS oracle.
|
||||
- Gatekeeper rule metadata is already metadata-only and should not become dynamic rule execution.
|
||||
- Composer is the final expression layer and must not leak raw Executor JSON.
|
||||
|
||||
## Question Pool
|
||||
|
||||
| ID | Dimension | Mode | Question | Status |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | Terminology | evidence-driven | What should the new audit metadata be called? | Resolved |
|
||||
| Q2 | Boundary | evidence-driven | Does this require public API or schema changes? | Resolved |
|
||||
| Q3 | Acceptance | evidence-driven | Which current assets define deterministic acceptance? | Resolved |
|
||||
| Q4 | Technical | evidence-driven | Where should prompt version metadata live with minimal implementation risk? | Resolved |
|
||||
| Q5 | Scope | user-interview | Should live E2E be mandatory for all scenarios? | Confirmed by objective as conditional |
|
||||
|
||||
## Evidence-driven Conclusions
|
||||
|
||||
- Q1 conclusion: use `prompt_audit` for prompt version metadata and keep existing `gatekeeper_result.rule_set_version`.
|
||||
- Q2 conclusion: keep this as an internal trace/self-evaluation contract change. Do not add endpoints, tables, or new Agent roles.
|
||||
- Q3 conclusion: `DiagnosisTraceEvaluatorTest`, baseline reports, fixed fixtures, and demo scripts define current acceptance style.
|
||||
- Q4 conclusion: add a small Chat prompt audit catalog near `ChatService` prompt loading and persist a compact snapshot with verifier evaluation.
|
||||
- Q5 conclusion: run live E2E with `mvp-demo` profile if dependencies are available; otherwise record the blocker and rely on deterministic eval/unit evidence.
|
||||
|
||||
## Specify / Alignment
|
||||
|
||||
| Check | Status | Notes |
|
||||
|---|---|---|
|
||||
| proposal goals -> proposal | Aligned | Proposal covers demo preflight, eval expansion, prompt audit, and docs. |
|
||||
| proposal scope/constraints -> design | Aligned | Design records no public API/schema changes, prompt audit shape, eval fields, and demo script behavior. |
|
||||
| design decisions -> specs/tasks | Aligned | Specs cover persisted prompt audit, evaluator checks, baseline, and demo script outputs. |
|
||||
| specs observable behavior -> tasks | Aligned | Each requirement has implementation and verification tasks. |
|
||||
|
||||
## Audit
|
||||
|
||||
Input -> processing -> output chain:
|
||||
|
||||
```text
|
||||
prompt resource metadata
|
||||
-> ChatService / PromptAudit snapshot
|
||||
-> verifier_evaluation.prompt_audit
|
||||
-> Trace API / eval fixtures
|
||||
-> DiagnosisTraceEvaluator baseline
|
||||
|
||||
run-interview-demo-check.ps1
|
||||
-> service readiness
|
||||
-> chat / trace / feedback
|
||||
-> mvp/demo/output summary
|
||||
```
|
||||
|
||||
Architecture risk assessment:
|
||||
|
||||
1. The audit shape is intentionally compact and internal; storing full prompt text would create noisy traces and possible sensitive-content risk.
|
||||
2. Eval should assert versions by explicit metadata, not by prompt content hashes that churn during local prompt edits.
|
||||
3. Live demo checks may still be LOW_CONFID because LLM output is not deterministic; deterministic fixtures remain the regression source of truth.
|
||||
4. No devflow/OpenSpec conflict found.
|
||||
|
||||
## Commit Gate
|
||||
|
||||
- `openspec validate interview-demo-quality-audit --strict`: passed.
|
||||
- File completeness:
|
||||
- `proposal.md`: present.
|
||||
- `design.md`: present.
|
||||
- `specs/`: present for `chat-verifier-agent`, `diagnosis-eval-harness`, and `mvp-demo-trace-acceptance`.
|
||||
- `tasks.md`: present.
|
||||
- Consistency:
|
||||
- Proposal goals map to design sections.
|
||||
- Design decisions map to spec requirements and executable tasks.
|
||||
- Task acceptance checks are verifiable.
|
||||
- `.committed` marker created.
|
||||
|
||||
## Current Checkpoint
|
||||
|
||||
- Commit completed.
|
||||
- Apply is authorized by the original objective: "完成后归档提交".
|
||||
|
||||
## Pre-apply Research
|
||||
|
||||
- Capability source: sm-flow built-in apply protocol. `openspec-apply-change` was not invoked directly in this session.
|
||||
- Repository semantic search/LSP note: the requested `codebase-retrieval` and LSP tools were not available in the exposed toolset, so impact analysis used `rg`, direct file reads, OpenSpec/devflow artifacts, and targeted tests.
|
||||
- Reference implementation and reuse:
|
||||
- `ChatService.persistVerifierEvaluation(...)` is the single persistence point for Chat verifier/composer audit data; prompt audit was added there to cover normal, fallback, and degraded Composer paths.
|
||||
- `DiagnosisTraceEvaluator` and `DiagnosisEvalReportWriter` are the deterministic eval extension points; no LLM judge was introduced.
|
||||
- `mvp/demo/scripts/run-payment-timeout-demo.ps1` provided the request/trace/feedback flow reused by the new interview preflight script.
|
||||
- Interface impact remains L2 internal trace contract: `verifier_evaluation.prompt_audit` and eval report fields are added; no public endpoint, table, or request DTO changed.
|
||||
|
||||
## Apply Notes
|
||||
|
||||
- Added compact Chat prompt audit metadata: `chat-prompts-v1`, with planner/executor/verifier/composer prompt versions and resource paths.
|
||||
- Extended diagnosis eval schema, result reporting, baseline fixtures, JSON report, and Markdown report for Prompt audit and Gatekeeper rule metadata.
|
||||
- Added two fixture-backed audit cases:
|
||||
- `prompt-gatekeeper-audit-closure`
|
||||
- `audit-metadata-low-confid`
|
||||
- Added `mvp/demo/scripts/run-interview-demo-check.ps1` to run service readiness, Chat, Trace, feedback, and summary output.
|
||||
- Updated MVP demo/eval/architecture docs to explain `prompt_audit.version`, `gatekeeper_result.rule_set_version`, and deterministic fixture baseline.
|
||||
@@ -0,0 +1,58 @@
|
||||
# Evidence
|
||||
|
||||
## Context Files Read
|
||||
|
||||
- `devflow/index.md`
|
||||
- `devflow/glossary/CONTEXT.md`
|
||||
- `devflow/projects/2026-07-08-diagnosis-eval-demo-gatekeeper-closure/decisions.md`
|
||||
- `devflow/projects/2026-07-08-executor-composer-final-answer/decisions.md`
|
||||
- `mvp/architecture/current-mvp-architecture.md`
|
||||
- `mvp/architecture/agent-orchestration.md`
|
||||
- `mvp/architecture/executor-evidence-pipeline-refactor.md`
|
||||
- `mvp/architecture/harness-quality-gates.md`
|
||||
- `mvp/demo/README.md`
|
||||
- `mvp/demo/ten-minute-interview-demo.md`
|
||||
- `mvp/eval/README.md`
|
||||
- `mvp/eval/cases/diagnosis-cases.json`
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ExecutorGatekeeperService.java`
|
||||
- `src/main/java/com/superbiz/agent/eval/DiagnosisTraceEvaluator.java`
|
||||
- `src/main/resources/gatekeeper/gatekeeper-rules.json`
|
||||
|
||||
## Tooling Note
|
||||
|
||||
The required `codebase-retrieval` and LSP tools were not exposed in this session. Impact analysis used `rg`, direct file reads, existing OpenSpec/devflow artifacts, and targeted tests instead.
|
||||
|
||||
## Implementation Evidence
|
||||
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- Adds `prompt_audit` under `verifier_evaluation` through the shared `persistVerifierEvaluation(...)` path.
|
||||
- Uses compact metadata only: audit version, prompt names, prompt versions, and resource paths.
|
||||
- `src/main/java/com/superbiz/agent/eval/DiagnosisTraceEvaluator.java`
|
||||
- Adds deterministic checks for `requirePromptAudit`, `expectedPromptAuditVersion`, `expectedPromptVersions`, and `requireGatekeeperRules`.
|
||||
- `src/main/java/com/superbiz/agent/eval/DiagnosisEvalReportWriter.java`
|
||||
- Adds Prompt Audit and Gatekeeper rule count columns to Markdown reports.
|
||||
- `mvp/eval/cases/diagnosis-cases.json`
|
||||
- Expands fixed baseline to 12 fixture-backed cases.
|
||||
- `mvp/eval/fixtures/prompt-gatekeeper-audit-closure-pass.json`
|
||||
- Positive PASS fixture proving Prompt audit and Gatekeeper rule metadata closure.
|
||||
- `mvp/eval/fixtures/audit-metadata-low-confid.json`
|
||||
- LOW_CONFID fixture proving safe answer behavior while audit metadata remains present.
|
||||
- `mvp/demo/scripts/run-interview-demo-check.ps1`
|
||||
- Adds service readiness, Chat, Trace, feedback, and summary output for interview preflight.
|
||||
|
||||
## Verification Evidence
|
||||
|
||||
- OpenSpec:
|
||||
- `openspec validate interview-demo-quality-audit --strict`: passed before archive.
|
||||
- `openspec validate --specs --strict`: 10 specs passed after merging deltas into main specs.
|
||||
- Unit/eval:
|
||||
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`: passed.
|
||||
- `mvn -q "-Dtest=ChatServiceSequentialAgentTest" test`: passed.
|
||||
- `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ChatServiceSequentialAgentTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest" test`: passed.
|
||||
- Compile:
|
||||
- `mvn -q -DskipTests compile`: passed.
|
||||
- E2E:
|
||||
- Started `mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo`.
|
||||
- Ran `mvp/demo/scripts/run-interview-demo-check.ps1` against `http://localhost:9900`.
|
||||
- Summary recorded `chatSuccess=true`, `verdict=LOW_CONFID`, `gatekeeperRuleSetVersion=gatekeeper-rules-v1`, and `promptAuditVersion=chat-prompts-v1`.
|
||||
@@ -0,0 +1,103 @@
|
||||
# Acceptance
|
||||
|
||||
## 实现结果
|
||||
|
||||
- OpenSpec tasks: `42/42` complete。
|
||||
- Phase commits:
|
||||
- `52bf030 feat(trace): add session run isolation schema`
|
||||
- `6fdbd34 docs(openspec): tighten run isolation contract`
|
||||
- `26d5529 feat(trace): isolate chat runs`
|
||||
- `027aed1 feat(trace): add run-scoped trace reads`
|
||||
- `d928a19 feat(trace): bind feedback to runs`
|
||||
- `78c1477 feat(trace): isolate aiops runs`
|
||||
- `f9df943 feat(trace): finish run-aware demo verification`
|
||||
- OpenSpec archive: `openspec/changes/archive/2026-07-10-session-run-trace-isolation`
|
||||
|
||||
## 静态验证
|
||||
|
||||
```powershell
|
||||
node --check src\main\resources\static\app.js
|
||||
node --check src\main\resources\static\trace.js
|
||||
openspec validate session-run-trace-isolation --strict
|
||||
git diff --check -- . ':!devflow/index.md'
|
||||
```
|
||||
|
||||
结果:通过。
|
||||
|
||||
## 脚本验证
|
||||
|
||||
PowerShell demo 脚本解析:
|
||||
|
||||
```powershell
|
||||
$scripts = @(
|
||||
'mvp\demo\scripts\run-payment-timeout-demo.ps1',
|
||||
'mvp\demo\scripts\run-interview-demo-check.ps1'
|
||||
)
|
||||
foreach ($script in $scripts) {
|
||||
[scriptblock]::Create((Get-Content -Raw -Encoding UTF8 $script)) | Out-Null
|
||||
}
|
||||
```
|
||||
|
||||
结果:通过。
|
||||
|
||||
Focused tests:
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=ChatControllerTest,DiagnosisTraceServiceTest,FeedbackControllerTest,FeedbackServiceTest,AiOpsServiceTest" test
|
||||
```
|
||||
|
||||
结果:通过。
|
||||
|
||||
Baseline / regression:
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest" test
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
|
||||
```
|
||||
|
||||
结果:通过,无 baseline drift。
|
||||
|
||||
## E2E 验证
|
||||
|
||||
使用 Maven 启动:
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
|
||||
```
|
||||
|
||||
E2E 使用同一 `sessionId` 连续两轮 Chat:
|
||||
|
||||
- `sessionId`: `e2e-phase6-chat-codex-20260710-2120`
|
||||
- `run1`: `run-e2a97696-4398-4abc-90e4-28f45c838f92`
|
||||
- `run2`: `run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172`
|
||||
|
||||
验证结果:
|
||||
|
||||
- Chat1 / Chat2 均成功。
|
||||
- run1 exact trace 返回 run1。
|
||||
- run2 exact trace 返回 run2。
|
||||
- session-only latest trace 返回 run2。
|
||||
- DB 中同一 session 有两条 `diagnosis_run`。
|
||||
- step/tool rows 按 `run_id` 隔离,mixed row check 为 0。
|
||||
- `chat_session.message_pair_count = 2`。
|
||||
|
||||
## 日志验证
|
||||
|
||||
检查:
|
||||
|
||||
- `target/e2e/phase6-mvn-20260710-211831.out.log`
|
||||
- `logs/application.log`
|
||||
- `logs/chat.log`
|
||||
|
||||
结果:能找到 E2E `sessionId`、两个 `runId`、Chat execution、run persistence 和 trace lookup 相关日志。
|
||||
|
||||
## 浏览器/人工验证
|
||||
|
||||
未单独进行浏览器点击验证。Trace UI 的本次验收通过静态语法检查、URL/runId 参数代码审查和后端 exact trace E2E 共同覆盖。建议后续手动打开 `trace.html?sessionId=...&runId=...` 做展示层冒烟。
|
||||
|
||||
## 剩余风险 / 后续事项
|
||||
|
||||
- 缺少 `runId` 的 Feedback fallback 是短期兼容路径,客户端全部迁移后可收紧。
|
||||
- `diagnosis_session` 仍保留为历史兼容和回滚表,后续需要观察窗口后再评估约束收紧或归档策略。
|
||||
- `case_library.diagnosis_id` 仍是过渡字段,旧值可能为 `session_id`,新自动值为 `run_id`。
|
||||
- 历史 mixed trace 不能恢复真实多轮边界,只能按 compatibility run 查询。
|
||||
@@ -0,0 +1,41 @@
|
||||
# Session / Run / Trace Isolation
|
||||
|
||||
## 背景
|
||||
|
||||
同一个 `sessionId` 以前同时代表多轮 Chat 上下文和一次持久化诊断 Trace。端到端验证发现,同一 `sessionId` 连续两轮 Chat 时,Redis 多轮上下文是正确的,但 MySQL 中 `diagnosis_session` 会被后一轮覆盖,`agent_step` 和 `tool_invocation` 会按同一个 `session_id` 混在一起。
|
||||
|
||||
这会导致 Trace 回放、Verifier/Evaluation 读数、Feedback 绑定和 `case_library` 来源都可能跨轮污染。
|
||||
|
||||
## 目标
|
||||
|
||||
- 将会话态和运行态拆开:`chat_session` 保存会话元数据,`diagnosis_run` 保存一次诊断运行。
|
||||
- 引入正式 API 字段 `runId`,作为一次可回放诊断执行的边界。
|
||||
- `agent_step` 和 `tool_invocation` 保留原 Trace 明细角色,新增 `run_id` 并按 run 隔离读写。
|
||||
- Trace、Feedback、CaseLibrary、AIOps、demo 脚本和 Trace UI 都支持 run-aware 流程。
|
||||
- 保留旧 `diagnosis_session` 作为历史兼容和回滚表。
|
||||
- 完成 Maven E2E、DB 检查、日志检查和 baseline drift 验证。
|
||||
|
||||
## 范围
|
||||
|
||||
- Flyway/JPA 增加 `chat_session`、`diagnosis_run`,并给 `agent_step`、`tool_invocation` 增加 `run_id`。
|
||||
- Chat 每次有效执行创建一个新的 `diagnosis_run`,响应返回 `sessionId + runId`。
|
||||
- Trace API 支持 latest-run fallback 和 exact-run 查询:`GET /api/diagnosis/{sessionId}/trace?runId=...`。
|
||||
- 新增 run list API:`GET /api/chat/session/{sessionId}/runs`。
|
||||
- Feedback 优先绑定 `runId`,缺省时短期 fallback 到 latest run 并返回 `fallbackToLatestRun=true`。
|
||||
- AIOps 每次有效执行创建并透出 `runId`,SSE 保持 `message` event name 并发送 `type=metadata`。
|
||||
- MVP demo、Trace UI、表文档和架构文档统一为 `chat_session -> diagnosis_run -> trace detail(run_id)`。
|
||||
|
||||
## 非目标
|
||||
|
||||
- 不新增 `diagnosis_trace` 或 `trace_event` 主表。
|
||||
- 不实现完整 run-list UI。
|
||||
- 不删除旧 `diagnosis_session`。
|
||||
- 不改变 Redis 对话历史窗口策略。
|
||||
- 不把完整多轮正文历史持久化到 MySQL。
|
||||
- 不尝试把历史混合 trace 还原成真实多轮边界。
|
||||
|
||||
## 关联
|
||||
|
||||
- OpenSpec: `openspec/changes/archive/2026-07-10-session-run-trace-isolation`
|
||||
- Change slug: `session-run-trace-isolation`
|
||||
- 分档: complex
|
||||
@@ -0,0 +1,48 @@
|
||||
# Decisions
|
||||
|
||||
## 核心决策
|
||||
|
||||
| 决策 | 选择 | 理由 |
|
||||
|---|---|---|
|
||||
| 领域拆分 | 新增 `chat_session` 和 `diagnosis_run` | 会话元数据和一次诊断执行的生命周期不同,继续塞在一张表会导致上下文膨胀和边界混淆 |
|
||||
| Trace 明细 | 复用 `agent_step` / `tool_invocation`,增加 `run_id` | 现有明细表已经能表达 Trace,隔离需要 run key,不需要新事件模型 |
|
||||
| API 身份 | `runId = "run-" + UUID` | 外部 ID 不依赖数据库自增 ID,碰撞风险低 |
|
||||
| Trace 兼容 | 缺少 `runId` 时按 `created_at DESC, id DESC` 解析 latest run | 保留旧客户端兼容性,避免 feedback/eval 更新 `updated_at` 后改变 latest 判定 |
|
||||
| 历史迁移 | 每条旧 `diagnosis_session` 生成一条 compatibility run | 旧混合数据没有真实轮次边界,不能伪造多 run 历史 |
|
||||
| Feedback fallback | 缺少 `runId` 时短期绑定 latest run 并返回 `fallbackToLatestRun=true` | 老客户端可继续工作,同时让歧义可观测 |
|
||||
| Case provenance | 新自动案例写 `case_library.diagnosis_id = run_id` | 保留旧列,文档声明过渡语义 |
|
||||
| AIOps 范围 | 同一个 change 内完成 AIOps run isolation | AIOps 是一等 Trace 入口,不能留下同类混合 trace bug |
|
||||
| 所有权校验 | 服务层校验 run/session ownership,暂不加 DB 外键 | 兼容历史 orphan rows 和回滚窗口 |
|
||||
|
||||
## 用户确认
|
||||
|
||||
- 选择拆 `chat_session` 和 `diagnosis_run`,不只是在旧表加字段。
|
||||
- `chat_session` 第一阶段只保存元数据,不保存完整对话正文。
|
||||
- 完整多轮对话历史继续放在 Redis `SessionContext.messageHistory`。
|
||||
- `runId` 是正式 API 字段。
|
||||
- Trace 缺少 `runId` 时短期默认查 latest run。
|
||||
- Feedback 缺少 `runId` 时短期 fallback,长期可再收紧。
|
||||
- 每次有效 Chat/AIOps 都创建 run。
|
||||
- 不新增 `diagnosis_trace` / `trace_event` 主表。
|
||||
- 旧 `diagnosis_session` 保留用于历史和回滚,新代码不再写新执行态。
|
||||
- demo 脚本和 Trace UI 做最小 `runId` 支持。
|
||||
|
||||
## 接口影响
|
||||
|
||||
级别:L4。
|
||||
|
||||
- 新 API 响应字段:`runId`。
|
||||
- Trace API 新 query 参数:`runId`。
|
||||
- 新 API:`GET /api/chat/session/{sessionId}/runs`。
|
||||
- Feedback request 新增 optional/preferred `runId`。
|
||||
- Feedback response 新增 bound `runId` 和 `fallbackToLatestRun`。
|
||||
- `/api/ai_ops` SSE 保持 event name `message`,新增 `type=metadata` 消息。
|
||||
- DB contract 新增两张表和两个 `run_id` 列。
|
||||
- 旧 `sessionId` only 调用仍兼容,但 fallback 必须可观测。
|
||||
|
||||
## 风险接受
|
||||
|
||||
- 历史混合 trace 无法真实拆分,只能作为 compatibility run。
|
||||
- 上下文传播同时依赖 `RunnableConfig.metadata` 和 `SessionContextHolder`,后续改动必须注意 `sessionId/runId` 同步。
|
||||
- `case_library.diagnosis_id` 在过渡期存在 `session_id` 和 `run_id` 两种语义。
|
||||
- 缺少 `runId` 的 Feedback 仍有歧义,后续客户端迁移完成后可收紧为参数错误。
|
||||
@@ -0,0 +1,76 @@
|
||||
# Evidence
|
||||
|
||||
## 上下文证据
|
||||
|
||||
- `SessionContext.messageHistory` 和 `getMessagePairCount()` 证明 Redis 承载热对话历史;MySQL 只需要长期审计的会话目录和运行记录。
|
||||
- `CaseLibraryService.createFromSession` 原先按 `DiagnosisSession.sessionId` 去重并映射 query/answer,因此 run 隔离后需要新增 `createFromRun`。
|
||||
- 旧 `mvp/architecture/data-model.md` 把 `case_library.diagnosis_id` 解释为 `diagnosis_session.session_id`,本次改为过渡语义:旧数据可能是 `session_id`,新自动案例是 `run_id`。
|
||||
- 既有 Trace OpenSpec 要求 `GET /api/diagnosis/{sessionId}/trace` 是只读端点;latest-run 和 exact-run 查询都必须保持只读。
|
||||
- ISS-010 的 E2E 事实显示同一 `sessionId` 两轮 Chat 会产生 MySQL Trace 混合,是本 change 的直接触发证据。
|
||||
|
||||
## 实现证据
|
||||
|
||||
- Phase 1 增加 `V011__add_session_run_isolation.sql`,创建 `chat_session`、`diagnosis_run`,并为 `agent_step` / `tool_invocation` 增加 nullable `run_id`。
|
||||
- Phase 2 将 Chat 写路径切到 `chat_session + diagnosis_run`,并让 Hook/Tool/Evaluation/Gatekeeper 使用 run-scoped 数据。
|
||||
- Phase 3 将 Trace API 改为 latest-run / exact-run 双模式,并加入 lightweight run summaries。
|
||||
- Phase 4 将 Feedback 和 CaseLibrary 绑定到 run,保留没有 run-backed 数据时的 legacy fallback。
|
||||
- Phase 5 将 AIOps 接入 run isolation,SSE metadata 暴露 `sessionId + runId`。
|
||||
- Phase 6 更新 demo 脚本、Trace UI、MVP 架构文档和表文档,并修正 review 后发现的 session-only 文档残留。
|
||||
|
||||
## E2E 证据
|
||||
|
||||
Maven 启动命令:
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
|
||||
```
|
||||
|
||||
日志:
|
||||
|
||||
- `target/e2e/phase6-mvn-20260710-211831.out.log`
|
||||
- `target/e2e/phase6-mvn-20260710-211831.err.log`
|
||||
- `logs/application.log`
|
||||
- `logs/chat.log`
|
||||
|
||||
E2E session:
|
||||
|
||||
- `sessionId`: `e2e-phase6-chat-codex-20260710-2120`
|
||||
- `run1`: `run-e2a97696-4398-4abc-90e4-28f45c838f92`
|
||||
- `run2`: `run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172`
|
||||
|
||||
结果:
|
||||
|
||||
- 两轮 Chat 都成功,并复用同一个 `sessionId`。
|
||||
- 两轮返回不同 `runId`。
|
||||
- run1 exact trace 只返回 run1。
|
||||
- run2 exact trace 只返回 run2。
|
||||
- session-only Trace latest fallback 返回 run2。
|
||||
- `chat_session.message_pair_count = 2`,证明多轮上下文连续。
|
||||
|
||||
## DB 证据
|
||||
|
||||
通过 `scripts/query_mysql.py` 检查:
|
||||
|
||||
- `diagnosis_run` 中该 E2E session 有 2 条 `SUCCESS / CHAT` 运行。
|
||||
- `agent_step` 按 run 分组:run1 `10` 行,run2 `9` 行。
|
||||
- `tool_invocation` 按 run 分组:run1 `14` 行,run2 `8` 行。
|
||||
- mixed row check 为 `0`,没有 NULL 或 unexpected `run_id` 混入该 E2E session。
|
||||
|
||||
## Baseline 证据
|
||||
|
||||
运行:
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest" test
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
|
||||
```
|
||||
|
||||
结果:
|
||||
|
||||
- 两组 baseline / regression 命令通过。
|
||||
- baseline harness 使用离线 fixture,不依赖 live DB/session tables。
|
||||
- 未观察到 baseline drift。
|
||||
|
||||
## 工具限制
|
||||
|
||||
AGENTS 要求的 `codebase-retrieval` 和 LSP 工具在本会话不可用。替代验证使用 OpenSpec、`rg`、定向阅读、 focused tests、E2E、DB 查询和日志检查。
|
||||
+28
-18
@@ -1,8 +1,8 @@
|
||||
# SuperBizAgent MVP 文档
|
||||
|
||||
**更新日期**:2026-07-05
|
||||
**更新日期**:2026-07-10
|
||||
|
||||
本目录保存 MVP 阶段的架构、问题、演示、评测和数据表说明。当前架构入口已经整理到 `mvp/architecture/`,旧版架构材料已归档,避免继续把历史方案当成当前实现。
|
||||
本目录保存 MVP 阶段的架构、问题、演示、评测和数据表说明。当前材料按“当前入口”和“历史归档”拆开,避免把早期设计稿当成当前实现。
|
||||
|
||||
## 当前入口
|
||||
|
||||
@@ -12,6 +12,7 @@
|
||||
| [architecture/current-mvp-architecture.md](architecture/current-mvp-architecture.md) | 当前可运行系统架构 |
|
||||
| [architecture/interview-one-pager.md](architecture/interview-one-pager.md) | 面试一页式架构讲解 |
|
||||
| [architecture/agent-orchestration.md](architecture/agent-orchestration.md) | Agent 编排架构 |
|
||||
| [architecture/executor-evidence-pipeline-refactor.md](architecture/executor-evidence-pipeline-refactor.md) | Executor 证据链路改造记录 |
|
||||
| [architecture/harness-quality-gates.md](architecture/harness-quality-gates.md) | Harness 与质量门禁 |
|
||||
| [architecture/rag-architecture.md](architecture/rag-architecture.md) | RAG/知识检索新架构 |
|
||||
| [architecture/retrieval-observability.md](architecture/retrieval-observability.md) | 检索与可观测性架构 |
|
||||
@@ -20,15 +21,16 @@
|
||||
| [architecture/knowledge-base-authoring.md](architecture/knowledge-base-authoring.md) | 知识库文档编写与维护 |
|
||||
| [architecture/data-model.md](architecture/data-model.md) | 数据模型总览 |
|
||||
| [architecture/evolution-roadmap.md](architecture/evolution-roadmap.md) | Agent 架构演进路线 |
|
||||
| [issues/rag-refactor-plan.md](issues/rag-refactor-plan.md) | RAG 重构计划和阶段拆解 |
|
||||
| [issues/README.md](issues/README.md) | MVP issue 索引 |
|
||||
| [issues/active/rag-refactor-plan.md](issues/active/rag-refactor-plan.md) | RAG 重构计划和阶段拆解 |
|
||||
| [tables/README.md](tables/README.md) | 当前 MySQL 表说明 |
|
||||
| [demo/README.md](demo/README.md) | Demo 运行和面试演示材料 |
|
||||
| [demo/ten-minute-interview-demo.md](demo/ten-minute-interview-demo.md) | 10 分钟面试演示脚本 |
|
||||
| [eval/README.md](eval/README.md) | 诊断评测材料 |
|
||||
| [issues/README.md](issues/README.md) | MVP issue 索引 |
|
||||
|
||||
## 当前系统一句话
|
||||
|
||||
SuperBizAgent MVP 是一个可追踪的故障诊断 Agent:Chat 和 AIOps 入口进入 Agent 编排,Executor 显式调用知识库、日志、指标等工具收集证据,诊断过程落到 `diagnosis_session`、`agent_step`、`tool_invocation`,最终通过 Trace API、Verifier 和评测脚本证明结果可解释、可回放、可对比。
|
||||
SuperBizAgent MVP 是一个可追踪的故障诊断 Agent:Chat 和 AIOps 入口进入 Agent 编排,Executor 显式调用知识库、日志、指标等工具收集证据;多轮会话元数据落到 `chat_session`,每次诊断运行落到 `diagnosis_run`,步骤和工具明细通过 `agent_step.run_id`、`tool_invocation.run_id` 关联,最终通过 Trace API、Verifier 和评测脚本证明结果可解释、可回放、可对比。
|
||||
|
||||
## 文档结构
|
||||
|
||||
@@ -39,6 +41,7 @@ mvp/
|
||||
current-mvp-architecture.md
|
||||
interview-one-pager.md
|
||||
agent-orchestration.md
|
||||
executor-evidence-pipeline-refactor.md
|
||||
harness-quality-gates.md
|
||||
rag-architecture.md
|
||||
retrieval-observability.md
|
||||
@@ -47,12 +50,17 @@ mvp/
|
||||
knowledge-base-authoring.md
|
||||
data-model.md
|
||||
evolution-roadmap.md
|
||||
archive/2026-07-05-legacy/
|
||||
archive/
|
||||
issues/
|
||||
README.md
|
||||
rag-refactor-plan.md
|
||||
ISS-*.md
|
||||
rag-*.md
|
||||
active/
|
||||
archived/
|
||||
design-notes/
|
||||
rag/
|
||||
tables/
|
||||
README.md
|
||||
*表-*.md
|
||||
archive/
|
||||
demo/
|
||||
README.md
|
||||
ten-minute-interview-demo.md
|
||||
@@ -65,9 +73,7 @@ mvp/
|
||||
cases/
|
||||
fixtures/
|
||||
reports/
|
||||
notes/
|
||||
plan/
|
||||
tables/
|
||||
archive/
|
||||
```
|
||||
|
||||
## 当前核心设计
|
||||
@@ -77,7 +83,8 @@ mvp/
|
||||
- `VectorSearchService` 是检索稳定门面。
|
||||
- Spring AI VectorStore 是当前读取主路径,Milvus SDK 保留为 fallback。
|
||||
- AIOps payload 会生成推荐知识库 query,保留业务语义。
|
||||
- Trace API 聚合 session、step、tool invocation 和 self evaluation。
|
||||
- `sessionId` 表示多轮会话上下文,`runId` 表示一次可回放诊断运行。
|
||||
- Trace API 聚合 `diagnosis_run`、`agent_step.run_id`、`tool_invocation.run_id` 和 self evaluation。
|
||||
- RAG 行为通过 offline baseline 和 live acceptance 脚本做回归验证。
|
||||
|
||||
## 关键运行链路
|
||||
@@ -87,7 +94,8 @@ Chat
|
||||
-> ChatService
|
||||
-> Planner / Executor / Verifier
|
||||
-> evidence tools
|
||||
-> diagnosis_session / agent_step / tool_invocation
|
||||
-> chat_session / diagnosis_run
|
||||
-> agent_step.run_id / tool_invocation.run_id
|
||||
-> DiagnosisTraceService
|
||||
|
||||
AIOps
|
||||
@@ -96,6 +104,7 @@ AIOps
|
||||
-> Planner / Executor
|
||||
-> Prometheus / logs / lookup_knowledge
|
||||
-> AiOpsRuleEvaluationService
|
||||
-> diagnosis_run(agent_flow=AI_OPS)
|
||||
-> DiagnosisTraceService
|
||||
|
||||
RAG
|
||||
@@ -107,10 +116,11 @@ RAG
|
||||
-> tool_invocation
|
||||
```
|
||||
|
||||
## 旧文档说明
|
||||
## 归档说明
|
||||
|
||||
旧版架构文档已移动到:
|
||||
历史材料分两类:
|
||||
|
||||
- [architecture/archive/2026-07-05-legacy/](architecture/archive/2026-07-05-legacy/)
|
||||
- 旧架构文档:[architecture/archive/2026-07-05-legacy/](architecture/archive/2026-07-05-legacy/)
|
||||
- 本次文档清理归档:[archive/2026-07-09-doc-cleanup/](archive/2026-07-09-doc-cleanup/)
|
||||
|
||||
归档文档只用于追溯设计历史。当前实现和后续规划以 `architecture/current-mvp-architecture.md` 与 `architecture/rag-architecture.md` 为准。
|
||||
归档文档只用于追溯设计历史。当前实现和后续规划以 `architecture/`、`issues/README.md`、`tables/README.md` 和 OpenSpec/devflow 的最新记录为准。
|
||||
|
||||
+14
-12
@@ -1,6 +1,6 @@
|
||||
# MVP 架构文档
|
||||
|
||||
**更新日期**:2026-07-06
|
||||
**更新日期**:2026-07-10
|
||||
|
||||
这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到:
|
||||
|
||||
@@ -15,7 +15,8 @@
|
||||
| [current-mvp-architecture.md](current-mvp-architecture.md) | 当前可运行 MVP 的总体架构、链路、持久化和质量门禁 |
|
||||
| [interview-one-pager.md](interview-one-pager.md) | 面试一页式架构讲解,包含总图、亮点、取舍和追问回答 |
|
||||
| [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat SequentialAgent、AIOps SupervisorAgent、工具边界 |
|
||||
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Verifier、评测基线组成的质量门禁 |
|
||||
| [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md) | Chat 证据链路当前数据契约,覆盖 Executor V2、Gatekeeper、Verifier、Composer、`evidence_refs` |
|
||||
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Gatekeeper、Verifier、Composer、评测基线组成的质量门禁 |
|
||||
| [rag-architecture.md](rag-architecture.md) | RAG/知识检索新架构,覆盖 L0 hint、VectorStore 主路径、SDK fallback、证据追踪 |
|
||||
| [modular-rag-pipeline.md](modular-rag-pipeline.md) | `lookup_knowledge` 模块化 RAG 落地架构,覆盖 pipeline、fallback、evidence-first contract、trace |
|
||||
| [rag-eval-closure.md](rag-eval-closure.md) | RAG 评测闭环,覆盖 offline baseline、baseline diff、diagnosis eval 和 live acceptance |
|
||||
@@ -28,19 +29,20 @@
|
||||
|
||||
## 当前架构一句话
|
||||
|
||||
SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,执行过程落到 `diagnosis_session`、`agent_step`、`tool_invocation`,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
|
||||
SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,Chat 链路由 Gatekeeper 做引用真实性校验、Verifier 做可推导性判断、Composer 生成最终表达;多轮会话元数据落到 `chat_session`,每次诊断运行落到 `diagnosis_run`,步骤和工具明细通过 `agent_step.run_id`、`tool_invocation.run_id` 关联,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
|
||||
|
||||
## 阅读顺序
|
||||
|
||||
1. 先读 [current-mvp-architecture.md](current-mvp-architecture.md),理解系统边界和主链路。
|
||||
2. 面试前读 [interview-one-pager.md](interview-one-pager.md),准备 2-5 分钟讲解。
|
||||
3. 再读 [agent-orchestration.md](agent-orchestration.md),理解当前 Agent 如何协作。
|
||||
4. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。
|
||||
5. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。
|
||||
6. 继续读 [modular-rag-pipeline.md](modular-rag-pipeline.md),看 `lookup_knowledge` 的模块化落地和 evidence-first contract。
|
||||
7. 再读 [rag-eval-closure.md](rag-eval-closure.md),看 RAG baseline 如何形成质量闭环。
|
||||
8. 然后读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
|
||||
9. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
|
||||
10. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
|
||||
11. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
|
||||
12. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
|
||||
4. 接着读 [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md),理解 Chat 证据链路的数据结构和验真边界。
|
||||
5. 然后读 [harness-quality-gates.md](harness-quality-gates.md),理解为什么系统可追踪、可验证。
|
||||
6. 再读 [rag-architecture.md](rag-architecture.md),理解当前 RAG 为什么保留显式 `lookup_knowledge`,以及 Spring AI VectorStore 如何接入。
|
||||
7. 继续读 [modular-rag-pipeline.md](modular-rag-pipeline.md),看 `lookup_knowledge` 的模块化落地和 evidence-first contract。
|
||||
8. 再读 [rag-eval-closure.md](rag-eval-closure.md),看 RAG baseline 如何形成质量闭环。
|
||||
9. 然后读 [retrieval-observability.md](retrieval-observability.md),看检索细节和质量回归方式。
|
||||
10. 再读 [feedback-architecture.md](feedback-architecture.md),理解 self_evaluation、用户反馈和案例沉淀。
|
||||
11. 按需读 [session-trace-lifecycle.md](session-trace-lifecycle.md)、[knowledge-base-authoring.md](knowledge-base-authoring.md)、[data-model.md](data-model.md),补齐运行生命周期、知识库维护和数据关系。
|
||||
12. 最后读 [evolution-roadmap.md](evolution-roadmap.md),区分后续演进和当前实现。
|
||||
13. 需要追溯旧方案时,再进入 `archive/2026-07-05-legacy/`。
|
||||
|
||||
@@ -1,14 +1,14 @@
|
||||
# Agent 编排架构
|
||||
|
||||
**更新日期**:2026-07-05
|
||||
**状态**:当前可运行架构
|
||||
**更新日期**:2026-07-08
|
||||
**状态**:当前可运行架构
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
||||
|
||||
## 1. 设计定位
|
||||
|
||||
旧版 Agent 架构把系统描述为 Supervisor、Planner、SubAgent、Verifier 的团队协作。当前 MVP 保留这个核心思想,但实现更收敛:
|
||||
|
||||
- Chat 链路使用固定顺序工作流:`Planner -> Executor -> Verifier`。
|
||||
- Chat 链路使用固定顺序工作流:`Planner -> Executor -> Gatekeeper -> Verifier -> Composer`。
|
||||
- AIOps 链路使用 `SupervisorAgent` 调度 `Planner + Executor`,最终由规则评估器做轻量验证。
|
||||
- 当前没有拆分 ExternalApiSubAgent、InternalErrorSubAgent、DatabaseSubAgent;这些作为后续演进方向保留。
|
||||
- 证据工具不直接散落在各个 Agent 里,而是通过 Spring AI ToolCallback / `@Tool` 统一暴露。
|
||||
@@ -23,9 +23,11 @@ flowchart TB
|
||||
ChatPlanner --> ChatExecutor["chat_executor"]
|
||||
ChatExecutor --> ChatTools["evidence tools"]
|
||||
ChatTools --> ChatExecutor
|
||||
ChatExecutor --> ChatVerifier["chat_verifier"]
|
||||
ChatExecutor --> ChatGatekeeper["ExecutorGatekeeperService"]
|
||||
ChatGatekeeper --> ChatVerifier["chat_verifier"]
|
||||
ChatVerifier --> ChatDecision{"PASS / LOW_CONFID / REJECT"}
|
||||
ChatDecision --> ChatAnswer["final answer"]
|
||||
ChatDecision --> ChatComposer["chat_composer"]
|
||||
ChatComposer --> ChatAnswer["final answer"]
|
||||
end
|
||||
|
||||
subgraph AiOps["AIOps diagnosis"]
|
||||
@@ -40,20 +42,25 @@ flowchart TB
|
||||
end
|
||||
|
||||
subgraph Trace["Trace persistence"]
|
||||
Session["diagnosis_session"]
|
||||
ChatSession["chat_session"]
|
||||
Run["diagnosis_run"]
|
||||
Step["agent_step"]
|
||||
Invocation["tool_invocation"]
|
||||
SelfEval["self_evaluation"]
|
||||
end
|
||||
|
||||
ChatService --> Session
|
||||
ChatService --> ChatSession
|
||||
ChatService --> Run
|
||||
ChatPlanner --> Step
|
||||
ChatExecutor --> Step
|
||||
ChatGatekeeper --> SelfEval
|
||||
ChatVerifier --> Step
|
||||
ChatTools --> Invocation
|
||||
ChatDecision --> SelfEval
|
||||
ChatComposer --> Step
|
||||
|
||||
AiOpsService --> Session
|
||||
AiOpsService --> ChatSession
|
||||
AiOpsService --> Run
|
||||
AiOpsPlanner --> Step
|
||||
AiOpsExecutor --> Step
|
||||
AiOpsTools --> Invocation
|
||||
@@ -68,9 +75,13 @@ Chat 复杂诊断采用 `SequentialAgent`,顺序固定:
|
||||
chat_planner
|
||||
-> chat_executor
|
||||
-> lookup_knowledge / query_logs / query_metrics / date_time
|
||||
-> outputs executor_evidence_v2
|
||||
-> VerifierInputHook / ExecutorGatekeeperService
|
||||
-> validates source_invocation_id / raw_path / evidence_excerpt
|
||||
-> chat_verifier
|
||||
-> reads tool_trace_summary
|
||||
-> outputs verifier JSON
|
||||
-> judges whether verified evidence can derive claims
|
||||
-> chat_composer
|
||||
-> writes final user-facing answer
|
||||
```
|
||||
|
||||
关键行为:
|
||||
@@ -78,8 +89,10 @@ chat_planner
|
||||
| 角色 | 当前职责 | 输出 |
|
||||
|---|---|---|
|
||||
| `chat_planner` | 拆解问题,注入知识域地图和对话历史,给出排查方向 | `planner_plan` |
|
||||
| `chat_executor` | 按计划调用证据工具,组合工具返回形成诊断答复 | `executor_feedback` |
|
||||
| `chat_verifier` | 只基于已有证据校验 Executor 答案,不做新检索 | `verifier_output` |
|
||||
| `chat_executor` | 按计划调用证据工具,抽取带 `source_invocation_id + raw_path + evidence_excerpt` 的微观事实 | `executor_evidence_v2` |
|
||||
| `ExecutorGatekeeperService` | 在 Verifier 前做代码级引用验真,拒绝伪造 ID、错配 raw_path、错配 excerpt | `gatekeeper_result` |
|
||||
| `chat_verifier` | 只判断已验真 evidence excerpt 是否能推出 claim,不做新检索 | `verifier_output` |
|
||||
| `chat_composer` | 只表达 Verifier 允许输出的 claims、缺口和建议,生成最终用户答复 | `composer_output` |
|
||||
|
||||
Chat 链路最多支持两轮验证:
|
||||
|
||||
@@ -90,22 +103,28 @@ sequenceDiagram
|
||||
participant P as chat_planner
|
||||
participant E as chat_executor
|
||||
participant T as tools
|
||||
participant G as gatekeeper
|
||||
participant V as chat_verifier
|
||||
participant S as diagnosis_session
|
||||
participant M as chat_composer
|
||||
participant R as diagnosis_run
|
||||
|
||||
C->>P: 原始问题 + history + retry_context
|
||||
P-->>C: planner_plan
|
||||
C->>E: planner_plan + 上下文
|
||||
E->>T: 调用证据工具
|
||||
T-->>E: 证据结果
|
||||
E-->>C: executor_feedback
|
||||
C->>V: executor_final_answer + tool_trace_summary
|
||||
E-->>C: executor_evidence_v2
|
||||
C->>G: executor_structured_output + tool_invocation.evidence_refs
|
||||
G-->>C: gatekeeper_result
|
||||
C->>V: executor_structured_output + gatekeeper_result + tool_trace_summary
|
||||
V-->>C: PASS / LOW_CONFID / REJECT
|
||||
C->>S: 写入 verifier_evaluation
|
||||
C->>R: 写入 verifier_evaluation
|
||||
alt LOW_CONFID 且允许补证据
|
||||
C->>P: retry_context: 仅补缺失证据
|
||||
else PASS 或 REJECT
|
||||
C-->>S: 保存最终 answer
|
||||
C->>M: allowed_claims + missing_info + recommended_actions
|
||||
M-->>C: composer_output
|
||||
C->>R: 保存 Composer 最终 answer
|
||||
end
|
||||
```
|
||||
|
||||
@@ -113,7 +132,7 @@ sequenceDiagram
|
||||
|
||||
| Verdict | 行为 |
|
||||
|---|---|
|
||||
| `PASS` | 输出 Executor 答案 |
|
||||
| `PASS` | 把 Verifier 允许表达的 claims 交给 Composer 输出 |
|
||||
| `LOW_CONFID` | 如果分数低于阈值且仍有轮次,构造 `retry_context` 补证据;否则输出低置信提示 |
|
||||
| `REJECT` | 输出降级答复,只保留已确认信息和下一步建议 |
|
||||
|
||||
@@ -177,21 +196,25 @@ flowchart LR
|
||||
SkillBody --> Executor
|
||||
Executor --> EvidenceTools["lookup_knowledge / logs / metrics"]
|
||||
EvidenceTools --> ToolTrace["tool_invocation evidence"]
|
||||
Executor --> Verifier["Verifier"]
|
||||
Executor --> Gatekeeper["Gatekeeper"]
|
||||
Gatekeeper --> Verifier["Verifier"]
|
||||
ToolTrace --> Verifier
|
||||
Verifier --> Composer["Composer"]
|
||||
```
|
||||
|
||||
| 角色 | Skill 可见性 | 工具权限 |
|
||||
|---|---|---|
|
||||
| Planner | 只看 skill name / description,并输出 `selected_skill` | 不暴露 `read_skill` |
|
||||
| Executor | 读取 Planner 选中的 skill 正文 | 暴露官方 `read_skill` 和证据工具 |
|
||||
| Verifier | 不看 skill catalog,也不读 skill 正文 | 只读取 `tool_trace_summary` |
|
||||
| Gatekeeper | 不看 skill catalog,也不读 skill 正文 | 只读取 Executor 输出和 `tool_invocation.retrieval_details.evidence_refs` |
|
||||
| Verifier | 不看 skill catalog,也不读 skill 正文 | 只读取 Gatekeeper 结果、结构化 claims 和 trace summary |
|
||||
| Composer | 不看 skill catalog,也不读 skill 正文 | 只读取 Verifier 允许表达的内容 |
|
||||
|
||||
## 7. 与旧版设计的差异
|
||||
|
||||
| 旧版设想 | 当前实现 |
|
||||
|---|---|
|
||||
| Supervisor + Planner + 多个专科 SubAgent + Verifier | Chat: Planner + Executor + Verifier;AIOps: Supervisor + Planner + Executor |
|
||||
| Supervisor + Planner + 多个专科 SubAgent + Verifier | Chat: Planner + Executor + Gatekeeper + Verifier + Composer;AIOps: Supervisor + Planner + Executor |
|
||||
| ExternalApiSubAgent / InternalErrorSubAgent / DatabaseSubAgent | 暂未拆分,能力通过通用 Executor + 工具 + Prompt 约束实现 |
|
||||
| 每个 SubAgent 专属工具集 | 当前 Executor 持有统一证据工具集合 |
|
||||
| Verifier 支持 PASS / REVISE / REJECT | 当前 Chat Verifier 输出 PASS / LOW_CONFID / REJECT |
|
||||
|
||||
@@ -154,4 +154,4 @@ ALTER TABLE diagnosis_session ADD COLUMN answer LONGTEXT COMMENT 'Agent 返回
|
||||
|
||||
- **LLM 观点层**:在 `selfEvaluation` 的 `llm_opinion` 字段叠加 LLM 结构化观点(has_root_cause、has_solution 等),作为独立 factors,不改变现有规则逻辑
|
||||
- **案例结构化字段**:useful 触发时自动提取 faultCategory / errorCode,替代暂时的 GENERAL
|
||||
- **重复召回问题**:Executor Prompt 约束或工具层 session 维度去重(见 [ISS-001](../issues/ISS-001-duplicate-retrieval.md))
|
||||
- **重复召回问题**:Executor Prompt 约束或工具层 session 维度去重(见 [ISS-001](../../../issues/archived/ISS-001-duplicate-retrieval.md))
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
# 当前 MVP 架构
|
||||
|
||||
**更新日期**:2026-07-05
|
||||
**状态**:当前可运行架构
|
||||
**更新日期**:2026-07-08
|
||||
**状态**:当前可运行架构
|
||||
**适用范围**:Demo、面试讲解、后续迭代规划
|
||||
|
||||
## 1. 系统定位
|
||||
@@ -38,7 +38,9 @@ flowchart TB
|
||||
Supervisor["Supervisor"]
|
||||
Planner["Planner"]
|
||||
Executor["Executor"]
|
||||
Gatekeeper["Gatekeeper"]
|
||||
Verifier["Verifier"]
|
||||
Composer["Composer"]
|
||||
end
|
||||
|
||||
subgraph Tools["Evidence Tools"]
|
||||
@@ -63,7 +65,8 @@ flowchart TB
|
||||
end
|
||||
|
||||
subgraph Store["Persistence and Trace"]
|
||||
Session["diagnosis_session"]
|
||||
ChatSession["chat_session"]
|
||||
Run["diagnosis_run"]
|
||||
Step["agent_step"]
|
||||
Invocation["tool_invocation"]
|
||||
ApiDoc["api_document"]
|
||||
@@ -105,7 +108,9 @@ Agent Orchestration
|
||||
-> Supervisor
|
||||
-> Planner
|
||||
-> Executor
|
||||
-> Gatekeeper
|
||||
-> Verifier
|
||||
-> Composer
|
||||
|
||||
Evidence Tools
|
||||
-> lookup_knowledge
|
||||
@@ -126,13 +131,15 @@ RAG Retrieval
|
||||
-> Milvus SDK fallback
|
||||
|
||||
Persistence
|
||||
-> diagnosis_session
|
||||
-> agent_step
|
||||
-> tool_invocation
|
||||
-> chat_session
|
||||
-> diagnosis_run
|
||||
-> agent_step.run_id
|
||||
-> tool_invocation.run_id
|
||||
-> api_document
|
||||
-> Milvus/Zilliz collection
|
||||
|
||||
Quality Gates
|
||||
-> executor gatekeeper
|
||||
-> chat verifier
|
||||
-> AIOps rule evaluation
|
||||
-> diagnosis eval baseline
|
||||
@@ -150,23 +157,30 @@ sequenceDiagram
|
||||
participant Planner as Planner Agent
|
||||
participant Executor as Executor Agent
|
||||
participant Tool as Evidence Tools
|
||||
participant Gatekeeper as Gatekeeper Hook
|
||||
participant Verifier as Verifier Agent
|
||||
participant Composer as Composer Agent
|
||||
participant DB as Trace Tables
|
||||
participant Trace as Trace API
|
||||
|
||||
User->>API: 提交诊断问题
|
||||
API->>Chat: execute chat strategy
|
||||
Chat->>DB: 创建 chat_session metadata + diagnosis_run(runId)
|
||||
Chat->>Planner: 复杂问题进入规划
|
||||
Planner->>DB: 写入 agent_step
|
||||
Planner->>DB: 写入 agent_step.run_id
|
||||
Planner->>Executor: 下发排查方向
|
||||
Executor->>Tool: lookup_knowledge / logs / metrics
|
||||
Tool->>DB: 写入 tool_invocation
|
||||
Tool->>DB: 写入 tool_invocation.run_id
|
||||
Tool-->>Executor: 返回证据
|
||||
Executor->>Verifier: 生成候选诊断并校验
|
||||
Verifier->>DB: 合并 self_evaluation.verifier_evaluation
|
||||
Chat->>DB: 保存 diagnosis_session.answer
|
||||
User->>Trace: GET /api/diagnosis/{sessionId}/trace
|
||||
Trace->>DB: 聚合 session / step / tool
|
||||
Executor->>Gatekeeper: 输出 executor_evidence_v2
|
||||
Gatekeeper->>DB: 读取 tool_invocation.evidence_refs 并校验引用
|
||||
Gatekeeper->>Verifier: 传入已验真的 claims / excerpts
|
||||
Verifier->>DB: 合并 diagnosis_run.self_evaluation.verifier_evaluation
|
||||
Verifier->>Composer: 传入 allowed_claims / missing_info / actions
|
||||
Composer->>Chat: 生成最终用户答复
|
||||
Chat->>DB: 保存 diagnosis_run.answer
|
||||
User->>Trace: GET /api/diagnosis/{sessionId}/trace?runId=...
|
||||
Trace->>DB: 聚合 run / step / tool
|
||||
Trace-->>User: 返回可回放诊断链路
|
||||
```
|
||||
|
||||
@@ -180,14 +194,17 @@ POST /api/chat
|
||||
-> lookup_knowledge
|
||||
-> query_logs
|
||||
-> query_metrics
|
||||
-> Verifier 校验最终诊断
|
||||
-> 保存 diagnosis_session
|
||||
-> 保存 agent_step
|
||||
-> 保存 tool_invocation
|
||||
-> 合并 self_evaluation.verifier_evaluation
|
||||
-> Gatekeeper 校验 Executor 证据引用真实性
|
||||
-> Verifier 判断 claim 是否能由已核验证据推出
|
||||
-> Composer 生成最终用户答复
|
||||
-> 保存 chat_session metadata
|
||||
-> 保存 diagnosis_run
|
||||
-> 保存 agent_step.run_id
|
||||
-> 保存 tool_invocation.run_id
|
||||
-> 合并 diagnosis_run.self_evaluation.verifier_evaluation
|
||||
```
|
||||
|
||||
Chat 链路的质量门禁是 LLM Verifier。Verifier 输出合并到 `diagnosis_session.self_evaluation.verifier_evaluation`,Trace API 会展示该验证结果。
|
||||
Chat 链路的质量门禁由三段组成:Gatekeeper 先做代码级引用验真,Verifier 再做 LLM 可推导性判断,Composer 最后控制对用户的表达边界。Gatekeeper、Verifier、Composer 的输出合并到当前 `diagnosis_run.self_evaluation.verifier_evaluation`,Trace API 会展示该验证结果。
|
||||
|
||||
Agent 编排细节见 [agent-orchestration.md](agent-orchestration.md)。
|
||||
|
||||
@@ -242,7 +259,7 @@ POST /api/ai_ops
|
||||
-> Prometheus / logs / knowledge tools
|
||||
-> 生成告警分析报告
|
||||
-> AiOpsRuleEvaluationService
|
||||
-> 合并 self_evaluation.aiops_rule_evaluation
|
||||
-> 合并 diagnosis_run.self_evaluation.aiops_rule_evaluation
|
||||
-> Trace API 可查看全链路
|
||||
```
|
||||
|
||||
@@ -304,32 +321,41 @@ RAG 总体设计见 [rag-architecture.md](rag-architecture.md),检索运行细
|
||||
|
||||
## 6. 持久化模型
|
||||
|
||||
当前诊断持久化以三张表为核心:
|
||||
当前诊断持久化以 session/run/trace 明细为核心:
|
||||
|
||||
```text
|
||||
diagnosis_session
|
||||
-> 一次诊断会话的主记录
|
||||
chat_session
|
||||
-> 多轮会话目录和元数据
|
||||
-> session_id / status / message_pair_count
|
||||
|
||||
diagnosis_run
|
||||
-> 一次诊断运行的主记录
|
||||
-> run_id / session_id
|
||||
-> query / status / agent_flow / answer
|
||||
-> self_evaluation
|
||||
-> step_count / tool_call_count / duration
|
||||
|
||||
agent_step
|
||||
-> Agent 模型调用步骤
|
||||
-> session_id / run_id
|
||||
-> step_index / agent_name
|
||||
-> model_input / model_output / thought
|
||||
-> duration / token_count
|
||||
|
||||
tool_invocation
|
||||
-> 工具调用事实
|
||||
-> session_id / run_id
|
||||
-> tool_name / input_params / output_preview
|
||||
-> retrieval_layer / retrieval_details
|
||||
-> retrieval_details.evidence_refs
|
||||
-> relevance_level / dedup_reason
|
||||
-> duration / success
|
||||
```
|
||||
|
||||
说明:
|
||||
|
||||
- 旧的 `diagnosis_record` 已不是当前主模型,迁移脚本中已经由 `diagnosis_session + agent_step + tool_invocation` 取代。
|
||||
- 旧的 `diagnosis_record` 已不是当前主模型。
|
||||
- `diagnosis_session` 已降级为历史兼容和回滚表,新执行写入 `chat_session + diagnosis_run`。
|
||||
- `api_document` 仍用于文档元数据管理。
|
||||
- 文档向量内容存放在 Milvus/Zilliz collection 中。
|
||||
|
||||
@@ -339,19 +365,20 @@ tool_invocation
|
||||
|
||||
```text
|
||||
GET /api/diagnosis/{sessionId}/trace
|
||||
GET /api/diagnosis/{sessionId}/trace?runId=run-...
|
||||
```
|
||||
|
||||
Trace API 聚合:
|
||||
|
||||
- 会话状态和最终报告。
|
||||
- 会话元数据、运行状态和最终报告。
|
||||
- Agent step 序列。
|
||||
- 工具调用和检索细节。
|
||||
- Chat verifier 结果。
|
||||
- Chat Gatekeeper / Verifier / Composer 结果。
|
||||
- AIOps rule evaluation 结果。
|
||||
|
||||
Trace 是本项目区别于普通问答系统的关键:答案不是孤立文本,而是可以追溯到 Agent 决策、工具调用和证据来源。
|
||||
|
||||
Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gates.md](harness-quality-gates.md),用户反馈与 `self_evaluation` 闭环见 [feedback-architecture.md](feedback-architecture.md)。
|
||||
Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明见 [harness-quality-gates.md](harness-quality-gates.md),用户反馈与 `self_evaluation` 闭环见 [feedback-architecture.md](feedback-architecture.md)。
|
||||
|
||||
## 8. 质量门禁
|
||||
|
||||
@@ -359,7 +386,9 @@ Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gate
|
||||
|
||||
| 门禁 | 位置 | 作用 |
|
||||
|---|---|---|
|
||||
| Chat Verifier | `ChatService` | 校验普通诊断回答质量 |
|
||||
| Executor Gatekeeper | `VerifierInputHook` / `ExecutorGatekeeperService` | 校验 Executor 引用的 invocation、`raw_path`、`evidence_excerpt` 是否真实 |
|
||||
| Chat Verifier | `ChatService` | 判断已验真证据是否能推出 Executor claims |
|
||||
| Chat Composer | `ChatService` | 只表达 Verifier 允许输出的内容,避免把 no-evidence 说成已排除 |
|
||||
| AIOps Rule Evaluation | `AiOpsRuleEvaluationService` | 校验告警诊断是否聚焦 payload 并使用证据 |
|
||||
| Diagnosis Eval Baseline | `mvp/eval/` | 固化诊断 trace 和报告行为 |
|
||||
| RAG Retrieval Baseline | `eval/rag-retrieval/` | 固化检索召回行为,避免 RAG 重构回退 |
|
||||
@@ -379,6 +408,11 @@ Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gate
|
||||
- `title`、`breadcrumb`、`content` 参与 embedding 文本。
|
||||
- `tool_invocation` 记录检索层、relevance level、dedup reason。
|
||||
- Chat verifier 和 AIOps rule evaluation 合并进 `self_evaluation`。
|
||||
- Chat Executor 结构化输出 `executor_evidence_v2`,不再直接承担最终用户答复。
|
||||
- `tool_invocation.retrieval_details.evidence_refs` 支持 `raw_path` 精确引用和 `$.no_evidence` 负向证据。
|
||||
- Gatekeeper 对 Executor 引用做代码级验真,并在审计中记录 `rule_set_version` 和规则元数据摘要。
|
||||
- Verifier 只判断可推导性。
|
||||
- Composer 在 Verifier 之后生成最终用户表达,并限制 negative observation 过度表述。
|
||||
- RAG offline baseline 和 live acceptance 脚本。
|
||||
|
||||
暂不作为当前已完成能力声明:
|
||||
@@ -406,4 +440,5 @@ Prompt、Hook、Verifier 和评测门禁的完整说明见 [harness-quality-gate
|
||||
| Spring AI VectorStore 配置辅助 | `SpringAiVectorStoreSidecarService` |
|
||||
| Trace 聚合 | `DiagnosisTraceService` |
|
||||
| 工具调用记录 | `ToolInvocationRecorder` |
|
||||
| Executor 引用验真 | `ExecutorGatekeeperService`, `VerifierInputHook` |
|
||||
| self_evaluation 合并 | `SelfEvaluationMergeService` |
|
||||
|
||||
+81
-151
@@ -1,31 +1,44 @@
|
||||
# 数据模型总览
|
||||
|
||||
**更新日期**:2026-07-05
|
||||
**更新日期**:2026-07-10
|
||||
**状态**:当前可运行架构
|
||||
|
||||
## 1. 定位
|
||||
|
||||
本文从架构角度说明当前 MVP 的核心数据模型。详细字段仍以 Flyway migration 和 `mvp/tables/` 为准。
|
||||
本文从架构角度说明当前 MVP 的核心数据模型。详细字段以 Flyway migration、实体类和 `mvp/tables/` 为准。
|
||||
|
||||
核心数据分三组:
|
||||
|
||||
- 诊断 Trace:`diagnosis_session`、`agent_step`、`tool_invocation`
|
||||
- 会话与诊断 Trace:`chat_session`、`diagnosis_run`、`agent_step`、`tool_invocation`
|
||||
- 知识库:`api_document`、`knowledge_domain`、Milvus/Zilliz metadata
|
||||
- 反馈沉淀:`case_library`
|
||||
|
||||
`diagnosis_session` 仍保留为历史兼容和回滚表,不再是新执行写入的主模型。
|
||||
|
||||
## 2. 总体关系
|
||||
|
||||
```mermaid
|
||||
erDiagram
|
||||
diagnosis_session ||--o{ agent_step : has
|
||||
diagnosis_session ||--o{ tool_invocation : has
|
||||
diagnosis_session ||--o| case_library : creates_when_useful
|
||||
chat_session ||--o{ diagnosis_run : owns
|
||||
diagnosis_run ||--o{ agent_step : has
|
||||
diagnosis_run ||--o{ tool_invocation : has
|
||||
diagnosis_run ||--o| case_library : creates_when_useful
|
||||
api_document ||--o{ milvus_chunk : indexed_as
|
||||
knowledge_domain ||--o{ api_document : groups
|
||||
|
||||
diagnosis_session {
|
||||
chat_session {
|
||||
bigint id
|
||||
varchar session_id
|
||||
varchar status
|
||||
int message_pair_count
|
||||
datetime last_active_at
|
||||
datetime expires_at
|
||||
}
|
||||
|
||||
diagnosis_run {
|
||||
bigint id
|
||||
varchar run_id
|
||||
varchar session_id
|
||||
text query
|
||||
varchar status
|
||||
varchar agent_flow
|
||||
@@ -37,42 +50,23 @@ erDiagram
|
||||
agent_step {
|
||||
bigint id
|
||||
varchar session_id
|
||||
varchar run_id
|
||||
int step_index
|
||||
varchar agent_name
|
||||
text model_input
|
||||
text model_output
|
||||
text thought
|
||||
boolean has_tool_call
|
||||
}
|
||||
|
||||
tool_invocation {
|
||||
bigint id
|
||||
varchar session_id
|
||||
varchar run_id
|
||||
bigint step_id
|
||||
varchar tool_name
|
||||
json input_params
|
||||
text output_preview
|
||||
varchar retrieval_layer
|
||||
json retrieval_details
|
||||
varchar relevance_level
|
||||
varchar dedup_reason
|
||||
}
|
||||
|
||||
api_document {
|
||||
bigint id
|
||||
varchar doc_id
|
||||
varchar file_name
|
||||
varchar file_path
|
||||
varchar status
|
||||
int chunk_count
|
||||
text metadata
|
||||
}
|
||||
|
||||
knowledge_domain {
|
||||
bigint id
|
||||
varchar domain_id
|
||||
varchar description
|
||||
text when_to_retrieve
|
||||
int document_count
|
||||
}
|
||||
|
||||
case_library {
|
||||
@@ -84,123 +78,87 @@ erDiagram
|
||||
text root_cause
|
||||
text solution
|
||||
}
|
||||
|
||||
milvus_chunk {
|
||||
varchar id
|
||||
text content
|
||||
json metadata
|
||||
vector vector
|
||||
}
|
||||
```
|
||||
|
||||
说明:Milvus/Zilliz collection 不是 MySQL 表,图中的 `milvus_chunk` 是逻辑模型。
|
||||
|
||||
## 3. 诊断 Trace 模型
|
||||
## 3. 会话与运行模型
|
||||
|
||||
### diagnosis_session
|
||||
### chat_session
|
||||
|
||||
会话级主记录。
|
||||
|
||||
关键字段:
|
||||
`chat_session` 是会话目录表,保存 `sessionId` 的元数据:
|
||||
|
||||
| 字段 | 说明 |
|
||||
|---|---|
|
||||
| `session_id` | 外部关联键,Trace 和 Feedback 都使用它 |
|
||||
| `query` | 用户原始问题或 AIOps 输入摘要 |
|
||||
| `status` | 执行状态 |
|
||||
| `session_id` | 外部会话 ID,用于多轮上下文和 run 列表 |
|
||||
| `status` | 会话目录状态 |
|
||||
| `message_pair_count` | Redis 对话轮次数快照 |
|
||||
| `last_active_at` | 最近活跃时间 |
|
||||
| `expires_at` | 可为空的目录 TTL 元数据 |
|
||||
|
||||
它不保存完整对话历史,正文消息仍由 Redis `SessionContext.messageHistory` 管理。
|
||||
|
||||
### diagnosis_run
|
||||
|
||||
`diagnosis_run` 是一次可回放诊断执行的主记录:
|
||||
|
||||
| 字段 | 说明 |
|
||||
|---|---|
|
||||
| `run_id` | 运行 ID,格式为 `run-` + UUID |
|
||||
| `session_id` | 所属 `chat_session.session_id` |
|
||||
| `query` | 本次 Chat 问题或 AIOps 告警摘要 |
|
||||
| `status` | 本次执行状态 |
|
||||
| `agent_flow` | `CHAT` / `AI_OPS` |
|
||||
| `answer` | 最终答复或告警报告 |
|
||||
| `self_evaluation` | rule/verifier/aiops 自评估容器 |
|
||||
| `feedback` | 用户反馈 |
|
||||
| `answer` | 本次运行最终答复或告警报告 |
|
||||
| `self_evaluation` | 本次运行的 rule/verifier/aiops 自评估容器 |
|
||||
| `feedback` | 本次运行的用户反馈 |
|
||||
|
||||
同一个 `sessionId` 可以有多个 `runId`。Trace、反馈、评测和案例沉淀都应优先使用 `runId`,避免多轮同 session 下的数据混合。
|
||||
|
||||
## 4. Trace 明细模型
|
||||
|
||||
### agent_step
|
||||
|
||||
记录模型调用步骤。
|
||||
|
||||
用途:
|
||||
|
||||
- 回放 Agent 推理过程。
|
||||
- 查看 Planner / Executor / Verifier 的输入输出摘要。
|
||||
- 统计 step count、duration、token count。
|
||||
`agent_step` 记录模型调用步骤。新写入同时保留 `session_id` 和 `run_id`,其中 `run_id` 是回放边界。Trace 页面和评测应先按 `run_id` 隔离取数,展示顺序以 Trace API 返回顺序为准。
|
||||
|
||||
### tool_invocation
|
||||
|
||||
记录工具调用事实。
|
||||
`tool_invocation` 记录显式工具调用事实。`retrieval_details.evidence_refs` 是 Chat 证据链路的关键字段:
|
||||
|
||||
用途:
|
||||
|
||||
- 给 Trace API 展示证据。
|
||||
- 给 Verifier 构造 `tool_trace_summary`。
|
||||
- 给 `EvaluationService` 计算 evidence score。
|
||||
- 给 RAG eval 和人工排查提供检索细节。
|
||||
|
||||
## 4. 知识库模型
|
||||
|
||||
### api_document
|
||||
|
||||
MySQL 中的文档元数据表。
|
||||
|
||||
职责:
|
||||
|
||||
- 管理上传文件。
|
||||
- 保存 file hash,用于去重。
|
||||
- 记录索引状态和 chunk 数量。
|
||||
- 保存 frontmatter JSON。
|
||||
|
||||
### knowledge_domain
|
||||
|
||||
领域级元数据。
|
||||
|
||||
职责:
|
||||
|
||||
- 按 category 聚合文档。
|
||||
- 存储领域描述。
|
||||
- 存储 `when_to_retrieve`,辅助 Planner/Executor 判断什么时候检索该领域。
|
||||
|
||||
### Milvus/Zilliz metadata
|
||||
|
||||
向量 collection 中每个 chunk 的 metadata 主要包括:
|
||||
|
||||
```text
|
||||
docId
|
||||
_source
|
||||
chunkIndex
|
||||
totalChunks
|
||||
title
|
||||
breadcrumb
|
||||
category
|
||||
```json
|
||||
{
|
||||
"evidence_status": "supported",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.logs[0]",
|
||||
"text": "2026-07-08 23:05:28 ERROR order-service HikariPool-1 - Connection is not available..."
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
这些字段支撑:
|
||||
|
||||
- category filter。
|
||||
- source 展示。
|
||||
- breadcrumb 上下文。
|
||||
- docId 删除和重建索引。
|
||||
- evidence block 构造。
|
||||
`$.no_evidence` 只表示“本次工具查询未检索到匹配证据”,不能被解释为“问题不存在”或“根因已排除”。
|
||||
|
||||
## 5. 反馈沉淀模型
|
||||
|
||||
### case_library
|
||||
|
||||
`useful` 反馈会触发 `CaseLibraryService.createFromSession`。
|
||||
`useful` 反馈会触发 `CaseLibraryService.createFromRun`。
|
||||
|
||||
当前自动映射:
|
||||
|
||||
| 字段 | 来源 |
|
||||
|---|---|
|
||||
| `case_id` | UUID |
|
||||
| `diagnosis_id` | `diagnosis_session.session_id` |
|
||||
| `diagnosis_id` | 新数据为 `diagnosis_run.run_id`;历史数据可能为 `diagnosis_session.session_id` |
|
||||
| `source_type` | `AUTO` |
|
||||
| `fault_category` | 当前默认 `GENERAL` |
|
||||
| `title` | session query 前 100 字符 |
|
||||
| `root_cause` | session answer |
|
||||
| `solution` | session answer |
|
||||
| `title` | run query 前 100 字符 |
|
||||
| `root_cause` | run answer |
|
||||
| `solution` | run answer |
|
||||
| `created_by` | `system` |
|
||||
|
||||
## 6. self_evaluation 结构
|
||||
|
||||
`diagnosis_session.self_evaluation` 是 JSON 容器:
|
||||
`diagnosis_run.self_evaluation` 是运行级 JSON 容器:
|
||||
|
||||
```json
|
||||
{
|
||||
@@ -210,49 +168,21 @@ category
|
||||
}
|
||||
```
|
||||
|
||||
边界:
|
||||
Chat 通常写入 `rule_evaluation` 和 `verifier_evaluation`;AIOps 写入 `aiops_rule_evaluation`。
|
||||
|
||||
- `rule_evaluation` 评估证据收集充分度。
|
||||
- `verifier_evaluation` 评估 Chat 答案关键事实是否有证据支撑。
|
||||
- `aiops_rule_evaluation` 评估 AIOps 报告是否聚焦告警并使用证据。
|
||||
|
||||
## 7. 数据写入时序
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant API as API
|
||||
participant Svc as ChatService/AiOpsService
|
||||
participant Session as diagnosis_session
|
||||
participant Agent as Agent
|
||||
participant Step as agent_step
|
||||
participant Tool as tool_invocation
|
||||
participant Eval as self_evaluation
|
||||
participant Feedback as case_library
|
||||
|
||||
API->>Svc: request
|
||||
Svc->>Session: create/update RUNNING
|
||||
Agent->>Step: before/after model
|
||||
Agent->>Tool: tool call record
|
||||
Svc->>Session: SUCCESS/FAILED + answer
|
||||
Svc->>Eval: merge evaluation
|
||||
API->>Svc: feedback useful
|
||||
Svc->>Feedback: create case
|
||||
```
|
||||
|
||||
## 8. 当前边界和后续
|
||||
## 7. 当前边界和后续
|
||||
|
||||
当前边界:
|
||||
|
||||
- `agent_step.session_id` 和 `tool_invocation.session_id` 通过 sessionId 关联,不强制外键。
|
||||
- `tool_invocation.step_id` 可为空。
|
||||
- Milvus chunk 与 `api_document` 通过 metadata.docId 逻辑关联。
|
||||
- `case_library` 与 session 通过 `diagnosis_id=session_id` 关联。
|
||||
- `chat_session` 只存会话元数据,不存完整正文历史。
|
||||
- `diagnosis_run` 存一次运行的长期审计状态。
|
||||
- `agent_step.run_id` 和 `tool_invocation.run_id` 是 Trace、Verifier、Eval 的运行边界。
|
||||
- 当前实现主要使用逻辑关联,不依赖数据库外键。
|
||||
- `case_library.diagnosis_id` 是过渡字段,新值按 `run_id` 解释,旧值可能按 `session_id` 解释。
|
||||
- `diagnosis_session` 只作为历史兼容和回滚表保留。
|
||||
|
||||
后续可增强:
|
||||
|
||||
1. 增加 run id,支持同 session 多次独立诊断。
|
||||
2. 强化 `tool_invocation.step_id` 关联。
|
||||
3. 将 evidence block 结构化保存。
|
||||
4. 将 `case_library` 的 rootCause/solution 从完整 answer 中结构化抽取。
|
||||
|
||||
1. 强化 `tool_invocation.step_id` 关联。
|
||||
2. 将 Gatekeeper 规则配置化时的规则元数据保存为可审计版本。
|
||||
3. 将 `case_library` 的 rootCause/solution 从完整 answer 中结构化抽取。
|
||||
|
||||
@@ -1,78 +1,44 @@
|
||||
# Current Chat Agent Data Contracts
|
||||
# Chat Evidence Pipeline Contracts
|
||||
|
||||
**状态**:当前实现
|
||||
**日期**:2026-07-07
|
||||
**范围**:当前 Chat 复杂诊断链路的数据结构定义
|
||||
**状态**:当前实现
|
||||
**更新日期**:2026-07-08
|
||||
**范围**:Chat 复杂诊断链路中的 Planner、Executor、Gatekeeper、Verifier、Composer 数据契约
|
||||
|
||||
当前代码实现是三 Agent 顺序链路:
|
||||
当前 Chat 复杂诊断链路是:
|
||||
|
||||
```text
|
||||
chat_planner -> chat_executor -> chat_verifier
|
||||
chat_planner
|
||||
-> chat_executor
|
||||
-> VerifierInputHook / ExecutorGatekeeperService
|
||||
-> chat_verifier
|
||||
-> chat_composer
|
||||
-> final answer
|
||||
```
|
||||
|
||||
对应 `ChatService.executeChatComplex(...)` 中的 `SequentialAgent`。
|
||||
设计原则:
|
||||
|
||||
- Planner 暂不输出 `scope_contract`。
|
||||
- Executor 只做证据收集和微观事实提炼,不生成最终用户答案。
|
||||
- Gatekeeper 在 Verifier 前做代码级引用真实性校验。
|
||||
- Verifier 判断 claim 是否能由已核验证据推出。
|
||||
- Composer 只表达 Verifier 允许输出的内容。
|
||||
|
||||
---
|
||||
|
||||
## 1. Workflow Input
|
||||
## 1. Planner
|
||||
|
||||
由 `ChatService.buildWorkflowInput(...)` 构造,传给 `chat_workflow`。
|
||||
|
||||
```text
|
||||
请按固定工作流完成本轮 Planner -> Executor -> Verifier。
|
||||
|
||||
--- 用户问题 ---
|
||||
{question}
|
||||
|
||||
--- retry_context ---
|
||||
{retry_context}
|
||||
|
||||
Verifier 完成后由外层代码读取 verifier_output 并决定最终用户输出。
|
||||
```
|
||||
|
||||
| 字段 | 来源 | 定义 |
|
||||
|---|---|---|
|
||||
| `question` | 用户输入 | 用户本轮原始问题 |
|
||||
| `retry_context` | ChatService | 第二轮补证据约束;首轮为空 |
|
||||
|
||||
---
|
||||
|
||||
## 2. chat_planner
|
||||
|
||||
### 2.1 Input
|
||||
|
||||
`chat_planner` 的输入来自 workflow input 和 system prompt 追加上下文。
|
||||
Planner 当前保持不变,输出 `planner_plan`:
|
||||
|
||||
```json
|
||||
{
|
||||
"question": "用户原始问题",
|
||||
"history": [],
|
||||
"available_knowledge_domains": "...",
|
||||
"skill_catalog": {},
|
||||
"retry_context": null
|
||||
"selected_skill": "diagnose-mysql-connection-pool",
|
||||
"selection_reason": "选择该 skill 的原因",
|
||||
"plan": ["步骤1", "步骤2"],
|
||||
"reasoning": "规划思路"
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 来源 | 定义 |
|
||||
|---|---|---|
|
||||
| `question` | workflow input | 用户原始问题 |
|
||||
| `history` | `ChatService.buildChatPlannerAgent(...)` | 对话历史,拼接到 planner system prompt |
|
||||
| `available_knowledge_domains` | `KnowledgeDomainService.buildKnowledgeMap()` | 可用知识域地图,拼接到 planner system prompt |
|
||||
| `skill_catalog` | `PlannerSkillMetadataHook` | Planner 可见的 skill name/description 元数据 |
|
||||
| `retry_context` | `ChatService` | Verifier 低置信后构造的补证据上下文 |
|
||||
|
||||
### 2.2 Output:`planner_plan`
|
||||
|
||||
当前 prompt 要求输出 JSON:
|
||||
|
||||
```json
|
||||
{
|
||||
"selected_skill": "匹配的 skill 名称;如果没有匹配则为 null",
|
||||
"selection_reason": "选择该 skill 的原因;如果没有匹配则说明不使用 skill",
|
||||
"plan": ["步骤1描述", "步骤2描述", "步骤3描述"],
|
||||
"reasoning": "规划思路说明"
|
||||
}
|
||||
```
|
||||
字段定义:
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
@@ -81,434 +47,397 @@ Verifier 完成后由外层代码读取 verifier_output 并决定最终用户输
|
||||
| `plan` | array | 给 Executor 的执行步骤 |
|
||||
| `reasoning` | string | 规划思路说明 |
|
||||
|
||||
运行态输出 key:
|
||||
当前边界:
|
||||
|
||||
```text
|
||||
planner_plan
|
||||
```
|
||||
- 不新增 `scope_contract`。
|
||||
- 不要求 Planner 显式列出 forbidden actions。
|
||||
- 窄范围控制先由 Executor Prompt 约束,后续如仍不稳定再引入 Planner contract。
|
||||
|
||||
---
|
||||
|
||||
## 3. chat_executor
|
||||
## 2. Executor
|
||||
|
||||
### 3.1 Input
|
||||
Executor 输出 `executor_evidence_v2`。它不是最终答复,而是给 Gatekeeper、Verifier、Composer 使用的结构化诊断材料。
|
||||
|
||||
`chat_executor` 接收前序 `planner_plan`,并通过 system prompt 获得历史、skill 读取约束、retry 约束和工具权限。
|
||||
### 2.1 输出结构
|
||||
|
||||
```json
|
||||
{
|
||||
"planner_plan": {},
|
||||
"history": [],
|
||||
"retry_context": null,
|
||||
"tool_permissions": {
|
||||
"method_tools": ["dateTimeTools", "lookupKnowledgeTool", "queryMetricsTools", "queryLogsTools"],
|
||||
"tool_callbacks": []
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 来源 | 定义 |
|
||||
|---|---|---|
|
||||
| `planner_plan` | `chat_planner` | Planner 输出的计划 |
|
||||
| `history` | `ChatService.buildChatExecutorAgent(...)` | 对话历史,拼接到 executor system prompt |
|
||||
| `retry_context` | `ChatService` | 本轮补证据约束 |
|
||||
| `method_tools` | `ChatService.buildMethodToolsArray()` | Executor 可直接调用的本地工具 |
|
||||
| `tool_callbacks` | `ToolCallback[]` | 框架发现或外部注入工具 |
|
||||
| `read_skill` | `SkillsAgentHook` | 当存在 skillRegistry 时,Executor 可读取 Planner 选中的 skill |
|
||||
|
||||
### 3.2 Output:`executor_feedback`
|
||||
|
||||
当前 `chat-executor-prompt.md` 要求输出一个 JSON 对象,即 `executor_evidence_v1`。
|
||||
|
||||
```json
|
||||
{
|
||||
"answer_version": "executor_evidence_v1",
|
||||
"diagnosis_summary": "1-2句话总结,仅包含有证据支撑的事实和证据边界",
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "root_cause",
|
||||
"claim_text": "事实断言或有限结论",
|
||||
"claim_type": "observation",
|
||||
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"source_type": "tool_trace",
|
||||
"source_id": "工具返回中的 evidence block id、trace_ref 或可定位标识",
|
||||
"tool_name": "lookup_knowledge/query_logs/query_metrics/read_skill 等",
|
||||
"source_invocation_ids": [],
|
||||
"evidence_excerpt": "从工具返回中摘取的原话、指标值、日志片段或关键数据"
|
||||
"source_id": "",
|
||||
"tool_name": "query_metrics",
|
||||
"source_invocation_id": 517,
|
||||
"raw_path": "$.alerts[0]",
|
||||
"evidence_excerpt": "HighCPUUsage, service=payment-service, state=firing, current=92%, duration=25m"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [
|
||||
{
|
||||
"hypothesis_text": "未被证实但值得排查的方向",
|
||||
"basis": "它基于哪些已知证据或为什么只是推测",
|
||||
"needed_evidence": ["需要补充的证据"]
|
||||
}
|
||||
],
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "建议动作",
|
||||
"reason": "为什么建议做这个动作",
|
||||
"evidence_bindings": []
|
||||
}
|
||||
],
|
||||
"missing_info": [
|
||||
"导致无法确认完整根因的证据缺口"
|
||||
],
|
||||
"user_facing_answer": "面向用户的中文回答。必须与 claims/hypotheses/recommended_actions/missing_info 一致。"
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [],
|
||||
"missing_info": []
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
字段定义:
|
||||
|
||||
| 字段 | 类型 | 必填 | 定义 |
|
||||
|---|---|---:|---|
|
||||
| `answer_version` | string | 是 | 固定为 `executor_evidence_v2` |
|
||||
| `claims` | array | 是 | Executor 提出的待验证事实断言 |
|
||||
| `claims[].claim_id` | string | 是 | claim 标识 |
|
||||
| `claims[].claim_type` | string | 是 | `observation`、`negative_observation`、`symptom`、`root_cause` 等;窄范围任务只允许前两者 |
|
||||
| `claims[].claim_text` | string | 是 | 事实断言文本 |
|
||||
| `claims[].support_level` | string | 是 | `direct` 或 `indirect` |
|
||||
| `claims[].evidence_bindings` | array | 是 | 支撑该 claim 的证据绑定,不能为空 |
|
||||
| `evidence_bindings[].source_type` | string | 否 | 当前通常为 `tool_trace` |
|
||||
| `evidence_bindings[].source_id` | string | 否 | 兼容字段,不作为精确引用主键 |
|
||||
| `evidence_bindings[].tool_name` | string | 是 | `query_logs`、`query_metrics`、`lookup_knowledge` 等 |
|
||||
| `evidence_bindings[].source_invocation_id` | number/null | 是 | 来源 `tool_invocation.id`;缺失时 Gatekeeper 只在能唯一匹配时回填 |
|
||||
| `evidence_bindings[].raw_path` | string | 是 | 工具返回中的稳定定位路径 |
|
||||
| `evidence_bindings[].evidence_excerpt` | string | 是 | 工具返回中的原文片段或系统抽取的最小证据文本 |
|
||||
| `hypotheses` | array | 是 | 未证实但值得排查的方向,不是 confirmed fact |
|
||||
| `recommended_actions` | array | 是 | 下一步动作;本期只允许证据收集或继续排查动作 |
|
||||
| `missing_info` | array | 是 | 无法确认结论所缺少的证据 |
|
||||
|
||||
禁止字段:
|
||||
|
||||
- `diagnosis_summary`
|
||||
- `user_facing_answer`
|
||||
- `source_invocation_ids` 作为主引用字段
|
||||
|
||||
### 2.2 raw_path
|
||||
|
||||
当前支持的精确路径:
|
||||
|
||||
| 工具 | 正向证据路径 | 负向证据路径 |
|
||||
|---|---|---|
|
||||
| `answer_version` | string | 当前固定为 `executor_evidence_v1` |
|
||||
| `diagnosis_summary` | string | 有证据边界的简短诊断摘要 |
|
||||
| `claims` | array | 已证实或有明确间接支撑的事实断言 |
|
||||
| `claims[].claim_id` | string | claim 标识 |
|
||||
| `claims[].claim_type` | string | claim 类型,例如 `root_cause`、`symptom`、`impact` |
|
||||
| `claims[].claim_text` | string | 事实断言文本 |
|
||||
| `claims[].support_level` | string | `direct` 或 `indirect` |
|
||||
| `claims[].evidence_bindings` | array | 支撑 claim 的证据绑定,不能为空 |
|
||||
| `evidence_bindings[].source_type` | string | 证据来源类型,例如 `tool_trace` |
|
||||
| `evidence_bindings[].source_id` | string | evidence block id、trace_ref 或其它定位标识 |
|
||||
| `evidence_bindings[].tool_name` | string | 来源工具名 |
|
||||
| `evidence_bindings[].source_invocation_ids` | array | 来源 `tool_invocation.id` |
|
||||
| `evidence_bindings[].evidence_excerpt` | string | 工具返回中的原话、指标值、日志片段或关键数据 |
|
||||
| `hypotheses` | array | 未证实但值得排查的方向 |
|
||||
| `hypotheses[].hypothesis_text` | string | 假设文本 |
|
||||
| `hypotheses[].basis` | string | 假设依据和未证实原因 |
|
||||
| `hypotheses[].needed_evidence` | array | 确认该假设还需要的证据 |
|
||||
| `recommended_actions` | array | 建议动作 |
|
||||
| `recommended_actions[].action_text` | string | 建议动作文本 |
|
||||
| `recommended_actions[].reason` | string | 建议原因 |
|
||||
| `recommended_actions[].evidence_bindings` | array | 建议动作关联证据,可为空 |
|
||||
| `missing_info` | array | 证据缺口 |
|
||||
| `user_facing_answer` | string | 候选用户答案,PASS 时由 ChatService 提取输出 |
|
||||
| `query_metrics` | `$.alerts[i]` | `$.no_evidence` |
|
||||
| `query_logs` | `$.logs[i]` | `$.no_evidence` |
|
||||
| `lookup_knowledge` | `$.evidence_blocks[i]` | `$.no_evidence` |
|
||||
|
||||
运行态输出 key:
|
||||
约束:
|
||||
|
||||
```text
|
||||
executor_feedback
|
||||
- `raw_path` 必须指向数组条目或 `$.no_evidence`。
|
||||
- 禁止字段级子路径,例如 `$.alerts[0].state`、`$.logs[0].message`。
|
||||
- 同一条工具数组项只能绑定一次;多个字段应合并进同一个 `evidence_excerpt`。
|
||||
|
||||
### 2.3 negative_observation
|
||||
|
||||
当工具明确返回 no-hit / no-evidence 时,Executor 可以输出 `negative_observation`:
|
||||
|
||||
```json
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "negative_observation",
|
||||
"claim_text": "当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"source_invocation_id": 517,
|
||||
"raw_path": "$.no_evidence",
|
||||
"evidence_excerpt": "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
语义边界:
|
||||
|
||||
- `$.no_evidence` 只表示“该工具对当前查询返回无匹配证据”。
|
||||
- 不表示“问题绝对不存在”。
|
||||
- 不表示“根因被排除”。
|
||||
- 不表示“系统已经健康”。
|
||||
- `negative_observation` 的 `evidence_bindings` 只能绑定 `$.no_evidence`,不能混绑其它服务的正向日志。
|
||||
|
||||
### 2.4 窄范围任务
|
||||
|
||||
窄范围任务指用户只要求确认某个服务、告警、日志、错误、订单或时间窗口。
|
||||
|
||||
Executor 必须遵守:
|
||||
|
||||
- 只输出 `observation` / `negative_observation`。
|
||||
- claim 数量通常 1 条,最多 2 条。
|
||||
- claim 数量限制不限制 `evidence_bindings` 数量。
|
||||
- 不输出根因、风险、修复建议、经验推断。
|
||||
- 不把 Runbook / Skill / 知识库通用知识写成当前环境事实。
|
||||
- 精确查询返回 no-evidence 后,不得放宽关键词、删除服务名或扩大服务范围继续查。
|
||||
|
||||
---
|
||||
|
||||
## 4. chat_verifier
|
||||
## 3. Tool Invocation Evidence Refs
|
||||
|
||||
### 4.1 Input
|
||||
工具调用落库到 `tool_invocation`,其中 `retrieval_details.evidence_refs` 是 Gatekeeper 的主校验源。
|
||||
|
||||
`VerifierInputHook` 会在 Verifier 调用前替换消息历史,构造显式 JSON payload。
|
||||
### 3.1 正向证据
|
||||
|
||||
```json
|
||||
{
|
||||
"evidence_status": "supported",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.logs[0]",
|
||||
"text": "2026-07-08 23:05:28 ERROR order-service HikariPool-1 - Connection is not available..."
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
### 3.2 负向证据
|
||||
|
||||
```json
|
||||
{
|
||||
"evidence_status": "no_evidence",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.no_evidence",
|
||||
"text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
字段定义:
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `evidence_status` | string | `supported`、`no_evidence`、`deduped`、`failed` |
|
||||
| `evidence_refs[].raw_path` | string | 证据在工具返回中的稳定定位符 |
|
||||
| `evidence_refs[].text` | string | 系统抽取的最小证据文本,供 Gatekeeper 和 Verifier 使用 |
|
||||
|
||||
---
|
||||
|
||||
## 4. Gatekeeper
|
||||
|
||||
Gatekeeper 位于 Verifier 前,由 `VerifierInputHook` 触发,负责代码级引用真实性校验。
|
||||
|
||||
### 4.1 输入
|
||||
|
||||
- `sessionId`
|
||||
- `executor_structured_output`
|
||||
- 当前 session 的 `tool_invocation`
|
||||
|
||||
### 4.2 输出
|
||||
|
||||
```json
|
||||
{
|
||||
"status": "pass",
|
||||
"severity": "none",
|
||||
"rule_set_version": "gatekeeper-rules-v1",
|
||||
"rules": [
|
||||
{
|
||||
"id": "evidence.raw_path",
|
||||
"description": "raw_path must exist in retrieval_details.evidence_refs",
|
||||
"enabled": true,
|
||||
"default_severity": "reject"
|
||||
}
|
||||
],
|
||||
"checked_bindings": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"tool_name": "query_logs",
|
||||
"source_invocation_id": 517,
|
||||
"raw_path": "$.no_evidence",
|
||||
"matched_text": "query_logs returned no evidence; ...",
|
||||
"status": "pass"
|
||||
}
|
||||
],
|
||||
"failed_rules": [],
|
||||
"warnings": [],
|
||||
"errors": []
|
||||
}
|
||||
```
|
||||
|
||||
字段定义:
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `status` | string | `pass` 或 `fail` |
|
||||
| `severity` | string | `none`、`low_confid`、`reject` |
|
||||
| `rule_set_version` | string | 当前加载的 Gatekeeper 规则集版本 |
|
||||
| `rules` | array | 已启用规则的轻量元数据摘要 |
|
||||
| `checked_bindings` | array | 每条证据绑定的校验结果 |
|
||||
| `failed_rules` | array | 失败规则 id |
|
||||
| `warnings` | array | 自动回填等非阻断信息 |
|
||||
| `errors` | array | 失败明细 |
|
||||
|
||||
校验规则:
|
||||
|
||||
- `answer_version` 必须是 `executor_evidence_v2`。
|
||||
- 不允许 `diagnosis_summary` / `user_facing_answer`。
|
||||
- 每个 claim 必须有非空 `evidence_bindings`。
|
||||
- `tool_name` 必须和真实 invocation 对齐。
|
||||
- `source_invocation_id` 必须存在;缺失时只在 `tool_name + raw_path + evidence_excerpt` 能唯一匹配真实 invocation 时回填。
|
||||
- `raw_path` 必须存在于 `retrieval_details.evidence_refs`。
|
||||
- `evidence_excerpt` 必须由 `evidence_refs[].text` 支撑。
|
||||
- `negative_observation` 只能绑定 `$.no_evidence`。
|
||||
|
||||
规则配置:
|
||||
|
||||
- 当前规则元数据位于 `src/main/resources/gatekeeper/gatekeeper-rules.json`。
|
||||
- 规则实现仍是确定性 Java 代码,不执行动态脚本。
|
||||
- 当前配置只承载规则 id、描述、默认 severity、启用状态和简单参数,例如 excerpt token overlap 阈值。
|
||||
|
||||
失败分级:
|
||||
|
||||
| 场景 | severity |
|
||||
|---|---|
|
||||
| 伪造 invocation id | `reject` |
|
||||
| tool_name 与 invocation 不匹配 | `reject` |
|
||||
| raw_path 不存在 | `reject` |
|
||||
| excerpt 与 matched_text 不匹配 | `reject` |
|
||||
| negative_observation 绑定正向日志 | `reject` |
|
||||
| 缺少 raw_path / invocation id 且无法唯一回填 | `low_confid` |
|
||||
| 旧 invocation 没有 `evidence_refs` | `low_confid` |
|
||||
|
||||
---
|
||||
|
||||
## 5. Verifier
|
||||
|
||||
Verifier 输入由 `VerifierInputHook` 构造:
|
||||
|
||||
```json
|
||||
{
|
||||
"original_query": "用户原始问题",
|
||||
"executor_final_answer": "{...executor_feedback raw text...}",
|
||||
"executor_structured_output": {},
|
||||
"executor_final_answer": "{...executor raw text for debug/fallback only...}",
|
||||
"executor_structured_output": {
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": []
|
||||
},
|
||||
"executor_output_parse_status": {
|
||||
"status": "valid",
|
||||
"detail": "parsed executor evidence contract"
|
||||
},
|
||||
"tool_trace_summary": [],
|
||||
"gatekeeper_result": {},
|
||||
"retry_context": null
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 来源 | 定义 |
|
||||
|---|---|---|
|
||||
| `original_query` | `VerifierContextHolder` | 用户原始问题 |
|
||||
| `executor_final_answer` | `VerifierContextHolder` 或上一条 AssistantMessage | Executor 原始输出文本 |
|
||||
| `executor_structured_output` | `VerifierInputHook.parseExecutorOutput(...)` | Executor 输出可解析且包含 `claims` 时的 JSON 对象;否则为 null |
|
||||
| `executor_output_parse_status.status` | `VerifierInputHook` | `valid` / `missing` / `malformed` |
|
||||
| `executor_output_parse_status.detail` | `VerifierInputHook` | 解析状态说明 |
|
||||
| `tool_trace_summary` | `ToolTraceSummaryService.buildVerifierTraceSummary(...)` | 基于真实 `tool_invocation` 构建的证据索引 |
|
||||
| `retry_context` | `VerifierContextHolder` | 当前补证据上下文 |
|
||||
Verifier 职责:
|
||||
|
||||
### 4.2 `tool_trace_summary`
|
||||
- 不调用工具。
|
||||
- 不读 skill。
|
||||
- 不逐字核验 excerpt 真伪;这由 Gatekeeper 完成。
|
||||
- 只判断 `claim_text` 是否能由已核验的 `evidence_excerpt` 推出。
|
||||
- 结构化输出有效时,不得从 `executor_final_answer` 抽取额外确认事实。
|
||||
- 对 `gatekeeper_result.severity=reject` 不得输出 `PASS`。
|
||||
- 对 `gatekeeper_result.severity=low_confid` 不得输出 `PASS`。
|
||||
|
||||
`ToolTraceSummaryService` 聚合 evidence tools:
|
||||
|
||||
```text
|
||||
lookup_knowledge, query_logs, query_metrics, query_order
|
||||
```
|
||||
|
||||
输出项结构:
|
||||
|
||||
```json
|
||||
{
|
||||
"trace_ref": "trace-1",
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"input_summary": "query=payment-service timeout",
|
||||
"output_summary": "log_evidence: ...",
|
||||
"evidence_level": "direct",
|
||||
"topic_domain": "general",
|
||||
"source_invocation_ids": [394],
|
||||
"invocation_count": 1,
|
||||
"failed_invocation_count": 0,
|
||||
"no_hit_invocation_count": 0,
|
||||
"query_samples": ["payment-service timeout"],
|
||||
"retrieval_layers": [],
|
||||
"relevance_levels": [],
|
||||
"source_documents": []
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `trace_ref` | string | Verifier 可引用的证据摘要编号 |
|
||||
| `tool_name` | string | 聚合后的工具名 |
|
||||
| `success` | boolean | 是否存在可用证据 |
|
||||
| `input_summary` | string | 工具输入摘要 |
|
||||
| `output_summary` | string | 工具输出摘要 |
|
||||
| `evidence_level` | string | `direct` / `indirect` / `none` |
|
||||
| `topic_domain` | string | 主题域,优先来自 `retrieval_details.retrieved_domains` |
|
||||
| `source_invocation_ids` | array | 聚合的 `tool_invocation.id` |
|
||||
| `invocation_count` | number | 聚合调用次数 |
|
||||
| `failed_invocation_count` | number | 失败调用次数 |
|
||||
| `no_hit_invocation_count` | number | 无证据或去重调用次数 |
|
||||
| `query_samples` | array | 查询样例 |
|
||||
| `retrieval_layers` | array | 检索层级 |
|
||||
| `relevance_levels` | array | 相关性等级 |
|
||||
| `source_documents` | array | 来源文档标签 |
|
||||
|
||||
### 4.3 Output:`verifier_output`
|
||||
|
||||
当前 `chat-verifier-prompt.md` 要求输出:
|
||||
输出:
|
||||
|
||||
```json
|
||||
{
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 0.8,
|
||||
"critical_fact_count": 2,
|
||||
"facts_checked": [
|
||||
"groundedness_score": 1.0,
|
||||
"critical_fact_count": 1,
|
||||
"claim_checks": [],
|
||||
"facts_checked": [],
|
||||
"rationale": "..."
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6. Composer
|
||||
|
||||
Composer 位于 Verifier 之后,输入是 ChatService 过滤后的允许表达材料。
|
||||
|
||||
输入概念:
|
||||
|
||||
| 字段 | 定义 |
|
||||
|---|---|
|
||||
| `original_query` | 用户原始问题 |
|
||||
| `verdict` | `PASS` / `LOW_CONFID` / `REJECT` |
|
||||
| `allowed_claims` | Verifier 允许表达的 claims |
|
||||
| `allowed_hypotheses` | Verifier 允许表达的假设 |
|
||||
| `missing_info` | 证据缺口 |
|
||||
| `recommended_actions` | 允许表达的建议动作 |
|
||||
| `rationale` | Verifier 判定理由 |
|
||||
|
||||
输出:
|
||||
|
||||
```json
|
||||
{
|
||||
"answer_summary": "一句话概括",
|
||||
"recommended_actions": [
|
||||
{
|
||||
"fact": "ERR_TIMEOUT 表示请求超时",
|
||||
"is_critical": true,
|
||||
"verification": "direct_evidence",
|
||||
"detail": "知识库文档明确给出该错误码定义",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"trace_ref": "trace-1",
|
||||
"tool_name": "lookup_knowledge",
|
||||
"topic_domain": "api",
|
||||
"source_invocation_ids": [101, 104],
|
||||
"note": "trace-1 的文档摘要直接给出错误码定义"
|
||||
}
|
||||
]
|
||||
"action_text": "下一步动作",
|
||||
"reason": "原因"
|
||||
}
|
||||
],
|
||||
"rationale": "所有关键事实均有支撑,且至少一条具有直接证据"
|
||||
"user_facing_answer": "最终给用户看的中文答案"
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `verdict` | string | `PASS` / `LOW_CONFID` / `REJECT` |
|
||||
| `groundedness_score` | number | 关键事实证据支撑评分 |
|
||||
| `critical_fact_count` | number | `facts_checked` 中 `is_critical=true` 的数量 |
|
||||
| `facts_checked` | array | Verifier 校验过的事实列表 |
|
||||
| `facts_checked[].fact` | string | 被校验事实 |
|
||||
| `facts_checked[].is_critical` | boolean | 是否关键事实 |
|
||||
| `facts_checked[].verification` | string | `direct_evidence` / `indirect_support` / `no_evidence` / `contradicted` |
|
||||
| `facts_checked[].detail` | string | 校验说明 |
|
||||
| `facts_checked[].evidence_refs` | array | 证据引用 |
|
||||
| `evidence_refs[].trace_ref` | string | 引用的 `tool_trace_summary.trace_ref` |
|
||||
| `evidence_refs[].tool_name` | string | 引用工具 |
|
||||
| `evidence_refs[].topic_domain` | string | 引用主题域 |
|
||||
| `evidence_refs[].source_invocation_ids` | array | 引用的 `tool_invocation.id` |
|
||||
| `evidence_refs[].note` | string | 引用说明 |
|
||||
| `rationale` | string | verdict 判定理由 |
|
||||
表达边界:
|
||||
|
||||
运行态输出 key:
|
||||
|
||||
```text
|
||||
verifier_output
|
||||
```
|
||||
- Composer 不补事实、不补根因、不调用工具。
|
||||
- 只表达 `allowed_claims`、`allowed_hypotheses`、`missing_info`、`recommended_actions`。
|
||||
- 当 claim 是 `negative_observation` 或证据来自 `$.no_evidence` 时,只能表达“当前查询未检索到 / 本次检索未发现匹配证据”。
|
||||
- 禁止表达“问题不存在”“已排除该问题”“确认没有”“日志层面已排除”等过度结论。
|
||||
|
||||
---
|
||||
|
||||
## 5. VerifierDecision
|
||||
## 7. Trace Persistence
|
||||
|
||||
`ChatService.parseVerifierDecision(...)` 将 `verifier_output` 解析为内部 record:
|
||||
|
||||
```json
|
||||
{
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundednessScore": 0.5,
|
||||
"criticalFactCount": 2,
|
||||
"factsChecked": [],
|
||||
"rationale": "证据不足",
|
||||
"round": 1
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `verdict` | string | Verifier verdict |
|
||||
| `groundednessScore` | number | groundedness score |
|
||||
| `criticalFactCount` | number | 关键事实数量 |
|
||||
| `factsChecked` | array | 解析后的 facts_checked |
|
||||
| `rationale` | string | 判定理由 |
|
||||
| `round` | number | 当前验证轮次 |
|
||||
|
||||
---
|
||||
|
||||
## 6. retry_context
|
||||
|
||||
当 `LOW_CONFID` 且满足重试条件时,`ChatService.buildRetryContext(...)` 构造:
|
||||
|
||||
```json
|
||||
{
|
||||
"round": 1,
|
||||
"missing_evidence_facts": [
|
||||
"某关键事实:缺少直接证据"
|
||||
],
|
||||
"instruction": "仅补充以上断言相关证据,不要重复已完成检索"
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `round` | number | 触发 retry 的轮次 |
|
||||
| `missing_evidence_facts` | array | 来自 Verifier 的证据缺口 |
|
||||
| `instruction` | string | 补证据约束 |
|
||||
|
||||
---
|
||||
|
||||
## 7. diagnosis_session.self_evaluation.verifier_evaluation
|
||||
|
||||
`ChatService.persistVerifierEvaluation(...)` 将 Verifier 结果合并进 `diagnosis_session.self_evaluation`。
|
||||
`diagnosis_run.self_evaluation.verifier_evaluation` 持久化:
|
||||
|
||||
```json
|
||||
{
|
||||
"verifier_evaluation": {
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.5,
|
||||
"critical_fact_count": 2,
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 1.0,
|
||||
"critical_fact_count": 1,
|
||||
"claim_checks": [],
|
||||
"facts_checked": [],
|
||||
"rationale": "证据不足",
|
||||
"rationale": "...",
|
||||
"round": 1,
|
||||
"traceability_version": "v1",
|
||||
"executor_output_parse_status": {
|
||||
"status": "valid",
|
||||
"detail": "parsed executor evidence contract"
|
||||
},
|
||||
"executor_output_parse_status": {},
|
||||
"executor_structured_output": {},
|
||||
"gatekeeper_result": {
|
||||
"rule_set_version": "gatekeeper-rules-v1"
|
||||
},
|
||||
"composer_output": {},
|
||||
"tool_trace_summary": []
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `verifier_evaluation.verdict` | string | Verifier verdict |
|
||||
| `verifier_evaluation.groundedness_score` | number | groundedness score |
|
||||
| `verifier_evaluation.critical_fact_count` | number | 关键事实数量 |
|
||||
| `verifier_evaluation.facts_checked` | array | 校验事实列表 |
|
||||
| `verifier_evaluation.rationale` | string | 判定理由 |
|
||||
| `verifier_evaluation.round` | number | 验证轮次 |
|
||||
| `verifier_evaluation.traceability_version` | string | 当前固定为 `v1` |
|
||||
| `verifier_evaluation.executor_output_parse_status` | object | Executor 输出解析状态 |
|
||||
| `verifier_evaluation.executor_structured_output` | object/null | 解析后的 Executor 结构化输出 |
|
||||
| `verifier_evaluation.tool_trace_summary` | array | Verifier 使用的工具证据索引 |
|
||||
Trace API 可用于回放:
|
||||
|
||||
- Executor 输出了哪些 claim。
|
||||
- 每个 claim 引用了哪些 `source_invocation_id + raw_path + evidence_excerpt`。
|
||||
- Gatekeeper 是否通过、是否自动回填。
|
||||
- Verifier 如何判断可推导性。
|
||||
- Composer 最终如何表达给用户。
|
||||
|
||||
---
|
||||
|
||||
## 8. Final Answer Rendering
|
||||
## 8. 当前已验证样例
|
||||
|
||||
ChatService 根据 Verifier verdict 决定最终 `diagnosis_session.answer`。
|
||||
|
||||
| Verdict | 当前行为 |
|
||||
|---|---|
|
||||
| `PASS` | 优先提取 `executor_feedback.user_facing_answer`;提取失败则使用 executor 原文 |
|
||||
| `LOW_CONFID` | 输出低置信模板:已确认信息、当前缺口、建议下一步 |
|
||||
| `REJECT` | 输出降级模板:已确认信息、证据缺口、建议下一步 |
|
||||
|
||||
低置信模板使用:
|
||||
|
||||
```text
|
||||
以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。
|
||||
|
||||
已确认信息:
|
||||
- ...
|
||||
|
||||
当前缺口:
|
||||
- ...
|
||||
|
||||
建议下一步:
|
||||
- ...
|
||||
```
|
||||
|
||||
拒绝模板使用:
|
||||
|
||||
```text
|
||||
当前无法基于已获取证据生成可靠结论。
|
||||
|
||||
已确认信息:
|
||||
- ...
|
||||
|
||||
证据缺口:
|
||||
- ...
|
||||
|
||||
建议下一步:
|
||||
- ...
|
||||
```
|
||||
| 场景 | sessionId | 结果 |
|
||||
|---|---|---|
|
||||
| HighCPUUsage 窄范围正向确认 | `iss008-narrow-highcpu-rerun-20260708-215510` | `PASS`,1 条 `observation`,无越界 claim |
|
||||
| HikariCP negative_observation | `iss009-hikari-negative-latest-20260708-232428` | `PASS`,`raw_path=$.no_evidence`,无过度表达 |
|
||||
|
||||
---
|
||||
|
||||
## 9. Trace Persistence Data
|
||||
## 9. 仍需记录或后续补强
|
||||
|
||||
### 9.1 diagnosis_session
|
||||
当前架构文档已记录主链路、数据契约和语义边界。后续如果继续实现,建议再补:
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `session_id` | string | 会话 id |
|
||||
| `query` | text | 用户问题 |
|
||||
| `status` | string | 会话状态 |
|
||||
| `agent_flow` | string | 当前 Chat 链路为 `CHAT` |
|
||||
| `total_duration_ms` | number | 总耗时 |
|
||||
| `total_token_count` | number | 总 token |
|
||||
| `step_count` | number | agent step 数 |
|
||||
| `tool_call_count` | number | tool invocation 数 |
|
||||
| `answer` | longtext | 最终用户答案 |
|
||||
| `self_evaluation` | json | 包含 verifier_evaluation |
|
||||
| `feedback` | string | 用户反馈 |
|
||||
|
||||
### 9.2 agent_step
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `session_id` | string | 会话 id |
|
||||
| `step_index` | number | 步骤序号 |
|
||||
| `agent_name` | string | `planner` / `executor` / `verifier` |
|
||||
| `model_input` | text | 模型输入摘要 |
|
||||
| `model_output` | text | 模型输出摘要 |
|
||||
| `thought` | text | hook 记录的摘要信息 |
|
||||
| `has_tool_call` | boolean | 是否包含工具调用 |
|
||||
| `duration_ms` | number | 模型调用耗时 |
|
||||
| `token_count` | number | token 数 |
|
||||
|
||||
### 9.3 tool_invocation
|
||||
|
||||
| 字段 | 类型 | 定义 |
|
||||
|---|---|---|
|
||||
| `id` | number | 工具调用 id |
|
||||
| `session_id` | string | 会话 id |
|
||||
| `step_id` | number | 对应 agent_step id |
|
||||
| `tool_name` | string | 工具名 |
|
||||
| `input_params` | json | 工具输入参数 |
|
||||
| `output_preview` | text | 工具输出预览 |
|
||||
| `output_length` | number | 原始输出长度 |
|
||||
| `retrieval_layer` | string | 检索层 |
|
||||
| `l0_match_count` | number | L0 命中数 |
|
||||
| `l1_match_count` | number | L1 命中数 |
|
||||
| `is_truncated` | boolean | 输出是否截断 |
|
||||
| `relevance_level` | string | 相关性等级 |
|
||||
| `dedup_reason` | string | 去重原因 |
|
||||
| `retrieval_details` | json | 检索细节 |
|
||||
| `duration_ms` | number | 工具耗时 |
|
||||
| `success` | boolean | 是否成功 |
|
||||
| `error_message` | text | 错误信息 |
|
||||
1. Planner `scope_contract` 的 ADR:只有当 Prompt-first 无法稳定控制越界时再引入。
|
||||
2. 更完整的 Gatekeeper 规则配置化:当前只有本地轻量 metadata/catalog,后续如果做索引层、元数据层、远程规则层,需要单独记录加载顺序、变更审批和回滚策略。
|
||||
3. Prompt version 记录:当前 prompt 变更没有版本号,后续如果需要回滚和对比,应记录 prompt version。
|
||||
|
||||
@@ -1,15 +1,15 @@
|
||||
# 反馈与自评估架构
|
||||
|
||||
**更新日期**:2026-07-05
|
||||
**状态**:当前可运行架构
|
||||
**更新日期**:2026-07-10
|
||||
**状态**:当前可运行架构
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/confidence-feedback.md`
|
||||
|
||||
## 1. 定位
|
||||
|
||||
反馈架构包含两条闭环:
|
||||
|
||||
1. 系统自评估:基于工具调用、Verifier、AIOps 规则检查,写入 `diagnosis_session.self_evaluation`。
|
||||
2. 用户反馈:用户标记 `useful` 或 `not_useful`,写入 `diagnosis_session.feedback`,其中 `useful` 会沉淀案例。
|
||||
1. 系统自评估:基于当前 run 的工具调用、Gatekeeper、Verifier、Composer、AIOps 规则检查,写入 `diagnosis_run.self_evaluation`。
|
||||
2. 用户反馈:用户标记 `useful` 或 `not_useful`,优先写入 `diagnosis_run.feedback`,其中 `useful` 会沉淀案例。
|
||||
|
||||
当前重要边界:
|
||||
|
||||
@@ -21,13 +21,18 @@
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Answer["Chat / AIOps final answer"] --> Session["diagnosis_session.answer"]
|
||||
Answer["Chat / AIOps final answer"] --> Run["diagnosis_run.answer"]
|
||||
|
||||
subgraph SelfEval["Self evaluation"]
|
||||
Invocation["tool_invocation"] --> RuleEval["EvaluationService: rule_evaluation"]
|
||||
Invocation --> EvidenceRefs["evidence_refs"]
|
||||
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
Invocation --> TraceSummary["ToolTraceSummaryService"]
|
||||
TraceSummary --> Verifier["chat_verifier"]
|
||||
Gatekeeper --> Verifier["chat_verifier"]
|
||||
TraceSummary --> Verifier
|
||||
Verifier --> VerifierEval["verifier_evaluation"]
|
||||
Verifier --> Composer["chat_composer"]
|
||||
Composer --> VerifierEval
|
||||
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
|
||||
AiOpsRule --> AiOpsEval["aiops_rule_evaluation"]
|
||||
end
|
||||
@@ -35,24 +40,24 @@ flowchart TD
|
||||
RuleEval --> Merge["SelfEvaluationMergeService"]
|
||||
VerifierEval --> Merge
|
||||
AiOpsEval --> Merge
|
||||
Merge --> SelfJson["diagnosis_session.self_evaluation"]
|
||||
Merge --> SelfJson["diagnosis_run.self_evaluation"]
|
||||
|
||||
subgraph UserFeedback["User feedback"]
|
||||
UI["Feedback bar"] --> API["POST /api/feedback"]
|
||||
API --> FeedbackService["FeedbackService"]
|
||||
FeedbackService --> FeedbackField["diagnosis_session.feedback"]
|
||||
FeedbackService --> FeedbackField["diagnosis_run.feedback"]
|
||||
FeedbackService --> Useful{"feedback == useful?"}
|
||||
Useful -->|yes| CaseService["CaseLibraryService.createFromSession"]
|
||||
Useful -->|yes| CaseService["CaseLibraryService.createFromRun"]
|
||||
CaseService --> Case["case_library"]
|
||||
Useful -->|no| BadCase["Bad case by feedback=not_useful"]
|
||||
end
|
||||
|
||||
Session --> UI
|
||||
Run --> UI
|
||||
```
|
||||
|
||||
## 3. self_evaluation JSON
|
||||
|
||||
`SelfEvaluationMergeService` 统一维护 `diagnosis_session.self_evaluation`。
|
||||
`SelfEvaluationMergeService` 统一维护当前运行的 `diagnosis_run.self_evaluation`。历史兼容数据可能仍存在于 `diagnosis_session.self_evaluation`,但新 Chat/AIOps 执行不再写旧表。
|
||||
|
||||
当前结构:
|
||||
|
||||
@@ -67,10 +72,15 @@ flowchart TD
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 0.8,
|
||||
"critical_fact_count": 2,
|
||||
"claim_checks": [],
|
||||
"facts_checked": [],
|
||||
"rationale": "...",
|
||||
"round": 1,
|
||||
"traceability_version": "v1",
|
||||
"executor_output_parse_status": {},
|
||||
"executor_structured_output": {},
|
||||
"gatekeeper_result": {},
|
||||
"composer_output": {},
|
||||
"tool_trace_summary": []
|
||||
},
|
||||
"aiops_rule_evaluation": {
|
||||
@@ -117,17 +127,29 @@ flowchart TD
|
||||
|
||||
## 5. Chat Verifier 自评估
|
||||
|
||||
Chat Verifier 校验 Executor 的最终答案是否被证据支撑。
|
||||
Chat 自评估分三步:
|
||||
|
||||
1. Gatekeeper 用代码校验 Executor 的引用是否真实。
|
||||
2. Verifier 判断已验真的 `evidence_excerpt` 是否能推出 `claim_text`。
|
||||
3. Composer 只把 Verifier 允许表达的内容写成最终用户答复。
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Answer["executor_final_answer"] --> Verifier["chat_verifier"]
|
||||
ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
Invocation["tool_invocation"] --> Summary["ToolTraceSummaryService"]
|
||||
Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
|
||||
EvidenceRefs --> Gatekeeper
|
||||
Gatekeeper --> GateResult["gatekeeper_result"]
|
||||
Summary --> Evidence["tool_trace_summary"]
|
||||
GateResult --> Verifier["chat_verifier"]
|
||||
ExecutorOutput --> Verifier
|
||||
Evidence --> Verifier
|
||||
Verifier --> Output["verifier_output JSON"]
|
||||
Output --> Composer["chat_composer"]
|
||||
Composer --> ComposerOutput["composer_output"]
|
||||
Output --> Merge["SelfEvaluationMergeService.mergeVerifierEvaluation"]
|
||||
Merge --> Session["diagnosis_session.self_evaluation.verifier_evaluation"]
|
||||
ComposerOutput --> Merge
|
||||
Merge --> Run["diagnosis_run.self_evaluation.verifier_evaluation"]
|
||||
```
|
||||
|
||||
Verifier 输出:
|
||||
@@ -137,16 +159,25 @@ Verifier 输出:
|
||||
| `verdict` | `PASS` / `LOW_CONFID` / `REJECT` |
|
||||
| `groundedness_score` | 关键事实证据支撑度 |
|
||||
| `critical_fact_count` | 关键事实数量 |
|
||||
| `claim_checks` | 对 Executor 结构化 claims 的逐条可推导性判断 |
|
||||
| `facts_checked` | 逐条事实校验 |
|
||||
| `rationale` | 判定原因 |
|
||||
| `tool_trace_summary` | 本次校验使用的证据索引 |
|
||||
| `executor_structured_output` | Executor 输出的结构化 claims 与证据绑定 |
|
||||
| `gatekeeper_result` | 引用真实性校验结果 |
|
||||
| `composer_output` | 最终表达的解析状态和摘要 |
|
||||
| `tool_trace_summary` | 本次校验使用的工具调用导航索引 |
|
||||
|
||||
ChatService 根据 verdict 决定:
|
||||
|
||||
- `PASS`:输出 Executor 答案。
|
||||
- `PASS`:把允许表达的 claims 交给 Composer 输出。
|
||||
- `LOW_CONFID`:必要时构造 `retry_context` 补证据;否则输出低置信提示。
|
||||
- `REJECT`:降级输出,只保留已确认信息。
|
||||
|
||||
边界:
|
||||
|
||||
- `executor_final_answer` 只作为 debug/fallback 上下文;结构化输出有效时,Verifier 不得从中抽取额外确认事实。
|
||||
- `$.no_evidence` 只能表达“当前查询未检索到匹配证据”,不能表达“已排除/确认没有”。
|
||||
|
||||
## 6. AIOps 规则自评估
|
||||
|
||||
AIOps 当前使用 `AiOpsRuleEvaluationService`,结果写入 `aiops_rule_evaluation`。
|
||||
@@ -167,6 +198,7 @@ POST /api/feedback
|
||||
Content-Type: application/json
|
||||
|
||||
{
|
||||
"runId": "run-xxx",
|
||||
"sessionId": "xxx",
|
||||
"feedback": "useful" | "not_useful"
|
||||
}
|
||||
@@ -178,6 +210,8 @@ Content-Type: application/json
|
||||
{
|
||||
"success": true,
|
||||
"message": "反馈已记录",
|
||||
"runId": "run-xxx",
|
||||
"fallbackToLatestRun": false,
|
||||
"caseId": "uuid 或 null"
|
||||
}
|
||||
```
|
||||
@@ -186,10 +220,16 @@ Content-Type: application/json
|
||||
|
||||
| feedback | 行为 |
|
||||
|---|---|
|
||||
| `useful` | 写入 `DiagnosisSession.feedback`,调用 `CaseLibraryService.createFromSession` |
|
||||
| `not_useful` | 写入 `DiagnosisSession.feedback`,不改变 session status |
|
||||
| `useful` | 写入 `DiagnosisRun.feedback`,调用 `CaseLibraryService.createFromRun` |
|
||||
| `not_useful` | 写入 `DiagnosisRun.feedback`,不改变 run status |
|
||||
| 其他值 | 返回 HTTP 400 |
|
||||
|
||||
兼容行为:
|
||||
|
||||
- 请求带 `runId` 时,后端验证 `runId` 属于 `sessionId`。
|
||||
- 请求缺少 `runId` 且存在 run-backed 数据时,后端绑定 latest run,并返回 `fallbackToLatestRun=true` 和实际 `runId`。
|
||||
- 仅当没有 `diagnosis_run` 但存在历史 `diagnosis_session` 时,才使用历史 fallback;该路径不声明 latest-run fallback。
|
||||
|
||||
## 8. 案例沉淀
|
||||
|
||||
`useful` 反馈会生成或复用 `case_library` 记录。
|
||||
@@ -199,7 +239,7 @@ Content-Type: application/json
|
||||
| CaseLibrary 字段 | 来源 |
|
||||
|---|---|
|
||||
| `caseId` | UUID |
|
||||
| `diagnosisId` | `DiagnosisSession.sessionId` |
|
||||
| `diagnosisId` | 新数据为 `DiagnosisRun.runId`;历史数据可能为 `DiagnosisSession.sessionId` |
|
||||
| `sourceType` | `AUTO` |
|
||||
| `faultCategory` | 当前固定为 `GENERAL` |
|
||||
| `title` | `query` 前 100 字符 |
|
||||
@@ -210,7 +250,7 @@ Content-Type: application/json
|
||||
幂等性:
|
||||
|
||||
```text
|
||||
case_library.diagnosisId == sessionId
|
||||
case_library.diagnosisId == runId
|
||||
-> existing case: return existing
|
||||
-> missing case: create new
|
||||
```
|
||||
@@ -229,23 +269,23 @@ Trace API 会展示:
|
||||
|
||||
| 视角 | 数据来源 |
|
||||
|---|---|
|
||||
| 执行是否成功 | `diagnosis_session.status` |
|
||||
| 执行是否成功 | `diagnosis_run.status` |
|
||||
| 证据是否充分 | `self_evaluation.rule_evaluation` / `verifier_evaluation` |
|
||||
| 用户是否认可 | `diagnosis_session.feedback` |
|
||||
| 用户是否认可 | `diagnosis_run.feedback` |
|
||||
|
||||
## 10. 后续增强
|
||||
|
||||
近期优先:
|
||||
|
||||
1. 将 `rule_evaluation` 与 `verifier_evaluation` 在 Trace API 中结构化展示。
|
||||
2. `not_useful` 反馈沉淀 bad case,而不是只写字段。
|
||||
3. useful 案例自动提取 faultCategory、errorCode、service、rootCause、solution。
|
||||
4. AIOps 引入 LLM Verifier。
|
||||
5. 把反馈和 eval baseline 打通,形成可回归的质量改进闭环。
|
||||
2. 将 ISS-008 / ISS-009 这类 E2E 通过样例固化进 diagnosis eval fixtures。
|
||||
3. `not_useful` 反馈沉淀 bad case,而不是只写字段。
|
||||
4. useful 案例自动提取 faultCategory、errorCode、service、rootCause、solution。
|
||||
5. AIOps 引入 LLM Verifier。
|
||||
6. 把反馈和 eval baseline 打通,形成可回归的质量改进闭环。
|
||||
|
||||
暂不优先:
|
||||
|
||||
- 用用户反馈直接修改 session status。
|
||||
- 仅凭 `evidence_score` 判断答案正确。
|
||||
- 在没有人工审核时自动把 bad case 反向写入 Prompt。
|
||||
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
# Harness 与质量门禁架构
|
||||
|
||||
**更新日期**:2026-07-05
|
||||
**状态**:当前可运行架构 + 后续门禁规划
|
||||
**更新日期**:2026-07-08
|
||||
**状态**:当前可运行架构 + 后续门禁规划
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
||||
|
||||
## 1. 设计目标
|
||||
@@ -21,6 +21,7 @@ Prompt contract
|
||||
+ Tool boundary
|
||||
+ Agent hooks
|
||||
+ Trace persistence
|
||||
+ Gatekeeper deterministic validation
|
||||
+ Verifier / rule evaluation
|
||||
+ Eval baseline
|
||||
```
|
||||
@@ -30,21 +31,25 @@ Prompt contract
|
||||
```mermaid
|
||||
flowchart TB
|
||||
Input["User / AIOps input"] --> Prompt["Prompt contract"]
|
||||
Prompt --> Agent["Planner / Executor / Verifier"]
|
||||
Prompt --> Agent["Planner / Executor / Verifier / Composer"]
|
||||
Agent --> Tools["Evidence tools"]
|
||||
Tools --> Invocation["tool_invocation"]
|
||||
Agent --> StepHook["AgentLoggingHook"]
|
||||
StepHook --> Step["agent_step"]
|
||||
Agent --> Session["diagnosis_session"]
|
||||
Agent --> Run["diagnosis_run"]
|
||||
|
||||
Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
|
||||
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
Agent --> Gatekeeper
|
||||
Invocation --> TraceSummary["ToolTraceSummaryService"]
|
||||
TraceSummary --> Verifier["chat_verifier"]
|
||||
Gatekeeper --> Verifier["chat_verifier"]
|
||||
TraceSummary --> Verifier
|
||||
Verifier --> SelfEval["self_evaluation.verifier_evaluation"]
|
||||
|
||||
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
|
||||
AiOpsRule --> AiOpsEval["self_evaluation.aiops_rule_evaluation"]
|
||||
|
||||
Session --> TraceAPI["DiagnosisTraceService"]
|
||||
Run --> TraceAPI["DiagnosisTraceService"]
|
||||
Step --> TraceAPI
|
||||
Invocation --> TraceAPI
|
||||
SelfEval --> TraceAPI
|
||||
@@ -63,17 +68,37 @@ flowchart TB
|
||||
| `planner-prompt.md` | AIOps Planner 规划、再规划、输出告警报告 |
|
||||
| `executor-prompt.md` | AIOps Executor 按步骤调用工具 |
|
||||
| `chat-planner-prompt.md` | Chat 复杂问题规划 |
|
||||
| `chat-executor-prompt.md` | Chat 执行工具并形成诊断答复 |
|
||||
| `chat-verifier-prompt.md` | 校验 Executor 答案是否被工具证据支撑 |
|
||||
| `chat-executor-prompt.md` | Chat 执行工具并输出 `executor_evidence_v2` 微观事实 |
|
||||
| `chat-verifier-prompt.md` | 基于 Gatekeeper 已验真的证据判断 claims 是否可推出 |
|
||||
| `chat-composer-prompt.md` | 基于 Verifier 允许表达的内容生成最终用户答复 |
|
||||
|
||||
Prompt 层当前承担的门禁:
|
||||
|
||||
- 禁止凭记忆回答错误码、接口定义、排障步骤。
|
||||
- 需要外部信息时必须调用工具。
|
||||
- 工具连续失败或返回空结果时,最终报告必须诚实说明。
|
||||
- Chat Verifier 不允许做新检索,只能校验已有证据。
|
||||
- Chat Executor 不允许在窄范围问题中扩展根因、风险或修复建议。
|
||||
- Chat Verifier 不允许做新检索,只能判断已验真证据是否可推出 claims。
|
||||
- Chat Composer 不允许补事实,尤其不能把 `$.no_evidence` 表达为“已排除/确认没有”。
|
||||
- AIOps payload 模式必须聚焦输入告警。
|
||||
|
||||
Chat 链路还会在 `verifier_evaluation.prompt_audit` 中持久化紧凑 Prompt 审计快照:
|
||||
|
||||
```json
|
||||
{
|
||||
"version": "chat-prompts-v1",
|
||||
"prompts": [
|
||||
{
|
||||
"name": "chat_executor",
|
||||
"version": "chat-executor-v2",
|
||||
"resource": "prompts/chat-executor-prompt.md"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
该快照只保存版本和资源路径,不保存完整 Prompt 文本。它用于面试演示、trace 回放和离线 baseline 解释“本次诊断使用了哪套 Prompt 契约”。
|
||||
|
||||
## 4. Trace Hooks
|
||||
|
||||
`AgentLoggingHook` 是当前 Agent step 可观测性的核心。
|
||||
@@ -85,10 +110,10 @@ sequenceDiagram
|
||||
participant H as AgentLoggingHook
|
||||
participant DB as agent_step
|
||||
|
||||
A->>H: before_model(messages, sessionId)
|
||||
H->>DB: 写入 model_input / step_index / agent_name
|
||||
A->>H: before_model(messages, sessionId, runId)
|
||||
H->>DB: 写入 session_id / run_id / model_input / step_index / agent_name
|
||||
A-->>A: LLM 推理
|
||||
A->>H: after_model(messages, sessionId)
|
||||
A->>H: after_model(messages, sessionId, runId)
|
||||
H->>DB: 回填 model_output / thought / has_tool_call / duration / token_count
|
||||
```
|
||||
|
||||
@@ -101,6 +126,8 @@ sequenceDiagram
|
||||
- token count。
|
||||
- Verifier 的 JSON 输出摘要。
|
||||
|
||||
新写入必须带 `run_id`;`session_id` 仍保留用于粗粒度排查和历史兼容。
|
||||
|
||||
## 5. Tool Invocation 门禁
|
||||
|
||||
工具调用记录由 `ToolInvocationRecorder` 和具体工具共同完成。
|
||||
@@ -115,6 +142,7 @@ retrieval_layer
|
||||
l0_match_count
|
||||
l1_match_count
|
||||
retrieval_details
|
||||
-> evidence_refs
|
||||
relevance_level
|
||||
dedup_reason
|
||||
duration_ms
|
||||
@@ -128,23 +156,44 @@ error_message
|
||||
- 检索结果归一化为 `PRECISE`、`HIGHLY_RELEVANT`、`REFERENCE`。
|
||||
- 同 session 内重复文档会被 `RetrievedDocTracker` 去重。
|
||||
- dedup、no evidence、failed 等状态进入 `retrieval_details.evidence_status`。
|
||||
- `retrieval_details.evidence_refs` 记录可被 Executor 引用的最小证据文本,格式为 `raw_path + text`。
|
||||
- no-hit / no-evidence 工具结果会生成 `raw_path=$.no_evidence` 的负向证据引用,语义仅限“本次查询未检索到匹配证据”。
|
||||
|
||||
## 6. Verifier 门禁
|
||||
## 6. Gatekeeper 与 Verifier 门禁
|
||||
|
||||
Chat Verifier 的输入不是原始工具日志,而是 `ToolTraceSummaryService` 构造的证据索引。
|
||||
Chat Verifier 前置一层 Gatekeeper。Gatekeeper 不调用 LLM,只用代码检查 Executor 输出的证据引用是否真实存在。
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Invocation["tool_invocation"] --> Summary["ToolTraceSummaryService"]
|
||||
Invocation["tool_invocation"] --> EvidenceRefs["retrieval_details.evidence_refs"]
|
||||
ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
EvidenceRefs --> Gatekeeper
|
||||
Gatekeeper --> GateResult["gatekeeper_result"]
|
||||
Invocation --> Summary["ToolTraceSummaryService"]
|
||||
Summary --> EvidenceIndex["tool_trace_summary"]
|
||||
EvidenceIndex --> Verifier["chat_verifier"]
|
||||
ExecutorAnswer["executor_final_answer"] --> Verifier
|
||||
GateResult --> Verifier["chat_verifier"]
|
||||
ExecutorOutput --> Verifier
|
||||
EvidenceIndex --> Verifier
|
||||
Verifier --> Verdict{"verdict"}
|
||||
Verdict -->|PASS| Pass["输出原答案"]
|
||||
Verdict -->|PASS| Composer["chat_composer"]
|
||||
Composer --> Pass["输出最终答复"]
|
||||
Verdict -->|LOW_CONFID| Low["补证据或低置信输出"]
|
||||
Verdict -->|REJECT| Reject["降级输出"]
|
||||
```
|
||||
|
||||
Gatekeeper 检查:
|
||||
|
||||
| 检查 | 失败语义 |
|
||||
|---|---|
|
||||
| `answer_version=executor_evidence_v2` | 非结构化或旧结构输出降为低置信 |
|
||||
| `source_invocation_id` 真实存在 | 伪造 ID 直接拒绝 |
|
||||
| `tool_name` 与 invocation 对齐 | 张冠李戴直接拒绝 |
|
||||
| `raw_path` 存在于 `evidence_refs` | 无中生有直接拒绝 |
|
||||
| `evidence_excerpt` 由 `evidence_refs[].text` 支撑 | excerpt 编造或错配直接拒绝 |
|
||||
| `negative_observation` 只能引用 `$.no_evidence` | 用正向日志证明“没查到”直接拒绝 |
|
||||
|
||||
Gatekeeper 审计还会记录 `rule_set_version` 和已启用规则元数据摘要。当前规则元数据来自本地 `gatekeeper-rules.json`,规则执行仍是确定性 Java 代码。
|
||||
|
||||
Verifier 输出:
|
||||
|
||||
```json
|
||||
@@ -152,17 +201,22 @@ Verifier 输出:
|
||||
"verdict": "PASS|LOW_CONFID|REJECT",
|
||||
"groundedness_score": 0.8,
|
||||
"critical_fact_count": 2,
|
||||
"claim_checks": [],
|
||||
"facts_checked": [],
|
||||
"rationale": "..."
|
||||
}
|
||||
```
|
||||
|
||||
Verifier 不再逐字核验 excerpt 真伪;这由 Gatekeeper 完成。Verifier 只回答一个问题:`claim_text` 是否能由已经验真的 `evidence_excerpt` 推导出来。
|
||||
|
||||
结果写入:
|
||||
|
||||
```text
|
||||
diagnosis_session.self_evaluation.verifier_evaluation
|
||||
diagnosis_run.self_evaluation.verifier_evaluation
|
||||
```
|
||||
|
||||
其中同时持久化 `executor_structured_output`、`gatekeeper_result`、`tool_trace_summary`、`prompt_audit` 和 `composer_output`,用于 Trace 回放。
|
||||
|
||||
## 7. AIOps 规则门禁
|
||||
|
||||
AIOps 当前不走 Chat Verifier,而是用 `AiOpsRuleEvaluationService` 做轻量检查。
|
||||
@@ -177,7 +231,7 @@ AIOps 当前不走 Chat Verifier,而是用 `AiOpsRuleEvaluationService` 做轻
|
||||
结果写入:
|
||||
|
||||
```text
|
||||
diagnosis_session.self_evaluation.aiops_rule_evaluation
|
||||
diagnosis_run.self_evaluation.aiops_rule_evaluation
|
||||
```
|
||||
|
||||
## 8. Eval Baseline
|
||||
@@ -197,9 +251,8 @@ diagnosis_session.self_evaluation.aiops_rule_evaluation
|
||||
- 工具参数 schema 校验。
|
||||
- 同一工具调用次数上限。
|
||||
- 工具超时的统一熔断。
|
||||
- 报告中的数值与工具返回值自动对齐校验。
|
||||
- Prompt 版本记录和回滚。
|
||||
- Gatekeeper 规则远程化或三层分离:索引层、元数据层、规则实现层。
|
||||
- Prompt 版本回滚和更细粒度变更审计。
|
||||
- Verifier 对 AIOps 报告的 LLM 级事实校验。
|
||||
|
||||
这些应在评测集扩大后逐步加入,避免一次性把诊断流程卡得过死。
|
||||
|
||||
|
||||
@@ -5,7 +5,7 @@
|
||||
|
||||
## 1. 一句话
|
||||
|
||||
SuperBizAgent 是一个面向企业故障诊断的可追踪 Agent 系统:它把用户问题或 AIOps 告警转换成 Planner、Executor、Verifier 的诊断链路,所有工具证据、模型步骤、最终答案、自评估和用户反馈都能通过同一个 `sessionId` 回放。
|
||||
SuperBizAgent 是一个面向企业故障诊断的可追踪 Agent 系统:它把用户问题或 AIOps 告警转换成 Planner、Executor、Gatekeeper、Verifier、Composer 的诊断链路,`sessionId` 保留多轮上下文,`runId` 精确绑定一次诊断运行;所有工具证据、模型步骤、最终答案、自评估和用户反馈都能按 `sessionId + runId` 回放。
|
||||
|
||||
## 2. 一张图
|
||||
|
||||
@@ -16,7 +16,7 @@ flowchart TB
|
||||
API --> Chat["ChatService"]
|
||||
API --> AiOps["AiOpsService"]
|
||||
|
||||
Chat --> ChatFlow["Chat: Planner -> Executor -> Verifier"]
|
||||
Chat --> ChatFlow["Chat: Planner -> Executor -> Gatekeeper -> Verifier -> Composer"]
|
||||
AiOps --> AiOpsFlow["AIOps: Supervisor -> Planner / Executor"]
|
||||
|
||||
ChatFlow --> Tools["Evidence Tools"]
|
||||
@@ -34,14 +34,15 @@ flowchart TB
|
||||
AiOpsFlow --> Trace
|
||||
Tools --> Trace
|
||||
|
||||
Trace --> Session["diagnosis_session"]
|
||||
Trace --> ChatSession["chat_session"]
|
||||
Trace --> Run["diagnosis_run"]
|
||||
Trace --> Step["agent_step"]
|
||||
Trace --> Invocation["tool_invocation"]
|
||||
|
||||
Invocation --> Verifier["Verifier / Rule Evaluation"]
|
||||
Verifier --> SelfEval["self_evaluation"]
|
||||
|
||||
Session --> TraceAPI["GET /api/diagnosis/{sessionId}/trace"]
|
||||
Run --> TraceAPI["GET /api/diagnosis/{sessionId}/trace?runId=..."]
|
||||
Step --> TraceAPI
|
||||
Invocation --> TraceAPI
|
||||
SelfEval --> TraceAPI
|
||||
@@ -55,24 +56,24 @@ flowchart TB
|
||||
```text
|
||||
这个项目不是把问题直接丢给大模型,而是把诊断拆成可审计的执行链路。
|
||||
|
||||
Chat 复杂问题走 Planner -> Executor -> Verifier:
|
||||
Planner 负责拆解,Executor 负责调用知识库、日志和指标工具,Verifier 只基于已有工具证据校验最终答案。
|
||||
Chat 复杂问题走 Planner -> Executor -> Gatekeeper -> Verifier -> Composer:
|
||||
Planner 负责拆解,Executor 只负责调用知识库、日志和指标工具并提炼带证据引用的微观事实;Gatekeeper 用代码核对 invocation、raw_path 和 excerpt 是否真实;Verifier 判断这些事实能否由已验真的证据推出;Composer 只把允许表达的结论写成最终答案。
|
||||
|
||||
AIOps 告警入口走 Supervisor 调度 Planner/Executor:
|
||||
如果请求里有 alert payload,系统会进入 PAYLOAD_TARGETED 模式,报告必须聚焦这个告警,而不是被当前环境中的其他活跃告警带偏。
|
||||
|
||||
所有过程都会落到 diagnosis_session、agent_step、tool_invocation。
|
||||
所以我可以用一个 sessionId 回放:模型怎么规划、调了哪些工具、工具返回什么、Verifier 怎么判定、用户最后是否反馈有用。
|
||||
会话元数据会落到 chat_session,每次诊断运行会落到 diagnosis_run,步骤和工具明细通过 run_id 关联。
|
||||
所以我可以用 sessionId + runId 精确回放:模型怎么规划、调了哪些工具、工具返回什么、Gatekeeper 怎么验真、Verifier 怎么判定、Composer 最后怎么表达、用户最后是否反馈有用。
|
||||
```
|
||||
|
||||
## 4. 五个亮点
|
||||
|
||||
| 亮点 | 怎么讲 |
|
||||
|---|---|
|
||||
| 可追踪 Agent | 每次诊断都有 `sessionId`,Trace API 可以回放 session、step、tool |
|
||||
| 可追踪 Agent | 每次诊断都有 `runId`,Trace API 可以回放 run、step、tool;同一 `sessionId` 可有多次独立 run |
|
||||
| 显式工具证据链 | `lookup_knowledge`、日志、指标都记录到 `tool_invocation` |
|
||||
| RAG 工程化 | L0 降级为 hint,Spring AI VectorStore 做主检索,SDK fallback 保底 |
|
||||
| 质量门禁 | Chat Verifier 校验 groundedness,AIOps rule evaluation 控制告警聚焦 |
|
||||
| 质量门禁 | Chat Gatekeeper 验引用、Verifier 判可推导、Composer 控表达,AIOps rule evaluation 控制告警聚焦 |
|
||||
| 反馈闭环 | useful 反馈沉淀 `case_library`,not_useful 保留 bad case 信号 |
|
||||
|
||||
## 5. 三个关键取舍
|
||||
@@ -99,15 +100,14 @@ aiops_rule_evaluation -> AIOps 报告是否聚焦告警并使用证据
|
||||
|
||||
| 追问 | 回答方向 |
|
||||
|---|---|
|
||||
| 怎么防止幻觉? | Executor 必须用工具;Verifier 只基于 `tool_trace_summary` 校验;LOW_CONFID/REJECT 会降级输出 |
|
||||
| 怎么防止幻觉? | Executor 输出 `executor_evidence_v2`,每个 claim 绑定 `source_invocation_id + raw_path + evidence_excerpt`;Gatekeeper 用 `tool_invocation.retrieval_details.evidence_refs` 核验引用真实性;Verifier 只判断可推导性;Composer 防止把 no-evidence 说成已排除 |
|
||||
| RAG 质量怎么保证? | offline golden cases + live acceptance + trace inspection 三层验证 |
|
||||
| 为什么 L0 不直接返回? | L0 子串命中不等于语义相关,当前只做 domain/entity hint 和 metadata filter |
|
||||
| AIOps 如何避免跑偏? | payload 模式生成 recommended query,并用 rule evaluation 检查报告聚焦输入告警 |
|
||||
| 下一步怎么演进? | evidence block、邻居 chunk、Playbook、AIOps LLM Verifier、MCP 工具协议化 |
|
||||
| 下一步怎么演进? | 固化 E2E fixture、Prompt version、Gatekeeper 规则配置化、邻居 chunk、AIOps LLM Verifier、MCP 工具协议化 |
|
||||
|
||||
## 7. 现场演示入口
|
||||
|
||||
- Demo 脚本:`mvp/demo/ten-minute-interview-demo.md`
|
||||
- 故事案例:`interview/story-cases.md`
|
||||
- 架构细节:`mvp/architecture/README.md`
|
||||
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
**更新日期**:2026-07-06
|
||||
**状态**:当前主架构 + 后续演进边界
|
||||
**关联计划**:`mvp/issues/rag-refactor-plan.md`
|
||||
**关联计划**:[`mvp/issues/active/rag-refactor-plan.md`](../issues/active/rag-refactor-plan.md)
|
||||
|
||||
## 1. 架构目标
|
||||
|
||||
|
||||
@@ -141,7 +141,8 @@ post-retrieval 层再把检索候选归一为:
|
||||
|
||||
- 给 Agent 输出 completeness hint。
|
||||
- 写入 `tool_invocation.relevance_level`。
|
||||
- 给 Verifier 构造 `tool_trace_summary`。
|
||||
- 给 Gatekeeper 提供 `evidence_refs` 引用验真源。
|
||||
- 给 Verifier 构造 `tool_trace_summary` 审计导航。
|
||||
- 供 EvaluationService 计算 evidence score。
|
||||
|
||||
## 6. 文档切片和 metadata
|
||||
@@ -178,6 +179,7 @@ flowchart LR
|
||||
LookupResult --> Recorder["ToolInvocationRecorder"]
|
||||
Recorder --> Invocation["tool_invocation"]
|
||||
Invocation --> Trace["DiagnosisTraceService"]
|
||||
Invocation --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
Invocation --> Summary["ToolTraceSummaryService"]
|
||||
Summary --> Verifier["chat_verifier"]
|
||||
Invocation --> Eval["EvaluationService / RAG eval"]
|
||||
@@ -205,9 +207,25 @@ success
|
||||
- evidence status。
|
||||
- dedup reason。
|
||||
- evidence block summaries。
|
||||
- evidence refs:`raw_path + text`,用于核对 Executor 的 `evidence_excerpt`。
|
||||
- context pack summary。
|
||||
- rerank trace。
|
||||
|
||||
其中 `evidence_refs` 是当前 Chat 证据链路的精确引用源:
|
||||
|
||||
```json
|
||||
{
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.evidence_blocks[0]",
|
||||
"text": "最小证据文本"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
如果检索返回 no evidence,应使用 `raw_path=$.no_evidence` 记录负向证据。它只能说明“本次检索没有匹配证据”,不能作为“问题不存在”的证明。
|
||||
|
||||
## 8. 去重与行动记忆
|
||||
|
||||
当前 session 级去重由 `RetrievedDocTracker` 负责。
|
||||
|
||||
@@ -1,68 +1,68 @@
|
||||
# 会话与 Trace 生命周期
|
||||
|
||||
**更新日期**:2026-07-05
|
||||
**状态**:当前可运行架构
|
||||
**更新日期**:2026-07-10
|
||||
**状态**:当前可运行架构
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/session-management.md`
|
||||
|
||||
## 1. 定位
|
||||
|
||||
旧版会话设计以 Redis 会话为主,MySQL 作为可选长期沉淀。当前 MVP 的可追踪诊断已经转为 MySQL Trace 三表为主:
|
||||
当前 MVP 把“会话态”和“运行态”拆开:
|
||||
|
||||
```text
|
||||
diagnosis_session
|
||||
-> agent_step
|
||||
-> tool_invocation
|
||||
chat_session(sessionId)
|
||||
-> diagnosis_run(runId)
|
||||
-> agent_step(runId)
|
||||
-> tool_invocation(runId)
|
||||
```
|
||||
|
||||
因此本文描述的是当前可运行链路:
|
||||
|
||||
- `sessionId` 是一次诊断和后续 trace/feedback 的关联键。
|
||||
- `diagnosis_session` 保存会话级状态、问题、答案、自评估和反馈。
|
||||
- `agent_step` 保存每个 Agent 模型调用。
|
||||
- `tool_invocation` 保存工具调用事实。
|
||||
- `DiagnosisTraceService` 聚合三类记录,形成可回放 trace。
|
||||
- `sessionId` 表示多轮会话目录和 Redis 上下文。
|
||||
- `runId` 表示一次可回放诊断执行。
|
||||
- `DiagnosisTraceService` 聚合一个 run 的主记录、步骤和工具调用,形成可回放 Trace。
|
||||
- `diagnosis_session` 只保留为历史兼容和回滚表。
|
||||
|
||||
## 2. 生命周期总图
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Start["request: chat / ai_ops"] --> Resolve["resolve sessionId"]
|
||||
Resolve --> Create["create or reset diagnosis_session"]
|
||||
Create --> Running["status = RUNNING"]
|
||||
Resolve --> Session["ensure chat_session metadata"]
|
||||
Session --> Run["create diagnosis_run(runId)"]
|
||||
Run --> Running["run.status = RUNNING"]
|
||||
|
||||
Running --> Agent["Agent workflow"]
|
||||
Agent --> StepHook["AgentLoggingHook"]
|
||||
StepHook --> Step["agent_step"]
|
||||
Agent --> Tool["Evidence tools"]
|
||||
Tool --> Invocation["tool_invocation"]
|
||||
Agent --> Context["execution context(sessionId, runId)"]
|
||||
Context --> StepHook["AgentLoggingHook"]
|
||||
StepHook --> Step["agent_step(session_id, run_id)"]
|
||||
Context --> Tool["Evidence tools"]
|
||||
Tool --> Invocation["tool_invocation(session_id, run_id)"]
|
||||
Invocation --> Gatekeeper["Gatekeeper evidence validation"]
|
||||
|
||||
Agent --> Final{"workflow result"}
|
||||
Final -->|success| Success["status = SUCCESS, answer saved"]
|
||||
Final -->|failed| Failed["status = FAILED"]
|
||||
Final -->|success| Success["run.status = SUCCESS, answer saved"]
|
||||
Final -->|failed| Failed["run.status = FAILED"]
|
||||
|
||||
Success --> Evaluation["self_evaluation merge"]
|
||||
Success --> Evaluation["diagnosis_run.self_evaluation merge"]
|
||||
Failed --> Evaluation
|
||||
Evaluation --> Trace["GET /api/diagnosis/{sessionId}/trace"]
|
||||
Success --> Feedback["POST /api/feedback"]
|
||||
Feedback --> Case["useful -> case_library"]
|
||||
Evaluation --> Trace["GET /api/diagnosis/{sessionId}/trace?runId=..."]
|
||||
Success --> Feedback["POST /api/feedback(sessionId, runId)"]
|
||||
Feedback --> Case["useful -> case_library(run_id)"]
|
||||
```
|
||||
|
||||
## 3. sessionId 规则
|
||||
## 3. ID 规则
|
||||
|
||||
| 链路 | sessionId 来源 |
|
||||
|---|---|
|
||||
| Chat | 如果请求带 sessionId,则复用;否则生成短 UUID |
|
||||
| AIOps | 如果 payload 带 sessionId,则复用;否则生成 UUID |
|
||||
| Trace | URL path 中的 `{sessionId}` |
|
||||
| Feedback | request body 中的 `sessionId` |
|
||||
| ID | 来源 | 含义 |
|
||||
|---|---|---|
|
||||
| `sessionId` | Chat request `Id`、AIOps payload `sessionId`,缺失时由服务生成 | 多轮会话目录和 Redis 上下文 |
|
||||
| `runId` | 每次有效 Chat/AIOps 执行创建 | 一次诊断运行和 Trace 回放边界 |
|
||||
|
||||
设计含义:
|
||||
|
||||
- 同一个 `sessionId` 可以贯穿诊断、trace 查询和用户反馈。
|
||||
- 当前诊断开始时会重置当前 session 的运行态字段,例如 answer、duration、step/tool count。
|
||||
- `sessionId` 是业务关联键,不依赖数据库自增 ID 暴露给外部。
|
||||
- 同一个 `sessionId` 可以贯穿多轮 Chat。
|
||||
- 每次有效 Chat/AIOps 执行都会创建新的 `runId`。
|
||||
- Trace 和 Feedback 新客户端应传 `runId`;只传 `sessionId` 时兼容解析 latest run。
|
||||
- latest run 排序使用 `diagnosis_run.created_at DESC, id DESC`,不使用 `updated_at`。
|
||||
|
||||
## 4. 状态流转
|
||||
## 4. 运行状态流转
|
||||
|
||||
```mermaid
|
||||
stateDiagram-v2
|
||||
@@ -76,14 +76,14 @@ stateDiagram-v2
|
||||
|
||||
字段边界:
|
||||
|
||||
| 字段 | 含义 |
|
||||
|---|---|
|
||||
| `status` | 执行状态:`PENDING` / `RUNNING` / `SUCCESS` / `FAILED` |
|
||||
| `answer` | Agent 最终返回给用户的报告或答复 |
|
||||
| `self_evaluation` | 系统自评估 JSON |
|
||||
| `feedback` | 用户反馈:`useful` / `not_useful` / null |
|
||||
| 字段 | 所属表 | 含义 |
|
||||
|---|---|---|
|
||||
| `status` | `diagnosis_run` | 单次运行执行状态 |
|
||||
| `answer` | `diagnosis_run` | 本次运行最终报告或答复 |
|
||||
| `self_evaluation` | `diagnosis_run` | 本次运行系统自评估 JSON |
|
||||
| `feedback` | `diagnosis_run` | 本次运行用户反馈 |
|
||||
|
||||
`feedback` 不修改 `status`。一个执行成功但用户标记 `not_useful` 的 session,仍然应该是 `SUCCESS + feedback=not_useful`。
|
||||
`feedback` 不修改 `status`。一个执行成功但用户标记 `not_useful` 的 run,仍然应该是 `SUCCESS + feedback=not_useful`。
|
||||
|
||||
## 5. agent_step 写入
|
||||
|
||||
@@ -92,105 +92,69 @@ stateDiagram-v2
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant Agent as ReactAgent
|
||||
participant Agent as Agent
|
||||
participant Hook as AgentLoggingHook
|
||||
participant DB as agent_step
|
||||
|
||||
Agent->>Hook: before_model(messages, sessionId)
|
||||
Hook->>DB: insert step_index / agent_name / model_input
|
||||
Agent-->>Agent: model call
|
||||
Agent->>Hook: after_model(messages, sessionId)
|
||||
Hook->>DB: update model_output / thought / has_tool_call / duration / token_count
|
||||
Agent->>Hook: before_model(messages, sessionId, runId)
|
||||
Hook->>DB: insert step(session_id, run_id, model_input, step_index)
|
||||
Agent->>Hook: after_model(output, sessionId, runId)
|
||||
Hook->>DB: update model_output, duration, token_count, has_tool_call
|
||||
```
|
||||
|
||||
当前记录:
|
||||
|
||||
- `session_id`
|
||||
- `step_index`
|
||||
- `agent_name`
|
||||
- `model_input`
|
||||
- `model_output`
|
||||
- `thought`
|
||||
- `has_tool_call`
|
||||
- `duration_ms`
|
||||
- `token_count`
|
||||
新写入必须带 `run_id`,同时保留 `session_id` 便于粗粒度排查。
|
||||
|
||||
## 6. tool_invocation 写入
|
||||
|
||||
工具调用记录真实工具事实,不记录模型猜测。
|
||||
|
||||
关键字段:
|
||||
工具调用记录同样通过执行上下文拿到 `sessionId + runId`:
|
||||
|
||||
```text
|
||||
session_id
|
||||
step_id
|
||||
tool_name
|
||||
input_params
|
||||
output_preview
|
||||
output_length
|
||||
retrieval_layer
|
||||
l0_match_count
|
||||
l1_match_count
|
||||
retrieval_details
|
||||
relevance_level
|
||||
dedup_reason
|
||||
duration_ms
|
||||
success
|
||||
error_message
|
||||
ToolInvocationRecorder
|
||||
-> tool_invocation.session_id
|
||||
-> tool_invocation.run_id
|
||||
-> retrieval_details / evidence_refs
|
||||
```
|
||||
|
||||
对 `lookup_knowledge`,`retrieval_details` 会承载 L0/L1、领域、证据状态、去重等检索细节。对非检索工具,检索字段可以为空。
|
||||
Verifier、Gatekeeper 和 EvaluationService 应按 `run_id` 读取工具调用,避免同一 session 的其他 run 参与评分或证据校验。
|
||||
|
||||
## 7. Trace API 聚合
|
||||
|
||||
```text
|
||||
GET /api/diagnosis/{sessionId}/trace
|
||||
GET /api/diagnosis/{sessionId}/trace?runId=run-...
|
||||
```
|
||||
|
||||
聚合逻辑:
|
||||
|
||||
```text
|
||||
diagnosis_session by sessionId
|
||||
+ agent_step ordered by step_index
|
||||
+ tool_invocation ordered by id
|
||||
diagnosis_run by sessionId + runId
|
||||
+ chat_session metadata when available
|
||||
+ agent_step where run_id = runId, ordered by the Trace API
|
||||
+ tool_invocation where run_id = runId order by id
|
||||
-> DiagnosisTraceResponse
|
||||
```
|
||||
|
||||
Trace 视图回答的问题:
|
||||
|
||||
- 这次诊断是否成功?
|
||||
- 哪些 Agent 参与了?
|
||||
- 每一步模型输入输出是什么摘要?
|
||||
- 调用了哪些工具?
|
||||
- 工具返回了什么证据?
|
||||
- Verifier / AIOps rule 是否通过?
|
||||
- 用户是否反馈有用?
|
||||
当 `runId` 缺失时,Trace API 为兼容旧客户端解析最新 run,并在响应中返回 resolved `runId`。当 `runId` 属于其他 `sessionId` 时,API 必须拒绝,不能泄漏其他会话的 Trace。
|
||||
|
||||
## 8. Chat 与 AIOps 差异
|
||||
|
||||
| 维度 | Chat | AIOps |
|
||||
|---|---|---|
|
||||
| `agent_flow` | `CHAT` | `AI_OPS` |
|
||||
| 编排方式 | `SequentialAgent`: Planner -> Executor -> Verifier | `SupervisorAgent`: Planner + Executor |
|
||||
| 编排方式 | `SequentialAgent`: Planner -> Executor -> Gatekeeper -> Verifier -> Composer | `SupervisorAgent`: Planner + Executor |
|
||||
| 自评估 | `rule_evaluation` + `verifier_evaluation` | `aiops_rule_evaluation` |
|
||||
| 答案字段 | Chat 最终答复 | 告警分析报告 |
|
||||
| payload | 用户自然语言 + history | alert payload 或 auto-discovery |
|
||||
| runId 暴露 | `/api/chat` JSON response | `/api/ai_ops` SSE metadata message |
|
||||
|
||||
## 9. 清理与边界
|
||||
|
||||
当前会话持久化边界:
|
||||
|
||||
- MySQL Trace 记录是主要可回放来源。
|
||||
- Chat 历史仍可作为请求上下文传入 Agent,但不是本文档的主持久化模型。
|
||||
- Redis 主会话存储是历史设计,不作为当前架构事实。
|
||||
- `RetrievedDocTracker` 是 session 级运行时去重状态,诊断结束后清理。
|
||||
- Redis 会话历史用于多轮上下文,不是长期审计记录。
|
||||
- MySQL `diagnosis_run + agent_step + tool_invocation` 是主要可回放来源。
|
||||
- `chat_session.expires_at` 只是目录元数据;Redis 消息历史可独立过期。
|
||||
- `RetrievedDocTracker` 仍是 session 级运行时去重状态,诊断结束后清理。
|
||||
|
||||
## 10. 后续增强
|
||||
|
||||
可考虑:
|
||||
|
||||
1. Trace API 增加更结构化的 `self_evaluation` 展示。
|
||||
2. `agent_step` 与 `tool_invocation.step_id` 建立更严格关联。
|
||||
3. 对多轮同 session 诊断增加 run id,避免复用 session 时历史记录混杂。
|
||||
4. 为 Trace 增加导出能力,服务面试演示和回归分析。
|
||||
|
||||
3. 旧 `diagnosis_session` 只读观察期结束后,再评估数据库层面的约束收紧或归档策略。
|
||||
|
||||
@@ -0,0 +1,17 @@
|
||||
# 2026-07-09 MVP 文档清理归档
|
||||
|
||||
本目录保存本次清理中从当前入口移出的历史设计材料。这些文档仍有追溯价值,但不再代表当前可运行实现。
|
||||
|
||||
## 归档内容
|
||||
|
||||
| 目录 | 内容 | 归档原因 |
|
||||
|---|---|---|
|
||||
| `discuss/` | 早期 Executor Prompt、L0、RAG 讨论稿 | 已被当前 architecture、OpenSpec change 和 issue 取代 |
|
||||
| `plan/` | `session-storage-design.md` | 会话存储已实现,当前表以 Flyway 和 `mvp/tables/` 为准 |
|
||||
| `notes/` | 早期工程决策和 Demo Trace 验收笔记 | 相关内容已沉淀到 architecture、demo、eval 和 devflow |
|
||||
|
||||
## 使用原则
|
||||
|
||||
- 当前架构以 `mvp/architecture/` 为准。
|
||||
- 当前表结构以 `mvp/tables/`、Flyway migration 和实体类为准。
|
||||
- 当前问题入口以 `mvp/issues/README.md` 为准。
|
||||
+45
-9
@@ -6,9 +6,15 @@
|
||||
|
||||
- `ten-minute-interview-demo.md`:10 分钟现场演示脚本。
|
||||
- `interview-walkthrough.md`:面试讲解话术。
|
||||
- `evidence-pipeline-scenarios.md`:PASS / LOW_CONFID / REJECT / no-evidence 场景矩阵。
|
||||
- `trace-inspection-checklist.md`:Trace 字段检查清单。
|
||||
- `scripts/run-interview-demo-check.ps1`:面试预检脚本,包含服务可达性、Chat、Trace、反馈和 summary 输出。
|
||||
- `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。
|
||||
- `interview-q-and-a.md`:面试追问回答,覆盖 Agent 工程取舍、审计和评测。
|
||||
- `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。
|
||||
- `requests/narrow-highcpu-chat.json`:窄范围正向观察请求。
|
||||
- `requests/hikari-no-evidence-chat.json`:no-evidence 负向观察请求。
|
||||
- `requests/safety-unsupported-claim-chat.json`:安全降级讨论请求。
|
||||
|
||||
## 1. 前置条件
|
||||
|
||||
@@ -33,7 +39,7 @@ http://localhost:9900
|
||||
最快方式:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1
|
||||
```
|
||||
|
||||
脚本会生成:
|
||||
@@ -42,6 +48,7 @@ powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-de
|
||||
mvp/demo/output/chat-response.json
|
||||
mvp/demo/output/trace-response.json
|
||||
mvp/demo/output/feedback-response.json
|
||||
mvp/demo/output/interview-demo-summary.json
|
||||
```
|
||||
|
||||
手动请求:
|
||||
@@ -60,10 +67,23 @@ Invoke-RestMethod `
|
||||
-Body $body
|
||||
```
|
||||
|
||||
如果要继续手动查询同一次诊断运行,先保留响应中的 run id:
|
||||
|
||||
```powershell
|
||||
$chat = Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/chat" `
|
||||
-ContentType "application/json" `
|
||||
-Body $body
|
||||
|
||||
$runId = $chat.data.runId
|
||||
```
|
||||
|
||||
期望结果:
|
||||
|
||||
- `data.success = true`
|
||||
- `data.sessionId = mvp-demo-payment-timeout-001`
|
||||
- `data.runId` 为本次诊断运行的唯一 ID
|
||||
- `data.answer` 包含诊断答复
|
||||
|
||||
## 4. 查询 Trace
|
||||
@@ -71,22 +91,27 @@ Invoke-RestMethod `
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace"
|
||||
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace?runId=$runId"
|
||||
```
|
||||
|
||||
期望结果:
|
||||
|
||||
- `code = 200`
|
||||
- `data.runId` 等于 `$runId`
|
||||
- `data.session.sessionId` 等于 Chat session id
|
||||
- `data.run.runId` 等于 `$runId`
|
||||
- `data.steps` 包含 planner / executor / verifier 等步骤
|
||||
- `data.toolInvocations` 包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具
|
||||
- `data.session.selfEvaluation` 包含 verifier 或 rule evaluation
|
||||
- Chat V2 链路中,`data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` 记录 Chat Prompt 审计版本
|
||||
- Chat V2 链路中,`data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` 记录 Gatekeeper 规则集版本
|
||||
|
||||
## 5. 提交反馈
|
||||
|
||||
```powershell
|
||||
$feedback = @{
|
||||
sessionId = $sessionId
|
||||
runId = $runId
|
||||
feedback = "useful"
|
||||
} | ConvertTo-Json
|
||||
|
||||
@@ -100,7 +125,8 @@ Invoke-RestMethod `
|
||||
期望结果:
|
||||
|
||||
- `success = true`
|
||||
- 后续 Trace 中 `data.session.feedback = useful`
|
||||
- `runId = $runId`
|
||||
- 后续精确 Trace 中 `data.session.feedback = useful`
|
||||
- useful 反馈会尝试沉淀 `case_library`
|
||||
|
||||
## 6. AIOps 告警诊断 Demo
|
||||
@@ -126,19 +152,19 @@ Invoke-WebRequest `
|
||||
|
||||
期望结果:
|
||||
|
||||
- SSE 首条包含 `session` 消息,sessionId 为 `mvp-demo-aiops-payment-cpu-001`
|
||||
- SSE 首条是 `type=metadata` 的 `message` 事件,包含 sessionId `mvp-demo-aiops-payment-cpu-001` 和本次 AIOps `runId`
|
||||
- 后续流式输出包含 AIOps 告警分析报告
|
||||
- 报告聚焦输入的 `HighCPUUsage/payment-service`
|
||||
- 同一 session 的 Trace 中 `data.session.agentFlow = AI_OPS`
|
||||
- 精确 Trace 中 `data.session.agentFlow = AI_OPS`
|
||||
- `data.session.answer` 包含最终告警报告
|
||||
- `data.toolInvocations` 包含证据工具调用
|
||||
|
||||
查询 AIOps Trace:
|
||||
查询 AIOps Trace 时优先使用 SSE metadata 中的 runId:
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace"
|
||||
-Uri "http://localhost:9900/api/diagnosis/$aiopsSessionId/trace?runId=$aiopsRunId"
|
||||
```
|
||||
|
||||
## 7. Demo 主线
|
||||
@@ -146,7 +172,7 @@ Invoke-RestMethod `
|
||||
Chat 主线:
|
||||
|
||||
```text
|
||||
一个 session id
|
||||
一个 session id + 一个 run id
|
||||
-> 用户问题
|
||||
-> 多 Agent 执行
|
||||
-> 证据工具
|
||||
@@ -159,7 +185,7 @@ Chat 主线:
|
||||
AIOps 主线:
|
||||
|
||||
```text
|
||||
一个 session id
|
||||
一个 session id + 一个 run id
|
||||
-> 告警 payload
|
||||
-> AIOps Planner / Executor
|
||||
-> 证据工具
|
||||
@@ -167,3 +193,13 @@ AIOps 主线:
|
||||
-> AIOps rule evaluation
|
||||
-> Trace API 回放
|
||||
```
|
||||
|
||||
## 8. Evidence Pipeline 场景矩阵
|
||||
|
||||
面试时不要把所有安全场景都压到 live LLM 现场表现上。建议使用:
|
||||
|
||||
- `scripts/run-interview-demo-check.ps1` 跑主路径和预检 summary。
|
||||
- `evidence-pipeline-scenarios.md` 讲解 PASS / LOW_CONFID / REJECT / no-evidence 矩阵。
|
||||
- `mvp/eval/reports/baseline-report.md` 证明固定 fixture 12/12 通过。
|
||||
|
||||
这样可以同时展示真实链路和确定性回归能力。
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
## 1. 目标
|
||||
|
||||
验证旧版 `/api/ai_ops` 入口可以作为可追踪的告警触发诊断入口,并且 payload 模式下报告聚焦输入告警。
|
||||
验证 `/api/ai_ops` 入口可以作为可追踪的告警触发诊断入口,并且 payload 模式下报告聚焦输入告警。
|
||||
|
||||
## 2. 输入
|
||||
|
||||
@@ -25,14 +25,14 @@
|
||||
|
||||
## 3. 验收标准
|
||||
|
||||
1. SSE 流输出 `session` 消息,且包含请求中的 session id。
|
||||
2. AIOps 执行创建或更新 `diagnosis_session`,并写入 `agent_flow = AI_OPS`。
|
||||
3. 持久化的 session query 包含告警名、服务名、等级、时间范围和描述。
|
||||
4. 如果生成最终报告,`diagnosis_session.answer` 包含该报告。
|
||||
5. `GET /api/diagnosis/{sessionId}/trace` 返回 AIOps session、按顺序排列的 agent steps 和 tool invocations。
|
||||
1. SSE 流首条输出 `type=metadata` 的 `message` 事件,且包含请求中的 session id 和本次 AIOps run id。
|
||||
2. AIOps 执行创建 `diagnosis_run`,并写入 `agent_flow = AI_OPS`。
|
||||
3. 持久化的 run query 包含告警名、服务名、等级、时间范围和描述。
|
||||
4. 如果生成最终报告,`diagnosis_run.answer` 包含该报告。
|
||||
5. `GET /api/diagnosis/{sessionId}/trace?runId=...` 返回 AIOps run、按顺序排列的 agent steps 和 tool invocations。
|
||||
6. payload 模式下,报告主线聚焦 `HighCPUUsage/payment-service`。
|
||||
7. 其他活跃告警最多作为相关风险或上下文出现,不应展开成完整独立根因章节。
|
||||
8. `self_evaluation.aiops_rule_evaluation` 存在,并能反映报告完整性、payload 聚焦和证据工具覆盖情况。
|
||||
8. `diagnosis_run.self_evaluation.aiops_rule_evaluation` 存在,并能反映报告完整性、payload 聚焦和证据工具覆盖情况。
|
||||
|
||||
## 4. 已知边界
|
||||
|
||||
|
||||
@@ -0,0 +1,63 @@
|
||||
# Evidence Pipeline Demo Scenarios
|
||||
|
||||
这份清单用于面试时说明 Chat 证据链路如何覆盖 `PASS`、`LOW_CONFID`、`REJECT` 和 no-evidence 场景。
|
||||
|
||||
重点区别:
|
||||
|
||||
- Live demo 证明本地服务、工具、Trace、Feedback 主链路能跑通。
|
||||
- Fixture-backed eval 证明固定安全场景可以确定性回归,不依赖 LLM 当场随机输出。
|
||||
|
||||
## Scenario Matrix
|
||||
|
||||
| 场景 | 类型 | 输入/证据 | 期望讲点 |
|
||||
|---|---|---|---|
|
||||
| Payment timeout | Live 主路径 | `requests/payment-timeout-chat.json` | 完整 Chat -> Trace -> Feedback 闭环 |
|
||||
| Narrow HighCPU observation | Live 可尝试 + fixture-backed | `requests/narrow-highcpu-chat.json` / `mvp/eval/fixtures/narrow-highcpu-observation-pass.json` | Executor 只输出观察类 claim,Gatekeeper 验引用,Verifier PASS |
|
||||
| Hikari no-evidence | Live 可尝试 + fixture-backed | `requests/hikari-no-evidence-chat.json` / `mvp/eval/fixtures/hikari-no-evidence-negative-observation-pass.json` | `$.no_evidence` 只表示本次查询无匹配证据,Composer 不说“已排除” |
|
||||
| Unsupported claim filtering | Fixture-backed | `requests/safety-unsupported-claim-chat.json` / `mvp/eval/fixtures/unsupported-claim-filtering-low-confid.json` | Verifier 将 unsupported claim 降为 LOW_CONFID,最终答案不确认“主库故障” |
|
||||
| Fabricated invocation reject | Fixture-backed | `mvp/eval/fixtures/gatekeeper-fabricated-invocation-reject.json` | Gatekeeper 拦截伪造 invocation,最终 REJECT/降级 |
|
||||
| Composer fallback | Fixture-backed | `mvp/eval/fixtures/composer-fallback-no-raw-json-low-confid.json` | 即使 Composer 输出异常,也不能把 Executor JSON 泄漏给用户 |
|
||||
|
||||
## Trace Fields To Inspect
|
||||
|
||||
| 能力 | JSON path |
|
||||
|---|---|
|
||||
| Executor V2 输出 | `data.session.selfEvaluation.verifier_evaluation.executor_structured_output` |
|
||||
| Gatekeeper 结果 | `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.status` |
|
||||
| Gatekeeper 规则版本 | `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` |
|
||||
| 证据绑定校验 | `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.checked_bindings` |
|
||||
| Verifier claim checks | `data.session.selfEvaluation.verifier_evaluation.claim_checks` |
|
||||
| Composer 输出 | `data.session.selfEvaluation.verifier_evaluation.composer_output` |
|
||||
| 工具证据引用 | `data.toolInvocations[*].retrievalDetails.evidence_refs` |
|
||||
|
||||
## How To Present It
|
||||
|
||||
```text
|
||||
我把现场 demo 和固定 eval 分开。
|
||||
现场 demo 证明系统能跑通真实链路;
|
||||
fixture-backed eval 证明反幻觉安全场景可以稳定回归。
|
||||
Gatekeeper 的规则版本也进入 trace,所以后续调整阈值或规则时可以审计。
|
||||
```
|
||||
|
||||
## Optional Live Requests
|
||||
|
||||
手动发送某个请求样例:
|
||||
|
||||
```powershell
|
||||
$body = Get-Content -Raw -Encoding UTF8 "mvp/demo/requests/narrow-highcpu-chat.json"
|
||||
Invoke-RestMethod `
|
||||
-Method Post `
|
||||
-Uri "http://localhost:9900/api/chat" `
|
||||
-ContentType "application/json" `
|
||||
-Body $body
|
||||
```
|
||||
|
||||
然后保留响应里的 `runId`,查询同一 run 的 Trace:
|
||||
|
||||
```powershell
|
||||
Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "http://localhost:9900/api/diagnosis/mvp-demo-narrow-highcpu-001/trace?runId=$runId"
|
||||
```
|
||||
|
||||
注意:除 payment-timeout 主路径外,其它 live 请求是“可尝试”的演示入口;稳定验收以 `mvp/eval` fixture 和 baseline 为准。
|
||||
@@ -0,0 +1,33 @@
|
||||
# 面试追问 Q&A
|
||||
|
||||
## 为什么不用普通 Chatbot?
|
||||
|
||||
这个项目的重点不是生成一段诊断文本,而是把诊断拆成可审计链路:Planner 拆解问题,Executor 调工具拿证据,Gatekeeper 用代码核验证据引用,Verifier 判断可推导性,Composer 生成最终表达。`sessionId` 保留多轮上下文,`runId` 精确绑定一次诊断运行,Trace 和 Feedback 都可以按 `sessionId + runId` 回放和定位。
|
||||
|
||||
## 为什么 RAG 要做成显式工具?
|
||||
|
||||
`lookup_knowledge` 保持显式工具调用,才能在 `tool_invocation` 里看到 Agent 查了什么、命中了什么、相关性等级是什么,以及最终答案是否真的使用了这些证据。隐式 Advisor 更方便,但不利于审计 Agent 决策。
|
||||
|
||||
## 怎么防止 Executor 幻觉?
|
||||
|
||||
Executor 不直接负责最终用户答案,而是输出 `executor_evidence_v2` 的微观事实和证据引用。Gatekeeper 会校验 `source_invocation_id`、`raw_path`、`evidence_excerpt` 是否真实存在;Verifier 再判断 claim 是否能由已验真的证据推出;Composer 只表达 Verifier 允许输出的内容。
|
||||
|
||||
## LOW_CONFID 是失败吗?
|
||||
|
||||
不是。`LOW_CONFID` 表示当前证据不足以支撑强结论,但系统仍然可以安全表达已确认事实和缺失信息。面试时可以把它作为“没有证据就不强答”的质量门禁,而不是模型能力失败。
|
||||
|
||||
## Prompt 改了怎么审计?
|
||||
|
||||
Chat verifier evaluation 里会记录 `prompt_audit.version`,并列出 planner、executor、verifier、composer 的 Prompt 版本和资源路径。它不保存完整 Prompt 文本,只保留用于回放和回归解释的紧凑元数据。
|
||||
|
||||
## Gatekeeper 改了怎么审计?
|
||||
|
||||
Gatekeeper 结果里记录 `gatekeeper_result.rule_set_version` 和已启用规则元数据摘要。规则执行仍是确定性 Java 代码,版本和规则元数据用于解释“这次引用验真用的是哪套规则”。
|
||||
|
||||
## 为什么现在不拆 SubAgent?
|
||||
|
||||
当前 MVP 的主要风险不是 Agent 数量不够,而是证据、验证和回归是否稳定。文档里的演进路线把 SubAgent 放在 P2:等故障类型、工具权限和评测集足够明确后再拆,避免只是移动复杂度。
|
||||
|
||||
## 为什么 baseline 比 live demo 更重要?
|
||||
|
||||
live demo 证明链路在当前环境能跑通,但 LLM 和外部依赖会波动。`mvp/eval` 的固定 fixture baseline 是确定性回归来源,用来判断 Prompt、工具、Gatekeeper、Verifier 或 Composer 的改动有没有让系统退化。
|
||||
@@ -13,7 +13,7 @@
|
||||
关键主张不是“模型回答了一次”,而是:
|
||||
|
||||
```text
|
||||
系统能展示用了什么证据、答案如何被检查、如何用 sessionId 回放整次诊断。
|
||||
系统能展示用了什么证据、答案如何被检查、如何用 sessionId + runId 精确回放这次诊断。
|
||||
```
|
||||
|
||||
## 2. Demo 流程
|
||||
@@ -23,7 +23,8 @@
|
||||
3. 打开 `mvp/demo/output/chat-response.json`。
|
||||
4. 打开 `mvp/demo/output/trace-response.json`。
|
||||
5. 指出证据工具和 verifier evaluation。
|
||||
6. 提交 feedback,并展示它挂在同一个 session 上。
|
||||
6. 提交 feedback,并展示它挂在当前 run 上。
|
||||
7. 打开 `evidence-pipeline-scenarios.md`,说明 PASS / LOW_CONFID / REJECT / no-evidence 的固定回归矩阵。
|
||||
|
||||
## 3. 命令
|
||||
|
||||
@@ -58,7 +59,7 @@ mvp/demo/output/chat-response.json
|
||||
话术:
|
||||
|
||||
```text
|
||||
这是用户看到的答案。这里的 sessionId 是稳定的,所以我后面可以追踪这一次回答是怎么来的。
|
||||
这是用户看到的答案。这里的 sessionId 是稳定的,同时响应里会返回 runId,所以我后面可以精确追踪这一次回答是怎么来的。
|
||||
```
|
||||
|
||||
### 4.2 证据 Trace
|
||||
@@ -111,7 +112,7 @@ mvp/demo/output/feedback-response.json
|
||||
话术:
|
||||
|
||||
```text
|
||||
feedback 会挂在同一个 diagnosis session 上。
|
||||
feedback 会挂在当前 diagnosis run 上。
|
||||
这让后续挖掘 useful case 或 not_useful bad case 成为可能。
|
||||
```
|
||||
|
||||
@@ -125,6 +126,14 @@ Demo 证明真实链路能跑通,offline eval baseline 证明固定 case 可
|
||||
这两者分开是有意的:Demo 面向人类审阅,eval 面向自动化信号。
|
||||
```
|
||||
|
||||
如果被问到怎么防止证据归因幻觉,可以补充:
|
||||
|
||||
```text
|
||||
Executor 的 claim 必须绑定 source_invocation_id、raw_path 和 evidence_excerpt。
|
||||
Gatekeeper 用代码核验这些引用,并把 rule_set_version 写进 trace。
|
||||
Verifier 只判断已核验证据能否推出 claim,Composer 只表达允许输出的内容。
|
||||
```
|
||||
|
||||
## 5. 强面试表达
|
||||
|
||||
```text
|
||||
@@ -141,4 +150,3 @@ traceability、evidence persistence、verifier gating、feedback 和 regression
|
||||
mvp-demo profile mock 了日志和指标,但不是完整生产运行环境。
|
||||
密钥清理、默认隔离测试和生产可靠性是后续 hardening 工作。
|
||||
```
|
||||
|
||||
|
||||
@@ -12,14 +12,16 @@
|
||||
|
||||
## 3. 验收标准
|
||||
|
||||
1. Chat 返回成功答复,且 session id 与请求一致。
|
||||
2. Trace API 返回 session 元数据、最终答案、按顺序排列的 agent steps 和 tool invocations。
|
||||
1. Chat 返回成功答复,且 session id 与请求一致,并返回本次诊断的 run id。
|
||||
2. Trace API 使用 `sessionId + runId` 返回会话元数据、运行摘要、最终答案、按顺序排列的 agent steps 和 tool invocations。
|
||||
3. Trace 中有足够证据说明用了哪些工具,以及 verifier / self-evaluation 是否已持久化。
|
||||
4. 可以使用同一个 session id 提交反馈。
|
||||
5. 后续 Trace 查询能看到已持久化的 feedback 值。
|
||||
4. 可以使用同一个 session id 和本次 run id 提交反馈。
|
||||
5. 后续精确 Trace 查询能看到已持久化的 feedback 值。
|
||||
|
||||
## 4. 需要检查的 Trace 字段
|
||||
|
||||
- `data.runId`
|
||||
- `data.run.runId`
|
||||
- `data.session.query`
|
||||
- `data.session.answer`
|
||||
- `data.session.selfEvaluation`
|
||||
|
||||
@@ -0,0 +1,4 @@
|
||||
{
|
||||
"Id": "mvp-demo-hikari-no-evidence-001",
|
||||
"Question": "确认 inventory-service 当前是否有 HikariCP 连接池耗尽日志。"
|
||||
}
|
||||
@@ -0,0 +1,4 @@
|
||||
{
|
||||
"Id": "mvp-demo-narrow-highcpu-001",
|
||||
"Question": "确认 payment-service 当前是否存在 HighCPUUsage 告警。"
|
||||
}
|
||||
@@ -0,0 +1,4 @@
|
||||
{
|
||||
"Id": "mvp-demo-safety-unsupported-001",
|
||||
"Question": "订单超时是否可以确认由数据库主库故障导致?请只基于当前证据回答。"
|
||||
}
|
||||
@@ -0,0 +1,168 @@
|
||||
param(
|
||||
[string]$BaseUrl = "http://localhost:9900",
|
||||
[string]$SessionId = "mvp-demo-interview-payment-timeout-001",
|
||||
[string]$RequestFile = "$PSScriptRoot/../requests/payment-timeout-chat.json",
|
||||
[string]$OutputDir = "$PSScriptRoot/../output"
|
||||
)
|
||||
|
||||
$ErrorActionPreference = "Stop"
|
||||
|
||||
function Test-ServiceReachable {
|
||||
param([string]$Url)
|
||||
|
||||
try {
|
||||
$request = [System.Net.WebRequest]::Create($Url)
|
||||
$request.Method = "GET"
|
||||
$request.Timeout = 5000
|
||||
$response = $request.GetResponse()
|
||||
$response.Close()
|
||||
return $true
|
||||
} catch [System.Net.WebException] {
|
||||
if ($_.Exception.Response -ne $null) {
|
||||
$_.Exception.Response.Close()
|
||||
return $true
|
||||
}
|
||||
return $false
|
||||
}
|
||||
}
|
||||
|
||||
function Get-TraceData {
|
||||
param($TraceResponse)
|
||||
|
||||
if ($TraceResponse.PSObject.Properties.Name -contains "data") {
|
||||
return $TraceResponse.data
|
||||
}
|
||||
return $TraceResponse
|
||||
}
|
||||
|
||||
function Get-SelfEvaluation {
|
||||
param($TraceData)
|
||||
|
||||
if ($null -eq $TraceData -or $null -eq $TraceData.session) {
|
||||
return $null
|
||||
}
|
||||
return $TraceData.session.selfEvaluation
|
||||
}
|
||||
|
||||
function Get-ToolNames {
|
||||
param($TraceData)
|
||||
|
||||
if ($null -eq $TraceData -or $null -eq $TraceData.toolInvocations) {
|
||||
return @()
|
||||
}
|
||||
return @($TraceData.toolInvocations | ForEach-Object { $_.toolName } | Where-Object { $_ } | Sort-Object -Unique)
|
||||
}
|
||||
|
||||
New-Item -ItemType Directory -Force -Path $OutputDir | Out-Null
|
||||
|
||||
Write-Host "Running interview demo preflight..."
|
||||
Write-Host "BaseUrl: $BaseUrl"
|
||||
Write-Host "SessionId: $SessionId"
|
||||
|
||||
if (-not (Test-ServiceReachable -Url $BaseUrl)) {
|
||||
throw "Service is not reachable: $BaseUrl. Start the app with mvp-demo profile first: mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo"
|
||||
}
|
||||
|
||||
$request = Get-Content -Raw -Encoding UTF8 -Path $RequestFile | ConvertFrom-Json
|
||||
$request.Id = $SessionId
|
||||
$body = $request | ConvertTo-Json -Depth 8
|
||||
|
||||
$chatRequest = @{
|
||||
Method = "Post"
|
||||
Uri = "$BaseUrl/api/chat"
|
||||
ContentType = "application/json; charset=utf-8"
|
||||
Body = $body
|
||||
}
|
||||
$chat = Invoke-RestMethod @chatRequest
|
||||
|
||||
$chatPath = Join-Path $OutputDir "chat-response.json"
|
||||
$chat | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $chatPath
|
||||
|
||||
$runId = $chat.data.runId
|
||||
if (-not $runId) {
|
||||
throw "Chat response did not include runId; exact trace verification cannot continue."
|
||||
}
|
||||
|
||||
$traceRequest = @{
|
||||
Method = "Get"
|
||||
Uri = "$BaseUrl/api/diagnosis/$SessionId/trace?runId=$([System.Uri]::EscapeDataString($runId))"
|
||||
}
|
||||
$trace = Invoke-RestMethod @traceRequest
|
||||
|
||||
$tracePath = Join-Path $OutputDir "trace-response.json"
|
||||
$trace | ConvertTo-Json -Depth 80 | Set-Content -Encoding UTF8 -Path $tracePath
|
||||
|
||||
$feedbackBody = @{
|
||||
sessionId = $SessionId
|
||||
runId = $runId
|
||||
feedback = "useful"
|
||||
} | ConvertTo-Json
|
||||
|
||||
$feedbackRequest = @{
|
||||
Method = "Post"
|
||||
Uri = "$BaseUrl/api/feedback"
|
||||
ContentType = "application/json; charset=utf-8"
|
||||
Body = $feedbackBody
|
||||
}
|
||||
$feedback = Invoke-RestMethod @feedbackRequest
|
||||
|
||||
$feedbackPath = Join-Path $OutputDir "feedback-response.json"
|
||||
$feedback | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $feedbackPath
|
||||
|
||||
$traceData = Get-TraceData -TraceResponse $trace
|
||||
$selfEvaluation = Get-SelfEvaluation -TraceData $traceData
|
||||
$verifierEvaluation = $null
|
||||
if ($null -ne $selfEvaluation) {
|
||||
$verifierEvaluation = $selfEvaluation.verifier_evaluation
|
||||
}
|
||||
|
||||
$gatekeeperResult = $null
|
||||
$promptAudit = $null
|
||||
if ($null -ne $verifierEvaluation) {
|
||||
$gatekeeperResult = $verifierEvaluation.gatekeeper_result
|
||||
$promptAudit = $verifierEvaluation.prompt_audit
|
||||
}
|
||||
|
||||
$verdict = $null
|
||||
$gatekeeperStatus = $null
|
||||
$gatekeeperRuleSetVersion = $null
|
||||
$promptAuditVersion = $null
|
||||
if ($null -ne $verifierEvaluation) {
|
||||
$verdict = $verifierEvaluation.verdict
|
||||
}
|
||||
if ($null -ne $gatekeeperResult) {
|
||||
$gatekeeperStatus = $gatekeeperResult.status
|
||||
$gatekeeperRuleSetVersion = $gatekeeperResult.rule_set_version
|
||||
}
|
||||
if ($null -ne $promptAudit) {
|
||||
$promptAuditVersion = $promptAudit.version
|
||||
}
|
||||
$toolNames = Get-ToolNames -TraceData $traceData
|
||||
$summaryPath = Join-Path $OutputDir "interview-demo-summary.json"
|
||||
|
||||
$summary = [ordered]@{
|
||||
sessionId = $SessionId
|
||||
runId = $runId
|
||||
baseUrl = $BaseUrl
|
||||
chatSuccess = $chat.data.success
|
||||
verdict = $verdict
|
||||
gatekeeperStatus = $gatekeeperStatus
|
||||
gatekeeperRuleSetVersion = $gatekeeperRuleSetVersion
|
||||
promptAuditVersion = $promptAuditVersion
|
||||
toolNames = $toolNames
|
||||
paths = [ordered]@{
|
||||
chat = $chatPath
|
||||
trace = $tracePath
|
||||
feedback = $feedbackPath
|
||||
summary = $summaryPath
|
||||
}
|
||||
}
|
||||
|
||||
$summary | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path $summaryPath
|
||||
|
||||
Write-Host ""
|
||||
Write-Host "Interview demo preflight completed."
|
||||
Write-Host "Verdict: $($summary.verdict)"
|
||||
Write-Host "Gatekeeper rules: $($summary.gatekeeperRuleSetVersion)"
|
||||
Write-Host "Prompt audit: $($summary.promptAuditVersion)"
|
||||
Write-Host "Summary: $summaryPath"
|
||||
@@ -26,15 +26,22 @@ $chat = Invoke-RestMethod `
|
||||
$chat | ConvertTo-Json -Depth 20 | Set-Content -Encoding UTF8 -Path "$OutputDir/chat-response.json"
|
||||
Write-Host "已保存 Chat 响应: $OutputDir/chat-response.json"
|
||||
|
||||
$runId = $chat.data.runId
|
||||
if (-not $runId) {
|
||||
throw "Chat 响应缺少 runId,无法查询精确 Trace。"
|
||||
}
|
||||
Write-Host "RunId: $runId"
|
||||
|
||||
$trace = Invoke-RestMethod `
|
||||
-Method Get `
|
||||
-Uri "$BaseUrl/api/diagnosis/$SessionId/trace"
|
||||
-Uri "$BaseUrl/api/diagnosis/$SessionId/trace?runId=$([System.Uri]::EscapeDataString($runId))"
|
||||
|
||||
$trace | ConvertTo-Json -Depth 50 | Set-Content -Encoding UTF8 -Path "$OutputDir/trace-response.json"
|
||||
Write-Host "已保存 Trace 响应: $OutputDir/trace-response.json"
|
||||
|
||||
$feedbackBody = @{
|
||||
sessionId = $SessionId
|
||||
runId = $runId
|
||||
feedback = "useful"
|
||||
} | ConvertTo-Json
|
||||
|
||||
|
||||
@@ -37,7 +37,7 @@ http://localhost:9900
|
||||
推荐使用固定脚本:
|
||||
|
||||
```powershell
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-demo.ps1
|
||||
powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-interview-demo-check.ps1
|
||||
```
|
||||
|
||||
脚本会写出:
|
||||
@@ -46,13 +46,14 @@ powershell -ExecutionPolicy Bypass -File mvp/demo/scripts/run-payment-timeout-de
|
||||
mvp/demo/output/chat-response.json
|
||||
mvp/demo/output/trace-response.json
|
||||
mvp/demo/output/feedback-response.json
|
||||
mvp/demo/output/interview-demo-summary.json
|
||||
```
|
||||
|
||||
现场话术:
|
||||
|
||||
```text
|
||||
这里我用固定 sessionId 跑一个支付接口超时问题。
|
||||
固定 sessionId 的好处是,后面 trace 和 feedback 都能关联到同一次诊断。
|
||||
固定 sessionId 的好处是保留多轮上下文;每次诊断还会返回 runId,后面 trace 和 feedback 都用这个 runId 精确关联到同一次运行。
|
||||
```
|
||||
|
||||
## 3. 展示用户答案
|
||||
@@ -67,6 +68,7 @@ mvp/demo/output/chat-response.json
|
||||
|
||||
```text
|
||||
data.sessionId
|
||||
data.runId
|
||||
data.answer
|
||||
```
|
||||
|
||||
@@ -75,7 +77,7 @@ data.answer
|
||||
```text
|
||||
这是用户看到的答案。
|
||||
但这个项目的重点不是这段文字,而是这段文字是否有证据链。
|
||||
接下来我用同一个 sessionId 查 trace。
|
||||
接下来我用同一个 sessionId 加 runId 查 trace。
|
||||
```
|
||||
|
||||
## 4. 展示 Trace
|
||||
@@ -98,6 +100,8 @@ data.toolInvocations[*].outputPreview
|
||||
data.toolInvocations[*].retrievalLayer
|
||||
data.toolInvocations[*].relevanceLevel
|
||||
data.summary.hasVerifierEvaluation
|
||||
data.session.selfEvaluation.verifier_evaluation.prompt_audit.version
|
||||
data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version
|
||||
```
|
||||
|
||||
现场话术:
|
||||
@@ -155,6 +159,10 @@ Verifier 不做新检索,只看工具 trace 汇总。
|
||||
如果 PASS,就输出原答案。
|
||||
如果 LOW_CONFID,可以补证据或加低置信提示。
|
||||
如果 REJECT,就降级输出,只保留已确认信息。
|
||||
|
||||
Prompt 和 Gatekeeper 的版本也会进入 trace。
|
||||
`prompt_audit.version` 用于说明本次 Chat 使用哪套 Prompt 契约,`gatekeeper_result.rule_set_version` 用于说明引用验真的规则版本。
|
||||
固定 fixture baseline 是回归判断来源,live demo 主要证明当前环境链路可跑通。
|
||||
```
|
||||
|
||||
## 7. 展示反馈闭环
|
||||
@@ -175,7 +183,7 @@ caseId
|
||||
现场话术:
|
||||
|
||||
```text
|
||||
用户反馈 useful 会写回同一个 diagnosis_session。
|
||||
用户反馈 useful 会写回当前 diagnosis_run。
|
||||
后端会把这次诊断自动沉淀到 case_library,后续可以做案例检索或 bad case 分析。
|
||||
|
||||
这里 status 和 feedback 是分开的:
|
||||
@@ -226,6 +234,7 @@ AIOps 有两个模式。
|
||||
mvp/demo/output/chat-response.json
|
||||
mvp/demo/output/trace-response.json
|
||||
mvp/demo/output/feedback-response.json
|
||||
mvp/demo/output/interview-demo-summary.json
|
||||
```
|
||||
|
||||
降级话术:
|
||||
|
||||
@@ -1,16 +1,20 @@
|
||||
# Trace 检查清单
|
||||
|
||||
运行 `scripts/run-payment-timeout-demo.ps1` 后,用这份清单检查 `trace-response.json`。
|
||||
运行 `scripts/run-interview-demo-check.ps1` 后,用这份清单检查 `trace-response.json` 和 `interview-demo-summary.json`。
|
||||
|
||||
## 1. Session
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
| `data.session.sessionId` | 是否等于 `mvp-demo-payment-timeout-001` | 一个 session id 串起 chat、工具、verifier、feedback 和 trace |
|
||||
| `data.runId` / `data.run.runId` | 是否等于 demo 响应中的 `runId` | `runId` 精确绑定这一次诊断运行 |
|
||||
| `data.session.sessionId` | 是否等于 `mvp-demo-payment-timeout-001` | `sessionId` 保留多轮上下文,Trace 精确回放依赖 `runId` |
|
||||
| `data.session.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
|
||||
| `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
|
||||
| `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
|
||||
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在同一次诊断上 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | 如果是 Chat V2 链路,是否记录 Gatekeeper 规则版本 | 安全规则可审计、可回归 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` | 如果是 Chat V2 链路,是否记录 Prompt 审计版本 | Prompt 变更可解释、可回归 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.prompts[*].version` | 是否记录 planner / executor / verifier / composer 版本 | 便于定位 Prompt 变更影响 |
|
||||
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在当前 diagnosis run 上 |
|
||||
|
||||
## 2. Agent 步骤
|
||||
|
||||
@@ -30,6 +34,7 @@
|
||||
| `data.toolInvocations[*].outputPreview` | 是否有受控长度的证据预览 | 保留证据但不倾倒巨大 payload |
|
||||
| `data.toolInvocations[*].success` | 是否区分成功和失败 | 工具失败对 Verifier 和 reviewer 可见 |
|
||||
| `data.toolInvocations[*].retrievalDetails` | 是否包含检索 metadata | 检索质量可事后检查 |
|
||||
| `data.toolInvocations[*].retrievalDetails.evidence_refs` | 是否包含 `raw_path + text` | Gatekeeper 可以用代码核对 Executor 引用 |
|
||||
| `data.toolInvocations[*].relevanceLevel` | 是否有相关性等级 | 可解释检索结果强弱 |
|
||||
|
||||
## 4. Summary
|
||||
@@ -44,10 +49,10 @@
|
||||
## 5. 好的结果长什么样
|
||||
|
||||
```text
|
||||
同一个 session id
|
||||
同一个 session id + run id
|
||||
-> 最终答案
|
||||
-> 持久化 agent steps
|
||||
-> 持久化 evidence tool calls
|
||||
-> verifier / self-evaluation
|
||||
-> feedback attached to the same session
|
||||
-> feedback attached to the same run
|
||||
```
|
||||
|
||||
+13
-5
@@ -29,18 +29,23 @@ The baseline evaluates saved trace fixtures. It does not start the application a
|
||||
The committed baseline currently contains:
|
||||
|
||||
```text
|
||||
8 fixed cases
|
||||
8 passing fixture evaluations
|
||||
2 PASS verdicts
|
||||
5 LOW_CONFID verdicts
|
||||
12 fixed cases
|
||||
12 passing fixture evaluations
|
||||
5 PASS verdicts
|
||||
6 LOW_CONFID verdicts
|
||||
1 REJECT verdict
|
||||
```
|
||||
|
||||
The three V2 audit-closure cases cover:
|
||||
The V2 evidence-pipeline matrix covers:
|
||||
|
||||
- Positive supported evidence for a narrow HighCPU observation.
|
||||
- No-evidence `negative_observation` using `$.no_evidence`.
|
||||
- Gatekeeper failure for a fabricated tool invocation reference.
|
||||
- Unsupported claim filtering before the final answer.
|
||||
- Composer fallback rendering without raw Executor JSON leakage.
|
||||
- Gatekeeper rule set version audit for new matrix fixtures.
|
||||
- Prompt audit version checks for planner, executor, verifier, and composer prompts.
|
||||
- Gatekeeper rule metadata checks for enabled rule id and default severity.
|
||||
|
||||
## Verification
|
||||
|
||||
@@ -76,6 +81,9 @@ Stage 5 adds these V2 checks:
|
||||
- Required V2 fixtures must include `gatekeeper_result`, `claim_checks`, and `composer_output`.
|
||||
- `claim_checks` must be structurally auditable.
|
||||
- Composer output must record whether normal parsing or fallback rendering was used.
|
||||
- Gatekeeper rule set version can be asserted per fixture.
|
||||
- Prompt audit version and per-prompt versions can be asserted per fixture.
|
||||
- Gatekeeper rule metadata can be required per fixture.
|
||||
- Final answers must not leak raw Executor protocol markers such as `executor_evidence_v2`, `answer_version`, `evidence_bindings`, or `claim_id`.
|
||||
- Configured unsupported claim keywords must not appear as confirmed final-answer content.
|
||||
|
||||
|
||||
@@ -1,4 +1,67 @@
|
||||
[
|
||||
{
|
||||
"id": "narrow-highcpu-observation",
|
||||
"title": "Narrow HighCPU observation",
|
||||
"question": "确认 payment-service 当前是否存在 HighCPUUsage 告警。",
|
||||
"traceFixture": "narrow-highcpu-observation-pass.json",
|
||||
"expectedRootCauseKeywords": ["payment-service", "HighCPUUsage", "92%"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_metrics"],
|
||||
"allowedVerdicts": ["PASS"],
|
||||
"forbiddenAnswerKeywords": ["根因", "修复建议", "通常情况下"],
|
||||
"requireV2AuditClosure": true,
|
||||
"requireClaimChecks": true,
|
||||
"requireComposerOutput": true,
|
||||
"expectedGatekeeperStatuses": ["pass"],
|
||||
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
|
||||
"expectedComposerStatuses": ["valid"],
|
||||
"forbiddenConfirmedClaimKeywords": ["数据库连接池"]
|
||||
},
|
||||
{
|
||||
"id": "prompt-gatekeeper-audit-closure",
|
||||
"title": "Prompt and Gatekeeper audit closure",
|
||||
"question": "确认 payment-service 当前是否存在 HighCPUUsage 告警,并检查审计元数据是否完整。",
|
||||
"traceFixture": "prompt-gatekeeper-audit-closure-pass.json",
|
||||
"expectedRootCauseKeywords": ["payment-service", "HighCPUUsage", "92%"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_metrics"],
|
||||
"allowedVerdicts": ["PASS"],
|
||||
"forbiddenAnswerKeywords": ["根因", "修复建议", "通常情况下"],
|
||||
"requireV2AuditClosure": true,
|
||||
"requireClaimChecks": true,
|
||||
"requireComposerOutput": true,
|
||||
"expectedGatekeeperStatuses": ["pass"],
|
||||
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
|
||||
"expectedComposerStatuses": ["valid"],
|
||||
"forbiddenConfirmedClaimKeywords": ["数据库连接池"],
|
||||
"requirePromptAudit": true,
|
||||
"expectedPromptAuditVersion": "chat-prompts-v1",
|
||||
"expectedPromptVersions": {
|
||||
"chat_planner": "chat-planner-v1",
|
||||
"chat_executor": "chat-executor-v2",
|
||||
"chat_verifier": "chat-verifier-v2",
|
||||
"chat_composer": "chat-composer-v1"
|
||||
},
|
||||
"requireGatekeeperRules": true
|
||||
},
|
||||
{
|
||||
"id": "hikari-no-evidence-negative-observation",
|
||||
"title": "Hikari no-evidence negative observation",
|
||||
"question": "确认 inventory-service 当前是否有 HikariCP 连接池耗尽日志。",
|
||||
"traceFixture": "hikari-no-evidence-negative-observation-pass.json",
|
||||
"expectedRootCauseKeywords": ["未检索到", "HikariCP", "匹配证据"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_logs"],
|
||||
"allowedVerdicts": ["PASS"],
|
||||
"forbiddenAnswerKeywords": ["已排除", "确认没有", "日志层面已排除"],
|
||||
"requireV2AuditClosure": true,
|
||||
"requireClaimChecks": true,
|
||||
"requireComposerOutput": true,
|
||||
"expectedGatekeeperStatuses": ["pass"],
|
||||
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
|
||||
"expectedComposerStatuses": ["valid"],
|
||||
"forbiddenConfirmedClaimKeywords": ["已排除 HikariCP"]
|
||||
},
|
||||
{
|
||||
"id": "payment-timeout",
|
||||
"title": "Payment API timeout",
|
||||
@@ -88,6 +151,33 @@
|
||||
"expectedComposerStatuses": ["valid"],
|
||||
"forbiddenConfirmedClaimKeywords": ["主库故障"]
|
||||
},
|
||||
{
|
||||
"id": "audit-metadata-low-confid",
|
||||
"title": "Audit metadata low confidence",
|
||||
"question": "订单超时是否可以确认由数据库主库故障导致,并检查审计元数据是否完整?",
|
||||
"traceFixture": "audit-metadata-low-confid.json",
|
||||
"expectedRootCauseKeywords": ["超时", "证据"],
|
||||
"minKeywordMatches": 2,
|
||||
"requiredEvidenceTools": ["query_logs"],
|
||||
"allowedVerdicts": ["LOW_CONFID"],
|
||||
"forbiddenAnswerKeywords": ["已经确认"],
|
||||
"requireV2AuditClosure": true,
|
||||
"requireClaimChecks": true,
|
||||
"requireComposerOutput": true,
|
||||
"expectedGatekeeperStatuses": ["pass"],
|
||||
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
|
||||
"expectedComposerStatuses": ["valid"],
|
||||
"forbiddenConfirmedClaimKeywords": ["主库故障"],
|
||||
"requirePromptAudit": true,
|
||||
"expectedPromptAuditVersion": "chat-prompts-v1",
|
||||
"expectedPromptVersions": {
|
||||
"chat_planner": "chat-planner-v1",
|
||||
"chat_executor": "chat-executor-v2",
|
||||
"chat_verifier": "chat-verifier-v2",
|
||||
"chat_composer": "chat-composer-v1"
|
||||
},
|
||||
"requireGatekeeperRules": true
|
||||
},
|
||||
{
|
||||
"id": "composer-fallback-no-raw-json",
|
||||
"title": "Composer fallback no raw JSON",
|
||||
|
||||
@@ -0,0 +1,192 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-audit-metadata-low-confid",
|
||||
"query": "订单超时是否可以确认由数据库主库故障导致,并检查审计元数据是否完整?",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 45000,
|
||||
"toolCallCount": 1,
|
||||
"answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:日志显示订单接口出现超时。\n\n仍需补充信息:目前没有数据库故障日志或主库状态证据,不能把该方向写成确认根因。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "LOW_CONFID",
|
||||
"groundedness_score": 0.42,
|
||||
"critical_fact_count": 2,
|
||||
"prompt_audit": {
|
||||
"version": "chat-prompts-v1",
|
||||
"prompts": [
|
||||
{
|
||||
"name": "chat_planner",
|
||||
"version": "chat-planner-v1",
|
||||
"resource": "prompts/chat-planner-prompt.md"
|
||||
},
|
||||
{
|
||||
"name": "chat_executor",
|
||||
"version": "chat-executor-v2",
|
||||
"resource": "prompts/chat-executor-prompt.md"
|
||||
},
|
||||
{
|
||||
"name": "chat_verifier",
|
||||
"version": "chat-verifier-v2",
|
||||
"resource": "prompts/chat-verifier-prompt.md"
|
||||
},
|
||||
{
|
||||
"name": "chat_composer",
|
||||
"version": "chat-composer-v1",
|
||||
"resource": "prompts/chat-composer-prompt.md"
|
||||
}
|
||||
]
|
||||
},
|
||||
"gatekeeper_result": {
|
||||
"status": "pass",
|
||||
"severity": "none",
|
||||
"rule_set_version": "gatekeeper-rules-v1",
|
||||
"rules": [
|
||||
{
|
||||
"id": "evidence.invocation",
|
||||
"description": "source_invocation_id must reference an existing tool invocation",
|
||||
"enabled": true,
|
||||
"default_severity": "reject"
|
||||
},
|
||||
{
|
||||
"id": "evidence.excerpt",
|
||||
"description": "evidence_excerpt must be supported by recorded evidence text",
|
||||
"enabled": true,
|
||||
"default_severity": "reject"
|
||||
}
|
||||
],
|
||||
"checked_bindings": [
|
||||
{
|
||||
"claim_id": "claim-timeout",
|
||||
"tool_name": "query_logs",
|
||||
"source_invocation_id": 22,
|
||||
"raw_path": "$.logs[0]",
|
||||
"matched_text": "order api timeout",
|
||||
"status": "pass"
|
||||
}
|
||||
],
|
||||
"failed_rules": [],
|
||||
"warnings": [],
|
||||
"errors": []
|
||||
},
|
||||
"executor_structured_output": {
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-timeout",
|
||||
"claim_type": "symptom",
|
||||
"claim_text": "订单接口出现超时",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"source_type": "tool_trace",
|
||||
"tool_name": "query_logs",
|
||||
"source_invocation_id": 22,
|
||||
"raw_path": "$.logs[0]",
|
||||
"evidence_excerpt": "order api timeout"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"claim_id": "claim-db-primary",
|
||||
"claim_type": "root_cause",
|
||||
"claim_text": "数据库主库故障导致订单超时",
|
||||
"support_level": "weak",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"source_type": "tool_trace",
|
||||
"tool_name": "query_logs",
|
||||
"source_invocation_id": 22,
|
||||
"raw_path": "$.logs[0]",
|
||||
"evidence_excerpt": "order api timeout"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "补充查询数据库主库状态和错误日志",
|
||||
"reason": "当前只有订单接口超时日志"
|
||||
}
|
||||
],
|
||||
"missing_info": ["数据库主库状态", "数据库错误日志"]
|
||||
},
|
||||
"claim_checks": [
|
||||
{
|
||||
"claim_id": "claim-timeout",
|
||||
"claim_text": "订单接口出现超时",
|
||||
"claim_type": "symptom",
|
||||
"verification": "direct_observation",
|
||||
"detail": "日志直接记录 order api timeout",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"source_invocation_id": 22,
|
||||
"raw_path": "$.logs[0]"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"claim_id": "claim-db-primary",
|
||||
"claim_text": "数据库主库故障导致订单超时",
|
||||
"claim_type": "root_cause",
|
||||
"verification": "unsupported",
|
||||
"detail": "日志只能证明订单接口超时,不能证明数据库主库故障",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"source_invocation_id": 22,
|
||||
"raw_path": "$.logs[0]"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"facts_checked": [],
|
||||
"composer_output": {
|
||||
"status": "valid",
|
||||
"answer_summary": "日志显示订单接口超时,但数据库方向证据不足。",
|
||||
"recommended_actions": [
|
||||
{
|
||||
"action_text": "补充查询数据库主库状态和错误日志",
|
||||
"reason": "当前只有订单接口超时日志"
|
||||
}
|
||||
],
|
||||
"user_facing_answer": "以下结论基于当前已获取证据,仍存在部分证据缺口,请谨慎参考。\n\n已确认信息:日志显示订单接口出现超时。\n\n仍需补充信息:目前没有数据库故障日志或主库状态证据,不能把该方向写成确认根因。"
|
||||
},
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 22,
|
||||
"sessionId": "eval-audit-metadata-low-confid",
|
||||
"toolName": "query_logs",
|
||||
"outputPreview": "order api timeout",
|
||||
"retrievalDetails": {
|
||||
"evidence_status": "supported",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.logs[0]",
|
||||
"text": "order api timeout"
|
||||
}
|
||||
]
|
||||
},
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 1,
|
||||
"returnedToolCallCount": 1,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,126 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-hikari-no-evidence-negative-observation",
|
||||
"query": "确认 inventory-service 当前是否有 HikariCP 连接池耗尽日志。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 21000,
|
||||
"toolCallCount": 1,
|
||||
"answer": "本次查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志;这只表示当前查询没有匹配证据,仍不能据此判断系统一定健康。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 1.0,
|
||||
"critical_fact_count": 1,
|
||||
"gatekeeper_result": {
|
||||
"status": "pass",
|
||||
"severity": "none",
|
||||
"rule_set_version": "gatekeeper-rules-v1",
|
||||
"rules": [
|
||||
{
|
||||
"id": "evidence.raw_path",
|
||||
"description": "raw_path must exist in retrieval_details.evidence_refs",
|
||||
"enabled": true,
|
||||
"default_severity": "reject"
|
||||
}
|
||||
],
|
||||
"checked_bindings": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"tool_name": "query_logs",
|
||||
"source_invocation_id": 12,
|
||||
"raw_path": "$.no_evidence",
|
||||
"matched_text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志",
|
||||
"status": "pass"
|
||||
}
|
||||
],
|
||||
"failed_rules": [],
|
||||
"warnings": [],
|
||||
"errors": []
|
||||
},
|
||||
"executor_structured_output": {
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "negative_observation",
|
||||
"claim_text": "当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"source_type": "tool_trace",
|
||||
"tool_name": "query_logs",
|
||||
"source_invocation_id": 12,
|
||||
"raw_path": "$.no_evidence",
|
||||
"evidence_excerpt": "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [],
|
||||
"missing_info": [
|
||||
"仅查询了 application-logs 中 inventory-service HikariCP 相关日志"
|
||||
]
|
||||
},
|
||||
"claim_checks": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_text": "当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。",
|
||||
"claim_type": "negative_observation",
|
||||
"verification": "direct_observation",
|
||||
"detail": "$.no_evidence 只支持本次查询未检索到匹配证据。",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"source_invocation_id": 12,
|
||||
"raw_path": "$.no_evidence"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"facts_checked": [],
|
||||
"composer_output": {
|
||||
"status": "valid",
|
||||
"answer_summary": "本次查询未检索到匹配日志。",
|
||||
"recommended_actions": [],
|
||||
"user_facing_answer": "本次查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志;这只表示当前查询没有匹配证据,仍不能据此判断系统一定健康。"
|
||||
},
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"success": true,
|
||||
"source_invocation_ids": [12],
|
||||
"evidence_level": "no_evidence"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 12,
|
||||
"sessionId": "eval-hikari-no-evidence-negative-observation",
|
||||
"toolName": "query_logs",
|
||||
"outputPreview": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志",
|
||||
"retrievalDetails": {
|
||||
"evidence_status": "no_evidence",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.no_evidence",
|
||||
"text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志"
|
||||
}
|
||||
]
|
||||
},
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 1,
|
||||
"returnedToolCallCount": 1,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,124 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-narrow-highcpu-observation",
|
||||
"query": "确认 payment-service 当前是否存在 HighCPUUsage 告警。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 18000,
|
||||
"toolCallCount": 1,
|
||||
"answer": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 1.0,
|
||||
"critical_fact_count": 1,
|
||||
"gatekeeper_result": {
|
||||
"status": "pass",
|
||||
"severity": "none",
|
||||
"rule_set_version": "gatekeeper-rules-v1",
|
||||
"rules": [
|
||||
{
|
||||
"id": "evidence.raw_path",
|
||||
"description": "raw_path must exist in retrieval_details.evidence_refs",
|
||||
"enabled": true,
|
||||
"default_severity": "reject"
|
||||
}
|
||||
],
|
||||
"checked_bindings": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"tool_name": "query_metrics",
|
||||
"source_invocation_id": 11,
|
||||
"raw_path": "$.alerts[0]",
|
||||
"matched_text": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m",
|
||||
"status": "pass"
|
||||
}
|
||||
],
|
||||
"failed_rules": [],
|
||||
"warnings": [],
|
||||
"errors": []
|
||||
},
|
||||
"executor_structured_output": {
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "observation",
|
||||
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"source_type": "tool_trace",
|
||||
"tool_name": "query_metrics",
|
||||
"source_invocation_id": 11,
|
||||
"raw_path": "$.alerts[0]",
|
||||
"evidence_excerpt": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [],
|
||||
"missing_info": []
|
||||
},
|
||||
"claim_checks": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
|
||||
"claim_type": "observation",
|
||||
"verification": "direct_observation",
|
||||
"detail": "已核验的指标证据直接包含服务名、告警名和 CPU 当前值。",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"source_invocation_id": 11,
|
||||
"raw_path": "$.alerts[0]"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"facts_checked": [],
|
||||
"composer_output": {
|
||||
"status": "valid",
|
||||
"answer_summary": "payment-service 当前存在 HighCPUUsage 告警。",
|
||||
"recommended_actions": [],
|
||||
"user_facing_answer": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。"
|
||||
},
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_metrics",
|
||||
"success": true,
|
||||
"source_invocation_ids": [11],
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 11,
|
||||
"sessionId": "eval-narrow-highcpu-observation",
|
||||
"toolName": "query_metrics",
|
||||
"outputPreview": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m",
|
||||
"retrievalDetails": {
|
||||
"evidence_status": "supported",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.alerts[0]",
|
||||
"text": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m"
|
||||
}
|
||||
]
|
||||
},
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 1,
|
||||
"returnedToolCallCount": 1,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,155 @@
|
||||
{
|
||||
"session": {
|
||||
"sessionId": "eval-prompt-gatekeeper-audit-closure",
|
||||
"query": "确认 payment-service 当前是否存在 HighCPUUsage 告警,并检查审计元数据是否完整。",
|
||||
"status": "SUCCESS",
|
||||
"agentFlow": "CHAT",
|
||||
"totalDurationMs": 19000,
|
||||
"toolCallCount": 1,
|
||||
"answer": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
|
||||
"selfEvaluation": {
|
||||
"verifier_evaluation": {
|
||||
"verdict": "PASS",
|
||||
"groundedness_score": 1.0,
|
||||
"critical_fact_count": 1,
|
||||
"prompt_audit": {
|
||||
"version": "chat-prompts-v1",
|
||||
"prompts": [
|
||||
{
|
||||
"name": "chat_planner",
|
||||
"version": "chat-planner-v1",
|
||||
"resource": "prompts/chat-planner-prompt.md"
|
||||
},
|
||||
{
|
||||
"name": "chat_executor",
|
||||
"version": "chat-executor-v2",
|
||||
"resource": "prompts/chat-executor-prompt.md"
|
||||
},
|
||||
{
|
||||
"name": "chat_verifier",
|
||||
"version": "chat-verifier-v2",
|
||||
"resource": "prompts/chat-verifier-prompt.md"
|
||||
},
|
||||
{
|
||||
"name": "chat_composer",
|
||||
"version": "chat-composer-v1",
|
||||
"resource": "prompts/chat-composer-prompt.md"
|
||||
}
|
||||
]
|
||||
},
|
||||
"gatekeeper_result": {
|
||||
"status": "pass",
|
||||
"severity": "none",
|
||||
"rule_set_version": "gatekeeper-rules-v1",
|
||||
"rules": [
|
||||
{
|
||||
"id": "evidence.invocation",
|
||||
"description": "source_invocation_id must reference an existing tool invocation",
|
||||
"enabled": true,
|
||||
"default_severity": "reject"
|
||||
},
|
||||
{
|
||||
"id": "evidence.raw_path",
|
||||
"description": "raw_path must exist in retrieval_details.evidence_refs",
|
||||
"enabled": true,
|
||||
"default_severity": "reject"
|
||||
}
|
||||
],
|
||||
"checked_bindings": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"tool_name": "query_metrics",
|
||||
"source_invocation_id": 21,
|
||||
"raw_path": "$.alerts[0]",
|
||||
"matched_text": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m",
|
||||
"status": "pass"
|
||||
}
|
||||
],
|
||||
"failed_rules": [],
|
||||
"warnings": [],
|
||||
"errors": []
|
||||
},
|
||||
"executor_structured_output": {
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_type": "observation",
|
||||
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
|
||||
"support_level": "direct",
|
||||
"evidence_bindings": [
|
||||
{
|
||||
"source_type": "tool_trace",
|
||||
"tool_name": "query_metrics",
|
||||
"source_invocation_id": 21,
|
||||
"raw_path": "$.alerts[0]",
|
||||
"evidence_excerpt": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"hypotheses": [],
|
||||
"recommended_actions": [],
|
||||
"missing_info": []
|
||||
},
|
||||
"claim_checks": [
|
||||
{
|
||||
"claim_id": "claim-1",
|
||||
"claim_text": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。",
|
||||
"claim_type": "observation",
|
||||
"verification": "direct_observation",
|
||||
"detail": "已核验的指标证据直接包含服务名、告警名和 CPU 当前值。",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"source_invocation_id": 21,
|
||||
"raw_path": "$.alerts[0]"
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"facts_checked": [],
|
||||
"composer_output": {
|
||||
"status": "valid",
|
||||
"answer_summary": "payment-service 当前存在 HighCPUUsage 告警。",
|
||||
"recommended_actions": [],
|
||||
"user_facing_answer": "payment-service 当前存在 HighCPUUsage 告警,CPU 使用率为 92%。"
|
||||
},
|
||||
"tool_trace_summary": [
|
||||
{
|
||||
"tool_name": "query_metrics",
|
||||
"success": true,
|
||||
"source_invocation_ids": [21],
|
||||
"evidence_level": "direct"
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"steps": [],
|
||||
"toolInvocations": [
|
||||
{
|
||||
"id": 21,
|
||||
"sessionId": "eval-prompt-gatekeeper-audit-closure",
|
||||
"toolName": "query_metrics",
|
||||
"outputPreview": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m",
|
||||
"retrievalDetails": {
|
||||
"evidence_status": "supported",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.alerts[0]",
|
||||
"text": "HighCPUUsage firing, service=payment-service, current=92%, duration=25m"
|
||||
}
|
||||
]
|
||||
},
|
||||
"success": true
|
||||
}
|
||||
],
|
||||
"summary": {
|
||||
"persistedStepCount": 3,
|
||||
"returnedStepCount": 3,
|
||||
"persistedToolCallCount": 1,
|
||||
"returnedToolCallCount": 1,
|
||||
"hasVerifierEvaluation": true,
|
||||
"hasFeedback": false
|
||||
}
|
||||
}
|
||||
@@ -1,15 +1,72 @@
|
||||
{
|
||||
"totalCases" : 8,
|
||||
"passedCases" : 8,
|
||||
"totalCases" : 12,
|
||||
"passedCases" : 12,
|
||||
"passRate" : 1.0,
|
||||
"verdictDistribution" : {
|
||||
"PASS" : 2,
|
||||
"LOW_CONFID" : 5,
|
||||
"PASS" : 5,
|
||||
"LOW_CONFID" : 6,
|
||||
"REJECT" : 1
|
||||
},
|
||||
"averageToolCallCount" : 1.625,
|
||||
"averageDurationMs" : 44875.0,
|
||||
"averageToolCallCount" : 1.4166666666666667,
|
||||
"averageDurationMs" : 38500.0,
|
||||
"results" : [ {
|
||||
"caseId" : "narrow-highcpu-observation",
|
||||
"title" : "Narrow HighCPU observation",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "PASS",
|
||||
"matchedKeywordCount" : 3,
|
||||
"requiredKeywordCount" : 3,
|
||||
"evidenceCoverage" : {
|
||||
"query_metrics" : true
|
||||
},
|
||||
"gatekeeperStatus" : "pass",
|
||||
"gatekeeperRuleSetVersion" : "gatekeeper-rules-v1",
|
||||
"promptAuditVersion" : null,
|
||||
"composerStatus" : "valid",
|
||||
"claimCheckCount" : 1,
|
||||
"gatekeeperRuleCount" : 1,
|
||||
"toolCallCount" : 1,
|
||||
"durationMs" : 18000
|
||||
}, {
|
||||
"caseId" : "prompt-gatekeeper-audit-closure",
|
||||
"title" : "Prompt and Gatekeeper audit closure",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "PASS",
|
||||
"matchedKeywordCount" : 3,
|
||||
"requiredKeywordCount" : 3,
|
||||
"evidenceCoverage" : {
|
||||
"query_metrics" : true
|
||||
},
|
||||
"gatekeeperStatus" : "pass",
|
||||
"gatekeeperRuleSetVersion" : "gatekeeper-rules-v1",
|
||||
"promptAuditVersion" : "chat-prompts-v1",
|
||||
"composerStatus" : "valid",
|
||||
"claimCheckCount" : 1,
|
||||
"gatekeeperRuleCount" : 2,
|
||||
"toolCallCount" : 1,
|
||||
"durationMs" : 19000
|
||||
}, {
|
||||
"caseId" : "hikari-no-evidence-negative-observation",
|
||||
"title" : "Hikari no-evidence negative observation",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "PASS",
|
||||
"matchedKeywordCount" : 3,
|
||||
"requiredKeywordCount" : 3,
|
||||
"evidenceCoverage" : {
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : "pass",
|
||||
"gatekeeperRuleSetVersion" : "gatekeeper-rules-v1",
|
||||
"promptAuditVersion" : null,
|
||||
"composerStatus" : "valid",
|
||||
"claimCheckCount" : 1,
|
||||
"gatekeeperRuleCount" : 1,
|
||||
"toolCallCount" : 1,
|
||||
"durationMs" : 21000
|
||||
}, {
|
||||
"caseId" : "payment-timeout",
|
||||
"title" : "Payment API timeout",
|
||||
"passed" : true,
|
||||
@@ -23,8 +80,11 @@
|
||||
"query_metrics" : true
|
||||
},
|
||||
"gatekeeperStatus" : null,
|
||||
"gatekeeperRuleSetVersion" : null,
|
||||
"promptAuditVersion" : null,
|
||||
"composerStatus" : null,
|
||||
"claimCheckCount" : null,
|
||||
"gatekeeperRuleCount" : null,
|
||||
"toolCallCount" : 3,
|
||||
"durationMs" : 42000
|
||||
}, {
|
||||
@@ -40,8 +100,11 @@
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : null,
|
||||
"gatekeeperRuleSetVersion" : null,
|
||||
"promptAuditVersion" : null,
|
||||
"composerStatus" : null,
|
||||
"claimCheckCount" : null,
|
||||
"gatekeeperRuleCount" : null,
|
||||
"toolCallCount" : 2,
|
||||
"durationMs" : 51000
|
||||
}, {
|
||||
@@ -56,8 +119,11 @@
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : null,
|
||||
"gatekeeperRuleSetVersion" : null,
|
||||
"promptAuditVersion" : null,
|
||||
"composerStatus" : null,
|
||||
"claimCheckCount" : null,
|
||||
"gatekeeperRuleCount" : null,
|
||||
"toolCallCount" : 1,
|
||||
"durationMs" : 36000
|
||||
}, {
|
||||
@@ -73,8 +139,11 @@
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : null,
|
||||
"gatekeeperRuleSetVersion" : null,
|
||||
"promptAuditVersion" : null,
|
||||
"composerStatus" : null,
|
||||
"claimCheckCount" : null,
|
||||
"gatekeeperRuleCount" : null,
|
||||
"toolCallCount" : 2,
|
||||
"durationMs" : 47000
|
||||
}, {
|
||||
@@ -90,8 +159,11 @@
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : null,
|
||||
"gatekeeperRuleSetVersion" : null,
|
||||
"promptAuditVersion" : null,
|
||||
"composerStatus" : null,
|
||||
"claimCheckCount" : null,
|
||||
"gatekeeperRuleCount" : null,
|
||||
"toolCallCount" : 2,
|
||||
"durationMs" : 53000
|
||||
}, {
|
||||
@@ -106,8 +178,11 @@
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : "fail",
|
||||
"gatekeeperRuleSetVersion" : null,
|
||||
"promptAuditVersion" : null,
|
||||
"composerStatus" : "valid",
|
||||
"claimCheckCount" : 1,
|
||||
"gatekeeperRuleCount" : null,
|
||||
"toolCallCount" : 1,
|
||||
"durationMs" : 39000
|
||||
}, {
|
||||
@@ -122,10 +197,32 @@
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : "pass",
|
||||
"gatekeeperRuleSetVersion" : null,
|
||||
"promptAuditVersion" : null,
|
||||
"composerStatus" : "valid",
|
||||
"claimCheckCount" : 2,
|
||||
"gatekeeperRuleCount" : null,
|
||||
"toolCallCount" : 1,
|
||||
"durationMs" : 44000
|
||||
}, {
|
||||
"caseId" : "audit-metadata-low-confid",
|
||||
"title" : "Audit metadata low confidence",
|
||||
"passed" : true,
|
||||
"failedChecks" : [ ],
|
||||
"verdict" : "LOW_CONFID",
|
||||
"matchedKeywordCount" : 2,
|
||||
"requiredKeywordCount" : 2,
|
||||
"evidenceCoverage" : {
|
||||
"query_logs" : true
|
||||
},
|
||||
"gatekeeperStatus" : "pass",
|
||||
"gatekeeperRuleSetVersion" : "gatekeeper-rules-v1",
|
||||
"promptAuditVersion" : "chat-prompts-v1",
|
||||
"composerStatus" : "valid",
|
||||
"claimCheckCount" : 2,
|
||||
"gatekeeperRuleCount" : 2,
|
||||
"toolCallCount" : 1,
|
||||
"durationMs" : 45000
|
||||
}, {
|
||||
"caseId" : "composer-fallback-no-raw-json",
|
||||
"title" : "Composer fallback no raw JSON",
|
||||
@@ -138,9 +235,12 @@
|
||||
"query_metrics" : true
|
||||
},
|
||||
"gatekeeperStatus" : "pass",
|
||||
"gatekeeperRuleSetVersion" : null,
|
||||
"promptAuditVersion" : null,
|
||||
"composerStatus" : "composer_malformed",
|
||||
"claimCheckCount" : 2,
|
||||
"gatekeeperRuleCount" : null,
|
||||
"toolCallCount" : 1,
|
||||
"durationMs" : 47000
|
||||
} ]
|
||||
}
|
||||
}
|
||||
@@ -1,26 +1,30 @@
|
||||
# Diagnosis Eval Report
|
||||
|
||||
- Total cases: 8
|
||||
- Passed cases: 8
|
||||
- Total cases: 12
|
||||
- Passed cases: 12
|
||||
- Pass rate: 100.00%
|
||||
- Average tool calls: 1.63
|
||||
- Average duration ms: 44875.00
|
||||
- Average tool calls: 1.42
|
||||
- Average duration ms: 38500.00
|
||||
|
||||
## Verdict Distribution
|
||||
|
||||
- PASS: 2
|
||||
- LOW_CONFID: 5
|
||||
- PASS: 5
|
||||
- LOW_CONFID: 6
|
||||
- REJECT: 1
|
||||
|
||||
## Cases
|
||||
|
||||
| Case | Result | Verdict | Gatekeeper | Composer | Claim Checks | Keywords | Tool Calls | Duration ms | Failed Checks |
|
||||
| --- | --- | --- | --- | --- | ---: | --- | ---: | ---: | --- |
|
||||
| payment-timeout | PASS | PASS | - | - | - | 3/3 | 3 | 42000 | - |
|
||||
| mysql-pool-exhausted | PASS | LOW_CONFID | - | - | - | 3/3 | 2 | 51000 | - |
|
||||
| redis-timeout | PASS | LOW_CONFID | - | - | - | 2/2 | 1 | 36000 | - |
|
||||
| slow-response | PASS | PASS | - | - | - | 2/2 | 2 | 47000 | - |
|
||||
| jvm-memory-risk | PASS | LOW_CONFID | - | - | - | 3/3 | 2 | 53000 | - |
|
||||
| gatekeeper-fabricated-invocation | PASS | REJECT | fail | valid | 1 | 3/3 | 1 | 39000 | - |
|
||||
| unsupported-claim-filtering | PASS | LOW_CONFID | pass | valid | 2 | 2/2 | 1 | 44000 | - |
|
||||
| composer-fallback-no-raw-json | PASS | LOW_CONFID | pass | composer_malformed | 2 | 2/2 | 1 | 47000 | - |
|
||||
| Case | Result | Verdict | Gatekeeper | Rule Set | Prompt Audit | Composer | Claim Checks | Rules | Keywords | Tool Calls | Duration ms | Failed Checks |
|
||||
| --- | --- | --- | --- | --- | --- | --- | ---: | ---: | --- | ---: | ---: | --- |
|
||||
| narrow-highcpu-observation | PASS | PASS | pass | gatekeeper-rules-v1 | - | valid | 1 | 1 | 3/3 | 1 | 18000 | - |
|
||||
| prompt-gatekeeper-audit-closure | PASS | PASS | pass | gatekeeper-rules-v1 | chat-prompts-v1 | valid | 1 | 2 | 3/3 | 1 | 19000 | - |
|
||||
| hikari-no-evidence-negative-observation | PASS | PASS | pass | gatekeeper-rules-v1 | - | valid | 1 | 1 | 3/3 | 1 | 21000 | - |
|
||||
| payment-timeout | PASS | PASS | - | - | - | - | - | - | 3/3 | 3 | 42000 | - |
|
||||
| mysql-pool-exhausted | PASS | LOW_CONFID | - | - | - | - | - | - | 3/3 | 2 | 51000 | - |
|
||||
| redis-timeout | PASS | LOW_CONFID | - | - | - | - | - | - | 2/2 | 1 | 36000 | - |
|
||||
| slow-response | PASS | PASS | - | - | - | - | - | - | 2/2 | 2 | 47000 | - |
|
||||
| jvm-memory-risk | PASS | LOW_CONFID | - | - | - | - | - | - | 3/3 | 2 | 53000 | - |
|
||||
| gatekeeper-fabricated-invocation | PASS | REJECT | fail | - | - | valid | 1 | - | 3/3 | 1 | 39000 | - |
|
||||
| unsupported-claim-filtering | PASS | LOW_CONFID | pass | - | - | valid | 2 | - | 2/2 | 1 | 44000 | - |
|
||||
| audit-metadata-low-confid | PASS | LOW_CONFID | pass | gatekeeper-rules-v1 | chat-prompts-v1 | valid | 2 | 2 | 2/2 | 1 | 45000 | - |
|
||||
| composer-fallback-no-raw-json | PASS | LOW_CONFID | pass | - | - | composer_malformed | 2 | - | 2/2 | 1 | 47000 | - |
|
||||
|
||||
+28
-1
@@ -32,8 +32,18 @@ baseline report:整套固定集当前认可的结果
|
||||
"requireClaimChecks": true,
|
||||
"requireComposerOutput": true,
|
||||
"expectedGatekeeperStatuses": ["pass"],
|
||||
"expectedGatekeeperRuleSetVersion": "gatekeeper-rules-v1",
|
||||
"expectedComposerStatuses": ["valid"],
|
||||
"forbiddenConfirmedClaimKeywords": ["主库故障"]
|
||||
"forbiddenConfirmedClaimKeywords": ["主库故障"],
|
||||
"requirePromptAudit": true,
|
||||
"expectedPromptAuditVersion": "chat-prompts-v1",
|
||||
"expectedPromptVersions": {
|
||||
"chat_planner": "chat-planner-v1",
|
||||
"chat_executor": "chat-executor-v2",
|
||||
"chat_verifier": "chat-verifier-v2",
|
||||
"chat_composer": "chat-composer-v1"
|
||||
},
|
||||
"requireGatekeeperRules": true
|
||||
}
|
||||
```
|
||||
|
||||
@@ -54,8 +64,13 @@ baseline report:整套固定集当前认可的结果
|
||||
| `requireClaimChecks` | 是否要求 `claim_checks` | 要求 claim check 数组存在且非空 |
|
||||
| `requireComposerOutput` | 是否要求 `composer_output` | 要求 Composer 审计存在并带 `status` |
|
||||
| `expectedGatekeeperStatuses` | 允许的 Gatekeeper 状态 | 实际 `gatekeeper_result.status` 不在列表中则失败 |
|
||||
| `expectedGatekeeperRuleSetVersion` | 期望的 Gatekeeper 规则集版本 | 配置后校验 `gatekeeper_result.rule_set_version` |
|
||||
| `expectedComposerStatuses` | 允许的 Composer 状态 | 实际 `composer_output.status` 不在列表中则失败 |
|
||||
| `forbiddenConfirmedClaimKeywords` | 不得进入最终答案的未支持结论关键词 | 用于证明 unsupported/external_unknown claim 被过滤 |
|
||||
| `requirePromptAudit` | 是否要求 Prompt 审计元数据 | 要求 `prompt_audit.version` 存在 |
|
||||
| `expectedPromptAuditVersion` | 期望的 Prompt 审计目录版本 | 配置后校验 `prompt_audit.version` |
|
||||
| `expectedPromptVersions` | 期望的各角色 Prompt 版本 | 校验 `prompt_audit.prompts[*].name/version` |
|
||||
| `requireGatekeeperRules` | 是否要求 Gatekeeper 规则元数据 | 要求 `gatekeeper_result.rules` 非空,且每条规则有 `id`、`enabled`、`default_severity` |
|
||||
|
||||
## 2. Trace Fixture
|
||||
|
||||
@@ -69,6 +84,10 @@ fixture 是一次 Agent 运行后的 trace 快照。评测器只读取当前规
|
||||
| `session.totalDurationMs` | 运行耗时 | 进入报告 |
|
||||
| `session.selfEvaluation.verifier_evaluation.verdict` | Verifier 判定 | 必须存在并符合 case 的 `allowedVerdicts` |
|
||||
| `session.selfEvaluation.verifier_evaluation.gatekeeper_result.status` | Gatekeeper 结果 | V2 case 必须存在;`fail` 不允许搭配 `PASS` |
|
||||
| `session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | Gatekeeper 规则集版本 | 新矩阵 case 可显式断言该版本 |
|
||||
| `session.selfEvaluation.verifier_evaluation.gatekeeper_result.rules` | Gatekeeper 规则元数据摘要 | 审计 case 可要求规则列表非空且字段完整 |
|
||||
| `session.selfEvaluation.verifier_evaluation.prompt_audit.version` | Chat Prompt 审计目录版本 | 审计 case 可显式断言该版本 |
|
||||
| `session.selfEvaluation.verifier_evaluation.prompt_audit.prompts` | 各 Chat Prompt 名称、版本和资源路径 | 审计 case 可断言 planner、executor、verifier、composer 版本 |
|
||||
| `session.selfEvaluation.verifier_evaluation.claim_checks` | Verifier V2 claim 级校验 | V2 case 必须存在;每项需要 `claim_id`、`verification`、`detail` |
|
||||
| `session.selfEvaluation.verifier_evaluation.composer_output.status` | Composer 渲染状态 | V2 case 必须存在;记录 `valid`、`composer_malformed` 等 |
|
||||
| `session.selfEvaluation.verifier_evaluation.executor_structured_output.claims[*].evidence_bindings` | Executor claim 证据绑定 | 如果结构化输出存在,每条 claim 需要证据绑定 |
|
||||
@@ -91,8 +110,11 @@ Java 类型:`DiagnosisEvalResult`
|
||||
| `requiredKeywordCount` | case 配置的关键词数量 |
|
||||
| `evidenceCoverage` | 每个必需工具是否出现 |
|
||||
| `gatekeeperStatus` | 读到的 `gatekeeper_result.status` |
|
||||
| `gatekeeperRuleSetVersion` | 读到的 `gatekeeper_result.rule_set_version` |
|
||||
| `promptAuditVersion` | 读到的 `prompt_audit.version` |
|
||||
| `composerStatus` | 读到的 `composer_output.status` |
|
||||
| `claimCheckCount` | `claim_checks` 数量 |
|
||||
| `gatekeeperRuleCount` | `gatekeeper_result.rules` 数量 |
|
||||
| `toolCallCount` | trace 中工具调用总数 |
|
||||
| `durationMs` | trace 总耗时 |
|
||||
|
||||
@@ -125,6 +147,11 @@ Executor structured output
|
||||
|
||||
- V2 case 必须有 `gatekeeper_result`、`claim_checks`、`composer_output`。
|
||||
- `gatekeeper_result.status = fail` 时,Verifier verdict 不能是 `PASS`。
|
||||
- 配置 `expectedGatekeeperRuleSetVersion` 的 case 必须匹配 `gatekeeper_result.rule_set_version`。
|
||||
- 配置 `requirePromptAudit` 的 case 必须包含 `prompt_audit.version`。
|
||||
- 配置 `expectedPromptAuditVersion` 的 case 必须匹配 `prompt_audit.version`。
|
||||
- 配置 `expectedPromptVersions` 的 case 必须能在 `prompt_audit.prompts` 中找到对应角色和版本。
|
||||
- 配置 `requireGatekeeperRules` 的 case 必须包含非空 `gatekeeper_result.rules`,且每条规则有 `id`、`enabled`、`default_severity`。
|
||||
- `claim_checks[*].verification` 只能是 `direct_observation`、`reasonable_inference`、`overstated`、`unsupported`、`external_unknown`、`contradicted`。
|
||||
- Composer 输出必须记录 `status`。
|
||||
- 最终答案不能泄漏 `executor_evidence_v2`、`answer_version`、`evidence_bindings`、`claim_id`。
|
||||
|
||||
+56
-36
@@ -1,45 +1,65 @@
|
||||
# 已知问题记录
|
||||
# MVP Issues 索引
|
||||
|
||||
| # | 标题 | 严重程度 | 状态 | 文件 |
|
||||
|---|---|---|---|---|
|
||||
| ISS-001 | Executor 重复召回同一文档 | 中 | 已修复 | [ISS-001-duplicate-retrieval.md](ISS-001-duplicate-retrieval.md) |
|
||||
| ISS-002 | Executor 无约束重复调用 lookup_knowledge | 中 | 已修复 | [ISS-002-executor-unconstrained-lookup.md](ISS-002-executor-unconstrained-lookup.md) |
|
||||
| ISS-003 | MVP 设计与实现 Review 收敛 | 高 | 待规划 | [ISS-003-mvp-design-implementation-review.md](ISS-003-mvp-design-implementation-review.md) |
|
||||
| ISS-004 | Executor 域级检索水位控制(Phase 2) | 低 | 待规划 | [ISS-004-executor-domain-hard-limit.md](ISS-004-executor-domain-hard-limit.md) |
|
||||
| ISS-005 | 证据链补齐与降级契约收敛 | 高 | 已归档 | [ISS-005-evidence-trace-hardening.md](ISS-005-evidence-trace-hardening.md) |
|
||||
| ISS-006 | 固定诊断评测集与回归 Harness | 高 | 已归档 | [ISS-006-diagnosis-eval-harness.md](ISS-006-diagnosis-eval-harness.md) |
|
||||
| ISS-007 | Verifier 证据摘要保真与工具命中质量问题 | 高 | 已实施 | [ISS-007-verifier-evidence-summary-fidelity.md](ISS-007-verifier-evidence-summary-fidelity.md) |
|
||||
| executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [executor-evidence-attribution-hallucination.md](executor-evidence-attribution-hallucination.md) |
|
||||
| executor-self-evidence-loop-design-note | Executor 自证循环与证据摘要链路设计记录 | 高 | 已形成方向 | [executor-self-evidence-loop-design-note.md](executor-self-evidence-loop-design-note.md) |
|
||||
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 中 | 已归档 | [expand-diagnosis-eval-fixtures.md](expand-diagnosis-eval-fixtures.md) |
|
||||
| diagnosis-eval-baseline-diff | 诊断评测 baseline diff 与回归判断 | 中 | 已归档 | [diagnosis-eval-baseline-diff.md](diagnosis-eval-baseline-diff.md) |
|
||||
| mvp-demo-interview-runbook | Plan C 面试可复现 Demo 包 | 中 | 已归档 | [mvp-demo-interview-runbook.md](mvp-demo-interview-runbook.md) |
|
||||
**更新日期**:2026-07-10
|
||||
**状态**:按活跃问题、设计笔记、RAG 问题集和已归档问题整理
|
||||
|
||||
## RAG 重构计划
|
||||
## 目录约定
|
||||
|
||||
| 目录 | 用途 |
|
||||
|---|---|
|
||||
| [active/](active/) | 仍需要规划或实现的问题 |
|
||||
| [design-notes/](design-notes/) | 已形成方向、用于指导后续实现的设计记录 |
|
||||
| [rag/](rag/) | RAG 子问题集合;多数已合并到 RAG 重构计划 |
|
||||
| [archived/](archived/) | 已修复、已实施或已归档的问题 |
|
||||
|
||||
## 活跃问题
|
||||
|
||||
| 名称 | 标题 | 严重程度 | 状态 | 文件 |
|
||||
|---|---|---|---|---|
|
||||
| rag-refactor-plan | RAG 检索重构计划 | 高 | 待规划 | [rag-refactor-plan.md](rag-refactor-plan.md) |
|
||||
| ISS-003 | MVP 设计与实现 Review 收敛 | 高 | 待规划 | [active/ISS-003-mvp-design-implementation-review.md](active/ISS-003-mvp-design-implementation-review.md) |
|
||||
| ISS-004 | Executor 域级检索水位控制 | 低 | 待规划 | [active/ISS-004-executor-domain-hard-limit.md](active/ISS-004-executor-domain-hard-limit.md) |
|
||||
| executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [active/executor-evidence-attribution-hallucination.md](active/executor-evidence-attribution-hallucination.md) |
|
||||
| rag-refactor-plan | RAG 检索重构计划 | 高 | 待规划 | [active/rag-refactor-plan.md](active/rag-refactor-plan.md) |
|
||||
|
||||
## RAG 检索问题
|
||||
## 设计笔记
|
||||
|
||||
| 名称 | 标题 | 严重程度 | 状态 | 文件 |
|
||||
|---|---|---|---|---|
|
||||
| chunk-context-reconstruction | RAG 切片上下文重建缺失 | 高 | 已合并到重构计划 | [rag-chunk-context-reconstruction.md](rag-chunk-context-reconstruction.md) |
|
||||
| breadcrumb-embedding-gap | RAG breadcrumb 未参与向量语义 | 高 | 已合并到重构计划 | [rag-breadcrumb-embedding-gap.md](rag-breadcrumb-embedding-gap.md) |
|
||||
| l0-l1-fusion-ranking | RAG L0 和 L1 未真正融合排序 | 中 | 已合并到重构计划 | [rag-l0-l1-fusion-ranking.md](rag-l0-l1-fusion-ranking.md) |
|
||||
| l0-keyword-matching-quality | RAG L0 关键词匹配质量不足 | 中 | 已合并到重构计划 | [rag-l0-keyword-matching-quality.md](rag-l0-keyword-matching-quality.md) |
|
||||
| l1-score-calibration | RAG L1 分数阈值未校准 | 中 | 已合并到重构计划 | [rag-l1-score-calibration.md](rag-l1-score-calibration.md) |
|
||||
| context-packing-and-reranking | RAG 缺少上下文打包和 Rerank | 中 | 已合并到重构计划 | [rag-context-packing-and-reranking.md](rag-context-packing-and-reranking.md) |
|
||||
| upload-chunk-parameter-drift | RAG 上传切片参数未真正生效 | 低 | 已合并到重构计划 | [rag-upload-chunk-parameter-drift.md](rag-upload-chunk-parameter-drift.md) |
|
||||
| query-rewrite-gap | RAG 查询改写能力薄弱 | 中 | 已合并到重构计划 | [rag-query-rewrite-gap.md](rag-query-rewrite-gap.md) |
|
||||
| 名称 | 标题 | 状态 | 文件 |
|
||||
|---|---|---|---|
|
||||
| executor-self-evidence-loop-design-note | Executor 自证循环与证据摘要链路设计记录 | 已形成方向 | [design-notes/executor-self-evidence-loop-design-note.md](design-notes/executor-self-evidence-loop-design-note.md) |
|
||||
| executor-structured-output-v2 | Executor 结构化输出 V2 阶段设计 | 部分已实施,保留为后续改造参考 | [design-notes/executor-structured-output-v2.md](design-notes/executor-structured-output-v2.md) |
|
||||
|
||||
## RAG 框架化改造
|
||||
## RAG 问题集
|
||||
|
||||
| 名称 | 标题 | 严重程度 | 状态 | 文件 |
|
||||
|---|---|---|---|---|
|
||||
| spring-ai-vectorstore-migration | RAG 迁移到 Spring AI VectorStore 检索抽象 | 高 | 已合并到重构计划 | [rag-spring-ai-vectorstore-migration.md](rag-spring-ai-vectorstore-migration.md) |
|
||||
| spring-ai-query-transformer | RAG 接入 Spring AI Query Transformer | 中 | 已合并到重构计划 | [rag-spring-ai-query-transformer.md](rag-spring-ai-query-transformer.md) |
|
||||
| spring-ai-document-postprocessor | RAG 使用 DocumentPostProcessor 做后处理 | 中 | 已合并到重构计划 | [rag-spring-ai-document-postprocessor.md](rag-spring-ai-document-postprocessor.md) |
|
||||
| l0-domain-entity-hint | RAG 将 L0 降级为领域和实体 Hint | 中 | 已合并到重构计划 | [rag-l0-domain-entity-hint.md](rag-l0-domain-entity-hint.md) |
|
||||
| spring-ai-advisor-boundary | RAG 明确 Spring AI Advisor 与 Agent Tool 的边界 | 中 | 已合并到重构计划 | [rag-spring-ai-advisor-boundary.md](rag-spring-ai-advisor-boundary.md) |
|
||||
这些问题已经收敛到 [active/rag-refactor-plan.md](active/rag-refactor-plan.md),单个文件保留用于追溯原始问题和设计背景。
|
||||
|
||||
| 名称 | 标题 | 状态 | 文件 |
|
||||
|---|---|---|---|
|
||||
| breadcrumb-embedding-gap | RAG breadcrumb 未参与向量语义 | 已合并到重构计划 | [rag/rag-breadcrumb-embedding-gap.md](rag/rag-breadcrumb-embedding-gap.md) |
|
||||
| chunk-context-reconstruction | RAG 切片上下文重建缺失 | 已合并到重构计划 | [rag/rag-chunk-context-reconstruction.md](rag/rag-chunk-context-reconstruction.md) |
|
||||
| context-packing-and-reranking | RAG 缺少上下文打包和 Rerank | 已合并到重构计划 | [rag/rag-context-packing-and-reranking.md](rag/rag-context-packing-and-reranking.md) |
|
||||
| l0-domain-entity-hint | RAG 将 L0 降级为领域和实体 Hint | 已合并到重构计划 | [rag/rag-l0-domain-entity-hint.md](rag/rag-l0-domain-entity-hint.md) |
|
||||
| l0-keyword-matching-quality | RAG L0 关键词匹配质量不足 | 已合并到重构计划 | [rag/rag-l0-keyword-matching-quality.md](rag/rag-l0-keyword-matching-quality.md) |
|
||||
| l0-l1-fusion-ranking | RAG L0 和 L1 未真正融合排序 | 已合并到重构计划 | [rag/rag-l0-l1-fusion-ranking.md](rag/rag-l0-l1-fusion-ranking.md) |
|
||||
| l1-score-calibration | RAG L1 分数阈值未校准 | 已合并到重构计划 | [rag/rag-l1-score-calibration.md](rag/rag-l1-score-calibration.md) |
|
||||
| query-rewrite-gap | RAG 查询改写能力薄弱 | 已合并到重构计划 | [rag/rag-query-rewrite-gap.md](rag/rag-query-rewrite-gap.md) |
|
||||
| spring-ai-advisor-boundary | Spring AI Advisor 与 Agent Tool 边界 | 已合并到重构计划 | [rag/rag-spring-ai-advisor-boundary.md](rag/rag-spring-ai-advisor-boundary.md) |
|
||||
| spring-ai-document-postprocessor | 使用 DocumentPostProcessor 做后处理 | 已合并到重构计划 | [rag/rag-spring-ai-document-postprocessor.md](rag/rag-spring-ai-document-postprocessor.md) |
|
||||
| spring-ai-query-transformer | 接入 Spring AI Query Transformer | 已合并到重构计划 | [rag/rag-spring-ai-query-transformer.md](rag/rag-spring-ai-query-transformer.md) |
|
||||
| spring-ai-vectorstore-migration | 迁移到 Spring AI VectorStore 检索抽象 | 已合并到重构计划 | [rag/rag-spring-ai-vectorstore-migration.md](rag/rag-spring-ai-vectorstore-migration.md) |
|
||||
| upload-chunk-parameter-drift | 上传切片参数未真正生效 | 已合并到重构计划 | [rag/rag-upload-chunk-parameter-drift.md](rag/rag-upload-chunk-parameter-drift.md) |
|
||||
|
||||
## 已归档问题
|
||||
|
||||
| 名称 | 标题 | 状态 | 文件 |
|
||||
|---|---|---|---|
|
||||
| ISS-001 | Executor 重复召回同一文档 | 已修复 | [archived/ISS-001-duplicate-retrieval.md](archived/ISS-001-duplicate-retrieval.md) |
|
||||
| ISS-002 | Executor 无约束重复调用 lookup_knowledge | 已修复 | [archived/ISS-002-executor-unconstrained-lookup.md](archived/ISS-002-executor-unconstrained-lookup.md) |
|
||||
| ISS-005 | 证据链补齐与降级契约收敛 | 已归档 | [archived/ISS-005-evidence-trace-hardening.md](archived/ISS-005-evidence-trace-hardening.md) |
|
||||
| ISS-006 | 固定诊断评测集与回归 Harness | 已归档 | [archived/ISS-006-diagnosis-eval-harness.md](archived/ISS-006-diagnosis-eval-harness.md) |
|
||||
| ISS-007 | Verifier 证据摘要保真与工具命中质量问题 | 已实施 | [archived/ISS-007-verifier-evidence-summary-fidelity.md](archived/ISS-007-verifier-evidence-summary-fidelity.md) |
|
||||
| ISS-008 | Executor 窄范围查询越界 | 已修复 | [archived/ISS-008-executor-narrow-scope-overreach.md](archived/ISS-008-executor-narrow-scope-overreach.md) |
|
||||
| ISS-009 | negative_observation 精确引用 no-evidence 结果 | 已修复 | [archived/ISS-009-negative-observation-no-evidence-reference.md](archived/ISS-009-negative-observation-no-evidence-reference.md) |
|
||||
| ISS-010 | 同 session 多轮诊断 Trace 隔离 | 已归档 | [archived/ISS-010-session-run-trace-isolation.md](archived/ISS-010-session-run-trace-isolation.md) |
|
||||
| diagnosis-eval-baseline-diff | 诊断评测 baseline diff 与回归判断 | 已归档 | [archived/diagnosis-eval-baseline-diff.md](archived/diagnosis-eval-baseline-diff.md) |
|
||||
| expand-diagnosis-eval-fixtures | 补齐固定诊断评测 fixture 与 baseline | 已归档 | [archived/expand-diagnosis-eval-fixtures.md](archived/expand-diagnosis-eval-fixtures.md) |
|
||||
| mvp-demo-interview-runbook | Plan C 面试可复现 Demo 包 | 已归档 | [archived/mvp-demo-interview-runbook.md](archived/mvp-demo-interview-runbook.md) |
|
||||
|
||||
@@ -347,19 +347,19 @@ RAG、Agent、AIOps、数据库记录互相关联,必须分阶段推进,每
|
||||
|
||||
本计划合并以下问题和改造方向:
|
||||
|
||||
- [rag-chunk-context-reconstruction.md](rag-chunk-context-reconstruction.md)
|
||||
- [rag-breadcrumb-embedding-gap.md](rag-breadcrumb-embedding-gap.md)
|
||||
- [rag-l0-l1-fusion-ranking.md](rag-l0-l1-fusion-ranking.md)
|
||||
- [rag-l0-keyword-matching-quality.md](rag-l0-keyword-matching-quality.md)
|
||||
- [rag-l1-score-calibration.md](rag-l1-score-calibration.md)
|
||||
- [rag-context-packing-and-reranking.md](rag-context-packing-and-reranking.md)
|
||||
- [rag-upload-chunk-parameter-drift.md](rag-upload-chunk-parameter-drift.md)
|
||||
- [rag-query-rewrite-gap.md](rag-query-rewrite-gap.md)
|
||||
- [rag-spring-ai-vectorstore-migration.md](rag-spring-ai-vectorstore-migration.md)
|
||||
- [rag-spring-ai-query-transformer.md](rag-spring-ai-query-transformer.md)
|
||||
- [rag-spring-ai-document-postprocessor.md](rag-spring-ai-document-postprocessor.md)
|
||||
- [rag-l0-domain-entity-hint.md](rag-l0-domain-entity-hint.md)
|
||||
- [rag-spring-ai-advisor-boundary.md](rag-spring-ai-advisor-boundary.md)
|
||||
- [rag-chunk-context-reconstruction.md](../rag/rag-chunk-context-reconstruction.md)
|
||||
- [rag-breadcrumb-embedding-gap.md](../rag/rag-breadcrumb-embedding-gap.md)
|
||||
- [rag-l0-l1-fusion-ranking.md](../rag/rag-l0-l1-fusion-ranking.md)
|
||||
- [rag-l0-keyword-matching-quality.md](../rag/rag-l0-keyword-matching-quality.md)
|
||||
- [rag-l1-score-calibration.md](../rag/rag-l1-score-calibration.md)
|
||||
- [rag-context-packing-and-reranking.md](../rag/rag-context-packing-and-reranking.md)
|
||||
- [rag-upload-chunk-parameter-drift.md](../rag/rag-upload-chunk-parameter-drift.md)
|
||||
- [rag-query-rewrite-gap.md](../rag/rag-query-rewrite-gap.md)
|
||||
- [rag-spring-ai-vectorstore-migration.md](../rag/rag-spring-ai-vectorstore-migration.md)
|
||||
- [rag-spring-ai-query-transformer.md](../rag/rag-spring-ai-query-transformer.md)
|
||||
- [rag-spring-ai-document-postprocessor.md](../rag/rag-spring-ai-document-postprocessor.md)
|
||||
- [rag-l0-domain-entity-hint.md](../rag/rag-l0-domain-entity-hint.md)
|
||||
- [rag-spring-ai-advisor-boundary.md](../rag/rag-spring-ai-advisor-boundary.md)
|
||||
|
||||
---
|
||||
|
||||
+1
-1
@@ -4,7 +4,7 @@
|
||||
**严重程度**:中(影响 token 消耗和上下文质量,不影响功能正确性)
|
||||
**发现时间**:2026-06-30
|
||||
**修复版本**:session-dedup-knowledge-map
|
||||
**历史架构文档**:[会话级去重与知识域地图](../architecture/archive/2026-07-05-legacy/session-dedup-knowledge-map.md)
|
||||
**历史架构文档**:[会话级去重与知识域地图](../../architecture/archive/2026-07-05-legacy/session-dedup-knowledge-map.md)
|
||||
|
||||
---
|
||||
|
||||
@@ -0,0 +1,308 @@
|
||||
# ISS-008 Executor 窄范围查询越界
|
||||
|
||||
**严重程度**:中
|
||||
**状态**:已修复
|
||||
**发现时间**:2026-07-08
|
||||
**关联**:
|
||||
- `ISS-007-verifier-evidence-summary-fidelity`
|
||||
- `executor-structured-output-v2`
|
||||
- `executor-evidence-attribution-hallucination`
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
当前 Chat 诊断链路已经演进为:
|
||||
|
||||
```text
|
||||
Planner
|
||||
-> Executor
|
||||
-> VerifierInputHook / Gatekeeper
|
||||
-> Verifier
|
||||
-> Composer
|
||||
```
|
||||
|
||||
其中 Executor 的定位已经从“生成最终诊断答案”收敛为:
|
||||
|
||||
```text
|
||||
证据收集 + 微观事实提炼
|
||||
```
|
||||
|
||||
但在窄范围问题中,Executor 仍可能把用户只要求确认的一件事扩展成多条 claim,例如用户只问 `HighCPUUsage`,Executor 可能顺手输出内存、连接池、数据库或修复建议相关内容。
|
||||
|
||||
这类问题不一定是证据伪造。很多时候工具返回里确实有其它信息,但它们不属于当前用户问题的范围。Gatekeeper 只能校验证据引用真假,不能完整承担“用户意图范围控制”;Verifier 虽然可以降级,但会增加链路负担。
|
||||
|
||||
因此本 issue 采用低成本的 Prompt-first 修复:先收紧 Executor prompt,不改 Planner,不引入 `scope_contract`。
|
||||
|
||||
---
|
||||
|
||||
## 问题类型
|
||||
|
||||
### 1. 窄范围查询越界
|
||||
|
||||
用户问题只要求确认一个服务、告警、日志、订单或时间窗口,但 Executor 输出了用户未要求的 claim。
|
||||
|
||||
示例:
|
||||
|
||||
```text
|
||||
只确认 payment-service 是否存在 HighCPUUsage,不要分析订单、OOM、数据库慢查询、连接池或 user-service。
|
||||
```
|
||||
|
||||
错误输出包括:
|
||||
|
||||
- `HighMemoryUsage`
|
||||
- `SlowResponse`
|
||||
- `order-123`
|
||||
- `HikariCP`
|
||||
- `DB / database`
|
||||
- `user-service`
|
||||
|
||||
### 2. Observation 变成 Diagnosis
|
||||
|
||||
Executor 本应输出观察事实,却输出根因、风险、修复建议或经验推断。
|
||||
|
||||
错误输出包括:
|
||||
|
||||
- “CPU 过高是请求超时的根因”
|
||||
- “建议扩容”
|
||||
- “通常这种情况是数据库慢查询导致”
|
||||
- “存在内存泄漏风险”
|
||||
|
||||
### 3. Runbook 通用知识变成当前事实
|
||||
|
||||
Runbook、Skill、知识库可以指导要查什么,但不能直接变成本次环境已发生的事实。
|
||||
|
||||
错误输出包括:
|
||||
|
||||
```text
|
||||
Runbook 中说 HighCPUUsage 常见原因是流量突增,所以当前环境发生了流量突增。
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 修复决策
|
||||
|
||||
本期只修 Executor prompt。
|
||||
|
||||
### 本期做
|
||||
|
||||
1. 强化 Executor 单一职责:证据收集 + 微观事实提炼。
|
||||
2. 增加 `角色边界 HARD-GATE`。
|
||||
3. 增加 `窄范围确认任务 HARD-GATE`。
|
||||
4. 增加工具使用边界,避免为补全故事而扩展检索。
|
||||
5. 增加输出前自检,要求输出 JSON 前删除越界 claim。
|
||||
|
||||
### 本期不做
|
||||
|
||||
1. 不改 Planner。
|
||||
2. 不新增 `scope_contract`。
|
||||
3. 不解析 Planner 输出中的 scope。
|
||||
4. 不做 Gatekeeper scope 校验。
|
||||
5. 不改多 Agent 编排。
|
||||
|
||||
---
|
||||
|
||||
## 设计原则
|
||||
|
||||
### Claim 要少,Evidence 可以多
|
||||
|
||||
窄范围任务下,Executor 应输出最少必要 claim,通常 1 条,最多 2 条。
|
||||
|
||||
但 claim 数量限制不限制 `evidence_bindings` 数量。一条核心 claim 可以绑定多条直接相关证据。
|
||||
|
||||
```text
|
||||
正确:
|
||||
1 条 claim + 多条 evidence_bindings
|
||||
|
||||
错误:
|
||||
为了展示多条证据,把同一个观察事实拆成多条 claim
|
||||
```
|
||||
|
||||
### 只输出当前问题范围内的 Observation
|
||||
|
||||
窄范围任务下,`claims` 只能使用:
|
||||
|
||||
- `observation`
|
||||
- `negative_observation`
|
||||
|
||||
禁止使用:
|
||||
|
||||
- `root_cause`
|
||||
- `risk`
|
||||
- `recommendation`
|
||||
- 其它建议类或诊断类 claim
|
||||
|
||||
### 证据不足时不要补故事
|
||||
|
||||
如果工具没有返回可被精确引用的证据:
|
||||
|
||||
```text
|
||||
source_invocation_id + raw_path + evidence_excerpt
|
||||
```
|
||||
|
||||
Executor 不应生成 confirmed claim,应写入 `missing_info`。
|
||||
|
||||
---
|
||||
|
||||
## Prompt 修复点
|
||||
|
||||
已更新:
|
||||
|
||||
- `src/main/resources/prompts/chat-executor-prompt.md`
|
||||
|
||||
核心新增约束:
|
||||
|
||||
1. `角色边界 HARD-GATE`
|
||||
2. `窄范围确认任务 HARD-GATE`
|
||||
3. `工具使用边界`
|
||||
4. `输出前自检`
|
||||
5. 条目级 `raw_path` 强约束:同一条工具数组项只能绑定一次,禁止输出 `$.alerts[0].alert_name`、`$.alerts[0].state` 等字段级子路径。
|
||||
|
||||
---
|
||||
|
||||
## 验证结果
|
||||
|
||||
### 2026-07-08 E2E 验证
|
||||
|
||||
输入:
|
||||
|
||||
```text
|
||||
只确认 payment-service 是否存在 HighCPUUsage,不要分析订单123、OOM、数据库慢查询、连接池或 user-service。
|
||||
```
|
||||
|
||||
第一次验证发现:
|
||||
|
||||
- Executor 已经只输出 `payment-service + HighCPUUsage` 相关 observation,没有输出越界 claim。
|
||||
- 但 Executor 额外生成了字段级 `raw_path`:
|
||||
- `$.alerts[0].alert_name`
|
||||
- `$.alerts[0].state`
|
||||
- 当前 Gatekeeper 只支持条目级路径 `$.alerts[i]` / `$.logs[i]` / `$.evidence_blocks[i]`,因此判定为 `REJECT`。
|
||||
|
||||
已追加 prompt 约束:
|
||||
|
||||
```text
|
||||
同一条工具数组项只能绑定一次。
|
||||
不要为了引用其中多个字段而拆成多个 evidence_bindings。
|
||||
raw_path 禁止指向字段级子路径。
|
||||
```
|
||||
|
||||
第二次验证结果:
|
||||
|
||||
```text
|
||||
sessionId: iss008-narrow-highcpu-rerun-20260708-215510
|
||||
verdict: PASS
|
||||
groundedness_score: 1.0
|
||||
gatekeeper_result.status: pass
|
||||
gatekeeper_result.severity: none
|
||||
claim_count: 1
|
||||
claim_type: observation
|
||||
forbidden_hits: none
|
||||
```
|
||||
|
||||
Executor claim:
|
||||
|
||||
```text
|
||||
payment-service 当前存在 HighCPUUsage 告警,CPU 使用率持续超过 80%,当前值为 92%,告警状态为 firing,已持续 25 分钟。
|
||||
```
|
||||
|
||||
最终答案未出现以下排除项:
|
||||
|
||||
- `HighMemoryUsage`
|
||||
- `SlowResponse`
|
||||
- `order-123`
|
||||
- `订单123`
|
||||
- `OOM`
|
||||
- `DB / database`
|
||||
- `HikariCP`
|
||||
- `connection pool / 连接池`
|
||||
- `user-service`
|
||||
|
||||
---
|
||||
|
||||
## 验收标准
|
||||
|
||||
### 1. HighCPUUsage 窄范围
|
||||
|
||||
输入:
|
||||
|
||||
```text
|
||||
只确认 payment-service 是否存在 HighCPUUsage,不要分析订单123、OOM、数据库慢查询、连接池或 user-service。
|
||||
```
|
||||
|
||||
期望:
|
||||
|
||||
- `claims` 只围绕 `payment-service + HighCPUUsage`。
|
||||
- 不出现 `HighMemoryUsage`。
|
||||
- 不出现 `SlowResponse`。
|
||||
- 不出现 `order-123`。
|
||||
- 不出现 `OOM`。
|
||||
- 不出现 `DB / database`。
|
||||
- 不出现 `HikariCP / connection pool`。
|
||||
- 不出现 `user-service`。
|
||||
|
||||
### 2. HighMemoryUsage 窄范围
|
||||
|
||||
输入:
|
||||
|
||||
```text
|
||||
只确认 order-service 是否存在 HighMemoryUsage。
|
||||
```
|
||||
|
||||
期望:
|
||||
|
||||
- 可以输出内存使用率、告警状态、持续时间等观察事实。
|
||||
- 不输出“内存泄漏已确认”。
|
||||
- 不输出扩容、重启、修改 JVM 参数等修复建议。
|
||||
|
||||
### 3. SlowResponse 窄范围
|
||||
|
||||
输入:
|
||||
|
||||
```text
|
||||
只确认 user-service 是否存在 SlowResponse 告警和慢请求日志。
|
||||
```
|
||||
|
||||
期望:
|
||||
|
||||
- 可以绑定 alert 和 logs 多条证据。
|
||||
- 不推断数据库连接池耗尽。
|
||||
- 不推断下游服务故障。
|
||||
|
||||
### 4. 用户明确排除项
|
||||
|
||||
输入:
|
||||
|
||||
```text
|
||||
只看 order-service 支付失败日志,不要分析 HikariCP。
|
||||
```
|
||||
|
||||
期望:
|
||||
|
||||
- `claim_text` 不出现 HikariCP 确认结论。
|
||||
- 最终答案不出现 HikariCP 确认结论。
|
||||
|
||||
### 5. 证据不足
|
||||
|
||||
输入:
|
||||
|
||||
```text
|
||||
只确认 inventory-service 是否存在 HikariCP 连接池耗尽日志。
|
||||
```
|
||||
|
||||
期望:
|
||||
|
||||
- 如果工具返回 `logs=[]`,Executor 不编造 positive claim。
|
||||
- 输出 `negative_observation` 或 `missing_info`。
|
||||
- 不返回 `generic-service` 占位事实。
|
||||
|
||||
---
|
||||
|
||||
## 后续增强
|
||||
|
||||
如果 Prompt-first 后仍不稳定,再考虑:
|
||||
|
||||
1. Planner 输出 `scope_contract`。
|
||||
2. Gatekeeper 增加 scope 校验。
|
||||
3. eval fixture 增加 forbidden claim 自动断言。
|
||||
|
||||
本期暂不进入这些改造。
|
||||
@@ -0,0 +1,239 @@
|
||||
# ISS-009 negative_observation 精确引用 no-evidence 结果
|
||||
|
||||
**严重程度**:中
|
||||
**状态**:已修复
|
||||
**发现时间**:2026-07-08
|
||||
**关联**:
|
||||
- `ISS-007-verifier-evidence-summary-fidelity`
|
||||
- `ISS-008-executor-narrow-scope-overreach`
|
||||
- `executor-structured-output-v2`
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
ISS-007 已经把正向证据引用收敛为:
|
||||
|
||||
```text
|
||||
source_invocation_id + raw_path + evidence_excerpt
|
||||
```
|
||||
|
||||
Gatekeeper 通过 `tool_invocation.retrieval_details.evidence_refs` 校验 Executor 引用是否真实存在。
|
||||
|
||||
但负向观察存在一个缺口:当工具明确返回“没查到”时,结果通常是空数组:
|
||||
|
||||
```json
|
||||
{
|
||||
"logs": [],
|
||||
"total": 0,
|
||||
"message": "未找到匹配的日志"
|
||||
}
|
||||
```
|
||||
|
||||
这时没有 `$.logs[0]`、`$.alerts[0]` 或 `$.evidence_blocks[0]` 可以引用。Executor 如果输出 `negative_observation`,Gatekeeper 无法稳定验证它引用的“无证据结果”,容易降级为 `LOW_CONFID` 或 `REJECT`。
|
||||
|
||||
---
|
||||
|
||||
## 问题
|
||||
|
||||
用户问:
|
||||
|
||||
```text
|
||||
只确认 inventory-service 是否存在 HikariCP 连接池耗尽日志。
|
||||
```
|
||||
|
||||
工具返回:
|
||||
|
||||
```json
|
||||
{
|
||||
"success": false,
|
||||
"logs": [],
|
||||
"total": 0,
|
||||
"message": "未找到匹配的日志"
|
||||
}
|
||||
```
|
||||
|
||||
合理 claim 是:
|
||||
|
||||
```json
|
||||
{
|
||||
"claim_type": "negative_observation",
|
||||
"claim_text": "未检索到 inventory-service 的 HikariCP 连接池耗尽日志。"
|
||||
}
|
||||
```
|
||||
|
||||
但旧设计只支持正向数组项:
|
||||
|
||||
```text
|
||||
$.alerts[i]
|
||||
$.logs[i]
|
||||
$.evidence_blocks[i]
|
||||
```
|
||||
|
||||
因此负向观察缺少可回溯的精确引用点。
|
||||
|
||||
---
|
||||
|
||||
## 修复决策
|
||||
|
||||
给 no-hit / no-evidence 结果增加一等证据引用:
|
||||
|
||||
```json
|
||||
{
|
||||
"raw_path": "$.no_evidence",
|
||||
"text": "query_logs returned no evidence; evidence_status=no_evidence; query=inventory-service HikariCP; topic=application-logs; total=0; message=未找到匹配的日志"
|
||||
}
|
||||
```
|
||||
|
||||
Executor 可以引用:
|
||||
|
||||
```json
|
||||
{
|
||||
"tool_name": "query_logs",
|
||||
"source_invocation_id": 123,
|
||||
"raw_path": "$.no_evidence",
|
||||
"evidence_excerpt": "query_logs returned no evidence; query=inventory-service HikariCP; total=0; evidence_status=no_evidence"
|
||||
}
|
||||
```
|
||||
|
||||
### 语义边界
|
||||
|
||||
`$.no_evidence` 只表示:
|
||||
|
||||
```text
|
||||
该工具对当前查询返回无匹配证据。
|
||||
```
|
||||
|
||||
它不表示:
|
||||
|
||||
- 问题绝对不存在。
|
||||
- 根因被排除。
|
||||
- 系统已经健康。
|
||||
- 没有必要继续排查。
|
||||
|
||||
---
|
||||
|
||||
## 实施范围
|
||||
|
||||
### 已修改
|
||||
|
||||
- `ToolInvocationRecorder`
|
||||
- `query_logs` no-hit 时生成 `evidence_refs[0].raw_path="$.no_evidence"`。
|
||||
- `query_metrics` no-hit 时生成 `evidence_refs[0].raw_path="$.no_evidence"`。
|
||||
- `lookup_knowledge` no-hit 且无 evidence blocks 时生成 `evidence_refs[0].raw_path="$.no_evidence"`。
|
||||
- `ExecutorGatekeeperService`
|
||||
- 复用既有 `evidence_refs` 校验逻辑,无需新增特殊分支。
|
||||
- `$.no_evidence` 和普通 raw_path 一样必须存在于 `retrieval_details.evidence_refs`。
|
||||
- `chat-executor-prompt.md`
|
||||
- 明确 `negative_observation` 必须引用 `$.no_evidence`。
|
||||
- 明确没有实际工具调用时禁止使用 `$.no_evidence`。
|
||||
|
||||
### 未修改
|
||||
|
||||
- 不新增数据库表。
|
||||
- 不新增复杂 metadata。
|
||||
- 不改变 Planner。
|
||||
- 不改变 Agent 编排。
|
||||
|
||||
---
|
||||
|
||||
## 验收标准
|
||||
|
||||
1. `query_logs` 返回 `logs=[] / total=0 / evidence_status=no_evidence` 时,`tool_invocation.retrieval_details.evidence_refs` 包含:
|
||||
|
||||
```json
|
||||
{
|
||||
"raw_path": "$.no_evidence"
|
||||
}
|
||||
```
|
||||
|
||||
2. Executor 输出 `negative_observation` 并引用 `$.no_evidence` 时,Gatekeeper 可以校验通过。
|
||||
|
||||
3. Executor 如果用 `$.no_evidence` 搭配正向证据文本,例如 `HikariCP active=50/50`,Gatekeeper 必须拒绝。
|
||||
|
||||
4. `$.no_evidence` 不得被解释为“问题绝对不存在”,只能表达“当前查询未检索到匹配证据”。
|
||||
|
||||
5. HikariCP negative E2E:
|
||||
|
||||
```text
|
||||
只确认 inventory-service 是否存在 HikariCP 连接池耗尽日志。
|
||||
```
|
||||
|
||||
期望:
|
||||
|
||||
- 工具返回 no-hit。
|
||||
- 不返回 `generic-service`。
|
||||
- Executor 输出 `negative_observation`。
|
||||
- `raw_path="$.no_evidence"`。
|
||||
- Gatekeeper `pass/none`。
|
||||
- Verifier 不误判为正向 HikariCP 证据。
|
||||
|
||||
---
|
||||
|
||||
## 验证记录
|
||||
|
||||
### 2026-07-08 单元测试
|
||||
|
||||
命令:
|
||||
|
||||
```text
|
||||
mvn '-Dtest=ToolInvocationRecorderTest,ExecutorGatekeeperServiceTest,QueryLogsToolsTest' test
|
||||
```
|
||||
|
||||
结果:
|
||||
|
||||
```text
|
||||
Tests run: 23, Failures: 0, Errors: 0, Skipped: 0
|
||||
BUILD SUCCESS
|
||||
```
|
||||
|
||||
覆盖点:
|
||||
|
||||
- `query_logs` no-hit 生成 `$.no_evidence`。
|
||||
- `query_metrics` no-hit 生成 `$.no_evidence`。
|
||||
- `lookup_knowledge` no-hit 生成 `$.no_evidence`。
|
||||
- Gatekeeper 可以校验 `$.no_evidence`。
|
||||
- 多个 no-evidence 调用存在时,Gatekeeper 可按 `tool_name + raw_path + evidence_excerpt` 唯一回填 `source_invocation_id`。
|
||||
- `negative_observation` 混绑正向 `$.logs[i]` 会被拒绝。
|
||||
- HikariCP negative mock 不返回 `generic-service`。
|
||||
|
||||
### 2026-07-08 E2E 验证
|
||||
|
||||
输入:
|
||||
|
||||
```text
|
||||
只确认 inventory-service 是否存在 HikariCP 连接池耗尽日志。
|
||||
```
|
||||
|
||||
最终通过 session:
|
||||
|
||||
```text
|
||||
sessionId: iss009-hikari-negative-latest-20260708-232428
|
||||
verdict: PASS
|
||||
groundedness_score: 1.0
|
||||
gatekeeper_result.status: pass
|
||||
gatekeeper_result.severity: none
|
||||
claim_count: 1
|
||||
claim_type: negative_observation
|
||||
raw_path: $.no_evidence
|
||||
generic_service_hit: false
|
||||
overstate_hit: false
|
||||
```
|
||||
|
||||
Executor claim:
|
||||
|
||||
```text
|
||||
当前查询未检索到 inventory-service 的 HikariCP 连接池耗尽日志。
|
||||
```
|
||||
|
||||
最终答案:
|
||||
|
||||
```text
|
||||
本次查询在 inventory-service 中未发现 HikariCP 连接池耗尽的日志记录,检索结果未匹配到相关证据。
|
||||
```
|
||||
|
||||
说明:
|
||||
|
||||
- Executor 仍可能输出多个 `$.no_evidence` binding。
|
||||
- 如果 `source_invocation_id` 缺失,Gatekeeper 会按 `tool_name + raw_path + evidence_excerpt` 唯一匹配真实 invocation 并写入 warning。
|
||||
- 最终答案不使用“排除”“确认没有”“不存在该问题”等过度表达。
|
||||
@@ -0,0 +1,586 @@
|
||||
# ISS-010 同 session 多轮诊断 Trace 隔离
|
||||
|
||||
**状态**:已归档
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-10
|
||||
**来源**:同一 `sessionId` 多轮 Chat E2E 验证
|
||||
|
||||
**归档日期**:2026-07-10
|
||||
**OpenSpec**:`openspec/changes/archive/2026-07-10-session-run-trace-isolation`
|
||||
**实现提交**:`52bf030`、`26d5529`、`027aed1`、`d928a19`、`78c1477`、`f9df943`
|
||||
**归档提交**:`3578709`
|
||||
|
||||
---
|
||||
|
||||
## 背景
|
||||
|
||||
归档结论:当前 MVP 已将“会话态”和“运行态”拆开。`sessionId` 表示多轮会话目录和 Redis 上下文;`runId` 表示一次可回放诊断执行。Trace、Feedback、Evaluation、AIOps 和案例沉淀的新路径都按 `runId` 隔离。
|
||||
|
||||
原问题中 Chat 链路同时存在两类“会话”语义:
|
||||
|
||||
```text
|
||||
Redis SessionContext
|
||||
-> 保存同一 sessionId 的多轮对话历史
|
||||
-> 用于下一轮模型上下文
|
||||
|
||||
MySQL diagnosis_session / agent_step / tool_invocation
|
||||
-> 保存诊断 Trace
|
||||
-> 用于 Trace API、Verifier、Evidence score、Feedback 和评测
|
||||
```
|
||||
|
||||
多轮对话需要继续复用 `sessionId`,否则无法保留上下文。但一次诊断 Trace 应该是可独立回放、可独立评分、可独立反馈的执行单元。
|
||||
|
||||
当前实现只按 `sessionId` 关联 Trace,导致同一个 `sessionId` 下多轮诊断的 step/tool 记录混在一起。
|
||||
|
||||
当前实现已改为:
|
||||
|
||||
```text
|
||||
chat_session(sessionId)
|
||||
-> diagnosis_run(runId)
|
||||
-> agent_step.run_id
|
||||
-> tool_invocation.run_id
|
||||
```
|
||||
|
||||
旧 `diagnosis_session` 保留为历史兼容和回滚表,新 Chat/AIOps 执行不再写入新的运行态。
|
||||
|
||||
## 归档结果
|
||||
|
||||
- OpenSpec 已归档到 `openspec/changes/archive/2026-07-10-session-run-trace-isolation`。
|
||||
- 主规格已同步到 `openspec/specs/session-run-trace-isolation/spec.md`。
|
||||
- `mvp/architecture/` 和 `mvp/tables/` 已更新为 `chat_session -> diagnosis_run -> agent_step/tool_invocation(run_id)` 模型。
|
||||
- Demo 脚本和 Trace UI 已支持 `sessionId + runId` 精确 Trace 和 Feedback。
|
||||
- Maven E2E、`scripts/query_mysql.py` DB 检查、`logs/` 日志检查和 baseline drift 检查均已通过;未观察到 baseline drift。
|
||||
- `devflow/projects/2026-07-10-session-run-trace-isolation/` 已保存 brief、evidence、decisions、acceptance。
|
||||
|
||||
---
|
||||
|
||||
## E2E 证据
|
||||
|
||||
本次使用 `mvp-demo` profile 通过 Maven 启动服务,并用同一个 `sessionId` 连续请求两轮 `/api/chat`:
|
||||
|
||||
```text
|
||||
sessionId = e2e-multiturn-codex-20260710-1615
|
||||
round 1 = 支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。
|
||||
round 2 = 基于上一轮结论,只列出目前最缺的三类证据,以及下一步应该优先查哪个系统。
|
||||
```
|
||||
|
||||
验证结果:
|
||||
|
||||
- 第一轮成功,走多 Agent:`planner -> executor -> verifier -> composer`。
|
||||
- 第二轮成功,日志显示进入请求时 `会话历史消息对数: 1`,说明 Redis 历史上下文被复用。
|
||||
- `/api/chat/session/{sessionId}` 返回 `messagePairCount=2`。
|
||||
- `diagnosis_session` 只有一行,`query` 被第二轮问题覆盖。
|
||||
- `agent_step` 返回 14 行,包含第一轮多 Agent step 和第二轮 `intelligent_assistant` step。
|
||||
- `tool_invocation` 返回 19 行,包含两轮工具调用。
|
||||
- `self_evaluation.verifier_evaluation` 仍保留第一轮 Verifier 结果;第二轮简单问答没有新的 Verifier,但 rule evaluation 会基于同 session 全部工具调用重新计算。
|
||||
|
||||
关键入库形态:
|
||||
|
||||
```text
|
||||
diagnosis_session
|
||||
session_id = e2e-multiturn-codex-20260710-1615
|
||||
query = round 2 question
|
||||
status = SUCCESS
|
||||
step_count = 14
|
||||
tool_call_count = 19
|
||||
|
||||
agent_step
|
||||
round 1: planner, executor..., verifier, composer
|
||||
round 2: intelligent_assistant...
|
||||
|
||||
tool_invocation
|
||||
round 1 tools + round 2 tools all under same session_id
|
||||
```
|
||||
|
||||
本次验证产物保存在:
|
||||
|
||||
- `target/e2e/request-round1.json`
|
||||
- `target/e2e/response-round1.json`
|
||||
- `target/e2e/request-round2.json`
|
||||
- `target/e2e/response-round2.json`
|
||||
- `target/e2e/trace-after-round2.json`
|
||||
|
||||
---
|
||||
|
||||
## 核心问题
|
||||
|
||||
### P0:Trace 不是单次诊断的稳定回放
|
||||
|
||||
`GET /api/diagnosis/{sessionId}/trace` 会聚合同一 `sessionId` 下所有 `agent_step` 和 `tool_invocation`。
|
||||
|
||||
多轮之后,Trace 不再表示某一轮诊断,而是混合历史执行轨迹。
|
||||
|
||||
### P0:Verifier 和评分可能读取跨轮证据
|
||||
|
||||
Verifier、Gatekeeper、`ToolTraceSummaryService` 和 `EvaluationService` 当前主要按 `sessionId` 查询工具调用。
|
||||
|
||||
如果上一轮和当前轮证据混在一起,当前轮可能引用或评分到历史工具结果。
|
||||
|
||||
### P1:反馈语义不清晰
|
||||
|
||||
`feedback` 当前在 `diagnosis_session` 上按 `sessionId` 保存。
|
||||
|
||||
多轮之后,用户反馈的是哪一轮答案不再明确。`useful` 反馈沉淀到 `case_library` 时也可能关联到最新主表答案,而不是用户实际评价的那一轮。
|
||||
|
||||
### P1:`diagnosis_session` 字段被覆盖但子表追加
|
||||
|
||||
主表 `query/answer/status/self_evaluation/step_count/tool_call_count` 表示最新运行或混合统计,子表却保留多轮历史。
|
||||
|
||||
这会让 Trace summary、数据库统计和人工排查产生歧义。
|
||||
|
||||
---
|
||||
|
||||
## 已确认决策
|
||||
|
||||
### D1:`runId` 是正式 API 字段
|
||||
|
||||
`/api/chat` 和 `/api/ai_ops` 的响应或 SSE 消息需要暴露本次执行的 `runId`。
|
||||
|
||||
```text
|
||||
sessionId = 多轮对话上下文 ID
|
||||
runId = 本轮诊断执行 ID
|
||||
```
|
||||
|
||||
新客户端应优先用 `runId` 查询 Trace 和提交 Feedback。旧客户端只传 `sessionId` 时,服务端兼容解析该 session 的最新 run。
|
||||
|
||||
### D2:拆分会话态和运行态
|
||||
|
||||
不再把 session、run、trace 全部塞进 `diagnosis_session` 一张主表。
|
||||
|
||||
新增两张主表:
|
||||
|
||||
```text
|
||||
chat_session
|
||||
-> 多轮对话上下文主表
|
||||
|
||||
diagnosis_run
|
||||
-> 单次诊断执行主表
|
||||
```
|
||||
|
||||
Trace 继续使用现有明细表表达:
|
||||
|
||||
```text
|
||||
agent_step
|
||||
tool_invocation
|
||||
```
|
||||
|
||||
暂不新增单独的 `diagnosis_trace` 或 `trace_event` 主表。
|
||||
|
||||
### D3:`runId` 格式
|
||||
|
||||
使用 `run-` + UUID 全量字符串。
|
||||
|
||||
```text
|
||||
run-550e8400-e29b-41d4-a716-446655440000
|
||||
```
|
||||
|
||||
### D4:Trace API 兼容旧路径
|
||||
|
||||
```text
|
||||
GET /api/diagnosis/{sessionId}/trace
|
||||
-> 查该 session 最新 run
|
||||
|
||||
GET /api/diagnosis/{sessionId}/trace?runId=run-xxx
|
||||
-> 查指定 run
|
||||
```
|
||||
|
||||
指定 `runId` 时必须校验该 run 属于 path 中的 `sessionId`。
|
||||
|
||||
最新 run 建议按 `diagnosis_run.created_at DESC, id DESC` 解析,避免旧 run 因反馈或异步评分更新 `updated_at` 后被误认为最新。
|
||||
|
||||
### D5:Feedback 优先绑定 run
|
||||
|
||||
Feedback request 支持 `runId`。
|
||||
|
||||
- 有 `runId`:绑定指定 run。
|
||||
- 无 `runId`:短期兼容绑定该 `sessionId` 最新 run,并显式标记 fallback。
|
||||
- `case_library.diagnosis_id` 新数据保存 `run_id`。
|
||||
|
||||
兼容语义:历史 `case_library.diagnosis_id` 可能保存 `diagnosis_session.session_id`;本 change 之后自动沉淀的新数据保存 `diagnosis_run.run_id`。查询、幂等和文档需要在过渡期识别两种来源,避免把旧案例误判为无效数据。
|
||||
|
||||
### D6:所有 `/api/chat` 执行请求都创建 run
|
||||
|
||||
只要请求通过参数校验并进入 `ChatService.executeChatWithStrategy`,就创建新的 diagnosis run。
|
||||
|
||||
- 简单问答也创建 run。
|
||||
- 复杂诊断也创建 run。
|
||||
- 空问题等参数校验失败不创建 run。
|
||||
|
||||
### D7:AIOps 同步纳入 run 隔离
|
||||
|
||||
每次 `/api/ai_ops` 执行也创建新的 diagnosis run。AIOps 的 step、tool invocation 和 rule evaluation 都按 `runId` 隔离。
|
||||
|
||||
阶段说明:AIOps 可作为独立实现切片排在 Chat 之后,但必须在本 change 整体完成前落地;Chat-only 的中间状态只能作为过渡验证状态,不能作为生产完成状态归档。
|
||||
|
||||
### D8:`chat_session` 第一阶段只保存会话元数据
|
||||
|
||||
`chat_session` 是会话目录/索引表,不保存完整对话历史正文。
|
||||
|
||||
建议保存:
|
||||
|
||||
```text
|
||||
session_id
|
||||
status
|
||||
message_pair_count
|
||||
created_at
|
||||
last_active_at
|
||||
expires_at
|
||||
```
|
||||
|
||||
完整多轮对话历史继续放在 Redis `SessionContext.messageHistory`,用于下一轮 prompt 上下文。
|
||||
|
||||
每轮需要长期审计的用户问题和最终回答保存到 `diagnosis_run.query` / `diagnosis_run.answer`。
|
||||
|
||||
`chat_session.expires_at` 只表示 MySQL 会话目录的过期/清理元数据;Redis TTL 到期后,`SessionContext.messageHistory` 可能不再存在,但已经持久化的 `diagnosis_run`、`agent_step` 和 `tool_invocation` 仍作为审计记录保留。
|
||||
|
||||
如果未来需要长期保存完整聊天历史,再单独设计 `chat_message` 表,不在本阶段引入。
|
||||
|
||||
### D9:旧 `diagnosis_session` 表保留但新代码不再写入
|
||||
|
||||
新增 `diagnosis_run` 后,旧 `diagnosis_session` 不立即删除、不立即改造成 view、不直接重命名。
|
||||
|
||||
迁移策略:
|
||||
|
||||
1. 新增 `chat_session` / `diagnosis_run`。
|
||||
2. 为 `agent_step` / `tool_invocation` 新增 nullable `run_id`。
|
||||
3. 将旧 `diagnosis_session` 数据迁移/复制为 `diagnosis_run` 兼容记录。
|
||||
4. 为旧 `agent_step` / `tool_invocation` 回填对应 `run_id`。
|
||||
5. 增加必要索引和查询方法,先保持兼容读取。
|
||||
6. 新代码切换为只写 `chat_session` 和 `diagnosis_run`,并为新 step/tool 写入 `run_id`。
|
||||
7. 验证新旧数据 `run_id` 覆盖情况后,再将新写路径要求 `run_id` 非空,并补充索引/约束。
|
||||
8. 旧 `diagnosis_session` 暂时保留,用于历史核对和回滚窗口。
|
||||
9. 后续确认无依赖后,再单独归档或删除旧表。
|
||||
|
||||
### D10:提供轻量 run 列表 API
|
||||
|
||||
新增轻量查询接口,用于查看一个 Chat Session 下有哪些 Diagnosis Run。
|
||||
|
||||
```text
|
||||
GET /api/chat/session/{sessionId}/runs
|
||||
```
|
||||
|
||||
建议返回字段:
|
||||
|
||||
```text
|
||||
runId
|
||||
sessionId
|
||||
query
|
||||
status
|
||||
agentFlow
|
||||
answerPreview
|
||||
stepCount
|
||||
toolCallCount
|
||||
createdAt
|
||||
updatedAt
|
||||
```
|
||||
|
||||
该接口只读 `diagnosis_run` 主表,不展开 `agent_step` / `tool_invocation` 大字段。
|
||||
|
||||
### D11:Feedback 缺少 `runId` 时短期兼容,长期收紧
|
||||
|
||||
Feedback 新协议优先要求 `runId`。
|
||||
|
||||
短期兼容策略:
|
||||
|
||||
- 有 `runId`:绑定指定 run。
|
||||
- 无 `runId`:绑定该 `sessionId` 最新 run。
|
||||
- 无 `runId` fallback 时,在响应或日志中明确标记 `fallbackToLatestRun=true`,并返回实际绑定的 `runId`。
|
||||
|
||||
长期收紧策略:
|
||||
|
||||
- 当前端、demo 脚本和外部调用方都完成 `runId` 传递后,再评估是否将缺少 `runId` 改为参数错误。
|
||||
|
||||
### D12:同步更新 demo 脚本和 Trace UI 的 `runId` 最小支持
|
||||
|
||||
本 issue 实施范围包含 demo 脚本和 Trace UI 的最小协议适配。
|
||||
|
||||
范围:
|
||||
|
||||
- Demo 脚本读取 `/api/chat` 或 `/api/ai_ops` 返回的 `runId`。
|
||||
- Demo 脚本查询 Trace 时传 `?runId=...`。
|
||||
- Trace UI 支持 URL 参数 `?sessionId=...&runId=...`。
|
||||
- Trace UI 查询时如果有 `runId`,带上 `runId`。
|
||||
- 不在本阶段实现完整 run 列表 UI。
|
||||
|
||||
---
|
||||
|
||||
## 目标语义
|
||||
|
||||
引入明确的 `sessionId` / `runId` 分层:
|
||||
|
||||
```text
|
||||
sessionId = 多轮对话上下文
|
||||
runId = 单次诊断执行 / 单次可回放 Trace
|
||||
```
|
||||
|
||||
目标关系:
|
||||
|
||||
```text
|
||||
chat_session(sessionId)
|
||||
-> one conversation context
|
||||
-> conversation metadata / TTL / last active state
|
||||
|
||||
Redis SessionContext(sessionId)
|
||||
-> hot messageHistory cache
|
||||
-> supports prompt context window
|
||||
|
||||
diagnosis_run(runId, sessionId)
|
||||
-> one diagnosis run
|
||||
|
||||
agent_step(runId, sessionId)
|
||||
-> steps of one run
|
||||
|
||||
tool_invocation(runId, sessionId)
|
||||
-> tool calls of one run
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 建议方案
|
||||
|
||||
采用“拆分主表 + 复用现有 Trace 明细表”的方案:
|
||||
|
||||
1. 新增 `chat_session`。
|
||||
2. 新增 `diagnosis_run`。
|
||||
3. 逐步迁移当前 `diagnosis_session` 语义到 `diagnosis_run`。
|
||||
4. `agent_step` 新增 `run_id`,继续保留 `session_id` 作为冗余筛选和兼容字段。
|
||||
5. `tool_invocation` 新增 `run_id`,继续保留 `session_id` 作为冗余筛选和兼容字段。
|
||||
6. Trace API 聚合 `diagnosis_run + agent_step + tool_invocation`。
|
||||
|
||||
建议核心字段:
|
||||
|
||||
```text
|
||||
chat_session
|
||||
id
|
||||
session_id unique
|
||||
status
|
||||
message_pair_count
|
||||
created_at
|
||||
last_active_at
|
||||
expires_at
|
||||
|
||||
diagnosis_run
|
||||
id
|
||||
run_id unique
|
||||
session_id
|
||||
query
|
||||
status
|
||||
agent_flow
|
||||
answer
|
||||
self_evaluation
|
||||
feedback
|
||||
total_duration_ms
|
||||
total_token_count
|
||||
step_count
|
||||
tool_call_count
|
||||
created_at
|
||||
updated_at
|
||||
|
||||
agent_step(run_id, step_index)
|
||||
tool_invocation(run_id, id)
|
||||
```
|
||||
|
||||
理由:
|
||||
|
||||
- `chat_session` 只表达会话态,避免会话上下文和诊断结果混在一起。
|
||||
- `diagnosis_run` 只表达一次执行,天然隔离每轮 Trace、评分和反馈。
|
||||
- `agent_step` / `tool_invocation` 已足够表达 Trace 明细,暂不需要额外 trace 主表。
|
||||
- 后续如果需要统一时间线,再增加 `trace_event`,不阻塞本次隔离。
|
||||
|
||||
---
|
||||
|
||||
## 分阶段计划
|
||||
|
||||
### Phase 0:协议基线和数据边界
|
||||
|
||||
目标:先把语义定死,避免实现中反复。
|
||||
|
||||
已确认基线:
|
||||
|
||||
1. `/api/chat` 是否返回 `runId`。
|
||||
2. `GET /api/diagnosis/{sessionId}/trace` 默认查最新 run 还是要求显式传 `runId`。
|
||||
3. Feedback 是否优先绑定 `runId`,只有旧请求缺失 `runId` 时才回退最新 run。
|
||||
4. AIOps 是否和 Chat 同步接入 `runId`。
|
||||
5. 简单问答是否也创建 diagnosis run。
|
||||
|
||||
建议默认:
|
||||
|
||||
- `/api/chat` 返回 `sessionId + runId`。
|
||||
- `GET /api/diagnosis/{sessionId}/trace` 兼容查最新 run。
|
||||
- `GET /api/diagnosis/{sessionId}/trace?runId=...` 查指定 run。
|
||||
- Feedback 优先按 `runId` 绑定。
|
||||
- Chat 简单问答也创建 run。
|
||||
- AIOps 同步接入 run 隔离。
|
||||
|
||||
### Phase 1:Schema 迁移和历史数据兼容
|
||||
|
||||
目标:引入 `chat_session` / `diagnosis_run`,并保留旧数据可查询。
|
||||
|
||||
任务:
|
||||
|
||||
- Flyway 新增 `chat_session`。
|
||||
- Flyway 新增 `diagnosis_run`。
|
||||
- 为旧 `diagnosis_session` 生成兼容 `diagnosis_run` 记录。
|
||||
- 为旧 `agent_step` / `tool_invocation` 回填对应 `run_id`。
|
||||
- 增加 `find latest run by sessionId` 查询。
|
||||
- 增加 `find by runId` 查询。
|
||||
- 保留旧 `diagnosis_session` 一段时间,新代码不再写入。
|
||||
|
||||
验收:
|
||||
|
||||
- 旧 session 的 Trace 仍可查。
|
||||
- 新索引存在。
|
||||
- 不改变旧 `/api/chat` 必需字段。
|
||||
- `agent_step` / `tool_invocation` 支持 nullable `run_id` 并完成旧数据回填。
|
||||
- 新增 repository 查询可以按 `sessionId` 找最新 run、按 `runId` 找指定 run。
|
||||
|
||||
### Phase 2:Chat 写入切到 runId
|
||||
|
||||
目标:每轮 `/api/chat` 创建一个新的 run,step/tool 按 run 隔离。
|
||||
|
||||
任务:
|
||||
|
||||
- `ChatService` 每次执行生成新的 `runId`。
|
||||
- `ChatController` 确保 `chat_session` 存在并更新会话态。
|
||||
- `ChatService` 按 `runId` 创建 `diagnosis_run`。
|
||||
- `AgentLoggingHook` 写入 `agent_step.run_id`。
|
||||
- `ToolInvocationRecorder` 写入 `tool_invocation.run_id`。
|
||||
- `SessionContextHolder` 或新的上下文 holder 同时携带 `sessionId + runId`。
|
||||
- `backfillSessionMetrics` 按 `runId` 统计。
|
||||
- `EvaluationService` 按 `runId` 读取工具调用。
|
||||
|
||||
验收:
|
||||
|
||||
- 同一 `sessionId` 连续两轮后,`diagnosis_run` 有两行不同 `run_id`。
|
||||
- `/api/chat` 响应增加正式字段 `runId`。
|
||||
- 两轮 `agent_step` / `tool_invocation` 分别按各自 `run_id` 查询。
|
||||
- Redis `messagePairCount` 仍为 2,证明上下文不被破坏。
|
||||
|
||||
### Phase 3:Trace API 兼容和精确查询
|
||||
|
||||
目标:Trace API 可查最新 run,也可查指定 run,并能列出一个 session 下的 run。
|
||||
|
||||
任务:
|
||||
|
||||
- `GET /api/diagnosis/{sessionId}/trace` 从 `diagnosis_run` 默认解析最新 run。
|
||||
- 增加 `runId` query 参数。
|
||||
- Trace response 增加 `runId`。
|
||||
- 新增 `GET /api/chat/session/{sessionId}/runs`。
|
||||
|
||||
验收:
|
||||
|
||||
- 不传 `runId` 返回最新 run。
|
||||
- 传第一轮 `runId` 只返回第一轮 step/tool。
|
||||
- 传第二轮 `runId` 只返回第二轮 step/tool。
|
||||
- run 列表 API 只返回轻量 run 摘要,不展开 trace 明细。
|
||||
- Demo 脚本和 Trace UI 的 `runId` 最小适配按 OpenSpec tasks 放到 Phase 6,避免 Phase 3 同时混入前端/脚本范围。
|
||||
|
||||
### Phase 4:Feedback 和 CaseLibrary 绑定 run
|
||||
|
||||
目标:反馈明确评价哪一轮诊断。
|
||||
|
||||
任务:
|
||||
|
||||
- Feedback request 支持 `runId`。
|
||||
- 旧请求只有 `sessionId` 时短期绑定最新 run,并显式标记 fallback。
|
||||
- `case_library.diagnosis_id` 新数据保存 `run_id`。
|
||||
- `CaseLibraryService` 以 run 为来源生成 case,并用 `run_id` 做新数据幂等键。
|
||||
- 文档说明 `diagnosis_id` 的过渡语义:旧数据可能是 `session_id`,新数据是 `run_id`。
|
||||
|
||||
验收:
|
||||
|
||||
- 同 session 多轮后,对第一轮提交 feedback 不会覆盖第二轮。
|
||||
- useful 生成 case 时能定位到对应 run 的 query/answer。
|
||||
|
||||
### Phase 5:AIOps 同步 run 隔离
|
||||
|
||||
目标:AIOps 使用同样的 run 语义,避免另一条入口继续混杂。
|
||||
|
||||
任务:
|
||||
|
||||
- `AiOpsService` 生成并返回/透出 `runId`。
|
||||
- AIOps `agent_step` / `tool_invocation` 按 `runId` 隔离。
|
||||
- AIOps rule evaluation 写入当前 run 的 `diagnosis_run.self_evaluation.aiops_rule_evaluation`。
|
||||
- AIOps Trace 查询兼容 `sessionId + runId`。
|
||||
|
||||
验收:
|
||||
|
||||
- 同一 AIOps `sessionId` 重跑不会混合 step/tool。
|
||||
- AIOps rule evaluation 只读取当前 run 工具调用。
|
||||
|
||||
---
|
||||
|
||||
## 暂不做
|
||||
|
||||
1. 暂不新增 `diagnosis_trace` 或 `trace_event` 主表。
|
||||
2. 暂不做完整 run 列表 UI。
|
||||
3. 暂不删除历史 Trace 数据。
|
||||
4. 暂不改变 Redis 多轮上下文窗口策略。
|
||||
5. 暂不立即物理删除旧 `diagnosis_session` 表。
|
||||
|
||||
---
|
||||
|
||||
## 风险
|
||||
|
||||
### 1. 兼容风险
|
||||
|
||||
现有脚本、Trace 页面和反馈接口可能只知道 `sessionId`。
|
||||
|
||||
缓解:保留 `sessionId` 默认查最新 run 的行为。
|
||||
|
||||
### 2. 异步上下文风险
|
||||
|
||||
工具调用和 Agent hook 依赖 ThreadLocal / RunnableConfig 传递上下文。
|
||||
|
||||
缓解:统一上下文对象,明确 `sessionId` 和 `runId` 必须同时传递。
|
||||
|
||||
### 3. 历史数据回填风险
|
||||
|
||||
旧数据没有真实 run 边界,只能按当前 `diagnosis_session` 生成一条兼容 `diagnosis_run`。
|
||||
|
||||
缓解:旧数据视为单 run,不尝试拆分历史混合数据。
|
||||
|
||||
### 4. 评分口径变化风险
|
||||
|
||||
按 `runId` 隔离后,工具调用数和 evidence score 可能下降,但语义更正确。
|
||||
|
||||
缓解:更新 eval fixture 和 baseline,记录这是预期行为变化。
|
||||
|
||||
---
|
||||
|
||||
## 已收敛问题
|
||||
|
||||
本 issue 当前已经收敛以下设计边界:
|
||||
|
||||
- `runId` 是正式 API 字段。
|
||||
- `chat_session` 和 `diagnosis_run` 拆分为两张主表。
|
||||
- Trace 明细继续由 `agent_step` / `tool_invocation` 承载。
|
||||
- `chat_session` 只保存元数据,不保存完整对话历史。
|
||||
- 旧 `diagnosis_session` 保留但新代码不再写入。
|
||||
- 提供轻量 run 列表 API。
|
||||
- Feedback 缺少 `runId` 时短期兼容、长期收紧。
|
||||
- Demo 脚本和 Trace UI 做 `runId` 最小支持。
|
||||
|
||||
---
|
||||
|
||||
## 相关文件
|
||||
|
||||
- `mvp/architecture/session-trace-lifecycle.md`
|
||||
- `mvp/architecture/data-model.md`
|
||||
- `mvp/tables/聊天会话表-chat_session.md`
|
||||
- `mvp/tables/诊断运行表-diagnosis_run.md`
|
||||
- `mvp/tables/诊断会话表-diagnosis_session.md`
|
||||
- `mvp/tables/Agent步骤表-agent_step.md`
|
||||
- `mvp/tables/工具调用表-tool_invocation.md`
|
||||
- `mvp/tables/案例库表-case_library.md`
|
||||
- `src/main/java/com/superbiz/agent/controller/ChatController.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ChatService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/AiOpsService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/DiagnosisTraceService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/EvaluationService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/FeedbackService.java`
|
||||
- `src/main/java/com/superbiz/agent/service/CaseLibraryService.java`
|
||||
- `src/main/java/com/superbiz/agent/hook/AgentLoggingHook.java`
|
||||
- `src/main/java/com/superbiz/agent/service/ToolInvocationRecorder.java`
|
||||
- `src/main/java/com/superbiz/agent/util/SessionContextHolder.java`
|
||||
- `src/main/resources/db/migration/V005__create_session_storage.sql`
|
||||
@@ -0,0 +1,45 @@
|
||||
# Agent 步骤表:agent_step
|
||||
|
||||
**状态**:当前表
|
||||
**来源**:`V005__create_session_storage.sql`、`V006__fix_agent_step_json_to_text.sql`、`AgentStep`
|
||||
|
||||
## 定位
|
||||
|
||||
`agent_step` 记录一次诊断运行中每个 Agent 步骤的模型输入、输出、耗时和 Token 消耗。`run_id` 是执行隔离边界;Trace 页面展示顺序以 Trace API 返回顺序为准。
|
||||
|
||||
## 字段
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
|---|---|---|---|
|
||||
| `id` | BIGINT | 是 | 自增主键 |
|
||||
| `session_id` | VARCHAR(64) | 是 | 所属会话目录 ID,保留用于粗粒度过滤和兼容 |
|
||||
| `run_id` | VARCHAR(64) | 否 | 所属 `diagnosis_run.run_id`;新执行应写入 |
|
||||
| `step_index` | INT | 是 | 步骤序号,从 0 开始 |
|
||||
| `agent_name` | VARCHAR(32) | 是 | Agent 名称,例如 planner、executor、verifier、composer |
|
||||
| `model_input` | TEXT | 否 | 模型输入摘要;`V006` 已从 JSON 改为 TEXT |
|
||||
| `model_output` | TEXT | 否 | 模型输出摘要;`V006` 已从 JSON 改为 TEXT |
|
||||
| `thought` | TEXT | 否 | Agent 思考过程或调试摘要 |
|
||||
| `has_tool_call` | BOOLEAN | 否 | 本步骤是否触发工具调用 |
|
||||
| `duration_ms` | INT | 否 | 本步骤耗时 |
|
||||
| `token_count` | INT | 否 | 本步骤 Token 消耗 |
|
||||
| `created_at` | DATETIME | 是 | 创建时间 |
|
||||
|
||||
## 索引
|
||||
|
||||
| 索引 | 字段 | 用途 |
|
||||
|---|---|---|
|
||||
| `idx_session_step` | `session_id, step_index` | 历史兼容和粗粒度排查 |
|
||||
| `idx_agent_step_run_step` | `run_id, step_index` | 按运行筛选步骤并辅助顺序查询 |
|
||||
| `idx_agent_name` | `agent_name` | 按 Agent 类型筛选 |
|
||||
|
||||
## 关系
|
||||
|
||||
- `agent_step.run_id` 逻辑关联 `diagnosis_run.run_id`。
|
||||
- `agent_step.session_id` 保留为 `chat_session.session_id` 的冗余关联,便于粗粒度过滤和兼容查询。
|
||||
- `tool_invocation.step_id` 可关联 `agent_step.id`,但当前允许为空且不强制外键。
|
||||
|
||||
## 注意点
|
||||
|
||||
- 前端展示步骤时应使用 Trace API 返回顺序;服务端会在同一 `run_id` 范围内整理步骤顺序。
|
||||
- 新 Trace、Verifier 和评测读路径应按 `run_id` 取数,避免同一 `sessionId` 多轮诊断混入。
|
||||
- Verifier 应在 Executor 循环完成后出现;如果 `step_index` 中 Verifier 提前,通常意味着编排或记录顺序有问题。
|
||||
@@ -0,0 +1,46 @@
|
||||
# MVP 数据表索引
|
||||
|
||||
**更新日期**:2026-07-10
|
||||
**状态**:当前表文档入口
|
||||
|
||||
本目录保存当前 MVP 使用的数据表说明。详细结构以 Flyway migration 和实体类为准;本目录用于面试讲解、排查索引和快速理解数据流。
|
||||
|
||||
## 当前表
|
||||
|
||||
| 表 | 用途 | 文档 |
|
||||
|---|---|---|
|
||||
| `chat_session` | 会话目录元数据,保存同一个 `sessionId` 的多轮会话状态快照 | [聊天会话表-chat_session.md](聊天会话表-chat_session.md) |
|
||||
| `diagnosis_run` | 运行级主记录,保存一次 Chat/AIOps 诊断的 query、状态、答案、自评估和反馈 | [诊断运行表-diagnosis_run.md](诊断运行表-diagnosis_run.md) |
|
||||
| `agent_step` | Agent 步骤记录,按 `run_id` 隔离回放执行链路 | [Agent步骤表-agent_step.md](Agent步骤表-agent_step.md) |
|
||||
| `tool_invocation` | 工具调用记录,按 `run_id` 支撑 Trace、Verifier 和评测 | [工具调用表-tool_invocation.md](工具调用表-tool_invocation.md) |
|
||||
| `api_document` | 知识库文档元数据,和向量库 chunk 通过 `doc_id` 关联 | [文档元数据表-api_document.md](文档元数据表-api_document.md) |
|
||||
| `knowledge_domain` | 知识域元数据,支撑 RAG domain hint 和检索策略 | [知识域表-knowledge_domain.md](知识域表-knowledge_domain.md) |
|
||||
| `case_library` | 用户反馈沉淀出的高质量诊断案例 | [案例库表-case_library.md](案例库表-case_library.md) |
|
||||
| `diagnosis_session` | 历史兼容和回滚表,新执行写入不再依赖它 | [诊断会话表-diagnosis_session.md](诊断会话表-diagnosis_session.md) |
|
||||
|
||||
## 已归档表
|
||||
|
||||
| 表 | 归档原因 | 文档 |
|
||||
|---|---|---|
|
||||
| `diagnosis_record` | 已由 `V007` 删除,历史上被 `diagnosis_session + agent_step + tool_invocation` 替代;当前新模型是 `chat_session + diagnosis_run + trace detail` | [archive/2026-07-09-doc-cleanup/旧诊断记录表-diagnosis_record.md](archive/2026-07-09-doc-cleanup/旧诊断记录表-diagnosis_record.md) |
|
||||
|
||||
## 核心关系
|
||||
|
||||
```text
|
||||
chat_session.session_id
|
||||
-> diagnosis_run.session_id
|
||||
-> agent_step.run_id
|
||||
-> tool_invocation.run_id
|
||||
-> case_library.diagnosis_id (new AUTO cases use run_id)
|
||||
|
||||
diagnosis_session.session_id
|
||||
-> historical compatibility / rollback only
|
||||
|
||||
api_document.doc_id
|
||||
-> vector chunk metadata.docId / doc_id
|
||||
|
||||
knowledge_domain.domain_id
|
||||
-> api_document metadata.category / vector chunk metadata.category
|
||||
```
|
||||
|
||||
当前实现主要使用逻辑关联,不依赖数据库外键。
|
||||
@@ -1,332 +0,0 @@
|
||||
# api_document - 文档元数据表
|
||||
|
||||
## 表定位
|
||||
|
||||
**文档管理表**:管理接口文档的元信息,不负责文档检索(检索由 Milvus 负责)
|
||||
|
||||
## 设计理念
|
||||
|
||||
### 文档管理,不是文档检索
|
||||
|
||||
**核心定位**:
|
||||
- MySQL 负责文档元数据管理(状态、版本、去重)
|
||||
- Milvus 负责文档内容存储和检索
|
||||
- 通过 doc_id 关联两者
|
||||
|
||||
**MVP版本原则**:
|
||||
- ✅ 最简字段,满足基本管理需求
|
||||
- ✅ 文件去重(基于 file_hash)
|
||||
- ✅ 状态追踪(索引进度)
|
||||
- ✅ 硬删除(同步删除 Milvus 数据)
|
||||
- ❌ 暂不支持:软删除、启用开关、版本管理(Phase 2)
|
||||
|
||||
---
|
||||
|
||||
## 表结构(MVP版)
|
||||
|
||||
```sql
|
||||
CREATE TABLE api_document (
|
||||
-- 主键
|
||||
id BIGINT PRIMARY KEY AUTO_INCREMENT,
|
||||
doc_id VARCHAR(64) UNIQUE NOT NULL COMMENT '文档唯一ID(UUID),关联Milvus',
|
||||
|
||||
-- 文档分类
|
||||
fault_category VARCHAR(32) DEFAULT 'EXTERNAL_API' COMMENT '文档类别',
|
||||
fault_source VARCHAR(128) COMMENT '文档归属(省份/服务名)',
|
||||
api_name VARCHAR(128) COMMENT '接口名称',
|
||||
version VARCHAR(32) DEFAULT 'v1.0' COMMENT '文档版本',
|
||||
|
||||
-- 文件信息
|
||||
file_name VARCHAR(256) NOT NULL COMMENT '原始文件名',
|
||||
file_path VARCHAR(512) COMMENT '文件存储路径',
|
||||
file_hash VARCHAR(64) COMMENT '文件MD5 hash(用于去重)',
|
||||
file_size BIGINT COMMENT '文件大小(字节)',
|
||||
|
||||
-- 索引状态
|
||||
status VARCHAR(16) DEFAULT 'PENDING' COMMENT '索引状态(PENDING/PROCESSING/INDEXED/FAILED)',
|
||||
chunk_count INT DEFAULT 0 COMMENT '分块数量',
|
||||
error_message TEXT COMMENT '失败原因',
|
||||
|
||||
-- 时间字段
|
||||
indexed_at DATETIME COMMENT '索引完成时间',
|
||||
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
|
||||
updated_at DATETIME DEFAULT CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP,
|
||||
|
||||
-- 索引
|
||||
UNIQUE INDEX uk_file_hash (file_hash),
|
||||
INDEX idx_doc_id (doc_id),
|
||||
INDEX idx_fault_source (fault_source),
|
||||
INDEX idx_status (status),
|
||||
INDEX idx_created_at (created_at)
|
||||
) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4 COMMENT='文档元数据表(MVP版)';
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 字段说明
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
|------|------|------|------|
|
||||
| doc_id | VARCHAR(64) | 是 | **核心**:文档唯一ID,关联 Milvus |
|
||||
| fault_category | VARCHAR(32) | 否 | 文档类别 |
|
||||
| fault_source | VARCHAR(128) | 否 | 文档归属(省份/服务名)|
|
||||
| api_name | VARCHAR(128) | 否 | 接口名称 |
|
||||
| version | VARCHAR(32) | 否 | 文档版本 |
|
||||
| file_name | VARCHAR(256) | 是 | 原始文件名 |
|
||||
| file_path | VARCHAR(512) | 否 | 文件存储路径 |
|
||||
| file_hash | VARCHAR(64) | 否 | **去重关键**:文件MD5 |
|
||||
| file_size | BIGINT | 否 | 文件大小 |
|
||||
| status | VARCHAR(16) | 是 | **状态追踪**:PENDING/PROCESSING/INDEXED/FAILED |
|
||||
| chunk_count | INT | 否 | 分块数量 |
|
||||
| error_message | TEXT | 否 | 失败原因 |
|
||||
| indexed_at | DATETIME | 否 | 索引完成时间 |
|
||||
|
||||
---
|
||||
|
||||
## 核心设计决策
|
||||
|
||||
### 1. doc_id:MySQL 与 Milvus 的桥梁
|
||||
|
||||
```
|
||||
作用:
|
||||
- MySQL:通过 doc_id 管理文档元数据
|
||||
- Milvus:每个 chunk 的 metadata 中携带 doc_id
|
||||
|
||||
关联关系:
|
||||
api_document (MySQL)
|
||||
doc_id: doc-001
|
||||
↓ 1:N
|
||||
Milvus chunks
|
||||
chunk_1: {doc_id: 'doc-001', text: '...', vector: [...]}
|
||||
chunk_2: {doc_id: 'doc-001', text: '...', vector: [...]}
|
||||
|
||||
管理操作:
|
||||
- 删除文档:
|
||||
DELETE FROM milvus_collection WHERE metadata["doc_id"] == 'doc-001';
|
||||
DELETE FROM api_document WHERE doc_id = 'doc-001';
|
||||
```
|
||||
|
||||
### 2. file_hash:文件去重
|
||||
|
||||
```
|
||||
去重流程:
|
||||
1. 用户上传文件
|
||||
↓
|
||||
2. 计算文件 MD5
|
||||
file_hash = md5(file_content)
|
||||
↓
|
||||
3. 检查是否已存在
|
||||
SELECT * FROM api_document WHERE file_hash = 'abc123...';
|
||||
↓
|
||||
4a. 如果存在 → 提示"文档已存在"
|
||||
4b. 如果不存在 → 继续导入
|
||||
|
||||
唯一约束:UNIQUE INDEX uk_file_hash (file_hash)
|
||||
```
|
||||
|
||||
### 3. status:状态追踪
|
||||
|
||||
```
|
||||
状态流转:
|
||||
PENDING (待处理)
|
||||
↓
|
||||
PROCESSING (处理中)
|
||||
↓ 成功
|
||||
INDEXED (已索引)
|
||||
↓ 失败
|
||||
FAILED (失败)
|
||||
|
||||
用途:
|
||||
- 批量导入时监控进度
|
||||
- 失败重试
|
||||
- 统计索引成功率
|
||||
```
|
||||
|
||||
### 4. 硬删除策略(MVP)
|
||||
|
||||
```
|
||||
删除文档时:
|
||||
1. 删除 Milvus 中的所有分块
|
||||
2. 删除 MySQL 元数据
|
||||
3. 可选:删除原始文件
|
||||
|
||||
特点:
|
||||
- 简单直接
|
||||
- 数据彻底删除
|
||||
- 不可恢复(需谨慎)
|
||||
|
||||
Phase 2 可增强:
|
||||
- 软删除(archived_at)
|
||||
- 启用开关(enabled)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 数据流
|
||||
|
||||
### 场景1:导入新文档
|
||||
|
||||
```
|
||||
1. 用户上传文件
|
||||
↓
|
||||
2. 计算 hash
|
||||
↓
|
||||
3. 检查去重(MySQL)
|
||||
↓
|
||||
4. 插入元数据(status=PROCESSING)
|
||||
↓
|
||||
5. 后台处理:解析 → 分块 → 向量化 → 存入 Milvus
|
||||
↓
|
||||
6. 更新状态(status=INDEXED, chunk_count=15)
|
||||
```
|
||||
|
||||
### 场景2:删除文档
|
||||
|
||||
```
|
||||
1. 用户删除文档
|
||||
↓
|
||||
2. 删除 Milvus 数据(WHERE metadata["doc_id"] == 'xxx')
|
||||
↓
|
||||
3. 删除 MySQL 元数据
|
||||
↓
|
||||
4. 可选:删除原始文件
|
||||
```
|
||||
|
||||
### 场景3:重新索引
|
||||
|
||||
```
|
||||
1. 删除旧数据(Milvus + MySQL)
|
||||
↓
|
||||
2. 重新导入(同场景1)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 典型查询
|
||||
|
||||
```sql
|
||||
-- 查看文档列表
|
||||
SELECT doc_id, file_name, version, status, chunk_count, indexed_at
|
||||
FROM api_document
|
||||
WHERE fault_source = '广东'
|
||||
AND status = 'INDEXED'
|
||||
ORDER BY indexed_at DESC;
|
||||
|
||||
-- 查询失败的文档
|
||||
SELECT doc_id, file_name, error_message
|
||||
FROM api_document
|
||||
WHERE status = 'FAILED';
|
||||
|
||||
-- 统计各状态文档数量
|
||||
SELECT status, COUNT(*) as count
|
||||
FROM api_document
|
||||
GROUP BY status;
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 与 Milvus 的协作
|
||||
|
||||
### Milvus Collection Schema
|
||||
|
||||
```python
|
||||
{
|
||||
"collection_name": "api_doc_collection",
|
||||
"fields": [
|
||||
{"name": "id", "type": "VARCHAR", "is_primary": true},
|
||||
{"name": "content", "type": "VARCHAR"},
|
||||
{"name": "vector", "type": "FLOAT_VECTOR", "dim": 1536},
|
||||
{"name": "metadata", "type": "JSON"}
|
||||
]
|
||||
}
|
||||
|
||||
# metadata 结构
|
||||
{
|
||||
"doc_id": "doc-001", # 关联 MySQL
|
||||
"_source": "/path/to/file",
|
||||
"_file_name": "xxx.docx",
|
||||
"chunkIndex": 0,
|
||||
"totalChunks": 15
|
||||
}
|
||||
```
|
||||
|
||||
### Java 代码示例
|
||||
|
||||
```java
|
||||
// 插入时携带 doc_id
|
||||
Map<String, Object> metadata = new HashMap<>();
|
||||
metadata.put("doc_id", docId); // 关联 MySQL
|
||||
metadata.put("_source", filePath);
|
||||
metadata.put("chunkIndex", chunkIndex);
|
||||
|
||||
// 删除文档的所有分块
|
||||
String expr = String.format("metadata[\"doc_id\"] == \"%s\"", docId);
|
||||
milvusClient.delete(DeleteParam.newBuilder()
|
||||
.withCollectionName(COLLECTION_NAME)
|
||||
.withExpr(expr)
|
||||
.build());
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 数据示例
|
||||
|
||||
```sql
|
||||
-- 外部接口文档
|
||||
INSERT INTO api_document VALUES
|
||||
(1, 'doc-001', 'EXTERNAL_API', '广东', '社保查询', 'v2.1',
|
||||
'广东社保查询v2.1.docx', '/docs/guangdong/social-v2.1.docx',
|
||||
'abc123...', 1048576,
|
||||
'INDEXED', 15, NULL, '2024-06-15 10:30:00', NOW(), NOW());
|
||||
|
||||
-- 内部服务文档
|
||||
INSERT INTO api_document VALUES
|
||||
(2, 'doc-002', 'INTERNAL_ERROR', 'order-service', '订单服务API', 'v1.0',
|
||||
'订单服务API文档.pdf', '/docs/internal/order-service-api.pdf',
|
||||
'def456...', 2097152,
|
||||
'INDEXED', 20, NULL, '2024-06-14 15:20:00', NOW(), NOW());
|
||||
|
||||
-- 处理失败的文档
|
||||
INSERT INTO api_document VALUES
|
||||
(3, 'doc-003', 'EXTERNAL_API', '江苏', '公积金查询', 'v1.5',
|
||||
'江苏公积金查询.html', '/docs/jiangsu/fund-v1.5.html',
|
||||
'ghi789...', 512000,
|
||||
'FAILED', 0, '不支持HTML格式', NULL, NOW(), NOW());
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 数据量预估
|
||||
|
||||
```
|
||||
预估:100-200 条
|
||||
- 外部接口文档:50-100 条
|
||||
- 内部服务文档:20-50 条
|
||||
- 其他文档:30-50 条
|
||||
|
||||
存储:
|
||||
- 单条记录:约 1KB
|
||||
- 200 条:约 200KB
|
||||
|
||||
结论:数据量很小
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## MVP 版本的简化
|
||||
|
||||
```
|
||||
Phase 1(当前):
|
||||
✅ 基础字段和表结构
|
||||
✅ 文件去重(file_hash)
|
||||
✅ 状态追踪(status)
|
||||
✅ 硬删除
|
||||
✅ 通过 doc_id 关联 Milvus
|
||||
|
||||
Phase 2(未来增强):
|
||||
❌ enabled(启用开关)
|
||||
❌ archived_at(软删除)
|
||||
❌ batch_id(批次管理)
|
||||
❌ status 细化
|
||||
❌ tags(标签分类)
|
||||
```
|
||||
+3
-1
@@ -1,4 +1,6 @@
|
||||
# diagnosis_record - 诊断记录表
|
||||
# diagnosis_record - 旧诊断记录表
|
||||
|
||||
> 归档说明:`diagnosis_record` 已在 `V007__drop_diagnosis_record.sql` 中删除,当前主模型是 `diagnosis_session + agent_step + tool_invocation`。本文只用于追溯早期设计。
|
||||
|
||||
## 表定位
|
||||
|
||||
@@ -1,265 +0,0 @@
|
||||
# case_library - 案例库表
|
||||
|
||||
## 表定位
|
||||
|
||||
**知识沉淀表**:存储高质量诊断案例,支持相似案例推荐
|
||||
|
||||
## 设计理念
|
||||
|
||||
### 知识沉淀,系统越用越智能
|
||||
|
||||
**核心价值**:
|
||||
- 质量过滤:只存储高质量案例(成功诊断 + 用户反馈有用)
|
||||
- 知识沉淀:历史诊断经验可复用
|
||||
- 提升准确率:相似问题提供历史参考
|
||||
- 加速诊断:快速推荐相似案例
|
||||
|
||||
**MVP版本设计原则**:
|
||||
- ✅ 能用:满足基本案例推荐功能
|
||||
- ✅ 简单:字段不多,逻辑清晰
|
||||
- ✅ 可扩展:后续可增加字段
|
||||
|
||||
---
|
||||
|
||||
## 表结构(MVP版)
|
||||
|
||||
```sql
|
||||
CREATE TABLE case_library (
|
||||
-- 主键
|
||||
id BIGINT PRIMARY KEY AUTO_INCREMENT,
|
||||
case_id VARCHAR(64) UNIQUE NOT NULL COMMENT '案例唯一ID(UUID)',
|
||||
|
||||
-- 来源关联
|
||||
diagnosis_id VARCHAR(64) COMMENT '关联诊断记录(可选,人工录入时为空)',
|
||||
source_type VARCHAR(16) DEFAULT 'AUTO' COMMENT '来源类型(AUTO:自动生成/MANUAL:人工录入)',
|
||||
|
||||
-- 案例分类
|
||||
fault_category VARCHAR(32) COMMENT '故障类别(EXTERNAL_API/INTERNAL_ERROR/DATABASE...)',
|
||||
fault_source VARCHAR(128) COMMENT '故障源(省份/服务名/类名...)',
|
||||
fault_target VARCHAR(256) COMMENT '故障目标(接口URL/方法名/SQL...)',
|
||||
error_code VARCHAR(64) COMMENT '错误码',
|
||||
|
||||
-- 案例内容
|
||||
title VARCHAR(256) NOT NULL COMMENT '案例标题(简短描述)',
|
||||
root_cause TEXT NOT NULL COMMENT '根因分析',
|
||||
solution TEXT NOT NULL COMMENT '解决方案',
|
||||
|
||||
-- 简单统计
|
||||
reference_count INT DEFAULT 0 COMMENT '引用次数(被推荐的次数)',
|
||||
|
||||
-- 元数据
|
||||
created_by VARCHAR(64) COMMENT '创建人',
|
||||
created_at DATETIME DEFAULT CURRENT_TIMESTAMP,
|
||||
updated_at DATETIME DEFAULT CURRENT_TIMESTAMP ON UPDATE CURRENT_TIMESTAMP,
|
||||
|
||||
-- 索引
|
||||
INDEX idx_fault_category (fault_category),
|
||||
INDEX idx_error_code (error_code),
|
||||
INDEX idx_fault_source (fault_source),
|
||||
INDEX idx_fault_target (fault_target(100)),
|
||||
INDEX idx_diagnosis_id (diagnosis_id),
|
||||
INDEX idx_reference_count (reference_count),
|
||||
INDEX idx_created_at (created_at)
|
||||
) ENGINE=InnoDB DEFAULT CHARSET=utf8mb4 COMMENT='案例库表(MVP版)';
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 字段说明
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
|------|------|------|------|
|
||||
| case_id | VARCHAR(64) | 是 | 案例唯一标识(UUID)|
|
||||
| diagnosis_id | VARCHAR(64) | 否 | 关联诊断记录(人工录入时为空)|
|
||||
| source_type | VARCHAR(16) | 是 | 来源:AUTO(自动)/MANUAL(人工)|
|
||||
| fault_category | VARCHAR(32) | 否 | 故障类别 |
|
||||
| fault_source | VARCHAR(128) | 否 | 故障源 |
|
||||
| fault_target | VARCHAR(256) | 否 | 故障目标(与 diagnosis_record 一致)|
|
||||
| error_code | VARCHAR(64) | 否 | 错误码 |
|
||||
| title | VARCHAR(256) | 是 | 案例标题 |
|
||||
| root_cause | TEXT | 是 | 根因分析(核心内容)|
|
||||
| solution | TEXT | 是 | 解决方案(核心内容)|
|
||||
| reference_count | INT | 是 | 引用次数(用于排序)|
|
||||
|
||||
---
|
||||
|
||||
## 核心设计决策
|
||||
|
||||
### 1. 案例来源
|
||||
|
||||
```
|
||||
来源1:自动生成(source_type=AUTO)
|
||||
├─ 触发条件:诊断成功 + 用户反馈"有用"
|
||||
├─ 关联诊断:diagnosis_id 不为空
|
||||
└─ 质量保证:用户验证过
|
||||
|
||||
来源2:人工录入(source_type=MANUAL)
|
||||
├─ 运维团队总结的经典案例
|
||||
├─ diagnosis_id 为空
|
||||
└─ 质量最高
|
||||
|
||||
注意:诊断失败或用户反馈"无用"的不自动生成案例
|
||||
```
|
||||
|
||||
### 2. 简化的评分机制(MVP)
|
||||
|
||||
```
|
||||
MVP版本:只按 reference_count 排序
|
||||
- 引用次数多的排前面
|
||||
- 简单有效
|
||||
|
||||
Phase 2 可增强:
|
||||
- 增加 useful_count(用户反馈有用次数)
|
||||
- 增加 score(综合评分)
|
||||
- 增加 is_featured(人工标记的经典案例)
|
||||
```
|
||||
|
||||
### 3. 与 diagnosis_record 的关系
|
||||
|
||||
```
|
||||
关系:一对一(可选)
|
||||
- 一次诊断 → 可以生成一个案例
|
||||
- 通过 diagnosis_id 关联
|
||||
- diagnosis_id 可为空(人工录入案例)
|
||||
|
||||
流程:
|
||||
diagnosis_record(成功)
|
||||
↓
|
||||
用户反馈"有用"
|
||||
↓
|
||||
自动生成 case_library
|
||||
↓
|
||||
后续可人工修正、合并相似案例
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 数据示例
|
||||
|
||||
### 示例1:外部接口故障案例
|
||||
```sql
|
||||
INSERT INTO case_library VALUES
|
||||
(1, 'case-001', 'diag-001', 'AUTO', 'EXTERNAL_API', '广东', '/api/v1/guangdong/social-security', '40003',
|
||||
'广东社保查询idCard字段缺失',
|
||||
'请求报文中未传入idCard字段,导致参数校验失败',
|
||||
'前端表单增加idCard必填校验;后端增加参数校验提示',
|
||||
15, 'system', NOW(), NOW());
|
||||
```
|
||||
|
||||
### 示例2:内部错误案例
|
||||
```sql
|
||||
INSERT INTO case_library VALUES
|
||||
(2, 'case-002', 'diag-045', 'AUTO', 'INTERNAL_ERROR', 'order-service', 'OrderController.createOrder()', 'NullPointerException',
|
||||
'订单服务创建订单空指针异常',
|
||||
'OrderController.createOrder()方法中user对象为null,未做空判断',
|
||||
'在第45行添加空判断:if (user == null) throw new BizException("用户信息不存在")',
|
||||
8, 'system', NOW(), NOW());
|
||||
```
|
||||
|
||||
### 示例3:人工录入案例
|
||||
```sql
|
||||
INSERT INTO case_library VALUES
|
||||
(3, 'case-003', NULL, 'MANUAL', 'DATABASE', 'mysql-master-01', 'UPDATE orders SET status=? WHERE order_id=?', '1213',
|
||||
'订单库存更新死锁通用处理',
|
||||
'两个事务互相等待对方释放锁',
|
||||
'调整事务加锁顺序:统一先锁订单,再锁库存;或使用乐观锁',
|
||||
3, 'admin', NOW(), NOW());
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 典型查询
|
||||
|
||||
### 精确匹配查询
|
||||
```sql
|
||||
-- 按错误码查询
|
||||
SELECT * FROM case_library
|
||||
WHERE error_code = '40003'
|
||||
ORDER BY reference_count DESC
|
||||
LIMIT 5;
|
||||
|
||||
-- 按故障类别 + 错误码 + 故障目标查询
|
||||
SELECT * FROM case_library
|
||||
WHERE fault_category = 'INTERNAL_ERROR'
|
||||
AND error_code = 'NullPointerException'
|
||||
AND fault_target = 'OrderController.createOrder()'
|
||||
ORDER BY reference_count DESC
|
||||
LIMIT 5;
|
||||
```
|
||||
|
||||
### 统计分析
|
||||
```sql
|
||||
-- 统计案例分布
|
||||
SELECT
|
||||
fault_category,
|
||||
COUNT(*) as count,
|
||||
AVG(reference_count) as avg_reference
|
||||
FROM case_library
|
||||
GROUP BY fault_category
|
||||
ORDER BY count DESC;
|
||||
|
||||
-- Top 引用案例
|
||||
SELECT title, reference_count, created_at
|
||||
FROM case_library
|
||||
ORDER BY reference_count DESC
|
||||
LIMIT 10;
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 与 Milvus 的配合
|
||||
|
||||
### 混合检索策略
|
||||
|
||||
```
|
||||
1. 精确匹配(MySQL)
|
||||
- 按 error_code 查询
|
||||
- 按 fault_category + fault_source 查询
|
||||
- 优点:快速、准确
|
||||
|
||||
2. 语义检索(Milvus)
|
||||
- 将案例内容向量化
|
||||
- 按语义相似度查询
|
||||
- 优点:能找到相似但不同错误码的案例
|
||||
|
||||
3. 混合策略(推荐)
|
||||
Step 1: 先精确匹配(MySQL)
|
||||
Step 2: 如果结果 < 3 个,补充语义检索(Milvus)
|
||||
Step 3: 合并去重,按 reference_count 排序
|
||||
Step 4: 返回 Top 5
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 数据量预估
|
||||
|
||||
```
|
||||
预估:500-1000 条
|
||||
- 初期:每月新增 10-20 条
|
||||
- 稳定期:每月新增 5-10 条
|
||||
- 总量:1-2 年达到稳定
|
||||
|
||||
存储:
|
||||
- 单条记录:约 2KB
|
||||
- 1000 条:约 2MB
|
||||
|
||||
结论:数据量很小
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## MVP 版本的简化
|
||||
|
||||
```
|
||||
Phase 1(当前):
|
||||
✅ 基础字段和表结构
|
||||
✅ 自动生成案例
|
||||
✅ 人工录入案例
|
||||
✅ 按 reference_count 简单排序
|
||||
|
||||
Phase 2(未来增强):
|
||||
❌ useful_count + score(复杂评分)
|
||||
❌ 版本管理
|
||||
❌ 标签分类(tags)
|
||||
❌ 案例合并功能
|
||||
```
|
||||
@@ -0,0 +1,69 @@
|
||||
# 工具调用表:tool_invocation
|
||||
|
||||
**状态**:当前表
|
||||
**来源**:`V005__create_session_storage.sql`、`V010__add_relevance_level_to_tool_invocation.sql`、`ToolInvocation`
|
||||
|
||||
## 定位
|
||||
|
||||
`tool_invocation` 记录 Agent 在一次诊断运行中显式调用工具的事实,包括工具名、入参、输出摘要、检索层级、证据引用和失败信息。它是 Trace、Verifier、评测和人工排查的共同数据源。
|
||||
|
||||
## 字段
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
|---|---|---|---|
|
||||
| `id` | BIGINT | 是 | 自增主键 |
|
||||
| `session_id` | VARCHAR(64) | 是 | 所属会话目录 ID,保留用于粗粒度过滤和兼容 |
|
||||
| `run_id` | VARCHAR(64) | 否 | 所属 `diagnosis_run.run_id`;新执行应写入 |
|
||||
| `step_id` | BIGINT | 否 | 可关联 `agent_step.id` |
|
||||
| `tool_name` | VARCHAR(64) | 是 | 工具名称,例如 `lookup_knowledge`、日志查询、指标查询 |
|
||||
| `input_params` | JSON | 是 | 工具入参 |
|
||||
| `output_preview` | TEXT | 否 | 工具输出摘要或前缀 |
|
||||
| `output_length` | INT | 否 | 工具输出字符数 |
|
||||
| `retrieval_layer` | VARCHAR(8) | 否 | 检索层级,例如 `L0`、`L1`、`L0+L1` |
|
||||
| `l0_match_count` | INT | 否 | L0 命中数量 |
|
||||
| `l1_match_count` | INT | 否 | L1 命中数量 |
|
||||
| `is_truncated` | BOOLEAN | 否 | 输出是否被截断 |
|
||||
| `relevance_level` | VARCHAR(20) | 否 | 归一化质量等级:`PRECISE`、`HIGHLY_RELEVANT`、`REFERENCE`、`DEDUPED` |
|
||||
| `dedup_reason` | VARCHAR(32) | 否 | 去重原因,例如 `doc_retrieved`、`domain_retrieved` |
|
||||
| `retrieval_details` | JSON | 否 | 检索明细、证据引用、Gatekeeper 可用导航信息 |
|
||||
| `duration_ms` | INT | 否 | 工具耗时 |
|
||||
| `success` | BOOLEAN | 否 | 工具是否成功 |
|
||||
| `error_message` | TEXT | 否 | 失败原因 |
|
||||
| `created_at` | DATETIME | 是 | 创建时间 |
|
||||
|
||||
## 索引
|
||||
|
||||
| 索引 | 字段 | 用途 |
|
||||
|---|---|---|
|
||||
| `idx_session_id` | `session_id` | 历史兼容和粗粒度排查 |
|
||||
| `idx_tool_invocation_run_id` | `run_id, id` | Trace、Verifier、评测按运行查询工具调用 |
|
||||
| `idx_tool_name` | `tool_name` | 按工具类型排查 |
|
||||
| `idx_retrieval_layer` | `retrieval_layer` | 观察 RAG L0/L1 行为 |
|
||||
|
||||
## 关系
|
||||
|
||||
- `tool_invocation.run_id` 逻辑关联 `diagnosis_run.run_id`。
|
||||
- `tool_invocation.session_id` 保留为 `chat_session.session_id` 的冗余关联,便于粗粒度过滤和兼容查询。
|
||||
- `tool_invocation.step_id` 可关联 `agent_step.id`,但当前不强制。
|
||||
|
||||
## 关键 JSON
|
||||
|
||||
`retrieval_details` 是扩展字段。当前重要结构包括:
|
||||
|
||||
```json
|
||||
{
|
||||
"evidence_status": "supported",
|
||||
"evidence_refs": [
|
||||
{
|
||||
"raw_path": "$.logs[0]",
|
||||
"text": "工具返回中可核对的最小证据文本"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
## 注意点
|
||||
|
||||
- Verifier 不应只信任 RAG 证据;所有工具只要能提供 `evidence_refs`,都应该进入可校验证据链。
|
||||
- `output_preview` 只适合展示和排查,不应被当成完整原始输出。
|
||||
- `$.no_evidence` 只代表“本次工具未命中证据”,不能推导为“故障不存在”。
|
||||
@@ -0,0 +1,51 @@
|
||||
# 文档元数据表:api_document
|
||||
|
||||
**状态**:当前表
|
||||
**来源**:`V003__create_api_document.sql`、`V004__add_metadata_to_api_document.sql`、`ApiDocument`
|
||||
|
||||
## 定位
|
||||
|
||||
`api_document` 是知识库文档的 MySQL 元数据表。它不保存向量正文,正文切片和向量检索由 Milvus/Zilliz collection 承担;两侧通过 `doc_id` 和 chunk metadata 逻辑关联。
|
||||
|
||||
## 字段
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
|---|---|---|---|
|
||||
| `id` | BIGINT | 是 | 自增主键 |
|
||||
| `doc_id` | VARCHAR(64) | 是 | 文档唯一 ID,关联向量库 chunk metadata |
|
||||
| `fault_category` | VARCHAR(32) | 否 | 文档类别,默认 `EXTERNAL_API`;实体侧使用 `FaultCategory` |
|
||||
| `fault_source` | VARCHAR(128) | 否 | 文档归属,例如服务名、省份或系统来源 |
|
||||
| `api_name` | VARCHAR(128) | 否 | 接口或文档主题名称 |
|
||||
| `version` | VARCHAR(32) | 否 | 文档版本,默认 `v1.0` |
|
||||
| `file_name` | VARCHAR(256) | 是 | 原始文件名 |
|
||||
| `file_path` | VARCHAR(512) | 否 | 文件存储路径 |
|
||||
| `file_hash` | VARCHAR(64) | 否 | 文件 MD5,用于去重 |
|
||||
| `file_size` | BIGINT | 否 | 文件大小,单位字节 |
|
||||
| `status` | VARCHAR(16) | 否 | 索引状态:`PENDING`、`PROCESSING`、`INDEXED`、`FAILED` |
|
||||
| `chunk_count` | INT | 否 | 向量库切片数量 |
|
||||
| `error_message` | TEXT | 否 | 索引失败原因 |
|
||||
| `metadata` | TEXT | 否 | frontmatter 元数据 JSON 字符串 |
|
||||
| `indexed_at` | DATETIME | 否 | 索引完成时间 |
|
||||
| `created_at` | DATETIME | 是 | 创建时间 |
|
||||
| `updated_at` | DATETIME | 是 | 更新时间 |
|
||||
|
||||
## 索引
|
||||
|
||||
| 索引 | 字段 | 用途 |
|
||||
|---|---|---|
|
||||
| `uk_file_hash` | `file_hash` | 文件去重 |
|
||||
| `idx_doc_id` | `doc_id` | 按文档 ID 查询 |
|
||||
| `idx_fault_source` | `fault_source` | 按来源筛选 |
|
||||
| `idx_status` | `status` | 查看索引状态 |
|
||||
| `idx_created_at` | `created_at` | 按上传时间排序 |
|
||||
|
||||
## 关系
|
||||
|
||||
- `api_document.doc_id` 与向量库 chunk metadata 中的 `docId` / `doc_id` 逻辑关联。
|
||||
- `knowledge_domain.domain_id` 与文档 metadata 中的 `category` 形成领域聚合关系;当前没有数据库外键。
|
||||
|
||||
## 注意点
|
||||
|
||||
- 删除文档时需要同时处理 MySQL 元数据和向量库 chunk。
|
||||
- `metadata` 是 JSON 字符串,不是 MySQL JSON 列。
|
||||
- 表字段以 Flyway 为准;实体默认值和枚举可能与迁移脚本的 SQL 默认值存在历史差异,排查时优先看实际迁移和数据库结构。
|
||||
@@ -0,0 +1,51 @@
|
||||
# 案例库表:case_library
|
||||
|
||||
**状态**:当前表
|
||||
**来源**:`V002__create_case_library.sql`、`CaseLibrary`
|
||||
|
||||
## 定位
|
||||
|
||||
`case_library` 保存高质量诊断案例,用于后续相似案例推荐和知识沉淀。当前自动沉淀路径来自 `useful` 用户反馈:新数据把 `diagnosis_run` 中的 query 和 answer 映射为案例内容。
|
||||
|
||||
## 字段
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
|---|---|---|---|
|
||||
| `id` | BIGINT | 是 | 自增主键 |
|
||||
| `case_id` | VARCHAR(64) | 是 | 案例唯一 ID |
|
||||
| `diagnosis_id` | VARCHAR(64) | 否 | 关联诊断来源;新自动生成时存 `diagnosis_run.run_id`,历史数据可能是 `diagnosis_session.session_id` |
|
||||
| `source_type` | VARCHAR(16) | 否 | 来源类型:`AUTO` 或 `MANUAL` |
|
||||
| `fault_category` | VARCHAR(32) | 否 | 故障类别,实体侧使用 `FaultCategory` |
|
||||
| `fault_source` | VARCHAR(128) | 否 | 故障源,例如服务、系统或省份 |
|
||||
| `fault_target` | VARCHAR(256) | 否 | 故障目标,例如接口、方法、SQL 或组件 |
|
||||
| `error_code` | VARCHAR(64) | 否 | 错误码或异常类型 |
|
||||
| `title` | VARCHAR(256) | 是 | 案例标题 |
|
||||
| `root_cause` | TEXT | 是 | 根因分析 |
|
||||
| `solution` | TEXT | 是 | 解决方案 |
|
||||
| `reference_count` | INT | 否 | 被推荐次数 |
|
||||
| `created_by` | VARCHAR(64) | 否 | 创建人 |
|
||||
| `created_at` | DATETIME | 是 | 创建时间 |
|
||||
| `updated_at` | DATETIME | 是 | 更新时间 |
|
||||
|
||||
## 索引
|
||||
|
||||
| 索引 | 字段 | 用途 |
|
||||
|---|---|---|
|
||||
| `idx_fault_category` | `fault_category` | 按故障类别筛选 |
|
||||
| `idx_error_code` | `error_code` | 按错误码精确匹配 |
|
||||
| `idx_fault_source` | `fault_source` | 按故障源筛选 |
|
||||
| `idx_fault_target` | `fault_target(100)` | 按故障目标筛选 |
|
||||
| `idx_diagnosis_id` | `diagnosis_id` | 追溯来源运行或历史会话 |
|
||||
| `idx_reference_count` | `reference_count` | 推荐排序 |
|
||||
| `idx_created_at` | `created_at` | 时间排序 |
|
||||
|
||||
## 关系
|
||||
|
||||
- `case_library.diagnosis_id` 是过渡字段:新自动案例逻辑关联 `diagnosis_run.run_id`,历史自动案例可能仍是 `diagnosis_session.session_id`。
|
||||
- 人工录入案例可以不填写 `diagnosis_id`。
|
||||
|
||||
## 注意点
|
||||
|
||||
- 旧文档里提到的 `diagnosis_record` 已被 `V007` 删除,不再是当前主模型。
|
||||
- 查询新自动案例时优先按 `run_id` 追溯;遇到旧值时再按历史 `session_id` 解释。
|
||||
- 当前自动沉淀仍比较粗:`root_cause` 和 `solution` 都可能来自完整 answer。后续可从结构化结论中拆分根因、证据和修复建议。
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user