docs(openspec): archive session run isolation
This commit is contained in:
+31
-30
@@ -2,33 +2,34 @@
|
||||
|
||||
## 项目
|
||||
|
||||
| 日期 | slug | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|
||||
|---|---|---|---|---|---|
|
||||
| 2026-07-09 | interview-demo-quality-audit | Agent eval/demo/Prompt audit | interview demo preflight, prompt_audit, gatekeeper rules, diagnosis baseline, 12 fixtures | openspec/changes/archive/2026-07-09-interview-demo-quality-audit | archived |
|
||||
| 2026-07-08 | executor-composer-final-answer | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
|
||||
| 2026-07-08 | diagnosis-eval-demo-gatekeeper-closure | Agent eval/demo/Gatekeeper | diagnosis eval matrix, stable demo scenarios, Gatekeeper rule set version, audit metadata | openspec/changes/archive/2026-07-08-diagnosis-eval-demo-gatekeeper-closure | archived |
|
||||
| 2026-07-08 | verifier-evidence-reference-fidelity | Chat质量门禁/证据归因 | evidence_refs, raw_path, Gatekeeper severity, verifier evidence excerpt, HikariCP mock, no_evidence | openspec/changes/archive/2026-07-08-verifier-evidence-reference-fidelity | archived |
|
||||
| 2026-07-07 | executor-evidence-output-contract | Chat质量门禁/证据归因 | Executor structured output, evidence bindings, Verifier structured claims, LOW_CONFID, hallucination | openspec/changes/archive/2026-07-07-executor-evidence-output-contract | archived |
|
||||
| 2026-07-07 | executor-v2-output-contract | Chat质量门禁/证据归因 | executor_evidence_v2, user_facing_answer removal, diagnosis_summary removal, structured renderer | openspec/changes/archive/2026-07-07-executor-v2-output-contract | archived |
|
||||
| 2026-07-07 | executor-gatekeeper-hook | Chat质量门禁/证据归因 | Gatekeeper, verifier payload, source_invocation_ids, tool_name match, self_evaluation | openspec/changes/archive/2026-07-07-executor-gatekeeper-hook | archived |
|
||||
| 2026-07-07 | executor-verifier-claim-checks | Chat质量门禁/证据归因 | Verifier claim_checks, facts_checked compatibility, effective verdict guardrail, malformed output downgrade | openspec/changes/archive/2026-07-07-executor-verifier-claim-checks | archived |
|
||||
| 2026-07-06 | rag-eval-pipeline-closure | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
|
||||
| 2026-07-06 | modular-rag-pipeline | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
|
||||
| 2026-07-05 | diagnosis-playbook-skills | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented |
|
||||
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
|
||||
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
|
||||
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
|
||||
| 2026-07-04 | diagnosis-eval-harness | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
|
||||
| 2026-07-04 | evidence-trace-hardening | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
|
||||
| 2026-07-04 | aiops-traceable-diagnosis-entry | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
|
||||
| 2026-07-04 | aiops-alert-scope-control | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
|
||||
| 2026-07-03 | mvp-demo-trace-acceptance | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
|
||||
| 2026-07-02 | chat-verifier-agent | Chat质量门禁/可追溯验证 | Verifier, groundedness_score, facts_checked, evidence_refs, tool_trace_summary, self_evaluation | openspec/changes/archive/2026-07-03-chat-verifier-agent | archived |
|
||||
| 2026-07-01 | executor-action-memory-relevance | 检索质量/行动记忆 | relevanceLevel, completenessHint, Min-Max归一化, RetrievedDocTracker域级记录, Executor检索约束, ISS-002 | openspec/changes/archive/2026-07-01-executor-action-memory-relevance | archived |
|
||||
| 2026-06-30 | session-dedup-knowledge-map | 去重/知识图谱 | RetrievedDocTracker, KnowledgeDomainService, knowledge_domain, covers, whenToRetrieve, Planner注入, ISS-001 | openspec/changes/archive/2026-06-30-session-dedup-knowledge-map | archived |
|
||||
| 2026-06-29 | confidence-feedback | 质量评估/反馈机制 | evidence_score, selfEvaluation, feedback, useful, not_useful, case_library, BAD_CASE, tool_invocation规则引擎, 反馈按钮, sessionId回传 | openspec/changes/confidence-feedback | archived |
|
||||
| 2026-06-26 | session-storage | 会话存储/可观测 | diagnosis_session, agent_step, tool_invocation, token追踪, 多Agent路由 | openspec/changes/session-storage | archived |
|
||||
| 2026-06-25 | doc-management-ui | 前端开发/文档管理 | 文档管理页面, CRUD, 状态监控, 纯静态页面, API集成 | archived |
|
||||
| 2026-06-24 | lookup-knowledge-integration | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | archived |
|
||||
| 2026-06-23 | phase1-infrastructure | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | archived |
|
||||
| 2026-05-29 | chatmodel-abstraction | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | archived |
|
||||
| 日期 | slug | 说明 | 领域 | 关键词 | 关联 OpenSpec | 状态 |
|
||||
|---|---|---|---|---|---|---|
|
||||
| 2026-07-10 | session-run-trace-isolation | 拆分会话态和运行态,引入 runId 隔离 Trace、Feedback、AIOps 和 demo 链路。 | Trace/session/run isolation | chat_session, diagnosis_run, runId, trace exact run, feedback fallback, AIOps SSE metadata, baseline drift | openspec/changes/archive/2026-07-10-session-run-trace-isolation | archived |
|
||||
| 2026-07-09 | interview-demo-quality-audit | 增加面试演示前置质量审计,覆盖 prompt、Gatekeeper 和评测基线。 | Agent eval/demo/Prompt audit | interview demo preflight, prompt_audit, gatekeeper rules, diagnosis baseline, 12 fixtures | openspec/changes/archive/2026-07-09-interview-demo-quality-audit | archived |
|
||||
| 2026-07-08 | executor-composer-final-answer | 引入 Composer 生成最终回答,只使用 Verifier 允许的结论材料。 | Chat quality gate/evidence attribution | chat_composer, final answer, allowed_claims, allowed_hypotheses, safe fallback, composer_output | openspec/changes/archive/2026-07-08-executor-composer-final-answer | archived |
|
||||
| 2026-07-08 | diagnosis-eval-demo-gatekeeper-closure | 收敛诊断评测、稳定 demo 场景和 Gatekeeper 审计元数据。 | Agent eval/demo/Gatekeeper | diagnosis eval matrix, stable demo scenarios, Gatekeeper rule set version, audit metadata | openspec/changes/archive/2026-07-08-diagnosis-eval-demo-gatekeeper-closure | archived |
|
||||
| 2026-07-08 | verifier-evidence-reference-fidelity | 强化 Verifier 对 evidence_refs、raw_path 和 no_evidence 的保真校验。 | Chat质量门禁/证据归因 | evidence_refs, raw_path, Gatekeeper severity, verifier evidence excerpt, HikariCP mock, no_evidence | openspec/changes/archive/2026-07-08-verifier-evidence-reference-fidelity | archived |
|
||||
| 2026-07-07 | executor-evidence-output-contract | 设计 Executor 结构化证据输出,解决证据归因幻觉和 LOW_CONFID 问题。 | Chat质量门禁/证据归因 | Executor structured output, evidence bindings, Verifier structured claims, LOW_CONFID, hallucination | openspec/changes/archive/2026-07-07-executor-evidence-output-contract | archived |
|
||||
| 2026-07-07 | executor-v2-output-contract | 将 Executor 输出升级为 V2 契约,移除面向用户的最终回答字段。 | Chat质量门禁/证据归因 | executor_evidence_v2, user_facing_answer removal, diagnosis_summary removal, structured renderer | openspec/changes/archive/2026-07-07-executor-v2-output-contract | archived |
|
||||
| 2026-07-07 | executor-gatekeeper-hook | 在 Executor 与 Verifier 之间接入 Gatekeeper,校验证据绑定来源。 | Chat质量门禁/证据归因 | Gatekeeper, verifier payload, source_invocation_ids, tool_name match, self_evaluation | openspec/changes/archive/2026-07-07-executor-gatekeeper-hook | archived |
|
||||
| 2026-07-07 | executor-verifier-claim-checks | 增加 Verifier claim_checks 和事实校验兼容逻辑。 | Chat质量门禁/证据归因 | Verifier claim_checks, facts_checked compatibility, effective verdict guardrail, malformed output downgrade | openspec/changes/archive/2026-07-07-executor-verifier-claim-checks | archived |
|
||||
| 2026-07-06 | rag-eval-pipeline-closure | 建立 RAG 评测闭环,加入 fixture、快照和 baseline diff。 | RAG/评测/回归闭环 | lookupResult fixture, LookupKnowledgeTool snapshot, evidenceBlocks, contextPack, retrievalTrace, rerankTrace, baseline diff, fallback case | devflow/projects/2026-07-06-rag-eval-pipeline-closure | archived |
|
||||
| 2026-07-06 | modular-rag-pipeline | 将 lookup_knowledge 改造成模块化 RAG 管线,补齐证据块和检索追踪。 | RAG/Agent工具/证据链 | modular RAG, lookup_knowledge, evidenceBlocks, contextPack, rerank, retrievalTrace, L0 hint, unfiltered retry | openspec/changes/archive/2026-07-06-modular-rag-pipeline | archived |
|
||||
| 2026-07-05 | diagnosis-playbook-skills | 增加诊断 Playbook Skill,沉淀支付超时、MySQL 池、Redis 超时等套路。 | Agent Skill/Playbook | read_skill, diagnosis playbook, progressive disclosure, payment timeout, MySQL pool, Redis timeout | openspec/changes/diagnosis-playbook-skills | implemented |
|
||||
| 2026-07-05 | mvp-demo-interview-runbook | 准备可复现的 MVP 面试演示包、运行手册和 Trace 检查清单。 | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
|
||||
| 2026-07-05 | diagnosis-eval-baseline-diff | 增加诊断评测 baseline diff,用于判断回归和证据覆盖变化。 | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
|
||||
| 2026-07-04 | expand-diagnosis-eval-fixtures | 扩充诊断评测 fixture,覆盖 Redis、慢响应和 JVM 内存风险。 | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
|
||||
| 2026-07-04 | diagnosis-eval-harness | 建立固定诊断评测 Harness,输出 trace、证据覆盖和 verdict 分布。 | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
|
||||
| 2026-07-04 | evidence-trace-hardening | 强化工具调用证据链、降级契约和离线验证能力。 | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
|
||||
| 2026-07-04 | aiops-traceable-diagnosis-entry | 增加可追踪的 AIOps 告警诊断入口,打通 sessionId 和 Trace API。 | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
|
||||
| 2026-07-04 | aiops-alert-scope-control | 收敛 AIOps 告警诊断范围,区分 payload 定向和自动发现模式。 | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
|
||||
| 2026-07-03 | mvp-demo-trace-acceptance | 增加 MVP demo 的 Trace 验收,覆盖会话、步骤、工具和反馈链路。 | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
|
||||
| 2026-07-02 | chat-verifier-agent | 增加 Chat Verifier Agent,用 groundedness 和 evidence_refs 校验回答。 | Chat质量门禁/可追溯验证 | Verifier, groundedness_score, facts_checked, evidence_refs, tool_trace_summary, self_evaluation | openspec/changes/archive/2026-07-03-chat-verifier-agent | archived |
|
||||
| 2026-07-01 | executor-action-memory-relevance | 增加行动记忆和相关性信号,约束 Executor 重复检索。 | 检索质量/行动记忆 | relevanceLevel, completenessHint, Min-Max归一化, RetrievedDocTracker域级记录, Executor检索约束, ISS-002 | openspec/changes/archive/2026-07-01-executor-action-memory-relevance | archived |
|
||||
| 2026-06-30 | session-dedup-knowledge-map | 引入会话级去重和知识域地图,减少重复召回。 | 去重/知识图谱 | RetrievedDocTracker, KnowledgeDomainService, knowledge_domain, covers, whenToRetrieve, Planner注入, ISS-001 | openspec/changes/archive/2026-06-30-session-dedup-knowledge-map | archived |
|
||||
| 2026-06-29 | confidence-feedback | 建立质量评估和用户反馈机制,并把有用反馈沉淀为案例。 | 质量评估/反馈机制 | evidence_score, selfEvaluation, feedback, useful, not_useful, case_library, BAD_CASE, tool_invocation规则引擎, 反馈按钮, sessionId回传 | openspec/changes/confidence-feedback | archived |
|
||||
| 2026-06-26 | session-storage | 建立通用会话存储,记录 session、agent step 和 tool invocation。 | 会话存储/可观测 | diagnosis_session, agent_step, tool_invocation, token追踪, 多Agent路由 | openspec/changes/session-storage | archived |
|
||||
| 2026-06-25 | doc-management-ui | 实现文档管理页面,支持文档 CRUD、状态监控和 API 集成。 | 前端开发/文档管理 | 文档管理页面, CRUD, 状态监控, 纯静态页面, API集成 | - | archived |
|
||||
| 2026-06-24 | lookup-knowledge-integration | 接入知识库检索,支持 L0 精确匹配和 L1 语义检索。 | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | - | archived |
|
||||
| 2026-06-23 | phase1-infrastructure | 搭建第一阶段基础设施,包括 MySQL、Redis、Milvus、Flyway 和 JPA。 | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | - | archived |
|
||||
| 2026-05-29 | chatmodel-abstraction | 抽象 ChatModel 和 EmbeddingModel,支持多模型路由。 | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | - | archived |
|
||||
|
||||
@@ -0,0 +1,103 @@
|
||||
# Acceptance
|
||||
|
||||
## 实现结果
|
||||
|
||||
- OpenSpec tasks: `42/42` complete。
|
||||
- Phase commits:
|
||||
- `52bf030 feat(trace): add session run isolation schema`
|
||||
- `6fdbd34 docs(openspec): tighten run isolation contract`
|
||||
- `26d5529 feat(trace): isolate chat runs`
|
||||
- `027aed1 feat(trace): add run-scoped trace reads`
|
||||
- `d928a19 feat(trace): bind feedback to runs`
|
||||
- `78c1477 feat(trace): isolate aiops runs`
|
||||
- `f9df943 feat(trace): finish run-aware demo verification`
|
||||
- OpenSpec archive: `openspec/changes/archive/2026-07-10-session-run-trace-isolation`
|
||||
|
||||
## 静态验证
|
||||
|
||||
```powershell
|
||||
node --check src\main\resources\static\app.js
|
||||
node --check src\main\resources\static\trace.js
|
||||
openspec validate session-run-trace-isolation --strict
|
||||
git diff --check -- . ':!devflow/index.md'
|
||||
```
|
||||
|
||||
结果:通过。
|
||||
|
||||
## 脚本验证
|
||||
|
||||
PowerShell demo 脚本解析:
|
||||
|
||||
```powershell
|
||||
$scripts = @(
|
||||
'mvp\demo\scripts\run-payment-timeout-demo.ps1',
|
||||
'mvp\demo\scripts\run-interview-demo-check.ps1'
|
||||
)
|
||||
foreach ($script in $scripts) {
|
||||
[scriptblock]::Create((Get-Content -Raw -Encoding UTF8 $script)) | Out-Null
|
||||
}
|
||||
```
|
||||
|
||||
结果:通过。
|
||||
|
||||
Focused tests:
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=ChatControllerTest,DiagnosisTraceServiceTest,FeedbackControllerTest,FeedbackServiceTest,AiOpsServiceTest" test
|
||||
```
|
||||
|
||||
结果:通过。
|
||||
|
||||
Baseline / regression:
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest" test
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
|
||||
```
|
||||
|
||||
结果:通过,无 baseline drift。
|
||||
|
||||
## E2E 验证
|
||||
|
||||
使用 Maven 启动:
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
|
||||
```
|
||||
|
||||
E2E 使用同一 `sessionId` 连续两轮 Chat:
|
||||
|
||||
- `sessionId`: `e2e-phase6-chat-codex-20260710-2120`
|
||||
- `run1`: `run-e2a97696-4398-4abc-90e4-28f45c838f92`
|
||||
- `run2`: `run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172`
|
||||
|
||||
验证结果:
|
||||
|
||||
- Chat1 / Chat2 均成功。
|
||||
- run1 exact trace 返回 run1。
|
||||
- run2 exact trace 返回 run2。
|
||||
- session-only latest trace 返回 run2。
|
||||
- DB 中同一 session 有两条 `diagnosis_run`。
|
||||
- step/tool rows 按 `run_id` 隔离,mixed row check 为 0。
|
||||
- `chat_session.message_pair_count = 2`。
|
||||
|
||||
## 日志验证
|
||||
|
||||
检查:
|
||||
|
||||
- `target/e2e/phase6-mvn-20260710-211831.out.log`
|
||||
- `logs/application.log`
|
||||
- `logs/chat.log`
|
||||
|
||||
结果:能找到 E2E `sessionId`、两个 `runId`、Chat execution、run persistence 和 trace lookup 相关日志。
|
||||
|
||||
## 浏览器/人工验证
|
||||
|
||||
未单独进行浏览器点击验证。Trace UI 的本次验收通过静态语法检查、URL/runId 参数代码审查和后端 exact trace E2E 共同覆盖。建议后续手动打开 `trace.html?sessionId=...&runId=...` 做展示层冒烟。
|
||||
|
||||
## 剩余风险 / 后续事项
|
||||
|
||||
- 缺少 `runId` 的 Feedback fallback 是短期兼容路径,客户端全部迁移后可收紧。
|
||||
- `diagnosis_session` 仍保留为历史兼容和回滚表,后续需要观察窗口后再评估约束收紧或归档策略。
|
||||
- `case_library.diagnosis_id` 仍是过渡字段,旧值可能为 `session_id`,新自动值为 `run_id`。
|
||||
- 历史 mixed trace 不能恢复真实多轮边界,只能按 compatibility run 查询。
|
||||
@@ -0,0 +1,41 @@
|
||||
# Session / Run / Trace Isolation
|
||||
|
||||
## 背景
|
||||
|
||||
同一个 `sessionId` 以前同时代表多轮 Chat 上下文和一次持久化诊断 Trace。端到端验证发现,同一 `sessionId` 连续两轮 Chat 时,Redis 多轮上下文是正确的,但 MySQL 中 `diagnosis_session` 会被后一轮覆盖,`agent_step` 和 `tool_invocation` 会按同一个 `session_id` 混在一起。
|
||||
|
||||
这会导致 Trace 回放、Verifier/Evaluation 读数、Feedback 绑定和 `case_library` 来源都可能跨轮污染。
|
||||
|
||||
## 目标
|
||||
|
||||
- 将会话态和运行态拆开:`chat_session` 保存会话元数据,`diagnosis_run` 保存一次诊断运行。
|
||||
- 引入正式 API 字段 `runId`,作为一次可回放诊断执行的边界。
|
||||
- `agent_step` 和 `tool_invocation` 保留原 Trace 明细角色,新增 `run_id` 并按 run 隔离读写。
|
||||
- Trace、Feedback、CaseLibrary、AIOps、demo 脚本和 Trace UI 都支持 run-aware 流程。
|
||||
- 保留旧 `diagnosis_session` 作为历史兼容和回滚表。
|
||||
- 完成 Maven E2E、DB 检查、日志检查和 baseline drift 验证。
|
||||
|
||||
## 范围
|
||||
|
||||
- Flyway/JPA 增加 `chat_session`、`diagnosis_run`,并给 `agent_step`、`tool_invocation` 增加 `run_id`。
|
||||
- Chat 每次有效执行创建一个新的 `diagnosis_run`,响应返回 `sessionId + runId`。
|
||||
- Trace API 支持 latest-run fallback 和 exact-run 查询:`GET /api/diagnosis/{sessionId}/trace?runId=...`。
|
||||
- 新增 run list API:`GET /api/chat/session/{sessionId}/runs`。
|
||||
- Feedback 优先绑定 `runId`,缺省时短期 fallback 到 latest run 并返回 `fallbackToLatestRun=true`。
|
||||
- AIOps 每次有效执行创建并透出 `runId`,SSE 保持 `message` event name 并发送 `type=metadata`。
|
||||
- MVP demo、Trace UI、表文档和架构文档统一为 `chat_session -> diagnosis_run -> trace detail(run_id)`。
|
||||
|
||||
## 非目标
|
||||
|
||||
- 不新增 `diagnosis_trace` 或 `trace_event` 主表。
|
||||
- 不实现完整 run-list UI。
|
||||
- 不删除旧 `diagnosis_session`。
|
||||
- 不改变 Redis 对话历史窗口策略。
|
||||
- 不把完整多轮正文历史持久化到 MySQL。
|
||||
- 不尝试把历史混合 trace 还原成真实多轮边界。
|
||||
|
||||
## 关联
|
||||
|
||||
- OpenSpec: `openspec/changes/archive/2026-07-10-session-run-trace-isolation`
|
||||
- Change slug: `session-run-trace-isolation`
|
||||
- 分档: complex
|
||||
@@ -0,0 +1,48 @@
|
||||
# Decisions
|
||||
|
||||
## 核心决策
|
||||
|
||||
| 决策 | 选择 | 理由 |
|
||||
|---|---|---|
|
||||
| 领域拆分 | 新增 `chat_session` 和 `diagnosis_run` | 会话元数据和一次诊断执行的生命周期不同,继续塞在一张表会导致上下文膨胀和边界混淆 |
|
||||
| Trace 明细 | 复用 `agent_step` / `tool_invocation`,增加 `run_id` | 现有明细表已经能表达 Trace,隔离需要 run key,不需要新事件模型 |
|
||||
| API 身份 | `runId = "run-" + UUID` | 外部 ID 不依赖数据库自增 ID,碰撞风险低 |
|
||||
| Trace 兼容 | 缺少 `runId` 时按 `created_at DESC, id DESC` 解析 latest run | 保留旧客户端兼容性,避免 feedback/eval 更新 `updated_at` 后改变 latest 判定 |
|
||||
| 历史迁移 | 每条旧 `diagnosis_session` 生成一条 compatibility run | 旧混合数据没有真实轮次边界,不能伪造多 run 历史 |
|
||||
| Feedback fallback | 缺少 `runId` 时短期绑定 latest run 并返回 `fallbackToLatestRun=true` | 老客户端可继续工作,同时让歧义可观测 |
|
||||
| Case provenance | 新自动案例写 `case_library.diagnosis_id = run_id` | 保留旧列,文档声明过渡语义 |
|
||||
| AIOps 范围 | 同一个 change 内完成 AIOps run isolation | AIOps 是一等 Trace 入口,不能留下同类混合 trace bug |
|
||||
| 所有权校验 | 服务层校验 run/session ownership,暂不加 DB 外键 | 兼容历史 orphan rows 和回滚窗口 |
|
||||
|
||||
## 用户确认
|
||||
|
||||
- 选择拆 `chat_session` 和 `diagnosis_run`,不只是在旧表加字段。
|
||||
- `chat_session` 第一阶段只保存元数据,不保存完整对话正文。
|
||||
- 完整多轮对话历史继续放在 Redis `SessionContext.messageHistory`。
|
||||
- `runId` 是正式 API 字段。
|
||||
- Trace 缺少 `runId` 时短期默认查 latest run。
|
||||
- Feedback 缺少 `runId` 时短期 fallback,长期可再收紧。
|
||||
- 每次有效 Chat/AIOps 都创建 run。
|
||||
- 不新增 `diagnosis_trace` / `trace_event` 主表。
|
||||
- 旧 `diagnosis_session` 保留用于历史和回滚,新代码不再写新执行态。
|
||||
- demo 脚本和 Trace UI 做最小 `runId` 支持。
|
||||
|
||||
## 接口影响
|
||||
|
||||
级别:L4。
|
||||
|
||||
- 新 API 响应字段:`runId`。
|
||||
- Trace API 新 query 参数:`runId`。
|
||||
- 新 API:`GET /api/chat/session/{sessionId}/runs`。
|
||||
- Feedback request 新增 optional/preferred `runId`。
|
||||
- Feedback response 新增 bound `runId` 和 `fallbackToLatestRun`。
|
||||
- `/api/ai_ops` SSE 保持 event name `message`,新增 `type=metadata` 消息。
|
||||
- DB contract 新增两张表和两个 `run_id` 列。
|
||||
- 旧 `sessionId` only 调用仍兼容,但 fallback 必须可观测。
|
||||
|
||||
## 风险接受
|
||||
|
||||
- 历史混合 trace 无法真实拆分,只能作为 compatibility run。
|
||||
- 上下文传播同时依赖 `RunnableConfig.metadata` 和 `SessionContextHolder`,后续改动必须注意 `sessionId/runId` 同步。
|
||||
- `case_library.diagnosis_id` 在过渡期存在 `session_id` 和 `run_id` 两种语义。
|
||||
- 缺少 `runId` 的 Feedback 仍有歧义,后续客户端迁移完成后可收紧为参数错误。
|
||||
@@ -0,0 +1,76 @@
|
||||
# Evidence
|
||||
|
||||
## 上下文证据
|
||||
|
||||
- `SessionContext.messageHistory` 和 `getMessagePairCount()` 证明 Redis 承载热对话历史;MySQL 只需要长期审计的会话目录和运行记录。
|
||||
- `CaseLibraryService.createFromSession` 原先按 `DiagnosisSession.sessionId` 去重并映射 query/answer,因此 run 隔离后需要新增 `createFromRun`。
|
||||
- 旧 `mvp/architecture/data-model.md` 把 `case_library.diagnosis_id` 解释为 `diagnosis_session.session_id`,本次改为过渡语义:旧数据可能是 `session_id`,新自动案例是 `run_id`。
|
||||
- 既有 Trace OpenSpec 要求 `GET /api/diagnosis/{sessionId}/trace` 是只读端点;latest-run 和 exact-run 查询都必须保持只读。
|
||||
- ISS-010 的 E2E 事实显示同一 `sessionId` 两轮 Chat 会产生 MySQL Trace 混合,是本 change 的直接触发证据。
|
||||
|
||||
## 实现证据
|
||||
|
||||
- Phase 1 增加 `V011__add_session_run_isolation.sql`,创建 `chat_session`、`diagnosis_run`,并为 `agent_step` / `tool_invocation` 增加 nullable `run_id`。
|
||||
- Phase 2 将 Chat 写路径切到 `chat_session + diagnosis_run`,并让 Hook/Tool/Evaluation/Gatekeeper 使用 run-scoped 数据。
|
||||
- Phase 3 将 Trace API 改为 latest-run / exact-run 双模式,并加入 lightweight run summaries。
|
||||
- Phase 4 将 Feedback 和 CaseLibrary 绑定到 run,保留没有 run-backed 数据时的 legacy fallback。
|
||||
- Phase 5 将 AIOps 接入 run isolation,SSE metadata 暴露 `sessionId + runId`。
|
||||
- Phase 6 更新 demo 脚本、Trace UI、MVP 架构文档和表文档,并修正 review 后发现的 session-only 文档残留。
|
||||
|
||||
## E2E 证据
|
||||
|
||||
Maven 启动命令:
|
||||
|
||||
```powershell
|
||||
mvn spring-boot:run -Dspring-boot.run.profiles=mvp-demo
|
||||
```
|
||||
|
||||
日志:
|
||||
|
||||
- `target/e2e/phase6-mvn-20260710-211831.out.log`
|
||||
- `target/e2e/phase6-mvn-20260710-211831.err.log`
|
||||
- `logs/application.log`
|
||||
- `logs/chat.log`
|
||||
|
||||
E2E session:
|
||||
|
||||
- `sessionId`: `e2e-phase6-chat-codex-20260710-2120`
|
||||
- `run1`: `run-e2a97696-4398-4abc-90e4-28f45c838f92`
|
||||
- `run2`: `run-76ce6a6e-92ab-40c9-800a-eca0c1bb5172`
|
||||
|
||||
结果:
|
||||
|
||||
- 两轮 Chat 都成功,并复用同一个 `sessionId`。
|
||||
- 两轮返回不同 `runId`。
|
||||
- run1 exact trace 只返回 run1。
|
||||
- run2 exact trace 只返回 run2。
|
||||
- session-only Trace latest fallback 返回 run2。
|
||||
- `chat_session.message_pair_count = 2`,证明多轮上下文连续。
|
||||
|
||||
## DB 证据
|
||||
|
||||
通过 `scripts/query_mysql.py` 检查:
|
||||
|
||||
- `diagnosis_run` 中该 E2E session 有 2 条 `SUCCESS / CHAT` 运行。
|
||||
- `agent_step` 按 run 分组:run1 `10` 行,run2 `9` 行。
|
||||
- `tool_invocation` 按 run 分组:run1 `14` 行,run2 `8` 行。
|
||||
- mixed row check 为 `0`,没有 NULL 或 unexpected `run_id` 混入该 E2E session。
|
||||
|
||||
## Baseline 证据
|
||||
|
||||
运行:
|
||||
|
||||
```powershell
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest" test
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
|
||||
```
|
||||
|
||||
结果:
|
||||
|
||||
- 两组 baseline / regression 命令通过。
|
||||
- baseline harness 使用离线 fixture,不依赖 live DB/session tables。
|
||||
- 未观察到 baseline drift。
|
||||
|
||||
## 工具限制
|
||||
|
||||
AGENTS 要求的 `codebase-retrieval` 和 LSP 工具在本会话不可用。替代验证使用 OpenSpec、`rg`、定向阅读、 focused tests、E2E、DB 查询和日志检查。
|
||||
@@ -0,0 +1 @@
|
||||
archive-ready
|
||||
@@ -1,24 +1,33 @@
|
||||
## Purpose
|
||||
|
||||
Provide a repeatable MVP demo flow that can run a chat diagnosis, expose its persisted execution trace, and submit feedback for the same session id.
|
||||
Provide a repeatable MVP demo flow that can run a chat diagnosis, expose its persisted execution trace, and submit feedback for the same diagnosis run.
|
||||
## Requirements
|
||||
### Requirement: Diagnosis trace can be queried by session id
|
||||
The system SHALL expose a read-only HTTP endpoint `GET /api/diagnosis/{sessionId}/trace` that returns the persisted diagnosis trace for the requested session id.
|
||||
The system SHALL expose a read-only HTTP endpoint `GET /api/diagnosis/{sessionId}/trace` that returns the persisted diagnosis trace for the requested session id. When `runId` is omitted, the endpoint SHALL return the latest diagnosis run for compatibility. When `runId` is provided, the endpoint SHALL return that exact run after validating it belongs to the path `sessionId`.
|
||||
|
||||
#### Scenario: Existing session trace is returned
|
||||
- **WHEN** a caller requests trace data for a session id that exists in `diagnosis_session`
|
||||
- **THEN** the system returns a success response containing the session summary, ordered agent steps, ordered tool invocations, self-evaluation data, final answer, and feedback
|
||||
#### Scenario: Existing session latest trace is returned
|
||||
- **WHEN** a caller requests trace data for a session id that has at least one `diagnosis_run`
|
||||
- **THEN** the system returns a success response containing the resolved run id, session summary, run summary, ordered agent steps, ordered tool invocations, self-evaluation data, final answer, and feedback for the latest run
|
||||
|
||||
#### Scenario: Existing session exact trace is returned
|
||||
- **WHEN** a caller requests trace data with `GET /api/diagnosis/{sessionId}/trace?runId=run-xxx`
|
||||
- **THEN** the system validates that `runId` belongs to `sessionId`
|
||||
- **AND** it returns a success response containing only the trace data for that run
|
||||
|
||||
#### Scenario: Missing session returns not found
|
||||
- **WHEN** a caller requests trace data for a session id that does not exist in `diagnosis_session`
|
||||
- **WHEN** a caller requests trace data for a session id that does not exist in `chat_session`, `diagnosis_run`, or historical compatibility data
|
||||
- **THEN** the system returns a 404 response using the existing session-not-found error contract
|
||||
|
||||
### Requirement: Trace aggregation is read-only
|
||||
The system MUST build trace output from existing persisted diagnosis tables and MUST NOT mutate diagnosis sessions, agent steps, tool invocations, feedback, or chat session state while serving the trace request.
|
||||
The system MUST build trace output from existing persisted diagnosis tables and MUST NOT mutate chat sessions, diagnosis runs, agent steps, tool invocations, feedback, or chat session state while serving the trace request.
|
||||
|
||||
#### Scenario: Trace query does not change persisted state
|
||||
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace`
|
||||
- **THEN** the system reads `diagnosis_session`, `agent_step`, and `tool_invocation` records and returns an aggregate without saving any of those records
|
||||
- **THEN** the system reads `diagnosis_run`, `agent_step`, and `tool_invocation` records and returns an aggregate without saving any of those records
|
||||
|
||||
#### Scenario: Exact trace query does not change persisted state
|
||||
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace?runId=run-xxx`
|
||||
- **THEN** the system reads the specified run and its trace detail records without saving any of those records
|
||||
|
||||
### Requirement: MVP demo profile is available
|
||||
The system SHALL provide an `mvp-demo` Spring profile that documents the demo runtime intent and keeps mock log and metric providers enabled for repeatable diagnosis demonstrations.
|
||||
@@ -28,11 +37,11 @@ The system SHALL provide an `mvp-demo` Spring profile that documents the demo ru
|
||||
- **THEN** `prometheus.mock-enabled` and `cls.mock-enabled` are enabled by profile configuration
|
||||
|
||||
### Requirement: End-to-end MVP acceptance case is documented
|
||||
The project SHALL include an end-to-end acceptance case that demonstrates start-up, chat diagnosis, trace query, and feedback submission using the same session id.
|
||||
The project SHALL include an end-to-end acceptance case that demonstrates start-up, chat diagnosis, trace query, and feedback submission using the same `sessionId + runId`.
|
||||
|
||||
#### Scenario: Reviewer follows the acceptance case
|
||||
- **WHEN** a reviewer follows the documented MVP demo acceptance steps
|
||||
- **THEN** they can run the application, submit a diagnosis question, query the trace endpoint, and submit feedback for the same session id
|
||||
- **THEN** they can run the application, submit a diagnosis question, query the exact trace endpoint, and submit feedback for the same run
|
||||
|
||||
### Requirement: MVP demo SHALL provide an interview runbook
|
||||
The MVP demo SHALL include a concise interview runbook that explains how to demonstrate the Agent flow and how to narrate the engineering value.
|
||||
@@ -51,10 +60,11 @@ The MVP demo SHALL provide scripts and request payloads for running the payment-
|
||||
#### Scenario: Demo script sends the fixed diagnosis request
|
||||
- **WHEN** the demo script is executed against a running local service
|
||||
- **THEN** it SHALL send the fixed payment-timeout chat request with a stable session id
|
||||
- **AND** it SHALL read the returned run id for exact trace and feedback calls
|
||||
|
||||
#### Scenario: Demo script captures review artifacts
|
||||
- **WHEN** the demo script finishes successfully
|
||||
- **THEN** it SHALL write chat, trace, and feedback responses under a demo output directory
|
||||
- **THEN** it SHALL write chat, exact trace, and feedback responses under a demo output directory
|
||||
|
||||
### Requirement: MVP demo SHALL be reproducible for interviews
|
||||
The MVP demo SHALL provide a repeatable way to show a diagnosis answer, trace, verifier evaluation, and feedback.
|
||||
@@ -62,8 +72,8 @@ The MVP demo SHALL provide a repeatable way to show a diagnosis answer, trace, v
|
||||
#### Scenario: interview demo check script records an evidence bundle
|
||||
- **WHEN** the user runs the interview demo check script against a running `mvp-demo` service
|
||||
- **THEN** the script SHALL submit a fixed Chat diagnosis request
|
||||
- **AND** it SHALL fetch the trace for the same session id
|
||||
- **AND** it SHALL submit useful feedback for that session
|
||||
- **AND** it SHALL fetch the trace for the same `sessionId + runId`
|
||||
- **AND** it SHALL submit useful feedback for that run
|
||||
- **AND** it SHALL write chat, trace, feedback, and summary outputs under `mvp/demo/output/`
|
||||
|
||||
#### Scenario: interview demo check fails with actionable readiness output
|
||||
@@ -77,7 +87,7 @@ The MVP demo SHALL provide a repeatable way to show a diagnosis answer, trace, v
|
||||
- **AND** it SHALL explain that deterministic eval fixtures are the regression source of truth
|
||||
|
||||
### Requirement: MVP demo SHALL provide a trace inspection checklist
|
||||
The MVP demo SHALL document which trace fields to inspect for evidence, verifier behavior, and session-level auditability.
|
||||
The MVP demo SHALL document which trace fields to inspect for evidence, verifier behavior, and run-level auditability.
|
||||
|
||||
#### Scenario: Checklist maps fields to interview claims
|
||||
- **WHEN** a developer reviews a trace response
|
||||
@@ -85,7 +95,7 @@ The MVP demo SHALL document which trace fields to inspect for evidence, verifier
|
||||
|
||||
### Requirement: MVP demo SHALL provide a browser trace workbench
|
||||
The MVP demo SHALL provide a browser-accessible static page for inspecting one
|
||||
diagnosis trace by session id using the existing read-only Trace API.
|
||||
diagnosis trace by session id and optional run id using the existing read-only Trace API.
|
||||
|
||||
#### Scenario: Existing trace renders in the workbench
|
||||
- **WHEN** a reviewer opens the Trace workbench with a session id that exists
|
||||
@@ -93,6 +103,11 @@ diagnosis trace by session id using the existing read-only Trace API.
|
||||
- **AND** it SHALL render session summary, ordered agent steps, ordered tool
|
||||
invocations, verifier evaluation, and final answer when present
|
||||
|
||||
#### Scenario: Exact trace renders in the workbench
|
||||
- **WHEN** a reviewer opens the Trace workbench with `?sessionId=...&runId=...`
|
||||
- **THEN** the page SHALL request `GET /api/diagnosis/{sessionId}/trace?runId=...`
|
||||
- **AND** it SHALL render only that run's trace data
|
||||
|
||||
#### Scenario: Trace workbench handles missing or failed traces
|
||||
- **WHEN** the Trace API returns an error or the session id is empty
|
||||
- **THEN** the page SHALL show a clear error or empty state without mutating any
|
||||
|
||||
@@ -0,0 +1,181 @@
|
||||
# session-run-trace-isolation Specification
|
||||
|
||||
## Purpose
|
||||
|
||||
Separate multi-turn conversation metadata from per-execution diagnosis state. `sessionId` identifies the conversation context, while `runId` identifies one replayable diagnosis execution and scopes Trace, Feedback, Evaluation, AIOps, and case-library provenance.
|
||||
|
||||
## Requirements
|
||||
|
||||
### Requirement: Conversation metadata SHALL be separated from diagnosis runs
|
||||
The system SHALL persist multi-turn conversation metadata in `chat_session` and one execution's auditable state in `diagnosis_run`.
|
||||
|
||||
#### Scenario: Valid Chat execution creates session metadata and a run
|
||||
- **WHEN** a valid `/api/chat` request enters the Chat execution path and the service resolves an effective `sessionId`
|
||||
- **THEN** the system SHALL ensure a `chat_session` row exists for the effective `sessionId`
|
||||
- **AND** it SHALL create a new `diagnosis_run` row with a unique `run_id`
|
||||
- **AND** the `diagnosis_run.session_id` SHALL equal the effective `sessionId`
|
||||
|
||||
#### Scenario: Invalid Chat request does not create a run
|
||||
- **WHEN** a `/api/chat` request fails parameter validation before execution
|
||||
- **THEN** the system SHALL NOT create a `diagnosis_run`
|
||||
|
||||
#### Scenario: Chat session stores metadata only
|
||||
- **WHEN** a Chat request completes
|
||||
- **THEN** `chat_session` SHALL store metadata such as status, message pair count, created time, last active time, and optional expiration time
|
||||
- **AND** it SHALL NOT store full conversation message history
|
||||
|
||||
### Requirement: Chat responses SHALL expose run identity
|
||||
The system SHALL expose the current execution `runId` to clients that submit Chat requests.
|
||||
|
||||
#### Scenario: Chat response includes runId
|
||||
- **WHEN** `/api/chat` returns a successful response
|
||||
- **THEN** the response SHALL include `sessionId`
|
||||
- **AND** the response SHALL include `runId` for the created diagnosis run
|
||||
|
||||
#### Scenario: Multi-turn Chat keeps one session and multiple runs
|
||||
- **WHEN** two valid `/api/chat` requests use the same `sessionId`
|
||||
- **THEN** the system SHALL preserve multi-turn Redis context for that `sessionId`
|
||||
- **AND** it SHALL persist two distinct `diagnosis_run.run_id` values
|
||||
|
||||
### Requirement: Trace details SHALL be scoped by run
|
||||
The system SHALL write and read `agent_step` and `tool_invocation` rows using `run_id` as the execution boundary.
|
||||
|
||||
#### Scenario: Agent steps are recorded with runId
|
||||
- **WHEN** an Agent model step is persisted during a diagnosis run
|
||||
- **THEN** the `agent_step` row SHALL include the current `run_id`
|
||||
- **AND** it SHALL retain the current `session_id`
|
||||
|
||||
#### Scenario: Tool invocations are recorded with runId
|
||||
- **WHEN** an evidence tool invocation is persisted during a diagnosis run
|
||||
- **THEN** the `tool_invocation` row SHALL include the current `run_id`
|
||||
- **AND** it SHALL retain the current `session_id`
|
||||
|
||||
#### Scenario: Run metrics count only current run rows
|
||||
- **WHEN** a diagnosis run completes
|
||||
- **THEN** its step and tool counts SHALL be calculated from rows matching that `run_id`
|
||||
- **AND** rows from other runs in the same `sessionId` SHALL NOT be counted
|
||||
|
||||
### Requirement: Trace API SHALL support latest-run and exact-run queries
|
||||
The system SHALL allow callers to query a diagnosis trace by `sessionId` alone for compatibility or by `sessionId + runId` for exact run replay.
|
||||
|
||||
#### Scenario: Trace without runId resolves latest run
|
||||
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace` without `runId`
|
||||
- **THEN** the system SHALL resolve the latest run for that session by `diagnosis_run.created_at DESC, id DESC`
|
||||
- **AND** the response SHALL include the resolved `runId`
|
||||
|
||||
#### Scenario: Trace with runId returns exact run
|
||||
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace?runId=run-xxx`
|
||||
- **THEN** the system SHALL validate that `runId` belongs to the path `sessionId`
|
||||
- **AND** it SHALL return only the session summary, run summary, agent steps, tool invocations, self-evaluation, answer, and feedback for that run
|
||||
- **AND** the session summary SHALL come from `chat_session` metadata when available, while the run summary SHALL come from `diagnosis_run`
|
||||
|
||||
#### Scenario: Trace rejects run from another session
|
||||
- **WHEN** a caller requests a `runId` that belongs to a different `sessionId`
|
||||
- **THEN** the system SHALL return an error instead of leaking trace data from the other session
|
||||
|
||||
### Requirement: Session runs SHALL be listable without expanding trace details
|
||||
The system SHALL provide a lightweight run-list API for a Chat Session.
|
||||
|
||||
#### Scenario: Run list returns summaries
|
||||
- **WHEN** a caller requests `GET /api/chat/session/{sessionId}/runs`
|
||||
- **THEN** the system SHALL return run summaries from `diagnosis_run` inside the existing API response wrapper
|
||||
- **AND** each summary SHALL include `runId`, `sessionId`, `query`, `status`, `agentFlow`, `answerPreview`, `stepCount`, `toolCallCount`, `createdAt`, and `updatedAt`
|
||||
- **AND** the response SHALL NOT expand `agent_step` or `tool_invocation` detail rows
|
||||
|
||||
#### Scenario: Run list handles session without runs
|
||||
- **WHEN** a caller requests `GET /api/chat/session/{sessionId}/runs` for an existing `chat_session` with no runs
|
||||
- **THEN** the system SHALL return a successful empty list
|
||||
|
||||
#### Scenario: Run list rejects missing session
|
||||
- **WHEN** a caller requests `GET /api/chat/session/{sessionId}/runs` for a session that does not exist in `chat_session` or `diagnosis_run`
|
||||
- **THEN** the system SHALL use the existing not-found/error response behavior
|
||||
|
||||
### Requirement: Feedback SHALL bind to diagnosis runs
|
||||
The system SHALL bind new feedback to a diagnosis run rather than an ambiguous multi-turn session.
|
||||
|
||||
#### Scenario: Feedback with runId updates specified run
|
||||
- **WHEN** a feedback request includes `sessionId` and `runId`
|
||||
- **THEN** the system SHALL validate that the run belongs to the session
|
||||
- **AND** it SHALL update feedback on that run
|
||||
- **AND** the response SHALL include the actual bound `runId`
|
||||
- **AND** the response SHALL include `fallbackToLatestRun=false`
|
||||
|
||||
#### Scenario: Feedback without runId falls back observably
|
||||
- **WHEN** a legacy feedback request includes `sessionId` but omits `runId`
|
||||
- **AND** at least one `diagnosis_run` exists for that session
|
||||
- **THEN** the system SHALL bind feedback to the latest run for that session
|
||||
- **AND** the response SHALL include `fallbackToLatestRun=true`
|
||||
- **AND** the response SHALL include the actual bound `runId`
|
||||
|
||||
#### Scenario: Historical feedback without run-backed data remains compatible
|
||||
- **WHEN** a legacy feedback request includes `sessionId` but omits `runId`
|
||||
- **AND** no `diagnosis_run` exists for that session
|
||||
- **AND** a historical `diagnosis_session` row exists for that session
|
||||
- **THEN** the system MAY bind feedback to the historical session row for migration compatibility
|
||||
- **AND** the response SHALL NOT claim latest-run fallback
|
||||
- **AND** the response MAY omit `runId`
|
||||
|
||||
#### Scenario: Feedback rejects run from another session
|
||||
- **WHEN** a feedback request includes a `runId` that belongs to a different `sessionId`
|
||||
- **THEN** the system SHALL return a failed feedback response instead of updating either run
|
||||
|
||||
#### Scenario: Useful feedback creates case from run
|
||||
- **WHEN** feedback for a run is `useful`
|
||||
- **THEN** the system SHALL create or reuse a `case_library` row using that run's query and answer
|
||||
- **AND** new automatic case data SHALL store `case_library.diagnosis_id` as the `run_id`
|
||||
|
||||
### Requirement: AIOps executions SHALL use run isolation
|
||||
The system SHALL create and expose a diagnosis run for every valid `/api/ai_ops` execution.
|
||||
|
||||
#### Scenario: AIOps creates run
|
||||
- **WHEN** `/api/ai_ops` starts a valid execution
|
||||
- **THEN** the system SHALL create a `diagnosis_run` with `agent_flow=AI_OPS`
|
||||
- **AND** AIOps agent steps, tool invocations, and rule evaluation SHALL be associated with that `run_id`
|
||||
- **AND** AIOps rule evaluation SHALL be stored under `diagnosis_run.self_evaluation.aiops_rule_evaluation` for the current run
|
||||
|
||||
#### Scenario: AIOps SSE exposes runId
|
||||
- **WHEN** `/api/ai_ops` streams response metadata to the caller
|
||||
- **THEN** the stream SHALL send a compatible metadata message before report content
|
||||
- **AND** the SSE event name SHALL remain `message`
|
||||
- **AND** the message type SHALL be `metadata`
|
||||
- **AND** the metadata payload SHALL expose the resolved `sessionId`
|
||||
- **AND** the metadata payload SHALL expose the created `runId`
|
||||
- **AND** report content SHALL continue to use the existing content message shape
|
||||
|
||||
### Requirement: Migration SHALL preserve historical trace access
|
||||
The system SHALL migrate historical diagnosis data into compatibility runs without deleting the old `diagnosis_session` table.
|
||||
|
||||
#### Scenario: Historical session gets compatibility run
|
||||
- **WHEN** migration runs on an existing `diagnosis_session` row
|
||||
- **THEN** the system SHALL create a compatible `diagnosis_run` row for that session
|
||||
- **AND** old `agent_step` and `tool_invocation` rows for that session SHALL be backfilled to that `run_id` when possible
|
||||
|
||||
#### Scenario: Old table is retained
|
||||
- **WHEN** the migration completes
|
||||
- **THEN** the `diagnosis_session` table SHALL remain available for historical comparison and rollback
|
||||
- **AND** new execution writes SHALL target `chat_session` and `diagnosis_run`
|
||||
|
||||
### Requirement: Demo and Trace UI SHALL support runId
|
||||
The demo tooling and Trace UI SHALL support minimal run-aware workflows.
|
||||
|
||||
#### Scenario: Demo script queries exact trace
|
||||
- **WHEN** a demo script receives a `/api/chat` or `/api/ai_ops` response containing `runId`
|
||||
- **THEN** it SHALL include `runId` when querying the Trace API
|
||||
|
||||
#### Scenario: Trace UI honors URL runId
|
||||
- **WHEN** the Trace UI is opened with `?sessionId=...&runId=...`
|
||||
- **THEN** it SHALL query `GET /api/diagnosis/{sessionId}/trace?runId=...`
|
||||
|
||||
### Requirement: Run isolation SHALL be verified against baselines
|
||||
The change SHALL verify both runtime behavior and evaluation baseline impact.
|
||||
|
||||
#### Scenario: Multi-turn E2E proves run isolation
|
||||
- **WHEN** E2E verification sends two valid Chat requests with the same `sessionId`
|
||||
- **THEN** database inspection SHALL show two `diagnosis_run` rows
|
||||
- **AND** each run SHALL have only its own step and tool rows when queried by `run_id`
|
||||
- **AND** Redis session metadata SHALL still show multi-turn context continuity
|
||||
|
||||
#### Scenario: Baseline drift is checked
|
||||
- **WHEN** verification is complete
|
||||
- **THEN** the project SHALL run or explicitly evaluate the relevant baseline diff command
|
||||
- **AND** any drift caused by run isolation SHALL be documented as expected or investigated as a regression
|
||||
Reference in New Issue
Block a user