feat(graph): complete stategraph cleanup and acceptance
This commit is contained in:
@@ -1,6 +1,6 @@
|
||||
# MVP 架构文档
|
||||
|
||||
**更新日期**:2026-07-10
|
||||
**更新日期**:2026-07-17
|
||||
|
||||
这里是 MVP 当前架构的唯一入口。旧版设计、早期拆解和已经被新实现替代的方案已归档到:
|
||||
|
||||
@@ -14,9 +14,9 @@
|
||||
|---|---|
|
||||
| [current-mvp-architecture.md](current-mvp-architecture.md) | 当前可运行 MVP 的总体架构、链路、持久化和质量门禁 |
|
||||
| [interview-one-pager.md](interview-one-pager.md) | 面试一页式架构讲解,包含总图、亮点、取舍和追问回答 |
|
||||
| [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat SequentialAgent、AIOps SupervisorAgent、工具边界 |
|
||||
| [agent-orchestration.md](agent-orchestration.md) | Agent 编排细节,覆盖 Chat bounded StateGraph、AIOps SupervisorAgent、工具边界 |
|
||||
| [executor-evidence-pipeline-refactor.md](executor-evidence-pipeline-refactor.md) | Chat 证据链路当前数据契约,覆盖 Executor V2、Gatekeeper、Verifier、Composer、`evidence_refs` |
|
||||
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、Hook、Trace、Gatekeeper、Verifier、Composer、评测基线组成的质量门禁 |
|
||||
| [harness-quality-gates.md](harness-quality-gates.md) | Prompt、StateGraph、Trace、Gatekeeper、Verifier、Composer、评测基线组成的质量门禁 |
|
||||
| [rag-architecture.md](rag-architecture.md) | RAG/知识检索新架构,覆盖 L0 hint、VectorStore 主路径、SDK fallback、证据追踪 |
|
||||
| [modular-rag-pipeline.md](modular-rag-pipeline.md) | `lookup_knowledge` 模块化 RAG 落地架构,覆盖 pipeline、fallback、evidence-first contract、trace |
|
||||
| [rag-eval-closure.md](rag-eval-closure.md) | RAG 评测闭环,覆盖 offline baseline、baseline diff、diagnosis eval 和 live acceptance |
|
||||
@@ -29,7 +29,7 @@
|
||||
|
||||
## 当前架构一句话
|
||||
|
||||
SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:Chat 和 AIOps 入口统一进入 Agent 编排,Executor 通过显式工具收集日志、指标和知识库证据,Chat 链路由 Gatekeeper 做引用真实性校验、Verifier 做可推导性判断、Composer 生成最终表达;多轮会话元数据落到 `chat_session`,每次诊断运行落到 `diagnosis_run`,步骤和工具明细通过 `agent_step.run_id`、`tool_invocation.run_id` 关联,最终通过 Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
|
||||
SuperBizAgent MVP 是一个面向故障诊断的可追踪 Agent 系统:复杂 Chat 由有界 StateGraph 显式编排 Planner、Executor、Gatekeeper、Verified Input、Verifier、Composer 与安全 Fallback,Executor 通过工具收集日志、指标和知识库证据;多轮会话元数据落到 `chat_session`,每次诊断运行落到 `diagnosis_run`,步骤和工具明细通过 `run_id` 关联,Graph 路由摘要独立保存为 `orchestration_trace`,最终通过精确 Run Trace API 和评测脚本证明诊断链路可解释、可回放、可对比。
|
||||
|
||||
## 阅读顺序
|
||||
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# Agent 编排架构
|
||||
|
||||
**更新日期**:2026-07-08
|
||||
**更新日期**:2026-07-17
|
||||
**状态**:当前可运行架构
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
||||
|
||||
@@ -8,7 +8,7 @@
|
||||
|
||||
旧版 Agent 架构把系统描述为 Supervisor、Planner、SubAgent、Verifier 的团队协作。当前 MVP 保留这个核心思想,但实现更收敛:
|
||||
|
||||
- Chat 链路使用固定顺序工作流:`Planner -> Executor -> Gatekeeper -> Verifier -> Composer`。
|
||||
- Chat 复杂诊断使用有递归上限的显式 StateGraph;正常路径是 `Planner -> Executor -> Gatekeeper -> Verified Input -> Verifier -> Composer`,条件边负责有限技术重试、一次补证据和安全 Fallback。
|
||||
- AIOps 链路使用 `SupervisorAgent` 调度 `Planner + Executor`,最终由规则评估器做轻量验证。
|
||||
- 当前没有拆分 ExternalApiSubAgent、InternalErrorSubAgent、DatabaseSubAgent;这些作为后续演进方向保留。
|
||||
- 证据工具不直接散落在各个 Agent 里,而是通过 Spring AI ToolCallback / `@Tool` 统一暴露。
|
||||
@@ -19,15 +19,19 @@
|
||||
flowchart TB
|
||||
subgraph Chat["Chat diagnosis"]
|
||||
ChatIn["POST /api/chat"] --> ChatService["ChatService"]
|
||||
ChatService --> ChatPlanner["chat_planner"]
|
||||
ChatService --> ChatGraph["ChatDiagnosisGraphRuntime / StateGraph"]
|
||||
ChatGraph --> ChatPlanner["Planner Node"]
|
||||
ChatPlanner --> ChatExecutor["chat_executor"]
|
||||
ChatExecutor --> ChatTools["evidence tools"]
|
||||
ChatTools --> ChatExecutor
|
||||
ChatExecutor --> ChatGatekeeper["ExecutorGatekeeperService"]
|
||||
ChatGatekeeper --> ChatVerifier["chat_verifier"]
|
||||
ChatExecutor --> ChatGatekeeper["Gatekeeper Node"]
|
||||
ChatGatekeeper --> VerifiedInput["Verified Input Node"]
|
||||
VerifiedInput --> ChatVerifier["Verifier Node"]
|
||||
ChatVerifier --> ChatDecision{"PASS / LOW_CONFID / REJECT"}
|
||||
ChatDecision --> ChatComposer["chat_composer"]
|
||||
ChatDecision --> ChatComposer["Composer Node"]
|
||||
ChatDecision --> ChatFallback["Fallback Node"]
|
||||
ChatComposer --> ChatAnswer["final answer"]
|
||||
ChatFallback --> ChatAnswer
|
||||
end
|
||||
|
||||
subgraph AiOps["AIOps diagnosis"]
|
||||
@@ -54,6 +58,7 @@ flowchart TB
|
||||
ChatPlanner --> Step
|
||||
ChatExecutor --> Step
|
||||
ChatGatekeeper --> SelfEval
|
||||
ChatGraph --> Run
|
||||
ChatVerifier --> Step
|
||||
ChatTools --> Invocation
|
||||
ChatDecision --> SelfEval
|
||||
@@ -69,19 +74,13 @@ flowchart TB
|
||||
|
||||
## 3. Chat 编排
|
||||
|
||||
Chat 复杂诊断采用 `SequentialAgent`,顺序固定:
|
||||
Chat 复杂诊断采用 `ChatDiagnosisGraphRuntime` 编译的 bounded StateGraph。它有一条正常路径和显式条件边,不再依赖固定顺序 Agent 或 Verifier Hook:
|
||||
|
||||
```text
|
||||
chat_planner
|
||||
-> chat_executor
|
||||
-> lookup_knowledge / query_logs / query_metrics / date_time
|
||||
-> outputs executor_evidence_v2
|
||||
-> VerifierInputHook / ExecutorGatekeeperService
|
||||
-> validates source_invocation_id / raw_path / evidence_excerpt
|
||||
-> chat_verifier
|
||||
-> judges whether verified evidence can derive claims
|
||||
-> chat_composer
|
||||
-> writes final user-facing answer
|
||||
START -> PLANNER -> EXECUTOR -> GATEKEEPER -> VERIFIED_INPUT -> VERIFIER -> COMPOSER -> END
|
||||
| | | | |
|
||||
+ retry + fallback + fallback + retry + retry/fallback
|
||||
+ EVIDENCE_RETRY -> PLANNER (最多一次)
|
||||
```
|
||||
|
||||
关键行为:
|
||||
@@ -90,16 +89,18 @@ chat_planner
|
||||
|---|---|---|
|
||||
| `chat_planner` | 拆解问题,注入知识域地图和对话历史,给出排查方向 | `planner_plan` |
|
||||
| `chat_executor` | 按计划调用证据工具,抽取带 `source_invocation_id + raw_path + evidence_excerpt` 的微观事实 | `executor_evidence_v2` |
|
||||
| `ExecutorGatekeeperService` | 在 Verifier 前做代码级引用验真,拒绝伪造 ID、错配 raw_path、错配 excerpt | `gatekeeper_result` |
|
||||
| `chat_verifier` | 只判断已验真 evidence excerpt 是否能推出 claim,不做新检索 | `verifier_output` |
|
||||
| `GatekeeperNode` / `ExecutorGatekeeperService` | 按当前 `runId` 做代码级引用验真,拒绝伪造 ID、错配 raw_path、错配 excerpt | `gatekeeper_result` |
|
||||
| `VerifiedInputNode` | 只投影 Gatekeeper 通过的 claims 与 matched evidence,隔离完整工具 Trace | `verified_executor_output`、`verified_evidence` |
|
||||
| `chat_verifier` | 只判断已验真的 evidence excerpt 是否能推出 claim,不做新检索、不读取完整工具 Trace | `verifier_output` |
|
||||
| `chat_composer` | 只表达 Verifier 允许输出的 claims、缺口和建议,生成最终用户答复 | `composer_output` |
|
||||
| `FallbackNode` | 在不可恢复失败或路由上限触发时生成非空安全答复 | `final_answer`、degraded trace |
|
||||
|
||||
Chat 链路最多支持两轮验证:
|
||||
Chat Graph 支持有限技术重试,并只允许一次 evidence retry;所有分支最终进入 Composer 或 Fallback:
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant C as ChatService
|
||||
participant C as ChatService / StateGraph
|
||||
participant P as chat_planner
|
||||
participant E as chat_executor
|
||||
participant T as tools
|
||||
@@ -114,18 +115,22 @@ sequenceDiagram
|
||||
E->>T: 调用证据工具
|
||||
T-->>E: 证据结果
|
||||
E-->>C: executor_evidence_v2
|
||||
C->>G: executor_structured_output + tool_invocation.evidence_refs
|
||||
C->>G: executor_output + run-owned tool_invocation.evidence_refs
|
||||
G-->>C: gatekeeper_result
|
||||
C->>V: executor_structured_output + gatekeeper_result + tool_trace_summary
|
||||
C->>C: VerifiedInputNode projects passed claims/evidence
|
||||
C->>V: verified_executor_output + verified_evidence + gatekeeper_audit
|
||||
V-->>C: PASS / LOW_CONFID / REJECT
|
||||
C->>R: 写入 verifier_evaluation
|
||||
alt LOW_CONFID 且允许补证据
|
||||
alt LOW_CONFID 且允许一次补证据
|
||||
C->>P: retry_context: 仅补缺失证据
|
||||
else PASS 或 REJECT
|
||||
else PASS / LOW_CONFID 可输出
|
||||
C->>M: allowed_claims + missing_info + recommended_actions
|
||||
M-->>C: composer_output
|
||||
C->>R: 保存 Composer 最终 answer
|
||||
else 不可恢复失败
|
||||
C->>R: Fallback 安全答复
|
||||
end
|
||||
C->>R: 保存 orchestration_trace(version/transitions/final_node/termination_reason/degraded/evidence_retry_count)
|
||||
```
|
||||
|
||||
决策语义:
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# 当前 MVP 架构
|
||||
|
||||
**更新日期**:2026-07-08
|
||||
**更新日期**:2026-07-17
|
||||
**状态**:当前可运行架构
|
||||
**适用范围**:Demo、面试讲解、后续迭代规划
|
||||
|
||||
@@ -36,11 +36,14 @@ flowchart TB
|
||||
|
||||
subgraph Agent["Agent Orchestration"]
|
||||
Supervisor["Supervisor"]
|
||||
StateGraph["Chat Diagnosis StateGraph"]
|
||||
Planner["Planner"]
|
||||
Executor["Executor"]
|
||||
Gatekeeper["Gatekeeper"]
|
||||
VerifiedInput["Verified Input"]
|
||||
Verifier["Verifier"]
|
||||
Composer["Composer"]
|
||||
Fallback["Fallback"]
|
||||
end
|
||||
|
||||
subgraph Tools["Evidence Tools"]
|
||||
@@ -74,7 +77,14 @@ flowchart TB
|
||||
end
|
||||
|
||||
API --> App
|
||||
ChatService --> Agent
|
||||
ChatService --> StateGraph
|
||||
StateGraph --> Planner
|
||||
StateGraph --> Executor
|
||||
StateGraph --> Gatekeeper
|
||||
StateGraph --> VerifiedInput
|
||||
StateGraph --> Verifier
|
||||
StateGraph --> Composer
|
||||
StateGraph --> Fallback
|
||||
AiOpsService --> Agent
|
||||
SkillRegistry --> PlannerSkillHook
|
||||
PlannerSkillHook --> Planner
|
||||
@@ -374,11 +384,12 @@ Trace API 聚合:
|
||||
- Agent step 序列。
|
||||
- 工具调用和检索细节。
|
||||
- Chat Gatekeeper / Verifier / Composer 结果。
|
||||
- Chat `run.orchestrationTrace` 路由摘要,独立于 self-evaluation 和步骤/工具明细。
|
||||
- AIOps rule evaluation 结果。
|
||||
|
||||
Trace 是本项目区别于普通问答系统的关键:答案不是孤立文本,而是可以追溯到 Agent 决策、工具调用和证据来源。
|
||||
|
||||
Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明见 [harness-quality-gates.md](harness-quality-gates.md),用户反馈与 `self_evaluation` 闭环见 [feedback-architecture.md](feedback-architecture.md)。
|
||||
Prompt、StateGraph、Gatekeeper、Verifier、Composer 和评测门禁的完整说明见 [harness-quality-gates.md](harness-quality-gates.md),用户反馈与 `self_evaluation` 闭环见 [feedback-architecture.md](feedback-architecture.md)。
|
||||
|
||||
## 8. 质量门禁
|
||||
|
||||
@@ -386,9 +397,11 @@ Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明
|
||||
|
||||
| 门禁 | 位置 | 作用 |
|
||||
|---|---|---|
|
||||
| Executor Gatekeeper | `VerifierInputHook` / `ExecutorGatekeeperService` | 校验 Executor 引用的 invocation、`raw_path`、`evidence_excerpt` 是否真实 |
|
||||
| Chat Verifier | `ChatService` | 判断已验真证据是否能推出 Executor claims |
|
||||
| Chat Composer | `ChatService` | 只表达 Verifier 允许输出的内容,避免把 no-evidence 说成已排除 |
|
||||
| Executor Gatekeeper | `GatekeeperNode` / `ExecutorGatekeeperService` | 校验当前 Run 的 invocation、`raw_path`、`evidence_excerpt` 是否真实 |
|
||||
| Verified Input | `VerifiedInputNode` | 仅投影 Gatekeeper 通过的 claims/evidence,阻断完整工具 Trace 进入 Verifier |
|
||||
| Chat Verifier | `VerifierNodeAdapter` | 判断已验真证据是否能推出 Executor claims |
|
||||
| Chat Composer / Fallback | `ComposerNodeAdapter` / `FallbackNode` | 输出受控答复;异常分支也必须安全终止 |
|
||||
| Graph routing | `DiagnosisGraphWorkflowTest` / `run.orchestrationTrace` | 验证条件边、有限重试、最终节点和终止原因 |
|
||||
| AIOps Rule Evaluation | `AiOpsRuleEvaluationService` | 校验告警诊断是否聚焦 payload 并使用证据 |
|
||||
| Diagnosis Eval Baseline | `mvp/eval/` | 固化诊断 trace 和报告行为 |
|
||||
| RAG Retrieval Baseline | `eval/rag-retrieval/` | 固化检索召回行为,避免 RAG 重构回退 |
|
||||
@@ -399,6 +412,7 @@ Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明
|
||||
已经完成:
|
||||
|
||||
- Chat 和 AIOps 两条入口链路。
|
||||
- Chat 复杂诊断已单轨切换到 bounded StateGraph,并持久化 Run-owned `orchestration_trace`。
|
||||
- 显式 `lookup_knowledge` Agent Tool。
|
||||
- L0 从最终决策降级为 domain/entity hint。
|
||||
- `VectorSearchService` 作为稳定检索门面。
|
||||
@@ -430,6 +444,7 @@ Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明
|
||||
| 能力 | 代码 |
|
||||
|---|---|
|
||||
| Chat 入口与编排 | `ChatController`, `ChatService` |
|
||||
| Chat StateGraph | `ChatDiagnosisGraphRuntime`, `DiagnosisGraphFactory`, `DiagnosisRealGraphActionsFactory` |
|
||||
| AIOps 入口与编排 | `ChatController.aiOps`, `AiOpsService` |
|
||||
| AIOps 规则验证 | `AiOpsRuleEvaluationService` |
|
||||
| 知识库工具 | `LookupKnowledgeTool` |
|
||||
@@ -440,5 +455,5 @@ Prompt、Hook、Gatekeeper、Verifier、Composer 和评测门禁的完整说明
|
||||
| Spring AI VectorStore 配置辅助 | `SpringAiVectorStoreSidecarService` |
|
||||
| Trace 聚合 | `DiagnosisTraceService` |
|
||||
| 工具调用记录 | `ToolInvocationRecorder` |
|
||||
| Executor 引用验真 | `ExecutorGatekeeperService`, `VerifierInputHook` |
|
||||
| Executor 引用验真与投影 | `GatekeeperNode`, `ExecutorGatekeeperService`, `VerifiedInputNode` |
|
||||
| self_evaluation 合并 | `SelfEvaluationMergeService` |
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
# Chat Evidence Pipeline Contracts
|
||||
|
||||
**状态**:当前实现
|
||||
**更新日期**:2026-07-08
|
||||
**更新日期**:2026-07-17
|
||||
**范围**:Chat 复杂诊断链路中的 Planner、Executor、Gatekeeper、Verifier、Composer 数据契约
|
||||
|
||||
当前 Chat 复杂诊断链路是:
|
||||
@@ -9,7 +9,8 @@
|
||||
```text
|
||||
chat_planner
|
||||
-> chat_executor
|
||||
-> VerifierInputHook / ExecutorGatekeeperService
|
||||
-> GatekeeperNode / ExecutorGatekeeperService
|
||||
-> VerifiedInputNode
|
||||
-> chat_verifier
|
||||
-> chat_composer
|
||||
-> final answer
|
||||
@@ -219,13 +220,13 @@ Executor 必须遵守:
|
||||
|
||||
## 4. Gatekeeper
|
||||
|
||||
Gatekeeper 位于 Verifier 前,由 `VerifierInputHook` 触发,负责代码级引用真实性校验。
|
||||
Gatekeeper 是 StateGraph 中的显式 Node,调用 `ExecutorGatekeeperService` 对当前 Run 的证据引用做代码级真实性校验。
|
||||
|
||||
### 4.1 输入
|
||||
|
||||
- `sessionId`
|
||||
- `sessionId + runId`(来自 `RunnableConfig`,工具查询以 `runId` 为边界)
|
||||
- `executor_structured_output`
|
||||
- 当前 session 的 `tool_invocation`
|
||||
- 当前 run 的 `tool_invocation`
|
||||
|
||||
### 4.2 输出
|
||||
|
||||
@@ -304,22 +305,20 @@ Gatekeeper 位于 Verifier 前,由 `VerifierInputHook` 触发,负责代码
|
||||
|
||||
## 5. Verifier
|
||||
|
||||
Verifier 输入由 `VerifierInputHook` 构造:
|
||||
`VerifiedInputNode` 只保留 Gatekeeper 检查通过的 claim/binding,并为 Verifier 构造最小输入:
|
||||
|
||||
```json
|
||||
{
|
||||
"original_query": "用户原始问题",
|
||||
"executor_final_answer": "{...executor raw text for debug/fallback only...}",
|
||||
"executor_structured_output": {
|
||||
"diagnosis_context": {
|
||||
"query": "用户原始问题"
|
||||
},
|
||||
"verified_executor_output": {
|
||||
"answer_version": "executor_evidence_v2",
|
||||
"claims": []
|
||||
},
|
||||
"executor_output_parse_status": {
|
||||
"status": "valid",
|
||||
"detail": "parsed executor evidence contract"
|
||||
},
|
||||
"tool_trace_summary": [],
|
||||
"gatekeeper_result": {},
|
||||
"verified_evidence": [],
|
||||
"gatekeeper_audit": {},
|
||||
"verdict_ceiling": "PASS",
|
||||
"retry_context": null
|
||||
}
|
||||
```
|
||||
@@ -330,7 +329,7 @@ Verifier 职责:
|
||||
- 不读 skill。
|
||||
- 不逐字核验 excerpt 真伪;这由 Gatekeeper 完成。
|
||||
- 只判断 `claim_text` 是否能由已核验的 `evidence_excerpt` 推出。
|
||||
- 结构化输出有效时,不得从 `executor_final_answer` 抽取额外确认事实。
|
||||
- 不读取 Executor 原始答复或完整工具 Trace,只读取 verified projection。
|
||||
- 对 `gatekeeper_result.severity=reject` 不得输出 `PASS`。
|
||||
- 对 `gatekeeper_result.severity=low_confid` 不得输出 `PASS`。
|
||||
|
||||
@@ -410,7 +409,7 @@ Composer 位于 Verifier 之后,输入是 ChatService 过滤后的允许表达
|
||||
"rule_set_version": "gatekeeper-rules-v1"
|
||||
},
|
||||
"composer_output": {},
|
||||
"tool_trace_summary": []
|
||||
"verified_evidence": []
|
||||
}
|
||||
}
|
||||
```
|
||||
@@ -422,6 +421,9 @@ Trace API 可用于回放:
|
||||
- Gatekeeper 是否通过、是否自动回填。
|
||||
- Verifier 如何判断可推导性。
|
||||
- Composer 最终如何表达给用户。
|
||||
- `run.orchestrationTrace` 如何经过条件边、有限重试并终止。
|
||||
|
||||
历史 Run/fixture 的 `verifier_evaluation.tool_trace_summary` 仍可被 Trace UI 或离线评测只读解析,但它是旧链路兼容字段,不是当前 Verifier 输入,也不再由生产链路生成。
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# 反馈与自评估架构
|
||||
|
||||
**更新日期**:2026-07-10
|
||||
**更新日期**:2026-07-17
|
||||
**状态**:当前可运行架构
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/confidence-feedback.md`
|
||||
|
||||
@@ -27,9 +27,8 @@ flowchart TD
|
||||
Invocation["tool_invocation"] --> RuleEval["EvaluationService: rule_evaluation"]
|
||||
Invocation --> EvidenceRefs["evidence_refs"]
|
||||
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
Invocation --> TraceSummary["ToolTraceSummaryService"]
|
||||
Gatekeeper --> Verifier["chat_verifier"]
|
||||
TraceSummary --> Verifier
|
||||
Gatekeeper --> Projection["VerifiedInputNode"]
|
||||
Projection --> Verifier["chat_verifier"]
|
||||
Verifier --> VerifierEval["verifier_evaluation"]
|
||||
Verifier --> Composer["chat_composer"]
|
||||
Composer --> VerifierEval
|
||||
@@ -81,7 +80,7 @@ flowchart TD
|
||||
"executor_structured_output": {},
|
||||
"gatekeeper_result": {},
|
||||
"composer_output": {},
|
||||
"tool_trace_summary": []
|
||||
"verified_evidence": []
|
||||
},
|
||||
"aiops_rule_evaluation": {
|
||||
"verdict": "...",
|
||||
@@ -136,14 +135,13 @@ Chat 自评估分三步:
|
||||
```mermaid
|
||||
flowchart LR
|
||||
ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
Invocation["tool_invocation"] --> Summary["ToolTraceSummaryService"]
|
||||
Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
|
||||
Invocation["tool_invocation"] --> EvidenceRefs["retrieval_details.evidence_refs"]
|
||||
EvidenceRefs --> Gatekeeper
|
||||
Gatekeeper --> GateResult["gatekeeper_result"]
|
||||
Summary --> Evidence["tool_trace_summary"]
|
||||
GateResult --> Verifier["chat_verifier"]
|
||||
ExecutorOutput --> Verifier
|
||||
Evidence --> Verifier
|
||||
GateResult --> Projection["VerifiedInputNode"]
|
||||
ExecutorOutput --> Projection
|
||||
Projection --> VerifiedOutput["verified_executor_output + verified_evidence"]
|
||||
VerifiedOutput --> Verifier["chat_verifier"]
|
||||
Verifier --> Output["verifier_output JSON"]
|
||||
Output --> Composer["chat_composer"]
|
||||
Composer --> ComposerOutput["composer_output"]
|
||||
@@ -163,9 +161,9 @@ Verifier 输出:
|
||||
| `facts_checked` | 逐条事实校验 |
|
||||
| `rationale` | 判定原因 |
|
||||
| `executor_structured_output` | Executor 输出的结构化 claims 与证据绑定 |
|
||||
| `verified_evidence` | Gatekeeper 通过并投影给 Verifier 的最小 matched evidence |
|
||||
| `gatekeeper_result` | 引用真实性校验结果 |
|
||||
| `composer_output` | 最终表达的解析状态和摘要 |
|
||||
| `tool_trace_summary` | 本次校验使用的工具调用导航索引 |
|
||||
|
||||
ChatService 根据 verdict 决定:
|
||||
|
||||
@@ -177,6 +175,7 @@ ChatService 根据 verdict 决定:
|
||||
|
||||
- `executor_final_answer` 只作为 debug/fallback 上下文;结构化输出有效时,Verifier 不得从中抽取额外确认事实。
|
||||
- `$.no_evidence` 只能表达“当前查询未检索到匹配证据”,不能表达“已排除/确认没有”。
|
||||
- `run.orchestrationTrace` 是独立的 StateGraph 路由摘要,不属于 `self_evaluation`;历史 `tool_trace_summary` 仅用于旧 Run/fixture 只读兼容,不是当前 Verifier 输入。
|
||||
|
||||
## 6. AIOps 规则自评估
|
||||
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# Harness 与质量门禁架构
|
||||
|
||||
**更新日期**:2026-07-08
|
||||
**更新日期**:2026-07-17
|
||||
**状态**:当前可运行架构 + 后续门禁规划
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/agent-architecture.md`
|
||||
|
||||
@@ -20,6 +20,7 @@ Agent 系统的核心风险不是“没有答案”,而是:
|
||||
Prompt contract
|
||||
+ Tool boundary
|
||||
+ Agent hooks
|
||||
+ StateGraph routing contract
|
||||
+ Trace persistence
|
||||
+ Gatekeeper deterministic validation
|
||||
+ Verifier / rule evaluation
|
||||
@@ -31,7 +32,7 @@ Prompt contract
|
||||
```mermaid
|
||||
flowchart TB
|
||||
Input["User / AIOps input"] --> Prompt["Prompt contract"]
|
||||
Prompt --> Agent["Planner / Executor / Verifier / Composer"]
|
||||
Prompt --> Agent["Diagnosis StateGraph Nodes"]
|
||||
Agent --> Tools["Evidence tools"]
|
||||
Tools --> Invocation["tool_invocation"]
|
||||
Agent --> StepHook["AgentLoggingHook"]
|
||||
@@ -41,15 +42,16 @@ flowchart TB
|
||||
Invocation --> EvidenceRefs["retrieval_details.evidence_refs"]
|
||||
EvidenceRefs --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
Agent --> Gatekeeper
|
||||
Invocation --> TraceSummary["ToolTraceSummaryService"]
|
||||
Gatekeeper --> Verifier["chat_verifier"]
|
||||
TraceSummary --> Verifier
|
||||
Gatekeeper --> Projection["VerifiedInputNode"]
|
||||
Projection --> Verifier["chat_verifier"]
|
||||
Verifier --> SelfEval["self_evaluation.verifier_evaluation"]
|
||||
|
||||
Invocation --> AiOpsRule["AiOpsRuleEvaluationService"]
|
||||
AiOpsRule --> AiOpsEval["self_evaluation.aiops_rule_evaluation"]
|
||||
|
||||
Run --> TraceAPI["DiagnosisTraceService"]
|
||||
Agent --> Routing["diagnosis_run.orchestration_trace"]
|
||||
Routing --> TraceAPI
|
||||
Step --> TraceAPI
|
||||
Invocation --> TraceAPI
|
||||
SelfEval --> TraceAPI
|
||||
@@ -161,7 +163,7 @@ error_message
|
||||
|
||||
## 6. Gatekeeper 与 Verifier 门禁
|
||||
|
||||
Chat Verifier 前置一层 Gatekeeper。Gatekeeper 不调用 LLM,只用代码检查 Executor 输出的证据引用是否真实存在。
|
||||
Chat StateGraph 在 Verifier 前显式执行 Gatekeeper 和 Verified Input。Gatekeeper 不调用 LLM,只用代码检查 Executor 输出的证据引用是否真实存在;Verified Input 只投影通过的 binding。
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
@@ -169,11 +171,12 @@ flowchart LR
|
||||
ExecutorOutput["executor_evidence_v2"] --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
EvidenceRefs --> Gatekeeper
|
||||
Gatekeeper --> GateResult["gatekeeper_result"]
|
||||
Invocation --> Summary["ToolTraceSummaryService"]
|
||||
Summary --> EvidenceIndex["tool_trace_summary"]
|
||||
GateResult --> Projection["VerifiedInputNode"]
|
||||
Projection --> VerifiedClaims["verified_executor_output"]
|
||||
Projection --> VerifiedEvidence["verified_evidence"]
|
||||
GateResult --> Verifier["chat_verifier"]
|
||||
ExecutorOutput --> Verifier
|
||||
EvidenceIndex --> Verifier
|
||||
VerifiedClaims --> Verifier
|
||||
VerifiedEvidence --> Verifier
|
||||
Verifier --> Verdict{"verdict"}
|
||||
Verdict -->|PASS| Composer["chat_composer"]
|
||||
Composer --> Pass["输出最终答复"]
|
||||
@@ -215,7 +218,7 @@ Verifier 不再逐字核验 excerpt 真伪;这由 Gatekeeper 完成。Verifier
|
||||
diagnosis_run.self_evaluation.verifier_evaluation
|
||||
```
|
||||
|
||||
其中同时持久化 `executor_structured_output`、`gatekeeper_result`、`tool_trace_summary`、`prompt_audit` 和 `composer_output`,用于 Trace 回放。
|
||||
其中持久化 verified `executor_structured_output`、`verified_evidence`、`gatekeeper_result`、`prompt_audit` 和 `composer_output`,用于 Trace 回放。Graph 路由另存 `diagnosis_run.orchestration_trace`;历史 `tool_trace_summary` 只作为旧 Run/fixture 的读取兼容字段,不属于当前 Verifier 输入。
|
||||
|
||||
## 7. AIOps 规则门禁
|
||||
|
||||
|
||||
@@ -142,7 +142,7 @@ post-retrieval 层再把检索候选归一为:
|
||||
- 给 Agent 输出 completeness hint。
|
||||
- 写入 `tool_invocation.relevance_level`。
|
||||
- 给 Gatekeeper 提供 `evidence_refs` 引用验真源。
|
||||
- 给 Verifier 构造 `tool_trace_summary` 审计导航。
|
||||
- 由 Gatekeeper 核验后,经 `VerifiedInputNode` 给 Verifier 构造最小 `verified_evidence` 投影。
|
||||
- 供 EvaluationService 计算 evidence score。
|
||||
|
||||
## 6. 文档切片和 metadata
|
||||
@@ -180,11 +180,13 @@ flowchart LR
|
||||
Recorder --> Invocation["tool_invocation"]
|
||||
Invocation --> Trace["DiagnosisTraceService"]
|
||||
Invocation --> Gatekeeper["ExecutorGatekeeperService"]
|
||||
Invocation --> Summary["ToolTraceSummaryService"]
|
||||
Summary --> Verifier["chat_verifier"]
|
||||
Gatekeeper --> Projection["VerifiedInputNode"]
|
||||
Projection --> Verifier["chat_verifier"]
|
||||
Invocation --> Eval["EvaluationService / RAG eval"]
|
||||
```
|
||||
|
||||
旧 Trace/fixture 中的 `tool_trace_summary` 只保留读取兼容;当前 StateGraph 不再生成它,也不会把完整工具调用摘要输入 Verifier。
|
||||
|
||||
`tool_invocation` 中与检索相关的字段:
|
||||
|
||||
```text
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# 会话与 Trace 生命周期
|
||||
|
||||
**更新日期**:2026-07-10
|
||||
**更新日期**:2026-07-17
|
||||
**状态**:当前可运行架构
|
||||
**参考历史文档**:`archive/2026-07-05-legacy/session-management.md`
|
||||
|
||||
@@ -29,7 +29,7 @@ flowchart TD
|
||||
Session --> Run["create diagnosis_run(runId)"]
|
||||
Run --> Running["run.status = RUNNING"]
|
||||
|
||||
Running --> Agent["Agent workflow"]
|
||||
Running --> Agent["Chat StateGraph / AIOps workflow"]
|
||||
Agent --> Context["execution context(sessionId, runId)"]
|
||||
Context --> StepHook["AgentLoggingHook"]
|
||||
StepHook --> Step["agent_step(session_id, run_id)"]
|
||||
@@ -37,7 +37,8 @@ flowchart TD
|
||||
Tool --> Invocation["tool_invocation(session_id, run_id)"]
|
||||
Invocation --> Gatekeeper["Gatekeeper evidence validation"]
|
||||
|
||||
Agent --> Final{"workflow result"}
|
||||
Agent --> GraphTrace["Chat: save orchestration_trace"]
|
||||
GraphTrace --> Final{"workflow result"}
|
||||
Final -->|success| Success["run.status = SUCCESS, answer saved"]
|
||||
Final -->|failed| Failed["run.status = FAILED"]
|
||||
|
||||
@@ -81,6 +82,7 @@ stateDiagram-v2
|
||||
| `status` | `diagnosis_run` | 单次运行执行状态 |
|
||||
| `answer` | `diagnosis_run` | 本次运行最终报告或答复 |
|
||||
| `self_evaluation` | `diagnosis_run` | 本次运行系统自评估 JSON |
|
||||
| `orchestration_trace` | `diagnosis_run` | Chat StateGraph 路由摘要;包含 version、transitions、final node、termination reason、degraded 和 evidence retry count |
|
||||
| `feedback` | `diagnosis_run` | 本次运行用户反馈 |
|
||||
|
||||
`feedback` 不修改 `status`。一个执行成功但用户标记 `not_useful` 的 run,仍然应该是 `SUCCESS + feedback=not_useful`。
|
||||
@@ -115,7 +117,7 @@ ToolInvocationRecorder
|
||||
-> retrieval_details / evidence_refs
|
||||
```
|
||||
|
||||
Verifier、Gatekeeper 和 EvaluationService 应按 `run_id` 读取工具调用,避免同一 session 的其他 run 参与评分或证据校验。
|
||||
Gatekeeper 和 EvaluationService 按 `run_id` 读取工具调用,避免同一 session 的其他 run 参与评分或证据校验。Verifier 只读取 `VerifiedInputNode` 生成的 verified projection,不直接读取完整工具调用列表。
|
||||
|
||||
## 7. Trace API 聚合
|
||||
|
||||
@@ -128,6 +130,7 @@ GET /api/diagnosis/{sessionId}/trace?runId=run-...
|
||||
|
||||
```text
|
||||
diagnosis_run by sessionId + runId
|
||||
+ run.orchestrationTrace parsed from diagnosis_run.orchestration_trace
|
||||
+ chat_session metadata when available
|
||||
+ agent_step where run_id = runId, ordered by the Trace API
|
||||
+ tool_invocation where run_id = runId order by id
|
||||
@@ -136,16 +139,20 @@ diagnosis_run by sessionId + runId
|
||||
|
||||
当 `runId` 缺失时,Trace API 为兼容旧客户端解析最新 run,并在响应中返回 resolved `runId`。当 `runId` 属于其他 `sessionId` 时,API 必须拒绝,不能泄漏其他会话的 Trace。
|
||||
|
||||
`run.orchestrationTrace` 只属于精确 Run 投影,不复制到顶层或 `session`。它解释 Graph 路由;`selfEvaluation` 解释证据/答案质量;`steps` 和 `toolInvocations` 保存详细执行证据,三者职责互不替代。历史 Run 的该字段可以为空。
|
||||
|
||||
## 8. Chat 与 AIOps 差异
|
||||
|
||||
| 维度 | Chat | AIOps |
|
||||
|---|---|---|
|
||||
| `agent_flow` | `CHAT` | `AI_OPS` |
|
||||
| 编排方式 | `SequentialAgent`: Planner -> Executor -> Gatekeeper -> Verifier -> Composer | `SupervisorAgent`: Planner + Executor |
|
||||
| 编排方式 | bounded `StateGraph`: Planner / Executor / Gatekeeper / Verified Input / Verifier / Composer / Fallback | `SupervisorAgent`: Planner + Executor |
|
||||
| 自评估 | `rule_evaluation` + `verifier_evaluation` | `aiops_rule_evaluation` |
|
||||
| 答案字段 | Chat 最终答复 | 告警分析报告 |
|
||||
| runId 暴露 | `/api/chat` JSON response | `/api/ai_ops` SSE metadata message |
|
||||
|
||||
Chat StateGraph 的权威自动化验收分三层:`DiagnosisGraphWorkflowTest` 验证路由,`DiagnosisGraphNodeContractTest` 验证真实 Node 输入输出,`ChatServiceGraphIntegrationTest` 验证 Run 生命周期、Trace 持久化和对外集成。
|
||||
|
||||
## 9. 清理与边界
|
||||
|
||||
- Redis 会话历史用于多轮上下文,不是长期审计记录。
|
||||
|
||||
+9
-2
@@ -8,7 +8,7 @@
|
||||
- `interview-walkthrough.md`:面试讲解话术。
|
||||
- `evidence-pipeline-scenarios.md`:PASS / LOW_CONFID / REJECT / no-evidence 场景矩阵。
|
||||
- `trace-inspection-checklist.md`:Trace 字段检查清单。
|
||||
- `scripts/run-interview-demo-check.ps1`:面试预检脚本,包含服务可达性、Chat、Trace、反馈和 summary 输出。
|
||||
- `scripts/run-interview-demo-check.ps1`:面试预检脚本,绑定 exact runId,强制校验 Run orchestration trace,并输出 Chat、Trace、反馈和 summary。
|
||||
- `scripts/run-payment-timeout-demo.ps1`:本地可执行 Demo 脚本。
|
||||
- `interview-q-and-a.md`:面试追问回答,覆盖 Agent 工程取舍、审计和评测。
|
||||
- `requests/payment-timeout-chat.json`:固定 Chat 请求 payload。
|
||||
@@ -51,6 +51,8 @@ mvp/demo/output/feedback-response.json
|
||||
mvp/demo/output/interview-demo-summary.json
|
||||
```
|
||||
|
||||
自动化验收应传入唯一 `-SessionId`,并用 `-OutputDir target/...` 避免覆盖仓库样例。脚本从 Chat 响应取得 exact `runId`,缺少 `data.run.orchestrationTrace` 或 version/final node/termination reason/transitions/degraded/evidence retry count 时会立即失败。summary 额外包含 `orchestrationVersion`、`finalNode`、`terminationReason`、`degraded`、`transitionCount` 和 `evidenceRetryCount`。
|
||||
|
||||
手动请求:
|
||||
|
||||
```powershell
|
||||
@@ -100,6 +102,10 @@ Invoke-RestMethod `
|
||||
- `data.runId` 等于 `$runId`
|
||||
- `data.session.sessionId` 等于 Chat session id
|
||||
- `data.run.runId` 等于 `$runId`
|
||||
- `data.run.orchestrationTrace.version` 非空
|
||||
- `data.run.orchestrationTrace.final_node` 和 `termination_reason` 非空
|
||||
- `data.run.orchestrationTrace.transitions` 是本次 Graph 的条件边记录
|
||||
- `data.run.orchestrationTrace.degraded` 和 `evidence_retry_count` 记录安全降级与补证据次数
|
||||
- `data.steps` 包含 planner / executor / verifier 等步骤
|
||||
- `data.toolInvocations` 包含 `lookup_knowledge`、`query_logs`、`query_metrics` 等证据工具
|
||||
- `data.session.selfEvaluation` 包含 verifier 或 rule evaluation
|
||||
@@ -174,9 +180,10 @@ Chat 主线:
|
||||
```text
|
||||
一个 session id + 一个 run id
|
||||
-> 用户问题
|
||||
-> 多 Agent 执行
|
||||
-> bounded StateGraph(Planner / Executor / Gatekeeper / Verified Input / Verifier / Composer / Fallback)
|
||||
-> 证据工具
|
||||
-> Verifier / self_evaluation
|
||||
-> run.orchestrationTrace 路由摘要
|
||||
-> 最终答案
|
||||
-> 用户反馈
|
||||
-> Trace API 回放
|
||||
|
||||
@@ -79,6 +79,15 @@ $chatPath = Join-Path $OutputDir "chat-response.json"
|
||||
$chat | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $chatPath
|
||||
|
||||
$runId = $chat.data.runId
|
||||
if ($chat.data.success -ne $true) {
|
||||
throw "Chat response was not successful."
|
||||
}
|
||||
if ([string]::IsNullOrWhiteSpace([string]$chat.data.answer)) {
|
||||
throw "Chat response did not include a non-empty answer."
|
||||
}
|
||||
if ($chat.data.sessionId -ne $SessionId) {
|
||||
throw "Chat response sessionId '$($chat.data.sessionId)' did not match requested sessionId '$SessionId'."
|
||||
}
|
||||
if (-not $runId) {
|
||||
throw "Chat response did not include runId; exact trace verification cannot continue."
|
||||
}
|
||||
@@ -92,6 +101,51 @@ $trace = Invoke-RestMethod @traceRequest
|
||||
$tracePath = Join-Path $OutputDir "trace-response.json"
|
||||
$trace | ConvertTo-Json -Depth 80 | Set-Content -Encoding UTF8 -Path $tracePath
|
||||
|
||||
$traceData = Get-TraceData -TraceResponse $trace
|
||||
if ($null -eq $traceData -or $null -eq $traceData.run) {
|
||||
throw "Exact Trace response did not include data.run."
|
||||
}
|
||||
if ($traceData.runId -ne $runId -or $traceData.run.runId -ne $runId) {
|
||||
throw "Exact Trace runId did not match Chat runId '$runId'."
|
||||
}
|
||||
if ($traceData.run.sessionId -ne $SessionId) {
|
||||
throw "Exact Trace run did not belong to requested sessionId '$SessionId'."
|
||||
}
|
||||
|
||||
$orchestrationTrace = $traceData.run.orchestrationTrace
|
||||
if ($null -eq $orchestrationTrace) {
|
||||
throw "Exact Trace data.run.orchestrationTrace is missing."
|
||||
}
|
||||
foreach ($field in @("version", "final_node", "termination_reason")) {
|
||||
if (-not ($orchestrationTrace.PSObject.Properties.Name -contains $field) -or
|
||||
[string]::IsNullOrWhiteSpace([string]$orchestrationTrace.$field)) {
|
||||
throw "Exact Trace data.run.orchestrationTrace.$field is missing."
|
||||
}
|
||||
}
|
||||
foreach ($field in @("transitions", "degraded", "evidence_retry_count")) {
|
||||
if (-not ($orchestrationTrace.PSObject.Properties.Name -contains $field)) {
|
||||
throw "Exact Trace data.run.orchestrationTrace.$field is missing."
|
||||
}
|
||||
}
|
||||
if ($null -eq $orchestrationTrace.transitions) {
|
||||
throw "Exact Trace data.run.orchestrationTrace.transitions must be an array."
|
||||
}
|
||||
if ([int]$orchestrationTrace.evidence_retry_count -lt 0) {
|
||||
throw "Exact Trace data.run.orchestrationTrace.evidence_retry_count must not be negative."
|
||||
}
|
||||
if ($traceData.run.status -ne "SUCCESS" -or $traceData.run.agentFlow -ne "CHAT") {
|
||||
throw "Exact Trace run must be CHAT/SUCCESS."
|
||||
}
|
||||
if ([string]::IsNullOrWhiteSpace([string]$traceData.run.answer)) {
|
||||
throw "Exact Trace run did not include a non-empty answer."
|
||||
}
|
||||
if (@($traceData.steps).Count -eq 0 -or @($traceData.toolInvocations).Count -eq 0) {
|
||||
throw "Exact Trace did not include both Agent steps and tool invocation evidence."
|
||||
}
|
||||
if ($null -eq $traceData.run.selfEvaluation) {
|
||||
throw "Exact Trace run did not include selfEvaluation."
|
||||
}
|
||||
|
||||
$feedbackBody = @{
|
||||
sessionId = $SessionId
|
||||
runId = $runId
|
||||
@@ -105,11 +159,13 @@ $feedbackRequest = @{
|
||||
Body = $feedbackBody
|
||||
}
|
||||
$feedback = Invoke-RestMethod @feedbackRequest
|
||||
if ($feedback.success -ne $true) {
|
||||
throw "Feedback request was not successful for runId '$runId'."
|
||||
}
|
||||
|
||||
$feedbackPath = Join-Path $OutputDir "feedback-response.json"
|
||||
$feedback | ConvertTo-Json -Depth 30 | Set-Content -Encoding UTF8 -Path $feedbackPath
|
||||
|
||||
$traceData = Get-TraceData -TraceResponse $trace
|
||||
$selfEvaluation = Get-SelfEvaluation -TraceData $traceData
|
||||
$verifierEvaluation = $null
|
||||
if ($null -ne $selfEvaluation) {
|
||||
@@ -138,6 +194,7 @@ if ($null -ne $promptAudit) {
|
||||
$promptAuditVersion = $promptAudit.version
|
||||
}
|
||||
$toolNames = Get-ToolNames -TraceData $traceData
|
||||
$transitionCount = @($orchestrationTrace.transitions).Count
|
||||
$summaryPath = Join-Path $OutputDir "interview-demo-summary.json"
|
||||
|
||||
$summary = [ordered]@{
|
||||
@@ -149,6 +206,12 @@ $summary = [ordered]@{
|
||||
gatekeeperStatus = $gatekeeperStatus
|
||||
gatekeeperRuleSetVersion = $gatekeeperRuleSetVersion
|
||||
promptAuditVersion = $promptAuditVersion
|
||||
orchestrationVersion = $orchestrationTrace.version
|
||||
finalNode = $orchestrationTrace.final_node
|
||||
terminationReason = $orchestrationTrace.termination_reason
|
||||
degraded = [bool]$orchestrationTrace.degraded
|
||||
transitionCount = $transitionCount
|
||||
evidenceRetryCount = [int]$orchestrationTrace.evidence_retry_count
|
||||
toolNames = $toolNames
|
||||
paths = [ordered]@{
|
||||
chat = $chatPath
|
||||
@@ -165,4 +228,6 @@ Write-Host "Interview demo preflight completed."
|
||||
Write-Host "Verdict: $($summary.verdict)"
|
||||
Write-Host "Gatekeeper rules: $($summary.gatekeeperRuleSetVersion)"
|
||||
Write-Host "Prompt audit: $($summary.promptAuditVersion)"
|
||||
Write-Host "Graph final node: $($summary.finalNode)"
|
||||
Write-Host "Graph termination: $($summary.terminationReason)"
|
||||
Write-Host "Summary: $summaryPath"
|
||||
|
||||
@@ -7,16 +7,30 @@
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
| `data.runId` / `data.run.runId` | 是否等于 demo 响应中的 `runId` | `runId` 精确绑定这一次诊断运行 |
|
||||
| `data.session.sessionId` | 是否等于 `mvp-demo-payment-timeout-001` | `sessionId` 保留多轮上下文,Trace 精确回放依赖 `runId` |
|
||||
| `data.session.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
|
||||
| `data.session.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
|
||||
| `data.session.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | 如果是 Chat V2 链路,是否记录 Gatekeeper 规则版本 | 安全规则可审计、可回归 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.version` | 如果是 Chat V2 链路,是否记录 Prompt 审计版本 | Prompt 变更可解释、可回归 |
|
||||
| `data.session.selfEvaluation.verifier_evaluation.prompt_audit.prompts[*].version` | 是否记录 planner / executor / verifier / composer 版本 | 便于定位 Prompt 变更影响 |
|
||||
| `data.session.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在当前 diagnosis run 上 |
|
||||
| `data.run.sessionId` | 是否等于本次 Chat 请求的唯一 sessionId | `sessionId` 保留多轮上下文,Trace 精确回放依赖 `runId` |
|
||||
| `data.run.query` | 是否包含支付超时问题 | Trace 记录了原始用户意图 |
|
||||
| `data.run.answer` | 是否包含最终诊断答案 | 最终答案没有脱离 Trace |
|
||||
| `data.run.selfEvaluation` | 是否包含 verifier 或 rule evaluation | 答案经过质量门,不只是模型原始输出 |
|
||||
| `data.run.selfEvaluation.verifier_evaluation.gatekeeper_result.rule_set_version` | 是否记录 Gatekeeper 规则版本 | 安全规则可审计、可回归 |
|
||||
| `data.run.selfEvaluation.verifier_evaluation.prompt_audit.version` | 是否记录 Prompt 审计版本 | Prompt 变更可解释、可回归 |
|
||||
| `data.run.selfEvaluation.verifier_evaluation.prompt_audit.prompts[*].version` | 是否记录 planner / executor / verifier / composer 版本 | 便于定位 Prompt 变更影响 |
|
||||
| `data.run.feedback` | 提交反馈后是否变为 `useful` | 用户反馈挂在当前 diagnosis run 上 |
|
||||
|
||||
## 2. Agent 步骤
|
||||
## 2. StateGraph 路由
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
| `data.run.orchestrationTrace.version` | 是否存在当前 trace contract 版本 | 路由摘要可演进、可兼容 |
|
||||
| `data.run.orchestrationTrace.transitions[*]` | 是否记录实际经过的 Node 和 route | Graph 条件边不是从日志推断 |
|
||||
| `data.run.orchestrationTrace.final_node` | 最终是 Composer 还是 Fallback | 正常输出与安全降级明确区分 |
|
||||
| `data.run.orchestrationTrace.termination_reason` | 是否给出终止原因 | 每次 Run 都有可解释终点 |
|
||||
| `data.run.orchestrationTrace.degraded` | 是否发生安全降级 | fallback 是可审计行为 |
|
||||
| `data.run.orchestrationTrace.evidence_retry_count` | 是否为 0 或 1 | 补证据循环有硬上限 |
|
||||
| `interview-demo-summary.json.finalNode` 等摘要字段 | 是否与 exact Trace 一致 | summary 只消费 Run 路由真理源 |
|
||||
|
||||
`orchestrationTrace` 负责路由;`selfEvaluation` 负责证据和答案质量;AgentStep/ToolInvocation 负责详细执行与工具证据。三者不能互相替代。
|
||||
|
||||
## 3. Agent 步骤
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
@@ -25,7 +39,7 @@
|
||||
| `data.steps[*].durationMs` | 是否有步骤耗时 | Trace 可用于耗时分析 |
|
||||
| `data.steps[*].tokenCount` | 如可用,是否记录 token | Trace 可用于模型成本分析 |
|
||||
|
||||
## 3. 工具证据
|
||||
## 4. 工具证据
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
@@ -37,7 +51,7 @@
|
||||
| `data.toolInvocations[*].retrievalDetails.evidence_refs` | 是否包含 `raw_path + text` | Gatekeeper 可以用代码核对 Executor 引用 |
|
||||
| `data.toolInvocations[*].relevanceLevel` | 是否有相关性等级 | 可解释检索结果强弱 |
|
||||
|
||||
## 4. Summary
|
||||
## 5. Summary
|
||||
|
||||
| JSON path | 检查点 | 面试讲点 |
|
||||
|---|---|---|
|
||||
@@ -46,7 +60,7 @@
|
||||
| `data.summary.hasVerifierEvaluation` | 是否存在 Verifier 结果 | 最终答案经过质量门 |
|
||||
| `data.summary.hasFeedback` | 提交反馈后是否为 true | 人类反馈闭环完成 |
|
||||
|
||||
## 5. 好的结果长什么样
|
||||
## 6. 好的结果长什么样
|
||||
|
||||
```text
|
||||
同一个 session id + run id
|
||||
@@ -54,5 +68,6 @@
|
||||
-> 持久化 agent steps
|
||||
-> 持久化 evidence tool calls
|
||||
-> verifier / self-evaluation
|
||||
-> run.orchestrationTrace routing summary
|
||||
-> feedback attached to the same run
|
||||
```
|
||||
|
||||
+7
-5
@@ -4,13 +4,15 @@ This folder contains the fixed offline regression set for the MVP diagnosis Agen
|
||||
|
||||
## Background
|
||||
|
||||
The current diagnosis chain is:
|
||||
The current complex Chat diagnosis chain is a bounded StateGraph:
|
||||
|
||||
```text
|
||||
Planner -> Executor -> Gatekeeper -> Verifier -> Composer -> final answer
|
||||
Planner -> Executor -> Gatekeeper -> Verified Input -> Verifier -> Composer -> final answer
|
||||
| |
|
||||
+ bounded evidence retry + safe Fallback
|
||||
```
|
||||
|
||||
Stages 1-4 introduced Executor V2 structured output, deterministic Gatekeeper audit, Verifier `claim_checks`, and Composer final-answer rendering. Stage 5 makes those audit fields part of the offline regression harness so future prompt, tool, or chain changes can be checked without relying on a one-off demo.
|
||||
Executor V2 structured output, deterministic Gatekeeper audit, verified-only Verifier input, `claim_checks`, Composer rendering, and StateGraph routing are covered by deterministic tests so future prompt, tool, or graph changes can be checked without relying on a one-off demo.
|
||||
|
||||
## Scope
|
||||
|
||||
@@ -55,10 +57,10 @@ Run the focused evaluator test:
|
||||
mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test
|
||||
```
|
||||
|
||||
Run the broader phase-5 regression set:
|
||||
Run the authoritative Graph layers plus the fixed evaluator checks:
|
||||
|
||||
```powershell
|
||||
mvn "-Dtest=DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
|
||||
mvn -q "-Dtest=DiagnosisGraphWorkflowTest,DiagnosisGraphNodeContractTest,ChatServiceGraphIntegrationTest,DiagnosisTraceEvaluatorTest,DiagnosisEvalBaselineDiffTest,ExecutorGatekeeperServiceTest" test
|
||||
```
|
||||
|
||||
When fixtures or evaluator rules change, regenerate both baseline reports from the same case file and fixture directory, then update JSON and Markdown together.
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# MVP Issues 索引
|
||||
|
||||
**更新日期**:2026-07-16
|
||||
**更新日期**:2026-07-20
|
||||
**状态**:按活跃问题、设计笔记、RAG 问题集和已归档问题整理
|
||||
|
||||
## 目录约定
|
||||
@@ -16,7 +16,6 @@
|
||||
|
||||
| 名称 | 标题 | 严重程度 | 状态 | 文件 |
|
||||
|---|---|---|---|---|
|
||||
| ISS-011 | Chat 诊断 StateGraph 编排改造 | 高 | 待实现 | [active/ISS-011-chat-diagnosis-stategraph-orchestration.md](active/ISS-011-chat-diagnosis-stategraph-orchestration.md) |
|
||||
| ISS-003 | MVP 设计与实现 Review 收敛 | 高 | 待规划 | [active/ISS-003-mvp-design-implementation-review.md](active/ISS-003-mvp-design-implementation-review.md) |
|
||||
| ISS-004 | Executor 域级检索水位控制 | 低 | 待规划 | [active/ISS-004-executor-domain-hard-limit.md](active/ISS-004-executor-domain-hard-limit.md) |
|
||||
| executor-evidence-attribution-hallucination | Executor 证据归因幻觉 | 高 | 待规划 | [active/executor-evidence-attribution-hallucination.md](active/executor-evidence-attribution-hallucination.md) |
|
||||
@@ -53,6 +52,7 @@
|
||||
|
||||
| 名称 | 标题 | 状态 | 文件 |
|
||||
|---|---|---|---|
|
||||
| ISS-011 | Chat 诊断 StateGraph 编排改造 | 已归档 | [archived/ISS-011-chat-diagnosis-stategraph-orchestration.md](archived/ISS-011-chat-diagnosis-stategraph-orchestration.md) |
|
||||
| ISS-001 | Executor 重复召回同一文档 | 已修复 | [archived/ISS-001-duplicate-retrieval.md](archived/ISS-001-duplicate-retrieval.md) |
|
||||
| ISS-002 | Executor 无约束重复调用 lookup_knowledge | 已修复 | [archived/ISS-002-executor-unconstrained-lookup.md](archived/ISS-002-executor-unconstrained-lookup.md) |
|
||||
| ISS-005 | 证据链补齐与降级契约收敛 | 已归档 | [archived/ISS-005-evidence-trace-hardening.md](archived/ISS-005-evidence-trace-hardening.md) |
|
||||
|
||||
+56
-55
@@ -1,8 +1,9 @@
|
||||
# ISS-011 Chat 诊断 StateGraph 编排改造
|
||||
|
||||
**状态**:待实现
|
||||
**状态**:已归档
|
||||
**严重程度**:高
|
||||
**发现时间**:2026-07-16
|
||||
**完成时间**:2026-07-20
|
||||
**来源**:OnCall / Agent 编排模拟面试、当前 Chat 复杂诊断调用链复核
|
||||
**预计实施周期**:2–3 个工作日
|
||||
|
||||
@@ -871,69 +872,69 @@ Eval baseline
|
||||
|
||||
### 编排
|
||||
|
||||
- [ ] 每次进入 Planner 阶段时,INVALID_OUTPUT / RETRYABLE_FAILED 最多触发一次技术重试。
|
||||
- [ ] Planner NON_RETRYABLE_FAILED 或当前阶段第二次技术失败直接进入 Fallback。
|
||||
- [ ] Planner 技术重试不增加 `evidence_retry_count`,补证据重新进入 Planner 时重置当前阶段的 `planner_retry_count`。
|
||||
- [ ] Executor FAILED / TOOL_BLOCKED 后不会执行 Gatekeeper 和 Verifier。
|
||||
- [ ] Executor INVALID_OUTPUT 不重试,不执行 Gatekeeper、Verifier 和模型 Composer。
|
||||
- [ ] TOOL_BLOCKED 只用于工具层明确阻断且不存在合法 Executor 输出的场景。
|
||||
- [ ] 工具空结果或工具失败后仍形成合法 Executor 输出时状态为 COMPLETED,并继续 Gatekeeper。
|
||||
- [ ] Executor 合法 no-evidence 会继续执行 Gatekeeper 和 Verifier。
|
||||
- [ ] Gatekeeper REJECT 直接进入 Fallback,不执行 Verifier。
|
||||
- [ ] Gatekeeper LOW_CONFID 且零条已验真 binding 时直接进入 Fallback。
|
||||
- [ ] Gatekeeper LOW_CONFID 且存在已验真 binding 时,Verifier 只接收通过校验的 binding。
|
||||
- [ ] Gatekeeper PASS 和可继续的 LOW_CONFID 都经过 Verifier Input Builder。
|
||||
- [ ] Verifier 只接收通过 binding 对应的 `verified_evidence`,不接收完整 `tool_trace_summary`。
|
||||
- [ ] 未被 Executor 引用或未通过 Gatekeeper 的工具结果不能进入 Verifier 输入。
|
||||
- [ ] Gatekeeper LOW_CONFID 路径的 `effective_verdict` 不得升级为 PASS。
|
||||
- [ ] Gatekeeper 原始 pass/fail + severity 正确标准化为 PASS / LOW_CONFID / REJECT,未知状态安全映射为 REJECT。
|
||||
- [ ] Verifier 执行状态与诊断 verdict 分离,任何失败状态不得出现在 model/effective verdict 中。
|
||||
- [ ] Composer 和 Graph 条件边只读取 `effective_verdict`。
|
||||
- [ ] Verifier INVALID_OUTPUT / RETRYABLE_FAILED 使用相同 verified input 最多技术重试一次,且不重新执行 Gatekeeper、Executor 或工具。
|
||||
- [ ] Composer INVALID_OUTPUT / RETRYABLE_FAILED 使用相同安全输入最多技术重试一次,且不重新执行 Verifier 或前序节点。
|
||||
- [ ] Verifier 第二次技术失败或 NON_RETRYABLE_FAILED 的 Fallback 不输出 Executor claim。
|
||||
- [ ] Composer 第二次技术失败或 NON_RETRYABLE_FAILED 使用确定性安全模板。
|
||||
- [ ] `verifier_retry_count`、`composer_retry_count` 和 `evidence_retry_count` 互相独立。
|
||||
- [ ] Verifier LOW_CONFID 最多触发一次 Planner 补证据。
|
||||
- [ ] LOW_CONFID 补证据循环受一次补查上限和 Graph recursion limit 限制。
|
||||
- [ ] Gatekeeper verdict ceiling 导致的 LOW_CONFID 不触发补证据。
|
||||
- [ ] 无法从 `facts_checked` 提取有效 `evidence_gaps` 时不触发补证据。
|
||||
- [ ] 第二轮 Planner 只输出增量计划,不扩大诊断范围或重复成功查询。
|
||||
- [ ] 第二轮 Executor 只执行增量查询,但输出完整 `executor_evidence_v2` 快照,而不是仅输出新增片段。
|
||||
- [ ] 第二轮完整快照包含需要保留的第一轮可信 claims,并由 Gatekeeper 对全部 binding 重新验真。
|
||||
- [ ] Java 编排层不对两轮 claim 文本进行语义合并。
|
||||
- [ ] Composer 技术重试耗尽或不可重试失败时使用固定模板结束。
|
||||
- [x] 每次进入 Planner 阶段时,INVALID_OUTPUT / RETRYABLE_FAILED 最多触发一次技术重试。
|
||||
- [x] Planner NON_RETRYABLE_FAILED 或当前阶段第二次技术失败直接进入 Fallback。
|
||||
- [x] Planner 技术重试不增加 `evidence_retry_count`,补证据重新进入 Planner 时重置当前阶段的 `planner_retry_count`。
|
||||
- [x] Executor FAILED / TOOL_BLOCKED 后不会执行 Gatekeeper 和 Verifier。
|
||||
- [x] Executor INVALID_OUTPUT 不重试,不执行 Gatekeeper、Verifier 和模型 Composer。
|
||||
- [x] TOOL_BLOCKED 只用于工具层明确阻断且不存在合法 Executor 输出的场景。
|
||||
- [x] 工具空结果或工具失败后仍形成合法 Executor 输出时状态为 COMPLETED,并继续 Gatekeeper。
|
||||
- [x] Executor 合法 no-evidence 会继续执行 Gatekeeper 和 Verifier。
|
||||
- [x] Gatekeeper REJECT 直接进入 Fallback,不执行 Verifier。
|
||||
- [x] Gatekeeper LOW_CONFID 且零条已验真 binding 时直接进入 Fallback。
|
||||
- [x] Gatekeeper LOW_CONFID 且存在已验真 binding 时,Verifier 只接收通过校验的 binding。
|
||||
- [x] Gatekeeper PASS 和可继续的 LOW_CONFID 都经过 Verifier Input Builder。
|
||||
- [x] Verifier 只接收通过 binding 对应的 `verified_evidence`,不接收完整 `tool_trace_summary`。
|
||||
- [x] 未被 Executor 引用或未通过 Gatekeeper 的工具结果不能进入 Verifier 输入。
|
||||
- [x] Gatekeeper LOW_CONFID 路径的 `effective_verdict` 不得升级为 PASS。
|
||||
- [x] Gatekeeper 原始 pass/fail + severity 正确标准化为 PASS / LOW_CONFID / REJECT,未知状态安全映射为 REJECT。
|
||||
- [x] Verifier 执行状态与诊断 verdict 分离,任何失败状态不得出现在 model/effective verdict 中。
|
||||
- [x] Composer 和 Graph 条件边只读取 `effective_verdict`。
|
||||
- [x] Verifier INVALID_OUTPUT / RETRYABLE_FAILED 使用相同 verified input 最多技术重试一次,且不重新执行 Gatekeeper、Executor 或工具。
|
||||
- [x] Composer INVALID_OUTPUT / RETRYABLE_FAILED 使用相同安全输入最多技术重试一次,且不重新执行 Verifier 或前序节点。
|
||||
- [x] Verifier 第二次技术失败或 NON_RETRYABLE_FAILED 的 Fallback 不输出 Executor claim。
|
||||
- [x] Composer 第二次技术失败或 NON_RETRYABLE_FAILED 使用确定性安全模板。
|
||||
- [x] `verifier_retry_count`、`composer_retry_count` 和 `evidence_retry_count` 互相独立。
|
||||
- [x] Verifier LOW_CONFID 最多触发一次 Planner 补证据。
|
||||
- [x] LOW_CONFID 补证据循环受一次补查上限和 Graph recursion limit 限制。
|
||||
- [x] Gatekeeper verdict ceiling 导致的 LOW_CONFID 不触发补证据。
|
||||
- [x] 无法从 `facts_checked` 提取有效 `evidence_gaps` 时不触发补证据。
|
||||
- [x] 第二轮 Planner 只输出增量计划,不扩大诊断范围或重复成功查询。
|
||||
- [x] 第二轮 Executor 只执行增量查询,但输出完整 `executor_evidence_v2` 快照,而不是仅输出新增片段。
|
||||
- [x] 第二轮完整快照包含需要保留的第一轮可信 claims,并由 Gatekeeper 对全部 binding 重新验真。
|
||||
- [x] Java 编排层不对两轮 claim 文本进行语义合并。
|
||||
- [x] Composer 技术重试耗尽或不可重试失败时使用固定模板结束。
|
||||
|
||||
### 证据和安全
|
||||
|
||||
- [ ] Gatekeeper 规则语义不放宽。
|
||||
- [ ] Verifier 只消费已验真证据。
|
||||
- [ ] no-evidence 不得表达为已排除或问题不存在。
|
||||
- [ ] REJECT 降级不泄漏 Executor 原始答案和未验证根因。
|
||||
- [ ] Executor INVALID_OUTPUT、Gatekeeper REJECT 和零条可信 binding 的固定 Fallback 不输出任何 Executor claim。
|
||||
- [ ] 前置验证失败 Fallback 只展示校验状态、工具执行概况、诊断限制和人工复核建议。
|
||||
- [x] Gatekeeper 规则语义不放宽。
|
||||
- [x] Verifier 只消费已验真证据。
|
||||
- [x] no-evidence 不得表达为已排除或问题不存在。
|
||||
- [x] REJECT 降级不泄漏 Executor 原始答案和未验证根因。
|
||||
- [x] Executor INVALID_OUTPUT、Gatekeeper REJECT 和零条可信 binding 的固定 Fallback 不输出任何 Executor claim。
|
||||
- [x] 前置验证失败 Fallback 只展示校验状态、工具执行概况、诊断限制和人工复核建议。
|
||||
|
||||
### 数据与审计
|
||||
|
||||
- [ ] Graph 使用 runId 作为 threadId。
|
||||
- [ ] Agent step、tool invocation 和 self_evaluation 仍绑定正确 runId。
|
||||
- [ ] `orchestration_trace` 只写入当前 diagnosis run,不污染其他 run 或 session 级数据。
|
||||
- [ ] `orchestration_trace` 不包含 Prompt、模型思考、工具原文和 Graph State 快照。
|
||||
- [ ] `orchestration_trace.transitions` 由有界 `orchestration_events` 生成,与实际节点执行顺序一致。
|
||||
- [ ] 可处理异常发生时,已经产生的 orchestration events 能够 best-effort 写入当前 run。
|
||||
- [ ] Trace 能展示实际节点路径、重试原因和终止原因。
|
||||
- [ ] 每个新 StateGraph Chat run 的 `run.orchestrationTrace` 非空,且顶层和兼容 `session` 投影不重复该字段。
|
||||
- [ ] Run 最终状态、答案、耗时、Token 和工具调用数正确回填。
|
||||
- [ ] 所有成功生成安全响应的终止路径将 Run 标记为 SUCCESS,并通过 verdict 或 `orchestrationTrace.degraded` 表达质量。
|
||||
- [ ] 只有未处理异常、持久化失败或无法生成安全响应时将 Run 标记为 FAILED。
|
||||
- [x] Graph 使用 runId 作为 threadId。
|
||||
- [x] Agent step、tool invocation 和 self_evaluation 仍绑定正确 runId。
|
||||
- [x] `orchestration_trace` 只写入当前 diagnosis run,不污染其他 run 或 session 级数据。
|
||||
- [x] `orchestration_trace` 不包含 Prompt、模型思考、工具原文和 Graph State 快照。
|
||||
- [x] `orchestration_trace.transitions` 由有界 `orchestration_events` 生成,与实际节点执行顺序一致。
|
||||
- [x] 可处理异常发生时,已经产生的 orchestration events 能够 best-effort 写入当前 run。
|
||||
- [x] Trace 能展示实际节点路径、重试原因和终止原因。
|
||||
- [x] 每个新 StateGraph Chat run 的 `run.orchestrationTrace` 非空,且顶层和兼容 `session` 投影不重复该字段。
|
||||
- [x] Run 最终状态、答案、耗时、Token 和工具调用数正确回填。
|
||||
- [x] 所有成功生成安全响应的终止路径将 Run 标记为 SUCCESS,并通过 verdict 或 `orchestrationTrace.degraded` 表达质量。
|
||||
- [x] 只有未处理异常、持久化失败或无法生成安全响应时将 Run 标记为 FAILED。
|
||||
|
||||
### 工程质量
|
||||
|
||||
- [ ] 新 Graph 测试覆盖所有分支。
|
||||
- [ ] `ChatServiceSequentialAgentTest` 已由新测试替换。
|
||||
- [ ] 不保留长期重复的 Sequential 和 Graph 两套实现。
|
||||
- [ ] 数据库 schema 仅新增 `diagnosis_run.orchestration_trace` nullable JSON 字段。
|
||||
- [ ] `/api/chat` 和证据协议不变;Trace API 仅在 `run` 对象新增必有的 `orchestrationTrace` 字段。
|
||||
- [x] 新 Graph 测试覆盖所有分支。
|
||||
- [x] `ChatServiceSequentialAgentTest` 已由新测试替换。
|
||||
- [x] 不保留长期重复的 Sequential 和 Graph 两套实现。
|
||||
- [x] 数据库 schema 仅新增 `diagnosis_run.orchestration_trace` nullable JSON 字段。
|
||||
- [x] `/api/chat` 和证据协议不变;Trace API 仅在 `run` 对象新增必有的 `orchestrationTrace` 字段。
|
||||
|
||||
---
|
||||
|
||||
Reference in New Issue
Block a user