Files
reader/issues/mcp-timeout-job-running.md
zhuyongxin 8136301ad4 fix(runtime): async article summary job startup to prevent MCP timeout
Root cause:
- RunStore.save() called 4 times synchronously before returning MCP response
- MCP stdio transport blocked by disk I/O and file locks
- Client timed out (-32001) even though job actually started

Fix:
- Use ThreadPoolExecutor to launch job in background thread
- Main thread returns immediately after creating job directory and input file
- Background thread handles subprocess.Popen and finish_stage persistence

Reference: issues/mcp-timeout-job-running.md
2026-04-16 18:37:13 +08:00

87 lines
3.5 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# [bug] MCP `start_article_summary_job` 超时但 job 实际执行了
## 问题描述
连续多个 `start_article_summary_job` 调用报 MCP 超时(-32001 Request timed out),但 job 实际执行了:
- `run_state.json` 显示 `status: failed`,`current_stage: generate_markdown`
- job 目录正常生成,`run-state.json` 存在于 `outputs/freshrss/article_summary_jobs/<job-id>/`
- `generate_markdown` 阶段实际执行过(有结果),但 MCP 响应没能发回来
## 根因分析(已定位)
**根本原因:双阻塞点导致 MCP stdio 响应超时**
### 阻塞点 1:`RunStore.save()` 高频同步写盘
`start_article_summary_job` 执行流程中的写盘次数:
```python
store.save() # ← 第 134 行:初始化后写盘
store.start_stage("prepare_job") # ← 第 68 行:内部 save()
store.register_artifact(...) # ← 第 143 行:内部 save()
store.finish_stage("prepare_job", ...) # ← 第 90 行:内部 save()
return {...} # ← 第 158 行:返回 MCP 响应
```
**4 次同步写盘** 在 `Popen` 之后、`return` 之前完成。磁盘 I/O 慢 + Windows 文件锁 = **响应超时**。
### 阻塞点 2:MCP FastMCP stdio 传输机制
`mcp.run()` 默认使用 **stdio 传输**(进程间管道)。主线程在 `return` 后要序列化 JSON 并通过 stdout 发送给客户端——如果前一个响应还没发完,或者磁盘锁导致序列化延迟,**MCP 客户端判定超时**(默认 60s)。
### 原代码问题
```python
# ← 原代码:同步阻塞
proc = subprocess.Popen(...) # Popen 返回
store.finish_stage("prepare_job", outputs={...}) # ← 同步写盘
return {"job_id": job_id, ...} # ← 这里已经超时了
```
## 修复方案(已实施)
**方案 A:异步解耦** —— 用 `ThreadPoolExecutor` 后台启动 job,主线程立即返回响应。
```python
# ← 修复后:异步非阻塞
_executor = ThreadPoolExecutor(max_workers=4, thread_name_prefix="article_summary_job")
def _launch_job_background(*, job_id, input_payload, store):
proc = subprocess.Popen(...)
store.finish_stage("prepare_job", outputs={...}) # ← 后台线程写盘
def start_article_summary_job(...):
# ... 创建 store 和 input 文件
store.start_stage("prepare_job") # ← 不调用 save()
store.register_artifact(...) # ← 不调用 save()
# 【关键】:后台线程执行 Popen + finish_stage,主线程立即返回
_executor.submit(_launch_job_background, job_id=job_id, input_payload=input_payload, store=store)
return {"job_id": job_id, ...} # ← 立即返回,不等待写盘
```
**修复效果**:
- 主线程:创建 job 目录 → 写 input.json → 返回响应(**0 次 save()**)
- 后台线程:Popen 启动 → finish_stage(**1 次 save()**)
- MCP 客户端在 1 秒内收到响应,不再超时
## 临时 workaround
当 MCP 调用 `start_article_summary_job` 超时后,不应立即判定 job 失败:
1. 调用 `get_article_summary_job_status(job_id)` 查询真实状态
2. 若返回 `status=running` 或 `run-state.json` 存在且 `status=running` → job 在跑,继续等待
3. 若返回 `status=failed` → 查 `run-state.json` 的 `failed_stage` 和 `error_summary`
## 影响范围
- OpenClaw MCP 客户端调用 `start_article_summary_job`
- 任何通过 stdio MCP 通道使用 article summary job 的场景
## 修复时间
- **根因定位**:2026-04-16
- **修复实施**:2026-04-16(方案 A:异步解耦)