Files
reader/docs/openclaw/openclaw-orchestration-flow.md
T

341 lines
8.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# OpenClaw → reader MCP 标准编排流程
## 1. 文档目的
本文档只回答一个问题:OpenClaw 在正式环境里应该如何编排 reader。
这里不重复介绍 reader 内部实现,只定义正式控制面:
- 如何启动日报
- 如何轮询 job
- 如何读取 run 结果
- 如何判断是否恢复
- 如何走异步恢复
- 什么时候直接新开 run 或人工介入
## 2. 当前正式入口
### 2.1 新 run
正式生产入口:
- `start_freshrss_pipeline_job`
- `get_freshrss_pipeline_job_status`
- `get_freshrss_pipeline_job_result`
同步入口:
- `run_freshrss_openclaw_pipeline`
同步入口只保留给 debug / fallback,不再是正式编排默认路径。
### 2.2 run 级读取
正式 run 级读取接口:
- `get_run_status`
- `list_runs`
- `list_run_artifacts`
- `get_delivery_payload`
- `get_run_report`
### 2.3 恢复
正式恢复入口:
- `inspect_resume_plan`
- `start_resume_job`
- `get_resume_job_status`
- `get_resume_job_result`
同步恢复入口:
- `resume_run`
`resume_run` 只保留给 debug / fallback。
## 3. 编排基本原则
### 3.1 OpenClaw 不手拼路径
OpenClaw 不应自己推导这些路径:
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
需要路径时,只消费 MCP 返回值:
- `output_dir`
- `artifact.path`
- `delivery_output`
- `report_output`
### 3.2 顶层 `status` 才是分支依据
`get_run_status` 和 job status 接口都可能做状态收敛。
因此:
- 优先使用顶层 `status`
- `status_source` 用来解释状态来自原始 state 还是收敛结果
- `state_conflict=true` 说明底层状态文件已经落后于真实产物
不要再拿旧的 `raw_status`、`raw_current_stage` 或早期阶段名重新做分支。
### 3.3 默认生产语义
正式生产运行默认:
- `mark_read=true`
- `debug_artifacts=false`
只有 debug / test / validation 时才放宽。
## 4. 标准 Happy Path
### Step 1: 启动新 job
调用:
- `start_freshrss_pipeline_job`
推荐参数:
```json
{
"limit": 5,
"mark_read": true,
"include_read": false,
"debug_artifacts": false,
"timeout_seconds": 60,
"max_retries": 2
}
```
预期:
- 立即返回 `job_id`
- 后续由 OpenClaw 轮询 job,而不是同步等待整条流水线
### Step 2: 轮询 job
调用:
- `get_freshrss_pipeline_job_status(job_id=...)`
根据返回:
- `status=running`:继续轮询
- `status=success`:读取 job result
- `status=failed`:进入失败处理
额外规则:
- 如果 `status_source=linked_run_reconciliation`,说明 outer job state 已落后,但 linked run 已经给出可用终态
- 如果 `status_source=stale_job_state_timeout`,把它当成终态失败,不要继续无限轮询
### Step 3: 读取 job result
调用:
- `get_freshrss_pipeline_job_result(job_id=...)`
预期读取:
- `run_id`
- `output_dir`
- `delivery_output`
- `report_output`
从这一刻开始,`run_id` 是正式的稳定句柄。
### Step 4: 读取 run 级状态与结果
调用:
- `get_run_status(run_id=...)`
- `get_delivery_payload(run_id=...)`
- `get_run_report(run_id=...)`
根据 `get_run_status`:
- `status=running`:继续观察
- `status=success`:继续下游 digest / 发布 / 汇报
- `status=failed`:进入恢复或重跑决策
- `status=partial`:优先检查 report、artifacts 和 recovery
如果 `status_source=run_report_reconciliation`,说明 `run-state.json` 已经过期,但 reader 已经根据终态产物收敛出有效状态。
如果 `status_source=stale_run_state_timeout`,说明 reader 认为该 run 长时间未收敛且没有终态产物,应按失败处理。
## 5. 恢复决策
### 5.1 先看预检,不要直接恢复
恢复前固定动作:
- 先调用 `inspect_resume_plan(run_id)`
只在以下条件同时成立时才启动恢复:
- `can_resume=true`
- `recommended_action=resume`
重点字段:
- `requested_resume_from_stage`
- `resume_from_stage`
- `resume_decision_source`
- `artifact_resume_from_stage`
- `artifact_snapshot`
### 5.2 正式恢复路径
正式恢复控制面:
1. `start_resume_job(run_id)`
2. `get_resume_job_status(job_id)`
3. `get_resume_job_result(job_id)`
不要再把同步 `resume_run(run_id)` 当成正式恢复入口。
### 5.3 当前支持范围
当前只支持:
- 带有效 `run-state.json` 的 `freshrss_daily_digest` run
- 从以下阶段恢复:
- `generate_summaries`
- `apply_filters`
- `build_delivery_payload`
- `write_run_report`
当前不支持:
- `fetch_feed`
- `extract_articles`
正式生产恢复优先依赖:
- `summary/summary-batch.json`
- `candidates/candidate-batch.json`
### 5.4 什么时候不要恢复
以下情况直接新开 run 更合理:
- `recommended_action=start_new_run`
- `recommended_action=read_terminal_result`
- 没有有效 `run-state.json`
- 恢复所需关键 artifacts 缺失
- 连续恢复失败
## 6. 状态到动作映射
| 接口 | 状态 | OpenClaw 动作 |
| --- | --- | --- |
| `get_freshrss_pipeline_job_status` | `running` | 继续轮询 job |
| `get_freshrss_pipeline_job_status` | `success` | 读取 `get_freshrss_pipeline_job_result` |
| `get_freshrss_pipeline_job_status` | `failed` | 结束本次 job,必要时读 linked run |
| `get_run_status` | `running` | 继续观察 run |
| `get_run_status` | `success` | 读取 `get_delivery_payload` / `get_run_report` |
| `get_run_status` | `failed` | 先看 `inspect_resume_plan` |
| `inspect_resume_plan` | `recommended_action=resume` | 启动 `start_resume_job` |
| `inspect_resume_plan` | `recommended_action=read_terminal_result` | 直接读 run 结果,不恢复 |
| `inspect_resume_plan` | `recommended_action=start_new_run` | 新开 run 或人工介入 |
## 7. 人工介入条件
出现以下任一情况时,建议不要自动编排:
- 连续恢复失败
- payload / report 结构不符合预期
- `get_run_status` 与实际产物长期明显冲突
- FreshRSS、LLM 或外部依赖异常
- 恢复判定结果和编排预期不一致
## 8. 结论
当前 OpenClaw 的正式调用方式已经收口为两条异步控制面:
- 主日报:`start_freshrss_pipeline_job -> poll -> get result -> run reads`
- 恢复:`inspect_resume_plan -> start_resume_job -> poll -> get result`
同步 `run_freshrss_openclaw_pipeline` 和 `resume_run` 仅用于 debug / fallback,不应再作为默认正式编排路径。
### 7.2 `get_run_report`
用途:
- 获取 run 的结果摘要与关键元信息
- 用于状态判断、排障、补充上下文
OpenClaw 应做:
- 作为诊断与编排辅助信息读取
- 不把 raw file path 解析逻辑继续散落到 skill 里
---
## 8. 最小编排动作表
### 8.1 标准生产执行
1. 调 `run_freshrss_openclaw_pipeline`
2. 拿 `run_id`
3. 轮询 `get_run_status`
4. 若 `success`:
- 调 `get_delivery_payload`
- 调 `get_run_report`
5. 进入 digest / Hugo / chat / IMA 下游编排
### 8.2 失败恢复执行
1. 调 `get_run_status`
2. 若 `failed && recovery.resumable=true`:
- 调 `start_resume_job`
- 轮询 `get_resume_job_status`
- 读取 `get_resume_job_result`
3. 恢复后再次:
- 调 `get_run_status`
- 若成功,再读 payload / report
4. 若恢复失败或明确不可恢复:
- 新开 run 或人工介入
---
## 9. 不推荐做法
以下做法不应再作为正式主路径:
- 让 OpenClaw 直接长时间 `exec` reader CLI 作为主要生产入口
- 让 OpenClaw 自己拼 reader 输出路径来判断成功/失败
- 让 OpenClaw 自己读取 `outputs/.../*.json` 作为正式结果源
- 在未确认恢复支持范围外的失败点上强行恢复
CLI 现在的定位是:
- debug
- fallback
- 人工排障
而不是正式生产主入口。
---
## 10. 当前已知局限
- 恢复能力仍是最小实现,不支持任意 stage 任意重入
- 历史无 `run-state.json` 的 run 不支持正式恢复
- 极旧 run 的结果读取仍可能依赖保守目录扫描
- `write_run_report` 若涉及重新 `mark_read`,仍依赖 FreshRSS 环境和可用凭据
---
## 11. 一句话结论
OpenClaw 当前应把 reader 当作正式 MCP workflow service 使用:
**启动用 `start_freshrss_pipeline_job`,观测用 `get_run_status`,结果读取用 `get_delivery_payload` / `get_run_report`,恢复默认用 `inspect_resume_plan` + `start_resume_job`,不要再把 reader 当成长 CLI 任务和路径拼接仓库来驱动。**