feat: add async resume jobs and doc navigation

This commit is contained in:
root
2026-04-14 15:55:02 +08:00
parent 6705613aa4
commit b8727f1885
26 changed files with 2791 additions and 665 deletions
+51 -39
View File
@@ -1,6 +1,25 @@
# 文档索引
## 当前目录结构
## 当前最短阅读路径
1. `README.md`
- 仓库入口与常用脚本
2. `docs/current/context-reset-brief.md`
- 当前状态的最短摘要
3. `docs/openclaw/README.md`
- OpenClaw 集成文档导航
4. `docs/openclaw/openclaw-handoff.md`
- OpenClaw 接手总览
5. `docs/openclaw/openclaw-orchestration-flow.md`
- OpenClaw 正式编排手册
6. `docs/design/README.md`
- 设计文档导航,区分当前有效设计与背景草案
7. `plans/README.md`
- 规划文档导航
8. `TODO.md`
- 当前任务状态
## 目录结构
- `docs/README.md`
- 文档总索引
@@ -8,45 +27,18 @@
- 当前状态、收束入口、阶段导航
- `docs/design/`
- 当前实现的设计文档
- `docs/design/README.md`
- 设计文档导航
- `docs/openclaw/`
- OpenClaw 日报聚合与下游对象设计
- OpenClaw 集成文档、对象规范与历史归档
- `docs/notes/`
- 较上层的方案笔记与非最终设计
- `docs/notes/README.md`
- notes 导航
- `docs/archive/`
- 历史归档,不作为最新事实来源
## 当前推荐阅读顺序
1. `docs/current/context-reset-brief.md`
- 当前真实进度与下一步入口
2. `docs/openclaw/openclaw-handoff.md`
- 给 OpenClaw 的接手说明、环境变量、MCP 调用方式与已知限制
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
- 提供给 OpenClaw 的单篇结构化输入字段说明
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
- 提供给 OpenClaw 的批量投递 envelope 说明
5. `docs/design/summary-mcp-service-design.md`
- 当前 MCP 服务的职责、接口和边界
6. `docs/design/filter-rule-engine-design.md`
- 过滤层的输入输出、规则结构与当前实现
7. `docs/design/filter-rule-engine-usage.md`
- 规则怎么写、怎么跑、结果怎么解读的使用说明
8. `docs/design/daily-keyword-index-design.md`
- 日报级词元库与周期性词元清洗 skill 设计
9. `docs/design/markdown-sink-design.md`
- 第一版 Markdown sink 的输入输出、目录结构与落地方式
10. `docs/openclaw/openclaw-daily-digest-refactor.md`
- 为什么要从单篇入库改成 OpenClaw 日报聚合链路
11. `docs/openclaw/article-candidate-daily-digest-schema.md`
- `ArticleCandidateRecord`、`OpenClawCandidateInput` 与 `DailyDigest` 的正式设计
12. `docs/design/source-schema-design.md`
- `source -> item -> document` 的对象设计
13. `docs/notes/reading-pipeline-design-notes.md`
- 更上层的阅读流方案与阶段划分
14. `docs/design/summary-loop-explained.md`
- 当前 LLM 摘要校验闭环的解释
## 当前文档分层
## 按主题阅读
### 1. 当前状态与导航
@@ -58,11 +50,15 @@
- 当前阶段状态的最短摘要
- `docs/README.md`
- 文档索引与阅读顺序
- `plans/README.md`
- 规划文档导航
### 2. 当前实现设计
- `docs/design/README.md`
- 设计文档导航与状态说明
- `docs/design/summary-mcp-service-design.md`
- 当前内容提取 MCP 的真实设计
- 早期 content-extract MCP 设计草案,现主要保留背景参考价值
- `docs/design/summary-core-interface-design.md`
- 摘要/提取内核的接口抽象
- `docs/design/source-schema-design.md`
@@ -82,19 +78,23 @@
### 3. OpenClaw 与下游设计
- `docs/openclaw/README.md`
- OpenClaw 相关文档导航与归档边界
- `docs/openclaw/openclaw-handoff.md`
- OpenClaw 接手所需的运行说明、工具入口与已知限制
- `docs/openclaw/openclaw-orchestration-flow.md`
- OpenClaw 编排层的正式运行手册
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
- 提供给 OpenClaw 的单篇结构化输入字段说明
- `docs/openclaw/openclaw-delivery-payload-spec.md`
- 提供给 OpenClaw 的批量投递 envelope 说明
- `docs/openclaw/openclaw-daily-digest-refactor.md`
- 改造为 OpenClaw 日报聚合链路的原因与目标结构
- `docs/openclaw/article-candidate-daily-digest-schema.md`
- `ArticleCandidateRecord`、`OpenClawCandidateInput` 与 `DailyDigest` 的字段设计与对象关系
### 4. 方案笔记
- `docs/notes/README.md`
- notes 导航与使用边界
- `docs/notes/reading-pipeline-design-notes.md`
- 整体阅读流、规则、sink、push 的方案笔记
@@ -102,11 +102,23 @@
- `docs/archive/content-extract-mcp-mvp-archive.md`
- MVP 阶段归档,部分状态已被后续进展覆盖
- `docs/openclaw/archive/README.md`
- OpenClaw 历史文档归档说明
- `docs/openclaw/archive/formalization-summary-2026-04-07.md`
- 第一阶段正式化总结
- `docs/openclaw/archive/openclaw-daily-digest-refactor.md`
- 早期日报聚合改造背景
- `docs/openclaw/archive/digest-optimization-summary.md`
- 早期 digest 优化总结
- `docs/openclaw/archive/p1-status-reconciliation-plan-2026-04-14.md`
- `resume` / 状态收敛问题的阶段修复计划与回填
## 当前文档维护原则
## 维护原则
- `docs/current/context-reset-brief.md` 记录当前最新状态
- `TODO.md` 记录任务优先级与下一步
- `plans/README.md` 负责规划文档分层与导航
- `docs/design/README.md` 负责设计文档分层与导航
- `outputs/README.md` 记录当前输出目录约定
- `docs/archive/content-extract-mcp-mvp-archive.md` 只当历史快照,不再作为最新事实来源
- 新增阶段性进展,优先更新 `README.md`、`TODO.md`、`docs/current/context-reset-brief.md`
- 新增阶段性进展,优先更新 `README.md`、`TODO.md`、`docs/current/context-reset-brief.md`
+73 -121
View File
@@ -2,152 +2,104 @@
## 当前结论
当前仓库已经具备交付给 OpenClaw 的基础条件。
当前仓库已经具备作为 OpenClaw 上游服务的正式基础能力。
当前主链路是:
当前正式主链路是:
`FreshRSS 未读 -> RSS 内容提取 -> LLM 总结 -> 规则过滤 -> OpenClaw delivery payload`
OpenClaw 应通过 MCP 工具 `run_freshrss_openclaw_pipeline` 调用这条链路,而不是自行拼接脚本。
当前正式控制面已经收口为异步 job:
## 当前已完成
- 主日报:`start_freshrss_pipeline_job -> poll -> get result`
- 恢复:`inspect_resume_plan -> start_resume_job -> poll -> get result`
- 单篇总结:`start_article_summary_job -> poll -> get result`
- 已完成 FreshRSS `greader` API 接入与未读拉取
- 已完成 FreshRSS 条目到标准化 `item` 的映射
- 已完成 RSS-first 提取策略
- 已完成 LLM 总结与校验闭环
- 已完成规则引擎过滤
- 已完成 `ArticleCandidateRecord` 与 `OpenClawCandidateInput` 分层
- 已完成 `OpenClawDeliveryPayload` 批量投递结构
- 已完成 FreshRSS 已读状态回写
- 已完成“仅在最终 payload 成功写盘后再标记已读”的语义
- 已完成 MCP 工具 `run_freshrss_openclaw_pipeline`
- 已完成默认精简输出模式,减少中间文件
- 已完成日报级 `keywords` 词元库与全局词频统计
- 已完成 `keyword-cleanup-review` skill 骨架与 review bundle 脚本
- 已完成低复杂治理层:`term_cleanup_policy` / `term_watchlist` / `term_change_log`
- 已完成采纳建议写回脚本 `scripts/apply_term_suggestions.py`
同步 `run_freshrss_openclaw_pipeline`、`resume_run`、`generate_article_summaries` 仍保留,但只用于 debug / fallback。
## 当前 MCP 工具
## 当前权威入口
当前服务入口:
先看这些文档:
- `src/summary_mcp/server.py`
1. `README.md`
2. `docs/README.md`
3. `docs/openclaw/README.md`
4. `docs/openclaw/openclaw-handoff.md`
5. `docs/openclaw/openclaw-orchestration-flow.md`
6. `plans/README.md`
7. `TODO.md`
当前暴露的 MCP 工具:
如果问题是 OpenClaw 集成、状态分支或恢复策略,优先看 `docs/openclaw/`,不要先翻历史计划。
- `extract_url_content`
- `extract_item_content`
- `filter_summary_result`
- `run_freshrss_openclaw_pipeline`
## 当前正式能力
其中生产主入口是:
- FreshRSS 主日报 run 会落地 `run-state.json`
- `get_run_status` / `list_runs` / `list_run_artifacts` 提供 run 级观测
- `get_delivery_payload` / `get_run_report` 提供正式结果读取
- 查询层已经支持 stale state 与终态 artifacts 的状态收敛
- 主日报正式启动已切到 async job
- `resume` 已切到 async job,并在执行前先做 `inspect_resume_plan`
- 生产恢复依赖 `summary/summary-batch.json` 与 `candidates/candidate-batch.json`
- 单篇总结也已补齐 async job 形态
- `run_freshrss_openclaw_pipeline`
## 当前关键代码入口
## 当前关键文件
- MCP 服务入口:`src/summary_mcp/server.py`
- 主日报 workflow:`src/summary_mcp/workflows/freshrss_pipeline.py`
- run / artifact 查询:`src/summary_mcp/runtime/query_service.py`
- 主日报 async job:`src/summary_mcp/runtime/freshrss_pipeline_jobs.py`
- resume 预检与恢复:`src/summary_mcp/runtime/resume_service.py`
- resume async job:`src/summary_mcp/runtime/resume_jobs.py`
- 单篇总结 async job:`src/summary_mcp/runtime/article_summary_jobs.py`
- 关键词治理:`src/summary_mcp/core/keyword_index.py`
- MCP 服务入口
- `src/summary_mcp/server.py`
- FreshRSS 统一工作流
- `src/summary_mcp/workflows/freshrss_pipeline.py`
- 词元统计核心
- `src/summary_mcp/core/keyword_index.py`
- 词元统计模型
- `src/summary_mcp/models/keyword_index.py`
- 摘要循环
- `src/summary_mcp/core/summary_loop.py`
- 提取主流程
- `src/summary_mcp/core/pipeline.py`
- FreshRSS 集成
- `src/summary_mcp/integrations/freshrss.py`
- 规则引擎
- `src/summary_mcp/filters/engine.py`
- LLM 结果校验
- `src/summary_mcp/validators/llm_result.py`
- OpenClaw candidate 模型
- `src/summary_mcp/models/article_candidate.py`
- OpenClaw delivery 模型
- `src/summary_mcp/models/openclaw_delivery.py`
- 生产脚本入口
- `scripts/run_freshrss_pipeline.py`
- 词元统计重建脚本
- `scripts/build_keyword_index.py`
- 词元清洗 skill
- `skills/keyword-cleanup-review/SKILL.md`
- skill review bundle 脚本
- `skills/keyword-cleanup-review/scripts/build_review_bundle.py`
- 采纳建议写回脚本
- `scripts/apply_term_suggestions.py`
- 清洗治理配置
- `configs/term_cleanup_policy.json`
- `configs/term_watchlist.json`
- `configs/term_change_log.json`
- OpenClaw 交接说明
- `docs/openclaw/openclaw-handoff.md`
## 当前核心产物
## 当前输出规则
主日报稳定产物:
默认生产模式只输出:
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
- `outputs/freshrss/rerun/<run_dir>/raw/freshrss.raw.json`
- `outputs/freshrss/rerun/<run_dir>/summary/summary-batch.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/candidate-batch.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/digest-brief.json`
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
- `outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json`
- `raw/freshrss.raw.json`
- `candidates/openclaw-delivery-payload.json`
- `run-report.json`
job 状态目录:
同时会更新本地运行数据:
- `outputs/freshrss/pipeline_jobs/<job_id>/`
- `outputs/freshrss/resume_jobs/<job_id>/`
- `outputs/freshrss/article_summary_jobs/<job_id>/`
关键词运行数据:
- `data/term_index/daily/YYYY-MM-DD.json`
- `data/term_index/term_stats.json`
如果需要词元清洗审阅输入,可额外生成:
## 当前已验证
- `outputs/term_index/review/keyword-cleanup-bundle.json`
- FreshRSS 未读拉取与已读回写可用
- RSS-first 提取策略可用
- 主日报 MCP 主链路可触发并写出正式产物
- 状态查询与结果读取接口可用
- stale state / artifacts 收敛逻辑已落地
- `resume` 的 artifact-first 判定已落地
- `start_resume_job -> poll -> result` 已做本地 synthetic 验证
- 单篇总结 async job 可跑通
- 关键词 review bundle 与建议写回脚本可用
如果需要在人工确认后把建议正式写入 watchlist / change log,可使用:
## 当前主要限制
- `scripts/apply_term_suggestions.py`
- 某些源 RSS 正文不足时会被直接跳过
- 规则仍然偏保守,部分内容会落到 `review`
- `paywall` 启发式对中文仍可能误判
- 关键词治理还没有接入周期性调度
- `digest-brief.json` 仍没有独立 MCP 读取工具
- `resume` 目前的剩余主风险不再是恢复点判定,而是缺少真实生产环境的完整恢复验证
如果需要排障,可开启:
## 当前建议
- `debug_artifacts=true`
- 或脚本参数 `--debug-artifacts`
这样才会额外输出逐条中间文件。
## 当前验证状态
已经验证通过:
- FreshRSS 未读拉取成功
- 已读回写成功
- MCP 工具入口可直接触发完整链路
- 微信公众号样本可直接使用 RSS 提供的 `summary` 内容提取,不再回源抓网页
- 精简输出模式已实际跑通
- 日报级词元统计已通过离线样例验证,确认别名、停用词、非 `drop` 过滤和 rerun 覆盖逻辑正常
- `keyword-cleanup-review` skill 已通过 `quick_validate.py` 结构校验
- review bundle 脚本已实际跑通
- `apply_term_suggestions.py` 已通过 dry-run 与临时副本写回验证
## 当前已知限制
- 当前对 FreshRSS 条目采用 RSS-first 策略,不再回源抓原网页
- 如果 RSS 中没有足够正文内容,该条会直接跳过,不会进入后续总结
- 某些规则仍偏保守,部分内容可能落到 `review`
- `paywall` 相关启发式仍可能误判中文文本
- Webhook / 主动投递到 OpenClaw 外部接口尚未实现,当前是由 OpenClaw 通过 MCP 主动调用
- 词元清洗 skill 当前已支持“bundle 构建 -> 建议审阅 -> 人工确认写回 watchlist/change_log”,但尚未接入周期性调度
- 当前词元统计仍以前置 `OpenClawDeliveryPayload` 作为日报前代理输入,真实 `DailyDigest` 接入后还需切换上游
## 当前最建议的交接阅读顺序
1. `README.md`
2. `docs/openclaw/openclaw-handoff.md`
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
5. `docs/design/daily-keyword-index-design.md`
6. `skills/keyword-cleanup-review/SKILL.md`
7. `TODO.md`
## 一句话结论
当前仓库已经从“提取 MCP 原型”演进到“可供 OpenClaw 调用的 FreshRSS -> OpenClaw payload 上游处理器”,并已补上第一阶段的日报级词元统计能力和词元清洗 skill 骨架;后续重点转向 skill 周期调度、知识库状态流转和 webhook 接线。
- 把 `docs/openclaw/openclaw-orchestration-flow.md` 当成正式编排手册
- 把 `docs/openclaw/openclaw-handoff.md` 当成接手总览
- 把 `plans/README.md` 当成规划文档导航
- 把 `docs/openclaw/archive/` 和 `docs/archive/` 当成历史资料,不要当当前事实源
+56
View File
@@ -0,0 +1,56 @@
# 设计文档导航
## 使用原则
`docs/design/` 目录同时包含两类文档:
- 当前实现仍然有效的设计说明
- 早期架构草案和背景设计
不要默认把这里所有文档都当成当前生产事实。
当前生产事实仍以这些入口为准:
1. `README.md`
2. `docs/current/context-reset-brief.md`
3. `docs/openclaw/README.md`
4. `docs/openclaw/openclaw-handoff.md`
5. `docs/openclaw/openclaw-orchestration-flow.md`
6. `TODO.md`
## 当前实现仍然有效
- `filter-rule-engine-design.md`
- 规则过滤层的设计与职责边界
- `filter-rule-engine-usage.md`
- 规则引擎的使用说明
- `daily-keyword-index-design.md`
- 关键词索引与清洗治理设计
- `summary-loop-explained.md`
- LLM 摘要校验闭环说明
- `markdown-sink-design.md`
- Markdown sink 设计
## 当前仍有参考价值,但不是生产真相入口
- `summary-mcp-service-design.md`
- 早期 MCP 服务设计草案,部分定位已被后续 workflow service 演进覆盖
- `source-schema-design.md`
- 更偏对象建模和来源抽象的背景设计
- `summary-core-interface-design.md`
- 更偏早期摘要内核接口抽象
## 建议阅读顺序
如果你是在理解当前实现:
1. `filter-rule-engine-design.md`
2. `filter-rule-engine-usage.md`
3. `daily-keyword-index-design.md`
4. `summary-loop-explained.md`
5. `markdown-sink-design.md`
如果你是在回看背景设计:
1. `summary-mcp-service-design.md`
2. `source-schema-design.md`
3. `summary-core-interface-design.md`
+11
View File
@@ -1,5 +1,16 @@
# Source Schema 设计草案
## 状态说明
本文件偏对象建模和来源抽象,主要用于解释早期 schema 设计思路。
它不是当前生产运行手册,也不是当前 workflow service 的唯一真相来源。
如果你关注当前 OpenClaw 集成或运行状态,应优先看:
- `docs/current/context-reset-brief.md`
- `docs/openclaw/README.md`
- `docs/openclaw/openclaw-orchestration-flow.md`
## 1. 文档目的
本文档用于定义阅读流系统中的来源与内容对象模型,目标是把“来源分类”的讨论收敛成一套可执行的数据结构,供后续的抓取、摘要、过滤、入库和推送流程统一使用。
@@ -1,5 +1,17 @@
# Summary Core Interface 设计草案
## 状态说明
本文件记录的是较早期的摘要内核接口抽象。
它更适合用于理解背景设计,不应直接当成当前生产接口契约。
当前接口与编排真相请优先看:
- `README.md`
- `docs/current/context-reset-brief.md`
- `docs/openclaw/README.md`
- `docs/openclaw/openclaw-handoff.md`
## 1. 文档目的
本文档用于定义 `summary-core` 的输入输出接口,目标是把“页面摘要能力”从概念讨论收敛成一套稳定、可复用、可封装的数据接口。
+14 -1
View File
@@ -1,5 +1,18 @@
# Content Extract MCP Service 设计草案
## 状态说明
本文件主要记录早期 “content extract MCP” 的设计抽象。
它仍有背景参考价值,但不是当前生产事实入口。
当前生产能力已经演进为更完整的 workflow service,正式口径请优先看:
- `README.md`
- `docs/current/context-reset-brief.md`
- `docs/openclaw/README.md`
- `docs/openclaw/openclaw-handoff.md`
- `docs/openclaw/openclaw-orchestration-flow.md`
## 1. 文档目的
本文档用于定义当前仓库中已经落地的 MCP 服务设计,即“内容提取 MCP”。
@@ -12,7 +25,7 @@
- validator 与 LLM 摘要如何接在 MCP 之后
- 当前 MVP 已完成到哪一层
这份文档描述的是当前真实实现,而不是早期“摘要 MCP”设想。
这份文档主要记录当时实现阶段的设计取向,而不是当前生产阶段的唯一事实来源。
---
+22
View File
@@ -0,0 +1,22 @@
# Notes 导航
`docs/notes/` 保存的是更早期、讨论型、背景型方案笔记。
这些文档的用途是:
- 理解项目最初的问题空间
- 回看为什么会形成现在的对象分层和流程划分
这些文档不是当前生产事实来源。
当前如需判断“现在到底怎么跑”,优先看:
- `README.md`
- `docs/current/context-reset-brief.md`
- `docs/openclaw/README.md`
- `docs/openclaw/openclaw-orchestration-flow.md`
当前 notes:
- `reading-pipeline-design-notes.md`
- 早期阅读流方案讨论纪要
+36
View File
@@ -0,0 +1,36 @@
# OpenClaw 文档导航
## 当前有效文档
- `openclaw-handoff.md`
- 面向接手者的总览文档
- 说明 reader 的职责边界、MCP 工具面、环境变量和正式集成约束
- `openclaw-orchestration-flow.md`
- 面向 OpenClaw 编排层的正式运行手册
- 说明启动、轮询、读结果、恢复和人工介入的标准动作
- `openclaw-candidate-input-field-spec.md`
- 单篇 `OpenClawCandidateInput` 字段规范
- `openclaw-delivery-payload-spec.md`
- 批量 `OpenClawDeliveryPayload` 字段规范
- `article-candidate-daily-digest-schema.md`
- 对象分层设计说明
- 用于理解 `ArticleCandidateRecord` / `OpenClawCandidateInput` / `DailyDigest` 的关系
## 当前推荐阅读顺序
1. `openclaw-handoff.md`
2. `openclaw-orchestration-flow.md`
3. `openclaw-candidate-input-field-spec.md`
4. `openclaw-delivery-payload-spec.md`
5. `article-candidate-daily-digest-schema.md`
## 归档说明
`archive/` 下的文档保留历史决策、阶段总结和排障规划,但不再作为当前事实来源。
当前已归档:
- `archive/formalization-summary-2026-04-07.md`
- `archive/openclaw-daily-digest-refactor.md`
- `archive/digest-optimization-summary.md`
- `archive/p1-status-reconciliation-plan-2026-04-14.md`
+9
View File
@@ -0,0 +1,9 @@
# OpenClaw 历史归档
本目录只保留阶段性总结、设计演进记录和排障计划。
使用原则:
- 需要了解“为什么会这样设计”时再看
- 不要把这里的描述当成当前生产事实
- 当前正式口径以 `docs/openclaw/README.md`、`docs/openclaw/openclaw-handoff.md`、`docs/openclaw/openclaw-orchestration-flow.md` 为准
@@ -0,0 +1,367 @@
# Reader 日报链路 P1 状态收敛问题:规划与修复清单(2026-04-14)
## 背景
在 2026-04-14 的 reader 日报正式运行中,出现了以下现象:
- `openclaw-delivery-payload.json`、`digest-brief.json`、`run-report.json` 已真实落盘
- 但 `get_freshrss_pipeline_job_status` / `get_run_status` 仍可能显示:
- `running`
- `failed`
- 或 `current_stage=generate_summaries`
- `resume_run` 在这种状态下可能直接超时
这说明当前 reader 的**状态层(job/run-state)**与**产物层(artifacts/report)**之间没有稳定收敛。
---
## 本次确认的核心结论
### 1. job status 与 run status 是两套独立状态系统
- **job 层状态**:`src/summary_mcp/runtime/freshrss_pipeline_jobs.py`
- `start_freshrss_pipeline_job()`
- `run_freshrss_pipeline_job()`
- `get_freshrss_pipeline_job_status()`
- 状态文件位于:`outputs/freshrss/pipeline_jobs/<job_id>/run-state.json`
- 只有 4 个粗粒度 stage:
- `prepare_job`
- `load_input`
- `run_pipeline`
- `write_result`
- **run 层状态**:`src/summary_mcp/workflows/freshrss_pipeline.py`
- `run_freshrss_pipeline()`
- 由 `src/summary_mcp/runtime/query_service.py:get_run_status()` 查询
- 状态文件位于:`outputs/freshrss/rerun/<run_dir>/run-state.json`
- 包含 6 个细粒度 stage:
- `fetch_feed`
- `extract_articles`
- `generate_summaries`
- `apply_filters`
- `build_delivery_payload`
- `write_run_report`
**问题:** 两套状态没有统一收敛规则,用户可以同时看到两套不同口径的“当前进度”。
---
### 2. 查询层目前优先信 run-state,不会用 artifacts / run-report 纠偏
代码位置:`src/summary_mcp/runtime/query_service.py`
关键行为:
- `_resolve_run_record()` 只要发现 `run-state.json` 存在,就优先使用 `RunStore.load(...)`
- 即使 `run-report.json`、`delivery_payload`、`digest_brief` 已存在,也不会自动纠偏状态
**结果:**
- 一旦 `run-state.json` 因中断、超时、外层 SIGTERM 或写回未完成而停留在旧值
- `get_run_status()` 就会持续返回过期状态
- 造成“产物已完成,但状态仍显示 running/failed/卡在 summary”的错觉
---
### 3. `generate_summaries` 假卡住,本质上更像 stale state,不像真实业务卡住
代码位置:`src/summary_mcp/workflows/freshrss_pipeline.py`
从执行顺序看:
1. `start_stage(generate_summaries)`
2. summary 循环
3. `finish_stage(generate_summaries)`
4. `start_stage(apply_filters)`
5. `finish_stage(apply_filters)`
6. `start_stage(build_delivery_payload)`
7. 写 payload / digest brief
8. `finish_stage(build_delivery_payload)`
9. `start_stage(write_run_report)`
10. 写 run-report
11. `finish_stage(write_run_report)`
12. `finish_run(...)`
**判断:**
如果 payload / digest brief / run-report 都已经存在,那么“仍显示卡在 `generate_summaries`”更可能是:
- `run-state.json` 没来得及写回最终状态
- 或查询时读到了旧状态
而不是 summary 阶段真实没有跑过去。
---
### 4. `resume_run` 不是轻量恢复,而是同步继续跑工作流
代码位置:`src/summary_mcp/runtime/resume_service.py`
关键行为:
- `resume_run()` 会根据 `resume_from_stage` 直接继续执行:
- `_run_summary_stage(...)`
- `_run_filter_stage(...)`
- `_run_delivery_stage(...)`
- `_run_report_stage(...)`
这意味着它不是“修状态”的工具,而是“同步继续跑剩余工作流”的工具。
**问题:**
- 如果 stale state 把 `resume_from_stage` 定在 `generate_summaries`
- 那么 `resume_run` 会从一个过早阶段重新跑
- 在 MCP 包装层下非常容易超时
---
## 问题分类
### A. 真实 bug
1. **查询层过度信任 stale `run-state.json`**
- 文件:`src/summary_mcp/runtime/query_service.py`
- 影响:产物已完成但状态仍错误
2. **`resume_run` 过度依赖 stale `current_stage` / recovery 信息**
- 文件:`src/summary_mcp/runtime/resume_service.py`
- 影响:从过早阶段重跑,放大 timeout 风险
### B. 状态设计缺陷
3. **job 层与 run 层两套状态源没有统一收敛规则**
- 文件:`src/summary_mcp/runtime/freshrss_pipeline_jobs.py`
- 文件:`src/summary_mcp/runtime/query_service.py`
- 影响:用户看到两个互相打架的状态解释
4. **状态系统完全依赖显式写回,不会按产物反推修正**
- 文件:`src/summary_mcp/runtime/run_store.py`
- 影响:一旦中断,状态比产物更容易脏
### C. 调用层误判
5. **把 `resume_run` 当成轻量恢复接口使用**
- 实际上它更接近“同步恢复执行器”
- 影响:在长链路场景下超时是高概率事件
---
## 修复目标
## 当前落地状态(回填)
- [x] Phase 1 已落地:`get_run_status()` 会基于 `run-report.json` 与关键产物做终态收敛,并暴露 `status_source` / `state_conflict`
- [x] Phase 2 已落地第一阶段:`resume_run()` 会拒绝对已有终态 `run-report.json` 的 run 继续恢复
- [x] Phase 2 已继续增强:恢复起点现在会优先根据 artifacts 重算,而不是直接盲信 `run-state.recovery.resume_from_stage`
- [x] 新增 `inspect_resume_plan(run_id)` 作为恢复前置判定接口,避免调用方用 `resume_run` 探路
- [x] Phase 2 已补齐生产恢复 artifacts:正式 run 会稳定写出 `summary/summary-batch.json` 与 `candidates/candidate-batch.json`,`resume_run` / `inspect_resume_plan` 会优先使用它们,而不是依赖 debug per-item 文件
- [x] Phase 3 已落地:job 状态与结果读取会基于 linked run 做收敛,避免 outer job stale state 卡住编排
### 一级目标(必须达成)
1. 当 `run-report.json` / `delivery_payload` / `digest_brief` 已存在时,`get_run_status()` 不应继续盲目展示明显过期的 stage 状态;对调用方暴露的 `status` 必须直接收敛为可用终态,而不是只附加 hint
2. 当状态层与产物层冲突时,查询结果必须显式标注“状态冲突 / stale state”
3. `resume_run()` 在恢复前应优先基于现有 artifacts 判断真实可恢复起点,避免从过早阶段重跑
### 二级目标(建议达成)
4. job 层状态结果中增加对 linked run 的补充解释,避免“job running 但 run 产物已齐”这种情况毫无说明
5. 为后续编排层提供明确可消费的“状态可信度/冲突提示”字段
---
## 最小修复方案
### Phase 1|先修 run 查询层(优先级最高)
#### 目标
让 `get_run_status()` 至少能正确识别:
- run-state 是旧的
- 但关键产物已经齐了
#### 建议改动点
文件:`src/summary_mcp/runtime/query_service.py`
#### 建议动作
- [x] 在 `_resolve_run_record()` 或 `_build_status_response()` 中增加“关键产物存在性检查”
- `run-report.json`
- `candidates/openclaw-delivery-payload.json`
- `candidates/digest-brief.json`
- [x] 如果 `run-state.current_stage` 仍停留在早期阶段,但关键产物已齐:
- 不要继续原样输出为可信最终态
- 应直接把对外 `status` / `current_stage` / `recovery` 收敛成终态语义
- 同时新增解释字段,例如:
- `state_conflict: true`
- `state_conflict_reason: "run_state indicates generate_summaries but run-report.json already proves the workflow reached a terminal state"`
- `status_source: "run_report_reconciliation"`
- [x] 保留 `state_source=run_state`,但增加 `status_source` / `state_quality` / `state_conflict` 之类解释字段
#### 预期收益
- OpenClaw 继续按 `status` 分支时也不会卡住
- 第一时间减少“明明产物齐了却还像没跑完”的误判
- 不需要立刻动 workflow 主链路
---
### Phase 2|修 `resume_run` 的恢复起点判断
#### 目标
避免 stale state 让恢复逻辑从 `generate_summaries` 这类过早阶段重跑。
#### 建议改动点
文件:`src/summary_mcp/runtime/resume_service.py`
#### 建议动作
- [x] 在 `_resolve_resume_from_stage()` 之前/之后加入真实 artifacts 检查
- [x] 如果以下文件已存在:
- `openclaw-delivery-payload.json`
- `digest-brief.json`
- `run-report.json`
则不要再从 `generate_summaries` 或 `apply_filters` 起跑
- [x] 为 `resume_run()` 增加“恢复起点是基于 artifacts 重算还是基于 state 推断”的返回说明
- [x] 必要时增加更保守逻辑:
- `run-report.json` 已存在时,默认拒绝继续 resume,并提示“产物已完成,请先检查状态一致性”
- 补充:默认生产模式下,主链路会稳定写出 `summary-batch` / `candidate-batch`,恢复逻辑优先消费这两个 batch artifacts;若它们缺失或不稳定,才回退到更早的安全 stage 或直接拒绝恢复
- 补充:调用方可先走 `inspect_resume_plan`,只有 `recommended_action=resume` 时再调用 `resume_run`
#### 预期收益
- 降低无意义重跑和 timeout 风险
- 让 `resume_run` 更接近真正的恢复工具,而不是误重跑工具
---
### Phase 3|补 job/run 双状态解释层
#### 目标
让 `get_freshrss_pipeline_job_status()` 和 `get_run_status()` 的关系对调用方更可理解。
#### 建议改动点
文件:`src/summary_mcp/runtime/freshrss_pipeline_jobs.py`
#### 建议动作
- [x] 在 `get_freshrss_pipeline_job_status()` 中,读取 linked run 的关键产物存在性(轻量即可)
- [x] 若 job 仍显示 `run_pipeline`,但 linked run 已有 report/payload/digest 产物:
- 不仅增加解释字段,还应直接把 job 对外 `status` 收敛为终态,避免外层永远轮询
- 例如:
- `status_source: "linked_run_reconciliation"`
- `status_note: "linked run artifacts are complete; the job can be treated as completed"`
- [x] 若 `result.json` 缺失,但 linked run 已有 `run-report.json` 与 delivery 产物:
- `get_freshrss_pipeline_job_result()` 应能基于 linked run 产物合成最小结果,至少稳定返回 `run_id`
- [x] 明确文档:job status 是外层异步任务态,不等于内部 workflow 细粒度状态
#### 预期收益
- 减少“job running / run finished”口径冲突带来的误解
- 避免 OpenClaw 因 outer job stale state 卡死在轮询和 result 读取前
---
### Phase 4|把 `resume_run` 改成异步恢复 job
#### 目标
解决当前剩余的核心问题:`resume_run` 虽然恢复判定已经安全,但执行模型仍是同步 MCP 调用,长链路恢复时依然可能超时,导致 OpenClaw 编排层“看起来像又卡住了”。
#### 建议改动点
文件:
- `src/summary_mcp/runtime/resume_jobs.py`(新)
- `scripts/run_resume_job.py`(新)
- `src/summary_mcp/server.py`
- `src/summary_mcp/runtime/__init__.py`
- `src/summary_mcp/runtime/resume_service.py`
#### 建议动作
- [x] 新增最小异步恢复接口:
- `start_resume_job(run_id)`
- `get_resume_job_status(job_id)`
- `get_resume_job_result(job_id)`
- [x] job 目录固定落到:
- `outputs/freshrss/resume_jobs/<job_id>/`
- [x] 最少产物约定:
- `run-state.json`
- `input.json`
- `result.json`(成功时)
- `job-report.json`
- [x] `start_resume_job` 内部先调用 `inspect_resume_plan`
- 只有 `recommended_action=resume` 才允许真正启动
- `read_terminal_result` / `start_new_run` 要直接在 job 输入校验阶段返回,不进入执行器
- [x] 后台执行时复用现有 `_resume_freshrss_run(...)`
- 不重写恢复业务逻辑
- 只把同步入口拆成异步 job 外壳
- [x] `resume_run(run_id)` 保留,但降级为 debug / fallback
- 文档中明确:OpenClaw 编排默认应走 resume async job,而不是同步 `resume_run`
- [x] job result 里至少稳定返回:
- `run_id`
- `resume_from_stage`
- `status`
- `result_source`
- `delivery_output` / `report_output`(若存在)
#### 预期收益
- 彻底切掉恢复阶段的 MCP 同步超时风险
- 让 OpenClaw 对“启动恢复 / 轮询恢复 / 读取恢复结果”的控制面与主 pipeline async job 保持一致
- 把“恢复判定”与“恢复执行”分层,减少误调用和卡住错觉
---
## 不建议现在就做的事
- [ ] **不要先做自动 fallback 修状态**
- 例如:看到 artifacts 齐了就直接把 run-state 强行改成 success
- 原因:这会掩盖真正的状态写回问题
- [ ] **不要先大改 workflow 主链路**
- 当前更像查询层与恢复层的状态解释缺陷
- 先修读取与恢复判断,收益更大、风险更低
---
## 建议执行顺序
1. **先改 `query_service.py`**
- 让 `get_run_status()` 能暴露 stale state / artifact conflict
2. **再改 `resume_service.py`**
- 避免从错误阶段重跑
3. **最后看 `freshrss_pipeline_jobs.py`**
- 给 job status 加 linked run 补充说明
4. **收尾改 `resume async job`**
- 让恢复执行也走正式异步控制面,避免同步恢复再把编排卡住
---
## 验收标准
### 验收 1:状态冲突识别
构造一个场景:
- `run-state.json` 留在 `generate_summaries`
- 但 payload / digest brief / run-report 已存在
期望:
- `get_run_status()` 不再只回“卡在 generate_summaries”
- 会显式返回冲突提示字段
### 验收 2:恢复起点修正
构造一个场景:
- `run-state` 指向 `generate_summaries`
- 但 `delivery_payload` / `run-report` 已存在
期望:
- `resume_run()` 不应再从 summary 阶段重跑
- 至少应拒绝恢复并提示“产物已完成,优先检查状态一致性”
### 验收 3:job/run 双层说明
构造一个场景:
- job status 仍在 `run_pipeline`
- linked run 已有关键产物
期望:
- `get_freshrss_pipeline_job_status()` 能返回补充说明,不再只有生硬 running
### 验收 4:恢复执行不再阻塞编排
构造一个场景:
- run 可恢复
- 恢复点为 `generate_summaries` 或 `apply_filters`
- 恢复执行耗时超过单次 MCP 同步窗口
期望:
- OpenClaw 调用的是 `start_resume_job(...)`,而不是同步 `resume_run(...)`
- `get_resume_job_status(job_id)` 可稳定轮询到终态
- `get_resume_job_result(job_id)` 至少稳定返回 `run_id`、`resume_from_stage` 与最终产物引用
- 即使恢复失败,也能在 job-report / result 中看清失败点,而不是只表现为调用超时
---
## 备注
截至 2026-04-14,本文件中的 Phase 1 / 2 / 3 / 4 已完成主要落地;当前 `resume` 链路已经从“状态收敛 + 安全恢复点判定”进一步补齐到“正式异步恢复执行”。
+172 -263
View File
@@ -1,49 +1,54 @@
# OpenClaw Handoff
## Role
This file is the integration overview for OpenClaw maintainers.
Use it for:
- reader capability boundary
- production MCP entrypoints
- environment requirements
- integration rules and limitations
Do not use it as the step-by-step runbook.
For formal orchestration, read `docs/openclaw/openclaw-orchestration-flow.md`.
For field contracts, read:
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
- `docs/openclaw/openclaw-delivery-payload-spec.md`
Historical plans and incident documents live under `docs/openclaw/archive/`.
## Purpose
This repository provides a FreshRSS-first reading pipeline for OpenClaw:
reader is the upstream FreshRSS processing service for OpenClaw:
`FreshRSS unread items -> RSS content extraction -> LLM summary -> rule engine -> OpenClaw delivery payload`
OpenClaw should treat this repository as an MCP-backed upstream content processor.
This repository is responsible only for upstream reading-pipeline work:
reader is responsible for:
- FreshRSS pull
- content extraction
- LLM summary generation/validation
- LLM summary generation and validation
- rule-based filtering
- OpenClaw delivery payload generation
- selected-article summary capability based on existing extracted text
- run-state persistence, status lookup, result lookup, and minimal resume for the FreshRSS workflow
- run-state persistence and run/result lookup
- async resume control for the FreshRSS workflow
- async selected-article summary generation from existing extracted files
This repository should **not** take over downstream orchestration responsibilities that belong to OpenClaw / skills, such as:
reader is not responsible for:
- Hugo publishing
- chat reporting
- user confirmation handling
- IMA upload orchestration
## Production Entrypoint
## Production Surface
reader 当前正式工作流服务启动入口是 MCP tool:
Current MCP tool count: 21.
- `start_freshrss_pipeline_job`
OpenClaw 应先拿到 `job_id`,轮询 job 状态,再在成功后读取 `run_id` 作为正式后续句柄。
`run_freshrss_openclaw_pipeline` 仍保留,但定位是同步 debug / fallback 路径,而不是正式生产启动入口。
OpenClaw should treat the returned `run_id` from `get_freshrss_pipeline_job_result` as the only stable handle for follow-up reads. Do not hand-build `outputs/freshrss/rerun/...` paths in OpenClaw.
Job state is written under `outputs/freshrss/pipeline_jobs/<job_id>/` and will minimally contain `run-state.json`, `input.json`, `result.json` on success, and `job-report.json`.
## Supported MCP Tools
Current MCP tools: 17 total, including the FreshRSS workflow set plus async job tools for both the main pipeline and article-summary flow.
Workflow service tools:
Main daily workflow:
- `start_freshrss_pipeline_job`
- `get_freshrss_pipeline_job_status`
@@ -53,41 +58,139 @@ Workflow service tools:
- `list_run_artifacts`
- `get_delivery_payload`
- `get_run_report`
Resume workflow:
- `inspect_resume_plan`
- `start_resume_job`
- `get_resume_job_status`
- `get_resume_job_result`
- `resume_run`
Article-summary tools:
Selected-article summary workflow:
- `start_article_summary_job`
- `get_article_summary_job_status`
- `get_article_summary_job_result`
- `generate_article_summaries`
Single-step / debug tools:
Debug / single-step tools:
- `run_freshrss_openclaw_pipeline`(同步模式,仅适合 debug / fallback)
- `run_freshrss_openclaw_pipeline`
- `extract_url_content`
- `extract_item_content`
- `filter_summary_result`
- `generate_article_summaries`(同步模式,仅适合轻量调试)
## Recommended Selected-Article Flow
Production rules:
For OpenClaw selected-article follow-up, prefer the async job path:
- main production start path is `start_freshrss_pipeline_job`
- production resume path is `inspect_resume_plan -> start_resume_job -> get_resume_job_status -> get_resume_job_result`
- `run_freshrss_openclaw_pipeline` is sync debug / fallback only
- `resume_run` is sync debug / fallback only
- `generate_article_summaries` is sync debug / fallback only
1. Call `start_article_summary_job` with a real extracted file path plus a non-empty `selected_ids` list.
2. Poll `get_article_summary_job_status` until `status` becomes `success` or `failed`.
3. On success, call `get_article_summary_job_result` and continue downstream processing from `written_paths`.
4. Use `generate_article_summaries` only as a synchronous debug fallback, not as the default production path.
## Production Contract
Job state is written under `outputs/freshrss/article_summary_jobs/<job_id>/` and will minimally contain:
OpenClaw should treat the returned `run_id` from `get_freshrss_pipeline_job_result` as the only stable handle for follow-up reads.
OpenClaw should not hand-build these paths:
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
If filesystem access is needed for debugging, only consume paths returned by MCP:
- `output_dir`
- `artifact.path`
- `delivery_output`
- `report_output`
Top-level `status` is the only status field callers should branch on.
`status_source` and `state_conflict` are explanatory fields for reconciled status.
## Minimal Production Sequence
Daily workflow:
1. Call `start_freshrss_pipeline_job`
2. Poll `get_freshrss_pipeline_job_status`
3. On success, read `get_freshrss_pipeline_job_result`
4. Persist the returned `run_id`
5. Use `get_run_status`, `get_delivery_payload`, and `get_run_report` for follow-up reads
Resume workflow:
1. Call `inspect_resume_plan(run_id)`
2. Only if `can_resume=true` and `recommended_action=resume`, call `start_resume_job`
3. Poll `get_resume_job_status`
4. Read `get_resume_job_result`
Selected-article summary workflow:
1. Call `start_article_summary_job` with a real extracted file path and non-empty `selected_ids`
2. Poll `get_article_summary_job_status`
3. Read `get_article_summary_job_result`
## Capability Boundary
Formal workflow boundary:
- only workflow `freshrss_daily_digest`
- every current FreshRSS run writes `run-state.json`
- `get_run_status` / `list_runs` / `list_run_artifacts` can still infer basic state for older runs without `run-state.json`
- resume requires a valid `run-state.json`; inferred historical runs are not resumable
Resume boundary:
- resume in place on the original `run_id`
- supported resume points:
- `generate_summaries`
- `apply_filters`
- `build_delivery_payload`
- `write_run_report`
- unsupported resume points:
- `fetch_feed`
- `extract_articles`
- production resume prefers:
- `summary/summary-batch.json`
- `candidates/candidate-batch.json`
- if required artifacts are missing, recovery should return non-resumable instead of silently falling back
Selected-article summary boundary:
- uses existing extracted files as input
- should not re-fetch original URLs
## Output Expectations
Main daily pipeline core artifacts:
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
- `outputs/freshrss/rerun/<run_dir>/raw/freshrss.raw.json`
- `outputs/freshrss/rerun/<run_dir>/summary/summary-batch.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/candidate-batch.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/digest-brief.json`
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
- `outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json`
Async job state directories:
- main pipeline job: `outputs/freshrss/pipeline_jobs/<job_id>/`
- resume job: `outputs/freshrss/resume_jobs/<job_id>/`
- article-summary job: `outputs/freshrss/article_summary_jobs/<job_id>/`
Each job directory minimally contains:
- `run-state.json`
- `input.json`
- `result.json` on success
- `job-report.json`
## Required Environment Variables
## Environment And Startup
The MCP server process must have these variables available:
Required environment variables:
- `FRESHRSS_API_BASE_URL`
- `FRESHRSS_USERNAME`
@@ -96,54 +199,13 @@ The MCP server process must have these variables available:
- `LLM_API_KEY`
- `LLM_MODEL`
Example:
```powershell
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=osiman
set FRESHRSS_API_PASSWORD=your-api-password
set LLM_API_URL=https://api.deepseek.com
set LLM_API_KEY=your-llm-api-key
set LLM_MODEL=deepseek-chat
```
## Server Startup
Install dependencies:
Startup:
```bash
pip install -e .
```
Start the MCP server:
```bash
summary-mcp
```
## Recommended MCP Workflow
Recommended production path:
1. Call `start_freshrss_pipeline_job` and persist the returned `job_id`
2. Poll `get_freshrss_pipeline_job_status(job_id)` until `status` becomes `success` or `failed`
3. On success, call `get_freshrss_pipeline_job_result(job_id)` and persist the returned `run_id`
4. Use `get_run_status(run_id)` as the authoritative run-state read for status, stage, artifacts, and recovery
5. Use `list_runs(...)` when OpenClaw needs recent-run discovery or high-level inspection
6. Use `list_run_artifacts(run_id)` when OpenClaw needs to inspect what this run actually produced
7. Use `get_delivery_payload(run_id)` and `get_run_report(run_id)` as the formal result-reading APIs
8. Use `resume_run(run_id)` only when the run falls inside the minimal supported resume scope
OpenClaw should not directly derive or hardcode:
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
If filesystem access is needed for debugging, consume only paths returned by MCP such as `output_dir`, `artifact.path`, `delivery_output`, or `report_output`.
## Recommended MCP Call
Recommended production start call:
```json
@@ -157,204 +219,51 @@ Recommended production start call:
}
```
Recommended semantics:
## Data And Content Policy
- Use `mark_read=true` for normal production runs.
- Use `mark_read=false` only for debug, test, or validation runs.
- Keep `debug_artifacts=false` for routine production runs.
- Set `debug_artifacts=true` only when troubleshooting a bad batch.
- Treat the returned `job_id` as the startup handle, and the later `run_id` from `get_freshrss_pipeline_job_result` as the stable identifier for all follow-up run reads.
- If no real `openclaw-delivery-payload.json` was produced, OpenClaw should stop instead of generating a digest from placeholders or examples.
- Use `run_freshrss_openclaw_pipeline` only when a synchronous debug / fallback path is explicitly needed.
## Formal Capability Boundary
reader 当前正式 MCP workflow service 的边界如下:
- formal workflow: only `freshrss_daily_digest`
- run truth: every FreshRSS run writes `run-state.json`
- main production start path: `start_freshrss_pipeline_job` / `get_freshrss_pipeline_job_status` / `get_freshrss_pipeline_job_result`
- state query tools: `get_run_status`, `list_runs`, `list_run_artifacts`
- result read tools: `get_delivery_payload`, `get_run_report`
- `digest-brief.json` is generated and registered as an artifact, but there is no standalone `get_digest_brief` tool yet
- `run_freshrss_openclaw_pipeline` is still supported, but only as a synchronous debug / fallback path
- the FreshRSS daily workflow now has a minimal background job model backed by a detached runner process, not a full queue / worker system
- article summary now has a minimal asynchronous job model with `start_article_summary_job` / `get_article_summary_job_status` / `get_article_summary_job_result`
- `generate_article_summaries` is still supported, but it is a synchronous debug path and outside the formal `resume_run` scope
Historical compatibility note:
- `get_run_status` / `list_runs` / `list_run_artifacts` can still infer basic state for older run directories without `run-state.json`
- `resume_run` does **not** support those inferred historical runs; it requires a valid `run-state.json`
## What The Async Job Returns
Primary return fields from `get_freshrss_pipeline_job_result`:
- `job_id`
- `run_id`
- `output_dir`
- `raw_output`
- `delivery_output`
- `report_output`
- `digest_brief_output`
- `pulled_count`
- `delivered_count`
- `marked_read_count`
- `status_counts`
- `keyword_index`
Optional:
- `items`
- returned only when `include_item_reports=true`
`run_freshrss_openclaw_pipeline` still returns the same synchronous payload for debug / fallback use.
Follow-up structured reads should use MCP tools rather than re-reading these files directly.
## Minimal Output Files
By default the pipeline writes these core artifacts:
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
- `outputs/freshrss/rerun/<run_dir>/raw/freshrss.raw.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/digest-brief.json`
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
- `outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json` (one per item)
It also updates local runtime keyword data:
- `data/term_index/daily/YYYY-MM-DD.json`
- `data/term_index/term_stats.json`
Per-item extracted files live under `extracted/` and are always written.
If `debug_artifacts=true`, the pipeline additionally writes normalized items, summaries, filter decisions, candidate records, and candidate inputs.
The main pipeline does not emit a batch-level `freshrss.extracted.json` file by default.
## `resume_run` Minimal Scope
`resume_run` currently supports only the minimum resume contract:
- only runs with a valid `run-state.json`
- only workflow `freshrss_daily_digest`
- resume in place on the original `run_id`
- supported resume points: `generate_summaries`, `apply_filters`, `build_delivery_payload`, `write_run_report`
- unsupported resume points: `fetch_feed`, `extract_articles`
- if required artifacts are missing, the tool returns a non-resumable response instead of silently falling back to an earlier stage
Artifact expectations by resume point:
- `generate_summaries`: requires `raw/freshrss.raw.json` and `extracted/`
- `apply_filters`: requires the above plus per-item summary outputs
- `build_delivery_payload`: requires per-item candidate inputs consistent with filter-stage output
- `write_run_report`: requires `candidates/openclaw-delivery-payload.json`; if `mark_read=true`, raw input must still be present
## Payload Specs
OpenClaw payload field specs live here:
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
- `docs/openclaw/openclaw-delivery-payload-spec.md`
## Read-State Semantics
The pipeline reads from FreshRSS unread items by default.
If `mark_read=true`:
- items are marked as read only after the final `openclaw-delivery-payload.json` has been written successfully
- only successfully delivered items are marked as read
- failed or skipped items remain unread
## FreshRSS Content Policy
For FreshRSS items, the pipeline is RSS-first and does not re-crawl webpages.
Behavior:
FreshRSS processing is RSS-first:
- use `item.raw_content` first
- if missing, use `item.raw_summary`
- if neither contains usable content, skip the item
- do not fetch the original webpage again for FreshRSS items
This is intentional.
Read-state policy:
## Keyword Cleanup Governance
- items are marked read only after successful delivery payload write
- only successfully delivered items are marked read
This repository also includes a lightweight keyword-governance flow for downstream review.
Downstream boundary:
Current pieces:
- runtime keyword stats
- `data/term_index/daily/YYYY-MM-DD.json`
- `data/term_index/term_stats.json`
- governance config
- `configs/term_cleanup_policy.json`
- `configs/term_watchlist.json`
- `configs/term_change_log.json`
- review bundle builder
- `skills/keyword-cleanup-review/scripts/build_review_bundle.py`
- accepted-suggestion writer
- `scripts/apply_term_suggestions.py`
Current status:
- OpenClaw can read the keyword review bundle as maintenance input
- accepted suggestions still require explicit human confirmation
- the repository can write accepted watch / alias / stopword / interest-keyword changes after confirmation
- this governance flow is not yet wired into a periodic scheduler inside the repository
Boundary:
- keyword cleanup is a maintenance flow, not the production RSS ingestion path
- the repository does not auto-apply cleanup suggestions without confirmation
- current keyword stats are built from the delivered candidate payload, not yet from a final `DailyDigest`
## Known Limitations
- Some sources put only partial content in RSS; those items may be skipped if RSS content is insufficient.
- WeChat articles often block direct crawling, but this pipeline now avoids that path for FreshRSS items and uses RSS-provided content when available.
- Rule behavior is still conservative in some cases; many items may land in `review` depending on current rules.
- Paywall heuristics may produce false positives for some Chinese text patterns.
- Keyword cleanup governance is usable now, but periodic scheduling and before/after evaluation are not implemented yet.
## Files OpenClaw Should Read First
Recommended reading order for a new maintainer:
1. `README.md`
2. `docs/openclaw/openclaw-handoff.md`
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
5. `docs/design/daily-keyword-index-design.md`
6. `skills/keyword-cleanup-review/SKILL.md`
7. `docs/current/context-reset-brief.md`
## Downstream Boundary Rules
For the daily-digest workflow:
- the digest should go to Hugo and chat reporting, not directly into IMA
- the full daily digest should **not** be uploaded to IMA
- the daily digest goes to Hugo and chat reporting
- the full daily digest should not be uploaded to IMA
- only explicitly user-selected article summaries should be uploaded to IMA
- selected-article summaries should be generated from existing extracted text, not by re-fetching original URLs
## Current Recommendation
## Related Maintenance Flow
For integration handoff, the repository is usable now.
Keyword cleanup exists as a separate maintenance flow, not the main RSS ingestion path.
The minimum you need to give OpenClaw is:
- the repository code
- the MCP server startup command
- the required environment variables in the target environment
- the instruction to call `run_freshrss_openclaw_pipeline`
- the rule that follow-up state/result reads must go through MCP tools first, not handwritten filesystem paths
If OpenClaw will also participate in keyword-governance review, additionally point it to:
Relevant files:
- `docs/design/daily-keyword-index-design.md`
- `skills/keyword-cleanup-review/SKILL.md`
- `scripts/apply_term_suggestions.py`
## Known Limitations
- some sources expose only partial RSS content; those items may be skipped
- rule behavior is still conservative; many items may land in `review`
- paywall heuristics may still produce false positives on some Chinese text
- keyword cleanup governance is usable but not yet wired to periodic scheduling
## Read First
Recommended reading order for a new maintainer:
1. `README.md`
2. `docs/openclaw/README.md`
3. `docs/openclaw/openclaw-handoff.md`
4. `docs/openclaw/openclaw-orchestration-flow.md`
5. `docs/openclaw/openclaw-candidate-input-field-spec.md`
6. `docs/openclaw/openclaw-delivery-payload-spec.md`
7. `docs/current/context-reset-brief.md`
+167 -134
View File
@@ -2,79 +2,86 @@
## 1. 文档目的
本文档定义 OpenClaw 在正式环境中如何调用 reader 作为上游 MCP workflow service。
本文档只回答一个问题:OpenClaw 在正式环境里应该如何编排 reader。
目标不是描述 reader 内部实现,而是明确 OpenClaw 的编排动作:
这里不重复介绍 reader 内部实现,只定义正式控制面:
- 什么时候启动新 run
- 什么时候查询状态
- 什么时候读取结果
- 什么时候尝试恢复
- 什么时候直接新开 run
- 什么时候需要人工介入
- 如何启动日报
- 如何轮询 job
- 如何读取 run 结果
- 如何判断是否恢复
- 如何走异步恢复
- 什么时候直接新开 run 或人工介入
本文档基于 reader 当前**已真实落地**的能力编写,不描述尚未实现的未来接口。
## 2. 当前正式入口
---
### 2.1 新 run
## 2. 当前 reader 已正式支持的 MCP 能力
正式生产入口:
当前可用能力:
- `start_freshrss_pipeline_job`
- `get_freshrss_pipeline_job_status`
- `get_freshrss_pipeline_job_result`
同步入口:
- `run_freshrss_openclaw_pipeline`
同步入口只保留给 debug / fallback,不再是正式编排默认路径。
### 2.2 run 级读取
正式 run 级读取接口:
- `get_run_status`
- `list_runs`
- `list_run_artifacts`
- `get_delivery_payload`
- `get_run_report`
### 2.3 恢复
正式恢复入口:
- `inspect_resume_plan`
- `start_resume_job`
- `get_resume_job_status`
- `get_resume_job_result`
同步恢复入口:
- `resume_run`
其中:
- `run_freshrss_openclaw_pipeline` 是当前正式启动入口
- `get_run_status` / `list_runs` / `list_run_artifacts` 用于观测
- `get_delivery_payload` / `get_run_report` 用于读取正式结果
- `resume_run` 用于最小恢复能力
---
`resume_run` 只保留给 debug / fallback。
## 3. 编排基本原则
### 3.1 OpenClaw 不再手拼路径
### 3.1 OpenClaw 不手拼路径
OpenClaw 不应再自己拼 reader 输出路径来判断运行状态或读取核心结果。
OpenClaw 不应自己推导这些路径:
优先使用 MCP:
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
- 查状态 → `get_run_status`
- 读 payload → `get_delivery_payload`
- 读 report → `get_run_report`
- 做恢复 → `resume_run`
需要路径时,只消费 MCP 返回值:
只有在排障/人工核查时,才回退到直接看 reader run 目录。
- `output_dir`
- `artifact.path`
- `delivery_output`
- `report_output`
### 3.2 reader 是上游 workflow engine
### 3.2 顶层 `status` 才是分支依据
reader 负责:
`get_run_status` 和 job status 接口都可能做状态收敛。
- FreshRSS 拉取
- 内容提取
- 摘要
- 过滤
- payload 生成
- run 状态记录
- 最小恢复
因此:
OpenClaw 负责:
- 优先使用顶层 `status`
- `status_source` 用来解释状态来自原始 state 还是收敛结果
- `state_conflict=true` 说明底层状态文件已经落后于真实产物
- 触发执行
- 轮询状态
- 读取结果
- 生成 digest markdown
- Hugo 发布
- 聊天汇报
- 用户确认精选
- IMA 编排
不要再拿旧的 `raw_status`、`raw_current_stage` 或早期阶段名重新做分支。
### 3.3 默认生产语义
@@ -82,19 +89,18 @@ OpenClaw 负责:
- `mark_read=true`
- `debug_artifacts=false`
- 只在 debug/test/validation 时显式放宽
---
只有 debug / test / validation 时才放宽。
## 4. 标准 Happy Path
### Step 1: 启动新 run
### Step 1: 启动新 job
调用:
- `run_freshrss_openclaw_pipeline`
- `start_freshrss_pipeline_job`
推荐参数示例:
推荐参数:
```json
{
@@ -107,131 +113,156 @@ OpenClaw 负责:
}
```
期望:
预期:
- 获得 `run_id`
- 获得 `output_dir`
- 获得初始结果摘要
- 立即返回 `job_id`
- 后续由 OpenClaw 轮询 job,而不是同步等待整条流水线
如果启动阶段直接抛错:
### Step 2: 轮询 job
- 直接判为启动失败
- 不进入后续查询
调用:
### Step 2: 查询运行状态
- `get_freshrss_pipeline_job_status(job_id=...)`
根据返回:
- `status=running`:继续轮询
- `status=success`:读取 job result
- `status=failed`:进入失败处理
额外规则:
- 如果 `status_source=linked_run_reconciliation`,说明 outer job state 已落后,但 linked run 已经给出可用终态
- 如果 `status_source=stale_job_state_timeout`,把它当成终态失败,不要继续无限轮询
### Step 3: 读取 job result
调用:
- `get_freshrss_pipeline_job_result(job_id=...)`
预期读取:
- `run_id`
- `output_dir`
- `delivery_output`
- `report_output`
从这一刻开始,`run_id` 是正式的稳定句柄。
### Step 4: 读取 run 级状态与结果
调用:
- `get_run_status(run_id=...)`
根据返回:
- `status=running` → 继续轮询
- `status=success` → 进入结果读取
- `status=failed` → 进入失败处理
- `status=partial` → 视为未完成,优先看 `recovery` 和当前阶段
### Step 3: 读取正式结果
成功后读取:
- `get_delivery_payload(run_id=...)`
- `get_run_report(run_id=...)`
后续 OpenClaw 编排应以这两个接口为正式结果源,而不是自己拼路径读取 JSON。
根据 `get_run_status`:
### Step 4: 进入下游编排
- `status=running`:继续观察
- `status=success`:继续下游 digest / 发布 / 汇报
- `status=failed`:进入恢复或重跑决策
- `status=partial`:优先检查 report、artifacts 和 recovery
OpenClaw 在拿到正式 payload / report 后,继续执行:
如果 `status_source=run_report_reconciliation`,说明 `run-state.json` 已经过期,但 reader 已经根据终态产物收敛出有效状态。
- public/internal digest 生成
- Hugo 发布
- 聊天汇报
- 用户确认精选
- IMA 沉淀
如果 `status_source=stale_run_state_timeout`,说明 reader 认为该 run 长时间未收敛且没有终态产物,应按失败处理。
---
## 5. 恢复决策
## 5. 状态 → 动作映射
### 5.1 先看预检,不要直接恢复
| reader 状态 | OpenClaw 动作 |
|---|---|
| `running` | 继续轮询 `get_run_status` |
| `success` | 读取 `get_delivery_payload` 和 `get_run_report` |
| `failed` 且 `recovery.resumable=true` | 评估是否调用 `resume_run` |
| `failed` 且 `recovery.resumable=false` | 直接判失败,通常新开 run 或人工介入 |
| `partial` | 先读状态详情和 recovery,再决定继续等 / 恢复 / 人工介入 |
恢复前固定动作:
---
- 先调用 `inspect_resume_plan(run_id)`
## 6. 失败处理与恢复决策
只在以下条件同时成立时才启动恢复:
### 6.1 什么时候优先尝试 `resume_run`
- `can_resume=true`
- `recommended_action=resume`
满足以下条件时,优先考虑恢复而不是新开 run:
重点字段:
- `get_run_status` 返回 `failed`
- `recovery.resumable=true`
- 当前 run 对应的是 freshrss workflow
- 当前失败点在 reader 第一版支持的恢复范围内
- `requested_resume_from_stage`
- `resume_from_stage`
- `resume_decision_source`
- `artifact_resume_from_stage`
- `artifact_snapshot`
### 6.2 `resume_run` 当前支持范围
### 5.2 正式恢复路径
当前最小实现仅支持:
正式恢复控制面:
- 仅对带 `run-state.json` 的 freshrss run
- 仅从最近可恢复点继续
- 支持的恢复点:
1. `start_resume_job(run_id)`
2. `get_resume_job_status(job_id)`
3. `get_resume_job_result(job_id)`
不要再把同步 `resume_run(run_id)` 当成正式恢复入口。
### 5.3 当前支持范围
当前只支持:
- 带有效 `run-state.json` 的 `freshrss_daily_digest` run
- 从以下阶段恢复:
- `generate_summaries`
- `apply_filters`
- `build_delivery_payload`
- `write_run_report`
明确不支持:
当前不支持:
- `fetch_feed`
- `extract_articles`
### 6.3 什么时候不要恢复,直接新开 run
正式生产恢复优先依赖:
以下情况不建议 `resume_run`:
- `summary/summary-batch.json`
- `candidates/candidate-batch.json`
- `recovery.resumable=false`
- run 没有 `run-state.json`
- 失败点是 `fetch_feed` 或 `extract_articles`
- 恢复所需关键产物缺失
- 恢复点语义不明确或结果存在明显漂移风险
### 5.4 什么时候不要恢复
这时更合理的动作通常是:
以下情况直接新开 run 更合理:
- 直接新开 run
- 或人工介入排查
- `recommended_action=start_new_run`
- `recommended_action=read_terminal_result`
- 没有有效 `run-state.json`
- 恢复所需关键 artifacts 缺失
- 连续恢复失败
### 6.4 什么时候需要人工介入
## 6. 状态到动作映射
出现以下任一情况时,建议人工介入:
| 接口 | 状态 | OpenClaw 动作 |
| --- | --- | --- |
| `get_freshrss_pipeline_job_status` | `running` | 继续轮询 job |
| `get_freshrss_pipeline_job_status` | `success` | 读取 `get_freshrss_pipeline_job_result` |
| `get_freshrss_pipeline_job_status` | `failed` | 结束本次 job,必要时读 linked run |
| `get_run_status` | `running` | 继续观察 run |
| `get_run_status` | `success` | 读取 `get_delivery_payload` / `get_run_report` |
| `get_run_status` | `failed` | 先看 `inspect_resume_plan` |
| `inspect_resume_plan` | `recommended_action=resume` | 启动 `start_resume_job` |
| `inspect_resume_plan` | `recommended_action=read_terminal_result` | 直接读 run 结果,不恢复 |
| `inspect_resume_plan` | `recommended_action=start_new_run` | 新开 run 或人工介入 |
## 7. 人工介入条件
出现以下任一情况时,建议不要自动编排:
- 连续恢复失败
- `get_run_status` 与实际产物明显不一致
- payload/report 结构不符合预期
- 恢复依赖的关键文件缺失且原因不明
- FreshRSS / LLM / 外部环境异常
- payload / report 结构不符合预期
- `get_run_status` 与实际产物长期明显冲突
- FreshRSS、LLM 或外部依赖异常
- 恢复判定结果和编排预期不一致
---
## 8. 结论
## 7. 读取结果的标准动作
当前 OpenClaw 的正式调用方式已经收口为两条异步控制面:
### 7.1 `get_delivery_payload`
- 主日报:`start_freshrss_pipeline_job -> poll -> get result -> run reads`
- 恢复:`inspect_resume_plan -> start_resume_job -> poll -> get result`
用途:
- 获取正式交付给 OpenClaw 的 payload
- 后续 digest 生成应以该返回为准
OpenClaw 应做:
- 读取后直接进入 digest 生成
- 不再自己拼 `candidates/openclaw-delivery-payload.json`
同步 `run_freshrss_openclaw_pipeline` 和 `resume_run` 仅用于 debug / fallback,不应再作为默认正式编排路径。
### 7.2 `get_run_report`
@@ -263,7 +294,9 @@ OpenClaw 应做:
1. 调 `get_run_status`
2. 若 `failed && recovery.resumable=true`:
- 调 `resume_run`
- 调 `start_resume_job`
- 轮询 `get_resume_job_status`
- 读取 `get_resume_job_result`
3. 恢复后再次:
- 调 `get_run_status`
- 若成功,再读 payload / report
@@ -279,7 +312,7 @@ OpenClaw 应做:
- 让 OpenClaw 直接长时间 `exec` reader CLI 作为主要生产入口
- 让 OpenClaw 自己拼 reader 输出路径来判断成功/失败
- 让 OpenClaw 自己读取 `outputs/.../*.json` 作为正式结果源
- 在未确认 `resume_run` 支持范围外的失败点上强行恢复
- 在未确认恢复支持范围外的失败点上强行恢复
CLI 现在的定位是:
@@ -293,7 +326,7 @@ CLI 现在的定位是:
## 10. 当前已知局限
- `resume_run` 仍是最小实现,不支持任意 stage 任意重入
- 恢复能力仍是最小实现,不支持任意 stage 任意重入
- 历史无 `run-state.json` 的 run 不支持正式恢复
- 极旧 run 的结果读取仍可能依赖保守目录扫描
- `write_run_report` 若涉及重新 `mark_read`,仍依赖 FreshRSS 环境和可用凭据
@@ -304,4 +337,4 @@ CLI 现在的定位是:
OpenClaw 当前应把 reader 当作正式 MCP workflow service 使用:
**启动用 `run_freshrss_openclaw_pipeline`,观测用 `get_run_status`,结果读取用 `get_delivery_payload` / `get_run_report`,恢复仅在 `resume_run` 最小支持范围内启用;不要再把 reader 当成长 CLI 任务和路径拼接仓库来驱动。**
**启动用 `start_freshrss_pipeline_job`,观测用 `get_run_status`,结果读取用 `get_delivery_payload` / `get_run_report`,恢复默认用 `inspect_resume_plan` + `start_resume_job`,不要再把 reader 当成长 CLI 任务和路径拼接仓库来驱动。**