feat: add async resume jobs and doc navigation
This commit is contained in:
@@ -1,6 +1,6 @@
|
|||||||
# Reader MCP Workflow Service
|
# Reader MCP Workflow Service
|
||||||
|
|
||||||
reader 当前已经收口为面向 OpenClaw 的 MCP workflow service。正式能力边界以 FreshRSS 日报工作流为准:启动 run、写入 `run-state.json`、查询运行状态、读取结构化结果,以及最小可用的 `resume_run`。CLI 仍保留,但定位为 debug / fallback,而不是正式集成入口。
|
reader 当前已经收口为面向 OpenClaw 的 MCP workflow service。正式能力边界以 FreshRSS 日报工作流为准:启动 run、写入 `run-state.json`、查询运行状态、读取结构化结果,以及异步恢复 job。CLI 与同步入口仍保留,但定位为 debug / fallback,而不是正式集成入口。
|
||||||
|
|
||||||
## 运行
|
## 运行
|
||||||
|
|
||||||
@@ -9,62 +9,54 @@ pip install -e .
|
|||||||
summary-mcp
|
summary-mcp
|
||||||
```
|
```
|
||||||
|
|
||||||
服务当前暴露 17 个工具。
|
服务当前暴露 21 个工具。
|
||||||
|
|
||||||
正式 workflow service 相关工具:
|
正式集成摘要:
|
||||||
|
|
||||||
- `start_freshrss_pipeline_job`
|
- 主日报正式入口:`start_freshrss_pipeline_job`
|
||||||
- `get_freshrss_pipeline_job_status`
|
- 主日报正式读取:`get_run_status`、`get_delivery_payload`、`get_run_report`
|
||||||
- `get_freshrss_pipeline_job_result`
|
- 恢复正式入口:`inspect_resume_plan`、`start_resume_job`、`get_resume_job_status`、`get_resume_job_result`
|
||||||
- `get_run_status`
|
- 单篇总结正式入口:`start_article_summary_job`、`get_article_summary_job_status`、`get_article_summary_job_result`
|
||||||
- `list_runs`
|
- `run_freshrss_openclaw_pipeline`、`resume_run`、`generate_article_summaries` 仅用于同步 debug / fallback
|
||||||
- `list_run_artifacts`
|
|
||||||
- `get_delivery_payload`
|
|
||||||
- `get_run_report`
|
|
||||||
- `resume_run`
|
|
||||||
|
|
||||||
单步处理 / 调试相关工具:
|
## 文档入口
|
||||||
|
|
||||||
- `run_freshrss_openclaw_pipeline`(同步 debug / fallback)
|
如果你在做 OpenClaw 集成,不要只看这个 README,优先看:
|
||||||
- `extract_url_content`
|
|
||||||
- `extract_item_content`
|
- `docs/openclaw/README.md`
|
||||||
- `filter_summary_result`
|
- `docs/openclaw/openclaw-handoff.md`
|
||||||
- `generate_article_summaries`(同步模式)
|
- `docs/openclaw/openclaw-orchestration-flow.md`
|
||||||
- `start_article_summary_job`
|
|
||||||
- `get_article_summary_job_status`
|
字段契约见:
|
||||||
- `get_article_summary_job_result`
|
|
||||||
|
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||||
|
- `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||||
|
|
||||||
|
文档总索引见:
|
||||||
|
|
||||||
|
- `docs/README.md`
|
||||||
|
- `docs/current/context-reset-brief.md`
|
||||||
|
- `docs/design/README.md`
|
||||||
|
- `plans/README.md`
|
||||||
|
|
||||||
## 正式能力边界
|
## 正式能力边界
|
||||||
|
|
||||||
- 当前正式 workflow 只有 `freshrss_daily_digest`
|
- 当前正式 workflow 只有 `freshrss_daily_digest`
|
||||||
|
- 当前生产编排默认走异步 job,而不是同步 MCP / CLI
|
||||||
- 每次 FreshRSS 主流水线 run 都会在 `outputs/freshrss/rerun/<run_dir>/run-state.json` 落地运行真相
|
- 每次 FreshRSS 主流水线 run 都会在 `outputs/freshrss/rerun/<run_dir>/run-state.json` 落地运行真相
|
||||||
- FreshRSS 主日报的正式生产启动路径已切到最小异步 job:`start_freshrss_pipeline_job` -> `get_freshrss_pipeline_job_status` -> `get_freshrss_pipeline_job_result`
|
- OpenClaw 正式读取结果应优先使用 MCP 返回的 `run_id`、`output_dir`、`delivery_output`、`report_output`
|
||||||
- OpenClaw 正式读取结果应优先使用 `get_delivery_payload` 与 `get_run_report`,而不是自己拼输出目录路径
|
- 正式恢复只支持带有效 `run-state.json` 的当前 run,不处理历史推断 run
|
||||||
- `digest-brief.json` 当前会随主流水线产出,但还没有独立的 MCP 读取工具;如需定位它,应通过 `list_run_artifacts` 或 `get_run_report` 返回的信息发现
|
- 正式生产恢复依赖 `summary/summary-batch.json` 与 `candidates/candidate-batch.json`
|
||||||
- `run_freshrss_openclaw_pipeline` 仍保留,但定位已降级为同步 debug / fallback 路径,不再是 OpenClaw 的默认生产启动入口
|
|
||||||
- 单篇总结已补上最小异步 job 形态:`start_article_summary_job` / `get_article_summary_job_status` / `get_article_summary_job_result`
|
|
||||||
- `generate_article_summaries` 仍保留,但定位是同步 debug 路径,而不是 OpenClaw 的正式生产集成入口
|
|
||||||
|
|
||||||
## OpenClaw 推荐调用路径
|
## OpenClaw 最短调用路径
|
||||||
|
|
||||||
1. 调用 `start_freshrss_pipeline_job` 启动正式日报 job,并保存返回的 `job_id`
|
1. 调 `start_freshrss_pipeline_job`
|
||||||
2. 轮询 `get_freshrss_pipeline_job_status(job_id)`,直到 `status` 变成 `success` 或 `failed`
|
2. 轮询 `get_freshrss_pipeline_job_status`
|
||||||
3. 成功后调用 `get_freshrss_pipeline_job_result(job_id)` 读取 `run_id` 与关键产物路径
|
3. 成功后读 `get_freshrss_pipeline_job_result`,拿 `run_id`
|
||||||
4. 后续所有 run 级状态判断都基于 `get_run_status(run_id)` 或 `list_runs(...)`
|
4. 用 `get_run_status`、`get_delivery_payload`、`get_run_report` 做后续读取
|
||||||
5. 需要看产物列表时用 `list_run_artifacts(run_id)`,不要在 OpenClaw 里硬编码 `outputs/freshrss/rerun/...`
|
5. 如需恢复,先调 `inspect_resume_plan`,只有 `recommended_action=resume` 才走 `start_resume_job`
|
||||||
6. 需要消费正式结果时优先用 `get_delivery_payload(run_id)` 与 `get_run_report(run_id)`
|
|
||||||
7. 仅当 `resume_run` 的最小恢复范围满足时,才对失败 run 调用 `resume_run(run_id)`;否则应重启一个新 run
|
|
||||||
|
|
||||||
主日报 async job 的自身状态目录固定在 `outputs/freshrss/pipeline_jobs/<job_id>/`,最少包含 `run-state.json`、`input.json`、`result.json`(成功时)和 `job-report.json`。
|
更完整的状态分支、恢复策略和人工介入条件见 `docs/openclaw/openclaw-orchestration-flow.md`。
|
||||||
|
|
||||||
## `resume_run` 当前最小范围
|
|
||||||
|
|
||||||
- 只支持带有效 `run-state.json` 的 run
|
|
||||||
- 只支持 workflow `freshrss_daily_digest`
|
|
||||||
- 恢复时继续沿用原 `run_id`,不会新建 retry run
|
|
||||||
- 当前支持的恢复起点只有:`generate_summaries`、`apply_filters`、`build_delivery_payload`、`write_run_report`
|
|
||||||
- 当前明确不支持从 `fetch_feed`、`extract_articles` 恢复;这类失败应新开 run
|
|
||||||
- 恢复前会校验关键中间产物是否齐备,缺失时直接返回不可恢复,而不会自动回退到更早 stage
|
|
||||||
|
|
||||||
## 单篇文章总结后处理(可选使用独立 LLM)
|
## 单篇文章总结后处理(可选使用独立 LLM)
|
||||||
|
|
||||||
@@ -83,10 +75,10 @@ summary-mcp
|
|||||||
|
|
||||||
相关能力:
|
相关能力:
|
||||||
|
|
||||||
- `start_article_summary_job` / `get_article_summary_job_status` / `get_article_summary_job_result`(正式推荐的最小异步 job 路径)
|
- 正式路径:`start_article_summary_job` / `get_article_summary_job_status` / `get_article_summary_job_result`
|
||||||
- `generate_article_summaries` MCP 工具(同步 debug 路径)
|
- 同步 debug:`generate_article_summaries`
|
||||||
- `scripts/run_article_summaries.py` CLI 辅助脚本
|
- CLI:`scripts/run_article_summaries.py`
|
||||||
- `scripts/run_article_summary_job.py` 后台 runner 入口
|
- 后台 runner:`scripts/run_article_summary_job.py`
|
||||||
|
|
||||||
## 校验 LLM 摘要结果
|
## 校验 LLM 摘要结果
|
||||||
|
|
||||||
|
|||||||
@@ -12,11 +12,16 @@
|
|||||||
|
|
||||||
开始编码前必须阅读:
|
开始编码前必须阅读:
|
||||||
|
|
||||||
1. `plans/reader-mcp-architecture-design.md`
|
1. `README.md`
|
||||||
2. `plans/reader-mcp-implementation-plan.md`
|
2. `docs/README.md`
|
||||||
3. 本文件
|
3. `docs/openclaw/README.md`
|
||||||
4. `plans/issues/2026-04-06-reader-digest-sigterm.md`
|
4. `docs/openclaw/openclaw-handoff.md`
|
||||||
5. `docs/openclaw/openclaw-handoff.md`
|
5. `docs/openclaw/openclaw-orchestration-flow.md`
|
||||||
|
6. `plans/README.md`
|
||||||
|
7. `plans/reader-mcp-architecture-design.md`
|
||||||
|
8. `plans/reader-mcp-implementation-plan.md`
|
||||||
|
9. 本文件
|
||||||
|
10. `plans/issues/2026-04-06-reader-digest-sigterm.md`
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -184,6 +189,43 @@
|
|||||||
- 2026-04-07:已完成 `resume_run` minimal design 与现有 runtime/workflow/server 代码对齐分析,开始实现最小恢复链路。
|
- 2026-04-07:已完成 `resume_run` minimal design 与现有 runtime/workflow/server 代码对齐分析,开始实现最小恢复链路。
|
||||||
- 2026-04-07:已完成 `resume_run` 最小实现编码,新增 runtime 恢复服务并接入 MCP server;当前进入设计对齐与本地自检。
|
- 2026-04-07:已完成 `resume_run` 最小实现编码,新增 runtime 恢复服务并接入 MCP server;当前进入设计对齐与本地自检。
|
||||||
- 2026-04-07:已完成 `resume_run` 架构对齐与本地自检;已验证 `write_run_report` 可恢复,且 `extract_articles` 会被明确拒绝恢复。
|
- 2026-04-07:已完成 `resume_run` 架构对齐与本地自检;已验证 `write_run_report` 可恢复,且 `extract_articles` 会被明确拒绝恢复。
|
||||||
|
- 2026-04-14:已补 `inspect_resume_plan`、artifact-first 恢复判定,以及生产模式下稳定 `summary-batch` / `candidate-batch` artifacts;当前 `resume` 的剩余主问题不再是恢复点判断,而是同步执行模型仍可能让 OpenClaw 恢复阶段超时。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### [DONE][P1] 把 `resume_run` 升级为最小真异步 job
|
||||||
|
|
||||||
|
目标:
|
||||||
|
- 解决 `resume_run` 在 OpenClaw → MCP 同步链路里仍可能超时的问题
|
||||||
|
- 让恢复也具备“启动 / 轮询 / 读取结果”的正式控制面
|
||||||
|
|
||||||
|
要求:
|
||||||
|
- 新增最小异步接口:
|
||||||
|
- `start_resume_job`
|
||||||
|
- `get_resume_job_status`
|
||||||
|
- `get_resume_job_result`
|
||||||
|
- 状态目录固定落到:
|
||||||
|
- `outputs/freshrss/resume_jobs/<job_id>/`
|
||||||
|
- 至少包含:
|
||||||
|
- `run-state.json`
|
||||||
|
- `input.json`
|
||||||
|
- `result.json`(成功时)
|
||||||
|
- `job-report.json`
|
||||||
|
- 启动前必须先走 `inspect_resume_plan`
|
||||||
|
- 业务执行继续复用现有 `resume_service`,不要重写恢复主逻辑
|
||||||
|
- `resume_run` 保留为同步 debug / fallback 路径,但不再作为 OpenClaw 的默认恢复入口
|
||||||
|
|
||||||
|
完成标准:
|
||||||
|
- 可恢复 run 上,`start_resume_job` 能成功返回 `job_id`
|
||||||
|
- `get_resume_job_status` 能稳定反映恢复 job 生命周期
|
||||||
|
- `get_resume_job_result` 能稳定返回 `run_id`、`resume_from_stage`、最终状态与关键产物路径
|
||||||
|
- 恢复耗时超过单次 MCP 同步窗口时,OpenClaw 仍不会因为同步调用挂住
|
||||||
|
|
||||||
|
进展备注:
|
||||||
|
- 2026-04-14:已落地 `src/summary_mcp/runtime/resume_jobs.py` 与 `scripts/run_resume_job.py`,新增 `start_resume_job` / `get_resume_job_status` / `get_resume_job_result`
|
||||||
|
- 2026-04-14:启动前会先走 `inspect_resume_plan`;不可恢复 run 会在 job 输入校验阶段直接失败,不进入后台恢复执行
|
||||||
|
- 2026-04-14:后台执行复用现有 `_resume_freshrss_run(...)`,没有重写恢复主逻辑
|
||||||
|
- 2026-04-14:已完成本地 synthetic 验证:`write_run_report` 恢复可通过 `start -> poll -> result` 闭环成功收敛
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -304,6 +346,7 @@
|
|||||||
- 2026-04-07:新增架构设计文档 `plans/reader-mcp-architecture-design.md`
|
- 2026-04-07:新增架构设计文档 `plans/reader-mcp-architecture-design.md`
|
||||||
- 2026-04-07:新增实施计划文档 `plans/reader-mcp-implementation-plan.md`
|
- 2026-04-07:新增实施计划文档 `plans/reader-mcp-implementation-plan.md`
|
||||||
- 2026-04-07:完成 `freshrss` pipeline 的 run-state 基础设施,新增 `runtime` 包并覆盖关键 stages 状态持久化。
|
- 2026-04-07:完成 `freshrss` pipeline 的 run-state 基础设施,新增 `runtime` 包并覆盖关键 stages 状态持久化。
|
||||||
|
- 2026-04-14:已补 OpenClaw 文档导航、历史归档、design/notes/plans 导航,并统一当前正式口径为 async job 编排入口。
|
||||||
|
|
||||||
### 风险提醒
|
### 风险提醒
|
||||||
|
|
||||||
|
|||||||
+51
-39
@@ -1,6 +1,25 @@
|
|||||||
# 文档索引
|
# 文档索引
|
||||||
|
|
||||||
## 当前目录结构
|
## 当前最短阅读路径
|
||||||
|
|
||||||
|
1. `README.md`
|
||||||
|
- 仓库入口与常用脚本
|
||||||
|
2. `docs/current/context-reset-brief.md`
|
||||||
|
- 当前状态的最短摘要
|
||||||
|
3. `docs/openclaw/README.md`
|
||||||
|
- OpenClaw 集成文档导航
|
||||||
|
4. `docs/openclaw/openclaw-handoff.md`
|
||||||
|
- OpenClaw 接手总览
|
||||||
|
5. `docs/openclaw/openclaw-orchestration-flow.md`
|
||||||
|
- OpenClaw 正式编排手册
|
||||||
|
6. `docs/design/README.md`
|
||||||
|
- 设计文档导航,区分当前有效设计与背景草案
|
||||||
|
7. `plans/README.md`
|
||||||
|
- 规划文档导航
|
||||||
|
8. `TODO.md`
|
||||||
|
- 当前任务状态
|
||||||
|
|
||||||
|
## 目录结构
|
||||||
|
|
||||||
- `docs/README.md`
|
- `docs/README.md`
|
||||||
- 文档总索引
|
- 文档总索引
|
||||||
@@ -8,45 +27,18 @@
|
|||||||
- 当前状态、收束入口、阶段导航
|
- 当前状态、收束入口、阶段导航
|
||||||
- `docs/design/`
|
- `docs/design/`
|
||||||
- 当前实现的设计文档
|
- 当前实现的设计文档
|
||||||
|
- `docs/design/README.md`
|
||||||
|
- 设计文档导航
|
||||||
- `docs/openclaw/`
|
- `docs/openclaw/`
|
||||||
- OpenClaw 日报聚合与下游对象设计
|
- OpenClaw 集成文档、对象规范与历史归档
|
||||||
- `docs/notes/`
|
- `docs/notes/`
|
||||||
- 较上层的方案笔记与非最终设计
|
- 较上层的方案笔记与非最终设计
|
||||||
|
- `docs/notes/README.md`
|
||||||
|
- notes 导航
|
||||||
- `docs/archive/`
|
- `docs/archive/`
|
||||||
- 历史归档,不作为最新事实来源
|
- 历史归档,不作为最新事实来源
|
||||||
|
|
||||||
## 当前推荐阅读顺序
|
## 按主题阅读
|
||||||
|
|
||||||
1. `docs/current/context-reset-brief.md`
|
|
||||||
- 当前真实进度与下一步入口
|
|
||||||
2. `docs/openclaw/openclaw-handoff.md`
|
|
||||||
- 给 OpenClaw 的接手说明、环境变量、MCP 调用方式与已知限制
|
|
||||||
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
|
||||||
- 提供给 OpenClaw 的单篇结构化输入字段说明
|
|
||||||
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
|
|
||||||
- 提供给 OpenClaw 的批量投递 envelope 说明
|
|
||||||
5. `docs/design/summary-mcp-service-design.md`
|
|
||||||
- 当前 MCP 服务的职责、接口和边界
|
|
||||||
6. `docs/design/filter-rule-engine-design.md`
|
|
||||||
- 过滤层的输入输出、规则结构与当前实现
|
|
||||||
7. `docs/design/filter-rule-engine-usage.md`
|
|
||||||
- 规则怎么写、怎么跑、结果怎么解读的使用说明
|
|
||||||
8. `docs/design/daily-keyword-index-design.md`
|
|
||||||
- 日报级词元库与周期性词元清洗 skill 设计
|
|
||||||
9. `docs/design/markdown-sink-design.md`
|
|
||||||
- 第一版 Markdown sink 的输入输出、目录结构与落地方式
|
|
||||||
10. `docs/openclaw/openclaw-daily-digest-refactor.md`
|
|
||||||
- 为什么要从单篇入库改成 OpenClaw 日报聚合链路
|
|
||||||
11. `docs/openclaw/article-candidate-daily-digest-schema.md`
|
|
||||||
- `ArticleCandidateRecord`、`OpenClawCandidateInput` 与 `DailyDigest` 的正式设计
|
|
||||||
12. `docs/design/source-schema-design.md`
|
|
||||||
- `source -> item -> document` 的对象设计
|
|
||||||
13. `docs/notes/reading-pipeline-design-notes.md`
|
|
||||||
- 更上层的阅读流方案与阶段划分
|
|
||||||
14. `docs/design/summary-loop-explained.md`
|
|
||||||
- 当前 LLM 摘要校验闭环的解释
|
|
||||||
|
|
||||||
## 当前文档分层
|
|
||||||
|
|
||||||
### 1. 当前状态与导航
|
### 1. 当前状态与导航
|
||||||
|
|
||||||
@@ -58,11 +50,15 @@
|
|||||||
- 当前阶段状态的最短摘要
|
- 当前阶段状态的最短摘要
|
||||||
- `docs/README.md`
|
- `docs/README.md`
|
||||||
- 文档索引与阅读顺序
|
- 文档索引与阅读顺序
|
||||||
|
- `plans/README.md`
|
||||||
|
- 规划文档导航
|
||||||
|
|
||||||
### 2. 当前实现设计
|
### 2. 当前实现设计
|
||||||
|
|
||||||
|
- `docs/design/README.md`
|
||||||
|
- 设计文档导航与状态说明
|
||||||
- `docs/design/summary-mcp-service-design.md`
|
- `docs/design/summary-mcp-service-design.md`
|
||||||
- 当前内容提取 MCP 的真实设计
|
- 早期 content-extract MCP 设计草案,现主要保留背景参考价值
|
||||||
- `docs/design/summary-core-interface-design.md`
|
- `docs/design/summary-core-interface-design.md`
|
||||||
- 摘要/提取内核的接口抽象
|
- 摘要/提取内核的接口抽象
|
||||||
- `docs/design/source-schema-design.md`
|
- `docs/design/source-schema-design.md`
|
||||||
@@ -82,19 +78,23 @@
|
|||||||
|
|
||||||
### 3. OpenClaw 与下游设计
|
### 3. OpenClaw 与下游设计
|
||||||
|
|
||||||
|
- `docs/openclaw/README.md`
|
||||||
|
- OpenClaw 相关文档导航与归档边界
|
||||||
- `docs/openclaw/openclaw-handoff.md`
|
- `docs/openclaw/openclaw-handoff.md`
|
||||||
- OpenClaw 接手所需的运行说明、工具入口与已知限制
|
- OpenClaw 接手所需的运行说明、工具入口与已知限制
|
||||||
|
- `docs/openclaw/openclaw-orchestration-flow.md`
|
||||||
|
- OpenClaw 编排层的正式运行手册
|
||||||
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||||
- 提供给 OpenClaw 的单篇结构化输入字段说明
|
- 提供给 OpenClaw 的单篇结构化输入字段说明
|
||||||
- `docs/openclaw/openclaw-delivery-payload-spec.md`
|
- `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||||
- 提供给 OpenClaw 的批量投递 envelope 说明
|
- 提供给 OpenClaw 的批量投递 envelope 说明
|
||||||
- `docs/openclaw/openclaw-daily-digest-refactor.md`
|
|
||||||
- 改造为 OpenClaw 日报聚合链路的原因与目标结构
|
|
||||||
- `docs/openclaw/article-candidate-daily-digest-schema.md`
|
- `docs/openclaw/article-candidate-daily-digest-schema.md`
|
||||||
- `ArticleCandidateRecord`、`OpenClawCandidateInput` 与 `DailyDigest` 的字段设计与对象关系
|
- `ArticleCandidateRecord`、`OpenClawCandidateInput` 与 `DailyDigest` 的字段设计与对象关系
|
||||||
|
|
||||||
### 4. 方案笔记
|
### 4. 方案笔记
|
||||||
|
|
||||||
|
- `docs/notes/README.md`
|
||||||
|
- notes 导航与使用边界
|
||||||
- `docs/notes/reading-pipeline-design-notes.md`
|
- `docs/notes/reading-pipeline-design-notes.md`
|
||||||
- 整体阅读流、规则、sink、push 的方案笔记
|
- 整体阅读流、规则、sink、push 的方案笔记
|
||||||
|
|
||||||
@@ -102,11 +102,23 @@
|
|||||||
|
|
||||||
- `docs/archive/content-extract-mcp-mvp-archive.md`
|
- `docs/archive/content-extract-mcp-mvp-archive.md`
|
||||||
- MVP 阶段归档,部分状态已被后续进展覆盖
|
- MVP 阶段归档,部分状态已被后续进展覆盖
|
||||||
|
- `docs/openclaw/archive/README.md`
|
||||||
|
- OpenClaw 历史文档归档说明
|
||||||
|
- `docs/openclaw/archive/formalization-summary-2026-04-07.md`
|
||||||
|
- 第一阶段正式化总结
|
||||||
|
- `docs/openclaw/archive/openclaw-daily-digest-refactor.md`
|
||||||
|
- 早期日报聚合改造背景
|
||||||
|
- `docs/openclaw/archive/digest-optimization-summary.md`
|
||||||
|
- 早期 digest 优化总结
|
||||||
|
- `docs/openclaw/archive/p1-status-reconciliation-plan-2026-04-14.md`
|
||||||
|
- `resume` / 状态收敛问题的阶段修复计划与回填
|
||||||
|
|
||||||
## 当前文档维护原则
|
## 维护原则
|
||||||
|
|
||||||
- `docs/current/context-reset-brief.md` 记录当前最新状态
|
- `docs/current/context-reset-brief.md` 记录当前最新状态
|
||||||
- `TODO.md` 记录任务优先级与下一步
|
- `TODO.md` 记录任务优先级与下一步
|
||||||
|
- `plans/README.md` 负责规划文档分层与导航
|
||||||
|
- `docs/design/README.md` 负责设计文档分层与导航
|
||||||
- `outputs/README.md` 记录当前输出目录约定
|
- `outputs/README.md` 记录当前输出目录约定
|
||||||
- `docs/archive/content-extract-mcp-mvp-archive.md` 只当历史快照,不再作为最新事实来源
|
- `docs/archive/content-extract-mcp-mvp-archive.md` 只当历史快照,不再作为最新事实来源
|
||||||
- 新增阶段性进展,优先更新 `README.md`、`TODO.md`、`docs/current/context-reset-brief.md`
|
- 新增阶段性进展,优先更新 `README.md`、`TODO.md`、`docs/current/context-reset-brief.md`
|
||||||
|
|||||||
@@ -2,152 +2,104 @@
|
|||||||
|
|
||||||
## 当前结论
|
## 当前结论
|
||||||
|
|
||||||
当前仓库已经具备交付给 OpenClaw 的基础条件。
|
当前仓库已经具备作为 OpenClaw 上游服务的正式基础能力。
|
||||||
|
|
||||||
当前主链路是:
|
当前正式主链路是:
|
||||||
|
|
||||||
`FreshRSS 未读 -> RSS 内容提取 -> LLM 总结 -> 规则过滤 -> OpenClaw delivery payload`
|
`FreshRSS 未读 -> RSS 内容提取 -> LLM 总结 -> 规则过滤 -> OpenClaw delivery payload`
|
||||||
|
|
||||||
OpenClaw 应通过 MCP 工具 `run_freshrss_openclaw_pipeline` 调用这条链路,而不是自行拼接脚本。
|
当前正式控制面已经收口为异步 job:
|
||||||
|
|
||||||
## 当前已完成
|
- 主日报:`start_freshrss_pipeline_job -> poll -> get result`
|
||||||
|
- 恢复:`inspect_resume_plan -> start_resume_job -> poll -> get result`
|
||||||
|
- 单篇总结:`start_article_summary_job -> poll -> get result`
|
||||||
|
|
||||||
- 已完成 FreshRSS `greader` API 接入与未读拉取
|
同步 `run_freshrss_openclaw_pipeline`、`resume_run`、`generate_article_summaries` 仍保留,但只用于 debug / fallback。
|
||||||
- 已完成 FreshRSS 条目到标准化 `item` 的映射
|
|
||||||
- 已完成 RSS-first 提取策略
|
|
||||||
- 已完成 LLM 总结与校验闭环
|
|
||||||
- 已完成规则引擎过滤
|
|
||||||
- 已完成 `ArticleCandidateRecord` 与 `OpenClawCandidateInput` 分层
|
|
||||||
- 已完成 `OpenClawDeliveryPayload` 批量投递结构
|
|
||||||
- 已完成 FreshRSS 已读状态回写
|
|
||||||
- 已完成“仅在最终 payload 成功写盘后再标记已读”的语义
|
|
||||||
- 已完成 MCP 工具 `run_freshrss_openclaw_pipeline`
|
|
||||||
- 已完成默认精简输出模式,减少中间文件
|
|
||||||
- 已完成日报级 `keywords` 词元库与全局词频统计
|
|
||||||
- 已完成 `keyword-cleanup-review` skill 骨架与 review bundle 脚本
|
|
||||||
- 已完成低复杂治理层:`term_cleanup_policy` / `term_watchlist` / `term_change_log`
|
|
||||||
- 已完成采纳建议写回脚本 `scripts/apply_term_suggestions.py`
|
|
||||||
|
|
||||||
## 当前 MCP 工具
|
## 当前权威入口
|
||||||
|
|
||||||
当前服务入口:
|
先看这些文档:
|
||||||
|
|
||||||
- `src/summary_mcp/server.py`
|
1. `README.md`
|
||||||
|
2. `docs/README.md`
|
||||||
|
3. `docs/openclaw/README.md`
|
||||||
|
4. `docs/openclaw/openclaw-handoff.md`
|
||||||
|
5. `docs/openclaw/openclaw-orchestration-flow.md`
|
||||||
|
6. `plans/README.md`
|
||||||
|
7. `TODO.md`
|
||||||
|
|
||||||
当前暴露的 MCP 工具:
|
如果问题是 OpenClaw 集成、状态分支或恢复策略,优先看 `docs/openclaw/`,不要先翻历史计划。
|
||||||
|
|
||||||
- `extract_url_content`
|
## 当前正式能力
|
||||||
- `extract_item_content`
|
|
||||||
- `filter_summary_result`
|
|
||||||
- `run_freshrss_openclaw_pipeline`
|
|
||||||
|
|
||||||
其中生产主入口是:
|
- FreshRSS 主日报 run 会落地 `run-state.json`
|
||||||
|
- `get_run_status` / `list_runs` / `list_run_artifacts` 提供 run 级观测
|
||||||
|
- `get_delivery_payload` / `get_run_report` 提供正式结果读取
|
||||||
|
- 查询层已经支持 stale state 与终态 artifacts 的状态收敛
|
||||||
|
- 主日报正式启动已切到 async job
|
||||||
|
- `resume` 已切到 async job,并在执行前先做 `inspect_resume_plan`
|
||||||
|
- 生产恢复依赖 `summary/summary-batch.json` 与 `candidates/candidate-batch.json`
|
||||||
|
- 单篇总结也已补齐 async job 形态
|
||||||
|
|
||||||
- `run_freshrss_openclaw_pipeline`
|
## 当前关键代码入口
|
||||||
|
|
||||||
## 当前关键文件
|
- MCP 服务入口:`src/summary_mcp/server.py`
|
||||||
|
- 主日报 workflow:`src/summary_mcp/workflows/freshrss_pipeline.py`
|
||||||
|
- run / artifact 查询:`src/summary_mcp/runtime/query_service.py`
|
||||||
|
- 主日报 async job:`src/summary_mcp/runtime/freshrss_pipeline_jobs.py`
|
||||||
|
- resume 预检与恢复:`src/summary_mcp/runtime/resume_service.py`
|
||||||
|
- resume async job:`src/summary_mcp/runtime/resume_jobs.py`
|
||||||
|
- 单篇总结 async job:`src/summary_mcp/runtime/article_summary_jobs.py`
|
||||||
|
- 关键词治理:`src/summary_mcp/core/keyword_index.py`
|
||||||
|
|
||||||
- MCP 服务入口
|
## 当前核心产物
|
||||||
- `src/summary_mcp/server.py`
|
|
||||||
- FreshRSS 统一工作流
|
|
||||||
- `src/summary_mcp/workflows/freshrss_pipeline.py`
|
|
||||||
- 词元统计核心
|
|
||||||
- `src/summary_mcp/core/keyword_index.py`
|
|
||||||
- 词元统计模型
|
|
||||||
- `src/summary_mcp/models/keyword_index.py`
|
|
||||||
- 摘要循环
|
|
||||||
- `src/summary_mcp/core/summary_loop.py`
|
|
||||||
- 提取主流程
|
|
||||||
- `src/summary_mcp/core/pipeline.py`
|
|
||||||
- FreshRSS 集成
|
|
||||||
- `src/summary_mcp/integrations/freshrss.py`
|
|
||||||
- 规则引擎
|
|
||||||
- `src/summary_mcp/filters/engine.py`
|
|
||||||
- LLM 结果校验
|
|
||||||
- `src/summary_mcp/validators/llm_result.py`
|
|
||||||
- OpenClaw candidate 模型
|
|
||||||
- `src/summary_mcp/models/article_candidate.py`
|
|
||||||
- OpenClaw delivery 模型
|
|
||||||
- `src/summary_mcp/models/openclaw_delivery.py`
|
|
||||||
- 生产脚本入口
|
|
||||||
- `scripts/run_freshrss_pipeline.py`
|
|
||||||
- 词元统计重建脚本
|
|
||||||
- `scripts/build_keyword_index.py`
|
|
||||||
- 词元清洗 skill
|
|
||||||
- `skills/keyword-cleanup-review/SKILL.md`
|
|
||||||
- skill review bundle 脚本
|
|
||||||
- `skills/keyword-cleanup-review/scripts/build_review_bundle.py`
|
|
||||||
- 采纳建议写回脚本
|
|
||||||
- `scripts/apply_term_suggestions.py`
|
|
||||||
- 清洗治理配置
|
|
||||||
- `configs/term_cleanup_policy.json`
|
|
||||||
- `configs/term_watchlist.json`
|
|
||||||
- `configs/term_change_log.json`
|
|
||||||
- OpenClaw 交接说明
|
|
||||||
- `docs/openclaw/openclaw-handoff.md`
|
|
||||||
|
|
||||||
## 当前输出规则
|
主日报稳定产物:
|
||||||
|
|
||||||
默认生产模式只输出:
|
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
|
||||||
|
- `outputs/freshrss/rerun/<run_dir>/raw/freshrss.raw.json`
|
||||||
|
- `outputs/freshrss/rerun/<run_dir>/summary/summary-batch.json`
|
||||||
|
- `outputs/freshrss/rerun/<run_dir>/candidates/candidate-batch.json`
|
||||||
|
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
|
||||||
|
- `outputs/freshrss/rerun/<run_dir>/candidates/digest-brief.json`
|
||||||
|
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
|
||||||
|
- `outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json`
|
||||||
|
|
||||||
- `raw/freshrss.raw.json`
|
job 状态目录:
|
||||||
- `candidates/openclaw-delivery-payload.json`
|
|
||||||
- `run-report.json`
|
|
||||||
|
|
||||||
同时会更新本地运行数据:
|
- `outputs/freshrss/pipeline_jobs/<job_id>/`
|
||||||
|
- `outputs/freshrss/resume_jobs/<job_id>/`
|
||||||
|
- `outputs/freshrss/article_summary_jobs/<job_id>/`
|
||||||
|
|
||||||
|
关键词运行数据:
|
||||||
|
|
||||||
- `data/term_index/daily/YYYY-MM-DD.json`
|
- `data/term_index/daily/YYYY-MM-DD.json`
|
||||||
- `data/term_index/term_stats.json`
|
- `data/term_index/term_stats.json`
|
||||||
|
|
||||||
如果需要词元清洗审阅输入,可额外生成:
|
## 当前已验证
|
||||||
|
|
||||||
- `outputs/term_index/review/keyword-cleanup-bundle.json`
|
- FreshRSS 未读拉取与已读回写可用
|
||||||
|
- RSS-first 提取策略可用
|
||||||
|
- 主日报 MCP 主链路可触发并写出正式产物
|
||||||
|
- 状态查询与结果读取接口可用
|
||||||
|
- stale state / artifacts 收敛逻辑已落地
|
||||||
|
- `resume` 的 artifact-first 判定已落地
|
||||||
|
- `start_resume_job -> poll -> result` 已做本地 synthetic 验证
|
||||||
|
- 单篇总结 async job 可跑通
|
||||||
|
- 关键词 review bundle 与建议写回脚本可用
|
||||||
|
|
||||||
如果需要在人工确认后把建议正式写入 watchlist / change log,可使用:
|
## 当前主要限制
|
||||||
|
|
||||||
- `scripts/apply_term_suggestions.py`
|
- 某些源 RSS 正文不足时会被直接跳过
|
||||||
|
- 规则仍然偏保守,部分内容会落到 `review`
|
||||||
|
- `paywall` 启发式对中文仍可能误判
|
||||||
|
- 关键词治理还没有接入周期性调度
|
||||||
|
- `digest-brief.json` 仍没有独立 MCP 读取工具
|
||||||
|
- `resume` 目前的剩余主风险不再是恢复点判定,而是缺少真实生产环境的完整恢复验证
|
||||||
|
|
||||||
如果需要排障,可开启:
|
## 当前建议
|
||||||
|
|
||||||
- `debug_artifacts=true`
|
- 把 `docs/openclaw/openclaw-orchestration-flow.md` 当成正式编排手册
|
||||||
- 或脚本参数 `--debug-artifacts`
|
- 把 `docs/openclaw/openclaw-handoff.md` 当成接手总览
|
||||||
|
- 把 `plans/README.md` 当成规划文档导航
|
||||||
这样才会额外输出逐条中间文件。
|
- 把 `docs/openclaw/archive/` 和 `docs/archive/` 当成历史资料,不要当当前事实源
|
||||||
|
|
||||||
## 当前验证状态
|
|
||||||
|
|
||||||
已经验证通过:
|
|
||||||
|
|
||||||
- FreshRSS 未读拉取成功
|
|
||||||
- 已读回写成功
|
|
||||||
- MCP 工具入口可直接触发完整链路
|
|
||||||
- 微信公众号样本可直接使用 RSS 提供的 `summary` 内容提取,不再回源抓网页
|
|
||||||
- 精简输出模式已实际跑通
|
|
||||||
- 日报级词元统计已通过离线样例验证,确认别名、停用词、非 `drop` 过滤和 rerun 覆盖逻辑正常
|
|
||||||
- `keyword-cleanup-review` skill 已通过 `quick_validate.py` 结构校验
|
|
||||||
- review bundle 脚本已实际跑通
|
|
||||||
- `apply_term_suggestions.py` 已通过 dry-run 与临时副本写回验证
|
|
||||||
|
|
||||||
## 当前已知限制
|
|
||||||
|
|
||||||
- 当前对 FreshRSS 条目采用 RSS-first 策略,不再回源抓原网页
|
|
||||||
- 如果 RSS 中没有足够正文内容,该条会直接跳过,不会进入后续总结
|
|
||||||
- 某些规则仍偏保守,部分内容可能落到 `review`
|
|
||||||
- `paywall` 相关启发式仍可能误判中文文本
|
|
||||||
- Webhook / 主动投递到 OpenClaw 外部接口尚未实现,当前是由 OpenClaw 通过 MCP 主动调用
|
|
||||||
- 词元清洗 skill 当前已支持“bundle 构建 -> 建议审阅 -> 人工确认写回 watchlist/change_log”,但尚未接入周期性调度
|
|
||||||
- 当前词元统计仍以前置 `OpenClawDeliveryPayload` 作为日报前代理输入,真实 `DailyDigest` 接入后还需切换上游
|
|
||||||
|
|
||||||
## 当前最建议的交接阅读顺序
|
|
||||||
|
|
||||||
1. `README.md`
|
|
||||||
2. `docs/openclaw/openclaw-handoff.md`
|
|
||||||
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
|
||||||
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
|
|
||||||
5. `docs/design/daily-keyword-index-design.md`
|
|
||||||
6. `skills/keyword-cleanup-review/SKILL.md`
|
|
||||||
7. `TODO.md`
|
|
||||||
|
|
||||||
## 一句话结论
|
|
||||||
|
|
||||||
当前仓库已经从“提取 MCP 原型”演进到“可供 OpenClaw 调用的 FreshRSS -> OpenClaw payload 上游处理器”,并已补上第一阶段的日报级词元统计能力和词元清洗 skill 骨架;后续重点转向 skill 周期调度、知识库状态流转和 webhook 接线。
|
|
||||||
|
|||||||
@@ -0,0 +1,56 @@
|
|||||||
|
# 设计文档导航
|
||||||
|
|
||||||
|
## 使用原则
|
||||||
|
|
||||||
|
`docs/design/` 目录同时包含两类文档:
|
||||||
|
|
||||||
|
- 当前实现仍然有效的设计说明
|
||||||
|
- 早期架构草案和背景设计
|
||||||
|
|
||||||
|
不要默认把这里所有文档都当成当前生产事实。
|
||||||
|
当前生产事实仍以这些入口为准:
|
||||||
|
|
||||||
|
1. `README.md`
|
||||||
|
2. `docs/current/context-reset-brief.md`
|
||||||
|
3. `docs/openclaw/README.md`
|
||||||
|
4. `docs/openclaw/openclaw-handoff.md`
|
||||||
|
5. `docs/openclaw/openclaw-orchestration-flow.md`
|
||||||
|
6. `TODO.md`
|
||||||
|
|
||||||
|
## 当前实现仍然有效
|
||||||
|
|
||||||
|
- `filter-rule-engine-design.md`
|
||||||
|
- 规则过滤层的设计与职责边界
|
||||||
|
- `filter-rule-engine-usage.md`
|
||||||
|
- 规则引擎的使用说明
|
||||||
|
- `daily-keyword-index-design.md`
|
||||||
|
- 关键词索引与清洗治理设计
|
||||||
|
- `summary-loop-explained.md`
|
||||||
|
- LLM 摘要校验闭环说明
|
||||||
|
- `markdown-sink-design.md`
|
||||||
|
- Markdown sink 设计
|
||||||
|
|
||||||
|
## 当前仍有参考价值,但不是生产真相入口
|
||||||
|
|
||||||
|
- `summary-mcp-service-design.md`
|
||||||
|
- 早期 MCP 服务设计草案,部分定位已被后续 workflow service 演进覆盖
|
||||||
|
- `source-schema-design.md`
|
||||||
|
- 更偏对象建模和来源抽象的背景设计
|
||||||
|
- `summary-core-interface-design.md`
|
||||||
|
- 更偏早期摘要内核接口抽象
|
||||||
|
|
||||||
|
## 建议阅读顺序
|
||||||
|
|
||||||
|
如果你是在理解当前实现:
|
||||||
|
|
||||||
|
1. `filter-rule-engine-design.md`
|
||||||
|
2. `filter-rule-engine-usage.md`
|
||||||
|
3. `daily-keyword-index-design.md`
|
||||||
|
4. `summary-loop-explained.md`
|
||||||
|
5. `markdown-sink-design.md`
|
||||||
|
|
||||||
|
如果你是在回看背景设计:
|
||||||
|
|
||||||
|
1. `summary-mcp-service-design.md`
|
||||||
|
2. `source-schema-design.md`
|
||||||
|
3. `summary-core-interface-design.md`
|
||||||
@@ -1,5 +1,16 @@
|
|||||||
# Source Schema 设计草案
|
# Source Schema 设计草案
|
||||||
|
|
||||||
|
## 状态说明
|
||||||
|
|
||||||
|
本文件偏对象建模和来源抽象,主要用于解释早期 schema 设计思路。
|
||||||
|
|
||||||
|
它不是当前生产运行手册,也不是当前 workflow service 的唯一真相来源。
|
||||||
|
如果你关注当前 OpenClaw 集成或运行状态,应优先看:
|
||||||
|
|
||||||
|
- `docs/current/context-reset-brief.md`
|
||||||
|
- `docs/openclaw/README.md`
|
||||||
|
- `docs/openclaw/openclaw-orchestration-flow.md`
|
||||||
|
|
||||||
## 1. 文档目的
|
## 1. 文档目的
|
||||||
|
|
||||||
本文档用于定义阅读流系统中的来源与内容对象模型,目标是把“来源分类”的讨论收敛成一套可执行的数据结构,供后续的抓取、摘要、过滤、入库和推送流程统一使用。
|
本文档用于定义阅读流系统中的来源与内容对象模型,目标是把“来源分类”的讨论收敛成一套可执行的数据结构,供后续的抓取、摘要、过滤、入库和推送流程统一使用。
|
||||||
|
|||||||
@@ -1,5 +1,17 @@
|
|||||||
# Summary Core Interface 设计草案
|
# Summary Core Interface 设计草案
|
||||||
|
|
||||||
|
## 状态说明
|
||||||
|
|
||||||
|
本文件记录的是较早期的摘要内核接口抽象。
|
||||||
|
|
||||||
|
它更适合用于理解背景设计,不应直接当成当前生产接口契约。
|
||||||
|
当前接口与编排真相请优先看:
|
||||||
|
|
||||||
|
- `README.md`
|
||||||
|
- `docs/current/context-reset-brief.md`
|
||||||
|
- `docs/openclaw/README.md`
|
||||||
|
- `docs/openclaw/openclaw-handoff.md`
|
||||||
|
|
||||||
## 1. 文档目的
|
## 1. 文档目的
|
||||||
|
|
||||||
本文档用于定义 `summary-core` 的输入输出接口,目标是把“页面摘要能力”从概念讨论收敛成一套稳定、可复用、可封装的数据接口。
|
本文档用于定义 `summary-core` 的输入输出接口,目标是把“页面摘要能力”从概念讨论收敛成一套稳定、可复用、可封装的数据接口。
|
||||||
|
|||||||
@@ -1,5 +1,18 @@
|
|||||||
# Content Extract MCP Service 设计草案
|
# Content Extract MCP Service 设计草案
|
||||||
|
|
||||||
|
## 状态说明
|
||||||
|
|
||||||
|
本文件主要记录早期 “content extract MCP” 的设计抽象。
|
||||||
|
|
||||||
|
它仍有背景参考价值,但不是当前生产事实入口。
|
||||||
|
当前生产能力已经演进为更完整的 workflow service,正式口径请优先看:
|
||||||
|
|
||||||
|
- `README.md`
|
||||||
|
- `docs/current/context-reset-brief.md`
|
||||||
|
- `docs/openclaw/README.md`
|
||||||
|
- `docs/openclaw/openclaw-handoff.md`
|
||||||
|
- `docs/openclaw/openclaw-orchestration-flow.md`
|
||||||
|
|
||||||
## 1. 文档目的
|
## 1. 文档目的
|
||||||
|
|
||||||
本文档用于定义当前仓库中已经落地的 MCP 服务设计,即“内容提取 MCP”。
|
本文档用于定义当前仓库中已经落地的 MCP 服务设计,即“内容提取 MCP”。
|
||||||
@@ -12,7 +25,7 @@
|
|||||||
- validator 与 LLM 摘要如何接在 MCP 之后
|
- validator 与 LLM 摘要如何接在 MCP 之后
|
||||||
- 当前 MVP 已完成到哪一层
|
- 当前 MVP 已完成到哪一层
|
||||||
|
|
||||||
这份文档描述的是当前真实实现,而不是早期“摘要 MCP”设想。
|
这份文档主要记录当时实现阶段的设计取向,而不是当前生产阶段的唯一事实来源。
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,22 @@
|
|||||||
|
# Notes 导航
|
||||||
|
|
||||||
|
`docs/notes/` 保存的是更早期、讨论型、背景型方案笔记。
|
||||||
|
|
||||||
|
这些文档的用途是:
|
||||||
|
|
||||||
|
- 理解项目最初的问题空间
|
||||||
|
- 回看为什么会形成现在的对象分层和流程划分
|
||||||
|
|
||||||
|
这些文档不是当前生产事实来源。
|
||||||
|
|
||||||
|
当前如需判断“现在到底怎么跑”,优先看:
|
||||||
|
|
||||||
|
- `README.md`
|
||||||
|
- `docs/current/context-reset-brief.md`
|
||||||
|
- `docs/openclaw/README.md`
|
||||||
|
- `docs/openclaw/openclaw-orchestration-flow.md`
|
||||||
|
|
||||||
|
当前 notes:
|
||||||
|
|
||||||
|
- `reading-pipeline-design-notes.md`
|
||||||
|
- 早期阅读流方案讨论纪要
|
||||||
@@ -0,0 +1,36 @@
|
|||||||
|
# OpenClaw 文档导航
|
||||||
|
|
||||||
|
## 当前有效文档
|
||||||
|
|
||||||
|
- `openclaw-handoff.md`
|
||||||
|
- 面向接手者的总览文档
|
||||||
|
- 说明 reader 的职责边界、MCP 工具面、环境变量和正式集成约束
|
||||||
|
- `openclaw-orchestration-flow.md`
|
||||||
|
- 面向 OpenClaw 编排层的正式运行手册
|
||||||
|
- 说明启动、轮询、读结果、恢复和人工介入的标准动作
|
||||||
|
- `openclaw-candidate-input-field-spec.md`
|
||||||
|
- 单篇 `OpenClawCandidateInput` 字段规范
|
||||||
|
- `openclaw-delivery-payload-spec.md`
|
||||||
|
- 批量 `OpenClawDeliveryPayload` 字段规范
|
||||||
|
- `article-candidate-daily-digest-schema.md`
|
||||||
|
- 对象分层设计说明
|
||||||
|
- 用于理解 `ArticleCandidateRecord` / `OpenClawCandidateInput` / `DailyDigest` 的关系
|
||||||
|
|
||||||
|
## 当前推荐阅读顺序
|
||||||
|
|
||||||
|
1. `openclaw-handoff.md`
|
||||||
|
2. `openclaw-orchestration-flow.md`
|
||||||
|
3. `openclaw-candidate-input-field-spec.md`
|
||||||
|
4. `openclaw-delivery-payload-spec.md`
|
||||||
|
5. `article-candidate-daily-digest-schema.md`
|
||||||
|
|
||||||
|
## 归档说明
|
||||||
|
|
||||||
|
`archive/` 下的文档保留历史决策、阶段总结和排障规划,但不再作为当前事实来源。
|
||||||
|
|
||||||
|
当前已归档:
|
||||||
|
|
||||||
|
- `archive/formalization-summary-2026-04-07.md`
|
||||||
|
- `archive/openclaw-daily-digest-refactor.md`
|
||||||
|
- `archive/digest-optimization-summary.md`
|
||||||
|
- `archive/p1-status-reconciliation-plan-2026-04-14.md`
|
||||||
@@ -0,0 +1,9 @@
|
|||||||
|
# OpenClaw 历史归档
|
||||||
|
|
||||||
|
本目录只保留阶段性总结、设计演进记录和排障计划。
|
||||||
|
|
||||||
|
使用原则:
|
||||||
|
|
||||||
|
- 需要了解“为什么会这样设计”时再看
|
||||||
|
- 不要把这里的描述当成当前生产事实
|
||||||
|
- 当前正式口径以 `docs/openclaw/README.md`、`docs/openclaw/openclaw-handoff.md`、`docs/openclaw/openclaw-orchestration-flow.md` 为准
|
||||||
@@ -0,0 +1,367 @@
|
|||||||
|
# Reader 日报链路 P1 状态收敛问题:规划与修复清单(2026-04-14)
|
||||||
|
|
||||||
|
## 背景
|
||||||
|
|
||||||
|
在 2026-04-14 的 reader 日报正式运行中,出现了以下现象:
|
||||||
|
|
||||||
|
- `openclaw-delivery-payload.json`、`digest-brief.json`、`run-report.json` 已真实落盘
|
||||||
|
- 但 `get_freshrss_pipeline_job_status` / `get_run_status` 仍可能显示:
|
||||||
|
- `running`
|
||||||
|
- `failed`
|
||||||
|
- 或 `current_stage=generate_summaries`
|
||||||
|
- `resume_run` 在这种状态下可能直接超时
|
||||||
|
|
||||||
|
这说明当前 reader 的**状态层(job/run-state)**与**产物层(artifacts/report)**之间没有稳定收敛。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 本次确认的核心结论
|
||||||
|
|
||||||
|
### 1. job status 与 run status 是两套独立状态系统
|
||||||
|
|
||||||
|
- **job 层状态**:`src/summary_mcp/runtime/freshrss_pipeline_jobs.py`
|
||||||
|
- `start_freshrss_pipeline_job()`
|
||||||
|
- `run_freshrss_pipeline_job()`
|
||||||
|
- `get_freshrss_pipeline_job_status()`
|
||||||
|
- 状态文件位于:`outputs/freshrss/pipeline_jobs/<job_id>/run-state.json`
|
||||||
|
- 只有 4 个粗粒度 stage:
|
||||||
|
- `prepare_job`
|
||||||
|
- `load_input`
|
||||||
|
- `run_pipeline`
|
||||||
|
- `write_result`
|
||||||
|
|
||||||
|
- **run 层状态**:`src/summary_mcp/workflows/freshrss_pipeline.py`
|
||||||
|
- `run_freshrss_pipeline()`
|
||||||
|
- 由 `src/summary_mcp/runtime/query_service.py:get_run_status()` 查询
|
||||||
|
- 状态文件位于:`outputs/freshrss/rerun/<run_dir>/run-state.json`
|
||||||
|
- 包含 6 个细粒度 stage:
|
||||||
|
- `fetch_feed`
|
||||||
|
- `extract_articles`
|
||||||
|
- `generate_summaries`
|
||||||
|
- `apply_filters`
|
||||||
|
- `build_delivery_payload`
|
||||||
|
- `write_run_report`
|
||||||
|
|
||||||
|
**问题:** 两套状态没有统一收敛规则,用户可以同时看到两套不同口径的“当前进度”。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### 2. 查询层目前优先信 run-state,不会用 artifacts / run-report 纠偏
|
||||||
|
|
||||||
|
代码位置:`src/summary_mcp/runtime/query_service.py`
|
||||||
|
|
||||||
|
关键行为:
|
||||||
|
- `_resolve_run_record()` 只要发现 `run-state.json` 存在,就优先使用 `RunStore.load(...)`
|
||||||
|
- 即使 `run-report.json`、`delivery_payload`、`digest_brief` 已存在,也不会自动纠偏状态
|
||||||
|
|
||||||
|
**结果:**
|
||||||
|
- 一旦 `run-state.json` 因中断、超时、外层 SIGTERM 或写回未完成而停留在旧值
|
||||||
|
- `get_run_status()` 就会持续返回过期状态
|
||||||
|
- 造成“产物已完成,但状态仍显示 running/failed/卡在 summary”的错觉
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### 3. `generate_summaries` 假卡住,本质上更像 stale state,不像真实业务卡住
|
||||||
|
|
||||||
|
代码位置:`src/summary_mcp/workflows/freshrss_pipeline.py`
|
||||||
|
|
||||||
|
从执行顺序看:
|
||||||
|
1. `start_stage(generate_summaries)`
|
||||||
|
2. summary 循环
|
||||||
|
3. `finish_stage(generate_summaries)`
|
||||||
|
4. `start_stage(apply_filters)`
|
||||||
|
5. `finish_stage(apply_filters)`
|
||||||
|
6. `start_stage(build_delivery_payload)`
|
||||||
|
7. 写 payload / digest brief
|
||||||
|
8. `finish_stage(build_delivery_payload)`
|
||||||
|
9. `start_stage(write_run_report)`
|
||||||
|
10. 写 run-report
|
||||||
|
11. `finish_stage(write_run_report)`
|
||||||
|
12. `finish_run(...)`
|
||||||
|
|
||||||
|
**判断:**
|
||||||
|
如果 payload / digest brief / run-report 都已经存在,那么“仍显示卡在 `generate_summaries`”更可能是:
|
||||||
|
- `run-state.json` 没来得及写回最终状态
|
||||||
|
- 或查询时读到了旧状态
|
||||||
|
|
||||||
|
而不是 summary 阶段真实没有跑过去。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### 4. `resume_run` 不是轻量恢复,而是同步继续跑工作流
|
||||||
|
|
||||||
|
代码位置:`src/summary_mcp/runtime/resume_service.py`
|
||||||
|
|
||||||
|
关键行为:
|
||||||
|
- `resume_run()` 会根据 `resume_from_stage` 直接继续执行:
|
||||||
|
- `_run_summary_stage(...)`
|
||||||
|
- `_run_filter_stage(...)`
|
||||||
|
- `_run_delivery_stage(...)`
|
||||||
|
- `_run_report_stage(...)`
|
||||||
|
|
||||||
|
这意味着它不是“修状态”的工具,而是“同步继续跑剩余工作流”的工具。
|
||||||
|
|
||||||
|
**问题:**
|
||||||
|
- 如果 stale state 把 `resume_from_stage` 定在 `generate_summaries`
|
||||||
|
- 那么 `resume_run` 会从一个过早阶段重新跑
|
||||||
|
- 在 MCP 包装层下非常容易超时
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 问题分类
|
||||||
|
|
||||||
|
### A. 真实 bug
|
||||||
|
|
||||||
|
1. **查询层过度信任 stale `run-state.json`**
|
||||||
|
- 文件:`src/summary_mcp/runtime/query_service.py`
|
||||||
|
- 影响:产物已完成但状态仍错误
|
||||||
|
|
||||||
|
2. **`resume_run` 过度依赖 stale `current_stage` / recovery 信息**
|
||||||
|
- 文件:`src/summary_mcp/runtime/resume_service.py`
|
||||||
|
- 影响:从过早阶段重跑,放大 timeout 风险
|
||||||
|
|
||||||
|
### B. 状态设计缺陷
|
||||||
|
|
||||||
|
3. **job 层与 run 层两套状态源没有统一收敛规则**
|
||||||
|
- 文件:`src/summary_mcp/runtime/freshrss_pipeline_jobs.py`
|
||||||
|
- 文件:`src/summary_mcp/runtime/query_service.py`
|
||||||
|
- 影响:用户看到两个互相打架的状态解释
|
||||||
|
|
||||||
|
4. **状态系统完全依赖显式写回,不会按产物反推修正**
|
||||||
|
- 文件:`src/summary_mcp/runtime/run_store.py`
|
||||||
|
- 影响:一旦中断,状态比产物更容易脏
|
||||||
|
|
||||||
|
### C. 调用层误判
|
||||||
|
|
||||||
|
5. **把 `resume_run` 当成轻量恢复接口使用**
|
||||||
|
- 实际上它更接近“同步恢复执行器”
|
||||||
|
- 影响:在长链路场景下超时是高概率事件
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 修复目标
|
||||||
|
|
||||||
|
## 当前落地状态(回填)
|
||||||
|
|
||||||
|
- [x] Phase 1 已落地:`get_run_status()` 会基于 `run-report.json` 与关键产物做终态收敛,并暴露 `status_source` / `state_conflict`
|
||||||
|
- [x] Phase 2 已落地第一阶段:`resume_run()` 会拒绝对已有终态 `run-report.json` 的 run 继续恢复
|
||||||
|
- [x] Phase 2 已继续增强:恢复起点现在会优先根据 artifacts 重算,而不是直接盲信 `run-state.recovery.resume_from_stage`
|
||||||
|
- [x] 新增 `inspect_resume_plan(run_id)` 作为恢复前置判定接口,避免调用方用 `resume_run` 探路
|
||||||
|
- [x] Phase 2 已补齐生产恢复 artifacts:正式 run 会稳定写出 `summary/summary-batch.json` 与 `candidates/candidate-batch.json`,`resume_run` / `inspect_resume_plan` 会优先使用它们,而不是依赖 debug per-item 文件
|
||||||
|
- [x] Phase 3 已落地:job 状态与结果读取会基于 linked run 做收敛,避免 outer job stale state 卡住编排
|
||||||
|
|
||||||
|
### 一级目标(必须达成)
|
||||||
|
|
||||||
|
1. 当 `run-report.json` / `delivery_payload` / `digest_brief` 已存在时,`get_run_status()` 不应继续盲目展示明显过期的 stage 状态;对调用方暴露的 `status` 必须直接收敛为可用终态,而不是只附加 hint
|
||||||
|
2. 当状态层与产物层冲突时,查询结果必须显式标注“状态冲突 / stale state”
|
||||||
|
3. `resume_run()` 在恢复前应优先基于现有 artifacts 判断真实可恢复起点,避免从过早阶段重跑
|
||||||
|
|
||||||
|
### 二级目标(建议达成)
|
||||||
|
|
||||||
|
4. job 层状态结果中增加对 linked run 的补充解释,避免“job running 但 run 产物已齐”这种情况毫无说明
|
||||||
|
5. 为后续编排层提供明确可消费的“状态可信度/冲突提示”字段
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 最小修复方案
|
||||||
|
|
||||||
|
### Phase 1|先修 run 查询层(优先级最高)
|
||||||
|
|
||||||
|
#### 目标
|
||||||
|
让 `get_run_status()` 至少能正确识别:
|
||||||
|
- run-state 是旧的
|
||||||
|
- 但关键产物已经齐了
|
||||||
|
|
||||||
|
#### 建议改动点
|
||||||
|
文件:`src/summary_mcp/runtime/query_service.py`
|
||||||
|
|
||||||
|
#### 建议动作
|
||||||
|
- [x] 在 `_resolve_run_record()` 或 `_build_status_response()` 中增加“关键产物存在性检查”
|
||||||
|
- `run-report.json`
|
||||||
|
- `candidates/openclaw-delivery-payload.json`
|
||||||
|
- `candidates/digest-brief.json`
|
||||||
|
- [x] 如果 `run-state.current_stage` 仍停留在早期阶段,但关键产物已齐:
|
||||||
|
- 不要继续原样输出为可信最终态
|
||||||
|
- 应直接把对外 `status` / `current_stage` / `recovery` 收敛成终态语义
|
||||||
|
- 同时新增解释字段,例如:
|
||||||
|
- `state_conflict: true`
|
||||||
|
- `state_conflict_reason: "run_state indicates generate_summaries but run-report.json already proves the workflow reached a terminal state"`
|
||||||
|
- `status_source: "run_report_reconciliation"`
|
||||||
|
- [x] 保留 `state_source=run_state`,但增加 `status_source` / `state_quality` / `state_conflict` 之类解释字段
|
||||||
|
|
||||||
|
#### 预期收益
|
||||||
|
- OpenClaw 继续按 `status` 分支时也不会卡住
|
||||||
|
- 第一时间减少“明明产物齐了却还像没跑完”的误判
|
||||||
|
- 不需要立刻动 workflow 主链路
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Phase 2|修 `resume_run` 的恢复起点判断
|
||||||
|
|
||||||
|
#### 目标
|
||||||
|
避免 stale state 让恢复逻辑从 `generate_summaries` 这类过早阶段重跑。
|
||||||
|
|
||||||
|
#### 建议改动点
|
||||||
|
文件:`src/summary_mcp/runtime/resume_service.py`
|
||||||
|
|
||||||
|
#### 建议动作
|
||||||
|
- [x] 在 `_resolve_resume_from_stage()` 之前/之后加入真实 artifacts 检查
|
||||||
|
- [x] 如果以下文件已存在:
|
||||||
|
- `openclaw-delivery-payload.json`
|
||||||
|
- `digest-brief.json`
|
||||||
|
- `run-report.json`
|
||||||
|
则不要再从 `generate_summaries` 或 `apply_filters` 起跑
|
||||||
|
- [x] 为 `resume_run()` 增加“恢复起点是基于 artifacts 重算还是基于 state 推断”的返回说明
|
||||||
|
- [x] 必要时增加更保守逻辑:
|
||||||
|
- `run-report.json` 已存在时,默认拒绝继续 resume,并提示“产物已完成,请先检查状态一致性”
|
||||||
|
- 补充:默认生产模式下,主链路会稳定写出 `summary-batch` / `candidate-batch`,恢复逻辑优先消费这两个 batch artifacts;若它们缺失或不稳定,才回退到更早的安全 stage 或直接拒绝恢复
|
||||||
|
- 补充:调用方可先走 `inspect_resume_plan`,只有 `recommended_action=resume` 时再调用 `resume_run`
|
||||||
|
|
||||||
|
#### 预期收益
|
||||||
|
- 降低无意义重跑和 timeout 风险
|
||||||
|
- 让 `resume_run` 更接近真正的恢复工具,而不是误重跑工具
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Phase 3|补 job/run 双状态解释层
|
||||||
|
|
||||||
|
#### 目标
|
||||||
|
让 `get_freshrss_pipeline_job_status()` 和 `get_run_status()` 的关系对调用方更可理解。
|
||||||
|
|
||||||
|
#### 建议改动点
|
||||||
|
文件:`src/summary_mcp/runtime/freshrss_pipeline_jobs.py`
|
||||||
|
|
||||||
|
#### 建议动作
|
||||||
|
- [x] 在 `get_freshrss_pipeline_job_status()` 中,读取 linked run 的关键产物存在性(轻量即可)
|
||||||
|
- [x] 若 job 仍显示 `run_pipeline`,但 linked run 已有 report/payload/digest 产物:
|
||||||
|
- 不仅增加解释字段,还应直接把 job 对外 `status` 收敛为终态,避免外层永远轮询
|
||||||
|
- 例如:
|
||||||
|
- `status_source: "linked_run_reconciliation"`
|
||||||
|
- `status_note: "linked run artifacts are complete; the job can be treated as completed"`
|
||||||
|
- [x] 若 `result.json` 缺失,但 linked run 已有 `run-report.json` 与 delivery 产物:
|
||||||
|
- `get_freshrss_pipeline_job_result()` 应能基于 linked run 产物合成最小结果,至少稳定返回 `run_id`
|
||||||
|
- [x] 明确文档:job status 是外层异步任务态,不等于内部 workflow 细粒度状态
|
||||||
|
|
||||||
|
#### 预期收益
|
||||||
|
- 减少“job running / run finished”口径冲突带来的误解
|
||||||
|
- 避免 OpenClaw 因 outer job stale state 卡死在轮询和 result 读取前
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Phase 4|把 `resume_run` 改成异步恢复 job
|
||||||
|
|
||||||
|
#### 目标
|
||||||
|
解决当前剩余的核心问题:`resume_run` 虽然恢复判定已经安全,但执行模型仍是同步 MCP 调用,长链路恢复时依然可能超时,导致 OpenClaw 编排层“看起来像又卡住了”。
|
||||||
|
|
||||||
|
#### 建议改动点
|
||||||
|
文件:
|
||||||
|
- `src/summary_mcp/runtime/resume_jobs.py`(新)
|
||||||
|
- `scripts/run_resume_job.py`(新)
|
||||||
|
- `src/summary_mcp/server.py`
|
||||||
|
- `src/summary_mcp/runtime/__init__.py`
|
||||||
|
- `src/summary_mcp/runtime/resume_service.py`
|
||||||
|
|
||||||
|
#### 建议动作
|
||||||
|
- [x] 新增最小异步恢复接口:
|
||||||
|
- `start_resume_job(run_id)`
|
||||||
|
- `get_resume_job_status(job_id)`
|
||||||
|
- `get_resume_job_result(job_id)`
|
||||||
|
- [x] job 目录固定落到:
|
||||||
|
- `outputs/freshrss/resume_jobs/<job_id>/`
|
||||||
|
- [x] 最少产物约定:
|
||||||
|
- `run-state.json`
|
||||||
|
- `input.json`
|
||||||
|
- `result.json`(成功时)
|
||||||
|
- `job-report.json`
|
||||||
|
- [x] `start_resume_job` 内部先调用 `inspect_resume_plan`
|
||||||
|
- 只有 `recommended_action=resume` 才允许真正启动
|
||||||
|
- `read_terminal_result` / `start_new_run` 要直接在 job 输入校验阶段返回,不进入执行器
|
||||||
|
- [x] 后台执行时复用现有 `_resume_freshrss_run(...)`
|
||||||
|
- 不重写恢复业务逻辑
|
||||||
|
- 只把同步入口拆成异步 job 外壳
|
||||||
|
- [x] `resume_run(run_id)` 保留,但降级为 debug / fallback
|
||||||
|
- 文档中明确:OpenClaw 编排默认应走 resume async job,而不是同步 `resume_run`
|
||||||
|
- [x] job result 里至少稳定返回:
|
||||||
|
- `run_id`
|
||||||
|
- `resume_from_stage`
|
||||||
|
- `status`
|
||||||
|
- `result_source`
|
||||||
|
- `delivery_output` / `report_output`(若存在)
|
||||||
|
|
||||||
|
#### 预期收益
|
||||||
|
- 彻底切掉恢复阶段的 MCP 同步超时风险
|
||||||
|
- 让 OpenClaw 对“启动恢复 / 轮询恢复 / 读取恢复结果”的控制面与主 pipeline async job 保持一致
|
||||||
|
- 把“恢复判定”与“恢复执行”分层,减少误调用和卡住错觉
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 不建议现在就做的事
|
||||||
|
|
||||||
|
- [ ] **不要先做自动 fallback 修状态**
|
||||||
|
- 例如:看到 artifacts 齐了就直接把 run-state 强行改成 success
|
||||||
|
- 原因:这会掩盖真正的状态写回问题
|
||||||
|
|
||||||
|
- [ ] **不要先大改 workflow 主链路**
|
||||||
|
- 当前更像查询层与恢复层的状态解释缺陷
|
||||||
|
- 先修读取与恢复判断,收益更大、风险更低
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 建议执行顺序
|
||||||
|
|
||||||
|
1. **先改 `query_service.py`**
|
||||||
|
- 让 `get_run_status()` 能暴露 stale state / artifact conflict
|
||||||
|
2. **再改 `resume_service.py`**
|
||||||
|
- 避免从错误阶段重跑
|
||||||
|
3. **最后看 `freshrss_pipeline_jobs.py`**
|
||||||
|
- 给 job status 加 linked run 补充说明
|
||||||
|
4. **收尾改 `resume async job`**
|
||||||
|
- 让恢复执行也走正式异步控制面,避免同步恢复再把编排卡住
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 验收标准
|
||||||
|
|
||||||
|
### 验收 1:状态冲突识别
|
||||||
|
构造一个场景:
|
||||||
|
- `run-state.json` 留在 `generate_summaries`
|
||||||
|
- 但 payload / digest brief / run-report 已存在
|
||||||
|
|
||||||
|
期望:
|
||||||
|
- `get_run_status()` 不再只回“卡在 generate_summaries”
|
||||||
|
- 会显式返回冲突提示字段
|
||||||
|
|
||||||
|
### 验收 2:恢复起点修正
|
||||||
|
构造一个场景:
|
||||||
|
- `run-state` 指向 `generate_summaries`
|
||||||
|
- 但 `delivery_payload` / `run-report` 已存在
|
||||||
|
|
||||||
|
期望:
|
||||||
|
- `resume_run()` 不应再从 summary 阶段重跑
|
||||||
|
- 至少应拒绝恢复并提示“产物已完成,优先检查状态一致性”
|
||||||
|
|
||||||
|
### 验收 3:job/run 双层说明
|
||||||
|
构造一个场景:
|
||||||
|
- job status 仍在 `run_pipeline`
|
||||||
|
- linked run 已有关键产物
|
||||||
|
|
||||||
|
期望:
|
||||||
|
- `get_freshrss_pipeline_job_status()` 能返回补充说明,不再只有生硬 running
|
||||||
|
|
||||||
|
### 验收 4:恢复执行不再阻塞编排
|
||||||
|
构造一个场景:
|
||||||
|
- run 可恢复
|
||||||
|
- 恢复点为 `generate_summaries` 或 `apply_filters`
|
||||||
|
- 恢复执行耗时超过单次 MCP 同步窗口
|
||||||
|
|
||||||
|
期望:
|
||||||
|
- OpenClaw 调用的是 `start_resume_job(...)`,而不是同步 `resume_run(...)`
|
||||||
|
- `get_resume_job_status(job_id)` 可稳定轮询到终态
|
||||||
|
- `get_resume_job_result(job_id)` 至少稳定返回 `run_id`、`resume_from_stage` 与最终产物引用
|
||||||
|
- 即使恢复失败,也能在 job-report / result 中看清失败点,而不是只表现为调用超时
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 备注
|
||||||
|
|
||||||
|
截至 2026-04-14,本文件中的 Phase 1 / 2 / 3 / 4 已完成主要落地;当前 `resume` 链路已经从“状态收敛 + 安全恢复点判定”进一步补齐到“正式异步恢复执行”。
|
||||||
+172
-263
@@ -1,49 +1,54 @@
|
|||||||
# OpenClaw Handoff
|
# OpenClaw Handoff
|
||||||
|
|
||||||
|
## Role
|
||||||
|
|
||||||
|
This file is the integration overview for OpenClaw maintainers.
|
||||||
|
|
||||||
|
Use it for:
|
||||||
|
|
||||||
|
- reader capability boundary
|
||||||
|
- production MCP entrypoints
|
||||||
|
- environment requirements
|
||||||
|
- integration rules and limitations
|
||||||
|
|
||||||
|
Do not use it as the step-by-step runbook.
|
||||||
|
For formal orchestration, read `docs/openclaw/openclaw-orchestration-flow.md`.
|
||||||
|
For field contracts, read:
|
||||||
|
|
||||||
|
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||||
|
- `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||||
|
|
||||||
|
Historical plans and incident documents live under `docs/openclaw/archive/`.
|
||||||
|
|
||||||
## Purpose
|
## Purpose
|
||||||
|
|
||||||
This repository provides a FreshRSS-first reading pipeline for OpenClaw:
|
reader is the upstream FreshRSS processing service for OpenClaw:
|
||||||
|
|
||||||
`FreshRSS unread items -> RSS content extraction -> LLM summary -> rule engine -> OpenClaw delivery payload`
|
`FreshRSS unread items -> RSS content extraction -> LLM summary -> rule engine -> OpenClaw delivery payload`
|
||||||
|
|
||||||
OpenClaw should treat this repository as an MCP-backed upstream content processor.
|
reader is responsible for:
|
||||||
|
|
||||||
This repository is responsible only for upstream reading-pipeline work:
|
|
||||||
|
|
||||||
- FreshRSS pull
|
- FreshRSS pull
|
||||||
- content extraction
|
- content extraction
|
||||||
- LLM summary generation/validation
|
- LLM summary generation and validation
|
||||||
- rule-based filtering
|
- rule-based filtering
|
||||||
- OpenClaw delivery payload generation
|
- OpenClaw delivery payload generation
|
||||||
- selected-article summary capability based on existing extracted text
|
- run-state persistence and run/result lookup
|
||||||
- run-state persistence, status lookup, result lookup, and minimal resume for the FreshRSS workflow
|
- async resume control for the FreshRSS workflow
|
||||||
|
- async selected-article summary generation from existing extracted files
|
||||||
|
|
||||||
This repository should **not** take over downstream orchestration responsibilities that belong to OpenClaw / skills, such as:
|
reader is not responsible for:
|
||||||
|
|
||||||
- Hugo publishing
|
- Hugo publishing
|
||||||
- chat reporting
|
- chat reporting
|
||||||
- user confirmation handling
|
- user confirmation handling
|
||||||
- IMA upload orchestration
|
- IMA upload orchestration
|
||||||
|
|
||||||
## Production Entrypoint
|
## Production Surface
|
||||||
|
|
||||||
reader 当前正式工作流服务启动入口是 MCP tool:
|
Current MCP tool count: 21.
|
||||||
|
|
||||||
- `start_freshrss_pipeline_job`
|
Main daily workflow:
|
||||||
|
|
||||||
OpenClaw 应先拿到 `job_id`,轮询 job 状态,再在成功后读取 `run_id` 作为正式后续句柄。
|
|
||||||
|
|
||||||
`run_freshrss_openclaw_pipeline` 仍保留,但定位是同步 debug / fallback 路径,而不是正式生产启动入口。
|
|
||||||
|
|
||||||
OpenClaw should treat the returned `run_id` from `get_freshrss_pipeline_job_result` as the only stable handle for follow-up reads. Do not hand-build `outputs/freshrss/rerun/...` paths in OpenClaw.
|
|
||||||
|
|
||||||
Job state is written under `outputs/freshrss/pipeline_jobs/<job_id>/` and will minimally contain `run-state.json`, `input.json`, `result.json` on success, and `job-report.json`.
|
|
||||||
|
|
||||||
## Supported MCP Tools
|
|
||||||
|
|
||||||
Current MCP tools: 17 total, including the FreshRSS workflow set plus async job tools for both the main pipeline and article-summary flow.
|
|
||||||
|
|
||||||
Workflow service tools:
|
|
||||||
|
|
||||||
- `start_freshrss_pipeline_job`
|
- `start_freshrss_pipeline_job`
|
||||||
- `get_freshrss_pipeline_job_status`
|
- `get_freshrss_pipeline_job_status`
|
||||||
@@ -53,41 +58,139 @@ Workflow service tools:
|
|||||||
- `list_run_artifacts`
|
- `list_run_artifacts`
|
||||||
- `get_delivery_payload`
|
- `get_delivery_payload`
|
||||||
- `get_run_report`
|
- `get_run_report`
|
||||||
|
|
||||||
|
Resume workflow:
|
||||||
|
|
||||||
|
- `inspect_resume_plan`
|
||||||
|
- `start_resume_job`
|
||||||
|
- `get_resume_job_status`
|
||||||
|
- `get_resume_job_result`
|
||||||
- `resume_run`
|
- `resume_run`
|
||||||
|
|
||||||
Article-summary tools:
|
Selected-article summary workflow:
|
||||||
|
|
||||||
- `start_article_summary_job`
|
- `start_article_summary_job`
|
||||||
- `get_article_summary_job_status`
|
- `get_article_summary_job_status`
|
||||||
- `get_article_summary_job_result`
|
- `get_article_summary_job_result`
|
||||||
|
- `generate_article_summaries`
|
||||||
|
|
||||||
Single-step / debug tools:
|
Debug / single-step tools:
|
||||||
|
|
||||||
- `run_freshrss_openclaw_pipeline`(同步模式,仅适合 debug / fallback)
|
- `run_freshrss_openclaw_pipeline`
|
||||||
- `extract_url_content`
|
- `extract_url_content`
|
||||||
- `extract_item_content`
|
- `extract_item_content`
|
||||||
- `filter_summary_result`
|
- `filter_summary_result`
|
||||||
- `generate_article_summaries`(同步模式,仅适合轻量调试)
|
|
||||||
|
|
||||||
## Recommended Selected-Article Flow
|
Production rules:
|
||||||
|
|
||||||
For OpenClaw selected-article follow-up, prefer the async job path:
|
- main production start path is `start_freshrss_pipeline_job`
|
||||||
|
- production resume path is `inspect_resume_plan -> start_resume_job -> get_resume_job_status -> get_resume_job_result`
|
||||||
|
- `run_freshrss_openclaw_pipeline` is sync debug / fallback only
|
||||||
|
- `resume_run` is sync debug / fallback only
|
||||||
|
- `generate_article_summaries` is sync debug / fallback only
|
||||||
|
|
||||||
1. Call `start_article_summary_job` with a real extracted file path plus a non-empty `selected_ids` list.
|
## Production Contract
|
||||||
2. Poll `get_article_summary_job_status` until `status` becomes `success` or `failed`.
|
|
||||||
3. On success, call `get_article_summary_job_result` and continue downstream processing from `written_paths`.
|
|
||||||
4. Use `generate_article_summaries` only as a synchronous debug fallback, not as the default production path.
|
|
||||||
|
|
||||||
Job state is written under `outputs/freshrss/article_summary_jobs/<job_id>/` and will minimally contain:
|
OpenClaw should treat the returned `run_id` from `get_freshrss_pipeline_job_result` as the only stable handle for follow-up reads.
|
||||||
|
|
||||||
|
OpenClaw should not hand-build these paths:
|
||||||
|
|
||||||
|
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
|
||||||
|
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
|
||||||
|
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
|
||||||
|
|
||||||
|
If filesystem access is needed for debugging, only consume paths returned by MCP:
|
||||||
|
|
||||||
|
- `output_dir`
|
||||||
|
- `artifact.path`
|
||||||
|
- `delivery_output`
|
||||||
|
- `report_output`
|
||||||
|
|
||||||
|
Top-level `status` is the only status field callers should branch on.
|
||||||
|
`status_source` and `state_conflict` are explanatory fields for reconciled status.
|
||||||
|
|
||||||
|
## Minimal Production Sequence
|
||||||
|
|
||||||
|
Daily workflow:
|
||||||
|
|
||||||
|
1. Call `start_freshrss_pipeline_job`
|
||||||
|
2. Poll `get_freshrss_pipeline_job_status`
|
||||||
|
3. On success, read `get_freshrss_pipeline_job_result`
|
||||||
|
4. Persist the returned `run_id`
|
||||||
|
5. Use `get_run_status`, `get_delivery_payload`, and `get_run_report` for follow-up reads
|
||||||
|
|
||||||
|
Resume workflow:
|
||||||
|
|
||||||
|
1. Call `inspect_resume_plan(run_id)`
|
||||||
|
2. Only if `can_resume=true` and `recommended_action=resume`, call `start_resume_job`
|
||||||
|
3. Poll `get_resume_job_status`
|
||||||
|
4. Read `get_resume_job_result`
|
||||||
|
|
||||||
|
Selected-article summary workflow:
|
||||||
|
|
||||||
|
1. Call `start_article_summary_job` with a real extracted file path and non-empty `selected_ids`
|
||||||
|
2. Poll `get_article_summary_job_status`
|
||||||
|
3. Read `get_article_summary_job_result`
|
||||||
|
|
||||||
|
## Capability Boundary
|
||||||
|
|
||||||
|
Formal workflow boundary:
|
||||||
|
|
||||||
|
- only workflow `freshrss_daily_digest`
|
||||||
|
- every current FreshRSS run writes `run-state.json`
|
||||||
|
- `get_run_status` / `list_runs` / `list_run_artifacts` can still infer basic state for older runs without `run-state.json`
|
||||||
|
- resume requires a valid `run-state.json`; inferred historical runs are not resumable
|
||||||
|
|
||||||
|
Resume boundary:
|
||||||
|
|
||||||
|
- resume in place on the original `run_id`
|
||||||
|
- supported resume points:
|
||||||
|
- `generate_summaries`
|
||||||
|
- `apply_filters`
|
||||||
|
- `build_delivery_payload`
|
||||||
|
- `write_run_report`
|
||||||
|
- unsupported resume points:
|
||||||
|
- `fetch_feed`
|
||||||
|
- `extract_articles`
|
||||||
|
- production resume prefers:
|
||||||
|
- `summary/summary-batch.json`
|
||||||
|
- `candidates/candidate-batch.json`
|
||||||
|
- if required artifacts are missing, recovery should return non-resumable instead of silently falling back
|
||||||
|
|
||||||
|
Selected-article summary boundary:
|
||||||
|
|
||||||
|
- uses existing extracted files as input
|
||||||
|
- should not re-fetch original URLs
|
||||||
|
|
||||||
|
## Output Expectations
|
||||||
|
|
||||||
|
Main daily pipeline core artifacts:
|
||||||
|
|
||||||
|
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
|
||||||
|
- `outputs/freshrss/rerun/<run_dir>/raw/freshrss.raw.json`
|
||||||
|
- `outputs/freshrss/rerun/<run_dir>/summary/summary-batch.json`
|
||||||
|
- `outputs/freshrss/rerun/<run_dir>/candidates/candidate-batch.json`
|
||||||
|
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
|
||||||
|
- `outputs/freshrss/rerun/<run_dir>/candidates/digest-brief.json`
|
||||||
|
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
|
||||||
|
- `outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json`
|
||||||
|
|
||||||
|
Async job state directories:
|
||||||
|
|
||||||
|
- main pipeline job: `outputs/freshrss/pipeline_jobs/<job_id>/`
|
||||||
|
- resume job: `outputs/freshrss/resume_jobs/<job_id>/`
|
||||||
|
- article-summary job: `outputs/freshrss/article_summary_jobs/<job_id>/`
|
||||||
|
|
||||||
|
Each job directory minimally contains:
|
||||||
|
|
||||||
- `run-state.json`
|
- `run-state.json`
|
||||||
- `input.json`
|
- `input.json`
|
||||||
- `result.json` on success
|
- `result.json` on success
|
||||||
- `job-report.json`
|
- `job-report.json`
|
||||||
|
|
||||||
## Required Environment Variables
|
## Environment And Startup
|
||||||
|
|
||||||
The MCP server process must have these variables available:
|
Required environment variables:
|
||||||
|
|
||||||
- `FRESHRSS_API_BASE_URL`
|
- `FRESHRSS_API_BASE_URL`
|
||||||
- `FRESHRSS_USERNAME`
|
- `FRESHRSS_USERNAME`
|
||||||
@@ -96,54 +199,13 @@ The MCP server process must have these variables available:
|
|||||||
- `LLM_API_KEY`
|
- `LLM_API_KEY`
|
||||||
- `LLM_MODEL`
|
- `LLM_MODEL`
|
||||||
|
|
||||||
Example:
|
Startup:
|
||||||
|
|
||||||
```powershell
|
|
||||||
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
|
||||||
set FRESHRSS_USERNAME=osiman
|
|
||||||
set FRESHRSS_API_PASSWORD=your-api-password
|
|
||||||
set LLM_API_URL=https://api.deepseek.com
|
|
||||||
set LLM_API_KEY=your-llm-api-key
|
|
||||||
set LLM_MODEL=deepseek-chat
|
|
||||||
```
|
|
||||||
|
|
||||||
## Server Startup
|
|
||||||
|
|
||||||
Install dependencies:
|
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
pip install -e .
|
pip install -e .
|
||||||
```
|
|
||||||
|
|
||||||
Start the MCP server:
|
|
||||||
|
|
||||||
```bash
|
|
||||||
summary-mcp
|
summary-mcp
|
||||||
```
|
```
|
||||||
|
|
||||||
## Recommended MCP Workflow
|
|
||||||
|
|
||||||
Recommended production path:
|
|
||||||
|
|
||||||
1. Call `start_freshrss_pipeline_job` and persist the returned `job_id`
|
|
||||||
2. Poll `get_freshrss_pipeline_job_status(job_id)` until `status` becomes `success` or `failed`
|
|
||||||
3. On success, call `get_freshrss_pipeline_job_result(job_id)` and persist the returned `run_id`
|
|
||||||
4. Use `get_run_status(run_id)` as the authoritative run-state read for status, stage, artifacts, and recovery
|
|
||||||
5. Use `list_runs(...)` when OpenClaw needs recent-run discovery or high-level inspection
|
|
||||||
6. Use `list_run_artifacts(run_id)` when OpenClaw needs to inspect what this run actually produced
|
|
||||||
7. Use `get_delivery_payload(run_id)` and `get_run_report(run_id)` as the formal result-reading APIs
|
|
||||||
8. Use `resume_run(run_id)` only when the run falls inside the minimal supported resume scope
|
|
||||||
|
|
||||||
OpenClaw should not directly derive or hardcode:
|
|
||||||
|
|
||||||
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
|
|
||||||
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
|
|
||||||
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
|
|
||||||
|
|
||||||
If filesystem access is needed for debugging, consume only paths returned by MCP such as `output_dir`, `artifact.path`, `delivery_output`, or `report_output`.
|
|
||||||
|
|
||||||
## Recommended MCP Call
|
|
||||||
|
|
||||||
Recommended production start call:
|
Recommended production start call:
|
||||||
|
|
||||||
```json
|
```json
|
||||||
@@ -157,204 +219,51 @@ Recommended production start call:
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
Recommended semantics:
|
## Data And Content Policy
|
||||||
|
|
||||||
- Use `mark_read=true` for normal production runs.
|
FreshRSS processing is RSS-first:
|
||||||
- Use `mark_read=false` only for debug, test, or validation runs.
|
|
||||||
- Keep `debug_artifacts=false` for routine production runs.
|
|
||||||
- Set `debug_artifacts=true` only when troubleshooting a bad batch.
|
|
||||||
- Treat the returned `job_id` as the startup handle, and the later `run_id` from `get_freshrss_pipeline_job_result` as the stable identifier for all follow-up run reads.
|
|
||||||
- If no real `openclaw-delivery-payload.json` was produced, OpenClaw should stop instead of generating a digest from placeholders or examples.
|
|
||||||
- Use `run_freshrss_openclaw_pipeline` only when a synchronous debug / fallback path is explicitly needed.
|
|
||||||
|
|
||||||
## Formal Capability Boundary
|
|
||||||
|
|
||||||
reader 当前正式 MCP workflow service 的边界如下:
|
|
||||||
|
|
||||||
- formal workflow: only `freshrss_daily_digest`
|
|
||||||
- run truth: every FreshRSS run writes `run-state.json`
|
|
||||||
- main production start path: `start_freshrss_pipeline_job` / `get_freshrss_pipeline_job_status` / `get_freshrss_pipeline_job_result`
|
|
||||||
- state query tools: `get_run_status`, `list_runs`, `list_run_artifacts`
|
|
||||||
- result read tools: `get_delivery_payload`, `get_run_report`
|
|
||||||
- `digest-brief.json` is generated and registered as an artifact, but there is no standalone `get_digest_brief` tool yet
|
|
||||||
- `run_freshrss_openclaw_pipeline` is still supported, but only as a synchronous debug / fallback path
|
|
||||||
- the FreshRSS daily workflow now has a minimal background job model backed by a detached runner process, not a full queue / worker system
|
|
||||||
- article summary now has a minimal asynchronous job model with `start_article_summary_job` / `get_article_summary_job_status` / `get_article_summary_job_result`
|
|
||||||
- `generate_article_summaries` is still supported, but it is a synchronous debug path and outside the formal `resume_run` scope
|
|
||||||
|
|
||||||
Historical compatibility note:
|
|
||||||
|
|
||||||
- `get_run_status` / `list_runs` / `list_run_artifacts` can still infer basic state for older run directories without `run-state.json`
|
|
||||||
- `resume_run` does **not** support those inferred historical runs; it requires a valid `run-state.json`
|
|
||||||
|
|
||||||
## What The Async Job Returns
|
|
||||||
|
|
||||||
Primary return fields from `get_freshrss_pipeline_job_result`:
|
|
||||||
|
|
||||||
- `job_id`
|
|
||||||
- `run_id`
|
|
||||||
- `output_dir`
|
|
||||||
- `raw_output`
|
|
||||||
- `delivery_output`
|
|
||||||
- `report_output`
|
|
||||||
- `digest_brief_output`
|
|
||||||
- `pulled_count`
|
|
||||||
- `delivered_count`
|
|
||||||
- `marked_read_count`
|
|
||||||
- `status_counts`
|
|
||||||
- `keyword_index`
|
|
||||||
|
|
||||||
Optional:
|
|
||||||
|
|
||||||
- `items`
|
|
||||||
- returned only when `include_item_reports=true`
|
|
||||||
|
|
||||||
`run_freshrss_openclaw_pipeline` still returns the same synchronous payload for debug / fallback use.
|
|
||||||
|
|
||||||
Follow-up structured reads should use MCP tools rather than re-reading these files directly.
|
|
||||||
|
|
||||||
## Minimal Output Files
|
|
||||||
|
|
||||||
By default the pipeline writes these core artifacts:
|
|
||||||
|
|
||||||
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
|
|
||||||
- `outputs/freshrss/rerun/<run_dir>/raw/freshrss.raw.json`
|
|
||||||
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
|
|
||||||
- `outputs/freshrss/rerun/<run_dir>/candidates/digest-brief.json`
|
|
||||||
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
|
|
||||||
- `outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json` (one per item)
|
|
||||||
|
|
||||||
It also updates local runtime keyword data:
|
|
||||||
|
|
||||||
- `data/term_index/daily/YYYY-MM-DD.json`
|
|
||||||
- `data/term_index/term_stats.json`
|
|
||||||
|
|
||||||
Per-item extracted files live under `extracted/` and are always written.
|
|
||||||
If `debug_artifacts=true`, the pipeline additionally writes normalized items, summaries, filter decisions, candidate records, and candidate inputs.
|
|
||||||
The main pipeline does not emit a batch-level `freshrss.extracted.json` file by default.
|
|
||||||
|
|
||||||
## `resume_run` Minimal Scope
|
|
||||||
|
|
||||||
`resume_run` currently supports only the minimum resume contract:
|
|
||||||
|
|
||||||
- only runs with a valid `run-state.json`
|
|
||||||
- only workflow `freshrss_daily_digest`
|
|
||||||
- resume in place on the original `run_id`
|
|
||||||
- supported resume points: `generate_summaries`, `apply_filters`, `build_delivery_payload`, `write_run_report`
|
|
||||||
- unsupported resume points: `fetch_feed`, `extract_articles`
|
|
||||||
- if required artifacts are missing, the tool returns a non-resumable response instead of silently falling back to an earlier stage
|
|
||||||
|
|
||||||
Artifact expectations by resume point:
|
|
||||||
|
|
||||||
- `generate_summaries`: requires `raw/freshrss.raw.json` and `extracted/`
|
|
||||||
- `apply_filters`: requires the above plus per-item summary outputs
|
|
||||||
- `build_delivery_payload`: requires per-item candidate inputs consistent with filter-stage output
|
|
||||||
- `write_run_report`: requires `candidates/openclaw-delivery-payload.json`; if `mark_read=true`, raw input must still be present
|
|
||||||
|
|
||||||
## Payload Specs
|
|
||||||
|
|
||||||
OpenClaw payload field specs live here:
|
|
||||||
|
|
||||||
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
|
||||||
- `docs/openclaw/openclaw-delivery-payload-spec.md`
|
|
||||||
|
|
||||||
## Read-State Semantics
|
|
||||||
|
|
||||||
The pipeline reads from FreshRSS unread items by default.
|
|
||||||
|
|
||||||
If `mark_read=true`:
|
|
||||||
|
|
||||||
- items are marked as read only after the final `openclaw-delivery-payload.json` has been written successfully
|
|
||||||
- only successfully delivered items are marked as read
|
|
||||||
- failed or skipped items remain unread
|
|
||||||
|
|
||||||
## FreshRSS Content Policy
|
|
||||||
|
|
||||||
For FreshRSS items, the pipeline is RSS-first and does not re-crawl webpages.
|
|
||||||
|
|
||||||
Behavior:
|
|
||||||
|
|
||||||
- use `item.raw_content` first
|
- use `item.raw_content` first
|
||||||
- if missing, use `item.raw_summary`
|
- if missing, use `item.raw_summary`
|
||||||
- if neither contains usable content, skip the item
|
- if neither contains usable content, skip the item
|
||||||
- do not fetch the original webpage again for FreshRSS items
|
- do not fetch the original webpage again for FreshRSS items
|
||||||
|
|
||||||
This is intentional.
|
Read-state policy:
|
||||||
|
|
||||||
## Keyword Cleanup Governance
|
- items are marked read only after successful delivery payload write
|
||||||
|
- only successfully delivered items are marked read
|
||||||
|
|
||||||
This repository also includes a lightweight keyword-governance flow for downstream review.
|
Downstream boundary:
|
||||||
|
|
||||||
Current pieces:
|
- the daily digest goes to Hugo and chat reporting
|
||||||
|
- the full daily digest should not be uploaded to IMA
|
||||||
- runtime keyword stats
|
|
||||||
- `data/term_index/daily/YYYY-MM-DD.json`
|
|
||||||
- `data/term_index/term_stats.json`
|
|
||||||
- governance config
|
|
||||||
- `configs/term_cleanup_policy.json`
|
|
||||||
- `configs/term_watchlist.json`
|
|
||||||
- `configs/term_change_log.json`
|
|
||||||
- review bundle builder
|
|
||||||
- `skills/keyword-cleanup-review/scripts/build_review_bundle.py`
|
|
||||||
- accepted-suggestion writer
|
|
||||||
- `scripts/apply_term_suggestions.py`
|
|
||||||
|
|
||||||
Current status:
|
|
||||||
|
|
||||||
- OpenClaw can read the keyword review bundle as maintenance input
|
|
||||||
- accepted suggestions still require explicit human confirmation
|
|
||||||
- the repository can write accepted watch / alias / stopword / interest-keyword changes after confirmation
|
|
||||||
- this governance flow is not yet wired into a periodic scheduler inside the repository
|
|
||||||
|
|
||||||
Boundary:
|
|
||||||
|
|
||||||
- keyword cleanup is a maintenance flow, not the production RSS ingestion path
|
|
||||||
- the repository does not auto-apply cleanup suggestions without confirmation
|
|
||||||
- current keyword stats are built from the delivered candidate payload, not yet from a final `DailyDigest`
|
|
||||||
|
|
||||||
## Known Limitations
|
|
||||||
|
|
||||||
- Some sources put only partial content in RSS; those items may be skipped if RSS content is insufficient.
|
|
||||||
- WeChat articles often block direct crawling, but this pipeline now avoids that path for FreshRSS items and uses RSS-provided content when available.
|
|
||||||
- Rule behavior is still conservative in some cases; many items may land in `review` depending on current rules.
|
|
||||||
- Paywall heuristics may produce false positives for some Chinese text patterns.
|
|
||||||
- Keyword cleanup governance is usable now, but periodic scheduling and before/after evaluation are not implemented yet.
|
|
||||||
|
|
||||||
## Files OpenClaw Should Read First
|
|
||||||
|
|
||||||
Recommended reading order for a new maintainer:
|
|
||||||
|
|
||||||
1. `README.md`
|
|
||||||
2. `docs/openclaw/openclaw-handoff.md`
|
|
||||||
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
|
||||||
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
|
|
||||||
5. `docs/design/daily-keyword-index-design.md`
|
|
||||||
6. `skills/keyword-cleanup-review/SKILL.md`
|
|
||||||
7. `docs/current/context-reset-brief.md`
|
|
||||||
|
|
||||||
## Downstream Boundary Rules
|
|
||||||
|
|
||||||
For the daily-digest workflow:
|
|
||||||
|
|
||||||
- the digest should go to Hugo and chat reporting, not directly into IMA
|
|
||||||
- the full daily digest should **not** be uploaded to IMA
|
|
||||||
- only explicitly user-selected article summaries should be uploaded to IMA
|
- only explicitly user-selected article summaries should be uploaded to IMA
|
||||||
- selected-article summaries should be generated from existing extracted text, not by re-fetching original URLs
|
|
||||||
|
|
||||||
## Current Recommendation
|
## Related Maintenance Flow
|
||||||
|
|
||||||
For integration handoff, the repository is usable now.
|
Keyword cleanup exists as a separate maintenance flow, not the main RSS ingestion path.
|
||||||
|
|
||||||
The minimum you need to give OpenClaw is:
|
Relevant files:
|
||||||
|
|
||||||
- the repository code
|
|
||||||
- the MCP server startup command
|
|
||||||
- the required environment variables in the target environment
|
|
||||||
- the instruction to call `run_freshrss_openclaw_pipeline`
|
|
||||||
- the rule that follow-up state/result reads must go through MCP tools first, not handwritten filesystem paths
|
|
||||||
|
|
||||||
If OpenClaw will also participate in keyword-governance review, additionally point it to:
|
|
||||||
|
|
||||||
- `docs/design/daily-keyword-index-design.md`
|
- `docs/design/daily-keyword-index-design.md`
|
||||||
- `skills/keyword-cleanup-review/SKILL.md`
|
- `skills/keyword-cleanup-review/SKILL.md`
|
||||||
- `scripts/apply_term_suggestions.py`
|
- `scripts/apply_term_suggestions.py`
|
||||||
|
|
||||||
|
## Known Limitations
|
||||||
|
|
||||||
|
- some sources expose only partial RSS content; those items may be skipped
|
||||||
|
- rule behavior is still conservative; many items may land in `review`
|
||||||
|
- paywall heuristics may still produce false positives on some Chinese text
|
||||||
|
- keyword cleanup governance is usable but not yet wired to periodic scheduling
|
||||||
|
|
||||||
|
## Read First
|
||||||
|
|
||||||
|
Recommended reading order for a new maintainer:
|
||||||
|
|
||||||
|
1. `README.md`
|
||||||
|
2. `docs/openclaw/README.md`
|
||||||
|
3. `docs/openclaw/openclaw-handoff.md`
|
||||||
|
4. `docs/openclaw/openclaw-orchestration-flow.md`
|
||||||
|
5. `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||||
|
6. `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||||
|
7. `docs/current/context-reset-brief.md`
|
||||||
|
|||||||
@@ -2,79 +2,86 @@
|
|||||||
|
|
||||||
## 1. 文档目的
|
## 1. 文档目的
|
||||||
|
|
||||||
本文档定义 OpenClaw 在正式环境中如何调用 reader 作为上游 MCP workflow service。
|
本文档只回答一个问题:OpenClaw 在正式环境里应该如何编排 reader。
|
||||||
|
|
||||||
目标不是描述 reader 内部实现,而是明确 OpenClaw 的编排动作:
|
这里不重复介绍 reader 内部实现,只定义正式控制面:
|
||||||
|
|
||||||
- 什么时候启动新 run
|
- 如何启动日报
|
||||||
- 什么时候查询状态
|
- 如何轮询 job
|
||||||
- 什么时候读取结果
|
- 如何读取 run 结果
|
||||||
- 什么时候尝试恢复
|
- 如何判断是否恢复
|
||||||
- 什么时候直接新开 run
|
- 如何走异步恢复
|
||||||
- 什么时候需要人工介入
|
- 什么时候直接新开 run 或人工介入
|
||||||
|
|
||||||
本文档基于 reader 当前**已真实落地**的能力编写,不描述尚未实现的未来接口。
|
## 2. 当前正式入口
|
||||||
|
|
||||||
---
|
### 2.1 新 run
|
||||||
|
|
||||||
## 2. 当前 reader 已正式支持的 MCP 能力
|
正式生产入口:
|
||||||
|
|
||||||
当前可用能力:
|
- `start_freshrss_pipeline_job`
|
||||||
|
- `get_freshrss_pipeline_job_status`
|
||||||
|
- `get_freshrss_pipeline_job_result`
|
||||||
|
|
||||||
|
同步入口:
|
||||||
|
|
||||||
- `run_freshrss_openclaw_pipeline`
|
- `run_freshrss_openclaw_pipeline`
|
||||||
|
|
||||||
|
同步入口只保留给 debug / fallback,不再是正式编排默认路径。
|
||||||
|
|
||||||
|
### 2.2 run 级读取
|
||||||
|
|
||||||
|
正式 run 级读取接口:
|
||||||
|
|
||||||
- `get_run_status`
|
- `get_run_status`
|
||||||
- `list_runs`
|
- `list_runs`
|
||||||
- `list_run_artifacts`
|
- `list_run_artifacts`
|
||||||
- `get_delivery_payload`
|
- `get_delivery_payload`
|
||||||
- `get_run_report`
|
- `get_run_report`
|
||||||
|
|
||||||
|
### 2.3 恢复
|
||||||
|
|
||||||
|
正式恢复入口:
|
||||||
|
|
||||||
|
- `inspect_resume_plan`
|
||||||
|
- `start_resume_job`
|
||||||
|
- `get_resume_job_status`
|
||||||
|
- `get_resume_job_result`
|
||||||
|
|
||||||
|
同步恢复入口:
|
||||||
|
|
||||||
- `resume_run`
|
- `resume_run`
|
||||||
|
|
||||||
其中:
|
`resume_run` 只保留给 debug / fallback。
|
||||||
|
|
||||||
- `run_freshrss_openclaw_pipeline` 是当前正式启动入口
|
|
||||||
- `get_run_status` / `list_runs` / `list_run_artifacts` 用于观测
|
|
||||||
- `get_delivery_payload` / `get_run_report` 用于读取正式结果
|
|
||||||
- `resume_run` 用于最小恢复能力
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 3. 编排基本原则
|
## 3. 编排基本原则
|
||||||
|
|
||||||
### 3.1 OpenClaw 不再手拼路径
|
### 3.1 OpenClaw 不手拼路径
|
||||||
|
|
||||||
OpenClaw 不应再自己拼 reader 输出路径来判断运行状态或读取核心结果。
|
OpenClaw 不应自己推导这些路径:
|
||||||
|
|
||||||
优先使用 MCP:
|
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
|
||||||
|
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
|
||||||
|
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
|
||||||
|
|
||||||
- 查状态 → `get_run_status`
|
需要路径时,只消费 MCP 返回值:
|
||||||
- 读 payload → `get_delivery_payload`
|
|
||||||
- 读 report → `get_run_report`
|
|
||||||
- 做恢复 → `resume_run`
|
|
||||||
|
|
||||||
只有在排障/人工核查时,才回退到直接看 reader run 目录。
|
- `output_dir`
|
||||||
|
- `artifact.path`
|
||||||
|
- `delivery_output`
|
||||||
|
- `report_output`
|
||||||
|
|
||||||
### 3.2 reader 是上游 workflow engine
|
### 3.2 顶层 `status` 才是分支依据
|
||||||
|
|
||||||
reader 负责:
|
`get_run_status` 和 job status 接口都可能做状态收敛。
|
||||||
|
|
||||||
- FreshRSS 拉取
|
因此:
|
||||||
- 内容提取
|
|
||||||
- 摘要
|
|
||||||
- 过滤
|
|
||||||
- payload 生成
|
|
||||||
- run 状态记录
|
|
||||||
- 最小恢复
|
|
||||||
|
|
||||||
OpenClaw 负责:
|
- 优先使用顶层 `status`
|
||||||
|
- `status_source` 用来解释状态来自原始 state 还是收敛结果
|
||||||
|
- `state_conflict=true` 说明底层状态文件已经落后于真实产物
|
||||||
|
|
||||||
- 触发执行
|
不要再拿旧的 `raw_status`、`raw_current_stage` 或早期阶段名重新做分支。
|
||||||
- 轮询状态
|
|
||||||
- 读取结果
|
|
||||||
- 生成 digest markdown
|
|
||||||
- Hugo 发布
|
|
||||||
- 聊天汇报
|
|
||||||
- 用户确认精选
|
|
||||||
- IMA 编排
|
|
||||||
|
|
||||||
### 3.3 默认生产语义
|
### 3.3 默认生产语义
|
||||||
|
|
||||||
@@ -82,19 +89,18 @@ OpenClaw 负责:
|
|||||||
|
|
||||||
- `mark_read=true`
|
- `mark_read=true`
|
||||||
- `debug_artifacts=false`
|
- `debug_artifacts=false`
|
||||||
- 只在 debug/test/validation 时显式放宽
|
|
||||||
|
|
||||||
---
|
只有 debug / test / validation 时才放宽。
|
||||||
|
|
||||||
## 4. 标准 Happy Path
|
## 4. 标准 Happy Path
|
||||||
|
|
||||||
### Step 1: 启动新 run
|
### Step 1: 启动新 job
|
||||||
|
|
||||||
调用:
|
调用:
|
||||||
|
|
||||||
- `run_freshrss_openclaw_pipeline`
|
- `start_freshrss_pipeline_job`
|
||||||
|
|
||||||
推荐参数示例:
|
推荐参数:
|
||||||
|
|
||||||
```json
|
```json
|
||||||
{
|
{
|
||||||
@@ -107,131 +113,156 @@ OpenClaw 负责:
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
期望:
|
预期:
|
||||||
|
|
||||||
- 获得 `run_id`
|
- 立即返回 `job_id`
|
||||||
- 获得 `output_dir`
|
- 后续由 OpenClaw 轮询 job,而不是同步等待整条流水线
|
||||||
- 获得初始结果摘要
|
|
||||||
|
|
||||||
如果启动阶段直接抛错:
|
### Step 2: 轮询 job
|
||||||
|
|
||||||
- 直接判为启动失败
|
调用:
|
||||||
- 不进入后续查询
|
|
||||||
|
|
||||||
### Step 2: 查询运行状态
|
- `get_freshrss_pipeline_job_status(job_id=...)`
|
||||||
|
|
||||||
|
根据返回:
|
||||||
|
|
||||||
|
- `status=running`:继续轮询
|
||||||
|
- `status=success`:读取 job result
|
||||||
|
- `status=failed`:进入失败处理
|
||||||
|
|
||||||
|
额外规则:
|
||||||
|
|
||||||
|
- 如果 `status_source=linked_run_reconciliation`,说明 outer job state 已落后,但 linked run 已经给出可用终态
|
||||||
|
- 如果 `status_source=stale_job_state_timeout`,把它当成终态失败,不要继续无限轮询
|
||||||
|
|
||||||
|
### Step 3: 读取 job result
|
||||||
|
|
||||||
|
调用:
|
||||||
|
|
||||||
|
- `get_freshrss_pipeline_job_result(job_id=...)`
|
||||||
|
|
||||||
|
预期读取:
|
||||||
|
|
||||||
|
- `run_id`
|
||||||
|
- `output_dir`
|
||||||
|
- `delivery_output`
|
||||||
|
- `report_output`
|
||||||
|
|
||||||
|
从这一刻开始,`run_id` 是正式的稳定句柄。
|
||||||
|
|
||||||
|
### Step 4: 读取 run 级状态与结果
|
||||||
|
|
||||||
调用:
|
调用:
|
||||||
|
|
||||||
- `get_run_status(run_id=...)`
|
- `get_run_status(run_id=...)`
|
||||||
|
|
||||||
根据返回:
|
|
||||||
|
|
||||||
- `status=running` → 继续轮询
|
|
||||||
- `status=success` → 进入结果读取
|
|
||||||
- `status=failed` → 进入失败处理
|
|
||||||
- `status=partial` → 视为未完成,优先看 `recovery` 和当前阶段
|
|
||||||
|
|
||||||
### Step 3: 读取正式结果
|
|
||||||
|
|
||||||
成功后读取:
|
|
||||||
|
|
||||||
- `get_delivery_payload(run_id=...)`
|
- `get_delivery_payload(run_id=...)`
|
||||||
- `get_run_report(run_id=...)`
|
- `get_run_report(run_id=...)`
|
||||||
|
|
||||||
后续 OpenClaw 编排应以这两个接口为正式结果源,而不是自己拼路径读取 JSON。
|
根据 `get_run_status`:
|
||||||
|
|
||||||
### Step 4: 进入下游编排
|
- `status=running`:继续观察
|
||||||
|
- `status=success`:继续下游 digest / 发布 / 汇报
|
||||||
|
- `status=failed`:进入恢复或重跑决策
|
||||||
|
- `status=partial`:优先检查 report、artifacts 和 recovery
|
||||||
|
|
||||||
OpenClaw 在拿到正式 payload / report 后,继续执行:
|
如果 `status_source=run_report_reconciliation`,说明 `run-state.json` 已经过期,但 reader 已经根据终态产物收敛出有效状态。
|
||||||
|
|
||||||
- public/internal digest 生成
|
如果 `status_source=stale_run_state_timeout`,说明 reader 认为该 run 长时间未收敛且没有终态产物,应按失败处理。
|
||||||
- Hugo 发布
|
|
||||||
- 聊天汇报
|
|
||||||
- 用户确认精选
|
|
||||||
- IMA 沉淀
|
|
||||||
|
|
||||||
---
|
## 5. 恢复决策
|
||||||
|
|
||||||
## 5. 状态 → 动作映射
|
### 5.1 先看预检,不要直接恢复
|
||||||
|
|
||||||
| reader 状态 | OpenClaw 动作 |
|
恢复前固定动作:
|
||||||
|---|---|
|
|
||||||
| `running` | 继续轮询 `get_run_status` |
|
|
||||||
| `success` | 读取 `get_delivery_payload` 和 `get_run_report` |
|
|
||||||
| `failed` 且 `recovery.resumable=true` | 评估是否调用 `resume_run` |
|
|
||||||
| `failed` 且 `recovery.resumable=false` | 直接判失败,通常新开 run 或人工介入 |
|
|
||||||
| `partial` | 先读状态详情和 recovery,再决定继续等 / 恢复 / 人工介入 |
|
|
||||||
|
|
||||||
---
|
- 先调用 `inspect_resume_plan(run_id)`
|
||||||
|
|
||||||
## 6. 失败处理与恢复决策
|
只在以下条件同时成立时才启动恢复:
|
||||||
|
|
||||||
### 6.1 什么时候优先尝试 `resume_run`
|
- `can_resume=true`
|
||||||
|
- `recommended_action=resume`
|
||||||
|
|
||||||
满足以下条件时,优先考虑恢复而不是新开 run:
|
重点字段:
|
||||||
|
|
||||||
- `get_run_status` 返回 `failed`
|
- `requested_resume_from_stage`
|
||||||
- `recovery.resumable=true`
|
- `resume_from_stage`
|
||||||
- 当前 run 对应的是 freshrss workflow
|
- `resume_decision_source`
|
||||||
- 当前失败点在 reader 第一版支持的恢复范围内
|
- `artifact_resume_from_stage`
|
||||||
|
- `artifact_snapshot`
|
||||||
|
|
||||||
### 6.2 `resume_run` 当前支持范围
|
### 5.2 正式恢复路径
|
||||||
|
|
||||||
当前最小实现仅支持:
|
正式恢复控制面:
|
||||||
|
|
||||||
- 仅对带 `run-state.json` 的 freshrss run
|
1. `start_resume_job(run_id)`
|
||||||
- 仅从最近可恢复点继续
|
2. `get_resume_job_status(job_id)`
|
||||||
- 支持的恢复点:
|
3. `get_resume_job_result(job_id)`
|
||||||
|
|
||||||
|
不要再把同步 `resume_run(run_id)` 当成正式恢复入口。
|
||||||
|
|
||||||
|
### 5.3 当前支持范围
|
||||||
|
|
||||||
|
当前只支持:
|
||||||
|
|
||||||
|
- 带有效 `run-state.json` 的 `freshrss_daily_digest` run
|
||||||
|
- 从以下阶段恢复:
|
||||||
- `generate_summaries`
|
- `generate_summaries`
|
||||||
- `apply_filters`
|
- `apply_filters`
|
||||||
- `build_delivery_payload`
|
- `build_delivery_payload`
|
||||||
- `write_run_report`
|
- `write_run_report`
|
||||||
|
|
||||||
明确不支持:
|
当前不支持:
|
||||||
|
|
||||||
- `fetch_feed`
|
- `fetch_feed`
|
||||||
- `extract_articles`
|
- `extract_articles`
|
||||||
|
|
||||||
### 6.3 什么时候不要恢复,直接新开 run
|
正式生产恢复优先依赖:
|
||||||
|
|
||||||
以下情况不建议 `resume_run`:
|
- `summary/summary-batch.json`
|
||||||
|
- `candidates/candidate-batch.json`
|
||||||
|
|
||||||
- `recovery.resumable=false`
|
### 5.4 什么时候不要恢复
|
||||||
- run 没有 `run-state.json`
|
|
||||||
- 失败点是 `fetch_feed` 或 `extract_articles`
|
|
||||||
- 恢复所需关键产物缺失
|
|
||||||
- 恢复点语义不明确或结果存在明显漂移风险
|
|
||||||
|
|
||||||
这时更合理的动作通常是:
|
以下情况直接新开 run 更合理:
|
||||||
|
|
||||||
- 直接新开 run
|
- `recommended_action=start_new_run`
|
||||||
- 或人工介入排查
|
- `recommended_action=read_terminal_result`
|
||||||
|
- 没有有效 `run-state.json`
|
||||||
|
- 恢复所需关键 artifacts 缺失
|
||||||
|
- 连续恢复失败
|
||||||
|
|
||||||
### 6.4 什么时候需要人工介入
|
## 6. 状态到动作映射
|
||||||
|
|
||||||
出现以下任一情况时,建议人工介入:
|
| 接口 | 状态 | OpenClaw 动作 |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `get_freshrss_pipeline_job_status` | `running` | 继续轮询 job |
|
||||||
|
| `get_freshrss_pipeline_job_status` | `success` | 读取 `get_freshrss_pipeline_job_result` |
|
||||||
|
| `get_freshrss_pipeline_job_status` | `failed` | 结束本次 job,必要时读 linked run |
|
||||||
|
| `get_run_status` | `running` | 继续观察 run |
|
||||||
|
| `get_run_status` | `success` | 读取 `get_delivery_payload` / `get_run_report` |
|
||||||
|
| `get_run_status` | `failed` | 先看 `inspect_resume_plan` |
|
||||||
|
| `inspect_resume_plan` | `recommended_action=resume` | 启动 `start_resume_job` |
|
||||||
|
| `inspect_resume_plan` | `recommended_action=read_terminal_result` | 直接读 run 结果,不恢复 |
|
||||||
|
| `inspect_resume_plan` | `recommended_action=start_new_run` | 新开 run 或人工介入 |
|
||||||
|
|
||||||
|
## 7. 人工介入条件
|
||||||
|
|
||||||
|
出现以下任一情况时,建议不要自动编排:
|
||||||
|
|
||||||
- 连续恢复失败
|
- 连续恢复失败
|
||||||
- `get_run_status` 与实际产物明显不一致
|
- payload / report 结构不符合预期
|
||||||
- payload/report 结构不符合预期
|
- `get_run_status` 与实际产物长期明显冲突
|
||||||
- 恢复依赖的关键文件缺失且原因不明
|
- FreshRSS、LLM 或外部依赖异常
|
||||||
- FreshRSS / LLM / 外部环境异常
|
- 恢复判定结果和编排预期不一致
|
||||||
|
|
||||||
---
|
## 8. 结论
|
||||||
|
|
||||||
## 7. 读取结果的标准动作
|
当前 OpenClaw 的正式调用方式已经收口为两条异步控制面:
|
||||||
|
|
||||||
### 7.1 `get_delivery_payload`
|
- 主日报:`start_freshrss_pipeline_job -> poll -> get result -> run reads`
|
||||||
|
- 恢复:`inspect_resume_plan -> start_resume_job -> poll -> get result`
|
||||||
|
|
||||||
用途:
|
同步 `run_freshrss_openclaw_pipeline` 和 `resume_run` 仅用于 debug / fallback,不应再作为默认正式编排路径。
|
||||||
|
|
||||||
- 获取正式交付给 OpenClaw 的 payload
|
|
||||||
- 后续 digest 生成应以该返回为准
|
|
||||||
|
|
||||||
OpenClaw 应做:
|
|
||||||
|
|
||||||
- 读取后直接进入 digest 生成
|
|
||||||
- 不再自己拼 `candidates/openclaw-delivery-payload.json`
|
|
||||||
|
|
||||||
### 7.2 `get_run_report`
|
### 7.2 `get_run_report`
|
||||||
|
|
||||||
@@ -263,7 +294,9 @@ OpenClaw 应做:
|
|||||||
|
|
||||||
1. 调 `get_run_status`
|
1. 调 `get_run_status`
|
||||||
2. 若 `failed && recovery.resumable=true`:
|
2. 若 `failed && recovery.resumable=true`:
|
||||||
- 调 `resume_run`
|
- 调 `start_resume_job`
|
||||||
|
- 轮询 `get_resume_job_status`
|
||||||
|
- 读取 `get_resume_job_result`
|
||||||
3. 恢复后再次:
|
3. 恢复后再次:
|
||||||
- 调 `get_run_status`
|
- 调 `get_run_status`
|
||||||
- 若成功,再读 payload / report
|
- 若成功,再读 payload / report
|
||||||
@@ -279,7 +312,7 @@ OpenClaw 应做:
|
|||||||
- 让 OpenClaw 直接长时间 `exec` reader CLI 作为主要生产入口
|
- 让 OpenClaw 直接长时间 `exec` reader CLI 作为主要生产入口
|
||||||
- 让 OpenClaw 自己拼 reader 输出路径来判断成功/失败
|
- 让 OpenClaw 自己拼 reader 输出路径来判断成功/失败
|
||||||
- 让 OpenClaw 自己读取 `outputs/.../*.json` 作为正式结果源
|
- 让 OpenClaw 自己读取 `outputs/.../*.json` 作为正式结果源
|
||||||
- 在未确认 `resume_run` 支持范围外的失败点上强行恢复
|
- 在未确认恢复支持范围外的失败点上强行恢复
|
||||||
|
|
||||||
CLI 现在的定位是:
|
CLI 现在的定位是:
|
||||||
|
|
||||||
@@ -293,7 +326,7 @@ CLI 现在的定位是:
|
|||||||
|
|
||||||
## 10. 当前已知局限
|
## 10. 当前已知局限
|
||||||
|
|
||||||
- `resume_run` 仍是最小实现,不支持任意 stage 任意重入
|
- 恢复能力仍是最小实现,不支持任意 stage 任意重入
|
||||||
- 历史无 `run-state.json` 的 run 不支持正式恢复
|
- 历史无 `run-state.json` 的 run 不支持正式恢复
|
||||||
- 极旧 run 的结果读取仍可能依赖保守目录扫描
|
- 极旧 run 的结果读取仍可能依赖保守目录扫描
|
||||||
- `write_run_report` 若涉及重新 `mark_read`,仍依赖 FreshRSS 环境和可用凭据
|
- `write_run_report` 若涉及重新 `mark_read`,仍依赖 FreshRSS 环境和可用凭据
|
||||||
@@ -304,4 +337,4 @@ CLI 现在的定位是:
|
|||||||
|
|
||||||
OpenClaw 当前应把 reader 当作正式 MCP workflow service 使用:
|
OpenClaw 当前应把 reader 当作正式 MCP workflow service 使用:
|
||||||
|
|
||||||
**启动用 `run_freshrss_openclaw_pipeline`,观测用 `get_run_status`,结果读取用 `get_delivery_payload` / `get_run_report`,恢复仅在 `resume_run` 最小支持范围内启用;不要再把 reader 当成长 CLI 任务和路径拼接仓库来驱动。**
|
**启动用 `start_freshrss_pipeline_job`,观测用 `get_run_status`,结果读取用 `get_delivery_payload` / `get_run_report`,恢复默认用 `inspect_resume_plan` + `start_resume_job`,不要再把 reader 当成长 CLI 任务和路径拼接仓库来驱动。**
|
||||||
|
|||||||
@@ -0,0 +1,73 @@
|
|||||||
|
# 规划文档导航
|
||||||
|
|
||||||
|
## 使用原则
|
||||||
|
|
||||||
|
`plans/` 目录保存架构设计、实施计划、专题方案和历史问题分析。
|
||||||
|
|
||||||
|
不要把所有 plan 都当成当前权威事实。
|
||||||
|
当前事实优先级应是:
|
||||||
|
|
||||||
|
1. `README.md`
|
||||||
|
2. `docs/README.md`
|
||||||
|
3. `docs/openclaw/README.md`
|
||||||
|
4. `docs/openclaw/openclaw-handoff.md`
|
||||||
|
5. `docs/openclaw/openclaw-orchestration-flow.md`
|
||||||
|
6. `TODO.md`
|
||||||
|
|
||||||
|
`plans/` 更适合回答:
|
||||||
|
|
||||||
|
- 为什么这样设计
|
||||||
|
- 某个能力是怎么分阶段落地的
|
||||||
|
- 某次事故当时是怎么分析的
|
||||||
|
|
||||||
|
## 当前权威规划
|
||||||
|
|
||||||
|
这些文档仍然是当前协作时应优先阅读的规划基线:
|
||||||
|
|
||||||
|
- `reader-mcp-architecture-design.md`
|
||||||
|
- MCP workflow service 的总体架构方向
|
||||||
|
- `reader-mcp-implementation-plan.md`
|
||||||
|
- 实施分阶段计划
|
||||||
|
- `../TODO.md`
|
||||||
|
- 当前任务状态与落地进展
|
||||||
|
|
||||||
|
## 已完成能力的专题方案
|
||||||
|
|
||||||
|
这些方案主要用于回看设计取舍,相关能力已经基本落地:
|
||||||
|
|
||||||
|
- `freshrss-pipeline-async-job-plan.md`
|
||||||
|
- 主日报 async job 方案
|
||||||
|
- `article-summary-async-job-plan.md`
|
||||||
|
- 单篇总结 async job 方案
|
||||||
|
- `resume-run-minimal-design.md`
|
||||||
|
- `resume_run` 最小恢复语义设计
|
||||||
|
|
||||||
|
## 仍有参考价值的专题设计
|
||||||
|
|
||||||
|
- `keyword-cleanup-artifact-slimming-v1.md`
|
||||||
|
- `keyword-cleanup-review-suggestions-layer-design.md`
|
||||||
|
- `article-summary-prompt-independent.md`
|
||||||
|
- `article-deep-summary-skill.md`
|
||||||
|
- `docker-deployment-plan.md`
|
||||||
|
|
||||||
|
## 历史问题分析
|
||||||
|
|
||||||
|
- `issues/2026-04-06-reader-digest-sigterm.md`
|
||||||
|
- 一次真实运行事故的分析
|
||||||
|
|
||||||
|
## 建议阅读顺序
|
||||||
|
|
||||||
|
如果是新接手维护:
|
||||||
|
|
||||||
|
1. `reader-mcp-architecture-design.md`
|
||||||
|
2. `reader-mcp-implementation-plan.md`
|
||||||
|
3. `../TODO.md`
|
||||||
|
4. `../docs/openclaw/README.md`
|
||||||
|
5. `../docs/openclaw/openclaw-handoff.md`
|
||||||
|
|
||||||
|
如果是在排查某一类能力:
|
||||||
|
|
||||||
|
- 主日报启动/轮询:看 `freshrss-pipeline-async-job-plan.md`
|
||||||
|
- 恢复:看 `resume-run-minimal-design.md`
|
||||||
|
- 单篇总结:看 `article-summary-async-job-plan.md`
|
||||||
|
- 历史故障:看 `issues/2026-04-06-reader-digest-sigterm.md`
|
||||||
@@ -0,0 +1,24 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import sys
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||||
|
SRC_ROOT = REPO_ROOT / "src"
|
||||||
|
|
||||||
|
if str(SRC_ROOT) not in sys.path:
|
||||||
|
sys.path.insert(0, str(SRC_ROOT))
|
||||||
|
|
||||||
|
from summary_mcp.runtime.resume_jobs import run_resume_job
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
parser = argparse.ArgumentParser(description="Run a background resume job by job_id.")
|
||||||
|
parser.add_argument("--job-id", required=True, help="Resume job id")
|
||||||
|
args = parser.parse_args()
|
||||||
|
run_resume_job(job_id=args.job_id)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -1,17 +1,21 @@
|
|||||||
|
from .run_store import RunStore
|
||||||
|
from .state_models import ArtifactRecord, RecoveryState, RunError, RunState, StageState
|
||||||
from .freshrss_pipeline_jobs import (
|
from .freshrss_pipeline_jobs import (
|
||||||
get_freshrss_pipeline_job_result,
|
get_freshrss_pipeline_job_result,
|
||||||
get_freshrss_pipeline_job_status,
|
get_freshrss_pipeline_job_status,
|
||||||
start_freshrss_pipeline_job,
|
start_freshrss_pipeline_job,
|
||||||
)
|
)
|
||||||
from .query_service import get_delivery_payload, get_run_report, get_run_status, list_run_artifacts, list_runs
|
from .query_service import get_delivery_payload, get_run_report, get_run_status, list_run_artifacts, list_runs
|
||||||
from .run_store import RunStore
|
from .resume_jobs import get_resume_job_result, get_resume_job_status, start_resume_job
|
||||||
from .state_models import ArtifactRecord, RecoveryState, RunError, RunState, StageState
|
from .resume_service import inspect_resume_plan, resume_run
|
||||||
|
|
||||||
__all__ = [
|
__all__ = [
|
||||||
"ArtifactRecord",
|
"ArtifactRecord",
|
||||||
"get_delivery_payload",
|
"get_delivery_payload",
|
||||||
"get_freshrss_pipeline_job_result",
|
"get_freshrss_pipeline_job_result",
|
||||||
"get_freshrss_pipeline_job_status",
|
"get_freshrss_pipeline_job_status",
|
||||||
|
"get_resume_job_result",
|
||||||
|
"get_resume_job_status",
|
||||||
"get_run_report",
|
"get_run_report",
|
||||||
"RecoveryState",
|
"RecoveryState",
|
||||||
"RunError",
|
"RunError",
|
||||||
@@ -19,7 +23,10 @@ __all__ = [
|
|||||||
"RunStore",
|
"RunStore",
|
||||||
"StageState",
|
"StageState",
|
||||||
"get_run_status",
|
"get_run_status",
|
||||||
|
"inspect_resume_plan",
|
||||||
"list_run_artifacts",
|
"list_run_artifacts",
|
||||||
"list_runs",
|
"list_runs",
|
||||||
|
"resume_run",
|
||||||
|
"start_resume_job",
|
||||||
"start_freshrss_pipeline_job",
|
"start_freshrss_pipeline_job",
|
||||||
]
|
]
|
||||||
|
|||||||
@@ -10,6 +10,7 @@ from typing import Any
|
|||||||
from uuid import uuid4
|
from uuid import uuid4
|
||||||
|
|
||||||
from .run_store import RunStore
|
from .run_store import RunStore
|
||||||
|
from .query_service import _resolve_run_record
|
||||||
|
|
||||||
REPO_ROOT = Path(__file__).resolve().parents[3]
|
REPO_ROOT = Path(__file__).resolve().parents[3]
|
||||||
OUTPUT_ROOT = REPO_ROOT / "outputs" / "freshrss"
|
OUTPUT_ROOT = REPO_ROOT / "outputs" / "freshrss"
|
||||||
@@ -26,6 +27,8 @@ DEFAULT_STAGES = [
|
|||||||
"run_pipeline",
|
"run_pipeline",
|
||||||
"write_result",
|
"write_result",
|
||||||
]
|
]
|
||||||
|
MIN_JOB_STALE_SECONDS = 30 * 60
|
||||||
|
MAX_JOB_STALE_SECONDS = 6 * 60 * 60
|
||||||
|
|
||||||
|
|
||||||
def _now() -> datetime:
|
def _now() -> datetime:
|
||||||
@@ -369,33 +372,34 @@ def run_freshrss_pipeline_job(*, job_id: str) -> dict[str, Any]:
|
|||||||
def get_freshrss_pipeline_job_status(*, job_id: str) -> dict[str, Any]:
|
def get_freshrss_pipeline_job_status(*, job_id: str) -> dict[str, Any]:
|
||||||
store = _load_run_store(job_id)
|
store = _load_run_store(job_id)
|
||||||
state = store.state
|
state = store.state
|
||||||
completed_stage_count = sum(1 for s in state.stages if s.status == "success")
|
effective = _resolve_effective_job_view(job_id=job_id, state=state)
|
||||||
running_stage_count = sum(1 for s in state.stages if s.status == "running")
|
progress = _build_effective_job_progress(state=state, effective_status=effective["status"])
|
||||||
failed_stage_count = sum(1 for s in state.stages if s.status == "failed")
|
|
||||||
pending_stage_count = sum(1 for s in state.stages if s.status == "pending")
|
|
||||||
linked_run_id = state.input.get("run_id") if isinstance(state.input, dict) else None
|
linked_run_id = state.input.get("run_id") if isinstance(state.input, dict) else None
|
||||||
linked_output_dir = _normalize_repo_path_value(state.input.get("output_dir")) if isinstance(state.input, dict) else None
|
linked_output_dir = _normalize_repo_path_value(state.input.get("output_dir")) if isinstance(state.input, dict) else None
|
||||||
return {
|
return {
|
||||||
"job_id": state.run_id,
|
"job_id": state.run_id,
|
||||||
"workflow": state.workflow,
|
"workflow": state.workflow,
|
||||||
"run_type": state.run_type,
|
"run_type": state.run_type,
|
||||||
"status": state.status,
|
"status": effective["status"],
|
||||||
"current_stage": state.current_stage,
|
"current_stage": effective["current_stage"],
|
||||||
"started_at": state.started_at.isoformat(),
|
"started_at": state.started_at.isoformat(),
|
||||||
"updated_at": state.updated_at.isoformat(),
|
"updated_at": effective["updated_at"],
|
||||||
"finished_at": state.finished_at.isoformat() if state.finished_at else None,
|
"finished_at": effective["finished_at"],
|
||||||
"output_dir": _normalize_repo_path(_job_dir(job_id)),
|
"output_dir": _normalize_repo_path(_job_dir(job_id)),
|
||||||
"linked_run_id": linked_run_id,
|
"linked_run_id": linked_run_id,
|
||||||
"linked_output_dir": linked_output_dir,
|
"linked_output_dir": linked_output_dir,
|
||||||
"progress": {
|
"progress": progress,
|
||||||
"completed_stage_count": completed_stage_count,
|
|
||||||
"running_stage_count": running_stage_count,
|
|
||||||
"failed_stage_count": failed_stage_count,
|
|
||||||
"pending_stage_count": pending_stage_count,
|
|
||||||
"total_stage_count": len(state.stages),
|
|
||||||
},
|
|
||||||
"artifacts": [artifact.model_dump(mode="json") for artifact in state.artifacts],
|
"artifacts": [artifact.model_dump(mode="json") for artifact in state.artifacts],
|
||||||
"error_summary": state.error.model_dump(mode="json") if state.error else None,
|
"error_summary": state.error.model_dump(mode="json") if state.error else None,
|
||||||
|
"status_source": effective["status_source"],
|
||||||
|
"state_quality": effective["state_quality"],
|
||||||
|
"state_conflict": effective["state_conflict"],
|
||||||
|
"state_conflict_reason": effective["state_conflict_reason"],
|
||||||
|
"status_note": effective["status_note"],
|
||||||
|
"raw_status": state.status,
|
||||||
|
"raw_current_stage": state.current_stage,
|
||||||
|
"linked_run_status": effective["linked_run_status"],
|
||||||
|
"linked_run_state_conflict": effective["linked_run_state_conflict"],
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
@@ -403,30 +407,244 @@ def get_freshrss_pipeline_job_result(*, job_id: str) -> dict[str, Any]:
|
|||||||
store = _load_run_store(job_id)
|
store = _load_run_store(job_id)
|
||||||
state = store.state
|
state = store.state
|
||||||
result = _load_result(job_id)
|
result = _load_result(job_id)
|
||||||
if state.status != "success" or result is None:
|
effective = _resolve_effective_job_view(job_id=job_id, state=state, result=result)
|
||||||
|
synthesized_result = result or effective["result"]
|
||||||
|
if effective["status"] != "success" or synthesized_result is None:
|
||||||
|
message = "FreshRSS pipeline job result is not ready."
|
||||||
|
if effective["status"] == "failed":
|
||||||
|
message = "FreshRSS pipeline job did not complete successfully, so no terminal job result is available."
|
||||||
return {
|
return {
|
||||||
"job_id": state.run_id,
|
"job_id": state.run_id,
|
||||||
"status": state.status,
|
"status": effective["status"],
|
||||||
"message": "FreshRSS pipeline job result is not ready.",
|
"message": message,
|
||||||
"linked_run_id": state.input.get("run_id") if isinstance(state.input, dict) else None,
|
"linked_run_id": state.input.get("run_id") if isinstance(state.input, dict) else None,
|
||||||
"error_summary": state.error.model_dump(mode="json") if state.error else None,
|
"error_summary": state.error.model_dump(mode="json") if state.error else None,
|
||||||
|
"status_source": effective["status_source"],
|
||||||
|
"status_note": effective["status_note"],
|
||||||
}
|
}
|
||||||
|
|
||||||
artifact = next((a.model_dump(mode="json") for a in state.artifacts if a.name == "job_result"), None)
|
artifact = next((a.model_dump(mode="json") for a in state.artifacts if a.name == "job_result"), None)
|
||||||
return {
|
return {
|
||||||
"job_id": state.run_id,
|
"job_id": state.run_id,
|
||||||
"status": state.status,
|
"status": effective["status"],
|
||||||
"run_id": result.get("run_id"),
|
"run_id": synthesized_result.get("run_id"),
|
||||||
"output_dir": result.get("output_dir"),
|
"output_dir": synthesized_result.get("output_dir"),
|
||||||
"raw_output": result.get("raw_output"),
|
"raw_output": synthesized_result.get("raw_output"),
|
||||||
"delivery_output": result.get("delivery_output"),
|
"delivery_output": synthesized_result.get("delivery_output"),
|
||||||
"digest_brief_output": result.get("digest_brief_output"),
|
"digest_brief_output": synthesized_result.get("digest_brief_output"),
|
||||||
"report_output": result.get("report_output"),
|
"report_output": synthesized_result.get("report_output"),
|
||||||
"pulled_count": result.get("pulled_count"),
|
"pulled_count": synthesized_result.get("pulled_count"),
|
||||||
"delivered_count": result.get("delivered_count"),
|
"delivered_count": synthesized_result.get("delivered_count"),
|
||||||
"marked_read_count": result.get("marked_read_count"),
|
"marked_read_count": synthesized_result.get("marked_read_count"),
|
||||||
"status_counts": result.get("status_counts"),
|
"status_counts": synthesized_result.get("status_counts"),
|
||||||
"keyword_index": result.get("keyword_index"),
|
"keyword_index": synthesized_result.get("keyword_index"),
|
||||||
"artifact": artifact,
|
"artifact": artifact,
|
||||||
"result": result,
|
"result": synthesized_result,
|
||||||
|
"status_source": effective["status_source"],
|
||||||
|
"status_note": effective["status_note"],
|
||||||
|
"result_source": synthesized_result.get("result_source", "job_result"),
|
||||||
|
"linked_run_status": effective["linked_run_status"],
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _resolve_effective_job_view(
|
||||||
|
*,
|
||||||
|
job_id: str,
|
||||||
|
state: Any,
|
||||||
|
result: dict[str, Any] | None = None,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
resolved_result = result or _load_result(job_id)
|
||||||
|
linked_run_record = _load_linked_run_record(state)
|
||||||
|
raw_status = state.status
|
||||||
|
raw_current_stage = state.current_stage
|
||||||
|
effective_status = raw_status
|
||||||
|
effective_current_stage = raw_current_stage
|
||||||
|
effective_updated_at = state.updated_at.isoformat()
|
||||||
|
effective_finished_at = state.finished_at.isoformat() if state.finished_at else None
|
||||||
|
status_source = "job_run_state"
|
||||||
|
state_quality = "trusted"
|
||||||
|
state_conflict = False
|
||||||
|
state_conflict_reason = None
|
||||||
|
status_note = None
|
||||||
|
|
||||||
|
if resolved_result is not None:
|
||||||
|
completed_at = resolved_result.get("completed_at")
|
||||||
|
effective_status = "success"
|
||||||
|
effective_current_stage = None
|
||||||
|
effective_updated_at = completed_at or effective_updated_at
|
||||||
|
effective_finished_at = completed_at or effective_finished_at
|
||||||
|
if raw_status != "success" or raw_current_stage is not None:
|
||||||
|
status_source = "job_result_reconciliation"
|
||||||
|
state_quality = "reconciled"
|
||||||
|
state_conflict = True
|
||||||
|
state_conflict_reason = "result.json already exists, but the job run-state did not converge to success."
|
||||||
|
status_note = "job result exists; the job can be treated as completed."
|
||||||
|
elif linked_run_record is not None:
|
||||||
|
if _linked_run_result_available(linked_run_record):
|
||||||
|
effective_status = "success"
|
||||||
|
effective_current_stage = None
|
||||||
|
effective_updated_at = linked_run_record.get("updated_at") or effective_updated_at
|
||||||
|
effective_finished_at = linked_run_record.get("finished_at") or effective_finished_at
|
||||||
|
status_source = "linked_run_reconciliation"
|
||||||
|
state_quality = "reconciled"
|
||||||
|
state_conflict = raw_status != "success" or raw_current_stage is not None
|
||||||
|
state_conflict_reason = (
|
||||||
|
"The linked run already has terminal artifacts, but the outer job run-state did not converge."
|
||||||
|
)
|
||||||
|
status_note = "linked run artifacts are complete; the job can be treated as completed."
|
||||||
|
resolved_result = _build_result_from_linked_run(job_id=job_id, state=state, linked_run_record=linked_run_record)
|
||||||
|
elif linked_run_record["status"] == "failed" and raw_status == "running":
|
||||||
|
effective_status = "failed"
|
||||||
|
effective_current_stage = None
|
||||||
|
effective_updated_at = linked_run_record.get("updated_at") or effective_updated_at
|
||||||
|
effective_finished_at = linked_run_record.get("finished_at") or effective_finished_at
|
||||||
|
status_source = "linked_run_reconciliation"
|
||||||
|
state_quality = "reconciled"
|
||||||
|
state_conflict = True
|
||||||
|
state_conflict_reason = "The linked run is already failed, but the outer job still reports running."
|
||||||
|
status_note = "linked run failed; the outer job state appears stale."
|
||||||
|
elif raw_status == "running":
|
||||||
|
stale_reason = _build_stale_job_reason(state)
|
||||||
|
if stale_reason is not None:
|
||||||
|
effective_status = "failed"
|
||||||
|
effective_current_stage = raw_current_stage
|
||||||
|
effective_finished_at = effective_updated_at
|
||||||
|
status_source = "stale_job_state_timeout"
|
||||||
|
state_quality = "reconciled"
|
||||||
|
state_conflict = True
|
||||||
|
state_conflict_reason = stale_reason
|
||||||
|
status_note = stale_reason
|
||||||
|
|
||||||
|
return {
|
||||||
|
"status": effective_status,
|
||||||
|
"current_stage": effective_current_stage,
|
||||||
|
"updated_at": effective_updated_at,
|
||||||
|
"finished_at": effective_finished_at,
|
||||||
|
"status_source": status_source,
|
||||||
|
"state_quality": state_quality,
|
||||||
|
"state_conflict": state_conflict,
|
||||||
|
"state_conflict_reason": state_conflict_reason,
|
||||||
|
"status_note": status_note,
|
||||||
|
"linked_run_status": linked_run_record["status"] if linked_run_record is not None else None,
|
||||||
|
"linked_run_state_conflict": linked_run_record.get("state_conflict") if linked_run_record is not None else None,
|
||||||
|
"result": resolved_result,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _build_effective_job_progress(*, state: Any, effective_status: str) -> dict[str, int]:
|
||||||
|
total_stage_count = len(state.stages)
|
||||||
|
if effective_status == "success":
|
||||||
|
return {
|
||||||
|
"completed_stage_count": total_stage_count,
|
||||||
|
"running_stage_count": 0,
|
||||||
|
"failed_stage_count": 0,
|
||||||
|
"pending_stage_count": 0,
|
||||||
|
"total_stage_count": total_stage_count,
|
||||||
|
}
|
||||||
|
if effective_status == "failed" and state.status == "running":
|
||||||
|
completed_stage_count = sum(1 for s in state.stages if s.status == "success")
|
||||||
|
pending_stage_count = max(total_stage_count - completed_stage_count - 1, 0)
|
||||||
|
return {
|
||||||
|
"completed_stage_count": completed_stage_count,
|
||||||
|
"running_stage_count": 0,
|
||||||
|
"failed_stage_count": 1,
|
||||||
|
"pending_stage_count": pending_stage_count,
|
||||||
|
"total_stage_count": total_stage_count,
|
||||||
|
}
|
||||||
|
return {
|
||||||
|
"completed_stage_count": sum(1 for s in state.stages if s.status == "success"),
|
||||||
|
"running_stage_count": sum(1 for s in state.stages if s.status == "running"),
|
||||||
|
"failed_stage_count": sum(1 for s in state.stages if s.status == "failed"),
|
||||||
|
"pending_stage_count": sum(1 for s in state.stages if s.status == "pending"),
|
||||||
|
"total_stage_count": total_stage_count,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _load_linked_run_record(state: Any) -> dict[str, Any] | None:
|
||||||
|
if not isinstance(state.input, dict):
|
||||||
|
return None
|
||||||
|
linked_run_id = state.input.get("run_id")
|
||||||
|
if not isinstance(linked_run_id, str) or not linked_run_id.strip():
|
||||||
|
return None
|
||||||
|
try:
|
||||||
|
return _resolve_run_record(linked_run_id)
|
||||||
|
except FileNotFoundError:
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def _linked_run_result_available(record: dict[str, Any]) -> bool:
|
||||||
|
artifact_presence = record.get("artifact_presence")
|
||||||
|
if not isinstance(artifact_presence, dict):
|
||||||
|
return False
|
||||||
|
return bool(artifact_presence.get("run_report")) and bool(artifact_presence.get("delivery_payload"))
|
||||||
|
|
||||||
|
|
||||||
|
def _build_result_from_linked_run(
|
||||||
|
*,
|
||||||
|
job_id: str,
|
||||||
|
state: Any,
|
||||||
|
linked_run_record: dict[str, Any],
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
report = linked_run_record.get("report")
|
||||||
|
if not isinstance(report, dict):
|
||||||
|
raise RuntimeError("Cannot synthesize job result because linked run-report.json is missing.")
|
||||||
|
|
||||||
|
result = {
|
||||||
|
"job_id": job_id,
|
||||||
|
"run_id": linked_run_record["run_id"],
|
||||||
|
"output_dir": _normalize_repo_path_value(str(linked_run_record["run_dir"])),
|
||||||
|
"raw_output": _normalize_repo_path_value(report.get("raw_output")),
|
||||||
|
"delivery_output": _normalize_repo_path_value(report.get("delivery_output")),
|
||||||
|
"digest_brief_output": _normalize_repo_path_value(report.get("digest_brief_output")),
|
||||||
|
"report_output": _normalize_repo_path(linked_run_record["run_dir"] / "run-report.json"),
|
||||||
|
"keyword_index": _normalize_keyword_index(report.get("keyword_index")),
|
||||||
|
"pulled_count": report.get("pulled_count"),
|
||||||
|
"delivered_count": report.get("delivered_count"),
|
||||||
|
"marked_read_count": report.get("marked_read_count"),
|
||||||
|
"status_counts": report.get("status_counts"),
|
||||||
|
"debug_artifacts": report.get("debug_artifacts"),
|
||||||
|
"completed_at": report.get("completed_at") or linked_run_record.get("finished_at"),
|
||||||
|
"result_source": "linked_run_report",
|
||||||
|
}
|
||||||
|
if bool(state.input.get("include_item_reports")):
|
||||||
|
result["items"] = report.get("items", [])
|
||||||
|
return result
|
||||||
|
|
||||||
|
|
||||||
|
def _build_stale_job_reason(state: Any) -> str | None:
|
||||||
|
updated_at = state.updated_at
|
||||||
|
stale_after_seconds = _estimate_job_stale_seconds(state)
|
||||||
|
age_seconds = (datetime.now(tz=updated_at.tzinfo) - updated_at).total_seconds()
|
||||||
|
if age_seconds < stale_after_seconds:
|
||||||
|
return None
|
||||||
|
|
||||||
|
current_stage = state.current_stage or "unknown_stage"
|
||||||
|
age_minutes = int(age_seconds // 60)
|
||||||
|
stale_after_minutes = int(stale_after_seconds // 60)
|
||||||
|
return (
|
||||||
|
f"job run-state has remained in running state at {current_stage} for about {age_minutes} minutes "
|
||||||
|
f"without result.json; it exceeded the stale threshold of {stale_after_minutes} minutes."
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _estimate_job_stale_seconds(state: Any) -> int:
|
||||||
|
input_payload = state.input if isinstance(state.input, dict) else {}
|
||||||
|
limit = _safe_int(input_payload.get("limit"), default=5)
|
||||||
|
timeout_seconds = _safe_float(input_payload.get("timeout_seconds"), default=60.0)
|
||||||
|
max_retries = _safe_int(input_payload.get("max_retries"), default=2)
|
||||||
|
estimated = int((limit * max(timeout_seconds, 1.0) * max(max_retries, 1)) + 20 * 60)
|
||||||
|
return max(MIN_JOB_STALE_SECONDS, min(MAX_JOB_STALE_SECONDS, estimated))
|
||||||
|
|
||||||
|
|
||||||
|
def _safe_int(value: Any, *, default: int) -> int:
|
||||||
|
try:
|
||||||
|
return int(value)
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
return default
|
||||||
|
|
||||||
|
|
||||||
|
def _safe_float(value: Any, *, default: float) -> float:
|
||||||
|
try:
|
||||||
|
return float(value)
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
return default
|
||||||
|
|||||||
@@ -45,6 +45,8 @@ DISCOVERED_ARTIFACTS = [
|
|||||||
"relative_path": Path("candidates/digest-brief.json"),
|
"relative_path": Path("candidates/digest-brief.json"),
|
||||||
},
|
},
|
||||||
]
|
]
|
||||||
|
MIN_RUNNING_STALE_SECONDS = 30 * 60
|
||||||
|
MAX_RUNNING_STALE_SECONDS = 6 * 60 * 60
|
||||||
|
|
||||||
|
|
||||||
def get_run_status(*, run_id: str) -> dict[str, Any]:
|
def get_run_status(*, run_id: str) -> dict[str, Any]:
|
||||||
@@ -186,7 +188,7 @@ def _build_run_record(run_dir: Path) -> dict[str, Any]:
|
|||||||
run_store = RunStore.load(path=state_path, repo_root=REPO_ROOT)
|
run_store = RunStore.load(path=state_path, repo_root=REPO_ROOT)
|
||||||
state = run_store.state
|
state = run_store.state
|
||||||
run_id = state.run_id
|
run_id = state.run_id
|
||||||
return {
|
record = {
|
||||||
"run_id": run_id,
|
"run_id": run_id,
|
||||||
"workflow": state.workflow,
|
"workflow": state.workflow,
|
||||||
"run_type": state.run_type,
|
"run_type": state.run_type,
|
||||||
@@ -204,12 +206,13 @@ def _build_run_record(run_dir: Path) -> dict[str, Any]:
|
|||||||
"state": state,
|
"state": state,
|
||||||
"report": _load_json(report_path) if report_path.exists() else None,
|
"report": _load_json(report_path) if report_path.exists() else None,
|
||||||
}
|
}
|
||||||
|
return _reconcile_run_record(record)
|
||||||
|
|
||||||
report = _load_json(report_path) if report_path.exists() else None
|
report = _load_json(report_path) if report_path.exists() else None
|
||||||
run_id = str(report.get("run_id")) if isinstance(report, dict) and report.get("run_id") else run_dir.name
|
run_id = str(report.get("run_id")) if isinstance(report, dict) and report.get("run_id") else run_dir.name
|
||||||
inferred_record = _infer_run_record_from_directory(run_dir=run_dir, report=report, run_id=run_id)
|
inferred_record = _infer_run_record_from_directory(run_dir=run_dir, report=report, run_id=run_id)
|
||||||
inferred_record["aliases"] = {run_id, run_dir.name}
|
inferred_record["aliases"] = {run_id, run_dir.name}
|
||||||
return inferred_record
|
return _reconcile_run_record(inferred_record)
|
||||||
|
|
||||||
|
|
||||||
def _infer_run_record_from_directory(*, run_dir: Path, report: dict[str, Any] | None, run_id: str) -> dict[str, Any]:
|
def _infer_run_record_from_directory(*, run_dir: Path, report: dict[str, Any] | None, run_id: str) -> dict[str, Any]:
|
||||||
@@ -298,6 +301,13 @@ def _build_status_response(record: dict[str, Any]) -> dict[str, Any]:
|
|||||||
"artifacts": _collect_artifacts(record),
|
"artifacts": _collect_artifacts(record),
|
||||||
"recovery": record["recovery"],
|
"recovery": record["recovery"],
|
||||||
"state_source": record["state_source"],
|
"state_source": record["state_source"],
|
||||||
|
"status_source": record["status_source"],
|
||||||
|
"state_quality": record["state_quality"],
|
||||||
|
"state_conflict": record["state_conflict"],
|
||||||
|
"state_conflict_reason": record["state_conflict_reason"],
|
||||||
|
"raw_status": record["raw_status"],
|
||||||
|
"raw_current_stage": record["raw_current_stage"],
|
||||||
|
"artifact_presence": record["artifact_presence"],
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
@@ -310,6 +320,10 @@ def _build_run_lookup_response(record: dict[str, Any]) -> dict[str, Any]:
|
|||||||
"status": record["status"],
|
"status": record["status"],
|
||||||
"output_dir": _normalize_repo_path(record["run_dir"]),
|
"output_dir": _normalize_repo_path(record["run_dir"]),
|
||||||
"state_source": record["state_source"],
|
"state_source": record["state_source"],
|
||||||
|
"status_source": record["status_source"],
|
||||||
|
"state_quality": record["state_quality"],
|
||||||
|
"state_conflict": record["state_conflict"],
|
||||||
|
"state_conflict_reason": record["state_conflict_reason"],
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
@@ -330,9 +344,191 @@ def _build_list_response(record: dict[str, Any]) -> dict[str, Any]:
|
|||||||
"recovery": record["recovery"],
|
"recovery": record["recovery"],
|
||||||
"artifact_count": len(_collect_artifacts(record)),
|
"artifact_count": len(_collect_artifacts(record)),
|
||||||
"state_source": record["state_source"],
|
"state_source": record["state_source"],
|
||||||
|
"status_source": record["status_source"],
|
||||||
|
"state_quality": record["state_quality"],
|
||||||
|
"state_conflict": record["state_conflict"],
|
||||||
|
"state_conflict_reason": record["state_conflict_reason"],
|
||||||
|
"raw_status": record["raw_status"],
|
||||||
|
"raw_current_stage": record["raw_current_stage"],
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _reconcile_run_record(record: dict[str, Any]) -> dict[str, Any]:
|
||||||
|
raw_status = record["status"]
|
||||||
|
raw_current_stage = record["current_stage"]
|
||||||
|
raw_stages = record["stages"]
|
||||||
|
raw_recovery = record["recovery"]
|
||||||
|
artifact_presence = _build_artifact_presence(record["run_dir"], report=record.get("report"))
|
||||||
|
|
||||||
|
reconciled = dict(record)
|
||||||
|
reconciled["raw_status"] = raw_status
|
||||||
|
reconciled["raw_current_stage"] = raw_current_stage
|
||||||
|
reconciled["status_source"] = record["state_source"]
|
||||||
|
reconciled["state_quality"] = "trusted"
|
||||||
|
reconciled["state_conflict"] = False
|
||||||
|
reconciled["state_conflict_reason"] = None
|
||||||
|
reconciled["artifact_presence"] = artifact_presence
|
||||||
|
|
||||||
|
report = record.get("report")
|
||||||
|
if not isinstance(report, dict):
|
||||||
|
stale_reason = _build_stale_running_reason(record)
|
||||||
|
if stale_reason is not None:
|
||||||
|
reconciled["status"] = "failed"
|
||||||
|
reconciled["finished_at"] = record["updated_at"]
|
||||||
|
reconciled["stages"] = _build_stale_failed_stages(raw_stages, raw_current_stage)
|
||||||
|
reconciled["error"] = {
|
||||||
|
"type": "StaleRunState",
|
||||||
|
"message": stale_reason,
|
||||||
|
"stage": raw_current_stage,
|
||||||
|
"details": {
|
||||||
|
"raw_status": raw_status,
|
||||||
|
"raw_current_stage": raw_current_stage,
|
||||||
|
},
|
||||||
|
}
|
||||||
|
reconciled["status_source"] = "stale_run_state_timeout"
|
||||||
|
reconciled["state_quality"] = "reconciled"
|
||||||
|
reconciled["state_conflict"] = True
|
||||||
|
reconciled["state_conflict_reason"] = stale_reason
|
||||||
|
return reconciled
|
||||||
|
|
||||||
|
effective_status = _infer_status_from_report(report)
|
||||||
|
report_completed_at = _maybe_iso(report.get("completed_at"))
|
||||||
|
conflict = (
|
||||||
|
raw_status != effective_status
|
||||||
|
or raw_current_stage is not None
|
||||||
|
or any(stage["status"] in {"running", "failed"} for stage in raw_stages)
|
||||||
|
or bool(raw_recovery.get("resumable"))
|
||||||
|
)
|
||||||
|
|
||||||
|
if not conflict:
|
||||||
|
reconciled["artifact_presence"] = artifact_presence
|
||||||
|
return reconciled
|
||||||
|
|
||||||
|
reconciled["status"] = effective_status
|
||||||
|
reconciled["current_stage"] = None
|
||||||
|
reconciled["updated_at"] = report_completed_at or record["updated_at"]
|
||||||
|
reconciled["finished_at"] = report_completed_at or record["finished_at"]
|
||||||
|
reconciled["stages"] = _build_terminal_success_stages(raw_stages)
|
||||||
|
reconciled["error"] = None
|
||||||
|
reconciled["recovery"] = {
|
||||||
|
"resumable": False,
|
||||||
|
"resume_from_stage": None,
|
||||||
|
"last_success_stage": DEFAULT_STAGES[-1],
|
||||||
|
}
|
||||||
|
reconciled["status_source"] = "run_report_reconciliation"
|
||||||
|
reconciled["state_quality"] = "reconciled"
|
||||||
|
reconciled["state_conflict"] = True
|
||||||
|
reconciled["state_conflict_reason"] = (
|
||||||
|
"run-state.json did not converge, but run-report.json already proves the workflow reached a terminal state."
|
||||||
|
)
|
||||||
|
return reconciled
|
||||||
|
|
||||||
|
|
||||||
|
def _build_artifact_presence(run_dir: Path, *, report: dict[str, Any] | None) -> dict[str, bool]:
|
||||||
|
return {
|
||||||
|
"run_report": isinstance(report, dict) or (run_dir / "run-report.json").exists(),
|
||||||
|
"delivery_payload": (run_dir / "candidates" / "openclaw-delivery-payload.json").exists(),
|
||||||
|
"digest_brief": (run_dir / "candidates" / "digest-brief.json").exists(),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _build_terminal_success_stages(stages: list[dict[str, Any]]) -> list[dict[str, Any]]:
|
||||||
|
stage_by_name = {stage["name"]: stage for stage in stages}
|
||||||
|
reconciled_stages: list[dict[str, Any]] = []
|
||||||
|
for stage_name in DEFAULT_STAGES:
|
||||||
|
existing = stage_by_name.get(stage_name)
|
||||||
|
if existing is None:
|
||||||
|
reconciled_stages.append(_stage_dict(name=stage_name, status="success"))
|
||||||
|
continue
|
||||||
|
reconciled_stages.append(
|
||||||
|
{
|
||||||
|
"name": stage_name,
|
||||||
|
"status": "success",
|
||||||
|
"started_at": existing.get("started_at"),
|
||||||
|
"finished_at": existing.get("finished_at"),
|
||||||
|
"outputs": existing.get("outputs", {}),
|
||||||
|
"error": None,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
return reconciled_stages
|
||||||
|
|
||||||
|
|
||||||
|
def _build_stale_failed_stages(stages: list[dict[str, Any]], current_stage: str | None) -> list[dict[str, Any]]:
|
||||||
|
stage_by_name = {stage["name"]: stage for stage in stages}
|
||||||
|
reconciled_stages: list[dict[str, Any]] = []
|
||||||
|
for stage_name in DEFAULT_STAGES:
|
||||||
|
existing = stage_by_name.get(stage_name)
|
||||||
|
if existing is None:
|
||||||
|
reconciled_stages.append(_stage_dict(name=stage_name, status="pending"))
|
||||||
|
continue
|
||||||
|
|
||||||
|
status = existing.get("status")
|
||||||
|
if status == "running" or (current_stage is not None and stage_name == current_stage):
|
||||||
|
status = "failed"
|
||||||
|
|
||||||
|
reconciled_stages.append(
|
||||||
|
{
|
||||||
|
"name": stage_name,
|
||||||
|
"status": status,
|
||||||
|
"started_at": existing.get("started_at"),
|
||||||
|
"finished_at": existing.get("finished_at") or existing.get("started_at"),
|
||||||
|
"outputs": existing.get("outputs", {}),
|
||||||
|
"error": existing.get("error"),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
return reconciled_stages
|
||||||
|
|
||||||
|
|
||||||
|
def _build_stale_running_reason(record: dict[str, Any]) -> str | None:
|
||||||
|
if record["status"] != "running":
|
||||||
|
return None
|
||||||
|
|
||||||
|
updated_at = _parse_iso_datetime(record.get("updated_at"))
|
||||||
|
if updated_at is None:
|
||||||
|
return None
|
||||||
|
|
||||||
|
stale_after_seconds = _estimate_running_stale_seconds(record)
|
||||||
|
age_seconds = (datetime.now(tz=updated_at.tzinfo) - updated_at).total_seconds()
|
||||||
|
if age_seconds < stale_after_seconds:
|
||||||
|
return None
|
||||||
|
|
||||||
|
current_stage = record.get("current_stage") or "unknown_stage"
|
||||||
|
age_minutes = int(age_seconds // 60)
|
||||||
|
stale_after_minutes = int(stale_after_seconds // 60)
|
||||||
|
return (
|
||||||
|
f"run-state.json has remained in running state at {current_stage} for about {age_minutes} minutes "
|
||||||
|
f"without terminal artifacts; it exceeded the stale threshold of {stale_after_minutes} minutes."
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _estimate_running_stale_seconds(record: dict[str, Any]) -> int:
|
||||||
|
state = record.get("state")
|
||||||
|
input_payload = state.input if isinstance(state, RunState) and isinstance(state.input, dict) else {}
|
||||||
|
limit = _safe_int(input_payload.get("limit"), default=5)
|
||||||
|
timeout_seconds = _safe_float(input_payload.get("timeout_seconds"), default=60.0)
|
||||||
|
max_retries = _safe_int(input_payload.get("max_retries"), default=2)
|
||||||
|
|
||||||
|
expected_items = limit
|
||||||
|
current_stage = record.get("current_stage")
|
||||||
|
if current_stage == "generate_summaries":
|
||||||
|
expected_items = _stage_output_from_record(record, "generate_summaries", "expected_items") or limit
|
||||||
|
elif current_stage == "extract_articles":
|
||||||
|
expected_items = _stage_output_from_record(record, "extract_articles", "expected_items") or limit
|
||||||
|
|
||||||
|
estimated = int((expected_items * max(timeout_seconds, 1.0) * max(max_retries, 1)) + 15 * 60)
|
||||||
|
return max(MIN_RUNNING_STALE_SECONDS, min(MAX_RUNNING_STALE_SECONDS, estimated))
|
||||||
|
|
||||||
|
|
||||||
|
def _stage_output_from_record(record: dict[str, Any], stage_name: str, key: str) -> Any:
|
||||||
|
for stage in record.get("stages", []):
|
||||||
|
if stage.get("name") != stage_name:
|
||||||
|
continue
|
||||||
|
outputs = stage.get("outputs")
|
||||||
|
if isinstance(outputs, dict):
|
||||||
|
return outputs.get(key)
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
def _collect_artifacts(record: dict[str, Any]) -> list[dict[str, Any]]:
|
def _collect_artifacts(record: dict[str, Any]) -> list[dict[str, Any]]:
|
||||||
artifacts: list[dict[str, Any]] = []
|
artifacts: list[dict[str, Any]] = []
|
||||||
seen_names: set[str] = set()
|
seen_names: set[str] = set()
|
||||||
@@ -587,6 +783,29 @@ def _maybe_iso(value: Any) -> str | None:
|
|||||||
return None
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def _parse_iso_datetime(value: Any) -> datetime | None:
|
||||||
|
if not isinstance(value, str) or not value.strip():
|
||||||
|
return None
|
||||||
|
try:
|
||||||
|
return datetime.fromisoformat(value)
|
||||||
|
except ValueError:
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def _safe_int(value: Any, *, default: int) -> int:
|
||||||
|
try:
|
||||||
|
return int(value)
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
return default
|
||||||
|
|
||||||
|
|
||||||
|
def _safe_float(value: Any, *, default: float) -> float:
|
||||||
|
try:
|
||||||
|
return float(value)
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
return default
|
||||||
|
|
||||||
|
|
||||||
def _record_sort_key(record: dict[str, Any]) -> tuple[str, str]:
|
def _record_sort_key(record: dict[str, Any]) -> tuple[str, str]:
|
||||||
return (record.get("updated_at") or "", record["run_dir"].name)
|
return (record.get("updated_at") or "", record["run_dir"].name)
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,595 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import subprocess
|
||||||
|
import sys
|
||||||
|
from datetime import datetime
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
from uuid import uuid4
|
||||||
|
|
||||||
|
from .query_service import _resolve_run_record
|
||||||
|
from .resume_service import (
|
||||||
|
SUPPORTED_RESUME_STAGES,
|
||||||
|
UNSUPPORTED_RESUME_STAGES,
|
||||||
|
_build_resume_plan,
|
||||||
|
_resume_freshrss_run,
|
||||||
|
_validate_resume_artifacts,
|
||||||
|
inspect_resume_plan,
|
||||||
|
)
|
||||||
|
from .run_store import RunStore
|
||||||
|
from .state_models import RunState
|
||||||
|
|
||||||
|
REPO_ROOT = Path(__file__).resolve().parents[3]
|
||||||
|
OUTPUT_ROOT = REPO_ROOT / "outputs" / "freshrss"
|
||||||
|
RESUME_JOBS_ROOT = OUTPUT_ROOT / "resume_jobs"
|
||||||
|
WORKFLOW_NAME = "freshrss_resume_job"
|
||||||
|
RUN_TYPE = "resume_job"
|
||||||
|
RUN_STATE_FILENAME = "run-state.json"
|
||||||
|
INPUT_FILENAME = "input.json"
|
||||||
|
RESULT_FILENAME = "result.json"
|
||||||
|
JOB_REPORT_FILENAME = "job-report.json"
|
||||||
|
DEFAULT_STAGES = [
|
||||||
|
"prepare_job",
|
||||||
|
"validate_resume_plan",
|
||||||
|
"resume_run",
|
||||||
|
"write_result",
|
||||||
|
]
|
||||||
|
MIN_JOB_STALE_SECONDS = 30 * 60
|
||||||
|
MAX_JOB_STALE_SECONDS = 6 * 60 * 60
|
||||||
|
|
||||||
|
|
||||||
|
def _now() -> datetime:
|
||||||
|
return datetime.now().astimezone()
|
||||||
|
|
||||||
|
|
||||||
|
def _new_job_id() -> str:
|
||||||
|
ts = _now().strftime("%Y%m%d-%H%M%S")
|
||||||
|
return f"freshrss-resume-job-{ts}-{uuid4().hex[:8]}"
|
||||||
|
|
||||||
|
|
||||||
|
def _job_dir(job_id: str) -> Path:
|
||||||
|
return RESUME_JOBS_ROOT / job_id
|
||||||
|
|
||||||
|
|
||||||
|
def _run_state_path(job_id: str) -> Path:
|
||||||
|
return _job_dir(job_id) / RUN_STATE_FILENAME
|
||||||
|
|
||||||
|
|
||||||
|
def _input_path(job_id: str) -> Path:
|
||||||
|
return _job_dir(job_id) / INPUT_FILENAME
|
||||||
|
|
||||||
|
|
||||||
|
def _result_path(job_id: str) -> Path:
|
||||||
|
return _job_dir(job_id) / RESULT_FILENAME
|
||||||
|
|
||||||
|
|
||||||
|
def _job_report_path(job_id: str) -> Path:
|
||||||
|
return _job_dir(job_id) / JOB_REPORT_FILENAME
|
||||||
|
|
||||||
|
|
||||||
|
def _normalize_repo_path(path: Path) -> str:
|
||||||
|
try:
|
||||||
|
return str(path.resolve().relative_to(REPO_ROOT.resolve()))
|
||||||
|
except ValueError:
|
||||||
|
return str(path)
|
||||||
|
|
||||||
|
|
||||||
|
def _normalize_repo_path_value(path_value: str | None) -> str | None:
|
||||||
|
if not path_value:
|
||||||
|
return None
|
||||||
|
return _normalize_repo_path(Path(path_value))
|
||||||
|
|
||||||
|
|
||||||
|
def _load_run_store(job_id: str) -> RunStore:
|
||||||
|
return RunStore.load(path=_run_state_path(job_id), repo_root=REPO_ROOT)
|
||||||
|
|
||||||
|
|
||||||
|
def _load_result(job_id: str) -> dict[str, Any] | None:
|
||||||
|
path = _result_path(job_id)
|
||||||
|
if not path.exists():
|
||||||
|
return None
|
||||||
|
return json.loads(path.read_text(encoding="utf-8-sig"))
|
||||||
|
|
||||||
|
|
||||||
|
def _load_job_report(job_id: str) -> dict[str, Any] | None:
|
||||||
|
path = _job_report_path(job_id)
|
||||||
|
if not path.exists():
|
||||||
|
return None
|
||||||
|
return json.loads(path.read_text(encoding="utf-8-sig"))
|
||||||
|
|
||||||
|
|
||||||
|
def _write_job_report(*, job_id: str, payload: dict[str, Any]) -> Path:
|
||||||
|
report_file = _job_report_path(job_id)
|
||||||
|
report_file.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||||
|
return report_file
|
||||||
|
|
||||||
|
|
||||||
|
def _build_result_payload(
|
||||||
|
*,
|
||||||
|
job_id: str,
|
||||||
|
run_id: str,
|
||||||
|
run_dir: Path,
|
||||||
|
resume_plan: dict[str, Any],
|
||||||
|
resume_result: dict[str, Any],
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
return {
|
||||||
|
"job_id": job_id,
|
||||||
|
"run_id": run_id,
|
||||||
|
"resume_from_stage": resume_plan["effective_resume_from_stage"],
|
||||||
|
"requested_resume_from_stage": resume_plan["requested_resume_from_stage"],
|
||||||
|
"resume_decision_source": resume_plan["decision_source"],
|
||||||
|
"artifact_resume_from_stage": resume_plan["artifact_resume_from_stage"],
|
||||||
|
"artifact_snapshot": resume_plan["artifact_snapshot"],
|
||||||
|
"status": resume_result["status"],
|
||||||
|
"output_dir": _normalize_repo_path(run_dir),
|
||||||
|
"delivery_output": _normalize_repo_path_value(resume_result.get("delivery_output")),
|
||||||
|
"report_output": _normalize_repo_path_value(resume_result.get("report_output")),
|
||||||
|
"digest_brief_output": _normalize_repo_path_value(resume_result.get("digest_brief_output")),
|
||||||
|
"pulled_count": resume_result.get("pulled_count"),
|
||||||
|
"delivered_count": resume_result.get("delivered_count"),
|
||||||
|
"marked_read_count": resume_result.get("marked_read_count"),
|
||||||
|
"status_counts": resume_result.get("status_counts"),
|
||||||
|
"keyword_index": resume_result.get("keyword_index"),
|
||||||
|
"completed_at": _now().isoformat(),
|
||||||
|
"result_source": "resume_job_result",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _build_result_from_linked_run(*, job_id: str, linked_run_record: dict[str, Any]) -> dict[str, Any] | None:
|
||||||
|
report = linked_run_record.get("report")
|
||||||
|
if not isinstance(report, dict):
|
||||||
|
return None
|
||||||
|
|
||||||
|
recovery = linked_run_record.get("recovery")
|
||||||
|
requested_resume_from_stage = None
|
||||||
|
if isinstance(recovery, dict):
|
||||||
|
requested_resume_from_stage = recovery.get("resume_from_stage")
|
||||||
|
|
||||||
|
return {
|
||||||
|
"job_id": job_id,
|
||||||
|
"run_id": linked_run_record["run_id"],
|
||||||
|
"resume_from_stage": requested_resume_from_stage,
|
||||||
|
"requested_resume_from_stage": requested_resume_from_stage,
|
||||||
|
"resume_decision_source": "linked_run_report",
|
||||||
|
"artifact_resume_from_stage": None,
|
||||||
|
"artifact_snapshot": None,
|
||||||
|
"status": linked_run_record["status"],
|
||||||
|
"output_dir": _normalize_repo_path(linked_run_record["run_dir"]),
|
||||||
|
"delivery_output": _normalize_repo_path_value(report.get("delivery_output")),
|
||||||
|
"report_output": _normalize_repo_path(linked_run_record["run_dir"] / "run-report.json"),
|
||||||
|
"digest_brief_output": _normalize_repo_path_value(report.get("digest_brief_output")),
|
||||||
|
"pulled_count": report.get("pulled_count"),
|
||||||
|
"delivered_count": report.get("delivered_count"),
|
||||||
|
"marked_read_count": report.get("marked_read_count"),
|
||||||
|
"status_counts": report.get("status_counts"),
|
||||||
|
"keyword_index": report.get("keyword_index"),
|
||||||
|
"completed_at": report.get("completed_at"),
|
||||||
|
"result_source": "linked_run_report",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _linked_run_result_available(linked_run_record: dict[str, Any]) -> bool:
|
||||||
|
return linked_run_record["status"] in {"success", "partial"} and isinstance(linked_run_record.get("report"), dict)
|
||||||
|
|
||||||
|
|
||||||
|
def _load_linked_run_record(state: Any) -> dict[str, Any] | None:
|
||||||
|
if not isinstance(state.input, dict):
|
||||||
|
return None
|
||||||
|
run_id = state.input.get("run_id")
|
||||||
|
if not isinstance(run_id, str) or not run_id.strip():
|
||||||
|
return None
|
||||||
|
try:
|
||||||
|
return _resolve_run_record(run_id)
|
||||||
|
except FileNotFoundError:
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def _build_stale_job_reason(state: Any) -> str | None:
|
||||||
|
if state.status != "running" or state.current_stage is None or state.finished_at is not None:
|
||||||
|
return None
|
||||||
|
now = _now()
|
||||||
|
age_seconds = max(0.0, (now - state.updated_at).total_seconds())
|
||||||
|
if age_seconds < MIN_JOB_STALE_SECONDS:
|
||||||
|
return None
|
||||||
|
if age_seconds >= MAX_JOB_STALE_SECONDS:
|
||||||
|
return (
|
||||||
|
f"Resume job has remained in stage '{state.current_stage}' for more than {int(MAX_JOB_STALE_SECONDS)} seconds "
|
||||||
|
"without producing a terminal result; treating the job state as stale."
|
||||||
|
)
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def start_resume_job(*, run_id: str) -> dict[str, Any]:
|
||||||
|
resume_view = inspect_resume_plan(run_id=run_id)
|
||||||
|
job_id = _new_job_id()
|
||||||
|
job_dir = _job_dir(job_id)
|
||||||
|
job_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
started_at = _now()
|
||||||
|
|
||||||
|
input_payload = {
|
||||||
|
"run_id": run_id,
|
||||||
|
"requested_resume_from_stage": resume_view.get("requested_resume_from_stage"),
|
||||||
|
"resume_from_stage": resume_view.get("resume_from_stage"),
|
||||||
|
"resume_decision_source": resume_view.get("resume_decision_source"),
|
||||||
|
"recommended_action": resume_view.get("recommended_action"),
|
||||||
|
"artifact_resume_from_stage": resume_view.get("artifact_resume_from_stage"),
|
||||||
|
"artifact_snapshot": resume_view.get("artifact_snapshot"),
|
||||||
|
"launcher_pid": os.getpid(),
|
||||||
|
}
|
||||||
|
|
||||||
|
store = RunStore.create(
|
||||||
|
path=_run_state_path(job_id),
|
||||||
|
run_id=job_id,
|
||||||
|
workflow=WORKFLOW_NAME,
|
||||||
|
run_type=RUN_TYPE,
|
||||||
|
started_at=started_at,
|
||||||
|
input_payload=input_payload,
|
||||||
|
repo_root=REPO_ROOT,
|
||||||
|
)
|
||||||
|
for stage_name in DEFAULT_STAGES:
|
||||||
|
store._get_or_create_stage(stage_name)
|
||||||
|
store.save()
|
||||||
|
|
||||||
|
store.start_stage("prepare_job")
|
||||||
|
input_file = _input_path(job_id)
|
||||||
|
input_file.write_text(json.dumps(input_payload, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||||
|
store.register_artifact(name="job_input", path=input_file, kind="json", stage="prepare_job")
|
||||||
|
|
||||||
|
if not resume_view.get("can_resume") or resume_view.get("recommended_action") != "resume":
|
||||||
|
report_file = _write_job_report(
|
||||||
|
job_id=job_id,
|
||||||
|
payload={
|
||||||
|
"job_id": job_id,
|
||||||
|
"status": "failed",
|
||||||
|
"error_type": "ResumePreflightRejected",
|
||||||
|
"error_message": resume_view.get("message"),
|
||||||
|
"failed_stage": "validate_resume_plan",
|
||||||
|
"linked_run_id": run_id,
|
||||||
|
"resume_from_stage": resume_view.get("resume_from_stage"),
|
||||||
|
"requested_resume_from_stage": resume_view.get("requested_resume_from_stage"),
|
||||||
|
"recommended_action": resume_view.get("recommended_action"),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
store.register_artifact(name="job_report", path=report_file, kind="json", stage="prepare_job")
|
||||||
|
store.finish_stage("prepare_job", outputs={"linked_run_id": run_id})
|
||||||
|
store.fail_stage("validate_resume_plan", error=ValueError(str(resume_view.get("message") or "Resume preflight rejected.")))
|
||||||
|
return {
|
||||||
|
"job_id": job_id,
|
||||||
|
"workflow": WORKFLOW_NAME,
|
||||||
|
"run_type": RUN_TYPE,
|
||||||
|
"status": "failed",
|
||||||
|
"output_dir": _normalize_repo_path(job_dir),
|
||||||
|
"linked_run_id": run_id,
|
||||||
|
"resume_from_stage": resume_view.get("resume_from_stage"),
|
||||||
|
"recommended_action": resume_view.get("recommended_action"),
|
||||||
|
"message": resume_view.get("message"),
|
||||||
|
}
|
||||||
|
|
||||||
|
runner_script = REPO_ROOT / "scripts" / "run_resume_job.py"
|
||||||
|
cmd = [sys.executable, str(runner_script), "--job-id", job_id]
|
||||||
|
try:
|
||||||
|
proc = subprocess.Popen(
|
||||||
|
cmd,
|
||||||
|
cwd=str(REPO_ROOT),
|
||||||
|
stdout=subprocess.DEVNULL,
|
||||||
|
stderr=subprocess.DEVNULL,
|
||||||
|
start_new_session=True,
|
||||||
|
)
|
||||||
|
except Exception as exc:
|
||||||
|
report_file = _write_job_report(
|
||||||
|
job_id=job_id,
|
||||||
|
payload={
|
||||||
|
"job_id": job_id,
|
||||||
|
"status": "failed",
|
||||||
|
"error_type": type(exc).__name__,
|
||||||
|
"error_message": str(exc),
|
||||||
|
"failed_stage": "prepare_job",
|
||||||
|
"linked_run_id": run_id,
|
||||||
|
"resume_from_stage": resume_view.get("resume_from_stage"),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
store.register_artifact(name="job_report", path=report_file, kind="json", stage="prepare_job")
|
||||||
|
store.fail_stage("prepare_job", error=exc)
|
||||||
|
raise
|
||||||
|
|
||||||
|
store.finish_stage(
|
||||||
|
"prepare_job",
|
||||||
|
outputs={
|
||||||
|
"runner_pid": proc.pid,
|
||||||
|
"runner_command": cmd,
|
||||||
|
"linked_run_id": run_id,
|
||||||
|
"resume_from_stage": resume_view.get("resume_from_stage"),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
return {
|
||||||
|
"job_id": job_id,
|
||||||
|
"workflow": WORKFLOW_NAME,
|
||||||
|
"run_type": RUN_TYPE,
|
||||||
|
"status": "running",
|
||||||
|
"output_dir": _normalize_repo_path(job_dir),
|
||||||
|
"linked_run_id": run_id,
|
||||||
|
"resume_from_stage": resume_view.get("resume_from_stage"),
|
||||||
|
"message": "Resume job started successfully. Use get_resume_job_status to poll progress.",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def run_resume_job(*, job_id: str) -> dict[str, Any]:
|
||||||
|
store = _load_run_store(job_id)
|
||||||
|
input_payload = json.loads(_input_path(job_id).read_text(encoding="utf-8-sig"))
|
||||||
|
current_stage = "resume_run"
|
||||||
|
run_id = str(input_payload["run_id"])
|
||||||
|
|
||||||
|
try:
|
||||||
|
store.start_stage("validate_resume_plan")
|
||||||
|
record = _resolve_run_record(run_id)
|
||||||
|
if record["state_source"] != "run_state" or not isinstance(record.get("state"), RunState):
|
||||||
|
raise RuntimeError("This run cannot be resumed because run-state.json is missing or could not be loaded.")
|
||||||
|
|
||||||
|
run_store = RunStore.load(path=record["run_dir"] / "run-state.json", repo_root=REPO_ROOT)
|
||||||
|
resume_plan = _build_resume_plan(record=record, state=run_store.state)
|
||||||
|
if resume_plan["decision"] != "resume":
|
||||||
|
raise RuntimeError(str(resume_plan["message"]))
|
||||||
|
|
||||||
|
resume_from_stage = resume_plan["effective_resume_from_stage"]
|
||||||
|
if resume_from_stage in UNSUPPORTED_RESUME_STAGES:
|
||||||
|
raise RuntimeError(f"This run cannot be resumed from {resume_from_stage} in the current implementation.")
|
||||||
|
if resume_from_stage not in SUPPORTED_RESUME_STAGES:
|
||||||
|
raise RuntimeError(f"This run cannot be resumed because stage '{resume_from_stage}' is not supported.")
|
||||||
|
|
||||||
|
missing_artifacts = _validate_resume_artifacts(
|
||||||
|
record=record,
|
||||||
|
state=run_store.state,
|
||||||
|
resume_from_stage=resume_from_stage,
|
||||||
|
)
|
||||||
|
if missing_artifacts:
|
||||||
|
raise RuntimeError(
|
||||||
|
f"This run cannot be resumed from {resume_from_stage} because required artifacts are missing: {missing_artifacts}"
|
||||||
|
)
|
||||||
|
|
||||||
|
store.finish_stage(
|
||||||
|
"validate_resume_plan",
|
||||||
|
outputs={
|
||||||
|
"linked_run_id": run_id,
|
||||||
|
"resume_from_stage": resume_from_stage,
|
||||||
|
"requested_resume_from_stage": resume_plan["requested_resume_from_stage"],
|
||||||
|
"resume_decision_source": resume_plan["decision_source"],
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
current_stage = "resume_run"
|
||||||
|
store.start_stage("resume_run")
|
||||||
|
resume_result = _resume_freshrss_run(record=record, run_store=run_store, resume_from_stage=resume_from_stage)
|
||||||
|
store.finish_stage(
|
||||||
|
"resume_run",
|
||||||
|
outputs={
|
||||||
|
"linked_run_id": run_id,
|
||||||
|
"resume_from_stage": resume_from_stage,
|
||||||
|
"linked_run_status": resume_result["status"],
|
||||||
|
"delivery_output": _normalize_repo_path_value(resume_result.get("delivery_output")),
|
||||||
|
"report_output": _normalize_repo_path_value(resume_result.get("report_output")),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
store.start_stage("write_result")
|
||||||
|
result = _build_result_payload(
|
||||||
|
job_id=job_id,
|
||||||
|
run_id=run_id,
|
||||||
|
run_dir=record["run_dir"],
|
||||||
|
resume_plan=resume_plan,
|
||||||
|
resume_result=resume_result,
|
||||||
|
)
|
||||||
|
result_file = _result_path(job_id)
|
||||||
|
result_file.write_text(json.dumps(result, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||||
|
store.register_artifact(name="job_result", path=result_file, kind="json", stage="write_result")
|
||||||
|
|
||||||
|
report_file = _write_job_report(
|
||||||
|
job_id=job_id,
|
||||||
|
payload={
|
||||||
|
"job_id": job_id,
|
||||||
|
"status": "success",
|
||||||
|
"run_id": run_id,
|
||||||
|
"resume_from_stage": result["resume_from_stage"],
|
||||||
|
"delivery_output": result["delivery_output"],
|
||||||
|
"report_output": result["report_output"],
|
||||||
|
"digest_brief_output": result["digest_brief_output"],
|
||||||
|
},
|
||||||
|
)
|
||||||
|
store.register_artifact(name="job_report", path=report_file, kind="json", stage="write_result")
|
||||||
|
store.finish_stage(
|
||||||
|
"write_result",
|
||||||
|
outputs={
|
||||||
|
"result_path": _normalize_repo_path(result_file),
|
||||||
|
"linked_run_id": run_id,
|
||||||
|
},
|
||||||
|
)
|
||||||
|
store.finish_run(status="success")
|
||||||
|
return result
|
||||||
|
except Exception as exc:
|
||||||
|
current_stage = store.state.current_stage or current_stage
|
||||||
|
report_file = _write_job_report(
|
||||||
|
job_id=job_id,
|
||||||
|
payload={
|
||||||
|
"job_id": job_id,
|
||||||
|
"status": "failed",
|
||||||
|
"error_type": type(exc).__name__,
|
||||||
|
"error_message": str(exc),
|
||||||
|
"failed_stage": current_stage,
|
||||||
|
"linked_run_id": run_id,
|
||||||
|
},
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
store.register_artifact(name="job_report", path=report_file, kind="json", stage=current_stage)
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
store.fail_stage(current_stage, error=exc)
|
||||||
|
raise
|
||||||
|
|
||||||
|
|
||||||
|
def _resolve_effective_job_view(
|
||||||
|
*,
|
||||||
|
job_id: str,
|
||||||
|
state: Any,
|
||||||
|
result: dict[str, Any] | None = None,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
resolved_result = result or _load_result(job_id)
|
||||||
|
linked_run_record = _load_linked_run_record(state)
|
||||||
|
raw_status = state.status
|
||||||
|
raw_current_stage = state.current_stage
|
||||||
|
effective_status = raw_status
|
||||||
|
effective_current_stage = raw_current_stage
|
||||||
|
effective_updated_at = state.updated_at.isoformat()
|
||||||
|
effective_finished_at = state.finished_at.isoformat() if state.finished_at else None
|
||||||
|
status_source = "job_run_state"
|
||||||
|
state_quality = "trusted"
|
||||||
|
state_conflict = False
|
||||||
|
state_conflict_reason = None
|
||||||
|
status_note = None
|
||||||
|
|
||||||
|
if resolved_result is not None:
|
||||||
|
completed_at = resolved_result.get("completed_at")
|
||||||
|
effective_status = "success"
|
||||||
|
effective_current_stage = None
|
||||||
|
effective_updated_at = completed_at or effective_updated_at
|
||||||
|
effective_finished_at = completed_at or effective_finished_at
|
||||||
|
if raw_status != "success" or raw_current_stage is not None:
|
||||||
|
status_source = "job_result_reconciliation"
|
||||||
|
state_quality = "reconciled"
|
||||||
|
state_conflict = True
|
||||||
|
state_conflict_reason = "result.json already exists, but the resume job run-state did not converge to success."
|
||||||
|
status_note = "resume job result exists; the job can be treated as completed."
|
||||||
|
elif linked_run_record is not None:
|
||||||
|
if _linked_run_result_available(linked_run_record) and raw_status == "running":
|
||||||
|
effective_status = "success"
|
||||||
|
effective_current_stage = None
|
||||||
|
effective_updated_at = linked_run_record.get("updated_at") or effective_updated_at
|
||||||
|
effective_finished_at = linked_run_record.get("finished_at") or effective_finished_at
|
||||||
|
status_source = "linked_run_reconciliation"
|
||||||
|
state_quality = "reconciled"
|
||||||
|
state_conflict = raw_status != "success" or raw_current_stage is not None
|
||||||
|
state_conflict_reason = "The linked run already has terminal artifacts, but the resume job run-state did not converge."
|
||||||
|
status_note = "linked run artifacts are complete; the resume job can be treated as completed."
|
||||||
|
resolved_result = _build_result_from_linked_run(job_id=job_id, linked_run_record=linked_run_record)
|
||||||
|
elif raw_status == "running":
|
||||||
|
stale_reason = _build_stale_job_reason(state)
|
||||||
|
if stale_reason is not None:
|
||||||
|
effective_status = "failed"
|
||||||
|
effective_current_stage = raw_current_stage
|
||||||
|
effective_finished_at = effective_updated_at
|
||||||
|
status_source = "stale_job_state_timeout"
|
||||||
|
state_quality = "reconciled"
|
||||||
|
state_conflict = True
|
||||||
|
state_conflict_reason = stale_reason
|
||||||
|
status_note = stale_reason
|
||||||
|
|
||||||
|
return {
|
||||||
|
"status": effective_status,
|
||||||
|
"current_stage": effective_current_stage,
|
||||||
|
"updated_at": effective_updated_at,
|
||||||
|
"finished_at": effective_finished_at,
|
||||||
|
"status_source": status_source,
|
||||||
|
"state_quality": state_quality,
|
||||||
|
"state_conflict": state_conflict,
|
||||||
|
"state_conflict_reason": state_conflict_reason,
|
||||||
|
"status_note": status_note,
|
||||||
|
"result": resolved_result,
|
||||||
|
"linked_run_status": linked_run_record["status"] if linked_run_record is not None else None,
|
||||||
|
"linked_run_state_conflict": linked_run_record.get("state_conflict") if linked_run_record is not None else None,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _build_effective_job_progress(*, state: Any, effective_status: str) -> dict[str, int]:
|
||||||
|
completed_stage_count = sum(1 for stage in state.stages if stage.status == "success")
|
||||||
|
running_stage_count = sum(1 for stage in state.stages if stage.status == "running")
|
||||||
|
failed_stage_count = sum(1 for stage in state.stages if stage.status == "failed")
|
||||||
|
pending_stage_count = sum(1 for stage in state.stages if stage.status == "pending")
|
||||||
|
if effective_status == "success" and running_stage_count > 0:
|
||||||
|
pending_stage_count += running_stage_count
|
||||||
|
running_stage_count = 0
|
||||||
|
return {
|
||||||
|
"completed_stage_count": completed_stage_count,
|
||||||
|
"running_stage_count": running_stage_count,
|
||||||
|
"failed_stage_count": failed_stage_count,
|
||||||
|
"pending_stage_count": pending_stage_count,
|
||||||
|
"total_stage_count": len(state.stages),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def get_resume_job_status(*, job_id: str) -> dict[str, Any]:
|
||||||
|
store = _load_run_store(job_id)
|
||||||
|
state = store.state
|
||||||
|
effective = _resolve_effective_job_view(job_id=job_id, state=state)
|
||||||
|
linked_run_id = state.input.get("run_id") if isinstance(state.input, dict) else None
|
||||||
|
return {
|
||||||
|
"job_id": state.run_id,
|
||||||
|
"workflow": state.workflow,
|
||||||
|
"run_type": state.run_type,
|
||||||
|
"status": effective["status"],
|
||||||
|
"current_stage": effective["current_stage"],
|
||||||
|
"started_at": state.started_at.isoformat(),
|
||||||
|
"updated_at": effective["updated_at"],
|
||||||
|
"finished_at": effective["finished_at"],
|
||||||
|
"output_dir": _normalize_repo_path(_job_dir(job_id)),
|
||||||
|
"linked_run_id": linked_run_id,
|
||||||
|
"progress": _build_effective_job_progress(state=state, effective_status=effective["status"]),
|
||||||
|
"artifacts": [artifact.model_dump(mode="json") for artifact in state.artifacts],
|
||||||
|
"error_summary": state.error.model_dump(mode="json") if state.error else None,
|
||||||
|
"status_source": effective["status_source"],
|
||||||
|
"state_quality": effective["state_quality"],
|
||||||
|
"state_conflict": effective["state_conflict"],
|
||||||
|
"state_conflict_reason": effective["state_conflict_reason"],
|
||||||
|
"status_note": effective["status_note"],
|
||||||
|
"raw_status": state.status,
|
||||||
|
"raw_current_stage": state.current_stage,
|
||||||
|
"linked_run_status": effective["linked_run_status"],
|
||||||
|
"linked_run_state_conflict": effective["linked_run_state_conflict"],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def get_resume_job_result(*, job_id: str) -> dict[str, Any]:
|
||||||
|
store = _load_run_store(job_id)
|
||||||
|
state = store.state
|
||||||
|
result = _load_result(job_id)
|
||||||
|
effective = _resolve_effective_job_view(job_id=job_id, state=state, result=result)
|
||||||
|
synthesized_result = result or effective["result"]
|
||||||
|
if effective["status"] != "success" or synthesized_result is None:
|
||||||
|
message = "Resume job result is not ready."
|
||||||
|
if effective["status"] == "failed":
|
||||||
|
message = "Resume job did not complete successfully, so no terminal job result is available."
|
||||||
|
return {
|
||||||
|
"job_id": state.run_id,
|
||||||
|
"status": effective["status"],
|
||||||
|
"message": message,
|
||||||
|
"linked_run_id": state.input.get("run_id") if isinstance(state.input, dict) else None,
|
||||||
|
"error_summary": state.error.model_dump(mode="json") if state.error else None,
|
||||||
|
"status_source": effective["status_source"],
|
||||||
|
"status_note": effective["status_note"],
|
||||||
|
}
|
||||||
|
|
||||||
|
artifact = next((a.model_dump(mode="json") for a in state.artifacts if a.name == "job_result"), None)
|
||||||
|
return {
|
||||||
|
"job_id": state.run_id,
|
||||||
|
"status": effective["status"],
|
||||||
|
"run_id": synthesized_result.get("run_id"),
|
||||||
|
"resume_from_stage": synthesized_result.get("resume_from_stage"),
|
||||||
|
"requested_resume_from_stage": synthesized_result.get("requested_resume_from_stage"),
|
||||||
|
"resume_decision_source": synthesized_result.get("resume_decision_source"),
|
||||||
|
"artifact_resume_from_stage": synthesized_result.get("artifact_resume_from_stage"),
|
||||||
|
"artifact_snapshot": synthesized_result.get("artifact_snapshot"),
|
||||||
|
"output_dir": synthesized_result.get("output_dir"),
|
||||||
|
"delivery_output": synthesized_result.get("delivery_output"),
|
||||||
|
"report_output": synthesized_result.get("report_output"),
|
||||||
|
"digest_brief_output": synthesized_result.get("digest_brief_output"),
|
||||||
|
"pulled_count": synthesized_result.get("pulled_count"),
|
||||||
|
"delivered_count": synthesized_result.get("delivered_count"),
|
||||||
|
"marked_read_count": synthesized_result.get("marked_read_count"),
|
||||||
|
"status_counts": synthesized_result.get("status_counts"),
|
||||||
|
"keyword_index": synthesized_result.get("keyword_index"),
|
||||||
|
"artifact": artifact,
|
||||||
|
"result": synthesized_result,
|
||||||
|
"status_source": effective["status_source"],
|
||||||
|
"status_note": effective["status_note"],
|
||||||
|
"result_source": synthesized_result.get("result_source", "job_result"),
|
||||||
|
"linked_run_status": effective["linked_run_status"],
|
||||||
|
}
|
||||||
@@ -25,6 +25,7 @@ from summary_mcp.models.openclaw_delivery import (
|
|||||||
)
|
)
|
||||||
from summary_mcp.models.summary_io import ExtractionOutput
|
from summary_mcp.models.summary_io import ExtractionOutput
|
||||||
from summary_mcp.workflows.freshrss_pipeline import (
|
from summary_mcp.workflows.freshrss_pipeline import (
|
||||||
|
CANDIDATE_BATCH_ARTIFACT,
|
||||||
DEFAULT_PROMPT_PATH,
|
DEFAULT_PROMPT_PATH,
|
||||||
DEFAULT_RULES_PATH,
|
DEFAULT_RULES_PATH,
|
||||||
DEFAULT_TERM_ALIASES_PATH,
|
DEFAULT_TERM_ALIASES_PATH,
|
||||||
@@ -37,13 +38,18 @@ from summary_mcp.workflows.freshrss_pipeline import (
|
|||||||
FILTER_STAGE,
|
FILTER_STAGE,
|
||||||
REPORT_STAGE,
|
REPORT_STAGE,
|
||||||
REPO_ROOT,
|
REPO_ROOT,
|
||||||
|
SUMMARY_BATCH_ARTIFACT,
|
||||||
SUMMARY_STAGE,
|
SUMMARY_STAGE,
|
||||||
WORKFLOW_NAME,
|
WORKFLOW_NAME,
|
||||||
_build_item_context,
|
_build_item_context,
|
||||||
|
_candidate_batch_output,
|
||||||
_final_run_status,
|
_final_run_status,
|
||||||
_load_json,
|
_load_json,
|
||||||
_load_required_env,
|
_load_required_env,
|
||||||
|
_persist_candidate_batch_artifact,
|
||||||
|
_persist_summary_batch_artifact,
|
||||||
_save_json,
|
_save_json,
|
||||||
|
_summary_batch_output,
|
||||||
)
|
)
|
||||||
|
|
||||||
from .query_service import _normalize_repo_path, _resolve_repo_path, _resolve_run_record
|
from .query_service import _normalize_repo_path, _resolve_repo_path, _resolve_run_record
|
||||||
@@ -62,6 +68,12 @@ UNSUPPORTED_RESUME_STAGES = {
|
|||||||
}
|
}
|
||||||
UTC = timezone.utc
|
UTC = timezone.utc
|
||||||
DEFAULT_STREAM_ID = "user/-/state/com.google/reading-list"
|
DEFAULT_STREAM_ID = "user/-/state/com.google/reading-list"
|
||||||
|
RESUME_STAGE_ORDER = {
|
||||||
|
SUMMARY_STAGE: 1,
|
||||||
|
FILTER_STAGE: 2,
|
||||||
|
DELIVERY_STAGE: 3,
|
||||||
|
REPORT_STAGE: 4,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
def resume_run(*, run_id: str) -> dict[str, Any]:
|
def resume_run(*, run_id: str) -> dict[str, Any]:
|
||||||
@@ -71,37 +83,68 @@ def resume_run(*, run_id: str) -> dict[str, Any]:
|
|||||||
"workflow": record["workflow"],
|
"workflow": record["workflow"],
|
||||||
"status": record["status"],
|
"status": record["status"],
|
||||||
"output_dir": _normalize_repo_path(record["run_dir"]),
|
"output_dir": _normalize_repo_path(record["run_dir"]),
|
||||||
|
"state_source": record["state_source"],
|
||||||
|
"status_source": record.get("status_source"),
|
||||||
}
|
}
|
||||||
|
|
||||||
|
if record["status"] in {"success", "partial"} and isinstance(record.get("report"), dict):
|
||||||
|
return {
|
||||||
|
**base_response,
|
||||||
|
"resumed": False,
|
||||||
|
"resume_from_stage": None,
|
||||||
|
"requested_resume_from_stage": None,
|
||||||
|
"resume_decision_source": "artifacts",
|
||||||
|
"recommended_action": "read_terminal_result",
|
||||||
|
"message": "This run already has a terminal run-report.json; prefer reading get_run_status/get_run_report instead of resuming.",
|
||||||
|
"missing_artifacts": [],
|
||||||
|
"state_conflict": record.get("state_conflict", False),
|
||||||
|
"state_conflict_reason": record.get("state_conflict_reason"),
|
||||||
|
}
|
||||||
|
|
||||||
if record["state_source"] != "run_state" or not isinstance(record.get("state"), RunState):
|
if record["state_source"] != "run_state" or not isinstance(record.get("state"), RunState):
|
||||||
return {
|
return {
|
||||||
**base_response,
|
**base_response,
|
||||||
"resumed": False,
|
"resumed": False,
|
||||||
"resume_from_stage": None,
|
"resume_from_stage": None,
|
||||||
|
"requested_resume_from_stage": None,
|
||||||
|
"resume_decision_source": "unavailable",
|
||||||
|
"recommended_action": "start_new_run",
|
||||||
"message": "This run cannot be resumed because run-state.json is missing or could not be loaded.",
|
"message": "This run cannot be resumed because run-state.json is missing or could not be loaded.",
|
||||||
"missing_artifacts": [],
|
"missing_artifacts": [],
|
||||||
}
|
}
|
||||||
|
|
||||||
run_store = RunStore.load(path=record["run_dir"] / "run-state.json", repo_root=REPO_ROOT)
|
run_store = RunStore.load(path=record["run_dir"] / "run-state.json", repo_root=REPO_ROOT)
|
||||||
state = run_store.state
|
state = run_store.state
|
||||||
resume_from_stage = _resolve_resume_from_stage(state)
|
resume_plan = _build_resume_plan(record=record, state=state)
|
||||||
|
requested_resume_from_stage = resume_plan["requested_resume_from_stage"]
|
||||||
|
resume_from_stage = resume_plan["effective_resume_from_stage"]
|
||||||
|
|
||||||
if state.workflow != WORKFLOW_NAME:
|
if state.workflow != WORKFLOW_NAME:
|
||||||
return {
|
return {
|
||||||
**base_response,
|
**base_response,
|
||||||
"resumed": False,
|
"resumed": False,
|
||||||
"resume_from_stage": resume_from_stage,
|
"resume_from_stage": resume_from_stage,
|
||||||
|
"requested_resume_from_stage": requested_resume_from_stage,
|
||||||
|
"resume_decision_source": resume_plan["decision_source"],
|
||||||
|
"recommended_action": "start_new_run",
|
||||||
"message": f"This run cannot be resumed because workflow '{state.workflow}' is not supported by the minimal resume_run implementation.",
|
"message": f"This run cannot be resumed because workflow '{state.workflow}' is not supported by the minimal resume_run implementation.",
|
||||||
"missing_artifacts": [],
|
"missing_artifacts": [],
|
||||||
}
|
}
|
||||||
|
|
||||||
if resume_from_stage is None:
|
if resume_plan["decision"] != "resume":
|
||||||
return {
|
return {
|
||||||
**base_response,
|
**base_response,
|
||||||
"resumed": False,
|
"resumed": False,
|
||||||
"resume_from_stage": None,
|
"resume_from_stage": resume_from_stage,
|
||||||
"message": "This run does not expose a recoverable stage in run-state.json.",
|
"requested_resume_from_stage": requested_resume_from_stage,
|
||||||
"missing_artifacts": [],
|
"resume_decision_source": resume_plan["decision_source"],
|
||||||
|
"recommended_action": resume_plan["recommended_action"],
|
||||||
|
"message": resume_plan["message"],
|
||||||
|
"missing_artifacts": resume_plan["missing_artifacts"],
|
||||||
|
"artifact_resume_from_stage": resume_plan["artifact_resume_from_stage"],
|
||||||
|
"artifact_snapshot": resume_plan["artifact_snapshot"],
|
||||||
|
"state_conflict": record.get("state_conflict", False),
|
||||||
|
"state_conflict_reason": record.get("state_conflict_reason"),
|
||||||
}
|
}
|
||||||
|
|
||||||
if resume_from_stage in UNSUPPORTED_RESUME_STAGES:
|
if resume_from_stage in UNSUPPORTED_RESUME_STAGES:
|
||||||
@@ -109,6 +152,9 @@ def resume_run(*, run_id: str) -> dict[str, Any]:
|
|||||||
**base_response,
|
**base_response,
|
||||||
"resumed": False,
|
"resumed": False,
|
||||||
"resume_from_stage": resume_from_stage,
|
"resume_from_stage": resume_from_stage,
|
||||||
|
"requested_resume_from_stage": requested_resume_from_stage,
|
||||||
|
"resume_decision_source": resume_plan["decision_source"],
|
||||||
|
"recommended_action": "start_new_run",
|
||||||
"message": f"This run cannot be resumed from {resume_from_stage} in the current minimal implementation.",
|
"message": f"This run cannot be resumed from {resume_from_stage} in the current minimal implementation.",
|
||||||
"missing_artifacts": [],
|
"missing_artifacts": [],
|
||||||
}
|
}
|
||||||
@@ -118,6 +164,9 @@ def resume_run(*, run_id: str) -> dict[str, Any]:
|
|||||||
**base_response,
|
**base_response,
|
||||||
"resumed": False,
|
"resumed": False,
|
||||||
"resume_from_stage": resume_from_stage,
|
"resume_from_stage": resume_from_stage,
|
||||||
|
"requested_resume_from_stage": requested_resume_from_stage,
|
||||||
|
"resume_decision_source": resume_plan["decision_source"],
|
||||||
|
"recommended_action": "start_new_run",
|
||||||
"message": f"This run cannot be resumed because stage '{resume_from_stage}' is not supported.",
|
"message": f"This run cannot be resumed because stage '{resume_from_stage}' is not supported.",
|
||||||
"missing_artifacts": [],
|
"missing_artifacts": [],
|
||||||
}
|
}
|
||||||
@@ -128,8 +177,13 @@ def resume_run(*, run_id: str) -> dict[str, Any]:
|
|||||||
**base_response,
|
**base_response,
|
||||||
"resumed": False,
|
"resumed": False,
|
||||||
"resume_from_stage": resume_from_stage,
|
"resume_from_stage": resume_from_stage,
|
||||||
|
"requested_resume_from_stage": requested_resume_from_stage,
|
||||||
|
"resume_decision_source": resume_plan["decision_source"],
|
||||||
|
"recommended_action": "start_new_run",
|
||||||
"message": f"This run cannot be resumed from {resume_from_stage} because required artifacts are missing.",
|
"message": f"This run cannot be resumed from {resume_from_stage} because required artifacts are missing.",
|
||||||
"missing_artifacts": missing_artifacts,
|
"missing_artifacts": missing_artifacts,
|
||||||
|
"artifact_resume_from_stage": resume_plan["artifact_resume_from_stage"],
|
||||||
|
"artifact_snapshot": resume_plan["artifact_snapshot"],
|
||||||
}
|
}
|
||||||
|
|
||||||
try:
|
try:
|
||||||
@@ -141,10 +195,15 @@ def resume_run(*, run_id: str) -> dict[str, Any]:
|
|||||||
**base_response,
|
**base_response,
|
||||||
"resumed": True,
|
"resumed": True,
|
||||||
"resume_from_stage": resume_from_stage,
|
"resume_from_stage": resume_from_stage,
|
||||||
|
"requested_resume_from_stage": requested_resume_from_stage,
|
||||||
|
"resume_decision_source": resume_plan["decision_source"],
|
||||||
|
"recommended_action": "inspect_error",
|
||||||
"status": run_store.state.status,
|
"status": run_store.state.status,
|
||||||
"message": f"Run resumed from {resume_from_stage} but failed again at {failed_stage}: {error}",
|
"message": f"Run resumed from {resume_from_stage} but failed again at {failed_stage}: {error}",
|
||||||
"missing_artifacts": [],
|
"missing_artifacts": [],
|
||||||
"error_summary": run_store.state.error.model_dump(mode="json") if run_store.state.error else None,
|
"error_summary": run_store.state.error.model_dump(mode="json") if run_store.state.error else None,
|
||||||
|
"artifact_resume_from_stage": resume_plan["artifact_resume_from_stage"],
|
||||||
|
"artifact_snapshot": resume_plan["artifact_snapshot"],
|
||||||
}
|
}
|
||||||
|
|
||||||
return {
|
return {
|
||||||
@@ -152,6 +211,9 @@ def resume_run(*, run_id: str) -> dict[str, Any]:
|
|||||||
"workflow": run_store.state.workflow,
|
"workflow": run_store.state.workflow,
|
||||||
"resumed": True,
|
"resumed": True,
|
||||||
"resume_from_stage": resume_from_stage,
|
"resume_from_stage": resume_from_stage,
|
||||||
|
"requested_resume_from_stage": requested_resume_from_stage,
|
||||||
|
"resume_decision_source": resume_plan["decision_source"],
|
||||||
|
"recommended_action": "none",
|
||||||
"status": result["status"],
|
"status": result["status"],
|
||||||
"output_dir": _normalize_repo_path(record["run_dir"]),
|
"output_dir": _normalize_repo_path(record["run_dir"]),
|
||||||
"message": f"Run resumed from {resume_from_stage} and completed with status {result['status']}.",
|
"message": f"Run resumed from {resume_from_stage} and completed with status {result['status']}.",
|
||||||
@@ -166,9 +228,92 @@ def resume_run(*, run_id: str) -> dict[str, Any]:
|
|||||||
"marked_read_count": result.get("marked_read_count"),
|
"marked_read_count": result.get("marked_read_count"),
|
||||||
"status_counts": result.get("status_counts"),
|
"status_counts": result.get("status_counts"),
|
||||||
"missing_artifacts": [],
|
"missing_artifacts": [],
|
||||||
|
"artifact_resume_from_stage": resume_plan["artifact_resume_from_stage"],
|
||||||
|
"artifact_snapshot": resume_plan["artifact_snapshot"],
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def inspect_resume_plan(*, run_id: str) -> dict[str, Any]:
|
||||||
|
record = _resolve_run_record(run_id)
|
||||||
|
base_response = {
|
||||||
|
"run_id": record["run_id"],
|
||||||
|
"workflow": record["workflow"],
|
||||||
|
"status": record["status"],
|
||||||
|
"output_dir": _normalize_repo_path(record["run_dir"]),
|
||||||
|
"state_source": record["state_source"],
|
||||||
|
"status_source": record.get("status_source"),
|
||||||
|
"state_conflict": record.get("state_conflict", False),
|
||||||
|
"state_conflict_reason": record.get("state_conflict_reason"),
|
||||||
|
}
|
||||||
|
|
||||||
|
if record["status"] in {"success", "partial"} and isinstance(record.get("report"), dict):
|
||||||
|
return {
|
||||||
|
**base_response,
|
||||||
|
"can_resume": False,
|
||||||
|
"resume_from_stage": None,
|
||||||
|
"requested_resume_from_stage": None,
|
||||||
|
"resume_decision_source": "artifacts",
|
||||||
|
"recommended_action": "read_terminal_result",
|
||||||
|
"message": "This run already has a terminal run-report.json; prefer reading get_run_status/get_run_report instead of resuming.",
|
||||||
|
"missing_artifacts": [],
|
||||||
|
}
|
||||||
|
|
||||||
|
if record["state_source"] != "run_state" or not isinstance(record.get("state"), RunState):
|
||||||
|
return {
|
||||||
|
**base_response,
|
||||||
|
"can_resume": False,
|
||||||
|
"resume_from_stage": None,
|
||||||
|
"requested_resume_from_stage": None,
|
||||||
|
"resume_decision_source": "unavailable",
|
||||||
|
"recommended_action": "start_new_run",
|
||||||
|
"message": "This run cannot be resumed because run-state.json is missing or could not be loaded.",
|
||||||
|
"missing_artifacts": [],
|
||||||
|
}
|
||||||
|
|
||||||
|
state = record["state"]
|
||||||
|
resume_plan = _build_resume_plan(record=record, state=state)
|
||||||
|
response = {
|
||||||
|
**base_response,
|
||||||
|
"can_resume": False,
|
||||||
|
"resume_from_stage": resume_plan["effective_resume_from_stage"],
|
||||||
|
"requested_resume_from_stage": resume_plan["requested_resume_from_stage"],
|
||||||
|
"resume_decision_source": resume_plan["decision_source"],
|
||||||
|
"recommended_action": resume_plan["recommended_action"],
|
||||||
|
"message": resume_plan["message"],
|
||||||
|
"missing_artifacts": resume_plan["missing_artifacts"],
|
||||||
|
"artifact_resume_from_stage": resume_plan["artifact_resume_from_stage"],
|
||||||
|
"artifact_snapshot": resume_plan["artifact_snapshot"],
|
||||||
|
}
|
||||||
|
|
||||||
|
if state.workflow != WORKFLOW_NAME:
|
||||||
|
response["message"] = (
|
||||||
|
f"This run cannot be resumed because workflow '{state.workflow}' is not supported by the minimal resume_run implementation."
|
||||||
|
)
|
||||||
|
return response
|
||||||
|
|
||||||
|
resume_from_stage = resume_plan["effective_resume_from_stage"]
|
||||||
|
if resume_plan["decision"] != "resume":
|
||||||
|
return response
|
||||||
|
if resume_from_stage in UNSUPPORTED_RESUME_STAGES:
|
||||||
|
response["message"] = f"This run cannot be resumed from {resume_from_stage} in the current minimal implementation."
|
||||||
|
response["recommended_action"] = "start_new_run"
|
||||||
|
return response
|
||||||
|
if resume_from_stage not in SUPPORTED_RESUME_STAGES:
|
||||||
|
response["message"] = f"This run cannot be resumed because stage '{resume_from_stage}' is not supported."
|
||||||
|
response["recommended_action"] = "start_new_run"
|
||||||
|
return response
|
||||||
|
|
||||||
|
missing_artifacts = _validate_resume_artifacts(record=record, state=state, resume_from_stage=resume_from_stage)
|
||||||
|
if missing_artifacts:
|
||||||
|
response["message"] = f"This run cannot be resumed from {resume_from_stage} because required artifacts are missing."
|
||||||
|
response["missing_artifacts"] = missing_artifacts
|
||||||
|
response["recommended_action"] = "start_new_run"
|
||||||
|
return response
|
||||||
|
|
||||||
|
response["can_resume"] = True
|
||||||
|
return response
|
||||||
|
|
||||||
|
|
||||||
def _resume_freshrss_run(*, record: dict[str, Any], run_store: RunStore, resume_from_stage: str) -> dict[str, Any]:
|
def _resume_freshrss_run(*, record: dict[str, Any], run_store: RunStore, resume_from_stage: str) -> dict[str, Any]:
|
||||||
state = run_store.state
|
state = run_store.state
|
||||||
run_dir = record["run_dir"]
|
run_dir = record["run_dir"]
|
||||||
@@ -304,6 +449,75 @@ def _build_resume_config(*, state: RunState, run_dir: Path) -> dict[str, Any]:
|
|||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _find_summary_batch_path(*, run_dir: Path, state: RunState) -> Path | None:
|
||||||
|
return _find_artifact_path(
|
||||||
|
run_dir=run_dir,
|
||||||
|
state=state,
|
||||||
|
artifact_name=SUMMARY_BATCH_ARTIFACT,
|
||||||
|
relative_path=Path("summary/summary-batch.json"),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _find_candidate_batch_path(*, run_dir: Path, state: RunState) -> Path | None:
|
||||||
|
return _find_artifact_path(
|
||||||
|
run_dir=run_dir,
|
||||||
|
state=state,
|
||||||
|
artifact_name=CANDIDATE_BATCH_ARTIFACT,
|
||||||
|
relative_path=Path("candidates/candidate-batch.json"),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _load_summary_batch_lookup(path: Path | None) -> tuple[bool, dict[str, dict[str, Any]]]:
|
||||||
|
if path is None or not path.exists():
|
||||||
|
return False, {}
|
||||||
|
try:
|
||||||
|
payload = _load_json(path)
|
||||||
|
except Exception:
|
||||||
|
return False, {}
|
||||||
|
|
||||||
|
items = payload.get("items")
|
||||||
|
if not isinstance(items, list):
|
||||||
|
return False, {}
|
||||||
|
|
||||||
|
summaries_by_item_key: dict[str, dict[str, Any]] = {}
|
||||||
|
for entry in items:
|
||||||
|
if not isinstance(entry, dict):
|
||||||
|
return False, {}
|
||||||
|
item_key = entry.get("item_key")
|
||||||
|
summary = entry.get("summary")
|
||||||
|
if not isinstance(item_key, str) or not item_key.strip() or not isinstance(summary, dict):
|
||||||
|
return False, {}
|
||||||
|
summaries_by_item_key[item_key] = summary
|
||||||
|
return True, summaries_by_item_key
|
||||||
|
|
||||||
|
|
||||||
|
def _load_candidate_batch_lookup(path: Path | None) -> tuple[bool, dict[str, OpenClawCandidateInput]]:
|
||||||
|
if path is None or not path.exists():
|
||||||
|
return False, {}
|
||||||
|
try:
|
||||||
|
payload = _load_json(path)
|
||||||
|
except Exception:
|
||||||
|
return False, {}
|
||||||
|
|
||||||
|
items = payload.get("items")
|
||||||
|
if not isinstance(items, list):
|
||||||
|
return False, {}
|
||||||
|
|
||||||
|
candidates_by_item_key: dict[str, OpenClawCandidateInput] = {}
|
||||||
|
try:
|
||||||
|
for entry in items:
|
||||||
|
if not isinstance(entry, dict):
|
||||||
|
return False, {}
|
||||||
|
item_key = entry.get("item_key")
|
||||||
|
candidate_payload = entry.get("candidate")
|
||||||
|
if not isinstance(item_key, str) or not item_key.strip() or not isinstance(candidate_payload, dict):
|
||||||
|
return False, {}
|
||||||
|
candidates_by_item_key[item_key] = OpenClawCandidateInput.model_validate(candidate_payload)
|
||||||
|
except Exception:
|
||||||
|
return False, {}
|
||||||
|
return True, candidates_by_item_key
|
||||||
|
|
||||||
|
|
||||||
def _load_item_contexts(
|
def _load_item_contexts(
|
||||||
*,
|
*,
|
||||||
items: list[Any],
|
items: list[Any],
|
||||||
@@ -313,6 +527,8 @@ def _load_item_contexts(
|
|||||||
include_candidates: bool = False,
|
include_candidates: bool = False,
|
||||||
delivery_candidates_by_id: dict[str, OpenClawCandidateInput] | None = None,
|
delivery_candidates_by_id: dict[str, OpenClawCandidateInput] | None = None,
|
||||||
) -> list[dict[str, Any]]:
|
) -> list[dict[str, Any]]:
|
||||||
|
summary_batch_valid, summary_batch_by_item_key = _load_summary_batch_lookup(_summary_batch_output(run_dir))
|
||||||
|
candidate_batch_valid, candidate_batch_by_item_key = _load_candidate_batch_lookup(_candidate_batch_output(run_dir))
|
||||||
contexts: list[dict[str, Any]] = []
|
contexts: list[dict[str, Any]] = []
|
||||||
for index, item in enumerate(items, start=1):
|
for index, item in enumerate(items, start=1):
|
||||||
context = _build_item_context(
|
context = _build_item_context(
|
||||||
@@ -334,16 +550,24 @@ def _load_item_contexts(
|
|||||||
else:
|
else:
|
||||||
context["item_report"]["status"] = "extracted"
|
context["item_report"]["status"] = "extracted"
|
||||||
|
|
||||||
if include_summaries and context["summary_output"] is not None and context["summary_output"].exists():
|
if include_summaries:
|
||||||
summary_payload = _load_json(context["summary_output"])
|
summary_payload: dict[str, Any] | None = None
|
||||||
context["summary_payload"] = summary_payload
|
if context["summary_output"] is not None and context["summary_output"].exists():
|
||||||
context["item_report"]["status"] = "summarized"
|
summary_payload = _load_json(context["summary_output"])
|
||||||
elif include_summaries and context.get("extraction") is not None and context["extraction"].success:
|
elif summary_batch_valid:
|
||||||
context["item_report"]["status"] = "summary_failed"
|
summary_payload = summary_batch_by_item_key.get(context["item_key"])
|
||||||
|
|
||||||
|
if summary_payload is not None:
|
||||||
|
context["summary_payload"] = summary_payload
|
||||||
|
context["item_report"]["status"] = "summarized"
|
||||||
|
elif context.get("extraction") is not None and context["extraction"].success:
|
||||||
|
context["item_report"]["status"] = "summary_failed"
|
||||||
|
|
||||||
candidate: OpenClawCandidateInput | None = None
|
candidate: OpenClawCandidateInput | None = None
|
||||||
if include_candidates and context["openclaw_path"] is not None and context["openclaw_path"].exists():
|
if include_candidates and context["openclaw_path"] is not None and context["openclaw_path"].exists():
|
||||||
candidate = OpenClawCandidateInput.model_validate(_load_json(context["openclaw_path"]))
|
candidate = OpenClawCandidateInput.model_validate(_load_json(context["openclaw_path"]))
|
||||||
|
elif include_candidates and candidate_batch_valid:
|
||||||
|
candidate = candidate_batch_by_item_key.get(context["item_key"])
|
||||||
elif delivery_candidates_by_id is not None and context.get("extraction") is not None and context["extraction"].success:
|
elif delivery_candidates_by_id is not None and context.get("extraction") is not None and context["extraction"].success:
|
||||||
candidate_id = candidate_id_for(item, context["extraction"].article)
|
candidate_id = candidate_id_for(item, context["extraction"].article)
|
||||||
candidate = delivery_candidates_by_id.get(candidate_id)
|
candidate = delivery_candidates_by_id.get(candidate_id)
|
||||||
@@ -406,6 +630,11 @@ def _run_summary_stage(*, run_store: RunStore, item_contexts: list[dict[str, Any
|
|||||||
},
|
},
|
||||||
)
|
)
|
||||||
|
|
||||||
|
summary_batch_output = _persist_summary_batch_artifact(
|
||||||
|
run_store=run_store,
|
||||||
|
run_dir=config["run_dir"],
|
||||||
|
item_contexts=item_contexts,
|
||||||
|
)
|
||||||
if config["debug_artifacts"] and (config["run_dir"] / "summary").exists():
|
if config["debug_artifacts"] and (config["run_dir"] / "summary").exists():
|
||||||
run_store.register_artifact(name="summary_dir", path=config["run_dir"] / "summary", kind="directory", stage=SUMMARY_STAGE)
|
run_store.register_artifact(name="summary_dir", path=config["run_dir"] / "summary", kind="directory", stage=SUMMARY_STAGE)
|
||||||
run_store.finish_stage(
|
run_store.finish_stage(
|
||||||
@@ -415,6 +644,7 @@ def _run_summary_stage(*, run_store: RunStore, item_contexts: list[dict[str, Any
|
|||||||
"completed_items": summary_success_count + summary_failed_count,
|
"completed_items": summary_success_count + summary_failed_count,
|
||||||
"success_count": summary_success_count,
|
"success_count": summary_success_count,
|
||||||
"failed_count": summary_failed_count,
|
"failed_count": summary_failed_count,
|
||||||
|
"summary_batch_output": str(summary_batch_output),
|
||||||
},
|
},
|
||||||
)
|
)
|
||||||
|
|
||||||
@@ -500,6 +730,11 @@ def _run_filter_stage(*, run_store: RunStore, item_contexts: list[dict[str, Any]
|
|||||||
},
|
},
|
||||||
)
|
)
|
||||||
|
|
||||||
|
candidate_batch_output = _persist_candidate_batch_artifact(
|
||||||
|
run_store=run_store,
|
||||||
|
run_dir=config["run_dir"],
|
||||||
|
item_contexts=item_contexts,
|
||||||
|
)
|
||||||
if config["debug_artifacts"] and (config["run_dir"] / "candidates").exists():
|
if config["debug_artifacts"] and (config["run_dir"] / "candidates").exists():
|
||||||
run_store.register_artifact(name="candidate_dir", path=config["run_dir"] / "candidates", kind="directory", stage=FILTER_STAGE)
|
run_store.register_artifact(name="candidate_dir", path=config["run_dir"] / "candidates", kind="directory", stage=FILTER_STAGE)
|
||||||
run_store.finish_stage(
|
run_store.finish_stage(
|
||||||
@@ -511,6 +746,7 @@ def _run_filter_stage(*, run_store: RunStore, item_contexts: list[dict[str, Any]
|
|||||||
"keep_count": keep_count,
|
"keep_count": keep_count,
|
||||||
"review_count": review_count,
|
"review_count": review_count,
|
||||||
"drop_count": drop_count,
|
"drop_count": drop_count,
|
||||||
|
"candidate_batch_output": str(candidate_batch_output),
|
||||||
},
|
},
|
||||||
)
|
)
|
||||||
|
|
||||||
@@ -682,6 +918,165 @@ def _build_keyword_index_result(
|
|||||||
return result
|
return result
|
||||||
|
|
||||||
|
|
||||||
|
def _build_resume_plan(*, record: dict[str, Any], state: RunState) -> dict[str, Any]:
|
||||||
|
requested_resume_from_stage = _resolve_resume_from_stage(state)
|
||||||
|
artifact_snapshot = _collect_resume_artifact_snapshot(record=record, state=state)
|
||||||
|
artifact_resume_from_stage = _resolve_resume_stage_from_artifacts(artifact_snapshot)
|
||||||
|
|
||||||
|
if artifact_snapshot["has_run_report"]:
|
||||||
|
return {
|
||||||
|
"decision": "reject_terminal",
|
||||||
|
"decision_source": "artifacts",
|
||||||
|
"requested_resume_from_stage": requested_resume_from_stage,
|
||||||
|
"effective_resume_from_stage": None,
|
||||||
|
"artifact_resume_from_stage": artifact_resume_from_stage,
|
||||||
|
"recommended_action": "read_terminal_result",
|
||||||
|
"message": "This run already has a terminal run-report.json; prefer reading get_run_status/get_run_report instead of resuming.",
|
||||||
|
"missing_artifacts": [],
|
||||||
|
"artifact_snapshot": artifact_snapshot,
|
||||||
|
}
|
||||||
|
|
||||||
|
if artifact_resume_from_stage is None:
|
||||||
|
return {
|
||||||
|
"decision": "reject_unrecoverable",
|
||||||
|
"decision_source": "artifacts",
|
||||||
|
"requested_resume_from_stage": requested_resume_from_stage,
|
||||||
|
"effective_resume_from_stage": None,
|
||||||
|
"artifact_resume_from_stage": None,
|
||||||
|
"recommended_action": "start_new_run",
|
||||||
|
"message": "This run does not expose a safe artifact-backed resume point; start a new run instead.",
|
||||||
|
"missing_artifacts": artifact_snapshot["missing_for_next_resume"],
|
||||||
|
"artifact_snapshot": artifact_snapshot,
|
||||||
|
}
|
||||||
|
|
||||||
|
decision_source = "artifacts"
|
||||||
|
message = f"Resume will continue from {artifact_resume_from_stage} based on available artifacts."
|
||||||
|
if requested_resume_from_stage == artifact_resume_from_stage:
|
||||||
|
decision_source = "state_and_artifacts"
|
||||||
|
message = f"Resume stage {artifact_resume_from_stage} was confirmed by both run-state.json and artifacts."
|
||||||
|
elif requested_resume_from_stage is not None:
|
||||||
|
requested_rank = RESUME_STAGE_ORDER.get(requested_resume_from_stage, -1)
|
||||||
|
artifact_rank = RESUME_STAGE_ORDER.get(artifact_resume_from_stage, -1)
|
||||||
|
if artifact_rank > requested_rank:
|
||||||
|
message = (
|
||||||
|
f"Resume stage was advanced from {requested_resume_from_stage} to {artifact_resume_from_stage} "
|
||||||
|
f"because artifacts prove the run already progressed further."
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
message = (
|
||||||
|
f"Resume stage was moved back from {requested_resume_from_stage} to {artifact_resume_from_stage} "
|
||||||
|
f"because later-stage artifacts are not stable enough for a safe resume."
|
||||||
|
)
|
||||||
|
|
||||||
|
return {
|
||||||
|
"decision": "resume",
|
||||||
|
"decision_source": decision_source,
|
||||||
|
"requested_resume_from_stage": requested_resume_from_stage,
|
||||||
|
"effective_resume_from_stage": artifact_resume_from_stage,
|
||||||
|
"artifact_resume_from_stage": artifact_resume_from_stage,
|
||||||
|
"recommended_action": "resume",
|
||||||
|
"message": message,
|
||||||
|
"missing_artifacts": [],
|
||||||
|
"artifact_snapshot": artifact_snapshot,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _collect_resume_artifact_snapshot(*, record: dict[str, Any], state: RunState) -> dict[str, Any]:
|
||||||
|
run_dir = record["run_dir"]
|
||||||
|
raw_output = _find_artifact_path(run_dir=run_dir, state=state, artifact_name="raw_output", relative_path=Path("raw/freshrss.raw.json"))
|
||||||
|
extracted_dir = _find_artifact_path(run_dir=run_dir, state=state, artifact_name="extracted_dir", relative_path=Path("extracted"))
|
||||||
|
summary_batch_output = _find_summary_batch_path(run_dir=run_dir, state=state)
|
||||||
|
candidate_batch_output = _find_candidate_batch_path(run_dir=run_dir, state=state)
|
||||||
|
delivery_output = _find_artifact_path(
|
||||||
|
run_dir=run_dir,
|
||||||
|
state=state,
|
||||||
|
artifact_name="delivery_payload",
|
||||||
|
relative_path=Path("candidates/openclaw-delivery-payload.json"),
|
||||||
|
)
|
||||||
|
digest_brief_output = _find_artifact_path(
|
||||||
|
run_dir=run_dir,
|
||||||
|
state=state,
|
||||||
|
artifact_name="digest_brief",
|
||||||
|
relative_path=Path("candidates/digest-brief.json"),
|
||||||
|
)
|
||||||
|
report_output = run_dir / "run-report.json"
|
||||||
|
|
||||||
|
items: list[Any] = []
|
||||||
|
raw_output_valid = False
|
||||||
|
if raw_output is not None:
|
||||||
|
try:
|
||||||
|
items = _load_items(raw_output)
|
||||||
|
raw_output_valid = True
|
||||||
|
except Exception:
|
||||||
|
items = []
|
||||||
|
raw_output_valid = False
|
||||||
|
|
||||||
|
summary_batch_valid, _ = _load_summary_batch_lookup(summary_batch_output)
|
||||||
|
candidate_batch_valid, _ = _load_candidate_batch_lookup(candidate_batch_output)
|
||||||
|
item_contexts = _load_item_contexts(
|
||||||
|
items=items,
|
||||||
|
run_dir=run_dir,
|
||||||
|
debug_artifacts=bool(state.input.get("debug_artifacts", False)),
|
||||||
|
include_summaries=True,
|
||||||
|
include_candidates=True,
|
||||||
|
)
|
||||||
|
extraction_success_count = sum(
|
||||||
|
1 for context in item_contexts if context.get("extraction") is not None and context["extraction"].success
|
||||||
|
)
|
||||||
|
summary_count = sum(1 for context in item_contexts if context.get("summary_payload") is not None)
|
||||||
|
candidate_count = sum(1 for context in item_contexts if context.get("candidate") is not None)
|
||||||
|
extracted_complete = raw_output_valid and bool(items) and all(context["extracted_path"].exists() for context in item_contexts)
|
||||||
|
stable_summary_outputs = summary_batch_valid or (extraction_success_count > 0 and extraction_success_count == summary_count)
|
||||||
|
stable_candidate_outputs = candidate_batch_valid or (summary_count > 0 and summary_count == candidate_count)
|
||||||
|
|
||||||
|
missing_for_next_resume: list[str] = []
|
||||||
|
if raw_output is None or not raw_output_valid:
|
||||||
|
missing_for_next_resume.append("raw/freshrss.raw.json")
|
||||||
|
if extracted_dir is None or not extracted_complete:
|
||||||
|
missing_for_next_resume.append("extracted/")
|
||||||
|
if extraction_success_count > 0 and not stable_summary_outputs and not stable_candidate_outputs:
|
||||||
|
missing_for_next_resume.append("summary/summary-batch.json")
|
||||||
|
if extraction_success_count > 0 and stable_summary_outputs and not stable_candidate_outputs:
|
||||||
|
missing_for_next_resume.append("candidates/candidate-batch.json")
|
||||||
|
|
||||||
|
return {
|
||||||
|
"has_raw_output": raw_output is not None,
|
||||||
|
"raw_output_valid": raw_output_valid,
|
||||||
|
"has_extracted_dir": extracted_dir is not None,
|
||||||
|
"has_summary_batch": summary_batch_output is not None,
|
||||||
|
"summary_batch_valid": summary_batch_valid,
|
||||||
|
"has_candidate_batch": candidate_batch_output is not None,
|
||||||
|
"candidate_batch_valid": candidate_batch_valid,
|
||||||
|
"has_delivery_payload": delivery_output is not None,
|
||||||
|
"has_digest_brief": digest_brief_output is not None,
|
||||||
|
"has_run_report": report_output.exists(),
|
||||||
|
"item_count": len(items),
|
||||||
|
"extraction_success_count": extraction_success_count,
|
||||||
|
"summary_count": summary_count,
|
||||||
|
"candidate_count": candidate_count,
|
||||||
|
"extracted_complete": extracted_complete,
|
||||||
|
"stable_summary_outputs": stable_summary_outputs,
|
||||||
|
"stable_candidate_outputs": stable_candidate_outputs,
|
||||||
|
"missing_for_next_resume": missing_for_next_resume,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _resolve_resume_stage_from_artifacts(snapshot: dict[str, Any]) -> str | None:
|
||||||
|
if snapshot["has_run_report"]:
|
||||||
|
return None
|
||||||
|
if snapshot["has_delivery_payload"]:
|
||||||
|
return REPORT_STAGE
|
||||||
|
if snapshot["stable_candidate_outputs"]:
|
||||||
|
return DELIVERY_STAGE
|
||||||
|
if snapshot["has_raw_output"] and snapshot["raw_output_valid"] and snapshot["has_extracted_dir"] and snapshot["extracted_complete"]:
|
||||||
|
if snapshot["extraction_success_count"] <= 0:
|
||||||
|
return SUMMARY_STAGE
|
||||||
|
if snapshot["stable_summary_outputs"]:
|
||||||
|
return FILTER_STAGE
|
||||||
|
return SUMMARY_STAGE
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
def _load_filter_context(config: dict[str, Any]) -> FilterContext:
|
def _load_filter_context(config: dict[str, Any]) -> FilterContext:
|
||||||
if config["context"] is not None:
|
if config["context"] is not None:
|
||||||
return FilterContext.model_validate(config["context"])
|
return FilterContext.model_validate(config["context"])
|
||||||
@@ -695,6 +1090,8 @@ def _validate_resume_artifacts(*, record: dict[str, Any], state: RunState, resum
|
|||||||
missing_artifacts: list[str] = []
|
missing_artifacts: list[str] = []
|
||||||
raw_output = _find_artifact_path(run_dir=run_dir, state=state, artifact_name="raw_output", relative_path=Path("raw/freshrss.raw.json"))
|
raw_output = _find_artifact_path(run_dir=run_dir, state=state, artifact_name="raw_output", relative_path=Path("raw/freshrss.raw.json"))
|
||||||
extracted_dir = _find_artifact_path(run_dir=run_dir, state=state, artifact_name="extracted_dir", relative_path=Path("extracted"))
|
extracted_dir = _find_artifact_path(run_dir=run_dir, state=state, artifact_name="extracted_dir", relative_path=Path("extracted"))
|
||||||
|
summary_batch_output = _find_summary_batch_path(run_dir=run_dir, state=state)
|
||||||
|
candidate_batch_output = _find_candidate_batch_path(run_dir=run_dir, state=state)
|
||||||
if resume_from_stage in {SUMMARY_STAGE, FILTER_STAGE, DELIVERY_STAGE} and raw_output is None:
|
if resume_from_stage in {SUMMARY_STAGE, FILTER_STAGE, DELIVERY_STAGE} and raw_output is None:
|
||||||
missing_artifacts.append("raw/freshrss.raw.json")
|
missing_artifacts.append("raw/freshrss.raw.json")
|
||||||
if resume_from_stage in {SUMMARY_STAGE, FILTER_STAGE} and extracted_dir is None:
|
if resume_from_stage in {SUMMARY_STAGE, FILTER_STAGE} and extracted_dir is None:
|
||||||
@@ -711,6 +1108,8 @@ def _validate_resume_artifacts(*, record: dict[str, Any], state: RunState, resum
|
|||||||
include_summaries=resume_from_stage in {FILTER_STAGE, DELIVERY_STAGE},
|
include_summaries=resume_from_stage in {FILTER_STAGE, DELIVERY_STAGE},
|
||||||
include_candidates=resume_from_stage == DELIVERY_STAGE,
|
include_candidates=resume_from_stage == DELIVERY_STAGE,
|
||||||
)
|
)
|
||||||
|
summary_batch_valid, _ = _load_summary_batch_lookup(summary_batch_output)
|
||||||
|
candidate_batch_valid, _ = _load_candidate_batch_lookup(candidate_batch_output)
|
||||||
|
|
||||||
if resume_from_stage == SUMMARY_STAGE:
|
if resume_from_stage == SUMMARY_STAGE:
|
||||||
for context in item_contexts:
|
for context in item_contexts:
|
||||||
@@ -718,17 +1117,26 @@ def _validate_resume_artifacts(*, record: dict[str, Any], state: RunState, resum
|
|||||||
missing_artifacts.append(_normalize_repo_path(context["extracted_path"]))
|
missing_artifacts.append(_normalize_repo_path(context["extracted_path"]))
|
||||||
|
|
||||||
if resume_from_stage == FILTER_STAGE:
|
if resume_from_stage == FILTER_STAGE:
|
||||||
for context in item_contexts:
|
expected_summary_count = int(_stage_output(state, SUMMARY_STAGE, "success_count") or 0)
|
||||||
if context.get("extraction") is not None and context["extraction"].success:
|
actual_summary_count = sum(1 for context in item_contexts if context.get("summary_payload") is not None)
|
||||||
if context["summary_output"] is None or not context["summary_output"].exists():
|
if summary_batch_valid:
|
||||||
missing_artifacts.append(
|
if expected_summary_count != actual_summary_count:
|
||||||
_normalize_repo_path(context["summary_output"] or (run_dir / "summary" / context["item_key"] / "result.loop.json"))
|
missing_artifacts.append(_normalize_repo_path(summary_batch_output or _summary_batch_output(run_dir)))
|
||||||
)
|
else:
|
||||||
|
for context in item_contexts:
|
||||||
|
if context.get("extraction") is not None and context["extraction"].success:
|
||||||
|
if context["summary_output"] is None or not context["summary_output"].exists():
|
||||||
|
missing_artifacts.append(
|
||||||
|
_normalize_repo_path(context["summary_output"] or (run_dir / "summary" / context["item_key"] / "result.loop.json"))
|
||||||
|
)
|
||||||
|
|
||||||
if resume_from_stage == DELIVERY_STAGE:
|
if resume_from_stage == DELIVERY_STAGE:
|
||||||
expected_candidate_count = int(_stage_output(state, FILTER_STAGE, "candidate_count") or 0)
|
expected_candidate_count = int(_stage_output(state, FILTER_STAGE, "candidate_count") or 0)
|
||||||
actual_candidate_count = sum(1 for context in item_contexts if context.get("candidate") is not None)
|
actual_candidate_count = sum(1 for context in item_contexts if context.get("candidate") is not None)
|
||||||
if expected_candidate_count != actual_candidate_count:
|
if candidate_batch_valid:
|
||||||
|
if expected_candidate_count != actual_candidate_count:
|
||||||
|
missing_artifacts.append(_normalize_repo_path(candidate_batch_output or _candidate_batch_output(run_dir)))
|
||||||
|
elif expected_candidate_count != actual_candidate_count:
|
||||||
for context in item_contexts:
|
for context in item_contexts:
|
||||||
if context.get("summary_payload") is not None and (context["openclaw_path"] is None or not context["openclaw_path"].exists()):
|
if context.get("summary_payload") is not None and (context["openclaw_path"] is None or not context["openclaw_path"].exists()):
|
||||||
missing_artifacts.append(
|
missing_artifacts.append(
|
||||||
|
|||||||
@@ -18,10 +18,14 @@ from summary_mcp.models.summary_io import ExtractionInput
|
|||||||
from summary_mcp.runtime import get_delivery_payload as load_delivery_payload
|
from summary_mcp.runtime import get_delivery_payload as load_delivery_payload
|
||||||
from summary_mcp.runtime import get_freshrss_pipeline_job_result as load_freshrss_pipeline_job_result
|
from summary_mcp.runtime import get_freshrss_pipeline_job_result as load_freshrss_pipeline_job_result
|
||||||
from summary_mcp.runtime import get_freshrss_pipeline_job_status as load_freshrss_pipeline_job_status
|
from summary_mcp.runtime import get_freshrss_pipeline_job_status as load_freshrss_pipeline_job_status
|
||||||
|
from summary_mcp.runtime import get_resume_job_result as load_resume_job_result
|
||||||
|
from summary_mcp.runtime import get_resume_job_status as load_resume_job_status
|
||||||
from summary_mcp.runtime import get_run_status as load_run_status
|
from summary_mcp.runtime import get_run_status as load_run_status
|
||||||
from summary_mcp.runtime import get_run_report as load_run_report
|
from summary_mcp.runtime import get_run_report as load_run_report
|
||||||
|
from summary_mcp.runtime import inspect_resume_plan as load_resume_plan
|
||||||
from summary_mcp.runtime import list_run_artifacts as load_run_artifacts
|
from summary_mcp.runtime import list_run_artifacts as load_run_artifacts
|
||||||
from summary_mcp.runtime import list_runs as load_runs
|
from summary_mcp.runtime import list_runs as load_runs
|
||||||
|
from summary_mcp.runtime import start_resume_job as launch_resume_job
|
||||||
from summary_mcp.runtime import start_freshrss_pipeline_job as launch_freshrss_pipeline_job
|
from summary_mcp.runtime import start_freshrss_pipeline_job as launch_freshrss_pipeline_job
|
||||||
from summary_mcp.runtime.article_summary_jobs import (
|
from summary_mcp.runtime.article_summary_jobs import (
|
||||||
get_article_summary_job_result as load_article_summary_job_result,
|
get_article_summary_job_result as load_article_summary_job_result,
|
||||||
@@ -232,6 +236,30 @@ def resume_run(run_id: str) -> dict:
|
|||||||
return resume_existing_run(run_id=run_id)
|
return resume_existing_run(run_id=run_id)
|
||||||
|
|
||||||
|
|
||||||
|
@mcp.tool()
|
||||||
|
def inspect_resume_plan(run_id: str) -> dict:
|
||||||
|
"""Inspect the effective resume plan for a FreshRSS workflow run without executing it."""
|
||||||
|
return load_resume_plan(run_id=run_id)
|
||||||
|
|
||||||
|
|
||||||
|
@mcp.tool()
|
||||||
|
def start_resume_job(run_id: str) -> dict:
|
||||||
|
"""Start an asynchronous resume job for a resumable FreshRSS workflow run."""
|
||||||
|
return launch_resume_job(run_id=run_id)
|
||||||
|
|
||||||
|
|
||||||
|
@mcp.tool()
|
||||||
|
def get_resume_job_status(job_id: str) -> dict:
|
||||||
|
"""Get the current status of an asynchronous resume job."""
|
||||||
|
return load_resume_job_status(job_id=job_id)
|
||||||
|
|
||||||
|
|
||||||
|
@mcp.tool()
|
||||||
|
def get_resume_job_result(job_id: str) -> dict:
|
||||||
|
"""Get the final result of an asynchronous resume job."""
|
||||||
|
return load_resume_job_result(job_id=job_id)
|
||||||
|
|
||||||
|
|
||||||
@mcp.tool()
|
@mcp.tool()
|
||||||
def start_article_summary_job(
|
def start_article_summary_job(
|
||||||
*,
|
*,
|
||||||
|
|||||||
@@ -51,6 +51,10 @@ SUMMARY_STAGE = "generate_summaries"
|
|||||||
FILTER_STAGE = "apply_filters"
|
FILTER_STAGE = "apply_filters"
|
||||||
DELIVERY_STAGE = "build_delivery_payload"
|
DELIVERY_STAGE = "build_delivery_payload"
|
||||||
REPORT_STAGE = "write_run_report"
|
REPORT_STAGE = "write_run_report"
|
||||||
|
SUMMARY_BATCH_ARTIFACT = "summary_batch"
|
||||||
|
CANDIDATE_BATCH_ARTIFACT = "candidate_batch"
|
||||||
|
SUMMARY_BATCH_FILENAME = "summary-batch.json"
|
||||||
|
CANDIDATE_BATCH_FILENAME = "candidate-batch.json"
|
||||||
|
|
||||||
|
|
||||||
def _save_json(path: Path, payload: dict[str, Any] | list[Any]) -> None:
|
def _save_json(path: Path, payload: dict[str, Any] | list[Any]) -> None:
|
||||||
@@ -137,6 +141,77 @@ def _build_item_context(*, index: int, item: Any, resolved_output_dir: Path, deb
|
|||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _summary_batch_output(run_dir: Path) -> Path:
|
||||||
|
return run_dir / "summary" / SUMMARY_BATCH_FILENAME
|
||||||
|
|
||||||
|
|
||||||
|
def _candidate_batch_output(run_dir: Path) -> Path:
|
||||||
|
return run_dir / "candidates" / CANDIDATE_BATCH_FILENAME
|
||||||
|
|
||||||
|
|
||||||
|
def _build_summary_batch_payload(*, run_id: str, item_contexts: list[dict[str, Any]]) -> dict[str, Any]:
|
||||||
|
items: list[dict[str, Any]] = []
|
||||||
|
for item_context in item_contexts:
|
||||||
|
summary_payload = item_context.get("summary_payload")
|
||||||
|
if summary_payload is None:
|
||||||
|
continue
|
||||||
|
item = item_context["item"]
|
||||||
|
items.append(
|
||||||
|
{
|
||||||
|
"item_key": item_context["item_key"],
|
||||||
|
"item_id": item.item_id,
|
||||||
|
"summary": summary_payload,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
return {
|
||||||
|
"run_id": run_id,
|
||||||
|
"summary_count": len(items),
|
||||||
|
"items": items,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _build_candidate_batch_payload(*, run_id: str, item_contexts: list[dict[str, Any]]) -> dict[str, Any]:
|
||||||
|
items: list[dict[str, Any]] = []
|
||||||
|
for item_context in item_contexts:
|
||||||
|
candidate = item_context.get("candidate")
|
||||||
|
if candidate is None:
|
||||||
|
continue
|
||||||
|
item = item_context["item"]
|
||||||
|
items.append(
|
||||||
|
{
|
||||||
|
"item_key": item_context["item_key"],
|
||||||
|
"item_id": item.item_id,
|
||||||
|
"candidate_id": candidate.candidate_id,
|
||||||
|
"candidate": candidate.model_dump(mode="json"),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
return {
|
||||||
|
"run_id": run_id,
|
||||||
|
"candidate_count": len(items),
|
||||||
|
"items": items,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _persist_summary_batch_artifact(*, run_store: RunStore, run_dir: Path, item_contexts: list[dict[str, Any]]) -> Path:
|
||||||
|
output_path = _summary_batch_output(run_dir)
|
||||||
|
_save_json(
|
||||||
|
output_path,
|
||||||
|
_build_summary_batch_payload(run_id=run_store.state.run_id, item_contexts=item_contexts),
|
||||||
|
)
|
||||||
|
run_store.register_artifact(name=SUMMARY_BATCH_ARTIFACT, path=output_path, kind="json", stage=SUMMARY_STAGE)
|
||||||
|
return output_path
|
||||||
|
|
||||||
|
|
||||||
|
def _persist_candidate_batch_artifact(*, run_store: RunStore, run_dir: Path, item_contexts: list[dict[str, Any]]) -> Path:
|
||||||
|
output_path = _candidate_batch_output(run_dir)
|
||||||
|
_save_json(
|
||||||
|
output_path,
|
||||||
|
_build_candidate_batch_payload(run_id=run_store.state.run_id, item_contexts=item_contexts),
|
||||||
|
)
|
||||||
|
run_store.register_artifact(name=CANDIDATE_BATCH_ARTIFACT, path=output_path, kind="json", stage=FILTER_STAGE)
|
||||||
|
return output_path
|
||||||
|
|
||||||
|
|
||||||
def _build_run_report(
|
def _build_run_report(
|
||||||
*,
|
*,
|
||||||
resolved_run_id: str,
|
resolved_run_id: str,
|
||||||
@@ -418,6 +493,11 @@ def run_freshrss_pipeline(
|
|||||||
},
|
},
|
||||||
)
|
)
|
||||||
|
|
||||||
|
summary_batch_output = _persist_summary_batch_artifact(
|
||||||
|
run_store=run_store,
|
||||||
|
run_dir=resolved_output_dir,
|
||||||
|
item_contexts=item_contexts,
|
||||||
|
)
|
||||||
if debug_artifacts and (resolved_output_dir / "summary").exists():
|
if debug_artifacts and (resolved_output_dir / "summary").exists():
|
||||||
run_store.register_artifact(name="summary_dir", path=resolved_output_dir / "summary", kind="directory", stage=SUMMARY_STAGE)
|
run_store.register_artifact(name="summary_dir", path=resolved_output_dir / "summary", kind="directory", stage=SUMMARY_STAGE)
|
||||||
run_store.finish_stage(
|
run_store.finish_stage(
|
||||||
@@ -427,6 +507,7 @@ def run_freshrss_pipeline(
|
|||||||
"completed_items": summary_success_count + summary_failed_count,
|
"completed_items": summary_success_count + summary_failed_count,
|
||||||
"success_count": summary_success_count,
|
"success_count": summary_success_count,
|
||||||
"failed_count": summary_failed_count,
|
"failed_count": summary_failed_count,
|
||||||
|
"summary_batch_output": str(summary_batch_output),
|
||||||
},
|
},
|
||||||
)
|
)
|
||||||
|
|
||||||
@@ -510,6 +591,11 @@ def run_freshrss_pipeline(
|
|||||||
},
|
},
|
||||||
)
|
)
|
||||||
|
|
||||||
|
candidate_batch_output = _persist_candidate_batch_artifact(
|
||||||
|
run_store=run_store,
|
||||||
|
run_dir=resolved_output_dir,
|
||||||
|
item_contexts=item_contexts,
|
||||||
|
)
|
||||||
if debug_artifacts and (resolved_output_dir / "candidates").exists():
|
if debug_artifacts and (resolved_output_dir / "candidates").exists():
|
||||||
run_store.register_artifact(name="candidate_dir", path=resolved_output_dir / "candidates", kind="directory", stage=FILTER_STAGE)
|
run_store.register_artifact(name="candidate_dir", path=resolved_output_dir / "candidates", kind="directory", stage=FILTER_STAGE)
|
||||||
run_store.finish_stage(
|
run_store.finish_stage(
|
||||||
@@ -521,6 +607,7 @@ def run_freshrss_pipeline(
|
|||||||
"keep_count": keep_count,
|
"keep_count": keep_count,
|
||||||
"review_count": review_count,
|
"review_count": review_count,
|
||||||
"drop_count": drop_count,
|
"drop_count": drop_count,
|
||||||
|
"candidate_batch_output": str(candidate_batch_output),
|
||||||
},
|
},
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user