feat(reader): add async article summary jobs

This commit is contained in:
root
2026-04-10 16:37:43 +08:00
parent 87d18e4263
commit 2053cccfef
8 changed files with 1269 additions and 15 deletions
+59 -11
View File
@@ -9,7 +9,7 @@ pip install -e .
summary-mcp
```
服务当前暴露 11 个工具。
服务当前暴露 14 个工具。
正式 workflow service 相关工具:
@@ -26,7 +26,10 @@ summary-mcp
- `extract_url_content`
- `extract_item_content`
- `filter_summary_result`
- `generate_article_summaries`
- `generate_article_summaries`(同步模式)
- `start_article_summary_job`
- `get_article_summary_job_status`
- `get_article_summary_job_result`
## 正式能力边界
@@ -34,7 +37,9 @@ summary-mcp
- 每次 FreshRSS 主流水线 run 都会在 `outputs/freshrss/rerun/<run_dir>/run-state.json` 落地运行真相
- OpenClaw 正式读取结果应优先使用 `get_delivery_payload` 与 `get_run_report`,而不是自己拼输出目录路径
- `digest-brief.json` 当前会随主流水线产出,但还没有独立的 MCP 读取工具;如需定位它,应通过 `list_run_artifacts` 或 `get_run_report` 返回的信息发现
- `run_freshrss_openclaw_pipeline` 与 `resume_run` 当前都是同步 MCP 调用;仓库里还没有后台队列 / worker / 异步任务管理
- `run_freshrss_openclaw_pipeline` 与 `resume_run` 当前都是同步 MCP 调用;仓库里还没有针对 FreshRSS 主流程的后台队列 / worker / 异步任务管理
- 单篇总结已补上最小异步 job 形态:`start_article_summary_job` / `get_article_summary_job_status` / `get_article_summary_job_result`
- `generate_article_summaries` 仍保留,但定位是同步 debug 路径,而不是 OpenClaw 的正式生产集成入口
## OpenClaw 推荐调用路径
@@ -70,8 +75,10 @@ summary-mcp
相关能力:
- `generate_article_summaries` MCP 工具(基于已有 extracted payload 做单篇总结)
- `start_article_summary_job` / `get_article_summary_job_status` / `get_article_summary_job_result`(正式推荐的最小异步 job 路径)
- `generate_article_summaries` MCP 工具(同步 debug 路径)
- `scripts/run_article_summaries.py` CLI 辅助脚本
- `scripts/run_article_summary_job.py` 后台 runner 入口
## 校验 LLM 摘要结果
@@ -257,7 +264,7 @@ python scripts/build_keyword_index.py ^
- `skills/keyword-cleanup-review/`
构建给 LLM skill 使用的评审数据包(review bundle):
构建给关键词治理流程使用的评审数据包(review bundle,临时工作文件):
```bash
python skills/keyword-cleanup-review/scripts/build_review_bundle.py ^
@@ -266,17 +273,44 @@ python skills/keyword-cleanup-review/scripts/build_review_bundle.py ^
--output outputs/term_index/review/keyword-cleanup-bundle.json
```
现在这个评审数据包(review bundle)还会额外携带治理上下文:
这个评审数据包(review bundle)会额外携带治理上下文:
- 来自 `configs/term_cleanup_policy.json` 的清理阈值
- 当前 watch list(`configs/term_watchlist.json`)
- 最近已应用的变更(`configs/term_change_log.json`)
这个 skill 只负责生成 review 输入与建议,不会自动修改:
接下来可以把 bundle 渲染成正式建议产物(默认走确定性规则,不把 LLM 作为默认路径):
- `term_aliases`
- `term_stopwords`
- `filter_context.personal.json`
```bash
python scripts/generate_term_cleanup_suggestions.py ^
--bundle outputs/term_index/review/keyword-cleanup-bundle.json
```
默认只生成:
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`(唯一正式建议产物,建议短期保留)
如需人工审阅展示稿,再显式加:
```bash
python scripts/generate_term_cleanup_suggestions.py ^
--bundle outputs/term_index/review/keyword-cleanup-bundle.json ^
--emit-markdown
```
这时才会额外生成:
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`(临时展示稿,可按需生成,不必作为长期资产保留)
其中本轮最小版本优先覆盖 `interest_keyword_suggestions` 和 `watch_terms` 主链路;`alias_suggestions` / `stopword_suggestions` 先保持保守。
产物保留策略建议:
- `data/term_index/daily/YYYY-MM-DD.json`、`data/term_index/term_stats.json` 作为事实层长期保留
- `configs/filter_context.personal.json`、`configs/term_watchlist.json`、`configs/term_aliases.json`、`configs/term_stopwords.json`、`configs/term_change_log.json` 作为状态层长期保留
- `term-cleanup-suggestions-YYYY-MM-DD.json` 作为正式建议产物短期保留(最近少量几份或仅保留已应用过的)
- `keyword-cleanup-bundle.json` 仅作为临时工作文件,默认只保留当前最新一份
- `term-cleanup-suggestions-YYYY-MM-DD.md` 仅作为临时展示稿,优先按需生成,不建议默认长期归档
如果你想先预览已接受建议,再决定是否写配置文件:
@@ -324,12 +358,26 @@ python scripts/run_article_summaries.py ^
也可以通过 `summary_mcp.server` 暴露的 MCP 工具 `generate_article_summaries` 调用:
- `extracted_path`(string):单篇 extracted JSON 路径(例如 `outputs/freshrss/rerun/<run_id>/extracted/item-01.extracted.json`),或者包含 `results` 数组的 batch extracted JSON
- `selected_ids`(array of strings):要总结的一个或多个 `item_id`。如果传空数组,则对文件中的全部条目做总结
- `selected_ids`(array of strings):必填,至少传一个 `item_id`;如果传入的 ID 在 extracted payload 中一个都匹配不到,会直接报错
- `output_dir`(optional string):Markdown 输出目录;若不传,则默认写到 extracted 文件旁边的 `single_summaries/` 目录
- `llm_api_key` / `llm_model` / `llm_api_url`(optional strings):单篇总结 LLM 的覆盖配置;不传时会按前文规则回退到 `ARTICLE_SUMMARY_*` 或主 `LLM_*`
该工具返回一个 JSON 数组,内容为生成好的 Markdown 文件路径。
OpenClaw / 正式集成建议优先走异步 job:
- 调 `start_article_summary_job` 启动任务,立即拿到 `job_id`
- 轮询 `get_article_summary_job_status(job_id)`,直到 `status` 进入 `success` 或 `failed`
- 成功后调用 `get_article_summary_job_result(job_id)` 读取 `written_paths` 与结构化结果
- 失败时优先查看 `error_summary` 与 job 目录中的 `job-report.json`
异步 job 状态目录固定落在 `outputs/freshrss/article_summary_jobs/<job_id>/`,最小会包含:
- `run-state.json`
- `input.json`
- `result.json`(成功时)
- `job-report.json`
单篇总结使用独立 prompt:`outputs/prompts/article-summary-prompt.txt`。
它与日报 prompt 完全独立,输出的是中文结构化知识笔记,包含这些部分: