391 lines
16 KiB
Markdown
391 lines
16 KiB
Markdown
# Reader MCP Workflow Service
|
||
|
||
reader 当前已经收口为面向 OpenClaw 的 MCP workflow service。正式能力边界以 FreshRSS 日报工作流为准:启动 run、写入 `run-state.json`、查询运行状态、读取结构化结果,以及异步恢复 job。CLI 与同步入口仍保留,但定位为 debug / fallback,而不是正式集成入口。
|
||
|
||
## 运行
|
||
|
||
```bash
|
||
pip install -e .
|
||
summary-mcp
|
||
```
|
||
|
||
服务当前暴露 21 个工具。
|
||
|
||
正式集成摘要:
|
||
|
||
- 主日报正式入口:`start_freshrss_pipeline_job`
|
||
- 主日报正式读取:`get_run_status`、`get_delivery_payload`、`get_run_report`
|
||
- 恢复正式入口:`inspect_resume_plan`、`start_resume_job`、`get_resume_job_status`、`get_resume_job_result`
|
||
- 单篇总结正式入口:`start_article_summary_job`、`get_article_summary_job_status`、`get_article_summary_job_result`
|
||
- `run_freshrss_openclaw_pipeline`、`resume_run`、`generate_article_summaries` 仅用于同步 debug / fallback
|
||
|
||
## 文档入口
|
||
|
||
如果你在做 OpenClaw 集成,不要只看这个 README,优先看:
|
||
|
||
- `docs/openclaw/README.md`
|
||
- `docs/openclaw/openclaw-handoff.md`
|
||
- `docs/openclaw/openclaw-orchestration-flow.md`
|
||
|
||
字段契约见:
|
||
|
||
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||
- `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||
|
||
文档总索引见:
|
||
|
||
- `docs/README.md`
|
||
- `docs/current/context-reset-brief.md`
|
||
- `docs/design/README.md`
|
||
- `plans/README.md`
|
||
|
||
## 正式能力边界
|
||
|
||
- 当前正式 workflow 只有 `freshrss_daily_digest`
|
||
- 当前生产编排默认走异步 job,而不是同步 MCP / CLI
|
||
- 每次 FreshRSS 主流水线 run 都会在 `outputs/freshrss/rerun/<run_dir>/run-state.json` 落地运行真相
|
||
- OpenClaw 正式读取结果应优先使用 MCP 返回的 `run_id`、`output_dir`、`delivery_output`、`report_output`
|
||
- 正式恢复只支持带有效 `run-state.json` 的当前 run,不处理历史推断 run
|
||
- 正式生产恢复依赖 `summary/summary-batch.json` 与 `candidates/candidate-batch.json`
|
||
|
||
## OpenClaw 最短调用路径
|
||
|
||
1. 调 `start_freshrss_pipeline_job`
|
||
2. 轮询 `get_freshrss_pipeline_job_status`
|
||
3. 成功后读 `get_freshrss_pipeline_job_result`,拿 `run_id`
|
||
4. 用 `get_run_status`、`get_delivery_payload`、`get_run_report` 做后续读取
|
||
5. 如需恢复,先调 `inspect_resume_plan`,只有 `recommended_action=resume` 才走 `start_resume_job`
|
||
|
||
更完整的状态分支、恢复策略和人工介入条件见 `docs/openclaw/openclaw-orchestration-flow.md`。
|
||
|
||
## 单篇文章总结后处理(可选使用独立 LLM)
|
||
|
||
### 生产环境推荐输入
|
||
|
||
- 单篇总结的正式生产输入,优先使用 FreshRSS 主流水线输出的**单篇 extracted 文件**:
|
||
- `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json`
|
||
- 这些**逐条 extracted 文件**是下游单篇总结的**正式默认产物**。
|
||
- 像 `outputs/freshrss/extracted/freshrss.extracted.json` 这样的**批量 extracted 文件**,只作为临时场景、兼容旧流程的输入形态保留,**不是首选生产默认**。
|
||
|
||
### daily 知识库默认配置
|
||
|
||
- `IMA_DAILY_KNOWLEDGE_BASE_ID` —— 单篇日报总结默认上传的 IMA 知识库 ID
|
||
- `IMA_DAILY_KNOWLEDGE_BASE_NAME` —— 默认知识库名称(预期值:`daily`)
|
||
- 上传逻辑在运行时应先校验目标知识库;若配置的目标不存在,应先按名称查找,仍不存在则创建 `daily`
|
||
|
||
相关能力:
|
||
|
||
- 正式路径:`start_article_summary_job` / `get_article_summary_job_status` / `get_article_summary_job_result`
|
||
- 同步 debug:`generate_article_summaries`
|
||
- CLI:`scripts/run_article_summaries.py`
|
||
- 后台 runner:`scripts/run_article_summary_job.py`
|
||
|
||
## 校验 LLM 摘要结果
|
||
|
||
```bash
|
||
validate-llm-result outputs/reference/summary/result.json --extracted outputs/reference/extracted/read-flow-2026.extracted.json
|
||
```
|
||
|
||
## 跑最小 extraction → summary 循环
|
||
|
||
```bash
|
||
python scripts/run_summary_loop.py ^
|
||
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||
--prompt outputs/prompts/llm-summary-prompt.txt ^
|
||
--output outputs/reference/summary/result.loop.json
|
||
```
|
||
|
||
## 拉取 FreshRSS 条目并映射为标准化 `item`
|
||
|
||
```bash
|
||
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||
set FRESHRSS_USERNAME=bot
|
||
set FRESHRSS_API_PASSWORD=your-api-password
|
||
python scripts/pull_freshrss_items.py --limit 5 --mark-read
|
||
```
|
||
|
||
默认会排除已经带 `read` 标签的条目。
|
||
如果你想拿到完整阅读列表,可以加 `--include-read`。
|
||
启用 `--mark-read` 后,脚本会在执行成功后把本次抓到的条目标记为已读。
|
||
|
||
脚本会写出:
|
||
|
||
- `outputs/freshrss/raw/freshrss.raw.json`
|
||
- `outputs/freshrss/items/freshrss.items.json`
|
||
|
||
## 拉取 FreshRSS 条目并逐条做内容提取
|
||
|
||
```bash
|
||
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||
set FRESHRSS_USERNAME=osiman
|
||
set FRESHRSS_API_PASSWORD=your-api-password
|
||
python scripts/run_freshrss_extract.py --limit 1 --mark-read
|
||
```
|
||
|
||
默认会排除已经带 `read` 标签的条目。
|
||
启用 `--mark-read` 后,只有提取成功的条目才会被标记为已读。
|
||
|
||
脚本会写出:
|
||
|
||
- `outputs/freshrss/raw/freshrss.raw.json`
|
||
- `outputs/freshrss/items/freshrss.items.json`
|
||
- `outputs/freshrss/extracted/freshrss.extracted.json`
|
||
|
||
注意:这个**批量 extracted 文件**主要用于独立提取场景和旧流程兼容。下游单篇总结的正式生产默认输入,仍然是 `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json` 这类**逐条 extracted 文件**。
|
||
|
||
## 跑完整 FreshRSS 流水线,并在最终 delivery payload 写盘成功后再标记已读
|
||
|
||
```bash
|
||
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||
set FRESHRSS_USERNAME=osiman
|
||
set FRESHRSS_API_PASSWORD=your-api-password
|
||
set LLM_API_URL=https://api.deepseek.com
|
||
set LLM_API_KEY=your-llm-api-key
|
||
set LLM_MODEL=deepseek-chat
|
||
python scripts/run_freshrss_pipeline.py --limit 5 --mark-read
|
||
```
|
||
|
||
如果你希望过滤时引入个人工程兴趣 / AI Agent 兴趣画像,可以传入 context 文件:
|
||
|
||
```bash
|
||
python scripts/run_freshrss_pipeline.py ^
|
||
--limit 5 ^
|
||
--context configs/filter_context.personal.json ^
|
||
--mark-read
|
||
```
|
||
|
||
这条 CLI 与 MCP `run_freshrss_openclaw_pipeline` / `start_freshrss_pipeline_job` 共用同一条主流水线逻辑,但正式生产集成应优先走 async MCP job;CLI 与同步 MCP 入口仅用于本地 debug / fallback。默认会写出这些产物:
|
||
|
||
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
|
||
- `outputs/freshrss/rerun/<run_dir>/raw/freshrss.raw.json`
|
||
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
|
||
- `outputs/freshrss/rerun/<run_dir>/candidates/digest-brief.json`(给 OpenClaw 生成 public digest 用的轻量输入,仅包含 `keep` 候选)
|
||
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
|
||
- `outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json`(每篇一份)
|
||
|
||
同时还会更新每日关键词索引运行数据:
|
||
|
||
- `data/term_index/daily/YYYY-MM-DD.json`
|
||
- `data/term_index/term_stats.json`
|
||
|
||
主流水线默认**不会**产出批量级的 `freshrss.extracted.json`。
|
||
如果你需要更多逐条中间产物,例如标准化 items、摘要结果、过滤决策、candidate record、candidate input,可以加 `--debug-artifacts`。
|
||
|
||
当 OpenClaw 接入这个 MCP 服务后,应先通过 `start_freshrss_pipeline_job` 启动任务,轮询 `get_freshrss_pipeline_job_status`,再从 `get_freshrss_pipeline_job_result` 读取稳定的 `run_id`。拿到 `run_id` 之后,再通过 `get_run_status` / `get_delivery_payload` / `get_run_report` 读取状态与结果,而不是直接拼接目录路径。`run_freshrss_openclaw_pipeline` 仅保留为同步 debug / fallback 路径。
|
||
在排查复杂问题时,也可以把 `debug_artifacts=true` 打开,并结合 `list_run_artifacts` 查看该 run 下实际产物。
|
||
|
||
## 对结构化摘要结果执行确定性过滤规则
|
||
|
||
```bash
|
||
python scripts/run_filter_rules.py ^
|
||
--summary outputs/reference/summary/result.loop.json ^
|
||
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||
--output outputs/reference/filter/filter-decision.json
|
||
```
|
||
|
||
如果你希望注入兴趣主题或来源标签,也可以额外传入 context 文件:
|
||
|
||
```bash
|
||
python scripts/run_filter_rules.py ^
|
||
--summary outputs/reference/summary/result.loop.json ^
|
||
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||
--context outputs/reference/filter/filter-context.json ^
|
||
--output outputs/reference/filter/filter-decision.with-context.json
|
||
```
|
||
|
||
规则引擎设计和规则编写说明见:
|
||
|
||
- `docs/design/filter-rule-engine-design.md`
|
||
- `docs/design/filter-rule-engine-usage.md`
|
||
|
||
## 将过滤结果写入 Markdown sink
|
||
|
||
```bash
|
||
python scripts/run_markdown_sink.py ^
|
||
--summary outputs/reference/summary/result.loop.json ^
|
||
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||
--filter outputs/reference/filter/filter-decision.json
|
||
```
|
||
|
||
脚本会把 Markdown 笔记写到 `knowledge-base/` 下。
|
||
|
||
## 构建内部 `ArticleCandidateRecord` 与精简版 `OpenClawCandidateInput`
|
||
|
||
```bash
|
||
python scripts/run_article_candidate.py ^
|
||
--summary outputs/reference/summary/result.loop.json ^
|
||
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||
--filter outputs/reference/filter/filter-decision.json ^
|
||
--section-hint tools_and_workflows
|
||
```
|
||
|
||
默认会写出:
|
||
|
||
- `outputs/reference/candidates/article-candidate-record.json`
|
||
- `outputs/reference/candidates/openclaw-candidate-input.json`
|
||
|
||
## 构建批量 OpenClaw delivery payload
|
||
|
||
```bash
|
||
python scripts/build_openclaw_delivery.py ^
|
||
--input-dir outputs/freshrss/candidates/batch ^
|
||
--sort-by-rank ^
|
||
--date 2026-03-25
|
||
```
|
||
|
||
默认会写出:
|
||
|
||
- `outputs/reference/candidates/openclaw-delivery-payload.json`
|
||
|
||
输出目录布局说明见 `outputs/README.md`。
|
||
|
||
## 关键词索引默认配置
|
||
|
||
相关配置文件位于:
|
||
|
||
- `configs/term_aliases.json`
|
||
- `configs/term_stopwords.json`
|
||
- `configs/term_cleanup_policy.json`
|
||
- `configs/term_watchlist.json`
|
||
- `configs/term_change_log.json`
|
||
|
||
你也可以基于已有 delivery payload 重新构建关键词索引:
|
||
|
||
```bash
|
||
python scripts/build_keyword_index.py ^
|
||
--input outputs/reference/candidates/openclaw-delivery-payload.json
|
||
```
|
||
|
||
运行期关键词数据存放在:
|
||
|
||
- `data/term_index/`
|
||
|
||
关键词清理评审 skill 位于:
|
||
|
||
- `skills/keyword-cleanup-review/`
|
||
|
||
构建给关键词治理流程使用的评审数据包(review bundle,临时工作文件):
|
||
|
||
```bash
|
||
python skills/keyword-cleanup-review/scripts/build_review_bundle.py ^
|
||
--days 7 ^
|
||
--top 50 ^
|
||
--output outputs/term_index/review/keyword-cleanup-bundle.json
|
||
```
|
||
|
||
这个评审数据包(review bundle)会额外携带治理上下文:
|
||
|
||
- 来自 `configs/term_cleanup_policy.json` 的清理阈值
|
||
- 当前 watch list(`configs/term_watchlist.json`)
|
||
- 最近已应用的变更(`configs/term_change_log.json`)
|
||
|
||
接下来可以把 bundle 渲染成正式建议产物(默认走确定性规则,不把 LLM 作为默认路径):
|
||
|
||
```bash
|
||
python scripts/generate_term_cleanup_suggestions.py ^
|
||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json
|
||
```
|
||
|
||
默认只生成:
|
||
|
||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`(唯一正式建议产物,建议短期保留)
|
||
|
||
如需人工审阅展示稿,再显式加:
|
||
|
||
```bash
|
||
python scripts/generate_term_cleanup_suggestions.py ^
|
||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json ^
|
||
--emit-markdown
|
||
```
|
||
|
||
这时才会额外生成:
|
||
|
||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`(临时展示稿,可按需生成,不必作为长期资产保留)
|
||
|
||
其中本轮最小版本优先覆盖 `interest_keyword_suggestions` 和 `watch_terms` 主链路;`alias_suggestions` / `stopword_suggestions` 先保持保守。
|
||
|
||
产物保留策略建议:
|
||
|
||
- `data/term_index/daily/YYYY-MM-DD.json`、`data/term_index/term_stats.json` 作为事实层长期保留
|
||
- `configs/filter_context.personal.json`、`configs/term_watchlist.json`、`configs/term_aliases.json`、`configs/term_stopwords.json`、`configs/term_change_log.json` 作为状态层长期保留
|
||
- `term-cleanup-suggestions-YYYY-MM-DD.json` 作为正式建议产物短期保留(最近少量几份或仅保留已应用过的)
|
||
- `keyword-cleanup-bundle.json` 仅作为临时工作文件,默认只保留当前最新一份
|
||
- `term-cleanup-suggestions-YYYY-MM-DD.md` 仅作为临时展示稿,优先按需生成,不建议默认长期归档
|
||
|
||
如果你想先预览已接受建议,再决定是否写配置文件:
|
||
|
||
```bash
|
||
python scripts/apply_term_suggestions.py ^
|
||
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json ^
|
||
--accept-watch Cron Heartbeat Memory ^
|
||
--dry-run
|
||
```
|
||
|
||
去掉 `--dry-run` 后才会真正写文件。
|
||
这个脚本也支持通过 `--accept-alias`、`--accept-stopword`、`--accept-interest` 应用 alias / stopword / interest keyword 变更。
|
||
已接受的 watch 词会写入 `configs/term_watchlist.json`,每次应用动作也会被追加到 `configs/term_change_log.json`。
|
||
|
||
## 单篇总结 LLM 配置
|
||
|
||
如果你希望单篇总结后处理使用独立模型,而不影响主流水线,可以设置:
|
||
|
||
- `ARTICLE_SUMMARY_LLM_API_URL`
|
||
- `ARTICLE_SUMMARY_LLM_MODEL`
|
||
- `ARTICLE_SUMMARY_LLM_API_KEY`
|
||
|
||
如果这些变量未设置,单篇总结会回退使用主流程中的 `LLM_*` / `OPENAI_*` 配置。
|
||
如果显式指定 DeepSeek 作为单篇总结模型且请求超时或失败,当前实现会自动再用主流程默认模型配置重试一次。
|
||
|
||
示例(PowerShell 风格):
|
||
|
||
```bash
|
||
set ARTICLE_SUMMARY_LLM_API_URL=https://api.deepseek.com
|
||
set ARTICLE_SUMMARY_LLM_API_KEY=your-article-summary-key
|
||
set ARTICLE_SUMMARY_LLM_MODEL=deepseek-chat
|
||
```
|
||
|
||
然后可以这样调用 CLI。
|
||
正式生产环境建议优先使用 `outputs/freshrss/rerun/<run_id>/extracted/` 下的**逐条 extracted 文件**;下面这个**批量 extracted** 示例仅保留为兼容旧流程 / 临时场景输入:
|
||
|
||
```bash
|
||
python scripts/run_article_summaries.py ^
|
||
--extracted outputs/freshrss/extracted/freshrss.extracted.json ^
|
||
--ids 12345 67890 ^
|
||
--output-dir outputs/freshrss/single_summaries ^
|
||
--timeout 120
|
||
```
|
||
|
||
也可以通过 `summary_mcp.server` 暴露的 MCP 工具 `generate_article_summaries` 调用:
|
||
|
||
- `extracted_path`(string):单篇 extracted JSON 路径(例如 `outputs/freshrss/rerun/<run_id>/extracted/item-01.extracted.json`),或者包含 `results` 数组的 batch extracted JSON
|
||
- `selected_ids`(array of strings):必填,至少传一个 `item_id`;如果传入的 ID 在 extracted payload 中一个都匹配不到,会直接报错
|
||
- `output_dir`(optional string):Markdown 输出目录;若不传,则默认写到 extracted 文件旁边的 `single_summaries/` 目录
|
||
- `llm_api_key` / `llm_model` / `llm_api_url`(optional strings):单篇总结 LLM 的覆盖配置;不传时会按前文规则回退到 `ARTICLE_SUMMARY_*` 或主 `LLM_*`
|
||
|
||
该工具返回一个 JSON 数组,内容为生成好的 Markdown 文件路径。
|
||
|
||
OpenClaw / 正式集成建议优先走异步 job:
|
||
|
||
- 调 `start_article_summary_job` 启动任务,立即拿到 `job_id`
|
||
- 轮询 `get_article_summary_job_status(job_id)`,直到 `status` 进入 `success` 或 `failed`
|
||
- 成功后调用 `get_article_summary_job_result(job_id)` 读取 `written_paths` 与结构化结果
|
||
- 失败时优先查看 `error_summary` 与 job 目录中的 `job-report.json`
|
||
|
||
异步 job 状态目录固定落在 `outputs/freshrss/article_summary_jobs/<job_id>/`,最小会包含:
|
||
|
||
- `run-state.json`
|
||
- `input.json`
|
||
- `result.json`(成功时)
|
||
- `job-report.json`
|
||
|
||
单篇总结使用独立 prompt:`outputs/prompts/article-summary-prompt.txt`。
|
||
它与日报 prompt 完全独立,输出的是中文结构化知识笔记,包含这些部分:
|
||
|
||
- 核心结论
|
||
- 主要论点
|
||
- 关键方法 / 机制
|
||
- 重要细节
|
||
- 可复用启发
|
||
- 关键词
|
||
- 主题
|