diff --git a/README.md b/README.md index 7a25030..150f09f 100644 --- a/README.md +++ b/README.md @@ -1,6 +1,6 @@ -# 内容提取 MCP +# Reader MCP Workflow Service -这是一个基于 Python 的 MCP 项目骨架,用于完成文章内容提取、结构化摘要校验、确定性过滤,以及 Markdown 输出落盘。 +reader 当前已经收口为面向 OpenClaw 的 MCP workflow service。正式能力边界以 FreshRSS 日报工作流为准:启动 run、写入 `run-state.json`、查询运行状态、读取结构化结果,以及最小可用的 `resume_run`。CLI 仍保留,但定位为 debug / fallback,而不是正式集成入口。 ## 运行 @@ -9,14 +9,50 @@ pip install -e . summary-mcp ``` -服务当前暴露 5 个工具: +服务当前暴露 11 个工具。 + +正式 workflow service 相关工具: + +- `run_freshrss_openclaw_pipeline` +- `get_run_status` +- `list_runs` +- `list_run_artifacts` +- `get_delivery_payload` +- `get_run_report` +- `resume_run` + +单步处理 / 调试相关工具: - `extract_url_content` - `extract_item_content` - `filter_summary_result` -- `run_freshrss_openclaw_pipeline` - `generate_article_summaries` +## 正式能力边界 + +- 当前正式 workflow 只有 `freshrss_daily_digest` +- 每次 FreshRSS 主流水线 run 都会在 `outputs/freshrss/rerun//run-state.json` 落地运行真相 +- OpenClaw 正式读取结果应优先使用 `get_delivery_payload` 与 `get_run_report`,而不是自己拼输出目录路径 +- `digest-brief.json` 当前会随主流水线产出,但还没有独立的 MCP 读取工具;如需定位它,应通过 `list_run_artifacts` 或 `get_run_report` 返回的信息发现 +- `run_freshrss_openclaw_pipeline` 与 `resume_run` 当前都是同步 MCP 调用;仓库里还没有后台队列 / worker / 异步任务管理 + +## OpenClaw 推荐调用路径 + +1. 调用 `run_freshrss_openclaw_pipeline` 启动正式日报 run,并保存返回的 `run_id` +2. 后续所有状态判断都基于 `get_run_status(run_id)` 或 `list_runs(...)` +3. 需要看产物列表时用 `list_run_artifacts(run_id)`,不要在 OpenClaw 里硬编码 `outputs/freshrss/rerun/...` +4. 需要消费正式结果时优先用 `get_delivery_payload(run_id)` 与 `get_run_report(run_id)` +5. 仅当 `resume_run` 的最小恢复范围满足时,才对失败 run 调用 `resume_run(run_id)`;否则应重启一个新 run + +## `resume_run` 当前最小范围 + +- 只支持带有效 `run-state.json` 的 run +- 只支持 workflow `freshrss_daily_digest` +- 恢复时继续沿用原 `run_id`,不会新建 retry run +- 当前支持的恢复起点只有:`generate_summaries`、`apply_filters`、`build_delivery_payload`、`write_run_report` +- 当前明确不支持从 `fetch_feed`、`extract_articles` 恢复;这类失败应新开 run +- 恢复前会校验关键中间产物是否齐备,缺失时直接返回不可恢复,而不会自动回退到更早 stage + ## 单篇文章总结后处理(可选使用独立 LLM) ### 生产环境推荐输入 @@ -34,7 +70,7 @@ summary-mcp 相关能力: -- `article-summary` MCP 工具(基于已有 extracted payload 做单篇总结) +- `generate_article_summaries` MCP 工具(基于已有 extracted payload 做单篇总结) - `scripts/run_article_summaries.py` CLI 辅助脚本 ## 校验 LLM 摘要结果 @@ -111,13 +147,14 @@ python scripts/run_freshrss_pipeline.py ^ --mark-read ``` -这是当前推荐的**正式生产入口**。默认只写出这些产物: +这条 CLI 与 MCP `run_freshrss_openclaw_pipeline` 共用同一条主流水线逻辑,但正式生产集成应优先走 MCP;CLI 仅用于本地 debug / fallback。默认会写出这些产物: -- `outputs/freshrss/rerun//raw/freshrss.raw.json` -- `outputs/freshrss/rerun//candidates/openclaw-delivery-payload.json` -- `outputs/freshrss/rerun//candidates/digest-brief.json`(给 OpenClaw 生成 public digest 用的轻量输入,仅包含 `keep` 候选) -- `outputs/freshrss/rerun//run-report.json` -- `outputs/freshrss/rerun//extracted/item-XX.extracted.json`(每篇一份) +- `outputs/freshrss/rerun//run-state.json` +- `outputs/freshrss/rerun//raw/freshrss.raw.json` +- `outputs/freshrss/rerun//candidates/openclaw-delivery-payload.json` +- `outputs/freshrss/rerun//candidates/digest-brief.json`(给 OpenClaw 生成 public digest 用的轻量输入,仅包含 `keep` 候选) +- `outputs/freshrss/rerun//run-report.json` +- `outputs/freshrss/rerun//extracted/item-XX.extracted.json`(每篇一份) 同时还会更新每日关键词索引运行数据: @@ -127,8 +164,8 @@ python scripts/run_freshrss_pipeline.py ^ 主流水线默认**不会**产出批量级的 `freshrss.extracted.json`。 如果你需要更多逐条中间产物,例如标准化 items、摘要结果、过滤决策、candidate record、candidate input,可以加 `--debug-artifacts`。 -当 OpenClaw 接入这个 MCP 服务后,应直接调用 `run_freshrss_openclaw_pipeline` 来获得同样行为。 -在排查复杂问题时,也可以把 `debug_artifacts=true` 打开。 +当 OpenClaw 接入这个 MCP 服务后,应以 `run_id` 作为稳定句柄:先调用 `run_freshrss_openclaw_pipeline`,再通过 `get_run_status` / `get_delivery_payload` / `get_run_report` 读取状态与结果,而不是直接拼接目录路径。 +在排查复杂问题时,也可以把 `debug_artifacts=true` 打开,并结合 `list_run_artifacts` 查看该 run 下实际产物。 ## 对结构化摘要结果执行确定性过滤规则 diff --git a/TODO.md b/TODO.md index 2c35214..bab5c9d 100644 --- a/TODO.md +++ b/TODO.md @@ -202,12 +202,23 @@ --- -### [TODO][P3] 更新 README / handoff / docs,明确 MCP 为正式入口 +### [DONE][P3] 更新 README / handoff / docs,明确 MCP 为正式入口 目标: - 把生产建议从 CLI 迁移到 MCP - CLI 明确降级为 debug / fallback +完成情况: +- 已更新 `README.md`,补齐 reader 作为正式 MCP workflow service 的当前能力边界、推荐调用路径、已支持 tools 与最小 `resume_run` 范围 +- 已更新 `docs/openclaw/openclaw-handoff.md`,明确 OpenClaw 应优先通过 MCP 读取 run 状态与结果,不再自己拼接 reader 输出路径 + +改动文件: +- `README.md` +- `docs/openclaw/openclaw-handoff.md` + +遗留风险: +- 当前仍无独立 `get_digest_brief` tool;若下游确实需要该产物,仍应先通过 `list_run_artifacts` / `get_run_report` 发现,而不是写死路径 + --- ## 3. 记录区 diff --git a/docs/openclaw/openclaw-handoff.md b/docs/openclaw/openclaw-handoff.md index f23ec39..ebe4a3c 100644 --- a/docs/openclaw/openclaw-handoff.md +++ b/docs/openclaw/openclaw-handoff.md @@ -16,6 +16,7 @@ This repository is responsible only for upstream reading-pipeline work: - rule-based filtering - OpenClaw delivery payload generation - selected-article summary capability based on existing extracted text +- run-state persistence, status lookup, result lookup, and minimal resume for the FreshRSS workflow This repository should **not** take over downstream orchestration responsibilities that belong to OpenClaw / skills, such as: @@ -26,11 +27,34 @@ This repository should **not** take over downstream orchestration responsibiliti ## Production Entrypoint -OpenClaw should call the MCP tool: +reader 当前正式工作流服务入口是 MCP tool: - `run_freshrss_openclaw_pipeline` -This is the canonical entrypoint for production use. +It starts the only formally supported workflow today: `freshrss_daily_digest`. + +OpenClaw should treat the returned `run_id` as the only stable handle for follow-up reads. Do not hand-build `outputs/freshrss/rerun/...` paths in OpenClaw. + +## Supported MCP Tools + +Current MCP tools: 11 total. + +Workflow service tools: + +- `run_freshrss_openclaw_pipeline` +- `get_run_status` +- `list_runs` +- `list_run_artifacts` +- `get_delivery_payload` +- `get_run_report` +- `resume_run` + +Single-step / debug tools: + +- `extract_url_content` +- `extract_item_content` +- `filter_summary_result` +- `generate_article_summaries` ## Required Environment Variables @@ -68,9 +92,28 @@ Start the MCP server: summary-mcp ``` +## Recommended MCP Workflow + +Recommended production path: + +1. Call `run_freshrss_openclaw_pipeline` and persist the returned `run_id` +2. Use `get_run_status(run_id)` as the authoritative run-state read for status, stage, artifacts, and recovery +3. Use `list_runs(...)` when OpenClaw needs recent-run discovery or high-level inspection +4. Use `list_run_artifacts(run_id)` when OpenClaw needs to inspect what this run actually produced +5. Use `get_delivery_payload(run_id)` and `get_run_report(run_id)` as the formal result-reading APIs +6. Use `resume_run(run_id)` only when the run falls inside the minimal supported resume scope + +OpenClaw should not directly derive or hardcode: + +- `outputs/freshrss/rerun//run-state.json` +- `outputs/freshrss/rerun//candidates/openclaw-delivery-payload.json` +- `outputs/freshrss/rerun//run-report.json` + +If filesystem access is needed for debugging, consume only paths returned by MCP such as `output_dir`, `artifact.path`, `delivery_output`, or `report_output`. + ## Recommended MCP Call -Recommended production call: +Recommended production start call: ```json { @@ -89,17 +132,36 @@ Recommended semantics: - Use `mark_read=false` only for debug, test, or validation runs. - Keep `debug_artifacts=false` for routine production runs. - Set `debug_artifacts=true` only when troubleshooting a bad batch. +- Treat the returned `run_id` as the stable identifier for all follow-up MCP reads. - If no real `openclaw-delivery-payload.json` was produced, OpenClaw should stop instead of generating a digest from placeholders or examples. +## Formal Capability Boundary + +reader 当前正式 MCP workflow service 的边界如下: + +- formal workflow: only `freshrss_daily_digest` +- run truth: every FreshRSS run writes `run-state.json` +- state query tools: `get_run_status`, `list_runs`, `list_run_artifacts` +- result read tools: `get_delivery_payload`, `get_run_report` +- `digest-brief.json` is generated and registered as an artifact, but there is no standalone `get_digest_brief` tool yet +- `run_freshrss_openclaw_pipeline` and `resume_run` are synchronous MCP calls today; there is no background queue / worker model yet +- `generate_article_summaries` is supported, but it is outside the formal `resume_run` scope and not part of the FreshRSS workflow-state model + +Historical compatibility note: + +- `get_run_status` / `list_runs` / `list_run_artifacts` can still infer basic state for older run directories without `run-state.json` +- `resume_run` does **not** support those inferred historical runs; it requires a valid `run-state.json` + ## What The Tool Returns -Primary return fields: +Primary return fields from `run_freshrss_openclaw_pipeline`: - `run_id` - `output_dir` - `raw_output` - `delivery_output` - `report_output` +- `digest_brief_output` - `pulled_count` - `delivered_count` - `marked_read_count` @@ -110,16 +172,20 @@ Primary return fields: Optional: - `items` - - Returned only when `include_item_reports=true` + - returned only when `include_item_reports=true` + +Follow-up structured reads should use MCP tools rather than re-reading these files directly. ## Minimal Output Files -By default the pipeline writes only: +By default the pipeline writes these core artifacts: -- `outputs/freshrss/rerun//raw/freshrss.raw.json` -- `outputs/freshrss/rerun//candidates/openclaw-delivery-payload.json` -- `outputs/freshrss/rerun//run-report.json` -- `outputs/freshrss/rerun//extracted/item-XX.extracted.json` (one per item) +- `outputs/freshrss/rerun//run-state.json` +- `outputs/freshrss/rerun//raw/freshrss.raw.json` +- `outputs/freshrss/rerun//candidates/openclaw-delivery-payload.json` +- `outputs/freshrss/rerun//candidates/digest-brief.json` +- `outputs/freshrss/rerun//run-report.json` +- `outputs/freshrss/rerun//extracted/item-XX.extracted.json` (one per item) It also updates local runtime keyword data: @@ -130,6 +196,24 @@ Per-item extracted files live under `extracted/` and are always written. If `debug_artifacts=true`, the pipeline additionally writes normalized items, summaries, filter decisions, candidate records, and candidate inputs. The main pipeline does not emit a batch-level `freshrss.extracted.json` file by default. +## `resume_run` Minimal Scope + +`resume_run` currently supports only the minimum resume contract: + +- only runs with a valid `run-state.json` +- only workflow `freshrss_daily_digest` +- resume in place on the original `run_id` +- supported resume points: `generate_summaries`, `apply_filters`, `build_delivery_payload`, `write_run_report` +- unsupported resume points: `fetch_feed`, `extract_articles` +- if required artifacts are missing, the tool returns a non-resumable response instead of silently falling back to an earlier stage + +Artifact expectations by resume point: + +- `generate_summaries`: requires `raw/freshrss.raw.json` and `extracted/` +- `apply_filters`: requires the above plus per-item summary outputs +- `build_delivery_payload`: requires per-item candidate inputs consistent with filter-stage output +- `write_run_report`: requires `candidates/openclaw-delivery-payload.json`; if `mark_read=true`, raw input must still be present + ## Payload Specs OpenClaw payload field specs live here: @@ -230,9 +314,10 @@ The minimum you need to give OpenClaw is: - the MCP server startup command - the required environment variables in the target environment - the instruction to call `run_freshrss_openclaw_pipeline` +- the rule that follow-up state/result reads must go through MCP tools first, not handwritten filesystem paths If OpenClaw will also participate in keyword-governance review, additionally point it to: - `docs/design/daily-keyword-index-design.md` - `skills/keyword-cleanup-review/SKILL.md` -- `scripts/apply_term_suggestions.py` \ No newline at end of file +- `scripts/apply_term_suggestions.py`