docs: formalize reader MCP workflow service handoff

This commit is contained in:
root
2026-04-07 15:32:46 +08:00
parent c622bc6247
commit 4f219ef92f
3 changed files with 158 additions and 25 deletions
+96 -11
View File
@@ -16,6 +16,7 @@ This repository is responsible only for upstream reading-pipeline work:
- rule-based filtering
- OpenClaw delivery payload generation
- selected-article summary capability based on existing extracted text
- run-state persistence, status lookup, result lookup, and minimal resume for the FreshRSS workflow
This repository should **not** take over downstream orchestration responsibilities that belong to OpenClaw / skills, such as:
@@ -26,11 +27,34 @@ This repository should **not** take over downstream orchestration responsibiliti
## Production Entrypoint
OpenClaw should call the MCP tool:
reader 当前正式工作流服务入口是 MCP tool:
- `run_freshrss_openclaw_pipeline`
This is the canonical entrypoint for production use.
It starts the only formally supported workflow today: `freshrss_daily_digest`.
OpenClaw should treat the returned `run_id` as the only stable handle for follow-up reads. Do not hand-build `outputs/freshrss/rerun/...` paths in OpenClaw.
## Supported MCP Tools
Current MCP tools: 11 total.
Workflow service tools:
- `run_freshrss_openclaw_pipeline`
- `get_run_status`
- `list_runs`
- `list_run_artifacts`
- `get_delivery_payload`
- `get_run_report`
- `resume_run`
Single-step / debug tools:
- `extract_url_content`
- `extract_item_content`
- `filter_summary_result`
- `generate_article_summaries`
## Required Environment Variables
@@ -68,9 +92,28 @@ Start the MCP server:
summary-mcp
```
## Recommended MCP Workflow
Recommended production path:
1. Call `run_freshrss_openclaw_pipeline` and persist the returned `run_id`
2. Use `get_run_status(run_id)` as the authoritative run-state read for status, stage, artifacts, and recovery
3. Use `list_runs(...)` when OpenClaw needs recent-run discovery or high-level inspection
4. Use `list_run_artifacts(run_id)` when OpenClaw needs to inspect what this run actually produced
5. Use `get_delivery_payload(run_id)` and `get_run_report(run_id)` as the formal result-reading APIs
6. Use `resume_run(run_id)` only when the run falls inside the minimal supported resume scope
OpenClaw should not directly derive or hardcode:
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
If filesystem access is needed for debugging, consume only paths returned by MCP such as `output_dir`, `artifact.path`, `delivery_output`, or `report_output`.
## Recommended MCP Call
Recommended production call:
Recommended production start call:
```json
{
@@ -89,17 +132,36 @@ Recommended semantics:
- Use `mark_read=false` only for debug, test, or validation runs.
- Keep `debug_artifacts=false` for routine production runs.
- Set `debug_artifacts=true` only when troubleshooting a bad batch.
- Treat the returned `run_id` as the stable identifier for all follow-up MCP reads.
- If no real `openclaw-delivery-payload.json` was produced, OpenClaw should stop instead of generating a digest from placeholders or examples.
## Formal Capability Boundary
reader 当前正式 MCP workflow service 的边界如下:
- formal workflow: only `freshrss_daily_digest`
- run truth: every FreshRSS run writes `run-state.json`
- state query tools: `get_run_status`, `list_runs`, `list_run_artifacts`
- result read tools: `get_delivery_payload`, `get_run_report`
- `digest-brief.json` is generated and registered as an artifact, but there is no standalone `get_digest_brief` tool yet
- `run_freshrss_openclaw_pipeline` and `resume_run` are synchronous MCP calls today; there is no background queue / worker model yet
- `generate_article_summaries` is supported, but it is outside the formal `resume_run` scope and not part of the FreshRSS workflow-state model
Historical compatibility note:
- `get_run_status` / `list_runs` / `list_run_artifacts` can still infer basic state for older run directories without `run-state.json`
- `resume_run` does **not** support those inferred historical runs; it requires a valid `run-state.json`
## What The Tool Returns
Primary return fields:
Primary return fields from `run_freshrss_openclaw_pipeline`:
- `run_id`
- `output_dir`
- `raw_output`
- `delivery_output`
- `report_output`
- `digest_brief_output`
- `pulled_count`
- `delivered_count`
- `marked_read_count`
@@ -110,16 +172,20 @@ Primary return fields:
Optional:
- `items`
- Returned only when `include_item_reports=true`
- returned only when `include_item_reports=true`
Follow-up structured reads should use MCP tools rather than re-reading these files directly.
## Minimal Output Files
By default the pipeline writes only:
By default the pipeline writes these core artifacts:
- `outputs/freshrss/rerun/<run_id>/raw/freshrss.raw.json`
- `outputs/freshrss/rerun/<run_id>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_id>/run-report.json`
- `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json` (one per item)
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
- `outputs/freshrss/rerun/<run_dir>/raw/freshrss.raw.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/digest-brief.json`
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
- `outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json` (one per item)
It also updates local runtime keyword data:
@@ -130,6 +196,24 @@ Per-item extracted files live under `extracted/` and are always written.
If `debug_artifacts=true`, the pipeline additionally writes normalized items, summaries, filter decisions, candidate records, and candidate inputs.
The main pipeline does not emit a batch-level `freshrss.extracted.json` file by default.
## `resume_run` Minimal Scope
`resume_run` currently supports only the minimum resume contract:
- only runs with a valid `run-state.json`
- only workflow `freshrss_daily_digest`
- resume in place on the original `run_id`
- supported resume points: `generate_summaries`, `apply_filters`, `build_delivery_payload`, `write_run_report`
- unsupported resume points: `fetch_feed`, `extract_articles`
- if required artifacts are missing, the tool returns a non-resumable response instead of silently falling back to an earlier stage
Artifact expectations by resume point:
- `generate_summaries`: requires `raw/freshrss.raw.json` and `extracted/`
- `apply_filters`: requires the above plus per-item summary outputs
- `build_delivery_payload`: requires per-item candidate inputs consistent with filter-stage output
- `write_run_report`: requires `candidates/openclaw-delivery-payload.json`; if `mark_read=true`, raw input must still be present
## Payload Specs
OpenClaw payload field specs live here:
@@ -230,9 +314,10 @@ The minimum you need to give OpenClaw is:
- the MCP server startup command
- the required environment variables in the target environment
- the instruction to call `run_freshrss_openclaw_pipeline`
- the rule that follow-up state/result reads must go through MCP tools first, not handwritten filesystem paths
If OpenClaw will also participate in keyword-governance review, additionally point it to:
- `docs/design/daily-keyword-index-design.md`
- `skills/keyword-cleanup-review/SKILL.md`
- `scripts/apply_term_suggestions.py`
- `scripts/apply_term_suggestions.py`