11 KiB
OpenClaw Handoff
Purpose
This repository provides a FreshRSS-first reading pipeline for OpenClaw:
FreshRSS unread items -> RSS content extraction -> LLM summary -> rule engine -> OpenClaw delivery payload
OpenClaw should treat this repository as an MCP-backed upstream content processor.
This repository is responsible only for upstream reading-pipeline work:
- FreshRSS pull
- content extraction
- LLM summary generation/validation
- rule-based filtering
- OpenClaw delivery payload generation
- selected-article summary capability based on existing extracted text
- run-state persistence, status lookup, result lookup, and minimal resume for the FreshRSS workflow
This repository should not take over downstream orchestration responsibilities that belong to OpenClaw / skills, such as:
- Hugo publishing
- chat reporting
- user confirmation handling
- IMA upload orchestration
Production Entrypoint
reader 当前正式工作流服务入口是 MCP tool:
run_freshrss_openclaw_pipeline
It starts the only formally supported workflow today: freshrss_daily_digest.
OpenClaw should treat the returned run_id as the only stable handle for follow-up reads. Do not hand-build outputs/freshrss/rerun/... paths in OpenClaw.
Supported MCP Tools
Current MCP tools: 11 total.
Workflow service tools:
run_freshrss_openclaw_pipelineget_run_statuslist_runslist_run_artifactsget_delivery_payloadget_run_reportresume_run
Single-step / debug tools:
extract_url_contentextract_item_contentfilter_summary_resultgenerate_article_summaries
Required Environment Variables
The MCP server process must have these variables available:
FRESHRSS_API_BASE_URLFRESHRSS_USERNAMEFRESHRSS_API_PASSWORDLLM_API_URLLLM_API_KEYLLM_MODEL
Example:
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=osiman
set FRESHRSS_API_PASSWORD=your-api-password
set LLM_API_URL=https://api.deepseek.com
set LLM_API_KEY=your-llm-api-key
set LLM_MODEL=deepseek-chat
Server Startup
Install dependencies:
pip install -e .
Start the MCP server:
summary-mcp
Recommended MCP Workflow
Recommended production path:
- Call
run_freshrss_openclaw_pipelineand persist the returnedrun_id - Use
get_run_status(run_id)as the authoritative run-state read for status, stage, artifacts, and recovery - Use
list_runs(...)when OpenClaw needs recent-run discovery or high-level inspection - Use
list_run_artifacts(run_id)when OpenClaw needs to inspect what this run actually produced - Use
get_delivery_payload(run_id)andget_run_report(run_id)as the formal result-reading APIs - Use
resume_run(run_id)only when the run falls inside the minimal supported resume scope
OpenClaw should not directly derive or hardcode:
outputs/freshrss/rerun/<run_dir>/run-state.jsonoutputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.jsonoutputs/freshrss/rerun/<run_dir>/run-report.json
If filesystem access is needed for debugging, consume only paths returned by MCP such as output_dir, artifact.path, delivery_output, or report_output.
Recommended MCP Call
Recommended production start call:
{
"limit": 5,
"mark_read": true,
"include_read": false,
"debug_artifacts": false,
"timeout_seconds": 60,
"max_retries": 2
}
Recommended semantics:
- Use
mark_read=truefor normal production runs. - Use
mark_read=falseonly for debug, test, or validation runs. - Keep
debug_artifacts=falsefor routine production runs. - Set
debug_artifacts=trueonly when troubleshooting a bad batch. - Treat the returned
run_idas the stable identifier for all follow-up MCP reads. - If no real
openclaw-delivery-payload.jsonwas produced, OpenClaw should stop instead of generating a digest from placeholders or examples.
Formal Capability Boundary
reader 当前正式 MCP workflow service 的边界如下:
- formal workflow: only
freshrss_daily_digest - run truth: every FreshRSS run writes
run-state.json - state query tools:
get_run_status,list_runs,list_run_artifacts - result read tools:
get_delivery_payload,get_run_report digest-brief.jsonis generated and registered as an artifact, but there is no standaloneget_digest_brieftool yetrun_freshrss_openclaw_pipelineandresume_runare synchronous MCP calls today; there is no background queue / worker model yetgenerate_article_summariesis supported, but it is outside the formalresume_runscope and not part of the FreshRSS workflow-state model
Historical compatibility note:
get_run_status/list_runs/list_run_artifactscan still infer basic state for older run directories withoutrun-state.jsonresume_rundoes not support those inferred historical runs; it requires a validrun-state.json
What The Tool Returns
Primary return fields from run_freshrss_openclaw_pipeline:
run_idoutput_dirraw_outputdelivery_outputreport_outputdigest_brief_outputpulled_countdelivered_countmarked_read_countstatus_countsdelivery_payloadkeyword_index
Optional:
items- returned only when
include_item_reports=true
- returned only when
Follow-up structured reads should use MCP tools rather than re-reading these files directly.
Minimal Output Files
By default the pipeline writes these core artifacts:
outputs/freshrss/rerun/<run_dir>/run-state.jsonoutputs/freshrss/rerun/<run_dir>/raw/freshrss.raw.jsonoutputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.jsonoutputs/freshrss/rerun/<run_dir>/candidates/digest-brief.jsonoutputs/freshrss/rerun/<run_dir>/run-report.jsonoutputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json(one per item)
It also updates local runtime keyword data:
data/term_index/daily/YYYY-MM-DD.jsondata/term_index/term_stats.json
Per-item extracted files live under extracted/ and are always written.
If debug_artifacts=true, the pipeline additionally writes normalized items, summaries, filter decisions, candidate records, and candidate inputs.
The main pipeline does not emit a batch-level freshrss.extracted.json file by default.
resume_run Minimal Scope
resume_run currently supports only the minimum resume contract:
- only runs with a valid
run-state.json - only workflow
freshrss_daily_digest - resume in place on the original
run_id - supported resume points:
generate_summaries,apply_filters,build_delivery_payload,write_run_report - unsupported resume points:
fetch_feed,extract_articles - if required artifacts are missing, the tool returns a non-resumable response instead of silently falling back to an earlier stage
Artifact expectations by resume point:
generate_summaries: requiresraw/freshrss.raw.jsonandextracted/apply_filters: requires the above plus per-item summary outputsbuild_delivery_payload: requires per-item candidate inputs consistent with filter-stage outputwrite_run_report: requirescandidates/openclaw-delivery-payload.json; ifmark_read=true, raw input must still be present
Payload Specs
OpenClaw payload field specs live here:
docs/openclaw/openclaw-candidate-input-field-spec.mddocs/openclaw/openclaw-delivery-payload-spec.md
Read-State Semantics
The pipeline reads from FreshRSS unread items by default.
If mark_read=true:
- items are marked as read only after the final
openclaw-delivery-payload.jsonhas been written successfully - only successfully delivered items are marked as read
- failed or skipped items remain unread
FreshRSS Content Policy
For FreshRSS items, the pipeline is RSS-first and does not re-crawl webpages.
Behavior:
- use
item.raw_contentfirst - if missing, use
item.raw_summary - if neither contains usable content, skip the item
- do not fetch the original webpage again for FreshRSS items
This is intentional.
Keyword Cleanup Governance
This repository also includes a lightweight keyword-governance flow for downstream review.
Current pieces:
- runtime keyword stats
data/term_index/daily/YYYY-MM-DD.jsondata/term_index/term_stats.json
- governance config
configs/term_cleanup_policy.jsonconfigs/term_watchlist.jsonconfigs/term_change_log.json
- review bundle builder
skills/keyword-cleanup-review/scripts/build_review_bundle.py
- accepted-suggestion writer
scripts/apply_term_suggestions.py
Current status:
- OpenClaw can read the keyword review bundle as maintenance input
- accepted suggestions still require explicit human confirmation
- the repository can write accepted watch / alias / stopword / interest-keyword changes after confirmation
- this governance flow is not yet wired into a periodic scheduler inside the repository
Boundary:
- keyword cleanup is a maintenance flow, not the production RSS ingestion path
- the repository does not auto-apply cleanup suggestions without confirmation
- current keyword stats are built from the delivered candidate payload, not yet from a final
DailyDigest
Known Limitations
- Some sources put only partial content in RSS; those items may be skipped if RSS content is insufficient.
- WeChat articles often block direct crawling, but this pipeline now avoids that path for FreshRSS items and uses RSS-provided content when available.
- Rule behavior is still conservative in some cases; many items may land in
reviewdepending on current rules. - Paywall heuristics may produce false positives for some Chinese text patterns.
- Keyword cleanup governance is usable now, but periodic scheduling and before/after evaluation are not implemented yet.
Files OpenClaw Should Read First
Recommended reading order for a new maintainer:
README.mddocs/openclaw/openclaw-handoff.mddocs/openclaw/openclaw-candidate-input-field-spec.mddocs/openclaw/openclaw-delivery-payload-spec.mddocs/design/daily-keyword-index-design.mdskills/keyword-cleanup-review/SKILL.mddocs/current/context-reset-brief.md
Downstream Boundary Rules
For the daily-digest workflow:
- the digest should go to Hugo and chat reporting, not directly into IMA
- the full daily digest should not be uploaded to IMA
- only explicitly user-selected article summaries should be uploaded to IMA
- selected-article summaries should be generated from existing extracted text, not by re-fetching original URLs
Current Recommendation
For integration handoff, the repository is usable now.
The minimum you need to give OpenClaw is:
- the repository code
- the MCP server startup command
- the required environment variables in the target environment
- the instruction to call
run_freshrss_openclaw_pipeline - the rule that follow-up state/result reads must go through MCP tools first, not handwritten filesystem paths
If OpenClaw will also participate in keyword-governance review, additionally point it to:
docs/design/daily-keyword-index-design.mdskills/keyword-cleanup-review/SKILL.mdscripts/apply_term_suggestions.py