# OpenClaw Handoff ## Purpose This repository provides a FreshRSS-first reading pipeline for OpenClaw: `FreshRSS unread items -> RSS content extraction -> LLM summary -> rule engine -> OpenClaw delivery payload` OpenClaw should treat this repository as an MCP-backed upstream content processor. ## Production Entrypoint OpenClaw should call the MCP tool: - `run_freshrss_openclaw_pipeline` This is the canonical entrypoint for production use. ## Required Environment Variables The MCP server process must have these variables available: - `FRESHRSS_API_BASE_URL` - `FRESHRSS_USERNAME` - `FRESHRSS_API_PASSWORD` - `LLM_API_URL` - `LLM_API_KEY` - `LLM_MODEL` Example: ```powershell set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php set FRESHRSS_USERNAME=osiman set FRESHRSS_API_PASSWORD=your-api-password set LLM_API_URL=https://api.deepseek.com set LLM_API_KEY=your-llm-api-key set LLM_MODEL=deepseek-chat ``` ## Server Startup Install dependencies: ```bash pip install -e . ``` Start the MCP server: ```bash summary-mcp ``` ## Recommended MCP Call Recommended production call: ```json { "limit": 5, "mark_read": true, "include_read": false, "debug_artifacts": false, "timeout_seconds": 60, "max_retries": 2 } ``` Recommended semantics: - Use `mark_read=true` for normal production runs. - Use `mark_read=false` only for debug, test, or validation runs. - Keep `debug_artifacts=false` for routine production runs. - Set `debug_artifacts=true` only when troubleshooting a bad batch. ## What The Tool Returns Primary return fields: - `run_id` - `output_dir` - `raw_output` - `delivery_output` - `report_output` - `pulled_count` - `delivered_count` - `marked_read_count` - `status_counts` - `delivery_payload` - `keyword_index` Optional: - `items` - Returned only when `include_item_reports=true` ## Minimal Output Files By default the pipeline writes only: - `outputs/freshrss/rerun//raw/freshrss.raw.json` - `outputs/freshrss/rerun//candidates/openclaw-delivery-payload.json` - `outputs/freshrss/rerun//run-report.json` - `outputs/freshrss/rerun//extracted/item-XX.extracted.json` (one per item) It also updates local runtime keyword data: - `data/term_index/daily/YYYY-MM-DD.json` - `data/term_index/term_stats.json` Per-item extracted files live under `extracted/` and are always written. If `debug_artifacts=true`, the pipeline additionally writes normalized items, summaries, filter decisions, candidate records, and candidate inputs. The main pipeline does not emit a batch-level `freshrss.extracted.json` file by default. ## Payload Specs OpenClaw payload field specs live here: - `docs/openclaw/openclaw-candidate-input-field-spec.md` - `docs/openclaw/openclaw-delivery-payload-spec.md` ## Read-State Semantics The pipeline reads from FreshRSS unread items by default. If `mark_read=true`: - items are marked as read only after the final `openclaw-delivery-payload.json` has been written successfully - only successfully delivered items are marked as read - failed or skipped items remain unread ## FreshRSS Content Policy For FreshRSS items, the pipeline is RSS-first and does not re-crawl webpages. Behavior: - use `item.raw_content` first - if missing, use `item.raw_summary` - if neither contains usable content, skip the item - do not fetch the original webpage again for FreshRSS items This is intentional. ## Keyword Cleanup Governance This repository also includes a lightweight keyword-governance flow for downstream review. Current pieces: - runtime keyword stats - `data/term_index/daily/YYYY-MM-DD.json` - `data/term_index/term_stats.json` - governance config - `configs/term_cleanup_policy.json` - `configs/term_watchlist.json` - `configs/term_change_log.json` - review bundle builder - `skills/keyword-cleanup-review/scripts/build_review_bundle.py` - accepted-suggestion writer - `scripts/apply_term_suggestions.py` Current status: - OpenClaw can read the keyword review bundle as maintenance input - accepted suggestions still require explicit human confirmation - the repository can write accepted watch / alias / stopword / interest-keyword changes after confirmation - this governance flow is not yet wired into a periodic scheduler inside the repository Boundary: - keyword cleanup is a maintenance flow, not the production RSS ingestion path - the repository does not auto-apply cleanup suggestions without confirmation - current keyword stats are built from the delivered candidate payload, not yet from a final `DailyDigest` ## Known Limitations - Some sources put only partial content in RSS; those items may be skipped if RSS content is insufficient. - WeChat articles often block direct crawling, but this pipeline now avoids that path for FreshRSS items and uses RSS-provided content when available. - Rule behavior is still conservative in some cases; many items may land in `review` depending on current rules. - Paywall heuristics may produce false positives for some Chinese text patterns. - Keyword cleanup governance is usable now, but periodic scheduling and before/after evaluation are not implemented yet. ## Files OpenClaw Should Read First Recommended reading order for a new maintainer: 1. `README.md` 2. `docs/openclaw/openclaw-handoff.md` 3. `docs/openclaw/openclaw-candidate-input-field-spec.md` 4. `docs/openclaw/openclaw-delivery-payload-spec.md` 5. `docs/design/daily-keyword-index-design.md` 6. `skills/keyword-cleanup-review/SKILL.md` 7. `docs/current/context-reset-brief.md` ## Current Recommendation For integration handoff, the repository is usable now. The minimum you need to give OpenClaw is: - the repository code - the MCP server startup command - the required environment variables in the target environment - the instruction to call `run_freshrss_openclaw_pipeline` If OpenClaw will also participate in keyword-governance review, additionally point it to: - `docs/design/daily-keyword-index-design.md` - `skills/keyword-cleanup-review/SKILL.md` - `scripts/apply_term_suggestions.py`