Files
reader/docs/openclaw/openclaw-handoff.md
T

212 lines
5.9 KiB
Markdown

# OpenClaw Handoff
## Purpose
This repository provides a FreshRSS-first reading pipeline for OpenClaw:
`FreshRSS unread items -> RSS content extraction -> LLM summary -> rule engine -> OpenClaw delivery payload`
OpenClaw should treat this repository as an MCP-backed upstream content processor.
## Production Entrypoint
OpenClaw should call the MCP tool:
- `run_freshrss_openclaw_pipeline`
This is the canonical entrypoint for production use.
## Required Environment Variables
The MCP server process must have these variables available:
- `FRESHRSS_API_BASE_URL`
- `FRESHRSS_USERNAME`
- `FRESHRSS_API_PASSWORD`
- `LLM_API_URL`
- `LLM_API_KEY`
- `LLM_MODEL`
Example:
```powershell
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=osiman
set FRESHRSS_API_PASSWORD=your-api-password
set LLM_API_URL=https://api.deepseek.com
set LLM_API_KEY=your-llm-api-key
set LLM_MODEL=deepseek-chat
```
## Server Startup
Install dependencies:
```bash
pip install -e .
```
Start the MCP server:
```bash
summary-mcp
```
## Recommended MCP Call
Recommended production call:
```json
{
"limit": 5,
"mark_read": true,
"include_read": false,
"debug_artifacts": false,
"timeout_seconds": 60,
"max_retries": 2
}
```
Recommended semantics:
- Use `mark_read=true` for normal production runs.
- Use `mark_read=false` only for debug, test, or validation runs.
- Keep `debug_artifacts=false` for routine production runs.
- Set `debug_artifacts=true` only when troubleshooting a bad batch.
## What The Tool Returns
Primary return fields:
- `run_id`
- `output_dir`
- `raw_output`
- `delivery_output`
- `report_output`
- `pulled_count`
- `delivered_count`
- `marked_read_count`
- `status_counts`
- `delivery_payload`
- `keyword_index`
Optional:
- `items`
- Returned only when `include_item_reports=true`
## Minimal Output Files
By default the pipeline writes only:
- `outputs/freshrss/rerun/<run_id>/raw/freshrss.raw.json`
- `outputs/freshrss/rerun/<run_id>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_id>/run-report.json`
- `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json` (one per item)
It also updates local runtime keyword data:
- `data/term_index/daily/YYYY-MM-DD.json`
- `data/term_index/term_stats.json`
Per-item extracted files live under `extracted/` and are always written.
If `debug_artifacts=true`, the pipeline additionally writes normalized items, summaries, filter decisions, candidate records, and candidate inputs.
The main pipeline does not emit a batch-level `freshrss.extracted.json` file by default.
## Payload Specs
OpenClaw payload field specs live here:
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
- `docs/openclaw/openclaw-delivery-payload-spec.md`
## Read-State Semantics
The pipeline reads from FreshRSS unread items by default.
If `mark_read=true`:
- items are marked as read only after the final `openclaw-delivery-payload.json` has been written successfully
- only successfully delivered items are marked as read
- failed or skipped items remain unread
## FreshRSS Content Policy
For FreshRSS items, the pipeline is RSS-first and does not re-crawl webpages.
Behavior:
- use `item.raw_content` first
- if missing, use `item.raw_summary`
- if neither contains usable content, skip the item
- do not fetch the original webpage again for FreshRSS items
This is intentional.
## Keyword Cleanup Governance
This repository also includes a lightweight keyword-governance flow for downstream review.
Current pieces:
- runtime keyword stats
- `data/term_index/daily/YYYY-MM-DD.json`
- `data/term_index/term_stats.json`
- governance config
- `configs/term_cleanup_policy.json`
- `configs/term_watchlist.json`
- `configs/term_change_log.json`
- review bundle builder
- `skills/keyword-cleanup-review/scripts/build_review_bundle.py`
- accepted-suggestion writer
- `scripts/apply_term_suggestions.py`
Current status:
- OpenClaw can read the keyword review bundle as maintenance input
- accepted suggestions still require explicit human confirmation
- the repository can write accepted watch / alias / stopword / interest-keyword changes after confirmation
- this governance flow is not yet wired into a periodic scheduler inside the repository
Boundary:
- keyword cleanup is a maintenance flow, not the production RSS ingestion path
- the repository does not auto-apply cleanup suggestions without confirmation
- current keyword stats are built from the delivered candidate payload, not yet from a final `DailyDigest`
## Known Limitations
- Some sources put only partial content in RSS; those items may be skipped if RSS content is insufficient.
- WeChat articles often block direct crawling, but this pipeline now avoids that path for FreshRSS items and uses RSS-provided content when available.
- Rule behavior is still conservative in some cases; many items may land in `review` depending on current rules.
- Paywall heuristics may produce false positives for some Chinese text patterns.
- Keyword cleanup governance is usable now, but periodic scheduling and before/after evaluation are not implemented yet.
## Files OpenClaw Should Read First
Recommended reading order for a new maintainer:
1. `README.md`
2. `docs/openclaw/openclaw-handoff.md`
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
5. `docs/design/daily-keyword-index-design.md`
6. `skills/keyword-cleanup-review/SKILL.md`
7. `docs/current/context-reset-brief.md`
## Current Recommendation
For integration handoff, the repository is usable now.
The minimum you need to give OpenClaw is:
- the repository code
- the MCP server startup command
- the required environment variables in the target environment
- the instruction to call `run_freshrss_openclaw_pipeline`
If OpenClaw will also participate in keyword-governance review, additionally point it to:
- `docs/design/daily-keyword-index-design.md`
- `skills/keyword-cleanup-review/SKILL.md`
- `scripts/apply_term_suggestions.py`