feat: add freshrss openclaw pipeline and clean repo
This commit is contained in:
@@ -0,0 +1,163 @@
|
||||
# OpenClaw Handoff
|
||||
|
||||
## Purpose
|
||||
|
||||
This repository provides a FreshRSS-first reading pipeline for OpenClaw:
|
||||
|
||||
`FreshRSS unread items -> RSS content extraction -> LLM summary -> rule engine -> OpenClaw delivery payload`
|
||||
|
||||
OpenClaw should treat this repository as an MCP-backed upstream content processor.
|
||||
|
||||
## Production Entrypoint
|
||||
|
||||
OpenClaw should call the MCP tool:
|
||||
|
||||
- `run_freshrss_openclaw_pipeline`
|
||||
|
||||
This is the canonical entrypoint for production use.
|
||||
|
||||
## Required Environment Variables
|
||||
|
||||
The MCP server process must have these variables available:
|
||||
|
||||
- `FRESHRSS_API_BASE_URL`
|
||||
- `FRESHRSS_USERNAME`
|
||||
- `FRESHRSS_API_PASSWORD`
|
||||
- `LLM_API_URL`
|
||||
- `LLM_API_KEY`
|
||||
- `LLM_MODEL`
|
||||
|
||||
Example:
|
||||
|
||||
```powershell
|
||||
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||||
set FRESHRSS_USERNAME=osiman
|
||||
set FRESHRSS_API_PASSWORD=your-api-password
|
||||
set LLM_API_URL=https://api.deepseek.com
|
||||
set LLM_API_KEY=your-llm-api-key
|
||||
set LLM_MODEL=deepseek-chat
|
||||
```
|
||||
|
||||
## Server Startup
|
||||
|
||||
Install dependencies:
|
||||
|
||||
```bash
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
Start the MCP server:
|
||||
|
||||
```bash
|
||||
summary-mcp
|
||||
```
|
||||
|
||||
## Recommended MCP Call
|
||||
|
||||
Recommended default call:
|
||||
|
||||
```json
|
||||
{
|
||||
"limit": 5,
|
||||
"mark_read": false,
|
||||
"include_read": false,
|
||||
"debug_artifacts": false,
|
||||
"timeout_seconds": 60,
|
||||
"max_retries": 2
|
||||
}
|
||||
```
|
||||
|
||||
Recommended semantics:
|
||||
|
||||
- Use `mark_read=false` while validating integration.
|
||||
- Use `mark_read=true` only after confirming OpenClaw will consume the returned payload successfully.
|
||||
- Keep `debug_artifacts=false` for routine production runs.
|
||||
- Set `debug_artifacts=true` only when troubleshooting a bad batch.
|
||||
|
||||
## What The Tool Returns
|
||||
|
||||
Primary return fields:
|
||||
|
||||
- `run_id`
|
||||
- `output_dir`
|
||||
- `raw_output`
|
||||
- `delivery_output`
|
||||
- `report_output`
|
||||
- `pulled_count`
|
||||
- `delivered_count`
|
||||
- `marked_read_count`
|
||||
- `status_counts`
|
||||
- `delivery_payload`
|
||||
|
||||
Optional:
|
||||
|
||||
- `items`
|
||||
- Returned only when `include_item_reports=true`
|
||||
|
||||
## Minimal Output Files
|
||||
|
||||
By default the pipeline writes only:
|
||||
|
||||
- `outputs/freshrss/rerun/<run_id>/raw/freshrss.raw.json`
|
||||
- `outputs/freshrss/rerun/<run_id>/candidates/openclaw-delivery-payload.json`
|
||||
- `outputs/freshrss/rerun/<run_id>/run-report.json`
|
||||
|
||||
If `debug_artifacts=true`, the pipeline additionally writes per-item intermediate files.
|
||||
|
||||
## Payload Specs
|
||||
|
||||
OpenClaw payload field specs live here:
|
||||
|
||||
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||
- `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||
|
||||
## Read-State Semantics
|
||||
|
||||
The pipeline reads from FreshRSS unread items by default.
|
||||
|
||||
If `mark_read=true`:
|
||||
|
||||
- items are marked as read only after the final `openclaw-delivery-payload.json` has been written successfully
|
||||
- only successfully delivered items are marked as read
|
||||
- failed or skipped items remain unread
|
||||
|
||||
## FreshRSS Content Policy
|
||||
|
||||
For FreshRSS items, the pipeline is RSS-first and does not re-crawl webpages.
|
||||
|
||||
Behavior:
|
||||
|
||||
- use `item.raw_content` first
|
||||
- if missing, use `item.raw_summary`
|
||||
- if neither contains usable content, skip the item
|
||||
- do not fetch the original webpage again for FreshRSS items
|
||||
|
||||
This is intentional.
|
||||
|
||||
## Known Limitations
|
||||
|
||||
- Some sources put only partial content in RSS; those items may be skipped if RSS content is insufficient.
|
||||
- WeChat articles often block direct crawling, but this pipeline now avoids that path for FreshRSS items and uses RSS-provided content when available.
|
||||
- Rule behavior is still conservative in some cases; many items may land in `review` depending on current rules.
|
||||
- Paywall heuristics may produce false positives for some Chinese text patterns.
|
||||
|
||||
## Files OpenClaw Should Read First
|
||||
|
||||
Recommended reading order for a new maintainer:
|
||||
|
||||
1. `README.md`
|
||||
2. `docs/openclaw/openclaw-handoff.md`
|
||||
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||
5. `docs/current/context-reset-brief.md`
|
||||
|
||||
## Current Recommendation
|
||||
|
||||
For integration handoff, the repository is usable now.
|
||||
|
||||
The minimum you need to give OpenClaw is:
|
||||
|
||||
- the repository code
|
||||
- the MCP server startup command
|
||||
- the required environment variables in the target environment
|
||||
- the instruction to call `run_freshrss_openclaw_pipeline`
|
||||
Reference in New Issue
Block a user