361 lines
13 KiB
Markdown
361 lines
13 KiB
Markdown
# OpenClaw Handoff
|
||
|
||
## Purpose
|
||
|
||
This repository provides a FreshRSS-first reading pipeline for OpenClaw:
|
||
|
||
`FreshRSS unread items -> RSS content extraction -> LLM summary -> rule engine -> OpenClaw delivery payload`
|
||
|
||
OpenClaw should treat this repository as an MCP-backed upstream content processor.
|
||
|
||
This repository is responsible only for upstream reading-pipeline work:
|
||
|
||
- FreshRSS pull
|
||
- content extraction
|
||
- LLM summary generation/validation
|
||
- rule-based filtering
|
||
- OpenClaw delivery payload generation
|
||
- selected-article summary capability based on existing extracted text
|
||
- run-state persistence, status lookup, result lookup, and minimal resume for the FreshRSS workflow
|
||
|
||
This repository should **not** take over downstream orchestration responsibilities that belong to OpenClaw / skills, such as:
|
||
|
||
- Hugo publishing
|
||
- chat reporting
|
||
- user confirmation handling
|
||
- IMA upload orchestration
|
||
|
||
## Production Entrypoint
|
||
|
||
reader 当前正式工作流服务启动入口是 MCP tool:
|
||
|
||
- `start_freshrss_pipeline_job`
|
||
|
||
OpenClaw 应先拿到 `job_id`,轮询 job 状态,再在成功后读取 `run_id` 作为正式后续句柄。
|
||
|
||
`run_freshrss_openclaw_pipeline` 仍保留,但定位是同步 debug / fallback 路径,而不是正式生产启动入口。
|
||
|
||
OpenClaw should treat the returned `run_id` from `get_freshrss_pipeline_job_result` as the only stable handle for follow-up reads. Do not hand-build `outputs/freshrss/rerun/...` paths in OpenClaw.
|
||
|
||
Job state is written under `outputs/freshrss/pipeline_jobs/<job_id>/` and will minimally contain `run-state.json`, `input.json`, `result.json` on success, and `job-report.json`.
|
||
|
||
## Supported MCP Tools
|
||
|
||
Current MCP tools: 17 total, including the FreshRSS workflow set plus async job tools for both the main pipeline and article-summary flow.
|
||
|
||
Workflow service tools:
|
||
|
||
- `start_freshrss_pipeline_job`
|
||
- `get_freshrss_pipeline_job_status`
|
||
- `get_freshrss_pipeline_job_result`
|
||
- `get_run_status`
|
||
- `list_runs`
|
||
- `list_run_artifacts`
|
||
- `get_delivery_payload`
|
||
- `get_run_report`
|
||
- `resume_run`
|
||
|
||
Article-summary tools:
|
||
|
||
- `start_article_summary_job`
|
||
- `get_article_summary_job_status`
|
||
- `get_article_summary_job_result`
|
||
|
||
Single-step / debug tools:
|
||
|
||
- `run_freshrss_openclaw_pipeline`(同步模式,仅适合 debug / fallback)
|
||
- `extract_url_content`
|
||
- `extract_item_content`
|
||
- `filter_summary_result`
|
||
- `generate_article_summaries`(同步模式,仅适合轻量调试)
|
||
|
||
## Recommended Selected-Article Flow
|
||
|
||
For OpenClaw selected-article follow-up, prefer the async job path:
|
||
|
||
1. Call `start_article_summary_job` with a real extracted file path plus a non-empty `selected_ids` list.
|
||
2. Poll `get_article_summary_job_status` until `status` becomes `success` or `failed`.
|
||
3. On success, call `get_article_summary_job_result` and continue downstream processing from `written_paths`.
|
||
4. Use `generate_article_summaries` only as a synchronous debug fallback, not as the default production path.
|
||
|
||
Job state is written under `outputs/freshrss/article_summary_jobs/<job_id>/` and will minimally contain:
|
||
|
||
- `run-state.json`
|
||
- `input.json`
|
||
- `result.json` on success
|
||
- `job-report.json`
|
||
|
||
## Required Environment Variables
|
||
|
||
The MCP server process must have these variables available:
|
||
|
||
- `FRESHRSS_API_BASE_URL`
|
||
- `FRESHRSS_USERNAME`
|
||
- `FRESHRSS_API_PASSWORD`
|
||
- `LLM_API_URL`
|
||
- `LLM_API_KEY`
|
||
- `LLM_MODEL`
|
||
|
||
Example:
|
||
|
||
```powershell
|
||
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||
set FRESHRSS_USERNAME=osiman
|
||
set FRESHRSS_API_PASSWORD=your-api-password
|
||
set LLM_API_URL=https://api.deepseek.com
|
||
set LLM_API_KEY=your-llm-api-key
|
||
set LLM_MODEL=deepseek-chat
|
||
```
|
||
|
||
## Server Startup
|
||
|
||
Install dependencies:
|
||
|
||
```bash
|
||
pip install -e .
|
||
```
|
||
|
||
Start the MCP server:
|
||
|
||
```bash
|
||
summary-mcp
|
||
```
|
||
|
||
## Recommended MCP Workflow
|
||
|
||
Recommended production path:
|
||
|
||
1. Call `start_freshrss_pipeline_job` and persist the returned `job_id`
|
||
2. Poll `get_freshrss_pipeline_job_status(job_id)` until `status` becomes `success` or `failed`
|
||
3. On success, call `get_freshrss_pipeline_job_result(job_id)` and persist the returned `run_id`
|
||
4. Use `get_run_status(run_id)` as the authoritative run-state read for status, stage, artifacts, and recovery
|
||
5. Use `list_runs(...)` when OpenClaw needs recent-run discovery or high-level inspection
|
||
6. Use `list_run_artifacts(run_id)` when OpenClaw needs to inspect what this run actually produced
|
||
7. Use `get_delivery_payload(run_id)` and `get_run_report(run_id)` as the formal result-reading APIs
|
||
8. Use `resume_run(run_id)` only when the run falls inside the minimal supported resume scope
|
||
|
||
OpenClaw should not directly derive or hardcode:
|
||
|
||
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
|
||
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
|
||
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
|
||
|
||
If filesystem access is needed for debugging, consume only paths returned by MCP such as `output_dir`, `artifact.path`, `delivery_output`, or `report_output`.
|
||
|
||
## Recommended MCP Call
|
||
|
||
Recommended production start call:
|
||
|
||
```json
|
||
{
|
||
"limit": 5,
|
||
"mark_read": true,
|
||
"include_read": false,
|
||
"debug_artifacts": false,
|
||
"timeout_seconds": 60,
|
||
"max_retries": 2
|
||
}
|
||
```
|
||
|
||
Recommended semantics:
|
||
|
||
- Use `mark_read=true` for normal production runs.
|
||
- Use `mark_read=false` only for debug, test, or validation runs.
|
||
- Keep `debug_artifacts=false` for routine production runs.
|
||
- Set `debug_artifacts=true` only when troubleshooting a bad batch.
|
||
- Treat the returned `job_id` as the startup handle, and the later `run_id` from `get_freshrss_pipeline_job_result` as the stable identifier for all follow-up run reads.
|
||
- If no real `openclaw-delivery-payload.json` was produced, OpenClaw should stop instead of generating a digest from placeholders or examples.
|
||
- Use `run_freshrss_openclaw_pipeline` only when a synchronous debug / fallback path is explicitly needed.
|
||
|
||
## Formal Capability Boundary
|
||
|
||
reader 当前正式 MCP workflow service 的边界如下:
|
||
|
||
- formal workflow: only `freshrss_daily_digest`
|
||
- run truth: every FreshRSS run writes `run-state.json`
|
||
- main production start path: `start_freshrss_pipeline_job` / `get_freshrss_pipeline_job_status` / `get_freshrss_pipeline_job_result`
|
||
- state query tools: `get_run_status`, `list_runs`, `list_run_artifacts`
|
||
- result read tools: `get_delivery_payload`, `get_run_report`
|
||
- `digest-brief.json` is generated and registered as an artifact, but there is no standalone `get_digest_brief` tool yet
|
||
- `run_freshrss_openclaw_pipeline` is still supported, but only as a synchronous debug / fallback path
|
||
- the FreshRSS daily workflow now has a minimal background job model backed by a detached runner process, not a full queue / worker system
|
||
- article summary now has a minimal asynchronous job model with `start_article_summary_job` / `get_article_summary_job_status` / `get_article_summary_job_result`
|
||
- `generate_article_summaries` is still supported, but it is a synchronous debug path and outside the formal `resume_run` scope
|
||
|
||
Historical compatibility note:
|
||
|
||
- `get_run_status` / `list_runs` / `list_run_artifacts` can still infer basic state for older run directories without `run-state.json`
|
||
- `resume_run` does **not** support those inferred historical runs; it requires a valid `run-state.json`
|
||
|
||
## What The Async Job Returns
|
||
|
||
Primary return fields from `get_freshrss_pipeline_job_result`:
|
||
|
||
- `job_id`
|
||
- `run_id`
|
||
- `output_dir`
|
||
- `raw_output`
|
||
- `delivery_output`
|
||
- `report_output`
|
||
- `digest_brief_output`
|
||
- `pulled_count`
|
||
- `delivered_count`
|
||
- `marked_read_count`
|
||
- `status_counts`
|
||
- `keyword_index`
|
||
|
||
Optional:
|
||
|
||
- `items`
|
||
- returned only when `include_item_reports=true`
|
||
|
||
`run_freshrss_openclaw_pipeline` still returns the same synchronous payload for debug / fallback use.
|
||
|
||
Follow-up structured reads should use MCP tools rather than re-reading these files directly.
|
||
|
||
## Minimal Output Files
|
||
|
||
By default the pipeline writes these core artifacts:
|
||
|
||
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
|
||
- `outputs/freshrss/rerun/<run_dir>/raw/freshrss.raw.json`
|
||
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
|
||
- `outputs/freshrss/rerun/<run_dir>/candidates/digest-brief.json`
|
||
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
|
||
- `outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json` (one per item)
|
||
|
||
It also updates local runtime keyword data:
|
||
|
||
- `data/term_index/daily/YYYY-MM-DD.json`
|
||
- `data/term_index/term_stats.json`
|
||
|
||
Per-item extracted files live under `extracted/` and are always written.
|
||
If `debug_artifacts=true`, the pipeline additionally writes normalized items, summaries, filter decisions, candidate records, and candidate inputs.
|
||
The main pipeline does not emit a batch-level `freshrss.extracted.json` file by default.
|
||
|
||
## `resume_run` Minimal Scope
|
||
|
||
`resume_run` currently supports only the minimum resume contract:
|
||
|
||
- only runs with a valid `run-state.json`
|
||
- only workflow `freshrss_daily_digest`
|
||
- resume in place on the original `run_id`
|
||
- supported resume points: `generate_summaries`, `apply_filters`, `build_delivery_payload`, `write_run_report`
|
||
- unsupported resume points: `fetch_feed`, `extract_articles`
|
||
- if required artifacts are missing, the tool returns a non-resumable response instead of silently falling back to an earlier stage
|
||
|
||
Artifact expectations by resume point:
|
||
|
||
- `generate_summaries`: requires `raw/freshrss.raw.json` and `extracted/`
|
||
- `apply_filters`: requires the above plus per-item summary outputs
|
||
- `build_delivery_payload`: requires per-item candidate inputs consistent with filter-stage output
|
||
- `write_run_report`: requires `candidates/openclaw-delivery-payload.json`; if `mark_read=true`, raw input must still be present
|
||
|
||
## Payload Specs
|
||
|
||
OpenClaw payload field specs live here:
|
||
|
||
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||
- `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||
|
||
## Read-State Semantics
|
||
|
||
The pipeline reads from FreshRSS unread items by default.
|
||
|
||
If `mark_read=true`:
|
||
|
||
- items are marked as read only after the final `openclaw-delivery-payload.json` has been written successfully
|
||
- only successfully delivered items are marked as read
|
||
- failed or skipped items remain unread
|
||
|
||
## FreshRSS Content Policy
|
||
|
||
For FreshRSS items, the pipeline is RSS-first and does not re-crawl webpages.
|
||
|
||
Behavior:
|
||
|
||
- use `item.raw_content` first
|
||
- if missing, use `item.raw_summary`
|
||
- if neither contains usable content, skip the item
|
||
- do not fetch the original webpage again for FreshRSS items
|
||
|
||
This is intentional.
|
||
|
||
## Keyword Cleanup Governance
|
||
|
||
This repository also includes a lightweight keyword-governance flow for downstream review.
|
||
|
||
Current pieces:
|
||
|
||
- runtime keyword stats
|
||
- `data/term_index/daily/YYYY-MM-DD.json`
|
||
- `data/term_index/term_stats.json`
|
||
- governance config
|
||
- `configs/term_cleanup_policy.json`
|
||
- `configs/term_watchlist.json`
|
||
- `configs/term_change_log.json`
|
||
- review bundle builder
|
||
- `skills/keyword-cleanup-review/scripts/build_review_bundle.py`
|
||
- accepted-suggestion writer
|
||
- `scripts/apply_term_suggestions.py`
|
||
|
||
Current status:
|
||
|
||
- OpenClaw can read the keyword review bundle as maintenance input
|
||
- accepted suggestions still require explicit human confirmation
|
||
- the repository can write accepted watch / alias / stopword / interest-keyword changes after confirmation
|
||
- this governance flow is not yet wired into a periodic scheduler inside the repository
|
||
|
||
Boundary:
|
||
|
||
- keyword cleanup is a maintenance flow, not the production RSS ingestion path
|
||
- the repository does not auto-apply cleanup suggestions without confirmation
|
||
- current keyword stats are built from the delivered candidate payload, not yet from a final `DailyDigest`
|
||
|
||
## Known Limitations
|
||
|
||
- Some sources put only partial content in RSS; those items may be skipped if RSS content is insufficient.
|
||
- WeChat articles often block direct crawling, but this pipeline now avoids that path for FreshRSS items and uses RSS-provided content when available.
|
||
- Rule behavior is still conservative in some cases; many items may land in `review` depending on current rules.
|
||
- Paywall heuristics may produce false positives for some Chinese text patterns.
|
||
- Keyword cleanup governance is usable now, but periodic scheduling and before/after evaluation are not implemented yet.
|
||
|
||
## Files OpenClaw Should Read First
|
||
|
||
Recommended reading order for a new maintainer:
|
||
|
||
1. `README.md`
|
||
2. `docs/openclaw/openclaw-handoff.md`
|
||
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||
5. `docs/design/daily-keyword-index-design.md`
|
||
6. `skills/keyword-cleanup-review/SKILL.md`
|
||
7. `docs/current/context-reset-brief.md`
|
||
|
||
## Downstream Boundary Rules
|
||
|
||
For the daily-digest workflow:
|
||
|
||
- the digest should go to Hugo and chat reporting, not directly into IMA
|
||
- the full daily digest should **not** be uploaded to IMA
|
||
- only explicitly user-selected article summaries should be uploaded to IMA
|
||
- selected-article summaries should be generated from existing extracted text, not by re-fetching original URLs
|
||
|
||
## Current Recommendation
|
||
|
||
For integration handoff, the repository is usable now.
|
||
|
||
The minimum you need to give OpenClaw is:
|
||
|
||
- the repository code
|
||
- the MCP server startup command
|
||
- the required environment variables in the target environment
|
||
- the instruction to call `run_freshrss_openclaw_pipeline`
|
||
- the rule that follow-up state/result reads must go through MCP tools first, not handwritten filesystem paths
|
||
|
||
If OpenClaw will also participate in keyword-governance review, additionally point it to:
|
||
|
||
- `docs/design/daily-keyword-index-design.md`
|
||
- `skills/keyword-cleanup-review/SKILL.md`
|
||
- `scripts/apply_term_suggestions.py`
|