Files
reader/docs/openclaw/openclaw-handoff.md
T

361 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# OpenClaw Handoff
## Purpose
This repository provides a FreshRSS-first reading pipeline for OpenClaw:
`FreshRSS unread items -> RSS content extraction -> LLM summary -> rule engine -> OpenClaw delivery payload`
OpenClaw should treat this repository as an MCP-backed upstream content processor.
This repository is responsible only for upstream reading-pipeline work:
- FreshRSS pull
- content extraction
- LLM summary generation/validation
- rule-based filtering
- OpenClaw delivery payload generation
- selected-article summary capability based on existing extracted text
- run-state persistence, status lookup, result lookup, and minimal resume for the FreshRSS workflow
This repository should **not** take over downstream orchestration responsibilities that belong to OpenClaw / skills, such as:
- Hugo publishing
- chat reporting
- user confirmation handling
- IMA upload orchestration
## Production Entrypoint
reader 当前正式工作流服务启动入口是 MCP tool:
- `start_freshrss_pipeline_job`
OpenClaw 应先拿到 `job_id`,轮询 job 状态,再在成功后读取 `run_id` 作为正式后续句柄。
`run_freshrss_openclaw_pipeline` 仍保留,但定位是同步 debug / fallback 路径,而不是正式生产启动入口。
OpenClaw should treat the returned `run_id` from `get_freshrss_pipeline_job_result` as the only stable handle for follow-up reads. Do not hand-build `outputs/freshrss/rerun/...` paths in OpenClaw.
Job state is written under `outputs/freshrss/pipeline_jobs/<job_id>/` and will minimally contain `run-state.json`, `input.json`, `result.json` on success, and `job-report.json`.
## Supported MCP Tools
Current MCP tools: 17 total, including the FreshRSS workflow set plus async job tools for both the main pipeline and article-summary flow.
Workflow service tools:
- `start_freshrss_pipeline_job`
- `get_freshrss_pipeline_job_status`
- `get_freshrss_pipeline_job_result`
- `get_run_status`
- `list_runs`
- `list_run_artifacts`
- `get_delivery_payload`
- `get_run_report`
- `resume_run`
Article-summary tools:
- `start_article_summary_job`
- `get_article_summary_job_status`
- `get_article_summary_job_result`
Single-step / debug tools:
- `run_freshrss_openclaw_pipeline`(同步模式,仅适合 debug / fallback)
- `extract_url_content`
- `extract_item_content`
- `filter_summary_result`
- `generate_article_summaries`(同步模式,仅适合轻量调试)
## Recommended Selected-Article Flow
For OpenClaw selected-article follow-up, prefer the async job path:
1. Call `start_article_summary_job` with a real extracted file path plus a non-empty `selected_ids` list.
2. Poll `get_article_summary_job_status` until `status` becomes `success` or `failed`.
3. On success, call `get_article_summary_job_result` and continue downstream processing from `written_paths`.
4. Use `generate_article_summaries` only as a synchronous debug fallback, not as the default production path.
Job state is written under `outputs/freshrss/article_summary_jobs/<job_id>/` and will minimally contain:
- `run-state.json`
- `input.json`
- `result.json` on success
- `job-report.json`
## Required Environment Variables
The MCP server process must have these variables available:
- `FRESHRSS_API_BASE_URL`
- `FRESHRSS_USERNAME`
- `FRESHRSS_API_PASSWORD`
- `LLM_API_URL`
- `LLM_API_KEY`
- `LLM_MODEL`
Example:
```powershell
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
set FRESHRSS_USERNAME=osiman
set FRESHRSS_API_PASSWORD=your-api-password
set LLM_API_URL=https://api.deepseek.com
set LLM_API_KEY=your-llm-api-key
set LLM_MODEL=deepseek-chat
```
## Server Startup
Install dependencies:
```bash
pip install -e .
```
Start the MCP server:
```bash
summary-mcp
```
## Recommended MCP Workflow
Recommended production path:
1. Call `start_freshrss_pipeline_job` and persist the returned `job_id`
2. Poll `get_freshrss_pipeline_job_status(job_id)` until `status` becomes `success` or `failed`
3. On success, call `get_freshrss_pipeline_job_result(job_id)` and persist the returned `run_id`
4. Use `get_run_status(run_id)` as the authoritative run-state read for status, stage, artifacts, and recovery
5. Use `list_runs(...)` when OpenClaw needs recent-run discovery or high-level inspection
6. Use `list_run_artifacts(run_id)` when OpenClaw needs to inspect what this run actually produced
7. Use `get_delivery_payload(run_id)` and `get_run_report(run_id)` as the formal result-reading APIs
8. Use `resume_run(run_id)` only when the run falls inside the minimal supported resume scope
OpenClaw should not directly derive or hardcode:
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
If filesystem access is needed for debugging, consume only paths returned by MCP such as `output_dir`, `artifact.path`, `delivery_output`, or `report_output`.
## Recommended MCP Call
Recommended production start call:
```json
{
"limit": 5,
"mark_read": true,
"include_read": false,
"debug_artifacts": false,
"timeout_seconds": 60,
"max_retries": 2
}
```
Recommended semantics:
- Use `mark_read=true` for normal production runs.
- Use `mark_read=false` only for debug, test, or validation runs.
- Keep `debug_artifacts=false` for routine production runs.
- Set `debug_artifacts=true` only when troubleshooting a bad batch.
- Treat the returned `job_id` as the startup handle, and the later `run_id` from `get_freshrss_pipeline_job_result` as the stable identifier for all follow-up run reads.
- If no real `openclaw-delivery-payload.json` was produced, OpenClaw should stop instead of generating a digest from placeholders or examples.
- Use `run_freshrss_openclaw_pipeline` only when a synchronous debug / fallback path is explicitly needed.
## Formal Capability Boundary
reader 当前正式 MCP workflow service 的边界如下:
- formal workflow: only `freshrss_daily_digest`
- run truth: every FreshRSS run writes `run-state.json`
- main production start path: `start_freshrss_pipeline_job` / `get_freshrss_pipeline_job_status` / `get_freshrss_pipeline_job_result`
- state query tools: `get_run_status`, `list_runs`, `list_run_artifacts`
- result read tools: `get_delivery_payload`, `get_run_report`
- `digest-brief.json` is generated and registered as an artifact, but there is no standalone `get_digest_brief` tool yet
- `run_freshrss_openclaw_pipeline` is still supported, but only as a synchronous debug / fallback path
- the FreshRSS daily workflow now has a minimal background job model backed by a detached runner process, not a full queue / worker system
- article summary now has a minimal asynchronous job model with `start_article_summary_job` / `get_article_summary_job_status` / `get_article_summary_job_result`
- `generate_article_summaries` is still supported, but it is a synchronous debug path and outside the formal `resume_run` scope
Historical compatibility note:
- `get_run_status` / `list_runs` / `list_run_artifacts` can still infer basic state for older run directories without `run-state.json`
- `resume_run` does **not** support those inferred historical runs; it requires a valid `run-state.json`
## What The Async Job Returns
Primary return fields from `get_freshrss_pipeline_job_result`:
- `job_id`
- `run_id`
- `output_dir`
- `raw_output`
- `delivery_output`
- `report_output`
- `digest_brief_output`
- `pulled_count`
- `delivered_count`
- `marked_read_count`
- `status_counts`
- `keyword_index`
Optional:
- `items`
- returned only when `include_item_reports=true`
`run_freshrss_openclaw_pipeline` still returns the same synchronous payload for debug / fallback use.
Follow-up structured reads should use MCP tools rather than re-reading these files directly.
## Minimal Output Files
By default the pipeline writes these core artifacts:
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
- `outputs/freshrss/rerun/<run_dir>/raw/freshrss.raw.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/digest-brief.json`
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
- `outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json` (one per item)
It also updates local runtime keyword data:
- `data/term_index/daily/YYYY-MM-DD.json`
- `data/term_index/term_stats.json`
Per-item extracted files live under `extracted/` and are always written.
If `debug_artifacts=true`, the pipeline additionally writes normalized items, summaries, filter decisions, candidate records, and candidate inputs.
The main pipeline does not emit a batch-level `freshrss.extracted.json` file by default.
## `resume_run` Minimal Scope
`resume_run` currently supports only the minimum resume contract:
- only runs with a valid `run-state.json`
- only workflow `freshrss_daily_digest`
- resume in place on the original `run_id`
- supported resume points: `generate_summaries`, `apply_filters`, `build_delivery_payload`, `write_run_report`
- unsupported resume points: `fetch_feed`, `extract_articles`
- if required artifacts are missing, the tool returns a non-resumable response instead of silently falling back to an earlier stage
Artifact expectations by resume point:
- `generate_summaries`: requires `raw/freshrss.raw.json` and `extracted/`
- `apply_filters`: requires the above plus per-item summary outputs
- `build_delivery_payload`: requires per-item candidate inputs consistent with filter-stage output
- `write_run_report`: requires `candidates/openclaw-delivery-payload.json`; if `mark_read=true`, raw input must still be present
## Payload Specs
OpenClaw payload field specs live here:
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
- `docs/openclaw/openclaw-delivery-payload-spec.md`
## Read-State Semantics
The pipeline reads from FreshRSS unread items by default.
If `mark_read=true`:
- items are marked as read only after the final `openclaw-delivery-payload.json` has been written successfully
- only successfully delivered items are marked as read
- failed or skipped items remain unread
## FreshRSS Content Policy
For FreshRSS items, the pipeline is RSS-first and does not re-crawl webpages.
Behavior:
- use `item.raw_content` first
- if missing, use `item.raw_summary`
- if neither contains usable content, skip the item
- do not fetch the original webpage again for FreshRSS items
This is intentional.
## Keyword Cleanup Governance
This repository also includes a lightweight keyword-governance flow for downstream review.
Current pieces:
- runtime keyword stats
- `data/term_index/daily/YYYY-MM-DD.json`
- `data/term_index/term_stats.json`
- governance config
- `configs/term_cleanup_policy.json`
- `configs/term_watchlist.json`
- `configs/term_change_log.json`
- review bundle builder
- `skills/keyword-cleanup-review/scripts/build_review_bundle.py`
- accepted-suggestion writer
- `scripts/apply_term_suggestions.py`
Current status:
- OpenClaw can read the keyword review bundle as maintenance input
- accepted suggestions still require explicit human confirmation
- the repository can write accepted watch / alias / stopword / interest-keyword changes after confirmation
- this governance flow is not yet wired into a periodic scheduler inside the repository
Boundary:
- keyword cleanup is a maintenance flow, not the production RSS ingestion path
- the repository does not auto-apply cleanup suggestions without confirmation
- current keyword stats are built from the delivered candidate payload, not yet from a final `DailyDigest`
## Known Limitations
- Some sources put only partial content in RSS; those items may be skipped if RSS content is insufficient.
- WeChat articles often block direct crawling, but this pipeline now avoids that path for FreshRSS items and uses RSS-provided content when available.
- Rule behavior is still conservative in some cases; many items may land in `review` depending on current rules.
- Paywall heuristics may produce false positives for some Chinese text patterns.
- Keyword cleanup governance is usable now, but periodic scheduling and before/after evaluation are not implemented yet.
## Files OpenClaw Should Read First
Recommended reading order for a new maintainer:
1. `README.md`
2. `docs/openclaw/openclaw-handoff.md`
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
5. `docs/design/daily-keyword-index-design.md`
6. `skills/keyword-cleanup-review/SKILL.md`
7. `docs/current/context-reset-brief.md`
## Downstream Boundary Rules
For the daily-digest workflow:
- the digest should go to Hugo and chat reporting, not directly into IMA
- the full daily digest should **not** be uploaded to IMA
- only explicitly user-selected article summaries should be uploaded to IMA
- selected-article summaries should be generated from existing extracted text, not by re-fetching original URLs
## Current Recommendation
For integration handoff, the repository is usable now.
The minimum you need to give OpenClaw is:
- the repository code
- the MCP server startup command
- the required environment variables in the target environment
- the instruction to call `run_freshrss_openclaw_pipeline`
- the rule that follow-up state/result reads must go through MCP tools first, not handwritten filesystem paths
If OpenClaw will also participate in keyword-governance review, additionally point it to:
- `docs/design/daily-keyword-index-design.md`
- `skills/keyword-cleanup-review/SKILL.md`
- `scripts/apply_term_suggestions.py`