Files
reader/docs/openclaw/openclaw-handoff.md
T

270 lines
7.4 KiB
Markdown

# OpenClaw Handoff
## Role
This file is the integration overview for OpenClaw maintainers.
Use it for:
- reader capability boundary
- production MCP entrypoints
- environment requirements
- integration rules and limitations
Do not use it as the step-by-step runbook.
For formal orchestration, read `docs/openclaw/openclaw-orchestration-flow.md`.
For field contracts, read:
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
- `docs/openclaw/openclaw-delivery-payload-spec.md`
Historical plans and incident documents live under `docs/openclaw/archive/`.
## Purpose
reader is the upstream FreshRSS processing service for OpenClaw:
`FreshRSS unread items -> RSS content extraction -> LLM summary -> rule engine -> OpenClaw delivery payload`
reader is responsible for:
- FreshRSS pull
- content extraction
- LLM summary generation and validation
- rule-based filtering
- OpenClaw delivery payload generation
- run-state persistence and run/result lookup
- async resume control for the FreshRSS workflow
- async selected-article summary generation from existing extracted files
reader is not responsible for:
- Hugo publishing
- chat reporting
- user confirmation handling
- IMA upload orchestration
## Production Surface
Current MCP tool count: 21.
Main daily workflow:
- `start_freshrss_pipeline_job`
- `get_freshrss_pipeline_job_status`
- `get_freshrss_pipeline_job_result`
- `get_run_status`
- `list_runs`
- `list_run_artifacts`
- `get_delivery_payload`
- `get_run_report`
Resume workflow:
- `inspect_resume_plan`
- `start_resume_job`
- `get_resume_job_status`
- `get_resume_job_result`
- `resume_run`
Selected-article summary workflow:
- `start_article_summary_job`
- `get_article_summary_job_status`
- `get_article_summary_job_result`
- `generate_article_summaries`
Debug / single-step tools:
- `run_freshrss_openclaw_pipeline`
- `extract_url_content`
- `extract_item_content`
- `filter_summary_result`
Production rules:
- main production start path is `start_freshrss_pipeline_job`
- production resume path is `inspect_resume_plan -> start_resume_job -> get_resume_job_status -> get_resume_job_result`
- `run_freshrss_openclaw_pipeline` is sync debug / fallback only
- `resume_run` is sync debug / fallback only
- `generate_article_summaries` is sync debug / fallback only
## Production Contract
OpenClaw should treat the returned `run_id` from `get_freshrss_pipeline_job_result` as the only stable handle for follow-up reads.
OpenClaw should not hand-build these paths:
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
If filesystem access is needed for debugging, only consume paths returned by MCP:
- `output_dir`
- `artifact.path`
- `delivery_output`
- `report_output`
Top-level `status` is the only status field callers should branch on.
`status_source` and `state_conflict` are explanatory fields for reconciled status.
## Minimal Production Sequence
Daily workflow:
1. Call `start_freshrss_pipeline_job`
2. Poll `get_freshrss_pipeline_job_status`
3. On success, read `get_freshrss_pipeline_job_result`
4. Persist the returned `run_id`
5. Use `get_run_status`, `get_delivery_payload`, and `get_run_report` for follow-up reads
Resume workflow:
1. Call `inspect_resume_plan(run_id)`
2. Only if `can_resume=true` and `recommended_action=resume`, call `start_resume_job`
3. Poll `get_resume_job_status`
4. Read `get_resume_job_result`
Selected-article summary workflow:
1. Call `start_article_summary_job` with a real extracted file path and non-empty `selected_ids`
2. Poll `get_article_summary_job_status`
3. Read `get_article_summary_job_result`
## Capability Boundary
Formal workflow boundary:
- only workflow `freshrss_daily_digest`
- every current FreshRSS run writes `run-state.json`
- `get_run_status` / `list_runs` / `list_run_artifacts` can still infer basic state for older runs without `run-state.json`
- resume requires a valid `run-state.json`; inferred historical runs are not resumable
Resume boundary:
- resume in place on the original `run_id`
- supported resume points:
- `generate_summaries`
- `apply_filters`
- `build_delivery_payload`
- `write_run_report`
- unsupported resume points:
- `fetch_feed`
- `extract_articles`
- production resume prefers:
- `summary/summary-batch.json`
- `candidates/candidate-batch.json`
- if required artifacts are missing, recovery should return non-resumable instead of silently falling back
Selected-article summary boundary:
- uses existing extracted files as input
- should not re-fetch original URLs
## Output Expectations
Main daily pipeline core artifacts:
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
- `outputs/freshrss/rerun/<run_dir>/raw/freshrss.raw.json`
- `outputs/freshrss/rerun/<run_dir>/summary/summary-batch.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/candidate-batch.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_dir>/candidates/digest-brief.json`
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
- `outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json`
Async job state directories:
- main pipeline job: `outputs/freshrss/pipeline_jobs/<job_id>/`
- resume job: `outputs/freshrss/resume_jobs/<job_id>/`
- article-summary job: `outputs/freshrss/article_summary_jobs/<job_id>/`
Each job directory minimally contains:
- `run-state.json`
- `input.json`
- `result.json` on success
- `job-report.json`
## Environment And Startup
Required environment variables:
- `FRESHRSS_API_BASE_URL`
- `FRESHRSS_USERNAME`
- `FRESHRSS_API_PASSWORD`
- `LLM_API_URL`
- `LLM_API_KEY`
- `LLM_MODEL`
Startup:
```bash
pip install -e .
summary-mcp
```
Recommended production start call:
```json
{
"limit": 5,
"mark_read": true,
"include_read": false,
"debug_artifacts": false,
"timeout_seconds": 60,
"max_retries": 2
}
```
## Data And Content Policy
FreshRSS processing is RSS-first:
- use `item.raw_content` first
- if missing, use `item.raw_summary`
- if neither contains usable content, skip the item
- do not fetch the original webpage again for FreshRSS items
Read-state policy:
- items are marked read only after successful delivery payload write
- only successfully delivered items are marked read
Downstream boundary:
- the daily digest goes to Hugo and chat reporting
- the full daily digest should not be uploaded to IMA
- only explicitly user-selected article summaries should be uploaded to IMA
## Related Maintenance Flow
Keyword cleanup exists as a separate maintenance flow, not the main RSS ingestion path.
Relevant files:
- `docs/design/daily-keyword-index-design.md`
- `skills/keyword-cleanup-review/SKILL.md`
- `scripts/apply_term_suggestions.py`
## Known Limitations
- some sources expose only partial RSS content; those items may be skipped
- rule behavior is still conservative; many items may land in `review`
- paywall heuristics may still produce false positives on some Chinese text
- keyword cleanup governance is usable but not yet wired to periodic scheduling
## Read First
Recommended reading order for a new maintainer:
1. `README.md`
2. `docs/openclaw/README.md`
3. `docs/openclaw/openclaw-handoff.md`
4. `docs/openclaw/openclaw-orchestration-flow.md`
5. `docs/openclaw/openclaw-candidate-input-field-spec.md`
6. `docs/openclaw/openclaw-delivery-payload-spec.md`
7. `docs/current/context-reset-brief.md`