# Reader Digest Flow Reference ## Purpose Concrete operational checklist for the `reader-digest-flow` skill. ## Default Operating Model ### Layering - `reader` layer: - FreshRSS pull - extraction - summary/filter/payload generation - selected-article summary capability - OpenClaw / skill layer: - public digest generation - internal review digest generation - Hugo publishing - chat reporting - user confirmation handling - calling selected-article summaries - IMA upload orchestration - Hugo layer: - public digest browsing and archive only - IMA layer: - long-term storage for selected article notes only ### Hard rules - Do not upload the full digest to IMA. - Upload only explicitly user-selected articles to IMA. - Do not generate a digest without a real payload. - Generate two views from the same payload: a public digest for Hugo and an internal review digest for chat/operator workflow. - Do not expose internal review states or operator-facing labels in the public digest. - Do not re-fetch original URLs for selected summaries; use existing extracted text. - Prefer `ARTICLE_SUMMARY_*` for selected article summarization, with fallback to main `LLM_*` only if needed. - Actively report progress after each completed phase. ## Step-by-step checklist ### 1. Run reader pipeline Use the formal MCP workflow path as the default production route. Prefer MCP run/status/result operations over direct path stitching. Only fall back to CLI or direct file inspection for debug / manual troubleshooting. Formal production startup sequence: 1. `start_freshrss_pipeline_job` 2. `get_freshrss_pipeline_job_status` 3. `get_freshrss_pipeline_job_result` 4. after success, continue with `run_id` via `get_run_status` / `get_delivery_payload` / `get_run_report` Treat the old synchronous `run_freshrss_openclaw_pipeline` as debug / light validation / fallback only. Project root: ```bash /home/ubuntu/zhu/github/reader ``` Default behavior for a normal production run: - if the user did not specify a count, randomly choose a limit between 5 and 10 items for that run - run with mark-read enabled - do not enable `debug_artifacts` - only skip mark-read if the user explicitly says the run is debug, test, or validation - only enable `debug_artifacts` if the user explicitly says the run is debug, test, validation, or troubleshooting Typical artifacts to inspect after a successful run: ```text outputs/freshrss/rerun//candidates/openclaw-delivery-payload.json outputs/freshrss/rerun//candidates/digest-brief.json outputs/freshrss/rerun//run-report.json outputs/freshrss/rerun//extracted/item-XX.extracted.json ``` Notes: - `digest-brief.json` is the preferred input for **public digest** generation. - It is a lighter public-only view and currently includes only `keep` candidates. - If `digest-brief.json` is missing, fall back to `openclaw-delivery-payload.json`. - If the synchronous MCP wrapper times out but a real reader run was still created, do not discard the run; continue from run truth using `list_runs`, `get_run_report`, and `get_delivery_payload`. ### 2. Generate digest markdown Generate two output views from the same run, preferably in one model call: 1. a **public digest** for Hugo / public readers 2. an **internal review digest** for chat / operator workflow Input preference: - **public digest**: prefer `candidates/digest-brief.json` - **internal review digest**: use `candidates/openclaw-delivery-payload.json` Recommended generation pattern: - Pass the public brief and the full payload as two explicitly labeled input blocks. - Ask the model to return both outputs in one response. - Prefer a structured response shape (for example JSON with `public_digest_markdown` and `internal_review_digest_markdown`) when post-processing is needed. Before publishing, persist the generated digest artifacts back into the same reader run directory: ```text outputs/freshrss/rerun//digest/public_digest.md outputs/freshrss/rerun//digest/internal_review_digest.md outputs/freshrss/rerun//digest/combined.json ``` Then write only the public digest into Hugo here: ```text /home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md ``` Recommended public-digest front matter: ```toml +++ title = "AI 日报 · YYYY-MM-DD" date = YYYY-MM-DDTHH:MM:SS+08:00 summary = "当日日报摘要" +++ ``` Recommended **public digest** structure: - `今日概览` - `今日重点` - `趋势观察` - `延伸阅读` Public digest constraints: - Use only the public brief view when present. - Keep public wording free of internal workflow language. - Optimize for concise public readability with solid information density. - For each `今日重点` item, add one short editor-style sentence explaining why the item matters in today's digest (for example: `这篇内容更值得关注的原因在于……`). - Prefer rendering highlight points as a short public label such as `值得关注:` followed by one bullet per line. - **Article numbering: MUST use `1.` `2.` `3.` (Arabic numeral + period). Do NOT use `① ② ③`,`一、二、三`,`第一条` or any other variant.** - **All four sections are required. Missing any one is a format violation.** - **Hugo digest must include ALL keep articles from `digest-brief.json`. The IMA deposition subset is a separate downstream step.** Recommended **internal review digest** structure: - `今日候选概况` - `已入选重点` - `待你确认` - `建议沉淀到 IMA` - `原始候选清单` Internal review digest constraints: - Do not show `rank` values. - Replace machine states with Chinese labels such as `已入选` / `待确认` / `暂不纳入`. - For `已入选重点`, include a fuller summary plus a short judgment paragraph. - For `待你确认`, include a fuller summary, reason, and recommendation. - Keep it readable as an operator review draft, not a raw payload dump. ### 3. Publish Hugo Publish only the public digest to Hugo. Treat Hugo publication as the default continuation of a successful normal daily digest run. Do not ask for a second confirmation before generating/writing the public digest and publishing it, unless the user explicitly requests not to publish to Hugo. Hard execution steps: 1. write the public digest markdown to: - `/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md` 2. redeploy Hugo immediately after writing: - `cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh` 3. verify all three URLs before continuing: - `http://127.0.0.1:14322/` - `http://127.0.0.1:14322/daily/` - `http://127.0.0.1:14322/daily/YYYY-MM-DD/` Expected verification targets: - homepage works - `/daily/` works - `/daily/YYYY-MM-DD/` works ### 4. Report digest in chat Provide the internal review digest in chat and ask which articles should be retained. Hard reporting rule: - do not send only article titles - for each article, include at least a one-sentence summary and a short recommendation / judgment so the user can decide without reopening the source ### 5. Generate selected article summaries Only do this after the user explicitly confirms which articles to retain. Preferred MCP path: - `start_article_summary_job` - `get_article_summary_job_status` - `get_article_summary_job_result` Expected inputs: - `extracted_path` - `selected_ids` - optional `output_dir` - optional article-summary LLM overrides Recommended output layout: - `outputs/freshrss/single_summaries/YYYY-MM-DD/` - async job state: `outputs/freshrss/article_summary_jobs//` Recommended production sequence: 1. call `start_article_summary_job` 2. poll `get_article_summary_job_status` until `status` becomes `success` or `failed` 3. on success, call `get_article_summary_job_result` and continue downstream from `written_paths` Hard fallback rule: - If the async MCP job path returns timeout / transport failure / job-launch failure (for example MCP timeout while the reader article-summary workflow itself is still healthy), do not treat that as article-summary business failure. - Immediately retry through the local reader environment under `/home/ubuntu/zhu/github/reader` using the project `.venv`, calling the article-summary workflow directly. - The production goal is successful generation of the selected-article Markdown files; async MCP job is preferred, but local `.venv` execution is the required fallback path. Synchronous helper: - `generate_article_summaries` remains available for debug / light validation only, not as the default production path. CLI fallback: ```bash python scripts/run_article_summaries.py \ --extracted outputs/freshrss/rerun//extracted/item-01.extracted.json \ --ids \ --output-dir outputs/freshrss/single_summaries/YYYY-MM-DD ``` ### 6. Upload selected summaries to IMA Upload only the generated markdown files for the selected articles. Once the user has selected the articles to retain, treat that selection itself as the authorization to continue the IMA deposition step; do not ask for a second confirmation about uploading into the knowledge base. Hard execution rules before upload: 1. reformat/check the generated markdown into IMA-facing final content 2. normalize the final upload filename to `<文章标题>.md` 3. do not use internal temp names such as `ima-*`, `item-*`, `summary-*`, or English slug filenames as the final uploaded object name 4. if the knowledge base already contains the same filename, append a timestamp suffix before `.md` 5. if local work needs internal temp names, create a final upload copy with the user-facing title before calling IMA APIs Default target knowledge base: - `daily` - Read `IMA_DAILY_KNOWLEDGE_BASE_ID` / `IMA_DAILY_KNOWLEDGE_BASE_NAME` from reader `.env` - Verify the configured target at runtime before upload - If the configured target is unavailable, resolve by name `daily`; if still absent, create `daily` ## Hard rules recap - Never generate a digest from placeholder or example data when a real run is expected. - For normal runs, mark processed FreshRSS items as read unless the user explicitly requested a debug/test/validation run. - Public digest goes to Hugo; internal review digest goes to chat; neither full digest goes to IMA. - Only explicitly user-selected articles go to IMA. - All daily IMA deposition must go directly into the IMA knowledge-base path as Markdown knowledge items (`media_type=7`), not through the IMA notes path. - Uploading to IMA notes, or creating notes first and then linking them into a knowledge base, does not count as SOP completion. - Selected article summaries use extracted text, not live refetch. - Prefer `ARTICLE_SUMMARY_*` for selected article summarization. - Final IMA upload filenames must use user-facing article titles, not internal slugs or workflow temp names.