- New flow: pipeline → report candidates → user selects Hugo articles → generate public digest → publish Hugo → user selects IMA articles → IMA deposition - digest-brief.json is now the sole data source (concise, less token consumption) - Removed 来源 line from public-digest-example.md to match SKILL hard rule - Public digest no longer auto-generated with all keep items
21 KiB
name, description
| name | description |
|---|---|
| reader-digest-flow | Orchestrate the end-to-end daily digest workflow around the reader project. Use when the user wants to run the AI daily digest flow, report candidates for review, let the user decide which articles go into the Hugo daily digest, then generate and publish the public digest, and optionally summarize selected articles into IMA knowledge notes. Triggers include requests like '跑今天日报', '生成日报', '汇报今天内容', '把选中的文章沉淀', '更新 Hugo', or any request to operate the reader → report → user-selection → Hugo → knowledge-base flow. |
Reader Digest Flow
Run the reader-based daily digest as a fixed SOP. Treat this skill as the orchestrator for the workflow; do not move Hugo, Feishu reporting, or IMA upload logic into the reader project itself.
Core Rules
- Run
readeras the formal upstream MCP workflow service for fetching, extraction, filtering, payload generation, result reading, and selected-article post-processing. - Daily digest production runs should default to a random article count between 5 and 10 unless the user explicitly specifies a count.
- Prefer MCP workflow operations over direct long-lived CLI execution. Treat CLI as debug / fallback, not the default production path.
- Main FreshRSS production runs should default to the async MCP job path:
start_freshrss_pipeline_job→get_freshrss_pipeline_job_status→get_freshrss_pipeline_job_result. Use the synchronousrun_freshrss_openclaw_pipelineonly for debug / light validation / fallback. - If the main pipeline job fails, do not immediately abandon the run. First inspect the linked run through
get_run_statusandinspect_resume_plan; when the plan returnsrecommended_action=resume, continue throughstart_resume_job→get_resume_job_status→get_resume_job_result. - Operational note from validation: before the main pipeline async job existed, a debug/test run could still complete successfully inside reader even when the synchronous MCP wrapper returned timeout. For historical/debug cases, continue from real run truth using MCP status/result query tools rather than treating the whole flow as failed.
- For reader run observation and result reading, prefer MCP tools such as status / payload / report queries instead of having OpenClaw or this skill hand-build reader output paths.
- When reader status APIs return reconciled status, always branch on the top-level
status. Treatstatus_sourceandstate_conflictas explanatory metadata; do not re-derive flow control from stale stage names. - For normal production runs, mark processed FreshRSS items as read. Only skip mark-read when the user explicitly says the run is debug/test/validation.
- For normal production runs, do not pass
debug_artifacts=true. Only enable debug artifacts when the user explicitly says the run is debug/test/validation or when troubleshooting is the goal. - Do not generate the daily digest from examples or placeholder data. Always require a real payload first.
- Generate the daily digest markdown from the payload in OpenClaw.
- Split outputs into a public digest for Hugo and an internal review digest for chat / operator decision-making.
- Public digest and internal review digest are orchestration-layer outputs, and by default are generated by the current OpenClaw session model rather than inheriting reader
LLM_*settings. - Publish only the public digest to Hugo, and only after the user has explicitly selected which articles go into the daily digest.
- The default flow is: pipeline → report candidates → user selects Hugo articles → generate public digest → publish Hugo → (optional) user selects IMA articles → IMA deposition.
- Do NOT generate the public digest or publish Hugo before the user has confirmed the Hugo article selection.
- Report the internal review digest back to the user in chat.
- Do not upload the full daily digest to IMA.
- Do not upload any selected article summary to IMA until the user has explicitly confirmed the selection. Once the user has confirmed which articles to keep, proceed directly with IMA deposition and do not ask for a second confirmation about uploading to the knowledge base.
- Upload only user-selected articles to IMA as individual knowledge notes.
- The formal IMA upload object must be the generated single-article summary content, uploaded directly into the target IMA knowledge base as a Markdown knowledge item (
media_type=7), not the original article webpage URL. - For daily deposition, the target must explicitly be the IMA "knowledge-base" type path, not the IMA "notes" type path.
- Do not use webpage URL import or note-then-link (
media_type=11) as the default daily deposition path. Importing original article URLs into IMA can be used only as temporary source collection, and does not count as completing the daily summary deposition flow. - Adding content as an IMA note, or uploading to any non-knowledge-base container first and then linking it indirectly, does not count as completing the SOP. Completion requires direct knowledge-base ingestion.
- Default IMA target for daily single-article summaries is the
dailyknowledge base. - Resolve that default target from reader
.env(IMA_DAILY_KNOWLEDGE_BASE_ID,IMA_DAILY_KNOWLEDGE_BASE_NAME) and verify it at runtime before upload. - If the configured daily knowledge base is missing, try to locate it by name; if still missing, create
dailyand continue. - For selected article notes, use the dedicated article-summary flow in
reader. - Prefer
ARTICLE_SUMMARY_*LLM settings for selected article summaries; fall back to the mainLLM_*settings only if the dedicated settings are absent. - Do not re-fetch original article URLs for selected summaries; always use the existing extracted article text.
- Default selected-article summary output should be organized by date, for example under
outputs/freshrss/single_summaries/YYYY-MM-DD/.
Fixed Flow
Phase 1: Run reader pipeline
In the reader project, run the FreshRSS pipeline and obtain a real payload.
Formal production start path:
start_freshrss_pipeline_jobget_freshrss_pipeline_job_statusget_freshrss_pipeline_job_result
After the async job succeeds, treat the returned run_id as the stable handle for downstream get_run_status / get_delivery_payload / get_run_report reads.
Minimum expected artifacts:
outputs/freshrss/rerun/<run-id>/candidates/digest-brief.json— 精简候选清单(唯一要读取的文件,含 keep + review)outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json— 完整 payload(pipeline 内部产物,digest-brief 由此推导,流程中无需读取)outputs/freshrss/rerun/<run-id>/run-report.json- extracted article data:
outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json等(每个候选一篇,格式为{"success": true, "article": {...}},Phase 6 使用)
If the main pipeline job fails:
- read the linked run through
get_run_status - call
inspect_resume_plan(run_id) - if
recommended_action=resume, continue with:start_resume_jobget_resume_job_statusget_resume_job_result
- if
recommended_action=read_terminal_result, continue from the terminal run result instead of retrying - if
recommended_action=start_new_run, stop and report the exact failure point
Phase 2: Report candidates to the user (internal review digest only)
After the pipeline run succeeds, do NOT generate the public digest or publish Hugo yet. Read candidates and produce only the internal review digest in a concise format (not the full payload dump).
Recommended internal review digest structure:
今日候选概况已入选重点待你确认原始候选清单
Data source: use candidates/digest-brief.json (top_candidates array) — it now includes both keep and review articles in a concise, simplified format (title, source, summary, highlights, category). No need to read the full payload.
Internal review digest writing rules:
- Keep this as a human-readable concise review draft, not a raw machine dump.
- Do not display
rankordigest_rankvalues. - Convert machine states to Chinese operator-facing labels:
keep→已入选review→待确认drop→暂不纳入
- For every item under
已入选重点, include:- title + source
- status (
已入选) - summary paragraph (use the concise version from digest-brief)
- highlights (from digest-brief)
- a short judgment paragraph explaining why it matters
- For every item under
待你确认, include:- title + source
- status (
待确认) - summary paragraph (concise)
- highlights (if available)
- reason(为什么要你确认)
- recommendation(建议采纳/不采纳)
- In
原始候选清单, present all articles as a brief list (title + status + one-liner), using Chinese status labels.
Hard reporting requirement: the chat report must not be only a title list. For every reported article, include at least:
- title
- one-sentence summary
- a short reason explaining why it matters / why it is recommended or pending confirmation
At this step:
- the public digest is not generated yet
- Hugo is not published yet
- the internal review digest is sent to the user in chat
- the user decides: which articles go into the Hugo daily digest AND separately which articles go into IMA deposition
Phase 3: User selects Hugo articles
The user reviews the internal digest and specifies which articles should appear in the Hugo daily digest.
- Confirm the selection explicitly before proceeding.
- If the user wants to include some
reviewcandidate articles, respect that choice. - Only the user-selected articles will appear in the public digest.
Phase 4: Generate and publish Hugo public digest
Generate the public digest markdown only for the user-selected articles, then publish to Hugo.
⚠️ 格式基准 — 开始生成任何 digest 内容之前,必须先完整阅读 references/public-digest-example.md 并以此为格式基准,不得凭记忆或直觉写作。
⚠️ HARD PRE-WRITE CHECKLIST — 对照 references/public-digest-example.md 逐项确认后再开始写:
- frontmatter
summary字段已填写(完整句子,不是关键词罗列,参考格式:"围绕 XXX 的当日深度观察。") - 编号只用
1.2.3.,不用中文数字或罗马数字 - 四个 section 全部存在:
今日概览今日重点趋势观察延伸阅读 - 每篇
今日重点下有:- 标题(无"来源:"字样,来源仅在延伸阅读标注)
- 摘要段落
- "值得关注:"要点列表(每篇 3 条)
- "这篇更值得关注的原因在于:"段落
延伸阅读每条含来源标注:- [标题](url)|来源
Public digest source material: use candidates/digest-brief.json — it has all the info needed (title, summary, highlights, source, url). No need to read the full payload.
Persist the generated digest under:
outputs/freshrss/rerun/<run-id>/digest/public_digest.md
Then publish to Hugo:
- write the public digest markdown to:
/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md
- immediately redeploy Hugo:
cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh
- verify before moving on:
- homepage works
- list page works
- detail page works
- latest digest is visible
Do not treat "markdown file written" as equivalent to publish success. Hugo publish for this SOP is complete only after redeploy and page verification pass.
Do not block on style polish unless the user explicitly asks.
The public digest is the browsing layer, not the long-term knowledge layer. It must not expose internal workflow states or operator-facing review labels.
Public digest writing rules:
- Use article content from
candidates/digest-brief.jsonfor the user-selected subset. - Hugo public digest includes ONLY the articles the user selected in Phase 3. Not all
keepitems — only the user's explicit selection. - Keep the tone suitable for public browsing and Hugo publishing.
- Style should follow
references/public-digest-example.mdas the default public-writing example. - Public digest is a public reading draft / editor-style public note, not a workflow report.
- In
今日概览, focus on the day's topic lines, shared signals, and broader industry movement; do not describe filtering mechanics or internal selection process. - Do not expose internal workflow labels or operator language such as
待确认,建议沉淀到 IMA,keep/review/drop, orselection_decision. - Explicitly avoid wording such as
共筛出,候选,保留,入选,待确认,建议沉淀in public digest. - Prefer concise but information-dense writing.
- For each item under
今日重点, include not only summary and highlights, but also one short editor-style value sentence, for example:这篇内容更值得关注的原因在于……. - When rendering highlights in public digest, prefer a short label such as
值得关注:followed by one item per line, instead of packing multiple points into a single long sentence. - In
延伸阅读, every item must include source attribution in the form:- [标题](url)|来源. - Article numbering: MUST use
1.2.3.(Arabic numeral + period). Do NOT use① ② ③,一、二、三,第一条or any other numbering variant. - All four sections are required:
今日概览,今日重点,趋势观察,延伸阅读. Missing any section is a format violation.
Phase 5: User selects articles for IMA deposition
After Hugo publication, ask the user which articles should be retained for long-term knowledge (separate from the Hugo selection).
- The user can select from all
keepandreviewcandidates, including or excluding articles that were already put in Hugo. - Confirm the selection explicitly before proceeding.
- Only the user-selected articles will be summarized and uploaded to IMA.
At this step:
- the public digest is already in Hugo
- the user decides which articles are worth preserving as knowledge notes
- the digest is not uploaded to IMA as a whole
Phase 6: Summarize selected articles
For every article explicitly selected by the user:
- use the
readerarticle-summary capability - ⚠️ 关键:extracted_path 必须传入 pipeline 产生的 individual item 文件,路径为:
不要传
outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json outputs/freshrss/rerun/<run-id>/extracted/item-02.extracted.json ...candidates/openclaw-delivery-payload.json或outputs/freshrss/extracted/(这两个路径的 payload 格式与 article-summary workflow 不兼容,会报错** - for each selected article, run one summary job per extracted item file
- pass the corresponding
item_idfrom thecandidatesarray 去掉cand:前缀 - generate one markdown summary per article
⚠️ item_id 前缀注意: candidates 里的 ID 格式是 cand:sha256:xxx,但 item 文件里的 item_id 是 sha256:xxx(无 cand: 前缀)。传 selected_ids 时必须去掉前缀,否则匹配失败。
⚠️ item_id 与 extracted 文件编号的对应关系:candidates 数组顺序 ≠ extracted 文件编号顺序。pipeline 提取阶段和 LLM 筛选阶段是两套独立顺序,不能按"第几篇"的位置来映射。
正确做法: 传 selected_ids 前,必须先从 candidates 数组找到目标文章的 item_id(去前缀),再到对应的 item-XX.extracted.json 文件里读取其内部的 article.item_id 做交叉验证,确认匹配后再传。禁止仅凭"文章在 candidates 里排第几"来推断应该用哪个 item-XX 文件。
⚠️ single-item 输入约束: 当 extracted_path 指向 outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json 这类单篇文件时,selected_ids 只应包含这一个文件对应的单个 item_id。不要对单个 extracted 文件传多个 IDs。
Preferred routes:
- async MCP job path:
start_article_summary_jobget_article_summary_job_statusget_article_summary_job_result
- local fallback: run the reader article-summary workflow directly inside the project
.venv - CLI fallback:
scripts/run_article_summaries.py
Formal production rule:
- For normal production deposition, selected-article summary generation should default to the async MCP job path instead of synchronous
generate_article_summaries. - Start the job, poll status until
success/failed, then readwritten_pathsfrom the job result. - Treat synchronous
generate_article_summariesas a debug / light-weight helper, not the default production entry.
Hard fallback rule:
- If the async article-summary job path fails because of MCP/tool-layer timeout, transport failure, or job-launch failure, do not stop the daily deposition flow.
- In that case, immediately fall back to running the reader article-summary path locally inside
/home/ubuntu/zhu/github/readerwith the project.venv. - Treat a successful local article-summary run as equivalent completion for the summary-generation phase; async MCP job is the preferred entry, not a single point of failure.
Use the real extracted JSON structure already produced by the project. Do not invent alternative inputs.
Phase 7: Upload selected article notes to IMA
Upload only the generated single-article markdown summaries to IMA.
Hard gate before upload: even if the markdown file was generated by reader, do not upload it to IMA as-is. You must first reformat/check it against the IMA-facing Markdown rules in this skill, then upload the formatted version. Treat reader output as article-summary source material, not automatically as final IMA-ready Markdown.
Hard naming rule before upload: the final uploaded Markdown filename must use the article's user-facing natural title (normally the original Chinese article title) plus .md. Do not use internal workflow names, slugs, prefixes, or temp filenames such as ima-*, item-*, summary-*, or English-only shorthand as the final IMA object name.
Default target knowledge base for this phase:
daily- Read from reader
.envviaIMA_DAILY_KNOWLEDGE_BASE_IDandIMA_DAILY_KNOWLEDGE_BASE_NAME - Verify the target at runtime through IMA APIs / skill lookups before upload
- If the configured target does not exist, try to find
dailyby name; if still absent, create it and continue
Completion criteria for this phase:
- a selected article has been summarized from extracted content
- a single-article Markdown summary file has been generated
- the final upload filename has been normalized to a user-facing article title (not an internal slug / temp name)
- the Markdown file is uploaded directly into the target knowledge base as a Markdown knowledge item (
media_type=7) - the uploaded object preserves source link context and uses IMA-friendly layout for readability
- the upload target is the
dailyknowledge base unless the user explicitly requests otherwise
Do not treat the following as completion of summary deposition:
- importing the original article webpage URL into IMA
- storing the original article only as source material without the generated summary content
- creating a note first and then linking that note into the knowledge base as the default path
Do not upload:
- the full daily digest
- raw payloads
- raw extraction output
IMA Markdown layout guidance for selected article deposition:
- Prefer direct Markdown file upload into the knowledge base (
media_type=7). - Before every upload, open and check the actual markdown file that will be uploaded; do not assume the generator already matched IMA style.
- The upload target must be an IMA-facing formatted markdown file, not the raw default output from
readerif the styles differ. - Remove stiff metadata headers such as
Source:/Category:when preparing the IMA-facing Markdown. - Keep source traceability by placing
原文链接:near the top, followed by the original URL on the next line. - Break long prose under
核心结论and主要论点into short paragraphs for IMA readability instead of relying on platform auto-formatting. - The final upload filename should normally be
<文章标题>.md; only when a same-name file already exists should you append a timestamp suffix before.md. - If local working files use internal slugs or prefixes for convenience, create or rename a final upload copy before calling IMA upload APIs.
- Use
references/ima-markdown-example.mdas the formatting example whenever preparing the final upload file. - Completion for IMA upload requires both: (1) upload API success, and (2) the uploaded markdown having passed the above format checks.
Operational Guidance
- Prefer real run outputs over examples.
- Verify at each boundary with real files or accessible URLs.
- When validating selected article summaries, confirm that a markdown file is actually generated.
- If Codex or another coding agent is asked to implement workflow changes inside
reader, keep the project boundary clean:- workflow logic in
reader - orchestration logic in this skill / OpenClaw
- workflow logic in
Key Paths
Reader project
/home/ubuntu/zhu/github/reader
Hugo project
/home/ubuntu/zhu/apps/hugo-site- digest content root:
/home/ubuntu/zhu/apps/hugo-site/content/daily/
References
Read references/flow.md when you need the concrete step-by-step command checklist and file expectations.