274 lines
16 KiB
Markdown
274 lines
16 KiB
Markdown
---
|
||
name: reader-digest-flow
|
||
description: Orchestrate the end-to-end daily digest workflow around the reader project. Use when the user wants to run the AI daily digest flow, generate a daily report from reader payloads, publish the digest to Hugo, report the digest back in chat, select valuable articles, and then summarize only the selected articles into IMA knowledge notes. Triggers include requests like '跑今天日报', '生成日报', '汇报今天内容', '把选中的文章沉淀', '更新 Hugo', or any request to operate the reader → digest → selection → knowledge-base flow.
|
||
---
|
||
|
||
# Reader Digest Flow
|
||
|
||
Run the reader-based daily digest as a fixed SOP. Treat this skill as the orchestrator for the workflow; do not move Hugo, Feishu reporting, or IMA upload logic into the reader project itself.
|
||
|
||
## Core Rules
|
||
|
||
- Run `reader` as the formal upstream MCP workflow service for fetching, extraction, filtering, payload generation, result reading, and selected-article post-processing.
|
||
- **Daily digest production runs should default to a random article count between 5 and 10 unless the user explicitly specifies a count.**
|
||
- Prefer MCP workflow operations over direct long-lived CLI execution. Treat CLI as debug / fallback, not the default production path.
|
||
- **Operational note from validation:** a debug/test run may still complete successfully inside reader even when the synchronous MCP wrapper returns timeout. In that case, use MCP status/result query tools (`list_runs`, `get_run_report`, `get_delivery_payload`, article-summary async job status/result tools) to continue from the real run truth, rather than treating the whole flow as failed.
|
||
- For reader run observation and result reading, prefer MCP tools such as status / payload / report queries instead of having OpenClaw or this skill hand-build reader output paths.
|
||
- **For normal production runs, mark processed FreshRSS items as read. Only skip mark-read when the user explicitly says the run is debug/test/validation.**
|
||
- **For normal production runs, do not pass `debug_artifacts=true`. Only enable debug artifacts when the user explicitly says the run is debug/test/validation or when troubleshooting is the goal.**
|
||
- **Do not generate the daily digest from examples or placeholder data. Always require a real payload first.**
|
||
- Generate the daily digest markdown from the payload in OpenClaw.
|
||
- Split outputs into a **public digest** for Hugo and an **internal review digest** for chat / operator decision-making.
|
||
- **Public digest and internal review digest are orchestration-layer outputs, and by default are generated by the current OpenClaw session model rather than inheriting reader `LLM_*` settings.**
|
||
- Publish only the public digest to Hugo.
|
||
- Once a normal daily digest run succeeds and a public digest is generated, proceed directly with Hugo publication as the default action; do not ask the user for a separate confirmation about generating or publishing the Hugo page unless the user explicitly says to skip Hugo.
|
||
- Report the internal review digest back to the user in chat.
|
||
- **Do not upload the full daily digest to IMA.**
|
||
- **Do not upload any selected article summary to IMA until the user has explicitly confirmed the selection. Once the user has confirmed which articles to keep, proceed directly with IMA deposition and do not ask for a second confirmation about uploading to the knowledge base.**
|
||
- Upload **only user-selected articles** to IMA as individual knowledge notes.
|
||
- **The formal IMA upload object must be the generated single-article summary content, uploaded directly into the target IMA knowledge base as a Markdown knowledge item (`media_type=7`), not the original article webpage URL.**
|
||
- **For daily deposition, the target must explicitly be the IMA "knowledge-base" type path, not the IMA "notes" type path.**
|
||
- **Do not use webpage URL import or note-then-link (`media_type=11`) as the default daily deposition path. Importing original article URLs into IMA can be used only as temporary source collection, and does not count as completing the daily summary deposition flow.**
|
||
- **Adding content as an IMA note, or uploading to any non-knowledge-base container first and then linking it indirectly, does not count as completing the SOP. Completion requires direct knowledge-base ingestion.**
|
||
- Default IMA target for daily single-article summaries is the `daily` knowledge base.
|
||
- Resolve that default target from reader `.env` (`IMA_DAILY_KNOWLEDGE_BASE_ID`, `IMA_DAILY_KNOWLEDGE_BASE_NAME`) and verify it at runtime before upload.
|
||
- If the configured daily knowledge base is missing, try to locate it by name; if still missing, create `daily` and continue.
|
||
- For selected article notes, use the dedicated article-summary flow in `reader`.
|
||
- Prefer `ARTICLE_SUMMARY_*` LLM settings for selected article summaries; fall back to the main `LLM_*` settings only if the dedicated settings are absent.
|
||
- Do not re-fetch original article URLs for selected summaries; always use the existing extracted article text.
|
||
- Default selected-article summary output should be organized by date, for example under `outputs/freshrss/single_summaries/YYYY-MM-DD/`.
|
||
|
||
## Fixed Flow
|
||
|
||
### Phase 1: Run reader pipeline
|
||
|
||
In the `reader` project, run the FreshRSS pipeline and obtain a real payload.
|
||
|
||
Minimum expected artifacts:
|
||
|
||
- `outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json`
|
||
- `outputs/freshrss/rerun/<run-id>/candidates/digest-brief.json`(public digest 优先读取的轻量输入,仅包含 `keep` 候选)
|
||
- `outputs/freshrss/rerun/<run-id>/run-report.json`
|
||
- extracted article data, such as:
|
||
- `outputs/freshrss/extracted/freshrss.extracted.json`
|
||
|
||
If the pipeline fails, stop and report the exact failure point.
|
||
|
||
### Phase 2: Generate daily digest markdown
|
||
|
||
Read the real run outputs and generate digest markdown for the day.
|
||
If there is no real payload, stop instead of writing a fake or example digest.
|
||
|
||
For **public digest**, prefer `candidates/digest-brief.json` when present; fall back to `candidates/openclaw-delivery-payload.json` only if the brief file is missing.
|
||
For **internal review digest**, continue reading the full `candidates/openclaw-delivery-payload.json`.
|
||
|
||
Produce two output views from the same run, preferably in **one generation step**:
|
||
|
||
1. **Public digest** — for Hugo / public browsing
|
||
2. **Internal review digest** — for chat reporting and operator decisions
|
||
|
||
Before publishing, persist digest outputs back into the same reader run directory under:
|
||
|
||
- `outputs/freshrss/rerun/<run-id>/digest/public_digest.md`
|
||
- `outputs/freshrss/rerun/<run-id>/digest/internal_review_digest.md`
|
||
- `outputs/freshrss/rerun/<run-id>/digest/combined.json`
|
||
|
||
Then write only the public digest into Hugo using this structure:
|
||
|
||
- `content/daily/YYYY-MM-DD/index.md`
|
||
|
||
The public digest is the browsing layer, not the long-term knowledge layer.
|
||
It must not expose internal workflow states or operator-facing review labels.
|
||
|
||
Recommended **public digest** structure:
|
||
|
||
- `今日概览`
|
||
- `今日重点`
|
||
- `趋势观察`
|
||
- `延伸阅读`
|
||
|
||
Public digest writing rules:
|
||
|
||
- Use `digest-brief.json` as the default source when available.
|
||
- Treat it as a public-only view that already excludes non-keep items.
|
||
- Keep the tone suitable for public browsing and Hugo publishing.
|
||
- Style should follow `references/public-digest-example.md` as the default public-writing example.
|
||
- Public digest is a **public reading draft / editor-style public note**, not a workflow report.
|
||
- In `今日概览`, focus on the day’s topic lines, shared signals, and broader industry movement; do **not** describe filtering mechanics or internal selection process.
|
||
- Do **not** expose internal workflow labels or operator language such as `待确认`, `建议沉淀到 IMA`, `keep/review/drop`, or `selection_decision`.
|
||
- Explicitly avoid wording such as `共筛出`, `候选`, `保留`, `入选`, `待确认`, `建议沉淀` in public digest.
|
||
- Prefer concise but information-dense writing.
|
||
- For each item under `今日重点`, include not only summary and highlights, but also one short editor-style value sentence, for example: `这篇内容更值得关注的原因在于……`.
|
||
- When rendering highlights in public digest, prefer a short label such as `值得关注:` followed by one item per line, instead of packing multiple points into a single long sentence.
|
||
- In `延伸阅读`, every item must include source attribution in the form: `- [标题](url)|来源`.
|
||
|
||
Recommended **internal review digest** structure:
|
||
|
||
- `今日候选概况`
|
||
- `已入选重点`
|
||
- `待你确认`
|
||
- `建议沉淀到 IMA`
|
||
- `原始候选清单`
|
||
|
||
Internal review digest writing rules:
|
||
|
||
- Use the full `openclaw-delivery-payload.json`.
|
||
- Keep this as a human-readable review draft rather than a raw machine dump.
|
||
- Do **not** display `rank` values.
|
||
- Convert machine states to Chinese operator-facing labels:
|
||
- `keep` → `已入选`
|
||
- `review` → `待确认`
|
||
- `drop` → `暂不纳入`
|
||
- For every item under `已入选重点`, include:
|
||
- title + source
|
||
- status
|
||
- a fuller summary paragraph
|
||
- a short judgment paragraph explaining why it matters in today's digest
|
||
- For every item under `待你确认`, include:
|
||
- title + source
|
||
- status
|
||
- a fuller summary paragraph
|
||
- reason
|
||
- recommendation
|
||
- In `原始候选清单`, also use Chinese status labels instead of raw machine values.
|
||
|
||
### Phase 3: Publish to Hugo
|
||
|
||
Publish only the public digest to Hugo.
|
||
|
||
Required execution steps:
|
||
|
||
1. write the public digest markdown to:
|
||
- `/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md`
|
||
2. immediately redeploy Hugo:
|
||
- `cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh`
|
||
3. verify before moving on:
|
||
- homepage works
|
||
- list page works
|
||
- detail page works
|
||
- latest digest is visible
|
||
|
||
Do not treat “markdown file written” as equivalent to publish success. Hugo publish for this SOP is complete only after redeploy and page verification pass.
|
||
|
||
Do not block on style polish unless the user explicitly asks.
|
||
|
||
### Phase 4: Report digest back to the user
|
||
|
||
Send the internal review digest in chat and ask the user which articles should be retained for long-term knowledge.
|
||
|
||
**Hard reporting requirement:** the chat report must not be only a title list. For every reported article, include at least:
|
||
- title
|
||
- one-sentence summary
|
||
- a short reason explaining why it matters / why it is recommended or pending confirmation
|
||
|
||
At this step:
|
||
|
||
- the public digest is already in Hugo
|
||
- the internal review digest stays in chat / operator workflow
|
||
- the digest is **not** uploaded to IMA
|
||
- the user decides which articles are worth preserving
|
||
|
||
### Phase 5: Summarize selected articles
|
||
|
||
For every article explicitly selected by the user:
|
||
|
||
1. use the `reader` article-summary capability
|
||
2. point it at the existing extracted payload
|
||
3. pass the selected `item_id` values
|
||
4. generate one markdown summary per article
|
||
|
||
Preferred routes:
|
||
|
||
- async MCP job path:
|
||
- `start_article_summary_job`
|
||
- `get_article_summary_job_status`
|
||
- `get_article_summary_job_result`
|
||
- local fallback: run the reader article-summary workflow directly inside the project `.venv`
|
||
- CLI fallback: `scripts/run_article_summaries.py`
|
||
|
||
Formal production rule:
|
||
- For normal production deposition, selected-article summary generation should default to the async MCP job path instead of synchronous `generate_article_summaries`.
|
||
- Start the job, poll status until `success` / `failed`, then read `written_paths` from the job result.
|
||
- Treat synchronous `generate_article_summaries` as a debug / light-weight helper, not the default production entry.
|
||
|
||
Hard fallback rule:
|
||
- If the async article-summary job path fails because of MCP/tool-layer timeout, transport failure, or job-launch failure, do **not** stop the daily deposition flow.
|
||
- In that case, immediately fall back to running the reader article-summary path locally inside `/home/ubuntu/zhu/github/reader` with the project `.venv`.
|
||
- Treat a successful local article-summary run as equivalent completion for the summary-generation phase; async MCP job is the preferred entry, not a single point of failure.
|
||
|
||
Use the real extracted JSON structure already produced by the project. Do not invent alternative inputs.
|
||
|
||
### Phase 6: Upload selected article notes to IMA
|
||
|
||
Upload only the generated single-article markdown summaries to IMA.
|
||
|
||
**Hard gate before upload:** even if the markdown file was generated by `reader`, do **not** upload it to IMA as-is. You must first reformat/check it against the IMA-facing Markdown rules in this skill, then upload the formatted version. Treat `reader` output as article-summary source material, not automatically as final IMA-ready Markdown.
|
||
|
||
**Hard naming rule before upload:** the final uploaded Markdown filename must use the article's user-facing natural title (normally the original Chinese article title) plus `.md`. Do **not** use internal workflow names, slugs, prefixes, or temp filenames such as `ima-*`, `item-*`, `summary-*`, or English-only shorthand as the final IMA object name.
|
||
|
||
Default target knowledge base for this phase:
|
||
|
||
- `daily`
|
||
- Read from reader `.env` via `IMA_DAILY_KNOWLEDGE_BASE_ID` and `IMA_DAILY_KNOWLEDGE_BASE_NAME`
|
||
- Verify the target at runtime through IMA APIs / skill lookups before upload
|
||
- If the configured target does not exist, try to find `daily` by name; if still absent, create it and continue
|
||
|
||
Completion criteria for this phase:
|
||
|
||
- a selected article has been summarized from extracted content
|
||
- a single-article Markdown summary file has been generated
|
||
- the final upload filename has been normalized to a user-facing article title (not an internal slug / temp name)
|
||
- the Markdown file is uploaded directly into the target knowledge base as a Markdown knowledge item (`media_type=7`)
|
||
- the uploaded object preserves source link context and uses IMA-friendly layout for readability
|
||
- the upload target is the `daily` knowledge base unless the user explicitly requests otherwise
|
||
|
||
Do not treat the following as completion of summary deposition:
|
||
|
||
- importing the original article webpage URL into IMA
|
||
- storing the original article only as source material without the generated summary content
|
||
- creating a note first and then linking that note into the knowledge base as the default path
|
||
|
||
Do not upload:
|
||
|
||
- the full daily digest
|
||
- raw payloads
|
||
- raw extraction output
|
||
|
||
IMA Markdown layout guidance for selected article deposition:
|
||
|
||
- Prefer direct Markdown file upload into the knowledge base (`media_type=7`).
|
||
- Before every upload, open and check the actual markdown file that will be uploaded; do not assume the generator already matched IMA style.
|
||
- The upload target must be an IMA-facing formatted markdown file, not the raw default output from `reader` if the styles differ.
|
||
- Remove stiff metadata headers such as `Source:` / `Category:` when preparing the IMA-facing Markdown.
|
||
- Keep source traceability by placing `原文链接:` near the top, followed by the original URL on the next line.
|
||
- Break long prose under `核心结论` and `主要论点` into short paragraphs for IMA readability instead of relying on platform auto-formatting.
|
||
- The final upload filename should normally be `<文章标题>.md`; only when a same-name file already exists should you append a timestamp suffix before `.md`.
|
||
- If local working files use internal slugs or prefixes for convenience, create or rename a final upload copy before calling IMA upload APIs.
|
||
- Use `references/ima-markdown-example.md` as the formatting example whenever preparing the final upload file.
|
||
- Completion for IMA upload requires both: (1) upload API success, and (2) the uploaded markdown having passed the above format checks.
|
||
|
||
## Operational Guidance
|
||
|
||
- Prefer real run outputs over examples.
|
||
- Verify at each boundary with real files or accessible URLs.
|
||
- When validating selected article summaries, confirm that a markdown file is actually generated.
|
||
- If Codex or another coding agent is asked to implement workflow changes inside `reader`, keep the project boundary clean:
|
||
- workflow logic in `reader`
|
||
- orchestration logic in this skill / OpenClaw
|
||
|
||
## Key Paths
|
||
|
||
### Reader project
|
||
|
||
- `/home/ubuntu/zhu/github/reader`
|
||
|
||
### Hugo project
|
||
|
||
- `/home/ubuntu/zhu/apps/hugo-site`
|
||
- digest content root:
|
||
- `/home/ubuntu/zhu/apps/hugo-site/content/daily/`
|
||
|
||
## References
|
||
|
||
Read `references/flow.md` when you need the concrete step-by-step command checklist and file expectations.
|