197 lines
8.4 KiB
Markdown
197 lines
8.4 KiB
Markdown
---
|
||
name: reader-digest-flow
|
||
description: Orchestrate the end-to-end daily digest workflow around the reader project. Use when the user wants to run the AI daily digest flow, generate a daily report from reader payloads, publish the digest to Hugo, report the digest back in chat, select valuable articles, and then summarize only the selected articles into IMA knowledge notes. Triggers include requests like '跑今天日报', '生成日报', '汇报今天内容', '把选中的文章沉淀', '更新 Hugo', or any request to operate the reader → digest → selection → knowledge-base flow.
|
||
---
|
||
|
||
# Reader Digest Flow
|
||
|
||
Run the reader-based daily digest as a fixed SOP. Treat this skill as the orchestrator for the workflow; do not move Hugo, Feishu reporting, or IMA upload logic into the reader project itself.
|
||
|
||
## Core Rules
|
||
|
||
- Run `reader` / MCP for upstream fetching, extraction, filtering, payload generation, and selected-article post-processing.
|
||
- **For normal production runs, mark processed FreshRSS items as read. Only skip mark-read when the user explicitly says the run is debug/test/validation.**
|
||
- **Do not generate the daily digest from examples or placeholder data. Always require a real payload first.**
|
||
- Generate the daily digest markdown from the payload in OpenClaw.
|
||
- Split outputs into a **public digest** for Hugo and an **internal review digest** for chat / operator decision-making.
|
||
- Publish only the public digest to Hugo.
|
||
- Report the internal review digest back to the user in chat.
|
||
- **Do not upload the full daily digest to IMA.**
|
||
- **Do not upload any selected article summary to IMA until the user has explicitly confirmed the selection.**
|
||
- Upload **only user-selected articles** to IMA as individual knowledge notes.
|
||
- Default IMA target for daily single-article summaries is the `daily` knowledge base.
|
||
- Resolve that default target from reader `.env` (`IMA_DAILY_KNOWLEDGE_BASE_ID`, `IMA_DAILY_KNOWLEDGE_BASE_NAME`) and verify it at runtime before upload.
|
||
- If the configured daily knowledge base is missing, try to locate it by name; if still missing, create `daily` and continue.
|
||
- For selected article notes, use the dedicated article-summary flow in `reader`.
|
||
- Prefer `ARTICLE_SUMMARY_*` LLM settings for selected article summaries; fall back to the main `LLM_*` settings only if the dedicated settings are absent.
|
||
- Do not re-fetch original article URLs for selected summaries; always use the existing extracted article text.
|
||
- Default selected-article summary output should be organized by date, for example under `outputs/freshrss/single_summaries/YYYY-MM-DD/`.
|
||
|
||
## Fixed Flow
|
||
|
||
### Phase 1: Run reader pipeline
|
||
|
||
In the `reader` project, run the FreshRSS pipeline and obtain a real payload.
|
||
|
||
Minimum expected artifacts:
|
||
|
||
- `outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json`
|
||
- `outputs/freshrss/rerun/<run-id>/candidates/digest-brief.json`(public digest 优先读取的轻量输入,仅包含 `keep` 候选)
|
||
- `outputs/freshrss/rerun/<run-id>/run-report.json`
|
||
- extracted article data, such as:
|
||
- `outputs/freshrss/extracted/freshrss.extracted.json`
|
||
|
||
If the pipeline fails, stop and report the exact failure point.
|
||
|
||
### Phase 2: Generate daily digest markdown
|
||
|
||
Read the real run outputs and generate digest markdown for the day.
|
||
If there is no real payload, stop instead of writing a fake or example digest.
|
||
|
||
For **public digest**, prefer `candidates/digest-brief.json` when present; fall back to `candidates/openclaw-delivery-payload.json` only if the brief file is missing.
|
||
For **internal review digest**, continue reading the full `candidates/openclaw-delivery-payload.json`.
|
||
|
||
Produce two output views from the same run, preferably in **one generation step**:
|
||
|
||
1. **Public digest** — for Hugo / public browsing
|
||
2. **Internal review digest** — for chat reporting and operator decisions
|
||
|
||
Before publishing, persist digest outputs back into the same reader run directory under:
|
||
|
||
- `outputs/freshrss/rerun/<run-id>/digest/public_digest.md`
|
||
- `outputs/freshrss/rerun/<run-id>/digest/internal_review_digest.md`
|
||
- `outputs/freshrss/rerun/<run-id>/digest/combined.json`
|
||
|
||
Then write only the public digest into Hugo using this structure:
|
||
|
||
- `content/daily/YYYY-MM-DD/index.md`
|
||
|
||
The public digest is the browsing layer, not the long-term knowledge layer.
|
||
It must not expose internal workflow states or operator-facing review labels.
|
||
|
||
Recommended **public digest** structure:
|
||
|
||
- `今日概览`
|
||
- `今日重点`
|
||
- `趋势观察`
|
||
- `延伸阅读`
|
||
|
||
Public digest writing rules:
|
||
|
||
- Use `digest-brief.json` as the default source when available.
|
||
- Treat it as a public-only view that already excludes non-keep items.
|
||
- Keep the tone suitable for public browsing and Hugo publishing.
|
||
- Do **not** expose internal workflow labels or operator language such as `待确认`, `建议沉淀到 IMA`, `keep/review/drop`, or `selection_decision`.
|
||
- Prefer concise but information-dense writing.
|
||
- For each item under `今日重点`, include not only summary and highlights, but also one short editor-style value sentence, for example: `这篇内容更值得关注的原因在于……`.
|
||
- When rendering highlights in public digest, prefer a short label such as `值得关注:` followed by one item per line, instead of packing multiple points into a single long sentence.
|
||
|
||
Recommended **internal review digest** structure:
|
||
|
||
- `今日候选概况`
|
||
- `已入选重点`
|
||
- `待你确认`
|
||
- `建议沉淀到 IMA`
|
||
- `原始候选清单`
|
||
|
||
Internal review digest writing rules:
|
||
|
||
- Use the full `openclaw-delivery-payload.json`.
|
||
- Keep this as a human-readable review draft rather than a raw machine dump.
|
||
- Do **not** display `rank` values.
|
||
- Convert machine states to Chinese operator-facing labels:
|
||
- `keep` → `已入选`
|
||
- `review` → `待确认`
|
||
- `drop` → `暂不纳入`
|
||
- For every item under `已入选重点`, include:
|
||
- title + source
|
||
- status
|
||
- a fuller summary paragraph
|
||
- a short judgment paragraph explaining why it matters in today's digest
|
||
- For every item under `待你确认`, include:
|
||
- title + source
|
||
- status
|
||
- a fuller summary paragraph
|
||
- reason
|
||
- recommendation
|
||
- In `原始候选清单`, also use Chinese status labels instead of raw machine values.
|
||
|
||
### Phase 3: Publish to Hugo
|
||
|
||
Publish only the public digest to Hugo and verify:
|
||
|
||
- list page works
|
||
- detail page works
|
||
- latest digest is visible
|
||
|
||
Do not block on style polish unless the user explicitly asks.
|
||
|
||
### Phase 4: Report digest back to the user
|
||
|
||
Send the internal review digest in chat and ask the user which articles should be retained for long-term knowledge.
|
||
|
||
At this step:
|
||
|
||
- the public digest is already in Hugo
|
||
- the internal review digest stays in chat / operator workflow
|
||
- the digest is **not** uploaded to IMA
|
||
- the user decides which articles are worth preserving
|
||
|
||
### Phase 5: Summarize selected articles
|
||
|
||
For every article explicitly selected by the user:
|
||
|
||
1. use the `reader` article-summary capability
|
||
2. point it at the existing extracted payload
|
||
3. pass the selected `item_id` values
|
||
4. generate one markdown summary per article
|
||
|
||
Preferred routes:
|
||
|
||
- MCP tool: `generate_article_summaries`
|
||
- CLI fallback: `scripts/run_article_summaries.py`
|
||
|
||
Use the real extracted JSON structure already produced by the project. Do not invent alternative inputs.
|
||
|
||
### Phase 6: Upload selected article notes to IMA
|
||
|
||
Upload only the generated single-article markdown summaries to IMA.
|
||
|
||
Default target knowledge base for this phase:
|
||
|
||
- `daily`
|
||
- Read from reader `.env` via `IMA_DAILY_KNOWLEDGE_BASE_ID` and `IMA_DAILY_KNOWLEDGE_BASE_NAME`
|
||
- Verify the target at runtime through IMA APIs / skill lookups before upload
|
||
- If the configured target does not exist, try to find `daily` by name; if still absent, create it and continue
|
||
|
||
Do not upload:
|
||
|
||
- the full daily digest
|
||
- raw payloads
|
||
- raw extraction output
|
||
|
||
## Operational Guidance
|
||
|
||
- Prefer real run outputs over examples.
|
||
- Verify at each boundary with real files or accessible URLs.
|
||
- When validating selected article summaries, confirm that a markdown file is actually generated.
|
||
- If Codex or another coding agent is asked to implement workflow changes inside `reader`, keep the project boundary clean:
|
||
- workflow logic in `reader`
|
||
- orchestration logic in this skill / OpenClaw
|
||
|
||
## Key Paths
|
||
|
||
### Reader project
|
||
|
||
- `/home/ubuntu/zhu/github/reader`
|
||
|
||
### Hugo project
|
||
|
||
- `/home/ubuntu/zhu/apps/hugo-site`
|
||
- digest content root:
|
||
- `/home/ubuntu/zhu/apps/hugo-site/content/daily/`
|
||
|
||
## References
|
||
|
||
Read `references/flow.md` when you need the concrete step-by-step command checklist and file expectations.
|