Files
openclaw-skills/reader-digest-flow/references/flow.md
T

227 lines
7.5 KiB
Markdown

# Reader Digest Flow Reference
## Purpose
Concrete operational checklist for the `reader-digest-flow` skill.
## Default Operating Model
### Layering
- `reader` layer:
- FreshRSS pull
- extraction
- summary/filter/payload generation
- selected-article summary capability
- OpenClaw / skill layer:
- public digest generation
- internal review digest generation
- Hugo publishing
- chat reporting
- user confirmation handling
- calling selected-article summaries
- IMA upload orchestration
- Hugo layer:
- public digest browsing and archive only
- IMA layer:
- long-term storage for selected article notes only
### Hard rules
- Do not upload the full digest to IMA.
- Upload only explicitly user-selected articles to IMA.
- Do not generate a digest without a real payload.
- Generate two views from the same payload: a public digest for Hugo and an internal review digest for chat/operator workflow.
- Do not expose internal review states or operator-facing labels in the public digest.
- Do not re-fetch original URLs for selected summaries; use existing extracted text.
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization, with fallback to main `LLM_*` only if needed.
- Actively report progress after each completed phase.
## Step-by-step checklist
### 1. Run reader pipeline
Use the formal MCP workflow path as the default production route.
Prefer MCP run/status/result operations over direct path stitching. Only fall back to CLI or direct file inspection for debug / manual troubleshooting.
Project root:
```bash
/home/ubuntu/zhu/github/reader
```
Default behavior for a normal production run:
- run with mark-read enabled
- only skip mark-read if the user explicitly says the run is debug, test, or validation
Typical artifacts to inspect after a successful run:
```text
outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json
outputs/freshrss/rerun/<run-id>/candidates/digest-brief.json
outputs/freshrss/rerun/<run-id>/run-report.json
outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json
```
Notes:
- `digest-brief.json` is the preferred input for **public digest** generation.
- It is a lighter public-only view and currently includes only `keep` candidates.
- If `digest-brief.json` is missing, fall back to `openclaw-delivery-payload.json`.
### 2. Generate digest markdown
Generate two output views from the same run, preferably in one model call:
1. a **public digest** for Hugo / public readers
2. an **internal review digest** for chat / operator workflow
Input preference:
- **public digest**: prefer `candidates/digest-brief.json`
- **internal review digest**: use `candidates/openclaw-delivery-payload.json`
Recommended generation pattern:
- Pass the public brief and the full payload as two explicitly labeled input blocks.
- Ask the model to return both outputs in one response.
- Prefer a structured response shape (for example JSON with `public_digest_markdown` and `internal_review_digest_markdown`) when post-processing is needed.
Before publishing, persist the generated digest artifacts back into the same reader run directory:
```text
outputs/freshrss/rerun/<run-id>/digest/public_digest.md
outputs/freshrss/rerun/<run-id>/digest/internal_review_digest.md
outputs/freshrss/rerun/<run-id>/digest/combined.json
```
Then write only the public digest into Hugo here:
```text
/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md
```
Recommended public-digest front matter:
```toml
+++
title = "AI 日报 · YYYY-MM-DD"
date = YYYY-MM-DDTHH:MM:SS+08:00
summary = "当日日报摘要"
+++
```
Recommended **public digest** structure:
- `今日概览`
- `今日重点`
- `趋势观察`
- `延伸阅读`
Public digest constraints:
- Use only the public brief view when present.
- Keep public wording free of internal workflow language.
- Optimize for concise public readability with solid information density.
- For each `今日重点` item, add one short editor-style sentence explaining why the item matters in today's digest (for example: `这篇内容更值得关注的原因在于……`).
- Prefer rendering highlight points as a short public label such as `值得关注:` followed by one bullet per line.
Recommended **internal review digest** structure:
- `今日候选概况`
- `已入选重点`
- `待你确认`
- `建议沉淀到 IMA`
- `原始候选清单`
Internal review digest constraints:
- Do not show `rank` values.
- Replace machine states with Chinese labels such as `已入选` / `待确认` / `暂不纳入`.
- For `已入选重点`, include a fuller summary plus a short judgment paragraph.
- For `待你确认`, include a fuller summary, reason, and recommendation.
- Keep it readable as an operator review draft, not a raw payload dump.
### 3. Publish Hugo
Publish only the public digest to Hugo.
Hard execution steps:
1. write the public digest markdown to:
- `/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md`
2. redeploy Hugo immediately after writing:
- `cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh`
3. verify all three URLs before continuing:
- `http://127.0.0.1:14322/`
- `http://127.0.0.1:14322/daily/`
- `http://127.0.0.1:14322/daily/YYYY-MM-DD/`
Expected verification targets:
- homepage works
- `/daily/` works
- `/daily/YYYY-MM-DD/` works
### 4. Report digest in chat
Provide the internal review digest in chat and ask which articles should be retained.
### 5. Generate selected article summaries
Only do this after the user explicitly confirms which articles to retain.
Preferred MCP tool:
- `generate_article_summaries`
Expected inputs:
- `extracted_path`
- `selected_ids`
- optional `output_dir`
- optional article-summary LLM overrides
Recommended output layout:
- `outputs/freshrss/single_summaries/YYYY-MM-DD/`
CLI fallback:
```bash
python scripts/run_article_summaries.py \
--extracted outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json \
--ids <item_id_1> <item_id_2> \
--output-dir outputs/freshrss/single_summaries/YYYY-MM-DD
```
### 6. Upload selected summaries to IMA
Upload only the generated markdown files for the selected articles.
Hard execution rules before upload:
1. reformat/check the generated markdown into IMA-facing final content
2. normalize the final upload filename to `<文章标题>.md`
3. do not use internal temp names such as `ima-*`, `item-*`, `summary-*`, or English slug filenames as the final uploaded object name
4. if the knowledge base already contains the same filename, append a timestamp suffix before `.md`
5. if local work needs internal temp names, create a final upload copy with the user-facing title before calling IMA APIs
Default target knowledge base:
- `daily`
- Read `IMA_DAILY_KNOWLEDGE_BASE_ID` / `IMA_DAILY_KNOWLEDGE_BASE_NAME` from reader `.env`
- Verify the configured target at runtime before upload
- If the configured target is unavailable, resolve by name `daily`; if still absent, create `daily`
## Hard rules recap
- Never generate a digest from placeholder or example data when a real run is expected.
- For normal runs, mark processed FreshRSS items as read unless the user explicitly requested a debug/test/validation run.
- Public digest goes to Hugo; internal review digest goes to chat; neither full digest goes to IMA.
- Only explicitly user-selected articles go to IMA.
- Selected article summaries use extracted text, not live refetch.
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization.
- Final IMA upload filenames must use user-facing article titles, not internal slugs or workflow temp names.