196 lines
5.9 KiB
Markdown
196 lines
5.9 KiB
Markdown
# Reader Digest Flow Reference
|
|
|
|
## Purpose
|
|
|
|
Concrete operational checklist for the `reader-digest-flow` skill.
|
|
|
|
## Default Operating Model
|
|
|
|
### Layering
|
|
|
|
- `reader` layer:
|
|
- FreshRSS pull
|
|
- extraction
|
|
- summary/filter/payload generation
|
|
- selected-article summary capability
|
|
- OpenClaw / skill layer:
|
|
- public digest generation
|
|
- internal review digest generation
|
|
- Hugo publishing
|
|
- chat reporting
|
|
- user confirmation handling
|
|
- calling selected-article summaries
|
|
- IMA upload orchestration
|
|
- Hugo layer:
|
|
- public digest browsing and archive only
|
|
- IMA layer:
|
|
- long-term storage for selected article notes only
|
|
|
|
### Hard rules
|
|
|
|
- Do not upload the full digest to IMA.
|
|
- Upload only explicitly user-selected articles to IMA.
|
|
- Do not generate a digest without a real payload.
|
|
- Generate two views from the same payload: a public digest for Hugo and an internal review digest for chat/operator workflow.
|
|
- Do not expose internal review states or operator-facing labels in the public digest.
|
|
- Do not re-fetch original URLs for selected summaries; use existing extracted text.
|
|
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization, with fallback to main `LLM_*` only if needed.
|
|
- Actively report progress after each completed phase.
|
|
|
|
## Step-by-step checklist
|
|
|
|
### 1. Run reader pipeline
|
|
|
|
Project root:
|
|
|
|
```bash
|
|
/home/ubuntu/zhu/github/reader
|
|
```
|
|
|
|
Default behavior for a normal production run:
|
|
|
|
- run with mark-read enabled
|
|
- only skip mark-read if the user explicitly says the run is debug, test, or validation
|
|
|
|
Typical artifacts to inspect after a successful run:
|
|
|
|
```text
|
|
outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json
|
|
outputs/freshrss/rerun/<run-id>/candidates/digest-brief.json
|
|
outputs/freshrss/rerun/<run-id>/run-report.json
|
|
outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json
|
|
```
|
|
|
|
Notes:
|
|
|
|
- `digest-brief.json` is the preferred input for **public digest** generation.
|
|
- It is a lighter public-only view and currently includes only `keep` candidates.
|
|
- If `digest-brief.json` is missing, fall back to `openclaw-delivery-payload.json`.
|
|
|
|
### 2. Generate digest markdown
|
|
|
|
Generate two output views from the same run, preferably in one model call:
|
|
|
|
1. a **public digest** for Hugo / public readers
|
|
2. an **internal review digest** for chat / operator workflow
|
|
|
|
Input preference:
|
|
|
|
- **public digest**: prefer `candidates/digest-brief.json`
|
|
- **internal review digest**: use `candidates/openclaw-delivery-payload.json`
|
|
|
|
Recommended generation pattern:
|
|
|
|
- Pass the public brief and the full payload as two explicitly labeled input blocks.
|
|
- Ask the model to return both outputs in one response.
|
|
- Prefer a structured response shape (for example JSON with `public_digest_markdown` and `internal_review_digest_markdown`) when post-processing is needed.
|
|
|
|
Write only the public digest into Hugo here:
|
|
|
|
```text
|
|
/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md
|
|
```
|
|
|
|
Recommended public-digest front matter:
|
|
|
|
```toml
|
|
+++
|
|
title = "AI 日报 · YYYY-MM-DD"
|
|
date = YYYY-MM-DDTHH:MM:SS+08:00
|
|
summary = "当日日报摘要"
|
|
+++
|
|
```
|
|
|
|
Recommended **public digest** structure:
|
|
|
|
- `今日概览`
|
|
- `今日重点`
|
|
- `趋势观察`
|
|
- `延伸阅读`
|
|
- `信息来源`
|
|
|
|
Public digest constraints:
|
|
|
|
- Use only the public brief view when present.
|
|
- Keep public wording free of internal workflow language.
|
|
- Optimize for concise public readability with solid information density.
|
|
- For each `今日重点` item, add one short editor-style sentence explaining why the item matters in today's digest (for example: `这篇内容更值得关注的原因在于……`).
|
|
|
|
Recommended **internal review digest** structure:
|
|
|
|
- `今日候选概况`
|
|
- `已入选重点`
|
|
- `待你确认`
|
|
- `建议沉淀到 IMA`
|
|
- `原始候选清单`
|
|
|
|
Internal review digest constraints:
|
|
|
|
- Do not show `rank` values.
|
|
- Replace machine states with Chinese labels such as `已入选` / `待确认` / `暂不纳入`.
|
|
- For `已入选重点`, include a fuller summary plus a short judgment paragraph.
|
|
- For `待你确认`, include a fuller summary, reason, and recommendation.
|
|
- Keep it readable as an operator review draft, not a raw payload dump.
|
|
|
|
### 3. Publish Hugo
|
|
|
|
Publish only the public digest to Hugo.
|
|
|
|
Expected verification targets:
|
|
|
|
- homepage works
|
|
- `/daily/` works
|
|
- `/daily/YYYY-MM-DD/` works
|
|
|
|
### 4. Report digest in chat
|
|
|
|
Provide the internal review digest in chat and ask which articles should be retained.
|
|
|
|
### 5. Generate selected article summaries
|
|
|
|
Only do this after the user explicitly confirms which articles to retain.
|
|
|
|
Preferred MCP tool:
|
|
|
|
- `generate_article_summaries`
|
|
|
|
Expected inputs:
|
|
|
|
- `extracted_path`
|
|
- `selected_ids`
|
|
- optional `output_dir`
|
|
- optional article-summary LLM overrides
|
|
|
|
Recommended output layout:
|
|
|
|
- `outputs/freshrss/single_summaries/YYYY-MM-DD/`
|
|
|
|
CLI fallback:
|
|
|
|
```bash
|
|
python scripts/run_article_summaries.py \
|
|
--extracted outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json \
|
|
--ids <item_id_1> <item_id_2> \
|
|
--output-dir outputs/freshrss/single_summaries/YYYY-MM-DD
|
|
```
|
|
|
|
### 6. Upload selected summaries to IMA
|
|
|
|
Upload only the generated markdown files for the selected articles.
|
|
|
|
Default target knowledge base:
|
|
|
|
- `daily`
|
|
- Read `IMA_DAILY_KNOWLEDGE_BASE_ID` / `IMA_DAILY_KNOWLEDGE_BASE_NAME` from reader `.env`
|
|
- Verify the configured target at runtime before upload
|
|
- If the configured target is unavailable, resolve by name `daily`; if still absent, create `daily`
|
|
|
|
## Hard rules recap
|
|
|
|
- Never generate a digest from placeholder or example data when a real run is expected.
|
|
- For normal runs, mark processed FreshRSS items as read unless the user explicitly requested a debug/test/validation run.
|
|
- Public digest goes to Hugo; internal review digest goes to chat; neither full digest goes to IMA.
|
|
- Only explicitly user-selected articles go to IMA.
|
|
- Selected article summaries use extracted text, not live refetch.
|
|
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization.
|