163 lines
4.3 KiB
Markdown
163 lines
4.3 KiB
Markdown
# Reader Digest Flow Reference
|
|
|
|
## Purpose
|
|
|
|
Concrete operational checklist for the `reader-digest-flow` skill.
|
|
|
|
## Default Operating Model
|
|
|
|
### Layering
|
|
|
|
- `reader` layer:
|
|
- FreshRSS pull
|
|
- extraction
|
|
- summary/filter/payload generation
|
|
- selected-article summary capability
|
|
- OpenClaw / skill layer:
|
|
- public digest generation
|
|
- internal review digest generation
|
|
- Hugo publishing
|
|
- chat reporting
|
|
- user confirmation handling
|
|
- calling selected-article summaries
|
|
- IMA upload orchestration
|
|
- Hugo layer:
|
|
- public digest browsing and archive only
|
|
- IMA layer:
|
|
- long-term storage for selected article notes only
|
|
|
|
### Hard rules
|
|
|
|
- Do not upload the full digest to IMA.
|
|
- Upload only explicitly user-selected articles to IMA.
|
|
- Do not generate a digest without a real payload.
|
|
- Generate two views from the same payload: a public digest for Hugo and an internal review digest for chat/operator workflow.
|
|
- Do not expose internal review states or operator-facing labels in the public digest.
|
|
- Do not re-fetch original URLs for selected summaries; use existing extracted text.
|
|
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization, with fallback to main `LLM_*` only if needed.
|
|
- Actively report progress after each completed phase.
|
|
|
|
## Step-by-step checklist
|
|
|
|
### 1. Run reader pipeline
|
|
|
|
Project root:
|
|
|
|
```bash
|
|
/home/ubuntu/zhu/github/reader
|
|
```
|
|
|
|
Default behavior for a normal production run:
|
|
|
|
- run with mark-read enabled
|
|
- only skip mark-read if the user explicitly says the run is debug, test, or validation
|
|
|
|
Typical artifacts to inspect after a successful run:
|
|
|
|
```text
|
|
outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json
|
|
outputs/freshrss/rerun/<run-id>/run-report.json
|
|
outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json
|
|
```
|
|
|
|
### 2. Generate digest markdown
|
|
|
|
Generate two output views from the same payload:
|
|
|
|
1. a **public digest** for Hugo / public readers
|
|
2. an **internal review digest** for chat / operator workflow
|
|
|
|
Write only the public digest into Hugo here:
|
|
|
|
```text
|
|
/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md
|
|
```
|
|
|
|
Recommended public-digest front matter:
|
|
|
|
```toml
|
|
+++
|
|
title = "AI 日报 · YYYY-MM-DD"
|
|
date = YYYY-MM-DDTHH:MM:SS+08:00
|
|
summary = "当日日报摘要"
|
|
+++
|
|
```
|
|
|
|
Recommended **public digest** structure:
|
|
|
|
- `今日概览`
|
|
- `今日重点`
|
|
- `趋势观察`
|
|
- `延伸阅读`
|
|
- `信息来源`
|
|
|
|
Recommended **internal review digest** structure:
|
|
|
|
- `今日候选概况`
|
|
- `已入选重点`
|
|
- `待你确认`
|
|
- `建议沉淀到 IMA`
|
|
- `原始候选清单`
|
|
|
|
### 3. Publish Hugo
|
|
|
|
Publish only the public digest to Hugo.
|
|
|
|
Expected verification targets:
|
|
|
|
- homepage works
|
|
- `/daily/` works
|
|
- `/daily/YYYY-MM-DD/` works
|
|
|
|
### 4. Report digest in chat
|
|
|
|
Provide the internal review digest in chat and ask which articles should be retained.
|
|
|
|
### 5. Generate selected article summaries
|
|
|
|
Only do this after the user explicitly confirms which articles to retain.
|
|
|
|
Preferred MCP tool:
|
|
|
|
- `generate_article_summaries`
|
|
|
|
Expected inputs:
|
|
|
|
- `extracted_path`
|
|
- `selected_ids`
|
|
- optional `output_dir`
|
|
- optional article-summary LLM overrides
|
|
|
|
Recommended output layout:
|
|
|
|
- `outputs/freshrss/single_summaries/YYYY-MM-DD/`
|
|
|
|
CLI fallback:
|
|
|
|
```bash
|
|
python scripts/run_article_summaries.py \
|
|
--extracted outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json \
|
|
--ids <item_id_1> <item_id_2> \
|
|
--output-dir outputs/freshrss/single_summaries/YYYY-MM-DD
|
|
```
|
|
|
|
### 6. Upload selected summaries to IMA
|
|
|
|
Upload only the generated markdown files for the selected articles.
|
|
|
|
Default target knowledge base:
|
|
|
|
- `daily`
|
|
- Read `IMA_DAILY_KNOWLEDGE_BASE_ID` / `IMA_DAILY_KNOWLEDGE_BASE_NAME` from reader `.env`
|
|
- Verify the configured target at runtime before upload
|
|
- If the configured target is unavailable, resolve by name `daily`; if still absent, create `daily`
|
|
|
|
## Hard rules recap
|
|
|
|
- Never generate a digest from placeholder or example data when a real run is expected.
|
|
- For normal runs, mark processed FreshRSS items as read unless the user explicitly requested a debug/test/validation run.
|
|
- Public digest goes to Hugo; internal review digest goes to chat; neither full digest goes to IMA.
|
|
- Only explicitly user-selected articles go to IMA.
|
|
- Selected article summaries use extracted text, not live refetch.
|
|
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization.
|