- 编号强制使用 1. 2. 3.(禁止 ① ② ③ 等变体) - 四个 section 缺一不可:今日概览/今日重点/趋势观察/延伸阅读 - Hugo 公开版必须包含 digest-brief.json 全部 keep 文章 - IMA 沉淀为下游独立步骤,不影响 Hugo 内容 - public-digest-example.md 顶部加格式警告注释
270 lines
11 KiB
Markdown
270 lines
11 KiB
Markdown
# Reader Digest Flow Reference
|
||
|
||
## Purpose
|
||
|
||
Concrete operational checklist for the `reader-digest-flow` skill.
|
||
|
||
## Default Operating Model
|
||
|
||
### Layering
|
||
|
||
- `reader` layer:
|
||
- FreshRSS pull
|
||
- extraction
|
||
- summary/filter/payload generation
|
||
- selected-article summary capability
|
||
- OpenClaw / skill layer:
|
||
- public digest generation
|
||
- internal review digest generation
|
||
- Hugo publishing
|
||
- chat reporting
|
||
- user confirmation handling
|
||
- calling selected-article summaries
|
||
- IMA upload orchestration
|
||
- Hugo layer:
|
||
- public digest browsing and archive only
|
||
- IMA layer:
|
||
- long-term storage for selected article notes only
|
||
|
||
### Hard rules
|
||
|
||
- Do not upload the full digest to IMA.
|
||
- Upload only explicitly user-selected articles to IMA.
|
||
- Do not generate a digest without a real payload.
|
||
- Generate two views from the same payload: a public digest for Hugo and an internal review digest for chat/operator workflow.
|
||
- Do not expose internal review states or operator-facing labels in the public digest.
|
||
- Do not re-fetch original URLs for selected summaries; use existing extracted text.
|
||
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization, with fallback to main `LLM_*` only if needed.
|
||
- Actively report progress after each completed phase.
|
||
|
||
## Step-by-step checklist
|
||
|
||
### 1. Run reader pipeline
|
||
|
||
Use the formal MCP workflow path as the default production route.
|
||
Prefer MCP run/status/result operations over direct path stitching. Only fall back to CLI or direct file inspection for debug / manual troubleshooting.
|
||
|
||
Formal production startup sequence:
|
||
|
||
1. `start_freshrss_pipeline_job`
|
||
2. `get_freshrss_pipeline_job_status`
|
||
3. `get_freshrss_pipeline_job_result`
|
||
4. after success, continue with `run_id` via `get_run_status` / `get_delivery_payload` / `get_run_report`
|
||
|
||
Treat the old synchronous `run_freshrss_openclaw_pipeline` as debug / light validation / fallback only.
|
||
|
||
Project root:
|
||
|
||
```bash
|
||
/home/ubuntu/zhu/github/reader
|
||
```
|
||
|
||
Default behavior for a normal production run:
|
||
|
||
- if the user did not specify a count, randomly choose a limit between 5 and 10 items for that run
|
||
- run with mark-read enabled
|
||
- do not enable `debug_artifacts`
|
||
- only skip mark-read if the user explicitly says the run is debug, test, or validation
|
||
- only enable `debug_artifacts` if the user explicitly says the run is debug, test, validation, or troubleshooting
|
||
|
||
Typical artifacts to inspect after a successful run:
|
||
|
||
```text
|
||
outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json
|
||
outputs/freshrss/rerun/<run-id>/candidates/digest-brief.json
|
||
outputs/freshrss/rerun/<run-id>/run-report.json
|
||
outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json
|
||
```
|
||
|
||
Notes:
|
||
|
||
- `digest-brief.json` is the preferred input for **public digest** generation.
|
||
- It is a lighter public-only view and currently includes only `keep` candidates.
|
||
- If `digest-brief.json` is missing, fall back to `openclaw-delivery-payload.json`.
|
||
- If the synchronous MCP wrapper times out but a real reader run was still created, do not discard the run; continue from run truth using `list_runs`, `get_run_report`, and `get_delivery_payload`.
|
||
|
||
### 2. Generate digest markdown
|
||
|
||
Generate two output views from the same run, preferably in one model call:
|
||
|
||
1. a **public digest** for Hugo / public readers
|
||
2. an **internal review digest** for chat / operator workflow
|
||
|
||
Input preference:
|
||
|
||
- **public digest**: prefer `candidates/digest-brief.json`
|
||
- **internal review digest**: use `candidates/openclaw-delivery-payload.json`
|
||
|
||
Recommended generation pattern:
|
||
|
||
- Pass the public brief and the full payload as two explicitly labeled input blocks.
|
||
- Ask the model to return both outputs in one response.
|
||
- Prefer a structured response shape (for example JSON with `public_digest_markdown` and `internal_review_digest_markdown`) when post-processing is needed.
|
||
|
||
Before publishing, persist the generated digest artifacts back into the same reader run directory:
|
||
|
||
```text
|
||
outputs/freshrss/rerun/<run-id>/digest/public_digest.md
|
||
outputs/freshrss/rerun/<run-id>/digest/internal_review_digest.md
|
||
outputs/freshrss/rerun/<run-id>/digest/combined.json
|
||
```
|
||
|
||
Then write only the public digest into Hugo here:
|
||
|
||
```text
|
||
/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md
|
||
```
|
||
|
||
Recommended public-digest front matter:
|
||
|
||
```toml
|
||
+++
|
||
title = "AI 日报 · YYYY-MM-DD"
|
||
date = YYYY-MM-DDTHH:MM:SS+08:00
|
||
summary = "当日日报摘要"
|
||
+++
|
||
```
|
||
|
||
Recommended **public digest** structure:
|
||
|
||
- `今日概览`
|
||
- `今日重点`
|
||
- `趋势观察`
|
||
- `延伸阅读`
|
||
|
||
Public digest constraints:
|
||
|
||
- Use only the public brief view when present.
|
||
- Keep public wording free of internal workflow language.
|
||
- Optimize for concise public readability with solid information density.
|
||
- For each `今日重点` item, add one short editor-style sentence explaining why the item matters in today's digest (for example: `这篇内容更值得关注的原因在于……`).
|
||
- Prefer rendering highlight points as a short public label such as `值得关注:` followed by one bullet per line.
|
||
- **Article numbering: MUST use `1.` `2.` `3.` (Arabic numeral + period). Do NOT use `① ② ③`,`一、二、三`,`第一条` or any other variant.**
|
||
- **All four sections are required. Missing any one is a format violation.**
|
||
- **Hugo digest must include ALL keep articles from `digest-brief.json`. The IMA deposition subset is a separate downstream step.**
|
||
|
||
Recommended **internal review digest** structure:
|
||
|
||
- `今日候选概况`
|
||
- `已入选重点`
|
||
- `待你确认`
|
||
- `建议沉淀到 IMA`
|
||
- `原始候选清单`
|
||
|
||
Internal review digest constraints:
|
||
|
||
- Do not show `rank` values.
|
||
- Replace machine states with Chinese labels such as `已入选` / `待确认` / `暂不纳入`.
|
||
- For `已入选重点`, include a fuller summary plus a short judgment paragraph.
|
||
- For `待你确认`, include a fuller summary, reason, and recommendation.
|
||
- Keep it readable as an operator review draft, not a raw payload dump.
|
||
|
||
### 3. Publish Hugo
|
||
|
||
Publish only the public digest to Hugo.
|
||
Treat Hugo publication as the default continuation of a successful normal daily digest run. Do not ask for a second confirmation before generating/writing the public digest and publishing it, unless the user explicitly requests not to publish to Hugo.
|
||
|
||
Hard execution steps:
|
||
|
||
1. write the public digest markdown to:
|
||
- `/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md`
|
||
2. redeploy Hugo immediately after writing:
|
||
- `cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh`
|
||
3. verify all three URLs before continuing:
|
||
- `http://127.0.0.1:14322/`
|
||
- `http://127.0.0.1:14322/daily/`
|
||
- `http://127.0.0.1:14322/daily/YYYY-MM-DD/`
|
||
|
||
Expected verification targets:
|
||
|
||
- homepage works
|
||
- `/daily/` works
|
||
- `/daily/YYYY-MM-DD/` works
|
||
|
||
### 4. Report digest in chat
|
||
|
||
Provide the internal review digest in chat and ask which articles should be retained.
|
||
|
||
Hard reporting rule:
|
||
- do not send only article titles
|
||
- for each article, include at least a one-sentence summary and a short recommendation / judgment so the user can decide without reopening the source
|
||
|
||
### 5. Generate selected article summaries
|
||
|
||
Only do this after the user explicitly confirms which articles to retain.
|
||
|
||
Preferred MCP path:
|
||
|
||
- `start_article_summary_job`
|
||
- `get_article_summary_job_status`
|
||
- `get_article_summary_job_result`
|
||
|
||
Expected inputs:
|
||
|
||
- `extracted_path`
|
||
- `selected_ids`
|
||
- optional `output_dir`
|
||
- optional article-summary LLM overrides
|
||
|
||
Recommended output layout:
|
||
|
||
- `outputs/freshrss/single_summaries/YYYY-MM-DD/`
|
||
- async job state: `outputs/freshrss/article_summary_jobs/<job_id>/`
|
||
|
||
Recommended production sequence:
|
||
|
||
1. call `start_article_summary_job`
|
||
2. poll `get_article_summary_job_status` until `status` becomes `success` or `failed`
|
||
3. on success, call `get_article_summary_job_result` and continue downstream from `written_paths`
|
||
|
||
Hard fallback rule:
|
||
|
||
- If the async MCP job path returns timeout / transport failure / job-launch failure (for example MCP timeout while the reader article-summary workflow itself is still healthy), do not treat that as article-summary business failure.
|
||
- Immediately retry through the local reader environment under `/home/ubuntu/zhu/github/reader` using the project `.venv`, calling the article-summary workflow directly.
|
||
- The production goal is successful generation of the selected-article Markdown files; async MCP job is preferred, but local `.venv` execution is the required fallback path.
|
||
|
||
Synchronous helper:
|
||
|
||
- `generate_article_summaries` remains available for debug / light validation only, not as the default production path.
|
||
|
||
CLI fallback:
|
||
|
||
```bash
|
||
python scripts/run_article_summaries.py \
|
||
--extracted outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json \
|
||
--ids <item_id_1> <item_id_2> \
|
||
--output-dir outputs/freshrss/single_summaries/YYYY-MM-DD
|
||
```
|
||
|
||
### 6. Upload selected summaries to IMA
|
||
|
||
Upload only the generated markdown files for the selected articles.
|
||
Once the user has selected the articles to retain, treat that selection itself as the authorization to continue the IMA deposition step; do not ask for a second confirmation about uploading into the knowledge base.
|
||
|
||
Hard execution rules before upload:
|
||
|
||
1. reformat/check the generated markdown into IMA-facing final content
|
||
2. normalize the final upload filename to `<文章标题>.md`
|
||
3. do not use internal temp names such as `ima-*`, `item-*`, `summary-*`, or English slug filenames as the final uploaded object name
|
||
4. if the knowledge base already contains the same filename, append a timestamp suffix before `.md`
|
||
5. if local work needs internal temp names, create a final upload copy with the user-facing title before calling IMA APIs
|
||
|
||
Default target knowledge base:
|
||
|
||
- `daily`
|
||
- Read `IMA_DAILY_KNOWLEDGE_BASE_ID` / `IMA_DAILY_KNOWLEDGE_BASE_NAME` from reader `.env`
|
||
- Verify the configured target at runtime before upload
|
||
- If the configured target is unavailable, resolve by name `daily`; if still absent, create `daily`
|
||
|
||
## Hard rules recap
|
||
|
||
- Never generate a digest from placeholder or example data when a real run is expected.
|
||
- For normal runs, mark processed FreshRSS items as read unless the user explicitly requested a debug/test/validation run.
|
||
- Public digest goes to Hugo; internal review digest goes to chat; neither full digest goes to IMA.
|
||
- Only explicitly user-selected articles go to IMA.
|
||
- All daily IMA deposition must go directly into the IMA knowledge-base path as Markdown knowledge items (`media_type=7`), not through the IMA notes path.
|
||
- Uploading to IMA notes, or creating notes first and then linking them into a knowledge base, does not count as SOP completion.
|
||
- Selected article summaries use extracted text, not live refetch.
|
||
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization.
|
||
- Final IMA upload filenames must use user-facing article titles, not internal slugs or workflow temp names.
|