docs: align reader-digest-flow with async resume flow

This commit is contained in:
root
2026-04-14 16:29:56 +08:00
parent 8dd76a1f86
commit 22c7f1638e
2 changed files with 59 additions and 7 deletions
+39 -6
View File
@@ -13,8 +13,10 @@ Run the reader-based daily digest as a fixed SOP. Treat this skill as the orches
- **Daily digest production runs should default to a random article count between 5 and 10 unless the user explicitly specifies a count.** - **Daily digest production runs should default to a random article count between 5 and 10 unless the user explicitly specifies a count.**
- Prefer MCP workflow operations over direct long-lived CLI execution. Treat CLI as debug / fallback, not the default production path. - Prefer MCP workflow operations over direct long-lived CLI execution. Treat CLI as debug / fallback, not the default production path.
- **Main FreshRSS production runs should default to the async MCP job path**: `start_freshrss_pipeline_job` → `get_freshrss_pipeline_job_status` → `get_freshrss_pipeline_job_result`. Use the synchronous `run_freshrss_openclaw_pipeline` only for debug / light validation / fallback. - **Main FreshRSS production runs should default to the async MCP job path**: `start_freshrss_pipeline_job` → `get_freshrss_pipeline_job_status` → `get_freshrss_pipeline_job_result`. Use the synchronous `run_freshrss_openclaw_pipeline` only for debug / light validation / fallback.
- **If the main pipeline job fails, do not immediately abandon the run.** First inspect the linked run through `get_run_status` and `inspect_resume_plan`; when the plan returns `recommended_action=resume`, continue through `start_resume_job` → `get_resume_job_status` → `get_resume_job_result`.
- **Operational note from validation:** before the main pipeline async job existed, a debug/test run could still complete successfully inside reader even when the synchronous MCP wrapper returned timeout. For historical/debug cases, continue from real run truth using MCP status/result query tools rather than treating the whole flow as failed. - **Operational note from validation:** before the main pipeline async job existed, a debug/test run could still complete successfully inside reader even when the synchronous MCP wrapper returned timeout. For historical/debug cases, continue from real run truth using MCP status/result query tools rather than treating the whole flow as failed.
- For reader run observation and result reading, prefer MCP tools such as status / payload / report queries instead of having OpenClaw or this skill hand-build reader output paths. - For reader run observation and result reading, prefer MCP tools such as status / payload / report queries instead of having OpenClaw or this skill hand-build reader output paths.
- **When reader status APIs return reconciled status, always branch on the top-level `status`.** Treat `status_source` and `state_conflict` as explanatory metadata; do not re-derive flow control from stale stage names.
- **For normal production runs, mark processed FreshRSS items as read. Only skip mark-read when the user explicitly says the run is debug/test/validation.** - **For normal production runs, mark processed FreshRSS items as read. Only skip mark-read when the user explicitly says the run is debug/test/validation.**
- **For normal production runs, do not pass `debug_artifacts=true`. Only enable debug artifacts when the user explicitly says the run is debug/test/validation or when troubleshooting is the goal.** - **For normal production runs, do not pass `debug_artifacts=true`. Only enable debug artifacts when the user explicitly says the run is debug/test/validation or when troubleshooting is the goal.**
- **Do not generate the daily digest from examples or placeholder data. Always require a real payload first.** - **Do not generate the daily digest from examples or placeholder data. Always require a real payload first.**
@@ -58,16 +60,36 @@ Minimum expected artifacts:
- `outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json` - `outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run-id>/candidates/digest-brief.json`(public digest 优先读取的轻量输入,仅包含 `keep` 候选) - `outputs/freshrss/rerun/<run-id>/candidates/digest-brief.json`(public digest 优先读取的轻量输入,仅包含 `keep` 候选)
- `outputs/freshrss/rerun/<run-id>/run-report.json` - `outputs/freshrss/rerun/<run-id>/run-report.json`
- extracted article data, such as: - extracted article data: `outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json` 等(每个候选一篇,格式为 `{"success": true, "article": {...}}`)
- `outputs/freshrss/extracted/freshrss.extracted.json`
If the pipeline fails, stop and report the exact failure point. If the main pipeline job fails:
- read the linked run through `get_run_status`
- call `inspect_resume_plan(run_id)`
- if `recommended_action=resume`, continue with:
- `start_resume_job`
- `get_resume_job_status`
- `get_resume_job_result`
- if `recommended_action=read_terminal_result`, continue from the terminal run result instead of retrying
- if `recommended_action=start_new_run`, stop and report the exact failure point
### Phase 2: Generate daily digest markdown ### Phase 2: Generate daily digest markdown
Read the real run outputs and generate digest markdown for the day. Read the real run outputs and generate digest markdown for the day.
If there is no real payload, stop instead of writing a fake or example digest. If there is no real payload, stop instead of writing a fake or example digest.
**⚠️ HARD PRE-WRITE CHECKLIST — 对照 `references/public-digest-example.md` 逐项确认后再开始写:**
1. [ ] frontmatter `summary` 字段已填写(一句话概括今日核心主题)
2. [ ] 编号只用 `1.` `2.` `3.`,不用中文数字或罗马数字
3. [ ] 四个 section 全部存在:`今日概览` `今日重点` `趋势观察` `延伸阅读`
4. [ ] 每篇 `今日重点` 下有:
- [ ] 标题 + 来源
- [ ] 摘要段落
- [ ] "更值得关注的原因在于:"段落
- [ ] "值得关注:"要点列表(每篇 3 条)
5. [ ] `延伸阅读` 每条含来源标注:`- [标题](url)|来源`
For **public digest**, prefer `candidates/digest-brief.json` when present; fall back to `candidates/openclaw-delivery-payload.json` only if the brief file is missing. For **public digest**, prefer `candidates/digest-brief.json` when present; fall back to `candidates/openclaw-delivery-payload.json` only if the brief file is missing.
For **internal review digest**, continue reading the full `candidates/openclaw-delivery-payload.json`. For **internal review digest**, continue reading the full `candidates/openclaw-delivery-payload.json`.
@@ -185,9 +207,20 @@ At this step:
For every article explicitly selected by the user: For every article explicitly selected by the user:
1. use the `reader` article-summary capability 1. use the `reader` article-summary capability
2. point it at the existing extracted payload 2. **⚠️ 关键:extracted_path 必须传入 pipeline 产生的 individual item 文件,路径为:**
3. pass the selected `item_id` values ```
4. generate one markdown summary per article outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json
outputs/freshrss/rerun/<run-id>/extracted/item-02.extracted.json
...
```
**不要**传 `candidates/openclaw-delivery-payload.json` 或 `outputs/freshrss/extracted/`(这两个路径的 payload 格式与 article-summary workflow 不兼容,会报错**
3. for each selected article, run **one summary job per extracted item file**
4. pass the corresponding `item_id` from the `candidates` array **去掉 `cand:` 前缀**
5. generate one markdown summary per article
**⚠️ item_id 前缀注意:** `candidates` 里的 ID 格式是 `cand:sha256:xxx`,但 item 文件里的 `item_id` 是 `sha256:xxx`(无 `cand:` 前缀)。传 selected_ids 时**必须去掉前缀**,否则匹配失败。
**⚠️ single-item 输入约束:** 当 `extracted_path` 指向 `outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json` 这类单篇文件时,`selected_ids` 只应包含这一个文件对应的单个 `item_id`。不要对单个 extracted 文件传多个 IDs。
Preferred routes: Preferred routes:
+20 -1
View File
@@ -34,6 +34,8 @@ Concrete operational checklist for the `reader-digest-flow` skill.
- Generate two views from the same payload: a public digest for Hugo and an internal review digest for chat/operator workflow. - Generate two views from the same payload: a public digest for Hugo and an internal review digest for chat/operator workflow.
- Do not expose internal review states or operator-facing labels in the public digest. - Do not expose internal review states or operator-facing labels in the public digest.
- Do not re-fetch original URLs for selected summaries; use existing extracted text. - Do not re-fetch original URLs for selected summaries; use existing extracted text.
- If the main pipeline fails, inspect the run first; when `inspect_resume_plan` says `recommended_action=resume`, continue via the async resume job path instead of stopping immediately.
- Always branch on reader's top-level reconciled `status`; treat `status_source` and `state_conflict` only as explanatory metadata.
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization, with fallback to main `LLM_*` only if needed. - Prefer `ARTICLE_SUMMARY_*` for selected article summarization, with fallback to main `LLM_*` only if needed.
- Actively report progress after each completed phase. - Actively report progress after each completed phase.
@@ -51,6 +53,17 @@ Formal production startup sequence:
3. `get_freshrss_pipeline_job_result` 3. `get_freshrss_pipeline_job_result`
4. after success, continue with `run_id` via `get_run_status` / `get_delivery_payload` / `get_run_report` 4. after success, continue with `run_id` via `get_run_status` / `get_delivery_payload` / `get_run_report`
If the main pipeline job ends in `failed`:
1. inspect the linked run with `get_run_status`
2. call `inspect_resume_plan(run_id)`
3. if `recommended_action=resume`, continue with:
- `start_resume_job`
- `get_resume_job_status`
- `get_resume_job_result`
4. if `recommended_action=read_terminal_result`, continue from the terminal run result
5. if `recommended_action=start_new_run`, stop and report the failure
Treat the old synchronous `run_freshrss_openclaw_pipeline` as debug / light validation / fallback only. Treat the old synchronous `run_freshrss_openclaw_pipeline` as debug / light validation / fallback only.
Project root: Project root:
@@ -206,6 +219,12 @@ Expected inputs:
- optional `output_dir` - optional `output_dir`
- optional article-summary LLM overrides - optional article-summary LLM overrides
Single-item rule:
- when `extracted_path` is `outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json`, call one summary job per file
- in that case, `selected_ids` should contain only the matching single `item_id`
- if the candidate ID came from the delivery payload, strip the `cand:` prefix before passing it
Recommended output layout: Recommended output layout:
- `outputs/freshrss/single_summaries/YYYY-MM-DD/` - `outputs/freshrss/single_summaries/YYYY-MM-DD/`
@@ -232,7 +251,7 @@ CLI fallback:
```bash ```bash
python scripts/run_article_summaries.py \ python scripts/run_article_summaries.py \
--extracted outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json \ --extracted outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json \
--ids <item_id_1> <item_id_2> \ --ids <item_id_without_cand_prefix> \
--output-dir outputs/freshrss/single_summaries/YYYY-MM-DD --output-dir outputs/freshrss/single_summaries/YYYY-MM-DD
``` ```