docs: align reader-digest-flow with async resume flow

This commit is contained in:
root
2026-04-14 16:29:56 +08:00
parent 8dd76a1f86
commit 22c7f1638e
2 changed files with 59 additions and 7 deletions
+39 -6
View File
@@ -13,8 +13,10 @@ Run the reader-based daily digest as a fixed SOP. Treat this skill as the orches
- **Daily digest production runs should default to a random article count between 5 and 10 unless the user explicitly specifies a count.**
- Prefer MCP workflow operations over direct long-lived CLI execution. Treat CLI as debug / fallback, not the default production path.
- **Main FreshRSS production runs should default to the async MCP job path**: `start_freshrss_pipeline_job` → `get_freshrss_pipeline_job_status` → `get_freshrss_pipeline_job_result`. Use the synchronous `run_freshrss_openclaw_pipeline` only for debug / light validation / fallback.
- **If the main pipeline job fails, do not immediately abandon the run.** First inspect the linked run through `get_run_status` and `inspect_resume_plan`; when the plan returns `recommended_action=resume`, continue through `start_resume_job` → `get_resume_job_status` → `get_resume_job_result`.
- **Operational note from validation:** before the main pipeline async job existed, a debug/test run could still complete successfully inside reader even when the synchronous MCP wrapper returned timeout. For historical/debug cases, continue from real run truth using MCP status/result query tools rather than treating the whole flow as failed.
- For reader run observation and result reading, prefer MCP tools such as status / payload / report queries instead of having OpenClaw or this skill hand-build reader output paths.
- **When reader status APIs return reconciled status, always branch on the top-level `status`.** Treat `status_source` and `state_conflict` as explanatory metadata; do not re-derive flow control from stale stage names.
- **For normal production runs, mark processed FreshRSS items as read. Only skip mark-read when the user explicitly says the run is debug/test/validation.**
- **For normal production runs, do not pass `debug_artifacts=true`. Only enable debug artifacts when the user explicitly says the run is debug/test/validation or when troubleshooting is the goal.**
- **Do not generate the daily digest from examples or placeholder data. Always require a real payload first.**
@@ -58,16 +60,36 @@ Minimum expected artifacts:
- `outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run-id>/candidates/digest-brief.json`(public digest 优先读取的轻量输入,仅包含 `keep` 候选)
- `outputs/freshrss/rerun/<run-id>/run-report.json`
- extracted article data, such as:
- `outputs/freshrss/extracted/freshrss.extracted.json`
- extracted article data: `outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json` 等(每个候选一篇,格式为 `{"success": true, "article": {...}}`)
If the pipeline fails, stop and report the exact failure point.
If the main pipeline job fails:
- read the linked run through `get_run_status`
- call `inspect_resume_plan(run_id)`
- if `recommended_action=resume`, continue with:
- `start_resume_job`
- `get_resume_job_status`
- `get_resume_job_result`
- if `recommended_action=read_terminal_result`, continue from the terminal run result instead of retrying
- if `recommended_action=start_new_run`, stop and report the exact failure point
### Phase 2: Generate daily digest markdown
Read the real run outputs and generate digest markdown for the day.
If there is no real payload, stop instead of writing a fake or example digest.
**⚠️ HARD PRE-WRITE CHECKLIST — 对照 `references/public-digest-example.md` 逐项确认后再开始写:**
1. [ ] frontmatter `summary` 字段已填写(一句话概括今日核心主题)
2. [ ] 编号只用 `1.` `2.` `3.`,不用中文数字或罗马数字
3. [ ] 四个 section 全部存在:`今日概览` `今日重点` `趋势观察` `延伸阅读`
4. [ ] 每篇 `今日重点` 下有:
- [ ] 标题 + 来源
- [ ] 摘要段落
- [ ] "更值得关注的原因在于:"段落
- [ ] "值得关注:"要点列表(每篇 3 条)
5. [ ] `延伸阅读` 每条含来源标注:`- [标题](url)|来源`
For **public digest**, prefer `candidates/digest-brief.json` when present; fall back to `candidates/openclaw-delivery-payload.json` only if the brief file is missing.
For **internal review digest**, continue reading the full `candidates/openclaw-delivery-payload.json`.
@@ -185,9 +207,20 @@ At this step:
For every article explicitly selected by the user:
1. use the `reader` article-summary capability
2. point it at the existing extracted payload
3. pass the selected `item_id` values
4. generate one markdown summary per article
2. **⚠️ 关键:extracted_path 必须传入 pipeline 产生的 individual item 文件,路径为:**
```
outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json
outputs/freshrss/rerun/<run-id>/extracted/item-02.extracted.json
...
```
**不要**传 `candidates/openclaw-delivery-payload.json` 或 `outputs/freshrss/extracted/`(这两个路径的 payload 格式与 article-summary workflow 不兼容,会报错**
3. for each selected article, run **one summary job per extracted item file**
4. pass the corresponding `item_id` from the `candidates` array **去掉 `cand:` 前缀**
5. generate one markdown summary per article
**⚠️ item_id 前缀注意:** `candidates` 里的 ID 格式是 `cand:sha256:xxx`,但 item 文件里的 `item_id` 是 `sha256:xxx`(无 `cand:` 前缀)。传 selected_ids 时**必须去掉前缀**,否则匹配失败。
**⚠️ single-item 输入约束:** 当 `extracted_path` 指向 `outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json` 这类单篇文件时,`selected_ids` 只应包含这一个文件对应的单个 `item_id`。不要对单个 extracted 文件传多个 IDs。
Preferred routes:
+20 -1
View File
@@ -34,6 +34,8 @@ Concrete operational checklist for the `reader-digest-flow` skill.
- Generate two views from the same payload: a public digest for Hugo and an internal review digest for chat/operator workflow.
- Do not expose internal review states or operator-facing labels in the public digest.
- Do not re-fetch original URLs for selected summaries; use existing extracted text.
- If the main pipeline fails, inspect the run first; when `inspect_resume_plan` says `recommended_action=resume`, continue via the async resume job path instead of stopping immediately.
- Always branch on reader's top-level reconciled `status`; treat `status_source` and `state_conflict` only as explanatory metadata.
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization, with fallback to main `LLM_*` only if needed.
- Actively report progress after each completed phase.
@@ -51,6 +53,17 @@ Formal production startup sequence:
3. `get_freshrss_pipeline_job_result`
4. after success, continue with `run_id` via `get_run_status` / `get_delivery_payload` / `get_run_report`
If the main pipeline job ends in `failed`:
1. inspect the linked run with `get_run_status`
2. call `inspect_resume_plan(run_id)`
3. if `recommended_action=resume`, continue with:
- `start_resume_job`
- `get_resume_job_status`
- `get_resume_job_result`
4. if `recommended_action=read_terminal_result`, continue from the terminal run result
5. if `recommended_action=start_new_run`, stop and report the failure
Treat the old synchronous `run_freshrss_openclaw_pipeline` as debug / light validation / fallback only.
Project root:
@@ -206,6 +219,12 @@ Expected inputs:
- optional `output_dir`
- optional article-summary LLM overrides
Single-item rule:
- when `extracted_path` is `outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json`, call one summary job per file
- in that case, `selected_ids` should contain only the matching single `item_id`
- if the candidate ID came from the delivery payload, strip the `cand:` prefix before passing it
Recommended output layout:
- `outputs/freshrss/single_summaries/YYYY-MM-DD/`
@@ -232,7 +251,7 @@ CLI fallback:
```bash
python scripts/run_article_summaries.py \
--extracted outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json \
--ids <item_id_1> <item_id_2> \
--ids <item_id_without_cand_prefix> \
--output-dir outputs/freshrss/single_summaries/YYYY-MM-DD
```