Compare commits

..
2 Commits
Author SHA1 Message Date
root 1be0c4a8b1 feat: add keyword maintenance skill 2026-04-14 17:09:40 +08:00
root 22c7f1638e docs: align reader-digest-flow with async resume flow 2026-04-14 16:29:56 +08:00
6 changed files with 381 additions and 7 deletions
+39 -6
View File
@@ -13,8 +13,10 @@ Run the reader-based daily digest as a fixed SOP. Treat this skill as the orches
- **Daily digest production runs should default to a random article count between 5 and 10 unless the user explicitly specifies a count.** - **Daily digest production runs should default to a random article count between 5 and 10 unless the user explicitly specifies a count.**
- Prefer MCP workflow operations over direct long-lived CLI execution. Treat CLI as debug / fallback, not the default production path. - Prefer MCP workflow operations over direct long-lived CLI execution. Treat CLI as debug / fallback, not the default production path.
- **Main FreshRSS production runs should default to the async MCP job path**: `start_freshrss_pipeline_job` → `get_freshrss_pipeline_job_status` → `get_freshrss_pipeline_job_result`. Use the synchronous `run_freshrss_openclaw_pipeline` only for debug / light validation / fallback. - **Main FreshRSS production runs should default to the async MCP job path**: `start_freshrss_pipeline_job` → `get_freshrss_pipeline_job_status` → `get_freshrss_pipeline_job_result`. Use the synchronous `run_freshrss_openclaw_pipeline` only for debug / light validation / fallback.
- **If the main pipeline job fails, do not immediately abandon the run.** First inspect the linked run through `get_run_status` and `inspect_resume_plan`; when the plan returns `recommended_action=resume`, continue through `start_resume_job` → `get_resume_job_status` → `get_resume_job_result`.
- **Operational note from validation:** before the main pipeline async job existed, a debug/test run could still complete successfully inside reader even when the synchronous MCP wrapper returned timeout. For historical/debug cases, continue from real run truth using MCP status/result query tools rather than treating the whole flow as failed. - **Operational note from validation:** before the main pipeline async job existed, a debug/test run could still complete successfully inside reader even when the synchronous MCP wrapper returned timeout. For historical/debug cases, continue from real run truth using MCP status/result query tools rather than treating the whole flow as failed.
- For reader run observation and result reading, prefer MCP tools such as status / payload / report queries instead of having OpenClaw or this skill hand-build reader output paths. - For reader run observation and result reading, prefer MCP tools such as status / payload / report queries instead of having OpenClaw or this skill hand-build reader output paths.
- **When reader status APIs return reconciled status, always branch on the top-level `status`.** Treat `status_source` and `state_conflict` as explanatory metadata; do not re-derive flow control from stale stage names.
- **For normal production runs, mark processed FreshRSS items as read. Only skip mark-read when the user explicitly says the run is debug/test/validation.** - **For normal production runs, mark processed FreshRSS items as read. Only skip mark-read when the user explicitly says the run is debug/test/validation.**
- **For normal production runs, do not pass `debug_artifacts=true`. Only enable debug artifacts when the user explicitly says the run is debug/test/validation or when troubleshooting is the goal.** - **For normal production runs, do not pass `debug_artifacts=true`. Only enable debug artifacts when the user explicitly says the run is debug/test/validation or when troubleshooting is the goal.**
- **Do not generate the daily digest from examples or placeholder data. Always require a real payload first.** - **Do not generate the daily digest from examples or placeholder data. Always require a real payload first.**
@@ -58,16 +60,36 @@ Minimum expected artifacts:
- `outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json` - `outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run-id>/candidates/digest-brief.json`(public digest 优先读取的轻量输入,仅包含 `keep` 候选) - `outputs/freshrss/rerun/<run-id>/candidates/digest-brief.json`(public digest 优先读取的轻量输入,仅包含 `keep` 候选)
- `outputs/freshrss/rerun/<run-id>/run-report.json` - `outputs/freshrss/rerun/<run-id>/run-report.json`
- extracted article data, such as: - extracted article data: `outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json` 等(每个候选一篇,格式为 `{"success": true, "article": {...}}`)
- `outputs/freshrss/extracted/freshrss.extracted.json`
If the pipeline fails, stop and report the exact failure point. If the main pipeline job fails:
- read the linked run through `get_run_status`
- call `inspect_resume_plan(run_id)`
- if `recommended_action=resume`, continue with:
- `start_resume_job`
- `get_resume_job_status`
- `get_resume_job_result`
- if `recommended_action=read_terminal_result`, continue from the terminal run result instead of retrying
- if `recommended_action=start_new_run`, stop and report the exact failure point
### Phase 2: Generate daily digest markdown ### Phase 2: Generate daily digest markdown
Read the real run outputs and generate digest markdown for the day. Read the real run outputs and generate digest markdown for the day.
If there is no real payload, stop instead of writing a fake or example digest. If there is no real payload, stop instead of writing a fake or example digest.
**⚠️ HARD PRE-WRITE CHECKLIST — 对照 `references/public-digest-example.md` 逐项确认后再开始写:**
1. [ ] frontmatter `summary` 字段已填写(一句话概括今日核心主题)
2. [ ] 编号只用 `1.` `2.` `3.`,不用中文数字或罗马数字
3. [ ] 四个 section 全部存在:`今日概览` `今日重点` `趋势观察` `延伸阅读`
4. [ ] 每篇 `今日重点` 下有:
- [ ] 标题 + 来源
- [ ] 摘要段落
- [ ] "更值得关注的原因在于:"段落
- [ ] "值得关注:"要点列表(每篇 3 条)
5. [ ] `延伸阅读` 每条含来源标注:`- [标题](url)|来源`
For **public digest**, prefer `candidates/digest-brief.json` when present; fall back to `candidates/openclaw-delivery-payload.json` only if the brief file is missing. For **public digest**, prefer `candidates/digest-brief.json` when present; fall back to `candidates/openclaw-delivery-payload.json` only if the brief file is missing.
For **internal review digest**, continue reading the full `candidates/openclaw-delivery-payload.json`. For **internal review digest**, continue reading the full `candidates/openclaw-delivery-payload.json`.
@@ -185,9 +207,20 @@ At this step:
For every article explicitly selected by the user: For every article explicitly selected by the user:
1. use the `reader` article-summary capability 1. use the `reader` article-summary capability
2. point it at the existing extracted payload 2. **⚠️ 关键:extracted_path 必须传入 pipeline 产生的 individual item 文件,路径为:**
3. pass the selected `item_id` values ```
4. generate one markdown summary per article outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json
outputs/freshrss/rerun/<run-id>/extracted/item-02.extracted.json
...
```
**不要**传 `candidates/openclaw-delivery-payload.json` 或 `outputs/freshrss/extracted/`(这两个路径的 payload 格式与 article-summary workflow 不兼容,会报错**
3. for each selected article, run **one summary job per extracted item file**
4. pass the corresponding `item_id` from the `candidates` array **去掉 `cand:` 前缀**
5. generate one markdown summary per article
**⚠️ item_id 前缀注意:** `candidates` 里的 ID 格式是 `cand:sha256:xxx`,但 item 文件里的 `item_id` 是 `sha256:xxx`(无 `cand:` 前缀)。传 selected_ids 时**必须去掉前缀**,否则匹配失败。
**⚠️ single-item 输入约束:** 当 `extracted_path` 指向 `outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json` 这类单篇文件时,`selected_ids` 只应包含这一个文件对应的单个 `item_id`。不要对单个 extracted 文件传多个 IDs。
Preferred routes: Preferred routes:
+20 -1
View File
@@ -34,6 +34,8 @@ Concrete operational checklist for the `reader-digest-flow` skill.
- Generate two views from the same payload: a public digest for Hugo and an internal review digest for chat/operator workflow. - Generate two views from the same payload: a public digest for Hugo and an internal review digest for chat/operator workflow.
- Do not expose internal review states or operator-facing labels in the public digest. - Do not expose internal review states or operator-facing labels in the public digest.
- Do not re-fetch original URLs for selected summaries; use existing extracted text. - Do not re-fetch original URLs for selected summaries; use existing extracted text.
- If the main pipeline fails, inspect the run first; when `inspect_resume_plan` says `recommended_action=resume`, continue via the async resume job path instead of stopping immediately.
- Always branch on reader's top-level reconciled `status`; treat `status_source` and `state_conflict` only as explanatory metadata.
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization, with fallback to main `LLM_*` only if needed. - Prefer `ARTICLE_SUMMARY_*` for selected article summarization, with fallback to main `LLM_*` only if needed.
- Actively report progress after each completed phase. - Actively report progress after each completed phase.
@@ -51,6 +53,17 @@ Formal production startup sequence:
3. `get_freshrss_pipeline_job_result` 3. `get_freshrss_pipeline_job_result`
4. after success, continue with `run_id` via `get_run_status` / `get_delivery_payload` / `get_run_report` 4. after success, continue with `run_id` via `get_run_status` / `get_delivery_payload` / `get_run_report`
If the main pipeline job ends in `failed`:
1. inspect the linked run with `get_run_status`
2. call `inspect_resume_plan(run_id)`
3. if `recommended_action=resume`, continue with:
- `start_resume_job`
- `get_resume_job_status`
- `get_resume_job_result`
4. if `recommended_action=read_terminal_result`, continue from the terminal run result
5. if `recommended_action=start_new_run`, stop and report the failure
Treat the old synchronous `run_freshrss_openclaw_pipeline` as debug / light validation / fallback only. Treat the old synchronous `run_freshrss_openclaw_pipeline` as debug / light validation / fallback only.
Project root: Project root:
@@ -206,6 +219,12 @@ Expected inputs:
- optional `output_dir` - optional `output_dir`
- optional article-summary LLM overrides - optional article-summary LLM overrides
Single-item rule:
- when `extracted_path` is `outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json`, call one summary job per file
- in that case, `selected_ids` should contain only the matching single `item_id`
- if the candidate ID came from the delivery payload, strip the `cand:` prefix before passing it
Recommended output layout: Recommended output layout:
- `outputs/freshrss/single_summaries/YYYY-MM-DD/` - `outputs/freshrss/single_summaries/YYYY-MM-DD/`
@@ -232,7 +251,7 @@ CLI fallback:
```bash ```bash
python scripts/run_article_summaries.py \ python scripts/run_article_summaries.py \
--extracted outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json \ --extracted outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json \
--ids <item_id_1> <item_id_2> \ --ids <item_id_without_cand_prefix> \
--output-dir outputs/freshrss/single_summaries/YYYY-MM-DD --output-dir outputs/freshrss/single_summaries/YYYY-MM-DD
``` ```
+128
View File
@@ -0,0 +1,128 @@
---
name: reader-keyword-maintenance
description: 维护 reader 项目的关键词 review 目录与低频清理 SOP。当用户要求清理或整理 reader 的关键词 review 目录、控制哪些产物长期保留、检查 suggestions JSON 是否可 apply、或执行“只保留当前 bundle / 仅保留最近少量 suggestions JSON / 删除旧 markdown”等低频维护动作时使用。不要用于正式关键词 review 生成,不要用于 reader 日报主链路。
---
# Reader Keyword Maintenance
这个 skill 负责 reader 关键词治理里的**低频维护动作**,目标是把清理策略、展示稿生成、dry-run 校验和 review 目录维护从 reader 主链路里拆出来,由 OpenClaw 单独承接。
## 角色边界
这个 skill 负责:
- 检查 `outputs/term_index/review/` 当前有哪些产物
- 判断哪些 review 产物应长期保留、短期保留或可清理
- 用 `apply_term_suggestions.py --dry-run` 做 apply 前校验
- 执行低频清理和收尾动作
这个 skill 不负责:
- 正式生成关键词 review 建议
- 充当 alias / stopword 正式 review 决策器
- 替代 reader 侧的 `keyword-cleanup-review`
如果用户要的是“做一轮正式关键词 review / 生成正式 suggestions JSON”,应由 OpenClaw 去调用 reader 侧 `keyword-cleanup-review`,而不是由本 skill 直接承担。
## 适用场景
当用户要求以下事情时使用:
- 查看或解释 `reader/outputs/term_index/review/` 下有哪些文件
- 清理旧的 review bundle、old suggestions、old markdown
- 只保留当前 bundle 或最近少量 suggestions JSON
- 检查 suggestions JSON 是否可被 `apply_term_suggestions.py` 消费
- 执行一次人工 review 准备动作,但不改变 reader 主链路设计
不要用于:
- 日报正式生产运行
- 生成 daily term index / term_stats 主链路
- 自动写配置作为默认行为
- 修改 reader 仓库里的主 SOP 作为低频维护动作的默认入口
- 正式生成 alias / stopword / interest / watch 建议
## 默认口径
长期保留:
- `data/term_index/daily/*.json`
- `data/term_index/term_stats.json`
- `configs/filter_context.personal.json`
- `configs/term_watchlist.json`
- `configs/term_aliases.json`
- `configs/term_stopwords.json`
- `configs/term_change_log.json`
正式建议产物(短期保留):
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`
临时工作文件 / 展示层:
- `outputs/term_index/review/keyword-cleanup-bundle.json`
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`
更细的保留/清理判断,读取:
- `references/cleanup-policy.md`
## 固定 SOP
### 1. 检查 review 目录
- 查看 `outputs/term_index/review/` 当前有哪些文件
- 分类成:
- 长期保留
- 短期保留
- 临时工作文件
### 2. 准备人工审阅
- 确认最新 bundle 存在
- 确认最新 suggestions JSON 存在
- 如需正式生成 suggestions 或 Markdown 展示稿,应转到 reader 侧 `keyword-cleanup-review`
### 3. dry-run 校验 apply
- 仅通过 `scripts/apply_term_suggestions.py --dry-run` 校验 suggestions JSON 是否可消费
- 不直接落配置
### 4. 清理 review 产物
- 删除旧 markdown 展示稿
- 默认只保留当前/latest bundle
- 保守清理旧 suggestions JSON
- 不碰 facts/state/config files
## 常用动作
### 1. 验证 suggestions JSON 能否被 apply 消费
```bash
cd /home/ubuntu/zhu/github/reader && \
/usr/bin/python3.11 scripts/apply_term_suggestions.py \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--accept-interest "Claude Code" \
--dry-run
```
## 清理策略
保守策略:
- bundle:默认只保留当前最新一份
- markdown:默认不归档,需要时现生成
- suggestions JSON:只保留最近少量几份或已应用过的记录
## 维护原则
- 优先减少中间产物,不要让 review 工作文件变成长期资产
- JSON suggestions 是 review / apply 之间唯一正式输入
- Markdown 只是展示层,不是系统真相
- 任何清理动作前,先确认不会影响当前待审阅或待 apply 的 suggestions JSON
- 正式关键词 review 的生成入口不在这里,而在 reader 侧 `keyword-cleanup-review`
执行清理或 review-prep 前,先读取:
- `references/maintenance-checklist.md`
@@ -0,0 +1,33 @@
# Alias Review Example
Use this as the default style reference when preparing a low-frequency alias review report.
## Suggested merges
- `Claude code -> Claude Code`
- 理由:明显属于同一产品名,仅是大小写写法不一致。
- 证据:`Claude Code` 在最近多日持续出现,而小写写法只是在少量上下文中作为变体出现。
- `Sub-Agent -> SubAgent`
- 理由:更像词形差异,不构成新的独立概念。
- 证据:两者都围绕同一 agent 架构语境出现,且没有稳定区分语义。
## Not recommended to merge
- `Skills ↔ Agent Skills`
- 理由:前者过泛,后者更具体,当前强行归并会损失粒度。
- `Anthropic ↔ Claude`
- 理由:公司名与产品名并不等价,不应直接视为一个关键词。
## Needs human judgment
- `AI助手 ↔ AI Agent`
- 风险点:语义可能接近,但中文表述范围更宽,是否并入需要结合实际使用语境判断。
## Style notes
- Keep the report concise.
- Give judgment first, then evidence.
- Do not claim anything is already applied.
- Prefer conservative proposals over broad semantic merging.
@@ -0,0 +1,106 @@
# Cleanup Policy
## Purpose
This reference defines how `reader-keyword-maintenance` should treat keyword governance artifacts.
The goal is to keep long-term assets small and stable while allowing review-time working files to exist when needed.
This policy is for low-frequency maintenance only.
It does not replace the formal keyword review generation flow in reader.
## Artifact classes
### Long-term assets
Keep these by default:
- `data/term_index/daily/*.json`
- `data/term_index/term_stats.json`
- `configs/filter_context.personal.json`
- `configs/term_watchlist.json`
- `configs/term_aliases.json`
- `configs/term_stopwords.json`
- `configs/term_change_log.json`
These are facts or active state.
### Short-term decision artifacts
Keep selectively:
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`
Recommended policy:
- keep only recent few files, or
- keep only files that were actually used for apply decisions
### Temporary working/display files
Treat as disposable unless the user explicitly asks to archive them:
- `outputs/term_index/review/keyword-cleanup-bundle.json`
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`
Recommended policy:
- bundle: keep only latest current file
- markdown: generate on demand, do not archive by default
## Default maintenance actions
### Safe checks
Before deleting anything:
1. confirm the target is not the current bundle under active review
2. confirm the target suggestions JSON is not the one about to be applied
3. never delete configs or facts during review cleanup
### Safe cleanup order
1. remove or overwrite old markdown display drafts
2. keep only latest bundle file
3. prune old suggestions JSON files conservatively
## Decision rules
### When user says "clean review artifacts"
Default action:
- keep facts/state untouched
- keep current suggestions JSON
- remove markdown drafts if they are old and clearly derived display files
- keep bundle only as current working file
### When user says "prepare manual review"
Default action:
- ensure latest bundle exists
- ensure latest suggestions JSON exists
- do not generate formal suggestions or markdown here by default
- if the user wants formal review material, route to reader-side `keyword-cleanup-review`
### When user says "review alias candidates" or "review stopword candidates"
Default action:
- explain that formal keyword review generation belongs to reader-side `keyword-cleanup-review`
- keep this skill focused on maintenance, cleanup, display generation, and dry-run validation
- only proceed here if the user explicitly wants a low-frequency maintenance view rather than the formal review flow
### When user says "what can be deleted"
Explain in three buckets:
- must keep
- can keep temporarily
- safe to regenerate/delete
## Non-goals
This policy does not change reader production logic.
It only governs low-frequency maintenance and cleanup decisions in OpenClaw.
@@ -0,0 +1,55 @@
# Maintenance Checklist
Use this checklist before doing any cleanup or review-maintenance action for reader keyword artifacts.
## Pre-check
1. Confirm the current repo root is `/home/ubuntu/zhu/github/reader`
2. Confirm the user asked for a maintenance / cleanup / review-prep action
3. If the user actually wants formal keyword review generation, route to reader-side `keyword-cleanup-review` instead of using this maintenance skill
4. Identify whether the action targets:
- current bundle
- suggestions JSON
- markdown display draft
- old review outputs
## Safety check
Before deleting or overwriting anything:
1. Do not touch:
- `data/term_index/daily/*.json`
- `data/term_index/term_stats.json`
- `configs/filter_context.personal.json`
- `configs/term_watchlist.json`
- `configs/term_aliases.json`
- `configs/term_stopwords.json`
- `configs/term_change_log.json`
2. Confirm the suggestions JSON to keep is not the one about to be applied
3. Treat markdown drafts as disposable only after confirming they are display-only artifacts
## Review-prep flow
When preparing manual review:
1. Ensure the latest bundle exists
2. Ensure the latest suggestions JSON exists
3. Generate markdown only if the user explicitly wants human-readable review material
4. Prefer showing conclusions in chat before creating more files
## Cleanup flow
Recommended order:
1. remove old markdown display drafts
2. keep only the current/latest bundle
3. prune old suggestions JSON conservatively
4. leave facts/state/config untouched
## Post-check
After the action:
1. verify the expected kept files still exist
2. verify no config file was accidentally changed
3. summarize what was kept vs removed