refactor: simplify reader digest skill

This commit is contained in:
zhuyongxin
2026-07-28 19:14:09 +08:00
parent 5eb390e3ed
commit 6dd8cef347
14 changed files with 388 additions and 1859 deletions
+105 -283
View File
@@ -1,328 +1,150 @@
# Reader Digest Flow Reference
# Reader Digest Flow 操作参考
## Purpose
本文件承载 `reader-digest-flow` 的具体执行步骤。正式行为边界以 `../SKILL.md` 为准。
Concrete operational checklist for the `reader-digest-flow` skill.
## 1. 日报 Pipeline
## Default Operating Model
### 默认参数
### Layering
- `reader` layer:
- FreshRSS pull
- extraction
- summary/filter/payload generation
- selected-article summary capability
- OpenClaw / skill layer:
- public digest generation
- internal review digest generation
- Hugo publishing
- chat reporting
- user confirmation handling
- calling selected-article summaries
- IMA upload orchestration
- Hugo layer:
- public digest browsing and archive only
- IMA layer:
- long-term storage for selected article notes only
### Hard rules
- Do not upload the full digest to IMA.
- Upload only explicitly user-selected articles to IMA.
- Do not generate a digest without a real payload.
- Generate two views from the same payload: a public digest for Hugo and an internal review digest for chat/operator workflow.
- Do not expose internal review states or operator-facing labels in the public digest.
- Do not re-fetch original URLs for selected summaries; use existing extracted text.
- If the main pipeline fails, inspect the run first; when `inspect_resume_plan` says `recommended_action=resume`, continue via the async resume job path instead of stopping immediately.
- Always branch on reader's top-level reconciled `status`; treat `status_source` and `state_conflict` only as explanatory metadata.
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization, with fallback to main `LLM_*` only if needed.
- Actively report progress after each completed phase.
## Step-by-step checklist
### 1. Run reader pipeline
Use the formal MCP workflow path as the default production route.
Prefer MCP run/status/result operations over direct path stitching. Only fall back to CLI or direct file inspection for debug / manual troubleshooting.
Formal production startup sequence:
1. `start_freshrss_pipeline_job`
2. `get_freshrss_pipeline_job_status`
3. `get_freshrss_pipeline_job_result`
4. after success, continue with `run_id` via `get_run_status` / `get_delivery_payload` / `get_run_report`
If the main pipeline job ends in `failed`:
1. inspect the linked run with `get_run_status`
2. call `inspect_resume_plan(run_id)`
3. if `recommended_action=resume`, continue with:
- `start_resume_job`
- `get_resume_job_status`
- `get_resume_job_result`
4. if `recommended_action=read_terminal_result`, continue from the terminal run result
5. if `recommended_action=start_new_run`, stop and report the failure
Treat the old synchronous `run_freshrss_openclaw_pipeline` as debug / light validation / fallback only.
Project root:
```bash
/home/ubuntu/zhu/github/reader
```json
{
"limit": 7,
"include_read": false,
"mark_read": true,
"debug_artifacts": false,
"timeout_seconds": 60,
"max_retries": 2
}
```
Default behavior for a normal production run:
- if the user did not specify a count, randomly choose a limit between 5 and 10 items for that run
- run with mark-read enabled
- do not enable `debug_artifacts`
- only skip mark-read if the user explicitly says the run is debug, test, or validation
- only enable `debug_artifacts` if the user explicitly says the run is debug, test, validation, or troubleshooting
Typical artifacts to inspect after a successful run:
### 正式调用序列
```text
outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json
outputs/freshrss/rerun/<run-id>/candidates/digest-brief.json
outputs/freshrss/rerun/<run-id>/run-report.json
outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json
start_freshrss_pipeline_job
→ get_freshrss_pipeline_job_status
→ get_freshrss_pipeline_job_result
→ get_run_status
→ get_delivery_payload / get_run_report
```
Notes:
状态动作:
- `digest-brief.json` is the preferred input for **public digest** generation.
- It is a lighter public-only view and currently includes only `keep` candidates.
- If `digest-brief.json` is missing, fall back to `openclaw-delivery-payload.json`.
- If the synchronous MCP wrapper times out but a real reader run was still created, do not discard the run; continue from run truth using `list_runs`, `get_run_report`, and `get_delivery_payload`.
- `running`:按合理间隔继续轮询;
- `success`:读取结果,保存 `run_id`;
- `failed`:读取关联 Run 并执行 Resume Plan;
- `partial`:优先检查 Run Report、Artifact 和 Recovery 信息。
**⚠️ Output path divergence when run from Feishu context:**
When triggered via MCP from inside a Feishu session, the extracted files may be written to an MCP-managed temp path instead of `outputs/freshrss/rerun/<run-id>/extracted/`. Verify the actual output path from the pipeline result before continuing to Phase 5/6. If the project-relative path doesn't exist, use the MCP result's `extracted_path` directly or copy the files to the expected project location.
### 2. Generate digest markdown
Generate two output views from the same run, preferably in one model call:
1. a **public digest** for Hugo / public readers
2. an **internal review digest** for chat / operator workflow
Input preference:
- **public digest**: prefer `candidates/digest-brief.json`
- **internal review digest**: use `candidates/openclaw-delivery-payload.json`
Recommended generation pattern:
- Pass the public brief and the full payload as two explicitly labeled input blocks.
- Ask the model to return both outputs in one response.
- Prefer a structured response shape (for example JSON with `public_digest_markdown` and `internal_review_digest_markdown`) when post-processing is needed.
Before publishing, persist the generated digest artifacts back into the same reader run directory:
恢复序列:
```text
outputs/freshrss/rerun/<run-id>/digest/public_digest.md
outputs/freshrss/rerun/<run-id>/digest/internal_review_digest.md
outputs/freshrss/rerun/<run-id>/digest/combined.json
inspect_resume_plan
→ recommended_action=resume
→ start_resume_job
→ get_resume_job_status
→ get_resume_job_result
```
Then write only the public digest into Hugo here:
如果建议为 `read_terminal_result`,直接读取结果;如果为 `start_new_run`,停止并等待用户决定。
不要根据 `run_id` 拼目录。后续只消费调用返回的 `output_dir`、`artifact.path`、`delivery_output`、`report_output`。
## 2. 候选汇报与选择
优先使用当前 Run 的 digest brief Artifact;不可用时使用 `get_delivery_payload` 返回的候选。
展示规则:
1. 使用候选数组原始顺序并从 1 编号;
2. 不因 `keep/review` 分组而重新编号;
3. 每篇展示标题、来源、摘要和判断理由;
4. 用户编号只映射当前候选数组;
5. 跨 Artifact 读取详情时按 URL 或完整 `item_id` 匹配。
不要按 `summary-batch`、`item-XX` 和候选数组的相同位置推断它们是同一篇文章。
## 3. Hugo 日报
用户确认发布文章后:
1. 读取 `public-digest-example.md`;
2. 仅使用用户选中的文章生成公开内容;
3. 写入 Hugo 当日页面;
4. 前台执行部署,避免把构建日志作为聊天通知;
5. 验证首页、列表页和详情页。
当前部署位置:
```text
/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md
```
Recommended public-digest front matter:
验证至少覆盖:
```toml
+++
title = "AI 日报 · YYYY-MM-DD"
date = YYYY-MM-DDTHH:MM:SS+08:00
summary = "当日日报摘要"
+++
```text
http://127.0.0.1:14322/
http://127.0.0.1:14322/daily/
http://127.0.0.1:14322/daily/YYYY-MM-DD/
```
Recommended **public digest** structure:
页面不存在时检查源文件、构建结果、容器状态和页面日期。需要 Docker/Hugo 深度排障时再读取部署项目自己的文档,不把历史事故规则复制回主 Skill。
- `今日概览`
- `今日重点`
- `趋势观察`
- `延伸阅读`
## 4. 单篇知识笔记
Public digest constraints:
用户确认 IMA 文章后:
- Use only the public brief view when present.
- Keep public wording free of internal workflow language.
- Optimize for concise public readability with solid information density.
- For each `今日重点` item, add one short editor-style sentence explaining why the item matters in today's digest (for example: `这篇内容更值得关注的原因在于……`).
- Prefer rendering highlight points as a short public label such as `值得关注:` followed by one bullet per line.
- **Article numbering: MUST use `1.` `2.` `3.` (Arabic numeral + period). Do NOT use `① ② ③`,`一、二、三`,`第一条` or any other variant.**
- **All four sections are required. Missing any one is a format violation.**
- **Hugo digest must include ALL keep articles from `digest-brief.json`. The IMA deposition subset is a separate downstream step.**
1. 从候选中取得 URL 和完整 `item_id`;
2. 从 Run Artifact 中找到匹配的 extracted 文件;
3. 交叉验证 `article.item_id` 或 URL;
4. 对每个单篇文件调用 `start_article_summary_job(extracted_path=<返回的 Artifact 路径>, selected_ids=[<完整 item_id>])`;
5. `selected_ids` 传完整、非空、去除 `cand:` 前缀后的 `item_id`;
6. 从 Job 结果的 `written_paths` 读取 Markdown。
Recommended **internal review digest** structure:
正式序列:
- `今日候选概况`
- `已入选重点`
- `待你确认`
- `建议沉淀到 IMA`
- `原始候选清单`
Internal review digest constraints:
- Do not show `rank` values.
- Replace machine states with Chinese labels such as `已入选` / `待确认` / `暂不纳入`.
- For `已入选重点`, include a fuller summary plus a short judgment paragraph.
- For `待你确认`, include a fuller summary, reason, and recommendation.
- Keep it readable as an operator review draft, not a raw payload dump.
- **Feishu compatibility: never use markdown tables in the digest report.**
When delivered via Feishu, a single table forces the whole message to plain
text. Use lists and sections instead. See `references/feishu-format-notes.md`.
### 3. Publish Hugo
Publish only the public digest to Hugo.
Treat Hugo publication as the default continuation of a successful normal daily digest run. Do not ask for a second confirmation before generating/writing the public digest and publishing it, unless the user explicitly requests not to publish to Hugo.
**⚠️ Hugo redeploy: always use foreground terminal, never background+notify_on_complete.**
The Hugo redeploy runs via `./redeploy.sh` and completes in ~10 seconds. **Do NOT use `terminal(background=true, notify_on_complete=true)`** for this step — the Gateway will push the raw Docker build output (compiler logs, layered build output, nginx config) to Feishu as a notification. This output is machine-readable, not human-readable, and the user has explicitly said this is noise.
**Correct approach:**
```python
# ✅ Foreground terminal with adequate timeout
cmd = "cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh"
result = terminal(cmd, timeout=120)
# Then verify 200 OK on the detail page
```text
start_article_summary_job
→ get_article_summary_job_status
→ get_article_summary_job_result
```
**Wrong approach (DO NOT use):**
内容依据是 extracted Artifact 中的 `article.plain_text`。不得重新抓取原 URL,不得使用原文之外的知识扩写。
```python
# ❌ Background + notify sends raw build logs to Feishu
terminal("cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh", background=True, notify_on_complete=True)
```
This rule only applies to Hugo redeploy (fast, deterministic output). For long-running Reader pipeline jobs (>60s), background + notify_on_complete is still appropriate since the output is meaningful content (article summaries, pipeline stats).
1. **Pre-check: verify Hugo content directory exists**
```bash
ls -d /home/ubuntu/zhu/apps/hugo-site/content/daily/ 2>/dev/null || mkdir -p /home/ubuntu/zhu/apps/hugo-site/content/daily/
```
If the entire `content/` directory is missing, create it before proceeding.
**Do not assume the Hugo site has a standard structure.**
2. write the public digest markdown to:
- `/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md`
3. redeploy Hugo immediately after writing:
- `cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh`
4. verify all three URLs before continuing:
- `http://127.0.0.1:14322/`
- `http://127.0.0.1:14322/daily/`
- `http://127.0.0.1:14322/daily/YYYY-MM-DD/`
Expected verification targets:
- homepage works
- `/daily/` works
- `/daily/YYYY-MM-DD/` works
### 4. Report digest in chat
Provide the internal review digest in chat and ask which articles should be retained.
Hard reporting rule:
- do not send only article titles
- for each article, include at least a one-sentence summary and a short recommendation / judgment so the user can decide without reopening the source
- **Feishu**: avoid markdown tables entirely. Use lists with headings.
### 5. Generate selected article summaries
Only do this after the user explicitly confirms which articles to retain.
Preferred MCP path:
- `start_article_summary_job`
- `get_article_summary_job_status`
- `get_article_summary_job_result`
Expected inputs:
- `extracted_path`
- `selected_ids`
**⚠️ item_id 前缀注意:**
The `candidates` array IDs use `cand:sha256:xxx` format, but individual extracted item files use `sha256:xxx` (no `cand:` prefix). When passing `selected_ids` to article-summary tools, strip the `cand:` prefix. If the ID doesn't match, the summary tool won't find the article.
**Path resolution for article-summary:**
After a pipeline run via MCP (Feishu context), the extracted files may live at an MCP-managed temp path, not the expected `outputs/freshrss/rerun/<run-id>/extracted/`. Before calling `generate_article_summaries` or `start_article_summary_job`, verify the extracted path exists. If not, use the `run_id` from the pipeline result to locate the actual output directory through `get_run_report`, or copy the files from the MCP-managed path.
Single-item rule:
- when `extracted_path` is `outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json`, call one summary job per file
- in that case, `selected_ids` should contain only the matching single `item_id`
- if the candidate ID came from the delivery payload, strip the `cand:` prefix before passing it
Recommended output layout:
- `outputs/freshrss/single_summaries/YYYY-MM-DD/`
- async job state: `outputs/freshrss/article_summary_jobs/<job_id>/`
Recommended production sequence:
1. call `start_article_summary_job`
2. poll `get_article_summary_job_status` until `status` becomes `success` or `failed`
3. on success, call `get_article_summary_job_result` and continue downstream from `written_paths`
Hard fallback rule:
- If the async MCP job path returns timeout / transport failure / job-launch failure (for example MCP timeout while the reader article-summary workflow itself is still healthy), do not treat that as article-summary business failure.
- Immediately retry through the local reader environment under `/home/ubuntu/zhu/github/reader` using the project `.venv`, calling the article-summary workflow directly.
- The production goal is successful generation of the selected-article Markdown files; async MCP job is preferred, but local `.venv` execution is the required fallback path.
Synchronous helper:
- `generate_article_summaries` remains available for debug / light validation only, not as the default production path.
CLI fallback:
CLI 仅在异步 MCP 不可用或人工排障时使用:
```bash
python scripts/run_article_summaries.py \
--extracted outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json \
--ids <item_id_without_cand_prefix> \
--output-dir outputs/freshrss/single_summaries/YYYY-MM-DD
--extracted <returned-extracted-path> \
--ids <full-item-id> \
--output-dir <explicit-output-dir>
```
### 6. Upload selected summaries to IMA
## 5. IMA 上传
Upload only the generated markdown files for the selected articles.
Once the user has selected the articles to retain, treat that selection itself as the authorization to continue the IMA deposition step; do not ask for a second confirmation about uploading into the knowledge base.
用户在知识沉淀阶段的文章选择即为上传授权。
**⚠️ IMA upload 500 retry: Always wrap IMA uploads in a retry loop.**
IMA's OpenAPI may return HTTP 500 on the first attempt. If the first upload fails, wait ~3 seconds and retry. Normally the second attempt succeeds. If `cos-upload.cjs` is used (for CDN-backed uploads), check subprocess stderr even when exit code is 0 — an HTTP error in the upload service may still produce exit code 0.
执行顺序:
Hard execution rules before upload:
1. 检查生成的 Markdown 与来源;
2. 文件名规范化为 `<完整文章标题>.md`;
3. 确认目标为 `daily` knowledge base;
4. 执行 preflight、create_media、COS upload、add_knowledge;
5. 验证知识库条目存在。
1. reformat/check the generated markdown into IMA-facing final content
2. normalize the final upload filename to `<文章标题>.md`
3. do not use internal temp names such as `ima-*`, `item-*`, `summary-*`, or English slug filenames as the final uploaded object name
4. if the knowledge base already contains the same filename, append a timestamp suffix before `.md`
5. if local work needs internal temp names, create a final upload copy with the user-facing title before calling IMA APIs
上传格式与 API 参数分别见:
Default target knowledge base:
- `ima-format-quickref.md`
- `ima-upload-api.md`
- `ima-credential-chain.md`
- `daily`
- Read `IMA_DAILY_KNOWLEDGE_BASE_ID` / `IMA_DAILY_KNOWLEDGE_BASE_NAME` from reader `.env`
- Verify the configured target at runtime before upload
- If the configured target is unavailable, resolve by name `daily`; if still absent, create `daily`
失败时记录发生在 preflight、create_media、COS upload 或 add_knowledge 的具体阶段。不要把失败的 Markdown 改写成 URL 导入或 Notes 类型来绕过错误。
## Hard rules recap
## 6. CLI fallback 原则
- Never generate a digest from placeholder or example data when a real run is expected.
- For normal runs, mark processed FreshRSS items as read unless the user explicitly requested a debug/test/validation run.
- Public digest goes to Hugo; internal review digest goes to chat; neither full digest goes to IMA.
- Only explicitly user-selected articles go to IMA.
- All daily IMA deposition must go directly into the IMA knowledge-base path as Markdown knowledge items (`media_type=7`), not through the IMA notes path.
- Uploading to IMA notes, or creating notes first and then linking them into a knowledge base, does not count as SOP completion.
- Selected article summaries use extracted text, not live refetch.
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization.
- Final IMA upload filenames must use user-facing article titles, not internal slugs or workflow temp names.
CLI 仅在以下场景使用:
- MCP 服务不可用;
- Tool transport/launch 失败且无法取得有效 Job;
- 用户明确要求本地调试;
- 人工排障需要直接检查脚本输出。
如果已取得 `job_id` 或 `run_id`,先查询真实状态,避免因响应超时重复启动任务。