refactor: simplify reader digest skill
This commit is contained in:
@@ -1,328 +1,150 @@
|
||||
# Reader Digest Flow Reference
|
||||
# Reader Digest Flow 操作参考
|
||||
|
||||
## Purpose
|
||||
本文件承载 `reader-digest-flow` 的具体执行步骤。正式行为边界以 `../SKILL.md` 为准。
|
||||
|
||||
Concrete operational checklist for the `reader-digest-flow` skill.
|
||||
## 1. 日报 Pipeline
|
||||
|
||||
## Default Operating Model
|
||||
### 默认参数
|
||||
|
||||
### Layering
|
||||
|
||||
- `reader` layer:
|
||||
- FreshRSS pull
|
||||
- extraction
|
||||
- summary/filter/payload generation
|
||||
- selected-article summary capability
|
||||
- OpenClaw / skill layer:
|
||||
- public digest generation
|
||||
- internal review digest generation
|
||||
- Hugo publishing
|
||||
- chat reporting
|
||||
- user confirmation handling
|
||||
- calling selected-article summaries
|
||||
- IMA upload orchestration
|
||||
- Hugo layer:
|
||||
- public digest browsing and archive only
|
||||
- IMA layer:
|
||||
- long-term storage for selected article notes only
|
||||
|
||||
### Hard rules
|
||||
|
||||
- Do not upload the full digest to IMA.
|
||||
- Upload only explicitly user-selected articles to IMA.
|
||||
- Do not generate a digest without a real payload.
|
||||
- Generate two views from the same payload: a public digest for Hugo and an internal review digest for chat/operator workflow.
|
||||
- Do not expose internal review states or operator-facing labels in the public digest.
|
||||
- Do not re-fetch original URLs for selected summaries; use existing extracted text.
|
||||
- If the main pipeline fails, inspect the run first; when `inspect_resume_plan` says `recommended_action=resume`, continue via the async resume job path instead of stopping immediately.
|
||||
- Always branch on reader's top-level reconciled `status`; treat `status_source` and `state_conflict` only as explanatory metadata.
|
||||
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization, with fallback to main `LLM_*` only if needed.
|
||||
- Actively report progress after each completed phase.
|
||||
|
||||
## Step-by-step checklist
|
||||
|
||||
### 1. Run reader pipeline
|
||||
|
||||
Use the formal MCP workflow path as the default production route.
|
||||
Prefer MCP run/status/result operations over direct path stitching. Only fall back to CLI or direct file inspection for debug / manual troubleshooting.
|
||||
|
||||
Formal production startup sequence:
|
||||
|
||||
1. `start_freshrss_pipeline_job`
|
||||
2. `get_freshrss_pipeline_job_status`
|
||||
3. `get_freshrss_pipeline_job_result`
|
||||
4. after success, continue with `run_id` via `get_run_status` / `get_delivery_payload` / `get_run_report`
|
||||
|
||||
If the main pipeline job ends in `failed`:
|
||||
|
||||
1. inspect the linked run with `get_run_status`
|
||||
2. call `inspect_resume_plan(run_id)`
|
||||
3. if `recommended_action=resume`, continue with:
|
||||
- `start_resume_job`
|
||||
- `get_resume_job_status`
|
||||
- `get_resume_job_result`
|
||||
4. if `recommended_action=read_terminal_result`, continue from the terminal run result
|
||||
5. if `recommended_action=start_new_run`, stop and report the failure
|
||||
|
||||
Treat the old synchronous `run_freshrss_openclaw_pipeline` as debug / light validation / fallback only.
|
||||
|
||||
Project root:
|
||||
|
||||
```bash
|
||||
/home/ubuntu/zhu/github/reader
|
||||
```json
|
||||
{
|
||||
"limit": 7,
|
||||
"include_read": false,
|
||||
"mark_read": true,
|
||||
"debug_artifacts": false,
|
||||
"timeout_seconds": 60,
|
||||
"max_retries": 2
|
||||
}
|
||||
```
|
||||
|
||||
Default behavior for a normal production run:
|
||||
|
||||
- if the user did not specify a count, randomly choose a limit between 5 and 10 items for that run
|
||||
- run with mark-read enabled
|
||||
- do not enable `debug_artifacts`
|
||||
- only skip mark-read if the user explicitly says the run is debug, test, or validation
|
||||
- only enable `debug_artifacts` if the user explicitly says the run is debug, test, validation, or troubleshooting
|
||||
|
||||
Typical artifacts to inspect after a successful run:
|
||||
### 正式调用序列
|
||||
|
||||
```text
|
||||
outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json
|
||||
outputs/freshrss/rerun/<run-id>/candidates/digest-brief.json
|
||||
outputs/freshrss/rerun/<run-id>/run-report.json
|
||||
outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json
|
||||
start_freshrss_pipeline_job
|
||||
→ get_freshrss_pipeline_job_status
|
||||
→ get_freshrss_pipeline_job_result
|
||||
→ get_run_status
|
||||
→ get_delivery_payload / get_run_report
|
||||
```
|
||||
|
||||
Notes:
|
||||
状态动作:
|
||||
|
||||
- `digest-brief.json` is the preferred input for **public digest** generation.
|
||||
- It is a lighter public-only view and currently includes only `keep` candidates.
|
||||
- If `digest-brief.json` is missing, fall back to `openclaw-delivery-payload.json`.
|
||||
- If the synchronous MCP wrapper times out but a real reader run was still created, do not discard the run; continue from run truth using `list_runs`, `get_run_report`, and `get_delivery_payload`.
|
||||
- `running`:按合理间隔继续轮询;
|
||||
- `success`:读取结果,保存 `run_id`;
|
||||
- `failed`:读取关联 Run 并执行 Resume Plan;
|
||||
- `partial`:优先检查 Run Report、Artifact 和 Recovery 信息。
|
||||
|
||||
**⚠️ Output path divergence when run from Feishu context:**
|
||||
When triggered via MCP from inside a Feishu session, the extracted files may be written to an MCP-managed temp path instead of `outputs/freshrss/rerun/<run-id>/extracted/`. Verify the actual output path from the pipeline result before continuing to Phase 5/6. If the project-relative path doesn't exist, use the MCP result's `extracted_path` directly or copy the files to the expected project location.
|
||||
|
||||
### 2. Generate digest markdown
|
||||
|
||||
Generate two output views from the same run, preferably in one model call:
|
||||
|
||||
1. a **public digest** for Hugo / public readers
|
||||
2. an **internal review digest** for chat / operator workflow
|
||||
|
||||
Input preference:
|
||||
|
||||
- **public digest**: prefer `candidates/digest-brief.json`
|
||||
- **internal review digest**: use `candidates/openclaw-delivery-payload.json`
|
||||
|
||||
Recommended generation pattern:
|
||||
|
||||
- Pass the public brief and the full payload as two explicitly labeled input blocks.
|
||||
- Ask the model to return both outputs in one response.
|
||||
- Prefer a structured response shape (for example JSON with `public_digest_markdown` and `internal_review_digest_markdown`) when post-processing is needed.
|
||||
|
||||
Before publishing, persist the generated digest artifacts back into the same reader run directory:
|
||||
恢复序列:
|
||||
|
||||
```text
|
||||
outputs/freshrss/rerun/<run-id>/digest/public_digest.md
|
||||
outputs/freshrss/rerun/<run-id>/digest/internal_review_digest.md
|
||||
outputs/freshrss/rerun/<run-id>/digest/combined.json
|
||||
inspect_resume_plan
|
||||
→ recommended_action=resume
|
||||
→ start_resume_job
|
||||
→ get_resume_job_status
|
||||
→ get_resume_job_result
|
||||
```
|
||||
|
||||
Then write only the public digest into Hugo here:
|
||||
如果建议为 `read_terminal_result`,直接读取结果;如果为 `start_new_run`,停止并等待用户决定。
|
||||
|
||||
不要根据 `run_id` 拼目录。后续只消费调用返回的 `output_dir`、`artifact.path`、`delivery_output`、`report_output`。
|
||||
|
||||
## 2. 候选汇报与选择
|
||||
|
||||
优先使用当前 Run 的 digest brief Artifact;不可用时使用 `get_delivery_payload` 返回的候选。
|
||||
|
||||
展示规则:
|
||||
|
||||
1. 使用候选数组原始顺序并从 1 编号;
|
||||
2. 不因 `keep/review` 分组而重新编号;
|
||||
3. 每篇展示标题、来源、摘要和判断理由;
|
||||
4. 用户编号只映射当前候选数组;
|
||||
5. 跨 Artifact 读取详情时按 URL 或完整 `item_id` 匹配。
|
||||
|
||||
不要按 `summary-batch`、`item-XX` 和候选数组的相同位置推断它们是同一篇文章。
|
||||
|
||||
## 3. Hugo 日报
|
||||
|
||||
用户确认发布文章后:
|
||||
|
||||
1. 读取 `public-digest-example.md`;
|
||||
2. 仅使用用户选中的文章生成公开内容;
|
||||
3. 写入 Hugo 当日页面;
|
||||
4. 前台执行部署,避免把构建日志作为聊天通知;
|
||||
5. 验证首页、列表页和详情页。
|
||||
|
||||
当前部署位置:
|
||||
|
||||
```text
|
||||
/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md
|
||||
```
|
||||
|
||||
Recommended public-digest front matter:
|
||||
验证至少覆盖:
|
||||
|
||||
```toml
|
||||
+++
|
||||
title = "AI 日报 · YYYY-MM-DD"
|
||||
date = YYYY-MM-DDTHH:MM:SS+08:00
|
||||
summary = "当日日报摘要"
|
||||
+++
|
||||
```text
|
||||
http://127.0.0.1:14322/
|
||||
http://127.0.0.1:14322/daily/
|
||||
http://127.0.0.1:14322/daily/YYYY-MM-DD/
|
||||
```
|
||||
|
||||
Recommended **public digest** structure:
|
||||
页面不存在时检查源文件、构建结果、容器状态和页面日期。需要 Docker/Hugo 深度排障时再读取部署项目自己的文档,不把历史事故规则复制回主 Skill。
|
||||
|
||||
- `今日概览`
|
||||
- `今日重点`
|
||||
- `趋势观察`
|
||||
- `延伸阅读`
|
||||
## 4. 单篇知识笔记
|
||||
|
||||
Public digest constraints:
|
||||
用户确认 IMA 文章后:
|
||||
|
||||
- Use only the public brief view when present.
|
||||
- Keep public wording free of internal workflow language.
|
||||
- Optimize for concise public readability with solid information density.
|
||||
- For each `今日重点` item, add one short editor-style sentence explaining why the item matters in today's digest (for example: `这篇内容更值得关注的原因在于……`).
|
||||
- Prefer rendering highlight points as a short public label such as `值得关注:` followed by one bullet per line.
|
||||
- **Article numbering: MUST use `1.` `2.` `3.` (Arabic numeral + period). Do NOT use `① ② ③`,`一、二、三`,`第一条` or any other variant.**
|
||||
- **All four sections are required. Missing any one is a format violation.**
|
||||
- **Hugo digest must include ALL keep articles from `digest-brief.json`. The IMA deposition subset is a separate downstream step.**
|
||||
1. 从候选中取得 URL 和完整 `item_id`;
|
||||
2. 从 Run Artifact 中找到匹配的 extracted 文件;
|
||||
3. 交叉验证 `article.item_id` 或 URL;
|
||||
4. 对每个单篇文件调用 `start_article_summary_job(extracted_path=<返回的 Artifact 路径>, selected_ids=[<完整 item_id>])`;
|
||||
5. `selected_ids` 传完整、非空、去除 `cand:` 前缀后的 `item_id`;
|
||||
6. 从 Job 结果的 `written_paths` 读取 Markdown。
|
||||
|
||||
Recommended **internal review digest** structure:
|
||||
正式序列:
|
||||
|
||||
- `今日候选概况`
|
||||
- `已入选重点`
|
||||
- `待你确认`
|
||||
- `建议沉淀到 IMA`
|
||||
- `原始候选清单`
|
||||
|
||||
Internal review digest constraints:
|
||||
|
||||
- Do not show `rank` values.
|
||||
- Replace machine states with Chinese labels such as `已入选` / `待确认` / `暂不纳入`.
|
||||
- For `已入选重点`, include a fuller summary plus a short judgment paragraph.
|
||||
- For `待你确认`, include a fuller summary, reason, and recommendation.
|
||||
- Keep it readable as an operator review draft, not a raw payload dump.
|
||||
- **Feishu compatibility: never use markdown tables in the digest report.**
|
||||
When delivered via Feishu, a single table forces the whole message to plain
|
||||
text. Use lists and sections instead. See `references/feishu-format-notes.md`.
|
||||
|
||||
### 3. Publish Hugo
|
||||
|
||||
Publish only the public digest to Hugo.
|
||||
Treat Hugo publication as the default continuation of a successful normal daily digest run. Do not ask for a second confirmation before generating/writing the public digest and publishing it, unless the user explicitly requests not to publish to Hugo.
|
||||
|
||||
**⚠️ Hugo redeploy: always use foreground terminal, never background+notify_on_complete.**
|
||||
|
||||
The Hugo redeploy runs via `./redeploy.sh` and completes in ~10 seconds. **Do NOT use `terminal(background=true, notify_on_complete=true)`** for this step — the Gateway will push the raw Docker build output (compiler logs, layered build output, nginx config) to Feishu as a notification. This output is machine-readable, not human-readable, and the user has explicitly said this is noise.
|
||||
|
||||
**Correct approach:**
|
||||
|
||||
```python
|
||||
# ✅ Foreground terminal with adequate timeout
|
||||
cmd = "cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh"
|
||||
result = terminal(cmd, timeout=120)
|
||||
# Then verify 200 OK on the detail page
|
||||
```text
|
||||
start_article_summary_job
|
||||
→ get_article_summary_job_status
|
||||
→ get_article_summary_job_result
|
||||
```
|
||||
|
||||
**Wrong approach (DO NOT use):**
|
||||
内容依据是 extracted Artifact 中的 `article.plain_text`。不得重新抓取原 URL,不得使用原文之外的知识扩写。
|
||||
|
||||
```python
|
||||
# ❌ Background + notify sends raw build logs to Feishu
|
||||
terminal("cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh", background=True, notify_on_complete=True)
|
||||
```
|
||||
|
||||
This rule only applies to Hugo redeploy (fast, deterministic output). For long-running Reader pipeline jobs (>60s), background + notify_on_complete is still appropriate since the output is meaningful content (article summaries, pipeline stats).
|
||||
|
||||
1. **Pre-check: verify Hugo content directory exists**
|
||||
```bash
|
||||
ls -d /home/ubuntu/zhu/apps/hugo-site/content/daily/ 2>/dev/null || mkdir -p /home/ubuntu/zhu/apps/hugo-site/content/daily/
|
||||
```
|
||||
If the entire `content/` directory is missing, create it before proceeding.
|
||||
**Do not assume the Hugo site has a standard structure.**
|
||||
|
||||
2. write the public digest markdown to:
|
||||
- `/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md`
|
||||
3. redeploy Hugo immediately after writing:
|
||||
- `cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh`
|
||||
4. verify all three URLs before continuing:
|
||||
- `http://127.0.0.1:14322/`
|
||||
- `http://127.0.0.1:14322/daily/`
|
||||
- `http://127.0.0.1:14322/daily/YYYY-MM-DD/`
|
||||
|
||||
Expected verification targets:
|
||||
|
||||
- homepage works
|
||||
- `/daily/` works
|
||||
- `/daily/YYYY-MM-DD/` works
|
||||
|
||||
### 4. Report digest in chat
|
||||
|
||||
Provide the internal review digest in chat and ask which articles should be retained.
|
||||
|
||||
Hard reporting rule:
|
||||
- do not send only article titles
|
||||
- for each article, include at least a one-sentence summary and a short recommendation / judgment so the user can decide without reopening the source
|
||||
- **Feishu**: avoid markdown tables entirely. Use lists with headings.
|
||||
|
||||
### 5. Generate selected article summaries
|
||||
|
||||
Only do this after the user explicitly confirms which articles to retain.
|
||||
|
||||
Preferred MCP path:
|
||||
|
||||
- `start_article_summary_job`
|
||||
- `get_article_summary_job_status`
|
||||
- `get_article_summary_job_result`
|
||||
|
||||
Expected inputs:
|
||||
|
||||
- `extracted_path`
|
||||
- `selected_ids`
|
||||
**⚠️ item_id 前缀注意:**
|
||||
The `candidates` array IDs use `cand:sha256:xxx` format, but individual extracted item files use `sha256:xxx` (no `cand:` prefix). When passing `selected_ids` to article-summary tools, strip the `cand:` prefix. If the ID doesn't match, the summary tool won't find the article.
|
||||
|
||||
**Path resolution for article-summary:**
|
||||
After a pipeline run via MCP (Feishu context), the extracted files may live at an MCP-managed temp path, not the expected `outputs/freshrss/rerun/<run-id>/extracted/`. Before calling `generate_article_summaries` or `start_article_summary_job`, verify the extracted path exists. If not, use the `run_id` from the pipeline result to locate the actual output directory through `get_run_report`, or copy the files from the MCP-managed path.
|
||||
|
||||
Single-item rule:
|
||||
|
||||
- when `extracted_path` is `outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json`, call one summary job per file
|
||||
- in that case, `selected_ids` should contain only the matching single `item_id`
|
||||
- if the candidate ID came from the delivery payload, strip the `cand:` prefix before passing it
|
||||
|
||||
Recommended output layout:
|
||||
|
||||
- `outputs/freshrss/single_summaries/YYYY-MM-DD/`
|
||||
- async job state: `outputs/freshrss/article_summary_jobs/<job_id>/`
|
||||
|
||||
Recommended production sequence:
|
||||
|
||||
1. call `start_article_summary_job`
|
||||
2. poll `get_article_summary_job_status` until `status` becomes `success` or `failed`
|
||||
3. on success, call `get_article_summary_job_result` and continue downstream from `written_paths`
|
||||
|
||||
Hard fallback rule:
|
||||
|
||||
- If the async MCP job path returns timeout / transport failure / job-launch failure (for example MCP timeout while the reader article-summary workflow itself is still healthy), do not treat that as article-summary business failure.
|
||||
- Immediately retry through the local reader environment under `/home/ubuntu/zhu/github/reader` using the project `.venv`, calling the article-summary workflow directly.
|
||||
- The production goal is successful generation of the selected-article Markdown files; async MCP job is preferred, but local `.venv` execution is the required fallback path.
|
||||
|
||||
Synchronous helper:
|
||||
|
||||
- `generate_article_summaries` remains available for debug / light validation only, not as the default production path.
|
||||
|
||||
CLI fallback:
|
||||
CLI 仅在异步 MCP 不可用或人工排障时使用:
|
||||
|
||||
```bash
|
||||
python scripts/run_article_summaries.py \
|
||||
--extracted outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json \
|
||||
--ids <item_id_without_cand_prefix> \
|
||||
--output-dir outputs/freshrss/single_summaries/YYYY-MM-DD
|
||||
--extracted <returned-extracted-path> \
|
||||
--ids <full-item-id> \
|
||||
--output-dir <explicit-output-dir>
|
||||
```
|
||||
|
||||
### 6. Upload selected summaries to IMA
|
||||
## 5. IMA 上传
|
||||
|
||||
Upload only the generated markdown files for the selected articles.
|
||||
Once the user has selected the articles to retain, treat that selection itself as the authorization to continue the IMA deposition step; do not ask for a second confirmation about uploading into the knowledge base.
|
||||
用户在知识沉淀阶段的文章选择即为上传授权。
|
||||
|
||||
**⚠️ IMA upload 500 retry: Always wrap IMA uploads in a retry loop.**
|
||||
IMA's OpenAPI may return HTTP 500 on the first attempt. If the first upload fails, wait ~3 seconds and retry. Normally the second attempt succeeds. If `cos-upload.cjs` is used (for CDN-backed uploads), check subprocess stderr even when exit code is 0 — an HTTP error in the upload service may still produce exit code 0.
|
||||
执行顺序:
|
||||
|
||||
Hard execution rules before upload:
|
||||
1. 检查生成的 Markdown 与来源;
|
||||
2. 文件名规范化为 `<完整文章标题>.md`;
|
||||
3. 确认目标为 `daily` knowledge base;
|
||||
4. 执行 preflight、create_media、COS upload、add_knowledge;
|
||||
5. 验证知识库条目存在。
|
||||
|
||||
1. reformat/check the generated markdown into IMA-facing final content
|
||||
2. normalize the final upload filename to `<文章标题>.md`
|
||||
3. do not use internal temp names such as `ima-*`, `item-*`, `summary-*`, or English slug filenames as the final uploaded object name
|
||||
4. if the knowledge base already contains the same filename, append a timestamp suffix before `.md`
|
||||
5. if local work needs internal temp names, create a final upload copy with the user-facing title before calling IMA APIs
|
||||
上传格式与 API 参数分别见:
|
||||
|
||||
Default target knowledge base:
|
||||
- `ima-format-quickref.md`
|
||||
- `ima-upload-api.md`
|
||||
- `ima-credential-chain.md`
|
||||
|
||||
- `daily`
|
||||
- Read `IMA_DAILY_KNOWLEDGE_BASE_ID` / `IMA_DAILY_KNOWLEDGE_BASE_NAME` from reader `.env`
|
||||
- Verify the configured target at runtime before upload
|
||||
- If the configured target is unavailable, resolve by name `daily`; if still absent, create `daily`
|
||||
失败时记录发生在 preflight、create_media、COS upload 或 add_knowledge 的具体阶段。不要把失败的 Markdown 改写成 URL 导入或 Notes 类型来绕过错误。
|
||||
|
||||
## Hard rules recap
|
||||
## 6. CLI fallback 原则
|
||||
|
||||
- Never generate a digest from placeholder or example data when a real run is expected.
|
||||
- For normal runs, mark processed FreshRSS items as read unless the user explicitly requested a debug/test/validation run.
|
||||
- Public digest goes to Hugo; internal review digest goes to chat; neither full digest goes to IMA.
|
||||
- Only explicitly user-selected articles go to IMA.
|
||||
- All daily IMA deposition must go directly into the IMA knowledge-base path as Markdown knowledge items (`media_type=7`), not through the IMA notes path.
|
||||
- Uploading to IMA notes, or creating notes first and then linking them into a knowledge base, does not count as SOP completion.
|
||||
- Selected article summaries use extracted text, not live refetch.
|
||||
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization.
|
||||
- Final IMA upload filenames must use user-facing article titles, not internal slugs or workflow temp names.
|
||||
CLI 仅在以下场景使用:
|
||||
|
||||
- MCP 服务不可用;
|
||||
- Tool transport/launch 失败且无法取得有效 Job;
|
||||
- 用户明确要求本地调试;
|
||||
- 人工排障需要直接检查脚本输出。
|
||||
|
||||
如果已取得 `job_id` 或 `run_id`,先查询真实状态,避免因响应超时重复启动任务。
|
||||
|
||||
Reference in New Issue
Block a user