--- name: reader-digest-flow description: Orchestrate the end-to-end daily digest workflow around the reader project. Use when the user wants to run the AI daily digest flow, report candidates for review, let the user decide which articles go into the Hugo daily digest, then generate and publish the public digest, and optionally summarize selected articles into IMA knowledge notes. Triggers include requests like '跑今天日报', '生成日报', '汇报今天内容', '把选中的文章沉淀', '更新 Hugo', or any request to operate the reader → report → user-selection → Hugo → knowledge-base flow. --- **🚨 强制前置规则 [MANDATORY] — 每次涉及日报/IMM 沉淀操作前必须执行:** - **Step 0: 完整加载本 skill(`skill_view(name='reader-digest-flow')`),阅读到本行之后** - **Step 0.2: 加载 IMA 格式基准文件(`skill_view(name='reader-digest-flow', file_path='references/ima-format-reference.md')`)** - **Step 0.5: 加载 `ima-skill` 的 knowledge-base 子模块(`skill_view(name='ima-skill', file_path='knowledge-base/SKILL.md')`)** - 🔴 不读 skill 就干活 = 必错,历史 100% 验证 - 读完 skill 中的「铁律」和引用文件后再动手,禁止凭记忆操作 # Reader Digest Flow ## 🔴 铁律:只跑一次 pipeline,且只能按用户指令跑 ### 🚨 唯一合法的触发条件 **只有一种情况可以运行 pipeline:用户明确说了「跑日报」或「再跑一篇」「再跑一次」。** **禁止因以下任何原因运行 pipeline:** - ❌ "验证"数据 → 读已有文件 - ❌ "重新确认"候选排序 → 读已有文件 - ❌ "看看会不会有不同结果" → 读已有文件 - ❌ 用户问"为什么第X篇是XX" → 读已有文件 - ❌ 用户质疑编号 → 读已有文件 - ❌ 任何"让我先查一下"的冲动 → **先停,读已有文件** **用户明确指令(2026-07-07):** > **"只有当我要求重新跑一次日报时候,才去重新跑一次"** ### 🚨 同一天多次重跑的收敛陷阱(2026-07-27 教训) 当一天内连续请求多次重跑时,FreshRSS feed 会返回同一批文章(因为当天的新内容就那么多),pipeline 排序也趋于稳定。此时**继续重跑不会带来新候选**——不是算法的问题,是原料池已见底。 **处理建议:** - 如果用户第一次说"质量太低重跑",用 `--include-read` 扩大候选池(skill 已有此规则) - 如果第二次重跑后仍是同一批 keep 文章,**直接说明情况**:今天 feed 里质量最好的就这些,没有更好的替换候选了 - 不要无限重跑,也不要自作主张降低质量门槛把 review 提成 keep ### 🚨 用户要求的重跑 vs 私自重跑(2026-07-20 教训) **只有以下两种跑法:** | 场景 | 是否允许 | 示例 | |:----|:-------:|:----| | 用户明确要求"再跑一篇""换一批""重新跑一次" | ✅ **允许** | 老大说"再跑一篇,这批质量一般" | | 你自己想"验证一下""再看看有没有更好的结果" | ❌ **禁止** | 不能因候选质量不好就自作主张重跑 | **判断标准:** 只有用户亲口说出明确的重新执行指令才算。你心里觉得"这批不够好"不算。 **重跑后编号规则:** 新批次的候选从 1 重新编号(新 run_id = 新序号体系),与旧批次完全独立。不要在同一个 conversation 里混用两批次的编号。 ### 🚨 每次运行必须记录 run ID Pipeline 跑完后,**立即将 run_id 存入 memory**: ``` memory(action="add", content="当前日报 run_id: 20260707-012841,候选 7 篇") ``` 这样后续所有操作都指向同一个数据集,不会产生混淆。 ### 🚨 编号规则(用户引用的唯一来源) **用户的编号永远指向 digest-brief.json 的 top_candidates 数组(1-based index)。** - 显示给用户的候选编号 = `top_candidates` 数组顺序 - 用户说"第2篇" = `top_candidates[1]` - **绝对禁止用 `summary-batch.json` 或 `extracted/` 文件编号来理解用户的输入** ### 🚨 候选展示时 **必须带编号** 向用户展示候选列表时,必须打印 `digest-brief.json` 中 `top_candidates` 数组的编号和标题,并且给用户选择的编号必须与这个编号一致。 正确做法: ```bash python3 -c "import json d = json.load(open('...digest-brief.json')) for i, c in enumerate(d['top_candidates'], 1): print(f'{i}. {c[\"title\"][:60]} | {c[\"source_name\"]}')" ``` 错误做法: - ❌ 按其他文件(如 summary-batch.json)的顺序展示 - ❌ 只展示标题不带编号 - ❌ 展示编号但与 digest-brief.json 的数组索引不匹配 ### 🚨 永远不要重新分组/重编号候选 — 保持原始 pipeline 序号 **教训(2026-07-20):用户质问「为什么待确认,你给我的序号是重新排?」** 错误演示(按 keep/review 分组后各自编号 1. 2. …): ``` 🟢 已入选(4 篇) 1. PagePilot ... 2. Harness ... 🟡 待确认(2 篇) 1. Kimi K3 ... 2. Kimi K3 不同来源... ``` **铁律:** 1. **永远维护原始编号。** 候选列表的编号是 `enumerate(top_candidates, 1)` 的结果,用户只认这个编号体系。 2. **不要按 keep/review 分组各自编号。** 如果分组展示,每篇候选也必须保留其在 `top_candidates` 数组中的原始序号。 3. **推荐的展示方式:** 平铺编号列表,每行用 emoji 状态指示器(🟢已入选 / 🟡待确认)标记决策状态,不改变序号。 4. **重新跑 pipeline 后:** 新批次的候选从 1 重新编号(新 run_id = 新序号体系),与旧批次完全独立。 ### 🚨 当用户说"还是错的"时 1. **不跑新的 pipeline** 2. **立即打印 digest-brief.json 的 top_candidates 编号+标题,展示给用户确认** 3. 用 URL 交叉引用匹配候选与摘要 4. 如果对不上,承认错误 ### 匹配候选与其摘要的唯一正确方式:按 URL 交叉引用 绝不按数组位置索引把 `summary-batch.json` 和 `digest-brief.json` 对齐——它们用不同的排序逻辑,位置从来不匹配。 ### 🚨 沟通纪律:只说用户问的 **跟 pipeline 技术同等重要。** 用户问了一个具体问题,只回答那个问题。禁止: - ❌ 主动分析"为什么会发生这种情况的历史原因" - ❌ 补充"顺便说一下,以前也…" - ❌ 解释用户没问的场景的细节 - ❌ 在道歉/认错后面跟长段分析 **用户原话(2026-07-07):** > "为什么你还要提关于 7-1 的 16 次原因????我没有再次问你阿???" **正确做法:** 用户问「为什么第二是AICR」→ 答「我跑了两遍pipeline,排序变了」→ 停止。用户没问的就不说。 ### 🚨 HARD PRE-UPLOAD CHECKLIST — IMA Phase 7 Before uploading any article to IMA, you MUST run through this checklist. A single unchecked item means DO NOT UPLOAD — fix it first. 1. [ ] **skill 已加载** — `skill_view(name='reader-digest-flow')` 已执行并阅读到文件末尾 2. [ ] **ima-skill 已加载** — `skill_view(name='ima-skill', file_path='knowledge-base/SKILL.md')` 已执行 3. [ ] **格式基准文件已读** — `skill_view(name='reader-digest-flow', file_path='references/ima-format-reference.md')` 已执行并理解其格式要求 4. [ ] **使用 media_type=7** — 代码使用 `preflight → create_media → COS → add_knowledge(media_type=7)`,不是 `import_doc`(media_type=11) 5. [ ] **文件名 = 完整原始标题** — 文件名为 `<完整原始文章标题>.md`,如 `从Vibe Coding到Harness—— 一套大仓AI工程化实战.md` 6. [ ] **文件名未缩写** — 对比 pipeline 中 `digest-brief.json` 的 `title` 字段与上传文件名,必须完全一致(不含 .md 后缀比对) 7. [ ] **内容来自 pipeline 数据** — 正文从 `extracted/item-XX.extracted.json` 的 `article.content` 或 `summary-batch.json` 的 `summary` 字段读取 8. [ ] **内容不小于 500 字** — 正文长度 >= 500 字符(中文字算1字) 9. [ ] **`add_knowledge` 的 `title` 字段正确** — 传完整标题,不加 .md 后缀 10. [ ] **文件内容符合 ima-format-reference.md 标准** — 必须包含:顶部 `Source:`/`Category:` 元数据、5 个 section、底部 `关键词`/`主题`。其中 **核心结论和主要论点是段落**(非分点),**关键方法/机制、重要细节、可复用启发是分点**。每个 section 展开到完整知识点级别。目标文件大小 **4,000+ bytes** **执行前必须逐项打勾**,少一项就停。历史教训(2026-07-08):两次 IMA 上传都踩了这些坑——先是 media_type=11 笔记格式,后是文件名缩写+自编摘要。pre-upload checklist 就是为了防这个。 ## Core Rules for fetching, extraction, filtering, payload generation, result reading, and selected-article post-processing. - **Daily digest production runs should default to a random article count between 5 and 10 unless the user explicitly specifies a count.** - Prefer **CLI path** (`scripts/run_freshrss_pipeline.py` via `.venv/bin/python3`) for production runs — it bypasses the MCP transport-layer 120s timeout and has no circular import. Use MCP workflow operations only for inline Feishu/tool queries where CLI is unavailable. - **Main FreshRSS production runs should default to the async MCP job path**: `start_freshrss_pipeline_job` → `get_freshrss_pipeline_job_status` → `get_freshrss_pipeline_job_result`. Use the synchronous `run_freshrss_openclaw_pipeline` only for debug / light validation / fallback. - **If the main pipeline job fails, do not immediately abandon the run.** First inspect the linked run through `get_run_status` and `inspect_resume_plan`; when the plan returns `recommended_action=resume`, continue through `start_resume_job` → `get_resume_job_status` → `get_resume_job_result`. - **Operational note from validation:** before the main pipeline async job existed, a debug/test run could still complete successfully inside reader even when the synchronous MCP wrapper returned timeout. For historical/debug cases, continue from real run truth using MCP status/result query tools rather than treating the whole flow as failed. - For reader run observation and result reading, prefer MCP tools such as status / payload / report queries instead of having OpenClaw or this skill hand-build reader output paths. - **When reader status APIs return reconciled status, always branch on the top-level `status`.** Treat `status_source` and `state_conflict` as explanatory metadata; do not re-derive flow control from stale stage names. - **For normal production runs, mark processed FreshRSS items as read. Only skip mark-read when the user explicitly says the run is debug/test/validation.** - **For normal production runs, do not pass `debug_artifacts=true`. Only enable debug artifacts when the user explicitly says the run is debug/test/validation or when troubleshooting is the goal.** - **Do not generate the daily digest from examples or placeholder data. Always require a real payload first.** - Generate the daily digest markdown from the payload in OpenClaw. - Split outputs into a **public digest** for Hugo and an **internal review digest** for chat / operator decision-making. - **Public digest and internal review digest are orchestration-layer outputs, and by default are generated by the current OpenClaw session model rather than inheriting reader `LLM_*` settings.** - **Publish only the public digest to Hugo, and only after the user has explicitly selected which articles go into the daily digest.** - The default flow is: pipeline → report candidates → **user selects Hugo articles** → generate public digest → publish Hugo → (optional) user selects IMA articles → IMA deposition. - Do NOT generate the public digest or publish Hugo before the user has confirmed the Hugo article selection. - Report the internal review digest back to the user in chat. - **Do not upload the full daily digest to IMA.** - **Do not upload any selected article summary to IMA until the user has explicitly confirmed the selection. Once the user has confirmed which articles to keep, proceed directly with IMA deposition and do not ask for a second confirmation about uploading to the knowledge base.** - Upload **only user-selected articles** to IMA as individual knowledge notes. - **The formal IMA upload object must be the generated single-article summary content, uploaded directly into the target IMA knowledge base as a Markdown knowledge item (`media_type=7`), not the original article webpage URL.** - **For daily deposition, the target must explicitly be the IMA "knowledge-base" type path, not the IMA "notes" type path.** - **Default deposition: Markdown 文件上传(media_type=7)** - 必须用完整的文件上传流程: **preflight → create_media → COS upload → add_knowledge(media_type=7)** - 先将文章内容写入本地 `.md` 文件,再通过文件上传流程提交 - 文件命名规则:`<文章标题>.md`(使用用户界面看到的原始中文标题) - **Do NOT use `import_doc` + `media_type=11`** — the user explicitly rejected this format (called it "笔记格式"). 只有 `media_type=7`(Markdown 文件上传)才是正确的沉淀格式。 - **Do NOT use `import_urls`** — the user explicitly rejected this (called it "公众号格式"). - **Adding content as an IMA note, or uploading to any non-knowledge-base container first and then linking it indirectly, does not count as completing the SOP. Completion requires direct knowledge-base ingestion.** - Default IMA target for daily single-article summaries is the `daily` knowledge base. - Resolve that default target from reader `.env` (`IMA_DAILY_KNOWLEDGE_BASE_ID`, `IMA_DAILY_KNOWLEDGE_BASE_NAME`) and verify it at runtime before upload. - If the configured daily knowledge base is missing, try to locate it by name; if still missing, create `daily` and continue. - For selected article notes, use the dedicated article-summary flow in `reader`. - Prefer `ARTICLE_SUMMARY_*` LLM settings for selected article summaries; fall back to the main `LLM_*` settings only if the dedicated settings are absent. - Do not re-fetch original article URLs for selected summaries; always use the existing extracted article text. - Default selected-article summary output should be organized by date, for example under `outputs/freshrss/single_summaries/YYYY-MM-DD/`. ## Fixed Flow ### Phase 1: Run reader pipeline In the `reader` project, run the FreshRSS pipeline and obtain a real payload. Formal production start path: - `start_freshrss_pipeline_job` - `get_freshrss_pipeline_job_status` - `get_freshrss_pipeline_job_result` After the async job succeeds, treat the returned `run_id` as the stable handle for downstream `get_run_status` / `get_delivery_payload` / `get_run_report` reads. Minimum expected artifacts: - `outputs/freshrss/rerun//candidates/digest-brief.json` — **精简候选清单**(唯一要读取的文件,含 keep + review) - `outputs/freshrss/rerun//candidates/openclaw-delivery-payload.json` — 完整 payload(pipeline 内部产物,digest-brief 由此推导,流程中无需读取) - `outputs/freshrss/rerun//run-report.json` - extracted article data: `outputs/freshrss/rerun//extracted/item-01.extracted.json` 等(每个候选一篇,格式为 `{"success": true, "article": {...}}`,Phase 6 使用) If the main pipeline job fails: - read the linked run through `get_run_status` - call `inspect_resume_plan(run_id)` - if `recommended_action=resume`, continue with: - `start_resume_job` - `get_resume_job_status` - `get_resume_job_result` - if `recommended_action=read_terminal_result`, continue from the terminal run result instead of retrying - if `recommended_action=start_new_run`, stop and report the exact failure point - **if the resume fails with `Missing required value` errors** (e.g. `FRESHRSS_API_BASE_URL`), this is MCP server environment isolation — the MCP process doesn't source `.env`. **Do NOT retry resume.** Fall back immediately to the synchronous `run_freshrss_openclaw_pipeline`, which handles env isolation differently and completes successfully. ### Phase 2: Report candidates to the user (internal review digest only) **⛔ USER PREFERENCE (2026-07-01): Do NOT re-summarize the project background. The user has run this flow many times. Report ONLY the day's candidates — no pipeline overview, no project context, no "as you know" framing. If the previous pipeline results were already reported earlier in the same conversation, omit the candidate details too — just report the status and any anomalies. The user said "你已经向我总结了3次了" — respect this.** After the pipeline run succeeds, do NOT generate the public digest or publish Hugo yet. Read candidates and produce **only the internal review digest** in a concise format (not the full payload dump). Recommended **internal review digest** structure: - `今日候选概况` (count table + one-liner per article) - `已入选重点` (only keep articles) - `待你确认` (only review articles) Data source: use `candidates/digest-brief.json` (`top_candidates` array) — it now includes both keep and review articles in a concise, simplified format (title, source, summary, highlights, category). No need to read the full payload. **⚠️ 每篇文章简述长度要求:** 每篇文章的描述必须包含 2-3 句话(含核心问题、核心方法/结论、业务价值/效果),不能只有一句话或关键词堆砌。老大需要足够的信息来判断哪些文章值得发布和沉淀。示例: > **🟢 [已入选] 得物推荐系统诊断 Agent:「推查查」** | 得物技术 > 得物自研"推查查"诊断Agent,融合Highway确定性流水线与ATV自主推理双模式,实现推荐系统异常快速定位与复杂问题根因分析。核心创新是将排查经验通过Skill/Story原子化封装与进化层持续沉淀,推动排查从人工经验向自动化智能转变。该方案在得物生产环境中验证了从小时级定位缩短到分钟级的效果。 Internal review digest writing rules: - **Keep this as a terse bullet-style report.** Do not write essay-length paragraphs. - **Each article must include a 1-2 line summary describing what it's about**, not just the title. The user needs enough context to decide whether to include it in the public digest. - **⚠️ Summary content MUST come from the SAME data source as the candidate ordering.** Read `candidates/digest-brief.json` first and display candidates in its array order — that's the user's reference ordering. Then, when you need detailed summaries (for IMA notes, Hugo writing, etc.), cross-reference by URL/title, NOT by positional index. Never read `summary-batch.json` items by index and assume they align with `digest-brief.json`'s ordering — they are independently sorted and WILL diverge. - Do **not** display `rank` or `digest_rank` values. - **Use the reported format the user has validated:** numbered list (`1. 2. 3.`) with emoji status indicator (`🟢 keep` / `🟡 review`) and source attribution. Example: ``` 🟢 [keep] 如何在 ChatGPT 中投放广告?|赛博禅心 摘要: 本文详细介绍了ChatGPT广告投放的流程、出价方式... ``` - Convert machine states to Chinese operator-facing labels: - `keep` → `已入选` - `review` → `待确认` - `drop` → `暂不纳入` - For every item under `已入选重点`, include: - title + source - one-sentence summary (max 1 line) - For every item under `待你确认`, include: - title + source - one-line reason why it needs review - brief recommendation - **Do NOT include "worth关注" bullet lists in the internal report.** Three-bullet highlights are for the public digest, not for the operator chat. The user has already seen the article at hand — repeating its highlights is redundant. - **The HARD RULE: never write more than 3-4 sentences per article in the internal chat report.** Anything longer wastes the user's time. At this step: - the public digest is **not** generated yet - Hugo is **not** published yet - the internal review digest is sent to the user in chat - the user decides: **which articles go into the Hugo daily digest** AND separately which articles go into IMA deposition ### Phase 3: User selects Hugo articles The user reviews the internal digest and specifies which articles should appear in the Hugo daily digest. - Confirm the selection explicitly before proceeding. - If the user wants to include some `review` candidate articles, respect that choice. - Only the user-selected articles will appear in the public digest. **⛔ DO NOT SKIP — Article numbering enforcement.** When the user selects articles by position number (e.g. "2, 3, 4, 7"), they are ALWAYS referring to the **order in `candidates/digest-brief.json`'s `top_candidates` array** (1-based). NOT the `item-XX` extracted file numbering, NOT the order in `summary-batch.json`, NOT the payload file. **MANDATORY pre-write verification step (DO NOT SKIP):** 1. Read `candidates/digest-brief.json` and print the array-index order: ```bash python3 -c "import json d = json.load(open('...digest-brief.json')) for i, c in enumerate(d['top_candidates'], 1): print(f'{i}. {c[\"title\"][:60]} | {c[\"source_name\"]}')" ``` 2. This is the authoritative user reference — position 1 = first item in `top_candidates` array. 3. THEN cross-reference by URL for detailed summaries from `summary-batch.json`, but DO NOT use summary-batch ordering for position mapping. 4. Triple-check: user's "article 3" IS `top_candidates[2]` — verify by title match, not by filename. 5. **Before writing the digest, print the final mapping** (user-number → article title) and confirm it matches the user's selection. **Common failure (observed 2026-07-02):** Wrote 0702 digest with wrong articles because position 1 in candidates array was confused with a different article. Always read `candidates` array directly for position mapping. **If the user says "还是错的" after digest is published, re-read the `candidates` array, re-print with positions, and show the user before rewriting.** **⚠️ User may decline to publish:** If the user judges the day's candidate articles as low quality (e.g. too many "资讯" items, lack of substantive technical content), they may say "今天的文章质量低,先不发日报" or similar. **Respect this decision immediately.** Do NOT push back, do NOT ask if they're sure, do NOT suggest alternatives. Skip directly to the end of the flow for that day — do NOT proceed to Phases 4-7. The run still creates useful `digest-brief.json` and term index artifacts; they just don't get published. ### Phase 4: Generate and publish Hugo public digest Generate the public digest markdown **only for the user-selected articles**, then publish to Hugo. **⚠️ 格式基准 — 开始生成任何 digest 内容之前,必须先完整阅读 `references/public-digest-example.md` 并以此为格式基准,不得凭记忆或直觉写作。** **⚠️ HARD PRE-WRITE CHECKLIST — 对照 `references/public-digest-example.md` 逐项确认后再开始写:** 1. [ ] frontmatter `summary` 字段已填写(**完整句子**,不是关键词罗列,参考格式:"围绕 XXX 的当日深度观察。") 2. [ ] 编号只用 `1.` `2.` `3.`,不用中文数字或罗马数字 3. [ ] 四个 section 全部存在:`今日概览` `今日重点` `趋势观察` `延伸阅读` 4. [ ] 每篇 `今日重点` 下有: - [ ] 标题(无"来源:"字样,来源仅在延伸阅读标注) - [ ] 摘要段落 - [ ] "值得关注:"要点列表(每篇 3 条) - [ ] "这篇更值得关注的原因在于:"段落 5. [ ] `tags` 字段已写入 frontmatter(至少 4 个标签,覆盖当日核心主题) - **🔴 优先级规则(2026-07-23 教训):先选已有 tag → 次选 term_index 补充 → 禁止自编** - **第一步**:从已有的通用 tag 池中选出能覆盖今日文章的 tag。已验证池包括但不限于:`RAG`, `Agent工程`, `多Agent`, `AI应用`, `Harness工程`, `LLM`, `MCP`, `Skill`, `Prompt Engineering`, `循环工程`, `上下文管理`, `AI Coding Agent`, `Code Review` 等。 - **第二步**:如果已有 tag 池无法完整覆盖所有文章主题,再从今天 `term_index` 中选取补充 tag。 - **补充 tag 必须过三关**(Agent 工程化 / AI 后端 / LLM 应用)+ **专有名词检查**(不是某个项目的特有命名)。 - **🔴 前置步骤(MANDATORY):** 写 Hugo frontmatter 前,必须先读 keyword engine 输出: ```bash python3 -c " import json d = json.load(open('/home/ubuntu/zhu/github/reader/data/term_index/daily/YYYY-MM-DD.json')) for t in d['terms']: print(t['term']) ``` - **`term_index` 的角色是「防自编验证」而非「唯一来源」**。写 tag 前读出它,是为了检查有没有遗漏的可能主题词,以及确保用词规范性。但最终选哪些 tag,以「通用性优先、覆盖今日主题、复用已有 tag」为准。已有 tag 即使不在今日 term_index 里也照用不误(因为它们是跨日报积累的经过验证的通用概念)。 ``` - 从 term_index 的 `terms` 列表中,选出 4-8 个与当日文章强相关的主题词作为 tags - **🟢 相关度门槛(MANDATORY):** 每个候选 tag 必须跟 Agent 工程化、AI 后端、LLM 应用三者之一直接相关。不相关的一律跳过,哪怕它在 term_index 里。 - **注意:用语义匹配而非精确匹配** — tag `Harness Engineering` 如果在 term_index 里是 `Harness工程化`,语义等价,可保留。同样 `工程化` ↔ `Harness工程化`,`代码质量` ↔ `Code Review` 等。 - ✅ 可选的:`多Agent`, `Multi-Agent架构`, `MCP`, `ACP`, `Harness工程化`, `Loop Engineering`, `Prompt Engineering`, `强化学习`, `AI Agent`, `AI Coding Agent` - ❌ 跳过的:跨境电商, 多语言客服, 国产芯片, 可解释性技术, 多租户, 成本优化, 消息队列, 限流架构, Serverless数据库, AI视频生成, SWE-Bench, 端云协同, 具身智能, 开源(模型发布策略), 浏览器自动化(具体测试手段), 组件化知识库(PagePilot专有术语), 产品名(Kimi K3, K3, Lychee-FD等), 模型名(Fable 5, GPT-5.6 Sol等) - **🔴 硬门槛:优先复用历史已验证的通用 tag,从已有 tag 集合开始选**,而不是每次从 term_index 翻生僻新词。已验证的通用 tag 包括但不限于:`RAG`, `Agent工程`, `多Agent`, `AI应用`, `Harness工程`, `LLM`, `MCP`, `Skill`, `Prompt Engineering`, `循环工程` 等。仅当已有 tag 集合无法覆盖当日文章时,才从 term_index 补充,且补充词必须过三关+专有名词检查。 - **🔴 选出候选 tags 后,必须逐一审查,问一遍「这个 tag 跟 Agent 工程化 / AI 后端 / LLM 应用直接相关吗?」。只要有一个不是,就删掉。宁可 4 个精准的,不要 7 个凑数的。** - **🔴 第二道门(2026-07-22 教训):term_index 里有 ≠ 能用。** 即使 tag 在 term_index 里,如果它只是某篇文章的专有项目名/方法论名(如 `端到端交付2.0`、`页面评论`),不是通用工程概念,也要排除。检查标准:去掉这个 tag,换成通用的上位概念,内容描述还是完整的吗?如果是,说明它是冗余的专有词,排除。 - 关键词引擎已经通过 `configs/term_aliases.json`(120+ 条归一规则)和 `configs/term_stopwords.json`(100+ 条过滤规则)处理过,产出已过滤产品名、模型名、人名等噪声 - 如果 term_index 文件不存在 → 检查 pipeline 是否跑完(`build_keyword_index.py` 作为 pipeline 末环节自动执行) - 如果文件存在但 tags 过少 → 从 `summary-batch.json` 的每篇 `keywords` 中补充手工挑选 **🔴 关键词引擎配置修正流程(mandatory):** 当需要批量增加 aliases/stopwords 时,必须先跑 `build_review_bundle.py` + `generate_term_cleanup_suggestions.py` 查看统计建议,然后再结合语义判断手动改配置。详见 `references/keyword-engine-maintenance.md`。 **🚨 关键词清洗触发条件(HARD RULE):** **只有用户明确说了以下指令时才能触发关键词清洗流程:** - ✅ 「清理关键词」「清理tag」「跑关键词清洗」「词库治理」 - ❌ 以下情况**禁止**触发:发现 tag 质量不高时、发现词库不干净时、无特别指令时 - 两个辅助 skill(`reader-keyword-cleanup` 调度入口 + `reader-keyword-maintenance` 操作手册)仅在用户发出明确指令时加载和使用。 **项目脚本 vs 人工判断的分工:** - ✅ 脚本负责:大小写变体/单复数/空格差异/高频词推荐→用 `apply_term_suggestions.py` 落地 - ❌ 脚本不负责:产品名/模型名/人名的语义识别→需人工改 `term_aliases.json`/`term_stopwords.json` 6. [ ] `延伸阅读` 每条含来源标注:`- [标题](url)|来源` 7. **MANDATORY** [ ] `> ⚠️ 格式规范...` blockquote **已从最终内容中删除**。`public-digest-example.md` 中的 `> 这个 blockquote 是 agent 指令,不是页面内容。复制格式时不小心把它写进正文会暴露到公网。写完文件后必须检查:`grep '格式规范\|生成 Hugo 时必须遵守' ` 确认无匹配。**这条是强制项,遗漏会被用户指出。** Public digest source material: use `candidates/digest-brief.json` — it has all the info needed (title, summary, highlights, source, url). No need to read the full payload. Persist the generated digest under: - `outputs/freshrss/rerun//digest/public_digest.md` Then publish to Hugo: 1. write the public digest markdown to: - `/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md` 2. immediately redeploy Hugo: - `cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh` - **Use foreground `terminal()` with `timeout=120`.** Do NOT use `background=true` or `notify_on_complete=true` — Hugo builds finish in ~10s and produce machine-only Docker build logs. The user has explicitly complained about background notifications pushing raw build output to the platform (Feishu/Telegram). 3. verify before moving on: 3. verify before moving on: - homepage works (`curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:14322/`) - list page works (`curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:14322/daily/`) - **detail page works** (`curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:14322/daily/YYYY-MM-DD/`) - latest digest is visible in the list page 4. **If step (3) detail page returns 404 despite a successful redeploy** — Docker build cache may be stale. Force a clean rebuild: `docker compose build --no-cache && docker compose down --remove-orphans && docker compose up -d --remove-orphans`, then re-run verification step (3). Do not treat "markdown file written" as equivalent to publish success. Hugo publish for this SOP is complete only after redeploy and page verification pass. Do not block on style polish unless the user explicitly asks. The public digest is the browsing layer, not the long-term knowledge layer. It must not expose internal workflow states or operator-facing review labels. Public digest writing rules: - Use article content from `candidates/digest-brief.json` for the user-selected subset. - **Hugo public digest includes ONLY the articles the user selected in Phase 3.** Not all `keep` items — only the user's explicit selection. - Keep the tone suitable for public browsing and Hugo publishing. - Style should follow `references/public-digest-example.md` as the default public-writing example. - Public digest is a **public reading draft / editor-style public note**, not a workflow report. - In `今日概览`, focus on the day's topic lines, shared signals, and broader industry movement; do **not** describe filtering mechanics or internal selection process. - Do **not** expose internal workflow labels or operator language such as `待确认`, `建议沉淀到 IMA`, `keep/review/drop`, or `selection_decision`. - Explicitly avoid wording such as `共筛出`, `候选`, `保留`, `入选`, `待确认`, `建议沉淀` in public digest. - Prefer concise but information-dense writing. - For each item under `今日重点`, include not only summary and highlights, but also one short editor-style value sentence, for example: `这篇内容更值得关注的原因在于……`. - When rendering highlights in public digest, prefer a short label such as `值得关注:` followed by one item per line, instead of packing multiple points into a single long sentence. - In `延伸阅读`, every item must include source attribution in the form: `- [标题](url)|来源`. - **`延伸阅读` MUST contain the selected articles' original links** — one entry per article that was featured in `今日重点`. This is the reader's source attribution section. Do NOT put non-selected articles (like pipeline `drop` items or skipped candidates) in 延伸阅读. Every entry must link back to the original article URL from the pipeline data, not the Hugo page or a summary page. - **Article numbering: MUST use `1.` `2.` `3.` (Arabic numeral + period). Do NOT use `① ② ③`,`一、二、三`,`第一条` or any other numbering variant.** - **All four sections are required: `今日概览`, `今日重点`, `趋势观察`, `延伸阅读`. Missing any section is a format violation.** ### Phase 5: User selects articles for IMA deposition After Hugo publication, ask the user which articles should be retained for long-term knowledge (separate from the Hugo selection). - The user can select from all `keep` and `review` candidates, including or excluding articles that were already put in Hugo. - Confirm the selection explicitly before proceeding. - Only the user-selected articles will be summarized and uploaded to IMA. At this step: - the public digest is already in Hugo - the user decides which articles are worth preserving as knowledge notes - the digest is **not** uploaded to IMA as a whole ### Phase 6: Summarize selected articles For every article explicitly selected by the user: 1. use the `reader` article-summary capability 2. **⚠️ 关键:extracted_path 必须传入 pipeline 产生的 individual item 文件,路径为:** ``` outputs/freshrss/rerun//extracted/item-01.extracted.json outputs/freshrss/rerun//extracted/item-02.extracted.json ... ``` **不要**传 `candidates/openclaw-delivery-payload.json` 或 `outputs/freshrss/extracted/`(这两个路径的 payload 格式与 article-summary workflow 不兼容,会报错** 3. for each selected article, run **one summary job per extracted item file** 4. pass the corresponding `item_id` from the `candidates` array **去掉 `cand:` 前缀** 5. generate one markdown summary per article **⚠️ item_id 截断陷阱(2026-07-17 教训):** 上面 bash 扫描脚本用的 `[:20]` 只显示前 20 字符,**绝对不能拿这个截断值去传 `selected_ids`。** `start_article_summary_job` 的 `selected_ids` 要求完整的 64 字符 item_id(`sha256:xxx...` 整串),截断版本会导致 `ValueError: No extracted entries matched selected_ids`。 **正确做法:** 从 bash 扫描中拿到 item-XX 的映射关系后,**必须再用 python 完整读取一次 `item-XX.extracted.json` 中的完整 `item_id`**,或者将 bash 扫描的 `[:20]` 改为 `[:]` 显示全部 64 字符(但输出会很长)。最稳妥的方案: ```bash # 读取完整 item_id(不截断),配合 grep 提取 python3 -c " import json, glob, os for f in sorted(glob.glob('outputs/freshrss/rerun//extracted/item-*.extracted.json')): d = json.load(open(f)) art = d.get('article', d) print(f'{os.path.basename(f)}: {art[\"item_id\"]}') " ``` **⚠️ item_id 与 extracted 文件编号的对应关系:candidates 数组顺序 ≠ extracted 文件编号顺序。pipeline 提取阶段和 LLM 筛选阶段是两套独立顺序,不能按"第几篇"的位置来映射。** **正确做法:** 传 selected_ids 前,必须先从 `candidates` 数组找到目标文章的 `item_id`(去前缀),再到对应的 `item-XX.extracted.json` 文件里读取其内部的 `article.item_id` 做交叉验证,确认匹配后再传。禁止仅凭"文章在 candidates 里排第几"来推断应该用哪个 item-XX 文件。 **批量扫描技巧:** 用一段 bash 循环一秒扫完所有 item-XX 文件,拿到 title ↔ item_id ↔ item-XX 的完整映射: ```bash workdir=/home/ubuntu/zhu/github/reader for i in 1 2 3 4 5 6 7; do f=$workdir/outputs/freshrss/rerun//extracted/item-$(printf '%02d' $i).extracted.json python3 -c " import json d = json.load(open('$f')) print(f'item-{i:02d}: {d[\"article\"][\"item_id\"][:20]}… | {d[\"article\"][\"title\"][:50]}') " done ``` 输出类似: ``` item-01: sha256:036a5fbad61… | 沙发搬到线上:火山引擎视频云如何用 RTC+直播打造一场"云上陪看房"? item-02: sha256:47cf049738… | Loop Engineering 实践指南:在 Code Buddy 中构建自主循环系统 ... ``` 拿到映射后即可精确传参给 `generate_article_summaries`。 **⚠️ single-item 输入约束:** 当 `extracted_path` 指向 `outputs/freshrss/rerun//extracted/item-XX.extracted.json` 这类单篇文件时,`selected_ids` 只应包含这一个文件对应的单个 `item_id`。不要对单个 extracted 文件传多个 IDs。 Preferred routes: - async MCP job path: - `start_article_summary_job` - `get_article_summary_job_status` - `get_article_summary_job_result` - local fallback: run the reader article-summary workflow directly inside the project `.venv` - CLI fallback: `scripts/run_article_summaries.py` **⚠️ Now available (July 2026):** The CLI path `scripts/run_article_summaries.py` via `.venv/bin/python3` now works — the circular import (`CANDIDATE_BATCH_ARTIFACT` in `freshrss_pipeline.py`) was fixed by making `RunStore` a lazy import. Use the CLI path confidently for article summaries when the MCP async job path is unavailable. **⚠️ generate_article_summaries 调用模式:** 传入单篇 extracted 文件 + 空的 `selected_ids=[]`(而非填 item_id)即可正常触发单篇总结。当 `extracted_path` 是 `item-XX.extracted.json` 且 `selected_ids=[]` 时,工具会自动匹配文件内的 `article.item_id`。 **⚠️ output_dir 必须显式传入:** `generate_article_summaries` 默认将输出写到当前输入文件同级 `single_summaries/` 目录(即 `extracted/single_summaries/`),但本 skill 的标准路径是 `outputs/freshrss/single_summaries/YYYY-MM-DD/`。每次调用时必须显式传 `output_dir` 覆盖默认值。 ```python # 正确模式 generate_article_summaries( extracted_path="outputs/.../extracted/item-XX.extracted.json", selected_ids=[], # 空列表 = 自动匹配 output_dir="outputs/freshrss/single_summaries/YYYY-MM-DD", # ⚠️ 必传!工具默认输出到 extracted/single_summaries/ timeout_seconds=180 ) # 如传 selected_ids,必须使用完整的 64 字符 item_id(sha256:xxx... 整串) ``` **⚠️ async start_article_summary_job 仅接受精简 payload 格式:** 当传入 batch 文件时,顶层必须是 `{"items": [...]}` 数组或单文件 `{"article": {...}}` 结构。不要包裹额外的 `{"success": true, "results": {"items": [...]}}` 层——工具不支持该格式。 **⚠️ async start_article_summary_job 不支持 `selected_ids=[]` 自动匹配:** 与同步 `generate_article_summaries` 不同,异步 `start_article_summary_job` 的 `selected_ids` 参数必须是非空数组。传空数组 `[]` 会立即报错 `selected_ids must not be empty`。使用异步路径时,必须先从 extracted JSON 中提取完整的 64 字符 item_id(`sha256:xxx...`)并显式传入。 **⚠️ COS 上传必须使用 subprocess.run:** 通过 reader-digest-flow 编排 IMA 上传时,调用 `cos-upload.cjs` 脚本必须用 Python `subprocess.run()` 以 args list 方式执行,不能通过 bash `terminal()`。详见 `ima-skill` SKILL.md 中的 COS 上传执行陷阱章节。 Formal production rule: - For normal production deposition, selected-article summary generation should default to the async MCP job path instead of synchronous `generate_article_summaries`. - Start the job, poll status until `success` / `failed`, then read `written_paths` from the job result. - Treat synchronous `generate_article_summaries` as a debug / light-weight helper, not the default production entry. Hard fallback rule:\n- If the async article-summary job path fails because of MCP/tool-layer timeout, transport failure, or job-launch failure, do **not** stop the daily deposition flow.\n- In that case, immediately fall back to running the reader article-summary path locally inside `/home/ubuntu/zhu/github/reader` with the project `.venv`.\n- Treat a successful local article-summary run as equivalent completion for the summary-generation phase; async MCP job is the preferred entry, not a single point of failure.\n- **⚠️ Circular import fixed (July 2026):** The `scripts/run_freshrss_pipeline.py` CLI path now works — the circular import was fixed by making `RunStore` a lazy import (`_get_runstore()`) inside `freshrss_pipeline.py`. The CLI path is now the **preferred production path**, bypassing the MCP transport-layer 120s timeout entirely. **⚠️ Sync pipeline timeout (MCP path only):** `run_freshrss_openclaw_pipeline` defaults to `timeout_seconds=60`, which is often too short. The pipeline involves LLM calls for summarization and can take 2-3 minutes. **Always pass `timeout_seconds=180` explicitly** when calling the synchronous pipeline from a Feishu session. Without this, the MCP transport-level 120s timeout may fire before the pipeline completes, forcing a redundant retry. **Recommended production path: CLI.** Use `scripts/run_freshrss_pipeline.py --limit 7 --include-read --mark-read --timeout 300` via `.venv/bin/python3` — no MCP transport caps, no circular import, full `limit=7` support. Use the real extracted JSON structure already produced by the project. Do not invent alternative inputs. ### Phase 7: Upload selected article notes to IMA Upload only the generated single-article markdown summaries to IMA. **Hard gate before upload:** even if the markdown file was generated by `reader`, do **not** upload it to IMA as-is. You must first reformat/check it against the IMA-facing Markdown rules in this skill, then upload the formatted version. Treat `reader` output as article-summary source material, not automatically as final IMA-ready Markdown. **Hard naming rule before upload:** the final uploaded Markdown filename must use the article's user-facing natural title (normally the original Chinese article title) plus `.md`. Do **not** use internal workflow names, slugs, prefixes, or temp filenames such as `ima-*`, `item-*`, `summary-*`, or English-only shorthand as the final IMA object name. Default target knowledge base for this phase: - `daily` - Read from reader `.env` via `IMA_DAILY_KNOWLEDGE_BASE_ID` and `IMA_DAILY_KNOWLEDGE_BASE_NAME` - Verify the target at runtime through IMA APIs / skill lookups before upload - If the configured target does not exist, try to find `daily` by name; if still absent, create it and continue Completion criteria for this phase: - a selected article has been summarized from extracted content - a single-article Markdown **file** has been written to `/tmp/ima_upload/<标题>.md` - the final upload filename uses the article's user-facing natural Chinese title (not an internal slug / temp name) - the Markdown file is uploaded via the file upload flow: **preflight → create_media → COS → add_knowledge(media_type=7)** - the uploaded object preserves source link context and uses IMA-friendly layout for readability - the upload target is the `daily` knowledge base unless the user explicitly requests otherwise Do not treat the following as completion of summary deposition: - importing the original article webpage URL into IMA - storing the original article only as source material without the generated summary content - creating a note via `import_doc` and linking it with `media_type=11` — **this is explicitly rejected ("笔记格式")** - uploading a Markdown file but using an internal temp name as the file title Do not upload: - the full daily digest - raw payloads - raw extraction output IMA Markdown layout guidance for selected article deposition: - Prefer direct Markdown file upload into the knowledge base (`media_type=7`). - Before every upload, open and check the actual markdown file that will be uploaded; do not assume the generator already matched IMA style. - The upload target must be an IMA-facing formatted markdown file, not the raw default output from `reader` if the styles differ. - **Keep `Source:` and `Category:` metadata headers** near the top of the file (after the title, before the first section). See `references/ima-format-reference.md` for the exact format. - Keep source traceability via the `Source:` metadata line (bare URL, no label prefix). Do NOT add a separate `原文链接:` block — the `Source:` line already carries the original URL. - Each of the 5 sections (`核心结论`, `主要论点`, `关键方法 / 机制`, `重要细节`, `可复用启发`) must be expanded to full-knowledge-point level — not keyword lists. See `references/ima-format-reference.md` for the expected depth and granularity. - Include `## 关键词` and `## 主题` sections at the bottom. - Target total file size: **4,000+ bytes** per article (reference: `ai-agent-的-skill-系统设计.md` is 4,782B). - **When generated summary is below 4,000 bytes (common with WeChat articles where `article.content` is empty):** Expand the content by enriching each of the 5 sections. Pull additional detail from `digest-brief.json`'s `highlights` array, the article `summary` field, and general knowledge about the topic. The goal is deeper explanations per bullet point, not keyword padding. - Break long prose under `核心结论` and `主要论点` into short paragraphs for IMA readability instead of relying on platform auto-formatting. - The final upload filename should normally be `<文章标题>.md`; only when a same-name file already exists should you append a timestamp suffix before `.md`. - If local working files use internal slugs or prefixes for convenience, create or rename a final upload copy before calling IMA upload APIs. - **Format benchmark file (MUST READ before every IMA write):** `references/ima-format-reference.md` — this is the authoritative format template. Do NOT write IMA markdown without reading this file first. ## Operational Guidance - Prefer real run outputs over examples. - Verify at each boundary with real files or accessible URLs. - **When the user asks a question about how the pipeline works (content_source, extraction vs re-fetch, stage semantics), read the relevant reference file first** — `references/content-extraction.md` covers content sourcing, `references/flow.md` covers stage details. Do not answer from memory; the reference files were written to capture the answers exactly. - When validating selected article summaries, confirm that a markdown file is actually generated. - If Codex or another coding agent is asked to implement workflow changes inside `reader`, keep the project boundary clean: - workflow logic in `reader` - orchestration logic in this skill / OpenClaw ## Key Paths ### Reader project - `/home/ubuntu/zhu/github/reader` ### Hugo project - `/home/ubuntu/zhu/apps/hugo-site` - digest content root: - `/home/ubuntu/zhu/apps/hugo-site/content/daily/` ## Pitfalls ### Output path divergence when running via MCP (Feishu context) When the pipeline is triggered via `run_freshrss_openclaw_pipeline` (or the async `start_freshrss_pipeline_job`) **from inside a Feishu session**, two path divergence issues occur: **A) Extracted files diverge:** The extracted article files may be written to an **MCP-managed temp path**, NOT the standard project path `outputs/freshrss/rerun//extracted/`. This means Phase 6's article-summary step (`generate_article_summaries` with `extracted_path=outputs/freshrss/...`) will fail because the files don't exist at the expected project path. **B) output_dir relative paths diverge:** When `generate_article_summaries` receives a **relative** `output_dir` (e.g. `outputs/freshrss/single_summaries/YYYY-MM-DD`), the MCP server resolves it against its own working directory — NOT the reader project root — and may write files to `/root/.hermes/outputs/...` instead of `/home/ubuntu/zhu/github/reader/outputs/...`. The tool's `written_paths` return value will report the resolved path, but the files won't be at the project-relative location you intended. **Workaround:** 1. **CLI mode preferred**: Run the pipeline via CLI terminal (e.g. `hermes terminal`) instead of Feishu, so all files land in the project directory. 2. **Feishu fallback**: If you must run from Feishu: - Use **absolute paths** for `output_dir` (e.g. `/home/ubuntu/zhu/github/reader/outputs/freshrss/single_summaries/2026-06-11`) instead of relative paths - After pipeline completion, check `written_paths` from the result and if files landed at `/root/.hermes/outputs/...`, copy them to the project path: `cp -r /root/.hermes/outputs/freshrss/single_summaries/YYYY-MM-DD /home/ubuntu/zhu/github/reader/outputs/freshrss/single_summaries/` **Root cause**: MCP tools and the Feishu gateway run with a different working directory than CLI Hermes. The reader project's `run_freshrss_pipeline.py` writes to `outputs/freshrss/` which is relative to `workdir`, and the MCP server resolves it differently than the CLI session would. ### Hugo site structure missing (content/ directory) The Hugo site at `/home/ubuntu/zhu/apps/hugo-site/` may NOT have a standard `content/` directory. Before writing any Hugo digest: 1. Verify `content/daily/` exists: ```bash ls -d /home/ubuntu/zhu/apps/hugo-site/content/daily/ 2>/dev/null ``` 2. If missing, create it: ```bash mkdir -p /home/ubuntu/zhu/apps/hugo-site/content/daily/ ``` 3. If the entire `content/` directory is missing, also verify `hugo new site` isn't needed first. **Do not assume the Hugo site has a standard structure.** Always verify before write. ### IMA upload 500 errors IMA's OpenAPI occasionally returns HTTP 500 on the first upload attempt. **Always wrap IMA uploads in a retry loop** (at least 2 attempts, 3-second delay). If `cos-upload.cjs` is involved, capture the subprocess exit code and stderr — a 500 from the upload service may still return exit code 0 if it exits cleanly after the HTTP error. ### IMA upload: `create_media` content_type must omit charset suffix When calling `create_media` for Markdown file uploads, the `content_type` field must be pure MIME type without parameters. The charset suffix (`; charset=utf-8`) causes code 220001 `"invalid media_type"` despite the actual media type being correct. ```json # ✅ Correct "content_type": "text/markdown" # ❌ Fails with code 220001 "invalid media_type" "content_type": "text/markdown; charset=utf-8" ``` This is a documented quirk of the IMA API — it rejects `content_type` values with MIME parameters. Strip them before passing. ### IMA COS credential redaction (Hermes redact_secrets) When `security.redact_secrets: true` in Hermes config (default), `create_media`'s `cos_credential.token` gets replaced with `***` in `terminal()` output, causing COS upload to fail with HTTP 403. **Do NOT pass COS credentials through terminal().** Use `execute_code` + `subprocess.run` — save the raw API response to a file first via terminal(), then read it from Python and call `cos-upload.cjs` with `subprocess.run()`. See `references/ima-cos-credential-handling.md` for the complete workaround. ### IMA upload: chained commands produce concatenated JSON (parse failure) When chaining multiple IMA API calls in one `terminal()` command (`preflight && check_repeated && create_media > file`), the output file contains multiple JSON objects concatenated without delimiters. Subsequent `execute_code` / Python `subprocess.run` calls that read this file will fail with `JSONDecodeError: Extra data` because the file isn't a single valid JSON object. **Observed (2026-07-21):** Running all three steps in one bash line produced a 1509-byte file containing three concatenated JSON responses. Python `json.loads()` on the raw file failed because the first JSON didn't terminate before the second began. **Fix:** Run each IMA API call into its own clean file. Do NOT chain IMA calls that produce output: ```bash # ❌ WRONG — concatenated JSON in output file preflight ... && check_repeated ... && create_media ... > file.json # ✅ CORRECT — each API call gets its own file preflight ...> /dev/null && check_repeated ...> /dev/null create_media ... > /tmp/create_media_clean.json ``` **Exception:** `preflight` and `check_repeated` output is only checked for exit code and `pass`/`is_repeated` fields, so they can safely redirect to `/dev/null`. Only `create_media` output needs to be saved to a file for COS credential extraction. ### ⛔ 铁律:IMA 必须用 **Markdown 文件上传**(media_type=7) **唯一允许的沉淀方式:文件上传流程** 1. 写文章 `.md` 文件到 `/tmp/ima_upload/<标题>.md` 2. `preflight-check.cjs` 验证文件类型 3. `create_media` 获取 COS 凭证 4. `cos-upload.cjs` 上传文件到 COS 5. `add_knowledge` 关联到 daily KB(`media_type=7`, 传 `file_info`) **🔴 铁律:IMA 文件的正文内容必须使用 pipeline 产出的数据,禁止自编摘要** **正文内容来源(优先级从高到低):** 1. `extracted/item-XX.extracted.json` → `article.content`(原始文章正文) 2. `summary-batch.json` → 对应文章的 `summary` 字段(LLM 完整摘要) 3. `digest-brief.json` → `highlights` + `summary`(候选摘要) **⛔ 禁止:** 用自己写的两三句话作为正文。正文长度不足 500 字的直接不合格。 **`add_knowledge` 的 `title` 字段**:必须传完整的原始文章中文标题,不能缩写,不能加 .md 后缀。 **文件名规则:** 必须使用文章**完整原始标题**,不得缩写。后缀为 `.md`。 **格式基准文件(必读):** `references/ima-format-reference.md`(即 `ai-agent-的-skill-系统设计.md`)包含了正确的 IMA 笔记格式模板,每次写 IMA 沉淀前必须先读此文件,严格按其中的内容量、颗粒度和 section 结构生成。 **笔记内容模板(必须包含 5 个 section,且每个 section 需要像 ima-format-reference.md 一样展开到完整知识点级别,不能只列关键词):** ```markdown # 文章完整标题 Source: https://... Category: 分类 ## 核心结论 (一段总结文章核心发现或主张的段落,不是分点) ## 主要论点 (一段概括文章核心论述的段落,综合多个论点形成连贯叙述,不是分点列表) ## 关键方法 / 机制 - **方法名**:详细说明 - **方法名**:详细说明 (基于 pipeline 产出的 highlights 展开,提取核心方法) ## 重要细节 - 每条一个带解释的完整知识点 - 每条一个带解释的完整知识点 (从 pipeline 数据中提取有参考价值的细节信息) ## 可复用启发 - 每条 actionable 的实践启示,附应用场景说明 - 每条 actionable 的实践启示,附应用场景说明 (从文章中提炼可迁移到其他场景的实践启示) ``` - ✅ 正确:`从Vibe Coding到Harness—— 一套大仓AI工程化实战.md` - ❌ 错误:`大仓AI工程化实战.md` - ❌ 错误:`Harness-工程实践.md` **绝对禁止:** - ❌ `import_urls` 传公众号链接 → 用户拒绝的"公众号格式" - ❌ `import_doc` + `media_type=11` → 用户拒绝的"笔记格式",不是真正的 md 文件 - ❌ 跳过摘要直接丢链接 **Do NOT skip the summarization step.** Even when the user says nothing about summaries, the expected format includes a properly structured Markdown note with clear sections — not just an article URL dumped into the KB. The authoritative format template is `references/ima-format-reference.md`. The Hugo redeploy step runs `./redeploy.sh` and completes in ~10 seconds. Using `terminal(background=true, notify_on_complete=true)` causes the Gateway to push the full Docker build output (compiler logs, layer cache hits, nginx config) to Feishu as a notification — machine garbage from the user's perspective. **Fix:** Always use foreground terminal with `timeout=120` for Hugo redeploy. The deploy is fast enough that no background notification is needed. See `references/flow.md` §Publish Hugo for the exact pattern. ### Candidate ordering confusion: digest-brief vs summary-batch **Critical pitfall (observed 2026-07-06/07):** The pipeline produces TWO files with article data, each with its OWN ordering: - `candidates/digest-brief.json` — `top_candidates` array is sorted by **quality/digest_rank** (most relevant first). **This is the user's reference ordering.** - `summary/summary-batch.json` — `items` array is in **raw FreshRSS fetch order** (item-01 ~ item-07). **This order is UNRELATED to digest-brief.json ordering.** These two orderings ALWAYS differ. Reading `summary-batch.json` items by positional index and treating them as matching `digest-brief.json` candidates by position will produce WRONG article assignment — every article's summary will be shifted to the wrong title. **Enforcement rule:** 1. **Display candidates using `digest-brief.json` `top_candidates` array order only.** This is the list the user sees and references by number. 2. **When the user selects articles by number** (e.g. "2,3,4,7"), apply the selection to `digest-brief.json`'s `top_candidates` array — NOT to any other file. 3. **When fetching detailed summaries for selected articles**, match by URL/title cross-reference — never by positional index. Use a lookup like: ```python # Build URL→summary map from summary-batch, then match digest-brief candidates by URL url_to_summary = {} for si in summary_batch['items']: sd = si.get('summary', {}) if isinstance(sd, dict) and sd.get('url'): url_to_summary[sd['url']] = sd ``` 4. **Verify before write:** After mapping selected articles to their summaries, print candidate title + matched summary title side by side. If they don't refer to the same article, the mapping is wrong. **How to detect mismatch before the user does:** Print a verification table: ```python for i, c in enumerate(selected_candidates, 1): sd = url_to_summary.get(c['url'], {}) if sd.get('title') != c['title']: print(f"⚠️ MISMATCH #{i}: candidate[{c['title']}] != summary[{sd.get('title')}]") ``` **One-pipeline one-run rule:** If the pipeline runs only once, there is exactly ONE set of article data. The confusion is entirely between which file's ordering you use for which purpose. Use `digest-brief` ordering for user-facing numbering; cross-reference by URL for all data lookups. Hugo by default does NOT render pages with a `date` value set in the **future** (relative to the server clock). This is silent — no error, no warning, the page just doesn't appear. **Observed failure (2026-06-29):** The digest frontmatter used `date = 2026-06-29T16:55:00+08:00`, but the server time was 09:27 CST. The `hugo list all` command showed the page existed with the correct URL, but no `index.html` was written to the output directory, and the detail page returned 404. **Detection:** ```bash # Check what Hugo actually output docker exec hugo-site ls /usr/share/nginx/html/daily/ | grep YYYY-MM-DD # If the directory is missing despite a clean build, suspect future date # Check server time vs frontmatter date date '+%Y-%m-%d %H:%M:%S %z' ``` **Fix:** Ensure the frontmatter `date` value is in the past, not the future. Use a time slightly before the current moment (e.g. `09:25` when the current time is `09:27`): ```toml title = "AI 日报 · 2026-06-29" date = 2026-06-29T09:25:00+08:00 # ✅ Must be BEFORE current server time ``` **Prevention:** When writing the frontmatter date, always check the server's current time (`date '+%Y-%m-%d %H:%M:%S %z'`) and set the `date` field to a value at least 30 seconds in the past. Never hardcode a time like `16:55` (late afternoon) unless it's truly before the build time. ### Phase 4: Docker build cache not picking up content changes on re-deploy When fixing already-published content (e.g. removing a formatting error from an existing daily page), `./redeploy.sh` may use Docker's builder cache and serve stale HTML — even though the source `.md` files on disk are correct. The build output shows `CACHED` for the COPY and RUN steps. **Detection:** After a `./redeploy.sh`, verify the fix is live by checking the rendered HTML through the Nginx reverse proxy (not just the Docker internal port). Use `curl -s https://osiman.site/daily/YYYY-MM-DD/ | grep 'fixed-text-or-pattern'`. **Root cause:** Docker layer caching. The COPY step detects file changes correctly, but the Hugo build step (`RUN hugo --destination /out`) may re-use a cached output if the COPIED content hash is considered unchanged by Docker BuildKit's internal heuristics. **Fix sequence:** 1. `cd /home/ubuntu/zhu/apps/hugo-site` 2. `docker compose build --no-cache` — force a clean rebuild 3. `docker compose down --remove-orphans && docker compose up -d --remove-orphans` — restart with new image 4. Verify via Nginx proxy URL (not localhost:14322) **⚠️ `docker compose up -d` triggers the tool's long-lived-process detector.** To work around this: run `docker compose build --no-cache` first as a foreground command (it returns when the build completes), then stop and remove the old container (`docker stop hugo-site && docker rm hugo-site`), then start the new container: `docker compose -f /path/to/docker-compose.yml up -d`. Then verify with a separate `terminal()` call checking `curl -s -o /dev/null -w '%{http_code}' http://localhost:14322/daily/YYYY-MM-DD/`. ### Phase 4: Skipping the pre-write checklist (Hugo format drift) The skill's Phase 4 contains a **HARD PRE-WRITE CHECKLIST** with checkboxes referencing `references/public-digest-example.md`. This is not optional or aspirational — it is a **runtime requirement** that must be checked off item by item before writing any Hugo digest. **Common failure (observed 2026-06-10):** The agent reads the checklist but skips reading public-digest-example.md, assuming the format from memory is close enough. This produces a flat article list instead of the required four-section structure (`今日概览` / `今日重点` / `趋势观察` / `延伸阅读`), with missing `summary` field, wrong frontmatter format (`---` instead of `+++`), inline source labels (`**来源:**`) instead of `延伸阅读` attribution, and no `值得关注:` / `这篇更值得关注的原因在于:` per-article structure. **Enforcement rule:** Before writing any Phase 4 output, you MUST: 1. Read `references/public-digest-example.md` in full 2. Copy its exact structure (four sections, TOML frontmatter, `|来源` format, etc.) — do not paraphrase from memory 3. Check off all items in the HARD PRE-WRITE CHECKLIST 4. Only then proceed to write A Hugo digest that does not match the checklist is a format violation and will be corrected by the user. Do not skip this step. ### Keyword engine maintenance: don't bypass the project scripts **🔴 铁律(observed 2026-07-16):当需要清理关键词库时,不要直接改 `term_aliases.json` / `term_stopwords.json`。** 必须按以下流程走: 1. 先跑 `build_review_bundle.py` 重建 bundle 2. 再跑 `generate_term_cleanup_suggestions.py` 看统计建议 3. 再跑 `generate_term_cleanup_semantic_suggestions.py` 用 LLM 找语义问题 4. 最后结合两者的输出,确认后再改配置 不走这个流程的后果:会漏掉项目内置的统计发现(如大小写变体 aliases),也会错过 LLM 发现的语义级问题(如 `AI Harness`→`Harness Engineering`、`自动化`该停用)。 详见 `references/keyword-engine-maintenance.md` 的「两阶段清洗 SOP」。 ### Phase 6: Missing `output_dir` parameter for `generate_article_summaries` The MCP tool `generate_article_summaries` defaults to writing output files in the **`extracted/single_summaries/` subdirectory** of the same run directory as the input files. This is NOT the path the skill expects. **Expected path:** `outputs/freshrss/single_summaries/YYYY-MM-DD/` **Actual default path:** `outputs/freshrss/rerun//extracted/single_summaries/` **Fix:** Pass the `output_dir` parameter explicitly — and use an **absolute path** when running from Feishu to avoid the MCP working-directory divergence (see the "Output path divergence" pitfall): ```python # Correct: absolute path for Feishu context generate_article_summaries( extracted_path="/home/ubuntu/zhu/github/reader/outputs/freshrss/rerun//extracted/item-XX.extracted.json", selected_ids=[], output_dir="/home/ubuntu/zhu/github/reader/outputs/freshrss/single_summaries/2026-06-11", # ABSOLUTE path required from Feishu timeout_seconds=180 ) # Also correct for CLI context (relative works here): generate_article_summaries( extracted_path="outputs/freshrss/rerun//extracted/item-XX.extracted.json", selected_ids=[], output_dir="outputs/freshrss/single_summaries/YYYY-MM-DD", timeout_seconds=180 ) ``` After generation, verify the files exist at the expected path by either checking `written_paths` from the result or listing the target directory with `terminal()`. If running from Feishu and files landed at `/root/.hermes/outputs/...`, copy them to the project path. ### Phase 1: Async pipeline job / resume hangs without error (MCP env isolation) The async `start_freshrss_pipeline_job` runs inside the MCP server process, which operates in its own environment and does NOT automatically source the reader project's `.env` file. This manifests in **two distinct failure patterns**: **Pattern A — Stuck at `generate_summaries` (no error, now fixed):** The linked run shows `status=running`, `current_stage=generate_summaries` with no progress. Previously observed to hang for 8+ hours with no error message. **Root cause (fixed 2026-07-01):** The async pipeline's subprocess (`scripts/run_freshrss_pipeline_job.py`) was crashing at **import time** due to the circular import chain: `freshrss_pipeline.py` → module-level `from summary_mcp.runtime import RunStore` → `runtime.__init__` → `resume_jobs` → `resume_service` → `freshrss_pipeline` (partially initialized). The subprocess's stderr was redirected to `/dev/null`, so the `ImportError` was invisible — the process exited immediately, became a zombie (status `Zs`), and the job status showed `running` forever because `run-state.json` was never updated. **Fix:** `RunStore` import in `freshrss_pipeline.py` changed to lazy import (`_get_runstore()`), breaking the circular chain at its source. Type hints for `RunStore` parameters simplified to untyped parameters. **Post-fix state (verified 2026-07-01):** The async pipeline now completes successfully in ~1-2 minutes. Tested with `limit=3` — `get_freshrss_pipeline_job_result` returns `status: success`. The previous "hang" was entirely caused by the circular import, not by a genuine pipeline stall. **Pattern B — Resume fails with explicit error:** Calling `resume_run` crashes at `write_run_report` with errors like: > Missing required value 'FRESHRSS_API_BASE_URL': not passed as argument, not set as environment variable, and not found in /home/ubuntu/zhu/github/reader/.env. Both patterns share the same root cause: MCP server env isolation — the env vars are correctly set in the project `.env`, but the MCP process doesn't have access to them. **Detection heuristic for Pattern A:** If `get_run_status` shows `completed_stage_count=2` (fetch_feed, extract_articles), `current_stage=generate_summaries`, wait at least **90 seconds** from the `updated_at` timestamp before declaring it stuck. The async pipeline now completes `generate_summaries` within 1-2 minutes (post-circular-import-fix). If `updated_at` has not advanced after 90 seconds, diagnose with: ```bash # Check if the subprocess is alive or a zombie ps aux | grep run_freshrss | grep -v grep # Status 'Z' = zombie (subprocess exited, parent didn't reap) # No match = subprocess already dead, check job-report.json for error # Running with CPU > 0 = still working, wait longer ``` If the subprocess is a zombie or missing, skip resume and fall back immediately to the CLI path. **Fallback for both patterns:** Do NOT retry resume. Check if the subprocess is a zombie first (`ps aux | grep run_freshrss | grep -v grep` — look for status `Z`). If zombie, the subprocess crashed — check `job-report.json` for error. Fall back to the CLI path (`scripts/run_freshrss_pipeline.py` via `.venv/bin/python3` with `--timeout 300`), which now has no circular import and no transport-layer timeout. Only use `run_freshrss_openclaw_pipeline` (MCP sync) when running from Feishu/tool context where CLI is unavailable. **⚠️ Timeout trap in sync pipeline fallback (transport-layer 120s hard cap):** - The `run_freshrss_openclaw_pipeline` MCP tool's `timeout_seconds` parameter defaults to 60 — which is **too short** for production use. - **Critical: the MCP transport layer has its own hard timeout (120s) that `timeout_seconds` cannot override.** Even with `timeout_seconds=180`, if the total end-to-end pipeline time exceeds ~120s, the transport layer kills the call with: `MCP call timed out after 120.0s`. - **Empirical timing (July 2026):** Each article in the sync pipeline takes roughly 20-25 seconds total (fetch + extract + LLM summarize). With `limit=5` the pipeline completes within the 120s window; with `limit=7` it reliably times out at the transport layer. - **Fix:** When falling back to the synchronous pipeline, **reduce `limit` to 5** to stay within the 120s transport window. `timeout_seconds=180` is still recommended to give the pipeline internal breathing room, but it alone cannot solve the transport-layer cap. ```python # Correct production fallback call (reduced limit to fit within 120s transport cap): run_freshrss_openclaw_pipeline( limit=5, # ⚠️ REQUIRED — 7 triggers transport 120s timeout mark_read=True, timeout_seconds=180 ) ``` **Trade-off:** Reducing `limit` means fewer articles per run. For full `limit=7` production runs, the preferred path is now the CLI (`scripts/run_freshrss_pipeline.py`) which has no transport-layer cap. The async MCP path also works (confirmed July 2026) but adds subprocess overhead. When using the MCP sync path (`run_freshrss_openclaw_pipeline`), stick to `limit=5` to fit within the 120s transport window. **⚠️ `include_read=true` as a day-start strategy:** When the unread feed (`include_read=false`, the default) is dominated by sources the extractor cannot handle (WeChat `mp.weixin.qq.com` consistently returns `CONTENT_EXTRACTION_FAILED`), running the first pipeline call of the day with `include_read=true` can unlock a completely different candidate set — older articles from sources the extractor CAN handle. This is not a bug workaround; it is a legitimate morning strategy to evaluate when the first `include_read=false` pass produces few or zero keep/review articles. The trade-off is that previously read articles get marked as read again, which is benign for already-processed content. **Additional note for stuck-run chaining:** If a previous pipeline run from earlier in the same day is still at `generate_summaries` after 60+ seconds, calling `start_freshrss_pipeline_job` will link the new job to the same stuck run. Check `linked_run_id` and its `updated_at` timestamp immediately after starting the async job — if the linked run has not advanced for 60+ seconds, skip resume entirely and go straight to the CLI path fallback. **Actual state (July 2026):** The critical env vars ARE already injected in `/root/.hermes/config.yaml` under `mcp_servers.reader.env` (`LLM_API_KEY`, `LLM_API_URL`, `LLM_MODEL`, etc.). The reader project's `.env` also has matching entries. Subprocesses inherit os.environ by default. **Testing on 2026-07-01 confirmed both paths work:** - **Async MCP path** (`start_freshrss_pipeline_job`): completes in ~1-2 minutes. `get_freshrss_pipeline_job_result` returns `status: success`. The previous "hang" (Pattern A) was confirmed to be the circular import causing the subprocess to crash silently. - **CLI path** (`scripts/run_freshrss_pipeline.py`): verified working with `--limit 7 --include-read --mark-read --timeout 300`, completing in ~24 seconds for 7 articles. **This is the preferred production path** because it has no MCP transport-layer timeout (120s cap) and no subprocess overhead. **The pipeline summarization is parallel** — `ThreadPoolExecutor(max_workers=4)` was added in the 2026-07-01 session. 7 articles complete summarization in ~24 seconds (3-4s per article, 4 concurrent workers). Before this fix, summarization was serial (one article at a time), which exacerbated the perceived "hang" when combined with the circular import crash. **Subprocess env injection (2026-07-01):** The `start_freshrss_pipeline_job` spawns a subprocess via `subprocess.Popen(cmd, cwd=..., stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL, start_new_session=True)`. Without explicit `env=` parameter, the subprocess inherits `os.environ` from the parent MCP server. The Hermes config env block (`mcp_servers.reader.env`) IS correctly inherited. The circular import caused the subprocess to crash at import time (before any env var was read), with stderr going to `/dev/null`. The fix was lazy-importing `RunStore` in `freshrss_pipeline.py` — no explicit env injection into the subprocess was needed because env vars propagate naturally through `os.environ` inheritance. **Debug mode:** Pass `--debug-artifacts` flag to get per-item intermediate files in the output directory — useful for troubleshooting extraction failures or filter decisions. Debug mode adds ~0.5% overhead (just file writes). **Long-term fix:** Track as a deeper investigation. Current workaround (sync pipeline fallback) is reliable. ## References Read `references/content-extraction.md` to understand how the pipeline selects content sources (RSS-only for FreshRSS, never re-fetches original URLs) and what `content_source` values mean for digest quality. Read `references/flow.md` when you need the concrete step-by-step command checklist and file expectations. Read `references/feishu-format-notes.md` for Feishu markdown formatting rules. Read `references/ima-credential-chain.md` for where IMA credentials come from (reader `.env` → `~/.config/ima/` → Hermes `.env`) and recovery when they're missing. Read `references/ima-upload-api.md` for the complete IMA upload API reference (create_media → COS → add_knowledge flow, credentials, curl examples, and markdown reformatting requirements). Read `references/ima-format-reference.md` — **the authoritative IMA note format template (MUST READ before every IMA write).** Read `references/ima-format-quickref.md` before any IMA upload — concrete correct/incorrect examples for filenames, titles, and content source. Read `references/memory-drift-recovery.md` when the memory tool refuses writes due to MEMORY.md file drift — recovery procedure for the recurring `"file on disk has content that wouldn't round-trip"` error. Read `references/keyword-engine-maintenance.md` for the complete keyword engine maintenance workflow — `build_review_bundle.py` → `generate_term_cleanup_suggestions.py` → `generate_term_cleanup_semantic_suggestions.py` → `apply_term_suggestions.py`, plus the distinction between script-driven (statistical) and LLM-driven (semantic) cleanup.