docs: 添加 Agent Skill 到项目仓库

- 复制 reader-digest-flow skill 到 skills/ 目录(含 SKILL.md + references/)
- README 新增 Agent Skill 章节说明供 Agent 使用的工作流
This commit is contained in:
root
2026-07-28 18:38:54 +08:00
parent 7b791ac947
commit 5eb390e3ed
16 changed files with 2112 additions and 0 deletions
+942
View File
@@ -0,0 +1,942 @@
---
name: reader-digest-flow
description: Orchestrate the end-to-end daily digest workflow around the reader project. Use when the user wants to run the AI daily digest flow, report candidates for review, let the user decide which articles go into the Hugo daily digest, then generate and publish the public digest, and optionally summarize selected articles into IMA knowledge notes. Triggers include requests like '跑今天日报', '生成日报', '汇报今天内容', '把选中的文章沉淀', '更新 Hugo', or any request to operate the reader → report → user-selection → Hugo → knowledge-base flow.
---
**🚨 强制前置规则 [MANDATORY] — 每次涉及日报/IMM 沉淀操作前必须执行:**
- **Step 0: 完整加载本 skill(`skill_view(name='reader-digest-flow')`),阅读到本行之后**
- **Step 0.2: 加载 IMA 格式基准文件(`skill_view(name='reader-digest-flow', file_path='references/ima-format-reference.md')`)**
- **Step 0.5: 加载 `ima-skill` 的 knowledge-base 子模块(`skill_view(name='ima-skill', file_path='knowledge-base/SKILL.md')`)**
- 🔴 不读 skill 就干活 = 必错,历史 100% 验证
- 读完 skill 中的「铁律」和引用文件后再动手,禁止凭记忆操作
# Reader Digest Flow
## 🔴 铁律:只跑一次 pipeline,且只能按用户指令跑
### 🚨 唯一合法的触发条件
**只有一种情况可以运行 pipeline:用户明确说了「跑日报」或「再跑一篇」「再跑一次」。**
**禁止因以下任何原因运行 pipeline:**
- ❌ "验证"数据 → 读已有文件
- ❌ "重新确认"候选排序 → 读已有文件
- ❌ "看看会不会有不同结果" → 读已有文件
- ❌ 用户问"为什么第X篇是XX" → 读已有文件
- ❌ 用户质疑编号 → 读已有文件
- ❌ 任何"让我先查一下"的冲动 → **先停,读已有文件**
**用户明确指令(2026-07-07):**
> **"只有当我要求重新跑一次日报时候,才去重新跑一次"**
### 🚨 同一天多次重跑的收敛陷阱(2026-07-27 教训)
当一天内连续请求多次重跑时,FreshRSS feed 会返回同一批文章(因为当天的新内容就那么多),pipeline 排序也趋于稳定。此时**继续重跑不会带来新候选**——不是算法的问题,是原料池已见底。
**处理建议:**
- 如果用户第一次说"质量太低重跑",用 `--include-read` 扩大候选池(skill 已有此规则)
- 如果第二次重跑后仍是同一批 keep 文章,**直接说明情况**:今天 feed 里质量最好的就这些,没有更好的替换候选了
- 不要无限重跑,也不要自作主张降低质量门槛把 review 提成 keep
### 🚨 用户要求的重跑 vs 私自重跑(2026-07-20 教训)
**只有以下两种跑法:**
| 场景 | 是否允许 | 示例 |
|:----|:-------:|:----|
| 用户明确要求"再跑一篇""换一批""重新跑一次" | ✅ **允许** | 老大说"再跑一篇,这批质量一般" |
| 你自己想"验证一下""再看看有没有更好的结果" | ❌ **禁止** | 不能因候选质量不好就自作主张重跑 |
**判断标准:** 只有用户亲口说出明确的重新执行指令才算。你心里觉得"这批不够好"不算。
**重跑后编号规则:** 新批次的候选从 1 重新编号(新 run_id = 新序号体系),与旧批次完全独立。不要在同一个 conversation 里混用两批次的编号。
### 🚨 每次运行必须记录 run ID
Pipeline 跑完后,**立即将 run_id 存入 memory**:
```
memory(action="add", content="当前日报 run_id: 20260707-012841,候选 7 篇")
```
这样后续所有操作都指向同一个数据集,不会产生混淆。
### 🚨 编号规则(用户引用的唯一来源)
**用户的编号永远指向 digest-brief.json 的 top_candidates 数组(1-based index)。**
- 显示给用户的候选编号 = `top_candidates` 数组顺序
- 用户说"第2篇" = `top_candidates[1]`
- **绝对禁止用 `summary-batch.json` 或 `extracted/` 文件编号来理解用户的输入**
### 🚨 候选展示时 **必须带编号**
向用户展示候选列表时,必须打印 `digest-brief.json` 中 `top_candidates` 数组的编号和标题,并且给用户选择的编号必须与这个编号一致。
正确做法:
```bash
python3 -c "import json
d = json.load(open('...digest-brief.json'))
for i, c in enumerate(d['top_candidates'], 1):
print(f'{i}. {c[\"title\"][:60]} | {c[\"source_name\"]}')"
```
错误做法:
- ❌ 按其他文件(如 summary-batch.json)的顺序展示
- ❌ 只展示标题不带编号
- ❌ 展示编号但与 digest-brief.json 的数组索引不匹配
### 🚨 永远不要重新分组/重编号候选 — 保持原始 pipeline 序号
**教训(2026-07-20):用户质问「为什么待确认,你给我的序号是重新排?」**
错误演示(按 keep/review 分组后各自编号 1. 2. …):
```
🟢 已入选(4 篇)
1. PagePilot ...
2. Harness ...
🟡 待确认(2 篇)
1. Kimi K3 ...
2. Kimi K3 不同来源...
```
**铁律:**
1. **永远维护原始编号。** 候选列表的编号是 `enumerate(top_candidates, 1)` 的结果,用户只认这个编号体系。
2. **不要按 keep/review 分组各自编号。** 如果分组展示,每篇候选也必须保留其在 `top_candidates` 数组中的原始序号。
3. **推荐的展示方式:** 平铺编号列表,每行用 emoji 状态指示器(🟢已入选 / 🟡待确认)标记决策状态,不改变序号。
4. **重新跑 pipeline 后:** 新批次的候选从 1 重新编号(新 run_id = 新序号体系),与旧批次完全独立。
### 🚨 当用户说"还是错的"时
1. **不跑新的 pipeline**
2. **立即打印 digest-brief.json 的 top_candidates 编号+标题,展示给用户确认**
3. 用 URL 交叉引用匹配候选与摘要
4. 如果对不上,承认错误
### 匹配候选与其摘要的唯一正确方式:按 URL 交叉引用
绝不按数组位置索引把 `summary-batch.json` 和 `digest-brief.json` 对齐——它们用不同的排序逻辑,位置从来不匹配。
### 🚨 沟通纪律:只说用户问的
**跟 pipeline 技术同等重要。** 用户问了一个具体问题,只回答那个问题。禁止:
- ❌ 主动分析"为什么会发生这种情况的历史原因"
- ❌ 补充"顺便说一下,以前也…"
- ❌ 解释用户没问的场景的细节
- ❌ 在道歉/认错后面跟长段分析
**用户原话(2026-07-07):**
> "为什么你还要提关于 7-1 的 16 次原因????我没有再次问你阿???"
**正确做法:** 用户问「为什么第二是AICR」→ 答「我跑了两遍pipeline,排序变了」→ 停止。用户没问的就不说。
### 🚨 HARD PRE-UPLOAD CHECKLIST — IMA Phase 7
Before uploading any article to IMA, you MUST run through this checklist. A single unchecked item means DO NOT UPLOAD — fix it first.
1. [ ] **skill 已加载** — `skill_view(name='reader-digest-flow')` 已执行并阅读到文件末尾
2. [ ] **ima-skill 已加载** — `skill_view(name='ima-skill', file_path='knowledge-base/SKILL.md')` 已执行
3. [ ] **格式基准文件已读** — `skill_view(name='reader-digest-flow', file_path='references/ima-format-reference.md')` 已执行并理解其格式要求
4. [ ] **使用 media_type=7** — 代码使用 `preflight → create_media → COS → add_knowledge(media_type=7)`,不是 `import_doc`(media_type=11)
5. [ ] **文件名 = 完整原始标题** — 文件名为 `<完整原始文章标题>.md`,如 `从Vibe Coding到Harness—— 一套大仓AI工程化实战.md`
6. [ ] **文件名未缩写** — 对比 pipeline 中 `digest-brief.json` 的 `title` 字段与上传文件名,必须完全一致(不含 .md 后缀比对)
7. [ ] **内容来自 pipeline 数据** — 正文从 `extracted/item-XX.extracted.json` 的 `article.content` 或 `summary-batch.json` 的 `summary` 字段读取
8. [ ] **内容不小于 500 字** — 正文长度 >= 500 字符(中文字算1字)
9. [ ] **`add_knowledge` 的 `title` 字段正确** — 传完整标题,不加 .md 后缀
10. [ ] **文件内容符合 ima-format-reference.md 标准** — 必须包含:顶部 `Source:`/`Category:` 元数据、5 个 section、底部 `关键词`/`主题`。其中 **核心结论和主要论点是段落**(非分点),**关键方法/机制、重要细节、可复用启发是分点**。每个 section 展开到完整知识点级别。目标文件大小 **4,000+ bytes**
**执行前必须逐项打勾**,少一项就停。历史教训(2026-07-08):两次 IMA 上传都踩了这些坑——先是 media_type=11 笔记格式,后是文件名缩写+自编摘要。pre-upload checklist 就是为了防这个。
## Core Rules for fetching, extraction, filtering, payload generation, result reading, and selected-article post-processing.
- **Daily digest production runs should default to a random article count between 5 and 10 unless the user explicitly specifies a count.**
- Prefer **CLI path** (`scripts/run_freshrss_pipeline.py` via `.venv/bin/python3`) for production runs — it bypasses the MCP transport-layer 120s timeout and has no circular import. Use MCP workflow operations only for inline Feishu/tool queries where CLI is unavailable.
- **Main FreshRSS production runs should default to the async MCP job path**: `start_freshrss_pipeline_job` → `get_freshrss_pipeline_job_status` → `get_freshrss_pipeline_job_result`. Use the synchronous `run_freshrss_openclaw_pipeline` only for debug / light validation / fallback.
- **If the main pipeline job fails, do not immediately abandon the run.** First inspect the linked run through `get_run_status` and `inspect_resume_plan`; when the plan returns `recommended_action=resume`, continue through `start_resume_job` → `get_resume_job_status` → `get_resume_job_result`.
- **Operational note from validation:** before the main pipeline async job existed, a debug/test run could still complete successfully inside reader even when the synchronous MCP wrapper returned timeout. For historical/debug cases, continue from real run truth using MCP status/result query tools rather than treating the whole flow as failed.
- For reader run observation and result reading, prefer MCP tools such as status / payload / report queries instead of having OpenClaw or this skill hand-build reader output paths.
- **When reader status APIs return reconciled status, always branch on the top-level `status`.** Treat `status_source` and `state_conflict` as explanatory metadata; do not re-derive flow control from stale stage names.
- **For normal production runs, mark processed FreshRSS items as read. Only skip mark-read when the user explicitly says the run is debug/test/validation.**
- **For normal production runs, do not pass `debug_artifacts=true`. Only enable debug artifacts when the user explicitly says the run is debug/test/validation or when troubleshooting is the goal.**
- **Do not generate the daily digest from examples or placeholder data. Always require a real payload first.**
- Generate the daily digest markdown from the payload in OpenClaw.
- Split outputs into a **public digest** for Hugo and an **internal review digest** for chat / operator decision-making.
- **Public digest and internal review digest are orchestration-layer outputs, and by default are generated by the current OpenClaw session model rather than inheriting reader `LLM_*` settings.**
- **Publish only the public digest to Hugo, and only after the user has explicitly selected which articles go into the daily digest.**
- The default flow is: pipeline → report candidates → **user selects Hugo articles** → generate public digest → publish Hugo → (optional) user selects IMA articles → IMA deposition.
- Do NOT generate the public digest or publish Hugo before the user has confirmed the Hugo article selection.
- Report the internal review digest back to the user in chat.
- **Do not upload the full daily digest to IMA.**
- **Do not upload any selected article summary to IMA until the user has explicitly confirmed the selection. Once the user has confirmed which articles to keep, proceed directly with IMA deposition and do not ask for a second confirmation about uploading to the knowledge base.**
- Upload **only user-selected articles** to IMA as individual knowledge notes.
- **The formal IMA upload object must be the generated single-article summary content, uploaded directly into the target IMA knowledge base as a Markdown knowledge item (`media_type=7`), not the original article webpage URL.**
- **For daily deposition, the target must explicitly be the IMA "knowledge-base" type path, not the IMA "notes" type path.**
- **Default deposition: Markdown 文件上传(media_type=7)**
- 必须用完整的文件上传流程: **preflight → create_media → COS upload → add_knowledge(media_type=7)**
- 先将文章内容写入本地 `.md` 文件,再通过文件上传流程提交
- 文件命名规则:`<文章标题>.md`(使用用户界面看到的原始中文标题)
- **Do NOT use `import_doc` + `media_type=11`** — the user explicitly rejected this format (called it "笔记格式"). 只有 `media_type=7`(Markdown 文件上传)才是正确的沉淀格式。
- **Do NOT use `import_urls`** — the user explicitly rejected this (called it "公众号格式").
- **Adding content as an IMA note, or uploading to any non-knowledge-base container first and then linking it indirectly, does not count as completing the SOP. Completion requires direct knowledge-base ingestion.**
- Default IMA target for daily single-article summaries is the `daily` knowledge base.
- Resolve that default target from reader `.env` (`IMA_DAILY_KNOWLEDGE_BASE_ID`, `IMA_DAILY_KNOWLEDGE_BASE_NAME`) and verify it at runtime before upload.
- If the configured daily knowledge base is missing, try to locate it by name; if still missing, create `daily` and continue.
- For selected article notes, use the dedicated article-summary flow in `reader`.
- Prefer `ARTICLE_SUMMARY_*` LLM settings for selected article summaries; fall back to the main `LLM_*` settings only if the dedicated settings are absent.
- Do not re-fetch original article URLs for selected summaries; always use the existing extracted article text.
- Default selected-article summary output should be organized by date, for example under `outputs/freshrss/single_summaries/YYYY-MM-DD/`.
## Fixed Flow
### Phase 1: Run reader pipeline
In the `reader` project, run the FreshRSS pipeline and obtain a real payload.
Formal production start path:
- `start_freshrss_pipeline_job`
- `get_freshrss_pipeline_job_status`
- `get_freshrss_pipeline_job_result`
After the async job succeeds, treat the returned `run_id` as the stable handle for downstream `get_run_status` / `get_delivery_payload` / `get_run_report` reads.
Minimum expected artifacts:
- `outputs/freshrss/rerun/<run-id>/candidates/digest-brief.json` — **精简候选清单**(唯一要读取的文件,含 keep + review)
- `outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json` — 完整 payload(pipeline 内部产物,digest-brief 由此推导,流程中无需读取)
- `outputs/freshrss/rerun/<run-id>/run-report.json`
- extracted article data: `outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json` 等(每个候选一篇,格式为 `{"success": true, "article": {...}}`,Phase 6 使用)
If the main pipeline job fails:
- read the linked run through `get_run_status`
- call `inspect_resume_plan(run_id)`
- if `recommended_action=resume`, continue with:
- `start_resume_job`
- `get_resume_job_status`
- `get_resume_job_result`
- if `recommended_action=read_terminal_result`, continue from the terminal run result instead of retrying
- if `recommended_action=start_new_run`, stop and report the exact failure point
- **if the resume fails with `Missing required value` errors** (e.g. `FRESHRSS_API_BASE_URL`), this is MCP server environment isolation — the MCP process doesn't source `.env`. **Do NOT retry resume.** Fall back immediately to the synchronous `run_freshrss_openclaw_pipeline`, which handles env isolation differently and completes successfully.
### Phase 2: Report candidates to the user (internal review digest only)
**⛔ USER PREFERENCE (2026-07-01): Do NOT re-summarize the project background. The user has run this flow many times. Report ONLY the day's candidates — no pipeline overview, no project context, no "as you know" framing. If the previous pipeline results were already reported earlier in the same conversation, omit the candidate details too — just report the status and any anomalies. The user said "你已经向我总结了3次了" — respect this.**
After the pipeline run succeeds, do NOT generate the public digest or publish Hugo yet.
Read candidates and produce **only the internal review digest** in a concise format (not the full payload dump).
Recommended **internal review digest** structure:
- `今日候选概况` (count table + one-liner per article)
- `已入选重点` (only keep articles)
- `待你确认` (only review articles)
Data source: use `candidates/digest-brief.json` (`top_candidates` array) — it now includes both keep and review articles in a concise, simplified format (title, source, summary, highlights, category). No need to read the full payload.
**⚠️ 每篇文章简述长度要求:** 每篇文章的描述必须包含 2-3 句话(含核心问题、核心方法/结论、业务价值/效果),不能只有一句话或关键词堆砌。老大需要足够的信息来判断哪些文章值得发布和沉淀。示例:
> **🟢 [已入选] 得物推荐系统诊断 Agent:「推查查」** | 得物技术
> 得物自研"推查查"诊断Agent,融合Highway确定性流水线与ATV自主推理双模式,实现推荐系统异常快速定位与复杂问题根因分析。核心创新是将排查经验通过Skill/Story原子化封装与进化层持续沉淀,推动排查从人工经验向自动化智能转变。该方案在得物生产环境中验证了从小时级定位缩短到分钟级的效果。
Internal review digest writing rules:
- **Keep this as a terse bullet-style report.** Do not write essay-length paragraphs.
- **Each article must include a 1-2 line summary describing what it's about**, not just the title. The user needs enough context to decide whether to include it in the public digest.
- **⚠️ Summary content MUST come from the SAME data source as the candidate ordering.** Read `candidates/digest-brief.json` first and display candidates in its array order — that's the user's reference ordering. Then, when you need detailed summaries (for IMA notes, Hugo writing, etc.), cross-reference by URL/title, NOT by positional index. Never read `summary-batch.json` items by index and assume they align with `digest-brief.json`'s ordering — they are independently sorted and WILL diverge.
- Do **not** display `rank` or `digest_rank` values.
- **Use the reported format the user has validated:** numbered list (`1. 2. 3.`) with emoji status indicator (`🟢 keep` / `🟡 review`) and source attribution. Example:
```
🟢 [keep] 如何在 ChatGPT 中投放广告?|赛博禅心
摘要: 本文详细介绍了ChatGPT广告投放的流程、出价方式...
```
- Convert machine states to Chinese operator-facing labels:
- `keep` → `已入选`
- `review` → `待确认`
- `drop` → `暂不纳入`
- For every item under `已入选重点`, include:
- title + source
- one-sentence summary (max 1 line)
- For every item under `待你确认`, include:
- title + source
- one-line reason why it needs review
- brief recommendation
- **Do NOT include "worth关注" bullet lists in the internal report.** Three-bullet highlights are for the public digest, not for the operator chat. The user has already seen the article at hand — repeating its highlights is redundant.
- **The HARD RULE: never write more than 3-4 sentences per article in the internal chat report.** Anything longer wastes the user's time.
At this step:
- the public digest is **not** generated yet
- Hugo is **not** published yet
- the internal review digest is sent to the user in chat
- the user decides: **which articles go into the Hugo daily digest** AND separately which articles go into IMA deposition
### Phase 3: User selects Hugo articles
The user reviews the internal digest and specifies which articles should appear in the Hugo daily digest.
- Confirm the selection explicitly before proceeding.
- If the user wants to include some `review` candidate articles, respect that choice.
- Only the user-selected articles will appear in the public digest.
**⛔ DO NOT SKIP — Article numbering enforcement.** When the user selects articles by position number (e.g. "2, 3, 4, 7"), they are ALWAYS referring to the **order in `candidates/digest-brief.json`'s `top_candidates` array** (1-based). NOT the `item-XX` extracted file numbering, NOT the order in `summary-batch.json`, NOT the payload file.
**MANDATORY pre-write verification step (DO NOT SKIP):**
1. Read `candidates/digest-brief.json` and print the array-index order:
```bash
python3 -c "import json
d = json.load(open('...digest-brief.json'))
for i, c in enumerate(d['top_candidates'], 1):
print(f'{i}. {c[\"title\"][:60]} | {c[\"source_name\"]}')"
```
2. This is the authoritative user reference — position 1 = first item in `top_candidates` array.
3. THEN cross-reference by URL for detailed summaries from `summary-batch.json`, but DO NOT use summary-batch ordering for position mapping.
4. Triple-check: user's "article 3" IS `top_candidates[2]` — verify by title match, not by filename.
5. **Before writing the digest, print the final mapping** (user-number → article title) and confirm it matches the user's selection.
**Common failure (observed 2026-07-02):** Wrote 0702 digest with wrong articles because position 1 in candidates array was confused with a different article. Always read `candidates` array directly for position mapping.
**If the user says "还是错的" after digest is published, re-read the `candidates` array, re-print with positions, and show the user before rewriting.**
**⚠️ User may decline to publish:** If the user judges the day's candidate articles as low quality (e.g. too many "资讯" items, lack of substantive technical content), they may say "今天的文章质量低,先不发日报" or similar. **Respect this decision immediately.** Do NOT push back, do NOT ask if they're sure, do NOT suggest alternatives. Skip directly to the end of the flow for that day — do NOT proceed to Phases 4-7. The run still creates useful `digest-brief.json` and term index artifacts; they just don't get published.
### Phase 4: Generate and publish Hugo public digest
Generate the public digest markdown **only for the user-selected articles**, then publish to Hugo.
**⚠️ 格式基准 — 开始生成任何 digest 内容之前,必须先完整阅读 `references/public-digest-example.md` 并以此为格式基准,不得凭记忆或直觉写作。**
**⚠️ HARD PRE-WRITE CHECKLIST — 对照 `references/public-digest-example.md` 逐项确认后再开始写:**
1. [ ] frontmatter `summary` 字段已填写(**完整句子**,不是关键词罗列,参考格式:"围绕 XXX 的当日深度观察。")
2. [ ] 编号只用 `1.` `2.` `3.`,不用中文数字或罗马数字
3. [ ] 四个 section 全部存在:`今日概览` `今日重点` `趋势观察` `延伸阅读`
4. [ ] 每篇 `今日重点` 下有:
- [ ] 标题(无"来源:"字样,来源仅在延伸阅读标注)
- [ ] 摘要段落
- [ ] "值得关注:"要点列表(每篇 3 条)
- [ ] "这篇更值得关注的原因在于:"段落
5. [ ] `tags` 字段已写入 frontmatter(至少 4 个标签,覆盖当日核心主题)
- **🔴 优先级规则(2026-07-23 教训):先选已有 tag → 次选 term_index 补充 → 禁止自编**
- **第一步**:从已有的通用 tag 池中选出能覆盖今日文章的 tag。已验证池包括但不限于:`RAG`, `Agent工程`, `多Agent`, `AI应用`, `Harness工程`, `LLM`, `MCP`, `Skill`, `Prompt Engineering`, `循环工程`, `上下文管理`, `AI Coding Agent`, `Code Review` 等。
- **第二步**:如果已有 tag 池无法完整覆盖所有文章主题,再从今天 `term_index` 中选取补充 tag。
- **补充 tag 必须过三关**(Agent 工程化 / AI 后端 / LLM 应用)+ **专有名词检查**(不是某个项目的特有命名)。
- **🔴 前置步骤(MANDATORY):** 写 Hugo frontmatter 前,必须先读 keyword engine 输出:
```bash
python3 -c "
import json
d = json.load(open('/home/ubuntu/zhu/github/reader/data/term_index/daily/YYYY-MM-DD.json'))
for t in d['terms']:
print(t['term'])
```
- **`term_index` 的角色是「防自编验证」而非「唯一来源」**。写 tag 前读出它,是为了检查有没有遗漏的可能主题词,以及确保用词规范性。但最终选哪些 tag,以「通用性优先、覆盖今日主题、复用已有 tag」为准。已有 tag 即使不在今日 term_index 里也照用不误(因为它们是跨日报积累的经过验证的通用概念)。
```
- 从 term_index 的 `terms` 列表中,选出 4-8 个与当日文章强相关的主题词作为 tags
- **🟢 相关度门槛(MANDATORY):** 每个候选 tag 必须跟 Agent 工程化、AI 后端、LLM 应用三者之一直接相关。不相关的一律跳过,哪怕它在 term_index 里。
- **注意:用语义匹配而非精确匹配** — tag `Harness Engineering` 如果在 term_index 里是 `Harness工程化`,语义等价,可保留。同样 `工程化` ↔ `Harness工程化`,`代码质量` ↔ `Code Review` 等。
- ✅ 可选的:`多Agent`, `Multi-Agent架构`, `MCP`, `ACP`, `Harness工程化`, `Loop Engineering`, `Prompt Engineering`, `强化学习`, `AI Agent`, `AI Coding Agent`
- ❌ 跳过的:跨境电商, 多语言客服, 国产芯片, 可解释性技术, 多租户, 成本优化, 消息队列, 限流架构, Serverless数据库, AI视频生成, SWE-Bench, 端云协同, 具身智能, 开源(模型发布策略), 浏览器自动化(具体测试手段), 组件化知识库(PagePilot专有术语), 产品名(Kimi K3, K3, Lychee-FD等), 模型名(Fable 5, GPT-5.6 Sol等)
- **🔴 硬门槛:优先复用历史已验证的通用 tag,从已有 tag 集合开始选**,而不是每次从 term_index 翻生僻新词。已验证的通用 tag 包括但不限于:`RAG`, `Agent工程`, `多Agent`, `AI应用`, `Harness工程`, `LLM`, `MCP`, `Skill`, `Prompt Engineering`, `循环工程` 等。仅当已有 tag 集合无法覆盖当日文章时,才从 term_index 补充,且补充词必须过三关+专有名词检查。
- **🔴 选出候选 tags 后,必须逐一审查,问一遍「这个 tag 跟 Agent 工程化 / AI 后端 / LLM 应用直接相关吗?」。只要有一个不是,就删掉。宁可 4 个精准的,不要 7 个凑数的。**
- **🔴 第二道门(2026-07-22 教训):term_index 里有 ≠ 能用。** 即使 tag 在 term_index 里,如果它只是某篇文章的专有项目名/方法论名(如 `端到端交付2.0`、`页面评论`),不是通用工程概念,也要排除。检查标准:去掉这个 tag,换成通用的上位概念,内容描述还是完整的吗?如果是,说明它是冗余的专有词,排除。
- 关键词引擎已经通过 `configs/term_aliases.json`(120+ 条归一规则)和 `configs/term_stopwords.json`(100+ 条过滤规则)处理过,产出已过滤产品名、模型名、人名等噪声
- 如果 term_index 文件不存在 → 检查 pipeline 是否跑完(`build_keyword_index.py` 作为 pipeline 末环节自动执行)
- 如果文件存在但 tags 过少 → 从 `summary-batch.json` 的每篇 `keywords` 中补充手工挑选
**🔴 关键词引擎配置修正流程(mandatory):** 当需要批量增加 aliases/stopwords 时,必须先跑 `build_review_bundle.py` + `generate_term_cleanup_suggestions.py` 查看统计建议,然后再结合语义判断手动改配置。详见 `references/keyword-engine-maintenance.md`。
**🚨 关键词清洗触发条件(HARD RULE):**
**只有用户明确说了以下指令时才能触发关键词清洗流程:**
- ✅ 「清理关键词」「清理tag」「跑关键词清洗」「词库治理」
- ❌ 以下情况**禁止**触发:发现 tag 质量不高时、发现词库不干净时、无特别指令时
- 两个辅助 skill(`reader-keyword-cleanup` 调度入口 + `reader-keyword-maintenance` 操作手册)仅在用户发出明确指令时加载和使用。
**项目脚本 vs 人工判断的分工:**
- ✅ 脚本负责:大小写变体/单复数/空格差异/高频词推荐→用 `apply_term_suggestions.py` 落地
- ❌ 脚本不负责:产品名/模型名/人名的语义识别→需人工改 `term_aliases.json`/`term_stopwords.json`
6. [ ] `延伸阅读` 每条含来源标注:`- [标题](url)|来源`
7. **MANDATORY** [ ] `> ⚠️ 格式规范...` blockquote **已从最终内容中删除**。`public-digest-example.md` 中的 `> 这个 blockquote 是 agent 指令,不是页面内容。复制格式时不小心把它写进正文会暴露到公网。写完文件后必须检查:`grep '格式规范\|生成 Hugo 时必须遵守' <output-file>` 确认无匹配。**这条是强制项,遗漏会被用户指出。**
Public digest source material: use `candidates/digest-brief.json` — it has all the info needed (title, summary, highlights, source, url). No need to read the full payload.
Persist the generated digest under:
- `outputs/freshrss/rerun/<run-id>/digest/public_digest.md`
Then publish to Hugo:
1. write the public digest markdown to:
- `/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md`
2. immediately redeploy Hugo:
- `cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh`
- **Use foreground `terminal()` with `timeout=120`.** Do NOT use `background=true` or `notify_on_complete=true` — Hugo builds finish in ~10s and produce machine-only Docker build logs. The user has explicitly complained about background notifications pushing raw build output to the platform (Feishu/Telegram).
3. verify before moving on:
3. verify before moving on:
- homepage works (`curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:14322/`)
- list page works (`curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:14322/daily/`)
- **detail page works** (`curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:14322/daily/YYYY-MM-DD/`)
- latest digest is visible in the list page
4. **If step (3) detail page returns 404 despite a successful redeploy** — Docker build cache may be stale. Force a clean rebuild: `docker compose build --no-cache && docker compose down --remove-orphans && docker compose up -d --remove-orphans`, then re-run verification step (3).
Do not treat "markdown file written" as equivalent to publish success. Hugo publish for this SOP is complete only after redeploy and page verification pass.
Do not block on style polish unless the user explicitly asks.
The public digest is the browsing layer, not the long-term knowledge layer.
It must not expose internal workflow states or operator-facing review labels.
Public digest writing rules:
- Use article content from `candidates/digest-brief.json` for the user-selected subset.
- **Hugo public digest includes ONLY the articles the user selected in Phase 3.** Not all `keep` items — only the user's explicit selection.
- Keep the tone suitable for public browsing and Hugo publishing.
- Style should follow `references/public-digest-example.md` as the default public-writing example.
- Public digest is a **public reading draft / editor-style public note**, not a workflow report.
- In `今日概览`, focus on the day's topic lines, shared signals, and broader industry movement; do **not** describe filtering mechanics or internal selection process.
- Do **not** expose internal workflow labels or operator language such as `待确认`, `建议沉淀到 IMA`, `keep/review/drop`, or `selection_decision`.
- Explicitly avoid wording such as `共筛出`, `候选`, `保留`, `入选`, `待确认`, `建议沉淀` in public digest.
- Prefer concise but information-dense writing.
- For each item under `今日重点`, include not only summary and highlights, but also one short editor-style value sentence, for example: `这篇内容更值得关注的原因在于……`.
- When rendering highlights in public digest, prefer a short label such as `值得关注:` followed by one item per line, instead of packing multiple points into a single long sentence.
- In `延伸阅读`, every item must include source attribution in the form: `- [标题](url)|来源`.
- **`延伸阅读` MUST contain the selected articles' original links** — one entry per article that was featured in `今日重点`. This is the reader's source attribution section. Do NOT put non-selected articles (like pipeline `drop` items or skipped candidates) in 延伸阅读. Every entry must link back to the original article URL from the pipeline data, not the Hugo page or a summary page.
- **Article numbering: MUST use `1.` `2.` `3.` (Arabic numeral + period). Do NOT use `① ② ③`,`一、二、三`,`第一条` or any other numbering variant.**
- **All four sections are required: `今日概览`, `今日重点`, `趋势观察`, `延伸阅读`. Missing any section is a format violation.**
### Phase 5: User selects articles for IMA deposition
After Hugo publication, ask the user which articles should be retained for long-term knowledge (separate from the Hugo selection).
- The user can select from all `keep` and `review` candidates, including or excluding articles that were already put in Hugo.
- Confirm the selection explicitly before proceeding.
- Only the user-selected articles will be summarized and uploaded to IMA.
At this step:
- the public digest is already in Hugo
- the user decides which articles are worth preserving as knowledge notes
- the digest is **not** uploaded to IMA as a whole
### Phase 6: Summarize selected articles
For every article explicitly selected by the user:
1. use the `reader` article-summary capability
2. **⚠️ 关键:extracted_path 必须传入 pipeline 产生的 individual item 文件,路径为:**
```
outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json
outputs/freshrss/rerun/<run-id>/extracted/item-02.extracted.json
...
```
**不要**传 `candidates/openclaw-delivery-payload.json` 或 `outputs/freshrss/extracted/`(这两个路径的 payload 格式与 article-summary workflow 不兼容,会报错**
3. for each selected article, run **one summary job per extracted item file**
4. pass the corresponding `item_id` from the `candidates` array **去掉 `cand:` 前缀**
5. generate one markdown summary per article
**⚠️ item_id 截断陷阱(2026-07-17 教训):** 上面 bash 扫描脚本用的 `[:20]` 只显示前 20 字符,**绝对不能拿这个截断值去传 `selected_ids`。** `start_article_summary_job` 的 `selected_ids` 要求完整的 64 字符 item_id(`sha256:xxx...` 整串),截断版本会导致 `ValueError: No extracted entries matched selected_ids`。
**正确做法:** 从 bash 扫描中拿到 item-XX 的映射关系后,**必须再用 python 完整读取一次 `item-XX.extracted.json` 中的完整 `item_id`**,或者将 bash 扫描的 `[:20]` 改为 `[:]` 显示全部 64 字符(但输出会很长)。最稳妥的方案:
```bash
# 读取完整 item_id(不截断),配合 grep 提取
python3 -c "
import json, glob, os
for f in sorted(glob.glob('outputs/freshrss/rerun/<run-id>/extracted/item-*.extracted.json')):
d = json.load(open(f))
art = d.get('article', d)
print(f'{os.path.basename(f)}: {art[\"item_id\"]}')
"
```
**⚠️ item_id 与 extracted 文件编号的对应关系:candidates 数组顺序 ≠ extracted 文件编号顺序。pipeline 提取阶段和 LLM 筛选阶段是两套独立顺序,不能按"第几篇"的位置来映射。**
**正确做法:** 传 selected_ids 前,必须先从 `candidates` 数组找到目标文章的 `item_id`(去前缀),再到对应的 `item-XX.extracted.json` 文件里读取其内部的 `article.item_id` 做交叉验证,确认匹配后再传。禁止仅凭"文章在 candidates 里排第几"来推断应该用哪个 item-XX 文件。
**批量扫描技巧:** 用一段 bash 循环一秒扫完所有 item-XX 文件,拿到 title ↔ item_id ↔ item-XX 的完整映射:
```bash
workdir=/home/ubuntu/zhu/github/reader
for i in 1 2 3 4 5 6 7; do
f=$workdir/outputs/freshrss/rerun/<run-id>/extracted/item-$(printf '%02d' $i).extracted.json
python3 -c "
import json
d = json.load(open('$f'))
print(f'item-{i:02d}: {d[\"article\"][\"item_id\"][:20]}… | {d[\"article\"][\"title\"][:50]}')
"
done
```
输出类似:
```
item-01: sha256:036a5fbad61… | 沙发搬到线上:火山引擎视频云如何用 RTC+直播打造一场"云上陪看房"?
item-02: sha256:47cf049738… | Loop Engineering 实践指南:在 Code Buddy 中构建自主循环系统
...
```
拿到映射后即可精确传参给 `generate_article_summaries`。
**⚠️ single-item 输入约束:** 当 `extracted_path` 指向 `outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json` 这类单篇文件时,`selected_ids` 只应包含这一个文件对应的单个 `item_id`。不要对单个 extracted 文件传多个 IDs。
Preferred routes:
- async MCP job path:
- `start_article_summary_job`
- `get_article_summary_job_status`
- `get_article_summary_job_result`
- local fallback: run the reader article-summary workflow directly inside the project `.venv`
- CLI fallback: `scripts/run_article_summaries.py`
**⚠️ Now available (July 2026):** The CLI path `scripts/run_article_summaries.py` via `.venv/bin/python3` now works — the circular import (`CANDIDATE_BATCH_ARTIFACT` in `freshrss_pipeline.py`) was fixed by making `RunStore` a lazy import. Use the CLI path confidently for article summaries when the MCP async job path is unavailable.
**⚠️ generate_article_summaries 调用模式:** 传入单篇 extracted 文件 + 空的 `selected_ids=[]`(而非填 item_id)即可正常触发单篇总结。当 `extracted_path` 是 `item-XX.extracted.json` 且 `selected_ids=[]` 时,工具会自动匹配文件内的 `article.item_id`。
**⚠️ output_dir 必须显式传入:** `generate_article_summaries` 默认将输出写到当前输入文件同级 `single_summaries/` 目录(即 `extracted/single_summaries/`),但本 skill 的标准路径是 `outputs/freshrss/single_summaries/YYYY-MM-DD/`。每次调用时必须显式传 `output_dir` 覆盖默认值。
```python
# 正确模式
generate_article_summaries(
extracted_path="outputs/.../extracted/item-XX.extracted.json",
selected_ids=[], # 空列表 = 自动匹配
output_dir="outputs/freshrss/single_summaries/YYYY-MM-DD", # ⚠️ 必传!工具默认输出到 extracted/single_summaries/
timeout_seconds=180
)
# 如传 selected_ids,必须使用完整的 64 字符 item_id(sha256:xxx... 整串)
```
**⚠️ async start_article_summary_job 仅接受精简 payload 格式:** 当传入 batch 文件时,顶层必须是 `{"items": [...]}` 数组或单文件 `{"article": {...}}` 结构。不要包裹额外的 `{"success": true, "results": {"items": [...]}}` 层——工具不支持该格式。
**⚠️ async start_article_summary_job 不支持 `selected_ids=[]` 自动匹配:** 与同步 `generate_article_summaries` 不同,异步 `start_article_summary_job` 的 `selected_ids` 参数必须是非空数组。传空数组 `[]` 会立即报错 `selected_ids must not be empty`。使用异步路径时,必须先从 extracted JSON 中提取完整的 64 字符 item_id(`sha256:xxx...`)并显式传入。
**⚠️ COS 上传必须使用 subprocess.run:** 通过 reader-digest-flow 编排 IMA 上传时,调用 `cos-upload.cjs` 脚本必须用 Python `subprocess.run()` 以 args list 方式执行,不能通过 bash `terminal()`。详见 `ima-skill` SKILL.md 中的 COS 上传执行陷阱章节。
Formal production rule:
- For normal production deposition, selected-article summary generation should default to the async MCP job path instead of synchronous `generate_article_summaries`.
- Start the job, poll status until `success` / `failed`, then read `written_paths` from the job result.
- Treat synchronous `generate_article_summaries` as a debug / light-weight helper, not the default production entry.
Hard fallback rule:\n- If the async article-summary job path fails because of MCP/tool-layer timeout, transport failure, or job-launch failure, do **not** stop the daily deposition flow.\n- In that case, immediately fall back to running the reader article-summary path locally inside `/home/ubuntu/zhu/github/reader` with the project `.venv`.\n- Treat a successful local article-summary run as equivalent completion for the summary-generation phase; async MCP job is the preferred entry, not a single point of failure.\n- **⚠️ Circular import fixed (July 2026):** The `scripts/run_freshrss_pipeline.py` CLI path now works — the circular import was fixed by making `RunStore` a lazy import (`_get_runstore()`) inside `freshrss_pipeline.py`. The CLI path is now the **preferred production path**, bypassing the MCP transport-layer 120s timeout entirely.
**⚠️ Sync pipeline timeout (MCP path only):** `run_freshrss_openclaw_pipeline` defaults to `timeout_seconds=60`, which is often too short. The pipeline involves LLM calls for summarization and can take 2-3 minutes. **Always pass `timeout_seconds=180` explicitly** when calling the synchronous pipeline from a Feishu session. Without this, the MCP transport-level 120s timeout may fire before the pipeline completes, forcing a redundant retry.
**Recommended production path: CLI.** Use `scripts/run_freshrss_pipeline.py --limit 7 --include-read --mark-read --timeout 300` via `.venv/bin/python3` — no MCP transport caps, no circular import, full `limit=7` support.
Use the real extracted JSON structure already produced by the project. Do not invent alternative inputs.
### Phase 7: Upload selected article notes to IMA
Upload only the generated single-article markdown summaries to IMA.
**Hard gate before upload:** even if the markdown file was generated by `reader`, do **not** upload it to IMA as-is. You must first reformat/check it against the IMA-facing Markdown rules in this skill, then upload the formatted version. Treat `reader` output as article-summary source material, not automatically as final IMA-ready Markdown.
**Hard naming rule before upload:** the final uploaded Markdown filename must use the article's user-facing natural title (normally the original Chinese article title) plus `.md`. Do **not** use internal workflow names, slugs, prefixes, or temp filenames such as `ima-*`, `item-*`, `summary-*`, or English-only shorthand as the final IMA object name.
Default target knowledge base for this phase:
- `daily`
- Read from reader `.env` via `IMA_DAILY_KNOWLEDGE_BASE_ID` and `IMA_DAILY_KNOWLEDGE_BASE_NAME`
- Verify the target at runtime through IMA APIs / skill lookups before upload
- If the configured target does not exist, try to find `daily` by name; if still absent, create it and continue
Completion criteria for this phase:
- a selected article has been summarized from extracted content
- a single-article Markdown **file** has been written to `/tmp/ima_upload/<标题>.md`
- the final upload filename uses the article's user-facing natural Chinese title (not an internal slug / temp name)
- the Markdown file is uploaded via the file upload flow: **preflight → create_media → COS → add_knowledge(media_type=7)**
- the uploaded object preserves source link context and uses IMA-friendly layout for readability
- the upload target is the `daily` knowledge base unless the user explicitly requests otherwise
Do not treat the following as completion of summary deposition:
- importing the original article webpage URL into IMA
- storing the original article only as source material without the generated summary content
- creating a note via `import_doc` and linking it with `media_type=11` — **this is explicitly rejected ("笔记格式")**
- uploading a Markdown file but using an internal temp name as the file title
Do not upload:
- the full daily digest
- raw payloads
- raw extraction output
IMA Markdown layout guidance for selected article deposition:
- Prefer direct Markdown file upload into the knowledge base (`media_type=7`).
- Before every upload, open and check the actual markdown file that will be uploaded; do not assume the generator already matched IMA style.
- The upload target must be an IMA-facing formatted markdown file, not the raw default output from `reader` if the styles differ.
- **Keep `Source:` and `Category:` metadata headers** near the top of the file (after the title, before the first section). See `references/ima-format-reference.md` for the exact format.
- Keep source traceability via the `Source:` metadata line (bare URL, no label prefix). Do NOT add a separate `原文链接:` block — the `Source:` line already carries the original URL.
- Each of the 5 sections (`核心结论`, `主要论点`, `关键方法 / 机制`, `重要细节`, `可复用启发`) must be expanded to full-knowledge-point level — not keyword lists. See `references/ima-format-reference.md` for the expected depth and granularity.
- Include `## 关键词` and `## 主题` sections at the bottom.
- Target total file size: **4,000+ bytes** per article (reference: `ai-agent-的-skill-系统设计.md` is 4,782B).
- **When generated summary is below 4,000 bytes (common with WeChat articles where `article.content` is empty):** Expand the content by enriching each of the 5 sections. Pull additional detail from `digest-brief.json`'s `highlights` array, the article `summary` field, and general knowledge about the topic. The goal is deeper explanations per bullet point, not keyword padding.
- Break long prose under `核心结论` and `主要论点` into short paragraphs for IMA readability instead of relying on platform auto-formatting.
- The final upload filename should normally be `<文章标题>.md`; only when a same-name file already exists should you append a timestamp suffix before `.md`.
- If local working files use internal slugs or prefixes for convenience, create or rename a final upload copy before calling IMA upload APIs.
- **Format benchmark file (MUST READ before every IMA write):** `references/ima-format-reference.md` — this is the authoritative format template. Do NOT write IMA markdown without reading this file first.
## Operational Guidance
- Prefer real run outputs over examples.
- Verify at each boundary with real files or accessible URLs.
- **When the user asks a question about how the pipeline works (content_source, extraction vs re-fetch, stage semantics), read the relevant reference file first** — `references/content-extraction.md` covers content sourcing, `references/flow.md` covers stage details. Do not answer from memory; the reference files were written to capture the answers exactly.
- When validating selected article summaries, confirm that a markdown file is actually generated.
- If Codex or another coding agent is asked to implement workflow changes inside `reader`, keep the project boundary clean:
- workflow logic in `reader`
- orchestration logic in this skill / OpenClaw
## Key Paths
### Reader project
- `/home/ubuntu/zhu/github/reader`
### Hugo project
- `/home/ubuntu/zhu/apps/hugo-site`
- digest content root:
- `/home/ubuntu/zhu/apps/hugo-site/content/daily/`
## Pitfalls
### Output path divergence when running via MCP (Feishu context)
When the pipeline is triggered via `run_freshrss_openclaw_pipeline` (or the async `start_freshrss_pipeline_job`) **from inside a Feishu session**, two path divergence issues occur:
**A) Extracted files diverge:** The extracted article files may be written to an **MCP-managed temp path**, NOT the standard project path `outputs/freshrss/rerun/<run-id>/extracted/`. This means Phase 6's article-summary step (`generate_article_summaries` with `extracted_path=outputs/freshrss/...`) will fail because the files don't exist at the expected project path.
**B) output_dir relative paths diverge:** When `generate_article_summaries` receives a **relative** `output_dir` (e.g. `outputs/freshrss/single_summaries/YYYY-MM-DD`), the MCP server resolves it against its own working directory — NOT the reader project root — and may write files to `/root/.hermes/outputs/...` instead of `/home/ubuntu/zhu/github/reader/outputs/...`. The tool's `written_paths` return value will report the resolved path, but the files won't be at the project-relative location you intended.
**Workaround:**
1. **CLI mode preferred**: Run the pipeline via CLI terminal (e.g. `hermes terminal`) instead of Feishu, so all files land in the project directory.
2. **Feishu fallback**: If you must run from Feishu:
- Use **absolute paths** for `output_dir` (e.g. `/home/ubuntu/zhu/github/reader/outputs/freshrss/single_summaries/2026-06-11`) instead of relative paths
- After pipeline completion, check `written_paths` from the result and if files landed at `/root/.hermes/outputs/...`, copy them to the project path: `cp -r /root/.hermes/outputs/freshrss/single_summaries/YYYY-MM-DD /home/ubuntu/zhu/github/reader/outputs/freshrss/single_summaries/`
**Root cause**: MCP tools and the Feishu gateway run with a different working directory than CLI Hermes. The reader project's `run_freshrss_pipeline.py` writes to `outputs/freshrss/` which is relative to `workdir`, and the MCP server resolves it differently than the CLI session would.
### Hugo site structure missing (content/ directory)
The Hugo site at `/home/ubuntu/zhu/apps/hugo-site/` may NOT have a standard `content/` directory. Before writing any Hugo digest:
1. Verify `content/daily/` exists:
```bash
ls -d /home/ubuntu/zhu/apps/hugo-site/content/daily/ 2>/dev/null
```
2. If missing, create it:
```bash
mkdir -p /home/ubuntu/zhu/apps/hugo-site/content/daily/
```
3. If the entire `content/` directory is missing, also verify `hugo new site` isn't needed first.
**Do not assume the Hugo site has a standard structure.** Always verify before write.
### IMA upload 500 errors
IMA's OpenAPI occasionally returns HTTP 500 on the first upload attempt. **Always wrap IMA uploads in a retry loop** (at least 2 attempts, 3-second delay). If `cos-upload.cjs` is involved, capture the subprocess exit code and stderr — a 500 from the upload service may still return exit code 0 if it exits cleanly after the HTTP error.
### IMA upload: `create_media` content_type must omit charset suffix
When calling `create_media` for Markdown file uploads, the `content_type` field must be pure MIME type without parameters. The charset suffix (`; charset=utf-8`) causes code 220001 `"invalid media_type"` despite the actual media type being correct.
```json
# ✅ Correct
"content_type": "text/markdown"
# ❌ Fails with code 220001 "invalid media_type"
"content_type": "text/markdown; charset=utf-8"
```
This is a documented quirk of the IMA API — it rejects `content_type` values with MIME parameters. Strip them before passing.
### IMA COS credential redaction (Hermes redact_secrets)
When `security.redact_secrets: true` in Hermes config (default), `create_media`'s `cos_credential.token` gets replaced with `***` in `terminal()` output, causing COS upload to fail with HTTP 403.
**Do NOT pass COS credentials through terminal().** Use `execute_code` + `subprocess.run` — save the raw API response to a file first via terminal(), then read it from Python and call `cos-upload.cjs` with `subprocess.run()`.
See `references/ima-cos-credential-handling.md` for the complete workaround.
### IMA upload: chained commands produce concatenated JSON (parse failure)
When chaining multiple IMA API calls in one `terminal()` command (`preflight && check_repeated && create_media > file`), the output file contains multiple JSON objects concatenated without delimiters. Subsequent `execute_code` / Python `subprocess.run` calls that read this file will fail with `JSONDecodeError: Extra data` because the file isn't a single valid JSON object.
**Observed (2026-07-21):** Running all three steps in one bash line produced a 1509-byte file containing three concatenated JSON responses. Python `json.loads()` on the raw file failed because the first JSON didn't terminate before the second began.
**Fix:** Run each IMA API call into its own clean file. Do NOT chain IMA calls that produce output:
```bash
# ❌ WRONG — concatenated JSON in output file
preflight ... && check_repeated ... && create_media ... > file.json
# ✅ CORRECT — each API call gets its own file
preflight ...> /dev/null && check_repeated ...> /dev/null
create_media ... > /tmp/create_media_clean.json
```
**Exception:** `preflight` and `check_repeated` output is only checked for exit code and `pass`/`is_repeated` fields, so they can safely redirect to `/dev/null`. Only `create_media` output needs to be saved to a file for COS credential extraction.
### ⛔ 铁律:IMA 必须用 **Markdown 文件上传**(media_type=7)
**唯一允许的沉淀方式:文件上传流程**
1. 写文章 `.md` 文件到 `/tmp/ima_upload/<标题>.md`
2. `preflight-check.cjs` 验证文件类型
3. `create_media` 获取 COS 凭证
4. `cos-upload.cjs` 上传文件到 COS
5. `add_knowledge` 关联到 daily KB(`media_type=7`, 传 `file_info`)
**🔴 铁律:IMA 文件的正文内容必须使用 pipeline 产出的数据,禁止自编摘要**
**正文内容来源(优先级从高到低):**
1. `extracted/item-XX.extracted.json` → `article.content`(原始文章正文)
2. `summary-batch.json` → 对应文章的 `summary` 字段(LLM 完整摘要)
3. `digest-brief.json` → `highlights` + `summary`(候选摘要)
**⛔ 禁止:** 用自己写的两三句话作为正文。正文长度不足 500 字的直接不合格。
**`add_knowledge` 的 `title` 字段**:必须传完整的原始文章中文标题,不能缩写,不能加 .md 后缀。
**文件名规则:** 必须使用文章**完整原始标题**,不得缩写。后缀为 `.md`。
**格式基准文件(必读):** `references/ima-format-reference.md`(即 `ai-agent-的-skill-系统设计.md`)包含了正确的 IMA 笔记格式模板,每次写 IMA 沉淀前必须先读此文件,严格按其中的内容量、颗粒度和 section 结构生成。
**笔记内容模板(必须包含 5 个 section,且每个 section 需要像 ima-format-reference.md 一样展开到完整知识点级别,不能只列关键词):**
```markdown
# 文章完整标题
Source: https://...
Category: 分类
## 核心结论
(一段总结文章核心发现或主张的段落,不是分点)
## 主要论点
(一段概括文章核心论述的段落,综合多个论点形成连贯叙述,不是分点列表)
## 关键方法 / 机制
- **方法名**:详细说明
- **方法名**:详细说明
(基于 pipeline 产出的 highlights 展开,提取核心方法)
## 重要细节
- 每条一个带解释的完整知识点
- 每条一个带解释的完整知识点
(从 pipeline 数据中提取有参考价值的细节信息)
## 可复用启发
- 每条 actionable 的实践启示,附应用场景说明
- 每条 actionable 的实践启示,附应用场景说明
(从文章中提炼可迁移到其他场景的实践启示)
```
- ✅ 正确:`从Vibe Coding到Harness—— 一套大仓AI工程化实战.md`
- ❌ 错误:`大仓AI工程化实战.md`
- ❌ 错误:`Harness-工程实践.md`
**绝对禁止:**
- ❌ `import_urls` 传公众号链接 → 用户拒绝的"公众号格式"
- ❌ `import_doc` + `media_type=11` → 用户拒绝的"笔记格式",不是真正的 md 文件
- ❌ 跳过摘要直接丢链接
**Do NOT skip the summarization step.** Even when the user says nothing about summaries, the expected format includes a properly structured Markdown note with clear sections — not just an article URL dumped into the KB. The authoritative format template is `references/ima-format-reference.md`.
The Hugo redeploy step runs `./redeploy.sh` and completes in ~10 seconds. Using `terminal(background=true, notify_on_complete=true)` causes the Gateway to push the full Docker build output (compiler logs, layer cache hits, nginx config) to Feishu as a notification — machine garbage from the user's perspective.
**Fix:** Always use foreground terminal with `timeout=120` for Hugo redeploy. The deploy is fast enough that no background notification is needed. See `references/flow.md` §Publish Hugo for the exact pattern.
### Candidate ordering confusion: digest-brief vs summary-batch
**Critical pitfall (observed 2026-07-06/07):** The pipeline produces TWO files with article data, each with its OWN ordering:
- `candidates/digest-brief.json` — `top_candidates` array is sorted by **quality/digest_rank** (most relevant first). **This is the user's reference ordering.**
- `summary/summary-batch.json` — `items` array is in **raw FreshRSS fetch order** (item-01 ~ item-07). **This order is UNRELATED to digest-brief.json ordering.**
These two orderings ALWAYS differ. Reading `summary-batch.json` items by positional index and treating them as matching `digest-brief.json` candidates by position will produce WRONG article assignment — every article's summary will be shifted to the wrong title.
**Enforcement rule:**
1. **Display candidates using `digest-brief.json` `top_candidates` array order only.** This is the list the user sees and references by number.
2. **When the user selects articles by number** (e.g. "2,3,4,7"), apply the selection to `digest-brief.json`'s `top_candidates` array — NOT to any other file.
3. **When fetching detailed summaries for selected articles**, match by URL/title cross-reference — never by positional index. Use a lookup like:
```python
# Build URL→summary map from summary-batch, then match digest-brief candidates by URL
url_to_summary = {}
for si in summary_batch['items']:
sd = si.get('summary', {})
if isinstance(sd, dict) and sd.get('url'):
url_to_summary[sd['url']] = sd
```
4. **Verify before write:** After mapping selected articles to their summaries, print candidate title + matched summary title side by side. If they don't refer to the same article, the mapping is wrong.
**How to detect mismatch before the user does:** Print a verification table:
```python
for i, c in enumerate(selected_candidates, 1):
sd = url_to_summary.get(c['url'], {})
if sd.get('title') != c['title']:
print(f"⚠️ MISMATCH #{i}: candidate[{c['title']}] != summary[{sd.get('title')}]")
```
**One-pipeline one-run rule:** If the pipeline runs only once, there is exactly ONE set of article data. The confusion is entirely between which file's ordering you use for which purpose. Use `digest-brief` ordering for user-facing numbering; cross-reference by URL for all data lookups.
Hugo by default does NOT render pages with a `date` value set in the **future** (relative to the server clock). This is silent — no error, no warning, the page just doesn't appear.
**Observed failure (2026-06-29):** The digest frontmatter used `date = 2026-06-29T16:55:00+08:00`, but the server time was 09:27 CST. The `hugo list all` command showed the page existed with the correct URL, but no `index.html` was written to the output directory, and the detail page returned 404.
**Detection:**
```bash
# Check what Hugo actually output
docker exec hugo-site ls /usr/share/nginx/html/daily/ | grep YYYY-MM-DD
# If the directory is missing despite a clean build, suspect future date
# Check server time vs frontmatter date
date '+%Y-%m-%d %H:%M:%S %z'
```
**Fix:** Ensure the frontmatter `date` value is in the past, not the future. Use a time slightly before the current moment (e.g. `09:25` when the current time is `09:27`):
```toml
title = "AI 日报 · 2026-06-29"
date = 2026-06-29T09:25:00+08:00 # ✅ Must be BEFORE current server time
```
**Prevention:** When writing the frontmatter date, always check the server's current time (`date '+%Y-%m-%d %H:%M:%S %z'`) and set the `date` field to a value at least 30 seconds in the past. Never hardcode a time like `16:55` (late afternoon) unless it's truly before the build time.
### Phase 4: Docker build cache not picking up content changes on re-deploy
When fixing already-published content (e.g. removing a formatting error from an existing daily page), `./redeploy.sh` may use Docker's builder cache and serve stale HTML — even though the source `.md` files on disk are correct. The build output shows `CACHED` for the COPY and RUN steps.
**Detection:** After a `./redeploy.sh`, verify the fix is live by checking the rendered HTML through the Nginx reverse proxy (not just the Docker internal port). Use `curl -s https://osiman.site/daily/YYYY-MM-DD/ | grep 'fixed-text-or-pattern'`.
**Root cause:** Docker layer caching. The COPY step detects file changes correctly, but the Hugo build step (`RUN hugo --destination /out`) may re-use a cached output if the COPIED content hash is considered unchanged by Docker BuildKit's internal heuristics.
**Fix sequence:**
1. `cd /home/ubuntu/zhu/apps/hugo-site`
2. `docker compose build --no-cache` — force a clean rebuild
3. `docker compose down --remove-orphans && docker compose up -d --remove-orphans` — restart with new image
4. Verify via Nginx proxy URL (not localhost:14322)
**⚠️ `docker compose up -d` triggers the tool's long-lived-process detector.** To work around this: run `docker compose build --no-cache` first as a foreground command (it returns when the build completes), then stop and remove the old container (`docker stop hugo-site && docker rm hugo-site`), then start the new container: `docker compose -f /path/to/docker-compose.yml up -d`. Then verify with a separate `terminal()` call checking `curl -s -o /dev/null -w '%{http_code}' http://localhost:14322/daily/YYYY-MM-DD/`.
### Phase 4: Skipping the pre-write checklist (Hugo format drift)
The skill's Phase 4 contains a **HARD PRE-WRITE CHECKLIST** with checkboxes referencing `references/public-digest-example.md`. This is not optional or aspirational — it is a **runtime requirement** that must be checked off item by item before writing any Hugo digest.
**Common failure (observed 2026-06-10):** The agent reads the checklist but skips reading public-digest-example.md, assuming the format from memory is close enough. This produces a flat article list instead of the required four-section structure (`今日概览` / `今日重点` / `趋势观察` / `延伸阅读`), with missing `summary` field, wrong frontmatter format (`---` instead of `+++`), inline source labels (`**来源:**`) instead of `延伸阅读` attribution, and no `值得关注:` / `这篇更值得关注的原因在于:` per-article structure.
**Enforcement rule:** Before writing any Phase 4 output, you MUST:
1. Read `references/public-digest-example.md` in full
2. Copy its exact structure (four sections, TOML frontmatter, `|来源` format, etc.) — do not paraphrase from memory
3. Check off all items in the HARD PRE-WRITE CHECKLIST
4. Only then proceed to write
A Hugo digest that does not match the checklist is a format violation and will be corrected by the user. Do not skip this step.
### Keyword engine maintenance: don't bypass the project scripts
**🔴 铁律(observed 2026-07-16):当需要清理关键词库时,不要直接改 `term_aliases.json` / `term_stopwords.json`。** 必须按以下流程走:
1. 先跑 `build_review_bundle.py` 重建 bundle
2. 再跑 `generate_term_cleanup_suggestions.py` 看统计建议
3. 再跑 `generate_term_cleanup_semantic_suggestions.py` 用 LLM 找语义问题
4. 最后结合两者的输出,确认后再改配置
不走这个流程的后果:会漏掉项目内置的统计发现(如大小写变体 aliases),也会错过 LLM 发现的语义级问题(如 `AI Harness`→`Harness Engineering`、`自动化`该停用)。
详见 `references/keyword-engine-maintenance.md` 的「两阶段清洗 SOP」。
### Phase 6: Missing `output_dir` parameter for `generate_article_summaries`
The MCP tool `generate_article_summaries` defaults to writing output files in the **`extracted/single_summaries/` subdirectory** of the same run directory as the input files. This is NOT the path the skill expects.
**Expected path:** `outputs/freshrss/single_summaries/YYYY-MM-DD/`
**Actual default path:** `outputs/freshrss/rerun/<run-id>/extracted/single_summaries/`
**Fix:** Pass the `output_dir` parameter explicitly — and use an **absolute path** when running from Feishu to avoid the MCP working-directory divergence (see the "Output path divergence" pitfall):
```python
# Correct: absolute path for Feishu context
generate_article_summaries(
extracted_path="/home/ubuntu/zhu/github/reader/outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json",
selected_ids=[],
output_dir="/home/ubuntu/zhu/github/reader/outputs/freshrss/single_summaries/2026-06-11", # ABSOLUTE path required from Feishu
timeout_seconds=180
)
# Also correct for CLI context (relative works here):
generate_article_summaries(
extracted_path="outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json",
selected_ids=[],
output_dir="outputs/freshrss/single_summaries/YYYY-MM-DD",
timeout_seconds=180
)
```
After generation, verify the files exist at the expected path by either checking `written_paths` from the result or listing the target directory with `terminal()`. If running from Feishu and files landed at `/root/.hermes/outputs/...`, copy them to the project path.
### Phase 1: Async pipeline job / resume hangs without error (MCP env isolation)
The async `start_freshrss_pipeline_job` runs inside the MCP server process, which operates in its own environment and does NOT automatically source the reader project's `.env` file. This manifests in **two distinct failure patterns**:
**Pattern A — Stuck at `generate_summaries` (no error, now fixed):** The linked run shows `status=running`, `current_stage=generate_summaries` with no progress. Previously observed to hang for 8+ hours with no error message.
**Root cause (fixed 2026-07-01):** The async pipeline's subprocess (`scripts/run_freshrss_pipeline_job.py`) was crashing at **import time** due to the circular import chain: `freshrss_pipeline.py` → module-level `from summary_mcp.runtime import RunStore` → `runtime.__init__` → `resume_jobs` → `resume_service` → `freshrss_pipeline` (partially initialized). The subprocess's stderr was redirected to `/dev/null`, so the `ImportError` was invisible — the process exited immediately, became a zombie (status `Zs`), and the job status showed `running` forever because `run-state.json` was never updated.
**Fix:** `RunStore` import in `freshrss_pipeline.py` changed to lazy import (`_get_runstore()`), breaking the circular chain at its source. Type hints for `RunStore` parameters simplified to untyped parameters.
**Post-fix state (verified 2026-07-01):** The async pipeline now completes successfully in ~1-2 minutes. Tested with `limit=3` — `get_freshrss_pipeline_job_result` returns `status: success`. The previous "hang" was entirely caused by the circular import, not by a genuine pipeline stall.
**Pattern B — Resume fails with explicit error:** Calling `resume_run` crashes at `write_run_report` with errors like:
> Missing required value 'FRESHRSS_API_BASE_URL': not passed as argument, not set as environment variable, and not found in /home/ubuntu/zhu/github/reader/.env.
Both patterns share the same root cause: MCP server env isolation — the env vars are correctly set in the project `.env`, but the MCP process doesn't have access to them.
**Detection heuristic for Pattern A:** If `get_run_status` shows `completed_stage_count=2` (fetch_feed, extract_articles), `current_stage=generate_summaries`, wait at least **90 seconds** from the `updated_at` timestamp before declaring it stuck. The async pipeline now completes `generate_summaries` within 1-2 minutes (post-circular-import-fix). If `updated_at` has not advanced after 90 seconds, diagnose with:
```bash
# Check if the subprocess is alive or a zombie
ps aux | grep run_freshrss | grep -v grep
# Status 'Z' = zombie (subprocess exited, parent didn't reap)
# No match = subprocess already dead, check job-report.json for error
# Running with CPU > 0 = still working, wait longer
```
If the subprocess is a zombie or missing, skip resume and fall back immediately to the CLI path.
**Fallback for both patterns:** Do NOT retry resume. Check if the subprocess is a zombie first (`ps aux | grep run_freshrss | grep -v grep` — look for status `Z`). If zombie, the subprocess crashed — check `job-report.json` for error. Fall back to the CLI path (`scripts/run_freshrss_pipeline.py` via `.venv/bin/python3` with `--timeout 300`), which now has no circular import and no transport-layer timeout. Only use `run_freshrss_openclaw_pipeline` (MCP sync) when running from Feishu/tool context where CLI is unavailable.
**⚠️ Timeout trap in sync pipeline fallback (transport-layer 120s hard cap):**
- The `run_freshrss_openclaw_pipeline` MCP tool's `timeout_seconds` parameter defaults to 60 — which is **too short** for production use.
- **Critical: the MCP transport layer has its own hard timeout (120s) that `timeout_seconds` cannot override.** Even with `timeout_seconds=180`, if the total end-to-end pipeline time exceeds ~120s, the transport layer kills the call with: `MCP call timed out after 120.0s`.
- **Empirical timing (July 2026):** Each article in the sync pipeline takes roughly 20-25 seconds total (fetch + extract + LLM summarize). With `limit=5` the pipeline completes within the 120s window; with `limit=7` it reliably times out at the transport layer.
- **Fix:** When falling back to the synchronous pipeline, **reduce `limit` to 5** to stay within the 120s transport window. `timeout_seconds=180` is still recommended to give the pipeline internal breathing room, but it alone cannot solve the transport-layer cap.
```python
# Correct production fallback call (reduced limit to fit within 120s transport cap):
run_freshrss_openclaw_pipeline(
limit=5, # ⚠️ REQUIRED — 7 triggers transport 120s timeout
mark_read=True,
timeout_seconds=180
)
```
**Trade-off:** Reducing `limit` means fewer articles per run. For full `limit=7` production runs, the preferred path is now the CLI (`scripts/run_freshrss_pipeline.py`) which has no transport-layer cap. The async MCP path also works (confirmed July 2026) but adds subprocess overhead. When using the MCP sync path (`run_freshrss_openclaw_pipeline`), stick to `limit=5` to fit within the 120s transport window.
**⚠️ `include_read=true` as a day-start strategy:** When the unread feed (`include_read=false`, the default) is dominated by sources the extractor cannot handle (WeChat `mp.weixin.qq.com` consistently returns `CONTENT_EXTRACTION_FAILED`), running the first pipeline call of the day with `include_read=true` can unlock a completely different candidate set — older articles from sources the extractor CAN handle. This is not a bug workaround; it is a legitimate morning strategy to evaluate when the first `include_read=false` pass produces few or zero keep/review articles. The trade-off is that previously read articles get marked as read again, which is benign for already-processed content.
**Additional note for stuck-run chaining:** If a previous pipeline run from earlier in the same day is still at `generate_summaries` after 60+ seconds, calling `start_freshrss_pipeline_job` will link the new job to the same stuck run. Check `linked_run_id` and its `updated_at` timestamp immediately after starting the async job — if the linked run has not advanced for 60+ seconds, skip resume entirely and go straight to the CLI path fallback.
**Actual state (July 2026):** The critical env vars ARE already injected in `/root/.hermes/config.yaml` under `mcp_servers.reader.env` (`LLM_API_KEY`, `LLM_API_URL`, `LLM_MODEL`, etc.). The reader project's `.env` also has matching entries. Subprocesses inherit os.environ by default. **Testing on 2026-07-01 confirmed both paths work:**
- **Async MCP path** (`start_freshrss_pipeline_job`): completes in ~1-2 minutes. `get_freshrss_pipeline_job_result` returns `status: success`. The previous "hang" (Pattern A) was confirmed to be the circular import causing the subprocess to crash silently.
- **CLI path** (`scripts/run_freshrss_pipeline.py`): verified working with `--limit 7 --include-read --mark-read --timeout 300`, completing in ~24 seconds for 7 articles. **This is the preferred production path** because it has no MCP transport-layer timeout (120s cap) and no subprocess overhead.
**The pipeline summarization is parallel** — `ThreadPoolExecutor(max_workers=4)` was added in the 2026-07-01 session. 7 articles complete summarization in ~24 seconds (3-4s per article, 4 concurrent workers). Before this fix, summarization was serial (one article at a time), which exacerbated the perceived "hang" when combined with the circular import crash.
**Subprocess env injection (2026-07-01):** The `start_freshrss_pipeline_job` spawns a subprocess via `subprocess.Popen(cmd, cwd=..., stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL, start_new_session=True)`. Without explicit `env=` parameter, the subprocess inherits `os.environ` from the parent MCP server. The Hermes config env block (`mcp_servers.reader.env`) IS correctly inherited. The circular import caused the subprocess to crash at import time (before any env var was read), with stderr going to `/dev/null`. The fix was lazy-importing `RunStore` in `freshrss_pipeline.py` — no explicit env injection into the subprocess was needed because env vars propagate naturally through `os.environ` inheritance.
**Debug mode:** Pass `--debug-artifacts` flag to get per-item intermediate files in the output directory — useful for troubleshooting extraction failures or filter decisions. Debug mode adds ~0.5% overhead (just file writes).
**Long-term fix:** Track as a deeper investigation. Current workaround (sync pipeline fallback) is reliable.
## References
Read `references/content-extraction.md` to understand how the pipeline selects content sources (RSS-only for FreshRSS, never re-fetches original URLs) and what `content_source` values mean for digest quality.
Read `references/flow.md` when you need the concrete step-by-step command checklist and file expectations.
Read `references/feishu-format-notes.md` for Feishu markdown formatting rules.
Read `references/ima-credential-chain.md` for where IMA credentials come from (reader `.env` → `~/.config/ima/` → Hermes `.env`) and recovery when they're missing.
Read `references/ima-upload-api.md` for the complete IMA upload API reference (create_media → COS → add_knowledge flow, credentials, curl examples, and markdown reformatting requirements).
Read `references/ima-format-reference.md` — **the authoritative IMA note format template (MUST READ before every IMA write).**
Read `references/ima-format-quickref.md` before any IMA upload — concrete correct/incorrect examples for filenames, titles, and content source.
Read `references/memory-drift-recovery.md` when the memory tool refuses writes due to MEMORY.md file drift — recovery procedure for the recurring `"file on disk has content that wouldn't round-trip"` error.
Read `references/keyword-engine-maintenance.md` for the complete keyword engine maintenance workflow — `build_review_bundle.py` → `generate_term_cleanup_suggestions.py` → `generate_term_cleanup_semantic_suggestions.py` → `apply_term_suggestions.py`, plus the distinction between script-driven (statistical) and LLM-driven (semantic) cleanup.
@@ -0,0 +1,48 @@
# Content Extraction Pipeline
How the pipeline turns FreshRSS items into extractable article text.
## Core Rule: FreshRSS items never re-fetch the original URL
**FreshRSS is an RSS-only upstream.** The pipeline *never* makes an HTTP request to the original article URL for a FreshRSS item. This is enforced by `RSS_ONLY_UPSTREAMS = {"freshrss"}` in `pipeline.py`.
The only exception: non-FreshRSS upstreams (future sources that don't set `upstream: freshrss`) may trigger `fetch_html()` as a fallback.
## Content source priority chain
`content_loader.py` tries sources in this order, using the **first one with ≥500 readable characters**:
| Priority | Source | Meaning |
|----------|--------|---------|
| 1 | `raw_html` | HTML pre-injected via `ExtractionInput.raw_html`. Rarely used in normal FreshRSS runs. |
| 2 | `item.raw_content` | RSS `<content:encoded>` — the full article body. Some feeds provide this; many don't. |
| 3 | `item.raw_summary` | RSS `<description>` — the summary/snippet field. **This is the most common source in current runs.** |
| 4 | `rss_content` | RSS content from non-item sources. |
| — | `none` | Nothing usable → raises `RSS_CONTENT_MISSING` for FreshRSS items (because fetch is skipped). |
## How `content_source` maps to actual text quality
The `content_source` field in every `item-XX.extracted.json` tells you what the pipeline actually used:
- **`item.raw_content`** → Full article text from RSS `<content:encoded>`. Best quality, same as reading the original page.
- **`item.raw_summary`** → RSS summary/description only. **Not the full article.** Length varies wildly by source (300-2000 chars typical). The AI summary is based on this snippet, not the complete text.
- **`rss_content`** → From standalone RSS content. Quality depends on the feed.
- **`fetched_html`** → HTML fetched from the original URL (**never happens for FreshRSS**; only for non-FreshRSS upstreams).
## What this means for digest quality
If you see `content_source: item.raw_summary` in the extracted files (current norm), the AI is summarizing from a **feed summary/snippet**, not the full article body. Articles that seem shallow in the daily digest may simply have short RSS descriptions.
To improve quality: either find feeds that provide full `<content:encoded>`, or switch the feed source to a system that provides full-text RSS (e.g., RSS-proxy with full-text extraction, or a third-party service like FiveFilters).
## Quick check
```bash
# Check content_source for latest run
grep -h "content_source" outputs/freshrss/rerun/*/extracted/item-*.json | sort | uniq -c
```
## Relevant code paths
- `src/summary_mcp/core/pipeline.py` — `RSS_ONLY_UPSTREAMS`, `_should_skip_fetch()`, `extract_content()`
- `src/summary_mcp/core/content_loader.py` — `choose_inline_content()` priority chain, `fetch_html()` (never called for FreshRSS)
@@ -0,0 +1,26 @@
# Feishu Markdown Format Notes
## Background
Hermes' Feishu gateway (`gateway/platforms/feishu.py`) sends outbound messages
through `_build_outbound_payload`, which checks content for markdown patterns
to decide how to send:
- Content matches `_MARKDOWN_HINT_RE` (bold, lists, code, links, etc.)
→ sent as Feishu `post` type using `md` elements → renders correctly.
- Content matches `_MARKDOWN_TABLE_RE` (a markdown table header + separator)
→ **entire message** forced to `text` type (plain text) → no rendering.
The root cause is that the `_build_markdown_post_payload` helper wraps content
in `{"tag": "md", "text": "..."}` elements, and Feishu's `md` element does not
support table rendering. There is no table-to-native-Feishu-table conversion.
## Rules for Feishu output
- **Never use markdown tables** in any message delivered via Feishu.
A single table anywhere in the message forces the whole message to plain text.
- Prefer bullet lists, sections with headings, or inline formatting instead.
- Bold (`**bold**`), inline code (`` `code` ``), unordered lists (`- item`),
ordered lists (`1. item`), and links all work correctly.
- Code fences (``` ``` ```) work but may have edge cases with trailing content.
@@ -0,0 +1,328 @@
# Reader Digest Flow Reference
## Purpose
Concrete operational checklist for the `reader-digest-flow` skill.
## Default Operating Model
### Layering
- `reader` layer:
- FreshRSS pull
- extraction
- summary/filter/payload generation
- selected-article summary capability
- OpenClaw / skill layer:
- public digest generation
- internal review digest generation
- Hugo publishing
- chat reporting
- user confirmation handling
- calling selected-article summaries
- IMA upload orchestration
- Hugo layer:
- public digest browsing and archive only
- IMA layer:
- long-term storage for selected article notes only
### Hard rules
- Do not upload the full digest to IMA.
- Upload only explicitly user-selected articles to IMA.
- Do not generate a digest without a real payload.
- Generate two views from the same payload: a public digest for Hugo and an internal review digest for chat/operator workflow.
- Do not expose internal review states or operator-facing labels in the public digest.
- Do not re-fetch original URLs for selected summaries; use existing extracted text.
- If the main pipeline fails, inspect the run first; when `inspect_resume_plan` says `recommended_action=resume`, continue via the async resume job path instead of stopping immediately.
- Always branch on reader's top-level reconciled `status`; treat `status_source` and `state_conflict` only as explanatory metadata.
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization, with fallback to main `LLM_*` only if needed.
- Actively report progress after each completed phase.
## Step-by-step checklist
### 1. Run reader pipeline
Use the formal MCP workflow path as the default production route.
Prefer MCP run/status/result operations over direct path stitching. Only fall back to CLI or direct file inspection for debug / manual troubleshooting.
Formal production startup sequence:
1. `start_freshrss_pipeline_job`
2. `get_freshrss_pipeline_job_status`
3. `get_freshrss_pipeline_job_result`
4. after success, continue with `run_id` via `get_run_status` / `get_delivery_payload` / `get_run_report`
If the main pipeline job ends in `failed`:
1. inspect the linked run with `get_run_status`
2. call `inspect_resume_plan(run_id)`
3. if `recommended_action=resume`, continue with:
- `start_resume_job`
- `get_resume_job_status`
- `get_resume_job_result`
4. if `recommended_action=read_terminal_result`, continue from the terminal run result
5. if `recommended_action=start_new_run`, stop and report the failure
Treat the old synchronous `run_freshrss_openclaw_pipeline` as debug / light validation / fallback only.
Project root:
```bash
/home/ubuntu/zhu/github/reader
```
Default behavior for a normal production run:
- if the user did not specify a count, randomly choose a limit between 5 and 10 items for that run
- run with mark-read enabled
- do not enable `debug_artifacts`
- only skip mark-read if the user explicitly says the run is debug, test, or validation
- only enable `debug_artifacts` if the user explicitly says the run is debug, test, validation, or troubleshooting
Typical artifacts to inspect after a successful run:
```text
outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json
outputs/freshrss/rerun/<run-id>/candidates/digest-brief.json
outputs/freshrss/rerun/<run-id>/run-report.json
outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json
```
Notes:
- `digest-brief.json` is the preferred input for **public digest** generation.
- It is a lighter public-only view and currently includes only `keep` candidates.
- If `digest-brief.json` is missing, fall back to `openclaw-delivery-payload.json`.
- If the synchronous MCP wrapper times out but a real reader run was still created, do not discard the run; continue from run truth using `list_runs`, `get_run_report`, and `get_delivery_payload`.
**⚠️ Output path divergence when run from Feishu context:**
When triggered via MCP from inside a Feishu session, the extracted files may be written to an MCP-managed temp path instead of `outputs/freshrss/rerun/<run-id>/extracted/`. Verify the actual output path from the pipeline result before continuing to Phase 5/6. If the project-relative path doesn't exist, use the MCP result's `extracted_path` directly or copy the files to the expected project location.
### 2. Generate digest markdown
Generate two output views from the same run, preferably in one model call:
1. a **public digest** for Hugo / public readers
2. an **internal review digest** for chat / operator workflow
Input preference:
- **public digest**: prefer `candidates/digest-brief.json`
- **internal review digest**: use `candidates/openclaw-delivery-payload.json`
Recommended generation pattern:
- Pass the public brief and the full payload as two explicitly labeled input blocks.
- Ask the model to return both outputs in one response.
- Prefer a structured response shape (for example JSON with `public_digest_markdown` and `internal_review_digest_markdown`) when post-processing is needed.
Before publishing, persist the generated digest artifacts back into the same reader run directory:
```text
outputs/freshrss/rerun/<run-id>/digest/public_digest.md
outputs/freshrss/rerun/<run-id>/digest/internal_review_digest.md
outputs/freshrss/rerun/<run-id>/digest/combined.json
```
Then write only the public digest into Hugo here:
```text
/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md
```
Recommended public-digest front matter:
```toml
+++
title = "AI 日报 · YYYY-MM-DD"
date = YYYY-MM-DDTHH:MM:SS+08:00
summary = "当日日报摘要"
+++
```
Recommended **public digest** structure:
- `今日概览`
- `今日重点`
- `趋势观察`
- `延伸阅读`
Public digest constraints:
- Use only the public brief view when present.
- Keep public wording free of internal workflow language.
- Optimize for concise public readability with solid information density.
- For each `今日重点` item, add one short editor-style sentence explaining why the item matters in today's digest (for example: `这篇内容更值得关注的原因在于……`).
- Prefer rendering highlight points as a short public label such as `值得关注:` followed by one bullet per line.
- **Article numbering: MUST use `1.` `2.` `3.` (Arabic numeral + period). Do NOT use `① ② ③`,`一、二、三`,`第一条` or any other variant.**
- **All four sections are required. Missing any one is a format violation.**
- **Hugo digest must include ALL keep articles from `digest-brief.json`. The IMA deposition subset is a separate downstream step.**
Recommended **internal review digest** structure:
- `今日候选概况`
- `已入选重点`
- `待你确认`
- `建议沉淀到 IMA`
- `原始候选清单`
Internal review digest constraints:
- Do not show `rank` values.
- Replace machine states with Chinese labels such as `已入选` / `待确认` / `暂不纳入`.
- For `已入选重点`, include a fuller summary plus a short judgment paragraph.
- For `待你确认`, include a fuller summary, reason, and recommendation.
- Keep it readable as an operator review draft, not a raw payload dump.
- **Feishu compatibility: never use markdown tables in the digest report.**
When delivered via Feishu, a single table forces the whole message to plain
text. Use lists and sections instead. See `references/feishu-format-notes.md`.
### 3. Publish Hugo
Publish only the public digest to Hugo.
Treat Hugo publication as the default continuation of a successful normal daily digest run. Do not ask for a second confirmation before generating/writing the public digest and publishing it, unless the user explicitly requests not to publish to Hugo.
**⚠️ Hugo redeploy: always use foreground terminal, never background+notify_on_complete.**
The Hugo redeploy runs via `./redeploy.sh` and completes in ~10 seconds. **Do NOT use `terminal(background=true, notify_on_complete=true)`** for this step — the Gateway will push the raw Docker build output (compiler logs, layered build output, nginx config) to Feishu as a notification. This output is machine-readable, not human-readable, and the user has explicitly said this is noise.
**Correct approach:**
```python
# ✅ Foreground terminal with adequate timeout
cmd = "cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh"
result = terminal(cmd, timeout=120)
# Then verify 200 OK on the detail page
```
**Wrong approach (DO NOT use):**
```python
# ❌ Background + notify sends raw build logs to Feishu
terminal("cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh", background=True, notify_on_complete=True)
```
This rule only applies to Hugo redeploy (fast, deterministic output). For long-running Reader pipeline jobs (>60s), background + notify_on_complete is still appropriate since the output is meaningful content (article summaries, pipeline stats).
1. **Pre-check: verify Hugo content directory exists**
```bash
ls -d /home/ubuntu/zhu/apps/hugo-site/content/daily/ 2>/dev/null || mkdir -p /home/ubuntu/zhu/apps/hugo-site/content/daily/
```
If the entire `content/` directory is missing, create it before proceeding.
**Do not assume the Hugo site has a standard structure.**
2. write the public digest markdown to:
- `/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md`
3. redeploy Hugo immediately after writing:
- `cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh`
4. verify all three URLs before continuing:
- `http://127.0.0.1:14322/`
- `http://127.0.0.1:14322/daily/`
- `http://127.0.0.1:14322/daily/YYYY-MM-DD/`
Expected verification targets:
- homepage works
- `/daily/` works
- `/daily/YYYY-MM-DD/` works
### 4. Report digest in chat
Provide the internal review digest in chat and ask which articles should be retained.
Hard reporting rule:
- do not send only article titles
- for each article, include at least a one-sentence summary and a short recommendation / judgment so the user can decide without reopening the source
- **Feishu**: avoid markdown tables entirely. Use lists with headings.
### 5. Generate selected article summaries
Only do this after the user explicitly confirms which articles to retain.
Preferred MCP path:
- `start_article_summary_job`
- `get_article_summary_job_status`
- `get_article_summary_job_result`
Expected inputs:
- `extracted_path`
- `selected_ids`
**⚠️ item_id 前缀注意:**
The `candidates` array IDs use `cand:sha256:xxx` format, but individual extracted item files use `sha256:xxx` (no `cand:` prefix). When passing `selected_ids` to article-summary tools, strip the `cand:` prefix. If the ID doesn't match, the summary tool won't find the article.
**Path resolution for article-summary:**
After a pipeline run via MCP (Feishu context), the extracted files may live at an MCP-managed temp path, not the expected `outputs/freshrss/rerun/<run-id>/extracted/`. Before calling `generate_article_summaries` or `start_article_summary_job`, verify the extracted path exists. If not, use the `run_id` from the pipeline result to locate the actual output directory through `get_run_report`, or copy the files from the MCP-managed path.
Single-item rule:
- when `extracted_path` is `outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json`, call one summary job per file
- in that case, `selected_ids` should contain only the matching single `item_id`
- if the candidate ID came from the delivery payload, strip the `cand:` prefix before passing it
Recommended output layout:
- `outputs/freshrss/single_summaries/YYYY-MM-DD/`
- async job state: `outputs/freshrss/article_summary_jobs/<job_id>/`
Recommended production sequence:
1. call `start_article_summary_job`
2. poll `get_article_summary_job_status` until `status` becomes `success` or `failed`
3. on success, call `get_article_summary_job_result` and continue downstream from `written_paths`
Hard fallback rule:
- If the async MCP job path returns timeout / transport failure / job-launch failure (for example MCP timeout while the reader article-summary workflow itself is still healthy), do not treat that as article-summary business failure.
- Immediately retry through the local reader environment under `/home/ubuntu/zhu/github/reader` using the project `.venv`, calling the article-summary workflow directly.
- The production goal is successful generation of the selected-article Markdown files; async MCP job is preferred, but local `.venv` execution is the required fallback path.
Synchronous helper:
- `generate_article_summaries` remains available for debug / light validation only, not as the default production path.
CLI fallback:
```bash
python scripts/run_article_summaries.py \
--extracted outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json \
--ids <item_id_without_cand_prefix> \
--output-dir outputs/freshrss/single_summaries/YYYY-MM-DD
```
### 6. Upload selected summaries to IMA
Upload only the generated markdown files for the selected articles.
Once the user has selected the articles to retain, treat that selection itself as the authorization to continue the IMA deposition step; do not ask for a second confirmation about uploading into the knowledge base.
**⚠️ IMA upload 500 retry: Always wrap IMA uploads in a retry loop.**
IMA's OpenAPI may return HTTP 500 on the first attempt. If the first upload fails, wait ~3 seconds and retry. Normally the second attempt succeeds. If `cos-upload.cjs` is used (for CDN-backed uploads), check subprocess stderr even when exit code is 0 — an HTTP error in the upload service may still produce exit code 0.
Hard execution rules before upload:
1. reformat/check the generated markdown into IMA-facing final content
2. normalize the final upload filename to `<文章标题>.md`
3. do not use internal temp names such as `ima-*`, `item-*`, `summary-*`, or English slug filenames as the final uploaded object name
4. if the knowledge base already contains the same filename, append a timestamp suffix before `.md`
5. if local work needs internal temp names, create a final upload copy with the user-facing title before calling IMA APIs
Default target knowledge base:
- `daily`
- Read `IMA_DAILY_KNOWLEDGE_BASE_ID` / `IMA_DAILY_KNOWLEDGE_BASE_NAME` from reader `.env`
- Verify the configured target at runtime before upload
- If the configured target is unavailable, resolve by name `daily`; if still absent, create `daily`
## Hard rules recap
- Never generate a digest from placeholder or example data when a real run is expected.
- For normal runs, mark processed FreshRSS items as read unless the user explicitly requested a debug/test/validation run.
- Public digest goes to Hugo; internal review digest goes to chat; neither full digest goes to IMA.
- Only explicitly user-selected articles go to IMA.
- All daily IMA deposition must go directly into the IMA knowledge-base path as Markdown knowledge items (`media_type=7`), not through the IMA notes path.
- Uploading to IMA notes, or creating notes first and then linking them into a knowledge base, does not count as SOP completion.
- Selected article summaries use extracted text, not live refetch.
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization.
- Final IMA upload filenames must use user-facing article titles, not internal slugs or workflow temp names.
@@ -0,0 +1,55 @@
# IMA COS 上传凭证处理(Hermes redact_secrets 兼容模式)
## 问题
Hermes 配置 `security.redact_secrets: true` 时,`create_media` API 返回的 `cos_credential.token` 字段在 `terminal()` 输出中被替换为 `***`。
直接通过 `terminal()` 调用 `cos-upload.cjs` 会因 token 截断而失败(HTTP 403 InvalidAccessKeyId)。
## 安全的工作流(execute_code + subprocess.run)
不要用 `terminal()` 传递 COS 凭证。改用 `execute_code()` + `subprocess.run()` 模式:
```python
# Phase A: terminal() 中保存原始响应到文件
result = terminal("""
bash -c '
set -a
source /home/ubuntu/zhu/github/reader/.env
set +a
OPTS=$(printf "%s" "{\\"clientId\\":\\""$IMA_OPENAPI_CLIENTID"\\",\\\"apiKey\\":\\\""$IMA_OPENAPI_APIKEY"\\\"}")
RESP=$(node /root/.hermes/skills/openclaw-imports/ima-skill/ima_api.cjs "openapi/wiki/v1/create_media" "{...}" "$OPTS" 2>/dev/null)
echo "$RESP" > /tmp/create_media_raw.json
echo "saved"
'
""")
# Phase B: execute_code 中从文件读取凭证
import json, subprocess
with open("/tmp/create_media_raw.json") as f:
data = json.load(f)
cred = data["data"]["cos_credential"]
# Phase C: subprocess.run 直接调用,不经过 terminal()
args = ["node", f"{SKILL_DIR}/knowledge-base/scripts/cos-upload.cjs",
"--file", FILE_PATH, "--secret-id", cred["secret_id"], "--secret-key", cred["secret_key"],
"--token", cred["token"], "--bucket", cred["bucket_name"], "--region", cred["region"],
"--cos-key", cred["cos_key"], "--content-type", "text/markdown",
"--start-time", cred["start_time"], "--expired-time", cred["expired_time"], "--timeout", "300000"]
r = subprocess.run(args, capture_output=True, text=True, timeout=310)
```
## 对比:错误的做法(terminal 直接传递)
```bash
# ⛔ 这样不行!token 会被 redact_secrets 替换为 ***
TOKEN=$(echo "$RESP" | jq -r '.data.cos_credential.token')
node ... --token "$TOKEN" ... # 会收到 HTTP 403
```
## 关键原则
- `terminal()` 输出中的敏感字段会被自动脱敏,但不影响底层 JSON 文件写入
- `execute_code` 中 `terminal()` 返回的 `output` 已经是脱敏后的文本
- **唯一可靠的凭证源**是直接写入磁盘的原始 JSON 文件
- `subprocess.run` 在 `execute_code` 中绕过脱敏,因为凭证在 Python 内存中直接被传递给子进程,不经过 Hermes 的 stdout 脱敏管道
@@ -0,0 +1,46 @@
# IMA 凭证链:从 reader `.env` 到 IMA 上传
## 凭证来源
IMA 上传所需的凭证存储在多个位置,优先级如下:
| 优先级 | 位置 | 说明 |
|--------|------|------|
| 1 | 环境变量 `IMA_OPENAPI_CLIENTID` / `IMA_OPENAPI_APIKEY` | Hermes session 级 |
| 2 | `~/.config/ima/client_id` / `api_key` | ima-skill 的默认检查路径 |
| 3 | `/home/ubuntu/zhu/github/reader/.env` | reader 项目配置,含完整的 IMA 凭证和 KB ID |
## 凭证内容(reader .env 中)
```
IMA_OPENAPI_CLIENTID=<32位hex>
IMA_OPENAPI_APIKEY=<base64编码的API密钥>
IMA_DAILY_KNOWLEDGE_BASE_ID=<base64编码的KB ID>
IMA_DAILY_KNOWLEDGE_BASE_NAME=daily
```
## 缺失时的处理流程
当 IMA 上传失败(`-100` 凭证缺失错误)时:
1. 从 reader `.env` 读取凭证:
```
grep -E '^(IMA_OPENAPI_CLIENTID|IMA_OPENAPI_APIKEY)=' /home/ubuntu/zhu/github/reader/.env
```
2. 同步到 ima-skill 默认检查路径:
```
echo "<client_id>" > ~/.config/ima/client_id
echo "<api_key>" > ~/.config/ima/api_key
```
3. (可选)追加到 Hermes `.env` 以全局生效:
```
echo "IMA_OPENAPI_CLIENTID=<client_id>" >> /root/.hermes/.env
echo "IMA_OPENAPI_APIKEY=<api_key>" >> /root/.hermes/.env
echo "IMA_DAILY_KNOWLEDGE_BASE_ID=<kb_id>" >> /root/.hermes/.env
echo "IMA_DAILY_KNOWLEDGE_BASE_NAME=daily" >> /root/.hermes/.env
```
## 执行注意事项
- **COS 凭证红线**:`create_media` 返回的 `cos_credential` 中的 `token`/`secret_id`/`secret_key` 在 `terminal()` 输出中会被 Hermes 替换为 `***`。必须用 `subprocess.run()` 捕获原始输出,或用 `execute_code` 内联操作。
- **凭证格式**:`api_key` 是 base64 字符串(76 字符),`kb_id` 也是 base64 字符串。不要截断或转码。
@@ -0,0 +1,97 @@
# IMA Markdown 文件上传流程(media_type=7)
## 什么时候用此流程
当用户说"沉淀到 IMA"时,必须用此文件上传流程,而不是 URL 导入或笔记导入。
## 完整流程
### 1. 写 .md 文件
```bash
mkdir -p /tmp/ima_upload
cat > /tmp/ima_upload/文章标题.md << 'EOF'
# 文章标题
> 来源:XXX
## 摘要
...
## 核心亮点
- ...
[原文链接](url)
EOF
```
### 2. preflight 检查
```python
pf = subprocess.run(
["node", f"{SKILL_DIR}/knowledge-base/scripts/preflight-check.cjs",
"--file", filepath],
capture_output=True, text=True
)
meta = json.loads(pf.stdout)
# meta = {pass, file_name, file_ext, file_size, media_type, content_type}
```
### 3. create_media(获取 COS 凭证)
```python
r = ima_api("openapi/wiki/v1/create_media", {
"file_name": fname,
"file_size": meta["file_size"],
"content_type": meta["content_type"],
"knowledge_base_id": kb_id,
"file_ext": meta["file_ext"]
})
cos = r["data"]["cos_credential"]
media_id = r["data"]["media_id"]
```
### 4. COS 上传
```python
cu = subprocess.run([
"node", f"{SKILL_DIR}/knowledge-base/scripts/cos-upload.cjs",
"--file", filepath,
"--secret-id", cos["secret_id"],
"--secret-key", cos["secret_key"],
"--token", cos["token"],
"--bucket", cos["bucket_name"],
"--region", cos["region"],
"--cos-key", cos["cos_key"],
"--content-type", meta["content_type"],
"--start-time", str(cos["start_time"]),
"--expired-time", str(cos["expired_time"]),
"--timeout", "300000"
], capture_output=True, text=True, timeout=30)
# 非0退出 = 上传失败
```
### 5. add_knowledge
```python
r = ima_api("openapi/wiki/v1/add_knowledge", {
"media_type": 7,
"media_id": media_id,
"title": title, # 文章中文标题
"knowledge_base_id": kb_id,
"file_info": {
"cos_key": cos["cos_key"],
"file_size": meta["file_size"],
"file_name": fname
}
})
```
## 注意事项
- **COS 凭证不能通过 terminal() 读取**(Hermes redact_secrets 会把 token 替换成 `***`)。必须用 Python `subprocess.run()` 捕获原始 stdout。
- **create_media 的 content_type 不能带 charset 参数**(如 `text/markdown; charset=utf-8` 会被拒)。用纯 MIME 类型 `text/markdown`。
- **文件命名 = `<文章中文标题>.md`**。不要用英文 slug 或 temp name。
- **media_type=7** 是 Markdown 文件。**绝对不要用 media_type=11**(笔记)或 `import_urls`。
@@ -0,0 +1,42 @@
# IMA 上传格式速查表(日报沉淀专用)
## 文件名 vs 标题 vs 内容对照表
| 项目 | ✅ 正确 | ❌ 错误 |
|------|---------|---------|
| **文件名** | `从Vibe Coding到Harness—— 一套大仓AI工程化实战.md` | `大仓AI工程化实战.md` |
| **add_knowledge title** | `从Vibe Coding到Harness—— 一套大仓AI工程化实战` | `大仓AI工程化实战` / `大仓AI工程化实战.md` |
| **上传方式** | `preflight` → `create_media` → `cos-upload` → `add_knowledge(media_type=7)` | `import_doc` + `add_knowledge(media_type=11)` |
| **正文内容** | pipeline `extracted/` 的 `article.content` 或 `summary-batch.json` 的 `summary` | 自己写的两三句话 |
| **正文长度** | ≥ 500 字 | < 500 字 |
## 文件名常见错误模式
| 原始标题 | ❌ 错误文件名 | ✅ 正确文件名 |
|----------|-------------|-------------|
| `从Vibe Coding到Harness—— 一套大仓AI工程化实战` | `大仓AI工程化实战.md` | `从Vibe Coding到Harness—— 一套大仓AI工程化实战.md` |
| `契约化多端架构:基于领域模型的Harness实践` | `契约化多端架构Harness实践.md` | `契约化多端架构:基于领域模型的Harness实践.md` |
| `Loop Engineering 实战:从日志扫描到预发部署的全自主闭环` | `LoopEngineering实战.md` | `Loop Engineering 实战:从日志扫描到预发部署的全自主闭环.md` |
## 5 Section 格式要求
| Section | 格式 | 说明 |
|---------|------|------|
| **核心结论** | 一段连贯段落(非分点) | 总结文章核心发现或主张 |
| **主要论点** | 一段连贯段落(非分点) | 综合多个论点形成连贯叙述 |
| **关键方法 / 机制** | `- **方法名**:详细说明` 分点 | 每条展开到能理解原理的程度 |
| **重要细节** | `- 每条一个带解释的完整知识点` 分点 | 每条是一个完整知识点,不是关键词 |
| **可复用启发** | `- 每条 actionable 的实践启示` 分点 | 附应用场景说明 |
## 底部额外 section
- `## 关键词` — 标签式关键词列表
- `## 主题` — 主题分类
## 速查口诀
> 文件名 = 完整标题.md
> 内容取 pipeline,不自己写
> 5 个 section,核心结论和主要论点是段落
> media_type = 7,不是 11
> title 不加 .md
@@ -0,0 +1,38 @@
# AI Agent 的 Skill 系统设计
Source: https://mp.weixin.qq.com/s?__biz=MzAxNDEwNjk5OQ==&mid=2650544717&idx=1&sn=b578abf5a81034670900a3b8eb874296
Category: 方法论
## 核心结论
好的 Skill 应是一个小而准的行为系统,通过触发、加载、执行、约束、验证和迭代的组织,将通用 Agent 转化为在特定任务上稳定可靠的专用 Agent。核心原则是上下文窗口是公共资源,必须采用渐进披露、按任务风险设置自由度,并通过真实任务前向测试来证明行为改变。
## 主要论点
Skill 设计的本质是行为编程而非文档编写,需要将期望行为转化为 Agent 能稳定执行的结构化工作流。为此,必须同时解决发现(正确场景触发)、加载(最小上下文)、执行(合适自由度)和验证(真实任务测试)四件事,并通过门控、脚本外化、测试防合理化等机制确保 Agent 在复杂压力下不走捷径。
## 关键方法 / 机制
- 渐进披露的三层内容加载:元数据(frontmatter 的 name/description)用于发现,正文(SKILL.md)用于执行,资源(scripts/references/assets)按需读取。description 只做路由器,不包含完整流程,防止 Agent 凭印象执行。
- 门控机制(HARD-GATE):在低自由度任务中,使用明确的 <HARD-GATE> 标签禁止 Agent 在条件满足前执行后续动作,减少解释空间,让关键路径更像程序而非建议。
- 脚本外化降低上下文消耗和行为漂移:将需要确定性的操作(如 PDF 旋转)封装为 scripts/ 中的可执行脚本,避免 Agent 每次临时生成代码,提升可靠性和节省 token。
- 基于 TDD 的前向测试方法:用子代理模拟真实用户任务,只给原始任务和最少上下文,不泄露预期结论;观察行为轨迹、输出文件等原始证据,发现并封堵 Agent 的违规行为。
- 反模式自查表与检查表:交付前检查触发条件、自由度设置、资源引用、验证流程等关键项,确保 Skill 不是一份草稿而是一个可用的能力包。
- 跨平台适配与优雅降级:Skill 应写行为规则(如 TodoWrite),再通过平台层映射到具体工具名(如 todowrite);平台能力不足时优雅降级,保持 Skill 的可迁移性。
## 重要细节
- SKILL.md 的 frontmatter 和正文职责分离:name/description 用于发现(Agent 触发前可见),正文用于执行(触发后加载)。如果触发条件写在正文里,Agent 在决定是否触发时根本读不到。
- 命名规范:短、可触发、动词优先,例如 create-skill 比 skill-creation 更好,这本质上是路由质量——Agent 在技能库里找能力时,name/description 是第一层索引。
- 资源组织原则——“信息只放一个地方”:不要在 SKILL.md 和 references/ 中重复同一段规则,重复会带来漂移,导致 Agent 在两个版本间自行解释,增加维护成本。
- 门控类型示例:先决条件门控(先理解例子再编辑)、并发冲突门控、未保存内容门控、敏感操作门控(已创建 Skill 需处理影响再修改)。
- 流程图用 GraphViz DOT 嵌入 Markdown:对于包含非线性判断、循环、回退的步骤,流程图比纯文本更稳定,能防止 Agent 遗漏关键分支。
- 验证时防“合理化”问题:AI Agent 在压力下会为跳过规则编造理由,Skill 需要提前写出这些借口并给出反驳;审查循环应围绕真实失败风险而非措辞偏好。
## 可复用启发
- “上下文窗口是公共资源”原则:设计任何 Agent 指令时,每段内容都要质疑“Agent 真的需要这段解释吗?”和“值得占用的 token 成本吗?”,这适用于提示词、系统消息等所有 Agent 输入设计。
- 先收集具体例子再抽象 Skill:不要从抽象能力开始写,而是先收集用户会怎么触发、哪些请求应该触发/不应该触发、成功输出是什么等具体场景,避免写出宽泛不可执行的指令。
- 用脚本固化确定性任务、用门控防止关键路径走捷径:对于高脆弱、低变化空间的任务(如文件格式转换),应使用脚本而非描述性建议;对于必须按顺序执行的步骤,用门控打断 Agent 的“合理化”冲动。
- 设计防合理化的测试流程:用子代理模拟真实用户,只给原始任务,不泄露预期答案;观察是否存在只有看到结论才能成功的情况——如果这样,说明 Skill 不够清楚或测试设置泄露答案。
## 关键词
SKILL.md、YAML、Markdown、DOT、GraphViz、TDD、HARD-GATE、quick_validate.py
## 主题
AI Agent、行为编程、Token 经济、系统设计、约束机制
@@ -0,0 +1,33 @@
# IMA 笔记格式参考
> ⚠️ 格式基准文件为 `references/ima-format-reference.md`(`ai-agent-的-skill-系统设计.md`),每次生成 IMA 沉淀前必须先读该文件。
## 标准结构
### 顶部元数据
```markdown
# 文章完整标题
Source: https://原文链接(纯 URL,不加 "原文链接:" 标签)
Category: 分类
```
### 5 个必含 Section
1. **核心结论** — 一段总结文章核心发现或主张的段落(不是分点),像基准文件一样是一段连贯文字
2. **主要论点** — 一段概括文章核心论述的段落(不是分点列表),综合多个论点形成连贯叙述
3. **关键方法 / 机制** — `**方法名**:详细说明` 的格式,每条展开到能理解其原理的程度
4. **重要细节** — 每条一个带解释的完整知识点,不是关键词
5. **可复用启发** — 每条 actionable 的实践启示,附应用场景说明
### 底部额外 Section
- **## 关键词** — 标签式关键词列表
- **## 主题** — 主题分类
### 质量要求
- 每篇文章总内容量应达到 **4,000+ bytes**(基准文件 4,782B)
- 每个 section 展开到完整的知识点级别,不能只列关键词
- 内容来源:优先 `item-XX.extracted.json` 的 `article.content`,其次 `summary-batch.json` 的 `summary` 字段
@@ -0,0 +1,80 @@
# IMA Upload API Reference
Full API flow for uploading markdown articles to the IMA `daily` knowledge base. Used in Phase 7 of the reader-digest-flow.
## Credentials
```
IMA_OPENAPI_CLIENTID - from reader .env or user-provided
IMA_OPENAPI_APIKEY - from reader .env or user-provided
IMA_DAILY_KNOWLEDGE_BASE_ID - daily KB UUID
```
The `ima-skill` v1.1.7+ ships with `ima_api.cjs` for credential loading.
Legacy auth header: `ima-openapi-ctx: skill_version=1.1.7`.
Do NOT export the full API key in shell commands — use `execute_code` with `subprocess.run` and Python string variables.
## Flow (3 steps)
### 1. create_media
```
POST https://ima.qq.com/openapi/wiki/v1/create_media
Headers: ima-openapi-clientid, ima-openapi-apikey, Content-Type: application/json
Body: { file_name, file_size, content_type, knowledge_base_id, file_ext }
Returns: { code: 0, data: { media_id, cos_credential: { secret_id, secret_key, token, bucket_name, region, cos_key, start_time, expired_time } } }
```
`file_ext` is without the dot (e.g. `md` not `.md`).
`file_name` must be the user-facing article title + `.md`.
`content_type` for markdown is `text/markdown`; media_type=7.
### 2. COS upload
Use `cos-upload.cjs` from `ima-skill/knowledge-base/scripts/`:
```
node <skill_dir>/knowledge-base/scripts/cos-upload.cjs \
--file <local_md_file> \
--secret-id <from create_media> \
--secret-key <from create_media> \
--token <from create_media> \
--bucket <bucket_name> \
--region <region> \
--cos-key <cos_key> \
--content-type text/markdown \
--start-time <start_time> \
--expired-time <expired_time>
```
⚠️ Must use Python `subprocess.run(args=[...])` to avoid shell parameter mangling.
⚠️ Always capture `returncode` and `stderr` — COS may return exit 0 on HTTP 500.
### 3. add_knowledge
```
POST https://ima.qq.com/openapi/wiki/v1/add_knowledge
Headers: same as create_media
Body: { media_type: 7, media_id, title: "<file_name>", knowledge_base_id, file_info: { cos_key, file_size, file_name } }
```
`media_type=7` for markdown. `title` MUST equal `file_name`.
## Article Markdown reformatting (before upload)
Generated summaries from `reader` have `Source:` and `Category:` header lines.
Before uploading, reformat to IMA style:
```
原文链接:<original article URL>
## 核心结论
...
## 主要论点
...
```
Remove `Source:`, `Category:` lines. Keep `原文链接:` at top with the URL on the next line.
Break long prose (>200 chars per paragraph) into shorter paragraphs for IMA readability.
@@ -0,0 +1,195 @@
# 关键词引擎维护流程
## 概述
reader 项目内置了完整的关键词清洗链路,用于管理 `term_aliases.json`、`term_stopwords.json`、`filter_context.personal.json` 等配置。本文件说明链路各环节的职责与使用方式。
## 完整数据流
```
每日日报 pipeline
│
▼
term_index/daily/YYYY-MM-DD.json ← build_keyword_index.py (pipeline 末环节自动)
│
▼
term_index/term_stats.json ← 全量汇总
│
▼
build_review_bundle.py ← 打包审查数据包 (手动触发)
│
▼
generate_term_cleanup_suggestions.py ← 生成建议 (手动触发)
│
▼
apply_term_suggestions.py ← 人工确认后落盘 (手动触发)
```
## 各环节命令
### 1. 重建每日 term_index(pipeline 末环节自动执行,也可手动)
```bash
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 scripts/build_keyword_index.py \
--input outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json
```
### 2. 重建 review bundle(打包当前配置+统计供审查)
```bash
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
--days 365 \
--top 200 \
--output outputs/term_index/review/keyword-cleanup-bundle.json
```
参数:
- `--days 365` — 考虑最近多少天的统计数据
- `--top 200` — 取前 N 个高频词纳入 bundle
### 3. 生成清洗建议
```bash
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 scripts/generate_term_cleanup_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
--emit-markdown
```
输出:
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json` — 正式建议产物
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md` — 人工阅读展示稿
当前脚本能生成的建议类型:
| 类型 | 生成规则 | 当前产出 |
|:----|:---------|:---------|
| interest_keyword_suggestions | 高频未覆盖词(total≥3, days≥2) | ✅ 按阈值产出 |
| watch_terms | 中频观察词(total≤2, days≤2) | ✅ 按阈值产出 |
| alias_suggestions | 大小写变体/单复数/空格连词符差异 | ✅ 统计规则产出 |
| stopword_suggestions | — | ❌ 脚本硬编码为空 |
### 4. 应用建议(dry-run → review → apply)
```bash
# 先 dry-run 预览
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 scripts/apply_term_suggestions.py \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--accept-interest 强化学习 ReAct CLI \
--dry-run
# 确认后正式 apply(去掉 --dry-run)
.venv/bin/python3 scripts/apply_term_suggestions.py \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--accept-interest 强化学习 ReAct CLI \
```
支持的 accept 参数:
- `--accept-interest 词1 词2 ...` — 添加到 interest_keywords
- `--accept-watch 词1 词2 ...` — 添加到 watchlist
- `--accept-alias 源词1 源词2 ...` — 添加 alias 映射
- `--accept-stopword 词1 词2 ...` — 添加停用词
## 脚本能做什么 vs 不能做什么
### ✅ 脚本能做的(统计级清洗)
- 发现大小写变体(`vibe coding` → `Vibe Coding`)
- 发现单复数差异(`Agent Skill` → `Agent Skills`)
- 发现空格/连词符差异
- 按频次推荐 should-be-interest / should-be-watch 的词
- 批量 apply 到配置文件,自动记录变更日志
### ❌ 脚本不能做的(语义级清洗,需要人工判断)
- 识别产品名(`WorkBuddy`、`LibTV Agent`、`飞书妙搭`)→ 应加 stopword
- 识别模型名(`Qwen3-30B-A3B`、`GLM 5.2`、`Opus 4.8`)→ 应加 stopword
- 识别人名/地名(`Andrej Karpathy`、`成都天府长岛`)→ 应加 stopword
- 英文专有名词归一化到中文主题词(`Multi-Agent` → `多Agent`)
- 长产品名归一化(`火山云数据库PostgreSQL Serverless版` → `Serverless数据库`)
- 判断某个词对 Hugo tag 是否有用(`1688`、`OPC训练营` → 无效)
## 语义级清洗:LLM 辅助(generate_term_cleanup_semantic_suggestions.py)
除了统计规则的清洗,项目还有一个 **LLM 驱动的语义级清洗脚本**,能发现统计规则做不到的事情:
```bash
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 scripts/generate_term_cleanup_semantic_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json
```
参数:
- `--bundle` — review bundle(先跑 build_review_bundle.py)
- `--suggestions` — 规则层的建议 JSON(可选,供 LLM 避免重复)
- `--output` — 输出路径(自动生成)
- `--dry-run` — 只打印 prompt 不调 LLM
### 语义脚本能做的(而统计规则不能做的)
| 类型 | LLM 能发现什么 | 示例 |
|:----|:--------------|:-----|
| **semantic_alias** | 中英文映射、简称↔全称、近义词 | `AI智能体`→`AI Agent`, `AI Harness`→`Harness Engineering` |
| **stopword** | 语义判断哪些词太泛/不相干 | `自动化`(太泛), `AlphaFold`(生物领域), `陌生化`(穿透词) |
| **promote_to_interest** | 与用户关注领域对齐的新词 | `强化学习`, `Loop Engineering`, `MoE`, `ChatGPT` |
### 两阶段清洗 SOP
当需要清洗关键词时,按此顺序操作:
**阶段一:规则统计清洗(脚本发现 + 人工确认)**
1. `build_review_bundle.py` → 重建 bundle
2. `generate_term_cleanup_suggestions.py` → 产出统计级建议
3. 检查建议,决定哪些 accept
4. `apply_term_suggestions.py --dry-run` → 预览
5. `apply_term_suggestions.py` → 正式落地
**阶段二:语义级清洗(LLM 发现 + 人工确认)**
1. `generate_term_cleanup_semantic_suggestions.py` → 调 LLM 产生语义建议
2. 检查 LLM 的 semantic_alias / stopword / promote_to_interest
3. 人工编辑 term_aliases.json / term_stopwords.json(语义级判断不能自动化)
**🔴 不要绕过确认流程直接改配置。** 先跑脚本看看发现了什么,再针对性改。未经审视的批量改配置会引入新问题。
## 完整的语义级清洗操作流程
当需要大量添加 aliases/stopwords 时,推荐流程:
1. **跑一次 `build_review_bundle.py` + `generate_term_cleanup_suggestions.py`** 看看统计建议
2. **走一遍实际 pipeline 产出**,收集所有 unique keywords:
```bash
for rid in $(ls outputs/freshrss/rerun/); do
python3 -c "
import json
d = json.load(open('outputs/freshrss/rerun/$rid/summary/summary-batch.json'))
for item in d['items']:
for kw in item['summary'].get('keywords', []):
print(kw)
"
done | sort -u
```
3. 按类别分组:正常词 / 产品名(stopword) / 模型名(stopword) / 需要 alias 的
4. 分别更新 `term_aliases.json` 和 `term_stopwords.json`
5. **重建 term_index**:对每个受影响日期的 run 跑 `build_keyword_index.py`
6. **更新 Hugo tags**:从重建后的 `data/term_index/daily/YYYY-MM-DD.json` 取
## 当前配置 (2026-07-16)
| 文件 | 条目数 |
|:----|:------|
| `configs/term_aliases.json` | 142 |
| `configs/term_stopwords.json` | 106 |
| `configs/filter_context.personal.json` | 54 (interest_keywords) |
| `configs/term_watchlist.json` | 6 |
## 关键词/别名/停用词配置的更新规范
- `term_aliases.json` — 只放"源词→目标词"的映射,目标是让不同写法的同一概念归一到标准形式
- `term_stopwords.json` — 放产品名、模型名、人名、地名等对 Hugo tag 无价值的词
- **随时可以加**,加了后重建 term_index 即可生效
- 不涉及 pipeline 重新跑——只影响下游展示
@@ -0,0 +1,60 @@
# Memory Drift Recovery
## Symptom
`store_memory(action="add", ...)` fails with:
> Refusing to write MEMORY.md: file on disk has content that wouldn't round-trip through the memory tool...
A `.bak` snapshot is created: `/root/.hermes/memories/MEMORY.md.bak.<timestamp>`
## Root Cause
The MEMORY.md file format doesn't match what the memory tool expects — likely because the file was modified externally (by `patch`, `write_file`, shell `>>` append, or a concurrent session). The tool uses a `§` (section sign) delimited format internally and its serialization/deserialization doesn't match the on-disk content.
## Recovery Procedure
### Step 1: Read the backup and the current file
```bash
diff /root/.hermes/memories/MEMORY.md.bak.<timestamp> /root/.hermes/memories/MEMORY.md
```
### Step 2: Extract missing entries (if any)
```bash
# List entries from the backup
grep '^§' /root/.hermes/memories/MEMORY.md.bak.<timestamp>
```
### Step 3: Re-add each missing entry via store_memory
For each entry that was in the backup but is now gone from the current file:
```bash
store_memory(action="add", content="<entry text>", target="memory")
```
### Step 4: Reset to a clean state
If the file is completely corrupted, the cleanest path is:
1. Save any new entries from the backup you want to keep
2. Rewrite the file as a clean `§`-delimited list (one entry per `§` line)
3. The format is: `§<content>\n` per entry, with `---` or blank line separators
```bash
# Example clean format:
echo '§当前重要条目一
§当前重要条目二
§当前重要条目三' > /root/.hermes/memories/MEMORY.md
```
### Prevention
- Do NOT use `write_file` or `patch` to modify MEMORY.md directly — always use `store_memory()`
- Do NOT use shell `>>` to append to MEMORY.md
- If you must bulk-import, use `store_memory` per-entry, not file-level operations
## Environment
- Host: Linux (5.15)
- Hermes home: `/root/.hermes`
- Memory files: `~/.hermes/memories/MEMORY.md`, `~/.hermes/memories/USER.md`
@@ -0,0 +1,41 @@
+++
title = "AI 日报 · 示例"
date = 2026-04-01T16:55:00+08:00
summary = "围绕 Agent 架构分层、Skills 标准化与桌面 Agent 工程实践的当日观察。"
+++
> ⚠️ 格式规范(生成 Hugo 时必须遵守):
> - 文章编号:`1.` `2.` `3.`(阿拉伯数字 + 点),禁止 `① ② ③` / `一、二、三` 等变体
> - 四个 section 缺一不可:`今日概览` → `今日重点` → `趋势观察` → `延伸阅读`
> - 每篇文章结构:标题 → 摘要段 → "值得关注:"三点 → "这篇更值得关注的理由"段
# 今日概览
今天的公开候选主要集中在 AI Agent 的架构演进、工具化落地与工程化实践三条线索上。相比早期偏概念展示的讨论,这一批内容更强调模块化能力栈、真实部署路径与系统可维护性,说明行业关注点正在从“模型能做什么”转向“系统如何稳定落地并持续复用”。
## 今日重点
### 1. 学习笔记:从 Agent 到 Skills — AI 智能体架构的范式转变
文章分析了 AI 智能体架构从单体 Agent 向模块化 Skills 的范式转变。Anthropic 先后推出 MCP 和 Agent Skills 开放标准,构建了知识、工具、协作和运行分层架构。文章通过一个自动化美化相册的真实项目,对比了 Claude Code 与 OpenClaw 两种实现方案,验证了新架构的可复用性与灵活性。
值得关注:
- Anthropic 在 14 个月内先后推出 MCP 和 Agent Skills 两个开放标准,推动 AI 智能体架构分层化。
- 新范式核心是构建薄 Agent 引擎与可组合的 Skills 库,取代为每个用例定制单体 Agent。
- 文章通过自动化美化相册项目,实操演示了 Skills、MCP、OpenClaw 和 A2A 协议如何协同工作。
这篇内容更值得关注的原因在于,它不只是提出了“Agent 要模块化”这个判断,而是把开放标准、分层架构和真实项目案例串成了一条完整论证链,能直接支撑今天日报的主线。
## 趋势观察
1. Agent 正在从单体能力转向可组合的模块化体系。无论是 Skills、MCP、记忆还是运行时编排,这批内容都在强调解耦与复用,而不是把智能体继续当成一个不可拆分的黑箱。
2. 工程化正在变成 AI 应用竞争的主战场。桌面 Agent、企业级架构和部署实践类内容增多,说明真正的差异化开始落在接入现有流程、控制风险和提升可维护性上。
3. AI 能力的竞争点正在上移。模型本身仍重要,但真正可持续的优势越来越来自系统设计、工作流整合和对业务场景的理解。
## 延伸阅读
- [学习笔记:从 Agent 到 Skills — AI 智能体架构的范式转变](https://example.com/a)|阿里云开发者
- [Agent Skills:打通可复用专业领域知识的最后一公里](https://example.com/b)|阿里云开发者
- [CoPaw深度解析:源码架构和功能实践](https://example.com/c)|阿里云开发者
> ⚠️ 延伸阅读必须包含当天所有入选文章的原文链接(`- [标题](url)|来源`),一条对应一篇今日重点文章。不要放未入选或 pipeline drop 的文章。这是日报读者获取原文的入口,不是"其他相关阅读"区。
@@ -0,0 +1,64 @@
# Tag 去重合并计划 (2026-07-22)
## 问题
https://osiman.site/tags/ 页面存在大量重叠 tag,如 `Harness工程化` / `Harness工程` / `Harness Engineering` 三个 tag 指向同一概念。
## 方案:数据层合并(Option A)
不修改展示层,直接合并所有日报 md 文件中的 tags。Hugo 自动重建 tag 页面,旧 tag 自动废弃。
## 规范映射表
| 废弃 tag | → | 规范名 | 说明 |
|:---------|:---|:-------|:----|
| `Harness工程化`, `Harness Engineering` | → | `Harness工程` | 三合一 |
| `Skills` | → | `Skill` | 单复数 |
| `AI Coding`, `AI代码生成`, `AI工程化` | → | `AI Coding Agent` | 统一为 Agent 维度 |
| `Agent 框架`, `Agentic架构`, `Agent工程` | → | `Agent工程` | 三合一 |
| `Prompt` | → | `Prompt Engineering` | 从简写改全称 |
| `多Agent`, `Multi-Agent架构`, `多Agent协作` | → | `多Agent` | 三合一 |
| `LLM`, `LLM评估`, `LLM训练`, `大模型应用开发` | → | `LLM` | 四合一 |
| `循环工程`, `Loop Engineering`, `Agent Loop` | → | `循环工程` | 三合一 |
| `上下文管理`, `Context工程` | → | `上下文管理` | 统一中文 |
| `Code Review`, `代码质量` | → | `Code Review` | 统一英文 |
| `推理加速`, `长文本推理`, `多步推理` | → | `推理加速` | 三合一 |
| `安全`, `安全防御` | → | `安全` | 二合一 |
| `技能系统`, `知识工程`, `知识管理` | → | `知识管理` | 三合一 |
| `Agent`, `AI` | → | (删除) | 太泛,无信息量 |
| `工程化` | → | (删除) | 冗余 |
| `架构` | → | (删除) | 冗余 |
## 保留的独立 tag
`MCP`, `RAG`, `ACP`, `KV Cache`, `MoE架构`, `State Lake`, `RL决策训练`, `Token管理`, `Vibe Coding`, `Spec工程`, `Skill流水线`, `Hook 链`, `全双工语音交互`, `多端架构`, `契约化架构`, `数据`, `工具链`, `工作流`, `评测`, `部署`, `搜索`, `AI搜索`, `AI原生研发`, `AI Coding Agent`, `企业落地`, `记忆`, `开源`
## 执行方式
用 Python 脚本扫描 `content/daily/` 下所有 md 文件,对每个文件的 `tags = [...]` 替换为规范名版本。脚本参考:
```python
import re, os, json
MAPPING = {
"Harness工程化": "Harness工程",
"Harness Engineering": "Harness工程",
"Skills": "Skill",
# ... 完整映射
}
REMOVE = {"Agent", "AI", "工程化", "架构", ...}
hugo_dir = "/home/ubuntu/zhu/apps/hugo-site/content/daily"
for root, _, files in os.walk(hugo_dir):
for fname in files:
if not fname.endswith(".md"):
continue
path = os.path.join(root, fname)
with open(path, 'r') as f:
content = f.read()
# parse tags from frontmatter
# replace deprecated → canonical
# remove items in REMOVE
# deduplicate
# write back
```