refactor: simplify reader digest skill

This commit is contained in:
zhuyongxin
2026-07-28 19:14:09 +08:00
parent 5eb390e3ed
commit 6dd8cef347
14 changed files with 388 additions and 1859 deletions
@@ -37,10 +37,7 @@ To improve quality: either find feeds that provide full `<content:encoded>`, or
## Quick check
```bash
# Check content_source for latest run
grep -h "content_source" outputs/freshrss/rerun/*/extracted/item-*.json | sort | uniq -c
```
先调用 `list_run_artifacts(run_id)`,再读取返回的 extracted Artifact 路径,检查 `content_source` 和 `article.plain_text`。不要按 `run_id` 或“最新目录”手拼路径。
## Relevant code paths
+105 -283
View File
@@ -1,328 +1,150 @@
# Reader Digest Flow Reference
# Reader Digest Flow 操作参考
## Purpose
本文件承载 `reader-digest-flow` 的具体执行步骤。正式行为边界以 `../SKILL.md` 为准。
Concrete operational checklist for the `reader-digest-flow` skill.
## 1. 日报 Pipeline
## Default Operating Model
### 默认参数
### Layering
- `reader` layer:
- FreshRSS pull
- extraction
- summary/filter/payload generation
- selected-article summary capability
- OpenClaw / skill layer:
- public digest generation
- internal review digest generation
- Hugo publishing
- chat reporting
- user confirmation handling
- calling selected-article summaries
- IMA upload orchestration
- Hugo layer:
- public digest browsing and archive only
- IMA layer:
- long-term storage for selected article notes only
### Hard rules
- Do not upload the full digest to IMA.
- Upload only explicitly user-selected articles to IMA.
- Do not generate a digest without a real payload.
- Generate two views from the same payload: a public digest for Hugo and an internal review digest for chat/operator workflow.
- Do not expose internal review states or operator-facing labels in the public digest.
- Do not re-fetch original URLs for selected summaries; use existing extracted text.
- If the main pipeline fails, inspect the run first; when `inspect_resume_plan` says `recommended_action=resume`, continue via the async resume job path instead of stopping immediately.
- Always branch on reader's top-level reconciled `status`; treat `status_source` and `state_conflict` only as explanatory metadata.
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization, with fallback to main `LLM_*` only if needed.
- Actively report progress after each completed phase.
## Step-by-step checklist
### 1. Run reader pipeline
Use the formal MCP workflow path as the default production route.
Prefer MCP run/status/result operations over direct path stitching. Only fall back to CLI or direct file inspection for debug / manual troubleshooting.
Formal production startup sequence:
1. `start_freshrss_pipeline_job`
2. `get_freshrss_pipeline_job_status`
3. `get_freshrss_pipeline_job_result`
4. after success, continue with `run_id` via `get_run_status` / `get_delivery_payload` / `get_run_report`
If the main pipeline job ends in `failed`:
1. inspect the linked run with `get_run_status`
2. call `inspect_resume_plan(run_id)`
3. if `recommended_action=resume`, continue with:
- `start_resume_job`
- `get_resume_job_status`
- `get_resume_job_result`
4. if `recommended_action=read_terminal_result`, continue from the terminal run result
5. if `recommended_action=start_new_run`, stop and report the failure
Treat the old synchronous `run_freshrss_openclaw_pipeline` as debug / light validation / fallback only.
Project root:
```bash
/home/ubuntu/zhu/github/reader
```json
{
"limit": 7,
"include_read": false,
"mark_read": true,
"debug_artifacts": false,
"timeout_seconds": 60,
"max_retries": 2
}
```
Default behavior for a normal production run:
- if the user did not specify a count, randomly choose a limit between 5 and 10 items for that run
- run with mark-read enabled
- do not enable `debug_artifacts`
- only skip mark-read if the user explicitly says the run is debug, test, or validation
- only enable `debug_artifacts` if the user explicitly says the run is debug, test, validation, or troubleshooting
Typical artifacts to inspect after a successful run:
### 正式调用序列
```text
outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json
outputs/freshrss/rerun/<run-id>/candidates/digest-brief.json
outputs/freshrss/rerun/<run-id>/run-report.json
outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json
start_freshrss_pipeline_job
→ get_freshrss_pipeline_job_status
→ get_freshrss_pipeline_job_result
→ get_run_status
→ get_delivery_payload / get_run_report
```
Notes:
状态动作:
- `digest-brief.json` is the preferred input for **public digest** generation.
- It is a lighter public-only view and currently includes only `keep` candidates.
- If `digest-brief.json` is missing, fall back to `openclaw-delivery-payload.json`.
- If the synchronous MCP wrapper times out but a real reader run was still created, do not discard the run; continue from run truth using `list_runs`, `get_run_report`, and `get_delivery_payload`.
- `running`:按合理间隔继续轮询;
- `success`:读取结果,保存 `run_id`;
- `failed`:读取关联 Run 并执行 Resume Plan;
- `partial`:优先检查 Run Report、Artifact 和 Recovery 信息。
**⚠️ Output path divergence when run from Feishu context:**
When triggered via MCP from inside a Feishu session, the extracted files may be written to an MCP-managed temp path instead of `outputs/freshrss/rerun/<run-id>/extracted/`. Verify the actual output path from the pipeline result before continuing to Phase 5/6. If the project-relative path doesn't exist, use the MCP result's `extracted_path` directly or copy the files to the expected project location.
### 2. Generate digest markdown
Generate two output views from the same run, preferably in one model call:
1. a **public digest** for Hugo / public readers
2. an **internal review digest** for chat / operator workflow
Input preference:
- **public digest**: prefer `candidates/digest-brief.json`
- **internal review digest**: use `candidates/openclaw-delivery-payload.json`
Recommended generation pattern:
- Pass the public brief and the full payload as two explicitly labeled input blocks.
- Ask the model to return both outputs in one response.
- Prefer a structured response shape (for example JSON with `public_digest_markdown` and `internal_review_digest_markdown`) when post-processing is needed.
Before publishing, persist the generated digest artifacts back into the same reader run directory:
恢复序列:
```text
outputs/freshrss/rerun/<run-id>/digest/public_digest.md
outputs/freshrss/rerun/<run-id>/digest/internal_review_digest.md
outputs/freshrss/rerun/<run-id>/digest/combined.json
inspect_resume_plan
→ recommended_action=resume
→ start_resume_job
→ get_resume_job_status
→ get_resume_job_result
```
Then write only the public digest into Hugo here:
如果建议为 `read_terminal_result`,直接读取结果;如果为 `start_new_run`,停止并等待用户决定。
不要根据 `run_id` 拼目录。后续只消费调用返回的 `output_dir`、`artifact.path`、`delivery_output`、`report_output`。
## 2. 候选汇报与选择
优先使用当前 Run 的 digest brief Artifact;不可用时使用 `get_delivery_payload` 返回的候选。
展示规则:
1. 使用候选数组原始顺序并从 1 编号;
2. 不因 `keep/review` 分组而重新编号;
3. 每篇展示标题、来源、摘要和判断理由;
4. 用户编号只映射当前候选数组;
5. 跨 Artifact 读取详情时按 URL 或完整 `item_id` 匹配。
不要按 `summary-batch`、`item-XX` 和候选数组的相同位置推断它们是同一篇文章。
## 3. Hugo 日报
用户确认发布文章后:
1. 读取 `public-digest-example.md`;
2. 仅使用用户选中的文章生成公开内容;
3. 写入 Hugo 当日页面;
4. 前台执行部署,避免把构建日志作为聊天通知;
5. 验证首页、列表页和详情页。
当前部署位置:
```text
/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md
```
Recommended public-digest front matter:
验证至少覆盖:
```toml
+++
title = "AI 日报 · YYYY-MM-DD"
date = YYYY-MM-DDTHH:MM:SS+08:00
summary = "当日日报摘要"
+++
```text
http://127.0.0.1:14322/
http://127.0.0.1:14322/daily/
http://127.0.0.1:14322/daily/YYYY-MM-DD/
```
Recommended **public digest** structure:
页面不存在时检查源文件、构建结果、容器状态和页面日期。需要 Docker/Hugo 深度排障时再读取部署项目自己的文档,不把历史事故规则复制回主 Skill。
- `今日概览`
- `今日重点`
- `趋势观察`
- `延伸阅读`
## 4. 单篇知识笔记
Public digest constraints:
用户确认 IMA 文章后:
- Use only the public brief view when present.
- Keep public wording free of internal workflow language.
- Optimize for concise public readability with solid information density.
- For each `今日重点` item, add one short editor-style sentence explaining why the item matters in today's digest (for example: `这篇内容更值得关注的原因在于……`).
- Prefer rendering highlight points as a short public label such as `值得关注:` followed by one bullet per line.
- **Article numbering: MUST use `1.` `2.` `3.` (Arabic numeral + period). Do NOT use `① ② ③`,`一、二、三`,`第一条` or any other variant.**
- **All four sections are required. Missing any one is a format violation.**
- **Hugo digest must include ALL keep articles from `digest-brief.json`. The IMA deposition subset is a separate downstream step.**
1. 从候选中取得 URL 和完整 `item_id`;
2. 从 Run Artifact 中找到匹配的 extracted 文件;
3. 交叉验证 `article.item_id` 或 URL;
4. 对每个单篇文件调用 `start_article_summary_job(extracted_path=<返回的 Artifact 路径>, selected_ids=[<完整 item_id>])`;
5. `selected_ids` 传完整、非空、去除 `cand:` 前缀后的 `item_id`;
6. 从 Job 结果的 `written_paths` 读取 Markdown。
Recommended **internal review digest** structure:
正式序列:
- `今日候选概况`
- `已入选重点`
- `待你确认`
- `建议沉淀到 IMA`
- `原始候选清单`
Internal review digest constraints:
- Do not show `rank` values.
- Replace machine states with Chinese labels such as `已入选` / `待确认` / `暂不纳入`.
- For `已入选重点`, include a fuller summary plus a short judgment paragraph.
- For `待你确认`, include a fuller summary, reason, and recommendation.
- Keep it readable as an operator review draft, not a raw payload dump.
- **Feishu compatibility: never use markdown tables in the digest report.**
When delivered via Feishu, a single table forces the whole message to plain
text. Use lists and sections instead. See `references/feishu-format-notes.md`.
### 3. Publish Hugo
Publish only the public digest to Hugo.
Treat Hugo publication as the default continuation of a successful normal daily digest run. Do not ask for a second confirmation before generating/writing the public digest and publishing it, unless the user explicitly requests not to publish to Hugo.
**⚠️ Hugo redeploy: always use foreground terminal, never background+notify_on_complete.**
The Hugo redeploy runs via `./redeploy.sh` and completes in ~10 seconds. **Do NOT use `terminal(background=true, notify_on_complete=true)`** for this step — the Gateway will push the raw Docker build output (compiler logs, layered build output, nginx config) to Feishu as a notification. This output is machine-readable, not human-readable, and the user has explicitly said this is noise.
**Correct approach:**
```python
# ✅ Foreground terminal with adequate timeout
cmd = "cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh"
result = terminal(cmd, timeout=120)
# Then verify 200 OK on the detail page
```text
start_article_summary_job
→ get_article_summary_job_status
→ get_article_summary_job_result
```
**Wrong approach (DO NOT use):**
内容依据是 extracted Artifact 中的 `article.plain_text`。不得重新抓取原 URL,不得使用原文之外的知识扩写。
```python
# ❌ Background + notify sends raw build logs to Feishu
terminal("cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh", background=True, notify_on_complete=True)
```
This rule only applies to Hugo redeploy (fast, deterministic output). For long-running Reader pipeline jobs (>60s), background + notify_on_complete is still appropriate since the output is meaningful content (article summaries, pipeline stats).
1. **Pre-check: verify Hugo content directory exists**
```bash
ls -d /home/ubuntu/zhu/apps/hugo-site/content/daily/ 2>/dev/null || mkdir -p /home/ubuntu/zhu/apps/hugo-site/content/daily/
```
If the entire `content/` directory is missing, create it before proceeding.
**Do not assume the Hugo site has a standard structure.**
2. write the public digest markdown to:
- `/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md`
3. redeploy Hugo immediately after writing:
- `cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh`
4. verify all three URLs before continuing:
- `http://127.0.0.1:14322/`
- `http://127.0.0.1:14322/daily/`
- `http://127.0.0.1:14322/daily/YYYY-MM-DD/`
Expected verification targets:
- homepage works
- `/daily/` works
- `/daily/YYYY-MM-DD/` works
### 4. Report digest in chat
Provide the internal review digest in chat and ask which articles should be retained.
Hard reporting rule:
- do not send only article titles
- for each article, include at least a one-sentence summary and a short recommendation / judgment so the user can decide without reopening the source
- **Feishu**: avoid markdown tables entirely. Use lists with headings.
### 5. Generate selected article summaries
Only do this after the user explicitly confirms which articles to retain.
Preferred MCP path:
- `start_article_summary_job`
- `get_article_summary_job_status`
- `get_article_summary_job_result`
Expected inputs:
- `extracted_path`
- `selected_ids`
**⚠️ item_id 前缀注意:**
The `candidates` array IDs use `cand:sha256:xxx` format, but individual extracted item files use `sha256:xxx` (no `cand:` prefix). When passing `selected_ids` to article-summary tools, strip the `cand:` prefix. If the ID doesn't match, the summary tool won't find the article.
**Path resolution for article-summary:**
After a pipeline run via MCP (Feishu context), the extracted files may live at an MCP-managed temp path, not the expected `outputs/freshrss/rerun/<run-id>/extracted/`. Before calling `generate_article_summaries` or `start_article_summary_job`, verify the extracted path exists. If not, use the `run_id` from the pipeline result to locate the actual output directory through `get_run_report`, or copy the files from the MCP-managed path.
Single-item rule:
- when `extracted_path` is `outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json`, call one summary job per file
- in that case, `selected_ids` should contain only the matching single `item_id`
- if the candidate ID came from the delivery payload, strip the `cand:` prefix before passing it
Recommended output layout:
- `outputs/freshrss/single_summaries/YYYY-MM-DD/`
- async job state: `outputs/freshrss/article_summary_jobs/<job_id>/`
Recommended production sequence:
1. call `start_article_summary_job`
2. poll `get_article_summary_job_status` until `status` becomes `success` or `failed`
3. on success, call `get_article_summary_job_result` and continue downstream from `written_paths`
Hard fallback rule:
- If the async MCP job path returns timeout / transport failure / job-launch failure (for example MCP timeout while the reader article-summary workflow itself is still healthy), do not treat that as article-summary business failure.
- Immediately retry through the local reader environment under `/home/ubuntu/zhu/github/reader` using the project `.venv`, calling the article-summary workflow directly.
- The production goal is successful generation of the selected-article Markdown files; async MCP job is preferred, but local `.venv` execution is the required fallback path.
Synchronous helper:
- `generate_article_summaries` remains available for debug / light validation only, not as the default production path.
CLI fallback:
CLI 仅在异步 MCP 不可用或人工排障时使用:
```bash
python scripts/run_article_summaries.py \
--extracted outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json \
--ids <item_id_without_cand_prefix> \
--output-dir outputs/freshrss/single_summaries/YYYY-MM-DD
--extracted <returned-extracted-path> \
--ids <full-item-id> \
--output-dir <explicit-output-dir>
```
### 6. Upload selected summaries to IMA
## 5. IMA 上传
Upload only the generated markdown files for the selected articles.
Once the user has selected the articles to retain, treat that selection itself as the authorization to continue the IMA deposition step; do not ask for a second confirmation about uploading into the knowledge base.
用户在知识沉淀阶段的文章选择即为上传授权。
**⚠️ IMA upload 500 retry: Always wrap IMA uploads in a retry loop.**
IMA's OpenAPI may return HTTP 500 on the first attempt. If the first upload fails, wait ~3 seconds and retry. Normally the second attempt succeeds. If `cos-upload.cjs` is used (for CDN-backed uploads), check subprocess stderr even when exit code is 0 — an HTTP error in the upload service may still produce exit code 0.
执行顺序:
Hard execution rules before upload:
1. 检查生成的 Markdown 与来源;
2. 文件名规范化为 `<完整文章标题>.md`;
3. 确认目标为 `daily` knowledge base;
4. 执行 preflight、create_media、COS upload、add_knowledge;
5. 验证知识库条目存在。
1. reformat/check the generated markdown into IMA-facing final content
2. normalize the final upload filename to `<文章标题>.md`
3. do not use internal temp names such as `ima-*`, `item-*`, `summary-*`, or English slug filenames as the final uploaded object name
4. if the knowledge base already contains the same filename, append a timestamp suffix before `.md`
5. if local work needs internal temp names, create a final upload copy with the user-facing title before calling IMA APIs
上传格式与 API 参数分别见:
Default target knowledge base:
- `ima-format-quickref.md`
- `ima-upload-api.md`
- `ima-credential-chain.md`
- `daily`
- Read `IMA_DAILY_KNOWLEDGE_BASE_ID` / `IMA_DAILY_KNOWLEDGE_BASE_NAME` from reader `.env`
- Verify the configured target at runtime before upload
- If the configured target is unavailable, resolve by name `daily`; if still absent, create `daily`
失败时记录发生在 preflight、create_media、COS upload 或 add_knowledge 的具体阶段。不要把失败的 Markdown 改写成 URL 导入或 Notes 类型来绕过错误。
## Hard rules recap
## 6. CLI fallback 原则
- Never generate a digest from placeholder or example data when a real run is expected.
- For normal runs, mark processed FreshRSS items as read unless the user explicitly requested a debug/test/validation run.
- Public digest goes to Hugo; internal review digest goes to chat; neither full digest goes to IMA.
- Only explicitly user-selected articles go to IMA.
- All daily IMA deposition must go directly into the IMA knowledge-base path as Markdown knowledge items (`media_type=7`), not through the IMA notes path.
- Uploading to IMA notes, or creating notes first and then linking them into a knowledge base, does not count as SOP completion.
- Selected article summaries use extracted text, not live refetch.
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization.
- Final IMA upload filenames must use user-facing article titles, not internal slugs or workflow temp names.
CLI 仅在以下场景使用:
- MCP 服务不可用;
- Tool transport/launch 失败且无法取得有效 Job;
- 用户明确要求本地调试;
- 人工排障需要直接检查脚本输出。
如果已取得 `job_id` 或 `run_id`,先查询真实状态,避免因响应超时重复启动任务。
@@ -1,55 +0,0 @@
# IMA COS 上传凭证处理(Hermes redact_secrets 兼容模式)
## 问题
Hermes 配置 `security.redact_secrets: true` 时,`create_media` API 返回的 `cos_credential.token` 字段在 `terminal()` 输出中被替换为 `***`。
直接通过 `terminal()` 调用 `cos-upload.cjs` 会因 token 截断而失败(HTTP 403 InvalidAccessKeyId)。
## 安全的工作流(execute_code + subprocess.run)
不要用 `terminal()` 传递 COS 凭证。改用 `execute_code()` + `subprocess.run()` 模式:
```python
# Phase A: terminal() 中保存原始响应到文件
result = terminal("""
bash -c '
set -a
source /home/ubuntu/zhu/github/reader/.env
set +a
OPTS=$(printf "%s" "{\\"clientId\\":\\""$IMA_OPENAPI_CLIENTID"\\",\\\"apiKey\\":\\\""$IMA_OPENAPI_APIKEY"\\\"}")
RESP=$(node /root/.hermes/skills/openclaw-imports/ima-skill/ima_api.cjs "openapi/wiki/v1/create_media" "{...}" "$OPTS" 2>/dev/null)
echo "$RESP" > /tmp/create_media_raw.json
echo "saved"
'
""")
# Phase B: execute_code 中从文件读取凭证
import json, subprocess
with open("/tmp/create_media_raw.json") as f:
data = json.load(f)
cred = data["data"]["cos_credential"]
# Phase C: subprocess.run 直接调用,不经过 terminal()
args = ["node", f"{SKILL_DIR}/knowledge-base/scripts/cos-upload.cjs",
"--file", FILE_PATH, "--secret-id", cred["secret_id"], "--secret-key", cred["secret_key"],
"--token", cred["token"], "--bucket", cred["bucket_name"], "--region", cred["region"],
"--cos-key", cred["cos_key"], "--content-type", "text/markdown",
"--start-time", cred["start_time"], "--expired-time", cred["expired_time"], "--timeout", "300000"]
r = subprocess.run(args, capture_output=True, text=True, timeout=310)
```
## 对比:错误的做法(terminal 直接传递)
```bash
# ⛔ 这样不行!token 会被 redact_secrets 替换为 ***
TOKEN=$(echo "$RESP" | jq -r '.data.cos_credential.token')
node ... --token "$TOKEN" ... # 会收到 HTTP 403
```
## 关键原则
- `terminal()` 输出中的敏感字段会被自动脱敏,但不影响底层 JSON 文件写入
- `execute_code` 中 `terminal()` 返回的 `output` 已经是脱敏后的文本
- **唯一可靠的凭证源**是直接写入磁盘的原始 JSON 文件
- `subprocess.run` 在 `execute_code` 中绕过脱敏,因为凭证在 Python 内存中直接被传递给子进程,不经过 Hermes 的 stdout 脱敏管道
@@ -1,46 +1,27 @@
# IMA 凭证链:从 reader `.env` 到 IMA 上传
# IMA 凭证与安全边界
## 凭证来源
## 必需配置
IMA 上传所需的凭证存储在多个位置,优先级如下:
- `IMA_OPENAPI_CLIENTID`
- `IMA_OPENAPI_APIKEY`
- `IMA_DAILY_KNOWLEDGE_BASE_ID`
- `IMA_DAILY_KNOWLEDGE_BASE_NAME=daily`
| 优先级 | 位置 | 说明 |
|--------|------|------|
| 1 | 环境变量 `IMA_OPENAPI_CLIENTID` / `IMA_OPENAPI_APIKEY` | Hermes session 级 |
| 2 | `~/.config/ima/client_id` / `api_key` | ima-skill 的默认检查路径 |
| 3 | `/home/ubuntu/zhu/github/reader/.env` | reader 项目配置,含完整的 IMA 凭证和 KB ID |
优先使用当前进程环境和 IMA Skill 已支持的凭证加载机制。不要在本 Skill 中复制、迁移或重写密钥文件。
## 凭证内容(reader .env 中)
## 缺失处理
```
IMA_OPENAPI_CLIENTID=<32位hex>
IMA_OPENAPI_APIKEY=<base64编码的API密钥>
IMA_DAILY_KNOWLEDGE_BASE_ID=<base64编码的KB ID>
IMA_DAILY_KNOWLEDGE_BASE_NAME=daily
```
Preflight 返回凭证缺失或目标知识库无法解析时:
## 缺失时的处理流程
1. 停止上传;
2. 只报告缺失的变量名或配置项;
3. 等待用户或运行环境补齐配置;
4. 配置恢复后重新执行 preflight,不重复生成知识笔记。
当 IMA 上传失败(`-100` 凭证缺失错误)时:
## 安全边界
1. 从 reader `.env` 读取凭证:
```
grep -E '^(IMA_OPENAPI_CLIENTID|IMA_OPENAPI_APIKEY)=' /home/ubuntu/zhu/github/reader/.env
```
2. 同步到 ima-skill 默认检查路径:
```
echo "<client_id>" > ~/.config/ima/client_id
echo "<api_key>" > ~/.config/ima/api_key
```
3. (可选)追加到 Hermes `.env` 以全局生效:
```
echo "IMA_OPENAPI_CLIENTID=<client_id>" >> /root/.hermes/.env
echo "IMA_OPENAPI_APIKEY=<api_key>" >> /root/.hermes/.env
echo "IMA_DAILY_KNOWLEDGE_BASE_ID=<kb_id>" >> /root/.hermes/.env
echo "IMA_DAILY_KNOWLEDGE_BASE_NAME=daily" >> /root/.hermes/.env
```
## 执行注意事项
- **COS 凭证红线**:`create_media` 返回的 `cos_credential` 中的 `token`/`secret_id`/`secret_key` 在 `terminal()` 输出中会被 Hermes 替换为 `***`。必须用 `subprocess.run()` 捕获原始输出,或用 `execute_code` 内联操作。
- **凭证格式**:`api_key` 是 base64 字符串(76 字符),`kb_id` 也是 base64 字符串。不要截断或转码。
- 不在聊天、日志或命令输出中打印完整 API key、KB ID 或 COS 临时凭证;
- `create_media` 返回的 COS 凭证仅在同一受控进程内传给上传工具,不写入磁盘;
- 不通过拼接 Shell 字符串传递凭证,使用参数数组或 IMA Skill 的封装;
- 不绕过 Hermes 的脱敏机制;若现有工具链无法安全传递凭证,停止并报告;
- 上传结束后不持久化 COS 临时凭证。
@@ -1,97 +0,0 @@
# IMA Markdown 文件上传流程(media_type=7)
## 什么时候用此流程
当用户说"沉淀到 IMA"时,必须用此文件上传流程,而不是 URL 导入或笔记导入。
## 完整流程
### 1. 写 .md 文件
```bash
mkdir -p /tmp/ima_upload
cat > /tmp/ima_upload/文章标题.md << 'EOF'
# 文章标题
> 来源:XXX
## 摘要
...
## 核心亮点
- ...
[原文链接](url)
EOF
```
### 2. preflight 检查
```python
pf = subprocess.run(
["node", f"{SKILL_DIR}/knowledge-base/scripts/preflight-check.cjs",
"--file", filepath],
capture_output=True, text=True
)
meta = json.loads(pf.stdout)
# meta = {pass, file_name, file_ext, file_size, media_type, content_type}
```
### 3. create_media(获取 COS 凭证)
```python
r = ima_api("openapi/wiki/v1/create_media", {
"file_name": fname,
"file_size": meta["file_size"],
"content_type": meta["content_type"],
"knowledge_base_id": kb_id,
"file_ext": meta["file_ext"]
})
cos = r["data"]["cos_credential"]
media_id = r["data"]["media_id"]
```
### 4. COS 上传
```python
cu = subprocess.run([
"node", f"{SKILL_DIR}/knowledge-base/scripts/cos-upload.cjs",
"--file", filepath,
"--secret-id", cos["secret_id"],
"--secret-key", cos["secret_key"],
"--token", cos["token"],
"--bucket", cos["bucket_name"],
"--region", cos["region"],
"--cos-key", cos["cos_key"],
"--content-type", meta["content_type"],
"--start-time", str(cos["start_time"]),
"--expired-time", str(cos["expired_time"]),
"--timeout", "300000"
], capture_output=True, text=True, timeout=30)
# 非0退出 = 上传失败
```
### 5. add_knowledge
```python
r = ima_api("openapi/wiki/v1/add_knowledge", {
"media_type": 7,
"media_id": media_id,
"title": title, # 文章中文标题
"knowledge_base_id": kb_id,
"file_info": {
"cos_key": cos["cos_key"],
"file_size": meta["file_size"],
"file_name": fname
}
})
```
## 注意事项
- **COS 凭证不能通过 terminal() 读取**(Hermes redact_secrets 会把 token 替换成 `***`)。必须用 Python `subprocess.run()` 捕获原始 stdout。
- **create_media 的 content_type 不能带 charset 参数**(如 `text/markdown; charset=utf-8` 会被拒)。用纯 MIME 类型 `text/markdown`。
- **文件命名 = `<文章中文标题>.md`**。不要用英文 slug 或 temp name。
- **media_type=7** 是 Markdown 文件。**绝对不要用 media_type=11**(笔记)或 `import_urls`。
@@ -1,42 +1,47 @@
# IMA 上传格式速查表(日报沉淀专用)
# IMA Markdown 格式速查
## 文件名 vs 标题 vs 内容对照表
## 文件与标题
| 项目 | ✅ 正确 | ❌ 错误 |
|------|---------|---------|
| **文件名** | `从Vibe Coding到Harness—— 一套大仓AI工程化实战.md` | `大仓AI工程化实战.md` |
| **add_knowledge title** | `从Vibe Coding到Harness—— 一套大仓AI工程化实战` | `大仓AI工程化实战` / `大仓AI工程化实战.md` |
| **上传方式** | `preflight` → `create_media` → `cos-upload` → `add_knowledge(media_type=7)` | `import_doc` + `add_knowledge(media_type=11)` |
| **正文内容** | pipeline `extracted/` 的 `article.content` 或 `summary-batch.json` 的 `summary` | 自己写的两三句话 |
| **正文长度** | ≥ 500 字 | < 500 字 |
- 文件名:`<完整文章标题>.md`
- `add_knowledge.title`:完整文章标题,不包含 `.md`
- 上传类型:Markdown 文件,`media_type=7`
- 目标:`daily` knowledge base
## 文件名常见错误模式
## 内容来源
| 原始标题 | ❌ 错误文件名 | ✅ 正确文件名 |
|----------|-------------|-------------|
| `从Vibe Coding到Harness—— 一套大仓AI工程化实战` | `大仓AI工程化实战.md` | `从Vibe Coding到Harness—— 一套大仓AI工程化实战.md` |
| `契约化多端架构:基于领域模型的Harness实践` | `契约化多端架构Harness实践.md` | `契约化多端架构:基于领域模型的Harness实践.md` |
| `Loop Engineering 实战:从日志扫描到预发部署的全自主闭环` | `LoopEngineering实战.md` | `Loop Engineering 实战:从日志扫描到预发部署的全自主闭环.md` |
只能使用当前 Run 的可追踪内容:
## 5 Section 格式要求
1. extracted Artifact 的 `article.plain_text`;
2. 对应文章的结构化摘要;
3. digest brief 的 summary 与 highlights。
| Section | 格式 | 说明 |
|---------|------|------|
| **核心结论** | 一段连贯段落(非分点) | 总结文章核心发现或主张 |
| **主要论点** | 一段连贯段落(非分点) | 综合多个论点形成连贯叙述 |
| **关键方法 / 机制** | `- **方法名**:详细说明` 分点 | 每条展开到能理解原理的程度 |
| **重要细节** | `- 每条一个带解释的完整知识点` 分点 | 每条是一个完整知识点,不是关键词 |
| **可复用启发** | `- 每条 actionable 的实践启示` 分点 | 附应用场景说明 |
不得使用通用知识补写原文没有的信息,不设置固定字数或字节数门槛。内容较短时保持简洁并忠于来源。
## 底部额外 section
## 标准结构
- `## 关键词` — 标签式关键词列表
- `## 主题` — 主题分类
```markdown
# 完整文章标题
## 速查口诀
Source: https://原文链接
Category: 分类
> 文件名 = 完整标题.md
> 内容取 pipeline,不自己写
> 5 个 section,核心结论和主要论点是段落
> media_type = 7,不是 11
> title 不加 .md
## 核心结论
## 主要论点
## 关键方法 / 机制
## 重要细节
## 可复用启发
## 关键词
## 主题
```
- 核心结论和主要论点使用连贯段落;
- 方法、细节和启发按完整知识点分项;
- 没有来源支持的 Section 可以简写,不得编造内容填充。
上传 API 见 `ima-upload-api.md`。
@@ -1,38 +0,0 @@
# AI Agent 的 Skill 系统设计
Source: https://mp.weixin.qq.com/s?__biz=MzAxNDEwNjk5OQ==&mid=2650544717&idx=1&sn=b578abf5a81034670900a3b8eb874296
Category: 方法论
## 核心结论
好的 Skill 应是一个小而准的行为系统,通过触发、加载、执行、约束、验证和迭代的组织,将通用 Agent 转化为在特定任务上稳定可靠的专用 Agent。核心原则是上下文窗口是公共资源,必须采用渐进披露、按任务风险设置自由度,并通过真实任务前向测试来证明行为改变。
## 主要论点
Skill 设计的本质是行为编程而非文档编写,需要将期望行为转化为 Agent 能稳定执行的结构化工作流。为此,必须同时解决发现(正确场景触发)、加载(最小上下文)、执行(合适自由度)和验证(真实任务测试)四件事,并通过门控、脚本外化、测试防合理化等机制确保 Agent 在复杂压力下不走捷径。
## 关键方法 / 机制
- 渐进披露的三层内容加载:元数据(frontmatter 的 name/description)用于发现,正文(SKILL.md)用于执行,资源(scripts/references/assets)按需读取。description 只做路由器,不包含完整流程,防止 Agent 凭印象执行。
- 门控机制(HARD-GATE):在低自由度任务中,使用明确的 <HARD-GATE> 标签禁止 Agent 在条件满足前执行后续动作,减少解释空间,让关键路径更像程序而非建议。
- 脚本外化降低上下文消耗和行为漂移:将需要确定性的操作(如 PDF 旋转)封装为 scripts/ 中的可执行脚本,避免 Agent 每次临时生成代码,提升可靠性和节省 token。
- 基于 TDD 的前向测试方法:用子代理模拟真实用户任务,只给原始任务和最少上下文,不泄露预期结论;观察行为轨迹、输出文件等原始证据,发现并封堵 Agent 的违规行为。
- 反模式自查表与检查表:交付前检查触发条件、自由度设置、资源引用、验证流程等关键项,确保 Skill 不是一份草稿而是一个可用的能力包。
- 跨平台适配与优雅降级:Skill 应写行为规则(如 TodoWrite),再通过平台层映射到具体工具名(如 todowrite);平台能力不足时优雅降级,保持 Skill 的可迁移性。
## 重要细节
- SKILL.md 的 frontmatter 和正文职责分离:name/description 用于发现(Agent 触发前可见),正文用于执行(触发后加载)。如果触发条件写在正文里,Agent 在决定是否触发时根本读不到。
- 命名规范:短、可触发、动词优先,例如 create-skill 比 skill-creation 更好,这本质上是路由质量——Agent 在技能库里找能力时,name/description 是第一层索引。
- 资源组织原则——“信息只放一个地方”:不要在 SKILL.md 和 references/ 中重复同一段规则,重复会带来漂移,导致 Agent 在两个版本间自行解释,增加维护成本。
- 门控类型示例:先决条件门控(先理解例子再编辑)、并发冲突门控、未保存内容门控、敏感操作门控(已创建 Skill 需处理影响再修改)。
- 流程图用 GraphViz DOT 嵌入 Markdown:对于包含非线性判断、循环、回退的步骤,流程图比纯文本更稳定,能防止 Agent 遗漏关键分支。
- 验证时防“合理化”问题:AI Agent 在压力下会为跳过规则编造理由,Skill 需要提前写出这些借口并给出反驳;审查循环应围绕真实失败风险而非措辞偏好。
## 可复用启发
- “上下文窗口是公共资源”原则:设计任何 Agent 指令时,每段内容都要质疑“Agent 真的需要这段解释吗?”和“值得占用的 token 成本吗?”,这适用于提示词、系统消息等所有 Agent 输入设计。
- 先收集具体例子再抽象 Skill:不要从抽象能力开始写,而是先收集用户会怎么触发、哪些请求应该触发/不应该触发、成功输出是什么等具体场景,避免写出宽泛不可执行的指令。
- 用脚本固化确定性任务、用门控防止关键路径走捷径:对于高脆弱、低变化空间的任务(如文件格式转换),应使用脚本而非描述性建议;对于必须按顺序执行的步骤,用门控打断 Agent 的“合理化”冲动。
- 设计防合理化的测试流程:用子代理模拟真实用户,只给原始任务,不泄露预期答案;观察是否存在只有看到结论才能成功的情况——如果这样,说明 Skill 不够清楚或测试设置泄露答案。
## 关键词
SKILL.md、YAML、Markdown、DOT、GraphViz、TDD、HARD-GATE、quick_validate.py
## 主题
AI Agent、行为编程、Token 经济、系统设计、约束机制
@@ -1,33 +0,0 @@
# IMA 笔记格式参考
> ⚠️ 格式基准文件为 `references/ima-format-reference.md`(`ai-agent-的-skill-系统设计.md`),每次生成 IMA 沉淀前必须先读该文件。
## 标准结构
### 顶部元数据
```markdown
# 文章完整标题
Source: https://原文链接(纯 URL,不加 "原文链接:" 标签)
Category: 分类
```
### 5 个必含 Section
1. **核心结论** — 一段总结文章核心发现或主张的段落(不是分点),像基准文件一样是一段连贯文字
2. **主要论点** — 一段概括文章核心论述的段落(不是分点列表),综合多个论点形成连贯叙述
3. **关键方法 / 机制** — `**方法名**:详细说明` 的格式,每条展开到能理解其原理的程度
4. **重要细节** — 每条一个带解释的完整知识点,不是关键词
5. **可复用启发** — 每条 actionable 的实践启示,附应用场景说明
### 底部额外 Section
- **## 关键词** — 标签式关键词列表
- **## 主题** — 主题分类
### 质量要求
- 每篇文章总内容量应达到 **4,000+ bytes**(基准文件 4,782B)
- 每个 section 展开到完整的知识点级别,不能只列关键词
- 内容来源:优先 `item-XX.extracted.json` 的 `article.content`,其次 `summary-batch.json` 的 `summary` 字段
@@ -1,80 +1,91 @@
# IMA Upload API Reference
# IMA Markdown 上传 API
Full API flow for uploading markdown articles to the IMA `daily` knowledge base. Used in Phase 7 of the reader-digest-flow.
用于将用户选中的单篇 Markdown 知识笔记上传到 `daily` knowledge base。
## Credentials
## 凭证
```
IMA_OPENAPI_CLIENTID - from reader .env or user-provided
IMA_OPENAPI_APIKEY - from reader .env or user-provided
IMA_DAILY_KNOWLEDGE_BASE_ID - daily KB UUID
- `IMA_OPENAPI_CLIENTID`
- `IMA_OPENAPI_APIKEY`
- `IMA_DAILY_KNOWLEDGE_BASE_ID`
- `IMA_DAILY_KNOWLEDGE_BASE_NAME`
凭证定位和恢复见 `ima-credential-chain.md`。不要在终端输出完整密钥。
## 上传前检查
- 文件名为 `<完整文章标题>.md`;
- `title` 为完整文章标题,不带 `.md`;
- Markdown 符合 `ima-format-quickref.md`;
- 内容可追溯到当前 Run Artifact;
- 用户已经明确选择该文章;
- 目标知识库已经解析并验证。
## 1. Preflight
调用 IMA Skill 的 `preflight-check.cjs` 检查文件类型、扩展名、大小和 MIME。
预期:
```text
file_ext=md
content_type=text/markdown
media_type=7
```
The `ima-skill` v1.1.7+ ships with `ima_api.cjs` for credential loading.
Legacy auth header: `ima-openapi-ctx: skill_version=1.1.7`.
## 2. Create Media
Do NOT export the full API key in shell commands — use `execute_code` with `subprocess.run` and Python string variables.
## Flow (3 steps)
### 1. create_media
```
POST https://ima.qq.com/openapi/wiki/v1/create_media
Headers: ima-openapi-clientid, ima-openapi-apikey, Content-Type: application/json
Body: { file_name, file_size, content_type, knowledge_base_id, file_ext }
Returns: { code: 0, data: { media_id, cos_credential: { secret_id, secret_key, token, bucket_name, region, cos_key, start_time, expired_time } } }
```text
POST /openapi/wiki/v1/create_media
```
`file_ext` is without the dot (e.g. `md` not `.md`).
`file_name` must be the user-facing article title + `.md`.
`content_type` for markdown is `text/markdown`; media_type=7.
请求核心字段:
### 2. COS upload
Use `cos-upload.cjs` from `ima-skill/knowledge-base/scripts/`:
```
node <skill_dir>/knowledge-base/scripts/cos-upload.cjs \
--file <local_md_file> \
--secret-id <from create_media> \
--secret-key <from create_media> \
--token <from create_media> \
--bucket <bucket_name> \
--region <region> \
--cos-key <cos_key> \
--content-type text/markdown \
--start-time <start_time> \
--expired-time <expired_time>
```json
{
"file_name": "<完整文章标题>.md",
"file_size": 0,
"content_type": "text/markdown",
"knowledge_base_id": "<daily-kb-id>",
"file_ext": "md"
}
```
⚠️ Must use Python `subprocess.run(args=[...])` to avoid shell parameter mangling.
⚠️ Always capture `returncode` and `stderr` — COS may return exit 0 on HTTP 500.
保存返回的 `media_id` 和 `cos_credential`。COS 临时凭证只在进程内传递,不打印到聊天或日志。
### 3. add_knowledge
## 3. COS Upload
```
POST https://ima.qq.com/openapi/wiki/v1/add_knowledge
Headers: same as create_media
Body: { media_type: 7, media_id, title: "<file_name>", knowledge_base_id, file_info: { cos_key, file_size, file_name } }
使用 IMA Skill 提供的 `cos-upload.cjs`,通过参数数组调用并检查:
- 进程 `returncode`;
- `stderr`;
- HTTP 上传结果。
不要拼接包含凭证的 Shell 字符串,也不要把多条 JSON 响应重定向到同一个文件。
## 4. Add Knowledge
```text
POST /openapi/wiki/v1/add_knowledge
```
`media_type=7` for markdown. `title` MUST equal `file_name`.
核心字段:
## Article Markdown reformatting (before upload)
Generated summaries from `reader` have `Source:` and `Category:` header lines.
Before uploading, reformat to IMA style:
```
原文链接:<original article URL>
## 核心结论
...
## 主要论点
...
```json
{
"media_type": 7,
"media_id": "<media-id>",
"title": "<完整文章标题>",
"knowledge_base_id": "<daily-kb-id>",
"file_info": {
"cos_key": "<cos-key>",
"file_size": 0,
"file_name": "<完整文章标题>.md"
}
}
```
Remove `Source:`, `Category:` lines. Keep `原文链接:` at top with the URL on the next line.
Break long prose (>200 chars per paragraph) into shorter paragraphs for IMA readability.
## 5. 验证
上传成功不能只依据本地命令退出码。验证目标 knowledge base 中存在对应标题或返回对象,并向用户报告最终结果。
失败时明确报告发生在 Preflight、Create Media、COS Upload 或 Add Knowledge 的哪一步,不改用 URL 导入或 Notes 类型绕过错误。
@@ -1,195 +1,29 @@
# 关键词引擎维护流程
# 关键词治理路由
## 概述
关键词治理不属于 `reader-digest-flow` 的日常执行阶段。
reader 项目内置了完整的关键词清洗链路,用于管理 `term_aliases.json`、`term_stopwords.json`、`filter_context.personal.json` 等配置。本文件说明链路各环节的职责与使用方式。
仅当用户明确要求“清理关键词”“词库治理”“生成关键词建议”时,委托:
## 完整数据流
```
每日日报 pipeline
│
▼
term_index/daily/YYYY-MM-DD.json ← build_keyword_index.py (pipeline 末环节自动)
│
▼
term_index/term_stats.json ← 全量汇总
│
▼
build_review_bundle.py ← 打包审查数据包 (手动触发)
│
▼
generate_term_cleanup_suggestions.py ← 生成建议 (手动触发)
│
▼
apply_term_suggestions.py ← 人工确认后落盘 (手动触发)
```text
skills/keyword-cleanup-review/SKILL.md
```
## 各环节命令
Review 输入生成由该 Skill 定义,后续确认与 Apply 由当前编排层负责:
### 1. 重建每日 term_index(pipeline 末环节自动执行,也可手动)
```bash
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 scripts/build_keyword_index.py \
--input outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json
```text
build review bundle
→ generate rule suggestions
→ optional semantic suggestions
→ human review
→ dry-run
→ apply accepted suggestions
```
### 2. 重建 review bundle(打包当前配置+统计供审查)
约束:
```bash
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
--days 365 \
--top 200 \
--output outputs/term_index/review/keyword-cleanup-bundle.json
```
- Suggestions JSON 是 Review 与 Apply 之间的正式契约;
- LLM 语义建议不能自动 Apply;
- 不直接编辑 aliases、stopwords、watchlist 或 interest 配置;
- 不在日报主流程中因 tag 质量不佳自动触发治理。
参数:
- `--days 365` — 考虑最近多少天的统计数据
- `--top 200` — 取前 N 个高频词纳入 bundle
### 3. 生成清洗建议
```bash
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 scripts/generate_term_cleanup_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
--emit-markdown
```
输出:
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json` — 正式建议产物
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md` — 人工阅读展示稿
当前脚本能生成的建议类型:
| 类型 | 生成规则 | 当前产出 |
|:----|:---------|:---------|
| interest_keyword_suggestions | 高频未覆盖词(total≥3, days≥2) | ✅ 按阈值产出 |
| watch_terms | 中频观察词(total≤2, days≤2) | ✅ 按阈值产出 |
| alias_suggestions | 大小写变体/单复数/空格连词符差异 | ✅ 统计规则产出 |
| stopword_suggestions | — | ❌ 脚本硬编码为空 |
### 4. 应用建议(dry-run → review → apply)
```bash
# 先 dry-run 预览
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 scripts/apply_term_suggestions.py \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--accept-interest 强化学习 ReAct CLI \
--dry-run
# 确认后正式 apply(去掉 --dry-run)
.venv/bin/python3 scripts/apply_term_suggestions.py \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--accept-interest 强化学习 ReAct CLI \
```
支持的 accept 参数:
- `--accept-interest 词1 词2 ...` — 添加到 interest_keywords
- `--accept-watch 词1 词2 ...` — 添加到 watchlist
- `--accept-alias 源词1 源词2 ...` — 添加 alias 映射
- `--accept-stopword 词1 词2 ...` — 添加停用词
## 脚本能做什么 vs 不能做什么
### ✅ 脚本能做的(统计级清洗)
- 发现大小写变体(`vibe coding` → `Vibe Coding`)
- 发现单复数差异(`Agent Skill` → `Agent Skills`)
- 发现空格/连词符差异
- 按频次推荐 should-be-interest / should-be-watch 的词
- 批量 apply 到配置文件,自动记录变更日志
### ❌ 脚本不能做的(语义级清洗,需要人工判断)
- 识别产品名(`WorkBuddy`、`LibTV Agent`、`飞书妙搭`)→ 应加 stopword
- 识别模型名(`Qwen3-30B-A3B`、`GLM 5.2`、`Opus 4.8`)→ 应加 stopword
- 识别人名/地名(`Andrej Karpathy`、`成都天府长岛`)→ 应加 stopword
- 英文专有名词归一化到中文主题词(`Multi-Agent` → `多Agent`)
- 长产品名归一化(`火山云数据库PostgreSQL Serverless版` → `Serverless数据库`)
- 判断某个词对 Hugo tag 是否有用(`1688`、`OPC训练营` → 无效)
## 语义级清洗:LLM 辅助(generate_term_cleanup_semantic_suggestions.py)
除了统计规则的清洗,项目还有一个 **LLM 驱动的语义级清洗脚本**,能发现统计规则做不到的事情:
```bash
cd /home/ubuntu/zhu/github/reader
.venv/bin/python3 scripts/generate_term_cleanup_semantic_suggestions.py \
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
--output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json
```
参数:
- `--bundle` — review bundle(先跑 build_review_bundle.py)
- `--suggestions` — 规则层的建议 JSON(可选,供 LLM 避免重复)
- `--output` — 输出路径(自动生成)
- `--dry-run` — 只打印 prompt 不调 LLM
### 语义脚本能做的(而统计规则不能做的)
| 类型 | LLM 能发现什么 | 示例 |
|:----|:--------------|:-----|
| **semantic_alias** | 中英文映射、简称↔全称、近义词 | `AI智能体`→`AI Agent`, `AI Harness`→`Harness Engineering` |
| **stopword** | 语义判断哪些词太泛/不相干 | `自动化`(太泛), `AlphaFold`(生物领域), `陌生化`(穿透词) |
| **promote_to_interest** | 与用户关注领域对齐的新词 | `强化学习`, `Loop Engineering`, `MoE`, `ChatGPT` |
### 两阶段清洗 SOP
当需要清洗关键词时,按此顺序操作:
**阶段一:规则统计清洗(脚本发现 + 人工确认)**
1. `build_review_bundle.py` → 重建 bundle
2. `generate_term_cleanup_suggestions.py` → 产出统计级建议
3. 检查建议,决定哪些 accept
4. `apply_term_suggestions.py --dry-run` → 预览
5. `apply_term_suggestions.py` → 正式落地
**阶段二:语义级清洗(LLM 发现 + 人工确认)**
1. `generate_term_cleanup_semantic_suggestions.py` → 调 LLM 产生语义建议
2. 检查 LLM 的 semantic_alias / stopword / promote_to_interest
3. 人工编辑 term_aliases.json / term_stopwords.json(语义级判断不能自动化)
**🔴 不要绕过确认流程直接改配置。** 先跑脚本看看发现了什么,再针对性改。未经审视的批量改配置会引入新问题。
## 完整的语义级清洗操作流程
当需要大量添加 aliases/stopwords 时,推荐流程:
1. **跑一次 `build_review_bundle.py` + `generate_term_cleanup_suggestions.py`** 看看统计建议
2. **走一遍实际 pipeline 产出**,收集所有 unique keywords:
```bash
for rid in $(ls outputs/freshrss/rerun/); do
python3 -c "
import json
d = json.load(open('outputs/freshrss/rerun/$rid/summary/summary-batch.json'))
for item in d['items']:
for kw in item['summary'].get('keywords', []):
print(kw)
"
done | sort -u
```
3. 按类别分组:正常词 / 产品名(stopword) / 模型名(stopword) / 需要 alias 的
4. 分别更新 `term_aliases.json` 和 `term_stopwords.json`
5. **重建 term_index**:对每个受影响日期的 run 跑 `build_keyword_index.py`
6. **更新 Hugo tags**:从重建后的 `data/term_index/daily/YYYY-MM-DD.json` 取
## 当前配置 (2026-07-16)
| 文件 | 条目数 |
|:----|:------|
| `configs/term_aliases.json` | 142 |
| `configs/term_stopwords.json` | 106 |
| `configs/filter_context.personal.json` | 54 (interest_keywords) |
| `configs/term_watchlist.json` | 6 |
## 关键词/别名/停用词配置的更新规范
- `term_aliases.json` — 只放"源词→目标词"的映射,目标是让不同写法的同一概念归一到标准形式
- `term_stopwords.json` — 放产品名、模型名、人名、地名等对 Hugo tag 无价值的词
- **随时可以加**,加了后重建 term_index 即可生效
- 不涉及 pipeline 重新跑——只影响下游展示
Review bundle、Suggestions、Schema 和产物保留策略以 `keyword-cleanup-review` 为唯一事实来源;该 Skill 不直接 Apply 配置。
@@ -1,60 +0,0 @@
# Memory Drift Recovery
## Symptom
`store_memory(action="add", ...)` fails with:
> Refusing to write MEMORY.md: file on disk has content that wouldn't round-trip through the memory tool...
A `.bak` snapshot is created: `/root/.hermes/memories/MEMORY.md.bak.<timestamp>`
## Root Cause
The MEMORY.md file format doesn't match what the memory tool expects — likely because the file was modified externally (by `patch`, `write_file`, shell `>>` append, or a concurrent session). The tool uses a `§` (section sign) delimited format internally and its serialization/deserialization doesn't match the on-disk content.
## Recovery Procedure
### Step 1: Read the backup and the current file
```bash
diff /root/.hermes/memories/MEMORY.md.bak.<timestamp> /root/.hermes/memories/MEMORY.md
```
### Step 2: Extract missing entries (if any)
```bash
# List entries from the backup
grep '^§' /root/.hermes/memories/MEMORY.md.bak.<timestamp>
```
### Step 3: Re-add each missing entry via store_memory
For each entry that was in the backup but is now gone from the current file:
```bash
store_memory(action="add", content="<entry text>", target="memory")
```
### Step 4: Reset to a clean state
If the file is completely corrupted, the cleanest path is:
1. Save any new entries from the backup you want to keep
2. Rewrite the file as a clean `§`-delimited list (one entry per `§` line)
3. The format is: `§<content>\n` per entry, with `---` or blank line separators
```bash
# Example clean format:
echo '§当前重要条目一
§当前重要条目二
§当前重要条目三' > /root/.hermes/memories/MEMORY.md
```
### Prevention
- Do NOT use `write_file` or `patch` to modify MEMORY.md directly — always use `store_memory()`
- Do NOT use shell `>>` to append to MEMORY.md
- If you must bulk-import, use `store_memory` per-entry, not file-level operations
## Environment
- Host: Linux (5.15)
- Hermes home: `/root/.hermes`
- Memory files: `~/.hermes/memories/MEMORY.md`, `~/.hermes/memories/USER.md`
@@ -1,41 +1,33 @@
+++
title = "AI 日报 · 示例"
date = 2026-04-01T16:55:00+08:00
summary = "围绕 Agent 架构分层、Skills 标准化与桌面 Agent 工程实践的当日观察。"
date = 2026-04-01T09:00:00+08:00
summary = "围绕 Agent 架构分层、Skills 标准化与工程实践的当日观察。"
+++
> ⚠️ 格式规范(生成 Hugo 时必须遵守):
> - 文章编号:`1.` `2.` `3.`(阿拉伯数字 + 点),禁止 `① ② ③` / `一、二、三` 等变体
> - 四个 section 缺一不可:`今日概览` → `今日重点` → `趋势观察` → `延伸阅读`
> - 每篇文章结构:标题 → 摘要段 → "值得关注:"三点 → "这篇更值得关注的理由"段
# 今日概览
今天的公开候选主要集中在 AI Agent 的架构演进、工具化落地与工程化实践三条线索上。相比早期偏概念展示的讨论,这一批内容更强调模块化能力栈、真实部署路径与系统可维护性,说明行业关注点正在从“模型能做什么”转向“系统如何稳定落地并持续复用”。
今天的公开内容主要集中在 AI Agent 架构演进、工具化落地与工程实践三条线索。行业关注点正在从“模型能做什么”转向“系统如何稳定落地并持续复用”。
## 今日重点
### 1. 学习笔记:从 Agent 到 Skills — AI 智能体架构的范式转变
### 1. 从 Agent 到 Skills:AI 智能体架构的范式转变
文章分析了 AI 智能体架构从单体 Agent 向模块化 Skills 的范式转变。Anthropic 先后推出 MCP 和 Agent Skills 开放标准,构建了知识、工具、协作和运行分层架构。文章通过一个自动化美化相册的真实项目,对比了 Claude Code 与 OpenClaw 两种实现方案,验证了新架构的可复用性与灵活性。
文章分析了 AI 智能体从单体 Agent 向模块化 Skills 的演进,并结合 MCP、Skills 和真实项目说明能力分层与复用方式。
值得关注:
- Anthropic 在 14 个月内先后推出 MCP 和 Agent Skills 两个开放标准,推动 AI 智能体架构分层化。
- 新范式核心是构建薄 Agent 引擎与可组合的 Skills 库,取代为每个用例定制单体 Agent。
- 文章通过自动化美化相册项目,实操演示了 Skills、MCP、OpenClaw 和 A2A 协议如何协同工作。
这篇内容更值得关注的原因在于,它不只是提出了“Agent 要模块化”这个判断,而是把开放标准、分层架构和真实项目案例串成了一条完整论证链,能直接支撑今天日报的主线。
- Skills 将领域流程从 Agent 主体中拆出,便于复用和维护。
- MCP 为 Agent 与外部工具提供标准化连接方式。
- 工程竞争点逐渐从模型调用转向状态、工具和工作流设计。
这篇内容值得关注的原因在于,它把开放协议、分层架构和真实落地案例连接成了完整论证链。
## 趋势观察
1. Agent 正在从单体能力转向可组合的模块化体系。无论是 Skills、MCP、记忆还是运行时编排,这批内容都在强调解耦与复用,而不是把智能体继续当成一个不可拆分的黑箱。
2. 工程化正在变成 AI 应用竞争的主战场。桌面 Agent、企业级架构和部署实践类内容增多,说明真正的差异化开始落在接入现有流程、控制风险和提升可维护性上。
3. AI 能力的竞争点正在上移。模型本身仍重要,但真正可持续的优势越来越来自系统设计、工作流整合和对业务场景的理解。
1. Agent 正在从单体能力转向可组合的模块化体系。
2. 工具契约、状态管理和验证机制正在成为 AI 应用的核心工程能力。
3. Human-in-the-loop 仍是控制高风险副作用的重要边界。
## 延伸阅读
- [学习笔记:从 Agent 到 Skills — AI 智能体架构的范式转变](https://example.com/a)|阿里云开发者
- [Agent Skills:打通可复用专业领域知识的最后一公里](https://example.com/b)|阿里云开发者
- [CoPaw深度解析:源码架构和功能实践](https://example.com/c)|阿里云开发者
> ⚠️ 延伸阅读必须包含当天所有入选文章的原文链接(`- [标题](url)|来源`),一条对应一篇今日重点文章。不要放未入选或 pipeline drop 的文章。这是日报读者获取原文的入口,不是"其他相关阅读"区。
- [从 Agent 到 Skills:AI 智能体架构的范式转变](https://example.com/a)|示例来源
@@ -1,64 +0,0 @@
# Tag 去重合并计划 (2026-07-22)
## 问题
https://osiman.site/tags/ 页面存在大量重叠 tag,如 `Harness工程化` / `Harness工程` / `Harness Engineering` 三个 tag 指向同一概念。
## 方案:数据层合并(Option A)
不修改展示层,直接合并所有日报 md 文件中的 tags。Hugo 自动重建 tag 页面,旧 tag 自动废弃。
## 规范映射表
| 废弃 tag | → | 规范名 | 说明 |
|:---------|:---|:-------|:----|
| `Harness工程化`, `Harness Engineering` | → | `Harness工程` | 三合一 |
| `Skills` | → | `Skill` | 单复数 |
| `AI Coding`, `AI代码生成`, `AI工程化` | → | `AI Coding Agent` | 统一为 Agent 维度 |
| `Agent 框架`, `Agentic架构`, `Agent工程` | → | `Agent工程` | 三合一 |
| `Prompt` | → | `Prompt Engineering` | 从简写改全称 |
| `多Agent`, `Multi-Agent架构`, `多Agent协作` | → | `多Agent` | 三合一 |
| `LLM`, `LLM评估`, `LLM训练`, `大模型应用开发` | → | `LLM` | 四合一 |
| `循环工程`, `Loop Engineering`, `Agent Loop` | → | `循环工程` | 三合一 |
| `上下文管理`, `Context工程` | → | `上下文管理` | 统一中文 |
| `Code Review`, `代码质量` | → | `Code Review` | 统一英文 |
| `推理加速`, `长文本推理`, `多步推理` | → | `推理加速` | 三合一 |
| `安全`, `安全防御` | → | `安全` | 二合一 |
| `技能系统`, `知识工程`, `知识管理` | → | `知识管理` | 三合一 |
| `Agent`, `AI` | → | (删除) | 太泛,无信息量 |
| `工程化` | → | (删除) | 冗余 |
| `架构` | → | (删除) | 冗余 |
## 保留的独立 tag
`MCP`, `RAG`, `ACP`, `KV Cache`, `MoE架构`, `State Lake`, `RL决策训练`, `Token管理`, `Vibe Coding`, `Spec工程`, `Skill流水线`, `Hook 链`, `全双工语音交互`, `多端架构`, `契约化架构`, `数据`, `工具链`, `工作流`, `评测`, `部署`, `搜索`, `AI搜索`, `AI原生研发`, `AI Coding Agent`, `企业落地`, `记忆`, `开源`
## 执行方式
用 Python 脚本扫描 `content/daily/` 下所有 md 文件,对每个文件的 `tags = [...]` 替换为规范名版本。脚本参考:
```python
import re, os, json
MAPPING = {
"Harness工程化": "Harness工程",
"Harness Engineering": "Harness工程",
"Skills": "Skill",
# ... 完整映射
}
REMOVE = {"Agent", "AI", "工程化", "架构", ...}
hugo_dir = "/home/ubuntu/zhu/apps/hugo-site/content/daily"
for root, _, files in os.walk(hugo_dir):
for fname in files:
if not fname.endswith(".md"):
continue
path = os.path.join(root, fname)
with open(path, 'r') as f:
content = f.read()
# parse tags from frontmatter
# replace deprecated → canonical
# remove items in REMOVE
# deduplicate
# write back
```