docs: sync reader-digest-flow skill with Hermes version (absolute path, extract_failed handling, rerun include_read)
This commit is contained in:
@@ -123,6 +123,7 @@ Hugo 发布选择与知识沉淀选择相互独立。询问用户哪些文章值
|
|||||||
3. 轮询 `get_article_summary_job_status`,成功后读取 `get_article_summary_job_result`;
|
3. 轮询 `get_article_summary_job_status`,成功后读取 `get_article_summary_job_result`;
|
||||||
4. 只使用现有 `article.plain_text`,不重新抓取原 URL;
|
4. 只使用现有 `article.plain_text`,不重新抓取原 URL;
|
||||||
5. 使用返回的 `written_paths` 定位结果并检查 Markdown 内容。
|
5. 使用返回的 `written_paths` 定位结果并检查 Markdown 内容。
|
||||||
|
6. **`extracted_path` 必须传绝对路径**(前缀 `/home/ubuntu/zhu/github/reader/`):summary-mcp 工作目录是 `/root/.hermes`,相对路径会报 `extracted_path does not exist`。同一工具连续 3 次失败会触发 MCP 冷却(约 45-60s,报 `MCP server 'reader' is unreachable`),等待冷却后再重试,不要循环重试同一调用。
|
||||||
|
|
||||||
异步 MCP 不可用时才使用项目 CLI fallback。不要因为内容较短而引入原文之外的知识。
|
异步 MCP 不可用时才使用项目 CLI fallback。不要因为内容较短而引入原文之外的知识。
|
||||||
|
|
||||||
@@ -131,8 +132,9 @@ Hugo 发布选择与知识沉淀选择相互独立。询问用户哪些文章值
|
|||||||
上传前按需读取:
|
上传前按需读取:
|
||||||
|
|
||||||
- 格式规则:`references/ima-format-quickref.md`
|
- 格式规则:`references/ima-format-quickref.md`
|
||||||
- API 和上传步骤:`references/ima-upload-api.md`
|
- API 和上传步骤:`references/ima-upload-api.md`(含 `-200` 版本拦截修复)
|
||||||
- 凭证定位:`references/ima-credential-chain.md`
|
- 凭证定位:`references/ima-credential-chain.md`
|
||||||
|
- 批量上传脚本:`scripts/ima_upload_one.py`(Python 编排,规避中文文件名 bash 引号问题)
|
||||||
|
|
||||||
硬规则:
|
硬规则:
|
||||||
|
|
||||||
@@ -154,6 +156,12 @@ Hugo 发布选择与知识沉淀选择相互独立。询问用户哪些文章值
|
|||||||
|
|
||||||
`reader-digest-flow` 不直接编辑 `term_aliases.json`、`term_stopwords.json` 或兴趣配置。简要路由见 `references/keyword-engine-maintenance.md`。
|
`reader-digest-flow` 不直接编辑 `term_aliases.json`、`term_stopwords.json` 或兴趣配置。简要路由见 `references/keyword-engine-maintenance.md`。
|
||||||
|
|
||||||
|
## 环境坑位(本机部署)
|
||||||
|
|
||||||
|
- **MCP 相对路径陷阱**:`summary-mcp` 服务工作目录是 `/root/.hermes`,不是 reader 项目根。MCP 返回的 `output_dir`/`artifact.path` 是相对路径,直接传给 `start_article_summary_job(extracted_path=...)` 会报 `extracted_path does not exist`。传入前必须拼绝对路径前缀 `/home/ubuntu/zhu/github/reader/`。
|
||||||
|
- **提取失败不等于运行失败**:`status_counts.extract_failed` 的条目(`CONTENT_EXTRACTION_FAILED`,`retryable=false`)跳过即可并如实汇报;失败文章常是推广/活动等低价值内容,不因此自行重跑。`linked_run_status=partial` 时先读 run-report 的 item 级 `error` 确认原因。
|
||||||
|
- **用户要求"重新跑一批"**:候选质量低(用户主动提出)时重跑,应 `include_read=true` 并调高 `limit`(如 10),否则默认 `include_read=false` 会拉回同一批未读文章。重跑是新 Run,候选编号体系重新建立,汇报时提醒用户按新列表选择。
|
||||||
|
|
||||||
## 停止与人工介入
|
## 停止与人工介入
|
||||||
|
|
||||||
出现以下任一情况时停止自动流程并报告:
|
出现以下任一情况时停止自动流程并报告:
|
||||||
@@ -167,11 +175,12 @@ Hugo 发布选择与知识沉淀选择相互独立。询问用户哪些文章值
|
|||||||
|
|
||||||
## Reference 路由
|
## Reference 路由
|
||||||
|
|
||||||
- `references/flow.md`:具体 MCP 调用序列、候选映射、Hugo 发布和 IMA 主步骤。
|
- `references/flow.md`:具体 MCP 调用序列、候选映射(含 extracted_path 绝对路径、候选≠文件名顺序)、Hugo 发布和 IMA 主步骤。
|
||||||
- `references/content-extraction.md`:FreshRSS 内容来源与 `plain_text` 质量判断。
|
- `references/content-extraction.md`:FreshRSS 内容来源与 `plain_text` 质量判断。
|
||||||
- `references/public-digest-example.md`:可直接参考的 Hugo 最终页面结构。
|
- `references/public-digest-example.md`:可直接参考的 Hugo 最终页面结构。
|
||||||
- `references/feishu-format-notes.md`:Feishu 输出格式限制。
|
- `references/feishu-format-notes.md`:Feishu 输出格式限制。
|
||||||
- `references/ima-format-quickref.md`:IMA Markdown 格式规则。
|
- `references/ima-format-quickref.md`:IMA Markdown 格式规则。
|
||||||
- `references/ima-upload-api.md`:IMA Markdown 文件上传 API。
|
- `references/ima-upload-api.md`:IMA Markdown 文件上传 API(含 `-200` 版本拦截修复)。
|
||||||
- `references/ima-credential-chain.md`:IMA 凭证与知识库配置定位。
|
- `references/ima-credential-chain.md`:IMA 凭证与知识库配置定位。
|
||||||
- `references/keyword-engine-maintenance.md`:关键词治理 Skill 路由。
|
- `references/keyword-engine-maintenance.md`:关键词治理 Skill 路由。
|
||||||
|
- `scripts/ima_upload_one.py`:单篇 Markdown 上传 daily 知识库的完整 Python 脚本(preflight→重名→create_media→COS→add_knowledge)。
|
||||||
|
|||||||
@@ -62,6 +62,12 @@ inspect_resume_plan
|
|||||||
|
|
||||||
不要按 `summary-batch`、`item-XX` 和候选数组的相同位置推断它们是同一篇文章。
|
不要按 `summary-batch`、`item-XX` 和候选数组的相同位置推断它们是同一篇文章。
|
||||||
|
|
||||||
|
### extracted 文件与候选编号错位(实测 2026-07-31)
|
||||||
|
|
||||||
|
- `digest-brief.json` 的 `top_candidates` **没有 `item_id` 字段**,只有 `url`;
|
||||||
|
- `extracted/item-XX.extracted.json` 的文件名序号与候选编号**可能不一致**(实例:候选2 = item-05、候选3 = item-02);
|
||||||
|
- 正确做法:用 **URL 交叉匹配**(归一化 `%3D`→`=` 后逐条比对),或用完整 `item_id`(从 candidate-batch.json 的 `items[i].item_id` 按候选数组顺序取)在 extracted 文件里反查;两者都能验证时优先 item_id。
|
||||||
|
|
||||||
## 3. Hugo 日报
|
## 3. Hugo 日报
|
||||||
|
|
||||||
用户确认发布文章后:
|
用户确认发布文章后:
|
||||||
@@ -99,6 +105,27 @@ http://127.0.0.1:14322/daily/YYYY-MM-DD/
|
|||||||
5. `selected_ids` 传完整、非空、去除 `cand:` 前缀后的 `item_id`;
|
5. `selected_ids` 传完整、非空、去除 `cand:` 前缀后的 `item_id`;
|
||||||
6. 从 Job 结果的 `written_paths` 读取 Markdown。
|
6. 从 Job 结果的 `written_paths` 读取 Markdown。
|
||||||
|
|
||||||
|
### ⚠️ extracted_path 必须用绝对路径
|
||||||
|
|
||||||
|
`summary-mcp` 进程的工作目录是 `/root/.hermes`(不是 reader 项目根)。传相对路径(如 `outputs/freshrss/...`)会直接报 `extracted_path does not exist`。必须传绝对路径:
|
||||||
|
|
||||||
|
```text
|
||||||
|
/home/ubuntu/zhu/github/reader/outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json
|
||||||
|
```
|
||||||
|
|
||||||
|
### ⚠️ 候选编号 ≠ extracted 文件名顺序
|
||||||
|
|
||||||
|
候选数组顺序与 extracted 文件名(`item-01`…`item-07`)**不一定对齐**(实测候选2 落在 item-05)。`digest-brief.json` 的候选**没有 `item_id` 字段**,只有 URL。可靠匹配方法:
|
||||||
|
|
||||||
|
1. 从 `candidate-batch.json` 取每项完整 `item_id`(在 `candidate` 嵌套对象里,顶层 `item_key` 只是 `item-XX` 文件名序号);
|
||||||
|
2. 或按 URL 匹配:归一化(`%3D`→`=`)后与每个 extracted 文件的 `article.url` / `article.canonical_url` 比对;
|
||||||
|
3. 绝不要按候选位置对应 extracted 文件序号。
|
||||||
|
|
||||||
|
```python
|
||||||
|
def norm(u): return u.replace('%3D','=').replace('%3d','=').strip()
|
||||||
|
# 对每个 extracted 文件取 norm(article.url),与候选 norm(url) 精确比对
|
||||||
|
```
|
||||||
|
|
||||||
正式序列:
|
正式序列:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
|
|||||||
@@ -32,6 +32,22 @@ content_type=text/markdown
|
|||||||
media_type=7
|
media_type=7
|
||||||
```
|
```
|
||||||
|
|
||||||
|
### ⚠️ IMA skill 版本拦截(-200)
|
||||||
|
|
||||||
|
`ima_api.cjs` 每天首次调用会检查更新,若检测到新版(如 1.1.8 > 当前 1.1.7)会以 `code=-200` 拦截原请求。注意:**官方 zip 包内的 `meta.json` 可能没同步版本号**(下载 1.1.8 zip 后 meta 仍写 1.1.7),所以光替换文件无法跳过拦截。
|
||||||
|
|
||||||
|
快速修复(脚本本身已是新版,只差版本号):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd /root/.hermes/skills/openclaw-imports/ima-skill && python3 -c "
|
||||||
|
import json
|
||||||
|
m = json.load(open('meta.json')); m['version'] = '1.1.8'
|
||||||
|
json.dump(m, open('meta.json','w'), ensure_ascii=False, indent=2)
|
||||||
|
"
|
||||||
|
```
|
||||||
|
|
||||||
|
先用 `diff -rq` 对比 zip 与安装目录:若只有 `.DS_Store`/meta 差异,说明代码已是最新,直接改 meta.json 版本号即可;若脚本有实质差异才需要整体替换。
|
||||||
|
|
||||||
## 2. Create Media
|
## 2. Create Media
|
||||||
|
|
||||||
```text
|
```text
|
||||||
@@ -89,3 +105,22 @@ POST /openapi/wiki/v1/add_knowledge
|
|||||||
上传成功不能只依据本地命令退出码。验证目标 knowledge base 中存在对应标题或返回对象,并向用户报告最终结果。
|
上传成功不能只依据本地命令退出码。验证目标 knowledge base 中存在对应标题或返回对象,并向用户报告最终结果。
|
||||||
|
|
||||||
失败时明确报告发生在 Preflight、Create Media、COS Upload 或 Add Knowledge 的哪一步,不改用 URL 导入或 Notes 类型绕过错误。
|
失败时明确报告发生在 Preflight、Create Media、COS Upload 或 Add Knowledge 的哪一步,不改用 URL 导入或 Notes 类型绕过错误。
|
||||||
|
|
||||||
|
## 6. 已知坑:IMA skill 版本拦截(-200)
|
||||||
|
|
||||||
|
`ima_api.cjs` 每天首次调用会检查远端版本,若发现新版本(如 1.1.8 > 1.1.7)会以 exit=1 + stderr `{"code":-200}` 拦截**所有** API 调用,原请求不发送。此前遇到过。
|
||||||
|
|
||||||
|
处理方式(不必整包替换):
|
||||||
|
|
||||||
|
1. 按 stderr 提示下载新版 zip(如 `https://app-dl.ima.qq.com/skills/ima-skills-1.1.8.zip`)并解压;
|
||||||
|
2. 对比新旧 `ima_api.cjs` 的 md5——zip 内核心脚本常与本地一致,只是 `meta.json` 的 `version` 未同步(zip 内仍写 1.1.7);
|
||||||
|
3. 若 `ima_api.cjs` 一致,只需把本地 `meta.json` 的 `version` 改为远端版本号即可跳过拦截,无需替换文件。
|
||||||
|
|
||||||
|
调用成功后再执行本文件前面的上传流程。
|
||||||
|
|
||||||
|
## 6. 版本拦截与批量上传实测(2026-07-31)
|
||||||
|
|
||||||
|
- **`-200` skill 更新拦截**:`ima_api.cjs` 每天首次调用检查版本,发现新版时以 code -200 退出并提示更新。下载 zip 后**先对比 `ima_api.cjs` 的 md5**——实测 zip 内脚本与已装版本完全一致,只是 `meta.json` 版本号未同步。此时只需把 `~/.hermes/skills/openclaw-imports/ima-skill/meta.json` 的 `version` 改为最新版即可跳过拦截,无需替换任何脚本。
|
||||||
|
- **Python 脚本编排上传**比 bash 可靠:bash 拼接含中文文件名/凭证的 curl 易出错。用 `subprocess` 参数数组依次调 `preflight-check.cjs` → `ima_api.cjs check_repeated_names` → `create_media` → `cos-upload.cjs`(`--secret-id/--secret-key/--token` 走参数数组,不打印)→ `add_knowledge`,每步解析返回 JSON,失败即停。
|
||||||
|
- **批量上传**:4 篇逐个跑同一脚本即可;同名文件先 `check_repeated_names` 确认无重复。
|
||||||
|
- 凭证从 `IMA_OPENAPI_CLIENTID` / `IMA_OPENAPI_APIKEY` 环境变量读取(`ima_api.cjs` 自动加载),KB ID 用 `IMA_DAILY_KNOWLEDGE_BASE_ID`。
|
||||||
|
|||||||
@@ -0,0 +1,96 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""IMA 上传单篇 Markdown 知识笔记到 daily 知识库。
|
||||||
|
用法: python3 ima_upload_one.py "<绝对路径/summary.md>" "<完整文章标题>"
|
||||||
|
依赖环境变量: IMA_OPENAPI_CLIENTID / IMA_OPENAPI_APIKEY / IMA_DAILY_KNOWLEDGE_BASE_ID
|
||||||
|
流程: preflight -> check_repeated_names -> create_media -> cos-upload -> add_knowledge
|
||||||
|
说明: 用 Python 而非 bash 编排,避免中文文件名/引号转义问题。
|
||||||
|
退出码 2 = 文件名重复(需与用户确认保留双方或取消),非 0 均为失败。
|
||||||
|
"""
|
||||||
|
import json, os, subprocess, sys
|
||||||
|
|
||||||
|
SKILL_DIR = "/root/.hermes/skills/openclaw-imports/ima-skill"
|
||||||
|
IMA_API = os.path.join(SKILL_DIR, "ima_api.cjs")
|
||||||
|
COS_UPLOAD = os.path.join(SKILL_DIR, "knowledge-base/scripts/cos-upload.cjs")
|
||||||
|
PREFLIGHT = os.path.join(SKILL_DIR, "knowledge-base/scripts/preflight-check.cjs")
|
||||||
|
|
||||||
|
def run_node(script, args):
|
||||||
|
r = subprocess.run(["node", script] + args, capture_output=True, text=True, timeout=120)
|
||||||
|
if r.returncode != 0:
|
||||||
|
raise RuntimeError(f"{script} exit={r.returncode} stderr={r.stderr[:500]}")
|
||||||
|
return json.loads(r.stdout)
|
||||||
|
|
||||||
|
def ima_api(api_path, body):
|
||||||
|
r = subprocess.run(["node", IMA_API, api_path, json.dumps(body, ensure_ascii=False)],
|
||||||
|
capture_output=True, text=True, timeout=120)
|
||||||
|
if r.returncode != 0:
|
||||||
|
raise RuntimeError(f"ima_api {api_path} exit={r.returncode} stderr={r.stderr[:500]}")
|
||||||
|
resp = json.loads(r.stdout)
|
||||||
|
if resp.get("code") != 0:
|
||||||
|
raise RuntimeError(f"ima_api {api_path} code={resp.get('code')} msg={resp.get('msg')}")
|
||||||
|
return resp.get("data", {})
|
||||||
|
|
||||||
|
def main():
|
||||||
|
kb_id = os.environ["IMA_DAILY_KNOWLEDGE_BASE_ID"]
|
||||||
|
file_path = sys.argv[1]
|
||||||
|
title = sys.argv[2] # 完整文章标题(不含 .md)
|
||||||
|
|
||||||
|
pf = run_node(PREFLIGHT, ["--file", file_path])
|
||||||
|
if not pf.get("pass"):
|
||||||
|
raise RuntimeError(f"preflight failed: {pf}")
|
||||||
|
file_name = pf["file_name"]; media_type = pf["media_type"]
|
||||||
|
content_type = pf["content_type"]; file_size = pf["file_size"]; file_ext = pf["file_ext"]
|
||||||
|
print(f"[preflight] ok file={file_name} ext={file_ext} size={file_size} media_type={media_type}")
|
||||||
|
|
||||||
|
dup = ima_api("openapi/wiki/v1/check_repeated_names", {
|
||||||
|
"params": [{"name": file_name, "media_type": media_type}],
|
||||||
|
"knowledge_base_id": kb_id
|
||||||
|
})
|
||||||
|
is_rep = dup.get("results", [{}])[0].get("is_repeated", False) if dup.get("results") else False
|
||||||
|
if is_rep:
|
||||||
|
print(f"[check_repeated_names] REPEATED: {file_name} — 需要处理")
|
||||||
|
sys.exit(2)
|
||||||
|
print("[check_repeated_names] no duplicate")
|
||||||
|
|
||||||
|
cm = ima_api("openapi/wiki/v1/create_media", {
|
||||||
|
"file_name": file_name,
|
||||||
|
"file_size": file_size,
|
||||||
|
"content_type": content_type,
|
||||||
|
"knowledge_base_id": kb_id,
|
||||||
|
"file_ext": file_ext
|
||||||
|
})
|
||||||
|
media_id = cm["media_id"]
|
||||||
|
cos = cm["cos_credential"]
|
||||||
|
print(f"[create_media] media_id={media_id} cos_bucket={cos.get('bucket_name')} cos_key={cos.get('cos_key','')[:40]}")
|
||||||
|
|
||||||
|
r = subprocess.run(["node", COS_UPLOAD,
|
||||||
|
"--file", file_path,
|
||||||
|
"--secret-id", cos["secret_id"],
|
||||||
|
"--secret-key", cos["secret_key"],
|
||||||
|
"--token", cos["token"],
|
||||||
|
"--bucket", cos["bucket_name"],
|
||||||
|
"--region", cos["region"],
|
||||||
|
"--cos-key", cos["cos_key"],
|
||||||
|
"--content-type", content_type,
|
||||||
|
"--start-time", str(cos.get("start_time", "")),
|
||||||
|
"--expired-time", str(cos.get("expired_time", "")),
|
||||||
|
"--timeout", "300000"
|
||||||
|
], capture_output=True, text=True, timeout=360)
|
||||||
|
if r.returncode != 0:
|
||||||
|
raise RuntimeError(f"cos-upload exit={r.returncode} stderr={r.stderr[:800]}")
|
||||||
|
print(f"[cos-upload] ok rc=0 stdout={r.stdout.strip()[:200]}")
|
||||||
|
|
||||||
|
ak = ima_api("openapi/wiki/v1/add_knowledge", {
|
||||||
|
"media_type": media_type,
|
||||||
|
"media_id": media_id,
|
||||||
|
"title": title,
|
||||||
|
"knowledge_base_id": kb_id,
|
||||||
|
"file_info": {
|
||||||
|
"cos_key": cos["cos_key"],
|
||||||
|
"file_size": file_size,
|
||||||
|
"file_name": file_name
|
||||||
|
}
|
||||||
|
})
|
||||||
|
print(f"[add_knowledge] ok media_id={ak.get('media_id') or media_id}")
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
Reference in New Issue
Block a user