docs: sync reader-digest-flow skill with Hermes version (absolute path, extract_failed handling, rerun include_read)

This commit is contained in:
root
2026-08-03 09:55:19 +08:00
parent 807027976e
commit 8e27cca166
4 changed files with 170 additions and 3 deletions
@@ -62,6 +62,12 @@ inspect_resume_plan
不要按 `summary-batch`、`item-XX` 和候选数组的相同位置推断它们是同一篇文章。
### extracted 文件与候选编号错位(实测 2026-07-31)
- `digest-brief.json` 的 `top_candidates` **没有 `item_id` 字段**,只有 `url`;
- `extracted/item-XX.extracted.json` 的文件名序号与候选编号**可能不一致**(实例:候选2 = item-05、候选3 = item-02);
- 正确做法:用 **URL 交叉匹配**(归一化 `%3D`→`=` 后逐条比对),或用完整 `item_id`(从 candidate-batch.json 的 `items[i].item_id` 按候选数组顺序取)在 extracted 文件里反查;两者都能验证时优先 item_id。
## 3. Hugo 日报
用户确认发布文章后:
@@ -99,6 +105,27 @@ http://127.0.0.1:14322/daily/YYYY-MM-DD/
5. `selected_ids` 传完整、非空、去除 `cand:` 前缀后的 `item_id`;
6. 从 Job 结果的 `written_paths` 读取 Markdown。
### ⚠️ extracted_path 必须用绝对路径
`summary-mcp` 进程的工作目录是 `/root/.hermes`(不是 reader 项目根)。传相对路径(如 `outputs/freshrss/...`)会直接报 `extracted_path does not exist`。必须传绝对路径:
```text
/home/ubuntu/zhu/github/reader/outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json
```
### ⚠️ 候选编号 ≠ extracted 文件名顺序
候选数组顺序与 extracted 文件名(`item-01`…`item-07`)**不一定对齐**(实测候选2 落在 item-05)。`digest-brief.json` 的候选**没有 `item_id` 字段**,只有 URL。可靠匹配方法:
1. 从 `candidate-batch.json` 取每项完整 `item_id`(在 `candidate` 嵌套对象里,顶层 `item_key` 只是 `item-XX` 文件名序号);
2. 或按 URL 匹配:归一化(`%3D`→`=`)后与每个 extracted 文件的 `article.url` / `article.canonical_url` 比对;
3. 绝不要按候选位置对应 extracted 文件序号。
```python
def norm(u): return u.replace('%3D','=').replace('%3d','=').strip()
# 对每个 extracted 文件取 norm(article.url),与候选 norm(url) 精确比对
```
正式序列:
```text