docs: sync reader-digest-flow skill with Hermes version (absolute path, extract_failed handling, rerun include_read)
This commit is contained in:
@@ -62,6 +62,12 @@ inspect_resume_plan
|
||||
|
||||
不要按 `summary-batch`、`item-XX` 和候选数组的相同位置推断它们是同一篇文章。
|
||||
|
||||
### extracted 文件与候选编号错位(实测 2026-07-31)
|
||||
|
||||
- `digest-brief.json` 的 `top_candidates` **没有 `item_id` 字段**,只有 `url`;
|
||||
- `extracted/item-XX.extracted.json` 的文件名序号与候选编号**可能不一致**(实例:候选2 = item-05、候选3 = item-02);
|
||||
- 正确做法:用 **URL 交叉匹配**(归一化 `%3D`→`=` 后逐条比对),或用完整 `item_id`(从 candidate-batch.json 的 `items[i].item_id` 按候选数组顺序取)在 extracted 文件里反查;两者都能验证时优先 item_id。
|
||||
|
||||
## 3. Hugo 日报
|
||||
|
||||
用户确认发布文章后:
|
||||
@@ -99,6 +105,27 @@ http://127.0.0.1:14322/daily/YYYY-MM-DD/
|
||||
5. `selected_ids` 传完整、非空、去除 `cand:` 前缀后的 `item_id`;
|
||||
6. 从 Job 结果的 `written_paths` 读取 Markdown。
|
||||
|
||||
### ⚠️ extracted_path 必须用绝对路径
|
||||
|
||||
`summary-mcp` 进程的工作目录是 `/root/.hermes`(不是 reader 项目根)。传相对路径(如 `outputs/freshrss/...`)会直接报 `extracted_path does not exist`。必须传绝对路径:
|
||||
|
||||
```text
|
||||
/home/ubuntu/zhu/github/reader/outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json
|
||||
```
|
||||
|
||||
### ⚠️ 候选编号 ≠ extracted 文件名顺序
|
||||
|
||||
候选数组顺序与 extracted 文件名(`item-01`…`item-07`)**不一定对齐**(实测候选2 落在 item-05)。`digest-brief.json` 的候选**没有 `item_id` 字段**,只有 URL。可靠匹配方法:
|
||||
|
||||
1. 从 `candidate-batch.json` 取每项完整 `item_id`(在 `candidate` 嵌套对象里,顶层 `item_key` 只是 `item-XX` 文件名序号);
|
||||
2. 或按 URL 匹配:归一化(`%3D`→`=`)后与每个 extracted 文件的 `article.url` / `article.canonical_url` 比对;
|
||||
3. 绝不要按候选位置对应 extracted 文件序号。
|
||||
|
||||
```python
|
||||
def norm(u): return u.replace('%3D','=').replace('%3d','=').strip()
|
||||
# 对每个 extracted 文件取 norm(article.url),与候选 norm(url) 精确比对
|
||||
```
|
||||
|
||||
正式序列:
|
||||
|
||||
```text
|
||||
|
||||
Reference in New Issue
Block a user