178 lines
6.0 KiB
Markdown
178 lines
6.0 KiB
Markdown
# Reader Digest Flow 操作参考
|
||
|
||
本文件承载 `reader-digest-flow` 的具体执行步骤。正式行为边界以 `../SKILL.md` 为准。
|
||
|
||
## 1. 日报 Pipeline
|
||
|
||
### 默认参数
|
||
|
||
```json
|
||
{
|
||
"limit": 7,
|
||
"include_read": false,
|
||
"mark_read": true,
|
||
"debug_artifacts": false,
|
||
"timeout_seconds": 60,
|
||
"max_retries": 2
|
||
}
|
||
```
|
||
|
||
### 正式调用序列
|
||
|
||
```text
|
||
start_freshrss_pipeline_job
|
||
→ get_freshrss_pipeline_job_status
|
||
→ get_freshrss_pipeline_job_result
|
||
→ get_run_status
|
||
→ get_delivery_payload / get_run_report
|
||
```
|
||
|
||
状态动作:
|
||
|
||
- `running`:按合理间隔继续轮询;
|
||
- `success`:读取结果,保存 `run_id`;
|
||
- `failed`:读取关联 Run 并执行 Resume Plan;
|
||
- `partial`:优先检查 Run Report、Artifact 和 Recovery 信息。
|
||
|
||
恢复序列:
|
||
|
||
```text
|
||
inspect_resume_plan
|
||
→ recommended_action=resume
|
||
→ start_resume_job
|
||
→ get_resume_job_status
|
||
→ get_resume_job_result
|
||
```
|
||
|
||
如果建议为 `read_terminal_result`,直接读取结果;如果为 `start_new_run`,停止并等待用户决定。
|
||
|
||
不要根据 `run_id` 拼目录。后续只消费调用返回的 `output_dir`、`artifact.path`、`delivery_output`、`report_output`。
|
||
|
||
## 2. 候选汇报与选择
|
||
|
||
优先使用当前 Run 的 digest brief Artifact;不可用时使用 `get_delivery_payload` 返回的候选。
|
||
|
||
展示规则:
|
||
|
||
1. 使用候选数组原始顺序并从 1 编号;
|
||
2. 不因 `keep/review` 分组而重新编号;
|
||
3. 每篇展示标题、来源、摘要和判断理由;
|
||
4. 用户编号只映射当前候选数组;
|
||
5. 跨 Artifact 读取详情时按 URL 或完整 `item_id` 匹配。
|
||
|
||
不要按 `summary-batch`、`item-XX` 和候选数组的相同位置推断它们是同一篇文章。
|
||
|
||
### extracted 文件与候选编号错位(实测 2026-07-31)
|
||
|
||
- `digest-brief.json` 的 `top_candidates` **没有 `item_id` 字段**,只有 `url`;
|
||
- `extracted/item-XX.extracted.json` 的文件名序号与候选编号**可能不一致**(实例:候选2 = item-05、候选3 = item-02);
|
||
- 正确做法:用 **URL 交叉匹配**(归一化 `%3D`→`=` 后逐条比对),或用完整 `item_id`(从 candidate-batch.json 的 `items[i].item_id` 按候选数组顺序取)在 extracted 文件里反查;两者都能验证时优先 item_id。
|
||
|
||
## 3. Hugo 日报
|
||
|
||
用户确认发布文章后:
|
||
|
||
1. 读取 `public-digest-example.md`;
|
||
2. 仅使用用户选中的文章生成公开内容;
|
||
3. 写入 Hugo 当日页面;
|
||
4. 前台执行部署,避免把构建日志作为聊天通知;
|
||
5. 验证首页、列表页和详情页。
|
||
|
||
当前部署位置:
|
||
|
||
```text
|
||
/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md
|
||
```
|
||
|
||
验证至少覆盖:
|
||
|
||
```text
|
||
http://127.0.0.1:14322/
|
||
http://127.0.0.1:14322/daily/
|
||
http://127.0.0.1:14322/daily/YYYY-MM-DD/
|
||
```
|
||
|
||
页面不存在时检查源文件、构建结果、容器状态和页面日期。需要 Docker/Hugo 深度排障时再读取部署项目自己的文档,不把历史事故规则复制回主 Skill。
|
||
|
||
## 4. 单篇知识笔记
|
||
|
||
用户确认 IMA 文章后:
|
||
|
||
1. 从候选中取得 URL 和完整 `item_id`;
|
||
2. 从 Run Artifact 中找到匹配的 extracted 文件;
|
||
3. 交叉验证 `article.item_id` 或 URL;
|
||
4. 对每个单篇文件调用 `start_article_summary_job(extracted_path=<返回的 Artifact 路径>, selected_ids=[<完整 item_id>])`;
|
||
5. `selected_ids` 传完整、非空、去除 `cand:` 前缀后的 `item_id`;
|
||
6. 从 Job 结果的 `written_paths` 读取 Markdown。
|
||
|
||
### ⚠️ extracted_path 必须用绝对路径
|
||
|
||
`summary-mcp` 进程的工作目录是 `/root/.hermes`(不是 reader 项目根)。传相对路径(如 `outputs/freshrss/...`)会直接报 `extracted_path does not exist`。必须传绝对路径:
|
||
|
||
```text
|
||
/home/ubuntu/zhu/github/reader/outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json
|
||
```
|
||
|
||
### ⚠️ 候选编号 ≠ extracted 文件名顺序
|
||
|
||
候选数组顺序与 extracted 文件名(`item-01`…`item-07`)**不一定对齐**(实测候选2 落在 item-05)。`digest-brief.json` 的候选**没有 `item_id` 字段**,只有 URL。可靠匹配方法:
|
||
|
||
1. 从 `candidate-batch.json` 取每项完整 `item_id`(在 `candidate` 嵌套对象里,顶层 `item_key` 只是 `item-XX` 文件名序号);
|
||
2. 或按 URL 匹配:归一化(`%3D`→`=`)后与每个 extracted 文件的 `article.url` / `article.canonical_url` 比对;
|
||
3. 绝不要按候选位置对应 extracted 文件序号。
|
||
|
||
```python
|
||
def norm(u): return u.replace('%3D','=').replace('%3d','=').strip()
|
||
# 对每个 extracted 文件取 norm(article.url),与候选 norm(url) 精确比对
|
||
```
|
||
|
||
正式序列:
|
||
|
||
```text
|
||
start_article_summary_job
|
||
→ get_article_summary_job_status
|
||
→ get_article_summary_job_result
|
||
```
|
||
|
||
内容依据是 extracted Artifact 中的 `article.plain_text`。不得重新抓取原 URL,不得使用原文之外的知识扩写。
|
||
|
||
CLI 仅在异步 MCP 不可用或人工排障时使用:
|
||
|
||
```bash
|
||
python scripts/run_article_summaries.py \
|
||
--extracted <returned-extracted-path> \
|
||
--ids <full-item-id> \
|
||
--output-dir <explicit-output-dir>
|
||
```
|
||
|
||
## 5. IMA 上传
|
||
|
||
用户在知识沉淀阶段的文章选择即为上传授权。
|
||
|
||
执行顺序:
|
||
|
||
1. 检查生成的 Markdown 与来源;
|
||
2. 文件名规范化为 `<完整文章标题>.md`;
|
||
3. 确认目标为 `daily` knowledge base;
|
||
4. 执行 preflight、create_media、COS upload、add_knowledge;
|
||
5. 验证知识库条目存在。
|
||
|
||
上传格式与 API 参数分别见:
|
||
|
||
- `ima-format-quickref.md`
|
||
- `ima-upload-api.md`
|
||
- `ima-credential-chain.md`
|
||
|
||
失败时记录发生在 preflight、create_media、COS upload 或 add_knowledge 的具体阶段。不要把失败的 Markdown 改写成 URL 导入或 Notes 类型来绕过错误。
|
||
|
||
## 6. CLI fallback 原则
|
||
|
||
CLI 仅在以下场景使用:
|
||
|
||
- MCP 服务不可用;
|
||
- Tool transport/launch 失败且无法取得有效 Job;
|
||
- 用户明确要求本地调试;
|
||
- 人工排障需要直接检查脚本输出。
|
||
|
||
如果已取得 `job_id` 或 `run_id`,先查询真实状态,避免因响应超时重复启动任务。
|