Compare commits
24
Commits
2df0af5b82
...
main
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
8e27cca166 | ||
|
|
807027976e | ||
|
|
6dd8cef347 | ||
|
|
5eb390e3ed | ||
|
|
7b791ac947 | ||
|
|
cdbcdcd485 | ||
|
|
590d050218 | ||
|
|
4399c9ca90 | ||
|
|
8136301ad4 | ||
|
|
b29cd8f934 | ||
|
|
4c4d1a45e6 | ||
|
|
b8727f1885 | ||
|
|
6705613aa4 | ||
|
|
416414ae1d | ||
|
|
f3e7488fc8 | ||
|
|
e3663f681d | ||
|
|
a06f2a1d08 | ||
|
|
c0194647d8 | ||
|
|
c528e0abc7 | ||
|
|
52ce6bfdf5 | ||
|
|
c58d8114cd | ||
|
|
10f9cd088f | ||
|
|
2053cccfef | ||
|
|
87d18e4263 |
@@ -1,340 +1,214 @@
|
||||
# Reader MCP Workflow Service
|
||||
# Reader · AI 日报引擎
|
||||
|
||||
reader 当前已经收口为面向 OpenClaw 的 MCP workflow service。正式能力边界以 FreshRSS 日报工作流为准:启动 run、写入 `run-state.json`、查询运行状态、读取结构化结果,以及最小可用的 `resume_run`。CLI 仍保留,但定位为 debug / fallback,而不是正式集成入口。
|
||||
> 从 FreshRSS 到 AI 日报的自动化流水线,为个人知识管理生成每日 AI 工程化简报。
|
||||
|
||||
## 运行
|
||||
Reader 是一个端到端的 AI 日报生产系统,定时从自建 FreshRSS 的 RSS 订阅源拉取文章,经过内容提取、LLM 筛选与摘要、关键词索引构建,最终产出两个输出:
|
||||
|
||||
1. **公开日报** — 推送到 [Hugo 站点](https://osiman.site/daily/) 的精选技术简报
|
||||
2. **知识沉淀** — 单篇结构化摘要上传到 IMA 知识库(`daily` KB)
|
||||
|
||||
整个流程由 OpenClaw 编排,作为 MCP Workflow Service 对外暴露。
|
||||
|
||||
---
|
||||
|
||||
## ✨ 核心能力
|
||||
|
||||
| 能力 | 说明 |
|
||||
|:----|:------|
|
||||
| **RSS 拉取** | 从 FreshRSS API 拉取订阅文章,支持增量读取与已读标记 |
|
||||
| **内容提取** | 自动提取文章正文、标题、来源等结构化字段 |
|
||||
| **LLM 筛选** | 基于个人兴趣画像(`filter_context.personal.json`)自动评估文章质量,分为 keep / review / drop 三档 |
|
||||
| **LLM 摘要** | 并行生成每篇文章的结构化摘要(4 路并发,约 24 秒完成 7 篇) |
|
||||
| **关键词索引** | 自动构建每日关键词索引,支持别名映射与停用词过滤 |
|
||||
| **候选简报** | 生成 `digest-brief.json` 供编排层(OpenClaw)决策 |
|
||||
| **单篇沉淀** | 对选中的文章生成结构化知识笔记,上传到 IMA 知识库 |
|
||||
| **异步 Job** | 全部生产流程走异步 job,支持恢复与状态查询 |
|
||||
|
||||
---
|
||||
|
||||
## 🏗 架构概览
|
||||
|
||||
```
|
||||
FreshRSS ──→ 拉取 ──→ 内容提取 ──→ LLM 筛选 ──→ 关键词索引
|
||||
│
|
||||
digest-brief.json
|
||||
│
|
||||
┌──────────────┼──────────────┐
|
||||
▼ ▼ ▼
|
||||
Hugo 日报 IMA 知识库 term_index
|
||||
(公开简报) (单篇沉淀) (关键词数据)
|
||||
```
|
||||
|
||||
### MCP 工具层
|
||||
|
||||
Reader 通过 Hermes MCP 暴露 20+ 个工具,分为三类:
|
||||
|
||||
**日报流水线:**
|
||||
- `start_freshrss_pipeline_job` → 启动异步日报 Job
|
||||
- `get_freshrss_pipeline_job_status` / `get_freshrss_pipeline_job_result` → 轮询结果
|
||||
|
||||
**状态查询:**
|
||||
- `get_run_status` / `get_delivery_payload` / `get_run_report` → 读取运行结果
|
||||
- `list_runs` / `list_run_artifacts` → 浏览运行历史
|
||||
|
||||
**恢复与单篇总结:**
|
||||
- `inspect_resume_plan` / `start_resume_job` → 恢复失败 Job
|
||||
- `start_article_summary_job` / `generate_article_summaries` → 单篇文章摘要
|
||||
|
||||
### CLI 入口
|
||||
|
||||
同步入口,适合本地 debug / fallback:
|
||||
|
||||
```bash
|
||||
# 完整日报流水线
|
||||
python scripts/run_freshrss_pipeline.py --limit 7 --mark-read --timeout 300
|
||||
|
||||
# 单篇文章摘要
|
||||
python scripts/run_article_summaries.py \
|
||||
--extracted outputs/freshrss/rerun/<run_id>/extracted/item-01.extracted.json \
|
||||
--output-dir outputs/freshrss/single_summaries/YYYY-MM-DD
|
||||
|
||||
# 关键词维护
|
||||
python scripts/build_keyword_index.py
|
||||
python scripts/generate_term_cleanup_suggestions.py
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🚀 快速开始
|
||||
|
||||
### 环境变量
|
||||
|
||||
```
|
||||
# FreshRSS
|
||||
FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||||
FRESHRSS_USERNAME=bot
|
||||
FRESHRSS_API_PASSWORD=xxx
|
||||
|
||||
# LLM(主流水线)
|
||||
LLM_API_URL=https://api.deepseek.com
|
||||
LLM_API_KEY=xxx
|
||||
LLM_MODEL=deepseek-chat
|
||||
|
||||
# LLM(可选,单篇摘要独立模型)
|
||||
ARTICLE_SUMMARY_LLM_API_URL=https://api.deepseek.com
|
||||
ARTICLE_SUMMARY_LLM_API_KEY=xxx
|
||||
ARTICLE_SUMMARY_LLM_MODEL=deepseek-chat
|
||||
|
||||
# IMA 知识库(可选,仅沉淀时需要)
|
||||
IMA_DAILY_KNOWLEDGE_BASE_ID=xxx
|
||||
IMA_DAILY_KNOWLEDGE_BASE_NAME=daily
|
||||
```
|
||||
|
||||
### 运行
|
||||
|
||||
```bash
|
||||
# 安装
|
||||
pip install -e .
|
||||
|
||||
# 跑日报流水线(CLI 模式)
|
||||
python scripts/run_freshrss_pipeline.py --limit 7 --mark-read --timeout 300
|
||||
|
||||
# 启动 MCP 服务(OpenClaw 集成用)
|
||||
summary-mcp
|
||||
```
|
||||
|
||||
服务当前暴露 11 个工具。
|
||||
---
|
||||
|
||||
正式 workflow service 相关工具:
|
||||
## 📁 项目结构
|
||||
|
||||
- `run_freshrss_openclaw_pipeline`
|
||||
- `get_run_status`
|
||||
- `list_runs`
|
||||
- `list_run_artifacts`
|
||||
- `get_delivery_payload`
|
||||
- `get_run_report`
|
||||
- `resume_run`
|
||||
|
||||
单步处理 / 调试相关工具:
|
||||
|
||||
- `extract_url_content`
|
||||
- `extract_item_content`
|
||||
- `filter_summary_result`
|
||||
- `generate_article_summaries`
|
||||
|
||||
## 正式能力边界
|
||||
|
||||
- 当前正式 workflow 只有 `freshrss_daily_digest`
|
||||
- 每次 FreshRSS 主流水线 run 都会在 `outputs/freshrss/rerun/<run_dir>/run-state.json` 落地运行真相
|
||||
- OpenClaw 正式读取结果应优先使用 `get_delivery_payload` 与 `get_run_report`,而不是自己拼输出目录路径
|
||||
- `digest-brief.json` 当前会随主流水线产出,但还没有独立的 MCP 读取工具;如需定位它,应通过 `list_run_artifacts` 或 `get_run_report` 返回的信息发现
|
||||
- `run_freshrss_openclaw_pipeline` 与 `resume_run` 当前都是同步 MCP 调用;仓库里还没有后台队列 / worker / 异步任务管理
|
||||
|
||||
## OpenClaw 推荐调用路径
|
||||
|
||||
1. 调用 `run_freshrss_openclaw_pipeline` 启动正式日报 run,并保存返回的 `run_id`
|
||||
2. 后续所有状态判断都基于 `get_run_status(run_id)` 或 `list_runs(...)`
|
||||
3. 需要看产物列表时用 `list_run_artifacts(run_id)`,不要在 OpenClaw 里硬编码 `outputs/freshrss/rerun/...`
|
||||
4. 需要消费正式结果时优先用 `get_delivery_payload(run_id)` 与 `get_run_report(run_id)`
|
||||
5. 仅当 `resume_run` 的最小恢复范围满足时,才对失败 run 调用 `resume_run(run_id)`;否则应重启一个新 run
|
||||
|
||||
## `resume_run` 当前最小范围
|
||||
|
||||
- 只支持带有效 `run-state.json` 的 run
|
||||
- 只支持 workflow `freshrss_daily_digest`
|
||||
- 恢复时继续沿用原 `run_id`,不会新建 retry run
|
||||
- 当前支持的恢复起点只有:`generate_summaries`、`apply_filters`、`build_delivery_payload`、`write_run_report`
|
||||
- 当前明确不支持从 `fetch_feed`、`extract_articles` 恢复;这类失败应新开 run
|
||||
- 恢复前会校验关键中间产物是否齐备,缺失时直接返回不可恢复,而不会自动回退到更早 stage
|
||||
|
||||
## 单篇文章总结后处理(可选使用独立 LLM)
|
||||
|
||||
### 生产环境推荐输入
|
||||
|
||||
- 单篇总结的正式生产输入,优先使用 FreshRSS 主流水线输出的**单篇 extracted 文件**:
|
||||
- `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json`
|
||||
- 这些**逐条 extracted 文件**是下游单篇总结的**正式默认产物**。
|
||||
- 像 `outputs/freshrss/extracted/freshrss.extracted.json` 这样的**批量 extracted 文件**,只作为临时场景、兼容旧流程的输入形态保留,**不是首选生产默认**。
|
||||
|
||||
### daily 知识库默认配置
|
||||
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_ID` —— 单篇日报总结默认上传的 IMA 知识库 ID
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_NAME` —— 默认知识库名称(预期值:`daily`)
|
||||
- 上传逻辑在运行时应先校验目标知识库;若配置的目标不存在,应先按名称查找,仍不存在则创建 `daily`
|
||||
|
||||
相关能力:
|
||||
|
||||
- `generate_article_summaries` MCP 工具(基于已有 extracted payload 做单篇总结)
|
||||
- `scripts/run_article_summaries.py` CLI 辅助脚本
|
||||
|
||||
## 校验 LLM 摘要结果
|
||||
|
||||
```bash
|
||||
validate-llm-result outputs/reference/summary/result.json --extracted outputs/reference/extracted/read-flow-2026.extracted.json
|
||||
```
|
||||
reader/
|
||||
├── configs/ # 配置
|
||||
│ ├── filter_context.personal.json # 个人兴趣画像
|
||||
│ ├── term_aliases.json # 关键词别名映射(149 条)
|
||||
│ ├── term_stopwords.json # 关键词停用词(132 条)
|
||||
│ └── term_cleanup_policy.json # 关键词清理策略
|
||||
├── src/
|
||||
│ └── summary_mcp/ # MCP 服务核心
|
||||
│ ├── server.py # MCP 服务入口
|
||||
│ ├── runtime/ # 运行时(Job 管理、状态持久化)
|
||||
│ └── workflows/ # 工作流(日报流水线逻辑)
|
||||
├── scripts/ # CLI 入口
|
||||
├── outputs/ # 运行时产出
|
||||
│ └── freshrss/
|
||||
│ ├── rerun/<run_id>/ # 每次运行的全量产物
|
||||
│ │ ├── candidates/ # digest-brief.json, delivery payload
|
||||
│ │ ├── extracted/ # item-XX.extracted.json
|
||||
│ │ └── run-state.json # 运行状态
|
||||
│ └── single_summaries/ # 单篇摘要输出
|
||||
├── data/
|
||||
│ └── term_index/ # 关键词索引数据
|
||||
│ ├── daily/YYYY-MM-DD.json
|
||||
│ └── term_stats.json
|
||||
├── docs/ # 设计文档
|
||||
└── prompts/ # LLM Prompt 模板
|
||||
```
|
||||
|
||||
## 跑最小 extraction → summary 循环
|
||||
---
|
||||
|
||||
## ⚙️ 关键技术决策
|
||||
|
||||
| 决策 | 选择 | 原因 |
|
||||
|:----|:----|:------|
|
||||
| 运行模式 | **异步 Job** 为主,CLI fallback | 避免 MCP 传输层 120s 超时限制 |
|
||||
| 摘要并发 | **ThreadPoolExecutor(max_workers=4)** | LLM 调用是 I/O 密集型,4 路并行将 7 篇摘要从 2-3 分钟压到 ~24 秒 |
|
||||
| 环境变量 | **子进程显式注入 .env** | 解决 MCP 服务器环境隔离导致子进程读取不到 LLM_API_KEY 的问题 |
|
||||
| 关键词过滤 | **别名映射 + 停用词 + 语义清洗** | 先用 `term_aliases.json` 归一化,再用 `term_stopwords.json` 过滤噪声,最后通过 LLM 做语义级清洗 |
|
||||
| Tag 选择 | **复用已有通用 Tag**,不从 term_index 翻生僻词 | 保持 Hugo 站点 /tags/ 页面整洁,避免大量一次性专有名词 |
|
||||
|
||||
---
|
||||
|
||||
## 🔧 关键词治理
|
||||
|
||||
配置治理走四步流程(`scripts/` 下脚本):
|
||||
|
||||
```bash
|
||||
python scripts/run_summary_loop.py ^
|
||||
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||
--prompt outputs/prompts/llm-summary-prompt.txt ^
|
||||
--output outputs/reference/summary/result.loop.json
|
||||
# 1. 构建评审数据包
|
||||
python skills/keyword-cleanup-review/scripts/build_review_bundle.py --days 7 --top 50
|
||||
|
||||
# 2. 统计规则级建议(大小写、单复数、频次阈值)
|
||||
python scripts/generate_term_cleanup_suggestions.py
|
||||
|
||||
# 3. LLM 语义级建议(中英映射、简称-全称、近义词)
|
||||
python scripts/generate_term_cleanup_semantic_suggestions.py
|
||||
|
||||
# 4. 确认后写入配置
|
||||
python scripts/apply_term_suggestions.py --accept-watch ... --dry-run
|
||||
```
|
||||
|
||||
## 拉取 FreshRSS 条目并映射为标准化 `item`
|
||||
详见 `docs/design/keyword-engine-maintenance.md`。
|
||||
|
||||
```bash
|
||||
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||||
set FRESHRSS_USERNAME=bot
|
||||
set FRESHRSS_API_PASSWORD=your-api-password
|
||||
python scripts/pull_freshrss_items.py --limit 5 --mark-read
|
||||
```
|
||||
---
|
||||
|
||||
默认会排除已经带 `read` 标签的条目。
|
||||
如果你想拿到完整阅读列表,可以加 `--include-read`。
|
||||
启用 `--mark-read` 后,脚本会在执行成功后把本次抓到的条目标记为已读。
|
||||
## 🤖 Agent Skill
|
||||
|
||||
脚本会写出:
|
||||
Reader 附带一个完整的 OpenClaw Agent Skill,位于 `skills/reader-digest-flow/`,供 AI Agent(Hermes / Claude Code 等)编排每日日报流程使用。
|
||||
|
||||
- `outputs/freshrss/raw/freshrss.raw.json`
|
||||
- `outputs/freshrss/items/freshrss.items.json`
|
||||
Skill 包含完整的 7 阶段工作流定义:
|
||||
1. **Phase 1** — 跑 Pipeline(FreshRSS → 提取 → LLM 筛选)
|
||||
2. **Phase 2** — 汇报候选(展示候选文章给用户决策)
|
||||
3. **Phase 3** — 用户选文(选择 Hugo 发布文章)
|
||||
4. **Phase 4** — 生成并发布 Hugo 日报
|
||||
5. **Phase 5** — 用户选 IMA 沉淀文章
|
||||
6. **Phase 6** — LLM 摘要生成
|
||||
7. **Phase 7** — IMA 知识库上传
|
||||
|
||||
## 拉取 FreshRSS 条目并逐条做内容提取
|
||||
以及海量铁律(不重跑 pipeline、编号规则、Tag 选择规范、IMA 上传流程等)和参考文件(`references/` 目录)。
|
||||
|
||||
```bash
|
||||
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||||
set FRESHRSS_USERNAME=osiman
|
||||
set FRESHRSS_API_PASSWORD=your-api-password
|
||||
python scripts/run_freshrss_extract.py --limit 1 --mark-read
|
||||
```
|
||||
---
|
||||
|
||||
默认会排除已经带 `read` 标签的条目。
|
||||
启用 `--mark-read` 后,只有提取成功的条目才会被标记为已读。
|
||||
## 📄 文档
|
||||
|
||||
脚本会写出:
|
||||
- `docs/openclaw/README.md` — OpenClaw 集成指南
|
||||
- `docs/openclaw/openclaw-orchestration-flow.md` — 编排流程
|
||||
- `docs/openclaw/openclaw-delivery-payload-spec.md` — 字段契约
|
||||
- `docs/design/README.md` — 设计文档总索引
|
||||
- `docs/design/filter-rule-engine-design.md` — 过滤规则引擎设计
|
||||
- `docs/design/filter-rule-engine-usage.md` — 过滤规则使用说明
|
||||
|
||||
- `outputs/freshrss/raw/freshrss.raw.json`
|
||||
- `outputs/freshrss/items/freshrss.items.json`
|
||||
- `outputs/freshrss/extracted/freshrss.extracted.json`
|
||||
---
|
||||
|
||||
注意:这个**批量 extracted 文件**主要用于独立提取场景和旧流程兼容。下游单篇总结的正式生产默认输入,仍然是 `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json` 这类**逐条 extracted 文件**。
|
||||
## 📝 License
|
||||
|
||||
## 跑完整 FreshRSS 流水线,并在最终 delivery payload 写盘成功后再标记已读
|
||||
|
||||
```bash
|
||||
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||||
set FRESHRSS_USERNAME=osiman
|
||||
set FRESHRSS_API_PASSWORD=your-api-password
|
||||
set LLM_API_URL=https://api.deepseek.com
|
||||
set LLM_API_KEY=your-llm-api-key
|
||||
set LLM_MODEL=deepseek-chat
|
||||
python scripts/run_freshrss_pipeline.py --limit 5 --mark-read
|
||||
```
|
||||
|
||||
如果你希望过滤时引入个人工程兴趣 / AI Agent 兴趣画像,可以传入 context 文件:
|
||||
|
||||
```bash
|
||||
python scripts/run_freshrss_pipeline.py ^
|
||||
--limit 5 ^
|
||||
--context configs/filter_context.personal.json ^
|
||||
--mark-read
|
||||
```
|
||||
|
||||
这条 CLI 与 MCP `run_freshrss_openclaw_pipeline` 共用同一条主流水线逻辑,但正式生产集成应优先走 MCP;CLI 仅用于本地 debug / fallback。默认会写出这些产物:
|
||||
|
||||
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/raw/freshrss.raw.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/candidates/digest-brief.json`(给 OpenClaw 生成 public digest 用的轻量输入,仅包含 `keep` 候选)
|
||||
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json`(每篇一份)
|
||||
|
||||
同时还会更新每日关键词索引运行数据:
|
||||
|
||||
- `data/term_index/daily/YYYY-MM-DD.json`
|
||||
- `data/term_index/term_stats.json`
|
||||
|
||||
主流水线默认**不会**产出批量级的 `freshrss.extracted.json`。
|
||||
如果你需要更多逐条中间产物,例如标准化 items、摘要结果、过滤决策、candidate record、candidate input,可以加 `--debug-artifacts`。
|
||||
|
||||
当 OpenClaw 接入这个 MCP 服务后,应以 `run_id` 作为稳定句柄:先调用 `run_freshrss_openclaw_pipeline`,再通过 `get_run_status` / `get_delivery_payload` / `get_run_report` 读取状态与结果,而不是直接拼接目录路径。
|
||||
在排查复杂问题时,也可以把 `debug_artifacts=true` 打开,并结合 `list_run_artifacts` 查看该 run 下实际产物。
|
||||
|
||||
## 对结构化摘要结果执行确定性过滤规则
|
||||
|
||||
```bash
|
||||
python scripts/run_filter_rules.py ^
|
||||
--summary outputs/reference/summary/result.loop.json ^
|
||||
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||
--output outputs/reference/filter/filter-decision.json
|
||||
```
|
||||
|
||||
如果你希望注入兴趣主题或来源标签,也可以额外传入 context 文件:
|
||||
|
||||
```bash
|
||||
python scripts/run_filter_rules.py ^
|
||||
--summary outputs/reference/summary/result.loop.json ^
|
||||
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||
--context outputs/reference/filter/filter-context.json ^
|
||||
--output outputs/reference/filter/filter-decision.with-context.json
|
||||
```
|
||||
|
||||
规则引擎设计和规则编写说明见:
|
||||
|
||||
- `docs/design/filter-rule-engine-design.md`
|
||||
- `docs/design/filter-rule-engine-usage.md`
|
||||
|
||||
## 将过滤结果写入 Markdown sink
|
||||
|
||||
```bash
|
||||
python scripts/run_markdown_sink.py ^
|
||||
--summary outputs/reference/summary/result.loop.json ^
|
||||
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||
--filter outputs/reference/filter/filter-decision.json
|
||||
```
|
||||
|
||||
脚本会把 Markdown 笔记写到 `knowledge-base/` 下。
|
||||
|
||||
## 构建内部 `ArticleCandidateRecord` 与精简版 `OpenClawCandidateInput`
|
||||
|
||||
```bash
|
||||
python scripts/run_article_candidate.py ^
|
||||
--summary outputs/reference/summary/result.loop.json ^
|
||||
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||
--filter outputs/reference/filter/filter-decision.json ^
|
||||
--section-hint tools_and_workflows
|
||||
```
|
||||
|
||||
默认会写出:
|
||||
|
||||
- `outputs/reference/candidates/article-candidate-record.json`
|
||||
- `outputs/reference/candidates/openclaw-candidate-input.json`
|
||||
|
||||
## 构建批量 OpenClaw delivery payload
|
||||
|
||||
```bash
|
||||
python scripts/build_openclaw_delivery.py ^
|
||||
--input-dir outputs/freshrss/candidates/batch ^
|
||||
--sort-by-rank ^
|
||||
--date 2026-03-25
|
||||
```
|
||||
|
||||
默认会写出:
|
||||
|
||||
- `outputs/reference/candidates/openclaw-delivery-payload.json`
|
||||
|
||||
输出目录布局说明见 `outputs/README.md`。
|
||||
|
||||
## 关键词索引默认配置
|
||||
|
||||
相关配置文件位于:
|
||||
|
||||
- `configs/term_aliases.json`
|
||||
- `configs/term_stopwords.json`
|
||||
- `configs/term_cleanup_policy.json`
|
||||
- `configs/term_watchlist.json`
|
||||
- `configs/term_change_log.json`
|
||||
|
||||
你也可以基于已有 delivery payload 重新构建关键词索引:
|
||||
|
||||
```bash
|
||||
python scripts/build_keyword_index.py ^
|
||||
--input outputs/reference/candidates/openclaw-delivery-payload.json
|
||||
```
|
||||
|
||||
运行期关键词数据存放在:
|
||||
|
||||
- `data/term_index/`
|
||||
|
||||
关键词清理评审 skill 位于:
|
||||
|
||||
- `skills/keyword-cleanup-review/`
|
||||
|
||||
构建给 LLM skill 使用的评审数据包(review bundle):
|
||||
|
||||
```bash
|
||||
python skills/keyword-cleanup-review/scripts/build_review_bundle.py ^
|
||||
--days 7 ^
|
||||
--top 50 ^
|
||||
--output outputs/term_index/review/keyword-cleanup-bundle.json
|
||||
```
|
||||
|
||||
现在这个评审数据包(review bundle)还会额外携带治理上下文:
|
||||
|
||||
- 来自 `configs/term_cleanup_policy.json` 的清理阈值
|
||||
- 当前 watch list(`configs/term_watchlist.json`)
|
||||
- 最近已应用的变更(`configs/term_change_log.json`)
|
||||
|
||||
这个 skill 只负责生成 review 输入与建议,不会自动修改:
|
||||
|
||||
- `term_aliases`
|
||||
- `term_stopwords`
|
||||
- `filter_context.personal.json`
|
||||
|
||||
如果你想先预览已接受建议,再决定是否写配置文件:
|
||||
|
||||
```bash
|
||||
python scripts/apply_term_suggestions.py ^
|
||||
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json ^
|
||||
--accept-watch Cron Heartbeat Memory ^
|
||||
--dry-run
|
||||
```
|
||||
|
||||
去掉 `--dry-run` 后才会真正写文件。
|
||||
这个脚本也支持通过 `--accept-alias`、`--accept-stopword`、`--accept-interest` 应用 alias / stopword / interest keyword 变更。
|
||||
已接受的 watch 词会写入 `configs/term_watchlist.json`,每次应用动作也会被追加到 `configs/term_change_log.json`。
|
||||
|
||||
## 单篇总结 LLM 配置
|
||||
|
||||
如果你希望单篇总结后处理使用独立模型,而不影响主流水线,可以设置:
|
||||
|
||||
- `ARTICLE_SUMMARY_LLM_API_URL`
|
||||
- `ARTICLE_SUMMARY_LLM_MODEL`
|
||||
- `ARTICLE_SUMMARY_LLM_API_KEY`
|
||||
|
||||
如果这些变量未设置,单篇总结会回退使用主流程中的 `LLM_*` / `OPENAI_*` 配置。
|
||||
|
||||
示例(PowerShell 风格):
|
||||
|
||||
```bash
|
||||
set ARTICLE_SUMMARY_LLM_API_URL=https://api.deepseek.com
|
||||
set ARTICLE_SUMMARY_LLM_API_KEY=your-article-summary-key
|
||||
set ARTICLE_SUMMARY_LLM_MODEL=deepseek-chat
|
||||
```
|
||||
|
||||
然后可以这样调用 CLI。
|
||||
正式生产环境建议优先使用 `outputs/freshrss/rerun/<run_id>/extracted/` 下的**逐条 extracted 文件**;下面这个**批量 extracted** 示例仅保留为兼容旧流程 / 临时场景输入:
|
||||
|
||||
```bash
|
||||
python scripts/run_article_summaries.py ^
|
||||
--extracted outputs/freshrss/extracted/freshrss.extracted.json ^
|
||||
--ids 12345 67890 ^
|
||||
--output-dir outputs/freshrss/single_summaries
|
||||
```
|
||||
|
||||
也可以通过 `summary_mcp.server` 暴露的 MCP 工具 `generate_article_summaries` 调用:
|
||||
|
||||
- `extracted_path`(string):单篇 extracted JSON 路径(例如 `outputs/freshrss/rerun/<run_id>/extracted/item-01.extracted.json`),或者包含 `results` 数组的 batch extracted JSON
|
||||
- `selected_ids`(array of strings):要总结的一个或多个 `item_id`。如果传空数组,则对文件中的全部条目做总结
|
||||
- `output_dir`(optional string):Markdown 输出目录;若不传,则默认写到 extracted 文件旁边的 `single_summaries/` 目录
|
||||
- `llm_api_key` / `llm_model` / `llm_api_url`(optional strings):单篇总结 LLM 的覆盖配置;不传时会按前文规则回退到 `ARTICLE_SUMMARY_*` 或主 `LLM_*`
|
||||
|
||||
该工具返回一个 JSON 数组,内容为生成好的 Markdown 文件路径。
|
||||
|
||||
单篇总结使用独立 prompt:`outputs/prompts/article-summary-prompt.txt`。
|
||||
它与日报 prompt 完全独立,输出的是中文结构化知识笔记,包含这些部分:
|
||||
|
||||
- 核心结论
|
||||
- 主要论点
|
||||
- 关键方法 / 机制
|
||||
- 重要细节
|
||||
- 可复用启发
|
||||
- 关键词
|
||||
- 主题
|
||||
MIT
|
||||
|
||||
@@ -12,11 +12,16 @@
|
||||
|
||||
开始编码前必须阅读:
|
||||
|
||||
1. `plans/reader-mcp-architecture-design.md`
|
||||
2. `plans/reader-mcp-implementation-plan.md`
|
||||
3. 本文件
|
||||
4. `plans/issues/2026-04-06-reader-digest-sigterm.md`
|
||||
5. `docs/openclaw/openclaw-handoff.md`
|
||||
1. `README.md`
|
||||
2. `docs/README.md`
|
||||
3. `docs/openclaw/README.md`
|
||||
4. `docs/openclaw/openclaw-handoff.md`
|
||||
5. `docs/openclaw/openclaw-orchestration-flow.md`
|
||||
6. `plans/README.md`
|
||||
7. `plans/reader-mcp-architecture-design.md`
|
||||
8. `plans/reader-mcp-implementation-plan.md`
|
||||
9. 本文件
|
||||
10. `plans/issues/2026-04-06-reader-digest-sigterm.md`
|
||||
|
||||
---
|
||||
|
||||
@@ -184,6 +189,119 @@
|
||||
- 2026-04-07:已完成 `resume_run` minimal design 与现有 runtime/workflow/server 代码对齐分析,开始实现最小恢复链路。
|
||||
- 2026-04-07:已完成 `resume_run` 最小实现编码,新增 runtime 恢复服务并接入 MCP server;当前进入设计对齐与本地自检。
|
||||
- 2026-04-07:已完成 `resume_run` 架构对齐与本地自检;已验证 `write_run_report` 可恢复,且 `extract_articles` 会被明确拒绝恢复。
|
||||
- 2026-04-14:已补 `inspect_resume_plan`、artifact-first 恢复判定,以及生产模式下稳定 `summary-batch` / `candidate-batch` artifacts;当前 `resume` 的剩余主问题不再是恢复点判断,而是同步执行模型仍可能让 OpenClaw 恢复阶段超时。
|
||||
|
||||
---
|
||||
|
||||
### [DONE][P1] 把 `resume_run` 升级为最小真异步 job
|
||||
|
||||
目标:
|
||||
- 解决 `resume_run` 在 OpenClaw → MCP 同步链路里仍可能超时的问题
|
||||
- 让恢复也具备“启动 / 轮询 / 读取结果”的正式控制面
|
||||
|
||||
要求:
|
||||
- 新增最小异步接口:
|
||||
- `start_resume_job`
|
||||
- `get_resume_job_status`
|
||||
- `get_resume_job_result`
|
||||
- 状态目录固定落到:
|
||||
- `outputs/freshrss/resume_jobs/<job_id>/`
|
||||
- 至少包含:
|
||||
- `run-state.json`
|
||||
- `input.json`
|
||||
- `result.json`(成功时)
|
||||
- `job-report.json`
|
||||
- 启动前必须先走 `inspect_resume_plan`
|
||||
- 业务执行继续复用现有 `resume_service`,不要重写恢复主逻辑
|
||||
- `resume_run` 保留为同步 debug / fallback 路径,但不再作为 OpenClaw 的默认恢复入口
|
||||
|
||||
完成标准:
|
||||
- 可恢复 run 上,`start_resume_job` 能成功返回 `job_id`
|
||||
- `get_resume_job_status` 能稳定反映恢复 job 生命周期
|
||||
- `get_resume_job_result` 能稳定返回 `run_id`、`resume_from_stage`、最终状态与关键产物路径
|
||||
- 恢复耗时超过单次 MCP 同步窗口时,OpenClaw 仍不会因为同步调用挂住
|
||||
|
||||
进展备注:
|
||||
- 2026-04-14:已落地 `src/summary_mcp/runtime/resume_jobs.py` 与 `scripts/run_resume_job.py`,新增 `start_resume_job` / `get_resume_job_status` / `get_resume_job_result`
|
||||
- 2026-04-14:启动前会先走 `inspect_resume_plan`;不可恢复 run 会在 job 输入校验阶段直接失败,不进入后台恢复执行
|
||||
- 2026-04-14:后台执行复用现有 `_resume_freshrss_run(...)`,没有重写恢复主逻辑
|
||||
- 2026-04-14:已完成本地 synthetic 验证:`write_run_report` 恢复可通过 `start -> poll -> result` 闭环成功收敛
|
||||
|
||||
---
|
||||
|
||||
### [DONE][P1] 单篇总结改为最小真异步 job
|
||||
|
||||
目标:
|
||||
- 解决 `generate_article_summaries` 在 OpenClaw → MCP 同步链路里易 timeout 的问题
|
||||
- 将单篇总结正式升级为可启动、可轮询、可读取结果的异步 job
|
||||
|
||||
要求:
|
||||
- 新增最小异步接口:
|
||||
- `start_article_summary_job`
|
||||
- `get_article_summary_job_status`
|
||||
- `get_article_summary_job_result`
|
||||
- 状态目录固定落到:
|
||||
- `outputs/freshrss/article_summary_jobs/<job_id>/`
|
||||
- 至少包含:
|
||||
- `run-state.json`
|
||||
- `input.json`
|
||||
- `result.json`(成功时)
|
||||
- `job-report.json`
|
||||
- 执行模型优先使用后台子进程,不使用线程
|
||||
- 业务逻辑继续复用 `summarize_selected_articles(...)`,不要重写正文总结核心逻辑
|
||||
- 对 OpenClaw / reader-digest-flow 而言,异步 job 成功后应可继续接 IMA 沉淀闭环
|
||||
|
||||
当前进展:
|
||||
- 2026-04-10:已完成方案文档 `plans/article-summary-async-job-plan.md`
|
||||
- 2026-04-10:已落地最小代码骨架:
|
||||
- `src/summary_mcp/runtime/article_summary_jobs.py`
|
||||
- `scripts/run_article_summary_job.py`
|
||||
- `src/summary_mcp/server.py` 已新增 3 个 async job tools
|
||||
- 2026-04-10:已用真实 extracted 文件验证最小异步链路可跑通,job 能成功进入 `running → success`,并可读回结果
|
||||
- 2026-04-10:已补 README / OpenClaw handoff 文档,并对外统一为 `start_article_summary_job` / `get_article_summary_job_status` / `get_article_summary_job_result`
|
||||
- 2026-04-10:已完成聚焦自检:
|
||||
- MCP tool 注册名校验通过
|
||||
- stubbed async job 成功路径通过
|
||||
- stubbed async job 失败路径通过,`error_summary` 与 `job-report.json` 可回读
|
||||
|
||||
下一步:
|
||||
- 将 `reader-digest-flow` 正式默认路径切到 async job
|
||||
- 用真实 LLM 配置再做一次非 stub 的服务端冒烟验证
|
||||
|
||||
---
|
||||
|
||||
### [DOING][P0] FreshRSS 主日报 run 改为最小真异步 job
|
||||
|
||||
目标:
|
||||
- 解决 `run_freshrss_openclaw_pipeline` 在正式生产链路里仍为同步 MCP 调用、易超时的问题
|
||||
- 将 FreshRSS 主日报启动路径升级为可启动、可轮询、可读取结果的异步 job
|
||||
|
||||
要求:
|
||||
- 新增最小异步接口:
|
||||
- `start_freshrss_pipeline_job`
|
||||
- `get_freshrss_pipeline_job_status`
|
||||
- `get_freshrss_pipeline_job_result`
|
||||
- 状态目录固定落到:
|
||||
- `outputs/freshrss/pipeline_jobs/<job_id>/`
|
||||
- 至少包含:
|
||||
- `run-state.json`
|
||||
- `input.json`
|
||||
- `result.json`(成功时)
|
||||
- `job-report.json`
|
||||
- 执行模型优先使用后台子进程,不使用线程
|
||||
- 主业务逻辑继续复用 `run_freshrss_pipeline(...)`,不要重写日报核心逻辑
|
||||
- job 成功后结果中必须带回 `run_id` 与关键产物路径
|
||||
- README / handoff / OpenClaw 生产建议路径需要同步改成 async start path
|
||||
|
||||
当前进展:
|
||||
- 2026-04-11:问题定位完成,确认之前异步化的是 article-summary,不是主日报 run
|
||||
- 2026-04-11:已新增方案文档 `plans/freshrss-pipeline-async-job-plan.md`
|
||||
|
||||
下一步:
|
||||
- 复用 article-summary job runtime 骨架实现主日报 async job
|
||||
- 新增后台 runner 脚本
|
||||
- 暴露 3 个 MCP tools
|
||||
- 用真实 MCP 冒烟验证 `start -> status -> result`
|
||||
|
||||
---
|
||||
|
||||
@@ -202,6 +320,23 @@
|
||||
|
||||
---
|
||||
|
||||
### [DONE][P1] interest/watch 候选引擎从固定阈值改为百分位排名 + 增速因子
|
||||
|
||||
目标:
|
||||
- 解决固定阈值(total_count>=3)不随数据量自适应的问题
|
||||
- 引入趋势信号(growth 因子),识别近期集中爆发的词
|
||||
- 支持 7 天、41 天、200 天数据量下取同样的 top 5%/5%-20% 而不需调阈值
|
||||
|
||||
要求:
|
||||
- `build_review_bundle.py`:新增 percentile 和 growth 计算函数;候选池从固定阈值改为百分位 + 增速
|
||||
- `configs/term_cleanup_policy.json`:升级为 v2 schema,percentile/growth 替代绝对阈值
|
||||
- 不改 `generate_term_cleanup_suggestions.py` 和 `apply_term_suggestions.py`
|
||||
- 全量跑一次对比新旧产出,确认差异合理
|
||||
|
||||
方案文档:`plans/keyword-cleanup-interest-watch-engine-improvement.md`
|
||||
|
||||
---
|
||||
|
||||
### [DONE][P3] 更新 README / handoff / docs,明确 MCP 为正式入口
|
||||
|
||||
目标:
|
||||
@@ -228,6 +363,7 @@
|
||||
- 2026-04-07:新增架构设计文档 `plans/reader-mcp-architecture-design.md`
|
||||
- 2026-04-07:新增实施计划文档 `plans/reader-mcp-implementation-plan.md`
|
||||
- 2026-04-07:完成 `freshrss` pipeline 的 run-state 基础设施,新增 `runtime` 包并覆盖关键 stages 状态持久化。
|
||||
- 2026-04-14:已补 OpenClaw 文档导航、历史归档、design/notes/plans 导航,并统一当前正式口径为 async job 编排入口。
|
||||
|
||||
### 风险提醒
|
||||
|
||||
|
||||
@@ -23,34 +23,66 @@
|
||||
"前沿科技"
|
||||
],
|
||||
"interest_keywords": [
|
||||
"Java",
|
||||
"Go",
|
||||
"Python",
|
||||
"Spring",
|
||||
"Agent",
|
||||
"Agent Skills",
|
||||
"AgentScope",
|
||||
"AI Agent",
|
||||
"AI Coding Agent",
|
||||
"AliSQL",
|
||||
"Anthropic",
|
||||
"Claude",
|
||||
"Claude Code",
|
||||
"CLAUDE.md",
|
||||
"CLI",
|
||||
"Context Engineering",
|
||||
"Cursor",
|
||||
"ChatGPT",
|
||||
"DeepSeek",
|
||||
"FastAPI",
|
||||
"Gin",
|
||||
"Go",
|
||||
"gRPC",
|
||||
"MySQL",
|
||||
"PostgreSQL",
|
||||
"Redis",
|
||||
"Harness Engineering",
|
||||
"Hermes Agent",
|
||||
"Java",
|
||||
"Kafka",
|
||||
"微服务",
|
||||
"可观测性",
|
||||
"Kubernetes",
|
||||
"云原生",
|
||||
"AI Agent",
|
||||
"Agent",
|
||||
"LLM",
|
||||
"RAG",
|
||||
"Loop Engineering",
|
||||
"MCP",
|
||||
"Prompt Engineering",
|
||||
"Workflow",
|
||||
"向量数据库",
|
||||
"知识库",
|
||||
"MoE",
|
||||
"MySQL",
|
||||
"MySQL复制延迟",
|
||||
"OpenAI",
|
||||
"DeepSeek",
|
||||
"OpenClaw",
|
||||
"AliSQL",
|
||||
"MySQL复制延迟"
|
||||
"PostgreSQL",
|
||||
"Prompt Engineering",
|
||||
"Python",
|
||||
"RAG",
|
||||
"ReAct",
|
||||
"ReActAgent",
|
||||
"Redis",
|
||||
"Skill",
|
||||
"SKILL.md",
|
||||
"Skills",
|
||||
"Spring",
|
||||
"SubAgent",
|
||||
"TypeScript",
|
||||
"Vibe Coding",
|
||||
"Workflow",
|
||||
"上下文压缩",
|
||||
"上下文工程",
|
||||
"上下文管理",
|
||||
"云原生",
|
||||
"代码审查",
|
||||
"可观测性",
|
||||
"向量数据库",
|
||||
"多Agent协作",
|
||||
"大模型",
|
||||
"子Agent",
|
||||
"强化学习",
|
||||
"微服务",
|
||||
"渐进式披露",
|
||||
"知识库"
|
||||
]
|
||||
}
|
||||
+146
-2
@@ -1,5 +1,149 @@
|
||||
{
|
||||
"AI助手": "AI Agent",
|
||||
"Agent": "Agent",
|
||||
"Agent框架": "Agent",
|
||||
"Agent能力": "Agent Skills",
|
||||
"智能体": "AI Agent",
|
||||
"Agentic架构": "Agentic架构",
|
||||
"多Agent协作": "多Agent协作",
|
||||
"多智能体架构": "多Agent协作",
|
||||
"Multi-Agent": "多Agent",
|
||||
"Subagent": "子Agent",
|
||||
"Sub Agents验证": "子Agent",
|
||||
"子Agent": "子Agent",
|
||||
"子智能体": "子Agent",
|
||||
"Coding Agent": "AI Coding Agent",
|
||||
"AI编程": "AI Coding Agent",
|
||||
"AI辅助编程": "AI Coding Agent",
|
||||
"代码生成": "AI代码生成",
|
||||
"代码审查": "Code Review",
|
||||
"Prompt": "Prompt Engineering",
|
||||
"Prompt Caching": "提示缓存",
|
||||
"RAG": "RAG",
|
||||
"图文RAG": "RAG",
|
||||
"Prompt架构": "Prompt Engineering"
|
||||
"向量检索": "向量检索",
|
||||
"向量嵌入": "向量嵌入",
|
||||
"Multi-Token Prediction": "多Token预测",
|
||||
"Pair-In Pair-Out": "PIPO架构",
|
||||
"PIPO": "PIPO架构",
|
||||
"上下文管理": "上下文管理",
|
||||
"上下文卸载": "上下文卸载",
|
||||
"Self-GC": "上下文压缩",
|
||||
"记忆管理": "上下文管理",
|
||||
"会话管理": "上下文管理",
|
||||
"Harness Engineering": "Harness工程化",
|
||||
"Harness架构": "Harness工程化",
|
||||
"Harness": "Harness工程化",
|
||||
"Loop Engineering": "Loop Engineering",
|
||||
"推理加速": "推理加速",
|
||||
"推理深度": "推理深度",
|
||||
"长链路推理": "长链路推理",
|
||||
"RLVR": "RLVR",
|
||||
"GRPO": "GRPO",
|
||||
"强化学习": "强化学习",
|
||||
"Multi-Agent RL": "多Agent强化学习",
|
||||
"Viking AI搜索": "AI搜索",
|
||||
"Viking AI Search": "AI搜索",
|
||||
"智能搜索": "AI搜索",
|
||||
"SearchCLI": "CLI搜索",
|
||||
"视频生成": "AI视频生成",
|
||||
"视频生成模型": "AI视频生成",
|
||||
"LingBot-Video": "AI视频生成",
|
||||
"视觉自回归模型": "AI视频生成",
|
||||
"火山云数据库PostgreSQL Serverless版": "Serverless数据库",
|
||||
"PostgreSQL": "PostgreSQL",
|
||||
"MySQL": "MySQL",
|
||||
"OceanBase": "OceanBase",
|
||||
"StarRocks": "StarRocks",
|
||||
"Milvus": "Milvus",
|
||||
"Seal AI Zone": "AI安全",
|
||||
"NEX沙箱": "沙箱隔离",
|
||||
"MicroVM": "沙箱隔离",
|
||||
"安全左移": "安全左移",
|
||||
"安全中台": "AI安全",
|
||||
"成本降低": "成本优化",
|
||||
"成本杠杆": "成本优化",
|
||||
"Scale-to-Zero": "弹性伸缩",
|
||||
"Data as Git": "数据分支管理",
|
||||
"Schema Diff": "Schema对比",
|
||||
"Time Travel": "数据回溯",
|
||||
"多端架构": "多端架构",
|
||||
"契约化": "契约化架构",
|
||||
"大仓": "大仓工程化",
|
||||
"Vibe Coding": "Vibe Coding",
|
||||
"LLM Judge": "LLM评估",
|
||||
"SWE-Bench": "SWE-Bench",
|
||||
"SWE Bench Pro": "SWE-Bench",
|
||||
"SWE-Bench Pro": "SWE-Bench",
|
||||
"Verification Agent": "验证Agent",
|
||||
"CLI工具": "CLI",
|
||||
"CLI": "CLI",
|
||||
"漏桶算法": "限流架构",
|
||||
"固定窗口限流": "限流架构",
|
||||
"Suspend消费控制": "限流架构",
|
||||
"RocketMQ LiteTopic": "消息队列",
|
||||
"LLM Wiki": "LLM知识库",
|
||||
"知识工程": "知识工程",
|
||||
"语义资产": "语义资产管理",
|
||||
"知识图谱": "知识图谱",
|
||||
"知识库沉淀": "知识管理",
|
||||
"Skill": "Skill",
|
||||
"Skill Hub": "技能生态",
|
||||
"具身智能": "具身智能",
|
||||
"Open X-Embodiment": "具身智能",
|
||||
"YOLO Classifier": "目标检测",
|
||||
"MCP": "MCP",
|
||||
"MCP连接器": "MCP",
|
||||
"缓存击穿": "缓存优化",
|
||||
"GPU算力调度": "算力调度",
|
||||
"异构资源": "异构计算",
|
||||
"XPU": "异构计算",
|
||||
"弹性RDMA": "RDMA网络",
|
||||
"国内主流GPU": "国产芯片",
|
||||
"国产AI芯片": "国产芯片",
|
||||
"Paxos协议": "分布式一致性",
|
||||
"Token": "Token管理",
|
||||
"Token效率": "Token管理",
|
||||
"百万token上下文": "长上下文",
|
||||
"MoE": "MoE架构",
|
||||
"MoE架构": "MoE架构",
|
||||
"思维链": "思维链",
|
||||
"CoT Distillation": "思维链蒸馏",
|
||||
"自然语言驱动": "自然语言交互",
|
||||
"NL2SQL": "NL2SQL",
|
||||
"AI对齐": "AI对齐",
|
||||
"注意力机制": "注意力机制",
|
||||
"多模态": "多模态",
|
||||
"音视频工作台": "音视频处理",
|
||||
"AI助手": "AI Agent",
|
||||
"Agent架构": "AI Agent",
|
||||
"Agent专业化": "AI Agent",
|
||||
"Agent Teams": "多Agent协作",
|
||||
"Agentic Engineering": "AI Agent",
|
||||
"AI智能体": "AI Agent",
|
||||
"LLM Agent": "AI Agent",
|
||||
"AI Harness": "Harness Engineering",
|
||||
"AI代码生成": "AI Coding Agent",
|
||||
"Memory管理": "上下文管理",
|
||||
"Agent Skill": "Agent Skills",
|
||||
"Binlog": "binlog",
|
||||
"vibe coding": "Vibe Coding",
|
||||
"Agent组织化协作平台": "Agent协作平台",
|
||||
"Anthropic": "Anthropic",
|
||||
"OpenClaw": "OpenClaw",
|
||||
"WorkBuddy": "WorkBuddy",
|
||||
"Claude": "Claude",
|
||||
"ChatGPT": "ChatGPT",
|
||||
"GPT": "GPT",
|
||||
"Opus": "Opus",
|
||||
"Sonnet": "Sonnet",
|
||||
"Grok": "Grok",
|
||||
"Qwen": "Qwen",
|
||||
"GLM": "GLM",
|
||||
"Claude Code": "Claude Code",
|
||||
"Cursor": "Cursor",
|
||||
"Codex": "Codex",
|
||||
"Pi": "Pi",
|
||||
"CoT": "CoT",
|
||||
"SVG": "SVG",
|
||||
"TTS": "TTS"
|
||||
}
|
||||
@@ -1,4 +1,595 @@
|
||||
{
|
||||
"schema_version": "v1",
|
||||
"entries": []
|
||||
"entries": [
|
||||
{
|
||||
"applied_at": "2026-04-08T02:34:14.194320Z",
|
||||
"action": "add_watch_term",
|
||||
"term": "A2A",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=0.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"suggestion_date": "2026-04-08",
|
||||
"based_on_days": 7
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-04-08T02:34:14.194320Z",
|
||||
"action": "add_watch_term",
|
||||
"term": "Agentic Loop",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=1.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"suggestion_date": "2026-04-08",
|
||||
"based_on_days": 7
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-04-08T02:34:14.194320Z",
|
||||
"action": "add_watch_term",
|
||||
"term": "AI Gateway",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=1.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"suggestion_date": "2026-04-08",
|
||||
"based_on_days": 7
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-04-08T02:34:14.194320Z",
|
||||
"action": "add_watch_term",
|
||||
"term": "Claude Skills",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=1.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"suggestion_date": "2026-04-08",
|
||||
"based_on_days": 7
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-04-08T02:34:14.194320Z",
|
||||
"action": "add_watch_term",
|
||||
"term": "Cron",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=0.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"suggestion_date": "2026-04-08",
|
||||
"based_on_days": 7
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-04-08T02:34:14.194320Z",
|
||||
"action": "add_watch_term",
|
||||
"term": "CoPaw",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=0.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"suggestion_date": "2026-04-08",
|
||||
"based_on_days": 7
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-04-08T02:34:14.194320Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Claude Code",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=7, days_seen=4, recent_count=4.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"suggestion_date": "2026-04-08",
|
||||
"based_on_days": 7
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-04-08T02:34:14.194320Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Agent Skills",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=3, recent_count=0.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"suggestion_date": "2026-04-08",
|
||||
"based_on_days": 7
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-04-08T02:34:14.194320Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "SubAgent",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=4, recent_count=2.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"suggestion_date": "2026-04-08",
|
||||
"based_on_days": 7
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-04-08T02:34:14.194320Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "AgentScope",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=4, recent_count=1.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"suggestion_date": "2026-04-08",
|
||||
"based_on_days": 7
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-04-08T02:34:14.194320Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "ReActAgent",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=3, days_seen=3, recent_count=1.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"suggestion_date": "2026-04-08",
|
||||
"based_on_days": 7
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Anthropic",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=13, days_seen=10, recent_count=13.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Harness Engineering",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=12, days_seen=11, recent_count=12.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Skill",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=11, days_seen=9, recent_count=11.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "上下文工程",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=6, recent_count=6.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "多Agent协作",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=6, recent_count=6.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Claude",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=5, recent_count=6.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "上下文管理",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=5, recent_count=6.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "渐进式披露",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=5, days_seen=5, recent_count=5.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Skills",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=5, days_seen=4, recent_count=5.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "SKILL.md",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=4, recent_count=4.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "CLAUDE.md",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=4, recent_count=4.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "上下文压缩",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=4, recent_count=4.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "AI Coding Agent",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=3, days_seen=3, recent_count=3.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Hermes Agent",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=3, days_seen=3, recent_count=3.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Vibe Coding",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=5, days_seen=4, recent_count=5.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Context Engineering",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=3, days_seen=3, recent_count=3.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Cursor",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=3, days_seen=3, recent_count=3.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "大模型",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=3, recent_count=4.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "TypeScript",
|
||||
"reason": "Core language for AI agent development (e.g., Claude Code, Cursor) and backend engineering, complements existing Python/Java/Go keywords.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "代码审查",
|
||||
"reason": "Chinese term for 'code review', a key practice in backend engineering and AI agent development workflows.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "Channels",
|
||||
"reason": "Too generic; could refer to communication channels, YouTube channels, or software channels, not specific to user's focus areas.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "Memory",
|
||||
"reason": "Extremely broad term; could refer to computer memory, human memory, or memory in various contexts, not discriminative enough.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "Prompt",
|
||||
"reason": "Already covered by 'Prompt Engineering' as a more specific term; 'Prompt' alone is too broad and matches many unrelated articles.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "AGI",
|
||||
"reason": "Too broad and speculative; not directly actionable for the user's practical engineering focus areas.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "AI日报",
|
||||
"reason": "Generic news term; not a technical concept or tool, would add noise to the keyword index.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "AIHOT",
|
||||
"reason": "Unclear meaning, likely a brand or aggregator, not a specific technical term.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "All In Code",
|
||||
"reason": "Too vague; could refer to a podcast, a philosophy, or a project, not a specific technical concept.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "auto-twitter-campaign",
|
||||
"reason": "Too specific to a single project/tool, not a general interest keyword for the user's focus areas.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "ChangeSet",
|
||||
"reason": "Generic term used in version control and databases; too broad to be a useful filter.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "Lumina",
|
||||
"reason": "Unclear reference; could be a product, framework, or brand, not clearly aligned with user's focus.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "OpenViking",
|
||||
"reason": "Unclear reference; not a known tool or concept in the user's stated focus areas.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "Seedance 2.0",
|
||||
"reason": "Unclear reference; likely a product or version, not a general technical term.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "质量门禁",
|
||||
"reason": "Chinese term for 'quality gate', too generic in software engineering; not specific to user's focus areas.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Agent Skill",
|
||||
"reason": "Singular variant",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "Agent Skills",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Binlog",
|
||||
"reason": "Case variant (auto-ranked)",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "binlog",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Coding Agent",
|
||||
"reason": "Abbreviated form of 'AI Coding Agent', referring to the same concept.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "AI Coding Agent",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Subagent",
|
||||
"reason": "Case variant",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "SubAgent",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Subagents",
|
||||
"reason": "Plural variant",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "SubAgent",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "vibe coding",
|
||||
"reason": "Case variant",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "Vibe Coding",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Agent架构",
|
||||
"reason": "Chinese translation of 'Agent architecture', a core concept in AI Agent engineering.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "AI Agent",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Agent专业化",
|
||||
"reason": "Chinese term for 'Agent specialization', directly related to Agent engineering.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "AI Agent",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Agent Teams",
|
||||
"reason": "English equivalent of 'Multi-Agent collaboration', same concept.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "多Agent协作",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Agentic Engineering",
|
||||
"reason": "Broader term for engineering with AI agents, closely related to Agent engineering focus.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "AI Agent",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "CLI工具",
|
||||
"reason": "Chinese translation of 'CLI tool', same concept.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "CLI",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "AI编程",
|
||||
"reason": "Chinese term for 'AI programming', closely related to AI Coding Agent.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "AI Coding Agent",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "记忆管理",
|
||||
"reason": "Chinese term for 'memory management', closely related to context management in LLM applications.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "上下文管理",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "会话管理",
|
||||
"reason": "Chinese term for 'session management', related to context management in LLM applications.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "上下文管理",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-07-15T02:17:50.155586Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "强化学习",
|
||||
"reason": "top 0.8% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=10, days_seen=10, recent_count=10.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
|
||||
"suggestion_date": "2026-07-15",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-07-15T02:17:50.155586Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "ReAct",
|
||||
"reason": "top 1.1% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=8, days_seen=8, recent_count=8.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
|
||||
"suggestion_date": "2026-07-15",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-07-15T02:17:50.155586Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "CLI",
|
||||
"reason": "top 1.1% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=8, days_seen=7, recent_count=8.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
|
||||
"suggestion_date": "2026-07-15",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-07-15T02:17:50.155586Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "子Agent",
|
||||
"reason": "top 1.3% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=7, days_seen=6, recent_count=7.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
|
||||
"suggestion_date": "2026-07-15",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-07-15T02:17:50.155586Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Loop Engineering",
|
||||
"reason": "top 1.5% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=6, recent_count=6.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
|
||||
"suggestion_date": "2026-07-15",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-07-15T02:17:50.155586Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "MoE",
|
||||
"reason": "top 2.0% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=5, days_seen=4, recent_count=5.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
|
||||
"suggestion_date": "2026-07-15",
|
||||
"based_on_days": 365
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -1,14 +1,12 @@
|
||||
{
|
||||
"schema_version": "v1",
|
||||
"schema_version": "v2",
|
||||
"interest_keyword_review": {
|
||||
"min_total_count": 3,
|
||||
"min_days_seen": 2
|
||||
"percentile_max": 0.05,
|
||||
"growth_promotion": 0.5
|
||||
},
|
||||
"watch_term_review": {
|
||||
"min_total_count": 1,
|
||||
"min_days_seen": 1,
|
||||
"max_total_count": 2,
|
||||
"max_days_seen": 2
|
||||
"percentile_min": 0.05,
|
||||
"percentile_max": 0.20
|
||||
},
|
||||
"alias_review": {
|
||||
"min_total_count": 2,
|
||||
@@ -19,7 +17,10 @@
|
||||
"max_days_seen": 2
|
||||
},
|
||||
"notes": [
|
||||
"当前阶段采用保守阈值,避免在低样本条件下直接扩充 interest_keywords。",
|
||||
"watch_terms 先用于观察,后续再决定是否升格为 interest_keywords 或进入 alias/stopword 配置。"
|
||||
"v2: interest/watch 使用百分位排名 + 增速因子替代固定阈值",
|
||||
"percentile 越小表示排名越高(top 5% = percentile 0.05)",
|
||||
"growth = recent_count / total_count,衡量近期活跃度",
|
||||
"watch_term_review 的 percentile_min 可理解为兴趣边界下限,低于此值的词归入 interest 候选",
|
||||
"growth_promotion(默认 0.5)用于识别近期集中爆发词,即使排位不高也主动推荐确认"
|
||||
]
|
||||
}
|
||||
+121
-1
@@ -1,6 +1,126 @@
|
||||
[
|
||||
"1688",
|
||||
"AGI",
|
||||
"AIHOT",
|
||||
"AI日报",
|
||||
"All In Code",
|
||||
"Andrej Karpathy",
|
||||
"Anthropic",
|
||||
"auto-twitter-campaign",
|
||||
"Boundaries",
|
||||
"ChangeSet",
|
||||
"Channels",
|
||||
"Claude Fable 5",
|
||||
"Claude Mythos",
|
||||
"Cohere",
|
||||
"Confidence Head",
|
||||
"Cosmos 3",
|
||||
"Databricks",
|
||||
"DINOv2",
|
||||
"domain-mapping",
|
||||
"Dropbox",
|
||||
"EchoGen",
|
||||
"FLUX.1-dev VAE",
|
||||
"GB300 GPU",
|
||||
"GLM 5.2",
|
||||
"GLM5.0",
|
||||
"GPT-5.5",
|
||||
"GPT-Live",
|
||||
"GPT5.5",
|
||||
"Grok 4.5",
|
||||
"GrowBrain",
|
||||
"iMedImage",
|
||||
"iMedLoop",
|
||||
"iMedMaaS",
|
||||
"iMedStudio",
|
||||
"J-space",
|
||||
"JLens",
|
||||
"John Jumper",
|
||||
"J空间",
|
||||
"KAIROS",
|
||||
"KubeRay",
|
||||
"LibTV Agent",
|
||||
"LingBot-Video",
|
||||
"Lumina",
|
||||
"Markdown",
|
||||
"Marvis",
|
||||
"MDASH",
|
||||
"Meta Superintelligence Labs",
|
||||
"MTS",
|
||||
"Muse Image",
|
||||
"Muse Video",
|
||||
"N-gram Embedding",
|
||||
"OCP China",
|
||||
"OCP China 2026",
|
||||
"On-Policy Distillation",
|
||||
"OPC训练营",
|
||||
"OpenAI",
|
||||
"OpenBMC",
|
||||
"OpenClaw",
|
||||
"OpenViking",
|
||||
"Opus 4.8",
|
||||
"Qwen3",
|
||||
"Qwen3-30B-A3B",
|
||||
"RAS API",
|
||||
"Redfish",
|
||||
"ScMoE",
|
||||
"Seal AI Zone",
|
||||
"SealRouter",
|
||||
"Seedance 2.0",
|
||||
"Sonnet 5",
|
||||
"Spec模式",
|
||||
"STE固件团队",
|
||||
"Three.js",
|
||||
"Unity AI Gateway",
|
||||
"Vant Weapp",
|
||||
"WeTV",
|
||||
"WorkBuddy",
|
||||
"wpc",
|
||||
"YOLO Classifier",
|
||||
"一人公司",
|
||||
"中国科学技术大学",
|
||||
"五大扶持体系",
|
||||
"出门问问",
|
||||
"分镜",
|
||||
"剧本",
|
||||
"奋斗文化",
|
||||
"字节跳动",
|
||||
"小银",
|
||||
"得力",
|
||||
"德适科技",
|
||||
"成都天府长岛",
|
||||
"扣子",
|
||||
"星云平台",
|
||||
"火山引擎",
|
||||
"百度百舸",
|
||||
"百炼网关",
|
||||
"科大讯飞",
|
||||
"腾讯云开发者社区",
|
||||
"腾讯混元Hy3",
|
||||
"蚂蚁灵波",
|
||||
"贝尔实验室",
|
||||
"质量门禁",
|
||||
"配乐",
|
||||
"配音",
|
||||
"银行客户经理",
|
||||
"飞书妙搭",
|
||||
"飞盘物理",
|
||||
"奋斗文化"
|
||||
"自动化",
|
||||
"定时任务",
|
||||
"开源模型",
|
||||
"陌生化",
|
||||
"AlphaFold",
|
||||
"Brand Kit",
|
||||
"DataWorks",
|
||||
"Enhance-Nanocodec",
|
||||
"IRIS Codec",
|
||||
"Lovart",
|
||||
"MiniMax M3",
|
||||
"Gemini 3.5 Flash",
|
||||
"Codex",
|
||||
"CodeBuddy",
|
||||
"Claude Cowork",
|
||||
"AGENTS.md",
|
||||
"Claude",
|
||||
"RLVR"
|
||||
]
|
||||
@@ -1,5 +1,48 @@
|
||||
{
|
||||
"schema_version": "v1",
|
||||
"updated_at": "2026-03-27T00:00:00Z",
|
||||
"terms": []
|
||||
"updated_at": "2026-07-15T02:17:50.155586Z",
|
||||
"terms": [
|
||||
{
|
||||
"term": "A2A",
|
||||
"added_at": "2026-04-08T02:34:14.194320Z",
|
||||
"source": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=0.",
|
||||
"status": "watching"
|
||||
},
|
||||
{
|
||||
"term": "Agentic Loop",
|
||||
"added_at": "2026-04-08T02:34:14.194320Z",
|
||||
"source": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=1.",
|
||||
"status": "watching"
|
||||
},
|
||||
{
|
||||
"term": "AI Gateway",
|
||||
"added_at": "2026-04-08T02:34:14.194320Z",
|
||||
"source": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=1.",
|
||||
"status": "watching"
|
||||
},
|
||||
{
|
||||
"term": "Claude Skills",
|
||||
"added_at": "2026-04-08T02:34:14.194320Z",
|
||||
"source": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=1.",
|
||||
"status": "watching"
|
||||
},
|
||||
{
|
||||
"term": "CoPaw",
|
||||
"added_at": "2026-04-08T02:34:14.194320Z",
|
||||
"source": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=0.",
|
||||
"status": "watching"
|
||||
},
|
||||
{
|
||||
"term": "Cron",
|
||||
"added_at": "2026-04-08T02:34:14.194320Z",
|
||||
"source": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=0.",
|
||||
"status": "watching"
|
||||
}
|
||||
]
|
||||
}
|
||||
+50
-38
@@ -1,6 +1,25 @@
|
||||
# 文档索引
|
||||
|
||||
## 当前目录结构
|
||||
## 当前最短阅读路径
|
||||
|
||||
1. `README.md`
|
||||
- 仓库入口与常用脚本
|
||||
2. `docs/current/context-reset-brief.md`
|
||||
- 当前状态的最短摘要
|
||||
3. `docs/openclaw/README.md`
|
||||
- OpenClaw 集成文档导航
|
||||
4. `docs/openclaw/openclaw-handoff.md`
|
||||
- OpenClaw 接手总览
|
||||
5. `docs/openclaw/openclaw-orchestration-flow.md`
|
||||
- OpenClaw 正式编排手册
|
||||
6. `docs/design/README.md`
|
||||
- 设计文档导航,区分当前有效设计与背景草案
|
||||
7. `plans/README.md`
|
||||
- 规划文档导航
|
||||
8. `TODO.md`
|
||||
- 当前任务状态
|
||||
|
||||
## 目录结构
|
||||
|
||||
- `docs/README.md`
|
||||
- 文档总索引
|
||||
@@ -8,45 +27,18 @@
|
||||
- 当前状态、收束入口、阶段导航
|
||||
- `docs/design/`
|
||||
- 当前实现的设计文档
|
||||
- `docs/design/README.md`
|
||||
- 设计文档导航
|
||||
- `docs/openclaw/`
|
||||
- OpenClaw 日报聚合与下游对象设计
|
||||
- OpenClaw 集成文档、对象规范与历史归档
|
||||
- `docs/notes/`
|
||||
- 较上层的方案笔记与非最终设计
|
||||
- `docs/notes/README.md`
|
||||
- notes 导航
|
||||
- `docs/archive/`
|
||||
- 历史归档,不作为最新事实来源
|
||||
|
||||
## 当前推荐阅读顺序
|
||||
|
||||
1. `docs/current/context-reset-brief.md`
|
||||
- 当前真实进度与下一步入口
|
||||
2. `docs/openclaw/openclaw-handoff.md`
|
||||
- 给 OpenClaw 的接手说明、环境变量、MCP 调用方式与已知限制
|
||||
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||
- 提供给 OpenClaw 的单篇结构化输入字段说明
|
||||
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||
- 提供给 OpenClaw 的批量投递 envelope 说明
|
||||
5. `docs/design/summary-mcp-service-design.md`
|
||||
- 当前 MCP 服务的职责、接口和边界
|
||||
6. `docs/design/filter-rule-engine-design.md`
|
||||
- 过滤层的输入输出、规则结构与当前实现
|
||||
7. `docs/design/filter-rule-engine-usage.md`
|
||||
- 规则怎么写、怎么跑、结果怎么解读的使用说明
|
||||
8. `docs/design/daily-keyword-index-design.md`
|
||||
- 日报级词元库与周期性词元清洗 skill 设计
|
||||
9. `docs/design/markdown-sink-design.md`
|
||||
- 第一版 Markdown sink 的输入输出、目录结构与落地方式
|
||||
10. `docs/openclaw/openclaw-daily-digest-refactor.md`
|
||||
- 为什么要从单篇入库改成 OpenClaw 日报聚合链路
|
||||
11. `docs/openclaw/article-candidate-daily-digest-schema.md`
|
||||
- `ArticleCandidateRecord`、`OpenClawCandidateInput` 与 `DailyDigest` 的正式设计
|
||||
12. `docs/design/source-schema-design.md`
|
||||
- `source -> item -> document` 的对象设计
|
||||
13. `docs/notes/reading-pipeline-design-notes.md`
|
||||
- 更上层的阅读流方案与阶段划分
|
||||
14. `docs/design/summary-loop-explained.md`
|
||||
- 当前 LLM 摘要校验闭环的解释
|
||||
|
||||
## 当前文档分层
|
||||
## 按主题阅读
|
||||
|
||||
### 1. 当前状态与导航
|
||||
|
||||
@@ -58,11 +50,15 @@
|
||||
- 当前阶段状态的最短摘要
|
||||
- `docs/README.md`
|
||||
- 文档索引与阅读顺序
|
||||
- `plans/README.md`
|
||||
- 规划文档导航
|
||||
|
||||
### 2. 当前实现设计
|
||||
|
||||
- `docs/design/README.md`
|
||||
- 设计文档导航与状态说明
|
||||
- `docs/design/summary-mcp-service-design.md`
|
||||
- 当前内容提取 MCP 的真实设计
|
||||
- 早期 content-extract MCP 设计草案,现主要保留背景参考价值
|
||||
- `docs/design/summary-core-interface-design.md`
|
||||
- 摘要/提取内核的接口抽象
|
||||
- `docs/design/source-schema-design.md`
|
||||
@@ -82,19 +78,23 @@
|
||||
|
||||
### 3. OpenClaw 与下游设计
|
||||
|
||||
- `docs/openclaw/README.md`
|
||||
- OpenClaw 相关文档导航与归档边界
|
||||
- `docs/openclaw/openclaw-handoff.md`
|
||||
- OpenClaw 接手所需的运行说明、工具入口与已知限制
|
||||
- `docs/openclaw/openclaw-orchestration-flow.md`
|
||||
- OpenClaw 编排层的正式运行手册
|
||||
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||
- 提供给 OpenClaw 的单篇结构化输入字段说明
|
||||
- `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||
- 提供给 OpenClaw 的批量投递 envelope 说明
|
||||
- `docs/openclaw/openclaw-daily-digest-refactor.md`
|
||||
- 改造为 OpenClaw 日报聚合链路的原因与目标结构
|
||||
- `docs/openclaw/article-candidate-daily-digest-schema.md`
|
||||
- `ArticleCandidateRecord`、`OpenClawCandidateInput` 与 `DailyDigest` 的字段设计与对象关系
|
||||
|
||||
### 4. 方案笔记
|
||||
|
||||
- `docs/notes/README.md`
|
||||
- notes 导航与使用边界
|
||||
- `docs/notes/reading-pipeline-design-notes.md`
|
||||
- 整体阅读流、规则、sink、push 的方案笔记
|
||||
|
||||
@@ -102,11 +102,23 @@
|
||||
|
||||
- `docs/archive/content-extract-mcp-mvp-archive.md`
|
||||
- MVP 阶段归档,部分状态已被后续进展覆盖
|
||||
- `docs/openclaw/archive/README.md`
|
||||
- OpenClaw 历史文档归档说明
|
||||
- `docs/openclaw/archive/formalization-summary-2026-04-07.md`
|
||||
- 第一阶段正式化总结
|
||||
- `docs/openclaw/archive/openclaw-daily-digest-refactor.md`
|
||||
- 早期日报聚合改造背景
|
||||
- `docs/openclaw/archive/digest-optimization-summary.md`
|
||||
- 早期 digest 优化总结
|
||||
- `docs/openclaw/archive/p1-status-reconciliation-plan-2026-04-14.md`
|
||||
- `resume` / 状态收敛问题的阶段修复计划与回填
|
||||
|
||||
## 当前文档维护原则
|
||||
## 维护原则
|
||||
|
||||
- `docs/current/context-reset-brief.md` 记录当前最新状态
|
||||
- `TODO.md` 记录任务优先级与下一步
|
||||
- `plans/README.md` 负责规划文档分层与导航
|
||||
- `docs/design/README.md` 负责设计文档分层与导航
|
||||
- `outputs/README.md` 记录当前输出目录约定
|
||||
- `docs/archive/content-extract-mcp-mvp-archive.md` 只当历史快照,不再作为最新事实来源
|
||||
- 新增阶段性进展,优先更新 `README.md`、`TODO.md`、`docs/current/context-reset-brief.md`
|
||||
@@ -2,152 +2,104 @@
|
||||
|
||||
## 当前结论
|
||||
|
||||
当前仓库已经具备交付给 OpenClaw 的基础条件。
|
||||
当前仓库已经具备作为 OpenClaw 上游服务的正式基础能力。
|
||||
|
||||
当前主链路是:
|
||||
当前正式主链路是:
|
||||
|
||||
`FreshRSS 未读 -> RSS 内容提取 -> LLM 总结 -> 规则过滤 -> OpenClaw delivery payload`
|
||||
|
||||
OpenClaw 应通过 MCP 工具 `run_freshrss_openclaw_pipeline` 调用这条链路,而不是自行拼接脚本。
|
||||
当前正式控制面已经收口为异步 job:
|
||||
|
||||
## 当前已完成
|
||||
- 主日报:`start_freshrss_pipeline_job -> poll -> get result`
|
||||
- 恢复:`inspect_resume_plan -> start_resume_job -> poll -> get result`
|
||||
- 单篇总结:`start_article_summary_job -> poll -> get result`
|
||||
|
||||
- 已完成 FreshRSS `greader` API 接入与未读拉取
|
||||
- 已完成 FreshRSS 条目到标准化 `item` 的映射
|
||||
- 已完成 RSS-first 提取策略
|
||||
- 已完成 LLM 总结与校验闭环
|
||||
- 已完成规则引擎过滤
|
||||
- 已完成 `ArticleCandidateRecord` 与 `OpenClawCandidateInput` 分层
|
||||
- 已完成 `OpenClawDeliveryPayload` 批量投递结构
|
||||
- 已完成 FreshRSS 已读状态回写
|
||||
- 已完成“仅在最终 payload 成功写盘后再标记已读”的语义
|
||||
- 已完成 MCP 工具 `run_freshrss_openclaw_pipeline`
|
||||
- 已完成默认精简输出模式,减少中间文件
|
||||
- 已完成日报级 `keywords` 词元库与全局词频统计
|
||||
- 已完成 `keyword-cleanup-review` skill 骨架与 review bundle 脚本
|
||||
- 已完成低复杂治理层:`term_cleanup_policy` / `term_watchlist` / `term_change_log`
|
||||
- 已完成采纳建议写回脚本 `scripts/apply_term_suggestions.py`
|
||||
同步 `run_freshrss_openclaw_pipeline`、`resume_run`、`generate_article_summaries` 仍保留,但只用于 debug / fallback。
|
||||
|
||||
## 当前 MCP 工具
|
||||
## 当前权威入口
|
||||
|
||||
当前服务入口:
|
||||
先看这些文档:
|
||||
|
||||
- `src/summary_mcp/server.py`
|
||||
1. `README.md`
|
||||
2. `docs/README.md`
|
||||
3. `docs/openclaw/README.md`
|
||||
4. `docs/openclaw/openclaw-handoff.md`
|
||||
5. `docs/openclaw/openclaw-orchestration-flow.md`
|
||||
6. `plans/README.md`
|
||||
7. `TODO.md`
|
||||
|
||||
当前暴露的 MCP 工具:
|
||||
如果问题是 OpenClaw 集成、状态分支或恢复策略,优先看 `docs/openclaw/`,不要先翻历史计划。
|
||||
|
||||
- `extract_url_content`
|
||||
- `extract_item_content`
|
||||
- `filter_summary_result`
|
||||
- `run_freshrss_openclaw_pipeline`
|
||||
## 当前正式能力
|
||||
|
||||
其中生产主入口是:
|
||||
- FreshRSS 主日报 run 会落地 `run-state.json`
|
||||
- `get_run_status` / `list_runs` / `list_run_artifacts` 提供 run 级观测
|
||||
- `get_delivery_payload` / `get_run_report` 提供正式结果读取
|
||||
- 查询层已经支持 stale state 与终态 artifacts 的状态收敛
|
||||
- 主日报正式启动已切到 async job
|
||||
- `resume` 已切到 async job,并在执行前先做 `inspect_resume_plan`
|
||||
- 生产恢复依赖 `summary/summary-batch.json` 与 `candidates/candidate-batch.json`
|
||||
- 单篇总结也已补齐 async job 形态
|
||||
|
||||
- `run_freshrss_openclaw_pipeline`
|
||||
## 当前关键代码入口
|
||||
|
||||
## 当前关键文件
|
||||
- MCP 服务入口:`src/summary_mcp/server.py`
|
||||
- 主日报 workflow:`src/summary_mcp/workflows/freshrss_pipeline.py`
|
||||
- run / artifact 查询:`src/summary_mcp/runtime/query_service.py`
|
||||
- 主日报 async job:`src/summary_mcp/runtime/freshrss_pipeline_jobs.py`
|
||||
- resume 预检与恢复:`src/summary_mcp/runtime/resume_service.py`
|
||||
- resume async job:`src/summary_mcp/runtime/resume_jobs.py`
|
||||
- 单篇总结 async job:`src/summary_mcp/runtime/article_summary_jobs.py`
|
||||
- 关键词治理:`src/summary_mcp/core/keyword_index.py`
|
||||
|
||||
- MCP 服务入口
|
||||
- `src/summary_mcp/server.py`
|
||||
- FreshRSS 统一工作流
|
||||
- `src/summary_mcp/workflows/freshrss_pipeline.py`
|
||||
- 词元统计核心
|
||||
- `src/summary_mcp/core/keyword_index.py`
|
||||
- 词元统计模型
|
||||
- `src/summary_mcp/models/keyword_index.py`
|
||||
- 摘要循环
|
||||
- `src/summary_mcp/core/summary_loop.py`
|
||||
- 提取主流程
|
||||
- `src/summary_mcp/core/pipeline.py`
|
||||
- FreshRSS 集成
|
||||
- `src/summary_mcp/integrations/freshrss.py`
|
||||
- 规则引擎
|
||||
- `src/summary_mcp/filters/engine.py`
|
||||
- LLM 结果校验
|
||||
- `src/summary_mcp/validators/llm_result.py`
|
||||
- OpenClaw candidate 模型
|
||||
- `src/summary_mcp/models/article_candidate.py`
|
||||
- OpenClaw delivery 模型
|
||||
- `src/summary_mcp/models/openclaw_delivery.py`
|
||||
- 生产脚本入口
|
||||
- `scripts/run_freshrss_pipeline.py`
|
||||
- 词元统计重建脚本
|
||||
- `scripts/build_keyword_index.py`
|
||||
- 词元清洗 skill
|
||||
- `skills/keyword-cleanup-review/SKILL.md`
|
||||
- skill review bundle 脚本
|
||||
- `skills/keyword-cleanup-review/scripts/build_review_bundle.py`
|
||||
- 采纳建议写回脚本
|
||||
- `scripts/apply_term_suggestions.py`
|
||||
- 清洗治理配置
|
||||
- `configs/term_cleanup_policy.json`
|
||||
- `configs/term_watchlist.json`
|
||||
- `configs/term_change_log.json`
|
||||
- OpenClaw 交接说明
|
||||
- `docs/openclaw/openclaw-handoff.md`
|
||||
## 当前核心产物
|
||||
|
||||
## 当前输出规则
|
||||
主日报稳定产物:
|
||||
|
||||
默认生产模式只输出:
|
||||
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/raw/freshrss.raw.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/summary/summary-batch.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/candidates/candidate-batch.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/candidates/digest-brief.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json`
|
||||
|
||||
- `raw/freshrss.raw.json`
|
||||
- `candidates/openclaw-delivery-payload.json`
|
||||
- `run-report.json`
|
||||
job 状态目录:
|
||||
|
||||
同时会更新本地运行数据:
|
||||
- `outputs/freshrss/pipeline_jobs/<job_id>/`
|
||||
- `outputs/freshrss/resume_jobs/<job_id>/`
|
||||
- `outputs/freshrss/article_summary_jobs/<job_id>/`
|
||||
|
||||
关键词运行数据:
|
||||
|
||||
- `data/term_index/daily/YYYY-MM-DD.json`
|
||||
- `data/term_index/term_stats.json`
|
||||
|
||||
如果需要词元清洗审阅输入,可额外生成:
|
||||
## 当前已验证
|
||||
|
||||
- `outputs/term_index/review/keyword-cleanup-bundle.json`
|
||||
- FreshRSS 未读拉取与已读回写可用
|
||||
- RSS-first 提取策略可用
|
||||
- 主日报 MCP 主链路可触发并写出正式产物
|
||||
- 状态查询与结果读取接口可用
|
||||
- stale state / artifacts 收敛逻辑已落地
|
||||
- `resume` 的 artifact-first 判定已落地
|
||||
- `start_resume_job -> poll -> result` 已做本地 synthetic 验证
|
||||
- 单篇总结 async job 可跑通
|
||||
- 关键词 review bundle 与建议写回脚本可用
|
||||
|
||||
如果需要在人工确认后把建议正式写入 watchlist / change log,可使用:
|
||||
## 当前主要限制
|
||||
|
||||
- `scripts/apply_term_suggestions.py`
|
||||
- 某些源 RSS 正文不足时会被直接跳过
|
||||
- 规则仍然偏保守,部分内容会落到 `review`
|
||||
- `paywall` 启发式对中文仍可能误判
|
||||
- 关键词治理还没有接入周期性调度
|
||||
- `digest-brief.json` 仍没有独立 MCP 读取工具
|
||||
- `resume` 目前的剩余主风险不再是恢复点判定,而是缺少真实生产环境的完整恢复验证
|
||||
|
||||
如果需要排障,可开启:
|
||||
## 当前建议
|
||||
|
||||
- `debug_artifacts=true`
|
||||
- 或脚本参数 `--debug-artifacts`
|
||||
|
||||
这样才会额外输出逐条中间文件。
|
||||
|
||||
## 当前验证状态
|
||||
|
||||
已经验证通过:
|
||||
|
||||
- FreshRSS 未读拉取成功
|
||||
- 已读回写成功
|
||||
- MCP 工具入口可直接触发完整链路
|
||||
- 微信公众号样本可直接使用 RSS 提供的 `summary` 内容提取,不再回源抓网页
|
||||
- 精简输出模式已实际跑通
|
||||
- 日报级词元统计已通过离线样例验证,确认别名、停用词、非 `drop` 过滤和 rerun 覆盖逻辑正常
|
||||
- `keyword-cleanup-review` skill 已通过 `quick_validate.py` 结构校验
|
||||
- review bundle 脚本已实际跑通
|
||||
- `apply_term_suggestions.py` 已通过 dry-run 与临时副本写回验证
|
||||
|
||||
## 当前已知限制
|
||||
|
||||
- 当前对 FreshRSS 条目采用 RSS-first 策略,不再回源抓原网页
|
||||
- 如果 RSS 中没有足够正文内容,该条会直接跳过,不会进入后续总结
|
||||
- 某些规则仍偏保守,部分内容可能落到 `review`
|
||||
- `paywall` 相关启发式仍可能误判中文文本
|
||||
- Webhook / 主动投递到 OpenClaw 外部接口尚未实现,当前是由 OpenClaw 通过 MCP 主动调用
|
||||
- 词元清洗 skill 当前已支持“bundle 构建 -> 建议审阅 -> 人工确认写回 watchlist/change_log”,但尚未接入周期性调度
|
||||
- 当前词元统计仍以前置 `OpenClawDeliveryPayload` 作为日报前代理输入,真实 `DailyDigest` 接入后还需切换上游
|
||||
|
||||
## 当前最建议的交接阅读顺序
|
||||
|
||||
1. `README.md`
|
||||
2. `docs/openclaw/openclaw-handoff.md`
|
||||
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||
5. `docs/design/daily-keyword-index-design.md`
|
||||
6. `skills/keyword-cleanup-review/SKILL.md`
|
||||
7. `TODO.md`
|
||||
|
||||
## 一句话结论
|
||||
|
||||
当前仓库已经从“提取 MCP 原型”演进到“可供 OpenClaw 调用的 FreshRSS -> OpenClaw payload 上游处理器”,并已补上第一阶段的日报级词元统计能力和词元清洗 skill 骨架;后续重点转向 skill 周期调度、知识库状态流转和 webhook 接线。
|
||||
- 把 `docs/openclaw/openclaw-orchestration-flow.md` 当成正式编排手册
|
||||
- 把 `docs/openclaw/openclaw-handoff.md` 当成接手总览
|
||||
- 把 `plans/README.md` 当成规划文档导航
|
||||
- 把 `docs/openclaw/archive/` 和 `docs/archive/` 当成历史资料,不要当当前事实源
|
||||
|
||||
@@ -0,0 +1,56 @@
|
||||
# 设计文档导航
|
||||
|
||||
## 使用原则
|
||||
|
||||
`docs/design/` 目录同时包含两类文档:
|
||||
|
||||
- 当前实现仍然有效的设计说明
|
||||
- 早期架构草案和背景设计
|
||||
|
||||
不要默认把这里所有文档都当成当前生产事实。
|
||||
当前生产事实仍以这些入口为准:
|
||||
|
||||
1. `README.md`
|
||||
2. `docs/current/context-reset-brief.md`
|
||||
3. `docs/openclaw/README.md`
|
||||
4. `docs/openclaw/openclaw-handoff.md`
|
||||
5. `docs/openclaw/openclaw-orchestration-flow.md`
|
||||
6. `TODO.md`
|
||||
|
||||
## 当前实现仍然有效
|
||||
|
||||
- `filter-rule-engine-design.md`
|
||||
- 规则过滤层的设计与职责边界
|
||||
- `filter-rule-engine-usage.md`
|
||||
- 规则引擎的使用说明
|
||||
- `daily-keyword-index-design.md`
|
||||
- 关键词索引与清洗治理设计
|
||||
- `summary-loop-explained.md`
|
||||
- LLM 摘要校验闭环说明
|
||||
- `markdown-sink-design.md`
|
||||
- Markdown sink 设计
|
||||
|
||||
## 当前仍有参考价值,但不是生产真相入口
|
||||
|
||||
- `summary-mcp-service-design.md`
|
||||
- 早期 MCP 服务设计草案,部分定位已被后续 workflow service 演进覆盖
|
||||
- `source-schema-design.md`
|
||||
- 更偏对象建模和来源抽象的背景设计
|
||||
- `summary-core-interface-design.md`
|
||||
- 更偏早期摘要内核接口抽象
|
||||
|
||||
## 建议阅读顺序
|
||||
|
||||
如果你是在理解当前实现:
|
||||
|
||||
1. `filter-rule-engine-design.md`
|
||||
2. `filter-rule-engine-usage.md`
|
||||
3. `daily-keyword-index-design.md`
|
||||
4. `summary-loop-explained.md`
|
||||
5. `markdown-sink-design.md`
|
||||
|
||||
如果你是在回看背景设计:
|
||||
|
||||
1. `summary-mcp-service-design.md`
|
||||
2. `source-schema-design.md`
|
||||
3. `summary-core-interface-design.md`
|
||||
@@ -326,8 +326,13 @@ LLM 可以帮助做清洗建议,但不适合直接维护主词元库。
|
||||
|
||||
skill 不直接修改配置文件,而是生成建议文件,例如:
|
||||
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`
|
||||
|
||||
其中建议语义为:
|
||||
|
||||
- JSON 是 review / apply 之间的唯一正式建议产物
|
||||
- Markdown 是人工临时审阅展示稿,不是长期真相来源
|
||||
|
||||
低复杂治理层建议补充三类输入:
|
||||
|
||||
@@ -403,8 +408,9 @@ skill 不直接修改配置文件,而是生成建议文件,例如:
|
||||
2. 程序更新 `daily/YYYY-MM-DD.json`
|
||||
3. 程序更新 `term_stats.json`
|
||||
4. 每周或人工触发一次词元清洗 skill
|
||||
5. skill 输出建议
|
||||
6. 人工确认后再更新配置文件
|
||||
5. skill 生成 suggestions JSON(正式建议产物)
|
||||
6. 如需要人工阅读,再临时生成 Markdown 展示稿
|
||||
7. 人工确认后再更新配置文件
|
||||
|
||||
## 14. 与规则引擎的关系
|
||||
|
||||
@@ -445,7 +451,8 @@ skill 不直接修改配置文件,而是生成建议文件,例如:
|
||||
再补治理层:
|
||||
|
||||
- 增加词元清洗 skill
|
||||
- 输出建议文件
|
||||
- 输出建议文件(以 JSON 为正式产物)
|
||||
- Markdown 仅作为按需生成的人工展示层
|
||||
- 人工确认后更新配置
|
||||
|
||||
### Phase 3
|
||||
@@ -459,3 +466,9 @@ skill 不直接修改配置文件,而是生成建议文件,例如:
|
||||
## 16. 一句话结论
|
||||
|
||||
这套设计选择“只统计日报中的 `keywords`,由程序维护轻量词元库,再由独立 skill 周期性做清洗建议”,目的是在控制数据规模的前提下,为规则配置和长期兴趣演化提供稳定、可审计、可扩展的基础设施。
|
||||
|
||||
补充的产物策略是:
|
||||
|
||||
- facts/state 长期保留
|
||||
- suggestions JSON 作为正式建议产物短期保留
|
||||
- review bundle 与 Markdown 展示稿降级为临时工作文件 / 展示层
|
||||
|
||||
@@ -0,0 +1,151 @@
|
||||
# 关键词清洗流程概述
|
||||
|
||||
> 2026-05-14 初版
|
||||
> 从"数据记录"到"人工确认落盘"的完整链路
|
||||
|
||||
---
|
||||
|
||||
## 整体数据流
|
||||
|
||||
```
|
||||
每日日报 pipeline
|
||||
│
|
||||
▼
|
||||
term_index/daily/YYYY-MM-DD.json ← 每天一篇候选文章的热词统计
|
||||
│
|
||||
▼
|
||||
term_index/term_stats.json ← 所有 daily 的汇总(1070 个词)
|
||||
│
|
||||
├──── build_review_bundle.py ← 打包为审查数据包
|
||||
│ │
|
||||
│ ▼
|
||||
│ review/keyword-cleanup-bundle.json
|
||||
│ │
|
||||
│ ▼
|
||||
│ generate_term_cleanup_suggestions.py
|
||||
│ │
|
||||
│ ▼
|
||||
│ review/term-cleanup-suggestions-YYYY-MM-DD.json ← 正式建议产物
|
||||
│ │
|
||||
│ ▼
|
||||
│ (可选) review/term-cleanup-suggestions-YYYY-MM-DD.md ← 展示稿
|
||||
│
|
||||
├──── 人工确认哪些建议 accept
|
||||
│
|
||||
▼
|
||||
apply_term_suggestions.py ← 写入配置
|
||||
│
|
||||
├── configs/filter_context.personal.json ← interest_keywords
|
||||
├── configs/term_aliases.json ← alias
|
||||
├── configs/term_stopwords.json ← stopword
|
||||
├── configs/term_watchlist.json ← watch
|
||||
└── configs/term_change_log.json ← 变更日志
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 各环节说明
|
||||
|
||||
### 阶段 1:数据记录(每日自动)
|
||||
|
||||
```bash
|
||||
# FreshRSS pipeline 跑完后自动产出
|
||||
data/term_index/daily/2026-05-14.json
|
||||
```
|
||||
|
||||
- 每天一篇,记录当天候选文章中出现的热词
|
||||
- 包含 term、total_count、days_seen 等信息
|
||||
- 目前累计 **41 天**,共 **1070 个独立词**
|
||||
|
||||
### 阶段 2:全量汇总(每日自动)
|
||||
|
||||
```bash
|
||||
data/term_index/term_stats.json
|
||||
```
|
||||
|
||||
- 从所有 daily 文件重建,会覆盖重跑
|
||||
- 按 total_count 排序,前 5 名:OpenClaw(30)、Claude Code(25)、AI Agent(17)、Anthropic(13)、MCP(13)
|
||||
|
||||
### 阶段 3:构建审查数据包(手动触发)
|
||||
|
||||
```bash
|
||||
python skills/keyword-cleanup-review/scripts/build_review_bundle.py \
|
||||
--days 365 \
|
||||
--top 100 \
|
||||
--output outputs/term_index/review/keyword-cleanup-bundle.json
|
||||
```
|
||||
|
||||
- 把 term_stats + 当前配置打成一包,方便后续处理
|
||||
- 输出:`review/keyword-cleanup-bundle.json`
|
||||
|
||||
### 阶段 4:生成建议(手动触发)
|
||||
|
||||
```bash
|
||||
python scripts/generate_term_cleanup_suggestions.py \
|
||||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
|
||||
--emit-markdown
|
||||
```
|
||||
|
||||
#### 当前产出能力
|
||||
|
||||
| 建议类型 | 状态 | 当前阈值 | 说明 |
|
||||
|---------|------|----------|------|
|
||||
| interest_keyword_suggestions | ✅ **已实现** | total≥3, days≥2 | 产出 20 条 |
|
||||
| watch_terms | ✅ **已实现** | total≤2, days≤2 | 本次 0 条 |
|
||||
| alias_suggestions | ❌ **硬编码为空** | policy 有阈值(total≥2, days≥2)但脚本未实现 | |
|
||||
| stopword_suggestions | ❌ **硬编码为空** | policy 有阈值(total≤2, days≤2)但脚本未实现 | |
|
||||
|
||||
**关键发现:** alias 和 stopword 不是"阈值太保守",是 **generate 脚本里压根没写对应的生成函数**。policy 文件里阈值已经配好了(alias: min_total=2/min_days=2,stopword: max_total=2/max_days=2),但脚本第 376-380 行直接硬编码为 `[]` 和 `0`。
|
||||
|
||||
### 阶段 5:人工确认(手动)
|
||||
|
||||
```
|
||||
OpenClaw 把建议列给你 → 你确认哪些 accept → 我执行 apply
|
||||
```
|
||||
|
||||
本次模式:
|
||||
- 高频(≥5次/5天以上)→ 强烈推荐 ✅
|
||||
- 中频(3-4次)→ 附带建议 ✅
|
||||
- 泛词 → 建议跳过 ❌
|
||||
|
||||
### 阶段 6:落盘配置(手动)
|
||||
|
||||
```bash
|
||||
python scripts/apply_term_suggestions.py \
|
||||
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
|
||||
--accept-interest 词1 词2 ...
|
||||
```
|
||||
|
||||
- dry-run 预览 → 确认后正式 apply
|
||||
- 写入 `configs/filter_context.personal.json`
|
||||
- 同步记录到 `term_change_log.json`
|
||||
- **不备份原始配置**(待优化)
|
||||
- **apply 后不自动清理 review 目录**(待优化)
|
||||
|
||||
### 阶段 7:维护清理(按需)
|
||||
|
||||
由 OpenClaw 侧 `reader-keyword-maintenance` skill 处理:
|
||||
- 删除旧 markdown 展示稿
|
||||
- 保留最近一份 bundle
|
||||
- 保守保留 suggestions JSON
|
||||
|
||||
---
|
||||
|
||||
## 当前配置资产
|
||||
|
||||
| 文件 | 内容 | 数据量 |
|
||||
|------|------|--------|
|
||||
| `filter_context.personal.json` | interest_keywords | 52 个 |
|
||||
| `term_aliases.json` | 别名映射 | 0 组(未启用) |
|
||||
| `term_stopwords.json` | 停用词 | 0 个(未启用) |
|
||||
| `term_watchlist.json` | 观察词 | 6 个 |
|
||||
| `term_change_log.json` | 所有变更记录 | 已记录 |
|
||||
|
||||
---
|
||||
|
||||
## 待优化项
|
||||
|
||||
1. **alias/stopword 建议生成为空** — generate 脚本硬编码缺实现,policy 已有阈值,需要补函数
|
||||
2. **apply 前无配置备份** — 建议 apply 前自动 cp 备份
|
||||
3. **apply 后无自动收尾** — 建议 apply 后自动删旧 markdown 和 bundle
|
||||
4. **alias 识别依赖规则而非 LLM** — 当前全靠统计阈值,无法做语义级判断(如中英文映射、缩写展开)。如果需要高级 alias 识别,可以用 LLM 生成候选,规则脚本做 apply
|
||||
@@ -1,5 +1,16 @@
|
||||
# Source Schema 设计草案
|
||||
|
||||
## 状态说明
|
||||
|
||||
本文件偏对象建模和来源抽象,主要用于解释早期 schema 设计思路。
|
||||
|
||||
它不是当前生产运行手册,也不是当前 workflow service 的唯一真相来源。
|
||||
如果你关注当前 OpenClaw 集成或运行状态,应优先看:
|
||||
|
||||
- `docs/current/context-reset-brief.md`
|
||||
- `docs/openclaw/README.md`
|
||||
- `docs/openclaw/openclaw-orchestration-flow.md`
|
||||
|
||||
## 1. 文档目的
|
||||
|
||||
本文档用于定义阅读流系统中的来源与内容对象模型,目标是把“来源分类”的讨论收敛成一套可执行的数据结构,供后续的抓取、摘要、过滤、入库和推送流程统一使用。
|
||||
|
||||
@@ -1,5 +1,17 @@
|
||||
# Summary Core Interface 设计草案
|
||||
|
||||
## 状态说明
|
||||
|
||||
本文件记录的是较早期的摘要内核接口抽象。
|
||||
|
||||
它更适合用于理解背景设计,不应直接当成当前生产接口契约。
|
||||
当前接口与编排真相请优先看:
|
||||
|
||||
- `README.md`
|
||||
- `docs/current/context-reset-brief.md`
|
||||
- `docs/openclaw/README.md`
|
||||
- `docs/openclaw/openclaw-handoff.md`
|
||||
|
||||
## 1. 文档目的
|
||||
|
||||
本文档用于定义 `summary-core` 的输入输出接口,目标是把“页面摘要能力”从概念讨论收敛成一套稳定、可复用、可封装的数据接口。
|
||||
|
||||
@@ -1,5 +1,18 @@
|
||||
# Content Extract MCP Service 设计草案
|
||||
|
||||
## 状态说明
|
||||
|
||||
本文件主要记录早期 “content extract MCP” 的设计抽象。
|
||||
|
||||
它仍有背景参考价值,但不是当前生产事实入口。
|
||||
当前生产能力已经演进为更完整的 workflow service,正式口径请优先看:
|
||||
|
||||
- `README.md`
|
||||
- `docs/current/context-reset-brief.md`
|
||||
- `docs/openclaw/README.md`
|
||||
- `docs/openclaw/openclaw-handoff.md`
|
||||
- `docs/openclaw/openclaw-orchestration-flow.md`
|
||||
|
||||
## 1. 文档目的
|
||||
|
||||
本文档用于定义当前仓库中已经落地的 MCP 服务设计,即“内容提取 MCP”。
|
||||
@@ -12,7 +25,7 @@
|
||||
- validator 与 LLM 摘要如何接在 MCP 之后
|
||||
- 当前 MVP 已完成到哪一层
|
||||
|
||||
这份文档描述的是当前真实实现,而不是早期“摘要 MCP”设想。
|
||||
这份文档主要记录当时实现阶段的设计取向,而不是当前生产阶段的唯一事实来源。
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -0,0 +1,22 @@
|
||||
# Notes 导航
|
||||
|
||||
`docs/notes/` 保存的是更早期、讨论型、背景型方案笔记。
|
||||
|
||||
这些文档的用途是:
|
||||
|
||||
- 理解项目最初的问题空间
|
||||
- 回看为什么会形成现在的对象分层和流程划分
|
||||
|
||||
这些文档不是当前生产事实来源。
|
||||
|
||||
当前如需判断“现在到底怎么跑”,优先看:
|
||||
|
||||
- `README.md`
|
||||
- `docs/current/context-reset-brief.md`
|
||||
- `docs/openclaw/README.md`
|
||||
- `docs/openclaw/openclaw-orchestration-flow.md`
|
||||
|
||||
当前 notes:
|
||||
|
||||
- `reading-pipeline-design-notes.md`
|
||||
- 早期阅读流方案讨论纪要
|
||||
@@ -0,0 +1,36 @@
|
||||
# OpenClaw 文档导航
|
||||
|
||||
## 当前有效文档
|
||||
|
||||
- `openclaw-handoff.md`
|
||||
- 面向接手者的总览文档
|
||||
- 说明 reader 的职责边界、MCP 工具面、环境变量和正式集成约束
|
||||
- `openclaw-orchestration-flow.md`
|
||||
- 面向 OpenClaw 编排层的正式运行手册
|
||||
- 说明启动、轮询、读结果、恢复和人工介入的标准动作
|
||||
- `openclaw-candidate-input-field-spec.md`
|
||||
- 单篇 `OpenClawCandidateInput` 字段规范
|
||||
- `openclaw-delivery-payload-spec.md`
|
||||
- 批量 `OpenClawDeliveryPayload` 字段规范
|
||||
- `article-candidate-daily-digest-schema.md`
|
||||
- 对象分层设计说明
|
||||
- 用于理解 `ArticleCandidateRecord` / `OpenClawCandidateInput` / `DailyDigest` 的关系
|
||||
|
||||
## 当前推荐阅读顺序
|
||||
|
||||
1. `openclaw-handoff.md`
|
||||
2. `openclaw-orchestration-flow.md`
|
||||
3. `openclaw-candidate-input-field-spec.md`
|
||||
4. `openclaw-delivery-payload-spec.md`
|
||||
5. `article-candidate-daily-digest-schema.md`
|
||||
|
||||
## 归档说明
|
||||
|
||||
`archive/` 下的文档保留历史决策、阶段总结和排障规划,但不再作为当前事实来源。
|
||||
|
||||
当前已归档:
|
||||
|
||||
- `archive/formalization-summary-2026-04-07.md`
|
||||
- `archive/openclaw-daily-digest-refactor.md`
|
||||
- `archive/digest-optimization-summary.md`
|
||||
- `archive/p1-status-reconciliation-plan-2026-04-14.md`
|
||||
@@ -0,0 +1,9 @@
|
||||
# OpenClaw 历史归档
|
||||
|
||||
本目录只保留阶段性总结、设计演进记录和排障计划。
|
||||
|
||||
使用原则:
|
||||
|
||||
- 需要了解“为什么会这样设计”时再看
|
||||
- 不要把这里的描述当成当前生产事实
|
||||
- 当前正式口径以 `docs/openclaw/README.md`、`docs/openclaw/openclaw-handoff.md`、`docs/openclaw/openclaw-orchestration-flow.md` 为准
|
||||
@@ -0,0 +1,163 @@
|
||||
# reader MCP workflow service formalization summary (2026-04-07)
|
||||
|
||||
## Overview
|
||||
|
||||
On 2026-04-07, the reader project was formally advanced from a script-first integration model into a reader-centric MCP workflow service model.
|
||||
|
||||
The key shift is:
|
||||
|
||||
- before: OpenClaw primarily relied on long CLI / exec flows and direct output-path stitching
|
||||
- now: reader exposes a formal workflow-oriented MCP surface with run-state, status queries, result reads, and minimal recovery
|
||||
|
||||
This document records the main outcomes and commits for the first formalization phase.
|
||||
|
||||
---
|
||||
|
||||
## Completed capability set
|
||||
|
||||
### 1. Run-state persistence
|
||||
|
||||
Commit:
|
||||
|
||||
- `72a6853` — `Add run-state persistence for FreshRSS pipeline`
|
||||
|
||||
Delivered:
|
||||
|
||||
- `run-state.json`
|
||||
- `RunState / StageState / ArtifactRecord`
|
||||
- stage-level state persistence for the FreshRSS pipeline
|
||||
|
||||
### 2. Architecture / implementation docs
|
||||
|
||||
Commit:
|
||||
|
||||
- `7563aa8` — `docs: add reader MCP architecture and implementation plan`
|
||||
|
||||
Delivered:
|
||||
|
||||
- architecture design
|
||||
- implementation plan
|
||||
- TODO-driven collaboration model
|
||||
|
||||
### 3. MCP run-status query tools
|
||||
|
||||
Commit:
|
||||
|
||||
- `d9173fb` — `feat: add MCP run status query tools`
|
||||
|
||||
Delivered:
|
||||
|
||||
- `get_run_status`
|
||||
- `list_runs`
|
||||
- `list_run_artifacts`
|
||||
|
||||
### 4. MCP result-read tools
|
||||
|
||||
Commit:
|
||||
|
||||
- `4a02894` — `Add MCP delivery payload and run report queries`
|
||||
|
||||
Delivered:
|
||||
|
||||
- `get_delivery_payload`
|
||||
- `get_run_report`
|
||||
|
||||
### 5. Minimal resume design
|
||||
|
||||
Commit:
|
||||
|
||||
- `d91cdbc` — `docs: narrow resume_run minimal recovery design`
|
||||
|
||||
Delivered:
|
||||
|
||||
- narrowed design for `resume_run`
|
||||
- explicit supported / unsupported recovery points
|
||||
|
||||
### 6. Minimal `resume_run`
|
||||
|
||||
Commit:
|
||||
|
||||
- `c622bc6` — `Implement minimal resume_run for freshrss runs`
|
||||
|
||||
Delivered:
|
||||
|
||||
- minimal `resume_run`
|
||||
- supports only freshrss runs with `run-state.json`
|
||||
- supports only recent resumable points
|
||||
- explicitly rejects `fetch_feed` and `extract_articles`
|
||||
|
||||
### 7. Formal handoff / workflow docs
|
||||
|
||||
Commits:
|
||||
|
||||
- `4f219ef` — `docs: formalize reader MCP workflow service handoff`
|
||||
- `2df0af5` — `docs: add openclaw orchestration flow for reader MCP`
|
||||
|
||||
Delivered:
|
||||
|
||||
- formal handoff aligned to actual implementation
|
||||
- OpenClaw orchestration runbook
|
||||
- explicit rule that OpenClaw should stop hand-stitching reader paths in the normal production flow
|
||||
|
||||
---
|
||||
|
||||
## Current formal MCP workflow surface
|
||||
|
||||
The current reader MCP workflow surface now includes:
|
||||
|
||||
- `run_freshrss_openclaw_pipeline`
|
||||
- `get_run_status`
|
||||
- `list_runs`
|
||||
- `list_run_artifacts`
|
||||
- `get_delivery_payload`
|
||||
- `get_run_report`
|
||||
- `resume_run` (minimal version)
|
||||
|
||||
---
|
||||
|
||||
## Current boundary
|
||||
|
||||
reader is now the upstream workflow engine for:
|
||||
|
||||
- FreshRSS pull
|
||||
- extraction
|
||||
- summary
|
||||
- filter
|
||||
- payload generation
|
||||
- run-state persistence
|
||||
- result read
|
||||
- minimal recovery
|
||||
|
||||
OpenClaw / skill remains responsible for:
|
||||
|
||||
- digest markdown generation
|
||||
- Hugo publishing
|
||||
- chat reporting
|
||||
- user confirmation
|
||||
- IMA orchestration
|
||||
|
||||
---
|
||||
|
||||
## Current limitations
|
||||
|
||||
The first formalization phase is complete, but some constraints remain:
|
||||
|
||||
- `resume_run` is still minimal and does not support arbitrary stage re-entry
|
||||
- historical runs without `run-state.json` are not formally recoverable
|
||||
- some very old runs may still require conservative artifact/path discovery
|
||||
- `rerun_stage` is not implemented
|
||||
- deeper runtime consolidation of `run_freshrss_openclaw_pipeline` can still be improved later
|
||||
|
||||
---
|
||||
|
||||
## Practical conclusion
|
||||
|
||||
The reader project should now be treated as a formal MCP workflow service rather than as a long-running CLI-first integration point.
|
||||
|
||||
For normal production orchestration:
|
||||
|
||||
- start via MCP
|
||||
- observe via MCP status tools
|
||||
- read results via MCP result tools
|
||||
- use `resume_run` only within the documented minimal recovery range
|
||||
- keep CLI for debug / fallback only
|
||||
@@ -0,0 +1,367 @@
|
||||
# Reader 日报链路 P1 状态收敛问题:规划与修复清单(2026-04-14)
|
||||
|
||||
## 背景
|
||||
|
||||
在 2026-04-14 的 reader 日报正式运行中,出现了以下现象:
|
||||
|
||||
- `openclaw-delivery-payload.json`、`digest-brief.json`、`run-report.json` 已真实落盘
|
||||
- 但 `get_freshrss_pipeline_job_status` / `get_run_status` 仍可能显示:
|
||||
- `running`
|
||||
- `failed`
|
||||
- 或 `current_stage=generate_summaries`
|
||||
- `resume_run` 在这种状态下可能直接超时
|
||||
|
||||
这说明当前 reader 的**状态层(job/run-state)**与**产物层(artifacts/report)**之间没有稳定收敛。
|
||||
|
||||
---
|
||||
|
||||
## 本次确认的核心结论
|
||||
|
||||
### 1. job status 与 run status 是两套独立状态系统
|
||||
|
||||
- **job 层状态**:`src/summary_mcp/runtime/freshrss_pipeline_jobs.py`
|
||||
- `start_freshrss_pipeline_job()`
|
||||
- `run_freshrss_pipeline_job()`
|
||||
- `get_freshrss_pipeline_job_status()`
|
||||
- 状态文件位于:`outputs/freshrss/pipeline_jobs/<job_id>/run-state.json`
|
||||
- 只有 4 个粗粒度 stage:
|
||||
- `prepare_job`
|
||||
- `load_input`
|
||||
- `run_pipeline`
|
||||
- `write_result`
|
||||
|
||||
- **run 层状态**:`src/summary_mcp/workflows/freshrss_pipeline.py`
|
||||
- `run_freshrss_pipeline()`
|
||||
- 由 `src/summary_mcp/runtime/query_service.py:get_run_status()` 查询
|
||||
- 状态文件位于:`outputs/freshrss/rerun/<run_dir>/run-state.json`
|
||||
- 包含 6 个细粒度 stage:
|
||||
- `fetch_feed`
|
||||
- `extract_articles`
|
||||
- `generate_summaries`
|
||||
- `apply_filters`
|
||||
- `build_delivery_payload`
|
||||
- `write_run_report`
|
||||
|
||||
**问题:** 两套状态没有统一收敛规则,用户可以同时看到两套不同口径的“当前进度”。
|
||||
|
||||
---
|
||||
|
||||
### 2. 查询层目前优先信 run-state,不会用 artifacts / run-report 纠偏
|
||||
|
||||
代码位置:`src/summary_mcp/runtime/query_service.py`
|
||||
|
||||
关键行为:
|
||||
- `_resolve_run_record()` 只要发现 `run-state.json` 存在,就优先使用 `RunStore.load(...)`
|
||||
- 即使 `run-report.json`、`delivery_payload`、`digest_brief` 已存在,也不会自动纠偏状态
|
||||
|
||||
**结果:**
|
||||
- 一旦 `run-state.json` 因中断、超时、外层 SIGTERM 或写回未完成而停留在旧值
|
||||
- `get_run_status()` 就会持续返回过期状态
|
||||
- 造成“产物已完成,但状态仍显示 running/failed/卡在 summary”的错觉
|
||||
|
||||
---
|
||||
|
||||
### 3. `generate_summaries` 假卡住,本质上更像 stale state,不像真实业务卡住
|
||||
|
||||
代码位置:`src/summary_mcp/workflows/freshrss_pipeline.py`
|
||||
|
||||
从执行顺序看:
|
||||
1. `start_stage(generate_summaries)`
|
||||
2. summary 循环
|
||||
3. `finish_stage(generate_summaries)`
|
||||
4. `start_stage(apply_filters)`
|
||||
5. `finish_stage(apply_filters)`
|
||||
6. `start_stage(build_delivery_payload)`
|
||||
7. 写 payload / digest brief
|
||||
8. `finish_stage(build_delivery_payload)`
|
||||
9. `start_stage(write_run_report)`
|
||||
10. 写 run-report
|
||||
11. `finish_stage(write_run_report)`
|
||||
12. `finish_run(...)`
|
||||
|
||||
**判断:**
|
||||
如果 payload / digest brief / run-report 都已经存在,那么“仍显示卡在 `generate_summaries`”更可能是:
|
||||
- `run-state.json` 没来得及写回最终状态
|
||||
- 或查询时读到了旧状态
|
||||
|
||||
而不是 summary 阶段真实没有跑过去。
|
||||
|
||||
---
|
||||
|
||||
### 4. `resume_run` 不是轻量恢复,而是同步继续跑工作流
|
||||
|
||||
代码位置:`src/summary_mcp/runtime/resume_service.py`
|
||||
|
||||
关键行为:
|
||||
- `resume_run()` 会根据 `resume_from_stage` 直接继续执行:
|
||||
- `_run_summary_stage(...)`
|
||||
- `_run_filter_stage(...)`
|
||||
- `_run_delivery_stage(...)`
|
||||
- `_run_report_stage(...)`
|
||||
|
||||
这意味着它不是“修状态”的工具,而是“同步继续跑剩余工作流”的工具。
|
||||
|
||||
**问题:**
|
||||
- 如果 stale state 把 `resume_from_stage` 定在 `generate_summaries`
|
||||
- 那么 `resume_run` 会从一个过早阶段重新跑
|
||||
- 在 MCP 包装层下非常容易超时
|
||||
|
||||
---
|
||||
|
||||
## 问题分类
|
||||
|
||||
### A. 真实 bug
|
||||
|
||||
1. **查询层过度信任 stale `run-state.json`**
|
||||
- 文件:`src/summary_mcp/runtime/query_service.py`
|
||||
- 影响:产物已完成但状态仍错误
|
||||
|
||||
2. **`resume_run` 过度依赖 stale `current_stage` / recovery 信息**
|
||||
- 文件:`src/summary_mcp/runtime/resume_service.py`
|
||||
- 影响:从过早阶段重跑,放大 timeout 风险
|
||||
|
||||
### B. 状态设计缺陷
|
||||
|
||||
3. **job 层与 run 层两套状态源没有统一收敛规则**
|
||||
- 文件:`src/summary_mcp/runtime/freshrss_pipeline_jobs.py`
|
||||
- 文件:`src/summary_mcp/runtime/query_service.py`
|
||||
- 影响:用户看到两个互相打架的状态解释
|
||||
|
||||
4. **状态系统完全依赖显式写回,不会按产物反推修正**
|
||||
- 文件:`src/summary_mcp/runtime/run_store.py`
|
||||
- 影响:一旦中断,状态比产物更容易脏
|
||||
|
||||
### C. 调用层误判
|
||||
|
||||
5. **把 `resume_run` 当成轻量恢复接口使用**
|
||||
- 实际上它更接近“同步恢复执行器”
|
||||
- 影响:在长链路场景下超时是高概率事件
|
||||
|
||||
---
|
||||
|
||||
## 修复目标
|
||||
|
||||
## 当前落地状态(回填)
|
||||
|
||||
- [x] Phase 1 已落地:`get_run_status()` 会基于 `run-report.json` 与关键产物做终态收敛,并暴露 `status_source` / `state_conflict`
|
||||
- [x] Phase 2 已落地第一阶段:`resume_run()` 会拒绝对已有终态 `run-report.json` 的 run 继续恢复
|
||||
- [x] Phase 2 已继续增强:恢复起点现在会优先根据 artifacts 重算,而不是直接盲信 `run-state.recovery.resume_from_stage`
|
||||
- [x] 新增 `inspect_resume_plan(run_id)` 作为恢复前置判定接口,避免调用方用 `resume_run` 探路
|
||||
- [x] Phase 2 已补齐生产恢复 artifacts:正式 run 会稳定写出 `summary/summary-batch.json` 与 `candidates/candidate-batch.json`,`resume_run` / `inspect_resume_plan` 会优先使用它们,而不是依赖 debug per-item 文件
|
||||
- [x] Phase 3 已落地:job 状态与结果读取会基于 linked run 做收敛,避免 outer job stale state 卡住编排
|
||||
|
||||
### 一级目标(必须达成)
|
||||
|
||||
1. 当 `run-report.json` / `delivery_payload` / `digest_brief` 已存在时,`get_run_status()` 不应继续盲目展示明显过期的 stage 状态;对调用方暴露的 `status` 必须直接收敛为可用终态,而不是只附加 hint
|
||||
2. 当状态层与产物层冲突时,查询结果必须显式标注“状态冲突 / stale state”
|
||||
3. `resume_run()` 在恢复前应优先基于现有 artifacts 判断真实可恢复起点,避免从过早阶段重跑
|
||||
|
||||
### 二级目标(建议达成)
|
||||
|
||||
4. job 层状态结果中增加对 linked run 的补充解释,避免“job running 但 run 产物已齐”这种情况毫无说明
|
||||
5. 为后续编排层提供明确可消费的“状态可信度/冲突提示”字段
|
||||
|
||||
---
|
||||
|
||||
## 最小修复方案
|
||||
|
||||
### Phase 1|先修 run 查询层(优先级最高)
|
||||
|
||||
#### 目标
|
||||
让 `get_run_status()` 至少能正确识别:
|
||||
- run-state 是旧的
|
||||
- 但关键产物已经齐了
|
||||
|
||||
#### 建议改动点
|
||||
文件:`src/summary_mcp/runtime/query_service.py`
|
||||
|
||||
#### 建议动作
|
||||
- [x] 在 `_resolve_run_record()` 或 `_build_status_response()` 中增加“关键产物存在性检查”
|
||||
- `run-report.json`
|
||||
- `candidates/openclaw-delivery-payload.json`
|
||||
- `candidates/digest-brief.json`
|
||||
- [x] 如果 `run-state.current_stage` 仍停留在早期阶段,但关键产物已齐:
|
||||
- 不要继续原样输出为可信最终态
|
||||
- 应直接把对外 `status` / `current_stage` / `recovery` 收敛成终态语义
|
||||
- 同时新增解释字段,例如:
|
||||
- `state_conflict: true`
|
||||
- `state_conflict_reason: "run_state indicates generate_summaries but run-report.json already proves the workflow reached a terminal state"`
|
||||
- `status_source: "run_report_reconciliation"`
|
||||
- [x] 保留 `state_source=run_state`,但增加 `status_source` / `state_quality` / `state_conflict` 之类解释字段
|
||||
|
||||
#### 预期收益
|
||||
- OpenClaw 继续按 `status` 分支时也不会卡住
|
||||
- 第一时间减少“明明产物齐了却还像没跑完”的误判
|
||||
- 不需要立刻动 workflow 主链路
|
||||
|
||||
---
|
||||
|
||||
### Phase 2|修 `resume_run` 的恢复起点判断
|
||||
|
||||
#### 目标
|
||||
避免 stale state 让恢复逻辑从 `generate_summaries` 这类过早阶段重跑。
|
||||
|
||||
#### 建议改动点
|
||||
文件:`src/summary_mcp/runtime/resume_service.py`
|
||||
|
||||
#### 建议动作
|
||||
- [x] 在 `_resolve_resume_from_stage()` 之前/之后加入真实 artifacts 检查
|
||||
- [x] 如果以下文件已存在:
|
||||
- `openclaw-delivery-payload.json`
|
||||
- `digest-brief.json`
|
||||
- `run-report.json`
|
||||
则不要再从 `generate_summaries` 或 `apply_filters` 起跑
|
||||
- [x] 为 `resume_run()` 增加“恢复起点是基于 artifacts 重算还是基于 state 推断”的返回说明
|
||||
- [x] 必要时增加更保守逻辑:
|
||||
- `run-report.json` 已存在时,默认拒绝继续 resume,并提示“产物已完成,请先检查状态一致性”
|
||||
- 补充:默认生产模式下,主链路会稳定写出 `summary-batch` / `candidate-batch`,恢复逻辑优先消费这两个 batch artifacts;若它们缺失或不稳定,才回退到更早的安全 stage 或直接拒绝恢复
|
||||
- 补充:调用方可先走 `inspect_resume_plan`,只有 `recommended_action=resume` 时再调用 `resume_run`
|
||||
|
||||
#### 预期收益
|
||||
- 降低无意义重跑和 timeout 风险
|
||||
- 让 `resume_run` 更接近真正的恢复工具,而不是误重跑工具
|
||||
|
||||
---
|
||||
|
||||
### Phase 3|补 job/run 双状态解释层
|
||||
|
||||
#### 目标
|
||||
让 `get_freshrss_pipeline_job_status()` 和 `get_run_status()` 的关系对调用方更可理解。
|
||||
|
||||
#### 建议改动点
|
||||
文件:`src/summary_mcp/runtime/freshrss_pipeline_jobs.py`
|
||||
|
||||
#### 建议动作
|
||||
- [x] 在 `get_freshrss_pipeline_job_status()` 中,读取 linked run 的关键产物存在性(轻量即可)
|
||||
- [x] 若 job 仍显示 `run_pipeline`,但 linked run 已有 report/payload/digest 产物:
|
||||
- 不仅增加解释字段,还应直接把 job 对外 `status` 收敛为终态,避免外层永远轮询
|
||||
- 例如:
|
||||
- `status_source: "linked_run_reconciliation"`
|
||||
- `status_note: "linked run artifacts are complete; the job can be treated as completed"`
|
||||
- [x] 若 `result.json` 缺失,但 linked run 已有 `run-report.json` 与 delivery 产物:
|
||||
- `get_freshrss_pipeline_job_result()` 应能基于 linked run 产物合成最小结果,至少稳定返回 `run_id`
|
||||
- [x] 明确文档:job status 是外层异步任务态,不等于内部 workflow 细粒度状态
|
||||
|
||||
#### 预期收益
|
||||
- 减少“job running / run finished”口径冲突带来的误解
|
||||
- 避免 OpenClaw 因 outer job stale state 卡死在轮询和 result 读取前
|
||||
|
||||
---
|
||||
|
||||
### Phase 4|把 `resume_run` 改成异步恢复 job
|
||||
|
||||
#### 目标
|
||||
解决当前剩余的核心问题:`resume_run` 虽然恢复判定已经安全,但执行模型仍是同步 MCP 调用,长链路恢复时依然可能超时,导致 OpenClaw 编排层“看起来像又卡住了”。
|
||||
|
||||
#### 建议改动点
|
||||
文件:
|
||||
- `src/summary_mcp/runtime/resume_jobs.py`(新)
|
||||
- `scripts/run_resume_job.py`(新)
|
||||
- `src/summary_mcp/server.py`
|
||||
- `src/summary_mcp/runtime/__init__.py`
|
||||
- `src/summary_mcp/runtime/resume_service.py`
|
||||
|
||||
#### 建议动作
|
||||
- [x] 新增最小异步恢复接口:
|
||||
- `start_resume_job(run_id)`
|
||||
- `get_resume_job_status(job_id)`
|
||||
- `get_resume_job_result(job_id)`
|
||||
- [x] job 目录固定落到:
|
||||
- `outputs/freshrss/resume_jobs/<job_id>/`
|
||||
- [x] 最少产物约定:
|
||||
- `run-state.json`
|
||||
- `input.json`
|
||||
- `result.json`(成功时)
|
||||
- `job-report.json`
|
||||
- [x] `start_resume_job` 内部先调用 `inspect_resume_plan`
|
||||
- 只有 `recommended_action=resume` 才允许真正启动
|
||||
- `read_terminal_result` / `start_new_run` 要直接在 job 输入校验阶段返回,不进入执行器
|
||||
- [x] 后台执行时复用现有 `_resume_freshrss_run(...)`
|
||||
- 不重写恢复业务逻辑
|
||||
- 只把同步入口拆成异步 job 外壳
|
||||
- [x] `resume_run(run_id)` 保留,但降级为 debug / fallback
|
||||
- 文档中明确:OpenClaw 编排默认应走 resume async job,而不是同步 `resume_run`
|
||||
- [x] job result 里至少稳定返回:
|
||||
- `run_id`
|
||||
- `resume_from_stage`
|
||||
- `status`
|
||||
- `result_source`
|
||||
- `delivery_output` / `report_output`(若存在)
|
||||
|
||||
#### 预期收益
|
||||
- 彻底切掉恢复阶段的 MCP 同步超时风险
|
||||
- 让 OpenClaw 对“启动恢复 / 轮询恢复 / 读取恢复结果”的控制面与主 pipeline async job 保持一致
|
||||
- 把“恢复判定”与“恢复执行”分层,减少误调用和卡住错觉
|
||||
|
||||
---
|
||||
|
||||
## 不建议现在就做的事
|
||||
|
||||
- [ ] **不要先做自动 fallback 修状态**
|
||||
- 例如:看到 artifacts 齐了就直接把 run-state 强行改成 success
|
||||
- 原因:这会掩盖真正的状态写回问题
|
||||
|
||||
- [ ] **不要先大改 workflow 主链路**
|
||||
- 当前更像查询层与恢复层的状态解释缺陷
|
||||
- 先修读取与恢复判断,收益更大、风险更低
|
||||
|
||||
---
|
||||
|
||||
## 建议执行顺序
|
||||
|
||||
1. **先改 `query_service.py`**
|
||||
- 让 `get_run_status()` 能暴露 stale state / artifact conflict
|
||||
2. **再改 `resume_service.py`**
|
||||
- 避免从错误阶段重跑
|
||||
3. **最后看 `freshrss_pipeline_jobs.py`**
|
||||
- 给 job status 加 linked run 补充说明
|
||||
4. **收尾改 `resume async job`**
|
||||
- 让恢复执行也走正式异步控制面,避免同步恢复再把编排卡住
|
||||
|
||||
---
|
||||
|
||||
## 验收标准
|
||||
|
||||
### 验收 1:状态冲突识别
|
||||
构造一个场景:
|
||||
- `run-state.json` 留在 `generate_summaries`
|
||||
- 但 payload / digest brief / run-report 已存在
|
||||
|
||||
期望:
|
||||
- `get_run_status()` 不再只回“卡在 generate_summaries”
|
||||
- 会显式返回冲突提示字段
|
||||
|
||||
### 验收 2:恢复起点修正
|
||||
构造一个场景:
|
||||
- `run-state` 指向 `generate_summaries`
|
||||
- 但 `delivery_payload` / `run-report` 已存在
|
||||
|
||||
期望:
|
||||
- `resume_run()` 不应再从 summary 阶段重跑
|
||||
- 至少应拒绝恢复并提示“产物已完成,优先检查状态一致性”
|
||||
|
||||
### 验收 3:job/run 双层说明
|
||||
构造一个场景:
|
||||
- job status 仍在 `run_pipeline`
|
||||
- linked run 已有关键产物
|
||||
|
||||
期望:
|
||||
- `get_freshrss_pipeline_job_status()` 能返回补充说明,不再只有生硬 running
|
||||
|
||||
### 验收 4:恢复执行不再阻塞编排
|
||||
构造一个场景:
|
||||
- run 可恢复
|
||||
- 恢复点为 `generate_summaries` 或 `apply_filters`
|
||||
- 恢复执行耗时超过单次 MCP 同步窗口
|
||||
|
||||
期望:
|
||||
- OpenClaw 调用的是 `start_resume_job(...)`,而不是同步 `resume_run(...)`
|
||||
- `get_resume_job_status(job_id)` 可稳定轮询到终态
|
||||
- `get_resume_job_result(job_id)` 至少稳定返回 `run_id`、`resume_from_stage` 与最终产物引用
|
||||
- 即使恢复失败,也能在 job-report / result 中看清失败点,而不是只表现为调用超时
|
||||
|
||||
---
|
||||
|
||||
## 备注
|
||||
|
||||
截至 2026-04-14,本文件中的 Phase 1 / 2 / 3 / 4 已完成主要落地;当前 `resume` 链路已经从“状态收敛 + 安全恢复点判定”进一步补齐到“正式异步恢复执行”。
|
||||
+189
-243
@@ -1,64 +1,196 @@
|
||||
# OpenClaw Handoff
|
||||
|
||||
## Role
|
||||
|
||||
This file is the integration overview for OpenClaw maintainers.
|
||||
|
||||
Use it for:
|
||||
|
||||
- reader capability boundary
|
||||
- production MCP entrypoints
|
||||
- environment requirements
|
||||
- integration rules and limitations
|
||||
|
||||
Do not use it as the step-by-step runbook.
|
||||
For formal orchestration, read `docs/openclaw/openclaw-orchestration-flow.md`.
|
||||
For field contracts, read:
|
||||
|
||||
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||
- `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||
|
||||
Historical plans and incident documents live under `docs/openclaw/archive/`.
|
||||
|
||||
## Purpose
|
||||
|
||||
This repository provides a FreshRSS-first reading pipeline for OpenClaw:
|
||||
reader is the upstream FreshRSS processing service for OpenClaw:
|
||||
|
||||
`FreshRSS unread items -> RSS content extraction -> LLM summary -> rule engine -> OpenClaw delivery payload`
|
||||
|
||||
OpenClaw should treat this repository as an MCP-backed upstream content processor.
|
||||
|
||||
This repository is responsible only for upstream reading-pipeline work:
|
||||
reader is responsible for:
|
||||
|
||||
- FreshRSS pull
|
||||
- content extraction
|
||||
- LLM summary generation/validation
|
||||
- LLM summary generation and validation
|
||||
- rule-based filtering
|
||||
- OpenClaw delivery payload generation
|
||||
- selected-article summary capability based on existing extracted text
|
||||
- run-state persistence, status lookup, result lookup, and minimal resume for the FreshRSS workflow
|
||||
- run-state persistence and run/result lookup
|
||||
- async resume control for the FreshRSS workflow
|
||||
- async selected-article summary generation from existing extracted files
|
||||
|
||||
This repository should **not** take over downstream orchestration responsibilities that belong to OpenClaw / skills, such as:
|
||||
reader is not responsible for:
|
||||
|
||||
- Hugo publishing
|
||||
- chat reporting
|
||||
- user confirmation handling
|
||||
- IMA upload orchestration
|
||||
|
||||
## Production Entrypoint
|
||||
## Production Surface
|
||||
|
||||
reader 当前正式工作流服务入口是 MCP tool:
|
||||
Current MCP tool count: 21.
|
||||
|
||||
- `run_freshrss_openclaw_pipeline`
|
||||
Main daily workflow:
|
||||
|
||||
It starts the only formally supported workflow today: `freshrss_daily_digest`.
|
||||
|
||||
OpenClaw should treat the returned `run_id` as the only stable handle for follow-up reads. Do not hand-build `outputs/freshrss/rerun/...` paths in OpenClaw.
|
||||
|
||||
## Supported MCP Tools
|
||||
|
||||
Current MCP tools: 11 total.
|
||||
|
||||
Workflow service tools:
|
||||
|
||||
- `run_freshrss_openclaw_pipeline`
|
||||
- `start_freshrss_pipeline_job`
|
||||
- `get_freshrss_pipeline_job_status`
|
||||
- `get_freshrss_pipeline_job_result`
|
||||
- `get_run_status`
|
||||
- `list_runs`
|
||||
- `list_run_artifacts`
|
||||
- `get_delivery_payload`
|
||||
- `get_run_report`
|
||||
|
||||
Resume workflow:
|
||||
|
||||
- `inspect_resume_plan`
|
||||
- `start_resume_job`
|
||||
- `get_resume_job_status`
|
||||
- `get_resume_job_result`
|
||||
- `resume_run`
|
||||
|
||||
Single-step / debug tools:
|
||||
Selected-article summary workflow:
|
||||
|
||||
- `start_article_summary_job`
|
||||
- `get_article_summary_job_status`
|
||||
- `get_article_summary_job_result`
|
||||
- `generate_article_summaries`
|
||||
|
||||
Debug / single-step tools:
|
||||
|
||||
- `run_freshrss_openclaw_pipeline`
|
||||
- `extract_url_content`
|
||||
- `extract_item_content`
|
||||
- `filter_summary_result`
|
||||
- `generate_article_summaries`
|
||||
|
||||
## Required Environment Variables
|
||||
Production rules:
|
||||
|
||||
The MCP server process must have these variables available:
|
||||
- main production start path is `start_freshrss_pipeline_job`
|
||||
- production resume path is `inspect_resume_plan -> start_resume_job -> get_resume_job_status -> get_resume_job_result`
|
||||
- `run_freshrss_openclaw_pipeline` is sync debug / fallback only
|
||||
- `resume_run` is sync debug / fallback only
|
||||
- `generate_article_summaries` is sync debug / fallback only
|
||||
|
||||
## Production Contract
|
||||
|
||||
OpenClaw should treat the returned `run_id` from `get_freshrss_pipeline_job_result` as the only stable handle for follow-up reads.
|
||||
|
||||
OpenClaw should not hand-build these paths:
|
||||
|
||||
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
|
||||
|
||||
If filesystem access is needed for debugging, only consume paths returned by MCP:
|
||||
|
||||
- `output_dir`
|
||||
- `artifact.path`
|
||||
- `delivery_output`
|
||||
- `report_output`
|
||||
|
||||
Top-level `status` is the only status field callers should branch on.
|
||||
`status_source` and `state_conflict` are explanatory fields for reconciled status.
|
||||
|
||||
## Minimal Production Sequence
|
||||
|
||||
Daily workflow:
|
||||
|
||||
1. Call `start_freshrss_pipeline_job`
|
||||
2. Poll `get_freshrss_pipeline_job_status`
|
||||
3. On success, read `get_freshrss_pipeline_job_result`
|
||||
4. Persist the returned `run_id`
|
||||
5. Use `get_run_status`, `get_delivery_payload`, and `get_run_report` for follow-up reads
|
||||
|
||||
Resume workflow:
|
||||
|
||||
1. Call `inspect_resume_plan(run_id)`
|
||||
2. Only if `can_resume=true` and `recommended_action=resume`, call `start_resume_job`
|
||||
3. Poll `get_resume_job_status`
|
||||
4. Read `get_resume_job_result`
|
||||
|
||||
Selected-article summary workflow:
|
||||
|
||||
1. Call `start_article_summary_job` with a real extracted file path and non-empty `selected_ids`
|
||||
2. Poll `get_article_summary_job_status`
|
||||
3. Read `get_article_summary_job_result`
|
||||
|
||||
## Capability Boundary
|
||||
|
||||
Formal workflow boundary:
|
||||
|
||||
- only workflow `freshrss_daily_digest`
|
||||
- every current FreshRSS run writes `run-state.json`
|
||||
- `get_run_status` / `list_runs` / `list_run_artifacts` can still infer basic state for older runs without `run-state.json`
|
||||
- resume requires a valid `run-state.json`; inferred historical runs are not resumable
|
||||
|
||||
Resume boundary:
|
||||
|
||||
- resume in place on the original `run_id`
|
||||
- supported resume points:
|
||||
- `generate_summaries`
|
||||
- `apply_filters`
|
||||
- `build_delivery_payload`
|
||||
- `write_run_report`
|
||||
- unsupported resume points:
|
||||
- `fetch_feed`
|
||||
- `extract_articles`
|
||||
- production resume prefers:
|
||||
- `summary/summary-batch.json`
|
||||
- `candidates/candidate-batch.json`
|
||||
- if required artifacts are missing, recovery should return non-resumable instead of silently falling back
|
||||
|
||||
Selected-article summary boundary:
|
||||
|
||||
- uses existing extracted files as input
|
||||
- should not re-fetch original URLs
|
||||
|
||||
## Output Expectations
|
||||
|
||||
Main daily pipeline core artifacts:
|
||||
|
||||
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/raw/freshrss.raw.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/summary/summary-batch.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/candidates/candidate-batch.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/candidates/digest-brief.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json`
|
||||
|
||||
Async job state directories:
|
||||
|
||||
- main pipeline job: `outputs/freshrss/pipeline_jobs/<job_id>/`
|
||||
- resume job: `outputs/freshrss/resume_jobs/<job_id>/`
|
||||
- article-summary job: `outputs/freshrss/article_summary_jobs/<job_id>/`
|
||||
|
||||
Each job directory minimally contains:
|
||||
|
||||
- `run-state.json`
|
||||
- `input.json`
|
||||
- `result.json` on success
|
||||
- `job-report.json`
|
||||
|
||||
## Environment And Startup
|
||||
|
||||
Required environment variables:
|
||||
|
||||
- `FRESHRSS_API_BASE_URL`
|
||||
- `FRESHRSS_USERNAME`
|
||||
@@ -67,52 +199,13 @@ The MCP server process must have these variables available:
|
||||
- `LLM_API_KEY`
|
||||
- `LLM_MODEL`
|
||||
|
||||
Example:
|
||||
|
||||
```powershell
|
||||
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||||
set FRESHRSS_USERNAME=osiman
|
||||
set FRESHRSS_API_PASSWORD=your-api-password
|
||||
set LLM_API_URL=https://api.deepseek.com
|
||||
set LLM_API_KEY=your-llm-api-key
|
||||
set LLM_MODEL=deepseek-chat
|
||||
```
|
||||
|
||||
## Server Startup
|
||||
|
||||
Install dependencies:
|
||||
Startup:
|
||||
|
||||
```bash
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
Start the MCP server:
|
||||
|
||||
```bash
|
||||
summary-mcp
|
||||
```
|
||||
|
||||
## Recommended MCP Workflow
|
||||
|
||||
Recommended production path:
|
||||
|
||||
1. Call `run_freshrss_openclaw_pipeline` and persist the returned `run_id`
|
||||
2. Use `get_run_status(run_id)` as the authoritative run-state read for status, stage, artifacts, and recovery
|
||||
3. Use `list_runs(...)` when OpenClaw needs recent-run discovery or high-level inspection
|
||||
4. Use `list_run_artifacts(run_id)` when OpenClaw needs to inspect what this run actually produced
|
||||
5. Use `get_delivery_payload(run_id)` and `get_run_report(run_id)` as the formal result-reading APIs
|
||||
6. Use `resume_run(run_id)` only when the run falls inside the minimal supported resume scope
|
||||
|
||||
OpenClaw should not directly derive or hardcode:
|
||||
|
||||
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
|
||||
|
||||
If filesystem access is needed for debugging, consume only paths returned by MCP such as `output_dir`, `artifact.path`, `delivery_output`, or `report_output`.
|
||||
|
||||
## Recommended MCP Call
|
||||
|
||||
Recommended production start call:
|
||||
|
||||
```json
|
||||
@@ -126,198 +219,51 @@ Recommended production start call:
|
||||
}
|
||||
```
|
||||
|
||||
Recommended semantics:
|
||||
## Data And Content Policy
|
||||
|
||||
- Use `mark_read=true` for normal production runs.
|
||||
- Use `mark_read=false` only for debug, test, or validation runs.
|
||||
- Keep `debug_artifacts=false` for routine production runs.
|
||||
- Set `debug_artifacts=true` only when troubleshooting a bad batch.
|
||||
- Treat the returned `run_id` as the stable identifier for all follow-up MCP reads.
|
||||
- If no real `openclaw-delivery-payload.json` was produced, OpenClaw should stop instead of generating a digest from placeholders or examples.
|
||||
|
||||
## Formal Capability Boundary
|
||||
|
||||
reader 当前正式 MCP workflow service 的边界如下:
|
||||
|
||||
- formal workflow: only `freshrss_daily_digest`
|
||||
- run truth: every FreshRSS run writes `run-state.json`
|
||||
- state query tools: `get_run_status`, `list_runs`, `list_run_artifacts`
|
||||
- result read tools: `get_delivery_payload`, `get_run_report`
|
||||
- `digest-brief.json` is generated and registered as an artifact, but there is no standalone `get_digest_brief` tool yet
|
||||
- `run_freshrss_openclaw_pipeline` and `resume_run` are synchronous MCP calls today; there is no background queue / worker model yet
|
||||
- `generate_article_summaries` is supported, but it is outside the formal `resume_run` scope and not part of the FreshRSS workflow-state model
|
||||
|
||||
Historical compatibility note:
|
||||
|
||||
- `get_run_status` / `list_runs` / `list_run_artifacts` can still infer basic state for older run directories without `run-state.json`
|
||||
- `resume_run` does **not** support those inferred historical runs; it requires a valid `run-state.json`
|
||||
|
||||
## What The Tool Returns
|
||||
|
||||
Primary return fields from `run_freshrss_openclaw_pipeline`:
|
||||
|
||||
- `run_id`
|
||||
- `output_dir`
|
||||
- `raw_output`
|
||||
- `delivery_output`
|
||||
- `report_output`
|
||||
- `digest_brief_output`
|
||||
- `pulled_count`
|
||||
- `delivered_count`
|
||||
- `marked_read_count`
|
||||
- `status_counts`
|
||||
- `delivery_payload`
|
||||
- `keyword_index`
|
||||
|
||||
Optional:
|
||||
|
||||
- `items`
|
||||
- returned only when `include_item_reports=true`
|
||||
|
||||
Follow-up structured reads should use MCP tools rather than re-reading these files directly.
|
||||
|
||||
## Minimal Output Files
|
||||
|
||||
By default the pipeline writes these core artifacts:
|
||||
|
||||
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/raw/freshrss.raw.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/candidates/digest-brief.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json` (one per item)
|
||||
|
||||
It also updates local runtime keyword data:
|
||||
|
||||
- `data/term_index/daily/YYYY-MM-DD.json`
|
||||
- `data/term_index/term_stats.json`
|
||||
|
||||
Per-item extracted files live under `extracted/` and are always written.
|
||||
If `debug_artifacts=true`, the pipeline additionally writes normalized items, summaries, filter decisions, candidate records, and candidate inputs.
|
||||
The main pipeline does not emit a batch-level `freshrss.extracted.json` file by default.
|
||||
|
||||
## `resume_run` Minimal Scope
|
||||
|
||||
`resume_run` currently supports only the minimum resume contract:
|
||||
|
||||
- only runs with a valid `run-state.json`
|
||||
- only workflow `freshrss_daily_digest`
|
||||
- resume in place on the original `run_id`
|
||||
- supported resume points: `generate_summaries`, `apply_filters`, `build_delivery_payload`, `write_run_report`
|
||||
- unsupported resume points: `fetch_feed`, `extract_articles`
|
||||
- if required artifacts are missing, the tool returns a non-resumable response instead of silently falling back to an earlier stage
|
||||
|
||||
Artifact expectations by resume point:
|
||||
|
||||
- `generate_summaries`: requires `raw/freshrss.raw.json` and `extracted/`
|
||||
- `apply_filters`: requires the above plus per-item summary outputs
|
||||
- `build_delivery_payload`: requires per-item candidate inputs consistent with filter-stage output
|
||||
- `write_run_report`: requires `candidates/openclaw-delivery-payload.json`; if `mark_read=true`, raw input must still be present
|
||||
|
||||
## Payload Specs
|
||||
|
||||
OpenClaw payload field specs live here:
|
||||
|
||||
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||
- `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||
|
||||
## Read-State Semantics
|
||||
|
||||
The pipeline reads from FreshRSS unread items by default.
|
||||
|
||||
If `mark_read=true`:
|
||||
|
||||
- items are marked as read only after the final `openclaw-delivery-payload.json` has been written successfully
|
||||
- only successfully delivered items are marked as read
|
||||
- failed or skipped items remain unread
|
||||
|
||||
## FreshRSS Content Policy
|
||||
|
||||
For FreshRSS items, the pipeline is RSS-first and does not re-crawl webpages.
|
||||
|
||||
Behavior:
|
||||
FreshRSS processing is RSS-first:
|
||||
|
||||
- use `item.raw_content` first
|
||||
- if missing, use `item.raw_summary`
|
||||
- if neither contains usable content, skip the item
|
||||
- do not fetch the original webpage again for FreshRSS items
|
||||
|
||||
This is intentional.
|
||||
Read-state policy:
|
||||
|
||||
## Keyword Cleanup Governance
|
||||
- items are marked read only after successful delivery payload write
|
||||
- only successfully delivered items are marked read
|
||||
|
||||
This repository also includes a lightweight keyword-governance flow for downstream review.
|
||||
Downstream boundary:
|
||||
|
||||
Current pieces:
|
||||
|
||||
- runtime keyword stats
|
||||
- `data/term_index/daily/YYYY-MM-DD.json`
|
||||
- `data/term_index/term_stats.json`
|
||||
- governance config
|
||||
- `configs/term_cleanup_policy.json`
|
||||
- `configs/term_watchlist.json`
|
||||
- `configs/term_change_log.json`
|
||||
- review bundle builder
|
||||
- `skills/keyword-cleanup-review/scripts/build_review_bundle.py`
|
||||
- accepted-suggestion writer
|
||||
- `scripts/apply_term_suggestions.py`
|
||||
|
||||
Current status:
|
||||
|
||||
- OpenClaw can read the keyword review bundle as maintenance input
|
||||
- accepted suggestions still require explicit human confirmation
|
||||
- the repository can write accepted watch / alias / stopword / interest-keyword changes after confirmation
|
||||
- this governance flow is not yet wired into a periodic scheduler inside the repository
|
||||
|
||||
Boundary:
|
||||
|
||||
- keyword cleanup is a maintenance flow, not the production RSS ingestion path
|
||||
- the repository does not auto-apply cleanup suggestions without confirmation
|
||||
- current keyword stats are built from the delivered candidate payload, not yet from a final `DailyDigest`
|
||||
|
||||
## Known Limitations
|
||||
|
||||
- Some sources put only partial content in RSS; those items may be skipped if RSS content is insufficient.
|
||||
- WeChat articles often block direct crawling, but this pipeline now avoids that path for FreshRSS items and uses RSS-provided content when available.
|
||||
- Rule behavior is still conservative in some cases; many items may land in `review` depending on current rules.
|
||||
- Paywall heuristics may produce false positives for some Chinese text patterns.
|
||||
- Keyword cleanup governance is usable now, but periodic scheduling and before/after evaluation are not implemented yet.
|
||||
|
||||
## Files OpenClaw Should Read First
|
||||
|
||||
Recommended reading order for a new maintainer:
|
||||
|
||||
1. `README.md`
|
||||
2. `docs/openclaw/openclaw-handoff.md`
|
||||
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||
5. `docs/design/daily-keyword-index-design.md`
|
||||
6. `skills/keyword-cleanup-review/SKILL.md`
|
||||
7. `docs/current/context-reset-brief.md`
|
||||
|
||||
## Downstream Boundary Rules
|
||||
|
||||
For the daily-digest workflow:
|
||||
|
||||
- the digest should go to Hugo and chat reporting, not directly into IMA
|
||||
- the full daily digest should **not** be uploaded to IMA
|
||||
- the daily digest goes to Hugo and chat reporting
|
||||
- the full daily digest should not be uploaded to IMA
|
||||
- only explicitly user-selected article summaries should be uploaded to IMA
|
||||
- selected-article summaries should be generated from existing extracted text, not by re-fetching original URLs
|
||||
|
||||
## Current Recommendation
|
||||
## Related Maintenance Flow
|
||||
|
||||
For integration handoff, the repository is usable now.
|
||||
Keyword cleanup exists as a separate maintenance flow, not the main RSS ingestion path.
|
||||
|
||||
The minimum you need to give OpenClaw is:
|
||||
|
||||
- the repository code
|
||||
- the MCP server startup command
|
||||
- the required environment variables in the target environment
|
||||
- the instruction to call `run_freshrss_openclaw_pipeline`
|
||||
- the rule that follow-up state/result reads must go through MCP tools first, not handwritten filesystem paths
|
||||
|
||||
If OpenClaw will also participate in keyword-governance review, additionally point it to:
|
||||
Relevant files:
|
||||
|
||||
- `docs/design/daily-keyword-index-design.md`
|
||||
- `skills/keyword-cleanup-review/SKILL.md`
|
||||
- `scripts/apply_term_suggestions.py`
|
||||
|
||||
## Known Limitations
|
||||
|
||||
- some sources expose only partial RSS content; those items may be skipped
|
||||
- rule behavior is still conservative; many items may land in `review`
|
||||
- paywall heuristics may still produce false positives on some Chinese text
|
||||
- keyword cleanup governance is usable but not yet wired to periodic scheduling
|
||||
|
||||
## Read First
|
||||
|
||||
Recommended reading order for a new maintainer:
|
||||
|
||||
1. `README.md`
|
||||
2. `docs/openclaw/README.md`
|
||||
3. `docs/openclaw/openclaw-handoff.md`
|
||||
4. `docs/openclaw/openclaw-orchestration-flow.md`
|
||||
5. `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||
6. `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||
7. `docs/current/context-reset-brief.md`
|
||||
|
||||
@@ -2,79 +2,86 @@
|
||||
|
||||
## 1. 文档目的
|
||||
|
||||
本文档定义 OpenClaw 在正式环境中如何调用 reader 作为上游 MCP workflow service。
|
||||
本文档只回答一个问题:OpenClaw 在正式环境里应该如何编排 reader。
|
||||
|
||||
目标不是描述 reader 内部实现,而是明确 OpenClaw 的编排动作:
|
||||
这里不重复介绍 reader 内部实现,只定义正式控制面:
|
||||
|
||||
- 什么时候启动新 run
|
||||
- 什么时候查询状态
|
||||
- 什么时候读取结果
|
||||
- 什么时候尝试恢复
|
||||
- 什么时候直接新开 run
|
||||
- 什么时候需要人工介入
|
||||
- 如何启动日报
|
||||
- 如何轮询 job
|
||||
- 如何读取 run 结果
|
||||
- 如何判断是否恢复
|
||||
- 如何走异步恢复
|
||||
- 什么时候直接新开 run 或人工介入
|
||||
|
||||
本文档基于 reader 当前**已真实落地**的能力编写,不描述尚未实现的未来接口。
|
||||
## 2. 当前正式入口
|
||||
|
||||
---
|
||||
### 2.1 新 run
|
||||
|
||||
## 2. 当前 reader 已正式支持的 MCP 能力
|
||||
正式生产入口:
|
||||
|
||||
当前可用能力:
|
||||
- `start_freshrss_pipeline_job`
|
||||
- `get_freshrss_pipeline_job_status`
|
||||
- `get_freshrss_pipeline_job_result`
|
||||
|
||||
同步入口:
|
||||
|
||||
- `run_freshrss_openclaw_pipeline`
|
||||
|
||||
同步入口只保留给 debug / fallback,不再是正式编排默认路径。
|
||||
|
||||
### 2.2 run 级读取
|
||||
|
||||
正式 run 级读取接口:
|
||||
|
||||
- `get_run_status`
|
||||
- `list_runs`
|
||||
- `list_run_artifacts`
|
||||
- `get_delivery_payload`
|
||||
- `get_run_report`
|
||||
|
||||
### 2.3 恢复
|
||||
|
||||
正式恢复入口:
|
||||
|
||||
- `inspect_resume_plan`
|
||||
- `start_resume_job`
|
||||
- `get_resume_job_status`
|
||||
- `get_resume_job_result`
|
||||
|
||||
同步恢复入口:
|
||||
|
||||
- `resume_run`
|
||||
|
||||
其中:
|
||||
|
||||
- `run_freshrss_openclaw_pipeline` 是当前正式启动入口
|
||||
- `get_run_status` / `list_runs` / `list_run_artifacts` 用于观测
|
||||
- `get_delivery_payload` / `get_run_report` 用于读取正式结果
|
||||
- `resume_run` 用于最小恢复能力
|
||||
|
||||
---
|
||||
`resume_run` 只保留给 debug / fallback。
|
||||
|
||||
## 3. 编排基本原则
|
||||
|
||||
### 3.1 OpenClaw 不再手拼路径
|
||||
### 3.1 OpenClaw 不手拼路径
|
||||
|
||||
OpenClaw 不应再自己拼 reader 输出路径来判断运行状态或读取核心结果。
|
||||
OpenClaw 不应自己推导这些路径:
|
||||
|
||||
优先使用 MCP:
|
||||
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
|
||||
|
||||
- 查状态 → `get_run_status`
|
||||
- 读 payload → `get_delivery_payload`
|
||||
- 读 report → `get_run_report`
|
||||
- 做恢复 → `resume_run`
|
||||
需要路径时,只消费 MCP 返回值:
|
||||
|
||||
只有在排障/人工核查时,才回退到直接看 reader run 目录。
|
||||
- `output_dir`
|
||||
- `artifact.path`
|
||||
- `delivery_output`
|
||||
- `report_output`
|
||||
|
||||
### 3.2 reader 是上游 workflow engine
|
||||
### 3.2 顶层 `status` 才是分支依据
|
||||
|
||||
reader 负责:
|
||||
`get_run_status` 和 job status 接口都可能做状态收敛。
|
||||
|
||||
- FreshRSS 拉取
|
||||
- 内容提取
|
||||
- 摘要
|
||||
- 过滤
|
||||
- payload 生成
|
||||
- run 状态记录
|
||||
- 最小恢复
|
||||
因此:
|
||||
|
||||
OpenClaw 负责:
|
||||
- 优先使用顶层 `status`
|
||||
- `status_source` 用来解释状态来自原始 state 还是收敛结果
|
||||
- `state_conflict=true` 说明底层状态文件已经落后于真实产物
|
||||
|
||||
- 触发执行
|
||||
- 轮询状态
|
||||
- 读取结果
|
||||
- 生成 digest markdown
|
||||
- Hugo 发布
|
||||
- 聊天汇报
|
||||
- 用户确认精选
|
||||
- IMA 编排
|
||||
不要再拿旧的 `raw_status`、`raw_current_stage` 或早期阶段名重新做分支。
|
||||
|
||||
### 3.3 默认生产语义
|
||||
|
||||
@@ -82,19 +89,18 @@ OpenClaw 负责:
|
||||
|
||||
- `mark_read=true`
|
||||
- `debug_artifacts=false`
|
||||
- 只在 debug/test/validation 时显式放宽
|
||||
|
||||
---
|
||||
只有 debug / test / validation 时才放宽。
|
||||
|
||||
## 4. 标准 Happy Path
|
||||
|
||||
### Step 1: 启动新 run
|
||||
### Step 1: 启动新 job
|
||||
|
||||
调用:
|
||||
|
||||
- `run_freshrss_openclaw_pipeline`
|
||||
- `start_freshrss_pipeline_job`
|
||||
|
||||
推荐参数示例:
|
||||
推荐参数:
|
||||
|
||||
```json
|
||||
{
|
||||
@@ -107,131 +113,156 @@ OpenClaw 负责:
|
||||
}
|
||||
```
|
||||
|
||||
期望:
|
||||
预期:
|
||||
|
||||
- 获得 `run_id`
|
||||
- 获得 `output_dir`
|
||||
- 获得初始结果摘要
|
||||
- 立即返回 `job_id`
|
||||
- 后续由 OpenClaw 轮询 job,而不是同步等待整条流水线
|
||||
|
||||
如果启动阶段直接抛错:
|
||||
### Step 2: 轮询 job
|
||||
|
||||
- 直接判为启动失败
|
||||
- 不进入后续查询
|
||||
调用:
|
||||
|
||||
### Step 2: 查询运行状态
|
||||
- `get_freshrss_pipeline_job_status(job_id=...)`
|
||||
|
||||
根据返回:
|
||||
|
||||
- `status=running`:继续轮询
|
||||
- `status=success`:读取 job result
|
||||
- `status=failed`:进入失败处理
|
||||
|
||||
额外规则:
|
||||
|
||||
- 如果 `status_source=linked_run_reconciliation`,说明 outer job state 已落后,但 linked run 已经给出可用终态
|
||||
- 如果 `status_source=stale_job_state_timeout`,把它当成终态失败,不要继续无限轮询
|
||||
|
||||
### Step 3: 读取 job result
|
||||
|
||||
调用:
|
||||
|
||||
- `get_freshrss_pipeline_job_result(job_id=...)`
|
||||
|
||||
预期读取:
|
||||
|
||||
- `run_id`
|
||||
- `output_dir`
|
||||
- `delivery_output`
|
||||
- `report_output`
|
||||
|
||||
从这一刻开始,`run_id` 是正式的稳定句柄。
|
||||
|
||||
### Step 4: 读取 run 级状态与结果
|
||||
|
||||
调用:
|
||||
|
||||
- `get_run_status(run_id=...)`
|
||||
|
||||
根据返回:
|
||||
|
||||
- `status=running` → 继续轮询
|
||||
- `status=success` → 进入结果读取
|
||||
- `status=failed` → 进入失败处理
|
||||
- `status=partial` → 视为未完成,优先看 `recovery` 和当前阶段
|
||||
|
||||
### Step 3: 读取正式结果
|
||||
|
||||
成功后读取:
|
||||
|
||||
- `get_delivery_payload(run_id=...)`
|
||||
- `get_run_report(run_id=...)`
|
||||
|
||||
后续 OpenClaw 编排应以这两个接口为正式结果源,而不是自己拼路径读取 JSON。
|
||||
根据 `get_run_status`:
|
||||
|
||||
### Step 4: 进入下游编排
|
||||
- `status=running`:继续观察
|
||||
- `status=success`:继续下游 digest / 发布 / 汇报
|
||||
- `status=failed`:进入恢复或重跑决策
|
||||
- `status=partial`:优先检查 report、artifacts 和 recovery
|
||||
|
||||
OpenClaw 在拿到正式 payload / report 后,继续执行:
|
||||
如果 `status_source=run_report_reconciliation`,说明 `run-state.json` 已经过期,但 reader 已经根据终态产物收敛出有效状态。
|
||||
|
||||
- public/internal digest 生成
|
||||
- Hugo 发布
|
||||
- 聊天汇报
|
||||
- 用户确认精选
|
||||
- IMA 沉淀
|
||||
如果 `status_source=stale_run_state_timeout`,说明 reader 认为该 run 长时间未收敛且没有终态产物,应按失败处理。
|
||||
|
||||
---
|
||||
## 5. 恢复决策
|
||||
|
||||
## 5. 状态 → 动作映射
|
||||
### 5.1 先看预检,不要直接恢复
|
||||
|
||||
| reader 状态 | OpenClaw 动作 |
|
||||
|---|---|
|
||||
| `running` | 继续轮询 `get_run_status` |
|
||||
| `success` | 读取 `get_delivery_payload` 和 `get_run_report` |
|
||||
| `failed` 且 `recovery.resumable=true` | 评估是否调用 `resume_run` |
|
||||
| `failed` 且 `recovery.resumable=false` | 直接判失败,通常新开 run 或人工介入 |
|
||||
| `partial` | 先读状态详情和 recovery,再决定继续等 / 恢复 / 人工介入 |
|
||||
恢复前固定动作:
|
||||
|
||||
---
|
||||
- 先调用 `inspect_resume_plan(run_id)`
|
||||
|
||||
## 6. 失败处理与恢复决策
|
||||
只在以下条件同时成立时才启动恢复:
|
||||
|
||||
### 6.1 什么时候优先尝试 `resume_run`
|
||||
- `can_resume=true`
|
||||
- `recommended_action=resume`
|
||||
|
||||
满足以下条件时,优先考虑恢复而不是新开 run:
|
||||
重点字段:
|
||||
|
||||
- `get_run_status` 返回 `failed`
|
||||
- `recovery.resumable=true`
|
||||
- 当前 run 对应的是 freshrss workflow
|
||||
- 当前失败点在 reader 第一版支持的恢复范围内
|
||||
- `requested_resume_from_stage`
|
||||
- `resume_from_stage`
|
||||
- `resume_decision_source`
|
||||
- `artifact_resume_from_stage`
|
||||
- `artifact_snapshot`
|
||||
|
||||
### 6.2 `resume_run` 当前支持范围
|
||||
### 5.2 正式恢复路径
|
||||
|
||||
当前最小实现仅支持:
|
||||
正式恢复控制面:
|
||||
|
||||
- 仅对带 `run-state.json` 的 freshrss run
|
||||
- 仅从最近可恢复点继续
|
||||
- 支持的恢复点:
|
||||
1. `start_resume_job(run_id)`
|
||||
2. `get_resume_job_status(job_id)`
|
||||
3. `get_resume_job_result(job_id)`
|
||||
|
||||
不要再把同步 `resume_run(run_id)` 当成正式恢复入口。
|
||||
|
||||
### 5.3 当前支持范围
|
||||
|
||||
当前只支持:
|
||||
|
||||
- 带有效 `run-state.json` 的 `freshrss_daily_digest` run
|
||||
- 从以下阶段恢复:
|
||||
- `generate_summaries`
|
||||
- `apply_filters`
|
||||
- `build_delivery_payload`
|
||||
- `write_run_report`
|
||||
|
||||
明确不支持:
|
||||
当前不支持:
|
||||
|
||||
- `fetch_feed`
|
||||
- `extract_articles`
|
||||
|
||||
### 6.3 什么时候不要恢复,直接新开 run
|
||||
正式生产恢复优先依赖:
|
||||
|
||||
以下情况不建议 `resume_run`:
|
||||
- `summary/summary-batch.json`
|
||||
- `candidates/candidate-batch.json`
|
||||
|
||||
- `recovery.resumable=false`
|
||||
- run 没有 `run-state.json`
|
||||
- 失败点是 `fetch_feed` 或 `extract_articles`
|
||||
- 恢复所需关键产物缺失
|
||||
- 恢复点语义不明确或结果存在明显漂移风险
|
||||
### 5.4 什么时候不要恢复
|
||||
|
||||
这时更合理的动作通常是:
|
||||
以下情况直接新开 run 更合理:
|
||||
|
||||
- 直接新开 run
|
||||
- 或人工介入排查
|
||||
- `recommended_action=start_new_run`
|
||||
- `recommended_action=read_terminal_result`
|
||||
- 没有有效 `run-state.json`
|
||||
- 恢复所需关键 artifacts 缺失
|
||||
- 连续恢复失败
|
||||
|
||||
### 6.4 什么时候需要人工介入
|
||||
## 6. 状态到动作映射
|
||||
|
||||
出现以下任一情况时,建议人工介入:
|
||||
| 接口 | 状态 | OpenClaw 动作 |
|
||||
| --- | --- | --- |
|
||||
| `get_freshrss_pipeline_job_status` | `running` | 继续轮询 job |
|
||||
| `get_freshrss_pipeline_job_status` | `success` | 读取 `get_freshrss_pipeline_job_result` |
|
||||
| `get_freshrss_pipeline_job_status` | `failed` | 结束本次 job,必要时读 linked run |
|
||||
| `get_run_status` | `running` | 继续观察 run |
|
||||
| `get_run_status` | `success` | 读取 `get_delivery_payload` / `get_run_report` |
|
||||
| `get_run_status` | `failed` | 先看 `inspect_resume_plan` |
|
||||
| `inspect_resume_plan` | `recommended_action=resume` | 启动 `start_resume_job` |
|
||||
| `inspect_resume_plan` | `recommended_action=read_terminal_result` | 直接读 run 结果,不恢复 |
|
||||
| `inspect_resume_plan` | `recommended_action=start_new_run` | 新开 run 或人工介入 |
|
||||
|
||||
## 7. 人工介入条件
|
||||
|
||||
出现以下任一情况时,建议不要自动编排:
|
||||
|
||||
- 连续恢复失败
|
||||
- `get_run_status` 与实际产物明显不一致
|
||||
- payload / report 结构不符合预期
|
||||
- 恢复依赖的关键文件缺失且原因不明
|
||||
- FreshRSS / LLM / 外部环境异常
|
||||
- `get_run_status` 与实际产物长期明显冲突
|
||||
- FreshRSS、LLM 或外部依赖异常
|
||||
- 恢复判定结果和编排预期不一致
|
||||
|
||||
---
|
||||
## 8. 结论
|
||||
|
||||
## 7. 读取结果的标准动作
|
||||
当前 OpenClaw 的正式调用方式已经收口为两条异步控制面:
|
||||
|
||||
### 7.1 `get_delivery_payload`
|
||||
- 主日报:`start_freshrss_pipeline_job -> poll -> get result -> run reads`
|
||||
- 恢复:`inspect_resume_plan -> start_resume_job -> poll -> get result`
|
||||
|
||||
用途:
|
||||
|
||||
- 获取正式交付给 OpenClaw 的 payload
|
||||
- 后续 digest 生成应以该返回为准
|
||||
|
||||
OpenClaw 应做:
|
||||
|
||||
- 读取后直接进入 digest 生成
|
||||
- 不再自己拼 `candidates/openclaw-delivery-payload.json`
|
||||
同步 `run_freshrss_openclaw_pipeline` 和 `resume_run` 仅用于 debug / fallback,不应再作为默认正式编排路径。
|
||||
|
||||
### 7.2 `get_run_report`
|
||||
|
||||
@@ -263,7 +294,9 @@ OpenClaw 应做:
|
||||
|
||||
1. 调 `get_run_status`
|
||||
2. 若 `failed && recovery.resumable=true`:
|
||||
- 调 `resume_run`
|
||||
- 调 `start_resume_job`
|
||||
- 轮询 `get_resume_job_status`
|
||||
- 读取 `get_resume_job_result`
|
||||
3. 恢复后再次:
|
||||
- 调 `get_run_status`
|
||||
- 若成功,再读 payload / report
|
||||
@@ -279,7 +312,7 @@ OpenClaw 应做:
|
||||
- 让 OpenClaw 直接长时间 `exec` reader CLI 作为主要生产入口
|
||||
- 让 OpenClaw 自己拼 reader 输出路径来判断成功/失败
|
||||
- 让 OpenClaw 自己读取 `outputs/.../*.json` 作为正式结果源
|
||||
- 在未确认 `resume_run` 支持范围外的失败点上强行恢复
|
||||
- 在未确认恢复支持范围外的失败点上强行恢复
|
||||
|
||||
CLI 现在的定位是:
|
||||
|
||||
@@ -293,7 +326,7 @@ CLI 现在的定位是:
|
||||
|
||||
## 10. 当前已知局限
|
||||
|
||||
- `resume_run` 仍是最小实现,不支持任意 stage 任意重入
|
||||
- 恢复能力仍是最小实现,不支持任意 stage 任意重入
|
||||
- 历史无 `run-state.json` 的 run 不支持正式恢复
|
||||
- 极旧 run 的结果读取仍可能依赖保守目录扫描
|
||||
- `write_run_report` 若涉及重新 `mark_read`,仍依赖 FreshRSS 环境和可用凭据
|
||||
@@ -304,4 +337,4 @@ CLI 现在的定位是:
|
||||
|
||||
OpenClaw 当前应把 reader 当作正式 MCP workflow service 使用:
|
||||
|
||||
**启动用 `run_freshrss_openclaw_pipeline`,观测用 `get_run_status`,结果读取用 `get_delivery_payload` / `get_run_report`,恢复仅在 `resume_run` 最小支持范围内启用;不要再把 reader 当成长 CLI 任务和路径拼接仓库来驱动。**
|
||||
**启动用 `start_freshrss_pipeline_job`,观测用 `get_run_status`,结果读取用 `get_delivery_payload` / `get_run_report`,恢复默认用 `inspect_resume_plan` + `start_resume_job`,不要再把 reader 当成长 CLI 任务和路径拼接仓库来驱动。**
|
||||
|
||||
@@ -0,0 +1,86 @@
|
||||
# [bug] MCP `start_article_summary_job` 超时但 job 实际执行了
|
||||
|
||||
## 问题描述
|
||||
|
||||
连续多个 `start_article_summary_job` 调用报 MCP 超时(-32001 Request timed out),但 job 实际执行了:
|
||||
|
||||
- `run_state.json` 显示 `status: failed`,`current_stage: generate_markdown`
|
||||
- job 目录正常生成,`run-state.json` 存在于 `outputs/freshrss/article_summary_jobs/<job-id>/`
|
||||
- `generate_markdown` 阶段实际执行过(有结果),但 MCP 响应没能发回来
|
||||
|
||||
## 根因分析(已定位)
|
||||
|
||||
**根本原因:双阻塞点导致 MCP stdio 响应超时**
|
||||
|
||||
### 阻塞点 1:`RunStore.save()` 高频同步写盘
|
||||
|
||||
`start_article_summary_job` 执行流程中的写盘次数:
|
||||
|
||||
```python
|
||||
store.save() # ← 第 134 行:初始化后写盘
|
||||
store.start_stage("prepare_job") # ← 第 68 行:内部 save()
|
||||
store.register_artifact(...) # ← 第 143 行:内部 save()
|
||||
store.finish_stage("prepare_job", ...) # ← 第 90 行:内部 save()
|
||||
return {...} # ← 第 158 行:返回 MCP 响应
|
||||
```
|
||||
|
||||
**4 次同步写盘** 在 `Popen` 之后、`return` 之前完成。磁盘 I/O 慢 + Windows 文件锁 = **响应超时**。
|
||||
|
||||
### 阻塞点 2:MCP FastMCP stdio 传输机制
|
||||
|
||||
`mcp.run()` 默认使用 **stdio 传输**(进程间管道)。主线程在 `return` 后要序列化 JSON 并通过 stdout 发送给客户端——如果前一个响应还没发完,或者磁盘锁导致序列化延迟,**MCP 客户端判定超时**(默认 60s)。
|
||||
|
||||
### 原代码问题
|
||||
|
||||
```python
|
||||
# ← 原代码:同步阻塞
|
||||
proc = subprocess.Popen(...) # Popen 返回
|
||||
store.finish_stage("prepare_job", outputs={...}) # ← 同步写盘
|
||||
return {"job_id": job_id, ...} # ← 这里已经超时了
|
||||
```
|
||||
|
||||
## 修复方案(已实施)
|
||||
|
||||
**方案 A:异步解耦** —— 用 `ThreadPoolExecutor` 后台启动 job,主线程立即返回响应。
|
||||
|
||||
```python
|
||||
# ← 修复后:异步非阻塞
|
||||
_executor = ThreadPoolExecutor(max_workers=4, thread_name_prefix="article_summary_job")
|
||||
|
||||
def _launch_job_background(*, job_id, input_payload, store):
|
||||
proc = subprocess.Popen(...)
|
||||
store.finish_stage("prepare_job", outputs={...}) # ← 后台线程写盘
|
||||
|
||||
def start_article_summary_job(...):
|
||||
# ... 创建 store 和 input 文件
|
||||
store.start_stage("prepare_job") # ← 不调用 save()
|
||||
store.register_artifact(...) # ← 不调用 save()
|
||||
|
||||
# 【关键】:后台线程执行 Popen + finish_stage,主线程立即返回
|
||||
_executor.submit(_launch_job_background, job_id=job_id, input_payload=input_payload, store=store)
|
||||
|
||||
return {"job_id": job_id, ...} # ← 立即返回,不等待写盘
|
||||
```
|
||||
|
||||
**修复效果**:
|
||||
- 主线程:创建 job 目录 → 写 input.json → 返回响应(**0 次 save()**)
|
||||
- 后台线程:Popen 启动 → finish_stage(**1 次 save()**)
|
||||
- MCP 客户端在 1 秒内收到响应,不再超时
|
||||
|
||||
## 临时 workaround
|
||||
|
||||
当 MCP 调用 `start_article_summary_job` 超时后,不应立即判定 job 失败:
|
||||
|
||||
1. 调用 `get_article_summary_job_status(job_id)` 查询真实状态
|
||||
2. 若返回 `status=running` 或 `run-state.json` 存在且 `status=running` → job 在跑,继续等待
|
||||
3. 若返回 `status=failed` → 查 `run-state.json` 的 `failed_stage` 和 `error_summary`
|
||||
|
||||
## 影响范围
|
||||
|
||||
- OpenClaw MCP 客户端调用 `start_article_summary_job`
|
||||
- 任何通过 stdio MCP 通道使用 article summary job 的场景
|
||||
|
||||
## 修复时间
|
||||
|
||||
- **根因定位**:2026-04-16
|
||||
- **修复实施**:2026-04-16(方案 A:异步解耦)
|
||||
@@ -0,0 +1,73 @@
|
||||
# 规划文档导航
|
||||
|
||||
## 使用原则
|
||||
|
||||
`plans/` 目录保存架构设计、实施计划、专题方案和历史问题分析。
|
||||
|
||||
不要把所有 plan 都当成当前权威事实。
|
||||
当前事实优先级应是:
|
||||
|
||||
1. `README.md`
|
||||
2. `docs/README.md`
|
||||
3. `docs/openclaw/README.md`
|
||||
4. `docs/openclaw/openclaw-handoff.md`
|
||||
5. `docs/openclaw/openclaw-orchestration-flow.md`
|
||||
6. `TODO.md`
|
||||
|
||||
`plans/` 更适合回答:
|
||||
|
||||
- 为什么这样设计
|
||||
- 某个能力是怎么分阶段落地的
|
||||
- 某次事故当时是怎么分析的
|
||||
|
||||
## 当前权威规划
|
||||
|
||||
这些文档仍然是当前协作时应优先阅读的规划基线:
|
||||
|
||||
- `reader-mcp-architecture-design.md`
|
||||
- MCP workflow service 的总体架构方向
|
||||
- `reader-mcp-implementation-plan.md`
|
||||
- 实施分阶段计划
|
||||
- `../TODO.md`
|
||||
- 当前任务状态与落地进展
|
||||
|
||||
## 已完成能力的专题方案
|
||||
|
||||
这些方案主要用于回看设计取舍,相关能力已经基本落地:
|
||||
|
||||
- `freshrss-pipeline-async-job-plan.md`
|
||||
- 主日报 async job 方案
|
||||
- `article-summary-async-job-plan.md`
|
||||
- 单篇总结 async job 方案
|
||||
- `resume-run-minimal-design.md`
|
||||
- `resume_run` 最小恢复语义设计
|
||||
|
||||
## 仍有参考价值的专题设计
|
||||
|
||||
- `keyword-cleanup-artifact-slimming-v1.md`
|
||||
- `keyword-cleanup-review-suggestions-layer-design.md`
|
||||
- `article-summary-prompt-independent.md`
|
||||
- `article-deep-summary-skill.md`
|
||||
- `docker-deployment-plan.md`
|
||||
|
||||
## 历史问题分析
|
||||
|
||||
- `issues/2026-04-06-reader-digest-sigterm.md`
|
||||
- 一次真实运行事故的分析
|
||||
|
||||
## 建议阅读顺序
|
||||
|
||||
如果是新接手维护:
|
||||
|
||||
1. `reader-mcp-architecture-design.md`
|
||||
2. `reader-mcp-implementation-plan.md`
|
||||
3. `../TODO.md`
|
||||
4. `../docs/openclaw/README.md`
|
||||
5. `../docs/openclaw/openclaw-handoff.md`
|
||||
|
||||
如果是在排查某一类能力:
|
||||
|
||||
- 主日报启动/轮询:看 `freshrss-pipeline-async-job-plan.md`
|
||||
- 恢复:看 `resume-run-minimal-design.md`
|
||||
- 单篇总结:看 `article-summary-async-job-plan.md`
|
||||
- 历史故障:看 `issues/2026-04-06-reader-digest-sigterm.md`
|
||||
@@ -0,0 +1,757 @@
|
||||
# 单篇总结异步 job 最小版落地方案
|
||||
|
||||
## 1. 背景与问题定义
|
||||
|
||||
当前 reader 已经把日更 FreshRSS 主流程做成了带 `run-state.json` 的 runtime 模型:
|
||||
|
||||
- `src/summary_mcp/runtime/state_models.py`
|
||||
- `src/summary_mcp/runtime/run_store.py`
|
||||
- `src/summary_mcp/runtime/query_service.py`
|
||||
- `src/summary_mcp/workflows/freshrss_pipeline.py`
|
||||
|
||||
这条主链路已经具备:
|
||||
|
||||
- run / stage / artifact 的结构化状态
|
||||
- MCP 查询接口:`get_run_status` / `list_runs` / `list_run_artifacts`
|
||||
- 结果读取接口:`get_delivery_payload` / `get_run_report`
|
||||
|
||||
但“单篇总结”这条线目前还是同步调用:
|
||||
|
||||
- 核心逻辑:`src/summary_mcp/workflows/article_summary.py`
|
||||
- MCP 暴露:`src/summary_mcp/server.py` 中的 `generate_article_summaries`
|
||||
- CLI 辅助:`scripts/run_article_summaries.py`
|
||||
|
||||
现状问题已经很明确:
|
||||
|
||||
- 在 OpenClaw → MCP tool 这条链路里,`generate_article_summaries` 可能因为 tool 调用时长而 timeout
|
||||
- 但 reader 项目本体在 `.venv` 下直接跑 article summary,大约 37.5 秒即可成功
|
||||
- 这说明问题不一定在 summary 业务本身,而更可能在“同步工具调用 + 上层等待模型”这个包装层
|
||||
|
||||
所以目标不是先继续调 timeout,而是把单篇总结也纳入 **真正异步、可轮询、可落盘、可恢复基本状态** 的最小 job 模型里。
|
||||
|
||||
---
|
||||
|
||||
## 2. 为什么同步 MCP 不适合这一步
|
||||
|
||||
`generate_article_summaries` 当前在 `server.py` 里直接同步执行 `summarize_selected_articles(...)`,调用方必须一直阻塞等待,直到:
|
||||
|
||||
1. 读取 extracted payload
|
||||
2. 调用 LLM 生成总结
|
||||
3. 可选 repair retry
|
||||
4. 渲染 Markdown
|
||||
5. 写文件完成
|
||||
6. MCP tool 返回生成路径
|
||||
|
||||
这个模式对“几十秒级、依赖外部 LLM、可能重试”的任务不稳,核心问题有三层:
|
||||
|
||||
### 2.1 tool 调用时长不可控
|
||||
|
||||
`run_loop_payload()` 内部会发起外部 HTTP 请求,还可能做 repair retry。即便单次平均 37.5 秒,也已经接近很多上层编排系统的心理和技术超时边界。
|
||||
|
||||
### 2.2 调用方看不到中间状态
|
||||
|
||||
现在如果卡住,调用方只能等:
|
||||
|
||||
- 不知道是在读输入
|
||||
- 不知道是在调 LLM
|
||||
- 不知道是在重试
|
||||
- 不知道是否已经写出部分结果
|
||||
|
||||
这也是同步接口最烦的点:失败时只能看到“tool timeout / tool failed”,而不是“业务跑到哪一步了”。
|
||||
|
||||
### 2.3 与 reader 已有 runtime 风格不一致
|
||||
|
||||
FreshRSS 主流程已经是“run truth + status query + artifact read”的思路,而单篇总结仍然是黑箱同步函数。继续维持两套风格,只会让 SOP 更复杂:
|
||||
|
||||
- 日报主链路用 `run_id`
|
||||
- 单篇总结却要么同步等,要么退回 CLI fallback
|
||||
|
||||
这不利于后续把 `reader-digest-flow` 稳定成正式 SOP。
|
||||
|
||||
---
|
||||
|
||||
## 3. 本轮目标:最小真异步,不做大而全
|
||||
|
||||
这次只做 **单篇总结异步 job 最小版**,目标是:
|
||||
|
||||
> 让 OpenClaw 或其他调用方能先“启动单篇总结 job”,立即拿到 `job_id`,再通过状态接口轮询,最后读取输出文件/结果。
|
||||
|
||||
### 3.1 本轮必须做到的范围
|
||||
|
||||
1. 新增单篇总结 job 的 start/status/result 最小接口
|
||||
2. job 真正在后台执行,而不是 MCP 请求线程里阻塞到完成
|
||||
3. 状态落盘到文件,遵循 reader 当前 runtime 风格
|
||||
4. 复用现有 `summarize_selected_articles` 逻辑,不重写业务
|
||||
5. 输出仍然是现有 Markdown 文件,不改知识内容 schema
|
||||
|
||||
### 3.2 本轮明确不做
|
||||
|
||||
1. **不做通用队列系统**
|
||||
2. **不做数据库**
|
||||
3. **不做多 worker / 分布式调度**
|
||||
4. **不做取消 job / kill job**
|
||||
5. **不做并发配额控制**
|
||||
6. **不做 resume/retry from stage**
|
||||
7. **不把 article summary 一次性并入 freshrss `resume_run` 体系**
|
||||
8. **不改 summary prompt / validator / 输出格式**
|
||||
9. **不处理批量高吞吐场景优化**
|
||||
|
||||
一句话:这轮只解“同步 tool 容易 timeout,但业务本身能跑完”这个问题,不顺手扩成任务调度平台。
|
||||
|
||||
---
|
||||
|
||||
## 4. 接入当前 reader 结构的建议
|
||||
|
||||
### 4.1 复用现有 runtime 设计,但单独建 article summary job 命名空间
|
||||
|
||||
不建议把 article summary job 粗暴塞进现有 `freshrss_daily_digest` run 查询里混用一个 schema;更合适的是:
|
||||
|
||||
- 复用 `RunState / StageState / ArtifactRecord / RunStore` 这套思维
|
||||
- 但给单篇总结定义独立 workflow 名称与存储目录
|
||||
|
||||
建议:
|
||||
|
||||
- workflow: `article_summary_job`
|
||||
- run_type: `article_summary`
|
||||
- output root: `outputs/freshrss/article_summary_jobs/<job_id>/`
|
||||
|
||||
这样有几个好处:
|
||||
|
||||
- 不污染 `outputs/freshrss/rerun/`
|
||||
- 语义清楚:这是独立 job,不是假装自己是日报 rerun
|
||||
- 查询和排查时更直观
|
||||
|
||||
### 4.2 job 与结果 Markdown 解耦
|
||||
|
||||
job 目录只负责:
|
||||
|
||||
- 状态文件
|
||||
- 输入快照
|
||||
- artifact 索引
|
||||
- 执行报告
|
||||
|
||||
真正生成的总结 Markdown,仍然可以写到用户指定的 `output_dir`(或默认 `single_summaries/`)。
|
||||
|
||||
这样不破坏当前下游 SOP:
|
||||
|
||||
- `reader-digest-flow` 依然从原来的单篇总结输出目录拿 `.md`
|
||||
- job 目录只提供状态与索引,不强迫下游改结果路径约定
|
||||
|
||||
---
|
||||
|
||||
## 5. 新增工具 / API 设计
|
||||
|
||||
本轮建议新增 3 个 MCP tool,名字尽量和当前 runtime 风格一致。
|
||||
|
||||
## 5.1 `start_article_summary_job`
|
||||
|
||||
### 作用
|
||||
|
||||
启动一个后台 job,立即返回 `job_id`,不等待总结完成。
|
||||
|
||||
### 输入建议
|
||||
|
||||
```json
|
||||
{
|
||||
"extracted_path": "outputs/freshrss/rerun/<run_id>/extracted/item-01.extracted.json",
|
||||
"selected_ids": ["12345"],
|
||||
"output_dir": "outputs/freshrss/single_summaries/2026-04-10",
|
||||
"max_retries": 2,
|
||||
"timeout_seconds": 120,
|
||||
"llm_api_key": null,
|
||||
"llm_model": null,
|
||||
"llm_api_url": null
|
||||
}
|
||||
```
|
||||
|
||||
### 返回建议
|
||||
|
||||
```json
|
||||
{
|
||||
"job_id": "article-summary-20260410-144500-ab12cd34",
|
||||
"workflow": "article_summary_job",
|
||||
"run_type": "article_summary",
|
||||
"status": "running",
|
||||
"output_dir": "outputs/freshrss/article_summary_jobs/article-summary-20260410-144500-ab12cd34",
|
||||
"message": "Article summary job started successfully. Use get_article_summary_job_status to poll progress."
|
||||
}
|
||||
```
|
||||
|
||||
### 约束建议
|
||||
|
||||
- `selected_ids` 第一版允许多个,但建议由 OpenClaw 每次只传一篇或少量篇,避免一个 job 干太多事
|
||||
- `extracted_path` 必须存在,否则直接拒绝启动
|
||||
- `output_dir` 不传则按当前默认逻辑推导
|
||||
|
||||
---
|
||||
|
||||
## 5.2 `get_article_summary_job_status`
|
||||
|
||||
### 作用
|
||||
|
||||
查询 job 当前状态、阶段、进度、错误摘要、已注册 artifact。
|
||||
|
||||
### 输入
|
||||
|
||||
```json
|
||||
{
|
||||
"job_id": "article-summary-20260410-144500-ab12cd34"
|
||||
}
|
||||
```
|
||||
|
||||
### 返回建议
|
||||
|
||||
```json
|
||||
{
|
||||
"job_id": "article-summary-20260410-144500-ab12cd34",
|
||||
"workflow": "article_summary_job",
|
||||
"run_type": "article_summary",
|
||||
"status": "running",
|
||||
"current_stage": "generate_markdown",
|
||||
"started_at": "2026-04-10T14:45:00+08:00",
|
||||
"updated_at": "2026-04-10T14:45:23+08:00",
|
||||
"finished_at": null,
|
||||
"output_dir": "outputs/freshrss/article_summary_jobs/article-summary-20260410-144500-ab12cd34",
|
||||
"progress": {
|
||||
"completed_stage_count": 2,
|
||||
"running_stage_count": 1,
|
||||
"failed_stage_count": 0,
|
||||
"pending_stage_count": 1,
|
||||
"total_stage_count": 4
|
||||
},
|
||||
"artifacts": [...],
|
||||
"error_summary": null
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 5.3 `get_article_summary_job_result`
|
||||
|
||||
### 作用
|
||||
|
||||
当 job 成功后,返回结构化结果,供 OpenClaw 继续下游 IMA 沉淀。
|
||||
|
||||
### 输入
|
||||
|
||||
```json
|
||||
{
|
||||
"job_id": "article-summary-20260410-144500-ab12cd34"
|
||||
}
|
||||
```
|
||||
|
||||
### 返回建议
|
||||
|
||||
```json
|
||||
{
|
||||
"job_id": "article-summary-20260410-144500-ab12cd34",
|
||||
"status": "success",
|
||||
"written_paths": [
|
||||
"outputs/freshrss/single_summaries/2026-04-10/某篇文章标题.md"
|
||||
],
|
||||
"artifact": {
|
||||
"name": "job_result",
|
||||
"path": "outputs/freshrss/article_summary_jobs/article-summary-20260410-144500-ab12cd34/result.json",
|
||||
"kind": "json",
|
||||
"stage": "write_result"
|
||||
},
|
||||
"result": {
|
||||
"selected_ids": ["12345"],
|
||||
"written_paths": [
|
||||
"outputs/freshrss/single_summaries/2026-04-10/某篇文章标题.md"
|
||||
]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 行为建议
|
||||
|
||||
- 若 job 还没完成,返回当前状态 + 提示“not ready”
|
||||
- 若 job 失败,返回失败摘要,不硬抛文件不存在异常
|
||||
|
||||
---
|
||||
|
||||
## 6. job 状态文件设计
|
||||
|
||||
建议直接复用现有 `RunState` 模型,不另造一套 schema。
|
||||
|
||||
job 目录示例:
|
||||
|
||||
```text
|
||||
outputs/freshrss/article_summary_jobs/
|
||||
article-summary-20260410-144500-ab12cd34/
|
||||
run-state.json
|
||||
input.json
|
||||
result.json
|
||||
job-report.json
|
||||
```
|
||||
|
||||
## 6.1 `run-state.json`
|
||||
|
||||
建议直接沿用当前字段:
|
||||
|
||||
```json
|
||||
{
|
||||
"run_id": "article-summary-20260410-144500-ab12cd34",
|
||||
"workflow": "article_summary_job",
|
||||
"run_type": "article_summary",
|
||||
"status": "running",
|
||||
"current_stage": "generate_markdown",
|
||||
"started_at": "2026-04-10T14:45:00+08:00",
|
||||
"updated_at": "2026-04-10T14:45:23+08:00",
|
||||
"finished_at": null,
|
||||
"input": {
|
||||
"extracted_path": "outputs/freshrss/rerun/<run_id>/extracted/item-01.extracted.json",
|
||||
"selected_ids": ["12345"],
|
||||
"output_dir": "outputs/freshrss/single_summaries/2026-04-10",
|
||||
"max_retries": 2,
|
||||
"timeout_seconds": 120
|
||||
},
|
||||
"stages": [...],
|
||||
"artifacts": [...],
|
||||
"error": null,
|
||||
"recovery": {
|
||||
"resumable": false,
|
||||
"resume_from_stage": null,
|
||||
"last_success_stage": "load_input"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 第一版 stage 建议
|
||||
|
||||
建议只切 4 个 stage,够看即可:
|
||||
|
||||
1. `prepare_job`
|
||||
- 校验输入
|
||||
- 解析路径
|
||||
- 写 `input.json`
|
||||
|
||||
2. `load_input`
|
||||
- 读取 extracted payload
|
||||
- 确认 `selected_ids` 可匹配条目
|
||||
|
||||
3. `generate_markdown`
|
||||
- 调 `summarize_selected_articles(...)`
|
||||
- 这是主要耗时阶段
|
||||
|
||||
4. `write_result`
|
||||
- 写 `result.json`
|
||||
- 注册输出 artifact
|
||||
|
||||
这里不要把 LLM 调用再拆更多细 stage,否则最小版反而过度设计。
|
||||
|
||||
---
|
||||
|
||||
## 6.2 `input.json`
|
||||
|
||||
作用:保留启动请求快照,便于排查。
|
||||
|
||||
建议内容与 `run-state.input` 基本一致。
|
||||
|
||||
---
|
||||
|
||||
## 6.3 `result.json`
|
||||
|
||||
成功时写:
|
||||
|
||||
```json
|
||||
{
|
||||
"job_id": "article-summary-20260410-144500-ab12cd34",
|
||||
"selected_ids": ["12345"],
|
||||
"written_paths": [
|
||||
"outputs/freshrss/single_summaries/2026-04-10/某篇文章标题.md"
|
||||
],
|
||||
"completed_at": "2026-04-10T14:45:41+08:00"
|
||||
}
|
||||
```
|
||||
|
||||
失败时可以不写,或只写失败快照都行。最小版建议:
|
||||
|
||||
- 成功写 `result.json`
|
||||
- 失败只依赖 `run-state.json`
|
||||
|
||||
避免双份失败状态不一致。
|
||||
|
||||
---
|
||||
|
||||
## 7. 执行模型建议:优先子进程,不建议线程
|
||||
|
||||
### 7.1 推荐:子进程后台执行
|
||||
|
||||
最小真异步推荐模型:
|
||||
|
||||
- `start_article_summary_job` 负责:
|
||||
- 创建 job 目录
|
||||
- 初始化 `run-state.json`
|
||||
- 通过 `subprocess.Popen(...)` 启动一个独立 Python 进程执行 job runner
|
||||
- 立即返回 `job_id`
|
||||
|
||||
后台 runner 再去:
|
||||
|
||||
- 读取 `input.json`
|
||||
- 用 `RunStore` 更新状态
|
||||
- 调用 `summarize_selected_articles(...)`
|
||||
- 写 `result.json`
|
||||
- finish/fail run
|
||||
|
||||
### 7.2 为什么不推荐线程
|
||||
|
||||
虽然线程实现看起来更省事,但不适合作为 reader 的正式最小异步落地:
|
||||
|
||||
1. **MCP server 进程重启后线程直接丢失**
|
||||
2. 线程状态不天然可恢复,容易出现“状态文件还在 running,但线程没了”
|
||||
3. 未来要做健康检查/孤儿 job 检测时,线程模型更难收口
|
||||
|
||||
### 7.3 为什么子进程更贴当前项目风格
|
||||
|
||||
reader 现在本来就偏“文件产物 + runtime 状态真相”风格。子进程模式有天然优势:
|
||||
|
||||
- 和 CLI/fallback 思维一致
|
||||
- 进程边界清楚
|
||||
- `run-state.json` 由实际执行者写,职责清晰
|
||||
- 将来如果要做 orphan detection / stale running job 修复,也容易补
|
||||
|
||||
### 7.4 本轮不做进程管理增强
|
||||
|
||||
最小版里,不要求:
|
||||
|
||||
- 记录 PID 后做 kill/cancel
|
||||
- 自动清理僵尸进程
|
||||
- 守护进程/worker 池
|
||||
|
||||
但建议在 `input.json` 或 `run-state.input` 里附带:
|
||||
|
||||
- `launcher_pid`
|
||||
- `runner_command`
|
||||
|
||||
方便排障。
|
||||
|
||||
---
|
||||
|
||||
## 8. 与现有 `summarize_selected_articles` 的复用关系
|
||||
|
||||
核心原则:**不重写总结业务,只包一层 job runner。**
|
||||
|
||||
### 8.1 直接复用的部分
|
||||
|
||||
`src/summary_mcp/workflows/article_summary.py` 已经做了:
|
||||
|
||||
- 读取 extracted payload
|
||||
- 根据 `selected_ids` 找条目
|
||||
- 调用 `run_loop_payload(...)`
|
||||
- 渲染 Markdown
|
||||
- 写到 `output_dir`
|
||||
- 返回 `list[Path]`
|
||||
|
||||
这些都继续用。
|
||||
|
||||
### 8.2 最小新增建议
|
||||
|
||||
建议只新增一层 runtime/service,例如:
|
||||
|
||||
- `src/summary_mcp/runtime/article_summary_jobs.py`
|
||||
|
||||
职责:
|
||||
|
||||
- 生成 `job_id`
|
||||
- 创建 job 目录
|
||||
- 初始化 `RunStore`
|
||||
- 启动 runner 子进程
|
||||
- 提供 status/result 查询
|
||||
- 在 runner 里调用 `summarize_selected_articles`
|
||||
|
||||
### 8.3 是否需要改 `summarize_selected_articles`
|
||||
|
||||
最小版尽量少改,只建议加两类低风险增强:
|
||||
|
||||
1. **可选输入校验增强**
|
||||
- 如果 `selected_ids` 一个都匹配不到,显式报错
|
||||
- 避免“成功返回空列表”却让 job 看起来像成功
|
||||
|
||||
2. **可选 hook / telemetry(非必须)**
|
||||
- 如果后面需要更细粒度写 stage 输出,可再加
|
||||
- 但第一版没必要为了观测性重构函数
|
||||
|
||||
结论:
|
||||
|
||||
- 第一版优先保持 `summarize_selected_articles` 基本不动
|
||||
- job 层只把它作为黑盒业务函数调用
|
||||
|
||||
---
|
||||
|
||||
## 9. 建议新增代码组织
|
||||
|
||||
建议新增文件:
|
||||
|
||||
```text
|
||||
src/summary_mcp/runtime/article_summary_jobs.py
|
||||
scripts/run_article_summary_job.py
|
||||
```
|
||||
|
||||
### 9.1 `article_summary_jobs.py`
|
||||
|
||||
建议包含:
|
||||
|
||||
- `start_article_summary_job(...)`
|
||||
- `run_article_summary_job(...)`
|
||||
- `get_article_summary_job_status(...)`
|
||||
- `get_article_summary_job_result(...)`
|
||||
- 若干私有 helper:
|
||||
- job id 生成
|
||||
- job dir 解析
|
||||
- result artifact 读取
|
||||
|
||||
### 9.2 `scripts/run_article_summary_job.py`
|
||||
|
||||
作用:作为子进程 runner 入口。
|
||||
|
||||
例如:
|
||||
|
||||
```bash
|
||||
python scripts/run_article_summary_job.py --job-id article-summary-...
|
||||
```
|
||||
|
||||
runner 只做一件事:
|
||||
|
||||
- 根据 `job_id` 找到 job 目录和 `input.json`
|
||||
- 真正执行 job
|
||||
|
||||
这样避免在 `Popen("python -c ...")` 里塞长字符串,也方便本地调试。
|
||||
|
||||
---
|
||||
|
||||
## 10. 对 `reader-digest-flow` SOP 的影响
|
||||
|
||||
这块是重点,因为老大的实际痛点就在这。
|
||||
|
||||
### 10.1 现状 SOP
|
||||
|
||||
当前 skill 已经有硬规则:
|
||||
|
||||
- 优先走 MCP `generate_article_summaries`
|
||||
- MCP timeout 时,立刻 fallback 到 reader 本地 `.venv`
|
||||
|
||||
这个 fallback 现在是必要的,但它本质是在补“单篇总结没有正式异步接口”。
|
||||
|
||||
### 10.2 引入异步 job 后的推荐 SOP
|
||||
|
||||
建议调整为:
|
||||
|
||||
1. OpenClaw 在用户确认保留文章后
|
||||
2. 调 `start_article_summary_job`
|
||||
3. 拿到 `job_id`
|
||||
4. 轮询 `get_article_summary_job_status`
|
||||
5. 成功后调 `get_article_summary_job_result`
|
||||
6. 再继续 IMA 格式化与上传
|
||||
|
||||
### 10.3 对 skill 文档的影响
|
||||
|
||||
`reader-digest-flow` 需要后续补一条新规则:
|
||||
|
||||
- 单篇总结的正式生产路径从“同步 MCP + 本地 fallback”升级为“异步 job + 状态轮询”
|
||||
- 本地 `.venv` CLI fallback 仍保留,但降级为:
|
||||
- job 启动失败
|
||||
- job runner 异常
|
||||
- reader 服务端出现系统性问题时的应急路径
|
||||
|
||||
### 10.4 用户体验改善
|
||||
|
||||
异步 job 后,OpenClaw 可给出更像正式系统的反馈:
|
||||
|
||||
- “已启动单篇总结任务,正在生成”
|
||||
- “当前状态:generate_markdown”
|
||||
- “已完成,准备继续沉淀到 IMA”
|
||||
|
||||
而不是现在的:
|
||||
|
||||
- 直接卡住几十秒
|
||||
- 然后 tool timeout
|
||||
- 再走一套 fallback
|
||||
|
||||
---
|
||||
|
||||
## 11. 验证方案
|
||||
|
||||
本轮验证不要追求大而全,按 4 层做就够。
|
||||
|
||||
### 11.1 单元级验证
|
||||
|
||||
目标:确保 job 状态文件和结果文件行为正确。
|
||||
|
||||
建议覆盖:
|
||||
|
||||
1. `start_article_summary_job` 能创建 job 目录与 `run-state.json`
|
||||
2. 输入路径不存在时,启动直接失败
|
||||
3. `get_article_summary_job_status` 能正确读取状态
|
||||
4. job 成功后 `get_article_summary_job_result` 返回 `written_paths`
|
||||
5. job 失败后状态为 `failed`,并带错误摘要
|
||||
|
||||
### 11.2 本地集成验证
|
||||
|
||||
用一个真实 extracted 文件跑:
|
||||
|
||||
1. 启动 job
|
||||
2. 轮询 status
|
||||
3. 成功后检查:
|
||||
- `result.json` 存在
|
||||
- Markdown 文件存在
|
||||
- 路径正确
|
||||
|
||||
### 11.3 OpenClaw 链路验证
|
||||
|
||||
在实际 `reader-digest-flow` 环节,用一篇已确认保留的文章做:
|
||||
|
||||
1. 启动 async job
|
||||
2. 等待成功
|
||||
3. 继续做 IMA markdown 整理与上传
|
||||
4. 确认整个链路不再因为同步 tool timeout 中断
|
||||
|
||||
### 11.4 异常验证
|
||||
|
||||
至少测 3 类异常:
|
||||
|
||||
1. `selected_ids` 不存在
|
||||
2. LLM 接口失败 / 超时
|
||||
3. runner 进程异常退出
|
||||
|
||||
预期:
|
||||
|
||||
- `run-state.json` 最终为 `failed`
|
||||
- `error_summary` 可读
|
||||
- 调用方能明确知道失败,而不是只看到 transport timeout
|
||||
|
||||
---
|
||||
|
||||
## 12. 风险与回滚方案
|
||||
|
||||
## 12.1 风险
|
||||
|
||||
### 风险 1:后台子进程成功启动,但状态长期卡在 running
|
||||
|
||||
常见原因:
|
||||
|
||||
- runner 进程崩了
|
||||
- server 重启时某些路径没写全
|
||||
- 子进程命令不对
|
||||
|
||||
缓解:
|
||||
|
||||
- runner 启动前就先写 `run-state.json`
|
||||
- runner 一进来先更新 `prepare_job` / `load_input`
|
||||
- 后续可加“stale running 超时判定”,但第一版先不做自动修复
|
||||
|
||||
### 风险 2:复用旧函数导致“空输出也算成功”
|
||||
|
||||
当前 `summarize_selected_articles()` 如果没匹配到文章或全部失败,存在返回空列表的可能。
|
||||
|
||||
缓解:
|
||||
|
||||
- job runner 里把“`written_paths` 为空”视为失败
|
||||
- 或者顺手在 `summarize_selected_articles` 里补显式校验
|
||||
|
||||
### 风险 3:同一时刻大量 job 并发,LLM 调用被打爆
|
||||
|
||||
第一版不解决系统级并发控制。
|
||||
|
||||
缓解:
|
||||
|
||||
- SOP 层先按单篇/少量串行用
|
||||
- skill 层避免一口气启动很多 job
|
||||
|
||||
### 风险 4:状态目录与结果目录分离,排查时容易迷路
|
||||
|
||||
缓解:
|
||||
|
||||
- 在 `result.json` 和 artifact 元数据里明确记录 `written_paths`
|
||||
- 在 `run-state.input.output_dir` 中保留结果目录
|
||||
|
||||
---
|
||||
|
||||
## 12.2 回滚方案
|
||||
|
||||
这个方案很好回滚,因为是“新增,不替换”。
|
||||
|
||||
### 回滚原则
|
||||
|
||||
- 保留现有 `generate_article_summaries`
|
||||
- 保留现有 `scripts/run_article_summaries.py`
|
||||
- 新增 async job 接口如果不稳定,直接停止在 OpenClaw 层使用即可
|
||||
|
||||
### 回滚路径
|
||||
|
||||
1. 停止调用 `start_article_summary_job`
|
||||
2. 恢复回原 SOP:
|
||||
- 先尝试同步 MCP `generate_article_summaries`
|
||||
- 失败则本地 `.venv` fallback
|
||||
3. async job 相关代码保留但不作为正式入口
|
||||
|
||||
也就是说,这轮改动不会堵死当前生产路径,风险可控。
|
||||
|
||||
---
|
||||
|
||||
## 13. 实施顺序(建议按这个顺序落地)
|
||||
|
||||
### Phase A:先把方案落地成最小代码骨架
|
||||
|
||||
1. 新增 `src/summary_mcp/runtime/article_summary_jobs.py`
|
||||
2. 新增 job 目录常量与 helper
|
||||
3. 新增 `scripts/run_article_summary_job.py`
|
||||
4. 实现 runner 内部对 `summarize_selected_articles` 的调用
|
||||
5. 先本地命令验证 job 可跑通
|
||||
|
||||
### Phase B:再把 MCP 接口接上
|
||||
|
||||
6. 在 `server.py` 新增:
|
||||
- `start_article_summary_job`
|
||||
- `get_article_summary_job_status`
|
||||
- `get_article_summary_job_result`
|
||||
7. 本地 MCP 调用验证
|
||||
|
||||
### Phase C:最后接 OpenClaw SOP
|
||||
|
||||
8. 更新 `reader-digest-flow` skill,把正式路径切到 async job
|
||||
9. 保留本地 `.venv` fallback 作为应急方案
|
||||
10. 跑一次真实日报保留文章沉淀闭环
|
||||
|
||||
---
|
||||
|
||||
## 14. 推荐的最小返回/状态语义
|
||||
|
||||
为了跟现有 runtime 风格一致,建议继续用:
|
||||
|
||||
- `status`: `running` / `success` / `failed`
|
||||
- `current_stage`: 当前 stage 名
|
||||
- `artifacts`: 注册产物列表
|
||||
- `error_summary`: 结构化错误
|
||||
|
||||
第一版不必引入:
|
||||
|
||||
- `queued`
|
||||
- `cancelled`
|
||||
- `retrying`
|
||||
- `paused`
|
||||
|
||||
避免状态机一开始就复杂化。
|
||||
|
||||
---
|
||||
|
||||
## 15. 结论 / 拍板建议
|
||||
|
||||
拍板建议很直接:
|
||||
|
||||
1. **这件事值得做,而且优先级高**,因为它正好卡在当前日报 SOP 的真实痛点上
|
||||
2. **不建议继续先修同步 timeout**,因为同步模型本身就不适合几十秒级、外部 LLM 驱动的任务
|
||||
3. **最小真异步应采用“文件状态 + 后台子进程 + start/status/result 三接口”**
|
||||
4. **业务层严格复用 `summarize_selected_articles`**,不要为异步化重写 summary 核心逻辑
|
||||
5. **先把单篇总结 job 做成独立小 runtime 命名空间**,不要急着并入 freshrss 主 run 的 resume 体系
|
||||
|
||||
如果只允许做一轮最小落地,我建议就做到:
|
||||
|
||||
- `start_article_summary_job`
|
||||
- `get_article_summary_job_status`
|
||||
- `get_article_summary_job_result`
|
||||
- `outputs/freshrss/article_summary_jobs/<job_id>/run-state.json`
|
||||
- 子进程 runner
|
||||
|
||||
这套已经足够把当前 OpenClaw timeout 问题从“同步等待”改成“正式异步轮询”,并且几乎不碰无关模块。
|
||||
@@ -0,0 +1,277 @@
|
||||
# reader MCP Docker 部署计划
|
||||
|
||||
## 1. 目标
|
||||
|
||||
将 reader 作为正式 MCP workflow service 以 Docker 方式部署,满足以下原则:
|
||||
|
||||
1. 服务运行在容器内
|
||||
2. 运行态与产物必须外置挂载,不闷在容器内
|
||||
3. 配置统一记录在 `.env`
|
||||
4. 读写行为与当前仓库约定保持一致
|
||||
5. OpenClaw 后续可将该服务作为正式上游 MCP 使用
|
||||
|
||||
---
|
||||
|
||||
## 2. 部署原则
|
||||
|
||||
### 2.1 容器职责
|
||||
|
||||
容器只负责:
|
||||
|
||||
- 提供 reader MCP 服务运行环境
|
||||
- 加载 reader 代码与依赖
|
||||
- 读取挂载进来的配置与状态目录
|
||||
- 对外暴露 MCP 服务入口
|
||||
|
||||
### 2.2 宿主机职责
|
||||
|
||||
宿主机负责持久化:
|
||||
|
||||
- 配置文件
|
||||
- 运行态
|
||||
- outputs 产物
|
||||
- 数据目录
|
||||
- configs
|
||||
|
||||
### 2.3 配置收口原则
|
||||
|
||||
所有环境配置统一放在 `.env`,避免:
|
||||
|
||||
- 零散写在 compose 内
|
||||
- 零散写在 shell 命令里
|
||||
- 零散写在 OpenClaw skill 里
|
||||
|
||||
---
|
||||
|
||||
## 3. 建议部署目录
|
||||
|
||||
建议在 reader 仓库内准备标准部署结构:
|
||||
|
||||
```text
|
||||
/home/ubuntu/zhu/github/reader/
|
||||
Dockerfile
|
||||
docker-compose.yml
|
||||
.env
|
||||
outputs/
|
||||
data/
|
||||
configs/
|
||||
knowledge-base/
|
||||
```
|
||||
|
||||
说明:
|
||||
|
||||
- `Dockerfile`:构建 reader MCP 服务镜像
|
||||
- `docker-compose.yml`:单服务部署编排
|
||||
- `.env`:统一环境变量
|
||||
- `outputs/`:产物、run-state、digest、payload 等外置持久化
|
||||
- `data/`:term index 等数据外置持久化
|
||||
- `configs/`:reader 运行配置外置持久化
|
||||
- `knowledge-base/`:如当前 reader/skill 仍会依赖本地知识目录,可继续挂载
|
||||
|
||||
---
|
||||
|
||||
## 4. 必须挂载的目录 / 文件
|
||||
|
||||
### 必须挂载
|
||||
|
||||
- `.env`
|
||||
- `outputs/`
|
||||
- `data/`
|
||||
- `configs/`
|
||||
|
||||
### 建议挂载
|
||||
|
||||
- `knowledge-base/`
|
||||
|
||||
### 通常不必挂载
|
||||
|
||||
- `docs/`
|
||||
- `plans/`
|
||||
- `.git/`
|
||||
|
||||
---
|
||||
|
||||
## 5. `.env` 统一配置建议
|
||||
|
||||
至少应包含以下配置:
|
||||
|
||||
### FreshRSS
|
||||
|
||||
- `FRESHRSS_API_BASE_URL`
|
||||
- `FRESHRSS_USERNAME`
|
||||
- `FRESHRSS_API_PASSWORD`
|
||||
|
||||
### 主 LLM
|
||||
|
||||
- `LLM_API_URL`
|
||||
- `LLM_API_KEY`
|
||||
- `LLM_MODEL`
|
||||
|
||||
### 单篇总结专用 LLM(如已使用)
|
||||
|
||||
- `ARTICLE_SUMMARY_API_URL`
|
||||
- `ARTICLE_SUMMARY_API_KEY`
|
||||
- `ARTICLE_SUMMARY_MODEL`
|
||||
|
||||
### IMA(如 reader / skill 仍依赖这些配置约定)
|
||||
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_ID`
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_NAME`
|
||||
|
||||
### 运行控制
|
||||
|
||||
- `PYTHONUNBUFFERED=1`
|
||||
- 视需要增加日志级别等配置
|
||||
|
||||
原则:
|
||||
|
||||
- 所有会影响服务行为的环境项,都优先进入 `.env`
|
||||
- compose 文件只引用 `.env`,不在 compose 里硬编码业务参数
|
||||
|
||||
---
|
||||
|
||||
## 6. Dockerfile 设计建议
|
||||
|
||||
### 目标
|
||||
|
||||
- 使用 Python 3.11
|
||||
- 安装 reader 依赖
|
||||
- 默认启动 MCP 服务入口
|
||||
|
||||
### 建议思路
|
||||
|
||||
1. 基于 `python:3.11-slim`
|
||||
2. 设置工作目录到 `/app`
|
||||
3. 复制仓库代码
|
||||
4. 安装依赖(如 `pip install -e .`)
|
||||
5. 默认启动 reader MCP 服务
|
||||
|
||||
### 启动入口
|
||||
|
||||
优先使用当前正式服务入口,例如:
|
||||
|
||||
- `summary-mcp`
|
||||
|
||||
如果后续 reader 明确切换到别的稳定入口,再同步更新。
|
||||
|
||||
---
|
||||
|
||||
## 7. docker-compose 设计建议
|
||||
|
||||
建议先保持单服务简单结构,例如:
|
||||
|
||||
- service 名称:`reader-mcp`
|
||||
- `env_file: .env`
|
||||
- 挂载:
|
||||
- `./outputs:/app/outputs`
|
||||
- `./data:/app/data`
|
||||
- `./configs:/app/configs`
|
||||
- `./knowledge-base:/app/knowledge-base`(如需要)
|
||||
- `./.env:/app/.env:ro`(可选,若程序直接读取文件)
|
||||
- `restart: unless-stopped`
|
||||
|
||||
如果当前 MCP 服务是 stdio 型而不是 HTTP 型,需要进一步明确:
|
||||
|
||||
- 它是由 OpenClaw 以本地进程方式拉起
|
||||
- 还是以常驻 sidecar / gateway adapter 方式挂接
|
||||
|
||||
因此 compose 的最终 command 需要结合实际接入方式确认。
|
||||
|
||||
---
|
||||
|
||||
## 8. 部署前确认项
|
||||
|
||||
在正式执行前,需要先确认以下问题:
|
||||
|
||||
### 8.1 MCP 连接方式
|
||||
|
||||
必须确认 reader MCP 服务的正式接入方式是:
|
||||
|
||||
1. **stdio 型**:OpenClaw/调用方本地拉起进程
|
||||
2. **HTTP/SSE 型**:服务常驻监听端口,OpenClaw 远程连接
|
||||
|
||||
这会直接影响:
|
||||
|
||||
- Docker command
|
||||
- 是否需要端口映射
|
||||
- OpenClaw 接入配置
|
||||
|
||||
### 8.2 当前 `summary-mcp` 的服务形态
|
||||
|
||||
需要确认:
|
||||
|
||||
- 现有 `summary-mcp` 是 FastMCP stdio 默认模式
|
||||
- 还是已有可直接 HTTP 化的运行方式
|
||||
|
||||
在这点没确认前,不要盲目写死端口暴露方案。
|
||||
|
||||
### 8.3 OpenClaw 侧接入点
|
||||
|
||||
部署完成后,还需要明确 OpenClaw 将如何引用该 MCP 服务:
|
||||
|
||||
- 本机命令型 MCP
|
||||
- Docker 内服务桥接
|
||||
- 或其它现有 OpenClaw MCP 配置方式
|
||||
|
||||
---
|
||||
|
||||
## 9. 执行顺序(建议)
|
||||
|
||||
### Phase A:部署方案落地
|
||||
|
||||
1. 确认 MCP 服务连接方式(stdio / HTTP)
|
||||
2. 确认最终 Dockerfile 启动命令
|
||||
3. 确认 compose 结构与挂载目录
|
||||
4. 整理 `.env` 字段
|
||||
|
||||
### Phase B:容器化实现
|
||||
|
||||
1. 新建/更新 `Dockerfile`
|
||||
2. 新建/更新 `docker-compose.yml`
|
||||
3. 检查 `.dockerignore`
|
||||
4. 核对路径是否与仓库内当前代码一致
|
||||
|
||||
### Phase C:本地部署验证
|
||||
|
||||
1. `docker compose build`
|
||||
2. `docker compose up -d`
|
||||
3. 验证服务启动
|
||||
4. 验证容器外 `outputs/`、`data/` 等是否正常落盘
|
||||
|
||||
### Phase D:OpenClaw 接入验证
|
||||
|
||||
1. 让 OpenClaw 通过正式 MCP 路径连接 reader
|
||||
2. 真跑一轮:
|
||||
- run
|
||||
- status
|
||||
- payload
|
||||
- report
|
||||
3. 如有需要,验证一次最小 `resume_run`
|
||||
|
||||
---
|
||||
|
||||
## 10. 当前不在本轮范围内的事
|
||||
|
||||
本轮部署计划不直接处理:
|
||||
|
||||
- `rerun_stage`
|
||||
- 更复杂的后台任务系统
|
||||
- 多实例部署
|
||||
- 横向扩展
|
||||
- 生产告警体系
|
||||
|
||||
本轮只做:
|
||||
|
||||
- 单实例
|
||||
- Docker 化
|
||||
- 配置收口
|
||||
- 挂载持久化
|
||||
- OpenClaw 可正式接入
|
||||
|
||||
---
|
||||
|
||||
## 11. 一句话结论
|
||||
|
||||
reader 的下一步不是继续堆内部接口,而是:
|
||||
|
||||
**以 Docker 正式部署成 MCP workflow service,配置进 `.env`,状态和产物目录挂载到宿主机,然后由 OpenClaw 按正式 MCP 编排路径真实接入和验证。**
|
||||
@@ -0,0 +1,229 @@
|
||||
# FreshRSS 主日报异步 job 方案
|
||||
|
||||
## 背景
|
||||
|
||||
当前 `run_freshrss_openclaw_pipeline` 虽然已经作为正式 MCP workflow 入口存在,但执行模型仍是**同步 MCP 调用**。这会带来几个现实问题:
|
||||
|
||||
1. OpenClaw / MCP wrapper 存在超时风险,尤其是 5-10 篇的正式日报批次。
|
||||
2. wrapper timeout 与真实 run 是否已落地,容易出现语义分离。
|
||||
3. 当前已有 `run-state.json`、`get_run_status`、`get_run_report`、`get_delivery_payload`,但**启动层**仍然是同步调用,不利于正式生产链路稳定运行。
|
||||
4. 单篇总结已经验证了“最小 async job + 轮询状态 + 读取结果”模型可行,主日报 run 应收敛到同一套运行模式。
|
||||
|
||||
## 目标
|
||||
|
||||
将 FreshRSS 主日报 run 改造成与 article-summary 类似的**最小真异步 job**:
|
||||
|
||||
- 启动即返回 `job_id`
|
||||
- 真正执行由后台子进程完成
|
||||
- 状态可轮询
|
||||
- 成功后可读取结构化结果
|
||||
- 业务逻辑继续复用既有 `run_freshrss_pipeline(...)`
|
||||
- 不推翻现有 run-state / result query 能力
|
||||
|
||||
## 非目标
|
||||
|
||||
本阶段不做:
|
||||
|
||||
- 分布式任务队列
|
||||
- 多 worker 调度
|
||||
- 任意 stage 的后台恢复编排
|
||||
- 并发控制中心
|
||||
- 主流程与 article-summary job 的通用抽象框架一次性大重构
|
||||
|
||||
先做最小可用。
|
||||
|
||||
## 设计原则
|
||||
|
||||
1. **启动层异步化,执行核心不重写**
|
||||
- `run_freshrss_pipeline(...)` 继续是主业务逻辑真相。
|
||||
- async job 只负责启动、状态持久化、结果回读。
|
||||
|
||||
2. **run truth 与 job truth 分层**
|
||||
- job truth:这次异步任务有没有启动、运行到哪一步、是否成功。
|
||||
- run truth:真正的 freshrss workflow 输出与 `run-state.json`。
|
||||
|
||||
3. **OpenClaw 正式生产默认改为 async start path**
|
||||
- 启动走 async job
|
||||
- 状态和结果优先先看 job
|
||||
- 真正业务产物仍由现有 run 查询工具承接
|
||||
|
||||
4. **与 article-summary job 尽量同构**
|
||||
- 目录结构
|
||||
- `run-state.json` / `input.json` / `result.json` / `job-report.json`
|
||||
- 后台 runner 脚本
|
||||
|
||||
## 拟新增能力
|
||||
|
||||
### MCP tools
|
||||
|
||||
新增 3 个工具:
|
||||
|
||||
- `start_freshrss_pipeline_job`
|
||||
- `get_freshrss_pipeline_job_status`
|
||||
- `get_freshrss_pipeline_job_result`
|
||||
|
||||
### job 目录
|
||||
|
||||
固定目录:
|
||||
|
||||
`outputs/freshrss/pipeline_jobs/<job_id>/`
|
||||
|
||||
至少包含:
|
||||
|
||||
- `run-state.json`
|
||||
- `input.json`
|
||||
- `result.json`(成功时)
|
||||
- `job-report.json`
|
||||
|
||||
### 执行模型
|
||||
|
||||
- `start_freshrss_pipeline_job` 写入 input + 初始化 job state
|
||||
- 后台 `subprocess.Popen(...)` 启动 runner
|
||||
- runner 内部调用 `run_freshrss_pipeline(...)`
|
||||
- 成功后把 `run_id`、核心产物路径、关键计数写入 `result.json`
|
||||
|
||||
## job 输入参数
|
||||
|
||||
与现有 `run_freshrss_openclaw_pipeline` 尽量对齐:
|
||||
|
||||
- `limit`
|
||||
- `mark_read`
|
||||
- `include_read`
|
||||
- `debug_artifacts`
|
||||
- `continuation`
|
||||
- `timeout_seconds`
|
||||
- `max_retries`
|
||||
- `stream_id`
|
||||
- `api_base_url`
|
||||
- `username`
|
||||
- `api_password`
|
||||
- `llm_api_key`
|
||||
- `llm_model`
|
||||
- `llm_api_url`
|
||||
- `context`
|
||||
- `run_id`
|
||||
- `date_value`
|
||||
- `output_dir`
|
||||
- `include_item_reports`
|
||||
|
||||
## 返回语义
|
||||
|
||||
### start
|
||||
|
||||
返回:
|
||||
|
||||
- `job_id`
|
||||
- `workflow`
|
||||
- `run_type`
|
||||
- `status=running`
|
||||
- `output_dir`
|
||||
- `message`
|
||||
|
||||
### status
|
||||
|
||||
返回:
|
||||
|
||||
- `job_id`
|
||||
- `status`
|
||||
- `current_stage`
|
||||
- `started_at` / `updated_at` / `finished_at`
|
||||
- `progress`
|
||||
- `artifacts`
|
||||
- `error_summary`
|
||||
- 若主 run 已创建,可附带 `linked_run_id`
|
||||
|
||||
### result
|
||||
|
||||
成功时返回:
|
||||
|
||||
- `job_id`
|
||||
- `status=success`
|
||||
- `run_id`
|
||||
- `delivery_output`
|
||||
- `report_output`
|
||||
- `digest_brief_output`
|
||||
- `pulled_count`
|
||||
- `delivered_count`
|
||||
- `marked_read_count`
|
||||
- `artifact`
|
||||
- `result`
|
||||
|
||||
## stages 建议
|
||||
|
||||
最小 job stages:
|
||||
|
||||
1. `prepare_job`
|
||||
2. `load_input`
|
||||
3. `run_pipeline`
|
||||
4. `write_result`
|
||||
|
||||
其中 `run_pipeline` 内部仍由现有 freshrss workflow 自己写它的 run-state。
|
||||
|
||||
## 与现有同步入口的关系
|
||||
|
||||
### 保留
|
||||
|
||||
`run_freshrss_openclaw_pipeline` 暂时保留,作为:
|
||||
|
||||
- debug / light path
|
||||
- 本地调试工具
|
||||
- 向后兼容路径
|
||||
|
||||
### 正式语义调整
|
||||
|
||||
文档与 OpenClaw handoff 中,主日报正式生产默认启动入口改为:
|
||||
|
||||
- `start_freshrss_pipeline_job`
|
||||
|
||||
同步入口降级为:
|
||||
|
||||
- debug / fallback
|
||||
- 小批量验证
|
||||
|
||||
## OpenClaw 编排建议
|
||||
|
||||
新的推荐路径:
|
||||
|
||||
1. `start_freshrss_pipeline_job`
|
||||
2. `get_freshrss_pipeline_job_status`
|
||||
3. 成功后 `get_freshrss_pipeline_job_result`
|
||||
4. 后续仍用:
|
||||
- `get_run_status`
|
||||
- `get_delivery_payload`
|
||||
- `get_run_report`
|
||||
- `list_run_artifacts`
|
||||
|
||||
## 风险点
|
||||
|
||||
1. **job 成功但 run 部分失败**
|
||||
- 允许,job 结果应以真实 `run_freshrss_pipeline(...)` 返回为准。
|
||||
- `run_id` + `report_output` 仍是最终真相。
|
||||
|
||||
2. **runner 崩溃但来不及写 result**
|
||||
- 需保证 `job-report.json` 至少能写下失败摘要。
|
||||
|
||||
3. **重复状态源导致混淆**
|
||||
- 文档必须明确:
|
||||
- job state 管“启动任务”
|
||||
- run state 管“业务工作流真相”
|
||||
|
||||
4. **同步 / 异步双入口长期漂移**
|
||||
- 必须要求 async job 内部直接复用 `run_freshrss_pipeline(...)`
|
||||
- 禁止再实现一套平行主流程
|
||||
|
||||
## 验收标准
|
||||
|
||||
1. 能通过 MCP 启动一个主日报 async job 并立即返回 `job_id`
|
||||
2. 能轮询到 `running -> success/failed`
|
||||
3. 成功后 `result.json` 含 `run_id` 与核心产物路径
|
||||
4. 对应 run 仍能通过既有 `get_run_status` / `get_run_report` / `get_delivery_payload` 正常读取
|
||||
5. README / handoff / TODO / plans 同步更新
|
||||
|
||||
## 建议实施顺序
|
||||
|
||||
1. 复制 article-summary job 骨架到 freshrss pipeline job
|
||||
2. 新增 runner 脚本
|
||||
3. server.py 暴露 3 个新工具
|
||||
4. 补 query/result 读法
|
||||
5. 更新 README / handoff
|
||||
6. 将 TODO 主任务切到“主日报 async job”
|
||||
@@ -0,0 +1,257 @@
|
||||
# keyword-cleanup 产物精简方案 v1
|
||||
|
||||
## 1. 背景
|
||||
|
||||
当前 keyword-cleanup 治理链已经从“只有 review bundle”演进到:
|
||||
|
||||
- term stats / daily term index
|
||||
- review bundle
|
||||
- suggestions json
|
||||
- suggestions markdown
|
||||
- apply -> config / watchlist / change_log
|
||||
|
||||
这说明链路已经打通,但也带来一个新问题:
|
||||
|
||||
> 中间产物偏多,容易让治理系统本身比被治理对象更重。
|
||||
|
||||
本方案的目标不是回退功能,而是重新划分:
|
||||
|
||||
- 哪些产物是长期资产
|
||||
- 哪些产物只是决策输入
|
||||
- 哪些产物只是运行时工作文件
|
||||
|
||||
从而把 keyword-cleanup 收敛成一个更轻的治理辅助层,而不是继续长成一个复杂子系统。
|
||||
|
||||
---
|
||||
|
||||
## 2. 设计目标
|
||||
|
||||
本轮精简目标:
|
||||
|
||||
1. 保留真正有长期价值的事实层与状态层数据
|
||||
2. 保留唯一正式建议产物,用于 review / apply
|
||||
3. 将 review bundle 和 markdown 展示稿降级为临时产物
|
||||
4. 让主链路收敛到:
|
||||
|
||||
`stats -> suggestions json -> apply -> config`
|
||||
|
||||
而不是长期依赖:
|
||||
|
||||
`stats -> bundle -> suggestions json + md -> review -> apply`
|
||||
|
||||
---
|
||||
|
||||
## 3. 产物分层建议
|
||||
|
||||
### 3.1 长期保留:事实层
|
||||
|
||||
这些文件是系统长期事实基础,应继续长期保留:
|
||||
|
||||
- `data/term_index/daily/YYYY-MM-DD.json`
|
||||
- `data/term_index/term_stats.json`
|
||||
|
||||
原因:
|
||||
|
||||
- daily 文件代表每日聚合观察结果
|
||||
- global stats 是治理决策的核心事实来源
|
||||
- 二者共同构成 term governance 的历史依据
|
||||
|
||||
### 3.2 长期保留:状态层
|
||||
|
||||
这些文件代表治理系统当前状态,应继续长期保留:
|
||||
|
||||
- `configs/filter_context.personal.json`
|
||||
- `configs/term_watchlist.json`
|
||||
- `configs/term_aliases.json`
|
||||
- `configs/term_stopwords.json`
|
||||
- `configs/term_change_log.json`
|
||||
|
||||
原因:
|
||||
|
||||
- 它们是已确认生效的治理结果
|
||||
- 后续 reader 行为依赖这些配置
|
||||
- `change_log` 负责回溯治理动作
|
||||
|
||||
### 3.3 短期保留:正式建议层
|
||||
|
||||
建议保留:
|
||||
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`
|
||||
|
||||
定位:
|
||||
|
||||
- 这是 review / apply 之间的唯一正式建议产物
|
||||
- 机器可消费
|
||||
- 可作为某次治理决策的外部依据
|
||||
|
||||
建议策略:
|
||||
|
||||
- 默认仅保留最近少量几份
|
||||
- 或仅保留已经 apply 过的 suggestions JSON
|
||||
- 避免无限累积所有历史 suggestions 文件
|
||||
|
||||
### 3.4 降级为临时产物:review bundle
|
||||
|
||||
建议降级:
|
||||
|
||||
- `outputs/term_index/review/keyword-cleanup-bundle.json`
|
||||
|
||||
定位:
|
||||
|
||||
- review 输入打包文件
|
||||
- 只服务于 suggestions 生成过程
|
||||
- 不属于长期治理资产
|
||||
|
||||
建议策略:
|
||||
|
||||
- 默认只保留当前最新一份
|
||||
- 或迁移到更明确的 working/tmp 目录语义
|
||||
- 不按日期长期积累
|
||||
|
||||
### 3.5 降级为临时产物:Markdown 展示稿
|
||||
|
||||
建议降级:
|
||||
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`
|
||||
|
||||
定位:
|
||||
|
||||
- 纯人工审阅展示层
|
||||
- 不是唯一真相
|
||||
- 不参与 apply 逻辑
|
||||
|
||||
建议策略:
|
||||
|
||||
- 默认不长期持久化
|
||||
- 需要人工审阅时临时生成
|
||||
- 优先在聊天/界面中直接展示,而不是默认写成长期文件
|
||||
|
||||
---
|
||||
|
||||
## 4. 精简后的主链路
|
||||
|
||||
建议主链路口径收敛为:
|
||||
|
||||
1. 更新 daily term index
|
||||
2. 更新 global term stats
|
||||
3. 生成 suggestions JSON
|
||||
4. 人工确认
|
||||
5. apply 到 config
|
||||
6. 记录 change log
|
||||
|
||||
即:
|
||||
|
||||
`stats -> suggestions json -> apply -> config`
|
||||
|
||||
其中:
|
||||
|
||||
- bundle = 内部工作层
|
||||
- markdown = 展示层
|
||||
- suggestions JSON = 唯一正式建议输入
|
||||
|
||||
---
|
||||
|
||||
## 5. 为什么这样收敛
|
||||
|
||||
### 5.1 避免中间层过多
|
||||
|
||||
如果 bundle / md / suggestions 都被长期持久化,就容易出现:
|
||||
|
||||
- 多份文件语义重叠
|
||||
- 不知道谁是“准的”
|
||||
- 哪些只是试跑产物,哪些是正式治理决策不清晰
|
||||
|
||||
### 5.2 保持系统重心正确
|
||||
|
||||
keyword-cleanup 的最终目的不是维护一个漂亮的 review 文件集合,而是:
|
||||
|
||||
- 持续积累稳定的关键词事实数据
|
||||
- 让 interest/watch/alias/stopword 演化有据可依
|
||||
- 让 reader 的长期偏好配置从真实日报里长出来
|
||||
|
||||
### 5.3 降低治理系统自身复杂度
|
||||
|
||||
治理系统应该比主系统更轻,而不是更重。
|
||||
如果 review 产物越积越多,最终会反过来增加维护和理解成本。
|
||||
|
||||
---
|
||||
|
||||
## 6. 对现有实现的影响
|
||||
|
||||
本轮不要求删除已有能力,而是重新定义口径。
|
||||
|
||||
### 6.1 保留
|
||||
|
||||
- `build_review_bundle.py`
|
||||
- `generate_term_cleanup_suggestions.py`
|
||||
- `apply_term_suggestions.py`
|
||||
|
||||
### 6.2 调整口径
|
||||
|
||||
- `keyword-cleanup-bundle.json` 从“默认产物”降级为“临时工作文件”
|
||||
- `term-cleanup-suggestions-YYYY-MM-DD.md` 从“正式产物”降级为“临时展示稿”
|
||||
- `term-cleanup-suggestions-YYYY-MM-DD.json` 作为唯一正式建议产物保留
|
||||
|
||||
### 6.3 后续可选实现动作
|
||||
|
||||
- 覆盖式写入 bundle,而不是长期累积
|
||||
- Markdown 按需生成,而不是默认总是落盘
|
||||
- 增加清理策略,只保留最近 N 个 suggestions JSON
|
||||
|
||||
---
|
||||
|
||||
## 7. SOP 调整建议
|
||||
|
||||
### 7.1 review 阶段
|
||||
|
||||
默认步骤:
|
||||
|
||||
1. 生成或更新 term stats
|
||||
2. 生成最新 bundle(临时)
|
||||
3. 生成 suggestions JSON(正式)
|
||||
4. 如需要人工阅读,再临时生成 Markdown 或直接在聊天展示
|
||||
|
||||
### 7.2 apply 阶段
|
||||
|
||||
apply 后以以下内容作为最终真相:
|
||||
|
||||
- config 文件当前值
|
||||
- `term_change_log.json`
|
||||
- 如需要,保留对应 suggestions JSON 作为决策依据
|
||||
|
||||
### 7.3 清理策略
|
||||
|
||||
建议:
|
||||
|
||||
- bundle:默认仅保留最新
|
||||
- markdown:默认不归档
|
||||
- suggestions JSON:保留少量最近记录或已应用记录
|
||||
|
||||
---
|
||||
|
||||
## 8. 非目标
|
||||
|
||||
本轮不做:
|
||||
|
||||
- 删除现有脚本
|
||||
- 重写治理链路
|
||||
- 一次性重构所有 review 文档
|
||||
- 自动 apply 所有 suggestions
|
||||
- 引入更复杂的存储系统
|
||||
|
||||
本轮只做一件事:
|
||||
|
||||
> 把 keyword-cleanup 的产物语义分清,长期保留该留的,临时化该临时的。
|
||||
|
||||
---
|
||||
|
||||
## 9. 一句话结论
|
||||
|
||||
keyword-cleanup 应该收敛为:
|
||||
|
||||
- **事实层长期保留**:daily / term_stats
|
||||
- **状态层长期保留**:interest / watch / alias / stopword / change_log
|
||||
- **正式建议层轻量保留**:suggestions JSON
|
||||
- **中间输入层与展示层临时化**:bundle / markdown
|
||||
|
||||
最终目标是让 reader 的关键词治理成为一个轻量、可持续、可回溯的偏好演化机制,而不是一个不断膨胀的中间文件系统。
|
||||
@@ -0,0 +1,290 @@
|
||||
# interest/watch 候选引擎改进方案
|
||||
|
||||
> 从固定阈值到自适应排位 + 趋势因子的演进
|
||||
|
||||
## 1. 背景
|
||||
|
||||
### 1.1 当前实现
|
||||
|
||||
`build_review_bundle.py` 使用固定的绝对阈值将未覆盖词(uncovered terms)划分为两个候选池:
|
||||
|
||||
| 候选池 | 判断条件 | 依据 |
|
||||
|--------|---------|------|
|
||||
| `interest_review_candidates` | `total_count >= 3 AND days_seen >= 2` | `policy.interest_keyword_review` |
|
||||
| `watch_review_candidates` | `total_count <= 2 AND days_seen <= 2` | `policy.watch_term_review` |
|
||||
|
||||
`generate_term_cleanup_suggestions.py` 则直接从这两个候选池过滤、去重、排序后输出。
|
||||
|
||||
### 1.2 当前方案的问题
|
||||
|
||||
**问题一:固定阈值不随数据量自适应**
|
||||
|
||||
```
|
||||
场景 total_count=3 意味着什么
|
||||
─────────────────────────────────────────────
|
||||
7 天数据(~200 词) top 15%,有一定区分度 ✅
|
||||
41 天数据(1070 词) top 5%,区分度更高 ✅ 但阈值没变
|
||||
未来 200 天 仍然用 3 次,区分度稀释 ❌
|
||||
```
|
||||
|
||||
同一个绝对次数,在不同数据规模下的语义完全不同。手工调阈值不可持续。
|
||||
|
||||
**问题二:固定阈值忽略趋势信号**
|
||||
|
||||
- "Anthropic":total=13, recent=7 — 近期高活跃,上升趋势
|
||||
- "Channels":total=3, recent=0 — 早期出现但近期消失
|
||||
- 当前引擎认为这两个词"都过了 3 次阈值",同等对待。实际一个是强烈买入信号,一个是过气词。
|
||||
|
||||
**问题三:interest 和 watch 的分界线是硬的**
|
||||
|
||||
total=3 → interest,total=2 → watch。一个词从 2 次变成 3 次就自动"升级",没有过渡、没有缓冲。
|
||||
|
||||
### 1.3 讨论结论
|
||||
|
||||
与老大讨论后确认:
|
||||
|
||||
1. interest/watch 是**统计判断**,不需要大模型介入,纯算法可以解决
|
||||
2. 当前引擎缺的不是大模型,而是**算法本身没写完**——自适应维度(排位、趋势)还没实现
|
||||
3. alias 和 stopword 需要语义判断,与 interest/watch 分属不同阶段,不在本方案范围内
|
||||
4. 修改量小,可以在 1 小时内落地
|
||||
|
||||
---
|
||||
|
||||
## 2. 设计方案
|
||||
|
||||
### 2.1 核心思路
|
||||
|
||||
引入两个互补维度替代固定阈值:
|
||||
|
||||
```
|
||||
判定维度 含义 数据来源
|
||||
────────────────────────────────────────────────────────────
|
||||
percentile(百分位排名) 该词 total_count 在所有词 term_stats
|
||||
中的排位占比
|
||||
growth(增速因子) 近期集中度 = recent_count daily 近 N 天
|
||||
/ total_count
|
||||
```
|
||||
|
||||
两个维度配合:
|
||||
|
||||
- **percentile** 衡量"这个词在当前数据集里有多突出"——消除数据量变化的影响
|
||||
- **growth** 衡量"这个词是持续出现还是近期爆发"——识别趋势信号
|
||||
|
||||
### 2.2 候选池划分逻辑
|
||||
|
||||
```
|
||||
percentile
|
||||
│
|
||||
┌─────────────────────┐
|
||||
│ top 5% │
|
||||
│ → 建议 interest │ ← 高频稳定词
|
||||
├─────────────────────┤
|
||||
│ top 5%-20% │
|
||||
│ → 建议 watch │ ← 有信号但未达 threshold
|
||||
├─────────────────────┤
|
||||
│ bottom 80% │
|
||||
│ → 暂不处理 │ ← 噪声/低频
|
||||
└─────────────────────┘
|
||||
|
||||
额外规则:
|
||||
如果词在 top 20% 之外,但 growth > 0.5(近期集中度高)
|
||||
→ 主动提升到 watch / 主动推 confirm
|
||||
```
|
||||
|
||||
这样就不需要关心"total_count 是 3 还是 5",只看数据自己说话。
|
||||
|
||||
### 2.3 接口变化
|
||||
|
||||
**`configs/term_cleanup_policy.json`**:
|
||||
|
||||
```json
|
||||
{
|
||||
"schema_version": "v2",
|
||||
"interest_keyword_review": {
|
||||
"percentile_max": 0.05,
|
||||
"growth_promotion": 0.5
|
||||
},
|
||||
"watch_term_review": {
|
||||
"percentile_min": 0.05,
|
||||
"percentile_max": 0.20
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
`v1` 的 `min_total_count`/`min_days_seen` 等绝对阈值字段不再使用。
|
||||
|
||||
**`build_review_bundle.py` 输出的候选项**:
|
||||
|
||||
```json
|
||||
{
|
||||
"term": "Anthropic",
|
||||
"total_count": 13,
|
||||
"days_seen": 10,
|
||||
"percentile": 0.012,
|
||||
"growth": 0.54,
|
||||
"reason": "top 1.2% by frequency, 54% of occurrences in recent window — strong signal."
|
||||
}
|
||||
```
|
||||
|
||||
### 2.4 不需要改动的部分
|
||||
|
||||
- `generate_term_cleanup_suggestions.py` — 它只消费候选池,不用改
|
||||
- `apply_term_suggestions.py` — 消费 suggestions JSON,不用改
|
||||
- `keyword-cleanup-bundle.json` 结构 — 向后兼容,新增 percentile/growth 字段
|
||||
|
||||
---
|
||||
|
||||
## 3. 实施计划
|
||||
|
||||
### 3.1 改动范围
|
||||
|
||||
| 文件 | 改动量 | 内容 |
|
||||
|------|--------|------|
|
||||
| `skills/keyword-cleanup-review/scripts/build_review_bundle.py` | ~40 行 | 新增 `_compute_percentile()` 和 `_compute_growth()` 函数;修改候选池生成逻辑;候选项中增加 percentile/growth |
|
||||
| `configs/term_cleanup_policy.json` | ~10 行 | schema v2:percentile/growth 替代绝对阈值 |
|
||||
|
||||
### 3.2 实施步骤
|
||||
|
||||
1. **build_review_bundle.py**:在 `top_global_terms` 生成后,增加 percentile 计算函数和 growth 计算函数
|
||||
2. **build_review_bundle.py**:修改 `interest_review_candidates` 和 `watch_review_candidates` 的生成逻辑,从固定阈值改为 percentile + growth
|
||||
3. **build_review_bundle.py**:候选项增加 `percentile` 和 `growth` 字段,更新 `reason` 文案
|
||||
4. **term_cleanup_policy.json**:更新为 v2 schema
|
||||
5. **验证**:全量跑一次(`--days 365 --top 100`),对比新旧两份输出的差异
|
||||
|
||||
### 3.3 验证方法
|
||||
|
||||
```bash
|
||||
# 1. 用旧版生成 baseline
|
||||
cd /home/ubuntu/zhu/github/reader
|
||||
python3.11 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
|
||||
--days 365 --top 100 \
|
||||
--output /tmp/bundle-baseline.json
|
||||
|
||||
# 2. 改代码后用新版生成
|
||||
python3.11 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
|
||||
--days 365 --top 100 \
|
||||
--output /tmp/bundle-new.json
|
||||
|
||||
# 3. 对比 governance_hints
|
||||
python3 -c "
|
||||
import json
|
||||
a = json.load(open('/tmp/bundle-baseline.json'))
|
||||
b = json.load(open('/tmp/bundle-new.json'))
|
||||
for key in ['interest_review_candidates', 'watch_review_candidates']:
|
||||
old = set(i['term'] for i in a['governance_hints'][key])
|
||||
new = set(i['term'] for i in b['governance_hints'][key])
|
||||
print(f'{key}: 新增={new-old}, 减少={old-new}')
|
||||
"
|
||||
```
|
||||
|
||||
### 3.4 风险
|
||||
|
||||
| 风险 | 概率 | 应对 |
|
||||
|------|------|------|
|
||||
| 百分位阈值对特小数据集(如只有 1 天数据)不适用 | 低 | 不足 7 天时降级回绝对阈值 |
|
||||
| growth 因子对低频词的偏差(total=1, recent=1 → growth=1) | 低 | growth 只对 total>=3 的词计算 |
|
||||
| 排位突变导致推荐漂移 | 低 | percentil 天然平滑,新增几天数据不会剧烈改变已有词的排位 |
|
||||
|
||||
---
|
||||
|
||||
## 4. alias/stopword 设计方案
|
||||
|
||||
### 4.1 核心判断
|
||||
|
||||
alias 和 stopword 需要语义理解,与 interest/watch(纯统计)性质不同。
|
||||
|
||||
| 类型 | 需要什么 | 判断方式 |
|
||||
|------|---------|----------|
|
||||
| 大小写变体 | 表层 | 规则:casefold 去重 |
|
||||
| 单复数 | 表层 | 规则:去/加 s 后缀匹配 |
|
||||
| 分词变体(空格/连字符) | 表层 | 规则:去空格归一 |
|
||||
| 简写全称(MCP→Model Context Protocol) | **语义** | LLM |
|
||||
| 中英文(上下文工程→Context Engineering) | **语义** | LLM |
|
||||
| 同义不同名(Rush→猿辅导 Rush 平台) | **语义** | LLM |
|
||||
| stopword(大模型、AI 太泛) | **语义** | LLM |
|
||||
|
||||
### 4.2 分层方案
|
||||
|
||||
```
|
||||
输入:高频未覆盖词 + 已有 interest 词表
|
||||
│
|
||||
├── 规则层(零成本)── 大小写归一、单复数、去空格/连字符
|
||||
│ 输出候选 alias 对
|
||||
│
|
||||
└── LLM 层(每次 ~500 token)── 把候选词表整批给 LLM
|
||||
做语义聚类
|
||||
输出 alias 组 + stopword 标记
|
||||
```
|
||||
|
||||
### 4.3 规则层设计
|
||||
|
||||
在 `generate_term_cleanup_suggestions.py` 中新增 `_prepare_alias_suggestions()` 函数:
|
||||
|
||||
```python
|
||||
def _prepare_alias_suggestions(top_terms, interest_keywords):
|
||||
"""
|
||||
基于表层规则生成 alias 建议。
|
||||
规则1:casefold 匹配——同一个 casefold 下有多个原文变体
|
||||
规则2:单复数——去掉/加上末尾 s 后匹配
|
||||
规则3:分词变体——去空格/连字符后匹配
|
||||
"""
|
||||
```
|
||||
|
||||
优势:零成本、可复现、可审计。直接写入 suggestions JSON,随 generate 一起输出。
|
||||
|
||||
### 4.4 LLM 层设计
|
||||
|
||||
单独脚本,非 generate 主链路的一部分。
|
||||
|
||||
```bash
|
||||
python scripts/generate_term_cleanup_semantic_suggestions.py \
|
||||
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
|
||||
--output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json
|
||||
```
|
||||
|
||||
LLM prompt 设计:
|
||||
|
||||
```
|
||||
你是一个关键词治理助手。以下是一个用户的 interest 关键词列表和一批未覆盖的高频词。
|
||||
请做三件事:
|
||||
|
||||
1. ALIAS:判断哪些未覆盖词是已有 interest 关键词的别名/变体
|
||||
2. STOPWORD:标记哪些词太宽泛/通用,建议排除
|
||||
3. PROMOTE:标记哪些新词与用户关注方向一致,建议加入 interest
|
||||
|
||||
用户关注方向:AI Agent 工程化、后端工程、开源工具、大模型落地
|
||||
```
|
||||
|
||||
LLM 层输出格式:
|
||||
|
||||
```json
|
||||
{
|
||||
"alias_suggestions": [
|
||||
{"from": "Context Engineering", "to": "上下文工程", "reason": "中英文对应同一概念"}
|
||||
],
|
||||
"stopword_suggestions": [
|
||||
{"term": "大模型", "reason": "过于宽泛,高频率但低区分度"}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
### 4.5 预期效果
|
||||
|
||||
| 覆盖类型 | 规则层 | LLM 层 |
|
||||
|---------|--------|--------|
|
||||
| 大小写变体 | ✅ | — |
|
||||
| 单复数 | ✅ | — |
|
||||
| 分词变体 | ✅ | — |
|
||||
| 简写全称 | — | ✅ |
|
||||
| 中英文映射 | — | ✅ |
|
||||
| 同义不同名 | — | ✅ |
|
||||
| stopword 判断 | — | ✅ |
|
||||
|
||||
---
|
||||
|
||||
## 5. 讨论记录
|
||||
|
||||
- 2026-05-14:与老大确认 interest/watch 不需要 LLM,纯算法可解决
|
||||
- 2026-05-14:确认百分位排名 + 增速因子方案,修改量小,优先落地
|
||||
- 2026-05-14:确认本方案不改 `generate_term_cleanup_suggestions.py` 和 `apply_term_suggestions.py`
|
||||
- 2026-05-14:确认 alias/stopword 采用规则层 + LLM 层分层方案,规则层零成本优先
|
||||
@@ -0,0 +1,534 @@
|
||||
# keyword-cleanup-review 建议产物补齐设计
|
||||
|
||||
## 1. 背景与目标
|
||||
|
||||
当前仓库已经具备 keyword cleanup review 的大部分基础设施:
|
||||
|
||||
- 已有 review bundle 构建脚本 `skills/keyword-cleanup-review/scripts/build_review_bundle.py`
|
||||
- 已有 term stats 与 daily term index 数据源
|
||||
- 已有治理输入:`configs/term_cleanup_policy.json`、`configs/term_watchlist.json`、`configs/term_change_log.json`
|
||||
- 已有建议落地脚本 `scripts/apply_term_suggestions.py`
|
||||
|
||||
当前缺口是:
|
||||
|
||||
> 缺少一层“把 `keyword-cleanup-bundle.json` 转成正式建议产物”的实现层。
|
||||
|
||||
也就是说,仓库现在能生成 review bundle,也能消费 suggestions JSON,但中间缺少稳定、可复用、可落盘的 suggestions 生成器。
|
||||
|
||||
本轮目标是补齐最小闭环,让仓库能够从 review bundle 稳定生成两份正式建议产物:
|
||||
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`
|
||||
|
||||
并保证 JSON 与 `skills/keyword-cleanup-review/references/suggestion-schema.md` 对齐,且能直接衔接 `scripts/apply_term_suggestions.py`。
|
||||
|
||||
补充口径:本方案中的 JSON 是 review / apply 之间的唯一正式建议产物;bundle 与 Markdown 主要作为运行时工作文件和临时展示层,不建议与 facts/configs 一样长期沉淀。
|
||||
|
||||
---
|
||||
|
||||
## 2. 当前现状
|
||||
|
||||
### 2.1 已有输入层
|
||||
|
||||
`build_review_bundle.py` 已经把以下输入聚合成单个 bundle:
|
||||
|
||||
- `data/term_index/term_stats.json`
|
||||
- `data/term_index/daily/*.json`
|
||||
- `configs/term_aliases.json`
|
||||
- `configs/term_stopwords.json`
|
||||
- `configs/filter_context.personal.json`
|
||||
- `configs/term_cleanup_policy.json`
|
||||
- `configs/term_watchlist.json`
|
||||
- `configs/term_change_log.json`
|
||||
|
||||
bundle 中已经包含:
|
||||
|
||||
- 当前配置快照
|
||||
- 最近 N 天热点词
|
||||
- uncovered terms
|
||||
- interest review candidates
|
||||
- watch review candidates
|
||||
|
||||
这些信息已经足够支撑“保守的、可审查的” suggestions 生成。
|
||||
|
||||
### 2.2 已有输出消费层
|
||||
|
||||
`scripts/apply_term_suggestions.py` 已经能消费 suggestions JSON,并将接受的建议写回:
|
||||
|
||||
- `configs/term_aliases.json`
|
||||
- `configs/term_stopwords.json`
|
||||
- `configs/filter_context.personal.json`
|
||||
- `configs/term_watchlist.json`
|
||||
- `configs/term_change_log.json`
|
||||
|
||||
这说明落地层已存在,缺的是中间的正式建议产物生成层。
|
||||
|
||||
---
|
||||
|
||||
## 3. 当前缺口
|
||||
|
||||
当前流程停在:
|
||||
|
||||
`build_review_bundle.py` -> `keyword-cleanup-bundle.json`
|
||||
|
||||
但缺少:
|
||||
|
||||
`keyword-cleanup-bundle.json` -> `term-cleanup-suggestions-YYYY-MM-DD.json/.md`
|
||||
|
||||
因此出现几个问题:
|
||||
|
||||
- README / 设计文档里已经引用 suggestions 产物,但仓库内没有稳定生成脚本
|
||||
- 人工审阅与后续 apply 之间没有统一的正式交付格式
|
||||
- 同一份 bundle 无法稳定、幂等地重放为同名 suggestions 产物
|
||||
- alias / stopword / interest / watch 四类建议缺少统一出入口
|
||||
|
||||
---
|
||||
|
||||
## 4. 推荐最小闭环架构
|
||||
|
||||
推荐新增一层独立脚本:
|
||||
|
||||
- `scripts/generate_term_cleanup_suggestions.py`
|
||||
|
||||
职责:
|
||||
|
||||
- 输入:`outputs/term_index/review/keyword-cleanup-bundle.json`
|
||||
- 输出:
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`
|
||||
|
||||
推荐最小数据流:
|
||||
|
||||
1. `build_review_bundle.py` 生成 bundle
|
||||
2. `generate_term_cleanup_suggestions.py` 读取 bundle
|
||||
3. 脚本基于治理 hints 生成 suggestions JSON
|
||||
4. 同时渲染人类可审阅的 Markdown
|
||||
5. 审阅后可用 `apply_term_suggestions.py` 选择性写回配置
|
||||
|
||||
本轮不引入默认 LLM 路径:
|
||||
|
||||
- 默认实现采用确定性规则生成
|
||||
- 如果未来需要 LLM 参与,应作为显式可选增强,而不是默认路径
|
||||
|
||||
---
|
||||
|
||||
## 5. 产物设计
|
||||
|
||||
### 5.1 JSON 产物
|
||||
|
||||
JSON 必须与 `suggestion-schema.md` 对齐,至少包含:
|
||||
|
||||
```json
|
||||
{
|
||||
"date": "2026-04-08",
|
||||
"based_on_days": 7,
|
||||
"alias_suggestions": [],
|
||||
"stopword_suggestions": [],
|
||||
"interest_keyword_suggestions": [
|
||||
{
|
||||
"term": "Claude Code",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered."
|
||||
}
|
||||
],
|
||||
"watch_terms": [
|
||||
{
|
||||
"term": "A2A",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed first."
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
在不破坏兼容性的前提下,可以补充少量元数据字段,建议仅限:
|
||||
|
||||
- `source_bundle`
|
||||
- `policy_schema_version`
|
||||
- `summary`
|
||||
|
||||
建议项字段口径:
|
||||
|
||||
- `alias_suggestions[]`
|
||||
- `from`
|
||||
- `to`
|
||||
- `reason`
|
||||
- `stopword_suggestions[]`
|
||||
- `term`
|
||||
- `reason`
|
||||
- `interest_keyword_suggestions[]`
|
||||
- `term`
|
||||
- `reason`
|
||||
- 可选:`total_count`、`days_seen`、`recent_count`
|
||||
- `watch_terms[]`
|
||||
- `term`
|
||||
- `reason`
|
||||
- 可选:`total_count`、`days_seen`、`recent_count`
|
||||
|
||||
兼容性要求:
|
||||
|
||||
- `apply_term_suggestions.py` 只依赖分类 bucket 与关键字段名
|
||||
- 因此额外证据字段只能追加,不能替换现有字段名
|
||||
|
||||
### 5.2 Markdown 产物
|
||||
|
||||
Markdown 推荐结构:
|
||||
|
||||
1. 标题与日期
|
||||
2. 输入 bundle 与策略摘要
|
||||
3. 当前现状摘要
|
||||
- top/global 观察
|
||||
- uncovered terms 概览
|
||||
- 当前 watchlist / interest 覆盖情况
|
||||
4. 建议摘要
|
||||
- interest 建议数量
|
||||
- watch 建议数量
|
||||
- alias 建议数量
|
||||
- stopword 建议数量
|
||||
5. `interest_keyword_suggestions`
|
||||
6. `watch_terms`
|
||||
7. `alias_suggestions`
|
||||
8. `stopword_suggestions`
|
||||
9. 应用方式
|
||||
- 指向生成的 JSON
|
||||
- 给出 `apply_term_suggestions.py` 的调用示例
|
||||
|
||||
这样可以保证:
|
||||
|
||||
- 人可以直接审阅
|
||||
- 机器可以直接消费同名 JSON
|
||||
- Markdown 与 JSON 始终一一对应
|
||||
|
||||
---
|
||||
|
||||
## 6. 建议生成策略
|
||||
|
||||
### 6.1 本轮主链路:interest / watch
|
||||
|
||||
本轮先实现最小可用主链路:
|
||||
|
||||
- `interest_keyword_suggestions`
|
||||
- `watch_terms`
|
||||
|
||||
直接复用 bundle 中已有的:
|
||||
|
||||
- `governance_hints.interest_review_candidates`
|
||||
- `governance_hints.watch_review_candidates`
|
||||
|
||||
原因:
|
||||
|
||||
- 这些候选已经与 policy 对齐
|
||||
- 这些候选已经排除了大部分已覆盖项
|
||||
- 能直接与现有 apply 脚本形成闭环
|
||||
|
||||
### 6.2 alias / stopword 保守处理
|
||||
|
||||
本轮边界明确如下:
|
||||
|
||||
- `alias_suggestions` 先保持保守,默认可为空
|
||||
- `stopword_suggestions` 先保持保守,默认可为空
|
||||
- 后续如果补充更强证据或人工审查规则,再逐步增强
|
||||
|
||||
这样可以避免在证据不足时误伤配置。
|
||||
|
||||
### 6.3 Phase 2 新方向:alias review 交给 LLM 整理
|
||||
|
||||
对于 alias,不再优先走程序规则匹配。
|
||||
Phase 2 建议改为:
|
||||
|
||||
- 程序继续负责准备 review 输入(term stats / daily / current aliases / stopwords / interest / watchlist)
|
||||
- LLM 负责整理 alias 候选
|
||||
- 默认先输出人工审阅汇报,而不是直接 apply
|
||||
- 人工确认后,再决定是否写入 `term_aliases.json`
|
||||
|
||||
这样做的原因:
|
||||
|
||||
- alias 更偏语义整理,而不是简单趋势筛选
|
||||
- 与 watch / interest 相比,alias 一旦错误归并,代价更高
|
||||
- 对当前低频治理场景来说,LLM + 人工确认更轻,也比在程序里持续堆复杂规则更合适
|
||||
|
||||
---
|
||||
|
||||
## 7. 错误处理
|
||||
|
||||
脚本应做显式校验,并在失败时给出明确错误:
|
||||
|
||||
### 7.1 输入错误
|
||||
|
||||
- bundle 文件不存在 -> 直接失败
|
||||
- bundle 不是 JSON object -> 直接失败
|
||||
- 缺少关键字段(如 `days`、`governance_hints`)-> 直接失败
|
||||
- 候选 bucket 结构错误 -> 直接失败
|
||||
|
||||
### 7.2 输出错误
|
||||
|
||||
- 输出目录不存在时自动创建
|
||||
- JSON / Markdown 写入失败时直接退出非 0
|
||||
|
||||
### 7.3 数据去重与冲突
|
||||
|
||||
- 同一 term 不能同时出现在 interest 与 watch 中
|
||||
- 优先级:`interest_keyword_suggestions` > `watch_terms`
|
||||
- 已在 bundle 当前配置中覆盖的 term 不重复输出
|
||||
|
||||
---
|
||||
|
||||
## 8. 幂等性
|
||||
|
||||
本轮要求具备基础幂等性:
|
||||
|
||||
- 同一份 bundle 多次运行,默认生成同名产物
|
||||
- 同一份 bundle 多次运行,JSON 内容顺序稳定
|
||||
- Markdown 内容顺序稳定
|
||||
|
||||
建议做法:
|
||||
|
||||
- 优先使用 bundle 的 `generated_at` 日期作为 suggestions 文件日期
|
||||
- term 排序按证据强度与 term 名稳定排序
|
||||
- 不在默认输出中写入“每次运行变化”的当前时间戳
|
||||
|
||||
这样可以让生成器作为可重放步骤存在于 review 流程中。
|
||||
|
||||
---
|
||||
|
||||
## 9. 与 apply_term_suggestions.py 的衔接
|
||||
|
||||
正式链路应变成:
|
||||
|
||||
1. `build_review_bundle.py`
|
||||
2. `generate_term_cleanup_suggestions.py`
|
||||
3. 人工审阅 Markdown
|
||||
4. `apply_term_suggestions.py --suggestions ...`
|
||||
|
||||
衔接要求:
|
||||
|
||||
- JSON bucket 名必须与 `apply_term_suggestions.py` 读取逻辑一致
|
||||
- `date` 与 `based_on_days` 字段保留,用于 change log 回写
|
||||
- 建议项中的 `reason` 直接沿用到 apply 后的 change log
|
||||
|
||||
这保证建议生成层不会成为孤立产物,而是正式进入 repo 治理闭环。
|
||||
|
||||
---
|
||||
|
||||
## 10. Python 3.11 依赖处理
|
||||
|
||||
### 10.1 当前现状
|
||||
|
||||
`build_review_bundle.py` 当前使用 `from datetime import UTC`,这要求 Python 3.11。
|
||||
|
||||
仓库整体 `pyproject.toml` 当前也声明 `requires-python = ">=3.11"`,因此短期内使用 `/usr/bin/python3.11` 运行是符合仓库现状的。
|
||||
|
||||
### 10.2 短期建议
|
||||
|
||||
短期先在文档与验证命令中明确:
|
||||
|
||||
- bundle 构建使用 `/usr/bin/python3.11`
|
||||
- suggestions 生成脚本也按仓库当前 3.11 基线运行
|
||||
|
||||
### 10.3 中期建议
|
||||
|
||||
如果后续希望把 keyword cleanup 工具链下探到 Python 3.10,可做兼容改造:
|
||||
|
||||
- 把 `datetime.UTC` 替换为 `datetime.timezone.utc`
|
||||
- 重新检查相关脚本是否还有其他 3.11-only 语法或库依赖
|
||||
|
||||
本轮不做超范围兼容重构,只在文档中把此约束说清楚。
|
||||
|
||||
---
|
||||
|
||||
## 11. 边界与非目标
|
||||
|
||||
本轮明确边界:
|
||||
|
||||
- 先实现 interest/watch 主链路
|
||||
- alias/stopword 先保持保守或留待后续增强
|
||||
- 不做超范围重构
|
||||
- 不把 LLM 作为默认生成路径
|
||||
- 不直接改 `filter_rules.json`
|
||||
- 不自动 apply 建议到配置
|
||||
|
||||
因此,本轮交付定义为:
|
||||
|
||||
- 补齐 bundle -> suggestions 的正式实现层
|
||||
- 让 review 流程可运行、可落盘、可审阅、可应用
|
||||
|
||||
而不是一次性做完所有高级治理逻辑。
|
||||
|
||||
---
|
||||
|
||||
## 12. 实施建议
|
||||
|
||||
建议按以下小步落地:
|
||||
|
||||
### Phase 1(已完成)
|
||||
|
||||
1. 新增 `scripts/generate_term_cleanup_suggestions.py`
|
||||
2. 读取 bundle 并做结构校验
|
||||
3. 生成稳定排序的 interest/watch suggestions JSON
|
||||
4. Markdown 改成按需生成
|
||||
5. README / skill 文档补一条生成命令
|
||||
6. 用 `/usr/bin/python3.11` 完整跑通 bundle -> suggestions
|
||||
|
||||
完成后,keyword cleanup review 的最小正式链路变为:
|
||||
|
||||
`review bundle` -> `suggestions json` -> `apply accepted suggestions`
|
||||
|
||||
其中 Markdown 只是按需生成的展示层。
|
||||
|
||||
### Phase 2(下一步)
|
||||
|
||||
1. 保持程序继续准备 review 输入
|
||||
2. 引入 LLM 做 alias 候选整理
|
||||
3. 默认先生成 alias review 汇报,而不是直接 apply
|
||||
4. 由人工确认后再决定是否写入 `term_aliases.json`
|
||||
|
||||
这样 alias review 会成为一个低频治理动作,而不是主链路里的自动归一步骤。
|
||||
|
||||
### Phase 3(下一步)
|
||||
|
||||
1. 保持程序继续准备 review 输入
|
||||
2. 引入 LLM 做 stopword 候选整理
|
||||
3. 默认先生成 stopword review 汇报,而不是直接 apply
|
||||
4. 由人工确认后再决定是否写入 `term_stopwords.json`
|
||||
|
||||
这样 stopword review 会成为一个低频减噪动作,而不是主链路里的自动过滤步骤。
|
||||
|
||||
#### Phase 3 输入建议
|
||||
|
||||
建议给 LLM 的输入包括:
|
||||
|
||||
- `data/term_index/term_stats.json` 中的高频词与 recent evidence
|
||||
- 最近 N 天 `data/term_index/daily/*.json` 的热点词上下文
|
||||
- 当前 `configs/term_stopwords.json`
|
||||
- 当前 `configs/term_aliases.json`
|
||||
- 当前 `configs/filter_context.personal.json` 中的 `interest_keywords`
|
||||
- 当前 `configs/term_watchlist.json`
|
||||
|
||||
程序层只负责把这些输入整理成紧凑 review context,不负责直接做 stopword 决策。
|
||||
|
||||
#### Phase 3 输出建议
|
||||
|
||||
建议 LLM 默认输出一份人工审阅汇报,而不是直接写配置:
|
||||
|
||||
- `建议加入 stopword`
|
||||
- 词
|
||||
- 简短理由
|
||||
- 证据(如 total_count / days_seen / recent_count)
|
||||
- `暂不建议加入 stopword`
|
||||
- 词
|
||||
- 为什么虽然偏泛,但当前还不能杀
|
||||
- `需要人工判断`
|
||||
- 词
|
||||
- 风险点:可能是噪声,也可能仍保留有价值信号
|
||||
|
||||
如需结构化输出,可额外补一份 `stopword_review_candidates.json`,但该文件只作为 review 输入,不直接作为自动 apply 指令。
|
||||
|
||||
#### Phase 3 审阅原则
|
||||
|
||||
- 宁可少删,不乱杀
|
||||
- 优先处理过泛、低辨识度、持续污染统计的词
|
||||
- 对可能仍承载有效技术语义的词保持保守
|
||||
- 默认先汇报,确认后再执行
|
||||
|
||||
#### Phase 3 汇报模板建议
|
||||
|
||||
建议 stopword review 默认按以下结构汇报给用户:
|
||||
|
||||
1. `建议加入 stopword`
|
||||
- 词
|
||||
- 理由:为什么这个词对治理帮助低、噪声高
|
||||
- 证据:`total_count` / `days_seen` / `recent_count` 或最近出现上下文
|
||||
2. `暂不建议加入 stopword`
|
||||
- 词
|
||||
- 理由:为什么当前不建议删掉
|
||||
3. `需要人工判断`
|
||||
- 词
|
||||
- 风险点:泛词与有效主题词之间边界不清等
|
||||
|
||||
推荐汇报风格:
|
||||
|
||||
- 简短、保守、可审阅
|
||||
- 先给判断,再给证据
|
||||
- 不输出机器式原始 dump
|
||||
- 不默认承诺“已应用”,只汇报“建议”
|
||||
|
||||
#### Phase 2 输入建议
|
||||
|
||||
建议给 LLM 的输入包括:
|
||||
|
||||
- `data/term_index/term_stats.json` 中的高频词与 recent evidence
|
||||
- 最近 N 天 `data/term_index/daily/*.json` 的热点词上下文
|
||||
- 当前 `configs/term_aliases.json`
|
||||
- 当前 `configs/term_stopwords.json`
|
||||
- 当前 `configs/filter_context.personal.json` 中的 `interest_keywords`
|
||||
- 当前 `configs/term_watchlist.json`
|
||||
|
||||
程序层只负责把这些输入整理成紧凑 review context,不负责直接做 alias 决策。
|
||||
|
||||
#### Phase 2 输出建议
|
||||
|
||||
建议 LLM 默认输出一份人工审阅汇报,而不是直接写配置:
|
||||
|
||||
- `建议合并`
|
||||
- `from -> to`
|
||||
- 简短理由
|
||||
- 证据(如 total_count / days_seen / recent_count)
|
||||
- `暂不建议合并`
|
||||
- 为什么不建议并掉
|
||||
- `需要人工判断`
|
||||
- 语义相近但风险较高的项
|
||||
|
||||
如需结构化输出,可额外补一份 `alias_review_candidates.json`,但该文件只作为 review 输入,不直接作为自动 apply 指令。
|
||||
|
||||
#### Phase 2 审阅原则
|
||||
|
||||
- 宁可少提,不乱提
|
||||
- 优先整理明显同义 / 同概念 / 词形差异
|
||||
- 不把公司名、产品名、泛概念词强行混并
|
||||
- 默认先汇报,确认后再执行
|
||||
|
||||
#### Phase 2 汇报模板建议
|
||||
|
||||
建议 alias review 默认按以下结构汇报给用户:
|
||||
|
||||
1. `建议合并`
|
||||
- `from -> to`
|
||||
- 理由:为什么判断为同一概念或更合适的标准词
|
||||
- 证据:`total_count` / `days_seen` / `recent_count` 或最近出现上下文
|
||||
2. `暂不建议合并`
|
||||
- 候选对
|
||||
- 理由:为什么虽然相近,但当前不建议并
|
||||
3. `需要人工判断`
|
||||
- 候选对
|
||||
- 风险点:歧义、范围差异、产品名/公司名混淆等
|
||||
|
||||
推荐汇报风格:
|
||||
|
||||
- 简短、保守、可审阅
|
||||
- 先给判断,再给证据
|
||||
- 不输出机器式原始 dump
|
||||
- 不默认承诺“已应用”,只汇报“建议”
|
||||
|
||||
#### Phase 2 示例输出
|
||||
|
||||
建议合并:
|
||||
|
||||
- `Claude code -> Claude Code`
|
||||
- 理由:明显属于同一产品名,仅是大小写写法不一致。
|
||||
- 证据:`Claude Code` 在最近多日持续出现,而小写写法只是在少量上下文中作为变体出现。
|
||||
|
||||
- `Sub-Agent -> SubAgent`
|
||||
- 理由:更像词形差异,不构成新的独立概念。
|
||||
- 证据:两者都围绕同一 agent 架构语境出现,且没有稳定区分语义。
|
||||
|
||||
暂不建议合并:
|
||||
|
||||
- `Skills ↔ Agent Skills`
|
||||
- 理由:前者过泛,后者更具体,当前强行归并会损失粒度。
|
||||
|
||||
- `Anthropic ↔ Claude`
|
||||
- 理由:公司名与产品名并不等价,不应直接视为一个关键词。
|
||||
|
||||
需要人工判断:
|
||||
|
||||
- `AI助手 ↔ AI Agent`
|
||||
- 风险点:语义可能接近,但中文表述范围更宽,是否并入需要结合你的使用语境判断。
|
||||
|
||||
@@ -0,0 +1,34 @@
|
||||
# 示例文章标题
|
||||
|
||||
原文链接:
|
||||
https://example.com/article
|
||||
|
||||
## 核心结论
|
||||
|
||||
这里先用一段短句概括最重要的判断。
|
||||
|
||||
如果结论较长,继续拆成第二个短段,而不是塞成一个大长段。
|
||||
|
||||
## 主要论点
|
||||
|
||||
先交代文章的核心主张。
|
||||
|
||||
再单独起一段解释支撑这个主张的关键论据。
|
||||
|
||||
如果还有补充判断,继续拆段,保证在 IMA 中阅读时不会挤成一整坨。
|
||||
|
||||
## 关键方法 / 机制
|
||||
|
||||
- 要点一
|
||||
- 要点二
|
||||
- 要点三
|
||||
|
||||
## 重要细节
|
||||
|
||||
- 细节一
|
||||
- 细节二
|
||||
|
||||
## 可复用启发
|
||||
|
||||
- 启发一
|
||||
- 启发二
|
||||
@@ -0,0 +1,40 @@
|
||||
+++
|
||||
title = "AI 日报 · 示例"
|
||||
date = 2026-04-01T16:55:00+08:00
|
||||
summary = "围绕 Agent 架构分层、Skills 标准化与桌面 Agent 工程实践的当日观察。"
|
||||
+++
|
||||
|
||||
> ⚠️ 格式规范(生成 Hugo 时必须遵守):
|
||||
> - 文章编号:`1.` `2.` `3.`(阿拉伯数字 + 点),禁止 `① ② ③` / `一、二、三` 等变体
|
||||
> - 四个 section 缺一不可:`今日概览` → `今日重点` → `趋势观察` → `延伸阅读`
|
||||
> - 每篇文章结构:标题来源 → 摘要段 → "值得关注:"三点 → "这篇更值得关注的理由"段
|
||||
|
||||
# 今日概览
|
||||
|
||||
今天的公开候选主要集中在 AI Agent 的架构演进、工具化落地与工程化实践三条线索上。相比早期偏概念展示的讨论,这一批内容更强调模块化能力栈、真实部署路径与系统可维护性,说明行业关注点正在从“模型能做什么”转向“系统如何稳定落地并持续复用”。
|
||||
|
||||
## 今日重点
|
||||
|
||||
### 1. 学习笔记:从 Agent 到 Skills — AI 智能体架构的范式转变
|
||||
来源:阿里云开发者
|
||||
|
||||
文章分析了 AI 智能体架构从单体 Agent 向模块化 Skills 的范式转变。Anthropic 先后推出 MCP 和 Agent Skills 开放标准,构建了知识、工具、协作和运行分层架构。文章通过一个自动化美化相册的真实项目,对比了 Claude Code 与 OpenClaw 两种实现方案,验证了新架构的可复用性与灵活性。
|
||||
|
||||
值得关注:
|
||||
- Anthropic 在 14 个月内先后推出 MCP 和 Agent Skills 两个开放标准,推动 AI 智能体架构分层化。
|
||||
- 新范式核心是构建薄 Agent 引擎与可组合的 Skills 库,取代为每个用例定制单体 Agent。
|
||||
- 文章通过自动化美化相册项目,实操演示了 Skills、MCP、OpenClaw 和 A2A 协议如何协同工作。
|
||||
|
||||
这篇内容更值得关注的原因在于,它不只是提出了“Agent 要模块化”这个判断,而是把开放标准、分层架构和真实项目案例串成了一条完整论证链,能直接支撑今天日报的主线。
|
||||
|
||||
## 趋势观察
|
||||
|
||||
1. Agent 正在从单体能力转向可组合的模块化体系。无论是 Skills、MCP、记忆还是运行时编排,这批内容都在强调解耦与复用,而不是把智能体继续当成一个不可拆分的黑箱。
|
||||
2. 工程化正在变成 AI 应用竞争的主战场。桌面 Agent、企业级架构和部署实践类内容增多,说明真正的差异化开始落在接入现有流程、控制风险和提升可维护性上。
|
||||
3. AI 能力的竞争点正在上移。模型本身仍重要,但真正可持续的优势越来越来自系统设计、工作流整合和对业务场景的理解。
|
||||
|
||||
## 延伸阅读
|
||||
|
||||
- [学习笔记:从 Agent 到 Skills — AI 智能体架构的范式转变](https://example.com/a)|阿里云开发者
|
||||
- [Agent Skills:打通可复用专业领域知识的最后一公里](https://example.com/b)|阿里云开发者
|
||||
- [CoPaw深度解析:源码架构和功能实践](https://example.com/c)|阿里云开发者
|
||||
+314
@@ -0,0 +1,314 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
Generate semantic keyword suggestions using LLM.
|
||||
|
||||
Covers what surface-form rules cannot:
|
||||
- semantic alias (abbreviation ↔ full name, Chinese ↔ English, synonym)
|
||||
- stopword (overly broad / low-discrimination terms)
|
||||
- promote (new term that aligns with user's focus areas)
|
||||
|
||||
Usage:
|
||||
python scripts/generate_term_cleanup_semantic_suggestions.py \
|
||||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
|
||||
--output outputs/term_index/review/term-cleanup-semantic-suggestions-2026-05-14.json
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
import time
|
||||
from datetime import datetime, timezone
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
from urllib.request import Request, urlopen
|
||||
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
DEFAULT_BUNDLE_PATH = REPO_ROOT / "outputs" / "term_index" / "review" / "keyword-cleanup-bundle.json"
|
||||
DEFAULT_OUTPUT_DIR = REPO_ROOT / "outputs" / "term_index" / "review"
|
||||
|
||||
|
||||
def _load_json(path: Path) -> Any:
|
||||
return json.loads(path.read_text(encoding="utf-8-sig"))
|
||||
|
||||
|
||||
def _save_json(path: Path, payload: dict[str, Any]) -> None:
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
path.write_text(json.dumps(payload, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
|
||||
|
||||
|
||||
def _load_env(path: Path) -> dict[str, str]:
|
||||
"""Load key=value pairs from .env file."""
|
||||
env: dict[str, str] = {}
|
||||
if not path.exists():
|
||||
return env
|
||||
for line in path.read_text(encoding="utf-8").splitlines():
|
||||
line = line.strip()
|
||||
if not line or line.startswith("#") or "=" not in line:
|
||||
continue
|
||||
key, _, value = line.partition("=")
|
||||
env[key.strip()] = value.strip().strip("\"'")
|
||||
return env
|
||||
|
||||
|
||||
def _build_prompt(
|
||||
interest_keywords: list[str],
|
||||
rule_alias_suggestions: list[dict[str, str]],
|
||||
candidate_terms: list[dict[str, Any]],
|
||||
relevant_watch_terms: list[dict[str, Any]],
|
||||
) -> str:
|
||||
"""Build the LLM prompt for semantic suggestions."""
|
||||
|
||||
interest_bullets = "\n".join(f" - {t}" for t in sorted(interest_keywords))
|
||||
candidate_bullets = "\n".join(
|
||||
f" - {t['term']} (count={t['total_count']}, days={t['days_seen']})"
|
||||
for t in candidate_terms[:40]
|
||||
)
|
||||
|
||||
# Alias from rule layer (for LLM to build on, not duplicate)
|
||||
rule_alias_text = ""
|
||||
if rule_alias_suggestions:
|
||||
rule_alias_text = "\nSurface-form alias (already identified, skip these):\n" + "\n".join(
|
||||
f" {a['from']} → {a['to']} ({a['reason']})"
|
||||
for a in rule_alias_suggestions
|
||||
)
|
||||
|
||||
watch_text = ""
|
||||
if relevant_watch_terms:
|
||||
watch_text = "\nWatch terms (low-frequency but potentially relevant):\n" + "\n".join(
|
||||
f" {t['term']} (count={t['total_count']}, days={t['days_seen']})"
|
||||
for t in relevant_watch_terms[:20]
|
||||
)
|
||||
|
||||
return f"""You are a keyword governance assistant for an AI engineer. Your job is to analyze keyword data and produce structured suggestions.
|
||||
|
||||
## User's focus areas
|
||||
- AI Agent engineering (Skills, Harness, MCP, Agent architecture)
|
||||
- Backend engineering (Java, Go, Kubernetes, MySQL, distributed systems)
|
||||
- Open source AI tools and practices (Claude Code, Cursor, DeepSeek, OpenClaw)
|
||||
- LLM application engineering (context engineering, RAG, prompt engineering)
|
||||
|
||||
## Interest keywords (52 already configured)
|
||||
{interest_bullets}
|
||||
|
||||
## Uncovered candidate terms (sorted by frequency)
|
||||
{candidate_bullets}
|
||||
{watch_text}{rule_alias_text}
|
||||
|
||||
## Task
|
||||
Analyze the candidate terms and output a JSON object with exactly three keys:
|
||||
|
||||
1. "semantic_alias": array of alias suggestions that SURFACE RULES CAN'T CATCH (e.g. abbreviation↔full name, Chinese↔English, different naming for the same concept).
|
||||
Format: [{{"from": "<variant>", "to": "<canonical interest keyword>", "reason": "<why>"}}]
|
||||
|
||||
2. "stopword": array of terms that are too broad/generic to be useful as filters. A stopword is a term that appears frequently but has LOW DISCRIMINATION — it matches too many unrelated articles and clutters the keyword index.
|
||||
Format: [{{"term": "<term>", "reason": "<why it should be a stopword>"}}]
|
||||
|
||||
3. "promote_to_interest": array of uncovered terms that align well with the user's focus areas and should be added as interest keywords.
|
||||
Format: [{{"term": "<term>", "reason": "<why it fits>"}}]
|
||||
|
||||
## Rules
|
||||
- Be conservative. When in doubt, leave it out.
|
||||
- Only suggest alias for terms that clearly refer to the SAME concept as an existing interest keyword.
|
||||
- Only suggest stopword for terms that are genuinely too broad (appear in many unrelated contexts).
|
||||
- Only suggest promote for terms that clearly match the user's stated focus areas.
|
||||
- Output valid JSON only, no markdown, no explanation outside the JSON."""
|
||||
|
||||
|
||||
def _call_llm(prompt: str, api_url: str, model: str, api_key: str) -> str:
|
||||
"""Call LLM API and return the response text."""
|
||||
payload = json.dumps({
|
||||
"model": model,
|
||||
"messages": [{"role": "user", "content": prompt}],
|
||||
"temperature": 0.1,
|
||||
"max_tokens": 2048,
|
||||
}).encode("utf-8")
|
||||
|
||||
req = Request(
|
||||
api_url.rstrip("/") + "/chat/completions",
|
||||
data=payload,
|
||||
headers={
|
||||
"Content-Type": "application/json",
|
||||
"Authorization": f"Bearer {api_key}",
|
||||
},
|
||||
)
|
||||
|
||||
max_retries = 3
|
||||
for attempt in range(max_retries):
|
||||
try:
|
||||
with urlopen(req, timeout=120) as resp:
|
||||
result = json.loads(resp.read().decode("utf-8"))
|
||||
return result["choices"][0]["message"]["content"]
|
||||
except Exception as e:
|
||||
if attempt < max_retries - 1:
|
||||
wait = 2 ** attempt
|
||||
print(f" LLM call failed (attempt {attempt+1}/{max_retries}): {e}", file=sys.stderr)
|
||||
print(f" Retrying in {wait}s...", file=sys.stderr)
|
||||
time.sleep(wait)
|
||||
else:
|
||||
raise
|
||||
|
||||
|
||||
def _parse_llm_response(text: str) -> dict[str, list[dict[str, str]]]:
|
||||
"""Extract JSON from LLM response (may contain markdown fences)."""
|
||||
# Try to find JSON block
|
||||
json_match = re.search(r"```(?:json)?\s*\n?(\{.*?\})\s*\n?```", text, re.DOTALL)
|
||||
if json_match:
|
||||
text = json_match.group(1)
|
||||
|
||||
# Clean up: remove any text before { or after }
|
||||
start = text.find("{")
|
||||
end = text.rfind("}")
|
||||
if start >= 0 and end > start:
|
||||
text = text[start : end + 1]
|
||||
|
||||
try:
|
||||
result = json.loads(text)
|
||||
except json.JSONDecodeError:
|
||||
# Try partial recovery
|
||||
print(f" Warning: LLM response not clean JSON, attempting recovery", file=sys.stderr)
|
||||
print(f" Raw: {text[:500]}", file=sys.stderr)
|
||||
return {"semantic_alias": [], "stopword": [], "promote_to_interest": []}
|
||||
|
||||
# Normalize keys
|
||||
normalized = {
|
||||
"semantic_alias": result.get("semantic_alias", result.get("alias", [])),
|
||||
"stopword": result.get("stopword", result.get("stopword_suggestions", [])),
|
||||
"promote_to_interest": result.get("promote_to_interest", result.get("promote", [])),
|
||||
}
|
||||
# Ensure each is a list
|
||||
for key in normalized:
|
||||
if not isinstance(normalized[key], list):
|
||||
normalized[key] = []
|
||||
return normalized
|
||||
|
||||
|
||||
def main() -> None:
|
||||
parser = argparse.ArgumentParser(description="Generate semantic keyword suggestions via LLM.")
|
||||
parser.add_argument("--bundle", type=Path, default=DEFAULT_BUNDLE_PATH, help="Review bundle JSON path")
|
||||
parser.add_argument("--suggestions", type=Path, default=None, help="Existing suggestions JSON (for rule alias context)")
|
||||
parser.add_argument("--output", type=Path, default=None, help="Output JSON path (auto-generated if omitted)")
|
||||
parser.add_argument("--llm-api-url", type=str, default=None, help="LLM API base URL")
|
||||
parser.add_argument("--llm-model", type=str, default=None, help="LLM model name")
|
||||
parser.add_argument("--llm-api-key", type=str, default=None, help="LLM API key")
|
||||
parser.add_argument("--dry-run", action="store_true", help="Print prompt and exit without calling LLM")
|
||||
args = parser.parse_args()
|
||||
|
||||
# Load config
|
||||
env_path = REPO_ROOT / ".env"
|
||||
env = _load_env(env_path) if env_path.exists() else {}
|
||||
|
||||
api_url = args.llm_api_url or os.environ.get("LLM_API_URL") or env.get("LLM_API_URL", "https://api.deepseek.com")
|
||||
# Map OpenClaw model aliases to actual API model names
|
||||
model_raw = args.llm_model or os.environ.get("LLM_MODEL") or env.get("LLM_MODEL", "deepseek-chat")
|
||||
MODEL_ALIAS_MAP = {
|
||||
"deepseek/deepseek-v4-flash": "deepseek-chat",
|
||||
"deepseek/deepseek-chat": "deepseek-chat",
|
||||
"deepseek-v4-flash": "deepseek-chat",
|
||||
"deepseek-chat": "deepseek-chat",
|
||||
}
|
||||
model = MODEL_ALIAS_MAP.get(model_raw, model_raw)
|
||||
api_key = args.llm_api_key or os.environ.get("LLM_API_KEY") or env.get("LLM_API_KEY", "")
|
||||
|
||||
if not api_key:
|
||||
print("Error: No LLM API key found. Set LLM_API_KEY in .env or pass --llm-api-key.", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
# Load bundle
|
||||
if not args.bundle.exists():
|
||||
print(f"Error: Bundle not found: {args.bundle}", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
bundle = _load_json(args.bundle)
|
||||
current_config = bundle.get("current_config", {})
|
||||
interest_keywords = current_config.get("interest_keywords", [])
|
||||
top_global_terms = bundle.get("top_global_terms", [])
|
||||
governance_hints = bundle.get("governance_hints", {})
|
||||
|
||||
# Build candidate list (uncovered terms from interest + watch candidates)
|
||||
candidate_terms = []
|
||||
for item in governance_hints.get("interest_review_candidates", []):
|
||||
if isinstance(item, dict):
|
||||
candidate_terms.append({
|
||||
"term": item.get("term", ""),
|
||||
"total_count": item.get("total_count", 0),
|
||||
"days_seen": item.get("days_seen", 0),
|
||||
"percentile": item.get("percentile", 0),
|
||||
"growth": item.get("growth", 0),
|
||||
})
|
||||
for item in governance_hints.get("watch_review_candidates", []):
|
||||
if isinstance(item, dict):
|
||||
# Avoid duplicates
|
||||
if not any(c["term"] == item.get("term") for c in candidate_terms):
|
||||
candidate_terms.append({
|
||||
"term": item.get("term", ""),
|
||||
"total_count": item.get("total_count", 0),
|
||||
"days_seen": item.get("days_seen", 0),
|
||||
"percentile": item.get("percentile", 0),
|
||||
"growth": item.get("growth", 0),
|
||||
})
|
||||
|
||||
# Sort by total_count descending
|
||||
candidate_terms.sort(key=lambda x: -x["total_count"])
|
||||
relevant_watch_terms = governance_hints.get("watch_review_candidates", [])[:20]
|
||||
|
||||
# Load rule-layer alias suggestions if available
|
||||
rule_alias = []
|
||||
if args.suggestions and args.suggestions.exists():
|
||||
s = _load_json(args.suggestions)
|
||||
rule_alias = s.get("alias_suggestions", [])
|
||||
|
||||
# Build prompt
|
||||
prompt = _build_prompt(
|
||||
interest_keywords=interest_keywords,
|
||||
rule_alias_suggestions=rule_alias,
|
||||
candidate_terms=candidate_terms,
|
||||
relevant_watch_terms=relevant_watch_terms,
|
||||
)
|
||||
|
||||
# Determine output path
|
||||
suggestion_date = datetime.now(timezone.utc).date().isoformat()
|
||||
output_path = args.output or (DEFAULT_OUTPUT_DIR / f"term-cleanup-semantic-suggestions-{suggestion_date}.json")
|
||||
|
||||
if args.dry_run:
|
||||
print("=== DRY RUN: Prompt ===")
|
||||
print(prompt)
|
||||
print("\n=== END ===")
|
||||
print(f"\nWould write to: {output_path}")
|
||||
return
|
||||
|
||||
# Call LLM
|
||||
print(f"Calling LLM ({model})...", file=sys.stderr)
|
||||
response = _call_llm(prompt, api_url, model, api_key)
|
||||
print(f"LLM response received ({len(response)} chars)", file=sys.stderr)
|
||||
|
||||
# Parse
|
||||
parsed = _parse_llm_response(response)
|
||||
|
||||
# Build output
|
||||
output = {
|
||||
"date": suggestion_date,
|
||||
"source_bundle": str(args.bundle),
|
||||
"model": model,
|
||||
"interest_keyword_count": len(interest_keywords),
|
||||
"candidate_count": len(candidate_terms),
|
||||
**parsed,
|
||||
}
|
||||
|
||||
_save_json(output_path, output)
|
||||
|
||||
summary = {
|
||||
"output": str(output_path),
|
||||
"semantic_alias": len(output.get("semantic_alias", [])),
|
||||
"stopword": len(output.get("stopword", [])),
|
||||
"promote_to_interest": len(output.get("promote_to_interest", [])),
|
||||
}
|
||||
print(json.dumps(summary, ensure_ascii=False, indent=2))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,536 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
from datetime import datetime, timezone
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
DEFAULT_BUNDLE_PATH = REPO_ROOT / "outputs" / "term_index" / "review" / "keyword-cleanup-bundle.json"
|
||||
DEFAULT_OUTPUT_DIR = REPO_ROOT / "outputs" / "term_index" / "review"
|
||||
|
||||
|
||||
def _load_json(path: Path) -> Any:
|
||||
return json.loads(path.read_text(encoding="utf-8-sig"))
|
||||
|
||||
|
||||
def _save_json(path: Path, payload: dict[str, Any]) -> None:
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
path.write_text(json.dumps(payload, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
|
||||
|
||||
|
||||
def _save_text(path: Path, content: str) -> None:
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
path.write_text(content, encoding="utf-8")
|
||||
|
||||
|
||||
def _term_key(value: str) -> str:
|
||||
return value.strip().casefold()
|
||||
|
||||
|
||||
def _utc_today() -> str:
|
||||
return datetime.now(timezone.utc).date().isoformat()
|
||||
|
||||
|
||||
def _require_dict(payload: Any, name: str) -> dict[str, Any]:
|
||||
if not isinstance(payload, dict):
|
||||
raise RuntimeError(f"{name} must be a JSON object.")
|
||||
return payload
|
||||
|
||||
|
||||
def _require_list(payload: Any, name: str) -> list[Any]:
|
||||
if not isinstance(payload, list):
|
||||
raise RuntimeError(f"{name} must be a JSON array.")
|
||||
return payload
|
||||
|
||||
|
||||
def _bundle_date(bundle: dict[str, Any]) -> str:
|
||||
generated_at = bundle.get("generated_at")
|
||||
if isinstance(generated_at, str) and generated_at.strip():
|
||||
normalized = generated_at.replace("Z", "+00:00")
|
||||
try:
|
||||
return datetime.fromisoformat(normalized).date().isoformat()
|
||||
except ValueError:
|
||||
pass
|
||||
return _utc_today()
|
||||
|
||||
|
||||
def _recent_count_map(top_global_terms: list[dict[str, Any]]) -> dict[str, int]:
|
||||
counts: dict[str, int] = {}
|
||||
for item in top_global_terms:
|
||||
term = item.get("term")
|
||||
recent_count = item.get("recent_count")
|
||||
if isinstance(term, str) and isinstance(recent_count, int):
|
||||
counts[term] = recent_count
|
||||
return counts
|
||||
|
||||
|
||||
def _covered_term_sets(bundle: dict[str, Any]) -> tuple[set[str], set[str], set[str]]:
|
||||
current_config = _require_dict(bundle.get("current_config"), "bundle.current_config")
|
||||
interest_keywords = _require_list(current_config.get("interest_keywords"), "bundle.current_config.interest_keywords")
|
||||
stopwords = _require_list(current_config.get("stopwords"), "bundle.current_config.stopwords")
|
||||
watchlist = _require_list(current_config.get("watchlist"), "bundle.current_config.watchlist")
|
||||
|
||||
interest_set = {_term_key(item) for item in interest_keywords if isinstance(item, str) and item.strip()}
|
||||
stopword_set = {_term_key(item) for item in stopwords if isinstance(item, str) and item.strip()}
|
||||
watch_set = {
|
||||
_term_key(str(item.get("term", "")))
|
||||
for item in watchlist
|
||||
if isinstance(item, dict) and isinstance(item.get("term"), str) and str(item.get("term", "")).strip()
|
||||
}
|
||||
return interest_set, stopword_set, watch_set
|
||||
|
||||
|
||||
def _sort_key(item: dict[str, Any]) -> tuple[int, int, int, str, str]:
|
||||
total_count = int(item.get("total_count") or 0)
|
||||
days_seen = int(item.get("days_seen") or 0)
|
||||
recent_count = int(item.get("recent_count") or 0)
|
||||
term = str(item.get("term") or "")
|
||||
return (-total_count, -days_seen, -recent_count, term.casefold(), term)
|
||||
|
||||
|
||||
def _prepare_interest_suggestions(bundle: dict[str, Any]) -> list[dict[str, Any]]:
|
||||
governance_hints = _require_dict(bundle.get("governance_hints"), "bundle.governance_hints")
|
||||
candidates = _require_list(
|
||||
governance_hints.get("interest_review_candidates"),
|
||||
"bundle.governance_hints.interest_review_candidates",
|
||||
)
|
||||
top_global_terms = _require_list(bundle.get("top_global_terms"), "bundle.top_global_terms")
|
||||
recent_counts = _recent_count_map([item for item in top_global_terms if isinstance(item, dict)])
|
||||
interest_set, stopword_set, watch_set = _covered_term_sets(bundle)
|
||||
|
||||
suggestions: list[dict[str, Any]] = []
|
||||
seen: set[str] = set()
|
||||
for item in candidates:
|
||||
if not isinstance(item, dict):
|
||||
continue
|
||||
term = item.get("term")
|
||||
if not isinstance(term, str) or not term.strip():
|
||||
continue
|
||||
term_key = _term_key(term)
|
||||
if term_key in seen or term_key in interest_set or term_key in stopword_set:
|
||||
continue
|
||||
total_count = int(item.get("total_count") or 0)
|
||||
days_seen = int(item.get("days_seen") or 0)
|
||||
recent_count = recent_counts.get(term, 0)
|
||||
base_reason = str(item.get("reason") or "Meets the configured interest-keyword review threshold.")
|
||||
if term_key in watch_set:
|
||||
base_reason += " It is currently in watchlist and is ready for promotion."
|
||||
reason = f"{base_reason} Evidence: total_count={total_count}, days_seen={days_seen}, recent_count={recent_count}."
|
||||
suggestions.append(
|
||||
{
|
||||
"term": term,
|
||||
"reason": reason,
|
||||
"total_count": total_count,
|
||||
"days_seen": days_seen,
|
||||
"recent_count": recent_count,
|
||||
}
|
||||
)
|
||||
seen.add(term_key)
|
||||
|
||||
suggestions.sort(key=_sort_key)
|
||||
return suggestions
|
||||
|
||||
|
||||
def _prepare_watch_suggestions(bundle: dict[str, Any], reserved_terms: set[str]) -> list[dict[str, Any]]:
|
||||
governance_hints = _require_dict(bundle.get("governance_hints"), "bundle.governance_hints")
|
||||
candidates = _require_list(
|
||||
governance_hints.get("watch_review_candidates"),
|
||||
"bundle.governance_hints.watch_review_candidates",
|
||||
)
|
||||
top_global_terms = _require_list(bundle.get("top_global_terms"), "bundle.top_global_terms")
|
||||
recent_counts = _recent_count_map([item for item in top_global_terms if isinstance(item, dict)])
|
||||
interest_set, stopword_set, watch_set = _covered_term_sets(bundle)
|
||||
|
||||
suggestions: list[dict[str, Any]] = []
|
||||
seen: set[str] = set(reserved_terms)
|
||||
for item in candidates:
|
||||
if not isinstance(item, dict):
|
||||
continue
|
||||
term = item.get("term")
|
||||
if not isinstance(term, str) or not term.strip():
|
||||
continue
|
||||
term_key = _term_key(term)
|
||||
if term_key in seen or term_key in interest_set or term_key in stopword_set or term_key in watch_set:
|
||||
continue
|
||||
total_count = int(item.get("total_count") or 0)
|
||||
days_seen = int(item.get("days_seen") or 0)
|
||||
recent_count = recent_counts.get(term, 0)
|
||||
base_reason = str(item.get("reason") or "Falls into the configured watch-term review range.")
|
||||
reason = f"{base_reason} Evidence: total_count={total_count}, days_seen={days_seen}, recent_count={recent_count}."
|
||||
suggestions.append(
|
||||
{
|
||||
"term": term,
|
||||
"reason": reason,
|
||||
"total_count": total_count,
|
||||
"days_seen": days_seen,
|
||||
"recent_count": recent_count,
|
||||
}
|
||||
)
|
||||
seen.add(term_key)
|
||||
|
||||
suggestions.sort(key=_sort_key)
|
||||
return suggestions
|
||||
|
||||
|
||||
def _prepare_alias_suggestions(
|
||||
bundle: dict[str, Any],
|
||||
all_terms: list[dict[str, Any]] | None = None,
|
||||
) -> list[dict[str, Any]]:
|
||||
"""
|
||||
Generate alias suggestions using surface-form rules (no LLM).
|
||||
|
||||
Rules:
|
||||
1. casefold match — same normalized form, different original casing
|
||||
2. trailing-s singularization — singular/plural variants
|
||||
3. whitespace/hyphen normalization — word boundary variants
|
||||
|
||||
Scans all_terms (full term_stats) if provided; otherwise falls back
|
||||
to top_global_terms from the bundle.
|
||||
"""
|
||||
current_config = _require_dict(bundle.get("current_config"), "bundle.current_config")
|
||||
interest_keywords = _require_list(
|
||||
current_config.get("interest_keywords"), "bundle.current_config.interest_keywords"
|
||||
)
|
||||
top_global_terms = _require_list(bundle.get("top_global_terms"), "bundle.top_global_terms")
|
||||
source_terms = all_terms if all_terms is not None else top_global_terms
|
||||
|
||||
interest_set = {_term_key(t) for t in interest_keywords if isinstance(t, str)}
|
||||
interest_originals: set[str] = {t for t in interest_keywords if isinstance(t, str)}
|
||||
|
||||
# Build full casefold → [original forms] map
|
||||
cf_map: dict[str, list[str]] = {}
|
||||
for item in source_terms:
|
||||
term = None
|
||||
if isinstance(item, dict):
|
||||
term = item.get("term")
|
||||
elif isinstance(item, str):
|
||||
term = item
|
||||
if not isinstance(term, str) or not term.strip():
|
||||
continue
|
||||
key = _term_key(term)
|
||||
if key not in cf_map:
|
||||
cf_map[key] = []
|
||||
if term not in cf_map[key]:
|
||||
cf_map[key].append(term)
|
||||
|
||||
suggestions: list[dict[str, Any]] = []
|
||||
seen_pairs: set[tuple[str, str]] = set()
|
||||
|
||||
def _add(from_term: str, to_term: str, reason: str) -> None:
|
||||
pair = (_term_key(from_term), _term_key(to_term))
|
||||
if pair in seen_pairs:
|
||||
return
|
||||
seen_pairs.add(pair)
|
||||
suggestions.append({"from": from_term, "to": to_term, "reason": reason})
|
||||
|
||||
# Build a set of all term keys from source for quick lookup
|
||||
source_keys = set(cf_map.keys())
|
||||
|
||||
# Rule 1: casefold match — same normalized form, different casing
|
||||
for key, variants in cf_map.items():
|
||||
if len(variants) < 2:
|
||||
continue
|
||||
canonical = None
|
||||
alt_forms = []
|
||||
for v in variants:
|
||||
if v in interest_originals:
|
||||
canonical = v
|
||||
else:
|
||||
alt_forms.append(v)
|
||||
if canonical and alt_forms:
|
||||
for alt in alt_forms:
|
||||
_add(alt, canonical, "Case variant")
|
||||
elif len(variants) >= 2 and not canonical:
|
||||
# None is canonical — suggest the highest-frequency form
|
||||
ranked = sorted(variants, key=lambda t: -(
|
||||
next(
|
||||
(it.get("total_count", 0) for it in top_global_terms if it.get("term") == t),
|
||||
0,
|
||||
)
|
||||
))
|
||||
for alt in ranked[1:]:
|
||||
_add(alt, ranked[0], "Case variant (auto-ranked)")
|
||||
|
||||
# Rule 2: singular/plural — trailing-s normalization
|
||||
# Check all source terms (not just interest keys) for bidirectional matching
|
||||
for key in source_keys:
|
||||
if key in interest_set:
|
||||
continue
|
||||
if key.endswith("s") and len(key) > 2:
|
||||
singular_key = key.rstrip("s")
|
||||
if singular_key in interest_set and singular_key != key:
|
||||
# Find canonical interest keyword
|
||||
canon = next((t for t in interest_keywords if _term_key(t) == singular_key), None)
|
||||
from_form = cf_map[key][0]
|
||||
if canon:
|
||||
_add(from_form, canon, "Plural variant")
|
||||
# singular form → interest has plural
|
||||
plural_key = key + "s"
|
||||
if plural_key in interest_set and plural_key != key:
|
||||
canon = next((t for t in interest_keywords if _term_key(t) == plural_key), None)
|
||||
from_form = cf_map[key][0]
|
||||
if canon:
|
||||
_add(from_form, canon, "Singular variant")
|
||||
|
||||
# Rule 3: whitespace/hyphen normalization
|
||||
for key in source_keys:
|
||||
if key in interest_set:
|
||||
continue
|
||||
normalized = key.replace("-", "").replace("_", "").replace(" ", "")
|
||||
if normalized in interest_set and normalized != key:
|
||||
canon = next((t for t in interest_keywords if _term_key(t) == normalized), None)
|
||||
from_form = cf_map[key][0]
|
||||
if canon:
|
||||
_add(from_form, canon, "Whitespace/punctuation variant")
|
||||
|
||||
suggestions.sort(key=lambda x: (x["from"].casefold(), x["to"].casefold()))
|
||||
return suggestions
|
||||
|
||||
|
||||
def _render_table(items: list[dict[str, Any]]) -> str:
|
||||
if not items:
|
||||
return "_None in this pass._\n"
|
||||
lines = [
|
||||
"| Term | Total | Days | Recent | Reason |",
|
||||
"| --- | ---: | ---: | ---: | --- |",
|
||||
]
|
||||
for item in items:
|
||||
term = str(item.get("term") or "")
|
||||
total_count = int(item.get("total_count") or 0)
|
||||
days_seen = int(item.get("days_seen") or 0)
|
||||
recent_count = int(item.get("recent_count") or 0)
|
||||
reason = str(item.get("reason") or "").replace("|", "\\|")
|
||||
lines.append(f"| {term} | {total_count} | {days_seen} | {recent_count} | {reason} |")
|
||||
return "\n".join(lines) + "\n"
|
||||
|
||||
|
||||
def _render_simple_table(items: list[dict[str, Any]], first_column: str) -> str:
|
||||
if not items:
|
||||
return "_None in this pass._\n"
|
||||
lines = [
|
||||
f"| {first_column} | Reason |",
|
||||
"| --- | --- |",
|
||||
]
|
||||
for item in items:
|
||||
value = str(item.get(first_column.casefold()) or item.get(first_column) or "")
|
||||
reason = str(item.get("reason") or "").replace("|", "\\|")
|
||||
lines.append(f"| {value} | {reason} |")
|
||||
return "\n".join(lines) + "\n"
|
||||
|
||||
|
||||
def _render_markdown(
|
||||
*,
|
||||
suggestion_date: str,
|
||||
bundle_path: Path,
|
||||
json_output_path: Path,
|
||||
bundle: dict[str, Any],
|
||||
suggestions: dict[str, Any],
|
||||
) -> str:
|
||||
policy = _require_dict(bundle.get("policy"), "bundle.policy")
|
||||
current_config = _require_dict(bundle.get("current_config"), "bundle.current_config")
|
||||
top_global_terms = _require_list(bundle.get("top_global_terms"), "bundle.top_global_terms")
|
||||
uncovered_terms = _require_list(bundle.get("uncovered_terms"), "bundle.uncovered_terms")
|
||||
top_preview = [item for item in top_global_terms if isinstance(item, dict)][:5]
|
||||
uncovered_preview = [item for item in uncovered_terms if isinstance(item, dict)][:5]
|
||||
|
||||
interest_items = suggestions["interest_keyword_suggestions"]
|
||||
watch_items = suggestions["watch_terms"]
|
||||
alias_items = suggestions["alias_suggestions"]
|
||||
stopword_items = suggestions["stopword_suggestions"]
|
||||
|
||||
lines = [
|
||||
f"# Term Cleanup Suggestions - {suggestion_date}",
|
||||
"",
|
||||
"## Review Context",
|
||||
"",
|
||||
f"- Source bundle: `{bundle_path}`",
|
||||
f"- Suggestions JSON: `{json_output_path}`",
|
||||
f"- Bundle generated_at: `{bundle.get('generated_at', 'unknown')}`",
|
||||
f"- Based on days: `{suggestions['based_on_days']}`",
|
||||
f"- Policy schema version: `{policy.get('schema_version', 'unknown')}`",
|
||||
"- Scope: implement `interest_keyword_suggestions` and `watch_terms` main path first; keep alias/stopword conservative in this pass.",
|
||||
"",
|
||||
"## Current State",
|
||||
"",
|
||||
f"- Interest keywords: `{current_config.get('interest_keyword_count', 0)}`",
|
||||
f"- Watch terms: `{current_config.get('watch_term_count', 0)}`",
|
||||
f"- Stopwords: `{current_config.get('stopword_count', 0)}`",
|
||||
f"- Aliases: `{current_config.get('alias_count', 0)}`",
|
||||
f"- Top global terms considered: `{len(top_global_terms)}`",
|
||||
f"- Uncovered terms considered: `{len(uncovered_terms)}`",
|
||||
"",
|
||||
"### Top Terms Snapshot",
|
||||
"",
|
||||
]
|
||||
|
||||
if top_preview:
|
||||
for item in top_preview:
|
||||
lines.append(
|
||||
f"- `{item.get('term', '')}`: total_count={item.get('total_count', 0)}, days_seen={item.get('days_seen', 0)}, recent_count={item.get('recent_count', 0)}"
|
||||
)
|
||||
else:
|
||||
lines.append("- No top terms available.")
|
||||
|
||||
lines.extend([
|
||||
"",
|
||||
"### Uncovered Terms Snapshot",
|
||||
"",
|
||||
])
|
||||
if uncovered_preview:
|
||||
for item in uncovered_preview:
|
||||
lines.append(
|
||||
f"- `{item.get('term', '')}`: total_count={item.get('total_count', 0)}, days_seen={item.get('days_seen', 0)}, recent_count={item.get('recent_count', 0)}"
|
||||
)
|
||||
else:
|
||||
lines.append("- No uncovered terms available.")
|
||||
|
||||
lines.extend([
|
||||
"",
|
||||
"## Suggestion Summary",
|
||||
"",
|
||||
f"- `interest_keyword_suggestions`: `{len(interest_items)}`",
|
||||
f"- `watch_terms`: `{len(watch_items)}`",
|
||||
f"- `alias_suggestions`: `{len(alias_items)}`",
|
||||
f"- `stopword_suggestions`: `{len(stopword_items)}`",
|
||||
"",
|
||||
"## Interest Keyword Suggestions",
|
||||
"",
|
||||
_render_table(interest_items).rstrip(),
|
||||
"",
|
||||
"## Watch Terms",
|
||||
"",
|
||||
_render_table(watch_items).rstrip(),
|
||||
"",
|
||||
"## Alias Suggestions",
|
||||
"",
|
||||
"_Conservative by design in this minimal version; no automatic alias suggestions are emitted yet._" if not alias_items else _render_simple_table(alias_items, "from").rstrip(),
|
||||
"",
|
||||
"## Stopword Suggestions",
|
||||
"",
|
||||
"_Conservative by design in this minimal version; no automatic stopword suggestions are emitted yet._" if not stopword_items else _render_simple_table(stopword_items, "term").rstrip(),
|
||||
"",
|
||||
"## Apply",
|
||||
"",
|
||||
"Review the Markdown first, then selectively apply accepted suggestions with the JSON file.",
|
||||
"",
|
||||
"```bash",
|
||||
f"python scripts/apply_term_suggestions.py \\",
|
||||
f" --suggestions {json_output_path} \\",
|
||||
" --accept-interest \"Claude Code\" \\",
|
||||
" --accept-watch \"A2A\" \\",
|
||||
" --dry-run",
|
||||
"```",
|
||||
"",
|
||||
])
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def _build_output_paths(
|
||||
*,
|
||||
output_dir: Path,
|
||||
suggestion_date: str,
|
||||
json_output: Path | None,
|
||||
markdown_output: Path | None,
|
||||
) -> tuple[Path, Path]:
|
||||
stem = f"term-cleanup-suggestions-{suggestion_date}"
|
||||
resolved_json = json_output or (output_dir / f"{stem}.json")
|
||||
resolved_markdown = markdown_output or (output_dir / f"{stem}.md")
|
||||
return resolved_json, resolved_markdown
|
||||
|
||||
|
||||
def main() -> None:
|
||||
parser = argparse.ArgumentParser(description="Generate term cleanup suggestions JSON and Markdown from review bundle.")
|
||||
parser.add_argument("--bundle", type=Path, default=DEFAULT_BUNDLE_PATH, help="Review bundle JSON file")
|
||||
parser.add_argument(
|
||||
"--output-dir",
|
||||
type=Path,
|
||||
default=DEFAULT_OUTPUT_DIR,
|
||||
help="Directory for generated suggestions outputs when explicit output paths are not provided",
|
||||
)
|
||||
parser.add_argument("--date", type=str, default=None, help="Override suggestions date (YYYY-MM-DD)")
|
||||
parser.add_argument("--json-output", type=Path, default=None, help="Explicit suggestions JSON output path")
|
||||
parser.add_argument("--markdown-output", type=Path, default=None, help="Explicit suggestions Markdown output path")
|
||||
parser.add_argument(
|
||||
"--emit-markdown",
|
||||
action="store_true",
|
||||
help="Also write the human-readable Markdown review draft. JSON suggestions are always written.",
|
||||
)
|
||||
args = parser.parse_args()
|
||||
|
||||
if not args.bundle.exists():
|
||||
raise RuntimeError(f"Bundle file not found: {args.bundle}")
|
||||
|
||||
bundle = _require_dict(_load_json(args.bundle), "bundle")
|
||||
days = bundle.get("days")
|
||||
if not isinstance(days, int):
|
||||
raise RuntimeError("bundle.days must be an integer.")
|
||||
|
||||
suggestion_date = args.date or _bundle_date(bundle)
|
||||
json_output_path, markdown_output_path = _build_output_paths(
|
||||
output_dir=args.output_dir,
|
||||
suggestion_date=suggestion_date,
|
||||
json_output=args.json_output,
|
||||
markdown_output=args.markdown_output,
|
||||
)
|
||||
|
||||
# Load full term_stats for alias scanning (bundle only has top N)
|
||||
stats_path = REPO_ROOT / "data" / "term_index" / "term_stats.json"
|
||||
all_stats_terms: list[str] = []
|
||||
if stats_path.exists():
|
||||
stats_payload = _load_json(stats_path)
|
||||
raw_terms = stats_payload.get("terms") if isinstance(stats_payload, dict) else []
|
||||
if isinstance(raw_terms, list):
|
||||
all_stats_terms = [str(t["term"]) for t in raw_terms if isinstance(t, dict) and isinstance(t.get("term"), str)]
|
||||
|
||||
interest_items = _prepare_interest_suggestions(bundle)
|
||||
reserved_terms = {_term_key(str(item.get("term") or "")) for item in interest_items}
|
||||
watch_items = _prepare_watch_suggestions(bundle, reserved_terms=reserved_terms)
|
||||
alias_items = _prepare_alias_suggestions(bundle, all_terms=all_stats_terms)
|
||||
|
||||
suggestions = {
|
||||
"date": suggestion_date,
|
||||
"based_on_days": days,
|
||||
"source_bundle": str(args.bundle),
|
||||
"policy_schema_version": _require_dict(bundle.get("policy"), "bundle.policy").get("schema_version", "unknown"),
|
||||
"summary": {
|
||||
"interest_keyword_suggestions": len(interest_items),
|
||||
"watch_terms": len(watch_items),
|
||||
"alias_suggestions": len(alias_items),
|
||||
"stopword_suggestions": 0,
|
||||
},
|
||||
"alias_suggestions": alias_items,
|
||||
"stopword_suggestions": [],
|
||||
"interest_keyword_suggestions": interest_items,
|
||||
"watch_terms": watch_items,
|
||||
}
|
||||
markdown = _render_markdown(
|
||||
suggestion_date=suggestion_date,
|
||||
bundle_path=args.bundle,
|
||||
json_output_path=json_output_path,
|
||||
bundle=bundle,
|
||||
suggestions=suggestions,
|
||||
)
|
||||
|
||||
_save_json(json_output_path, suggestions)
|
||||
if args.emit_markdown:
|
||||
_save_text(markdown_output_path, markdown)
|
||||
|
||||
summary = {
|
||||
"bundle": str(args.bundle),
|
||||
"date": suggestion_date,
|
||||
"json_output": str(json_output_path),
|
||||
"markdown_output": str(markdown_output_path) if args.emit_markdown else None,
|
||||
"interest_keyword_suggestions": len(interest_items),
|
||||
"watch_terms": len(watch_items),
|
||||
"alias_suggestions": len(alias_items),
|
||||
"stopword_suggestions": 0,
|
||||
"emit_markdown": args.emit_markdown,
|
||||
}
|
||||
print(json.dumps(summary, ensure_ascii=False, indent=2))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -37,7 +37,7 @@ def main() -> None:
|
||||
help="Directory to write per-article Markdown summaries",
|
||||
)
|
||||
parser.add_argument("--max-retries", type=int, default=2, help="Maximum LLM retry attempts per article")
|
||||
parser.add_argument("--timeout", type=float, default=60.0, help="LLM request timeout in seconds")
|
||||
parser.add_argument("--timeout", type=float, default=120.0, help="LLM request timeout in seconds")
|
||||
parser.add_argument("--api-key", type=str, default=None, help="Override article-summary LLM API key")
|
||||
parser.add_argument("--model", type=str, default=None, help="Override article-summary LLM model")
|
||||
parser.add_argument("--api-url", type=str, default=None, help="Override article-summary LLM API URL/base URL")
|
||||
|
||||
@@ -0,0 +1,24 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
SRC_ROOT = REPO_ROOT / "src"
|
||||
|
||||
if str(SRC_ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(SRC_ROOT))
|
||||
|
||||
from summary_mcp.runtime.article_summary_jobs import run_article_summary_job
|
||||
|
||||
|
||||
def main() -> None:
|
||||
parser = argparse.ArgumentParser(description="Run a background article-summary job by job_id.")
|
||||
parser.add_argument("--job-id", required=True, help="Article summary job id")
|
||||
args = parser.parse_args()
|
||||
run_article_summary_job(job_id=args.job_id)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,24 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
SRC_ROOT = REPO_ROOT / "src"
|
||||
|
||||
if str(SRC_ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(SRC_ROOT))
|
||||
|
||||
from summary_mcp.runtime.freshrss_pipeline_jobs import run_freshrss_pipeline_job
|
||||
|
||||
|
||||
def main() -> None:
|
||||
parser = argparse.ArgumentParser(description="Run a background FreshRSS pipeline job by job_id.")
|
||||
parser.add_argument("--job-id", required=True, help="FreshRSS pipeline job id")
|
||||
args = parser.parse_args()
|
||||
run_freshrss_pipeline_job(job_id=args.job_id)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,24 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
SRC_ROOT = REPO_ROOT / "src"
|
||||
|
||||
if str(SRC_ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(SRC_ROOT))
|
||||
|
||||
from summary_mcp.runtime.resume_jobs import run_resume_job
|
||||
|
||||
|
||||
def main() -> None:
|
||||
parser = argparse.ArgumentParser(description="Run a background resume job by job_id.")
|
||||
parser.add_argument("--job-id", required=True, help="Resume job id")
|
||||
args = parser.parse_args()
|
||||
run_resume_job(job_id=args.job_id)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -1,15 +1,32 @@
|
||||
---
|
||||
name: keyword-cleanup-review
|
||||
description: 审查和整理本仓库的每日关键词索引和频率统计。当用户需要检查 `data/term_index/term_stats.json`、最近的 `data/term_index/daily/*.json`、`configs/term_aliases.json`、`configs/term_stopwords.json` 或 `configs/filter_context.personal.json`,以提议别名合并、停用词、关注词或 `interest_keywords` 更新(但不直接修改配置)时使用。
|
||||
description: 生成 reader 项目的正式关键词 review 输入。当用户需要基于 `data/term_index/term_stats.json`、最近的 `data/term_index/daily/*.json` 和当前 configs 产出 bundle、正式 suggestions JSON,或按需生成审阅 Markdown 供 OpenClaw 汇报和等待确认时使用。不要用于低频清理 review 目录、删除旧产物或直接 apply 配置。
|
||||
---
|
||||
|
||||
# 关键词清理审查
|
||||
|
||||
使用此技能将仓库的关键词统计转化为可审查的清理建议。
|
||||
使用此技能将仓库的关键词统计转化为**正式 review 输入**,供 OpenClaw 后续做汇报、确认和 apply 编排。
|
||||
|
||||
## 角色边界
|
||||
|
||||
这个 skill 负责:
|
||||
|
||||
- 构建 review bundle
|
||||
- 生成正式 suggestions JSON
|
||||
- 按需生成人工审阅 Markdown
|
||||
- 给 OpenClaw 提供稳定的关键词 review 输入
|
||||
|
||||
这个 skill 不负责:
|
||||
|
||||
- 清理 `outputs/term_index/review/` 下的旧文件
|
||||
- 决定删除哪些历史 bundle / suggestions / markdown
|
||||
- 直接 apply `configs/term_aliases.json` / `configs/term_stopwords.json` / `configs/filter_context.personal.json`
|
||||
|
||||
低频维护、清理和 dry-run 校验应由 OpenClaw 侧 maintenance SOP 处理,而不是由本 skill 承担。
|
||||
|
||||
## 工作流程
|
||||
|
||||
1. 构建精简的审查数据包:
|
||||
### Phase 1:构建审查数据包
|
||||
|
||||
```bash
|
||||
python skills/keyword-cleanup-review/scripts/build_review_bundle.py
|
||||
@@ -17,25 +34,89 @@ python skills/keyword-cleanup-review/scripts/build_review_bundle.py
|
||||
|
||||
可选参数:
|
||||
|
||||
- `--days 7`
|
||||
- `--top 50`
|
||||
- `--days 7`(默认 7,建议传 365 覆盖全量)
|
||||
- `--top 100`(考虑的词数)
|
||||
- `--output outputs/term_index/review/keyword-cleanup-bundle.json`
|
||||
|
||||
2. 阅读生成的数据包和建议模式:
|
||||
#### 候选引擎策略
|
||||
|
||||
- `outputs/term_index/review/keyword-cleanup-bundle.json`
|
||||
- `skills/keyword-cleanup-review/references/suggestion-schema.md`
|
||||
根据 `configs/term_cleanup_policy.json` 的 `schema_version` 自动切换:
|
||||
|
||||
3. 生成两份输出:
|
||||
| 版本 | 策略 | 说明 |
|
||||
|------|------|------|
|
||||
| v1(旧) | 固定阈值(total≥3/days≥2 → interest) | 小数据集兼容 |
|
||||
| v2(当前默认) | 百分位排名 + 增速因子 | 自适应数据量,不需要手工调阈值 |
|
||||
|
||||
- 一份简短的供人工审阅的 Markdown 报告
|
||||
- 一份符合模式的 JSON 建议文件
|
||||
v2 策略说明:
|
||||
- **percentile**:total_count 在所有词里的排位占比。top 5% → interest 候选,5%-20% → watch 候选
|
||||
- **growth**:recent_count / total_count,衡量近期活跃度。growth≥0.5 的排位外词也会主动推荐
|
||||
|
||||
4. 严格保持边界:
|
||||
### Phase 2:生成建议(规则层)
|
||||
|
||||
```bash
|
||||
python scripts/generate_term_cleanup_suggestions.py \
|
||||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json
|
||||
```
|
||||
|
||||
如需人工审阅展示稿:
|
||||
|
||||
```bash
|
||||
python scripts/generate_term_cleanup_suggestions.py \
|
||||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
|
||||
--emit-markdown
|
||||
```
|
||||
|
||||
#### 产出能力
|
||||
|
||||
| 建议类型 | 状态 | 方法 |
|
||||
|---------|------|------|
|
||||
| interest 建议 | ✅ 已实现 | 百分位 top 5% + 增速促活 |
|
||||
| watch 建议 | ✅ 已实现 | 百分位 5%-20% |
|
||||
| alias 建议 | ✅ 已实现 | 规则层:大小写归一、单复数、去空格/连字符 |
|
||||
| stopword 建议 | ❌ 规则层空缺 | 见 Phase 3(LLM 层) |
|
||||
|
||||
默认生成:
|
||||
- `term-cleanup-suggestions-YYYY-MM-DD.json`(正式建议产物)
|
||||
|
||||
显式加 `--emit-markdown` 额外生成:
|
||||
- `term-cleanup-suggestions-YYYY-MM-DD.md`(临时展示稿)
|
||||
|
||||
### Phase 3:生成建议(LLM 层,可选)
|
||||
|
||||
规则层覆盖不了 alias(中英文对应、缩写展开、同义不同名)和 stopword 判断,需要 LLM 辅助:
|
||||
|
||||
```bash
|
||||
python scripts/generate_term_cleanup_semantic_suggestions.py \
|
||||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
|
||||
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
|
||||
--output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json
|
||||
```
|
||||
|
||||
从 `.env` 读取 LLM 配置(`LLM_API_URL` / `LLM_MODEL` / `LLM_API_KEY`),使用 DeepSeek API。
|
||||
|
||||
输出三部分:
|
||||
|
||||
| 输出 | 说明 |
|
||||
|------|------|
|
||||
| `semantic_alias` | 语义级别名(中英文、缩写、同义不同名) |
|
||||
| `stopword` | 泛词过滤建议(规则层做不了的需要语义判断的) |
|
||||
| `promote_to_interest` | 与用户关注方向一致的新词,建议加入 interest |
|
||||
|
||||
**注:LLM 层产物是候选,不应自动 apply,需要人工确认后由 OpenClaw 编排 apply。**
|
||||
|
||||
### Phase 4:输出给 OpenClaw 编排
|
||||
|
||||
- `suggestions JSON` = review / apply 之间唯一正式建议输入
|
||||
- `semantic-suggestions JSON` = LLM 补充建议,需要人工筛选后合并到 suggestions JSON 再 apply
|
||||
- Markdown = 临时展示层
|
||||
- 后续汇报、确认、dry-run、apply、收尾清理由 OpenClaw 编排层执行
|
||||
|
||||
### Phase 5:严格保持边界
|
||||
|
||||
- 建议 `configs/term_aliases.json` 的修改
|
||||
- 建议 `configs/term_stopwords.json` 的修改
|
||||
- 建议 `configs/filter_context.personal.json` 的新增
|
||||
- **LLM 层产出(semantic-suggestions)不自动 apply**,需人工确认后由 OpenClaw 编排层执行
|
||||
- 除非用户明确要求,否则不要直接编辑这些文件
|
||||
- 除非用户要求修改规则逻辑,否则不要建议直接编辑 `configs/filter_rules.json`
|
||||
|
||||
@@ -78,10 +159,47 @@ Markdown 输出应:
|
||||
- 分类别名、停用词、兴趣关键词和关注词建议
|
||||
- 用简短、具体的句子解释理由
|
||||
|
||||
说明:Markdown 主要用于人工临时审阅,不必默认当作长期资产保留。
|
||||
|
||||
JSON 输出应遵循:
|
||||
|
||||
- `references/suggestion-schema.md`
|
||||
|
||||
说明:JSON 是 review / apply 之间的唯一正式建议产物,应优先保留。
|
||||
|
||||
## 产物口径
|
||||
|
||||
长期保留:
|
||||
|
||||
- `data/term_index/daily/*.json`
|
||||
- `data/term_index/term_stats.json`
|
||||
- `configs/filter_context.personal.json`
|
||||
- `configs/term_watchlist.json`
|
||||
- `configs/term_aliases.json`
|
||||
- `configs/term_stopwords.json`
|
||||
- `configs/term_change_log.json`
|
||||
|
||||
短期保留:
|
||||
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`
|
||||
- `outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json`
|
||||
|
||||
临时产物:
|
||||
|
||||
- `outputs/term_index/review/keyword-cleanup-bundle.json`
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`
|
||||
|
||||
默认执行口径:
|
||||
|
||||
- bundle 只作为运行时工作文件,默认只保留当前最新一份
|
||||
- Markdown 只作为人工展示层,优先按需生成,不默认长期归档
|
||||
- JSON suggestions 是 review / apply 之间唯一正式建议输入
|
||||
|
||||
说明:
|
||||
|
||||
- “是否删除旧 bundle / 旧 markdown / 旧 suggestions” 不属于本 skill 的正式职责
|
||||
- 这类维护动作应由 OpenClaw 侧的 maintenance skill 处理
|
||||
|
||||
## 仓库说明
|
||||
|
||||
当前仓库行为:
|
||||
@@ -97,6 +215,9 @@ JSON 输出应遵循:
|
||||
## 资源
|
||||
|
||||
- 脚本:
|
||||
- `scripts/build_review_bundle.py`
|
||||
- `skills/keyword-cleanup-review/scripts/build_review_bundle.py`
|
||||
- `scripts/generate_term_cleanup_suggestions.py`
|
||||
- `scripts/generate_term_cleanup_semantic_suggestions.py`(LLM 层)
|
||||
- 参考文档:
|
||||
- `references/suggestion-schema.md`
|
||||
- `plans/keyword-cleanup-interest-watch-engine-improvement.md`(v2 引擎设计)
|
||||
|
||||
@@ -9,16 +9,15 @@ from typing import Any
|
||||
|
||||
|
||||
DEFAULT_POLICY: dict[str, Any] = {
|
||||
"schema_version": "v1",
|
||||
"schema_version": "v2",
|
||||
"interest_keyword_review": {
|
||||
"min_total_count": 3,
|
||||
"min_days_seen": 2,
|
||||
"percentile_min": 0.0,
|
||||
"percentile_max": 0.05,
|
||||
"growth_promotion": 0.5,
|
||||
},
|
||||
"watch_term_review": {
|
||||
"min_total_count": 1,
|
||||
"min_days_seen": 1,
|
||||
"max_total_count": 2,
|
||||
"max_days_seen": 2,
|
||||
"percentile_min": 0.05,
|
||||
"percentile_max": 0.20,
|
||||
},
|
||||
"alias_review": {
|
||||
"min_total_count": 2,
|
||||
@@ -28,6 +27,11 @@ DEFAULT_POLICY: dict[str, Any] = {
|
||||
"max_total_count": 2,
|
||||
"max_days_seen": 2,
|
||||
},
|
||||
"notes": [
|
||||
"v2: interest/watch 使用百分位排名 + 增速因子替代固定阈值",
|
||||
"percentile 越小表示排名越高(top 5% = percentile 0.05)",
|
||||
"growth = recent_count / total_count,衡量近期活跃度",
|
||||
],
|
||||
}
|
||||
|
||||
|
||||
@@ -126,6 +130,37 @@ def _within_watch_thresholds(item: dict[str, Any], thresholds: dict[str, Any]) -
|
||||
)
|
||||
|
||||
|
||||
def _compute_percentile(value: int, sorted_values: list[int]) -> float:
|
||||
"""
|
||||
Return the percentile rank of `value` in `sorted_values` (ascending).
|
||||
0.0 = highest frequency (top rank), 1.0 = lowest frequency (bottom rank).
|
||||
"""
|
||||
if not sorted_values:
|
||||
return 1.0
|
||||
# bisect_left — count of values strictly less than `value`
|
||||
lo, hi = 0, len(sorted_values)
|
||||
while lo < hi:
|
||||
mid = (lo + hi) // 2
|
||||
if sorted_values[mid] < value:
|
||||
lo = mid + 1
|
||||
else:
|
||||
hi = mid
|
||||
rank = lo
|
||||
# invert: smallest value → rank=0 → 1.0 (bottom)
|
||||
# largest value → rank=len → 0.0 (top)
|
||||
return 1.0 - (rank / len(sorted_values))
|
||||
|
||||
|
||||
def _compute_growth(recent_count: int, total_count: int) -> float:
|
||||
"""
|
||||
Return growth factor: recent_count / total_count.
|
||||
Only meaningful when total_count >= 3; returns 0.0 for small counts.
|
||||
"""
|
||||
if total_count < 3:
|
||||
return 0.0
|
||||
return recent_count / total_count
|
||||
|
||||
|
||||
def main() -> None:
|
||||
parser = argparse.ArgumentParser(
|
||||
description="Build a compact review bundle for the keyword-cleanup-review skill."
|
||||
@@ -235,6 +270,13 @@ def main() -> None:
|
||||
alias_values = _casefold_set(list(aliases.values()))
|
||||
watch_set = _casefold_set([str(item.get("term", "")) for item in watchlist])
|
||||
|
||||
# Build a sorted list of all total_counts for percentile computation
|
||||
all_total_counts = sorted(
|
||||
int(item.get("total_count") or 0)
|
||||
for item in stats_terms
|
||||
if isinstance(item, dict) and isinstance(item.get("term"), str)
|
||||
)
|
||||
|
||||
top_global_terms = []
|
||||
for item in stats_terms[: args.top]:
|
||||
if not isinstance(item, dict):
|
||||
@@ -256,42 +298,108 @@ def main() -> None:
|
||||
"is_alias_target": folded in alias_values,
|
||||
"in_watchlist": folded in watch_set,
|
||||
"recent_count": recent_counter.get(term, 0),
|
||||
"percentile": _compute_percentile(
|
||||
int(item.get("total_count") or 0), all_total_counts
|
||||
),
|
||||
"growth": _compute_growth(
|
||||
recent_counter.get(term, 0),
|
||||
int(item.get("total_count") or 0),
|
||||
),
|
||||
}
|
||||
)
|
||||
|
||||
# Keep more uncovered terms for percentile-based selection
|
||||
uncovered_terms = [
|
||||
item for item in top_global_terms if not item["in_interest_keywords"] and not item["is_stopword"]
|
||||
][:20]
|
||||
][:100]
|
||||
|
||||
policy_version = (policy.get("schema_version") if isinstance(policy, dict) else None) or "v1"
|
||||
interest_thresholds = policy.get("interest_keyword_review") if isinstance(policy, dict) else {}
|
||||
watch_thresholds = policy.get("watch_term_review") if isinstance(policy, dict) else {}
|
||||
|
||||
if policy_version == "v2" or "percentile_max" in interest_thresholds:
|
||||
# v2: percentile + growth based selection
|
||||
pct_min_interest = float(interest_thresholds.get("percentile_min", 0.0))
|
||||
pct_max_interest = float(interest_thresholds.get("percentile_max", 0.05))
|
||||
growth_promo = float(interest_thresholds.get("growth_promotion", 0.5))
|
||||
pct_min_watch = float(watch_thresholds.get("percentile_min", 0.05))
|
||||
pct_max_watch = float(watch_thresholds.get("percentile_max", 0.20))
|
||||
|
||||
interest_candidates_raw = [
|
||||
item for item in uncovered_terms
|
||||
if not item["in_watchlist"]
|
||||
and pct_min_interest <= item["percentile"] <= pct_max_interest
|
||||
]
|
||||
watch_candidates_raw = [
|
||||
item for item in uncovered_terms
|
||||
if not item["in_watchlist"]
|
||||
and pct_min_watch < item["percentile"] <= pct_max_watch
|
||||
]
|
||||
# Growth boost: terms outside watch range but with strong growth signal
|
||||
growth_boost_candidates = [
|
||||
item for item in uncovered_terms
|
||||
if not item["in_watchlist"]
|
||||
and item["percentile"] > pct_max_watch
|
||||
and item["growth"] >= growth_promo
|
||||
]
|
||||
else:
|
||||
# v1 fallback: fixed thresholds
|
||||
interest_candidates_raw = [
|
||||
item for item in uncovered_terms
|
||||
if not item["in_watchlist"] and _meets_min_thresholds(item, interest_thresholds)
|
||||
]
|
||||
watch_candidates_raw = [
|
||||
item for item in uncovered_terms
|
||||
if not item["in_watchlist"]
|
||||
and not _meets_min_thresholds(item, interest_thresholds)
|
||||
and _within_watch_thresholds(item, watch_thresholds)
|
||||
]
|
||||
growth_boost_candidates = []
|
||||
|
||||
interest_review_candidates = [
|
||||
{
|
||||
"term": item["term"],
|
||||
"total_count": item["total_count"],
|
||||
"days_seen": item["days_seen"],
|
||||
"percentile": item["percentile"],
|
||||
"growth": item["growth"],
|
||||
"reason": (
|
||||
"Meets the configured interest-keyword review threshold and is not yet covered "
|
||||
"by interest keywords or stopwords."
|
||||
f"top {item['percentile']:.1%} by frequency,"
|
||||
f"growth={item['growth']:.0%},"
|
||||
"not yet covered by interest keywords or stopwords."
|
||||
),
|
||||
}
|
||||
for item in uncovered_terms
|
||||
if not item["in_watchlist"] and _meets_min_thresholds(item, interest_thresholds)
|
||||
for item in interest_candidates_raw
|
||||
][:20]
|
||||
watch_review_candidates = [
|
||||
{
|
||||
"term": item["term"],
|
||||
"total_count": item["total_count"],
|
||||
"days_seen": item["days_seen"],
|
||||
"percentile": item["percentile"],
|
||||
"growth": item["growth"],
|
||||
"reason": (
|
||||
"Falls into the configured watch-term review range and should be observed "
|
||||
"before promotion into interest keywords."
|
||||
f"top {item['percentile']:.1%} by frequency,"
|
||||
f"growth={item['growth']:.0%},"
|
||||
"fell into watch-review range."
|
||||
),
|
||||
}
|
||||
for item in uncovered_terms
|
||||
if not item["in_watchlist"]
|
||||
and not _meets_min_thresholds(item, interest_thresholds)
|
||||
and _within_watch_thresholds(item, watch_thresholds)
|
||||
for item in watch_candidates_raw
|
||||
][:20]
|
||||
growth_boost_review_items = [
|
||||
{
|
||||
"term": item["term"],
|
||||
"total_count": item["total_count"],
|
||||
"days_seen": item["days_seen"],
|
||||
"percentile": item["percentile"],
|
||||
"growth": item["growth"],
|
||||
"reason": (
|
||||
f"growth spike: {item['growth']:.0%} of occurrences in recent window "
|
||||
f"(total={item['total_count']}, days={item['days_seen']})."
|
||||
),
|
||||
}
|
||||
for item in growth_boost_candidates
|
||||
][:5]
|
||||
recent_hot_terms = sorted(
|
||||
({"term": term, "recent_count": count} for term, count in recent_counter.items()),
|
||||
key=lambda item: (-item["recent_count"], item["term"].casefold(), item["term"]),
|
||||
@@ -330,6 +438,7 @@ def main() -> None:
|
||||
"governance_hints": {
|
||||
"interest_review_candidates": interest_review_candidates,
|
||||
"watch_review_candidates": watch_review_candidates,
|
||||
"growth_boost_review_items": growth_boost_review_items,
|
||||
},
|
||||
}
|
||||
_save_json(args.output, bundle)
|
||||
|
||||
@@ -0,0 +1,186 @@
|
||||
---
|
||||
name: reader-digest-flow
|
||||
description: 编排 reader 项目的端到端 AI 日报流程。仅在用户要求运行/重跑日报、汇报候选、发布 Hugo 日报、沉淀选中文章或继续已有日报任务时使用;覆盖异步 MCP 任务、候选确认、发布、单篇摘要和 IMA 知识库上传。不要因验证、排障冲动或候选质量不佳自行重跑。
|
||||
---
|
||||
|
||||
# Reader Digest Flow
|
||||
|
||||
## 职责边界
|
||||
|
||||
本 Skill 负责:
|
||||
|
||||
- 通过 reader MCP 启动、观察和恢复日报任务;
|
||||
- 向用户展示候选并保持稳定编号;
|
||||
- 根据用户选择生成并发布 Hugo 日报;
|
||||
- 对用户选中的文章生成知识笔记并编排 IMA 上传;
|
||||
- 在每个副作用边界执行确认和结果验证。
|
||||
|
||||
本 Skill 不负责:
|
||||
|
||||
- 实现 reader 内部抓取、摘要、过滤或恢复逻辑;
|
||||
- 通过手拼目录推导 Run 状态或 Artifact;
|
||||
- 未经用户要求自行重跑 Pipeline;
|
||||
- 未经用户确认发布日报或写入知识库;
|
||||
- 直接维护关键词配置;关键词治理委托给 `keyword-cleanup-review`。
|
||||
|
||||
## 核心规则
|
||||
|
||||
1. **只按用户指令运行。** 只有用户明确要求“跑日报”“重新跑”“再跑一次”时才启动新 Pipeline。验证、解释排序和排障默认读取已有 Run。
|
||||
2. **一次对话绑定一个当前 Run。** 以异步 Job 结果返回的 `run_id` 为稳定句柄;新 Run 产生新的候选编号体系,不混用历史编号。
|
||||
3. **状态以 MCP 返回为准。** Agent 只根据顶层 `status` 和 `recommended_action` 分支;`status_source`、`state_conflict` 仅用于解释。
|
||||
4. **路径以返回值为准。** 使用 `output_dir`、`artifact.path`、`delivery_output`、`report_output` 和 `written_paths`;不要根据 `run_id` 手拼 `outputs/...`。
|
||||
5. **候选编号保持稳定。** 用户编号永远对应当前候选列表的原始顺序(1-based);跨产物读取详情时按 URL 或完整 `item_id` 关联,不按数组位置关联。
|
||||
6. **副作用必须授权。** 用户确认 Hugo 文章后才能发布;用户确认 IMA 文章后才能生成并上传知识笔记。
|
||||
7. **内容必须有来源。** 日报和知识笔记只能基于当前 Run 的 `article.plain_text`、摘要、highlights 等 Artifact;不得使用通用知识补写原文没有的信息,也不为满足长度而扩写。
|
||||
|
||||
## 默认生产参数
|
||||
|
||||
用户未显式覆盖时使用:
|
||||
|
||||
```json
|
||||
{
|
||||
"limit": 7,
|
||||
"include_read": false,
|
||||
"mark_read": true,
|
||||
"debug_artifacts": false,
|
||||
"timeout_seconds": 60,
|
||||
"max_retries": 2
|
||||
}
|
||||
```
|
||||
|
||||
- 默认只处理未读文章。
|
||||
- `include_read=true` 仅在用户明确要求扩大到已读内容时使用。
|
||||
- debug/test/validation 才允许 `mark_read=false` 或 `debug_artifacts=true`。
|
||||
- 不随机生成文章数量;用户指定 `limit` 时按用户值执行。
|
||||
|
||||
## 正式流程
|
||||
|
||||
### Phase 1:启动并观察日报 Job
|
||||
|
||||
正式生产入口统一为异步 MCP:
|
||||
|
||||
1. 调用 `start_freshrss_pipeline_job`;
|
||||
2. 轮询 `get_freshrss_pipeline_job_status`;
|
||||
3. `status=success` 后调用 `get_freshrss_pipeline_job_result`;
|
||||
4. 保存返回的 `run_id` 和 Artifact 路径;
|
||||
5. 使用 `get_run_status`、`get_delivery_payload`、`get_run_report` 读取业务状态和结果。
|
||||
|
||||
失败时:
|
||||
|
||||
1. 使用 `get_run_status(run_id)` 读取关联 Run;
|
||||
2. 调用 `inspect_resume_plan(run_id)`;
|
||||
3. `recommended_action=resume` 时启动并轮询异步 Resume Job;
|
||||
4. `recommended_action=read_terminal_result` 时直接读取已有终态结果;
|
||||
5. `recommended_action=start_new_run` 时停止并向用户报告,不自行新建 Run。
|
||||
|
||||
CLI 仅用于 MCP 不可用时的 fallback、debug 或人工排障,不是默认生产入口。具体调用序列见 `references/flow.md`。
|
||||
|
||||
### Phase 2:汇报候选
|
||||
|
||||
- 使用当前 Run 返回的 Delivery Payload 或 digest brief Artifact;
|
||||
- 按候选原始顺序从 1 编号,状态可显示为“已入选/待确认”,但不得重新分组编号;
|
||||
- 每篇提供标题、来源、2-3 句摘要和筛选理由,避免原始 JSON dump;
|
||||
- 用户质疑编号或排序时读取当前 Run 产物核对,不重新运行 Pipeline;
|
||||
- 需要跨 Artifact 取详情时按 URL 或完整 `item_id` 交叉验证。
|
||||
|
||||
Feishu 输出不要使用 Markdown 表格,见 `references/feishu-format-notes.md`。
|
||||
|
||||
### Phase 3:等待 Hugo 选择
|
||||
|
||||
- 等待用户明确选择要发布的文章;
|
||||
- 用户编号映射到当前候选列表,不映射到 extracted 文件序号;
|
||||
- 用户拒绝发布时立即停止当日日报后续流程,不劝说、不自动换一批;
|
||||
- 用户明确要求重跑时才创建新 Run,并重新建立编号体系。
|
||||
|
||||
### Phase 4:生成并发布 Hugo 日报
|
||||
|
||||
发布前读取 `references/public-digest-example.md`,按其最终页面结构生成:
|
||||
|
||||
- `今日概览`
|
||||
- `今日重点`
|
||||
- `趋势观察`
|
||||
|
||||
每篇 `今日重点` 文章末尾必须添加 `来源:[来源名](原文 URL)`,来源链接跟随对应文章,不再生成独立的 `延伸阅读` 章节或重复链接。
|
||||
|
||||
仅发布用户在 Phase 3 选中的文章。公开页面不得出现 `keep/review/drop`、候选、待确认等内部状态。
|
||||
|
||||
写入 Hugo 后执行部署,并验证首页、日报列表页和当日详情页均可访问。命令和检查项见 `references/flow.md`。
|
||||
|
||||
### Phase 5:等待 IMA 选择
|
||||
|
||||
Hugo 发布选择与知识沉淀选择相互独立。询问用户哪些文章值得长期保存:
|
||||
|
||||
- 仅处理用户明确选择的文章;
|
||||
- 不把整份日报上传到 IMA;
|
||||
- 本次选择本身即授权后续单篇摘要和 IMA 上传,不重复确认。
|
||||
|
||||
### Phase 6:生成单篇知识笔记
|
||||
|
||||
对每篇选中文章:
|
||||
|
||||
1. 通过 URL/完整 `item_id` 找到对应 extracted Artifact;
|
||||
2. 使用 `start_article_summary_job(extracted_path=<返回的 Artifact 路径>, selected_ids=[<完整 item_id>])`;
|
||||
3. 轮询 `get_article_summary_job_status`,成功后读取 `get_article_summary_job_result`;
|
||||
4. 只使用现有 `article.plain_text`,不重新抓取原 URL;
|
||||
5. 使用返回的 `written_paths` 定位结果并检查 Markdown 内容。
|
||||
6. **`extracted_path` 必须传绝对路径**(前缀 `/home/ubuntu/zhu/github/reader/`):summary-mcp 工作目录是 `/root/.hermes`,相对路径会报 `extracted_path does not exist`。同一工具连续 3 次失败会触发 MCP 冷却(约 45-60s,报 `MCP server 'reader' is unreachable`),等待冷却后再重试,不要循环重试同一调用。
|
||||
|
||||
异步 MCP 不可用时才使用项目 CLI fallback。不要因为内容较短而引入原文之外的知识。
|
||||
|
||||
### Phase 7:上传到 IMA
|
||||
|
||||
上传前按需读取:
|
||||
|
||||
- 格式规则:`references/ima-format-quickref.md`
|
||||
- API 和上传步骤:`references/ima-upload-api.md`(含 `-200` 版本拦截修复)
|
||||
- 凭证定位:`references/ima-credential-chain.md`
|
||||
- 批量上传脚本:`scripts/ima_upload_one.py`(Python 编排,规避中文文件名 bash 引号问题)
|
||||
|
||||
硬规则:
|
||||
|
||||
- 使用完整文章标题作为文件名;
|
||||
- 以 Markdown 文件 `media_type=7` 上传到 `daily` knowledge base;
|
||||
- 保留原文 URL 和 Category,保证来源可追踪;
|
||||
- 不使用 URL 导入或 Notes 类型替代知识库文件;
|
||||
- 上传后验证目标知识库中存在对应条目;
|
||||
- 失败时报告具体阶段,不无限重试。
|
||||
|
||||
## 关键词治理路由
|
||||
|
||||
只有用户明确要求“清理关键词”“词库治理”等操作时才触发关键词治理。Review bundle 和 Suggestions 生成委托给 `keyword-cleanup-review`,确认与 Apply 仍由当前编排层负责:
|
||||
|
||||
1. 生成 review bundle;
|
||||
2. 生成规则层与可选语义层 Suggestions JSON;
|
||||
3. 等待人工确认;
|
||||
4. dry-run 后按精确 accept 参数 Apply。
|
||||
|
||||
`reader-digest-flow` 不直接编辑 `term_aliases.json`、`term_stopwords.json` 或兴趣配置。简要路由见 `references/keyword-engine-maintenance.md`。
|
||||
|
||||
## 环境坑位(本机部署)
|
||||
|
||||
- **MCP 相对路径陷阱**:`summary-mcp` 服务工作目录是 `/root/.hermes`,不是 reader 项目根。MCP 返回的 `output_dir`/`artifact.path` 是相对路径,直接传给 `start_article_summary_job(extracted_path=...)` 会报 `extracted_path does not exist`。传入前必须拼绝对路径前缀 `/home/ubuntu/zhu/github/reader/`。
|
||||
- **提取失败不等于运行失败**:`status_counts.extract_failed` 的条目(`CONTENT_EXTRACTION_FAILED`,`retryable=false`)跳过即可并如实汇报;失败文章常是推广/活动等低价值内容,不因此自行重跑。`linked_run_status=partial` 时先读 run-report 的 item 级 `error` 确认原因。
|
||||
- **用户要求"重新跑一批"**:候选质量低(用户主动提出)时重跑,应 `include_read=true` 并调高 `limit`(如 10),否则默认 `include_read=false` 会拉回同一批未读文章。重跑是新 Run,候选编号体系重新建立,汇报时提醒用户按新列表选择。
|
||||
|
||||
## 停止与人工介入
|
||||
|
||||
出现以下任一情况时停止自动流程并报告:
|
||||
|
||||
- 用户没有授权运行、发布或知识库写入;
|
||||
- Job/Run 返回不可恢复,或连续恢复失败;
|
||||
- Payload、候选 ID 或 Artifact 之间无法可靠关联;
|
||||
- 生成内容缺少可追踪来源;
|
||||
- Hugo 部署验证失败;
|
||||
- IMA 凭证、目标知识库或上传结果无法验证。
|
||||
|
||||
## Reference 路由
|
||||
|
||||
- `references/flow.md`:具体 MCP 调用序列、候选映射(含 extracted_path 绝对路径、候选≠文件名顺序)、Hugo 发布和 IMA 主步骤。
|
||||
- `references/content-extraction.md`:FreshRSS 内容来源与 `plain_text` 质量判断。
|
||||
- `references/public-digest-example.md`:可直接参考的 Hugo 最终页面结构。
|
||||
- `references/feishu-format-notes.md`:Feishu 输出格式限制。
|
||||
- `references/ima-format-quickref.md`:IMA Markdown 格式规则。
|
||||
- `references/ima-upload-api.md`:IMA Markdown 文件上传 API(含 `-200` 版本拦截修复)。
|
||||
- `references/ima-credential-chain.md`:IMA 凭证与知识库配置定位。
|
||||
- `references/keyword-engine-maintenance.md`:关键词治理 Skill 路由。
|
||||
- `scripts/ima_upload_one.py`:单篇 Markdown 上传 daily 知识库的完整 Python 脚本(preflight→重名→create_media→COS→add_knowledge)。
|
||||
@@ -0,0 +1,45 @@
|
||||
# 内容提取流程
|
||||
|
||||
本文说明处理流水线如何把 FreshRSS 条目转换为可供摘要使用的文章文本。
|
||||
|
||||
## 核心规则:FreshRSS 条目不重新抓取原文 URL
|
||||
|
||||
**FreshRSS 是仅提供 RSS 内容的上游。** 对于 FreshRSS 条目,流水线不会向文章原始 URL 发起 HTTP 请求。该行为由 `pipeline.py` 中的 `RSS_ONLY_UPSTREAMS = {"freshrss"}` 强制保证。
|
||||
|
||||
唯一例外是非 FreshRSS 上游。未来未设置 `upstream: freshrss` 的其他来源,可以在必要时使用 `fetch_html()` 作为回退。
|
||||
|
||||
## 内容来源优先级
|
||||
|
||||
`content_loader.py` 按以下顺序检查内容,并使用第一个包含 **至少 500 个可读字符** 的来源:
|
||||
|
||||
| 优先级 | 来源 | 含义 |
|
||||
|--------|------|------|
|
||||
| 1 | `raw_html` | 通过 `ExtractionInput.raw_html` 预先注入的 HTML;常规 FreshRSS 运行中很少使用。 |
|
||||
| 2 | `item.raw_content` | RSS `<content:encoded>` 中的文章正文;部分订阅源提供,部分不提供。 |
|
||||
| 3 | `item.raw_summary` | RSS `<description>` 中的摘要或片段;这是当前运行中最常见的来源。 |
|
||||
| 4 | `rss_content` | 来自非条目字段的独立 RSS 内容。 |
|
||||
| — | `none` | 没有可用内容;FreshRSS 不允许回源抓取,因此抛出 `RSS_CONTENT_MISSING`。 |
|
||||
|
||||
## `content_source` 与文本质量的关系
|
||||
|
||||
每个 `item-XX.extracted.json` 中的 `content_source` 字段表示流水线实际使用的内容来源:
|
||||
|
||||
- **`item.raw_content`**:RSS `<content:encoded>` 提供的文章正文,通常质量最好,接近直接阅读原文。
|
||||
- **`item.raw_summary`**:只有 RSS 摘要或描述,并非完整正文。不同来源长度差异较大,通常为 300-2000 个字符;AI 摘要基于该片段,而不是完整文章。
|
||||
- **`rss_content`**:来自独立 RSS 内容,质量取决于订阅源。
|
||||
- **`fetched_html`**:从原始 URL 抓取的 HTML。FreshRSS 条目不会出现该来源,只适用于非 FreshRSS 上游。
|
||||
|
||||
## 对日报质量的影响
|
||||
|
||||
如果提取结果文件中出现 `content_source: item.raw_summary`,说明 AI 使用的是订阅源摘要或片段,而不是完整正文。日报内容显得较浅时,原因可能只是 RSS 描述过短。
|
||||
|
||||
提高质量可以选择提供完整 `<content:encoded>` 的订阅源,或者把内容来源切换到支持全文 RSS 的系统,例如具备全文提取能力的 RSS 代理或 FiveFilters 等服务。
|
||||
|
||||
## 快速检查
|
||||
|
||||
先调用 `list_run_artifacts(run_id)`,再读取返回的提取结果产物路径,检查 `content_source` 和 `article.plain_text`。不要按 `run_id` 或“最新目录”手拼路径。
|
||||
|
||||
## 相关代码路径
|
||||
|
||||
- `src/summary_mcp/core/pipeline.py`:`RSS_ONLY_UPSTREAMS`、`_should_skip_fetch()`、`extract_content()`。
|
||||
- `src/summary_mcp/core/content_loader.py`:`choose_inline_content()` 的优先级链和 `fetch_html()`;FreshRSS 条目不会调用后者。
|
||||
@@ -0,0 +1,17 @@
|
||||
# 飞书 Markdown 格式说明
|
||||
|
||||
## 背景
|
||||
|
||||
Hermes 的飞书网关(`gateway/platforms/feishu.py`)通过 `_build_outbound_payload` 发送消息。该方法会检查内容中的 Markdown 特征,并据此决定消息类型:
|
||||
|
||||
- 内容匹配 `_MARKDOWN_HINT_RE`(加粗、列表、代码、链接等)时,使用包含 `md` 元素的飞书 `post` 类型发送,可以正常渲染。
|
||||
- 内容匹配 `_MARKDOWN_TABLE_RE`(Markdown 表头和分隔行)时,整条消息会被强制转换为 `text` 类型,即纯文本,不再渲染 Markdown。
|
||||
|
||||
原因是 `_build_markdown_post_payload` 会把内容包装为 `{"tag": "md", "text": "..."}` 元素,而飞书的 `md` 元素不支持表格,也没有把 Markdown 表格转换为飞书原生表格的逻辑。
|
||||
|
||||
## 飞书输出规则
|
||||
|
||||
- 通过飞书发送的消息不得使用 Markdown 表格;消息中只要出现一个表格,整条消息就会退化为纯文本。
|
||||
- 需要表达结构化信息时,优先使用分点列表、带标题的分节或行内格式。
|
||||
- 加粗(`**加粗**`)、行内代码(`` `代码` ``)、无序列表(`- 项目`)、有序列表(`1. 项目`)和链接均可正常使用。
|
||||
- 围栏式代码块可以使用,但代码块后的尾随内容可能存在渲染边界问题。
|
||||
@@ -0,0 +1,177 @@
|
||||
# Reader Digest Flow 操作参考
|
||||
|
||||
本文件承载 `reader-digest-flow` 的具体执行步骤。正式行为边界以 `../SKILL.md` 为准。
|
||||
|
||||
## 1. 日报 Pipeline
|
||||
|
||||
### 默认参数
|
||||
|
||||
```json
|
||||
{
|
||||
"limit": 7,
|
||||
"include_read": false,
|
||||
"mark_read": true,
|
||||
"debug_artifacts": false,
|
||||
"timeout_seconds": 60,
|
||||
"max_retries": 2
|
||||
}
|
||||
```
|
||||
|
||||
### 正式调用序列
|
||||
|
||||
```text
|
||||
start_freshrss_pipeline_job
|
||||
→ get_freshrss_pipeline_job_status
|
||||
→ get_freshrss_pipeline_job_result
|
||||
→ get_run_status
|
||||
→ get_delivery_payload / get_run_report
|
||||
```
|
||||
|
||||
状态动作:
|
||||
|
||||
- `running`:按合理间隔继续轮询;
|
||||
- `success`:读取结果,保存 `run_id`;
|
||||
- `failed`:读取关联 Run 并执行 Resume Plan;
|
||||
- `partial`:优先检查 Run Report、Artifact 和 Recovery 信息。
|
||||
|
||||
恢复序列:
|
||||
|
||||
```text
|
||||
inspect_resume_plan
|
||||
→ recommended_action=resume
|
||||
→ start_resume_job
|
||||
→ get_resume_job_status
|
||||
→ get_resume_job_result
|
||||
```
|
||||
|
||||
如果建议为 `read_terminal_result`,直接读取结果;如果为 `start_new_run`,停止并等待用户决定。
|
||||
|
||||
不要根据 `run_id` 拼目录。后续只消费调用返回的 `output_dir`、`artifact.path`、`delivery_output`、`report_output`。
|
||||
|
||||
## 2. 候选汇报与选择
|
||||
|
||||
优先使用当前 Run 的 digest brief Artifact;不可用时使用 `get_delivery_payload` 返回的候选。
|
||||
|
||||
展示规则:
|
||||
|
||||
1. 使用候选数组原始顺序并从 1 编号;
|
||||
2. 不因 `keep/review` 分组而重新编号;
|
||||
3. 每篇展示标题、来源、摘要和判断理由;
|
||||
4. 用户编号只映射当前候选数组;
|
||||
5. 跨 Artifact 读取详情时按 URL 或完整 `item_id` 匹配。
|
||||
|
||||
不要按 `summary-batch`、`item-XX` 和候选数组的相同位置推断它们是同一篇文章。
|
||||
|
||||
### extracted 文件与候选编号错位(实测 2026-07-31)
|
||||
|
||||
- `digest-brief.json` 的 `top_candidates` **没有 `item_id` 字段**,只有 `url`;
|
||||
- `extracted/item-XX.extracted.json` 的文件名序号与候选编号**可能不一致**(实例:候选2 = item-05、候选3 = item-02);
|
||||
- 正确做法:用 **URL 交叉匹配**(归一化 `%3D`→`=` 后逐条比对),或用完整 `item_id`(从 candidate-batch.json 的 `items[i].item_id` 按候选数组顺序取)在 extracted 文件里反查;两者都能验证时优先 item_id。
|
||||
|
||||
## 3. Hugo 日报
|
||||
|
||||
用户确认发布文章后:
|
||||
|
||||
1. 读取 `public-digest-example.md`;
|
||||
2. 仅使用用户选中的文章生成公开内容;
|
||||
3. 写入 Hugo 当日页面;
|
||||
4. 前台执行部署,避免把构建日志作为聊天通知;
|
||||
5. 验证首页、列表页和详情页。
|
||||
|
||||
当前部署位置:
|
||||
|
||||
```text
|
||||
/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md
|
||||
```
|
||||
|
||||
验证至少覆盖:
|
||||
|
||||
```text
|
||||
http://127.0.0.1:14322/
|
||||
http://127.0.0.1:14322/daily/
|
||||
http://127.0.0.1:14322/daily/YYYY-MM-DD/
|
||||
```
|
||||
|
||||
页面不存在时检查源文件、构建结果、容器状态和页面日期。需要 Docker/Hugo 深度排障时再读取部署项目自己的文档,不把历史事故规则复制回主 Skill。
|
||||
|
||||
## 4. 单篇知识笔记
|
||||
|
||||
用户确认 IMA 文章后:
|
||||
|
||||
1. 从候选中取得 URL 和完整 `item_id`;
|
||||
2. 从 Run Artifact 中找到匹配的 extracted 文件;
|
||||
3. 交叉验证 `article.item_id` 或 URL;
|
||||
4. 对每个单篇文件调用 `start_article_summary_job(extracted_path=<返回的 Artifact 路径>, selected_ids=[<完整 item_id>])`;
|
||||
5. `selected_ids` 传完整、非空、去除 `cand:` 前缀后的 `item_id`;
|
||||
6. 从 Job 结果的 `written_paths` 读取 Markdown。
|
||||
|
||||
### ⚠️ extracted_path 必须用绝对路径
|
||||
|
||||
`summary-mcp` 进程的工作目录是 `/root/.hermes`(不是 reader 项目根)。传相对路径(如 `outputs/freshrss/...`)会直接报 `extracted_path does not exist`。必须传绝对路径:
|
||||
|
||||
```text
|
||||
/home/ubuntu/zhu/github/reader/outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json
|
||||
```
|
||||
|
||||
### ⚠️ 候选编号 ≠ extracted 文件名顺序
|
||||
|
||||
候选数组顺序与 extracted 文件名(`item-01`…`item-07`)**不一定对齐**(实测候选2 落在 item-05)。`digest-brief.json` 的候选**没有 `item_id` 字段**,只有 URL。可靠匹配方法:
|
||||
|
||||
1. 从 `candidate-batch.json` 取每项完整 `item_id`(在 `candidate` 嵌套对象里,顶层 `item_key` 只是 `item-XX` 文件名序号);
|
||||
2. 或按 URL 匹配:归一化(`%3D`→`=`)后与每个 extracted 文件的 `article.url` / `article.canonical_url` 比对;
|
||||
3. 绝不要按候选位置对应 extracted 文件序号。
|
||||
|
||||
```python
|
||||
def norm(u): return u.replace('%3D','=').replace('%3d','=').strip()
|
||||
# 对每个 extracted 文件取 norm(article.url),与候选 norm(url) 精确比对
|
||||
```
|
||||
|
||||
正式序列:
|
||||
|
||||
```text
|
||||
start_article_summary_job
|
||||
→ get_article_summary_job_status
|
||||
→ get_article_summary_job_result
|
||||
```
|
||||
|
||||
内容依据是 extracted Artifact 中的 `article.plain_text`。不得重新抓取原 URL,不得使用原文之外的知识扩写。
|
||||
|
||||
CLI 仅在异步 MCP 不可用或人工排障时使用:
|
||||
|
||||
```bash
|
||||
python scripts/run_article_summaries.py \
|
||||
--extracted <returned-extracted-path> \
|
||||
--ids <full-item-id> \
|
||||
--output-dir <explicit-output-dir>
|
||||
```
|
||||
|
||||
## 5. IMA 上传
|
||||
|
||||
用户在知识沉淀阶段的文章选择即为上传授权。
|
||||
|
||||
执行顺序:
|
||||
|
||||
1. 检查生成的 Markdown 与来源;
|
||||
2. 文件名规范化为 `<完整文章标题>.md`;
|
||||
3. 确认目标为 `daily` knowledge base;
|
||||
4. 执行 preflight、create_media、COS upload、add_knowledge;
|
||||
5. 验证知识库条目存在。
|
||||
|
||||
上传格式与 API 参数分别见:
|
||||
|
||||
- `ima-format-quickref.md`
|
||||
- `ima-upload-api.md`
|
||||
- `ima-credential-chain.md`
|
||||
|
||||
失败时记录发生在 preflight、create_media、COS upload 或 add_knowledge 的具体阶段。不要把失败的 Markdown 改写成 URL 导入或 Notes 类型来绕过错误。
|
||||
|
||||
## 6. CLI fallback 原则
|
||||
|
||||
CLI 仅在以下场景使用:
|
||||
|
||||
- MCP 服务不可用;
|
||||
- Tool transport/launch 失败且无法取得有效 Job;
|
||||
- 用户明确要求本地调试;
|
||||
- 人工排障需要直接检查脚本输出。
|
||||
|
||||
如果已取得 `job_id` 或 `run_id`,先查询真实状态,避免因响应超时重复启动任务。
|
||||
@@ -0,0 +1,27 @@
|
||||
# IMA 凭证与安全边界
|
||||
|
||||
## 必需配置
|
||||
|
||||
- `IMA_OPENAPI_CLIENTID`
|
||||
- `IMA_OPENAPI_APIKEY`
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_ID`
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_NAME=daily`
|
||||
|
||||
优先使用当前进程环境和 IMA Skill 已支持的凭证加载机制。不要在本 Skill 中复制、迁移或重写密钥文件。
|
||||
|
||||
## 缺失处理
|
||||
|
||||
Preflight 返回凭证缺失或目标知识库无法解析时:
|
||||
|
||||
1. 停止上传;
|
||||
2. 只报告缺失的变量名或配置项;
|
||||
3. 等待用户或运行环境补齐配置;
|
||||
4. 配置恢复后重新执行 preflight,不重复生成知识笔记。
|
||||
|
||||
## 安全边界
|
||||
|
||||
- 不在聊天、日志或命令输出中打印完整 API key、KB ID 或 COS 临时凭证;
|
||||
- `create_media` 返回的 COS 凭证仅在同一受控进程内传给上传工具,不写入磁盘;
|
||||
- 不通过拼接 Shell 字符串传递凭证,使用参数数组或 IMA Skill 的封装;
|
||||
- 不绕过 Hermes 的脱敏机制;若现有工具链无法安全传递凭证,停止并报告;
|
||||
- 上传结束后不持久化 COS 临时凭证。
|
||||
@@ -0,0 +1,47 @@
|
||||
# IMA Markdown 格式速查
|
||||
|
||||
## 文件与标题
|
||||
|
||||
- 文件名:`<完整文章标题>.md`
|
||||
- `add_knowledge.title`:完整文章标题,不包含 `.md`
|
||||
- 上传类型:Markdown 文件,`media_type=7`
|
||||
- 目标:`daily` knowledge base
|
||||
|
||||
## 内容来源
|
||||
|
||||
只能使用当前 Run 的可追踪内容:
|
||||
|
||||
1. extracted Artifact 的 `article.plain_text`;
|
||||
2. 对应文章的结构化摘要;
|
||||
3. digest brief 的 summary 与 highlights。
|
||||
|
||||
不得使用通用知识补写原文没有的信息,不设置固定字数或字节数门槛。内容较短时保持简洁并忠于来源。
|
||||
|
||||
## 标准结构
|
||||
|
||||
```markdown
|
||||
# 完整文章标题
|
||||
|
||||
Source: https://原文链接
|
||||
Category: 分类
|
||||
|
||||
## 核心结论
|
||||
|
||||
## 主要论点
|
||||
|
||||
## 关键方法 / 机制
|
||||
|
||||
## 重要细节
|
||||
|
||||
## 可复用启发
|
||||
|
||||
## 关键词
|
||||
|
||||
## 主题
|
||||
```
|
||||
|
||||
- 核心结论和主要论点使用连贯段落;
|
||||
- 方法、细节和启发按完整知识点分项;
|
||||
- 没有来源支持的 Section 可以简写,不得编造内容填充。
|
||||
|
||||
上传 API 见 `ima-upload-api.md`。
|
||||
@@ -0,0 +1,126 @@
|
||||
# IMA Markdown 上传 API
|
||||
|
||||
用于将用户选中的单篇 Markdown 知识笔记上传到 `daily` knowledge base。
|
||||
|
||||
## 凭证
|
||||
|
||||
- `IMA_OPENAPI_CLIENTID`
|
||||
- `IMA_OPENAPI_APIKEY`
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_ID`
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_NAME`
|
||||
|
||||
凭证定位和恢复见 `ima-credential-chain.md`。不要在终端输出完整密钥。
|
||||
|
||||
## 上传前检查
|
||||
|
||||
- 文件名为 `<完整文章标题>.md`;
|
||||
- `title` 为完整文章标题,不带 `.md`;
|
||||
- Markdown 符合 `ima-format-quickref.md`;
|
||||
- 内容可追溯到当前 Run Artifact;
|
||||
- 用户已经明确选择该文章;
|
||||
- 目标知识库已经解析并验证。
|
||||
|
||||
## 1. Preflight
|
||||
|
||||
调用 IMA Skill 的 `preflight-check.cjs` 检查文件类型、扩展名、大小和 MIME。
|
||||
|
||||
预期:
|
||||
|
||||
```text
|
||||
file_ext=md
|
||||
content_type=text/markdown
|
||||
media_type=7
|
||||
```
|
||||
|
||||
### ⚠️ IMA skill 版本拦截(-200)
|
||||
|
||||
`ima_api.cjs` 每天首次调用会检查更新,若检测到新版(如 1.1.8 > 当前 1.1.7)会以 `code=-200` 拦截原请求。注意:**官方 zip 包内的 `meta.json` 可能没同步版本号**(下载 1.1.8 zip 后 meta 仍写 1.1.7),所以光替换文件无法跳过拦截。
|
||||
|
||||
快速修复(脚本本身已是新版,只差版本号):
|
||||
|
||||
```bash
|
||||
cd /root/.hermes/skills/openclaw-imports/ima-skill && python3 -c "
|
||||
import json
|
||||
m = json.load(open('meta.json')); m['version'] = '1.1.8'
|
||||
json.dump(m, open('meta.json','w'), ensure_ascii=False, indent=2)
|
||||
"
|
||||
```
|
||||
|
||||
先用 `diff -rq` 对比 zip 与安装目录:若只有 `.DS_Store`/meta 差异,说明代码已是最新,直接改 meta.json 版本号即可;若脚本有实质差异才需要整体替换。
|
||||
|
||||
## 2. Create Media
|
||||
|
||||
```text
|
||||
POST /openapi/wiki/v1/create_media
|
||||
```
|
||||
|
||||
请求核心字段:
|
||||
|
||||
```json
|
||||
{
|
||||
"file_name": "<完整文章标题>.md",
|
||||
"file_size": 0,
|
||||
"content_type": "text/markdown",
|
||||
"knowledge_base_id": "<daily-kb-id>",
|
||||
"file_ext": "md"
|
||||
}
|
||||
```
|
||||
|
||||
保存返回的 `media_id` 和 `cos_credential`。COS 临时凭证只在进程内传递,不打印到聊天或日志。
|
||||
|
||||
## 3. COS Upload
|
||||
|
||||
使用 IMA Skill 提供的 `cos-upload.cjs`,通过参数数组调用并检查:
|
||||
|
||||
- 进程 `returncode`;
|
||||
- `stderr`;
|
||||
- HTTP 上传结果。
|
||||
|
||||
不要拼接包含凭证的 Shell 字符串,也不要把多条 JSON 响应重定向到同一个文件。
|
||||
|
||||
## 4. Add Knowledge
|
||||
|
||||
```text
|
||||
POST /openapi/wiki/v1/add_knowledge
|
||||
```
|
||||
|
||||
核心字段:
|
||||
|
||||
```json
|
||||
{
|
||||
"media_type": 7,
|
||||
"media_id": "<media-id>",
|
||||
"title": "<完整文章标题>",
|
||||
"knowledge_base_id": "<daily-kb-id>",
|
||||
"file_info": {
|
||||
"cos_key": "<cos-key>",
|
||||
"file_size": 0,
|
||||
"file_name": "<完整文章标题>.md"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## 5. 验证
|
||||
|
||||
上传成功不能只依据本地命令退出码。验证目标 knowledge base 中存在对应标题或返回对象,并向用户报告最终结果。
|
||||
|
||||
失败时明确报告发生在 Preflight、Create Media、COS Upload 或 Add Knowledge 的哪一步,不改用 URL 导入或 Notes 类型绕过错误。
|
||||
|
||||
## 6. 已知坑:IMA skill 版本拦截(-200)
|
||||
|
||||
`ima_api.cjs` 每天首次调用会检查远端版本,若发现新版本(如 1.1.8 > 1.1.7)会以 exit=1 + stderr `{"code":-200}` 拦截**所有** API 调用,原请求不发送。此前遇到过。
|
||||
|
||||
处理方式(不必整包替换):
|
||||
|
||||
1. 按 stderr 提示下载新版 zip(如 `https://app-dl.ima.qq.com/skills/ima-skills-1.1.8.zip`)并解压;
|
||||
2. 对比新旧 `ima_api.cjs` 的 md5——zip 内核心脚本常与本地一致,只是 `meta.json` 的 `version` 未同步(zip 内仍写 1.1.7);
|
||||
3. 若 `ima_api.cjs` 一致,只需把本地 `meta.json` 的 `version` 改为远端版本号即可跳过拦截,无需替换文件。
|
||||
|
||||
调用成功后再执行本文件前面的上传流程。
|
||||
|
||||
## 6. 版本拦截与批量上传实测(2026-07-31)
|
||||
|
||||
- **`-200` skill 更新拦截**:`ima_api.cjs` 每天首次调用检查版本,发现新版时以 code -200 退出并提示更新。下载 zip 后**先对比 `ima_api.cjs` 的 md5**——实测 zip 内脚本与已装版本完全一致,只是 `meta.json` 版本号未同步。此时只需把 `~/.hermes/skills/openclaw-imports/ima-skill/meta.json` 的 `version` 改为最新版即可跳过拦截,无需替换任何脚本。
|
||||
- **Python 脚本编排上传**比 bash 可靠:bash 拼接含中文文件名/凭证的 curl 易出错。用 `subprocess` 参数数组依次调 `preflight-check.cjs` → `ima_api.cjs check_repeated_names` → `create_media` → `cos-upload.cjs`(`--secret-id/--secret-key/--token` 走参数数组,不打印)→ `add_knowledge`,每步解析返回 JSON,失败即停。
|
||||
- **批量上传**:4 篇逐个跑同一脚本即可;同名文件先 `check_repeated_names` 确认无重复。
|
||||
- 凭证从 `IMA_OPENAPI_CLIENTID` / `IMA_OPENAPI_APIKEY` 环境变量读取(`ima_api.cjs` 自动加载),KB ID 用 `IMA_DAILY_KNOWLEDGE_BASE_ID`。
|
||||
@@ -0,0 +1,29 @@
|
||||
# 关键词治理路由
|
||||
|
||||
关键词治理不属于 `reader-digest-flow` 的日常执行阶段。
|
||||
|
||||
仅当用户明确要求“清理关键词”“词库治理”“生成关键词建议”时,委托:
|
||||
|
||||
```text
|
||||
skills/keyword-cleanup-review/SKILL.md
|
||||
```
|
||||
|
||||
Review 输入生成由该 Skill 定义,后续确认与 Apply 由当前编排层负责:
|
||||
|
||||
```text
|
||||
build review bundle
|
||||
→ generate rule suggestions
|
||||
→ optional semantic suggestions
|
||||
→ human review
|
||||
→ dry-run
|
||||
→ apply accepted suggestions
|
||||
```
|
||||
|
||||
约束:
|
||||
|
||||
- Suggestions JSON 是 Review 与 Apply 之间的正式契约;
|
||||
- LLM 语义建议不能自动 Apply;
|
||||
- 不直接编辑 aliases、stopwords、watchlist 或 interest 配置;
|
||||
- 不在日报主流程中因 tag 质量不佳自动触发治理。
|
||||
|
||||
Review bundle、Suggestions、Schema 和产物保留策略以 `keyword-cleanup-review` 为唯一事实来源;该 Skill 不直接 Apply 配置。
|
||||
@@ -0,0 +1,31 @@
|
||||
+++
|
||||
title = "AI 日报 · 示例"
|
||||
date = 2026-04-01T09:00:00+08:00
|
||||
summary = "围绕 Agent 架构分层、Skills 标准化与工程实践的当日观察。"
|
||||
+++
|
||||
|
||||
# 今日概览
|
||||
|
||||
今天的公开内容主要集中在 AI Agent 架构演进、工具化落地与工程实践三条线索。行业关注点正在从“模型能做什么”转向“系统如何稳定落地并持续复用”。
|
||||
|
||||
## 今日重点
|
||||
|
||||
### 1. 从 Agent 到 Skills:AI 智能体架构的范式转变
|
||||
|
||||
文章分析了 AI 智能体从单体 Agent 向模块化 Skills 的演进,并结合 MCP、Skills 和真实项目说明能力分层与复用方式。
|
||||
|
||||
值得关注:
|
||||
|
||||
- Skills 将领域流程从 Agent 主体中拆出,便于复用和维护。
|
||||
- MCP 为 Agent 与外部工具提供标准化连接方式。
|
||||
- 工程竞争点逐渐从模型调用转向状态、工具和工作流设计。
|
||||
|
||||
这篇内容值得关注的原因在于,它把开放协议、分层架构和真实落地案例连接成了完整论证链。
|
||||
|
||||
来源:[示例来源](https://example.com/a)
|
||||
|
||||
## 趋势观察
|
||||
|
||||
1. Agent 正在从单体能力转向可组合的模块化体系。
|
||||
2. 工具契约、状态管理和验证机制正在成为 AI 应用的核心工程能力。
|
||||
3. Human-in-the-loop 仍是控制高风险副作用的重要边界。
|
||||
@@ -0,0 +1,96 @@
|
||||
#!/usr/bin/env python3
|
||||
"""IMA 上传单篇 Markdown 知识笔记到 daily 知识库。
|
||||
用法: python3 ima_upload_one.py "<绝对路径/summary.md>" "<完整文章标题>"
|
||||
依赖环境变量: IMA_OPENAPI_CLIENTID / IMA_OPENAPI_APIKEY / IMA_DAILY_KNOWLEDGE_BASE_ID
|
||||
流程: preflight -> check_repeated_names -> create_media -> cos-upload -> add_knowledge
|
||||
说明: 用 Python 而非 bash 编排,避免中文文件名/引号转义问题。
|
||||
退出码 2 = 文件名重复(需与用户确认保留双方或取消),非 0 均为失败。
|
||||
"""
|
||||
import json, os, subprocess, sys
|
||||
|
||||
SKILL_DIR = "/root/.hermes/skills/openclaw-imports/ima-skill"
|
||||
IMA_API = os.path.join(SKILL_DIR, "ima_api.cjs")
|
||||
COS_UPLOAD = os.path.join(SKILL_DIR, "knowledge-base/scripts/cos-upload.cjs")
|
||||
PREFLIGHT = os.path.join(SKILL_DIR, "knowledge-base/scripts/preflight-check.cjs")
|
||||
|
||||
def run_node(script, args):
|
||||
r = subprocess.run(["node", script] + args, capture_output=True, text=True, timeout=120)
|
||||
if r.returncode != 0:
|
||||
raise RuntimeError(f"{script} exit={r.returncode} stderr={r.stderr[:500]}")
|
||||
return json.loads(r.stdout)
|
||||
|
||||
def ima_api(api_path, body):
|
||||
r = subprocess.run(["node", IMA_API, api_path, json.dumps(body, ensure_ascii=False)],
|
||||
capture_output=True, text=True, timeout=120)
|
||||
if r.returncode != 0:
|
||||
raise RuntimeError(f"ima_api {api_path} exit={r.returncode} stderr={r.stderr[:500]}")
|
||||
resp = json.loads(r.stdout)
|
||||
if resp.get("code") != 0:
|
||||
raise RuntimeError(f"ima_api {api_path} code={resp.get('code')} msg={resp.get('msg')}")
|
||||
return resp.get("data", {})
|
||||
|
||||
def main():
|
||||
kb_id = os.environ["IMA_DAILY_KNOWLEDGE_BASE_ID"]
|
||||
file_path = sys.argv[1]
|
||||
title = sys.argv[2] # 完整文章标题(不含 .md)
|
||||
|
||||
pf = run_node(PREFLIGHT, ["--file", file_path])
|
||||
if not pf.get("pass"):
|
||||
raise RuntimeError(f"preflight failed: {pf}")
|
||||
file_name = pf["file_name"]; media_type = pf["media_type"]
|
||||
content_type = pf["content_type"]; file_size = pf["file_size"]; file_ext = pf["file_ext"]
|
||||
print(f"[preflight] ok file={file_name} ext={file_ext} size={file_size} media_type={media_type}")
|
||||
|
||||
dup = ima_api("openapi/wiki/v1/check_repeated_names", {
|
||||
"params": [{"name": file_name, "media_type": media_type}],
|
||||
"knowledge_base_id": kb_id
|
||||
})
|
||||
is_rep = dup.get("results", [{}])[0].get("is_repeated", False) if dup.get("results") else False
|
||||
if is_rep:
|
||||
print(f"[check_repeated_names] REPEATED: {file_name} — 需要处理")
|
||||
sys.exit(2)
|
||||
print("[check_repeated_names] no duplicate")
|
||||
|
||||
cm = ima_api("openapi/wiki/v1/create_media", {
|
||||
"file_name": file_name,
|
||||
"file_size": file_size,
|
||||
"content_type": content_type,
|
||||
"knowledge_base_id": kb_id,
|
||||
"file_ext": file_ext
|
||||
})
|
||||
media_id = cm["media_id"]
|
||||
cos = cm["cos_credential"]
|
||||
print(f"[create_media] media_id={media_id} cos_bucket={cos.get('bucket_name')} cos_key={cos.get('cos_key','')[:40]}")
|
||||
|
||||
r = subprocess.run(["node", COS_UPLOAD,
|
||||
"--file", file_path,
|
||||
"--secret-id", cos["secret_id"],
|
||||
"--secret-key", cos["secret_key"],
|
||||
"--token", cos["token"],
|
||||
"--bucket", cos["bucket_name"],
|
||||
"--region", cos["region"],
|
||||
"--cos-key", cos["cos_key"],
|
||||
"--content-type", content_type,
|
||||
"--start-time", str(cos.get("start_time", "")),
|
||||
"--expired-time", str(cos.get("expired_time", "")),
|
||||
"--timeout", "300000"
|
||||
], capture_output=True, text=True, timeout=360)
|
||||
if r.returncode != 0:
|
||||
raise RuntimeError(f"cos-upload exit={r.returncode} stderr={r.stderr[:800]}")
|
||||
print(f"[cos-upload] ok rc=0 stdout={r.stdout.strip()[:200]}")
|
||||
|
||||
ak = ima_api("openapi/wiki/v1/add_knowledge", {
|
||||
"media_type": media_type,
|
||||
"media_id": media_id,
|
||||
"title": title,
|
||||
"knowledge_base_id": kb_id,
|
||||
"file_info": {
|
||||
"cos_key": cos["cos_key"],
|
||||
"file_size": file_size,
|
||||
"file_name": file_name
|
||||
}
|
||||
})
|
||||
print(f"[add_knowledge] ok media_id={ak.get('media_id') or media_id}")
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -10,6 +10,7 @@ from pathlib import Path
|
||||
from typing import Any, Callable
|
||||
|
||||
import httpx
|
||||
from dotenv import dotenv_values
|
||||
|
||||
from summary_mcp.validators.llm_result import ValidationReport
|
||||
from summary_mcp.validators.llm_result import validate_llm_result as validate_llm_result_from_path
|
||||
@@ -18,6 +19,8 @@ from summary_mcp.validators.llm_result import validate_llm_result_payload
|
||||
|
||||
JSON_BLOCK_RE = re.compile(r"```(?:json)?\s*(\{.*\})\s*```", re.DOTALL)
|
||||
DEFAULT_CHAT_COMPLETIONS_URL = "https://api.openai.com/v1/chat/completions"
|
||||
REPO_ROOT = Path(__file__).resolve().parents[3]
|
||||
DEFAULT_DOTENV_PATH = REPO_ROOT / ".env"
|
||||
|
||||
|
||||
def load_text(path: Path) -> str:
|
||||
@@ -99,22 +102,60 @@ def normalize_chat_completions_url(api_url: str | None) -> str | None:
|
||||
return f"{normalized}/chat/completions"
|
||||
|
||||
|
||||
def _load_repo_dotenv() -> dict[str, str]:
|
||||
if not DEFAULT_DOTENV_PATH.exists():
|
||||
return {}
|
||||
return {
|
||||
key: value
|
||||
for key, value in dotenv_values(DEFAULT_DOTENV_PATH).items()
|
||||
if isinstance(key, str) and isinstance(value, str) and value
|
||||
}
|
||||
|
||||
|
||||
def _pick_value(*values: str | None) -> str | None:
|
||||
for value in values:
|
||||
if value:
|
||||
return value
|
||||
return None
|
||||
|
||||
|
||||
def resolve_llm_settings(
|
||||
*,
|
||||
api_key: str | None = None,
|
||||
model: str | None = None,
|
||||
api_url: str | None = None,
|
||||
) -> tuple[str, str, str]:
|
||||
resolved_api_key = api_key or os.environ.get("LLM_API_KEY") or os.environ.get("OPENAI_API_KEY")
|
||||
dotenv_map = _load_repo_dotenv()
|
||||
|
||||
resolved_api_key = _pick_value(
|
||||
api_key,
|
||||
os.environ.get("LLM_API_KEY"),
|
||||
os.environ.get("OPENAI_API_KEY"),
|
||||
dotenv_map.get("LLM_API_KEY"),
|
||||
dotenv_map.get("OPENAI_API_KEY"),
|
||||
)
|
||||
if not resolved_api_key:
|
||||
raise RuntimeError("Missing LLM_API_KEY or OPENAI_API_KEY, or pass an API key.")
|
||||
|
||||
resolved_model = model or os.environ.get("LLM_MODEL") or os.environ.get("OPENAI_MODEL")
|
||||
resolved_model = _pick_value(
|
||||
model,
|
||||
os.environ.get("LLM_MODEL"),
|
||||
os.environ.get("OPENAI_MODEL"),
|
||||
dotenv_map.get("LLM_MODEL"),
|
||||
dotenv_map.get("OPENAI_MODEL"),
|
||||
)
|
||||
if not resolved_model:
|
||||
raise RuntimeError("Missing LLM_MODEL or OPENAI_MODEL, or pass a model.")
|
||||
|
||||
resolved_api_url = normalize_chat_completions_url(
|
||||
api_url or os.environ.get("LLM_API_URL") or os.environ.get("OPENAI_API_URL") or DEFAULT_CHAT_COMPLETIONS_URL
|
||||
_pick_value(
|
||||
api_url,
|
||||
os.environ.get("LLM_API_URL"),
|
||||
os.environ.get("OPENAI_API_URL"),
|
||||
dotenv_map.get("LLM_API_URL"),
|
||||
dotenv_map.get("OPENAI_API_URL"),
|
||||
DEFAULT_CHAT_COMPLETIONS_URL,
|
||||
)
|
||||
)
|
||||
if not resolved_api_url:
|
||||
raise RuntimeError("Missing LLM_API_URL, OPENAI_API_URL, or pass an API URL.")
|
||||
|
||||
@@ -228,6 +228,7 @@ class FreshRSSClient:
|
||||
params: list[tuple[str, str | int]] = [
|
||||
("output", "json"),
|
||||
("n", limit),
|
||||
("r", "n")
|
||||
]
|
||||
if continuation:
|
||||
params.append(("c", continuation))
|
||||
|
||||
@@ -66,6 +66,7 @@ class ArticleCandidateRecord(BaseModel):
|
||||
|
||||
class OpenClawCandidateInput(BaseModel):
|
||||
candidate_id: str
|
||||
item_id: str | None = None
|
||||
title: str
|
||||
url: HttpUrl
|
||||
canonical_url: HttpUrl | None = None
|
||||
@@ -169,6 +170,7 @@ def build_openclaw_candidate_input(record: ArticleCandidateRecord) -> OpenClawCa
|
||||
|
||||
return OpenClawCandidateInput(
|
||||
candidate_id=record.candidate_id,
|
||||
item_id=item.item_id if item is not None else None,
|
||||
title=title,
|
||||
url=raw_url,
|
||||
canonical_url=normalize_candidate_url(raw_url),
|
||||
|
||||
@@ -1,12 +1,17 @@
|
||||
from __future__ import annotations
|
||||
|
||||
try:
|
||||
from datetime import UTC, date, datetime
|
||||
except ImportError: # Python < 3.11 compatibility
|
||||
from datetime import timezone, date, datetime
|
||||
|
||||
UTC = timezone.utc
|
||||
|
||||
from pydantic import BaseModel, Field
|
||||
|
||||
from .article_candidate import OpenClawCandidateInput
|
||||
|
||||
DIGEST_BRIEF_LIMIT = 5
|
||||
DIGEST_BRIEF_LIMIT: int | None = None
|
||||
DIGEST_BRIEF_HIGHLIGHT_LIMIT = 3
|
||||
|
||||
|
||||
@@ -76,11 +81,12 @@ def build_openclaw_delivery_payload(
|
||||
def build_openclaw_digest_brief(
|
||||
payload: OpenClawDeliveryPayload,
|
||||
*,
|
||||
limit: int = DIGEST_BRIEF_LIMIT,
|
||||
limit: int | None = DIGEST_BRIEF_LIMIT,
|
||||
highlight_limit: int = DIGEST_BRIEF_HIGHLIGHT_LIMIT,
|
||||
) -> OpenClawDigestBrief:
|
||||
keep_candidates = [candidate for candidate in payload.candidates if candidate.selection_decision == "keep"]
|
||||
keep_candidates = [candidate for candidate in payload.candidates if candidate.selection_decision in ("keep", "review")]
|
||||
sorted_candidates = sorted(keep_candidates, key=lambda candidate: candidate.digest_rank, reverse=True)
|
||||
selected_candidates = sorted_candidates if limit is None else sorted_candidates[:limit]
|
||||
top_candidates = [
|
||||
DigestBriefCandidate(
|
||||
title=candidate.title,
|
||||
@@ -92,7 +98,7 @@ def build_openclaw_digest_brief(
|
||||
selection_decision=candidate.selection_decision,
|
||||
url=str(candidate.url),
|
||||
)
|
||||
for candidate in sorted_candidates[:limit]
|
||||
for candidate in selected_candidates
|
||||
]
|
||||
return OpenClawDigestBrief(
|
||||
schema_version="digest-brief.v1",
|
||||
|
||||
@@ -1,10 +1,43 @@
|
||||
from .query_service import get_delivery_payload, get_run_report, get_run_status, list_run_artifacts, list_runs
|
||||
from .run_store import RunStore
|
||||
from .state_models import ArtifactRecord, RecoveryState, RunError, RunState, StageState
|
||||
from .freshrss_pipeline_jobs import (
|
||||
get_freshrss_pipeline_job_result,
|
||||
get_freshrss_pipeline_job_status,
|
||||
start_freshrss_pipeline_job,
|
||||
)
|
||||
from .query_service import get_delivery_payload, get_run_report, get_run_status, list_run_artifacts, list_runs
|
||||
|
||||
# NOTE: resume_jobs and resume_service are NOT eagerly imported here to avoid
|
||||
# a circular import chain:
|
||||
# workflows/freshrss_pipeline.py -> runtime -> resume_jobs -> resume_service
|
||||
# -> workflows/freshrss_pipeline.py (circular!)
|
||||
# They are lazy-loaded via __getattr__ when accessed as summary_mcp.runtime.*
|
||||
|
||||
|
||||
def __getattr__(name):
|
||||
import importlib
|
||||
|
||||
_LAZY = {
|
||||
"get_resume_job_result": ("resume_jobs", "get_resume_job_result"),
|
||||
"get_resume_job_status": ("resume_jobs", "get_resume_job_status"),
|
||||
"start_resume_job": ("resume_jobs", "start_resume_job"),
|
||||
"inspect_resume_plan": ("resume_service", "inspect_resume_plan"),
|
||||
"resume_run": ("resume_service", "resume_run"),
|
||||
}
|
||||
if name in _LAZY:
|
||||
mod_name, attr_name = _LAZY[name]
|
||||
mod = importlib.import_module(f".{mod_name}", __package__)
|
||||
return getattr(mod, attr_name)
|
||||
raise AttributeError(f"module {__name__!r} has no attribute {name!r}")
|
||||
|
||||
|
||||
__all__ = [
|
||||
"ArtifactRecord",
|
||||
"get_delivery_payload",
|
||||
"get_freshrss_pipeline_job_result",
|
||||
"get_freshrss_pipeline_job_status",
|
||||
"get_resume_job_result",
|
||||
"get_resume_job_status",
|
||||
"get_run_report",
|
||||
"RecoveryState",
|
||||
"RunError",
|
||||
@@ -12,6 +45,10 @@ __all__ = [
|
||||
"RunStore",
|
||||
"StageState",
|
||||
"get_run_status",
|
||||
"inspect_resume_plan",
|
||||
"list_run_artifacts",
|
||||
"list_runs",
|
||||
"resume_run",
|
||||
"start_resume_job",
|
||||
"start_freshrss_pipeline_job",
|
||||
]
|
||||
|
||||
@@ -0,0 +1,346 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import subprocess
|
||||
import sys
|
||||
from concurrent.futures import ThreadPoolExecutor
|
||||
from datetime import datetime
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
from uuid import uuid4
|
||||
|
||||
from .run_store import RunStore
|
||||
from .state_models import RunState
|
||||
|
||||
# 后台线程池:用于异步启动 job,避免阻塞 MCP stdio 响应
|
||||
# max_workers=4 足够应对并发请求,线程复用降低启动开销
|
||||
_executor = ThreadPoolExecutor(max_workers=4, thread_name_prefix="article_summary_job")
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[3]
|
||||
OUTPUT_ROOT = REPO_ROOT / "outputs" / "freshrss"
|
||||
ARTICLE_SUMMARY_JOBS_ROOT = OUTPUT_ROOT / "article_summary_jobs"
|
||||
WORKFLOW_NAME = "article_summary_job"
|
||||
RUN_TYPE = "article_summary"
|
||||
RUN_STATE_FILENAME = "run-state.json"
|
||||
INPUT_FILENAME = "input.json"
|
||||
RESULT_FILENAME = "result.json"
|
||||
JOB_REPORT_FILENAME = "job-report.json"
|
||||
DEFAULT_STAGES = [
|
||||
"prepare_job",
|
||||
"load_input",
|
||||
"generate_markdown",
|
||||
"write_result",
|
||||
]
|
||||
|
||||
|
||||
def _now() -> datetime:
|
||||
return datetime.now().astimezone()
|
||||
|
||||
|
||||
def _new_job_id() -> str:
|
||||
ts = _now().strftime("%Y%m%d-%H%M%S")
|
||||
return f"article-summary-{ts}-{uuid4().hex[:8]}"
|
||||
|
||||
|
||||
def _job_dir(job_id: str) -> Path:
|
||||
return ARTICLE_SUMMARY_JOBS_ROOT / job_id
|
||||
|
||||
|
||||
def _run_state_path(job_id: str) -> Path:
|
||||
return _job_dir(job_id) / RUN_STATE_FILENAME
|
||||
|
||||
|
||||
def _input_path(job_id: str) -> Path:
|
||||
return _job_dir(job_id) / INPUT_FILENAME
|
||||
|
||||
|
||||
def _result_path(job_id: str) -> Path:
|
||||
return _job_dir(job_id) / RESULT_FILENAME
|
||||
|
||||
|
||||
def _job_report_path(job_id: str) -> Path:
|
||||
return _job_dir(job_id) / JOB_REPORT_FILENAME
|
||||
|
||||
|
||||
def _normalize_repo_path(path: Path) -> str:
|
||||
try:
|
||||
return str(path.resolve().relative_to(REPO_ROOT.resolve()))
|
||||
except ValueError:
|
||||
return str(path)
|
||||
|
||||
|
||||
def _load_run_store(job_id: str) -> RunStore:
|
||||
return RunStore.load(path=_run_state_path(job_id), repo_root=REPO_ROOT)
|
||||
|
||||
|
||||
def _load_result(job_id: str) -> dict[str, Any] | None:
|
||||
path = _result_path(job_id)
|
||||
if not path.exists():
|
||||
return None
|
||||
return json.loads(path.read_text(encoding="utf-8-sig"))
|
||||
|
||||
|
||||
def _validate_article_summary_extracted_path(extracted_path: Path) -> None:
|
||||
if not extracted_path.exists():
|
||||
raise FileNotFoundError(f"extracted_path does not exist: {extracted_path}")
|
||||
if extracted_path.is_dir():
|
||||
raise ValueError(
|
||||
"extracted_path must be a JSON file, not a directory. "
|
||||
"Pass a single extracted file like outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json, "
|
||||
"or a batch extracted JSON file."
|
||||
)
|
||||
|
||||
|
||||
def _launch_job_background(*, job_id: str, input_payload: dict[str, Any], store: RunStore) -> None:
|
||||
"""后台线程执行:启动 subprocess 并完成 run_state 写盘。
|
||||
|
||||
主线程已经返回 MCP 响应,后台负责:
|
||||
1. 启动 runner 子进程(Popen)
|
||||
2. 记录 prepare_job 阶段完成
|
||||
3. 持久化 run-state.json
|
||||
"""
|
||||
runner_script = REPO_ROOT / "scripts" / "run_article_summary_job.py"
|
||||
cmd = [sys.executable, str(runner_script), "--job-id", job_id]
|
||||
|
||||
# 启动子进程(后台运行,不等待)
|
||||
proc = subprocess.Popen(
|
||||
cmd,
|
||||
cwd=str(REPO_ROOT),
|
||||
stdout=subprocess.DEVNULL,
|
||||
stderr=subprocess.DEVNULL,
|
||||
start_new_session=True,
|
||||
)
|
||||
|
||||
# 记录 prepare_job 完成
|
||||
store.finish_stage("prepare_job", outputs={"runner_pid": proc.pid, "runner_command": cmd})
|
||||
|
||||
|
||||
def start_article_summary_job(
|
||||
*,
|
||||
extracted_path: Path,
|
||||
selected_ids: list[str],
|
||||
output_dir: Path | None = None,
|
||||
max_retries: int = 2,
|
||||
timeout_seconds: float = 120.0,
|
||||
llm_api_key: str | None = None,
|
||||
llm_model: str | None = None,
|
||||
llm_api_url: str | None = None,
|
||||
) -> dict[str, Any]:
|
||||
"""Start an asynchronous article-summary job and return a job_id immediately.
|
||||
|
||||
关键:使用后台线程异步启动 job,主线程立即返回 MCP 响应,避免 stdio 阻塞。
|
||||
"""
|
||||
_validate_article_summary_extracted_path(extracted_path)
|
||||
if not selected_ids:
|
||||
raise ValueError("selected_ids must not be empty")
|
||||
|
||||
job_id = _new_job_id()
|
||||
job_dir = _job_dir(job_id)
|
||||
job_dir.mkdir(parents=True, exist_ok=True)
|
||||
started_at = _now()
|
||||
|
||||
resolved_output_dir = output_dir or (extracted_path.parent / "single_summaries")
|
||||
input_payload = {
|
||||
"extracted_path": str(extracted_path),
|
||||
"selected_ids": selected_ids,
|
||||
"output_dir": str(resolved_output_dir),
|
||||
"max_retries": max_retries,
|
||||
"timeout_seconds": timeout_seconds,
|
||||
"llm_api_key": llm_api_key,
|
||||
"llm_model": llm_model,
|
||||
"llm_api_url": llm_api_url,
|
||||
"launcher_pid": os.getpid(),
|
||||
}
|
||||
|
||||
# 创建 run-state(只写初始状态,不启动 subprocess)
|
||||
store = RunStore.create(
|
||||
path=_run_state_path(job_id),
|
||||
run_id=job_id,
|
||||
workflow=WORKFLOW_NAME,
|
||||
run_type=RUN_TYPE,
|
||||
started_at=started_at,
|
||||
input_payload=input_payload,
|
||||
repo_root=REPO_ROOT,
|
||||
)
|
||||
for stage_name in DEFAULT_STAGES:
|
||||
store._get_or_create_stage(stage_name)
|
||||
|
||||
# 注册输入文件
|
||||
input_file = _input_path(job_id)
|
||||
input_file.write_text(json.dumps(input_payload, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||
|
||||
# 启动 prepare_job 阶段(只写状态,不调用 save() —— 交给后台线程)
|
||||
store.start_stage("prepare_job")
|
||||
store.register_artifact(name="job_input", path=input_file, kind="json", stage="prepare_job")
|
||||
|
||||
# 【关键改造】:后台线程执行 Popen + finish_stage,主线程立即返回
|
||||
_executor.submit(_launch_job_background, job_id=job_id, input_payload=input_payload, store=store)
|
||||
|
||||
return {
|
||||
"job_id": job_id,
|
||||
"workflow": WORKFLOW_NAME,
|
||||
"run_type": RUN_TYPE,
|
||||
"status": "running",
|
||||
"output_dir": _normalize_repo_path(job_dir),
|
||||
"message": "Article summary job started successfully. Use get_article_summary_job_status to poll progress.",
|
||||
}
|
||||
|
||||
|
||||
def run_article_summary_job(*, job_id: str) -> dict[str, Any]:
|
||||
from summary_mcp.workflows.article_summary import ArticleSummaryConfig, summarize_selected_articles
|
||||
|
||||
store = _load_run_store(job_id)
|
||||
input_payload = json.loads(_input_path(job_id).read_text(encoding="utf-8-sig"))
|
||||
|
||||
try:
|
||||
store.start_stage("load_input")
|
||||
extracted_path = Path(input_payload["extracted_path"])
|
||||
output_dir = Path(input_payload["output_dir"])
|
||||
selected_ids = list(input_payload["selected_ids"])
|
||||
_validate_article_summary_extracted_path(extracted_path)
|
||||
if not selected_ids:
|
||||
raise ValueError("selected_ids must not be empty")
|
||||
store.finish_stage(
|
||||
"load_input",
|
||||
outputs={
|
||||
"extracted_path": _normalize_repo_path(extracted_path),
|
||||
"selected_id_count": len(selected_ids),
|
||||
"output_dir": _normalize_repo_path(output_dir),
|
||||
},
|
||||
)
|
||||
|
||||
store.start_stage("generate_markdown")
|
||||
config = ArticleSummaryConfig(
|
||||
max_retries=int(input_payload.get("max_retries") or 2),
|
||||
timeout_seconds=float(input_payload.get("timeout_seconds") or 120.0),
|
||||
)
|
||||
written_paths = summarize_selected_articles(
|
||||
extracted_path=extracted_path,
|
||||
selected_ids=selected_ids,
|
||||
output_dir=output_dir,
|
||||
config=config,
|
||||
api_key=input_payload.get("llm_api_key"),
|
||||
model=input_payload.get("llm_model"),
|
||||
api_url=input_payload.get("llm_api_url"),
|
||||
)
|
||||
if not written_paths:
|
||||
raise RuntimeError("Article summary job produced no Markdown outputs.")
|
||||
normalized_paths = [_normalize_repo_path(Path(p)) for p in written_paths]
|
||||
store.finish_stage(
|
||||
"generate_markdown",
|
||||
outputs={
|
||||
"written_count": len(written_paths),
|
||||
"written_paths": normalized_paths,
|
||||
},
|
||||
)
|
||||
for idx, p in enumerate(written_paths, start=1):
|
||||
store.register_artifact(
|
||||
name=f"summary_markdown_{idx}",
|
||||
path=Path(p),
|
||||
kind="markdown",
|
||||
stage="generate_markdown",
|
||||
metadata={"output_type": "single_summary"},
|
||||
)
|
||||
|
||||
store.start_stage("write_result")
|
||||
result = {
|
||||
"job_id": job_id,
|
||||
"selected_ids": selected_ids,
|
||||
"written_paths": normalized_paths,
|
||||
"completed_at": _now().isoformat(),
|
||||
}
|
||||
result_file = _result_path(job_id)
|
||||
result_file.write_text(json.dumps(result, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||
store.register_artifact(name="job_result", path=result_file, kind="json", stage="write_result")
|
||||
report_file = _job_report_path(job_id)
|
||||
report_file.write_text(
|
||||
json.dumps(
|
||||
{
|
||||
"job_id": job_id,
|
||||
"status": "success",
|
||||
"written_count": len(normalized_paths),
|
||||
"written_paths": normalized_paths,
|
||||
},
|
||||
ensure_ascii=False,
|
||||
indent=2,
|
||||
),
|
||||
encoding="utf-8",
|
||||
)
|
||||
store.register_artifact(name="job_report", path=report_file, kind="json", stage="write_result")
|
||||
store.finish_stage("write_result", outputs={"result_path": _normalize_repo_path(result_file)})
|
||||
store.finish_run(status="success")
|
||||
return result
|
||||
except Exception as exc:
|
||||
current_stage = store.state.current_stage or "generate_markdown"
|
||||
store.fail_stage(current_stage, error=exc)
|
||||
report_file = _job_report_path(job_id)
|
||||
report_file.write_text(
|
||||
json.dumps(
|
||||
{
|
||||
"job_id": job_id,
|
||||
"status": "failed",
|
||||
"error_type": type(exc).__name__,
|
||||
"error_message": str(exc),
|
||||
"failed_stage": current_stage,
|
||||
},
|
||||
ensure_ascii=False,
|
||||
indent=2,
|
||||
),
|
||||
encoding="utf-8",
|
||||
)
|
||||
try:
|
||||
store.register_artifact(name="job_report", path=report_file, kind="json", stage=current_stage)
|
||||
except Exception:
|
||||
pass
|
||||
raise
|
||||
|
||||
|
||||
def get_article_summary_job_status(*, job_id: str) -> dict[str, Any]:
|
||||
store = _load_run_store(job_id)
|
||||
state = store.state
|
||||
completed_stage_count = sum(1 for s in state.stages if s.status == "success")
|
||||
running_stage_count = sum(1 for s in state.stages if s.status == "running")
|
||||
failed_stage_count = sum(1 for s in state.stages if s.status == "failed")
|
||||
pending_stage_count = sum(1 for s in state.stages if s.status == "pending")
|
||||
return {
|
||||
"job_id": state.run_id,
|
||||
"workflow": state.workflow,
|
||||
"run_type": state.run_type,
|
||||
"status": state.status,
|
||||
"current_stage": state.current_stage,
|
||||
"started_at": state.started_at.isoformat(),
|
||||
"updated_at": state.updated_at.isoformat(),
|
||||
"finished_at": state.finished_at.isoformat() if state.finished_at else None,
|
||||
"output_dir": _normalize_repo_path(_job_dir(job_id)),
|
||||
"progress": {
|
||||
"completed_stage_count": completed_stage_count,
|
||||
"running_stage_count": running_stage_count,
|
||||
"failed_stage_count": failed_stage_count,
|
||||
"pending_stage_count": pending_stage_count,
|
||||
"total_stage_count": len(state.stages),
|
||||
},
|
||||
"artifacts": [artifact.model_dump(mode="json") for artifact in state.artifacts],
|
||||
"error_summary": state.error.model_dump(mode="json") if state.error else None,
|
||||
}
|
||||
|
||||
|
||||
def get_article_summary_job_result(*, job_id: str) -> dict[str, Any]:
|
||||
store = _load_run_store(job_id)
|
||||
state = store.state
|
||||
result = _load_result(job_id)
|
||||
if state.status != "success" or result is None:
|
||||
return {
|
||||
"job_id": state.run_id,
|
||||
"status": state.status,
|
||||
"message": "Article summary job result is not ready.",
|
||||
"error_summary": state.error.model_dump(mode="json") if state.error else None,
|
||||
}
|
||||
artifact = next((a.model_dump(mode="json") for a in state.artifacts if a.name == "job_result"), None)
|
||||
return {
|
||||
"job_id": state.run_id,
|
||||
"status": state.status,
|
||||
"written_paths": result.get("written_paths", []),
|
||||
"artifact": artifact,
|
||||
"result": result,
|
||||
}
|
||||
@@ -0,0 +1,672 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import subprocess
|
||||
import sys
|
||||
from datetime import UTC, datetime
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
from uuid import uuid4
|
||||
|
||||
from dotenv import dotenv_values
|
||||
|
||||
from .run_store import RunStore
|
||||
from .query_service import _resolve_run_record
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[3]
|
||||
OUTPUT_ROOT = REPO_ROOT / "outputs" / "freshrss"
|
||||
PIPELINE_JOBS_ROOT = OUTPUT_ROOT / "pipeline_jobs"
|
||||
WORKFLOW_NAME = "freshrss_pipeline_job"
|
||||
RUN_TYPE = "daily_digest"
|
||||
RUN_STATE_FILENAME = "run-state.json"
|
||||
INPUT_FILENAME = "input.json"
|
||||
RESULT_FILENAME = "result.json"
|
||||
JOB_REPORT_FILENAME = "job-report.json"
|
||||
DEFAULT_STAGES = [
|
||||
"prepare_job",
|
||||
"load_input",
|
||||
"run_pipeline",
|
||||
"write_result",
|
||||
]
|
||||
MIN_JOB_STALE_SECONDS = 30 * 60
|
||||
MAX_JOB_STALE_SECONDS = 6 * 60 * 60
|
||||
DEFAULT_DOTENV_PATH = REPO_ROOT / ".env"
|
||||
|
||||
|
||||
def _build_subprocess_env() -> dict[str, str]:
|
||||
"""Build an env dict for subprocess, merging parent env with .env values.
|
||||
|
||||
The subprocess inherits the Hermes MCP server's environment, but .env values
|
||||
may not be in os.environ at the time the subprocess is spawned. This function
|
||||
loads them from .env and merges them in so the child process sees all needed
|
||||
variables (LLM_API_KEY, LLM_MODEL, LLM_API_URL, FRESHRSS_*, etc.) directly
|
||||
in os.environ, avoiding any dotenv-loading timing issues inside the subprocess.
|
||||
"""
|
||||
env = os.environ.copy()
|
||||
if DEFAULT_DOTENV_PATH.exists():
|
||||
for key, value in dotenv_values(DEFAULT_DOTENV_PATH).items():
|
||||
if isinstance(key, str) and isinstance(value, str) and value:
|
||||
# Only set if not already present in parent env
|
||||
env.setdefault(key, value)
|
||||
return env
|
||||
|
||||
|
||||
def _now() -> datetime:
|
||||
return datetime.now().astimezone()
|
||||
|
||||
|
||||
def _new_job_id() -> str:
|
||||
ts = _now().strftime("%Y%m%d-%H%M%S")
|
||||
return f"freshrss-pipeline-job-{ts}-{uuid4().hex[:8]}"
|
||||
|
||||
|
||||
def _job_dir(job_id: str) -> Path:
|
||||
return PIPELINE_JOBS_ROOT / job_id
|
||||
|
||||
|
||||
def _run_state_path(job_id: str) -> Path:
|
||||
return _job_dir(job_id) / RUN_STATE_FILENAME
|
||||
|
||||
|
||||
def _input_path(job_id: str) -> Path:
|
||||
return _job_dir(job_id) / INPUT_FILENAME
|
||||
|
||||
|
||||
def _result_path(job_id: str) -> Path:
|
||||
return _job_dir(job_id) / RESULT_FILENAME
|
||||
|
||||
|
||||
def _job_report_path(job_id: str) -> Path:
|
||||
return _job_dir(job_id) / JOB_REPORT_FILENAME
|
||||
|
||||
|
||||
def _normalize_repo_path(path: Path) -> str:
|
||||
try:
|
||||
return str(path.resolve().relative_to(REPO_ROOT.resolve()))
|
||||
except ValueError:
|
||||
return str(path)
|
||||
|
||||
|
||||
def _normalize_repo_path_value(path_value: str | None) -> str | None:
|
||||
if not path_value:
|
||||
return None
|
||||
return _normalize_repo_path(Path(path_value))
|
||||
|
||||
|
||||
def _normalize_keyword_index(keyword_index: dict[str, Any] | None) -> dict[str, Any]:
|
||||
if not isinstance(keyword_index, dict):
|
||||
return {}
|
||||
normalized = dict(keyword_index)
|
||||
for field in ("daily_output", "stats_output"):
|
||||
normalized[field] = _normalize_repo_path_value(normalized.get(field))
|
||||
return normalized
|
||||
|
||||
|
||||
def _load_run_store(job_id: str) -> RunStore:
|
||||
return RunStore.load(path=_run_state_path(job_id), repo_root=REPO_ROOT)
|
||||
|
||||
|
||||
def _load_result(job_id: str) -> dict[str, Any] | None:
|
||||
path = _result_path(job_id)
|
||||
if not path.exists():
|
||||
return None
|
||||
return json.loads(path.read_text(encoding="utf-8-sig"))
|
||||
|
||||
|
||||
def _build_run_defaults(*, run_id: str | None, output_dir: Path | None) -> tuple[str, Path]:
|
||||
resolved_output_dir = output_dir
|
||||
resolved_run_id = run_id
|
||||
if resolved_run_id is not None and resolved_output_dir is not None:
|
||||
return resolved_run_id, resolved_output_dir
|
||||
|
||||
run_stamp = datetime.now(tz=UTC).strftime("%Y%m%d-%H%M%S")
|
||||
if resolved_run_id is None:
|
||||
resolved_run_id = f"freshrss-pipeline-{run_stamp}"
|
||||
if resolved_output_dir is None:
|
||||
resolved_output_dir = OUTPUT_ROOT / "rerun" / run_stamp
|
||||
return resolved_run_id, resolved_output_dir
|
||||
|
||||
|
||||
def _build_result_payload(*, job_id: str, pipeline_result: dict[str, Any], include_item_reports: bool) -> dict[str, Any]:
|
||||
result = {
|
||||
"job_id": job_id,
|
||||
"run_id": pipeline_result["run_id"],
|
||||
"output_dir": _normalize_repo_path_value(pipeline_result.get("output_dir")),
|
||||
"raw_output": _normalize_repo_path_value(pipeline_result.get("raw_output")),
|
||||
"delivery_output": _normalize_repo_path_value(pipeline_result.get("delivery_output")),
|
||||
"digest_brief_output": _normalize_repo_path_value(pipeline_result.get("digest_brief_output")),
|
||||
"report_output": _normalize_repo_path_value(pipeline_result.get("report_output")),
|
||||
"keyword_index": _normalize_keyword_index(pipeline_result.get("keyword_index")),
|
||||
"pulled_count": pipeline_result.get("pulled_count"),
|
||||
"delivered_count": pipeline_result.get("delivered_count"),
|
||||
"marked_read_count": pipeline_result.get("marked_read_count"),
|
||||
"status_counts": pipeline_result.get("status_counts"),
|
||||
"debug_artifacts": pipeline_result.get("debug_artifacts"),
|
||||
"completed_at": _now().isoformat(),
|
||||
}
|
||||
if include_item_reports:
|
||||
result["items"] = pipeline_result.get("items", [])
|
||||
return result
|
||||
|
||||
|
||||
def _write_job_report(
|
||||
*,
|
||||
job_id: str,
|
||||
payload: dict[str, Any],
|
||||
) -> Path:
|
||||
report_file = _job_report_path(job_id)
|
||||
report_file.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||
return report_file
|
||||
|
||||
|
||||
def start_freshrss_pipeline_job(
|
||||
*,
|
||||
limit: int = 5,
|
||||
mark_read: bool = False,
|
||||
include_read: bool = False,
|
||||
debug_artifacts: bool = False,
|
||||
continuation: str | None = None,
|
||||
timeout_seconds: float = 60.0,
|
||||
max_retries: int = 2,
|
||||
stream_id: str = "user/-/state/com.google/reading-list",
|
||||
api_base_url: str | None = None,
|
||||
username: str | None = None,
|
||||
api_password: str | None = None,
|
||||
llm_api_key: str | None = None,
|
||||
llm_model: str | None = None,
|
||||
llm_api_url: str | None = None,
|
||||
context: dict[str, Any] | None = None,
|
||||
run_id: str | None = None,
|
||||
date_value: str | None = None,
|
||||
output_dir: Path | None = None,
|
||||
include_item_reports: bool = False,
|
||||
) -> dict[str, Any]:
|
||||
job_id = _new_job_id()
|
||||
job_dir = _job_dir(job_id)
|
||||
job_dir.mkdir(parents=True, exist_ok=True)
|
||||
started_at = _now()
|
||||
resolved_run_id, resolved_output_dir = _build_run_defaults(run_id=run_id, output_dir=output_dir)
|
||||
|
||||
input_payload = {
|
||||
"limit": limit,
|
||||
"mark_read": mark_read,
|
||||
"include_read": include_read,
|
||||
"debug_artifacts": debug_artifacts,
|
||||
"continuation": continuation,
|
||||
"timeout_seconds": timeout_seconds,
|
||||
"max_retries": max_retries,
|
||||
"stream_id": stream_id,
|
||||
"api_base_url": api_base_url,
|
||||
"username": username,
|
||||
"api_password": api_password,
|
||||
"llm_api_key": llm_api_key,
|
||||
"llm_model": llm_model,
|
||||
"llm_api_url": llm_api_url,
|
||||
"context": context,
|
||||
"run_id": resolved_run_id,
|
||||
"date_value": date_value,
|
||||
"output_dir": str(resolved_output_dir),
|
||||
"include_item_reports": include_item_reports,
|
||||
"launcher_pid": os.getpid(),
|
||||
}
|
||||
|
||||
store = RunStore.create(
|
||||
path=_run_state_path(job_id),
|
||||
run_id=job_id,
|
||||
workflow=WORKFLOW_NAME,
|
||||
run_type=RUN_TYPE,
|
||||
started_at=started_at,
|
||||
input_payload=input_payload,
|
||||
repo_root=REPO_ROOT,
|
||||
)
|
||||
for stage_name in DEFAULT_STAGES:
|
||||
store._get_or_create_stage(stage_name)
|
||||
store.save()
|
||||
|
||||
store.start_stage("prepare_job")
|
||||
|
||||
input_file = _input_path(job_id)
|
||||
input_file.write_text(json.dumps(input_payload, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||
store.register_artifact(name="job_input", path=input_file, kind="json", stage="prepare_job")
|
||||
|
||||
runner_script = REPO_ROOT / "scripts" / "run_freshrss_pipeline_job.py"
|
||||
cmd = [sys.executable, str(runner_script), "--job-id", job_id]
|
||||
try:
|
||||
proc = subprocess.Popen(
|
||||
cmd,
|
||||
cwd=str(REPO_ROOT),
|
||||
stdout=subprocess.DEVNULL,
|
||||
stderr=subprocess.DEVNULL,
|
||||
start_new_session=True,
|
||||
env=_build_subprocess_env(),
|
||||
)
|
||||
except Exception as exc:
|
||||
report_file = _write_job_report(
|
||||
job_id=job_id,
|
||||
payload={
|
||||
"job_id": job_id,
|
||||
"status": "failed",
|
||||
"error_type": type(exc).__name__,
|
||||
"error_message": str(exc),
|
||||
"failed_stage": "prepare_job",
|
||||
"linked_run_id": resolved_run_id,
|
||||
"linked_output_dir": _normalize_repo_path(resolved_output_dir),
|
||||
},
|
||||
)
|
||||
store.register_artifact(name="job_report", path=report_file, kind="json", stage="prepare_job")
|
||||
store.fail_stage("prepare_job", error=exc)
|
||||
raise
|
||||
|
||||
store.finish_stage(
|
||||
"prepare_job",
|
||||
outputs={
|
||||
"runner_pid": proc.pid,
|
||||
"runner_command": cmd,
|
||||
"linked_run_id": resolved_run_id,
|
||||
"linked_output_dir": _normalize_repo_path(resolved_output_dir),
|
||||
},
|
||||
)
|
||||
|
||||
return {
|
||||
"job_id": job_id,
|
||||
"workflow": WORKFLOW_NAME,
|
||||
"run_type": RUN_TYPE,
|
||||
"status": "running",
|
||||
"output_dir": _normalize_repo_path(job_dir),
|
||||
"linked_run_id": resolved_run_id,
|
||||
"linked_output_dir": _normalize_repo_path(resolved_output_dir),
|
||||
"message": "FreshRSS pipeline job started successfully. Use get_freshrss_pipeline_job_status to poll progress.",
|
||||
}
|
||||
|
||||
|
||||
def run_freshrss_pipeline_job(*, job_id: str) -> dict[str, Any]:
|
||||
from datetime import date
|
||||
|
||||
from summary_mcp.workflows import run_freshrss_pipeline
|
||||
|
||||
store = _load_run_store(job_id)
|
||||
input_payload = json.loads(_input_path(job_id).read_text(encoding="utf-8-sig"))
|
||||
current_stage = "run_pipeline"
|
||||
|
||||
try:
|
||||
store.start_stage("load_input")
|
||||
resolved_output_dir = Path(input_payload["output_dir"])
|
||||
resolved_run_id = str(input_payload["run_id"])
|
||||
store.finish_stage(
|
||||
"load_input",
|
||||
outputs={
|
||||
"linked_run_id": resolved_run_id,
|
||||
"linked_output_dir": _normalize_repo_path(resolved_output_dir),
|
||||
"limit": int(input_payload.get("limit") or 5),
|
||||
},
|
||||
)
|
||||
|
||||
store.start_stage("run_pipeline")
|
||||
pipeline_result = run_freshrss_pipeline(
|
||||
api_base_url=input_payload.get("api_base_url"),
|
||||
username=input_payload.get("username"),
|
||||
api_password=input_payload.get("api_password"),
|
||||
stream_id=input_payload.get("stream_id") or "user/-/state/com.google/reading-list",
|
||||
limit=int(input_payload.get("limit") or 5),
|
||||
continuation=input_payload.get("continuation"),
|
||||
include_read=bool(input_payload.get("include_read")),
|
||||
mark_read=bool(input_payload.get("mark_read")),
|
||||
debug_artifacts=bool(input_payload.get("debug_artifacts")),
|
||||
context=input_payload.get("context"),
|
||||
max_retries=int(input_payload.get("max_retries") or 2),
|
||||
timeout_seconds=float(input_payload.get("timeout_seconds") or 60.0),
|
||||
llm_api_key=input_payload.get("llm_api_key"),
|
||||
llm_model=input_payload.get("llm_model"),
|
||||
llm_api_url=input_payload.get("llm_api_url"),
|
||||
run_id=resolved_run_id,
|
||||
delivery_date=date.fromisoformat(input_payload["date_value"]) if input_payload.get("date_value") else None,
|
||||
output_dir=resolved_output_dir,
|
||||
)
|
||||
store.finish_stage(
|
||||
"run_pipeline",
|
||||
outputs={
|
||||
"linked_run_id": pipeline_result.get("run_id"),
|
||||
"linked_output_dir": _normalize_repo_path_value(pipeline_result.get("output_dir")),
|
||||
"delivery_output": _normalize_repo_path_value(pipeline_result.get("delivery_output")),
|
||||
"report_output": _normalize_repo_path_value(pipeline_result.get("report_output")),
|
||||
"delivered_count": pipeline_result.get("delivered_count"),
|
||||
"marked_read_count": pipeline_result.get("marked_read_count"),
|
||||
},
|
||||
)
|
||||
|
||||
store.start_stage("write_result")
|
||||
result = _build_result_payload(
|
||||
job_id=job_id,
|
||||
pipeline_result=pipeline_result,
|
||||
include_item_reports=bool(input_payload.get("include_item_reports")),
|
||||
)
|
||||
result_file = _result_path(job_id)
|
||||
result_file.write_text(json.dumps(result, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||
store.register_artifact(name="job_result", path=result_file, kind="json", stage="write_result")
|
||||
|
||||
report_file = _write_job_report(
|
||||
job_id=job_id,
|
||||
payload={
|
||||
"job_id": job_id,
|
||||
"status": "success",
|
||||
"run_id": result["run_id"],
|
||||
"delivery_output": result["delivery_output"],
|
||||
"report_output": result["report_output"],
|
||||
"digest_brief_output": result["digest_brief_output"],
|
||||
"pulled_count": result["pulled_count"],
|
||||
"delivered_count": result["delivered_count"],
|
||||
"marked_read_count": result["marked_read_count"],
|
||||
},
|
||||
)
|
||||
store.register_artifact(name="job_report", path=report_file, kind="json", stage="write_result")
|
||||
store.finish_stage(
|
||||
"write_result",
|
||||
outputs={
|
||||
"result_path": _normalize_repo_path(result_file),
|
||||
"linked_run_id": result["run_id"],
|
||||
},
|
||||
)
|
||||
store.finish_run(status="success")
|
||||
return result
|
||||
except Exception as exc:
|
||||
current_stage = store.state.current_stage or current_stage
|
||||
report_file = _write_job_report(
|
||||
job_id=job_id,
|
||||
payload={
|
||||
"job_id": job_id,
|
||||
"status": "failed",
|
||||
"error_type": type(exc).__name__,
|
||||
"error_message": str(exc),
|
||||
"failed_stage": current_stage,
|
||||
"linked_run_id": input_payload.get("run_id"),
|
||||
"linked_output_dir": _normalize_repo_path_value(input_payload.get("output_dir")),
|
||||
},
|
||||
)
|
||||
try:
|
||||
store.register_artifact(name="job_report", path=report_file, kind="json", stage=current_stage)
|
||||
except Exception:
|
||||
pass
|
||||
store.fail_stage(current_stage, error=exc)
|
||||
raise
|
||||
|
||||
|
||||
def get_freshrss_pipeline_job_status(*, job_id: str) -> dict[str, Any]:
|
||||
store = _load_run_store(job_id)
|
||||
state = store.state
|
||||
effective = _resolve_effective_job_view(job_id=job_id, state=state)
|
||||
progress = _build_effective_job_progress(state=state, effective_status=effective["status"])
|
||||
linked_run_id = state.input.get("run_id") if isinstance(state.input, dict) else None
|
||||
linked_output_dir = _normalize_repo_path_value(state.input.get("output_dir")) if isinstance(state.input, dict) else None
|
||||
return {
|
||||
"job_id": state.run_id,
|
||||
"workflow": state.workflow,
|
||||
"run_type": state.run_type,
|
||||
"status": effective["status"],
|
||||
"current_stage": effective["current_stage"],
|
||||
"started_at": state.started_at.isoformat(),
|
||||
"updated_at": effective["updated_at"],
|
||||
"finished_at": effective["finished_at"],
|
||||
"output_dir": _normalize_repo_path(_job_dir(job_id)),
|
||||
"linked_run_id": linked_run_id,
|
||||
"linked_output_dir": linked_output_dir,
|
||||
"progress": progress,
|
||||
"artifacts": [artifact.model_dump(mode="json") for artifact in state.artifacts],
|
||||
"error_summary": state.error.model_dump(mode="json") if state.error else None,
|
||||
"status_source": effective["status_source"],
|
||||
"state_quality": effective["state_quality"],
|
||||
"state_conflict": effective["state_conflict"],
|
||||
"state_conflict_reason": effective["state_conflict_reason"],
|
||||
"status_note": effective["status_note"],
|
||||
"raw_status": state.status,
|
||||
"raw_current_stage": state.current_stage,
|
||||
"linked_run_status": effective["linked_run_status"],
|
||||
"linked_run_state_conflict": effective["linked_run_state_conflict"],
|
||||
}
|
||||
|
||||
|
||||
def get_freshrss_pipeline_job_result(*, job_id: str) -> dict[str, Any]:
|
||||
store = _load_run_store(job_id)
|
||||
state = store.state
|
||||
result = _load_result(job_id)
|
||||
effective = _resolve_effective_job_view(job_id=job_id, state=state, result=result)
|
||||
synthesized_result = result or effective["result"]
|
||||
if effective["status"] != "success" or synthesized_result is None:
|
||||
message = "FreshRSS pipeline job result is not ready."
|
||||
if effective["status"] == "failed":
|
||||
message = "FreshRSS pipeline job did not complete successfully, so no terminal job result is available."
|
||||
return {
|
||||
"job_id": state.run_id,
|
||||
"status": effective["status"],
|
||||
"message": message,
|
||||
"linked_run_id": state.input.get("run_id") if isinstance(state.input, dict) else None,
|
||||
"error_summary": state.error.model_dump(mode="json") if state.error else None,
|
||||
"status_source": effective["status_source"],
|
||||
"status_note": effective["status_note"],
|
||||
}
|
||||
|
||||
artifact = next((a.model_dump(mode="json") for a in state.artifacts if a.name == "job_result"), None)
|
||||
return {
|
||||
"job_id": state.run_id,
|
||||
"status": effective["status"],
|
||||
"run_id": synthesized_result.get("run_id"),
|
||||
"output_dir": synthesized_result.get("output_dir"),
|
||||
"raw_output": synthesized_result.get("raw_output"),
|
||||
"delivery_output": synthesized_result.get("delivery_output"),
|
||||
"digest_brief_output": synthesized_result.get("digest_brief_output"),
|
||||
"report_output": synthesized_result.get("report_output"),
|
||||
"pulled_count": synthesized_result.get("pulled_count"),
|
||||
"delivered_count": synthesized_result.get("delivered_count"),
|
||||
"marked_read_count": synthesized_result.get("marked_read_count"),
|
||||
"status_counts": synthesized_result.get("status_counts"),
|
||||
"keyword_index": synthesized_result.get("keyword_index"),
|
||||
"artifact": artifact,
|
||||
"result": synthesized_result,
|
||||
"status_source": effective["status_source"],
|
||||
"status_note": effective["status_note"],
|
||||
"result_source": synthesized_result.get("result_source", "job_result"),
|
||||
"linked_run_status": effective["linked_run_status"],
|
||||
}
|
||||
|
||||
|
||||
def _resolve_effective_job_view(
|
||||
*,
|
||||
job_id: str,
|
||||
state: Any,
|
||||
result: dict[str, Any] | None = None,
|
||||
) -> dict[str, Any]:
|
||||
resolved_result = result or _load_result(job_id)
|
||||
linked_run_record = _load_linked_run_record(state)
|
||||
raw_status = state.status
|
||||
raw_current_stage = state.current_stage
|
||||
effective_status = raw_status
|
||||
effective_current_stage = raw_current_stage
|
||||
effective_updated_at = state.updated_at.isoformat()
|
||||
effective_finished_at = state.finished_at.isoformat() if state.finished_at else None
|
||||
status_source = "job_run_state"
|
||||
state_quality = "trusted"
|
||||
state_conflict = False
|
||||
state_conflict_reason = None
|
||||
status_note = None
|
||||
|
||||
if resolved_result is not None:
|
||||
completed_at = resolved_result.get("completed_at")
|
||||
effective_status = "success"
|
||||
effective_current_stage = None
|
||||
effective_updated_at = completed_at or effective_updated_at
|
||||
effective_finished_at = completed_at or effective_finished_at
|
||||
if raw_status != "success" or raw_current_stage is not None:
|
||||
status_source = "job_result_reconciliation"
|
||||
state_quality = "reconciled"
|
||||
state_conflict = True
|
||||
state_conflict_reason = "result.json already exists, but the job run-state did not converge to success."
|
||||
status_note = "job result exists; the job can be treated as completed."
|
||||
elif linked_run_record is not None:
|
||||
if _linked_run_result_available(linked_run_record):
|
||||
effective_status = "success"
|
||||
effective_current_stage = None
|
||||
effective_updated_at = linked_run_record.get("updated_at") or effective_updated_at
|
||||
effective_finished_at = linked_run_record.get("finished_at") or effective_finished_at
|
||||
status_source = "linked_run_reconciliation"
|
||||
state_quality = "reconciled"
|
||||
state_conflict = raw_status != "success" or raw_current_stage is not None
|
||||
state_conflict_reason = (
|
||||
"The linked run already has terminal artifacts, but the outer job run-state did not converge."
|
||||
)
|
||||
status_note = "linked run artifacts are complete; the job can be treated as completed."
|
||||
resolved_result = _build_result_from_linked_run(job_id=job_id, state=state, linked_run_record=linked_run_record)
|
||||
elif linked_run_record["status"] == "failed" and raw_status == "running":
|
||||
effective_status = "failed"
|
||||
effective_current_stage = None
|
||||
effective_updated_at = linked_run_record.get("updated_at") or effective_updated_at
|
||||
effective_finished_at = linked_run_record.get("finished_at") or effective_finished_at
|
||||
status_source = "linked_run_reconciliation"
|
||||
state_quality = "reconciled"
|
||||
state_conflict = True
|
||||
state_conflict_reason = "The linked run is already failed, but the outer job still reports running."
|
||||
status_note = "linked run failed; the outer job state appears stale."
|
||||
elif raw_status == "running":
|
||||
stale_reason = _build_stale_job_reason(state)
|
||||
if stale_reason is not None:
|
||||
effective_status = "failed"
|
||||
effective_current_stage = raw_current_stage
|
||||
effective_finished_at = effective_updated_at
|
||||
status_source = "stale_job_state_timeout"
|
||||
state_quality = "reconciled"
|
||||
state_conflict = True
|
||||
state_conflict_reason = stale_reason
|
||||
status_note = stale_reason
|
||||
|
||||
return {
|
||||
"status": effective_status,
|
||||
"current_stage": effective_current_stage,
|
||||
"updated_at": effective_updated_at,
|
||||
"finished_at": effective_finished_at,
|
||||
"status_source": status_source,
|
||||
"state_quality": state_quality,
|
||||
"state_conflict": state_conflict,
|
||||
"state_conflict_reason": state_conflict_reason,
|
||||
"status_note": status_note,
|
||||
"linked_run_status": linked_run_record["status"] if linked_run_record is not None else None,
|
||||
"linked_run_state_conflict": linked_run_record.get("state_conflict") if linked_run_record is not None else None,
|
||||
"result": resolved_result,
|
||||
}
|
||||
|
||||
|
||||
def _build_effective_job_progress(*, state: Any, effective_status: str) -> dict[str, int]:
|
||||
total_stage_count = len(state.stages)
|
||||
if effective_status == "success":
|
||||
return {
|
||||
"completed_stage_count": total_stage_count,
|
||||
"running_stage_count": 0,
|
||||
"failed_stage_count": 0,
|
||||
"pending_stage_count": 0,
|
||||
"total_stage_count": total_stage_count,
|
||||
}
|
||||
if effective_status == "failed" and state.status == "running":
|
||||
completed_stage_count = sum(1 for s in state.stages if s.status == "success")
|
||||
pending_stage_count = max(total_stage_count - completed_stage_count - 1, 0)
|
||||
return {
|
||||
"completed_stage_count": completed_stage_count,
|
||||
"running_stage_count": 0,
|
||||
"failed_stage_count": 1,
|
||||
"pending_stage_count": pending_stage_count,
|
||||
"total_stage_count": total_stage_count,
|
||||
}
|
||||
return {
|
||||
"completed_stage_count": sum(1 for s in state.stages if s.status == "success"),
|
||||
"running_stage_count": sum(1 for s in state.stages if s.status == "running"),
|
||||
"failed_stage_count": sum(1 for s in state.stages if s.status == "failed"),
|
||||
"pending_stage_count": sum(1 for s in state.stages if s.status == "pending"),
|
||||
"total_stage_count": total_stage_count,
|
||||
}
|
||||
|
||||
|
||||
def _load_linked_run_record(state: Any) -> dict[str, Any] | None:
|
||||
if not isinstance(state.input, dict):
|
||||
return None
|
||||
linked_run_id = state.input.get("run_id")
|
||||
if not isinstance(linked_run_id, str) or not linked_run_id.strip():
|
||||
return None
|
||||
try:
|
||||
return _resolve_run_record(linked_run_id)
|
||||
except FileNotFoundError:
|
||||
return None
|
||||
|
||||
|
||||
def _linked_run_result_available(record: dict[str, Any]) -> bool:
|
||||
artifact_presence = record.get("artifact_presence")
|
||||
if not isinstance(artifact_presence, dict):
|
||||
return False
|
||||
return bool(artifact_presence.get("run_report")) and bool(artifact_presence.get("delivery_payload"))
|
||||
|
||||
|
||||
def _build_result_from_linked_run(
|
||||
*,
|
||||
job_id: str,
|
||||
state: Any,
|
||||
linked_run_record: dict[str, Any],
|
||||
) -> dict[str, Any]:
|
||||
report = linked_run_record.get("report")
|
||||
if not isinstance(report, dict):
|
||||
raise RuntimeError("Cannot synthesize job result because linked run-report.json is missing.")
|
||||
|
||||
result = {
|
||||
"job_id": job_id,
|
||||
"run_id": linked_run_record["run_id"],
|
||||
"output_dir": _normalize_repo_path_value(str(linked_run_record["run_dir"])),
|
||||
"raw_output": _normalize_repo_path_value(report.get("raw_output")),
|
||||
"delivery_output": _normalize_repo_path_value(report.get("delivery_output")),
|
||||
"digest_brief_output": _normalize_repo_path_value(report.get("digest_brief_output")),
|
||||
"report_output": _normalize_repo_path(linked_run_record["run_dir"] / "run-report.json"),
|
||||
"keyword_index": _normalize_keyword_index(report.get("keyword_index")),
|
||||
"pulled_count": report.get("pulled_count"),
|
||||
"delivered_count": report.get("delivered_count"),
|
||||
"marked_read_count": report.get("marked_read_count"),
|
||||
"status_counts": report.get("status_counts"),
|
||||
"debug_artifacts": report.get("debug_artifacts"),
|
||||
"completed_at": report.get("completed_at") or linked_run_record.get("finished_at"),
|
||||
"result_source": "linked_run_report",
|
||||
}
|
||||
if bool(state.input.get("include_item_reports")):
|
||||
result["items"] = report.get("items", [])
|
||||
return result
|
||||
|
||||
|
||||
def _build_stale_job_reason(state: Any) -> str | None:
|
||||
updated_at = state.updated_at
|
||||
stale_after_seconds = _estimate_job_stale_seconds(state)
|
||||
age_seconds = (datetime.now(tz=updated_at.tzinfo) - updated_at).total_seconds()
|
||||
if age_seconds < stale_after_seconds:
|
||||
return None
|
||||
|
||||
current_stage = state.current_stage or "unknown_stage"
|
||||
age_minutes = int(age_seconds // 60)
|
||||
stale_after_minutes = int(stale_after_seconds // 60)
|
||||
return (
|
||||
f"job run-state has remained in running state at {current_stage} for about {age_minutes} minutes "
|
||||
f"without result.json; it exceeded the stale threshold of {stale_after_minutes} minutes."
|
||||
)
|
||||
|
||||
|
||||
def _estimate_job_stale_seconds(state: Any) -> int:
|
||||
input_payload = state.input if isinstance(state.input, dict) else {}
|
||||
limit = _safe_int(input_payload.get("limit"), default=5)
|
||||
timeout_seconds = _safe_float(input_payload.get("timeout_seconds"), default=60.0)
|
||||
max_retries = _safe_int(input_payload.get("max_retries"), default=2)
|
||||
estimated = int((limit * max(timeout_seconds, 1.0) * max(max_retries, 1)) + 20 * 60)
|
||||
return max(MIN_JOB_STALE_SECONDS, min(MAX_JOB_STALE_SECONDS, estimated))
|
||||
|
||||
|
||||
def _safe_int(value: Any, *, default: int) -> int:
|
||||
try:
|
||||
return int(value)
|
||||
except (TypeError, ValueError):
|
||||
return default
|
||||
|
||||
|
||||
def _safe_float(value: Any, *, default: float) -> float:
|
||||
try:
|
||||
return float(value)
|
||||
except (TypeError, ValueError):
|
||||
return default
|
||||
@@ -45,6 +45,8 @@ DISCOVERED_ARTIFACTS = [
|
||||
"relative_path": Path("candidates/digest-brief.json"),
|
||||
},
|
||||
]
|
||||
MIN_RUNNING_STALE_SECONDS = 30 * 60
|
||||
MAX_RUNNING_STALE_SECONDS = 6 * 60 * 60
|
||||
|
||||
|
||||
def get_run_status(*, run_id: str) -> dict[str, Any]:
|
||||
@@ -186,7 +188,7 @@ def _build_run_record(run_dir: Path) -> dict[str, Any]:
|
||||
run_store = RunStore.load(path=state_path, repo_root=REPO_ROOT)
|
||||
state = run_store.state
|
||||
run_id = state.run_id
|
||||
return {
|
||||
record = {
|
||||
"run_id": run_id,
|
||||
"workflow": state.workflow,
|
||||
"run_type": state.run_type,
|
||||
@@ -204,12 +206,13 @@ def _build_run_record(run_dir: Path) -> dict[str, Any]:
|
||||
"state": state,
|
||||
"report": _load_json(report_path) if report_path.exists() else None,
|
||||
}
|
||||
return _reconcile_run_record(record)
|
||||
|
||||
report = _load_json(report_path) if report_path.exists() else None
|
||||
run_id = str(report.get("run_id")) if isinstance(report, dict) and report.get("run_id") else run_dir.name
|
||||
inferred_record = _infer_run_record_from_directory(run_dir=run_dir, report=report, run_id=run_id)
|
||||
inferred_record["aliases"] = {run_id, run_dir.name}
|
||||
return inferred_record
|
||||
return _reconcile_run_record(inferred_record)
|
||||
|
||||
|
||||
def _infer_run_record_from_directory(*, run_dir: Path, report: dict[str, Any] | None, run_id: str) -> dict[str, Any]:
|
||||
@@ -298,6 +301,13 @@ def _build_status_response(record: dict[str, Any]) -> dict[str, Any]:
|
||||
"artifacts": _collect_artifacts(record),
|
||||
"recovery": record["recovery"],
|
||||
"state_source": record["state_source"],
|
||||
"status_source": record["status_source"],
|
||||
"state_quality": record["state_quality"],
|
||||
"state_conflict": record["state_conflict"],
|
||||
"state_conflict_reason": record["state_conflict_reason"],
|
||||
"raw_status": record["raw_status"],
|
||||
"raw_current_stage": record["raw_current_stage"],
|
||||
"artifact_presence": record["artifact_presence"],
|
||||
}
|
||||
|
||||
|
||||
@@ -310,6 +320,10 @@ def _build_run_lookup_response(record: dict[str, Any]) -> dict[str, Any]:
|
||||
"status": record["status"],
|
||||
"output_dir": _normalize_repo_path(record["run_dir"]),
|
||||
"state_source": record["state_source"],
|
||||
"status_source": record["status_source"],
|
||||
"state_quality": record["state_quality"],
|
||||
"state_conflict": record["state_conflict"],
|
||||
"state_conflict_reason": record["state_conflict_reason"],
|
||||
}
|
||||
|
||||
|
||||
@@ -330,9 +344,191 @@ def _build_list_response(record: dict[str, Any]) -> dict[str, Any]:
|
||||
"recovery": record["recovery"],
|
||||
"artifact_count": len(_collect_artifacts(record)),
|
||||
"state_source": record["state_source"],
|
||||
"status_source": record["status_source"],
|
||||
"state_quality": record["state_quality"],
|
||||
"state_conflict": record["state_conflict"],
|
||||
"state_conflict_reason": record["state_conflict_reason"],
|
||||
"raw_status": record["raw_status"],
|
||||
"raw_current_stage": record["raw_current_stage"],
|
||||
}
|
||||
|
||||
|
||||
def _reconcile_run_record(record: dict[str, Any]) -> dict[str, Any]:
|
||||
raw_status = record["status"]
|
||||
raw_current_stage = record["current_stage"]
|
||||
raw_stages = record["stages"]
|
||||
raw_recovery = record["recovery"]
|
||||
artifact_presence = _build_artifact_presence(record["run_dir"], report=record.get("report"))
|
||||
|
||||
reconciled = dict(record)
|
||||
reconciled["raw_status"] = raw_status
|
||||
reconciled["raw_current_stage"] = raw_current_stage
|
||||
reconciled["status_source"] = record["state_source"]
|
||||
reconciled["state_quality"] = "trusted"
|
||||
reconciled["state_conflict"] = False
|
||||
reconciled["state_conflict_reason"] = None
|
||||
reconciled["artifact_presence"] = artifact_presence
|
||||
|
||||
report = record.get("report")
|
||||
if not isinstance(report, dict):
|
||||
stale_reason = _build_stale_running_reason(record)
|
||||
if stale_reason is not None:
|
||||
reconciled["status"] = "failed"
|
||||
reconciled["finished_at"] = record["updated_at"]
|
||||
reconciled["stages"] = _build_stale_failed_stages(raw_stages, raw_current_stage)
|
||||
reconciled["error"] = {
|
||||
"type": "StaleRunState",
|
||||
"message": stale_reason,
|
||||
"stage": raw_current_stage,
|
||||
"details": {
|
||||
"raw_status": raw_status,
|
||||
"raw_current_stage": raw_current_stage,
|
||||
},
|
||||
}
|
||||
reconciled["status_source"] = "stale_run_state_timeout"
|
||||
reconciled["state_quality"] = "reconciled"
|
||||
reconciled["state_conflict"] = True
|
||||
reconciled["state_conflict_reason"] = stale_reason
|
||||
return reconciled
|
||||
|
||||
effective_status = _infer_status_from_report(report)
|
||||
report_completed_at = _maybe_iso(report.get("completed_at"))
|
||||
conflict = (
|
||||
raw_status != effective_status
|
||||
or raw_current_stage is not None
|
||||
or any(stage["status"] in {"running", "failed"} for stage in raw_stages)
|
||||
or bool(raw_recovery.get("resumable"))
|
||||
)
|
||||
|
||||
if not conflict:
|
||||
reconciled["artifact_presence"] = artifact_presence
|
||||
return reconciled
|
||||
|
||||
reconciled["status"] = effective_status
|
||||
reconciled["current_stage"] = None
|
||||
reconciled["updated_at"] = report_completed_at or record["updated_at"]
|
||||
reconciled["finished_at"] = report_completed_at or record["finished_at"]
|
||||
reconciled["stages"] = _build_terminal_success_stages(raw_stages)
|
||||
reconciled["error"] = None
|
||||
reconciled["recovery"] = {
|
||||
"resumable": False,
|
||||
"resume_from_stage": None,
|
||||
"last_success_stage": DEFAULT_STAGES[-1],
|
||||
}
|
||||
reconciled["status_source"] = "run_report_reconciliation"
|
||||
reconciled["state_quality"] = "reconciled"
|
||||
reconciled["state_conflict"] = True
|
||||
reconciled["state_conflict_reason"] = (
|
||||
"run-state.json did not converge, but run-report.json already proves the workflow reached a terminal state."
|
||||
)
|
||||
return reconciled
|
||||
|
||||
|
||||
def _build_artifact_presence(run_dir: Path, *, report: dict[str, Any] | None) -> dict[str, bool]:
|
||||
return {
|
||||
"run_report": isinstance(report, dict) or (run_dir / "run-report.json").exists(),
|
||||
"delivery_payload": (run_dir / "candidates" / "openclaw-delivery-payload.json").exists(),
|
||||
"digest_brief": (run_dir / "candidates" / "digest-brief.json").exists(),
|
||||
}
|
||||
|
||||
|
||||
def _build_terminal_success_stages(stages: list[dict[str, Any]]) -> list[dict[str, Any]]:
|
||||
stage_by_name = {stage["name"]: stage for stage in stages}
|
||||
reconciled_stages: list[dict[str, Any]] = []
|
||||
for stage_name in DEFAULT_STAGES:
|
||||
existing = stage_by_name.get(stage_name)
|
||||
if existing is None:
|
||||
reconciled_stages.append(_stage_dict(name=stage_name, status="success"))
|
||||
continue
|
||||
reconciled_stages.append(
|
||||
{
|
||||
"name": stage_name,
|
||||
"status": "success",
|
||||
"started_at": existing.get("started_at"),
|
||||
"finished_at": existing.get("finished_at"),
|
||||
"outputs": existing.get("outputs", {}),
|
||||
"error": None,
|
||||
}
|
||||
)
|
||||
return reconciled_stages
|
||||
|
||||
|
||||
def _build_stale_failed_stages(stages: list[dict[str, Any]], current_stage: str | None) -> list[dict[str, Any]]:
|
||||
stage_by_name = {stage["name"]: stage for stage in stages}
|
||||
reconciled_stages: list[dict[str, Any]] = []
|
||||
for stage_name in DEFAULT_STAGES:
|
||||
existing = stage_by_name.get(stage_name)
|
||||
if existing is None:
|
||||
reconciled_stages.append(_stage_dict(name=stage_name, status="pending"))
|
||||
continue
|
||||
|
||||
status = existing.get("status")
|
||||
if status == "running" or (current_stage is not None and stage_name == current_stage):
|
||||
status = "failed"
|
||||
|
||||
reconciled_stages.append(
|
||||
{
|
||||
"name": stage_name,
|
||||
"status": status,
|
||||
"started_at": existing.get("started_at"),
|
||||
"finished_at": existing.get("finished_at") or existing.get("started_at"),
|
||||
"outputs": existing.get("outputs", {}),
|
||||
"error": existing.get("error"),
|
||||
}
|
||||
)
|
||||
return reconciled_stages
|
||||
|
||||
|
||||
def _build_stale_running_reason(record: dict[str, Any]) -> str | None:
|
||||
if record["status"] != "running":
|
||||
return None
|
||||
|
||||
updated_at = _parse_iso_datetime(record.get("updated_at"))
|
||||
if updated_at is None:
|
||||
return None
|
||||
|
||||
stale_after_seconds = _estimate_running_stale_seconds(record)
|
||||
age_seconds = (datetime.now(tz=updated_at.tzinfo) - updated_at).total_seconds()
|
||||
if age_seconds < stale_after_seconds:
|
||||
return None
|
||||
|
||||
current_stage = record.get("current_stage") or "unknown_stage"
|
||||
age_minutes = int(age_seconds // 60)
|
||||
stale_after_minutes = int(stale_after_seconds // 60)
|
||||
return (
|
||||
f"run-state.json has remained in running state at {current_stage} for about {age_minutes} minutes "
|
||||
f"without terminal artifacts; it exceeded the stale threshold of {stale_after_minutes} minutes."
|
||||
)
|
||||
|
||||
|
||||
def _estimate_running_stale_seconds(record: dict[str, Any]) -> int:
|
||||
state = record.get("state")
|
||||
input_payload = state.input if isinstance(state, RunState) and isinstance(state.input, dict) else {}
|
||||
limit = _safe_int(input_payload.get("limit"), default=5)
|
||||
timeout_seconds = _safe_float(input_payload.get("timeout_seconds"), default=60.0)
|
||||
max_retries = _safe_int(input_payload.get("max_retries"), default=2)
|
||||
|
||||
expected_items = limit
|
||||
current_stage = record.get("current_stage")
|
||||
if current_stage == "generate_summaries":
|
||||
expected_items = _stage_output_from_record(record, "generate_summaries", "expected_items") or limit
|
||||
elif current_stage == "extract_articles":
|
||||
expected_items = _stage_output_from_record(record, "extract_articles", "expected_items") or limit
|
||||
|
||||
estimated = int((expected_items * max(timeout_seconds, 1.0) * max(max_retries, 1)) + 15 * 60)
|
||||
return max(MIN_RUNNING_STALE_SECONDS, min(MAX_RUNNING_STALE_SECONDS, estimated))
|
||||
|
||||
|
||||
def _stage_output_from_record(record: dict[str, Any], stage_name: str, key: str) -> Any:
|
||||
for stage in record.get("stages", []):
|
||||
if stage.get("name") != stage_name:
|
||||
continue
|
||||
outputs = stage.get("outputs")
|
||||
if isinstance(outputs, dict):
|
||||
return outputs.get(key)
|
||||
return None
|
||||
|
||||
|
||||
def _collect_artifacts(record: dict[str, Any]) -> list[dict[str, Any]]:
|
||||
artifacts: list[dict[str, Any]] = []
|
||||
seen_names: set[str] = set()
|
||||
@@ -587,6 +783,29 @@ def _maybe_iso(value: Any) -> str | None:
|
||||
return None
|
||||
|
||||
|
||||
def _parse_iso_datetime(value: Any) -> datetime | None:
|
||||
if not isinstance(value, str) or not value.strip():
|
||||
return None
|
||||
try:
|
||||
return datetime.fromisoformat(value)
|
||||
except ValueError:
|
||||
return None
|
||||
|
||||
|
||||
def _safe_int(value: Any, *, default: int) -> int:
|
||||
try:
|
||||
return int(value)
|
||||
except (TypeError, ValueError):
|
||||
return default
|
||||
|
||||
|
||||
def _safe_float(value: Any, *, default: float) -> float:
|
||||
try:
|
||||
return float(value)
|
||||
except (TypeError, ValueError):
|
||||
return default
|
||||
|
||||
|
||||
def _record_sort_key(record: dict[str, Any]) -> tuple[str, str]:
|
||||
return (record.get("updated_at") or "", record["run_dir"].name)
|
||||
|
||||
|
||||
@@ -0,0 +1,614 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import subprocess
|
||||
import sys
|
||||
from datetime import datetime
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
from uuid import uuid4
|
||||
|
||||
from dotenv import dotenv_values
|
||||
|
||||
from .query_service import _resolve_run_record
|
||||
from .resume_service import (
|
||||
SUPPORTED_RESUME_STAGES,
|
||||
UNSUPPORTED_RESUME_STAGES,
|
||||
_build_resume_plan,
|
||||
_resume_freshrss_run,
|
||||
_validate_resume_artifacts,
|
||||
inspect_resume_plan,
|
||||
)
|
||||
from .run_store import RunStore
|
||||
from .state_models import RunState
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[3]
|
||||
OUTPUT_ROOT = REPO_ROOT / "outputs" / "freshrss"
|
||||
RESUME_JOBS_ROOT = OUTPUT_ROOT / "resume_jobs"
|
||||
WORKFLOW_NAME = "freshrss_resume_job"
|
||||
RUN_TYPE = "resume_job"
|
||||
RUN_STATE_FILENAME = "run-state.json"
|
||||
INPUT_FILENAME = "input.json"
|
||||
RESULT_FILENAME = "result.json"
|
||||
JOB_REPORT_FILENAME = "job-report.json"
|
||||
DEFAULT_STAGES = [
|
||||
"prepare_job",
|
||||
"validate_resume_plan",
|
||||
"resume_run",
|
||||
"write_result",
|
||||
]
|
||||
MIN_JOB_STALE_SECONDS = 30 * 60
|
||||
MAX_JOB_STALE_SECONDS = 6 * 60 * 60
|
||||
DEFAULT_DOTENV_PATH = REPO_ROOT / ".env"
|
||||
|
||||
|
||||
def _build_subprocess_env() -> dict[str, str]:
|
||||
"""Build an env dict for subprocess, merging parent env with .env values.
|
||||
|
||||
Ensures the subprocess sees all needed variables (LLM_API_KEY, LLM_MODEL,
|
||||
LLM_API_URL, FRESHRSS_*, etc.) directly in os.environ, avoiding dotenv
|
||||
timing issues in the child process.
|
||||
"""
|
||||
env = os.environ.copy()
|
||||
if DEFAULT_DOTENV_PATH.exists():
|
||||
for key, value in dotenv_values(DEFAULT_DOTENV_PATH).items():
|
||||
if isinstance(key, str) and isinstance(value, str) and value:
|
||||
env.setdefault(key, value)
|
||||
return env
|
||||
|
||||
|
||||
def _now() -> datetime:
|
||||
return datetime.now().astimezone()
|
||||
|
||||
|
||||
def _new_job_id() -> str:
|
||||
ts = _now().strftime("%Y%m%d-%H%M%S")
|
||||
return f"freshrss-resume-job-{ts}-{uuid4().hex[:8]}"
|
||||
|
||||
|
||||
def _job_dir(job_id: str) -> Path:
|
||||
return RESUME_JOBS_ROOT / job_id
|
||||
|
||||
|
||||
def _run_state_path(job_id: str) -> Path:
|
||||
return _job_dir(job_id) / RUN_STATE_FILENAME
|
||||
|
||||
|
||||
def _input_path(job_id: str) -> Path:
|
||||
return _job_dir(job_id) / INPUT_FILENAME
|
||||
|
||||
|
||||
def _result_path(job_id: str) -> Path:
|
||||
return _job_dir(job_id) / RESULT_FILENAME
|
||||
|
||||
|
||||
def _job_report_path(job_id: str) -> Path:
|
||||
return _job_dir(job_id) / JOB_REPORT_FILENAME
|
||||
|
||||
|
||||
def _normalize_repo_path(path: Path) -> str:
|
||||
try:
|
||||
return str(path.resolve().relative_to(REPO_ROOT.resolve()))
|
||||
except ValueError:
|
||||
return str(path)
|
||||
|
||||
|
||||
def _normalize_repo_path_value(path_value: str | None) -> str | None:
|
||||
if not path_value:
|
||||
return None
|
||||
return _normalize_repo_path(Path(path_value))
|
||||
|
||||
|
||||
def _load_run_store(job_id: str) -> RunStore:
|
||||
return RunStore.load(path=_run_state_path(job_id), repo_root=REPO_ROOT)
|
||||
|
||||
|
||||
def _load_result(job_id: str) -> dict[str, Any] | None:
|
||||
path = _result_path(job_id)
|
||||
if not path.exists():
|
||||
return None
|
||||
return json.loads(path.read_text(encoding="utf-8-sig"))
|
||||
|
||||
|
||||
def _load_job_report(job_id: str) -> dict[str, Any] | None:
|
||||
path = _job_report_path(job_id)
|
||||
if not path.exists():
|
||||
return None
|
||||
return json.loads(path.read_text(encoding="utf-8-sig"))
|
||||
|
||||
|
||||
def _write_job_report(*, job_id: str, payload: dict[str, Any]) -> Path:
|
||||
report_file = _job_report_path(job_id)
|
||||
report_file.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||
return report_file
|
||||
|
||||
|
||||
def _build_result_payload(
|
||||
*,
|
||||
job_id: str,
|
||||
run_id: str,
|
||||
run_dir: Path,
|
||||
resume_plan: dict[str, Any],
|
||||
resume_result: dict[str, Any],
|
||||
) -> dict[str, Any]:
|
||||
return {
|
||||
"job_id": job_id,
|
||||
"run_id": run_id,
|
||||
"resume_from_stage": resume_plan["effective_resume_from_stage"],
|
||||
"requested_resume_from_stage": resume_plan["requested_resume_from_stage"],
|
||||
"resume_decision_source": resume_plan["decision_source"],
|
||||
"artifact_resume_from_stage": resume_plan["artifact_resume_from_stage"],
|
||||
"artifact_snapshot": resume_plan["artifact_snapshot"],
|
||||
"status": resume_result["status"],
|
||||
"output_dir": _normalize_repo_path(run_dir),
|
||||
"delivery_output": _normalize_repo_path_value(resume_result.get("delivery_output")),
|
||||
"report_output": _normalize_repo_path_value(resume_result.get("report_output")),
|
||||
"digest_brief_output": _normalize_repo_path_value(resume_result.get("digest_brief_output")),
|
||||
"pulled_count": resume_result.get("pulled_count"),
|
||||
"delivered_count": resume_result.get("delivered_count"),
|
||||
"marked_read_count": resume_result.get("marked_read_count"),
|
||||
"status_counts": resume_result.get("status_counts"),
|
||||
"keyword_index": resume_result.get("keyword_index"),
|
||||
"completed_at": _now().isoformat(),
|
||||
"result_source": "resume_job_result",
|
||||
}
|
||||
|
||||
|
||||
def _build_result_from_linked_run(*, job_id: str, linked_run_record: dict[str, Any]) -> dict[str, Any] | None:
|
||||
report = linked_run_record.get("report")
|
||||
if not isinstance(report, dict):
|
||||
return None
|
||||
|
||||
recovery = linked_run_record.get("recovery")
|
||||
requested_resume_from_stage = None
|
||||
if isinstance(recovery, dict):
|
||||
requested_resume_from_stage = recovery.get("resume_from_stage")
|
||||
|
||||
return {
|
||||
"job_id": job_id,
|
||||
"run_id": linked_run_record["run_id"],
|
||||
"resume_from_stage": requested_resume_from_stage,
|
||||
"requested_resume_from_stage": requested_resume_from_stage,
|
||||
"resume_decision_source": "linked_run_report",
|
||||
"artifact_resume_from_stage": None,
|
||||
"artifact_snapshot": None,
|
||||
"status": linked_run_record["status"],
|
||||
"output_dir": _normalize_repo_path(linked_run_record["run_dir"]),
|
||||
"delivery_output": _normalize_repo_path_value(report.get("delivery_output")),
|
||||
"report_output": _normalize_repo_path(linked_run_record["run_dir"] / "run-report.json"),
|
||||
"digest_brief_output": _normalize_repo_path_value(report.get("digest_brief_output")),
|
||||
"pulled_count": report.get("pulled_count"),
|
||||
"delivered_count": report.get("delivered_count"),
|
||||
"marked_read_count": report.get("marked_read_count"),
|
||||
"status_counts": report.get("status_counts"),
|
||||
"keyword_index": report.get("keyword_index"),
|
||||
"completed_at": report.get("completed_at"),
|
||||
"result_source": "linked_run_report",
|
||||
}
|
||||
|
||||
|
||||
def _linked_run_result_available(linked_run_record: dict[str, Any]) -> bool:
|
||||
return linked_run_record["status"] in {"success", "partial"} and isinstance(linked_run_record.get("report"), dict)
|
||||
|
||||
|
||||
def _load_linked_run_record(state: Any) -> dict[str, Any] | None:
|
||||
if not isinstance(state.input, dict):
|
||||
return None
|
||||
run_id = state.input.get("run_id")
|
||||
if not isinstance(run_id, str) or not run_id.strip():
|
||||
return None
|
||||
try:
|
||||
return _resolve_run_record(run_id)
|
||||
except FileNotFoundError:
|
||||
return None
|
||||
|
||||
|
||||
def _build_stale_job_reason(state: Any) -> str | None:
|
||||
if state.status != "running" or state.current_stage is None or state.finished_at is not None:
|
||||
return None
|
||||
now = _now()
|
||||
age_seconds = max(0.0, (now - state.updated_at).total_seconds())
|
||||
if age_seconds < MIN_JOB_STALE_SECONDS:
|
||||
return None
|
||||
if age_seconds >= MAX_JOB_STALE_SECONDS:
|
||||
return (
|
||||
f"Resume job has remained in stage '{state.current_stage}' for more than {int(MAX_JOB_STALE_SECONDS)} seconds "
|
||||
"without producing a terminal result; treating the job state as stale."
|
||||
)
|
||||
return None
|
||||
|
||||
|
||||
def start_resume_job(*, run_id: str) -> dict[str, Any]:
|
||||
resume_view = inspect_resume_plan(run_id=run_id)
|
||||
job_id = _new_job_id()
|
||||
job_dir = _job_dir(job_id)
|
||||
job_dir.mkdir(parents=True, exist_ok=True)
|
||||
started_at = _now()
|
||||
|
||||
input_payload = {
|
||||
"run_id": run_id,
|
||||
"requested_resume_from_stage": resume_view.get("requested_resume_from_stage"),
|
||||
"resume_from_stage": resume_view.get("resume_from_stage"),
|
||||
"resume_decision_source": resume_view.get("resume_decision_source"),
|
||||
"recommended_action": resume_view.get("recommended_action"),
|
||||
"artifact_resume_from_stage": resume_view.get("artifact_resume_from_stage"),
|
||||
"artifact_snapshot": resume_view.get("artifact_snapshot"),
|
||||
"launcher_pid": os.getpid(),
|
||||
}
|
||||
|
||||
store = RunStore.create(
|
||||
path=_run_state_path(job_id),
|
||||
run_id=job_id,
|
||||
workflow=WORKFLOW_NAME,
|
||||
run_type=RUN_TYPE,
|
||||
started_at=started_at,
|
||||
input_payload=input_payload,
|
||||
repo_root=REPO_ROOT,
|
||||
)
|
||||
for stage_name in DEFAULT_STAGES:
|
||||
store._get_or_create_stage(stage_name)
|
||||
store.save()
|
||||
|
||||
store.start_stage("prepare_job")
|
||||
input_file = _input_path(job_id)
|
||||
input_file.write_text(json.dumps(input_payload, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||
store.register_artifact(name="job_input", path=input_file, kind="json", stage="prepare_job")
|
||||
|
||||
if not resume_view.get("can_resume") or resume_view.get("recommended_action") != "resume":
|
||||
report_file = _write_job_report(
|
||||
job_id=job_id,
|
||||
payload={
|
||||
"job_id": job_id,
|
||||
"status": "failed",
|
||||
"error_type": "ResumePreflightRejected",
|
||||
"error_message": resume_view.get("message"),
|
||||
"failed_stage": "validate_resume_plan",
|
||||
"linked_run_id": run_id,
|
||||
"resume_from_stage": resume_view.get("resume_from_stage"),
|
||||
"requested_resume_from_stage": resume_view.get("requested_resume_from_stage"),
|
||||
"recommended_action": resume_view.get("recommended_action"),
|
||||
},
|
||||
)
|
||||
store.register_artifact(name="job_report", path=report_file, kind="json", stage="prepare_job")
|
||||
store.finish_stage("prepare_job", outputs={"linked_run_id": run_id})
|
||||
store.fail_stage("validate_resume_plan", error=ValueError(str(resume_view.get("message") or "Resume preflight rejected.")))
|
||||
return {
|
||||
"job_id": job_id,
|
||||
"workflow": WORKFLOW_NAME,
|
||||
"run_type": RUN_TYPE,
|
||||
"status": "failed",
|
||||
"output_dir": _normalize_repo_path(job_dir),
|
||||
"linked_run_id": run_id,
|
||||
"resume_from_stage": resume_view.get("resume_from_stage"),
|
||||
"recommended_action": resume_view.get("recommended_action"),
|
||||
"message": resume_view.get("message"),
|
||||
}
|
||||
|
||||
runner_script = REPO_ROOT / "scripts" / "run_resume_job.py"
|
||||
cmd = [sys.executable, str(runner_script), "--job-id", job_id]
|
||||
try:
|
||||
proc = subprocess.Popen(
|
||||
cmd,
|
||||
cwd=str(REPO_ROOT),
|
||||
stdout=subprocess.DEVNULL,
|
||||
stderr=subprocess.DEVNULL,
|
||||
start_new_session=True,
|
||||
env=_build_subprocess_env(),
|
||||
)
|
||||
except Exception as exc:
|
||||
report_file = _write_job_report(
|
||||
job_id=job_id,
|
||||
payload={
|
||||
"job_id": job_id,
|
||||
"status": "failed",
|
||||
"error_type": type(exc).__name__,
|
||||
"error_message": str(exc),
|
||||
"failed_stage": "prepare_job",
|
||||
"linked_run_id": run_id,
|
||||
"resume_from_stage": resume_view.get("resume_from_stage"),
|
||||
},
|
||||
)
|
||||
store.register_artifact(name="job_report", path=report_file, kind="json", stage="prepare_job")
|
||||
store.fail_stage("prepare_job", error=exc)
|
||||
raise
|
||||
|
||||
store.finish_stage(
|
||||
"prepare_job",
|
||||
outputs={
|
||||
"runner_pid": proc.pid,
|
||||
"runner_command": cmd,
|
||||
"linked_run_id": run_id,
|
||||
"resume_from_stage": resume_view.get("resume_from_stage"),
|
||||
},
|
||||
)
|
||||
return {
|
||||
"job_id": job_id,
|
||||
"workflow": WORKFLOW_NAME,
|
||||
"run_type": RUN_TYPE,
|
||||
"status": "running",
|
||||
"output_dir": _normalize_repo_path(job_dir),
|
||||
"linked_run_id": run_id,
|
||||
"resume_from_stage": resume_view.get("resume_from_stage"),
|
||||
"message": "Resume job started successfully. Use get_resume_job_status to poll progress.",
|
||||
}
|
||||
|
||||
|
||||
def run_resume_job(*, job_id: str) -> dict[str, Any]:
|
||||
store = _load_run_store(job_id)
|
||||
input_payload = json.loads(_input_path(job_id).read_text(encoding="utf-8-sig"))
|
||||
current_stage = "resume_run"
|
||||
run_id = str(input_payload["run_id"])
|
||||
|
||||
try:
|
||||
store.start_stage("validate_resume_plan")
|
||||
record = _resolve_run_record(run_id)
|
||||
if record["state_source"] != "run_state" or not isinstance(record.get("state"), RunState):
|
||||
raise RuntimeError("This run cannot be resumed because run-state.json is missing or could not be loaded.")
|
||||
|
||||
run_store = RunStore.load(path=record["run_dir"] / "run-state.json", repo_root=REPO_ROOT)
|
||||
resume_plan = _build_resume_plan(record=record, state=run_store.state)
|
||||
if resume_plan["decision"] != "resume":
|
||||
raise RuntimeError(str(resume_plan["message"]))
|
||||
|
||||
resume_from_stage = resume_plan["effective_resume_from_stage"]
|
||||
if resume_from_stage in UNSUPPORTED_RESUME_STAGES:
|
||||
raise RuntimeError(f"This run cannot be resumed from {resume_from_stage} in the current implementation.")
|
||||
if resume_from_stage not in SUPPORTED_RESUME_STAGES:
|
||||
raise RuntimeError(f"This run cannot be resumed because stage '{resume_from_stage}' is not supported.")
|
||||
|
||||
missing_artifacts = _validate_resume_artifacts(
|
||||
record=record,
|
||||
state=run_store.state,
|
||||
resume_from_stage=resume_from_stage,
|
||||
)
|
||||
if missing_artifacts:
|
||||
raise RuntimeError(
|
||||
f"This run cannot be resumed from {resume_from_stage} because required artifacts are missing: {missing_artifacts}"
|
||||
)
|
||||
|
||||
store.finish_stage(
|
||||
"validate_resume_plan",
|
||||
outputs={
|
||||
"linked_run_id": run_id,
|
||||
"resume_from_stage": resume_from_stage,
|
||||
"requested_resume_from_stage": resume_plan["requested_resume_from_stage"],
|
||||
"resume_decision_source": resume_plan["decision_source"],
|
||||
},
|
||||
)
|
||||
|
||||
current_stage = "resume_run"
|
||||
store.start_stage("resume_run")
|
||||
resume_result = _resume_freshrss_run(record=record, run_store=run_store, resume_from_stage=resume_from_stage)
|
||||
store.finish_stage(
|
||||
"resume_run",
|
||||
outputs={
|
||||
"linked_run_id": run_id,
|
||||
"resume_from_stage": resume_from_stage,
|
||||
"linked_run_status": resume_result["status"],
|
||||
"delivery_output": _normalize_repo_path_value(resume_result.get("delivery_output")),
|
||||
"report_output": _normalize_repo_path_value(resume_result.get("report_output")),
|
||||
},
|
||||
)
|
||||
|
||||
store.start_stage("write_result")
|
||||
result = _build_result_payload(
|
||||
job_id=job_id,
|
||||
run_id=run_id,
|
||||
run_dir=record["run_dir"],
|
||||
resume_plan=resume_plan,
|
||||
resume_result=resume_result,
|
||||
)
|
||||
result_file = _result_path(job_id)
|
||||
result_file.write_text(json.dumps(result, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||
store.register_artifact(name="job_result", path=result_file, kind="json", stage="write_result")
|
||||
|
||||
report_file = _write_job_report(
|
||||
job_id=job_id,
|
||||
payload={
|
||||
"job_id": job_id,
|
||||
"status": "success",
|
||||
"run_id": run_id,
|
||||
"resume_from_stage": result["resume_from_stage"],
|
||||
"delivery_output": result["delivery_output"],
|
||||
"report_output": result["report_output"],
|
||||
"digest_brief_output": result["digest_brief_output"],
|
||||
},
|
||||
)
|
||||
store.register_artifact(name="job_report", path=report_file, kind="json", stage="write_result")
|
||||
store.finish_stage(
|
||||
"write_result",
|
||||
outputs={
|
||||
"result_path": _normalize_repo_path(result_file),
|
||||
"linked_run_id": run_id,
|
||||
},
|
||||
)
|
||||
store.finish_run(status="success")
|
||||
return result
|
||||
except Exception as exc:
|
||||
current_stage = store.state.current_stage or current_stage
|
||||
report_file = _write_job_report(
|
||||
job_id=job_id,
|
||||
payload={
|
||||
"job_id": job_id,
|
||||
"status": "failed",
|
||||
"error_type": type(exc).__name__,
|
||||
"error_message": str(exc),
|
||||
"failed_stage": current_stage,
|
||||
"linked_run_id": run_id,
|
||||
},
|
||||
)
|
||||
try:
|
||||
store.register_artifact(name="job_report", path=report_file, kind="json", stage=current_stage)
|
||||
except Exception:
|
||||
pass
|
||||
store.fail_stage(current_stage, error=exc)
|
||||
raise
|
||||
|
||||
|
||||
def _resolve_effective_job_view(
|
||||
*,
|
||||
job_id: str,
|
||||
state: Any,
|
||||
result: dict[str, Any] | None = None,
|
||||
) -> dict[str, Any]:
|
||||
resolved_result = result or _load_result(job_id)
|
||||
linked_run_record = _load_linked_run_record(state)
|
||||
raw_status = state.status
|
||||
raw_current_stage = state.current_stage
|
||||
effective_status = raw_status
|
||||
effective_current_stage = raw_current_stage
|
||||
effective_updated_at = state.updated_at.isoformat()
|
||||
effective_finished_at = state.finished_at.isoformat() if state.finished_at else None
|
||||
status_source = "job_run_state"
|
||||
state_quality = "trusted"
|
||||
state_conflict = False
|
||||
state_conflict_reason = None
|
||||
status_note = None
|
||||
|
||||
if resolved_result is not None:
|
||||
completed_at = resolved_result.get("completed_at")
|
||||
effective_status = "success"
|
||||
effective_current_stage = None
|
||||
effective_updated_at = completed_at or effective_updated_at
|
||||
effective_finished_at = completed_at or effective_finished_at
|
||||
if raw_status != "success" or raw_current_stage is not None:
|
||||
status_source = "job_result_reconciliation"
|
||||
state_quality = "reconciled"
|
||||
state_conflict = True
|
||||
state_conflict_reason = "result.json already exists, but the resume job run-state did not converge to success."
|
||||
status_note = "resume job result exists; the job can be treated as completed."
|
||||
elif linked_run_record is not None:
|
||||
if _linked_run_result_available(linked_run_record) and raw_status == "running":
|
||||
effective_status = "success"
|
||||
effective_current_stage = None
|
||||
effective_updated_at = linked_run_record.get("updated_at") or effective_updated_at
|
||||
effective_finished_at = linked_run_record.get("finished_at") or effective_finished_at
|
||||
status_source = "linked_run_reconciliation"
|
||||
state_quality = "reconciled"
|
||||
state_conflict = raw_status != "success" or raw_current_stage is not None
|
||||
state_conflict_reason = "The linked run already has terminal artifacts, but the resume job run-state did not converge."
|
||||
status_note = "linked run artifacts are complete; the resume job can be treated as completed."
|
||||
resolved_result = _build_result_from_linked_run(job_id=job_id, linked_run_record=linked_run_record)
|
||||
elif raw_status == "running":
|
||||
stale_reason = _build_stale_job_reason(state)
|
||||
if stale_reason is not None:
|
||||
effective_status = "failed"
|
||||
effective_current_stage = raw_current_stage
|
||||
effective_finished_at = effective_updated_at
|
||||
status_source = "stale_job_state_timeout"
|
||||
state_quality = "reconciled"
|
||||
state_conflict = True
|
||||
state_conflict_reason = stale_reason
|
||||
status_note = stale_reason
|
||||
|
||||
return {
|
||||
"status": effective_status,
|
||||
"current_stage": effective_current_stage,
|
||||
"updated_at": effective_updated_at,
|
||||
"finished_at": effective_finished_at,
|
||||
"status_source": status_source,
|
||||
"state_quality": state_quality,
|
||||
"state_conflict": state_conflict,
|
||||
"state_conflict_reason": state_conflict_reason,
|
||||
"status_note": status_note,
|
||||
"result": resolved_result,
|
||||
"linked_run_status": linked_run_record["status"] if linked_run_record is not None else None,
|
||||
"linked_run_state_conflict": linked_run_record.get("state_conflict") if linked_run_record is not None else None,
|
||||
}
|
||||
|
||||
|
||||
def _build_effective_job_progress(*, state: Any, effective_status: str) -> dict[str, int]:
|
||||
completed_stage_count = sum(1 for stage in state.stages if stage.status == "success")
|
||||
running_stage_count = sum(1 for stage in state.stages if stage.status == "running")
|
||||
failed_stage_count = sum(1 for stage in state.stages if stage.status == "failed")
|
||||
pending_stage_count = sum(1 for stage in state.stages if stage.status == "pending")
|
||||
if effective_status == "success" and running_stage_count > 0:
|
||||
pending_stage_count += running_stage_count
|
||||
running_stage_count = 0
|
||||
return {
|
||||
"completed_stage_count": completed_stage_count,
|
||||
"running_stage_count": running_stage_count,
|
||||
"failed_stage_count": failed_stage_count,
|
||||
"pending_stage_count": pending_stage_count,
|
||||
"total_stage_count": len(state.stages),
|
||||
}
|
||||
|
||||
|
||||
def get_resume_job_status(*, job_id: str) -> dict[str, Any]:
|
||||
store = _load_run_store(job_id)
|
||||
state = store.state
|
||||
effective = _resolve_effective_job_view(job_id=job_id, state=state)
|
||||
linked_run_id = state.input.get("run_id") if isinstance(state.input, dict) else None
|
||||
return {
|
||||
"job_id": state.run_id,
|
||||
"workflow": state.workflow,
|
||||
"run_type": state.run_type,
|
||||
"status": effective["status"],
|
||||
"current_stage": effective["current_stage"],
|
||||
"started_at": state.started_at.isoformat(),
|
||||
"updated_at": effective["updated_at"],
|
||||
"finished_at": effective["finished_at"],
|
||||
"output_dir": _normalize_repo_path(_job_dir(job_id)),
|
||||
"linked_run_id": linked_run_id,
|
||||
"progress": _build_effective_job_progress(state=state, effective_status=effective["status"]),
|
||||
"artifacts": [artifact.model_dump(mode="json") for artifact in state.artifacts],
|
||||
"error_summary": state.error.model_dump(mode="json") if state.error else None,
|
||||
"status_source": effective["status_source"],
|
||||
"state_quality": effective["state_quality"],
|
||||
"state_conflict": effective["state_conflict"],
|
||||
"state_conflict_reason": effective["state_conflict_reason"],
|
||||
"status_note": effective["status_note"],
|
||||
"raw_status": state.status,
|
||||
"raw_current_stage": state.current_stage,
|
||||
"linked_run_status": effective["linked_run_status"],
|
||||
"linked_run_state_conflict": effective["linked_run_state_conflict"],
|
||||
}
|
||||
|
||||
|
||||
def get_resume_job_result(*, job_id: str) -> dict[str, Any]:
|
||||
store = _load_run_store(job_id)
|
||||
state = store.state
|
||||
result = _load_result(job_id)
|
||||
effective = _resolve_effective_job_view(job_id=job_id, state=state, result=result)
|
||||
synthesized_result = result or effective["result"]
|
||||
if effective["status"] != "success" or synthesized_result is None:
|
||||
message = "Resume job result is not ready."
|
||||
if effective["status"] == "failed":
|
||||
message = "Resume job did not complete successfully, so no terminal job result is available."
|
||||
return {
|
||||
"job_id": state.run_id,
|
||||
"status": effective["status"],
|
||||
"message": message,
|
||||
"linked_run_id": state.input.get("run_id") if isinstance(state.input, dict) else None,
|
||||
"error_summary": state.error.model_dump(mode="json") if state.error else None,
|
||||
"status_source": effective["status_source"],
|
||||
"status_note": effective["status_note"],
|
||||
}
|
||||
|
||||
artifact = next((a.model_dump(mode="json") for a in state.artifacts if a.name == "job_result"), None)
|
||||
return {
|
||||
"job_id": state.run_id,
|
||||
"status": effective["status"],
|
||||
"run_id": synthesized_result.get("run_id"),
|
||||
"resume_from_stage": synthesized_result.get("resume_from_stage"),
|
||||
"requested_resume_from_stage": synthesized_result.get("requested_resume_from_stage"),
|
||||
"resume_decision_source": synthesized_result.get("resume_decision_source"),
|
||||
"artifact_resume_from_stage": synthesized_result.get("artifact_resume_from_stage"),
|
||||
"artifact_snapshot": synthesized_result.get("artifact_snapshot"),
|
||||
"output_dir": synthesized_result.get("output_dir"),
|
||||
"delivery_output": synthesized_result.get("delivery_output"),
|
||||
"report_output": synthesized_result.get("report_output"),
|
||||
"digest_brief_output": synthesized_result.get("digest_brief_output"),
|
||||
"pulled_count": synthesized_result.get("pulled_count"),
|
||||
"delivered_count": synthesized_result.get("delivered_count"),
|
||||
"marked_read_count": synthesized_result.get("marked_read_count"),
|
||||
"status_counts": synthesized_result.get("status_counts"),
|
||||
"keyword_index": synthesized_result.get("keyword_index"),
|
||||
"artifact": artifact,
|
||||
"result": synthesized_result,
|
||||
"status_source": effective["status_source"],
|
||||
"status_note": effective["status_note"],
|
||||
"result_source": synthesized_result.get("result_source", "job_result"),
|
||||
"linked_run_status": effective["linked_run_status"],
|
||||
}
|
||||
@@ -25,6 +25,7 @@ from summary_mcp.models.openclaw_delivery import (
|
||||
)
|
||||
from summary_mcp.models.summary_io import ExtractionOutput
|
||||
from summary_mcp.workflows.freshrss_pipeline import (
|
||||
CANDIDATE_BATCH_ARTIFACT,
|
||||
DEFAULT_PROMPT_PATH,
|
||||
DEFAULT_RULES_PATH,
|
||||
DEFAULT_TERM_ALIASES_PATH,
|
||||
@@ -37,13 +38,18 @@ from summary_mcp.workflows.freshrss_pipeline import (
|
||||
FILTER_STAGE,
|
||||
REPORT_STAGE,
|
||||
REPO_ROOT,
|
||||
SUMMARY_BATCH_ARTIFACT,
|
||||
SUMMARY_STAGE,
|
||||
WORKFLOW_NAME,
|
||||
_build_item_context,
|
||||
_candidate_batch_output,
|
||||
_final_run_status,
|
||||
_load_json,
|
||||
_load_required_env,
|
||||
_persist_candidate_batch_artifact,
|
||||
_persist_summary_batch_artifact,
|
||||
_save_json,
|
||||
_summary_batch_output,
|
||||
)
|
||||
|
||||
from .query_service import _normalize_repo_path, _resolve_repo_path, _resolve_run_record
|
||||
@@ -62,6 +68,12 @@ UNSUPPORTED_RESUME_STAGES = {
|
||||
}
|
||||
UTC = timezone.utc
|
||||
DEFAULT_STREAM_ID = "user/-/state/com.google/reading-list"
|
||||
RESUME_STAGE_ORDER = {
|
||||
SUMMARY_STAGE: 1,
|
||||
FILTER_STAGE: 2,
|
||||
DELIVERY_STAGE: 3,
|
||||
REPORT_STAGE: 4,
|
||||
}
|
||||
|
||||
|
||||
def resume_run(*, run_id: str) -> dict[str, Any]:
|
||||
@@ -71,6 +83,22 @@ def resume_run(*, run_id: str) -> dict[str, Any]:
|
||||
"workflow": record["workflow"],
|
||||
"status": record["status"],
|
||||
"output_dir": _normalize_repo_path(record["run_dir"]),
|
||||
"state_source": record["state_source"],
|
||||
"status_source": record.get("status_source"),
|
||||
}
|
||||
|
||||
if record["status"] in {"success", "partial"} and isinstance(record.get("report"), dict):
|
||||
return {
|
||||
**base_response,
|
||||
"resumed": False,
|
||||
"resume_from_stage": None,
|
||||
"requested_resume_from_stage": None,
|
||||
"resume_decision_source": "artifacts",
|
||||
"recommended_action": "read_terminal_result",
|
||||
"message": "This run already has a terminal run-report.json; prefer reading get_run_status/get_run_report instead of resuming.",
|
||||
"missing_artifacts": [],
|
||||
"state_conflict": record.get("state_conflict", False),
|
||||
"state_conflict_reason": record.get("state_conflict_reason"),
|
||||
}
|
||||
|
||||
if record["state_source"] != "run_state" or not isinstance(record.get("state"), RunState):
|
||||
@@ -78,30 +106,45 @@ def resume_run(*, run_id: str) -> dict[str, Any]:
|
||||
**base_response,
|
||||
"resumed": False,
|
||||
"resume_from_stage": None,
|
||||
"requested_resume_from_stage": None,
|
||||
"resume_decision_source": "unavailable",
|
||||
"recommended_action": "start_new_run",
|
||||
"message": "This run cannot be resumed because run-state.json is missing or could not be loaded.",
|
||||
"missing_artifacts": [],
|
||||
}
|
||||
|
||||
run_store = RunStore.load(path=record["run_dir"] / "run-state.json", repo_root=REPO_ROOT)
|
||||
state = run_store.state
|
||||
resume_from_stage = _resolve_resume_from_stage(state)
|
||||
resume_plan = _build_resume_plan(record=record, state=state)
|
||||
requested_resume_from_stage = resume_plan["requested_resume_from_stage"]
|
||||
resume_from_stage = resume_plan["effective_resume_from_stage"]
|
||||
|
||||
if state.workflow != WORKFLOW_NAME:
|
||||
return {
|
||||
**base_response,
|
||||
"resumed": False,
|
||||
"resume_from_stage": resume_from_stage,
|
||||
"requested_resume_from_stage": requested_resume_from_stage,
|
||||
"resume_decision_source": resume_plan["decision_source"],
|
||||
"recommended_action": "start_new_run",
|
||||
"message": f"This run cannot be resumed because workflow '{state.workflow}' is not supported by the minimal resume_run implementation.",
|
||||
"missing_artifacts": [],
|
||||
}
|
||||
|
||||
if resume_from_stage is None:
|
||||
if resume_plan["decision"] != "resume":
|
||||
return {
|
||||
**base_response,
|
||||
"resumed": False,
|
||||
"resume_from_stage": None,
|
||||
"message": "This run does not expose a recoverable stage in run-state.json.",
|
||||
"missing_artifacts": [],
|
||||
"resume_from_stage": resume_from_stage,
|
||||
"requested_resume_from_stage": requested_resume_from_stage,
|
||||
"resume_decision_source": resume_plan["decision_source"],
|
||||
"recommended_action": resume_plan["recommended_action"],
|
||||
"message": resume_plan["message"],
|
||||
"missing_artifacts": resume_plan["missing_artifacts"],
|
||||
"artifact_resume_from_stage": resume_plan["artifact_resume_from_stage"],
|
||||
"artifact_snapshot": resume_plan["artifact_snapshot"],
|
||||
"state_conflict": record.get("state_conflict", False),
|
||||
"state_conflict_reason": record.get("state_conflict_reason"),
|
||||
}
|
||||
|
||||
if resume_from_stage in UNSUPPORTED_RESUME_STAGES:
|
||||
@@ -109,6 +152,9 @@ def resume_run(*, run_id: str) -> dict[str, Any]:
|
||||
**base_response,
|
||||
"resumed": False,
|
||||
"resume_from_stage": resume_from_stage,
|
||||
"requested_resume_from_stage": requested_resume_from_stage,
|
||||
"resume_decision_source": resume_plan["decision_source"],
|
||||
"recommended_action": "start_new_run",
|
||||
"message": f"This run cannot be resumed from {resume_from_stage} in the current minimal implementation.",
|
||||
"missing_artifacts": [],
|
||||
}
|
||||
@@ -118,6 +164,9 @@ def resume_run(*, run_id: str) -> dict[str, Any]:
|
||||
**base_response,
|
||||
"resumed": False,
|
||||
"resume_from_stage": resume_from_stage,
|
||||
"requested_resume_from_stage": requested_resume_from_stage,
|
||||
"resume_decision_source": resume_plan["decision_source"],
|
||||
"recommended_action": "start_new_run",
|
||||
"message": f"This run cannot be resumed because stage '{resume_from_stage}' is not supported.",
|
||||
"missing_artifacts": [],
|
||||
}
|
||||
@@ -128,8 +177,13 @@ def resume_run(*, run_id: str) -> dict[str, Any]:
|
||||
**base_response,
|
||||
"resumed": False,
|
||||
"resume_from_stage": resume_from_stage,
|
||||
"requested_resume_from_stage": requested_resume_from_stage,
|
||||
"resume_decision_source": resume_plan["decision_source"],
|
||||
"recommended_action": "start_new_run",
|
||||
"message": f"This run cannot be resumed from {resume_from_stage} because required artifacts are missing.",
|
||||
"missing_artifacts": missing_artifacts,
|
||||
"artifact_resume_from_stage": resume_plan["artifact_resume_from_stage"],
|
||||
"artifact_snapshot": resume_plan["artifact_snapshot"],
|
||||
}
|
||||
|
||||
try:
|
||||
@@ -141,10 +195,15 @@ def resume_run(*, run_id: str) -> dict[str, Any]:
|
||||
**base_response,
|
||||
"resumed": True,
|
||||
"resume_from_stage": resume_from_stage,
|
||||
"requested_resume_from_stage": requested_resume_from_stage,
|
||||
"resume_decision_source": resume_plan["decision_source"],
|
||||
"recommended_action": "inspect_error",
|
||||
"status": run_store.state.status,
|
||||
"message": f"Run resumed from {resume_from_stage} but failed again at {failed_stage}: {error}",
|
||||
"missing_artifacts": [],
|
||||
"error_summary": run_store.state.error.model_dump(mode="json") if run_store.state.error else None,
|
||||
"artifact_resume_from_stage": resume_plan["artifact_resume_from_stage"],
|
||||
"artifact_snapshot": resume_plan["artifact_snapshot"],
|
||||
}
|
||||
|
||||
return {
|
||||
@@ -152,6 +211,9 @@ def resume_run(*, run_id: str) -> dict[str, Any]:
|
||||
"workflow": run_store.state.workflow,
|
||||
"resumed": True,
|
||||
"resume_from_stage": resume_from_stage,
|
||||
"requested_resume_from_stage": requested_resume_from_stage,
|
||||
"resume_decision_source": resume_plan["decision_source"],
|
||||
"recommended_action": "none",
|
||||
"status": result["status"],
|
||||
"output_dir": _normalize_repo_path(record["run_dir"]),
|
||||
"message": f"Run resumed from {resume_from_stage} and completed with status {result['status']}.",
|
||||
@@ -166,9 +228,92 @@ def resume_run(*, run_id: str) -> dict[str, Any]:
|
||||
"marked_read_count": result.get("marked_read_count"),
|
||||
"status_counts": result.get("status_counts"),
|
||||
"missing_artifacts": [],
|
||||
"artifact_resume_from_stage": resume_plan["artifact_resume_from_stage"],
|
||||
"artifact_snapshot": resume_plan["artifact_snapshot"],
|
||||
}
|
||||
|
||||
|
||||
def inspect_resume_plan(*, run_id: str) -> dict[str, Any]:
|
||||
record = _resolve_run_record(run_id)
|
||||
base_response = {
|
||||
"run_id": record["run_id"],
|
||||
"workflow": record["workflow"],
|
||||
"status": record["status"],
|
||||
"output_dir": _normalize_repo_path(record["run_dir"]),
|
||||
"state_source": record["state_source"],
|
||||
"status_source": record.get("status_source"),
|
||||
"state_conflict": record.get("state_conflict", False),
|
||||
"state_conflict_reason": record.get("state_conflict_reason"),
|
||||
}
|
||||
|
||||
if record["status"] in {"success", "partial"} and isinstance(record.get("report"), dict):
|
||||
return {
|
||||
**base_response,
|
||||
"can_resume": False,
|
||||
"resume_from_stage": None,
|
||||
"requested_resume_from_stage": None,
|
||||
"resume_decision_source": "artifacts",
|
||||
"recommended_action": "read_terminal_result",
|
||||
"message": "This run already has a terminal run-report.json; prefer reading get_run_status/get_run_report instead of resuming.",
|
||||
"missing_artifacts": [],
|
||||
}
|
||||
|
||||
if record["state_source"] != "run_state" or not isinstance(record.get("state"), RunState):
|
||||
return {
|
||||
**base_response,
|
||||
"can_resume": False,
|
||||
"resume_from_stage": None,
|
||||
"requested_resume_from_stage": None,
|
||||
"resume_decision_source": "unavailable",
|
||||
"recommended_action": "start_new_run",
|
||||
"message": "This run cannot be resumed because run-state.json is missing or could not be loaded.",
|
||||
"missing_artifacts": [],
|
||||
}
|
||||
|
||||
state = record["state"]
|
||||
resume_plan = _build_resume_plan(record=record, state=state)
|
||||
response = {
|
||||
**base_response,
|
||||
"can_resume": False,
|
||||
"resume_from_stage": resume_plan["effective_resume_from_stage"],
|
||||
"requested_resume_from_stage": resume_plan["requested_resume_from_stage"],
|
||||
"resume_decision_source": resume_plan["decision_source"],
|
||||
"recommended_action": resume_plan["recommended_action"],
|
||||
"message": resume_plan["message"],
|
||||
"missing_artifacts": resume_plan["missing_artifacts"],
|
||||
"artifact_resume_from_stage": resume_plan["artifact_resume_from_stage"],
|
||||
"artifact_snapshot": resume_plan["artifact_snapshot"],
|
||||
}
|
||||
|
||||
if state.workflow != WORKFLOW_NAME:
|
||||
response["message"] = (
|
||||
f"This run cannot be resumed because workflow '{state.workflow}' is not supported by the minimal resume_run implementation."
|
||||
)
|
||||
return response
|
||||
|
||||
resume_from_stage = resume_plan["effective_resume_from_stage"]
|
||||
if resume_plan["decision"] != "resume":
|
||||
return response
|
||||
if resume_from_stage in UNSUPPORTED_RESUME_STAGES:
|
||||
response["message"] = f"This run cannot be resumed from {resume_from_stage} in the current minimal implementation."
|
||||
response["recommended_action"] = "start_new_run"
|
||||
return response
|
||||
if resume_from_stage not in SUPPORTED_RESUME_STAGES:
|
||||
response["message"] = f"This run cannot be resumed because stage '{resume_from_stage}' is not supported."
|
||||
response["recommended_action"] = "start_new_run"
|
||||
return response
|
||||
|
||||
missing_artifacts = _validate_resume_artifacts(record=record, state=state, resume_from_stage=resume_from_stage)
|
||||
if missing_artifacts:
|
||||
response["message"] = f"This run cannot be resumed from {resume_from_stage} because required artifacts are missing."
|
||||
response["missing_artifacts"] = missing_artifacts
|
||||
response["recommended_action"] = "start_new_run"
|
||||
return response
|
||||
|
||||
response["can_resume"] = True
|
||||
return response
|
||||
|
||||
|
||||
def _resume_freshrss_run(*, record: dict[str, Any], run_store: RunStore, resume_from_stage: str) -> dict[str, Any]:
|
||||
state = run_store.state
|
||||
run_dir = record["run_dir"]
|
||||
@@ -304,6 +449,75 @@ def _build_resume_config(*, state: RunState, run_dir: Path) -> dict[str, Any]:
|
||||
}
|
||||
|
||||
|
||||
def _find_summary_batch_path(*, run_dir: Path, state: RunState) -> Path | None:
|
||||
return _find_artifact_path(
|
||||
run_dir=run_dir,
|
||||
state=state,
|
||||
artifact_name=SUMMARY_BATCH_ARTIFACT,
|
||||
relative_path=Path("summary/summary-batch.json"),
|
||||
)
|
||||
|
||||
|
||||
def _find_candidate_batch_path(*, run_dir: Path, state: RunState) -> Path | None:
|
||||
return _find_artifact_path(
|
||||
run_dir=run_dir,
|
||||
state=state,
|
||||
artifact_name=CANDIDATE_BATCH_ARTIFACT,
|
||||
relative_path=Path("candidates/candidate-batch.json"),
|
||||
)
|
||||
|
||||
|
||||
def _load_summary_batch_lookup(path: Path | None) -> tuple[bool, dict[str, dict[str, Any]]]:
|
||||
if path is None or not path.exists():
|
||||
return False, {}
|
||||
try:
|
||||
payload = _load_json(path)
|
||||
except Exception:
|
||||
return False, {}
|
||||
|
||||
items = payload.get("items")
|
||||
if not isinstance(items, list):
|
||||
return False, {}
|
||||
|
||||
summaries_by_item_key: dict[str, dict[str, Any]] = {}
|
||||
for entry in items:
|
||||
if not isinstance(entry, dict):
|
||||
return False, {}
|
||||
item_key = entry.get("item_key")
|
||||
summary = entry.get("summary")
|
||||
if not isinstance(item_key, str) or not item_key.strip() or not isinstance(summary, dict):
|
||||
return False, {}
|
||||
summaries_by_item_key[item_key] = summary
|
||||
return True, summaries_by_item_key
|
||||
|
||||
|
||||
def _load_candidate_batch_lookup(path: Path | None) -> tuple[bool, dict[str, OpenClawCandidateInput]]:
|
||||
if path is None or not path.exists():
|
||||
return False, {}
|
||||
try:
|
||||
payload = _load_json(path)
|
||||
except Exception:
|
||||
return False, {}
|
||||
|
||||
items = payload.get("items")
|
||||
if not isinstance(items, list):
|
||||
return False, {}
|
||||
|
||||
candidates_by_item_key: dict[str, OpenClawCandidateInput] = {}
|
||||
try:
|
||||
for entry in items:
|
||||
if not isinstance(entry, dict):
|
||||
return False, {}
|
||||
item_key = entry.get("item_key")
|
||||
candidate_payload = entry.get("candidate")
|
||||
if not isinstance(item_key, str) or not item_key.strip() or not isinstance(candidate_payload, dict):
|
||||
return False, {}
|
||||
candidates_by_item_key[item_key] = OpenClawCandidateInput.model_validate(candidate_payload)
|
||||
except Exception:
|
||||
return False, {}
|
||||
return True, candidates_by_item_key
|
||||
|
||||
|
||||
def _load_item_contexts(
|
||||
*,
|
||||
items: list[Any],
|
||||
@@ -313,6 +527,8 @@ def _load_item_contexts(
|
||||
include_candidates: bool = False,
|
||||
delivery_candidates_by_id: dict[str, OpenClawCandidateInput] | None = None,
|
||||
) -> list[dict[str, Any]]:
|
||||
summary_batch_valid, summary_batch_by_item_key = _load_summary_batch_lookup(_summary_batch_output(run_dir))
|
||||
candidate_batch_valid, candidate_batch_by_item_key = _load_candidate_batch_lookup(_candidate_batch_output(run_dir))
|
||||
contexts: list[dict[str, Any]] = []
|
||||
for index, item in enumerate(items, start=1):
|
||||
context = _build_item_context(
|
||||
@@ -334,16 +550,24 @@ def _load_item_contexts(
|
||||
else:
|
||||
context["item_report"]["status"] = "extracted"
|
||||
|
||||
if include_summaries and context["summary_output"] is not None and context["summary_output"].exists():
|
||||
if include_summaries:
|
||||
summary_payload: dict[str, Any] | None = None
|
||||
if context["summary_output"] is not None and context["summary_output"].exists():
|
||||
summary_payload = _load_json(context["summary_output"])
|
||||
elif summary_batch_valid:
|
||||
summary_payload = summary_batch_by_item_key.get(context["item_key"])
|
||||
|
||||
if summary_payload is not None:
|
||||
context["summary_payload"] = summary_payload
|
||||
context["item_report"]["status"] = "summarized"
|
||||
elif include_summaries and context.get("extraction") is not None and context["extraction"].success:
|
||||
elif context.get("extraction") is not None and context["extraction"].success:
|
||||
context["item_report"]["status"] = "summary_failed"
|
||||
|
||||
candidate: OpenClawCandidateInput | None = None
|
||||
if include_candidates and context["openclaw_path"] is not None and context["openclaw_path"].exists():
|
||||
candidate = OpenClawCandidateInput.model_validate(_load_json(context["openclaw_path"]))
|
||||
elif include_candidates and candidate_batch_valid:
|
||||
candidate = candidate_batch_by_item_key.get(context["item_key"])
|
||||
elif delivery_candidates_by_id is not None and context.get("extraction") is not None and context["extraction"].success:
|
||||
candidate_id = candidate_id_for(item, context["extraction"].article)
|
||||
candidate = delivery_candidates_by_id.get(candidate_id)
|
||||
@@ -406,6 +630,11 @@ def _run_summary_stage(*, run_store: RunStore, item_contexts: list[dict[str, Any
|
||||
},
|
||||
)
|
||||
|
||||
summary_batch_output = _persist_summary_batch_artifact(
|
||||
run_store=run_store,
|
||||
run_dir=config["run_dir"],
|
||||
item_contexts=item_contexts,
|
||||
)
|
||||
if config["debug_artifacts"] and (config["run_dir"] / "summary").exists():
|
||||
run_store.register_artifact(name="summary_dir", path=config["run_dir"] / "summary", kind="directory", stage=SUMMARY_STAGE)
|
||||
run_store.finish_stage(
|
||||
@@ -415,6 +644,7 @@ def _run_summary_stage(*, run_store: RunStore, item_contexts: list[dict[str, Any
|
||||
"completed_items": summary_success_count + summary_failed_count,
|
||||
"success_count": summary_success_count,
|
||||
"failed_count": summary_failed_count,
|
||||
"summary_batch_output": str(summary_batch_output),
|
||||
},
|
||||
)
|
||||
|
||||
@@ -500,6 +730,11 @@ def _run_filter_stage(*, run_store: RunStore, item_contexts: list[dict[str, Any]
|
||||
},
|
||||
)
|
||||
|
||||
candidate_batch_output = _persist_candidate_batch_artifact(
|
||||
run_store=run_store,
|
||||
run_dir=config["run_dir"],
|
||||
item_contexts=item_contexts,
|
||||
)
|
||||
if config["debug_artifacts"] and (config["run_dir"] / "candidates").exists():
|
||||
run_store.register_artifact(name="candidate_dir", path=config["run_dir"] / "candidates", kind="directory", stage=FILTER_STAGE)
|
||||
run_store.finish_stage(
|
||||
@@ -511,6 +746,7 @@ def _run_filter_stage(*, run_store: RunStore, item_contexts: list[dict[str, Any]
|
||||
"keep_count": keep_count,
|
||||
"review_count": review_count,
|
||||
"drop_count": drop_count,
|
||||
"candidate_batch_output": str(candidate_batch_output),
|
||||
},
|
||||
)
|
||||
|
||||
@@ -682,6 +918,165 @@ def _build_keyword_index_result(
|
||||
return result
|
||||
|
||||
|
||||
def _build_resume_plan(*, record: dict[str, Any], state: RunState) -> dict[str, Any]:
|
||||
requested_resume_from_stage = _resolve_resume_from_stage(state)
|
||||
artifact_snapshot = _collect_resume_artifact_snapshot(record=record, state=state)
|
||||
artifact_resume_from_stage = _resolve_resume_stage_from_artifacts(artifact_snapshot)
|
||||
|
||||
if artifact_snapshot["has_run_report"]:
|
||||
return {
|
||||
"decision": "reject_terminal",
|
||||
"decision_source": "artifacts",
|
||||
"requested_resume_from_stage": requested_resume_from_stage,
|
||||
"effective_resume_from_stage": None,
|
||||
"artifact_resume_from_stage": artifact_resume_from_stage,
|
||||
"recommended_action": "read_terminal_result",
|
||||
"message": "This run already has a terminal run-report.json; prefer reading get_run_status/get_run_report instead of resuming.",
|
||||
"missing_artifacts": [],
|
||||
"artifact_snapshot": artifact_snapshot,
|
||||
}
|
||||
|
||||
if artifact_resume_from_stage is None:
|
||||
return {
|
||||
"decision": "reject_unrecoverable",
|
||||
"decision_source": "artifacts",
|
||||
"requested_resume_from_stage": requested_resume_from_stage,
|
||||
"effective_resume_from_stage": None,
|
||||
"artifact_resume_from_stage": None,
|
||||
"recommended_action": "start_new_run",
|
||||
"message": "This run does not expose a safe artifact-backed resume point; start a new run instead.",
|
||||
"missing_artifacts": artifact_snapshot["missing_for_next_resume"],
|
||||
"artifact_snapshot": artifact_snapshot,
|
||||
}
|
||||
|
||||
decision_source = "artifacts"
|
||||
message = f"Resume will continue from {artifact_resume_from_stage} based on available artifacts."
|
||||
if requested_resume_from_stage == artifact_resume_from_stage:
|
||||
decision_source = "state_and_artifacts"
|
||||
message = f"Resume stage {artifact_resume_from_stage} was confirmed by both run-state.json and artifacts."
|
||||
elif requested_resume_from_stage is not None:
|
||||
requested_rank = RESUME_STAGE_ORDER.get(requested_resume_from_stage, -1)
|
||||
artifact_rank = RESUME_STAGE_ORDER.get(artifact_resume_from_stage, -1)
|
||||
if artifact_rank > requested_rank:
|
||||
message = (
|
||||
f"Resume stage was advanced from {requested_resume_from_stage} to {artifact_resume_from_stage} "
|
||||
f"because artifacts prove the run already progressed further."
|
||||
)
|
||||
else:
|
||||
message = (
|
||||
f"Resume stage was moved back from {requested_resume_from_stage} to {artifact_resume_from_stage} "
|
||||
f"because later-stage artifacts are not stable enough for a safe resume."
|
||||
)
|
||||
|
||||
return {
|
||||
"decision": "resume",
|
||||
"decision_source": decision_source,
|
||||
"requested_resume_from_stage": requested_resume_from_stage,
|
||||
"effective_resume_from_stage": artifact_resume_from_stage,
|
||||
"artifact_resume_from_stage": artifact_resume_from_stage,
|
||||
"recommended_action": "resume",
|
||||
"message": message,
|
||||
"missing_artifacts": [],
|
||||
"artifact_snapshot": artifact_snapshot,
|
||||
}
|
||||
|
||||
|
||||
def _collect_resume_artifact_snapshot(*, record: dict[str, Any], state: RunState) -> dict[str, Any]:
|
||||
run_dir = record["run_dir"]
|
||||
raw_output = _find_artifact_path(run_dir=run_dir, state=state, artifact_name="raw_output", relative_path=Path("raw/freshrss.raw.json"))
|
||||
extracted_dir = _find_artifact_path(run_dir=run_dir, state=state, artifact_name="extracted_dir", relative_path=Path("extracted"))
|
||||
summary_batch_output = _find_summary_batch_path(run_dir=run_dir, state=state)
|
||||
candidate_batch_output = _find_candidate_batch_path(run_dir=run_dir, state=state)
|
||||
delivery_output = _find_artifact_path(
|
||||
run_dir=run_dir,
|
||||
state=state,
|
||||
artifact_name="delivery_payload",
|
||||
relative_path=Path("candidates/openclaw-delivery-payload.json"),
|
||||
)
|
||||
digest_brief_output = _find_artifact_path(
|
||||
run_dir=run_dir,
|
||||
state=state,
|
||||
artifact_name="digest_brief",
|
||||
relative_path=Path("candidates/digest-brief.json"),
|
||||
)
|
||||
report_output = run_dir / "run-report.json"
|
||||
|
||||
items: list[Any] = []
|
||||
raw_output_valid = False
|
||||
if raw_output is not None:
|
||||
try:
|
||||
items = _load_items(raw_output)
|
||||
raw_output_valid = True
|
||||
except Exception:
|
||||
items = []
|
||||
raw_output_valid = False
|
||||
|
||||
summary_batch_valid, _ = _load_summary_batch_lookup(summary_batch_output)
|
||||
candidate_batch_valid, _ = _load_candidate_batch_lookup(candidate_batch_output)
|
||||
item_contexts = _load_item_contexts(
|
||||
items=items,
|
||||
run_dir=run_dir,
|
||||
debug_artifacts=bool(state.input.get("debug_artifacts", False)),
|
||||
include_summaries=True,
|
||||
include_candidates=True,
|
||||
)
|
||||
extraction_success_count = sum(
|
||||
1 for context in item_contexts if context.get("extraction") is not None and context["extraction"].success
|
||||
)
|
||||
summary_count = sum(1 for context in item_contexts if context.get("summary_payload") is not None)
|
||||
candidate_count = sum(1 for context in item_contexts if context.get("candidate") is not None)
|
||||
extracted_complete = raw_output_valid and bool(items) and all(context["extracted_path"].exists() for context in item_contexts)
|
||||
stable_summary_outputs = summary_batch_valid or (extraction_success_count > 0 and extraction_success_count == summary_count)
|
||||
stable_candidate_outputs = candidate_batch_valid or (summary_count > 0 and summary_count == candidate_count)
|
||||
|
||||
missing_for_next_resume: list[str] = []
|
||||
if raw_output is None or not raw_output_valid:
|
||||
missing_for_next_resume.append("raw/freshrss.raw.json")
|
||||
if extracted_dir is None or not extracted_complete:
|
||||
missing_for_next_resume.append("extracted/")
|
||||
if extraction_success_count > 0 and not stable_summary_outputs and not stable_candidate_outputs:
|
||||
missing_for_next_resume.append("summary/summary-batch.json")
|
||||
if extraction_success_count > 0 and stable_summary_outputs and not stable_candidate_outputs:
|
||||
missing_for_next_resume.append("candidates/candidate-batch.json")
|
||||
|
||||
return {
|
||||
"has_raw_output": raw_output is not None,
|
||||
"raw_output_valid": raw_output_valid,
|
||||
"has_extracted_dir": extracted_dir is not None,
|
||||
"has_summary_batch": summary_batch_output is not None,
|
||||
"summary_batch_valid": summary_batch_valid,
|
||||
"has_candidate_batch": candidate_batch_output is not None,
|
||||
"candidate_batch_valid": candidate_batch_valid,
|
||||
"has_delivery_payload": delivery_output is not None,
|
||||
"has_digest_brief": digest_brief_output is not None,
|
||||
"has_run_report": report_output.exists(),
|
||||
"item_count": len(items),
|
||||
"extraction_success_count": extraction_success_count,
|
||||
"summary_count": summary_count,
|
||||
"candidate_count": candidate_count,
|
||||
"extracted_complete": extracted_complete,
|
||||
"stable_summary_outputs": stable_summary_outputs,
|
||||
"stable_candidate_outputs": stable_candidate_outputs,
|
||||
"missing_for_next_resume": missing_for_next_resume,
|
||||
}
|
||||
|
||||
|
||||
def _resolve_resume_stage_from_artifacts(snapshot: dict[str, Any]) -> str | None:
|
||||
if snapshot["has_run_report"]:
|
||||
return None
|
||||
if snapshot["has_delivery_payload"]:
|
||||
return REPORT_STAGE
|
||||
if snapshot["stable_candidate_outputs"]:
|
||||
return DELIVERY_STAGE
|
||||
if snapshot["has_raw_output"] and snapshot["raw_output_valid"] and snapshot["has_extracted_dir"] and snapshot["extracted_complete"]:
|
||||
if snapshot["extraction_success_count"] <= 0:
|
||||
return SUMMARY_STAGE
|
||||
if snapshot["stable_summary_outputs"]:
|
||||
return FILTER_STAGE
|
||||
return SUMMARY_STAGE
|
||||
return None
|
||||
|
||||
|
||||
def _load_filter_context(config: dict[str, Any]) -> FilterContext:
|
||||
if config["context"] is not None:
|
||||
return FilterContext.model_validate(config["context"])
|
||||
@@ -695,6 +1090,8 @@ def _validate_resume_artifacts(*, record: dict[str, Any], state: RunState, resum
|
||||
missing_artifacts: list[str] = []
|
||||
raw_output = _find_artifact_path(run_dir=run_dir, state=state, artifact_name="raw_output", relative_path=Path("raw/freshrss.raw.json"))
|
||||
extracted_dir = _find_artifact_path(run_dir=run_dir, state=state, artifact_name="extracted_dir", relative_path=Path("extracted"))
|
||||
summary_batch_output = _find_summary_batch_path(run_dir=run_dir, state=state)
|
||||
candidate_batch_output = _find_candidate_batch_path(run_dir=run_dir, state=state)
|
||||
if resume_from_stage in {SUMMARY_STAGE, FILTER_STAGE, DELIVERY_STAGE} and raw_output is None:
|
||||
missing_artifacts.append("raw/freshrss.raw.json")
|
||||
if resume_from_stage in {SUMMARY_STAGE, FILTER_STAGE} and extracted_dir is None:
|
||||
@@ -711,6 +1108,8 @@ def _validate_resume_artifacts(*, record: dict[str, Any], state: RunState, resum
|
||||
include_summaries=resume_from_stage in {FILTER_STAGE, DELIVERY_STAGE},
|
||||
include_candidates=resume_from_stage == DELIVERY_STAGE,
|
||||
)
|
||||
summary_batch_valid, _ = _load_summary_batch_lookup(summary_batch_output)
|
||||
candidate_batch_valid, _ = _load_candidate_batch_lookup(candidate_batch_output)
|
||||
|
||||
if resume_from_stage == SUMMARY_STAGE:
|
||||
for context in item_contexts:
|
||||
@@ -718,6 +1117,12 @@ def _validate_resume_artifacts(*, record: dict[str, Any], state: RunState, resum
|
||||
missing_artifacts.append(_normalize_repo_path(context["extracted_path"]))
|
||||
|
||||
if resume_from_stage == FILTER_STAGE:
|
||||
expected_summary_count = int(_stage_output(state, SUMMARY_STAGE, "success_count") or 0)
|
||||
actual_summary_count = sum(1 for context in item_contexts if context.get("summary_payload") is not None)
|
||||
if summary_batch_valid:
|
||||
if expected_summary_count != actual_summary_count:
|
||||
missing_artifacts.append(_normalize_repo_path(summary_batch_output or _summary_batch_output(run_dir)))
|
||||
else:
|
||||
for context in item_contexts:
|
||||
if context.get("extraction") is not None and context["extraction"].success:
|
||||
if context["summary_output"] is None or not context["summary_output"].exists():
|
||||
@@ -728,7 +1133,10 @@ def _validate_resume_artifacts(*, record: dict[str, Any], state: RunState, resum
|
||||
if resume_from_stage == DELIVERY_STAGE:
|
||||
expected_candidate_count = int(_stage_output(state, FILTER_STAGE, "candidate_count") or 0)
|
||||
actual_candidate_count = sum(1 for context in item_contexts if context.get("candidate") is not None)
|
||||
if candidate_batch_valid:
|
||||
if expected_candidate_count != actual_candidate_count:
|
||||
missing_artifacts.append(_normalize_repo_path(candidate_batch_output or _candidate_batch_output(run_dir)))
|
||||
elif expected_candidate_count != actual_candidate_count:
|
||||
for context in item_contexts:
|
||||
if context.get("summary_payload") is not None and (context["openclaw_path"] is None or not context["openclaw_path"].exists()):
|
||||
missing_artifacts.append(
|
||||
|
||||
+135
-4
@@ -1,7 +1,7 @@
|
||||
from __future__ import annotations
|
||||
|
||||
# MCP 服务入口:将内容提取、过滤、FreshRSS 全链路管道暴露为 MCP 工具。
|
||||
# 生产主入口是 run_freshrss_openclaw_pipeline,其余工具供单步调试使用。
|
||||
# 主日报正式启动入口是 start_freshrss_pipeline_job;同步入口仅保留给 debug / fallback。
|
||||
|
||||
from datetime import date
|
||||
from pathlib import Path
|
||||
@@ -16,10 +16,23 @@ from summary_mcp.models.item import Item
|
||||
from summary_mcp.models.llm_result import LlmSummaryResult
|
||||
from summary_mcp.models.summary_io import ExtractionInput
|
||||
from summary_mcp.runtime import get_delivery_payload as load_delivery_payload
|
||||
from summary_mcp.runtime import get_freshrss_pipeline_job_result as load_freshrss_pipeline_job_result
|
||||
from summary_mcp.runtime import get_freshrss_pipeline_job_status as load_freshrss_pipeline_job_status
|
||||
from summary_mcp.runtime import get_resume_job_result as load_resume_job_result
|
||||
from summary_mcp.runtime import get_resume_job_status as load_resume_job_status
|
||||
from summary_mcp.runtime import get_run_status as load_run_status
|
||||
from summary_mcp.runtime import get_run_report as load_run_report
|
||||
from summary_mcp.runtime import inspect_resume_plan as load_resume_plan
|
||||
from summary_mcp.runtime import list_run_artifacts as load_run_artifacts
|
||||
from summary_mcp.runtime import list_runs as load_runs
|
||||
from summary_mcp.runtime import start_resume_job as launch_resume_job
|
||||
from summary_mcp.runtime import start_freshrss_pipeline_job as launch_freshrss_pipeline_job
|
||||
from summary_mcp.runtime.article_summary_jobs import (
|
||||
get_article_summary_job_result as load_article_summary_job_result,
|
||||
get_article_summary_job_status as load_article_summary_job_status,
|
||||
start_article_summary_job as launch_article_summary_job,
|
||||
_validate_article_summary_extracted_path,
|
||||
)
|
||||
from summary_mcp.runtime.resume_service import resume_run as resume_existing_run
|
||||
from summary_mcp.workflows import run_freshrss_pipeline
|
||||
from summary_mcp.workflows.article_summary import ArticleSummaryConfig, summarize_selected_articles
|
||||
@@ -125,6 +138,64 @@ def run_freshrss_openclaw_pipeline(
|
||||
return result
|
||||
|
||||
|
||||
@mcp.tool()
|
||||
def start_freshrss_pipeline_job(
|
||||
limit: int = 5,
|
||||
mark_read: bool = False,
|
||||
include_read: bool = False,
|
||||
debug_artifacts: bool = False,
|
||||
continuation: str | None = None,
|
||||
timeout_seconds: float = 60.0,
|
||||
max_retries: int = 2,
|
||||
stream_id: str = "user/-/state/com.google/reading-list",
|
||||
api_base_url: str | None = None,
|
||||
username: str | None = None,
|
||||
api_password: str | None = None,
|
||||
llm_api_key: str | None = None,
|
||||
llm_model: str | None = None,
|
||||
llm_api_url: str | None = None,
|
||||
context: dict | None = None,
|
||||
run_id: str | None = None,
|
||||
date_value: str | None = None,
|
||||
output_dir: str | None = None,
|
||||
include_item_reports: bool = False,
|
||||
) -> dict:
|
||||
"""Start an asynchronous FreshRSS pipeline job and return a job_id immediately."""
|
||||
return launch_freshrss_pipeline_job(
|
||||
limit=limit,
|
||||
mark_read=mark_read,
|
||||
include_read=include_read,
|
||||
debug_artifacts=debug_artifacts,
|
||||
continuation=continuation,
|
||||
timeout_seconds=timeout_seconds,
|
||||
max_retries=max_retries,
|
||||
stream_id=stream_id,
|
||||
api_base_url=api_base_url,
|
||||
username=username,
|
||||
api_password=api_password,
|
||||
llm_api_key=llm_api_key,
|
||||
llm_model=llm_model,
|
||||
llm_api_url=llm_api_url,
|
||||
context=context,
|
||||
run_id=run_id,
|
||||
date_value=date_value,
|
||||
output_dir=Path(output_dir) if output_dir else None,
|
||||
include_item_reports=include_item_reports,
|
||||
)
|
||||
|
||||
|
||||
@mcp.tool()
|
||||
def get_freshrss_pipeline_job_status(job_id: str) -> dict:
|
||||
"""Get the current status of an asynchronous FreshRSS pipeline job."""
|
||||
return load_freshrss_pipeline_job_status(job_id=job_id)
|
||||
|
||||
|
||||
@mcp.tool()
|
||||
def get_freshrss_pipeline_job_result(job_id: str) -> dict:
|
||||
"""Get the final result of an asynchronous FreshRSS pipeline job."""
|
||||
return load_freshrss_pipeline_job_result(job_id=job_id)
|
||||
|
||||
|
||||
@mcp.tool()
|
||||
def get_run_status(run_id: str) -> dict:
|
||||
"""Get the current status of a workflow run by run_id."""
|
||||
@@ -165,6 +236,67 @@ def resume_run(run_id: str) -> dict:
|
||||
return resume_existing_run(run_id=run_id)
|
||||
|
||||
|
||||
@mcp.tool()
|
||||
def inspect_resume_plan(run_id: str) -> dict:
|
||||
"""Inspect the effective resume plan for a FreshRSS workflow run without executing it."""
|
||||
return load_resume_plan(run_id=run_id)
|
||||
|
||||
|
||||
@mcp.tool()
|
||||
def start_resume_job(run_id: str) -> dict:
|
||||
"""Start an asynchronous resume job for a resumable FreshRSS workflow run."""
|
||||
return launch_resume_job(run_id=run_id)
|
||||
|
||||
|
||||
@mcp.tool()
|
||||
def get_resume_job_status(job_id: str) -> dict:
|
||||
"""Get the current status of an asynchronous resume job."""
|
||||
return load_resume_job_status(job_id=job_id)
|
||||
|
||||
|
||||
@mcp.tool()
|
||||
def get_resume_job_result(job_id: str) -> dict:
|
||||
"""Get the final result of an asynchronous resume job."""
|
||||
return load_resume_job_result(job_id=job_id)
|
||||
|
||||
|
||||
@mcp.tool()
|
||||
def start_article_summary_job(
|
||||
*,
|
||||
extracted_path: str,
|
||||
selected_ids: list[str],
|
||||
output_dir: str | None = None,
|
||||
max_retries: int = 2,
|
||||
timeout_seconds: float = 120.0,
|
||||
llm_api_key: str | None = None,
|
||||
llm_model: str | None = None,
|
||||
llm_api_url: str | None = None,
|
||||
) -> dict:
|
||||
"""Start an asynchronous article-summary job and return a job_id immediately."""
|
||||
return launch_article_summary_job(
|
||||
extracted_path=Path(extracted_path),
|
||||
selected_ids=selected_ids,
|
||||
output_dir=Path(output_dir) if output_dir else None,
|
||||
max_retries=max_retries,
|
||||
timeout_seconds=timeout_seconds,
|
||||
llm_api_key=llm_api_key,
|
||||
llm_model=llm_model,
|
||||
llm_api_url=llm_api_url,
|
||||
)
|
||||
|
||||
|
||||
@mcp.tool()
|
||||
def get_article_summary_job_status(job_id: str) -> dict:
|
||||
"""Get the current status of an asynchronous article-summary job."""
|
||||
return load_article_summary_job_status(job_id=job_id)
|
||||
|
||||
|
||||
@mcp.tool()
|
||||
def get_article_summary_job_result(job_id: str) -> dict:
|
||||
"""Get the final result of an asynchronous article-summary job."""
|
||||
return load_article_summary_job_result(job_id=job_id)
|
||||
|
||||
|
||||
@mcp.tool()
|
||||
def generate_article_summaries(
|
||||
*,
|
||||
@@ -172,7 +304,7 @@ def generate_article_summaries(
|
||||
selected_ids: list[str],
|
||||
output_dir: str | None = None,
|
||||
max_retries: int = 2,
|
||||
timeout_seconds: float = 60.0,
|
||||
timeout_seconds: float = 120.0,
|
||||
llm_api_key: str | None = None,
|
||||
llm_model: str | None = None,
|
||||
llm_api_url: str | None = None,
|
||||
@@ -197,8 +329,7 @@ def generate_article_summaries(
|
||||
"""
|
||||
|
||||
extracted_path_obj = Path(extracted_path)
|
||||
if not extracted_path_obj.exists():
|
||||
raise FileNotFoundError(f"extracted_path does not exist: {extracted_path}")
|
||||
_validate_article_summary_extracted_path(extracted_path_obj)
|
||||
|
||||
if output_dir is None:
|
||||
default_dir = extracted_path_obj.parent / "single_summaries"
|
||||
|
||||
@@ -4,6 +4,9 @@ from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
from typing import Iterable, Mapping, Sequence
|
||||
|
||||
import httpx
|
||||
from dotenv import dotenv_values
|
||||
|
||||
from summary_mcp.core.summary_loop import run_loop_payload
|
||||
from summary_mcp.validators.article_summary import validate_article_summary_payload
|
||||
|
||||
@@ -12,13 +15,14 @@ REPO_ROOT = Path(__file__).resolve().parents[3]
|
||||
OUTPUT_ROOT = REPO_ROOT / "outputs"
|
||||
FRESHRSS_OUTPUT_ROOT = OUTPUT_ROOT / "freshrss"
|
||||
DEFAULT_PROMPT_PATH = OUTPUT_ROOT / "prompts" / "article-summary-prompt.txt"
|
||||
DEFAULT_DOTENV_PATH = REPO_ROOT / ".env"
|
||||
|
||||
|
||||
@dataclass
|
||||
class ArticleSummaryConfig:
|
||||
prompt_path: Path = DEFAULT_PROMPT_PATH
|
||||
max_retries: int = 2
|
||||
timeout_seconds: float = 60.0
|
||||
timeout_seconds: float = 120.0
|
||||
llm_api_key: str | None = None
|
||||
llm_model: str | None = None
|
||||
llm_api_url: str | None = None
|
||||
@@ -35,6 +39,7 @@ def _resolve_article_llm_settings(
|
||||
Resolution order for each field:
|
||||
- explicit function argument
|
||||
- ARTICLE_SUMMARY_* environment variable
|
||||
- ARTICLE_SUMMARY_* in repo .env
|
||||
- main LLM_* / OPENAI_* environment variables (handled by summary_loop.resolve_llm_settings)
|
||||
|
||||
This helper intentionally does not validate presence; the summary loop
|
||||
@@ -43,9 +48,17 @@ def _resolve_article_llm_settings(
|
||||
|
||||
import os
|
||||
|
||||
resolved_api_key = api_key or os.environ.get("ARTICLE_SUMMARY_LLM_API_KEY")
|
||||
resolved_model = model or os.environ.get("ARTICLE_SUMMARY_LLM_MODEL")
|
||||
resolved_api_url = api_url or os.environ.get("ARTICLE_SUMMARY_LLM_API_URL")
|
||||
dotenv_map = {}
|
||||
if DEFAULT_DOTENV_PATH.exists():
|
||||
dotenv_map = {
|
||||
key: value
|
||||
for key, value in dotenv_values(DEFAULT_DOTENV_PATH).items()
|
||||
if isinstance(key, str) and isinstance(value, str) and value
|
||||
}
|
||||
|
||||
resolved_api_key = api_key or os.environ.get("ARTICLE_SUMMARY_LLM_API_KEY") or dotenv_map.get("ARTICLE_SUMMARY_LLM_API_KEY")
|
||||
resolved_model = model or os.environ.get("ARTICLE_SUMMARY_LLM_MODEL") or dotenv_map.get("ARTICLE_SUMMARY_LLM_MODEL")
|
||||
resolved_api_url = api_url or os.environ.get("ARTICLE_SUMMARY_LLM_API_URL") or dotenv_map.get("ARTICLE_SUMMARY_LLM_API_URL")
|
||||
return resolved_api_key, resolved_model, resolved_api_url
|
||||
|
||||
|
||||
@@ -97,6 +110,32 @@ def _iter_selected_items(
|
||||
}
|
||||
return
|
||||
|
||||
# Fallback: pipeline delivery payload with top-level "candidates" array.
|
||||
# Each entry has item_id (the raw FreshRSS item_id) and candidate_id (with cand: prefix).
|
||||
# Normalize selected_ids by stripping "cand:" prefix so they match item_id.
|
||||
candidates = extracted_payload.get("candidates")
|
||||
if isinstance(candidates, list):
|
||||
norm_selected = {
|
||||
sid.removeprefix("cand:") if sid.startswith("cand:") else sid
|
||||
for sid in selected_ids
|
||||
}
|
||||
for entry in candidates:
|
||||
if not isinstance(entry, Mapping):
|
||||
continue
|
||||
# Support both item_id field (new) and candidate_id (legacy fallback)
|
||||
raw_item_id = entry.get("item_id") or ""
|
||||
if not raw_item_id and entry.get("candidate_id"):
|
||||
raw_item_id = entry["candidate_id"].removeprefix("cand:")
|
||||
item_id = str(raw_item_id) if raw_item_id else None
|
||||
if not item_id or item_id not in norm_selected:
|
||||
continue
|
||||
# Candidates store article fields directly, not nested under "article"
|
||||
yield item_id, {
|
||||
"item": {},
|
||||
"extraction": {"article": entry, "warnings": []},
|
||||
}
|
||||
return
|
||||
|
||||
# Fallback: legacy payload with top-level "items" array.
|
||||
items = extracted_payload.get("items")
|
||||
if isinstance(items, list):
|
||||
@@ -107,11 +146,10 @@ def _iter_selected_items(
|
||||
item_id = str(raw_item_id) if raw_item_id is not None else None
|
||||
if not item_id or item_id not in selected_set:
|
||||
continue
|
||||
|
||||
yield item_id, item
|
||||
return
|
||||
|
||||
# Format 3: single-item extracted file produced by run_freshrss_pipeline debug mode.
|
||||
# Format 3: single-item extracted file produced by run_freshrss_pipeline.
|
||||
# Shape: {"success": bool, "article": {"item_id": "...", ...}, "warnings": [...]}
|
||||
article = extracted_payload.get("article")
|
||||
if isinstance(article, Mapping):
|
||||
@@ -165,10 +203,12 @@ def summarize_selected_articles(
|
||||
|
||||
output_dir.mkdir(parents=True, exist_ok=True)
|
||||
written_paths: list[Path] = []
|
||||
matched_ids: set[str] = set()
|
||||
|
||||
from summary_mcp.core.summary_loop import build_summary_input
|
||||
|
||||
for item_id, entry in _iter_selected_items(payload, selected_ids):
|
||||
matched_ids.add(item_id)
|
||||
# Prefer the real project format where each entry has ``item`` and
|
||||
# ``extraction.article``; fall back to legacy layout where the
|
||||
# article fields live directly on the element.
|
||||
@@ -187,6 +227,10 @@ def summarize_selected_articles(
|
||||
"warnings": warnings,
|
||||
}
|
||||
|
||||
summary_exit_code = 1
|
||||
summary_payload = None
|
||||
|
||||
try:
|
||||
summary_exit_code, summary_payload, _ = run_loop_payload(
|
||||
extracted_payload=extracted_payload,
|
||||
prompt_path=cfg.prompt_path,
|
||||
@@ -198,6 +242,30 @@ def summarize_selected_articles(
|
||||
output_path=None,
|
||||
validator=validate_article_summary_payload,
|
||||
)
|
||||
except httpx.TimeoutException:
|
||||
summary_exit_code = 1
|
||||
summary_payload = None
|
||||
|
||||
should_fallback = (
|
||||
(summary_exit_code != 0 or summary_payload is None)
|
||||
and resolved_model is not None
|
||||
and resolved_model == (cfg.llm_model or model or resolved_model)
|
||||
and resolved_model.startswith("deepseek")
|
||||
)
|
||||
|
||||
if should_fallback:
|
||||
summary_exit_code, summary_payload, _ = run_loop_payload(
|
||||
extracted_payload=extracted_payload,
|
||||
prompt_path=cfg.prompt_path,
|
||||
max_retries=cfg.max_retries,
|
||||
timeout_seconds=cfg.timeout_seconds,
|
||||
api_key=None,
|
||||
model=None,
|
||||
api_url=None,
|
||||
output_path=None,
|
||||
validator=validate_article_summary_payload,
|
||||
)
|
||||
|
||||
if summary_exit_code != 0 or summary_payload is None:
|
||||
continue
|
||||
|
||||
@@ -250,12 +318,19 @@ def summarize_selected_articles(
|
||||
lines.append("、".join(topics))
|
||||
lines.append("")
|
||||
|
||||
# Normalize: collapse spaces/slashes/underscores to single dash, strip punctuation, collapse multi-dashes
|
||||
normalized = str(title).lower()
|
||||
for sep in (" ", "/", "_", "——", "―", "‐"):
|
||||
normalized = normalized.replace(sep, "-")
|
||||
safe_title = "-".join(
|
||||
str(title).lower().strip().replace(" ", "-").split()
|
||||
part for part in normalized.split("-") if part
|
||||
)[:80]
|
||||
filename = f"{safe_title or item_id}.md"
|
||||
output_path = output_dir / filename
|
||||
output_path.write_text("\n".join(lines), encoding="utf-8")
|
||||
written_paths.append(output_path)
|
||||
|
||||
if not matched_ids:
|
||||
raise ValueError(f"No extracted entries matched selected_ids: {list(selected_ids)}")
|
||||
|
||||
return written_paths
|
||||
|
||||
@@ -2,10 +2,13 @@ from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||
from datetime import date, datetime, timezone
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
from dotenv import dotenv_values
|
||||
|
||||
from summary_mcp.core.keyword_index import persist_keyword_indexes
|
||||
from summary_mcp.core.pipeline import extract_content
|
||||
from summary_mcp.core.summary_loop import resolve_llm_settings, run_loop_payload
|
||||
@@ -40,6 +43,7 @@ DEFAULT_TERM_ALIASES_PATH = REPO_ROOT / "configs" / "term_aliases.json"
|
||||
DEFAULT_TERM_STOPWORDS_PATH = REPO_ROOT / "configs" / "term_stopwords.json"
|
||||
DEFAULT_TERM_DAILY_DIR = DATA_ROOT / "daily"
|
||||
DEFAULT_TERM_STATS_PATH = DATA_ROOT / "term_stats.json"
|
||||
DEFAULT_DOTENV_PATH = REPO_ROOT / ".env"
|
||||
WORKFLOW_NAME = "freshrss_daily_digest"
|
||||
RUN_TYPE = "daily_digest"
|
||||
FETCH_STAGE = "fetch_feed"
|
||||
@@ -48,6 +52,10 @@ SUMMARY_STAGE = "generate_summaries"
|
||||
FILTER_STAGE = "apply_filters"
|
||||
DELIVERY_STAGE = "build_delivery_payload"
|
||||
REPORT_STAGE = "write_run_report"
|
||||
SUMMARY_BATCH_ARTIFACT = "summary_batch"
|
||||
CANDIDATE_BATCH_ARTIFACT = "candidate_batch"
|
||||
SUMMARY_BATCH_FILENAME = "summary-batch.json"
|
||||
CANDIDATE_BATCH_FILENAME = "candidate-batch.json"
|
||||
|
||||
|
||||
def _save_json(path: Path, payload: dict[str, Any] | list[Any]) -> None:
|
||||
@@ -59,14 +67,27 @@ def _load_json(path: Path) -> dict[str, Any]:
|
||||
return json.loads(path.read_text(encoding="utf-8-sig"))
|
||||
|
||||
|
||||
def _load_required_env(name: str, value: str | None) -> str:
|
||||
def _load_repo_dotenv() -> dict[str, str]:
|
||||
if not DEFAULT_DOTENV_PATH.exists():
|
||||
return {}
|
||||
return {
|
||||
key: value
|
||||
for key, value in dotenv_values(DEFAULT_DOTENV_PATH).items()
|
||||
if isinstance(key, str) and isinstance(value, str) and value
|
||||
}
|
||||
|
||||
|
||||
def _load_required_env(name: str, value: str | None, dotenv_map: dict[str, str] | None = None) -> str:
|
||||
if value:
|
||||
return value
|
||||
env_value = os.environ.get(name)
|
||||
if env_value:
|
||||
return env_value
|
||||
dotenv_value = (dotenv_map or {}).get(name)
|
||||
if dotenv_value:
|
||||
return dotenv_value
|
||||
raise RuntimeError(
|
||||
f"Missing required value '{name}': not passed as argument and not set as environment variable."
|
||||
f"Missing required value '{name}': not passed as argument, not set as environment variable, and not found in {DEFAULT_DOTENV_PATH}."
|
||||
)
|
||||
|
||||
|
||||
@@ -121,6 +142,77 @@ def _build_item_context(*, index: int, item: Any, resolved_output_dir: Path, deb
|
||||
}
|
||||
|
||||
|
||||
def _summary_batch_output(run_dir: Path) -> Path:
|
||||
return run_dir / "summary" / SUMMARY_BATCH_FILENAME
|
||||
|
||||
|
||||
def _candidate_batch_output(run_dir: Path) -> Path:
|
||||
return run_dir / "candidates" / CANDIDATE_BATCH_FILENAME
|
||||
|
||||
|
||||
def _build_summary_batch_payload(*, run_id: str, item_contexts: list[dict[str, Any]]) -> dict[str, Any]:
|
||||
items: list[dict[str, Any]] = []
|
||||
for item_context in item_contexts:
|
||||
summary_payload = item_context.get("summary_payload")
|
||||
if summary_payload is None:
|
||||
continue
|
||||
item = item_context["item"]
|
||||
items.append(
|
||||
{
|
||||
"item_key": item_context["item_key"],
|
||||
"item_id": item.item_id,
|
||||
"summary": summary_payload,
|
||||
}
|
||||
)
|
||||
return {
|
||||
"run_id": run_id,
|
||||
"summary_count": len(items),
|
||||
"items": items,
|
||||
}
|
||||
|
||||
|
||||
def _build_candidate_batch_payload(*, run_id: str, item_contexts: list[dict[str, Any]]) -> dict[str, Any]:
|
||||
items: list[dict[str, Any]] = []
|
||||
for item_context in item_contexts:
|
||||
candidate = item_context.get("candidate")
|
||||
if candidate is None:
|
||||
continue
|
||||
item = item_context["item"]
|
||||
items.append(
|
||||
{
|
||||
"item_key": item_context["item_key"],
|
||||
"item_id": item.item_id,
|
||||
"candidate_id": candidate.candidate_id,
|
||||
"candidate": candidate.model_dump(mode="json"),
|
||||
}
|
||||
)
|
||||
return {
|
||||
"run_id": run_id,
|
||||
"candidate_count": len(items),
|
||||
"items": items,
|
||||
}
|
||||
|
||||
|
||||
def _persist_summary_batch_artifact(*, run_store, run_dir: Path, item_contexts: list[dict[str, Any]]) -> Path:
|
||||
output_path = _summary_batch_output(run_dir)
|
||||
_save_json(
|
||||
output_path,
|
||||
_build_summary_batch_payload(run_id=run_store.state.run_id, item_contexts=item_contexts),
|
||||
)
|
||||
run_store.register_artifact(name=SUMMARY_BATCH_ARTIFACT, path=output_path, kind="json", stage=SUMMARY_STAGE)
|
||||
return output_path
|
||||
|
||||
|
||||
def _persist_candidate_batch_artifact(*, run_store, run_dir: Path, item_contexts: list[dict[str, Any]]) -> Path:
|
||||
output_path = _candidate_batch_output(run_dir)
|
||||
_save_json(
|
||||
output_path,
|
||||
_build_candidate_batch_payload(run_id=run_store.state.run_id, item_contexts=item_contexts),
|
||||
)
|
||||
run_store.register_artifact(name=CANDIDATE_BATCH_ARTIFACT, path=output_path, kind="json", stage=FILTER_STAGE)
|
||||
return output_path
|
||||
|
||||
|
||||
def _build_run_report(
|
||||
*,
|
||||
resolved_run_id: str,
|
||||
@@ -243,9 +335,10 @@ def run_freshrss_pipeline(
|
||||
|
||||
try:
|
||||
run_store.start_stage(FETCH_STAGE, outputs={"output_dir": str(resolved_output_dir)})
|
||||
resolved_api_base_url = _load_required_env("FRESHRSS_API_BASE_URL", api_base_url)
|
||||
resolved_username = _load_required_env("FRESHRSS_USERNAME", username)
|
||||
resolved_api_password = _load_required_env("FRESHRSS_API_PASSWORD", api_password)
|
||||
dotenv_map = _load_repo_dotenv()
|
||||
resolved_api_base_url = _load_required_env("FRESHRSS_API_BASE_URL", api_base_url, dotenv_map)
|
||||
resolved_username = _load_required_env("FRESHRSS_USERNAME", username, dotenv_map)
|
||||
resolved_api_password = _load_required_env("FRESHRSS_API_PASSWORD", api_password, dotenv_map)
|
||||
resolved_llm_api_key, resolved_llm_model, resolved_llm_api_url = resolve_llm_settings(
|
||||
api_key=llm_api_key,
|
||||
model=llm_model,
|
||||
@@ -369,9 +462,12 @@ def run_freshrss_pipeline(
|
||||
summary_success_count = 0
|
||||
summary_failed_count = 0
|
||||
summary_candidates = [ctx for ctx in item_contexts if ctx["extraction"] is not None and ctx["extraction"].success]
|
||||
# Parallelize LLM summaries — I/O bound calls, independent per article
|
||||
with ThreadPoolExecutor(max_workers=min(len(summary_candidates) or 1, 4)) as pool:
|
||||
fut_map = {}
|
||||
for item_context in summary_candidates:
|
||||
item_report = item_context["item_report"]
|
||||
summary_exit_code, summary_payload, summary_report = run_loop_payload(
|
||||
fut = pool.submit(
|
||||
run_loop_payload,
|
||||
extracted_payload=item_context["extracted_payload"],
|
||||
prompt_path=resolved_prompt_path,
|
||||
output_path=item_context["summary_output"],
|
||||
@@ -381,6 +477,16 @@ def run_freshrss_pipeline(
|
||||
model=resolved_llm_model,
|
||||
api_url=resolved_llm_api_url,
|
||||
)
|
||||
fut_map[fut] = item_context
|
||||
|
||||
for fut in as_completed(fut_map):
|
||||
item_context = fut_map[fut]
|
||||
item_report = item_context["item_report"]
|
||||
try:
|
||||
summary_exit_code, summary_payload, summary_report = fut.result()
|
||||
except Exception as exc:
|
||||
summary_exit_code, summary_payload, summary_report = 1, None, None
|
||||
|
||||
if summary_exit_code != 0 or summary_payload is None:
|
||||
item_report["status"] = "summary_failed"
|
||||
if summary_report is not None:
|
||||
@@ -401,6 +507,11 @@ def run_freshrss_pipeline(
|
||||
},
|
||||
)
|
||||
|
||||
summary_batch_output = _persist_summary_batch_artifact(
|
||||
run_store=run_store,
|
||||
run_dir=resolved_output_dir,
|
||||
item_contexts=item_contexts,
|
||||
)
|
||||
if debug_artifacts and (resolved_output_dir / "summary").exists():
|
||||
run_store.register_artifact(name="summary_dir", path=resolved_output_dir / "summary", kind="directory", stage=SUMMARY_STAGE)
|
||||
run_store.finish_stage(
|
||||
@@ -410,6 +521,7 @@ def run_freshrss_pipeline(
|
||||
"completed_items": summary_success_count + summary_failed_count,
|
||||
"success_count": summary_success_count,
|
||||
"failed_count": summary_failed_count,
|
||||
"summary_batch_output": str(summary_batch_output),
|
||||
},
|
||||
)
|
||||
|
||||
@@ -493,6 +605,11 @@ def run_freshrss_pipeline(
|
||||
},
|
||||
)
|
||||
|
||||
candidate_batch_output = _persist_candidate_batch_artifact(
|
||||
run_store=run_store,
|
||||
run_dir=resolved_output_dir,
|
||||
item_contexts=item_contexts,
|
||||
)
|
||||
if debug_artifacts and (resolved_output_dir / "candidates").exists():
|
||||
run_store.register_artifact(name="candidate_dir", path=resolved_output_dir / "candidates", kind="directory", stage=FILTER_STAGE)
|
||||
run_store.finish_stage(
|
||||
@@ -504,6 +621,7 @@ def run_freshrss_pipeline(
|
||||
"keep_count": keep_count,
|
||||
"review_count": review_count,
|
||||
"drop_count": drop_count,
|
||||
"candidate_batch_output": str(candidate_batch_output),
|
||||
},
|
||||
)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user