Compare commits
7
Commits
4399c9ca90
..
main
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
8e27cca166 | ||
|
|
807027976e | ||
|
|
6dd8cef347 | ||
|
|
5eb390e3ed | ||
|
|
7b791ac947 | ||
|
|
cdbcdcd485 | ||
|
|
590d050218 |
@@ -1,390 +1,214 @@
|
||||
# Reader MCP Workflow Service
|
||||
# Reader · AI 日报引擎
|
||||
|
||||
reader 当前已经收口为面向 OpenClaw 的 MCP workflow service。正式能力边界以 FreshRSS 日报工作流为准:启动 run、写入 `run-state.json`、查询运行状态、读取结构化结果,以及异步恢复 job。CLI 与同步入口仍保留,但定位为 debug / fallback,而不是正式集成入口。
|
||||
> 从 FreshRSS 到 AI 日报的自动化流水线,为个人知识管理生成每日 AI 工程化简报。
|
||||
|
||||
## 运行
|
||||
Reader 是一个端到端的 AI 日报生产系统,定时从自建 FreshRSS 的 RSS 订阅源拉取文章,经过内容提取、LLM 筛选与摘要、关键词索引构建,最终产出两个输出:
|
||||
|
||||
1. **公开日报** — 推送到 [Hugo 站点](https://osiman.site/daily/) 的精选技术简报
|
||||
2. **知识沉淀** — 单篇结构化摘要上传到 IMA 知识库(`daily` KB)
|
||||
|
||||
整个流程由 OpenClaw 编排,作为 MCP Workflow Service 对外暴露。
|
||||
|
||||
---
|
||||
|
||||
## ✨ 核心能力
|
||||
|
||||
| 能力 | 说明 |
|
||||
|:----|:------|
|
||||
| **RSS 拉取** | 从 FreshRSS API 拉取订阅文章,支持增量读取与已读标记 |
|
||||
| **内容提取** | 自动提取文章正文、标题、来源等结构化字段 |
|
||||
| **LLM 筛选** | 基于个人兴趣画像(`filter_context.personal.json`)自动评估文章质量,分为 keep / review / drop 三档 |
|
||||
| **LLM 摘要** | 并行生成每篇文章的结构化摘要(4 路并发,约 24 秒完成 7 篇) |
|
||||
| **关键词索引** | 自动构建每日关键词索引,支持别名映射与停用词过滤 |
|
||||
| **候选简报** | 生成 `digest-brief.json` 供编排层(OpenClaw)决策 |
|
||||
| **单篇沉淀** | 对选中的文章生成结构化知识笔记,上传到 IMA 知识库 |
|
||||
| **异步 Job** | 全部生产流程走异步 job,支持恢复与状态查询 |
|
||||
|
||||
---
|
||||
|
||||
## 🏗 架构概览
|
||||
|
||||
```
|
||||
FreshRSS ──→ 拉取 ──→ 内容提取 ──→ LLM 筛选 ──→ 关键词索引
|
||||
│
|
||||
digest-brief.json
|
||||
│
|
||||
┌──────────────┼──────────────┐
|
||||
▼ ▼ ▼
|
||||
Hugo 日报 IMA 知识库 term_index
|
||||
(公开简报) (单篇沉淀) (关键词数据)
|
||||
```
|
||||
|
||||
### MCP 工具层
|
||||
|
||||
Reader 通过 Hermes MCP 暴露 20+ 个工具,分为三类:
|
||||
|
||||
**日报流水线:**
|
||||
- `start_freshrss_pipeline_job` → 启动异步日报 Job
|
||||
- `get_freshrss_pipeline_job_status` / `get_freshrss_pipeline_job_result` → 轮询结果
|
||||
|
||||
**状态查询:**
|
||||
- `get_run_status` / `get_delivery_payload` / `get_run_report` → 读取运行结果
|
||||
- `list_runs` / `list_run_artifacts` → 浏览运行历史
|
||||
|
||||
**恢复与单篇总结:**
|
||||
- `inspect_resume_plan` / `start_resume_job` → 恢复失败 Job
|
||||
- `start_article_summary_job` / `generate_article_summaries` → 单篇文章摘要
|
||||
|
||||
### CLI 入口
|
||||
|
||||
同步入口,适合本地 debug / fallback:
|
||||
|
||||
```bash
|
||||
# 完整日报流水线
|
||||
python scripts/run_freshrss_pipeline.py --limit 7 --mark-read --timeout 300
|
||||
|
||||
# 单篇文章摘要
|
||||
python scripts/run_article_summaries.py \
|
||||
--extracted outputs/freshrss/rerun/<run_id>/extracted/item-01.extracted.json \
|
||||
--output-dir outputs/freshrss/single_summaries/YYYY-MM-DD
|
||||
|
||||
# 关键词维护
|
||||
python scripts/build_keyword_index.py
|
||||
python scripts/generate_term_cleanup_suggestions.py
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🚀 快速开始
|
||||
|
||||
### 环境变量
|
||||
|
||||
```
|
||||
# FreshRSS
|
||||
FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||||
FRESHRSS_USERNAME=bot
|
||||
FRESHRSS_API_PASSWORD=xxx
|
||||
|
||||
# LLM(主流水线)
|
||||
LLM_API_URL=https://api.deepseek.com
|
||||
LLM_API_KEY=xxx
|
||||
LLM_MODEL=deepseek-chat
|
||||
|
||||
# LLM(可选,单篇摘要独立模型)
|
||||
ARTICLE_SUMMARY_LLM_API_URL=https://api.deepseek.com
|
||||
ARTICLE_SUMMARY_LLM_API_KEY=xxx
|
||||
ARTICLE_SUMMARY_LLM_MODEL=deepseek-chat
|
||||
|
||||
# IMA 知识库(可选,仅沉淀时需要)
|
||||
IMA_DAILY_KNOWLEDGE_BASE_ID=xxx
|
||||
IMA_DAILY_KNOWLEDGE_BASE_NAME=daily
|
||||
```
|
||||
|
||||
### 运行
|
||||
|
||||
```bash
|
||||
# 安装
|
||||
pip install -e .
|
||||
|
||||
# 跑日报流水线(CLI 模式)
|
||||
python scripts/run_freshrss_pipeline.py --limit 7 --mark-read --timeout 300
|
||||
|
||||
# 启动 MCP 服务(OpenClaw 集成用)
|
||||
summary-mcp
|
||||
```
|
||||
|
||||
服务当前暴露 21 个工具。
|
||||
---
|
||||
|
||||
正式集成摘要:
|
||||
## 📁 项目结构
|
||||
|
||||
- 主日报正式入口:`start_freshrss_pipeline_job`
|
||||
- 主日报正式读取:`get_run_status`、`get_delivery_payload`、`get_run_report`
|
||||
- 恢复正式入口:`inspect_resume_plan`、`start_resume_job`、`get_resume_job_status`、`get_resume_job_result`
|
||||
- 单篇总结正式入口:`start_article_summary_job`、`get_article_summary_job_status`、`get_article_summary_job_result`
|
||||
- `run_freshrss_openclaw_pipeline`、`resume_run`、`generate_article_summaries` 仅用于同步 debug / fallback
|
||||
|
||||
## 文档入口
|
||||
|
||||
如果你在做 OpenClaw 集成,不要只看这个 README,优先看:
|
||||
|
||||
- `docs/openclaw/README.md`
|
||||
- `docs/openclaw/openclaw-handoff.md`
|
||||
- `docs/openclaw/openclaw-orchestration-flow.md`
|
||||
|
||||
字段契约见:
|
||||
|
||||
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||
- `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||
|
||||
文档总索引见:
|
||||
|
||||
- `docs/README.md`
|
||||
- `docs/current/context-reset-brief.md`
|
||||
- `docs/design/README.md`
|
||||
- `plans/README.md`
|
||||
|
||||
## 正式能力边界
|
||||
|
||||
- 当前正式 workflow 只有 `freshrss_daily_digest`
|
||||
- 当前生产编排默认走异步 job,而不是同步 MCP / CLI
|
||||
- 每次 FreshRSS 主流水线 run 都会在 `outputs/freshrss/rerun/<run_dir>/run-state.json` 落地运行真相
|
||||
- OpenClaw 正式读取结果应优先使用 MCP 返回的 `run_id`、`output_dir`、`delivery_output`、`report_output`
|
||||
- 正式恢复只支持带有效 `run-state.json` 的当前 run,不处理历史推断 run
|
||||
- 正式生产恢复依赖 `summary/summary-batch.json` 与 `candidates/candidate-batch.json`
|
||||
|
||||
## OpenClaw 最短调用路径
|
||||
|
||||
1. 调 `start_freshrss_pipeline_job`
|
||||
2. 轮询 `get_freshrss_pipeline_job_status`
|
||||
3. 成功后读 `get_freshrss_pipeline_job_result`,拿 `run_id`
|
||||
4. 用 `get_run_status`、`get_delivery_payload`、`get_run_report` 做后续读取
|
||||
5. 如需恢复,先调 `inspect_resume_plan`,只有 `recommended_action=resume` 才走 `start_resume_job`
|
||||
|
||||
更完整的状态分支、恢复策略和人工介入条件见 `docs/openclaw/openclaw-orchestration-flow.md`。
|
||||
|
||||
## 单篇文章总结后处理(可选使用独立 LLM)
|
||||
|
||||
### 生产环境推荐输入
|
||||
|
||||
- 单篇总结的正式生产输入,优先使用 FreshRSS 主流水线输出的**单篇 extracted 文件**:
|
||||
- `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json`
|
||||
- 这些**逐条 extracted 文件**是下游单篇总结的**正式默认产物**。
|
||||
- 像 `outputs/freshrss/extracted/freshrss.extracted.json` 这样的**批量 extracted 文件**,只作为临时场景、兼容旧流程的输入形态保留,**不是首选生产默认**。
|
||||
|
||||
### daily 知识库默认配置
|
||||
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_ID` —— 单篇日报总结默认上传的 IMA 知识库 ID
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_NAME` —— 默认知识库名称(预期值:`daily`)
|
||||
- 上传逻辑在运行时应先校验目标知识库;若配置的目标不存在,应先按名称查找,仍不存在则创建 `daily`
|
||||
|
||||
相关能力:
|
||||
|
||||
- 正式路径:`start_article_summary_job` / `get_article_summary_job_status` / `get_article_summary_job_result`
|
||||
- 同步 debug:`generate_article_summaries`
|
||||
- CLI:`scripts/run_article_summaries.py`
|
||||
- 后台 runner:`scripts/run_article_summary_job.py`
|
||||
|
||||
## 校验 LLM 摘要结果
|
||||
|
||||
```bash
|
||||
validate-llm-result outputs/reference/summary/result.json --extracted outputs/reference/extracted/read-flow-2026.extracted.json
|
||||
```
|
||||
reader/
|
||||
├── configs/ # 配置
|
||||
│ ├── filter_context.personal.json # 个人兴趣画像
|
||||
│ ├── term_aliases.json # 关键词别名映射(149 条)
|
||||
│ ├── term_stopwords.json # 关键词停用词(132 条)
|
||||
│ └── term_cleanup_policy.json # 关键词清理策略
|
||||
├── src/
|
||||
│ └── summary_mcp/ # MCP 服务核心
|
||||
│ ├── server.py # MCP 服务入口
|
||||
│ ├── runtime/ # 运行时(Job 管理、状态持久化)
|
||||
│ └── workflows/ # 工作流(日报流水线逻辑)
|
||||
├── scripts/ # CLI 入口
|
||||
├── outputs/ # 运行时产出
|
||||
│ └── freshrss/
|
||||
│ ├── rerun/<run_id>/ # 每次运行的全量产物
|
||||
│ │ ├── candidates/ # digest-brief.json, delivery payload
|
||||
│ │ ├── extracted/ # item-XX.extracted.json
|
||||
│ │ └── run-state.json # 运行状态
|
||||
│ └── single_summaries/ # 单篇摘要输出
|
||||
├── data/
|
||||
│ └── term_index/ # 关键词索引数据
|
||||
│ ├── daily/YYYY-MM-DD.json
|
||||
│ └── term_stats.json
|
||||
├── docs/ # 设计文档
|
||||
└── prompts/ # LLM Prompt 模板
|
||||
```
|
||||
|
||||
## 跑最小 extraction → summary 循环
|
||||
---
|
||||
|
||||
## ⚙️ 关键技术决策
|
||||
|
||||
| 决策 | 选择 | 原因 |
|
||||
|:----|:----|:------|
|
||||
| 运行模式 | **异步 Job** 为主,CLI fallback | 避免 MCP 传输层 120s 超时限制 |
|
||||
| 摘要并发 | **ThreadPoolExecutor(max_workers=4)** | LLM 调用是 I/O 密集型,4 路并行将 7 篇摘要从 2-3 分钟压到 ~24 秒 |
|
||||
| 环境变量 | **子进程显式注入 .env** | 解决 MCP 服务器环境隔离导致子进程读取不到 LLM_API_KEY 的问题 |
|
||||
| 关键词过滤 | **别名映射 + 停用词 + 语义清洗** | 先用 `term_aliases.json` 归一化,再用 `term_stopwords.json` 过滤噪声,最后通过 LLM 做语义级清洗 |
|
||||
| Tag 选择 | **复用已有通用 Tag**,不从 term_index 翻生僻词 | 保持 Hugo 站点 /tags/ 页面整洁,避免大量一次性专有名词 |
|
||||
|
||||
---
|
||||
|
||||
## 🔧 关键词治理
|
||||
|
||||
配置治理走四步流程(`scripts/` 下脚本):
|
||||
|
||||
```bash
|
||||
python scripts/run_summary_loop.py ^
|
||||
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||
--prompt outputs/prompts/llm-summary-prompt.txt ^
|
||||
--output outputs/reference/summary/result.loop.json
|
||||
# 1. 构建评审数据包
|
||||
python skills/keyword-cleanup-review/scripts/build_review_bundle.py --days 7 --top 50
|
||||
|
||||
# 2. 统计规则级建议(大小写、单复数、频次阈值)
|
||||
python scripts/generate_term_cleanup_suggestions.py
|
||||
|
||||
# 3. LLM 语义级建议(中英映射、简称-全称、近义词)
|
||||
python scripts/generate_term_cleanup_semantic_suggestions.py
|
||||
|
||||
# 4. 确认后写入配置
|
||||
python scripts/apply_term_suggestions.py --accept-watch ... --dry-run
|
||||
```
|
||||
|
||||
## 拉取 FreshRSS 条目并映射为标准化 `item`
|
||||
详见 `docs/design/keyword-engine-maintenance.md`。
|
||||
|
||||
```bash
|
||||
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||||
set FRESHRSS_USERNAME=bot
|
||||
set FRESHRSS_API_PASSWORD=your-api-password
|
||||
python scripts/pull_freshrss_items.py --limit 5 --mark-read
|
||||
```
|
||||
---
|
||||
|
||||
默认会排除已经带 `read` 标签的条目。
|
||||
如果你想拿到完整阅读列表,可以加 `--include-read`。
|
||||
启用 `--mark-read` 后,脚本会在执行成功后把本次抓到的条目标记为已读。
|
||||
## 🤖 Agent Skill
|
||||
|
||||
脚本会写出:
|
||||
Reader 附带一个完整的 OpenClaw Agent Skill,位于 `skills/reader-digest-flow/`,供 AI Agent(Hermes / Claude Code 等)编排每日日报流程使用。
|
||||
|
||||
- `outputs/freshrss/raw/freshrss.raw.json`
|
||||
- `outputs/freshrss/items/freshrss.items.json`
|
||||
Skill 包含完整的 7 阶段工作流定义:
|
||||
1. **Phase 1** — 跑 Pipeline(FreshRSS → 提取 → LLM 筛选)
|
||||
2. **Phase 2** — 汇报候选(展示候选文章给用户决策)
|
||||
3. **Phase 3** — 用户选文(选择 Hugo 发布文章)
|
||||
4. **Phase 4** — 生成并发布 Hugo 日报
|
||||
5. **Phase 5** — 用户选 IMA 沉淀文章
|
||||
6. **Phase 6** — LLM 摘要生成
|
||||
7. **Phase 7** — IMA 知识库上传
|
||||
|
||||
## 拉取 FreshRSS 条目并逐条做内容提取
|
||||
以及海量铁律(不重跑 pipeline、编号规则、Tag 选择规范、IMA 上传流程等)和参考文件(`references/` 目录)。
|
||||
|
||||
```bash
|
||||
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||||
set FRESHRSS_USERNAME=osiman
|
||||
set FRESHRSS_API_PASSWORD=your-api-password
|
||||
python scripts/run_freshrss_extract.py --limit 1 --mark-read
|
||||
```
|
||||
---
|
||||
|
||||
默认会排除已经带 `read` 标签的条目。
|
||||
启用 `--mark-read` 后,只有提取成功的条目才会被标记为已读。
|
||||
## 📄 文档
|
||||
|
||||
脚本会写出:
|
||||
- `docs/openclaw/README.md` — OpenClaw 集成指南
|
||||
- `docs/openclaw/openclaw-orchestration-flow.md` — 编排流程
|
||||
- `docs/openclaw/openclaw-delivery-payload-spec.md` — 字段契约
|
||||
- `docs/design/README.md` — 设计文档总索引
|
||||
- `docs/design/filter-rule-engine-design.md` — 过滤规则引擎设计
|
||||
- `docs/design/filter-rule-engine-usage.md` — 过滤规则使用说明
|
||||
|
||||
- `outputs/freshrss/raw/freshrss.raw.json`
|
||||
- `outputs/freshrss/items/freshrss.items.json`
|
||||
- `outputs/freshrss/extracted/freshrss.extracted.json`
|
||||
---
|
||||
|
||||
注意:这个**批量 extracted 文件**主要用于独立提取场景和旧流程兼容。下游单篇总结的正式生产默认输入,仍然是 `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json` 这类**逐条 extracted 文件**。
|
||||
## 📝 License
|
||||
|
||||
## 跑完整 FreshRSS 流水线,并在最终 delivery payload 写盘成功后再标记已读
|
||||
|
||||
```bash
|
||||
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||||
set FRESHRSS_USERNAME=osiman
|
||||
set FRESHRSS_API_PASSWORD=your-api-password
|
||||
set LLM_API_URL=https://api.deepseek.com
|
||||
set LLM_API_KEY=your-llm-api-key
|
||||
set LLM_MODEL=deepseek-chat
|
||||
python scripts/run_freshrss_pipeline.py --limit 5 --mark-read
|
||||
```
|
||||
|
||||
如果你希望过滤时引入个人工程兴趣 / AI Agent 兴趣画像,可以传入 context 文件:
|
||||
|
||||
```bash
|
||||
python scripts/run_freshrss_pipeline.py ^
|
||||
--limit 5 ^
|
||||
--context configs/filter_context.personal.json ^
|
||||
--mark-read
|
||||
```
|
||||
|
||||
这条 CLI 与 MCP `run_freshrss_openclaw_pipeline` / `start_freshrss_pipeline_job` 共用同一条主流水线逻辑,但正式生产集成应优先走 async MCP job;CLI 与同步 MCP 入口仅用于本地 debug / fallback。默认会写出这些产物:
|
||||
|
||||
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/raw/freshrss.raw.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/candidates/digest-brief.json`(给 OpenClaw 生成 public digest 用的轻量输入,仅包含 `keep` 候选)
|
||||
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json`(每篇一份)
|
||||
|
||||
同时还会更新每日关键词索引运行数据:
|
||||
|
||||
- `data/term_index/daily/YYYY-MM-DD.json`
|
||||
- `data/term_index/term_stats.json`
|
||||
|
||||
主流水线默认**不会**产出批量级的 `freshrss.extracted.json`。
|
||||
如果你需要更多逐条中间产物,例如标准化 items、摘要结果、过滤决策、candidate record、candidate input,可以加 `--debug-artifacts`。
|
||||
|
||||
当 OpenClaw 接入这个 MCP 服务后,应先通过 `start_freshrss_pipeline_job` 启动任务,轮询 `get_freshrss_pipeline_job_status`,再从 `get_freshrss_pipeline_job_result` 读取稳定的 `run_id`。拿到 `run_id` 之后,再通过 `get_run_status` / `get_delivery_payload` / `get_run_report` 读取状态与结果,而不是直接拼接目录路径。`run_freshrss_openclaw_pipeline` 仅保留为同步 debug / fallback 路径。
|
||||
在排查复杂问题时,也可以把 `debug_artifacts=true` 打开,并结合 `list_run_artifacts` 查看该 run 下实际产物。
|
||||
|
||||
## 对结构化摘要结果执行确定性过滤规则
|
||||
|
||||
```bash
|
||||
python scripts/run_filter_rules.py ^
|
||||
--summary outputs/reference/summary/result.loop.json ^
|
||||
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||
--output outputs/reference/filter/filter-decision.json
|
||||
```
|
||||
|
||||
如果你希望注入兴趣主题或来源标签,也可以额外传入 context 文件:
|
||||
|
||||
```bash
|
||||
python scripts/run_filter_rules.py ^
|
||||
--summary outputs/reference/summary/result.loop.json ^
|
||||
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||
--context outputs/reference/filter/filter-context.json ^
|
||||
--output outputs/reference/filter/filter-decision.with-context.json
|
||||
```
|
||||
|
||||
规则引擎设计和规则编写说明见:
|
||||
|
||||
- `docs/design/filter-rule-engine-design.md`
|
||||
- `docs/design/filter-rule-engine-usage.md`
|
||||
|
||||
## 将过滤结果写入 Markdown sink
|
||||
|
||||
```bash
|
||||
python scripts/run_markdown_sink.py ^
|
||||
--summary outputs/reference/summary/result.loop.json ^
|
||||
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||
--filter outputs/reference/filter/filter-decision.json
|
||||
```
|
||||
|
||||
脚本会把 Markdown 笔记写到 `knowledge-base/` 下。
|
||||
|
||||
## 构建内部 `ArticleCandidateRecord` 与精简版 `OpenClawCandidateInput`
|
||||
|
||||
```bash
|
||||
python scripts/run_article_candidate.py ^
|
||||
--summary outputs/reference/summary/result.loop.json ^
|
||||
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||
--filter outputs/reference/filter/filter-decision.json ^
|
||||
--section-hint tools_and_workflows
|
||||
```
|
||||
|
||||
默认会写出:
|
||||
|
||||
- `outputs/reference/candidates/article-candidate-record.json`
|
||||
- `outputs/reference/candidates/openclaw-candidate-input.json`
|
||||
|
||||
## 构建批量 OpenClaw delivery payload
|
||||
|
||||
```bash
|
||||
python scripts/build_openclaw_delivery.py ^
|
||||
--input-dir outputs/freshrss/candidates/batch ^
|
||||
--sort-by-rank ^
|
||||
--date 2026-03-25
|
||||
```
|
||||
|
||||
默认会写出:
|
||||
|
||||
- `outputs/reference/candidates/openclaw-delivery-payload.json`
|
||||
|
||||
输出目录布局说明见 `outputs/README.md`。
|
||||
|
||||
## 关键词索引默认配置
|
||||
|
||||
相关配置文件位于:
|
||||
|
||||
- `configs/term_aliases.json`
|
||||
- `configs/term_stopwords.json`
|
||||
- `configs/term_cleanup_policy.json`
|
||||
- `configs/term_watchlist.json`
|
||||
- `configs/term_change_log.json`
|
||||
|
||||
你也可以基于已有 delivery payload 重新构建关键词索引:
|
||||
|
||||
```bash
|
||||
python scripts/build_keyword_index.py ^
|
||||
--input outputs/reference/candidates/openclaw-delivery-payload.json
|
||||
```
|
||||
|
||||
运行期关键词数据存放在:
|
||||
|
||||
- `data/term_index/`
|
||||
|
||||
关键词清理评审 skill 位于:
|
||||
|
||||
- `skills/keyword-cleanup-review/`
|
||||
|
||||
构建给关键词治理流程使用的评审数据包(review bundle,临时工作文件):
|
||||
|
||||
```bash
|
||||
python skills/keyword-cleanup-review/scripts/build_review_bundle.py ^
|
||||
--days 7 ^
|
||||
--top 50 ^
|
||||
--output outputs/term_index/review/keyword-cleanup-bundle.json
|
||||
```
|
||||
|
||||
这个评审数据包(review bundle)会额外携带治理上下文:
|
||||
|
||||
- 来自 `configs/term_cleanup_policy.json` 的清理阈值
|
||||
- 当前 watch list(`configs/term_watchlist.json`)
|
||||
- 最近已应用的变更(`configs/term_change_log.json`)
|
||||
|
||||
接下来可以把 bundle 渲染成正式建议产物(默认走确定性规则,不把 LLM 作为默认路径):
|
||||
|
||||
```bash
|
||||
python scripts/generate_term_cleanup_suggestions.py ^
|
||||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json
|
||||
```
|
||||
|
||||
默认只生成:
|
||||
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`(唯一正式建议产物,建议短期保留)
|
||||
|
||||
如需人工审阅展示稿,再显式加:
|
||||
|
||||
```bash
|
||||
python scripts/generate_term_cleanup_suggestions.py ^
|
||||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json ^
|
||||
--emit-markdown
|
||||
```
|
||||
|
||||
这时才会额外生成:
|
||||
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`(临时展示稿,可按需生成,不必作为长期资产保留)
|
||||
|
||||
其中本轮最小版本优先覆盖 `interest_keyword_suggestions` 和 `watch_terms` 主链路;`alias_suggestions` / `stopword_suggestions` 先保持保守。
|
||||
|
||||
产物保留策略建议:
|
||||
|
||||
- `data/term_index/daily/YYYY-MM-DD.json`、`data/term_index/term_stats.json` 作为事实层长期保留
|
||||
- `configs/filter_context.personal.json`、`configs/term_watchlist.json`、`configs/term_aliases.json`、`configs/term_stopwords.json`、`configs/term_change_log.json` 作为状态层长期保留
|
||||
- `term-cleanup-suggestions-YYYY-MM-DD.json` 作为正式建议产物短期保留(最近少量几份或仅保留已应用过的)
|
||||
- `keyword-cleanup-bundle.json` 仅作为临时工作文件,默认只保留当前最新一份
|
||||
- `term-cleanup-suggestions-YYYY-MM-DD.md` 仅作为临时展示稿,优先按需生成,不建议默认长期归档
|
||||
|
||||
如果你想先预览已接受建议,再决定是否写配置文件:
|
||||
|
||||
```bash
|
||||
python scripts/apply_term_suggestions.py ^
|
||||
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json ^
|
||||
--accept-watch Cron Heartbeat Memory ^
|
||||
--dry-run
|
||||
```
|
||||
|
||||
去掉 `--dry-run` 后才会真正写文件。
|
||||
这个脚本也支持通过 `--accept-alias`、`--accept-stopword`、`--accept-interest` 应用 alias / stopword / interest keyword 变更。
|
||||
已接受的 watch 词会写入 `configs/term_watchlist.json`,每次应用动作也会被追加到 `configs/term_change_log.json`。
|
||||
|
||||
## 单篇总结 LLM 配置
|
||||
|
||||
如果你希望单篇总结后处理使用独立模型,而不影响主流水线,可以设置:
|
||||
|
||||
- `ARTICLE_SUMMARY_LLM_API_URL`
|
||||
- `ARTICLE_SUMMARY_LLM_MODEL`
|
||||
- `ARTICLE_SUMMARY_LLM_API_KEY`
|
||||
|
||||
如果这些变量未设置,单篇总结会回退使用主流程中的 `LLM_*` / `OPENAI_*` 配置。
|
||||
如果显式指定 DeepSeek 作为单篇总结模型且请求超时或失败,当前实现会自动再用主流程默认模型配置重试一次。
|
||||
|
||||
示例(PowerShell 风格):
|
||||
|
||||
```bash
|
||||
set ARTICLE_SUMMARY_LLM_API_URL=https://api.deepseek.com
|
||||
set ARTICLE_SUMMARY_LLM_API_KEY=your-article-summary-key
|
||||
set ARTICLE_SUMMARY_LLM_MODEL=deepseek-chat
|
||||
```
|
||||
|
||||
然后可以这样调用 CLI。
|
||||
正式生产环境建议优先使用 `outputs/freshrss/rerun/<run_id>/extracted/` 下的**逐条 extracted 文件**;下面这个**批量 extracted** 示例仅保留为兼容旧流程 / 临时场景输入:
|
||||
|
||||
```bash
|
||||
python scripts/run_article_summaries.py ^
|
||||
--extracted outputs/freshrss/extracted/freshrss.extracted.json ^
|
||||
--ids 12345 67890 ^
|
||||
--output-dir outputs/freshrss/single_summaries ^
|
||||
--timeout 120
|
||||
```
|
||||
|
||||
也可以通过 `summary_mcp.server` 暴露的 MCP 工具 `generate_article_summaries` 调用:
|
||||
|
||||
- `extracted_path`(string):单篇 extracted JSON 路径(例如 `outputs/freshrss/rerun/<run_id>/extracted/item-01.extracted.json`),或者包含 `results` 数组的 batch extracted JSON
|
||||
- `selected_ids`(array of strings):必填,至少传一个 `item_id`;如果传入的 ID 在 extracted payload 中一个都匹配不到,会直接报错
|
||||
- `output_dir`(optional string):Markdown 输出目录;若不传,则默认写到 extracted 文件旁边的 `single_summaries/` 目录
|
||||
- `llm_api_key` / `llm_model` / `llm_api_url`(optional strings):单篇总结 LLM 的覆盖配置;不传时会按前文规则回退到 `ARTICLE_SUMMARY_*` 或主 `LLM_*`
|
||||
|
||||
该工具返回一个 JSON 数组,内容为生成好的 Markdown 文件路径。
|
||||
|
||||
OpenClaw / 正式集成建议优先走异步 job:
|
||||
|
||||
- 调 `start_article_summary_job` 启动任务,立即拿到 `job_id`
|
||||
- 轮询 `get_article_summary_job_status(job_id)`,直到 `status` 进入 `success` 或 `failed`
|
||||
- 成功后调用 `get_article_summary_job_result(job_id)` 读取 `written_paths` 与结构化结果
|
||||
- 失败时优先查看 `error_summary` 与 job 目录中的 `job-report.json`
|
||||
|
||||
异步 job 状态目录固定落在 `outputs/freshrss/article_summary_jobs/<job_id>/`,最小会包含:
|
||||
|
||||
- `run-state.json`
|
||||
- `input.json`
|
||||
- `result.json`(成功时)
|
||||
- `job-report.json`
|
||||
|
||||
单篇总结使用独立 prompt:`outputs/prompts/article-summary-prompt.txt`。
|
||||
它与日报 prompt 完全独立,输出的是中文结构化知识笔记,包含这些部分:
|
||||
|
||||
- 核心结论
|
||||
- 主要论点
|
||||
- 关键方法 / 机制
|
||||
- 重要细节
|
||||
- 可复用启发
|
||||
- 关键词
|
||||
- 主题
|
||||
MIT
|
||||
|
||||
@@ -320,6 +320,23 @@
|
||||
|
||||
---
|
||||
|
||||
### [DONE][P1] interest/watch 候选引擎从固定阈值改为百分位排名 + 增速因子
|
||||
|
||||
目标:
|
||||
- 解决固定阈值(total_count>=3)不随数据量自适应的问题
|
||||
- 引入趋势信号(growth 因子),识别近期集中爆发的词
|
||||
- 支持 7 天、41 天、200 天数据量下取同样的 top 5%/5%-20% 而不需调阈值
|
||||
|
||||
要求:
|
||||
- `build_review_bundle.py`:新增 percentile 和 growth 计算函数;候选池从固定阈值改为百分位 + 增速
|
||||
- `configs/term_cleanup_policy.json`:升级为 v2 schema,percentile/growth 替代绝对阈值
|
||||
- 不改 `generate_term_cleanup_suggestions.py` 和 `apply_term_suggestions.py`
|
||||
- 全量跑一次对比新旧产出,确认差异合理
|
||||
|
||||
方案文档:`plans/keyword-cleanup-interest-watch-engine-improvement.md`
|
||||
|
||||
---
|
||||
|
||||
### [DONE][P3] 更新 README / handoff / docs,明确 MCP 为正式入口
|
||||
|
||||
目标:
|
||||
|
||||
@@ -27,18 +27,30 @@
|
||||
"Agent Skills",
|
||||
"AgentScope",
|
||||
"AI Agent",
|
||||
"AI Coding Agent",
|
||||
"AliSQL",
|
||||
"Anthropic",
|
||||
"Claude",
|
||||
"Claude Code",
|
||||
"CLAUDE.md",
|
||||
"CLI",
|
||||
"Context Engineering",
|
||||
"Cursor",
|
||||
"ChatGPT",
|
||||
"DeepSeek",
|
||||
"FastAPI",
|
||||
"Gin",
|
||||
"Go",
|
||||
"gRPC",
|
||||
"Harness Engineering",
|
||||
"Hermes Agent",
|
||||
"Java",
|
||||
"Kafka",
|
||||
"Kubernetes",
|
||||
"LLM",
|
||||
"Loop Engineering",
|
||||
"MCP",
|
||||
"MoE",
|
||||
"MySQL",
|
||||
"MySQL复制延迟",
|
||||
"OpenAI",
|
||||
@@ -47,15 +59,30 @@
|
||||
"Prompt Engineering",
|
||||
"Python",
|
||||
"RAG",
|
||||
"ReAct",
|
||||
"ReActAgent",
|
||||
"Redis",
|
||||
"Skill",
|
||||
"SKILL.md",
|
||||
"Skills",
|
||||
"Spring",
|
||||
"SubAgent",
|
||||
"TypeScript",
|
||||
"Vibe Coding",
|
||||
"Workflow",
|
||||
"上下文压缩",
|
||||
"上下文工程",
|
||||
"上下文管理",
|
||||
"云原生",
|
||||
"代码审查",
|
||||
"可观测性",
|
||||
"向量数据库",
|
||||
"多Agent协作",
|
||||
"大模型",
|
||||
"子Agent",
|
||||
"强化学习",
|
||||
"微服务",
|
||||
"渐进式披露",
|
||||
"知识库"
|
||||
]
|
||||
}
|
||||
|
||||
+146
-2
@@ -1,5 +1,149 @@
|
||||
{
|
||||
"AI助手": "AI Agent",
|
||||
"Agent": "Agent",
|
||||
"Agent框架": "Agent",
|
||||
"Agent能力": "Agent Skills",
|
||||
"智能体": "AI Agent",
|
||||
"Agentic架构": "Agentic架构",
|
||||
"多Agent协作": "多Agent协作",
|
||||
"多智能体架构": "多Agent协作",
|
||||
"Multi-Agent": "多Agent",
|
||||
"Subagent": "子Agent",
|
||||
"Sub Agents验证": "子Agent",
|
||||
"子Agent": "子Agent",
|
||||
"子智能体": "子Agent",
|
||||
"Coding Agent": "AI Coding Agent",
|
||||
"AI编程": "AI Coding Agent",
|
||||
"AI辅助编程": "AI Coding Agent",
|
||||
"代码生成": "AI代码生成",
|
||||
"代码审查": "Code Review",
|
||||
"Prompt": "Prompt Engineering",
|
||||
"Prompt Caching": "提示缓存",
|
||||
"RAG": "RAG",
|
||||
"图文RAG": "RAG",
|
||||
"Prompt架构": "Prompt Engineering"
|
||||
"向量检索": "向量检索",
|
||||
"向量嵌入": "向量嵌入",
|
||||
"Multi-Token Prediction": "多Token预测",
|
||||
"Pair-In Pair-Out": "PIPO架构",
|
||||
"PIPO": "PIPO架构",
|
||||
"上下文管理": "上下文管理",
|
||||
"上下文卸载": "上下文卸载",
|
||||
"Self-GC": "上下文压缩",
|
||||
"记忆管理": "上下文管理",
|
||||
"会话管理": "上下文管理",
|
||||
"Harness Engineering": "Harness工程化",
|
||||
"Harness架构": "Harness工程化",
|
||||
"Harness": "Harness工程化",
|
||||
"Loop Engineering": "Loop Engineering",
|
||||
"推理加速": "推理加速",
|
||||
"推理深度": "推理深度",
|
||||
"长链路推理": "长链路推理",
|
||||
"RLVR": "RLVR",
|
||||
"GRPO": "GRPO",
|
||||
"强化学习": "强化学习",
|
||||
"Multi-Agent RL": "多Agent强化学习",
|
||||
"Viking AI搜索": "AI搜索",
|
||||
"Viking AI Search": "AI搜索",
|
||||
"智能搜索": "AI搜索",
|
||||
"SearchCLI": "CLI搜索",
|
||||
"视频生成": "AI视频生成",
|
||||
"视频生成模型": "AI视频生成",
|
||||
"LingBot-Video": "AI视频生成",
|
||||
"视觉自回归模型": "AI视频生成",
|
||||
"火山云数据库PostgreSQL Serverless版": "Serverless数据库",
|
||||
"PostgreSQL": "PostgreSQL",
|
||||
"MySQL": "MySQL",
|
||||
"OceanBase": "OceanBase",
|
||||
"StarRocks": "StarRocks",
|
||||
"Milvus": "Milvus",
|
||||
"Seal AI Zone": "AI安全",
|
||||
"NEX沙箱": "沙箱隔离",
|
||||
"MicroVM": "沙箱隔离",
|
||||
"安全左移": "安全左移",
|
||||
"安全中台": "AI安全",
|
||||
"成本降低": "成本优化",
|
||||
"成本杠杆": "成本优化",
|
||||
"Scale-to-Zero": "弹性伸缩",
|
||||
"Data as Git": "数据分支管理",
|
||||
"Schema Diff": "Schema对比",
|
||||
"Time Travel": "数据回溯",
|
||||
"多端架构": "多端架构",
|
||||
"契约化": "契约化架构",
|
||||
"大仓": "大仓工程化",
|
||||
"Vibe Coding": "Vibe Coding",
|
||||
"LLM Judge": "LLM评估",
|
||||
"SWE-Bench": "SWE-Bench",
|
||||
"SWE Bench Pro": "SWE-Bench",
|
||||
"SWE-Bench Pro": "SWE-Bench",
|
||||
"Verification Agent": "验证Agent",
|
||||
"CLI工具": "CLI",
|
||||
"CLI": "CLI",
|
||||
"漏桶算法": "限流架构",
|
||||
"固定窗口限流": "限流架构",
|
||||
"Suspend消费控制": "限流架构",
|
||||
"RocketMQ LiteTopic": "消息队列",
|
||||
"LLM Wiki": "LLM知识库",
|
||||
"知识工程": "知识工程",
|
||||
"语义资产": "语义资产管理",
|
||||
"知识图谱": "知识图谱",
|
||||
"知识库沉淀": "知识管理",
|
||||
"Skill": "Skill",
|
||||
"Skill Hub": "技能生态",
|
||||
"具身智能": "具身智能",
|
||||
"Open X-Embodiment": "具身智能",
|
||||
"YOLO Classifier": "目标检测",
|
||||
"MCP": "MCP",
|
||||
"MCP连接器": "MCP",
|
||||
"缓存击穿": "缓存优化",
|
||||
"GPU算力调度": "算力调度",
|
||||
"异构资源": "异构计算",
|
||||
"XPU": "异构计算",
|
||||
"弹性RDMA": "RDMA网络",
|
||||
"国内主流GPU": "国产芯片",
|
||||
"国产AI芯片": "国产芯片",
|
||||
"Paxos协议": "分布式一致性",
|
||||
"Token": "Token管理",
|
||||
"Token效率": "Token管理",
|
||||
"百万token上下文": "长上下文",
|
||||
"MoE": "MoE架构",
|
||||
"MoE架构": "MoE架构",
|
||||
"思维链": "思维链",
|
||||
"CoT Distillation": "思维链蒸馏",
|
||||
"自然语言驱动": "自然语言交互",
|
||||
"NL2SQL": "NL2SQL",
|
||||
"AI对齐": "AI对齐",
|
||||
"注意力机制": "注意力机制",
|
||||
"多模态": "多模态",
|
||||
"音视频工作台": "音视频处理",
|
||||
"AI助手": "AI Agent",
|
||||
"Agent架构": "AI Agent",
|
||||
"Agent专业化": "AI Agent",
|
||||
"Agent Teams": "多Agent协作",
|
||||
"Agentic Engineering": "AI Agent",
|
||||
"AI智能体": "AI Agent",
|
||||
"LLM Agent": "AI Agent",
|
||||
"AI Harness": "Harness Engineering",
|
||||
"AI代码生成": "AI Coding Agent",
|
||||
"Memory管理": "上下文管理",
|
||||
"Agent Skill": "Agent Skills",
|
||||
"Binlog": "binlog",
|
||||
"vibe coding": "Vibe Coding",
|
||||
"Agent组织化协作平台": "Agent协作平台",
|
||||
"Anthropic": "Anthropic",
|
||||
"OpenClaw": "OpenClaw",
|
||||
"WorkBuddy": "WorkBuddy",
|
||||
"Claude": "Claude",
|
||||
"ChatGPT": "ChatGPT",
|
||||
"GPT": "GPT",
|
||||
"Opus": "Opus",
|
||||
"Sonnet": "Sonnet",
|
||||
"Grok": "Grok",
|
||||
"Qwen": "Qwen",
|
||||
"GLM": "GLM",
|
||||
"Claude Code": "Claude Code",
|
||||
"Cursor": "Cursor",
|
||||
"Codex": "Codex",
|
||||
"Pi": "Pi",
|
||||
"CoT": "CoT",
|
||||
"SVG": "SVG",
|
||||
"TTS": "TTS"
|
||||
}
|
||||
|
||||
@@ -99,6 +99,497 @@
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"suggestion_date": "2026-04-08",
|
||||
"based_on_days": 7
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Anthropic",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=13, days_seen=10, recent_count=13.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Harness Engineering",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=12, days_seen=11, recent_count=12.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Skill",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=11, days_seen=9, recent_count=11.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "上下文工程",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=6, recent_count=6.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "多Agent协作",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=6, recent_count=6.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Claude",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=5, recent_count=6.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "上下文管理",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=5, recent_count=6.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "渐进式披露",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=5, days_seen=5, recent_count=5.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Skills",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=5, days_seen=4, recent_count=5.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "SKILL.md",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=4, recent_count=4.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "CLAUDE.md",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=4, recent_count=4.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "上下文压缩",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=4, recent_count=4.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "AI Coding Agent",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=3, days_seen=3, recent_count=3.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Hermes Agent",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=3, days_seen=3, recent_count=3.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Vibe Coding",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=5, days_seen=4, recent_count=5.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Context Engineering",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=3, days_seen=3, recent_count=3.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Cursor",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=3, days_seen=3, recent_count=3.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "大模型",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=3, recent_count=4.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "TypeScript",
|
||||
"reason": "Core language for AI agent development (e.g., Claude Code, Cursor) and backend engineering, complements existing Python/Java/Go keywords.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "代码审查",
|
||||
"reason": "Chinese term for 'code review', a key practice in backend engineering and AI agent development workflows.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "Channels",
|
||||
"reason": "Too generic; could refer to communication channels, YouTube channels, or software channels, not specific to user's focus areas.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "Memory",
|
||||
"reason": "Extremely broad term; could refer to computer memory, human memory, or memory in various contexts, not discriminative enough.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "Prompt",
|
||||
"reason": "Already covered by 'Prompt Engineering' as a more specific term; 'Prompt' alone is too broad and matches many unrelated articles.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "AGI",
|
||||
"reason": "Too broad and speculative; not directly actionable for the user's practical engineering focus areas.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "AI日报",
|
||||
"reason": "Generic news term; not a technical concept or tool, would add noise to the keyword index.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "AIHOT",
|
||||
"reason": "Unclear meaning, likely a brand or aggregator, not a specific technical term.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "All In Code",
|
||||
"reason": "Too vague; could refer to a podcast, a philosophy, or a project, not a specific technical concept.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "auto-twitter-campaign",
|
||||
"reason": "Too specific to a single project/tool, not a general interest keyword for the user's focus areas.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "ChangeSet",
|
||||
"reason": "Generic term used in version control and databases; too broad to be a useful filter.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "Lumina",
|
||||
"reason": "Unclear reference; could be a product, framework, or brand, not clearly aligned with user's focus.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "OpenViking",
|
||||
"reason": "Unclear reference; not a known tool or concept in the user's stated focus areas.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "Seedance 2.0",
|
||||
"reason": "Unclear reference; likely a product or version, not a general technical term.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "质量门禁",
|
||||
"reason": "Chinese term for 'quality gate', too generic in software engineering; not specific to user's focus areas.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Agent Skill",
|
||||
"reason": "Singular variant",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "Agent Skills",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Binlog",
|
||||
"reason": "Case variant (auto-ranked)",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "binlog",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Coding Agent",
|
||||
"reason": "Abbreviated form of 'AI Coding Agent', referring to the same concept.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "AI Coding Agent",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Subagent",
|
||||
"reason": "Case variant",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "SubAgent",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Subagents",
|
||||
"reason": "Plural variant",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "SubAgent",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "vibe coding",
|
||||
"reason": "Case variant",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "Vibe Coding",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Agent架构",
|
||||
"reason": "Chinese translation of 'Agent architecture', a core concept in AI Agent engineering.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "AI Agent",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Agent专业化",
|
||||
"reason": "Chinese term for 'Agent specialization', directly related to Agent engineering.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "AI Agent",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Agent Teams",
|
||||
"reason": "English equivalent of 'Multi-Agent collaboration', same concept.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "多Agent协作",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Agentic Engineering",
|
||||
"reason": "Broader term for engineering with AI agents, closely related to Agent engineering focus.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "AI Agent",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "CLI工具",
|
||||
"reason": "Chinese translation of 'CLI tool', same concept.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "CLI",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "AI编程",
|
||||
"reason": "Chinese term for 'AI programming', closely related to AI Coding Agent.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "AI Coding Agent",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "记忆管理",
|
||||
"reason": "Chinese term for 'memory management', closely related to context management in LLM applications.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "上下文管理",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "会话管理",
|
||||
"reason": "Chinese term for 'session management', related to context management in LLM applications.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "上下文管理",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-07-15T02:17:50.155586Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "强化学习",
|
||||
"reason": "top 0.8% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=10, days_seen=10, recent_count=10.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
|
||||
"suggestion_date": "2026-07-15",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-07-15T02:17:50.155586Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "ReAct",
|
||||
"reason": "top 1.1% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=8, days_seen=8, recent_count=8.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
|
||||
"suggestion_date": "2026-07-15",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-07-15T02:17:50.155586Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "CLI",
|
||||
"reason": "top 1.1% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=8, days_seen=7, recent_count=8.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
|
||||
"suggestion_date": "2026-07-15",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-07-15T02:17:50.155586Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "子Agent",
|
||||
"reason": "top 1.3% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=7, days_seen=6, recent_count=7.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
|
||||
"suggestion_date": "2026-07-15",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-07-15T02:17:50.155586Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Loop Engineering",
|
||||
"reason": "top 1.5% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=6, recent_count=6.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
|
||||
"suggestion_date": "2026-07-15",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-07-15T02:17:50.155586Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "MoE",
|
||||
"reason": "top 2.0% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=5, days_seen=4, recent_count=5.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
|
||||
"suggestion_date": "2026-07-15",
|
||||
"based_on_days": 365
|
||||
}
|
||||
]
|
||||
}
|
||||
|
||||
@@ -1,14 +1,12 @@
|
||||
{
|
||||
"schema_version": "v1",
|
||||
"schema_version": "v2",
|
||||
"interest_keyword_review": {
|
||||
"min_total_count": 3,
|
||||
"min_days_seen": 2
|
||||
"percentile_max": 0.05,
|
||||
"growth_promotion": 0.5
|
||||
},
|
||||
"watch_term_review": {
|
||||
"min_total_count": 1,
|
||||
"min_days_seen": 1,
|
||||
"max_total_count": 2,
|
||||
"max_days_seen": 2
|
||||
"percentile_min": 0.05,
|
||||
"percentile_max": 0.20
|
||||
},
|
||||
"alias_review": {
|
||||
"min_total_count": 2,
|
||||
@@ -19,7 +17,10 @@
|
||||
"max_days_seen": 2
|
||||
},
|
||||
"notes": [
|
||||
"当前阶段采用保守阈值,避免在低样本条件下直接扩充 interest_keywords。",
|
||||
"watch_terms 先用于观察,后续再决定是否升格为 interest_keywords 或进入 alias/stopword 配置。"
|
||||
"v2: interest/watch 使用百分位排名 + 增速因子替代固定阈值",
|
||||
"percentile 越小表示排名越高(top 5% = percentile 0.05)",
|
||||
"growth = recent_count / total_count,衡量近期活跃度",
|
||||
"watch_term_review 的 percentile_min 可理解为兴趣边界下限,低于此值的词归入 interest 候选",
|
||||
"growth_promotion(默认 0.5)用于识别近期集中爆发词,即使排位不高也主动推荐确认"
|
||||
]
|
||||
}
|
||||
+121
-1
@@ -1,6 +1,126 @@
|
||||
[
|
||||
"1688",
|
||||
"AGI",
|
||||
"AIHOT",
|
||||
"AI日报",
|
||||
"All In Code",
|
||||
"Andrej Karpathy",
|
||||
"Anthropic",
|
||||
"auto-twitter-campaign",
|
||||
"Boundaries",
|
||||
"ChangeSet",
|
||||
"Channels",
|
||||
"Claude Fable 5",
|
||||
"Claude Mythos",
|
||||
"Cohere",
|
||||
"Confidence Head",
|
||||
"Cosmos 3",
|
||||
"Databricks",
|
||||
"DINOv2",
|
||||
"domain-mapping",
|
||||
"Dropbox",
|
||||
"EchoGen",
|
||||
"FLUX.1-dev VAE",
|
||||
"GB300 GPU",
|
||||
"GLM 5.2",
|
||||
"GLM5.0",
|
||||
"GPT-5.5",
|
||||
"GPT-Live",
|
||||
"GPT5.5",
|
||||
"Grok 4.5",
|
||||
"GrowBrain",
|
||||
"iMedImage",
|
||||
"iMedLoop",
|
||||
"iMedMaaS",
|
||||
"iMedStudio",
|
||||
"J-space",
|
||||
"JLens",
|
||||
"John Jumper",
|
||||
"J空间",
|
||||
"KAIROS",
|
||||
"KubeRay",
|
||||
"LibTV Agent",
|
||||
"LingBot-Video",
|
||||
"Lumina",
|
||||
"Markdown",
|
||||
"Marvis",
|
||||
"MDASH",
|
||||
"Meta Superintelligence Labs",
|
||||
"MTS",
|
||||
"Muse Image",
|
||||
"Muse Video",
|
||||
"N-gram Embedding",
|
||||
"OCP China",
|
||||
"OCP China 2026",
|
||||
"On-Policy Distillation",
|
||||
"OPC训练营",
|
||||
"OpenAI",
|
||||
"OpenBMC",
|
||||
"OpenClaw",
|
||||
"OpenViking",
|
||||
"Opus 4.8",
|
||||
"Qwen3",
|
||||
"Qwen3-30B-A3B",
|
||||
"RAS API",
|
||||
"Redfish",
|
||||
"ScMoE",
|
||||
"Seal AI Zone",
|
||||
"SealRouter",
|
||||
"Seedance 2.0",
|
||||
"Sonnet 5",
|
||||
"Spec模式",
|
||||
"STE固件团队",
|
||||
"Three.js",
|
||||
"Unity AI Gateway",
|
||||
"Vant Weapp",
|
||||
"WeTV",
|
||||
"WorkBuddy",
|
||||
"wpc",
|
||||
"YOLO Classifier",
|
||||
"一人公司",
|
||||
"中国科学技术大学",
|
||||
"五大扶持体系",
|
||||
"出门问问",
|
||||
"分镜",
|
||||
"剧本",
|
||||
"奋斗文化",
|
||||
"字节跳动",
|
||||
"小银",
|
||||
"得力",
|
||||
"德适科技",
|
||||
"成都天府长岛",
|
||||
"扣子",
|
||||
"星云平台",
|
||||
"火山引擎",
|
||||
"百度百舸",
|
||||
"百炼网关",
|
||||
"科大讯飞",
|
||||
"腾讯云开发者社区",
|
||||
"腾讯混元Hy3",
|
||||
"蚂蚁灵波",
|
||||
"贝尔实验室",
|
||||
"质量门禁",
|
||||
"配乐",
|
||||
"配音",
|
||||
"银行客户经理",
|
||||
"飞盘物理"
|
||||
"飞书妙搭",
|
||||
"飞盘物理",
|
||||
"自动化",
|
||||
"定时任务",
|
||||
"开源模型",
|
||||
"陌生化",
|
||||
"AlphaFold",
|
||||
"Brand Kit",
|
||||
"DataWorks",
|
||||
"Enhance-Nanocodec",
|
||||
"IRIS Codec",
|
||||
"Lovart",
|
||||
"MiniMax M3",
|
||||
"Gemini 3.5 Flash",
|
||||
"Codex",
|
||||
"CodeBuddy",
|
||||
"Claude Cowork",
|
||||
"AGENTS.md",
|
||||
"Claude",
|
||||
"RLVR"
|
||||
]
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"schema_version": "v1",
|
||||
"updated_at": "2026-04-08T02:34:14.194320Z",
|
||||
"updated_at": "2026-07-15T02:17:50.155586Z",
|
||||
"terms": [
|
||||
{
|
||||
"term": "A2A",
|
||||
|
||||
@@ -0,0 +1,151 @@
|
||||
# 关键词清洗流程概述
|
||||
|
||||
> 2026-05-14 初版
|
||||
> 从"数据记录"到"人工确认落盘"的完整链路
|
||||
|
||||
---
|
||||
|
||||
## 整体数据流
|
||||
|
||||
```
|
||||
每日日报 pipeline
|
||||
│
|
||||
▼
|
||||
term_index/daily/YYYY-MM-DD.json ← 每天一篇候选文章的热词统计
|
||||
│
|
||||
▼
|
||||
term_index/term_stats.json ← 所有 daily 的汇总(1070 个词)
|
||||
│
|
||||
├──── build_review_bundle.py ← 打包为审查数据包
|
||||
│ │
|
||||
│ ▼
|
||||
│ review/keyword-cleanup-bundle.json
|
||||
│ │
|
||||
│ ▼
|
||||
│ generate_term_cleanup_suggestions.py
|
||||
│ │
|
||||
│ ▼
|
||||
│ review/term-cleanup-suggestions-YYYY-MM-DD.json ← 正式建议产物
|
||||
│ │
|
||||
│ ▼
|
||||
│ (可选) review/term-cleanup-suggestions-YYYY-MM-DD.md ← 展示稿
|
||||
│
|
||||
├──── 人工确认哪些建议 accept
|
||||
│
|
||||
▼
|
||||
apply_term_suggestions.py ← 写入配置
|
||||
│
|
||||
├── configs/filter_context.personal.json ← interest_keywords
|
||||
├── configs/term_aliases.json ← alias
|
||||
├── configs/term_stopwords.json ← stopword
|
||||
├── configs/term_watchlist.json ← watch
|
||||
└── configs/term_change_log.json ← 变更日志
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 各环节说明
|
||||
|
||||
### 阶段 1:数据记录(每日自动)
|
||||
|
||||
```bash
|
||||
# FreshRSS pipeline 跑完后自动产出
|
||||
data/term_index/daily/2026-05-14.json
|
||||
```
|
||||
|
||||
- 每天一篇,记录当天候选文章中出现的热词
|
||||
- 包含 term、total_count、days_seen 等信息
|
||||
- 目前累计 **41 天**,共 **1070 个独立词**
|
||||
|
||||
### 阶段 2:全量汇总(每日自动)
|
||||
|
||||
```bash
|
||||
data/term_index/term_stats.json
|
||||
```
|
||||
|
||||
- 从所有 daily 文件重建,会覆盖重跑
|
||||
- 按 total_count 排序,前 5 名:OpenClaw(30)、Claude Code(25)、AI Agent(17)、Anthropic(13)、MCP(13)
|
||||
|
||||
### 阶段 3:构建审查数据包(手动触发)
|
||||
|
||||
```bash
|
||||
python skills/keyword-cleanup-review/scripts/build_review_bundle.py \
|
||||
--days 365 \
|
||||
--top 100 \
|
||||
--output outputs/term_index/review/keyword-cleanup-bundle.json
|
||||
```
|
||||
|
||||
- 把 term_stats + 当前配置打成一包,方便后续处理
|
||||
- 输出:`review/keyword-cleanup-bundle.json`
|
||||
|
||||
### 阶段 4:生成建议(手动触发)
|
||||
|
||||
```bash
|
||||
python scripts/generate_term_cleanup_suggestions.py \
|
||||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
|
||||
--emit-markdown
|
||||
```
|
||||
|
||||
#### 当前产出能力
|
||||
|
||||
| 建议类型 | 状态 | 当前阈值 | 说明 |
|
||||
|---------|------|----------|------|
|
||||
| interest_keyword_suggestions | ✅ **已实现** | total≥3, days≥2 | 产出 20 条 |
|
||||
| watch_terms | ✅ **已实现** | total≤2, days≤2 | 本次 0 条 |
|
||||
| alias_suggestions | ❌ **硬编码为空** | policy 有阈值(total≥2, days≥2)但脚本未实现 | |
|
||||
| stopword_suggestions | ❌ **硬编码为空** | policy 有阈值(total≤2, days≤2)但脚本未实现 | |
|
||||
|
||||
**关键发现:** alias 和 stopword 不是"阈值太保守",是 **generate 脚本里压根没写对应的生成函数**。policy 文件里阈值已经配好了(alias: min_total=2/min_days=2,stopword: max_total=2/max_days=2),但脚本第 376-380 行直接硬编码为 `[]` 和 `0`。
|
||||
|
||||
### 阶段 5:人工确认(手动)
|
||||
|
||||
```
|
||||
OpenClaw 把建议列给你 → 你确认哪些 accept → 我执行 apply
|
||||
```
|
||||
|
||||
本次模式:
|
||||
- 高频(≥5次/5天以上)→ 强烈推荐 ✅
|
||||
- 中频(3-4次)→ 附带建议 ✅
|
||||
- 泛词 → 建议跳过 ❌
|
||||
|
||||
### 阶段 6:落盘配置(手动)
|
||||
|
||||
```bash
|
||||
python scripts/apply_term_suggestions.py \
|
||||
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
|
||||
--accept-interest 词1 词2 ...
|
||||
```
|
||||
|
||||
- dry-run 预览 → 确认后正式 apply
|
||||
- 写入 `configs/filter_context.personal.json`
|
||||
- 同步记录到 `term_change_log.json`
|
||||
- **不备份原始配置**(待优化)
|
||||
- **apply 后不自动清理 review 目录**(待优化)
|
||||
|
||||
### 阶段 7:维护清理(按需)
|
||||
|
||||
由 OpenClaw 侧 `reader-keyword-maintenance` skill 处理:
|
||||
- 删除旧 markdown 展示稿
|
||||
- 保留最近一份 bundle
|
||||
- 保守保留 suggestions JSON
|
||||
|
||||
---
|
||||
|
||||
## 当前配置资产
|
||||
|
||||
| 文件 | 内容 | 数据量 |
|
||||
|------|------|--------|
|
||||
| `filter_context.personal.json` | interest_keywords | 52 个 |
|
||||
| `term_aliases.json` | 别名映射 | 0 组(未启用) |
|
||||
| `term_stopwords.json` | 停用词 | 0 个(未启用) |
|
||||
| `term_watchlist.json` | 观察词 | 6 个 |
|
||||
| `term_change_log.json` | 所有变更记录 | 已记录 |
|
||||
|
||||
---
|
||||
|
||||
## 待优化项
|
||||
|
||||
1. **alias/stopword 建议生成为空** — generate 脚本硬编码缺实现,policy 已有阈值,需要补函数
|
||||
2. **apply 前无配置备份** — 建议 apply 前自动 cp 备份
|
||||
3. **apply 后无自动收尾** — 建议 apply 后自动删旧 markdown 和 bundle
|
||||
4. **alias 识别依赖规则而非 LLM** — 当前全靠统计阈值,无法做语义级判断(如中英文映射、缩写展开)。如果需要高级 alias 识别,可以用 LLM 生成候选,规则脚本做 apply
|
||||
@@ -0,0 +1,290 @@
|
||||
# interest/watch 候选引擎改进方案
|
||||
|
||||
> 从固定阈值到自适应排位 + 趋势因子的演进
|
||||
|
||||
## 1. 背景
|
||||
|
||||
### 1.1 当前实现
|
||||
|
||||
`build_review_bundle.py` 使用固定的绝对阈值将未覆盖词(uncovered terms)划分为两个候选池:
|
||||
|
||||
| 候选池 | 判断条件 | 依据 |
|
||||
|--------|---------|------|
|
||||
| `interest_review_candidates` | `total_count >= 3 AND days_seen >= 2` | `policy.interest_keyword_review` |
|
||||
| `watch_review_candidates` | `total_count <= 2 AND days_seen <= 2` | `policy.watch_term_review` |
|
||||
|
||||
`generate_term_cleanup_suggestions.py` 则直接从这两个候选池过滤、去重、排序后输出。
|
||||
|
||||
### 1.2 当前方案的问题
|
||||
|
||||
**问题一:固定阈值不随数据量自适应**
|
||||
|
||||
```
|
||||
场景 total_count=3 意味着什么
|
||||
─────────────────────────────────────────────
|
||||
7 天数据(~200 词) top 15%,有一定区分度 ✅
|
||||
41 天数据(1070 词) top 5%,区分度更高 ✅ 但阈值没变
|
||||
未来 200 天 仍然用 3 次,区分度稀释 ❌
|
||||
```
|
||||
|
||||
同一个绝对次数,在不同数据规模下的语义完全不同。手工调阈值不可持续。
|
||||
|
||||
**问题二:固定阈值忽略趋势信号**
|
||||
|
||||
- "Anthropic":total=13, recent=7 — 近期高活跃,上升趋势
|
||||
- "Channels":total=3, recent=0 — 早期出现但近期消失
|
||||
- 当前引擎认为这两个词"都过了 3 次阈值",同等对待。实际一个是强烈买入信号,一个是过气词。
|
||||
|
||||
**问题三:interest 和 watch 的分界线是硬的**
|
||||
|
||||
total=3 → interest,total=2 → watch。一个词从 2 次变成 3 次就自动"升级",没有过渡、没有缓冲。
|
||||
|
||||
### 1.3 讨论结论
|
||||
|
||||
与老大讨论后确认:
|
||||
|
||||
1. interest/watch 是**统计判断**,不需要大模型介入,纯算法可以解决
|
||||
2. 当前引擎缺的不是大模型,而是**算法本身没写完**——自适应维度(排位、趋势)还没实现
|
||||
3. alias 和 stopword 需要语义判断,与 interest/watch 分属不同阶段,不在本方案范围内
|
||||
4. 修改量小,可以在 1 小时内落地
|
||||
|
||||
---
|
||||
|
||||
## 2. 设计方案
|
||||
|
||||
### 2.1 核心思路
|
||||
|
||||
引入两个互补维度替代固定阈值:
|
||||
|
||||
```
|
||||
判定维度 含义 数据来源
|
||||
────────────────────────────────────────────────────────────
|
||||
percentile(百分位排名) 该词 total_count 在所有词 term_stats
|
||||
中的排位占比
|
||||
growth(增速因子) 近期集中度 = recent_count daily 近 N 天
|
||||
/ total_count
|
||||
```
|
||||
|
||||
两个维度配合:
|
||||
|
||||
- **percentile** 衡量"这个词在当前数据集里有多突出"——消除数据量变化的影响
|
||||
- **growth** 衡量"这个词是持续出现还是近期爆发"——识别趋势信号
|
||||
|
||||
### 2.2 候选池划分逻辑
|
||||
|
||||
```
|
||||
percentile
|
||||
│
|
||||
┌─────────────────────┐
|
||||
│ top 5% │
|
||||
│ → 建议 interest │ ← 高频稳定词
|
||||
├─────────────────────┤
|
||||
│ top 5%-20% │
|
||||
│ → 建议 watch │ ← 有信号但未达 threshold
|
||||
├─────────────────────┤
|
||||
│ bottom 80% │
|
||||
│ → 暂不处理 │ ← 噪声/低频
|
||||
└─────────────────────┘
|
||||
|
||||
额外规则:
|
||||
如果词在 top 20% 之外,但 growth > 0.5(近期集中度高)
|
||||
→ 主动提升到 watch / 主动推 confirm
|
||||
```
|
||||
|
||||
这样就不需要关心"total_count 是 3 还是 5",只看数据自己说话。
|
||||
|
||||
### 2.3 接口变化
|
||||
|
||||
**`configs/term_cleanup_policy.json`**:
|
||||
|
||||
```json
|
||||
{
|
||||
"schema_version": "v2",
|
||||
"interest_keyword_review": {
|
||||
"percentile_max": 0.05,
|
||||
"growth_promotion": 0.5
|
||||
},
|
||||
"watch_term_review": {
|
||||
"percentile_min": 0.05,
|
||||
"percentile_max": 0.20
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
`v1` 的 `min_total_count`/`min_days_seen` 等绝对阈值字段不再使用。
|
||||
|
||||
**`build_review_bundle.py` 输出的候选项**:
|
||||
|
||||
```json
|
||||
{
|
||||
"term": "Anthropic",
|
||||
"total_count": 13,
|
||||
"days_seen": 10,
|
||||
"percentile": 0.012,
|
||||
"growth": 0.54,
|
||||
"reason": "top 1.2% by frequency, 54% of occurrences in recent window — strong signal."
|
||||
}
|
||||
```
|
||||
|
||||
### 2.4 不需要改动的部分
|
||||
|
||||
- `generate_term_cleanup_suggestions.py` — 它只消费候选池,不用改
|
||||
- `apply_term_suggestions.py` — 消费 suggestions JSON,不用改
|
||||
- `keyword-cleanup-bundle.json` 结构 — 向后兼容,新增 percentile/growth 字段
|
||||
|
||||
---
|
||||
|
||||
## 3. 实施计划
|
||||
|
||||
### 3.1 改动范围
|
||||
|
||||
| 文件 | 改动量 | 内容 |
|
||||
|------|--------|------|
|
||||
| `skills/keyword-cleanup-review/scripts/build_review_bundle.py` | ~40 行 | 新增 `_compute_percentile()` 和 `_compute_growth()` 函数;修改候选池生成逻辑;候选项中增加 percentile/growth |
|
||||
| `configs/term_cleanup_policy.json` | ~10 行 | schema v2:percentile/growth 替代绝对阈值 |
|
||||
|
||||
### 3.2 实施步骤
|
||||
|
||||
1. **build_review_bundle.py**:在 `top_global_terms` 生成后,增加 percentile 计算函数和 growth 计算函数
|
||||
2. **build_review_bundle.py**:修改 `interest_review_candidates` 和 `watch_review_candidates` 的生成逻辑,从固定阈值改为 percentile + growth
|
||||
3. **build_review_bundle.py**:候选项增加 `percentile` 和 `growth` 字段,更新 `reason` 文案
|
||||
4. **term_cleanup_policy.json**:更新为 v2 schema
|
||||
5. **验证**:全量跑一次(`--days 365 --top 100`),对比新旧两份输出的差异
|
||||
|
||||
### 3.3 验证方法
|
||||
|
||||
```bash
|
||||
# 1. 用旧版生成 baseline
|
||||
cd /home/ubuntu/zhu/github/reader
|
||||
python3.11 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
|
||||
--days 365 --top 100 \
|
||||
--output /tmp/bundle-baseline.json
|
||||
|
||||
# 2. 改代码后用新版生成
|
||||
python3.11 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
|
||||
--days 365 --top 100 \
|
||||
--output /tmp/bundle-new.json
|
||||
|
||||
# 3. 对比 governance_hints
|
||||
python3 -c "
|
||||
import json
|
||||
a = json.load(open('/tmp/bundle-baseline.json'))
|
||||
b = json.load(open('/tmp/bundle-new.json'))
|
||||
for key in ['interest_review_candidates', 'watch_review_candidates']:
|
||||
old = set(i['term'] for i in a['governance_hints'][key])
|
||||
new = set(i['term'] for i in b['governance_hints'][key])
|
||||
print(f'{key}: 新增={new-old}, 减少={old-new}')
|
||||
"
|
||||
```
|
||||
|
||||
### 3.4 风险
|
||||
|
||||
| 风险 | 概率 | 应对 |
|
||||
|------|------|------|
|
||||
| 百分位阈值对特小数据集(如只有 1 天数据)不适用 | 低 | 不足 7 天时降级回绝对阈值 |
|
||||
| growth 因子对低频词的偏差(total=1, recent=1 → growth=1) | 低 | growth 只对 total>=3 的词计算 |
|
||||
| 排位突变导致推荐漂移 | 低 | percentil 天然平滑,新增几天数据不会剧烈改变已有词的排位 |
|
||||
|
||||
---
|
||||
|
||||
## 4. alias/stopword 设计方案
|
||||
|
||||
### 4.1 核心判断
|
||||
|
||||
alias 和 stopword 需要语义理解,与 interest/watch(纯统计)性质不同。
|
||||
|
||||
| 类型 | 需要什么 | 判断方式 |
|
||||
|------|---------|----------|
|
||||
| 大小写变体 | 表层 | 规则:casefold 去重 |
|
||||
| 单复数 | 表层 | 规则:去/加 s 后缀匹配 |
|
||||
| 分词变体(空格/连字符) | 表层 | 规则:去空格归一 |
|
||||
| 简写全称(MCP→Model Context Protocol) | **语义** | LLM |
|
||||
| 中英文(上下文工程→Context Engineering) | **语义** | LLM |
|
||||
| 同义不同名(Rush→猿辅导 Rush 平台) | **语义** | LLM |
|
||||
| stopword(大模型、AI 太泛) | **语义** | LLM |
|
||||
|
||||
### 4.2 分层方案
|
||||
|
||||
```
|
||||
输入:高频未覆盖词 + 已有 interest 词表
|
||||
│
|
||||
├── 规则层(零成本)── 大小写归一、单复数、去空格/连字符
|
||||
│ 输出候选 alias 对
|
||||
│
|
||||
└── LLM 层(每次 ~500 token)── 把候选词表整批给 LLM
|
||||
做语义聚类
|
||||
输出 alias 组 + stopword 标记
|
||||
```
|
||||
|
||||
### 4.3 规则层设计
|
||||
|
||||
在 `generate_term_cleanup_suggestions.py` 中新增 `_prepare_alias_suggestions()` 函数:
|
||||
|
||||
```python
|
||||
def _prepare_alias_suggestions(top_terms, interest_keywords):
|
||||
"""
|
||||
基于表层规则生成 alias 建议。
|
||||
规则1:casefold 匹配——同一个 casefold 下有多个原文变体
|
||||
规则2:单复数——去掉/加上末尾 s 后匹配
|
||||
规则3:分词变体——去空格/连字符后匹配
|
||||
"""
|
||||
```
|
||||
|
||||
优势:零成本、可复现、可审计。直接写入 suggestions JSON,随 generate 一起输出。
|
||||
|
||||
### 4.4 LLM 层设计
|
||||
|
||||
单独脚本,非 generate 主链路的一部分。
|
||||
|
||||
```bash
|
||||
python scripts/generate_term_cleanup_semantic_suggestions.py \
|
||||
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
|
||||
--output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json
|
||||
```
|
||||
|
||||
LLM prompt 设计:
|
||||
|
||||
```
|
||||
你是一个关键词治理助手。以下是一个用户的 interest 关键词列表和一批未覆盖的高频词。
|
||||
请做三件事:
|
||||
|
||||
1. ALIAS:判断哪些未覆盖词是已有 interest 关键词的别名/变体
|
||||
2. STOPWORD:标记哪些词太宽泛/通用,建议排除
|
||||
3. PROMOTE:标记哪些新词与用户关注方向一致,建议加入 interest
|
||||
|
||||
用户关注方向:AI Agent 工程化、后端工程、开源工具、大模型落地
|
||||
```
|
||||
|
||||
LLM 层输出格式:
|
||||
|
||||
```json
|
||||
{
|
||||
"alias_suggestions": [
|
||||
{"from": "Context Engineering", "to": "上下文工程", "reason": "中英文对应同一概念"}
|
||||
],
|
||||
"stopword_suggestions": [
|
||||
{"term": "大模型", "reason": "过于宽泛,高频率但低区分度"}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
### 4.5 预期效果
|
||||
|
||||
| 覆盖类型 | 规则层 | LLM 层 |
|
||||
|---------|--------|--------|
|
||||
| 大小写变体 | ✅ | — |
|
||||
| 单复数 | ✅ | — |
|
||||
| 分词变体 | ✅ | — |
|
||||
| 简写全称 | — | ✅ |
|
||||
| 中英文映射 | — | ✅ |
|
||||
| 同义不同名 | — | ✅ |
|
||||
| stopword 判断 | — | ✅ |
|
||||
|
||||
---
|
||||
|
||||
## 5. 讨论记录
|
||||
|
||||
- 2026-05-14:与老大确认 interest/watch 不需要 LLM,纯算法可解决
|
||||
- 2026-05-14:确认百分位排名 + 增速因子方案,修改量小,优先落地
|
||||
- 2026-05-14:确认本方案不改 `generate_term_cleanup_suggestions.py` 和 `apply_term_suggestions.py`
|
||||
- 2026-05-14:确认 alias/stopword 采用规则层 + LLM 层分层方案,规则层零成本优先
|
||||
+314
@@ -0,0 +1,314 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
Generate semantic keyword suggestions using LLM.
|
||||
|
||||
Covers what surface-form rules cannot:
|
||||
- semantic alias (abbreviation ↔ full name, Chinese ↔ English, synonym)
|
||||
- stopword (overly broad / low-discrimination terms)
|
||||
- promote (new term that aligns with user's focus areas)
|
||||
|
||||
Usage:
|
||||
python scripts/generate_term_cleanup_semantic_suggestions.py \
|
||||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
|
||||
--output outputs/term_index/review/term-cleanup-semantic-suggestions-2026-05-14.json
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
import time
|
||||
from datetime import datetime, timezone
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
from urllib.request import Request, urlopen
|
||||
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
DEFAULT_BUNDLE_PATH = REPO_ROOT / "outputs" / "term_index" / "review" / "keyword-cleanup-bundle.json"
|
||||
DEFAULT_OUTPUT_DIR = REPO_ROOT / "outputs" / "term_index" / "review"
|
||||
|
||||
|
||||
def _load_json(path: Path) -> Any:
|
||||
return json.loads(path.read_text(encoding="utf-8-sig"))
|
||||
|
||||
|
||||
def _save_json(path: Path, payload: dict[str, Any]) -> None:
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
path.write_text(json.dumps(payload, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
|
||||
|
||||
|
||||
def _load_env(path: Path) -> dict[str, str]:
|
||||
"""Load key=value pairs from .env file."""
|
||||
env: dict[str, str] = {}
|
||||
if not path.exists():
|
||||
return env
|
||||
for line in path.read_text(encoding="utf-8").splitlines():
|
||||
line = line.strip()
|
||||
if not line or line.startswith("#") or "=" not in line:
|
||||
continue
|
||||
key, _, value = line.partition("=")
|
||||
env[key.strip()] = value.strip().strip("\"'")
|
||||
return env
|
||||
|
||||
|
||||
def _build_prompt(
|
||||
interest_keywords: list[str],
|
||||
rule_alias_suggestions: list[dict[str, str]],
|
||||
candidate_terms: list[dict[str, Any]],
|
||||
relevant_watch_terms: list[dict[str, Any]],
|
||||
) -> str:
|
||||
"""Build the LLM prompt for semantic suggestions."""
|
||||
|
||||
interest_bullets = "\n".join(f" - {t}" for t in sorted(interest_keywords))
|
||||
candidate_bullets = "\n".join(
|
||||
f" - {t['term']} (count={t['total_count']}, days={t['days_seen']})"
|
||||
for t in candidate_terms[:40]
|
||||
)
|
||||
|
||||
# Alias from rule layer (for LLM to build on, not duplicate)
|
||||
rule_alias_text = ""
|
||||
if rule_alias_suggestions:
|
||||
rule_alias_text = "\nSurface-form alias (already identified, skip these):\n" + "\n".join(
|
||||
f" {a['from']} → {a['to']} ({a['reason']})"
|
||||
for a in rule_alias_suggestions
|
||||
)
|
||||
|
||||
watch_text = ""
|
||||
if relevant_watch_terms:
|
||||
watch_text = "\nWatch terms (low-frequency but potentially relevant):\n" + "\n".join(
|
||||
f" {t['term']} (count={t['total_count']}, days={t['days_seen']})"
|
||||
for t in relevant_watch_terms[:20]
|
||||
)
|
||||
|
||||
return f"""You are a keyword governance assistant for an AI engineer. Your job is to analyze keyword data and produce structured suggestions.
|
||||
|
||||
## User's focus areas
|
||||
- AI Agent engineering (Skills, Harness, MCP, Agent architecture)
|
||||
- Backend engineering (Java, Go, Kubernetes, MySQL, distributed systems)
|
||||
- Open source AI tools and practices (Claude Code, Cursor, DeepSeek, OpenClaw)
|
||||
- LLM application engineering (context engineering, RAG, prompt engineering)
|
||||
|
||||
## Interest keywords (52 already configured)
|
||||
{interest_bullets}
|
||||
|
||||
## Uncovered candidate terms (sorted by frequency)
|
||||
{candidate_bullets}
|
||||
{watch_text}{rule_alias_text}
|
||||
|
||||
## Task
|
||||
Analyze the candidate terms and output a JSON object with exactly three keys:
|
||||
|
||||
1. "semantic_alias": array of alias suggestions that SURFACE RULES CAN'T CATCH (e.g. abbreviation↔full name, Chinese↔English, different naming for the same concept).
|
||||
Format: [{{"from": "<variant>", "to": "<canonical interest keyword>", "reason": "<why>"}}]
|
||||
|
||||
2. "stopword": array of terms that are too broad/generic to be useful as filters. A stopword is a term that appears frequently but has LOW DISCRIMINATION — it matches too many unrelated articles and clutters the keyword index.
|
||||
Format: [{{"term": "<term>", "reason": "<why it should be a stopword>"}}]
|
||||
|
||||
3. "promote_to_interest": array of uncovered terms that align well with the user's focus areas and should be added as interest keywords.
|
||||
Format: [{{"term": "<term>", "reason": "<why it fits>"}}]
|
||||
|
||||
## Rules
|
||||
- Be conservative. When in doubt, leave it out.
|
||||
- Only suggest alias for terms that clearly refer to the SAME concept as an existing interest keyword.
|
||||
- Only suggest stopword for terms that are genuinely too broad (appear in many unrelated contexts).
|
||||
- Only suggest promote for terms that clearly match the user's stated focus areas.
|
||||
- Output valid JSON only, no markdown, no explanation outside the JSON."""
|
||||
|
||||
|
||||
def _call_llm(prompt: str, api_url: str, model: str, api_key: str) -> str:
|
||||
"""Call LLM API and return the response text."""
|
||||
payload = json.dumps({
|
||||
"model": model,
|
||||
"messages": [{"role": "user", "content": prompt}],
|
||||
"temperature": 0.1,
|
||||
"max_tokens": 2048,
|
||||
}).encode("utf-8")
|
||||
|
||||
req = Request(
|
||||
api_url.rstrip("/") + "/chat/completions",
|
||||
data=payload,
|
||||
headers={
|
||||
"Content-Type": "application/json",
|
||||
"Authorization": f"Bearer {api_key}",
|
||||
},
|
||||
)
|
||||
|
||||
max_retries = 3
|
||||
for attempt in range(max_retries):
|
||||
try:
|
||||
with urlopen(req, timeout=120) as resp:
|
||||
result = json.loads(resp.read().decode("utf-8"))
|
||||
return result["choices"][0]["message"]["content"]
|
||||
except Exception as e:
|
||||
if attempt < max_retries - 1:
|
||||
wait = 2 ** attempt
|
||||
print(f" LLM call failed (attempt {attempt+1}/{max_retries}): {e}", file=sys.stderr)
|
||||
print(f" Retrying in {wait}s...", file=sys.stderr)
|
||||
time.sleep(wait)
|
||||
else:
|
||||
raise
|
||||
|
||||
|
||||
def _parse_llm_response(text: str) -> dict[str, list[dict[str, str]]]:
|
||||
"""Extract JSON from LLM response (may contain markdown fences)."""
|
||||
# Try to find JSON block
|
||||
json_match = re.search(r"```(?:json)?\s*\n?(\{.*?\})\s*\n?```", text, re.DOTALL)
|
||||
if json_match:
|
||||
text = json_match.group(1)
|
||||
|
||||
# Clean up: remove any text before { or after }
|
||||
start = text.find("{")
|
||||
end = text.rfind("}")
|
||||
if start >= 0 and end > start:
|
||||
text = text[start : end + 1]
|
||||
|
||||
try:
|
||||
result = json.loads(text)
|
||||
except json.JSONDecodeError:
|
||||
# Try partial recovery
|
||||
print(f" Warning: LLM response not clean JSON, attempting recovery", file=sys.stderr)
|
||||
print(f" Raw: {text[:500]}", file=sys.stderr)
|
||||
return {"semantic_alias": [], "stopword": [], "promote_to_interest": []}
|
||||
|
||||
# Normalize keys
|
||||
normalized = {
|
||||
"semantic_alias": result.get("semantic_alias", result.get("alias", [])),
|
||||
"stopword": result.get("stopword", result.get("stopword_suggestions", [])),
|
||||
"promote_to_interest": result.get("promote_to_interest", result.get("promote", [])),
|
||||
}
|
||||
# Ensure each is a list
|
||||
for key in normalized:
|
||||
if not isinstance(normalized[key], list):
|
||||
normalized[key] = []
|
||||
return normalized
|
||||
|
||||
|
||||
def main() -> None:
|
||||
parser = argparse.ArgumentParser(description="Generate semantic keyword suggestions via LLM.")
|
||||
parser.add_argument("--bundle", type=Path, default=DEFAULT_BUNDLE_PATH, help="Review bundle JSON path")
|
||||
parser.add_argument("--suggestions", type=Path, default=None, help="Existing suggestions JSON (for rule alias context)")
|
||||
parser.add_argument("--output", type=Path, default=None, help="Output JSON path (auto-generated if omitted)")
|
||||
parser.add_argument("--llm-api-url", type=str, default=None, help="LLM API base URL")
|
||||
parser.add_argument("--llm-model", type=str, default=None, help="LLM model name")
|
||||
parser.add_argument("--llm-api-key", type=str, default=None, help="LLM API key")
|
||||
parser.add_argument("--dry-run", action="store_true", help="Print prompt and exit without calling LLM")
|
||||
args = parser.parse_args()
|
||||
|
||||
# Load config
|
||||
env_path = REPO_ROOT / ".env"
|
||||
env = _load_env(env_path) if env_path.exists() else {}
|
||||
|
||||
api_url = args.llm_api_url or os.environ.get("LLM_API_URL") or env.get("LLM_API_URL", "https://api.deepseek.com")
|
||||
# Map OpenClaw model aliases to actual API model names
|
||||
model_raw = args.llm_model or os.environ.get("LLM_MODEL") or env.get("LLM_MODEL", "deepseek-chat")
|
||||
MODEL_ALIAS_MAP = {
|
||||
"deepseek/deepseek-v4-flash": "deepseek-chat",
|
||||
"deepseek/deepseek-chat": "deepseek-chat",
|
||||
"deepseek-v4-flash": "deepseek-chat",
|
||||
"deepseek-chat": "deepseek-chat",
|
||||
}
|
||||
model = MODEL_ALIAS_MAP.get(model_raw, model_raw)
|
||||
api_key = args.llm_api_key or os.environ.get("LLM_API_KEY") or env.get("LLM_API_KEY", "")
|
||||
|
||||
if not api_key:
|
||||
print("Error: No LLM API key found. Set LLM_API_KEY in .env or pass --llm-api-key.", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
# Load bundle
|
||||
if not args.bundle.exists():
|
||||
print(f"Error: Bundle not found: {args.bundle}", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
bundle = _load_json(args.bundle)
|
||||
current_config = bundle.get("current_config", {})
|
||||
interest_keywords = current_config.get("interest_keywords", [])
|
||||
top_global_terms = bundle.get("top_global_terms", [])
|
||||
governance_hints = bundle.get("governance_hints", {})
|
||||
|
||||
# Build candidate list (uncovered terms from interest + watch candidates)
|
||||
candidate_terms = []
|
||||
for item in governance_hints.get("interest_review_candidates", []):
|
||||
if isinstance(item, dict):
|
||||
candidate_terms.append({
|
||||
"term": item.get("term", ""),
|
||||
"total_count": item.get("total_count", 0),
|
||||
"days_seen": item.get("days_seen", 0),
|
||||
"percentile": item.get("percentile", 0),
|
||||
"growth": item.get("growth", 0),
|
||||
})
|
||||
for item in governance_hints.get("watch_review_candidates", []):
|
||||
if isinstance(item, dict):
|
||||
# Avoid duplicates
|
||||
if not any(c["term"] == item.get("term") for c in candidate_terms):
|
||||
candidate_terms.append({
|
||||
"term": item.get("term", ""),
|
||||
"total_count": item.get("total_count", 0),
|
||||
"days_seen": item.get("days_seen", 0),
|
||||
"percentile": item.get("percentile", 0),
|
||||
"growth": item.get("growth", 0),
|
||||
})
|
||||
|
||||
# Sort by total_count descending
|
||||
candidate_terms.sort(key=lambda x: -x["total_count"])
|
||||
relevant_watch_terms = governance_hints.get("watch_review_candidates", [])[:20]
|
||||
|
||||
# Load rule-layer alias suggestions if available
|
||||
rule_alias = []
|
||||
if args.suggestions and args.suggestions.exists():
|
||||
s = _load_json(args.suggestions)
|
||||
rule_alias = s.get("alias_suggestions", [])
|
||||
|
||||
# Build prompt
|
||||
prompt = _build_prompt(
|
||||
interest_keywords=interest_keywords,
|
||||
rule_alias_suggestions=rule_alias,
|
||||
candidate_terms=candidate_terms,
|
||||
relevant_watch_terms=relevant_watch_terms,
|
||||
)
|
||||
|
||||
# Determine output path
|
||||
suggestion_date = datetime.now(timezone.utc).date().isoformat()
|
||||
output_path = args.output or (DEFAULT_OUTPUT_DIR / f"term-cleanup-semantic-suggestions-{suggestion_date}.json")
|
||||
|
||||
if args.dry_run:
|
||||
print("=== DRY RUN: Prompt ===")
|
||||
print(prompt)
|
||||
print("\n=== END ===")
|
||||
print(f"\nWould write to: {output_path}")
|
||||
return
|
||||
|
||||
# Call LLM
|
||||
print(f"Calling LLM ({model})...", file=sys.stderr)
|
||||
response = _call_llm(prompt, api_url, model, api_key)
|
||||
print(f"LLM response received ({len(response)} chars)", file=sys.stderr)
|
||||
|
||||
# Parse
|
||||
parsed = _parse_llm_response(response)
|
||||
|
||||
# Build output
|
||||
output = {
|
||||
"date": suggestion_date,
|
||||
"source_bundle": str(args.bundle),
|
||||
"model": model,
|
||||
"interest_keyword_count": len(interest_keywords),
|
||||
"candidate_count": len(candidate_terms),
|
||||
**parsed,
|
||||
}
|
||||
|
||||
_save_json(output_path, output)
|
||||
|
||||
summary = {
|
||||
"output": str(output_path),
|
||||
"semantic_alias": len(output.get("semantic_alias", [])),
|
||||
"stopword": len(output.get("stopword", [])),
|
||||
"promote_to_interest": len(output.get("promote_to_interest", [])),
|
||||
}
|
||||
print(json.dumps(summary, ensure_ascii=False, indent=2))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -175,6 +175,121 @@ def _prepare_watch_suggestions(bundle: dict[str, Any], reserved_terms: set[str])
|
||||
return suggestions
|
||||
|
||||
|
||||
def _prepare_alias_suggestions(
|
||||
bundle: dict[str, Any],
|
||||
all_terms: list[dict[str, Any]] | None = None,
|
||||
) -> list[dict[str, Any]]:
|
||||
"""
|
||||
Generate alias suggestions using surface-form rules (no LLM).
|
||||
|
||||
Rules:
|
||||
1. casefold match — same normalized form, different original casing
|
||||
2. trailing-s singularization — singular/plural variants
|
||||
3. whitespace/hyphen normalization — word boundary variants
|
||||
|
||||
Scans all_terms (full term_stats) if provided; otherwise falls back
|
||||
to top_global_terms from the bundle.
|
||||
"""
|
||||
current_config = _require_dict(bundle.get("current_config"), "bundle.current_config")
|
||||
interest_keywords = _require_list(
|
||||
current_config.get("interest_keywords"), "bundle.current_config.interest_keywords"
|
||||
)
|
||||
top_global_terms = _require_list(bundle.get("top_global_terms"), "bundle.top_global_terms")
|
||||
source_terms = all_terms if all_terms is not None else top_global_terms
|
||||
|
||||
interest_set = {_term_key(t) for t in interest_keywords if isinstance(t, str)}
|
||||
interest_originals: set[str] = {t for t in interest_keywords if isinstance(t, str)}
|
||||
|
||||
# Build full casefold → [original forms] map
|
||||
cf_map: dict[str, list[str]] = {}
|
||||
for item in source_terms:
|
||||
term = None
|
||||
if isinstance(item, dict):
|
||||
term = item.get("term")
|
||||
elif isinstance(item, str):
|
||||
term = item
|
||||
if not isinstance(term, str) or not term.strip():
|
||||
continue
|
||||
key = _term_key(term)
|
||||
if key not in cf_map:
|
||||
cf_map[key] = []
|
||||
if term not in cf_map[key]:
|
||||
cf_map[key].append(term)
|
||||
|
||||
suggestions: list[dict[str, Any]] = []
|
||||
seen_pairs: set[tuple[str, str]] = set()
|
||||
|
||||
def _add(from_term: str, to_term: str, reason: str) -> None:
|
||||
pair = (_term_key(from_term), _term_key(to_term))
|
||||
if pair in seen_pairs:
|
||||
return
|
||||
seen_pairs.add(pair)
|
||||
suggestions.append({"from": from_term, "to": to_term, "reason": reason})
|
||||
|
||||
# Build a set of all term keys from source for quick lookup
|
||||
source_keys = set(cf_map.keys())
|
||||
|
||||
# Rule 1: casefold match — same normalized form, different casing
|
||||
for key, variants in cf_map.items():
|
||||
if len(variants) < 2:
|
||||
continue
|
||||
canonical = None
|
||||
alt_forms = []
|
||||
for v in variants:
|
||||
if v in interest_originals:
|
||||
canonical = v
|
||||
else:
|
||||
alt_forms.append(v)
|
||||
if canonical and alt_forms:
|
||||
for alt in alt_forms:
|
||||
_add(alt, canonical, "Case variant")
|
||||
elif len(variants) >= 2 and not canonical:
|
||||
# None is canonical — suggest the highest-frequency form
|
||||
ranked = sorted(variants, key=lambda t: -(
|
||||
next(
|
||||
(it.get("total_count", 0) for it in top_global_terms if it.get("term") == t),
|
||||
0,
|
||||
)
|
||||
))
|
||||
for alt in ranked[1:]:
|
||||
_add(alt, ranked[0], "Case variant (auto-ranked)")
|
||||
|
||||
# Rule 2: singular/plural — trailing-s normalization
|
||||
# Check all source terms (not just interest keys) for bidirectional matching
|
||||
for key in source_keys:
|
||||
if key in interest_set:
|
||||
continue
|
||||
if key.endswith("s") and len(key) > 2:
|
||||
singular_key = key.rstrip("s")
|
||||
if singular_key in interest_set and singular_key != key:
|
||||
# Find canonical interest keyword
|
||||
canon = next((t for t in interest_keywords if _term_key(t) == singular_key), None)
|
||||
from_form = cf_map[key][0]
|
||||
if canon:
|
||||
_add(from_form, canon, "Plural variant")
|
||||
# singular form → interest has plural
|
||||
plural_key = key + "s"
|
||||
if plural_key in interest_set and plural_key != key:
|
||||
canon = next((t for t in interest_keywords if _term_key(t) == plural_key), None)
|
||||
from_form = cf_map[key][0]
|
||||
if canon:
|
||||
_add(from_form, canon, "Singular variant")
|
||||
|
||||
# Rule 3: whitespace/hyphen normalization
|
||||
for key in source_keys:
|
||||
if key in interest_set:
|
||||
continue
|
||||
normalized = key.replace("-", "").replace("_", "").replace(" ", "")
|
||||
if normalized in interest_set and normalized != key:
|
||||
canon = next((t for t in interest_keywords if _term_key(t) == normalized), None)
|
||||
from_form = cf_map[key][0]
|
||||
if canon:
|
||||
_add(from_form, canon, "Whitespace/punctuation variant")
|
||||
|
||||
suggestions.sort(key=lambda x: (x["from"].casefold(), x["to"].casefold()))
|
||||
return suggestions
|
||||
|
||||
|
||||
def _render_table(items: list[dict[str, Any]]) -> str:
|
||||
if not items:
|
||||
return "_None in this pass._\n"
|
||||
@@ -361,9 +476,19 @@ def main() -> None:
|
||||
markdown_output=args.markdown_output,
|
||||
)
|
||||
|
||||
# Load full term_stats for alias scanning (bundle only has top N)
|
||||
stats_path = REPO_ROOT / "data" / "term_index" / "term_stats.json"
|
||||
all_stats_terms: list[str] = []
|
||||
if stats_path.exists():
|
||||
stats_payload = _load_json(stats_path)
|
||||
raw_terms = stats_payload.get("terms") if isinstance(stats_payload, dict) else []
|
||||
if isinstance(raw_terms, list):
|
||||
all_stats_terms = [str(t["term"]) for t in raw_terms if isinstance(t, dict) and isinstance(t.get("term"), str)]
|
||||
|
||||
interest_items = _prepare_interest_suggestions(bundle)
|
||||
reserved_terms = {_term_key(str(item.get("term") or "")) for item in interest_items}
|
||||
watch_items = _prepare_watch_suggestions(bundle, reserved_terms=reserved_terms)
|
||||
alias_items = _prepare_alias_suggestions(bundle, all_terms=all_stats_terms)
|
||||
|
||||
suggestions = {
|
||||
"date": suggestion_date,
|
||||
@@ -373,10 +498,10 @@ def main() -> None:
|
||||
"summary": {
|
||||
"interest_keyword_suggestions": len(interest_items),
|
||||
"watch_terms": len(watch_items),
|
||||
"alias_suggestions": 0,
|
||||
"alias_suggestions": len(alias_items),
|
||||
"stopword_suggestions": 0,
|
||||
},
|
||||
"alias_suggestions": [],
|
||||
"alias_suggestions": alias_items,
|
||||
"stopword_suggestions": [],
|
||||
"interest_keyword_suggestions": interest_items,
|
||||
"watch_terms": watch_items,
|
||||
@@ -400,7 +525,7 @@ def main() -> None:
|
||||
"markdown_output": str(markdown_output_path) if args.emit_markdown else None,
|
||||
"interest_keyword_suggestions": len(interest_items),
|
||||
"watch_terms": len(watch_items),
|
||||
"alias_suggestions": 0,
|
||||
"alias_suggestions": len(alias_items),
|
||||
"stopword_suggestions": 0,
|
||||
"emit_markdown": args.emit_markdown,
|
||||
}
|
||||
|
||||
@@ -26,7 +26,7 @@ description: 生成 reader 项目的正式关键词 review 输入。当用户需
|
||||
|
||||
## 工作流程
|
||||
|
||||
1. 构建精简的审查数据包(临时工作文件):
|
||||
### Phase 1:构建审查数据包
|
||||
|
||||
```bash
|
||||
python skills/keyword-cleanup-review/scripts/build_review_bundle.py
|
||||
@@ -34,51 +34,92 @@ python skills/keyword-cleanup-review/scripts/build_review_bundle.py
|
||||
|
||||
可选参数:
|
||||
|
||||
- `--days 7`
|
||||
- `--top 50`
|
||||
- `--days 7`(默认 7,建议传 365 覆盖全量)
|
||||
- `--top 100`(考虑的词数)
|
||||
- `--output outputs/term_index/review/keyword-cleanup-bundle.json`
|
||||
|
||||
2. 阅读建议模式:
|
||||
#### 候选引擎策略
|
||||
|
||||
- `skills/keyword-cleanup-review/references/suggestion-schema.md`
|
||||
根据 `configs/term_cleanup_policy.json` 的 `schema_version` 自动切换:
|
||||
|
||||
3. 运行 suggestions 生成脚本:
|
||||
| 版本 | 策略 | 说明 |
|
||||
|------|------|------|
|
||||
| v1(旧) | 固定阈值(total≥3/days≥2 → interest) | 小数据集兼容 |
|
||||
| v2(当前默认) | 百分位排名 + 增速因子 | 自适应数据量,不需要手工调阈值 |
|
||||
|
||||
v2 策略说明:
|
||||
- **percentile**:total_count 在所有词里的排位占比。top 5% → interest 候选,5%-20% → watch 候选
|
||||
- **growth**:recent_count / total_count,衡量近期活跃度。growth≥0.5 的排位外词也会主动推荐
|
||||
|
||||
### Phase 2:生成建议(规则层)
|
||||
|
||||
```bash
|
||||
python scripts/generate_term_cleanup_suggestions.py ^
|
||||
python scripts/generate_term_cleanup_suggestions.py \
|
||||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json
|
||||
```
|
||||
|
||||
默认生成:
|
||||
|
||||
- 一份符合模式的 JSON 建议文件(正式建议产物,也是 review / apply 之间唯一正式输入)
|
||||
|
||||
如需人工审阅展示稿,再显式加:
|
||||
如需人工审阅展示稿:
|
||||
|
||||
```bash
|
||||
python scripts/generate_term_cleanup_suggestions.py ^
|
||||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json ^
|
||||
python scripts/generate_term_cleanup_suggestions.py \
|
||||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
|
||||
--emit-markdown
|
||||
```
|
||||
|
||||
这时才会额外生成:
|
||||
#### 产出能力
|
||||
|
||||
- 一份简短的供人工审阅的 Markdown 报告(临时展示稿)
|
||||
| 建议类型 | 状态 | 方法 |
|
||||
|---------|------|------|
|
||||
| interest 建议 | ✅ 已实现 | 百分位 top 5% + 增速促活 |
|
||||
| watch 建议 | ✅ 已实现 | 百分位 5%-20% |
|
||||
| alias 建议 | ✅ 已实现 | 规则层:大小写归一、单复数、去空格/连字符 |
|
||||
| stopword 建议 | ❌ 规则层空缺 | 见 Phase 3(LLM 层) |
|
||||
|
||||
4. 严格保持边界:
|
||||
默认生成:
|
||||
- `term-cleanup-suggestions-YYYY-MM-DD.json`(正式建议产物)
|
||||
|
||||
显式加 `--emit-markdown` 额外生成:
|
||||
- `term-cleanup-suggestions-YYYY-MM-DD.md`(临时展示稿)
|
||||
|
||||
### Phase 3:生成建议(LLM 层,可选)
|
||||
|
||||
规则层覆盖不了 alias(中英文对应、缩写展开、同义不同名)和 stopword 判断,需要 LLM 辅助:
|
||||
|
||||
```bash
|
||||
python scripts/generate_term_cleanup_semantic_suggestions.py \
|
||||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
|
||||
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
|
||||
--output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json
|
||||
```
|
||||
|
||||
从 `.env` 读取 LLM 配置(`LLM_API_URL` / `LLM_MODEL` / `LLM_API_KEY`),使用 DeepSeek API。
|
||||
|
||||
输出三部分:
|
||||
|
||||
| 输出 | 说明 |
|
||||
|------|------|
|
||||
| `semantic_alias` | 语义级别名(中英文、缩写、同义不同名) |
|
||||
| `stopword` | 泛词过滤建议(规则层做不了的需要语义判断的) |
|
||||
| `promote_to_interest` | 与用户关注方向一致的新词,建议加入 interest |
|
||||
|
||||
**注:LLM 层产物是候选,不应自动 apply,需要人工确认后由 OpenClaw 编排 apply。**
|
||||
|
||||
### Phase 4:输出给 OpenClaw 编排
|
||||
|
||||
- `suggestions JSON` = review / apply 之间唯一正式建议输入
|
||||
- `semantic-suggestions JSON` = LLM 补充建议,需要人工筛选后合并到 suggestions JSON 再 apply
|
||||
- Markdown = 临时展示层
|
||||
- 后续汇报、确认、dry-run、apply、收尾清理由 OpenClaw 编排层执行
|
||||
|
||||
### Phase 5:严格保持边界
|
||||
|
||||
- 建议 `configs/term_aliases.json` 的修改
|
||||
- 建议 `configs/term_stopwords.json` 的修改
|
||||
- 建议 `configs/filter_context.personal.json` 的新增
|
||||
- **LLM 层产出(semantic-suggestions)不自动 apply**,需人工确认后由 OpenClaw 编排层执行
|
||||
- 除非用户明确要求,否则不要直接编辑这些文件
|
||||
- 除非用户要求修改规则逻辑,否则不要建议直接编辑 `configs/filter_rules.json`
|
||||
|
||||
5. 输出交接口径:
|
||||
|
||||
- 将 JSON suggestions 视为正式 review 输入
|
||||
- 将 Markdown 视为可选展示层
|
||||
- 后续汇报、确认、dry-run apply、正式 apply、收尾清理应由 OpenClaw 编排层继续执行
|
||||
|
||||
## 审查启发式规则
|
||||
|
||||
优先考虑以下决策:
|
||||
@@ -141,6 +182,7 @@ JSON 输出应遵循:
|
||||
短期保留:
|
||||
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`
|
||||
- `outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json`
|
||||
|
||||
临时产物:
|
||||
|
||||
@@ -173,7 +215,9 @@ JSON 输出应遵循:
|
||||
## 资源
|
||||
|
||||
- 脚本:
|
||||
- `scripts/build_review_bundle.py`
|
||||
- `skills/keyword-cleanup-review/scripts/build_review_bundle.py`
|
||||
- `scripts/generate_term_cleanup_suggestions.py`
|
||||
- `scripts/generate_term_cleanup_semantic_suggestions.py`(LLM 层)
|
||||
- 参考文档:
|
||||
- `references/suggestion-schema.md`
|
||||
- `plans/keyword-cleanup-interest-watch-engine-improvement.md`(v2 引擎设计)
|
||||
|
||||
@@ -9,16 +9,15 @@ from typing import Any
|
||||
|
||||
|
||||
DEFAULT_POLICY: dict[str, Any] = {
|
||||
"schema_version": "v1",
|
||||
"schema_version": "v2",
|
||||
"interest_keyword_review": {
|
||||
"min_total_count": 3,
|
||||
"min_days_seen": 2,
|
||||
"percentile_min": 0.0,
|
||||
"percentile_max": 0.05,
|
||||
"growth_promotion": 0.5,
|
||||
},
|
||||
"watch_term_review": {
|
||||
"min_total_count": 1,
|
||||
"min_days_seen": 1,
|
||||
"max_total_count": 2,
|
||||
"max_days_seen": 2,
|
||||
"percentile_min": 0.05,
|
||||
"percentile_max": 0.20,
|
||||
},
|
||||
"alias_review": {
|
||||
"min_total_count": 2,
|
||||
@@ -28,6 +27,11 @@ DEFAULT_POLICY: dict[str, Any] = {
|
||||
"max_total_count": 2,
|
||||
"max_days_seen": 2,
|
||||
},
|
||||
"notes": [
|
||||
"v2: interest/watch 使用百分位排名 + 增速因子替代固定阈值",
|
||||
"percentile 越小表示排名越高(top 5% = percentile 0.05)",
|
||||
"growth = recent_count / total_count,衡量近期活跃度",
|
||||
],
|
||||
}
|
||||
|
||||
|
||||
@@ -126,6 +130,37 @@ def _within_watch_thresholds(item: dict[str, Any], thresholds: dict[str, Any]) -
|
||||
)
|
||||
|
||||
|
||||
def _compute_percentile(value: int, sorted_values: list[int]) -> float:
|
||||
"""
|
||||
Return the percentile rank of `value` in `sorted_values` (ascending).
|
||||
0.0 = highest frequency (top rank), 1.0 = lowest frequency (bottom rank).
|
||||
"""
|
||||
if not sorted_values:
|
||||
return 1.0
|
||||
# bisect_left — count of values strictly less than `value`
|
||||
lo, hi = 0, len(sorted_values)
|
||||
while lo < hi:
|
||||
mid = (lo + hi) // 2
|
||||
if sorted_values[mid] < value:
|
||||
lo = mid + 1
|
||||
else:
|
||||
hi = mid
|
||||
rank = lo
|
||||
# invert: smallest value → rank=0 → 1.0 (bottom)
|
||||
# largest value → rank=len → 0.0 (top)
|
||||
return 1.0 - (rank / len(sorted_values))
|
||||
|
||||
|
||||
def _compute_growth(recent_count: int, total_count: int) -> float:
|
||||
"""
|
||||
Return growth factor: recent_count / total_count.
|
||||
Only meaningful when total_count >= 3; returns 0.0 for small counts.
|
||||
"""
|
||||
if total_count < 3:
|
||||
return 0.0
|
||||
return recent_count / total_count
|
||||
|
||||
|
||||
def main() -> None:
|
||||
parser = argparse.ArgumentParser(
|
||||
description="Build a compact review bundle for the keyword-cleanup-review skill."
|
||||
@@ -235,6 +270,13 @@ def main() -> None:
|
||||
alias_values = _casefold_set(list(aliases.values()))
|
||||
watch_set = _casefold_set([str(item.get("term", "")) for item in watchlist])
|
||||
|
||||
# Build a sorted list of all total_counts for percentile computation
|
||||
all_total_counts = sorted(
|
||||
int(item.get("total_count") or 0)
|
||||
for item in stats_terms
|
||||
if isinstance(item, dict) and isinstance(item.get("term"), str)
|
||||
)
|
||||
|
||||
top_global_terms = []
|
||||
for item in stats_terms[: args.top]:
|
||||
if not isinstance(item, dict):
|
||||
@@ -256,42 +298,108 @@ def main() -> None:
|
||||
"is_alias_target": folded in alias_values,
|
||||
"in_watchlist": folded in watch_set,
|
||||
"recent_count": recent_counter.get(term, 0),
|
||||
"percentile": _compute_percentile(
|
||||
int(item.get("total_count") or 0), all_total_counts
|
||||
),
|
||||
"growth": _compute_growth(
|
||||
recent_counter.get(term, 0),
|
||||
int(item.get("total_count") or 0),
|
||||
),
|
||||
}
|
||||
)
|
||||
|
||||
# Keep more uncovered terms for percentile-based selection
|
||||
uncovered_terms = [
|
||||
item for item in top_global_terms if not item["in_interest_keywords"] and not item["is_stopword"]
|
||||
][:20]
|
||||
][:100]
|
||||
|
||||
policy_version = (policy.get("schema_version") if isinstance(policy, dict) else None) or "v1"
|
||||
interest_thresholds = policy.get("interest_keyword_review") if isinstance(policy, dict) else {}
|
||||
watch_thresholds = policy.get("watch_term_review") if isinstance(policy, dict) else {}
|
||||
|
||||
if policy_version == "v2" or "percentile_max" in interest_thresholds:
|
||||
# v2: percentile + growth based selection
|
||||
pct_min_interest = float(interest_thresholds.get("percentile_min", 0.0))
|
||||
pct_max_interest = float(interest_thresholds.get("percentile_max", 0.05))
|
||||
growth_promo = float(interest_thresholds.get("growth_promotion", 0.5))
|
||||
pct_min_watch = float(watch_thresholds.get("percentile_min", 0.05))
|
||||
pct_max_watch = float(watch_thresholds.get("percentile_max", 0.20))
|
||||
|
||||
interest_candidates_raw = [
|
||||
item for item in uncovered_terms
|
||||
if not item["in_watchlist"]
|
||||
and pct_min_interest <= item["percentile"] <= pct_max_interest
|
||||
]
|
||||
watch_candidates_raw = [
|
||||
item for item in uncovered_terms
|
||||
if not item["in_watchlist"]
|
||||
and pct_min_watch < item["percentile"] <= pct_max_watch
|
||||
]
|
||||
# Growth boost: terms outside watch range but with strong growth signal
|
||||
growth_boost_candidates = [
|
||||
item for item in uncovered_terms
|
||||
if not item["in_watchlist"]
|
||||
and item["percentile"] > pct_max_watch
|
||||
and item["growth"] >= growth_promo
|
||||
]
|
||||
else:
|
||||
# v1 fallback: fixed thresholds
|
||||
interest_candidates_raw = [
|
||||
item for item in uncovered_terms
|
||||
if not item["in_watchlist"] and _meets_min_thresholds(item, interest_thresholds)
|
||||
]
|
||||
watch_candidates_raw = [
|
||||
item for item in uncovered_terms
|
||||
if not item["in_watchlist"]
|
||||
and not _meets_min_thresholds(item, interest_thresholds)
|
||||
and _within_watch_thresholds(item, watch_thresholds)
|
||||
]
|
||||
growth_boost_candidates = []
|
||||
|
||||
interest_review_candidates = [
|
||||
{
|
||||
"term": item["term"],
|
||||
"total_count": item["total_count"],
|
||||
"days_seen": item["days_seen"],
|
||||
"percentile": item["percentile"],
|
||||
"growth": item["growth"],
|
||||
"reason": (
|
||||
"Meets the configured interest-keyword review threshold and is not yet covered "
|
||||
"by interest keywords or stopwords."
|
||||
f"top {item['percentile']:.1%} by frequency,"
|
||||
f"growth={item['growth']:.0%},"
|
||||
"not yet covered by interest keywords or stopwords."
|
||||
),
|
||||
}
|
||||
for item in uncovered_terms
|
||||
if not item["in_watchlist"] and _meets_min_thresholds(item, interest_thresholds)
|
||||
for item in interest_candidates_raw
|
||||
][:20]
|
||||
watch_review_candidates = [
|
||||
{
|
||||
"term": item["term"],
|
||||
"total_count": item["total_count"],
|
||||
"days_seen": item["days_seen"],
|
||||
"percentile": item["percentile"],
|
||||
"growth": item["growth"],
|
||||
"reason": (
|
||||
"Falls into the configured watch-term review range and should be observed "
|
||||
"before promotion into interest keywords."
|
||||
f"top {item['percentile']:.1%} by frequency,"
|
||||
f"growth={item['growth']:.0%},"
|
||||
"fell into watch-review range."
|
||||
),
|
||||
}
|
||||
for item in uncovered_terms
|
||||
if not item["in_watchlist"]
|
||||
and not _meets_min_thresholds(item, interest_thresholds)
|
||||
and _within_watch_thresholds(item, watch_thresholds)
|
||||
for item in watch_candidates_raw
|
||||
][:20]
|
||||
growth_boost_review_items = [
|
||||
{
|
||||
"term": item["term"],
|
||||
"total_count": item["total_count"],
|
||||
"days_seen": item["days_seen"],
|
||||
"percentile": item["percentile"],
|
||||
"growth": item["growth"],
|
||||
"reason": (
|
||||
f"growth spike: {item['growth']:.0%} of occurrences in recent window "
|
||||
f"(total={item['total_count']}, days={item['days_seen']})."
|
||||
),
|
||||
}
|
||||
for item in growth_boost_candidates
|
||||
][:5]
|
||||
recent_hot_terms = sorted(
|
||||
({"term": term, "recent_count": count} for term, count in recent_counter.items()),
|
||||
key=lambda item: (-item["recent_count"], item["term"].casefold(), item["term"]),
|
||||
@@ -330,6 +438,7 @@ def main() -> None:
|
||||
"governance_hints": {
|
||||
"interest_review_candidates": interest_review_candidates,
|
||||
"watch_review_candidates": watch_review_candidates,
|
||||
"growth_boost_review_items": growth_boost_review_items,
|
||||
},
|
||||
}
|
||||
_save_json(args.output, bundle)
|
||||
|
||||
@@ -0,0 +1,186 @@
|
||||
---
|
||||
name: reader-digest-flow
|
||||
description: 编排 reader 项目的端到端 AI 日报流程。仅在用户要求运行/重跑日报、汇报候选、发布 Hugo 日报、沉淀选中文章或继续已有日报任务时使用;覆盖异步 MCP 任务、候选确认、发布、单篇摘要和 IMA 知识库上传。不要因验证、排障冲动或候选质量不佳自行重跑。
|
||||
---
|
||||
|
||||
# Reader Digest Flow
|
||||
|
||||
## 职责边界
|
||||
|
||||
本 Skill 负责:
|
||||
|
||||
- 通过 reader MCP 启动、观察和恢复日报任务;
|
||||
- 向用户展示候选并保持稳定编号;
|
||||
- 根据用户选择生成并发布 Hugo 日报;
|
||||
- 对用户选中的文章生成知识笔记并编排 IMA 上传;
|
||||
- 在每个副作用边界执行确认和结果验证。
|
||||
|
||||
本 Skill 不负责:
|
||||
|
||||
- 实现 reader 内部抓取、摘要、过滤或恢复逻辑;
|
||||
- 通过手拼目录推导 Run 状态或 Artifact;
|
||||
- 未经用户要求自行重跑 Pipeline;
|
||||
- 未经用户确认发布日报或写入知识库;
|
||||
- 直接维护关键词配置;关键词治理委托给 `keyword-cleanup-review`。
|
||||
|
||||
## 核心规则
|
||||
|
||||
1. **只按用户指令运行。** 只有用户明确要求“跑日报”“重新跑”“再跑一次”时才启动新 Pipeline。验证、解释排序和排障默认读取已有 Run。
|
||||
2. **一次对话绑定一个当前 Run。** 以异步 Job 结果返回的 `run_id` 为稳定句柄;新 Run 产生新的候选编号体系,不混用历史编号。
|
||||
3. **状态以 MCP 返回为准。** Agent 只根据顶层 `status` 和 `recommended_action` 分支;`status_source`、`state_conflict` 仅用于解释。
|
||||
4. **路径以返回值为准。** 使用 `output_dir`、`artifact.path`、`delivery_output`、`report_output` 和 `written_paths`;不要根据 `run_id` 手拼 `outputs/...`。
|
||||
5. **候选编号保持稳定。** 用户编号永远对应当前候选列表的原始顺序(1-based);跨产物读取详情时按 URL 或完整 `item_id` 关联,不按数组位置关联。
|
||||
6. **副作用必须授权。** 用户确认 Hugo 文章后才能发布;用户确认 IMA 文章后才能生成并上传知识笔记。
|
||||
7. **内容必须有来源。** 日报和知识笔记只能基于当前 Run 的 `article.plain_text`、摘要、highlights 等 Artifact;不得使用通用知识补写原文没有的信息,也不为满足长度而扩写。
|
||||
|
||||
## 默认生产参数
|
||||
|
||||
用户未显式覆盖时使用:
|
||||
|
||||
```json
|
||||
{
|
||||
"limit": 7,
|
||||
"include_read": false,
|
||||
"mark_read": true,
|
||||
"debug_artifacts": false,
|
||||
"timeout_seconds": 60,
|
||||
"max_retries": 2
|
||||
}
|
||||
```
|
||||
|
||||
- 默认只处理未读文章。
|
||||
- `include_read=true` 仅在用户明确要求扩大到已读内容时使用。
|
||||
- debug/test/validation 才允许 `mark_read=false` 或 `debug_artifacts=true`。
|
||||
- 不随机生成文章数量;用户指定 `limit` 时按用户值执行。
|
||||
|
||||
## 正式流程
|
||||
|
||||
### Phase 1:启动并观察日报 Job
|
||||
|
||||
正式生产入口统一为异步 MCP:
|
||||
|
||||
1. 调用 `start_freshrss_pipeline_job`;
|
||||
2. 轮询 `get_freshrss_pipeline_job_status`;
|
||||
3. `status=success` 后调用 `get_freshrss_pipeline_job_result`;
|
||||
4. 保存返回的 `run_id` 和 Artifact 路径;
|
||||
5. 使用 `get_run_status`、`get_delivery_payload`、`get_run_report` 读取业务状态和结果。
|
||||
|
||||
失败时:
|
||||
|
||||
1. 使用 `get_run_status(run_id)` 读取关联 Run;
|
||||
2. 调用 `inspect_resume_plan(run_id)`;
|
||||
3. `recommended_action=resume` 时启动并轮询异步 Resume Job;
|
||||
4. `recommended_action=read_terminal_result` 时直接读取已有终态结果;
|
||||
5. `recommended_action=start_new_run` 时停止并向用户报告,不自行新建 Run。
|
||||
|
||||
CLI 仅用于 MCP 不可用时的 fallback、debug 或人工排障,不是默认生产入口。具体调用序列见 `references/flow.md`。
|
||||
|
||||
### Phase 2:汇报候选
|
||||
|
||||
- 使用当前 Run 返回的 Delivery Payload 或 digest brief Artifact;
|
||||
- 按候选原始顺序从 1 编号,状态可显示为“已入选/待确认”,但不得重新分组编号;
|
||||
- 每篇提供标题、来源、2-3 句摘要和筛选理由,避免原始 JSON dump;
|
||||
- 用户质疑编号或排序时读取当前 Run 产物核对,不重新运行 Pipeline;
|
||||
- 需要跨 Artifact 取详情时按 URL 或完整 `item_id` 交叉验证。
|
||||
|
||||
Feishu 输出不要使用 Markdown 表格,见 `references/feishu-format-notes.md`。
|
||||
|
||||
### Phase 3:等待 Hugo 选择
|
||||
|
||||
- 等待用户明确选择要发布的文章;
|
||||
- 用户编号映射到当前候选列表,不映射到 extracted 文件序号;
|
||||
- 用户拒绝发布时立即停止当日日报后续流程,不劝说、不自动换一批;
|
||||
- 用户明确要求重跑时才创建新 Run,并重新建立编号体系。
|
||||
|
||||
### Phase 4:生成并发布 Hugo 日报
|
||||
|
||||
发布前读取 `references/public-digest-example.md`,按其最终页面结构生成:
|
||||
|
||||
- `今日概览`
|
||||
- `今日重点`
|
||||
- `趋势观察`
|
||||
|
||||
每篇 `今日重点` 文章末尾必须添加 `来源:[来源名](原文 URL)`,来源链接跟随对应文章,不再生成独立的 `延伸阅读` 章节或重复链接。
|
||||
|
||||
仅发布用户在 Phase 3 选中的文章。公开页面不得出现 `keep/review/drop`、候选、待确认等内部状态。
|
||||
|
||||
写入 Hugo 后执行部署,并验证首页、日报列表页和当日详情页均可访问。命令和检查项见 `references/flow.md`。
|
||||
|
||||
### Phase 5:等待 IMA 选择
|
||||
|
||||
Hugo 发布选择与知识沉淀选择相互独立。询问用户哪些文章值得长期保存:
|
||||
|
||||
- 仅处理用户明确选择的文章;
|
||||
- 不把整份日报上传到 IMA;
|
||||
- 本次选择本身即授权后续单篇摘要和 IMA 上传,不重复确认。
|
||||
|
||||
### Phase 6:生成单篇知识笔记
|
||||
|
||||
对每篇选中文章:
|
||||
|
||||
1. 通过 URL/完整 `item_id` 找到对应 extracted Artifact;
|
||||
2. 使用 `start_article_summary_job(extracted_path=<返回的 Artifact 路径>, selected_ids=[<完整 item_id>])`;
|
||||
3. 轮询 `get_article_summary_job_status`,成功后读取 `get_article_summary_job_result`;
|
||||
4. 只使用现有 `article.plain_text`,不重新抓取原 URL;
|
||||
5. 使用返回的 `written_paths` 定位结果并检查 Markdown 内容。
|
||||
6. **`extracted_path` 必须传绝对路径**(前缀 `/home/ubuntu/zhu/github/reader/`):summary-mcp 工作目录是 `/root/.hermes`,相对路径会报 `extracted_path does not exist`。同一工具连续 3 次失败会触发 MCP 冷却(约 45-60s,报 `MCP server 'reader' is unreachable`),等待冷却后再重试,不要循环重试同一调用。
|
||||
|
||||
异步 MCP 不可用时才使用项目 CLI fallback。不要因为内容较短而引入原文之外的知识。
|
||||
|
||||
### Phase 7:上传到 IMA
|
||||
|
||||
上传前按需读取:
|
||||
|
||||
- 格式规则:`references/ima-format-quickref.md`
|
||||
- API 和上传步骤:`references/ima-upload-api.md`(含 `-200` 版本拦截修复)
|
||||
- 凭证定位:`references/ima-credential-chain.md`
|
||||
- 批量上传脚本:`scripts/ima_upload_one.py`(Python 编排,规避中文文件名 bash 引号问题)
|
||||
|
||||
硬规则:
|
||||
|
||||
- 使用完整文章标题作为文件名;
|
||||
- 以 Markdown 文件 `media_type=7` 上传到 `daily` knowledge base;
|
||||
- 保留原文 URL 和 Category,保证来源可追踪;
|
||||
- 不使用 URL 导入或 Notes 类型替代知识库文件;
|
||||
- 上传后验证目标知识库中存在对应条目;
|
||||
- 失败时报告具体阶段,不无限重试。
|
||||
|
||||
## 关键词治理路由
|
||||
|
||||
只有用户明确要求“清理关键词”“词库治理”等操作时才触发关键词治理。Review bundle 和 Suggestions 生成委托给 `keyword-cleanup-review`,确认与 Apply 仍由当前编排层负责:
|
||||
|
||||
1. 生成 review bundle;
|
||||
2. 生成规则层与可选语义层 Suggestions JSON;
|
||||
3. 等待人工确认;
|
||||
4. dry-run 后按精确 accept 参数 Apply。
|
||||
|
||||
`reader-digest-flow` 不直接编辑 `term_aliases.json`、`term_stopwords.json` 或兴趣配置。简要路由见 `references/keyword-engine-maintenance.md`。
|
||||
|
||||
## 环境坑位(本机部署)
|
||||
|
||||
- **MCP 相对路径陷阱**:`summary-mcp` 服务工作目录是 `/root/.hermes`,不是 reader 项目根。MCP 返回的 `output_dir`/`artifact.path` 是相对路径,直接传给 `start_article_summary_job(extracted_path=...)` 会报 `extracted_path does not exist`。传入前必须拼绝对路径前缀 `/home/ubuntu/zhu/github/reader/`。
|
||||
- **提取失败不等于运行失败**:`status_counts.extract_failed` 的条目(`CONTENT_EXTRACTION_FAILED`,`retryable=false`)跳过即可并如实汇报;失败文章常是推广/活动等低价值内容,不因此自行重跑。`linked_run_status=partial` 时先读 run-report 的 item 级 `error` 确认原因。
|
||||
- **用户要求"重新跑一批"**:候选质量低(用户主动提出)时重跑,应 `include_read=true` 并调高 `limit`(如 10),否则默认 `include_read=false` 会拉回同一批未读文章。重跑是新 Run,候选编号体系重新建立,汇报时提醒用户按新列表选择。
|
||||
|
||||
## 停止与人工介入
|
||||
|
||||
出现以下任一情况时停止自动流程并报告:
|
||||
|
||||
- 用户没有授权运行、发布或知识库写入;
|
||||
- Job/Run 返回不可恢复,或连续恢复失败;
|
||||
- Payload、候选 ID 或 Artifact 之间无法可靠关联;
|
||||
- 生成内容缺少可追踪来源;
|
||||
- Hugo 部署验证失败;
|
||||
- IMA 凭证、目标知识库或上传结果无法验证。
|
||||
|
||||
## Reference 路由
|
||||
|
||||
- `references/flow.md`:具体 MCP 调用序列、候选映射(含 extracted_path 绝对路径、候选≠文件名顺序)、Hugo 发布和 IMA 主步骤。
|
||||
- `references/content-extraction.md`:FreshRSS 内容来源与 `plain_text` 质量判断。
|
||||
- `references/public-digest-example.md`:可直接参考的 Hugo 最终页面结构。
|
||||
- `references/feishu-format-notes.md`:Feishu 输出格式限制。
|
||||
- `references/ima-format-quickref.md`:IMA Markdown 格式规则。
|
||||
- `references/ima-upload-api.md`:IMA Markdown 文件上传 API(含 `-200` 版本拦截修复)。
|
||||
- `references/ima-credential-chain.md`:IMA 凭证与知识库配置定位。
|
||||
- `references/keyword-engine-maintenance.md`:关键词治理 Skill 路由。
|
||||
- `scripts/ima_upload_one.py`:单篇 Markdown 上传 daily 知识库的完整 Python 脚本(preflight→重名→create_media→COS→add_knowledge)。
|
||||
@@ -0,0 +1,45 @@
|
||||
# 内容提取流程
|
||||
|
||||
本文说明处理流水线如何把 FreshRSS 条目转换为可供摘要使用的文章文本。
|
||||
|
||||
## 核心规则:FreshRSS 条目不重新抓取原文 URL
|
||||
|
||||
**FreshRSS 是仅提供 RSS 内容的上游。** 对于 FreshRSS 条目,流水线不会向文章原始 URL 发起 HTTP 请求。该行为由 `pipeline.py` 中的 `RSS_ONLY_UPSTREAMS = {"freshrss"}` 强制保证。
|
||||
|
||||
唯一例外是非 FreshRSS 上游。未来未设置 `upstream: freshrss` 的其他来源,可以在必要时使用 `fetch_html()` 作为回退。
|
||||
|
||||
## 内容来源优先级
|
||||
|
||||
`content_loader.py` 按以下顺序检查内容,并使用第一个包含 **至少 500 个可读字符** 的来源:
|
||||
|
||||
| 优先级 | 来源 | 含义 |
|
||||
|--------|------|------|
|
||||
| 1 | `raw_html` | 通过 `ExtractionInput.raw_html` 预先注入的 HTML;常规 FreshRSS 运行中很少使用。 |
|
||||
| 2 | `item.raw_content` | RSS `<content:encoded>` 中的文章正文;部分订阅源提供,部分不提供。 |
|
||||
| 3 | `item.raw_summary` | RSS `<description>` 中的摘要或片段;这是当前运行中最常见的来源。 |
|
||||
| 4 | `rss_content` | 来自非条目字段的独立 RSS 内容。 |
|
||||
| — | `none` | 没有可用内容;FreshRSS 不允许回源抓取,因此抛出 `RSS_CONTENT_MISSING`。 |
|
||||
|
||||
## `content_source` 与文本质量的关系
|
||||
|
||||
每个 `item-XX.extracted.json` 中的 `content_source` 字段表示流水线实际使用的内容来源:
|
||||
|
||||
- **`item.raw_content`**:RSS `<content:encoded>` 提供的文章正文,通常质量最好,接近直接阅读原文。
|
||||
- **`item.raw_summary`**:只有 RSS 摘要或描述,并非完整正文。不同来源长度差异较大,通常为 300-2000 个字符;AI 摘要基于该片段,而不是完整文章。
|
||||
- **`rss_content`**:来自独立 RSS 内容,质量取决于订阅源。
|
||||
- **`fetched_html`**:从原始 URL 抓取的 HTML。FreshRSS 条目不会出现该来源,只适用于非 FreshRSS 上游。
|
||||
|
||||
## 对日报质量的影响
|
||||
|
||||
如果提取结果文件中出现 `content_source: item.raw_summary`,说明 AI 使用的是订阅源摘要或片段,而不是完整正文。日报内容显得较浅时,原因可能只是 RSS 描述过短。
|
||||
|
||||
提高质量可以选择提供完整 `<content:encoded>` 的订阅源,或者把内容来源切换到支持全文 RSS 的系统,例如具备全文提取能力的 RSS 代理或 FiveFilters 等服务。
|
||||
|
||||
## 快速检查
|
||||
|
||||
先调用 `list_run_artifacts(run_id)`,再读取返回的提取结果产物路径,检查 `content_source` 和 `article.plain_text`。不要按 `run_id` 或“最新目录”手拼路径。
|
||||
|
||||
## 相关代码路径
|
||||
|
||||
- `src/summary_mcp/core/pipeline.py`:`RSS_ONLY_UPSTREAMS`、`_should_skip_fetch()`、`extract_content()`。
|
||||
- `src/summary_mcp/core/content_loader.py`:`choose_inline_content()` 的优先级链和 `fetch_html()`;FreshRSS 条目不会调用后者。
|
||||
@@ -0,0 +1,17 @@
|
||||
# 飞书 Markdown 格式说明
|
||||
|
||||
## 背景
|
||||
|
||||
Hermes 的飞书网关(`gateway/platforms/feishu.py`)通过 `_build_outbound_payload` 发送消息。该方法会检查内容中的 Markdown 特征,并据此决定消息类型:
|
||||
|
||||
- 内容匹配 `_MARKDOWN_HINT_RE`(加粗、列表、代码、链接等)时,使用包含 `md` 元素的飞书 `post` 类型发送,可以正常渲染。
|
||||
- 内容匹配 `_MARKDOWN_TABLE_RE`(Markdown 表头和分隔行)时,整条消息会被强制转换为 `text` 类型,即纯文本,不再渲染 Markdown。
|
||||
|
||||
原因是 `_build_markdown_post_payload` 会把内容包装为 `{"tag": "md", "text": "..."}` 元素,而飞书的 `md` 元素不支持表格,也没有把 Markdown 表格转换为飞书原生表格的逻辑。
|
||||
|
||||
## 飞书输出规则
|
||||
|
||||
- 通过飞书发送的消息不得使用 Markdown 表格;消息中只要出现一个表格,整条消息就会退化为纯文本。
|
||||
- 需要表达结构化信息时,优先使用分点列表、带标题的分节或行内格式。
|
||||
- 加粗(`**加粗**`)、行内代码(`` `代码` ``)、无序列表(`- 项目`)、有序列表(`1. 项目`)和链接均可正常使用。
|
||||
- 围栏式代码块可以使用,但代码块后的尾随内容可能存在渲染边界问题。
|
||||
@@ -0,0 +1,177 @@
|
||||
# Reader Digest Flow 操作参考
|
||||
|
||||
本文件承载 `reader-digest-flow` 的具体执行步骤。正式行为边界以 `../SKILL.md` 为准。
|
||||
|
||||
## 1. 日报 Pipeline
|
||||
|
||||
### 默认参数
|
||||
|
||||
```json
|
||||
{
|
||||
"limit": 7,
|
||||
"include_read": false,
|
||||
"mark_read": true,
|
||||
"debug_artifacts": false,
|
||||
"timeout_seconds": 60,
|
||||
"max_retries": 2
|
||||
}
|
||||
```
|
||||
|
||||
### 正式调用序列
|
||||
|
||||
```text
|
||||
start_freshrss_pipeline_job
|
||||
→ get_freshrss_pipeline_job_status
|
||||
→ get_freshrss_pipeline_job_result
|
||||
→ get_run_status
|
||||
→ get_delivery_payload / get_run_report
|
||||
```
|
||||
|
||||
状态动作:
|
||||
|
||||
- `running`:按合理间隔继续轮询;
|
||||
- `success`:读取结果,保存 `run_id`;
|
||||
- `failed`:读取关联 Run 并执行 Resume Plan;
|
||||
- `partial`:优先检查 Run Report、Artifact 和 Recovery 信息。
|
||||
|
||||
恢复序列:
|
||||
|
||||
```text
|
||||
inspect_resume_plan
|
||||
→ recommended_action=resume
|
||||
→ start_resume_job
|
||||
→ get_resume_job_status
|
||||
→ get_resume_job_result
|
||||
```
|
||||
|
||||
如果建议为 `read_terminal_result`,直接读取结果;如果为 `start_new_run`,停止并等待用户决定。
|
||||
|
||||
不要根据 `run_id` 拼目录。后续只消费调用返回的 `output_dir`、`artifact.path`、`delivery_output`、`report_output`。
|
||||
|
||||
## 2. 候选汇报与选择
|
||||
|
||||
优先使用当前 Run 的 digest brief Artifact;不可用时使用 `get_delivery_payload` 返回的候选。
|
||||
|
||||
展示规则:
|
||||
|
||||
1. 使用候选数组原始顺序并从 1 编号;
|
||||
2. 不因 `keep/review` 分组而重新编号;
|
||||
3. 每篇展示标题、来源、摘要和判断理由;
|
||||
4. 用户编号只映射当前候选数组;
|
||||
5. 跨 Artifact 读取详情时按 URL 或完整 `item_id` 匹配。
|
||||
|
||||
不要按 `summary-batch`、`item-XX` 和候选数组的相同位置推断它们是同一篇文章。
|
||||
|
||||
### extracted 文件与候选编号错位(实测 2026-07-31)
|
||||
|
||||
- `digest-brief.json` 的 `top_candidates` **没有 `item_id` 字段**,只有 `url`;
|
||||
- `extracted/item-XX.extracted.json` 的文件名序号与候选编号**可能不一致**(实例:候选2 = item-05、候选3 = item-02);
|
||||
- 正确做法:用 **URL 交叉匹配**(归一化 `%3D`→`=` 后逐条比对),或用完整 `item_id`(从 candidate-batch.json 的 `items[i].item_id` 按候选数组顺序取)在 extracted 文件里反查;两者都能验证时优先 item_id。
|
||||
|
||||
## 3. Hugo 日报
|
||||
|
||||
用户确认发布文章后:
|
||||
|
||||
1. 读取 `public-digest-example.md`;
|
||||
2. 仅使用用户选中的文章生成公开内容;
|
||||
3. 写入 Hugo 当日页面;
|
||||
4. 前台执行部署,避免把构建日志作为聊天通知;
|
||||
5. 验证首页、列表页和详情页。
|
||||
|
||||
当前部署位置:
|
||||
|
||||
```text
|
||||
/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md
|
||||
```
|
||||
|
||||
验证至少覆盖:
|
||||
|
||||
```text
|
||||
http://127.0.0.1:14322/
|
||||
http://127.0.0.1:14322/daily/
|
||||
http://127.0.0.1:14322/daily/YYYY-MM-DD/
|
||||
```
|
||||
|
||||
页面不存在时检查源文件、构建结果、容器状态和页面日期。需要 Docker/Hugo 深度排障时再读取部署项目自己的文档,不把历史事故规则复制回主 Skill。
|
||||
|
||||
## 4. 单篇知识笔记
|
||||
|
||||
用户确认 IMA 文章后:
|
||||
|
||||
1. 从候选中取得 URL 和完整 `item_id`;
|
||||
2. 从 Run Artifact 中找到匹配的 extracted 文件;
|
||||
3. 交叉验证 `article.item_id` 或 URL;
|
||||
4. 对每个单篇文件调用 `start_article_summary_job(extracted_path=<返回的 Artifact 路径>, selected_ids=[<完整 item_id>])`;
|
||||
5. `selected_ids` 传完整、非空、去除 `cand:` 前缀后的 `item_id`;
|
||||
6. 从 Job 结果的 `written_paths` 读取 Markdown。
|
||||
|
||||
### ⚠️ extracted_path 必须用绝对路径
|
||||
|
||||
`summary-mcp` 进程的工作目录是 `/root/.hermes`(不是 reader 项目根)。传相对路径(如 `outputs/freshrss/...`)会直接报 `extracted_path does not exist`。必须传绝对路径:
|
||||
|
||||
```text
|
||||
/home/ubuntu/zhu/github/reader/outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json
|
||||
```
|
||||
|
||||
### ⚠️ 候选编号 ≠ extracted 文件名顺序
|
||||
|
||||
候选数组顺序与 extracted 文件名(`item-01`…`item-07`)**不一定对齐**(实测候选2 落在 item-05)。`digest-brief.json` 的候选**没有 `item_id` 字段**,只有 URL。可靠匹配方法:
|
||||
|
||||
1. 从 `candidate-batch.json` 取每项完整 `item_id`(在 `candidate` 嵌套对象里,顶层 `item_key` 只是 `item-XX` 文件名序号);
|
||||
2. 或按 URL 匹配:归一化(`%3D`→`=`)后与每个 extracted 文件的 `article.url` / `article.canonical_url` 比对;
|
||||
3. 绝不要按候选位置对应 extracted 文件序号。
|
||||
|
||||
```python
|
||||
def norm(u): return u.replace('%3D','=').replace('%3d','=').strip()
|
||||
# 对每个 extracted 文件取 norm(article.url),与候选 norm(url) 精确比对
|
||||
```
|
||||
|
||||
正式序列:
|
||||
|
||||
```text
|
||||
start_article_summary_job
|
||||
→ get_article_summary_job_status
|
||||
→ get_article_summary_job_result
|
||||
```
|
||||
|
||||
内容依据是 extracted Artifact 中的 `article.plain_text`。不得重新抓取原 URL,不得使用原文之外的知识扩写。
|
||||
|
||||
CLI 仅在异步 MCP 不可用或人工排障时使用:
|
||||
|
||||
```bash
|
||||
python scripts/run_article_summaries.py \
|
||||
--extracted <returned-extracted-path> \
|
||||
--ids <full-item-id> \
|
||||
--output-dir <explicit-output-dir>
|
||||
```
|
||||
|
||||
## 5. IMA 上传
|
||||
|
||||
用户在知识沉淀阶段的文章选择即为上传授权。
|
||||
|
||||
执行顺序:
|
||||
|
||||
1. 检查生成的 Markdown 与来源;
|
||||
2. 文件名规范化为 `<完整文章标题>.md`;
|
||||
3. 确认目标为 `daily` knowledge base;
|
||||
4. 执行 preflight、create_media、COS upload、add_knowledge;
|
||||
5. 验证知识库条目存在。
|
||||
|
||||
上传格式与 API 参数分别见:
|
||||
|
||||
- `ima-format-quickref.md`
|
||||
- `ima-upload-api.md`
|
||||
- `ima-credential-chain.md`
|
||||
|
||||
失败时记录发生在 preflight、create_media、COS upload 或 add_knowledge 的具体阶段。不要把失败的 Markdown 改写成 URL 导入或 Notes 类型来绕过错误。
|
||||
|
||||
## 6. CLI fallback 原则
|
||||
|
||||
CLI 仅在以下场景使用:
|
||||
|
||||
- MCP 服务不可用;
|
||||
- Tool transport/launch 失败且无法取得有效 Job;
|
||||
- 用户明确要求本地调试;
|
||||
- 人工排障需要直接检查脚本输出。
|
||||
|
||||
如果已取得 `job_id` 或 `run_id`,先查询真实状态,避免因响应超时重复启动任务。
|
||||
@@ -0,0 +1,27 @@
|
||||
# IMA 凭证与安全边界
|
||||
|
||||
## 必需配置
|
||||
|
||||
- `IMA_OPENAPI_CLIENTID`
|
||||
- `IMA_OPENAPI_APIKEY`
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_ID`
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_NAME=daily`
|
||||
|
||||
优先使用当前进程环境和 IMA Skill 已支持的凭证加载机制。不要在本 Skill 中复制、迁移或重写密钥文件。
|
||||
|
||||
## 缺失处理
|
||||
|
||||
Preflight 返回凭证缺失或目标知识库无法解析时:
|
||||
|
||||
1. 停止上传;
|
||||
2. 只报告缺失的变量名或配置项;
|
||||
3. 等待用户或运行环境补齐配置;
|
||||
4. 配置恢复后重新执行 preflight,不重复生成知识笔记。
|
||||
|
||||
## 安全边界
|
||||
|
||||
- 不在聊天、日志或命令输出中打印完整 API key、KB ID 或 COS 临时凭证;
|
||||
- `create_media` 返回的 COS 凭证仅在同一受控进程内传给上传工具,不写入磁盘;
|
||||
- 不通过拼接 Shell 字符串传递凭证,使用参数数组或 IMA Skill 的封装;
|
||||
- 不绕过 Hermes 的脱敏机制;若现有工具链无法安全传递凭证,停止并报告;
|
||||
- 上传结束后不持久化 COS 临时凭证。
|
||||
@@ -0,0 +1,47 @@
|
||||
# IMA Markdown 格式速查
|
||||
|
||||
## 文件与标题
|
||||
|
||||
- 文件名:`<完整文章标题>.md`
|
||||
- `add_knowledge.title`:完整文章标题,不包含 `.md`
|
||||
- 上传类型:Markdown 文件,`media_type=7`
|
||||
- 目标:`daily` knowledge base
|
||||
|
||||
## 内容来源
|
||||
|
||||
只能使用当前 Run 的可追踪内容:
|
||||
|
||||
1. extracted Artifact 的 `article.plain_text`;
|
||||
2. 对应文章的结构化摘要;
|
||||
3. digest brief 的 summary 与 highlights。
|
||||
|
||||
不得使用通用知识补写原文没有的信息,不设置固定字数或字节数门槛。内容较短时保持简洁并忠于来源。
|
||||
|
||||
## 标准结构
|
||||
|
||||
```markdown
|
||||
# 完整文章标题
|
||||
|
||||
Source: https://原文链接
|
||||
Category: 分类
|
||||
|
||||
## 核心结论
|
||||
|
||||
## 主要论点
|
||||
|
||||
## 关键方法 / 机制
|
||||
|
||||
## 重要细节
|
||||
|
||||
## 可复用启发
|
||||
|
||||
## 关键词
|
||||
|
||||
## 主题
|
||||
```
|
||||
|
||||
- 核心结论和主要论点使用连贯段落;
|
||||
- 方法、细节和启发按完整知识点分项;
|
||||
- 没有来源支持的 Section 可以简写,不得编造内容填充。
|
||||
|
||||
上传 API 见 `ima-upload-api.md`。
|
||||
@@ -0,0 +1,126 @@
|
||||
# IMA Markdown 上传 API
|
||||
|
||||
用于将用户选中的单篇 Markdown 知识笔记上传到 `daily` knowledge base。
|
||||
|
||||
## 凭证
|
||||
|
||||
- `IMA_OPENAPI_CLIENTID`
|
||||
- `IMA_OPENAPI_APIKEY`
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_ID`
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_NAME`
|
||||
|
||||
凭证定位和恢复见 `ima-credential-chain.md`。不要在终端输出完整密钥。
|
||||
|
||||
## 上传前检查
|
||||
|
||||
- 文件名为 `<完整文章标题>.md`;
|
||||
- `title` 为完整文章标题,不带 `.md`;
|
||||
- Markdown 符合 `ima-format-quickref.md`;
|
||||
- 内容可追溯到当前 Run Artifact;
|
||||
- 用户已经明确选择该文章;
|
||||
- 目标知识库已经解析并验证。
|
||||
|
||||
## 1. Preflight
|
||||
|
||||
调用 IMA Skill 的 `preflight-check.cjs` 检查文件类型、扩展名、大小和 MIME。
|
||||
|
||||
预期:
|
||||
|
||||
```text
|
||||
file_ext=md
|
||||
content_type=text/markdown
|
||||
media_type=7
|
||||
```
|
||||
|
||||
### ⚠️ IMA skill 版本拦截(-200)
|
||||
|
||||
`ima_api.cjs` 每天首次调用会检查更新,若检测到新版(如 1.1.8 > 当前 1.1.7)会以 `code=-200` 拦截原请求。注意:**官方 zip 包内的 `meta.json` 可能没同步版本号**(下载 1.1.8 zip 后 meta 仍写 1.1.7),所以光替换文件无法跳过拦截。
|
||||
|
||||
快速修复(脚本本身已是新版,只差版本号):
|
||||
|
||||
```bash
|
||||
cd /root/.hermes/skills/openclaw-imports/ima-skill && python3 -c "
|
||||
import json
|
||||
m = json.load(open('meta.json')); m['version'] = '1.1.8'
|
||||
json.dump(m, open('meta.json','w'), ensure_ascii=False, indent=2)
|
||||
"
|
||||
```
|
||||
|
||||
先用 `diff -rq` 对比 zip 与安装目录:若只有 `.DS_Store`/meta 差异,说明代码已是最新,直接改 meta.json 版本号即可;若脚本有实质差异才需要整体替换。
|
||||
|
||||
## 2. Create Media
|
||||
|
||||
```text
|
||||
POST /openapi/wiki/v1/create_media
|
||||
```
|
||||
|
||||
请求核心字段:
|
||||
|
||||
```json
|
||||
{
|
||||
"file_name": "<完整文章标题>.md",
|
||||
"file_size": 0,
|
||||
"content_type": "text/markdown",
|
||||
"knowledge_base_id": "<daily-kb-id>",
|
||||
"file_ext": "md"
|
||||
}
|
||||
```
|
||||
|
||||
保存返回的 `media_id` 和 `cos_credential`。COS 临时凭证只在进程内传递,不打印到聊天或日志。
|
||||
|
||||
## 3. COS Upload
|
||||
|
||||
使用 IMA Skill 提供的 `cos-upload.cjs`,通过参数数组调用并检查:
|
||||
|
||||
- 进程 `returncode`;
|
||||
- `stderr`;
|
||||
- HTTP 上传结果。
|
||||
|
||||
不要拼接包含凭证的 Shell 字符串,也不要把多条 JSON 响应重定向到同一个文件。
|
||||
|
||||
## 4. Add Knowledge
|
||||
|
||||
```text
|
||||
POST /openapi/wiki/v1/add_knowledge
|
||||
```
|
||||
|
||||
核心字段:
|
||||
|
||||
```json
|
||||
{
|
||||
"media_type": 7,
|
||||
"media_id": "<media-id>",
|
||||
"title": "<完整文章标题>",
|
||||
"knowledge_base_id": "<daily-kb-id>",
|
||||
"file_info": {
|
||||
"cos_key": "<cos-key>",
|
||||
"file_size": 0,
|
||||
"file_name": "<完整文章标题>.md"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## 5. 验证
|
||||
|
||||
上传成功不能只依据本地命令退出码。验证目标 knowledge base 中存在对应标题或返回对象,并向用户报告最终结果。
|
||||
|
||||
失败时明确报告发生在 Preflight、Create Media、COS Upload 或 Add Knowledge 的哪一步,不改用 URL 导入或 Notes 类型绕过错误。
|
||||
|
||||
## 6. 已知坑:IMA skill 版本拦截(-200)
|
||||
|
||||
`ima_api.cjs` 每天首次调用会检查远端版本,若发现新版本(如 1.1.8 > 1.1.7)会以 exit=1 + stderr `{"code":-200}` 拦截**所有** API 调用,原请求不发送。此前遇到过。
|
||||
|
||||
处理方式(不必整包替换):
|
||||
|
||||
1. 按 stderr 提示下载新版 zip(如 `https://app-dl.ima.qq.com/skills/ima-skills-1.1.8.zip`)并解压;
|
||||
2. 对比新旧 `ima_api.cjs` 的 md5——zip 内核心脚本常与本地一致,只是 `meta.json` 的 `version` 未同步(zip 内仍写 1.1.7);
|
||||
3. 若 `ima_api.cjs` 一致,只需把本地 `meta.json` 的 `version` 改为远端版本号即可跳过拦截,无需替换文件。
|
||||
|
||||
调用成功后再执行本文件前面的上传流程。
|
||||
|
||||
## 6. 版本拦截与批量上传实测(2026-07-31)
|
||||
|
||||
- **`-200` skill 更新拦截**:`ima_api.cjs` 每天首次调用检查版本,发现新版时以 code -200 退出并提示更新。下载 zip 后**先对比 `ima_api.cjs` 的 md5**——实测 zip 内脚本与已装版本完全一致,只是 `meta.json` 版本号未同步。此时只需把 `~/.hermes/skills/openclaw-imports/ima-skill/meta.json` 的 `version` 改为最新版即可跳过拦截,无需替换任何脚本。
|
||||
- **Python 脚本编排上传**比 bash 可靠:bash 拼接含中文文件名/凭证的 curl 易出错。用 `subprocess` 参数数组依次调 `preflight-check.cjs` → `ima_api.cjs check_repeated_names` → `create_media` → `cos-upload.cjs`(`--secret-id/--secret-key/--token` 走参数数组,不打印)→ `add_knowledge`,每步解析返回 JSON,失败即停。
|
||||
- **批量上传**:4 篇逐个跑同一脚本即可;同名文件先 `check_repeated_names` 确认无重复。
|
||||
- 凭证从 `IMA_OPENAPI_CLIENTID` / `IMA_OPENAPI_APIKEY` 环境变量读取(`ima_api.cjs` 自动加载),KB ID 用 `IMA_DAILY_KNOWLEDGE_BASE_ID`。
|
||||
@@ -0,0 +1,29 @@
|
||||
# 关键词治理路由
|
||||
|
||||
关键词治理不属于 `reader-digest-flow` 的日常执行阶段。
|
||||
|
||||
仅当用户明确要求“清理关键词”“词库治理”“生成关键词建议”时,委托:
|
||||
|
||||
```text
|
||||
skills/keyword-cleanup-review/SKILL.md
|
||||
```
|
||||
|
||||
Review 输入生成由该 Skill 定义,后续确认与 Apply 由当前编排层负责:
|
||||
|
||||
```text
|
||||
build review bundle
|
||||
→ generate rule suggestions
|
||||
→ optional semantic suggestions
|
||||
→ human review
|
||||
→ dry-run
|
||||
→ apply accepted suggestions
|
||||
```
|
||||
|
||||
约束:
|
||||
|
||||
- Suggestions JSON 是 Review 与 Apply 之间的正式契约;
|
||||
- LLM 语义建议不能自动 Apply;
|
||||
- 不直接编辑 aliases、stopwords、watchlist 或 interest 配置;
|
||||
- 不在日报主流程中因 tag 质量不佳自动触发治理。
|
||||
|
||||
Review bundle、Suggestions、Schema 和产物保留策略以 `keyword-cleanup-review` 为唯一事实来源;该 Skill 不直接 Apply 配置。
|
||||
@@ -0,0 +1,31 @@
|
||||
+++
|
||||
title = "AI 日报 · 示例"
|
||||
date = 2026-04-01T09:00:00+08:00
|
||||
summary = "围绕 Agent 架构分层、Skills 标准化与工程实践的当日观察。"
|
||||
+++
|
||||
|
||||
# 今日概览
|
||||
|
||||
今天的公开内容主要集中在 AI Agent 架构演进、工具化落地与工程实践三条线索。行业关注点正在从“模型能做什么”转向“系统如何稳定落地并持续复用”。
|
||||
|
||||
## 今日重点
|
||||
|
||||
### 1. 从 Agent 到 Skills:AI 智能体架构的范式转变
|
||||
|
||||
文章分析了 AI 智能体从单体 Agent 向模块化 Skills 的演进,并结合 MCP、Skills 和真实项目说明能力分层与复用方式。
|
||||
|
||||
值得关注:
|
||||
|
||||
- Skills 将领域流程从 Agent 主体中拆出,便于复用和维护。
|
||||
- MCP 为 Agent 与外部工具提供标准化连接方式。
|
||||
- 工程竞争点逐渐从模型调用转向状态、工具和工作流设计。
|
||||
|
||||
这篇内容值得关注的原因在于,它把开放协议、分层架构和真实落地案例连接成了完整论证链。
|
||||
|
||||
来源:[示例来源](https://example.com/a)
|
||||
|
||||
## 趋势观察
|
||||
|
||||
1. Agent 正在从单体能力转向可组合的模块化体系。
|
||||
2. 工具契约、状态管理和验证机制正在成为 AI 应用的核心工程能力。
|
||||
3. Human-in-the-loop 仍是控制高风险副作用的重要边界。
|
||||
@@ -0,0 +1,96 @@
|
||||
#!/usr/bin/env python3
|
||||
"""IMA 上传单篇 Markdown 知识笔记到 daily 知识库。
|
||||
用法: python3 ima_upload_one.py "<绝对路径/summary.md>" "<完整文章标题>"
|
||||
依赖环境变量: IMA_OPENAPI_CLIENTID / IMA_OPENAPI_APIKEY / IMA_DAILY_KNOWLEDGE_BASE_ID
|
||||
流程: preflight -> check_repeated_names -> create_media -> cos-upload -> add_knowledge
|
||||
说明: 用 Python 而非 bash 编排,避免中文文件名/引号转义问题。
|
||||
退出码 2 = 文件名重复(需与用户确认保留双方或取消),非 0 均为失败。
|
||||
"""
|
||||
import json, os, subprocess, sys
|
||||
|
||||
SKILL_DIR = "/root/.hermes/skills/openclaw-imports/ima-skill"
|
||||
IMA_API = os.path.join(SKILL_DIR, "ima_api.cjs")
|
||||
COS_UPLOAD = os.path.join(SKILL_DIR, "knowledge-base/scripts/cos-upload.cjs")
|
||||
PREFLIGHT = os.path.join(SKILL_DIR, "knowledge-base/scripts/preflight-check.cjs")
|
||||
|
||||
def run_node(script, args):
|
||||
r = subprocess.run(["node", script] + args, capture_output=True, text=True, timeout=120)
|
||||
if r.returncode != 0:
|
||||
raise RuntimeError(f"{script} exit={r.returncode} stderr={r.stderr[:500]}")
|
||||
return json.loads(r.stdout)
|
||||
|
||||
def ima_api(api_path, body):
|
||||
r = subprocess.run(["node", IMA_API, api_path, json.dumps(body, ensure_ascii=False)],
|
||||
capture_output=True, text=True, timeout=120)
|
||||
if r.returncode != 0:
|
||||
raise RuntimeError(f"ima_api {api_path} exit={r.returncode} stderr={r.stderr[:500]}")
|
||||
resp = json.loads(r.stdout)
|
||||
if resp.get("code") != 0:
|
||||
raise RuntimeError(f"ima_api {api_path} code={resp.get('code')} msg={resp.get('msg')}")
|
||||
return resp.get("data", {})
|
||||
|
||||
def main():
|
||||
kb_id = os.environ["IMA_DAILY_KNOWLEDGE_BASE_ID"]
|
||||
file_path = sys.argv[1]
|
||||
title = sys.argv[2] # 完整文章标题(不含 .md)
|
||||
|
||||
pf = run_node(PREFLIGHT, ["--file", file_path])
|
||||
if not pf.get("pass"):
|
||||
raise RuntimeError(f"preflight failed: {pf}")
|
||||
file_name = pf["file_name"]; media_type = pf["media_type"]
|
||||
content_type = pf["content_type"]; file_size = pf["file_size"]; file_ext = pf["file_ext"]
|
||||
print(f"[preflight] ok file={file_name} ext={file_ext} size={file_size} media_type={media_type}")
|
||||
|
||||
dup = ima_api("openapi/wiki/v1/check_repeated_names", {
|
||||
"params": [{"name": file_name, "media_type": media_type}],
|
||||
"knowledge_base_id": kb_id
|
||||
})
|
||||
is_rep = dup.get("results", [{}])[0].get("is_repeated", False) if dup.get("results") else False
|
||||
if is_rep:
|
||||
print(f"[check_repeated_names] REPEATED: {file_name} — 需要处理")
|
||||
sys.exit(2)
|
||||
print("[check_repeated_names] no duplicate")
|
||||
|
||||
cm = ima_api("openapi/wiki/v1/create_media", {
|
||||
"file_name": file_name,
|
||||
"file_size": file_size,
|
||||
"content_type": content_type,
|
||||
"knowledge_base_id": kb_id,
|
||||
"file_ext": file_ext
|
||||
})
|
||||
media_id = cm["media_id"]
|
||||
cos = cm["cos_credential"]
|
||||
print(f"[create_media] media_id={media_id} cos_bucket={cos.get('bucket_name')} cos_key={cos.get('cos_key','')[:40]}")
|
||||
|
||||
r = subprocess.run(["node", COS_UPLOAD,
|
||||
"--file", file_path,
|
||||
"--secret-id", cos["secret_id"],
|
||||
"--secret-key", cos["secret_key"],
|
||||
"--token", cos["token"],
|
||||
"--bucket", cos["bucket_name"],
|
||||
"--region", cos["region"],
|
||||
"--cos-key", cos["cos_key"],
|
||||
"--content-type", content_type,
|
||||
"--start-time", str(cos.get("start_time", "")),
|
||||
"--expired-time", str(cos.get("expired_time", "")),
|
||||
"--timeout", "300000"
|
||||
], capture_output=True, text=True, timeout=360)
|
||||
if r.returncode != 0:
|
||||
raise RuntimeError(f"cos-upload exit={r.returncode} stderr={r.stderr[:800]}")
|
||||
print(f"[cos-upload] ok rc=0 stdout={r.stdout.strip()[:200]}")
|
||||
|
||||
ak = ima_api("openapi/wiki/v1/add_knowledge", {
|
||||
"media_type": media_type,
|
||||
"media_id": media_id,
|
||||
"title": title,
|
||||
"knowledge_base_id": kb_id,
|
||||
"file_info": {
|
||||
"cos_key": cos["cos_key"],
|
||||
"file_size": file_size,
|
||||
"file_name": file_name
|
||||
}
|
||||
})
|
||||
print(f"[add_knowledge] ok media_id={ak.get('media_id') or media_id}")
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -6,8 +6,30 @@ from .freshrss_pipeline_jobs import (
|
||||
start_freshrss_pipeline_job,
|
||||
)
|
||||
from .query_service import get_delivery_payload, get_run_report, get_run_status, list_run_artifacts, list_runs
|
||||
from .resume_jobs import get_resume_job_result, get_resume_job_status, start_resume_job
|
||||
from .resume_service import inspect_resume_plan, resume_run
|
||||
|
||||
# NOTE: resume_jobs and resume_service are NOT eagerly imported here to avoid
|
||||
# a circular import chain:
|
||||
# workflows/freshrss_pipeline.py -> runtime -> resume_jobs -> resume_service
|
||||
# -> workflows/freshrss_pipeline.py (circular!)
|
||||
# They are lazy-loaded via __getattr__ when accessed as summary_mcp.runtime.*
|
||||
|
||||
|
||||
def __getattr__(name):
|
||||
import importlib
|
||||
|
||||
_LAZY = {
|
||||
"get_resume_job_result": ("resume_jobs", "get_resume_job_result"),
|
||||
"get_resume_job_status": ("resume_jobs", "get_resume_job_status"),
|
||||
"start_resume_job": ("resume_jobs", "start_resume_job"),
|
||||
"inspect_resume_plan": ("resume_service", "inspect_resume_plan"),
|
||||
"resume_run": ("resume_service", "resume_run"),
|
||||
}
|
||||
if name in _LAZY:
|
||||
mod_name, attr_name = _LAZY[name]
|
||||
mod = importlib.import_module(f".{mod_name}", __package__)
|
||||
return getattr(mod, attr_name)
|
||||
raise AttributeError(f"module {__name__!r} has no attribute {name!r}")
|
||||
|
||||
|
||||
__all__ = [
|
||||
"ArtifactRecord",
|
||||
|
||||
@@ -9,6 +9,8 @@ from pathlib import Path
|
||||
from typing import Any
|
||||
from uuid import uuid4
|
||||
|
||||
from dotenv import dotenv_values
|
||||
|
||||
from .run_store import RunStore
|
||||
from .query_service import _resolve_run_record
|
||||
|
||||
@@ -29,6 +31,25 @@ DEFAULT_STAGES = [
|
||||
]
|
||||
MIN_JOB_STALE_SECONDS = 30 * 60
|
||||
MAX_JOB_STALE_SECONDS = 6 * 60 * 60
|
||||
DEFAULT_DOTENV_PATH = REPO_ROOT / ".env"
|
||||
|
||||
|
||||
def _build_subprocess_env() -> dict[str, str]:
|
||||
"""Build an env dict for subprocess, merging parent env with .env values.
|
||||
|
||||
The subprocess inherits the Hermes MCP server's environment, but .env values
|
||||
may not be in os.environ at the time the subprocess is spawned. This function
|
||||
loads them from .env and merges them in so the child process sees all needed
|
||||
variables (LLM_API_KEY, LLM_MODEL, LLM_API_URL, FRESHRSS_*, etc.) directly
|
||||
in os.environ, avoiding any dotenv-loading timing issues inside the subprocess.
|
||||
"""
|
||||
env = os.environ.copy()
|
||||
if DEFAULT_DOTENV_PATH.exists():
|
||||
for key, value in dotenv_values(DEFAULT_DOTENV_PATH).items():
|
||||
if isinstance(key, str) and isinstance(value, str) and value:
|
||||
# Only set if not already present in parent env
|
||||
env.setdefault(key, value)
|
||||
return env
|
||||
|
||||
|
||||
def _now() -> datetime:
|
||||
@@ -218,6 +239,7 @@ def start_freshrss_pipeline_job(
|
||||
stdout=subprocess.DEVNULL,
|
||||
stderr=subprocess.DEVNULL,
|
||||
start_new_session=True,
|
||||
env=_build_subprocess_env(),
|
||||
)
|
||||
except Exception as exc:
|
||||
report_file = _write_job_report(
|
||||
|
||||
@@ -9,6 +9,8 @@ from pathlib import Path
|
||||
from typing import Any
|
||||
from uuid import uuid4
|
||||
|
||||
from dotenv import dotenv_values
|
||||
|
||||
from .query_service import _resolve_run_record
|
||||
from .resume_service import (
|
||||
SUPPORTED_RESUME_STAGES,
|
||||
@@ -38,6 +40,22 @@ DEFAULT_STAGES = [
|
||||
]
|
||||
MIN_JOB_STALE_SECONDS = 30 * 60
|
||||
MAX_JOB_STALE_SECONDS = 6 * 60 * 60
|
||||
DEFAULT_DOTENV_PATH = REPO_ROOT / ".env"
|
||||
|
||||
|
||||
def _build_subprocess_env() -> dict[str, str]:
|
||||
"""Build an env dict for subprocess, merging parent env with .env values.
|
||||
|
||||
Ensures the subprocess sees all needed variables (LLM_API_KEY, LLM_MODEL,
|
||||
LLM_API_URL, FRESHRSS_*, etc.) directly in os.environ, avoiding dotenv
|
||||
timing issues in the child process.
|
||||
"""
|
||||
env = os.environ.copy()
|
||||
if DEFAULT_DOTENV_PATH.exists():
|
||||
for key, value in dotenv_values(DEFAULT_DOTENV_PATH).items():
|
||||
if isinstance(key, str) and isinstance(value, str) and value:
|
||||
env.setdefault(key, value)
|
||||
return env
|
||||
|
||||
|
||||
def _now() -> datetime:
|
||||
@@ -276,6 +294,7 @@ def start_resume_job(*, run_id: str) -> dict[str, Any]:
|
||||
stdout=subprocess.DEVNULL,
|
||||
stderr=subprocess.DEVNULL,
|
||||
start_new_session=True,
|
||||
env=_build_subprocess_env(),
|
||||
)
|
||||
except Exception as exc:
|
||||
report_file = _write_job_report(
|
||||
|
||||
@@ -2,6 +2,7 @@ from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||
from datetime import date, datetime, timezone
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
@@ -192,7 +193,7 @@ def _build_candidate_batch_payload(*, run_id: str, item_contexts: list[dict[str,
|
||||
}
|
||||
|
||||
|
||||
def _persist_summary_batch_artifact(*, run_store: RunStore, run_dir: Path, item_contexts: list[dict[str, Any]]) -> Path:
|
||||
def _persist_summary_batch_artifact(*, run_store, run_dir: Path, item_contexts: list[dict[str, Any]]) -> Path:
|
||||
output_path = _summary_batch_output(run_dir)
|
||||
_save_json(
|
||||
output_path,
|
||||
@@ -202,7 +203,7 @@ def _persist_summary_batch_artifact(*, run_store: RunStore, run_dir: Path, item_
|
||||
return output_path
|
||||
|
||||
|
||||
def _persist_candidate_batch_artifact(*, run_store: RunStore, run_dir: Path, item_contexts: list[dict[str, Any]]) -> Path:
|
||||
def _persist_candidate_batch_artifact(*, run_store, run_dir: Path, item_contexts: list[dict[str, Any]]) -> Path:
|
||||
output_path = _candidate_batch_output(run_dir)
|
||||
_save_json(
|
||||
output_path,
|
||||
@@ -461,9 +462,12 @@ def run_freshrss_pipeline(
|
||||
summary_success_count = 0
|
||||
summary_failed_count = 0
|
||||
summary_candidates = [ctx for ctx in item_contexts if ctx["extraction"] is not None and ctx["extraction"].success]
|
||||
# Parallelize LLM summaries — I/O bound calls, independent per article
|
||||
with ThreadPoolExecutor(max_workers=min(len(summary_candidates) or 1, 4)) as pool:
|
||||
fut_map = {}
|
||||
for item_context in summary_candidates:
|
||||
item_report = item_context["item_report"]
|
||||
summary_exit_code, summary_payload, summary_report = run_loop_payload(
|
||||
fut = pool.submit(
|
||||
run_loop_payload,
|
||||
extracted_payload=item_context["extracted_payload"],
|
||||
prompt_path=resolved_prompt_path,
|
||||
output_path=item_context["summary_output"],
|
||||
@@ -473,6 +477,16 @@ def run_freshrss_pipeline(
|
||||
model=resolved_llm_model,
|
||||
api_url=resolved_llm_api_url,
|
||||
)
|
||||
fut_map[fut] = item_context
|
||||
|
||||
for fut in as_completed(fut_map):
|
||||
item_context = fut_map[fut]
|
||||
item_report = item_context["item_report"]
|
||||
try:
|
||||
summary_exit_code, summary_payload, summary_report = fut.result()
|
||||
except Exception as exc:
|
||||
summary_exit_code, summary_payload, summary_report = 1, None, None
|
||||
|
||||
if summary_exit_code != 0 or summary_payload is None:
|
||||
item_report["status"] = "summary_failed"
|
||||
if summary_report is not None:
|
||||
|
||||
Reference in New Issue
Block a user