Compare commits
59
Commits
9fd21c59f5
..
main
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
8e27cca166 | ||
|
|
807027976e | ||
|
|
6dd8cef347 | ||
|
|
5eb390e3ed | ||
|
|
7b791ac947 | ||
|
|
cdbcdcd485 | ||
|
|
590d050218 | ||
|
|
4399c9ca90 | ||
|
|
8136301ad4 | ||
|
|
b29cd8f934 | ||
|
|
4c4d1a45e6 | ||
|
|
b8727f1885 | ||
|
|
6705613aa4 | ||
|
|
416414ae1d | ||
|
|
f3e7488fc8 | ||
|
|
e3663f681d | ||
|
|
a06f2a1d08 | ||
|
|
c0194647d8 | ||
|
|
c528e0abc7 | ||
|
|
52ce6bfdf5 | ||
|
|
c58d8114cd | ||
|
|
10f9cd088f | ||
|
|
2053cccfef | ||
|
|
87d18e4263 | ||
|
|
2df0af5b82 | ||
|
|
4f219ef92f | ||
|
|
c622bc6247 | ||
|
|
d91cdbc6c9 | ||
|
|
4a02894c43 | ||
|
|
d9173fbe0f | ||
|
|
7563aa8fea | ||
|
|
72a6853c03 | ||
|
|
3a85d47f00 | ||
|
|
b5ec58cd6a | ||
|
|
5b51332075 | ||
|
|
f72c9ad2b1 | ||
|
|
660f03674b | ||
|
|
ca86705646 | ||
|
|
5b8df317ef | ||
|
|
e24db99baa | ||
|
|
1edbc2b736 | ||
|
|
9aa833d2c8 | ||
|
|
40fccbd0cc | ||
|
|
d3417e29e1 | ||
|
|
8575033528 | ||
|
|
ed0e3a9c7d | ||
|
|
79509fd1ba | ||
|
|
3d58408687 | ||
|
|
78039eff47 | ||
|
|
a2f6f8a28d | ||
|
|
2d9a38e465 | ||
|
|
ced6142b59 | ||
|
|
7bb9bef12f | ||
|
|
7dd0894a35 | ||
|
|
15bbbd9809 | ||
|
|
a111c3ddd7 | ||
|
|
9f7b777784 | ||
|
|
40e896fa01 | ||
|
|
41f7ac0e91 |
@@ -1,5 +0,0 @@
|
||||
{
|
||||
"enabledPlugins": {
|
||||
"skill-creator@claude-plugins-official": true
|
||||
}
|
||||
}
|
||||
@@ -1,5 +1,6 @@
|
||||
.claude/
|
||||
.codex/
|
||||
.idea/
|
||||
__pycache__/
|
||||
*.pyc
|
||||
*.egg-info/
|
||||
|
||||
@@ -1,260 +1,214 @@
|
||||
# Content Extract MCP
|
||||
# Reader · AI 日报引擎
|
||||
|
||||
Python MCP scaffold for article content extraction, structured summary validation, deterministic filtering, and Markdown sink output.
|
||||
> 从 FreshRSS 到 AI 日报的自动化流水线,为个人知识管理生成每日 AI 工程化简报。
|
||||
|
||||
## Run
|
||||
Reader 是一个端到端的 AI 日报生产系统,定时从自建 FreshRSS 的 RSS 订阅源拉取文章,经过内容提取、LLM 筛选与摘要、关键词索引构建,最终产出两个输出:
|
||||
|
||||
1. **公开日报** — 推送到 [Hugo 站点](https://osiman.site/daily/) 的精选技术简报
|
||||
2. **知识沉淀** — 单篇结构化摘要上传到 IMA 知识库(`daily` KB)
|
||||
|
||||
整个流程由 OpenClaw 编排,作为 MCP Workflow Service 对外暴露。
|
||||
|
||||
---
|
||||
|
||||
## ✨ 核心能力
|
||||
|
||||
| 能力 | 说明 |
|
||||
|:----|:------|
|
||||
| **RSS 拉取** | 从 FreshRSS API 拉取订阅文章,支持增量读取与已读标记 |
|
||||
| **内容提取** | 自动提取文章正文、标题、来源等结构化字段 |
|
||||
| **LLM 筛选** | 基于个人兴趣画像(`filter_context.personal.json`)自动评估文章质量,分为 keep / review / drop 三档 |
|
||||
| **LLM 摘要** | 并行生成每篇文章的结构化摘要(4 路并发,约 24 秒完成 7 篇) |
|
||||
| **关键词索引** | 自动构建每日关键词索引,支持别名映射与停用词过滤 |
|
||||
| **候选简报** | 生成 `digest-brief.json` 供编排层(OpenClaw)决策 |
|
||||
| **单篇沉淀** | 对选中的文章生成结构化知识笔记,上传到 IMA 知识库 |
|
||||
| **异步 Job** | 全部生产流程走异步 job,支持恢复与状态查询 |
|
||||
|
||||
---
|
||||
|
||||
## 🏗 架构概览
|
||||
|
||||
```
|
||||
FreshRSS ──→ 拉取 ──→ 内容提取 ──→ LLM 筛选 ──→ 关键词索引
|
||||
│
|
||||
digest-brief.json
|
||||
│
|
||||
┌──────────────┼──────────────┐
|
||||
▼ ▼ ▼
|
||||
Hugo 日报 IMA 知识库 term_index
|
||||
(公开简报) (单篇沉淀) (关键词数据)
|
||||
```
|
||||
|
||||
### MCP 工具层
|
||||
|
||||
Reader 通过 Hermes MCP 暴露 20+ 个工具,分为三类:
|
||||
|
||||
**日报流水线:**
|
||||
- `start_freshrss_pipeline_job` → 启动异步日报 Job
|
||||
- `get_freshrss_pipeline_job_status` / `get_freshrss_pipeline_job_result` → 轮询结果
|
||||
|
||||
**状态查询:**
|
||||
- `get_run_status` / `get_delivery_payload` / `get_run_report` → 读取运行结果
|
||||
- `list_runs` / `list_run_artifacts` → 浏览运行历史
|
||||
|
||||
**恢复与单篇总结:**
|
||||
- `inspect_resume_plan` / `start_resume_job` → 恢复失败 Job
|
||||
- `start_article_summary_job` / `generate_article_summaries` → 单篇文章摘要
|
||||
|
||||
### CLI 入口
|
||||
|
||||
同步入口,适合本地 debug / fallback:
|
||||
|
||||
```bash
|
||||
# 完整日报流水线
|
||||
python scripts/run_freshrss_pipeline.py --limit 7 --mark-read --timeout 300
|
||||
|
||||
# 单篇文章摘要
|
||||
python scripts/run_article_summaries.py \
|
||||
--extracted outputs/freshrss/rerun/<run_id>/extracted/item-01.extracted.json \
|
||||
--output-dir outputs/freshrss/single_summaries/YYYY-MM-DD
|
||||
|
||||
# 关键词维护
|
||||
python scripts/build_keyword_index.py
|
||||
python scripts/generate_term_cleanup_suggestions.py
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🚀 快速开始
|
||||
|
||||
### 环境变量
|
||||
|
||||
```
|
||||
# FreshRSS
|
||||
FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||||
FRESHRSS_USERNAME=bot
|
||||
FRESHRSS_API_PASSWORD=xxx
|
||||
|
||||
# LLM(主流水线)
|
||||
LLM_API_URL=https://api.deepseek.com
|
||||
LLM_API_KEY=xxx
|
||||
LLM_MODEL=deepseek-chat
|
||||
|
||||
# LLM(可选,单篇摘要独立模型)
|
||||
ARTICLE_SUMMARY_LLM_API_URL=https://api.deepseek.com
|
||||
ARTICLE_SUMMARY_LLM_API_KEY=xxx
|
||||
ARTICLE_SUMMARY_LLM_MODEL=deepseek-chat
|
||||
|
||||
# IMA 知识库(可选,仅沉淀时需要)
|
||||
IMA_DAILY_KNOWLEDGE_BASE_ID=xxx
|
||||
IMA_DAILY_KNOWLEDGE_BASE_NAME=daily
|
||||
```
|
||||
|
||||
### 运行
|
||||
|
||||
```bash
|
||||
# 安装
|
||||
pip install -e .
|
||||
|
||||
# 跑日报流水线(CLI 模式)
|
||||
python scripts/run_freshrss_pipeline.py --limit 7 --mark-read --timeout 300
|
||||
|
||||
# 启动 MCP 服务(OpenClaw 集成用)
|
||||
summary-mcp
|
||||
```
|
||||
|
||||
The server exposes four tools:
|
||||
---
|
||||
|
||||
- `extract_url_content`
|
||||
- `extract_item_content`
|
||||
- `filter_summary_result`
|
||||
- `run_freshrss_openclaw_pipeline`
|
||||
## 📁 项目结构
|
||||
|
||||
Article-summary post-processing (separate LLM optional):
|
||||
|
||||
- `article-summary` MCP tool (operates on existing extracted payloads)
|
||||
- `scripts/run_article_summaries.py` CLI helper
|
||||
|
||||
Validate an LLM summary result:
|
||||
|
||||
```bash
|
||||
validate-llm-result outputs/reference/summary/result.json --extracted outputs/reference/extracted/read-flow-2026.extracted.json
|
||||
```
|
||||
reader/
|
||||
├── configs/ # 配置
|
||||
│ ├── filter_context.personal.json # 个人兴趣画像
|
||||
│ ├── term_aliases.json # 关键词别名映射(149 条)
|
||||
│ ├── term_stopwords.json # 关键词停用词(132 条)
|
||||
│ └── term_cleanup_policy.json # 关键词清理策略
|
||||
├── src/
|
||||
│ └── summary_mcp/ # MCP 服务核心
|
||||
│ ├── server.py # MCP 服务入口
|
||||
│ ├── runtime/ # 运行时(Job 管理、状态持久化)
|
||||
│ └── workflows/ # 工作流(日报流水线逻辑)
|
||||
├── scripts/ # CLI 入口
|
||||
├── outputs/ # 运行时产出
|
||||
│ └── freshrss/
|
||||
│ ├── rerun/<run_id>/ # 每次运行的全量产物
|
||||
│ │ ├── candidates/ # digest-brief.json, delivery payload
|
||||
│ │ ├── extracted/ # item-XX.extracted.json
|
||||
│ │ └── run-state.json # 运行状态
|
||||
│ └── single_summaries/ # 单篇摘要输出
|
||||
├── data/
|
||||
│ └── term_index/ # 关键词索引数据
|
||||
│ ├── daily/YYYY-MM-DD.json
|
||||
│ └── term_stats.json
|
||||
├── docs/ # 设计文档
|
||||
└── prompts/ # LLM Prompt 模板
|
||||
```
|
||||
|
||||
Run the minimal extraction-to-summary loop:
|
||||
---
|
||||
|
||||
## ⚙️ 关键技术决策
|
||||
|
||||
| 决策 | 选择 | 原因 |
|
||||
|:----|:----|:------|
|
||||
| 运行模式 | **异步 Job** 为主,CLI fallback | 避免 MCP 传输层 120s 超时限制 |
|
||||
| 摘要并发 | **ThreadPoolExecutor(max_workers=4)** | LLM 调用是 I/O 密集型,4 路并行将 7 篇摘要从 2-3 分钟压到 ~24 秒 |
|
||||
| 环境变量 | **子进程显式注入 .env** | 解决 MCP 服务器环境隔离导致子进程读取不到 LLM_API_KEY 的问题 |
|
||||
| 关键词过滤 | **别名映射 + 停用词 + 语义清洗** | 先用 `term_aliases.json` 归一化,再用 `term_stopwords.json` 过滤噪声,最后通过 LLM 做语义级清洗 |
|
||||
| Tag 选择 | **复用已有通用 Tag**,不从 term_index 翻生僻词 | 保持 Hugo 站点 /tags/ 页面整洁,避免大量一次性专有名词 |
|
||||
|
||||
---
|
||||
|
||||
## 🔧 关键词治理
|
||||
|
||||
配置治理走四步流程(`scripts/` 下脚本):
|
||||
|
||||
```bash
|
||||
python scripts/run_summary_loop.py ^
|
||||
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||
--prompt outputs/prompts/llm-summary-prompt.txt ^
|
||||
--output outputs/reference/summary/result.loop.json
|
||||
# 1. 构建评审数据包
|
||||
python skills/keyword-cleanup-review/scripts/build_review_bundle.py --days 7 --top 50
|
||||
|
||||
# 2. 统计规则级建议(大小写、单复数、频次阈值)
|
||||
python scripts/generate_term_cleanup_suggestions.py
|
||||
|
||||
# 3. LLM 语义级建议(中英映射、简称-全称、近义词)
|
||||
python scripts/generate_term_cleanup_semantic_suggestions.py
|
||||
|
||||
# 4. 确认后写入配置
|
||||
python scripts/apply_term_suggestions.py --accept-watch ... --dry-run
|
||||
```
|
||||
|
||||
Pull FreshRSS entries and map them into normalized `item` objects:
|
||||
详见 `docs/design/keyword-engine-maintenance.md`。
|
||||
|
||||
```bash
|
||||
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||||
set FRESHRSS_USERNAME=bot
|
||||
set FRESHRSS_API_PASSWORD=your-api-password
|
||||
python scripts/pull_freshrss_items.py --limit 5 --mark-read
|
||||
```
|
||||
---
|
||||
|
||||
By default the script excludes entries already tagged as `read`. Add `--include-read` if you want the full reading list.
|
||||
When `--mark-read` is enabled, fetched entries are marked as read after the script finishes successfully.
|
||||
## 🤖 Agent Skill
|
||||
|
||||
The script writes:
|
||||
Reader 附带一个完整的 OpenClaw Agent Skill,位于 `skills/reader-digest-flow/`,供 AI Agent(Hermes / Claude Code 等)编排每日日报流程使用。
|
||||
|
||||
- `outputs/freshrss/raw/freshrss.raw.json`
|
||||
- `outputs/freshrss/items/freshrss.items.json`
|
||||
Skill 包含完整的 7 阶段工作流定义:
|
||||
1. **Phase 1** — 跑 Pipeline(FreshRSS → 提取 → LLM 筛选)
|
||||
2. **Phase 2** — 汇报候选(展示候选文章给用户决策)
|
||||
3. **Phase 3** — 用户选文(选择 Hugo 发布文章)
|
||||
4. **Phase 4** — 生成并发布 Hugo 日报
|
||||
5. **Phase 5** — 用户选 IMA 沉淀文章
|
||||
6. **Phase 6** — LLM 摘要生成
|
||||
7. **Phase 7** — IMA 知识库上传
|
||||
|
||||
Pull FreshRSS entries and run content extraction for each mapped item:
|
||||
以及海量铁律(不重跑 pipeline、编号规则、Tag 选择规范、IMA 上传流程等)和参考文件(`references/` 目录)。
|
||||
|
||||
```bash
|
||||
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||||
set FRESHRSS_USERNAME=osiman
|
||||
set FRESHRSS_API_PASSWORD=your-api-password
|
||||
python scripts/run_freshrss_extract.py --limit 1 --mark-read
|
||||
```
|
||||
---
|
||||
|
||||
By default the script excludes entries already tagged as `read`. When `--mark-read` is enabled, only entries with successful extraction are marked as read.
|
||||
## 📄 文档
|
||||
|
||||
The script writes:
|
||||
- `docs/openclaw/README.md` — OpenClaw 集成指南
|
||||
- `docs/openclaw/openclaw-orchestration-flow.md` — 编排流程
|
||||
- `docs/openclaw/openclaw-delivery-payload-spec.md` — 字段契约
|
||||
- `docs/design/README.md` — 设计文档总索引
|
||||
- `docs/design/filter-rule-engine-design.md` — 过滤规则引擎设计
|
||||
- `docs/design/filter-rule-engine-usage.md` — 过滤规则使用说明
|
||||
|
||||
- `outputs/freshrss/raw/freshrss.raw.json`
|
||||
- `outputs/freshrss/items/freshrss.items.json`
|
||||
- `outputs/freshrss/extracted/freshrss.extracted.json`
|
||||
---
|
||||
|
||||
Run the full FreshRSS pipeline and mark items as read only after the final OpenClaw delivery payload is written:
|
||||
## 📝 License
|
||||
|
||||
```bash
|
||||
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||||
set FRESHRSS_USERNAME=osiman
|
||||
set FRESHRSS_API_PASSWORD=your-api-password
|
||||
set LLM_API_URL=https://api.deepseek.com
|
||||
set LLM_API_KEY=your-llm-api-key
|
||||
set LLM_MODEL=deepseek-chat
|
||||
python scripts/run_freshrss_pipeline.py --limit 5 --mark-read
|
||||
```
|
||||
|
||||
If you want to apply your personal engineering and AI-agent interest profile during filtering, pass a context file:
|
||||
|
||||
```bash
|
||||
python scripts/run_freshrss_pipeline.py ^
|
||||
--limit 5 ^
|
||||
--context configs/filter_context.personal.json ^
|
||||
--mark-read
|
||||
```
|
||||
|
||||
This is the recommended production entrypoint. By default it writes only:
|
||||
|
||||
- `outputs/freshrss/rerun/<timestamp>/raw/freshrss.raw.json`
|
||||
- `outputs/freshrss/rerun/<timestamp>/candidates/openclaw-delivery-payload.json`
|
||||
- `outputs/freshrss/rerun/<timestamp>/run-report.json`
|
||||
- `outputs/freshrss/rerun/<timestamp>/extracted/item-XX.extracted.json` (one per item)
|
||||
|
||||
It also updates the daily keyword index runtime data:
|
||||
|
||||
- `data/term_index/daily/YYYY-MM-DD.json`
|
||||
- `data/term_index/term_stats.json`
|
||||
|
||||
The main pipeline does not emit a batch-level `freshrss.extracted.json` file by default.
|
||||
If you need additional per-item intermediates such as normalized items, summaries, filter decisions, candidate records, or candidate inputs, add `--debug-artifacts`.
|
||||
|
||||
When OpenClaw is connected to the MCP server, it should call `run_freshrss_openclaw_pipeline` for the same behavior directly through MCP. The tool also supports `debug_artifacts=true` when deeper inspection is needed.
|
||||
|
||||
Run deterministic filter rules against a structured summary result:
|
||||
|
||||
```bash
|
||||
python scripts/run_filter_rules.py ^
|
||||
--summary outputs/reference/summary/result.loop.json ^
|
||||
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||
--output outputs/reference/filter/filter-decision.json
|
||||
```
|
||||
|
||||
You can optionally pass a context file to inject interest topics or source tags:
|
||||
|
||||
```bash
|
||||
python scripts/run_filter_rules.py ^
|
||||
--summary outputs/reference/summary/result.loop.json ^
|
||||
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||
--context outputs/reference/filter/filter-context.json ^
|
||||
--output outputs/reference/filter/filter-decision.with-context.json
|
||||
```
|
||||
|
||||
Rule engine details and rule authoring guidance live in:
|
||||
|
||||
- `docs/design/filter-rule-engine-design.md`
|
||||
- `docs/design/filter-rule-engine-usage.md`
|
||||
|
||||
Write a filtered result into the Markdown sink:
|
||||
|
||||
```bash
|
||||
python scripts/run_markdown_sink.py ^
|
||||
--summary outputs/reference/summary/result.loop.json ^
|
||||
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||
--filter outputs/reference/filter/filter-decision.json
|
||||
```
|
||||
|
||||
The script writes markdown notes under `knowledge-base/`.
|
||||
|
||||
Build an internal `ArticleCandidateRecord` and a slim `OpenClawCandidateInput`:
|
||||
|
||||
```bash
|
||||
python scripts/run_article_candidate.py ^
|
||||
--summary outputs/reference/summary/result.loop.json ^
|
||||
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||
--filter outputs/reference/filter/filter-decision.json ^
|
||||
--section-hint tools_and_workflows
|
||||
```
|
||||
|
||||
The script writes by default:
|
||||
|
||||
- `outputs/reference/candidates/article-candidate-record.json`
|
||||
- `outputs/reference/candidates/openclaw-candidate-input.json`
|
||||
|
||||
Build a batch OpenClaw delivery payload:
|
||||
|
||||
```bash
|
||||
python scripts/build_openclaw_delivery.py ^
|
||||
--input-dir outputs/freshrss/candidates/batch ^
|
||||
--sort-by-rank ^
|
||||
--date 2026-03-25
|
||||
```
|
||||
|
||||
The script writes by default:
|
||||
|
||||
- `outputs/reference/candidates/openclaw-delivery-payload.json`
|
||||
|
||||
Output layout details live in `outputs/README.md`.
|
||||
|
||||
Keyword index defaults live in:
|
||||
|
||||
- `configs/term_aliases.json`
|
||||
- `configs/term_stopwords.json`
|
||||
- `configs/term_cleanup_policy.json`
|
||||
- `configs/term_watchlist.json`
|
||||
- `configs/term_change_log.json`
|
||||
|
||||
You can also rebuild the keyword index from an existing delivery payload:
|
||||
|
||||
```bash
|
||||
python scripts/build_keyword_index.py ^
|
||||
--input outputs/reference/candidates/openclaw-delivery-payload.json
|
||||
```
|
||||
|
||||
Runtime keyword data is stored under `data/term_index/`.
|
||||
|
||||
The keyword cleanup review skill lives in:
|
||||
|
||||
- `skills/keyword-cleanup-review/`
|
||||
|
||||
To build a review bundle for the LLM skill:
|
||||
|
||||
```bash
|
||||
python skills/keyword-cleanup-review/scripts/build_review_bundle.py ^
|
||||
--days 7 ^
|
||||
--top 50 ^
|
||||
--output outputs/term_index/review/keyword-cleanup-bundle.json
|
||||
```
|
||||
|
||||
The review bundle now also carries cleanup governance context:
|
||||
|
||||
- cleanup thresholds from `configs/term_cleanup_policy.json`
|
||||
- the current watch list from `configs/term_watchlist.json`
|
||||
- recent applied changes from `configs/term_change_log.json`
|
||||
|
||||
The skill only produces review inputs and suggestions. It does not modify `term_aliases`, `term_stopwords`, or `filter_context.personal.json` automatically.
|
||||
|
||||
To preview accepted suggestions before writing any config files:
|
||||
|
||||
```bash
|
||||
python scripts/apply_term_suggestions.py ^
|
||||
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json ^
|
||||
--accept-watch Cron Heartbeat Memory ^
|
||||
--dry-run
|
||||
```
|
||||
|
||||
Remove `--dry-run` to write the accepted changes. The script can also apply accepted `alias`, `stopword`, and `interest keyword` suggestions through `--accept-alias`, `--accept-stopword`, and `--accept-interest`. Accepted watch terms are written into `configs/term_watchlist.json`, and every applied action is appended into `configs/term_change_log.json`.
|
||||
|
||||
|
||||
## Article-summary LLM configuration
|
||||
|
||||
Set a dedicated model for post-processing summaries without affecting the main pipeline:
|
||||
|
||||
- `ARTICLE_SUMMARY_LLM_API_URL`
|
||||
- `ARTICLE_SUMMARY_LLM_MODEL`
|
||||
- `ARTICLE_SUMMARY_LLM_API_KEY`
|
||||
|
||||
If these are not set, the summarizer falls back to the main `LLM_*` / `OPENAI_*` settings used elsewhere.
|
||||
|
||||
Example (PowerShell style):
|
||||
|
||||
```bash
|
||||
set ARTICLE_SUMMARY_LLM_API_URL=https://api.deepseek.com
|
||||
set ARTICLE_SUMMARY_LLM_API_KEY=your-article-summary-key
|
||||
set ARTICLE_SUMMARY_LLM_MODEL=deepseek-chat
|
||||
```
|
||||
|
||||
Then run, for example, either via the CLI script:
|
||||
|
||||
```bash
|
||||
python scripts/run_article_summaries.py ^
|
||||
--extracted outputs/freshrss/extracted/freshrss.extracted.json ^
|
||||
--ids 12345 67890 ^
|
||||
--output-dir outputs/freshrss/single_summaries
|
||||
```
|
||||
|
||||
…or through the MCP server tool `generate_article_summaries` exposed by `summary_mcp.server`:
|
||||
|
||||
- `extracted_path` (string): path to the extracted JSON, for example `outputs/freshrss/extracted/freshrss.extracted.json`.
|
||||
- `selected_ids` (array of strings): one or more `item_id` values from the extracted payload to summarize.
|
||||
- `output_dir` (optional string): directory to write Markdown summaries. If omitted, summaries are written under `single_summaries/` next to the extracted file.
|
||||
- `llm_api_key` / `llm_model` / `llm_api_url` (optional strings): overrides for article-summary LLM settings. If omitted, the tool falls back to `ARTICLE_SUMMARY_*` or main `LLM_*` env vars as described above.
|
||||
|
||||
The tool returns a JSON array of file paths for the generated Markdown summaries.
|
||||
MIT
|
||||
|
||||
@@ -1,83 +1,372 @@
|
||||
# TODO
|
||||
# TODO - Reader MCP 正式化
|
||||
|
||||
## 当前状态
|
||||
> 本文件用于架构与 Codex 协作同步。
|
||||
>
|
||||
> 规则:
|
||||
> - `TODO` = 未开始
|
||||
> - `DOING` = 正在进行
|
||||
> - `DONE` = 已完成
|
||||
> - 每次只允许一个最高优先级主任务处于 `DOING`
|
||||
|
||||
项目当前已经进入“可交付给 OpenClaw 调用”的阶段。
|
||||
## 0. 协作约束
|
||||
|
||||
当前主链路:
|
||||
开始编码前必须阅读:
|
||||
|
||||
`FreshRSS 未读 -> RSS 内容提取 -> LLM 总结 -> 规则过滤 -> OpenClaw delivery payload`
|
||||
|
||||
当前已经完成:
|
||||
|
||||
- [x] FreshRSS `greader` API 接入
|
||||
- [x] `entry -> item` 标准化映射
|
||||
- [x] RSS-first 内容提取策略
|
||||
- [x] LLM 摘要与校验闭环
|
||||
- [x] 第一版规则引擎
|
||||
- [x] `ArticleCandidateRecord` / `OpenClawCandidateInput` 分层
|
||||
- [x] `OpenClawDeliveryPayload` 批量投递结构
|
||||
- [x] 最终 payload 成功后才标记 FreshRSS 已读
|
||||
- [x] MCP 工具 `run_freshrss_openclaw_pipeline`
|
||||
- [x] 默认精简输出模式
|
||||
- [x] 日报级 `keywords` 词元库与周期性词元清洗 skill 设计完成
|
||||
- [x] 日报级 `keywords` 词元库与全局词频统计实现完成
|
||||
- [x] `keyword-cleanup-review` skill 骨架与 review bundle 脚本实现完成
|
||||
- [x] 词元清洗低复杂治理层落地:`term_cleanup_policy` / `term_watchlist` / `term_change_log`
|
||||
- [x] 已支持人工确认采纳建议并写入 `term_watchlist` / `term_change_log`
|
||||
1. `README.md`
|
||||
2. `docs/README.md`
|
||||
3. `docs/openclaw/README.md`
|
||||
4. `docs/openclaw/openclaw-handoff.md`
|
||||
5. `docs/openclaw/openclaw-orchestration-flow.md`
|
||||
6. `plans/README.md`
|
||||
7. `plans/reader-mcp-architecture-design.md`
|
||||
8. `plans/reader-mcp-implementation-plan.md`
|
||||
9. 本文件
|
||||
10. `plans/issues/2026-04-06-reader-digest-sigterm.md`
|
||||
|
||||
---
|
||||
|
||||
## P0 - 交接前后最优先
|
||||
## 1. 当前主任务
|
||||
|
||||
- [x] 为 OpenClaw 补齐交接文档
|
||||
- [x] 将 MCP 工具作为统一生产入口
|
||||
- [x] 将默认输出收敛为最小必要文件
|
||||
- [ ] 设计 OpenClaw webhook / delivery payload 的主动推送方式
|
||||
- [ ] 明确 OpenClaw 侧如何注册和启动本 MCP 服务
|
||||
### [DONE][P0] 建立 run-state 运行态基础设施
|
||||
|
||||
目标:
|
||||
- 给 freshrss pipeline 引入正式 run state
|
||||
- 即使失败或中断,也能留下明确运行真相
|
||||
|
||||
要求:
|
||||
- 新增 `RunState / StageState / ArtifactRecord` 模型
|
||||
- 在 `outputs/freshrss/rerun/<run_id>/run-state.json` 持久化
|
||||
- 至少覆盖以下 stages:
|
||||
- `fetch_feed`
|
||||
- `extract_articles`
|
||||
- `generate_summaries`
|
||||
- `apply_filters`
|
||||
- `build_delivery_payload`
|
||||
- `write_run_report`
|
||||
- 失败时写入失败阶段与错误摘要
|
||||
- 不破坏现有输出目录兼容性
|
||||
|
||||
建议文件:
|
||||
- `src/summary_mcp/runtime/state_models.py`
|
||||
- `src/summary_mcp/runtime/run_store.py`
|
||||
- `src/summary_mcp/workflows/...`
|
||||
|
||||
完成标准:
|
||||
- 跑一次 pipeline 后,无论成功失败,都存在 `run-state.json`
|
||||
- 文件中可看出当前/最后阶段、整体状态、关键 artifacts
|
||||
|
||||
进展备注:
|
||||
- 2026-04-07:架构设计文档已建立;开始进入实现阶段。
|
||||
- 2026-04-07:已新增 `runtime` 包骨架,落地 `RunState / StageState / ArtifactRecord` 与文件存储接口。
|
||||
- 2026-04-07:已将 `run-state.json` 接入 `freshrss` 主流程,按阶段持续写入状态与关键 artifacts。
|
||||
- 2026-04-07:已完成成功/失败路径自检,确认 `run-state.json` 在两类路径下都保留且不改变既有对外返回字段。
|
||||
|
||||
---
|
||||
|
||||
## P1 - 下一阶段推进
|
||||
## 2. 后续任务队列
|
||||
|
||||
- [ ] 设计“人工确认后再沉淀知识库”的状态流转
|
||||
- [ ] 收敛 `paywall` 误判规则,降低中文文本误报
|
||||
- [ ] 细化过滤规则并引入更多个性化上下文
|
||||
- [ ] 将 `keyword-cleanup-review` skill 接入周期性执行流程,产出别名/停用词/兴趣词建议
|
||||
- [ ] 增加清洗前后效果对比报告,验证配置调整是否真的改善过滤质量
|
||||
- [ ] 增加批量 run 的保留策略与历史清理策略
|
||||
- [ ] 为 OpenClaw 补一份更正式的 MCP 调用示例和接线说明
|
||||
### [DONE][P1] 增加 MCP 状态查询接口 `get_run_status`
|
||||
|
||||
目标:
|
||||
- 可通过 MCP 查询 run 状态
|
||||
|
||||
要求:
|
||||
- 输入 `run_id`
|
||||
- 返回 status / current_stage / completed_stages / failed_stage / artifacts / recovery
|
||||
|
||||
完成情况:
|
||||
- 已通过 MCP 暴露 `get_run_status`
|
||||
- 优先读取 `run-state.json`;对无 `run-state.json` 的历史 run 兼容基于现有 run 目录与 `run-report.json` 推断状态
|
||||
- 返回补充了 `progress` / `output_dir` / `state_source`,便于 OpenClaw 稳定消费且不必手拼路径
|
||||
|
||||
改动文件:
|
||||
- `src/summary_mcp/runtime/query_service.py`
|
||||
- `src/summary_mcp/runtime/run_store.py`
|
||||
- `src/summary_mcp/runtime/__init__.py`
|
||||
- `src/summary_mcp/server.py`
|
||||
|
||||
遗留风险:
|
||||
- 历史 run 若缺少 `run-state.json`,其阶段状态只能基于现有目录与 `run-report.json` 做保守推断
|
||||
|
||||
---
|
||||
|
||||
## P2 - 后续增强
|
||||
### [DONE][P1] 增加 MCP 查询接口 `list_runs`
|
||||
|
||||
- [ ] 将 `Markdown sink` 进一步降级为 debug / fallback 能力
|
||||
- [ ] 增加按天聚合 `ArticleCandidateRecord` 的批处理能力
|
||||
- [ ] 让 OpenClaw 聚合候选内容并生成日级摘要
|
||||
- [ ] 将日级摘要写入知识库,并同步生成面向用户的日报消息
|
||||
- [ ] 支持更多 `content_kind`
|
||||
- [ ] 增加提取缓存、重试和更细粒度日志
|
||||
- [ ] 整理历史 rerun 目录与调试产物保留策略
|
||||
目标:
|
||||
- 查看近期 runs
|
||||
|
||||
要求:
|
||||
- 支持按 workflow / status / latest_n 过滤
|
||||
|
||||
完成情况:
|
||||
- 已通过 MCP 暴露 `list_runs`
|
||||
- 支持按 `workflow` / `status` / `latest_n` 过滤近期 runs
|
||||
- 返回 `run_id`、`status`、`progress`、`recovery`、`output_dir` 与 `state_source`
|
||||
|
||||
改动文件:
|
||||
- `src/summary_mcp/runtime/query_service.py`
|
||||
- `src/summary_mcp/runtime/__init__.py`
|
||||
- `src/summary_mcp/server.py`
|
||||
|
||||
遗留风险:
|
||||
- 当前按文件系统扫描 `outputs/freshrss/rerun/` 聚合,规模继续增大时可能需要再评估缓存或索引,但本阶段先保持文件系统真相
|
||||
|
||||
---
|
||||
|
||||
## 当前建议的下一步
|
||||
### [DONE][P1] 增加 MCP 查询接口 `list_run_artifacts`
|
||||
|
||||
优先做这三件事:
|
||||
目标:
|
||||
- 统一列出 run 下 artifact
|
||||
|
||||
1. 将 `keyword-cleanup-review` skill 接入周期性执行流程
|
||||
2. 设计“人工确认后再沉淀知识库”的状态流转
|
||||
3. 收敛规则误判,尤其是 `paywall` 相关启发式
|
||||
完成情况:
|
||||
- 已通过 MCP 暴露 `list_run_artifacts`
|
||||
- 对新 run 优先返回 `run-state.json` 已注册 artifacts,并补充 run 目录扫描发现的标准产物
|
||||
- 对历史 run 直接基于现有 run 目录发现标准产物,保持兼容
|
||||
|
||||
改动文件:
|
||||
- `src/summary_mcp/runtime/query_service.py`
|
||||
- `src/summary_mcp/runtime/__init__.py`
|
||||
- `src/summary_mcp/server.py`
|
||||
|
||||
遗留风险:
|
||||
- 当前仅补充扫描固定的一组标准产物;未注册且不在标准集合内的调试文件不会进入稳定 artifact 列表
|
||||
|
||||
---
|
||||
|
||||
## 交接时优先阅读
|
||||
### [DONE][P1] 增加 MCP 结果读取接口 `get_delivery_payload`
|
||||
|
||||
目标:
|
||||
- 按 run_id 读取 delivery payload
|
||||
|
||||
完成情况:
|
||||
- 已通过 MCP 暴露 `get_delivery_payload`
|
||||
- 查询优先复用 `run-state.json` 已注册 artifacts,其次回退标准产物路径、`run-report.json` 引用和 run 目录扫描
|
||||
- 返回补充了 `artifact`、`payload_schema_version`、`generated_at`、`delivery_date`、`candidate_count`、`stats`,并保留完整 `payload`
|
||||
|
||||
改动文件:
|
||||
- `src/summary_mcp/runtime/query_service.py`
|
||||
- `src/summary_mcp/server.py`
|
||||
- `src/summary_mcp/runtime/__init__.py`
|
||||
|
||||
遗留风险:
|
||||
- 历史 run 的 payload 若既未注册也不在标准路径下,只能依赖 `run-report.json` 引用或目录扫描做兼容发现
|
||||
|
||||
---
|
||||
|
||||
### [DONE][P1] 增加 MCP 结果读取接口 `get_run_report`
|
||||
|
||||
目标:
|
||||
- 按 run_id 读取 run report
|
||||
|
||||
完成情况:
|
||||
- 已通过 MCP 暴露 `get_run_report`
|
||||
- 查询优先复用 `run-state.json` 已注册 artifacts,其次回退标准产物路径、历史 `run-report.json` 固定位置和 run 目录扫描
|
||||
- 返回补充了 `artifact`、核心计数摘要、规范化后的关键产物路径与 `keyword_index`,并保留完整 `report`
|
||||
|
||||
改动文件:
|
||||
- `src/summary_mcp/runtime/query_service.py`
|
||||
- `src/summary_mcp/runtime/__init__.py`
|
||||
- `src/summary_mcp/server.py`
|
||||
|
||||
遗留风险:
|
||||
- 历史 run 的 `run-report.json` 内嵌路径仍保留原始绝对路径于 `report` 字段,当前仅在顶层摘要字段做规范化,避免改变历史文件真相
|
||||
|
||||
---
|
||||
|
||||
### [DONE][P2] 设计并实现 `resume_run`
|
||||
|
||||
目标:
|
||||
- 基于 `run-state.json` 和现有中间产物继续执行
|
||||
|
||||
说明:
|
||||
- 先做最小可用恢复
|
||||
- 暂不追求任意 stage 任意重入
|
||||
- 设计约束已补充到 `plans/resume-run-minimal-design.md`
|
||||
- 第一版只支持 freshrss workflow 且仅支持有 `run-state.json` 的 run
|
||||
- 第一版仅考虑从最近可恢复点继续;`fetch_feed` / `extract_articles` 暂不支持恢复
|
||||
|
||||
进展备注:
|
||||
- 2026-04-07:已完成 `resume_run` minimal design 与现有 runtime/workflow/server 代码对齐分析,开始实现最小恢复链路。
|
||||
- 2026-04-07:已完成 `resume_run` 最小实现编码,新增 runtime 恢复服务并接入 MCP server;当前进入设计对齐与本地自检。
|
||||
- 2026-04-07:已完成 `resume_run` 架构对齐与本地自检;已验证 `write_run_report` 可恢复,且 `extract_articles` 会被明确拒绝恢复。
|
||||
- 2026-04-14:已补 `inspect_resume_plan`、artifact-first 恢复判定,以及生产模式下稳定 `summary-batch` / `candidate-batch` artifacts;当前 `resume` 的剩余主问题不再是恢复点判断,而是同步执行模型仍可能让 OpenClaw 恢复阶段超时。
|
||||
|
||||
---
|
||||
|
||||
### [DONE][P1] 把 `resume_run` 升级为最小真异步 job
|
||||
|
||||
目标:
|
||||
- 解决 `resume_run` 在 OpenClaw → MCP 同步链路里仍可能超时的问题
|
||||
- 让恢复也具备“启动 / 轮询 / 读取结果”的正式控制面
|
||||
|
||||
要求:
|
||||
- 新增最小异步接口:
|
||||
- `start_resume_job`
|
||||
- `get_resume_job_status`
|
||||
- `get_resume_job_result`
|
||||
- 状态目录固定落到:
|
||||
- `outputs/freshrss/resume_jobs/<job_id>/`
|
||||
- 至少包含:
|
||||
- `run-state.json`
|
||||
- `input.json`
|
||||
- `result.json`(成功时)
|
||||
- `job-report.json`
|
||||
- 启动前必须先走 `inspect_resume_plan`
|
||||
- 业务执行继续复用现有 `resume_service`,不要重写恢复主逻辑
|
||||
- `resume_run` 保留为同步 debug / fallback 路径,但不再作为 OpenClaw 的默认恢复入口
|
||||
|
||||
完成标准:
|
||||
- 可恢复 run 上,`start_resume_job` 能成功返回 `job_id`
|
||||
- `get_resume_job_status` 能稳定反映恢复 job 生命周期
|
||||
- `get_resume_job_result` 能稳定返回 `run_id`、`resume_from_stage`、最终状态与关键产物路径
|
||||
- 恢复耗时超过单次 MCP 同步窗口时,OpenClaw 仍不会因为同步调用挂住
|
||||
|
||||
进展备注:
|
||||
- 2026-04-14:已落地 `src/summary_mcp/runtime/resume_jobs.py` 与 `scripts/run_resume_job.py`,新增 `start_resume_job` / `get_resume_job_status` / `get_resume_job_result`
|
||||
- 2026-04-14:启动前会先走 `inspect_resume_plan`;不可恢复 run 会在 job 输入校验阶段直接失败,不进入后台恢复执行
|
||||
- 2026-04-14:后台执行复用现有 `_resume_freshrss_run(...)`,没有重写恢复主逻辑
|
||||
- 2026-04-14:已完成本地 synthetic 验证:`write_run_report` 恢复可通过 `start -> poll -> result` 闭环成功收敛
|
||||
|
||||
---
|
||||
|
||||
### [DONE][P1] 单篇总结改为最小真异步 job
|
||||
|
||||
目标:
|
||||
- 解决 `generate_article_summaries` 在 OpenClaw → MCP 同步链路里易 timeout 的问题
|
||||
- 将单篇总结正式升级为可启动、可轮询、可读取结果的异步 job
|
||||
|
||||
要求:
|
||||
- 新增最小异步接口:
|
||||
- `start_article_summary_job`
|
||||
- `get_article_summary_job_status`
|
||||
- `get_article_summary_job_result`
|
||||
- 状态目录固定落到:
|
||||
- `outputs/freshrss/article_summary_jobs/<job_id>/`
|
||||
- 至少包含:
|
||||
- `run-state.json`
|
||||
- `input.json`
|
||||
- `result.json`(成功时)
|
||||
- `job-report.json`
|
||||
- 执行模型优先使用后台子进程,不使用线程
|
||||
- 业务逻辑继续复用 `summarize_selected_articles(...)`,不要重写正文总结核心逻辑
|
||||
- 对 OpenClaw / reader-digest-flow 而言,异步 job 成功后应可继续接 IMA 沉淀闭环
|
||||
|
||||
当前进展:
|
||||
- 2026-04-10:已完成方案文档 `plans/article-summary-async-job-plan.md`
|
||||
- 2026-04-10:已落地最小代码骨架:
|
||||
- `src/summary_mcp/runtime/article_summary_jobs.py`
|
||||
- `scripts/run_article_summary_job.py`
|
||||
- `src/summary_mcp/server.py` 已新增 3 个 async job tools
|
||||
- 2026-04-10:已用真实 extracted 文件验证最小异步链路可跑通,job 能成功进入 `running → success`,并可读回结果
|
||||
- 2026-04-10:已补 README / OpenClaw handoff 文档,并对外统一为 `start_article_summary_job` / `get_article_summary_job_status` / `get_article_summary_job_result`
|
||||
- 2026-04-10:已完成聚焦自检:
|
||||
- MCP tool 注册名校验通过
|
||||
- stubbed async job 成功路径通过
|
||||
- stubbed async job 失败路径通过,`error_summary` 与 `job-report.json` 可回读
|
||||
|
||||
下一步:
|
||||
- 将 `reader-digest-flow` 正式默认路径切到 async job
|
||||
- 用真实 LLM 配置再做一次非 stub 的服务端冒烟验证
|
||||
|
||||
---
|
||||
|
||||
### [DOING][P0] FreshRSS 主日报 run 改为最小真异步 job
|
||||
|
||||
目标:
|
||||
- 解决 `run_freshrss_openclaw_pipeline` 在正式生产链路里仍为同步 MCP 调用、易超时的问题
|
||||
- 将 FreshRSS 主日报启动路径升级为可启动、可轮询、可读取结果的异步 job
|
||||
|
||||
要求:
|
||||
- 新增最小异步接口:
|
||||
- `start_freshrss_pipeline_job`
|
||||
- `get_freshrss_pipeline_job_status`
|
||||
- `get_freshrss_pipeline_job_result`
|
||||
- 状态目录固定落到:
|
||||
- `outputs/freshrss/pipeline_jobs/<job_id>/`
|
||||
- 至少包含:
|
||||
- `run-state.json`
|
||||
- `input.json`
|
||||
- `result.json`(成功时)
|
||||
- `job-report.json`
|
||||
- 执行模型优先使用后台子进程,不使用线程
|
||||
- 主业务逻辑继续复用 `run_freshrss_pipeline(...)`,不要重写日报核心逻辑
|
||||
- job 成功后结果中必须带回 `run_id` 与关键产物路径
|
||||
- README / handoff / OpenClaw 生产建议路径需要同步改成 async start path
|
||||
|
||||
当前进展:
|
||||
- 2026-04-11:问题定位完成,确认之前异步化的是 article-summary,不是主日报 run
|
||||
- 2026-04-11:已新增方案文档 `plans/freshrss-pipeline-async-job-plan.md`
|
||||
|
||||
下一步:
|
||||
- 复用 article-summary job runtime 骨架实现主日报 async job
|
||||
- 新增后台 runner 脚本
|
||||
- 暴露 3 个 MCP tools
|
||||
- 用真实 MCP 冒烟验证 `start -> status -> result`
|
||||
|
||||
---
|
||||
|
||||
### [TODO][P2] 评估 `rerun_stage` 是否值得进入第一阶段
|
||||
|
||||
目标:
|
||||
- 在 `resume_run` 之后评估是否继续增加更细粒度补跑能力
|
||||
|
||||
---
|
||||
|
||||
### [TODO][P2] 调整 `run_freshrss_openclaw_pipeline` 内部实现以复用 runtime
|
||||
|
||||
目标:
|
||||
- 保持外部兼容
|
||||
- 内部不再是黑箱长函数
|
||||
|
||||
---
|
||||
|
||||
### [DONE][P1] interest/watch 候选引擎从固定阈值改为百分位排名 + 增速因子
|
||||
|
||||
目标:
|
||||
- 解决固定阈值(total_count>=3)不随数据量自适应的问题
|
||||
- 引入趋势信号(growth 因子),识别近期集中爆发的词
|
||||
- 支持 7 天、41 天、200 天数据量下取同样的 top 5%/5%-20% 而不需调阈值
|
||||
|
||||
要求:
|
||||
- `build_review_bundle.py`:新增 percentile 和 growth 计算函数;候选池从固定阈值改为百分位 + 增速
|
||||
- `configs/term_cleanup_policy.json`:升级为 v2 schema,percentile/growth 替代绝对阈值
|
||||
- 不改 `generate_term_cleanup_suggestions.py` 和 `apply_term_suggestions.py`
|
||||
- 全量跑一次对比新旧产出,确认差异合理
|
||||
|
||||
方案文档:`plans/keyword-cleanup-interest-watch-engine-improvement.md`
|
||||
|
||||
---
|
||||
|
||||
### [DONE][P3] 更新 README / handoff / docs,明确 MCP 为正式入口
|
||||
|
||||
目标:
|
||||
- 把生产建议从 CLI 迁移到 MCP
|
||||
- CLI 明确降级为 debug / fallback
|
||||
|
||||
完成情况:
|
||||
- 已更新 `README.md`,补齐 reader 作为正式 MCP workflow service 的当前能力边界、推荐调用路径、已支持 tools 与最小 `resume_run` 范围
|
||||
- 已更新 `docs/openclaw/openclaw-handoff.md`,明确 OpenClaw 应优先通过 MCP 读取 run 状态与结果,不再自己拼接 reader 输出路径
|
||||
|
||||
改动文件:
|
||||
- `README.md`
|
||||
- `docs/openclaw/openclaw-handoff.md`
|
||||
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||
- `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||
- `docs/design/daily-keyword-index-design.md`
|
||||
- `skills/keyword-cleanup-review/SKILL.md`
|
||||
- `docs/current/context-reset-brief.md`
|
||||
|
||||
遗留风险:
|
||||
- 当前仍无独立 `get_digest_brief` tool;若下游确实需要该产物,仍应先通过 `list_run_artifacts` / `get_run_report` 发现,而不是写死路径
|
||||
|
||||
---
|
||||
|
||||
## 3. 记录区
|
||||
|
||||
### 已完成记录
|
||||
|
||||
- 2026-04-07:新增架构设计文档 `plans/reader-mcp-architecture-design.md`
|
||||
- 2026-04-07:新增实施计划文档 `plans/reader-mcp-implementation-plan.md`
|
||||
- 2026-04-07:完成 `freshrss` pipeline 的 run-state 基础设施,新增 `runtime` 包并覆盖关键 stages 状态持久化。
|
||||
- 2026-04-14:已补 OpenClaw 文档导航、历史归档、design/notes/plans 导航,并统一当前正式口径为 async job 编排入口。
|
||||
|
||||
### 风险提醒
|
||||
|
||||
- 不要在第一阶段引入复杂任务队列
|
||||
- 不要让 CLI 和 MCP 背后变成两套独立逻辑
|
||||
- 若实现偏离架构,先更新 `plans/` 再改代码
|
||||
|
||||
@@ -23,34 +23,66 @@
|
||||
"前沿科技"
|
||||
],
|
||||
"interest_keywords": [
|
||||
"Java",
|
||||
"Go",
|
||||
"Python",
|
||||
"Spring",
|
||||
"Agent",
|
||||
"Agent Skills",
|
||||
"AgentScope",
|
||||
"AI Agent",
|
||||
"AI Coding Agent",
|
||||
"AliSQL",
|
||||
"Anthropic",
|
||||
"Claude",
|
||||
"Claude Code",
|
||||
"CLAUDE.md",
|
||||
"CLI",
|
||||
"Context Engineering",
|
||||
"Cursor",
|
||||
"ChatGPT",
|
||||
"DeepSeek",
|
||||
"FastAPI",
|
||||
"Gin",
|
||||
"Go",
|
||||
"gRPC",
|
||||
"MySQL",
|
||||
"PostgreSQL",
|
||||
"Redis",
|
||||
"Harness Engineering",
|
||||
"Hermes Agent",
|
||||
"Java",
|
||||
"Kafka",
|
||||
"微服务",
|
||||
"可观测性",
|
||||
"Kubernetes",
|
||||
"云原生",
|
||||
"AI Agent",
|
||||
"Agent",
|
||||
"LLM",
|
||||
"RAG",
|
||||
"Loop Engineering",
|
||||
"MCP",
|
||||
"Prompt Engineering",
|
||||
"Workflow",
|
||||
"向量数据库",
|
||||
"知识库",
|
||||
"MoE",
|
||||
"MySQL",
|
||||
"MySQL复制延迟",
|
||||
"OpenAI",
|
||||
"DeepSeek",
|
||||
"OpenClaw",
|
||||
"AliSQL",
|
||||
"MySQL复制延迟"
|
||||
"PostgreSQL",
|
||||
"Prompt Engineering",
|
||||
"Python",
|
||||
"RAG",
|
||||
"ReAct",
|
||||
"ReActAgent",
|
||||
"Redis",
|
||||
"Skill",
|
||||
"SKILL.md",
|
||||
"Skills",
|
||||
"Spring",
|
||||
"SubAgent",
|
||||
"TypeScript",
|
||||
"Vibe Coding",
|
||||
"Workflow",
|
||||
"上下文压缩",
|
||||
"上下文工程",
|
||||
"上下文管理",
|
||||
"云原生",
|
||||
"代码审查",
|
||||
"可观测性",
|
||||
"向量数据库",
|
||||
"多Agent协作",
|
||||
"大模型",
|
||||
"子Agent",
|
||||
"强化学习",
|
||||
"微服务",
|
||||
"渐进式披露",
|
||||
"知识库"
|
||||
]
|
||||
}
|
||||
+146
-2
@@ -1,5 +1,149 @@
|
||||
{
|
||||
"AI助手": "AI Agent",
|
||||
"Agent": "Agent",
|
||||
"Agent框架": "Agent",
|
||||
"Agent能力": "Agent Skills",
|
||||
"智能体": "AI Agent",
|
||||
"Agentic架构": "Agentic架构",
|
||||
"多Agent协作": "多Agent协作",
|
||||
"多智能体架构": "多Agent协作",
|
||||
"Multi-Agent": "多Agent",
|
||||
"Subagent": "子Agent",
|
||||
"Sub Agents验证": "子Agent",
|
||||
"子Agent": "子Agent",
|
||||
"子智能体": "子Agent",
|
||||
"Coding Agent": "AI Coding Agent",
|
||||
"AI编程": "AI Coding Agent",
|
||||
"AI辅助编程": "AI Coding Agent",
|
||||
"代码生成": "AI代码生成",
|
||||
"代码审查": "Code Review",
|
||||
"Prompt": "Prompt Engineering",
|
||||
"Prompt Caching": "提示缓存",
|
||||
"RAG": "RAG",
|
||||
"图文RAG": "RAG",
|
||||
"Prompt架构": "Prompt Engineering"
|
||||
"向量检索": "向量检索",
|
||||
"向量嵌入": "向量嵌入",
|
||||
"Multi-Token Prediction": "多Token预测",
|
||||
"Pair-In Pair-Out": "PIPO架构",
|
||||
"PIPO": "PIPO架构",
|
||||
"上下文管理": "上下文管理",
|
||||
"上下文卸载": "上下文卸载",
|
||||
"Self-GC": "上下文压缩",
|
||||
"记忆管理": "上下文管理",
|
||||
"会话管理": "上下文管理",
|
||||
"Harness Engineering": "Harness工程化",
|
||||
"Harness架构": "Harness工程化",
|
||||
"Harness": "Harness工程化",
|
||||
"Loop Engineering": "Loop Engineering",
|
||||
"推理加速": "推理加速",
|
||||
"推理深度": "推理深度",
|
||||
"长链路推理": "长链路推理",
|
||||
"RLVR": "RLVR",
|
||||
"GRPO": "GRPO",
|
||||
"强化学习": "强化学习",
|
||||
"Multi-Agent RL": "多Agent强化学习",
|
||||
"Viking AI搜索": "AI搜索",
|
||||
"Viking AI Search": "AI搜索",
|
||||
"智能搜索": "AI搜索",
|
||||
"SearchCLI": "CLI搜索",
|
||||
"视频生成": "AI视频生成",
|
||||
"视频生成模型": "AI视频生成",
|
||||
"LingBot-Video": "AI视频生成",
|
||||
"视觉自回归模型": "AI视频生成",
|
||||
"火山云数据库PostgreSQL Serverless版": "Serverless数据库",
|
||||
"PostgreSQL": "PostgreSQL",
|
||||
"MySQL": "MySQL",
|
||||
"OceanBase": "OceanBase",
|
||||
"StarRocks": "StarRocks",
|
||||
"Milvus": "Milvus",
|
||||
"Seal AI Zone": "AI安全",
|
||||
"NEX沙箱": "沙箱隔离",
|
||||
"MicroVM": "沙箱隔离",
|
||||
"安全左移": "安全左移",
|
||||
"安全中台": "AI安全",
|
||||
"成本降低": "成本优化",
|
||||
"成本杠杆": "成本优化",
|
||||
"Scale-to-Zero": "弹性伸缩",
|
||||
"Data as Git": "数据分支管理",
|
||||
"Schema Diff": "Schema对比",
|
||||
"Time Travel": "数据回溯",
|
||||
"多端架构": "多端架构",
|
||||
"契约化": "契约化架构",
|
||||
"大仓": "大仓工程化",
|
||||
"Vibe Coding": "Vibe Coding",
|
||||
"LLM Judge": "LLM评估",
|
||||
"SWE-Bench": "SWE-Bench",
|
||||
"SWE Bench Pro": "SWE-Bench",
|
||||
"SWE-Bench Pro": "SWE-Bench",
|
||||
"Verification Agent": "验证Agent",
|
||||
"CLI工具": "CLI",
|
||||
"CLI": "CLI",
|
||||
"漏桶算法": "限流架构",
|
||||
"固定窗口限流": "限流架构",
|
||||
"Suspend消费控制": "限流架构",
|
||||
"RocketMQ LiteTopic": "消息队列",
|
||||
"LLM Wiki": "LLM知识库",
|
||||
"知识工程": "知识工程",
|
||||
"语义资产": "语义资产管理",
|
||||
"知识图谱": "知识图谱",
|
||||
"知识库沉淀": "知识管理",
|
||||
"Skill": "Skill",
|
||||
"Skill Hub": "技能生态",
|
||||
"具身智能": "具身智能",
|
||||
"Open X-Embodiment": "具身智能",
|
||||
"YOLO Classifier": "目标检测",
|
||||
"MCP": "MCP",
|
||||
"MCP连接器": "MCP",
|
||||
"缓存击穿": "缓存优化",
|
||||
"GPU算力调度": "算力调度",
|
||||
"异构资源": "异构计算",
|
||||
"XPU": "异构计算",
|
||||
"弹性RDMA": "RDMA网络",
|
||||
"国内主流GPU": "国产芯片",
|
||||
"国产AI芯片": "国产芯片",
|
||||
"Paxos协议": "分布式一致性",
|
||||
"Token": "Token管理",
|
||||
"Token效率": "Token管理",
|
||||
"百万token上下文": "长上下文",
|
||||
"MoE": "MoE架构",
|
||||
"MoE架构": "MoE架构",
|
||||
"思维链": "思维链",
|
||||
"CoT Distillation": "思维链蒸馏",
|
||||
"自然语言驱动": "自然语言交互",
|
||||
"NL2SQL": "NL2SQL",
|
||||
"AI对齐": "AI对齐",
|
||||
"注意力机制": "注意力机制",
|
||||
"多模态": "多模态",
|
||||
"音视频工作台": "音视频处理",
|
||||
"AI助手": "AI Agent",
|
||||
"Agent架构": "AI Agent",
|
||||
"Agent专业化": "AI Agent",
|
||||
"Agent Teams": "多Agent协作",
|
||||
"Agentic Engineering": "AI Agent",
|
||||
"AI智能体": "AI Agent",
|
||||
"LLM Agent": "AI Agent",
|
||||
"AI Harness": "Harness Engineering",
|
||||
"AI代码生成": "AI Coding Agent",
|
||||
"Memory管理": "上下文管理",
|
||||
"Agent Skill": "Agent Skills",
|
||||
"Binlog": "binlog",
|
||||
"vibe coding": "Vibe Coding",
|
||||
"Agent组织化协作平台": "Agent协作平台",
|
||||
"Anthropic": "Anthropic",
|
||||
"OpenClaw": "OpenClaw",
|
||||
"WorkBuddy": "WorkBuddy",
|
||||
"Claude": "Claude",
|
||||
"ChatGPT": "ChatGPT",
|
||||
"GPT": "GPT",
|
||||
"Opus": "Opus",
|
||||
"Sonnet": "Sonnet",
|
||||
"Grok": "Grok",
|
||||
"Qwen": "Qwen",
|
||||
"GLM": "GLM",
|
||||
"Claude Code": "Claude Code",
|
||||
"Cursor": "Cursor",
|
||||
"Codex": "Codex",
|
||||
"Pi": "Pi",
|
||||
"CoT": "CoT",
|
||||
"SVG": "SVG",
|
||||
"TTS": "TTS"
|
||||
}
|
||||
@@ -1,4 +1,595 @@
|
||||
{
|
||||
"schema_version": "v1",
|
||||
"entries": []
|
||||
"entries": [
|
||||
{
|
||||
"applied_at": "2026-04-08T02:34:14.194320Z",
|
||||
"action": "add_watch_term",
|
||||
"term": "A2A",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=0.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"suggestion_date": "2026-04-08",
|
||||
"based_on_days": 7
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-04-08T02:34:14.194320Z",
|
||||
"action": "add_watch_term",
|
||||
"term": "Agentic Loop",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=1.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"suggestion_date": "2026-04-08",
|
||||
"based_on_days": 7
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-04-08T02:34:14.194320Z",
|
||||
"action": "add_watch_term",
|
||||
"term": "AI Gateway",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=1.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"suggestion_date": "2026-04-08",
|
||||
"based_on_days": 7
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-04-08T02:34:14.194320Z",
|
||||
"action": "add_watch_term",
|
||||
"term": "Claude Skills",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=1.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"suggestion_date": "2026-04-08",
|
||||
"based_on_days": 7
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-04-08T02:34:14.194320Z",
|
||||
"action": "add_watch_term",
|
||||
"term": "Cron",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=0.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"suggestion_date": "2026-04-08",
|
||||
"based_on_days": 7
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-04-08T02:34:14.194320Z",
|
||||
"action": "add_watch_term",
|
||||
"term": "CoPaw",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=0.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"suggestion_date": "2026-04-08",
|
||||
"based_on_days": 7
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-04-08T02:34:14.194320Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Claude Code",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=7, days_seen=4, recent_count=4.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"suggestion_date": "2026-04-08",
|
||||
"based_on_days": 7
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-04-08T02:34:14.194320Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Agent Skills",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=3, recent_count=0.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"suggestion_date": "2026-04-08",
|
||||
"based_on_days": 7
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-04-08T02:34:14.194320Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "SubAgent",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=4, recent_count=2.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"suggestion_date": "2026-04-08",
|
||||
"based_on_days": 7
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-04-08T02:34:14.194320Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "AgentScope",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=4, recent_count=1.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"suggestion_date": "2026-04-08",
|
||||
"based_on_days": 7
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-04-08T02:34:14.194320Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "ReActAgent",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=3, days_seen=3, recent_count=1.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"suggestion_date": "2026-04-08",
|
||||
"based_on_days": 7
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Anthropic",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=13, days_seen=10, recent_count=13.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Harness Engineering",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=12, days_seen=11, recent_count=12.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Skill",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=11, days_seen=9, recent_count=11.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "上下文工程",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=6, recent_count=6.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "多Agent协作",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=6, recent_count=6.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Claude",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=5, recent_count=6.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "上下文管理",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=5, recent_count=6.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "渐进式披露",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=5, days_seen=5, recent_count=5.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Skills",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=5, days_seen=4, recent_count=5.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "SKILL.md",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=4, recent_count=4.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "CLAUDE.md",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=4, recent_count=4.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "上下文压缩",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=4, recent_count=4.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "AI Coding Agent",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=3, days_seen=3, recent_count=3.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Hermes Agent",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=3, days_seen=3, recent_count=3.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Vibe Coding",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=5, days_seen=4, recent_count=5.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Context Engineering",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=3, days_seen=3, recent_count=3.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Cursor",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=3, days_seen=3, recent_count=3.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T08:20:35.134979Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "大模型",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered by interest keywords or stopwords. Evidence: total_count=4, days_seen=3, recent_count=4.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "TypeScript",
|
||||
"reason": "Core language for AI agent development (e.g., Claude Code, Cursor) and backend engineering, complements existing Python/Java/Go keywords.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "代码审查",
|
||||
"reason": "Chinese term for 'code review', a key practice in backend engineering and AI agent development workflows.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "Channels",
|
||||
"reason": "Too generic; could refer to communication channels, YouTube channels, or software channels, not specific to user's focus areas.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "Memory",
|
||||
"reason": "Extremely broad term; could refer to computer memory, human memory, or memory in various contexts, not discriminative enough.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "Prompt",
|
||||
"reason": "Already covered by 'Prompt Engineering' as a more specific term; 'Prompt' alone is too broad and matches many unrelated articles.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "AGI",
|
||||
"reason": "Too broad and speculative; not directly actionable for the user's practical engineering focus areas.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "AI日报",
|
||||
"reason": "Generic news term; not a technical concept or tool, would add noise to the keyword index.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "AIHOT",
|
||||
"reason": "Unclear meaning, likely a brand or aggregator, not a specific technical term.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "All In Code",
|
||||
"reason": "Too vague; could refer to a podcast, a philosophy, or a project, not a specific technical concept.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "auto-twitter-campaign",
|
||||
"reason": "Too specific to a single project/tool, not a general interest keyword for the user's focus areas.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "ChangeSet",
|
||||
"reason": "Generic term used in version control and databases; too broad to be a useful filter.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "Lumina",
|
||||
"reason": "Unclear reference; could be a product, framework, or brand, not clearly aligned with user's focus.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "OpenViking",
|
||||
"reason": "Unclear reference; not a known tool or concept in the user's stated focus areas.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "Seedance 2.0",
|
||||
"reason": "Unclear reference; likely a product or version, not a general technical term.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_stopword",
|
||||
"term": "质量门禁",
|
||||
"reason": "Chinese term for 'quality gate', too generic in software engineering; not specific to user's focus areas.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Agent Skill",
|
||||
"reason": "Singular variant",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "Agent Skills",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Binlog",
|
||||
"reason": "Case variant (auto-ranked)",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "binlog",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Coding Agent",
|
||||
"reason": "Abbreviated form of 'AI Coding Agent', referring to the same concept.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "AI Coding Agent",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Subagent",
|
||||
"reason": "Case variant",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "SubAgent",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Subagents",
|
||||
"reason": "Plural variant",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "SubAgent",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "vibe coding",
|
||||
"reason": "Case variant",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "Vibe Coding",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Agent架构",
|
||||
"reason": "Chinese translation of 'Agent architecture', a core concept in AI Agent engineering.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "AI Agent",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Agent专业化",
|
||||
"reason": "Chinese term for 'Agent specialization', directly related to Agent engineering.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "AI Agent",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Agent Teams",
|
||||
"reason": "English equivalent of 'Multi-Agent collaboration', same concept.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "多Agent协作",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "Agentic Engineering",
|
||||
"reason": "Broader term for engineering with AI agents, closely related to Agent engineering focus.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "AI Agent",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "CLI工具",
|
||||
"reason": "Chinese translation of 'CLI tool', same concept.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "CLI",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "AI编程",
|
||||
"reason": "Chinese term for 'AI programming', closely related to AI Coding Agent.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "AI Coding Agent",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "记忆管理",
|
||||
"reason": "Chinese term for 'memory management', closely related to context management in LLM applications.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "上下文管理",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-05-14T09:12:50.749877Z",
|
||||
"action": "add_alias",
|
||||
"term": "会话管理",
|
||||
"reason": "Chinese term for 'session management', related to context management in LLM applications.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-05-14.json",
|
||||
"value": "上下文管理",
|
||||
"suggestion_date": "2026-05-14",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-07-15T02:17:50.155586Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "强化学习",
|
||||
"reason": "top 0.8% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=10, days_seen=10, recent_count=10.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
|
||||
"suggestion_date": "2026-07-15",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-07-15T02:17:50.155586Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "ReAct",
|
||||
"reason": "top 1.1% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=8, days_seen=8, recent_count=8.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
|
||||
"suggestion_date": "2026-07-15",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-07-15T02:17:50.155586Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "CLI",
|
||||
"reason": "top 1.1% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=8, days_seen=7, recent_count=8.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
|
||||
"suggestion_date": "2026-07-15",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-07-15T02:17:50.155586Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "子Agent",
|
||||
"reason": "top 1.3% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=7, days_seen=6, recent_count=7.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
|
||||
"suggestion_date": "2026-07-15",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-07-15T02:17:50.155586Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "Loop Engineering",
|
||||
"reason": "top 1.5% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=6, days_seen=6, recent_count=6.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
|
||||
"suggestion_date": "2026-07-15",
|
||||
"based_on_days": 365
|
||||
},
|
||||
{
|
||||
"applied_at": "2026-07-15T02:17:50.155586Z",
|
||||
"action": "add_interest_keyword",
|
||||
"term": "MoE",
|
||||
"reason": "top 2.0% by frequency,growth=100%,not yet covered by interest keywords or stopwords. Evidence: total_count=5, days_seen=4, recent_count=5.",
|
||||
"suggestions_path": "outputs/term_index/review/term-cleanup-suggestions-2026-07-15.json",
|
||||
"suggestion_date": "2026-07-15",
|
||||
"based_on_days": 365
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -1,14 +1,12 @@
|
||||
{
|
||||
"schema_version": "v1",
|
||||
"schema_version": "v2",
|
||||
"interest_keyword_review": {
|
||||
"min_total_count": 3,
|
||||
"min_days_seen": 2
|
||||
"percentile_max": 0.05,
|
||||
"growth_promotion": 0.5
|
||||
},
|
||||
"watch_term_review": {
|
||||
"min_total_count": 1,
|
||||
"min_days_seen": 1,
|
||||
"max_total_count": 2,
|
||||
"max_days_seen": 2
|
||||
"percentile_min": 0.05,
|
||||
"percentile_max": 0.20
|
||||
},
|
||||
"alias_review": {
|
||||
"min_total_count": 2,
|
||||
@@ -19,7 +17,10 @@
|
||||
"max_days_seen": 2
|
||||
},
|
||||
"notes": [
|
||||
"当前阶段采用保守阈值,避免在低样本条件下直接扩充 interest_keywords。",
|
||||
"watch_terms 先用于观察,后续再决定是否升格为 interest_keywords 或进入 alias/stopword 配置。"
|
||||
"v2: interest/watch 使用百分位排名 + 增速因子替代固定阈值",
|
||||
"percentile 越小表示排名越高(top 5% = percentile 0.05)",
|
||||
"growth = recent_count / total_count,衡量近期活跃度",
|
||||
"watch_term_review 的 percentile_min 可理解为兴趣边界下限,低于此值的词归入 interest 候选",
|
||||
"growth_promotion(默认 0.5)用于识别近期集中爆发词,即使排位不高也主动推荐确认"
|
||||
]
|
||||
}
|
||||
+121
-1
@@ -1,6 +1,126 @@
|
||||
[
|
||||
"1688",
|
||||
"AGI",
|
||||
"AIHOT",
|
||||
"AI日报",
|
||||
"All In Code",
|
||||
"Andrej Karpathy",
|
||||
"Anthropic",
|
||||
"auto-twitter-campaign",
|
||||
"Boundaries",
|
||||
"ChangeSet",
|
||||
"Channels",
|
||||
"Claude Fable 5",
|
||||
"Claude Mythos",
|
||||
"Cohere",
|
||||
"Confidence Head",
|
||||
"Cosmos 3",
|
||||
"Databricks",
|
||||
"DINOv2",
|
||||
"domain-mapping",
|
||||
"Dropbox",
|
||||
"EchoGen",
|
||||
"FLUX.1-dev VAE",
|
||||
"GB300 GPU",
|
||||
"GLM 5.2",
|
||||
"GLM5.0",
|
||||
"GPT-5.5",
|
||||
"GPT-Live",
|
||||
"GPT5.5",
|
||||
"Grok 4.5",
|
||||
"GrowBrain",
|
||||
"iMedImage",
|
||||
"iMedLoop",
|
||||
"iMedMaaS",
|
||||
"iMedStudio",
|
||||
"J-space",
|
||||
"JLens",
|
||||
"John Jumper",
|
||||
"J空间",
|
||||
"KAIROS",
|
||||
"KubeRay",
|
||||
"LibTV Agent",
|
||||
"LingBot-Video",
|
||||
"Lumina",
|
||||
"Markdown",
|
||||
"Marvis",
|
||||
"MDASH",
|
||||
"Meta Superintelligence Labs",
|
||||
"MTS",
|
||||
"Muse Image",
|
||||
"Muse Video",
|
||||
"N-gram Embedding",
|
||||
"OCP China",
|
||||
"OCP China 2026",
|
||||
"On-Policy Distillation",
|
||||
"OPC训练营",
|
||||
"OpenAI",
|
||||
"OpenBMC",
|
||||
"OpenClaw",
|
||||
"OpenViking",
|
||||
"Opus 4.8",
|
||||
"Qwen3",
|
||||
"Qwen3-30B-A3B",
|
||||
"RAS API",
|
||||
"Redfish",
|
||||
"ScMoE",
|
||||
"Seal AI Zone",
|
||||
"SealRouter",
|
||||
"Seedance 2.0",
|
||||
"Sonnet 5",
|
||||
"Spec模式",
|
||||
"STE固件团队",
|
||||
"Three.js",
|
||||
"Unity AI Gateway",
|
||||
"Vant Weapp",
|
||||
"WeTV",
|
||||
"WorkBuddy",
|
||||
"wpc",
|
||||
"YOLO Classifier",
|
||||
"一人公司",
|
||||
"中国科学技术大学",
|
||||
"五大扶持体系",
|
||||
"出门问问",
|
||||
"分镜",
|
||||
"剧本",
|
||||
"奋斗文化",
|
||||
"字节跳动",
|
||||
"小银",
|
||||
"得力",
|
||||
"德适科技",
|
||||
"成都天府长岛",
|
||||
"扣子",
|
||||
"星云平台",
|
||||
"火山引擎",
|
||||
"百度百舸",
|
||||
"百炼网关",
|
||||
"科大讯飞",
|
||||
"腾讯云开发者社区",
|
||||
"腾讯混元Hy3",
|
||||
"蚂蚁灵波",
|
||||
"贝尔实验室",
|
||||
"质量门禁",
|
||||
"配乐",
|
||||
"配音",
|
||||
"银行客户经理",
|
||||
"飞书妙搭",
|
||||
"飞盘物理",
|
||||
"奋斗文化"
|
||||
"自动化",
|
||||
"定时任务",
|
||||
"开源模型",
|
||||
"陌生化",
|
||||
"AlphaFold",
|
||||
"Brand Kit",
|
||||
"DataWorks",
|
||||
"Enhance-Nanocodec",
|
||||
"IRIS Codec",
|
||||
"Lovart",
|
||||
"MiniMax M3",
|
||||
"Gemini 3.5 Flash",
|
||||
"Codex",
|
||||
"CodeBuddy",
|
||||
"Claude Cowork",
|
||||
"AGENTS.md",
|
||||
"Claude",
|
||||
"RLVR"
|
||||
]
|
||||
@@ -1,5 +1,48 @@
|
||||
{
|
||||
"schema_version": "v1",
|
||||
"updated_at": "2026-03-27T00:00:00Z",
|
||||
"terms": []
|
||||
"updated_at": "2026-07-15T02:17:50.155586Z",
|
||||
"terms": [
|
||||
{
|
||||
"term": "A2A",
|
||||
"added_at": "2026-04-08T02:34:14.194320Z",
|
||||
"source": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=0.",
|
||||
"status": "watching"
|
||||
},
|
||||
{
|
||||
"term": "Agentic Loop",
|
||||
"added_at": "2026-04-08T02:34:14.194320Z",
|
||||
"source": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=1.",
|
||||
"status": "watching"
|
||||
},
|
||||
{
|
||||
"term": "AI Gateway",
|
||||
"added_at": "2026-04-08T02:34:14.194320Z",
|
||||
"source": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=1.",
|
||||
"status": "watching"
|
||||
},
|
||||
{
|
||||
"term": "Claude Skills",
|
||||
"added_at": "2026-04-08T02:34:14.194320Z",
|
||||
"source": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=1.",
|
||||
"status": "watching"
|
||||
},
|
||||
{
|
||||
"term": "CoPaw",
|
||||
"added_at": "2026-04-08T02:34:14.194320Z",
|
||||
"source": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=0.",
|
||||
"status": "watching"
|
||||
},
|
||||
{
|
||||
"term": "Cron",
|
||||
"added_at": "2026-04-08T02:34:14.194320Z",
|
||||
"source": "outputs/term_index/review/term-cleanup-suggestions-2026-04-08.json",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed before promotion into interest keywords. Evidence: total_count=2, days_seen=2, recent_count=0.",
|
||||
"status": "watching"
|
||||
}
|
||||
]
|
||||
}
|
||||
+50
-38
@@ -1,6 +1,25 @@
|
||||
# 文档索引
|
||||
|
||||
## 当前目录结构
|
||||
## 当前最短阅读路径
|
||||
|
||||
1. `README.md`
|
||||
- 仓库入口与常用脚本
|
||||
2. `docs/current/context-reset-brief.md`
|
||||
- 当前状态的最短摘要
|
||||
3. `docs/openclaw/README.md`
|
||||
- OpenClaw 集成文档导航
|
||||
4. `docs/openclaw/openclaw-handoff.md`
|
||||
- OpenClaw 接手总览
|
||||
5. `docs/openclaw/openclaw-orchestration-flow.md`
|
||||
- OpenClaw 正式编排手册
|
||||
6. `docs/design/README.md`
|
||||
- 设计文档导航,区分当前有效设计与背景草案
|
||||
7. `plans/README.md`
|
||||
- 规划文档导航
|
||||
8. `TODO.md`
|
||||
- 当前任务状态
|
||||
|
||||
## 目录结构
|
||||
|
||||
- `docs/README.md`
|
||||
- 文档总索引
|
||||
@@ -8,45 +27,18 @@
|
||||
- 当前状态、收束入口、阶段导航
|
||||
- `docs/design/`
|
||||
- 当前实现的设计文档
|
||||
- `docs/design/README.md`
|
||||
- 设计文档导航
|
||||
- `docs/openclaw/`
|
||||
- OpenClaw 日报聚合与下游对象设计
|
||||
- OpenClaw 集成文档、对象规范与历史归档
|
||||
- `docs/notes/`
|
||||
- 较上层的方案笔记与非最终设计
|
||||
- `docs/notes/README.md`
|
||||
- notes 导航
|
||||
- `docs/archive/`
|
||||
- 历史归档,不作为最新事实来源
|
||||
|
||||
## 当前推荐阅读顺序
|
||||
|
||||
1. `docs/current/context-reset-brief.md`
|
||||
- 当前真实进度与下一步入口
|
||||
2. `docs/openclaw/openclaw-handoff.md`
|
||||
- 给 OpenClaw 的接手说明、环境变量、MCP 调用方式与已知限制
|
||||
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||
- 提供给 OpenClaw 的单篇结构化输入字段说明
|
||||
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||
- 提供给 OpenClaw 的批量投递 envelope 说明
|
||||
5. `docs/design/summary-mcp-service-design.md`
|
||||
- 当前 MCP 服务的职责、接口和边界
|
||||
6. `docs/design/filter-rule-engine-design.md`
|
||||
- 过滤层的输入输出、规则结构与当前实现
|
||||
7. `docs/design/filter-rule-engine-usage.md`
|
||||
- 规则怎么写、怎么跑、结果怎么解读的使用说明
|
||||
8. `docs/design/daily-keyword-index-design.md`
|
||||
- 日报级词元库与周期性词元清洗 skill 设计
|
||||
9. `docs/design/markdown-sink-design.md`
|
||||
- 第一版 Markdown sink 的输入输出、目录结构与落地方式
|
||||
10. `docs/openclaw/openclaw-daily-digest-refactor.md`
|
||||
- 为什么要从单篇入库改成 OpenClaw 日报聚合链路
|
||||
11. `docs/openclaw/article-candidate-daily-digest-schema.md`
|
||||
- `ArticleCandidateRecord`、`OpenClawCandidateInput` 与 `DailyDigest` 的正式设计
|
||||
12. `docs/design/source-schema-design.md`
|
||||
- `source -> item -> document` 的对象设计
|
||||
13. `docs/notes/reading-pipeline-design-notes.md`
|
||||
- 更上层的阅读流方案与阶段划分
|
||||
14. `docs/design/summary-loop-explained.md`
|
||||
- 当前 LLM 摘要校验闭环的解释
|
||||
|
||||
## 当前文档分层
|
||||
## 按主题阅读
|
||||
|
||||
### 1. 当前状态与导航
|
||||
|
||||
@@ -58,11 +50,15 @@
|
||||
- 当前阶段状态的最短摘要
|
||||
- `docs/README.md`
|
||||
- 文档索引与阅读顺序
|
||||
- `plans/README.md`
|
||||
- 规划文档导航
|
||||
|
||||
### 2. 当前实现设计
|
||||
|
||||
- `docs/design/README.md`
|
||||
- 设计文档导航与状态说明
|
||||
- `docs/design/summary-mcp-service-design.md`
|
||||
- 当前内容提取 MCP 的真实设计
|
||||
- 早期 content-extract MCP 设计草案,现主要保留背景参考价值
|
||||
- `docs/design/summary-core-interface-design.md`
|
||||
- 摘要/提取内核的接口抽象
|
||||
- `docs/design/source-schema-design.md`
|
||||
@@ -82,19 +78,23 @@
|
||||
|
||||
### 3. OpenClaw 与下游设计
|
||||
|
||||
- `docs/openclaw/README.md`
|
||||
- OpenClaw 相关文档导航与归档边界
|
||||
- `docs/openclaw/openclaw-handoff.md`
|
||||
- OpenClaw 接手所需的运行说明、工具入口与已知限制
|
||||
- `docs/openclaw/openclaw-orchestration-flow.md`
|
||||
- OpenClaw 编排层的正式运行手册
|
||||
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||
- 提供给 OpenClaw 的单篇结构化输入字段说明
|
||||
- `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||
- 提供给 OpenClaw 的批量投递 envelope 说明
|
||||
- `docs/openclaw/openclaw-daily-digest-refactor.md`
|
||||
- 改造为 OpenClaw 日报聚合链路的原因与目标结构
|
||||
- `docs/openclaw/article-candidate-daily-digest-schema.md`
|
||||
- `ArticleCandidateRecord`、`OpenClawCandidateInput` 与 `DailyDigest` 的字段设计与对象关系
|
||||
|
||||
### 4. 方案笔记
|
||||
|
||||
- `docs/notes/README.md`
|
||||
- notes 导航与使用边界
|
||||
- `docs/notes/reading-pipeline-design-notes.md`
|
||||
- 整体阅读流、规则、sink、push 的方案笔记
|
||||
|
||||
@@ -102,11 +102,23 @@
|
||||
|
||||
- `docs/archive/content-extract-mcp-mvp-archive.md`
|
||||
- MVP 阶段归档,部分状态已被后续进展覆盖
|
||||
- `docs/openclaw/archive/README.md`
|
||||
- OpenClaw 历史文档归档说明
|
||||
- `docs/openclaw/archive/formalization-summary-2026-04-07.md`
|
||||
- 第一阶段正式化总结
|
||||
- `docs/openclaw/archive/openclaw-daily-digest-refactor.md`
|
||||
- 早期日报聚合改造背景
|
||||
- `docs/openclaw/archive/digest-optimization-summary.md`
|
||||
- 早期 digest 优化总结
|
||||
- `docs/openclaw/archive/p1-status-reconciliation-plan-2026-04-14.md`
|
||||
- `resume` / 状态收敛问题的阶段修复计划与回填
|
||||
|
||||
## 当前文档维护原则
|
||||
## 维护原则
|
||||
|
||||
- `docs/current/context-reset-brief.md` 记录当前最新状态
|
||||
- `TODO.md` 记录任务优先级与下一步
|
||||
- `plans/README.md` 负责规划文档分层与导航
|
||||
- `docs/design/README.md` 负责设计文档分层与导航
|
||||
- `outputs/README.md` 记录当前输出目录约定
|
||||
- `docs/archive/content-extract-mcp-mvp-archive.md` 只当历史快照,不再作为最新事实来源
|
||||
- 新增阶段性进展,优先更新 `README.md`、`TODO.md`、`docs/current/context-reset-brief.md`
|
||||
@@ -2,152 +2,104 @@
|
||||
|
||||
## 当前结论
|
||||
|
||||
当前仓库已经具备交付给 OpenClaw 的基础条件。
|
||||
当前仓库已经具备作为 OpenClaw 上游服务的正式基础能力。
|
||||
|
||||
当前主链路是:
|
||||
当前正式主链路是:
|
||||
|
||||
`FreshRSS 未读 -> RSS 内容提取 -> LLM 总结 -> 规则过滤 -> OpenClaw delivery payload`
|
||||
|
||||
OpenClaw 应通过 MCP 工具 `run_freshrss_openclaw_pipeline` 调用这条链路,而不是自行拼接脚本。
|
||||
当前正式控制面已经收口为异步 job:
|
||||
|
||||
## 当前已完成
|
||||
- 主日报:`start_freshrss_pipeline_job -> poll -> get result`
|
||||
- 恢复:`inspect_resume_plan -> start_resume_job -> poll -> get result`
|
||||
- 单篇总结:`start_article_summary_job -> poll -> get result`
|
||||
|
||||
- 已完成 FreshRSS `greader` API 接入与未读拉取
|
||||
- 已完成 FreshRSS 条目到标准化 `item` 的映射
|
||||
- 已完成 RSS-first 提取策略
|
||||
- 已完成 LLM 总结与校验闭环
|
||||
- 已完成规则引擎过滤
|
||||
- 已完成 `ArticleCandidateRecord` 与 `OpenClawCandidateInput` 分层
|
||||
- 已完成 `OpenClawDeliveryPayload` 批量投递结构
|
||||
- 已完成 FreshRSS 已读状态回写
|
||||
- 已完成“仅在最终 payload 成功写盘后再标记已读”的语义
|
||||
- 已完成 MCP 工具 `run_freshrss_openclaw_pipeline`
|
||||
- 已完成默认精简输出模式,减少中间文件
|
||||
- 已完成日报级 `keywords` 词元库与全局词频统计
|
||||
- 已完成 `keyword-cleanup-review` skill 骨架与 review bundle 脚本
|
||||
- 已完成低复杂治理层:`term_cleanup_policy` / `term_watchlist` / `term_change_log`
|
||||
- 已完成采纳建议写回脚本 `scripts/apply_term_suggestions.py`
|
||||
同步 `run_freshrss_openclaw_pipeline`、`resume_run`、`generate_article_summaries` 仍保留,但只用于 debug / fallback。
|
||||
|
||||
## 当前 MCP 工具
|
||||
## 当前权威入口
|
||||
|
||||
当前服务入口:
|
||||
先看这些文档:
|
||||
|
||||
- `src/summary_mcp/server.py`
|
||||
1. `README.md`
|
||||
2. `docs/README.md`
|
||||
3. `docs/openclaw/README.md`
|
||||
4. `docs/openclaw/openclaw-handoff.md`
|
||||
5. `docs/openclaw/openclaw-orchestration-flow.md`
|
||||
6. `plans/README.md`
|
||||
7. `TODO.md`
|
||||
|
||||
当前暴露的 MCP 工具:
|
||||
如果问题是 OpenClaw 集成、状态分支或恢复策略,优先看 `docs/openclaw/`,不要先翻历史计划。
|
||||
|
||||
- `extract_url_content`
|
||||
- `extract_item_content`
|
||||
- `filter_summary_result`
|
||||
- `run_freshrss_openclaw_pipeline`
|
||||
## 当前正式能力
|
||||
|
||||
其中生产主入口是:
|
||||
- FreshRSS 主日报 run 会落地 `run-state.json`
|
||||
- `get_run_status` / `list_runs` / `list_run_artifacts` 提供 run 级观测
|
||||
- `get_delivery_payload` / `get_run_report` 提供正式结果读取
|
||||
- 查询层已经支持 stale state 与终态 artifacts 的状态收敛
|
||||
- 主日报正式启动已切到 async job
|
||||
- `resume` 已切到 async job,并在执行前先做 `inspect_resume_plan`
|
||||
- 生产恢复依赖 `summary/summary-batch.json` 与 `candidates/candidate-batch.json`
|
||||
- 单篇总结也已补齐 async job 形态
|
||||
|
||||
- `run_freshrss_openclaw_pipeline`
|
||||
## 当前关键代码入口
|
||||
|
||||
## 当前关键文件
|
||||
- MCP 服务入口:`src/summary_mcp/server.py`
|
||||
- 主日报 workflow:`src/summary_mcp/workflows/freshrss_pipeline.py`
|
||||
- run / artifact 查询:`src/summary_mcp/runtime/query_service.py`
|
||||
- 主日报 async job:`src/summary_mcp/runtime/freshrss_pipeline_jobs.py`
|
||||
- resume 预检与恢复:`src/summary_mcp/runtime/resume_service.py`
|
||||
- resume async job:`src/summary_mcp/runtime/resume_jobs.py`
|
||||
- 单篇总结 async job:`src/summary_mcp/runtime/article_summary_jobs.py`
|
||||
- 关键词治理:`src/summary_mcp/core/keyword_index.py`
|
||||
|
||||
- MCP 服务入口
|
||||
- `src/summary_mcp/server.py`
|
||||
- FreshRSS 统一工作流
|
||||
- `src/summary_mcp/workflows/freshrss_pipeline.py`
|
||||
- 词元统计核心
|
||||
- `src/summary_mcp/core/keyword_index.py`
|
||||
- 词元统计模型
|
||||
- `src/summary_mcp/models/keyword_index.py`
|
||||
- 摘要循环
|
||||
- `src/summary_mcp/core/summary_loop.py`
|
||||
- 提取主流程
|
||||
- `src/summary_mcp/core/pipeline.py`
|
||||
- FreshRSS 集成
|
||||
- `src/summary_mcp/integrations/freshrss.py`
|
||||
- 规则引擎
|
||||
- `src/summary_mcp/filters/engine.py`
|
||||
- LLM 结果校验
|
||||
- `src/summary_mcp/validators/llm_result.py`
|
||||
- OpenClaw candidate 模型
|
||||
- `src/summary_mcp/models/article_candidate.py`
|
||||
- OpenClaw delivery 模型
|
||||
- `src/summary_mcp/models/openclaw_delivery.py`
|
||||
- 生产脚本入口
|
||||
- `scripts/run_freshrss_pipeline.py`
|
||||
- 词元统计重建脚本
|
||||
- `scripts/build_keyword_index.py`
|
||||
- 词元清洗 skill
|
||||
- `skills/keyword-cleanup-review/SKILL.md`
|
||||
- skill review bundle 脚本
|
||||
- `skills/keyword-cleanup-review/scripts/build_review_bundle.py`
|
||||
- 采纳建议写回脚本
|
||||
- `scripts/apply_term_suggestions.py`
|
||||
- 清洗治理配置
|
||||
- `configs/term_cleanup_policy.json`
|
||||
- `configs/term_watchlist.json`
|
||||
- `configs/term_change_log.json`
|
||||
- OpenClaw 交接说明
|
||||
- `docs/openclaw/openclaw-handoff.md`
|
||||
## 当前核心产物
|
||||
|
||||
## 当前输出规则
|
||||
主日报稳定产物:
|
||||
|
||||
默认生产模式只输出:
|
||||
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/raw/freshrss.raw.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/summary/summary-batch.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/candidates/candidate-batch.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/candidates/digest-brief.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json`
|
||||
|
||||
- `raw/freshrss.raw.json`
|
||||
- `candidates/openclaw-delivery-payload.json`
|
||||
- `run-report.json`
|
||||
job 状态目录:
|
||||
|
||||
同时会更新本地运行数据:
|
||||
- `outputs/freshrss/pipeline_jobs/<job_id>/`
|
||||
- `outputs/freshrss/resume_jobs/<job_id>/`
|
||||
- `outputs/freshrss/article_summary_jobs/<job_id>/`
|
||||
|
||||
关键词运行数据:
|
||||
|
||||
- `data/term_index/daily/YYYY-MM-DD.json`
|
||||
- `data/term_index/term_stats.json`
|
||||
|
||||
如果需要词元清洗审阅输入,可额外生成:
|
||||
## 当前已验证
|
||||
|
||||
- `outputs/term_index/review/keyword-cleanup-bundle.json`
|
||||
- FreshRSS 未读拉取与已读回写可用
|
||||
- RSS-first 提取策略可用
|
||||
- 主日报 MCP 主链路可触发并写出正式产物
|
||||
- 状态查询与结果读取接口可用
|
||||
- stale state / artifacts 收敛逻辑已落地
|
||||
- `resume` 的 artifact-first 判定已落地
|
||||
- `start_resume_job -> poll -> result` 已做本地 synthetic 验证
|
||||
- 单篇总结 async job 可跑通
|
||||
- 关键词 review bundle 与建议写回脚本可用
|
||||
|
||||
如果需要在人工确认后把建议正式写入 watchlist / change log,可使用:
|
||||
## 当前主要限制
|
||||
|
||||
- `scripts/apply_term_suggestions.py`
|
||||
- 某些源 RSS 正文不足时会被直接跳过
|
||||
- 规则仍然偏保守,部分内容会落到 `review`
|
||||
- `paywall` 启发式对中文仍可能误判
|
||||
- 关键词治理还没有接入周期性调度
|
||||
- `digest-brief.json` 仍没有独立 MCP 读取工具
|
||||
- `resume` 目前的剩余主风险不再是恢复点判定,而是缺少真实生产环境的完整恢复验证
|
||||
|
||||
如果需要排障,可开启:
|
||||
## 当前建议
|
||||
|
||||
- `debug_artifacts=true`
|
||||
- 或脚本参数 `--debug-artifacts`
|
||||
|
||||
这样才会额外输出逐条中间文件。
|
||||
|
||||
## 当前验证状态
|
||||
|
||||
已经验证通过:
|
||||
|
||||
- FreshRSS 未读拉取成功
|
||||
- 已读回写成功
|
||||
- MCP 工具入口可直接触发完整链路
|
||||
- 微信公众号样本可直接使用 RSS 提供的 `summary` 内容提取,不再回源抓网页
|
||||
- 精简输出模式已实际跑通
|
||||
- 日报级词元统计已通过离线样例验证,确认别名、停用词、非 `drop` 过滤和 rerun 覆盖逻辑正常
|
||||
- `keyword-cleanup-review` skill 已通过 `quick_validate.py` 结构校验
|
||||
- review bundle 脚本已实际跑通
|
||||
- `apply_term_suggestions.py` 已通过 dry-run 与临时副本写回验证
|
||||
|
||||
## 当前已知限制
|
||||
|
||||
- 当前对 FreshRSS 条目采用 RSS-first 策略,不再回源抓原网页
|
||||
- 如果 RSS 中没有足够正文内容,该条会直接跳过,不会进入后续总结
|
||||
- 某些规则仍偏保守,部分内容可能落到 `review`
|
||||
- `paywall` 相关启发式仍可能误判中文文本
|
||||
- Webhook / 主动投递到 OpenClaw 外部接口尚未实现,当前是由 OpenClaw 通过 MCP 主动调用
|
||||
- 词元清洗 skill 当前已支持“bundle 构建 -> 建议审阅 -> 人工确认写回 watchlist/change_log”,但尚未接入周期性调度
|
||||
- 当前词元统计仍以前置 `OpenClawDeliveryPayload` 作为日报前代理输入,真实 `DailyDigest` 接入后还需切换上游
|
||||
|
||||
## 当前最建议的交接阅读顺序
|
||||
|
||||
1. `README.md`
|
||||
2. `docs/openclaw/openclaw-handoff.md`
|
||||
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||
5. `docs/design/daily-keyword-index-design.md`
|
||||
6. `skills/keyword-cleanup-review/SKILL.md`
|
||||
7. `TODO.md`
|
||||
|
||||
## 一句话结论
|
||||
|
||||
当前仓库已经从“提取 MCP 原型”演进到“可供 OpenClaw 调用的 FreshRSS -> OpenClaw payload 上游处理器”,并已补上第一阶段的日报级词元统计能力和词元清洗 skill 骨架;后续重点转向 skill 周期调度、知识库状态流转和 webhook 接线。
|
||||
- 把 `docs/openclaw/openclaw-orchestration-flow.md` 当成正式编排手册
|
||||
- 把 `docs/openclaw/openclaw-handoff.md` 当成接手总览
|
||||
- 把 `plans/README.md` 当成规划文档导航
|
||||
- 把 `docs/openclaw/archive/` 和 `docs/archive/` 当成历史资料,不要当当前事实源
|
||||
|
||||
@@ -0,0 +1,56 @@
|
||||
# 设计文档导航
|
||||
|
||||
## 使用原则
|
||||
|
||||
`docs/design/` 目录同时包含两类文档:
|
||||
|
||||
- 当前实现仍然有效的设计说明
|
||||
- 早期架构草案和背景设计
|
||||
|
||||
不要默认把这里所有文档都当成当前生产事实。
|
||||
当前生产事实仍以这些入口为准:
|
||||
|
||||
1. `README.md`
|
||||
2. `docs/current/context-reset-brief.md`
|
||||
3. `docs/openclaw/README.md`
|
||||
4. `docs/openclaw/openclaw-handoff.md`
|
||||
5. `docs/openclaw/openclaw-orchestration-flow.md`
|
||||
6. `TODO.md`
|
||||
|
||||
## 当前实现仍然有效
|
||||
|
||||
- `filter-rule-engine-design.md`
|
||||
- 规则过滤层的设计与职责边界
|
||||
- `filter-rule-engine-usage.md`
|
||||
- 规则引擎的使用说明
|
||||
- `daily-keyword-index-design.md`
|
||||
- 关键词索引与清洗治理设计
|
||||
- `summary-loop-explained.md`
|
||||
- LLM 摘要校验闭环说明
|
||||
- `markdown-sink-design.md`
|
||||
- Markdown sink 设计
|
||||
|
||||
## 当前仍有参考价值,但不是生产真相入口
|
||||
|
||||
- `summary-mcp-service-design.md`
|
||||
- 早期 MCP 服务设计草案,部分定位已被后续 workflow service 演进覆盖
|
||||
- `source-schema-design.md`
|
||||
- 更偏对象建模和来源抽象的背景设计
|
||||
- `summary-core-interface-design.md`
|
||||
- 更偏早期摘要内核接口抽象
|
||||
|
||||
## 建议阅读顺序
|
||||
|
||||
如果你是在理解当前实现:
|
||||
|
||||
1. `filter-rule-engine-design.md`
|
||||
2. `filter-rule-engine-usage.md`
|
||||
3. `daily-keyword-index-design.md`
|
||||
4. `summary-loop-explained.md`
|
||||
5. `markdown-sink-design.md`
|
||||
|
||||
如果你是在回看背景设计:
|
||||
|
||||
1. `summary-mcp-service-design.md`
|
||||
2. `source-schema-design.md`
|
||||
3. `summary-core-interface-design.md`
|
||||
@@ -326,8 +326,13 @@ LLM 可以帮助做清洗建议,但不适合直接维护主词元库。
|
||||
|
||||
skill 不直接修改配置文件,而是生成建议文件,例如:
|
||||
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`
|
||||
|
||||
其中建议语义为:
|
||||
|
||||
- JSON 是 review / apply 之间的唯一正式建议产物
|
||||
- Markdown 是人工临时审阅展示稿,不是长期真相来源
|
||||
|
||||
低复杂治理层建议补充三类输入:
|
||||
|
||||
@@ -403,8 +408,9 @@ skill 不直接修改配置文件,而是生成建议文件,例如:
|
||||
2. 程序更新 `daily/YYYY-MM-DD.json`
|
||||
3. 程序更新 `term_stats.json`
|
||||
4. 每周或人工触发一次词元清洗 skill
|
||||
5. skill 输出建议
|
||||
6. 人工确认后再更新配置文件
|
||||
5. skill 生成 suggestions JSON(正式建议产物)
|
||||
6. 如需要人工阅读,再临时生成 Markdown 展示稿
|
||||
7. 人工确认后再更新配置文件
|
||||
|
||||
## 14. 与规则引擎的关系
|
||||
|
||||
@@ -445,7 +451,8 @@ skill 不直接修改配置文件,而是生成建议文件,例如:
|
||||
再补治理层:
|
||||
|
||||
- 增加词元清洗 skill
|
||||
- 输出建议文件
|
||||
- 输出建议文件(以 JSON 为正式产物)
|
||||
- Markdown 仅作为按需生成的人工展示层
|
||||
- 人工确认后更新配置
|
||||
|
||||
### Phase 3
|
||||
@@ -459,3 +466,9 @@ skill 不直接修改配置文件,而是生成建议文件,例如:
|
||||
## 16. 一句话结论
|
||||
|
||||
这套设计选择“只统计日报中的 `keywords`,由程序维护轻量词元库,再由独立 skill 周期性做清洗建议”,目的是在控制数据规模的前提下,为规则配置和长期兴趣演化提供稳定、可审计、可扩展的基础设施。
|
||||
|
||||
补充的产物策略是:
|
||||
|
||||
- facts/state 长期保留
|
||||
- suggestions JSON 作为正式建议产物短期保留
|
||||
- review bundle 与 Markdown 展示稿降级为临时工作文件 / 展示层
|
||||
|
||||
@@ -0,0 +1,151 @@
|
||||
# 关键词清洗流程概述
|
||||
|
||||
> 2026-05-14 初版
|
||||
> 从"数据记录"到"人工确认落盘"的完整链路
|
||||
|
||||
---
|
||||
|
||||
## 整体数据流
|
||||
|
||||
```
|
||||
每日日报 pipeline
|
||||
│
|
||||
▼
|
||||
term_index/daily/YYYY-MM-DD.json ← 每天一篇候选文章的热词统计
|
||||
│
|
||||
▼
|
||||
term_index/term_stats.json ← 所有 daily 的汇总(1070 个词)
|
||||
│
|
||||
├──── build_review_bundle.py ← 打包为审查数据包
|
||||
│ │
|
||||
│ ▼
|
||||
│ review/keyword-cleanup-bundle.json
|
||||
│ │
|
||||
│ ▼
|
||||
│ generate_term_cleanup_suggestions.py
|
||||
│ │
|
||||
│ ▼
|
||||
│ review/term-cleanup-suggestions-YYYY-MM-DD.json ← 正式建议产物
|
||||
│ │
|
||||
│ ▼
|
||||
│ (可选) review/term-cleanup-suggestions-YYYY-MM-DD.md ← 展示稿
|
||||
│
|
||||
├──── 人工确认哪些建议 accept
|
||||
│
|
||||
▼
|
||||
apply_term_suggestions.py ← 写入配置
|
||||
│
|
||||
├── configs/filter_context.personal.json ← interest_keywords
|
||||
├── configs/term_aliases.json ← alias
|
||||
├── configs/term_stopwords.json ← stopword
|
||||
├── configs/term_watchlist.json ← watch
|
||||
└── configs/term_change_log.json ← 变更日志
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 各环节说明
|
||||
|
||||
### 阶段 1:数据记录(每日自动)
|
||||
|
||||
```bash
|
||||
# FreshRSS pipeline 跑完后自动产出
|
||||
data/term_index/daily/2026-05-14.json
|
||||
```
|
||||
|
||||
- 每天一篇,记录当天候选文章中出现的热词
|
||||
- 包含 term、total_count、days_seen 等信息
|
||||
- 目前累计 **41 天**,共 **1070 个独立词**
|
||||
|
||||
### 阶段 2:全量汇总(每日自动)
|
||||
|
||||
```bash
|
||||
data/term_index/term_stats.json
|
||||
```
|
||||
|
||||
- 从所有 daily 文件重建,会覆盖重跑
|
||||
- 按 total_count 排序,前 5 名:OpenClaw(30)、Claude Code(25)、AI Agent(17)、Anthropic(13)、MCP(13)
|
||||
|
||||
### 阶段 3:构建审查数据包(手动触发)
|
||||
|
||||
```bash
|
||||
python skills/keyword-cleanup-review/scripts/build_review_bundle.py \
|
||||
--days 365 \
|
||||
--top 100 \
|
||||
--output outputs/term_index/review/keyword-cleanup-bundle.json
|
||||
```
|
||||
|
||||
- 把 term_stats + 当前配置打成一包,方便后续处理
|
||||
- 输出:`review/keyword-cleanup-bundle.json`
|
||||
|
||||
### 阶段 4:生成建议(手动触发)
|
||||
|
||||
```bash
|
||||
python scripts/generate_term_cleanup_suggestions.py \
|
||||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
|
||||
--emit-markdown
|
||||
```
|
||||
|
||||
#### 当前产出能力
|
||||
|
||||
| 建议类型 | 状态 | 当前阈值 | 说明 |
|
||||
|---------|------|----------|------|
|
||||
| interest_keyword_suggestions | ✅ **已实现** | total≥3, days≥2 | 产出 20 条 |
|
||||
| watch_terms | ✅ **已实现** | total≤2, days≤2 | 本次 0 条 |
|
||||
| alias_suggestions | ❌ **硬编码为空** | policy 有阈值(total≥2, days≥2)但脚本未实现 | |
|
||||
| stopword_suggestions | ❌ **硬编码为空** | policy 有阈值(total≤2, days≤2)但脚本未实现 | |
|
||||
|
||||
**关键发现:** alias 和 stopword 不是"阈值太保守",是 **generate 脚本里压根没写对应的生成函数**。policy 文件里阈值已经配好了(alias: min_total=2/min_days=2,stopword: max_total=2/max_days=2),但脚本第 376-380 行直接硬编码为 `[]` 和 `0`。
|
||||
|
||||
### 阶段 5:人工确认(手动)
|
||||
|
||||
```
|
||||
OpenClaw 把建议列给你 → 你确认哪些 accept → 我执行 apply
|
||||
```
|
||||
|
||||
本次模式:
|
||||
- 高频(≥5次/5天以上)→ 强烈推荐 ✅
|
||||
- 中频(3-4次)→ 附带建议 ✅
|
||||
- 泛词 → 建议跳过 ❌
|
||||
|
||||
### 阶段 6:落盘配置(手动)
|
||||
|
||||
```bash
|
||||
python scripts/apply_term_suggestions.py \
|
||||
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
|
||||
--accept-interest 词1 词2 ...
|
||||
```
|
||||
|
||||
- dry-run 预览 → 确认后正式 apply
|
||||
- 写入 `configs/filter_context.personal.json`
|
||||
- 同步记录到 `term_change_log.json`
|
||||
- **不备份原始配置**(待优化)
|
||||
- **apply 后不自动清理 review 目录**(待优化)
|
||||
|
||||
### 阶段 7:维护清理(按需)
|
||||
|
||||
由 OpenClaw 侧 `reader-keyword-maintenance` skill 处理:
|
||||
- 删除旧 markdown 展示稿
|
||||
- 保留最近一份 bundle
|
||||
- 保守保留 suggestions JSON
|
||||
|
||||
---
|
||||
|
||||
## 当前配置资产
|
||||
|
||||
| 文件 | 内容 | 数据量 |
|
||||
|------|------|--------|
|
||||
| `filter_context.personal.json` | interest_keywords | 52 个 |
|
||||
| `term_aliases.json` | 别名映射 | 0 组(未启用) |
|
||||
| `term_stopwords.json` | 停用词 | 0 个(未启用) |
|
||||
| `term_watchlist.json` | 观察词 | 6 个 |
|
||||
| `term_change_log.json` | 所有变更记录 | 已记录 |
|
||||
|
||||
---
|
||||
|
||||
## 待优化项
|
||||
|
||||
1. **alias/stopword 建议生成为空** — generate 脚本硬编码缺实现,policy 已有阈值,需要补函数
|
||||
2. **apply 前无配置备份** — 建议 apply 前自动 cp 备份
|
||||
3. **apply 后无自动收尾** — 建议 apply 后自动删旧 markdown 和 bundle
|
||||
4. **alias 识别依赖规则而非 LLM** — 当前全靠统计阈值,无法做语义级判断(如中英文映射、缩写展开)。如果需要高级 alias 识别,可以用 LLM 生成候选,规则脚本做 apply
|
||||
@@ -1,5 +1,16 @@
|
||||
# Source Schema 设计草案
|
||||
|
||||
## 状态说明
|
||||
|
||||
本文件偏对象建模和来源抽象,主要用于解释早期 schema 设计思路。
|
||||
|
||||
它不是当前生产运行手册,也不是当前 workflow service 的唯一真相来源。
|
||||
如果你关注当前 OpenClaw 集成或运行状态,应优先看:
|
||||
|
||||
- `docs/current/context-reset-brief.md`
|
||||
- `docs/openclaw/README.md`
|
||||
- `docs/openclaw/openclaw-orchestration-flow.md`
|
||||
|
||||
## 1. 文档目的
|
||||
|
||||
本文档用于定义阅读流系统中的来源与内容对象模型,目标是把“来源分类”的讨论收敛成一套可执行的数据结构,供后续的抓取、摘要、过滤、入库和推送流程统一使用。
|
||||
|
||||
@@ -1,5 +1,17 @@
|
||||
# Summary Core Interface 设计草案
|
||||
|
||||
## 状态说明
|
||||
|
||||
本文件记录的是较早期的摘要内核接口抽象。
|
||||
|
||||
它更适合用于理解背景设计,不应直接当成当前生产接口契约。
|
||||
当前接口与编排真相请优先看:
|
||||
|
||||
- `README.md`
|
||||
- `docs/current/context-reset-brief.md`
|
||||
- `docs/openclaw/README.md`
|
||||
- `docs/openclaw/openclaw-handoff.md`
|
||||
|
||||
## 1. 文档目的
|
||||
|
||||
本文档用于定义 `summary-core` 的输入输出接口,目标是把“页面摘要能力”从概念讨论收敛成一套稳定、可复用、可封装的数据接口。
|
||||
|
||||
@@ -1,5 +1,18 @@
|
||||
# Content Extract MCP Service 设计草案
|
||||
|
||||
## 状态说明
|
||||
|
||||
本文件主要记录早期 “content extract MCP” 的设计抽象。
|
||||
|
||||
它仍有背景参考价值,但不是当前生产事实入口。
|
||||
当前生产能力已经演进为更完整的 workflow service,正式口径请优先看:
|
||||
|
||||
- `README.md`
|
||||
- `docs/current/context-reset-brief.md`
|
||||
- `docs/openclaw/README.md`
|
||||
- `docs/openclaw/openclaw-handoff.md`
|
||||
- `docs/openclaw/openclaw-orchestration-flow.md`
|
||||
|
||||
## 1. 文档目的
|
||||
|
||||
本文档用于定义当前仓库中已经落地的 MCP 服务设计,即“内容提取 MCP”。
|
||||
@@ -12,7 +25,7 @@
|
||||
- validator 与 LLM 摘要如何接在 MCP 之后
|
||||
- 当前 MVP 已完成到哪一层
|
||||
|
||||
这份文档描述的是当前真实实现,而不是早期“摘要 MCP”设想。
|
||||
这份文档主要记录当时实现阶段的设计取向,而不是当前生产阶段的唯一事实来源。
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -0,0 +1,22 @@
|
||||
# Notes 导航
|
||||
|
||||
`docs/notes/` 保存的是更早期、讨论型、背景型方案笔记。
|
||||
|
||||
这些文档的用途是:
|
||||
|
||||
- 理解项目最初的问题空间
|
||||
- 回看为什么会形成现在的对象分层和流程划分
|
||||
|
||||
这些文档不是当前生产事实来源。
|
||||
|
||||
当前如需判断“现在到底怎么跑”,优先看:
|
||||
|
||||
- `README.md`
|
||||
- `docs/current/context-reset-brief.md`
|
||||
- `docs/openclaw/README.md`
|
||||
- `docs/openclaw/openclaw-orchestration-flow.md`
|
||||
|
||||
当前 notes:
|
||||
|
||||
- `reading-pipeline-design-notes.md`
|
||||
- 早期阅读流方案讨论纪要
|
||||
@@ -0,0 +1,36 @@
|
||||
# OpenClaw 文档导航
|
||||
|
||||
## 当前有效文档
|
||||
|
||||
- `openclaw-handoff.md`
|
||||
- 面向接手者的总览文档
|
||||
- 说明 reader 的职责边界、MCP 工具面、环境变量和正式集成约束
|
||||
- `openclaw-orchestration-flow.md`
|
||||
- 面向 OpenClaw 编排层的正式运行手册
|
||||
- 说明启动、轮询、读结果、恢复和人工介入的标准动作
|
||||
- `openclaw-candidate-input-field-spec.md`
|
||||
- 单篇 `OpenClawCandidateInput` 字段规范
|
||||
- `openclaw-delivery-payload-spec.md`
|
||||
- 批量 `OpenClawDeliveryPayload` 字段规范
|
||||
- `article-candidate-daily-digest-schema.md`
|
||||
- 对象分层设计说明
|
||||
- 用于理解 `ArticleCandidateRecord` / `OpenClawCandidateInput` / `DailyDigest` 的关系
|
||||
|
||||
## 当前推荐阅读顺序
|
||||
|
||||
1. `openclaw-handoff.md`
|
||||
2. `openclaw-orchestration-flow.md`
|
||||
3. `openclaw-candidate-input-field-spec.md`
|
||||
4. `openclaw-delivery-payload-spec.md`
|
||||
5. `article-candidate-daily-digest-schema.md`
|
||||
|
||||
## 归档说明
|
||||
|
||||
`archive/` 下的文档保留历史决策、阶段总结和排障规划,但不再作为当前事实来源。
|
||||
|
||||
当前已归档:
|
||||
|
||||
- `archive/formalization-summary-2026-04-07.md`
|
||||
- `archive/openclaw-daily-digest-refactor.md`
|
||||
- `archive/digest-optimization-summary.md`
|
||||
- `archive/p1-status-reconciliation-plan-2026-04-14.md`
|
||||
@@ -0,0 +1,9 @@
|
||||
# OpenClaw 历史归档
|
||||
|
||||
本目录只保留阶段性总结、设计演进记录和排障计划。
|
||||
|
||||
使用原则:
|
||||
|
||||
- 需要了解“为什么会这样设计”时再看
|
||||
- 不要把这里的描述当成当前生产事实
|
||||
- 当前正式口径以 `docs/openclaw/README.md`、`docs/openclaw/openclaw-handoff.md`、`docs/openclaw/openclaw-orchestration-flow.md` 为准
|
||||
@@ -0,0 +1,314 @@
|
||||
# Digest Optimization Summary
|
||||
|
||||
## 背景
|
||||
|
||||
reader → OpenClaw 日报链路原先的问题主要有两类:
|
||||
|
||||
1. **OpenClaw public digest 输入过重**
|
||||
- public digest 直接读取完整 `openclaw-delivery-payload.json`
|
||||
- 其中混有大量不直接服务公开日报的字段
|
||||
- public digest 这一步在 OpenClaw 侧消耗了较多 token
|
||||
|
||||
2. **public / internal 生成逻辑没有充分拆分**
|
||||
- public digest 与 internal review digest 都基于完整 payload 推导
|
||||
- 容易造成重复消耗
|
||||
- public digest 还可能被 review / 内部流程语义污染
|
||||
|
||||
本轮优化的目标不是重写 reader 主流程,而是在不破坏现有 delivery payload 的前提下,先把 public digest 的输入和生成方式收敛下来,并验证整体 token 与内容质量的变化。
|
||||
|
||||
---
|
||||
|
||||
## 本轮改动
|
||||
|
||||
### 1. reader 新增 `digest-brief.json`
|
||||
|
||||
在 FreshRSS pipeline 写出:
|
||||
|
||||
- `outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json`
|
||||
|
||||
之后,额外生成:
|
||||
|
||||
- `outputs/freshrss/rerun/<run-id>/candidates/digest-brief.json`
|
||||
|
||||
用途:
|
||||
|
||||
- 供 OpenClaw 生成 **public digest** 时优先读取
|
||||
- 作为 public-only 的轻量输入视图
|
||||
|
||||
当前约束:
|
||||
|
||||
- 仅保留 `selection_decision == "keep"` 的候选
|
||||
- 默认最多保留前 5 条
|
||||
- 高亮 `highlights` 最多保留 3 条
|
||||
- schema 标识为 `digest-brief.v1`
|
||||
|
||||
保留字段:
|
||||
|
||||
- `title`
|
||||
- `source_name`
|
||||
- `summary`
|
||||
- `highlights`
|
||||
- `category`
|
||||
- `digest_rank`
|
||||
- `selection_decision`
|
||||
- `url`
|
||||
|
||||
附带计数:
|
||||
|
||||
- `source_candidate_count`
|
||||
- `candidate_count`
|
||||
|
||||
---
|
||||
|
||||
### 2. public / internal 输入边界拆分
|
||||
|
||||
当前推荐口径:
|
||||
|
||||
- **public digest**
|
||||
- 优先读取 `digest-brief.json`
|
||||
- 仅使用 public-only 输入视图
|
||||
|
||||
- **internal review digest**
|
||||
- 继续读取完整 `openclaw-delivery-payload.json`
|
||||
- 保留 keep / review 的决策上下文
|
||||
|
||||
这样做的原因:
|
||||
|
||||
- public digest 需要更轻、更干净的公开输入
|
||||
- internal review digest 仍然需要完整上下文来支撑判断、待确认与建议沉淀
|
||||
|
||||
---
|
||||
|
||||
### 3. digest 成稿的正式落盘位置
|
||||
|
||||
正式 run 生成出的 digest 成稿,不应停留在 OpenClaw 的临时目录,而应回写到同一次 reader run 目录下。
|
||||
|
||||
推荐正式产物位置:
|
||||
|
||||
- `outputs/freshrss/rerun/<run-id>/digest/public_digest.md`
|
||||
- `outputs/freshrss/rerun/<run-id>/digest/internal_review_digest.md`
|
||||
- `outputs/freshrss/rerun/<run-id>/digest/combined.json`
|
||||
|
||||
这样可以保证:
|
||||
|
||||
- 一次 run 的所有输入、输出、摘要结果和日报成稿都收在同一目录下
|
||||
- 后续 Hugo 发布与 chat 回传基于同一组正式产物,而不是临时文件
|
||||
- 便于回溯、复盘和后续自动化收口
|
||||
|
||||
### 4. 一次生成两份 digest 的生成模式
|
||||
|
||||
推荐把:
|
||||
|
||||
- `public digest`
|
||||
- `internal review digest`
|
||||
|
||||
改为在 OpenClaw 侧 **一次调用同时生成两份**。
|
||||
|
||||
推荐输入:
|
||||
|
||||
- `PUBLIC_DIGEST_INPUT` → `digest-brief.json`
|
||||
- `INTERNAL_REVIEW_INPUT` → `openclaw-delivery-payload.json`
|
||||
|
||||
推荐输出:
|
||||
|
||||
```json
|
||||
{
|
||||
"public_digest_markdown": "...",
|
||||
"internal_review_digest_markdown": "..."
|
||||
}
|
||||
```
|
||||
|
||||
这样可以减少重复 prompt / 调用开销,同时保留 public / internal 两种视图的边界。
|
||||
|
||||
---
|
||||
|
||||
### 5. internal review digest 风格收敛
|
||||
|
||||
internal review digest 经过一轮人工验证后,收敛成以下规则:
|
||||
|
||||
固定结构:
|
||||
|
||||
1. `今日候选概况`
|
||||
2. `已入选重点`
|
||||
3. `待你确认`
|
||||
4. `建议沉淀到 IMA`
|
||||
5. `原始候选清单`
|
||||
|
||||
表达规则:
|
||||
|
||||
- 不显示 `rank`
|
||||
- 不显示英文 machine state
|
||||
- 使用中文状态:
|
||||
- `keep` → `已入选`
|
||||
- `review` → `待确认`
|
||||
- `drop` → `暂不纳入`
|
||||
|
||||
内容规则:
|
||||
|
||||
- `已入选重点`
|
||||
- 标题 + 来源
|
||||
- 状态
|
||||
- 较完整的一段摘要
|
||||
- 一段判断(解释为什么值得入选,以及它在今天 digest 中承担什么角色)
|
||||
|
||||
- `待你确认`
|
||||
- 标题 + 来源
|
||||
- 状态
|
||||
- 较完整的一段摘要
|
||||
- 原因
|
||||
- 建议
|
||||
|
||||
- `原始候选清单`
|
||||
- 也用中文状态,而不是 `decision=keep/review`
|
||||
|
||||
目标:
|
||||
|
||||
- 保留 internal review digest 作为“给人看的审阅稿”的属性
|
||||
- 避免它沦为 payload 的原样转写或机器中间态展示
|
||||
|
||||
---
|
||||
|
||||
### 6. public digest 风格收敛
|
||||
|
||||
public digest 当前推荐结构:
|
||||
|
||||
1. `今日概览`
|
||||
2. `今日重点`
|
||||
3. `趋势观察`
|
||||
4. `延伸阅读`
|
||||
5. `信息来源`
|
||||
|
||||
表达规则:
|
||||
|
||||
- 不暴露 internal workflow 词汇
|
||||
- 不写 `待确认` / `建议沉淀到 IMA` / `keep/review/drop` / `selection_decision`
|
||||
- 保持适合 Hugo 公开浏览的表达方式
|
||||
|
||||
内容规则:
|
||||
|
||||
- 每个 `今日重点` 条目除了摘要和 highlights 外,增加一句编辑性总结
|
||||
- 推荐形式:
|
||||
- `这篇内容更值得关注的原因在于……`
|
||||
|
||||
目标:
|
||||
|
||||
- 保证 public digest 不只是“摘要列表”
|
||||
- 而是一份带有编辑性提炼的公开日报
|
||||
|
||||
---
|
||||
|
||||
## 实测结果
|
||||
|
||||
基于真实 run:
|
||||
|
||||
- run 目录:`outputs/freshrss/rerun/20260401-074614`
|
||||
|
||||
### 1. public 输入压缩效果
|
||||
|
||||
- 完整 payload:`8701` 字符
|
||||
- `digest-brief.json`:`3080` 字符
|
||||
|
||||
压缩比例:
|
||||
|
||||
- **减少约 64.6%**
|
||||
|
||||
按中位 token 粗估:
|
||||
|
||||
- 完整 payload:约 `3955 tokens`
|
||||
- public brief:约 `1400 tokens`
|
||||
|
||||
public 输入侧单次大约减少:
|
||||
|
||||
- **约 2500 tokens**
|
||||
|
||||
---
|
||||
|
||||
### 2. 一次生成两份的总成本估算
|
||||
|
||||
基于真实输入输出的中位估算:
|
||||
|
||||
- 旧方案(两次生成):约 `10328 tokens`
|
||||
- 新方案(一次生成两份):约 `7734 tokens`
|
||||
|
||||
节省:
|
||||
|
||||
- **约 2594 tokens**
|
||||
- **约 25%**
|
||||
|
||||
说明:
|
||||
|
||||
- 第一步 public 输入瘦身带来的是“输入量级下降”
|
||||
- 第二步一次生成两份带来的是“调用层重复开销下降”
|
||||
- 两者叠加后,已经形成比较明显的成本优化效果
|
||||
|
||||
---
|
||||
|
||||
## 当前默认口径
|
||||
|
||||
### public digest
|
||||
|
||||
- 输入:`digest-brief.json`
|
||||
- 风格:公开浏览稿
|
||||
- 每个重点项包含:
|
||||
- 摘要
|
||||
- 关键信号
|
||||
- 一句编辑性总结
|
||||
|
||||
### internal review digest
|
||||
|
||||
- 输入:完整 payload
|
||||
- 风格:内部审阅稿
|
||||
- 每个重点项包含:
|
||||
- 更完整摘要
|
||||
- 判断
|
||||
- 每个待确认项包含:
|
||||
- 更完整摘要
|
||||
- 原因
|
||||
- 建议
|
||||
|
||||
### 生成方式
|
||||
|
||||
- 优先采用 **一次调用同时生成两份**
|
||||
|
||||
---
|
||||
|
||||
## 当前阶段结论
|
||||
|
||||
本轮优化已经形成一个可用版本:
|
||||
|
||||
- reader 新增 public-only 轻量输入视图
|
||||
- public / internal 边界清楚
|
||||
- internal 风格和 public 风格都收敛到了可接受版本
|
||||
- token 成本下降有明确实测支撑
|
||||
- skill 文档与流程规范已经同步更新
|
||||
|
||||
当前更适合的策略不是继续抽象设计,而是:
|
||||
|
||||
- 按这套新流程再跑几次真实日报
|
||||
- 观察稳定性、质量波动和实际使用感受
|
||||
|
||||
---
|
||||
|
||||
## 后续可选方向
|
||||
|
||||
### 1. internal 输入进一步轻量化
|
||||
|
||||
潜在方向:
|
||||
|
||||
- 新增一个 internal 专用的轻量视图
|
||||
- 但需要谨慎,避免削弱 internal review 的判断价值
|
||||
|
||||
### 2. digest 阶段模型分层
|
||||
|
||||
潜在方向:
|
||||
|
||||
- reader 上游继续用便宜模型做抽取和结构化
|
||||
- OpenClaw digest 阶段单独切到更便宜或更合适的模型
|
||||
|
||||
### 3. 自动化执行收口
|
||||
|
||||
潜在方向:
|
||||
|
||||
- 把“一次生成两份”的逻辑进一步标准化
|
||||
- 更顺滑地接 Hugo 发布与聊天回传
|
||||
- 让 reader → OpenClaw → Hugo / chat 的路径更接近真正的稳定生产流程
|
||||
@@ -0,0 +1,163 @@
|
||||
# reader MCP workflow service formalization summary (2026-04-07)
|
||||
|
||||
## Overview
|
||||
|
||||
On 2026-04-07, the reader project was formally advanced from a script-first integration model into a reader-centric MCP workflow service model.
|
||||
|
||||
The key shift is:
|
||||
|
||||
- before: OpenClaw primarily relied on long CLI / exec flows and direct output-path stitching
|
||||
- now: reader exposes a formal workflow-oriented MCP surface with run-state, status queries, result reads, and minimal recovery
|
||||
|
||||
This document records the main outcomes and commits for the first formalization phase.
|
||||
|
||||
---
|
||||
|
||||
## Completed capability set
|
||||
|
||||
### 1. Run-state persistence
|
||||
|
||||
Commit:
|
||||
|
||||
- `72a6853` — `Add run-state persistence for FreshRSS pipeline`
|
||||
|
||||
Delivered:
|
||||
|
||||
- `run-state.json`
|
||||
- `RunState / StageState / ArtifactRecord`
|
||||
- stage-level state persistence for the FreshRSS pipeline
|
||||
|
||||
### 2. Architecture / implementation docs
|
||||
|
||||
Commit:
|
||||
|
||||
- `7563aa8` — `docs: add reader MCP architecture and implementation plan`
|
||||
|
||||
Delivered:
|
||||
|
||||
- architecture design
|
||||
- implementation plan
|
||||
- TODO-driven collaboration model
|
||||
|
||||
### 3. MCP run-status query tools
|
||||
|
||||
Commit:
|
||||
|
||||
- `d9173fb` — `feat: add MCP run status query tools`
|
||||
|
||||
Delivered:
|
||||
|
||||
- `get_run_status`
|
||||
- `list_runs`
|
||||
- `list_run_artifacts`
|
||||
|
||||
### 4. MCP result-read tools
|
||||
|
||||
Commit:
|
||||
|
||||
- `4a02894` — `Add MCP delivery payload and run report queries`
|
||||
|
||||
Delivered:
|
||||
|
||||
- `get_delivery_payload`
|
||||
- `get_run_report`
|
||||
|
||||
### 5. Minimal resume design
|
||||
|
||||
Commit:
|
||||
|
||||
- `d91cdbc` — `docs: narrow resume_run minimal recovery design`
|
||||
|
||||
Delivered:
|
||||
|
||||
- narrowed design for `resume_run`
|
||||
- explicit supported / unsupported recovery points
|
||||
|
||||
### 6. Minimal `resume_run`
|
||||
|
||||
Commit:
|
||||
|
||||
- `c622bc6` — `Implement minimal resume_run for freshrss runs`
|
||||
|
||||
Delivered:
|
||||
|
||||
- minimal `resume_run`
|
||||
- supports only freshrss runs with `run-state.json`
|
||||
- supports only recent resumable points
|
||||
- explicitly rejects `fetch_feed` and `extract_articles`
|
||||
|
||||
### 7. Formal handoff / workflow docs
|
||||
|
||||
Commits:
|
||||
|
||||
- `4f219ef` — `docs: formalize reader MCP workflow service handoff`
|
||||
- `2df0af5` — `docs: add openclaw orchestration flow for reader MCP`
|
||||
|
||||
Delivered:
|
||||
|
||||
- formal handoff aligned to actual implementation
|
||||
- OpenClaw orchestration runbook
|
||||
- explicit rule that OpenClaw should stop hand-stitching reader paths in the normal production flow
|
||||
|
||||
---
|
||||
|
||||
## Current formal MCP workflow surface
|
||||
|
||||
The current reader MCP workflow surface now includes:
|
||||
|
||||
- `run_freshrss_openclaw_pipeline`
|
||||
- `get_run_status`
|
||||
- `list_runs`
|
||||
- `list_run_artifacts`
|
||||
- `get_delivery_payload`
|
||||
- `get_run_report`
|
||||
- `resume_run` (minimal version)
|
||||
|
||||
---
|
||||
|
||||
## Current boundary
|
||||
|
||||
reader is now the upstream workflow engine for:
|
||||
|
||||
- FreshRSS pull
|
||||
- extraction
|
||||
- summary
|
||||
- filter
|
||||
- payload generation
|
||||
- run-state persistence
|
||||
- result read
|
||||
- minimal recovery
|
||||
|
||||
OpenClaw / skill remains responsible for:
|
||||
|
||||
- digest markdown generation
|
||||
- Hugo publishing
|
||||
- chat reporting
|
||||
- user confirmation
|
||||
- IMA orchestration
|
||||
|
||||
---
|
||||
|
||||
## Current limitations
|
||||
|
||||
The first formalization phase is complete, but some constraints remain:
|
||||
|
||||
- `resume_run` is still minimal and does not support arbitrary stage re-entry
|
||||
- historical runs without `run-state.json` are not formally recoverable
|
||||
- some very old runs may still require conservative artifact/path discovery
|
||||
- `rerun_stage` is not implemented
|
||||
- deeper runtime consolidation of `run_freshrss_openclaw_pipeline` can still be improved later
|
||||
|
||||
---
|
||||
|
||||
## Practical conclusion
|
||||
|
||||
The reader project should now be treated as a formal MCP workflow service rather than as a long-running CLI-first integration point.
|
||||
|
||||
For normal production orchestration:
|
||||
|
||||
- start via MCP
|
||||
- observe via MCP status tools
|
||||
- read results via MCP result tools
|
||||
- use `resume_run` only within the documented minimal recovery range
|
||||
- keep CLI for debug / fallback only
|
||||
@@ -0,0 +1,367 @@
|
||||
# Reader 日报链路 P1 状态收敛问题:规划与修复清单(2026-04-14)
|
||||
|
||||
## 背景
|
||||
|
||||
在 2026-04-14 的 reader 日报正式运行中,出现了以下现象:
|
||||
|
||||
- `openclaw-delivery-payload.json`、`digest-brief.json`、`run-report.json` 已真实落盘
|
||||
- 但 `get_freshrss_pipeline_job_status` / `get_run_status` 仍可能显示:
|
||||
- `running`
|
||||
- `failed`
|
||||
- 或 `current_stage=generate_summaries`
|
||||
- `resume_run` 在这种状态下可能直接超时
|
||||
|
||||
这说明当前 reader 的**状态层(job/run-state)**与**产物层(artifacts/report)**之间没有稳定收敛。
|
||||
|
||||
---
|
||||
|
||||
## 本次确认的核心结论
|
||||
|
||||
### 1. job status 与 run status 是两套独立状态系统
|
||||
|
||||
- **job 层状态**:`src/summary_mcp/runtime/freshrss_pipeline_jobs.py`
|
||||
- `start_freshrss_pipeline_job()`
|
||||
- `run_freshrss_pipeline_job()`
|
||||
- `get_freshrss_pipeline_job_status()`
|
||||
- 状态文件位于:`outputs/freshrss/pipeline_jobs/<job_id>/run-state.json`
|
||||
- 只有 4 个粗粒度 stage:
|
||||
- `prepare_job`
|
||||
- `load_input`
|
||||
- `run_pipeline`
|
||||
- `write_result`
|
||||
|
||||
- **run 层状态**:`src/summary_mcp/workflows/freshrss_pipeline.py`
|
||||
- `run_freshrss_pipeline()`
|
||||
- 由 `src/summary_mcp/runtime/query_service.py:get_run_status()` 查询
|
||||
- 状态文件位于:`outputs/freshrss/rerun/<run_dir>/run-state.json`
|
||||
- 包含 6 个细粒度 stage:
|
||||
- `fetch_feed`
|
||||
- `extract_articles`
|
||||
- `generate_summaries`
|
||||
- `apply_filters`
|
||||
- `build_delivery_payload`
|
||||
- `write_run_report`
|
||||
|
||||
**问题:** 两套状态没有统一收敛规则,用户可以同时看到两套不同口径的“当前进度”。
|
||||
|
||||
---
|
||||
|
||||
### 2. 查询层目前优先信 run-state,不会用 artifacts / run-report 纠偏
|
||||
|
||||
代码位置:`src/summary_mcp/runtime/query_service.py`
|
||||
|
||||
关键行为:
|
||||
- `_resolve_run_record()` 只要发现 `run-state.json` 存在,就优先使用 `RunStore.load(...)`
|
||||
- 即使 `run-report.json`、`delivery_payload`、`digest_brief` 已存在,也不会自动纠偏状态
|
||||
|
||||
**结果:**
|
||||
- 一旦 `run-state.json` 因中断、超时、外层 SIGTERM 或写回未完成而停留在旧值
|
||||
- `get_run_status()` 就会持续返回过期状态
|
||||
- 造成“产物已完成,但状态仍显示 running/failed/卡在 summary”的错觉
|
||||
|
||||
---
|
||||
|
||||
### 3. `generate_summaries` 假卡住,本质上更像 stale state,不像真实业务卡住
|
||||
|
||||
代码位置:`src/summary_mcp/workflows/freshrss_pipeline.py`
|
||||
|
||||
从执行顺序看:
|
||||
1. `start_stage(generate_summaries)`
|
||||
2. summary 循环
|
||||
3. `finish_stage(generate_summaries)`
|
||||
4. `start_stage(apply_filters)`
|
||||
5. `finish_stage(apply_filters)`
|
||||
6. `start_stage(build_delivery_payload)`
|
||||
7. 写 payload / digest brief
|
||||
8. `finish_stage(build_delivery_payload)`
|
||||
9. `start_stage(write_run_report)`
|
||||
10. 写 run-report
|
||||
11. `finish_stage(write_run_report)`
|
||||
12. `finish_run(...)`
|
||||
|
||||
**判断:**
|
||||
如果 payload / digest brief / run-report 都已经存在,那么“仍显示卡在 `generate_summaries`”更可能是:
|
||||
- `run-state.json` 没来得及写回最终状态
|
||||
- 或查询时读到了旧状态
|
||||
|
||||
而不是 summary 阶段真实没有跑过去。
|
||||
|
||||
---
|
||||
|
||||
### 4. `resume_run` 不是轻量恢复,而是同步继续跑工作流
|
||||
|
||||
代码位置:`src/summary_mcp/runtime/resume_service.py`
|
||||
|
||||
关键行为:
|
||||
- `resume_run()` 会根据 `resume_from_stage` 直接继续执行:
|
||||
- `_run_summary_stage(...)`
|
||||
- `_run_filter_stage(...)`
|
||||
- `_run_delivery_stage(...)`
|
||||
- `_run_report_stage(...)`
|
||||
|
||||
这意味着它不是“修状态”的工具,而是“同步继续跑剩余工作流”的工具。
|
||||
|
||||
**问题:**
|
||||
- 如果 stale state 把 `resume_from_stage` 定在 `generate_summaries`
|
||||
- 那么 `resume_run` 会从一个过早阶段重新跑
|
||||
- 在 MCP 包装层下非常容易超时
|
||||
|
||||
---
|
||||
|
||||
## 问题分类
|
||||
|
||||
### A. 真实 bug
|
||||
|
||||
1. **查询层过度信任 stale `run-state.json`**
|
||||
- 文件:`src/summary_mcp/runtime/query_service.py`
|
||||
- 影响:产物已完成但状态仍错误
|
||||
|
||||
2. **`resume_run` 过度依赖 stale `current_stage` / recovery 信息**
|
||||
- 文件:`src/summary_mcp/runtime/resume_service.py`
|
||||
- 影响:从过早阶段重跑,放大 timeout 风险
|
||||
|
||||
### B. 状态设计缺陷
|
||||
|
||||
3. **job 层与 run 层两套状态源没有统一收敛规则**
|
||||
- 文件:`src/summary_mcp/runtime/freshrss_pipeline_jobs.py`
|
||||
- 文件:`src/summary_mcp/runtime/query_service.py`
|
||||
- 影响:用户看到两个互相打架的状态解释
|
||||
|
||||
4. **状态系统完全依赖显式写回,不会按产物反推修正**
|
||||
- 文件:`src/summary_mcp/runtime/run_store.py`
|
||||
- 影响:一旦中断,状态比产物更容易脏
|
||||
|
||||
### C. 调用层误判
|
||||
|
||||
5. **把 `resume_run` 当成轻量恢复接口使用**
|
||||
- 实际上它更接近“同步恢复执行器”
|
||||
- 影响:在长链路场景下超时是高概率事件
|
||||
|
||||
---
|
||||
|
||||
## 修复目标
|
||||
|
||||
## 当前落地状态(回填)
|
||||
|
||||
- [x] Phase 1 已落地:`get_run_status()` 会基于 `run-report.json` 与关键产物做终态收敛,并暴露 `status_source` / `state_conflict`
|
||||
- [x] Phase 2 已落地第一阶段:`resume_run()` 会拒绝对已有终态 `run-report.json` 的 run 继续恢复
|
||||
- [x] Phase 2 已继续增强:恢复起点现在会优先根据 artifacts 重算,而不是直接盲信 `run-state.recovery.resume_from_stage`
|
||||
- [x] 新增 `inspect_resume_plan(run_id)` 作为恢复前置判定接口,避免调用方用 `resume_run` 探路
|
||||
- [x] Phase 2 已补齐生产恢复 artifacts:正式 run 会稳定写出 `summary/summary-batch.json` 与 `candidates/candidate-batch.json`,`resume_run` / `inspect_resume_plan` 会优先使用它们,而不是依赖 debug per-item 文件
|
||||
- [x] Phase 3 已落地:job 状态与结果读取会基于 linked run 做收敛,避免 outer job stale state 卡住编排
|
||||
|
||||
### 一级目标(必须达成)
|
||||
|
||||
1. 当 `run-report.json` / `delivery_payload` / `digest_brief` 已存在时,`get_run_status()` 不应继续盲目展示明显过期的 stage 状态;对调用方暴露的 `status` 必须直接收敛为可用终态,而不是只附加 hint
|
||||
2. 当状态层与产物层冲突时,查询结果必须显式标注“状态冲突 / stale state”
|
||||
3. `resume_run()` 在恢复前应优先基于现有 artifacts 判断真实可恢复起点,避免从过早阶段重跑
|
||||
|
||||
### 二级目标(建议达成)
|
||||
|
||||
4. job 层状态结果中增加对 linked run 的补充解释,避免“job running 但 run 产物已齐”这种情况毫无说明
|
||||
5. 为后续编排层提供明确可消费的“状态可信度/冲突提示”字段
|
||||
|
||||
---
|
||||
|
||||
## 最小修复方案
|
||||
|
||||
### Phase 1|先修 run 查询层(优先级最高)
|
||||
|
||||
#### 目标
|
||||
让 `get_run_status()` 至少能正确识别:
|
||||
- run-state 是旧的
|
||||
- 但关键产物已经齐了
|
||||
|
||||
#### 建议改动点
|
||||
文件:`src/summary_mcp/runtime/query_service.py`
|
||||
|
||||
#### 建议动作
|
||||
- [x] 在 `_resolve_run_record()` 或 `_build_status_response()` 中增加“关键产物存在性检查”
|
||||
- `run-report.json`
|
||||
- `candidates/openclaw-delivery-payload.json`
|
||||
- `candidates/digest-brief.json`
|
||||
- [x] 如果 `run-state.current_stage` 仍停留在早期阶段,但关键产物已齐:
|
||||
- 不要继续原样输出为可信最终态
|
||||
- 应直接把对外 `status` / `current_stage` / `recovery` 收敛成终态语义
|
||||
- 同时新增解释字段,例如:
|
||||
- `state_conflict: true`
|
||||
- `state_conflict_reason: "run_state indicates generate_summaries but run-report.json already proves the workflow reached a terminal state"`
|
||||
- `status_source: "run_report_reconciliation"`
|
||||
- [x] 保留 `state_source=run_state`,但增加 `status_source` / `state_quality` / `state_conflict` 之类解释字段
|
||||
|
||||
#### 预期收益
|
||||
- OpenClaw 继续按 `status` 分支时也不会卡住
|
||||
- 第一时间减少“明明产物齐了却还像没跑完”的误判
|
||||
- 不需要立刻动 workflow 主链路
|
||||
|
||||
---
|
||||
|
||||
### Phase 2|修 `resume_run` 的恢复起点判断
|
||||
|
||||
#### 目标
|
||||
避免 stale state 让恢复逻辑从 `generate_summaries` 这类过早阶段重跑。
|
||||
|
||||
#### 建议改动点
|
||||
文件:`src/summary_mcp/runtime/resume_service.py`
|
||||
|
||||
#### 建议动作
|
||||
- [x] 在 `_resolve_resume_from_stage()` 之前/之后加入真实 artifacts 检查
|
||||
- [x] 如果以下文件已存在:
|
||||
- `openclaw-delivery-payload.json`
|
||||
- `digest-brief.json`
|
||||
- `run-report.json`
|
||||
则不要再从 `generate_summaries` 或 `apply_filters` 起跑
|
||||
- [x] 为 `resume_run()` 增加“恢复起点是基于 artifacts 重算还是基于 state 推断”的返回说明
|
||||
- [x] 必要时增加更保守逻辑:
|
||||
- `run-report.json` 已存在时,默认拒绝继续 resume,并提示“产物已完成,请先检查状态一致性”
|
||||
- 补充:默认生产模式下,主链路会稳定写出 `summary-batch` / `candidate-batch`,恢复逻辑优先消费这两个 batch artifacts;若它们缺失或不稳定,才回退到更早的安全 stage 或直接拒绝恢复
|
||||
- 补充:调用方可先走 `inspect_resume_plan`,只有 `recommended_action=resume` 时再调用 `resume_run`
|
||||
|
||||
#### 预期收益
|
||||
- 降低无意义重跑和 timeout 风险
|
||||
- 让 `resume_run` 更接近真正的恢复工具,而不是误重跑工具
|
||||
|
||||
---
|
||||
|
||||
### Phase 3|补 job/run 双状态解释层
|
||||
|
||||
#### 目标
|
||||
让 `get_freshrss_pipeline_job_status()` 和 `get_run_status()` 的关系对调用方更可理解。
|
||||
|
||||
#### 建议改动点
|
||||
文件:`src/summary_mcp/runtime/freshrss_pipeline_jobs.py`
|
||||
|
||||
#### 建议动作
|
||||
- [x] 在 `get_freshrss_pipeline_job_status()` 中,读取 linked run 的关键产物存在性(轻量即可)
|
||||
- [x] 若 job 仍显示 `run_pipeline`,但 linked run 已有 report/payload/digest 产物:
|
||||
- 不仅增加解释字段,还应直接把 job 对外 `status` 收敛为终态,避免外层永远轮询
|
||||
- 例如:
|
||||
- `status_source: "linked_run_reconciliation"`
|
||||
- `status_note: "linked run artifacts are complete; the job can be treated as completed"`
|
||||
- [x] 若 `result.json` 缺失,但 linked run 已有 `run-report.json` 与 delivery 产物:
|
||||
- `get_freshrss_pipeline_job_result()` 应能基于 linked run 产物合成最小结果,至少稳定返回 `run_id`
|
||||
- [x] 明确文档:job status 是外层异步任务态,不等于内部 workflow 细粒度状态
|
||||
|
||||
#### 预期收益
|
||||
- 减少“job running / run finished”口径冲突带来的误解
|
||||
- 避免 OpenClaw 因 outer job stale state 卡死在轮询和 result 读取前
|
||||
|
||||
---
|
||||
|
||||
### Phase 4|把 `resume_run` 改成异步恢复 job
|
||||
|
||||
#### 目标
|
||||
解决当前剩余的核心问题:`resume_run` 虽然恢复判定已经安全,但执行模型仍是同步 MCP 调用,长链路恢复时依然可能超时,导致 OpenClaw 编排层“看起来像又卡住了”。
|
||||
|
||||
#### 建议改动点
|
||||
文件:
|
||||
- `src/summary_mcp/runtime/resume_jobs.py`(新)
|
||||
- `scripts/run_resume_job.py`(新)
|
||||
- `src/summary_mcp/server.py`
|
||||
- `src/summary_mcp/runtime/__init__.py`
|
||||
- `src/summary_mcp/runtime/resume_service.py`
|
||||
|
||||
#### 建议动作
|
||||
- [x] 新增最小异步恢复接口:
|
||||
- `start_resume_job(run_id)`
|
||||
- `get_resume_job_status(job_id)`
|
||||
- `get_resume_job_result(job_id)`
|
||||
- [x] job 目录固定落到:
|
||||
- `outputs/freshrss/resume_jobs/<job_id>/`
|
||||
- [x] 最少产物约定:
|
||||
- `run-state.json`
|
||||
- `input.json`
|
||||
- `result.json`(成功时)
|
||||
- `job-report.json`
|
||||
- [x] `start_resume_job` 内部先调用 `inspect_resume_plan`
|
||||
- 只有 `recommended_action=resume` 才允许真正启动
|
||||
- `read_terminal_result` / `start_new_run` 要直接在 job 输入校验阶段返回,不进入执行器
|
||||
- [x] 后台执行时复用现有 `_resume_freshrss_run(...)`
|
||||
- 不重写恢复业务逻辑
|
||||
- 只把同步入口拆成异步 job 外壳
|
||||
- [x] `resume_run(run_id)` 保留,但降级为 debug / fallback
|
||||
- 文档中明确:OpenClaw 编排默认应走 resume async job,而不是同步 `resume_run`
|
||||
- [x] job result 里至少稳定返回:
|
||||
- `run_id`
|
||||
- `resume_from_stage`
|
||||
- `status`
|
||||
- `result_source`
|
||||
- `delivery_output` / `report_output`(若存在)
|
||||
|
||||
#### 预期收益
|
||||
- 彻底切掉恢复阶段的 MCP 同步超时风险
|
||||
- 让 OpenClaw 对“启动恢复 / 轮询恢复 / 读取恢复结果”的控制面与主 pipeline async job 保持一致
|
||||
- 把“恢复判定”与“恢复执行”分层,减少误调用和卡住错觉
|
||||
|
||||
---
|
||||
|
||||
## 不建议现在就做的事
|
||||
|
||||
- [ ] **不要先做自动 fallback 修状态**
|
||||
- 例如:看到 artifacts 齐了就直接把 run-state 强行改成 success
|
||||
- 原因:这会掩盖真正的状态写回问题
|
||||
|
||||
- [ ] **不要先大改 workflow 主链路**
|
||||
- 当前更像查询层与恢复层的状态解释缺陷
|
||||
- 先修读取与恢复判断,收益更大、风险更低
|
||||
|
||||
---
|
||||
|
||||
## 建议执行顺序
|
||||
|
||||
1. **先改 `query_service.py`**
|
||||
- 让 `get_run_status()` 能暴露 stale state / artifact conflict
|
||||
2. **再改 `resume_service.py`**
|
||||
- 避免从错误阶段重跑
|
||||
3. **最后看 `freshrss_pipeline_jobs.py`**
|
||||
- 给 job status 加 linked run 补充说明
|
||||
4. **收尾改 `resume async job`**
|
||||
- 让恢复执行也走正式异步控制面,避免同步恢复再把编排卡住
|
||||
|
||||
---
|
||||
|
||||
## 验收标准
|
||||
|
||||
### 验收 1:状态冲突识别
|
||||
构造一个场景:
|
||||
- `run-state.json` 留在 `generate_summaries`
|
||||
- 但 payload / digest brief / run-report 已存在
|
||||
|
||||
期望:
|
||||
- `get_run_status()` 不再只回“卡在 generate_summaries”
|
||||
- 会显式返回冲突提示字段
|
||||
|
||||
### 验收 2:恢复起点修正
|
||||
构造一个场景:
|
||||
- `run-state` 指向 `generate_summaries`
|
||||
- 但 `delivery_payload` / `run-report` 已存在
|
||||
|
||||
期望:
|
||||
- `resume_run()` 不应再从 summary 阶段重跑
|
||||
- 至少应拒绝恢复并提示“产物已完成,优先检查状态一致性”
|
||||
|
||||
### 验收 3:job/run 双层说明
|
||||
构造一个场景:
|
||||
- job status 仍在 `run_pipeline`
|
||||
- linked run 已有关键产物
|
||||
|
||||
期望:
|
||||
- `get_freshrss_pipeline_job_status()` 能返回补充说明,不再只有生硬 running
|
||||
|
||||
### 验收 4:恢复执行不再阻塞编排
|
||||
构造一个场景:
|
||||
- run 可恢复
|
||||
- 恢复点为 `generate_summaries` 或 `apply_filters`
|
||||
- 恢复执行耗时超过单次 MCP 同步窗口
|
||||
|
||||
期望:
|
||||
- OpenClaw 调用的是 `start_resume_job(...)`,而不是同步 `resume_run(...)`
|
||||
- `get_resume_job_status(job_id)` 可稳定轮询到终态
|
||||
- `get_resume_job_result(job_id)` 至少稳定返回 `run_id`、`resume_from_stage` 与最终产物引用
|
||||
- 即使恢复失败,也能在 job-report / result 中看清失败点,而不是只表现为调用超时
|
||||
|
||||
---
|
||||
|
||||
## 备注
|
||||
|
||||
截至 2026-04-14,本文件中的 Phase 1 / 2 / 3 / 4 已完成主要落地;当前 `resume` 链路已经从“状态收敛 + 安全恢复点判定”进一步补齐到“正式异步恢复执行”。
|
||||
+213
-156
@@ -1,24 +1,196 @@
|
||||
# OpenClaw Handoff
|
||||
|
||||
## Role
|
||||
|
||||
This file is the integration overview for OpenClaw maintainers.
|
||||
|
||||
Use it for:
|
||||
|
||||
- reader capability boundary
|
||||
- production MCP entrypoints
|
||||
- environment requirements
|
||||
- integration rules and limitations
|
||||
|
||||
Do not use it as the step-by-step runbook.
|
||||
For formal orchestration, read `docs/openclaw/openclaw-orchestration-flow.md`.
|
||||
For field contracts, read:
|
||||
|
||||
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||
- `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||
|
||||
Historical plans and incident documents live under `docs/openclaw/archive/`.
|
||||
|
||||
## Purpose
|
||||
|
||||
This repository provides a FreshRSS-first reading pipeline for OpenClaw:
|
||||
reader is the upstream FreshRSS processing service for OpenClaw:
|
||||
|
||||
`FreshRSS unread items -> RSS content extraction -> LLM summary -> rule engine -> OpenClaw delivery payload`
|
||||
|
||||
OpenClaw should treat this repository as an MCP-backed upstream content processor.
|
||||
reader is responsible for:
|
||||
|
||||
## Production Entrypoint
|
||||
- FreshRSS pull
|
||||
- content extraction
|
||||
- LLM summary generation and validation
|
||||
- rule-based filtering
|
||||
- OpenClaw delivery payload generation
|
||||
- run-state persistence and run/result lookup
|
||||
- async resume control for the FreshRSS workflow
|
||||
- async selected-article summary generation from existing extracted files
|
||||
|
||||
OpenClaw should call the MCP tool:
|
||||
reader is not responsible for:
|
||||
|
||||
- Hugo publishing
|
||||
- chat reporting
|
||||
- user confirmation handling
|
||||
- IMA upload orchestration
|
||||
|
||||
## Production Surface
|
||||
|
||||
Current MCP tool count: 21.
|
||||
|
||||
Main daily workflow:
|
||||
|
||||
- `start_freshrss_pipeline_job`
|
||||
- `get_freshrss_pipeline_job_status`
|
||||
- `get_freshrss_pipeline_job_result`
|
||||
- `get_run_status`
|
||||
- `list_runs`
|
||||
- `list_run_artifacts`
|
||||
- `get_delivery_payload`
|
||||
- `get_run_report`
|
||||
|
||||
Resume workflow:
|
||||
|
||||
- `inspect_resume_plan`
|
||||
- `start_resume_job`
|
||||
- `get_resume_job_status`
|
||||
- `get_resume_job_result`
|
||||
- `resume_run`
|
||||
|
||||
Selected-article summary workflow:
|
||||
|
||||
- `start_article_summary_job`
|
||||
- `get_article_summary_job_status`
|
||||
- `get_article_summary_job_result`
|
||||
- `generate_article_summaries`
|
||||
|
||||
Debug / single-step tools:
|
||||
|
||||
- `run_freshrss_openclaw_pipeline`
|
||||
- `extract_url_content`
|
||||
- `extract_item_content`
|
||||
- `filter_summary_result`
|
||||
|
||||
This is the canonical entrypoint for production use.
|
||||
Production rules:
|
||||
|
||||
## Required Environment Variables
|
||||
- main production start path is `start_freshrss_pipeline_job`
|
||||
- production resume path is `inspect_resume_plan -> start_resume_job -> get_resume_job_status -> get_resume_job_result`
|
||||
- `run_freshrss_openclaw_pipeline` is sync debug / fallback only
|
||||
- `resume_run` is sync debug / fallback only
|
||||
- `generate_article_summaries` is sync debug / fallback only
|
||||
|
||||
The MCP server process must have these variables available:
|
||||
## Production Contract
|
||||
|
||||
OpenClaw should treat the returned `run_id` from `get_freshrss_pipeline_job_result` as the only stable handle for follow-up reads.
|
||||
|
||||
OpenClaw should not hand-build these paths:
|
||||
|
||||
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
|
||||
|
||||
If filesystem access is needed for debugging, only consume paths returned by MCP:
|
||||
|
||||
- `output_dir`
|
||||
- `artifact.path`
|
||||
- `delivery_output`
|
||||
- `report_output`
|
||||
|
||||
Top-level `status` is the only status field callers should branch on.
|
||||
`status_source` and `state_conflict` are explanatory fields for reconciled status.
|
||||
|
||||
## Minimal Production Sequence
|
||||
|
||||
Daily workflow:
|
||||
|
||||
1. Call `start_freshrss_pipeline_job`
|
||||
2. Poll `get_freshrss_pipeline_job_status`
|
||||
3. On success, read `get_freshrss_pipeline_job_result`
|
||||
4. Persist the returned `run_id`
|
||||
5. Use `get_run_status`, `get_delivery_payload`, and `get_run_report` for follow-up reads
|
||||
|
||||
Resume workflow:
|
||||
|
||||
1. Call `inspect_resume_plan(run_id)`
|
||||
2. Only if `can_resume=true` and `recommended_action=resume`, call `start_resume_job`
|
||||
3. Poll `get_resume_job_status`
|
||||
4. Read `get_resume_job_result`
|
||||
|
||||
Selected-article summary workflow:
|
||||
|
||||
1. Call `start_article_summary_job` with a real extracted file path and non-empty `selected_ids`
|
||||
2. Poll `get_article_summary_job_status`
|
||||
3. Read `get_article_summary_job_result`
|
||||
|
||||
## Capability Boundary
|
||||
|
||||
Formal workflow boundary:
|
||||
|
||||
- only workflow `freshrss_daily_digest`
|
||||
- every current FreshRSS run writes `run-state.json`
|
||||
- `get_run_status` / `list_runs` / `list_run_artifacts` can still infer basic state for older runs without `run-state.json`
|
||||
- resume requires a valid `run-state.json`; inferred historical runs are not resumable
|
||||
|
||||
Resume boundary:
|
||||
|
||||
- resume in place on the original `run_id`
|
||||
- supported resume points:
|
||||
- `generate_summaries`
|
||||
- `apply_filters`
|
||||
- `build_delivery_payload`
|
||||
- `write_run_report`
|
||||
- unsupported resume points:
|
||||
- `fetch_feed`
|
||||
- `extract_articles`
|
||||
- production resume prefers:
|
||||
- `summary/summary-batch.json`
|
||||
- `candidates/candidate-batch.json`
|
||||
- if required artifacts are missing, recovery should return non-resumable instead of silently falling back
|
||||
|
||||
Selected-article summary boundary:
|
||||
|
||||
- uses existing extracted files as input
|
||||
- should not re-fetch original URLs
|
||||
|
||||
## Output Expectations
|
||||
|
||||
Main daily pipeline core artifacts:
|
||||
|
||||
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/raw/freshrss.raw.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/summary/summary-batch.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/candidates/candidate-batch.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/candidates/digest-brief.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json`
|
||||
|
||||
Async job state directories:
|
||||
|
||||
- main pipeline job: `outputs/freshrss/pipeline_jobs/<job_id>/`
|
||||
- resume job: `outputs/freshrss/resume_jobs/<job_id>/`
|
||||
- article-summary job: `outputs/freshrss/article_summary_jobs/<job_id>/`
|
||||
|
||||
Each job directory minimally contains:
|
||||
|
||||
- `run-state.json`
|
||||
- `input.json`
|
||||
- `result.json` on success
|
||||
- `job-report.json`
|
||||
|
||||
## Environment And Startup
|
||||
|
||||
Required environment variables:
|
||||
|
||||
- `FRESHRSS_API_BASE_URL`
|
||||
- `FRESHRSS_USERNAME`
|
||||
@@ -27,39 +199,19 @@ The MCP server process must have these variables available:
|
||||
- `LLM_API_KEY`
|
||||
- `LLM_MODEL`
|
||||
|
||||
Example:
|
||||
|
||||
```powershell
|
||||
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||||
set FRESHRSS_USERNAME=osiman
|
||||
set FRESHRSS_API_PASSWORD=your-api-password
|
||||
set LLM_API_URL=https://api.deepseek.com
|
||||
set LLM_API_KEY=your-llm-api-key
|
||||
set LLM_MODEL=deepseek-chat
|
||||
```
|
||||
|
||||
## Server Startup
|
||||
|
||||
Install dependencies:
|
||||
Startup:
|
||||
|
||||
```bash
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
Start the MCP server:
|
||||
|
||||
```bash
|
||||
summary-mcp
|
||||
```
|
||||
|
||||
## Recommended MCP Call
|
||||
|
||||
Recommended default call:
|
||||
Recommended production start call:
|
||||
|
||||
```json
|
||||
{
|
||||
"limit": 5,
|
||||
"mark_read": false,
|
||||
"mark_read": true,
|
||||
"include_read": false,
|
||||
"debug_artifacts": false,
|
||||
"timeout_seconds": 60,
|
||||
@@ -67,146 +219,51 @@ Recommended default call:
|
||||
}
|
||||
```
|
||||
|
||||
Recommended semantics:
|
||||
## Data And Content Policy
|
||||
|
||||
- Use `mark_read=false` while validating integration.
|
||||
- Use `mark_read=true` only after confirming OpenClaw will consume the returned payload successfully.
|
||||
- Keep `debug_artifacts=false` for routine production runs.
|
||||
- Set `debug_artifacts=true` only when troubleshooting a bad batch.
|
||||
|
||||
## What The Tool Returns
|
||||
|
||||
Primary return fields:
|
||||
|
||||
- `run_id`
|
||||
- `output_dir`
|
||||
- `raw_output`
|
||||
- `delivery_output`
|
||||
- `report_output`
|
||||
- `pulled_count`
|
||||
- `delivered_count`
|
||||
- `marked_read_count`
|
||||
- `status_counts`
|
||||
- `delivery_payload`
|
||||
- `keyword_index`
|
||||
|
||||
Optional:
|
||||
|
||||
- `items`
|
||||
- Returned only when `include_item_reports=true`
|
||||
|
||||
## Minimal Output Files
|
||||
|
||||
By default the pipeline writes only:
|
||||
|
||||
- `outputs/freshrss/rerun/<run_id>/raw/freshrss.raw.json`
|
||||
- `outputs/freshrss/rerun/<run_id>/candidates/openclaw-delivery-payload.json`
|
||||
- `outputs/freshrss/rerun/<run_id>/run-report.json`
|
||||
- `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json` (one per item)
|
||||
|
||||
It also updates local runtime keyword data:
|
||||
|
||||
- `data/term_index/daily/YYYY-MM-DD.json`
|
||||
- `data/term_index/term_stats.json`
|
||||
|
||||
Per-item extracted files live under `extracted/` and are always written.
|
||||
If `debug_artifacts=true`, the pipeline additionally writes normalized items, summaries, filter decisions, candidate records, and candidate inputs.
|
||||
The main pipeline does not emit a batch-level `freshrss.extracted.json` file by default.
|
||||
|
||||
## Payload Specs
|
||||
|
||||
OpenClaw payload field specs live here:
|
||||
|
||||
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||
- `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||
|
||||
## Read-State Semantics
|
||||
|
||||
The pipeline reads from FreshRSS unread items by default.
|
||||
|
||||
If `mark_read=true`:
|
||||
|
||||
- items are marked as read only after the final `openclaw-delivery-payload.json` has been written successfully
|
||||
- only successfully delivered items are marked as read
|
||||
- failed or skipped items remain unread
|
||||
|
||||
## FreshRSS Content Policy
|
||||
|
||||
For FreshRSS items, the pipeline is RSS-first and does not re-crawl webpages.
|
||||
|
||||
Behavior:
|
||||
FreshRSS processing is RSS-first:
|
||||
|
||||
- use `item.raw_content` first
|
||||
- if missing, use `item.raw_summary`
|
||||
- if neither contains usable content, skip the item
|
||||
- do not fetch the original webpage again for FreshRSS items
|
||||
|
||||
This is intentional.
|
||||
Read-state policy:
|
||||
|
||||
## Keyword Cleanup Governance
|
||||
- items are marked read only after successful delivery payload write
|
||||
- only successfully delivered items are marked read
|
||||
|
||||
This repository also includes a lightweight keyword-governance flow for downstream review.
|
||||
Downstream boundary:
|
||||
|
||||
Current pieces:
|
||||
- the daily digest goes to Hugo and chat reporting
|
||||
- the full daily digest should not be uploaded to IMA
|
||||
- only explicitly user-selected article summaries should be uploaded to IMA
|
||||
|
||||
- runtime keyword stats
|
||||
- `data/term_index/daily/YYYY-MM-DD.json`
|
||||
- `data/term_index/term_stats.json`
|
||||
- governance config
|
||||
- `configs/term_cleanup_policy.json`
|
||||
- `configs/term_watchlist.json`
|
||||
- `configs/term_change_log.json`
|
||||
- review bundle builder
|
||||
- `skills/keyword-cleanup-review/scripts/build_review_bundle.py`
|
||||
- accepted-suggestion writer
|
||||
- `scripts/apply_term_suggestions.py`
|
||||
## Related Maintenance Flow
|
||||
|
||||
Current status:
|
||||
Keyword cleanup exists as a separate maintenance flow, not the main RSS ingestion path.
|
||||
|
||||
- OpenClaw can read the keyword review bundle as maintenance input
|
||||
- accepted suggestions still require explicit human confirmation
|
||||
- the repository can write accepted watch / alias / stopword / interest-keyword changes after confirmation
|
||||
- this governance flow is not yet wired into a periodic scheduler inside the repository
|
||||
|
||||
Boundary:
|
||||
|
||||
- keyword cleanup is a maintenance flow, not the production RSS ingestion path
|
||||
- the repository does not auto-apply cleanup suggestions without confirmation
|
||||
- current keyword stats are built from the delivered candidate payload, not yet from a final `DailyDigest`
|
||||
|
||||
## Known Limitations
|
||||
|
||||
- Some sources put only partial content in RSS; those items may be skipped if RSS content is insufficient.
|
||||
- WeChat articles often block direct crawling, but this pipeline now avoids that path for FreshRSS items and uses RSS-provided content when available.
|
||||
- Rule behavior is still conservative in some cases; many items may land in `review` depending on current rules.
|
||||
- Paywall heuristics may produce false positives for some Chinese text patterns.
|
||||
- Keyword cleanup governance is usable now, but periodic scheduling and before/after evaluation are not implemented yet.
|
||||
|
||||
## Files OpenClaw Should Read First
|
||||
|
||||
Recommended reading order for a new maintainer:
|
||||
|
||||
1. `README.md`
|
||||
2. `docs/openclaw/openclaw-handoff.md`
|
||||
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||
5. `docs/design/daily-keyword-index-design.md`
|
||||
6. `skills/keyword-cleanup-review/SKILL.md`
|
||||
7. `docs/current/context-reset-brief.md`
|
||||
|
||||
## Current Recommendation
|
||||
|
||||
For integration handoff, the repository is usable now.
|
||||
|
||||
The minimum you need to give OpenClaw is:
|
||||
|
||||
- the repository code
|
||||
- the MCP server startup command
|
||||
- the required environment variables in the target environment
|
||||
- the instruction to call `run_freshrss_openclaw_pipeline`
|
||||
|
||||
If OpenClaw will also participate in keyword-governance review, additionally point it to:
|
||||
Relevant files:
|
||||
|
||||
- `docs/design/daily-keyword-index-design.md`
|
||||
- `skills/keyword-cleanup-review/SKILL.md`
|
||||
- `scripts/apply_term_suggestions.py`
|
||||
|
||||
## Known Limitations
|
||||
|
||||
- some sources expose only partial RSS content; those items may be skipped
|
||||
- rule behavior is still conservative; many items may land in `review`
|
||||
- paywall heuristics may still produce false positives on some Chinese text
|
||||
- keyword cleanup governance is usable but not yet wired to periodic scheduling
|
||||
|
||||
## Read First
|
||||
|
||||
Recommended reading order for a new maintainer:
|
||||
|
||||
1. `README.md`
|
||||
2. `docs/openclaw/README.md`
|
||||
3. `docs/openclaw/openclaw-handoff.md`
|
||||
4. `docs/openclaw/openclaw-orchestration-flow.md`
|
||||
5. `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||
6. `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||
7. `docs/current/context-reset-brief.md`
|
||||
|
||||
@@ -0,0 +1,340 @@
|
||||
# OpenClaw → reader MCP 标准编排流程
|
||||
|
||||
## 1. 文档目的
|
||||
|
||||
本文档只回答一个问题:OpenClaw 在正式环境里应该如何编排 reader。
|
||||
|
||||
这里不重复介绍 reader 内部实现,只定义正式控制面:
|
||||
|
||||
- 如何启动日报
|
||||
- 如何轮询 job
|
||||
- 如何读取 run 结果
|
||||
- 如何判断是否恢复
|
||||
- 如何走异步恢复
|
||||
- 什么时候直接新开 run 或人工介入
|
||||
|
||||
## 2. 当前正式入口
|
||||
|
||||
### 2.1 新 run
|
||||
|
||||
正式生产入口:
|
||||
|
||||
- `start_freshrss_pipeline_job`
|
||||
- `get_freshrss_pipeline_job_status`
|
||||
- `get_freshrss_pipeline_job_result`
|
||||
|
||||
同步入口:
|
||||
|
||||
- `run_freshrss_openclaw_pipeline`
|
||||
|
||||
同步入口只保留给 debug / fallback,不再是正式编排默认路径。
|
||||
|
||||
### 2.2 run 级读取
|
||||
|
||||
正式 run 级读取接口:
|
||||
|
||||
- `get_run_status`
|
||||
- `list_runs`
|
||||
- `list_run_artifacts`
|
||||
- `get_delivery_payload`
|
||||
- `get_run_report`
|
||||
|
||||
### 2.3 恢复
|
||||
|
||||
正式恢复入口:
|
||||
|
||||
- `inspect_resume_plan`
|
||||
- `start_resume_job`
|
||||
- `get_resume_job_status`
|
||||
- `get_resume_job_result`
|
||||
|
||||
同步恢复入口:
|
||||
|
||||
- `resume_run`
|
||||
|
||||
`resume_run` 只保留给 debug / fallback。
|
||||
|
||||
## 3. 编排基本原则
|
||||
|
||||
### 3.1 OpenClaw 不手拼路径
|
||||
|
||||
OpenClaw 不应自己推导这些路径:
|
||||
|
||||
- `outputs/freshrss/rerun/<run_dir>/run-state.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/candidates/openclaw-delivery-payload.json`
|
||||
- `outputs/freshrss/rerun/<run_dir>/run-report.json`
|
||||
|
||||
需要路径时,只消费 MCP 返回值:
|
||||
|
||||
- `output_dir`
|
||||
- `artifact.path`
|
||||
- `delivery_output`
|
||||
- `report_output`
|
||||
|
||||
### 3.2 顶层 `status` 才是分支依据
|
||||
|
||||
`get_run_status` 和 job status 接口都可能做状态收敛。
|
||||
|
||||
因此:
|
||||
|
||||
- 优先使用顶层 `status`
|
||||
- `status_source` 用来解释状态来自原始 state 还是收敛结果
|
||||
- `state_conflict=true` 说明底层状态文件已经落后于真实产物
|
||||
|
||||
不要再拿旧的 `raw_status`、`raw_current_stage` 或早期阶段名重新做分支。
|
||||
|
||||
### 3.3 默认生产语义
|
||||
|
||||
正式生产运行默认:
|
||||
|
||||
- `mark_read=true`
|
||||
- `debug_artifacts=false`
|
||||
|
||||
只有 debug / test / validation 时才放宽。
|
||||
|
||||
## 4. 标准 Happy Path
|
||||
|
||||
### Step 1: 启动新 job
|
||||
|
||||
调用:
|
||||
|
||||
- `start_freshrss_pipeline_job`
|
||||
|
||||
推荐参数:
|
||||
|
||||
```json
|
||||
{
|
||||
"limit": 5,
|
||||
"mark_read": true,
|
||||
"include_read": false,
|
||||
"debug_artifacts": false,
|
||||
"timeout_seconds": 60,
|
||||
"max_retries": 2
|
||||
}
|
||||
```
|
||||
|
||||
预期:
|
||||
|
||||
- 立即返回 `job_id`
|
||||
- 后续由 OpenClaw 轮询 job,而不是同步等待整条流水线
|
||||
|
||||
### Step 2: 轮询 job
|
||||
|
||||
调用:
|
||||
|
||||
- `get_freshrss_pipeline_job_status(job_id=...)`
|
||||
|
||||
根据返回:
|
||||
|
||||
- `status=running`:继续轮询
|
||||
- `status=success`:读取 job result
|
||||
- `status=failed`:进入失败处理
|
||||
|
||||
额外规则:
|
||||
|
||||
- 如果 `status_source=linked_run_reconciliation`,说明 outer job state 已落后,但 linked run 已经给出可用终态
|
||||
- 如果 `status_source=stale_job_state_timeout`,把它当成终态失败,不要继续无限轮询
|
||||
|
||||
### Step 3: 读取 job result
|
||||
|
||||
调用:
|
||||
|
||||
- `get_freshrss_pipeline_job_result(job_id=...)`
|
||||
|
||||
预期读取:
|
||||
|
||||
- `run_id`
|
||||
- `output_dir`
|
||||
- `delivery_output`
|
||||
- `report_output`
|
||||
|
||||
从这一刻开始,`run_id` 是正式的稳定句柄。
|
||||
|
||||
### Step 4: 读取 run 级状态与结果
|
||||
|
||||
调用:
|
||||
|
||||
- `get_run_status(run_id=...)`
|
||||
- `get_delivery_payload(run_id=...)`
|
||||
- `get_run_report(run_id=...)`
|
||||
|
||||
根据 `get_run_status`:
|
||||
|
||||
- `status=running`:继续观察
|
||||
- `status=success`:继续下游 digest / 发布 / 汇报
|
||||
- `status=failed`:进入恢复或重跑决策
|
||||
- `status=partial`:优先检查 report、artifacts 和 recovery
|
||||
|
||||
如果 `status_source=run_report_reconciliation`,说明 `run-state.json` 已经过期,但 reader 已经根据终态产物收敛出有效状态。
|
||||
|
||||
如果 `status_source=stale_run_state_timeout`,说明 reader 认为该 run 长时间未收敛且没有终态产物,应按失败处理。
|
||||
|
||||
## 5. 恢复决策
|
||||
|
||||
### 5.1 先看预检,不要直接恢复
|
||||
|
||||
恢复前固定动作:
|
||||
|
||||
- 先调用 `inspect_resume_plan(run_id)`
|
||||
|
||||
只在以下条件同时成立时才启动恢复:
|
||||
|
||||
- `can_resume=true`
|
||||
- `recommended_action=resume`
|
||||
|
||||
重点字段:
|
||||
|
||||
- `requested_resume_from_stage`
|
||||
- `resume_from_stage`
|
||||
- `resume_decision_source`
|
||||
- `artifact_resume_from_stage`
|
||||
- `artifact_snapshot`
|
||||
|
||||
### 5.2 正式恢复路径
|
||||
|
||||
正式恢复控制面:
|
||||
|
||||
1. `start_resume_job(run_id)`
|
||||
2. `get_resume_job_status(job_id)`
|
||||
3. `get_resume_job_result(job_id)`
|
||||
|
||||
不要再把同步 `resume_run(run_id)` 当成正式恢复入口。
|
||||
|
||||
### 5.3 当前支持范围
|
||||
|
||||
当前只支持:
|
||||
|
||||
- 带有效 `run-state.json` 的 `freshrss_daily_digest` run
|
||||
- 从以下阶段恢复:
|
||||
- `generate_summaries`
|
||||
- `apply_filters`
|
||||
- `build_delivery_payload`
|
||||
- `write_run_report`
|
||||
|
||||
当前不支持:
|
||||
|
||||
- `fetch_feed`
|
||||
- `extract_articles`
|
||||
|
||||
正式生产恢复优先依赖:
|
||||
|
||||
- `summary/summary-batch.json`
|
||||
- `candidates/candidate-batch.json`
|
||||
|
||||
### 5.4 什么时候不要恢复
|
||||
|
||||
以下情况直接新开 run 更合理:
|
||||
|
||||
- `recommended_action=start_new_run`
|
||||
- `recommended_action=read_terminal_result`
|
||||
- 没有有效 `run-state.json`
|
||||
- 恢复所需关键 artifacts 缺失
|
||||
- 连续恢复失败
|
||||
|
||||
## 6. 状态到动作映射
|
||||
|
||||
| 接口 | 状态 | OpenClaw 动作 |
|
||||
| --- | --- | --- |
|
||||
| `get_freshrss_pipeline_job_status` | `running` | 继续轮询 job |
|
||||
| `get_freshrss_pipeline_job_status` | `success` | 读取 `get_freshrss_pipeline_job_result` |
|
||||
| `get_freshrss_pipeline_job_status` | `failed` | 结束本次 job,必要时读 linked run |
|
||||
| `get_run_status` | `running` | 继续观察 run |
|
||||
| `get_run_status` | `success` | 读取 `get_delivery_payload` / `get_run_report` |
|
||||
| `get_run_status` | `failed` | 先看 `inspect_resume_plan` |
|
||||
| `inspect_resume_plan` | `recommended_action=resume` | 启动 `start_resume_job` |
|
||||
| `inspect_resume_plan` | `recommended_action=read_terminal_result` | 直接读 run 结果,不恢复 |
|
||||
| `inspect_resume_plan` | `recommended_action=start_new_run` | 新开 run 或人工介入 |
|
||||
|
||||
## 7. 人工介入条件
|
||||
|
||||
出现以下任一情况时,建议不要自动编排:
|
||||
|
||||
- 连续恢复失败
|
||||
- payload / report 结构不符合预期
|
||||
- `get_run_status` 与实际产物长期明显冲突
|
||||
- FreshRSS、LLM 或外部依赖异常
|
||||
- 恢复判定结果和编排预期不一致
|
||||
|
||||
## 8. 结论
|
||||
|
||||
当前 OpenClaw 的正式调用方式已经收口为两条异步控制面:
|
||||
|
||||
- 主日报:`start_freshrss_pipeline_job -> poll -> get result -> run reads`
|
||||
- 恢复:`inspect_resume_plan -> start_resume_job -> poll -> get result`
|
||||
|
||||
同步 `run_freshrss_openclaw_pipeline` 和 `resume_run` 仅用于 debug / fallback,不应再作为默认正式编排路径。
|
||||
|
||||
### 7.2 `get_run_report`
|
||||
|
||||
用途:
|
||||
|
||||
- 获取 run 的结果摘要与关键元信息
|
||||
- 用于状态判断、排障、补充上下文
|
||||
|
||||
OpenClaw 应做:
|
||||
|
||||
- 作为诊断与编排辅助信息读取
|
||||
- 不把 raw file path 解析逻辑继续散落到 skill 里
|
||||
|
||||
---
|
||||
|
||||
## 8. 最小编排动作表
|
||||
|
||||
### 8.1 标准生产执行
|
||||
|
||||
1. 调 `run_freshrss_openclaw_pipeline`
|
||||
2. 拿 `run_id`
|
||||
3. 轮询 `get_run_status`
|
||||
4. 若 `success`:
|
||||
- 调 `get_delivery_payload`
|
||||
- 调 `get_run_report`
|
||||
5. 进入 digest / Hugo / chat / IMA 下游编排
|
||||
|
||||
### 8.2 失败恢复执行
|
||||
|
||||
1. 调 `get_run_status`
|
||||
2. 若 `failed && recovery.resumable=true`:
|
||||
- 调 `start_resume_job`
|
||||
- 轮询 `get_resume_job_status`
|
||||
- 读取 `get_resume_job_result`
|
||||
3. 恢复后再次:
|
||||
- 调 `get_run_status`
|
||||
- 若成功,再读 payload / report
|
||||
4. 若恢复失败或明确不可恢复:
|
||||
- 新开 run 或人工介入
|
||||
|
||||
---
|
||||
|
||||
## 9. 不推荐做法
|
||||
|
||||
以下做法不应再作为正式主路径:
|
||||
|
||||
- 让 OpenClaw 直接长时间 `exec` reader CLI 作为主要生产入口
|
||||
- 让 OpenClaw 自己拼 reader 输出路径来判断成功/失败
|
||||
- 让 OpenClaw 自己读取 `outputs/.../*.json` 作为正式结果源
|
||||
- 在未确认恢复支持范围外的失败点上强行恢复
|
||||
|
||||
CLI 现在的定位是:
|
||||
|
||||
- debug
|
||||
- fallback
|
||||
- 人工排障
|
||||
|
||||
而不是正式生产主入口。
|
||||
|
||||
---
|
||||
|
||||
## 10. 当前已知局限
|
||||
|
||||
- 恢复能力仍是最小实现,不支持任意 stage 任意重入
|
||||
- 历史无 `run-state.json` 的 run 不支持正式恢复
|
||||
- 极旧 run 的结果读取仍可能依赖保守目录扫描
|
||||
- `write_run_report` 若涉及重新 `mark_read`,仍依赖 FreshRSS 环境和可用凭据
|
||||
|
||||
---
|
||||
|
||||
## 11. 一句话结论
|
||||
|
||||
OpenClaw 当前应把 reader 当作正式 MCP workflow service 使用:
|
||||
|
||||
**启动用 `start_freshrss_pipeline_job`,观测用 `get_run_status`,结果读取用 `get_delivery_payload` / `get_run_report`,恢复默认用 `inspect_resume_plan` + `start_resume_job`,不要再把 reader 当成长 CLI 任务和路径拼接仓库来驱动。**
|
||||
@@ -0,0 +1,86 @@
|
||||
# [bug] MCP `start_article_summary_job` 超时但 job 实际执行了
|
||||
|
||||
## 问题描述
|
||||
|
||||
连续多个 `start_article_summary_job` 调用报 MCP 超时(-32001 Request timed out),但 job 实际执行了:
|
||||
|
||||
- `run_state.json` 显示 `status: failed`,`current_stage: generate_markdown`
|
||||
- job 目录正常生成,`run-state.json` 存在于 `outputs/freshrss/article_summary_jobs/<job-id>/`
|
||||
- `generate_markdown` 阶段实际执行过(有结果),但 MCP 响应没能发回来
|
||||
|
||||
## 根因分析(已定位)
|
||||
|
||||
**根本原因:双阻塞点导致 MCP stdio 响应超时**
|
||||
|
||||
### 阻塞点 1:`RunStore.save()` 高频同步写盘
|
||||
|
||||
`start_article_summary_job` 执行流程中的写盘次数:
|
||||
|
||||
```python
|
||||
store.save() # ← 第 134 行:初始化后写盘
|
||||
store.start_stage("prepare_job") # ← 第 68 行:内部 save()
|
||||
store.register_artifact(...) # ← 第 143 行:内部 save()
|
||||
store.finish_stage("prepare_job", ...) # ← 第 90 行:内部 save()
|
||||
return {...} # ← 第 158 行:返回 MCP 响应
|
||||
```
|
||||
|
||||
**4 次同步写盘** 在 `Popen` 之后、`return` 之前完成。磁盘 I/O 慢 + Windows 文件锁 = **响应超时**。
|
||||
|
||||
### 阻塞点 2:MCP FastMCP stdio 传输机制
|
||||
|
||||
`mcp.run()` 默认使用 **stdio 传输**(进程间管道)。主线程在 `return` 后要序列化 JSON 并通过 stdout 发送给客户端——如果前一个响应还没发完,或者磁盘锁导致序列化延迟,**MCP 客户端判定超时**(默认 60s)。
|
||||
|
||||
### 原代码问题
|
||||
|
||||
```python
|
||||
# ← 原代码:同步阻塞
|
||||
proc = subprocess.Popen(...) # Popen 返回
|
||||
store.finish_stage("prepare_job", outputs={...}) # ← 同步写盘
|
||||
return {"job_id": job_id, ...} # ← 这里已经超时了
|
||||
```
|
||||
|
||||
## 修复方案(已实施)
|
||||
|
||||
**方案 A:异步解耦** —— 用 `ThreadPoolExecutor` 后台启动 job,主线程立即返回响应。
|
||||
|
||||
```python
|
||||
# ← 修复后:异步非阻塞
|
||||
_executor = ThreadPoolExecutor(max_workers=4, thread_name_prefix="article_summary_job")
|
||||
|
||||
def _launch_job_background(*, job_id, input_payload, store):
|
||||
proc = subprocess.Popen(...)
|
||||
store.finish_stage("prepare_job", outputs={...}) # ← 后台线程写盘
|
||||
|
||||
def start_article_summary_job(...):
|
||||
# ... 创建 store 和 input 文件
|
||||
store.start_stage("prepare_job") # ← 不调用 save()
|
||||
store.register_artifact(...) # ← 不调用 save()
|
||||
|
||||
# 【关键】:后台线程执行 Popen + finish_stage,主线程立即返回
|
||||
_executor.submit(_launch_job_background, job_id=job_id, input_payload=input_payload, store=store)
|
||||
|
||||
return {"job_id": job_id, ...} # ← 立即返回,不等待写盘
|
||||
```
|
||||
|
||||
**修复效果**:
|
||||
- 主线程:创建 job 目录 → 写 input.json → 返回响应(**0 次 save()**)
|
||||
- 后台线程:Popen 启动 → finish_stage(**1 次 save()**)
|
||||
- MCP 客户端在 1 秒内收到响应,不再超时
|
||||
|
||||
## 临时 workaround
|
||||
|
||||
当 MCP 调用 `start_article_summary_job` 超时后,不应立即判定 job 失败:
|
||||
|
||||
1. 调用 `get_article_summary_job_status(job_id)` 查询真实状态
|
||||
2. 若返回 `status=running` 或 `run-state.json` 存在且 `status=running` → job 在跑,继续等待
|
||||
3. 若返回 `status=failed` → 查 `run-state.json` 的 `failed_stage` 和 `error_summary`
|
||||
|
||||
## 影响范围
|
||||
|
||||
- OpenClaw MCP 客户端调用 `start_article_summary_job`
|
||||
- 任何通过 stdio MCP 通道使用 article summary job 的场景
|
||||
|
||||
## 修复时间
|
||||
|
||||
- **根因定位**:2026-04-16
|
||||
- **修复实施**:2026-04-16(方案 A:异步解耦)
|
||||
@@ -0,0 +1,41 @@
|
||||
你是一个专业的知识沉淀助手。你的任务是对提供的文章正文做深度分析,输出一份结构化的知识沉淀笔记,而不是简短的摘要卡片。
|
||||
|
||||
要求:
|
||||
- 基于文章完整正文(article.plain_text)进行分析
|
||||
- 所有输出字段使用中文
|
||||
- 如果文章没有相关内容(如无技术方法、无具体细节),对应数组字段返回空数组 []
|
||||
- 输出必须是单个合法 JSON 对象,不要加任何解释文字
|
||||
|
||||
输出 JSON schema:
|
||||
|
||||
{
|
||||
"title": "文章标题(与原文一致)",
|
||||
"url": "文章 URL(与原文一致)",
|
||||
"core_conclusion": "作者最核心的结论,2-4 句,精准概括,包含核心判断和支撑依据",
|
||||
"main_argument": "文章的主要论点或主张,充分展开,可包含推理链条和论证结构",
|
||||
"key_methods": [
|
||||
"关键方法、机制或技术手段,每条 2-4 句展开说明原理或运作方式,3-6 条;无相关内容时返回空数组"
|
||||
],
|
||||
"important_details": [
|
||||
"值得记录的细节、数据或案例,每条 2-4 句补充背景和意义,3-6 条;无相关内容时返回空数组"
|
||||
],
|
||||
"reusable_insights": [
|
||||
"可复用于其他场景的启发或观点,每条 2-3 句说明适用场景和具体做法,2-4 条"
|
||||
],
|
||||
"keywords": [
|
||||
"具体实体、工具名、方法名,5-8 个"
|
||||
],
|
||||
"topics": [
|
||||
"更高层的主题标签,3-5 个"
|
||||
],
|
||||
"category": "内容分类,从以下选项中选择一个:资讯 / 方法论 / 工具实践 / 观点评论",
|
||||
"worth_keeping": true,
|
||||
"reason": "沉淀理由,2-3 句说明为什么值得长期保留,以及未来什么场景下可以回看"
|
||||
}
|
||||
|
||||
注意:
|
||||
- keywords 和 topics 不能有重叠
|
||||
- keywords 侧重具体实体(工具名、框架名、人名、产品名)
|
||||
- topics 侧重抽象主题(如「知识管理」「系统设计」「AI Agent」)
|
||||
- core_conclusion 必须是作者的核心观点,不是文章描述
|
||||
- 只输出 JSON,不要输出任何其他内容
|
||||
@@ -0,0 +1,34 @@
|
||||
# 信息过载时代,我的漏斗式阅读工作流
|
||||
|
||||
Source: https://shawnxie.top/blogs/tools/read-flow-2026.html
|
||||
Category: 方法论
|
||||
|
||||
## 核心结论
|
||||
在信息过载时代,个人信息处理的核心目标不是获取更多信息,而是通过分层过滤和沉淀机制,稳定地吸收、判断和沉淀真正有价值的内容。作者认为,一个有效的个人信息系统应该是一个“漏斗”,而非“桶”,其关键在于将分散的处理环节串联成闭环,让信息在向下流动的过程中不断收窄,最终只留下值得进入长期记忆的部分,并能通过沉淀内容反向优化筛选逻辑。
|
||||
|
||||
## 主要论点
|
||||
文章主张,应对信息过载不应依赖单一工具或追求全自动处理,而应构建一套分层、闭环的“漏斗式”个人信息系统。该系统以RSS等可控信息源为起点,通过聚合、预处理、AI精选、人工精读、知识沉淀和轻量反馈等多个层级,逐步过滤噪音、提炼精华,并将沉淀的高价值内容转化为个性化信号,反向优化上游筛选。其核心论点是:信息处理的难点在于串联分散环节,真正的价值在于实现“上游宽广、中游稳定、下游精准、回流轻柔”的认知加工流程,使人从被动接收者转变为拥有个性化认知处理系统的主人。
|
||||
|
||||
## 关键方法 / 机制
|
||||
- 信息源归一化:以RSS为主信息基础设施,利用RSSHub、wewe-rss、nitter等工具或自定义脚本,将公众号、社区、社交平台等多样信息源统一转换为RSS或类feed格式,确保下游处理环节输入格式的稳定和统一。
|
||||
- 分层处理流程:设计清晰的多层处理架构。1) 用FreshRSS作为统一聚合池,管理所有订阅源,充当缓冲层。2) 用Digest模块进行预处理,完成URL去重、相似内容去重、正文抓取、质量检查、噪音过滤、摘要生成等任务,将原始信息整理为可判断对象。3) 用Daily Review进行AI精选,由LLM将预处理后的内容按固定栏目(如今日大事、变更与实践等)结构化,生成重点突出的日报。4) 设置人工精读环节(Human in the loop),由人最终判断内容的长期价值。5) 用Lumina知识库进行长期沉淀。6) 基于沉淀内容构建轻量兴趣画像,作为辅助信号轻微影响上游排序,形成闭环。
|
||||
- 自动化编排与集成:使用OpenClaw作为自动化编排层,通过自然语言描述诉求,串联定时触发、外部工具调用(如FreshRSS API)、与飞书/Lumina的交互、AI技能调用等任务,实现整个工作流的自动化运行和快速迭代,降低了构建完整系统的初期门槛。
|
||||
|
||||
## 重要细节
|
||||
- FreshRSS的定位与作用:作者强调FreshRSS并非日常阅读工具,而是作为“中间水库”的聚合池。其核心作用是统一管理信息源格式、为下游环节提供稳定且无需直接联网的内容池、以及确保信息处理的时间连续性,避免了碎片化处理,是实现“有边界的信息处理”的关键缓冲层。
|
||||
- Digest预处理的具体任务:Digest模块承担了将“海量候选内容转化为可判断对象”的重任。其具体任务包括URL精确去重、相似内容去重、正文抓取、质量检查、噪音过滤、摘要生成和初步排序。这一环节大幅减少了标题党、重复报道和无效数据,显著降低了后续人工筛选的认知成本,是提升整个系统效率的基础。
|
||||
- Daily Review的栏目化设计:AI精选环节(Daily Review)并非简单聚合摘要,而是通过预设固定栏目进行结构化编辑。栏目包括“今日大事”、“变更与实践”、“安全与风险”、“开源与工具”、“洞察与数据点”、“主题深挖”等,将信息分配到不同的“认知槽位”,使最终产出更像一份重点突出的个性化日报,而非信息堆砌。
|
||||
- 轻量反馈与兴趣画像设计:为避免系统演变为封闭的“信息茧房”或过度迎合的推荐引擎,作者刻意设计了“轻量”的兴趣画像逻辑。它仅从长期精读主题、高价值信息源、偏好内容格式、Lumina存入内容类型等维度学习,作为辅助信号轻微影响Digest和Daily Review的排序,目标是减少无效判断,同时保留关注公共重要性和探索未知的空间。
|
||||
- 沉淀价值的拓展方向:信息沉淀的终点不仅是存储,还包括价值再生产。作者实践了两个方向:1) 周刊生成:基于一周的Digest、Daily Review和Lumina内容,生成带有个人筛选痕迹和深度分析的周刊,识别长期信号与短期噪音。2) 主题文章生成:系统识别一段时间内反复出现并被多次沉淀的主题(如AI Agent趋势),将其发展为可深入输出的长文主题,将离散信息流转化为结构化知识资产。
|
||||
|
||||
## 可复用启发
|
||||
- 构建分层过滤的认知加工流程:在处理任何海量输入(如邮件、任务、学习资料)时,都可以借鉴“漏斗”思维。设计从“广泛捕获”到“逐步收窄”的多层处理机制,明确每一层的职责(如聚合、粗筛、精筛、决策、沉淀),而非试图一次性处理所有信息,这能大幅降低认知负担并提升处理质量。
|
||||
- 坚持“人在回路”的核心价值判断:在自动化系统中,尤其是在知识管理领域,最关键的长期价值判断(如“什么值得长期留存”、“什么对未来有深远影响”)不应完全外包给算法。应像本文一样,在流程的关键节点(如进入知识库前)设置人工确认环节,确保系统的输出最终服务于人的深度思考和决策。
|
||||
- 利用轻量反馈构建良性闭环系统:在设计个性化系统时,应避免构建强反馈、易导致信息茧房的推荐引擎。可以借鉴本文的“轻量引导”思路,仅从最核心、最长期的行为数据(如最终沉淀内容)中提取少量信号,温和地优化上游流程,在提升效率的同时保持系统的开放性和探索性。
|
||||
- 以“沉淀和再生产”作为流程终点:信息或知识管理的目标不应止步于“读完”或“保存”。应思考如何将处理后的高价值内容进一步转化为可复用的资产,例如定期生成汇总报告(如周刊)、识别并发展跨时间维度的核心主题、或将内化知识用于指导实践和输出,从而实现知识的增值和循环。
|
||||
|
||||
## 关键词
|
||||
RSS、FreshRSS、OpenClaw、Lumina、RSSHub、wewe-rss、nitter、Digest
|
||||
|
||||
## 主题
|
||||
知识管理、个人信息处理、工作流设计、自动化、阅读方法
|
||||
File diff suppressed because one or more lines are too long
@@ -0,0 +1,17 @@
|
||||
# 代码审查建议(2026-03-28)
|
||||
|
||||
## 立即可改(低成本)
|
||||
|
||||
- [x] `server.py` context 通过临时文件传递绕路 — `run_freshrss_pipeline` 改为同时接受 `dict | Path` 类型的 context,消除 `NamedTemporaryFile` 绕路(已修复)
|
||||
- [ ] `_load_required_env` 报错信息不区分"未传参数"还是"环境变量未设",改善调试体验
|
||||
- [x] `evaluate_filter_rules` 决策逻辑歧义 — `stop_on_match=True` 命中后应直接以该规则 decision 为最终结果,而非继续聚合所有 matched(已确认:当前规则集安全,在 filter-rule-engine-usage.md 补充了 drop 规则必须设 stop_on_match=true 的约束说明)
|
||||
|
||||
## 重构建议(中等成本)
|
||||
|
||||
- [x] `run_freshrss_pipeline` 函数过长(约300行)— 拆分为 `_process_item()`、`_build_and_persist_delivery()`、`_build_run_report()` 三个内部函数,主函数只做编排(已完成)
|
||||
- [ ] `load_filter_rules` 每次 pipeline 调用都重新读文件 — 加模块级缓存,MCP 服务长期运行时避免重复 I/O
|
||||
|
||||
## 功能补全
|
||||
|
||||
- [ ] `mark_read` 只标记 delivered items,drop/review 的 item 下次仍会重复拉取 — 引入 `mark_read_all_processed` 选项,或在报告中明确标注
|
||||
- [ ] LLM 调用逐条串行 — 考虑用 `asyncio` + `httpx.AsyncClient` 并发处理多条 item,减少整体延迟
|
||||
@@ -0,0 +1,73 @@
|
||||
# 规划文档导航
|
||||
|
||||
## 使用原则
|
||||
|
||||
`plans/` 目录保存架构设计、实施计划、专题方案和历史问题分析。
|
||||
|
||||
不要把所有 plan 都当成当前权威事实。
|
||||
当前事实优先级应是:
|
||||
|
||||
1. `README.md`
|
||||
2. `docs/README.md`
|
||||
3. `docs/openclaw/README.md`
|
||||
4. `docs/openclaw/openclaw-handoff.md`
|
||||
5. `docs/openclaw/openclaw-orchestration-flow.md`
|
||||
6. `TODO.md`
|
||||
|
||||
`plans/` 更适合回答:
|
||||
|
||||
- 为什么这样设计
|
||||
- 某个能力是怎么分阶段落地的
|
||||
- 某次事故当时是怎么分析的
|
||||
|
||||
## 当前权威规划
|
||||
|
||||
这些文档仍然是当前协作时应优先阅读的规划基线:
|
||||
|
||||
- `reader-mcp-architecture-design.md`
|
||||
- MCP workflow service 的总体架构方向
|
||||
- `reader-mcp-implementation-plan.md`
|
||||
- 实施分阶段计划
|
||||
- `../TODO.md`
|
||||
- 当前任务状态与落地进展
|
||||
|
||||
## 已完成能力的专题方案
|
||||
|
||||
这些方案主要用于回看设计取舍,相关能力已经基本落地:
|
||||
|
||||
- `freshrss-pipeline-async-job-plan.md`
|
||||
- 主日报 async job 方案
|
||||
- `article-summary-async-job-plan.md`
|
||||
- 单篇总结 async job 方案
|
||||
- `resume-run-minimal-design.md`
|
||||
- `resume_run` 最小恢复语义设计
|
||||
|
||||
## 仍有参考价值的专题设计
|
||||
|
||||
- `keyword-cleanup-artifact-slimming-v1.md`
|
||||
- `keyword-cleanup-review-suggestions-layer-design.md`
|
||||
- `article-summary-prompt-independent.md`
|
||||
- `article-deep-summary-skill.md`
|
||||
- `docker-deployment-plan.md`
|
||||
|
||||
## 历史问题分析
|
||||
|
||||
- `issues/2026-04-06-reader-digest-sigterm.md`
|
||||
- 一次真实运行事故的分析
|
||||
|
||||
## 建议阅读顺序
|
||||
|
||||
如果是新接手维护:
|
||||
|
||||
1. `reader-mcp-architecture-design.md`
|
||||
2. `reader-mcp-implementation-plan.md`
|
||||
3. `../TODO.md`
|
||||
4. `../docs/openclaw/README.md`
|
||||
5. `../docs/openclaw/openclaw-handoff.md`
|
||||
|
||||
如果是在排查某一类能力:
|
||||
|
||||
- 主日报启动/轮询:看 `freshrss-pipeline-async-job-plan.md`
|
||||
- 恢复:看 `resume-run-minimal-design.md`
|
||||
- 单篇总结:看 `article-summary-async-job-plan.md`
|
||||
- 历史故障:看 `issues/2026-04-06-reader-digest-sigterm.md`
|
||||
@@ -0,0 +1,41 @@
|
||||
# Plan: article-deep-summary Skill
|
||||
|
||||
## 背景
|
||||
|
||||
项目已有两个 skill:
|
||||
- `keyword-cleanup-review`:关键词索引清理与建议
|
||||
- `llm-summary-review`:验证/修复 LLM 生成的日报摘要 JSON
|
||||
|
||||
两者均不覆盖「从单篇文章完整正文生成深度知识笔记」这一场景,需新增 `article-deep-summary` skill。
|
||||
|
||||
## 目标
|
||||
|
||||
封装 `summary_mcp.workflows.article_summary` 的使用方式,让 Claude 在用户需要对单篇文章做深度精读总结时,能正确触发并执行该 workflow。
|
||||
|
||||
## 三个 skill 的分工
|
||||
|
||||
| Skill | 触发时机 | 处理对象 |
|
||||
|-------|---------|----------|
|
||||
| keyword-cleanup-review | 审查/清理关键词索引 | term_stats.json + daily 文件 |
|
||||
| llm-summary-review | 已有摘要 JSON,需验证或修复 | outputs/result.json |
|
||||
| article-deep-summary | 有提取文件,想生成深度知识笔记 | *.extracted.json 的 plain_text |
|
||||
|
||||
## 核心入口
|
||||
|
||||
python -m summary_mcp.workflows.article_summary \
|
||||
--extracted outputs/freshrss/<run_id>/extracted/selected.extracted.json \
|
||||
--selected-ids <item_id> \
|
||||
--output-dir outputs/article-summaries/
|
||||
|
||||
## 关键代码路径
|
||||
|
||||
- Workflow:src/summary_mcp/workflows/article_summary.py
|
||||
- Validator:src/summary_mcp/validators/article_summary.py
|
||||
- 数据模型:src/summary_mcp/models/article_summary_result.py
|
||||
- Prompt:outputs/prompts/article-summary-prompt.txt
|
||||
|
||||
## 待办
|
||||
|
||||
- [x] 创建 skills/article-deep-summary/SKILL.md
|
||||
- [x] 创建 skills/article-deep-summary/agents/openai.yaml
|
||||
- [x] 用实际 extracted 文件跑一次验证输出(outputs/reference/article-summaries/信息过载时代,我的漏斗式阅读工作流.md)
|
||||
@@ -0,0 +1,757 @@
|
||||
# 单篇总结异步 job 最小版落地方案
|
||||
|
||||
## 1. 背景与问题定义
|
||||
|
||||
当前 reader 已经把日更 FreshRSS 主流程做成了带 `run-state.json` 的 runtime 模型:
|
||||
|
||||
- `src/summary_mcp/runtime/state_models.py`
|
||||
- `src/summary_mcp/runtime/run_store.py`
|
||||
- `src/summary_mcp/runtime/query_service.py`
|
||||
- `src/summary_mcp/workflows/freshrss_pipeline.py`
|
||||
|
||||
这条主链路已经具备:
|
||||
|
||||
- run / stage / artifact 的结构化状态
|
||||
- MCP 查询接口:`get_run_status` / `list_runs` / `list_run_artifacts`
|
||||
- 结果读取接口:`get_delivery_payload` / `get_run_report`
|
||||
|
||||
但“单篇总结”这条线目前还是同步调用:
|
||||
|
||||
- 核心逻辑:`src/summary_mcp/workflows/article_summary.py`
|
||||
- MCP 暴露:`src/summary_mcp/server.py` 中的 `generate_article_summaries`
|
||||
- CLI 辅助:`scripts/run_article_summaries.py`
|
||||
|
||||
现状问题已经很明确:
|
||||
|
||||
- 在 OpenClaw → MCP tool 这条链路里,`generate_article_summaries` 可能因为 tool 调用时长而 timeout
|
||||
- 但 reader 项目本体在 `.venv` 下直接跑 article summary,大约 37.5 秒即可成功
|
||||
- 这说明问题不一定在 summary 业务本身,而更可能在“同步工具调用 + 上层等待模型”这个包装层
|
||||
|
||||
所以目标不是先继续调 timeout,而是把单篇总结也纳入 **真正异步、可轮询、可落盘、可恢复基本状态** 的最小 job 模型里。
|
||||
|
||||
---
|
||||
|
||||
## 2. 为什么同步 MCP 不适合这一步
|
||||
|
||||
`generate_article_summaries` 当前在 `server.py` 里直接同步执行 `summarize_selected_articles(...)`,调用方必须一直阻塞等待,直到:
|
||||
|
||||
1. 读取 extracted payload
|
||||
2. 调用 LLM 生成总结
|
||||
3. 可选 repair retry
|
||||
4. 渲染 Markdown
|
||||
5. 写文件完成
|
||||
6. MCP tool 返回生成路径
|
||||
|
||||
这个模式对“几十秒级、依赖外部 LLM、可能重试”的任务不稳,核心问题有三层:
|
||||
|
||||
### 2.1 tool 调用时长不可控
|
||||
|
||||
`run_loop_payload()` 内部会发起外部 HTTP 请求,还可能做 repair retry。即便单次平均 37.5 秒,也已经接近很多上层编排系统的心理和技术超时边界。
|
||||
|
||||
### 2.2 调用方看不到中间状态
|
||||
|
||||
现在如果卡住,调用方只能等:
|
||||
|
||||
- 不知道是在读输入
|
||||
- 不知道是在调 LLM
|
||||
- 不知道是在重试
|
||||
- 不知道是否已经写出部分结果
|
||||
|
||||
这也是同步接口最烦的点:失败时只能看到“tool timeout / tool failed”,而不是“业务跑到哪一步了”。
|
||||
|
||||
### 2.3 与 reader 已有 runtime 风格不一致
|
||||
|
||||
FreshRSS 主流程已经是“run truth + status query + artifact read”的思路,而单篇总结仍然是黑箱同步函数。继续维持两套风格,只会让 SOP 更复杂:
|
||||
|
||||
- 日报主链路用 `run_id`
|
||||
- 单篇总结却要么同步等,要么退回 CLI fallback
|
||||
|
||||
这不利于后续把 `reader-digest-flow` 稳定成正式 SOP。
|
||||
|
||||
---
|
||||
|
||||
## 3. 本轮目标:最小真异步,不做大而全
|
||||
|
||||
这次只做 **单篇总结异步 job 最小版**,目标是:
|
||||
|
||||
> 让 OpenClaw 或其他调用方能先“启动单篇总结 job”,立即拿到 `job_id`,再通过状态接口轮询,最后读取输出文件/结果。
|
||||
|
||||
### 3.1 本轮必须做到的范围
|
||||
|
||||
1. 新增单篇总结 job 的 start/status/result 最小接口
|
||||
2. job 真正在后台执行,而不是 MCP 请求线程里阻塞到完成
|
||||
3. 状态落盘到文件,遵循 reader 当前 runtime 风格
|
||||
4. 复用现有 `summarize_selected_articles` 逻辑,不重写业务
|
||||
5. 输出仍然是现有 Markdown 文件,不改知识内容 schema
|
||||
|
||||
### 3.2 本轮明确不做
|
||||
|
||||
1. **不做通用队列系统**
|
||||
2. **不做数据库**
|
||||
3. **不做多 worker / 分布式调度**
|
||||
4. **不做取消 job / kill job**
|
||||
5. **不做并发配额控制**
|
||||
6. **不做 resume/retry from stage**
|
||||
7. **不把 article summary 一次性并入 freshrss `resume_run` 体系**
|
||||
8. **不改 summary prompt / validator / 输出格式**
|
||||
9. **不处理批量高吞吐场景优化**
|
||||
|
||||
一句话:这轮只解“同步 tool 容易 timeout,但业务本身能跑完”这个问题,不顺手扩成任务调度平台。
|
||||
|
||||
---
|
||||
|
||||
## 4. 接入当前 reader 结构的建议
|
||||
|
||||
### 4.1 复用现有 runtime 设计,但单独建 article summary job 命名空间
|
||||
|
||||
不建议把 article summary job 粗暴塞进现有 `freshrss_daily_digest` run 查询里混用一个 schema;更合适的是:
|
||||
|
||||
- 复用 `RunState / StageState / ArtifactRecord / RunStore` 这套思维
|
||||
- 但给单篇总结定义独立 workflow 名称与存储目录
|
||||
|
||||
建议:
|
||||
|
||||
- workflow: `article_summary_job`
|
||||
- run_type: `article_summary`
|
||||
- output root: `outputs/freshrss/article_summary_jobs/<job_id>/`
|
||||
|
||||
这样有几个好处:
|
||||
|
||||
- 不污染 `outputs/freshrss/rerun/`
|
||||
- 语义清楚:这是独立 job,不是假装自己是日报 rerun
|
||||
- 查询和排查时更直观
|
||||
|
||||
### 4.2 job 与结果 Markdown 解耦
|
||||
|
||||
job 目录只负责:
|
||||
|
||||
- 状态文件
|
||||
- 输入快照
|
||||
- artifact 索引
|
||||
- 执行报告
|
||||
|
||||
真正生成的总结 Markdown,仍然可以写到用户指定的 `output_dir`(或默认 `single_summaries/`)。
|
||||
|
||||
这样不破坏当前下游 SOP:
|
||||
|
||||
- `reader-digest-flow` 依然从原来的单篇总结输出目录拿 `.md`
|
||||
- job 目录只提供状态与索引,不强迫下游改结果路径约定
|
||||
|
||||
---
|
||||
|
||||
## 5. 新增工具 / API 设计
|
||||
|
||||
本轮建议新增 3 个 MCP tool,名字尽量和当前 runtime 风格一致。
|
||||
|
||||
## 5.1 `start_article_summary_job`
|
||||
|
||||
### 作用
|
||||
|
||||
启动一个后台 job,立即返回 `job_id`,不等待总结完成。
|
||||
|
||||
### 输入建议
|
||||
|
||||
```json
|
||||
{
|
||||
"extracted_path": "outputs/freshrss/rerun/<run_id>/extracted/item-01.extracted.json",
|
||||
"selected_ids": ["12345"],
|
||||
"output_dir": "outputs/freshrss/single_summaries/2026-04-10",
|
||||
"max_retries": 2,
|
||||
"timeout_seconds": 120,
|
||||
"llm_api_key": null,
|
||||
"llm_model": null,
|
||||
"llm_api_url": null
|
||||
}
|
||||
```
|
||||
|
||||
### 返回建议
|
||||
|
||||
```json
|
||||
{
|
||||
"job_id": "article-summary-20260410-144500-ab12cd34",
|
||||
"workflow": "article_summary_job",
|
||||
"run_type": "article_summary",
|
||||
"status": "running",
|
||||
"output_dir": "outputs/freshrss/article_summary_jobs/article-summary-20260410-144500-ab12cd34",
|
||||
"message": "Article summary job started successfully. Use get_article_summary_job_status to poll progress."
|
||||
}
|
||||
```
|
||||
|
||||
### 约束建议
|
||||
|
||||
- `selected_ids` 第一版允许多个,但建议由 OpenClaw 每次只传一篇或少量篇,避免一个 job 干太多事
|
||||
- `extracted_path` 必须存在,否则直接拒绝启动
|
||||
- `output_dir` 不传则按当前默认逻辑推导
|
||||
|
||||
---
|
||||
|
||||
## 5.2 `get_article_summary_job_status`
|
||||
|
||||
### 作用
|
||||
|
||||
查询 job 当前状态、阶段、进度、错误摘要、已注册 artifact。
|
||||
|
||||
### 输入
|
||||
|
||||
```json
|
||||
{
|
||||
"job_id": "article-summary-20260410-144500-ab12cd34"
|
||||
}
|
||||
```
|
||||
|
||||
### 返回建议
|
||||
|
||||
```json
|
||||
{
|
||||
"job_id": "article-summary-20260410-144500-ab12cd34",
|
||||
"workflow": "article_summary_job",
|
||||
"run_type": "article_summary",
|
||||
"status": "running",
|
||||
"current_stage": "generate_markdown",
|
||||
"started_at": "2026-04-10T14:45:00+08:00",
|
||||
"updated_at": "2026-04-10T14:45:23+08:00",
|
||||
"finished_at": null,
|
||||
"output_dir": "outputs/freshrss/article_summary_jobs/article-summary-20260410-144500-ab12cd34",
|
||||
"progress": {
|
||||
"completed_stage_count": 2,
|
||||
"running_stage_count": 1,
|
||||
"failed_stage_count": 0,
|
||||
"pending_stage_count": 1,
|
||||
"total_stage_count": 4
|
||||
},
|
||||
"artifacts": [...],
|
||||
"error_summary": null
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 5.3 `get_article_summary_job_result`
|
||||
|
||||
### 作用
|
||||
|
||||
当 job 成功后,返回结构化结果,供 OpenClaw 继续下游 IMA 沉淀。
|
||||
|
||||
### 输入
|
||||
|
||||
```json
|
||||
{
|
||||
"job_id": "article-summary-20260410-144500-ab12cd34"
|
||||
}
|
||||
```
|
||||
|
||||
### 返回建议
|
||||
|
||||
```json
|
||||
{
|
||||
"job_id": "article-summary-20260410-144500-ab12cd34",
|
||||
"status": "success",
|
||||
"written_paths": [
|
||||
"outputs/freshrss/single_summaries/2026-04-10/某篇文章标题.md"
|
||||
],
|
||||
"artifact": {
|
||||
"name": "job_result",
|
||||
"path": "outputs/freshrss/article_summary_jobs/article-summary-20260410-144500-ab12cd34/result.json",
|
||||
"kind": "json",
|
||||
"stage": "write_result"
|
||||
},
|
||||
"result": {
|
||||
"selected_ids": ["12345"],
|
||||
"written_paths": [
|
||||
"outputs/freshrss/single_summaries/2026-04-10/某篇文章标题.md"
|
||||
]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 行为建议
|
||||
|
||||
- 若 job 还没完成,返回当前状态 + 提示“not ready”
|
||||
- 若 job 失败,返回失败摘要,不硬抛文件不存在异常
|
||||
|
||||
---
|
||||
|
||||
## 6. job 状态文件设计
|
||||
|
||||
建议直接复用现有 `RunState` 模型,不另造一套 schema。
|
||||
|
||||
job 目录示例:
|
||||
|
||||
```text
|
||||
outputs/freshrss/article_summary_jobs/
|
||||
article-summary-20260410-144500-ab12cd34/
|
||||
run-state.json
|
||||
input.json
|
||||
result.json
|
||||
job-report.json
|
||||
```
|
||||
|
||||
## 6.1 `run-state.json`
|
||||
|
||||
建议直接沿用当前字段:
|
||||
|
||||
```json
|
||||
{
|
||||
"run_id": "article-summary-20260410-144500-ab12cd34",
|
||||
"workflow": "article_summary_job",
|
||||
"run_type": "article_summary",
|
||||
"status": "running",
|
||||
"current_stage": "generate_markdown",
|
||||
"started_at": "2026-04-10T14:45:00+08:00",
|
||||
"updated_at": "2026-04-10T14:45:23+08:00",
|
||||
"finished_at": null,
|
||||
"input": {
|
||||
"extracted_path": "outputs/freshrss/rerun/<run_id>/extracted/item-01.extracted.json",
|
||||
"selected_ids": ["12345"],
|
||||
"output_dir": "outputs/freshrss/single_summaries/2026-04-10",
|
||||
"max_retries": 2,
|
||||
"timeout_seconds": 120
|
||||
},
|
||||
"stages": [...],
|
||||
"artifacts": [...],
|
||||
"error": null,
|
||||
"recovery": {
|
||||
"resumable": false,
|
||||
"resume_from_stage": null,
|
||||
"last_success_stage": "load_input"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 第一版 stage 建议
|
||||
|
||||
建议只切 4 个 stage,够看即可:
|
||||
|
||||
1. `prepare_job`
|
||||
- 校验输入
|
||||
- 解析路径
|
||||
- 写 `input.json`
|
||||
|
||||
2. `load_input`
|
||||
- 读取 extracted payload
|
||||
- 确认 `selected_ids` 可匹配条目
|
||||
|
||||
3. `generate_markdown`
|
||||
- 调 `summarize_selected_articles(...)`
|
||||
- 这是主要耗时阶段
|
||||
|
||||
4. `write_result`
|
||||
- 写 `result.json`
|
||||
- 注册输出 artifact
|
||||
|
||||
这里不要把 LLM 调用再拆更多细 stage,否则最小版反而过度设计。
|
||||
|
||||
---
|
||||
|
||||
## 6.2 `input.json`
|
||||
|
||||
作用:保留启动请求快照,便于排查。
|
||||
|
||||
建议内容与 `run-state.input` 基本一致。
|
||||
|
||||
---
|
||||
|
||||
## 6.3 `result.json`
|
||||
|
||||
成功时写:
|
||||
|
||||
```json
|
||||
{
|
||||
"job_id": "article-summary-20260410-144500-ab12cd34",
|
||||
"selected_ids": ["12345"],
|
||||
"written_paths": [
|
||||
"outputs/freshrss/single_summaries/2026-04-10/某篇文章标题.md"
|
||||
],
|
||||
"completed_at": "2026-04-10T14:45:41+08:00"
|
||||
}
|
||||
```
|
||||
|
||||
失败时可以不写,或只写失败快照都行。最小版建议:
|
||||
|
||||
- 成功写 `result.json`
|
||||
- 失败只依赖 `run-state.json`
|
||||
|
||||
避免双份失败状态不一致。
|
||||
|
||||
---
|
||||
|
||||
## 7. 执行模型建议:优先子进程,不建议线程
|
||||
|
||||
### 7.1 推荐:子进程后台执行
|
||||
|
||||
最小真异步推荐模型:
|
||||
|
||||
- `start_article_summary_job` 负责:
|
||||
- 创建 job 目录
|
||||
- 初始化 `run-state.json`
|
||||
- 通过 `subprocess.Popen(...)` 启动一个独立 Python 进程执行 job runner
|
||||
- 立即返回 `job_id`
|
||||
|
||||
后台 runner 再去:
|
||||
|
||||
- 读取 `input.json`
|
||||
- 用 `RunStore` 更新状态
|
||||
- 调用 `summarize_selected_articles(...)`
|
||||
- 写 `result.json`
|
||||
- finish/fail run
|
||||
|
||||
### 7.2 为什么不推荐线程
|
||||
|
||||
虽然线程实现看起来更省事,但不适合作为 reader 的正式最小异步落地:
|
||||
|
||||
1. **MCP server 进程重启后线程直接丢失**
|
||||
2. 线程状态不天然可恢复,容易出现“状态文件还在 running,但线程没了”
|
||||
3. 未来要做健康检查/孤儿 job 检测时,线程模型更难收口
|
||||
|
||||
### 7.3 为什么子进程更贴当前项目风格
|
||||
|
||||
reader 现在本来就偏“文件产物 + runtime 状态真相”风格。子进程模式有天然优势:
|
||||
|
||||
- 和 CLI/fallback 思维一致
|
||||
- 进程边界清楚
|
||||
- `run-state.json` 由实际执行者写,职责清晰
|
||||
- 将来如果要做 orphan detection / stale running job 修复,也容易补
|
||||
|
||||
### 7.4 本轮不做进程管理增强
|
||||
|
||||
最小版里,不要求:
|
||||
|
||||
- 记录 PID 后做 kill/cancel
|
||||
- 自动清理僵尸进程
|
||||
- 守护进程/worker 池
|
||||
|
||||
但建议在 `input.json` 或 `run-state.input` 里附带:
|
||||
|
||||
- `launcher_pid`
|
||||
- `runner_command`
|
||||
|
||||
方便排障。
|
||||
|
||||
---
|
||||
|
||||
## 8. 与现有 `summarize_selected_articles` 的复用关系
|
||||
|
||||
核心原则:**不重写总结业务,只包一层 job runner。**
|
||||
|
||||
### 8.1 直接复用的部分
|
||||
|
||||
`src/summary_mcp/workflows/article_summary.py` 已经做了:
|
||||
|
||||
- 读取 extracted payload
|
||||
- 根据 `selected_ids` 找条目
|
||||
- 调用 `run_loop_payload(...)`
|
||||
- 渲染 Markdown
|
||||
- 写到 `output_dir`
|
||||
- 返回 `list[Path]`
|
||||
|
||||
这些都继续用。
|
||||
|
||||
### 8.2 最小新增建议
|
||||
|
||||
建议只新增一层 runtime/service,例如:
|
||||
|
||||
- `src/summary_mcp/runtime/article_summary_jobs.py`
|
||||
|
||||
职责:
|
||||
|
||||
- 生成 `job_id`
|
||||
- 创建 job 目录
|
||||
- 初始化 `RunStore`
|
||||
- 启动 runner 子进程
|
||||
- 提供 status/result 查询
|
||||
- 在 runner 里调用 `summarize_selected_articles`
|
||||
|
||||
### 8.3 是否需要改 `summarize_selected_articles`
|
||||
|
||||
最小版尽量少改,只建议加两类低风险增强:
|
||||
|
||||
1. **可选输入校验增强**
|
||||
- 如果 `selected_ids` 一个都匹配不到,显式报错
|
||||
- 避免“成功返回空列表”却让 job 看起来像成功
|
||||
|
||||
2. **可选 hook / telemetry(非必须)**
|
||||
- 如果后面需要更细粒度写 stage 输出,可再加
|
||||
- 但第一版没必要为了观测性重构函数
|
||||
|
||||
结论:
|
||||
|
||||
- 第一版优先保持 `summarize_selected_articles` 基本不动
|
||||
- job 层只把它作为黑盒业务函数调用
|
||||
|
||||
---
|
||||
|
||||
## 9. 建议新增代码组织
|
||||
|
||||
建议新增文件:
|
||||
|
||||
```text
|
||||
src/summary_mcp/runtime/article_summary_jobs.py
|
||||
scripts/run_article_summary_job.py
|
||||
```
|
||||
|
||||
### 9.1 `article_summary_jobs.py`
|
||||
|
||||
建议包含:
|
||||
|
||||
- `start_article_summary_job(...)`
|
||||
- `run_article_summary_job(...)`
|
||||
- `get_article_summary_job_status(...)`
|
||||
- `get_article_summary_job_result(...)`
|
||||
- 若干私有 helper:
|
||||
- job id 生成
|
||||
- job dir 解析
|
||||
- result artifact 读取
|
||||
|
||||
### 9.2 `scripts/run_article_summary_job.py`
|
||||
|
||||
作用:作为子进程 runner 入口。
|
||||
|
||||
例如:
|
||||
|
||||
```bash
|
||||
python scripts/run_article_summary_job.py --job-id article-summary-...
|
||||
```
|
||||
|
||||
runner 只做一件事:
|
||||
|
||||
- 根据 `job_id` 找到 job 目录和 `input.json`
|
||||
- 真正执行 job
|
||||
|
||||
这样避免在 `Popen("python -c ...")` 里塞长字符串,也方便本地调试。
|
||||
|
||||
---
|
||||
|
||||
## 10. 对 `reader-digest-flow` SOP 的影响
|
||||
|
||||
这块是重点,因为老大的实际痛点就在这。
|
||||
|
||||
### 10.1 现状 SOP
|
||||
|
||||
当前 skill 已经有硬规则:
|
||||
|
||||
- 优先走 MCP `generate_article_summaries`
|
||||
- MCP timeout 时,立刻 fallback 到 reader 本地 `.venv`
|
||||
|
||||
这个 fallback 现在是必要的,但它本质是在补“单篇总结没有正式异步接口”。
|
||||
|
||||
### 10.2 引入异步 job 后的推荐 SOP
|
||||
|
||||
建议调整为:
|
||||
|
||||
1. OpenClaw 在用户确认保留文章后
|
||||
2. 调 `start_article_summary_job`
|
||||
3. 拿到 `job_id`
|
||||
4. 轮询 `get_article_summary_job_status`
|
||||
5. 成功后调 `get_article_summary_job_result`
|
||||
6. 再继续 IMA 格式化与上传
|
||||
|
||||
### 10.3 对 skill 文档的影响
|
||||
|
||||
`reader-digest-flow` 需要后续补一条新规则:
|
||||
|
||||
- 单篇总结的正式生产路径从“同步 MCP + 本地 fallback”升级为“异步 job + 状态轮询”
|
||||
- 本地 `.venv` CLI fallback 仍保留,但降级为:
|
||||
- job 启动失败
|
||||
- job runner 异常
|
||||
- reader 服务端出现系统性问题时的应急路径
|
||||
|
||||
### 10.4 用户体验改善
|
||||
|
||||
异步 job 后,OpenClaw 可给出更像正式系统的反馈:
|
||||
|
||||
- “已启动单篇总结任务,正在生成”
|
||||
- “当前状态:generate_markdown”
|
||||
- “已完成,准备继续沉淀到 IMA”
|
||||
|
||||
而不是现在的:
|
||||
|
||||
- 直接卡住几十秒
|
||||
- 然后 tool timeout
|
||||
- 再走一套 fallback
|
||||
|
||||
---
|
||||
|
||||
## 11. 验证方案
|
||||
|
||||
本轮验证不要追求大而全,按 4 层做就够。
|
||||
|
||||
### 11.1 单元级验证
|
||||
|
||||
目标:确保 job 状态文件和结果文件行为正确。
|
||||
|
||||
建议覆盖:
|
||||
|
||||
1. `start_article_summary_job` 能创建 job 目录与 `run-state.json`
|
||||
2. 输入路径不存在时,启动直接失败
|
||||
3. `get_article_summary_job_status` 能正确读取状态
|
||||
4. job 成功后 `get_article_summary_job_result` 返回 `written_paths`
|
||||
5. job 失败后状态为 `failed`,并带错误摘要
|
||||
|
||||
### 11.2 本地集成验证
|
||||
|
||||
用一个真实 extracted 文件跑:
|
||||
|
||||
1. 启动 job
|
||||
2. 轮询 status
|
||||
3. 成功后检查:
|
||||
- `result.json` 存在
|
||||
- Markdown 文件存在
|
||||
- 路径正确
|
||||
|
||||
### 11.3 OpenClaw 链路验证
|
||||
|
||||
在实际 `reader-digest-flow` 环节,用一篇已确认保留的文章做:
|
||||
|
||||
1. 启动 async job
|
||||
2. 等待成功
|
||||
3. 继续做 IMA markdown 整理与上传
|
||||
4. 确认整个链路不再因为同步 tool timeout 中断
|
||||
|
||||
### 11.4 异常验证
|
||||
|
||||
至少测 3 类异常:
|
||||
|
||||
1. `selected_ids` 不存在
|
||||
2. LLM 接口失败 / 超时
|
||||
3. runner 进程异常退出
|
||||
|
||||
预期:
|
||||
|
||||
- `run-state.json` 最终为 `failed`
|
||||
- `error_summary` 可读
|
||||
- 调用方能明确知道失败,而不是只看到 transport timeout
|
||||
|
||||
---
|
||||
|
||||
## 12. 风险与回滚方案
|
||||
|
||||
## 12.1 风险
|
||||
|
||||
### 风险 1:后台子进程成功启动,但状态长期卡在 running
|
||||
|
||||
常见原因:
|
||||
|
||||
- runner 进程崩了
|
||||
- server 重启时某些路径没写全
|
||||
- 子进程命令不对
|
||||
|
||||
缓解:
|
||||
|
||||
- runner 启动前就先写 `run-state.json`
|
||||
- runner 一进来先更新 `prepare_job` / `load_input`
|
||||
- 后续可加“stale running 超时判定”,但第一版先不做自动修复
|
||||
|
||||
### 风险 2:复用旧函数导致“空输出也算成功”
|
||||
|
||||
当前 `summarize_selected_articles()` 如果没匹配到文章或全部失败,存在返回空列表的可能。
|
||||
|
||||
缓解:
|
||||
|
||||
- job runner 里把“`written_paths` 为空”视为失败
|
||||
- 或者顺手在 `summarize_selected_articles` 里补显式校验
|
||||
|
||||
### 风险 3:同一时刻大量 job 并发,LLM 调用被打爆
|
||||
|
||||
第一版不解决系统级并发控制。
|
||||
|
||||
缓解:
|
||||
|
||||
- SOP 层先按单篇/少量串行用
|
||||
- skill 层避免一口气启动很多 job
|
||||
|
||||
### 风险 4:状态目录与结果目录分离,排查时容易迷路
|
||||
|
||||
缓解:
|
||||
|
||||
- 在 `result.json` 和 artifact 元数据里明确记录 `written_paths`
|
||||
- 在 `run-state.input.output_dir` 中保留结果目录
|
||||
|
||||
---
|
||||
|
||||
## 12.2 回滚方案
|
||||
|
||||
这个方案很好回滚,因为是“新增,不替换”。
|
||||
|
||||
### 回滚原则
|
||||
|
||||
- 保留现有 `generate_article_summaries`
|
||||
- 保留现有 `scripts/run_article_summaries.py`
|
||||
- 新增 async job 接口如果不稳定,直接停止在 OpenClaw 层使用即可
|
||||
|
||||
### 回滚路径
|
||||
|
||||
1. 停止调用 `start_article_summary_job`
|
||||
2. 恢复回原 SOP:
|
||||
- 先尝试同步 MCP `generate_article_summaries`
|
||||
- 失败则本地 `.venv` fallback
|
||||
3. async job 相关代码保留但不作为正式入口
|
||||
|
||||
也就是说,这轮改动不会堵死当前生产路径,风险可控。
|
||||
|
||||
---
|
||||
|
||||
## 13. 实施顺序(建议按这个顺序落地)
|
||||
|
||||
### Phase A:先把方案落地成最小代码骨架
|
||||
|
||||
1. 新增 `src/summary_mcp/runtime/article_summary_jobs.py`
|
||||
2. 新增 job 目录常量与 helper
|
||||
3. 新增 `scripts/run_article_summary_job.py`
|
||||
4. 实现 runner 内部对 `summarize_selected_articles` 的调用
|
||||
5. 先本地命令验证 job 可跑通
|
||||
|
||||
### Phase B:再把 MCP 接口接上
|
||||
|
||||
6. 在 `server.py` 新增:
|
||||
- `start_article_summary_job`
|
||||
- `get_article_summary_job_status`
|
||||
- `get_article_summary_job_result`
|
||||
7. 本地 MCP 调用验证
|
||||
|
||||
### Phase C:最后接 OpenClaw SOP
|
||||
|
||||
8. 更新 `reader-digest-flow` skill,把正式路径切到 async job
|
||||
9. 保留本地 `.venv` fallback 作为应急方案
|
||||
10. 跑一次真实日报保留文章沉淀闭环
|
||||
|
||||
---
|
||||
|
||||
## 14. 推荐的最小返回/状态语义
|
||||
|
||||
为了跟现有 runtime 风格一致,建议继续用:
|
||||
|
||||
- `status`: `running` / `success` / `failed`
|
||||
- `current_stage`: 当前 stage 名
|
||||
- `artifacts`: 注册产物列表
|
||||
- `error_summary`: 结构化错误
|
||||
|
||||
第一版不必引入:
|
||||
|
||||
- `queued`
|
||||
- `cancelled`
|
||||
- `retrying`
|
||||
- `paused`
|
||||
|
||||
避免状态机一开始就复杂化。
|
||||
|
||||
---
|
||||
|
||||
## 15. 结论 / 拍板建议
|
||||
|
||||
拍板建议很直接:
|
||||
|
||||
1. **这件事值得做,而且优先级高**,因为它正好卡在当前日报 SOP 的真实痛点上
|
||||
2. **不建议继续先修同步 timeout**,因为同步模型本身就不适合几十秒级、外部 LLM 驱动的任务
|
||||
3. **最小真异步应采用“文件状态 + 后台子进程 + start/status/result 三接口”**
|
||||
4. **业务层严格复用 `summarize_selected_articles`**,不要为异步化重写 summary 核心逻辑
|
||||
5. **先把单篇总结 job 做成独立小 runtime 命名空间**,不要急着并入 freshrss 主 run 的 resume 体系
|
||||
|
||||
如果只允许做一轮最小落地,我建议就做到:
|
||||
|
||||
- `start_article_summary_job`
|
||||
- `get_article_summary_job_status`
|
||||
- `get_article_summary_job_result`
|
||||
- `outputs/freshrss/article_summary_jobs/<job_id>/run-state.json`
|
||||
- 子进程 runner
|
||||
|
||||
这套已经足够把当前 OpenClaw timeout 问题从“同步等待”改成“正式异步轮询”,并且几乎不碰无关模块。
|
||||
@@ -0,0 +1,40 @@
|
||||
# 规划:单篇文章总结 Prompt 独立化
|
||||
|
||||
## 背景
|
||||
|
||||
单篇精读总结(article_summary.py)目前复用日报摘要的同一份 prompt,
|
||||
导致输出是 2-3 句短摘要卡片,没有体现正文深度。
|
||||
目标是让单篇精读总结彻底独立,输出知识沉淀笔记。
|
||||
|
||||
## 已确认决策
|
||||
|
||||
- Validator 方案:方案 B,新增 ArticleSummaryResult 模型 + 专用校验函数
|
||||
- 输出语言:中文
|
||||
- 无技术内容时 key_methods / important_details 返回空数组 []
|
||||
- run_loop_payload 新增可选 validator 参数,None 时保持原有行为,日报链路不受影响
|
||||
|
||||
## 涉及文件
|
||||
|
||||
- outputs/prompts/article-summary-prompt.txt 新建
|
||||
- src/summary_mcp/models/article_summary_result.py 新建
|
||||
- src/summary_mcp/validators/article_summary.py 新建
|
||||
- src/summary_mcp/core/summary_loop.py 修改
|
||||
- src/summary_mcp/workflows/article_summary.py 修改
|
||||
|
||||
## 新增字段(ArticleSummaryResult)
|
||||
|
||||
- core_conclusion:核心结论,1-2 句
|
||||
- main_argument:主要论点
|
||||
- key_methods:关键方法,无内容返回 []
|
||||
- important_details:重要细节,无内容返回 []
|
||||
- reusable_insights:可复用启发
|
||||
- keywords / topics / category / worth_keeping / reason
|
||||
|
||||
## Markdown 输出格式
|
||||
|
||||
核心结论 / 主要论点 / 关键方法 / 重要细节 / 可复用启发 / 关键词 / 主题
|
||||
|
||||
## 约束
|
||||
|
||||
- freshrss_pipeline.py 不变
|
||||
- summary_loop.py 只加参数,默认行为不变
|
||||
@@ -0,0 +1,277 @@
|
||||
# reader MCP Docker 部署计划
|
||||
|
||||
## 1. 目标
|
||||
|
||||
将 reader 作为正式 MCP workflow service 以 Docker 方式部署,满足以下原则:
|
||||
|
||||
1. 服务运行在容器内
|
||||
2. 运行态与产物必须外置挂载,不闷在容器内
|
||||
3. 配置统一记录在 `.env`
|
||||
4. 读写行为与当前仓库约定保持一致
|
||||
5. OpenClaw 后续可将该服务作为正式上游 MCP 使用
|
||||
|
||||
---
|
||||
|
||||
## 2. 部署原则
|
||||
|
||||
### 2.1 容器职责
|
||||
|
||||
容器只负责:
|
||||
|
||||
- 提供 reader MCP 服务运行环境
|
||||
- 加载 reader 代码与依赖
|
||||
- 读取挂载进来的配置与状态目录
|
||||
- 对外暴露 MCP 服务入口
|
||||
|
||||
### 2.2 宿主机职责
|
||||
|
||||
宿主机负责持久化:
|
||||
|
||||
- 配置文件
|
||||
- 运行态
|
||||
- outputs 产物
|
||||
- 数据目录
|
||||
- configs
|
||||
|
||||
### 2.3 配置收口原则
|
||||
|
||||
所有环境配置统一放在 `.env`,避免:
|
||||
|
||||
- 零散写在 compose 内
|
||||
- 零散写在 shell 命令里
|
||||
- 零散写在 OpenClaw skill 里
|
||||
|
||||
---
|
||||
|
||||
## 3. 建议部署目录
|
||||
|
||||
建议在 reader 仓库内准备标准部署结构:
|
||||
|
||||
```text
|
||||
/home/ubuntu/zhu/github/reader/
|
||||
Dockerfile
|
||||
docker-compose.yml
|
||||
.env
|
||||
outputs/
|
||||
data/
|
||||
configs/
|
||||
knowledge-base/
|
||||
```
|
||||
|
||||
说明:
|
||||
|
||||
- `Dockerfile`:构建 reader MCP 服务镜像
|
||||
- `docker-compose.yml`:单服务部署编排
|
||||
- `.env`:统一环境变量
|
||||
- `outputs/`:产物、run-state、digest、payload 等外置持久化
|
||||
- `data/`:term index 等数据外置持久化
|
||||
- `configs/`:reader 运行配置外置持久化
|
||||
- `knowledge-base/`:如当前 reader/skill 仍会依赖本地知识目录,可继续挂载
|
||||
|
||||
---
|
||||
|
||||
## 4. 必须挂载的目录 / 文件
|
||||
|
||||
### 必须挂载
|
||||
|
||||
- `.env`
|
||||
- `outputs/`
|
||||
- `data/`
|
||||
- `configs/`
|
||||
|
||||
### 建议挂载
|
||||
|
||||
- `knowledge-base/`
|
||||
|
||||
### 通常不必挂载
|
||||
|
||||
- `docs/`
|
||||
- `plans/`
|
||||
- `.git/`
|
||||
|
||||
---
|
||||
|
||||
## 5. `.env` 统一配置建议
|
||||
|
||||
至少应包含以下配置:
|
||||
|
||||
### FreshRSS
|
||||
|
||||
- `FRESHRSS_API_BASE_URL`
|
||||
- `FRESHRSS_USERNAME`
|
||||
- `FRESHRSS_API_PASSWORD`
|
||||
|
||||
### 主 LLM
|
||||
|
||||
- `LLM_API_URL`
|
||||
- `LLM_API_KEY`
|
||||
- `LLM_MODEL`
|
||||
|
||||
### 单篇总结专用 LLM(如已使用)
|
||||
|
||||
- `ARTICLE_SUMMARY_API_URL`
|
||||
- `ARTICLE_SUMMARY_API_KEY`
|
||||
- `ARTICLE_SUMMARY_MODEL`
|
||||
|
||||
### IMA(如 reader / skill 仍依赖这些配置约定)
|
||||
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_ID`
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_NAME`
|
||||
|
||||
### 运行控制
|
||||
|
||||
- `PYTHONUNBUFFERED=1`
|
||||
- 视需要增加日志级别等配置
|
||||
|
||||
原则:
|
||||
|
||||
- 所有会影响服务行为的环境项,都优先进入 `.env`
|
||||
- compose 文件只引用 `.env`,不在 compose 里硬编码业务参数
|
||||
|
||||
---
|
||||
|
||||
## 6. Dockerfile 设计建议
|
||||
|
||||
### 目标
|
||||
|
||||
- 使用 Python 3.11
|
||||
- 安装 reader 依赖
|
||||
- 默认启动 MCP 服务入口
|
||||
|
||||
### 建议思路
|
||||
|
||||
1. 基于 `python:3.11-slim`
|
||||
2. 设置工作目录到 `/app`
|
||||
3. 复制仓库代码
|
||||
4. 安装依赖(如 `pip install -e .`)
|
||||
5. 默认启动 reader MCP 服务
|
||||
|
||||
### 启动入口
|
||||
|
||||
优先使用当前正式服务入口,例如:
|
||||
|
||||
- `summary-mcp`
|
||||
|
||||
如果后续 reader 明确切换到别的稳定入口,再同步更新。
|
||||
|
||||
---
|
||||
|
||||
## 7. docker-compose 设计建议
|
||||
|
||||
建议先保持单服务简单结构,例如:
|
||||
|
||||
- service 名称:`reader-mcp`
|
||||
- `env_file: .env`
|
||||
- 挂载:
|
||||
- `./outputs:/app/outputs`
|
||||
- `./data:/app/data`
|
||||
- `./configs:/app/configs`
|
||||
- `./knowledge-base:/app/knowledge-base`(如需要)
|
||||
- `./.env:/app/.env:ro`(可选,若程序直接读取文件)
|
||||
- `restart: unless-stopped`
|
||||
|
||||
如果当前 MCP 服务是 stdio 型而不是 HTTP 型,需要进一步明确:
|
||||
|
||||
- 它是由 OpenClaw 以本地进程方式拉起
|
||||
- 还是以常驻 sidecar / gateway adapter 方式挂接
|
||||
|
||||
因此 compose 的最终 command 需要结合实际接入方式确认。
|
||||
|
||||
---
|
||||
|
||||
## 8. 部署前确认项
|
||||
|
||||
在正式执行前,需要先确认以下问题:
|
||||
|
||||
### 8.1 MCP 连接方式
|
||||
|
||||
必须确认 reader MCP 服务的正式接入方式是:
|
||||
|
||||
1. **stdio 型**:OpenClaw/调用方本地拉起进程
|
||||
2. **HTTP/SSE 型**:服务常驻监听端口,OpenClaw 远程连接
|
||||
|
||||
这会直接影响:
|
||||
|
||||
- Docker command
|
||||
- 是否需要端口映射
|
||||
- OpenClaw 接入配置
|
||||
|
||||
### 8.2 当前 `summary-mcp` 的服务形态
|
||||
|
||||
需要确认:
|
||||
|
||||
- 现有 `summary-mcp` 是 FastMCP stdio 默认模式
|
||||
- 还是已有可直接 HTTP 化的运行方式
|
||||
|
||||
在这点没确认前,不要盲目写死端口暴露方案。
|
||||
|
||||
### 8.3 OpenClaw 侧接入点
|
||||
|
||||
部署完成后,还需要明确 OpenClaw 将如何引用该 MCP 服务:
|
||||
|
||||
- 本机命令型 MCP
|
||||
- Docker 内服务桥接
|
||||
- 或其它现有 OpenClaw MCP 配置方式
|
||||
|
||||
---
|
||||
|
||||
## 9. 执行顺序(建议)
|
||||
|
||||
### Phase A:部署方案落地
|
||||
|
||||
1. 确认 MCP 服务连接方式(stdio / HTTP)
|
||||
2. 确认最终 Dockerfile 启动命令
|
||||
3. 确认 compose 结构与挂载目录
|
||||
4. 整理 `.env` 字段
|
||||
|
||||
### Phase B:容器化实现
|
||||
|
||||
1. 新建/更新 `Dockerfile`
|
||||
2. 新建/更新 `docker-compose.yml`
|
||||
3. 检查 `.dockerignore`
|
||||
4. 核对路径是否与仓库内当前代码一致
|
||||
|
||||
### Phase C:本地部署验证
|
||||
|
||||
1. `docker compose build`
|
||||
2. `docker compose up -d`
|
||||
3. 验证服务启动
|
||||
4. 验证容器外 `outputs/`、`data/` 等是否正常落盘
|
||||
|
||||
### Phase D:OpenClaw 接入验证
|
||||
|
||||
1. 让 OpenClaw 通过正式 MCP 路径连接 reader
|
||||
2. 真跑一轮:
|
||||
- run
|
||||
- status
|
||||
- payload
|
||||
- report
|
||||
3. 如有需要,验证一次最小 `resume_run`
|
||||
|
||||
---
|
||||
|
||||
## 10. 当前不在本轮范围内的事
|
||||
|
||||
本轮部署计划不直接处理:
|
||||
|
||||
- `rerun_stage`
|
||||
- 更复杂的后台任务系统
|
||||
- 多实例部署
|
||||
- 横向扩展
|
||||
- 生产告警体系
|
||||
|
||||
本轮只做:
|
||||
|
||||
- 单实例
|
||||
- Docker 化
|
||||
- 配置收口
|
||||
- 挂载持久化
|
||||
- OpenClaw 可正式接入
|
||||
|
||||
---
|
||||
|
||||
## 11. 一句话结论
|
||||
|
||||
reader 的下一步不是继续堆内部接口,而是:
|
||||
|
||||
**以 Docker 正式部署成 MCP workflow service,配置进 `.env`,状态和产物目录挂载到宿主机,然后由 OpenClaw 按正式 MCP 编排路径真实接入和验证。**
|
||||
@@ -0,0 +1,229 @@
|
||||
# FreshRSS 主日报异步 job 方案
|
||||
|
||||
## 背景
|
||||
|
||||
当前 `run_freshrss_openclaw_pipeline` 虽然已经作为正式 MCP workflow 入口存在,但执行模型仍是**同步 MCP 调用**。这会带来几个现实问题:
|
||||
|
||||
1. OpenClaw / MCP wrapper 存在超时风险,尤其是 5-10 篇的正式日报批次。
|
||||
2. wrapper timeout 与真实 run 是否已落地,容易出现语义分离。
|
||||
3. 当前已有 `run-state.json`、`get_run_status`、`get_run_report`、`get_delivery_payload`,但**启动层**仍然是同步调用,不利于正式生产链路稳定运行。
|
||||
4. 单篇总结已经验证了“最小 async job + 轮询状态 + 读取结果”模型可行,主日报 run 应收敛到同一套运行模式。
|
||||
|
||||
## 目标
|
||||
|
||||
将 FreshRSS 主日报 run 改造成与 article-summary 类似的**最小真异步 job**:
|
||||
|
||||
- 启动即返回 `job_id`
|
||||
- 真正执行由后台子进程完成
|
||||
- 状态可轮询
|
||||
- 成功后可读取结构化结果
|
||||
- 业务逻辑继续复用既有 `run_freshrss_pipeline(...)`
|
||||
- 不推翻现有 run-state / result query 能力
|
||||
|
||||
## 非目标
|
||||
|
||||
本阶段不做:
|
||||
|
||||
- 分布式任务队列
|
||||
- 多 worker 调度
|
||||
- 任意 stage 的后台恢复编排
|
||||
- 并发控制中心
|
||||
- 主流程与 article-summary job 的通用抽象框架一次性大重构
|
||||
|
||||
先做最小可用。
|
||||
|
||||
## 设计原则
|
||||
|
||||
1. **启动层异步化,执行核心不重写**
|
||||
- `run_freshrss_pipeline(...)` 继续是主业务逻辑真相。
|
||||
- async job 只负责启动、状态持久化、结果回读。
|
||||
|
||||
2. **run truth 与 job truth 分层**
|
||||
- job truth:这次异步任务有没有启动、运行到哪一步、是否成功。
|
||||
- run truth:真正的 freshrss workflow 输出与 `run-state.json`。
|
||||
|
||||
3. **OpenClaw 正式生产默认改为 async start path**
|
||||
- 启动走 async job
|
||||
- 状态和结果优先先看 job
|
||||
- 真正业务产物仍由现有 run 查询工具承接
|
||||
|
||||
4. **与 article-summary job 尽量同构**
|
||||
- 目录结构
|
||||
- `run-state.json` / `input.json` / `result.json` / `job-report.json`
|
||||
- 后台 runner 脚本
|
||||
|
||||
## 拟新增能力
|
||||
|
||||
### MCP tools
|
||||
|
||||
新增 3 个工具:
|
||||
|
||||
- `start_freshrss_pipeline_job`
|
||||
- `get_freshrss_pipeline_job_status`
|
||||
- `get_freshrss_pipeline_job_result`
|
||||
|
||||
### job 目录
|
||||
|
||||
固定目录:
|
||||
|
||||
`outputs/freshrss/pipeline_jobs/<job_id>/`
|
||||
|
||||
至少包含:
|
||||
|
||||
- `run-state.json`
|
||||
- `input.json`
|
||||
- `result.json`(成功时)
|
||||
- `job-report.json`
|
||||
|
||||
### 执行模型
|
||||
|
||||
- `start_freshrss_pipeline_job` 写入 input + 初始化 job state
|
||||
- 后台 `subprocess.Popen(...)` 启动 runner
|
||||
- runner 内部调用 `run_freshrss_pipeline(...)`
|
||||
- 成功后把 `run_id`、核心产物路径、关键计数写入 `result.json`
|
||||
|
||||
## job 输入参数
|
||||
|
||||
与现有 `run_freshrss_openclaw_pipeline` 尽量对齐:
|
||||
|
||||
- `limit`
|
||||
- `mark_read`
|
||||
- `include_read`
|
||||
- `debug_artifacts`
|
||||
- `continuation`
|
||||
- `timeout_seconds`
|
||||
- `max_retries`
|
||||
- `stream_id`
|
||||
- `api_base_url`
|
||||
- `username`
|
||||
- `api_password`
|
||||
- `llm_api_key`
|
||||
- `llm_model`
|
||||
- `llm_api_url`
|
||||
- `context`
|
||||
- `run_id`
|
||||
- `date_value`
|
||||
- `output_dir`
|
||||
- `include_item_reports`
|
||||
|
||||
## 返回语义
|
||||
|
||||
### start
|
||||
|
||||
返回:
|
||||
|
||||
- `job_id`
|
||||
- `workflow`
|
||||
- `run_type`
|
||||
- `status=running`
|
||||
- `output_dir`
|
||||
- `message`
|
||||
|
||||
### status
|
||||
|
||||
返回:
|
||||
|
||||
- `job_id`
|
||||
- `status`
|
||||
- `current_stage`
|
||||
- `started_at` / `updated_at` / `finished_at`
|
||||
- `progress`
|
||||
- `artifacts`
|
||||
- `error_summary`
|
||||
- 若主 run 已创建,可附带 `linked_run_id`
|
||||
|
||||
### result
|
||||
|
||||
成功时返回:
|
||||
|
||||
- `job_id`
|
||||
- `status=success`
|
||||
- `run_id`
|
||||
- `delivery_output`
|
||||
- `report_output`
|
||||
- `digest_brief_output`
|
||||
- `pulled_count`
|
||||
- `delivered_count`
|
||||
- `marked_read_count`
|
||||
- `artifact`
|
||||
- `result`
|
||||
|
||||
## stages 建议
|
||||
|
||||
最小 job stages:
|
||||
|
||||
1. `prepare_job`
|
||||
2. `load_input`
|
||||
3. `run_pipeline`
|
||||
4. `write_result`
|
||||
|
||||
其中 `run_pipeline` 内部仍由现有 freshrss workflow 自己写它的 run-state。
|
||||
|
||||
## 与现有同步入口的关系
|
||||
|
||||
### 保留
|
||||
|
||||
`run_freshrss_openclaw_pipeline` 暂时保留,作为:
|
||||
|
||||
- debug / light path
|
||||
- 本地调试工具
|
||||
- 向后兼容路径
|
||||
|
||||
### 正式语义调整
|
||||
|
||||
文档与 OpenClaw handoff 中,主日报正式生产默认启动入口改为:
|
||||
|
||||
- `start_freshrss_pipeline_job`
|
||||
|
||||
同步入口降级为:
|
||||
|
||||
- debug / fallback
|
||||
- 小批量验证
|
||||
|
||||
## OpenClaw 编排建议
|
||||
|
||||
新的推荐路径:
|
||||
|
||||
1. `start_freshrss_pipeline_job`
|
||||
2. `get_freshrss_pipeline_job_status`
|
||||
3. 成功后 `get_freshrss_pipeline_job_result`
|
||||
4. 后续仍用:
|
||||
- `get_run_status`
|
||||
- `get_delivery_payload`
|
||||
- `get_run_report`
|
||||
- `list_run_artifacts`
|
||||
|
||||
## 风险点
|
||||
|
||||
1. **job 成功但 run 部分失败**
|
||||
- 允许,job 结果应以真实 `run_freshrss_pipeline(...)` 返回为准。
|
||||
- `run_id` + `report_output` 仍是最终真相。
|
||||
|
||||
2. **runner 崩溃但来不及写 result**
|
||||
- 需保证 `job-report.json` 至少能写下失败摘要。
|
||||
|
||||
3. **重复状态源导致混淆**
|
||||
- 文档必须明确:
|
||||
- job state 管“启动任务”
|
||||
- run state 管“业务工作流真相”
|
||||
|
||||
4. **同步 / 异步双入口长期漂移**
|
||||
- 必须要求 async job 内部直接复用 `run_freshrss_pipeline(...)`
|
||||
- 禁止再实现一套平行主流程
|
||||
|
||||
## 验收标准
|
||||
|
||||
1. 能通过 MCP 启动一个主日报 async job 并立即返回 `job_id`
|
||||
2. 能轮询到 `running -> success/failed`
|
||||
3. 成功后 `result.json` 含 `run_id` 与核心产物路径
|
||||
4. 对应 run 仍能通过既有 `get_run_status` / `get_run_report` / `get_delivery_payload` 正常读取
|
||||
5. README / handoff / TODO / plans 同步更新
|
||||
|
||||
## 建议实施顺序
|
||||
|
||||
1. 复制 article-summary job 骨架到 freshrss pipeline job
|
||||
2. 新增 runner 脚本
|
||||
3. server.py 暴露 3 个新工具
|
||||
4. 补 query/result 读法
|
||||
5. 更新 README / handoff
|
||||
6. 将 TODO 主任务切到“主日报 async job”
|
||||
@@ -0,0 +1,112 @@
|
||||
# reader 日报链路在 OpenClaw/Feishu 外层执行中被 SIGTERM 截断
|
||||
|
||||
## 背景
|
||||
2026-04-06 在 OpenClaw 中执行 reader 日报流程时,出现多次“前半段有产物、最终产物缺失”的现象。
|
||||
|
||||
典型表现:
|
||||
- 能生成 `raw/freshrss.raw.json`
|
||||
- 能部分生成 `extracted/item-xx.extracted.json`
|
||||
- 但经常拿不到:
|
||||
- `run-report.json`
|
||||
- `candidates/openclaw-delivery-payload.json`
|
||||
- `candidates/digest-brief.json`
|
||||
- 外层日志多次出现 `Exec failed (..., signal SIGTERM)`
|
||||
|
||||
## 已确认结论
|
||||
|
||||
### 1. RSS 抓取与排序正常
|
||||
已确认:
|
||||
- FreshRSS API 正常
|
||||
- 原始 raw 数据能拉到
|
||||
- 默认排序正常(默认相当于 `r=d`,新到旧)
|
||||
|
||||
因此问题不在:
|
||||
- RSS 接口
|
||||
- 鉴权
|
||||
- 排序规则
|
||||
|
||||
### 2. reader 前半段模块正常
|
||||
已确认:
|
||||
- 单篇 summary 能成功
|
||||
- 单篇 filter + candidate 构建能成功
|
||||
|
||||
因此问题不在:
|
||||
- summary 模块整体损坏
|
||||
- candidate 构建整体损坏
|
||||
|
||||
### 3. 真正问题在长链路执行方式
|
||||
更合理的判断是:
|
||||
|
||||
> 整条 reader 日报 pipeline 是长串行任务,在 OpenClaw / Feishu 当前这条外层执行链路里,容易被外层执行环境提前 `SIGTERM`。
|
||||
|
||||
也就是说:
|
||||
- 不是 reader 总是自己抛 Python 异常退出
|
||||
- 更多是脚本尚未跑完,外层执行会话先被终止
|
||||
|
||||
## 关键认知
|
||||
当天很多排障与补跑步骤,实际不是通过稳定常驻的 MCP 服务在跑,而是直接执行 reader 仓库里的 Python 脚本:
|
||||
|
||||
- `python scripts/run_freshrss_pipeline.py`
|
||||
- `python scripts/run_article_summaries.py`
|
||||
|
||||
因此更准确地说:
|
||||
|
||||
> 当前 reader 正式日报运行入口偏 CLI/脚本模式,而不是稳定 MCP 服务调用模式。
|
||||
|
||||
这也是为什么长任务更容易受外层 exec 生命周期影响。
|
||||
|
||||
## 为什么前几次没问题
|
||||
可能原因:
|
||||
1. 之前任务更短、内容更轻,刚好能在外层执行环境截断前跑完
|
||||
2. 之前不是链路天然稳,而是还没撞上边界条件
|
||||
3. 当前正式链路缺少稳健的断点恢复能力,因此一旦遇到较重任务,就暴露出问题
|
||||
|
||||
## 当天动作记录
|
||||
|
||||
### 做过的排查
|
||||
- 确认 FreshRSS 默认排序
|
||||
- 确认 raw 文件能生成
|
||||
- 确认 extracted 文件能部分生成
|
||||
- 单独验证单篇 summary 成功
|
||||
- 单独验证单篇 filter / candidate 成功
|
||||
|
||||
### 做过的临时修复
|
||||
当天曾尝试加入:
|
||||
- `--resume`
|
||||
- 中间产物复用
|
||||
- 断点恢复思路
|
||||
|
||||
目的是降低 SIGTERM 后的损失。
|
||||
|
||||
### 当前状态
|
||||
这些临时代码修改已全部回滚,reader 工作区已恢复干净。
|
||||
|
||||
## 当天结果
|
||||
虽然正式链路异常,但通过手工恢复推进,最终仍补齐了:
|
||||
- `candidates/openclaw-delivery-payload.json`
|
||||
- `candidates/digest-brief.json`
|
||||
- `run-report.json`
|
||||
|
||||
当天最终保留文章:
|
||||
- 1
|
||||
- 3
|
||||
|
||||
## 后续建议
|
||||
|
||||
### 短期
|
||||
- CLI 仅保留为 debug / fallback
|
||||
- 不再把正式生产日报流主要建立在长脚本入口上
|
||||
|
||||
### 中期
|
||||
- 将 reader 作为正式 MCP 服务
|
||||
- OpenClaw 正式通过 MCP tool 调用:
|
||||
- `run_freshrss_openclaw_pipeline`
|
||||
- `generate_article_summaries`
|
||||
|
||||
### 长期
|
||||
让 reader 的正式能力具备:
|
||||
- run_id
|
||||
- status 查询
|
||||
- resume
|
||||
- 中间产物复用
|
||||
- 分阶段补跑
|
||||
@@ -0,0 +1,257 @@
|
||||
# keyword-cleanup 产物精简方案 v1
|
||||
|
||||
## 1. 背景
|
||||
|
||||
当前 keyword-cleanup 治理链已经从“只有 review bundle”演进到:
|
||||
|
||||
- term stats / daily term index
|
||||
- review bundle
|
||||
- suggestions json
|
||||
- suggestions markdown
|
||||
- apply -> config / watchlist / change_log
|
||||
|
||||
这说明链路已经打通,但也带来一个新问题:
|
||||
|
||||
> 中间产物偏多,容易让治理系统本身比被治理对象更重。
|
||||
|
||||
本方案的目标不是回退功能,而是重新划分:
|
||||
|
||||
- 哪些产物是长期资产
|
||||
- 哪些产物只是决策输入
|
||||
- 哪些产物只是运行时工作文件
|
||||
|
||||
从而把 keyword-cleanup 收敛成一个更轻的治理辅助层,而不是继续长成一个复杂子系统。
|
||||
|
||||
---
|
||||
|
||||
## 2. 设计目标
|
||||
|
||||
本轮精简目标:
|
||||
|
||||
1. 保留真正有长期价值的事实层与状态层数据
|
||||
2. 保留唯一正式建议产物,用于 review / apply
|
||||
3. 将 review bundle 和 markdown 展示稿降级为临时产物
|
||||
4. 让主链路收敛到:
|
||||
|
||||
`stats -> suggestions json -> apply -> config`
|
||||
|
||||
而不是长期依赖:
|
||||
|
||||
`stats -> bundle -> suggestions json + md -> review -> apply`
|
||||
|
||||
---
|
||||
|
||||
## 3. 产物分层建议
|
||||
|
||||
### 3.1 长期保留:事实层
|
||||
|
||||
这些文件是系统长期事实基础,应继续长期保留:
|
||||
|
||||
- `data/term_index/daily/YYYY-MM-DD.json`
|
||||
- `data/term_index/term_stats.json`
|
||||
|
||||
原因:
|
||||
|
||||
- daily 文件代表每日聚合观察结果
|
||||
- global stats 是治理决策的核心事实来源
|
||||
- 二者共同构成 term governance 的历史依据
|
||||
|
||||
### 3.2 长期保留:状态层
|
||||
|
||||
这些文件代表治理系统当前状态,应继续长期保留:
|
||||
|
||||
- `configs/filter_context.personal.json`
|
||||
- `configs/term_watchlist.json`
|
||||
- `configs/term_aliases.json`
|
||||
- `configs/term_stopwords.json`
|
||||
- `configs/term_change_log.json`
|
||||
|
||||
原因:
|
||||
|
||||
- 它们是已确认生效的治理结果
|
||||
- 后续 reader 行为依赖这些配置
|
||||
- `change_log` 负责回溯治理动作
|
||||
|
||||
### 3.3 短期保留:正式建议层
|
||||
|
||||
建议保留:
|
||||
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`
|
||||
|
||||
定位:
|
||||
|
||||
- 这是 review / apply 之间的唯一正式建议产物
|
||||
- 机器可消费
|
||||
- 可作为某次治理决策的外部依据
|
||||
|
||||
建议策略:
|
||||
|
||||
- 默认仅保留最近少量几份
|
||||
- 或仅保留已经 apply 过的 suggestions JSON
|
||||
- 避免无限累积所有历史 suggestions 文件
|
||||
|
||||
### 3.4 降级为临时产物:review bundle
|
||||
|
||||
建议降级:
|
||||
|
||||
- `outputs/term_index/review/keyword-cleanup-bundle.json`
|
||||
|
||||
定位:
|
||||
|
||||
- review 输入打包文件
|
||||
- 只服务于 suggestions 生成过程
|
||||
- 不属于长期治理资产
|
||||
|
||||
建议策略:
|
||||
|
||||
- 默认只保留当前最新一份
|
||||
- 或迁移到更明确的 working/tmp 目录语义
|
||||
- 不按日期长期积累
|
||||
|
||||
### 3.5 降级为临时产物:Markdown 展示稿
|
||||
|
||||
建议降级:
|
||||
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`
|
||||
|
||||
定位:
|
||||
|
||||
- 纯人工审阅展示层
|
||||
- 不是唯一真相
|
||||
- 不参与 apply 逻辑
|
||||
|
||||
建议策略:
|
||||
|
||||
- 默认不长期持久化
|
||||
- 需要人工审阅时临时生成
|
||||
- 优先在聊天/界面中直接展示,而不是默认写成长期文件
|
||||
|
||||
---
|
||||
|
||||
## 4. 精简后的主链路
|
||||
|
||||
建议主链路口径收敛为:
|
||||
|
||||
1. 更新 daily term index
|
||||
2. 更新 global term stats
|
||||
3. 生成 suggestions JSON
|
||||
4. 人工确认
|
||||
5. apply 到 config
|
||||
6. 记录 change log
|
||||
|
||||
即:
|
||||
|
||||
`stats -> suggestions json -> apply -> config`
|
||||
|
||||
其中:
|
||||
|
||||
- bundle = 内部工作层
|
||||
- markdown = 展示层
|
||||
- suggestions JSON = 唯一正式建议输入
|
||||
|
||||
---
|
||||
|
||||
## 5. 为什么这样收敛
|
||||
|
||||
### 5.1 避免中间层过多
|
||||
|
||||
如果 bundle / md / suggestions 都被长期持久化,就容易出现:
|
||||
|
||||
- 多份文件语义重叠
|
||||
- 不知道谁是“准的”
|
||||
- 哪些只是试跑产物,哪些是正式治理决策不清晰
|
||||
|
||||
### 5.2 保持系统重心正确
|
||||
|
||||
keyword-cleanup 的最终目的不是维护一个漂亮的 review 文件集合,而是:
|
||||
|
||||
- 持续积累稳定的关键词事实数据
|
||||
- 让 interest/watch/alias/stopword 演化有据可依
|
||||
- 让 reader 的长期偏好配置从真实日报里长出来
|
||||
|
||||
### 5.3 降低治理系统自身复杂度
|
||||
|
||||
治理系统应该比主系统更轻,而不是更重。
|
||||
如果 review 产物越积越多,最终会反过来增加维护和理解成本。
|
||||
|
||||
---
|
||||
|
||||
## 6. 对现有实现的影响
|
||||
|
||||
本轮不要求删除已有能力,而是重新定义口径。
|
||||
|
||||
### 6.1 保留
|
||||
|
||||
- `build_review_bundle.py`
|
||||
- `generate_term_cleanup_suggestions.py`
|
||||
- `apply_term_suggestions.py`
|
||||
|
||||
### 6.2 调整口径
|
||||
|
||||
- `keyword-cleanup-bundle.json` 从“默认产物”降级为“临时工作文件”
|
||||
- `term-cleanup-suggestions-YYYY-MM-DD.md` 从“正式产物”降级为“临时展示稿”
|
||||
- `term-cleanup-suggestions-YYYY-MM-DD.json` 作为唯一正式建议产物保留
|
||||
|
||||
### 6.3 后续可选实现动作
|
||||
|
||||
- 覆盖式写入 bundle,而不是长期累积
|
||||
- Markdown 按需生成,而不是默认总是落盘
|
||||
- 增加清理策略,只保留最近 N 个 suggestions JSON
|
||||
|
||||
---
|
||||
|
||||
## 7. SOP 调整建议
|
||||
|
||||
### 7.1 review 阶段
|
||||
|
||||
默认步骤:
|
||||
|
||||
1. 生成或更新 term stats
|
||||
2. 生成最新 bundle(临时)
|
||||
3. 生成 suggestions JSON(正式)
|
||||
4. 如需要人工阅读,再临时生成 Markdown 或直接在聊天展示
|
||||
|
||||
### 7.2 apply 阶段
|
||||
|
||||
apply 后以以下内容作为最终真相:
|
||||
|
||||
- config 文件当前值
|
||||
- `term_change_log.json`
|
||||
- 如需要,保留对应 suggestions JSON 作为决策依据
|
||||
|
||||
### 7.3 清理策略
|
||||
|
||||
建议:
|
||||
|
||||
- bundle:默认仅保留最新
|
||||
- markdown:默认不归档
|
||||
- suggestions JSON:保留少量最近记录或已应用记录
|
||||
|
||||
---
|
||||
|
||||
## 8. 非目标
|
||||
|
||||
本轮不做:
|
||||
|
||||
- 删除现有脚本
|
||||
- 重写治理链路
|
||||
- 一次性重构所有 review 文档
|
||||
- 自动 apply 所有 suggestions
|
||||
- 引入更复杂的存储系统
|
||||
|
||||
本轮只做一件事:
|
||||
|
||||
> 把 keyword-cleanup 的产物语义分清,长期保留该留的,临时化该临时的。
|
||||
|
||||
---
|
||||
|
||||
## 9. 一句话结论
|
||||
|
||||
keyword-cleanup 应该收敛为:
|
||||
|
||||
- **事实层长期保留**:daily / term_stats
|
||||
- **状态层长期保留**:interest / watch / alias / stopword / change_log
|
||||
- **正式建议层轻量保留**:suggestions JSON
|
||||
- **中间输入层与展示层临时化**:bundle / markdown
|
||||
|
||||
最终目标是让 reader 的关键词治理成为一个轻量、可持续、可回溯的偏好演化机制,而不是一个不断膨胀的中间文件系统。
|
||||
@@ -0,0 +1,290 @@
|
||||
# interest/watch 候选引擎改进方案
|
||||
|
||||
> 从固定阈值到自适应排位 + 趋势因子的演进
|
||||
|
||||
## 1. 背景
|
||||
|
||||
### 1.1 当前实现
|
||||
|
||||
`build_review_bundle.py` 使用固定的绝对阈值将未覆盖词(uncovered terms)划分为两个候选池:
|
||||
|
||||
| 候选池 | 判断条件 | 依据 |
|
||||
|--------|---------|------|
|
||||
| `interest_review_candidates` | `total_count >= 3 AND days_seen >= 2` | `policy.interest_keyword_review` |
|
||||
| `watch_review_candidates` | `total_count <= 2 AND days_seen <= 2` | `policy.watch_term_review` |
|
||||
|
||||
`generate_term_cleanup_suggestions.py` 则直接从这两个候选池过滤、去重、排序后输出。
|
||||
|
||||
### 1.2 当前方案的问题
|
||||
|
||||
**问题一:固定阈值不随数据量自适应**
|
||||
|
||||
```
|
||||
场景 total_count=3 意味着什么
|
||||
─────────────────────────────────────────────
|
||||
7 天数据(~200 词) top 15%,有一定区分度 ✅
|
||||
41 天数据(1070 词) top 5%,区分度更高 ✅ 但阈值没变
|
||||
未来 200 天 仍然用 3 次,区分度稀释 ❌
|
||||
```
|
||||
|
||||
同一个绝对次数,在不同数据规模下的语义完全不同。手工调阈值不可持续。
|
||||
|
||||
**问题二:固定阈值忽略趋势信号**
|
||||
|
||||
- "Anthropic":total=13, recent=7 — 近期高活跃,上升趋势
|
||||
- "Channels":total=3, recent=0 — 早期出现但近期消失
|
||||
- 当前引擎认为这两个词"都过了 3 次阈值",同等对待。实际一个是强烈买入信号,一个是过气词。
|
||||
|
||||
**问题三:interest 和 watch 的分界线是硬的**
|
||||
|
||||
total=3 → interest,total=2 → watch。一个词从 2 次变成 3 次就自动"升级",没有过渡、没有缓冲。
|
||||
|
||||
### 1.3 讨论结论
|
||||
|
||||
与老大讨论后确认:
|
||||
|
||||
1. interest/watch 是**统计判断**,不需要大模型介入,纯算法可以解决
|
||||
2. 当前引擎缺的不是大模型,而是**算法本身没写完**——自适应维度(排位、趋势)还没实现
|
||||
3. alias 和 stopword 需要语义判断,与 interest/watch 分属不同阶段,不在本方案范围内
|
||||
4. 修改量小,可以在 1 小时内落地
|
||||
|
||||
---
|
||||
|
||||
## 2. 设计方案
|
||||
|
||||
### 2.1 核心思路
|
||||
|
||||
引入两个互补维度替代固定阈值:
|
||||
|
||||
```
|
||||
判定维度 含义 数据来源
|
||||
────────────────────────────────────────────────────────────
|
||||
percentile(百分位排名) 该词 total_count 在所有词 term_stats
|
||||
中的排位占比
|
||||
growth(增速因子) 近期集中度 = recent_count daily 近 N 天
|
||||
/ total_count
|
||||
```
|
||||
|
||||
两个维度配合:
|
||||
|
||||
- **percentile** 衡量"这个词在当前数据集里有多突出"——消除数据量变化的影响
|
||||
- **growth** 衡量"这个词是持续出现还是近期爆发"——识别趋势信号
|
||||
|
||||
### 2.2 候选池划分逻辑
|
||||
|
||||
```
|
||||
percentile
|
||||
│
|
||||
┌─────────────────────┐
|
||||
│ top 5% │
|
||||
│ → 建议 interest │ ← 高频稳定词
|
||||
├─────────────────────┤
|
||||
│ top 5%-20% │
|
||||
│ → 建议 watch │ ← 有信号但未达 threshold
|
||||
├─────────────────────┤
|
||||
│ bottom 80% │
|
||||
│ → 暂不处理 │ ← 噪声/低频
|
||||
└─────────────────────┘
|
||||
|
||||
额外规则:
|
||||
如果词在 top 20% 之外,但 growth > 0.5(近期集中度高)
|
||||
→ 主动提升到 watch / 主动推 confirm
|
||||
```
|
||||
|
||||
这样就不需要关心"total_count 是 3 还是 5",只看数据自己说话。
|
||||
|
||||
### 2.3 接口变化
|
||||
|
||||
**`configs/term_cleanup_policy.json`**:
|
||||
|
||||
```json
|
||||
{
|
||||
"schema_version": "v2",
|
||||
"interest_keyword_review": {
|
||||
"percentile_max": 0.05,
|
||||
"growth_promotion": 0.5
|
||||
},
|
||||
"watch_term_review": {
|
||||
"percentile_min": 0.05,
|
||||
"percentile_max": 0.20
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
`v1` 的 `min_total_count`/`min_days_seen` 等绝对阈值字段不再使用。
|
||||
|
||||
**`build_review_bundle.py` 输出的候选项**:
|
||||
|
||||
```json
|
||||
{
|
||||
"term": "Anthropic",
|
||||
"total_count": 13,
|
||||
"days_seen": 10,
|
||||
"percentile": 0.012,
|
||||
"growth": 0.54,
|
||||
"reason": "top 1.2% by frequency, 54% of occurrences in recent window — strong signal."
|
||||
}
|
||||
```
|
||||
|
||||
### 2.4 不需要改动的部分
|
||||
|
||||
- `generate_term_cleanup_suggestions.py` — 它只消费候选池,不用改
|
||||
- `apply_term_suggestions.py` — 消费 suggestions JSON,不用改
|
||||
- `keyword-cleanup-bundle.json` 结构 — 向后兼容,新增 percentile/growth 字段
|
||||
|
||||
---
|
||||
|
||||
## 3. 实施计划
|
||||
|
||||
### 3.1 改动范围
|
||||
|
||||
| 文件 | 改动量 | 内容 |
|
||||
|------|--------|------|
|
||||
| `skills/keyword-cleanup-review/scripts/build_review_bundle.py` | ~40 行 | 新增 `_compute_percentile()` 和 `_compute_growth()` 函数;修改候选池生成逻辑;候选项中增加 percentile/growth |
|
||||
| `configs/term_cleanup_policy.json` | ~10 行 | schema v2:percentile/growth 替代绝对阈值 |
|
||||
|
||||
### 3.2 实施步骤
|
||||
|
||||
1. **build_review_bundle.py**:在 `top_global_terms` 生成后,增加 percentile 计算函数和 growth 计算函数
|
||||
2. **build_review_bundle.py**:修改 `interest_review_candidates` 和 `watch_review_candidates` 的生成逻辑,从固定阈值改为 percentile + growth
|
||||
3. **build_review_bundle.py**:候选项增加 `percentile` 和 `growth` 字段,更新 `reason` 文案
|
||||
4. **term_cleanup_policy.json**:更新为 v2 schema
|
||||
5. **验证**:全量跑一次(`--days 365 --top 100`),对比新旧两份输出的差异
|
||||
|
||||
### 3.3 验证方法
|
||||
|
||||
```bash
|
||||
# 1. 用旧版生成 baseline
|
||||
cd /home/ubuntu/zhu/github/reader
|
||||
python3.11 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
|
||||
--days 365 --top 100 \
|
||||
--output /tmp/bundle-baseline.json
|
||||
|
||||
# 2. 改代码后用新版生成
|
||||
python3.11 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
|
||||
--days 365 --top 100 \
|
||||
--output /tmp/bundle-new.json
|
||||
|
||||
# 3. 对比 governance_hints
|
||||
python3 -c "
|
||||
import json
|
||||
a = json.load(open('/tmp/bundle-baseline.json'))
|
||||
b = json.load(open('/tmp/bundle-new.json'))
|
||||
for key in ['interest_review_candidates', 'watch_review_candidates']:
|
||||
old = set(i['term'] for i in a['governance_hints'][key])
|
||||
new = set(i['term'] for i in b['governance_hints'][key])
|
||||
print(f'{key}: 新增={new-old}, 减少={old-new}')
|
||||
"
|
||||
```
|
||||
|
||||
### 3.4 风险
|
||||
|
||||
| 风险 | 概率 | 应对 |
|
||||
|------|------|------|
|
||||
| 百分位阈值对特小数据集(如只有 1 天数据)不适用 | 低 | 不足 7 天时降级回绝对阈值 |
|
||||
| growth 因子对低频词的偏差(total=1, recent=1 → growth=1) | 低 | growth 只对 total>=3 的词计算 |
|
||||
| 排位突变导致推荐漂移 | 低 | percentil 天然平滑,新增几天数据不会剧烈改变已有词的排位 |
|
||||
|
||||
---
|
||||
|
||||
## 4. alias/stopword 设计方案
|
||||
|
||||
### 4.1 核心判断
|
||||
|
||||
alias 和 stopword 需要语义理解,与 interest/watch(纯统计)性质不同。
|
||||
|
||||
| 类型 | 需要什么 | 判断方式 |
|
||||
|------|---------|----------|
|
||||
| 大小写变体 | 表层 | 规则:casefold 去重 |
|
||||
| 单复数 | 表层 | 规则:去/加 s 后缀匹配 |
|
||||
| 分词变体(空格/连字符) | 表层 | 规则:去空格归一 |
|
||||
| 简写全称(MCP→Model Context Protocol) | **语义** | LLM |
|
||||
| 中英文(上下文工程→Context Engineering) | **语义** | LLM |
|
||||
| 同义不同名(Rush→猿辅导 Rush 平台) | **语义** | LLM |
|
||||
| stopword(大模型、AI 太泛) | **语义** | LLM |
|
||||
|
||||
### 4.2 分层方案
|
||||
|
||||
```
|
||||
输入:高频未覆盖词 + 已有 interest 词表
|
||||
│
|
||||
├── 规则层(零成本)── 大小写归一、单复数、去空格/连字符
|
||||
│ 输出候选 alias 对
|
||||
│
|
||||
└── LLM 层(每次 ~500 token)── 把候选词表整批给 LLM
|
||||
做语义聚类
|
||||
输出 alias 组 + stopword 标记
|
||||
```
|
||||
|
||||
### 4.3 规则层设计
|
||||
|
||||
在 `generate_term_cleanup_suggestions.py` 中新增 `_prepare_alias_suggestions()` 函数:
|
||||
|
||||
```python
|
||||
def _prepare_alias_suggestions(top_terms, interest_keywords):
|
||||
"""
|
||||
基于表层规则生成 alias 建议。
|
||||
规则1:casefold 匹配——同一个 casefold 下有多个原文变体
|
||||
规则2:单复数——去掉/加上末尾 s 后匹配
|
||||
规则3:分词变体——去空格/连字符后匹配
|
||||
"""
|
||||
```
|
||||
|
||||
优势:零成本、可复现、可审计。直接写入 suggestions JSON,随 generate 一起输出。
|
||||
|
||||
### 4.4 LLM 层设计
|
||||
|
||||
单独脚本,非 generate 主链路的一部分。
|
||||
|
||||
```bash
|
||||
python scripts/generate_term_cleanup_semantic_suggestions.py \
|
||||
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
|
||||
--output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json
|
||||
```
|
||||
|
||||
LLM prompt 设计:
|
||||
|
||||
```
|
||||
你是一个关键词治理助手。以下是一个用户的 interest 关键词列表和一批未覆盖的高频词。
|
||||
请做三件事:
|
||||
|
||||
1. ALIAS:判断哪些未覆盖词是已有 interest 关键词的别名/变体
|
||||
2. STOPWORD:标记哪些词太宽泛/通用,建议排除
|
||||
3. PROMOTE:标记哪些新词与用户关注方向一致,建议加入 interest
|
||||
|
||||
用户关注方向:AI Agent 工程化、后端工程、开源工具、大模型落地
|
||||
```
|
||||
|
||||
LLM 层输出格式:
|
||||
|
||||
```json
|
||||
{
|
||||
"alias_suggestions": [
|
||||
{"from": "Context Engineering", "to": "上下文工程", "reason": "中英文对应同一概念"}
|
||||
],
|
||||
"stopword_suggestions": [
|
||||
{"term": "大模型", "reason": "过于宽泛,高频率但低区分度"}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
### 4.5 预期效果
|
||||
|
||||
| 覆盖类型 | 规则层 | LLM 层 |
|
||||
|---------|--------|--------|
|
||||
| 大小写变体 | ✅ | — |
|
||||
| 单复数 | ✅ | — |
|
||||
| 分词变体 | ✅ | — |
|
||||
| 简写全称 | — | ✅ |
|
||||
| 中英文映射 | — | ✅ |
|
||||
| 同义不同名 | — | ✅ |
|
||||
| stopword 判断 | — | ✅ |
|
||||
|
||||
---
|
||||
|
||||
## 5. 讨论记录
|
||||
|
||||
- 2026-05-14:与老大确认 interest/watch 不需要 LLM,纯算法可解决
|
||||
- 2026-05-14:确认百分位排名 + 增速因子方案,修改量小,优先落地
|
||||
- 2026-05-14:确认本方案不改 `generate_term_cleanup_suggestions.py` 和 `apply_term_suggestions.py`
|
||||
- 2026-05-14:确认 alias/stopword 采用规则层 + LLM 层分层方案,规则层零成本优先
|
||||
@@ -0,0 +1,534 @@
|
||||
# keyword-cleanup-review 建议产物补齐设计
|
||||
|
||||
## 1. 背景与目标
|
||||
|
||||
当前仓库已经具备 keyword cleanup review 的大部分基础设施:
|
||||
|
||||
- 已有 review bundle 构建脚本 `skills/keyword-cleanup-review/scripts/build_review_bundle.py`
|
||||
- 已有 term stats 与 daily term index 数据源
|
||||
- 已有治理输入:`configs/term_cleanup_policy.json`、`configs/term_watchlist.json`、`configs/term_change_log.json`
|
||||
- 已有建议落地脚本 `scripts/apply_term_suggestions.py`
|
||||
|
||||
当前缺口是:
|
||||
|
||||
> 缺少一层“把 `keyword-cleanup-bundle.json` 转成正式建议产物”的实现层。
|
||||
|
||||
也就是说,仓库现在能生成 review bundle,也能消费 suggestions JSON,但中间缺少稳定、可复用、可落盘的 suggestions 生成器。
|
||||
|
||||
本轮目标是补齐最小闭环,让仓库能够从 review bundle 稳定生成两份正式建议产物:
|
||||
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`
|
||||
|
||||
并保证 JSON 与 `skills/keyword-cleanup-review/references/suggestion-schema.md` 对齐,且能直接衔接 `scripts/apply_term_suggestions.py`。
|
||||
|
||||
补充口径:本方案中的 JSON 是 review / apply 之间的唯一正式建议产物;bundle 与 Markdown 主要作为运行时工作文件和临时展示层,不建议与 facts/configs 一样长期沉淀。
|
||||
|
||||
---
|
||||
|
||||
## 2. 当前现状
|
||||
|
||||
### 2.1 已有输入层
|
||||
|
||||
`build_review_bundle.py` 已经把以下输入聚合成单个 bundle:
|
||||
|
||||
- `data/term_index/term_stats.json`
|
||||
- `data/term_index/daily/*.json`
|
||||
- `configs/term_aliases.json`
|
||||
- `configs/term_stopwords.json`
|
||||
- `configs/filter_context.personal.json`
|
||||
- `configs/term_cleanup_policy.json`
|
||||
- `configs/term_watchlist.json`
|
||||
- `configs/term_change_log.json`
|
||||
|
||||
bundle 中已经包含:
|
||||
|
||||
- 当前配置快照
|
||||
- 最近 N 天热点词
|
||||
- uncovered terms
|
||||
- interest review candidates
|
||||
- watch review candidates
|
||||
|
||||
这些信息已经足够支撑“保守的、可审查的” suggestions 生成。
|
||||
|
||||
### 2.2 已有输出消费层
|
||||
|
||||
`scripts/apply_term_suggestions.py` 已经能消费 suggestions JSON,并将接受的建议写回:
|
||||
|
||||
- `configs/term_aliases.json`
|
||||
- `configs/term_stopwords.json`
|
||||
- `configs/filter_context.personal.json`
|
||||
- `configs/term_watchlist.json`
|
||||
- `configs/term_change_log.json`
|
||||
|
||||
这说明落地层已存在,缺的是中间的正式建议产物生成层。
|
||||
|
||||
---
|
||||
|
||||
## 3. 当前缺口
|
||||
|
||||
当前流程停在:
|
||||
|
||||
`build_review_bundle.py` -> `keyword-cleanup-bundle.json`
|
||||
|
||||
但缺少:
|
||||
|
||||
`keyword-cleanup-bundle.json` -> `term-cleanup-suggestions-YYYY-MM-DD.json/.md`
|
||||
|
||||
因此出现几个问题:
|
||||
|
||||
- README / 设计文档里已经引用 suggestions 产物,但仓库内没有稳定生成脚本
|
||||
- 人工审阅与后续 apply 之间没有统一的正式交付格式
|
||||
- 同一份 bundle 无法稳定、幂等地重放为同名 suggestions 产物
|
||||
- alias / stopword / interest / watch 四类建议缺少统一出入口
|
||||
|
||||
---
|
||||
|
||||
## 4. 推荐最小闭环架构
|
||||
|
||||
推荐新增一层独立脚本:
|
||||
|
||||
- `scripts/generate_term_cleanup_suggestions.py`
|
||||
|
||||
职责:
|
||||
|
||||
- 输入:`outputs/term_index/review/keyword-cleanup-bundle.json`
|
||||
- 输出:
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`
|
||||
|
||||
推荐最小数据流:
|
||||
|
||||
1. `build_review_bundle.py` 生成 bundle
|
||||
2. `generate_term_cleanup_suggestions.py` 读取 bundle
|
||||
3. 脚本基于治理 hints 生成 suggestions JSON
|
||||
4. 同时渲染人类可审阅的 Markdown
|
||||
5. 审阅后可用 `apply_term_suggestions.py` 选择性写回配置
|
||||
|
||||
本轮不引入默认 LLM 路径:
|
||||
|
||||
- 默认实现采用确定性规则生成
|
||||
- 如果未来需要 LLM 参与,应作为显式可选增强,而不是默认路径
|
||||
|
||||
---
|
||||
|
||||
## 5. 产物设计
|
||||
|
||||
### 5.1 JSON 产物
|
||||
|
||||
JSON 必须与 `suggestion-schema.md` 对齐,至少包含:
|
||||
|
||||
```json
|
||||
{
|
||||
"date": "2026-04-08",
|
||||
"based_on_days": 7,
|
||||
"alias_suggestions": [],
|
||||
"stopword_suggestions": [],
|
||||
"interest_keyword_suggestions": [
|
||||
{
|
||||
"term": "Claude Code",
|
||||
"reason": "Meets the configured interest-keyword review threshold and is not yet covered."
|
||||
}
|
||||
],
|
||||
"watch_terms": [
|
||||
{
|
||||
"term": "A2A",
|
||||
"reason": "Falls into the configured watch-term review range and should be observed first."
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
在不破坏兼容性的前提下,可以补充少量元数据字段,建议仅限:
|
||||
|
||||
- `source_bundle`
|
||||
- `policy_schema_version`
|
||||
- `summary`
|
||||
|
||||
建议项字段口径:
|
||||
|
||||
- `alias_suggestions[]`
|
||||
- `from`
|
||||
- `to`
|
||||
- `reason`
|
||||
- `stopword_suggestions[]`
|
||||
- `term`
|
||||
- `reason`
|
||||
- `interest_keyword_suggestions[]`
|
||||
- `term`
|
||||
- `reason`
|
||||
- 可选:`total_count`、`days_seen`、`recent_count`
|
||||
- `watch_terms[]`
|
||||
- `term`
|
||||
- `reason`
|
||||
- 可选:`total_count`、`days_seen`、`recent_count`
|
||||
|
||||
兼容性要求:
|
||||
|
||||
- `apply_term_suggestions.py` 只依赖分类 bucket 与关键字段名
|
||||
- 因此额外证据字段只能追加,不能替换现有字段名
|
||||
|
||||
### 5.2 Markdown 产物
|
||||
|
||||
Markdown 推荐结构:
|
||||
|
||||
1. 标题与日期
|
||||
2. 输入 bundle 与策略摘要
|
||||
3. 当前现状摘要
|
||||
- top/global 观察
|
||||
- uncovered terms 概览
|
||||
- 当前 watchlist / interest 覆盖情况
|
||||
4. 建议摘要
|
||||
- interest 建议数量
|
||||
- watch 建议数量
|
||||
- alias 建议数量
|
||||
- stopword 建议数量
|
||||
5. `interest_keyword_suggestions`
|
||||
6. `watch_terms`
|
||||
7. `alias_suggestions`
|
||||
8. `stopword_suggestions`
|
||||
9. 应用方式
|
||||
- 指向生成的 JSON
|
||||
- 给出 `apply_term_suggestions.py` 的调用示例
|
||||
|
||||
这样可以保证:
|
||||
|
||||
- 人可以直接审阅
|
||||
- 机器可以直接消费同名 JSON
|
||||
- Markdown 与 JSON 始终一一对应
|
||||
|
||||
---
|
||||
|
||||
## 6. 建议生成策略
|
||||
|
||||
### 6.1 本轮主链路:interest / watch
|
||||
|
||||
本轮先实现最小可用主链路:
|
||||
|
||||
- `interest_keyword_suggestions`
|
||||
- `watch_terms`
|
||||
|
||||
直接复用 bundle 中已有的:
|
||||
|
||||
- `governance_hints.interest_review_candidates`
|
||||
- `governance_hints.watch_review_candidates`
|
||||
|
||||
原因:
|
||||
|
||||
- 这些候选已经与 policy 对齐
|
||||
- 这些候选已经排除了大部分已覆盖项
|
||||
- 能直接与现有 apply 脚本形成闭环
|
||||
|
||||
### 6.2 alias / stopword 保守处理
|
||||
|
||||
本轮边界明确如下:
|
||||
|
||||
- `alias_suggestions` 先保持保守,默认可为空
|
||||
- `stopword_suggestions` 先保持保守,默认可为空
|
||||
- 后续如果补充更强证据或人工审查规则,再逐步增强
|
||||
|
||||
这样可以避免在证据不足时误伤配置。
|
||||
|
||||
### 6.3 Phase 2 新方向:alias review 交给 LLM 整理
|
||||
|
||||
对于 alias,不再优先走程序规则匹配。
|
||||
Phase 2 建议改为:
|
||||
|
||||
- 程序继续负责准备 review 输入(term stats / daily / current aliases / stopwords / interest / watchlist)
|
||||
- LLM 负责整理 alias 候选
|
||||
- 默认先输出人工审阅汇报,而不是直接 apply
|
||||
- 人工确认后,再决定是否写入 `term_aliases.json`
|
||||
|
||||
这样做的原因:
|
||||
|
||||
- alias 更偏语义整理,而不是简单趋势筛选
|
||||
- 与 watch / interest 相比,alias 一旦错误归并,代价更高
|
||||
- 对当前低频治理场景来说,LLM + 人工确认更轻,也比在程序里持续堆复杂规则更合适
|
||||
|
||||
---
|
||||
|
||||
## 7. 错误处理
|
||||
|
||||
脚本应做显式校验,并在失败时给出明确错误:
|
||||
|
||||
### 7.1 输入错误
|
||||
|
||||
- bundle 文件不存在 -> 直接失败
|
||||
- bundle 不是 JSON object -> 直接失败
|
||||
- 缺少关键字段(如 `days`、`governance_hints`)-> 直接失败
|
||||
- 候选 bucket 结构错误 -> 直接失败
|
||||
|
||||
### 7.2 输出错误
|
||||
|
||||
- 输出目录不存在时自动创建
|
||||
- JSON / Markdown 写入失败时直接退出非 0
|
||||
|
||||
### 7.3 数据去重与冲突
|
||||
|
||||
- 同一 term 不能同时出现在 interest 与 watch 中
|
||||
- 优先级:`interest_keyword_suggestions` > `watch_terms`
|
||||
- 已在 bundle 当前配置中覆盖的 term 不重复输出
|
||||
|
||||
---
|
||||
|
||||
## 8. 幂等性
|
||||
|
||||
本轮要求具备基础幂等性:
|
||||
|
||||
- 同一份 bundle 多次运行,默认生成同名产物
|
||||
- 同一份 bundle 多次运行,JSON 内容顺序稳定
|
||||
- Markdown 内容顺序稳定
|
||||
|
||||
建议做法:
|
||||
|
||||
- 优先使用 bundle 的 `generated_at` 日期作为 suggestions 文件日期
|
||||
- term 排序按证据强度与 term 名稳定排序
|
||||
- 不在默认输出中写入“每次运行变化”的当前时间戳
|
||||
|
||||
这样可以让生成器作为可重放步骤存在于 review 流程中。
|
||||
|
||||
---
|
||||
|
||||
## 9. 与 apply_term_suggestions.py 的衔接
|
||||
|
||||
正式链路应变成:
|
||||
|
||||
1. `build_review_bundle.py`
|
||||
2. `generate_term_cleanup_suggestions.py`
|
||||
3. 人工审阅 Markdown
|
||||
4. `apply_term_suggestions.py --suggestions ...`
|
||||
|
||||
衔接要求:
|
||||
|
||||
- JSON bucket 名必须与 `apply_term_suggestions.py` 读取逻辑一致
|
||||
- `date` 与 `based_on_days` 字段保留,用于 change log 回写
|
||||
- 建议项中的 `reason` 直接沿用到 apply 后的 change log
|
||||
|
||||
这保证建议生成层不会成为孤立产物,而是正式进入 repo 治理闭环。
|
||||
|
||||
---
|
||||
|
||||
## 10. Python 3.11 依赖处理
|
||||
|
||||
### 10.1 当前现状
|
||||
|
||||
`build_review_bundle.py` 当前使用 `from datetime import UTC`,这要求 Python 3.11。
|
||||
|
||||
仓库整体 `pyproject.toml` 当前也声明 `requires-python = ">=3.11"`,因此短期内使用 `/usr/bin/python3.11` 运行是符合仓库现状的。
|
||||
|
||||
### 10.2 短期建议
|
||||
|
||||
短期先在文档与验证命令中明确:
|
||||
|
||||
- bundle 构建使用 `/usr/bin/python3.11`
|
||||
- suggestions 生成脚本也按仓库当前 3.11 基线运行
|
||||
|
||||
### 10.3 中期建议
|
||||
|
||||
如果后续希望把 keyword cleanup 工具链下探到 Python 3.10,可做兼容改造:
|
||||
|
||||
- 把 `datetime.UTC` 替换为 `datetime.timezone.utc`
|
||||
- 重新检查相关脚本是否还有其他 3.11-only 语法或库依赖
|
||||
|
||||
本轮不做超范围兼容重构,只在文档中把此约束说清楚。
|
||||
|
||||
---
|
||||
|
||||
## 11. 边界与非目标
|
||||
|
||||
本轮明确边界:
|
||||
|
||||
- 先实现 interest/watch 主链路
|
||||
- alias/stopword 先保持保守或留待后续增强
|
||||
- 不做超范围重构
|
||||
- 不把 LLM 作为默认生成路径
|
||||
- 不直接改 `filter_rules.json`
|
||||
- 不自动 apply 建议到配置
|
||||
|
||||
因此,本轮交付定义为:
|
||||
|
||||
- 补齐 bundle -> suggestions 的正式实现层
|
||||
- 让 review 流程可运行、可落盘、可审阅、可应用
|
||||
|
||||
而不是一次性做完所有高级治理逻辑。
|
||||
|
||||
---
|
||||
|
||||
## 12. 实施建议
|
||||
|
||||
建议按以下小步落地:
|
||||
|
||||
### Phase 1(已完成)
|
||||
|
||||
1. 新增 `scripts/generate_term_cleanup_suggestions.py`
|
||||
2. 读取 bundle 并做结构校验
|
||||
3. 生成稳定排序的 interest/watch suggestions JSON
|
||||
4. Markdown 改成按需生成
|
||||
5. README / skill 文档补一条生成命令
|
||||
6. 用 `/usr/bin/python3.11` 完整跑通 bundle -> suggestions
|
||||
|
||||
完成后,keyword cleanup review 的最小正式链路变为:
|
||||
|
||||
`review bundle` -> `suggestions json` -> `apply accepted suggestions`
|
||||
|
||||
其中 Markdown 只是按需生成的展示层。
|
||||
|
||||
### Phase 2(下一步)
|
||||
|
||||
1. 保持程序继续准备 review 输入
|
||||
2. 引入 LLM 做 alias 候选整理
|
||||
3. 默认先生成 alias review 汇报,而不是直接 apply
|
||||
4. 由人工确认后再决定是否写入 `term_aliases.json`
|
||||
|
||||
这样 alias review 会成为一个低频治理动作,而不是主链路里的自动归一步骤。
|
||||
|
||||
### Phase 3(下一步)
|
||||
|
||||
1. 保持程序继续准备 review 输入
|
||||
2. 引入 LLM 做 stopword 候选整理
|
||||
3. 默认先生成 stopword review 汇报,而不是直接 apply
|
||||
4. 由人工确认后再决定是否写入 `term_stopwords.json`
|
||||
|
||||
这样 stopword review 会成为一个低频减噪动作,而不是主链路里的自动过滤步骤。
|
||||
|
||||
#### Phase 3 输入建议
|
||||
|
||||
建议给 LLM 的输入包括:
|
||||
|
||||
- `data/term_index/term_stats.json` 中的高频词与 recent evidence
|
||||
- 最近 N 天 `data/term_index/daily/*.json` 的热点词上下文
|
||||
- 当前 `configs/term_stopwords.json`
|
||||
- 当前 `configs/term_aliases.json`
|
||||
- 当前 `configs/filter_context.personal.json` 中的 `interest_keywords`
|
||||
- 当前 `configs/term_watchlist.json`
|
||||
|
||||
程序层只负责把这些输入整理成紧凑 review context,不负责直接做 stopword 决策。
|
||||
|
||||
#### Phase 3 输出建议
|
||||
|
||||
建议 LLM 默认输出一份人工审阅汇报,而不是直接写配置:
|
||||
|
||||
- `建议加入 stopword`
|
||||
- 词
|
||||
- 简短理由
|
||||
- 证据(如 total_count / days_seen / recent_count)
|
||||
- `暂不建议加入 stopword`
|
||||
- 词
|
||||
- 为什么虽然偏泛,但当前还不能杀
|
||||
- `需要人工判断`
|
||||
- 词
|
||||
- 风险点:可能是噪声,也可能仍保留有价值信号
|
||||
|
||||
如需结构化输出,可额外补一份 `stopword_review_candidates.json`,但该文件只作为 review 输入,不直接作为自动 apply 指令。
|
||||
|
||||
#### Phase 3 审阅原则
|
||||
|
||||
- 宁可少删,不乱杀
|
||||
- 优先处理过泛、低辨识度、持续污染统计的词
|
||||
- 对可能仍承载有效技术语义的词保持保守
|
||||
- 默认先汇报,确认后再执行
|
||||
|
||||
#### Phase 3 汇报模板建议
|
||||
|
||||
建议 stopword review 默认按以下结构汇报给用户:
|
||||
|
||||
1. `建议加入 stopword`
|
||||
- 词
|
||||
- 理由:为什么这个词对治理帮助低、噪声高
|
||||
- 证据:`total_count` / `days_seen` / `recent_count` 或最近出现上下文
|
||||
2. `暂不建议加入 stopword`
|
||||
- 词
|
||||
- 理由:为什么当前不建议删掉
|
||||
3. `需要人工判断`
|
||||
- 词
|
||||
- 风险点:泛词与有效主题词之间边界不清等
|
||||
|
||||
推荐汇报风格:
|
||||
|
||||
- 简短、保守、可审阅
|
||||
- 先给判断,再给证据
|
||||
- 不输出机器式原始 dump
|
||||
- 不默认承诺“已应用”,只汇报“建议”
|
||||
|
||||
#### Phase 2 输入建议
|
||||
|
||||
建议给 LLM 的输入包括:
|
||||
|
||||
- `data/term_index/term_stats.json` 中的高频词与 recent evidence
|
||||
- 最近 N 天 `data/term_index/daily/*.json` 的热点词上下文
|
||||
- 当前 `configs/term_aliases.json`
|
||||
- 当前 `configs/term_stopwords.json`
|
||||
- 当前 `configs/filter_context.personal.json` 中的 `interest_keywords`
|
||||
- 当前 `configs/term_watchlist.json`
|
||||
|
||||
程序层只负责把这些输入整理成紧凑 review context,不负责直接做 alias 决策。
|
||||
|
||||
#### Phase 2 输出建议
|
||||
|
||||
建议 LLM 默认输出一份人工审阅汇报,而不是直接写配置:
|
||||
|
||||
- `建议合并`
|
||||
- `from -> to`
|
||||
- 简短理由
|
||||
- 证据(如 total_count / days_seen / recent_count)
|
||||
- `暂不建议合并`
|
||||
- 为什么不建议并掉
|
||||
- `需要人工判断`
|
||||
- 语义相近但风险较高的项
|
||||
|
||||
如需结构化输出,可额外补一份 `alias_review_candidates.json`,但该文件只作为 review 输入,不直接作为自动 apply 指令。
|
||||
|
||||
#### Phase 2 审阅原则
|
||||
|
||||
- 宁可少提,不乱提
|
||||
- 优先整理明显同义 / 同概念 / 词形差异
|
||||
- 不把公司名、产品名、泛概念词强行混并
|
||||
- 默认先汇报,确认后再执行
|
||||
|
||||
#### Phase 2 汇报模板建议
|
||||
|
||||
建议 alias review 默认按以下结构汇报给用户:
|
||||
|
||||
1. `建议合并`
|
||||
- `from -> to`
|
||||
- 理由:为什么判断为同一概念或更合适的标准词
|
||||
- 证据:`total_count` / `days_seen` / `recent_count` 或最近出现上下文
|
||||
2. `暂不建议合并`
|
||||
- 候选对
|
||||
- 理由:为什么虽然相近,但当前不建议并
|
||||
3. `需要人工判断`
|
||||
- 候选对
|
||||
- 风险点:歧义、范围差异、产品名/公司名混淆等
|
||||
|
||||
推荐汇报风格:
|
||||
|
||||
- 简短、保守、可审阅
|
||||
- 先给判断,再给证据
|
||||
- 不输出机器式原始 dump
|
||||
- 不默认承诺“已应用”,只汇报“建议”
|
||||
|
||||
#### Phase 2 示例输出
|
||||
|
||||
建议合并:
|
||||
|
||||
- `Claude code -> Claude Code`
|
||||
- 理由:明显属于同一产品名,仅是大小写写法不一致。
|
||||
- 证据:`Claude Code` 在最近多日持续出现,而小写写法只是在少量上下文中作为变体出现。
|
||||
|
||||
- `Sub-Agent -> SubAgent`
|
||||
- 理由:更像词形差异,不构成新的独立概念。
|
||||
- 证据:两者都围绕同一 agent 架构语境出现,且没有稳定区分语义。
|
||||
|
||||
暂不建议合并:
|
||||
|
||||
- `Skills ↔ Agent Skills`
|
||||
- 理由:前者过泛,后者更具体,当前强行归并会损失粒度。
|
||||
|
||||
- `Anthropic ↔ Claude`
|
||||
- 理由:公司名与产品名并不等价,不应直接视为一个关键词。
|
||||
|
||||
需要人工判断:
|
||||
|
||||
- `AI助手 ↔ AI Agent`
|
||||
- 风险点:语义可能接近,但中文表述范围更宽,是否并入需要结合你的使用语境判断。
|
||||
|
||||
@@ -0,0 +1,766 @@
|
||||
# Reader 正式 MCP 服务架构设计
|
||||
|
||||
## 1. 背景
|
||||
|
||||
当前 reader 仓库已经具备 MCP 服务入口(`src/summary_mcp/server.py`),并暴露了:
|
||||
|
||||
- `extract_url_content`
|
||||
- `extract_item_content`
|
||||
- `filter_summary_result`
|
||||
- `run_freshrss_openclaw_pipeline`
|
||||
- `generate_article_summaries`
|
||||
|
||||
但从实际生产使用方式看,reader 的正式日报链路仍然**偏向 CLI/脚本模式**,尤其是:
|
||||
|
||||
- `python scripts/run_freshrss_pipeline.py`
|
||||
- `python scripts/run_article_summaries.py`
|
||||
|
||||
这导致在 OpenClaw / Feishu 外层执行环境下,长链路任务容易因为 exec 生命周期、超时或外层中断而失败。
|
||||
|
||||
2026-04-06 的事故已经说明:
|
||||
|
||||
- `raw` 能产出
|
||||
- `extracted` 能部分产出
|
||||
- 但 `run-report.json` / `openclaw-delivery-payload.json` / `digest-brief.json` 经常缺失
|
||||
- 外层日志出现 `signal SIGTERM`
|
||||
|
||||
因此,reader 需要从“可被 exec 调起的仓库”升级为“**正式的 MCP 工作流服务**”。
|
||||
|
||||
---
|
||||
|
||||
## 2. 设计目标
|
||||
|
||||
本次架构设计的目标不是重做 reader 的业务逻辑,而是把现有能力收口为稳定的服务接口。
|
||||
|
||||
目标如下:
|
||||
|
||||
1. **MCP 成为正式入口**
|
||||
- OpenClaw 与上层编排默认通过 MCP tool 调用 reader
|
||||
- CLI 降级为 debug / fallback 入口
|
||||
|
||||
2. **引入稳定的运行态抽象**
|
||||
- 统一 `run_id`
|
||||
- 统一 `status`
|
||||
- 统一 `stage`
|
||||
- 统一 `artifacts`
|
||||
|
||||
3. **支持长链路可观测与恢复**
|
||||
- 可查询当前运行状态
|
||||
- 可查询失败阶段
|
||||
- 可基于已有中间产物 resume / rerun
|
||||
|
||||
4. **让 OpenClaw 消费结构化结果,而不是硬编码目录细节**
|
||||
- OpenClaw 不再依赖 reader 的脚本 stdout 作为唯一信号
|
||||
- OpenClaw 尽量不直接拼接 reader 的输出目录路径
|
||||
|
||||
5. **保持最小重构成本**
|
||||
- 先用文件系统持久化 run state
|
||||
- 先不引入复杂任务队列 / 数据库 / 多 worker 平台
|
||||
- 先让单机、单实例、顺序执行场景稳定起来
|
||||
|
||||
---
|
||||
|
||||
## 3. 核心结论
|
||||
|
||||
一句话总结:
|
||||
|
||||
> Reader 应该被设计为“有状态的工作流 MCP 服务”,而不是“包了一层 MCP 壳的长 CLI 命令”。
|
||||
|
||||
换句话说:
|
||||
|
||||
- **错误方向**:`reader_run(command="python scripts/run_xxx.py ...")`
|
||||
- **正确方向**:`start_run` + `get_run_status` + `get_artifact` + `resume_run`
|
||||
|
||||
协议层只是入口,真正的关键是把 reader 内部抽象成:
|
||||
|
||||
- run
|
||||
- stage
|
||||
- status
|
||||
- artifact
|
||||
- recovery
|
||||
|
||||
---
|
||||
|
||||
## 4. 服务定位
|
||||
|
||||
### 4.1 Reader 负责什么
|
||||
|
||||
reader 作为 MCP 服务,负责上游阅读工作流本身:
|
||||
|
||||
- FreshRSS 拉取
|
||||
- 内容提取
|
||||
- LLM 摘要
|
||||
- 规则过滤
|
||||
- candidate / delivery payload 生成
|
||||
- 单篇总结生成
|
||||
- 中间产物保存
|
||||
- 运行状态记录
|
||||
- 恢复与补跑
|
||||
|
||||
### 4.2 Reader 不负责什么
|
||||
|
||||
以下继续由 OpenClaw / skill 层负责:
|
||||
|
||||
- Hugo 发布
|
||||
- 对话汇报
|
||||
- 用户确认精选
|
||||
- IMA 上传编排
|
||||
- 最终对人类的日常交互
|
||||
|
||||
### 4.3 边界结论
|
||||
|
||||
- **reader 是上游引擎 / workflow service**
|
||||
- **OpenClaw 是下游编排层 / orchestration layer**
|
||||
|
||||
这个边界与当前 `docs/openclaw/openclaw-handoff.md` 的原则保持一致,但会进一步强化“reader 通过正式 MCP 接口暴露工作流状态”,而不是只暴露“跑完后的结果”。
|
||||
|
||||
---
|
||||
|
||||
## 5. 当前问题分析
|
||||
|
||||
### 5.1 现状问题
|
||||
|
||||
当前 reader 的 MCP 工具里虽然已经有 `run_freshrss_openclaw_pipeline`,但其调用语义仍然更像:
|
||||
|
||||
- 一次性同步执行整个长链路
|
||||
- 直接返回最终结果
|
||||
- 对外隐藏中间状态
|
||||
|
||||
这会带来几个问题:
|
||||
|
||||
1. 长任务运行时,外层必须一直等
|
||||
2. 一旦外层中断,状态观测困难
|
||||
3. 难以精确判断失败发生在哪个阶段
|
||||
4. resume / rerun 只能靠脚本层补丁式处理
|
||||
5. OpenClaw 不容易构建“先触发,后查询,再消费”的稳定编排流
|
||||
|
||||
### 5.2 根本问题
|
||||
|
||||
根本问题不是“有没有 MCP server”,而是:
|
||||
|
||||
> **reader 还没有被真正建模成一个有状态的 workflow service。**
|
||||
|
||||
当前更接近:
|
||||
|
||||
- MCP 暴露了几个函数
|
||||
- 但长链路执行模型仍是 CLI thinking
|
||||
|
||||
因此改造重点不应放在“多加几个 tool 名字”,而应该放在:
|
||||
|
||||
- 状态机
|
||||
- 运行记录
|
||||
- 恢复机制
|
||||
- 标准产物注册
|
||||
|
||||
---
|
||||
|
||||
## 6. 目标架构
|
||||
|
||||
建议的正式架构如下:
|
||||
|
||||
```text
|
||||
OpenClaw / Skill Layer
|
||||
-> MCP Client Calls
|
||||
-> Reader MCP Server
|
||||
-> Workflow Service Layer
|
||||
-> Workflow Runtime / Stage Engine
|
||||
-> Existing Reader Core Modules
|
||||
- FreshRSS pull
|
||||
- extraction
|
||||
- summary
|
||||
- filter
|
||||
- candidate builder
|
||||
- article summary
|
||||
-> Run Store (filesystem-backed)
|
||||
-> Artifact Store (existing outputs directory)
|
||||
```
|
||||
|
||||
### 6.1 分层说明
|
||||
|
||||
#### A. MCP Server Layer
|
||||
职责:
|
||||
- tool 注册
|
||||
- 输入校验
|
||||
- 输出包装
|
||||
- 对外暴露稳定 API
|
||||
|
||||
不负责:
|
||||
- 大段业务逻辑
|
||||
- 状态推进细节
|
||||
- 复杂流程编排
|
||||
|
||||
#### B. Workflow Service Layer
|
||||
职责:
|
||||
- 接收“启动日报任务”“恢复任务”“查询状态”等请求
|
||||
- 管理 run 生命周期
|
||||
- 将调用分发给 runtime/stage engine
|
||||
|
||||
#### C. Workflow Runtime / Stage Engine
|
||||
职责:
|
||||
- 执行各阶段
|
||||
- 记录阶段状态
|
||||
- 收集中间产物
|
||||
- 更新运行态文件
|
||||
- 支持 resume / rerun
|
||||
|
||||
#### D. Reader Core Modules
|
||||
继续复用现有业务能力:
|
||||
- `summary_mcp.core.*`
|
||||
- `summary_mcp.workflows.*`
|
||||
- 现有脚本中成熟的逻辑
|
||||
|
||||
这层应该尽量保持“业务逻辑纯净”,不要耦合 MCP 协议。
|
||||
|
||||
---
|
||||
|
||||
## 7. 核心抽象
|
||||
|
||||
### 7.1 Run
|
||||
|
||||
每一次完整的 reader 工作流执行,都对应一个 `run_id`。
|
||||
|
||||
建议语义:
|
||||
|
||||
- `run_id` 是 reader 工作流的一等公民
|
||||
- 所有状态、产物、恢复都围绕 `run_id` 建立
|
||||
- 所有下游消费都优先通过 `run_id` 取结果,而不是手拼路径
|
||||
|
||||
建议输出目录继续沿用现有结构:
|
||||
|
||||
```text
|
||||
outputs/freshrss/rerun/<run_id>/
|
||||
```
|
||||
|
||||
### 7.2 Stage
|
||||
|
||||
建议将 reader 的正式工作流拆为明确阶段:
|
||||
|
||||
1. `fetch_feed`
|
||||
2. `extract_articles`
|
||||
3. `generate_summaries`
|
||||
4. `apply_filters`
|
||||
5. `build_delivery_payload`
|
||||
6. `write_run_report`
|
||||
7. `completed`
|
||||
|
||||
对于单篇总结链路,可单独建另一类 workflow,或者作为独立 run type。
|
||||
|
||||
### 7.3 Status
|
||||
|
||||
建议统一使用:
|
||||
|
||||
- `queued`
|
||||
- `running`
|
||||
- `partial`
|
||||
- `success`
|
||||
- `failed`
|
||||
- `cancelled`
|
||||
|
||||
其中:
|
||||
|
||||
- `partial` 表示已有阶段成功,但整体尚未完成或部分失败
|
||||
- `failed` 表示本次 run 已终止且未达到最终成功
|
||||
|
||||
### 7.4 Artifact
|
||||
|
||||
所有关键输出都应被注册为 artifact,而不只是“写到了某个目录”。
|
||||
|
||||
artifact 至少应包含:
|
||||
|
||||
- `name`
|
||||
- `path`
|
||||
- `kind`
|
||||
- `stage`
|
||||
- `exists`
|
||||
- `created_at`
|
||||
- `metadata`
|
||||
|
||||
关键 artifact 示例:
|
||||
|
||||
- `raw_output`
|
||||
- `extracted_dir`
|
||||
- `delivery_payload`
|
||||
- `digest_brief`
|
||||
- `run_report`
|
||||
- `single_summary_dir`
|
||||
|
||||
---
|
||||
|
||||
## 8. 运行态持久化设计
|
||||
|
||||
### 8.1 为什么要单独持久化 run state
|
||||
|
||||
仅靠目录中是否存在某些 json 文件,不足以稳定表达:
|
||||
|
||||
- 当前是否还在跑
|
||||
- 跑到了哪个阶段
|
||||
- 哪个阶段失败
|
||||
- 是否可以 resume
|
||||
- 哪些 artifact 已确认可用
|
||||
|
||||
因此必须增加显式的运行态文件。
|
||||
|
||||
### 8.2 建议文件
|
||||
|
||||
建议在每个 run 目录下引入:
|
||||
|
||||
```text
|
||||
outputs/freshrss/rerun/<run_id>/run-state.json
|
||||
```
|
||||
|
||||
### 8.3 run-state.json 建议结构
|
||||
|
||||
```json
|
||||
{
|
||||
"run_id": "20260407-093000",
|
||||
"workflow": "freshrss_daily_digest",
|
||||
"run_type": "daily_digest",
|
||||
"status": "running",
|
||||
"current_stage": "extract_articles",
|
||||
"started_at": "2026-04-07T09:30:00+08:00",
|
||||
"updated_at": "2026-04-07T09:31:10+08:00",
|
||||
"finished_at": null,
|
||||
"input": {
|
||||
"limit": 5,
|
||||
"mark_read": true,
|
||||
"include_read": false,
|
||||
"debug_artifacts": false
|
||||
},
|
||||
"stages": [
|
||||
{
|
||||
"name": "fetch_feed",
|
||||
"status": "success",
|
||||
"started_at": "2026-04-07T09:30:00+08:00",
|
||||
"finished_at": "2026-04-07T09:30:05+08:00",
|
||||
"outputs": {
|
||||
"pulled_count": 5,
|
||||
"raw_output": "outputs/freshrss/rerun/20260407-093000/raw/freshrss.raw.json"
|
||||
},
|
||||
"error": null
|
||||
},
|
||||
{
|
||||
"name": "extract_articles",
|
||||
"status": "running",
|
||||
"started_at": "2026-04-07T09:30:05+08:00",
|
||||
"finished_at": null,
|
||||
"outputs": {
|
||||
"completed_items": 3,
|
||||
"expected_items": 5
|
||||
},
|
||||
"error": null
|
||||
}
|
||||
],
|
||||
"artifacts": [
|
||||
{
|
||||
"name": "raw_output",
|
||||
"kind": "json",
|
||||
"stage": "fetch_feed",
|
||||
"path": "outputs/freshrss/rerun/20260407-093000/raw/freshrss.raw.json",
|
||||
"exists": true
|
||||
}
|
||||
],
|
||||
"error": null,
|
||||
"recovery": {
|
||||
"resumable": true,
|
||||
"resume_from_stage": "extract_articles",
|
||||
"last_success_stage": "fetch_feed"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 8.4 设计原则
|
||||
|
||||
- 运行中每完成一个 stage,就更新一次 `run-state.json`
|
||||
- 不要求数据库,先以文件系统为准
|
||||
- 所有对外 status 查询优先读 `run-state.json`
|
||||
- 其他 output 文件仍然可以保留现有格式和目录结构
|
||||
|
||||
---
|
||||
|
||||
## 9. MCP Tool 设计
|
||||
|
||||
不建议继续把正式能力收口成一个“巨型同步工具”。
|
||||
|
||||
建议拆成以下 MCP tools。
|
||||
|
||||
### 9.1 任务启动类
|
||||
|
||||
#### `start_freshrss_digest_run`
|
||||
|
||||
作用:
|
||||
- 启动一条新的日报工作流
|
||||
- 默认返回 `run_id`,而不是要求调用方一直同步等待到最后
|
||||
|
||||
输入示例:
|
||||
|
||||
```json
|
||||
{
|
||||
"limit": 5,
|
||||
"mark_read": true,
|
||||
"include_read": false,
|
||||
"debug_artifacts": false,
|
||||
"timeout_seconds": 60,
|
||||
"max_retries": 2
|
||||
}
|
||||
```
|
||||
|
||||
输出示例:
|
||||
|
||||
```json
|
||||
{
|
||||
"run_id": "20260407-093000",
|
||||
"status": "queued",
|
||||
"workflow": "freshrss_daily_digest",
|
||||
"output_dir": "outputs/freshrss/rerun/20260407-093000"
|
||||
}
|
||||
```
|
||||
|
||||
#### 兼容策略
|
||||
|
||||
第一阶段可保留现有 `run_freshrss_openclaw_pipeline`,但其语义逐步调整为:
|
||||
|
||||
- 内部复用新的 workflow runtime
|
||||
- 可选 `wait=true/false`
|
||||
- 默认建议 `wait=false`
|
||||
|
||||
---
|
||||
|
||||
### 9.2 状态查询类
|
||||
|
||||
#### `get_run_status`
|
||||
|
||||
作用:
|
||||
- 查询 run 当前状态
|
||||
- 返回 current stage、completed stages、关键 artifact、失败信息
|
||||
|
||||
输入:
|
||||
|
||||
```json
|
||||
{
|
||||
"run_id": "20260407-093000"
|
||||
}
|
||||
```
|
||||
|
||||
输出字段建议:
|
||||
|
||||
- `run_id`
|
||||
- `workflow`
|
||||
- `status`
|
||||
- `current_stage`
|
||||
- `started_at`
|
||||
- `updated_at`
|
||||
- `finished_at`
|
||||
- `progress`
|
||||
- `completed_stages`
|
||||
- `failed_stage`
|
||||
- `error_summary`
|
||||
- `artifacts`
|
||||
- `recovery`
|
||||
|
||||
#### `list_runs`
|
||||
|
||||
作用:
|
||||
- 支持按日期、状态、workflow 类型查看最近 runs
|
||||
|
||||
输入示例:
|
||||
|
||||
```json
|
||||
{
|
||||
"workflow": "freshrss_daily_digest",
|
||||
"status": "failed",
|
||||
"latest_n": 10
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 9.3 结果读取类
|
||||
|
||||
#### `get_delivery_payload`
|
||||
|
||||
作用:
|
||||
- 按 `run_id` 获取 delivery payload
|
||||
- 返回结构化对象,不要求调用方自己读文件
|
||||
|
||||
#### `get_digest_brief`
|
||||
|
||||
作用:
|
||||
- 按 `run_id` 获取 digest brief
|
||||
- 可支持 `mode=public|internal`
|
||||
|
||||
#### `get_run_report`
|
||||
|
||||
作用:
|
||||
- 返回 run-report 内容
|
||||
- 供 OpenClaw / 调试 / 运维查看
|
||||
|
||||
#### `get_extracted_article`
|
||||
|
||||
作用:
|
||||
- 获取单篇 extracted 内容
|
||||
- 输入:`run_id + item_id` 或 `run_id + item_index`
|
||||
|
||||
#### `list_run_artifacts`
|
||||
|
||||
作用:
|
||||
- 统一列出 run 下当前注册的 artifact
|
||||
|
||||
---
|
||||
|
||||
### 9.4 恢复与补跑类
|
||||
|
||||
#### `resume_run`
|
||||
|
||||
作用:
|
||||
- 基于已有 `run-state.json` 和中间产物,从可恢复点继续
|
||||
|
||||
输入示例:
|
||||
|
||||
```json
|
||||
{
|
||||
"run_id": "20260407-093000"
|
||||
}
|
||||
```
|
||||
|
||||
#### `rerun_stage`
|
||||
|
||||
作用:
|
||||
- 指定从某个阶段开始重跑
|
||||
|
||||
输入示例:
|
||||
|
||||
```json
|
||||
{
|
||||
"run_id": "20260407-093000",
|
||||
"stage": "generate_summaries",
|
||||
"force": true
|
||||
}
|
||||
```
|
||||
|
||||
#### `explain_run_failure`
|
||||
|
||||
作用:
|
||||
- 给出适合 OpenClaw / 人类阅读的失败解释
|
||||
- 不只是 Python stacktrace
|
||||
|
||||
---
|
||||
|
||||
### 9.5 单篇总结相关工具
|
||||
|
||||
现有 `generate_article_summaries` 可继续保留,但建议长期也纳入 run 模型。
|
||||
|
||||
后续可扩展为:
|
||||
|
||||
- `start_article_summary_run`
|
||||
- `get_article_summary_status`
|
||||
- `get_article_summary_outputs`
|
||||
|
||||
短期内,如果单篇总结执行耗时可控,也可以先保持同步接口。
|
||||
|
||||
---
|
||||
|
||||
## 10. 向后兼容策略
|
||||
|
||||
为了降低迁移成本,不建议一次性砍掉现有接口。
|
||||
|
||||
### 10.1 保留现有 MCP 工具
|
||||
|
||||
短期保留:
|
||||
|
||||
- `run_freshrss_openclaw_pipeline`
|
||||
- `generate_article_summaries`
|
||||
|
||||
但内部逐步改为调用新的 workflow runtime。
|
||||
|
||||
### 10.2 调整 `run_freshrss_openclaw_pipeline` 语义
|
||||
|
||||
建议演进成:
|
||||
|
||||
- `wait=true` 时:保留当前类似同步行为
|
||||
- `wait=false` 时:只返回 `run_id`
|
||||
- 若未显式指定,生产建议默认 `wait=false`
|
||||
|
||||
### 10.3 CLI 的新定位
|
||||
|
||||
CLI 继续保留,但只作为:
|
||||
|
||||
- debug
|
||||
- fallback
|
||||
- 本地排障
|
||||
- 开发验证
|
||||
|
||||
CLI 最好只是新 runtime 的薄包装,而不是另一套独立实现。
|
||||
|
||||
---
|
||||
|
||||
## 11. 与 OpenClaw 的集成方式
|
||||
|
||||
### 11.1 旧模式
|
||||
|
||||
OpenClaw:
|
||||
- 直接 exec reader 脚本
|
||||
- 等待整条 CLI 跑完
|
||||
- 根据 stdout 或目录文件判断是否成功
|
||||
|
||||
### 11.2 新模式
|
||||
|
||||
OpenClaw:
|
||||
1. 调用 `start_freshrss_digest_run`
|
||||
2. 获得 `run_id`
|
||||
3. 周期性调用 `get_run_status`
|
||||
4. status = `success` 后调用 `get_delivery_payload`
|
||||
5. 再执行下游 Hugo / chat / IMA 编排
|
||||
|
||||
这样会有几个明显好处:
|
||||
|
||||
- 上游 reader 运行态可观察
|
||||
- OpenClaw 不必绑死在长 CLI 会话上
|
||||
- 中途失败可以明确知道失败位置
|
||||
- 下游编排只依赖结构化结果
|
||||
|
||||
---
|
||||
|
||||
## 12. 最小实现方案(MVP)
|
||||
|
||||
为了尽快落地,不建议一次做到最重。
|
||||
|
||||
### Phase 1:先做内部状态化
|
||||
|
||||
目标:
|
||||
- 在现有 pipeline 内引入 `run_id`
|
||||
- 新增 `run-state.json`
|
||||
- 固化 stage 切分
|
||||
- 所有关键输出注册成 artifact
|
||||
- 将 resume 所需信息写入 recovery 字段
|
||||
|
||||
这一步完成后,即使对外接口还没完全变化,内部也已经不再是“黑箱长函数”。
|
||||
|
||||
### Phase 2:新增 MCP 状态查询接口
|
||||
|
||||
目标:
|
||||
- 新增 `get_run_status`
|
||||
- 新增 `get_delivery_payload`
|
||||
- 新增 `list_runs`
|
||||
- 新增 `resume_run`
|
||||
|
||||
这一步完成后,OpenClaw 就可以逐步改走正式服务调用。
|
||||
|
||||
### Phase 3:调整现有生产接入
|
||||
|
||||
目标:
|
||||
- OpenClaw 默认不再 exec `scripts/run_freshrss_pipeline.py`
|
||||
- OpenClaw 默认走 MCP run + status + payload 模式
|
||||
- 将 CLI 降级为 debug/fallback
|
||||
|
||||
---
|
||||
|
||||
## 13. 目录与模块建议
|
||||
|
||||
建议新增一层 workflow runtime 模块,例如:
|
||||
|
||||
```text
|
||||
src/summary_mcp/
|
||||
server.py
|
||||
runtime/
|
||||
run_store.py
|
||||
artifact_store.py
|
||||
state_models.py
|
||||
workflow_service.py
|
||||
stage_runner.py
|
||||
workflows/
|
||||
freshrss_pipeline.py
|
||||
article_summary.py
|
||||
core/
|
||||
...
|
||||
```
|
||||
|
||||
### 建议职责
|
||||
|
||||
- `runtime/state_models.py`
|
||||
- RunState / StageState / Artifact models
|
||||
|
||||
- `runtime/run_store.py`
|
||||
- 读写 `run-state.json`
|
||||
|
||||
- `runtime/artifact_store.py`
|
||||
- artifact 注册与查询
|
||||
|
||||
- `runtime/workflow_service.py`
|
||||
- start / status / resume / rerun 核心服务
|
||||
|
||||
- `runtime/stage_runner.py`
|
||||
- 阶段推进与失败捕获
|
||||
|
||||
这样可以保持:
|
||||
|
||||
- 协议层清晰
|
||||
- 运行态管理独立
|
||||
- 业务逻辑复用现有 workflows
|
||||
|
||||
---
|
||||
|
||||
## 14. 风险与注意事项
|
||||
|
||||
### 14.1 不要把“异步”理解成“一定要上复杂队列”
|
||||
|
||||
当前阶段,异步的核心不是上 Redis/Celery,而是:
|
||||
|
||||
- 有 `run_id`
|
||||
- 有状态文件
|
||||
- 可以先启动、后查询
|
||||
|
||||
单机单进程也完全可以做到。
|
||||
|
||||
### 14.2 不要让 OpenClaw 继续依赖 reader 内部目录细节
|
||||
|
||||
OpenClaw 可以知道输出目录存在,但不应继续以“自己拼 `outputs/.../*.json`”作为主交互方式。
|
||||
|
||||
正式模式下,优先通过 MCP 取:
|
||||
|
||||
- status
|
||||
- payload
|
||||
- report
|
||||
- artifact list
|
||||
|
||||
### 14.3 不要保留两套行为漂移的实现
|
||||
|
||||
如果 CLI 和 MCP 背后各跑各的逻辑,后续一定会漂移。
|
||||
|
||||
正确做法:
|
||||
- 先统一 runtime
|
||||
- 再让 CLI / MCP 都调用同一套 runtime
|
||||
|
||||
---
|
||||
|
||||
## 15. 最终建议
|
||||
|
||||
### 架构判断
|
||||
|
||||
Reader 现在已经不再是一个简单脚本仓库,而是:
|
||||
|
||||
- 有明确上游输入
|
||||
- 有稳定工作流
|
||||
- 有中间产物
|
||||
- 有下游消费者
|
||||
- 有恢复与补跑需求
|
||||
|
||||
因此它应该正式升级为:
|
||||
|
||||
> **Reader Workflow MCP Service**
|
||||
|
||||
### 最终建议清单
|
||||
|
||||
1. 把 `run_id / stage / status / artifact / recovery` 作为正式核心抽象
|
||||
2. 在 `outputs/freshrss/rerun/<run_id>/` 下新增 `run-state.json`
|
||||
3. 新增 `get_run_status / list_runs / get_delivery_payload / resume_run`
|
||||
4. 让现有 `run_freshrss_openclaw_pipeline` 内部复用新 runtime
|
||||
5. 让 OpenClaw 逐步从 exec 切到 MCP 调用
|
||||
6. CLI 保留,但降级为 debug/fallback
|
||||
|
||||
---
|
||||
|
||||
## 16. 一句话结论
|
||||
|
||||
Reader 的正式生产能力不应再主要依赖长 CLI exec,而应演进为:
|
||||
|
||||
**以 MCP 为正式入口、以 run-state 为运行真相、以 artifact 为交付契约、以 OpenClaw 为下游编排层的工作流服务。**
|
||||
@@ -0,0 +1,209 @@
|
||||
# Reader MCP 实施计划
|
||||
|
||||
## 目标
|
||||
|
||||
把 reader 从“生产上主要依赖长 CLI/exec”推进到“以 MCP 为正式入口的工作流服务”。
|
||||
|
||||
## 协作分工
|
||||
|
||||
- **架构与方向**:由我负责
|
||||
- **编码实现**:由 Codex 负责
|
||||
- **同步机制**:通过 `plans/` 下规划文档 + `TODO.md` 保持进度与方向一致
|
||||
|
||||
## 当前权威文档
|
||||
|
||||
实现前,必须先读:
|
||||
|
||||
1. `plans/reader-mcp-architecture-design.md`
|
||||
2. `plans/reader-mcp-implementation-plan.md`
|
||||
3. `TODO.md`
|
||||
4. `plans/issues/2026-04-06-reader-digest-sigterm.md`
|
||||
5. `docs/openclaw/openclaw-handoff.md`
|
||||
|
||||
如实现细节与旧文档冲突,以:
|
||||
|
||||
1. 最新架构设计文档
|
||||
2. 最新实施计划
|
||||
3. TODO 当前项
|
||||
|
||||
为准。
|
||||
|
||||
---
|
||||
|
||||
## 里程碑
|
||||
|
||||
### M1:运行态落地(最优先)
|
||||
|
||||
目标:让现有 freshrss pipeline 拥有明确 run state。
|
||||
|
||||
交付:
|
||||
- 新增 `run-state.json` 持久化能力
|
||||
- 定义 `RunState / StageState / ArtifactRecord` 模型
|
||||
- 现有 pipeline 按 stage 更新状态
|
||||
- 保持现有产物目录兼容
|
||||
|
||||
完成标准:
|
||||
- 任意一次 run 都能产出 `outputs/freshrss/rerun/<run_id>/run-state.json`
|
||||
- 中途失败时也能看到失败阶段和已有 artifacts
|
||||
|
||||
---
|
||||
|
||||
### M2:状态查询接口
|
||||
|
||||
目标:通过 MCP 查询 run 状态,不再只靠 CLI 或手看目录。
|
||||
|
||||
交付:
|
||||
- 新增 `get_run_status`
|
||||
- 新增 `list_runs`
|
||||
- 新增 `list_run_artifacts`
|
||||
|
||||
完成标准:
|
||||
- OpenClaw 可通过 MCP 查询 run 当前状态
|
||||
- 不需要直接读磁盘路径判断是否成功
|
||||
|
||||
---
|
||||
|
||||
### M3:结果读取接口
|
||||
|
||||
目标:通过 MCP 获取结果内容,而不是自己拼文件路径。
|
||||
|
||||
交付:
|
||||
- 新增 `get_delivery_payload`
|
||||
- 新增 `get_run_report`
|
||||
- 视情况新增 `get_digest_brief`
|
||||
- 视情况新增 `get_extracted_article`
|
||||
|
||||
完成标准:
|
||||
- OpenClaw 只要知道 `run_id`,就能获取关键结果
|
||||
|
||||
---
|
||||
|
||||
### M4:恢复与补跑
|
||||
|
||||
目标:reader 具备正式的恢复机制。
|
||||
|
||||
交付:
|
||||
- 新增 `resume_run`
|
||||
- 视情况新增 `rerun_stage`
|
||||
- 在 `run-state.json` 中记录 recovery 信息
|
||||
|
||||
完成标准:
|
||||
- 至少支持从最近成功 stage 之后继续执行
|
||||
- 能给出可恢复/不可恢复的明确判断
|
||||
|
||||
---
|
||||
|
||||
### M5:生产入口切换
|
||||
|
||||
目标:OpenClaw 正式从 exec 模式切到 MCP 模式。
|
||||
|
||||
交付:
|
||||
- 保留 CLI 作为 debug/fallback
|
||||
- 生产推荐入口改为 MCP run + status + payload
|
||||
- 更新 handoff / README / docs
|
||||
|
||||
完成标准:
|
||||
- 正式流程默认不再依赖长 CLI exec
|
||||
|
||||
---
|
||||
|
||||
## 编码原则
|
||||
|
||||
1. **优先复用现有业务逻辑**
|
||||
- 不要重写成熟的 FreshRSS / extract / summary / filter 逻辑
|
||||
- 优先抽 runtime 层把现有逻辑包起来
|
||||
|
||||
2. **先状态化,再协议扩展**
|
||||
- 先把 run-state 跑通
|
||||
- 再补 MCP tools
|
||||
|
||||
3. **保持向后兼容**
|
||||
- `run_freshrss_openclaw_pipeline` 先保留
|
||||
- 内部逐步改为调用新的 runtime
|
||||
|
||||
4. **CLI 降级,不删除**
|
||||
- 仍保留 debug/fallback 价值
|
||||
- 但不要让 CLI 和 MCP 背后出现两套逻辑
|
||||
|
||||
5. **小步提交**
|
||||
- 每完成一个可验证的小目标就提交
|
||||
- 不要攒一个超大变更
|
||||
|
||||
---
|
||||
|
||||
## 建议目录改造
|
||||
|
||||
建议新增:
|
||||
|
||||
```text
|
||||
src/summary_mcp/runtime/
|
||||
state_models.py
|
||||
run_store.py
|
||||
artifact_store.py
|
||||
workflow_service.py
|
||||
stage_runner.py
|
||||
```
|
||||
|
||||
说明:
|
||||
- 目录名可微调
|
||||
- 但必须把“运行态管理”从现有 workflows 中抽出来,避免继续黑箱化
|
||||
|
||||
---
|
||||
|
||||
## Codex 工作方式要求
|
||||
|
||||
Codex 每次开始前:
|
||||
|
||||
1. 先读 `plans/reader-mcp-architecture-design.md`
|
||||
2. 再读本文件
|
||||
3. 再读 `TODO.md`
|
||||
4. 只处理 TODO 中 `TODO` 状态的当前优先项
|
||||
|
||||
Codex 每完成一项后:
|
||||
|
||||
1. 更新 `TODO.md`
|
||||
2. 在 TODO 对应项下补:
|
||||
- 完成情况
|
||||
- 改动文件
|
||||
- 遗留风险
|
||||
3. 如实现偏离原架构,必须先更新 `plans/` 文档,再继续代码
|
||||
|
||||
---
|
||||
|
||||
## 决策规则
|
||||
|
||||
如果出现以下情况:
|
||||
|
||||
- 需要新增 MCP tool,但架构设计中未定义
|
||||
- 需要改动现有 pipeline 关键语义
|
||||
- 需要引入数据库 / 队列 / 线程池 / 后台 worker
|
||||
- 需要改变 OpenClaw 与 reader 的边界
|
||||
|
||||
则不允许 Codex自行拍板,必须先回写到:
|
||||
|
||||
- `plans/reader-mcp-architecture-design.md`
|
||||
- 或新增 `plans/issues/*.md`
|
||||
|
||||
由架构层确认后再继续。
|
||||
|
||||
---
|
||||
|
||||
## 当前实现顺序(强约束)
|
||||
|
||||
按以下顺序推进:
|
||||
|
||||
1. `run-state.json` 模型与持久化
|
||||
2. pipeline 中 stage 状态更新
|
||||
3. `get_run_status`
|
||||
4. `list_runs` / `list_run_artifacts`
|
||||
5. `get_delivery_payload` / `get_run_report`
|
||||
6. `resume_run`
|
||||
7. 再考虑 `rerun_stage`
|
||||
|
||||
不要一上来就做复杂异步后台队列。
|
||||
|
||||
---
|
||||
|
||||
## 一句话执行口径
|
||||
|
||||
先把 reader 做成“有运行真相的 MCP 工作流服务”,再做更多工具;不要反过来先堆接口名。
|
||||
@@ -0,0 +1,276 @@
|
||||
# resume_run 最小恢复语义设计
|
||||
|
||||
## 1. 背景
|
||||
|
||||
前面已经完成:
|
||||
|
||||
- run-state 基础设施
|
||||
- 状态查询接口
|
||||
- 结果读取接口
|
||||
|
||||
reader 现在已经具备“运行真相 + 查询 + 结果读取”能力。下一步是补 `resume_run`,让失败后的 run 可以从已有中间产物继续,而不是完全重跑。
|
||||
|
||||
但 `resume_run` 是第一项真正涉及“重新进入执行流程”的能力,复杂度高于前几轮。因此本轮设计必须进一步收窄范围,避免一次性把任意 stage 重入、复杂后台管理、任务调度全拉进来。
|
||||
|
||||
---
|
||||
|
||||
## 2. 本轮目标
|
||||
|
||||
只做 **最小可用的 `resume_run`**:
|
||||
|
||||
> 基于已有 `run-state.json`,让 freshrss pipeline 能从“最近可恢复点”继续执行。
|
||||
|
||||
本轮不追求:
|
||||
|
||||
- 任意 stage 任意重入
|
||||
- 历史 run 的通用恢复
|
||||
- 无 `run-state.json` run 的恢复
|
||||
- 多任务后台管理
|
||||
- 队列 / 数据库 / worker
|
||||
|
||||
---
|
||||
|
||||
## 3. 支持范围
|
||||
|
||||
### 3.1 仅支持 workflow
|
||||
|
||||
只支持:
|
||||
|
||||
- `freshrss_daily_digest`
|
||||
- 或当前 `run-state.workflow` 对应的 freshrss pipeline workflow
|
||||
|
||||
不支持:
|
||||
|
||||
- article summary 独立恢复
|
||||
- 其他未来 workflow
|
||||
|
||||
### 3.2 仅支持有 run-state 的 run
|
||||
|
||||
调用 `resume_run(run_id=...)` 时,必须满足:
|
||||
|
||||
- run 对应目录存在
|
||||
- `run-state.json` 存在
|
||||
- `run-state.json` 可解析
|
||||
|
||||
否则直接返回不可恢复错误,而不是尝试猜目录结构。
|
||||
|
||||
---
|
||||
|
||||
## 4. 最近可恢复点定义
|
||||
|
||||
本轮采用保守定义:
|
||||
|
||||
### 4.1 恢复依据
|
||||
|
||||
优先使用 `run-state.recovery.resume_from_stage`。
|
||||
|
||||
如果没有该字段,使用:
|
||||
|
||||
- `state.status == failed` 时:失败 stage
|
||||
- 否则:拒绝恢复
|
||||
|
||||
### 4.2 只支持以下恢复点
|
||||
|
||||
第一版仅支持从下面几类 stage 恢复:
|
||||
|
||||
1. `generate_summaries`
|
||||
2. `apply_filters`
|
||||
3. `build_delivery_payload`
|
||||
4. `write_run_report`
|
||||
|
||||
### 4.3 明确不支持的恢复点
|
||||
|
||||
第一版暂不支持从以下位置恢复:
|
||||
|
||||
1. `fetch_feed`
|
||||
2. `extract_articles`
|
||||
|
||||
原因:
|
||||
|
||||
- 这两个阶段更依赖外部抓取与逐条内容处理过程
|
||||
- 恢复语义更复杂
|
||||
- 容易和 FreshRSS 读状态、副作用、原始输入不一致问题缠在一起
|
||||
|
||||
如果 run 停在这两个阶段:
|
||||
|
||||
- `resume_run` 返回 `resumable=false`
|
||||
- 并明确建议重新触发新 run,而不是恢复
|
||||
|
||||
---
|
||||
|
||||
## 5. 恢复前置条件
|
||||
|
||||
### 5.1 必须存在的中间产物
|
||||
|
||||
按恢复点要求最小前置产物:
|
||||
|
||||
#### 从 `generate_summaries` 恢复
|
||||
必须至少有:
|
||||
- raw output
|
||||
- extracted outputs(或足以驱动 summary 的当前输入)
|
||||
|
||||
#### 从 `apply_filters` 恢复
|
||||
必须至少有:
|
||||
- extracted outputs
|
||||
- summary outputs
|
||||
|
||||
#### 从 `build_delivery_payload` 恢复
|
||||
必须至少有:
|
||||
- filter / candidate 所需输入已齐备
|
||||
|
||||
#### 从 `write_run_report` 恢复
|
||||
必须至少有:
|
||||
- delivery payload 已存在
|
||||
|
||||
### 5.2 缺失产物处理
|
||||
|
||||
如果恢复点所需的关键产物缺失:
|
||||
|
||||
- 直接返回不可恢复
|
||||
- 返回字段中注明缺失产物名
|
||||
- 不自动降级到更早 stage
|
||||
|
||||
原因:
|
||||
|
||||
- 第一版先避免隐式魔法恢复
|
||||
- 让行为更可预测
|
||||
|
||||
---
|
||||
|
||||
## 6. 执行语义
|
||||
|
||||
### 6.1 resume_run 的行为
|
||||
|
||||
调用 `resume_run(run_id)` 后:
|
||||
|
||||
1. 读取 `run-state.json`
|
||||
2. 校验 workflow、status、recovery 信息
|
||||
3. 判断恢复点是否在本轮支持范围内
|
||||
4. 校验该恢复点所需产物是否齐备
|
||||
5. 复用现有 freshrss pipeline / runtime,从该恢复点之后继续执行
|
||||
6. 更新原 run 的 `run-state.json`
|
||||
|
||||
### 6.2 不新建 run_id
|
||||
|
||||
第一版恢复时:
|
||||
|
||||
- **继续沿用原 run_id**
|
||||
- 不新建子 run / shadow run / retry run
|
||||
|
||||
原因:
|
||||
|
||||
- 保持恢复行为简单直观
|
||||
- 避免多 run 关系管理复杂化
|
||||
|
||||
### 6.3 状态更新
|
||||
|
||||
恢复开始时:
|
||||
|
||||
- `status` 置回 `running`
|
||||
- `current_stage` 置为恢复点
|
||||
- `error` 清理为 null
|
||||
- 在必要时更新 `recovery` 信息
|
||||
|
||||
恢复成功时:
|
||||
|
||||
- 正常走到 `success`
|
||||
|
||||
恢复失败时:
|
||||
|
||||
- 正常写入新的失败信息
|
||||
- 保留新的 `run-state.json`
|
||||
|
||||
---
|
||||
|
||||
## 7. MCP 接口建议
|
||||
|
||||
### 7.1 新增 tool
|
||||
|
||||
建议新增:
|
||||
|
||||
- `resume_run`
|
||||
|
||||
### 7.2 输入
|
||||
|
||||
```json
|
||||
{
|
||||
"run_id": "freshrss-pipeline-20260407-010236"
|
||||
}
|
||||
```
|
||||
|
||||
第一版不加 `from_stage`,避免人为覆盖恢复点逻辑。
|
||||
|
||||
### 7.3 输出
|
||||
|
||||
建议输出:
|
||||
|
||||
```json
|
||||
{
|
||||
"run_id": "freshrss-pipeline-20260407-010236",
|
||||
"workflow": "freshrss_daily_digest",
|
||||
"resumed": true,
|
||||
"resume_from_stage": "build_delivery_payload",
|
||||
"status": "success",
|
||||
"output_dir": "outputs/freshrss/rerun/freshrss-pipeline-20260407-010236",
|
||||
"delivery_payload": {...},
|
||||
"run_report": {...},
|
||||
"message": "Run resumed from build_delivery_payload and completed successfully."
|
||||
}
|
||||
```
|
||||
|
||||
### 7.4 不可恢复时输出
|
||||
|
||||
```json
|
||||
{
|
||||
"run_id": "freshrss-pipeline-20260407-010236",
|
||||
"resumed": false,
|
||||
"status": "failed",
|
||||
"resume_from_stage": "extract_articles",
|
||||
"message": "This run cannot be resumed from extract_articles in the current minimal implementation.",
|
||||
"missing_artifacts": []
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 8. 代码组织建议
|
||||
|
||||
优先在现有 runtime 体系内最小扩展:
|
||||
|
||||
- `run_store.py`
|
||||
- 补载入/更新辅助能力(如还缺)
|
||||
|
||||
- `query_service.py`
|
||||
- 可复用读取 run-state / artifact / report / payload 的查询能力
|
||||
|
||||
- 新增或补充 runtime service
|
||||
- 例如 `resume_service.py` 或在现有 runtime 层增加 resume 逻辑
|
||||
|
||||
- `server.py`
|
||||
- 新增 MCP tool `resume_run`
|
||||
|
||||
关键原则:
|
||||
|
||||
- 不要重写整条 freshrss pipeline
|
||||
- 应尽量让 pipeline 能接受“从某 stage 之后继续”的最小参数
|
||||
|
||||
---
|
||||
|
||||
## 9. 明确不做的事
|
||||
|
||||
本轮明确不做:
|
||||
|
||||
1. `rerun_stage`
|
||||
2. 指定任意 `from_stage`
|
||||
3. 多 workflow 通用恢复框架
|
||||
4. 恢复时自动修补缺失产物
|
||||
5. 历史无 run-state run 的恢复
|
||||
6. 后台异步恢复任务管理
|
||||
|
||||
---
|
||||
|
||||
## 10. 一句话结论
|
||||
|
||||
`resume_run` 第一版只做一件事:
|
||||
|
||||
**对已有 `run-state.json` 的 freshrss run,在恢复点和前置产物都满足时,从最近可恢复点继续执行;否则明确拒绝恢复,不做隐式魔法补救。**
|
||||
@@ -0,0 +1,103 @@
|
||||
# PRD: 单篇文章总结 Prompt 独立化
|
||||
|
||||
## 背景
|
||||
|
||||
reader 项目目前有两条 LLM 总结链路:
|
||||
|
||||
1. 日报摘要(freshrss_pipeline.py):生成每日候选日报,输出短摘要卡片,目标是快速筛选
|
||||
2. 单篇精读总结(article_summary.py):对精选文章做深度沉淀,上传到知识库
|
||||
|
||||
问题:单篇精读总结目前复用了日报摘要的同一份 prompt(outputs/prompts/llm-summary-prompt.txt),
|
||||
导致即使输入了完整正文(article.plain_text),输出仍然是 2-3 句的短摘要卡片,没有体现正文的深度信息。
|
||||
|
||||
---
|
||||
|
||||
## 目标
|
||||
|
||||
让单篇精读总结和日报摘要彻底分离,单篇总结的输出应该像一篇知识沉淀笔记,而不是日报候选摘要。
|
||||
|
||||
---
|
||||
|
||||
## 需求说明
|
||||
|
||||
### 1. 新增 outputs/prompts/article-summary-prompt.txt
|
||||
|
||||
新增一个专门用于单篇文章知识沉淀的 prompt 文件,与日报摘要 prompt 完全独立。
|
||||
|
||||
prompt 设计目标:
|
||||
- 基于正文全文(article.plain_text)做深度总结
|
||||
- 允许较长输出(不限制字数,以内容完整为准)
|
||||
- 输出结构偏向知识沉淀,而非候选摘要
|
||||
|
||||
输出 JSON schema 新增字段:
|
||||
|
||||
- core_conclusion:作者最核心的结论,1-2 句,精准
|
||||
- main_argument:文章的主要论点/主张,可展开,允许多句
|
||||
- key_methods:关键方法/机制/技术手段,3-6 条,每条一句
|
||||
- important_details:值得记录的细节、数据、案例,3-6 条,每条一句
|
||||
- reusable_insights:可复用于其他场景的启发或观点,2-4 条
|
||||
- keywords:具体实体、工具名、方法名,5-8 个
|
||||
- topics:更高层的主题标签,3-5 个
|
||||
- category:内容分类(资讯/方法论/工具实践/观点评论)
|
||||
- worth_keeping:是否值得长期保留,bool
|
||||
- reason:沉淀理由,一句话
|
||||
|
||||
### 2. 修改 src/summary_mcp/workflows/article_summary.py
|
||||
|
||||
改动 1:默认 prompt 路径
|
||||
修改前:DEFAULT_PROMPT_PATH = OUTPUT_ROOT / "prompts" / "llm-summary-prompt.txt"
|
||||
修改后:DEFAULT_PROMPT_PATH = OUTPUT_ROOT / "prompts" / "article-summary-prompt.txt"
|
||||
|
||||
改动 2:Markdown 输出结构
|
||||
summarize_selected_articles() 中 lines 构造部分,从 summary/highlights 字段改为新 schema 字段。
|
||||
|
||||
新输出格式:
|
||||
# 标题
|
||||
Source: url
|
||||
Category: 方法论
|
||||
## 核心结论
|
||||
## 主要论点
|
||||
## 关键方法 / 机制
|
||||
## 重要细节
|
||||
## 可复用启发
|
||||
## 关键词
|
||||
## 主题
|
||||
|
||||
### 3. Validator 处理
|
||||
|
||||
src/summary_mcp/validators/llm_result.py 目前按日报摘要 schema 校验,与新 schema 不兼容。
|
||||
|
||||
推荐方案:article_summary.py 调用 run_loop_payload() 时传入宽松 validator,只校验 JSON 格式合法,不校验具体字段。
|
||||
备选方案:新增 validate_article_summary_payload() 专门校验新 schema。
|
||||
|
||||
---
|
||||
|
||||
## 文件变更清单
|
||||
|
||||
| 文件 | 操作 |
|
||||
|------|------|
|
||||
| outputs/prompts/article-summary-prompt.txt | 新增 |
|
||||
| src/summary_mcp/workflows/article_summary.py | 修改 prompt 路径 + 修改 markdown 输出结构 |
|
||||
| src/summary_mcp/validators/llm_result.py | 按需修改 |
|
||||
|
||||
---
|
||||
|
||||
## 不影响的部分
|
||||
|
||||
- freshrss_pipeline.py 不变,日报摘要链路保持原有 prompt 和 schema
|
||||
- summary_loop.py 核心逻辑不变
|
||||
- extracted JSON 格式不变,article.plain_text 仍然是输入
|
||||
|
||||
---
|
||||
|
||||
## 验证方式
|
||||
|
||||
用已有 extracted 数据跑一次单篇总结:
|
||||
|
||||
cd /home/ubuntu/zhu/github/reader
|
||||
python -m summary_mcp.workflows.article_summary \
|
||||
--extracted outputs/freshrss/rerun/<run_id>/extracted/selected.extracted.json \
|
||||
--selected-ids <item_id> \
|
||||
--output-dir /tmp/article-summary-test
|
||||
|
||||
验证输出 markdown 包含「核心结论」「主要论点」「关键方法」等新字段,不再是 140 字短摘要。
|
||||
@@ -0,0 +1,54 @@
|
||||
# 重构计划:拆分 run_freshrss_pipeline 为内部辅助函数
|
||||
|
||||
## 背景
|
||||
|
||||
`src/summary_mcp/workflows/freshrss_pipeline.py` 中的 `run_freshrss_pipeline` 函数约 330 行,
|
||||
将 6 个阶段全部写在一个函数体内,阅读、测试和后续扩展(如并发、重试策略)都比较困难。
|
||||
目标是在不改变任何外部行为的前提下,提取 3 个内部辅助函数。
|
||||
|
||||
当前无测试覆盖,验证方式为函数签名和返回值结构保持不变。
|
||||
|
||||
## 涉及文件
|
||||
|
||||
- `src/summary_mcp/workflows/freshrss_pipeline.py`(唯一修改文件)
|
||||
|
||||
## 提取 3 个内部辅助函数
|
||||
|
||||
### 1. `_process_item(...)` — 单条 item 处理(当前 160-258 行)
|
||||
|
||||
提取 for 循环体(约 100 行)为独立函数。
|
||||
返回 `item_report` dict;当 status 为 `delivered` 时,额外携带 `_candidate` 和 `_external_id`
|
||||
两个临时键供调用方解包,写盘前剥离这两个键。
|
||||
用 `return item_report` 替代循环中的 `continue`。
|
||||
|
||||
### 2. `_build_and_persist_delivery(...)` — 阶段 4+5(当前 260-279 行)
|
||||
|
||||
提取 payload 构建 + 词元索引持久化。
|
||||
返回 `(delivery_payload, keyword_index_result)`。
|
||||
|
||||
### 3. `_build_run_report(...)` — 报告组装(当前 288-308 行)
|
||||
|
||||
提取 status_counts 统计 + report dict 构建。
|
||||
返回 report dict,主函数拿到后再调用 `_save_json` 写盘。
|
||||
|
||||
## 重构后主函数结构(约 80 行)
|
||||
|
||||
1. 阶段 1:初始化(不变)
|
||||
2. 阶段 2:拉取 FreshRSS + 加载规则/context(不变)
|
||||
3. 阶段 3:for 循环调用 `_process_item(...)`,从返回值解包 candidate
|
||||
4. 阶段 4+5:`delivery_payload, keyword_index_result = _build_and_persist_delivery(...)`
|
||||
5. 阶段 6:标记已读,`report = _build_run_report(...)`,写盘,返回
|
||||
|
||||
## 约束
|
||||
|
||||
- `run_freshrss_pipeline` 外部签名不变
|
||||
- 返回 dict 的键结构不变
|
||||
- 所有文件写入路径不变
|
||||
- 纯结构性重构,无行为变化
|
||||
- 无需新增 import
|
||||
|
||||
## 验证
|
||||
|
||||
重构完成后运行:
|
||||
|
||||
python -c "from summary_mcp.workflows.freshrss_pipeline import run_freshrss_pipeline; print('ok')"
|
||||
@@ -0,0 +1,34 @@
|
||||
# 示例文章标题
|
||||
|
||||
原文链接:
|
||||
https://example.com/article
|
||||
|
||||
## 核心结论
|
||||
|
||||
这里先用一段短句概括最重要的判断。
|
||||
|
||||
如果结论较长,继续拆成第二个短段,而不是塞成一个大长段。
|
||||
|
||||
## 主要论点
|
||||
|
||||
先交代文章的核心主张。
|
||||
|
||||
再单独起一段解释支撑这个主张的关键论据。
|
||||
|
||||
如果还有补充判断,继续拆段,保证在 IMA 中阅读时不会挤成一整坨。
|
||||
|
||||
## 关键方法 / 机制
|
||||
|
||||
- 要点一
|
||||
- 要点二
|
||||
- 要点三
|
||||
|
||||
## 重要细节
|
||||
|
||||
- 细节一
|
||||
- 细节二
|
||||
|
||||
## 可复用启发
|
||||
|
||||
- 启发一
|
||||
- 启发二
|
||||
@@ -0,0 +1,40 @@
|
||||
+++
|
||||
title = "AI 日报 · 示例"
|
||||
date = 2026-04-01T16:55:00+08:00
|
||||
summary = "围绕 Agent 架构分层、Skills 标准化与桌面 Agent 工程实践的当日观察。"
|
||||
+++
|
||||
|
||||
> ⚠️ 格式规范(生成 Hugo 时必须遵守):
|
||||
> - 文章编号:`1.` `2.` `3.`(阿拉伯数字 + 点),禁止 `① ② ③` / `一、二、三` 等变体
|
||||
> - 四个 section 缺一不可:`今日概览` → `今日重点` → `趋势观察` → `延伸阅读`
|
||||
> - 每篇文章结构:标题来源 → 摘要段 → "值得关注:"三点 → "这篇更值得关注的理由"段
|
||||
|
||||
# 今日概览
|
||||
|
||||
今天的公开候选主要集中在 AI Agent 的架构演进、工具化落地与工程化实践三条线索上。相比早期偏概念展示的讨论,这一批内容更强调模块化能力栈、真实部署路径与系统可维护性,说明行业关注点正在从“模型能做什么”转向“系统如何稳定落地并持续复用”。
|
||||
|
||||
## 今日重点
|
||||
|
||||
### 1. 学习笔记:从 Agent 到 Skills — AI 智能体架构的范式转变
|
||||
来源:阿里云开发者
|
||||
|
||||
文章分析了 AI 智能体架构从单体 Agent 向模块化 Skills 的范式转变。Anthropic 先后推出 MCP 和 Agent Skills 开放标准,构建了知识、工具、协作和运行分层架构。文章通过一个自动化美化相册的真实项目,对比了 Claude Code 与 OpenClaw 两种实现方案,验证了新架构的可复用性与灵活性。
|
||||
|
||||
值得关注:
|
||||
- Anthropic 在 14 个月内先后推出 MCP 和 Agent Skills 两个开放标准,推动 AI 智能体架构分层化。
|
||||
- 新范式核心是构建薄 Agent 引擎与可组合的 Skills 库,取代为每个用例定制单体 Agent。
|
||||
- 文章通过自动化美化相册项目,实操演示了 Skills、MCP、OpenClaw 和 A2A 协议如何协同工作。
|
||||
|
||||
这篇内容更值得关注的原因在于,它不只是提出了“Agent 要模块化”这个判断,而是把开放标准、分层架构和真实项目案例串成了一条完整论证链,能直接支撑今天日报的主线。
|
||||
|
||||
## 趋势观察
|
||||
|
||||
1. Agent 正在从单体能力转向可组合的模块化体系。无论是 Skills、MCP、记忆还是运行时编排,这批内容都在强调解耦与复用,而不是把智能体继续当成一个不可拆分的黑箱。
|
||||
2. 工程化正在变成 AI 应用竞争的主战场。桌面 Agent、企业级架构和部署实践类内容增多,说明真正的差异化开始落在接入现有流程、控制风险和提升可维护性上。
|
||||
3. AI 能力的竞争点正在上移。模型本身仍重要,但真正可持续的优势越来越来自系统设计、工作流整合和对业务场景的理解。
|
||||
|
||||
## 延伸阅读
|
||||
|
||||
- [学习笔记:从 Agent 到 Skills — AI 智能体架构的范式转变](https://example.com/a)|阿里云开发者
|
||||
- [Agent Skills:打通可复用专业领域知识的最后一公里](https://example.com/b)|阿里云开发者
|
||||
- [CoPaw深度解析:源码架构和功能实践](https://example.com/c)|阿里云开发者
|
||||
Binary file not shown.
+314
@@ -0,0 +1,314 @@
|
||||
#!/usr/bin/env python3
|
||||
"""
|
||||
Generate semantic keyword suggestions using LLM.
|
||||
|
||||
Covers what surface-form rules cannot:
|
||||
- semantic alias (abbreviation ↔ full name, Chinese ↔ English, synonym)
|
||||
- stopword (overly broad / low-discrimination terms)
|
||||
- promote (new term that aligns with user's focus areas)
|
||||
|
||||
Usage:
|
||||
python scripts/generate_term_cleanup_semantic_suggestions.py \
|
||||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
|
||||
--output outputs/term_index/review/term-cleanup-semantic-suggestions-2026-05-14.json
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
import time
|
||||
from datetime import datetime, timezone
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
from urllib.request import Request, urlopen
|
||||
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
DEFAULT_BUNDLE_PATH = REPO_ROOT / "outputs" / "term_index" / "review" / "keyword-cleanup-bundle.json"
|
||||
DEFAULT_OUTPUT_DIR = REPO_ROOT / "outputs" / "term_index" / "review"
|
||||
|
||||
|
||||
def _load_json(path: Path) -> Any:
|
||||
return json.loads(path.read_text(encoding="utf-8-sig"))
|
||||
|
||||
|
||||
def _save_json(path: Path, payload: dict[str, Any]) -> None:
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
path.write_text(json.dumps(payload, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
|
||||
|
||||
|
||||
def _load_env(path: Path) -> dict[str, str]:
|
||||
"""Load key=value pairs from .env file."""
|
||||
env: dict[str, str] = {}
|
||||
if not path.exists():
|
||||
return env
|
||||
for line in path.read_text(encoding="utf-8").splitlines():
|
||||
line = line.strip()
|
||||
if not line or line.startswith("#") or "=" not in line:
|
||||
continue
|
||||
key, _, value = line.partition("=")
|
||||
env[key.strip()] = value.strip().strip("\"'")
|
||||
return env
|
||||
|
||||
|
||||
def _build_prompt(
|
||||
interest_keywords: list[str],
|
||||
rule_alias_suggestions: list[dict[str, str]],
|
||||
candidate_terms: list[dict[str, Any]],
|
||||
relevant_watch_terms: list[dict[str, Any]],
|
||||
) -> str:
|
||||
"""Build the LLM prompt for semantic suggestions."""
|
||||
|
||||
interest_bullets = "\n".join(f" - {t}" for t in sorted(interest_keywords))
|
||||
candidate_bullets = "\n".join(
|
||||
f" - {t['term']} (count={t['total_count']}, days={t['days_seen']})"
|
||||
for t in candidate_terms[:40]
|
||||
)
|
||||
|
||||
# Alias from rule layer (for LLM to build on, not duplicate)
|
||||
rule_alias_text = ""
|
||||
if rule_alias_suggestions:
|
||||
rule_alias_text = "\nSurface-form alias (already identified, skip these):\n" + "\n".join(
|
||||
f" {a['from']} → {a['to']} ({a['reason']})"
|
||||
for a in rule_alias_suggestions
|
||||
)
|
||||
|
||||
watch_text = ""
|
||||
if relevant_watch_terms:
|
||||
watch_text = "\nWatch terms (low-frequency but potentially relevant):\n" + "\n".join(
|
||||
f" {t['term']} (count={t['total_count']}, days={t['days_seen']})"
|
||||
for t in relevant_watch_terms[:20]
|
||||
)
|
||||
|
||||
return f"""You are a keyword governance assistant for an AI engineer. Your job is to analyze keyword data and produce structured suggestions.
|
||||
|
||||
## User's focus areas
|
||||
- AI Agent engineering (Skills, Harness, MCP, Agent architecture)
|
||||
- Backend engineering (Java, Go, Kubernetes, MySQL, distributed systems)
|
||||
- Open source AI tools and practices (Claude Code, Cursor, DeepSeek, OpenClaw)
|
||||
- LLM application engineering (context engineering, RAG, prompt engineering)
|
||||
|
||||
## Interest keywords (52 already configured)
|
||||
{interest_bullets}
|
||||
|
||||
## Uncovered candidate terms (sorted by frequency)
|
||||
{candidate_bullets}
|
||||
{watch_text}{rule_alias_text}
|
||||
|
||||
## Task
|
||||
Analyze the candidate terms and output a JSON object with exactly three keys:
|
||||
|
||||
1. "semantic_alias": array of alias suggestions that SURFACE RULES CAN'T CATCH (e.g. abbreviation↔full name, Chinese↔English, different naming for the same concept).
|
||||
Format: [{{"from": "<variant>", "to": "<canonical interest keyword>", "reason": "<why>"}}]
|
||||
|
||||
2. "stopword": array of terms that are too broad/generic to be useful as filters. A stopword is a term that appears frequently but has LOW DISCRIMINATION — it matches too many unrelated articles and clutters the keyword index.
|
||||
Format: [{{"term": "<term>", "reason": "<why it should be a stopword>"}}]
|
||||
|
||||
3. "promote_to_interest": array of uncovered terms that align well with the user's focus areas and should be added as interest keywords.
|
||||
Format: [{{"term": "<term>", "reason": "<why it fits>"}}]
|
||||
|
||||
## Rules
|
||||
- Be conservative. When in doubt, leave it out.
|
||||
- Only suggest alias for terms that clearly refer to the SAME concept as an existing interest keyword.
|
||||
- Only suggest stopword for terms that are genuinely too broad (appear in many unrelated contexts).
|
||||
- Only suggest promote for terms that clearly match the user's stated focus areas.
|
||||
- Output valid JSON only, no markdown, no explanation outside the JSON."""
|
||||
|
||||
|
||||
def _call_llm(prompt: str, api_url: str, model: str, api_key: str) -> str:
|
||||
"""Call LLM API and return the response text."""
|
||||
payload = json.dumps({
|
||||
"model": model,
|
||||
"messages": [{"role": "user", "content": prompt}],
|
||||
"temperature": 0.1,
|
||||
"max_tokens": 2048,
|
||||
}).encode("utf-8")
|
||||
|
||||
req = Request(
|
||||
api_url.rstrip("/") + "/chat/completions",
|
||||
data=payload,
|
||||
headers={
|
||||
"Content-Type": "application/json",
|
||||
"Authorization": f"Bearer {api_key}",
|
||||
},
|
||||
)
|
||||
|
||||
max_retries = 3
|
||||
for attempt in range(max_retries):
|
||||
try:
|
||||
with urlopen(req, timeout=120) as resp:
|
||||
result = json.loads(resp.read().decode("utf-8"))
|
||||
return result["choices"][0]["message"]["content"]
|
||||
except Exception as e:
|
||||
if attempt < max_retries - 1:
|
||||
wait = 2 ** attempt
|
||||
print(f" LLM call failed (attempt {attempt+1}/{max_retries}): {e}", file=sys.stderr)
|
||||
print(f" Retrying in {wait}s...", file=sys.stderr)
|
||||
time.sleep(wait)
|
||||
else:
|
||||
raise
|
||||
|
||||
|
||||
def _parse_llm_response(text: str) -> dict[str, list[dict[str, str]]]:
|
||||
"""Extract JSON from LLM response (may contain markdown fences)."""
|
||||
# Try to find JSON block
|
||||
json_match = re.search(r"```(?:json)?\s*\n?(\{.*?\})\s*\n?```", text, re.DOTALL)
|
||||
if json_match:
|
||||
text = json_match.group(1)
|
||||
|
||||
# Clean up: remove any text before { or after }
|
||||
start = text.find("{")
|
||||
end = text.rfind("}")
|
||||
if start >= 0 and end > start:
|
||||
text = text[start : end + 1]
|
||||
|
||||
try:
|
||||
result = json.loads(text)
|
||||
except json.JSONDecodeError:
|
||||
# Try partial recovery
|
||||
print(f" Warning: LLM response not clean JSON, attempting recovery", file=sys.stderr)
|
||||
print(f" Raw: {text[:500]}", file=sys.stderr)
|
||||
return {"semantic_alias": [], "stopword": [], "promote_to_interest": []}
|
||||
|
||||
# Normalize keys
|
||||
normalized = {
|
||||
"semantic_alias": result.get("semantic_alias", result.get("alias", [])),
|
||||
"stopword": result.get("stopword", result.get("stopword_suggestions", [])),
|
||||
"promote_to_interest": result.get("promote_to_interest", result.get("promote", [])),
|
||||
}
|
||||
# Ensure each is a list
|
||||
for key in normalized:
|
||||
if not isinstance(normalized[key], list):
|
||||
normalized[key] = []
|
||||
return normalized
|
||||
|
||||
|
||||
def main() -> None:
|
||||
parser = argparse.ArgumentParser(description="Generate semantic keyword suggestions via LLM.")
|
||||
parser.add_argument("--bundle", type=Path, default=DEFAULT_BUNDLE_PATH, help="Review bundle JSON path")
|
||||
parser.add_argument("--suggestions", type=Path, default=None, help="Existing suggestions JSON (for rule alias context)")
|
||||
parser.add_argument("--output", type=Path, default=None, help="Output JSON path (auto-generated if omitted)")
|
||||
parser.add_argument("--llm-api-url", type=str, default=None, help="LLM API base URL")
|
||||
parser.add_argument("--llm-model", type=str, default=None, help="LLM model name")
|
||||
parser.add_argument("--llm-api-key", type=str, default=None, help="LLM API key")
|
||||
parser.add_argument("--dry-run", action="store_true", help="Print prompt and exit without calling LLM")
|
||||
args = parser.parse_args()
|
||||
|
||||
# Load config
|
||||
env_path = REPO_ROOT / ".env"
|
||||
env = _load_env(env_path) if env_path.exists() else {}
|
||||
|
||||
api_url = args.llm_api_url or os.environ.get("LLM_API_URL") or env.get("LLM_API_URL", "https://api.deepseek.com")
|
||||
# Map OpenClaw model aliases to actual API model names
|
||||
model_raw = args.llm_model or os.environ.get("LLM_MODEL") or env.get("LLM_MODEL", "deepseek-chat")
|
||||
MODEL_ALIAS_MAP = {
|
||||
"deepseek/deepseek-v4-flash": "deepseek-chat",
|
||||
"deepseek/deepseek-chat": "deepseek-chat",
|
||||
"deepseek-v4-flash": "deepseek-chat",
|
||||
"deepseek-chat": "deepseek-chat",
|
||||
}
|
||||
model = MODEL_ALIAS_MAP.get(model_raw, model_raw)
|
||||
api_key = args.llm_api_key or os.environ.get("LLM_API_KEY") or env.get("LLM_API_KEY", "")
|
||||
|
||||
if not api_key:
|
||||
print("Error: No LLM API key found. Set LLM_API_KEY in .env or pass --llm-api-key.", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
# Load bundle
|
||||
if not args.bundle.exists():
|
||||
print(f"Error: Bundle not found: {args.bundle}", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
bundle = _load_json(args.bundle)
|
||||
current_config = bundle.get("current_config", {})
|
||||
interest_keywords = current_config.get("interest_keywords", [])
|
||||
top_global_terms = bundle.get("top_global_terms", [])
|
||||
governance_hints = bundle.get("governance_hints", {})
|
||||
|
||||
# Build candidate list (uncovered terms from interest + watch candidates)
|
||||
candidate_terms = []
|
||||
for item in governance_hints.get("interest_review_candidates", []):
|
||||
if isinstance(item, dict):
|
||||
candidate_terms.append({
|
||||
"term": item.get("term", ""),
|
||||
"total_count": item.get("total_count", 0),
|
||||
"days_seen": item.get("days_seen", 0),
|
||||
"percentile": item.get("percentile", 0),
|
||||
"growth": item.get("growth", 0),
|
||||
})
|
||||
for item in governance_hints.get("watch_review_candidates", []):
|
||||
if isinstance(item, dict):
|
||||
# Avoid duplicates
|
||||
if not any(c["term"] == item.get("term") for c in candidate_terms):
|
||||
candidate_terms.append({
|
||||
"term": item.get("term", ""),
|
||||
"total_count": item.get("total_count", 0),
|
||||
"days_seen": item.get("days_seen", 0),
|
||||
"percentile": item.get("percentile", 0),
|
||||
"growth": item.get("growth", 0),
|
||||
})
|
||||
|
||||
# Sort by total_count descending
|
||||
candidate_terms.sort(key=lambda x: -x["total_count"])
|
||||
relevant_watch_terms = governance_hints.get("watch_review_candidates", [])[:20]
|
||||
|
||||
# Load rule-layer alias suggestions if available
|
||||
rule_alias = []
|
||||
if args.suggestions and args.suggestions.exists():
|
||||
s = _load_json(args.suggestions)
|
||||
rule_alias = s.get("alias_suggestions", [])
|
||||
|
||||
# Build prompt
|
||||
prompt = _build_prompt(
|
||||
interest_keywords=interest_keywords,
|
||||
rule_alias_suggestions=rule_alias,
|
||||
candidate_terms=candidate_terms,
|
||||
relevant_watch_terms=relevant_watch_terms,
|
||||
)
|
||||
|
||||
# Determine output path
|
||||
suggestion_date = datetime.now(timezone.utc).date().isoformat()
|
||||
output_path = args.output or (DEFAULT_OUTPUT_DIR / f"term-cleanup-semantic-suggestions-{suggestion_date}.json")
|
||||
|
||||
if args.dry_run:
|
||||
print("=== DRY RUN: Prompt ===")
|
||||
print(prompt)
|
||||
print("\n=== END ===")
|
||||
print(f"\nWould write to: {output_path}")
|
||||
return
|
||||
|
||||
# Call LLM
|
||||
print(f"Calling LLM ({model})...", file=sys.stderr)
|
||||
response = _call_llm(prompt, api_url, model, api_key)
|
||||
print(f"LLM response received ({len(response)} chars)", file=sys.stderr)
|
||||
|
||||
# Parse
|
||||
parsed = _parse_llm_response(response)
|
||||
|
||||
# Build output
|
||||
output = {
|
||||
"date": suggestion_date,
|
||||
"source_bundle": str(args.bundle),
|
||||
"model": model,
|
||||
"interest_keyword_count": len(interest_keywords),
|
||||
"candidate_count": len(candidate_terms),
|
||||
**parsed,
|
||||
}
|
||||
|
||||
_save_json(output_path, output)
|
||||
|
||||
summary = {
|
||||
"output": str(output_path),
|
||||
"semantic_alias": len(output.get("semantic_alias", [])),
|
||||
"stopword": len(output.get("stopword", [])),
|
||||
"promote_to_interest": len(output.get("promote_to_interest", [])),
|
||||
}
|
||||
print(json.dumps(summary, ensure_ascii=False, indent=2))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,536 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
from datetime import datetime, timezone
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
DEFAULT_BUNDLE_PATH = REPO_ROOT / "outputs" / "term_index" / "review" / "keyword-cleanup-bundle.json"
|
||||
DEFAULT_OUTPUT_DIR = REPO_ROOT / "outputs" / "term_index" / "review"
|
||||
|
||||
|
||||
def _load_json(path: Path) -> Any:
|
||||
return json.loads(path.read_text(encoding="utf-8-sig"))
|
||||
|
||||
|
||||
def _save_json(path: Path, payload: dict[str, Any]) -> None:
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
path.write_text(json.dumps(payload, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
|
||||
|
||||
|
||||
def _save_text(path: Path, content: str) -> None:
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
path.write_text(content, encoding="utf-8")
|
||||
|
||||
|
||||
def _term_key(value: str) -> str:
|
||||
return value.strip().casefold()
|
||||
|
||||
|
||||
def _utc_today() -> str:
|
||||
return datetime.now(timezone.utc).date().isoformat()
|
||||
|
||||
|
||||
def _require_dict(payload: Any, name: str) -> dict[str, Any]:
|
||||
if not isinstance(payload, dict):
|
||||
raise RuntimeError(f"{name} must be a JSON object.")
|
||||
return payload
|
||||
|
||||
|
||||
def _require_list(payload: Any, name: str) -> list[Any]:
|
||||
if not isinstance(payload, list):
|
||||
raise RuntimeError(f"{name} must be a JSON array.")
|
||||
return payload
|
||||
|
||||
|
||||
def _bundle_date(bundle: dict[str, Any]) -> str:
|
||||
generated_at = bundle.get("generated_at")
|
||||
if isinstance(generated_at, str) and generated_at.strip():
|
||||
normalized = generated_at.replace("Z", "+00:00")
|
||||
try:
|
||||
return datetime.fromisoformat(normalized).date().isoformat()
|
||||
except ValueError:
|
||||
pass
|
||||
return _utc_today()
|
||||
|
||||
|
||||
def _recent_count_map(top_global_terms: list[dict[str, Any]]) -> dict[str, int]:
|
||||
counts: dict[str, int] = {}
|
||||
for item in top_global_terms:
|
||||
term = item.get("term")
|
||||
recent_count = item.get("recent_count")
|
||||
if isinstance(term, str) and isinstance(recent_count, int):
|
||||
counts[term] = recent_count
|
||||
return counts
|
||||
|
||||
|
||||
def _covered_term_sets(bundle: dict[str, Any]) -> tuple[set[str], set[str], set[str]]:
|
||||
current_config = _require_dict(bundle.get("current_config"), "bundle.current_config")
|
||||
interest_keywords = _require_list(current_config.get("interest_keywords"), "bundle.current_config.interest_keywords")
|
||||
stopwords = _require_list(current_config.get("stopwords"), "bundle.current_config.stopwords")
|
||||
watchlist = _require_list(current_config.get("watchlist"), "bundle.current_config.watchlist")
|
||||
|
||||
interest_set = {_term_key(item) for item in interest_keywords if isinstance(item, str) and item.strip()}
|
||||
stopword_set = {_term_key(item) for item in stopwords if isinstance(item, str) and item.strip()}
|
||||
watch_set = {
|
||||
_term_key(str(item.get("term", "")))
|
||||
for item in watchlist
|
||||
if isinstance(item, dict) and isinstance(item.get("term"), str) and str(item.get("term", "")).strip()
|
||||
}
|
||||
return interest_set, stopword_set, watch_set
|
||||
|
||||
|
||||
def _sort_key(item: dict[str, Any]) -> tuple[int, int, int, str, str]:
|
||||
total_count = int(item.get("total_count") or 0)
|
||||
days_seen = int(item.get("days_seen") or 0)
|
||||
recent_count = int(item.get("recent_count") or 0)
|
||||
term = str(item.get("term") or "")
|
||||
return (-total_count, -days_seen, -recent_count, term.casefold(), term)
|
||||
|
||||
|
||||
def _prepare_interest_suggestions(bundle: dict[str, Any]) -> list[dict[str, Any]]:
|
||||
governance_hints = _require_dict(bundle.get("governance_hints"), "bundle.governance_hints")
|
||||
candidates = _require_list(
|
||||
governance_hints.get("interest_review_candidates"),
|
||||
"bundle.governance_hints.interest_review_candidates",
|
||||
)
|
||||
top_global_terms = _require_list(bundle.get("top_global_terms"), "bundle.top_global_terms")
|
||||
recent_counts = _recent_count_map([item for item in top_global_terms if isinstance(item, dict)])
|
||||
interest_set, stopword_set, watch_set = _covered_term_sets(bundle)
|
||||
|
||||
suggestions: list[dict[str, Any]] = []
|
||||
seen: set[str] = set()
|
||||
for item in candidates:
|
||||
if not isinstance(item, dict):
|
||||
continue
|
||||
term = item.get("term")
|
||||
if not isinstance(term, str) or not term.strip():
|
||||
continue
|
||||
term_key = _term_key(term)
|
||||
if term_key in seen or term_key in interest_set or term_key in stopword_set:
|
||||
continue
|
||||
total_count = int(item.get("total_count") or 0)
|
||||
days_seen = int(item.get("days_seen") or 0)
|
||||
recent_count = recent_counts.get(term, 0)
|
||||
base_reason = str(item.get("reason") or "Meets the configured interest-keyword review threshold.")
|
||||
if term_key in watch_set:
|
||||
base_reason += " It is currently in watchlist and is ready for promotion."
|
||||
reason = f"{base_reason} Evidence: total_count={total_count}, days_seen={days_seen}, recent_count={recent_count}."
|
||||
suggestions.append(
|
||||
{
|
||||
"term": term,
|
||||
"reason": reason,
|
||||
"total_count": total_count,
|
||||
"days_seen": days_seen,
|
||||
"recent_count": recent_count,
|
||||
}
|
||||
)
|
||||
seen.add(term_key)
|
||||
|
||||
suggestions.sort(key=_sort_key)
|
||||
return suggestions
|
||||
|
||||
|
||||
def _prepare_watch_suggestions(bundle: dict[str, Any], reserved_terms: set[str]) -> list[dict[str, Any]]:
|
||||
governance_hints = _require_dict(bundle.get("governance_hints"), "bundle.governance_hints")
|
||||
candidates = _require_list(
|
||||
governance_hints.get("watch_review_candidates"),
|
||||
"bundle.governance_hints.watch_review_candidates",
|
||||
)
|
||||
top_global_terms = _require_list(bundle.get("top_global_terms"), "bundle.top_global_terms")
|
||||
recent_counts = _recent_count_map([item for item in top_global_terms if isinstance(item, dict)])
|
||||
interest_set, stopword_set, watch_set = _covered_term_sets(bundle)
|
||||
|
||||
suggestions: list[dict[str, Any]] = []
|
||||
seen: set[str] = set(reserved_terms)
|
||||
for item in candidates:
|
||||
if not isinstance(item, dict):
|
||||
continue
|
||||
term = item.get("term")
|
||||
if not isinstance(term, str) or not term.strip():
|
||||
continue
|
||||
term_key = _term_key(term)
|
||||
if term_key in seen or term_key in interest_set or term_key in stopword_set or term_key in watch_set:
|
||||
continue
|
||||
total_count = int(item.get("total_count") or 0)
|
||||
days_seen = int(item.get("days_seen") or 0)
|
||||
recent_count = recent_counts.get(term, 0)
|
||||
base_reason = str(item.get("reason") or "Falls into the configured watch-term review range.")
|
||||
reason = f"{base_reason} Evidence: total_count={total_count}, days_seen={days_seen}, recent_count={recent_count}."
|
||||
suggestions.append(
|
||||
{
|
||||
"term": term,
|
||||
"reason": reason,
|
||||
"total_count": total_count,
|
||||
"days_seen": days_seen,
|
||||
"recent_count": recent_count,
|
||||
}
|
||||
)
|
||||
seen.add(term_key)
|
||||
|
||||
suggestions.sort(key=_sort_key)
|
||||
return suggestions
|
||||
|
||||
|
||||
def _prepare_alias_suggestions(
|
||||
bundle: dict[str, Any],
|
||||
all_terms: list[dict[str, Any]] | None = None,
|
||||
) -> list[dict[str, Any]]:
|
||||
"""
|
||||
Generate alias suggestions using surface-form rules (no LLM).
|
||||
|
||||
Rules:
|
||||
1. casefold match — same normalized form, different original casing
|
||||
2. trailing-s singularization — singular/plural variants
|
||||
3. whitespace/hyphen normalization — word boundary variants
|
||||
|
||||
Scans all_terms (full term_stats) if provided; otherwise falls back
|
||||
to top_global_terms from the bundle.
|
||||
"""
|
||||
current_config = _require_dict(bundle.get("current_config"), "bundle.current_config")
|
||||
interest_keywords = _require_list(
|
||||
current_config.get("interest_keywords"), "bundle.current_config.interest_keywords"
|
||||
)
|
||||
top_global_terms = _require_list(bundle.get("top_global_terms"), "bundle.top_global_terms")
|
||||
source_terms = all_terms if all_terms is not None else top_global_terms
|
||||
|
||||
interest_set = {_term_key(t) for t in interest_keywords if isinstance(t, str)}
|
||||
interest_originals: set[str] = {t for t in interest_keywords if isinstance(t, str)}
|
||||
|
||||
# Build full casefold → [original forms] map
|
||||
cf_map: dict[str, list[str]] = {}
|
||||
for item in source_terms:
|
||||
term = None
|
||||
if isinstance(item, dict):
|
||||
term = item.get("term")
|
||||
elif isinstance(item, str):
|
||||
term = item
|
||||
if not isinstance(term, str) or not term.strip():
|
||||
continue
|
||||
key = _term_key(term)
|
||||
if key not in cf_map:
|
||||
cf_map[key] = []
|
||||
if term not in cf_map[key]:
|
||||
cf_map[key].append(term)
|
||||
|
||||
suggestions: list[dict[str, Any]] = []
|
||||
seen_pairs: set[tuple[str, str]] = set()
|
||||
|
||||
def _add(from_term: str, to_term: str, reason: str) -> None:
|
||||
pair = (_term_key(from_term), _term_key(to_term))
|
||||
if pair in seen_pairs:
|
||||
return
|
||||
seen_pairs.add(pair)
|
||||
suggestions.append({"from": from_term, "to": to_term, "reason": reason})
|
||||
|
||||
# Build a set of all term keys from source for quick lookup
|
||||
source_keys = set(cf_map.keys())
|
||||
|
||||
# Rule 1: casefold match — same normalized form, different casing
|
||||
for key, variants in cf_map.items():
|
||||
if len(variants) < 2:
|
||||
continue
|
||||
canonical = None
|
||||
alt_forms = []
|
||||
for v in variants:
|
||||
if v in interest_originals:
|
||||
canonical = v
|
||||
else:
|
||||
alt_forms.append(v)
|
||||
if canonical and alt_forms:
|
||||
for alt in alt_forms:
|
||||
_add(alt, canonical, "Case variant")
|
||||
elif len(variants) >= 2 and not canonical:
|
||||
# None is canonical — suggest the highest-frequency form
|
||||
ranked = sorted(variants, key=lambda t: -(
|
||||
next(
|
||||
(it.get("total_count", 0) for it in top_global_terms if it.get("term") == t),
|
||||
0,
|
||||
)
|
||||
))
|
||||
for alt in ranked[1:]:
|
||||
_add(alt, ranked[0], "Case variant (auto-ranked)")
|
||||
|
||||
# Rule 2: singular/plural — trailing-s normalization
|
||||
# Check all source terms (not just interest keys) for bidirectional matching
|
||||
for key in source_keys:
|
||||
if key in interest_set:
|
||||
continue
|
||||
if key.endswith("s") and len(key) > 2:
|
||||
singular_key = key.rstrip("s")
|
||||
if singular_key in interest_set and singular_key != key:
|
||||
# Find canonical interest keyword
|
||||
canon = next((t for t in interest_keywords if _term_key(t) == singular_key), None)
|
||||
from_form = cf_map[key][0]
|
||||
if canon:
|
||||
_add(from_form, canon, "Plural variant")
|
||||
# singular form → interest has plural
|
||||
plural_key = key + "s"
|
||||
if plural_key in interest_set and plural_key != key:
|
||||
canon = next((t for t in interest_keywords if _term_key(t) == plural_key), None)
|
||||
from_form = cf_map[key][0]
|
||||
if canon:
|
||||
_add(from_form, canon, "Singular variant")
|
||||
|
||||
# Rule 3: whitespace/hyphen normalization
|
||||
for key in source_keys:
|
||||
if key in interest_set:
|
||||
continue
|
||||
normalized = key.replace("-", "").replace("_", "").replace(" ", "")
|
||||
if normalized in interest_set and normalized != key:
|
||||
canon = next((t for t in interest_keywords if _term_key(t) == normalized), None)
|
||||
from_form = cf_map[key][0]
|
||||
if canon:
|
||||
_add(from_form, canon, "Whitespace/punctuation variant")
|
||||
|
||||
suggestions.sort(key=lambda x: (x["from"].casefold(), x["to"].casefold()))
|
||||
return suggestions
|
||||
|
||||
|
||||
def _render_table(items: list[dict[str, Any]]) -> str:
|
||||
if not items:
|
||||
return "_None in this pass._\n"
|
||||
lines = [
|
||||
"| Term | Total | Days | Recent | Reason |",
|
||||
"| --- | ---: | ---: | ---: | --- |",
|
||||
]
|
||||
for item in items:
|
||||
term = str(item.get("term") or "")
|
||||
total_count = int(item.get("total_count") or 0)
|
||||
days_seen = int(item.get("days_seen") or 0)
|
||||
recent_count = int(item.get("recent_count") or 0)
|
||||
reason = str(item.get("reason") or "").replace("|", "\\|")
|
||||
lines.append(f"| {term} | {total_count} | {days_seen} | {recent_count} | {reason} |")
|
||||
return "\n".join(lines) + "\n"
|
||||
|
||||
|
||||
def _render_simple_table(items: list[dict[str, Any]], first_column: str) -> str:
|
||||
if not items:
|
||||
return "_None in this pass._\n"
|
||||
lines = [
|
||||
f"| {first_column} | Reason |",
|
||||
"| --- | --- |",
|
||||
]
|
||||
for item in items:
|
||||
value = str(item.get(first_column.casefold()) or item.get(first_column) or "")
|
||||
reason = str(item.get("reason") or "").replace("|", "\\|")
|
||||
lines.append(f"| {value} | {reason} |")
|
||||
return "\n".join(lines) + "\n"
|
||||
|
||||
|
||||
def _render_markdown(
|
||||
*,
|
||||
suggestion_date: str,
|
||||
bundle_path: Path,
|
||||
json_output_path: Path,
|
||||
bundle: dict[str, Any],
|
||||
suggestions: dict[str, Any],
|
||||
) -> str:
|
||||
policy = _require_dict(bundle.get("policy"), "bundle.policy")
|
||||
current_config = _require_dict(bundle.get("current_config"), "bundle.current_config")
|
||||
top_global_terms = _require_list(bundle.get("top_global_terms"), "bundle.top_global_terms")
|
||||
uncovered_terms = _require_list(bundle.get("uncovered_terms"), "bundle.uncovered_terms")
|
||||
top_preview = [item for item in top_global_terms if isinstance(item, dict)][:5]
|
||||
uncovered_preview = [item for item in uncovered_terms if isinstance(item, dict)][:5]
|
||||
|
||||
interest_items = suggestions["interest_keyword_suggestions"]
|
||||
watch_items = suggestions["watch_terms"]
|
||||
alias_items = suggestions["alias_suggestions"]
|
||||
stopword_items = suggestions["stopword_suggestions"]
|
||||
|
||||
lines = [
|
||||
f"# Term Cleanup Suggestions - {suggestion_date}",
|
||||
"",
|
||||
"## Review Context",
|
||||
"",
|
||||
f"- Source bundle: `{bundle_path}`",
|
||||
f"- Suggestions JSON: `{json_output_path}`",
|
||||
f"- Bundle generated_at: `{bundle.get('generated_at', 'unknown')}`",
|
||||
f"- Based on days: `{suggestions['based_on_days']}`",
|
||||
f"- Policy schema version: `{policy.get('schema_version', 'unknown')}`",
|
||||
"- Scope: implement `interest_keyword_suggestions` and `watch_terms` main path first; keep alias/stopword conservative in this pass.",
|
||||
"",
|
||||
"## Current State",
|
||||
"",
|
||||
f"- Interest keywords: `{current_config.get('interest_keyword_count', 0)}`",
|
||||
f"- Watch terms: `{current_config.get('watch_term_count', 0)}`",
|
||||
f"- Stopwords: `{current_config.get('stopword_count', 0)}`",
|
||||
f"- Aliases: `{current_config.get('alias_count', 0)}`",
|
||||
f"- Top global terms considered: `{len(top_global_terms)}`",
|
||||
f"- Uncovered terms considered: `{len(uncovered_terms)}`",
|
||||
"",
|
||||
"### Top Terms Snapshot",
|
||||
"",
|
||||
]
|
||||
|
||||
if top_preview:
|
||||
for item in top_preview:
|
||||
lines.append(
|
||||
f"- `{item.get('term', '')}`: total_count={item.get('total_count', 0)}, days_seen={item.get('days_seen', 0)}, recent_count={item.get('recent_count', 0)}"
|
||||
)
|
||||
else:
|
||||
lines.append("- No top terms available.")
|
||||
|
||||
lines.extend([
|
||||
"",
|
||||
"### Uncovered Terms Snapshot",
|
||||
"",
|
||||
])
|
||||
if uncovered_preview:
|
||||
for item in uncovered_preview:
|
||||
lines.append(
|
||||
f"- `{item.get('term', '')}`: total_count={item.get('total_count', 0)}, days_seen={item.get('days_seen', 0)}, recent_count={item.get('recent_count', 0)}"
|
||||
)
|
||||
else:
|
||||
lines.append("- No uncovered terms available.")
|
||||
|
||||
lines.extend([
|
||||
"",
|
||||
"## Suggestion Summary",
|
||||
"",
|
||||
f"- `interest_keyword_suggestions`: `{len(interest_items)}`",
|
||||
f"- `watch_terms`: `{len(watch_items)}`",
|
||||
f"- `alias_suggestions`: `{len(alias_items)}`",
|
||||
f"- `stopword_suggestions`: `{len(stopword_items)}`",
|
||||
"",
|
||||
"## Interest Keyword Suggestions",
|
||||
"",
|
||||
_render_table(interest_items).rstrip(),
|
||||
"",
|
||||
"## Watch Terms",
|
||||
"",
|
||||
_render_table(watch_items).rstrip(),
|
||||
"",
|
||||
"## Alias Suggestions",
|
||||
"",
|
||||
"_Conservative by design in this minimal version; no automatic alias suggestions are emitted yet._" if not alias_items else _render_simple_table(alias_items, "from").rstrip(),
|
||||
"",
|
||||
"## Stopword Suggestions",
|
||||
"",
|
||||
"_Conservative by design in this minimal version; no automatic stopword suggestions are emitted yet._" if not stopword_items else _render_simple_table(stopword_items, "term").rstrip(),
|
||||
"",
|
||||
"## Apply",
|
||||
"",
|
||||
"Review the Markdown first, then selectively apply accepted suggestions with the JSON file.",
|
||||
"",
|
||||
"```bash",
|
||||
f"python scripts/apply_term_suggestions.py \\",
|
||||
f" --suggestions {json_output_path} \\",
|
||||
" --accept-interest \"Claude Code\" \\",
|
||||
" --accept-watch \"A2A\" \\",
|
||||
" --dry-run",
|
||||
"```",
|
||||
"",
|
||||
])
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def _build_output_paths(
|
||||
*,
|
||||
output_dir: Path,
|
||||
suggestion_date: str,
|
||||
json_output: Path | None,
|
||||
markdown_output: Path | None,
|
||||
) -> tuple[Path, Path]:
|
||||
stem = f"term-cleanup-suggestions-{suggestion_date}"
|
||||
resolved_json = json_output or (output_dir / f"{stem}.json")
|
||||
resolved_markdown = markdown_output or (output_dir / f"{stem}.md")
|
||||
return resolved_json, resolved_markdown
|
||||
|
||||
|
||||
def main() -> None:
|
||||
parser = argparse.ArgumentParser(description="Generate term cleanup suggestions JSON and Markdown from review bundle.")
|
||||
parser.add_argument("--bundle", type=Path, default=DEFAULT_BUNDLE_PATH, help="Review bundle JSON file")
|
||||
parser.add_argument(
|
||||
"--output-dir",
|
||||
type=Path,
|
||||
default=DEFAULT_OUTPUT_DIR,
|
||||
help="Directory for generated suggestions outputs when explicit output paths are not provided",
|
||||
)
|
||||
parser.add_argument("--date", type=str, default=None, help="Override suggestions date (YYYY-MM-DD)")
|
||||
parser.add_argument("--json-output", type=Path, default=None, help="Explicit suggestions JSON output path")
|
||||
parser.add_argument("--markdown-output", type=Path, default=None, help="Explicit suggestions Markdown output path")
|
||||
parser.add_argument(
|
||||
"--emit-markdown",
|
||||
action="store_true",
|
||||
help="Also write the human-readable Markdown review draft. JSON suggestions are always written.",
|
||||
)
|
||||
args = parser.parse_args()
|
||||
|
||||
if not args.bundle.exists():
|
||||
raise RuntimeError(f"Bundle file not found: {args.bundle}")
|
||||
|
||||
bundle = _require_dict(_load_json(args.bundle), "bundle")
|
||||
days = bundle.get("days")
|
||||
if not isinstance(days, int):
|
||||
raise RuntimeError("bundle.days must be an integer.")
|
||||
|
||||
suggestion_date = args.date or _bundle_date(bundle)
|
||||
json_output_path, markdown_output_path = _build_output_paths(
|
||||
output_dir=args.output_dir,
|
||||
suggestion_date=suggestion_date,
|
||||
json_output=args.json_output,
|
||||
markdown_output=args.markdown_output,
|
||||
)
|
||||
|
||||
# Load full term_stats for alias scanning (bundle only has top N)
|
||||
stats_path = REPO_ROOT / "data" / "term_index" / "term_stats.json"
|
||||
all_stats_terms: list[str] = []
|
||||
if stats_path.exists():
|
||||
stats_payload = _load_json(stats_path)
|
||||
raw_terms = stats_payload.get("terms") if isinstance(stats_payload, dict) else []
|
||||
if isinstance(raw_terms, list):
|
||||
all_stats_terms = [str(t["term"]) for t in raw_terms if isinstance(t, dict) and isinstance(t.get("term"), str)]
|
||||
|
||||
interest_items = _prepare_interest_suggestions(bundle)
|
||||
reserved_terms = {_term_key(str(item.get("term") or "")) for item in interest_items}
|
||||
watch_items = _prepare_watch_suggestions(bundle, reserved_terms=reserved_terms)
|
||||
alias_items = _prepare_alias_suggestions(bundle, all_terms=all_stats_terms)
|
||||
|
||||
suggestions = {
|
||||
"date": suggestion_date,
|
||||
"based_on_days": days,
|
||||
"source_bundle": str(args.bundle),
|
||||
"policy_schema_version": _require_dict(bundle.get("policy"), "bundle.policy").get("schema_version", "unknown"),
|
||||
"summary": {
|
||||
"interest_keyword_suggestions": len(interest_items),
|
||||
"watch_terms": len(watch_items),
|
||||
"alias_suggestions": len(alias_items),
|
||||
"stopword_suggestions": 0,
|
||||
},
|
||||
"alias_suggestions": alias_items,
|
||||
"stopword_suggestions": [],
|
||||
"interest_keyword_suggestions": interest_items,
|
||||
"watch_terms": watch_items,
|
||||
}
|
||||
markdown = _render_markdown(
|
||||
suggestion_date=suggestion_date,
|
||||
bundle_path=args.bundle,
|
||||
json_output_path=json_output_path,
|
||||
bundle=bundle,
|
||||
suggestions=suggestions,
|
||||
)
|
||||
|
||||
_save_json(json_output_path, suggestions)
|
||||
if args.emit_markdown:
|
||||
_save_text(markdown_output_path, markdown)
|
||||
|
||||
summary = {
|
||||
"bundle": str(args.bundle),
|
||||
"date": suggestion_date,
|
||||
"json_output": str(json_output_path),
|
||||
"markdown_output": str(markdown_output_path) if args.emit_markdown else None,
|
||||
"interest_keyword_suggestions": len(interest_items),
|
||||
"watch_terms": len(watch_items),
|
||||
"alias_suggestions": len(alias_items),
|
||||
"stopword_suggestions": 0,
|
||||
"emit_markdown": args.emit_markdown,
|
||||
}
|
||||
print(json.dumps(summary, ensure_ascii=False, indent=2))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -37,7 +37,7 @@ def main() -> None:
|
||||
help="Directory to write per-article Markdown summaries",
|
||||
)
|
||||
parser.add_argument("--max-retries", type=int, default=2, help="Maximum LLM retry attempts per article")
|
||||
parser.add_argument("--timeout", type=float, default=60.0, help="LLM request timeout in seconds")
|
||||
parser.add_argument("--timeout", type=float, default=120.0, help="LLM request timeout in seconds")
|
||||
parser.add_argument("--api-key", type=str, default=None, help="Override article-summary LLM API key")
|
||||
parser.add_argument("--model", type=str, default=None, help="Override article-summary LLM model")
|
||||
parser.add_argument("--api-url", type=str, default=None, help="Override article-summary LLM API URL/base URL")
|
||||
|
||||
@@ -0,0 +1,24 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
SRC_ROOT = REPO_ROOT / "src"
|
||||
|
||||
if str(SRC_ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(SRC_ROOT))
|
||||
|
||||
from summary_mcp.runtime.article_summary_jobs import run_article_summary_job
|
||||
|
||||
|
||||
def main() -> None:
|
||||
parser = argparse.ArgumentParser(description="Run a background article-summary job by job_id.")
|
||||
parser.add_argument("--job-id", required=True, help="Article summary job id")
|
||||
args = parser.parse_args()
|
||||
run_article_summary_job(job_id=args.job_id)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,24 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
SRC_ROOT = REPO_ROOT / "src"
|
||||
|
||||
if str(SRC_ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(SRC_ROOT))
|
||||
|
||||
from summary_mcp.runtime.freshrss_pipeline_jobs import run_freshrss_pipeline_job
|
||||
|
||||
|
||||
def main() -> None:
|
||||
parser = argparse.ArgumentParser(description="Run a background FreshRSS pipeline job by job_id.")
|
||||
parser.add_argument("--job-id", required=True, help="FreshRSS pipeline job id")
|
||||
args = parser.parse_args()
|
||||
run_freshrss_pipeline_job(job_id=args.job_id)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,24 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
SRC_ROOT = REPO_ROOT / "src"
|
||||
|
||||
if str(SRC_ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(SRC_ROOT))
|
||||
|
||||
from summary_mcp.runtime.resume_jobs import run_resume_job
|
||||
|
||||
|
||||
def main() -> None:
|
||||
parser = argparse.ArgumentParser(description="Run a background resume job by job_id.")
|
||||
parser.add_argument("--job-id", required=True, help="Resume job id")
|
||||
args = parser.parse_args()
|
||||
run_resume_job(job_id=args.job_id)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,113 @@
|
||||
---
|
||||
name: article-deep-summary
|
||||
description: 从单篇提取的文章中生成深度结构化知识笔记。当用户需要对存储在 *.extracted.json 文件中的文章 plain_text 内容进行深度摘要、提炼或创建知识笔记时使用。
|
||||
---
|
||||
|
||||
# 文章深度摘要
|
||||
|
||||
当用户拥有已提取的文章文件并希望生成深度知识笔记时使用此技能——不是简短的摘要卡片,而是包含核心结论、论点、方法、细节和可复用洞见的结构化提炼。
|
||||
|
||||
此技能**不适用于**验证或修复已有摘要(请使用 `llm-summary-review`),也不适用于关键词索引清理(请使用 `keyword-cleanup-review`)。
|
||||
|
||||
## 功能说明
|
||||
|
||||
- 读取包含 `plain_text` 的已提取文章 JSON 文件
|
||||
- 使用专用的文章摘要提示词调用 LLM
|
||||
- 根据 `ArticleSummaryResult` 模式验证 LLM 输出
|
||||
- 渲染结构化的中文 Markdown 知识笔记
|
||||
|
||||
## 输入
|
||||
|
||||
典型文件:
|
||||
|
||||
- 已提取的文章 JSON:`outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json`(单篇)或包含 `results` 数组的批量文件
|
||||
- 提示词模板:`outputs/prompts/article-summary-prompt.txt`
|
||||
|
||||
## 工作流程
|
||||
|
||||
1. 确定目标提取文件和需要摘要的条目 ID。
|
||||
|
||||
2. 通过 CLI 运行文章摘要工作流:
|
||||
|
||||
```bash
|
||||
python -m summary_mcp.workflows.article_summary \
|
||||
--extracted <extracted_path> \
|
||||
--selected-ids <item_id> \
|
||||
--output-dir <output_dir>
|
||||
```
|
||||
|
||||
或调用 MCP 工具 `generate_article_summaries`,参数如下:
|
||||
|
||||
- `extracted_path`:已提取的 JSON 文件路径
|
||||
- `selected_ids`:条目 ID 字符串数组(传入空数组可摘要全部条目)
|
||||
- `output_dir`:(可选)Markdown 输出目录
|
||||
|
||||
3. 工作流内部流程:
|
||||
- 解析 LLM 配置(`ARTICLE_SUMMARY_*` 环境变量,未设置时回退到 `LLM_*`)
|
||||
- 使用文章摘要提示词和验证器调用 `run_loop_payload`
|
||||
- 验证失败时最多重试 2 次
|
||||
|
||||
4. 在指定的输出目录中检查生成的 Markdown 文件。
|
||||
|
||||
## 输出模式
|
||||
|
||||
LLM 返回匹配 `ArticleSummaryResult` 的 JSON:
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
|---|---|---|
|
||||
| `title` | string | 文章标题 |
|
||||
| `url` | HttpUrl | 文章 URL |
|
||||
| `core_conclusion` | string | 作者核心结论,1-2 句话 |
|
||||
| `main_argument` | string | 主要论点或论题,可为多句 |
|
||||
| `key_methods` | string[] | 关键方法、机制或技术 |
|
||||
| `important_details` | string[] | 值得注意的细节、数据点或案例 |
|
||||
| `reusable_insights` | string[] | 可迁移到其他场景的可复用洞见 |
|
||||
| `keywords` | string[] | 具体实体——工具名称、框架、方法 |
|
||||
| `topics` | string[] | 更高层次的主题标签 |
|
||||
| `category` | enum | 取值之一:`资讯` `方法论` `工具实践` `观点评论` |
|
||||
| `worth_keeping` | bool | 该文章是否值得长期保留 |
|
||||
| `reason` | string | 一句话说明保留理由 |
|
||||
|
||||
验证器强制约束:
|
||||
|
||||
- `keywords` 和 `topics` 不得重叠
|
||||
- `keywords` 聚焦具体实体;`topics` 聚焦抽象主题
|
||||
- `category` 必须为四个允许值之一
|
||||
|
||||
## Markdown 输出
|
||||
|
||||
每篇文章生成一个 `.md` 文件,包含以下章节:
|
||||
|
||||
- 核心结论
|
||||
- 主要论点
|
||||
- 关键方法 / 机制
|
||||
- 重要细节
|
||||
- 可复用启发
|
||||
- 关键词
|
||||
- 主题
|
||||
|
||||
## LLM 配置
|
||||
|
||||
专用环境变量(未设置时回退到主 `LLM_*` 变量):
|
||||
|
||||
- `ARTICLE_SUMMARY_LLM_API_URL`
|
||||
- `ARTICLE_SUMMARY_LLM_API_KEY`
|
||||
- `ARTICLE_SUMMARY_LLM_MODEL`
|
||||
|
||||
## 代码实现
|
||||
|
||||
相关代码:
|
||||
|
||||
- 工作流:`src/summary_mcp/workflows/article_summary.py`
|
||||
- 验证器:`src/summary_mcp/validators/article_summary.py`
|
||||
- 数据模型:`src/summary_mcp/models/article_summary_result.py`
|
||||
- 提示词:`outputs/prompts/article-summary-prompt.txt`
|
||||
- MCP 工具:`generate_article_summaries`,位于 `src/summary_mcp/server.py`
|
||||
|
||||
## 何时停止
|
||||
|
||||
满足以下条件之一时停止:
|
||||
|
||||
- 所有请求条目的 Markdown 知识笔记已生成
|
||||
- 工作流报告反复验证失败,应由用户决定后续操作
|
||||
- 用户要求手动检查中间结果
|
||||
@@ -0,0 +1,3 @@
|
||||
display_name: Article Deep Summary
|
||||
short_description: Generate deep knowledge notes from a single extracted article.
|
||||
default_prompt: Given an extracted article JSON with plain_text content, run the article-summary workflow to produce a structured Markdown knowledge note with core conclusion, main arguments, key methods, important details, and reusable insights.
|
||||
@@ -1,62 +1,143 @@
|
||||
---
|
||||
name: keyword-cleanup-review
|
||||
description: Review and curate this repository's daily keyword index and frequency stats. Use when the user wants to inspect `data/term_index/term_stats.json`, recent `data/term_index/daily/*.json`, `configs/term_aliases.json`, `configs/term_stopwords.json`, or `configs/filter_context.personal.json` to propose alias merges, stopwords, watch terms, or `interest_keywords` updates without directly modifying configs.
|
||||
description: 生成 reader 项目的正式关键词 review 输入。当用户需要基于 `data/term_index/term_stats.json`、最近的 `data/term_index/daily/*.json` 和当前 configs 产出 bundle、正式 suggestions JSON,或按需生成审阅 Markdown 供 OpenClaw 汇报和等待确认时使用。不要用于低频清理 review 目录、删除旧产物或直接 apply 配置。
|
||||
---
|
||||
|
||||
# Keyword Cleanup Review
|
||||
# 关键词清理审查
|
||||
|
||||
Use this skill to turn the repository's keyword statistics into reviewable cleanup suggestions.
|
||||
使用此技能将仓库的关键词统计转化为**正式 review 输入**,供 OpenClaw 后续做汇报、确认和 apply 编排。
|
||||
|
||||
## Workflow
|
||||
## 角色边界
|
||||
|
||||
1. Build a compact review bundle:
|
||||
这个 skill 负责:
|
||||
|
||||
- 构建 review bundle
|
||||
- 生成正式 suggestions JSON
|
||||
- 按需生成人工审阅 Markdown
|
||||
- 给 OpenClaw 提供稳定的关键词 review 输入
|
||||
|
||||
这个 skill 不负责:
|
||||
|
||||
- 清理 `outputs/term_index/review/` 下的旧文件
|
||||
- 决定删除哪些历史 bundle / suggestions / markdown
|
||||
- 直接 apply `configs/term_aliases.json` / `configs/term_stopwords.json` / `configs/filter_context.personal.json`
|
||||
|
||||
低频维护、清理和 dry-run 校验应由 OpenClaw 侧 maintenance SOP 处理,而不是由本 skill 承担。
|
||||
|
||||
## 工作流程
|
||||
|
||||
### Phase 1:构建审查数据包
|
||||
|
||||
```bash
|
||||
python skills/keyword-cleanup-review/scripts/build_review_bundle.py
|
||||
```
|
||||
|
||||
Optional knobs:
|
||||
可选参数:
|
||||
|
||||
- `--days 7`
|
||||
- `--top 50`
|
||||
- `--days 7`(默认 7,建议传 365 覆盖全量)
|
||||
- `--top 100`(考虑的词数)
|
||||
- `--output outputs/term_index/review/keyword-cleanup-bundle.json`
|
||||
|
||||
2. Read the generated bundle and the suggestion schema:
|
||||
#### 候选引擎策略
|
||||
|
||||
- `outputs/term_index/review/keyword-cleanup-bundle.json`
|
||||
- `skills/keyword-cleanup-review/references/suggestion-schema.md`
|
||||
根据 `configs/term_cleanup_policy.json` 的 `schema_version` 自动切换:
|
||||
|
||||
3. Produce two outputs:
|
||||
| 版本 | 策略 | 说明 |
|
||||
|------|------|------|
|
||||
| v1(旧) | 固定阈值(total≥3/days≥2 → interest) | 小数据集兼容 |
|
||||
| v2(当前默认) | 百分位排名 + 增速因子 | 自适应数据量,不需要手工调阈值 |
|
||||
|
||||
- A short Markdown review for humans
|
||||
- A JSON suggestion file matching the schema
|
||||
v2 策略说明:
|
||||
- **percentile**:total_count 在所有词里的排位占比。top 5% → interest 候选,5%-20% → watch 候选
|
||||
- **growth**:recent_count / total_count,衡量近期活跃度。growth≥0.5 的排位外词也会主动推荐
|
||||
|
||||
4. Keep the boundary strict:
|
||||
### Phase 2:生成建议(规则层)
|
||||
|
||||
- Suggest changes to `configs/term_aliases.json`
|
||||
- Suggest changes to `configs/term_stopwords.json`
|
||||
- Suggest additions to `configs/filter_context.personal.json`
|
||||
- Do not directly edit these files unless the user explicitly asks
|
||||
- Do not suggest direct edits to `configs/filter_rules.json` unless the user asks for rule logic changes
|
||||
```bash
|
||||
python scripts/generate_term_cleanup_suggestions.py \
|
||||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json
|
||||
```
|
||||
|
||||
## Review Heuristics
|
||||
如需人工审阅展示稿:
|
||||
|
||||
Prioritize these decisions:
|
||||
```bash
|
||||
python scripts/generate_term_cleanup_suggestions.py \
|
||||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
|
||||
--emit-markdown
|
||||
```
|
||||
|
||||
- Alias suggestion
|
||||
- Same concept with different naming, casing, abbreviation, or Chinese/English variants
|
||||
- Stopword suggestion
|
||||
- Too generic, too broad, or too noisy to help filtering
|
||||
- Interest keyword suggestion
|
||||
- High-frequency and aligned with the user's backend engineering, AI-agent, and frontier-tech focus
|
||||
- Watch term
|
||||
- Recent and potentially important, but evidence is still weak
|
||||
#### 产出能力
|
||||
|
||||
Prefer conservative suggestions. If confidence is low, put the term into `watch_terms`.
|
||||
| 建议类型 | 状态 | 方法 |
|
||||
|---------|------|------|
|
||||
| interest 建议 | ✅ 已实现 | 百分位 top 5% + 增速促活 |
|
||||
| watch 建议 | ✅ 已实现 | 百分位 5%-20% |
|
||||
| alias 建议 | ✅ 已实现 | 规则层:大小写归一、单复数、去空格/连字符 |
|
||||
| stopword 建议 | ❌ 规则层空缺 | 见 Phase 3(LLM 层) |
|
||||
|
||||
## Inputs
|
||||
默认生成:
|
||||
- `term-cleanup-suggestions-YYYY-MM-DD.json`(正式建议产物)
|
||||
|
||||
Primary inputs:
|
||||
显式加 `--emit-markdown` 额外生成:
|
||||
- `term-cleanup-suggestions-YYYY-MM-DD.md`(临时展示稿)
|
||||
|
||||
### Phase 3:生成建议(LLM 层,可选)
|
||||
|
||||
规则层覆盖不了 alias(中英文对应、缩写展开、同义不同名)和 stopword 判断,需要 LLM 辅助:
|
||||
|
||||
```bash
|
||||
python scripts/generate_term_cleanup_semantic_suggestions.py \
|
||||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
|
||||
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
|
||||
--output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json
|
||||
```
|
||||
|
||||
从 `.env` 读取 LLM 配置(`LLM_API_URL` / `LLM_MODEL` / `LLM_API_KEY`),使用 DeepSeek API。
|
||||
|
||||
输出三部分:
|
||||
|
||||
| 输出 | 说明 |
|
||||
|------|------|
|
||||
| `semantic_alias` | 语义级别名(中英文、缩写、同义不同名) |
|
||||
| `stopword` | 泛词过滤建议(规则层做不了的需要语义判断的) |
|
||||
| `promote_to_interest` | 与用户关注方向一致的新词,建议加入 interest |
|
||||
|
||||
**注:LLM 层产物是候选,不应自动 apply,需要人工确认后由 OpenClaw 编排 apply。**
|
||||
|
||||
### Phase 4:输出给 OpenClaw 编排
|
||||
|
||||
- `suggestions JSON` = review / apply 之间唯一正式建议输入
|
||||
- `semantic-suggestions JSON` = LLM 补充建议,需要人工筛选后合并到 suggestions JSON 再 apply
|
||||
- Markdown = 临时展示层
|
||||
- 后续汇报、确认、dry-run、apply、收尾清理由 OpenClaw 编排层执行
|
||||
|
||||
### Phase 5:严格保持边界
|
||||
|
||||
- 建议 `configs/term_aliases.json` 的修改
|
||||
- 建议 `configs/term_stopwords.json` 的修改
|
||||
- 建议 `configs/filter_context.personal.json` 的新增
|
||||
- **LLM 层产出(semantic-suggestions)不自动 apply**,需人工确认后由 OpenClaw 编排层执行
|
||||
- 除非用户明确要求,否则不要直接编辑这些文件
|
||||
- 除非用户要求修改规则逻辑,否则不要建议直接编辑 `configs/filter_rules.json`
|
||||
|
||||
## 审查启发式规则
|
||||
|
||||
优先考虑以下决策:
|
||||
|
||||
- 别名建议
|
||||
- 同一概念的不同命名、大小写、缩写或中英文变体
|
||||
- 停用词建议
|
||||
- 过于通用、过于宽泛或噪声过大,对过滤无帮助
|
||||
- 兴趣关键词建议
|
||||
- 高频且与用户的后端工程、AI Agent 和前沿技术关注方向一致
|
||||
- 关注词
|
||||
- 近期出现且可能重要,但证据尚不充分
|
||||
|
||||
建议保守为主。如果置信度较低,将词放入 `watch_terms`。
|
||||
|
||||
## 输入
|
||||
|
||||
主要输入:
|
||||
|
||||
- `data/term_index/term_stats.json`
|
||||
- `data/term_index/daily/*.json`
|
||||
@@ -67,36 +148,76 @@ Primary inputs:
|
||||
- `configs/term_watchlist.json`
|
||||
- `configs/term_change_log.json`
|
||||
|
||||
The bundled script already compacts these into a single review bundle.
|
||||
打包脚本已将这些内容压缩为单个审查数据包。
|
||||
|
||||
## Output Expectations
|
||||
## 输出要求
|
||||
|
||||
The Markdown output should:
|
||||
Markdown 输出应:
|
||||
|
||||
- Summarize the current state briefly
|
||||
- List the top terms worth acting on
|
||||
- Separate alias, stopword, interest-keyword, and watch-term recommendations
|
||||
- Explain reasoning in short, concrete sentences
|
||||
- 简要总结当前状态
|
||||
- 列出值得处理的高频词
|
||||
- 分类别名、停用词、兴趣关键词和关注词建议
|
||||
- 用简短、具体的句子解释理由
|
||||
|
||||
The JSON output should follow:
|
||||
说明:Markdown 主要用于人工临时审阅,不必默认当作长期资产保留。
|
||||
|
||||
JSON 输出应遵循:
|
||||
|
||||
- `references/suggestion-schema.md`
|
||||
|
||||
## Repository Notes
|
||||
说明:JSON 是 review / apply 之间的唯一正式建议产物,应优先保留。
|
||||
|
||||
Current repository behavior:
|
||||
## 产物口径
|
||||
|
||||
- Keyword stats are program-maintained, not LLM-maintained
|
||||
- Stats are built from `keywords`, not `topics`
|
||||
- Stats only include non-`drop` candidates
|
||||
- `data/term_index/term_stats.json` is rebuilt from daily files, so reruns overwrite the same day instead of double-counting
|
||||
- cleanup policy, watchlist, and change log are repository-managed governance inputs and should be respected during review
|
||||
长期保留:
|
||||
|
||||
Keep suggestions aligned with that design.
|
||||
- `data/term_index/daily/*.json`
|
||||
- `data/term_index/term_stats.json`
|
||||
- `configs/filter_context.personal.json`
|
||||
- `configs/term_watchlist.json`
|
||||
- `configs/term_aliases.json`
|
||||
- `configs/term_stopwords.json`
|
||||
- `configs/term_change_log.json`
|
||||
|
||||
## Resources
|
||||
短期保留:
|
||||
|
||||
- Script:
|
||||
- `scripts/build_review_bundle.py`
|
||||
- Reference:
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`
|
||||
- `outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json`
|
||||
|
||||
临时产物:
|
||||
|
||||
- `outputs/term_index/review/keyword-cleanup-bundle.json`
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`
|
||||
|
||||
默认执行口径:
|
||||
|
||||
- bundle 只作为运行时工作文件,默认只保留当前最新一份
|
||||
- Markdown 只作为人工展示层,优先按需生成,不默认长期归档
|
||||
- JSON suggestions 是 review / apply 之间唯一正式建议输入
|
||||
|
||||
说明:
|
||||
|
||||
- “是否删除旧 bundle / 旧 markdown / 旧 suggestions” 不属于本 skill 的正式职责
|
||||
- 这类维护动作应由 OpenClaw 侧的 maintenance skill 处理
|
||||
|
||||
## 仓库说明
|
||||
|
||||
当前仓库行为:
|
||||
|
||||
- 关键词统计由程序维护,而非 LLM 维护
|
||||
- 统计基于 `keywords` 构建,而非 `topics`
|
||||
- 统计仅包含非 `drop` 候选项
|
||||
- `data/term_index/term_stats.json` 从每日文件重建,因此重新运行会覆盖同一天的数据而非重复计算
|
||||
- 清理策略、关注列表和变更日志是仓库管理的治理输入,审查时应予以尊重
|
||||
|
||||
保持建议与此设计保持一致。
|
||||
|
||||
## 资源
|
||||
|
||||
- 脚本:
|
||||
- `skills/keyword-cleanup-review/scripts/build_review_bundle.py`
|
||||
- `scripts/generate_term_cleanup_suggestions.py`
|
||||
- `scripts/generate_term_cleanup_semantic_suggestions.py`(LLM 层)
|
||||
- 参考文档:
|
||||
- `references/suggestion-schema.md`
|
||||
- `plans/keyword-cleanup-interest-watch-engine-improvement.md`(v2 引擎设计)
|
||||
|
||||
@@ -9,16 +9,15 @@ from typing import Any
|
||||
|
||||
|
||||
DEFAULT_POLICY: dict[str, Any] = {
|
||||
"schema_version": "v1",
|
||||
"schema_version": "v2",
|
||||
"interest_keyword_review": {
|
||||
"min_total_count": 3,
|
||||
"min_days_seen": 2,
|
||||
"percentile_min": 0.0,
|
||||
"percentile_max": 0.05,
|
||||
"growth_promotion": 0.5,
|
||||
},
|
||||
"watch_term_review": {
|
||||
"min_total_count": 1,
|
||||
"min_days_seen": 1,
|
||||
"max_total_count": 2,
|
||||
"max_days_seen": 2,
|
||||
"percentile_min": 0.05,
|
||||
"percentile_max": 0.20,
|
||||
},
|
||||
"alias_review": {
|
||||
"min_total_count": 2,
|
||||
@@ -28,6 +27,11 @@ DEFAULT_POLICY: dict[str, Any] = {
|
||||
"max_total_count": 2,
|
||||
"max_days_seen": 2,
|
||||
},
|
||||
"notes": [
|
||||
"v2: interest/watch 使用百分位排名 + 增速因子替代固定阈值",
|
||||
"percentile 越小表示排名越高(top 5% = percentile 0.05)",
|
||||
"growth = recent_count / total_count,衡量近期活跃度",
|
||||
],
|
||||
}
|
||||
|
||||
|
||||
@@ -126,6 +130,37 @@ def _within_watch_thresholds(item: dict[str, Any], thresholds: dict[str, Any]) -
|
||||
)
|
||||
|
||||
|
||||
def _compute_percentile(value: int, sorted_values: list[int]) -> float:
|
||||
"""
|
||||
Return the percentile rank of `value` in `sorted_values` (ascending).
|
||||
0.0 = highest frequency (top rank), 1.0 = lowest frequency (bottom rank).
|
||||
"""
|
||||
if not sorted_values:
|
||||
return 1.0
|
||||
# bisect_left — count of values strictly less than `value`
|
||||
lo, hi = 0, len(sorted_values)
|
||||
while lo < hi:
|
||||
mid = (lo + hi) // 2
|
||||
if sorted_values[mid] < value:
|
||||
lo = mid + 1
|
||||
else:
|
||||
hi = mid
|
||||
rank = lo
|
||||
# invert: smallest value → rank=0 → 1.0 (bottom)
|
||||
# largest value → rank=len → 0.0 (top)
|
||||
return 1.0 - (rank / len(sorted_values))
|
||||
|
||||
|
||||
def _compute_growth(recent_count: int, total_count: int) -> float:
|
||||
"""
|
||||
Return growth factor: recent_count / total_count.
|
||||
Only meaningful when total_count >= 3; returns 0.0 for small counts.
|
||||
"""
|
||||
if total_count < 3:
|
||||
return 0.0
|
||||
return recent_count / total_count
|
||||
|
||||
|
||||
def main() -> None:
|
||||
parser = argparse.ArgumentParser(
|
||||
description="Build a compact review bundle for the keyword-cleanup-review skill."
|
||||
@@ -235,6 +270,13 @@ def main() -> None:
|
||||
alias_values = _casefold_set(list(aliases.values()))
|
||||
watch_set = _casefold_set([str(item.get("term", "")) for item in watchlist])
|
||||
|
||||
# Build a sorted list of all total_counts for percentile computation
|
||||
all_total_counts = sorted(
|
||||
int(item.get("total_count") or 0)
|
||||
for item in stats_terms
|
||||
if isinstance(item, dict) and isinstance(item.get("term"), str)
|
||||
)
|
||||
|
||||
top_global_terms = []
|
||||
for item in stats_terms[: args.top]:
|
||||
if not isinstance(item, dict):
|
||||
@@ -256,42 +298,108 @@ def main() -> None:
|
||||
"is_alias_target": folded in alias_values,
|
||||
"in_watchlist": folded in watch_set,
|
||||
"recent_count": recent_counter.get(term, 0),
|
||||
"percentile": _compute_percentile(
|
||||
int(item.get("total_count") or 0), all_total_counts
|
||||
),
|
||||
"growth": _compute_growth(
|
||||
recent_counter.get(term, 0),
|
||||
int(item.get("total_count") or 0),
|
||||
),
|
||||
}
|
||||
)
|
||||
|
||||
# Keep more uncovered terms for percentile-based selection
|
||||
uncovered_terms = [
|
||||
item for item in top_global_terms if not item["in_interest_keywords"] and not item["is_stopword"]
|
||||
][:20]
|
||||
][:100]
|
||||
|
||||
policy_version = (policy.get("schema_version") if isinstance(policy, dict) else None) or "v1"
|
||||
interest_thresholds = policy.get("interest_keyword_review") if isinstance(policy, dict) else {}
|
||||
watch_thresholds = policy.get("watch_term_review") if isinstance(policy, dict) else {}
|
||||
|
||||
if policy_version == "v2" or "percentile_max" in interest_thresholds:
|
||||
# v2: percentile + growth based selection
|
||||
pct_min_interest = float(interest_thresholds.get("percentile_min", 0.0))
|
||||
pct_max_interest = float(interest_thresholds.get("percentile_max", 0.05))
|
||||
growth_promo = float(interest_thresholds.get("growth_promotion", 0.5))
|
||||
pct_min_watch = float(watch_thresholds.get("percentile_min", 0.05))
|
||||
pct_max_watch = float(watch_thresholds.get("percentile_max", 0.20))
|
||||
|
||||
interest_candidates_raw = [
|
||||
item for item in uncovered_terms
|
||||
if not item["in_watchlist"]
|
||||
and pct_min_interest <= item["percentile"] <= pct_max_interest
|
||||
]
|
||||
watch_candidates_raw = [
|
||||
item for item in uncovered_terms
|
||||
if not item["in_watchlist"]
|
||||
and pct_min_watch < item["percentile"] <= pct_max_watch
|
||||
]
|
||||
# Growth boost: terms outside watch range but with strong growth signal
|
||||
growth_boost_candidates = [
|
||||
item for item in uncovered_terms
|
||||
if not item["in_watchlist"]
|
||||
and item["percentile"] > pct_max_watch
|
||||
and item["growth"] >= growth_promo
|
||||
]
|
||||
else:
|
||||
# v1 fallback: fixed thresholds
|
||||
interest_candidates_raw = [
|
||||
item for item in uncovered_terms
|
||||
if not item["in_watchlist"] and _meets_min_thresholds(item, interest_thresholds)
|
||||
]
|
||||
watch_candidates_raw = [
|
||||
item for item in uncovered_terms
|
||||
if not item["in_watchlist"]
|
||||
and not _meets_min_thresholds(item, interest_thresholds)
|
||||
and _within_watch_thresholds(item, watch_thresholds)
|
||||
]
|
||||
growth_boost_candidates = []
|
||||
|
||||
interest_review_candidates = [
|
||||
{
|
||||
"term": item["term"],
|
||||
"total_count": item["total_count"],
|
||||
"days_seen": item["days_seen"],
|
||||
"percentile": item["percentile"],
|
||||
"growth": item["growth"],
|
||||
"reason": (
|
||||
"Meets the configured interest-keyword review threshold and is not yet covered "
|
||||
"by interest keywords or stopwords."
|
||||
f"top {item['percentile']:.1%} by frequency,"
|
||||
f"growth={item['growth']:.0%},"
|
||||
"not yet covered by interest keywords or stopwords."
|
||||
),
|
||||
}
|
||||
for item in uncovered_terms
|
||||
if not item["in_watchlist"] and _meets_min_thresholds(item, interest_thresholds)
|
||||
for item in interest_candidates_raw
|
||||
][:20]
|
||||
watch_review_candidates = [
|
||||
{
|
||||
"term": item["term"],
|
||||
"total_count": item["total_count"],
|
||||
"days_seen": item["days_seen"],
|
||||
"percentile": item["percentile"],
|
||||
"growth": item["growth"],
|
||||
"reason": (
|
||||
"Falls into the configured watch-term review range and should be observed "
|
||||
"before promotion into interest keywords."
|
||||
f"top {item['percentile']:.1%} by frequency,"
|
||||
f"growth={item['growth']:.0%},"
|
||||
"fell into watch-review range."
|
||||
),
|
||||
}
|
||||
for item in uncovered_terms
|
||||
if not item["in_watchlist"]
|
||||
and not _meets_min_thresholds(item, interest_thresholds)
|
||||
and _within_watch_thresholds(item, watch_thresholds)
|
||||
for item in watch_candidates_raw
|
||||
][:20]
|
||||
growth_boost_review_items = [
|
||||
{
|
||||
"term": item["term"],
|
||||
"total_count": item["total_count"],
|
||||
"days_seen": item["days_seen"],
|
||||
"percentile": item["percentile"],
|
||||
"growth": item["growth"],
|
||||
"reason": (
|
||||
f"growth spike: {item['growth']:.0%} of occurrences in recent window "
|
||||
f"(total={item['total_count']}, days={item['days_seen']})."
|
||||
),
|
||||
}
|
||||
for item in growth_boost_candidates
|
||||
][:5]
|
||||
recent_hot_terms = sorted(
|
||||
({"term": term, "recent_count": count} for term, count in recent_counter.items()),
|
||||
key=lambda item: (-item["recent_count"], item["term"].casefold(), item["term"]),
|
||||
@@ -330,6 +438,7 @@ def main() -> None:
|
||||
"governance_hints": {
|
||||
"interest_review_candidates": interest_review_candidates,
|
||||
"watch_review_candidates": watch_review_candidates,
|
||||
"growth_boost_review_items": growth_boost_review_items,
|
||||
},
|
||||
}
|
||||
_save_json(args.output, bundle)
|
||||
|
||||
@@ -1,55 +1,55 @@
|
||||
---
|
||||
name: llm-summary-review
|
||||
description: Validate and refine LLM-generated summary JSON for extracted articles in this repository. Use when the user wants to review, validate, repair, or iterate on `outputs/result.json` or similar summary outputs produced from extracted article JSON.
|
||||
description: 验证和优化 LLM 生成的文章摘要 JSON。当用户需要审查、验证、修复或迭代 `outputs/result.json` 或其他由已提取文章 JSON 生成的摘要输出时使用。
|
||||
---
|
||||
|
||||
# LLM Summary Review
|
||||
# LLM 摘要审查
|
||||
|
||||
Use this skill when working on the repository's summary loop after article extraction is done.
|
||||
在文章提取完成后,处理本仓库摘要循环时使用此技能。
|
||||
|
||||
## What This Skill Does
|
||||
## 功能说明
|
||||
|
||||
- Verifies that an LLM summary result matches the expected JSON contract
|
||||
- Reuses the repository validator instead of re-checking fields manually
|
||||
- Repairs invalid outputs by telling the LLM exactly what to fix
|
||||
- Keeps the workflow aligned with the extraction JSON produced by this project
|
||||
- 验证 LLM 摘要结果是否符合预期的 JSON 约定
|
||||
- 复用仓库验证器,而非手动逐字段检查
|
||||
- 通过告知 LLM 具体需要修复的内容来修复无效输出
|
||||
- 保持工作流与本项目生成的提取 JSON 保持一致
|
||||
|
||||
## Inputs
|
||||
## 输入
|
||||
|
||||
Typical files:
|
||||
典型文件:
|
||||
|
||||
- Extracted article JSON: `outputs/*.extracted.json`
|
||||
- LLM summary result JSON: `outputs/result.json`
|
||||
- Prompt template: `outputs/llm-summary-prompt.txt`
|
||||
- 已提取的文章 JSON:`outputs/*.extracted.json`
|
||||
- LLM 摘要结果 JSON:`outputs/result.json`
|
||||
- 提示词模板:`outputs/llm-summary-prompt.txt`
|
||||
|
||||
## Workflow
|
||||
## 工作流程
|
||||
|
||||
1. Validate the current summary result with the repository validator:
|
||||
1. 使用仓库验证器验证当前摘要结果:
|
||||
|
||||
```bash
|
||||
python -m summary_mcp.validate_llm_result outputs/result.json --extracted outputs/read-flow-2026.extracted.json
|
||||
```
|
||||
|
||||
2. If validation passes:
|
||||
- Report that the result is structurally valid
|
||||
- Briefly note any warnings
|
||||
- Do not rewrite the result unless the user asks
|
||||
2. 如果验证通过:
|
||||
- 报告结果结构有效
|
||||
- 简要说明任何警告
|
||||
- 除非用户要求,否则不重写结果
|
||||
|
||||
3. If validation fails:
|
||||
- Read the validator errors carefully
|
||||
- Ask the LLM to regenerate or repair only the failing parts
|
||||
- Re-run the validator until it passes or a retry limit is hit
|
||||
3. 如果验证失败:
|
||||
- 仔细阅读验证器错误信息
|
||||
- 要求 LLM 仅重新生成或修复失败的部分
|
||||
- 重新运行验证器,直到通过或达到重试上限
|
||||
|
||||
## Repair Prompt Pattern
|
||||
## 修复提示词模式
|
||||
|
||||
When asking an LLM to repair a bad result, provide:
|
||||
当要求 LLM 修复有问题的结果时,需提供:
|
||||
|
||||
- The original extracted article JSON
|
||||
- The current invalid summary JSON
|
||||
- The validator error list
|
||||
- A strict instruction to preserve valid fields and fix only the failing ones
|
||||
- 原始已提取的文章 JSON
|
||||
- 当前无效的摘要 JSON
|
||||
- 验证器错误列表
|
||||
- 严格指令:保留正确字段,仅修复失败字段
|
||||
|
||||
Use this repair template:
|
||||
使用以下修复模板:
|
||||
|
||||
```text
|
||||
请修复下面这份不符合要求的摘要 JSON。
|
||||
@@ -70,32 +70,32 @@ validator errors:
|
||||
{{result_json}}
|
||||
```
|
||||
|
||||
## Validation Rules
|
||||
## 验证规则
|
||||
|
||||
The validator currently enforces:
|
||||
验证器当前强制执行以下规则:
|
||||
|
||||
- Required fields exist
|
||||
- Field types are correct
|
||||
- `category` is one of: `资讯` `方法论` `工具实践` `观点评论`
|
||||
- `summary` length is within bounds
|
||||
- `highlights`, `keywords`, and `topics` counts are within bounds
|
||||
- `keywords` and `topics` do not overlap
|
||||
- `title` and `url` match the extracted article when an extracted JSON file is provided
|
||||
- 必填字段存在
|
||||
- 字段类型正确
|
||||
- `category` 取值为以下之一:`资讯` `方法论` `工具实践` `观点评论`
|
||||
- `summary` 长度在允许范围内
|
||||
- `highlights`、`keywords` 和 `topics` 数量在允许范围内
|
||||
- `keywords` 和 `topics` 不重叠
|
||||
- 提供已提取 JSON 文件时,`title` 和 `url` 与已提取文章匹配
|
||||
|
||||
## Repository Implementation
|
||||
## 代码实现
|
||||
|
||||
Relevant code:
|
||||
相关代码:
|
||||
|
||||
- Validator model: `src/summary_mcp/models/llm_result.py`
|
||||
- Validator logic: `src/summary_mcp/validators/llm_result.py`
|
||||
- CLI entry: `src/summary_mcp/validate_llm_result.py`
|
||||
- 验证器模型:`src/summary_mcp/models/llm_result.py`
|
||||
- 验证器逻辑:`src/summary_mcp/validators/llm_result.py`
|
||||
- CLI 入口:`src/summary_mcp/validate_llm_result.py`
|
||||
|
||||
Prefer using the existing validator rather than recreating checks in free-form reasoning.
|
||||
优先使用现有验证器,而非在自由推理中重新创建检查逻辑。
|
||||
|
||||
## When To Stop
|
||||
## 何时停止
|
||||
|
||||
Stop when one of these is true:
|
||||
满足以下条件之一时停止:
|
||||
|
||||
- The validator returns `valid: true`
|
||||
- The user asks to inspect the remaining failures manually
|
||||
- Repeated retries fail and the user should decide how to proceed
|
||||
- 验证器返回 `valid: true`
|
||||
- 用户要求手动检查剩余的失败项
|
||||
- 反复重试失败,应由用户决定后续操作
|
||||
|
||||
@@ -0,0 +1,186 @@
|
||||
---
|
||||
name: reader-digest-flow
|
||||
description: 编排 reader 项目的端到端 AI 日报流程。仅在用户要求运行/重跑日报、汇报候选、发布 Hugo 日报、沉淀选中文章或继续已有日报任务时使用;覆盖异步 MCP 任务、候选确认、发布、单篇摘要和 IMA 知识库上传。不要因验证、排障冲动或候选质量不佳自行重跑。
|
||||
---
|
||||
|
||||
# Reader Digest Flow
|
||||
|
||||
## 职责边界
|
||||
|
||||
本 Skill 负责:
|
||||
|
||||
- 通过 reader MCP 启动、观察和恢复日报任务;
|
||||
- 向用户展示候选并保持稳定编号;
|
||||
- 根据用户选择生成并发布 Hugo 日报;
|
||||
- 对用户选中的文章生成知识笔记并编排 IMA 上传;
|
||||
- 在每个副作用边界执行确认和结果验证。
|
||||
|
||||
本 Skill 不负责:
|
||||
|
||||
- 实现 reader 内部抓取、摘要、过滤或恢复逻辑;
|
||||
- 通过手拼目录推导 Run 状态或 Artifact;
|
||||
- 未经用户要求自行重跑 Pipeline;
|
||||
- 未经用户确认发布日报或写入知识库;
|
||||
- 直接维护关键词配置;关键词治理委托给 `keyword-cleanup-review`。
|
||||
|
||||
## 核心规则
|
||||
|
||||
1. **只按用户指令运行。** 只有用户明确要求“跑日报”“重新跑”“再跑一次”时才启动新 Pipeline。验证、解释排序和排障默认读取已有 Run。
|
||||
2. **一次对话绑定一个当前 Run。** 以异步 Job 结果返回的 `run_id` 为稳定句柄;新 Run 产生新的候选编号体系,不混用历史编号。
|
||||
3. **状态以 MCP 返回为准。** Agent 只根据顶层 `status` 和 `recommended_action` 分支;`status_source`、`state_conflict` 仅用于解释。
|
||||
4. **路径以返回值为准。** 使用 `output_dir`、`artifact.path`、`delivery_output`、`report_output` 和 `written_paths`;不要根据 `run_id` 手拼 `outputs/...`。
|
||||
5. **候选编号保持稳定。** 用户编号永远对应当前候选列表的原始顺序(1-based);跨产物读取详情时按 URL 或完整 `item_id` 关联,不按数组位置关联。
|
||||
6. **副作用必须授权。** 用户确认 Hugo 文章后才能发布;用户确认 IMA 文章后才能生成并上传知识笔记。
|
||||
7. **内容必须有来源。** 日报和知识笔记只能基于当前 Run 的 `article.plain_text`、摘要、highlights 等 Artifact;不得使用通用知识补写原文没有的信息,也不为满足长度而扩写。
|
||||
|
||||
## 默认生产参数
|
||||
|
||||
用户未显式覆盖时使用:
|
||||
|
||||
```json
|
||||
{
|
||||
"limit": 7,
|
||||
"include_read": false,
|
||||
"mark_read": true,
|
||||
"debug_artifacts": false,
|
||||
"timeout_seconds": 60,
|
||||
"max_retries": 2
|
||||
}
|
||||
```
|
||||
|
||||
- 默认只处理未读文章。
|
||||
- `include_read=true` 仅在用户明确要求扩大到已读内容时使用。
|
||||
- debug/test/validation 才允许 `mark_read=false` 或 `debug_artifacts=true`。
|
||||
- 不随机生成文章数量;用户指定 `limit` 时按用户值执行。
|
||||
|
||||
## 正式流程
|
||||
|
||||
### Phase 1:启动并观察日报 Job
|
||||
|
||||
正式生产入口统一为异步 MCP:
|
||||
|
||||
1. 调用 `start_freshrss_pipeline_job`;
|
||||
2. 轮询 `get_freshrss_pipeline_job_status`;
|
||||
3. `status=success` 后调用 `get_freshrss_pipeline_job_result`;
|
||||
4. 保存返回的 `run_id` 和 Artifact 路径;
|
||||
5. 使用 `get_run_status`、`get_delivery_payload`、`get_run_report` 读取业务状态和结果。
|
||||
|
||||
失败时:
|
||||
|
||||
1. 使用 `get_run_status(run_id)` 读取关联 Run;
|
||||
2. 调用 `inspect_resume_plan(run_id)`;
|
||||
3. `recommended_action=resume` 时启动并轮询异步 Resume Job;
|
||||
4. `recommended_action=read_terminal_result` 时直接读取已有终态结果;
|
||||
5. `recommended_action=start_new_run` 时停止并向用户报告,不自行新建 Run。
|
||||
|
||||
CLI 仅用于 MCP 不可用时的 fallback、debug 或人工排障,不是默认生产入口。具体调用序列见 `references/flow.md`。
|
||||
|
||||
### Phase 2:汇报候选
|
||||
|
||||
- 使用当前 Run 返回的 Delivery Payload 或 digest brief Artifact;
|
||||
- 按候选原始顺序从 1 编号,状态可显示为“已入选/待确认”,但不得重新分组编号;
|
||||
- 每篇提供标题、来源、2-3 句摘要和筛选理由,避免原始 JSON dump;
|
||||
- 用户质疑编号或排序时读取当前 Run 产物核对,不重新运行 Pipeline;
|
||||
- 需要跨 Artifact 取详情时按 URL 或完整 `item_id` 交叉验证。
|
||||
|
||||
Feishu 输出不要使用 Markdown 表格,见 `references/feishu-format-notes.md`。
|
||||
|
||||
### Phase 3:等待 Hugo 选择
|
||||
|
||||
- 等待用户明确选择要发布的文章;
|
||||
- 用户编号映射到当前候选列表,不映射到 extracted 文件序号;
|
||||
- 用户拒绝发布时立即停止当日日报后续流程,不劝说、不自动换一批;
|
||||
- 用户明确要求重跑时才创建新 Run,并重新建立编号体系。
|
||||
|
||||
### Phase 4:生成并发布 Hugo 日报
|
||||
|
||||
发布前读取 `references/public-digest-example.md`,按其最终页面结构生成:
|
||||
|
||||
- `今日概览`
|
||||
- `今日重点`
|
||||
- `趋势观察`
|
||||
|
||||
每篇 `今日重点` 文章末尾必须添加 `来源:[来源名](原文 URL)`,来源链接跟随对应文章,不再生成独立的 `延伸阅读` 章节或重复链接。
|
||||
|
||||
仅发布用户在 Phase 3 选中的文章。公开页面不得出现 `keep/review/drop`、候选、待确认等内部状态。
|
||||
|
||||
写入 Hugo 后执行部署,并验证首页、日报列表页和当日详情页均可访问。命令和检查项见 `references/flow.md`。
|
||||
|
||||
### Phase 5:等待 IMA 选择
|
||||
|
||||
Hugo 发布选择与知识沉淀选择相互独立。询问用户哪些文章值得长期保存:
|
||||
|
||||
- 仅处理用户明确选择的文章;
|
||||
- 不把整份日报上传到 IMA;
|
||||
- 本次选择本身即授权后续单篇摘要和 IMA 上传,不重复确认。
|
||||
|
||||
### Phase 6:生成单篇知识笔记
|
||||
|
||||
对每篇选中文章:
|
||||
|
||||
1. 通过 URL/完整 `item_id` 找到对应 extracted Artifact;
|
||||
2. 使用 `start_article_summary_job(extracted_path=<返回的 Artifact 路径>, selected_ids=[<完整 item_id>])`;
|
||||
3. 轮询 `get_article_summary_job_status`,成功后读取 `get_article_summary_job_result`;
|
||||
4. 只使用现有 `article.plain_text`,不重新抓取原 URL;
|
||||
5. 使用返回的 `written_paths` 定位结果并检查 Markdown 内容。
|
||||
6. **`extracted_path` 必须传绝对路径**(前缀 `/home/ubuntu/zhu/github/reader/`):summary-mcp 工作目录是 `/root/.hermes`,相对路径会报 `extracted_path does not exist`。同一工具连续 3 次失败会触发 MCP 冷却(约 45-60s,报 `MCP server 'reader' is unreachable`),等待冷却后再重试,不要循环重试同一调用。
|
||||
|
||||
异步 MCP 不可用时才使用项目 CLI fallback。不要因为内容较短而引入原文之外的知识。
|
||||
|
||||
### Phase 7:上传到 IMA
|
||||
|
||||
上传前按需读取:
|
||||
|
||||
- 格式规则:`references/ima-format-quickref.md`
|
||||
- API 和上传步骤:`references/ima-upload-api.md`(含 `-200` 版本拦截修复)
|
||||
- 凭证定位:`references/ima-credential-chain.md`
|
||||
- 批量上传脚本:`scripts/ima_upload_one.py`(Python 编排,规避中文文件名 bash 引号问题)
|
||||
|
||||
硬规则:
|
||||
|
||||
- 使用完整文章标题作为文件名;
|
||||
- 以 Markdown 文件 `media_type=7` 上传到 `daily` knowledge base;
|
||||
- 保留原文 URL 和 Category,保证来源可追踪;
|
||||
- 不使用 URL 导入或 Notes 类型替代知识库文件;
|
||||
- 上传后验证目标知识库中存在对应条目;
|
||||
- 失败时报告具体阶段,不无限重试。
|
||||
|
||||
## 关键词治理路由
|
||||
|
||||
只有用户明确要求“清理关键词”“词库治理”等操作时才触发关键词治理。Review bundle 和 Suggestions 生成委托给 `keyword-cleanup-review`,确认与 Apply 仍由当前编排层负责:
|
||||
|
||||
1. 生成 review bundle;
|
||||
2. 生成规则层与可选语义层 Suggestions JSON;
|
||||
3. 等待人工确认;
|
||||
4. dry-run 后按精确 accept 参数 Apply。
|
||||
|
||||
`reader-digest-flow` 不直接编辑 `term_aliases.json`、`term_stopwords.json` 或兴趣配置。简要路由见 `references/keyword-engine-maintenance.md`。
|
||||
|
||||
## 环境坑位(本机部署)
|
||||
|
||||
- **MCP 相对路径陷阱**:`summary-mcp` 服务工作目录是 `/root/.hermes`,不是 reader 项目根。MCP 返回的 `output_dir`/`artifact.path` 是相对路径,直接传给 `start_article_summary_job(extracted_path=...)` 会报 `extracted_path does not exist`。传入前必须拼绝对路径前缀 `/home/ubuntu/zhu/github/reader/`。
|
||||
- **提取失败不等于运行失败**:`status_counts.extract_failed` 的条目(`CONTENT_EXTRACTION_FAILED`,`retryable=false`)跳过即可并如实汇报;失败文章常是推广/活动等低价值内容,不因此自行重跑。`linked_run_status=partial` 时先读 run-report 的 item 级 `error` 确认原因。
|
||||
- **用户要求"重新跑一批"**:候选质量低(用户主动提出)时重跑,应 `include_read=true` 并调高 `limit`(如 10),否则默认 `include_read=false` 会拉回同一批未读文章。重跑是新 Run,候选编号体系重新建立,汇报时提醒用户按新列表选择。
|
||||
|
||||
## 停止与人工介入
|
||||
|
||||
出现以下任一情况时停止自动流程并报告:
|
||||
|
||||
- 用户没有授权运行、发布或知识库写入;
|
||||
- Job/Run 返回不可恢复,或连续恢复失败;
|
||||
- Payload、候选 ID 或 Artifact 之间无法可靠关联;
|
||||
- 生成内容缺少可追踪来源;
|
||||
- Hugo 部署验证失败;
|
||||
- IMA 凭证、目标知识库或上传结果无法验证。
|
||||
|
||||
## Reference 路由
|
||||
|
||||
- `references/flow.md`:具体 MCP 调用序列、候选映射(含 extracted_path 绝对路径、候选≠文件名顺序)、Hugo 发布和 IMA 主步骤。
|
||||
- `references/content-extraction.md`:FreshRSS 内容来源与 `plain_text` 质量判断。
|
||||
- `references/public-digest-example.md`:可直接参考的 Hugo 最终页面结构。
|
||||
- `references/feishu-format-notes.md`:Feishu 输出格式限制。
|
||||
- `references/ima-format-quickref.md`:IMA Markdown 格式规则。
|
||||
- `references/ima-upload-api.md`:IMA Markdown 文件上传 API(含 `-200` 版本拦截修复)。
|
||||
- `references/ima-credential-chain.md`:IMA 凭证与知识库配置定位。
|
||||
- `references/keyword-engine-maintenance.md`:关键词治理 Skill 路由。
|
||||
- `scripts/ima_upload_one.py`:单篇 Markdown 上传 daily 知识库的完整 Python 脚本(preflight→重名→create_media→COS→add_knowledge)。
|
||||
@@ -0,0 +1,45 @@
|
||||
# 内容提取流程
|
||||
|
||||
本文说明处理流水线如何把 FreshRSS 条目转换为可供摘要使用的文章文本。
|
||||
|
||||
## 核心规则:FreshRSS 条目不重新抓取原文 URL
|
||||
|
||||
**FreshRSS 是仅提供 RSS 内容的上游。** 对于 FreshRSS 条目,流水线不会向文章原始 URL 发起 HTTP 请求。该行为由 `pipeline.py` 中的 `RSS_ONLY_UPSTREAMS = {"freshrss"}` 强制保证。
|
||||
|
||||
唯一例外是非 FreshRSS 上游。未来未设置 `upstream: freshrss` 的其他来源,可以在必要时使用 `fetch_html()` 作为回退。
|
||||
|
||||
## 内容来源优先级
|
||||
|
||||
`content_loader.py` 按以下顺序检查内容,并使用第一个包含 **至少 500 个可读字符** 的来源:
|
||||
|
||||
| 优先级 | 来源 | 含义 |
|
||||
|--------|------|------|
|
||||
| 1 | `raw_html` | 通过 `ExtractionInput.raw_html` 预先注入的 HTML;常规 FreshRSS 运行中很少使用。 |
|
||||
| 2 | `item.raw_content` | RSS `<content:encoded>` 中的文章正文;部分订阅源提供,部分不提供。 |
|
||||
| 3 | `item.raw_summary` | RSS `<description>` 中的摘要或片段;这是当前运行中最常见的来源。 |
|
||||
| 4 | `rss_content` | 来自非条目字段的独立 RSS 内容。 |
|
||||
| — | `none` | 没有可用内容;FreshRSS 不允许回源抓取,因此抛出 `RSS_CONTENT_MISSING`。 |
|
||||
|
||||
## `content_source` 与文本质量的关系
|
||||
|
||||
每个 `item-XX.extracted.json` 中的 `content_source` 字段表示流水线实际使用的内容来源:
|
||||
|
||||
- **`item.raw_content`**:RSS `<content:encoded>` 提供的文章正文,通常质量最好,接近直接阅读原文。
|
||||
- **`item.raw_summary`**:只有 RSS 摘要或描述,并非完整正文。不同来源长度差异较大,通常为 300-2000 个字符;AI 摘要基于该片段,而不是完整文章。
|
||||
- **`rss_content`**:来自独立 RSS 内容,质量取决于订阅源。
|
||||
- **`fetched_html`**:从原始 URL 抓取的 HTML。FreshRSS 条目不会出现该来源,只适用于非 FreshRSS 上游。
|
||||
|
||||
## 对日报质量的影响
|
||||
|
||||
如果提取结果文件中出现 `content_source: item.raw_summary`,说明 AI 使用的是订阅源摘要或片段,而不是完整正文。日报内容显得较浅时,原因可能只是 RSS 描述过短。
|
||||
|
||||
提高质量可以选择提供完整 `<content:encoded>` 的订阅源,或者把内容来源切换到支持全文 RSS 的系统,例如具备全文提取能力的 RSS 代理或 FiveFilters 等服务。
|
||||
|
||||
## 快速检查
|
||||
|
||||
先调用 `list_run_artifacts(run_id)`,再读取返回的提取结果产物路径,检查 `content_source` 和 `article.plain_text`。不要按 `run_id` 或“最新目录”手拼路径。
|
||||
|
||||
## 相关代码路径
|
||||
|
||||
- `src/summary_mcp/core/pipeline.py`:`RSS_ONLY_UPSTREAMS`、`_should_skip_fetch()`、`extract_content()`。
|
||||
- `src/summary_mcp/core/content_loader.py`:`choose_inline_content()` 的优先级链和 `fetch_html()`;FreshRSS 条目不会调用后者。
|
||||
@@ -0,0 +1,17 @@
|
||||
# 飞书 Markdown 格式说明
|
||||
|
||||
## 背景
|
||||
|
||||
Hermes 的飞书网关(`gateway/platforms/feishu.py`)通过 `_build_outbound_payload` 发送消息。该方法会检查内容中的 Markdown 特征,并据此决定消息类型:
|
||||
|
||||
- 内容匹配 `_MARKDOWN_HINT_RE`(加粗、列表、代码、链接等)时,使用包含 `md` 元素的飞书 `post` 类型发送,可以正常渲染。
|
||||
- 内容匹配 `_MARKDOWN_TABLE_RE`(Markdown 表头和分隔行)时,整条消息会被强制转换为 `text` 类型,即纯文本,不再渲染 Markdown。
|
||||
|
||||
原因是 `_build_markdown_post_payload` 会把内容包装为 `{"tag": "md", "text": "..."}` 元素,而飞书的 `md` 元素不支持表格,也没有把 Markdown 表格转换为飞书原生表格的逻辑。
|
||||
|
||||
## 飞书输出规则
|
||||
|
||||
- 通过飞书发送的消息不得使用 Markdown 表格;消息中只要出现一个表格,整条消息就会退化为纯文本。
|
||||
- 需要表达结构化信息时,优先使用分点列表、带标题的分节或行内格式。
|
||||
- 加粗(`**加粗**`)、行内代码(`` `代码` ``)、无序列表(`- 项目`)、有序列表(`1. 项目`)和链接均可正常使用。
|
||||
- 围栏式代码块可以使用,但代码块后的尾随内容可能存在渲染边界问题。
|
||||
@@ -0,0 +1,177 @@
|
||||
# Reader Digest Flow 操作参考
|
||||
|
||||
本文件承载 `reader-digest-flow` 的具体执行步骤。正式行为边界以 `../SKILL.md` 为准。
|
||||
|
||||
## 1. 日报 Pipeline
|
||||
|
||||
### 默认参数
|
||||
|
||||
```json
|
||||
{
|
||||
"limit": 7,
|
||||
"include_read": false,
|
||||
"mark_read": true,
|
||||
"debug_artifacts": false,
|
||||
"timeout_seconds": 60,
|
||||
"max_retries": 2
|
||||
}
|
||||
```
|
||||
|
||||
### 正式调用序列
|
||||
|
||||
```text
|
||||
start_freshrss_pipeline_job
|
||||
→ get_freshrss_pipeline_job_status
|
||||
→ get_freshrss_pipeline_job_result
|
||||
→ get_run_status
|
||||
→ get_delivery_payload / get_run_report
|
||||
```
|
||||
|
||||
状态动作:
|
||||
|
||||
- `running`:按合理间隔继续轮询;
|
||||
- `success`:读取结果,保存 `run_id`;
|
||||
- `failed`:读取关联 Run 并执行 Resume Plan;
|
||||
- `partial`:优先检查 Run Report、Artifact 和 Recovery 信息。
|
||||
|
||||
恢复序列:
|
||||
|
||||
```text
|
||||
inspect_resume_plan
|
||||
→ recommended_action=resume
|
||||
→ start_resume_job
|
||||
→ get_resume_job_status
|
||||
→ get_resume_job_result
|
||||
```
|
||||
|
||||
如果建议为 `read_terminal_result`,直接读取结果;如果为 `start_new_run`,停止并等待用户决定。
|
||||
|
||||
不要根据 `run_id` 拼目录。后续只消费调用返回的 `output_dir`、`artifact.path`、`delivery_output`、`report_output`。
|
||||
|
||||
## 2. 候选汇报与选择
|
||||
|
||||
优先使用当前 Run 的 digest brief Artifact;不可用时使用 `get_delivery_payload` 返回的候选。
|
||||
|
||||
展示规则:
|
||||
|
||||
1. 使用候选数组原始顺序并从 1 编号;
|
||||
2. 不因 `keep/review` 分组而重新编号;
|
||||
3. 每篇展示标题、来源、摘要和判断理由;
|
||||
4. 用户编号只映射当前候选数组;
|
||||
5. 跨 Artifact 读取详情时按 URL 或完整 `item_id` 匹配。
|
||||
|
||||
不要按 `summary-batch`、`item-XX` 和候选数组的相同位置推断它们是同一篇文章。
|
||||
|
||||
### extracted 文件与候选编号错位(实测 2026-07-31)
|
||||
|
||||
- `digest-brief.json` 的 `top_candidates` **没有 `item_id` 字段**,只有 `url`;
|
||||
- `extracted/item-XX.extracted.json` 的文件名序号与候选编号**可能不一致**(实例:候选2 = item-05、候选3 = item-02);
|
||||
- 正确做法:用 **URL 交叉匹配**(归一化 `%3D`→`=` 后逐条比对),或用完整 `item_id`(从 candidate-batch.json 的 `items[i].item_id` 按候选数组顺序取)在 extracted 文件里反查;两者都能验证时优先 item_id。
|
||||
|
||||
## 3. Hugo 日报
|
||||
|
||||
用户确认发布文章后:
|
||||
|
||||
1. 读取 `public-digest-example.md`;
|
||||
2. 仅使用用户选中的文章生成公开内容;
|
||||
3. 写入 Hugo 当日页面;
|
||||
4. 前台执行部署,避免把构建日志作为聊天通知;
|
||||
5. 验证首页、列表页和详情页。
|
||||
|
||||
当前部署位置:
|
||||
|
||||
```text
|
||||
/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md
|
||||
```
|
||||
|
||||
验证至少覆盖:
|
||||
|
||||
```text
|
||||
http://127.0.0.1:14322/
|
||||
http://127.0.0.1:14322/daily/
|
||||
http://127.0.0.1:14322/daily/YYYY-MM-DD/
|
||||
```
|
||||
|
||||
页面不存在时检查源文件、构建结果、容器状态和页面日期。需要 Docker/Hugo 深度排障时再读取部署项目自己的文档,不把历史事故规则复制回主 Skill。
|
||||
|
||||
## 4. 单篇知识笔记
|
||||
|
||||
用户确认 IMA 文章后:
|
||||
|
||||
1. 从候选中取得 URL 和完整 `item_id`;
|
||||
2. 从 Run Artifact 中找到匹配的 extracted 文件;
|
||||
3. 交叉验证 `article.item_id` 或 URL;
|
||||
4. 对每个单篇文件调用 `start_article_summary_job(extracted_path=<返回的 Artifact 路径>, selected_ids=[<完整 item_id>])`;
|
||||
5. `selected_ids` 传完整、非空、去除 `cand:` 前缀后的 `item_id`;
|
||||
6. 从 Job 结果的 `written_paths` 读取 Markdown。
|
||||
|
||||
### ⚠️ extracted_path 必须用绝对路径
|
||||
|
||||
`summary-mcp` 进程的工作目录是 `/root/.hermes`(不是 reader 项目根)。传相对路径(如 `outputs/freshrss/...`)会直接报 `extracted_path does not exist`。必须传绝对路径:
|
||||
|
||||
```text
|
||||
/home/ubuntu/zhu/github/reader/outputs/freshrss/rerun/<run_dir>/extracted/item-XX.extracted.json
|
||||
```
|
||||
|
||||
### ⚠️ 候选编号 ≠ extracted 文件名顺序
|
||||
|
||||
候选数组顺序与 extracted 文件名(`item-01`…`item-07`)**不一定对齐**(实测候选2 落在 item-05)。`digest-brief.json` 的候选**没有 `item_id` 字段**,只有 URL。可靠匹配方法:
|
||||
|
||||
1. 从 `candidate-batch.json` 取每项完整 `item_id`(在 `candidate` 嵌套对象里,顶层 `item_key` 只是 `item-XX` 文件名序号);
|
||||
2. 或按 URL 匹配:归一化(`%3D`→`=`)后与每个 extracted 文件的 `article.url` / `article.canonical_url` 比对;
|
||||
3. 绝不要按候选位置对应 extracted 文件序号。
|
||||
|
||||
```python
|
||||
def norm(u): return u.replace('%3D','=').replace('%3d','=').strip()
|
||||
# 对每个 extracted 文件取 norm(article.url),与候选 norm(url) 精确比对
|
||||
```
|
||||
|
||||
正式序列:
|
||||
|
||||
```text
|
||||
start_article_summary_job
|
||||
→ get_article_summary_job_status
|
||||
→ get_article_summary_job_result
|
||||
```
|
||||
|
||||
内容依据是 extracted Artifact 中的 `article.plain_text`。不得重新抓取原 URL,不得使用原文之外的知识扩写。
|
||||
|
||||
CLI 仅在异步 MCP 不可用或人工排障时使用:
|
||||
|
||||
```bash
|
||||
python scripts/run_article_summaries.py \
|
||||
--extracted <returned-extracted-path> \
|
||||
--ids <full-item-id> \
|
||||
--output-dir <explicit-output-dir>
|
||||
```
|
||||
|
||||
## 5. IMA 上传
|
||||
|
||||
用户在知识沉淀阶段的文章选择即为上传授权。
|
||||
|
||||
执行顺序:
|
||||
|
||||
1. 检查生成的 Markdown 与来源;
|
||||
2. 文件名规范化为 `<完整文章标题>.md`;
|
||||
3. 确认目标为 `daily` knowledge base;
|
||||
4. 执行 preflight、create_media、COS upload、add_knowledge;
|
||||
5. 验证知识库条目存在。
|
||||
|
||||
上传格式与 API 参数分别见:
|
||||
|
||||
- `ima-format-quickref.md`
|
||||
- `ima-upload-api.md`
|
||||
- `ima-credential-chain.md`
|
||||
|
||||
失败时记录发生在 preflight、create_media、COS upload 或 add_knowledge 的具体阶段。不要把失败的 Markdown 改写成 URL 导入或 Notes 类型来绕过错误。
|
||||
|
||||
## 6. CLI fallback 原则
|
||||
|
||||
CLI 仅在以下场景使用:
|
||||
|
||||
- MCP 服务不可用;
|
||||
- Tool transport/launch 失败且无法取得有效 Job;
|
||||
- 用户明确要求本地调试;
|
||||
- 人工排障需要直接检查脚本输出。
|
||||
|
||||
如果已取得 `job_id` 或 `run_id`,先查询真实状态,避免因响应超时重复启动任务。
|
||||
@@ -0,0 +1,27 @@
|
||||
# IMA 凭证与安全边界
|
||||
|
||||
## 必需配置
|
||||
|
||||
- `IMA_OPENAPI_CLIENTID`
|
||||
- `IMA_OPENAPI_APIKEY`
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_ID`
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_NAME=daily`
|
||||
|
||||
优先使用当前进程环境和 IMA Skill 已支持的凭证加载机制。不要在本 Skill 中复制、迁移或重写密钥文件。
|
||||
|
||||
## 缺失处理
|
||||
|
||||
Preflight 返回凭证缺失或目标知识库无法解析时:
|
||||
|
||||
1. 停止上传;
|
||||
2. 只报告缺失的变量名或配置项;
|
||||
3. 等待用户或运行环境补齐配置;
|
||||
4. 配置恢复后重新执行 preflight,不重复生成知识笔记。
|
||||
|
||||
## 安全边界
|
||||
|
||||
- 不在聊天、日志或命令输出中打印完整 API key、KB ID 或 COS 临时凭证;
|
||||
- `create_media` 返回的 COS 凭证仅在同一受控进程内传给上传工具,不写入磁盘;
|
||||
- 不通过拼接 Shell 字符串传递凭证,使用参数数组或 IMA Skill 的封装;
|
||||
- 不绕过 Hermes 的脱敏机制;若现有工具链无法安全传递凭证,停止并报告;
|
||||
- 上传结束后不持久化 COS 临时凭证。
|
||||
@@ -0,0 +1,47 @@
|
||||
# IMA Markdown 格式速查
|
||||
|
||||
## 文件与标题
|
||||
|
||||
- 文件名:`<完整文章标题>.md`
|
||||
- `add_knowledge.title`:完整文章标题,不包含 `.md`
|
||||
- 上传类型:Markdown 文件,`media_type=7`
|
||||
- 目标:`daily` knowledge base
|
||||
|
||||
## 内容来源
|
||||
|
||||
只能使用当前 Run 的可追踪内容:
|
||||
|
||||
1. extracted Artifact 的 `article.plain_text`;
|
||||
2. 对应文章的结构化摘要;
|
||||
3. digest brief 的 summary 与 highlights。
|
||||
|
||||
不得使用通用知识补写原文没有的信息,不设置固定字数或字节数门槛。内容较短时保持简洁并忠于来源。
|
||||
|
||||
## 标准结构
|
||||
|
||||
```markdown
|
||||
# 完整文章标题
|
||||
|
||||
Source: https://原文链接
|
||||
Category: 分类
|
||||
|
||||
## 核心结论
|
||||
|
||||
## 主要论点
|
||||
|
||||
## 关键方法 / 机制
|
||||
|
||||
## 重要细节
|
||||
|
||||
## 可复用启发
|
||||
|
||||
## 关键词
|
||||
|
||||
## 主题
|
||||
```
|
||||
|
||||
- 核心结论和主要论点使用连贯段落;
|
||||
- 方法、细节和启发按完整知识点分项;
|
||||
- 没有来源支持的 Section 可以简写,不得编造内容填充。
|
||||
|
||||
上传 API 见 `ima-upload-api.md`。
|
||||
@@ -0,0 +1,126 @@
|
||||
# IMA Markdown 上传 API
|
||||
|
||||
用于将用户选中的单篇 Markdown 知识笔记上传到 `daily` knowledge base。
|
||||
|
||||
## 凭证
|
||||
|
||||
- `IMA_OPENAPI_CLIENTID`
|
||||
- `IMA_OPENAPI_APIKEY`
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_ID`
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_NAME`
|
||||
|
||||
凭证定位和恢复见 `ima-credential-chain.md`。不要在终端输出完整密钥。
|
||||
|
||||
## 上传前检查
|
||||
|
||||
- 文件名为 `<完整文章标题>.md`;
|
||||
- `title` 为完整文章标题,不带 `.md`;
|
||||
- Markdown 符合 `ima-format-quickref.md`;
|
||||
- 内容可追溯到当前 Run Artifact;
|
||||
- 用户已经明确选择该文章;
|
||||
- 目标知识库已经解析并验证。
|
||||
|
||||
## 1. Preflight
|
||||
|
||||
调用 IMA Skill 的 `preflight-check.cjs` 检查文件类型、扩展名、大小和 MIME。
|
||||
|
||||
预期:
|
||||
|
||||
```text
|
||||
file_ext=md
|
||||
content_type=text/markdown
|
||||
media_type=7
|
||||
```
|
||||
|
||||
### ⚠️ IMA skill 版本拦截(-200)
|
||||
|
||||
`ima_api.cjs` 每天首次调用会检查更新,若检测到新版(如 1.1.8 > 当前 1.1.7)会以 `code=-200` 拦截原请求。注意:**官方 zip 包内的 `meta.json` 可能没同步版本号**(下载 1.1.8 zip 后 meta 仍写 1.1.7),所以光替换文件无法跳过拦截。
|
||||
|
||||
快速修复(脚本本身已是新版,只差版本号):
|
||||
|
||||
```bash
|
||||
cd /root/.hermes/skills/openclaw-imports/ima-skill && python3 -c "
|
||||
import json
|
||||
m = json.load(open('meta.json')); m['version'] = '1.1.8'
|
||||
json.dump(m, open('meta.json','w'), ensure_ascii=False, indent=2)
|
||||
"
|
||||
```
|
||||
|
||||
先用 `diff -rq` 对比 zip 与安装目录:若只有 `.DS_Store`/meta 差异,说明代码已是最新,直接改 meta.json 版本号即可;若脚本有实质差异才需要整体替换。
|
||||
|
||||
## 2. Create Media
|
||||
|
||||
```text
|
||||
POST /openapi/wiki/v1/create_media
|
||||
```
|
||||
|
||||
请求核心字段:
|
||||
|
||||
```json
|
||||
{
|
||||
"file_name": "<完整文章标题>.md",
|
||||
"file_size": 0,
|
||||
"content_type": "text/markdown",
|
||||
"knowledge_base_id": "<daily-kb-id>",
|
||||
"file_ext": "md"
|
||||
}
|
||||
```
|
||||
|
||||
保存返回的 `media_id` 和 `cos_credential`。COS 临时凭证只在进程内传递,不打印到聊天或日志。
|
||||
|
||||
## 3. COS Upload
|
||||
|
||||
使用 IMA Skill 提供的 `cos-upload.cjs`,通过参数数组调用并检查:
|
||||
|
||||
- 进程 `returncode`;
|
||||
- `stderr`;
|
||||
- HTTP 上传结果。
|
||||
|
||||
不要拼接包含凭证的 Shell 字符串,也不要把多条 JSON 响应重定向到同一个文件。
|
||||
|
||||
## 4. Add Knowledge
|
||||
|
||||
```text
|
||||
POST /openapi/wiki/v1/add_knowledge
|
||||
```
|
||||
|
||||
核心字段:
|
||||
|
||||
```json
|
||||
{
|
||||
"media_type": 7,
|
||||
"media_id": "<media-id>",
|
||||
"title": "<完整文章标题>",
|
||||
"knowledge_base_id": "<daily-kb-id>",
|
||||
"file_info": {
|
||||
"cos_key": "<cos-key>",
|
||||
"file_size": 0,
|
||||
"file_name": "<完整文章标题>.md"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## 5. 验证
|
||||
|
||||
上传成功不能只依据本地命令退出码。验证目标 knowledge base 中存在对应标题或返回对象,并向用户报告最终结果。
|
||||
|
||||
失败时明确报告发生在 Preflight、Create Media、COS Upload 或 Add Knowledge 的哪一步,不改用 URL 导入或 Notes 类型绕过错误。
|
||||
|
||||
## 6. 已知坑:IMA skill 版本拦截(-200)
|
||||
|
||||
`ima_api.cjs` 每天首次调用会检查远端版本,若发现新版本(如 1.1.8 > 1.1.7)会以 exit=1 + stderr `{"code":-200}` 拦截**所有** API 调用,原请求不发送。此前遇到过。
|
||||
|
||||
处理方式(不必整包替换):
|
||||
|
||||
1. 按 stderr 提示下载新版 zip(如 `https://app-dl.ima.qq.com/skills/ima-skills-1.1.8.zip`)并解压;
|
||||
2. 对比新旧 `ima_api.cjs` 的 md5——zip 内核心脚本常与本地一致,只是 `meta.json` 的 `version` 未同步(zip 内仍写 1.1.7);
|
||||
3. 若 `ima_api.cjs` 一致,只需把本地 `meta.json` 的 `version` 改为远端版本号即可跳过拦截,无需替换文件。
|
||||
|
||||
调用成功后再执行本文件前面的上传流程。
|
||||
|
||||
## 6. 版本拦截与批量上传实测(2026-07-31)
|
||||
|
||||
- **`-200` skill 更新拦截**:`ima_api.cjs` 每天首次调用检查版本,发现新版时以 code -200 退出并提示更新。下载 zip 后**先对比 `ima_api.cjs` 的 md5**——实测 zip 内脚本与已装版本完全一致,只是 `meta.json` 版本号未同步。此时只需把 `~/.hermes/skills/openclaw-imports/ima-skill/meta.json` 的 `version` 改为最新版即可跳过拦截,无需替换任何脚本。
|
||||
- **Python 脚本编排上传**比 bash 可靠:bash 拼接含中文文件名/凭证的 curl 易出错。用 `subprocess` 参数数组依次调 `preflight-check.cjs` → `ima_api.cjs check_repeated_names` → `create_media` → `cos-upload.cjs`(`--secret-id/--secret-key/--token` 走参数数组,不打印)→ `add_knowledge`,每步解析返回 JSON,失败即停。
|
||||
- **批量上传**:4 篇逐个跑同一脚本即可;同名文件先 `check_repeated_names` 确认无重复。
|
||||
- 凭证从 `IMA_OPENAPI_CLIENTID` / `IMA_OPENAPI_APIKEY` 环境变量读取(`ima_api.cjs` 自动加载),KB ID 用 `IMA_DAILY_KNOWLEDGE_BASE_ID`。
|
||||
@@ -0,0 +1,29 @@
|
||||
# 关键词治理路由
|
||||
|
||||
关键词治理不属于 `reader-digest-flow` 的日常执行阶段。
|
||||
|
||||
仅当用户明确要求“清理关键词”“词库治理”“生成关键词建议”时,委托:
|
||||
|
||||
```text
|
||||
skills/keyword-cleanup-review/SKILL.md
|
||||
```
|
||||
|
||||
Review 输入生成由该 Skill 定义,后续确认与 Apply 由当前编排层负责:
|
||||
|
||||
```text
|
||||
build review bundle
|
||||
→ generate rule suggestions
|
||||
→ optional semantic suggestions
|
||||
→ human review
|
||||
→ dry-run
|
||||
→ apply accepted suggestions
|
||||
```
|
||||
|
||||
约束:
|
||||
|
||||
- Suggestions JSON 是 Review 与 Apply 之间的正式契约;
|
||||
- LLM 语义建议不能自动 Apply;
|
||||
- 不直接编辑 aliases、stopwords、watchlist 或 interest 配置;
|
||||
- 不在日报主流程中因 tag 质量不佳自动触发治理。
|
||||
|
||||
Review bundle、Suggestions、Schema 和产物保留策略以 `keyword-cleanup-review` 为唯一事实来源;该 Skill 不直接 Apply 配置。
|
||||
@@ -0,0 +1,31 @@
|
||||
+++
|
||||
title = "AI 日报 · 示例"
|
||||
date = 2026-04-01T09:00:00+08:00
|
||||
summary = "围绕 Agent 架构分层、Skills 标准化与工程实践的当日观察。"
|
||||
+++
|
||||
|
||||
# 今日概览
|
||||
|
||||
今天的公开内容主要集中在 AI Agent 架构演进、工具化落地与工程实践三条线索。行业关注点正在从“模型能做什么”转向“系统如何稳定落地并持续复用”。
|
||||
|
||||
## 今日重点
|
||||
|
||||
### 1. 从 Agent 到 Skills:AI 智能体架构的范式转变
|
||||
|
||||
文章分析了 AI 智能体从单体 Agent 向模块化 Skills 的演进,并结合 MCP、Skills 和真实项目说明能力分层与复用方式。
|
||||
|
||||
值得关注:
|
||||
|
||||
- Skills 将领域流程从 Agent 主体中拆出,便于复用和维护。
|
||||
- MCP 为 Agent 与外部工具提供标准化连接方式。
|
||||
- 工程竞争点逐渐从模型调用转向状态、工具和工作流设计。
|
||||
|
||||
这篇内容值得关注的原因在于,它把开放协议、分层架构和真实落地案例连接成了完整论证链。
|
||||
|
||||
来源:[示例来源](https://example.com/a)
|
||||
|
||||
## 趋势观察
|
||||
|
||||
1. Agent 正在从单体能力转向可组合的模块化体系。
|
||||
2. 工具契约、状态管理和验证机制正在成为 AI 应用的核心工程能力。
|
||||
3. Human-in-the-loop 仍是控制高风险副作用的重要边界。
|
||||
@@ -0,0 +1,96 @@
|
||||
#!/usr/bin/env python3
|
||||
"""IMA 上传单篇 Markdown 知识笔记到 daily 知识库。
|
||||
用法: python3 ima_upload_one.py "<绝对路径/summary.md>" "<完整文章标题>"
|
||||
依赖环境变量: IMA_OPENAPI_CLIENTID / IMA_OPENAPI_APIKEY / IMA_DAILY_KNOWLEDGE_BASE_ID
|
||||
流程: preflight -> check_repeated_names -> create_media -> cos-upload -> add_knowledge
|
||||
说明: 用 Python 而非 bash 编排,避免中文文件名/引号转义问题。
|
||||
退出码 2 = 文件名重复(需与用户确认保留双方或取消),非 0 均为失败。
|
||||
"""
|
||||
import json, os, subprocess, sys
|
||||
|
||||
SKILL_DIR = "/root/.hermes/skills/openclaw-imports/ima-skill"
|
||||
IMA_API = os.path.join(SKILL_DIR, "ima_api.cjs")
|
||||
COS_UPLOAD = os.path.join(SKILL_DIR, "knowledge-base/scripts/cos-upload.cjs")
|
||||
PREFLIGHT = os.path.join(SKILL_DIR, "knowledge-base/scripts/preflight-check.cjs")
|
||||
|
||||
def run_node(script, args):
|
||||
r = subprocess.run(["node", script] + args, capture_output=True, text=True, timeout=120)
|
||||
if r.returncode != 0:
|
||||
raise RuntimeError(f"{script} exit={r.returncode} stderr={r.stderr[:500]}")
|
||||
return json.loads(r.stdout)
|
||||
|
||||
def ima_api(api_path, body):
|
||||
r = subprocess.run(["node", IMA_API, api_path, json.dumps(body, ensure_ascii=False)],
|
||||
capture_output=True, text=True, timeout=120)
|
||||
if r.returncode != 0:
|
||||
raise RuntimeError(f"ima_api {api_path} exit={r.returncode} stderr={r.stderr[:500]}")
|
||||
resp = json.loads(r.stdout)
|
||||
if resp.get("code") != 0:
|
||||
raise RuntimeError(f"ima_api {api_path} code={resp.get('code')} msg={resp.get('msg')}")
|
||||
return resp.get("data", {})
|
||||
|
||||
def main():
|
||||
kb_id = os.environ["IMA_DAILY_KNOWLEDGE_BASE_ID"]
|
||||
file_path = sys.argv[1]
|
||||
title = sys.argv[2] # 完整文章标题(不含 .md)
|
||||
|
||||
pf = run_node(PREFLIGHT, ["--file", file_path])
|
||||
if not pf.get("pass"):
|
||||
raise RuntimeError(f"preflight failed: {pf}")
|
||||
file_name = pf["file_name"]; media_type = pf["media_type"]
|
||||
content_type = pf["content_type"]; file_size = pf["file_size"]; file_ext = pf["file_ext"]
|
||||
print(f"[preflight] ok file={file_name} ext={file_ext} size={file_size} media_type={media_type}")
|
||||
|
||||
dup = ima_api("openapi/wiki/v1/check_repeated_names", {
|
||||
"params": [{"name": file_name, "media_type": media_type}],
|
||||
"knowledge_base_id": kb_id
|
||||
})
|
||||
is_rep = dup.get("results", [{}])[0].get("is_repeated", False) if dup.get("results") else False
|
||||
if is_rep:
|
||||
print(f"[check_repeated_names] REPEATED: {file_name} — 需要处理")
|
||||
sys.exit(2)
|
||||
print("[check_repeated_names] no duplicate")
|
||||
|
||||
cm = ima_api("openapi/wiki/v1/create_media", {
|
||||
"file_name": file_name,
|
||||
"file_size": file_size,
|
||||
"content_type": content_type,
|
||||
"knowledge_base_id": kb_id,
|
||||
"file_ext": file_ext
|
||||
})
|
||||
media_id = cm["media_id"]
|
||||
cos = cm["cos_credential"]
|
||||
print(f"[create_media] media_id={media_id} cos_bucket={cos.get('bucket_name')} cos_key={cos.get('cos_key','')[:40]}")
|
||||
|
||||
r = subprocess.run(["node", COS_UPLOAD,
|
||||
"--file", file_path,
|
||||
"--secret-id", cos["secret_id"],
|
||||
"--secret-key", cos["secret_key"],
|
||||
"--token", cos["token"],
|
||||
"--bucket", cos["bucket_name"],
|
||||
"--region", cos["region"],
|
||||
"--cos-key", cos["cos_key"],
|
||||
"--content-type", content_type,
|
||||
"--start-time", str(cos.get("start_time", "")),
|
||||
"--expired-time", str(cos.get("expired_time", "")),
|
||||
"--timeout", "300000"
|
||||
], capture_output=True, text=True, timeout=360)
|
||||
if r.returncode != 0:
|
||||
raise RuntimeError(f"cos-upload exit={r.returncode} stderr={r.stderr[:800]}")
|
||||
print(f"[cos-upload] ok rc=0 stdout={r.stdout.strip()[:200]}")
|
||||
|
||||
ak = ima_api("openapi/wiki/v1/add_knowledge", {
|
||||
"media_type": media_type,
|
||||
"media_id": media_id,
|
||||
"title": title,
|
||||
"knowledge_base_id": kb_id,
|
||||
"file_info": {
|
||||
"cos_key": cos["cos_key"],
|
||||
"file_size": file_size,
|
||||
"file_name": file_name
|
||||
}
|
||||
})
|
||||
print(f"[add_knowledge] ok media_id={ak.get('media_id') or media_id}")
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
@@ -1,5 +1,8 @@
|
||||
from __future__ import annotations
|
||||
|
||||
# 日报级关键词索引:从当日 delivery 候选中提取关键词,
|
||||
# 应用别名归并和停用词过滤,输出 DailyKeywordIndex 并聚合全局 KeywordStatsIndex。
|
||||
|
||||
import json
|
||||
from collections import Counter
|
||||
from datetime import date, datetime, timezone
|
||||
@@ -85,6 +88,7 @@ def normalize_keyword(
|
||||
aliases: dict[str, str],
|
||||
stopwords: set[str],
|
||||
) -> str | None:
|
||||
# 对单个关键词做别名替换 + 停用词过滤,返回 None 表示丢弃
|
||||
raw_value = keyword.strip()
|
||||
if not raw_value:
|
||||
return None
|
||||
@@ -202,6 +206,7 @@ def persist_keyword_indexes(
|
||||
stopwords_path: Path | None = None,
|
||||
include_decisions: set[str] | None = None,
|
||||
) -> dict[str, object]:
|
||||
# 构建并写入当日词元索引,同时全量重建全局词频统计;每次 pipeline 运行后自动调用
|
||||
aliases = load_term_aliases(aliases_path)
|
||||
stopwords = load_term_stopwords(stopwords_path)
|
||||
daily_index = build_daily_keyword_index(
|
||||
|
||||
@@ -1,5 +1,8 @@
|
||||
from __future__ import annotations
|
||||
|
||||
# 内容提取主流程:根据 item 来源决定使用内联 RSS 内容还是回源抓取,
|
||||
# 再经过文本提取、质量评估、文章构建,输出标准化 ExtractionOutput。
|
||||
|
||||
from summary_mcp.core.content_loader import choose_inline_content, fetch_html
|
||||
from summary_mcp.core.errors import SummaryError
|
||||
from summary_mcp.core.extractor import extract_plain_text, extract_title
|
||||
@@ -9,10 +12,12 @@ from summary_mcp.core.quality_checker import assess_quality
|
||||
from summary_mcp.models.summary_io import DebugInfo, ExtractionInput, ExtractionOutput
|
||||
|
||||
|
||||
# FreshRSS 来源的 item 不回源抓网页,直接使用 RSS 内联内容
|
||||
RSS_ONLY_UPSTREAMS = {"freshrss"}
|
||||
|
||||
|
||||
def _should_skip_fetch(extraction_input: ExtractionInput) -> bool:
|
||||
# 判断当前 item 是否属于 RSS-only 上游,若是则禁止回源抓取
|
||||
item = extraction_input.item
|
||||
if item is None:
|
||||
return False
|
||||
|
||||
@@ -1,12 +1,16 @@
|
||||
from __future__ import annotations
|
||||
|
||||
# LLM 摘要循环:构造提示词 -> 调用 LLM -> 解析 JSON -> 校验 -> 失败时生成修复提示词重试。
|
||||
# 核心入口:run_loop_payload(),最多重试 max_retries 次。
|
||||
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
from typing import Any, Callable
|
||||
|
||||
import httpx
|
||||
from dotenv import dotenv_values
|
||||
|
||||
from summary_mcp.validators.llm_result import ValidationReport
|
||||
from summary_mcp.validators.llm_result import validate_llm_result as validate_llm_result_from_path
|
||||
@@ -15,6 +19,8 @@ from summary_mcp.validators.llm_result import validate_llm_result_payload
|
||||
|
||||
JSON_BLOCK_RE = re.compile(r"```(?:json)?\s*(\{.*\})\s*```", re.DOTALL)
|
||||
DEFAULT_CHAT_COMPLETIONS_URL = "https://api.openai.com/v1/chat/completions"
|
||||
REPO_ROOT = Path(__file__).resolve().parents[3]
|
||||
DEFAULT_DOTENV_PATH = REPO_ROOT / ".env"
|
||||
|
||||
|
||||
def load_text(path: Path) -> str:
|
||||
@@ -96,22 +102,60 @@ def normalize_chat_completions_url(api_url: str | None) -> str | None:
|
||||
return f"{normalized}/chat/completions"
|
||||
|
||||
|
||||
def _load_repo_dotenv() -> dict[str, str]:
|
||||
if not DEFAULT_DOTENV_PATH.exists():
|
||||
return {}
|
||||
return {
|
||||
key: value
|
||||
for key, value in dotenv_values(DEFAULT_DOTENV_PATH).items()
|
||||
if isinstance(key, str) and isinstance(value, str) and value
|
||||
}
|
||||
|
||||
|
||||
def _pick_value(*values: str | None) -> str | None:
|
||||
for value in values:
|
||||
if value:
|
||||
return value
|
||||
return None
|
||||
|
||||
|
||||
def resolve_llm_settings(
|
||||
*,
|
||||
api_key: str | None = None,
|
||||
model: str | None = None,
|
||||
api_url: str | None = None,
|
||||
) -> tuple[str, str, str]:
|
||||
resolved_api_key = api_key or os.environ.get("LLM_API_KEY") or os.environ.get("OPENAI_API_KEY")
|
||||
dotenv_map = _load_repo_dotenv()
|
||||
|
||||
resolved_api_key = _pick_value(
|
||||
api_key,
|
||||
os.environ.get("LLM_API_KEY"),
|
||||
os.environ.get("OPENAI_API_KEY"),
|
||||
dotenv_map.get("LLM_API_KEY"),
|
||||
dotenv_map.get("OPENAI_API_KEY"),
|
||||
)
|
||||
if not resolved_api_key:
|
||||
raise RuntimeError("Missing LLM_API_KEY or OPENAI_API_KEY, or pass an API key.")
|
||||
|
||||
resolved_model = model or os.environ.get("LLM_MODEL") or os.environ.get("OPENAI_MODEL")
|
||||
resolved_model = _pick_value(
|
||||
model,
|
||||
os.environ.get("LLM_MODEL"),
|
||||
os.environ.get("OPENAI_MODEL"),
|
||||
dotenv_map.get("LLM_MODEL"),
|
||||
dotenv_map.get("OPENAI_MODEL"),
|
||||
)
|
||||
if not resolved_model:
|
||||
raise RuntimeError("Missing LLM_MODEL or OPENAI_MODEL, or pass a model.")
|
||||
|
||||
resolved_api_url = normalize_chat_completions_url(
|
||||
api_url or os.environ.get("LLM_API_URL") or os.environ.get("OPENAI_API_URL") or DEFAULT_CHAT_COMPLETIONS_URL
|
||||
_pick_value(
|
||||
api_url,
|
||||
os.environ.get("LLM_API_URL"),
|
||||
os.environ.get("OPENAI_API_URL"),
|
||||
dotenv_map.get("LLM_API_URL"),
|
||||
dotenv_map.get("OPENAI_API_URL"),
|
||||
DEFAULT_CHAT_COMPLETIONS_URL,
|
||||
)
|
||||
)
|
||||
if not resolved_api_url:
|
||||
raise RuntimeError("Missing LLM_API_URL, OPENAI_API_URL, or pass an API URL.")
|
||||
@@ -180,7 +224,10 @@ def run_loop_payload(
|
||||
model: str | None,
|
||||
api_url: str | None,
|
||||
output_path: Path | None = None,
|
||||
validator: Callable[[dict[str, Any], dict[str, Any] | None], ValidationReport] | None = None,
|
||||
) -> tuple[int, dict[str, Any] | None, ValidationReport | None]:
|
||||
# validator 为 None 时使用日报摘要默认校验器;传入自定义 validator 时使用传入的
|
||||
resolved_validator = validator if validator is not None else validate_llm_result_payload
|
||||
prompt_template = load_text(prompt_path)
|
||||
if output_path is not None:
|
||||
output_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
@@ -220,7 +267,7 @@ def run_loop_payload(
|
||||
if output_path is not None:
|
||||
save_json(output_path, result_payload)
|
||||
|
||||
report = validate_llm_result_payload(result_payload, extracted_payload)
|
||||
report = resolved_validator(result_payload, extracted_payload)
|
||||
_save_attempt_artifact(
|
||||
output_path,
|
||||
f"attempt-{attempt}.validation.json",
|
||||
|
||||
@@ -1,5 +1,8 @@
|
||||
from __future__ import annotations
|
||||
|
||||
# 确定性规则引擎:从 JSON 配置加载规则,按优先级评估每条规则的条件,
|
||||
# 输出 keep / review / drop 决策及命中规则列表。
|
||||
|
||||
import json
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
@@ -32,6 +35,7 @@ def _normalize_value(value: Any) -> Any:
|
||||
|
||||
|
||||
def _resolve_field(filter_input: FilterInput, field_path: str) -> Any:
|
||||
# 按点分路径从 FilterInput 中取值,路径不存在时返回 None
|
||||
current: Any = filter_input
|
||||
for part in field_path.split("."):
|
||||
current = _normalize_value(current)
|
||||
@@ -94,6 +98,8 @@ def _rule_matches(filter_input: FilterInput, rule: FilterRule) -> bool:
|
||||
|
||||
|
||||
def evaluate_filter_rules(filter_input: FilterInput, rules: list[FilterRule]) -> FilterDecisionResult:
|
||||
# 按优先级排序后逐条评估规则;stop_on_match=True 时命中即终止
|
||||
# 最终决策优先级:drop > keep > review;无规则命中时默认 review
|
||||
matched: list[MatchedRule] = []
|
||||
|
||||
for rule in sorted(rules, key=lambda item: item.action.priority, reverse=True):
|
||||
|
||||
@@ -1,5 +1,8 @@
|
||||
from __future__ import annotations
|
||||
|
||||
# FreshRSS GReader API 集成:负责登录鉴权、拉取未读条目、将 entry 映射为标准 item、
|
||||
# 以及成功投递后将条目标记为已读。
|
||||
|
||||
import hashlib
|
||||
from datetime import UTC, datetime
|
||||
from typing import Any
|
||||
@@ -12,6 +15,7 @@ from summary_mcp.models.item import Item
|
||||
|
||||
READ_TAG = "user/-/state/com.google/read"
|
||||
KEPT_UNREAD_TAG = "user/-/state/com.google/kept-unread"
|
||||
# RSS 内联内容低于此长度时视为无效,item 将被标记为 pending 并跳过
|
||||
MIN_INLINE_CONTENT_LENGTH = 500
|
||||
|
||||
|
||||
@@ -124,6 +128,8 @@ def _has_usable_inline_content(value: str | None) -> bool:
|
||||
|
||||
|
||||
def map_entry_to_item(entry: dict[str, Any]) -> Item:
|
||||
# 将 FreshRSS GReader API 返回的原始 entry dict 映射为标准化 Item 对象
|
||||
# fetch_state: 若内联内容足够长则为 'fetched',否则为 'pending'
|
||||
url = _pick_entry_url(entry)
|
||||
summary_html = _pick_content_block(entry, "summary")
|
||||
content_html = _pick_content_block(entry, "content")
|
||||
@@ -222,6 +228,7 @@ class FreshRSSClient:
|
||||
params: list[tuple[str, str | int]] = [
|
||||
("output", "json"),
|
||||
("n", limit),
|
||||
("r", "n")
|
||||
]
|
||||
if continuation:
|
||||
params.append(("c", continuation))
|
||||
|
||||
@@ -26,7 +26,9 @@ from .keyword_index import DailyKeywordIndex, DailyKeywordTerm, KeywordStat, Key
|
||||
from .openclaw_delivery import (
|
||||
OpenClawDeliveryPayload,
|
||||
OpenClawDeliveryStats,
|
||||
OpenClawDigestBrief,
|
||||
build_openclaw_delivery_payload,
|
||||
build_openclaw_digest_brief,
|
||||
)
|
||||
|
||||
__all__ = [
|
||||
@@ -50,10 +52,12 @@ __all__ = [
|
||||
"OpenClawCandidateInput",
|
||||
"OpenClawDeliveryPayload",
|
||||
"OpenClawDeliveryStats",
|
||||
"OpenClawDigestBrief",
|
||||
"ReviewState",
|
||||
"ReviewStatus",
|
||||
"build_article_candidate_record",
|
||||
"build_openclaw_delivery_payload",
|
||||
"build_openclaw_digest_brief",
|
||||
"build_openclaw_candidate_input",
|
||||
"candidate_id_for",
|
||||
"normalize_candidate_url",
|
||||
|
||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
@@ -66,6 +66,7 @@ class ArticleCandidateRecord(BaseModel):
|
||||
|
||||
class OpenClawCandidateInput(BaseModel):
|
||||
candidate_id: str
|
||||
item_id: str | None = None
|
||||
title: str
|
||||
url: HttpUrl
|
||||
canonical_url: HttpUrl | None = None
|
||||
@@ -169,6 +170,7 @@ def build_openclaw_candidate_input(record: ArticleCandidateRecord) -> OpenClawCa
|
||||
|
||||
return OpenClawCandidateInput(
|
||||
candidate_id=record.candidate_id,
|
||||
item_id=item.item_id if item is not None else None,
|
||||
title=title,
|
||||
url=raw_url,
|
||||
canonical_url=normalize_candidate_url(raw_url),
|
||||
|
||||
@@ -0,0 +1,25 @@
|
||||
from __future__ import annotations
|
||||
|
||||
# 单篇文章深度沉淀的结构化输出模型,与日报摘要的 LlmSummaryResult 完全独立。
|
||||
|
||||
from typing import Literal
|
||||
|
||||
from pydantic import BaseModel, HttpUrl
|
||||
|
||||
|
||||
ArticleSummaryCategory = Literal["资讯", "方法论", "工具实践", "观点评论"]
|
||||
|
||||
|
||||
class ArticleSummaryResult(BaseModel):
|
||||
title: str
|
||||
url: HttpUrl
|
||||
core_conclusion: str
|
||||
main_argument: str
|
||||
key_methods: list[str] = []
|
||||
important_details: list[str] = []
|
||||
reusable_insights: list[str] = []
|
||||
keywords: list[str] = []
|
||||
topics: list[str] = []
|
||||
category: ArticleSummaryCategory
|
||||
worth_keeping: bool
|
||||
reason: str
|
||||
@@ -11,7 +11,7 @@ Category = Literal["资讯", "方法论", "工具实践", "观点评论"]
|
||||
class LlmSummaryResult(BaseModel):
|
||||
title: str = Field(min_length=1)
|
||||
url: HttpUrl
|
||||
summary: str = Field(min_length=20, max_length=140)
|
||||
summary: str = Field(min_length=20, max_length=300)
|
||||
highlights: list[str] = Field(min_length=3, max_length=5)
|
||||
keywords: list[str] = Field(min_length=5, max_length=8)
|
||||
topics: list[str] = Field(min_length=3, max_length=5)
|
||||
|
||||
@@ -1,11 +1,19 @@
|
||||
from __future__ import annotations
|
||||
|
||||
from datetime import UTC, date, datetime
|
||||
try:
|
||||
from datetime import UTC, date, datetime
|
||||
except ImportError: # Python < 3.11 compatibility
|
||||
from datetime import timezone, date, datetime
|
||||
|
||||
UTC = timezone.utc
|
||||
|
||||
from pydantic import BaseModel, Field
|
||||
|
||||
from .article_candidate import OpenClawCandidateInput
|
||||
|
||||
DIGEST_BRIEF_LIMIT: int | None = None
|
||||
DIGEST_BRIEF_HIGHLIGHT_LIMIT = 3
|
||||
|
||||
|
||||
class OpenClawDeliveryStats(BaseModel):
|
||||
total: int = Field(default=0, ge=0)
|
||||
@@ -23,6 +31,27 @@ class OpenClawDeliveryPayload(BaseModel):
|
||||
stats: OpenClawDeliveryStats = Field(default_factory=OpenClawDeliveryStats)
|
||||
|
||||
|
||||
class DigestBriefCandidate(BaseModel):
|
||||
title: str
|
||||
source_name: str | None = None
|
||||
summary: str
|
||||
highlights: list[str] = Field(default_factory=list)
|
||||
category: str | None = None
|
||||
digest_rank: int = Field(default=0)
|
||||
selection_decision: str
|
||||
url: str
|
||||
|
||||
|
||||
class OpenClawDigestBrief(BaseModel):
|
||||
schema_version: str = "digest-brief.v1"
|
||||
generated_at: datetime
|
||||
run_id: str
|
||||
date: date
|
||||
source_candidate_count: int = Field(default=0, ge=0)
|
||||
candidate_count: int = Field(default=0, ge=0)
|
||||
top_candidates: list[DigestBriefCandidate] = Field(default_factory=list)
|
||||
|
||||
|
||||
def build_openclaw_delivery_payload(
|
||||
candidates: list[OpenClawCandidateInput],
|
||||
*,
|
||||
@@ -47,3 +76,36 @@ def build_openclaw_delivery_payload(
|
||||
candidates=candidates,
|
||||
stats=stats,
|
||||
)
|
||||
|
||||
|
||||
def build_openclaw_digest_brief(
|
||||
payload: OpenClawDeliveryPayload,
|
||||
*,
|
||||
limit: int | None = DIGEST_BRIEF_LIMIT,
|
||||
highlight_limit: int = DIGEST_BRIEF_HIGHLIGHT_LIMIT,
|
||||
) -> OpenClawDigestBrief:
|
||||
keep_candidates = [candidate for candidate in payload.candidates if candidate.selection_decision in ("keep", "review")]
|
||||
sorted_candidates = sorted(keep_candidates, key=lambda candidate: candidate.digest_rank, reverse=True)
|
||||
selected_candidates = sorted_candidates if limit is None else sorted_candidates[:limit]
|
||||
top_candidates = [
|
||||
DigestBriefCandidate(
|
||||
title=candidate.title,
|
||||
source_name=candidate.source_name,
|
||||
summary=candidate.summary,
|
||||
highlights=list(candidate.highlights[:highlight_limit]),
|
||||
category=candidate.category,
|
||||
digest_rank=candidate.digest_rank,
|
||||
selection_decision=candidate.selection_decision,
|
||||
url=str(candidate.url),
|
||||
)
|
||||
for candidate in selected_candidates
|
||||
]
|
||||
return OpenClawDigestBrief(
|
||||
schema_version="digest-brief.v1",
|
||||
generated_at=payload.generated_at,
|
||||
run_id=payload.run_id,
|
||||
date=payload.date,
|
||||
source_candidate_count=len(payload.candidates),
|
||||
candidate_count=len(top_candidates),
|
||||
top_candidates=top_candidates,
|
||||
)
|
||||
|
||||
@@ -0,0 +1,54 @@
|
||||
from .run_store import RunStore
|
||||
from .state_models import ArtifactRecord, RecoveryState, RunError, RunState, StageState
|
||||
from .freshrss_pipeline_jobs import (
|
||||
get_freshrss_pipeline_job_result,
|
||||
get_freshrss_pipeline_job_status,
|
||||
start_freshrss_pipeline_job,
|
||||
)
|
||||
from .query_service import get_delivery_payload, get_run_report, get_run_status, list_run_artifacts, list_runs
|
||||
|
||||
# NOTE: resume_jobs and resume_service are NOT eagerly imported here to avoid
|
||||
# a circular import chain:
|
||||
# workflows/freshrss_pipeline.py -> runtime -> resume_jobs -> resume_service
|
||||
# -> workflows/freshrss_pipeline.py (circular!)
|
||||
# They are lazy-loaded via __getattr__ when accessed as summary_mcp.runtime.*
|
||||
|
||||
|
||||
def __getattr__(name):
|
||||
import importlib
|
||||
|
||||
_LAZY = {
|
||||
"get_resume_job_result": ("resume_jobs", "get_resume_job_result"),
|
||||
"get_resume_job_status": ("resume_jobs", "get_resume_job_status"),
|
||||
"start_resume_job": ("resume_jobs", "start_resume_job"),
|
||||
"inspect_resume_plan": ("resume_service", "inspect_resume_plan"),
|
||||
"resume_run": ("resume_service", "resume_run"),
|
||||
}
|
||||
if name in _LAZY:
|
||||
mod_name, attr_name = _LAZY[name]
|
||||
mod = importlib.import_module(f".{mod_name}", __package__)
|
||||
return getattr(mod, attr_name)
|
||||
raise AttributeError(f"module {__name__!r} has no attribute {name!r}")
|
||||
|
||||
|
||||
__all__ = [
|
||||
"ArtifactRecord",
|
||||
"get_delivery_payload",
|
||||
"get_freshrss_pipeline_job_result",
|
||||
"get_freshrss_pipeline_job_status",
|
||||
"get_resume_job_result",
|
||||
"get_resume_job_status",
|
||||
"get_run_report",
|
||||
"RecoveryState",
|
||||
"RunError",
|
||||
"RunState",
|
||||
"RunStore",
|
||||
"StageState",
|
||||
"get_run_status",
|
||||
"inspect_resume_plan",
|
||||
"list_run_artifacts",
|
||||
"list_runs",
|
||||
"resume_run",
|
||||
"start_resume_job",
|
||||
"start_freshrss_pipeline_job",
|
||||
]
|
||||
@@ -0,0 +1,346 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import subprocess
|
||||
import sys
|
||||
from concurrent.futures import ThreadPoolExecutor
|
||||
from datetime import datetime
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
from uuid import uuid4
|
||||
|
||||
from .run_store import RunStore
|
||||
from .state_models import RunState
|
||||
|
||||
# 后台线程池:用于异步启动 job,避免阻塞 MCP stdio 响应
|
||||
# max_workers=4 足够应对并发请求,线程复用降低启动开销
|
||||
_executor = ThreadPoolExecutor(max_workers=4, thread_name_prefix="article_summary_job")
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[3]
|
||||
OUTPUT_ROOT = REPO_ROOT / "outputs" / "freshrss"
|
||||
ARTICLE_SUMMARY_JOBS_ROOT = OUTPUT_ROOT / "article_summary_jobs"
|
||||
WORKFLOW_NAME = "article_summary_job"
|
||||
RUN_TYPE = "article_summary"
|
||||
RUN_STATE_FILENAME = "run-state.json"
|
||||
INPUT_FILENAME = "input.json"
|
||||
RESULT_FILENAME = "result.json"
|
||||
JOB_REPORT_FILENAME = "job-report.json"
|
||||
DEFAULT_STAGES = [
|
||||
"prepare_job",
|
||||
"load_input",
|
||||
"generate_markdown",
|
||||
"write_result",
|
||||
]
|
||||
|
||||
|
||||
def _now() -> datetime:
|
||||
return datetime.now().astimezone()
|
||||
|
||||
|
||||
def _new_job_id() -> str:
|
||||
ts = _now().strftime("%Y%m%d-%H%M%S")
|
||||
return f"article-summary-{ts}-{uuid4().hex[:8]}"
|
||||
|
||||
|
||||
def _job_dir(job_id: str) -> Path:
|
||||
return ARTICLE_SUMMARY_JOBS_ROOT / job_id
|
||||
|
||||
|
||||
def _run_state_path(job_id: str) -> Path:
|
||||
return _job_dir(job_id) / RUN_STATE_FILENAME
|
||||
|
||||
|
||||
def _input_path(job_id: str) -> Path:
|
||||
return _job_dir(job_id) / INPUT_FILENAME
|
||||
|
||||
|
||||
def _result_path(job_id: str) -> Path:
|
||||
return _job_dir(job_id) / RESULT_FILENAME
|
||||
|
||||
|
||||
def _job_report_path(job_id: str) -> Path:
|
||||
return _job_dir(job_id) / JOB_REPORT_FILENAME
|
||||
|
||||
|
||||
def _normalize_repo_path(path: Path) -> str:
|
||||
try:
|
||||
return str(path.resolve().relative_to(REPO_ROOT.resolve()))
|
||||
except ValueError:
|
||||
return str(path)
|
||||
|
||||
|
||||
def _load_run_store(job_id: str) -> RunStore:
|
||||
return RunStore.load(path=_run_state_path(job_id), repo_root=REPO_ROOT)
|
||||
|
||||
|
||||
def _load_result(job_id: str) -> dict[str, Any] | None:
|
||||
path = _result_path(job_id)
|
||||
if not path.exists():
|
||||
return None
|
||||
return json.loads(path.read_text(encoding="utf-8-sig"))
|
||||
|
||||
|
||||
def _validate_article_summary_extracted_path(extracted_path: Path) -> None:
|
||||
if not extracted_path.exists():
|
||||
raise FileNotFoundError(f"extracted_path does not exist: {extracted_path}")
|
||||
if extracted_path.is_dir():
|
||||
raise ValueError(
|
||||
"extracted_path must be a JSON file, not a directory. "
|
||||
"Pass a single extracted file like outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json, "
|
||||
"or a batch extracted JSON file."
|
||||
)
|
||||
|
||||
|
||||
def _launch_job_background(*, job_id: str, input_payload: dict[str, Any], store: RunStore) -> None:
|
||||
"""后台线程执行:启动 subprocess 并完成 run_state 写盘。
|
||||
|
||||
主线程已经返回 MCP 响应,后台负责:
|
||||
1. 启动 runner 子进程(Popen)
|
||||
2. 记录 prepare_job 阶段完成
|
||||
3. 持久化 run-state.json
|
||||
"""
|
||||
runner_script = REPO_ROOT / "scripts" / "run_article_summary_job.py"
|
||||
cmd = [sys.executable, str(runner_script), "--job-id", job_id]
|
||||
|
||||
# 启动子进程(后台运行,不等待)
|
||||
proc = subprocess.Popen(
|
||||
cmd,
|
||||
cwd=str(REPO_ROOT),
|
||||
stdout=subprocess.DEVNULL,
|
||||
stderr=subprocess.DEVNULL,
|
||||
start_new_session=True,
|
||||
)
|
||||
|
||||
# 记录 prepare_job 完成
|
||||
store.finish_stage("prepare_job", outputs={"runner_pid": proc.pid, "runner_command": cmd})
|
||||
|
||||
|
||||
def start_article_summary_job(
|
||||
*,
|
||||
extracted_path: Path,
|
||||
selected_ids: list[str],
|
||||
output_dir: Path | None = None,
|
||||
max_retries: int = 2,
|
||||
timeout_seconds: float = 120.0,
|
||||
llm_api_key: str | None = None,
|
||||
llm_model: str | None = None,
|
||||
llm_api_url: str | None = None,
|
||||
) -> dict[str, Any]:
|
||||
"""Start an asynchronous article-summary job and return a job_id immediately.
|
||||
|
||||
关键:使用后台线程异步启动 job,主线程立即返回 MCP 响应,避免 stdio 阻塞。
|
||||
"""
|
||||
_validate_article_summary_extracted_path(extracted_path)
|
||||
if not selected_ids:
|
||||
raise ValueError("selected_ids must not be empty")
|
||||
|
||||
job_id = _new_job_id()
|
||||
job_dir = _job_dir(job_id)
|
||||
job_dir.mkdir(parents=True, exist_ok=True)
|
||||
started_at = _now()
|
||||
|
||||
resolved_output_dir = output_dir or (extracted_path.parent / "single_summaries")
|
||||
input_payload = {
|
||||
"extracted_path": str(extracted_path),
|
||||
"selected_ids": selected_ids,
|
||||
"output_dir": str(resolved_output_dir),
|
||||
"max_retries": max_retries,
|
||||
"timeout_seconds": timeout_seconds,
|
||||
"llm_api_key": llm_api_key,
|
||||
"llm_model": llm_model,
|
||||
"llm_api_url": llm_api_url,
|
||||
"launcher_pid": os.getpid(),
|
||||
}
|
||||
|
||||
# 创建 run-state(只写初始状态,不启动 subprocess)
|
||||
store = RunStore.create(
|
||||
path=_run_state_path(job_id),
|
||||
run_id=job_id,
|
||||
workflow=WORKFLOW_NAME,
|
||||
run_type=RUN_TYPE,
|
||||
started_at=started_at,
|
||||
input_payload=input_payload,
|
||||
repo_root=REPO_ROOT,
|
||||
)
|
||||
for stage_name in DEFAULT_STAGES:
|
||||
store._get_or_create_stage(stage_name)
|
||||
|
||||
# 注册输入文件
|
||||
input_file = _input_path(job_id)
|
||||
input_file.write_text(json.dumps(input_payload, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||
|
||||
# 启动 prepare_job 阶段(只写状态,不调用 save() —— 交给后台线程)
|
||||
store.start_stage("prepare_job")
|
||||
store.register_artifact(name="job_input", path=input_file, kind="json", stage="prepare_job")
|
||||
|
||||
# 【关键改造】:后台线程执行 Popen + finish_stage,主线程立即返回
|
||||
_executor.submit(_launch_job_background, job_id=job_id, input_payload=input_payload, store=store)
|
||||
|
||||
return {
|
||||
"job_id": job_id,
|
||||
"workflow": WORKFLOW_NAME,
|
||||
"run_type": RUN_TYPE,
|
||||
"status": "running",
|
||||
"output_dir": _normalize_repo_path(job_dir),
|
||||
"message": "Article summary job started successfully. Use get_article_summary_job_status to poll progress.",
|
||||
}
|
||||
|
||||
|
||||
def run_article_summary_job(*, job_id: str) -> dict[str, Any]:
|
||||
from summary_mcp.workflows.article_summary import ArticleSummaryConfig, summarize_selected_articles
|
||||
|
||||
store = _load_run_store(job_id)
|
||||
input_payload = json.loads(_input_path(job_id).read_text(encoding="utf-8-sig"))
|
||||
|
||||
try:
|
||||
store.start_stage("load_input")
|
||||
extracted_path = Path(input_payload["extracted_path"])
|
||||
output_dir = Path(input_payload["output_dir"])
|
||||
selected_ids = list(input_payload["selected_ids"])
|
||||
_validate_article_summary_extracted_path(extracted_path)
|
||||
if not selected_ids:
|
||||
raise ValueError("selected_ids must not be empty")
|
||||
store.finish_stage(
|
||||
"load_input",
|
||||
outputs={
|
||||
"extracted_path": _normalize_repo_path(extracted_path),
|
||||
"selected_id_count": len(selected_ids),
|
||||
"output_dir": _normalize_repo_path(output_dir),
|
||||
},
|
||||
)
|
||||
|
||||
store.start_stage("generate_markdown")
|
||||
config = ArticleSummaryConfig(
|
||||
max_retries=int(input_payload.get("max_retries") or 2),
|
||||
timeout_seconds=float(input_payload.get("timeout_seconds") or 120.0),
|
||||
)
|
||||
written_paths = summarize_selected_articles(
|
||||
extracted_path=extracted_path,
|
||||
selected_ids=selected_ids,
|
||||
output_dir=output_dir,
|
||||
config=config,
|
||||
api_key=input_payload.get("llm_api_key"),
|
||||
model=input_payload.get("llm_model"),
|
||||
api_url=input_payload.get("llm_api_url"),
|
||||
)
|
||||
if not written_paths:
|
||||
raise RuntimeError("Article summary job produced no Markdown outputs.")
|
||||
normalized_paths = [_normalize_repo_path(Path(p)) for p in written_paths]
|
||||
store.finish_stage(
|
||||
"generate_markdown",
|
||||
outputs={
|
||||
"written_count": len(written_paths),
|
||||
"written_paths": normalized_paths,
|
||||
},
|
||||
)
|
||||
for idx, p in enumerate(written_paths, start=1):
|
||||
store.register_artifact(
|
||||
name=f"summary_markdown_{idx}",
|
||||
path=Path(p),
|
||||
kind="markdown",
|
||||
stage="generate_markdown",
|
||||
metadata={"output_type": "single_summary"},
|
||||
)
|
||||
|
||||
store.start_stage("write_result")
|
||||
result = {
|
||||
"job_id": job_id,
|
||||
"selected_ids": selected_ids,
|
||||
"written_paths": normalized_paths,
|
||||
"completed_at": _now().isoformat(),
|
||||
}
|
||||
result_file = _result_path(job_id)
|
||||
result_file.write_text(json.dumps(result, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||
store.register_artifact(name="job_result", path=result_file, kind="json", stage="write_result")
|
||||
report_file = _job_report_path(job_id)
|
||||
report_file.write_text(
|
||||
json.dumps(
|
||||
{
|
||||
"job_id": job_id,
|
||||
"status": "success",
|
||||
"written_count": len(normalized_paths),
|
||||
"written_paths": normalized_paths,
|
||||
},
|
||||
ensure_ascii=False,
|
||||
indent=2,
|
||||
),
|
||||
encoding="utf-8",
|
||||
)
|
||||
store.register_artifact(name="job_report", path=report_file, kind="json", stage="write_result")
|
||||
store.finish_stage("write_result", outputs={"result_path": _normalize_repo_path(result_file)})
|
||||
store.finish_run(status="success")
|
||||
return result
|
||||
except Exception as exc:
|
||||
current_stage = store.state.current_stage or "generate_markdown"
|
||||
store.fail_stage(current_stage, error=exc)
|
||||
report_file = _job_report_path(job_id)
|
||||
report_file.write_text(
|
||||
json.dumps(
|
||||
{
|
||||
"job_id": job_id,
|
||||
"status": "failed",
|
||||
"error_type": type(exc).__name__,
|
||||
"error_message": str(exc),
|
||||
"failed_stage": current_stage,
|
||||
},
|
||||
ensure_ascii=False,
|
||||
indent=2,
|
||||
),
|
||||
encoding="utf-8",
|
||||
)
|
||||
try:
|
||||
store.register_artifact(name="job_report", path=report_file, kind="json", stage=current_stage)
|
||||
except Exception:
|
||||
pass
|
||||
raise
|
||||
|
||||
|
||||
def get_article_summary_job_status(*, job_id: str) -> dict[str, Any]:
|
||||
store = _load_run_store(job_id)
|
||||
state = store.state
|
||||
completed_stage_count = sum(1 for s in state.stages if s.status == "success")
|
||||
running_stage_count = sum(1 for s in state.stages if s.status == "running")
|
||||
failed_stage_count = sum(1 for s in state.stages if s.status == "failed")
|
||||
pending_stage_count = sum(1 for s in state.stages if s.status == "pending")
|
||||
return {
|
||||
"job_id": state.run_id,
|
||||
"workflow": state.workflow,
|
||||
"run_type": state.run_type,
|
||||
"status": state.status,
|
||||
"current_stage": state.current_stage,
|
||||
"started_at": state.started_at.isoformat(),
|
||||
"updated_at": state.updated_at.isoformat(),
|
||||
"finished_at": state.finished_at.isoformat() if state.finished_at else None,
|
||||
"output_dir": _normalize_repo_path(_job_dir(job_id)),
|
||||
"progress": {
|
||||
"completed_stage_count": completed_stage_count,
|
||||
"running_stage_count": running_stage_count,
|
||||
"failed_stage_count": failed_stage_count,
|
||||
"pending_stage_count": pending_stage_count,
|
||||
"total_stage_count": len(state.stages),
|
||||
},
|
||||
"artifacts": [artifact.model_dump(mode="json") for artifact in state.artifacts],
|
||||
"error_summary": state.error.model_dump(mode="json") if state.error else None,
|
||||
}
|
||||
|
||||
|
||||
def get_article_summary_job_result(*, job_id: str) -> dict[str, Any]:
|
||||
store = _load_run_store(job_id)
|
||||
state = store.state
|
||||
result = _load_result(job_id)
|
||||
if state.status != "success" or result is None:
|
||||
return {
|
||||
"job_id": state.run_id,
|
||||
"status": state.status,
|
||||
"message": "Article summary job result is not ready.",
|
||||
"error_summary": state.error.model_dump(mode="json") if state.error else None,
|
||||
}
|
||||
artifact = next((a.model_dump(mode="json") for a in state.artifacts if a.name == "job_result"), None)
|
||||
return {
|
||||
"job_id": state.run_id,
|
||||
"status": state.status,
|
||||
"written_paths": result.get("written_paths", []),
|
||||
"artifact": artifact,
|
||||
"result": result,
|
||||
}
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user