Compare commits
7
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
3940db1443 | ||
|
|
32ad584348 | ||
|
|
100044e1f7 | ||
|
|
27fe1e8882 | ||
|
|
cecf4d3ae7 | ||
|
|
ff3d15b8c4 | ||
|
|
c8d96705d1 |
+12
@@ -0,0 +1,12 @@
|
|||||||
|
.claude/
|
||||||
|
.codex/
|
||||||
|
__pycache__/
|
||||||
|
*.pyc
|
||||||
|
*.egg-info/
|
||||||
|
.env.local
|
||||||
|
findings.md
|
||||||
|
progress.md
|
||||||
|
task_plan.md
|
||||||
|
outputs/freshrss/
|
||||||
|
data/term_index/
|
||||||
|
outputs/term_index/
|
||||||
@@ -1,6 +1,6 @@
|
|||||||
# Content Extract MCP
|
# Content Extract MCP
|
||||||
|
|
||||||
Python MCP scaffold for article content extraction.
|
Python MCP scaffold for article content extraction, structured summary validation, deterministic filtering, and Markdown sink output.
|
||||||
|
|
||||||
## Run
|
## Run
|
||||||
|
|
||||||
@@ -9,22 +9,208 @@ pip install -e .
|
|||||||
summary-mcp
|
summary-mcp
|
||||||
```
|
```
|
||||||
|
|
||||||
The server exposes two tools:
|
The server exposes four tools:
|
||||||
|
|
||||||
- `extract_url_content`
|
- `extract_url_content`
|
||||||
- `extract_item_content`
|
- `extract_item_content`
|
||||||
|
- `filter_summary_result`
|
||||||
|
- `run_freshrss_openclaw_pipeline`
|
||||||
|
|
||||||
Validate an LLM summary result:
|
Validate an LLM summary result:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
validate-llm-result outputs/result.json --extracted outputs/read-flow-2026.extracted.json
|
validate-llm-result outputs/reference/summary/result.json --extracted outputs/reference/extracted/read-flow-2026.extracted.json
|
||||||
```
|
```
|
||||||
|
|
||||||
Run the minimal extraction-to-summary loop:
|
Run the minimal extraction-to-summary loop:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
python scripts/run_summary_loop.py ^
|
python scripts/run_summary_loop.py ^
|
||||||
--extracted outputs/read-flow-2026.extracted.json ^
|
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||||
--prompt outputs/llm-summary-prompt.txt ^
|
--prompt outputs/prompts/llm-summary-prompt.txt ^
|
||||||
--output outputs/result.json
|
--output outputs/reference/summary/result.loop.json
|
||||||
```
|
```
|
||||||
|
|
||||||
|
Pull FreshRSS entries and map them into normalized `item` objects:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||||||
|
set FRESHRSS_USERNAME=bot
|
||||||
|
set FRESHRSS_API_PASSWORD=your-api-password
|
||||||
|
python scripts/pull_freshrss_items.py --limit 5 --mark-read
|
||||||
|
```
|
||||||
|
|
||||||
|
By default the script excludes entries already tagged as `read`. Add `--include-read` if you want the full reading list.
|
||||||
|
When `--mark-read` is enabled, fetched entries are marked as read after the script finishes successfully.
|
||||||
|
|
||||||
|
The script writes:
|
||||||
|
|
||||||
|
- `outputs/freshrss/raw/freshrss.raw.json`
|
||||||
|
- `outputs/freshrss/items/freshrss.items.json`
|
||||||
|
|
||||||
|
Pull FreshRSS entries and run content extraction for each mapped item:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||||||
|
set FRESHRSS_USERNAME=osiman
|
||||||
|
set FRESHRSS_API_PASSWORD=your-api-password
|
||||||
|
python scripts/run_freshrss_extract.py --limit 1 --mark-read
|
||||||
|
```
|
||||||
|
|
||||||
|
By default the script excludes entries already tagged as `read`. When `--mark-read` is enabled, only entries with successful extraction are marked as read.
|
||||||
|
|
||||||
|
The script writes:
|
||||||
|
|
||||||
|
- `outputs/freshrss/raw/freshrss.raw.json`
|
||||||
|
- `outputs/freshrss/items/freshrss.items.json`
|
||||||
|
- `outputs/freshrss/extracted/freshrss.extracted.json`
|
||||||
|
|
||||||
|
Run the full FreshRSS pipeline and mark items as read only after the final OpenClaw delivery payload is written:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||||||
|
set FRESHRSS_USERNAME=osiman
|
||||||
|
set FRESHRSS_API_PASSWORD=your-api-password
|
||||||
|
set LLM_API_URL=https://api.deepseek.com
|
||||||
|
set LLM_API_KEY=your-llm-api-key
|
||||||
|
set LLM_MODEL=deepseek-chat
|
||||||
|
python scripts/run_freshrss_pipeline.py --limit 5 --mark-read
|
||||||
|
```
|
||||||
|
|
||||||
|
If you want to apply your personal engineering and AI-agent interest profile during filtering, pass a context file:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python scripts/run_freshrss_pipeline.py ^
|
||||||
|
--limit 5 ^
|
||||||
|
--context configs/filter_context.personal.json ^
|
||||||
|
--mark-read
|
||||||
|
```
|
||||||
|
|
||||||
|
This is the recommended production entrypoint. By default it writes only:
|
||||||
|
|
||||||
|
- `outputs/freshrss/rerun/<timestamp>/raw/freshrss.raw.json`
|
||||||
|
- `outputs/freshrss/rerun/<timestamp>/candidates/openclaw-delivery-payload.json`
|
||||||
|
- `outputs/freshrss/rerun/<timestamp>/run-report.json`
|
||||||
|
|
||||||
|
It also updates the daily keyword index runtime data:
|
||||||
|
|
||||||
|
- `data/term_index/daily/YYYY-MM-DD.json`
|
||||||
|
- `data/term_index/term_stats.json`
|
||||||
|
|
||||||
|
If you need per-item intermediates, add `--debug-artifacts`.
|
||||||
|
|
||||||
|
When OpenClaw is connected to the MCP server, it should call `run_freshrss_openclaw_pipeline` for the same behavior directly through MCP. The tool also supports `debug_artifacts=true` when deeper inspection is needed.
|
||||||
|
|
||||||
|
Run deterministic filter rules against a structured summary result:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python scripts/run_filter_rules.py ^
|
||||||
|
--summary outputs/reference/summary/result.loop.json ^
|
||||||
|
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||||
|
--output outputs/reference/filter/filter-decision.json
|
||||||
|
```
|
||||||
|
|
||||||
|
You can optionally pass a context file to inject interest topics or source tags:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python scripts/run_filter_rules.py ^
|
||||||
|
--summary outputs/reference/summary/result.loop.json ^
|
||||||
|
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||||
|
--context outputs/reference/filter/filter-context.json ^
|
||||||
|
--output outputs/reference/filter/filter-decision.with-context.json
|
||||||
|
```
|
||||||
|
|
||||||
|
Rule engine details and rule authoring guidance live in:
|
||||||
|
|
||||||
|
- `docs/design/filter-rule-engine-design.md`
|
||||||
|
- `docs/design/filter-rule-engine-usage.md`
|
||||||
|
|
||||||
|
Write a filtered result into the Markdown sink:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python scripts/run_markdown_sink.py ^
|
||||||
|
--summary outputs/reference/summary/result.loop.json ^
|
||||||
|
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||||
|
--filter outputs/reference/filter/filter-decision.json
|
||||||
|
```
|
||||||
|
|
||||||
|
The script writes markdown notes under `knowledge-base/`.
|
||||||
|
|
||||||
|
Build an internal `ArticleCandidateRecord` and a slim `OpenClawCandidateInput`:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python scripts/run_article_candidate.py ^
|
||||||
|
--summary outputs/reference/summary/result.loop.json ^
|
||||||
|
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||||
|
--filter outputs/reference/filter/filter-decision.json ^
|
||||||
|
--section-hint tools_and_workflows
|
||||||
|
```
|
||||||
|
|
||||||
|
The script writes by default:
|
||||||
|
|
||||||
|
- `outputs/reference/candidates/article-candidate-record.json`
|
||||||
|
- `outputs/reference/candidates/openclaw-candidate-input.json`
|
||||||
|
|
||||||
|
Build a batch OpenClaw delivery payload:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python scripts/build_openclaw_delivery.py ^
|
||||||
|
--input-dir outputs/freshrss/candidates/batch ^
|
||||||
|
--sort-by-rank ^
|
||||||
|
--date 2026-03-25
|
||||||
|
```
|
||||||
|
|
||||||
|
The script writes by default:
|
||||||
|
|
||||||
|
- `outputs/reference/candidates/openclaw-delivery-payload.json`
|
||||||
|
|
||||||
|
Output layout details live in `outputs/README.md`.
|
||||||
|
|
||||||
|
Keyword index defaults live in:
|
||||||
|
|
||||||
|
- `configs/term_aliases.json`
|
||||||
|
- `configs/term_stopwords.json`
|
||||||
|
- `configs/term_cleanup_policy.json`
|
||||||
|
- `configs/term_watchlist.json`
|
||||||
|
- `configs/term_change_log.json`
|
||||||
|
|
||||||
|
You can also rebuild the keyword index from an existing delivery payload:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python scripts/build_keyword_index.py ^
|
||||||
|
--input outputs/reference/candidates/openclaw-delivery-payload.json
|
||||||
|
```
|
||||||
|
|
||||||
|
Runtime keyword data is stored under `data/term_index/`.
|
||||||
|
|
||||||
|
The keyword cleanup review skill lives in:
|
||||||
|
|
||||||
|
- `skills/keyword-cleanup-review/`
|
||||||
|
|
||||||
|
To build a review bundle for the LLM skill:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python skills/keyword-cleanup-review/scripts/build_review_bundle.py ^
|
||||||
|
--days 7 ^
|
||||||
|
--top 50 ^
|
||||||
|
--output outputs/term_index/review/keyword-cleanup-bundle.json
|
||||||
|
```
|
||||||
|
|
||||||
|
The review bundle now also carries cleanup governance context:
|
||||||
|
|
||||||
|
- cleanup thresholds from `configs/term_cleanup_policy.json`
|
||||||
|
- the current watch list from `configs/term_watchlist.json`
|
||||||
|
- recent applied changes from `configs/term_change_log.json`
|
||||||
|
|
||||||
|
The skill only produces review inputs and suggestions. It does not modify `term_aliases`, `term_stopwords`, or `filter_context.personal.json` automatically.
|
||||||
|
|
||||||
|
To preview accepted suggestions before writing any config files:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python scripts/apply_term_suggestions.py ^
|
||||||
|
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json ^
|
||||||
|
--accept-watch Cron Heartbeat Memory ^
|
||||||
|
--dry-run
|
||||||
|
```
|
||||||
|
|
||||||
|
Remove `--dry-run` to write the accepted changes. The script can also apply accepted `alias`, `stopword`, and `interest keyword` suggestions through `--accept-alias`, `--accept-stopword`, and `--accept-interest`. Accepted watch terms are written into `configs/term_watchlist.json`, and every applied action is appended into `configs/term_change_log.json`.
|
||||||
|
|||||||
@@ -1,73 +1,83 @@
|
|||||||
# TODO
|
# TODO
|
||||||
|
|
||||||
## 当前状态
|
## 当前状态
|
||||||
|
|
||||||
项目当前处于 `Content Extract MCP` 的 MVP 完成阶段。
|
项目当前已经进入“可交付给 OpenClaw 调用”的阶段。
|
||||||
|
|
||||||
已完成:
|
当前主链路:
|
||||||
|
|
||||||
- [x] 明确整体阅读流链路
|
`FreshRSS 未读 -> RSS 内容提取 -> LLM 总结 -> 规则过滤 -> OpenClaw delivery payload`
|
||||||
- [x] 明确 `source -> item -> document/article` 的数据抽象
|
|
||||||
- [x] 明确 `MCP 负责提取,LLM 负责摘要` 的职责边界
|
当前已经完成:
|
||||||
- [x] 搭建 Python MCP 服务骨架
|
|
||||||
- [x] 实现 `extract_url_content`
|
- [x] FreshRSS `greader` API 接入
|
||||||
- [x] 实现 `extract_item_content`
|
- [x] `entry -> item` 标准化映射
|
||||||
- [x] 完成标题与正文提取
|
- [x] RSS-first 内容提取策略
|
||||||
- [x] 完成质量标记与结构化错误输出
|
- [x] LLM 摘要与校验闭环
|
||||||
- [x] 用参考文章完成真实提取测试
|
- [x] 第一版规则引擎
|
||||||
- [x] 输出结构化提取 JSON 文件
|
- [x] `ArticleCandidateRecord` / `OpenClawCandidateInput` 分层
|
||||||
- [x] 设计并迭代 LLM 摘要 prompt
|
- [x] `OpenClawDeliveryPayload` 批量投递结构
|
||||||
- [x] 验证 LLM 输出的 `result.json` 基本符合预期
|
- [x] 最终 payload 成功后才标记 FreshRSS 已读
|
||||||
- [x] 已封装 LLM 摘要校验 workflow skill
|
- [x] MCP 工具 `run_freshrss_openclaw_pipeline`
|
||||||
|
- [x] 默认精简输出模式
|
||||||
|
- [x] 日报级 `keywords` 词元库与周期性词元清洗 skill 设计完成
|
||||||
|
- [x] 日报级 `keywords` 词元库与全局词频统计实现完成
|
||||||
|
- [x] `keyword-cleanup-review` skill 骨架与 review bundle 脚本实现完成
|
||||||
|
- [x] 词元清洗低复杂治理层落地:`term_cleanup_policy` / `term_watchlist` / `term_change_log`
|
||||||
|
- [x] 已支持人工确认采纳建议并写入 `term_watchlist` / `term_change_log`
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## P0 - 近期必须完成
|
## P0 - 交接前后最优先
|
||||||
|
|
||||||
- [x] 为 LLM 摘要结果定义正式 JSON Schema
|
- [x] 为 OpenClaw 补齐交接文档
|
||||||
- [x] 增加一个本地校验脚本,自动校验 `result.json` 是否符合 schema
|
- [x] 将 MCP 工具作为统一生产入口
|
||||||
- [x] 把“提取 JSON -> LLM 摘要 JSON”串成一个标准化本地流程
|
- [x] 将默认输出收敛为最小必要文件
|
||||||
- [x] 清理或移除当前已不再使用的 `src/summary_mcp/core/summarizer.py`
|
- [ ] 设计 OpenClaw webhook / delivery payload 的主动推送方式
|
||||||
- [x] 更新旧设计文档中仍然残留的 `summary mcp` 描述,避免和当前实现冲突
|
- [ ] 明确 OpenClaw 侧如何注册和启动本 MCP 服务
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## P1 - 下一阶段推进
|
## P1 - 下一阶段推进
|
||||||
|
|
||||||
- [ ] 把上游 RSS 聚合结果映射成标准化 `item`
|
- [ ] 设计“人工确认后再沉淀知识库”的状态流转
|
||||||
- [ ] 补充 `item` 的实际样例文件
|
- [ ] 收敛 `paywall` 误判规则,降低中文文本误报
|
||||||
- [ ] 定义规则过滤层的输入输出 schema
|
- [ ] 细化过滤规则并引入更多个性化上下文
|
||||||
- [ ] 设计知识库入库格式
|
- [ ] 将 `keyword-cleanup-review` skill 接入周期性执行流程,产出别名/停用词/兴趣词建议
|
||||||
- [ ] 设计 webhook / 推送格式
|
- [ ] 增加清洗前后效果对比报告,验证配置调整是否真的改善过滤质量
|
||||||
- [ ] 给提取结果增加更多正文清洗策略,比如尾部噪音清理
|
- [ ] 增加批量 run 的保留策略与历史清理策略
|
||||||
|
- [ ] 为 OpenClaw 补一份更正式的 MCP 调用示例和接线说明
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## P2 - 后续增强项
|
## P2 - 后续增强
|
||||||
|
|
||||||
- [ ] 增加批量提取能力
|
- [ ] 将 `Markdown sink` 进一步降级为 debug / fallback 能力
|
||||||
- [ ] 引入 Playwright 作为动态页面兜底抓取方案
|
- [ ] 增加按天聚合 `ArticleCandidateRecord` 的批处理能力
|
||||||
- [ ] 支持更多 `content_kind`,例如 `release`、`thread`、`video`
|
- [ ] 让 OpenClaw 聚合候选内容并生成日级摘要
|
||||||
|
- [ ] 将日级摘要写入知识库,并同步生成面向用户的日报消息
|
||||||
|
- [ ] 支持更多 `content_kind`
|
||||||
- [ ] 增加提取缓存、重试和更细粒度日志
|
- [ ] 增加提取缓存、重试和更细粒度日志
|
||||||
- [ ] 重命名包和项目名,从 `summary_mcp` 调整为更符合当前职责的名称
|
- [ ] 整理历史 rerun 目录与调试产物保留策略
|
||||||
- [ ] 与 OpenClaw 做更正式的工作流编排整合
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## 当前建议的下一步
|
## 当前建议的下一步
|
||||||
|
|
||||||
|
优先做这三件事:
|
||||||
|
|
||||||
优先做这两件事:
|
1. 将 `keyword-cleanup-review` skill 接入周期性执行流程
|
||||||
|
2. 设计“人工确认后再沉淀知识库”的状态流转
|
||||||
1. 把上游 RSS 聚合结果映射成标准化 `item`
|
3. 收敛规则误判,尤其是 `paywall` 相关启发式
|
||||||
2. 补充一个真实 `item` 样例文件并开始设计规则过滤 schema
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## 收束上下文后建议先看
|
## 交接时优先阅读
|
||||||
|
|
||||||
- docs/context-reset-brief.md
|
|
||||||
- docs/content-extract-mcp-mvp-archive.md
|
|
||||||
- TODO.md
|
|
||||||
|
|
||||||
|
|
||||||
|
- `README.md`
|
||||||
|
- `docs/openclaw/openclaw-handoff.md`
|
||||||
|
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||||
|
- `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||||
|
- `docs/design/daily-keyword-index-design.md`
|
||||||
|
- `skills/keyword-cleanup-review/SKILL.md`
|
||||||
|
- `docs/current/context-reset-brief.md`
|
||||||
@@ -0,0 +1,56 @@
|
|||||||
|
{
|
||||||
|
"source_tags": [
|
||||||
|
"backend",
|
||||||
|
"ai-agent",
|
||||||
|
"frontier-tech"
|
||||||
|
],
|
||||||
|
"interest_topics": [
|
||||||
|
"后端工程",
|
||||||
|
"系统设计",
|
||||||
|
"分布式系统",
|
||||||
|
"数据库",
|
||||||
|
"缓存",
|
||||||
|
"消息队列",
|
||||||
|
"性能优化",
|
||||||
|
"可观测性",
|
||||||
|
"云原生",
|
||||||
|
"Kubernetes",
|
||||||
|
"AI Agent",
|
||||||
|
"LLM",
|
||||||
|
"RAG",
|
||||||
|
"MCP",
|
||||||
|
"工作流自动化",
|
||||||
|
"前沿科技"
|
||||||
|
],
|
||||||
|
"interest_keywords": [
|
||||||
|
"Java",
|
||||||
|
"Go",
|
||||||
|
"Python",
|
||||||
|
"Spring",
|
||||||
|
"FastAPI",
|
||||||
|
"Gin",
|
||||||
|
"gRPC",
|
||||||
|
"MySQL",
|
||||||
|
"PostgreSQL",
|
||||||
|
"Redis",
|
||||||
|
"Kafka",
|
||||||
|
"微服务",
|
||||||
|
"可观测性",
|
||||||
|
"Kubernetes",
|
||||||
|
"云原生",
|
||||||
|
"AI Agent",
|
||||||
|
"Agent",
|
||||||
|
"LLM",
|
||||||
|
"RAG",
|
||||||
|
"MCP",
|
||||||
|
"Prompt Engineering",
|
||||||
|
"Workflow",
|
||||||
|
"向量数据库",
|
||||||
|
"知识库",
|
||||||
|
"OpenAI",
|
||||||
|
"DeepSeek",
|
||||||
|
"OpenClaw",
|
||||||
|
"AliSQL",
|
||||||
|
"MySQL复制延迟"
|
||||||
|
]
|
||||||
|
}
|
||||||
@@ -0,0 +1,138 @@
|
|||||||
|
[
|
||||||
|
{
|
||||||
|
"rule_id": "drop-low-content",
|
||||||
|
"enabled": true,
|
||||||
|
"stop_on_match": true,
|
||||||
|
"conditions_all": [
|
||||||
|
{ "field": "article.quality_flags.is_low_content", "op": "eq", "value": true }
|
||||||
|
],
|
||||||
|
"action": {
|
||||||
|
"decision": "drop",
|
||||||
|
"reason": "Extracted article quality is too low.",
|
||||||
|
"labels": ["quality", "low-content"],
|
||||||
|
"priority": 100
|
||||||
|
}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"rule_id": "review-truncated-content",
|
||||||
|
"enabled": true,
|
||||||
|
"stop_on_match": false,
|
||||||
|
"conditions_all": [
|
||||||
|
{ "field": "article.quality_flags.is_truncated", "op": "eq", "value": true }
|
||||||
|
],
|
||||||
|
"action": {
|
||||||
|
"decision": "review",
|
||||||
|
"reason": "Extracted article may be truncated.",
|
||||||
|
"labels": ["quality", "truncated"],
|
||||||
|
"priority": 95
|
||||||
|
}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"rule_id": "review-paywalled-content",
|
||||||
|
"enabled": true,
|
||||||
|
"stop_on_match": false,
|
||||||
|
"conditions_all": [
|
||||||
|
{ "field": "article.quality_flags.is_paywalled", "op": "eq", "value": true }
|
||||||
|
],
|
||||||
|
"action": {
|
||||||
|
"decision": "review",
|
||||||
|
"reason": "Article may be behind a paywall.",
|
||||||
|
"labels": ["quality", "paywall"],
|
||||||
|
"priority": 95
|
||||||
|
}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"rule_id": "drop-ephemeral-news",
|
||||||
|
"enabled": true,
|
||||||
|
"stop_on_match": true,
|
||||||
|
"conditions_all": [
|
||||||
|
{ "field": "summary.category", "op": "eq", "value": "资讯" },
|
||||||
|
{ "field": "summary.worth_keeping", "op": "eq", "value": false }
|
||||||
|
],
|
||||||
|
"action": {
|
||||||
|
"decision": "drop",
|
||||||
|
"reason": "News-like content marked as not worth keeping.",
|
||||||
|
"labels": ["summary", "ephemeral-news"],
|
||||||
|
"priority": 90
|
||||||
|
}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"rule_id": "keep-worth-keeping-method",
|
||||||
|
"enabled": true,
|
||||||
|
"stop_on_match": false,
|
||||||
|
"conditions_all": [
|
||||||
|
{ "field": "summary.worth_keeping", "op": "eq", "value": true },
|
||||||
|
{ "field": "summary.category", "op": "in", "value": ["方法论", "工具实践"] }
|
||||||
|
],
|
||||||
|
"action": {
|
||||||
|
"decision": "keep",
|
||||||
|
"reason": "Structured summary marked the content as worth keeping in a durable category.",
|
||||||
|
"labels": ["summary", "durable"],
|
||||||
|
"priority": 80
|
||||||
|
}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"rule_id": "keep-interest-topic",
|
||||||
|
"enabled": true,
|
||||||
|
"stop_on_match": false,
|
||||||
|
"conditions_all": [
|
||||||
|
{ "field": "summary.worth_keeping", "op": "eq", "value": true },
|
||||||
|
{ "field": "summary.category", "op": "not_in", "value": ["资讯"] },
|
||||||
|
{ "field": "summary.topics", "op": "overlap", "value": { "from_field": "context.interest_topics" } }
|
||||||
|
],
|
||||||
|
"action": {
|
||||||
|
"decision": "keep",
|
||||||
|
"reason": "Topics overlap with the backend and AI-agent interest profile on non-news content.",
|
||||||
|
"labels": ["interest", "topic-match"],
|
||||||
|
"priority": 78
|
||||||
|
}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"rule_id": "keep-interest-keyword",
|
||||||
|
"enabled": true,
|
||||||
|
"stop_on_match": false,
|
||||||
|
"conditions_all": [
|
||||||
|
{ "field": "summary.worth_keeping", "op": "eq", "value": true },
|
||||||
|
{ "field": "summary.category", "op": "not_in", "value": ["资讯"] },
|
||||||
|
{ "field": "summary.keywords", "op": "overlap", "value": { "from_field": "context.interest_keywords" } }
|
||||||
|
],
|
||||||
|
"action": {
|
||||||
|
"decision": "keep",
|
||||||
|
"reason": "Keywords match the backend engineering or AI-agent watchlist on durable content.",
|
||||||
|
"labels": ["interest", "keyword-match", "engineering"],
|
||||||
|
"priority": 77
|
||||||
|
}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"rule_id": "review-interest-news-watch",
|
||||||
|
"enabled": true,
|
||||||
|
"stop_on_match": false,
|
||||||
|
"conditions_all": [
|
||||||
|
{ "field": "summary.category", "op": "eq", "value": "资讯" }
|
||||||
|
],
|
||||||
|
"conditions_any": [
|
||||||
|
{ "field": "summary.topics", "op": "overlap", "value": { "from_field": "context.interest_topics" } },
|
||||||
|
{ "field": "summary.keywords", "op": "overlap", "value": { "from_field": "context.interest_keywords" } }
|
||||||
|
],
|
||||||
|
"action": {
|
||||||
|
"decision": "review",
|
||||||
|
"reason": "Relevant frontier or engineering news should be watched first instead of auto-kept.",
|
||||||
|
"labels": ["interest", "news-watch"],
|
||||||
|
"priority": 70
|
||||||
|
}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"rule_id": "review-worth-keeping-other",
|
||||||
|
"enabled": true,
|
||||||
|
"stop_on_match": false,
|
||||||
|
"conditions_all": [
|
||||||
|
{ "field": "summary.worth_keeping", "op": "eq", "value": true }
|
||||||
|
],
|
||||||
|
"action": {
|
||||||
|
"decision": "review",
|
||||||
|
"reason": "Worth-keeping signal is positive but no stronger keep rule matched.",
|
||||||
|
"labels": ["summary", "needs-review"],
|
||||||
|
"priority": 60
|
||||||
|
}
|
||||||
|
}
|
||||||
|
]
|
||||||
@@ -0,0 +1,5 @@
|
|||||||
|
{
|
||||||
|
"AI助手": "AI Agent",
|
||||||
|
"图文RAG": "RAG",
|
||||||
|
"Prompt架构": "Prompt Engineering"
|
||||||
|
}
|
||||||
@@ -0,0 +1,4 @@
|
|||||||
|
{
|
||||||
|
"schema_version": "v1",
|
||||||
|
"entries": []
|
||||||
|
}
|
||||||
@@ -0,0 +1,25 @@
|
|||||||
|
{
|
||||||
|
"schema_version": "v1",
|
||||||
|
"interest_keyword_review": {
|
||||||
|
"min_total_count": 3,
|
||||||
|
"min_days_seen": 2
|
||||||
|
},
|
||||||
|
"watch_term_review": {
|
||||||
|
"min_total_count": 1,
|
||||||
|
"min_days_seen": 1,
|
||||||
|
"max_total_count": 2,
|
||||||
|
"max_days_seen": 2
|
||||||
|
},
|
||||||
|
"alias_review": {
|
||||||
|
"min_total_count": 2,
|
||||||
|
"min_days_seen": 2
|
||||||
|
},
|
||||||
|
"stopword_review": {
|
||||||
|
"max_total_count": 2,
|
||||||
|
"max_days_seen": 2
|
||||||
|
},
|
||||||
|
"notes": [
|
||||||
|
"当前阶段采用保守阈值,避免在低样本条件下直接扩充 interest_keywords。",
|
||||||
|
"watch_terms 先用于观察,后续再决定是否升格为 interest_keywords 或进入 alias/stopword 配置。"
|
||||||
|
]
|
||||||
|
}
|
||||||
@@ -0,0 +1,6 @@
|
|||||||
|
[
|
||||||
|
"小银",
|
||||||
|
"银行客户经理",
|
||||||
|
"飞盘物理",
|
||||||
|
"奋斗文化"
|
||||||
|
]
|
||||||
@@ -0,0 +1,5 @@
|
|||||||
|
{
|
||||||
|
"schema_version": "v1",
|
||||||
|
"updated_at": "2026-03-27T00:00:00Z",
|
||||||
|
"terms": []
|
||||||
|
}
|
||||||
+112
@@ -0,0 +1,112 @@
|
|||||||
|
# 文档索引
|
||||||
|
|
||||||
|
## 当前目录结构
|
||||||
|
|
||||||
|
- `docs/README.md`
|
||||||
|
- 文档总索引
|
||||||
|
- `docs/current/`
|
||||||
|
- 当前状态、收束入口、阶段导航
|
||||||
|
- `docs/design/`
|
||||||
|
- 当前实现的设计文档
|
||||||
|
- `docs/openclaw/`
|
||||||
|
- OpenClaw 日报聚合与下游对象设计
|
||||||
|
- `docs/notes/`
|
||||||
|
- 较上层的方案笔记与非最终设计
|
||||||
|
- `docs/archive/`
|
||||||
|
- 历史归档,不作为最新事实来源
|
||||||
|
|
||||||
|
## 当前推荐阅读顺序
|
||||||
|
|
||||||
|
1. `docs/current/context-reset-brief.md`
|
||||||
|
- 当前真实进度与下一步入口
|
||||||
|
2. `docs/openclaw/openclaw-handoff.md`
|
||||||
|
- 给 OpenClaw 的接手说明、环境变量、MCP 调用方式与已知限制
|
||||||
|
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||||
|
- 提供给 OpenClaw 的单篇结构化输入字段说明
|
||||||
|
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||||
|
- 提供给 OpenClaw 的批量投递 envelope 说明
|
||||||
|
5. `docs/design/summary-mcp-service-design.md`
|
||||||
|
- 当前 MCP 服务的职责、接口和边界
|
||||||
|
6. `docs/design/filter-rule-engine-design.md`
|
||||||
|
- 过滤层的输入输出、规则结构与当前实现
|
||||||
|
7. `docs/design/filter-rule-engine-usage.md`
|
||||||
|
- 规则怎么写、怎么跑、结果怎么解读的使用说明
|
||||||
|
8. `docs/design/daily-keyword-index-design.md`
|
||||||
|
- 日报级词元库与周期性词元清洗 skill 设计
|
||||||
|
9. `docs/design/markdown-sink-design.md`
|
||||||
|
- 第一版 Markdown sink 的输入输出、目录结构与落地方式
|
||||||
|
10. `docs/openclaw/openclaw-daily-digest-refactor.md`
|
||||||
|
- 为什么要从单篇入库改成 OpenClaw 日报聚合链路
|
||||||
|
11. `docs/openclaw/article-candidate-daily-digest-schema.md`
|
||||||
|
- `ArticleCandidateRecord`、`OpenClawCandidateInput` 与 `DailyDigest` 的正式设计
|
||||||
|
12. `docs/design/source-schema-design.md`
|
||||||
|
- `source -> item -> document` 的对象设计
|
||||||
|
13. `docs/notes/reading-pipeline-design-notes.md`
|
||||||
|
- 更上层的阅读流方案与阶段划分
|
||||||
|
14. `docs/design/summary-loop-explained.md`
|
||||||
|
- 当前 LLM 摘要校验闭环的解释
|
||||||
|
|
||||||
|
## 当前文档分层
|
||||||
|
|
||||||
|
### 1. 当前状态与导航
|
||||||
|
|
||||||
|
- `README.md`
|
||||||
|
- 仓库入口与脚本运行方式
|
||||||
|
- `TODO.md`
|
||||||
|
- 当前优先级、已完成项、下一阶段任务
|
||||||
|
- `docs/current/context-reset-brief.md`
|
||||||
|
- 当前阶段状态的最短摘要
|
||||||
|
- `docs/README.md`
|
||||||
|
- 文档索引与阅读顺序
|
||||||
|
|
||||||
|
### 2. 当前实现设计
|
||||||
|
|
||||||
|
- `docs/design/summary-mcp-service-design.md`
|
||||||
|
- 当前内容提取 MCP 的真实设计
|
||||||
|
- `docs/design/summary-core-interface-design.md`
|
||||||
|
- 摘要/提取内核的接口抽象
|
||||||
|
- `docs/design/source-schema-design.md`
|
||||||
|
- `source`、`item`、`document` 三层 schema
|
||||||
|
- `docs/design/summary-loop-explained.md`
|
||||||
|
- 提取 JSON -> LLM 摘要 JSON -> 校验 的闭环说明
|
||||||
|
- `docs/design/filter-rule-engine-design.md`
|
||||||
|
- 第一版规则过滤引擎设计与落地位置
|
||||||
|
- `docs/design/filter-rule-engine-usage.md`
|
||||||
|
- 规则配置、调用方式与结果解读
|
||||||
|
- `docs/design/daily-keyword-index-design.md`
|
||||||
|
- 日报级词元库与周期性词元清洗 skill 设计
|
||||||
|
- `scripts/apply_term_suggestions.py`
|
||||||
|
- 人工确认后将建议写回 watchlist / change log 的脚本入口(见 `README.md` 用法)
|
||||||
|
- `docs/design/markdown-sink-design.md`
|
||||||
|
- 第一版 Markdown sink 设计与落地位置
|
||||||
|
|
||||||
|
### 3. OpenClaw 与下游设计
|
||||||
|
|
||||||
|
- `docs/openclaw/openclaw-handoff.md`
|
||||||
|
- OpenClaw 接手所需的运行说明、工具入口与已知限制
|
||||||
|
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||||
|
- 提供给 OpenClaw 的单篇结构化输入字段说明
|
||||||
|
- `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||||
|
- 提供给 OpenClaw 的批量投递 envelope 说明
|
||||||
|
- `docs/openclaw/openclaw-daily-digest-refactor.md`
|
||||||
|
- 改造为 OpenClaw 日报聚合链路的原因与目标结构
|
||||||
|
- `docs/openclaw/article-candidate-daily-digest-schema.md`
|
||||||
|
- `ArticleCandidateRecord`、`OpenClawCandidateInput` 与 `DailyDigest` 的字段设计与对象关系
|
||||||
|
|
||||||
|
### 4. 方案笔记
|
||||||
|
|
||||||
|
- `docs/notes/reading-pipeline-design-notes.md`
|
||||||
|
- 整体阅读流、规则、sink、push 的方案笔记
|
||||||
|
|
||||||
|
### 5. 历史归档
|
||||||
|
|
||||||
|
- `docs/archive/content-extract-mcp-mvp-archive.md`
|
||||||
|
- MVP 阶段归档,部分状态已被后续进展覆盖
|
||||||
|
|
||||||
|
## 当前文档维护原则
|
||||||
|
|
||||||
|
- `docs/current/context-reset-brief.md` 记录当前最新状态
|
||||||
|
- `TODO.md` 记录任务优先级与下一步
|
||||||
|
- `outputs/README.md` 记录当前输出目录约定
|
||||||
|
- `docs/archive/content-extract-mcp-mvp-archive.md` 只当历史快照,不再作为最新事实来源
|
||||||
|
- 新增阶段性进展,优先更新 `README.md`、`TODO.md`、`docs/current/context-reset-brief.md`
|
||||||
@@ -1,65 +0,0 @@
|
|||||||
# 项目当前状态简报
|
|
||||||
|
|
||||||
## 当前已完成
|
|
||||||
|
|
||||||
- 已明确整体链路:`来源 -> 聚合池 -> 内容提取 MCP -> LLM 摘要 -> 校验 -> 过滤 -> 入库/推送`
|
|
||||||
- 已确定当前 MCP 的职责边界:
|
|
||||||
- 只负责内容提取
|
|
||||||
- 不负责摘要、分类、价值判断
|
|
||||||
- 已完成 Python MCP 骨架
|
|
||||||
- 已实现两个 tool:
|
|
||||||
- `extract_url_content`
|
|
||||||
- `extract_item_content`
|
|
||||||
- 已完成真实 URL 提取验证
|
|
||||||
- 已完成 LLM 摘要 prompt
|
|
||||||
- 已完成 LLM 摘要结果 schema 校验器
|
|
||||||
- 已完成“提取 JSON -> LLM 摘要 JSON -> 校验”的最小闭环脚本
|
|
||||||
- 已完成 LLM 摘要校验 skill 封装
|
|
||||||
- 已清理旧的启发式 `summarizer.py`
|
|
||||||
- 已将旧的 MCP 设计文档更新为当前“Content Extract MCP”语义
|
|
||||||
|
|
||||||
## 当前关键文件
|
|
||||||
|
|
||||||
- MCP 入口:
|
|
||||||
- `src/summary_mcp/server.py`
|
|
||||||
- 提取主流程:
|
|
||||||
- `src/summary_mcp/core/pipeline.py`
|
|
||||||
- LLM 结果模型:
|
|
||||||
- `src/summary_mcp/models/llm_result.py`
|
|
||||||
- LLM 校验器:
|
|
||||||
- `src/summary_mcp/validators/llm_result.py`
|
|
||||||
- 校验 CLI:
|
|
||||||
- `src/summary_mcp/validate_llm_result.py`
|
|
||||||
- 最小闭环脚本:
|
|
||||||
- `scripts/run_summary_loop.py`
|
|
||||||
- 当前提示词:
|
|
||||||
- `outputs/llm-summary-prompt.txt`
|
|
||||||
- MVP 归档:
|
|
||||||
- `docs/content-extract-mcp-mvp-archive.md`
|
|
||||||
- 当前 TODO:
|
|
||||||
- `TODO.md`
|
|
||||||
|
|
||||||
## 当前已经验证通过
|
|
||||||
|
|
||||||
- 参考文章 URL 可提取为结构化 JSON
|
|
||||||
- LLM 可根据提取结果生成摘要 JSON
|
|
||||||
- validator 可校验摘要 JSON
|
|
||||||
- 最小闭环脚本可直接调用 LLM 接口并产出通过校验的结果
|
|
||||||
|
|
||||||
## 当前未开始的下一阶段
|
|
||||||
|
|
||||||
- 让 FreshRSS 作为主聚合池
|
|
||||||
- 设计 FreshRSS entry -> `item` 的映射
|
|
||||||
- 准备真实 `item` 样例
|
|
||||||
- 开始定义规则过滤层 schema
|
|
||||||
|
|
||||||
## 收束后建议从这里继续
|
|
||||||
|
|
||||||
优先从这两个问题继续:
|
|
||||||
|
|
||||||
1. FreshRSS 的 entry 字段如何映射成 `item`
|
|
||||||
2. 先做一个真实 `item` 样例文件,再讨论规则过滤
|
|
||||||
|
|
||||||
## 一句话结论
|
|
||||||
|
|
||||||
当前 MVP 已完成,下一阶段不再是继续打磨 MCP,而是开始接入 FreshRSS 上游并建立 `item` 标准化入口。
|
|
||||||
@@ -0,0 +1,153 @@
|
|||||||
|
# 项目当前状态简报
|
||||||
|
|
||||||
|
## 当前结论
|
||||||
|
|
||||||
|
当前仓库已经具备交付给 OpenClaw 的基础条件。
|
||||||
|
|
||||||
|
当前主链路是:
|
||||||
|
|
||||||
|
`FreshRSS 未读 -> RSS 内容提取 -> LLM 总结 -> 规则过滤 -> OpenClaw delivery payload`
|
||||||
|
|
||||||
|
OpenClaw 应通过 MCP 工具 `run_freshrss_openclaw_pipeline` 调用这条链路,而不是自行拼接脚本。
|
||||||
|
|
||||||
|
## 当前已完成
|
||||||
|
|
||||||
|
- 已完成 FreshRSS `greader` API 接入与未读拉取
|
||||||
|
- 已完成 FreshRSS 条目到标准化 `item` 的映射
|
||||||
|
- 已完成 RSS-first 提取策略
|
||||||
|
- 已完成 LLM 总结与校验闭环
|
||||||
|
- 已完成规则引擎过滤
|
||||||
|
- 已完成 `ArticleCandidateRecord` 与 `OpenClawCandidateInput` 分层
|
||||||
|
- 已完成 `OpenClawDeliveryPayload` 批量投递结构
|
||||||
|
- 已完成 FreshRSS 已读状态回写
|
||||||
|
- 已完成“仅在最终 payload 成功写盘后再标记已读”的语义
|
||||||
|
- 已完成 MCP 工具 `run_freshrss_openclaw_pipeline`
|
||||||
|
- 已完成默认精简输出模式,减少中间文件
|
||||||
|
- 已完成日报级 `keywords` 词元库与全局词频统计
|
||||||
|
- 已完成 `keyword-cleanup-review` skill 骨架与 review bundle 脚本
|
||||||
|
- 已完成低复杂治理层:`term_cleanup_policy` / `term_watchlist` / `term_change_log`
|
||||||
|
- 已完成采纳建议写回脚本 `scripts/apply_term_suggestions.py`
|
||||||
|
|
||||||
|
## 当前 MCP 工具
|
||||||
|
|
||||||
|
当前服务入口:
|
||||||
|
|
||||||
|
- `src/summary_mcp/server.py`
|
||||||
|
|
||||||
|
当前暴露的 MCP 工具:
|
||||||
|
|
||||||
|
- `extract_url_content`
|
||||||
|
- `extract_item_content`
|
||||||
|
- `filter_summary_result`
|
||||||
|
- `run_freshrss_openclaw_pipeline`
|
||||||
|
|
||||||
|
其中生产主入口是:
|
||||||
|
|
||||||
|
- `run_freshrss_openclaw_pipeline`
|
||||||
|
|
||||||
|
## 当前关键文件
|
||||||
|
|
||||||
|
- MCP 服务入口
|
||||||
|
- `src/summary_mcp/server.py`
|
||||||
|
- FreshRSS 统一工作流
|
||||||
|
- `src/summary_mcp/workflows/freshrss_pipeline.py`
|
||||||
|
- 词元统计核心
|
||||||
|
- `src/summary_mcp/core/keyword_index.py`
|
||||||
|
- 词元统计模型
|
||||||
|
- `src/summary_mcp/models/keyword_index.py`
|
||||||
|
- 摘要循环
|
||||||
|
- `src/summary_mcp/core/summary_loop.py`
|
||||||
|
- 提取主流程
|
||||||
|
- `src/summary_mcp/core/pipeline.py`
|
||||||
|
- FreshRSS 集成
|
||||||
|
- `src/summary_mcp/integrations/freshrss.py`
|
||||||
|
- 规则引擎
|
||||||
|
- `src/summary_mcp/filters/engine.py`
|
||||||
|
- LLM 结果校验
|
||||||
|
- `src/summary_mcp/validators/llm_result.py`
|
||||||
|
- OpenClaw candidate 模型
|
||||||
|
- `src/summary_mcp/models/article_candidate.py`
|
||||||
|
- OpenClaw delivery 模型
|
||||||
|
- `src/summary_mcp/models/openclaw_delivery.py`
|
||||||
|
- 生产脚本入口
|
||||||
|
- `scripts/run_freshrss_pipeline.py`
|
||||||
|
- 词元统计重建脚本
|
||||||
|
- `scripts/build_keyword_index.py`
|
||||||
|
- 词元清洗 skill
|
||||||
|
- `skills/keyword-cleanup-review/SKILL.md`
|
||||||
|
- skill review bundle 脚本
|
||||||
|
- `skills/keyword-cleanup-review/scripts/build_review_bundle.py`
|
||||||
|
- 采纳建议写回脚本
|
||||||
|
- `scripts/apply_term_suggestions.py`
|
||||||
|
- 清洗治理配置
|
||||||
|
- `configs/term_cleanup_policy.json`
|
||||||
|
- `configs/term_watchlist.json`
|
||||||
|
- `configs/term_change_log.json`
|
||||||
|
- OpenClaw 交接说明
|
||||||
|
- `docs/openclaw/openclaw-handoff.md`
|
||||||
|
|
||||||
|
## 当前输出规则
|
||||||
|
|
||||||
|
默认生产模式只输出:
|
||||||
|
|
||||||
|
- `raw/freshrss.raw.json`
|
||||||
|
- `candidates/openclaw-delivery-payload.json`
|
||||||
|
- `run-report.json`
|
||||||
|
|
||||||
|
同时会更新本地运行数据:
|
||||||
|
|
||||||
|
- `data/term_index/daily/YYYY-MM-DD.json`
|
||||||
|
- `data/term_index/term_stats.json`
|
||||||
|
|
||||||
|
如果需要词元清洗审阅输入,可额外生成:
|
||||||
|
|
||||||
|
- `outputs/term_index/review/keyword-cleanup-bundle.json`
|
||||||
|
|
||||||
|
如果需要在人工确认后把建议正式写入 watchlist / change log,可使用:
|
||||||
|
|
||||||
|
- `scripts/apply_term_suggestions.py`
|
||||||
|
|
||||||
|
如果需要排障,可开启:
|
||||||
|
|
||||||
|
- `debug_artifacts=true`
|
||||||
|
- 或脚本参数 `--debug-artifacts`
|
||||||
|
|
||||||
|
这样才会额外输出逐条中间文件。
|
||||||
|
|
||||||
|
## 当前验证状态
|
||||||
|
|
||||||
|
已经验证通过:
|
||||||
|
|
||||||
|
- FreshRSS 未读拉取成功
|
||||||
|
- 已读回写成功
|
||||||
|
- MCP 工具入口可直接触发完整链路
|
||||||
|
- 微信公众号样本可直接使用 RSS 提供的 `summary` 内容提取,不再回源抓网页
|
||||||
|
- 精简输出模式已实际跑通
|
||||||
|
- 日报级词元统计已通过离线样例验证,确认别名、停用词、非 `drop` 过滤和 rerun 覆盖逻辑正常
|
||||||
|
- `keyword-cleanup-review` skill 已通过 `quick_validate.py` 结构校验
|
||||||
|
- review bundle 脚本已实际跑通
|
||||||
|
- `apply_term_suggestions.py` 已通过 dry-run 与临时副本写回验证
|
||||||
|
|
||||||
|
## 当前已知限制
|
||||||
|
|
||||||
|
- 当前对 FreshRSS 条目采用 RSS-first 策略,不再回源抓原网页
|
||||||
|
- 如果 RSS 中没有足够正文内容,该条会直接跳过,不会进入后续总结
|
||||||
|
- 某些规则仍偏保守,部分内容可能落到 `review`
|
||||||
|
- `paywall` 相关启发式仍可能误判中文文本
|
||||||
|
- Webhook / 主动投递到 OpenClaw 外部接口尚未实现,当前是由 OpenClaw 通过 MCP 主动调用
|
||||||
|
- 词元清洗 skill 当前已支持“bundle 构建 -> 建议审阅 -> 人工确认写回 watchlist/change_log”,但尚未接入周期性调度
|
||||||
|
- 当前词元统计仍以前置 `OpenClawDeliveryPayload` 作为日报前代理输入,真实 `DailyDigest` 接入后还需切换上游
|
||||||
|
|
||||||
|
## 当前最建议的交接阅读顺序
|
||||||
|
|
||||||
|
1. `README.md`
|
||||||
|
2. `docs/openclaw/openclaw-handoff.md`
|
||||||
|
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||||
|
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||||
|
5. `docs/design/daily-keyword-index-design.md`
|
||||||
|
6. `skills/keyword-cleanup-review/SKILL.md`
|
||||||
|
7. `TODO.md`
|
||||||
|
|
||||||
|
## 一句话结论
|
||||||
|
|
||||||
|
当前仓库已经从“提取 MCP 原型”演进到“可供 OpenClaw 调用的 FreshRSS -> OpenClaw payload 上游处理器”,并已补上第一阶段的日报级词元统计能力和词元清洗 skill 骨架;后续重点转向 skill 周期调度、知识库状态流转和 webhook 接线。
|
||||||
@@ -0,0 +1,461 @@
|
|||||||
|
# 日报级词元库与词元清洗 Skill 设计
|
||||||
|
|
||||||
|
## 1. 设计目标
|
||||||
|
|
||||||
|
当前项目已经具备:
|
||||||
|
|
||||||
|
`FreshRSS -> extraction -> LLM summary -> rule engine -> OpenClaw payload`
|
||||||
|
|
||||||
|
下一阶段希望新增“词元库”能力,目标不是做全文级检索索引,而是解决两件事:
|
||||||
|
|
||||||
|
1. 为后续规则配置提供稳定、轻量、可读的词元来源
|
||||||
|
2. 为周期性的词元清洗、合并和兴趣词补充提供数据基础
|
||||||
|
|
||||||
|
因此这套设计的核心原则是:
|
||||||
|
|
||||||
|
- 词元来源轻量化
|
||||||
|
- 数据粒度日报化
|
||||||
|
- 主链路程序化维护
|
||||||
|
- 清洗治理由独立 skill 周期性执行
|
||||||
|
- LLM 不直接修改规则或兴趣词配置
|
||||||
|
|
||||||
|
## 2. 为什么不用逐篇词元库
|
||||||
|
|
||||||
|
逐篇记录每篇文章的词元事件,虽然可追溯,但当前阶段成本过高,收益不足:
|
||||||
|
|
||||||
|
- 生成文件会很多
|
||||||
|
- 存储和调试负担更大
|
||||||
|
- 后续真正调规则时,用户更关心“最近日报里反复出现什么词”,而不是“某一篇文章具体抽到了什么词”
|
||||||
|
- 当前目标是服务日报与规则配置,不是做文章级分析平台
|
||||||
|
|
||||||
|
所以本项目不采用:
|
||||||
|
|
||||||
|
`article -> term event -> term store`
|
||||||
|
|
||||||
|
而采用:
|
||||||
|
|
||||||
|
`daily digest -> keyword aggregate -> daily term index -> global term stats`
|
||||||
|
|
||||||
|
## 3. 为什么只保留 `keywords`
|
||||||
|
|
||||||
|
当前 LLM 摘要结果里已经有两个候选字段:
|
||||||
|
|
||||||
|
- `keywords`
|
||||||
|
- `topics`
|
||||||
|
|
||||||
|
本设计只使用 `keywords` 进入词元库,不使用 `topics` 作为主来源。
|
||||||
|
|
||||||
|
原因:
|
||||||
|
|
||||||
|
- `keywords` 更具体,适合后续规则配置
|
||||||
|
- 例如 `Java`、`Go`、`Python`、`MCP`、`RAG`、`Kafka`
|
||||||
|
- `topics` 更宽泛,适合摘要展示,不适合作为精确规则命中基础
|
||||||
|
- 例如“后端工程”“AI Agent”“前沿科技”过于宽泛
|
||||||
|
- 只保留一个字段可以显著控制词元库规模
|
||||||
|
- 当前项目的真实需求是“词元可配置”,不是“主题聚类”
|
||||||
|
|
||||||
|
结论:
|
||||||
|
|
||||||
|
- `keywords` 进入词元库
|
||||||
|
- `topics` 继续保留在单篇摘要结果中,但不纳入词元统计主流程
|
||||||
|
|
||||||
|
## 4. 词元库的上游边界
|
||||||
|
|
||||||
|
词元库不直接读取所有候选文章,而只读取“最终进入日报”的内容。
|
||||||
|
|
||||||
|
也就是说,真正的词元来源是:
|
||||||
|
|
||||||
|
- `DailyDigest`
|
||||||
|
- 或等价的“已被日报选中”的 candidate 集合
|
||||||
|
|
||||||
|
不纳入词元库的内容:
|
||||||
|
|
||||||
|
- 被 `drop` 的内容
|
||||||
|
- 仅在中间候选层出现、但未进入日报的内容
|
||||||
|
- 调试产物中的临时摘要结果
|
||||||
|
|
||||||
|
这样做的好处是:
|
||||||
|
|
||||||
|
- 词元库只反映真正进入日级产物的内容
|
||||||
|
- 高频词更接近长期兴趣,而不是临时噪声
|
||||||
|
- 数据量更可控
|
||||||
|
|
||||||
|
## 5. 总体链路
|
||||||
|
|
||||||
|
目标链路调整为:
|
||||||
|
|
||||||
|
`DailyDigest -> keyword normalization -> daily term index -> global term stats -> cleaning skill review -> human confirm -> config update`
|
||||||
|
|
||||||
|
职责拆分如下:
|
||||||
|
|
||||||
|
### 5.1 程序负责
|
||||||
|
|
||||||
|
- 从日报中提取 `keywords`
|
||||||
|
- 归一化词元
|
||||||
|
- 应用别名映射
|
||||||
|
- 应用停用词过滤
|
||||||
|
- 生成每日词频
|
||||||
|
- 更新全局累计统计
|
||||||
|
|
||||||
|
### 5.2 Skill 负责
|
||||||
|
|
||||||
|
- 周期性读取词频结果
|
||||||
|
- 识别重复词、近义词、大小写变体
|
||||||
|
- 识别泛词、噪声词、低价值词
|
||||||
|
- 建议哪些词应合并、停用、加入兴趣词配置
|
||||||
|
|
||||||
|
### 5.3 人工负责
|
||||||
|
|
||||||
|
- 审核 skill 输出的建议
|
||||||
|
- 决定是否更新:
|
||||||
|
- `term_aliases`
|
||||||
|
- `term_stopwords`
|
||||||
|
- `filter_context.personal.json`
|
||||||
|
|
||||||
|
## 6. 数据文件设计
|
||||||
|
|
||||||
|
建议新增以下文件:
|
||||||
|
|
||||||
|
- `data/term_index/daily/YYYY-MM-DD.json`
|
||||||
|
- 某一天日报的词元聚合结果
|
||||||
|
- `data/term_index/term_stats.json`
|
||||||
|
- 全局累计词元统计
|
||||||
|
- `configs/term_aliases.json`
|
||||||
|
- 词元别名归一配置
|
||||||
|
- `configs/term_stopwords.json`
|
||||||
|
- 词元停用词配置
|
||||||
|
- `configs/term_cleanup_policy.json`
|
||||||
|
- 清洗阈值与治理策略配置
|
||||||
|
- `configs/term_watchlist.json`
|
||||||
|
- 当前处于观察状态的词元列表
|
||||||
|
- `configs/term_change_log.json`
|
||||||
|
- 已确认生效的词元治理变更记录
|
||||||
|
|
||||||
|
说明:
|
||||||
|
|
||||||
|
- 不新增逐篇 `term_events.jsonl`
|
||||||
|
- 不新增文章级明细文件
|
||||||
|
- 默认只保留“日报聚合结果 + 全局统计结果”
|
||||||
|
|
||||||
|
## 7. 每日词元文件结构
|
||||||
|
|
||||||
|
文件路径示例:
|
||||||
|
|
||||||
|
- `data/term_index/daily/2026-03-26.json`
|
||||||
|
|
||||||
|
建议结构:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"date": "2026-03-26",
|
||||||
|
"source": "daily_digest",
|
||||||
|
"digest_id": "digest-2026-03-26",
|
||||||
|
"generated_at": "2026-03-26T21:30:00+08:00",
|
||||||
|
"candidate_count": 8,
|
||||||
|
"terms": [
|
||||||
|
{
|
||||||
|
"term": "AI Agent",
|
||||||
|
"normalized_term": "AI Agent",
|
||||||
|
"count": 4
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"term": "MCP",
|
||||||
|
"normalized_term": "MCP",
|
||||||
|
"count": 3
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"term": "RAG",
|
||||||
|
"normalized_term": "RAG",
|
||||||
|
"count": 2
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
字段说明:
|
||||||
|
|
||||||
|
- `date`
|
||||||
|
- 日报日期
|
||||||
|
- `source`
|
||||||
|
- 固定标记为 `daily_digest`
|
||||||
|
- `digest_id`
|
||||||
|
- 日报对象唯一标识
|
||||||
|
- `generated_at`
|
||||||
|
- 该词元文件生成时间
|
||||||
|
- `candidate_count`
|
||||||
|
- 当天日报包含的条目数
|
||||||
|
- `terms[]`
|
||||||
|
- 当天词元聚合结果
|
||||||
|
|
||||||
|
其中单条 `terms[]` 只保留:
|
||||||
|
|
||||||
|
- `term`
|
||||||
|
- `normalized_term`
|
||||||
|
- `count`
|
||||||
|
|
||||||
|
不保留:
|
||||||
|
|
||||||
|
- 逐篇文章来源列表
|
||||||
|
- 逐条命中明细
|
||||||
|
- 本地文件路径
|
||||||
|
|
||||||
|
## 8. 全局统计文件结构
|
||||||
|
|
||||||
|
文件路径:
|
||||||
|
|
||||||
|
- `data/term_index/term_stats.json`
|
||||||
|
|
||||||
|
建议结构:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"schema_version": "v1",
|
||||||
|
"generated_at": "2026-03-26T21:30:00+08:00",
|
||||||
|
"terms": [
|
||||||
|
{
|
||||||
|
"term": "AI Agent",
|
||||||
|
"total_count": 18,
|
||||||
|
"days_seen": 6,
|
||||||
|
"first_seen": "2026-03-20",
|
||||||
|
"last_seen": "2026-03-26"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"term": "MCP",
|
||||||
|
"total_count": 12,
|
||||||
|
"days_seen": 5,
|
||||||
|
"first_seen": "2026-03-21",
|
||||||
|
"last_seen": "2026-03-26"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
字段说明:
|
||||||
|
|
||||||
|
- `term`
|
||||||
|
- 最终归一后的词元
|
||||||
|
- `total_count`
|
||||||
|
- 累计出现次数
|
||||||
|
- `days_seen`
|
||||||
|
- 出现过的天数
|
||||||
|
- `first_seen`
|
||||||
|
- 首次出现日期
|
||||||
|
- `last_seen`
|
||||||
|
- 最近出现日期
|
||||||
|
|
||||||
|
当前阶段不额外记录:
|
||||||
|
|
||||||
|
- 各 category 分桶统计
|
||||||
|
- keep/review/drop 分桶统计
|
||||||
|
- 文章级来源列表
|
||||||
|
|
||||||
|
原因是:日报级词元库的第一目标是轻量稳定,不是分析平台。
|
||||||
|
|
||||||
|
## 9. 标准化与归一规则
|
||||||
|
|
||||||
|
程序在写入日报词元前,应先做标准化。
|
||||||
|
|
||||||
|
### 9.1 基础标准化
|
||||||
|
|
||||||
|
- 去除首尾空白
|
||||||
|
- 保留中英文大小写风格中的稳定写法
|
||||||
|
- 去重
|
||||||
|
- 过滤空字符串
|
||||||
|
|
||||||
|
### 9.2 别名映射
|
||||||
|
|
||||||
|
通过 `configs/term_aliases.json` 做归一。
|
||||||
|
|
||||||
|
示例:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"Agent": "AI Agent",
|
||||||
|
"智能体": "AI Agent",
|
||||||
|
"Postgres": "PostgreSQL",
|
||||||
|
"Model Context Protocol": "MCP"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### 9.3 停用词过滤
|
||||||
|
|
||||||
|
通过 `configs/term_stopwords.json` 过滤过泛词和噪声词。
|
||||||
|
|
||||||
|
示例:
|
||||||
|
|
||||||
|
```json
|
||||||
|
[
|
||||||
|
"技术",
|
||||||
|
"系统",
|
||||||
|
"方案",
|
||||||
|
"实践",
|
||||||
|
"文章"
|
||||||
|
]
|
||||||
|
```
|
||||||
|
|
||||||
|
## 10. 为什么不让 LLM 直接维护词元库
|
||||||
|
|
||||||
|
LLM 可以帮助做清洗建议,但不适合直接维护主词元库。
|
||||||
|
|
||||||
|
原因:
|
||||||
|
|
||||||
|
- 主词元库更新应该稳定、低成本、可复现
|
||||||
|
- 词频统计属于纯程序逻辑,没必要消耗模型调用
|
||||||
|
- 如果让 LLM 直接写词元库,会引入不稳定和难审计问题
|
||||||
|
|
||||||
|
因此主流程固定为:
|
||||||
|
|
||||||
|
- LLM 只负责在摘要结果里输出 `keywords`
|
||||||
|
- 程序负责归一、聚合、统计
|
||||||
|
|
||||||
|
## 11. 词元清洗 Skill 设计
|
||||||
|
|
||||||
|
新增一个周期性清洗 skill,定位是“治理器”,不是“实时生产者”。
|
||||||
|
|
||||||
|
### 11.1 Skill 输入
|
||||||
|
|
||||||
|
建议输入:
|
||||||
|
|
||||||
|
- `data/term_index/term_stats.json`
|
||||||
|
- 最近 N 天的 `data/term_index/daily/*.json`
|
||||||
|
- `configs/term_aliases.json`
|
||||||
|
- `configs/term_stopwords.json`
|
||||||
|
- `configs/filter_context.personal.json`
|
||||||
|
|
||||||
|
### 11.2 Skill 输出
|
||||||
|
|
||||||
|
skill 不直接修改配置文件,而是生成建议文件,例如:
|
||||||
|
|
||||||
|
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md`
|
||||||
|
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json`
|
||||||
|
|
||||||
|
低复杂治理层建议补充三类输入:
|
||||||
|
|
||||||
|
- `term_cleanup_policy`
|
||||||
|
- 用来定义 watch 和 interest 的最小证据阈值
|
||||||
|
- `term_watchlist`
|
||||||
|
- 用来记录“先观察、暂不升级”的词元
|
||||||
|
- `term_change_log`
|
||||||
|
- 用来记录已经确认落地的配置变更,避免后续遗忘上下文
|
||||||
|
|
||||||
|
建议项包括:
|
||||||
|
|
||||||
|
- 建议合并的别名词
|
||||||
|
- 建议新增的停用词
|
||||||
|
- 建议加入 `interest_keywords` 的候选词
|
||||||
|
- 建议降权观察的热点词
|
||||||
|
|
||||||
|
### 11.3 Skill 允许做什么
|
||||||
|
|
||||||
|
- 发现重复词
|
||||||
|
- 发现大小写变体
|
||||||
|
- 发现中英文混用的近义词
|
||||||
|
- 发现持续高频但尚未进入兴趣配置的词
|
||||||
|
- 发现明显过泛的词
|
||||||
|
|
||||||
|
### 11.4 Skill 不允许做什么
|
||||||
|
|
||||||
|
- 直接改 `filter_rules.json`
|
||||||
|
- 直接改 `filter_context.personal.json`
|
||||||
|
- 直接覆盖 `term_stats.json`
|
||||||
|
- 在无人工确认的情况下自动生效
|
||||||
|
|
||||||
|
## 12. 建议的清洗建议文件结构
|
||||||
|
|
||||||
|
建议 JSON 文件结构如下:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"date": "2026-03-26",
|
||||||
|
"based_on_days": 7,
|
||||||
|
"alias_suggestions": [
|
||||||
|
{
|
||||||
|
"from": "Agent",
|
||||||
|
"to": "AI Agent",
|
||||||
|
"reason": "和现有高频词语义一致,建议归并。"
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"stopword_suggestions": [
|
||||||
|
{
|
||||||
|
"term": "系统",
|
||||||
|
"reason": "出现频繁但语义过泛,难以作为规则命中词。"
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"interest_keyword_suggestions": [
|
||||||
|
{
|
||||||
|
"term": "MCP",
|
||||||
|
"reason": "最近多日持续高频,且符合当前 AI Agent 学习方向。"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## 13. 周期与触发方式
|
||||||
|
|
||||||
|
建议触发周期:
|
||||||
|
|
||||||
|
- 词元统计更新:每天一次,跟随日报生成
|
||||||
|
- skill 清洗:每周一次,或人工手动触发
|
||||||
|
|
||||||
|
推荐流程:
|
||||||
|
|
||||||
|
1. 当天日报生成完成
|
||||||
|
2. 程序更新 `daily/YYYY-MM-DD.json`
|
||||||
|
3. 程序更新 `term_stats.json`
|
||||||
|
4. 每周或人工触发一次词元清洗 skill
|
||||||
|
5. skill 输出建议
|
||||||
|
6. 人工确认后再更新配置文件
|
||||||
|
|
||||||
|
## 14. 与规则引擎的关系
|
||||||
|
|
||||||
|
这套词元库设计不是规则引擎的替代品,而是规则配置的辅助层。
|
||||||
|
|
||||||
|
关系如下:
|
||||||
|
|
||||||
|
- `keywords`
|
||||||
|
- 是词元库来源
|
||||||
|
- `term_stats`
|
||||||
|
- 是观察与调优依据
|
||||||
|
- `filter_context.personal.json`
|
||||||
|
- 是真正给规则引擎使用的兴趣词配置
|
||||||
|
- `filter_rules.json`
|
||||||
|
- 是最终裁决逻辑
|
||||||
|
|
||||||
|
也就是说:
|
||||||
|
|
||||||
|
`日报词元统计 -> 清洗建议 -> 人工确认 -> 更新 interest_keywords -> 规则引擎命中`
|
||||||
|
|
||||||
|
而不是:
|
||||||
|
|
||||||
|
`日报词元统计 -> 自动改规则`
|
||||||
|
|
||||||
|
## 15. 分阶段落地建议
|
||||||
|
|
||||||
|
### Phase 1
|
||||||
|
|
||||||
|
先做最小可用版本:
|
||||||
|
|
||||||
|
- 只读取日报中的 `keywords`
|
||||||
|
- 生成每日词元文件
|
||||||
|
- 生成全局累计词频文件
|
||||||
|
- 支持 `term_aliases` 和 `term_stopwords`
|
||||||
|
|
||||||
|
### Phase 2
|
||||||
|
|
||||||
|
再补治理层:
|
||||||
|
|
||||||
|
- 增加词元清洗 skill
|
||||||
|
- 输出建议文件
|
||||||
|
- 人工确认后更新配置
|
||||||
|
|
||||||
|
### Phase 3
|
||||||
|
|
||||||
|
最后再考虑增强:
|
||||||
|
|
||||||
|
- 增加趋势分析
|
||||||
|
- 增加最近 7 天热点词视图
|
||||||
|
- 增加“建议加入兴趣词”的自动排序
|
||||||
|
|
||||||
|
## 16. 一句话结论
|
||||||
|
|
||||||
|
这套设计选择“只统计日报中的 `keywords`,由程序维护轻量词元库,再由独立 skill 周期性做清洗建议”,目的是在控制数据规模的前提下,为规则配置和长期兴趣演化提供稳定、可审计、可扩展的基础设施。
|
||||||
@@ -0,0 +1,225 @@
|
|||||||
|
# 规则过滤引擎设计
|
||||||
|
|
||||||
|
## 1. 目标
|
||||||
|
|
||||||
|
当前过滤层的定位是:
|
||||||
|
|
||||||
|
- 不让 LLM 直接做最终过滤决策
|
||||||
|
- 让 LLM 只产出结构化信号
|
||||||
|
- 由规则引擎输出最终 `keep / drop / review` 决策
|
||||||
|
|
||||||
|
当前链路是:
|
||||||
|
|
||||||
|
`item -> content extraction -> llm summary -> filter rule engine`
|
||||||
|
|
||||||
|
## 2. 输入输出
|
||||||
|
|
||||||
|
### 2.1 FilterInput
|
||||||
|
|
||||||
|
过滤层统一读取四类输入:
|
||||||
|
|
||||||
|
- `item`
|
||||||
|
- `article`
|
||||||
|
- `summary`
|
||||||
|
- `context`
|
||||||
|
|
||||||
|
其中:
|
||||||
|
|
||||||
|
- `item` 表示上游标准化候选条目
|
||||||
|
- `article` 表示正文提取结果
|
||||||
|
- `summary` 表示结构化 LLM 摘要结果
|
||||||
|
- `context` 表示额外的用户偏好或运行时上下文
|
||||||
|
|
||||||
|
### 2.2 FilterDecisionResult
|
||||||
|
|
||||||
|
过滤结果统一输出:
|
||||||
|
|
||||||
|
- `decision`
|
||||||
|
- `keep`
|
||||||
|
- `drop`
|
||||||
|
- `review`
|
||||||
|
- `matched_rules`
|
||||||
|
- `reasons`
|
||||||
|
- `labels`
|
||||||
|
- `priority`
|
||||||
|
- `matches`
|
||||||
|
|
||||||
|
这样后续的知识库 sink、推送层或人工审核都可以稳定消费。
|
||||||
|
|
||||||
|
## 3. 规则结构
|
||||||
|
|
||||||
|
当前规则文件位置:
|
||||||
|
|
||||||
|
- `configs/filter_rules.json`
|
||||||
|
|
||||||
|
单条规则结构为:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"rule_id": "keep-worth-keeping-method",
|
||||||
|
"enabled": true,
|
||||||
|
"stop_on_match": false,
|
||||||
|
"conditions_all": [
|
||||||
|
{ "field": "summary.worth_keeping", "op": "eq", "value": true },
|
||||||
|
{ "field": "summary.category", "op": "in", "value": ["方法论", "工具实践"] }
|
||||||
|
],
|
||||||
|
"action": {
|
||||||
|
"decision": "keep",
|
||||||
|
"reason": "Structured summary marked the content as worth keeping in a durable category.",
|
||||||
|
"labels": ["summary", "durable"],
|
||||||
|
"priority": 80
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## 4. 当前支持的操作符
|
||||||
|
|
||||||
|
- `eq`
|
||||||
|
- `ne`
|
||||||
|
- `in`
|
||||||
|
- `not_in`
|
||||||
|
- `contains`
|
||||||
|
- `overlap`
|
||||||
|
- `gte`
|
||||||
|
- `lte`
|
||||||
|
- `exists`
|
||||||
|
|
||||||
|
当前条件组合方式:
|
||||||
|
|
||||||
|
- `conditions_all`
|
||||||
|
- `conditions_any`
|
||||||
|
|
||||||
|
## 5. 动态上下文字段
|
||||||
|
|
||||||
|
规则支持从其他字段动态取值,例如:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{ "field": "summary.topics", "op": "overlap", "value": { "from_field": "context.interest_topics" } }
|
||||||
|
```
|
||||||
|
|
||||||
|
这允许上层 Agent 在调用 `filter_summary_result` 时,把当前关注主题动态注入,而不必把偏好硬编码在规则文件里。
|
||||||
|
|
||||||
|
## 6. 当前默认规则思路
|
||||||
|
|
||||||
|
当前默认规则分为三类:
|
||||||
|
|
||||||
|
- 质量拦截
|
||||||
|
- 低质量正文直接 `drop`
|
||||||
|
- 疑似截断或付费墙进入 `review`
|
||||||
|
- 摘要价值判断
|
||||||
|
- `资讯 + worth_keeping=false` 直接 `drop`
|
||||||
|
- `方法论/工具实践 + worth_keeping=true` 直接 `keep`
|
||||||
|
- 个性化补充信号
|
||||||
|
- `summary.topics` 与 `context.interest_topics` 重叠时提升为 `keep`
|
||||||
|
- 其他 `worth_keeping=true` 的结果默认进入 `review`
|
||||||
|
|
||||||
|
## 7. 规则执行流程
|
||||||
|
|
||||||
|
当前过滤流程按以下顺序执行:
|
||||||
|
|
||||||
|
1. 上层先产出结构化 `summary`
|
||||||
|
2. 可选附带 `item`、`article` 与 `context`
|
||||||
|
3. 规则引擎按 `priority` 从高到低遍历规则
|
||||||
|
4. 每条规则根据 `conditions_all` / `conditions_any` 判断是否命中
|
||||||
|
5. 命中的规则被收集为 `matches`
|
||||||
|
6. 最终根据命中结果收敛为 `keep / drop / review`
|
||||||
|
|
||||||
|
当前决策收敛规则是:
|
||||||
|
|
||||||
|
- 只要命中任意 `drop`,最终结果就是 `drop`
|
||||||
|
- 否则只要命中任意 `keep`,最终结果就是 `keep`
|
||||||
|
- 否则如果命中任意 `review`,最终结果就是 `review`
|
||||||
|
- 如果没有任何规则命中,默认回落到 `review`
|
||||||
|
|
||||||
|
这样设计的原因是:
|
||||||
|
|
||||||
|
- `drop` 应该拥有最高约束力
|
||||||
|
- `keep` 只在没有硬性淘汰时生效
|
||||||
|
- 默认不自动放行未知内容
|
||||||
|
|
||||||
|
## 8. 为什么让 LLM 调用规则 tool,而不是直接裁决
|
||||||
|
|
||||||
|
当前架构故意拆成两层:
|
||||||
|
|
||||||
|
- LLM 负责生成结构化信号
|
||||||
|
- 规则引擎负责输出最终过滤决策
|
||||||
|
|
||||||
|
原因是:
|
||||||
|
|
||||||
|
- LLM 适合做语义理解、归类、摘要和价值信号提取
|
||||||
|
- 规则引擎适合做稳定、可复现、可审计的最终判断
|
||||||
|
|
||||||
|
因此推荐的调用方式是:
|
||||||
|
|
||||||
|
1. LLM 先调用提取 tool
|
||||||
|
2. LLM 或本地脚本产出 `summary_result`
|
||||||
|
3. LLM 再调用 `filter_summary_result`
|
||||||
|
4. 后续根据过滤结果决定是否入库、推送或人工审核
|
||||||
|
|
||||||
|
这意味着:
|
||||||
|
|
||||||
|
- LLM 是编排者
|
||||||
|
- rule engine 是裁决器
|
||||||
|
|
||||||
|
而不是让 LLM 在过滤阶段再次自由发挥。
|
||||||
|
|
||||||
|
## 9. 调用示例
|
||||||
|
|
||||||
|
### 9.1 本地脚本
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python scripts/run_filter_rules.py ^
|
||||||
|
--summary outputs/reference/summary/result.loop.json ^
|
||||||
|
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||||
|
--output outputs/reference/filter/filter-decision.json
|
||||||
|
```
|
||||||
|
|
||||||
|
### 9.2 带上下文的调用
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"interest_topics": ["个人信息管理", "阅读工作流"]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
将这份上下文作为 `--context` 传入后,规则就可以根据当前关注主题做加权。
|
||||||
|
|
||||||
|
### 9.3 MCP tool
|
||||||
|
|
||||||
|
`filter_summary_result` 接收:
|
||||||
|
|
||||||
|
- `summary_result`
|
||||||
|
- 可选 `extracted_article`
|
||||||
|
- 可选 `item`
|
||||||
|
- 可选 `context`
|
||||||
|
|
||||||
|
返回:
|
||||||
|
|
||||||
|
- `decision`
|
||||||
|
- `matched_rules`
|
||||||
|
- `reasons`
|
||||||
|
- `labels`
|
||||||
|
- `priority`
|
||||||
|
- `matches`
|
||||||
|
|
||||||
|
## 10. 当前落地位置
|
||||||
|
|
||||||
|
- 过滤模型:
|
||||||
|
- `src/summary_mcp/models/filtering.py`
|
||||||
|
- 规则引擎:
|
||||||
|
- `src/summary_mcp/filters/engine.py`
|
||||||
|
- 默认规则:
|
||||||
|
- `configs/filter_rules.json`
|
||||||
|
- 本地运行脚本:
|
||||||
|
- `scripts/run_filter_rules.py`
|
||||||
|
- MCP tool:
|
||||||
|
- `filter_summary_result`
|
||||||
|
|
||||||
|
## 11. 当前阶段的设计结论
|
||||||
|
|
||||||
|
第一版过滤层采用“规则主导,LLM 提供信号”的架构:
|
||||||
|
|
||||||
|
- LLM 不直接做最终裁决
|
||||||
|
- 上层 Agent 可以决定是否调用过滤 tool
|
||||||
|
- 真正的过滤决策由规则引擎输出
|
||||||
|
- 灰区内容后续再考虑是否引入第二层 LLM 辅助判断
|
||||||
@@ -0,0 +1,328 @@
|
|||||||
|
# 规则引擎使用说明
|
||||||
|
|
||||||
|
## 1. 它解决什么问题
|
||||||
|
|
||||||
|
规则引擎负责把上游产出的结构化信号收敛成最终筛选决策:
|
||||||
|
|
||||||
|
`item -> content extraction -> llm summary -> rule engine -> candidate/openclaw`
|
||||||
|
|
||||||
|
这里有一个明确边界:
|
||||||
|
|
||||||
|
- LLM 负责理解正文、生成结构化摘要信号
|
||||||
|
- 规则引擎负责输出稳定、可复现、可审计的 `keep / drop / review`
|
||||||
|
|
||||||
|
也就是说,规则引擎不是“再让 LLM 判断一遍”,而是用确定性规则做最后裁决。
|
||||||
|
|
||||||
|
## 2. 相关文件
|
||||||
|
|
||||||
|
- 规则模型: `src/summary_mcp/models/filtering.py`
|
||||||
|
- 规则执行器: `src/summary_mcp/filters/engine.py`
|
||||||
|
- 默认规则: `configs/filter_rules.json`
|
||||||
|
- 本地脚本: `scripts/run_filter_rules.py`
|
||||||
|
- MCP tool: `filter_summary_result`
|
||||||
|
|
||||||
|
## 3. 输入与输出
|
||||||
|
|
||||||
|
规则引擎统一读取一个 `FilterInput`,包含四部分:
|
||||||
|
|
||||||
|
- `item`
|
||||||
|
- 标准化后的条目对象
|
||||||
|
- `article`
|
||||||
|
- 正文提取结果
|
||||||
|
- `summary`
|
||||||
|
- LLM 结构化摘要结果
|
||||||
|
- `context`
|
||||||
|
- 运行时注入的偏好信息
|
||||||
|
|
||||||
|
当前最常用的判断信号主要来自两类字段:
|
||||||
|
|
||||||
|
- `article.quality_flags.*`
|
||||||
|
- 如 `is_low_content`、`is_truncated`、`is_paywalled`
|
||||||
|
- `summary.*`
|
||||||
|
- 如 `category`、`worth_keeping`、`topics`
|
||||||
|
|
||||||
|
输出是 `FilterDecisionResult`:
|
||||||
|
|
||||||
|
- `decision`
|
||||||
|
- 最终决策,`keep / drop / review`
|
||||||
|
- `matched_rules`
|
||||||
|
- 命中的规则 ID 列表
|
||||||
|
- `reasons`
|
||||||
|
- 命中规则的原因说明
|
||||||
|
- `labels`
|
||||||
|
- 聚合后的标签
|
||||||
|
- `priority`
|
||||||
|
- 命中规则中的最高优先级
|
||||||
|
- `matches`
|
||||||
|
- 每条命中规则的明细
|
||||||
|
|
||||||
|
## 4. 规则文件怎么写
|
||||||
|
|
||||||
|
规则文件位置是 `configs/filter_rules.json`,顶层必须是一个 JSON 数组。
|
||||||
|
|
||||||
|
单条规则结构:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"rule_id": "keep-worth-keeping-method",
|
||||||
|
"enabled": true,
|
||||||
|
"stop_on_match": false,
|
||||||
|
"conditions_all": [
|
||||||
|
{ "field": "summary.worth_keeping", "op": "eq", "value": true },
|
||||||
|
{ "field": "summary.category", "op": "in", "value": ["方法论", "工具实践"] }
|
||||||
|
],
|
||||||
|
"action": {
|
||||||
|
"decision": "keep",
|
||||||
|
"reason": "Structured summary marked the content as worth keeping in a durable category.",
|
||||||
|
"labels": ["summary", "durable"],
|
||||||
|
"priority": 80
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
字段说明:
|
||||||
|
|
||||||
|
- `rule_id`
|
||||||
|
- 规则唯一标识,建议稳定命名
|
||||||
|
- `enabled`
|
||||||
|
- 是否启用
|
||||||
|
- `stop_on_match`
|
||||||
|
- 命中后是否立即停止继续匹配后续规则
|
||||||
|
- `conditions_all`
|
||||||
|
- 全部命中才算命中
|
||||||
|
- `conditions_any`
|
||||||
|
- 任意命中即可
|
||||||
|
- `action.decision`
|
||||||
|
- 该规则命中时产出的规则级决策
|
||||||
|
- `action.reason`
|
||||||
|
- 命中原因
|
||||||
|
- `action.labels`
|
||||||
|
- 打到结果里的标签
|
||||||
|
- `action.priority`
|
||||||
|
- 执行时排序优先级,越大越先执行
|
||||||
|
|
||||||
|
## 5. 当前支持的操作符
|
||||||
|
|
||||||
|
- `eq`
|
||||||
|
- `ne`
|
||||||
|
- `in`
|
||||||
|
- `not_in`
|
||||||
|
- `contains`
|
||||||
|
- `overlap`
|
||||||
|
- `gte`
|
||||||
|
- `lte`
|
||||||
|
- `exists`
|
||||||
|
|
||||||
|
常见例子:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{ "field": "summary.worth_keeping", "op": "eq", "value": true }
|
||||||
|
```
|
||||||
|
|
||||||
|
```json
|
||||||
|
{ "field": "summary.category", "op": "in", "value": ["方法论", "工具实践"] }
|
||||||
|
```
|
||||||
|
|
||||||
|
```json
|
||||||
|
{ "field": "summary.topics", "op": "overlap", "value": ["知识管理", "阅读工作流"] }
|
||||||
|
```
|
||||||
|
|
||||||
|
## 6. 执行顺序和收敛逻辑
|
||||||
|
|
||||||
|
执行顺序不是按文件书写顺序,而是按 `action.priority` 从高到低排序。
|
||||||
|
|
||||||
|
命中后会先收集所有规则,再做最终收敛:
|
||||||
|
|
||||||
|
- 只要命中过任意 `drop`,最终就是 `drop`
|
||||||
|
- 否则只要命中过任意 `keep`,最终就是 `keep`
|
||||||
|
- 否则只要命中过任意 `review`,最终就是 `review`
|
||||||
|
- 如果完全没有命中,默认 `review`
|
||||||
|
|
||||||
|
这意味着:
|
||||||
|
|
||||||
|
- `drop` 是硬拦截
|
||||||
|
- `keep` 只能在没有更高约束的 `drop` 时生效
|
||||||
|
- `review` 是默认灰区兜底
|
||||||
|
|
||||||
|
如果你希望某条高优规则一旦命中就不再继续匹配,把 `stop_on_match` 设为 `true`。
|
||||||
|
|
||||||
|
## 7. 动态上下文怎么用
|
||||||
|
|
||||||
|
规则支持从 `context` 动态取值,不需要把用户偏好写死进规则文件。
|
||||||
|
|
||||||
|
示例:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"field": "summary.topics",
|
||||||
|
"op": "overlap",
|
||||||
|
"value": { "from_field": "context.interest_topics" }
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
对应的 `context` 可以是:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"interest_topics": ["个人知识管理", "阅读工作流"]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
这样同一套规则就可以被不同用户、不同运行场景复用。
|
||||||
|
|
||||||
|
## 8. 本地怎么跑
|
||||||
|
|
||||||
|
最小调用:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python scripts/run_filter_rules.py ^
|
||||||
|
--summary outputs/reference/summary/result.loop.json ^
|
||||||
|
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||||
|
--output outputs/reference/filter/filter-decision.json
|
||||||
|
```
|
||||||
|
|
||||||
|
带上下文:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python scripts/run_filter_rules.py ^
|
||||||
|
--summary outputs/reference/summary/result.loop.json ^
|
||||||
|
--extracted outputs/reference/extracted/read-flow-2026.extracted.json ^
|
||||||
|
--context outputs/reference/filter/filter-context.json ^
|
||||||
|
--output outputs/reference/filter/filter-decision.with-context.json
|
||||||
|
```
|
||||||
|
|
||||||
|
也可以显式指定另一份规则文件:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python scripts/run_filter_rules.py ^
|
||||||
|
--summary outputs/reference/summary/result.loop.json ^
|
||||||
|
--rules configs/filter_rules.json ^
|
||||||
|
--output outputs/reference/filter/filter-decision.json
|
||||||
|
```
|
||||||
|
|
||||||
|
## 9. MCP 怎么调用
|
||||||
|
|
||||||
|
MCP tool 名称是 `filter_summary_result`。
|
||||||
|
|
||||||
|
输入参数:
|
||||||
|
|
||||||
|
- `summary_result`
|
||||||
|
- `extracted_article`
|
||||||
|
- `item`
|
||||||
|
- `context`
|
||||||
|
|
||||||
|
其中只有 `summary_result` 是必填,其余都是可选补充信号。
|
||||||
|
|
||||||
|
示例:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"summary_result": {
|
||||||
|
"title": "如何构建个人阅读工作流",
|
||||||
|
"url": "https://example.com/read-flow",
|
||||||
|
"summary": "文章介绍了从采集、提炼到沉淀的个人阅读工作流设计。",
|
||||||
|
"highlights": ["先采集再提炼", "用规则做稳定筛选", "日报再进入知识库"],
|
||||||
|
"keywords": ["阅读工作流", "知识管理", "RSS", "规则引擎", "日报"],
|
||||||
|
"topics": ["阅读工作流", "知识管理", "信息筛选"],
|
||||||
|
"category": "方法论",
|
||||||
|
"worth_keeping": true,
|
||||||
|
"reason": "提供了可复用的方法框架。"
|
||||||
|
},
|
||||||
|
"context": {
|
||||||
|
"interest_topics": ["阅读工作流", "知识管理"]
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## 10. 结果怎么看
|
||||||
|
|
||||||
|
一个典型结果会像这样:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"decision": "keep",
|
||||||
|
"matched_rules": [
|
||||||
|
"keep-worth-keeping-method",
|
||||||
|
"keep-interest-topic"
|
||||||
|
],
|
||||||
|
"reasons": [
|
||||||
|
"Structured summary marked the content as worth keeping in a durable category.",
|
||||||
|
"Topics overlap with current interest profile."
|
||||||
|
],
|
||||||
|
"labels": ["durable", "interest", "summary", "topic-match"],
|
||||||
|
"priority": 80,
|
||||||
|
"matches": [
|
||||||
|
{
|
||||||
|
"rule_id": "keep-worth-keeping-method",
|
||||||
|
"decision": "keep",
|
||||||
|
"reason": "Structured summary marked the content as worth keeping in a durable category.",
|
||||||
|
"labels": ["summary", "durable"],
|
||||||
|
"priority": 80
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
解读方式:
|
||||||
|
|
||||||
|
- 看 `decision`
|
||||||
|
- 最终裁决
|
||||||
|
- 看 `matched_rules`
|
||||||
|
- 哪些规则生效了
|
||||||
|
- 看 `reasons`
|
||||||
|
- 为什么做出这个判断
|
||||||
|
- 看 `matches`
|
||||||
|
- 需要排查时看完整命中明细
|
||||||
|
|
||||||
|
## 11. 在当前整条链路里的位置
|
||||||
|
|
||||||
|
当前生产链路里,规则引擎已经集成在 `run_freshrss_openclaw_pipeline` 中。
|
||||||
|
|
||||||
|
顺序是:
|
||||||
|
|
||||||
|
1. 从 FreshRSS 拉未读
|
||||||
|
2. 从 RSS 项目里读取正文
|
||||||
|
3. 调用 LLM 生成结构化摘要
|
||||||
|
4. 规则引擎输出 `keep / drop / review`
|
||||||
|
5. 构建 `ArticleCandidateRecord`
|
||||||
|
6. 压缩成 `OpenClawCandidateInput`
|
||||||
|
7. 生成 `openclaw-delivery-payload.json`
|
||||||
|
8. 如果开启 `mark_read`,最后再标记已读
|
||||||
|
|
||||||
|
所以在生产模式下,一般不需要单独跑 `run_filter_rules.py`,只有在调规则或排查命中逻辑时才单独跑。
|
||||||
|
|
||||||
|
## 12. 调规则时的建议
|
||||||
|
|
||||||
|
- 把“硬性淘汰”规则放高优先级
|
||||||
|
- 比如低质量正文、明显噪音内容
|
||||||
|
- 把“强 keep”规则放在中高优先级
|
||||||
|
- 比如 `worth_keeping=true` 且类别是方法论
|
||||||
|
- 把“兜底 review”规则放低一些
|
||||||
|
- 避免过早收敛
|
||||||
|
- 用户偏好尽量走 `context`
|
||||||
|
- 不要把临时兴趣直接硬编码到规则里
|
||||||
|
- `rule_id` 保持稳定
|
||||||
|
- 方便后续审计、统计和排障
|
||||||
|
|
||||||
|
## 13. 一个最常见的改法
|
||||||
|
|
||||||
|
如果你想增加一条“命中关注主题就 keep”的规则,可以直接在 `configs/filter_rules.json` 里加:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"rule_id": "keep-interest-topic",
|
||||||
|
"enabled": true,
|
||||||
|
"stop_on_match": false,
|
||||||
|
"conditions_all": [
|
||||||
|
{ "field": "summary.topics", "op": "overlap", "value": { "from_field": "context.interest_topics" } }
|
||||||
|
],
|
||||||
|
"action": {
|
||||||
|
"decision": "keep",
|
||||||
|
"reason": "Topics overlap with current interest profile.",
|
||||||
|
"labels": ["interest", "topic-match"],
|
||||||
|
"priority": 75
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
如果只是临时停用某条规则,最简单的是把它的 `enabled` 改成 `false`,不要先删规则。
|
||||||
@@ -0,0 +1,105 @@
|
|||||||
|
# Markdown Sink 设计
|
||||||
|
|
||||||
|
## 1. 目标
|
||||||
|
|
||||||
|
第一版 sink 不直接绑定外部平台,而是先落一个稳定的 `Markdown sink`。
|
||||||
|
|
||||||
|
原因:
|
||||||
|
|
||||||
|
- 可审计
|
||||||
|
- 可版本化
|
||||||
|
- 易迁移
|
||||||
|
- 可直接兼容 Obsidian / 通用 Markdown 知识库
|
||||||
|
|
||||||
|
## 2. 统一输入
|
||||||
|
|
||||||
|
Markdown sink 读取统一的 `SinkInput`:
|
||||||
|
|
||||||
|
- `item`
|
||||||
|
- `article`
|
||||||
|
- `summary`
|
||||||
|
- `filter_decision`
|
||||||
|
- `metadata`
|
||||||
|
|
||||||
|
这样未来扩展 Notion sink、Webhook sink 时,可以继续复用同一套上游输入结构。
|
||||||
|
|
||||||
|
## 3. 目录结构
|
||||||
|
|
||||||
|
当前建议目录结构:
|
||||||
|
|
||||||
|
```text
|
||||||
|
knowledge-base/
|
||||||
|
inbox/
|
||||||
|
review/
|
||||||
|
archive/
|
||||||
|
```
|
||||||
|
|
||||||
|
具体落盘规则:
|
||||||
|
|
||||||
|
- `keep` -> `knowledge-base/inbox/YYYY/YYYY-MM/`
|
||||||
|
- `review` -> `knowledge-base/review/YYYY/YYYY-MM/`
|
||||||
|
- `drop` -> `knowledge-base/archive/YYYY/YYYY-MM/`
|
||||||
|
|
||||||
|
## 4. 文件命名
|
||||||
|
|
||||||
|
文件名格式:
|
||||||
|
|
||||||
|
```text
|
||||||
|
YYYY-MM-DD-title-slug.md
|
||||||
|
```
|
||||||
|
|
||||||
|
例如:
|
||||||
|
|
||||||
|
```text
|
||||||
|
2026-03-21-kimi-cursor.md
|
||||||
|
```
|
||||||
|
|
||||||
|
## 5. Markdown 内容结构
|
||||||
|
|
||||||
|
输出采用:
|
||||||
|
|
||||||
|
- YAML frontmatter
|
||||||
|
- 摘要区
|
||||||
|
- 要点区
|
||||||
|
- 过滤决策区
|
||||||
|
- 原文信息区
|
||||||
|
- 正文摘录区
|
||||||
|
|
||||||
|
frontmatter 中保留:
|
||||||
|
|
||||||
|
- `title`
|
||||||
|
- `url`
|
||||||
|
- `source_id`
|
||||||
|
- `item_id`
|
||||||
|
- `extract_id`
|
||||||
|
- `category`
|
||||||
|
- `worth_keeping`
|
||||||
|
- `decision`
|
||||||
|
- `priority`
|
||||||
|
- `topics`
|
||||||
|
- `keywords`
|
||||||
|
- `labels`
|
||||||
|
- `published_at`
|
||||||
|
- `saved_at`
|
||||||
|
- `quality_flags`
|
||||||
|
|
||||||
|
## 6. 当前落地位置
|
||||||
|
|
||||||
|
- sink 模型:
|
||||||
|
- `src/summary_mcp/models/sink.py`
|
||||||
|
- Markdown sink:
|
||||||
|
- `src/summary_mcp/sinks/markdown.py`
|
||||||
|
- 本地运行脚本:
|
||||||
|
- `scripts/run_markdown_sink.py`
|
||||||
|
|
||||||
|
## 7. 当前阶段结论
|
||||||
|
|
||||||
|
第一版知识库不先绑定 Notion、Lumina 或数据库,而是先让过滤后的内容稳定进入 Markdown 知识库。
|
||||||
|
|
||||||
|
只要这层格式稳定,后续可以继续扩展:
|
||||||
|
|
||||||
|
- Obsidian Vault
|
||||||
|
- Logseq
|
||||||
|
- Notion
|
||||||
|
- Webhook
|
||||||
|
- 自定义数据库 sink
|
||||||
@@ -1,4 +1,4 @@
|
|||||||
# Summary Loop 工作说明
|
# Summary Loop 工作说明
|
||||||
|
|
||||||
## 1. 目的
|
## 1. 目的
|
||||||
|
|
||||||
@@ -15,7 +15,7 @@
|
|||||||
这个闭环由三部分组成:
|
这个闭环由三部分组成:
|
||||||
|
|
||||||
- 提取结果
|
- 提取结果
|
||||||
- 来自 `outputs/*.extracted.json`
|
- 来自 `outputs/reference/extracted/*.json` 或 `outputs/freshrss/extracted/*.json`
|
||||||
- 提供标题、链接、正文、质量标记
|
- 提供标题、链接、正文、质量标记
|
||||||
|
|
||||||
- LLM 摘要
|
- LLM 摘要
|
||||||
@@ -53,8 +53,8 @@
|
|||||||
|
|
||||||
脚本主要使用两个输入文件:
|
脚本主要使用两个输入文件:
|
||||||
|
|
||||||
- 提取结果:`outputs/read-flow-2026.extracted.json`
|
- 提取结果:`outputs/reference/extracted/read-flow-2026.extracted.json`
|
||||||
- 摘要 prompt:`outputs/llm-summary-prompt.txt`
|
- 摘要 prompt:`outputs/prompts/llm-summary-prompt.txt`
|
||||||
|
|
||||||
提取结果不会整包无差别塞给 LLM,而是先裁剪成更小的摘要输入:
|
提取结果不会整包无差别塞给 LLM,而是先裁剪成更小的摘要输入:
|
||||||
|
|
||||||
@@ -187,16 +187,16 @@ validator
|
|||||||
|
|
||||||
每次执行脚本都会在输出目录下留下调试痕迹:
|
每次执行脚本都会在输出目录下留下调试痕迹:
|
||||||
|
|
||||||
- `result.loop.attempt-1.raw.txt`
|
- `outputs/reference/summary/result.loop.attempt-1.raw.txt`
|
||||||
- 模型原始输出
|
- 模型原始输出
|
||||||
|
|
||||||
- `result.loop.attempt-1.json`
|
- `outputs/reference/summary/result.loop.attempt-1.json`
|
||||||
- 提取出的 JSON 结果
|
- 提取出的 JSON 结果
|
||||||
|
|
||||||
- `result.loop.attempt-1.validation.json`
|
- `outputs/reference/summary/result.loop.attempt-1.validation.json`
|
||||||
- 这一轮的校验报告
|
- 这一轮的校验报告
|
||||||
|
|
||||||
- `result.loop.json`
|
- `outputs/reference/summary/result.loop.json`
|
||||||
- 当前最终结果
|
- 当前最终结果
|
||||||
|
|
||||||
这些文件的价值是:
|
这些文件的价值是:
|
||||||
@@ -0,0 +1,488 @@
|
|||||||
|
# Article Candidate / OpenClaw Input / Daily Digest 设计
|
||||||
|
|
||||||
|
## 1. 文档目的
|
||||||
|
|
||||||
|
本文档正式定义 OpenClaw 日报链路中的三层对象:
|
||||||
|
|
||||||
|
- `ArticleCandidateRecord`
|
||||||
|
- `OpenClawCandidateInput`
|
||||||
|
- `DailyDigest`
|
||||||
|
|
||||||
|
目标不是只定义字段,而是明确每一层对象的职责边界,避免把“内部记录对象”和“发给 OpenClaw 的下游输入对象”混成同一个 payload。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. 这次为什么要改
|
||||||
|
|
||||||
|
前一版设计里,`article_candidate` 同时承担了三种职责:
|
||||||
|
|
||||||
|
1. 本地审计与追溯
|
||||||
|
2. 规则过滤后的内部候选记录
|
||||||
|
3. 发给 OpenClaw 的下游输入
|
||||||
|
|
||||||
|
这会直接带来三个问题:
|
||||||
|
|
||||||
|
- 字段过重
|
||||||
|
- `article.plain_text` 这类正文内容会显著增加 token 消耗
|
||||||
|
- 字段过细
|
||||||
|
- `filter_decision.matches`、`matched_rules`、`labels` 这类规则证据链更适合本地审计,不适合发给下游聚合器
|
||||||
|
- 字段过耦合
|
||||||
|
- `source_refs` 是本地文件系统路径,对 OpenClaw 没有意义,反而会让上下游绑定内部实现
|
||||||
|
|
||||||
|
批量跑完 `FreshRSS -> extraction -> LLM summary -> filter -> article_candidate` 后,这个问题已经非常清楚:
|
||||||
|
|
||||||
|
- 本地对象需要尽量保留信息,便于追溯误判
|
||||||
|
- OpenClaw 输入需要尽量精简,便于聚合日报和控制 token
|
||||||
|
|
||||||
|
因此,正确改法不是继续给一个 `article_candidate` 做加减法,而是把对象拆层。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. 设计结论
|
||||||
|
|
||||||
|
建议正式拆成三层:
|
||||||
|
|
||||||
|
### 3.1 `ArticleCandidateRecord`
|
||||||
|
|
||||||
|
定位:
|
||||||
|
|
||||||
|
- 内部候选记录对象
|
||||||
|
- 面向本地留档、审计、调试、复跑
|
||||||
|
|
||||||
|
它回答的问题是:
|
||||||
|
|
||||||
|
- 这条候选内容从哪里来
|
||||||
|
- 提取、摘要、过滤各阶段到底产出了什么
|
||||||
|
- 为什么规则引擎会给出当前决策
|
||||||
|
- 后续如果要重放或排查,该去哪里追溯
|
||||||
|
|
||||||
|
### 3.2 `OpenClawCandidateInput`
|
||||||
|
|
||||||
|
定位:
|
||||||
|
|
||||||
|
- 发给 OpenClaw 的精简输入对象
|
||||||
|
- 面向日报聚合与编排
|
||||||
|
|
||||||
|
它回答的问题是:
|
||||||
|
|
||||||
|
- 这条内容是什么
|
||||||
|
- 为什么值得放进候选池
|
||||||
|
- 应该如何排序和栏目化
|
||||||
|
|
||||||
|
它不负责保存本地调试信息,也不负责承载完整正文。
|
||||||
|
|
||||||
|
### 3.3 `DailyDigest`
|
||||||
|
|
||||||
|
定位:
|
||||||
|
|
||||||
|
- OpenClaw 聚合后的日级主产物
|
||||||
|
- 面向日报汇报、知识沉淀、人工 review
|
||||||
|
|
||||||
|
它回答的问题是:
|
||||||
|
|
||||||
|
- 今天最值得关注的内容是什么
|
||||||
|
- 这些内容能提炼出什么结论
|
||||||
|
- 哪些内容值得进入长期知识库
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. 推荐链路
|
||||||
|
|
||||||
|
推荐把链路明确为:
|
||||||
|
|
||||||
|
`item -> article -> summary -> filter_result -> ArticleCandidateRecord -> OpenClawCandidateInput -> DailyDigest`
|
||||||
|
|
||||||
|
这里有两个关键转换:
|
||||||
|
|
||||||
|
1. `summary + filter_result` 先生成 `ArticleCandidateRecord`
|
||||||
|
- 这是内部标准记录
|
||||||
|
2. `ArticleCandidateRecord` 再投影成 `OpenClawCandidateInput`
|
||||||
|
- 这是跨系统传输对象
|
||||||
|
|
||||||
|
这样做的本质是:
|
||||||
|
|
||||||
|
- 内部对象追求可追溯
|
||||||
|
- 下游对象追求低耦合、低 token、高可消费性
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. `ArticleCandidateRecord` 设计
|
||||||
|
|
||||||
|
### 5.1 角色定义
|
||||||
|
|
||||||
|
`ArticleCandidateRecord` 是当前项目内部的正式候选记录对象。
|
||||||
|
|
||||||
|
它不是:
|
||||||
|
|
||||||
|
- 发给 OpenClaw 的最终 payload
|
||||||
|
- 发给用户的日报消息
|
||||||
|
- 最终知识库对象
|
||||||
|
|
||||||
|
它是下游所有再加工动作之前的“内部事实底稿”。
|
||||||
|
|
||||||
|
### 5.2 设计原则
|
||||||
|
|
||||||
|
- 保留上游对象,便于调试和重放
|
||||||
|
- 保留规则证据链,便于解释误判
|
||||||
|
- 保留本地引用,便于审计
|
||||||
|
- 允许后续重新生成不同版本的下游 payload
|
||||||
|
|
||||||
|
### 5.3 建议字段
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"candidate_id": "cand:sha256:xxx",
|
||||||
|
"item": {"...": "Item"},
|
||||||
|
"article": {"...": "ExtractedArticle"},
|
||||||
|
"summary": {"...": "LlmSummaryResult"},
|
||||||
|
"filter_result": {"...": "FilterDecisionResult"},
|
||||||
|
"digest_section_hint": "insights",
|
||||||
|
"digest_rank": 60,
|
||||||
|
"review_state": "pending",
|
||||||
|
"rendered_markdown": null,
|
||||||
|
"metadata": {
|
||||||
|
"generated_at": "2026-03-25T10:00:00Z",
|
||||||
|
"pipeline_version": "v2",
|
||||||
|
"producer": "summary_mcp",
|
||||||
|
"run_id": "candidate-2026-03-25-001"
|
||||||
|
},
|
||||||
|
"source_refs": {
|
||||||
|
"item_path": "outputs/freshrss/items/batch/item-01.item.json",
|
||||||
|
"extracted_path": "outputs/freshrss/extracted/batch/item-01.extracted.json",
|
||||||
|
"summary_path": "outputs/freshrss/summary/batch/item-01/result.loop.json",
|
||||||
|
"filter_path": "outputs/freshrss/filter/batch/item-01.filter.json"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### 5.4 字段说明
|
||||||
|
|
||||||
|
- `candidate_id`
|
||||||
|
- 候选记录唯一标识
|
||||||
|
- 建议基于 `item_id` 或 `extract_id` 派生
|
||||||
|
|
||||||
|
- `item`
|
||||||
|
- 保留来源、标题、发布时间、作者等上游元数据
|
||||||
|
|
||||||
|
- `article`
|
||||||
|
- 保留正文提取结果与质量标记
|
||||||
|
- 供本地调试与重放使用
|
||||||
|
|
||||||
|
- `summary`
|
||||||
|
- 保留结构化摘要结果
|
||||||
|
- 是后续 OpenClaw 输入映射的主要来源
|
||||||
|
|
||||||
|
- `filter_result`
|
||||||
|
- 保留规则引擎的完整裁决与命中细节
|
||||||
|
- 主要用于解释和排查
|
||||||
|
|
||||||
|
- `digest_section_hint`
|
||||||
|
- 给 OpenClaw 的栏目建议
|
||||||
|
|
||||||
|
- `digest_rank`
|
||||||
|
- 给 OpenClaw 的排序信号
|
||||||
|
|
||||||
|
- `review_state`
|
||||||
|
- 记录人工复核状态
|
||||||
|
|
||||||
|
- `rendered_markdown`
|
||||||
|
- 可选的人类可读卡片
|
||||||
|
- 用于调试或 fallback 展示
|
||||||
|
|
||||||
|
- `metadata`
|
||||||
|
- 记录运行时元信息
|
||||||
|
|
||||||
|
- `source_refs`
|
||||||
|
- 仅用于本地追溯
|
||||||
|
- 不应进入跨系统 payload
|
||||||
|
|
||||||
|
### 5.5 为什么内部对象要保留正文和完整过滤结果
|
||||||
|
|
||||||
|
因为这层对象不是给 OpenClaw 直接消费的,而是给系统内部留档的。
|
||||||
|
|
||||||
|
如果这里过早裁掉字段,会丢失两种关键能力:
|
||||||
|
|
||||||
|
1. 误判排查能力
|
||||||
|
- 例如文章被误判为 `paywall` 时,必须能回看 `article` 与 `filter_result`
|
||||||
|
2. 下游重放能力
|
||||||
|
- 后面如果调整了 OpenClaw payload 或摘要策略,可以从这层重新投影,而不必重新抓取全文
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. `OpenClawCandidateInput` 设计
|
||||||
|
|
||||||
|
### 6.1 角色定义
|
||||||
|
|
||||||
|
`OpenClawCandidateInput` 是当前项目发给 OpenClaw 的正式输入对象。
|
||||||
|
|
||||||
|
它应当是:
|
||||||
|
|
||||||
|
- 扁平化的
|
||||||
|
- 精简的
|
||||||
|
- 稳定的
|
||||||
|
- 不依赖本地文件路径的
|
||||||
|
|
||||||
|
### 6.2 设计原则
|
||||||
|
|
||||||
|
- 不携带正文全文
|
||||||
|
- 不携带规则引擎的完整证据链
|
||||||
|
- 不携带本地 `source_refs`
|
||||||
|
- 只保留 OpenClaw 聚合日报真正需要的字段
|
||||||
|
|
||||||
|
### 6.3 建议字段
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"candidate_id": "cand:sha256:xxx",
|
||||||
|
"title": "套壳中国大模型撑起500亿美元估值?扒一扒 Cursor 的套壳疑云",
|
||||||
|
"url": "http://www.ruanyifeng.com/blog/2026/03/kimi-cursor.html",
|
||||||
|
"published_at": "2026-03-21T10:19:11Z",
|
||||||
|
"author": "阮一峰",
|
||||||
|
"source_name": "阮一峰博客",
|
||||||
|
"summary": "2 到 3 句话摘要",
|
||||||
|
"highlights": ["...", "...", "..."],
|
||||||
|
"keywords": ["Cursor", "Kimi K2.5"],
|
||||||
|
"topics": ["人工智能", "商业伦理"],
|
||||||
|
"category": "观点评论",
|
||||||
|
"worth_keeping": true,
|
||||||
|
"worth_reason": "有持续性的行业洞察价值",
|
||||||
|
"selection_decision": "review",
|
||||||
|
"selection_reason": "worth_keeping=true but no stronger keep rule matched",
|
||||||
|
"digest_section_hint": "insights",
|
||||||
|
"digest_rank": 60
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### 6.4 字段说明
|
||||||
|
|
||||||
|
- `candidate_id`
|
||||||
|
- 与内部记录保持同一主键,便于回溯
|
||||||
|
|
||||||
|
- `title` / `url` / `published_at` / `author` / `source_name`
|
||||||
|
- OpenClaw 做日报聚合所需的最小来源信息
|
||||||
|
|
||||||
|
- `summary` / `highlights` / `keywords` / `topics` / `category`
|
||||||
|
- OpenClaw 聚合时真正需要的语义材料
|
||||||
|
|
||||||
|
- `worth_keeping` / `worth_reason`
|
||||||
|
- 摘要层提供的长期价值判断信号
|
||||||
|
|
||||||
|
- `selection_decision` / `selection_reason`
|
||||||
|
- 规则层的压缩结果
|
||||||
|
- 用于告诉 OpenClaw:这条为什么进入候选池,以及应该怎样对待
|
||||||
|
|
||||||
|
- `digest_section_hint` / `digest_rank`
|
||||||
|
- OpenClaw 做栏目化和排序时最直接的信号
|
||||||
|
|
||||||
|
### 6.5 为什么不用正文全文
|
||||||
|
|
||||||
|
因为 OpenClaw 当前承担的是“日报聚合”角色,不是“再次全文阅读器”。
|
||||||
|
|
||||||
|
如果把 `article.plain_text` 一并发给 OpenClaw,会出现三个问题:
|
||||||
|
|
||||||
|
1. token 浪费
|
||||||
|
- 每条都携带全文,批量聚合时成本会迅速膨胀
|
||||||
|
2. 角色混乱
|
||||||
|
- OpenClaw 会被迫重新阅读原文,而不是消费上游已经压缩好的摘要结果
|
||||||
|
3. 输入不稳定
|
||||||
|
- 正文长度差异很大,会让聚合阶段的上下文更难控制
|
||||||
|
|
||||||
|
因此,正文应保留在 `ArticleCandidateRecord`,而不是进入 `OpenClawCandidateInput`。
|
||||||
|
|
||||||
|
### 6.6 为什么过滤结果要压缩
|
||||||
|
|
||||||
|
完整的 `filter_result` 适合本地审计,不适合跨系统传输。
|
||||||
|
|
||||||
|
OpenClaw 通常只需要知道:
|
||||||
|
|
||||||
|
- 这条是 `keep / review / drop` 中的哪一种
|
||||||
|
- 为什么会进入候选池
|
||||||
|
- 优先级大概是多少
|
||||||
|
|
||||||
|
它通常不需要知道:
|
||||||
|
|
||||||
|
- 具体命中了哪几条规则
|
||||||
|
- 每条规则携带了哪些 labels
|
||||||
|
- 完整的 `matches` 证据链
|
||||||
|
|
||||||
|
所以建议把规则层压缩为:
|
||||||
|
|
||||||
|
- `selection_decision`
|
||||||
|
- `selection_reason`
|
||||||
|
- `digest_rank`
|
||||||
|
|
||||||
|
### 6.7 为什么不要 `source_refs`
|
||||||
|
|
||||||
|
因为 `source_refs` 是本地实现细节,不是业务语义。
|
||||||
|
|
||||||
|
把它发给 OpenClaw 的问题有两个:
|
||||||
|
|
||||||
|
1. 没价值
|
||||||
|
- OpenClaw 无法消费本地磁盘路径
|
||||||
|
2. 高耦合
|
||||||
|
- 这会让 OpenClaw 输入隐式依赖当前项目的目录结构与输出命名方式
|
||||||
|
|
||||||
|
因此,`source_refs` 应只保留在 `ArticleCandidateRecord`。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. 两个对象之间的映射关系
|
||||||
|
|
||||||
|
建议映射规则如下:
|
||||||
|
|
||||||
|
- `OpenClawCandidateInput.candidate_id`
|
||||||
|
- 来自 `ArticleCandidateRecord.candidate_id`
|
||||||
|
- `title` / `url` / `published_at` / `author`
|
||||||
|
- 优先来自 `item`
|
||||||
|
- `source_name`
|
||||||
|
- 优先来自 `item.source_title`,没有时可退化为域名或作者来源
|
||||||
|
- `summary` / `highlights` / `keywords` / `topics` / `category`
|
||||||
|
- 来自 `summary`
|
||||||
|
- `worth_keeping` / `worth_reason`
|
||||||
|
- 来自 `summary.worth_keeping` 与 `summary.reason`
|
||||||
|
- `selection_decision`
|
||||||
|
- 来自 `filter_result.decision`
|
||||||
|
- `selection_reason`
|
||||||
|
- 可取 `filter_result.reasons[0]` 或压缩后的组合说明
|
||||||
|
- `digest_section_hint`
|
||||||
|
- 来自 `ArticleCandidateRecord.digest_section_hint`
|
||||||
|
- `digest_rank`
|
||||||
|
- 来自 `ArticleCandidateRecord.digest_rank`
|
||||||
|
|
||||||
|
简化理解:
|
||||||
|
|
||||||
|
- `ArticleCandidateRecord` 保留证据
|
||||||
|
- `OpenClawCandidateInput` 保留结论
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. `DailyDigest` 设计
|
||||||
|
|
||||||
|
### 8.1 角色定义
|
||||||
|
|
||||||
|
`DailyDigest` 是 OpenClaw 聚合多个 `OpenClawCandidateInput` 后形成的日级主产物。
|
||||||
|
|
||||||
|
它不应该再携带单篇全文,也不应该回退到内部调试结构。
|
||||||
|
|
||||||
|
### 8.2 建议字段
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"digest_id": "digest:2026-03-25",
|
||||||
|
"date": "2026-03-25",
|
||||||
|
"title": "2026-03-25 阅读日报",
|
||||||
|
"summary": "今天的输入主要围绕 AI 工具、工程实践和软件产品策略展开。",
|
||||||
|
"sections": [
|
||||||
|
{
|
||||||
|
"section_id": "insights",
|
||||||
|
"title": "观察与判断",
|
||||||
|
"summary": "今天更值得关注的是 AI 对软件行业护城河和商业包装的影响。",
|
||||||
|
"items": [
|
||||||
|
{
|
||||||
|
"candidate_id": "cand:sha256:xxx",
|
||||||
|
"title": "文章标题",
|
||||||
|
"summary": "单篇摘要压缩版",
|
||||||
|
"why_it_matters": "为什么今天值得被放进这个栏目",
|
||||||
|
"action": "值得继续跟进"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"top_items": ["cand:sha256:xxx"],
|
||||||
|
"key_takeaways": [
|
||||||
|
"摘要层和知识沉淀层应分开。",
|
||||||
|
"规则证据链应保留在本地,而不是直接下发给聚合器。"
|
||||||
|
],
|
||||||
|
"watchlist": [
|
||||||
|
"继续收敛 paywall 误判。"
|
||||||
|
],
|
||||||
|
"candidate_ids": ["cand:sha256:xxx"],
|
||||||
|
"source_refs": [
|
||||||
|
{
|
||||||
|
"candidate_id": "cand:sha256:xxx",
|
||||||
|
"url": "https://example.com/post/1"
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"editor_notes": "今天的输入以观点和周刊类内容为主。",
|
||||||
|
"stats": {
|
||||||
|
"candidate_total": 12,
|
||||||
|
"kept_total": 4,
|
||||||
|
"review_total": 6,
|
||||||
|
"dropped_total": 2
|
||||||
|
},
|
||||||
|
"metadata": {
|
||||||
|
"generated_at": "2026-03-25T21:00:00Z",
|
||||||
|
"producer": "openclaw",
|
||||||
|
"pipeline_version": "v2"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### 8.3 `sections.items[]` 子结构建议
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"candidate_id": "cand:sha256:xxx",
|
||||||
|
"title": "文章标题",
|
||||||
|
"summary": "单篇摘要压缩版",
|
||||||
|
"why_it_matters": "为什么这条内容今天值得看",
|
||||||
|
"action": "值得试用 / 值得跟进 / 值得收藏 / 待验证"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
这里最关键的不是原始摘要复述,而是 `why_it_matters`。
|
||||||
|
|
||||||
|
因为日报的价值不在于再列一遍新闻,而在于表达“今天为什么值得看”。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 9. 为什么这样设计更合理
|
||||||
|
|
||||||
|
### 9.1 控制 token 成本
|
||||||
|
|
||||||
|
把正文全文挡在 OpenClaw 之前,能显著降低聚合阶段的 token 开销。
|
||||||
|
|
||||||
|
### 9.2 保持对象职责单一
|
||||||
|
|
||||||
|
- `ArticleCandidateRecord` 负责本地审计
|
||||||
|
- `OpenClawCandidateInput` 负责下游消费
|
||||||
|
- `DailyDigest` 负责日级成品
|
||||||
|
|
||||||
|
每层只做一件事,后续更容易演进。
|
||||||
|
|
||||||
|
### 9.3 降低上下游耦合
|
||||||
|
|
||||||
|
OpenClaw 不需要知道本地 `outputs/` 目录结构,也不应该依赖规则引擎的内部细节。
|
||||||
|
|
||||||
|
### 9.4 保留调试与重放能力
|
||||||
|
|
||||||
|
内部对象依然保留全文、质量标记、完整过滤证据链,因此不会因为“下游精简”而损失工程可维护性。
|
||||||
|
|
||||||
|
### 9.5 为后续人审与知识库沉淀留出口
|
||||||
|
|
||||||
|
当日报和知识库真正接起来时:
|
||||||
|
|
||||||
|
- OpenClaw 基于精简输入做日报聚合
|
||||||
|
- 人工基于日报成品做最终判断
|
||||||
|
- 本地记录对象作为审计底稿长期保留
|
||||||
|
|
||||||
|
这个边界是清晰且可持续的。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 10. 当前阶段的实现建议
|
||||||
|
|
||||||
|
建议按这个顺序推进:
|
||||||
|
|
||||||
|
1. 在文档层正式采用这三个对象
|
||||||
|
2. 将当前代码里的 `article_candidate` 重命名或重新定位为 `ArticleCandidateRecord`
|
||||||
|
3. 新增 `OpenClawCandidateInput` 模型
|
||||||
|
4. 增加 `ArticleCandidateRecord -> OpenClawCandidateInput` 的转换逻辑
|
||||||
|
5. 让 OpenClaw 只消费精简输入对象
|
||||||
|
6. 后续再补 `DailyDigest` 的真实生成逻辑
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 11. 一句话结论
|
||||||
|
|
||||||
|
这次改动的本质,不是“删掉几个字段”,而是把“内部候选记录对象”和“发给 OpenClaw 的精简输入对象”彻底分层;只有这样,当前项目才能同时保留审计能力、控制 token 成本,并稳定服务 OpenClaw 的日报聚合。
|
||||||
@@ -0,0 +1,351 @@
|
|||||||
|
# OpenClaw Candidate Input 字段说明
|
||||||
|
|
||||||
|
## 1. 文档目的
|
||||||
|
|
||||||
|
本文档定义当前项目输出给 OpenClaw 的结构化输入对象:`OpenClawCandidateInput`。
|
||||||
|
|
||||||
|
这是一份面向 OpenClaw 消费侧的接口说明文档,不要求了解本项目内部的 `item`、`article`、`filter_result` 等实现细节。
|
||||||
|
|
||||||
|
当前定位是:
|
||||||
|
|
||||||
|
- 单篇候选内容的精简输入
|
||||||
|
- 用于 OpenClaw 做日报聚合、排序、栏目分配和后续汇报
|
||||||
|
- 不承载正文全文、不承载本地文件路径、不承载规则引擎完整证据链
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. 设计原则
|
||||||
|
|
||||||
|
这个 payload 的设计遵循以下原则:
|
||||||
|
|
||||||
|
1. 足够轻
|
||||||
|
- 不发送正文全文,避免浪费 token
|
||||||
|
2. 足够稳定
|
||||||
|
- 不暴露本地目录结构和内部中间文件路径
|
||||||
|
3. 足够可消费
|
||||||
|
- 字段扁平化,避免 OpenClaw 反复解析嵌套对象
|
||||||
|
4. 足够有判断信号
|
||||||
|
- 保留摘要、主题、价值判断、筛选结果、排序信号
|
||||||
|
5. 足够可去重
|
||||||
|
- 同时提供原始 `url` 与归一化后的 `canonical_url`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. 完整 JSON 示例
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"candidate_id": "cand:sha256:4139f277b8cb621b02b3398f8a7bd78e3f9eef7bcd320b70b8590e802540c6ca",
|
||||||
|
"title": "套壳中国大模型撑起500亿美元估值?扒一扒 Cursor 的\"套壳\"疑云",
|
||||||
|
"url": "http://www.ruanyifeng.com/blog/2026/03/kimi-cursor.html?utm_source=rss",
|
||||||
|
"canonical_url": "http://www.ruanyifeng.com/blog/2026/03/kimi-cursor.html",
|
||||||
|
"published_at": "2026-03-21T10:19:11Z",
|
||||||
|
"author": "阮一峰",
|
||||||
|
"source_name": "阮一峰的网络日志",
|
||||||
|
"language": "zh-CN",
|
||||||
|
"summary": "AI编程工具Cursor推出的Composer 2模型被证实套壳中国Kimi K2.5模型,引发侵权争议。Kimi官方确认Cursor通过Fireworks AI获得授权,不存在侵权。作者分析Cursor隐瞒事实是为了支撑其不断膨胀的估值,将其包装成大模型公司。",
|
||||||
|
"highlights": [
|
||||||
|
"Cursor的Composer 2模型被技术手段揭露实际调用的是Kimi K2.5模型。",
|
||||||
|
"Kimi官方确认Cursor通过Fireworks AI获得授权,因此不构成侵权。",
|
||||||
|
"Cursor隐瞒使用Kimi模型,被认为是为了支撑其高达500亿美元的估值。"
|
||||||
|
],
|
||||||
|
"keywords": ["Cursor", "Composer 2", "Kimi K2.5", "Fireworks AI", "AI编程工具"],
|
||||||
|
"topics": ["人工智能", "大模型", "商业伦理"],
|
||||||
|
"category": "观点评论",
|
||||||
|
"worth_keeping": true,
|
||||||
|
"worth_reason": "文章深入剖析了AI行业的热点事件,涉及技术真相、商业动机和行业趋势,具有较高的参考价值。",
|
||||||
|
"selection_decision": "review",
|
||||||
|
"selection_reason": "Worth-keeping signal is positive but no stronger keep rule matched.",
|
||||||
|
"digest_section_hint": "insights",
|
||||||
|
"digest_rank": 60
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. 字段总览
|
||||||
|
|
||||||
|
| 字段名 | 类型 | 必填 | 说明 |
|
||||||
|
| --- | --- | --- | --- |
|
||||||
|
| `candidate_id` | `string` | 是 | 单篇候选内容唯一标识 |
|
||||||
|
| `title` | `string` | 是 | 文章或内容标题 |
|
||||||
|
| `url` | `string` | 是 | 原始链接 |
|
||||||
|
| `canonical_url` | `string \| null` | 否 | 归一化链接,用于更稳定的去重 |
|
||||||
|
| `published_at` | `string \| null` | 否 | 发布时间,ISO 8601 格式 |
|
||||||
|
| `author` | `string \| null` | 否 | 作者名 |
|
||||||
|
| `source_name` | `string \| null` | 否 | 来源名称,例如博客名、站点名、Feed 名 |
|
||||||
|
| `language` | `string \| null` | 否 | 语言标记,例如 `zh-CN`、`en` |
|
||||||
|
| `summary` | `string` | 是 | 2 到 3 句话的精简摘要 |
|
||||||
|
| `highlights` | `string[]` | 是 | 3 到 5 条单句要点 |
|
||||||
|
| `keywords` | `string[]` | 是 | 5 到 8 个关键词 |
|
||||||
|
| `topics` | `string[]` | 是 | 3 到 5 个高层主题标签 |
|
||||||
|
| `category` | `string` | 是 | 内容分类 |
|
||||||
|
| `worth_keeping` | `boolean` | 是 | 摘要层给出的长期保留判断 |
|
||||||
|
| `worth_reason` | `string` | 是 | 对 `worth_keeping` 的解释 |
|
||||||
|
| `selection_decision` | `string` | 是 | 规则层压缩后的筛选决策 |
|
||||||
|
| `selection_reason` | `string \| null` | 否 | 对 `selection_decision` 的简短解释 |
|
||||||
|
| `digest_section_hint` | `string \| null` | 否 | 建议栏目 |
|
||||||
|
| `digest_rank` | `integer` | 是 | 建议排序分值,范围 `0-100` |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. 字段详细说明
|
||||||
|
|
||||||
|
### 5.1 `candidate_id`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 单篇候选内容的稳定唯一标识
|
||||||
|
|
||||||
|
约束:
|
||||||
|
- 同一条上游内容在当前系统中应保持稳定
|
||||||
|
- OpenClaw 可以把它作为去重、回溯、引用的主键
|
||||||
|
|
||||||
|
### 5.2 `title`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 候选内容标题
|
||||||
|
|
||||||
|
来源:
|
||||||
|
- 优先来自标准化 `item.title`
|
||||||
|
- 如果上游标题缺失,则退回摘要结果中的 `summary.title`
|
||||||
|
|
||||||
|
### 5.3 `url`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 原始内容链接
|
||||||
|
|
||||||
|
用途:
|
||||||
|
- OpenClaw 在日报中回链原文
|
||||||
|
- 后续人工 review 或知识库引用时回溯来源
|
||||||
|
|
||||||
|
### 5.4 `canonical_url`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 对原始 `url` 做轻量归一化后的链接
|
||||||
|
|
||||||
|
当前归一化策略:
|
||||||
|
- 去掉 fragment
|
||||||
|
- 去掉常见追踪参数,例如 `utm_*`、`fbclid`、`gclid`、`ref` 等
|
||||||
|
- 保留其他非追踪 query 参数
|
||||||
|
|
||||||
|
用途:
|
||||||
|
- 让 OpenClaw 在去重时更稳定
|
||||||
|
- 避免同一篇文章因为带不同追踪参数被当成两篇
|
||||||
|
|
||||||
|
说明:
|
||||||
|
- `url` 保留原始值,`canonical_url` 提供去重辅助
|
||||||
|
- OpenClaw 建议优先用 `canonical_url` 做去重键,不足时再结合 `candidate_id`
|
||||||
|
|
||||||
|
### 5.5 `published_at`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 内容发布时间
|
||||||
|
|
||||||
|
格式:
|
||||||
|
- ISO 8601,例如 `2026-03-21T10:19:11Z`
|
||||||
|
|
||||||
|
### 5.6 `author`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 作者名称
|
||||||
|
|
||||||
|
### 5.7 `source_name`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 内容来源名称
|
||||||
|
|
||||||
|
常见示例:
|
||||||
|
- `阮一峰的网络日志`
|
||||||
|
- `FreshRSS`
|
||||||
|
- 某个 Feed 标题
|
||||||
|
- 某个站点域名
|
||||||
|
|
||||||
|
### 5.8 `language`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 语言标记
|
||||||
|
|
||||||
|
常见示例:
|
||||||
|
- `zh`
|
||||||
|
- `zh-CN`
|
||||||
|
- `en`
|
||||||
|
|
||||||
|
用途:
|
||||||
|
- OpenClaw 后续处理多语言输入时做风格和分类控制
|
||||||
|
- 为知识库入库、多语言日报、翻译或过滤策略留出口
|
||||||
|
|
||||||
|
说明:
|
||||||
|
- 如果当前上游无法可靠识别语言,可以为 `null`
|
||||||
|
|
||||||
|
### 5.9 `summary`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 摘要主文本
|
||||||
|
|
||||||
|
约束:
|
||||||
|
- 当前上游约束是 2 到 3 句话
|
||||||
|
- 长度不超过 140 个中文字符
|
||||||
|
|
||||||
|
### 5.10 `highlights`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 关键要点列表
|
||||||
|
|
||||||
|
约束:
|
||||||
|
- 3 到 5 条
|
||||||
|
- 每条一条独立信息
|
||||||
|
|
||||||
|
### 5.11 `keywords`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 具体实体、工具名、方法名、关键概念
|
||||||
|
|
||||||
|
约束:
|
||||||
|
- 5 到 8 个
|
||||||
|
|
||||||
|
### 5.12 `topics`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 高层主题标签
|
||||||
|
|
||||||
|
约束:
|
||||||
|
- 3 到 5 个
|
||||||
|
- 语义层级高于 `keywords`
|
||||||
|
|
||||||
|
### 5.13 `category`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 内容分类
|
||||||
|
|
||||||
|
当前允许值:
|
||||||
|
- `资讯`
|
||||||
|
- `方法论`
|
||||||
|
- `工具实践`
|
||||||
|
- `观点评论`
|
||||||
|
|
||||||
|
### 5.14 `worth_keeping`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 这条内容是否具有持续保留价值
|
||||||
|
|
||||||
|
说明:
|
||||||
|
- 这是摘要层的价值判断,不是最终入库决定
|
||||||
|
- OpenClaw 可以把它当作优先筛选信号,而不是绝对真值
|
||||||
|
|
||||||
|
### 5.15 `worth_reason`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 对 `worth_keeping` 的解释
|
||||||
|
|
||||||
|
### 5.16 `selection_decision`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 规则引擎压缩后的筛选决策
|
||||||
|
|
||||||
|
当前允许值:
|
||||||
|
- `keep`
|
||||||
|
- `review`
|
||||||
|
- `drop`
|
||||||
|
|
||||||
|
### 5.17 `selection_reason`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 对 `selection_decision` 的简短解释
|
||||||
|
|
||||||
|
说明:
|
||||||
|
- 当前一般取规则引擎 reasons 的第一条压缩说明
|
||||||
|
- 不承诺包含全部规则命中细节
|
||||||
|
|
||||||
|
### 5.18 `digest_section_hint`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 建议日报栏目
|
||||||
|
|
||||||
|
当前允许值:
|
||||||
|
- `top_news`
|
||||||
|
- `tools_and_workflows`
|
||||||
|
- `risk_and_security`
|
||||||
|
- `open_source`
|
||||||
|
- `insights`
|
||||||
|
- `deep_dive`
|
||||||
|
- `null`
|
||||||
|
|
||||||
|
### 5.19 `digest_rank`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 建议排序分值
|
||||||
|
|
||||||
|
范围:
|
||||||
|
- `0` 到 `100`
|
||||||
|
|
||||||
|
建议解释:
|
||||||
|
- `80-100`:高优先级,建议优先关注
|
||||||
|
- `50-79`:中优先级,建议正常进入候选池
|
||||||
|
- `0-49`:低优先级,建议降权或仅保留参考
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. OpenClaw 消费建议
|
||||||
|
|
||||||
|
推荐 OpenClaw 主要消费这几组信号:
|
||||||
|
|
||||||
|
### 6.1 语义聚合信号
|
||||||
|
|
||||||
|
- `summary`
|
||||||
|
- `highlights`
|
||||||
|
- `keywords`
|
||||||
|
- `topics`
|
||||||
|
- `category`
|
||||||
|
|
||||||
|
### 6.2 选择与排序信号
|
||||||
|
|
||||||
|
- `worth_keeping`
|
||||||
|
- `worth_reason`
|
||||||
|
- `selection_decision`
|
||||||
|
- `selection_reason`
|
||||||
|
- `digest_rank`
|
||||||
|
- `digest_section_hint`
|
||||||
|
|
||||||
|
### 6.3 来源展示信号
|
||||||
|
|
||||||
|
- `title`
|
||||||
|
- `url`
|
||||||
|
- `canonical_url`
|
||||||
|
- `published_at`
|
||||||
|
- `author`
|
||||||
|
- `source_name`
|
||||||
|
- `language`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. 当前不会提供的字段
|
||||||
|
|
||||||
|
OpenClaw 当前不应期待以下字段:
|
||||||
|
|
||||||
|
- 正文全文,例如 `plain_text`
|
||||||
|
- 本地文件路径,例如 `source_refs`
|
||||||
|
- 完整规则证据链,例如 `filter_result.matches`
|
||||||
|
- 内部中间对象,例如完整的 `item`、`article`、`summary`、`filter_result`
|
||||||
|
- 人工确认状态,例如 `review_status`、`knowledge_decision`
|
||||||
|
|
||||||
|
最后一类字段不会放在这个对象中,原因是:
|
||||||
|
|
||||||
|
- `OpenClawCandidateInput` 表示的是上游候选事实
|
||||||
|
- 人工确认和知识沉淀属于下游流程状态
|
||||||
|
- 两者混在一起会让输入对象逐渐失真、变脏
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. 兼容性约定
|
||||||
|
|
||||||
|
当前版本约定:
|
||||||
|
|
||||||
|
- 这是单篇输入对象,不是批量 envelope
|
||||||
|
- 批量投递由上层 `OpenClawDeliveryPayload` 承载
|
||||||
|
- 本文档只约束 `candidates[]` 中单条对象的字段
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 9. 一句话结论
|
||||||
|
|
||||||
|
`OpenClawCandidateInput` 是“给 OpenClaw 聚合日报用的单篇精简对象”:它保留来源信息、摘要语义、价值判断、去重辅助和排序信号,但不会携带正文全文、本地路径、规则证据链或人工确认状态。
|
||||||
@@ -0,0 +1,222 @@
|
|||||||
|
# OpenClaw 日报聚合改造说明
|
||||||
|
|
||||||
|
## 1. 为什么要改
|
||||||
|
|
||||||
|
当前项目已经跑通了:
|
||||||
|
|
||||||
|
`FreshRSS -> item -> extraction -> LLM summary -> validator -> filter -> Markdown sink`
|
||||||
|
|
||||||
|
这条链路证明了上游拉取、结构化提取、摘要校验和规则过滤都是可行的。
|
||||||
|
|
||||||
|
但如果最终目标是:
|
||||||
|
|
||||||
|
- 每日提炼总结进入知识库
|
||||||
|
- OpenClaw 向用户汇报今天的日报
|
||||||
|
|
||||||
|
那么当前“单篇过滤后直接写入本地 Markdown 知识库”的主路径就不够准确了。
|
||||||
|
|
||||||
|
原因有四点:
|
||||||
|
|
||||||
|
1. 单篇文章更像原材料,不是最终成品
|
||||||
|
2. 用户要的是“日报级总结 + 知识沉淀”,不是“每篇都直接入库”
|
||||||
|
3. 参考文章中真正的终态是 `Digest -> Daily Review -> 人工确认 -> Lumina`
|
||||||
|
4. 当前 `article_candidate` 如果同时承担内部记录和下游投递,会导致 payload 过重、token 过高、边界混乱
|
||||||
|
|
||||||
|
所以当前项目要从“单篇直接入库”切换为“单篇候选材料 -> 日报聚合 -> 人工确认 -> 知识沉淀”。
|
||||||
|
|
||||||
|
## 2. 参考文章给出的真实结构
|
||||||
|
|
||||||
|
参考文章里的关键分层是:
|
||||||
|
|
||||||
|
1. `Digest`
|
||||||
|
- 去重、抓正文、质量检查、生成摘要
|
||||||
|
- 产出可判断的候选内容
|
||||||
|
2. `Daily Review`
|
||||||
|
- 从候选内容中做栏目化精选
|
||||||
|
- 产出当天的日报
|
||||||
|
3. `Human in the loop`
|
||||||
|
- 人工判断哪些内容值得长期保留
|
||||||
|
4. `Lumina`
|
||||||
|
- 只承接真正值得长期保留的内容
|
||||||
|
|
||||||
|
这意味着:
|
||||||
|
|
||||||
|
- 日报不是知识库的替代品
|
||||||
|
- 知识库也不应该承接所有单篇文章
|
||||||
|
- 单篇内容更适合作为日报生成前的候选材料
|
||||||
|
|
||||||
|
## 3. 当前设计哪里不够
|
||||||
|
|
||||||
|
当前的不足主要在下游:
|
||||||
|
|
||||||
|
### 3.1 `sink` 语义过重
|
||||||
|
|
||||||
|
现在的 `Markdown sink` 更像“最终入库器”。
|
||||||
|
|
||||||
|
但在新的目标里,它最多只应扮演:
|
||||||
|
|
||||||
|
- debug sink
|
||||||
|
- fallback sink
|
||||||
|
- 审计/归档 sink
|
||||||
|
|
||||||
|
主路径不应再把它视为最终知识库形态。
|
||||||
|
|
||||||
|
### 3.2 缺少“日级对象”
|
||||||
|
|
||||||
|
当前项目有:
|
||||||
|
|
||||||
|
- `item`
|
||||||
|
- `article`
|
||||||
|
- `summary`
|
||||||
|
- `filter_decision`
|
||||||
|
|
||||||
|
但还没有:
|
||||||
|
|
||||||
|
- `ArticleCandidateRecord`
|
||||||
|
- `OpenClawCandidateInput`
|
||||||
|
- `DailyDigest`
|
||||||
|
|
||||||
|
如果没有这三个对象,就无法稳定表达:
|
||||||
|
|
||||||
|
- 单篇内容如何作为内部候选材料存在
|
||||||
|
- 单篇内容如何以低 token 的形式发给 OpenClaw
|
||||||
|
- 一天的内容如何被聚合为一个正式产物
|
||||||
|
|
||||||
|
### 3.3 缺少人工确认层
|
||||||
|
|
||||||
|
规则引擎的 `keep` 只能表示“值得进入下一步”,不能直接等价于“正式入知识库”。
|
||||||
|
|
||||||
|
如果没有人工确认层,系统就会退化成“规则命中即永久沉淀”,这和参考文章强调的 `Human in the loop` 不一致。
|
||||||
|
|
||||||
|
### 3.4 下游输入边界不清楚
|
||||||
|
|
||||||
|
如果把正文全文、完整规则证据链、本地文件路径都发给 OpenClaw,会出现三个直接问题:
|
||||||
|
|
||||||
|
- token 浪费
|
||||||
|
- 对象职责混乱
|
||||||
|
- OpenClaw 输入与本地实现耦合
|
||||||
|
|
||||||
|
所以必须把“内部记录对象”和“下游精简输入对象”分开。
|
||||||
|
|
||||||
|
## 4. 改造后的主路径
|
||||||
|
|
||||||
|
建议主路径调整为:
|
||||||
|
|
||||||
|
`FreshRSS/RSS -> extraction -> LLM summary -> validator -> rule engine -> ArticleCandidateRecord -> OpenClawCandidateInput -> OpenClaw -> DailyDigest -> human review -> knowledge base`
|
||||||
|
|
||||||
|
这里要注意几件事:
|
||||||
|
|
||||||
|
1. 当前项目负责生成高质量候选材料与精简输入
|
||||||
|
2. OpenClaw 负责按天聚合、生成日报、编排后续动作
|
||||||
|
3. 知识库只接收日报和人工确认后的长期内容
|
||||||
|
|
||||||
|
## 5. 新的对象分层
|
||||||
|
|
||||||
|
### 5.1 `ArticleCandidateRecord`
|
||||||
|
|
||||||
|
定位:
|
||||||
|
|
||||||
|
- 单篇内容的内部候选记录
|
||||||
|
- 用于留档、追溯、重放、审计
|
||||||
|
|
||||||
|
最少应包含:
|
||||||
|
|
||||||
|
- `item`
|
||||||
|
- `article`
|
||||||
|
- `summary`
|
||||||
|
- `filter_result`
|
||||||
|
- `metadata`
|
||||||
|
- 可选 `source_refs`
|
||||||
|
- 可选 `rendered_markdown`
|
||||||
|
|
||||||
|
### 5.2 `OpenClawCandidateInput`
|
||||||
|
|
||||||
|
定位:
|
||||||
|
|
||||||
|
- 当前项目发给 OpenClaw 的标准投递对象
|
||||||
|
- 用于日报聚合与排序
|
||||||
|
|
||||||
|
最少应包含:
|
||||||
|
|
||||||
|
- `candidate_id`
|
||||||
|
- `title`
|
||||||
|
- `url`
|
||||||
|
- `published_at`
|
||||||
|
- `author`
|
||||||
|
- `summary`
|
||||||
|
- `highlights`
|
||||||
|
- `topics`
|
||||||
|
- `category`
|
||||||
|
- `worth_keeping`
|
||||||
|
- `selection_decision`
|
||||||
|
- `selection_reason`
|
||||||
|
- `digest_section_hint`
|
||||||
|
- `digest_rank`
|
||||||
|
|
||||||
|
它不应包含:
|
||||||
|
|
||||||
|
- `article.plain_text`
|
||||||
|
- `filter_result.matches`
|
||||||
|
- `source_refs`
|
||||||
|
|
||||||
|
### 5.3 `DailyDigest`
|
||||||
|
|
||||||
|
定位:
|
||||||
|
|
||||||
|
- 一天的主产物
|
||||||
|
- 同时服务于“知识沉淀”和“日报汇报”
|
||||||
|
|
||||||
|
建议至少包含:
|
||||||
|
|
||||||
|
- `date`
|
||||||
|
- `sections`
|
||||||
|
- `top_items`
|
||||||
|
- `key_takeaways`
|
||||||
|
- `watchlist`
|
||||||
|
- `candidate_ids`
|
||||||
|
- `source_refs`
|
||||||
|
- `editor_notes`
|
||||||
|
|
||||||
|
## 6. 为什么这样更合理
|
||||||
|
|
||||||
|
这样改造有几个直接收益:
|
||||||
|
|
||||||
|
1. 知识库不会被大量单篇摘要污染
|
||||||
|
2. OpenClaw 不需要再次消费全文,token 更可控
|
||||||
|
3. OpenClaw 的输入边界更清晰,不依赖本地目录结构和规则证据链
|
||||||
|
4. 当前项目仍然保留完整候选记录,因此误判排查和重放能力不会丢
|
||||||
|
5. 未来扩展周报、专题文章和反馈回流会更自然
|
||||||
|
|
||||||
|
本质上,这是把当前系统从“单篇落库工具”调整为“日报生产链路中的候选记录生产器 + 精简投递器”。
|
||||||
|
|
||||||
|
## 7. 对当前实现的影响
|
||||||
|
|
||||||
|
不需要推翻已有能力,主要是重新定位:
|
||||||
|
|
||||||
|
- `extraction` 保留
|
||||||
|
- `LLM summary validator` 保留
|
||||||
|
- `rule engine` 保留
|
||||||
|
- `FreshRSS integration` 保留
|
||||||
|
- `Markdown sink` 保留,但降级为调试/回退能力
|
||||||
|
|
||||||
|
真正新增的是:
|
||||||
|
|
||||||
|
- `ArticleCandidateRecord` schema
|
||||||
|
- `OpenClawCandidateInput` schema
|
||||||
|
- `ArticleCandidateRecord -> OpenClawCandidateInput` 映射逻辑
|
||||||
|
- `DailyDigest` schema
|
||||||
|
- 人工确认后的状态流转
|
||||||
|
|
||||||
|
## 8. 下一步应该做什么
|
||||||
|
|
||||||
|
建议按这个顺序推进:
|
||||||
|
|
||||||
|
1. 在代码里把当前 `article_candidate` 重新定位为 `ArticleCandidateRecord`
|
||||||
|
2. 新增 `OpenClawCandidateInput` 模型
|
||||||
|
3. 增加精简 payload 的本地输出脚本或转换函数
|
||||||
|
4. 让 OpenClaw 只消费精简输入对象
|
||||||
|
5. 再设计批量聚合和 `DailyDigest` 生成
|
||||||
|
|
||||||
|
## 9. 一句话结论
|
||||||
|
|
||||||
|
如果最终目标是“每天有日报汇报,同时把日级提炼沉淀进知识库”,那么当前项目就不该继续围绕“单篇直接入库”演进,而应拆成“内部候选记录对象 + OpenClaw 精简输入对象 + 日级成品对象”三层结构。
|
||||||
@@ -0,0 +1,195 @@
|
|||||||
|
# OpenClaw Delivery Payload 字段说明
|
||||||
|
|
||||||
|
## 1. 文档目的
|
||||||
|
|
||||||
|
本文档定义当前项目批量投递给 OpenClaw 的外层 envelope:`OpenClawDeliveryPayload`。
|
||||||
|
|
||||||
|
它不是单篇 candidate 对象,而是承载一批 `OpenClawCandidateInput` 的批次对象。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. 完整 JSON 示例
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"schema_version": "v1",
|
||||||
|
"generated_at": "2026-03-26T10:15:30Z",
|
||||||
|
"run_id": "openclaw-delivery-20260325-101530",
|
||||||
|
"date": "2026-03-25",
|
||||||
|
"candidates": [
|
||||||
|
{
|
||||||
|
"candidate_id": "cand:sha256:xxx",
|
||||||
|
"title": "文章标题",
|
||||||
|
"url": "https://example.com/post",
|
||||||
|
"canonical_url": "https://example.com/post",
|
||||||
|
"published_at": "2026-03-25T09:30:00Z",
|
||||||
|
"author": "作者",
|
||||||
|
"source_name": "来源",
|
||||||
|
"language": "zh-CN",
|
||||||
|
"summary": "2 到 3 句话摘要",
|
||||||
|
"highlights": ["要点1", "要点2", "要点3"],
|
||||||
|
"keywords": ["关键词1", "关键词2", "关键词3", "关键词4", "关键词5"],
|
||||||
|
"topics": ["主题1", "主题2", "主题3"],
|
||||||
|
"category": "观点评论",
|
||||||
|
"worth_keeping": true,
|
||||||
|
"worth_reason": "有长期参考价值",
|
||||||
|
"selection_decision": "review",
|
||||||
|
"selection_reason": "worth_keeping=true but no stronger keep rule matched",
|
||||||
|
"digest_section_hint": "insights",
|
||||||
|
"digest_rank": 60
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"stats": {
|
||||||
|
"total": 1,
|
||||||
|
"keep_total": 0,
|
||||||
|
"review_total": 1,
|
||||||
|
"drop_total": 0
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. 字段总览
|
||||||
|
|
||||||
|
| 字段名 | 类型 | 必填 | 说明 |
|
||||||
|
| --- | --- | --- | --- |
|
||||||
|
| `schema_version` | `string` | 是 | 当前 envelope 的 schema 版本 |
|
||||||
|
| `generated_at` | `string` | 是 | 本批数据的生成时间,ISO 8601 格式 |
|
||||||
|
| `run_id` | `string` | 是 | 本次批量投递的唯一运行标识 |
|
||||||
|
| `date` | `string` | 是 | 本批次所属日期,格式 `YYYY-MM-DD` |
|
||||||
|
| `candidates` | `OpenClawCandidateInput[]` | 是 | 单篇候选内容列表 |
|
||||||
|
| `stats` | `object` | 是 | 批次级统计信息 |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. 字段详细说明
|
||||||
|
|
||||||
|
### 4.1 `schema_version`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 当前批量 envelope 的结构版本
|
||||||
|
|
||||||
|
用途:
|
||||||
|
- 未来字段调整时做兼容处理
|
||||||
|
- 让 OpenClaw 能按版本选择不同解析逻辑
|
||||||
|
|
||||||
|
当前值:
|
||||||
|
- `v1`
|
||||||
|
|
||||||
|
### 4.2 `generated_at`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 本批数据的实际生成时间
|
||||||
|
|
||||||
|
格式:
|
||||||
|
- ISO 8601,例如 `2026-03-26T10:15:30Z`
|
||||||
|
|
||||||
|
用途:
|
||||||
|
- 排查数据新旧
|
||||||
|
- 增量处理
|
||||||
|
- 失败重跑和批次比对
|
||||||
|
|
||||||
|
### 4.3 `run_id`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 本次投递任务的唯一标识
|
||||||
|
|
||||||
|
用途:
|
||||||
|
- 失败重跑追踪
|
||||||
|
- 批次去重
|
||||||
|
- 运行日志关联
|
||||||
|
|
||||||
|
建议:
|
||||||
|
- 由上游在每次批量生成时唯一产生
|
||||||
|
- 例如:`openclaw-delivery-20260325-101530`
|
||||||
|
|
||||||
|
### 4.4 `date`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 本批次所属日期
|
||||||
|
|
||||||
|
格式:
|
||||||
|
- `YYYY-MM-DD`
|
||||||
|
|
||||||
|
用途:
|
||||||
|
- OpenClaw 进行日报归档
|
||||||
|
- 将多次投递映射到同一天的 digest 流程
|
||||||
|
|
||||||
|
### 4.5 `candidates`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 单篇候选输入对象列表
|
||||||
|
|
||||||
|
说明:
|
||||||
|
- 数组内每个元素都应满足 `OpenClawCandidateInput` 定义
|
||||||
|
- 单篇字段说明请参考:
|
||||||
|
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||||
|
|
||||||
|
### 4.6 `stats`
|
||||||
|
|
||||||
|
含义:
|
||||||
|
- 本批次的统计信息
|
||||||
|
|
||||||
|
当前字段包括:
|
||||||
|
- `total`
|
||||||
|
- `keep_total`
|
||||||
|
- `review_total`
|
||||||
|
- `drop_total`
|
||||||
|
|
||||||
|
用途:
|
||||||
|
- OpenClaw 在接收后快速感知这批输入的规模和分布
|
||||||
|
- 便于记录运行状态和对账
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. `stats` 子结构说明
|
||||||
|
|
||||||
|
| 字段名 | 类型 | 必填 | 说明 |
|
||||||
|
| --- | --- | --- | --- |
|
||||||
|
| `total` | `integer` | 是 | candidate 总数 |
|
||||||
|
| `keep_total` | `integer` | 是 | `selection_decision=keep` 的数量 |
|
||||||
|
| `review_total` | `integer` | 是 | `selection_decision=review` 的数量 |
|
||||||
|
| `drop_total` | `integer` | 是 | `selection_decision=drop` 的数量 |
|
||||||
|
|
||||||
|
说明:
|
||||||
|
- 这些统计值由上游根据 `candidates[]` 自动汇总
|
||||||
|
- OpenClaw 可以直接信任,也可以再次校验
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. 推荐消费方式
|
||||||
|
|
||||||
|
建议 OpenClaw 的批处理逻辑按以下顺序进行:
|
||||||
|
|
||||||
|
1. 先读取 `schema_version`
|
||||||
|
- 选择对应版本的解析逻辑
|
||||||
|
2. 再读取 `generated_at`、`run_id` 和 `date`
|
||||||
|
- 建立本次批处理上下文
|
||||||
|
3. 再读取 `stats`
|
||||||
|
- 快速判断这批输入规模与分布
|
||||||
|
4. 最后逐条消费 `candidates[]`
|
||||||
|
- 去重
|
||||||
|
- 聚类
|
||||||
|
- 分栏目
|
||||||
|
- 排序
|
||||||
|
- 产生日报和 review 清单
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. 当前不会放入 envelope 的字段
|
||||||
|
|
||||||
|
当前 envelope 不会额外放入:
|
||||||
|
|
||||||
|
- 完整内部候选记录列表
|
||||||
|
- 人工确认状态列表
|
||||||
|
- 知识库入库状态列表
|
||||||
|
- 执行日志文本
|
||||||
|
|
||||||
|
这些都属于其他层的对象,不应混进上游投递协议。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. 一句话结论
|
||||||
|
|
||||||
|
`OpenClawDeliveryPayload` 是“给 OpenClaw 的批量投递对象”:它用 `schema_version`、`generated_at`、`run_id`、`date`、`stats` 管理批次上下文,用 `candidates[]` 承载真正的单篇候选输入。
|
||||||
@@ -0,0 +1,209 @@
|
|||||||
|
# OpenClaw Handoff
|
||||||
|
|
||||||
|
## Purpose
|
||||||
|
|
||||||
|
This repository provides a FreshRSS-first reading pipeline for OpenClaw:
|
||||||
|
|
||||||
|
`FreshRSS unread items -> RSS content extraction -> LLM summary -> rule engine -> OpenClaw delivery payload`
|
||||||
|
|
||||||
|
OpenClaw should treat this repository as an MCP-backed upstream content processor.
|
||||||
|
|
||||||
|
## Production Entrypoint
|
||||||
|
|
||||||
|
OpenClaw should call the MCP tool:
|
||||||
|
|
||||||
|
- `run_freshrss_openclaw_pipeline`
|
||||||
|
|
||||||
|
This is the canonical entrypoint for production use.
|
||||||
|
|
||||||
|
## Required Environment Variables
|
||||||
|
|
||||||
|
The MCP server process must have these variables available:
|
||||||
|
|
||||||
|
- `FRESHRSS_API_BASE_URL`
|
||||||
|
- `FRESHRSS_USERNAME`
|
||||||
|
- `FRESHRSS_API_PASSWORD`
|
||||||
|
- `LLM_API_URL`
|
||||||
|
- `LLM_API_KEY`
|
||||||
|
- `LLM_MODEL`
|
||||||
|
|
||||||
|
Example:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||||||
|
set FRESHRSS_USERNAME=osiman
|
||||||
|
set FRESHRSS_API_PASSWORD=your-api-password
|
||||||
|
set LLM_API_URL=https://api.deepseek.com
|
||||||
|
set LLM_API_KEY=your-llm-api-key
|
||||||
|
set LLM_MODEL=deepseek-chat
|
||||||
|
```
|
||||||
|
|
||||||
|
## Server Startup
|
||||||
|
|
||||||
|
Install dependencies:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pip install -e .
|
||||||
|
```
|
||||||
|
|
||||||
|
Start the MCP server:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
summary-mcp
|
||||||
|
```
|
||||||
|
|
||||||
|
## Recommended MCP Call
|
||||||
|
|
||||||
|
Recommended default call:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"limit": 5,
|
||||||
|
"mark_read": false,
|
||||||
|
"include_read": false,
|
||||||
|
"debug_artifacts": false,
|
||||||
|
"timeout_seconds": 60,
|
||||||
|
"max_retries": 2
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Recommended semantics:
|
||||||
|
|
||||||
|
- Use `mark_read=false` while validating integration.
|
||||||
|
- Use `mark_read=true` only after confirming OpenClaw will consume the returned payload successfully.
|
||||||
|
- Keep `debug_artifacts=false` for routine production runs.
|
||||||
|
- Set `debug_artifacts=true` only when troubleshooting a bad batch.
|
||||||
|
|
||||||
|
## What The Tool Returns
|
||||||
|
|
||||||
|
Primary return fields:
|
||||||
|
|
||||||
|
- `run_id`
|
||||||
|
- `output_dir`
|
||||||
|
- `raw_output`
|
||||||
|
- `delivery_output`
|
||||||
|
- `report_output`
|
||||||
|
- `pulled_count`
|
||||||
|
- `delivered_count`
|
||||||
|
- `marked_read_count`
|
||||||
|
- `status_counts`
|
||||||
|
- `delivery_payload`
|
||||||
|
- `keyword_index`
|
||||||
|
|
||||||
|
Optional:
|
||||||
|
|
||||||
|
- `items`
|
||||||
|
- Returned only when `include_item_reports=true`
|
||||||
|
|
||||||
|
## Minimal Output Files
|
||||||
|
|
||||||
|
By default the pipeline writes only:
|
||||||
|
|
||||||
|
- `outputs/freshrss/rerun/<run_id>/raw/freshrss.raw.json`
|
||||||
|
- `outputs/freshrss/rerun/<run_id>/candidates/openclaw-delivery-payload.json`
|
||||||
|
- `outputs/freshrss/rerun/<run_id>/run-report.json`
|
||||||
|
|
||||||
|
It also updates local runtime keyword data:
|
||||||
|
|
||||||
|
- `data/term_index/daily/YYYY-MM-DD.json`
|
||||||
|
- `data/term_index/term_stats.json`
|
||||||
|
|
||||||
|
If `debug_artifacts=true`, the pipeline additionally writes per-item intermediate files.
|
||||||
|
|
||||||
|
## Payload Specs
|
||||||
|
|
||||||
|
OpenClaw payload field specs live here:
|
||||||
|
|
||||||
|
- `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||||
|
- `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||||
|
|
||||||
|
## Read-State Semantics
|
||||||
|
|
||||||
|
The pipeline reads from FreshRSS unread items by default.
|
||||||
|
|
||||||
|
If `mark_read=true`:
|
||||||
|
|
||||||
|
- items are marked as read only after the final `openclaw-delivery-payload.json` has been written successfully
|
||||||
|
- only successfully delivered items are marked as read
|
||||||
|
- failed or skipped items remain unread
|
||||||
|
|
||||||
|
## FreshRSS Content Policy
|
||||||
|
|
||||||
|
For FreshRSS items, the pipeline is RSS-first and does not re-crawl webpages.
|
||||||
|
|
||||||
|
Behavior:
|
||||||
|
|
||||||
|
- use `item.raw_content` first
|
||||||
|
- if missing, use `item.raw_summary`
|
||||||
|
- if neither contains usable content, skip the item
|
||||||
|
- do not fetch the original webpage again for FreshRSS items
|
||||||
|
|
||||||
|
This is intentional.
|
||||||
|
|
||||||
|
## Keyword Cleanup Governance
|
||||||
|
|
||||||
|
This repository also includes a lightweight keyword-governance flow for downstream review.
|
||||||
|
|
||||||
|
Current pieces:
|
||||||
|
|
||||||
|
- runtime keyword stats
|
||||||
|
- `data/term_index/daily/YYYY-MM-DD.json`
|
||||||
|
- `data/term_index/term_stats.json`
|
||||||
|
- governance config
|
||||||
|
- `configs/term_cleanup_policy.json`
|
||||||
|
- `configs/term_watchlist.json`
|
||||||
|
- `configs/term_change_log.json`
|
||||||
|
- review bundle builder
|
||||||
|
- `skills/keyword-cleanup-review/scripts/build_review_bundle.py`
|
||||||
|
- accepted-suggestion writer
|
||||||
|
- `scripts/apply_term_suggestions.py`
|
||||||
|
|
||||||
|
Current status:
|
||||||
|
|
||||||
|
- OpenClaw can read the keyword review bundle as maintenance input
|
||||||
|
- accepted suggestions still require explicit human confirmation
|
||||||
|
- the repository can write accepted watch / alias / stopword / interest-keyword changes after confirmation
|
||||||
|
- this governance flow is not yet wired into a periodic scheduler inside the repository
|
||||||
|
|
||||||
|
Boundary:
|
||||||
|
|
||||||
|
- keyword cleanup is a maintenance flow, not the production RSS ingestion path
|
||||||
|
- the repository does not auto-apply cleanup suggestions without confirmation
|
||||||
|
- current keyword stats are built from the delivered candidate payload, not yet from a final `DailyDigest`
|
||||||
|
|
||||||
|
## Known Limitations
|
||||||
|
|
||||||
|
- Some sources put only partial content in RSS; those items may be skipped if RSS content is insufficient.
|
||||||
|
- WeChat articles often block direct crawling, but this pipeline now avoids that path for FreshRSS items and uses RSS-provided content when available.
|
||||||
|
- Rule behavior is still conservative in some cases; many items may land in `review` depending on current rules.
|
||||||
|
- Paywall heuristics may produce false positives for some Chinese text patterns.
|
||||||
|
- Keyword cleanup governance is usable now, but periodic scheduling and before/after evaluation are not implemented yet.
|
||||||
|
|
||||||
|
## Files OpenClaw Should Read First
|
||||||
|
|
||||||
|
Recommended reading order for a new maintainer:
|
||||||
|
|
||||||
|
1. `README.md`
|
||||||
|
2. `docs/openclaw/openclaw-handoff.md`
|
||||||
|
3. `docs/openclaw/openclaw-candidate-input-field-spec.md`
|
||||||
|
4. `docs/openclaw/openclaw-delivery-payload-spec.md`
|
||||||
|
5. `docs/design/daily-keyword-index-design.md`
|
||||||
|
6. `skills/keyword-cleanup-review/SKILL.md`
|
||||||
|
7. `docs/current/context-reset-brief.md`
|
||||||
|
|
||||||
|
## Current Recommendation
|
||||||
|
|
||||||
|
For integration handoff, the repository is usable now.
|
||||||
|
|
||||||
|
The minimum you need to give OpenClaw is:
|
||||||
|
|
||||||
|
- the repository code
|
||||||
|
- the MCP server startup command
|
||||||
|
- the required environment variables in the target environment
|
||||||
|
- the instruction to call `run_freshrss_openclaw_pipeline`
|
||||||
|
|
||||||
|
If OpenClaw will also participate in keyword-governance review, additionally point it to:
|
||||||
|
|
||||||
|
- `docs/design/daily-keyword-index-design.md`
|
||||||
|
- `skills/keyword-cleanup-review/SKILL.md`
|
||||||
|
- `scripts/apply_term_suggestions.py`
|
||||||
@@ -0,0 +1,282 @@
|
|||||||
|
---
|
||||||
|
title: "信息过载时代,我的漏斗式阅读工作流"
|
||||||
|
url: "https://shawnxie.top/blogs/tools/read-flow-2026.html"
|
||||||
|
source_id: ""
|
||||||
|
item_id: ""
|
||||||
|
extract_id: "sha256:1ee177246bce46534c81718bee693fea43a4b491f545e61b3c41f22e20994057"
|
||||||
|
category: "方法论"
|
||||||
|
worth_keeping: true
|
||||||
|
decision: "keep"
|
||||||
|
priority: 80
|
||||||
|
topics:
|
||||||
|
- "个人信息管理"
|
||||||
|
- "阅读工作流"
|
||||||
|
- "知识沉淀"
|
||||||
|
keywords:
|
||||||
|
- "RSS"
|
||||||
|
- "FreshRSS"
|
||||||
|
- "OpenClaw"
|
||||||
|
- "Digest"
|
||||||
|
- "Daily Review"
|
||||||
|
- "Lumina"
|
||||||
|
- "信息漏斗"
|
||||||
|
- "人在回路"
|
||||||
|
labels:
|
||||||
|
- "durable"
|
||||||
|
- "needs-review"
|
||||||
|
- "summary"
|
||||||
|
published_at: ""
|
||||||
|
saved_at: "2026-03-25T02:24:05.936615+00:00"
|
||||||
|
quality_flags:
|
||||||
|
is_paywalled: false
|
||||||
|
is_truncated: false
|
||||||
|
is_low_content: false
|
||||||
|
---
|
||||||
|
|
||||||
|
# 摘要
|
||||||
|
|
||||||
|
作者为解决信息过载问题,构建了一套以RSS为上游、FreshRSS为聚合池、OpenClaw为编排层的漏斗式工作流,通过分层筛选、AI精选和人工精读,将信息逐步沉淀至Lumina知识库,形成稳定可控的个人信息处理闭环。
|
||||||
|
|
||||||
|
# 要点
|
||||||
|
|
||||||
|
- 以RSS为主统一信息源,通过转换工具将公众号、社交动态等非RSS内容标准化接入。
|
||||||
|
- Digest预处理负责去重、抓取正文和生成摘要,Daily Review结合AI将内容分类为日报栏目。
|
||||||
|
- 人工精读是长期价值判断的核心环节,只有经过筛选的内容才进入Lumina知识库沉淀。
|
||||||
|
- 通过轻量兴趣画像和反馈机制,系统能逐步优化筛选排序,同时避免形成信息茧房。
|
||||||
|
- 沉淀后的内容可进一步生成周刊或主题文章,将信息流转化为长期资产。
|
||||||
|
|
||||||
|
# 过滤决策
|
||||||
|
|
||||||
|
- decision: `keep`
|
||||||
|
- priority: `80`
|
||||||
|
- matched rules: keep-worth-keeping-method, review-worth-keeping-other
|
||||||
|
|
||||||
|
## reasons
|
||||||
|
|
||||||
|
- Structured summary marked the content as worth keeping in a durable category.
|
||||||
|
- Worth-keeping signal is positive but no stronger keep rule matched.
|
||||||
|
|
||||||
|
# 原文信息
|
||||||
|
|
||||||
|
- 标题: 信息过载时代,我的漏斗式阅读工作流
|
||||||
|
- 链接: https://shawnxie.top/blogs/tools/read-flow-2026.html
|
||||||
|
- 作者: unknown
|
||||||
|
- 发布时间: unknown
|
||||||
|
- 内容类型: article
|
||||||
|
|
||||||
|
# 正文摘录
|
||||||
|
|
||||||
|
信息过载时代,我的漏斗式阅读工作流
|
||||||
|
这两年我越来越强烈地感觉到,信息问题早就不是获取不到,而是处理不过来。
|
||||||
|
真正让我疲惫的,不是没东西看,而是每天都有太多东西值得看:公众号文章、技术博客、GitHub Release、AI 新闻、社区讨论、长文、短讯、碎片化观点等等,全都在争夺我的注意力。
|
||||||
|
如果不做点什么,一个人的信息生活很容易退化成这样:
|
||||||
|
- 收藏夹里躺满了文章,真正读完的却寥寥无几
|
||||||
|
- 每天浏览大量内容,能留在脑子里的却屈指可数
|
||||||
|
- 输入看似充实,输出却异常薄弱
|
||||||
|
- 以为自己一直在吸收信息,实际上只是在不同窗口间疲于切换
|
||||||
|
2026年,我开始认真整理自己的一套信息处理工作流。我不需要大而全的平台,也不想要那种号称能“一键替我读完互联网”的 AI 产品。
|
||||||
|
我只需要一套围绕自己运转的信息吸收漏斗:既能尽可能广泛地捕捉信息,又能提前过滤掉噪音和重复内容,只把真正值得投入时间的内容呈现在我面前。更重要的是,这套系统还要能将我精读后的高价值内容沉淀下来,反过来优化下一轮的信息筛选。
|
||||||
|
信息处理不追求看得更多,而是在于更稳定地吸收、判断和沉淀。
|
||||||
|
这篇文章,我将把这套工作流思路完整写下来,包括:
|
||||||
|
- 为什么我要这样做
|
||||||
|
- 它是怎么搭起来的
|
||||||
|
- 每一层分别承担什么职责
|
||||||
|
- 未来优化方向
|
||||||
|
如果你也在被信息过载折磨,或者你已经有不少输入工具,却还是觉得每天看了很多,脑子里没留下什么,那也许这篇文章会有点参考价值。
|
||||||
|
为什么要做信息吸收漏斗?
|
||||||
|
我以前也尝试过很多方法:
|
||||||
|
- 订阅很多 RSS,然后每天刷
|
||||||
|
- 把值得看的东西先扔进稍后阅读
|
||||||
|
- 用收藏、标签、笔记软件把文章存起来
|
||||||
|
- 靠搜索和回忆去找以前读过的内容
|
||||||
|
这些工具各有用处,但通常只能解决信息处理中的某一个环节。而个人信息处理真正的难点,从来都不在于单点突破,而在于如何将这些分散的环节串联起来。信息系统如果不能同时解决这两个问题,就很容易失控:
|
||||||
|
输入太多,内容太杂
|
||||||
|
几乎每一个你关心的领域,都在持续产生新内容。技术突破、产品迭代、AI进展、工具更新、商业动态、研究成果、行业变化……
|
||||||
|
而且在未经过滤之前,重要新闻、工具发布、经验复盘、广告软文、标题党、重复报道等往往都混在一起,让人难以分辨。
|
||||||
|
如果每一条内容都需要你从零开始判断其价值,那日积月累下来的认知负担会非常高。很多时候,人真正感到疲惫的不是读一篇长文,而是反复做这些微小但无意义的判断:
|
||||||
|
- 这篇文章值不值得点开?
|
||||||
|
- 这条消息是不是和刚才那条重复?
|
||||||
|
- 这个标题是不是在夸大其词?
|
||||||
|
- 这条内容到底在说什么?
|
||||||
|
- 这篇文章我只是需要了解一下,还是值得保存?
|
||||||
|
缺少精读回路
|
||||||
|
很多内容看完就过去了,既没有标记,也没有沉淀,更没有真正融入自己的知识体系。
|
||||||
|
如果信息系统永远只知道最新内容是什么,却不知道什么内容真正对你有帮助,那最多只能算是一个信息输送管道,而不是一个能越来越懂你的个性化系统。
|
||||||
|
所以我要的不是做一个更强的信息池,而是做一个更稳的信息漏斗。
|
||||||
|
桶的思路是往里装更多,漏斗的思路是让信息在往下走的过程中不断收窄,最后留下真正值得进入大脑和长期记忆的部分。
|
||||||
|
从实践效果来看,这套漏斗式信息处理工作流至少给我带来了几个显著的变化:
|
||||||
|
- 我不再需要每天从海量标题中艰难地筛选重点内容
|
||||||
|
- 低质量和重复的信息大部分在上游就被过滤掉了
|
||||||
|
- 真正值得精读的内容,会在更靠后的环节中呈现在我面前
|
||||||
|
- 我读过并认为有价值的内容,终于开始能够反过来优化后续的筛选逻辑
|
||||||
|
不是追求更快地刷完信息,而是实现更稳定地吸收知识。
|
||||||
|
为什么基于 OpenClaw 搭建?
|
||||||
|
由于流程尚处探索阶段,选择先用 OpenClaw 将整条链路串联起来,让系统先运行起来,验证可行性,发现问题后再逐步优化和完善。
|
||||||
|
OpenClaw 主要扮演编排层的角色,负责以下工作:
|
||||||
|
- 定时触发各类任务
|
||||||
|
- 串联不同的外部工具和系统
|
||||||
|
- 与飞书文档进行交互
|
||||||
|
- 与 Lumina 知识库进行联动
|
||||||
|
- 在需要的环节引入 AI 进行结构化内容生成
|
||||||
|
这种方式起步快、迭代快,非常适合边运行边优化,不需要一开始就投入大量精力构建完整的前后端系统。
|
||||||
|
如果你也想搭建一套类似的信息处理流程,需要做这些准备:
|
||||||
|
1. 可持续维护的信息源
|
||||||
|
这是整个系统的上游基础,需要先想清楚几个问题:
|
||||||
|
- 自己真正长期关注的主题领域是什么
|
||||||
|
- 这些领域的主要信息来源有哪些
|
||||||
|
- 哪些信息源值得长期订阅,哪些只是偶尔查看即可
|
||||||
|
这一步不需要追求大而全,但要尽量保持稳定和高质量。
|
||||||
|
如果你还没有构建自己的信息源,可以先参考我收集的 RSS 订阅源。
|
||||||
|
2. 统一的聚合池
|
||||||
|
我选择的聚合池是 FreshRSS,一个开源的RSS阅读器,支持多平台,支持自定义订阅,内容分类管理和API,非常适合用来进行内容存储和同步。
|
||||||
|
它并非最终阅读工具,而是信息汇聚的中转站,让所有信息源先汇聚到同一个地方,形成可持续消费的内容候选池。
|
||||||
|
3. 自动化编排层
|
||||||
|
这是 OpenClaw 最能发挥价值的地方,极大地降低了自动化流程实现的难度,通过自然语言描述诉求,在对话中逐步落地想法。我主要用它来完成以下任务:
|
||||||
|
- 定时执行 digest(自定义的预处理Skill)
|
||||||
|
- 定时执行 daily-review(自定义的内容精选Skill)
|
||||||
|
- 发布飞书文档
|
||||||
|
- 调用 Lumina(个人开源的知识库项目)
|
||||||
|
- 串联各个技能和脚本
|
||||||
|
4. 长期知识沉淀层
|
||||||
|
在我的工作流中,这一层使用 Lumina——个人开发的信息管理工作台。当然也可以使用别的笔记工具,如Notion、Obsidian等。
|
||||||
|
Lumina 只承接我经过筛选后确认值得长期保留的优质内容,这一步直接决定了后续反馈机制的质量。
|
||||||
|
一个系统最终能学到什么,很大程度上取决于你为它提供了什么样的正反馈样本。
|
||||||
|
信息处理工作流
|
||||||
|
一句话描述:用 RSS 尽可能广泛地捕捉信息,用 FreshRSS 进行稳定聚合,用 Digest 完成预处理,用 Daily Review 实现每日精选,通过人工精读判断内容的长期价值,用 Lumina 进行知识沉淀,最后将这些沉淀的价值反向转化为轻量的个性化信号。
|
||||||
|
这套流程的核心优势不是追求全自动,而是在于分层处理的设计理念。这意味着:
|
||||||
|
- 并非所有信息都值得投入时间精读
|
||||||
|
- 并非所有信息都值得长期留存
|
||||||
|
- 并非所有信息都需要进行个性化加权
|
||||||
|
- 并非所有信息都适合直接作为输出内容
|
||||||
|
一旦层次划分清晰,系统就不会退化为单一维度的推荐流,而会演变成一个真正的认知加工流程。漏斗的真正价值不在于让信息越来越少,而在于实现了这几个关键目标:
|
||||||
|
- 上游宽广:确保不会错过真正重要的信息变化
|
||||||
|
- 中游稳定:有效过滤噪音,避免其直接干扰注意力系统
|
||||||
|
- 下游精准:让我不必在不值得精读的内容上浪费时间
|
||||||
|
- 回流轻柔:通过轻量反馈机制,避免系统演变成封闭的信息茧房
|
||||||
|
个人信息系统核心能力不是如何接入更多信息源,而是要明确:哪些信息值得进入系统,哪些信息值得投入时间精读,哪些内容值得长期留存,哪些知识最终真正融入了你的思考和行动?
|
||||||
|
当能够清晰地回答这些问题时,就不再是互联网信息的被动接收者,你将拥有一套真正属于自己的认知处理系统。接下来将按顺序详细介绍每一层的工作原理。
|
||||||
|
信息源:以 RSS 为主,但不局限于原生 RSS
|
||||||
|
整个系统的最上游是信息源,主要依赖 RSS。
|
||||||
|
RSS 是最被低估的个人信息基础设施。其天然符合个人信息系统最核心的几个要求:
|
||||||
|
- 订阅权完全掌握在自己手中
|
||||||
|
- 信息来源清晰明确
|
||||||
|
- 更新内容结构化
|
||||||
|
- 不受平台推荐算法直接支配
|
||||||
|
- 便于程序接入和自动化处理
|
||||||
|
然而在现实中,有一个不可避免的问题:并非所有值得关注的信息源都提供 RSS 订阅。例如:部分公众号内容、某些社区的特定栏目、垂直网站的更新页面和社交平台上的账号动态。
|
||||||
|
因此,需要在信息源进入流程之前,尽可能将其统一转换为 RSS 或类似 feed 的格式。常见的方案有:
|
||||||
|
- RSSHub / RSS-Bridge 等转换工具:能够将大量原本不提供 RSS 订阅的内容源,转换为可以被订阅和程序自动化消费的 feed 格式;
|
||||||
|
- GitHub feed:项目的 Releases、Commits、Discussions 等,都天然提供了结构化的 feed 接口;
|
||||||
|
- wewe-rss:能够将公众号文章转换成RSS订阅源;
|
||||||
|
- nitter:将 Twitter/X 动态转换为RSS源;
|
||||||
|
- 自定义抓取脚本:对于没有现成 RSS 解决方案的页面,也可以自行编写轻量级的抓取脚本,将其转换为内部可用的 feed 格式。
|
||||||
|
这一步的原则很简单:上游信息源可以多种多样,但在进入系统之前,格式必须尽可能统一。只有这样,下游的预处理和筛选环节才能稳定可靠地运行。
|
||||||
|
更多信息源归一处理方案可参考之前文章碎片时间刷文章!懒人阅读方案分享。
|
||||||
|
聚合池:用 FreshRSS 打造稳定的“中间水库”
|
||||||
|
所有订阅源最终都会汇聚到 FreshRSS 中,它并非"我每天真正坐下来阅读的地方",而是流程的缓冲层,主要有以下作用:
|
||||||
|
实现信息来源的统一管理
|
||||||
|
无论内容来自博客、社区、GitHub 还是转换后的 feed,最终都以统一的格式呈现,成为可被消费的标准化对象。
|
||||||
|
让下游环节无需直接对接互联网
|
||||||
|
digest 模块无需再逐个访问各个网站抓取今日更新内容,只需从 FreshRSS 这个统一的内容池中获取未读候选即可。大大降低了系统各环节之间的耦合度。
|
||||||
|
确保了信息处理的时间连续性
|
||||||
|
信息系统最忌讳的是“今天临时查看一下、明天就忘了、后天又重新开始”的碎片化处理方式。FreshRSS 提供了一个稳定的时间窗口,使后续任务能够按照固定节奏有序运行。
|
||||||
|
我的目标不是追求无边界的信息获取,而是实现有边界的信息处理。
|
||||||
|
预处理:Digest 将海量候选内容转化为可判断对象
|
||||||
|
如果把 FreshRSS 比作蓄水池,那么 Digest 就是这套系统中的第一道加工厂,专注于完成内容预处理任务。
|
||||||
|
让人每天真正感到疲惫的,往往不是精读一篇高质量文章,而是反复判断大量低质量内容是否值得投入时间。
|
||||||
|
核心处理任务
|
||||||
|
- URL 精确去重
|
||||||
|
- 相似内容去重
|
||||||
|
- 正文抓取
|
||||||
|
- 质量检查
|
||||||
|
- 噪音过滤
|
||||||
|
- 摘要生成
|
||||||
|
- 初步排序与文档输出
|
||||||
|
在真正的“阅读”开始之前,先将互联网上天然混乱的信息整理成一批更具可读性的候选内容。
|
||||||
|
为什么重要?
|
||||||
|
如果没有预处理环节,后续的所有精选工作都将建立在一堆未经整理的原始标题之上。而经过 Digest 处理后,情况会改善:
|
||||||
|
- 标题党内容大幅减少
|
||||||
|
- 重复报道得到有效过滤
|
||||||
|
- 无法获取正文的无效数据明显减少
|
||||||
|
- 每条候选内容至少附带一个可供快速判断的摘要
|
||||||
|
这将显著降低后续阅读的认知成本。Digest 的职责并非编辑终稿,而是将候选内容池整理干净、结构化,为后续的精选环节奠定基础。
|
||||||
|
AI 精选:Daily Review 为我呈现重点关注内容
|
||||||
|
如果说 Digest 的作用是将原始候选内容处理得更具可读性,那么 Daily Review 则是从这些经过预处理的候选中,进一步提炼出当天真正值得关注的精华内容。
|
||||||
|
这一步让流程从预处理迈向编辑阶段。不再追求内容的广度,而是致力于打造重点更突出、结构更清晰的阅读体验,让最终产物更像一份真正意义上的日报,而非简单堆砌的摘要集合。
|
||||||
|
目前,我将 Daily Review 设计为几个固定栏目,将不同类型的信息分配到不同的认知槽位中,包含:
|
||||||
|
- 今日大事:聚焦具有公共重要性的事件
|
||||||
|
- 变更与实践:关注对个人有直接操作价值的内容
|
||||||
|
- 安全与风险:警惕潜在的风险因素
|
||||||
|
- 开源与工具:追踪工具生态的发展变化
|
||||||
|
- 洞察与数据点:把握行业趋势和关键数据
|
||||||
|
- 主题深挖:将单一新闻事件提升到趋势层面进行分析
|
||||||
|
LLM 赋能
|
||||||
|
从 Digest 输出的候选内容,到 Daily Review 最终的成稿层,中间有很多工作适合 AI 来完成:
|
||||||
|
- 合并同一事件的多个信息来源
|
||||||
|
- 提炼核心主题
|
||||||
|
- 为内容分类并分配到对应栏目
|
||||||
|
- 识别值得深入探讨的话题
|
||||||
|
- 将零散的候选内容重组为人类可以快速阅读的日报结构
|
||||||
|
但我对 AI 在这一环节的应用始终保持克制。不是让 AI 全自动生成日报,而是扮演结构化整理者的角色,从候选内容中筛选出当天和我相关的精华部分。
|
||||||
|
精读留存:将 Human in the loop 置于系统核心
|
||||||
|
前面所有流程,都是为将信息筛选到值得精读的阶段。真正让系统不至于退化为另一种自动化信息流的,正是这一关键环节:Human in the loop(人在回路中)。
|
||||||
|
我坚信,在个人知识系统中,最不能完全外包的是长期价值判断能力。
|
||||||
|
系统可以协助我完成许多任务,如信息收集、去重、摘要、聚类、排序和精选等。但它无法完全替我做出关键决策:
|
||||||
|
- 哪些内容真正值得纳入长期知识库
|
||||||
|
- 哪些内容在未来会持续发挥价值
|
||||||
|
- 哪些内容会对我的写作和判断产生深远影响
|
||||||
|
因此,在我的工作流中,Lumina 前面始终设有一道人工筛选门槛。我会从之前各个环节产生的内容中,挑选出真正值得精读和长期留存的内容。
|
||||||
|
这是整套系统中最关键的价值确认环节。能够进入 Lumina 的内容,不仅代表我看过,更意味着我认为这篇内容值得在未来持续为我所用。
|
||||||
|
个人画像:让系统学会识别对我真正有价值的内容
|
||||||
|
如果流程到 Lumina 就戛然而止,那么它仍然只是一个单纯的过滤和沉淀系统。只有当沉淀的内容开始反过来影响上游的信息选择时,整个系统才真正形成了闭环。
|
||||||
|
为此我加入了一层设计较为克制的兴趣画像逻辑。
|
||||||
|
刻意控制了影响力,不希望整个系统演变成另一个猜你喜欢的推荐引擎。我只需要它能稍微更懂我,但又不过度迎合我的偏好,保留对公共重要性内容的关注和探索未知领域的空间。
|
||||||
|
这层画像目前主要承担轻量 rerank 的功能,作为辅助信号,轻微影响 Digest 和 Daily Review 环节中候选内容的排序。主要从以下几个维度逐步学习我的偏好:
|
||||||
|
- 长期精读的主题领域
|
||||||
|
- 能稳定提供价值的信息源
|
||||||
|
- 偏好的内容格式
|
||||||
|
- 最终存入 Lumina 知识库的内容类型
|
||||||
|
轻量引导的设计,旨在减少无效信息对注意力的浪费,同时避免构建封闭的信息茧房。个性化推荐应帮助我们减少无意义的判断,而非让我们躲进舒适区。
|
||||||
|
沉淀:将信息流转化为长期内容资产
|
||||||
|
信息漏斗的终点不是读完,而是沉淀。目前从两个方向拓展沉淀的价值:
|
||||||
|
1. 周刊生成
|
||||||
|
当系统积累了一周的高质量内容后,就不应再局限于每天生成一份日报。一周的时间跨度非常适合进行复盘总结。此时,系统已经拥有了丰富的素材:
|
||||||
|
- 一周的 Digest 预处理结果
|
||||||
|
- 一周的 Daily Review 精选内容
|
||||||
|
- 若干经过价值确认的 Lumina 知识库内容
|
||||||
|
- 一些开始反复出现的热门主题
|
||||||
|
此时生成周刊,比单纯从网页上抓取热点更有价值。因为这些内容已经经过了个人筛选和沉淀,带有明显的个性化痕迹。周刊不应仅仅是本周发生了什么的简单罗列,更应该具备深度和价值:
|
||||||
|
- 本周有哪些主题值得重点关注和记忆
|
||||||
|
- 哪些变化只是短期噪音,哪些是值得重视的长期信号
|
||||||
|
- 哪些内容具有长期参考价值,值得反复回看
|
||||||
|
周刊合集👉🏻:肖恩技术周刊
|
||||||
|
2. 主题文章生成
|
||||||
|
另一个方向是让系统能够逐渐识别哪些主题已经积累了足够的素材,值得写成长篇文章。
|
||||||
|
虽然信息流中的内容看起来是离散的,但如果拉长时间维度观察,就会发现很多内容其实都在指向同一个核心主题,例如:
|
||||||
|
- AI Agent 工程的发展趋势
|
||||||
|
- 开源工具链的演变
|
||||||
|
- 内容平台分发机制的变革
|
||||||
|
- 隐私保护、合规要求与数据治理
|
||||||
|
一旦某个主题在一段时间内反复出现,并且我多次对相关内容进行精读、收藏和沉淀,那么它就不应再仅仅是多条零散的新闻,而应该逐渐发展成为一个可以深入挖掘和输出的长期主题。
|
||||||
|
内容合集👉🏻:今日观察
|
||||||
|
未来迭代方向
|
||||||
|
尽管工作流目前已经能够稳定运行,但它远未达到最终完成的状态。可以继续完善的方面有:
|
||||||
|
细化反馈机制
|
||||||
|
目前系统中最强的反馈信号是这篇内容是否被存入 Lumina。但在理想状态下,我希望系统能够逐渐识别更多层次的用户行为:
|
||||||
|
- 点开内容但未读完
|
||||||
|
- 读完内容但未收藏
|
||||||
|
- 内容值得精读
|
||||||
|
- 内容值得长期沉淀
|
||||||
|
- 内容最终影响了写作、决策或实际实现
|
||||||
|
- 不喜欢的内容(负反馈)
|
||||||
|
一旦这些层次的反馈机制得以完善,兴趣画像将变得更加精准,不再只是一个粗粒度的偏好集合。
|
||||||
|
进一步抽象流程
|
||||||
|
目前,我更倾向于继续在现有的 OpenClaw 体系内迭代优化,这是探索阶段最适合的方式。
|
||||||
|
但如果这套流程能够变得更加稳定,职责边界也更加清晰,那么抽象出其中的信息处理内核,会更有利于后续工程化迭代。
|
||||||
|
不过比起将其产品化,还是先让这套漏斗系统持续稳定地运转,越来越懂我。
|
||||||
|
结语
|
||||||
|
信息处理的关键,从来不是看到更多,而是让真正重要的信息被自己接住。
|
||||||
|
当信息经过筛选、精读、沉淀和反馈,真正融入知识结构时,我们不再是信息的被动消费者,将真正成为自己信息环境的主人。
|
||||||
|
感谢阅读
|
||||||
|
微信公众号「肖恩聊技术」
|
||||||
|
如果这篇文章对你有帮助,欢迎扫码关注,获取原创文章推送。
|
||||||
@@ -0,0 +1,83 @@
|
|||||||
|
# Outputs Layout
|
||||||
|
|
||||||
|
`outputs/` 按来源和运行批次组织。当前推荐区分生产产物与调试产物。
|
||||||
|
|
||||||
|
## 生产默认输出
|
||||||
|
|
||||||
|
对于 FreshRSS 主流水线,默认只保留这 3 类文件:
|
||||||
|
|
||||||
|
- `outputs/freshrss/rerun/<run_id>/raw/freshrss.raw.json`
|
||||||
|
- FreshRSS 原始响应快照
|
||||||
|
- `outputs/freshrss/rerun/<run_id>/candidates/openclaw-delivery-payload.json`
|
||||||
|
- 发给 OpenClaw 的最终批量 payload
|
||||||
|
- `outputs/freshrss/rerun/<run_id>/run-report.json`
|
||||||
|
- 本次批处理的运行报告
|
||||||
|
|
||||||
|
这是当前推荐的生产模式。
|
||||||
|
|
||||||
|
## 词元统计运行数据
|
||||||
|
|
||||||
|
词元统计不放在 `outputs/` 下,而放在单独的运行数据目录:
|
||||||
|
|
||||||
|
- `data/term_index/daily/YYYY-MM-DD.json`
|
||||||
|
- 某一天的日报级 `keywords` 聚合结果
|
||||||
|
- `data/term_index/term_stats.json`
|
||||||
|
- 全局累计词频统计
|
||||||
|
|
||||||
|
这两类文件属于可重建的本地运行数据,不作为仓库长期跟踪产物。
|
||||||
|
|
||||||
|
## 调试扩展输出
|
||||||
|
|
||||||
|
当启用 `debug_artifacts` 时,才会额外写出这些中间文件:
|
||||||
|
|
||||||
|
- `outputs/freshrss/rerun/<run_id>/items/`
|
||||||
|
- 标准化 `item` 中间文件
|
||||||
|
- `outputs/freshrss/rerun/<run_id>/extracted/`
|
||||||
|
- 提取结果
|
||||||
|
- `outputs/freshrss/rerun/<run_id>/summary/`
|
||||||
|
- LLM 总结结果及调试尝试文件
|
||||||
|
- `outputs/freshrss/rerun/<run_id>/filter/`
|
||||||
|
- 规则过滤结果
|
||||||
|
- `outputs/freshrss/rerun/<run_id>/candidates/*.article-candidate-record.json`
|
||||||
|
- 内部候选记录
|
||||||
|
- `outputs/freshrss/rerun/<run_id>/candidates/*.openclaw-candidate-input.json`
|
||||||
|
- 单篇 OpenClaw 输入对象
|
||||||
|
|
||||||
|
## 何时启用调试产物
|
||||||
|
|
||||||
|
只在这些场景启用:
|
||||||
|
|
||||||
|
- 某一批次有异常,需要定位是提取、总结还是过滤阶段出错
|
||||||
|
- 需要核对某条内容为什么被 `keep` / `review` / `drop`
|
||||||
|
- 需要为规则调整收集样本
|
||||||
|
|
||||||
|
平时不要默认开启。
|
||||||
|
|
||||||
|
## 参考目录
|
||||||
|
|
||||||
|
- `outputs/prompts/`
|
||||||
|
- Prompt 模板
|
||||||
|
- `outputs/reference/`
|
||||||
|
- 手工样例、固定参考输入输出
|
||||||
|
- `outputs/freshrss/`
|
||||||
|
- FreshRSS 实跑批次输出
|
||||||
|
|
||||||
|
## 命名原则
|
||||||
|
|
||||||
|
- 每次批处理使用独立目录:`outputs/freshrss/rerun/<timestamp>/`
|
||||||
|
- 默认优先保留最终产物,不把所有中间文件都当成长期资产
|
||||||
|
- 调试文件只在需要时生成
|
||||||
|
- `reference/` 只放可复用样例,不放真实生产批次数据
|
||||||
|
- `data/term_index/` 只保存本地词元统计运行数据,不纳入 git 跟踪
|
||||||
|
|
||||||
|
## 当前建议
|
||||||
|
|
||||||
|
如果目标是交给 OpenClaw 稳定消费,优先依赖:
|
||||||
|
|
||||||
|
- `openclaw-delivery-payload.json`
|
||||||
|
- `run-report.json`
|
||||||
|
- `freshrss.raw.json`
|
||||||
|
- `data/term_index/daily/YYYY-MM-DD.json`
|
||||||
|
- `data/term_index/term_stats.json`
|
||||||
|
|
||||||
|
其中前 3 个是本次批处理产物,后 2 个是持续累积的词元统计数据。
|
||||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1,40 @@
|
|||||||
|
{
|
||||||
|
"candidate_id": "cand:sha256:4139f277b8cb621b02b3398f8a7bd78e3f9eef7bcd320b70b8590e802540c6ca",
|
||||||
|
"title": "套壳中国大模型撑起500亿美元估值?扒一扒 Cursor 的\"套壳\"疑云",
|
||||||
|
"url": "http://www.ruanyifeng.com/blog/2026/03/kimi-cursor.html",
|
||||||
|
"published_at": "2026-03-21T10:19:11Z",
|
||||||
|
"author": "阮一峰",
|
||||||
|
"source_name": "阮一峰的网络日志",
|
||||||
|
"summary": "AI编程工具Cursor推出的Composer 2模型被证实套壳中国Kimi K2.5模型,引发侵权争议。Kimi官方确认Cursor通过Fireworks AI获得授权,不存在侵权。作者分析Cursor隐瞒事实是为了支撑其不断膨胀的估值,将其包装成大模型公司。",
|
||||||
|
"highlights": [
|
||||||
|
"Cursor的Composer 2模型被技术手段揭露实际调用的是Kimi K2.5模型。",
|
||||||
|
"Kimi官方确认Cursor通过Fireworks AI获得授权,因此不构成侵权。",
|
||||||
|
"Cursor隐瞒使用Kimi模型,被认为是为了支撑其高达500亿美元的估值。",
|
||||||
|
"事件显示中国大模型技术已具备输出能力,国产模型与国外旗舰差距缩小。",
|
||||||
|
"Composer 2性能低于GPT-5.4但成本最低,生成速度较快。"
|
||||||
|
],
|
||||||
|
"keywords": [
|
||||||
|
"Cursor",
|
||||||
|
"Composer 2",
|
||||||
|
"Kimi K2.5",
|
||||||
|
"Fireworks AI",
|
||||||
|
"套壳模型",
|
||||||
|
"模型授权",
|
||||||
|
"估值泡沫",
|
||||||
|
"AI编程工具"
|
||||||
|
],
|
||||||
|
"topics": [
|
||||||
|
"人工智能",
|
||||||
|
"大模型",
|
||||||
|
"商业伦理",
|
||||||
|
"技术争议",
|
||||||
|
"创业融资"
|
||||||
|
],
|
||||||
|
"category": "观点评论",
|
||||||
|
"worth_keeping": true,
|
||||||
|
"worth_reason": "文章深入剖析了AI行业的热点事件,涉及技术真相、商业动机和行业趋势,具有较高的参考价值。",
|
||||||
|
"selection_decision": "review",
|
||||||
|
"selection_reason": "Worth-keeping signal is positive but no stronger keep rule matched.",
|
||||||
|
"digest_section_hint": "insights",
|
||||||
|
"digest_rank": 60
|
||||||
|
}
|
||||||
@@ -0,0 +1,39 @@
|
|||||||
|
{
|
||||||
|
"decision": "keep",
|
||||||
|
"matched_rules": [
|
||||||
|
"keep-worth-keeping-method",
|
||||||
|
"review-worth-keeping-other"
|
||||||
|
],
|
||||||
|
"reasons": [
|
||||||
|
"Structured summary marked the content as worth keeping in a durable category.",
|
||||||
|
"Worth-keeping signal is positive but no stronger keep rule matched."
|
||||||
|
],
|
||||||
|
"labels": [
|
||||||
|
"durable",
|
||||||
|
"needs-review",
|
||||||
|
"summary"
|
||||||
|
],
|
||||||
|
"priority": 80,
|
||||||
|
"matches": [
|
||||||
|
{
|
||||||
|
"rule_id": "keep-worth-keeping-method",
|
||||||
|
"decision": "keep",
|
||||||
|
"reason": "Structured summary marked the content as worth keeping in a durable category.",
|
||||||
|
"labels": [
|
||||||
|
"summary",
|
||||||
|
"durable"
|
||||||
|
],
|
||||||
|
"priority": 80
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"rule_id": "review-worth-keeping-other",
|
||||||
|
"decision": "review",
|
||||||
|
"reason": "Worth-keeping signal is positive but no stronger keep rule matched.",
|
||||||
|
"labels": [
|
||||||
|
"summary",
|
||||||
|
"needs-review"
|
||||||
|
],
|
||||||
|
"priority": 60
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
@@ -0,0 +1,53 @@
|
|||||||
|
{
|
||||||
|
"decision": "keep",
|
||||||
|
"matched_rules": [
|
||||||
|
"keep-worth-keeping-method",
|
||||||
|
"keep-interest-topic",
|
||||||
|
"review-worth-keeping-other"
|
||||||
|
],
|
||||||
|
"reasons": [
|
||||||
|
"Structured summary marked the content as worth keeping in a durable category.",
|
||||||
|
"Topics overlap with current interest profile.",
|
||||||
|
"Worth-keeping signal is positive but no stronger keep rule matched."
|
||||||
|
],
|
||||||
|
"labels": [
|
||||||
|
"durable",
|
||||||
|
"interest",
|
||||||
|
"needs-review",
|
||||||
|
"summary",
|
||||||
|
"topic-match"
|
||||||
|
],
|
||||||
|
"priority": 80,
|
||||||
|
"matches": [
|
||||||
|
{
|
||||||
|
"rule_id": "keep-worth-keeping-method",
|
||||||
|
"decision": "keep",
|
||||||
|
"reason": "Structured summary marked the content as worth keeping in a durable category.",
|
||||||
|
"labels": [
|
||||||
|
"summary",
|
||||||
|
"durable"
|
||||||
|
],
|
||||||
|
"priority": 80
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"rule_id": "keep-interest-topic",
|
||||||
|
"decision": "keep",
|
||||||
|
"reason": "Topics overlap with current interest profile.",
|
||||||
|
"labels": [
|
||||||
|
"interest",
|
||||||
|
"topic-match"
|
||||||
|
],
|
||||||
|
"priority": 75
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"rule_id": "review-worth-keeping-other",
|
||||||
|
"decision": "review",
|
||||||
|
"reason": "Worth-keeping signal is positive but no stronger keep rule matched.",
|
||||||
|
"labels": [
|
||||||
|
"summary",
|
||||||
|
"needs-review"
|
||||||
|
],
|
||||||
|
"priority": 60
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
@@ -0,0 +1,402 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
from datetime import UTC, datetime
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
|
||||||
|
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||||
|
DEFAULT_ALIASES_PATH = REPO_ROOT / "configs" / "term_aliases.json"
|
||||||
|
DEFAULT_STOPWORDS_PATH = REPO_ROOT / "configs" / "term_stopwords.json"
|
||||||
|
DEFAULT_CONTEXT_PATH = REPO_ROOT / "configs" / "filter_context.personal.json"
|
||||||
|
DEFAULT_WATCHLIST_PATH = REPO_ROOT / "configs" / "term_watchlist.json"
|
||||||
|
DEFAULT_CHANGE_LOG_PATH = REPO_ROOT / "configs" / "term_change_log.json"
|
||||||
|
|
||||||
|
|
||||||
|
def _load_json(path: Path, default: Any) -> Any:
|
||||||
|
if not path.exists():
|
||||||
|
return default
|
||||||
|
return json.loads(path.read_text(encoding="utf-8-sig"))
|
||||||
|
|
||||||
|
|
||||||
|
def _save_json(path: Path, payload: dict[str, Any] | list[Any]) -> None:
|
||||||
|
path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
path.write_text(json.dumps(payload, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
|
def _term_key(value: str) -> str:
|
||||||
|
return value.strip().casefold()
|
||||||
|
|
||||||
|
|
||||||
|
def _utc_now() -> str:
|
||||||
|
return datetime.now(tz=UTC).isoformat().replace("+00:00", "Z")
|
||||||
|
|
||||||
|
|
||||||
|
def _load_suggestions(path: Path) -> dict[str, Any]:
|
||||||
|
payload = _load_json(path, {})
|
||||||
|
if not isinstance(payload, dict):
|
||||||
|
raise RuntimeError("Suggestions file must contain a JSON object.")
|
||||||
|
return payload
|
||||||
|
|
||||||
|
|
||||||
|
def _load_aliases(path: Path) -> dict[str, str]:
|
||||||
|
payload = _load_json(path, {})
|
||||||
|
if not isinstance(payload, dict):
|
||||||
|
raise RuntimeError("Aliases file must contain a JSON object.")
|
||||||
|
result: dict[str, str] = {}
|
||||||
|
for key, value in payload.items():
|
||||||
|
if isinstance(key, str) and isinstance(value, str) and key.strip() and value.strip():
|
||||||
|
result[key] = value
|
||||||
|
return result
|
||||||
|
|
||||||
|
|
||||||
|
def _load_stopwords(path: Path) -> list[str]:
|
||||||
|
payload = _load_json(path, [])
|
||||||
|
if not isinstance(payload, list):
|
||||||
|
raise RuntimeError("Stopwords file must contain a JSON array.")
|
||||||
|
return [item for item in payload if isinstance(item, str) and item.strip()]
|
||||||
|
|
||||||
|
|
||||||
|
def _load_context(path: Path) -> dict[str, Any]:
|
||||||
|
payload = _load_json(path, {})
|
||||||
|
if not isinstance(payload, dict):
|
||||||
|
raise RuntimeError("Context file must contain a JSON object.")
|
||||||
|
payload.setdefault("interest_keywords", [])
|
||||||
|
if not isinstance(payload["interest_keywords"], list):
|
||||||
|
raise RuntimeError("filter_context.personal.json interest_keywords must be a JSON array.")
|
||||||
|
payload["interest_keywords"] = [item for item in payload["interest_keywords"] if isinstance(item, str) and item.strip()]
|
||||||
|
return payload
|
||||||
|
|
||||||
|
|
||||||
|
def _load_watchlist(path: Path) -> dict[str, Any]:
|
||||||
|
payload = _load_json(
|
||||||
|
path,
|
||||||
|
{
|
||||||
|
"schema_version": "v1",
|
||||||
|
"updated_at": _utc_now(),
|
||||||
|
"terms": [],
|
||||||
|
},
|
||||||
|
)
|
||||||
|
if not isinstance(payload, dict):
|
||||||
|
raise RuntimeError("Watchlist file must contain a JSON object.")
|
||||||
|
payload.setdefault("schema_version", "v1")
|
||||||
|
payload.setdefault("updated_at", _utc_now())
|
||||||
|
terms = payload.get("terms", [])
|
||||||
|
if not isinstance(terms, list):
|
||||||
|
raise RuntimeError("Watchlist terms must be a JSON array.")
|
||||||
|
payload["terms"] = [item for item in terms if isinstance(item, dict) and isinstance(item.get("term"), str)]
|
||||||
|
return payload
|
||||||
|
|
||||||
|
|
||||||
|
def _load_change_log(path: Path) -> dict[str, Any]:
|
||||||
|
payload = _load_json(
|
||||||
|
path,
|
||||||
|
{
|
||||||
|
"schema_version": "v1",
|
||||||
|
"entries": [],
|
||||||
|
},
|
||||||
|
)
|
||||||
|
if not isinstance(payload, dict):
|
||||||
|
raise RuntimeError("Change log file must contain a JSON object.")
|
||||||
|
payload.setdefault("schema_version", "v1")
|
||||||
|
entries = payload.get("entries", [])
|
||||||
|
if not isinstance(entries, list):
|
||||||
|
raise RuntimeError("Change log entries must be a JSON array.")
|
||||||
|
payload["entries"] = [item for item in entries if isinstance(item, dict)]
|
||||||
|
return payload
|
||||||
|
|
||||||
|
|
||||||
|
def _build_index(entries: list[dict[str, Any]], field: str) -> dict[str, dict[str, Any]]:
|
||||||
|
index: dict[str, dict[str, Any]] = {}
|
||||||
|
for entry in entries:
|
||||||
|
value = entry.get(field)
|
||||||
|
if isinstance(value, str) and value.strip():
|
||||||
|
index[_term_key(value)] = entry
|
||||||
|
return index
|
||||||
|
|
||||||
|
|
||||||
|
def _append_change(
|
||||||
|
change_log: dict[str, Any],
|
||||||
|
*,
|
||||||
|
applied_at: str,
|
||||||
|
action: str,
|
||||||
|
term: str,
|
||||||
|
value: str | None,
|
||||||
|
reason: str,
|
||||||
|
suggestions_path: Path,
|
||||||
|
suggestion_date: str | None,
|
||||||
|
based_on_days: int | None,
|
||||||
|
) -> None:
|
||||||
|
entry: dict[str, Any] = {
|
||||||
|
"applied_at": applied_at,
|
||||||
|
"action": action,
|
||||||
|
"term": term,
|
||||||
|
"reason": reason,
|
||||||
|
"suggestions_path": str(suggestions_path),
|
||||||
|
}
|
||||||
|
if value is not None:
|
||||||
|
entry["value"] = value
|
||||||
|
if suggestion_date is not None:
|
||||||
|
entry["suggestion_date"] = suggestion_date
|
||||||
|
if based_on_days is not None:
|
||||||
|
entry["based_on_days"] = based_on_days
|
||||||
|
change_log["entries"].append(entry)
|
||||||
|
|
||||||
|
|
||||||
|
def _remove_from_watchlist(
|
||||||
|
watchlist: dict[str, Any],
|
||||||
|
*,
|
||||||
|
term: str,
|
||||||
|
) -> bool:
|
||||||
|
folded = _term_key(term)
|
||||||
|
original_len = len(watchlist["terms"])
|
||||||
|
watchlist["terms"] = [item for item in watchlist["terms"] if _term_key(str(item.get("term", ""))) != folded]
|
||||||
|
return len(watchlist["terms"]) != original_len
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
parser = argparse.ArgumentParser(description="Apply accepted term cleanup suggestions into repo config files.")
|
||||||
|
parser.add_argument("--suggestions", type=Path, required=True, help="Suggestion JSON file")
|
||||||
|
parser.add_argument("--aliases", type=Path, default=DEFAULT_ALIASES_PATH, help="Term aliases JSON file")
|
||||||
|
parser.add_argument("--stopwords", type=Path, default=DEFAULT_STOPWORDS_PATH, help="Term stopwords JSON file")
|
||||||
|
parser.add_argument("--context", type=Path, default=DEFAULT_CONTEXT_PATH, help="Personal filter context JSON file")
|
||||||
|
parser.add_argument("--watchlist", type=Path, default=DEFAULT_WATCHLIST_PATH, help="Watchlist JSON file")
|
||||||
|
parser.add_argument("--change-log", type=Path, default=DEFAULT_CHANGE_LOG_PATH, help="Change log JSON file")
|
||||||
|
parser.add_argument("--accept-watch", nargs="*", default=[], help="Accept watch_terms by term name")
|
||||||
|
parser.add_argument(
|
||||||
|
"--accept-interest", nargs="*", default=[], help="Accept interest_keyword_suggestions by term name"
|
||||||
|
)
|
||||||
|
parser.add_argument("--accept-stopword", nargs="*", default=[], help="Accept stopword_suggestions by term name")
|
||||||
|
parser.add_argument("--accept-alias", nargs="*", default=[], help="Accept alias_suggestions by source term")
|
||||||
|
parser.add_argument("--dry-run", action="store_true", help="Preview changes without writing files")
|
||||||
|
args = parser.parse_args()
|
||||||
|
|
||||||
|
suggestions = _load_suggestions(args.suggestions)
|
||||||
|
aliases = _load_aliases(args.aliases)
|
||||||
|
stopwords = _load_stopwords(args.stopwords)
|
||||||
|
context = _load_context(args.context)
|
||||||
|
watchlist = _load_watchlist(args.watchlist)
|
||||||
|
change_log = _load_change_log(args.change_log)
|
||||||
|
|
||||||
|
alias_suggestions = suggestions.get("alias_suggestions", [])
|
||||||
|
stopword_suggestions = suggestions.get("stopword_suggestions", [])
|
||||||
|
interest_suggestions = suggestions.get("interest_keyword_suggestions", [])
|
||||||
|
watch_suggestions = suggestions.get("watch_terms", [])
|
||||||
|
if not all(isinstance(bucket, list) for bucket in [alias_suggestions, stopword_suggestions, interest_suggestions, watch_suggestions]):
|
||||||
|
raise RuntimeError("Suggestion JSON buckets must all be arrays.")
|
||||||
|
|
||||||
|
alias_index = _build_index([item for item in alias_suggestions if isinstance(item, dict)], "from")
|
||||||
|
stopword_index = _build_index([item for item in stopword_suggestions if isinstance(item, dict)], "term")
|
||||||
|
interest_index = _build_index([item for item in interest_suggestions if isinstance(item, dict)], "term")
|
||||||
|
watch_index = _build_index([item for item in watch_suggestions if isinstance(item, dict)], "term")
|
||||||
|
|
||||||
|
watch_map = {_term_key(str(item.get("term", ""))): item for item in watchlist["terms"]}
|
||||||
|
stopword_set = {_term_key(item) for item in stopwords}
|
||||||
|
interest_set = {_term_key(item) for item in context["interest_keywords"]}
|
||||||
|
alias_source_set = {_term_key(key) for key in aliases}
|
||||||
|
|
||||||
|
applied_at = _utc_now()
|
||||||
|
suggestion_date = suggestions.get("date") if isinstance(suggestions.get("date"), str) else None
|
||||||
|
based_on_days = suggestions.get("based_on_days") if isinstance(suggestions.get("based_on_days"), int) else None
|
||||||
|
|
||||||
|
applied: list[dict[str, Any]] = []
|
||||||
|
skipped: list[dict[str, Any]] = []
|
||||||
|
|
||||||
|
def skip(action: str, term: str, reason: str) -> None:
|
||||||
|
skipped.append({"action": action, "term": term, "reason": reason})
|
||||||
|
|
||||||
|
def applied_entry(action: str, term: str, reason: str, value: str | None = None) -> None:
|
||||||
|
item: dict[str, Any] = {"action": action, "term": term, "reason": reason}
|
||||||
|
if value is not None:
|
||||||
|
item["value"] = value
|
||||||
|
applied.append(item)
|
||||||
|
|
||||||
|
for raw_term in args.accept_watch:
|
||||||
|
term = raw_term.strip()
|
||||||
|
suggestion = watch_index.get(_term_key(term))
|
||||||
|
if suggestion is None:
|
||||||
|
skip("add_watch_term", term, "Term not found in watch_terms suggestions.")
|
||||||
|
continue
|
||||||
|
if _term_key(term) in watch_map:
|
||||||
|
skip("add_watch_term", term, "Term already exists in watchlist.")
|
||||||
|
continue
|
||||||
|
reason = str(suggestion.get("reason") or "Accepted from watch_terms suggestion.")
|
||||||
|
watch_item = {
|
||||||
|
"term": str(suggestion.get("term") or term),
|
||||||
|
"added_at": applied_at,
|
||||||
|
"source": str(args.suggestions),
|
||||||
|
"reason": reason,
|
||||||
|
"status": "watching",
|
||||||
|
}
|
||||||
|
watchlist["terms"].append(watch_item)
|
||||||
|
watch_map[_term_key(watch_item["term"])] = watch_item
|
||||||
|
_append_change(
|
||||||
|
change_log,
|
||||||
|
applied_at=applied_at,
|
||||||
|
action="add_watch_term",
|
||||||
|
term=watch_item["term"],
|
||||||
|
value=None,
|
||||||
|
reason=reason,
|
||||||
|
suggestions_path=args.suggestions,
|
||||||
|
suggestion_date=suggestion_date,
|
||||||
|
based_on_days=based_on_days,
|
||||||
|
)
|
||||||
|
applied_entry("add_watch_term", watch_item["term"], reason)
|
||||||
|
|
||||||
|
for raw_term in args.accept_interest:
|
||||||
|
term = raw_term.strip()
|
||||||
|
suggestion = interest_index.get(_term_key(term))
|
||||||
|
if suggestion is None:
|
||||||
|
skip("add_interest_keyword", term, "Term not found in interest_keyword_suggestions.")
|
||||||
|
continue
|
||||||
|
resolved_term = str(suggestion.get("term") or term)
|
||||||
|
if _term_key(resolved_term) in interest_set:
|
||||||
|
skip("add_interest_keyword", resolved_term, "Term already exists in interest_keywords.")
|
||||||
|
continue
|
||||||
|
reason = str(suggestion.get("reason") or "Accepted from interest_keyword_suggestions.")
|
||||||
|
context["interest_keywords"].append(resolved_term)
|
||||||
|
interest_set.add(_term_key(resolved_term))
|
||||||
|
_append_change(
|
||||||
|
change_log,
|
||||||
|
applied_at=applied_at,
|
||||||
|
action="add_interest_keyword",
|
||||||
|
term=resolved_term,
|
||||||
|
value=None,
|
||||||
|
reason=reason,
|
||||||
|
suggestions_path=args.suggestions,
|
||||||
|
suggestion_date=suggestion_date,
|
||||||
|
based_on_days=based_on_days,
|
||||||
|
)
|
||||||
|
if _remove_from_watchlist(watchlist, term=resolved_term):
|
||||||
|
_append_change(
|
||||||
|
change_log,
|
||||||
|
applied_at=applied_at,
|
||||||
|
action="remove_watch_term",
|
||||||
|
term=resolved_term,
|
||||||
|
value="promoted_to_interest_keyword",
|
||||||
|
reason="Removed from watchlist after promotion into interest_keywords.",
|
||||||
|
suggestions_path=args.suggestions,
|
||||||
|
suggestion_date=suggestion_date,
|
||||||
|
based_on_days=based_on_days,
|
||||||
|
)
|
||||||
|
applied_entry("add_interest_keyword", resolved_term, reason)
|
||||||
|
|
||||||
|
for raw_term in args.accept_stopword:
|
||||||
|
term = raw_term.strip()
|
||||||
|
suggestion = stopword_index.get(_term_key(term))
|
||||||
|
if suggestion is None:
|
||||||
|
skip("add_stopword", term, "Term not found in stopword_suggestions.")
|
||||||
|
continue
|
||||||
|
resolved_term = str(suggestion.get("term") or term)
|
||||||
|
if _term_key(resolved_term) in stopword_set:
|
||||||
|
skip("add_stopword", resolved_term, "Term already exists in stopwords.")
|
||||||
|
continue
|
||||||
|
reason = str(suggestion.get("reason") or "Accepted from stopword_suggestions.")
|
||||||
|
stopwords.append(resolved_term)
|
||||||
|
stopword_set.add(_term_key(resolved_term))
|
||||||
|
_append_change(
|
||||||
|
change_log,
|
||||||
|
applied_at=applied_at,
|
||||||
|
action="add_stopword",
|
||||||
|
term=resolved_term,
|
||||||
|
value=None,
|
||||||
|
reason=reason,
|
||||||
|
suggestions_path=args.suggestions,
|
||||||
|
suggestion_date=suggestion_date,
|
||||||
|
based_on_days=based_on_days,
|
||||||
|
)
|
||||||
|
if _remove_from_watchlist(watchlist, term=resolved_term):
|
||||||
|
_append_change(
|
||||||
|
change_log,
|
||||||
|
applied_at=applied_at,
|
||||||
|
action="remove_watch_term",
|
||||||
|
term=resolved_term,
|
||||||
|
value="promoted_to_stopword",
|
||||||
|
reason="Removed from watchlist after being added to stopwords.",
|
||||||
|
suggestions_path=args.suggestions,
|
||||||
|
suggestion_date=suggestion_date,
|
||||||
|
based_on_days=based_on_days,
|
||||||
|
)
|
||||||
|
applied_entry("add_stopword", resolved_term, reason)
|
||||||
|
|
||||||
|
for raw_term in args.accept_alias:
|
||||||
|
term = raw_term.strip()
|
||||||
|
suggestion = alias_index.get(_term_key(term))
|
||||||
|
if suggestion is None:
|
||||||
|
skip("add_alias", term, "Source term not found in alias_suggestions.")
|
||||||
|
continue
|
||||||
|
source_term = str(suggestion.get("from") or term)
|
||||||
|
target_term = str(suggestion.get("to") or "").strip()
|
||||||
|
if not target_term:
|
||||||
|
skip("add_alias", source_term, "Alias suggestion target is empty.")
|
||||||
|
continue
|
||||||
|
existing_target = aliases.get(source_term)
|
||||||
|
if existing_target == target_term:
|
||||||
|
skip("add_alias", source_term, "Alias already exists with the same target.")
|
||||||
|
continue
|
||||||
|
if _term_key(source_term) in alias_source_set and existing_target != target_term:
|
||||||
|
skip("add_alias", source_term, f"Alias source already exists with a different target: {existing_target}")
|
||||||
|
continue
|
||||||
|
reason = str(suggestion.get("reason") or "Accepted from alias_suggestions.")
|
||||||
|
aliases[source_term] = target_term
|
||||||
|
alias_source_set.add(_term_key(source_term))
|
||||||
|
_append_change(
|
||||||
|
change_log,
|
||||||
|
applied_at=applied_at,
|
||||||
|
action="add_alias",
|
||||||
|
term=source_term,
|
||||||
|
value=target_term,
|
||||||
|
reason=reason,
|
||||||
|
suggestions_path=args.suggestions,
|
||||||
|
suggestion_date=suggestion_date,
|
||||||
|
based_on_days=based_on_days,
|
||||||
|
)
|
||||||
|
if _remove_from_watchlist(watchlist, term=source_term):
|
||||||
|
_append_change(
|
||||||
|
change_log,
|
||||||
|
applied_at=applied_at,
|
||||||
|
action="remove_watch_term",
|
||||||
|
term=source_term,
|
||||||
|
value="resolved_as_alias_source",
|
||||||
|
reason="Removed from watchlist after alias mapping was accepted.",
|
||||||
|
suggestions_path=args.suggestions,
|
||||||
|
suggestion_date=suggestion_date,
|
||||||
|
based_on_days=based_on_days,
|
||||||
|
)
|
||||||
|
applied_entry("add_alias", source_term, reason, value=target_term)
|
||||||
|
|
||||||
|
watchlist["terms"] = sorted(watchlist["terms"], key=lambda item: (_term_key(str(item.get("term", ""))), str(item.get("term", ""))))
|
||||||
|
watchlist["updated_at"] = applied_at
|
||||||
|
stopwords = sorted(stopwords, key=lambda item: (_term_key(item), item))
|
||||||
|
context["interest_keywords"] = sorted(context["interest_keywords"], key=lambda item: (_term_key(item), item))
|
||||||
|
|
||||||
|
summary = {
|
||||||
|
"dry_run": args.dry_run,
|
||||||
|
"suggestions": str(args.suggestions),
|
||||||
|
"applied_count": len(applied),
|
||||||
|
"skipped_count": len(skipped),
|
||||||
|
"applied": applied,
|
||||||
|
"skipped": skipped,
|
||||||
|
"output_paths": {
|
||||||
|
"aliases": str(args.aliases),
|
||||||
|
"stopwords": str(args.stopwords),
|
||||||
|
"context": str(args.context),
|
||||||
|
"watchlist": str(args.watchlist),
|
||||||
|
"change_log": str(args.change_log),
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
if not args.dry_run:
|
||||||
|
_save_json(args.aliases, aliases)
|
||||||
|
_save_json(args.stopwords, stopwords)
|
||||||
|
_save_json(args.context, context)
|
||||||
|
_save_json(args.watchlist, watchlist)
|
||||||
|
_save_json(args.change_log, change_log)
|
||||||
|
|
||||||
|
print(json.dumps(summary, ensure_ascii=False, indent=2))
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -0,0 +1,76 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import sys
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
|
||||||
|
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||||
|
SRC_ROOT = REPO_ROOT / "src"
|
||||||
|
|
||||||
|
if str(SRC_ROOT) not in sys.path:
|
||||||
|
sys.path.insert(0, str(SRC_ROOT))
|
||||||
|
|
||||||
|
from summary_mcp.core.keyword_index import persist_keyword_indexes
|
||||||
|
from summary_mcp.models.openclaw_delivery import OpenClawDeliveryPayload
|
||||||
|
|
||||||
|
|
||||||
|
def _load_payload(path: Path) -> OpenClawDeliveryPayload:
|
||||||
|
return OpenClawDeliveryPayload.model_validate_json(path.read_text(encoding="utf-8-sig"))
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
parser = argparse.ArgumentParser(
|
||||||
|
description="Build the daily keyword index and global keyword stats from an OpenClaw delivery payload."
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--input",
|
||||||
|
type=Path,
|
||||||
|
required=True,
|
||||||
|
help="OpenClaw delivery payload JSON file",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--daily-dir",
|
||||||
|
type=Path,
|
||||||
|
default=REPO_ROOT / "data" / "term_index" / "daily",
|
||||||
|
help="Directory for daily keyword index files",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--stats-output",
|
||||||
|
type=Path,
|
||||||
|
default=REPO_ROOT / "data" / "term_index" / "term_stats.json",
|
||||||
|
help="Global keyword stats output path",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--aliases",
|
||||||
|
type=Path,
|
||||||
|
default=REPO_ROOT / "configs" / "term_aliases.json",
|
||||||
|
help="Keyword alias config JSON file",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--stopwords",
|
||||||
|
type=Path,
|
||||||
|
default=REPO_ROOT / "configs" / "term_stopwords.json",
|
||||||
|
help="Keyword stopwords config JSON file",
|
||||||
|
)
|
||||||
|
args = parser.parse_args()
|
||||||
|
|
||||||
|
payload = _load_payload(args.input)
|
||||||
|
result = persist_keyword_indexes(
|
||||||
|
payload.candidates,
|
||||||
|
for_date=payload.date,
|
||||||
|
digest_id=payload.run_id,
|
||||||
|
source="openclaw_delivery_payload",
|
||||||
|
daily_dir=args.daily_dir,
|
||||||
|
stats_path=args.stats_output,
|
||||||
|
aliases_path=args.aliases,
|
||||||
|
stopwords_path=args.stopwords,
|
||||||
|
)
|
||||||
|
|
||||||
|
print(f"Saved daily keyword index to {result['daily_output']}")
|
||||||
|
print(f"Saved keyword stats to {result['stats_output']}")
|
||||||
|
print(f"Indexed {result['candidate_count']} candidates and {result['term_count']} keywords")
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -0,0 +1,85 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
import sys
|
||||||
|
from datetime import date
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
|
||||||
|
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||||
|
SRC_ROOT = REPO_ROOT / "src"
|
||||||
|
OUTPUT_ROOT = REPO_ROOT / "outputs"
|
||||||
|
REFERENCE_OUTPUT_ROOT = OUTPUT_ROOT / "reference"
|
||||||
|
|
||||||
|
if str(SRC_ROOT) not in sys.path:
|
||||||
|
sys.path.insert(0, str(SRC_ROOT))
|
||||||
|
|
||||||
|
from summary_mcp.models import OpenClawCandidateInput, build_openclaw_delivery_payload
|
||||||
|
|
||||||
|
|
||||||
|
def _load_json(path: Path) -> dict[str, Any]:
|
||||||
|
return json.loads(path.read_text(encoding="utf-8-sig"))
|
||||||
|
|
||||||
|
|
||||||
|
def _save_json(path: Path, payload: dict[str, Any]) -> None:
|
||||||
|
path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
path.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
|
def _collect_input_paths(inputs: list[Path], input_dir: Path | None, pattern: str) -> list[Path]:
|
||||||
|
resolved = list(inputs)
|
||||||
|
if input_dir is not None:
|
||||||
|
resolved.extend(sorted(input_dir.glob(pattern)))
|
||||||
|
seen: set[Path] = set()
|
||||||
|
unique_paths: list[Path] = []
|
||||||
|
for path in resolved:
|
||||||
|
normalized = path.resolve()
|
||||||
|
if normalized in seen:
|
||||||
|
continue
|
||||||
|
seen.add(normalized)
|
||||||
|
unique_paths.append(path)
|
||||||
|
return unique_paths
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
parser = argparse.ArgumentParser(description="Build an OpenClaw batch delivery payload from slim candidate JSON files.")
|
||||||
|
parser.add_argument("--inputs", nargs="*", type=Path, default=[], help="One or more OpenClaw candidate input JSON files")
|
||||||
|
parser.add_argument("--input-dir", type=Path, default=None, help="Directory containing OpenClaw candidate input JSON files")
|
||||||
|
parser.add_argument(
|
||||||
|
"--pattern",
|
||||||
|
type=str,
|
||||||
|
default="*.openclaw-candidate-input.json",
|
||||||
|
help="Glob pattern used together with --input-dir",
|
||||||
|
)
|
||||||
|
parser.add_argument("--run-id", type=str, default=None, help="Optional delivery run id")
|
||||||
|
parser.add_argument("--date", type=str, default=None, help="Optional delivery date in YYYY-MM-DD format")
|
||||||
|
parser.add_argument("--sort-by-rank", action="store_true", help="Sort candidates by digest_rank descending")
|
||||||
|
parser.add_argument(
|
||||||
|
"--output",
|
||||||
|
type=Path,
|
||||||
|
default=REFERENCE_OUTPUT_ROOT / "candidates" / "openclaw-delivery-payload.json",
|
||||||
|
help="Where to save the OpenClaw delivery payload",
|
||||||
|
)
|
||||||
|
args = parser.parse_args()
|
||||||
|
|
||||||
|
input_paths = _collect_input_paths(args.inputs, args.input_dir, args.pattern)
|
||||||
|
if not input_paths:
|
||||||
|
raise SystemExit("No candidate input files found. Pass --inputs or --input-dir.")
|
||||||
|
|
||||||
|
candidates = [OpenClawCandidateInput.model_validate(_load_json(path)) for path in input_paths]
|
||||||
|
if args.sort_by_rank:
|
||||||
|
candidates.sort(key=lambda candidate: candidate.digest_rank, reverse=True)
|
||||||
|
|
||||||
|
payload = build_openclaw_delivery_payload(
|
||||||
|
candidates,
|
||||||
|
run_id=args.run_id,
|
||||||
|
for_date=date.fromisoformat(args.date) if args.date else None,
|
||||||
|
)
|
||||||
|
_save_json(args.output, payload.model_dump(mode="json"))
|
||||||
|
print(f"Saved OpenClaw delivery payload with {len(candidates)} candidates to {args.output}")
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -0,0 +1,114 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import sys
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
|
||||||
|
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||||
|
SRC_ROOT = REPO_ROOT / "src"
|
||||||
|
OUTPUT_ROOT = REPO_ROOT / "outputs"
|
||||||
|
FRESHRSS_OUTPUT_ROOT = OUTPUT_ROOT / "freshrss"
|
||||||
|
|
||||||
|
if str(SRC_ROOT) not in sys.path:
|
||||||
|
sys.path.insert(0, str(SRC_ROOT))
|
||||||
|
|
||||||
|
from summary_mcp.integrations.freshrss import READ_TAG, FreshRSSClient, map_entry_to_item
|
||||||
|
|
||||||
|
|
||||||
|
def _save_json(path: Path, payload: dict[str, Any] | list[Any]) -> None:
|
||||||
|
path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
path.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
|
def _load_required(name: str, value: str | None) -> str:
|
||||||
|
resolved = value or os.environ.get(name)
|
||||||
|
if not resolved:
|
||||||
|
cli_name = name.lower().replace("_", "-")
|
||||||
|
raise RuntimeError(f"Missing required value: pass --{cli_name} or set {name}.")
|
||||||
|
return resolved
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
parser = argparse.ArgumentParser(description="Pull FreshRSS entries and map them into normalized item objects.")
|
||||||
|
parser.add_argument("--api-base-url", type=str, default=None, help="FreshRSS greader API base URL")
|
||||||
|
parser.add_argument("--username", type=str, default=None, help="FreshRSS API username")
|
||||||
|
parser.add_argument("--api-password", type=str, default=None, help="FreshRSS API password")
|
||||||
|
parser.add_argument(
|
||||||
|
"--stream-id",
|
||||||
|
type=str,
|
||||||
|
default="user/-/state/com.google/reading-list",
|
||||||
|
help="Google Reader API stream id",
|
||||||
|
)
|
||||||
|
parser.add_argument("--limit", type=int, default=10, help="Maximum number of entries to request")
|
||||||
|
parser.add_argument("--continuation", type=str, default=None, help="Continuation token for paging")
|
||||||
|
parser.add_argument(
|
||||||
|
"--include-read",
|
||||||
|
action="store_true",
|
||||||
|
help="Do not exclude items already tagged as read.",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--mark-read",
|
||||||
|
action="store_true",
|
||||||
|
help="Mark fetched entries as read after this script finishes successfully.",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--raw-output",
|
||||||
|
type=Path,
|
||||||
|
default=FRESHRSS_OUTPUT_ROOT / "raw" / "freshrss.raw.json",
|
||||||
|
help="Where to save the raw FreshRSS response",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--items-output",
|
||||||
|
type=Path,
|
||||||
|
default=FRESHRSS_OUTPUT_ROOT / "items" / "freshrss.items.json",
|
||||||
|
help="Where to save the mapped item list",
|
||||||
|
)
|
||||||
|
parser.add_argument("--timeout", type=float, default=20.0, help="Request timeout in seconds")
|
||||||
|
args = parser.parse_args()
|
||||||
|
|
||||||
|
api_base_url = _load_required("FRESHRSS_API_BASE_URL", args.api_base_url)
|
||||||
|
username = _load_required("FRESHRSS_USERNAME", args.username)
|
||||||
|
api_password = _load_required("FRESHRSS_API_PASSWORD", args.api_password)
|
||||||
|
|
||||||
|
client = FreshRSSClient(
|
||||||
|
api_base_url=api_base_url,
|
||||||
|
username=username,
|
||||||
|
api_password=api_password,
|
||||||
|
timeout_seconds=args.timeout,
|
||||||
|
)
|
||||||
|
auth_token = client.client_login()
|
||||||
|
payload = client.fetch_stream_contents(
|
||||||
|
auth_token=auth_token,
|
||||||
|
stream_id=args.stream_id,
|
||||||
|
limit=args.limit,
|
||||||
|
continuation=args.continuation,
|
||||||
|
exclude_targets=[] if args.include_read else [READ_TAG],
|
||||||
|
)
|
||||||
|
|
||||||
|
entries = payload.get("items")
|
||||||
|
if not isinstance(entries, list):
|
||||||
|
raise RuntimeError("FreshRSS stream response does not contain an items array.")
|
||||||
|
|
||||||
|
mapped_items = [map_entry_to_item(entry).model_dump(mode="json") for entry in entries]
|
||||||
|
|
||||||
|
_save_json(args.raw_output, payload)
|
||||||
|
_save_json(args.items_output, mapped_items)
|
||||||
|
|
||||||
|
marked_count = 0
|
||||||
|
if args.mark_read:
|
||||||
|
item_ids = [entry.get("external_id") for entry in mapped_items if isinstance(entry.get("external_id"), str)]
|
||||||
|
if item_ids:
|
||||||
|
client.mark_items_as_read(auth_token=auth_token, item_ids=item_ids)
|
||||||
|
marked_count = len(item_ids)
|
||||||
|
|
||||||
|
print(f"Saved {len(mapped_items)} mapped items to {args.items_output}")
|
||||||
|
if args.mark_read:
|
||||||
|
print(f"Marked {marked_count} FreshRSS entries as read")
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -0,0 +1,124 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
import sys
|
||||||
|
from datetime import UTC, datetime
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
|
||||||
|
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||||
|
SRC_ROOT = REPO_ROOT / "src"
|
||||||
|
OUTPUT_ROOT = REPO_ROOT / "outputs"
|
||||||
|
REFERENCE_OUTPUT_ROOT = OUTPUT_ROOT / "reference"
|
||||||
|
|
||||||
|
if str(SRC_ROOT) not in sys.path:
|
||||||
|
sys.path.insert(0, str(SRC_ROOT))
|
||||||
|
|
||||||
|
from summary_mcp.models.article_candidate import (
|
||||||
|
CandidateMetadata,
|
||||||
|
CandidateSourceRefs,
|
||||||
|
build_article_candidate_record,
|
||||||
|
build_openclaw_candidate_input,
|
||||||
|
)
|
||||||
|
from summary_mcp.models.document import ExtractedArticle
|
||||||
|
from summary_mcp.models.filtering import FilterDecisionResult
|
||||||
|
from summary_mcp.models.item import Item
|
||||||
|
from summary_mcp.models.llm_result import LlmSummaryResult
|
||||||
|
|
||||||
|
|
||||||
|
def _load_json(path: Path) -> dict[str, Any]:
|
||||||
|
return json.loads(path.read_text(encoding="utf-8-sig"))
|
||||||
|
|
||||||
|
|
||||||
|
def _save_json(path: Path, payload: dict[str, Any]) -> None:
|
||||||
|
path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
path.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
parser = argparse.ArgumentParser(
|
||||||
|
description="Build an internal ArticleCandidateRecord and a slim OpenClawCandidateInput from pipeline outputs."
|
||||||
|
)
|
||||||
|
parser.add_argument("--summary", type=Path, required=True, help="Structured LLM summary JSON file")
|
||||||
|
parser.add_argument("--extracted", type=Path, required=True, help="Extracted article JSON file")
|
||||||
|
parser.add_argument("--filter", type=Path, required=True, help="Filter decision JSON file")
|
||||||
|
parser.add_argument("--item", type=Path, default=None, help="Optional normalized item JSON file")
|
||||||
|
parser.add_argument(
|
||||||
|
"--section-hint",
|
||||||
|
default=None,
|
||||||
|
choices=[
|
||||||
|
"top_news",
|
||||||
|
"tools_and_workflows",
|
||||||
|
"risk_and_security",
|
||||||
|
"open_source",
|
||||||
|
"insights",
|
||||||
|
"deep_dive",
|
||||||
|
],
|
||||||
|
help="Optional digest section hint for downstream aggregation",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--rank",
|
||||||
|
type=int,
|
||||||
|
default=None,
|
||||||
|
help="Optional digest rank override. Defaults to filter priority when omitted.",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--rendered-markdown",
|
||||||
|
type=Path,
|
||||||
|
default=None,
|
||||||
|
help="Optional markdown file to embed as rendered_markdown",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--output",
|
||||||
|
dest="record_output",
|
||||||
|
type=Path,
|
||||||
|
default=REFERENCE_OUTPUT_ROOT / "candidates" / "article-candidate-record.json",
|
||||||
|
help="Where to save the internal ArticleCandidateRecord payload",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--openclaw-output",
|
||||||
|
type=Path,
|
||||||
|
default=REFERENCE_OUTPUT_ROOT / "candidates" / "openclaw-candidate-input.json",
|
||||||
|
help="Where to save the slim OpenClawCandidateInput payload",
|
||||||
|
)
|
||||||
|
args = parser.parse_args()
|
||||||
|
|
||||||
|
summary = LlmSummaryResult.model_validate(_load_json(args.summary))
|
||||||
|
extracted_payload = _load_json(args.extracted)
|
||||||
|
article = ExtractedArticle.model_validate(extracted_payload.get("article", extracted_payload))
|
||||||
|
decision = FilterDecisionResult.model_validate(_load_json(args.filter))
|
||||||
|
item = Item.model_validate(_load_json(args.item)) if args.item else None
|
||||||
|
rendered_markdown = args.rendered_markdown.read_text(encoding="utf-8-sig") if args.rendered_markdown else None
|
||||||
|
|
||||||
|
record = build_article_candidate_record(
|
||||||
|
summary=summary,
|
||||||
|
article=article,
|
||||||
|
filter_result=decision,
|
||||||
|
item=item,
|
||||||
|
digest_section_hint=args.section_hint,
|
||||||
|
digest_rank=args.rank,
|
||||||
|
rendered_markdown=rendered_markdown,
|
||||||
|
source_refs=CandidateSourceRefs(
|
||||||
|
item_path=str(args.item) if args.item else None,
|
||||||
|
extracted_path=str(args.extracted),
|
||||||
|
summary_path=str(args.summary),
|
||||||
|
filter_path=str(args.filter),
|
||||||
|
),
|
||||||
|
metadata=CandidateMetadata(
|
||||||
|
generated_at=datetime.now(tz=UTC),
|
||||||
|
producer="run_article_candidate.py",
|
||||||
|
run_id=datetime.now(tz=UTC).strftime("candidate-%Y%m%d-%H%M%S"),
|
||||||
|
),
|
||||||
|
)
|
||||||
|
openclaw_input = build_openclaw_candidate_input(record)
|
||||||
|
|
||||||
|
_save_json(args.record_output, record.model_dump(mode="json"))
|
||||||
|
_save_json(args.openclaw_output, openclaw_input.model_dump(mode="json"))
|
||||||
|
print(f"Saved article candidate record to {args.record_output}")
|
||||||
|
print(f"Saved OpenClaw candidate input to {args.openclaw_output}")
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -0,0 +1,78 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
import sys
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
|
||||||
|
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||||
|
SRC_ROOT = REPO_ROOT / "src"
|
||||||
|
OUTPUT_ROOT = REPO_ROOT / "outputs"
|
||||||
|
REFERENCE_OUTPUT_ROOT = OUTPUT_ROOT / "reference"
|
||||||
|
|
||||||
|
if str(SRC_ROOT) not in sys.path:
|
||||||
|
sys.path.insert(0, str(SRC_ROOT))
|
||||||
|
|
||||||
|
from summary_mcp.filters.engine import load_filter_rules, evaluate_filter_rules
|
||||||
|
from summary_mcp.models.document import ExtractedArticle
|
||||||
|
from summary_mcp.models.filtering import FilterContext, FilterInput
|
||||||
|
from summary_mcp.models.item import Item
|
||||||
|
from summary_mcp.models.llm_result import LlmSummaryResult
|
||||||
|
|
||||||
|
|
||||||
|
def _load_json(path: Path) -> dict[str, Any]:
|
||||||
|
return json.loads(path.read_text(encoding="utf-8-sig"))
|
||||||
|
|
||||||
|
|
||||||
|
def _save_json(path: Path, payload: dict[str, Any]) -> None:
|
||||||
|
path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
path.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
parser = argparse.ArgumentParser(description="Run deterministic filter rules against a structured summary result.")
|
||||||
|
parser.add_argument("--summary", type=Path, required=True, help="Structured LLM summary JSON file")
|
||||||
|
parser.add_argument("--extracted", type=Path, default=None, help="Extracted article JSON file")
|
||||||
|
parser.add_argument("--item", type=Path, default=None, help="Normalized item JSON file")
|
||||||
|
parser.add_argument("--context", type=Path, default=None, help="Optional filter context JSON file")
|
||||||
|
parser.add_argument(
|
||||||
|
"--rules",
|
||||||
|
type=Path,
|
||||||
|
default=REPO_ROOT / "configs" / "filter_rules.json",
|
||||||
|
help="Filter rule config JSON file",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--output",
|
||||||
|
type=Path,
|
||||||
|
default=REFERENCE_OUTPUT_ROOT / "filter" / "filter-decision.json",
|
||||||
|
help="Where to save the filter decision",
|
||||||
|
)
|
||||||
|
args = parser.parse_args()
|
||||||
|
|
||||||
|
summary = LlmSummaryResult.model_validate(_load_json(args.summary))
|
||||||
|
|
||||||
|
article = None
|
||||||
|
if args.extracted is not None:
|
||||||
|
extracted_payload = _load_json(args.extracted)
|
||||||
|
article_payload = extracted_payload.get("article", extracted_payload)
|
||||||
|
article = ExtractedArticle.model_validate(article_payload)
|
||||||
|
|
||||||
|
item = None
|
||||||
|
if args.item is not None:
|
||||||
|
item = Item.model_validate(_load_json(args.item))
|
||||||
|
|
||||||
|
context = FilterContext.model_validate(_load_json(args.context)) if args.context else FilterContext()
|
||||||
|
rules = load_filter_rules(args.rules)
|
||||||
|
decision = evaluate_filter_rules(
|
||||||
|
FilterInput(item=item, article=article, summary=summary, context=context),
|
||||||
|
rules,
|
||||||
|
)
|
||||||
|
|
||||||
|
_save_json(args.output, decision.model_dump(mode="json"))
|
||||||
|
print(f"Saved filter decision to {args.output}")
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -0,0 +1,148 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import sys
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
|
||||||
|
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||||
|
SRC_ROOT = REPO_ROOT / "src"
|
||||||
|
OUTPUT_ROOT = REPO_ROOT / "outputs"
|
||||||
|
FRESHRSS_OUTPUT_ROOT = OUTPUT_ROOT / "freshrss"
|
||||||
|
|
||||||
|
if str(SRC_ROOT) not in sys.path:
|
||||||
|
sys.path.insert(0, str(SRC_ROOT))
|
||||||
|
|
||||||
|
from summary_mcp.core.pipeline import extract_content
|
||||||
|
from summary_mcp.integrations.freshrss import READ_TAG, FreshRSSClient, map_entry_to_item
|
||||||
|
from summary_mcp.models.summary_io import ExtractionInput
|
||||||
|
|
||||||
|
|
||||||
|
def _save_json(path: Path, payload: dict[str, Any] | list[Any]) -> None:
|
||||||
|
path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
path.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
|
def _load_required(name: str, value: str | None) -> str:
|
||||||
|
resolved = value or os.environ.get(name)
|
||||||
|
if not resolved:
|
||||||
|
cli_name = name.lower().replace("_", "-")
|
||||||
|
raise RuntimeError(f"Missing required value: pass --{cli_name} or set {name}.")
|
||||||
|
return resolved
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
parser = argparse.ArgumentParser(description="Pull FreshRSS entries and run content extraction for each mapped item.")
|
||||||
|
parser.add_argument("--api-base-url", type=str, default=None, help="FreshRSS greader API base URL")
|
||||||
|
parser.add_argument("--username", type=str, default=None, help="FreshRSS API username")
|
||||||
|
parser.add_argument("--api-password", type=str, default=None, help="FreshRSS API password")
|
||||||
|
parser.add_argument(
|
||||||
|
"--stream-id",
|
||||||
|
type=str,
|
||||||
|
default="user/-/state/com.google/reading-list",
|
||||||
|
help="Google Reader API stream id",
|
||||||
|
)
|
||||||
|
parser.add_argument("--limit", type=int, default=5, help="Maximum number of entries to request")
|
||||||
|
parser.add_argument("--continuation", type=str, default=None, help="Continuation token for paging")
|
||||||
|
parser.add_argument(
|
||||||
|
"--include-read",
|
||||||
|
action="store_true",
|
||||||
|
help="Do not exclude items already tagged as read.",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--mark-read",
|
||||||
|
action="store_true",
|
||||||
|
help="Mark entries as read after successful extraction for each item.",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--raw-output",
|
||||||
|
type=Path,
|
||||||
|
default=FRESHRSS_OUTPUT_ROOT / "raw" / "freshrss.raw.json",
|
||||||
|
help="Where to save the raw FreshRSS response",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--items-output",
|
||||||
|
type=Path,
|
||||||
|
default=FRESHRSS_OUTPUT_ROOT / "items" / "freshrss.items.json",
|
||||||
|
help="Where to save the mapped item list",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--extracted-output",
|
||||||
|
type=Path,
|
||||||
|
default=FRESHRSS_OUTPUT_ROOT / "extracted" / "freshrss.extracted.json",
|
||||||
|
help="Where to save the extraction results",
|
||||||
|
)
|
||||||
|
parser.add_argument("--timeout", type=float, default=20.0, help="Request timeout in seconds")
|
||||||
|
args = parser.parse_args()
|
||||||
|
|
||||||
|
api_base_url = _load_required("FRESHRSS_API_BASE_URL", args.api_base_url)
|
||||||
|
username = _load_required("FRESHRSS_USERNAME", args.username)
|
||||||
|
api_password = _load_required("FRESHRSS_API_PASSWORD", args.api_password)
|
||||||
|
|
||||||
|
client = FreshRSSClient(
|
||||||
|
api_base_url=api_base_url,
|
||||||
|
username=username,
|
||||||
|
api_password=api_password,
|
||||||
|
timeout_seconds=args.timeout,
|
||||||
|
)
|
||||||
|
auth_token = client.client_login()
|
||||||
|
payload = client.fetch_stream_contents(
|
||||||
|
auth_token=auth_token,
|
||||||
|
stream_id=args.stream_id,
|
||||||
|
limit=args.limit,
|
||||||
|
continuation=args.continuation,
|
||||||
|
exclude_targets=[] if args.include_read else [READ_TAG],
|
||||||
|
)
|
||||||
|
|
||||||
|
entries = payload.get("items")
|
||||||
|
if not isinstance(entries, list):
|
||||||
|
raise RuntimeError("FreshRSS stream response does not contain an items array.")
|
||||||
|
|
||||||
|
mapped_items = [map_entry_to_item(entry) for entry in entries]
|
||||||
|
extraction_results: list[dict[str, Any]] = []
|
||||||
|
success_count = 0
|
||||||
|
mark_read_ids: list[str] = []
|
||||||
|
|
||||||
|
for item in mapped_items:
|
||||||
|
extraction = extract_content(ExtractionInput(item=item))
|
||||||
|
extraction_results.append(
|
||||||
|
{
|
||||||
|
"item": item.model_dump(mode="json"),
|
||||||
|
"extraction": extraction.model_dump(mode="json"),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
if extraction.success:
|
||||||
|
success_count += 1
|
||||||
|
if item.external_id:
|
||||||
|
mark_read_ids.append(item.external_id)
|
||||||
|
|
||||||
|
_save_json(args.raw_output, payload)
|
||||||
|
_save_json(args.items_output, [item.model_dump(mode="json") for item in mapped_items])
|
||||||
|
_save_json(
|
||||||
|
args.extracted_output,
|
||||||
|
{
|
||||||
|
"stream_id": args.stream_id,
|
||||||
|
"requested_limit": args.limit,
|
||||||
|
"entry_count": len(entries),
|
||||||
|
"extracted_success_count": success_count,
|
||||||
|
"results": extraction_results,
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
marked_count = 0
|
||||||
|
if args.mark_read and mark_read_ids:
|
||||||
|
client.mark_items_as_read(auth_token=auth_token, item_ids=mark_read_ids)
|
||||||
|
marked_count = len(mark_read_ids)
|
||||||
|
|
||||||
|
print(
|
||||||
|
f"Saved {len(mapped_items)} mapped items and {success_count} successful extractions to {args.extracted_output}"
|
||||||
|
)
|
||||||
|
if args.mark_read:
|
||||||
|
print(f"Marked {marked_count} FreshRSS entries as read")
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -0,0 +1,100 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import sys
|
||||||
|
from datetime import date
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
|
||||||
|
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||||
|
SRC_ROOT = REPO_ROOT / "src"
|
||||||
|
|
||||||
|
if str(SRC_ROOT) not in sys.path:
|
||||||
|
sys.path.insert(0, str(SRC_ROOT))
|
||||||
|
|
||||||
|
from summary_mcp.workflows import run_freshrss_pipeline
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
parser = argparse.ArgumentParser(
|
||||||
|
description="Run the full FreshRSS -> extract -> LLM -> filter -> OpenClaw payload pipeline."
|
||||||
|
)
|
||||||
|
parser.add_argument("--api-base-url", type=str, default=None, help="FreshRSS greader API base URL")
|
||||||
|
parser.add_argument("--username", type=str, default=None, help="FreshRSS API username")
|
||||||
|
parser.add_argument("--api-password", type=str, default=None, help="FreshRSS API password")
|
||||||
|
parser.add_argument(
|
||||||
|
"--stream-id",
|
||||||
|
type=str,
|
||||||
|
default="user/-/state/com.google/reading-list",
|
||||||
|
help="Google Reader API stream id",
|
||||||
|
)
|
||||||
|
parser.add_argument("--limit", type=int, default=5, help="Maximum number of entries to request")
|
||||||
|
parser.add_argument("--continuation", type=str, default=None, help="Continuation token for paging")
|
||||||
|
parser.add_argument(
|
||||||
|
"--include-read",
|
||||||
|
action="store_true",
|
||||||
|
help="Do not exclude items already tagged as read.",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--mark-read",
|
||||||
|
action="store_true",
|
||||||
|
help="Mark only successfully delivered items as read after the final payload is written.",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--debug-artifacts",
|
||||||
|
action="store_true",
|
||||||
|
help="Persist per-item intermediate files for debugging.",
|
||||||
|
)
|
||||||
|
parser.add_argument("--prompt", type=Path, default=None, help="LLM prompt template file")
|
||||||
|
parser.add_argument("--rules", type=Path, default=None, help="Filter rule config JSON file")
|
||||||
|
parser.add_argument("--context", type=Path, default=None, help="Optional filter context JSON file")
|
||||||
|
parser.add_argument("--max-retries", type=int, default=2, help="Number of repair retries after the initial attempt")
|
||||||
|
parser.add_argument("--timeout", type=float, default=60.0, help="Request timeout in seconds")
|
||||||
|
parser.add_argument("--llm-api-key", type=str, default=None, help="LLM API key")
|
||||||
|
parser.add_argument("--llm-model", type=str, default=None, help="LLM model name")
|
||||||
|
parser.add_argument("--llm-api-url", type=str, default=None, help="LLM chat completions API URL or base URL")
|
||||||
|
parser.add_argument("--run-id", type=str, default=None, help="Optional pipeline run id")
|
||||||
|
parser.add_argument("--date", type=str, default=None, help="Optional delivery date in YYYY-MM-DD format")
|
||||||
|
parser.add_argument(
|
||||||
|
"--output-dir",
|
||||||
|
type=Path,
|
||||||
|
default=None,
|
||||||
|
help="Run output directory. Defaults to outputs/freshrss/rerun/<timestamp>",
|
||||||
|
)
|
||||||
|
args = parser.parse_args()
|
||||||
|
|
||||||
|
result = run_freshrss_pipeline(
|
||||||
|
api_base_url=args.api_base_url,
|
||||||
|
username=args.username,
|
||||||
|
api_password=args.api_password,
|
||||||
|
stream_id=args.stream_id,
|
||||||
|
limit=args.limit,
|
||||||
|
continuation=args.continuation,
|
||||||
|
include_read=args.include_read,
|
||||||
|
mark_read=args.mark_read,
|
||||||
|
debug_artifacts=args.debug_artifacts,
|
||||||
|
prompt=args.prompt,
|
||||||
|
rules=args.rules,
|
||||||
|
context_path=args.context,
|
||||||
|
max_retries=args.max_retries,
|
||||||
|
timeout_seconds=args.timeout,
|
||||||
|
llm_api_key=args.llm_api_key,
|
||||||
|
llm_model=args.llm_model,
|
||||||
|
llm_api_url=args.llm_api_url,
|
||||||
|
run_id=args.run_id,
|
||||||
|
delivery_date=date.fromisoformat(args.date) if args.date else None,
|
||||||
|
output_dir=args.output_dir,
|
||||||
|
)
|
||||||
|
|
||||||
|
print(f"Saved FreshRSS pipeline run to {result['output_dir']}")
|
||||||
|
print(f"Pulled {result['pulled_count']} items, delivered {result['delivered_count']} candidates")
|
||||||
|
print(
|
||||||
|
"Saved keyword index to "
|
||||||
|
f"{result['keyword_index']['daily_output']} and {result['keyword_index']['stats_output']}"
|
||||||
|
)
|
||||||
|
if args.mark_read:
|
||||||
|
print(f"Marked {result['marked_read_count']} FreshRSS entries as read")
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -0,0 +1,61 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
import sys
|
||||||
|
from datetime import UTC, datetime
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
|
||||||
|
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||||
|
SRC_ROOT = REPO_ROOT / "src"
|
||||||
|
|
||||||
|
if str(SRC_ROOT) not in sys.path:
|
||||||
|
sys.path.insert(0, str(SRC_ROOT))
|
||||||
|
|
||||||
|
from summary_mcp.models.document import ExtractedArticle
|
||||||
|
from summary_mcp.models.filtering import FilterDecisionResult
|
||||||
|
from summary_mcp.models.item import Item
|
||||||
|
from summary_mcp.models.llm_result import LlmSummaryResult
|
||||||
|
from summary_mcp.models.sink import SinkInput, SinkMetadata
|
||||||
|
from summary_mcp.sinks.markdown import MarkdownSink
|
||||||
|
|
||||||
|
|
||||||
|
def _load_json(path: Path) -> dict[str, Any]:
|
||||||
|
return json.loads(path.read_text(encoding="utf-8-sig"))
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
parser = argparse.ArgumentParser(description="Write a filtered result into the Markdown sink.")
|
||||||
|
parser.add_argument("--summary", type=Path, required=True, help="Structured LLM summary JSON file")
|
||||||
|
parser.add_argument("--extracted", type=Path, required=True, help="Extracted article JSON file")
|
||||||
|
parser.add_argument("--filter", type=Path, required=True, help="Filter decision JSON file")
|
||||||
|
parser.add_argument("--item", type=Path, default=None, help="Optional normalized item JSON file")
|
||||||
|
parser.add_argument(
|
||||||
|
"--base-dir",
|
||||||
|
type=Path,
|
||||||
|
default=REPO_ROOT / "knowledge-base",
|
||||||
|
help="Markdown sink base directory",
|
||||||
|
)
|
||||||
|
args = parser.parse_args()
|
||||||
|
|
||||||
|
summary = LlmSummaryResult.model_validate(_load_json(args.summary))
|
||||||
|
extracted_payload = _load_json(args.extracted)
|
||||||
|
article = ExtractedArticle.model_validate(extracted_payload.get("article", extracted_payload))
|
||||||
|
decision = FilterDecisionResult.model_validate(_load_json(args.filter))
|
||||||
|
item = Item.model_validate(_load_json(args.item)) if args.item else None
|
||||||
|
|
||||||
|
sink_input = SinkInput(
|
||||||
|
item=item,
|
||||||
|
article=article,
|
||||||
|
summary=summary,
|
||||||
|
filter_decision=decision,
|
||||||
|
metadata=SinkMetadata(generated_at=datetime.now(tz=UTC), source="run_markdown_sink.py"),
|
||||||
|
)
|
||||||
|
result = MarkdownSink(args.base_dir).write(sink_input)
|
||||||
|
print(json.dumps(result.model_dump(mode="json"), ensure_ascii=False, indent=2))
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
+2
-193
@@ -1,14 +1,8 @@
|
|||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
import argparse
|
import argparse
|
||||||
import json
|
|
||||||
import os
|
|
||||||
import re
|
|
||||||
import sys
|
import sys
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from typing import Any
|
|
||||||
|
|
||||||
import httpx
|
|
||||||
|
|
||||||
|
|
||||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||||
@@ -17,192 +11,7 @@ SRC_ROOT = REPO_ROOT / "src"
|
|||||||
if str(SRC_ROOT) not in sys.path:
|
if str(SRC_ROOT) not in sys.path:
|
||||||
sys.path.insert(0, str(SRC_ROOT))
|
sys.path.insert(0, str(SRC_ROOT))
|
||||||
|
|
||||||
from summary_mcp.validators.llm_result import validate_llm_result
|
from summary_mcp.core.summary_loop import run_loop
|
||||||
|
|
||||||
|
|
||||||
JSON_BLOCK_RE = re.compile(r"```(?:json)?\s*(\{.*\})\s*```", re.DOTALL)
|
|
||||||
|
|
||||||
|
|
||||||
def load_text(path: Path) -> str:
|
|
||||||
return path.read_text(encoding="utf-8")
|
|
||||||
|
|
||||||
|
|
||||||
def load_json(path: Path) -> dict[str, Any]:
|
|
||||||
return json.loads(load_text(path))
|
|
||||||
|
|
||||||
|
|
||||||
def save_json(path: Path, payload: dict[str, Any]) -> None:
|
|
||||||
path.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
|
|
||||||
|
|
||||||
|
|
||||||
def build_summary_input(extracted: dict[str, Any]) -> dict[str, Any]:
|
|
||||||
article = extracted.get('article') or {}
|
|
||||||
return {
|
|
||||||
'article': {
|
|
||||||
'title': article.get('title'),
|
|
||||||
'url': article.get('url'),
|
|
||||||
'plain_text': article.get('plain_text'),
|
|
||||||
'quality_flags': article.get('quality_flags'),
|
|
||||||
},
|
|
||||||
'warnings': extracted.get('warnings', []),
|
|
||||||
}
|
|
||||||
|
|
||||||
|
|
||||||
def build_initial_prompt(prompt_template: str, extracted: dict[str, Any]) -> str:
|
|
||||||
summary_input = build_summary_input(extracted)
|
|
||||||
return (
|
|
||||||
f"{prompt_template}\n\n"
|
|
||||||
"Below is the structured extracted article input. Generate the final summary JSON from it.\n\n"
|
|
||||||
f"{json.dumps(summary_input, ensure_ascii=False, indent=2)}"
|
|
||||||
)
|
|
||||||
|
|
||||||
|
|
||||||
def build_repair_prompt(
|
|
||||||
errors: list[str],
|
|
||||||
extracted: dict[str, Any],
|
|
||||||
result_json: dict[str, Any],
|
|
||||||
) -> str:
|
|
||||||
summary_input = build_summary_input(extracted)
|
|
||||||
return (
|
|
||||||
"Please repair the following invalid summary JSON.\n\n"
|
|
||||||
"Requirements:\n"
|
|
||||||
"- Output valid JSON only\n"
|
|
||||||
"- Keep fields that are already correct\n"
|
|
||||||
"- Fix only the validator-reported errors\n"
|
|
||||||
"- Do not add explanations\n\n"
|
|
||||||
f"validator errors:\n{json.dumps(errors, ensure_ascii=False, indent=2)}\n\n"
|
|
||||||
f"Extracted article input:\n{json.dumps(summary_input, ensure_ascii=False, indent=2)}\n\n"
|
|
||||||
f"Current summary JSON:\n{json.dumps(result_json, ensure_ascii=False, indent=2)}\n"
|
|
||||||
)
|
|
||||||
|
|
||||||
|
|
||||||
def extract_json_text(raw_text: str) -> str:
|
|
||||||
fenced = JSON_BLOCK_RE.search(raw_text)
|
|
||||||
if fenced:
|
|
||||||
return fenced.group(1)
|
|
||||||
|
|
||||||
stripped = raw_text.strip()
|
|
||||||
start = stripped.find("{")
|
|
||||||
end = stripped.rfind("}")
|
|
||||||
if start == -1 or end == -1 or end <= start:
|
|
||||||
raise ValueError("Model output does not contain a JSON object.")
|
|
||||||
return stripped[start : end + 1]
|
|
||||||
|
|
||||||
|
|
||||||
def call_llm(
|
|
||||||
prompt: str,
|
|
||||||
timeout_seconds: float,
|
|
||||||
api_key: str | None,
|
|
||||||
model: str | None,
|
|
||||||
api_url: str | None,
|
|
||||||
) -> str:
|
|
||||||
api_key = api_key or os.environ.get("LLM_API_KEY") or os.environ.get("OPENAI_API_KEY")
|
|
||||||
model = model or os.environ.get("LLM_MODEL") or os.environ.get("OPENAI_MODEL")
|
|
||||||
api_url = api_url or os.environ.get("LLM_API_URL", "https://api.openai.com/v1/chat/completions")
|
|
||||||
|
|
||||||
if not api_key:
|
|
||||||
raise RuntimeError("Missing LLM_API_KEY or OPENAI_API_KEY, or pass --api-key.")
|
|
||||||
if not model:
|
|
||||||
raise RuntimeError("Missing LLM_MODEL or OPENAI_MODEL, or pass --model.")
|
|
||||||
|
|
||||||
headers = {
|
|
||||||
"Authorization": f"Bearer {api_key}",
|
|
||||||
"Content-Type": "application/json",
|
|
||||||
}
|
|
||||||
payload = {
|
|
||||||
"model": model,
|
|
||||||
"messages": [
|
|
||||||
{
|
|
||||||
"role": "system",
|
|
||||||
"content": "You are a precise JSON generator. Always output a single valid JSON object.",
|
|
||||||
},
|
|
||||||
{"role": "user", "content": prompt},
|
|
||||||
],
|
|
||||||
"temperature": 0.2,
|
|
||||||
}
|
|
||||||
|
|
||||||
with httpx.Client(timeout=timeout_seconds) as client:
|
|
||||||
response = client.post(api_url, headers=headers, json=payload)
|
|
||||||
response.raise_for_status()
|
|
||||||
data = response.json()
|
|
||||||
|
|
||||||
try:
|
|
||||||
return data["choices"][0]["message"]["content"]
|
|
||||||
except (KeyError, IndexError, TypeError) as exc:
|
|
||||||
raise RuntimeError(f"Unexpected LLM response shape: {json.dumps(data, ensure_ascii=False)[:1000]}") from exc
|
|
||||||
|
|
||||||
|
|
||||||
def run_loop(
|
|
||||||
extracted_path: Path,
|
|
||||||
prompt_path: Path,
|
|
||||||
output_path: Path,
|
|
||||||
max_retries: int,
|
|
||||||
timeout_seconds: float,
|
|
||||||
api_key: str | None,
|
|
||||||
model: str | None,
|
|
||||||
api_url: str | None,
|
|
||||||
) -> int:
|
|
||||||
extracted = load_json(extracted_path)
|
|
||||||
prompt_template = load_text(prompt_path)
|
|
||||||
output_path.parent.mkdir(parents=True, exist_ok=True)
|
|
||||||
|
|
||||||
last_errors: list[str] = []
|
|
||||||
last_result: dict[str, Any] | None = None
|
|
||||||
|
|
||||||
for attempt in range(1, max_retries + 2):
|
|
||||||
if attempt == 1:
|
|
||||||
prompt = build_initial_prompt(prompt_template, extracted)
|
|
||||||
else:
|
|
||||||
assert last_result is not None
|
|
||||||
prompt = build_repair_prompt(last_errors, extracted, last_result)
|
|
||||||
|
|
||||||
raw_output = call_llm(prompt, timeout_seconds, api_key, model, api_url)
|
|
||||||
raw_path = output_path.with_name(f"{output_path.stem}.attempt-{attempt}.raw.txt")
|
|
||||||
raw_path.write_text(raw_output, encoding="utf-8")
|
|
||||||
|
|
||||||
try:
|
|
||||||
result_payload = json.loads(extract_json_text(raw_output))
|
|
||||||
except (json.JSONDecodeError, ValueError) as exc:
|
|
||||||
last_errors = [f"Model output is not valid JSON: {exc}"]
|
|
||||||
last_result = {"raw_output": raw_output}
|
|
||||||
validation_path = output_path.with_name(f"{output_path.stem}.attempt-{attempt}.validation.json")
|
|
||||||
save_json(
|
|
||||||
validation_path,
|
|
||||||
{
|
|
||||||
"valid": False,
|
|
||||||
"errors": last_errors,
|
|
||||||
"warnings": [],
|
|
||||||
"normalized_result": None,
|
|
||||||
},
|
|
||||||
)
|
|
||||||
if attempt > max_retries:
|
|
||||||
output_path.write_text(raw_output, encoding="utf-8")
|
|
||||||
return 1
|
|
||||||
continue
|
|
||||||
|
|
||||||
attempt_path = output_path.with_name(f"{output_path.stem}.attempt-{attempt}.json")
|
|
||||||
save_json(attempt_path, result_payload)
|
|
||||||
save_json(output_path, result_payload)
|
|
||||||
|
|
||||||
report = validate_llm_result(output_path, extracted_path)
|
|
||||||
validation_path = output_path.with_name(f"{output_path.stem}.attempt-{attempt}.validation.json")
|
|
||||||
save_json(
|
|
||||||
validation_path,
|
|
||||||
{
|
|
||||||
"valid": report.valid,
|
|
||||||
"errors": report.errors,
|
|
||||||
"warnings": report.warnings,
|
|
||||||
"normalized_result": report.normalized_result,
|
|
||||||
},
|
|
||||||
)
|
|
||||||
|
|
||||||
if report.valid:
|
|
||||||
return 0
|
|
||||||
|
|
||||||
last_errors = report.errors
|
|
||||||
last_result = result_payload
|
|
||||||
|
|
||||||
return 1
|
|
||||||
|
|
||||||
|
|
||||||
def main() -> None:
|
def main() -> None:
|
||||||
@@ -214,7 +23,7 @@ def main() -> None:
|
|||||||
parser.add_argument("--timeout", type=float, default=60.0, help="LLM request timeout in seconds")
|
parser.add_argument("--timeout", type=float, default=60.0, help="LLM request timeout in seconds")
|
||||||
parser.add_argument("--api-key", type=str, default=None, help="LLM API key")
|
parser.add_argument("--api-key", type=str, default=None, help="LLM API key")
|
||||||
parser.add_argument("--model", type=str, default=None, help="LLM model name")
|
parser.add_argument("--model", type=str, default=None, help="LLM model name")
|
||||||
parser.add_argument("--api-url", type=str, default=None, help="LLM chat completions API URL")
|
parser.add_argument("--api-url", type=str, default=None, help="LLM chat completions API URL or base URL")
|
||||||
args = parser.parse_args()
|
args = parser.parse_args()
|
||||||
|
|
||||||
raise SystemExit(
|
raise SystemExit(
|
||||||
|
|||||||
@@ -0,0 +1,102 @@
|
|||||||
|
---
|
||||||
|
name: keyword-cleanup-review
|
||||||
|
description: Review and curate this repository's daily keyword index and frequency stats. Use when the user wants to inspect `data/term_index/term_stats.json`, recent `data/term_index/daily/*.json`, `configs/term_aliases.json`, `configs/term_stopwords.json`, or `configs/filter_context.personal.json` to propose alias merges, stopwords, watch terms, or `interest_keywords` updates without directly modifying configs.
|
||||||
|
---
|
||||||
|
|
||||||
|
# Keyword Cleanup Review
|
||||||
|
|
||||||
|
Use this skill to turn the repository's keyword statistics into reviewable cleanup suggestions.
|
||||||
|
|
||||||
|
## Workflow
|
||||||
|
|
||||||
|
1. Build a compact review bundle:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python skills/keyword-cleanup-review/scripts/build_review_bundle.py
|
||||||
|
```
|
||||||
|
|
||||||
|
Optional knobs:
|
||||||
|
|
||||||
|
- `--days 7`
|
||||||
|
- `--top 50`
|
||||||
|
- `--output outputs/term_index/review/keyword-cleanup-bundle.json`
|
||||||
|
|
||||||
|
2. Read the generated bundle and the suggestion schema:
|
||||||
|
|
||||||
|
- `outputs/term_index/review/keyword-cleanup-bundle.json`
|
||||||
|
- `skills/keyword-cleanup-review/references/suggestion-schema.md`
|
||||||
|
|
||||||
|
3. Produce two outputs:
|
||||||
|
|
||||||
|
- A short Markdown review for humans
|
||||||
|
- A JSON suggestion file matching the schema
|
||||||
|
|
||||||
|
4. Keep the boundary strict:
|
||||||
|
|
||||||
|
- Suggest changes to `configs/term_aliases.json`
|
||||||
|
- Suggest changes to `configs/term_stopwords.json`
|
||||||
|
- Suggest additions to `configs/filter_context.personal.json`
|
||||||
|
- Do not directly edit these files unless the user explicitly asks
|
||||||
|
- Do not suggest direct edits to `configs/filter_rules.json` unless the user asks for rule logic changes
|
||||||
|
|
||||||
|
## Review Heuristics
|
||||||
|
|
||||||
|
Prioritize these decisions:
|
||||||
|
|
||||||
|
- Alias suggestion
|
||||||
|
- Same concept with different naming, casing, abbreviation, or Chinese/English variants
|
||||||
|
- Stopword suggestion
|
||||||
|
- Too generic, too broad, or too noisy to help filtering
|
||||||
|
- Interest keyword suggestion
|
||||||
|
- High-frequency and aligned with the user's backend engineering, AI-agent, and frontier-tech focus
|
||||||
|
- Watch term
|
||||||
|
- Recent and potentially important, but evidence is still weak
|
||||||
|
|
||||||
|
Prefer conservative suggestions. If confidence is low, put the term into `watch_terms`.
|
||||||
|
|
||||||
|
## Inputs
|
||||||
|
|
||||||
|
Primary inputs:
|
||||||
|
|
||||||
|
- `data/term_index/term_stats.json`
|
||||||
|
- `data/term_index/daily/*.json`
|
||||||
|
- `configs/term_aliases.json`
|
||||||
|
- `configs/term_stopwords.json`
|
||||||
|
- `configs/filter_context.personal.json`
|
||||||
|
- `configs/term_cleanup_policy.json`
|
||||||
|
- `configs/term_watchlist.json`
|
||||||
|
- `configs/term_change_log.json`
|
||||||
|
|
||||||
|
The bundled script already compacts these into a single review bundle.
|
||||||
|
|
||||||
|
## Output Expectations
|
||||||
|
|
||||||
|
The Markdown output should:
|
||||||
|
|
||||||
|
- Summarize the current state briefly
|
||||||
|
- List the top terms worth acting on
|
||||||
|
- Separate alias, stopword, interest-keyword, and watch-term recommendations
|
||||||
|
- Explain reasoning in short, concrete sentences
|
||||||
|
|
||||||
|
The JSON output should follow:
|
||||||
|
|
||||||
|
- `references/suggestion-schema.md`
|
||||||
|
|
||||||
|
## Repository Notes
|
||||||
|
|
||||||
|
Current repository behavior:
|
||||||
|
|
||||||
|
- Keyword stats are program-maintained, not LLM-maintained
|
||||||
|
- Stats are built from `keywords`, not `topics`
|
||||||
|
- Stats only include non-`drop` candidates
|
||||||
|
- `data/term_index/term_stats.json` is rebuilt from daily files, so reruns overwrite the same day instead of double-counting
|
||||||
|
- cleanup policy, watchlist, and change log are repository-managed governance inputs and should be respected during review
|
||||||
|
|
||||||
|
Keep suggestions aligned with that design.
|
||||||
|
|
||||||
|
## Resources
|
||||||
|
|
||||||
|
- Script:
|
||||||
|
- `scripts/build_review_bundle.py`
|
||||||
|
- Reference:
|
||||||
|
- `references/suggestion-schema.md`
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
display_name: Keyword Cleanup Review
|
||||||
|
short_description: Review keyword stats and suggest cleanup updates.
|
||||||
|
default_prompt: Review the repository keyword index, identify duplicate or noisy terms, and produce alias, stopword, watch-term, and interest-keyword suggestions without modifying configs directly.
|
||||||
@@ -0,0 +1,47 @@
|
|||||||
|
# Suggestion Schema
|
||||||
|
|
||||||
|
When this skill produces review output, prefer two files:
|
||||||
|
|
||||||
|
- Markdown summary for humans
|
||||||
|
- JSON suggestions for deterministic follow-up edits
|
||||||
|
|
||||||
|
Recommended JSON shape:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"date": "2026-03-27",
|
||||||
|
"based_on_days": 7,
|
||||||
|
"alias_suggestions": [
|
||||||
|
{
|
||||||
|
"from": "Agent",
|
||||||
|
"to": "AI Agent",
|
||||||
|
"reason": "High overlap with existing repository terminology."
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"stopword_suggestions": [
|
||||||
|
{
|
||||||
|
"term": "??",
|
||||||
|
"reason": "Too generic to be useful for filtering."
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"interest_keyword_suggestions": [
|
||||||
|
{
|
||||||
|
"term": "MCP",
|
||||||
|
"reason": "Repeated high-frequency term aligned with current learning focus."
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"watch_terms": [
|
||||||
|
{
|
||||||
|
"term": "OpenClaw",
|
||||||
|
"reason": "Trending recently but not enough history yet."
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Rules:
|
||||||
|
|
||||||
|
- Only output suggestions, never claim they are already applied.
|
||||||
|
- Avoid suggesting changes that would directly edit `filter_rules.json`.
|
||||||
|
- Prefer small, reviewable batches over large refactors.
|
||||||
|
- If evidence is weak, put the term in `watch_terms` instead of alias or stopword suggestions.
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
Provide keyword cleanup guidance for this repository's daily keyword index. Use this skill when the user wants to review, merge, clean, or curate `data/term_index/term_stats.json`, recent `data/term_index/daily/*.json`, `configs/term_aliases.json`, `configs/term_stopwords.json`, or `configs/filter_context.personal.json`, especially to propose alias merges, stopwords, or `interest_keywords` updates without directly modifying configs.
|
||||||
@@ -0,0 +1,340 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
from collections import defaultdict
|
||||||
|
from datetime import UTC, datetime
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
|
||||||
|
DEFAULT_POLICY: dict[str, Any] = {
|
||||||
|
"schema_version": "v1",
|
||||||
|
"interest_keyword_review": {
|
||||||
|
"min_total_count": 3,
|
||||||
|
"min_days_seen": 2,
|
||||||
|
},
|
||||||
|
"watch_term_review": {
|
||||||
|
"min_total_count": 1,
|
||||||
|
"min_days_seen": 1,
|
||||||
|
"max_total_count": 2,
|
||||||
|
"max_days_seen": 2,
|
||||||
|
},
|
||||||
|
"alias_review": {
|
||||||
|
"min_total_count": 2,
|
||||||
|
"min_days_seen": 2,
|
||||||
|
},
|
||||||
|
"stopword_review": {
|
||||||
|
"max_total_count": 2,
|
||||||
|
"max_days_seen": 2,
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _load_json(path: Path) -> Any:
|
||||||
|
return json.loads(path.read_text(encoding="utf-8-sig"))
|
||||||
|
|
||||||
|
|
||||||
|
def _save_json(path: Path, payload: dict[str, Any]) -> None:
|
||||||
|
path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
path.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
|
def _load_mapping(path: Path) -> dict[str, str]:
|
||||||
|
if not path.exists():
|
||||||
|
return {}
|
||||||
|
payload = _load_json(path)
|
||||||
|
return payload if isinstance(payload, dict) else {}
|
||||||
|
|
||||||
|
|
||||||
|
def _load_list(path: Path) -> list[str]:
|
||||||
|
if not path.exists():
|
||||||
|
return []
|
||||||
|
payload = _load_json(path)
|
||||||
|
if isinstance(payload, list):
|
||||||
|
return [item for item in payload if isinstance(item, str)]
|
||||||
|
return []
|
||||||
|
|
||||||
|
|
||||||
|
def _load_interest_keywords(path: Path) -> list[str]:
|
||||||
|
if not path.exists():
|
||||||
|
return []
|
||||||
|
payload = _load_json(path)
|
||||||
|
if not isinstance(payload, dict):
|
||||||
|
return []
|
||||||
|
values = payload.get("interest_keywords")
|
||||||
|
if not isinstance(values, list):
|
||||||
|
return []
|
||||||
|
return [item for item in values if isinstance(item, str)]
|
||||||
|
|
||||||
|
|
||||||
|
def _casefold_set(values: list[str]) -> set[str]:
|
||||||
|
return {value.strip().casefold() for value in values if value.strip()}
|
||||||
|
|
||||||
|
|
||||||
|
def _load_policy(path: Path) -> dict[str, Any]:
|
||||||
|
if not path.exists():
|
||||||
|
return DEFAULT_POLICY.copy()
|
||||||
|
payload = _load_json(path)
|
||||||
|
return payload if isinstance(payload, dict) else DEFAULT_POLICY.copy()
|
||||||
|
|
||||||
|
|
||||||
|
def _load_watchlist(path: Path) -> list[dict[str, Any]]:
|
||||||
|
if not path.exists():
|
||||||
|
return []
|
||||||
|
payload = _load_json(path)
|
||||||
|
if not isinstance(payload, dict):
|
||||||
|
return []
|
||||||
|
terms = payload.get("terms")
|
||||||
|
if not isinstance(terms, list):
|
||||||
|
return []
|
||||||
|
return [item for item in terms if isinstance(item, dict) and isinstance(item.get("term"), str)]
|
||||||
|
|
||||||
|
|
||||||
|
def _load_change_log(path: Path) -> list[dict[str, Any]]:
|
||||||
|
if not path.exists():
|
||||||
|
return []
|
||||||
|
payload = _load_json(path)
|
||||||
|
if not isinstance(payload, dict):
|
||||||
|
return []
|
||||||
|
entries = payload.get("entries")
|
||||||
|
if not isinstance(entries, list):
|
||||||
|
return []
|
||||||
|
return [item for item in entries if isinstance(item, dict)]
|
||||||
|
|
||||||
|
|
||||||
|
def _meets_min_thresholds(item: dict[str, Any], thresholds: dict[str, Any]) -> bool:
|
||||||
|
min_total_count = int(thresholds.get("min_total_count", 1))
|
||||||
|
min_days_seen = int(thresholds.get("min_days_seen", 1))
|
||||||
|
total_count = int(item.get("total_count") or 0)
|
||||||
|
days_seen = int(item.get("days_seen") or 0)
|
||||||
|
return total_count >= min_total_count and days_seen >= min_days_seen
|
||||||
|
|
||||||
|
|
||||||
|
def _within_watch_thresholds(item: dict[str, Any], thresholds: dict[str, Any]) -> bool:
|
||||||
|
min_total_count = int(thresholds.get("min_total_count", 1))
|
||||||
|
min_days_seen = int(thresholds.get("min_days_seen", 1))
|
||||||
|
max_total_count = int(thresholds.get("max_total_count", 999999))
|
||||||
|
max_days_seen = int(thresholds.get("max_days_seen", 999999))
|
||||||
|
total_count = int(item.get("total_count") or 0)
|
||||||
|
days_seen = int(item.get("days_seen") or 0)
|
||||||
|
return (
|
||||||
|
total_count >= min_total_count
|
||||||
|
and days_seen >= min_days_seen
|
||||||
|
and total_count <= max_total_count
|
||||||
|
and days_seen <= max_days_seen
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
parser = argparse.ArgumentParser(
|
||||||
|
description="Build a compact review bundle for the keyword-cleanup-review skill."
|
||||||
|
)
|
||||||
|
repo_root = Path(__file__).resolve().parents[3]
|
||||||
|
parser.add_argument(
|
||||||
|
"--stats",
|
||||||
|
type=Path,
|
||||||
|
default=repo_root / "data" / "term_index" / "term_stats.json",
|
||||||
|
help="Global keyword stats file",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--daily-dir",
|
||||||
|
type=Path,
|
||||||
|
default=repo_root / "data" / "term_index" / "daily",
|
||||||
|
help="Directory containing daily keyword index JSON files",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--aliases",
|
||||||
|
type=Path,
|
||||||
|
default=repo_root / "configs" / "term_aliases.json",
|
||||||
|
help="Term aliases JSON file",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--stopwords",
|
||||||
|
type=Path,
|
||||||
|
default=repo_root / "configs" / "term_stopwords.json",
|
||||||
|
help="Term stopwords JSON file",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--context",
|
||||||
|
type=Path,
|
||||||
|
default=repo_root / "configs" / "filter_context.personal.json",
|
||||||
|
help="Personal filter context JSON file",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--policy",
|
||||||
|
type=Path,
|
||||||
|
default=repo_root / "configs" / "term_cleanup_policy.json",
|
||||||
|
help="Keyword cleanup policy JSON file",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--watchlist",
|
||||||
|
type=Path,
|
||||||
|
default=repo_root / "configs" / "term_watchlist.json",
|
||||||
|
help="Tracked watch terms JSON file",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--change-log",
|
||||||
|
type=Path,
|
||||||
|
default=repo_root / "configs" / "term_change_log.json",
|
||||||
|
help="Applied keyword cleanup change log JSON file",
|
||||||
|
)
|
||||||
|
parser.add_argument("--days", type=int, default=7, help="How many recent daily files to include")
|
||||||
|
parser.add_argument("--top", type=int, default=50, help="How many top global terms to include")
|
||||||
|
parser.add_argument(
|
||||||
|
"--output",
|
||||||
|
type=Path,
|
||||||
|
default=repo_root / "outputs" / "term_index" / "review" / "keyword-cleanup-bundle.json",
|
||||||
|
help="Output JSON file",
|
||||||
|
)
|
||||||
|
args = parser.parse_args()
|
||||||
|
|
||||||
|
stats_payload = _load_json(args.stats) if args.stats.exists() else {"terms": []}
|
||||||
|
stats_terms = stats_payload.get("terms") if isinstance(stats_payload, dict) else []
|
||||||
|
if not isinstance(stats_terms, list):
|
||||||
|
stats_terms = []
|
||||||
|
|
||||||
|
aliases = _load_mapping(args.aliases)
|
||||||
|
stopwords = _load_list(args.stopwords)
|
||||||
|
interest_keywords = _load_interest_keywords(args.context)
|
||||||
|
policy = _load_policy(args.policy)
|
||||||
|
watchlist = _load_watchlist(args.watchlist)
|
||||||
|
change_log = _load_change_log(args.change_log)
|
||||||
|
|
||||||
|
recent_daily_paths = sorted(args.daily_dir.glob("*.json"))[-args.days :] if args.daily_dir.exists() else []
|
||||||
|
recent_daily: list[dict[str, Any]] = []
|
||||||
|
recent_counter: defaultdict[str, int] = defaultdict(int)
|
||||||
|
for path in recent_daily_paths:
|
||||||
|
payload = _load_json(path)
|
||||||
|
if not isinstance(payload, dict):
|
||||||
|
continue
|
||||||
|
terms = payload.get("terms")
|
||||||
|
if not isinstance(terms, list):
|
||||||
|
terms = []
|
||||||
|
compact_terms = []
|
||||||
|
for term in terms:
|
||||||
|
if not isinstance(term, dict):
|
||||||
|
continue
|
||||||
|
normalized = term.get("normalized_term")
|
||||||
|
count = term.get("count")
|
||||||
|
if not isinstance(normalized, str) or not isinstance(count, int):
|
||||||
|
continue
|
||||||
|
compact_terms.append({"term": normalized, "count": count})
|
||||||
|
recent_counter[normalized] += count
|
||||||
|
recent_daily.append(
|
||||||
|
{
|
||||||
|
"date": payload.get("date"),
|
||||||
|
"candidate_count": payload.get("candidate_count"),
|
||||||
|
"top_terms": compact_terms[:20],
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
interest_set = _casefold_set(interest_keywords)
|
||||||
|
stopword_set = _casefold_set(stopwords)
|
||||||
|
alias_keys = _casefold_set(list(aliases.keys()))
|
||||||
|
alias_values = _casefold_set(list(aliases.values()))
|
||||||
|
watch_set = _casefold_set([str(item.get("term", "")) for item in watchlist])
|
||||||
|
|
||||||
|
top_global_terms = []
|
||||||
|
for item in stats_terms[: args.top]:
|
||||||
|
if not isinstance(item, dict):
|
||||||
|
continue
|
||||||
|
term = item.get("term")
|
||||||
|
if not isinstance(term, str):
|
||||||
|
continue
|
||||||
|
folded = term.strip().casefold()
|
||||||
|
top_global_terms.append(
|
||||||
|
{
|
||||||
|
"term": term,
|
||||||
|
"total_count": item.get("total_count"),
|
||||||
|
"days_seen": item.get("days_seen"),
|
||||||
|
"first_seen": item.get("first_seen"),
|
||||||
|
"last_seen": item.get("last_seen"),
|
||||||
|
"in_interest_keywords": folded in interest_set,
|
||||||
|
"is_stopword": folded in stopword_set,
|
||||||
|
"is_alias_source": folded in alias_keys,
|
||||||
|
"is_alias_target": folded in alias_values,
|
||||||
|
"in_watchlist": folded in watch_set,
|
||||||
|
"recent_count": recent_counter.get(term, 0),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
uncovered_terms = [
|
||||||
|
item for item in top_global_terms if not item["in_interest_keywords"] and not item["is_stopword"]
|
||||||
|
][:20]
|
||||||
|
interest_thresholds = policy.get("interest_keyword_review") if isinstance(policy, dict) else {}
|
||||||
|
watch_thresholds = policy.get("watch_term_review") if isinstance(policy, dict) else {}
|
||||||
|
interest_review_candidates = [
|
||||||
|
{
|
||||||
|
"term": item["term"],
|
||||||
|
"total_count": item["total_count"],
|
||||||
|
"days_seen": item["days_seen"],
|
||||||
|
"reason": (
|
||||||
|
"Meets the configured interest-keyword review threshold and is not yet covered "
|
||||||
|
"by interest keywords or stopwords."
|
||||||
|
),
|
||||||
|
}
|
||||||
|
for item in uncovered_terms
|
||||||
|
if not item["in_watchlist"] and _meets_min_thresholds(item, interest_thresholds)
|
||||||
|
][:20]
|
||||||
|
watch_review_candidates = [
|
||||||
|
{
|
||||||
|
"term": item["term"],
|
||||||
|
"total_count": item["total_count"],
|
||||||
|
"days_seen": item["days_seen"],
|
||||||
|
"reason": (
|
||||||
|
"Falls into the configured watch-term review range and should be observed "
|
||||||
|
"before promotion into interest keywords."
|
||||||
|
),
|
||||||
|
}
|
||||||
|
for item in uncovered_terms
|
||||||
|
if not item["in_watchlist"]
|
||||||
|
and not _meets_min_thresholds(item, interest_thresholds)
|
||||||
|
and _within_watch_thresholds(item, watch_thresholds)
|
||||||
|
][:20]
|
||||||
|
recent_hot_terms = sorted(
|
||||||
|
({"term": term, "recent_count": count} for term, count in recent_counter.items()),
|
||||||
|
key=lambda item: (-item["recent_count"], item["term"].casefold(), item["term"]),
|
||||||
|
)[:20]
|
||||||
|
|
||||||
|
bundle = {
|
||||||
|
"generated_at": datetime.now(tz=UTC).isoformat(),
|
||||||
|
"days": args.days,
|
||||||
|
"top": args.top,
|
||||||
|
"sources": {
|
||||||
|
"stats": str(args.stats),
|
||||||
|
"daily_dir": str(args.daily_dir),
|
||||||
|
"aliases": str(args.aliases),
|
||||||
|
"stopwords": str(args.stopwords),
|
||||||
|
"context": str(args.context),
|
||||||
|
"policy": str(args.policy),
|
||||||
|
"watchlist": str(args.watchlist),
|
||||||
|
"change_log": str(args.change_log),
|
||||||
|
},
|
||||||
|
"policy": policy,
|
||||||
|
"current_config": {
|
||||||
|
"alias_count": len(aliases),
|
||||||
|
"stopword_count": len(stopwords),
|
||||||
|
"interest_keyword_count": len(interest_keywords),
|
||||||
|
"watch_term_count": len(watchlist),
|
||||||
|
"aliases": aliases,
|
||||||
|
"stopwords": stopwords,
|
||||||
|
"interest_keywords": interest_keywords,
|
||||||
|
"watchlist": watchlist,
|
||||||
|
},
|
||||||
|
"change_log_tail": change_log[-10:],
|
||||||
|
"top_global_terms": top_global_terms,
|
||||||
|
"recent_daily": recent_daily,
|
||||||
|
"recent_hot_terms": recent_hot_terms,
|
||||||
|
"uncovered_terms": uncovered_terms,
|
||||||
|
"governance_hints": {
|
||||||
|
"interest_review_candidates": interest_review_candidates,
|
||||||
|
"watch_review_candidates": watch_review_candidates,
|
||||||
|
},
|
||||||
|
}
|
||||||
|
_save_json(args.output, bundle)
|
||||||
|
print(f"Saved keyword cleanup bundle to {args.output}")
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -1,27 +0,0 @@
|
|||||||
Metadata-Version: 2.4
|
|
||||||
Name: summary-mcp
|
|
||||||
Version: 0.1.0
|
|
||||||
Summary: MCP service scaffold for article summarization
|
|
||||||
Requires-Python: >=3.11
|
|
||||||
Description-Content-Type: text/markdown
|
|
||||||
Requires-Dist: beautifulsoup4>=4.12.0
|
|
||||||
Requires-Dist: httpx>=0.27.0
|
|
||||||
Requires-Dist: mcp>=1.17.0
|
|
||||||
Requires-Dist: pydantic>=2.9.0
|
|
||||||
Requires-Dist: trafilatura>=1.12.0
|
|
||||||
|
|
||||||
# Summary MCP
|
|
||||||
|
|
||||||
Python MCP scaffold for article summarization.
|
|
||||||
|
|
||||||
## Run
|
|
||||||
|
|
||||||
```bash
|
|
||||||
pip install -e .
|
|
||||||
summary-mcp
|
|
||||||
```
|
|
||||||
|
|
||||||
The server exposes two tools:
|
|
||||||
|
|
||||||
- `summarize_url`
|
|
||||||
- `summarize_item`
|
|
||||||
@@ -1,23 +0,0 @@
|
|||||||
README.md
|
|
||||||
pyproject.toml
|
|
||||||
src/summary_mcp/__init__.py
|
|
||||||
src/summary_mcp/server.py
|
|
||||||
src/summary_mcp.egg-info/PKG-INFO
|
|
||||||
src/summary_mcp.egg-info/SOURCES.txt
|
|
||||||
src/summary_mcp.egg-info/dependency_links.txt
|
|
||||||
src/summary_mcp.egg-info/entry_points.txt
|
|
||||||
src/summary_mcp.egg-info/requires.txt
|
|
||||||
src/summary_mcp.egg-info/top_level.txt
|
|
||||||
src/summary_mcp/core/__init__.py
|
|
||||||
src/summary_mcp/core/content_loader.py
|
|
||||||
src/summary_mcp/core/errors.py
|
|
||||||
src/summary_mcp/core/extractor.py
|
|
||||||
src/summary_mcp/core/mapper.py
|
|
||||||
src/summary_mcp/core/normalizer.py
|
|
||||||
src/summary_mcp/core/pipeline.py
|
|
||||||
src/summary_mcp/core/quality_checker.py
|
|
||||||
src/summary_mcp/core/summarizer.py
|
|
||||||
src/summary_mcp/models/__init__.py
|
|
||||||
src/summary_mcp/models/document.py
|
|
||||||
src/summary_mcp/models/item.py
|
|
||||||
src/summary_mcp/models/summary_io.py
|
|
||||||
@@ -1 +0,0 @@
|
|||||||
|
|
||||||
@@ -1,2 +0,0 @@
|
|||||||
[console_scripts]
|
|
||||||
summary-mcp = summary_mcp.server:main
|
|
||||||
@@ -1,5 +0,0 @@
|
|||||||
beautifulsoup4>=4.12.0
|
|
||||||
httpx>=0.27.0
|
|
||||||
mcp>=1.17.0
|
|
||||||
pydantic>=2.9.0
|
|
||||||
trafilatura>=1.12.0
|
|
||||||
@@ -1 +0,0 @@
|
|||||||
summary_mcp
|
|
||||||
@@ -6,14 +6,24 @@ from summary_mcp.core.errors import SummaryError
|
|||||||
from summary_mcp.models.summary_io import ExtractionInput
|
from summary_mcp.models.summary_io import ExtractionInput
|
||||||
|
|
||||||
|
|
||||||
|
MIN_INLINE_CONTENT_LENGTH = 500
|
||||||
|
|
||||||
|
|
||||||
|
def _has_usable_content(value: str | None) -> bool:
|
||||||
|
return bool(value and len(value.strip()) >= MIN_INLINE_CONTENT_LENGTH)
|
||||||
|
|
||||||
|
|
||||||
def choose_inline_content(extraction_input: ExtractionInput) -> tuple[str | None, str]:
|
def choose_inline_content(extraction_input: ExtractionInput) -> tuple[str | None, str]:
|
||||||
if extraction_input.raw_html:
|
if extraction_input.raw_html:
|
||||||
return extraction_input.raw_html, "raw_html"
|
return extraction_input.raw_html, "raw_html"
|
||||||
|
|
||||||
if extraction_input.item and extraction_input.item.raw_content and len(extraction_input.item.raw_content.strip()) >= 500:
|
if extraction_input.item and _has_usable_content(extraction_input.item.raw_content):
|
||||||
return extraction_input.item.raw_content, "item.raw_content"
|
return extraction_input.item.raw_content, "item.raw_content"
|
||||||
|
|
||||||
if extraction_input.rss_content and len(extraction_input.rss_content.strip()) >= 500:
|
if extraction_input.item and _has_usable_content(extraction_input.item.raw_summary):
|
||||||
|
return extraction_input.item.raw_summary, "item.raw_summary"
|
||||||
|
|
||||||
|
if _has_usable_content(extraction_input.rss_content):
|
||||||
return extraction_input.rss_content, "rss_content"
|
return extraction_input.rss_content, "rss_content"
|
||||||
|
|
||||||
return None, "none"
|
return None, "none"
|
||||||
|
|||||||
@@ -0,0 +1,228 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
from collections import Counter
|
||||||
|
from datetime import UTC, date, datetime
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Iterable
|
||||||
|
|
||||||
|
from summary_mcp.models.article_candidate import OpenClawCandidateInput
|
||||||
|
from summary_mcp.models.keyword_index import (
|
||||||
|
DailyKeywordIndex,
|
||||||
|
DailyKeywordTerm,
|
||||||
|
KeywordStat,
|
||||||
|
KeywordStatsIndex,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
REPO_ROOT = Path(__file__).resolve().parents[3]
|
||||||
|
DATA_ROOT = REPO_ROOT / "data" / "term_index"
|
||||||
|
DEFAULT_DAILY_DIR = DATA_ROOT / "daily"
|
||||||
|
DEFAULT_STATS_PATH = DATA_ROOT / "term_stats.json"
|
||||||
|
DEFAULT_ALIASES_PATH = REPO_ROOT / "configs" / "term_aliases.json"
|
||||||
|
DEFAULT_STOPWORDS_PATH = REPO_ROOT / "configs" / "term_stopwords.json"
|
||||||
|
DEFAULT_INCLUDE_DECISIONS = {"keep", "review"}
|
||||||
|
|
||||||
|
|
||||||
|
def _load_json(path: Path) -> object:
|
||||||
|
return json.loads(path.read_text(encoding="utf-8-sig"))
|
||||||
|
|
||||||
|
|
||||||
|
def _save_json(path: Path, payload: dict | list) -> None:
|
||||||
|
path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
path.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
|
def _term_key(value: str) -> str:
|
||||||
|
return value.strip().casefold()
|
||||||
|
|
||||||
|
|
||||||
|
def load_term_aliases(path: Path | None = None) -> dict[str, str]:
|
||||||
|
aliases_path = path or DEFAULT_ALIASES_PATH
|
||||||
|
if not aliases_path.exists():
|
||||||
|
return {}
|
||||||
|
|
||||||
|
payload = _load_json(aliases_path)
|
||||||
|
if not isinstance(payload, dict):
|
||||||
|
raise RuntimeError("Term aliases file must contain a JSON object.")
|
||||||
|
|
||||||
|
aliases: dict[str, str] = {}
|
||||||
|
for raw_key, raw_value in payload.items():
|
||||||
|
if not isinstance(raw_key, str) or not isinstance(raw_value, str):
|
||||||
|
continue
|
||||||
|
normalized_key = _term_key(raw_key)
|
||||||
|
normalized_value = raw_value.strip()
|
||||||
|
if not normalized_key or not normalized_value:
|
||||||
|
continue
|
||||||
|
aliases[normalized_key] = normalized_value
|
||||||
|
return aliases
|
||||||
|
|
||||||
|
|
||||||
|
def load_term_stopwords(path: Path | None = None) -> set[str]:
|
||||||
|
stopwords_path = path or DEFAULT_STOPWORDS_PATH
|
||||||
|
if not stopwords_path.exists():
|
||||||
|
return set()
|
||||||
|
|
||||||
|
payload = _load_json(stopwords_path)
|
||||||
|
if not isinstance(payload, list):
|
||||||
|
raise RuntimeError("Term stopwords file must contain a JSON array.")
|
||||||
|
|
||||||
|
values: set[str] = set()
|
||||||
|
for item in payload:
|
||||||
|
if not isinstance(item, str):
|
||||||
|
continue
|
||||||
|
normalized = _term_key(item)
|
||||||
|
if normalized:
|
||||||
|
values.add(normalized)
|
||||||
|
return values
|
||||||
|
|
||||||
|
|
||||||
|
def normalize_keyword(
|
||||||
|
keyword: str,
|
||||||
|
*,
|
||||||
|
aliases: dict[str, str],
|
||||||
|
stopwords: set[str],
|
||||||
|
) -> str | None:
|
||||||
|
raw_value = keyword.strip()
|
||||||
|
if not raw_value:
|
||||||
|
return None
|
||||||
|
|
||||||
|
aliased_value = aliases.get(_term_key(raw_value), raw_value).strip()
|
||||||
|
if not aliased_value:
|
||||||
|
return None
|
||||||
|
if _term_key(aliased_value) in stopwords:
|
||||||
|
return None
|
||||||
|
return aliased_value
|
||||||
|
|
||||||
|
|
||||||
|
def build_daily_keyword_index(
|
||||||
|
candidates: Iterable[OpenClawCandidateInput],
|
||||||
|
*,
|
||||||
|
for_date: date,
|
||||||
|
digest_id: str,
|
||||||
|
source: str = "openclaw_delivery_payload",
|
||||||
|
include_decisions: set[str] | None = None,
|
||||||
|
aliases: dict[str, str] | None = None,
|
||||||
|
stopwords: set[str] | None = None,
|
||||||
|
) -> DailyKeywordIndex:
|
||||||
|
allowed_decisions = include_decisions or DEFAULT_INCLUDE_DECISIONS
|
||||||
|
resolved_aliases = aliases or {}
|
||||||
|
resolved_stopwords = stopwords or set()
|
||||||
|
|
||||||
|
term_counter: Counter[str] = Counter()
|
||||||
|
candidate_count = 0
|
||||||
|
|
||||||
|
for candidate in candidates:
|
||||||
|
if candidate.selection_decision not in allowed_decisions:
|
||||||
|
continue
|
||||||
|
|
||||||
|
candidate_count += 1
|
||||||
|
seen_for_candidate: set[str] = set()
|
||||||
|
for keyword in candidate.keywords:
|
||||||
|
normalized = normalize_keyword(
|
||||||
|
keyword,
|
||||||
|
aliases=resolved_aliases,
|
||||||
|
stopwords=resolved_stopwords,
|
||||||
|
)
|
||||||
|
if normalized is None or normalized in seen_for_candidate:
|
||||||
|
continue
|
||||||
|
seen_for_candidate.add(normalized)
|
||||||
|
term_counter[normalized] += 1
|
||||||
|
|
||||||
|
terms = [
|
||||||
|
DailyKeywordTerm(term=term, normalized_term=term, count=count)
|
||||||
|
for term, count in sorted(term_counter.items(), key=lambda item: (-item[1], item[0].casefold(), item[0]))
|
||||||
|
]
|
||||||
|
|
||||||
|
return DailyKeywordIndex(
|
||||||
|
schema_version="v1",
|
||||||
|
date=for_date,
|
||||||
|
source=source,
|
||||||
|
digest_id=digest_id,
|
||||||
|
generated_at=datetime.now(tz=UTC),
|
||||||
|
candidate_count=candidate_count,
|
||||||
|
terms=terms,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def read_daily_keyword_index(path: Path) -> DailyKeywordIndex:
|
||||||
|
return DailyKeywordIndex.model_validate(_load_json(path))
|
||||||
|
|
||||||
|
|
||||||
|
def rebuild_keyword_stats(*, daily_dir: Path = DEFAULT_DAILY_DIR) -> KeywordStatsIndex:
|
||||||
|
term_totals: dict[str, int] = {}
|
||||||
|
first_seen: dict[str, date] = {}
|
||||||
|
last_seen: dict[str, date] = {}
|
||||||
|
days_seen: dict[str, int] = {}
|
||||||
|
|
||||||
|
if daily_dir.exists():
|
||||||
|
for path in sorted(daily_dir.glob("*.json")):
|
||||||
|
daily_index = read_daily_keyword_index(path)
|
||||||
|
seen_today: set[str] = set()
|
||||||
|
for term in daily_index.terms:
|
||||||
|
normalized = term.normalized_term
|
||||||
|
term_totals[normalized] = term_totals.get(normalized, 0) + term.count
|
||||||
|
if normalized not in first_seen or daily_index.date < first_seen[normalized]:
|
||||||
|
first_seen[normalized] = daily_index.date
|
||||||
|
if normalized not in last_seen or daily_index.date > last_seen[normalized]:
|
||||||
|
last_seen[normalized] = daily_index.date
|
||||||
|
if normalized not in seen_today:
|
||||||
|
days_seen[normalized] = days_seen.get(normalized, 0) + 1
|
||||||
|
seen_today.add(normalized)
|
||||||
|
|
||||||
|
terms = [
|
||||||
|
KeywordStat(
|
||||||
|
term=term,
|
||||||
|
total_count=term_totals[term],
|
||||||
|
days_seen=days_seen[term],
|
||||||
|
first_seen=first_seen[term],
|
||||||
|
last_seen=last_seen[term],
|
||||||
|
)
|
||||||
|
for term in sorted(term_totals, key=lambda item: (-term_totals[item], item.casefold(), item))
|
||||||
|
]
|
||||||
|
|
||||||
|
return KeywordStatsIndex(
|
||||||
|
schema_version="v1",
|
||||||
|
generated_at=datetime.now(tz=UTC),
|
||||||
|
terms=terms,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def persist_keyword_indexes(
|
||||||
|
candidates: Iterable[OpenClawCandidateInput],
|
||||||
|
*,
|
||||||
|
for_date: date,
|
||||||
|
digest_id: str,
|
||||||
|
source: str = "openclaw_delivery_payload",
|
||||||
|
daily_dir: Path = DEFAULT_DAILY_DIR,
|
||||||
|
stats_path: Path = DEFAULT_STATS_PATH,
|
||||||
|
aliases_path: Path | None = None,
|
||||||
|
stopwords_path: Path | None = None,
|
||||||
|
include_decisions: set[str] | None = None,
|
||||||
|
) -> dict[str, object]:
|
||||||
|
aliases = load_term_aliases(aliases_path)
|
||||||
|
stopwords = load_term_stopwords(stopwords_path)
|
||||||
|
daily_index = build_daily_keyword_index(
|
||||||
|
candidates,
|
||||||
|
for_date=for_date,
|
||||||
|
digest_id=digest_id,
|
||||||
|
source=source,
|
||||||
|
include_decisions=include_decisions,
|
||||||
|
aliases=aliases,
|
||||||
|
stopwords=stopwords,
|
||||||
|
)
|
||||||
|
|
||||||
|
daily_path = daily_dir / f"{for_date.isoformat()}.json"
|
||||||
|
_save_json(daily_path, daily_index.model_dump(mode="json"))
|
||||||
|
|
||||||
|
stats_index = rebuild_keyword_stats(daily_dir=daily_dir)
|
||||||
|
_save_json(stats_path, stats_index.model_dump(mode="json"))
|
||||||
|
|
||||||
|
return {
|
||||||
|
"daily_output": str(daily_path),
|
||||||
|
"stats_output": str(stats_path),
|
||||||
|
"source": source,
|
||||||
|
"candidate_count": daily_index.candidate_count,
|
||||||
|
"term_count": len(daily_index.terms),
|
||||||
|
"top_terms": [term.model_dump(mode="json") for term in daily_index.terms[:10]],
|
||||||
|
}
|
||||||
@@ -9,6 +9,19 @@ from summary_mcp.core.quality_checker import assess_quality
|
|||||||
from summary_mcp.models.summary_io import DebugInfo, ExtractionInput, ExtractionOutput
|
from summary_mcp.models.summary_io import DebugInfo, ExtractionInput, ExtractionOutput
|
||||||
|
|
||||||
|
|
||||||
|
RSS_ONLY_UPSTREAMS = {"freshrss"}
|
||||||
|
|
||||||
|
|
||||||
|
def _should_skip_fetch(extraction_input: ExtractionInput) -> bool:
|
||||||
|
item = extraction_input.item
|
||||||
|
if item is None:
|
||||||
|
return False
|
||||||
|
|
||||||
|
metadata = item.metadata if isinstance(item.metadata, dict) else {}
|
||||||
|
upstream = metadata.get("upstream")
|
||||||
|
return isinstance(upstream, str) and upstream in RSS_ONLY_UPSTREAMS
|
||||||
|
|
||||||
|
|
||||||
def extract_content(extraction_input: ExtractionInput) -> ExtractionOutput:
|
def extract_content(extraction_input: ExtractionInput) -> ExtractionOutput:
|
||||||
try:
|
try:
|
||||||
normalized = normalize_input(extraction_input)
|
normalized = normalize_input(extraction_input)
|
||||||
@@ -16,11 +29,20 @@ def extract_content(extraction_input: ExtractionInput) -> ExtractionOutput:
|
|||||||
|
|
||||||
inline_content, content_source = choose_inline_content(normalized)
|
inline_content, content_source = choose_inline_content(normalized)
|
||||||
if inline_content is None:
|
if inline_content is None:
|
||||||
|
if _should_skip_fetch(normalized):
|
||||||
|
raise SummaryError(
|
||||||
|
code="RSS_CONTENT_MISSING",
|
||||||
|
message="Skipping item because RSS content is unavailable.",
|
||||||
|
retryable=False,
|
||||||
|
stage="extract",
|
||||||
|
details={"url": str(normalized.item.url), "upstream": normalized.item.metadata.get("upstream")},
|
||||||
|
)
|
||||||
inline_content = fetch_html(str(normalized.item.url))
|
inline_content = fetch_html(str(normalized.item.url))
|
||||||
content_source = "fetched_html"
|
content_source = "fetched_html"
|
||||||
|
|
||||||
if normalized.item.title is None and (
|
if normalized.item.title is None and (
|
||||||
content_source in {"raw_html", "fetched_html"} or inline_content.lstrip().startswith("<")
|
content_source in {"raw_html", "fetched_html", "item.raw_content", "item.raw_summary", "rss_content"}
|
||||||
|
or inline_content.lstrip().startswith("<")
|
||||||
):
|
):
|
||||||
normalized.item.title = extract_title(inline_content)
|
normalized.item.title = extract_title(inline_content)
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,273 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import re
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
import httpx
|
||||||
|
|
||||||
|
from summary_mcp.validators.llm_result import ValidationReport
|
||||||
|
from summary_mcp.validators.llm_result import validate_llm_result as validate_llm_result_from_path
|
||||||
|
from summary_mcp.validators.llm_result import validate_llm_result_payload
|
||||||
|
|
||||||
|
|
||||||
|
JSON_BLOCK_RE = re.compile(r"```(?:json)?\s*(\{.*\})\s*```", re.DOTALL)
|
||||||
|
DEFAULT_CHAT_COMPLETIONS_URL = "https://api.openai.com/v1/chat/completions"
|
||||||
|
|
||||||
|
|
||||||
|
def load_text(path: Path) -> str:
|
||||||
|
return path.read_text(encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
|
def load_json(path: Path) -> dict[str, Any]:
|
||||||
|
return json.loads(load_text(path))
|
||||||
|
|
||||||
|
|
||||||
|
def save_json(path: Path, payload: dict[str, Any]) -> None:
|
||||||
|
path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
path.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
|
def build_summary_input(extracted: dict[str, Any]) -> dict[str, Any]:
|
||||||
|
article = extracted.get("article") or {}
|
||||||
|
return {
|
||||||
|
"article": {
|
||||||
|
"title": article.get("title"),
|
||||||
|
"url": article.get("url"),
|
||||||
|
"plain_text": article.get("plain_text"),
|
||||||
|
"quality_flags": article.get("quality_flags"),
|
||||||
|
},
|
||||||
|
"warnings": extracted.get("warnings", []),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def build_initial_prompt(prompt_template: str, extracted: dict[str, Any]) -> str:
|
||||||
|
summary_input = build_summary_input(extracted)
|
||||||
|
return (
|
||||||
|
f"{prompt_template}\n\n"
|
||||||
|
"Below is the structured extracted article input. Generate the final summary JSON from it.\n\n"
|
||||||
|
f"{json.dumps(summary_input, ensure_ascii=False, indent=2)}"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def build_repair_prompt(
|
||||||
|
errors: list[str],
|
||||||
|
extracted: dict[str, Any],
|
||||||
|
result_json: dict[str, Any],
|
||||||
|
) -> str:
|
||||||
|
summary_input = build_summary_input(extracted)
|
||||||
|
return (
|
||||||
|
"Please repair the following invalid summary JSON.\n\n"
|
||||||
|
"Requirements:\n"
|
||||||
|
"- Output valid JSON only\n"
|
||||||
|
"- Keep fields that are already correct\n"
|
||||||
|
"- Fix only the validator-reported errors\n"
|
||||||
|
"- Do not add explanations\n\n"
|
||||||
|
f"validator errors:\n{json.dumps(errors, ensure_ascii=False, indent=2)}\n\n"
|
||||||
|
f"Extracted article input:\n{json.dumps(summary_input, ensure_ascii=False, indent=2)}\n\n"
|
||||||
|
f"Current summary JSON:\n{json.dumps(result_json, ensure_ascii=False, indent=2)}\n"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def extract_json_text(raw_text: str) -> str:
|
||||||
|
fenced = JSON_BLOCK_RE.search(raw_text)
|
||||||
|
if fenced:
|
||||||
|
return fenced.group(1)
|
||||||
|
|
||||||
|
stripped = raw_text.strip()
|
||||||
|
start = stripped.find("{")
|
||||||
|
end = stripped.rfind("}")
|
||||||
|
if start == -1 or end == -1 or end <= start:
|
||||||
|
raise ValueError("Model output does not contain a JSON object.")
|
||||||
|
return stripped[start : end + 1]
|
||||||
|
|
||||||
|
|
||||||
|
def normalize_chat_completions_url(api_url: str | None) -> str | None:
|
||||||
|
if api_url is None:
|
||||||
|
return None
|
||||||
|
|
||||||
|
normalized = api_url.strip().rstrip("/")
|
||||||
|
if not normalized:
|
||||||
|
return None
|
||||||
|
if normalized.endswith("/chat/completions"):
|
||||||
|
return normalized
|
||||||
|
return f"{normalized}/chat/completions"
|
||||||
|
|
||||||
|
|
||||||
|
def resolve_llm_settings(
|
||||||
|
*,
|
||||||
|
api_key: str | None = None,
|
||||||
|
model: str | None = None,
|
||||||
|
api_url: str | None = None,
|
||||||
|
) -> tuple[str, str, str]:
|
||||||
|
resolved_api_key = api_key or os.environ.get("LLM_API_KEY") or os.environ.get("OPENAI_API_KEY")
|
||||||
|
if not resolved_api_key:
|
||||||
|
raise RuntimeError("Missing LLM_API_KEY or OPENAI_API_KEY, or pass an API key.")
|
||||||
|
|
||||||
|
resolved_model = model or os.environ.get("LLM_MODEL") or os.environ.get("OPENAI_MODEL")
|
||||||
|
if not resolved_model:
|
||||||
|
raise RuntimeError("Missing LLM_MODEL or OPENAI_MODEL, or pass a model.")
|
||||||
|
|
||||||
|
resolved_api_url = normalize_chat_completions_url(
|
||||||
|
api_url or os.environ.get("LLM_API_URL") or os.environ.get("OPENAI_API_URL") or DEFAULT_CHAT_COMPLETIONS_URL
|
||||||
|
)
|
||||||
|
if not resolved_api_url:
|
||||||
|
raise RuntimeError("Missing LLM_API_URL, OPENAI_API_URL, or pass an API URL.")
|
||||||
|
|
||||||
|
return resolved_api_key, resolved_model, resolved_api_url
|
||||||
|
|
||||||
|
|
||||||
|
def call_llm(
|
||||||
|
prompt: str,
|
||||||
|
timeout_seconds: float,
|
||||||
|
api_key: str | None,
|
||||||
|
model: str | None,
|
||||||
|
api_url: str | None,
|
||||||
|
) -> str:
|
||||||
|
resolved_api_key, resolved_model, resolved_api_url = resolve_llm_settings(
|
||||||
|
api_key=api_key,
|
||||||
|
model=model,
|
||||||
|
api_url=api_url,
|
||||||
|
)
|
||||||
|
|
||||||
|
headers = {
|
||||||
|
"Authorization": f"Bearer {resolved_api_key}",
|
||||||
|
"Content-Type": "application/json",
|
||||||
|
}
|
||||||
|
payload = {
|
||||||
|
"model": resolved_model,
|
||||||
|
"messages": [
|
||||||
|
{
|
||||||
|
"role": "system",
|
||||||
|
"content": "You are a precise JSON generator. Always output a single valid JSON object.",
|
||||||
|
},
|
||||||
|
{"role": "user", "content": prompt},
|
||||||
|
],
|
||||||
|
"temperature": 0.2,
|
||||||
|
}
|
||||||
|
|
||||||
|
with httpx.Client(timeout=timeout_seconds) as client:
|
||||||
|
response = client.post(resolved_api_url, headers=headers, json=payload)
|
||||||
|
response.raise_for_status()
|
||||||
|
data = response.json()
|
||||||
|
|
||||||
|
try:
|
||||||
|
return data["choices"][0]["message"]["content"]
|
||||||
|
except (KeyError, IndexError, TypeError) as exc:
|
||||||
|
raise RuntimeError(f"Unexpected LLM response shape: {json.dumps(data, ensure_ascii=False)[:1000]}") from exc
|
||||||
|
|
||||||
|
|
||||||
|
def _save_attempt_artifact(base_output_path: Path | None, suffix: str, payload: str | dict[str, Any]) -> None:
|
||||||
|
if base_output_path is None:
|
||||||
|
return
|
||||||
|
|
||||||
|
path = base_output_path.with_name(f"{base_output_path.stem}.{suffix}")
|
||||||
|
if isinstance(payload, str):
|
||||||
|
path.write_text(payload, encoding="utf-8")
|
||||||
|
else:
|
||||||
|
save_json(path, payload)
|
||||||
|
|
||||||
|
|
||||||
|
def run_loop_payload(
|
||||||
|
*,
|
||||||
|
extracted_payload: dict[str, Any],
|
||||||
|
prompt_path: Path,
|
||||||
|
max_retries: int,
|
||||||
|
timeout_seconds: float,
|
||||||
|
api_key: str | None,
|
||||||
|
model: str | None,
|
||||||
|
api_url: str | None,
|
||||||
|
output_path: Path | None = None,
|
||||||
|
) -> tuple[int, dict[str, Any] | None, ValidationReport | None]:
|
||||||
|
prompt_template = load_text(prompt_path)
|
||||||
|
if output_path is not None:
|
||||||
|
output_path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
|
||||||
|
last_errors: list[str] = []
|
||||||
|
last_result: dict[str, Any] | None = None
|
||||||
|
|
||||||
|
for attempt in range(1, max_retries + 2):
|
||||||
|
if attempt == 1:
|
||||||
|
prompt = build_initial_prompt(prompt_template, extracted_payload)
|
||||||
|
else:
|
||||||
|
assert last_result is not None
|
||||||
|
prompt = build_repair_prompt(last_errors, extracted_payload, last_result)
|
||||||
|
|
||||||
|
raw_output = call_llm(prompt, timeout_seconds, api_key, model, api_url)
|
||||||
|
_save_attempt_artifact(output_path, f"attempt-{attempt}.raw.txt", raw_output)
|
||||||
|
|
||||||
|
try:
|
||||||
|
result_payload = json.loads(extract_json_text(raw_output))
|
||||||
|
except (json.JSONDecodeError, ValueError) as exc:
|
||||||
|
last_errors = [f"Model output is not valid JSON: {exc}"]
|
||||||
|
last_result = {"raw_output": raw_output}
|
||||||
|
report_payload = {
|
||||||
|
"valid": False,
|
||||||
|
"errors": last_errors,
|
||||||
|
"warnings": [],
|
||||||
|
"normalized_result": None,
|
||||||
|
}
|
||||||
|
_save_attempt_artifact(output_path, f"attempt-{attempt}.validation.json", report_payload)
|
||||||
|
if attempt > max_retries:
|
||||||
|
if output_path is not None:
|
||||||
|
output_path.write_text(raw_output, encoding="utf-8")
|
||||||
|
return 1, None, ValidationReport(valid=False, errors=last_errors)
|
||||||
|
continue
|
||||||
|
|
||||||
|
_save_attempt_artifact(output_path, f"attempt-{attempt}.json", result_payload)
|
||||||
|
if output_path is not None:
|
||||||
|
save_json(output_path, result_payload)
|
||||||
|
|
||||||
|
report = validate_llm_result_payload(result_payload, extracted_payload)
|
||||||
|
_save_attempt_artifact(
|
||||||
|
output_path,
|
||||||
|
f"attempt-{attempt}.validation.json",
|
||||||
|
{
|
||||||
|
"valid": report.valid,
|
||||||
|
"errors": report.errors,
|
||||||
|
"warnings": report.warnings,
|
||||||
|
"normalized_result": report.normalized_result,
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
if report.valid:
|
||||||
|
return 0, result_payload, report
|
||||||
|
|
||||||
|
last_errors = report.errors
|
||||||
|
last_result = result_payload
|
||||||
|
|
||||||
|
return 1, None, ValidationReport(valid=False, errors=last_errors)
|
||||||
|
|
||||||
|
|
||||||
|
def run_loop(
|
||||||
|
extracted_path: Path,
|
||||||
|
prompt_path: Path,
|
||||||
|
output_path: Path,
|
||||||
|
max_retries: int,
|
||||||
|
timeout_seconds: float,
|
||||||
|
api_key: str | None,
|
||||||
|
model: str | None,
|
||||||
|
api_url: str | None,
|
||||||
|
) -> int:
|
||||||
|
extracted = load_json(extracted_path)
|
||||||
|
exit_code, _, report = run_loop_payload(
|
||||||
|
extracted_payload=extracted,
|
||||||
|
prompt_path=prompt_path,
|
||||||
|
output_path=output_path,
|
||||||
|
max_retries=max_retries,
|
||||||
|
timeout_seconds=timeout_seconds,
|
||||||
|
api_key=api_key,
|
||||||
|
model=model,
|
||||||
|
api_url=api_url,
|
||||||
|
)
|
||||||
|
|
||||||
|
if exit_code == 0:
|
||||||
|
return 0
|
||||||
|
|
||||||
|
if report is not None and output_path.exists() and extracted_path.exists():
|
||||||
|
fallback_report = validate_llm_result_from_path(output_path, extracted_path)
|
||||||
|
if fallback_report.valid:
|
||||||
|
return 0
|
||||||
|
return exit_code
|
||||||
@@ -0,0 +1,3 @@
|
|||||||
|
from .engine import DEFAULT_RULES_PATH, evaluate_filter_rules, load_filter_rules
|
||||||
|
|
||||||
|
__all__ = ["DEFAULT_RULES_PATH", "evaluate_filter_rules", "load_filter_rules"]
|
||||||
@@ -0,0 +1,136 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from summary_mcp.models.filtering import (
|
||||||
|
FieldCondition,
|
||||||
|
FilterDecisionResult,
|
||||||
|
FilterInput,
|
||||||
|
FilterRule,
|
||||||
|
MatchedRule,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
REPO_ROOT = Path(__file__).resolve().parents[3]
|
||||||
|
DEFAULT_RULES_PATH = REPO_ROOT / "configs" / "filter_rules.json"
|
||||||
|
|
||||||
|
|
||||||
|
def load_filter_rules(path: Path | None = None) -> list[FilterRule]:
|
||||||
|
rules_path = path or DEFAULT_RULES_PATH
|
||||||
|
payload = json.loads(rules_path.read_text(encoding="utf-8-sig"))
|
||||||
|
if not isinstance(payload, list):
|
||||||
|
raise RuntimeError("Filter rules file must contain a JSON array.")
|
||||||
|
return [FilterRule.model_validate(item) for item in payload]
|
||||||
|
|
||||||
|
|
||||||
|
def _normalize_value(value: Any) -> Any:
|
||||||
|
if hasattr(value, "model_dump"):
|
||||||
|
return value.model_dump(mode="json")
|
||||||
|
return value
|
||||||
|
|
||||||
|
|
||||||
|
def _resolve_field(filter_input: FilterInput, field_path: str) -> Any:
|
||||||
|
current: Any = filter_input
|
||||||
|
for part in field_path.split("."):
|
||||||
|
current = _normalize_value(current)
|
||||||
|
if isinstance(current, dict):
|
||||||
|
if part not in current:
|
||||||
|
return None
|
||||||
|
current = current[part]
|
||||||
|
continue
|
||||||
|
return None
|
||||||
|
return _normalize_value(current)
|
||||||
|
|
||||||
|
|
||||||
|
def _expected_value(filter_input: FilterInput, value: Any) -> Any:
|
||||||
|
if isinstance(value, dict) and "from_field" in value:
|
||||||
|
field_name = value.get("from_field")
|
||||||
|
if isinstance(field_name, str):
|
||||||
|
return _resolve_field(filter_input, field_name)
|
||||||
|
return value
|
||||||
|
|
||||||
|
|
||||||
|
def _match_condition(filter_input: FilterInput, condition: FieldCondition) -> bool:
|
||||||
|
current = _resolve_field(filter_input, condition.field)
|
||||||
|
expected = _expected_value(filter_input, condition.value)
|
||||||
|
|
||||||
|
if condition.op == "exists":
|
||||||
|
return (current is not None) if expected is not False else (current is None)
|
||||||
|
if condition.op == "eq":
|
||||||
|
return current == expected
|
||||||
|
if condition.op == "ne":
|
||||||
|
return current != expected
|
||||||
|
if condition.op == "in":
|
||||||
|
return current in expected if isinstance(expected, list) else False
|
||||||
|
if condition.op == "not_in":
|
||||||
|
return current not in expected if isinstance(expected, list) else False
|
||||||
|
if condition.op == "contains":
|
||||||
|
if isinstance(current, list):
|
||||||
|
return expected in current
|
||||||
|
if isinstance(current, str) and isinstance(expected, str):
|
||||||
|
return expected in current
|
||||||
|
return False
|
||||||
|
if condition.op == "overlap":
|
||||||
|
if isinstance(current, list) and isinstance(expected, list):
|
||||||
|
return bool(set(current) & set(expected))
|
||||||
|
return False
|
||||||
|
if condition.op == "gte":
|
||||||
|
return current is not None and expected is not None and current >= expected
|
||||||
|
if condition.op == "lte":
|
||||||
|
return current is not None and expected is not None and current <= expected
|
||||||
|
return False
|
||||||
|
|
||||||
|
|
||||||
|
def _rule_matches(filter_input: FilterInput, rule: FilterRule) -> bool:
|
||||||
|
if not rule.enabled:
|
||||||
|
return False
|
||||||
|
if rule.conditions_all and not all(_match_condition(filter_input, condition) for condition in rule.conditions_all):
|
||||||
|
return False
|
||||||
|
if rule.conditions_any and not any(_match_condition(filter_input, condition) for condition in rule.conditions_any):
|
||||||
|
return False
|
||||||
|
return bool(rule.conditions_all or rule.conditions_any)
|
||||||
|
|
||||||
|
|
||||||
|
def evaluate_filter_rules(filter_input: FilterInput, rules: list[FilterRule]) -> FilterDecisionResult:
|
||||||
|
matched: list[MatchedRule] = []
|
||||||
|
|
||||||
|
for rule in sorted(rules, key=lambda item: item.action.priority, reverse=True):
|
||||||
|
if not _rule_matches(filter_input, rule):
|
||||||
|
continue
|
||||||
|
|
||||||
|
matched_rule = MatchedRule(
|
||||||
|
rule_id=rule.rule_id,
|
||||||
|
decision=rule.action.decision,
|
||||||
|
reason=rule.action.reason,
|
||||||
|
labels=rule.action.labels,
|
||||||
|
priority=rule.action.priority,
|
||||||
|
)
|
||||||
|
matched.append(matched_rule)
|
||||||
|
if rule.stop_on_match:
|
||||||
|
break
|
||||||
|
|
||||||
|
if any(rule.decision == "drop" for rule in matched):
|
||||||
|
final_decision = "drop"
|
||||||
|
elif any(rule.decision == "keep" for rule in matched):
|
||||||
|
final_decision = "keep"
|
||||||
|
elif any(rule.decision == "review" for rule in matched):
|
||||||
|
final_decision = "review"
|
||||||
|
else:
|
||||||
|
final_decision = "review"
|
||||||
|
|
||||||
|
labels = sorted({label for rule in matched for label in rule.labels})
|
||||||
|
reasons = [rule.reason for rule in matched]
|
||||||
|
priorities = [rule.priority for rule in matched]
|
||||||
|
|
||||||
|
return FilterDecisionResult(
|
||||||
|
decision=final_decision,
|
||||||
|
matched_rules=[rule.rule_id for rule in matched],
|
||||||
|
reasons=reasons if reasons else ["No rule matched; defaulted to review."],
|
||||||
|
labels=labels,
|
||||||
|
priority=max(priorities, default=0),
|
||||||
|
matches=matched,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
"""Integration helpers for upstream content sources."""
|
||||||
@@ -0,0 +1,292 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import hashlib
|
||||||
|
from datetime import UTC, datetime
|
||||||
|
from typing import Any
|
||||||
|
from urllib.parse import urlencode, urlparse
|
||||||
|
|
||||||
|
import httpx
|
||||||
|
|
||||||
|
from summary_mcp.models.item import Item
|
||||||
|
|
||||||
|
|
||||||
|
READ_TAG = "user/-/state/com.google/read"
|
||||||
|
KEPT_UNREAD_TAG = "user/-/state/com.google/kept-unread"
|
||||||
|
MIN_INLINE_CONTENT_LENGTH = 500
|
||||||
|
|
||||||
|
|
||||||
|
def _trim_api_base_url(api_base_url: str) -> str:
|
||||||
|
return api_base_url.rstrip("/")
|
||||||
|
|
||||||
|
|
||||||
|
def _build_item_id(source_id: str, external_id: str | None, url: str, published_at: datetime | None) -> str:
|
||||||
|
seed = "|".join(
|
||||||
|
[
|
||||||
|
source_id,
|
||||||
|
external_id or "",
|
||||||
|
url,
|
||||||
|
published_at.isoformat() if published_at else "",
|
||||||
|
]
|
||||||
|
)
|
||||||
|
return f"sha256:{hashlib.sha256(seed.encode('utf-8')).hexdigest()}"
|
||||||
|
|
||||||
|
|
||||||
|
def _build_source_id(entry: dict[str, Any], url: str) -> str:
|
||||||
|
origin = entry.get("origin") or {}
|
||||||
|
stream_id = origin.get("streamId")
|
||||||
|
if isinstance(stream_id, str) and stream_id.strip():
|
||||||
|
digest = hashlib.sha256(stream_id.encode("utf-8")).hexdigest()[:16]
|
||||||
|
return f"freshrss:{digest}"
|
||||||
|
|
||||||
|
host = urlparse(url).netloc or "unknown-source"
|
||||||
|
return f"freshrss:{host}"
|
||||||
|
|
||||||
|
|
||||||
|
def _pick_entry_url(entry: dict[str, Any]) -> str:
|
||||||
|
candidates = [
|
||||||
|
entry.get("canonical"),
|
||||||
|
entry.get("alternate"),
|
||||||
|
]
|
||||||
|
|
||||||
|
for candidate_list in candidates:
|
||||||
|
if not isinstance(candidate_list, list):
|
||||||
|
continue
|
||||||
|
for candidate in candidate_list:
|
||||||
|
href = (candidate or {}).get("href")
|
||||||
|
if isinstance(href, str) and href.strip():
|
||||||
|
return href.strip()
|
||||||
|
|
||||||
|
entry_id = entry.get("id")
|
||||||
|
if isinstance(entry_id, str) and entry_id.startswith("tag:"):
|
||||||
|
return entry_id
|
||||||
|
|
||||||
|
raise ValueError("FreshRSS entry does not contain a usable URL.")
|
||||||
|
|
||||||
|
|
||||||
|
def _pick_content_block(entry: dict[str, Any], key: str) -> str | None:
|
||||||
|
block = entry.get(key)
|
||||||
|
if isinstance(block, dict):
|
||||||
|
content = block.get("content")
|
||||||
|
if isinstance(content, str) and content.strip():
|
||||||
|
return content
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def _pick_categories(entry: dict[str, Any]) -> list[str]:
|
||||||
|
categories = entry.get("categories")
|
||||||
|
if not isinstance(categories, list):
|
||||||
|
return []
|
||||||
|
|
||||||
|
values: list[str] = []
|
||||||
|
for category in categories:
|
||||||
|
if not isinstance(category, str):
|
||||||
|
continue
|
||||||
|
if category.startswith("user/-/label/"):
|
||||||
|
values.append(category.removeprefix("user/-/label/"))
|
||||||
|
elif category.startswith("user/-/state/com.google/"):
|
||||||
|
continue
|
||||||
|
else:
|
||||||
|
values.append(category)
|
||||||
|
return values
|
||||||
|
|
||||||
|
|
||||||
|
def _parse_datetime(timestamp: Any) -> datetime | None:
|
||||||
|
if timestamp is None:
|
||||||
|
return None
|
||||||
|
if isinstance(timestamp, (int, float)):
|
||||||
|
if timestamp > 10_000_000_000:
|
||||||
|
return datetime.fromtimestamp(timestamp / 1000, tz=UTC)
|
||||||
|
return datetime.fromtimestamp(timestamp, tz=UTC)
|
||||||
|
if isinstance(timestamp, str) and timestamp.isdigit():
|
||||||
|
value = int(timestamp)
|
||||||
|
if value > 10_000_000_000:
|
||||||
|
return datetime.fromtimestamp(value / 1000, tz=UTC)
|
||||||
|
return datetime.fromtimestamp(value, tz=UTC)
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def _dedupe_non_empty(values: list[str] | None) -> list[str]:
|
||||||
|
if not values:
|
||||||
|
return []
|
||||||
|
|
||||||
|
seen: set[str] = set()
|
||||||
|
deduped: list[str] = []
|
||||||
|
for value in values:
|
||||||
|
if not value or value in seen:
|
||||||
|
continue
|
||||||
|
seen.add(value)
|
||||||
|
deduped.append(value)
|
||||||
|
return deduped
|
||||||
|
|
||||||
|
|
||||||
|
def _has_usable_inline_content(value: str | None) -> bool:
|
||||||
|
return bool(value and len(value.strip()) >= MIN_INLINE_CONTENT_LENGTH)
|
||||||
|
|
||||||
|
|
||||||
|
def map_entry_to_item(entry: dict[str, Any]) -> Item:
|
||||||
|
url = _pick_entry_url(entry)
|
||||||
|
summary_html = _pick_content_block(entry, "summary")
|
||||||
|
content_html = _pick_content_block(entry, "content")
|
||||||
|
published_at = _parse_datetime(entry.get("published"))
|
||||||
|
source_id = _build_source_id(entry, url)
|
||||||
|
external_id = entry.get("id") if isinstance(entry.get("id"), str) else None
|
||||||
|
fetch_state = "fetched" if _has_usable_inline_content(content_html) or _has_usable_inline_content(summary_html) else "pending"
|
||||||
|
|
||||||
|
metadata = {
|
||||||
|
"upstream": "freshrss",
|
||||||
|
"origin": entry.get("origin") or {},
|
||||||
|
"categories": _pick_categories(entry),
|
||||||
|
"crawled_at": _parse_datetime(entry.get("crawlTimeMsec")),
|
||||||
|
"published_epoch": entry.get("published"),
|
||||||
|
}
|
||||||
|
|
||||||
|
return Item(
|
||||||
|
item_id=_build_item_id(source_id, external_id, url, published_at),
|
||||||
|
source_id=source_id,
|
||||||
|
external_id=external_id,
|
||||||
|
title=entry.get("title"),
|
||||||
|
url=url,
|
||||||
|
author=entry.get("author"),
|
||||||
|
published_at=published_at,
|
||||||
|
discovered_at=datetime.now(tz=UTC),
|
||||||
|
content_kind="article",
|
||||||
|
language=None,
|
||||||
|
raw_summary=summary_html,
|
||||||
|
raw_content=content_html,
|
||||||
|
metadata=metadata,
|
||||||
|
fetch_state=fetch_state,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
class FreshRSSClient:
|
||||||
|
def __init__(
|
||||||
|
self,
|
||||||
|
api_base_url: str,
|
||||||
|
username: str,
|
||||||
|
api_password: str,
|
||||||
|
timeout_seconds: float = 20.0,
|
||||||
|
) -> None:
|
||||||
|
self.api_base_url = _trim_api_base_url(api_base_url)
|
||||||
|
self.username = username
|
||||||
|
self.api_password = api_password
|
||||||
|
self.timeout_seconds = timeout_seconds
|
||||||
|
|
||||||
|
def _build_url(self, path: str) -> str:
|
||||||
|
return f"{self.api_base_url}/{path.lstrip('/')}"
|
||||||
|
|
||||||
|
def _build_auth_headers(self, auth_token: str) -> dict[str, str]:
|
||||||
|
return {"Authorization": f"GoogleLogin auth={auth_token}"}
|
||||||
|
|
||||||
|
def client_login(self) -> str:
|
||||||
|
data = {
|
||||||
|
"Email": self.username,
|
||||||
|
"Passwd": self.api_password,
|
||||||
|
}
|
||||||
|
with httpx.Client(timeout=self.timeout_seconds) as client:
|
||||||
|
response = client.post(self._build_url("accounts/ClientLogin"), data=data)
|
||||||
|
response.raise_for_status()
|
||||||
|
|
||||||
|
auth_token: str | None = None
|
||||||
|
for line in response.text.splitlines():
|
||||||
|
if line.startswith("Auth="):
|
||||||
|
auth_token = line.split("=", 1)[1].strip()
|
||||||
|
break
|
||||||
|
|
||||||
|
if not auth_token:
|
||||||
|
raise RuntimeError("FreshRSS ClientLogin succeeded but did not return an Auth token.")
|
||||||
|
|
||||||
|
return auth_token
|
||||||
|
|
||||||
|
def get_edit_token(self, auth_token: str) -> str:
|
||||||
|
with httpx.Client(
|
||||||
|
timeout=self.timeout_seconds,
|
||||||
|
headers=self._build_auth_headers(auth_token),
|
||||||
|
) as client:
|
||||||
|
response = client.get(self._build_url("reader/api/0/token"))
|
||||||
|
response.raise_for_status()
|
||||||
|
|
||||||
|
edit_token = response.text.strip()
|
||||||
|
if not edit_token:
|
||||||
|
raise RuntimeError("FreshRSS token request succeeded but did not return an edit token.")
|
||||||
|
return edit_token
|
||||||
|
|
||||||
|
def fetch_stream_contents(
|
||||||
|
self,
|
||||||
|
auth_token: str,
|
||||||
|
stream_id: str = "user/-/state/com.google/reading-list",
|
||||||
|
limit: int = 20,
|
||||||
|
continuation: str | None = None,
|
||||||
|
exclude_targets: list[str] | None = None,
|
||||||
|
include_targets: list[str] | None = None,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
params: list[tuple[str, str | int]] = [
|
||||||
|
("output", "json"),
|
||||||
|
("n", limit),
|
||||||
|
]
|
||||||
|
if continuation:
|
||||||
|
params.append(("c", continuation))
|
||||||
|
for target in _dedupe_non_empty(exclude_targets):
|
||||||
|
params.append(("xt", target))
|
||||||
|
for target in _dedupe_non_empty(include_targets):
|
||||||
|
params.append(("it", target))
|
||||||
|
|
||||||
|
with httpx.Client(
|
||||||
|
timeout=self.timeout_seconds,
|
||||||
|
headers=self._build_auth_headers(auth_token),
|
||||||
|
) as client:
|
||||||
|
response = client.get(
|
||||||
|
self._build_url(f"reader/api/0/stream/contents/{stream_id}"),
|
||||||
|
params=params,
|
||||||
|
)
|
||||||
|
response.raise_for_status()
|
||||||
|
return response.json()
|
||||||
|
|
||||||
|
def edit_tag(
|
||||||
|
self,
|
||||||
|
auth_token: str,
|
||||||
|
edit_token: str,
|
||||||
|
item_ids: list[str],
|
||||||
|
add_tags: list[str] | None = None,
|
||||||
|
remove_tags: list[str] | None = None,
|
||||||
|
) -> str:
|
||||||
|
normalized_item_ids = _dedupe_non_empty(item_ids)
|
||||||
|
if not normalized_item_ids:
|
||||||
|
raise ValueError("FreshRSS edit-tag requires at least one item id.")
|
||||||
|
|
||||||
|
payload: list[tuple[str, str]] = [("ac", "edit"), ("T", edit_token)]
|
||||||
|
for item_id in normalized_item_ids:
|
||||||
|
payload.append(("i", item_id))
|
||||||
|
for tag in _dedupe_non_empty(add_tags):
|
||||||
|
payload.append(("a", tag))
|
||||||
|
for tag in _dedupe_non_empty(remove_tags):
|
||||||
|
payload.append(("r", tag))
|
||||||
|
|
||||||
|
headers = self._build_auth_headers(auth_token)
|
||||||
|
headers["Content-Type"] = "application/x-www-form-urlencoded"
|
||||||
|
|
||||||
|
with httpx.Client(
|
||||||
|
timeout=self.timeout_seconds,
|
||||||
|
headers=headers,
|
||||||
|
) as client:
|
||||||
|
response = client.post(
|
||||||
|
self._build_url("reader/api/0/edit-tag"),
|
||||||
|
content=urlencode(payload).encode("utf-8"),
|
||||||
|
)
|
||||||
|
response.raise_for_status()
|
||||||
|
|
||||||
|
return response.text.strip()
|
||||||
|
|
||||||
|
def mark_items_as_read(
|
||||||
|
self,
|
||||||
|
auth_token: str,
|
||||||
|
item_ids: list[str],
|
||||||
|
edit_token: str | None = None,
|
||||||
|
) -> str:
|
||||||
|
resolved_edit_token = edit_token or self.get_edit_token(auth_token)
|
||||||
|
return self.edit_tag(
|
||||||
|
auth_token=auth_token,
|
||||||
|
edit_token=resolved_edit_token,
|
||||||
|
item_ids=item_ids,
|
||||||
|
add_tags=[READ_TAG],
|
||||||
|
remove_tags=[KEPT_UNREAD_TAG],
|
||||||
|
)
|
||||||
@@ -1 +1,60 @@
|
|||||||
"""Shared models for the summary MCP service."""
|
"""Shared models for the summary MCP service."""
|
||||||
|
|
||||||
|
from .article_candidate import (
|
||||||
|
ArticleCandidate,
|
||||||
|
ArticleCandidateRecord,
|
||||||
|
CandidateMetadata,
|
||||||
|
CandidateSourceRefs,
|
||||||
|
DigestSectionHint,
|
||||||
|
OpenClawCandidateInput,
|
||||||
|
ReviewState,
|
||||||
|
build_article_candidate_record,
|
||||||
|
build_openclaw_candidate_input,
|
||||||
|
candidate_id_for,
|
||||||
|
normalize_candidate_url,
|
||||||
|
)
|
||||||
|
from .candidate_review import CandidateReviewStatus, KnowledgeDecision, ReviewStatus
|
||||||
|
from .daily_digest import (
|
||||||
|
DailyDigest,
|
||||||
|
DailyDigestItem,
|
||||||
|
DailyDigestMetadata,
|
||||||
|
DailyDigestSection,
|
||||||
|
DailyDigestSourceRef,
|
||||||
|
DailyDigestStats,
|
||||||
|
)
|
||||||
|
from .keyword_index import DailyKeywordIndex, DailyKeywordTerm, KeywordStat, KeywordStatsIndex
|
||||||
|
from .openclaw_delivery import (
|
||||||
|
OpenClawDeliveryPayload,
|
||||||
|
OpenClawDeliveryStats,
|
||||||
|
build_openclaw_delivery_payload,
|
||||||
|
)
|
||||||
|
|
||||||
|
__all__ = [
|
||||||
|
"ArticleCandidate",
|
||||||
|
"ArticleCandidateRecord",
|
||||||
|
"CandidateMetadata",
|
||||||
|
"CandidateReviewStatus",
|
||||||
|
"CandidateSourceRefs",
|
||||||
|
"DailyDigest",
|
||||||
|
"DailyDigestItem",
|
||||||
|
"DailyDigestMetadata",
|
||||||
|
"DailyDigestSection",
|
||||||
|
"DailyDigestSourceRef",
|
||||||
|
"DailyDigestStats",
|
||||||
|
"DailyKeywordIndex",
|
||||||
|
"DailyKeywordTerm",
|
||||||
|
"DigestSectionHint",
|
||||||
|
"KeywordStat",
|
||||||
|
"KeywordStatsIndex",
|
||||||
|
"KnowledgeDecision",
|
||||||
|
"OpenClawCandidateInput",
|
||||||
|
"OpenClawDeliveryPayload",
|
||||||
|
"OpenClawDeliveryStats",
|
||||||
|
"ReviewState",
|
||||||
|
"ReviewStatus",
|
||||||
|
"build_article_candidate_record",
|
||||||
|
"build_openclaw_delivery_payload",
|
||||||
|
"build_openclaw_candidate_input",
|
||||||
|
"candidate_id_for",
|
||||||
|
"normalize_candidate_url",
|
||||||
|
]
|
||||||
|
|||||||
@@ -0,0 +1,190 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from datetime import datetime
|
||||||
|
from typing import Literal
|
||||||
|
from urllib.parse import parse_qsl, urlencode, urlparse, urlunparse
|
||||||
|
|
||||||
|
from pydantic import BaseModel, Field, HttpUrl
|
||||||
|
|
||||||
|
from .document import ExtractedArticle
|
||||||
|
from .filtering import FilterDecision, FilterDecisionResult
|
||||||
|
from .item import Item
|
||||||
|
from .llm_result import Category, LlmSummaryResult
|
||||||
|
|
||||||
|
|
||||||
|
ReviewState = Literal["pending", "accepted", "rejected", "deferred"]
|
||||||
|
DigestSectionHint = Literal[
|
||||||
|
"top_news",
|
||||||
|
"tools_and_workflows",
|
||||||
|
"risk_and_security",
|
||||||
|
"open_source",
|
||||||
|
"insights",
|
||||||
|
"deep_dive",
|
||||||
|
]
|
||||||
|
|
||||||
|
_TRACKING_QUERY_PARAMS = {
|
||||||
|
"fbclid",
|
||||||
|
"gclid",
|
||||||
|
"igshid",
|
||||||
|
"mc_cid",
|
||||||
|
"mc_eid",
|
||||||
|
"mkt_tok",
|
||||||
|
"ref",
|
||||||
|
"ref_src",
|
||||||
|
"si",
|
||||||
|
"spm",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
class CandidateSourceRefs(BaseModel):
|
||||||
|
item_path: str | None = None
|
||||||
|
extracted_path: str | None = None
|
||||||
|
summary_path: str | None = None
|
||||||
|
filter_path: str | None = None
|
||||||
|
|
||||||
|
|
||||||
|
class CandidateMetadata(BaseModel):
|
||||||
|
generated_at: datetime | None = None
|
||||||
|
pipeline_version: str = "v1"
|
||||||
|
producer: str | None = None
|
||||||
|
run_id: str | None = None
|
||||||
|
|
||||||
|
|
||||||
|
class ArticleCandidateRecord(BaseModel):
|
||||||
|
candidate_id: str
|
||||||
|
item: Item | None = None
|
||||||
|
article: ExtractedArticle
|
||||||
|
summary: LlmSummaryResult
|
||||||
|
filter_result: FilterDecisionResult
|
||||||
|
digest_section_hint: DigestSectionHint | None = None
|
||||||
|
digest_rank: int = Field(default=50, ge=0, le=100)
|
||||||
|
review_state: ReviewState = "pending"
|
||||||
|
source_refs: CandidateSourceRefs = Field(default_factory=CandidateSourceRefs)
|
||||||
|
rendered_markdown: str | None = None
|
||||||
|
metadata: CandidateMetadata = Field(default_factory=CandidateMetadata)
|
||||||
|
|
||||||
|
|
||||||
|
class OpenClawCandidateInput(BaseModel):
|
||||||
|
candidate_id: str
|
||||||
|
title: str
|
||||||
|
url: HttpUrl
|
||||||
|
canonical_url: HttpUrl | None = None
|
||||||
|
published_at: datetime | None = None
|
||||||
|
author: str | None = None
|
||||||
|
source_name: str | None = None
|
||||||
|
language: str | None = None
|
||||||
|
summary: str
|
||||||
|
highlights: list[str] = Field(default_factory=list)
|
||||||
|
keywords: list[str] = Field(default_factory=list)
|
||||||
|
topics: list[str] = Field(default_factory=list)
|
||||||
|
category: Category
|
||||||
|
worth_keeping: bool
|
||||||
|
worth_reason: str
|
||||||
|
selection_decision: FilterDecision
|
||||||
|
selection_reason: str | None = None
|
||||||
|
digest_section_hint: DigestSectionHint | None = None
|
||||||
|
digest_rank: int = Field(default=50, ge=0, le=100)
|
||||||
|
|
||||||
|
|
||||||
|
# Backward-compatible alias while the rest of the codebase migrates to the new name.
|
||||||
|
ArticleCandidate = ArticleCandidateRecord
|
||||||
|
|
||||||
|
|
||||||
|
def normalize_candidate_url(url: str) -> str:
|
||||||
|
parsed = urlparse(url)
|
||||||
|
filtered_query = [
|
||||||
|
(key, value)
|
||||||
|
for key, value in parse_qsl(parsed.query, keep_blank_values=True)
|
||||||
|
if key.lower() not in _TRACKING_QUERY_PARAMS and not key.lower().startswith("utm_")
|
||||||
|
]
|
||||||
|
normalized = parsed._replace(query=urlencode(filtered_query, doseq=True), fragment="")
|
||||||
|
return urlunparse(normalized)
|
||||||
|
|
||||||
|
|
||||||
|
def candidate_id_for(item: Item | None, article: ExtractedArticle) -> str:
|
||||||
|
if item is not None and item.item_id:
|
||||||
|
return f"cand:{item.item_id}"
|
||||||
|
if article.item_id:
|
||||||
|
return f"cand:{article.item_id}"
|
||||||
|
return f"cand:{article.extract_id}"
|
||||||
|
|
||||||
|
|
||||||
|
def build_article_candidate_record(
|
||||||
|
*,
|
||||||
|
summary: LlmSummaryResult,
|
||||||
|
article: ExtractedArticle,
|
||||||
|
filter_result: FilterDecisionResult,
|
||||||
|
item: Item | None = None,
|
||||||
|
digest_section_hint: DigestSectionHint | None = None,
|
||||||
|
digest_rank: int | None = None,
|
||||||
|
rendered_markdown: str | None = None,
|
||||||
|
source_refs: CandidateSourceRefs | None = None,
|
||||||
|
metadata: CandidateMetadata | None = None,
|
||||||
|
) -> ArticleCandidateRecord:
|
||||||
|
return ArticleCandidateRecord(
|
||||||
|
candidate_id=candidate_id_for(item, article),
|
||||||
|
item=item,
|
||||||
|
article=article,
|
||||||
|
summary=summary,
|
||||||
|
filter_result=filter_result,
|
||||||
|
digest_section_hint=digest_section_hint,
|
||||||
|
digest_rank=digest_rank if digest_rank is not None else filter_result.priority,
|
||||||
|
source_refs=source_refs or CandidateSourceRefs(),
|
||||||
|
rendered_markdown=rendered_markdown,
|
||||||
|
metadata=metadata or CandidateMetadata(),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _source_name_from_item(item: Item | None, article: ExtractedArticle) -> str | None:
|
||||||
|
if item is not None:
|
||||||
|
metadata = item.metadata if isinstance(item.metadata, dict) else {}
|
||||||
|
origin = metadata.get("origin")
|
||||||
|
if isinstance(origin, dict):
|
||||||
|
origin_title = origin.get("title")
|
||||||
|
if isinstance(origin_title, str) and origin_title.strip():
|
||||||
|
return origin_title.strip()
|
||||||
|
|
||||||
|
for key in ("source_title", "feed_title", "site_name", "source_name"):
|
||||||
|
value = metadata.get(key)
|
||||||
|
if isinstance(value, str) and value.strip():
|
||||||
|
return value.strip()
|
||||||
|
|
||||||
|
candidate_url = str(item.url) if item is not None else str(article.url)
|
||||||
|
hostname = urlparse(candidate_url).netloc
|
||||||
|
return hostname or None
|
||||||
|
|
||||||
|
|
||||||
|
def build_openclaw_candidate_input(record: ArticleCandidateRecord) -> OpenClawCandidateInput:
|
||||||
|
item = record.item
|
||||||
|
article = record.article
|
||||||
|
summary = record.summary
|
||||||
|
filter_result = record.filter_result
|
||||||
|
|
||||||
|
title = item.title if item is not None and item.title else summary.title
|
||||||
|
raw_url = str(item.url) if item is not None else str(summary.url)
|
||||||
|
published_at = item.published_at if item is not None else article.published_at
|
||||||
|
author = item.author if item is not None and item.author else article.author
|
||||||
|
language = item.language if item is not None and item.language else article.language
|
||||||
|
selection_reason = filter_result.reasons[0] if filter_result.reasons else None
|
||||||
|
|
||||||
|
return OpenClawCandidateInput(
|
||||||
|
candidate_id=record.candidate_id,
|
||||||
|
title=title,
|
||||||
|
url=raw_url,
|
||||||
|
canonical_url=normalize_candidate_url(raw_url),
|
||||||
|
published_at=published_at,
|
||||||
|
author=author,
|
||||||
|
source_name=_source_name_from_item(item, article),
|
||||||
|
language=language,
|
||||||
|
summary=summary.summary,
|
||||||
|
highlights=summary.highlights,
|
||||||
|
keywords=summary.keywords,
|
||||||
|
topics=summary.topics,
|
||||||
|
category=summary.category,
|
||||||
|
worth_keeping=summary.worth_keeping,
|
||||||
|
worth_reason=summary.reason,
|
||||||
|
selection_decision=filter_result.decision,
|
||||||
|
selection_reason=selection_reason,
|
||||||
|
digest_section_hint=record.digest_section_hint,
|
||||||
|
digest_rank=record.digest_rank,
|
||||||
|
)
|
||||||
@@ -0,0 +1,19 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from datetime import datetime
|
||||||
|
from typing import Literal
|
||||||
|
|
||||||
|
from pydantic import BaseModel
|
||||||
|
|
||||||
|
|
||||||
|
ReviewStatus = Literal["pending", "confirmed", "rejected", "deferred"]
|
||||||
|
KnowledgeDecision = Literal["pending", "ingest", "skip"]
|
||||||
|
|
||||||
|
|
||||||
|
class CandidateReviewStatus(BaseModel):
|
||||||
|
candidate_id: str
|
||||||
|
review_status: ReviewStatus = "pending"
|
||||||
|
reviewed_at: datetime | None = None
|
||||||
|
review_note: str | None = None
|
||||||
|
knowledge_decision: KnowledgeDecision = "pending"
|
||||||
|
knowledge_note: str | None = None
|
||||||
@@ -0,0 +1,58 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from datetime import date, datetime
|
||||||
|
|
||||||
|
from pydantic import BaseModel, Field, HttpUrl
|
||||||
|
|
||||||
|
from .article_candidate import DigestSectionHint
|
||||||
|
|
||||||
|
|
||||||
|
class DailyDigestItem(BaseModel):
|
||||||
|
candidate_id: str
|
||||||
|
title: str
|
||||||
|
summary: str | None = None
|
||||||
|
why_it_matters: str
|
||||||
|
action: str | None = None
|
||||||
|
|
||||||
|
|
||||||
|
class DailyDigestSection(BaseModel):
|
||||||
|
section_id: DigestSectionHint
|
||||||
|
title: str
|
||||||
|
summary: str
|
||||||
|
items: list[DailyDigestItem] = Field(default_factory=list)
|
||||||
|
|
||||||
|
|
||||||
|
class DailyDigestSourceRef(BaseModel):
|
||||||
|
candidate_id: str
|
||||||
|
url: HttpUrl | None = None
|
||||||
|
title: str | None = None
|
||||||
|
|
||||||
|
|
||||||
|
class DailyDigestStats(BaseModel):
|
||||||
|
candidate_total: int = Field(default=0, ge=0)
|
||||||
|
kept_total: int = Field(default=0, ge=0)
|
||||||
|
review_total: int = Field(default=0, ge=0)
|
||||||
|
dropped_total: int = Field(default=0, ge=0)
|
||||||
|
|
||||||
|
|
||||||
|
class DailyDigestMetadata(BaseModel):
|
||||||
|
generated_at: datetime | None = None
|
||||||
|
producer: str | None = None
|
||||||
|
pipeline_version: str = "v1"
|
||||||
|
run_id: str | None = None
|
||||||
|
|
||||||
|
|
||||||
|
class DailyDigest(BaseModel):
|
||||||
|
digest_id: str
|
||||||
|
date: date
|
||||||
|
title: str
|
||||||
|
summary: str
|
||||||
|
sections: list[DailyDigestSection] = Field(default_factory=list)
|
||||||
|
top_items: list[str] = Field(default_factory=list)
|
||||||
|
key_takeaways: list[str] = Field(default_factory=list)
|
||||||
|
watchlist: list[str] = Field(default_factory=list)
|
||||||
|
candidate_ids: list[str] = Field(default_factory=list)
|
||||||
|
source_refs: list[DailyDigestSourceRef] = Field(default_factory=list)
|
||||||
|
editor_notes: str | None = None
|
||||||
|
stats: DailyDigestStats = Field(default_factory=DailyDigestStats)
|
||||||
|
metadata: DailyDigestMetadata = Field(default_factory=DailyDigestMetadata)
|
||||||
@@ -0,0 +1,67 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from datetime import datetime
|
||||||
|
from typing import Any, Literal
|
||||||
|
|
||||||
|
from pydantic import BaseModel, Field
|
||||||
|
|
||||||
|
from .document import ExtractedArticle
|
||||||
|
from .item import Item
|
||||||
|
from .llm_result import LlmSummaryResult
|
||||||
|
|
||||||
|
|
||||||
|
FilterDecision = Literal["keep", "drop", "review"]
|
||||||
|
ConditionOp = Literal["eq", "ne", "in", "not_in", "contains", "overlap", "gte", "lte", "exists"]
|
||||||
|
|
||||||
|
|
||||||
|
class FilterContext(BaseModel):
|
||||||
|
source_tags: list[str] = Field(default_factory=list)
|
||||||
|
interest_topics: list[str] = Field(default_factory=list)
|
||||||
|
interest_keywords: list[str] = Field(default_factory=list)
|
||||||
|
now: datetime | None = None
|
||||||
|
|
||||||
|
|
||||||
|
class FieldCondition(BaseModel):
|
||||||
|
field: str
|
||||||
|
op: ConditionOp
|
||||||
|
value: Any = None
|
||||||
|
|
||||||
|
|
||||||
|
class FilterAction(BaseModel):
|
||||||
|
decision: FilterDecision
|
||||||
|
reason: str
|
||||||
|
labels: list[str] = Field(default_factory=list)
|
||||||
|
priority: int = 50
|
||||||
|
|
||||||
|
|
||||||
|
class FilterRule(BaseModel):
|
||||||
|
rule_id: str
|
||||||
|
enabled: bool = True
|
||||||
|
stop_on_match: bool = False
|
||||||
|
conditions_all: list[FieldCondition] = Field(default_factory=list)
|
||||||
|
conditions_any: list[FieldCondition] = Field(default_factory=list)
|
||||||
|
action: FilterAction
|
||||||
|
|
||||||
|
|
||||||
|
class FilterInput(BaseModel):
|
||||||
|
item: Item | None = None
|
||||||
|
article: ExtractedArticle | None = None
|
||||||
|
summary: LlmSummaryResult
|
||||||
|
context: FilterContext = Field(default_factory=FilterContext)
|
||||||
|
|
||||||
|
|
||||||
|
class MatchedRule(BaseModel):
|
||||||
|
rule_id: str
|
||||||
|
decision: FilterDecision
|
||||||
|
reason: str
|
||||||
|
labels: list[str] = Field(default_factory=list)
|
||||||
|
priority: int = 0
|
||||||
|
|
||||||
|
|
||||||
|
class FilterDecisionResult(BaseModel):
|
||||||
|
decision: FilterDecision
|
||||||
|
matched_rules: list[str] = Field(default_factory=list)
|
||||||
|
reasons: list[str] = Field(default_factory=list)
|
||||||
|
labels: list[str] = Field(default_factory=list)
|
||||||
|
priority: int = 0
|
||||||
|
matches: list[MatchedRule] = Field(default_factory=list)
|
||||||
@@ -0,0 +1,35 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from datetime import date, datetime
|
||||||
|
|
||||||
|
from pydantic import BaseModel, Field
|
||||||
|
|
||||||
|
|
||||||
|
class DailyKeywordTerm(BaseModel):
|
||||||
|
term: str
|
||||||
|
normalized_term: str
|
||||||
|
count: int = Field(ge=1)
|
||||||
|
|
||||||
|
|
||||||
|
class DailyKeywordIndex(BaseModel):
|
||||||
|
schema_version: str = "v1"
|
||||||
|
date: date
|
||||||
|
source: str
|
||||||
|
digest_id: str
|
||||||
|
generated_at: datetime
|
||||||
|
candidate_count: int = Field(default=0, ge=0)
|
||||||
|
terms: list[DailyKeywordTerm] = Field(default_factory=list)
|
||||||
|
|
||||||
|
|
||||||
|
class KeywordStat(BaseModel):
|
||||||
|
term: str
|
||||||
|
total_count: int = Field(default=0, ge=0)
|
||||||
|
days_seen: int = Field(default=0, ge=0)
|
||||||
|
first_seen: date
|
||||||
|
last_seen: date
|
||||||
|
|
||||||
|
|
||||||
|
class KeywordStatsIndex(BaseModel):
|
||||||
|
schema_version: str = "v1"
|
||||||
|
generated_at: datetime
|
||||||
|
terms: list[KeywordStat] = Field(default_factory=list)
|
||||||
@@ -0,0 +1,49 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from datetime import UTC, date, datetime
|
||||||
|
|
||||||
|
from pydantic import BaseModel, Field
|
||||||
|
|
||||||
|
from .article_candidate import OpenClawCandidateInput
|
||||||
|
|
||||||
|
|
||||||
|
class OpenClawDeliveryStats(BaseModel):
|
||||||
|
total: int = Field(default=0, ge=0)
|
||||||
|
keep_total: int = Field(default=0, ge=0)
|
||||||
|
review_total: int = Field(default=0, ge=0)
|
||||||
|
drop_total: int = Field(default=0, ge=0)
|
||||||
|
|
||||||
|
|
||||||
|
class OpenClawDeliveryPayload(BaseModel):
|
||||||
|
schema_version: str = "v1"
|
||||||
|
generated_at: datetime
|
||||||
|
run_id: str
|
||||||
|
date: date
|
||||||
|
candidates: list[OpenClawCandidateInput] = Field(default_factory=list)
|
||||||
|
stats: OpenClawDeliveryStats = Field(default_factory=OpenClawDeliveryStats)
|
||||||
|
|
||||||
|
|
||||||
|
def build_openclaw_delivery_payload(
|
||||||
|
candidates: list[OpenClawCandidateInput],
|
||||||
|
*,
|
||||||
|
run_id: str | None = None,
|
||||||
|
for_date: date | None = None,
|
||||||
|
) -> OpenClawDeliveryPayload:
|
||||||
|
now = datetime.now(tz=UTC)
|
||||||
|
payload_date = for_date or now.date()
|
||||||
|
resolved_run_id = run_id or now.strftime("openclaw-delivery-%Y%m%d-%H%M%S")
|
||||||
|
|
||||||
|
stats = OpenClawDeliveryStats(
|
||||||
|
total=len(candidates),
|
||||||
|
keep_total=sum(1 for candidate in candidates if candidate.selection_decision == "keep"),
|
||||||
|
review_total=sum(1 for candidate in candidates if candidate.selection_decision == "review"),
|
||||||
|
drop_total=sum(1 for candidate in candidates if candidate.selection_decision == "drop"),
|
||||||
|
)
|
||||||
|
return OpenClawDeliveryPayload(
|
||||||
|
schema_version="v1",
|
||||||
|
generated_at=now,
|
||||||
|
run_id=resolved_run_id,
|
||||||
|
date=payload_date,
|
||||||
|
candidates=candidates,
|
||||||
|
stats=stats,
|
||||||
|
)
|
||||||
@@ -0,0 +1,35 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from datetime import datetime
|
||||||
|
from typing import Literal
|
||||||
|
|
||||||
|
from pydantic import BaseModel, Field
|
||||||
|
|
||||||
|
from .document import ExtractedArticle
|
||||||
|
from .filtering import FilterDecisionResult
|
||||||
|
from .item import Item
|
||||||
|
from .llm_result import LlmSummaryResult
|
||||||
|
|
||||||
|
|
||||||
|
SinkTarget = Literal["markdown"]
|
||||||
|
|
||||||
|
|
||||||
|
class SinkMetadata(BaseModel):
|
||||||
|
generated_at: datetime | None = None
|
||||||
|
pipeline_version: str = "v1"
|
||||||
|
source: str | None = None
|
||||||
|
|
||||||
|
|
||||||
|
class SinkInput(BaseModel):
|
||||||
|
item: Item | None = None
|
||||||
|
article: ExtractedArticle
|
||||||
|
summary: LlmSummaryResult
|
||||||
|
filter_decision: FilterDecisionResult
|
||||||
|
metadata: SinkMetadata = Field(default_factory=SinkMetadata)
|
||||||
|
|
||||||
|
|
||||||
|
class SinkResult(BaseModel):
|
||||||
|
target: SinkTarget
|
||||||
|
success: bool
|
||||||
|
output_path: str | None = None
|
||||||
|
message: str | None = None
|
||||||
@@ -1,10 +1,20 @@
|
|||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
import tempfile
|
||||||
|
from datetime import date
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
from mcp.server.fastmcp import FastMCP
|
from mcp.server.fastmcp import FastMCP
|
||||||
|
|
||||||
from summary_mcp.core.pipeline import extract_content
|
from summary_mcp.core.pipeline import extract_content
|
||||||
|
from summary_mcp.filters.engine import evaluate_filter_rules, load_filter_rules
|
||||||
|
from summary_mcp.models.document import ExtractedArticle
|
||||||
|
from summary_mcp.models.filtering import FilterContext, FilterInput
|
||||||
from summary_mcp.models.item import Item
|
from summary_mcp.models.item import Item
|
||||||
|
from summary_mcp.models.llm_result import LlmSummaryResult
|
||||||
from summary_mcp.models.summary_io import ExtractionInput
|
from summary_mcp.models.summary_io import ExtractionInput
|
||||||
|
from summary_mcp.workflows import run_freshrss_pipeline
|
||||||
|
|
||||||
|
|
||||||
mcp = FastMCP(name="content-extract-mcp")
|
mcp = FastMCP(name="content-extract-mcp")
|
||||||
@@ -26,12 +36,95 @@ def extract_url_content(url: str, language_hint: str | None = None) -> dict:
|
|||||||
def extract_item_content(item: dict) -> dict:
|
def extract_item_content(item: dict) -> dict:
|
||||||
"""Extract structured article content from a normalized item object."""
|
"""Extract structured article content from a normalized item object."""
|
||||||
parsed_item = Item.model_validate(item)
|
parsed_item = Item.model_validate(item)
|
||||||
result = extract_content(
|
result = extract_content(ExtractionInput(item=parsed_item))
|
||||||
ExtractionInput(item=parsed_item)
|
|
||||||
)
|
|
||||||
return result.model_dump(mode="json")
|
return result.model_dump(mode="json")
|
||||||
|
|
||||||
|
|
||||||
|
@mcp.tool()
|
||||||
|
def filter_summary_result(
|
||||||
|
summary_result: dict,
|
||||||
|
extracted_article: dict | None = None,
|
||||||
|
item: dict | None = None,
|
||||||
|
context: dict | None = None,
|
||||||
|
) -> dict:
|
||||||
|
"""Apply deterministic filter rules to a structured summary result."""
|
||||||
|
parsed_summary = LlmSummaryResult.model_validate(summary_result)
|
||||||
|
parsed_article = ExtractedArticle.model_validate(extracted_article) if extracted_article else None
|
||||||
|
parsed_item = Item.model_validate(item) if item else None
|
||||||
|
parsed_context = FilterContext.model_validate(context or {})
|
||||||
|
rules = load_filter_rules()
|
||||||
|
decision = evaluate_filter_rules(
|
||||||
|
FilterInput(
|
||||||
|
item=parsed_item,
|
||||||
|
article=parsed_article,
|
||||||
|
summary=parsed_summary,
|
||||||
|
context=parsed_context,
|
||||||
|
),
|
||||||
|
rules,
|
||||||
|
)
|
||||||
|
return decision.model_dump(mode="json")
|
||||||
|
|
||||||
|
|
||||||
|
@mcp.tool()
|
||||||
|
def run_freshrss_openclaw_pipeline(
|
||||||
|
limit: int = 5,
|
||||||
|
mark_read: bool = False,
|
||||||
|
include_read: bool = False,
|
||||||
|
debug_artifacts: bool = False,
|
||||||
|
continuation: str | None = None,
|
||||||
|
timeout_seconds: float = 60.0,
|
||||||
|
max_retries: int = 2,
|
||||||
|
stream_id: str = "user/-/state/com.google/reading-list",
|
||||||
|
api_base_url: str | None = None,
|
||||||
|
username: str | None = None,
|
||||||
|
api_password: str | None = None,
|
||||||
|
llm_api_key: str | None = None,
|
||||||
|
llm_model: str | None = None,
|
||||||
|
llm_api_url: str | None = None,
|
||||||
|
context: dict | None = None,
|
||||||
|
run_id: str | None = None,
|
||||||
|
date_value: str | None = None,
|
||||||
|
output_dir: str | None = None,
|
||||||
|
include_item_reports: bool = False,
|
||||||
|
) -> dict:
|
||||||
|
"""Run the full FreshRSS -> extract -> LLM -> filter -> OpenClaw payload pipeline."""
|
||||||
|
temp_context_path: Path | None = None
|
||||||
|
|
||||||
|
try:
|
||||||
|
if context is not None:
|
||||||
|
with tempfile.NamedTemporaryFile("w", encoding="utf-8", suffix=".json", delete=False) as handle:
|
||||||
|
json.dump(context, handle, ensure_ascii=False, indent=2)
|
||||||
|
temp_context_path = Path(handle.name)
|
||||||
|
|
||||||
|
result = run_freshrss_pipeline(
|
||||||
|
api_base_url=api_base_url,
|
||||||
|
username=username,
|
||||||
|
api_password=api_password,
|
||||||
|
stream_id=stream_id,
|
||||||
|
limit=limit,
|
||||||
|
continuation=continuation,
|
||||||
|
include_read=include_read,
|
||||||
|
mark_read=mark_read,
|
||||||
|
debug_artifacts=debug_artifacts,
|
||||||
|
context_path=temp_context_path,
|
||||||
|
max_retries=max_retries,
|
||||||
|
timeout_seconds=timeout_seconds,
|
||||||
|
llm_api_key=llm_api_key,
|
||||||
|
llm_model=llm_model,
|
||||||
|
llm_api_url=llm_api_url,
|
||||||
|
run_id=run_id,
|
||||||
|
delivery_date=date.fromisoformat(date_value) if date_value else None,
|
||||||
|
output_dir=Path(output_dir) if output_dir else None,
|
||||||
|
)
|
||||||
|
finally:
|
||||||
|
if temp_context_path and temp_context_path.exists():
|
||||||
|
temp_context_path.unlink()
|
||||||
|
|
||||||
|
if not include_item_reports:
|
||||||
|
result = {key: value for key, value in result.items() if key != "items"}
|
||||||
|
return result
|
||||||
|
|
||||||
|
|
||||||
def main() -> None:
|
def main() -> None:
|
||||||
mcp.run()
|
mcp.run()
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,3 @@
|
|||||||
|
from .markdown import MarkdownSink
|
||||||
|
|
||||||
|
__all__ = ["MarkdownSink"]
|
||||||
@@ -0,0 +1,127 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from datetime import UTC, datetime
|
||||||
|
from pathlib import Path
|
||||||
|
import re
|
||||||
|
|
||||||
|
from summary_mcp.models.sink import SinkInput, SinkResult
|
||||||
|
|
||||||
|
|
||||||
|
_DECISION_DIR = {
|
||||||
|
"keep": "inbox",
|
||||||
|
"review": "review",
|
||||||
|
"drop": "archive",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _slugify(text: str) -> str:
|
||||||
|
slug = text.lower()
|
||||||
|
slug = re.sub(r"[^a-z0-9\u4e00-\u9fff]+", "-", slug)
|
||||||
|
slug = re.sub(r"-+", "-", slug).strip("-")
|
||||||
|
return slug or "untitled"
|
||||||
|
|
||||||
|
|
||||||
|
def _yaml_scalar(value: str) -> str:
|
||||||
|
escaped = value.replace('"', '\\"')
|
||||||
|
return f'"{escaped}"'
|
||||||
|
|
||||||
|
|
||||||
|
def _yaml_list(values: list[str]) -> str:
|
||||||
|
if not values:
|
||||||
|
return "[]"
|
||||||
|
return "\n".join(f" - {_yaml_scalar(value)}" for value in values)
|
||||||
|
|
||||||
|
|
||||||
|
def _frontmatter(payload: SinkInput) -> str:
|
||||||
|
article = payload.article
|
||||||
|
summary = payload.summary
|
||||||
|
decision = payload.filter_decision
|
||||||
|
generated_at = payload.metadata.generated_at or datetime.now(tz=UTC)
|
||||||
|
published_at = article.published_at.isoformat() if article.published_at else ""
|
||||||
|
source_id = article.source_id or (payload.item.source_id if payload.item else "")
|
||||||
|
item_id = article.item_id or (payload.item.item_id if payload.item else "")
|
||||||
|
|
||||||
|
return "\n".join(
|
||||||
|
[
|
||||||
|
"---",
|
||||||
|
f"title: {_yaml_scalar(summary.title)}",
|
||||||
|
f"url: {_yaml_scalar(str(summary.url))}",
|
||||||
|
f"source_id: {_yaml_scalar(source_id)}",
|
||||||
|
f"item_id: {_yaml_scalar(item_id)}",
|
||||||
|
f"extract_id: {_yaml_scalar(article.extract_id)}",
|
||||||
|
f"category: {_yaml_scalar(summary.category)}",
|
||||||
|
f"worth_keeping: {'true' if summary.worth_keeping else 'false'}",
|
||||||
|
f"decision: {_yaml_scalar(decision.decision)}",
|
||||||
|
f"priority: {decision.priority}",
|
||||||
|
"topics:",
|
||||||
|
_yaml_list(summary.topics),
|
||||||
|
"keywords:",
|
||||||
|
_yaml_list(summary.keywords),
|
||||||
|
"labels:",
|
||||||
|
_yaml_list(decision.labels),
|
||||||
|
f"published_at: {_yaml_scalar(published_at)}",
|
||||||
|
f"saved_at: {_yaml_scalar(generated_at.isoformat())}",
|
||||||
|
"quality_flags:",
|
||||||
|
f" is_paywalled: {'true' if article.quality_flags.is_paywalled else 'false'}",
|
||||||
|
f" is_truncated: {'true' if article.quality_flags.is_truncated else 'false'}",
|
||||||
|
f" is_low_content: {'true' if article.quality_flags.is_low_content else 'false'}",
|
||||||
|
"---",
|
||||||
|
]
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _render_markdown(payload: SinkInput) -> str:
|
||||||
|
summary = payload.summary
|
||||||
|
article = payload.article
|
||||||
|
decision = payload.filter_decision
|
||||||
|
highlights = "\n".join(f"- {item}" for item in summary.highlights)
|
||||||
|
reasons = "\n".join(f"- {item}" for item in decision.reasons)
|
||||||
|
|
||||||
|
return (
|
||||||
|
f"{_frontmatter(payload)}\n\n"
|
||||||
|
f"# 摘要\n\n{summary.summary}\n\n"
|
||||||
|
f"# 要点\n\n{highlights}\n\n"
|
||||||
|
f"# 过滤决策\n\n"
|
||||||
|
f"- decision: `{decision.decision}`\n"
|
||||||
|
f"- priority: `{decision.priority}`\n"
|
||||||
|
f"- matched rules: {', '.join(decision.matched_rules) if decision.matched_rules else 'none'}\n\n"
|
||||||
|
f"## reasons\n\n{reasons}\n\n"
|
||||||
|
f"# 原文信息\n\n"
|
||||||
|
f"- 标题: {summary.title}\n"
|
||||||
|
f"- 链接: {summary.url}\n"
|
||||||
|
f"- 作者: {article.author or 'unknown'}\n"
|
||||||
|
f"- 发布时间: {article.published_at.isoformat() if article.published_at else 'unknown'}\n"
|
||||||
|
f"- 内容类型: {article.content_kind}\n\n"
|
||||||
|
f"# 正文摘录\n\n{article.plain_text.strip()}\n"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
class MarkdownSink:
|
||||||
|
def __init__(self, base_dir: Path) -> None:
|
||||||
|
self.base_dir = base_dir
|
||||||
|
|
||||||
|
def _target_dir(self, decision: str, published_at: datetime | None) -> Path:
|
||||||
|
bucket = _DECISION_DIR.get(decision, "archive")
|
||||||
|
if published_at is None:
|
||||||
|
return self.base_dir / bucket / "unknown-date"
|
||||||
|
return self.base_dir / bucket / published_at.strftime("%Y") / published_at.strftime("%Y-%m")
|
||||||
|
|
||||||
|
def _target_name(self, payload: SinkInput) -> str:
|
||||||
|
article = payload.article
|
||||||
|
published_at = article.published_at or datetime.now(tz=UTC)
|
||||||
|
title = payload.summary.title or article.title or "untitled"
|
||||||
|
return f"{published_at.strftime('%Y-%m-%d')}-{_slugify(title)}.md"
|
||||||
|
|
||||||
|
def write(self, payload: SinkInput) -> SinkResult:
|
||||||
|
target_dir = self._target_dir(payload.filter_decision.decision, payload.article.published_at)
|
||||||
|
target_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
output_path = target_dir / self._target_name(payload)
|
||||||
|
output_path.write_text(_render_markdown(payload), encoding="utf-8")
|
||||||
|
return SinkResult(
|
||||||
|
target="markdown",
|
||||||
|
success=True,
|
||||||
|
output_path=str(output_path),
|
||||||
|
message="Markdown note written successfully.",
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
@@ -47,12 +47,31 @@ def _validate_business_rules(result: LlmSummaryResult, extracted: dict[str, Any]
|
|||||||
if extracted_url and str(result.url) != extracted_url:
|
if extracted_url and str(result.url) != extracted_url:
|
||||||
errors.append("`url` does not match extracted article url")
|
errors.append("`url` does not match extracted article url")
|
||||||
|
|
||||||
if result.category == "\u8d44\u8baf" and result.worth_keeping:
|
if result.category == "资讯" and result.worth_keeping:
|
||||||
warnings.append("`??` category marked as worth keeping; check if this is intentional.")
|
warnings.append("`资讯` category marked as worth keeping; check if this is intentional.")
|
||||||
|
|
||||||
return errors, warnings
|
return errors, warnings
|
||||||
|
|
||||||
|
|
||||||
|
def validate_llm_result_payload(
|
||||||
|
result_payload: dict[str, Any],
|
||||||
|
extracted_payload: dict[str, Any] | None = None,
|
||||||
|
) -> ValidationReport:
|
||||||
|
try:
|
||||||
|
parsed = LlmSummaryResult.model_validate(result_payload)
|
||||||
|
except ValidationError as exc:
|
||||||
|
errors = [f"{'.'.join(str(part) for part in error['loc'])}: {error['msg']}" for error in exc.errors()]
|
||||||
|
return ValidationReport(valid=False, errors=errors)
|
||||||
|
|
||||||
|
errors, warnings = _validate_business_rules(parsed, extracted_payload)
|
||||||
|
return ValidationReport(
|
||||||
|
valid=not errors,
|
||||||
|
errors=errors,
|
||||||
|
warnings=warnings,
|
||||||
|
normalized_result=parsed.model_dump(mode="json"),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
def validate_llm_result(
|
def validate_llm_result(
|
||||||
result_path: Path,
|
result_path: Path,
|
||||||
extracted_path: Path | None = None,
|
extracted_path: Path | None = None,
|
||||||
@@ -69,16 +88,4 @@ def validate_llm_result(
|
|||||||
except json.JSONDecodeError as exc:
|
except json.JSONDecodeError as exc:
|
||||||
return ValidationReport(valid=False, errors=[f"Invalid extracted JSON: {exc}"])
|
return ValidationReport(valid=False, errors=[f"Invalid extracted JSON: {exc}"])
|
||||||
|
|
||||||
try:
|
return validate_llm_result_payload(raw_result, extracted)
|
||||||
parsed = LlmSummaryResult.model_validate(raw_result)
|
|
||||||
except ValidationError as exc:
|
|
||||||
errors = [f"{'.'.join(str(part) for part in error['loc'])}: {error['msg']}" for error in exc.errors()]
|
|
||||||
return ValidationReport(valid=False, errors=errors)
|
|
||||||
|
|
||||||
errors, warnings = _validate_business_rules(parsed, extracted)
|
|
||||||
return ValidationReport(
|
|
||||||
valid=not errors,
|
|
||||||
errors=errors,
|
|
||||||
warnings=warnings,
|
|
||||||
normalized_result=parsed.model_dump(mode="json"),
|
|
||||||
)
|
|
||||||
|
|||||||
@@ -0,0 +1,3 @@
|
|||||||
|
from .freshrss_pipeline import default_output_dir, read_delivery_payload, run_freshrss_pipeline
|
||||||
|
|
||||||
|
__all__ = ["default_output_dir", "read_delivery_payload", "run_freshrss_pipeline"]
|
||||||
@@ -0,0 +1,303 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
from datetime import UTC, date, datetime
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from summary_mcp.core.keyword_index import persist_keyword_indexes
|
||||||
|
from summary_mcp.core.pipeline import extract_content
|
||||||
|
from summary_mcp.core.summary_loop import resolve_llm_settings, run_loop_payload
|
||||||
|
from summary_mcp.filters.engine import evaluate_filter_rules, load_filter_rules
|
||||||
|
from summary_mcp.integrations.freshrss import READ_TAG, FreshRSSClient, map_entry_to_item
|
||||||
|
from summary_mcp.models.article_candidate import (
|
||||||
|
CandidateMetadata,
|
||||||
|
CandidateSourceRefs,
|
||||||
|
OpenClawCandidateInput,
|
||||||
|
build_article_candidate_record,
|
||||||
|
build_openclaw_candidate_input,
|
||||||
|
)
|
||||||
|
from summary_mcp.models.filtering import FilterContext, FilterInput
|
||||||
|
from summary_mcp.models.llm_result import LlmSummaryResult
|
||||||
|
from summary_mcp.models.openclaw_delivery import OpenClawDeliveryPayload, build_openclaw_delivery_payload
|
||||||
|
from summary_mcp.models.summary_io import ExtractionInput
|
||||||
|
|
||||||
|
|
||||||
|
REPO_ROOT = Path(__file__).resolve().parents[3]
|
||||||
|
OUTPUT_ROOT = REPO_ROOT / "outputs"
|
||||||
|
FRESHRSS_OUTPUT_ROOT = OUTPUT_ROOT / "freshrss"
|
||||||
|
DATA_ROOT = REPO_ROOT / "data" / "term_index"
|
||||||
|
DEFAULT_PROMPT_PATH = OUTPUT_ROOT / "prompts" / "llm-summary-prompt.txt"
|
||||||
|
DEFAULT_RULES_PATH = REPO_ROOT / "configs" / "filter_rules.json"
|
||||||
|
DEFAULT_TERM_ALIASES_PATH = REPO_ROOT / "configs" / "term_aliases.json"
|
||||||
|
DEFAULT_TERM_STOPWORDS_PATH = REPO_ROOT / "configs" / "term_stopwords.json"
|
||||||
|
DEFAULT_TERM_DAILY_DIR = DATA_ROOT / "daily"
|
||||||
|
DEFAULT_TERM_STATS_PATH = DATA_ROOT / "term_stats.json"
|
||||||
|
|
||||||
|
|
||||||
|
def _save_json(path: Path, payload: dict[str, Any] | list[Any]) -> None:
|
||||||
|
path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
path.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
|
def _load_json(path: Path) -> dict[str, Any]:
|
||||||
|
return json.loads(path.read_text(encoding="utf-8-sig"))
|
||||||
|
|
||||||
|
|
||||||
|
def _load_required_env(name: str, value: str | None) -> str:
|
||||||
|
resolved = value or os.environ.get(name)
|
||||||
|
if not resolved:
|
||||||
|
raise RuntimeError(f"Missing required value: {name}.")
|
||||||
|
return resolved
|
||||||
|
|
||||||
|
|
||||||
|
def default_output_dir() -> Path:
|
||||||
|
return FRESHRSS_OUTPUT_ROOT / "rerun" / datetime.now(tz=UTC).strftime("%Y%m%d-%H%M%S")
|
||||||
|
|
||||||
|
|
||||||
|
def _maybe_path(enabled: bool, path: Path) -> Path | None:
|
||||||
|
return path if enabled else None
|
||||||
|
|
||||||
|
|
||||||
|
def run_freshrss_pipeline(
|
||||||
|
*,
|
||||||
|
api_base_url: str | None = None,
|
||||||
|
username: str | None = None,
|
||||||
|
api_password: str | None = None,
|
||||||
|
stream_id: str = "user/-/state/com.google/reading-list",
|
||||||
|
limit: int = 5,
|
||||||
|
continuation: str | None = None,
|
||||||
|
include_read: bool = False,
|
||||||
|
mark_read: bool = False,
|
||||||
|
debug_artifacts: bool = False,
|
||||||
|
prompt: Path | None = None,
|
||||||
|
rules: Path | None = None,
|
||||||
|
context_path: Path | None = None,
|
||||||
|
max_retries: int = 2,
|
||||||
|
timeout_seconds: float = 60.0,
|
||||||
|
llm_api_key: str | None = None,
|
||||||
|
llm_model: str | None = None,
|
||||||
|
llm_api_url: str | None = None,
|
||||||
|
run_id: str | None = None,
|
||||||
|
delivery_date: date | None = None,
|
||||||
|
output_dir: Path | None = None,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
started_at = datetime.now(tz=UTC)
|
||||||
|
resolved_output_dir = output_dir or default_output_dir()
|
||||||
|
run_stamp = started_at.strftime("%Y%m%d-%H%M%S")
|
||||||
|
resolved_run_id = run_id or f"freshrss-pipeline-{run_stamp}"
|
||||||
|
resolved_delivery_date = delivery_date or datetime.now(tz=UTC).date()
|
||||||
|
|
||||||
|
resolved_api_base_url = _load_required_env("FRESHRSS_API_BASE_URL", api_base_url)
|
||||||
|
resolved_username = _load_required_env("FRESHRSS_USERNAME", username)
|
||||||
|
resolved_api_password = _load_required_env("FRESHRSS_API_PASSWORD", api_password)
|
||||||
|
resolved_llm_api_key, resolved_llm_model, resolved_llm_api_url = resolve_llm_settings(
|
||||||
|
api_key=llm_api_key,
|
||||||
|
model=llm_model,
|
||||||
|
api_url=llm_api_url,
|
||||||
|
)
|
||||||
|
|
||||||
|
resolved_prompt_path = prompt or DEFAULT_PROMPT_PATH
|
||||||
|
resolved_rules_path = rules or DEFAULT_RULES_PATH
|
||||||
|
raw_output = resolved_output_dir / "raw" / "freshrss.raw.json"
|
||||||
|
delivery_output = resolved_output_dir / "candidates" / "openclaw-delivery-payload.json"
|
||||||
|
report_output = resolved_output_dir / "run-report.json"
|
||||||
|
items_list_output = _maybe_path(debug_artifacts, resolved_output_dir / "items" / "freshrss.items.json")
|
||||||
|
|
||||||
|
client = FreshRSSClient(
|
||||||
|
api_base_url=resolved_api_base_url,
|
||||||
|
username=resolved_username,
|
||||||
|
api_password=resolved_api_password,
|
||||||
|
timeout_seconds=timeout_seconds,
|
||||||
|
)
|
||||||
|
auth_token = client.client_login()
|
||||||
|
payload = client.fetch_stream_contents(
|
||||||
|
auth_token=auth_token,
|
||||||
|
stream_id=stream_id,
|
||||||
|
limit=limit,
|
||||||
|
continuation=continuation,
|
||||||
|
exclude_targets=[] if include_read else [READ_TAG],
|
||||||
|
)
|
||||||
|
entries = payload.get("items")
|
||||||
|
if not isinstance(entries, list):
|
||||||
|
raise RuntimeError("FreshRSS stream response does not contain an items array.")
|
||||||
|
|
||||||
|
_save_json(raw_output, payload)
|
||||||
|
|
||||||
|
items = [map_entry_to_item(entry) for entry in entries]
|
||||||
|
if items_list_output is not None:
|
||||||
|
_save_json(items_list_output, [item.model_dump(mode="json") for item in items])
|
||||||
|
|
||||||
|
loaded_rules = load_filter_rules(resolved_rules_path)
|
||||||
|
context = FilterContext.model_validate(_load_json(context_path)) if context_path else FilterContext()
|
||||||
|
|
||||||
|
delivered_candidates: list[OpenClawCandidateInput] = []
|
||||||
|
delivered_item_ids: list[str] = []
|
||||||
|
item_reports: list[dict[str, Any]] = []
|
||||||
|
|
||||||
|
for index, item in enumerate(items, start=1):
|
||||||
|
item_key = f"item-{index:02d}"
|
||||||
|
item_path = _maybe_path(debug_artifacts, resolved_output_dir / "items" / f"{item_key}.item.json")
|
||||||
|
extracted_path = _maybe_path(debug_artifacts, resolved_output_dir / "extracted" / f"{item_key}.extracted.json")
|
||||||
|
summary_output = _maybe_path(debug_artifacts, resolved_output_dir / "summary" / item_key / "result.loop.json")
|
||||||
|
filter_path = _maybe_path(debug_artifacts, resolved_output_dir / "filter" / f"{item_key}.filter.json")
|
||||||
|
record_path = _maybe_path(debug_artifacts, resolved_output_dir / "candidates" / f"{item_key}.article-candidate-record.json")
|
||||||
|
openclaw_path = _maybe_path(debug_artifacts, resolved_output_dir / "candidates" / f"{item_key}.openclaw-candidate-input.json")
|
||||||
|
|
||||||
|
if item_path is not None:
|
||||||
|
_save_json(item_path, item.model_dump(mode="json"))
|
||||||
|
|
||||||
|
item_report: dict[str, Any] = {
|
||||||
|
"item_key": item_key,
|
||||||
|
"item_id": item.item_id,
|
||||||
|
"external_id": item.external_id,
|
||||||
|
"url": str(item.url),
|
||||||
|
"title": item.title,
|
||||||
|
"status": "pulled",
|
||||||
|
}
|
||||||
|
if debug_artifacts:
|
||||||
|
item_report["paths"] = {
|
||||||
|
"item": str(item_path) if item_path else None,
|
||||||
|
"extracted": str(extracted_path) if extracted_path else None,
|
||||||
|
"summary": str(summary_output) if summary_output else None,
|
||||||
|
"filter": str(filter_path) if filter_path else None,
|
||||||
|
"article_candidate": str(record_path) if record_path else None,
|
||||||
|
"openclaw_candidate": str(openclaw_path) if openclaw_path else None,
|
||||||
|
}
|
||||||
|
|
||||||
|
extraction = extract_content(ExtractionInput(item=item))
|
||||||
|
extracted_payload = extraction.model_dump(mode="json")
|
||||||
|
if extracted_path is not None:
|
||||||
|
_save_json(extracted_path, extracted_payload)
|
||||||
|
if not extraction.success or extraction.article is None:
|
||||||
|
item_report["status"] = "extract_failed"
|
||||||
|
item_report["error"] = extraction.error.model_dump(mode="json") if extraction.error else None
|
||||||
|
item_reports.append(item_report)
|
||||||
|
continue
|
||||||
|
|
||||||
|
item_report["status"] = "extracted"
|
||||||
|
summary_exit_code, summary_payload, summary_report = run_loop_payload(
|
||||||
|
extracted_payload=extracted_payload,
|
||||||
|
prompt_path=resolved_prompt_path,
|
||||||
|
output_path=summary_output,
|
||||||
|
max_retries=max_retries,
|
||||||
|
timeout_seconds=timeout_seconds,
|
||||||
|
api_key=resolved_llm_api_key,
|
||||||
|
model=resolved_llm_model,
|
||||||
|
api_url=resolved_llm_api_url,
|
||||||
|
)
|
||||||
|
if summary_exit_code != 0 or summary_payload is None:
|
||||||
|
item_report["status"] = "summary_failed"
|
||||||
|
if summary_report is not None:
|
||||||
|
item_report["summary_errors"] = summary_report.errors
|
||||||
|
item_reports.append(item_report)
|
||||||
|
continue
|
||||||
|
|
||||||
|
summary = LlmSummaryResult.model_validate(summary_payload)
|
||||||
|
decision = evaluate_filter_rules(
|
||||||
|
FilterInput(item=item, article=extraction.article, summary=summary, context=context),
|
||||||
|
loaded_rules,
|
||||||
|
)
|
||||||
|
if filter_path is not None:
|
||||||
|
_save_json(filter_path, decision.model_dump(mode="json"))
|
||||||
|
|
||||||
|
record = build_article_candidate_record(
|
||||||
|
summary=summary,
|
||||||
|
article=extraction.article,
|
||||||
|
filter_result=decision,
|
||||||
|
item=item,
|
||||||
|
source_refs=CandidateSourceRefs(
|
||||||
|
item_path=str(item_path) if item_path else None,
|
||||||
|
extracted_path=str(extracted_path) if extracted_path else None,
|
||||||
|
summary_path=str(summary_output) if summary_output else None,
|
||||||
|
filter_path=str(filter_path) if filter_path else None,
|
||||||
|
),
|
||||||
|
metadata=CandidateMetadata(
|
||||||
|
generated_at=datetime.now(tz=UTC),
|
||||||
|
producer="run_freshrss_pipeline",
|
||||||
|
run_id=resolved_run_id,
|
||||||
|
),
|
||||||
|
)
|
||||||
|
openclaw_input = build_openclaw_candidate_input(record)
|
||||||
|
|
||||||
|
if record_path is not None:
|
||||||
|
_save_json(record_path, record.model_dump(mode="json"))
|
||||||
|
if openclaw_path is not None:
|
||||||
|
_save_json(openclaw_path, openclaw_input.model_dump(mode="json"))
|
||||||
|
|
||||||
|
delivered_candidates.append(openclaw_input)
|
||||||
|
if item.external_id:
|
||||||
|
delivered_item_ids.append(item.external_id)
|
||||||
|
|
||||||
|
item_report["status"] = "delivered"
|
||||||
|
item_report["selection_decision"] = decision.decision
|
||||||
|
item_report["candidate_id"] = openclaw_input.candidate_id
|
||||||
|
item_reports.append(item_report)
|
||||||
|
|
||||||
|
delivered_candidates.sort(key=lambda candidate: candidate.digest_rank, reverse=True)
|
||||||
|
delivery_payload = build_openclaw_delivery_payload(
|
||||||
|
delivered_candidates,
|
||||||
|
run_id=resolved_run_id,
|
||||||
|
for_date=resolved_delivery_date,
|
||||||
|
)
|
||||||
|
_save_json(delivery_output, delivery_payload.model_dump(mode="json"))
|
||||||
|
|
||||||
|
keyword_index_result = persist_keyword_indexes(
|
||||||
|
delivery_payload.candidates,
|
||||||
|
for_date=delivery_payload.date,
|
||||||
|
digest_id=delivery_payload.run_id,
|
||||||
|
source="openclaw_delivery_payload",
|
||||||
|
daily_dir=DEFAULT_TERM_DAILY_DIR,
|
||||||
|
stats_path=DEFAULT_TERM_STATS_PATH,
|
||||||
|
aliases_path=DEFAULT_TERM_ALIASES_PATH,
|
||||||
|
stopwords_path=DEFAULT_TERM_STOPWORDS_PATH,
|
||||||
|
)
|
||||||
|
|
||||||
|
marked_count = 0
|
||||||
|
if mark_read and delivered_item_ids:
|
||||||
|
client.mark_items_as_read(auth_token=auth_token, item_ids=delivered_item_ids)
|
||||||
|
marked_count = len({item_id for item_id in delivered_item_ids if item_id})
|
||||||
|
|
||||||
|
status_counts: dict[str, int] = {}
|
||||||
|
for item_report in item_reports:
|
||||||
|
status = str(item_report["status"])
|
||||||
|
status_counts[status] = status_counts.get(status, 0) + 1
|
||||||
|
|
||||||
|
report = {
|
||||||
|
"run_id": resolved_run_id,
|
||||||
|
"started_at": started_at.isoformat(),
|
||||||
|
"completed_at": datetime.now(tz=UTC).isoformat(),
|
||||||
|
"requested_limit": limit,
|
||||||
|
"pulled_count": len(items),
|
||||||
|
"delivered_count": len(delivered_candidates),
|
||||||
|
"marked_read_count": marked_count,
|
||||||
|
"mark_read_requested": mark_read,
|
||||||
|
"debug_artifacts": debug_artifacts,
|
||||||
|
"raw_output": str(raw_output),
|
||||||
|
"delivery_output": str(delivery_output),
|
||||||
|
"keyword_index": keyword_index_result,
|
||||||
|
"status_counts": status_counts,
|
||||||
|
"items": item_reports,
|
||||||
|
}
|
||||||
|
_save_json(report_output, report)
|
||||||
|
|
||||||
|
return {
|
||||||
|
"run_id": resolved_run_id,
|
||||||
|
"output_dir": str(resolved_output_dir),
|
||||||
|
"raw_output": str(raw_output),
|
||||||
|
"delivery_output": str(delivery_output),
|
||||||
|
"report_output": str(report_output),
|
||||||
|
"keyword_index": keyword_index_result,
|
||||||
|
"pulled_count": len(items),
|
||||||
|
"delivered_count": len(delivered_candidates),
|
||||||
|
"marked_read_count": marked_count,
|
||||||
|
"status_counts": status_counts,
|
||||||
|
"debug_artifacts": debug_artifacts,
|
||||||
|
"delivery_payload": delivery_payload.model_dump(mode="json"),
|
||||||
|
"items": item_reports,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def read_delivery_payload(path: Path) -> OpenClawDeliveryPayload:
|
||||||
|
return OpenClawDeliveryPayload.model_validate(_load_json(path))
|
||||||
Reference in New Issue
Block a user