Add article-deep-summary skill and enrich article summary prompt

- Create skills/article-deep-summary/SKILL.md and agents/openai.yaml
- Relax prompt constraints to allow 2-4 sentence elaboration per field
- Add validation output for reference article
- Update plan checklist as completed
This commit is contained in:
zhuyongxin
2026-03-30 18:44:30 +08:00
parent e24db99baa
commit 5b8df317ef
6 changed files with 160 additions and 15 deletions
+6 -6
View File
@@ -11,16 +11,16 @@
{ {
"title": "文章标题(与原文一致)", "title": "文章标题(与原文一致)",
"url": "文章 URL(与原文一致)", "url": "文章 URL(与原文一致)",
"core_conclusion": "作者最核心的结论,1-2 句,精准概括", "core_conclusion": "作者最核心的结论,2-4 句,精准概括,包含核心判断和支撑依据",
"main_argument": "文章的主要论点或主张,可展开,允许多句", "main_argument": "文章的主要论点或主张,充分展开,可包含推理链条和论证结构",
"key_methods": [ "key_methods": [
"关键方法、机制或技术手段,每条一句,3-6 条;无相关内容时返回空数组" "关键方法、机制或技术手段,每条 2-4 句展开说明原理或运作方式,3-6 条;无相关内容时返回空数组"
], ],
"important_details": [ "important_details": [
"值得记录的细节、数据或案例,每条一句,3-6 条;无相关内容时返回空数组" "值得记录的细节、数据或案例,每条 2-4 句补充背景和意义,3-6 条;无相关内容时返回空数组"
], ],
"reusable_insights": [ "reusable_insights": [
"可复用于其他场景的启发或观点,2-4 条" "可复用于其他场景的启发或观点,每条 2-3 句说明适用场景和具体做法,2-4 条"
], ],
"keywords": [ "keywords": [
"具体实体、工具名、方法名,5-8 个" "具体实体、工具名、方法名,5-8 个"
@@ -30,7 +30,7 @@
], ],
"category": "内容分类,从以下选项中选择一个:资讯 / 方法论 / 工具实践 / 观点评论", "category": "内容分类,从以下选项中选择一个:资讯 / 方法论 / 工具实践 / 观点评论",
"worth_keeping": true, "worth_keeping": true,
"reason": "沉淀理由,一句话说明为什么值得长期保留" "reason": "沉淀理由,2-3 句说明为什么值得长期保留,以及未来什么场景下可以回看"
} }
注意: 注意:
@@ -0,0 +1,34 @@
# 信息过载时代,我的漏斗式阅读工作流
Source: https://shawnxie.top/blogs/tools/read-flow-2026.html
Category: 方法论
## 核心结论
在信息过载时代,个人信息处理的核心目标不是获取更多信息,而是通过分层过滤和沉淀机制,稳定地吸收、判断和沉淀真正有价值的内容。作者认为,一个有效的个人信息系统应该是一个“漏斗”,而非“桶”,其关键在于将分散的处理环节串联成闭环,让信息在向下流动的过程中不断收窄,最终只留下值得进入长期记忆的部分,并能通过沉淀内容反向优化筛选逻辑。
## 主要论点
文章主张,应对信息过载不应依赖单一工具或追求全自动处理,而应构建一套分层、闭环的“漏斗式”个人信息系统。该系统以RSS等可控信息源为起点,通过聚合、预处理、AI精选、人工精读、知识沉淀和轻量反馈等多个层级,逐步过滤噪音、提炼精华,并将沉淀的高价值内容转化为个性化信号,反向优化上游筛选。其核心论点是:信息处理的难点在于串联分散环节,真正的价值在于实现“上游宽广、中游稳定、下游精准、回流轻柔”的认知加工流程,使人从被动接收者转变为拥有个性化认知处理系统的主人。
## 关键方法 / 机制
- 信息源归一化:以RSS为主信息基础设施,利用RSSHub、wewe-rss、nitter等工具或自定义脚本,将公众号、社区、社交平台等多样信息源统一转换为RSS或类feed格式,确保下游处理环节输入格式的稳定和统一。
- 分层处理流程:设计清晰的多层处理架构。1) 用FreshRSS作为统一聚合池,管理所有订阅源,充当缓冲层。2) 用Digest模块进行预处理,完成URL去重、相似内容去重、正文抓取、质量检查、噪音过滤、摘要生成等任务,将原始信息整理为可判断对象。3) 用Daily Review进行AI精选,由LLM将预处理后的内容按固定栏目(如今日大事、变更与实践等)结构化,生成重点突出的日报。4) 设置人工精读环节(Human in the loop),由人最终判断内容的长期价值。5) 用Lumina知识库进行长期沉淀。6) 基于沉淀内容构建轻量兴趣画像,作为辅助信号轻微影响上游排序,形成闭环。
- 自动化编排与集成:使用OpenClaw作为自动化编排层,通过自然语言描述诉求,串联定时触发、外部工具调用(如FreshRSS API)、与飞书/Lumina的交互、AI技能调用等任务,实现整个工作流的自动化运行和快速迭代,降低了构建完整系统的初期门槛。
## 重要细节
- FreshRSS的定位与作用:作者强调FreshRSS并非日常阅读工具,而是作为“中间水库”的聚合池。其核心作用是统一管理信息源格式、为下游环节提供稳定且无需直接联网的内容池、以及确保信息处理的时间连续性,避免了碎片化处理,是实现“有边界的信息处理”的关键缓冲层。
- Digest预处理的具体任务:Digest模块承担了将“海量候选内容转化为可判断对象”的重任。其具体任务包括URL精确去重、相似内容去重、正文抓取、质量检查、噪音过滤、摘要生成和初步排序。这一环节大幅减少了标题党、重复报道和无效数据,显著降低了后续人工筛选的认知成本,是提升整个系统效率的基础。
- Daily Review的栏目化设计:AI精选环节(Daily Review)并非简单聚合摘要,而是通过预设固定栏目进行结构化编辑。栏目包括“今日大事”、“变更与实践”、“安全与风险”、“开源与工具”、“洞察与数据点”、“主题深挖”等,将信息分配到不同的“认知槽位”,使最终产出更像一份重点突出的个性化日报,而非信息堆砌。
- 轻量反馈与兴趣画像设计:为避免系统演变为封闭的“信息茧房”或过度迎合的推荐引擎,作者刻意设计了“轻量”的兴趣画像逻辑。它仅从长期精读主题、高价值信息源、偏好内容格式、Lumina存入内容类型等维度学习,作为辅助信号轻微影响Digest和Daily Review的排序,目标是减少无效判断,同时保留关注公共重要性和探索未知的空间。
- 沉淀价值的拓展方向:信息沉淀的终点不仅是存储,还包括价值再生产。作者实践了两个方向:1) 周刊生成:基于一周的Digest、Daily Review和Lumina内容,生成带有个人筛选痕迹和深度分析的周刊,识别长期信号与短期噪音。2) 主题文章生成:系统识别一段时间内反复出现并被多次沉淀的主题(如AI Agent趋势),将其发展为可深入输出的长文主题,将离散信息流转化为结构化知识资产。
## 可复用启发
- 构建分层过滤的认知加工流程:在处理任何海量输入(如邮件、任务、学习资料)时,都可以借鉴“漏斗”思维。设计从“广泛捕获”到“逐步收窄”的多层处理机制,明确每一层的职责(如聚合、粗筛、精筛、决策、沉淀),而非试图一次性处理所有信息,这能大幅降低认知负担并提升处理质量。
- 坚持“人在回路”的核心价值判断:在自动化系统中,尤其是在知识管理领域,最关键的长期价值判断(如“什么值得长期留存”、“什么对未来有深远影响”)不应完全外包给算法。应像本文一样,在流程的关键节点(如进入知识库前)设置人工确认环节,确保系统的输出最终服务于人的深度思考和决策。
- 利用轻量反馈构建良性闭环系统:在设计个性化系统时,应避免构建强反馈、易导致信息茧房的推荐引擎。可以借鉴本文的“轻量引导”思路,仅从最核心、最长期的行为数据(如最终沉淀内容)中提取少量信号,温和地优化上游流程,在提升效率的同时保持系统的开放性和探索性。
- 以“沉淀和再生产”作为流程终点:信息或知识管理的目标不应止步于“读完”或“保存”。应思考如何将处理后的高价值内容进一步转化为可复用的资产,例如定期生成汇总报告(如周刊)、识别并发展跨时间维度的核心主题、或将内化知识用于指导实践和输出,从而实现知识的增值和循环。
## 关键词
RSS、FreshRSS、OpenClaw、Lumina、RSSHub、wewe-rss、nitter、Digest
## 主题
知识管理、个人信息处理、工作流设计、自动化、阅读方法
File diff suppressed because one or more lines are too long
+3 -9
View File
@@ -34,14 +34,8 @@ python -m summary_mcp.workflows.article_summary \
- 数据模型:src/summary_mcp/models/article_summary_result.py - 数据模型:src/summary_mcp/models/article_summary_result.py
- Prompt:outputs/prompts/article-summary-prompt.txt - Prompt:outputs/prompts/article-summary-prompt.txt
## Skill 文件结构(待创建)
skills/article-deep-summary/
SKILL.md
agents/openai.yaml
## 待办 ## 待办
- [ ] 创建 skills/article-deep-summary/SKILL.md - [x] 创建 skills/article-deep-summary/SKILL.md
- [ ] 创建 skills/article-deep-summary/agents/openai.yaml - [x] 创建 skills/article-deep-summary/agents/openai.yaml
- [ ] 用实际 extracted 文件跑一次验证输出 - [x] 用实际 extracted 文件跑一次验证输出(outputs/reference/article-summaries/信息过载时代,我的漏斗式阅读工作流.md)
+113
View File
@@ -0,0 +1,113 @@
---
name: article-deep-summary
description: Generate a deep structured knowledge note from a single extracted article. Use when the user wants to deeply summarize, distill, or create a knowledge note from an article's full plain_text content stored in an *.extracted.json file.
---
# Article Deep Summary
Use this skill when the user has an extracted article file and wants to generate a deep knowledge note — not a brief summary card, but a structured distillation with core conclusion, arguments, methods, details, and reusable insights.
This skill is **not** for validating or repairing existing summaries (use `llm-summary-review` for that), nor for keyword index cleanup (use `keyword-cleanup-review` for that).
## What This Skill Does
- Reads an extracted article JSON containing `plain_text`
- Calls the LLM with the dedicated article-summary prompt
- Validates the LLM output against the `ArticleSummaryResult` schema
- Renders a structured Markdown knowledge note in Chinese
## Inputs
Typical files:
- Extracted article JSON: `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json` (single-item) or a batch file with a `results` array
- Prompt template: `outputs/prompts/article-summary-prompt.txt`
## Workflow
1. Identify the target extracted file and the item IDs to summarize.
2. Run the article-summary workflow via CLI:
```bash
python -m summary_mcp.workflows.article_summary \
--extracted <extracted_path> \
--selected-ids <item_id> \
--output-dir <output_dir>
```
Or call the MCP tool `generate_article_summaries` with:
- `extracted_path`: path to the extracted JSON file
- `selected_ids`: array of item ID strings (pass empty array to summarize all)
- `output_dir`: (optional) directory for Markdown output
3. The workflow internally:
- Resolves LLM settings (`ARTICLE_SUMMARY_*` env vars, falling back to `LLM_*`)
- Calls `run_loop_payload` with the article-summary prompt and validator
- Retries up to 2 times on validation failure
4. Check the output Markdown files in the specified output directory.
## Output Schema
The LLM returns a JSON matching `ArticleSummaryResult`:
| Field | Type | Description |
|---|---|---|
| `title` | string | Article title |
| `url` | HttpUrl | Article URL |
| `core_conclusion` | string | Author's core conclusion, 1-2 sentences |
| `main_argument` | string | Main argument or thesis, can be multi-sentence |
| `key_methods` | string[] | Key methods, mechanisms, or techniques |
| `important_details` | string[] | Noteworthy details, data points, or cases |
| `reusable_insights` | string[] | Reusable insights transferable to other contexts |
| `keywords` | string[] | Specific entities — tool names, frameworks, methods |
| `topics` | string[] | Higher-level topic labels |
| `category` | enum | One of: `资讯` `方法论` `工具实践` `观点评论` |
| `worth_keeping` | bool | Whether the article is worth long-term retention |
| `reason` | string | One-sentence justification for retention |
Constraints enforced by the validator:
- `keywords` and `topics` must not overlap
- `keywords` focuses on concrete entities; `topics` focuses on abstract themes
- `category` must be one of the four allowed values
## Markdown Output
Each article produces one `.md` file with sections:
- 核心结论
- 主要论点
- 关键方法 / 机制
- 重要细节
- 可复用启发
- 关键词
- 主题
## LLM Configuration
Dedicated environment variables (fallback to main `LLM_*` if unset):
- `ARTICLE_SUMMARY_LLM_API_URL`
- `ARTICLE_SUMMARY_LLM_API_KEY`
- `ARTICLE_SUMMARY_LLM_MODEL`
## Repository Implementation
Relevant code:
- Workflow: `src/summary_mcp/workflows/article_summary.py`
- Validator: `src/summary_mcp/validators/article_summary.py`
- Data model: `src/summary_mcp/models/article_summary_result.py`
- Prompt: `outputs/prompts/article-summary-prompt.txt`
- MCP tool: `generate_article_summaries` in `src/summary_mcp/server.py`
## When To Stop
Stop when one of these is true:
- Markdown knowledge notes have been generated for all requested items
- The workflow reports repeated validation failures and the user should decide how to proceed
- The user asks to inspect intermediate results manually
@@ -0,0 +1,3 @@
display_name: Article Deep Summary
short_description: Generate deep knowledge notes from a single extracted article.
default_prompt: Given an extracted article JSON with plain_text content, run the article-summary workflow to produce a structured Markdown knowledge note with core conclusion, main arguments, key methods, important details, and reusable insights.