Translate skill docs (article-deep-summary, llm-summary-review, keyword-cleanup-review) from English to Chinese

This commit is contained in:
wdm
2026-03-30 23:14:22 +08:00
parent 5b8df317ef
commit ca86705646
3 changed files with 158 additions and 158 deletions
+60 -60
View File
@@ -1,33 +1,33 @@
---
name: article-deep-summary
description: Generate a deep structured knowledge note from a single extracted article. Use when the user wants to deeply summarize, distill, or create a knowledge note from an article's full plain_text content stored in an *.extracted.json file.
description: 从单篇提取的文章中生成深度结构化知识笔记。当用户需要对存储在 *.extracted.json 文件中的文章 plain_text 内容进行深度摘要、提炼或创建知识笔记时使用。
---
# Article Deep Summary
# 文章深度摘要
Use this skill when the user has an extracted article file and wants to generate a deep knowledge note — not a brief summary card, but a structured distillation with core conclusion, arguments, methods, details, and reusable insights.
当用户拥有已提取的文章文件并希望生成深度知识笔记时使用此技能——不是简短的摘要卡片,而是包含核心结论、论点、方法、细节和可复用洞见的结构化提炼。
This skill is **not** for validating or repairing existing summaries (use `llm-summary-review` for that), nor for keyword index cleanup (use `keyword-cleanup-review` for that).
此技能**不适用于**验证或修复已有摘要(请使用 `llm-summary-review`),也不适用于关键词索引清理(请使用 `keyword-cleanup-review`)。
## What This Skill Does
## 功能说明
- Reads an extracted article JSON containing `plain_text`
- Calls the LLM with the dedicated article-summary prompt
- Validates the LLM output against the `ArticleSummaryResult` schema
- Renders a structured Markdown knowledge note in Chinese
- 读取包含 `plain_text` 的已提取文章 JSON 文件
- 使用专用的文章摘要提示词调用 LLM
- 根据 `ArticleSummaryResult` 模式验证 LLM 输出
- 渲染结构化的中文 Markdown 知识笔记
## Inputs
## 输入
Typical files:
典型文件:
- Extracted article JSON: `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json` (single-item) or a batch file with a `results` array
- Prompt template: `outputs/prompts/article-summary-prompt.txt`
- 已提取的文章 JSON:`outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json`(单篇)或包含 `results` 数组的批量文件
- 提示词模板:`outputs/prompts/article-summary-prompt.txt`
## Workflow
## 工作流程
1. Identify the target extracted file and the item IDs to summarize.
1. 确定目标提取文件和需要摘要的条目 ID。
2. Run the article-summary workflow via CLI:
2. 通过 CLI 运行文章摘要工作流:
```bash
python -m summary_mcp.workflows.article_summary \
@@ -36,47 +36,47 @@ python -m summary_mcp.workflows.article_summary \
--output-dir <output_dir>
```
Or call the MCP tool `generate_article_summaries` with:
或调用 MCP 工具 `generate_article_summaries`,参数如下:
- `extracted_path`: path to the extracted JSON file
- `selected_ids`: array of item ID strings (pass empty array to summarize all)
- `output_dir`: (optional) directory for Markdown output
- `extracted_path`:已提取的 JSON 文件路径
- `selected_ids`:条目 ID 字符串数组(传入空数组可摘要全部条目)
- `output_dir`:(可选)Markdown 输出目录
3. The workflow internally:
- Resolves LLM settings (`ARTICLE_SUMMARY_*` env vars, falling back to `LLM_*`)
- Calls `run_loop_payload` with the article-summary prompt and validator
- Retries up to 2 times on validation failure
3. 工作流内部流程:
- 解析 LLM 配置(`ARTICLE_SUMMARY_*` 环境变量,未设置时回退到 `LLM_*`)
- 使用文章摘要提示词和验证器调用 `run_loop_payload`
- 验证失败时最多重试 2 次
4. Check the output Markdown files in the specified output directory.
4. 在指定的输出目录中检查生成的 Markdown 文件。
## Output Schema
## 输出模式
The LLM returns a JSON matching `ArticleSummaryResult`:
LLM 返回匹配 `ArticleSummaryResult` 的 JSON:
| Field | Type | Description |
| 字段 | 类型 | 说明 |
|---|---|---|
| `title` | string | Article title |
| `url` | HttpUrl | Article URL |
| `core_conclusion` | string | Author's core conclusion, 1-2 sentences |
| `main_argument` | string | Main argument or thesis, can be multi-sentence |
| `key_methods` | string[] | Key methods, mechanisms, or techniques |
| `important_details` | string[] | Noteworthy details, data points, or cases |
| `reusable_insights` | string[] | Reusable insights transferable to other contexts |
| `keywords` | string[] | Specific entities — tool names, frameworks, methods |
| `topics` | string[] | Higher-level topic labels |
| `category` | enum | One of: `资讯` `方法论` `工具实践` `观点评论` |
| `worth_keeping` | bool | Whether the article is worth long-term retention |
| `reason` | string | One-sentence justification for retention |
| `title` | string | 文章标题 |
| `url` | HttpUrl | 文章 URL |
| `core_conclusion` | string | 作者核心结论,1-2 句话 |
| `main_argument` | string | 主要论点或论题,可为多句 |
| `key_methods` | string[] | 关键方法、机制或技术 |
| `important_details` | string[] | 值得注意的细节、数据点或案例 |
| `reusable_insights` | string[] | 可迁移到其他场景的可复用洞见 |
| `keywords` | string[] | 具体实体——工具名称、框架、方法 |
| `topics` | string[] | 更高层次的主题标签 |
| `category` | enum | 取值之一:`资讯` `方法论` `工具实践` `观点评论` |
| `worth_keeping` | bool | 该文章是否值得长期保留 |
| `reason` | string | 一句话说明保留理由 |
Constraints enforced by the validator:
验证器强制约束:
- `keywords` and `topics` must not overlap
- `keywords` focuses on concrete entities; `topics` focuses on abstract themes
- `category` must be one of the four allowed values
- `keywords` 和 `topics` 不得重叠
- `keywords` 聚焦具体实体;`topics` 聚焦抽象主题
- `category` 必须为四个允许值之一
## Markdown Output
## Markdown 输出
Each article produces one `.md` file with sections:
每篇文章生成一个 `.md` 文件,包含以下章节:
- 核心结论
- 主要论点
@@ -86,28 +86,28 @@ Each article produces one `.md` file with sections:
- 关键词
- 主题
## LLM Configuration
## LLM 配置
Dedicated environment variables (fallback to main `LLM_*` if unset):
专用环境变量(未设置时回退到主 `LLM_*` 变量):
- `ARTICLE_SUMMARY_LLM_API_URL`
- `ARTICLE_SUMMARY_LLM_API_KEY`
- `ARTICLE_SUMMARY_LLM_MODEL`
## Repository Implementation
## 代码实现
Relevant code:
相关代码:
- Workflow: `src/summary_mcp/workflows/article_summary.py`
- Validator: `src/summary_mcp/validators/article_summary.py`
- Data model: `src/summary_mcp/models/article_summary_result.py`
- Prompt: `outputs/prompts/article-summary-prompt.txt`
- MCP tool: `generate_article_summaries` in `src/summary_mcp/server.py`
- 工作流:`src/summary_mcp/workflows/article_summary.py`
- 验证器:`src/summary_mcp/validators/article_summary.py`
- 数据模型:`src/summary_mcp/models/article_summary_result.py`
- 提示词:`outputs/prompts/article-summary-prompt.txt`
- MCP 工具:`generate_article_summaries`,位于 `src/summary_mcp/server.py`
## When To Stop
## 何时停止
Stop when one of these is true:
满足以下条件之一时停止:
- Markdown knowledge notes have been generated for all requested items
- The workflow reports repeated validation failures and the user should decide how to proceed
- The user asks to inspect intermediate results manually
- 所有请求条目的 Markdown 知识笔记已生成
- 工作流报告反复验证失败,应由用户决定后续操作
- 用户要求手动检查中间结果
+48 -48
View File
@@ -1,62 +1,62 @@
---
name: keyword-cleanup-review
description: Review and curate this repository's daily keyword index and frequency stats. Use when the user wants to inspect `data/term_index/term_stats.json`, recent `data/term_index/daily/*.json`, `configs/term_aliases.json`, `configs/term_stopwords.json`, or `configs/filter_context.personal.json` to propose alias merges, stopwords, watch terms, or `interest_keywords` updates without directly modifying configs.
description: 审查和整理本仓库的每日关键词索引和频率统计。当用户需要检查 `data/term_index/term_stats.json`、最近的 `data/term_index/daily/*.json`、`configs/term_aliases.json`、`configs/term_stopwords.json` 或 `configs/filter_context.personal.json`,以提议别名合并、停用词、关注词或 `interest_keywords` 更新(但不直接修改配置)时使用。
---
# Keyword Cleanup Review
# 关键词清理审查
Use this skill to turn the repository's keyword statistics into reviewable cleanup suggestions.
使用此技能将仓库的关键词统计转化为可审查的清理建议。
## Workflow
## 工作流程
1. Build a compact review bundle:
1. 构建精简的审查数据包:
```bash
python skills/keyword-cleanup-review/scripts/build_review_bundle.py
```
Optional knobs:
可选参数:
- `--days 7`
- `--top 50`
- `--output outputs/term_index/review/keyword-cleanup-bundle.json`
2. Read the generated bundle and the suggestion schema:
2. 阅读生成的数据包和建议模式:
- `outputs/term_index/review/keyword-cleanup-bundle.json`
- `skills/keyword-cleanup-review/references/suggestion-schema.md`
3. Produce two outputs:
3. 生成两份输出:
- A short Markdown review for humans
- A JSON suggestion file matching the schema
- 一份简短的供人工审阅的 Markdown 报告
- 一份符合模式的 JSON 建议文件
4. Keep the boundary strict:
4. 严格保持边界:
- Suggest changes to `configs/term_aliases.json`
- Suggest changes to `configs/term_stopwords.json`
- Suggest additions to `configs/filter_context.personal.json`
- Do not directly edit these files unless the user explicitly asks
- Do not suggest direct edits to `configs/filter_rules.json` unless the user asks for rule logic changes
- 建议 `configs/term_aliases.json` 的修改
- 建议 `configs/term_stopwords.json` 的修改
- 建议 `configs/filter_context.personal.json` 的新增
- 除非用户明确要求,否则不要直接编辑这些文件
- 除非用户要求修改规则逻辑,否则不要建议直接编辑 `configs/filter_rules.json`
## Review Heuristics
## 审查启发式规则
Prioritize these decisions:
优先考虑以下决策:
- Alias suggestion
- Same concept with different naming, casing, abbreviation, or Chinese/English variants
- Stopword suggestion
- Too generic, too broad, or too noisy to help filtering
- Interest keyword suggestion
- High-frequency and aligned with the user's backend engineering, AI-agent, and frontier-tech focus
- Watch term
- Recent and potentially important, but evidence is still weak
- 别名建议
- 同一概念的不同命名、大小写、缩写或中英文变体
- 停用词建议
- 过于通用、过于宽泛或噪声过大,对过滤无帮助
- 兴趣关键词建议
- 高频且与用户的后端工程、AI Agent 和前沿技术关注方向一致
- 关注词
- 近期出现且可能重要,但证据尚不充分
Prefer conservative suggestions. If confidence is low, put the term into `watch_terms`.
建议保守为主。如果置信度较低,将词放入 `watch_terms`。
## Inputs
## 输入
Primary inputs:
主要输入:
- `data/term_index/term_stats.json`
- `data/term_index/daily/*.json`
@@ -67,36 +67,36 @@ Primary inputs:
- `configs/term_watchlist.json`
- `configs/term_change_log.json`
The bundled script already compacts these into a single review bundle.
打包脚本已将这些内容压缩为单个审查数据包。
## Output Expectations
## 输出要求
The Markdown output should:
Markdown 输出应:
- Summarize the current state briefly
- List the top terms worth acting on
- Separate alias, stopword, interest-keyword, and watch-term recommendations
- Explain reasoning in short, concrete sentences
- 简要总结当前状态
- 列出值得处理的高频词
- 分类别名、停用词、兴趣关键词和关注词建议
- 用简短、具体的句子解释理由
The JSON output should follow:
JSON 输出应遵循:
- `references/suggestion-schema.md`
## Repository Notes
## 仓库说明
Current repository behavior:
当前仓库行为:
- Keyword stats are program-maintained, not LLM-maintained
- Stats are built from `keywords`, not `topics`
- Stats only include non-`drop` candidates
- `data/term_index/term_stats.json` is rebuilt from daily files, so reruns overwrite the same day instead of double-counting
- cleanup policy, watchlist, and change log are repository-managed governance inputs and should be respected during review
- 关键词统计由程序维护,而非 LLM 维护
- 统计基于 `keywords` 构建,而非 `topics`
- 统计仅包含非 `drop` 候选项
- `data/term_index/term_stats.json` 从每日文件重建,因此重新运行会覆盖同一天的数据而非重复计算
- 清理策略、关注列表和变更日志是仓库管理的治理输入,审查时应予以尊重
Keep suggestions aligned with that design.
保持建议与此设计保持一致。
## Resources
## 资源
- Script:
- 脚本:
- `scripts/build_review_bundle.py`
- Reference:
- 参考文档:
- `references/suggestion-schema.md`
+50 -50
View File
@@ -1,55 +1,55 @@
---
name: llm-summary-review
description: Validate and refine LLM-generated summary JSON for extracted articles in this repository. Use when the user wants to review, validate, repair, or iterate on `outputs/result.json` or similar summary outputs produced from extracted article JSON.
description: 验证和优化 LLM 生成的文章摘要 JSON。当用户需要审查、验证、修复或迭代 `outputs/result.json` 或其他由已提取文章 JSON 生成的摘要输出时使用。
---
# LLM Summary Review
# LLM 摘要审查
Use this skill when working on the repository's summary loop after article extraction is done.
在文章提取完成后,处理本仓库摘要循环时使用此技能。
## What This Skill Does
## 功能说明
- Verifies that an LLM summary result matches the expected JSON contract
- Reuses the repository validator instead of re-checking fields manually
- Repairs invalid outputs by telling the LLM exactly what to fix
- Keeps the workflow aligned with the extraction JSON produced by this project
- 验证 LLM 摘要结果是否符合预期的 JSON 约定
- 复用仓库验证器,而非手动逐字段检查
- 通过告知 LLM 具体需要修复的内容来修复无效输出
- 保持工作流与本项目生成的提取 JSON 保持一致
## Inputs
## 输入
Typical files:
典型文件:
- Extracted article JSON: `outputs/*.extracted.json`
- LLM summary result JSON: `outputs/result.json`
- Prompt template: `outputs/llm-summary-prompt.txt`
- 已提取的文章 JSON:`outputs/*.extracted.json`
- LLM 摘要结果 JSON:`outputs/result.json`
- 提示词模板:`outputs/llm-summary-prompt.txt`
## Workflow
## 工作流程
1. Validate the current summary result with the repository validator:
1. 使用仓库验证器验证当前摘要结果:
```bash
python -m summary_mcp.validate_llm_result outputs/result.json --extracted outputs/read-flow-2026.extracted.json
```
2. If validation passes:
- Report that the result is structurally valid
- Briefly note any warnings
- Do not rewrite the result unless the user asks
2. 如果验证通过:
- 报告结果结构有效
- 简要说明任何警告
- 除非用户要求,否则不重写结果
3. If validation fails:
- Read the validator errors carefully
- Ask the LLM to regenerate or repair only the failing parts
- Re-run the validator until it passes or a retry limit is hit
3. 如果验证失败:
- 仔细阅读验证器错误信息
- 要求 LLM 仅重新生成或修复失败的部分
- 重新运行验证器,直到通过或达到重试上限
## Repair Prompt Pattern
## 修复提示词模式
When asking an LLM to repair a bad result, provide:
当要求 LLM 修复有问题的结果时,需提供:
- The original extracted article JSON
- The current invalid summary JSON
- The validator error list
- A strict instruction to preserve valid fields and fix only the failing ones
- 原始已提取的文章 JSON
- 当前无效的摘要 JSON
- 验证器错误列表
- 严格指令:保留正确字段,仅修复失败字段
Use this repair template:
使用以下修复模板:
```text
请修复下面这份不符合要求的摘要 JSON。
@@ -70,32 +70,32 @@ validator errors:
{{result_json}}
```
## Validation Rules
## 验证规则
The validator currently enforces:
验证器当前强制执行以下规则:
- Required fields exist
- Field types are correct
- `category` is one of: `资讯` `方法论` `工具实践` `观点评论`
- `summary` length is within bounds
- `highlights`, `keywords`, and `topics` counts are within bounds
- `keywords` and `topics` do not overlap
- `title` and `url` match the extracted article when an extracted JSON file is provided
- 必填字段存在
- 字段类型正确
- `category` 取值为以下之一:`资讯` `方法论` `工具实践` `观点评论`
- `summary` 长度在允许范围内
- `highlights`、`keywords` 和 `topics` 数量在允许范围内
- `keywords` 和 `topics` 不重叠
- 提供已提取 JSON 文件时,`title` 和 `url` 与已提取文章匹配
## Repository Implementation
## 代码实现
Relevant code:
相关代码:
- Validator model: `src/summary_mcp/models/llm_result.py`
- Validator logic: `src/summary_mcp/validators/llm_result.py`
- CLI entry: `src/summary_mcp/validate_llm_result.py`
- 验证器模型:`src/summary_mcp/models/llm_result.py`
- 验证器逻辑:`src/summary_mcp/validators/llm_result.py`
- CLI 入口:`src/summary_mcp/validate_llm_result.py`
Prefer using the existing validator rather than recreating checks in free-form reasoning.
优先使用现有验证器,而非在自由推理中重新创建检查逻辑。
## When To Stop
## 何时停止
Stop when one of these is true:
满足以下条件之一时停止:
- The validator returns `valid: true`
- The user asks to inspect the remaining failures manually
- Repeated retries fail and the user should decide how to proceed
- 验证器返回 `valid: true`
- 用户要求手动检查剩余的失败项
- 反复重试失败,应由用户决定后续操作