Translate skill docs (article-deep-summary, llm-summary-review, keyword-cleanup-review) from English to Chinese

This commit is contained in:
wdm
2026-03-30 23:14:22 +08:00
parent 5b8df317ef
commit ca86705646
3 changed files with 158 additions and 158 deletions
+60 -60
View File
@@ -1,33 +1,33 @@
--- ---
name: article-deep-summary name: article-deep-summary
description: Generate a deep structured knowledge note from a single extracted article. Use when the user wants to deeply summarize, distill, or create a knowledge note from an article's full plain_text content stored in an *.extracted.json file. description: 从单篇提取的文章中生成深度结构化知识笔记。当用户需要对存储在 *.extracted.json 文件中的文章 plain_text 内容进行深度摘要、提炼或创建知识笔记时使用。
--- ---
# Article Deep Summary # 文章深度摘要
Use this skill when the user has an extracted article file and wants to generate a deep knowledge note — not a brief summary card, but a structured distillation with core conclusion, arguments, methods, details, and reusable insights. 当用户拥有已提取的文章文件并希望生成深度知识笔记时使用此技能——不是简短的摘要卡片,而是包含核心结论、论点、方法、细节和可复用洞见的结构化提炼。
This skill is **not** for validating or repairing existing summaries (use `llm-summary-review` for that), nor for keyword index cleanup (use `keyword-cleanup-review` for that). 此技能**不适用于**验证或修复已有摘要(请使用 `llm-summary-review`),也不适用于关键词索引清理(请使用 `keyword-cleanup-review`)。
## What This Skill Does ## 功能说明
- Reads an extracted article JSON containing `plain_text` - 读取包含 `plain_text` 的已提取文章 JSON 文件
- Calls the LLM with the dedicated article-summary prompt - 使用专用的文章摘要提示词调用 LLM
- Validates the LLM output against the `ArticleSummaryResult` schema - 根据 `ArticleSummaryResult` 模式验证 LLM 输出
- Renders a structured Markdown knowledge note in Chinese - 渲染结构化的中文 Markdown 知识笔记
## Inputs ## 输入
Typical files: 典型文件:
- Extracted article JSON: `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json` (single-item) or a batch file with a `results` array - 已提取的文章 JSON:`outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json`(单篇)或包含 `results` 数组的批量文件
- Prompt template: `outputs/prompts/article-summary-prompt.txt` - 提示词模板:`outputs/prompts/article-summary-prompt.txt`
## Workflow ## 工作流程
1. Identify the target extracted file and the item IDs to summarize. 1. 确定目标提取文件和需要摘要的条目 ID。
2. Run the article-summary workflow via CLI: 2. 通过 CLI 运行文章摘要工作流:
```bash ```bash
python -m summary_mcp.workflows.article_summary \ python -m summary_mcp.workflows.article_summary \
@@ -36,47 +36,47 @@ python -m summary_mcp.workflows.article_summary \
--output-dir <output_dir> --output-dir <output_dir>
``` ```
Or call the MCP tool `generate_article_summaries` with: 或调用 MCP 工具 `generate_article_summaries`,参数如下:
- `extracted_path`: path to the extracted JSON file - `extracted_path`:已提取的 JSON 文件路径
- `selected_ids`: array of item ID strings (pass empty array to summarize all) - `selected_ids`:条目 ID 字符串数组(传入空数组可摘要全部条目)
- `output_dir`: (optional) directory for Markdown output - `output_dir`:(可选)Markdown 输出目录
3. The workflow internally: 3. 工作流内部流程:
- Resolves LLM settings (`ARTICLE_SUMMARY_*` env vars, falling back to `LLM_*`) - 解析 LLM 配置(`ARTICLE_SUMMARY_*` 环境变量,未设置时回退到 `LLM_*`)
- Calls `run_loop_payload` with the article-summary prompt and validator - 使用文章摘要提示词和验证器调用 `run_loop_payload`
- Retries up to 2 times on validation failure - 验证失败时最多重试 2 次
4. Check the output Markdown files in the specified output directory. 4. 在指定的输出目录中检查生成的 Markdown 文件。
## Output Schema ## 输出模式
The LLM returns a JSON matching `ArticleSummaryResult`: LLM 返回匹配 `ArticleSummaryResult` 的 JSON:
| Field | Type | Description | | 字段 | 类型 | 说明 |
|---|---|---| |---|---|---|
| `title` | string | Article title | | `title` | string | 文章标题 |
| `url` | HttpUrl | Article URL | | `url` | HttpUrl | 文章 URL |
| `core_conclusion` | string | Author's core conclusion, 1-2 sentences | | `core_conclusion` | string | 作者核心结论,1-2 句话 |
| `main_argument` | string | Main argument or thesis, can be multi-sentence | | `main_argument` | string | 主要论点或论题,可为多句 |
| `key_methods` | string[] | Key methods, mechanisms, or techniques | | `key_methods` | string[] | 关键方法、机制或技术 |
| `important_details` | string[] | Noteworthy details, data points, or cases | | `important_details` | string[] | 值得注意的细节、数据点或案例 |
| `reusable_insights` | string[] | Reusable insights transferable to other contexts | | `reusable_insights` | string[] | 可迁移到其他场景的可复用洞见 |
| `keywords` | string[] | Specific entities — tool names, frameworks, methods | | `keywords` | string[] | 具体实体——工具名称、框架、方法 |
| `topics` | string[] | Higher-level topic labels | | `topics` | string[] | 更高层次的主题标签 |
| `category` | enum | One of: `资讯` `方法论` `工具实践` `观点评论` | | `category` | enum | 取值之一:`资讯` `方法论` `工具实践` `观点评论` |
| `worth_keeping` | bool | Whether the article is worth long-term retention | | `worth_keeping` | bool | 该文章是否值得长期保留 |
| `reason` | string | One-sentence justification for retention | | `reason` | string | 一句话说明保留理由 |
Constraints enforced by the validator: 验证器强制约束:
- `keywords` and `topics` must not overlap - `keywords` 和 `topics` 不得重叠
- `keywords` focuses on concrete entities; `topics` focuses on abstract themes - `keywords` 聚焦具体实体;`topics` 聚焦抽象主题
- `category` must be one of the four allowed values - `category` 必须为四个允许值之一
## Markdown Output ## Markdown 输出
Each article produces one `.md` file with sections: 每篇文章生成一个 `.md` 文件,包含以下章节:
- 核心结论 - 核心结论
- 主要论点 - 主要论点
@@ -86,28 +86,28 @@ Each article produces one `.md` file with sections:
- 关键词 - 关键词
- 主题 - 主题
## LLM Configuration ## LLM 配置
Dedicated environment variables (fallback to main `LLM_*` if unset): 专用环境变量(未设置时回退到主 `LLM_*` 变量):
- `ARTICLE_SUMMARY_LLM_API_URL` - `ARTICLE_SUMMARY_LLM_API_URL`
- `ARTICLE_SUMMARY_LLM_API_KEY` - `ARTICLE_SUMMARY_LLM_API_KEY`
- `ARTICLE_SUMMARY_LLM_MODEL` - `ARTICLE_SUMMARY_LLM_MODEL`
## Repository Implementation ## 代码实现
Relevant code: 相关代码:
- Workflow: `src/summary_mcp/workflows/article_summary.py` - 工作流:`src/summary_mcp/workflows/article_summary.py`
- Validator: `src/summary_mcp/validators/article_summary.py` - 验证器:`src/summary_mcp/validators/article_summary.py`
- Data model: `src/summary_mcp/models/article_summary_result.py` - 数据模型:`src/summary_mcp/models/article_summary_result.py`
- Prompt: `outputs/prompts/article-summary-prompt.txt` - 提示词:`outputs/prompts/article-summary-prompt.txt`
- MCP tool: `generate_article_summaries` in `src/summary_mcp/server.py` - MCP 工具:`generate_article_summaries`,位于 `src/summary_mcp/server.py`
## When To Stop ## 何时停止
Stop when one of these is true: 满足以下条件之一时停止:
- Markdown knowledge notes have been generated for all requested items - 所有请求条目的 Markdown 知识笔记已生成
- The workflow reports repeated validation failures and the user should decide how to proceed - 工作流报告反复验证失败,应由用户决定后续操作
- The user asks to inspect intermediate results manually - 用户要求手动检查中间结果
+48 -48
View File
@@ -1,62 +1,62 @@
--- ---
name: keyword-cleanup-review name: keyword-cleanup-review
description: Review and curate this repository's daily keyword index and frequency stats. Use when the user wants to inspect `data/term_index/term_stats.json`, recent `data/term_index/daily/*.json`, `configs/term_aliases.json`, `configs/term_stopwords.json`, or `configs/filter_context.personal.json` to propose alias merges, stopwords, watch terms, or `interest_keywords` updates without directly modifying configs. description: 审查和整理本仓库的每日关键词索引和频率统计。当用户需要检查 `data/term_index/term_stats.json`、最近的 `data/term_index/daily/*.json`、`configs/term_aliases.json`、`configs/term_stopwords.json` 或 `configs/filter_context.personal.json`,以提议别名合并、停用词、关注词或 `interest_keywords` 更新(但不直接修改配置)时使用。
--- ---
# Keyword Cleanup Review # 关键词清理审查
Use this skill to turn the repository's keyword statistics into reviewable cleanup suggestions. 使用此技能将仓库的关键词统计转化为可审查的清理建议。
## Workflow ## 工作流程
1. Build a compact review bundle: 1. 构建精简的审查数据包:
```bash ```bash
python skills/keyword-cleanup-review/scripts/build_review_bundle.py python skills/keyword-cleanup-review/scripts/build_review_bundle.py
``` ```
Optional knobs: 可选参数:
- `--days 7` - `--days 7`
- `--top 50` - `--top 50`
- `--output outputs/term_index/review/keyword-cleanup-bundle.json` - `--output outputs/term_index/review/keyword-cleanup-bundle.json`
2. Read the generated bundle and the suggestion schema: 2. 阅读生成的数据包和建议模式:
- `outputs/term_index/review/keyword-cleanup-bundle.json` - `outputs/term_index/review/keyword-cleanup-bundle.json`
- `skills/keyword-cleanup-review/references/suggestion-schema.md` - `skills/keyword-cleanup-review/references/suggestion-schema.md`
3. Produce two outputs: 3. 生成两份输出:
- A short Markdown review for humans - 一份简短的供人工审阅的 Markdown 报告
- A JSON suggestion file matching the schema - 一份符合模式的 JSON 建议文件
4. Keep the boundary strict: 4. 严格保持边界:
- Suggest changes to `configs/term_aliases.json` - 建议 `configs/term_aliases.json` 的修改
- Suggest changes to `configs/term_stopwords.json` - 建议 `configs/term_stopwords.json` 的修改
- Suggest additions to `configs/filter_context.personal.json` - 建议 `configs/filter_context.personal.json` 的新增
- Do not directly edit these files unless the user explicitly asks - 除非用户明确要求,否则不要直接编辑这些文件
- Do not suggest direct edits to `configs/filter_rules.json` unless the user asks for rule logic changes - 除非用户要求修改规则逻辑,否则不要建议直接编辑 `configs/filter_rules.json`
## Review Heuristics ## 审查启发式规则
Prioritize these decisions: 优先考虑以下决策:
- Alias suggestion - 别名建议
- Same concept with different naming, casing, abbreviation, or Chinese/English variants - 同一概念的不同命名、大小写、缩写或中英文变体
- Stopword suggestion - 停用词建议
- Too generic, too broad, or too noisy to help filtering - 过于通用、过于宽泛或噪声过大,对过滤无帮助
- Interest keyword suggestion - 兴趣关键词建议
- High-frequency and aligned with the user's backend engineering, AI-agent, and frontier-tech focus - 高频且与用户的后端工程、AI Agent 和前沿技术关注方向一致
- Watch term - 关注词
- Recent and potentially important, but evidence is still weak - 近期出现且可能重要,但证据尚不充分
Prefer conservative suggestions. If confidence is low, put the term into `watch_terms`. 建议保守为主。如果置信度较低,将词放入 `watch_terms`。
## Inputs ## 输入
Primary inputs: 主要输入:
- `data/term_index/term_stats.json` - `data/term_index/term_stats.json`
- `data/term_index/daily/*.json` - `data/term_index/daily/*.json`
@@ -67,36 +67,36 @@ Primary inputs:
- `configs/term_watchlist.json` - `configs/term_watchlist.json`
- `configs/term_change_log.json` - `configs/term_change_log.json`
The bundled script already compacts these into a single review bundle. 打包脚本已将这些内容压缩为单个审查数据包。
## Output Expectations ## 输出要求
The Markdown output should: Markdown 输出应:
- Summarize the current state briefly - 简要总结当前状态
- List the top terms worth acting on - 列出值得处理的高频词
- Separate alias, stopword, interest-keyword, and watch-term recommendations - 分类别名、停用词、兴趣关键词和关注词建议
- Explain reasoning in short, concrete sentences - 用简短、具体的句子解释理由
The JSON output should follow: JSON 输出应遵循:
- `references/suggestion-schema.md` - `references/suggestion-schema.md`
## Repository Notes ## 仓库说明
Current repository behavior: 当前仓库行为:
- Keyword stats are program-maintained, not LLM-maintained - 关键词统计由程序维护,而非 LLM 维护
- Stats are built from `keywords`, not `topics` - 统计基于 `keywords` 构建,而非 `topics`
- Stats only include non-`drop` candidates - 统计仅包含非 `drop` 候选项
- `data/term_index/term_stats.json` is rebuilt from daily files, so reruns overwrite the same day instead of double-counting - `data/term_index/term_stats.json` 从每日文件重建,因此重新运行会覆盖同一天的数据而非重复计算
- cleanup policy, watchlist, and change log are repository-managed governance inputs and should be respected during review - 清理策略、关注列表和变更日志是仓库管理的治理输入,审查时应予以尊重
Keep suggestions aligned with that design. 保持建议与此设计保持一致。
## Resources ## 资源
- Script: - 脚本:
- `scripts/build_review_bundle.py` - `scripts/build_review_bundle.py`
- Reference: - 参考文档:
- `references/suggestion-schema.md` - `references/suggestion-schema.md`
+50 -50
View File
@@ -1,55 +1,55 @@
--- ---
name: llm-summary-review name: llm-summary-review
description: Validate and refine LLM-generated summary JSON for extracted articles in this repository. Use when the user wants to review, validate, repair, or iterate on `outputs/result.json` or similar summary outputs produced from extracted article JSON. description: 验证和优化 LLM 生成的文章摘要 JSON。当用户需要审查、验证、修复或迭代 `outputs/result.json` 或其他由已提取文章 JSON 生成的摘要输出时使用。
--- ---
# LLM Summary Review # LLM 摘要审查
Use this skill when working on the repository's summary loop after article extraction is done. 在文章提取完成后,处理本仓库摘要循环时使用此技能。
## What This Skill Does ## 功能说明
- Verifies that an LLM summary result matches the expected JSON contract - 验证 LLM 摘要结果是否符合预期的 JSON 约定
- Reuses the repository validator instead of re-checking fields manually - 复用仓库验证器,而非手动逐字段检查
- Repairs invalid outputs by telling the LLM exactly what to fix - 通过告知 LLM 具体需要修复的内容来修复无效输出
- Keeps the workflow aligned with the extraction JSON produced by this project - 保持工作流与本项目生成的提取 JSON 保持一致
## Inputs ## 输入
Typical files: 典型文件:
- Extracted article JSON: `outputs/*.extracted.json` - 已提取的文章 JSON:`outputs/*.extracted.json`
- LLM summary result JSON: `outputs/result.json` - LLM 摘要结果 JSON:`outputs/result.json`
- Prompt template: `outputs/llm-summary-prompt.txt` - 提示词模板:`outputs/llm-summary-prompt.txt`
## Workflow ## 工作流程
1. Validate the current summary result with the repository validator: 1. 使用仓库验证器验证当前摘要结果:
```bash ```bash
python -m summary_mcp.validate_llm_result outputs/result.json --extracted outputs/read-flow-2026.extracted.json python -m summary_mcp.validate_llm_result outputs/result.json --extracted outputs/read-flow-2026.extracted.json
``` ```
2. If validation passes: 2. 如果验证通过:
- Report that the result is structurally valid - 报告结果结构有效
- Briefly note any warnings - 简要说明任何警告
- Do not rewrite the result unless the user asks - 除非用户要求,否则不重写结果
3. If validation fails: 3. 如果验证失败:
- Read the validator errors carefully - 仔细阅读验证器错误信息
- Ask the LLM to regenerate or repair only the failing parts - 要求 LLM 仅重新生成或修复失败的部分
- Re-run the validator until it passes or a retry limit is hit - 重新运行验证器,直到通过或达到重试上限
## Repair Prompt Pattern ## 修复提示词模式
When asking an LLM to repair a bad result, provide: 当要求 LLM 修复有问题的结果时,需提供:
- The original extracted article JSON - 原始已提取的文章 JSON
- The current invalid summary JSON - 当前无效的摘要 JSON
- The validator error list - 验证器错误列表
- A strict instruction to preserve valid fields and fix only the failing ones - 严格指令:保留正确字段,仅修复失败字段
Use this repair template: 使用以下修复模板:
```text ```text
请修复下面这份不符合要求的摘要 JSON。 请修复下面这份不符合要求的摘要 JSON。
@@ -70,32 +70,32 @@ validator errors:
{{result_json}} {{result_json}}
``` ```
## Validation Rules ## 验证规则
The validator currently enforces: 验证器当前强制执行以下规则:
- Required fields exist - 必填字段存在
- Field types are correct - 字段类型正确
- `category` is one of: `资讯` `方法论` `工具实践` `观点评论` - `category` 取值为以下之一:`资讯` `方法论` `工具实践` `观点评论`
- `summary` length is within bounds - `summary` 长度在允许范围内
- `highlights`, `keywords`, and `topics` counts are within bounds - `highlights`、`keywords` 和 `topics` 数量在允许范围内
- `keywords` and `topics` do not overlap - `keywords` 和 `topics` 不重叠
- `title` and `url` match the extracted article when an extracted JSON file is provided - 提供已提取 JSON 文件时,`title` 和 `url` 与已提取文章匹配
## Repository Implementation ## 代码实现
Relevant code: 相关代码:
- Validator model: `src/summary_mcp/models/llm_result.py` - 验证器模型:`src/summary_mcp/models/llm_result.py`
- Validator logic: `src/summary_mcp/validators/llm_result.py` - 验证器逻辑:`src/summary_mcp/validators/llm_result.py`
- CLI entry: `src/summary_mcp/validate_llm_result.py` - CLI 入口:`src/summary_mcp/validate_llm_result.py`
Prefer using the existing validator rather than recreating checks in free-form reasoning. 优先使用现有验证器,而非在自由推理中重新创建检查逻辑。
## When To Stop ## 何时停止
Stop when one of these is true: 满足以下条件之一时停止:
- The validator returns `valid: true` - 验证器返回 `valid: true`
- The user asks to inspect the remaining failures manually - 用户要求手动检查剩余的失败项
- Repeated retries fail and the user should decide how to proceed - 反复重试失败,应由用户决定后续操作