Add article-deep-summary skill and enrich article summary prompt

- Create skills/article-deep-summary/SKILL.md and agents/openai.yaml
- Relax prompt constraints to allow 2-4 sentence elaboration per field
- Add validation output for reference article
- Update plan checklist as completed
This commit is contained in:
zhuyongxin
2026-03-30 18:44:30 +08:00
parent e24db99baa
commit 5b8df317ef
6 changed files with 160 additions and 15 deletions
+113
View File
@@ -0,0 +1,113 @@
---
name: article-deep-summary
description: Generate a deep structured knowledge note from a single extracted article. Use when the user wants to deeply summarize, distill, or create a knowledge note from an article's full plain_text content stored in an *.extracted.json file.
---
# Article Deep Summary
Use this skill when the user has an extracted article file and wants to generate a deep knowledge note — not a brief summary card, but a structured distillation with core conclusion, arguments, methods, details, and reusable insights.
This skill is **not** for validating or repairing existing summaries (use `llm-summary-review` for that), nor for keyword index cleanup (use `keyword-cleanup-review` for that).
## What This Skill Does
- Reads an extracted article JSON containing `plain_text`
- Calls the LLM with the dedicated article-summary prompt
- Validates the LLM output against the `ArticleSummaryResult` schema
- Renders a structured Markdown knowledge note in Chinese
## Inputs
Typical files:
- Extracted article JSON: `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json` (single-item) or a batch file with a `results` array
- Prompt template: `outputs/prompts/article-summary-prompt.txt`
## Workflow
1. Identify the target extracted file and the item IDs to summarize.
2. Run the article-summary workflow via CLI:
```bash
python -m summary_mcp.workflows.article_summary \
--extracted <extracted_path> \
--selected-ids <item_id> \
--output-dir <output_dir>
```
Or call the MCP tool `generate_article_summaries` with:
- `extracted_path`: path to the extracted JSON file
- `selected_ids`: array of item ID strings (pass empty array to summarize all)
- `output_dir`: (optional) directory for Markdown output
3. The workflow internally:
- Resolves LLM settings (`ARTICLE_SUMMARY_*` env vars, falling back to `LLM_*`)
- Calls `run_loop_payload` with the article-summary prompt and validator
- Retries up to 2 times on validation failure
4. Check the output Markdown files in the specified output directory.
## Output Schema
The LLM returns a JSON matching `ArticleSummaryResult`:
| Field | Type | Description |
|---|---|---|
| `title` | string | Article title |
| `url` | HttpUrl | Article URL |
| `core_conclusion` | string | Author's core conclusion, 1-2 sentences |
| `main_argument` | string | Main argument or thesis, can be multi-sentence |
| `key_methods` | string[] | Key methods, mechanisms, or techniques |
| `important_details` | string[] | Noteworthy details, data points, or cases |
| `reusable_insights` | string[] | Reusable insights transferable to other contexts |
| `keywords` | string[] | Specific entities — tool names, frameworks, methods |
| `topics` | string[] | Higher-level topic labels |
| `category` | enum | One of: `资讯` `方法论` `工具实践` `观点评论` |
| `worth_keeping` | bool | Whether the article is worth long-term retention |
| `reason` | string | One-sentence justification for retention |
Constraints enforced by the validator:
- `keywords` and `topics` must not overlap
- `keywords` focuses on concrete entities; `topics` focuses on abstract themes
- `category` must be one of the four allowed values
## Markdown Output
Each article produces one `.md` file with sections:
- 核心结论
- 主要论点
- 关键方法 / 机制
- 重要细节
- 可复用启发
- 关键词
- 主题
## LLM Configuration
Dedicated environment variables (fallback to main `LLM_*` if unset):
- `ARTICLE_SUMMARY_LLM_API_URL`
- `ARTICLE_SUMMARY_LLM_API_KEY`
- `ARTICLE_SUMMARY_LLM_MODEL`
## Repository Implementation
Relevant code:
- Workflow: `src/summary_mcp/workflows/article_summary.py`
- Validator: `src/summary_mcp/validators/article_summary.py`
- Data model: `src/summary_mcp/models/article_summary_result.py`
- Prompt: `outputs/prompts/article-summary-prompt.txt`
- MCP tool: `generate_article_summaries` in `src/summary_mcp/server.py`
## When To Stop
Stop when one of these is true:
- Markdown knowledge notes have been generated for all requested items
- The workflow reports repeated validation failures and the user should decide how to proceed
- The user asks to inspect intermediate results manually
@@ -0,0 +1,3 @@
display_name: Article Deep Summary
short_description: Generate deep knowledge notes from a single extracted article.
default_prompt: Given an extracted article JSON with plain_text content, run the article-summary workflow to produce a structured Markdown knowledge note with core conclusion, main arguments, key methods, important details, and reusable insights.