Add article-deep-summary skill and enrich article summary prompt
- Create skills/article-deep-summary/SKILL.md and agents/openai.yaml - Relax prompt constraints to allow 2-4 sentence elaboration per field - Add validation output for reference article - Update plan checklist as completed
This commit is contained in:
@@ -0,0 +1,113 @@
|
||||
---
|
||||
name: article-deep-summary
|
||||
description: Generate a deep structured knowledge note from a single extracted article. Use when the user wants to deeply summarize, distill, or create a knowledge note from an article's full plain_text content stored in an *.extracted.json file.
|
||||
---
|
||||
|
||||
# Article Deep Summary
|
||||
|
||||
Use this skill when the user has an extracted article file and wants to generate a deep knowledge note — not a brief summary card, but a structured distillation with core conclusion, arguments, methods, details, and reusable insights.
|
||||
|
||||
This skill is **not** for validating or repairing existing summaries (use `llm-summary-review` for that), nor for keyword index cleanup (use `keyword-cleanup-review` for that).
|
||||
|
||||
## What This Skill Does
|
||||
|
||||
- Reads an extracted article JSON containing `plain_text`
|
||||
- Calls the LLM with the dedicated article-summary prompt
|
||||
- Validates the LLM output against the `ArticleSummaryResult` schema
|
||||
- Renders a structured Markdown knowledge note in Chinese
|
||||
|
||||
## Inputs
|
||||
|
||||
Typical files:
|
||||
|
||||
- Extracted article JSON: `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json` (single-item) or a batch file with a `results` array
|
||||
- Prompt template: `outputs/prompts/article-summary-prompt.txt`
|
||||
|
||||
## Workflow
|
||||
|
||||
1. Identify the target extracted file and the item IDs to summarize.
|
||||
|
||||
2. Run the article-summary workflow via CLI:
|
||||
|
||||
```bash
|
||||
python -m summary_mcp.workflows.article_summary \
|
||||
--extracted <extracted_path> \
|
||||
--selected-ids <item_id> \
|
||||
--output-dir <output_dir>
|
||||
```
|
||||
|
||||
Or call the MCP tool `generate_article_summaries` with:
|
||||
|
||||
- `extracted_path`: path to the extracted JSON file
|
||||
- `selected_ids`: array of item ID strings (pass empty array to summarize all)
|
||||
- `output_dir`: (optional) directory for Markdown output
|
||||
|
||||
3. The workflow internally:
|
||||
- Resolves LLM settings (`ARTICLE_SUMMARY_*` env vars, falling back to `LLM_*`)
|
||||
- Calls `run_loop_payload` with the article-summary prompt and validator
|
||||
- Retries up to 2 times on validation failure
|
||||
|
||||
4. Check the output Markdown files in the specified output directory.
|
||||
|
||||
## Output Schema
|
||||
|
||||
The LLM returns a JSON matching `ArticleSummaryResult`:
|
||||
|
||||
| Field | Type | Description |
|
||||
|---|---|---|
|
||||
| `title` | string | Article title |
|
||||
| `url` | HttpUrl | Article URL |
|
||||
| `core_conclusion` | string | Author's core conclusion, 1-2 sentences |
|
||||
| `main_argument` | string | Main argument or thesis, can be multi-sentence |
|
||||
| `key_methods` | string[] | Key methods, mechanisms, or techniques |
|
||||
| `important_details` | string[] | Noteworthy details, data points, or cases |
|
||||
| `reusable_insights` | string[] | Reusable insights transferable to other contexts |
|
||||
| `keywords` | string[] | Specific entities — tool names, frameworks, methods |
|
||||
| `topics` | string[] | Higher-level topic labels |
|
||||
| `category` | enum | One of: `资讯` `方法论` `工具实践` `观点评论` |
|
||||
| `worth_keeping` | bool | Whether the article is worth long-term retention |
|
||||
| `reason` | string | One-sentence justification for retention |
|
||||
|
||||
Constraints enforced by the validator:
|
||||
|
||||
- `keywords` and `topics` must not overlap
|
||||
- `keywords` focuses on concrete entities; `topics` focuses on abstract themes
|
||||
- `category` must be one of the four allowed values
|
||||
|
||||
## Markdown Output
|
||||
|
||||
Each article produces one `.md` file with sections:
|
||||
|
||||
- 核心结论
|
||||
- 主要论点
|
||||
- 关键方法 / 机制
|
||||
- 重要细节
|
||||
- 可复用启发
|
||||
- 关键词
|
||||
- 主题
|
||||
|
||||
## LLM Configuration
|
||||
|
||||
Dedicated environment variables (fallback to main `LLM_*` if unset):
|
||||
|
||||
- `ARTICLE_SUMMARY_LLM_API_URL`
|
||||
- `ARTICLE_SUMMARY_LLM_API_KEY`
|
||||
- `ARTICLE_SUMMARY_LLM_MODEL`
|
||||
|
||||
## Repository Implementation
|
||||
|
||||
Relevant code:
|
||||
|
||||
- Workflow: `src/summary_mcp/workflows/article_summary.py`
|
||||
- Validator: `src/summary_mcp/validators/article_summary.py`
|
||||
- Data model: `src/summary_mcp/models/article_summary_result.py`
|
||||
- Prompt: `outputs/prompts/article-summary-prompt.txt`
|
||||
- MCP tool: `generate_article_summaries` in `src/summary_mcp/server.py`
|
||||
|
||||
## When To Stop
|
||||
|
||||
Stop when one of these is true:
|
||||
|
||||
- Markdown knowledge notes have been generated for all requested items
|
||||
- The workflow reports repeated validation failures and the user should decide how to proceed
|
||||
- The user asks to inspect intermediate results manually
|
||||
Reference in New Issue
Block a user