- Create skills/article-deep-summary/SKILL.md and agents/openai.yaml - Relax prompt constraints to allow 2-4 sentence elaboration per field - Add validation output for reference article - Update plan checklist as completed
114 lines
4.1 KiB
Markdown
114 lines
4.1 KiB
Markdown
---
|
|
name: article-deep-summary
|
|
description: Generate a deep structured knowledge note from a single extracted article. Use when the user wants to deeply summarize, distill, or create a knowledge note from an article's full plain_text content stored in an *.extracted.json file.
|
|
---
|
|
|
|
# Article Deep Summary
|
|
|
|
Use this skill when the user has an extracted article file and wants to generate a deep knowledge note — not a brief summary card, but a structured distillation with core conclusion, arguments, methods, details, and reusable insights.
|
|
|
|
This skill is **not** for validating or repairing existing summaries (use `llm-summary-review` for that), nor for keyword index cleanup (use `keyword-cleanup-review` for that).
|
|
|
|
## What This Skill Does
|
|
|
|
- Reads an extracted article JSON containing `plain_text`
|
|
- Calls the LLM with the dedicated article-summary prompt
|
|
- Validates the LLM output against the `ArticleSummaryResult` schema
|
|
- Renders a structured Markdown knowledge note in Chinese
|
|
|
|
## Inputs
|
|
|
|
Typical files:
|
|
|
|
- Extracted article JSON: `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json` (single-item) or a batch file with a `results` array
|
|
- Prompt template: `outputs/prompts/article-summary-prompt.txt`
|
|
|
|
## Workflow
|
|
|
|
1. Identify the target extracted file and the item IDs to summarize.
|
|
|
|
2. Run the article-summary workflow via CLI:
|
|
|
|
```bash
|
|
python -m summary_mcp.workflows.article_summary \
|
|
--extracted <extracted_path> \
|
|
--selected-ids <item_id> \
|
|
--output-dir <output_dir>
|
|
```
|
|
|
|
Or call the MCP tool `generate_article_summaries` with:
|
|
|
|
- `extracted_path`: path to the extracted JSON file
|
|
- `selected_ids`: array of item ID strings (pass empty array to summarize all)
|
|
- `output_dir`: (optional) directory for Markdown output
|
|
|
|
3. The workflow internally:
|
|
- Resolves LLM settings (`ARTICLE_SUMMARY_*` env vars, falling back to `LLM_*`)
|
|
- Calls `run_loop_payload` with the article-summary prompt and validator
|
|
- Retries up to 2 times on validation failure
|
|
|
|
4. Check the output Markdown files in the specified output directory.
|
|
|
|
## Output Schema
|
|
|
|
The LLM returns a JSON matching `ArticleSummaryResult`:
|
|
|
|
| Field | Type | Description |
|
|
|---|---|---|
|
|
| `title` | string | Article title |
|
|
| `url` | HttpUrl | Article URL |
|
|
| `core_conclusion` | string | Author's core conclusion, 1-2 sentences |
|
|
| `main_argument` | string | Main argument or thesis, can be multi-sentence |
|
|
| `key_methods` | string[] | Key methods, mechanisms, or techniques |
|
|
| `important_details` | string[] | Noteworthy details, data points, or cases |
|
|
| `reusable_insights` | string[] | Reusable insights transferable to other contexts |
|
|
| `keywords` | string[] | Specific entities — tool names, frameworks, methods |
|
|
| `topics` | string[] | Higher-level topic labels |
|
|
| `category` | enum | One of: `资讯` `方法论` `工具实践` `观点评论` |
|
|
| `worth_keeping` | bool | Whether the article is worth long-term retention |
|
|
| `reason` | string | One-sentence justification for retention |
|
|
|
|
Constraints enforced by the validator:
|
|
|
|
- `keywords` and `topics` must not overlap
|
|
- `keywords` focuses on concrete entities; `topics` focuses on abstract themes
|
|
- `category` must be one of the four allowed values
|
|
|
|
## Markdown Output
|
|
|
|
Each article produces one `.md` file with sections:
|
|
|
|
- 核心结论
|
|
- 主要论点
|
|
- 关键方法 / 机制
|
|
- 重要细节
|
|
- 可复用启发
|
|
- 关键词
|
|
- 主题
|
|
|
|
## LLM Configuration
|
|
|
|
Dedicated environment variables (fallback to main `LLM_*` if unset):
|
|
|
|
- `ARTICLE_SUMMARY_LLM_API_URL`
|
|
- `ARTICLE_SUMMARY_LLM_API_KEY`
|
|
- `ARTICLE_SUMMARY_LLM_MODEL`
|
|
|
|
## Repository Implementation
|
|
|
|
Relevant code:
|
|
|
|
- Workflow: `src/summary_mcp/workflows/article_summary.py`
|
|
- Validator: `src/summary_mcp/validators/article_summary.py`
|
|
- Data model: `src/summary_mcp/models/article_summary_result.py`
|
|
- Prompt: `outputs/prompts/article-summary-prompt.txt`
|
|
- MCP tool: `generate_article_summaries` in `src/summary_mcp/server.py`
|
|
|
|
## When To Stop
|
|
|
|
Stop when one of these is true:
|
|
|
|
- Markdown knowledge notes have been generated for all requested items
|
|
- The workflow reports repeated validation failures and the user should decide how to proceed
|
|
- The user asks to inspect intermediate results manually
|