- 复制 reader-digest-flow skill 到 skills/ 目录(含 SKILL.md + references/) - README 新增 Agent Skill 章节说明供 Agent 使用的工作流
49 lines
2.8 KiB
Markdown
49 lines
2.8 KiB
Markdown
# Content Extraction Pipeline
|
|
|
|
How the pipeline turns FreshRSS items into extractable article text.
|
|
|
|
## Core Rule: FreshRSS items never re-fetch the original URL
|
|
|
|
**FreshRSS is an RSS-only upstream.** The pipeline *never* makes an HTTP request to the original article URL for a FreshRSS item. This is enforced by `RSS_ONLY_UPSTREAMS = {"freshrss"}` in `pipeline.py`.
|
|
|
|
The only exception: non-FreshRSS upstreams (future sources that don't set `upstream: freshrss`) may trigger `fetch_html()` as a fallback.
|
|
|
|
## Content source priority chain
|
|
|
|
`content_loader.py` tries sources in this order, using the **first one with ≥500 readable characters**:
|
|
|
|
| Priority | Source | Meaning |
|
|
|----------|--------|---------|
|
|
| 1 | `raw_html` | HTML pre-injected via `ExtractionInput.raw_html`. Rarely used in normal FreshRSS runs. |
|
|
| 2 | `item.raw_content` | RSS `<content:encoded>` — the full article body. Some feeds provide this; many don't. |
|
|
| 3 | `item.raw_summary` | RSS `<description>` — the summary/snippet field. **This is the most common source in current runs.** |
|
|
| 4 | `rss_content` | RSS content from non-item sources. |
|
|
| — | `none` | Nothing usable → raises `RSS_CONTENT_MISSING` for FreshRSS items (because fetch is skipped). |
|
|
|
|
## How `content_source` maps to actual text quality
|
|
|
|
The `content_source` field in every `item-XX.extracted.json` tells you what the pipeline actually used:
|
|
|
|
- **`item.raw_content`** → Full article text from RSS `<content:encoded>`. Best quality, same as reading the original page.
|
|
- **`item.raw_summary`** → RSS summary/description only. **Not the full article.** Length varies wildly by source (300-2000 chars typical). The AI summary is based on this snippet, not the complete text.
|
|
- **`rss_content`** → From standalone RSS content. Quality depends on the feed.
|
|
- **`fetched_html`** → HTML fetched from the original URL (**never happens for FreshRSS**; only for non-FreshRSS upstreams).
|
|
|
|
## What this means for digest quality
|
|
|
|
If you see `content_source: item.raw_summary` in the extracted files (current norm), the AI is summarizing from a **feed summary/snippet**, not the full article body. Articles that seem shallow in the daily digest may simply have short RSS descriptions.
|
|
|
|
To improve quality: either find feeds that provide full `<content:encoded>`, or switch the feed source to a system that provides full-text RSS (e.g., RSS-proxy with full-text extraction, or a third-party service like FiveFilters).
|
|
|
|
## Quick check
|
|
|
|
```bash
|
|
# Check content_source for latest run
|
|
grep -h "content_source" outputs/freshrss/rerun/*/extracted/item-*.json | sort | uniq -c
|
|
```
|
|
|
|
## Relevant code paths
|
|
|
|
- `src/summary_mcp/core/pipeline.py` — `RSS_ONLY_UPSTREAMS`, `_should_skip_fetch()`, `extract_content()`
|
|
- `src/summary_mcp/core/content_loader.py` — `choose_inline_content()` priority chain, `fetch_html()` (never called for FreshRSS)
|