Files
reader/skills/reader-digest-flow/references/content-extraction.md
T
root 5eb390e3ed docs: 添加 Agent Skill 到项目仓库
- 复制 reader-digest-flow skill 到 skills/ 目录(含 SKILL.md + references/)
- README 新增 Agent Skill 章节说明供 Agent 使用的工作流
2026-07-28 18:38:54 +08:00

49 lines
2.8 KiB
Markdown

# Content Extraction Pipeline
How the pipeline turns FreshRSS items into extractable article text.
## Core Rule: FreshRSS items never re-fetch the original URL
**FreshRSS is an RSS-only upstream.** The pipeline *never* makes an HTTP request to the original article URL for a FreshRSS item. This is enforced by `RSS_ONLY_UPSTREAMS = {"freshrss"}` in `pipeline.py`.
The only exception: non-FreshRSS upstreams (future sources that don't set `upstream: freshrss`) may trigger `fetch_html()` as a fallback.
## Content source priority chain
`content_loader.py` tries sources in this order, using the **first one with ≥500 readable characters**:
| Priority | Source | Meaning |
|----------|--------|---------|
| 1 | `raw_html` | HTML pre-injected via `ExtractionInput.raw_html`. Rarely used in normal FreshRSS runs. |
| 2 | `item.raw_content` | RSS `<content:encoded>` — the full article body. Some feeds provide this; many don't. |
| 3 | `item.raw_summary` | RSS `<description>` — the summary/snippet field. **This is the most common source in current runs.** |
| 4 | `rss_content` | RSS content from non-item sources. |
| — | `none` | Nothing usable → raises `RSS_CONTENT_MISSING` for FreshRSS items (because fetch is skipped). |
## How `content_source` maps to actual text quality
The `content_source` field in every `item-XX.extracted.json` tells you what the pipeline actually used:
- **`item.raw_content`** → Full article text from RSS `<content:encoded>`. Best quality, same as reading the original page.
- **`item.raw_summary`** → RSS summary/description only. **Not the full article.** Length varies wildly by source (300-2000 chars typical). The AI summary is based on this snippet, not the complete text.
- **`rss_content`** → From standalone RSS content. Quality depends on the feed.
- **`fetched_html`** → HTML fetched from the original URL (**never happens for FreshRSS**; only for non-FreshRSS upstreams).
## What this means for digest quality
If you see `content_source: item.raw_summary` in the extracted files (current norm), the AI is summarizing from a **feed summary/snippet**, not the full article body. Articles that seem shallow in the daily digest may simply have short RSS descriptions.
To improve quality: either find feeds that provide full `<content:encoded>`, or switch the feed source to a system that provides full-text RSS (e.g., RSS-proxy with full-text extraction, or a third-party service like FiveFilters).
## Quick check
```bash
# Check content_source for latest run
grep -h "content_source" outputs/freshrss/rerun/*/extracted/item-*.json | sort | uniq -c
```
## Relevant code paths
- `src/summary_mcp/core/pipeline.py` — `RSS_ONLY_UPSTREAMS`, `_should_skip_fetch()`, `extract_content()`
- `src/summary_mcp/core/content_loader.py` — `choose_inline_content()` priority chain, `fetch_html()` (never called for FreshRSS)