docs: update digest template and localize references

This commit is contained in:
zhuyongxin
2026-07-29 10:20:20 +08:00
parent 6dd8cef347
commit 807027976e
4 changed files with 43 additions and 53 deletions
@@ -1,45 +1,45 @@
# Content Extraction Pipeline
# 内容提取流程
How the pipeline turns FreshRSS items into extractable article text.
本文说明处理流水线如何把 FreshRSS 条目转换为可供摘要使用的文章文本。
## Core Rule: FreshRSS items never re-fetch the original URL
## 核心规则:FreshRSS 条目不重新抓取原文 URL
**FreshRSS is an RSS-only upstream.** The pipeline *never* makes an HTTP request to the original article URL for a FreshRSS item. This is enforced by `RSS_ONLY_UPSTREAMS = {"freshrss"}` in `pipeline.py`.
**FreshRSS 是仅提供 RSS 内容的上游。** 对于 FreshRSS 条目,流水线不会向文章原始 URL 发起 HTTP 请求。该行为由 `pipeline.py` 中的 `RSS_ONLY_UPSTREAMS = {"freshrss"}` 强制保证。
The only exception: non-FreshRSS upstreams (future sources that don't set `upstream: freshrss`) may trigger `fetch_html()` as a fallback.
唯一例外是非 FreshRSS 上游。未来未设置 `upstream: freshrss` 的其他来源,可以在必要时使用 `fetch_html()` 作为回退。
## Content source priority chain
## 内容来源优先级
`content_loader.py` tries sources in this order, using the **first one with ≥500 readable characters**:
`content_loader.py` 按以下顺序检查内容,并使用第一个包含 **至少 500 个可读字符** 的来源:
| Priority | Source | Meaning |
|----------|--------|---------|
| 1 | `raw_html` | HTML pre-injected via `ExtractionInput.raw_html`. Rarely used in normal FreshRSS runs. |
| 2 | `item.raw_content` | RSS `<content:encoded>` — the full article body. Some feeds provide this; many don't. |
| 3 | `item.raw_summary` | RSS `<description>` — the summary/snippet field. **This is the most common source in current runs.** |
| 4 | `rss_content` | RSS content from non-item sources. |
| — | `none` | Nothing usable → raises `RSS_CONTENT_MISSING` for FreshRSS items (because fetch is skipped). |
| 优先级 | 来源 | 含义 |
|--------|------|------|
| 1 | `raw_html` | 通过 `ExtractionInput.raw_html` 预先注入的 HTML;常规 FreshRSS 运行中很少使用。 |
| 2 | `item.raw_content` | RSS `<content:encoded>` 中的文章正文;部分订阅源提供,部分不提供。 |
| 3 | `item.raw_summary` | RSS `<description>` 中的摘要或片段;这是当前运行中最常见的来源。 |
| 4 | `rss_content` | 来自非条目字段的独立 RSS 内容。 |
| — | `none` | 没有可用内容;FreshRSS 不允许回源抓取,因此抛出 `RSS_CONTENT_MISSING`。 |
## How `content_source` maps to actual text quality
## `content_source` 与文本质量的关系
The `content_source` field in every `item-XX.extracted.json` tells you what the pipeline actually used:
每个 `item-XX.extracted.json` 中的 `content_source` 字段表示流水线实际使用的内容来源:
- **`item.raw_content`** → Full article text from RSS `<content:encoded>`. Best quality, same as reading the original page.
- **`item.raw_summary`** → RSS summary/description only. **Not the full article.** Length varies wildly by source (300-2000 chars typical). The AI summary is based on this snippet, not the complete text.
- **`rss_content`** → From standalone RSS content. Quality depends on the feed.
- **`fetched_html`** → HTML fetched from the original URL (**never happens for FreshRSS**; only for non-FreshRSS upstreams).
- **`item.raw_content`**:RSS `<content:encoded>` 提供的文章正文,通常质量最好,接近直接阅读原文。
- **`item.raw_summary`**:只有 RSS 摘要或描述,并非完整正文。不同来源长度差异较大,通常为 300-2000 个字符;AI 摘要基于该片段,而不是完整文章。
- **`rss_content`**:来自独立 RSS 内容,质量取决于订阅源。
- **`fetched_html`**:从原始 URL 抓取的 HTML。FreshRSS 条目不会出现该来源,只适用于非 FreshRSS 上游。
## What this means for digest quality
## 对日报质量的影响
If you see `content_source: item.raw_summary` in the extracted files (current norm), the AI is summarizing from a **feed summary/snippet**, not the full article body. Articles that seem shallow in the daily digest may simply have short RSS descriptions.
如果提取结果文件中出现 `content_source: item.raw_summary`,说明 AI 使用的是订阅源摘要或片段,而不是完整正文。日报内容显得较浅时,原因可能只是 RSS 描述过短。
To improve quality: either find feeds that provide full `<content:encoded>`, or switch the feed source to a system that provides full-text RSS (e.g., RSS-proxy with full-text extraction, or a third-party service like FiveFilters).
提高质量可以选择提供完整 `<content:encoded>` 的订阅源,或者把内容来源切换到支持全文 RSS 的系统,例如具备全文提取能力的 RSS 代理或 FiveFilters 等服务。
## Quick check
## 快速检查
先调用 `list_run_artifacts(run_id)`,再读取返回的 extracted Artifact 路径,检查 `content_source` 和 `article.plain_text`。不要按 `run_id` 或“最新目录”手拼路径。
先调用 `list_run_artifacts(run_id)`,再读取返回的提取结果产物路径,检查 `content_source` 和 `article.plain_text`。不要按 `run_id` 或“最新目录”手拼路径。
## Relevant code paths
## 相关代码路径
- `src/summary_mcp/core/pipeline.py` — `RSS_ONLY_UPSTREAMS`, `_should_skip_fetch()`, `extract_content()`
- `src/summary_mcp/core/content_loader.py` — `choose_inline_content()` priority chain, `fetch_html()` (never called for FreshRSS)
- `src/summary_mcp/core/pipeline.py`:`RSS_ONLY_UPSTREAMS`、`_should_skip_fetch()`、`extract_content()`。
- `src/summary_mcp/core/content_loader.py`:`choose_inline_content()` 的优先级链和 `fetch_html()`;FreshRSS 条目不会调用后者。