docs: translate README to Chinese
This commit is contained in:
@@ -1,15 +1,15 @@
|
||||
# Content Extract MCP
|
||||
# 内容提取 MCP
|
||||
|
||||
Python MCP scaffold for article content extraction, structured summary validation, deterministic filtering, and Markdown sink output.
|
||||
这是一个基于 Python 的 MCP 项目骨架,用于完成文章内容提取、结构化摘要校验、确定性过滤,以及 Markdown 输出落盘。
|
||||
|
||||
## Run
|
||||
## 运行
|
||||
|
||||
```bash
|
||||
pip install -e .
|
||||
summary-mcp
|
||||
```
|
||||
|
||||
The server exposes five tools:
|
||||
服务当前暴露 5 个工具:
|
||||
|
||||
- `extract_url_content`
|
||||
- `extract_item_content`
|
||||
@@ -17,31 +17,33 @@ The server exposes five tools:
|
||||
- `run_freshrss_openclaw_pipeline`
|
||||
- `generate_article_summaries`
|
||||
|
||||
Article-summary post-processing (separate LLM optional):
|
||||
## 单篇文章总结后处理(可选使用独立 LLM)
|
||||
|
||||
**Recommended production input for selected-article summaries**
|
||||
### 生产环境推荐输入
|
||||
|
||||
- Use the per-item extracted files written by the main FreshRSS pipeline:
|
||||
- 单篇总结的正式生产输入,优先使用 FreshRSS 主流水线输出的**单篇 extracted 文件**:
|
||||
- `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json`
|
||||
- Treat these per-item extracted files as the formal default artifacts for downstream selected-article summarization.
|
||||
- A batch extracted file such as `outputs/freshrss/extracted/freshrss.extracted.json` is only a compatible input shape for ad hoc or legacy workflows, not the preferred production default.
|
||||
- 这些**逐条 extracted 文件**是下游单篇总结的**正式默认产物**。
|
||||
- 像 `outputs/freshrss/extracted/freshrss.extracted.json` 这样的**批量 extracted 文件**,只作为临时场景、兼容旧流程的输入形态保留,**不是首选生产默认**。
|
||||
|
||||
Daily knowledge-base defaults:
|
||||
### daily 知识库默认配置
|
||||
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_ID` — default IMA knowledge base ID for daily single-article summaries
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_NAME` — default IMA knowledge base name (expected: `daily`)
|
||||
- Runtime upload logic should verify the configured target before upload; if unavailable, try resolving by name and create `daily` if needed
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_ID` —— 单篇日报总结默认上传的 IMA 知识库 ID
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_NAME` —— 默认知识库名称(预期值:`daily`)
|
||||
- 上传逻辑在运行时应先校验目标知识库;若配置的目标不存在,应先按名称查找,仍不存在则创建 `daily`
|
||||
|
||||
- `article-summary` MCP tool (operates on existing extracted payloads)
|
||||
- `scripts/run_article_summaries.py` CLI helper
|
||||
相关能力:
|
||||
|
||||
Validate an LLM summary result:
|
||||
- `article-summary` MCP 工具(基于已有 extracted payload 做单篇总结)
|
||||
- `scripts/run_article_summaries.py` CLI 辅助脚本
|
||||
|
||||
## 校验 LLM 摘要结果
|
||||
|
||||
```bash
|
||||
validate-llm-result outputs/reference/summary/result.json --extracted outputs/reference/extracted/read-flow-2026.extracted.json
|
||||
```
|
||||
|
||||
Run the minimal extraction-to-summary loop:
|
||||
## 跑最小 extraction → summary 循环
|
||||
|
||||
```bash
|
||||
python scripts/run_summary_loop.py ^
|
||||
@@ -50,7 +52,7 @@ python scripts/run_summary_loop.py ^
|
||||
--output outputs/reference/summary/result.loop.json
|
||||
```
|
||||
|
||||
Pull FreshRSS entries and map them into normalized `item` objects:
|
||||
## 拉取 FreshRSS 条目并映射为标准化 `item`
|
||||
|
||||
```bash
|
||||
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||||
@@ -59,15 +61,16 @@ set FRESHRSS_API_PASSWORD=your-api-password
|
||||
python scripts/pull_freshrss_items.py --limit 5 --mark-read
|
||||
```
|
||||
|
||||
By default the script excludes entries already tagged as `read`. Add `--include-read` if you want the full reading list.
|
||||
When `--mark-read` is enabled, fetched entries are marked as read after the script finishes successfully.
|
||||
默认会排除已经带 `read` 标签的条目。
|
||||
如果你想拿到完整阅读列表,可以加 `--include-read`。
|
||||
启用 `--mark-read` 后,脚本会在执行成功后把本次抓到的条目标记为已读。
|
||||
|
||||
The script writes:
|
||||
脚本会写出:
|
||||
|
||||
- `outputs/freshrss/raw/freshrss.raw.json`
|
||||
- `outputs/freshrss/items/freshrss.items.json`
|
||||
|
||||
Pull FreshRSS entries and run content extraction for each mapped item:
|
||||
## 拉取 FreshRSS 条目并逐条做内容提取
|
||||
|
||||
```bash
|
||||
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||||
@@ -76,17 +79,18 @@ set FRESHRSS_API_PASSWORD=your-api-password
|
||||
python scripts/run_freshrss_extract.py --limit 1 --mark-read
|
||||
```
|
||||
|
||||
By default the script excludes entries already tagged as `read`. When `--mark-read` is enabled, only entries with successful extraction are marked as read.
|
||||
默认会排除已经带 `read` 标签的条目。
|
||||
启用 `--mark-read` 后,只有提取成功的条目才会被标记为已读。
|
||||
|
||||
The script writes:
|
||||
脚本会写出:
|
||||
|
||||
- `outputs/freshrss/raw/freshrss.raw.json`
|
||||
- `outputs/freshrss/items/freshrss.items.json`
|
||||
- `outputs/freshrss/extracted/freshrss.extracted.json`
|
||||
|
||||
Note: this batch extracted file is mainly a compatible artifact for standalone extraction runs and older workflows. The formal production default for downstream selected-article summarization is still the per-item extracted output under `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json`.
|
||||
注意:这个**批量 extracted 文件**主要用于独立提取场景和旧流程兼容。下游单篇总结的正式生产默认输入,仍然是 `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json` 这类**逐条 extracted 文件**。
|
||||
|
||||
Run the full FreshRSS pipeline and mark items as read only after the final OpenClaw delivery payload is written:
|
||||
## 跑完整 FreshRSS 流水线,并在最终 delivery payload 写盘成功后再标记已读
|
||||
|
||||
```bash
|
||||
set FRESHRSS_API_BASE_URL=http://127.0.0.1:8081/api/greader.php
|
||||
@@ -98,7 +102,7 @@ set LLM_MODEL=deepseek-chat
|
||||
python scripts/run_freshrss_pipeline.py --limit 5 --mark-read
|
||||
```
|
||||
|
||||
If you want to apply your personal engineering and AI-agent interest profile during filtering, pass a context file:
|
||||
如果你希望过滤时引入个人工程兴趣 / AI Agent 兴趣画像,可以传入 context 文件:
|
||||
|
||||
```bash
|
||||
python scripts/run_freshrss_pipeline.py ^
|
||||
@@ -107,24 +111,25 @@ python scripts/run_freshrss_pipeline.py ^
|
||||
--mark-read
|
||||
```
|
||||
|
||||
This is the recommended production entrypoint. By default it writes only:
|
||||
这是当前推荐的**正式生产入口**。默认只写出这些产物:
|
||||
|
||||
- `outputs/freshrss/rerun/<timestamp>/raw/freshrss.raw.json`
|
||||
- `outputs/freshrss/rerun/<timestamp>/candidates/openclaw-delivery-payload.json`
|
||||
- `outputs/freshrss/rerun/<timestamp>/run-report.json`
|
||||
- `outputs/freshrss/rerun/<timestamp>/extracted/item-XX.extracted.json` (one per item)
|
||||
- `outputs/freshrss/rerun/<timestamp>/extracted/item-XX.extracted.json`(每篇一份)
|
||||
|
||||
It also updates the daily keyword index runtime data:
|
||||
同时还会更新每日关键词索引运行数据:
|
||||
|
||||
- `data/term_index/daily/YYYY-MM-DD.json`
|
||||
- `data/term_index/term_stats.json`
|
||||
|
||||
The main pipeline does not emit a batch-level `freshrss.extracted.json` file by default.
|
||||
If you need additional per-item intermediates such as normalized items, summaries, filter decisions, candidate records, or candidate inputs, add `--debug-artifacts`.
|
||||
主流水线默认**不会**产出批量级的 `freshrss.extracted.json`。
|
||||
如果你需要更多逐条中间产物,例如标准化 items、摘要结果、过滤决策、candidate record、candidate input,可以加 `--debug-artifacts`。
|
||||
|
||||
When OpenClaw is connected to the MCP server, it should call `run_freshrss_openclaw_pipeline` for the same behavior directly through MCP. The tool also supports `debug_artifacts=true` when deeper inspection is needed.
|
||||
当 OpenClaw 接入这个 MCP 服务后,应直接调用 `run_freshrss_openclaw_pipeline` 来获得同样行为。
|
||||
在排查复杂问题时,也可以把 `debug_artifacts=true` 打开。
|
||||
|
||||
Run deterministic filter rules against a structured summary result:
|
||||
## 对结构化摘要结果执行确定性过滤规则
|
||||
|
||||
```bash
|
||||
python scripts/run_filter_rules.py ^
|
||||
@@ -133,7 +138,7 @@ python scripts/run_filter_rules.py ^
|
||||
--output outputs/reference/filter/filter-decision.json
|
||||
```
|
||||
|
||||
You can optionally pass a context file to inject interest topics or source tags:
|
||||
如果你希望注入兴趣主题或来源标签,也可以额外传入 context 文件:
|
||||
|
||||
```bash
|
||||
python scripts/run_filter_rules.py ^
|
||||
@@ -143,12 +148,12 @@ python scripts/run_filter_rules.py ^
|
||||
--output outputs/reference/filter/filter-decision.with-context.json
|
||||
```
|
||||
|
||||
Rule engine details and rule authoring guidance live in:
|
||||
规则引擎设计和规则编写说明见:
|
||||
|
||||
- `docs/design/filter-rule-engine-design.md`
|
||||
- `docs/design/filter-rule-engine-usage.md`
|
||||
|
||||
Write a filtered result into the Markdown sink:
|
||||
## 将过滤结果写入 Markdown sink
|
||||
|
||||
```bash
|
||||
python scripts/run_markdown_sink.py ^
|
||||
@@ -157,9 +162,9 @@ python scripts/run_markdown_sink.py ^
|
||||
--filter outputs/reference/filter/filter-decision.json
|
||||
```
|
||||
|
||||
The script writes markdown notes under `knowledge-base/`.
|
||||
脚本会把 Markdown 笔记写到 `knowledge-base/` 下。
|
||||
|
||||
Build an internal `ArticleCandidateRecord` and a slim `OpenClawCandidateInput`:
|
||||
## 构建内部 `ArticleCandidateRecord` 与精简版 `OpenClawCandidateInput`
|
||||
|
||||
```bash
|
||||
python scripts/run_article_candidate.py ^
|
||||
@@ -169,12 +174,12 @@ python scripts/run_article_candidate.py ^
|
||||
--section-hint tools_and_workflows
|
||||
```
|
||||
|
||||
The script writes by default:
|
||||
默认会写出:
|
||||
|
||||
- `outputs/reference/candidates/article-candidate-record.json`
|
||||
- `outputs/reference/candidates/openclaw-candidate-input.json`
|
||||
|
||||
Build a batch OpenClaw delivery payload:
|
||||
## 构建批量 OpenClaw delivery payload
|
||||
|
||||
```bash
|
||||
python scripts/build_openclaw_delivery.py ^
|
||||
@@ -183,13 +188,15 @@ python scripts/build_openclaw_delivery.py ^
|
||||
--date 2026-03-25
|
||||
```
|
||||
|
||||
The script writes by default:
|
||||
默认会写出:
|
||||
|
||||
- `outputs/reference/candidates/openclaw-delivery-payload.json`
|
||||
|
||||
Output layout details live in `outputs/README.md`.
|
||||
输出目录布局说明见 `outputs/README.md`。
|
||||
|
||||
Keyword index defaults live in:
|
||||
## 关键词索引默认配置
|
||||
|
||||
相关配置文件位于:
|
||||
|
||||
- `configs/term_aliases.json`
|
||||
- `configs/term_stopwords.json`
|
||||
@@ -197,20 +204,22 @@ Keyword index defaults live in:
|
||||
- `configs/term_watchlist.json`
|
||||
- `configs/term_change_log.json`
|
||||
|
||||
You can also rebuild the keyword index from an existing delivery payload:
|
||||
你也可以基于已有 delivery payload 重新构建关键词索引:
|
||||
|
||||
```bash
|
||||
python scripts/build_keyword_index.py ^
|
||||
--input outputs/reference/candidates/openclaw-delivery-payload.json
|
||||
```
|
||||
|
||||
Runtime keyword data is stored under `data/term_index/`.
|
||||
运行期关键词数据存放在:
|
||||
|
||||
The keyword cleanup review skill lives in:
|
||||
- `data/term_index/`
|
||||
|
||||
关键词清理评审 skill 位于:
|
||||
|
||||
- `skills/keyword-cleanup-review/`
|
||||
|
||||
To build a review bundle for the LLM skill:
|
||||
构建给 LLM skill 使用的评审数据包(review bundle):
|
||||
|
||||
```bash
|
||||
python skills/keyword-cleanup-review/scripts/build_review_bundle.py ^
|
||||
@@ -219,15 +228,19 @@ python skills/keyword-cleanup-review/scripts/build_review_bundle.py ^
|
||||
--output outputs/term_index/review/keyword-cleanup-bundle.json
|
||||
```
|
||||
|
||||
The review bundle now also carries cleanup governance context:
|
||||
现在这个评审数据包(review bundle)还会额外携带治理上下文:
|
||||
|
||||
- cleanup thresholds from `configs/term_cleanup_policy.json`
|
||||
- the current watch list from `configs/term_watchlist.json`
|
||||
- recent applied changes from `configs/term_change_log.json`
|
||||
- 来自 `configs/term_cleanup_policy.json` 的清理阈值
|
||||
- 当前 watch list(`configs/term_watchlist.json`)
|
||||
- 最近已应用的变更(`configs/term_change_log.json`)
|
||||
|
||||
The skill only produces review inputs and suggestions. It does not modify `term_aliases`, `term_stopwords`, or `filter_context.personal.json` automatically.
|
||||
这个 skill 只负责生成 review 输入与建议,不会自动修改:
|
||||
|
||||
To preview accepted suggestions before writing any config files:
|
||||
- `term_aliases`
|
||||
- `term_stopwords`
|
||||
- `filter_context.personal.json`
|
||||
|
||||
如果你想先预览已接受建议,再决定是否写配置文件:
|
||||
|
||||
```bash
|
||||
python scripts/apply_term_suggestions.py ^
|
||||
@@ -236,20 +249,21 @@ python scripts/apply_term_suggestions.py ^
|
||||
--dry-run
|
||||
```
|
||||
|
||||
Remove `--dry-run` to write the accepted changes. The script can also apply accepted `alias`, `stopword`, and `interest keyword` suggestions through `--accept-alias`, `--accept-stopword`, and `--accept-interest`. Accepted watch terms are written into `configs/term_watchlist.json`, and every applied action is appended into `configs/term_change_log.json`.
|
||||
去掉 `--dry-run` 后才会真正写文件。
|
||||
这个脚本也支持通过 `--accept-alias`、`--accept-stopword`、`--accept-interest` 应用 alias / stopword / interest keyword 变更。
|
||||
已接受的 watch 词会写入 `configs/term_watchlist.json`,每次应用动作也会被追加到 `configs/term_change_log.json`。
|
||||
|
||||
## 单篇总结 LLM 配置
|
||||
|
||||
## Article-summary LLM configuration
|
||||
|
||||
Set a dedicated model for post-processing summaries without affecting the main pipeline:
|
||||
如果你希望单篇总结后处理使用独立模型,而不影响主流水线,可以设置:
|
||||
|
||||
- `ARTICLE_SUMMARY_LLM_API_URL`
|
||||
- `ARTICLE_SUMMARY_LLM_MODEL`
|
||||
- `ARTICLE_SUMMARY_LLM_API_KEY`
|
||||
|
||||
If these are not set, the summarizer falls back to the main `LLM_*` / `OPENAI_*` settings used elsewhere.
|
||||
如果这些变量未设置,单篇总结会回退使用主流程中的 `LLM_*` / `OPENAI_*` 配置。
|
||||
|
||||
Example (PowerShell style):
|
||||
示例(PowerShell 风格):
|
||||
|
||||
```bash
|
||||
set ARTICLE_SUMMARY_LLM_API_URL=https://api.deepseek.com
|
||||
@@ -257,8 +271,8 @@ set ARTICLE_SUMMARY_LLM_API_KEY=your-article-summary-key
|
||||
set ARTICLE_SUMMARY_LLM_MODEL=deepseek-chat
|
||||
```
|
||||
|
||||
Then run, for example, either via the CLI script.
|
||||
For normal production use, prefer a per-item extracted file from `outputs/freshrss/rerun/<run_id>/extracted/`. The batch extracted example below is kept only as a compatible legacy/ad hoc input shape:
|
||||
然后可以这样调用 CLI。
|
||||
正式生产环境建议优先使用 `outputs/freshrss/rerun/<run_id>/extracted/` 下的**逐条 extracted 文件**;下面这个**批量 extracted** 示例仅保留为兼容旧流程 / 临时场景输入:
|
||||
|
||||
```bash
|
||||
python scripts/run_article_summaries.py ^
|
||||
@@ -267,13 +281,22 @@ python scripts/run_article_summaries.py ^
|
||||
--output-dir outputs/freshrss/single_summaries
|
||||
```
|
||||
|
||||
…or through the MCP server tool `generate_article_summaries` exposed by `summary_mcp.server`:
|
||||
也可以通过 `summary_mcp.server` 暴露的 MCP 工具 `generate_article_summaries` 调用:
|
||||
|
||||
- `extracted_path` (string): path to a single-item extracted JSON (e.g. `outputs/freshrss/rerun/<run_id>/extracted/item-01.extracted.json`) or a batch extracted JSON containing a `results` array.
|
||||
- `selected_ids` (array of strings): one or more `item_id` values to summarize. Pass an empty array to summarize all items in the file.
|
||||
- `output_dir` (optional string): directory to write Markdown summaries. If omitted, summaries are written under `single_summaries/` next to the extracted file.
|
||||
- `llm_api_key` / `llm_model` / `llm_api_url` (optional strings): overrides for article-summary LLM settings. If omitted, the tool falls back to `ARTICLE_SUMMARY_*` or main `LLM_*` env vars as described above.
|
||||
- `extracted_path`(string):单篇 extracted JSON 路径(例如 `outputs/freshrss/rerun/<run_id>/extracted/item-01.extracted.json`),或者包含 `results` 数组的 batch extracted JSON
|
||||
- `selected_ids`(array of strings):要总结的一个或多个 `item_id`。如果传空数组,则对文件中的全部条目做总结
|
||||
- `output_dir`(optional string):Markdown 输出目录;若不传,则默认写到 extracted 文件旁边的 `single_summaries/` 目录
|
||||
- `llm_api_key` / `llm_model` / `llm_api_url`(optional strings):单篇总结 LLM 的覆盖配置;不传时会按前文规则回退到 `ARTICLE_SUMMARY_*` 或主 `LLM_*`
|
||||
|
||||
The tool returns a JSON array of file paths for the generated Markdown summaries.
|
||||
该工具返回一个 JSON 数组,内容为生成好的 Markdown 文件路径。
|
||||
|
||||
The article summary uses a dedicated prompt (`outputs/prompts/article-summary-prompt.txt`) that is completely independent from the daily digest prompt. It outputs a structured knowledge note in Chinese with sections: 核心结论、主要论点、关键方法 / 机制、重要细节、可复用启发、关键词、主题.
|
||||
单篇总结使用独立 prompt:`outputs/prompts/article-summary-prompt.txt`。
|
||||
它与日报 prompt 完全独立,输出的是中文结构化知识笔记,包含这些部分:
|
||||
|
||||
- 核心结论
|
||||
- 主要论点
|
||||
- 关键方法 / 机制
|
||||
- 重要细节
|
||||
- 可复用启发
|
||||
- 关键词
|
||||
- 主题
|
||||
|
||||
Reference in New Issue
Block a user