diff --git a/README.md b/README.md index db9ed20..49c2862 100644 --- a/README.md +++ b/README.md @@ -96,13 +96,15 @@ This is the recommended production entrypoint. By default it writes only: - `outputs/freshrss/rerun//raw/freshrss.raw.json` - `outputs/freshrss/rerun//candidates/openclaw-delivery-payload.json` - `outputs/freshrss/rerun//run-report.json` +- `outputs/freshrss/rerun//extracted/item-XX.extracted.json` (one per item) It also updates the daily keyword index runtime data: - `data/term_index/daily/YYYY-MM-DD.json` - `data/term_index/term_stats.json` -If you need per-item intermediates, add `--debug-artifacts`. +The main pipeline does not emit a batch-level `freshrss.extracted.json` file by default. +If you need additional per-item intermediates such as normalized items, summaries, filter decisions, candidate records, or candidate inputs, add `--debug-artifacts`. When OpenClaw is connected to the MCP server, it should call `run_freshrss_openclaw_pipeline` for the same behavior directly through MCP. The tool also supports `debug_artifacts=true` when deeper inspection is needed. diff --git a/docs/openclaw/openclaw-handoff.md b/docs/openclaw/openclaw-handoff.md index fd76796..d06ca4a 100644 --- a/docs/openclaw/openclaw-handoff.md +++ b/docs/openclaw/openclaw-handoff.md @@ -102,13 +102,16 @@ By default the pipeline writes only: - `outputs/freshrss/rerun//raw/freshrss.raw.json` - `outputs/freshrss/rerun//candidates/openclaw-delivery-payload.json` - `outputs/freshrss/rerun//run-report.json` +- `outputs/freshrss/rerun//extracted/item-XX.extracted.json` (one per item) It also updates local runtime keyword data: - `data/term_index/daily/YYYY-MM-DD.json` - `data/term_index/term_stats.json` -If `debug_artifacts=true`, the pipeline additionally writes per-item intermediate files. +Per-item extracted files live under `extracted/` and are always written. +If `debug_artifacts=true`, the pipeline additionally writes normalized items, summaries, filter decisions, candidate records, and candidate inputs. +The main pipeline does not emit a batch-level `freshrss.extracted.json` file by default. ## Payload Specs diff --git a/outputs/README.md b/outputs/README.md index a4eecf9..d9ee2a5 100644 --- a/outputs/README.md +++ b/outputs/README.md @@ -4,7 +4,7 @@ ## 生产默认输出 -对于 FreshRSS 主流水线,默认只保留这 3 类文件: +对于 FreshRSS 主流水线,默认保留这 4 类文件: - `outputs/freshrss/rerun//raw/freshrss.raw.json` - FreshRSS 原始响应快照 @@ -12,6 +12,8 @@ - 发给 OpenClaw 的最终批量 payload - `outputs/freshrss/rerun//run-report.json` - 本次批处理的运行报告 +- `outputs/freshrss/rerun//extracted/item-XX.extracted.json` + - 单篇文章的正文提取结果(每篇一份) 这是当前推荐的生产模式。 @@ -28,12 +30,15 @@ ## 调试扩展输出 +默认正式产物已经包含: + +- `outputs/freshrss/rerun//extracted/item-XX.extracted.json` + - 单篇正文提取结果 + 当启用 `debug_artifacts` 时,才会额外写出这些中间文件: - `outputs/freshrss/rerun//items/` - 标准化 `item` 中间文件 -- `outputs/freshrss/rerun//extracted/` - - 提取结果 - `outputs/freshrss/rerun//summary/` - LLM 总结结果及调试尝试文件 - `outputs/freshrss/rerun//filter/` @@ -43,6 +48,8 @@ - `outputs/freshrss/rerun//candidates/*.openclaw-candidate-input.json` - 单篇 OpenClaw 输入对象 +主流水线默认不产出批量聚合版 `freshrss.extracted.json`。 + ## 何时启用调试产物 只在这些场景启用: diff --git a/src/summary_mcp/workflows/freshrss_pipeline.py b/src/summary_mcp/workflows/freshrss_pipeline.py index eb7ee5f..c39e879 100644 --- a/src/summary_mcp/workflows/freshrss_pipeline.py +++ b/src/summary_mcp/workflows/freshrss_pipeline.py @@ -141,7 +141,7 @@ def run_freshrss_pipeline( for index, item in enumerate(items, start=1): item_key = f"item-{index:02d}" item_path = _maybe_path(debug_artifacts, resolved_output_dir / "items" / f"{item_key}.item.json") - extracted_path = _maybe_path(debug_artifacts, resolved_output_dir / "extracted" / f"{item_key}.extracted.json") + extracted_path = resolved_output_dir / "extracted" / f"{item_key}.extracted.json" summary_output = _maybe_path(debug_artifacts, resolved_output_dir / "summary" / item_key / "result.loop.json") filter_path = _maybe_path(debug_artifacts, resolved_output_dir / "filter" / f"{item_key}.filter.json") record_path = _maybe_path(debug_artifacts, resolved_output_dir / "candidates" / f"{item_key}.article-candidate-record.json")