feat: keep per-item extracted outputs by default

This commit is contained in:
root
2026-03-28 23:14:18 +08:00
parent 0a00a97724
commit 9fd21c59f5
4 changed files with 18 additions and 6 deletions
+3 -1
View File
@@ -96,13 +96,15 @@ This is the recommended production entrypoint. By default it writes only:
- `outputs/freshrss/rerun/<timestamp>/raw/freshrss.raw.json`
- `outputs/freshrss/rerun/<timestamp>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<timestamp>/run-report.json`
- `outputs/freshrss/rerun/<timestamp>/extracted/item-XX.extracted.json` (one per item)
It also updates the daily keyword index runtime data:
- `data/term_index/daily/YYYY-MM-DD.json`
- `data/term_index/term_stats.json`
If you need per-item intermediates, add `--debug-artifacts`.
The main pipeline does not emit a batch-level `freshrss.extracted.json` file by default.
If you need additional per-item intermediates such as normalized items, summaries, filter decisions, candidate records, or candidate inputs, add `--debug-artifacts`.
When OpenClaw is connected to the MCP server, it should call `run_freshrss_openclaw_pipeline` for the same behavior directly through MCP. The tool also supports `debug_artifacts=true` when deeper inspection is needed.
+4 -1
View File
@@ -102,13 +102,16 @@ By default the pipeline writes only:
- `outputs/freshrss/rerun/<run_id>/raw/freshrss.raw.json`
- `outputs/freshrss/rerun/<run_id>/candidates/openclaw-delivery-payload.json`
- `outputs/freshrss/rerun/<run_id>/run-report.json`
- `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json` (one per item)
It also updates local runtime keyword data:
- `data/term_index/daily/YYYY-MM-DD.json`
- `data/term_index/term_stats.json`
If `debug_artifacts=true`, the pipeline additionally writes per-item intermediate files.
Per-item extracted files live under `extracted/` and are always written.
If `debug_artifacts=true`, the pipeline additionally writes normalized items, summaries, filter decisions, candidate records, and candidate inputs.
The main pipeline does not emit a batch-level `freshrss.extracted.json` file by default.
## Payload Specs
+10 -3
View File
@@ -4,7 +4,7 @@
## 生产默认输出
对于 FreshRSS 主流水线,默认只保留这 3 类文件:
对于 FreshRSS 主流水线,默认保留这 4 类文件:
- `outputs/freshrss/rerun/<run_id>/raw/freshrss.raw.json`
- FreshRSS 原始响应快照
@@ -12,6 +12,8 @@
- 发给 OpenClaw 的最终批量 payload
- `outputs/freshrss/rerun/<run_id>/run-report.json`
- 本次批处理的运行报告
- `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json`
- 单篇文章的正文提取结果(每篇一份)
这是当前推荐的生产模式。
@@ -28,12 +30,15 @@
## 调试扩展输出
默认正式产物已经包含:
- `outputs/freshrss/rerun/<run_id>/extracted/item-XX.extracted.json`
- 单篇正文提取结果
当启用 `debug_artifacts` 时,才会额外写出这些中间文件:
- `outputs/freshrss/rerun/<run_id>/items/`
- 标准化 `item` 中间文件
- `outputs/freshrss/rerun/<run_id>/extracted/`
- 提取结果
- `outputs/freshrss/rerun/<run_id>/summary/`
- LLM 总结结果及调试尝试文件
- `outputs/freshrss/rerun/<run_id>/filter/`
@@ -43,6 +48,8 @@
- `outputs/freshrss/rerun/<run_id>/candidates/*.openclaw-candidate-input.json`
- 单篇 OpenClaw 输入对象
主流水线默认不产出批量聚合版 `freshrss.extracted.json`。
## 何时启用调试产物
只在这些场景启用:
@@ -141,7 +141,7 @@ def run_freshrss_pipeline(
for index, item in enumerate(items, start=1):
item_key = f"item-{index:02d}"
item_path = _maybe_path(debug_artifacts, resolved_output_dir / "items" / f"{item_key}.item.json")
extracted_path = _maybe_path(debug_artifacts, resolved_output_dir / "extracted" / f"{item_key}.extracted.json")
extracted_path = resolved_output_dir / "extracted" / f"{item_key}.extracted.json"
summary_output = _maybe_path(debug_artifacts, resolved_output_dir / "summary" / item_key / "result.loop.json")
filter_path = _maybe_path(debug_artifacts, resolved_output_dir / "filter" / f"{item_key}.filter.json")
record_path = _maybe_path(debug_artifacts, resolved_output_dir / "candidates" / f"{item_key}.article-candidate-record.json")