Files
reader/docs/context-reset-brief.md
T

100 lines
3.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 项目当前状态简报
## 当前已完成
- 已明确整体链路:`来源 -> 聚合池 -> 内容提取 MCP -> LLM 摘要 -> 校验 -> 过滤 -> 入库/推送`
- 已确定当前 MCP 的职责边界:
- 只负责内容提取
- 不负责摘要、分类、价值判断
- 已完成 Python MCP 骨架
- 已实现三个 tool:
- `extract_url_content`
- `extract_item_content`
- `filter_summary_result`
- 已完成真实 URL 提取验证
- 已完成 LLM 摘要 prompt
- 已完成 LLM 摘要结果 schema 校验器
- 已完成“提取 JSON -> LLM 摘要 JSON -> 校验”的最小闭环脚本
- 已完成 LLM 摘要校验 skill 封装
- 已清理旧的启发式 `summarizer.py`
- 已将旧的 MCP 设计文档更新为当前“Content Extract MCP”语义
- 已完成 FreshRSS `greader` API 接入
- 已完成 FreshRSS entry -> `item` 映射
- 已产出真实 `item` 样例文件
- 已跑通 `FreshRSS -> item -> content extraction` 单条链路
- 已定义过滤层输入输出 schema
- 已完成第一版规则引擎、本地脚本和 MCP tool
## 当前关键文件
- MCP 入口:
- `src/summary_mcp/server.py`
- 提取主流程:
- `src/summary_mcp/core/pipeline.py`
- LLM 结果模型:
- `src/summary_mcp/models/llm_result.py`
- LLM 校验器:
- `src/summary_mcp/validators/llm_result.py`
- 过滤模型:
- `src/summary_mcp/models/filtering.py`
- 过滤引擎:
- `src/summary_mcp/filters/engine.py`
- 默认过滤规则:
- `configs/filter_rules.json`
- 校验 CLI:
- `src/summary_mcp/validate_llm_result.py`
- 最小闭环脚本:
- `scripts/run_summary_loop.py`
- FreshRSS 拉取脚本:
- `scripts/pull_freshrss_items.py`
- FreshRSS 提取脚本:
- `scripts/run_freshrss_extract.py`
- 过滤脚本:
- `scripts/run_filter_rules.py`
- 当前提示词:
- `outputs/llm-summary-prompt.txt`
- FreshRSS 原始响应样例:
- `outputs/freshrss.raw.json`
- FreshRSS item 样例:
- `outputs/freshrss.items.json`
- FreshRSS 提取结果:
- `outputs/freshrss.extracted.json`
- 过滤结果样例:
- `outputs/filter-decision.json`
- 文档索引:
- `docs/README.md`
- MVP 归档:
- `docs/content-extract-mcp-mvp-archive.md`
- 当前 TODO:
- `TODO.md`
## 当前已经验证通过
- 参考文章 URL 可提取为结构化 JSON
- LLM 可根据提取结果生成摘要 JSON
- validator 可校验摘要 JSON
- 最小闭环脚本可直接调用 LLM 接口并产出通过校验的结果
- FreshRSS API 可拉取真实 entry
- 真实 entry 可映射为标准化 `item`
- 标准化 `item` 可继续进入内容提取流程
- 规则引擎可对结构化摘要结果输出 `keep / drop / review` 决策
- MCP tool 可承接“由上层 LLM/Agent 调用过滤”的模式
## 当前未开始的下一阶段
- 设计知识库入库格式
- 设计 webhook / 推送格式
- 把 FreshRSS 拉取与提取流程进一步批量化/调度化
- 迭代更细的过滤规则与个性化上下文
## 收束后建议从这里继续
优先从这两个问题继续:
1. 先设计知识库 sink 的输入输出格式
2. 再决定先接 webhook / 推送还是继续细化过滤规则
## 一句话结论
当前 MVP 已完成,FreshRSS 上游和第一版规则过滤层都已接通,下一阶段应转向下游 sink 与推送设计。