refactor: simplify reader digest skill
This commit is contained in:
+120
-886
File diff suppressed because it is too large
Load Diff
@@ -37,10 +37,7 @@ To improve quality: either find feeds that provide full `<content:encoded>`, or
|
||||
|
||||
## Quick check
|
||||
|
||||
```bash
|
||||
# Check content_source for latest run
|
||||
grep -h "content_source" outputs/freshrss/rerun/*/extracted/item-*.json | sort | uniq -c
|
||||
```
|
||||
先调用 `list_run_artifacts(run_id)`,再读取返回的 extracted Artifact 路径,检查 `content_source` 和 `article.plain_text`。不要按 `run_id` 或“最新目录”手拼路径。
|
||||
|
||||
## Relevant code paths
|
||||
|
||||
|
||||
@@ -1,328 +1,150 @@
|
||||
# Reader Digest Flow Reference
|
||||
# Reader Digest Flow 操作参考
|
||||
|
||||
## Purpose
|
||||
本文件承载 `reader-digest-flow` 的具体执行步骤。正式行为边界以 `../SKILL.md` 为准。
|
||||
|
||||
Concrete operational checklist for the `reader-digest-flow` skill.
|
||||
## 1. 日报 Pipeline
|
||||
|
||||
## Default Operating Model
|
||||
### 默认参数
|
||||
|
||||
### Layering
|
||||
|
||||
- `reader` layer:
|
||||
- FreshRSS pull
|
||||
- extraction
|
||||
- summary/filter/payload generation
|
||||
- selected-article summary capability
|
||||
- OpenClaw / skill layer:
|
||||
- public digest generation
|
||||
- internal review digest generation
|
||||
- Hugo publishing
|
||||
- chat reporting
|
||||
- user confirmation handling
|
||||
- calling selected-article summaries
|
||||
- IMA upload orchestration
|
||||
- Hugo layer:
|
||||
- public digest browsing and archive only
|
||||
- IMA layer:
|
||||
- long-term storage for selected article notes only
|
||||
|
||||
### Hard rules
|
||||
|
||||
- Do not upload the full digest to IMA.
|
||||
- Upload only explicitly user-selected articles to IMA.
|
||||
- Do not generate a digest without a real payload.
|
||||
- Generate two views from the same payload: a public digest for Hugo and an internal review digest for chat/operator workflow.
|
||||
- Do not expose internal review states or operator-facing labels in the public digest.
|
||||
- Do not re-fetch original URLs for selected summaries; use existing extracted text.
|
||||
- If the main pipeline fails, inspect the run first; when `inspect_resume_plan` says `recommended_action=resume`, continue via the async resume job path instead of stopping immediately.
|
||||
- Always branch on reader's top-level reconciled `status`; treat `status_source` and `state_conflict` only as explanatory metadata.
|
||||
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization, with fallback to main `LLM_*` only if needed.
|
||||
- Actively report progress after each completed phase.
|
||||
|
||||
## Step-by-step checklist
|
||||
|
||||
### 1. Run reader pipeline
|
||||
|
||||
Use the formal MCP workflow path as the default production route.
|
||||
Prefer MCP run/status/result operations over direct path stitching. Only fall back to CLI or direct file inspection for debug / manual troubleshooting.
|
||||
|
||||
Formal production startup sequence:
|
||||
|
||||
1. `start_freshrss_pipeline_job`
|
||||
2. `get_freshrss_pipeline_job_status`
|
||||
3. `get_freshrss_pipeline_job_result`
|
||||
4. after success, continue with `run_id` via `get_run_status` / `get_delivery_payload` / `get_run_report`
|
||||
|
||||
If the main pipeline job ends in `failed`:
|
||||
|
||||
1. inspect the linked run with `get_run_status`
|
||||
2. call `inspect_resume_plan(run_id)`
|
||||
3. if `recommended_action=resume`, continue with:
|
||||
- `start_resume_job`
|
||||
- `get_resume_job_status`
|
||||
- `get_resume_job_result`
|
||||
4. if `recommended_action=read_terminal_result`, continue from the terminal run result
|
||||
5. if `recommended_action=start_new_run`, stop and report the failure
|
||||
|
||||
Treat the old synchronous `run_freshrss_openclaw_pipeline` as debug / light validation / fallback only.
|
||||
|
||||
Project root:
|
||||
|
||||
```bash
|
||||
/home/ubuntu/zhu/github/reader
|
||||
```json
|
||||
{
|
||||
"limit": 7,
|
||||
"include_read": false,
|
||||
"mark_read": true,
|
||||
"debug_artifacts": false,
|
||||
"timeout_seconds": 60,
|
||||
"max_retries": 2
|
||||
}
|
||||
```
|
||||
|
||||
Default behavior for a normal production run:
|
||||
|
||||
- if the user did not specify a count, randomly choose a limit between 5 and 10 items for that run
|
||||
- run with mark-read enabled
|
||||
- do not enable `debug_artifacts`
|
||||
- only skip mark-read if the user explicitly says the run is debug, test, or validation
|
||||
- only enable `debug_artifacts` if the user explicitly says the run is debug, test, validation, or troubleshooting
|
||||
|
||||
Typical artifacts to inspect after a successful run:
|
||||
### 正式调用序列
|
||||
|
||||
```text
|
||||
outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json
|
||||
outputs/freshrss/rerun/<run-id>/candidates/digest-brief.json
|
||||
outputs/freshrss/rerun/<run-id>/run-report.json
|
||||
outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json
|
||||
start_freshrss_pipeline_job
|
||||
→ get_freshrss_pipeline_job_status
|
||||
→ get_freshrss_pipeline_job_result
|
||||
→ get_run_status
|
||||
→ get_delivery_payload / get_run_report
|
||||
```
|
||||
|
||||
Notes:
|
||||
状态动作:
|
||||
|
||||
- `digest-brief.json` is the preferred input for **public digest** generation.
|
||||
- It is a lighter public-only view and currently includes only `keep` candidates.
|
||||
- If `digest-brief.json` is missing, fall back to `openclaw-delivery-payload.json`.
|
||||
- If the synchronous MCP wrapper times out but a real reader run was still created, do not discard the run; continue from run truth using `list_runs`, `get_run_report`, and `get_delivery_payload`.
|
||||
- `running`:按合理间隔继续轮询;
|
||||
- `success`:读取结果,保存 `run_id`;
|
||||
- `failed`:读取关联 Run 并执行 Resume Plan;
|
||||
- `partial`:优先检查 Run Report、Artifact 和 Recovery 信息。
|
||||
|
||||
**⚠️ Output path divergence when run from Feishu context:**
|
||||
When triggered via MCP from inside a Feishu session, the extracted files may be written to an MCP-managed temp path instead of `outputs/freshrss/rerun/<run-id>/extracted/`. Verify the actual output path from the pipeline result before continuing to Phase 5/6. If the project-relative path doesn't exist, use the MCP result's `extracted_path` directly or copy the files to the expected project location.
|
||||
|
||||
### 2. Generate digest markdown
|
||||
|
||||
Generate two output views from the same run, preferably in one model call:
|
||||
|
||||
1. a **public digest** for Hugo / public readers
|
||||
2. an **internal review digest** for chat / operator workflow
|
||||
|
||||
Input preference:
|
||||
|
||||
- **public digest**: prefer `candidates/digest-brief.json`
|
||||
- **internal review digest**: use `candidates/openclaw-delivery-payload.json`
|
||||
|
||||
Recommended generation pattern:
|
||||
|
||||
- Pass the public brief and the full payload as two explicitly labeled input blocks.
|
||||
- Ask the model to return both outputs in one response.
|
||||
- Prefer a structured response shape (for example JSON with `public_digest_markdown` and `internal_review_digest_markdown`) when post-processing is needed.
|
||||
|
||||
Before publishing, persist the generated digest artifacts back into the same reader run directory:
|
||||
恢复序列:
|
||||
|
||||
```text
|
||||
outputs/freshrss/rerun/<run-id>/digest/public_digest.md
|
||||
outputs/freshrss/rerun/<run-id>/digest/internal_review_digest.md
|
||||
outputs/freshrss/rerun/<run-id>/digest/combined.json
|
||||
inspect_resume_plan
|
||||
→ recommended_action=resume
|
||||
→ start_resume_job
|
||||
→ get_resume_job_status
|
||||
→ get_resume_job_result
|
||||
```
|
||||
|
||||
Then write only the public digest into Hugo here:
|
||||
如果建议为 `read_terminal_result`,直接读取结果;如果为 `start_new_run`,停止并等待用户决定。
|
||||
|
||||
不要根据 `run_id` 拼目录。后续只消费调用返回的 `output_dir`、`artifact.path`、`delivery_output`、`report_output`。
|
||||
|
||||
## 2. 候选汇报与选择
|
||||
|
||||
优先使用当前 Run 的 digest brief Artifact;不可用时使用 `get_delivery_payload` 返回的候选。
|
||||
|
||||
展示规则:
|
||||
|
||||
1. 使用候选数组原始顺序并从 1 编号;
|
||||
2. 不因 `keep/review` 分组而重新编号;
|
||||
3. 每篇展示标题、来源、摘要和判断理由;
|
||||
4. 用户编号只映射当前候选数组;
|
||||
5. 跨 Artifact 读取详情时按 URL 或完整 `item_id` 匹配。
|
||||
|
||||
不要按 `summary-batch`、`item-XX` 和候选数组的相同位置推断它们是同一篇文章。
|
||||
|
||||
## 3. Hugo 日报
|
||||
|
||||
用户确认发布文章后:
|
||||
|
||||
1. 读取 `public-digest-example.md`;
|
||||
2. 仅使用用户选中的文章生成公开内容;
|
||||
3. 写入 Hugo 当日页面;
|
||||
4. 前台执行部署,避免把构建日志作为聊天通知;
|
||||
5. 验证首页、列表页和详情页。
|
||||
|
||||
当前部署位置:
|
||||
|
||||
```text
|
||||
/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md
|
||||
```
|
||||
|
||||
Recommended public-digest front matter:
|
||||
验证至少覆盖:
|
||||
|
||||
```toml
|
||||
+++
|
||||
title = "AI 日报 · YYYY-MM-DD"
|
||||
date = YYYY-MM-DDTHH:MM:SS+08:00
|
||||
summary = "当日日报摘要"
|
||||
+++
|
||||
```text
|
||||
http://127.0.0.1:14322/
|
||||
http://127.0.0.1:14322/daily/
|
||||
http://127.0.0.1:14322/daily/YYYY-MM-DD/
|
||||
```
|
||||
|
||||
Recommended **public digest** structure:
|
||||
页面不存在时检查源文件、构建结果、容器状态和页面日期。需要 Docker/Hugo 深度排障时再读取部署项目自己的文档,不把历史事故规则复制回主 Skill。
|
||||
|
||||
- `今日概览`
|
||||
- `今日重点`
|
||||
- `趋势观察`
|
||||
- `延伸阅读`
|
||||
## 4. 单篇知识笔记
|
||||
|
||||
Public digest constraints:
|
||||
用户确认 IMA 文章后:
|
||||
|
||||
- Use only the public brief view when present.
|
||||
- Keep public wording free of internal workflow language.
|
||||
- Optimize for concise public readability with solid information density.
|
||||
- For each `今日重点` item, add one short editor-style sentence explaining why the item matters in today's digest (for example: `这篇内容更值得关注的原因在于……`).
|
||||
- Prefer rendering highlight points as a short public label such as `值得关注:` followed by one bullet per line.
|
||||
- **Article numbering: MUST use `1.` `2.` `3.` (Arabic numeral + period). Do NOT use `① ② ③`,`一、二、三`,`第一条` or any other variant.**
|
||||
- **All four sections are required. Missing any one is a format violation.**
|
||||
- **Hugo digest must include ALL keep articles from `digest-brief.json`. The IMA deposition subset is a separate downstream step.**
|
||||
1. 从候选中取得 URL 和完整 `item_id`;
|
||||
2. 从 Run Artifact 中找到匹配的 extracted 文件;
|
||||
3. 交叉验证 `article.item_id` 或 URL;
|
||||
4. 对每个单篇文件调用 `start_article_summary_job(extracted_path=<返回的 Artifact 路径>, selected_ids=[<完整 item_id>])`;
|
||||
5. `selected_ids` 传完整、非空、去除 `cand:` 前缀后的 `item_id`;
|
||||
6. 从 Job 结果的 `written_paths` 读取 Markdown。
|
||||
|
||||
Recommended **internal review digest** structure:
|
||||
正式序列:
|
||||
|
||||
- `今日候选概况`
|
||||
- `已入选重点`
|
||||
- `待你确认`
|
||||
- `建议沉淀到 IMA`
|
||||
- `原始候选清单`
|
||||
|
||||
Internal review digest constraints:
|
||||
|
||||
- Do not show `rank` values.
|
||||
- Replace machine states with Chinese labels such as `已入选` / `待确认` / `暂不纳入`.
|
||||
- For `已入选重点`, include a fuller summary plus a short judgment paragraph.
|
||||
- For `待你确认`, include a fuller summary, reason, and recommendation.
|
||||
- Keep it readable as an operator review draft, not a raw payload dump.
|
||||
- **Feishu compatibility: never use markdown tables in the digest report.**
|
||||
When delivered via Feishu, a single table forces the whole message to plain
|
||||
text. Use lists and sections instead. See `references/feishu-format-notes.md`.
|
||||
|
||||
### 3. Publish Hugo
|
||||
|
||||
Publish only the public digest to Hugo.
|
||||
Treat Hugo publication as the default continuation of a successful normal daily digest run. Do not ask for a second confirmation before generating/writing the public digest and publishing it, unless the user explicitly requests not to publish to Hugo.
|
||||
|
||||
**⚠️ Hugo redeploy: always use foreground terminal, never background+notify_on_complete.**
|
||||
|
||||
The Hugo redeploy runs via `./redeploy.sh` and completes in ~10 seconds. **Do NOT use `terminal(background=true, notify_on_complete=true)`** for this step — the Gateway will push the raw Docker build output (compiler logs, layered build output, nginx config) to Feishu as a notification. This output is machine-readable, not human-readable, and the user has explicitly said this is noise.
|
||||
|
||||
**Correct approach:**
|
||||
|
||||
```python
|
||||
# ✅ Foreground terminal with adequate timeout
|
||||
cmd = "cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh"
|
||||
result = terminal(cmd, timeout=120)
|
||||
# Then verify 200 OK on the detail page
|
||||
```text
|
||||
start_article_summary_job
|
||||
→ get_article_summary_job_status
|
||||
→ get_article_summary_job_result
|
||||
```
|
||||
|
||||
**Wrong approach (DO NOT use):**
|
||||
内容依据是 extracted Artifact 中的 `article.plain_text`。不得重新抓取原 URL,不得使用原文之外的知识扩写。
|
||||
|
||||
```python
|
||||
# ❌ Background + notify sends raw build logs to Feishu
|
||||
terminal("cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh", background=True, notify_on_complete=True)
|
||||
```
|
||||
|
||||
This rule only applies to Hugo redeploy (fast, deterministic output). For long-running Reader pipeline jobs (>60s), background + notify_on_complete is still appropriate since the output is meaningful content (article summaries, pipeline stats).
|
||||
|
||||
1. **Pre-check: verify Hugo content directory exists**
|
||||
```bash
|
||||
ls -d /home/ubuntu/zhu/apps/hugo-site/content/daily/ 2>/dev/null || mkdir -p /home/ubuntu/zhu/apps/hugo-site/content/daily/
|
||||
```
|
||||
If the entire `content/` directory is missing, create it before proceeding.
|
||||
**Do not assume the Hugo site has a standard structure.**
|
||||
|
||||
2. write the public digest markdown to:
|
||||
- `/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md`
|
||||
3. redeploy Hugo immediately after writing:
|
||||
- `cd /home/ubuntu/zhu/apps/hugo-site && ./redeploy.sh`
|
||||
4. verify all three URLs before continuing:
|
||||
- `http://127.0.0.1:14322/`
|
||||
- `http://127.0.0.1:14322/daily/`
|
||||
- `http://127.0.0.1:14322/daily/YYYY-MM-DD/`
|
||||
|
||||
Expected verification targets:
|
||||
|
||||
- homepage works
|
||||
- `/daily/` works
|
||||
- `/daily/YYYY-MM-DD/` works
|
||||
|
||||
### 4. Report digest in chat
|
||||
|
||||
Provide the internal review digest in chat and ask which articles should be retained.
|
||||
|
||||
Hard reporting rule:
|
||||
- do not send only article titles
|
||||
- for each article, include at least a one-sentence summary and a short recommendation / judgment so the user can decide without reopening the source
|
||||
- **Feishu**: avoid markdown tables entirely. Use lists with headings.
|
||||
|
||||
### 5. Generate selected article summaries
|
||||
|
||||
Only do this after the user explicitly confirms which articles to retain.
|
||||
|
||||
Preferred MCP path:
|
||||
|
||||
- `start_article_summary_job`
|
||||
- `get_article_summary_job_status`
|
||||
- `get_article_summary_job_result`
|
||||
|
||||
Expected inputs:
|
||||
|
||||
- `extracted_path`
|
||||
- `selected_ids`
|
||||
**⚠️ item_id 前缀注意:**
|
||||
The `candidates` array IDs use `cand:sha256:xxx` format, but individual extracted item files use `sha256:xxx` (no `cand:` prefix). When passing `selected_ids` to article-summary tools, strip the `cand:` prefix. If the ID doesn't match, the summary tool won't find the article.
|
||||
|
||||
**Path resolution for article-summary:**
|
||||
After a pipeline run via MCP (Feishu context), the extracted files may live at an MCP-managed temp path, not the expected `outputs/freshrss/rerun/<run-id>/extracted/`. Before calling `generate_article_summaries` or `start_article_summary_job`, verify the extracted path exists. If not, use the `run_id` from the pipeline result to locate the actual output directory through `get_run_report`, or copy the files from the MCP-managed path.
|
||||
|
||||
Single-item rule:
|
||||
|
||||
- when `extracted_path` is `outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json`, call one summary job per file
|
||||
- in that case, `selected_ids` should contain only the matching single `item_id`
|
||||
- if the candidate ID came from the delivery payload, strip the `cand:` prefix before passing it
|
||||
|
||||
Recommended output layout:
|
||||
|
||||
- `outputs/freshrss/single_summaries/YYYY-MM-DD/`
|
||||
- async job state: `outputs/freshrss/article_summary_jobs/<job_id>/`
|
||||
|
||||
Recommended production sequence:
|
||||
|
||||
1. call `start_article_summary_job`
|
||||
2. poll `get_article_summary_job_status` until `status` becomes `success` or `failed`
|
||||
3. on success, call `get_article_summary_job_result` and continue downstream from `written_paths`
|
||||
|
||||
Hard fallback rule:
|
||||
|
||||
- If the async MCP job path returns timeout / transport failure / job-launch failure (for example MCP timeout while the reader article-summary workflow itself is still healthy), do not treat that as article-summary business failure.
|
||||
- Immediately retry through the local reader environment under `/home/ubuntu/zhu/github/reader` using the project `.venv`, calling the article-summary workflow directly.
|
||||
- The production goal is successful generation of the selected-article Markdown files; async MCP job is preferred, but local `.venv` execution is the required fallback path.
|
||||
|
||||
Synchronous helper:
|
||||
|
||||
- `generate_article_summaries` remains available for debug / light validation only, not as the default production path.
|
||||
|
||||
CLI fallback:
|
||||
CLI 仅在异步 MCP 不可用或人工排障时使用:
|
||||
|
||||
```bash
|
||||
python scripts/run_article_summaries.py \
|
||||
--extracted outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json \
|
||||
--ids <item_id_without_cand_prefix> \
|
||||
--output-dir outputs/freshrss/single_summaries/YYYY-MM-DD
|
||||
--extracted <returned-extracted-path> \
|
||||
--ids <full-item-id> \
|
||||
--output-dir <explicit-output-dir>
|
||||
```
|
||||
|
||||
### 6. Upload selected summaries to IMA
|
||||
## 5. IMA 上传
|
||||
|
||||
Upload only the generated markdown files for the selected articles.
|
||||
Once the user has selected the articles to retain, treat that selection itself as the authorization to continue the IMA deposition step; do not ask for a second confirmation about uploading into the knowledge base.
|
||||
用户在知识沉淀阶段的文章选择即为上传授权。
|
||||
|
||||
**⚠️ IMA upload 500 retry: Always wrap IMA uploads in a retry loop.**
|
||||
IMA's OpenAPI may return HTTP 500 on the first attempt. If the first upload fails, wait ~3 seconds and retry. Normally the second attempt succeeds. If `cos-upload.cjs` is used (for CDN-backed uploads), check subprocess stderr even when exit code is 0 — an HTTP error in the upload service may still produce exit code 0.
|
||||
执行顺序:
|
||||
|
||||
Hard execution rules before upload:
|
||||
1. 检查生成的 Markdown 与来源;
|
||||
2. 文件名规范化为 `<完整文章标题>.md`;
|
||||
3. 确认目标为 `daily` knowledge base;
|
||||
4. 执行 preflight、create_media、COS upload、add_knowledge;
|
||||
5. 验证知识库条目存在。
|
||||
|
||||
1. reformat/check the generated markdown into IMA-facing final content
|
||||
2. normalize the final upload filename to `<文章标题>.md`
|
||||
3. do not use internal temp names such as `ima-*`, `item-*`, `summary-*`, or English slug filenames as the final uploaded object name
|
||||
4. if the knowledge base already contains the same filename, append a timestamp suffix before `.md`
|
||||
5. if local work needs internal temp names, create a final upload copy with the user-facing title before calling IMA APIs
|
||||
上传格式与 API 参数分别见:
|
||||
|
||||
Default target knowledge base:
|
||||
- `ima-format-quickref.md`
|
||||
- `ima-upload-api.md`
|
||||
- `ima-credential-chain.md`
|
||||
|
||||
- `daily`
|
||||
- Read `IMA_DAILY_KNOWLEDGE_BASE_ID` / `IMA_DAILY_KNOWLEDGE_BASE_NAME` from reader `.env`
|
||||
- Verify the configured target at runtime before upload
|
||||
- If the configured target is unavailable, resolve by name `daily`; if still absent, create `daily`
|
||||
失败时记录发生在 preflight、create_media、COS upload 或 add_knowledge 的具体阶段。不要把失败的 Markdown 改写成 URL 导入或 Notes 类型来绕过错误。
|
||||
|
||||
## Hard rules recap
|
||||
## 6. CLI fallback 原则
|
||||
|
||||
- Never generate a digest from placeholder or example data when a real run is expected.
|
||||
- For normal runs, mark processed FreshRSS items as read unless the user explicitly requested a debug/test/validation run.
|
||||
- Public digest goes to Hugo; internal review digest goes to chat; neither full digest goes to IMA.
|
||||
- Only explicitly user-selected articles go to IMA.
|
||||
- All daily IMA deposition must go directly into the IMA knowledge-base path as Markdown knowledge items (`media_type=7`), not through the IMA notes path.
|
||||
- Uploading to IMA notes, or creating notes first and then linking them into a knowledge base, does not count as SOP completion.
|
||||
- Selected article summaries use extracted text, not live refetch.
|
||||
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization.
|
||||
- Final IMA upload filenames must use user-facing article titles, not internal slugs or workflow temp names.
|
||||
CLI 仅在以下场景使用:
|
||||
|
||||
- MCP 服务不可用;
|
||||
- Tool transport/launch 失败且无法取得有效 Job;
|
||||
- 用户明确要求本地调试;
|
||||
- 人工排障需要直接检查脚本输出。
|
||||
|
||||
如果已取得 `job_id` 或 `run_id`,先查询真实状态,避免因响应超时重复启动任务。
|
||||
|
||||
@@ -1,55 +0,0 @@
|
||||
# IMA COS 上传凭证处理(Hermes redact_secrets 兼容模式)
|
||||
|
||||
## 问题
|
||||
|
||||
Hermes 配置 `security.redact_secrets: true` 时,`create_media` API 返回的 `cos_credential.token` 字段在 `terminal()` 输出中被替换为 `***`。
|
||||
|
||||
直接通过 `terminal()` 调用 `cos-upload.cjs` 会因 token 截断而失败(HTTP 403 InvalidAccessKeyId)。
|
||||
|
||||
## 安全的工作流(execute_code + subprocess.run)
|
||||
|
||||
不要用 `terminal()` 传递 COS 凭证。改用 `execute_code()` + `subprocess.run()` 模式:
|
||||
|
||||
```python
|
||||
# Phase A: terminal() 中保存原始响应到文件
|
||||
result = terminal("""
|
||||
bash -c '
|
||||
set -a
|
||||
source /home/ubuntu/zhu/github/reader/.env
|
||||
set +a
|
||||
OPTS=$(printf "%s" "{\\"clientId\\":\\""$IMA_OPENAPI_CLIENTID"\\",\\\"apiKey\\":\\\""$IMA_OPENAPI_APIKEY"\\\"}")
|
||||
RESP=$(node /root/.hermes/skills/openclaw-imports/ima-skill/ima_api.cjs "openapi/wiki/v1/create_media" "{...}" "$OPTS" 2>/dev/null)
|
||||
echo "$RESP" > /tmp/create_media_raw.json
|
||||
echo "saved"
|
||||
'
|
||||
""")
|
||||
|
||||
# Phase B: execute_code 中从文件读取凭证
|
||||
import json, subprocess
|
||||
with open("/tmp/create_media_raw.json") as f:
|
||||
data = json.load(f)
|
||||
cred = data["data"]["cos_credential"]
|
||||
|
||||
# Phase C: subprocess.run 直接调用,不经过 terminal()
|
||||
args = ["node", f"{SKILL_DIR}/knowledge-base/scripts/cos-upload.cjs",
|
||||
"--file", FILE_PATH, "--secret-id", cred["secret_id"], "--secret-key", cred["secret_key"],
|
||||
"--token", cred["token"], "--bucket", cred["bucket_name"], "--region", cred["region"],
|
||||
"--cos-key", cred["cos_key"], "--content-type", "text/markdown",
|
||||
"--start-time", cred["start_time"], "--expired-time", cred["expired_time"], "--timeout", "300000"]
|
||||
r = subprocess.run(args, capture_output=True, text=True, timeout=310)
|
||||
```
|
||||
|
||||
## 对比:错误的做法(terminal 直接传递)
|
||||
|
||||
```bash
|
||||
# ⛔ 这样不行!token 会被 redact_secrets 替换为 ***
|
||||
TOKEN=$(echo "$RESP" | jq -r '.data.cos_credential.token')
|
||||
node ... --token "$TOKEN" ... # 会收到 HTTP 403
|
||||
```
|
||||
|
||||
## 关键原则
|
||||
|
||||
- `terminal()` 输出中的敏感字段会被自动脱敏,但不影响底层 JSON 文件写入
|
||||
- `execute_code` 中 `terminal()` 返回的 `output` 已经是脱敏后的文本
|
||||
- **唯一可靠的凭证源**是直接写入磁盘的原始 JSON 文件
|
||||
- `subprocess.run` 在 `execute_code` 中绕过脱敏,因为凭证在 Python 内存中直接被传递给子进程,不经过 Hermes 的 stdout 脱敏管道
|
||||
@@ -1,46 +1,27 @@
|
||||
# IMA 凭证链:从 reader `.env` 到 IMA 上传
|
||||
# IMA 凭证与安全边界
|
||||
|
||||
## 凭证来源
|
||||
## 必需配置
|
||||
|
||||
IMA 上传所需的凭证存储在多个位置,优先级如下:
|
||||
- `IMA_OPENAPI_CLIENTID`
|
||||
- `IMA_OPENAPI_APIKEY`
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_ID`
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_NAME=daily`
|
||||
|
||||
| 优先级 | 位置 | 说明 |
|
||||
|--------|------|------|
|
||||
| 1 | 环境变量 `IMA_OPENAPI_CLIENTID` / `IMA_OPENAPI_APIKEY` | Hermes session 级 |
|
||||
| 2 | `~/.config/ima/client_id` / `api_key` | ima-skill 的默认检查路径 |
|
||||
| 3 | `/home/ubuntu/zhu/github/reader/.env` | reader 项目配置,含完整的 IMA 凭证和 KB ID |
|
||||
优先使用当前进程环境和 IMA Skill 已支持的凭证加载机制。不要在本 Skill 中复制、迁移或重写密钥文件。
|
||||
|
||||
## 凭证内容(reader .env 中)
|
||||
## 缺失处理
|
||||
|
||||
```
|
||||
IMA_OPENAPI_CLIENTID=<32位hex>
|
||||
IMA_OPENAPI_APIKEY=<base64编码的API密钥>
|
||||
IMA_DAILY_KNOWLEDGE_BASE_ID=<base64编码的KB ID>
|
||||
IMA_DAILY_KNOWLEDGE_BASE_NAME=daily
|
||||
```
|
||||
Preflight 返回凭证缺失或目标知识库无法解析时:
|
||||
|
||||
## 缺失时的处理流程
|
||||
1. 停止上传;
|
||||
2. 只报告缺失的变量名或配置项;
|
||||
3. 等待用户或运行环境补齐配置;
|
||||
4. 配置恢复后重新执行 preflight,不重复生成知识笔记。
|
||||
|
||||
当 IMA 上传失败(`-100` 凭证缺失错误)时:
|
||||
## 安全边界
|
||||
|
||||
1. 从 reader `.env` 读取凭证:
|
||||
```
|
||||
grep -E '^(IMA_OPENAPI_CLIENTID|IMA_OPENAPI_APIKEY)=' /home/ubuntu/zhu/github/reader/.env
|
||||
```
|
||||
2. 同步到 ima-skill 默认检查路径:
|
||||
```
|
||||
echo "<client_id>" > ~/.config/ima/client_id
|
||||
echo "<api_key>" > ~/.config/ima/api_key
|
||||
```
|
||||
3. (可选)追加到 Hermes `.env` 以全局生效:
|
||||
```
|
||||
echo "IMA_OPENAPI_CLIENTID=<client_id>" >> /root/.hermes/.env
|
||||
echo "IMA_OPENAPI_APIKEY=<api_key>" >> /root/.hermes/.env
|
||||
echo "IMA_DAILY_KNOWLEDGE_BASE_ID=<kb_id>" >> /root/.hermes/.env
|
||||
echo "IMA_DAILY_KNOWLEDGE_BASE_NAME=daily" >> /root/.hermes/.env
|
||||
```
|
||||
|
||||
## 执行注意事项
|
||||
|
||||
- **COS 凭证红线**:`create_media` 返回的 `cos_credential` 中的 `token`/`secret_id`/`secret_key` 在 `terminal()` 输出中会被 Hermes 替换为 `***`。必须用 `subprocess.run()` 捕获原始输出,或用 `execute_code` 内联操作。
|
||||
- **凭证格式**:`api_key` 是 base64 字符串(76 字符),`kb_id` 也是 base64 字符串。不要截断或转码。
|
||||
- 不在聊天、日志或命令输出中打印完整 API key、KB ID 或 COS 临时凭证;
|
||||
- `create_media` 返回的 COS 凭证仅在同一受控进程内传给上传工具,不写入磁盘;
|
||||
- 不通过拼接 Shell 字符串传递凭证,使用参数数组或 IMA Skill 的封装;
|
||||
- 不绕过 Hermes 的脱敏机制;若现有工具链无法安全传递凭证,停止并报告;
|
||||
- 上传结束后不持久化 COS 临时凭证。
|
||||
|
||||
@@ -1,97 +0,0 @@
|
||||
# IMA Markdown 文件上传流程(media_type=7)
|
||||
|
||||
## 什么时候用此流程
|
||||
|
||||
当用户说"沉淀到 IMA"时,必须用此文件上传流程,而不是 URL 导入或笔记导入。
|
||||
|
||||
## 完整流程
|
||||
|
||||
### 1. 写 .md 文件
|
||||
|
||||
```bash
|
||||
mkdir -p /tmp/ima_upload
|
||||
cat > /tmp/ima_upload/文章标题.md << 'EOF'
|
||||
# 文章标题
|
||||
|
||||
> 来源:XXX
|
||||
|
||||
## 摘要
|
||||
|
||||
...
|
||||
|
||||
## 核心亮点
|
||||
|
||||
- ...
|
||||
|
||||
[原文链接](url)
|
||||
EOF
|
||||
```
|
||||
|
||||
### 2. preflight 检查
|
||||
|
||||
```python
|
||||
pf = subprocess.run(
|
||||
["node", f"{SKILL_DIR}/knowledge-base/scripts/preflight-check.cjs",
|
||||
"--file", filepath],
|
||||
capture_output=True, text=True
|
||||
)
|
||||
meta = json.loads(pf.stdout)
|
||||
# meta = {pass, file_name, file_ext, file_size, media_type, content_type}
|
||||
```
|
||||
|
||||
### 3. create_media(获取 COS 凭证)
|
||||
|
||||
```python
|
||||
r = ima_api("openapi/wiki/v1/create_media", {
|
||||
"file_name": fname,
|
||||
"file_size": meta["file_size"],
|
||||
"content_type": meta["content_type"],
|
||||
"knowledge_base_id": kb_id,
|
||||
"file_ext": meta["file_ext"]
|
||||
})
|
||||
cos = r["data"]["cos_credential"]
|
||||
media_id = r["data"]["media_id"]
|
||||
```
|
||||
|
||||
### 4. COS 上传
|
||||
|
||||
```python
|
||||
cu = subprocess.run([
|
||||
"node", f"{SKILL_DIR}/knowledge-base/scripts/cos-upload.cjs",
|
||||
"--file", filepath,
|
||||
"--secret-id", cos["secret_id"],
|
||||
"--secret-key", cos["secret_key"],
|
||||
"--token", cos["token"],
|
||||
"--bucket", cos["bucket_name"],
|
||||
"--region", cos["region"],
|
||||
"--cos-key", cos["cos_key"],
|
||||
"--content-type", meta["content_type"],
|
||||
"--start-time", str(cos["start_time"]),
|
||||
"--expired-time", str(cos["expired_time"]),
|
||||
"--timeout", "300000"
|
||||
], capture_output=True, text=True, timeout=30)
|
||||
# 非0退出 = 上传失败
|
||||
```
|
||||
|
||||
### 5. add_knowledge
|
||||
|
||||
```python
|
||||
r = ima_api("openapi/wiki/v1/add_knowledge", {
|
||||
"media_type": 7,
|
||||
"media_id": media_id,
|
||||
"title": title, # 文章中文标题
|
||||
"knowledge_base_id": kb_id,
|
||||
"file_info": {
|
||||
"cos_key": cos["cos_key"],
|
||||
"file_size": meta["file_size"],
|
||||
"file_name": fname
|
||||
}
|
||||
})
|
||||
```
|
||||
|
||||
## 注意事项
|
||||
|
||||
- **COS 凭证不能通过 terminal() 读取**(Hermes redact_secrets 会把 token 替换成 `***`)。必须用 Python `subprocess.run()` 捕获原始 stdout。
|
||||
- **create_media 的 content_type 不能带 charset 参数**(如 `text/markdown; charset=utf-8` 会被拒)。用纯 MIME 类型 `text/markdown`。
|
||||
- **文件命名 = `<文章中文标题>.md`**。不要用英文 slug 或 temp name。
|
||||
- **media_type=7** 是 Markdown 文件。**绝对不要用 media_type=11**(笔记)或 `import_urls`。
|
||||
@@ -1,42 +1,47 @@
|
||||
# IMA 上传格式速查表(日报沉淀专用)
|
||||
# IMA Markdown 格式速查
|
||||
|
||||
## 文件名 vs 标题 vs 内容对照表
|
||||
## 文件与标题
|
||||
|
||||
| 项目 | ✅ 正确 | ❌ 错误 |
|
||||
|------|---------|---------|
|
||||
| **文件名** | `从Vibe Coding到Harness—— 一套大仓AI工程化实战.md` | `大仓AI工程化实战.md` |
|
||||
| **add_knowledge title** | `从Vibe Coding到Harness—— 一套大仓AI工程化实战` | `大仓AI工程化实战` / `大仓AI工程化实战.md` |
|
||||
| **上传方式** | `preflight` → `create_media` → `cos-upload` → `add_knowledge(media_type=7)` | `import_doc` + `add_knowledge(media_type=11)` |
|
||||
| **正文内容** | pipeline `extracted/` 的 `article.content` 或 `summary-batch.json` 的 `summary` | 自己写的两三句话 |
|
||||
| **正文长度** | ≥ 500 字 | < 500 字 |
|
||||
- 文件名:`<完整文章标题>.md`
|
||||
- `add_knowledge.title`:完整文章标题,不包含 `.md`
|
||||
- 上传类型:Markdown 文件,`media_type=7`
|
||||
- 目标:`daily` knowledge base
|
||||
|
||||
## 文件名常见错误模式
|
||||
## 内容来源
|
||||
|
||||
| 原始标题 | ❌ 错误文件名 | ✅ 正确文件名 |
|
||||
|----------|-------------|-------------|
|
||||
| `从Vibe Coding到Harness—— 一套大仓AI工程化实战` | `大仓AI工程化实战.md` | `从Vibe Coding到Harness—— 一套大仓AI工程化实战.md` |
|
||||
| `契约化多端架构:基于领域模型的Harness实践` | `契约化多端架构Harness实践.md` | `契约化多端架构:基于领域模型的Harness实践.md` |
|
||||
| `Loop Engineering 实战:从日志扫描到预发部署的全自主闭环` | `LoopEngineering实战.md` | `Loop Engineering 实战:从日志扫描到预发部署的全自主闭环.md` |
|
||||
只能使用当前 Run 的可追踪内容:
|
||||
|
||||
## 5 Section 格式要求
|
||||
1. extracted Artifact 的 `article.plain_text`;
|
||||
2. 对应文章的结构化摘要;
|
||||
3. digest brief 的 summary 与 highlights。
|
||||
|
||||
| Section | 格式 | 说明 |
|
||||
|---------|------|------|
|
||||
| **核心结论** | 一段连贯段落(非分点) | 总结文章核心发现或主张 |
|
||||
| **主要论点** | 一段连贯段落(非分点) | 综合多个论点形成连贯叙述 |
|
||||
| **关键方法 / 机制** | `- **方法名**:详细说明` 分点 | 每条展开到能理解原理的程度 |
|
||||
| **重要细节** | `- 每条一个带解释的完整知识点` 分点 | 每条是一个完整知识点,不是关键词 |
|
||||
| **可复用启发** | `- 每条 actionable 的实践启示` 分点 | 附应用场景说明 |
|
||||
不得使用通用知识补写原文没有的信息,不设置固定字数或字节数门槛。内容较短时保持简洁并忠于来源。
|
||||
|
||||
## 底部额外 section
|
||||
## 标准结构
|
||||
|
||||
- `## 关键词` — 标签式关键词列表
|
||||
- `## 主题` — 主题分类
|
||||
```markdown
|
||||
# 完整文章标题
|
||||
|
||||
## 速查口诀
|
||||
Source: https://原文链接
|
||||
Category: 分类
|
||||
|
||||
> 文件名 = 完整标题.md
|
||||
> 内容取 pipeline,不自己写
|
||||
> 5 个 section,核心结论和主要论点是段落
|
||||
> media_type = 7,不是 11
|
||||
> title 不加 .md
|
||||
## 核心结论
|
||||
|
||||
## 主要论点
|
||||
|
||||
## 关键方法 / 机制
|
||||
|
||||
## 重要细节
|
||||
|
||||
## 可复用启发
|
||||
|
||||
## 关键词
|
||||
|
||||
## 主题
|
||||
```
|
||||
|
||||
- 核心结论和主要论点使用连贯段落;
|
||||
- 方法、细节和启发按完整知识点分项;
|
||||
- 没有来源支持的 Section 可以简写,不得编造内容填充。
|
||||
|
||||
上传 API 见 `ima-upload-api.md`。
|
||||
|
||||
@@ -1,38 +0,0 @@
|
||||
# AI Agent 的 Skill 系统设计
|
||||
|
||||
Source: https://mp.weixin.qq.com/s?__biz=MzAxNDEwNjk5OQ==&mid=2650544717&idx=1&sn=b578abf5a81034670900a3b8eb874296
|
||||
Category: 方法论
|
||||
|
||||
## 核心结论
|
||||
好的 Skill 应是一个小而准的行为系统,通过触发、加载、执行、约束、验证和迭代的组织,将通用 Agent 转化为在特定任务上稳定可靠的专用 Agent。核心原则是上下文窗口是公共资源,必须采用渐进披露、按任务风险设置自由度,并通过真实任务前向测试来证明行为改变。
|
||||
|
||||
## 主要论点
|
||||
Skill 设计的本质是行为编程而非文档编写,需要将期望行为转化为 Agent 能稳定执行的结构化工作流。为此,必须同时解决发现(正确场景触发)、加载(最小上下文)、执行(合适自由度)和验证(真实任务测试)四件事,并通过门控、脚本外化、测试防合理化等机制确保 Agent 在复杂压力下不走捷径。
|
||||
|
||||
## 关键方法 / 机制
|
||||
- 渐进披露的三层内容加载:元数据(frontmatter 的 name/description)用于发现,正文(SKILL.md)用于执行,资源(scripts/references/assets)按需读取。description 只做路由器,不包含完整流程,防止 Agent 凭印象执行。
|
||||
- 门控机制(HARD-GATE):在低自由度任务中,使用明确的 <HARD-GATE> 标签禁止 Agent 在条件满足前执行后续动作,减少解释空间,让关键路径更像程序而非建议。
|
||||
- 脚本外化降低上下文消耗和行为漂移:将需要确定性的操作(如 PDF 旋转)封装为 scripts/ 中的可执行脚本,避免 Agent 每次临时生成代码,提升可靠性和节省 token。
|
||||
- 基于 TDD 的前向测试方法:用子代理模拟真实用户任务,只给原始任务和最少上下文,不泄露预期结论;观察行为轨迹、输出文件等原始证据,发现并封堵 Agent 的违规行为。
|
||||
- 反模式自查表与检查表:交付前检查触发条件、自由度设置、资源引用、验证流程等关键项,确保 Skill 不是一份草稿而是一个可用的能力包。
|
||||
- 跨平台适配与优雅降级:Skill 应写行为规则(如 TodoWrite),再通过平台层映射到具体工具名(如 todowrite);平台能力不足时优雅降级,保持 Skill 的可迁移性。
|
||||
|
||||
## 重要细节
|
||||
- SKILL.md 的 frontmatter 和正文职责分离:name/description 用于发现(Agent 触发前可见),正文用于执行(触发后加载)。如果触发条件写在正文里,Agent 在决定是否触发时根本读不到。
|
||||
- 命名规范:短、可触发、动词优先,例如 create-skill 比 skill-creation 更好,这本质上是路由质量——Agent 在技能库里找能力时,name/description 是第一层索引。
|
||||
- 资源组织原则——“信息只放一个地方”:不要在 SKILL.md 和 references/ 中重复同一段规则,重复会带来漂移,导致 Agent 在两个版本间自行解释,增加维护成本。
|
||||
- 门控类型示例:先决条件门控(先理解例子再编辑)、并发冲突门控、未保存内容门控、敏感操作门控(已创建 Skill 需处理影响再修改)。
|
||||
- 流程图用 GraphViz DOT 嵌入 Markdown:对于包含非线性判断、循环、回退的步骤,流程图比纯文本更稳定,能防止 Agent 遗漏关键分支。
|
||||
- 验证时防“合理化”问题:AI Agent 在压力下会为跳过规则编造理由,Skill 需要提前写出这些借口并给出反驳;审查循环应围绕真实失败风险而非措辞偏好。
|
||||
|
||||
## 可复用启发
|
||||
- “上下文窗口是公共资源”原则:设计任何 Agent 指令时,每段内容都要质疑“Agent 真的需要这段解释吗?”和“值得占用的 token 成本吗?”,这适用于提示词、系统消息等所有 Agent 输入设计。
|
||||
- 先收集具体例子再抽象 Skill:不要从抽象能力开始写,而是先收集用户会怎么触发、哪些请求应该触发/不应该触发、成功输出是什么等具体场景,避免写出宽泛不可执行的指令。
|
||||
- 用脚本固化确定性任务、用门控防止关键路径走捷径:对于高脆弱、低变化空间的任务(如文件格式转换),应使用脚本而非描述性建议;对于必须按顺序执行的步骤,用门控打断 Agent 的“合理化”冲动。
|
||||
- 设计防合理化的测试流程:用子代理模拟真实用户,只给原始任务,不泄露预期答案;观察是否存在只有看到结论才能成功的情况——如果这样,说明 Skill 不够清楚或测试设置泄露答案。
|
||||
|
||||
## 关键词
|
||||
SKILL.md、YAML、Markdown、DOT、GraphViz、TDD、HARD-GATE、quick_validate.py
|
||||
|
||||
## 主题
|
||||
AI Agent、行为编程、Token 经济、系统设计、约束机制
|
||||
@@ -1,33 +0,0 @@
|
||||
# IMA 笔记格式参考
|
||||
|
||||
> ⚠️ 格式基准文件为 `references/ima-format-reference.md`(`ai-agent-的-skill-系统设计.md`),每次生成 IMA 沉淀前必须先读该文件。
|
||||
|
||||
## 标准结构
|
||||
|
||||
### 顶部元数据
|
||||
|
||||
```markdown
|
||||
# 文章完整标题
|
||||
|
||||
Source: https://原文链接(纯 URL,不加 "原文链接:" 标签)
|
||||
Category: 分类
|
||||
```
|
||||
|
||||
### 5 个必含 Section
|
||||
|
||||
1. **核心结论** — 一段总结文章核心发现或主张的段落(不是分点),像基准文件一样是一段连贯文字
|
||||
2. **主要论点** — 一段概括文章核心论述的段落(不是分点列表),综合多个论点形成连贯叙述
|
||||
3. **关键方法 / 机制** — `**方法名**:详细说明` 的格式,每条展开到能理解其原理的程度
|
||||
4. **重要细节** — 每条一个带解释的完整知识点,不是关键词
|
||||
5. **可复用启发** — 每条 actionable 的实践启示,附应用场景说明
|
||||
|
||||
### 底部额外 Section
|
||||
|
||||
- **## 关键词** — 标签式关键词列表
|
||||
- **## 主题** — 主题分类
|
||||
|
||||
### 质量要求
|
||||
|
||||
- 每篇文章总内容量应达到 **4,000+ bytes**(基准文件 4,782B)
|
||||
- 每个 section 展开到完整的知识点级别,不能只列关键词
|
||||
- 内容来源:优先 `item-XX.extracted.json` 的 `article.content`,其次 `summary-batch.json` 的 `summary` 字段
|
||||
@@ -1,80 +1,91 @@
|
||||
# IMA Upload API Reference
|
||||
# IMA Markdown 上传 API
|
||||
|
||||
Full API flow for uploading markdown articles to the IMA `daily` knowledge base. Used in Phase 7 of the reader-digest-flow.
|
||||
用于将用户选中的单篇 Markdown 知识笔记上传到 `daily` knowledge base。
|
||||
|
||||
## Credentials
|
||||
## 凭证
|
||||
|
||||
```
|
||||
IMA_OPENAPI_CLIENTID - from reader .env or user-provided
|
||||
IMA_OPENAPI_APIKEY - from reader .env or user-provided
|
||||
IMA_DAILY_KNOWLEDGE_BASE_ID - daily KB UUID
|
||||
- `IMA_OPENAPI_CLIENTID`
|
||||
- `IMA_OPENAPI_APIKEY`
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_ID`
|
||||
- `IMA_DAILY_KNOWLEDGE_BASE_NAME`
|
||||
|
||||
凭证定位和恢复见 `ima-credential-chain.md`。不要在终端输出完整密钥。
|
||||
|
||||
## 上传前检查
|
||||
|
||||
- 文件名为 `<完整文章标题>.md`;
|
||||
- `title` 为完整文章标题,不带 `.md`;
|
||||
- Markdown 符合 `ima-format-quickref.md`;
|
||||
- 内容可追溯到当前 Run Artifact;
|
||||
- 用户已经明确选择该文章;
|
||||
- 目标知识库已经解析并验证。
|
||||
|
||||
## 1. Preflight
|
||||
|
||||
调用 IMA Skill 的 `preflight-check.cjs` 检查文件类型、扩展名、大小和 MIME。
|
||||
|
||||
预期:
|
||||
|
||||
```text
|
||||
file_ext=md
|
||||
content_type=text/markdown
|
||||
media_type=7
|
||||
```
|
||||
|
||||
The `ima-skill` v1.1.7+ ships with `ima_api.cjs` for credential loading.
|
||||
Legacy auth header: `ima-openapi-ctx: skill_version=1.1.7`.
|
||||
## 2. Create Media
|
||||
|
||||
Do NOT export the full API key in shell commands — use `execute_code` with `subprocess.run` and Python string variables.
|
||||
|
||||
## Flow (3 steps)
|
||||
|
||||
### 1. create_media
|
||||
|
||||
```
|
||||
POST https://ima.qq.com/openapi/wiki/v1/create_media
|
||||
Headers: ima-openapi-clientid, ima-openapi-apikey, Content-Type: application/json
|
||||
Body: { file_name, file_size, content_type, knowledge_base_id, file_ext }
|
||||
Returns: { code: 0, data: { media_id, cos_credential: { secret_id, secret_key, token, bucket_name, region, cos_key, start_time, expired_time } } }
|
||||
```text
|
||||
POST /openapi/wiki/v1/create_media
|
||||
```
|
||||
|
||||
`file_ext` is without the dot (e.g. `md` not `.md`).
|
||||
`file_name` must be the user-facing article title + `.md`.
|
||||
`content_type` for markdown is `text/markdown`; media_type=7.
|
||||
请求核心字段:
|
||||
|
||||
### 2. COS upload
|
||||
|
||||
Use `cos-upload.cjs` from `ima-skill/knowledge-base/scripts/`:
|
||||
|
||||
```
|
||||
node <skill_dir>/knowledge-base/scripts/cos-upload.cjs \
|
||||
--file <local_md_file> \
|
||||
--secret-id <from create_media> \
|
||||
--secret-key <from create_media> \
|
||||
--token <from create_media> \
|
||||
--bucket <bucket_name> \
|
||||
--region <region> \
|
||||
--cos-key <cos_key> \
|
||||
--content-type text/markdown \
|
||||
--start-time <start_time> \
|
||||
--expired-time <expired_time>
|
||||
```json
|
||||
{
|
||||
"file_name": "<完整文章标题>.md",
|
||||
"file_size": 0,
|
||||
"content_type": "text/markdown",
|
||||
"knowledge_base_id": "<daily-kb-id>",
|
||||
"file_ext": "md"
|
||||
}
|
||||
```
|
||||
|
||||
⚠️ Must use Python `subprocess.run(args=[...])` to avoid shell parameter mangling.
|
||||
⚠️ Always capture `returncode` and `stderr` — COS may return exit 0 on HTTP 500.
|
||||
保存返回的 `media_id` 和 `cos_credential`。COS 临时凭证只在进程内传递,不打印到聊天或日志。
|
||||
|
||||
### 3. add_knowledge
|
||||
## 3. COS Upload
|
||||
|
||||
```
|
||||
POST https://ima.qq.com/openapi/wiki/v1/add_knowledge
|
||||
Headers: same as create_media
|
||||
Body: { media_type: 7, media_id, title: "<file_name>", knowledge_base_id, file_info: { cos_key, file_size, file_name } }
|
||||
使用 IMA Skill 提供的 `cos-upload.cjs`,通过参数数组调用并检查:
|
||||
|
||||
- 进程 `returncode`;
|
||||
- `stderr`;
|
||||
- HTTP 上传结果。
|
||||
|
||||
不要拼接包含凭证的 Shell 字符串,也不要把多条 JSON 响应重定向到同一个文件。
|
||||
|
||||
## 4. Add Knowledge
|
||||
|
||||
```text
|
||||
POST /openapi/wiki/v1/add_knowledge
|
||||
```
|
||||
|
||||
`media_type=7` for markdown. `title` MUST equal `file_name`.
|
||||
核心字段:
|
||||
|
||||
## Article Markdown reformatting (before upload)
|
||||
|
||||
Generated summaries from `reader` have `Source:` and `Category:` header lines.
|
||||
Before uploading, reformat to IMA style:
|
||||
|
||||
```
|
||||
原文链接:<original article URL>
|
||||
|
||||
## 核心结论
|
||||
...
|
||||
|
||||
## 主要论点
|
||||
...
|
||||
```json
|
||||
{
|
||||
"media_type": 7,
|
||||
"media_id": "<media-id>",
|
||||
"title": "<完整文章标题>",
|
||||
"knowledge_base_id": "<daily-kb-id>",
|
||||
"file_info": {
|
||||
"cos_key": "<cos-key>",
|
||||
"file_size": 0,
|
||||
"file_name": "<完整文章标题>.md"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Remove `Source:`, `Category:` lines. Keep `原文链接:` at top with the URL on the next line.
|
||||
Break long prose (>200 chars per paragraph) into shorter paragraphs for IMA readability.
|
||||
## 5. 验证
|
||||
|
||||
上传成功不能只依据本地命令退出码。验证目标 knowledge base 中存在对应标题或返回对象,并向用户报告最终结果。
|
||||
|
||||
失败时明确报告发生在 Preflight、Create Media、COS Upload 或 Add Knowledge 的哪一步,不改用 URL 导入或 Notes 类型绕过错误。
|
||||
|
||||
@@ -1,195 +1,29 @@
|
||||
# 关键词引擎维护流程
|
||||
# 关键词治理路由
|
||||
|
||||
## 概述
|
||||
关键词治理不属于 `reader-digest-flow` 的日常执行阶段。
|
||||
|
||||
reader 项目内置了完整的关键词清洗链路,用于管理 `term_aliases.json`、`term_stopwords.json`、`filter_context.personal.json` 等配置。本文件说明链路各环节的职责与使用方式。
|
||||
仅当用户明确要求“清理关键词”“词库治理”“生成关键词建议”时,委托:
|
||||
|
||||
## 完整数据流
|
||||
|
||||
```
|
||||
每日日报 pipeline
|
||||
│
|
||||
▼
|
||||
term_index/daily/YYYY-MM-DD.json ← build_keyword_index.py (pipeline 末环节自动)
|
||||
│
|
||||
▼
|
||||
term_index/term_stats.json ← 全量汇总
|
||||
│
|
||||
▼
|
||||
build_review_bundle.py ← 打包审查数据包 (手动触发)
|
||||
│
|
||||
▼
|
||||
generate_term_cleanup_suggestions.py ← 生成建议 (手动触发)
|
||||
│
|
||||
▼
|
||||
apply_term_suggestions.py ← 人工确认后落盘 (手动触发)
|
||||
```text
|
||||
skills/keyword-cleanup-review/SKILL.md
|
||||
```
|
||||
|
||||
## 各环节命令
|
||||
Review 输入生成由该 Skill 定义,后续确认与 Apply 由当前编排层负责:
|
||||
|
||||
### 1. 重建每日 term_index(pipeline 末环节自动执行,也可手动)
|
||||
|
||||
```bash
|
||||
cd /home/ubuntu/zhu/github/reader
|
||||
.venv/bin/python3 scripts/build_keyword_index.py \
|
||||
--input outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json
|
||||
```text
|
||||
build review bundle
|
||||
→ generate rule suggestions
|
||||
→ optional semantic suggestions
|
||||
→ human review
|
||||
→ dry-run
|
||||
→ apply accepted suggestions
|
||||
```
|
||||
|
||||
### 2. 重建 review bundle(打包当前配置+统计供审查)
|
||||
约束:
|
||||
|
||||
```bash
|
||||
cd /home/ubuntu/zhu/github/reader
|
||||
.venv/bin/python3 skills/keyword-cleanup-review/scripts/build_review_bundle.py \
|
||||
--days 365 \
|
||||
--top 200 \
|
||||
--output outputs/term_index/review/keyword-cleanup-bundle.json
|
||||
```
|
||||
- Suggestions JSON 是 Review 与 Apply 之间的正式契约;
|
||||
- LLM 语义建议不能自动 Apply;
|
||||
- 不直接编辑 aliases、stopwords、watchlist 或 interest 配置;
|
||||
- 不在日报主流程中因 tag 质量不佳自动触发治理。
|
||||
|
||||
参数:
|
||||
- `--days 365` — 考虑最近多少天的统计数据
|
||||
- `--top 200` — 取前 N 个高频词纳入 bundle
|
||||
|
||||
### 3. 生成清洗建议
|
||||
|
||||
```bash
|
||||
cd /home/ubuntu/zhu/github/reader
|
||||
.venv/bin/python3 scripts/generate_term_cleanup_suggestions.py \
|
||||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
|
||||
--emit-markdown
|
||||
```
|
||||
|
||||
输出:
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json` — 正式建议产物
|
||||
- `outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.md` — 人工阅读展示稿
|
||||
|
||||
当前脚本能生成的建议类型:
|
||||
|
||||
| 类型 | 生成规则 | 当前产出 |
|
||||
|:----|:---------|:---------|
|
||||
| interest_keyword_suggestions | 高频未覆盖词(total≥3, days≥2) | ✅ 按阈值产出 |
|
||||
| watch_terms | 中频观察词(total≤2, days≤2) | ✅ 按阈值产出 |
|
||||
| alias_suggestions | 大小写变体/单复数/空格连词符差异 | ✅ 统计规则产出 |
|
||||
| stopword_suggestions | — | ❌ 脚本硬编码为空 |
|
||||
|
||||
### 4. 应用建议(dry-run → review → apply)
|
||||
|
||||
```bash
|
||||
# 先 dry-run 预览
|
||||
cd /home/ubuntu/zhu/github/reader
|
||||
.venv/bin/python3 scripts/apply_term_suggestions.py \
|
||||
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
|
||||
--accept-interest 强化学习 ReAct CLI \
|
||||
--dry-run
|
||||
|
||||
# 确认后正式 apply(去掉 --dry-run)
|
||||
.venv/bin/python3 scripts/apply_term_suggestions.py \
|
||||
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
|
||||
--accept-interest 强化学习 ReAct CLI \
|
||||
```
|
||||
|
||||
支持的 accept 参数:
|
||||
- `--accept-interest 词1 词2 ...` — 添加到 interest_keywords
|
||||
- `--accept-watch 词1 词2 ...` — 添加到 watchlist
|
||||
- `--accept-alias 源词1 源词2 ...` — 添加 alias 映射
|
||||
- `--accept-stopword 词1 词2 ...` — 添加停用词
|
||||
|
||||
## 脚本能做什么 vs 不能做什么
|
||||
|
||||
### ✅ 脚本能做的(统计级清洗)
|
||||
|
||||
- 发现大小写变体(`vibe coding` → `Vibe Coding`)
|
||||
- 发现单复数差异(`Agent Skill` → `Agent Skills`)
|
||||
- 发现空格/连词符差异
|
||||
- 按频次推荐 should-be-interest / should-be-watch 的词
|
||||
- 批量 apply 到配置文件,自动记录变更日志
|
||||
|
||||
### ❌ 脚本不能做的(语义级清洗,需要人工判断)
|
||||
|
||||
- 识别产品名(`WorkBuddy`、`LibTV Agent`、`飞书妙搭`)→ 应加 stopword
|
||||
- 识别模型名(`Qwen3-30B-A3B`、`GLM 5.2`、`Opus 4.8`)→ 应加 stopword
|
||||
- 识别人名/地名(`Andrej Karpathy`、`成都天府长岛`)→ 应加 stopword
|
||||
- 英文专有名词归一化到中文主题词(`Multi-Agent` → `多Agent`)
|
||||
- 长产品名归一化(`火山云数据库PostgreSQL Serverless版` → `Serverless数据库`)
|
||||
- 判断某个词对 Hugo tag 是否有用(`1688`、`OPC训练营` → 无效)
|
||||
|
||||
## 语义级清洗:LLM 辅助(generate_term_cleanup_semantic_suggestions.py)
|
||||
|
||||
除了统计规则的清洗,项目还有一个 **LLM 驱动的语义级清洗脚本**,能发现统计规则做不到的事情:
|
||||
|
||||
```bash
|
||||
cd /home/ubuntu/zhu/github/reader
|
||||
.venv/bin/python3 scripts/generate_term_cleanup_semantic_suggestions.py \
|
||||
--bundle outputs/term_index/review/keyword-cleanup-bundle.json \
|
||||
--suggestions outputs/term_index/review/term-cleanup-suggestions-YYYY-MM-DD.json \
|
||||
--output outputs/term_index/review/term-cleanup-semantic-suggestions-YYYY-MM-DD.json
|
||||
```
|
||||
|
||||
参数:
|
||||
- `--bundle` — review bundle(先跑 build_review_bundle.py)
|
||||
- `--suggestions` — 规则层的建议 JSON(可选,供 LLM 避免重复)
|
||||
- `--output` — 输出路径(自动生成)
|
||||
- `--dry-run` — 只打印 prompt 不调 LLM
|
||||
|
||||
### 语义脚本能做的(而统计规则不能做的)
|
||||
|
||||
| 类型 | LLM 能发现什么 | 示例 |
|
||||
|:----|:--------------|:-----|
|
||||
| **semantic_alias** | 中英文映射、简称↔全称、近义词 | `AI智能体`→`AI Agent`, `AI Harness`→`Harness Engineering` |
|
||||
| **stopword** | 语义判断哪些词太泛/不相干 | `自动化`(太泛), `AlphaFold`(生物领域), `陌生化`(穿透词) |
|
||||
| **promote_to_interest** | 与用户关注领域对齐的新词 | `强化学习`, `Loop Engineering`, `MoE`, `ChatGPT` |
|
||||
|
||||
### 两阶段清洗 SOP
|
||||
|
||||
当需要清洗关键词时,按此顺序操作:
|
||||
|
||||
**阶段一:规则统计清洗(脚本发现 + 人工确认)**
|
||||
1. `build_review_bundle.py` → 重建 bundle
|
||||
2. `generate_term_cleanup_suggestions.py` → 产出统计级建议
|
||||
3. 检查建议,决定哪些 accept
|
||||
4. `apply_term_suggestions.py --dry-run` → 预览
|
||||
5. `apply_term_suggestions.py` → 正式落地
|
||||
|
||||
**阶段二:语义级清洗(LLM 发现 + 人工确认)**
|
||||
1. `generate_term_cleanup_semantic_suggestions.py` → 调 LLM 产生语义建议
|
||||
2. 检查 LLM 的 semantic_alias / stopword / promote_to_interest
|
||||
3. 人工编辑 term_aliases.json / term_stopwords.json(语义级判断不能自动化)
|
||||
|
||||
**🔴 不要绕过确认流程直接改配置。** 先跑脚本看看发现了什么,再针对性改。未经审视的批量改配置会引入新问题。
|
||||
|
||||
## 完整的语义级清洗操作流程
|
||||
|
||||
当需要大量添加 aliases/stopwords 时,推荐流程:
|
||||
|
||||
1. **跑一次 `build_review_bundle.py` + `generate_term_cleanup_suggestions.py`** 看看统计建议
|
||||
2. **走一遍实际 pipeline 产出**,收集所有 unique keywords:
|
||||
```bash
|
||||
for rid in $(ls outputs/freshrss/rerun/); do
|
||||
python3 -c "
|
||||
import json
|
||||
d = json.load(open('outputs/freshrss/rerun/$rid/summary/summary-batch.json'))
|
||||
for item in d['items']:
|
||||
for kw in item['summary'].get('keywords', []):
|
||||
print(kw)
|
||||
"
|
||||
done | sort -u
|
||||
```
|
||||
3. 按类别分组:正常词 / 产品名(stopword) / 模型名(stopword) / 需要 alias 的
|
||||
4. 分别更新 `term_aliases.json` 和 `term_stopwords.json`
|
||||
5. **重建 term_index**:对每个受影响日期的 run 跑 `build_keyword_index.py`
|
||||
6. **更新 Hugo tags**:从重建后的 `data/term_index/daily/YYYY-MM-DD.json` 取
|
||||
|
||||
## 当前配置 (2026-07-16)
|
||||
|
||||
| 文件 | 条目数 |
|
||||
|:----|:------|
|
||||
| `configs/term_aliases.json` | 142 |
|
||||
| `configs/term_stopwords.json` | 106 |
|
||||
| `configs/filter_context.personal.json` | 54 (interest_keywords) |
|
||||
| `configs/term_watchlist.json` | 6 |
|
||||
|
||||
## 关键词/别名/停用词配置的更新规范
|
||||
|
||||
- `term_aliases.json` — 只放"源词→目标词"的映射,目标是让不同写法的同一概念归一到标准形式
|
||||
- `term_stopwords.json` — 放产品名、模型名、人名、地名等对 Hugo tag 无价值的词
|
||||
- **随时可以加**,加了后重建 term_index 即可生效
|
||||
- 不涉及 pipeline 重新跑——只影响下游展示
|
||||
Review bundle、Suggestions、Schema 和产物保留策略以 `keyword-cleanup-review` 为唯一事实来源;该 Skill 不直接 Apply 配置。
|
||||
|
||||
@@ -1,60 +0,0 @@
|
||||
# Memory Drift Recovery
|
||||
|
||||
## Symptom
|
||||
|
||||
`store_memory(action="add", ...)` fails with:
|
||||
> Refusing to write MEMORY.md: file on disk has content that wouldn't round-trip through the memory tool...
|
||||
|
||||
A `.bak` snapshot is created: `/root/.hermes/memories/MEMORY.md.bak.<timestamp>`
|
||||
|
||||
## Root Cause
|
||||
|
||||
The MEMORY.md file format doesn't match what the memory tool expects — likely because the file was modified externally (by `patch`, `write_file`, shell `>>` append, or a concurrent session). The tool uses a `§` (section sign) delimited format internally and its serialization/deserialization doesn't match the on-disk content.
|
||||
|
||||
## Recovery Procedure
|
||||
|
||||
### Step 1: Read the backup and the current file
|
||||
|
||||
```bash
|
||||
diff /root/.hermes/memories/MEMORY.md.bak.<timestamp> /root/.hermes/memories/MEMORY.md
|
||||
```
|
||||
|
||||
### Step 2: Extract missing entries (if any)
|
||||
|
||||
```bash
|
||||
# List entries from the backup
|
||||
grep '^§' /root/.hermes/memories/MEMORY.md.bak.<timestamp>
|
||||
```
|
||||
|
||||
### Step 3: Re-add each missing entry via store_memory
|
||||
|
||||
For each entry that was in the backup but is now gone from the current file:
|
||||
```bash
|
||||
store_memory(action="add", content="<entry text>", target="memory")
|
||||
```
|
||||
|
||||
### Step 4: Reset to a clean state
|
||||
|
||||
If the file is completely corrupted, the cleanest path is:
|
||||
1. Save any new entries from the backup you want to keep
|
||||
2. Rewrite the file as a clean `§`-delimited list (one entry per `§` line)
|
||||
3. The format is: `§<content>\n` per entry, with `---` or blank line separators
|
||||
|
||||
```bash
|
||||
# Example clean format:
|
||||
echo '§当前重要条目一
|
||||
§当前重要条目二
|
||||
§当前重要条目三' > /root/.hermes/memories/MEMORY.md
|
||||
```
|
||||
|
||||
### Prevention
|
||||
|
||||
- Do NOT use `write_file` or `patch` to modify MEMORY.md directly — always use `store_memory()`
|
||||
- Do NOT use shell `>>` to append to MEMORY.md
|
||||
- If you must bulk-import, use `store_memory` per-entry, not file-level operations
|
||||
|
||||
## Environment
|
||||
|
||||
- Host: Linux (5.15)
|
||||
- Hermes home: `/root/.hermes`
|
||||
- Memory files: `~/.hermes/memories/MEMORY.md`, `~/.hermes/memories/USER.md`
|
||||
@@ -1,41 +1,33 @@
|
||||
+++
|
||||
title = "AI 日报 · 示例"
|
||||
date = 2026-04-01T16:55:00+08:00
|
||||
summary = "围绕 Agent 架构分层、Skills 标准化与桌面 Agent 工程实践的当日观察。"
|
||||
date = 2026-04-01T09:00:00+08:00
|
||||
summary = "围绕 Agent 架构分层、Skills 标准化与工程实践的当日观察。"
|
||||
+++
|
||||
|
||||
> ⚠️ 格式规范(生成 Hugo 时必须遵守):
|
||||
> - 文章编号:`1.` `2.` `3.`(阿拉伯数字 + 点),禁止 `① ② ③` / `一、二、三` 等变体
|
||||
> - 四个 section 缺一不可:`今日概览` → `今日重点` → `趋势观察` → `延伸阅读`
|
||||
> - 每篇文章结构:标题 → 摘要段 → "值得关注:"三点 → "这篇更值得关注的理由"段
|
||||
|
||||
# 今日概览
|
||||
|
||||
今天的公开候选主要集中在 AI Agent 的架构演进、工具化落地与工程化实践三条线索上。相比早期偏概念展示的讨论,这一批内容更强调模块化能力栈、真实部署路径与系统可维护性,说明行业关注点正在从“模型能做什么”转向“系统如何稳定落地并持续复用”。
|
||||
今天的公开内容主要集中在 AI Agent 架构演进、工具化落地与工程实践三条线索。行业关注点正在从“模型能做什么”转向“系统如何稳定落地并持续复用”。
|
||||
|
||||
## 今日重点
|
||||
|
||||
### 1. 学习笔记:从 Agent 到 Skills — AI 智能体架构的范式转变
|
||||
### 1. 从 Agent 到 Skills:AI 智能体架构的范式转变
|
||||
|
||||
文章分析了 AI 智能体架构从单体 Agent 向模块化 Skills 的范式转变。Anthropic 先后推出 MCP 和 Agent Skills 开放标准,构建了知识、工具、协作和运行分层架构。文章通过一个自动化美化相册的真实项目,对比了 Claude Code 与 OpenClaw 两种实现方案,验证了新架构的可复用性与灵活性。
|
||||
文章分析了 AI 智能体从单体 Agent 向模块化 Skills 的演进,并结合 MCP、Skills 和真实项目说明能力分层与复用方式。
|
||||
|
||||
值得关注:
|
||||
- Anthropic 在 14 个月内先后推出 MCP 和 Agent Skills 两个开放标准,推动 AI 智能体架构分层化。
|
||||
- 新范式核心是构建薄 Agent 引擎与可组合的 Skills 库,取代为每个用例定制单体 Agent。
|
||||
- 文章通过自动化美化相册项目,实操演示了 Skills、MCP、OpenClaw 和 A2A 协议如何协同工作。
|
||||
|
||||
这篇内容更值得关注的原因在于,它不只是提出了“Agent 要模块化”这个判断,而是把开放标准、分层架构和真实项目案例串成了一条完整论证链,能直接支撑今天日报的主线。
|
||||
- Skills 将领域流程从 Agent 主体中拆出,便于复用和维护。
|
||||
- MCP 为 Agent 与外部工具提供标准化连接方式。
|
||||
- 工程竞争点逐渐从模型调用转向状态、工具和工作流设计。
|
||||
|
||||
这篇内容值得关注的原因在于,它把开放协议、分层架构和真实落地案例连接成了完整论证链。
|
||||
|
||||
## 趋势观察
|
||||
|
||||
1. Agent 正在从单体能力转向可组合的模块化体系。无论是 Skills、MCP、记忆还是运行时编排,这批内容都在强调解耦与复用,而不是把智能体继续当成一个不可拆分的黑箱。
|
||||
2. 工程化正在变成 AI 应用竞争的主战场。桌面 Agent、企业级架构和部署实践类内容增多,说明真正的差异化开始落在接入现有流程、控制风险和提升可维护性上。
|
||||
3. AI 能力的竞争点正在上移。模型本身仍重要,但真正可持续的优势越来越来自系统设计、工作流整合和对业务场景的理解。
|
||||
1. Agent 正在从单体能力转向可组合的模块化体系。
|
||||
2. 工具契约、状态管理和验证机制正在成为 AI 应用的核心工程能力。
|
||||
3. Human-in-the-loop 仍是控制高风险副作用的重要边界。
|
||||
|
||||
## 延伸阅读
|
||||
|
||||
- [学习笔记:从 Agent 到 Skills — AI 智能体架构的范式转变](https://example.com/a)|阿里云开发者
|
||||
- [Agent Skills:打通可复用专业领域知识的最后一公里](https://example.com/b)|阿里云开发者
|
||||
- [CoPaw深度解析:源码架构和功能实践](https://example.com/c)|阿里云开发者
|
||||
|
||||
> ⚠️ 延伸阅读必须包含当天所有入选文章的原文链接(`- [标题](url)|来源`),一条对应一篇今日重点文章。不要放未入选或 pipeline drop 的文章。这是日报读者获取原文的入口,不是"其他相关阅读"区。
|
||||
- [从 Agent 到 Skills:AI 智能体架构的范式转变](https://example.com/a)|示例来源
|
||||
|
||||
@@ -1,64 +0,0 @@
|
||||
# Tag 去重合并计划 (2026-07-22)
|
||||
|
||||
## 问题
|
||||
https://osiman.site/tags/ 页面存在大量重叠 tag,如 `Harness工程化` / `Harness工程` / `Harness Engineering` 三个 tag 指向同一概念。
|
||||
|
||||
## 方案:数据层合并(Option A)
|
||||
|
||||
不修改展示层,直接合并所有日报 md 文件中的 tags。Hugo 自动重建 tag 页面,旧 tag 自动废弃。
|
||||
|
||||
## 规范映射表
|
||||
|
||||
| 废弃 tag | → | 规范名 | 说明 |
|
||||
|:---------|:---|:-------|:----|
|
||||
| `Harness工程化`, `Harness Engineering` | → | `Harness工程` | 三合一 |
|
||||
| `Skills` | → | `Skill` | 单复数 |
|
||||
| `AI Coding`, `AI代码生成`, `AI工程化` | → | `AI Coding Agent` | 统一为 Agent 维度 |
|
||||
| `Agent 框架`, `Agentic架构`, `Agent工程` | → | `Agent工程` | 三合一 |
|
||||
| `Prompt` | → | `Prompt Engineering` | 从简写改全称 |
|
||||
| `多Agent`, `Multi-Agent架构`, `多Agent协作` | → | `多Agent` | 三合一 |
|
||||
| `LLM`, `LLM评估`, `LLM训练`, `大模型应用开发` | → | `LLM` | 四合一 |
|
||||
| `循环工程`, `Loop Engineering`, `Agent Loop` | → | `循环工程` | 三合一 |
|
||||
| `上下文管理`, `Context工程` | → | `上下文管理` | 统一中文 |
|
||||
| `Code Review`, `代码质量` | → | `Code Review` | 统一英文 |
|
||||
| `推理加速`, `长文本推理`, `多步推理` | → | `推理加速` | 三合一 |
|
||||
| `安全`, `安全防御` | → | `安全` | 二合一 |
|
||||
| `技能系统`, `知识工程`, `知识管理` | → | `知识管理` | 三合一 |
|
||||
| `Agent`, `AI` | → | (删除) | 太泛,无信息量 |
|
||||
| `工程化` | → | (删除) | 冗余 |
|
||||
| `架构` | → | (删除) | 冗余 |
|
||||
|
||||
## 保留的独立 tag
|
||||
|
||||
`MCP`, `RAG`, `ACP`, `KV Cache`, `MoE架构`, `State Lake`, `RL决策训练`, `Token管理`, `Vibe Coding`, `Spec工程`, `Skill流水线`, `Hook 链`, `全双工语音交互`, `多端架构`, `契约化架构`, `数据`, `工具链`, `工作流`, `评测`, `部署`, `搜索`, `AI搜索`, `AI原生研发`, `AI Coding Agent`, `企业落地`, `记忆`, `开源`
|
||||
|
||||
## 执行方式
|
||||
|
||||
用 Python 脚本扫描 `content/daily/` 下所有 md 文件,对每个文件的 `tags = [...]` 替换为规范名版本。脚本参考:
|
||||
|
||||
```python
|
||||
import re, os, json
|
||||
|
||||
MAPPING = {
|
||||
"Harness工程化": "Harness工程",
|
||||
"Harness Engineering": "Harness工程",
|
||||
"Skills": "Skill",
|
||||
# ... 完整映射
|
||||
}
|
||||
|
||||
REMOVE = {"Agent", "AI", "工程化", "架构", ...}
|
||||
|
||||
hugo_dir = "/home/ubuntu/zhu/apps/hugo-site/content/daily"
|
||||
for root, _, files in os.walk(hugo_dir):
|
||||
for fname in files:
|
||||
if not fname.endswith(".md"):
|
||||
continue
|
||||
path = os.path.join(root, fname)
|
||||
with open(path, 'r') as f:
|
||||
content = f.read()
|
||||
# parse tags from frontmatter
|
||||
# replace deprecated → canonical
|
||||
# remove items in REMOVE
|
||||
# deduplicate
|
||||
# write back
|
||||
```
|
||||
Reference in New Issue
Block a user