first commit
This commit is contained in:
@@ -0,0 +1,241 @@
|
||||
---
|
||||
name: ima-skill
|
||||
description: |
|
||||
统一的 IMA OpenAPI 技能,支持笔记管理和知识库操作。
|
||||
当用户提到知识库、资料库、笔记、备忘录、记事,或者想要上传文件、添加网页到知识库、
|
||||
搜索知识库内容、搜索/浏览/创建/编辑笔记时,使用此 skill。
|
||||
即使用户没有明确说"知识库"或"笔记",只要意图涉及文件上传到知识库、网页收藏、
|
||||
知识搜索、个人文档存取(如"帮我记一下"、"搜一下知识库里有没有XX"),也应触发此 skill。
|
||||
homepage: https://ima.qq.com
|
||||
metadata:
|
||||
openclaw:
|
||||
emoji: '🔧'
|
||||
requires: { env: ['IMA_OPENAPI_CLIENTID', 'IMA_OPENAPI_APIKEY'] }
|
||||
primaryEnv: 'IMA_OPENAPI_CLIENTID'
|
||||
security:
|
||||
credentials_usage: |
|
||||
This skill requires user-provisioned IMA OpenAPI credentials (Client ID and API Key)
|
||||
to authenticate with the official IMA API at https://ima.qq.com.
|
||||
Credentials are ONLY sent to the official IMA API endpoint (ima.qq.com) as HTTP headers.
|
||||
No credentials are logged, stored in files, or transmitted to any other destination.
|
||||
allowed_domains:
|
||||
- ima.qq.com
|
||||
---
|
||||
|
||||
# ima-skill
|
||||
|
||||
Unified IMA OpenAPI skill. Currently supports: **notes**, **knowledge-base**.
|
||||
|
||||
## Setup
|
||||
|
||||
> **Security note:** This skill authenticates with the **official IMA API** (`ima.qq.com`) — the same service the user already uses. Credentials are only sent as HTTP headers to `ima.qq.com` and never to any other domain, file, or log.
|
||||
|
||||
1. 打开 https://ima.qq.com/agent-interface 获取 **Client ID** 和 **API Key**
|
||||
2. 存储凭证(三选一):
|
||||
|
||||
**方式 A — OpenClaw 配置(推荐,能让 skill 直接变为 ready):**
|
||||
|
||||
在 `~/.openclaw/openclaw.json` 中加入:
|
||||
|
||||
```json
|
||||
{
|
||||
"skills": {
|
||||
"entries": {
|
||||
"ima-skill": {
|
||||
"enabled": true,
|
||||
"env": {
|
||||
"IMA_OPENAPI_CLIENTID": "your_client_id",
|
||||
"IMA_OPENAPI_APIKEY": "your_api_key"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**方式 B — 配置文件:**
|
||||
|
||||
```bash
|
||||
mkdir -p ~/.config/ima
|
||||
echo "your_client_id" > ~/.config/ima/client_id
|
||||
echo "your_api_key" > ~/.config/ima/api_key
|
||||
```
|
||||
|
||||
**方式 C — 环境变量:**
|
||||
|
||||
```bash
|
||||
export IMA_OPENAPI_CLIENTID="your_client_id"
|
||||
export IMA_OPENAPI_APIKEY="your_api_key"
|
||||
```
|
||||
|
||||
> 说明:如果你希望 `openclaw skills info ima-skill` 显示为 ready,优先使用 **方式 A** 或 **方式 C**;仅使用配置文件时,skill 运行时能读取,但 OpenClaw 的就绪检查未必会把它判定为已满足环境变量要求。
|
||||
|
||||
Agent 会按优先级依次尝试:环境变量 → 配置文件。
|
||||
|
||||
## 凭证预检
|
||||
|
||||
每次调用 API 前,先确认凭证可用。如果两个值都为空,停止操作并提示用户按 Setup 步骤配置。
|
||||
|
||||
```bash
|
||||
# Load user-provisioned IMA credentials (used ONLY for ima.qq.com API authentication)
|
||||
IMA_CLIENT_ID="${IMA_OPENAPI_CLIENTID:-$(cat ~/.config/ima/client_id 2>/dev/null)}"
|
||||
IMA_API_KEY="${IMA_OPENAPI_APIKEY:-$(cat ~/.config/ima/api_key 2>/dev/null)}"
|
||||
if [ -z "$IMA_CLIENT_ID" ] || [ -z "$IMA_API_KEY" ]; then
|
||||
echo "缺少 IMA 凭证,请按 Setup 步骤配置 Client ID 和 API Key"
|
||||
exit 1
|
||||
fi
|
||||
```
|
||||
|
||||
## API 调用模板
|
||||
|
||||
所有请求统一为 **HTTP POST + JSON Body**,仅发往官方 Base URL `https://ima.qq.com`。
|
||||
|
||||
定义辅助函数避免重复 header — 每个模块传入完整路径:
|
||||
|
||||
```bash
|
||||
# All requests go ONLY to the official IMA API (ima.qq.com)
|
||||
ima_api() {
|
||||
local path="$1" body="$2"
|
||||
curl -s -X POST "https://ima.qq.com/$path" \
|
||||
-H "ima-openapi-clientid: $IMA_CLIENT_ID" \
|
||||
-H "ima-openapi-apikey: $IMA_API_KEY" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d "$body"
|
||||
}
|
||||
```
|
||||
|
||||
> **Note:** All IMA OpenAPI endpoints currently use HTTP POST. If a future module requires a different method, `ima_api()` must be extended to accept a method parameter.
|
||||
|
||||
## 模块决策表
|
||||
|
||||
| 用户意图 | 模块 | 读取 |
|
||||
| ------------------------------------------------------------------------------------------ | -------------- | ------------------------- |
|
||||
| 搜索笔记、浏览笔记本、获取笔记内容、创建笔记、追加内容 | notes | `references/notes.md` |
|
||||
| 上传文件、添加网页链接、搜索知识库、浏览知识库内容、获取知识库信息、获取可添加的知识库列表 | knowledge-base | `references/knowledge-base.md` |
|
||||
|
||||
### ⚠️ 易混淆场景
|
||||
|
||||
以下场景容易误判模块,需特别注意:
|
||||
|
||||
| 用户说的 | 实际意图 | 正确路由 |
|
||||
| -------------------------------------------------------- | -------------------------- | -------------------------------------------------------------------- |
|
||||
| "把这段内容添加到知识库XX里的笔记YY" | 往已有**笔记**追加内容 | **notes** — 先搜索笔记获取 `doc_id`,再用 `append_doc` |
|
||||
| "把这个写到XX笔记里"、"记到XX笔记" | 往已有**笔记**追加内容 | **notes** — `append_doc` |
|
||||
| "把这篇笔记添加到知识库" | 将笔记关联到**知识库** | **knowledge-base** — `add_knowledge` with `media_type=11` |
|
||||
| "上传文件到知识库" | 上传**文件**到知识库 | **knowledge-base** — `create_media` → COS → `add_knowledge` |
|
||||
| "新建一篇笔记记录这些内容" | **创建**新笔记 | **notes** — `import_doc` |
|
||||
| "帮我记一下"、"记录一下"、"保存为笔记"(未指定已有笔记) | 意图不明确,**需要确认** | **notes** — 先询问用户是创建新笔记还是追加到哪篇已有笔记,再决定接口 |
|
||||
| "添加到笔记里"(未指定具体哪篇) | 意图不明确,**需要确认** | **notes** — 先询问用户是创建新笔记还是追加到哪篇已有笔记,再决定接口 |
|
||||
| "把知识库里的XX内容记到笔记" | 先从知识库读取,再写入笔记 | **多模块** — knowledge-base 搜索/读取 → notes 创建/追加 |
|
||||
|
||||
**核心判断规则**:
|
||||
|
||||
- 目标是**笔记的内容**(读、写、追加)→ notes 模块
|
||||
- 目标是**知识库的条目**(上传文件、添加链接、关联笔记到知识库)→ knowledge-base 模块
|
||||
- 用户提到"知识库"只是在**描述笔记的位置**(如"知识库里的那篇笔记"),真正操作对象仍是笔记 → notes 模块
|
||||
|
||||
> **多模块任务**:当用户意图涉及多个模块时(如"从知识库搜索内容并记到笔记"),按意图顺序依次读取对应的模块文档并逐步执行。先完成前一个模块的操作,再进入下一个模块。
|
||||
|
||||
## 注意事项
|
||||
|
||||
- **UTF-8 编码(仅 notes 模块)**:见下方「⚠️ UTF-8 编码强制要求」章节。notes 模块的所有写入操作前**必须**完成 UTF-8 编码校验,否则会导致内容乱码且无法修复。
|
||||
- **文件上传保持原样(knowledge-base 模块)**:当用户要求上传文件到知识库时,**必须保持文件原始内容不变**,不得进行任何编码转换。文件以二进制方式上传,服务端会自行处理编码。擅自转码可能破坏文件内容(如 PDF、图片、Excel 等非文本文件,或用户有意使用特定编码的文本文件)。
|
||||
- **PowerShell 5.1 环境(所有模块)**:见下方「⚠️ PowerShell 5.1 环境检测」章节。此问题影响**所有** API 调用(notes、knowledge-base 等),PowerShell 5.1 会静默将请求 Body 转为 GBK 编码导致乱码。
|
||||
|
||||
## ⚠️ UTF-8 编码强制要求(CRITICAL — 仅适用于 notes 模块)
|
||||
|
||||
> **此规则为强制性要求,不可跳过。** 非法编码会导致内容在 IMA 中显示为乱码,且无法修复,必须重新写入。
|
||||
>
|
||||
> **适用范围:notes 模块**(`import_doc`、`append_doc` 等文本写入 API)。
|
||||
>
|
||||
> **不适用于 knowledge-base 模块的文件上传**:上传文件时必须保持文件原始内容,不得转码。文件以二进制方式上传,服务端自行处理。
|
||||
|
||||
**每次调用 notes 写入类 API(`import_doc`/`append_doc`)之前,必须对 `content`、`title` 等所有字符串字段执行 UTF-8 编码校验/转换。** 无论内容来源如何——用户直接输入、从文件读取、WebFetch 抓取、剪贴板粘贴、外部 API 返回——都不能假设已经是合法 UTF-8,必须显式确认。
|
||||
|
||||
### 强制检查清单(notes 模块写入前)
|
||||
|
||||
在构造 notes 写入请求的 body **之前**,完成以下步骤:
|
||||
|
||||
1. **来自文件的内容**:先检测文件编码,转为 UTF-8 后再读入变量(注意:这是指读取文件内容作为笔记正文写入,不是上传文件到知识库)
|
||||
2. **来自 WebFetch / HTTP 请求的内容**:响应可能为 GBK/Latin-1 等,必须转码
|
||||
3. **来自用户输入或变量拼接的内容**:清洗非法 UTF-8 字节(`\xff\xfe` 等)
|
||||
4. **标题字段同理**:`title` 也必须为合法 UTF-8
|
||||
|
||||
### 各环境转码方法
|
||||
|
||||
**Python(推荐,几乎所有环境都有):**
|
||||
|
||||
```bash
|
||||
# 读取文件,自动检测编码并转为 UTF-8
|
||||
content=$(python3 -c "
|
||||
import sys
|
||||
data = open('tmpfile', 'rb').read()
|
||||
for enc in ['utf-8', 'gbk', 'gb2312', 'big5', 'latin-1']:
|
||||
try:
|
||||
sys.stdout.write(data.decode(enc))
|
||||
break
|
||||
except (UnicodeDecodeError, LookupError):
|
||||
continue
|
||||
" 2>/dev/null)
|
||||
|
||||
# 如果内容已在变量中,清洗非法 UTF-8 字节
|
||||
content=$(printf '%s' "$content" | python3 -c "import sys; sys.stdout.write(sys.stdin.buffer.read().decode('utf-8','ignore'))")
|
||||
```
|
||||
|
||||
**Node.js:**
|
||||
|
||||
```bash
|
||||
content=$(node -e "const fs=require('fs');const buf=fs.readFileSync('tmpfile');process.stdout.write(buf.toString('utf8'))")
|
||||
# 已知编码(如 GBK):
|
||||
content=$(node -e "const fs=require('fs');process.stdout.write(new TextDecoder('gbk').decode(fs.readFileSync('tmpfile')))")
|
||||
```
|
||||
|
||||
**Unix (macOS/Linux):**
|
||||
|
||||
```bash
|
||||
content=$(iconv -f "$(file -b --mime-encoding tmpfile)" -t UTF-8 tmpfile 2>/dev/null || cat tmpfile)
|
||||
```
|
||||
|
||||
**Windows PowerShell:**
|
||||
|
||||
```powershell
|
||||
# 读取非 UTF-8 文件并转码
|
||||
$content = [System.IO.File]::ReadAllText('tmpfile', [System.Text.Encoding]::Default)
|
||||
[System.IO.File]::WriteAllText('tmpfile.utf8', $content, [System.Text.Encoding]::UTF8)
|
||||
```
|
||||
|
||||
### ⚠️ PowerShell 5.1 环境检测(CRITICAL — 适用于所有模块)
|
||||
|
||||
> **此问题影响所有 API 调用(notes、knowledge-base 等)**
|
||||
>
|
||||
> **此问题极其隐蔽:PowerShell 5.1 下 `Invoke-RestMethod` 会静默将请求 Body 从 UTF-8 转为系统 ANSI 编码(中文 Windows 为 GBK),即使设置了 `Content-Type: charset=utf-8` 也无效。结果是请求看起来发送成功,但服务端收到的内容已经是乱码,且无任何错误提示。**
|
||||
|
||||
**当 agent 运行在 PowerShell 环境时,必须在首次 API 调用前检测版本:**
|
||||
|
||||
```powershell
|
||||
# 检测 PowerShell 版本 — 在任何 API 调用之前执行(notes 和 knowledge-base 都需要)
|
||||
if ($PSVersionTable.PSVersion.Major -le 5) {
|
||||
Write-Host "⚠️ 检测到 PowerShell 5.1,将使用 UTF-8 字节数组模式发送请求"
|
||||
$useUtf8Bytes = $true
|
||||
} else {
|
||||
Write-Host "✅ PowerShell 7+,默认 UTF-8,无需额外处理"
|
||||
$useUtf8Bytes = $false
|
||||
}
|
||||
```
|
||||
|
||||
**PowerShell 5.1 下必须使用以下方式发送请求**(用 `ConvertTo-Json` 构建 JSON 以避免手动拼接的转义风险,再显式转为 UTF-8 字节数组):
|
||||
|
||||
```powershell
|
||||
# PowerShell 5.1 安全请求模板(适用于所有模块的所有 API 调用)
|
||||
$body = @{ title = "标题"; content = $content; content_format = 1 } | ConvertTo-Json -Depth 10
|
||||
if ($useUtf8Bytes) {
|
||||
# CRITICAL: 必须转为字节数组,否则中文/非ASCII内容会变成乱码
|
||||
$utf8Bytes = [System.Text.Encoding]::UTF8.GetBytes($body)
|
||||
Invoke-RestMethod -Uri $url -Method Post -Body $utf8Bytes -ContentType "application/json; charset=utf-8" -Headers $headers
|
||||
} else {
|
||||
# PowerShell 7+ 可直接传字符串
|
||||
Invoke-RestMethod -Uri $url -Method Post -Body $body -ContentType "application/json; charset=utf-8" -Headers $headers
|
||||
}
|
||||
```
|
||||
|
||||
> **总结:** 在 PowerShell 5.1 环境中,**所有** API 调用(无论 notes 还是 knowledge-base)都必须将 Body 显式转为 UTF-8 字节数组。不检测版本直接发请求 = 中文内容必乱码。这是 PowerShell 5.1 的已知设计缺陷,不是 bug 可以被修复。
|
||||
@@ -0,0 +1,446 @@
|
||||
# IMA知识库 API
|
||||
|
||||
## ⚠️ 必读约束
|
||||
|
||||
### 🌐 服务信息
|
||||
|
||||
- **Base URL **:`https://ima.qq.com`
|
||||
- **Base Path**:`/openapi/wiki/v1`
|
||||
- **协议**:HTTP POST,JSON body
|
||||
- **完整示例**:`POST https://ima.qq.com/openapi/wiki/v1/get_knowledge_base`
|
||||
|
||||
### 🔒 认证
|
||||
|
||||
所有请求必须携带 Header:
|
||||
|
||||
| Header | 说明 |
|
||||
| ---------------------- | ------------------ |
|
||||
| `ima-openapi-clientid` | Client ID |
|
||||
| `ima-openapi-apikey` | API Key |
|
||||
| `Content-Type` | `application/json` |
|
||||
|
||||
---
|
||||
|
||||
## 快速决策
|
||||
|
||||
| 用户意图 | 接口 |
|
||||
| ------------------------------- | ---------------------------------------------------------------------- |
|
||||
| 「上传文件到知识库」 | `check_repeated_names` → `create_media` → COS Upload → `add_knowledge` |
|
||||
| 「上传文件到指定文件夹」 | 先定位文件夹 → 同上(传入 `folder_id`) |
|
||||
| 「添加网页/微信文章到知识库」 | `import_urls` |
|
||||
| 「获取知识库信息」 | `get_knowledge_base` |
|
||||
| 「浏览知识库内容 / 浏览文件夹」 | `get_knowledge_list`(可传 `folder_id` 进入子文件夹) |
|
||||
| 「在知识库中搜索」 | `search_knowledge` |
|
||||
| 「搜索知识库列表」 | `search_knowledge_base` |
|
||||
| 「获取可添加的知识库列表」 | `get_addable_knowledge_base_list` |
|
||||
| 「检查文件名是否重复」 | `check_repeated_names` |
|
||||
|
||||
---
|
||||
|
||||
## 数据结构
|
||||
|
||||
### KnowledgeBaseInfo(知识库信息)
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ----------------------- | -------- | ------------- |
|
||||
| `id` | string | 知识库唯一 ID |
|
||||
| `name` | string | 知识库名称 |
|
||||
| `cover_url` | string | 封面图 URL |
|
||||
| `description` | string | 描述 |
|
||||
| `recommended_questions` | string[] | 推荐问题列表 |
|
||||
|
||||
### KnowledgeInfo(知识条目)
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ------------------ | ------ | ------------- |
|
||||
| `media_id` | string | 媒体 ID |
|
||||
| `title` | string | 标题 |
|
||||
| `parent_folder_id` | string | 所属文件夹 ID |
|
||||
|
||||
### FolderInfo(文件夹条目)
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ------------------ | ------ | ----------- |
|
||||
| `folder_id` | string | 文件夹 ID |
|
||||
| `name` | string | 文件夹名称 |
|
||||
| `file_number` | int64 | 文件数 |
|
||||
| `folder_number` | int64 | 子文件夹数 |
|
||||
| `parent_folder_id` | string | 父文件夹 ID |
|
||||
| `is_top` | bool | 是否置顶 |
|
||||
|
||||
### AddableKnowledgeBaseInfo(可添加的知识库信息)
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ------ | ------ | ---------- |
|
||||
| `id` | string | 知识库 ID |
|
||||
| `name` | string | 知识库名称 |
|
||||
|
||||
### SearchedKnowledgeBaseInfo(搜索到的知识库信息)
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ----------- | ------ | ---------- |
|
||||
| `id` | string | 知识库 ID |
|
||||
| `name` | string | 知识库名称 |
|
||||
| `cover_url` | string | 封面图 URL |
|
||||
|
||||
### SearchedKnowledgeInfo(搜索到的知识条目)
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ------------------- | ------ | -------------------------- |
|
||||
| `media_id` | string | 媒体 ID |
|
||||
| `title` | string | 标题 |
|
||||
| `parent_folder_id` | string | 所属文件夹 ID |
|
||||
| `highlight_content` | string | 高亮内容(内容匹配时返回) |
|
||||
|
||||
### ContentInfo(内容信息)
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ------------ | ------ | ----------------------- |
|
||||
| `content_id` | string | 内容 ID(网页时为 URL) |
|
||||
|
||||
### ImportURLData(URL 导入结果)
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ---------- | ------ | ----------------------- |
|
||||
| `url` | string | 导入的 URL |
|
||||
| `ret_code` | int32 | 0=成功,非 0=失败 |
|
||||
| `media_id` | string | 导入成功后返回的媒体 ID |
|
||||
|
||||
### FileInfo(文件信息)
|
||||
|
||||
`add_knowledge` 文件上传时使用:
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ------------------ | ------ | -------------------------- |
|
||||
| `cos_key` | string | COS 对象 Key |
|
||||
| `file_size` | uint64 | 文件大小(字节) |
|
||||
| `last_modify_time` | int64 | 最后修改时间(秒级时间戳) |
|
||||
| `password` | string | 文件密码(如有) |
|
||||
| `file_name` | string | 文件名称 |
|
||||
|
||||
### Credential(COS 上传凭证)
|
||||
|
||||
`create_media` 返回,用于上传文件到腾讯云 COS:
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| --------------- | ------ | -------------------------- |
|
||||
| `token` | string | 临时 TOKEN |
|
||||
| `secret_id` | string | 临时 Secret ID |
|
||||
| `secret_key` | string | 临时 Secret Key |
|
||||
| `start_time` | int64 | 凭证开始时间(秒级时间戳) |
|
||||
| `expired_time` | int64 | 凭证过期时间(秒级时间戳) |
|
||||
| `appid` | string | COS AppID |
|
||||
| `bucket_name` | string | COS 桶名称 |
|
||||
| `region` | string | COS 桶所在区域 |
|
||||
| `custom_domain` | string | 自定义域名 |
|
||||
| `cos_key` | string | COS 对象 Key |
|
||||
|
||||
### MediaType(媒体类型枚举)
|
||||
|
||||
| 值 | 名称 | content_type / 说明 |
|
||||
| --- | -------------- | ------------------------------------------------------------------------------------------------------------- |
|
||||
| 1 | PDF | `application/pdf` |
|
||||
| 2 | 网页 | N/A(直接 AddKnowledge,`web_info.content_id=<url>`) |
|
||||
| 3 | Word | `application/msword` / `application/vnd.openxmlformats-officedocument.wordprocessingml.document` |
|
||||
| 4 | PPT | `application/vnd.ms-powerpoint` / `application/vnd.openxmlformats-officedocument.presentationml.presentation` |
|
||||
| 5 | Excel | `application/vnd.ms-excel` / `application/vnd.openxmlformats-officedocument.spreadsheetml.sheet` / `text/csv` |
|
||||
| 6 | 微信公众号文章 | N/A(直接 AddKnowledge,`web_info.content_id=<url>`,URL 匹配 `mp.weixin.qq.com/s`) |
|
||||
| 7 | MarkDown | `text/markdown` / `text/x-markdown` / `application/md` / `application/markdown` |
|
||||
| 9 | 图片 | `image/png`, `image/jpeg`, `image/webp` |
|
||||
| 11 | 笔记 | N/A(直接 AddKnowledge,`note_info.content_id=<doc_id>`) |
|
||||
| 12 | AI会话 | N/A(直接 AddKnowledge,`session_info.content_id=<session_id>`) |
|
||||
| 13 | TXT | `text/plain` |
|
||||
| 14 | Xmind | `application/x-xmind` / `application/vnd.xmind.workbook` / `application/zip` |
|
||||
| 15 | 录音 | `audio/mpeg`(mp3), `audio/x-m4a`(m4a), `audio/wav`(wav), `audio/aac`(aac) |
|
||||
| 16 | 视频解析 | **不支持通过 skill 添加**。Bilibili/YouTube/本地HTML等仅支持在 ima 桌面端内添加进知识库 |
|
||||
|
||||
---
|
||||
|
||||
## 接口详情
|
||||
|
||||
### 1. 创建媒体
|
||||
|
||||
POST /openapi/wiki/v1/create_media
|
||||
|
||||
**触发场景**:上传文件到知识库的第一步,获取 COS 上传凭证。
|
||||
|
||||
#### 请求参数
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
| ------------------- | ------ | ---- | ------------------------------ |
|
||||
| `file_name` | string | 是 | 文件名称(最长 1024 字符) |
|
||||
| `file_size` | uint64 | 是 | 文件大小(字节) |
|
||||
| `content_type` | string | 是 | MIME 类型 |
|
||||
| `knowledge_base_id` | string | 是 | 知识库 ID |
|
||||
| `file_ext` | string | 是 | 文件后缀名(无点号,如 `pdf`) |
|
||||
|
||||
#### 返回字段
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ---------------- | ---------- | ------------ |
|
||||
| `media_id` | string | 媒体 ID |
|
||||
| `cos_credential` | Credential | COS 上传凭证 |
|
||||
|
||||
---
|
||||
|
||||
### 2. 添加知识
|
||||
|
||||
POST /openapi/wiki/v1/add_knowledge
|
||||
|
||||
**触发场景**:上传文件到知识库的最后一步,或直接添加网页 URL。
|
||||
|
||||
#### 请求参数
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
| --------------------- | ----------- | -------- | --------------------------------------- |
|
||||
| `media_type` | int32 | 是 | 媒体类型 |
|
||||
| `media_id` | string | 否 | 文件上传时必填,CreateMedia 返回的 ID |
|
||||
| `title` | string | 是 | 标题 |
|
||||
| `knowledge_base_id` | string | 是 | 知识库 ID |
|
||||
| `folder_id` | string | 否 | 文件夹 ID(省略则添加到根目录) |
|
||||
| `note_info` | ContentInfo | 否 | 笔记内容信息 |
|
||||
| `web_info` | ContentInfo | 否 | 网页内容信息(media_type=2 时必填) |
|
||||
| `web_info.content_id` | string | 条件必填 | 网页 URL(media_type=2 时必填) |
|
||||
| `session_info` | ContentInfo | 否 | 会话内容信息 |
|
||||
| `file_info` | FileInfo | 否 | 文件信息(文件上传时必填,见 FileInfo) |
|
||||
|
||||
#### 返回字段
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ---------- | ------ | ------- |
|
||||
| `media_id` | string | 媒体 ID |
|
||||
|
||||
---
|
||||
|
||||
### 3. 获取知识库信息
|
||||
|
||||
POST /openapi/wiki/v1/get_knowledge_base
|
||||
|
||||
#### 请求参数
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
| ----- | -------- | ---- | --------------------------------- |
|
||||
| `ids` | string[] | 是 | 知识库 ID 列表(1-20 个,不重复) |
|
||||
|
||||
#### 返回字段
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ------- | -------------------------------- | -------------- |
|
||||
| `infos` | map\<string, KnowledgeBaseInfo\> | 知识库信息映射 |
|
||||
|
||||
---
|
||||
|
||||
### 4. 浏览知识库内容
|
||||
|
||||
POST /openapi/wiki/v1/get_knowledge_list
|
||||
|
||||
#### 请求参数
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
| ------------------- | ------ | ---- | ----------------------------- |
|
||||
| `cursor` | string | 是 | 游标,首次传空字符串 |
|
||||
| `limit` | uint64 | 是 | 数量限制(1-50) |
|
||||
| `knowledge_base_id` | string | 是 | 知识库 ID |
|
||||
| `folder_id` | string | 否 | 文件夹 ID(省略则列出根目录) |
|
||||
|
||||
#### 返回字段
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ---------------- | --------------- | ---------------- |
|
||||
| `knowledge_list` | KnowledgeInfo[] | 知识条目列表 |
|
||||
| `is_end` | bool | 是否到达列表末尾 |
|
||||
| `next_cursor` | string | 下页游标 |
|
||||
| `current_path` | FolderInfo[] | 当前路径 |
|
||||
|
||||
---
|
||||
|
||||
### 5. 搜索知识库内容
|
||||
|
||||
POST /openapi/wiki/v1/search_knowledge
|
||||
|
||||
#### 请求参数
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
| ------------------- | ------ | ---- | -------------------- |
|
||||
| `query` | string | 是 | 搜索关键词 |
|
||||
| `cursor` | string | 是 | 游标,首次传空字符串 |
|
||||
| `knowledge_base_id` | string | 是 | 知识库 ID |
|
||||
|
||||
#### 返回字段
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ------------- | ----------------------- | ------------------------------------------------------------------------ |
|
||||
| `info_list` | SearchedKnowledgeInfo[] | 搜索结果(`media_id`, `title`, `parent_folder_id`, `highlight_content`) |
|
||||
| `is_end` | bool | 是否到达列表末尾 |
|
||||
| `next_cursor` | string | 下页游标 |
|
||||
|
||||
---
|
||||
|
||||
### 6. 搜索知识库列表
|
||||
|
||||
POST /openapi/wiki/v1/search_knowledge_base
|
||||
|
||||
#### 请求参数
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
| -------- | ------ | ---- | -------------------- |
|
||||
| `query` | string | 是 | 搜索关键词 |
|
||||
| `cursor` | string | 是 | 游标,首次传空字符串 |
|
||||
| `limit` | uint64 | 是 | 数量限制(1-50) |
|
||||
|
||||
#### 返回字段
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ------------- | --------------------------- | ------------------------------------- |
|
||||
| `info_list` | SearchedKnowledgeBaseInfo[] | 搜索结果(`id`, `name`, `cover_url`) |
|
||||
| `is_end` | bool | 是否到达列表末尾 |
|
||||
| `next_cursor` | string | 下页游标 |
|
||||
|
||||
---
|
||||
|
||||
### 7. 获取可添加的知识库列表
|
||||
|
||||
POST /openapi/wiki/v1/get_addable_knowledge_base_list
|
||||
|
||||
**触发场景**:用户想上传文件或添加内容到知识库,但不确定可以添加到哪些知识库时,列出当前用户有权限添加内容的知识库。
|
||||
|
||||
#### 请求参数
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
| -------- | ------ | ---- | -------------------- |
|
||||
| `cursor` | string | 是 | 游标,首次传空字符串 |
|
||||
| `limit` | uint64 | 是 | 数量限制(1-50) |
|
||||
|
||||
#### 返回字段
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ----------------------------- | -------------------------- | ---------------------- |
|
||||
| `addable_knowledge_base_list` | AddableKnowledgeBaseInfo[] | 可添加内容的知识库列表 |
|
||||
| `next_cursor` | string | 下页游标 |
|
||||
| `is_end` | bool | 是否到达列表末尾 |
|
||||
|
||||
---
|
||||
|
||||
### 8. 检查文件名重复
|
||||
|
||||
POST /openapi/wiki/v1/check_repeated_names
|
||||
|
||||
**触发场景**:上传文件到知识库前,检查目标知识库(及文件夹)中是否已存在同名文件。仅用于文件类型(media_type 1/3/4/5/7/9/13/14),不用于网页(2/6)、笔记(11)等。
|
||||
|
||||
#### 请求参数
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
| ------------------- | ------------------------- | ---- | ----------------------------- |
|
||||
| `params` | CheckRepeatedNamesParam[] | 是 | 待检查的文件列表(1-2000 个) |
|
||||
| `knowledge_base_id` | string | 是 | 知识库 ID |
|
||||
| `folder_id` | string | 否 | 文件夹 ID(省略则检查根目录) |
|
||||
|
||||
**CheckRepeatedNamesParam:**
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ------------ | ------ | ----------------------------- |
|
||||
| `name` | string | 文件名称 |
|
||||
| `media_type` | int32 | 媒体类型(见 MediaType 枚举) |
|
||||
|
||||
#### 返回字段
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| --------- | -------------------------- | -------- |
|
||||
| `results` | CheckRepeatedNamesResult[] | 检查结果 |
|
||||
|
||||
**CheckRepeatedNamesResult:**
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ------------- | ------ | ------------------------- |
|
||||
| `name` | string | 文件名称 |
|
||||
| `is_repeated` | bool | `true` 表示同名文件已存在 |
|
||||
|
||||
---
|
||||
|
||||
### 9. 导入 URL
|
||||
|
||||
POST /openapi/wiki/v1/import_urls
|
||||
|
||||
**触发场景**:添加网页或微信公众号文章到知识库。替代 `add_knowledge` 的 `media_type=2/6` 用法,支持批量导入,服务端自动识别 URL 类型。
|
||||
|
||||
#### 请求参数
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
| ------------------- | -------- | ---- | ----------------------------------- |
|
||||
| `knowledge_base_id` | string | 是 | 知识库 ID |
|
||||
| `folder_id` | string | 是 | 文件夹 ID |
|
||||
| `urls` | string[] | 是 | URL 列表(1-10 个,每个非空字符串) |
|
||||
|
||||
#### 返回字段
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| --------- | ---------------------------- | ----------------------------------------- |
|
||||
| `results` | map\<string, ImportURLData\> | URL→结果映射(含 `ret_code`、`media_id`) |
|
||||
|
||||
---
|
||||
|
||||
## 文件夹说明
|
||||
|
||||
知识库内容以文件夹层级结构组织。文件夹是一种特殊的知识条目:
|
||||
|
||||
- `get_knowledge_list` 返回结果中同时包含 **文件**(`KnowledgeInfo`)和 **文件夹**(`FolderInfo`),通过 `current_path` 字段可获取当前路径的面包屑信息
|
||||
- `search_knowledge` 搜索结果中也会包含匹配的文件夹
|
||||
- 所有支持 `folder_id` 参数的接口(`add_knowledge`、`import_urls`、`get_knowledge_list`、`check_repeated_names`),省略 `folder_id` 则操作根目录。**根目录的 folder_id 等于 knowledge_base_id**,当接口要求 `folder_id` 必填时(如 `import_urls`),传 `knowledge_base_id` 的值即可表示根目录
|
||||
- **定位文件夹**:当用户只提供文件夹名称时,使用 `search_knowledge` 按名称搜索,或用 `get_knowledge_list` 逐级浏览,从返回结果中找到目标文件夹的 ID
|
||||
|
||||
---
|
||||
|
||||
## 文件大小限制
|
||||
|
||||
上传前必须校验文件大小,超限文件应在上传前拦截:
|
||||
|
||||
| 文件类型 | media_type | 最大大小 |
|
||||
| --------------------------- | ----------- | -------- |
|
||||
| Excel、TXT、Xmind、Markdown | 5/13/14/7 | 10 MB |
|
||||
| 图片 | 9 | 30 MB |
|
||||
| PDF、Word、PPT、音频及其他 | 1/3/4/15 等 | 200 MB |
|
||||
|
||||
网页(2/6)、笔记(11)等非文件类型无大小限制。音频文件额外限制:最长 2 小时。
|
||||
|
||||
---
|
||||
|
||||
## 响应格式
|
||||
|
||||
所有 API 返回统一结构:
|
||||
|
||||
```json
|
||||
{
|
||||
"retcode": 0,
|
||||
"errmsg": "成功",
|
||||
"data": { ... }
|
||||
}
|
||||
```
|
||||
|
||||
- `retcode=0`:成功,从 `data` 提取业务字段
|
||||
- `retcode≠0`:失败,**直接将 `errmsg` 展示给用户**,无需自行翻译错误码
|
||||
|
||||
---
|
||||
|
||||
## 游标翻页使用规范
|
||||
|
||||
1. **首次请求**:`cursor` 传空字符串 `""`
|
||||
2. 检查返回的 `is_end`:`false` 表示还有更多数据
|
||||
3. 将返回的 `next_cursor` 作为下次请求的 `cursor`
|
||||
4. `is_end = true` 时停止翻页
|
||||
|
||||
---
|
||||
|
||||
## 错误码
|
||||
|
||||
| 错误码 | 说明 | 建议处理 |
|
||||
| ------ | ------------ | --------------------------- |
|
||||
| 0 | 成功 | — |
|
||||
| 110001 | 参数非法 | 检查请求参数(详见 errmsg) |
|
||||
| 110002 | 配置非法 | 检查服务配置 |
|
||||
| 110010 | 下游网络错误 | 可重试 |
|
||||
| 110011 | 下游逻辑错误 | 不可重试,详见 errmsg |
|
||||
| 110012 | 接口无效 | 检查接口路径 |
|
||||
| 110013 | 客户端取消 | 检查请求是否超时 |
|
||||
| 110020 | 安全打击 | 检查内容是否违规 |
|
||||
| 110021 | 请求频控 | 降低请求频率后重试 |
|
||||
| 110030 | 无权限 | 确认操作权限 |
|
||||
@@ -0,0 +1,521 @@
|
||||
# Knowledge Base (知识库)
|
||||
|
||||
> Prerequisites: see root `../SKILL.md` for setup, credentials, and `ima_api()` helper.
|
||||
|
||||
API base path: `openapi/wiki/v1`
|
||||
|
||||
通过 IMA Wiki OpenAPI 管理用户知识库,支持上传文件、添加网页链接、搜索知识库内容、浏览知识库列表和获取知识库详情。
|
||||
|
||||
完整的数据结构和接口参数详见 `knowledge-base-api.md`。
|
||||
|
||||
## 接口决策表
|
||||
|
||||
| 用户意图 | 调用接口 | 关键参数 |
|
||||
| --------------------------------------------- | ---------------------------------------------------------------------- | ------------------------------------------------------------------------ |
|
||||
| 上传文件到知识库 | `check_repeated_names` → `create_media` → COS Upload → `add_knowledge` | `media_type`(按扩展名),`knowledge_base_id`,`file_name`,`file_size` |
|
||||
| 上传文件到知识库的某个文件夹 | 先定位文件夹 → 同上(`folder_id` 传入目标文件夹 ID) | 见「文件夹操作」章节 |
|
||||
| 添加网页/微信文章到知识库 | `import_urls` | `urls`(1-10 个),`knowledge_base_id`,可选 `folder_id`(省略则根目录) |
|
||||
| 添加笔记到知识库 | `add_knowledge` | `media_type=11`,`note_info.content_id=<doc_id>`,`knowledge_base_id` |
|
||||
| 添加 URL(文件型)到知识库 | `check_repeated_names` → 下载文件 → 走"上传文件"流程 | URL 指向 PDF/Word/PPT 等文件时,按文件方式处理 |
|
||||
| 检查文件名是否重复 | `check_repeated_names` | `params[].name`,`params[].media_type`,`knowledge_base_id`,`folder_id` |
|
||||
| 获取知识库信息 | `get_knowledge_base` | `ids`(1-20 个,不重复) |
|
||||
| 浏览知识库内容列表 / 浏览文件夹 | `get_knowledge_list` | `knowledge_base_id`,`cursor`,`limit`(1~50),可选 `folder_id` |
|
||||
| 在知识库中搜索(含文件和文件夹) | `search_knowledge` | `query`,`knowledge_base_id`,`cursor` |
|
||||
| 按关键词查找知识库(用户知道名字但不知道 ID) | `search_knowledge_base` | `query`,`cursor`,`limit`(1~50) |
|
||||
| 查看/了解自己有哪些知识库 | `search_knowledge_base`(`query` 传空字符串) | `query: ""`,`cursor`,`limit`(1~50) |
|
||||
| 添加内容但**未指定**目标知识库 | `get_addable_knowledge_base_list` → 展示列表让用户选择 | `cursor`,`limit`(1~50) |
|
||||
|
||||
### `search_knowledge_base` vs `get_addable_knowledge_base_list` 选择指南
|
||||
|
||||
这两个接口容易混淆,选择规则:
|
||||
|
||||
| 场景 | 使用接口 | 原因 |
|
||||
| ------------------------------------------------ | ---------------------------------------------- | ---------------------------------- |
|
||||
| 用户说了知识库名称(如"添加到产品文档库") | `search_knowledge_base` | 按名称搜索,找到 ID 后继续操作 |
|
||||
| 用户想浏览/了解某个知识库 | `search_knowledge_base` → `get_knowledge_base` | 先搜到 ID,再获取详情 |
|
||||
| 用户想查看自己有哪些知识库(无具体关键词) | `search_knowledge_base`(`query: ""`) | 空 query 返回用户的所有知识库列表 |
|
||||
| 用户要添加内容但**没说添加到哪个知识库** | `get_addable_knowledge_base_list` | 列出有权限添加的知识库,让用户选择 |
|
||||
| 用户说"添加到知识库"但上下文中无法确定哪个知识库 | `get_addable_knowledge_base_list` | 同上,不要猜测,让用户选择 |
|
||||
|
||||
**绝不要**在用户已明确指定知识库名称时调用 `get_addable_knowledge_base_list`,直接用 `search_knowledge_base` 按名称搜索即可。
|
||||
|
||||
## 文件类型检测
|
||||
|
||||
使用 `scripts/preflight-check.cjs` 脚本自动完成类型检测和大小校验。脚本按以下优先级解析:
|
||||
|
||||
1. **`--content-type` 已提供且可识别** → content-type 优先,直接使用
|
||||
2. **`--content-type` 不可识别** → 回退到扩展名
|
||||
3. **未提供 `--content-type`** → 使用扩展名
|
||||
4. **两者都无法识别** → 拒绝处理
|
||||
|
||||
```bash
|
||||
# 有扩展名(自动推断)
|
||||
node ../scripts/preflight-check.cjs --file report.pdf
|
||||
|
||||
# 无扩展名或扩展名不可识别(需传入 content-type,如从 HTTP HEAD 获取)
|
||||
node ../scripts/preflight-check.cjs --file downloaded_file --content-type application/pdf
|
||||
```
|
||||
|
||||
扩展名与类型的对应关系:
|
||||
|
||||
| 扩展名 | media_type | content_type |
|
||||
| ------------------- | ---------- | ---------------------------------------------------------------------------- |
|
||||
| `.pdf` | 1 | `application/pdf` |
|
||||
| `.doc` | 3 | `application/msword` |
|
||||
| `.docx` | 3 | `application/vnd.openxmlformats-officedocument.wordprocessingml.document` |
|
||||
| `.ppt` | 4 | `application/vnd.ms-powerpoint` |
|
||||
| `.pptx` | 4 | `application/vnd.openxmlformats-officedocument.presentationml.presentation` |
|
||||
| `.xls` | 5 | `application/vnd.ms-excel` |
|
||||
| `.xlsx` | 5 | `application/vnd.openxmlformats-officedocument.spreadsheetml.sheet` |
|
||||
| `.csv` | 5 | `text/csv` |
|
||||
| `.md` / `.markdown` | 7 | `text/markdown` |
|
||||
| `.png` | 9 | `image/png` |
|
||||
| `.jpg` / `.jpeg` | 9 | `image/jpeg` |
|
||||
| `.webp` | 9 | `image/webp` |
|
||||
| `.txt` | 13 | `text/plain` |
|
||||
| `.xmind` | 14 | `application/x-xmind` / `application/vnd.xmind.workbook` / `application/zip` |
|
||||
| `.mp3` | 15 | `audio/mpeg` |
|
||||
| `.m4a` | 15 | `audio/x-m4a` |
|
||||
| `.wav` | 15 | `audio/wav` |
|
||||
| `.aac` | 15 | `audio/aac` |
|
||||
|
||||
未识别的扩展名或无扩展名:**直接告知用户该文件类型不被支持,立即终止操作**。不要猜测或默认为某个类型,**不要询问用户是否仍要上传**。
|
||||
|
||||
> **不支持的类型**:视频文件(`.mp4`、`.avi`、`.mov` 等)、Bilibili(`bilibili.com/video/`)和 YouTube(`youtube.com/watch`)链接、本地 HTML 文件(`file://`)**无法**通过 skill 添加到知识库。直接告知用户「该文件类型不支持,仅支持在 ima 桌面端内添加进知识库」,**不要提供上传选项或询问是否继续**。
|
||||
|
||||
## URL 类型检测
|
||||
|
||||
添加 URL 到知识库时,需要根据 URL 模式和 Content-Type 判断类型。检测按以下优先级进行:
|
||||
|
||||
**1. Content-Type 为 `text/html` 时,按 URL 模式区分:**
|
||||
|
||||
| URL 模式 | media_type | 类型 | 处理方式 |
|
||||
| --------------------------------------------------- | ---------- | -------------- | --------------------------------------------------------- |
|
||||
| 匹配 `mp.weixin.qq.com/s/` 或 `mp.weixin.qq.com/s?` | 6 | 微信公众号文章 | 使用 `import_urls` |
|
||||
| 以 `https://www.bilibili.com/video/` 开头 | ❌ 16 | 视频网页 | **不支持**,告知用户「仅支持在 ima 桌面端内添加进知识库」 |
|
||||
| 以 `https://www.youtube.com/watch` 开头 | ❌ 16 | 视频网页 | **不支持**,告知用户「仅支持在 ima 桌面端内添加进知识库」 |
|
||||
| 以 `file://` 开头 | ❌ | 本地 HTML | **不支持**,告知用户「仅支持在 ima 桌面端内添加进知识库」 |
|
||||
| 其他 `text/html` 页面 | 2 | 普通网页 | 使用 `import_urls` |
|
||||
|
||||
**2. Content-Type 为文件类型时:** 按文件类型检测表处理(PDF、Word、Excel 等)。
|
||||
|
||||
**3. 其他:** 告知用户该类型不被支持。
|
||||
|
||||
## 添加前置检查(Pre-flight Check)
|
||||
|
||||
在执行任何添加知识到知识库的操作前(`add_knowledge`、`import_urls`、上传文件流程),**必须按以下顺序逐项检查**,任一项不通过则**立即终止并告知用户,不要询问是否仍要尝试上传**:
|
||||
|
||||
### 1. 类型支持检查
|
||||
|
||||
| 检查项 | 条件 | 不通过时的处理 |
|
||||
| ------------------ | ------------------------------------------------------------------------- | --------------------------------------------- |
|
||||
| 文件扩展名是否支持 | 扩展名不在「文件类型检测」表中 | 告知用户该文件类型不被支持 |
|
||||
| 视频文件 | `.mp4`、`.avi`、`.mov` 等视频扩展名 | 告知用户「仅支持在 ima 桌面端内添加进知识库」 |
|
||||
| 视频网页 URL | `https://www.bilibili.com/video/` 或 `https://www.youtube.com/watch` 开头 | 告知用户「仅支持在 ima 桌面端内添加进知识库」 |
|
||||
| 本地 HTML 文件 | `file://` 协议 | 告知用户「仅支持在 ima 桌面端内添加进知识库」 |
|
||||
|
||||
### 2. 文件大小检查
|
||||
|
||||
上传前必须校验文件大小,超限文件应**在上传前拦截**,不要发起请求:
|
||||
|
||||
| 文件类型 | media_type | 最大大小 |
|
||||
| --------------------------- | ----------- | -------- |
|
||||
| Excel、TXT、Xmind、Markdown | 5/13/14/7 | 10 MB |
|
||||
| 图片 | 9 | 30 MB |
|
||||
| PDF、Word、PPT、音频及其他 | 1/3/4/15 等 | 200 MB |
|
||||
|
||||
### 3. 音频时长检查
|
||||
|
||||
音频文件(media_type=15)额外限制:**最长 2 小时**。超过时告知用户。
|
||||
|
||||
### 4. 文件名重复检查
|
||||
|
||||
仅适用于文件类型(media_type 1/3/4/5/7/9/13/14/15),不适用于网页(2/6)、笔记(11)等:
|
||||
|
||||
- 调用 `check_repeated_names` 检查
|
||||
- `is_repeated=true`:询问用户是否保留两者(追加时间戳)或取消
|
||||
- 不支持"替换"操作
|
||||
|
||||
> **检查顺序很重要**:先做类型和大小检查(本地即可判断),通过后再调用远程接口检查重名。避免对不支持或超限的文件发起不必要的网络请求。
|
||||
|
||||
## 常用工作流
|
||||
|
||||
### 上传文件到知识库
|
||||
|
||||
完成「添加前置检查」后,执行以下步骤:创建媒体 → 上传 COS → 添加知识。
|
||||
|
||||
> 前置检查(类型检测、大小校验、重名检查)见「添加前置检查」章节。
|
||||
|
||||
```bash
|
||||
# 1. 前置检查 — 类型、大小一步完成
|
||||
# 有扩展名时:自动从扩展名推断 media_type 和 content_type
|
||||
# 无扩展名时:需通过 --content-type 传入(如从 HTTP HEAD 获取)
|
||||
PREFLIGHT=$(node ../scripts/preflight-check.cjs \
|
||||
--file "/path/to/report.pdf")
|
||||
echo "$PREFLIGHT"
|
||||
# pass=false 时直接终止,将 reason 展示给用户
|
||||
|
||||
# 2. 从 preflight 结果提取字段(用于后续 API 调用)
|
||||
FILE_NAME=$(echo "$PREFLIGHT" | node -e "const d=JSON.parse(require('fs').readFileSync(0,'utf8'));process.stdout.write(d.file_name)")
|
||||
FILE_EXT=$(echo "$PREFLIGHT" | node -e "const d=JSON.parse(require('fs').readFileSync(0,'utf8'));process.stdout.write(d.file_ext)")
|
||||
FILE_SIZE=$(echo "$PREFLIGHT" | node -e "const d=JSON.parse(require('fs').readFileSync(0,'utf8'));process.stdout.write(String(d.file_size))")
|
||||
MEDIA_TYPE=$(echo "$PREFLIGHT" | node -e "const d=JSON.parse(require('fs').readFileSync(0,'utf8'));process.stdout.write(String(d.media_type))")
|
||||
CONTENT_TYPE=$(echo "$PREFLIGHT" | node -e "const d=JSON.parse(require('fs').readFileSync(0,'utf8'));process.stdout.write(d.content_type)")
|
||||
|
||||
# 3. 重名检查(仅文件类型,见「添加前置检查」第 4 步)
|
||||
|
||||
# 4. create_media — 获取 media_id 和 COS 上传凭证
|
||||
ima_api "openapi/wiki/v1/create_media" "{
|
||||
\"file_name\": \"$FILE_NAME\",
|
||||
\"file_size\": $FILE_SIZE,
|
||||
\"content_type\": \"$CONTENT_TYPE\",
|
||||
\"knowledge_base_id\": \"<kb_id>\",
|
||||
\"file_ext\": \"$FILE_EXT\"
|
||||
}"
|
||||
# 从返回值提取 media_id 和 cos_credential 各字段
|
||||
|
||||
# 5. 上传文件到 COS
|
||||
node ../scripts/cos-upload.cjs \
|
||||
--file "/path/to/report.pdf" \
|
||||
--secret-id "<cos_credential.secret_id>" \
|
||||
--secret-key "<cos_credential.secret_key>" \
|
||||
--token "<cos_credential.token>" \
|
||||
--bucket "<cos_credential.bucket_name>" \
|
||||
--region "<cos_credential.region>" \
|
||||
--cos-key "<cos_credential.cos_key>" \
|
||||
--content-type "$CONTENT_TYPE" \
|
||||
--start-time "<cos_credential.start_time>" \
|
||||
--expired-time "<cos_credential.expired_time>"
|
||||
|
||||
# 6. add_knowledge — 将已上传的文件关联到知识库
|
||||
ima_api "openapi/wiki/v1/add_knowledge" "{
|
||||
\"media_type\": $MEDIA_TYPE,
|
||||
\"media_id\": \"<media_id>\",
|
||||
\"title\": \"$FILE_NAME\",
|
||||
\"knowledge_base_id\": \"<kb_id>\",
|
||||
\"file_info\": {
|
||||
\"cos_key\": \"<cos_credential.cos_key>\",
|
||||
\"file_size\": $FILE_SIZE,
|
||||
\"file_name\": \"$FILE_NAME\"
|
||||
}
|
||||
}"
|
||||
```
|
||||
|
||||
#### 批量上传时的重复处理
|
||||
|
||||
当上传多个文件时,可一次性检查所有文件名(最多 2000 个):
|
||||
|
||||
```bash
|
||||
# 批量检查
|
||||
ima_api "openapi/wiki/v1/check_repeated_names" '{
|
||||
"params": [
|
||||
{"name": "report.pdf", "media_type": 1},
|
||||
{"name": "slides.pptx", "media_type": 4},
|
||||
{"name": "data.xlsx", "media_type": 5}
|
||||
],
|
||||
"knowledge_base_id": "<kb_id>",
|
||||
"folder_id": "<folder_id>"
|
||||
}'
|
||||
# 注意:如果是根目录,省略 folder_id 字段
|
||||
# 遍历 results,对 is_repeated=true 的文件询问用户:
|
||||
# - "以下文件在知识库中已存在同名文件:report.pdf、data.xlsx。是否保留两者?(不支持替换)"
|
||||
# - 用户选择"保留两者"的文件:追加时间戳后继续上传
|
||||
# - 用户选择"取消"的文件:从上传列表中移除
|
||||
```
|
||||
|
||||
**时间戳命名规则**:在文件名(不含扩展名)末尾追加 `_YYYYMMDDHHmmss`,例如 `report_20260317153000.pdf`。
|
||||
|
||||
### 添加网页/微信文章到知识库
|
||||
|
||||
使用 `import_urls` 批量导入网页和微信公众号文章(1-10 个 URL),服务端自动识别类型:
|
||||
|
||||
```bash
|
||||
# 添加到根目录(不传 folder_id)
|
||||
ima_api "openapi/wiki/v1/import_urls" '{
|
||||
"knowledge_base_id": "<kb_id>",
|
||||
"urls": [
|
||||
"https://example.com/article",
|
||||
"https://mp.weixin.qq.com/s/xxxxx"
|
||||
]
|
||||
}'
|
||||
|
||||
# 添加到指定文件夹(传 folder_id,以 folder_ 开头)
|
||||
ima_api "openapi/wiki/v1/import_urls" '{
|
||||
"knowledge_base_id": "<kb_id>",
|
||||
"folder_id": "<folder_id>",
|
||||
"urls": [
|
||||
"https://example.com/article"
|
||||
]
|
||||
}'
|
||||
# 返回 results 映射:{ "<url>": { url, ret_code, media_id } }
|
||||
# ret_code=0 表示成功,非 0 查看 errmsg
|
||||
```
|
||||
|
||||
### 添加笔记到知识库
|
||||
|
||||
将已有笔记(通过 `doc_id` 引用)直接关联到知识库,无需下载内容:
|
||||
|
||||
```bash
|
||||
ima_api "openapi/wiki/v1/add_knowledge" '{
|
||||
"media_type": 11,
|
||||
"note_info": { "content_id": "<doc_id>" },
|
||||
"title": "笔记标题",
|
||||
"knowledge_base_id": "<kb_id>"
|
||||
}'
|
||||
```
|
||||
|
||||
### 添加 URL 到知识库(自动检测文件型 URL)
|
||||
|
||||
当用户提供 URL 时,需先判断该 URL 指向的是网页还是可下载文件(PDF、Word、PPT 等)。
|
||||
|
||||
**判断规则**(按优先级):
|
||||
|
||||
1. **URL 路径包含文件扩展名**:如 `https://arxiv.org/pdf/2603.12268` 以 `/pdf/` 开头,或 `https://example.com/report.pdf` 以 `.pdf` 结尾 → 文件型
|
||||
2. **发送 HEAD 请求检查 Content-Type**:`curl -sI -L <url>` 查看响应头
|
||||
- `application/pdf` → PDF 文件
|
||||
- `application/msword` 或 `application/vnd.openxmlformats-*` → Word/PPT/Excel 文件
|
||||
- `text/html` → 网页
|
||||
3. **已知文件型 URL 模式**:
|
||||
- `arxiv.org/pdf/*` → PDF
|
||||
- `*.pdf`、`*.docx`、`*.pptx`、`*.xlsx` 结尾 → 对应文件类型
|
||||
- GitHub raw 文件链接 → 按扩展名判断
|
||||
|
||||
**文件型 URL 处理流程**:
|
||||
|
||||
```bash
|
||||
# 1. 探测 URL 类型
|
||||
CONTENT_TYPE=$(curl -sI -L "https://arxiv.org/pdf/2603.12268" | grep -i "^content-type:" | tail -1 | awk '{print $2}' | tr -d '\r')
|
||||
# 结果如 application/pdf → 文件型
|
||||
|
||||
# 2. 下载文件到临时目录
|
||||
TEMP_DIR=$(mktemp -d)
|
||||
# 根据 Content-Type 或 URL 推断文件名和扩展名
|
||||
curl -sL -o "$TEMP_DIR/paper.pdf" "https://arxiv.org/pdf/2603.12268"
|
||||
|
||||
# 3. 前置检查(传入 content-type 作为备用,文件名有扩展名时会优先用扩展名)
|
||||
PREFLIGHT=$(node ../scripts/preflight-check.cjs \
|
||||
--file "$TEMP_DIR/paper.pdf" --content-type "$CONTENT_TYPE")
|
||||
echo "$PREFLIGHT"
|
||||
# pass=false 时直接终止
|
||||
|
||||
# 4. 按"上传文件到知识库"流程处理:create_media → COS Upload → add_knowledge
|
||||
# (参见上方"上传文件到知识库"工作流)
|
||||
|
||||
# 5. 清理临时文件
|
||||
rm -rf "$TEMP_DIR"
|
||||
```
|
||||
|
||||
**文件名推断**:
|
||||
|
||||
- 优先从 `Content-Disposition` 响应头提取文件名
|
||||
- 其次从 URL 路径中提取(如 `/pdf/2603.12268` → `2603.12268.pdf`)
|
||||
- 最后使用 URL 的最后一段路径 + 根据 Content-Type 补充扩展名
|
||||
|
||||
### 文件夹操作
|
||||
|
||||
知识库内容以文件夹结构组织。**文件夹本身也是一种知识条目**,在 `get_knowledge_list` 和 `search_knowledge` 的返回结果中会同时包含文件和文件夹。
|
||||
|
||||
#### 核心概念
|
||||
|
||||
- `folder_id`:文件夹的唯一标识,**始终以 `folder_` 前缀开头**(如 `folder_abc123`),在 `add_knowledge`、`import_urls`、`get_knowledge_list`、`check_repeated_names` 等接口中用于指定目标文件夹
|
||||
- **操作根目录时,不要传 `folder_id` 参数**(直接省略该字段),不要将 `knowledge_base_id` 作为 `folder_id` 传入
|
||||
- `get_knowledge_list` 返回的 `current_path`(`FolderInfo[]`)表示当前浏览位置的完整路径(面包屑)
|
||||
|
||||
#### 定位文件夹(用户提到文件夹名时)
|
||||
|
||||
当用户说「添加到 XX 文件夹」但只给了文件夹名称时,需要先找到 `folder_id`:
|
||||
|
||||
```bash
|
||||
# 方法 1:搜索知识库内容(推荐,可直接按名称搜索文件夹)
|
||||
ima_api "openapi/wiki/v1/search_knowledge" '{
|
||||
"query": "文件夹名称",
|
||||
"knowledge_base_id": "<kb_id>",
|
||||
"cursor": ""
|
||||
}'
|
||||
# 从返回的 info_list 中找到匹配的文件夹条目,取其 media_id 作为 folder_id
|
||||
|
||||
# 方法 2:浏览根目录列表逐级查找
|
||||
ima_api "openapi/wiki/v1/get_knowledge_list" '{
|
||||
"knowledge_base_id": "<kb_id>",
|
||||
"cursor": "",
|
||||
"limit": 50
|
||||
}'
|
||||
# 从返回的 knowledge_list 中找到目标文件夹,取其 media_id 作为 folder_id
|
||||
# 如果文件夹在子目录中,需要用返回的 folder_id 逐级深入
|
||||
```
|
||||
|
||||
#### 添加内容到指定文件夹
|
||||
|
||||
所有写入接口(`add_knowledge`、`import_urls`、`check_repeated_names`)都支持 `folder_id` 参数:
|
||||
|
||||
```bash
|
||||
# 上传文件到指定文件夹
|
||||
ima_api "openapi/wiki/v1/add_knowledge" '{
|
||||
"media_type": 1,
|
||||
"media_id": "<media_id>",
|
||||
"title": "report.pdf",
|
||||
"knowledge_base_id": "<kb_id>",
|
||||
"folder_id": "<folder_id>",
|
||||
"file_info": { "cos_key": "...", "file_size": 12345, "file_name": "report.pdf" }
|
||||
}'
|
||||
|
||||
# 导入网页到指定文件夹
|
||||
ima_api "openapi/wiki/v1/import_urls" '{
|
||||
"knowledge_base_id": "<kb_id>",
|
||||
"folder_id": "<folder_id>",
|
||||
"urls": ["https://example.com/article"]
|
||||
}'
|
||||
|
||||
# 添加到根目录时,直接省略 folder_id
|
||||
ima_api "openapi/wiki/v1/add_knowledge" '{
|
||||
"media_type": 1,
|
||||
"media_id": "<media_id>",
|
||||
"title": "report.pdf",
|
||||
"knowledge_base_id": "<kb_id>",
|
||||
"file_info": { "cos_key": "...", "file_size": 12345, "file_name": "report.pdf" }
|
||||
}'
|
||||
```
|
||||
|
||||
### 获取知识库信息
|
||||
|
||||
```bash
|
||||
ima_api "openapi/wiki/v1/get_knowledge_base" '{"ids": ["<kb_id>"]}'
|
||||
# 返回 infos 映射:{ "<kb_id>": { id, name, cover_url, description, recommended_questions } }
|
||||
```
|
||||
|
||||
### 浏览知识库内容
|
||||
|
||||
```bash
|
||||
# 浏览根目录
|
||||
ima_api "openapi/wiki/v1/get_knowledge_list" '{"knowledge_base_id": "<kb_id>", "cursor": "", "limit": 20}'
|
||||
|
||||
# 浏览指定文件夹
|
||||
ima_api "openapi/wiki/v1/get_knowledge_list" '{"knowledge_base_id": "<kb_id>", "folder_id": "<folder_id>", "cursor": "", "limit": 20}'
|
||||
# 翻页:用 next_cursor,is_end=true 时停止
|
||||
```
|
||||
|
||||
### 在知识库中搜索
|
||||
|
||||
```bash
|
||||
ima_api "openapi/wiki/v1/search_knowledge" '{"query": "搜索关键词", "knowledge_base_id": "<kb_id>", "cursor": ""}'
|
||||
```
|
||||
|
||||
### 搜索知识库列表
|
||||
|
||||
```bash
|
||||
# 按关键词搜索
|
||||
ima_api "openapi/wiki/v1/search_knowledge_base" '{"query": "搜索关键词", "cursor": "", "limit": 20}'
|
||||
|
||||
# 查看所有知识库(空 query)
|
||||
ima_api "openapi/wiki/v1/search_knowledge_base" '{"query": "", "cursor": "", "limit": 20}'
|
||||
```
|
||||
|
||||
### 获取可添加的知识库列表
|
||||
|
||||
**仅当用户要添加内容但未指定目标知识库时使用**。如果用户已给出知识库名称,应使用 `search_knowledge_base` 按名称搜索,而非此接口。
|
||||
|
||||
```bash
|
||||
# 首次请求
|
||||
ima_api "openapi/wiki/v1/get_addable_knowledge_base_list" '{"cursor": "", "limit": 20}'
|
||||
# 翻页:用 next_cursor,is_end=true 时停止
|
||||
```
|
||||
|
||||
## 核心响应字段
|
||||
|
||||
知识条目(`KnowledgeInfo`)关键字段:`media_id`(媒体ID)、`title`、`parent_folder_id`。
|
||||
|
||||
搜索到的知识条目(`SearchedKnowledgeInfo`)关键字段:`media_id`、`title`、`parent_folder_id`、`highlight_content`(高亮内容,内容匹配时返回)。
|
||||
|
||||
知识库信息(`KnowledgeBaseInfo`)关键字段:`id`、`name`、`cover_url`、`description`、`recommended_questions`。
|
||||
|
||||
文件夹条目(`FolderInfo`)关键字段:`folder_id`、`name`、`file_number`、`folder_number`、`parent_folder_id`、`is_top`。
|
||||
|
||||
完整字段定义见 `knowledge-base-api.md`。
|
||||
|
||||
## 分页
|
||||
|
||||
所有列表和搜索接口使用**游标分页**:
|
||||
|
||||
1. 首次请求:`cursor: ""`
|
||||
2. 检查返回的 `is_end`:`false` 表示还有更多数据
|
||||
3. 将返回的 `next_cursor` 作为下次请求的 `cursor`
|
||||
4. `is_end = true` 时停止翻页
|
||||
|
||||
## 响应处理
|
||||
|
||||
所有 API 返回统一结构 `{ "retcode": 0, "errmsg": "...", "data": { ... } }`:
|
||||
|
||||
- `retcode=0`:成功,从 `data` 提取业务字段
|
||||
- `retcode≠0`:失败,**直接将 `errmsg` 展示给用户**即可,不需要自行翻译错误码
|
||||
|
||||
## 用户体验
|
||||
|
||||
- **隐藏内部 ID**:面向用户的展示中**永远不要暴露 `knowledge_base_id`、`media_id`、`folder_id` 等内部 ID**。始终使用知识库名称、文件标题、文件夹名称等用户可读信息。ID 仅用于后续 API 调用,不展示给用户。
|
||||
- ✅ `"已添加到知识库「产品文档库」✓"`
|
||||
- ❌ `"已添加到知识库 abc123def456 ✓"`
|
||||
- 需要引用知识库时,先通过 `get_knowledge_base` 获取名称,再展示
|
||||
- **精简进度**:不要逐步暴露内部操作(如"正在创建媒体…正在上传 COS…")。只报告用户关心的信息:
|
||||
- 上传文件:`"正在上传 report.pdf…"` → `"已添加到知识库「产品文档库」✓"`
|
||||
- 添加网页:`"正在添加…"` → `"已添加到「产品文档库」✓"`
|
||||
- 失败时展示 `errmsg` 即可
|
||||
- **批量操作**:汇总结果,如 `"3 个文件已添加到「产品文档库」,1 个失败(data.xlsx: 文件大小超限)"`
|
||||
- **格式化展示**:读取类操作的结果应以结构化格式展示给用户,而非原始 JSON:
|
||||
|
||||
**知识库列表**(`search_knowledge_base` / `get_addable_knowledge_base_list`):
|
||||
|
||||
> 搜索知识库后,用返回的 ID 列表调用 `get_knowledge_base` 获取描述信息,一并展示。
|
||||
|
||||
```
|
||||
📚 搜索结果(共 3 个知识库):
|
||||
1. **产品文档库** — 存放产品相关的所有文档资料
|
||||
2. **技术方案库** — 各项目技术方案汇总
|
||||
3. **竞品分析库**
|
||||
```
|
||||
|
||||
**知识库内容列表**(`get_knowledge_list`):
|
||||
|
||||
```
|
||||
📂 知识库「产品文档库」内容:
|
||||
📁 设计文档/ (3 个文件, 1 个子文件夹)
|
||||
📁 会议纪要/ (12 个文件)
|
||||
📄 产品需求文档.pdf
|
||||
📄 技术方案.docx
|
||||
📄 数据分析.xlsx
|
||||
--- 第 1 页,还有更多内容 ---
|
||||
```
|
||||
|
||||
**搜索结果**(`search_knowledge`):
|
||||
|
||||
> `search_knowledge` 返回的条目包含 `media_id`、`title`、`parent_folder_id`、`highlight_content`(内容匹配时返回高亮片段),
|
||||
|
||||
```
|
||||
🔍 在知识库「产品文档库」中搜索「排期」的结果:
|
||||
|
||||
1. 📄 Q1排期表.xlsx (文件夹: 项目管理/)
|
||||
> ...包含**排期**计划的详细信息...
|
||||
2. 📄 开发排期讨论.pdf (文件夹: 会议纪要/)
|
||||
3. 📁 排期模板/ (文件夹: 根目录)
|
||||
```
|
||||
|
||||
**知识库详情**(`get_knowledge_base`):
|
||||
|
||||
```
|
||||
📚 产品文档库
|
||||
📝 描述:存放产品相关的所有文档资料
|
||||
💡 推荐问题:
|
||||
- 最新的产品需求是什么?
|
||||
- 技术方案有哪些?
|
||||
```
|
||||
|
||||
## 注意事项
|
||||
|
||||
- `get_knowledge_base` 接受 1-20 个 ID;单个 ID 也需包装为数组
|
||||
- `get_knowledge_list` 的 `limit` 范围为 1~50
|
||||
- **文件夹是知识条目的一种**:`get_knowledge_list` 和 `search_knowledge` 的返回结果中同时包含文件和文件夹,需通过字段区分(文件夹有 `folder_id`/`name`/`file_number`/`folder_number`,文件有 `media_id`/`title`)
|
||||
- **用户提到文件夹时**:如果用户只给了文件夹名称(而非 ID),必须先通过 `search_knowledge` 或 `get_knowledge_list` 找到对应的 `folder_id`,再执行后续操作
|
||||
- `folder_id` 在 `add_knowledge`、`import_urls`、`get_knowledge_list`、`check_repeated_names` 中均为可选字段,**操作根目录时直接省略 `folder_id`,不要传该参数**。`folder_id` 的值始终以 `folder_` 前缀开头(如 `folder_abc123`),**不要将 `knowledge_base_id` 作为 `folder_id` 传入**
|
||||
- **文件上传时 `title` 必须等于 `file_name`**:调用 `add_knowledge` 添加文件时,`title` 字段**必须使用文件的原始完整文件名(含扩展名)**,不要自行拟定标题。`file_name` 和 `title` 传同一个值。**禁止缩短、翻译、重命名或省略任何部分**。例如文件名为 `音频.mp3`,则 `file_name` 和 `title` 都必须传 `音频.mp3`
|
||||
- 文件扩展名必须正确提取,用于 `media_type` 检测和 `file_ext` 字段(无点号,如 `pdf`)
|
||||
- COS 上传脚本失败(非零退出码)时,不要继续调用 `add_knowledge`
|
||||
- COS 上传时 `--content-type` 应传入文件的实际 MIME 类型(如 `application/pdf`),而非通用的 `application/octet-stream`
|
||||
- 当用户提供 URL 添加到知识库时,必须先检测 URL 是否指向文件(通过 URL 路径扩展名 + HEAD 请求 Content-Type),文件型 URL 需下载后走上传流程;网页/微信文章型 URL 使用 `import_urls`
|
||||
@@ -0,0 +1,341 @@
|
||||
# IMA笔记 API
|
||||
|
||||
## ⚠️ 必读约束
|
||||
|
||||
### 🔒 认证
|
||||
|
||||
所有请求必须携带 Header:
|
||||
|
||||
```
|
||||
ima-openapi-clientid: {IMA_OPENAPI_CLIENTID}
|
||||
ima-openapi-apikey: {IMA_OPENAPI_APIKEY}
|
||||
Content-Type: application/json
|
||||
```
|
||||
|
||||
### 🔒 安全规则
|
||||
|
||||
- 笔记属于用户隐私,**不要在群聊中主动展示笔记内容**。
|
||||
- 仅响应授权用户的笔记操作请求。
|
||||
|
||||
---
|
||||
|
||||
## 快速决策
|
||||
|
||||
| 用户意图 | 接口别名 |
|
||||
| ------------------------------------ | --------------------------------------------- |
|
||||
| 「搜索笔记」「找包含XX的笔记」 | `/openapi/note/v1/search_note_book` |
|
||||
| 「列出笔记本」「有哪些笔记本」 | `/openapi/note/v1/list_note_folder_by_cursor` |
|
||||
| 「查看XX笔记本里的笔记」 | `/openapi/note/v1/list_note_by_folder_id` |
|
||||
| 「从markdown新建笔记」「导入笔记」「创建笔记」「生成笔记」| `/openapi/note/v1/import_doc` |
|
||||
| 「追加内容到笔记」「在笔记末尾添加」 | `/openapi/note/v1/append_doc` |
|
||||
| 「获取笔记纯文本」「读取笔记内容」 | `/openapi/note/v1/get_doc_content` |
|
||||
|
||||
---
|
||||
|
||||
## 数据结构
|
||||
|
||||
---
|
||||
|
||||
#### DocBasicInfo
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ------------ | -------- | ------------------------ |
|
||||
| `basic_info` | DocBasic | 见 [DocBasic](#docbasic) |
|
||||
|
||||
---
|
||||
|
||||
#### DocBasic
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| --------------- | --------------------- | ------------------------------ |
|
||||
| `docid` | string | 文章 id |
|
||||
| `title` | string | 标题 |
|
||||
| `summary` | string | 简介 |
|
||||
| `create_time` | int64 | |
|
||||
| `modify_time` | int64 | |
|
||||
| `status` | DocStatus | 文章状态,`0`=正常,`1`=已删除 |
|
||||
| `folder_id` | string | 文件夹 id |
|
||||
| `folder_name` | string | 文件夹名称 |
|
||||
| `summary_style` | map\<string, string\> | 简介样式 |
|
||||
|
||||
---
|
||||
|
||||
### FolderItem(笔记本条目)
|
||||
|
||||
`list_note_folder_by_cursor` 返回的笔记本对象,字段如下:
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ------------------ | ------ | -------------------------------------------- |
|
||||
| `folder_id` | string | 笔记本唯一 ID |
|
||||
| `name` | string | 笔记本名称 |
|
||||
| `note_number` | int64 | 笔记本内笔记数量 |
|
||||
| `create_time` | int64 | 创建时间(Unix 毫秒) |
|
||||
| `modify_time` | int64 | 修改时间(Unix 毫秒) |
|
||||
| `parent_folder_id` | string | 上级笔记本 ID(支持嵌套) |
|
||||
| `folder_type` | int | 类型:`0`=用户自建,`1`=全部笔记,`2`=未分类 |
|
||||
| `status` | int | 状态:`0`=正常,`1`=已删除 |
|
||||
|
||||
---
|
||||
|
||||
#### QueryInfo
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| --------- | ------ | ---------- |
|
||||
| `title` | string | 标题 query |
|
||||
| `content` | string | 正文 query |
|
||||
|
||||
---
|
||||
|
||||
#### SearchedDoc
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ---------------- | --------------------- | ------------------------------------------------------------------------------------------ |
|
||||
| `doc` | DocBasicInfo | 笔记 basic 数据,见 [DocBasicInfo](#docbasicinfo) |
|
||||
| `highlight_info` | map\<string, string\> | 该条笔记匹配的高亮词,key: `doc_title`(文档标题),value: 包含 `<em>高亮词</em>` 的字段值 |
|
||||
|
||||
---
|
||||
|
||||
#### NoteBookFolder
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| -------- | ----------------------- | -------------------------------------------------------------------------------- |
|
||||
| `folder` | NoteBookFolderBasicInfo | 笔记本信息,非笔记本为空,见 [NoteBookFolderBasicInfo](#notebookfolderbasicinfo) |
|
||||
|
||||
---
|
||||
|
||||
#### NoteBookFolderBasicInfo
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ------------ | ------------------- | ---------------------------------------------- |
|
||||
| `basic_info` | NoteBookFolderBasic | 见 [NoteBookFolderBasic](#notebookfolderbasic) |
|
||||
|
||||
---
|
||||
|
||||
#### NoteBookFolderBasic
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ------------- | ---------- | -------------------------------------------------- |
|
||||
| `folder_id` | string | 文件夹 id |
|
||||
| `name` | string | 笔记本名称 |
|
||||
| `status` | DocStatus | 笔记本状态,`0`=正常,`1`=已删除 |
|
||||
| `create_time` | int64 | 创建时间 |
|
||||
| `modify_time` | int64 | 修改时间 |
|
||||
| `note_number` | int64 | 笔记数量 |
|
||||
| `folder_type` | FolderType | 文件夹类型:`0`=用户自建,`1`=全部笔记,`2`=未分类 |
|
||||
|
||||
---
|
||||
|
||||
#### NoteBookInfo
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ------------ | ------------ | ---------------------------------------------- |
|
||||
| `basic_info` | DocBasicInfo | 笔记基础信息,见 [DocBasicInfo](#docbasicinfo) |
|
||||
|
||||
---
|
||||
|
||||
## 接口详情
|
||||
|
||||
### 1. 搜索笔记
|
||||
|
||||
POST /openapi/note/v1/search_note_book
|
||||
|
||||
**触发场景**:用户说「搜索」「找笔记」「查找包含XX的内容」
|
||||
|
||||
#### 请求参数
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
| ------------- | ---------- | ---- | ------------------------------------------------------------------------ |
|
||||
| `search_type` | SearchType | 否 | 检索方式,默认为标题,`0`=标题,`1`=正文 |
|
||||
| `sort_type` | SortType | 否 | 排序方式,默认为更新时间,`0`=更新时间,`1`=创建时间,`2`=标题,`3`=大小 |
|
||||
| `query_info` | QueryInfo | 否 | 用户 query,见 [QueryInfo](#queryinfo) |
|
||||
| `start` | int64 | 是 | 翻页字段 |
|
||||
| `end` | int64 | 是 | 翻页字段 |
|
||||
| `query_id` | string | 否 | queryid |
|
||||
|
||||
#### 返回字段
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| --------------- | ------------- | ------------------------------------------------- |
|
||||
| `docs` | SearchedDoc[] | 检索到的笔记 list,见 [SearchedDoc](#searcheddoc) |
|
||||
| `is_end` | bool | 是否为最后一批数据 |
|
||||
| `total_hit_num` | int64 | 检索命中结果总数 |
|
||||
|
||||
---
|
||||
|
||||
### 2. 列出笔记本
|
||||
|
||||
POST /openapi/note/v1/list_note_folder_by_cursor
|
||||
|
||||
**触发场景**:用户说「列出笔记本」「有哪些分类」「查看笔记本目录」
|
||||
|
||||
#### 请求参数
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
| -------- | ------ | ------ | ---------------------------------------- |
|
||||
| `cursor` | string | **是** | 游标,第一页传 `"0"`,后续传后台返回的值 |
|
||||
| `limit` | uint64 | **是** | 获取笔记数量限制 |
|
||||
|
||||
#### 返回字段
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ------------------- | ---------------- | ------------------------------------ |
|
||||
| `note_book_folders` | NoteBookFolder[] | 见 [NoteBookFolder](#notebookfolder) |
|
||||
| `next_cursor` | string | 下次请求的起始游标 |
|
||||
| `is_end` | bool | 是否为最后一批数据 |
|
||||
|
||||
---
|
||||
|
||||
### 3. 按笔记本拉取笔记列表
|
||||
|
||||
POST /openapi/note/v1/list_note_by_folder_id
|
||||
|
||||
**触发场景**:用户说「查看XX笔记本的笔记」「列出这个笔记本里的内容」
|
||||
|
||||
> 全部笔记根目录的 `folder_id` 为 `user_list_{userid}`,可从「列出笔记本」返回的 `folder_id` 获取。
|
||||
|
||||
#### 请求参数
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
| ----------- | ------ | ------ | ----------------------------- |
|
||||
| `folder_id` | string | 否 | 笔记本 ID,根目录为空 |
|
||||
| `cursor` | string | **是** | 当前游标,首次传空字符串 `""` |
|
||||
| `limit` | uint64 | **是** | 获取笔记数量限制 |
|
||||
|
||||
#### 返回字段
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| ---------------- | -------------- | -------------------------------- |
|
||||
| `note_book_list` | NoteBookInfo[] | 见 [NoteBookInfo](#notebookinfo) |
|
||||
| `next_cursor` | string | 下次请求的起始游标 |
|
||||
| `is_end` | bool | 是否为最后一批数据 |
|
||||
|
||||
---
|
||||
|
||||
### 4. 从 Markdown 新建笔记
|
||||
|
||||
POST /openapi/note/v1/import_doc
|
||||
|
||||
**触发场景**:用户说「从 Markdown 新建笔记」「导入笔记」「把这段 Markdown 保存为笔记」
|
||||
|
||||
#### 请求参数
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
| ---------------- | ------ | ------ | --------------------------------------------------------------- |
|
||||
| `content_format` | int | **是** | 文本类型:`1`=Markdown(默认)目前仅支持 `MARKDOWN`(值为 `1`) |
|
||||
| `content` | string | **是** | 笔记正文内容, 只支持markdown格式 |
|
||||
| `folder_id` | string | 否 | 关联的笔记本id |
|
||||
|
||||
#### 返回字段
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| -------- | ------ | ------------- |
|
||||
| `doc_id` | string | 新doc的唯一ID |
|
||||
|
||||
---
|
||||
|
||||
### 5. 追加内容到笔记
|
||||
|
||||
POST /openapi/note/v1/append_doc
|
||||
|
||||
**触发场景**:用户说「在这篇笔记末尾追加内容」「把 XX 添加到笔记里」
|
||||
|
||||
> ⚠️ **敏感操作**:追加会不可撤销地修改已有笔记。如果用户没有明确指定目标笔记(提供 `doc_id` 或笔记标题),**必须先向用户确认目标笔记**,不得自行猜测。模糊场景应优先建议用户使用 `import_doc` 新建笔记。
|
||||
|
||||
#### 请求参数
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
| ---------------- | ------ | ------ | --------------------------------------------------------------- |
|
||||
| `doc_id` | string | **是** | 目标笔记的唯一ID, 需要是本人的笔记 |
|
||||
| `content_format` | int | **是** | 文本类型:`1`=Markdown(默认)目前仅支持 `MARKDOWN`(值为 `1`) |
|
||||
| `content` | string | **是** | 要追加的文本内容, 只支持markdown格式 |
|
||||
|
||||
#### 返回字段
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| -------- | ------ | ---------------- |
|
||||
| `doc_id` | string | 目标笔记的唯一ID |
|
||||
|
||||
---
|
||||
|
||||
### 6. 获取笔记纯文本
|
||||
|
||||
POST /openapi/note/v1/get_doc_content
|
||||
|
||||
**触发场景**:用户说「读取笔记内容」「获取这篇笔记的纯文本」「把笔记转成 Markdown」
|
||||
|
||||
#### 请求参数
|
||||
|
||||
| 字段 | 类型 | 必填 | 说明 |
|
||||
| ----------------------- | ------ | ------ | ------------------------------------------------------------------ |
|
||||
| `doc_id` | string | **是** | 目标笔记的唯一 ID, 需要是本人的笔记 |
|
||||
| `target_content_format` | int | **是** | 目标文本类型:`0`=纯文本(推荐),`1`=Markdown(不支持),`2`=JSON |
|
||||
|
||||
#### 返回字段
|
||||
|
||||
| 字段 | 类型 | 说明 |
|
||||
| --------- | ------ | ----------------------------------------------------- |
|
||||
| `content` | string | 笔记的文本内容(按 `target_content_format` 格式返回) |
|
||||
|
||||
---
|
||||
|
||||
## 枚举值
|
||||
|
||||
### `sort_type`(排序方式)
|
||||
|
||||
| 值 | 说明 |
|
||||
| --- | ---------------- |
|
||||
| `0` | 更新时间(默认) |
|
||||
| `1` | 创建时间 |
|
||||
| `2` | 标题 |
|
||||
| `3` | 大小 |
|
||||
|
||||
### `search_type`(检索方式)
|
||||
|
||||
| 值 | 说明 |
|
||||
| --- | ---------------- |
|
||||
| `0` | 标题检索(默认) |
|
||||
| `1` | 正文检索 |
|
||||
|
||||
### `content_format`(文本类型)
|
||||
|
||||
| 值 | 说明 |
|
||||
| --- | ------------------------ |
|
||||
| `0` | PLAINTEXT - 纯文本 |
|
||||
| `1` | MARKDOWN - Markdown 格式 |
|
||||
| `2` | JSON - JSON 格式 |
|
||||
|
||||
### `FolderType`
|
||||
|
||||
| 值 | 说明 |
|
||||
| --- | -------- |
|
||||
| `0` | 用户自建 |
|
||||
| `1` | 全部笔记 |
|
||||
| `2` | 未分类 |
|
||||
|
||||
---
|
||||
|
||||
## 游标翻页使用规范
|
||||
|
||||
1. **首次请求**:`cursor` 传空字符串 `""`
|
||||
2. 检查返回的 `is_end`:`false` 表示还有更多数据
|
||||
3. 将返回的 `next_cursor` 作为下次请求的 `cursor`
|
||||
4. `is_end = true` 时停止翻页
|
||||
|
||||
---
|
||||
|
||||
## 错误码
|
||||
|
||||
| 错误码 | 说明 |
|
||||
| ------ | -------------------------------------------- |
|
||||
| 0 | 成功 |
|
||||
| 100001 | 参数错误 |
|
||||
| 100002 | 携带无效的 ID |
|
||||
| 100003 | 服务器内部错误 |
|
||||
| 100004 | 拉取的 size 不合法(超出范围)/ 用户空间不够 |
|
||||
| 100005 | 不能获取私有笔记的访客信息 / 不是笔记的作者 |
|
||||
| 100006 | 笔记已被删除 |
|
||||
| 100008 | 版本冲突 |
|
||||
| 100009 | 单篇笔记超过最大限制 |
|
||||
| 310001 | 笔记本不存在 |
|
||||
| 20002 | apiKey超过最大限频 |
|
||||
| 20004 | apikey鉴权失败 |
|
||||
@@ -0,0 +1,177 @@
|
||||
# Notes (笔记)
|
||||
|
||||
> Prerequisites: see root `../SKILL.md` for setup, credentials, and `ima_api()` helper.
|
||||
|
||||
API base path: `openapi/note/v1`
|
||||
|
||||
通过 IMA OpenAPI 管理用户个人笔记,支持读取(搜索、列表、获取内容)和写入(新建、追加)。
|
||||
|
||||
完整的数据结构和接口参数详见 `notes-api.md`。
|
||||
|
||||
> **隐私规则:** 笔记内容属于用户隐私,在群聊场景中只展示标题和摘要,禁止展示笔记正文。
|
||||
|
||||
## 接口决策表
|
||||
|
||||
| 用户意图 | 调用接口 | 关键参数 |
|
||||
| --------------------------------------------------------------------------------------------------------- | ---------------------------- | ----------------------------------------------------------------------------- |
|
||||
| 搜索/查找笔记 | `search_note_book` | `query_info`(QueryInfo 对象) |
|
||||
| 查看笔记本列表 | `list_note_folder_by_cursor` | `cursor`(必填,首页传`"0"`) + `limit`(必填) |
|
||||
| 浏览某笔记本里的笔记,当用户表述"最新"、"最近"之类的通用限定,没有指明笔记本时,都应该直接在全部笔记里去拉 | `list_note_by_folder_id` | `folder_id`(选填,空为全部笔记本) + `cursor`(必填,首次传`""`) + `limit`(必填) |
|
||||
| 读取笔记正文 | `get_doc_content` | `doc_id` + `target_content_format`(必填,推荐`0`纯文本) |
|
||||
| 新建一篇笔记(用户明确说"新建/创建笔记"时走此接口) | `import_doc` | `content` + `content_format`(必填,固定`1`) + 可选 `folder_id` |
|
||||
| 往已有笔记追加内容(⚠️ **敏感操作**:用户必须明确指定目标笔记,否则先确认再操作) | `append_doc` | `doc_id` + `content` + `content_format`(必填,固定`1`) |
|
||||
|
||||
## ⚠️ 新建 vs. 追加 — 行为规则
|
||||
|
||||
**新建笔记(`import_doc`)** 和 **追加内容到已有笔记(`append_doc`)** 是两个完全不同的操作,务必正确区分:
|
||||
|
||||
### 明确走新建的信号词
|
||||
|
||||
用户说以下任一表述时,**直接调用 `import_doc` 创建新笔记**:
|
||||
|
||||
- "**新建**笔记"、"**创建**笔记"、"**写一篇**笔记"
|
||||
- "**新建**一篇笔记记录这些内容"
|
||||
|
||||
### 明确走追加的信号词
|
||||
|
||||
用户说以下任一表述时,**调用 `append_doc` 追加到已有笔记**(但仍需确认目标笔记,见下方规则):
|
||||
|
||||
- "把这段话**追加到**《XX》笔记里"
|
||||
- "在那篇笔记**末尾加上**这段内容"
|
||||
|
||||
### 模糊场景 — 必须先询问用户
|
||||
|
||||
以下表述**既可能是新建、也可能是追加**,agent **不得自行假设**,必须先向用户确认:
|
||||
|
||||
- "帮我记一下"、"记录一下"、"保存为笔记"、"存成笔记"
|
||||
- "把这段内容记到笔记里"
|
||||
- "添加到笔记里"
|
||||
- 任何其他未明确表达"新建"或"追加"意图的表述
|
||||
|
||||
询问示例:
|
||||
> "您是想**创建一篇新笔记**,还是**追加到某篇已有笔记**?"
|
||||
|
||||
### 追加到已有笔记是敏感操作
|
||||
|
||||
`append_doc` 会**不可撤销地修改**用户的现有笔记,因此必须谨慎处理:
|
||||
|
||||
1. **用户明确指定了目标笔记** — 可以直接追加。例如:
|
||||
- "把这段话追加到《会议纪要》笔记里"
|
||||
- "在那篇笔记末尾加上这段内容"(上下文中已有明确的笔记对象)
|
||||
|
||||
2. **用户没有明确指定目标笔记** — **必须先向用户确认**,不要自行猜测。例如:
|
||||
- 用户说"添加到笔记里" → 询问:"您想追加到哪篇已有笔记?请提供笔记标题或让我帮您搜索。"
|
||||
- 用户说"把这个加到之前那篇笔记" → 如果上下文中有多篇笔记或不确定是哪篇 → 列出候选笔记让用户选择
|
||||
|
||||
> **原则**:不确定时,先问。宁可多问一句,也不要误改用户的已有笔记或自作主张创建新笔记。
|
||||
|
||||
### 🖼️ 本地图片不支持
|
||||
|
||||
`import_doc` 和 `append_doc` 的 `content` 字段仅支持纯文本/Markdown,**不支持本地图片**。
|
||||
|
||||
写入笔记内容前,必须检查并处理图片引用:
|
||||
|
||||
1. **过滤本地图片** — 如果用户提供的内容中包含本地图片路径(如 ``, ``, `` 等),**移除这些图片引用**,不要将其写入笔记。
|
||||
2. **告知用户** — 移除后主动提醒用户:
|
||||
> "笔记接口暂不支持上传本地图片,以下图片已被过滤:`xxx.png`、`yyy.jpg`。您可以先将图片上传到网络,再用网络链接插入笔记。"
|
||||
3. **保留网络图片** — 以 `http://` 或 `https://` 开头的图片链接可以正常保留。
|
||||
|
||||
## 常用工作流
|
||||
|
||||
### 查找并阅读笔记
|
||||
|
||||
先搜索获取 `docid`,再用 `get_doc_content` 读取正文:
|
||||
|
||||
```bash
|
||||
# 1. 按标题搜索
|
||||
ima_api "openapi/note/v1/search_note_book" '{"search_type": 0, "query_info": {"title": "会议纪要"}, "start": 0, "end": 20}'
|
||||
# 从返回的 docs[].doc.basic_info.docid 中取目标笔记 ID
|
||||
|
||||
# 2. 读取正文(纯文本格式,Markdown 格式目前不支持)
|
||||
ima_api "openapi/note/v1/get_doc_content" '{"doc_id": "目标docid", "target_content_format": 0}'
|
||||
```
|
||||
|
||||
### 浏览笔记本里的笔记
|
||||
|
||||
先拉笔记本列表获取 `folder_id`,再拉该笔记本下的笔记:
|
||||
|
||||
```bash
|
||||
# 1. 列出笔记本(首页 cursor 传 "0")
|
||||
ima_api "openapi/note/v1/list_note_folder_by_cursor" '{"cursor": "0", "limit": 20}'
|
||||
|
||||
# 2. 拉取指定笔记本的笔记(首页 cursor 传 "")
|
||||
ima_api "openapi/note/v1/list_note_by_folder_id" '{"folder_id": "user_list_xxx", "cursor": "", "limit": 20}'
|
||||
```
|
||||
|
||||
### 新建笔记
|
||||
|
||||
```bash
|
||||
# 新建到默认位置
|
||||
ima_api "openapi/note/v1/import_doc" '{"content_format": 1, "content": "# 标题\n\n正文内容"}'
|
||||
|
||||
# 新建到指定笔记本
|
||||
ima_api "openapi/note/v1/import_doc" '{"content_format": 1, "content": "# 标题\n\n正文内容", "folder_id": "笔记本ID"}'
|
||||
# 返回 doc_id,后续可用于 append_doc
|
||||
```
|
||||
|
||||
### 追加内容到已有笔记
|
||||
|
||||
```bash
|
||||
ima_api "openapi/note/v1/append_doc" '{"doc_id": "笔记ID", "content_format": 1, "content": "\n## 补充内容\n\n追加的文本"}'
|
||||
```
|
||||
|
||||
### 按正文搜索
|
||||
|
||||
```bash
|
||||
ima_api "openapi/note/v1/search_note_book" '{"search_type": 1, "query_info": {"content": "项目排期"}, "start": 0, "end": 20}'
|
||||
```
|
||||
|
||||
## 核心响应字段
|
||||
|
||||
**搜索结果**(`SearchedDoc`):笔记信息路径为 `doc.basic_info`(DocBasic),关键字段:`docid`、`title`、`summary`、`folder_id`、`folder_name`、`create_time`(Unix 毫秒)、`modify_time`、`status`。额外包含 `highlight_info`(高亮匹配,key 为 `doc_title`,value 含 `<em>高亮词</em>`)。
|
||||
|
||||
**笔记本条目**(`NoteBookFolder`):信息路径为 `folder.basic_info`(NoteBookFolderBasic),关键字段:`folder_id`、`name`、`note_number`、`create_time`、`modify_time`、`folder_type`(`0`=用户自建,`1`=全部笔记,`2`=未分类)、`status`。
|
||||
|
||||
**笔记列表条目**(`NoteBookInfo`):信息路径为 `basic_info.basic_info`(DocBasicInfo → DocBasic),关键字段:`docid`、`title`、`summary`、`folder_id`、`folder_name`、`create_time`、`modify_time`、`status`。
|
||||
|
||||
**写入结果**(`import_doc`/`append_doc`):返回 `doc_id`(新建或目标笔记的唯一 ID)。
|
||||
|
||||
完整字段定义见 `notes-api.md`。
|
||||
|
||||
## 分页
|
||||
|
||||
- **游标分页 — 笔记本列表**(`list_note_folder_by_cursor`):首次 `cursor: "0"`,后续用 `next_cursor`,`is_end=true` 时停止。
|
||||
- **游标分页 — 笔记列表**(`list_note_by_folder_id`):首次 `cursor: ""`,后续用 `next_cursor`,`is_end=true` 时停止。
|
||||
- **偏移量分页**(`search_note_book`):首次 `start: 0, end: 20`,翻页时递增,`is_end=true` 时停止。
|
||||
|
||||
## 枚举值
|
||||
|
||||
- **`content_format`:** `0`=纯文本,`1`=Markdown,`2`=JSON。写入(`import_doc`/`append_doc`)目前仅支持 `1`(Markdown)。读取(`get_doc_content`)推荐 `0`(纯文本),Markdown 格式不支持。
|
||||
- **`search_type`:** `0`=标题检索(默认),`1`=正文检索
|
||||
- **`sort_type`:** `0`=更新时间(默认),`1`=创建时间,`2`=标题,`3`=大小(仅 `search_note_book` 使用)
|
||||
- **`folder_type`:** `0`=用户自建,`1`=全部笔记(根目录),`2`=未分类
|
||||
|
||||
## 注意事项
|
||||
|
||||
- `folder_id` 不可为 `"0"`,根目录 ID 格式为 `user_list_{userid}`(从 `folder_type=1` 的笔记本条目获取)
|
||||
- 笔记内容有大小上限,超过时返回 `100009`,可拆分为多次 `append_doc` 写入
|
||||
- 写入内容不支持本地图片,写入前必须过滤本地图片路径并告知用户(详见"🖼️ 本地图片不支持"规则)
|
||||
- 展示笔记列表时只展示标题、摘要和修改时间,不要主动展示正文
|
||||
- 时间字段是 Unix 毫秒时间戳,展示时转为可读格式
|
||||
- 返回数据为嵌套结构:搜索结果取 `docs[].doc.basic_info.docid`,笔记本取 `note_book_folders[].folder.basic_info.folder_id`,笔记列表取 `note_book_list[].basic_info.basic_info.docid`,注意按层级解析
|
||||
|
||||
## 错误处理
|
||||
|
||||
| 错误码 | 含义 | 建议处理 |
|
||||
| ------ | ---------------------- | ---------------------------- |
|
||||
| 100001 | 参数错误 | 检查请求参数格式和必填字段 |
|
||||
| 100002 | 无效 ID | 检查凭证配置 |
|
||||
| 100003 | 服务器内部错误 | 等待后重试 |
|
||||
| 100004 | size 不合法 / 空间不够 | 检查参数范围 |
|
||||
| 100005 | 无权限 | 确认操作的是用户自己的笔记 |
|
||||
| 100006 | 笔记已删除 | 告知用户该笔记不存在 |
|
||||
| 100008 | 版本冲突 | 重新获取内容后再操作 |
|
||||
| 100009 | 超过大小限制 | 拆分为多次 `append_doc` 写入 |
|
||||
| 310001 | 笔记本不存在 | 检查 `folder_id` 是否正确 |
|
||||
| 20002 | apiKey超过最大限频 |
|
||||
| 20004 | apikey 鉴权失败 | 检查凭证配置是否正确 |
|
||||
@@ -0,0 +1,161 @@
|
||||
#!/usr/bin/env node
|
||||
'use strict';
|
||||
|
||||
const crypto = require('node:crypto');
|
||||
const fs = require('node:fs');
|
||||
const https = require('node:https');
|
||||
|
||||
// --- Argument parsing ---
|
||||
function parseArgs(argv) {
|
||||
const args = {};
|
||||
for (let i = 2; i < argv.length; i += 2) {
|
||||
const key = argv[i].replace(/^--/, '');
|
||||
const val = argv[i + 1];
|
||||
if (!val || val.startsWith('--')) {
|
||||
console.error(`Missing value for --${key}`);
|
||||
process.exit(1);
|
||||
}
|
||||
args[key] = val;
|
||||
}
|
||||
return args;
|
||||
}
|
||||
|
||||
const REQUIRED = ['file', 'secret-id', 'secret-key', 'token', 'bucket', 'region', 'cos-key'];
|
||||
|
||||
// --- Crypto helpers ---
|
||||
function hmacSha1(key, data) {
|
||||
return crypto.createHmac('sha1', key).update(data).digest('hex');
|
||||
}
|
||||
|
||||
function sha1(data) {
|
||||
return crypto.createHash('sha1').update(data).digest('hex');
|
||||
}
|
||||
|
||||
// --- COS Authorization header (PUT Object) ---
|
||||
// Reference: https://cloud.tencent.com/document/product/436/7778
|
||||
function buildAuthorization({ secretId, secretKey, method, pathname, headers, startTime, expiredTime }) {
|
||||
const keyTime = `${startTime};${expiredTime}`;
|
||||
|
||||
// 1. SignKey = HMAC-SHA1(SecretKey, KeyTime)
|
||||
const signKey = hmacSha1(secretKey, keyTime);
|
||||
|
||||
// 2. HttpString = method\npathname\nparams\nheaders\n
|
||||
// For PUT, no query params; headers we sign: host, content-length
|
||||
const headerKeys = Object.keys(headers).sort();
|
||||
const httpHeaders = headerKeys.map((k) => `${k.toLowerCase()}=${encodeURIComponent(headers[k])}`).join('&');
|
||||
const httpString = `${method.toLowerCase()}\n${pathname}\n\n${httpHeaders}\n`;
|
||||
|
||||
// 3. StringToSign = sha1\nKeyTime\nSHA1(HttpString)\n
|
||||
const stringToSign = `sha1\n${keyTime}\n${sha1(httpString)}\n`;
|
||||
|
||||
// 4. Signature = HMAC-SHA1(SignKey, StringToSign)
|
||||
const signature = hmacSha1(signKey, stringToSign);
|
||||
|
||||
// 5. Build Authorization
|
||||
const headerList = headerKeys.map((k) => k.toLowerCase()).join(';');
|
||||
return [
|
||||
`q-sign-algorithm=sha1`,
|
||||
`q-ak=${secretId}`,
|
||||
`q-sign-time=${keyTime}`,
|
||||
`q-key-time=${keyTime}`,
|
||||
`q-header-list=${headerList}`,
|
||||
`q-url-param-list=`,
|
||||
`q-signature=${signature}`,
|
||||
].join('&');
|
||||
}
|
||||
|
||||
// --- Upload via PUT Object ---
|
||||
function upload(args) {
|
||||
const secretId = args['secret-id'];
|
||||
const secretKey = args['secret-key'];
|
||||
const { token } = args;
|
||||
const { bucket } = args;
|
||||
const { region } = args;
|
||||
const cosKey = args['cos-key'];
|
||||
const filePath = args.file;
|
||||
|
||||
const startTime = args['start-time'] || String(Math.floor(Date.now() / 1000));
|
||||
const expiredTime = args['expired-time'] || String(Math.floor(Date.now() / 1000) + 3600);
|
||||
|
||||
const fileContent = fs.readFileSync(filePath);
|
||||
const hostname = `${bucket}.cos.${region}.myqcloud.com`;
|
||||
const pathname = `/${cosKey}`;
|
||||
|
||||
// Headers to sign
|
||||
const signHeaders = {
|
||||
'content-length': String(fileContent.length),
|
||||
host: hostname,
|
||||
};
|
||||
|
||||
const authorization = buildAuthorization({
|
||||
secretId,
|
||||
secretKey,
|
||||
method: 'PUT',
|
||||
pathname,
|
||||
headers: signHeaders,
|
||||
startTime,
|
||||
expiredTime,
|
||||
});
|
||||
|
||||
// Use the actual file content type if provided, otherwise fall back to octet-stream
|
||||
const contentType = args['content-type'] || 'application/octet-stream';
|
||||
|
||||
const options = {
|
||||
hostname,
|
||||
port: 443,
|
||||
path: pathname,
|
||||
method: 'PUT',
|
||||
headers: {
|
||||
'Content-Type': contentType,
|
||||
'Content-Length': fileContent.length,
|
||||
Authorization: authorization,
|
||||
'x-cos-security-token': token,
|
||||
},
|
||||
};
|
||||
|
||||
const req = https.request(options, (res) => {
|
||||
let body = '';
|
||||
res.on('data', (chunk) => (body += chunk));
|
||||
res.on('end', () => {
|
||||
if (res.statusCode >= 200 && res.statusCode < 300) {
|
||||
console.log(`Upload successful (HTTP ${res.statusCode})`);
|
||||
process.exit(0);
|
||||
} else {
|
||||
console.error(`COS upload failed (HTTP ${res.statusCode}): ${body}`);
|
||||
process.exit(1);
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
req.on('error', (err) => {
|
||||
console.error(`COS upload error: ${err.message}`);
|
||||
process.exit(1);
|
||||
});
|
||||
|
||||
req.write(fileContent);
|
||||
req.end();
|
||||
}
|
||||
|
||||
// --- Main ---
|
||||
function main() {
|
||||
const args = parseArgs(process.argv);
|
||||
|
||||
const missing = REQUIRED.filter((k) => !args[k]);
|
||||
if (missing.length) {
|
||||
console.error(`Missing required arguments: ${missing.map((k) => `--${k}`).join(', ')}`);
|
||||
console.error(
|
||||
`Usage: node cos-upload.cjs --file <path> --secret-id <sid> --secret-key <skey> --token <token> --bucket <bucket> --region <region> --cos-key <key> [--content-type <mime>] [--start-time <ts>] [--expired-time <ts>]`,
|
||||
);
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
const filePath = args.file;
|
||||
if (!fs.existsSync(filePath)) {
|
||||
console.error(`File not found: ${filePath}`);
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
upload(args);
|
||||
}
|
||||
|
||||
main();
|
||||
@@ -0,0 +1,252 @@
|
||||
#!/usr/bin/env node
|
||||
'use strict';
|
||||
|
||||
/**
|
||||
* Preflight check for uploading a file to IMA Knowledge Base.
|
||||
*
|
||||
* Validates file type, size, and extracts all metadata needed for
|
||||
* create_media and add_knowledge API calls.
|
||||
*
|
||||
* Resolution priority:
|
||||
* 1. If --content-type is provided and recognized → use it (content-type rules over extension)
|
||||
* 2. If --content-type is unrecognized, fall back to extension
|
||||
* 3. If no --content-type, use extension
|
||||
* 4. If neither can resolve → fail
|
||||
*
|
||||
* Usage:
|
||||
* node preflight-check.cjs --file /path/to/report.pdf
|
||||
* node preflight-check.cjs --file /path/to/downloaded_file --content-type application/pdf
|
||||
*
|
||||
* Output (JSON, always to stdout):
|
||||
*
|
||||
* Pass:
|
||||
* {
|
||||
* "pass": true,
|
||||
* "file_path": "/absolute/path/to/report.pdf",
|
||||
* "file_name": "report.pdf",
|
||||
* "file_ext": "pdf",
|
||||
* "file_size": 123456,
|
||||
* "media_type": 1,
|
||||
* "content_type": "application/pdf"
|
||||
* }
|
||||
*
|
||||
* Pass (no extension, content_type provided):
|
||||
* {
|
||||
* "pass": true,
|
||||
* "file_path": "/absolute/path/to/downloaded_file",
|
||||
* "file_name": "downloaded_file",
|
||||
* "file_ext": "",
|
||||
* "file_size": 123456,
|
||||
* "media_type": 1,
|
||||
* "content_type": "application/pdf"
|
||||
* }
|
||||
*
|
||||
* Fail:
|
||||
* {
|
||||
* "pass": false,
|
||||
* "file_path": "/absolute/path/to/video.mp4",
|
||||
* "file_name": "video.mp4",
|
||||
* "file_ext": "mp4",
|
||||
* "reason": "Video files (.mp4) are not supported. ..."
|
||||
* }
|
||||
*
|
||||
* Exit codes:
|
||||
* 0 = pass — file is ready for upload
|
||||
* 1 = fail — file rejected (unsupported type, over size limit, etc.)
|
||||
* 2 = error — file not found, usage error, etc.
|
||||
*/
|
||||
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
// ─── Extension → media_type + content_type ──────────────────────────────────
|
||||
|
||||
const EXT_MAP = {
|
||||
pdf: { media_type: 1, content_type: 'application/pdf' },
|
||||
doc: { media_type: 3, content_type: 'application/msword' },
|
||||
docx: { media_type: 3, content_type: 'application/vnd.openxmlformats-officedocument.wordprocessingml.document' },
|
||||
ppt: { media_type: 4, content_type: 'application/vnd.ms-powerpoint' },
|
||||
pptx: { media_type: 4, content_type: 'application/vnd.openxmlformats-officedocument.presentationml.presentation' },
|
||||
xls: { media_type: 5, content_type: 'application/vnd.ms-excel' },
|
||||
xlsx: { media_type: 5, content_type: 'application/vnd.openxmlformats-officedocument.spreadsheetml.sheet' },
|
||||
csv: { media_type: 5, content_type: 'text/csv' },
|
||||
md: { media_type: 7, content_type: 'text/markdown' },
|
||||
markdown: { media_type: 7, content_type: 'text/markdown' },
|
||||
png: { media_type: 9, content_type: 'image/png' },
|
||||
jpg: { media_type: 9, content_type: 'image/jpeg' },
|
||||
jpeg: { media_type: 9, content_type: 'image/jpeg' },
|
||||
webp: { media_type: 9, content_type: 'image/webp' },
|
||||
txt: { media_type: 13, content_type: 'text/plain' },
|
||||
xmind: { media_type: 14, content_type: 'application/x-xmind' },
|
||||
mp3: { media_type: 15, content_type: 'audio/mpeg' },
|
||||
m4a: { media_type: 15, content_type: 'audio/x-m4a' },
|
||||
wav: { media_type: 15, content_type: 'audio/wav' },
|
||||
aac: { media_type: 15, content_type: 'audio/aac' },
|
||||
};
|
||||
|
||||
// ─── Content-Type → media_type (reverse lookup) ─────────────────────────────
|
||||
|
||||
const CONTENT_TYPE_MAP = {};
|
||||
for (const [, value] of Object.entries(EXT_MAP)) {
|
||||
// First entry wins — keeps the canonical content_type per media_type
|
||||
if (!CONTENT_TYPE_MAP[value.content_type]) {
|
||||
CONTENT_TYPE_MAP[value.content_type] = value.media_type;
|
||||
}
|
||||
}
|
||||
// Extra aliases not covered by EXT_MAP
|
||||
Object.assign(CONTENT_TYPE_MAP, {
|
||||
'text/x-markdown': 7,
|
||||
'application/md': 7,
|
||||
'application/markdown': 7,
|
||||
'application/vnd.xmind.workbook': 14,
|
||||
'application/zip': 14, // xmind can be zip
|
||||
});
|
||||
|
||||
// ─── Size limits by media_type (bytes) ──────────────────────────────────────
|
||||
|
||||
const MB = 1024 * 1024;
|
||||
const SIZE_LIMITS = {
|
||||
5: 10 * MB, // Excel / CSV
|
||||
7: 10 * MB, // Markdown
|
||||
13: 10 * MB, // TXT
|
||||
14: 10 * MB, // Xmind
|
||||
9: 30 * MB, // Image
|
||||
};
|
||||
const DEFAULT_SIZE_LIMIT = 200 * MB; // PDF, Word, PPT, Audio, etc.
|
||||
|
||||
// ─── Explicitly unsupported extensions ──────────────────────────────────────
|
||||
|
||||
const UNSUPPORTED_VIDEO_EXT = new Set(['mp4', 'avi', 'mov', 'mkv', 'wmv', 'flv', 'webm', 'm4v', 'rmvb', 'rm', '3gp']);
|
||||
|
||||
const UNSUPPORTED_VIDEO_CT = new Set([
|
||||
'video/mp4',
|
||||
'video/x-msvideo',
|
||||
'video/quicktime',
|
||||
'video/x-matroska',
|
||||
'video/x-ms-wmv',
|
||||
'video/x-flv',
|
||||
'video/webm',
|
||||
]);
|
||||
|
||||
// ─── Helpers ────────────────────────────────────────────────────────────────
|
||||
|
||||
function fail(result) {
|
||||
console.log(JSON.stringify({ pass: false, ...result }));
|
||||
process.exit(1);
|
||||
}
|
||||
|
||||
function formatSize(bytes) {
|
||||
if (bytes < MB) return `${(bytes / 1024).toFixed(1)} KB`;
|
||||
return `${(bytes / MB).toFixed(1)} MB`;
|
||||
}
|
||||
|
||||
// ─── Argument parsing ───────────────────────────────────────────────────────
|
||||
|
||||
function parseArgs(argv) {
|
||||
const args = {};
|
||||
for (let i = 2; i < argv.length; i++) {
|
||||
if (argv[i].startsWith('--') && i + 1 < argv.length) {
|
||||
args[argv[i].replace(/^--/, '')] = argv[i + 1];
|
||||
i += 1;
|
||||
}
|
||||
}
|
||||
return args;
|
||||
}
|
||||
|
||||
// ─── Main ───────────────────────────────────────────────────────────────────
|
||||
|
||||
const args = parseArgs(process.argv);
|
||||
|
||||
if (!args.file) {
|
||||
console.error('Usage: node preflight-check.cjs --file <path> [--content-type <mime>]');
|
||||
process.exit(2);
|
||||
}
|
||||
|
||||
const filePath = path.resolve(args.file);
|
||||
const fileName = path.basename(filePath);
|
||||
const extMatch = fileName.match(/\.([^.]+)$/);
|
||||
const ext = extMatch ? extMatch[1].toLowerCase() : '';
|
||||
const inputContentType = args['content-type'] || '';
|
||||
|
||||
const base = { file_path: filePath, file_name: fileName, file_ext: ext };
|
||||
|
||||
// 1. Check file exists
|
||||
let stat;
|
||||
try {
|
||||
stat = fs.statSync(filePath);
|
||||
} catch (err) {
|
||||
if (err.code === 'ENOENT') {
|
||||
console.error(`File not found: ${filePath}`);
|
||||
process.exit(2);
|
||||
}
|
||||
throw err;
|
||||
}
|
||||
|
||||
// 2. Check not an unsupported video type (by ext or content-type)
|
||||
if (UNSUPPORTED_VIDEO_EXT.has(ext)) {
|
||||
fail({ ...base, reason: `Video files (.${ext}) are not supported. Only supported in IMA desktop app.` });
|
||||
}
|
||||
if (UNSUPPORTED_VIDEO_CT.has(inputContentType)) {
|
||||
fail({ ...base, reason: `Video files (${inputContentType}) are not supported. Only supported in IMA desktop app.` });
|
||||
}
|
||||
|
||||
// 3. Resolve media_type and content_type
|
||||
// Priority: content-type first, then fall back to extension
|
||||
let mediaType = null;
|
||||
let contentType = null;
|
||||
|
||||
const ctMediaType = inputContentType ? CONTENT_TYPE_MAP[inputContentType] : undefined;
|
||||
const extMapping = ext ? EXT_MAP[ext] : undefined;
|
||||
|
||||
if (ctMediaType != null) {
|
||||
// Content-type recognized — always wins
|
||||
mediaType = ctMediaType;
|
||||
contentType = inputContentType;
|
||||
} else if (inputContentType) {
|
||||
// Content-type provided but unrecognized — try extension fallback
|
||||
if (extMapping) {
|
||||
mediaType = extMapping.media_type;
|
||||
contentType = extMapping.content_type;
|
||||
} else {
|
||||
fail({
|
||||
...base,
|
||||
reason: `Unrecognized content type ${inputContentType}${ext ? ` and file extension .${ext}` : ''}. This file type is not supported.`,
|
||||
});
|
||||
}
|
||||
} else {
|
||||
// No content-type provided — fall back to extension
|
||||
if (extMapping) {
|
||||
mediaType = extMapping.media_type;
|
||||
contentType = extMapping.content_type;
|
||||
} else if (ext) {
|
||||
fail({ ...base, reason: `Unrecognized file extension .${ext}. This file type is not supported.` });
|
||||
} else {
|
||||
fail({ ...base, reason: 'File has no extension and no --content-type provided. Cannot determine file type.' });
|
||||
}
|
||||
}
|
||||
|
||||
// 4. Check file size
|
||||
const fileSize = stat.size;
|
||||
const sizeLimit = SIZE_LIMITS[mediaType] || DEFAULT_SIZE_LIMIT;
|
||||
|
||||
if (fileSize > sizeLimit) {
|
||||
fail({
|
||||
...base,
|
||||
file_size: fileSize,
|
||||
media_type: mediaType,
|
||||
content_type: contentType,
|
||||
reason: `File size ${formatSize(fileSize)} exceeds the ${formatSize(sizeLimit)} limit for this file type.`,
|
||||
});
|
||||
}
|
||||
|
||||
// 5. All checks passed
|
||||
console.log(
|
||||
JSON.stringify({
|
||||
pass: true,
|
||||
...base,
|
||||
file_size: fileSize,
|
||||
media_type: mediaType,
|
||||
content_type: contentType,
|
||||
}),
|
||||
);
|
||||
process.exit(0);
|
||||
@@ -0,0 +1,102 @@
|
||||
---
|
||||
name: openmaic
|
||||
description: Guided SOP for setting up and using OpenMAIC from OpenClaw. Use when the user wants to clone the OpenMAIC repo, choose a startup mode, configure recommended API keys, start the service, or generate a classroom from requirements or a PDF. Run one phase at a time and ask for confirmation before each state-changing step.
|
||||
user-invocable: true
|
||||
metadata: { "openclaw": { "emoji": "🏫" } }
|
||||
---
|
||||
|
||||
# OpenMAIC Skill
|
||||
|
||||
Use this as a guided, confirmation-heavy SOP. Do not compress the whole setup into one reply and do not perform state-changing actions without explicit user confirmation.
|
||||
|
||||
## Core Rules
|
||||
|
||||
- Move one phase at a time.
|
||||
- Before any state-changing action, ask for confirmation.
|
||||
- If local state already exists, show what you found and ask whether to keep it.
|
||||
- Do not assume the OpenClaw agent's own model or API key will be reused by OpenMAIC.
|
||||
- OpenMAIC classroom generation uses OpenMAIC server-side provider config.
|
||||
- This skill must not rely on any request-time model or provider overrides.
|
||||
- Only OpenMAIC server-side config files may control provider selection and defaults.
|
||||
- Do not default to asking the user to paste API keys into chat.
|
||||
- Prefer guiding the user to edit local config files themselves.
|
||||
- Do not offer to write API keys into config files on the user's behalf.
|
||||
- Once setup is complete and the user clearly asks to generate a classroom, do not ask for a second confirmation before submitting the generation job.
|
||||
- Keep confirmations for local file reads such as reading a PDF from disk.
|
||||
|
||||
## Optional Skill Config
|
||||
|
||||
If present, read defaults from `~/.openclaw/openclaw.json` under:
|
||||
|
||||
```jsonc
|
||||
{
|
||||
"skills": {
|
||||
"entries": {
|
||||
"openmaic": {
|
||||
"enabled": true,
|
||||
"config": {
|
||||
"accessCode": "sk-xxx",
|
||||
"repoDir": "/path/to/OpenMAIC",
|
||||
"url": "http://localhost:3000"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
- If `accessCode` is present, default to hosted mode and skip the mode-selection prompt.
|
||||
- Use `repoDir` and `url` only as defaults for local mode.
|
||||
- Still confirm before acting.
|
||||
|
||||
## SOP Phases
|
||||
|
||||
### 0. Choose Mode
|
||||
|
||||
First check skill config for `accessCode`. If present, announce that a stored access code was found and proceed directly to hosted mode (load [references/hosted-mode.md](references/hosted-mode.md), skip phases 1–4). Do not ask the user to paste the code again.
|
||||
|
||||
If no `accessCode` in config, ask the user how they want to use OpenMAIC:
|
||||
|
||||
1. **Use hosted OpenMAIC** (recommended for quick start) — Requires an access code from open.maic.chat. No local setup needed.
|
||||
2. **Run locally** — Clone the repo, configure provider keys, and run on your machine.
|
||||
|
||||
If the user chooses hosted mode, load [references/hosted-mode.md](references/hosted-mode.md) and skip phases 1–4.
|
||||
If the user chooses local mode, proceed to phase 1 as usual.
|
||||
|
||||
### 1. Clone Or Reuse Existing Repo
|
||||
|
||||
Load [references/clone.md](references/clone.md).
|
||||
|
||||
Use this when the user has not installed OpenMAIC yet or when you need to confirm which local checkout to use.
|
||||
|
||||
### 2. Choose Startup Mode
|
||||
|
||||
Load [references/startup-modes.md](references/startup-modes.md).
|
||||
|
||||
Use this after the repo location is confirmed. Present the available startup modes, recommend one, and wait for the user's choice.
|
||||
|
||||
### 3. Configure Provider Keys
|
||||
|
||||
Load [references/provider-keys.md](references/provider-keys.md).
|
||||
|
||||
Use this before starting classroom generation. Recommend a provider path and tell the user exactly which config file to edit themselves. If generation later fails due to provider/model/auth issues, return to this phase and direct the user to update the same server-side config files.
|
||||
|
||||
After the core LLM key is configured, ask the user if they want to enable optional features (web search, image generation, video generation, TTS). Each requires its own provider key — see the "Optional Features" section in provider-keys.md.
|
||||
|
||||
### 4. Start And Verify OpenMAIC
|
||||
|
||||
After the user has chosen a startup mode and configured keys, start OpenMAIC using the chosen method, then verify the service with `GET {url}/api/health`.
|
||||
|
||||
### 5. Generate A Classroom
|
||||
|
||||
Load [references/generate-flow.md](references/generate-flow.md).
|
||||
|
||||
Use this only after the service is healthy. Confirm before reading local PDFs. If the user has already clearly asked to generate, do not ask for a second confirmation before submitting the generation job, and then follow the polling loop until it succeeds or fails. Only send the supported content fields for generation requests. For long-running jobs, prefer sparse polling and tell the user to check back later if the turn ends before completion.
|
||||
|
||||
## Response Style
|
||||
|
||||
- Keep each step short and explicit.
|
||||
- Prefer 2-3 concrete options when the user must choose.
|
||||
- Always include the recommended option first and explain why in one sentence.
|
||||
- After a step completes, say what changed and what the next confirmation is for.
|
||||
- When returning a classroom link, place the raw absolute URL on its own line with no bold, markdown link syntax, code formatting, or tables.
|
||||
@@ -0,0 +1,38 @@
|
||||
# Clone Or Reuse Existing Repo
|
||||
|
||||
## Goal
|
||||
|
||||
Establish which OpenMAIC checkout will be used for setup and runtime actions.
|
||||
|
||||
## Procedure
|
||||
|
||||
1. Check whether OpenMAIC already exists locally.
|
||||
2. If a checkout exists, show the path and ask whether to reuse it.
|
||||
3. If no checkout exists, propose cloning the repo and ask for confirmation.
|
||||
4. After clone, confirm dependency installation separately.
|
||||
|
||||
## Recommended Path
|
||||
|
||||
- Recommended: reuse an existing checkout if it is already on the target branch.
|
||||
- Otherwise: clone a fresh checkout from GitHub, then install dependencies.
|
||||
|
||||
## Commands
|
||||
|
||||
Clone:
|
||||
|
||||
```bash
|
||||
git clone https://github.com/THU-MAIC/OpenMAIC.git
|
||||
cd OpenMAIC
|
||||
```
|
||||
|
||||
Install dependencies:
|
||||
|
||||
```bash
|
||||
pnpm install
|
||||
```
|
||||
|
||||
## Confirmation Requirements
|
||||
|
||||
- Ask before `git clone`.
|
||||
- Ask before `pnpm install`.
|
||||
- If the repo is dirty, tell the user and ask whether to continue with that checkout.
|
||||
@@ -0,0 +1,170 @@
|
||||
# Generate Flow
|
||||
|
||||
## Preconditions
|
||||
|
||||
- Repo path is confirmed
|
||||
- Startup mode has been chosen
|
||||
- OpenMAIC is healthy at the selected `url`
|
||||
- Provider keys are configured
|
||||
|
||||
> **Hosted mode**: If using hosted OpenMAIC (open.maic.chat), all
|
||||
> preconditions (repo, startup, provider keys) are already satisfied.
|
||||
> Include `Authorization: Bearer <access-code>` header on all requests below.
|
||||
> See [hosted-mode.md](hosted-mode.md) for details.
|
||||
|
||||
## Requirement-Only Generation
|
||||
|
||||
If the user has already clearly asked to generate the classroom and the preconditions are satisfied, submit the generation job immediately. Do not ask for a second confirmation just before calling `/api/generate-classroom`.
|
||||
|
||||
Submit the job with:
|
||||
|
||||
```text
|
||||
POST {url}/api/generate-classroom
|
||||
```
|
||||
|
||||
Request body:
|
||||
|
||||
```json
|
||||
{
|
||||
"requirement": "Create an introductory classroom on quantum mechanics for high school students"
|
||||
}
|
||||
```
|
||||
|
||||
Only send supported content fields:
|
||||
|
||||
- `requirement` (required)
|
||||
- optional `pdfContent`
|
||||
- optional `language` (`"zh-CN"` | `"en-US"`, defaults to `"zh-CN"`) — any other value silently falls back to `"zh-CN"`
|
||||
- optional `enableWebSearch` (boolean) — include web search context in outline generation
|
||||
- optional `enableImageGeneration` (boolean) — allow image generation metadata in outlines
|
||||
- optional `enableVideoGeneration` (boolean) — allow video generation metadata in outlines
|
||||
- optional `enableTTS` (boolean) — reserved for future server-side TTS generation
|
||||
- optional `agentMode` (`"default"` | `"generate"`) — controls agent profile strategy:
|
||||
- `"default"` (or omitted): uses built-in default agents
|
||||
- `"generate"`: uses LLM to generate custom agent profiles tailored to the course content
|
||||
|
||||
All optional boolean fields default to `false` when omitted. Omitting them preserves backward compatibility.
|
||||
|
||||
### Feature Detection
|
||||
|
||||
Before sending optional feature flags, query `GET {url}/api/health` and check the `capabilities` object:
|
||||
|
||||
```json
|
||||
{
|
||||
"status": "ok",
|
||||
"version": "...",
|
||||
"capabilities": {
|
||||
"webSearch": true,
|
||||
"imageGeneration": false,
|
||||
"videoGeneration": false,
|
||||
"tts": false
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Only set a feature flag to `true` if the corresponding capability is `true`. If the server does not return `capabilities` (older version), do not send the new fields.
|
||||
|
||||
Do not rely on request-time model or provider override parameters.
|
||||
|
||||
Treat the `POST` response as job submission only. Expect fields such as:
|
||||
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"jobId": "abc123",
|
||||
"status": "queued",
|
||||
"step": "queued",
|
||||
"pollUrl": "http://localhost:3000/api/generate-classroom/abc123",
|
||||
"pollIntervalMs": 5000
|
||||
}
|
||||
```
|
||||
|
||||
## PDF-Based Generation
|
||||
|
||||
1. Resolve the absolute path to the PDF.
|
||||
2. Confirm before reading the file.
|
||||
3. Parse the PDF first:
|
||||
|
||||
```text
|
||||
POST {url}/api/parse-pdf
|
||||
```
|
||||
|
||||
4. Then send `requirement` plus `pdfContent` to:
|
||||
|
||||
```text
|
||||
POST {url}/api/generate-classroom
|
||||
```
|
||||
|
||||
## Polling Loop
|
||||
|
||||
After the job is submitted:
|
||||
|
||||
1. Save `jobId`, `pollUrl`, and `pollIntervalMs`.
|
||||
2. Do not submit another generation job while this one is still `queued` or `running`.
|
||||
3. Poll:
|
||||
|
||||
```text
|
||||
GET {pollUrl}
|
||||
```
|
||||
|
||||
4. Prefer a conservative polling cadence of about 60 seconds between polls for classroom generation jobs, even if `pollIntervalMs` is shorter.
|
||||
5. Treat `queued` and `running` as in-progress states.
|
||||
6. Stop only when `status` becomes `succeeded` or `failed`.
|
||||
|
||||
### Reliability Rules
|
||||
|
||||
- Never restart the job just because a poll request fails once.
|
||||
- If a poll request returns a transient network error or `5xx`, wait about 60 seconds and retry the same `pollUrl`.
|
||||
- If the job is still running after many polls, tell the user it is still in progress and continue polling instead of resubmitting.
|
||||
- Prefer fewer poll attempts over aggressive polling. Long-running jobs are more likely to survive agent-loop limits if the tool-call cadence stays low.
|
||||
- Within a single agent turn, cap active polling to about 10 minutes. If the job is still not finished, tell the user it is still running and include the `jobId` and `pollUrl` so a later turn can continue checking without resubmitting.
|
||||
- Report progress to the user only when `status`, `step`, or visible progress meaningfully changes. Do not spam every poll result.
|
||||
- Do not try to recover from auth, provider, model, or base URL errors by changing request parameters. Tell the user to fix OpenMAIC server-side config and retry only after they confirm.
|
||||
- On `failed`, surface the server error and include the `jobId`.
|
||||
- On `succeeded`, use `result.classroomId` and `result.url` from the final poll response.
|
||||
|
||||
## If The Loop Ends First
|
||||
|
||||
If the job is still running when you stop active polling for this turn, tell the user that the classroom generation is still running in the background and invite them to come back a little later to continue checking the same job.
|
||||
|
||||
Use natural phrasing such as:
|
||||
|
||||
```text
|
||||
The classroom generation is still running in the background.
|
||||
Job ID: abc123
|
||||
|
||||
Check back with me in a little while and I can continue tracking this same job without starting over.
|
||||
```
|
||||
|
||||
## What To Return
|
||||
|
||||
Return the generated classroom ID plus a directly clickable classroom URL.
|
||||
|
||||
Output the URL as a raw absolute URL on its own line.
|
||||
|
||||
Do not wrap the URL in:
|
||||
|
||||
- bold markers such as `**...**`
|
||||
- markdown links such as `[title](url)`
|
||||
- code formatting such as `` `...` ``
|
||||
- angle brackets such as `<...>`
|
||||
- markdown tables
|
||||
|
||||
Use a compact format like:
|
||||
|
||||
```text
|
||||
Classroom ID: Uyh82Y32ZK
|
||||
Classroom URL:
|
||||
http://localhost:3001/classroom/Uyh82Y32ZK
|
||||
```
|
||||
|
||||
If the job fails, return the job ID plus the server error.
|
||||
|
||||
If generation fails, surface the server error directly instead of paraphrasing it away.
|
||||
|
||||
If the error suggests a provider or model configuration problem, explicitly tell the user to update `.env.local` or `server-providers.yml` instead of attempting a runtime override.
|
||||
|
||||
## Confirmation Requirements
|
||||
|
||||
- Ask before reading a local PDF.
|
||||
- Do not ask for a second confirmation before the generation request if the user has already clearly asked you to generate the classroom.
|
||||
@@ -0,0 +1,42 @@
|
||||
# Hosted Mode
|
||||
|
||||
Use this when the user has an access code from open.maic.chat and wants to skip local setup.
|
||||
|
||||
## Access Code Setup
|
||||
|
||||
1. Read `accessCode` from skill config (`~/.openclaw/openclaw.json` → `skills.entries.openmaic.config.accessCode`).
|
||||
2. If found, use it directly. Do not ask the user to paste the code into chat.
|
||||
3. If not found, tell the user to add their access code to the config file:
|
||||
```
|
||||
Edit ~/.openclaw/openclaw.json and set skills.entries.openmaic.config.accessCode to your access code (starts with sk-).
|
||||
```
|
||||
Wait for the user to confirm before continuing. Do not ask them to paste the code in chat.
|
||||
4. Verify connectivity: `GET https://open.maic.chat/api/health` with `Authorization: Bearer <access-code>`
|
||||
- On success: confirm connection and proceed to generation.
|
||||
- On failure (401): access code is invalid, ask the user to check or regenerate at open.maic.chat and update the config file.
|
||||
- On failure (network): suggest checking network or trying local mode.
|
||||
|
||||
## Generating a Classroom
|
||||
|
||||
Follow the same generation flow as [generate-flow.md](generate-flow.md) with these differences:
|
||||
|
||||
- **Base URL**: `https://open.maic.chat` (hardcoded, not configurable)
|
||||
- **Authorization**: Include header `Authorization: Bearer <access-code>` on all API requests
|
||||
- **Classroom URL**: `https://open.maic.chat/classroom/{id}`
|
||||
|
||||
### Feature Detection in Hosted Mode
|
||||
|
||||
Before generating, query `GET https://open.maic.chat/api/health` (with auth header) to check `capabilities`. Automatically include optional feature flags (`enableWebSearch`, `enableImageGeneration`, etc.) based on what the server supports. Do not send new fields if the server does not return `capabilities` (older version). This ensures forward compatibility — the hosted instance may update on a different schedule than the local codebase.
|
||||
|
||||
## Quota
|
||||
|
||||
- 10 generations per day, independent of web UI quota
|
||||
- If generation returns 403 with `Daily quota exhausted`, inform the user of the daily limit and that it resets at midnight.
|
||||
|
||||
## Error Handling
|
||||
|
||||
| HTTP Status | Meaning | Action |
|
||||
|-------------|---------|--------|
|
||||
| 401 | Invalid access code | Ask user to check their code or generate a new one at open.maic.chat |
|
||||
| 403 | Quota exhausted | Inform daily limit (10), suggest trying tomorrow |
|
||||
| 500 | Server error | Suggest retrying later or switching to local mode |
|
||||
@@ -0,0 +1,179 @@
|
||||
# Provider Keys
|
||||
|
||||
## Critical Boundary
|
||||
|
||||
OpenMAIC generation does not automatically reuse the OpenClaw agent's current model or API key.
|
||||
|
||||
OpenMAIC server APIs resolve their own model and provider keys from OpenMAIC server-side config.
|
||||
|
||||
This skill does not rely on runtime overrides for model, provider, API key, base URL, or provider type.
|
||||
|
||||
If the user wants to change any of those, they must edit OpenMAIC server-side config files.
|
||||
|
||||
## Interaction Policy
|
||||
|
||||
- Do not begin by asking the user to paste an API key into chat.
|
||||
- First, recommend a provider path.
|
||||
- Then ask how the user wants to configure it.
|
||||
- The user should edit `.env.local` or `server-providers.yml` themselves.
|
||||
- Do not offer to write the key for them.
|
||||
- Do not ask for the literal key in chat.
|
||||
- Do not suggest temporary request-time overrides.
|
||||
- If generation fails because of auth, provider, or model selection, direct the user back to server-side config files.
|
||||
|
||||
## Preferred User Flow
|
||||
|
||||
1. Recommend a provider option.
|
||||
2. Ask where the user wants to configure it:
|
||||
- `.env.local` (recommended for most users)
|
||||
- `server-providers.yml`
|
||||
3. Tell the user exactly which variables or YAML fields to edit.
|
||||
4. Wait for the user to confirm they finished editing before continuing.
|
||||
|
||||
## Recommendation Paths
|
||||
|
||||
### 1. Lowest-Friction Setup
|
||||
|
||||
Recommended when the user wants the smallest amount of configuration.
|
||||
|
||||
Set:
|
||||
|
||||
```env
|
||||
ANTHROPIC_API_KEY=sk-ant-...
|
||||
```
|
||||
|
||||
Why:
|
||||
|
||||
- OpenMAIC server fallback is currently `gpt-4o-mini` if `DEFAULT_MODEL` is unset.
|
||||
- If the user wants Anthropic or Google by default, they should set `DEFAULT_MODEL` explicitly.
|
||||
|
||||
### 2. Better Speed / Cost Balance
|
||||
|
||||
Recommended when the user is willing to set one extra variable.
|
||||
|
||||
Set:
|
||||
|
||||
```env
|
||||
GOOGLE_API_KEY=...
|
||||
DEFAULT_MODEL=google:gemini-3-flash-preview
|
||||
```
|
||||
|
||||
Why:
|
||||
|
||||
- Good quality-to-speed balance
|
||||
- Matches the repo's current recommendation direction better than the default fallback
|
||||
- The `google:` prefix is important. Without a provider prefix, model parsing defaults to OpenAI.
|
||||
|
||||
### 3. Existing Provider Reuse
|
||||
|
||||
Use when the user already has OpenAI or another supported provider configured and wants to stick with it.
|
||||
|
||||
Examples:
|
||||
|
||||
```env
|
||||
OPENAI_API_KEY=sk-...
|
||||
DEFAULT_MODEL=openai:gpt-4o-mini
|
||||
```
|
||||
|
||||
```env
|
||||
DEEPSEEK_API_KEY=...
|
||||
DEFAULT_MODEL=deepseek:deepseek-chat
|
||||
```
|
||||
|
||||
## Model String Rule
|
||||
|
||||
When recommending or showing `DEFAULT_MODEL`, always include the provider prefix:
|
||||
|
||||
- `google:gemini-3-flash-preview`
|
||||
- `anthropic:claude-3-5-haiku-20241022`
|
||||
- `openai:gpt-4o-mini`
|
||||
- `deepseek:deepseek-chat`
|
||||
|
||||
Do not recommend bare model IDs such as `gemini-3-flash-preview` by themselves, because OpenMAIC will otherwise parse them as OpenAI models.
|
||||
|
||||
Do not work around a wrong `DEFAULT_MODEL` by changing request parameters. The user should fix the server-side config instead.
|
||||
|
||||
## Preferred Config Method
|
||||
|
||||
For first setup, prefer `.env.local`:
|
||||
|
||||
```bash
|
||||
cp .env.example .env.local
|
||||
```
|
||||
|
||||
Then fill the chosen keys.
|
||||
|
||||
Alternative: `server-providers.yml`
|
||||
|
||||
```yaml
|
||||
providers:
|
||||
anthropic:
|
||||
apiKey: sk-ant-...
|
||||
|
||||
google:
|
||||
apiKey: ...
|
||||
|
||||
openai:
|
||||
apiKey: sk-...
|
||||
```
|
||||
|
||||
If using a non-default provider for classroom generation, also set the model selection explicitly:
|
||||
|
||||
```env
|
||||
DEFAULT_MODEL=google:gemini-3-flash-preview
|
||||
```
|
||||
|
||||
## Recommended Prompts To The User
|
||||
|
||||
Preferred:
|
||||
|
||||
- "I recommend configuring OpenMAIC through `.env.local` first. Please edit that file locally and tell me when you're done."
|
||||
- "For the simplest setup, I recommend Anthropic. For better speed/cost balance, I recommend Google plus `DEFAULT_MODEL=google:gemini-3-flash-preview`. Which path do you want?"
|
||||
|
||||
Avoid as the first move:
|
||||
|
||||
- "Send me your API key"
|
||||
- "Paste your API key here"
|
||||
- "Do you want me to write the key for you?"
|
||||
|
||||
## Confirmation Requirements
|
||||
|
||||
- Recommend one provider path first.
|
||||
- Ask the user which config-file path they want.
|
||||
- Instruct the user to modify the file themselves.
|
||||
- Wait for the user to confirm they finished editing before continuing.
|
||||
- Do not request the literal key.
|
||||
- If provider/model/auth errors happen later, tell the user exactly which config entry to fix and wait for confirmation before retrying.
|
||||
|
||||
## Optional Features
|
||||
|
||||
These features require additional provider keys beyond the core LLM provider. Ask the user if they want to enable any of these after the core LLM key is configured.
|
||||
|
||||
| Feature | Env Variable(s) | Description |
|
||||
|---------|-----------------|-------------|
|
||||
| Web Search | `TAVILY_API_KEY` | Enriches outlines with real-time web research |
|
||||
| Image Generation | `IMAGE_SEEDREAM_API_KEY`, `IMAGE_QWEN_IMAGE_API_KEY`, `IMAGE_NANO_BANANA_API_KEY` | Generates images for slides (any one suffices) |
|
||||
| Video Generation | `VIDEO_SEEDANCE_API_KEY`, `VIDEO_KLING_API_KEY`, `VIDEO_VEO_API_KEY`, `VIDEO_SORA_API_KEY` | Generates short videos (any one suffices) |
|
||||
| TTS | `TTS_OPENAI_API_KEY`, `TTS_AZURE_API_KEY`, `TTS_GLM_API_KEY`, `TTS_QWEN_API_KEY` | Text-to-speech narration (any one suffices) |
|
||||
|
||||
These are all optional. The classroom generation works without them — they only unlock richer content.
|
||||
|
||||
Alternatively, configure via `server-providers.yml`:
|
||||
|
||||
```yaml
|
||||
web-search:
|
||||
tavily:
|
||||
apiKey: tvly-...
|
||||
|
||||
image:
|
||||
seedream:
|
||||
apiKey: ...
|
||||
|
||||
video:
|
||||
seedance:
|
||||
apiKey: ...
|
||||
|
||||
tts:
|
||||
openai-tts:
|
||||
apiKey: sk-...
|
||||
```
|
||||
@@ -0,0 +1,69 @@
|
||||
# Startup Modes
|
||||
|
||||
## Goal
|
||||
|
||||
Help the user choose how OpenMAIC should run before you start anything.
|
||||
|
||||
## Options
|
||||
|
||||
### 1. Development Mode
|
||||
|
||||
Recommended for first-time setup and debugging.
|
||||
|
||||
```bash
|
||||
pnpm dev
|
||||
```
|
||||
|
||||
Tradeoff:
|
||||
|
||||
- Fastest feedback loop
|
||||
- Best for validating config changes
|
||||
- Not representative of production startup
|
||||
|
||||
### 2. Production-Like Local Mode
|
||||
|
||||
Recommended when the user wants behavior closer to a deployed server.
|
||||
|
||||
```bash
|
||||
pnpm build && pnpm start
|
||||
```
|
||||
|
||||
Tradeoff:
|
||||
|
||||
- Closer to production
|
||||
- Slower startup than `pnpm dev`
|
||||
|
||||
### 3. Docker Compose
|
||||
|
||||
Use only when the user explicitly wants containerized startup or wants to avoid local Node setup details.
|
||||
|
||||
```bash
|
||||
docker compose up --build
|
||||
```
|
||||
|
||||
Tradeoff:
|
||||
|
||||
- Cleaner isolation
|
||||
- Heavier and slower
|
||||
- Harder to debug application-level issues quickly
|
||||
|
||||
## Recommendation Order
|
||||
|
||||
1. `pnpm dev`
|
||||
2. `pnpm build && pnpm start`
|
||||
3. `docker compose up --build`
|
||||
|
||||
## Health Check
|
||||
|
||||
After startup, verify:
|
||||
|
||||
```bash
|
||||
curl -fsS http://localhost:3000/api/health
|
||||
```
|
||||
|
||||
If the skill config provides a custom `url`, use that instead.
|
||||
|
||||
## Confirmation Requirements
|
||||
|
||||
- Ask the user to choose one startup mode.
|
||||
- Ask again before running the selected command.
|
||||
@@ -0,0 +1,155 @@
|
||||
---
|
||||
name: reader-digest-flow
|
||||
description: Orchestrate the end-to-end daily digest workflow around the reader project. Use when the user wants to run the AI daily digest flow, generate a daily report from reader payloads, publish the digest to Hugo, report the digest back in chat, select valuable articles, and then summarize only the selected articles into IMA knowledge notes. Triggers include requests like '跑今天日报', '生成日报', '汇报今天内容', '把选中的文章沉淀', '更新 Hugo', or any request to operate the reader → digest → selection → knowledge-base flow.
|
||||
---
|
||||
|
||||
# Reader Digest Flow
|
||||
|
||||
Run the reader-based daily digest as a fixed SOP. Treat this skill as the orchestrator for the workflow; do not move Hugo, Feishu reporting, or IMA upload logic into the reader project itself.
|
||||
|
||||
## Core Rules
|
||||
|
||||
- Run `reader` / MCP for upstream fetching, extraction, filtering, payload generation, and selected-article post-processing.
|
||||
- **For normal production runs, mark processed FreshRSS items as read. Only skip mark-read when the user explicitly says the run is debug/test/validation.**
|
||||
- **Do not generate the daily digest from examples or placeholder data. Always require a real payload first.**
|
||||
- Generate the daily digest markdown from the payload in OpenClaw.
|
||||
- Split outputs into a **public digest** for Hugo and an **internal review digest** for chat / operator decision-making.
|
||||
- Publish only the public digest to Hugo.
|
||||
- Report the internal review digest back to the user in chat.
|
||||
- **Do not upload the full daily digest to IMA.**
|
||||
- **Do not upload any selected article summary to IMA until the user has explicitly confirmed the selection.**
|
||||
- Upload **only user-selected articles** to IMA as individual knowledge notes.
|
||||
- Default IMA target for daily single-article summaries is the `daily` knowledge base.
|
||||
- Resolve that default target from reader `.env` (`IMA_DAILY_KNOWLEDGE_BASE_ID`, `IMA_DAILY_KNOWLEDGE_BASE_NAME`) and verify it at runtime before upload.
|
||||
- If the configured daily knowledge base is missing, try to locate it by name; if still missing, create `daily` and continue.
|
||||
- For selected article notes, use the dedicated article-summary flow in `reader`.
|
||||
- Prefer `ARTICLE_SUMMARY_*` LLM settings for selected article summaries; fall back to the main `LLM_*` settings only if the dedicated settings are absent.
|
||||
- Do not re-fetch original article URLs for selected summaries; always use the existing extracted article text.
|
||||
- Default selected-article summary output should be organized by date, for example under `outputs/freshrss/single_summaries/YYYY-MM-DD/`.
|
||||
|
||||
## Fixed Flow
|
||||
|
||||
### Phase 1: Run reader pipeline
|
||||
|
||||
In the `reader` project, run the FreshRSS pipeline and obtain a real payload.
|
||||
|
||||
Minimum expected artifacts:
|
||||
|
||||
- `outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json`
|
||||
- `outputs/freshrss/rerun/<run-id>/run-report.json`
|
||||
- extracted article data, such as:
|
||||
- `outputs/freshrss/extracted/freshrss.extracted.json`
|
||||
|
||||
If the pipeline fails, stop and report the exact failure point.
|
||||
|
||||
### Phase 2: Generate daily digest markdown
|
||||
|
||||
Read the real payload and generate digest markdown for the day.
|
||||
If there is no real payload, stop instead of writing a fake or example digest.
|
||||
|
||||
Produce two output views from the same payload:
|
||||
|
||||
1. **Public digest** — for Hugo / public browsing
|
||||
2. **Internal review digest** — for chat reporting and operator decisions
|
||||
|
||||
Write only the public digest into Hugo using this structure:
|
||||
|
||||
- `content/daily/YYYY-MM-DD/index.md`
|
||||
|
||||
The public digest is the browsing layer, not the long-term knowledge layer.
|
||||
It must not expose internal workflow states or operator-facing review labels.
|
||||
|
||||
Recommended **public digest** structure:
|
||||
|
||||
- `今日概览`
|
||||
- `今日重点`
|
||||
- `趋势观察`
|
||||
- `延伸阅读`
|
||||
- `信息来源`
|
||||
|
||||
Recommended **internal review digest** structure:
|
||||
|
||||
- `今日候选概况`
|
||||
- `已入选重点`
|
||||
- `待你确认`
|
||||
- `建议沉淀到 IMA`
|
||||
- `原始候选清单`
|
||||
|
||||
### Phase 3: Publish to Hugo
|
||||
|
||||
Publish only the public digest to Hugo and verify:
|
||||
|
||||
- list page works
|
||||
- detail page works
|
||||
- latest digest is visible
|
||||
|
||||
Do not block on style polish unless the user explicitly asks.
|
||||
|
||||
### Phase 4: Report digest back to the user
|
||||
|
||||
Send the internal review digest in chat and ask the user which articles should be retained for long-term knowledge.
|
||||
|
||||
At this step:
|
||||
|
||||
- the public digest is already in Hugo
|
||||
- the internal review digest stays in chat / operator workflow
|
||||
- the digest is **not** uploaded to IMA
|
||||
- the user decides which articles are worth preserving
|
||||
|
||||
### Phase 5: Summarize selected articles
|
||||
|
||||
For every article explicitly selected by the user:
|
||||
|
||||
1. use the `reader` article-summary capability
|
||||
2. point it at the existing extracted payload
|
||||
3. pass the selected `item_id` values
|
||||
4. generate one markdown summary per article
|
||||
|
||||
Preferred routes:
|
||||
|
||||
- MCP tool: `generate_article_summaries`
|
||||
- CLI fallback: `scripts/run_article_summaries.py`
|
||||
|
||||
Use the real extracted JSON structure already produced by the project. Do not invent alternative inputs.
|
||||
|
||||
### Phase 6: Upload selected article notes to IMA
|
||||
|
||||
Upload only the generated single-article markdown summaries to IMA.
|
||||
|
||||
Default target knowledge base for this phase:
|
||||
|
||||
- `daily`
|
||||
- Read from reader `.env` via `IMA_DAILY_KNOWLEDGE_BASE_ID` and `IMA_DAILY_KNOWLEDGE_BASE_NAME`
|
||||
- Verify the target at runtime through IMA APIs / skill lookups before upload
|
||||
- If the configured target does not exist, try to find `daily` by name; if still absent, create it and continue
|
||||
|
||||
Do not upload:
|
||||
|
||||
- the full daily digest
|
||||
- raw payloads
|
||||
- raw extraction output
|
||||
|
||||
## Operational Guidance
|
||||
|
||||
- Prefer real run outputs over examples.
|
||||
- Verify at each boundary with real files or accessible URLs.
|
||||
- When validating selected article summaries, confirm that a markdown file is actually generated.
|
||||
- If Codex or another coding agent is asked to implement workflow changes inside `reader`, keep the project boundary clean:
|
||||
- workflow logic in `reader`
|
||||
- orchestration logic in this skill / OpenClaw
|
||||
|
||||
## Key Paths
|
||||
|
||||
### Reader project
|
||||
|
||||
- `/home/ubuntu/zhu/github/reader`
|
||||
|
||||
### Hugo project
|
||||
|
||||
- `/home/ubuntu/zhu/apps/hugo-site`
|
||||
- digest content root:
|
||||
- `/home/ubuntu/zhu/apps/hugo-site/content/daily/`
|
||||
|
||||
## References
|
||||
|
||||
Read `references/flow.md` when you need the concrete step-by-step command checklist and file expectations.
|
||||
@@ -0,0 +1,162 @@
|
||||
# Reader Digest Flow Reference
|
||||
|
||||
## Purpose
|
||||
|
||||
Concrete operational checklist for the `reader-digest-flow` skill.
|
||||
|
||||
## Default Operating Model
|
||||
|
||||
### Layering
|
||||
|
||||
- `reader` layer:
|
||||
- FreshRSS pull
|
||||
- extraction
|
||||
- summary/filter/payload generation
|
||||
- selected-article summary capability
|
||||
- OpenClaw / skill layer:
|
||||
- public digest generation
|
||||
- internal review digest generation
|
||||
- Hugo publishing
|
||||
- chat reporting
|
||||
- user confirmation handling
|
||||
- calling selected-article summaries
|
||||
- IMA upload orchestration
|
||||
- Hugo layer:
|
||||
- public digest browsing and archive only
|
||||
- IMA layer:
|
||||
- long-term storage for selected article notes only
|
||||
|
||||
### Hard rules
|
||||
|
||||
- Do not upload the full digest to IMA.
|
||||
- Upload only explicitly user-selected articles to IMA.
|
||||
- Do not generate a digest without a real payload.
|
||||
- Generate two views from the same payload: a public digest for Hugo and an internal review digest for chat/operator workflow.
|
||||
- Do not expose internal review states or operator-facing labels in the public digest.
|
||||
- Do not re-fetch original URLs for selected summaries; use existing extracted text.
|
||||
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization, with fallback to main `LLM_*` only if needed.
|
||||
- Actively report progress after each completed phase.
|
||||
|
||||
## Step-by-step checklist
|
||||
|
||||
### 1. Run reader pipeline
|
||||
|
||||
Project root:
|
||||
|
||||
```bash
|
||||
/home/ubuntu/zhu/github/reader
|
||||
```
|
||||
|
||||
Default behavior for a normal production run:
|
||||
|
||||
- run with mark-read enabled
|
||||
- only skip mark-read if the user explicitly says the run is debug, test, or validation
|
||||
|
||||
Typical artifacts to inspect after a successful run:
|
||||
|
||||
```text
|
||||
outputs/freshrss/rerun/<run-id>/candidates/openclaw-delivery-payload.json
|
||||
outputs/freshrss/rerun/<run-id>/run-report.json
|
||||
outputs/freshrss/rerun/<run-id>/extracted/item-XX.extracted.json
|
||||
```
|
||||
|
||||
### 2. Generate digest markdown
|
||||
|
||||
Generate two output views from the same payload:
|
||||
|
||||
1. a **public digest** for Hugo / public readers
|
||||
2. an **internal review digest** for chat / operator workflow
|
||||
|
||||
Write only the public digest into Hugo here:
|
||||
|
||||
```text
|
||||
/home/ubuntu/zhu/apps/hugo-site/content/daily/YYYY-MM-DD/index.md
|
||||
```
|
||||
|
||||
Recommended public-digest front matter:
|
||||
|
||||
```toml
|
||||
+++
|
||||
title = "AI 日报 · YYYY-MM-DD"
|
||||
date = YYYY-MM-DDTHH:MM:SS+08:00
|
||||
summary = "当日日报摘要"
|
||||
+++
|
||||
```
|
||||
|
||||
Recommended **public digest** structure:
|
||||
|
||||
- `今日概览`
|
||||
- `今日重点`
|
||||
- `趋势观察`
|
||||
- `延伸阅读`
|
||||
- `信息来源`
|
||||
|
||||
Recommended **internal review digest** structure:
|
||||
|
||||
- `今日候选概况`
|
||||
- `已入选重点`
|
||||
- `待你确认`
|
||||
- `建议沉淀到 IMA`
|
||||
- `原始候选清单`
|
||||
|
||||
### 3. Publish Hugo
|
||||
|
||||
Publish only the public digest to Hugo.
|
||||
|
||||
Expected verification targets:
|
||||
|
||||
- homepage works
|
||||
- `/daily/` works
|
||||
- `/daily/YYYY-MM-DD/` works
|
||||
|
||||
### 4. Report digest in chat
|
||||
|
||||
Provide the internal review digest in chat and ask which articles should be retained.
|
||||
|
||||
### 5. Generate selected article summaries
|
||||
|
||||
Only do this after the user explicitly confirms which articles to retain.
|
||||
|
||||
Preferred MCP tool:
|
||||
|
||||
- `generate_article_summaries`
|
||||
|
||||
Expected inputs:
|
||||
|
||||
- `extracted_path`
|
||||
- `selected_ids`
|
||||
- optional `output_dir`
|
||||
- optional article-summary LLM overrides
|
||||
|
||||
Recommended output layout:
|
||||
|
||||
- `outputs/freshrss/single_summaries/YYYY-MM-DD/`
|
||||
|
||||
CLI fallback:
|
||||
|
||||
```bash
|
||||
python scripts/run_article_summaries.py \
|
||||
--extracted outputs/freshrss/rerun/<run-id>/extracted/item-01.extracted.json \
|
||||
--ids <item_id_1> <item_id_2> \
|
||||
--output-dir outputs/freshrss/single_summaries/YYYY-MM-DD
|
||||
```
|
||||
|
||||
### 6. Upload selected summaries to IMA
|
||||
|
||||
Upload only the generated markdown files for the selected articles.
|
||||
|
||||
Default target knowledge base:
|
||||
|
||||
- `daily`
|
||||
- Read `IMA_DAILY_KNOWLEDGE_BASE_ID` / `IMA_DAILY_KNOWLEDGE_BASE_NAME` from reader `.env`
|
||||
- Verify the configured target at runtime before upload
|
||||
- If the configured target is unavailable, resolve by name `daily`; if still absent, create `daily`
|
||||
|
||||
## Hard rules recap
|
||||
|
||||
- Never generate a digest from placeholder or example data when a real run is expected.
|
||||
- For normal runs, mark processed FreshRSS items as read unless the user explicitly requested a debug/test/validation run.
|
||||
- Public digest goes to Hugo; internal review digest goes to chat; neither full digest goes to IMA.
|
||||
- Only explicitly user-selected articles go to IMA.
|
||||
- Selected article summaries use extracted text, not live refetch.
|
||||
- Prefer `ARTICLE_SUMMARY_*` for selected article summarization.
|
||||
Reference in New Issue
Block a user