feat(phase1): 完成文本提取和文档分块服务

Task 5.1: TextExtractor 服务
- 创建 TextExtractorService(仅支持 .md 和 .txt)
- 其他格式(.docx、.pdf)需通过外部转换服务先转为 Markdown
- 支持 UTF-8 编码的纯文本提取
- 提供文件格式验证方法

Task 5.2: 文档分块服务适配
- 修改 DocumentChunk DTO:添加 @Builder 支持
- 字段重命名:startIndex/endIndex → startOffset/endOffset
- 修复 DocumentChunkService 中的三处构造调用
- 使用 builder 模式替代构造函数

技术决策:
- 简化文本提取,只支持 Markdown 和纯文本
- 复杂格式转换由外部服务处理(分离关注点)
- 统一使用 Lombok @Builder 简化对象构建

编译验证:BUILD SUCCESS

Progress: 25/33 tasks completed (76%)
This commit is contained in:
zhuyongxin
2026-06-23 15:15:10 +08:00
parent 360e4febae
commit ea77518880
4 changed files with 126 additions and 55 deletions
@@ -1,59 +1,41 @@
package com.superbiz.agent.dto;
import lombok.Getter;
import lombok.Setter;
import lombok.AllArgsConstructor;
import lombok.Builder;
import lombok.Data;
import lombok.NoArgsConstructor;
/**
* 文档分片
*/
@Setter
@Getter
@Data
@Builder
@NoArgsConstructor
@AllArgsConstructor
public class DocumentChunk {
// Getters and Setters
/**
* 分片内容
*/
private String content;
/**
* 分片在原文档中的起始位置
*/
private int startIndex;
private int startOffset;
/**
* 分片在原文档中的结束位置
*/
private int endIndex;
private int endOffset;
/**
* 分片序号(从0开始)
*/
private int chunkIndex;
/**
* 分片标题或上下文信息
*/
private String title;
public DocumentChunk() {
}
public DocumentChunk(String content, int startIndex, int endIndex, int chunkIndex) {
this.content = content;
this.startIndex = startIndex;
this.endIndex = endIndex;
this.chunkIndex = chunkIndex;
}
@Override
public String toString() {
return "DocumentChunk{" +
"chunkIndex=" + chunkIndex +
", title='" + title + '\'' +
", contentLength=" + (content != null ? content.length() : 0) +
", startIndex=" + startIndex +
", endIndex=" + endIndex +
'}';
}
}