feat: archive mvp demo trace acceptance
This commit is contained in:
@@ -4,6 +4,7 @@
|
|||||||
|
|
||||||
| 日期 | slug | 领域 | 关键词 | 状态 |
|
| 日期 | slug | 领域 | 关键词 | 状态 |
|
||||||
|---|---|---|---|---|
|
|---|---|---|---|---|
|
||||||
|
| 2026-07-03 | mvp-demo-trace-acceptance | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
|
||||||
| 2026-05-29 | chatmodel-abstraction | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | archived |
|
| 2026-05-29 | chatmodel-abstraction | 解耦/多模型路由 | ChatModel, EmbeddingModel, DeepSeek, BGE-M3, SiliconFlow, Spring AI | archived |
|
||||||
| 2026-06-23 | phase1-infrastructure | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | archived |
|
| 2026-06-23 | phase1-infrastructure | 基础设施/文档管理 | MySQL, Redis, Milvus, Flyway, JPA, 向量检索, 类别过滤 | archived |
|
||||||
| 2026-06-24 | lookup-knowledge-integration | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | archived |
|
| 2026-06-24 | lookup-knowledge-integration | 知识库检索 | L0精确匹配, L1语义检索, frontmatter, 混合检索 | archived |
|
||||||
|
|||||||
@@ -0,0 +1,65 @@
|
|||||||
|
# MVP Demo Trace Acceptance
|
||||||
|
|
||||||
|
## Result
|
||||||
|
|
||||||
|
Accepted for implementation scope.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
### Static Verification
|
||||||
|
|
||||||
|
- Command: `mvn -q -DskipTests compile`
|
||||||
|
- Result: passed
|
||||||
|
- Notes: New trace controller, service, DTO, profile, verifier fallback, and test sources compile with the project.
|
||||||
|
|
||||||
|
### Script Verification
|
||||||
|
|
||||||
|
- Command: `mvn -q "-Dtest=DiagnosisTraceServiceTest,ChatServiceSupervisorAgentTest" test`
|
||||||
|
- Result: passed
|
||||||
|
- Notes: Covers successful trace aggregation, missing-session 404 path via `SessionNotFoundException`, low-confidence no-retry behavior, method-tool injection, and verifier fallback when Supervisor skips `chat_verifier`.
|
||||||
|
|
||||||
|
### OpenSpec Verification
|
||||||
|
|
||||||
|
- Command: `openspec validate mvp-demo-trace-acceptance --strict`
|
||||||
|
- Result: passed
|
||||||
|
|
||||||
|
### GitNexus Verification
|
||||||
|
|
||||||
|
- Result: skipped by user decision
|
||||||
|
- Notes: User requested subsequent project flow to bypass GitNexus.
|
||||||
|
|
||||||
|
### Manual / Runtime Verification
|
||||||
|
|
||||||
|
- Steps: Follow `mvp/demo/README.md` with `--spring.profiles.active=mvp-demo`.
|
||||||
|
- Result: passed
|
||||||
|
- Notes:
|
||||||
|
- Session `mvp-demo-payment-timeout-20260703-rerun2` completed as `SUCCESS`.
|
||||||
|
- Chat request returned `code=200`, `success=true`, and the same `sessionId`.
|
||||||
|
- Chat duration was `96316 ms`; persisted session duration was `95028 ms`.
|
||||||
|
- Trace API returned `code=200`, `returnedSteps=13`, `returnedTools=12`, `hasVerifier=true`, and `verifierVerdict=LOW_CONFID`.
|
||||||
|
- Trace agents included `planner,executor,verifier`.
|
||||||
|
- Trace tools included `lookup_knowledge,query_logs,query_metrics`.
|
||||||
|
- Feedback submission returned success, and a follow-up trace query showed `feedback=useful`.
|
||||||
|
- MySQL verification confirmed `agent_step` count `13` with agents `executor,planner,verifier`.
|
||||||
|
- MySQL verification confirmed `tool_invocation` count `12` with tools `lookup_knowledge,query_logs,query_metrics`.
|
||||||
|
|
||||||
|
## Completed Scope
|
||||||
|
|
||||||
|
- Added `GET /api/diagnosis/{sessionId}/trace`.
|
||||||
|
- Added read-only trace aggregation from persisted diagnosis tables.
|
||||||
|
- Added `mvp-demo` profile overlay.
|
||||||
|
- Added payment-timeout demo acceptance documentation.
|
||||||
|
- Added MVP note for interview storytelling.
|
||||||
|
- Added verifier fallback so runtime trace remains complete when Supervisor returns without `verifier_output`.
|
||||||
|
|
||||||
|
## Known Limits
|
||||||
|
|
||||||
|
- `mvp-demo` is not a fully offline mock runtime.
|
||||||
|
- Runtime still depends on available MySQL, Redis, Milvus/Zilliz, model, and embedding configuration.
|
||||||
|
- Sensitive configuration cleanup remains intentionally deferred.
|
||||||
|
- Supervisor can still make inefficient routing choices inside a single round; `ChatService` now invokes `chat_verifier` as a fallback when Supervisor returns without `verifier_output`, so trace completeness is preserved for the MVP demo.
|
||||||
|
|
||||||
|
## Handoff
|
||||||
|
|
||||||
|
- Runtime demo passed with current infrastructure.
|
||||||
|
- OpenSpec archive confirmation: requested by user after successful rerun.
|
||||||
@@ -0,0 +1,35 @@
|
|||||||
|
# MVP Demo Trace Acceptance Brief
|
||||||
|
|
||||||
|
## Background
|
||||||
|
|
||||||
|
- User goal: make the MVP runnable, observable, and explainable for an Agent Engineer interview.
|
||||||
|
- Current problem: the system can execute diagnosis, but reviewers need a simple way to replay one session from final answer back to agent steps and tool evidence.
|
||||||
|
- Associated OpenSpec: `openspec/changes/mvp-demo-trace-acceptance/`
|
||||||
|
- Devflow scale: standard-light.
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
- In scope:
|
||||||
|
- `mvp-demo` Spring profile overlay.
|
||||||
|
- `GET /api/diagnosis/{sessionId}/trace` read-only API.
|
||||||
|
- Trace aggregation DTO/service/controller.
|
||||||
|
- Focused service tests.
|
||||||
|
- Demo and acceptance documentation.
|
||||||
|
- Out of scope:
|
||||||
|
- Sensitive configuration cleanup.
|
||||||
|
- Full offline LLM/vector/database mock runtime.
|
||||||
|
- Database schema migration.
|
||||||
|
- Changes to chat execution, verifier routing, upload, or feedback behavior.
|
||||||
|
- Impact area:
|
||||||
|
- `src/main/java/com/superbiz/agent/controller`
|
||||||
|
- `src/main/java/com/superbiz/agent/service`
|
||||||
|
- `src/main/java/com/superbiz/agent/dto`
|
||||||
|
- `src/main/resources/application-mvp-demo.yml`
|
||||||
|
- `mvp/demo`
|
||||||
|
- `mvp/notes`
|
||||||
|
|
||||||
|
## OpenSpec Alignment
|
||||||
|
|
||||||
|
- proposal coverage: covered
|
||||||
|
- specs coverage: covered
|
||||||
|
- tasks coverage: covered
|
||||||
@@ -0,0 +1,87 @@
|
|||||||
|
# MVP Demo Trace Acceptance Decisions
|
||||||
|
|
||||||
|
## Clarify
|
||||||
|
|
||||||
|
- Entry summary: continue the MVP toward a runnable and explainable demo by adding an `mvp-demo` profile, an end-to-end acceptance case, and a trace query API.
|
||||||
|
- Slug: `mvp-demo-trace-acceptance`
|
||||||
|
- Devflow scale: standard-light. The change adds a public read-only API and documentation, but does not alter core chat execution or persistence schemas.
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
- `devflow/index.md` was checked. Relevant history includes `session-storage`, `confidence-feedback`, `executor-action-memory-relevance`, and `chat-verifier-agent`.
|
||||||
|
- `mvp/notes/agent-engineering-decisions.md` already recommends the next phase as "可复现 MVP Demo", including `mvp-demo` profile, fixed diagnosis case, one-click request, and `GET /api/diagnosis/{sessionId}/trace`.
|
||||||
|
- `mvp/issues/ISS-003-mvp-design-implementation-review.md` identifies test stability, session traceability, verifier evidence chain, upload path, and SupervisorAgent consistency as recent MVP concerns. Security cleanup is intentionally deferred by user decision.
|
||||||
|
|
||||||
|
## Question Pool
|
||||||
|
|
||||||
|
| # | Dimension | Question | Mode | Status |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| Q1 | Terminology | Should "trace" mean persisted diagnosis execution evidence instead of transient frontend chat history? | evidence-driven | Resolved |
|
||||||
|
| Q2 | Boundary | Should this change modify chat execution or only expose existing persisted evidence? | evidence-driven | Resolved |
|
||||||
|
| Q3 | Acceptance | What proves the MVP flow is end-to-end enough for demo/interview use? | evidence-driven | Resolved |
|
||||||
|
| Q4 | Interface | What is the API impact level for `GET /api/diagnosis/{sessionId}/trace`? | evidence-driven | Resolved |
|
||||||
|
|
||||||
|
## Evidence-driven
|
||||||
|
|
||||||
|
| Conclusion | Evidence Source | Reported To User |
|
||||||
|
|---|---|---|
|
||||||
|
| Trace should aggregate persisted diagnosis evidence, not Redis-only chat history. | `DiagnosisSession`, `AgentStep`, `ToolInvocation` entities and repositories | Reported in progress update |
|
||||||
|
| Core chat execution does not need to change for this slice. | Existing unified chat path and SupervisorAgent commits; requested scope is demo/profile/trace/acceptance | Reported in progress update |
|
||||||
|
| End-to-end acceptance should cover start -> chat -> trace -> feedback. | `ChatController`, `FeedbackController`, traceable session id decision in MVP notes | Reported in progress update |
|
||||||
|
| Trace API is additive L3 because it is a new HTTP API for frontend/demo consumers. | sm-flow interface impact rules | Recorded in OpenSpec design |
|
||||||
|
|
||||||
|
## User-interview
|
||||||
|
|
||||||
|
| Question | User Words | Confirmation | OpenSpec Writeback |
|
||||||
|
|---|---|---|---|
|
||||||
|
| Should security/sensitive config cleanup be included? | "安全问题先不考虑"; "敏感配置先不做" | Confirmed | Non-goal |
|
||||||
|
| Should this be implemented under sm-flow? | "按照 sm-flow 的流程来实现吧" | Confirmed | This change follows sm-flow artifacts |
|
||||||
|
|
||||||
|
## Key Decisions
|
||||||
|
|
||||||
|
- Decision: Add a new trace API instead of embedding trace details in `/api/chat`.
|
||||||
|
- Reason: Chat execution and observability should stay decoupled.
|
||||||
|
- Impact: Demo can query trace after any successful chat request using the same session id.
|
||||||
|
- Risk accepted: Response shape is new and should be treated as demo-facing contract.
|
||||||
|
|
||||||
|
- Decision: Keep `mvp-demo` profile as configuration overlay, not a fully mocked standalone runtime.
|
||||||
|
- Reason: The current MVP still depends on real DB/Redis/Milvus/LLM for full chat execution; this change avoids inventing a fake runtime that hides integration behavior.
|
||||||
|
- Impact: Demo profile improves repeatability for logs/metrics, while docs remain explicit about required external services.
|
||||||
|
- Risk accepted: End-to-end acceptance may still require valid infrastructure and keys.
|
||||||
|
|
||||||
|
## Cross-Artifact Alignment
|
||||||
|
|
||||||
|
| Upstream -> Downstream | Check | Status |
|
||||||
|
|---|---|---|
|
||||||
|
| brief/prd -> proposal | Goal, scope, non-goals, and acceptance expectation are in proposal | Aligned |
|
||||||
|
| proposal -> design | Scope, constraints, and API impact are in design | Aligned |
|
||||||
|
| design -> specs/tasks | Trace DTO, controller/service, demo profile, and docs are represented | Aligned |
|
||||||
|
| specs -> tasks | Observable behavior is covered by executable tasks | Aligned |
|
||||||
|
|
||||||
|
## Architecture Audit
|
||||||
|
|
||||||
|
- Data path: HTTP trace request -> controller -> trace service -> repositories -> aggregate DTO -> `Result.success`.
|
||||||
|
- The service is read-only and does not mutate diagnosis, step, tool, or feedback state.
|
||||||
|
- No schema change is needed because all required fields already exist in `diagnosis_session`, `agent_step`, and `tool_invocation`.
|
||||||
|
- Main risk is response size for large sessions; MVP mitigates by returning previews already persisted by tools rather than raw external logs.
|
||||||
|
- The additive API is acceptable for MVP because old callers remain unaffected.
|
||||||
|
|
||||||
|
## Pre-apply Research
|
||||||
|
|
||||||
|
- Reference implementations read:
|
||||||
|
- `ChatController` for `/api` controller conventions.
|
||||||
|
- `FeedbackController` for simple API controller shape.
|
||||||
|
- `GlobalExceptionHandler` and `SessionNotFoundException` for 404 handling.
|
||||||
|
- `DiagnosisSessionRepository`, `AgentStepRepository`, `ToolInvocationRepository` for available queries.
|
||||||
|
- `DiagnosisSession`, `AgentStep`, `ToolInvocation` for fields.
|
||||||
|
- Impact analysis:
|
||||||
|
- `DiagnosisSessionRepository`: LOW, direct imports in service/controller paths.
|
||||||
|
- `AgentStepRepository`: HIGH because it participates in chat/AiOps flows. This change only consumes existing query methods and does not modify the repository.
|
||||||
|
- `ToolInvocationRepository`: LOW.
|
||||||
|
|
||||||
|
## Commit Gate
|
||||||
|
|
||||||
|
- OpenSpec proposal/design/specs/tasks exist.
|
||||||
|
- API impact: L3 additive collaboration API, documented in design and spec.
|
||||||
|
- User-confirmed non-goal: sensitive configuration cleanup remains out of scope.
|
||||||
|
- No unresolved user-interview questions remain for this slice.
|
||||||
@@ -0,0 +1,25 @@
|
|||||||
|
# MVP Demo Trace Acceptance Evidence
|
||||||
|
|
||||||
|
## Evidence
|
||||||
|
|
||||||
|
| Source | Evidence | Conclusion | Reported |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `DiagnosisSessionRepository` | Existing `findBySessionId(String)` query | Trace can locate the session without new repository methods | Yes |
|
||||||
|
| `AgentStepRepository` | Existing `findBySessionIdOrderByStepIndex(String)` query | Agent steps can be returned in execution order | Yes |
|
||||||
|
| `ToolInvocationRepository` | Existing `findBySessionIdOrderByIdAsc(String)` query | Tool evidence can be returned in persisted order | Yes |
|
||||||
|
| `GlobalExceptionHandler` | Handles `SessionNotFoundException` as HTTP 404 with `Result.error(404, ...)` | Missing trace can reuse existing error contract | Yes |
|
||||||
|
| `mvn -q "-Dtest=DiagnosisTraceServiceTest" test` | Command passed | Trace aggregation behavior is covered offline | Yes |
|
||||||
|
| `mvn -q -DskipTests compile` | Command passed | New code compiles with the full project | Yes |
|
||||||
|
| `gitnexus detect-changes --repo SuperBizAgent-java` | Command completed with `No changes detected` and line-ending warnings | Required GitNexus check ran; output likely does not capture newly added files | Yes |
|
||||||
|
|
||||||
|
## Evidence-driven Conclusions
|
||||||
|
|
||||||
|
- Conclusion: No database migration is required.
|
||||||
|
- Evidence: All trace fields are available from existing `diagnosis_session`, `agent_step`, and `tool_invocation` entities.
|
||||||
|
- Risk: Response shape becomes a new API contract.
|
||||||
|
- User confirmation: Not required; additive L3 API recorded in OpenSpec.
|
||||||
|
|
||||||
|
- Conclusion: Trace aggregation can be tested without external infrastructure.
|
||||||
|
- Evidence: `DiagnosisTraceServiceTest` uses mocked repositories and an `ObjectMapper`.
|
||||||
|
- Risk: Runtime integration still depends on configured infrastructure.
|
||||||
|
- User confirmation: Not required; limitation recorded in acceptance docs.
|
||||||
@@ -0,0 +1,94 @@
|
|||||||
|
# MVP Demo Runbook
|
||||||
|
|
||||||
|
This demo proves the MVP flow from user question to persisted diagnosis trace.
|
||||||
|
|
||||||
|
## Prerequisites
|
||||||
|
|
||||||
|
- MySQL, Redis, Milvus/Zilliz, and LLM/embedding configuration are available through the current project configuration.
|
||||||
|
- Security and secret cleanup are intentionally out of scope for this MVP slice.
|
||||||
|
- The `mvp-demo` profile enables mock Prometheus and CLS providers so log and metric tools can return repeatable evidence.
|
||||||
|
|
||||||
|
## Start
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
mvn spring-boot:run "-Dspring-boot.run.profiles=mvp-demo"
|
||||||
|
```
|
||||||
|
|
||||||
|
The service listens on:
|
||||||
|
|
||||||
|
```text
|
||||||
|
http://localhost:9900
|
||||||
|
```
|
||||||
|
|
||||||
|
## 1. Run Chat Diagnosis
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
$sessionId = "mvp-demo-payment-timeout-001"
|
||||||
|
$body = @{
|
||||||
|
Id = $sessionId
|
||||||
|
Question = "支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。"
|
||||||
|
} | ConvertTo-Json
|
||||||
|
|
||||||
|
Invoke-RestMethod `
|
||||||
|
-Method Post `
|
||||||
|
-Uri "http://localhost:9900/api/chat" `
|
||||||
|
-ContentType "application/json" `
|
||||||
|
-Body $body
|
||||||
|
```
|
||||||
|
|
||||||
|
Expected result:
|
||||||
|
|
||||||
|
- `data.success` is `true`.
|
||||||
|
- `data.sessionId` equals `mvp-demo-payment-timeout-001`.
|
||||||
|
- `data.answer` contains a diagnosis answer.
|
||||||
|
|
||||||
|
## 2. Query Trace
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
Invoke-RestMethod `
|
||||||
|
-Method Get `
|
||||||
|
-Uri "http://localhost:9900/api/diagnosis/$sessionId/trace"
|
||||||
|
```
|
||||||
|
|
||||||
|
Expected result:
|
||||||
|
|
||||||
|
- `code` is `200`.
|
||||||
|
- `data.session.sessionId` equals the chat session id.
|
||||||
|
- `data.steps` contains planner/executor/verifier records for complex questions.
|
||||||
|
- `data.toolInvocations` contains evidence tool calls such as `lookup_knowledge`, `query_logs`, or `query_metrics`.
|
||||||
|
- `data.session.selfEvaluation` contains verifier or rule evaluation when available.
|
||||||
|
|
||||||
|
## 3. Submit Feedback
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
$feedback = @{
|
||||||
|
sessionId = $sessionId
|
||||||
|
feedback = "useful"
|
||||||
|
} | ConvertTo-Json
|
||||||
|
|
||||||
|
Invoke-RestMethod `
|
||||||
|
-Method Post `
|
||||||
|
-Uri "http://localhost:9900/api/feedback" `
|
||||||
|
-ContentType "application/json" `
|
||||||
|
-Body $feedback
|
||||||
|
```
|
||||||
|
|
||||||
|
Expected result:
|
||||||
|
|
||||||
|
- `success` is `true`.
|
||||||
|
- A later trace query shows `data.session.feedback` as `useful`.
|
||||||
|
|
||||||
|
## Demo Story
|
||||||
|
|
||||||
|
The important interview story is:
|
||||||
|
|
||||||
|
```text
|
||||||
|
one session id
|
||||||
|
-> user question
|
||||||
|
-> multi-agent execution
|
||||||
|
-> evidence tools
|
||||||
|
-> verifier/self-evaluation
|
||||||
|
-> final answer
|
||||||
|
-> feedback
|
||||||
|
-> trace API for replay and audit
|
||||||
|
```
|
||||||
@@ -0,0 +1,39 @@
|
|||||||
|
# Payment Timeout Acceptance Case
|
||||||
|
|
||||||
|
## Goal
|
||||||
|
|
||||||
|
Validate that the MVP can diagnose a payment timeout incident and expose the complete trace for replay.
|
||||||
|
|
||||||
|
## Input
|
||||||
|
|
||||||
|
- Session id: `mvp-demo-payment-timeout-001`
|
||||||
|
- Question: `支付接口最近出现超时,请结合知识库、日志和指标判断可能原因,并给出修复建议。`
|
||||||
|
- Profile: `mvp-demo`
|
||||||
|
|
||||||
|
## Acceptance Criteria
|
||||||
|
|
||||||
|
1. Chat returns a successful answer with the same session id.
|
||||||
|
2. Trace API returns session metadata, final answer, ordered agent steps, and ordered tool invocations.
|
||||||
|
3. Trace contains enough evidence to explain which tools were used and whether verifier/self-evaluation was persisted.
|
||||||
|
4. Feedback can be submitted for the same session id.
|
||||||
|
5. A follow-up trace query shows the persisted feedback value.
|
||||||
|
|
||||||
|
## Trace Fields To Inspect
|
||||||
|
|
||||||
|
- `data.session.query`
|
||||||
|
- `data.session.answer`
|
||||||
|
- `data.session.selfEvaluation`
|
||||||
|
- `data.session.feedback`
|
||||||
|
- `data.steps[*].agentName`
|
||||||
|
- `data.steps[*].thought`
|
||||||
|
- `data.toolInvocations[*].toolName`
|
||||||
|
- `data.toolInvocations[*].inputParams`
|
||||||
|
- `data.toolInvocations[*].outputPreview`
|
||||||
|
- `data.toolInvocations[*].retrievalDetails`
|
||||||
|
- `data.summary`
|
||||||
|
|
||||||
|
## Known Limits
|
||||||
|
|
||||||
|
- This case is not a full offline test. It still requires valid infrastructure for chat, persistence, vector search, and model calls.
|
||||||
|
- Mock logs and metrics are enabled by the `mvp-demo` profile to make those evidence tools repeatable.
|
||||||
|
- Sensitive configuration cleanup is deferred by current MVP priority.
|
||||||
@@ -0,0 +1,268 @@
|
|||||||
|
# MVP Agent 工程决策记录
|
||||||
|
|
||||||
|
本文记录 MVP 实现过程中已经落地的一些关键修复、取舍和工程判断。目标不是写流水账,而是沉淀面试时可以讲清楚的 Agent 工程思路。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. 统一流式与非流式 Chat 主链路
|
||||||
|
|
||||||
|
### 背景
|
||||||
|
|
||||||
|
早期 `/api/chat` 和 `/api/chat_stream` 是两条不同实现:
|
||||||
|
|
||||||
|
- 非流式接口会走复杂度判断,并可能进入 Planner / Executor / Verifier 多 Agent 流程。
|
||||||
|
- 流式接口直接创建单个 ReactAgent,然后 `agent.stream()` 输出 token。
|
||||||
|
|
||||||
|
这导致两个接口表面都是 chat,实际能力不一致:流式接口不会进入 verifier、不会沉淀完整诊断链路,也不容易和 `diagnosis_session`、`tool_invocation` 对齐。
|
||||||
|
|
||||||
|
### 决策
|
||||||
|
|
||||||
|
将两个接口统一到同一条核心链路:
|
||||||
|
|
||||||
|
```text
|
||||||
|
getOrCreateSession
|
||||||
|
-> 读取会话历史
|
||||||
|
-> ChatService.executeChatWithStrategy(...)
|
||||||
|
-> 写回会话历史
|
||||||
|
```
|
||||||
|
|
||||||
|
接口差异只保留在传输层:
|
||||||
|
|
||||||
|
- `/api/chat` 返回完整 JSON。
|
||||||
|
- `/api/chat_stream` 通过 SSE 分块发送最终答案。
|
||||||
|
|
||||||
|
### 取舍
|
||||||
|
|
||||||
|
这样会牺牲原来的 token 级实时流式体验,但换来业务行为一致、诊断链路一致、Verifier 和 evidence trace 一致。
|
||||||
|
|
||||||
|
对 MVP 来说,优先保证“同一个问题不因接口不同而进入不同智能链路”,比 token 级流式更重要。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. 会话 ID 与诊断链路统一
|
||||||
|
|
||||||
|
### 背景
|
||||||
|
|
||||||
|
原实现中:
|
||||||
|
|
||||||
|
- `ChatController` 用前端传入的 `Id` 在 JVM 内存里维护历史消息。
|
||||||
|
- `ChatService` 每次执行又生成新的 8 位 sessionId,作为 `diagnosis_session` 和工具调用追踪 ID。
|
||||||
|
|
||||||
|
这会造成前端会话、后端诊断会话、工具证据链三者分裂。
|
||||||
|
|
||||||
|
### 决策
|
||||||
|
|
||||||
|
将前端 chat session id 作为后端诊断链路的主 session id:
|
||||||
|
|
||||||
|
- Redis `SessionContext` 保存聊天历史。
|
||||||
|
- `diagnosis_session.session_id` 复用同一个 id。
|
||||||
|
- `RunnableConfig.metadata.sessionId` 和 `SessionContextHolder` 也使用同一个 id。
|
||||||
|
- `tool_invocation`、`agent_step`、verifier evaluation 都可按同一 session id 串起来。
|
||||||
|
|
||||||
|
### 企业级意义
|
||||||
|
|
||||||
|
Agent 系统最怕“答得出来但查不清”。统一 session id 后,一次用户请求可以完整追踪:
|
||||||
|
|
||||||
|
```text
|
||||||
|
用户问题 -> Agent 步骤 -> 工具调用 -> Verifier 判断 -> 最终答案 -> 用户反馈
|
||||||
|
```
|
||||||
|
|
||||||
|
这是可观测、可审计、可复盘的基础。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. 引入统一 ToolInvocationRecorder
|
||||||
|
|
||||||
|
### 背景
|
||||||
|
|
||||||
|
Verifier 需要结构化证据链,但原实现只有 `lookup_knowledge` 主动写入 `tool_invocation`。
|
||||||
|
|
||||||
|
`query_logs`、`query_metrics` 虽然返回 JSON,但没有统一落库,导致 verifier 看不到日志、指标等 evidence tool 的稳定记录。
|
||||||
|
|
||||||
|
### 决策
|
||||||
|
|
||||||
|
新增 `ToolInvocationRecorder`,作为所有 evidence tool 的统一落库入口。
|
||||||
|
|
||||||
|
当前接入:
|
||||||
|
|
||||||
|
- `lookup_knowledge`
|
||||||
|
- `query_logs`
|
||||||
|
- `query_metrics`
|
||||||
|
|
||||||
|
记录字段包括:
|
||||||
|
|
||||||
|
- tool name
|
||||||
|
- input params
|
||||||
|
- output preview
|
||||||
|
- output length
|
||||||
|
- success
|
||||||
|
- error message
|
||||||
|
- duration
|
||||||
|
- trace id / domain details
|
||||||
|
|
||||||
|
### 企业级意义
|
||||||
|
|
||||||
|
这一步把 Agent 从“模型说它查过”推进到“系统能证明它查过”。
|
||||||
|
|
||||||
|
后续 verifier 不应该依赖模型自由文本回忆工具调用,而应该消费结构化 trace summary。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. Verifier 作为事实约束层
|
||||||
|
|
||||||
|
### 背景
|
||||||
|
|
||||||
|
普通 Agent 很容易在工具调用后直接生成答案,但企业场景更关心:
|
||||||
|
|
||||||
|
- 关键结论有没有证据
|
||||||
|
- 证据是直接证据还是间接支持
|
||||||
|
- 哪些事实缺口需要人工介入
|
||||||
|
- 工具失败时是否诚实降级
|
||||||
|
|
||||||
|
### 决策
|
||||||
|
|
||||||
|
保留 Planner / Executor / Verifier 三角色:
|
||||||
|
|
||||||
|
- Planner 负责拆解问题。
|
||||||
|
- Executor 负责执行查询与形成初稿。
|
||||||
|
- Verifier 负责基于 `tool_trace_summary` 做事实核查。
|
||||||
|
|
||||||
|
Verifier 输出结构化 JSON,包括:
|
||||||
|
|
||||||
|
- verdict
|
||||||
|
- groundedness_score
|
||||||
|
- critical_fact_count
|
||||||
|
- facts_checked
|
||||||
|
- rationale
|
||||||
|
|
||||||
|
### 取舍
|
||||||
|
|
||||||
|
Verifier 会增加一次模型调用成本,但换来可解释性和质量约束。对企业级 Agent 来说,这是值得的。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. 从手写编排切换到 SupervisorAgent
|
||||||
|
|
||||||
|
### 背景
|
||||||
|
|
||||||
|
之前 `ChatService.executeChatComplex()` 中构建了 `SupervisorAgent`,但实际仍然手写调用:
|
||||||
|
|
||||||
|
```text
|
||||||
|
planner -> executor -> verifier
|
||||||
|
```
|
||||||
|
|
||||||
|
这会造成代码与设计不一致,维护者容易误以为当前已经由 Supervisor 调度。
|
||||||
|
|
||||||
|
### 决策
|
||||||
|
|
||||||
|
复杂问题真正切换到 `SupervisorAgent.invoke(...)`。
|
||||||
|
|
||||||
|
Supervisor 负责路由:
|
||||||
|
|
||||||
|
```text
|
||||||
|
chat_supervisor -> chat_planner
|
||||||
|
chat_supervisor -> chat_executor
|
||||||
|
chat_supervisor -> chat_verifier
|
||||||
|
chat_supervisor -> FINISH
|
||||||
|
```
|
||||||
|
|
||||||
|
外层仍保留:
|
||||||
|
|
||||||
|
- verifier 输出解析
|
||||||
|
- PASS / LOW_CONFID / REJECT 判定
|
||||||
|
- retry context
|
||||||
|
- fallback
|
||||||
|
- evaluation 入库
|
||||||
|
|
||||||
|
### 验证
|
||||||
|
|
||||||
|
新增离线专项测试 `ChatServiceSupervisorAgentTest`,使用 scripted `ChatModel` 验证真实 SupervisorAgent 路由顺序,不依赖真实 LLM、MySQL、Redis。
|
||||||
|
|
||||||
|
### 企业级意义
|
||||||
|
|
||||||
|
这让项目不只是“自己写 if/else 多 Agent”,而是使用框架原生 multi-agent orchestration,同时保留业务层的质量门控。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. 文档上传路径语义统一
|
||||||
|
|
||||||
|
### 背景
|
||||||
|
|
||||||
|
上传文档时,`DocumentManagementService.saveToLocal()` 返回带 `knowledge_base` 前缀的路径。
|
||||||
|
|
||||||
|
而 `KnowledgeIndexService.readDocument()` 又执行:
|
||||||
|
|
||||||
|
```java
|
||||||
|
Paths.get(knowledgeBasePath, filePath)
|
||||||
|
```
|
||||||
|
|
||||||
|
这可能拼出:
|
||||||
|
|
||||||
|
```text
|
||||||
|
knowledge_base/knowledge_base/...
|
||||||
|
```
|
||||||
|
|
||||||
|
最终表现为 L0 命中文档,但读取原文失败。
|
||||||
|
|
||||||
|
### 决策
|
||||||
|
|
||||||
|
统一路径语义:
|
||||||
|
|
||||||
|
- 新上传文档存相对 `knowledge.base-path` 的路径,例如 `payment/runbook.md`。
|
||||||
|
- `readDocument()` 兼容新旧路径:
|
||||||
|
- 相对路径
|
||||||
|
- 已带 base path 的旧相对路径
|
||||||
|
- 绝对路径
|
||||||
|
|
||||||
|
### 企业级意义
|
||||||
|
|
||||||
|
知识库检索不能只看“命中”,还要保证命中后的内容可读、可引用、可追踪。
|
||||||
|
|
||||||
|
这是 RAG / Agent 系统里很典型的工程细节:检索质量问题不一定来自模型,也可能来自路径、元数据、索引和原文之间的语义不一致。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. MVP 阶段的优先级取舍
|
||||||
|
|
||||||
|
当前主动暂缓的问题:
|
||||||
|
|
||||||
|
- 敏感配置外置与密钥轮换
|
||||||
|
- CORS / Redis 反序列化安全边界
|
||||||
|
- 默认 `mvn test` 离线化
|
||||||
|
|
||||||
|
原因不是这些不重要,而是当前目标是先跑通并讲清楚 MVP Agent 工程闭环。
|
||||||
|
|
||||||
|
短期优先目标:
|
||||||
|
|
||||||
|
```text
|
||||||
|
可演示 -> 可观测 -> 可验证 -> 可复盘
|
||||||
|
```
|
||||||
|
|
||||||
|
安全和完整测试体系属于企业落地必须项,但可以在 MVP 主链路稳定后作为下一阶段补齐。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. 后续建议
|
||||||
|
|
||||||
|
下一阶段建议聚焦“可复现 MVP Demo”:
|
||||||
|
|
||||||
|
1. 增加 `local-demo` 或 `mvp-demo` profile。
|
||||||
|
2. 准备固定诊断 case,例如“支付接口超时”。
|
||||||
|
3. 提供一键初始化知识库样例。
|
||||||
|
4. 提供一键触发复杂诊断请求的脚本。
|
||||||
|
5. 增加 trace 查询接口:
|
||||||
|
|
||||||
|
```text
|
||||||
|
GET /api/diagnosis/{sessionId}/trace
|
||||||
|
```
|
||||||
|
|
||||||
|
该接口聚合:
|
||||||
|
|
||||||
|
- diagnosis_session
|
||||||
|
- agent_step
|
||||||
|
- tool_invocation
|
||||||
|
- verifier evaluation
|
||||||
|
- final answer
|
||||||
|
- feedback
|
||||||
|
|
||||||
|
这样 MVP 就能从“功能实现”升级为“企业级 Agent 工程作品”。
|
||||||
@@ -0,0 +1,39 @@
|
|||||||
|
# MVP Demo Profile 与 Trace 查询接口
|
||||||
|
|
||||||
|
## 背景
|
||||||
|
|
||||||
|
MVP 已经能跑多 Agent 诊断、工具调用、Verifier 和反馈,但对外展示时仍然缺少一个稳定的复盘入口。面试官或评审如果想确认一次 Agent 回答是否可信,不能只看最终答案,还需要看到用户原始问题、Agent 步骤顺序、工具调用证据、Verifier / self-evaluation、最终答案和用户反馈。
|
||||||
|
|
||||||
|
## 决策
|
||||||
|
|
||||||
|
新增 `mvp-demo` profile 和 trace 查询接口:
|
||||||
|
|
||||||
|
```text
|
||||||
|
GET /api/diagnosis/{sessionId}/trace
|
||||||
|
```
|
||||||
|
|
||||||
|
接口聚合:
|
||||||
|
|
||||||
|
- `diagnosis_session`
|
||||||
|
- `agent_step`
|
||||||
|
- `tool_invocation`
|
||||||
|
- `self_evaluation`
|
||||||
|
- `feedback`
|
||||||
|
|
||||||
|
同时在 `mvp/demo` 下沉淀端到端验收 case,把启动、提问、查 trace、提交 feedback 串成一条可演示路径。
|
||||||
|
|
||||||
|
## 取舍
|
||||||
|
|
||||||
|
`mvp-demo` profile 不是完整离线 mock 环境,仍然复用当前真实 DB / Redis / Milvus / LLM 配置,只显式打开日志和指标 mock。原因是当前阶段目标是展示企业级 Agent 工程闭环,不是隐藏真实集成复杂度。
|
||||||
|
|
||||||
|
这让 MVP 的讲述从“我实现了一个聊天接口”升级为:
|
||||||
|
|
||||||
|
```text
|
||||||
|
我实现了一条可执行、可观测、可验收、可复盘的 Agent 诊断链路。
|
||||||
|
```
|
||||||
|
|
||||||
|
## 面试表达
|
||||||
|
|
||||||
|
- 我没有把 trace 塞进 chat 返回值,而是做成独立只读观测接口,保持执行链路和观测链路解耦。
|
||||||
|
- Trace API 复用已经沉淀的 `diagnosis_session`、`agent_step`、`tool_invocation` 三张表,没有引入新的 schema 风险。
|
||||||
|
- Demo profile 只做最小 overlay,让日志和指标工具可重复,保留真实基础设施集成,方便说明 MVP 与生产化之间的差距。
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
mvp-demo-trace-acceptance committed on 2026-07-03
|
||||||
@@ -0,0 +1,2 @@
|
|||||||
|
schema: spec-driven
|
||||||
|
created: 2026-07-03
|
||||||
@@ -0,0 +1,29 @@
|
|||||||
|
{
|
||||||
|
"id": "mvp-demo-trace-acceptance",
|
||||||
|
"metadata": {
|
||||||
|
"status": "committed",
|
||||||
|
"created_at": "2026-07-03",
|
||||||
|
"updated_at": "2026-07-03",
|
||||||
|
"implementation_status": "implemented"
|
||||||
|
},
|
||||||
|
"summary": "Add an MVP demo profile, a read-only diagnosis trace API, and an end-to-end acceptance case.",
|
||||||
|
"artifacts": {
|
||||||
|
"proposal": "proposal.md",
|
||||||
|
"design": "design.md",
|
||||||
|
"tasks": "tasks.md",
|
||||||
|
"specs": [
|
||||||
|
"specs/mvp-demo-trace-acceptance/spec.md"
|
||||||
|
],
|
||||||
|
"devflow": "devflow/projects/2026-07-03-mvp-demo-trace-acceptance"
|
||||||
|
},
|
||||||
|
"tasks": [
|
||||||
|
"Add DiagnosisTraceResponse DTO",
|
||||||
|
"Add DiagnosisTraceService aggregation",
|
||||||
|
"Add DiagnosisTraceController endpoint",
|
||||||
|
"Add mvp-demo profile",
|
||||||
|
"Add MVP demo acceptance documentation",
|
||||||
|
"Add focused trace service tests",
|
||||||
|
"Run targeted verification and GitNexus change detection",
|
||||||
|
"Update MVP notes and devflow acceptance"
|
||||||
|
]
|
||||||
|
}
|
||||||
@@ -0,0 +1,70 @@
|
|||||||
|
## Context
|
||||||
|
|
||||||
|
The MVP already persists diagnosis execution data across three tables:
|
||||||
|
|
||||||
|
- `diagnosis_session`: query, status, answer, counts, feedback, and `self_evaluation`.
|
||||||
|
- `agent_step`: ordered agent execution records.
|
||||||
|
- `tool_invocation`: evidence tool calls and retrieval metadata.
|
||||||
|
|
||||||
|
Recent work unified chat session ids and persisted tool invocations, so a single session id can now connect user input, agent steps, evidence tools, verifier evaluation, final answer, and feedback. The missing piece is a read-only aggregation API and a documented demo profile/workflow that a reviewer can run without reading database tables manually.
|
||||||
|
|
||||||
|
## Goals / Non-Goals
|
||||||
|
|
||||||
|
**Goals:**
|
||||||
|
|
||||||
|
- Add a trace API that returns one aggregated view for a diagnosis session.
|
||||||
|
- Keep the trace API read-only and based on existing persistence tables.
|
||||||
|
- Add an `mvp-demo` profile that makes the demo intent explicit and keeps mock log/metric tools enabled.
|
||||||
|
- Add a documented end-to-end acceptance case for start, chat, trace query, and feedback.
|
||||||
|
- Add focused tests for trace aggregation.
|
||||||
|
|
||||||
|
**Non-Goals:**
|
||||||
|
|
||||||
|
- Do not clean up committed sensitive configuration in this change.
|
||||||
|
- Do not add database migrations.
|
||||||
|
- Do not alter `/api/chat`, `/api/chat_stream`, verifier routing, feedback, or document upload behavior.
|
||||||
|
- Do not create a fully offline fake LLM runtime.
|
||||||
|
|
||||||
|
## Decisions
|
||||||
|
|
||||||
|
| Decision | Choice | Alternative Considered | Rationale |
|
||||||
|
|---|---|---|---|
|
||||||
|
| Trace API shape | Add `GET /api/diagnosis/{sessionId}/trace` | Extend `/api/chat` response | Trace is an observability concern and should not make chat responses larger or change chat clients. |
|
||||||
|
| Aggregation ownership | New `DiagnosisTraceService` | Put aggregation in controller | Keeps controller thin and allows focused unit tests with mocked repositories. |
|
||||||
|
| Response DTO | Dedicated nested DTO | Return raw entities or maps | DTO avoids leaking JPA entity details and gives a stable demo-facing contract. |
|
||||||
|
| Missing session handling | Throw `SessionNotFoundException` and use existing global 404 handler | Return empty success payload | A missing trace is a real lookup miss and should be visible to callers. |
|
||||||
|
| `self_evaluation` handling | Return raw JSON string and best-effort parsed JSON | Parse only, or ignore parse failures | Raw value preserves evidence even if JSON shape evolves; parsed value improves frontend/demo readability. |
|
||||||
|
| Demo profile | Add `application-mvp-demo.yml` overlay | Change default `application.yml` | Overlay avoids disturbing current runtime and keeps demo choices explicit. |
|
||||||
|
|
||||||
|
## Interface Impact
|
||||||
|
|
||||||
|
- Level: L3 collaboration API.
|
||||||
|
- Reason: This adds a new HTTP endpoint and response contract intended for frontend/demo/reviewer consumption.
|
||||||
|
- Compatibility: Additive only. Existing callers do not need to change.
|
||||||
|
- Documentation: The endpoint is documented in the MVP demo acceptance case.
|
||||||
|
|
||||||
|
## Data Structures
|
||||||
|
|
||||||
|
The trace response contains:
|
||||||
|
|
||||||
|
- `session`: session id, query, status, flow, counts, timing, created/updated time, final answer, raw self-evaluation JSON, parsed self-evaluation object, and feedback.
|
||||||
|
- `steps`: ordered agent steps with step index, agent name, model input/output, thought, tool flag, duration, token count, and created time.
|
||||||
|
- `toolInvocations`: ordered tool records with id, step id, tool name, input params, output preview, retrieval metadata, duration, success, error, and created time.
|
||||||
|
- `summary`: counts derived from the returned collections and session fields.
|
||||||
|
|
||||||
|
## Risks / Trade-offs
|
||||||
|
|
||||||
|
- [Risk] Trace responses may become large for long sessions. -> Mitigation: the MVP returns persisted previews and structured metadata, not raw full external logs.
|
||||||
|
- [Risk] `self_evaluation` JSON shape may evolve. -> Mitigation: return both raw and best-effort parsed forms.
|
||||||
|
- [Risk] Demo profile still depends on real DB/Redis/Milvus/LLM. -> Mitigation: document prerequisites and keep mock logs/metrics enabled for repeatable tool evidence.
|
||||||
|
- [Risk] New endpoint becomes a de facto frontend contract. -> Mitigation: use a dedicated DTO and document L3 additive API impact.
|
||||||
|
|
||||||
|
## Migration Plan
|
||||||
|
|
||||||
|
- Deploying this change requires only application restart with the new code.
|
||||||
|
- No database migration is required.
|
||||||
|
- Rollback is deleting the new endpoint/profile/docs; persisted data remains unchanged.
|
||||||
|
|
||||||
|
## Open Questions
|
||||||
|
|
||||||
|
- None for this slice. Security and full offline test profile remain deferred by explicit user decision.
|
||||||
@@ -0,0 +1,29 @@
|
|||||||
|
## Why
|
||||||
|
|
||||||
|
The MVP can already execute multi-agent diagnosis, persist session traces, and collect feedback, but it is still hard to demonstrate as a complete enterprise-style workflow. A demo profile, a trace query API, and an explicit end-to-end acceptance case make the project runnable, observable, and explainable for interview and portfolio review.
|
||||||
|
|
||||||
|
## What Changes
|
||||||
|
|
||||||
|
- Add an `mvp-demo` Spring profile that keeps the existing external infrastructure contract but turns on mock log and metric providers for repeatable demonstrations.
|
||||||
|
- Add a read-only trace query API: `GET /api/diagnosis/{sessionId}/trace`.
|
||||||
|
- Aggregate `diagnosis_session`, `agent_step`, `tool_invocation`, verifier/self-evaluation, final answer, and feedback into one trace response.
|
||||||
|
- Add an end-to-end MVP acceptance case that documents startup, chat request, trace query, and feedback submission.
|
||||||
|
- Add focused service tests for trace aggregation without requiring MySQL, Redis, Milvus, or a real LLM.
|
||||||
|
- Record the design decision in MVP notes for interview storytelling.
|
||||||
|
|
||||||
|
## Capabilities
|
||||||
|
|
||||||
|
### New Capabilities
|
||||||
|
|
||||||
|
- `mvp-demo-trace-acceptance`: Covers the MVP demo profile, trace query API, and end-to-end acceptance workflow for a reproducible agent diagnosis demo.
|
||||||
|
|
||||||
|
### Modified Capabilities
|
||||||
|
|
||||||
|
- None.
|
||||||
|
|
||||||
|
## Impact
|
||||||
|
|
||||||
|
- Affected code: new trace controller/service/DTOs, `application-mvp-demo.yml`, unit tests, MVP demo documentation.
|
||||||
|
- Affected API: adds `GET /api/diagnosis/{sessionId}/trace`. This is an additive L3 collaboration API because it is intended for frontend, demo, and external reviewer consumption.
|
||||||
|
- Affected runtime behavior: no change to chat execution, verifier, feedback, document upload, or persistence semantics.
|
||||||
|
- Non-goals: no sensitive configuration cleanup, no database schema migration, no replacement of existing chat endpoints, no full offline mock LLM implementation.
|
||||||
+33
@@ -0,0 +1,33 @@
|
|||||||
|
## ADDED Requirements
|
||||||
|
|
||||||
|
### Requirement: Diagnosis trace can be queried by session id
|
||||||
|
The system SHALL expose a read-only HTTP endpoint `GET /api/diagnosis/{sessionId}/trace` that returns the persisted diagnosis trace for the requested session id.
|
||||||
|
|
||||||
|
#### Scenario: Existing session trace is returned
|
||||||
|
- **WHEN** a caller requests trace data for a session id that exists in `diagnosis_session`
|
||||||
|
- **THEN** the system returns a success response containing the session summary, ordered agent steps, ordered tool invocations, self-evaluation data, final answer, and feedback
|
||||||
|
|
||||||
|
#### Scenario: Missing session returns not found
|
||||||
|
- **WHEN** a caller requests trace data for a session id that does not exist in `diagnosis_session`
|
||||||
|
- **THEN** the system returns a 404 response using the existing session-not-found error contract
|
||||||
|
|
||||||
|
### Requirement: Trace aggregation is read-only
|
||||||
|
The system MUST build trace output from existing persisted diagnosis tables and MUST NOT mutate diagnosis sessions, agent steps, tool invocations, feedback, or chat session state while serving the trace request.
|
||||||
|
|
||||||
|
#### Scenario: Trace query does not change persisted state
|
||||||
|
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace`
|
||||||
|
- **THEN** the system reads `diagnosis_session`, `agent_step`, and `tool_invocation` records and returns an aggregate without saving any of those records
|
||||||
|
|
||||||
|
### Requirement: MVP demo profile is available
|
||||||
|
The system SHALL provide an `mvp-demo` Spring profile that documents the demo runtime intent and keeps mock log and metric providers enabled for repeatable diagnosis demonstrations.
|
||||||
|
|
||||||
|
#### Scenario: Demo profile loads mock evidence providers
|
||||||
|
- **WHEN** the application starts with `--spring.profiles.active=mvp-demo`
|
||||||
|
- **THEN** `prometheus.mock-enabled` and `cls.mock-enabled` are enabled by profile configuration
|
||||||
|
|
||||||
|
### Requirement: End-to-end MVP acceptance case is documented
|
||||||
|
The project SHALL include an end-to-end acceptance case that demonstrates start-up, chat diagnosis, trace query, and feedback submission using the same session id.
|
||||||
|
|
||||||
|
#### Scenario: Reviewer follows the acceptance case
|
||||||
|
- **WHEN** a reviewer follows the documented MVP demo acceptance steps
|
||||||
|
- **THEN** they can run the application, submit a diagnosis question, query the trace endpoint, and submit feedback for the same session id
|
||||||
@@ -0,0 +1,25 @@
|
|||||||
|
## 1. Trace Query API
|
||||||
|
|
||||||
|
- [x] 1.1 Add a `DiagnosisTraceResponse` DTO that represents session summary, ordered agent steps, ordered tool invocations, and derived summary counts.
|
||||||
|
- [x] 1.2 Add `DiagnosisTraceService` that loads `DiagnosisSession`, `AgentStep`, and `ToolInvocation` records by session id and builds the response.
|
||||||
|
- [x] 1.3 Add `DiagnosisTraceController` with `GET /api/diagnosis/{sessionId}/trace`.
|
||||||
|
- [x] 1.4 Return 404 through `SessionNotFoundException` when the requested diagnosis session does not exist.
|
||||||
|
|
||||||
|
## 2. Demo Profile And Acceptance Case
|
||||||
|
|
||||||
|
- [x] 2.1 Add `src/main/resources/application-mvp-demo.yml` with MVP demo profile overlays and mock logs/metrics enabled.
|
||||||
|
- [x] 2.2 Add `mvp/demo/README.md` documenting prerequisites, startup, chat request, trace query, and feedback submission.
|
||||||
|
- [x] 2.3 Add a concrete payment-timeout acceptance case with request/response expectations.
|
||||||
|
|
||||||
|
## 3. Tests And Verification
|
||||||
|
|
||||||
|
- [x] 3.1 Add focused unit tests for `DiagnosisTraceService` success and missing-session behavior.
|
||||||
|
- [x] 3.2 Run targeted tests for the new trace service.
|
||||||
|
- [x] 3.3 Run compile verification.
|
||||||
|
- [x] 3.4 Run GitNexus change detection before commit or handoff.
|
||||||
|
|
||||||
|
## 4. Notes And Flow Records
|
||||||
|
|
||||||
|
- [x] 4.1 Update MVP engineering notes with the demo/trace decision.
|
||||||
|
- [x] 4.2 Update OpenSpec tasks as work completes.
|
||||||
|
- [x] 4.3 Record verification results in devflow acceptance notes.
|
||||||
@@ -0,0 +1,37 @@
|
|||||||
|
## Purpose
|
||||||
|
|
||||||
|
Provide a repeatable MVP demo flow that can run a chat diagnosis, expose its persisted execution trace, and submit feedback for the same session id.
|
||||||
|
|
||||||
|
## Requirements
|
||||||
|
|
||||||
|
### Requirement: Diagnosis trace can be queried by session id
|
||||||
|
The system SHALL expose a read-only HTTP endpoint `GET /api/diagnosis/{sessionId}/trace` that returns the persisted diagnosis trace for the requested session id.
|
||||||
|
|
||||||
|
#### Scenario: Existing session trace is returned
|
||||||
|
- **WHEN** a caller requests trace data for a session id that exists in `diagnosis_session`
|
||||||
|
- **THEN** the system returns a success response containing the session summary, ordered agent steps, ordered tool invocations, self-evaluation data, final answer, and feedback
|
||||||
|
|
||||||
|
#### Scenario: Missing session returns not found
|
||||||
|
- **WHEN** a caller requests trace data for a session id that does not exist in `diagnosis_session`
|
||||||
|
- **THEN** the system returns a 404 response using the existing session-not-found error contract
|
||||||
|
|
||||||
|
### Requirement: Trace aggregation is read-only
|
||||||
|
The system MUST build trace output from existing persisted diagnosis tables and MUST NOT mutate diagnosis sessions, agent steps, tool invocations, feedback, or chat session state while serving the trace request.
|
||||||
|
|
||||||
|
#### Scenario: Trace query does not change persisted state
|
||||||
|
- **WHEN** a caller requests `GET /api/diagnosis/{sessionId}/trace`
|
||||||
|
- **THEN** the system reads `diagnosis_session`, `agent_step`, and `tool_invocation` records and returns an aggregate without saving any of those records
|
||||||
|
|
||||||
|
### Requirement: MVP demo profile is available
|
||||||
|
The system SHALL provide an `mvp-demo` Spring profile that documents the demo runtime intent and keeps mock log and metric providers enabled for repeatable diagnosis demonstrations.
|
||||||
|
|
||||||
|
#### Scenario: Demo profile loads mock evidence providers
|
||||||
|
- **WHEN** the application starts with `--spring.profiles.active=mvp-demo`
|
||||||
|
- **THEN** `prometheus.mock-enabled` and `cls.mock-enabled` are enabled by profile configuration
|
||||||
|
|
||||||
|
### Requirement: End-to-end MVP acceptance case is documented
|
||||||
|
The project SHALL include an end-to-end acceptance case that demonstrates start-up, chat diagnosis, trace query, and feedback submission using the same session id.
|
||||||
|
|
||||||
|
#### Scenario: Reviewer follows the acceptance case
|
||||||
|
- **WHEN** a reviewer follows the documented MVP demo acceptance steps
|
||||||
|
- **THEN** they can run the application, submit a diagnosis question, query the trace endpoint, and submit feedback for the same session id
|
||||||
@@ -0,0 +1,24 @@
|
|||||||
|
package com.superbiz.agent.controller;
|
||||||
|
|
||||||
|
import com.superbiz.agent.dto.DiagnosisTraceResponse;
|
||||||
|
import com.superbiz.agent.dto.Result;
|
||||||
|
import com.superbiz.agent.service.DiagnosisTraceService;
|
||||||
|
import lombok.RequiredArgsConstructor;
|
||||||
|
import org.springframework.http.ResponseEntity;
|
||||||
|
import org.springframework.web.bind.annotation.GetMapping;
|
||||||
|
import org.springframework.web.bind.annotation.PathVariable;
|
||||||
|
import org.springframework.web.bind.annotation.RequestMapping;
|
||||||
|
import org.springframework.web.bind.annotation.RestController;
|
||||||
|
|
||||||
|
@RestController
|
||||||
|
@RequestMapping("/api/diagnosis")
|
||||||
|
@RequiredArgsConstructor
|
||||||
|
public class DiagnosisTraceController {
|
||||||
|
|
||||||
|
private final DiagnosisTraceService diagnosisTraceService;
|
||||||
|
|
||||||
|
@GetMapping("/{sessionId}/trace")
|
||||||
|
public ResponseEntity<Result<DiagnosisTraceResponse>> getTrace(@PathVariable String sessionId) {
|
||||||
|
return ResponseEntity.ok(Result.success(diagnosisTraceService.getTrace(sessionId)));
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,102 @@
|
|||||||
|
package com.superbiz.agent.dto;
|
||||||
|
|
||||||
|
import lombok.AllArgsConstructor;
|
||||||
|
import lombok.Builder;
|
||||||
|
import lombok.Data;
|
||||||
|
import lombok.NoArgsConstructor;
|
||||||
|
|
||||||
|
import java.time.LocalDateTime;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.Map;
|
||||||
|
|
||||||
|
@Data
|
||||||
|
@Builder
|
||||||
|
@NoArgsConstructor
|
||||||
|
@AllArgsConstructor
|
||||||
|
public class DiagnosisTraceResponse {
|
||||||
|
|
||||||
|
private SessionTrace session;
|
||||||
|
private List<AgentStepTrace> steps;
|
||||||
|
private List<ToolInvocationTrace> toolInvocations;
|
||||||
|
private TraceSummary summary;
|
||||||
|
|
||||||
|
@Data
|
||||||
|
@Builder
|
||||||
|
@NoArgsConstructor
|
||||||
|
@AllArgsConstructor
|
||||||
|
public static class SessionTrace {
|
||||||
|
private Long id;
|
||||||
|
private String sessionId;
|
||||||
|
private String query;
|
||||||
|
private String status;
|
||||||
|
private String agentFlow;
|
||||||
|
private Integer totalDurationMs;
|
||||||
|
private Integer totalTokenCount;
|
||||||
|
private Integer stepCount;
|
||||||
|
private Integer toolCallCount;
|
||||||
|
private String answer;
|
||||||
|
private String selfEvaluationRaw;
|
||||||
|
private Map<String, Object> selfEvaluation;
|
||||||
|
private String feedback;
|
||||||
|
private LocalDateTime createdAt;
|
||||||
|
private LocalDateTime updatedAt;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Data
|
||||||
|
@Builder
|
||||||
|
@NoArgsConstructor
|
||||||
|
@AllArgsConstructor
|
||||||
|
public static class AgentStepTrace {
|
||||||
|
private Long id;
|
||||||
|
private String sessionId;
|
||||||
|
private Integer stepIndex;
|
||||||
|
private String agentName;
|
||||||
|
private String modelInput;
|
||||||
|
private String modelOutput;
|
||||||
|
private String thought;
|
||||||
|
private Boolean hasToolCall;
|
||||||
|
private Integer durationMs;
|
||||||
|
private Integer tokenCount;
|
||||||
|
private LocalDateTime createdAt;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Data
|
||||||
|
@Builder
|
||||||
|
@NoArgsConstructor
|
||||||
|
@AllArgsConstructor
|
||||||
|
public static class ToolInvocationTrace {
|
||||||
|
private Long id;
|
||||||
|
private String sessionId;
|
||||||
|
private Long stepId;
|
||||||
|
private String toolName;
|
||||||
|
private String inputParamsRaw;
|
||||||
|
private Map<String, Object> inputParams;
|
||||||
|
private String outputPreview;
|
||||||
|
private Integer outputLength;
|
||||||
|
private String retrievalLayer;
|
||||||
|
private Integer l0MatchCount;
|
||||||
|
private Integer l1MatchCount;
|
||||||
|
private Boolean truncated;
|
||||||
|
private String relevanceLevel;
|
||||||
|
private String dedupReason;
|
||||||
|
private String retrievalDetailsRaw;
|
||||||
|
private Map<String, Object> retrievalDetails;
|
||||||
|
private Integer durationMs;
|
||||||
|
private Boolean success;
|
||||||
|
private String errorMessage;
|
||||||
|
private LocalDateTime createdAt;
|
||||||
|
}
|
||||||
|
|
||||||
|
@Data
|
||||||
|
@Builder
|
||||||
|
@NoArgsConstructor
|
||||||
|
@AllArgsConstructor
|
||||||
|
public static class TraceSummary {
|
||||||
|
private int persistedStepCount;
|
||||||
|
private int returnedStepCount;
|
||||||
|
private int persistedToolCallCount;
|
||||||
|
private int returnedToolCallCount;
|
||||||
|
private boolean hasVerifierEvaluation;
|
||||||
|
private boolean hasFeedback;
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -416,6 +416,9 @@ public class ChatService {
|
|||||||
answer = extractStateText(stateOptional, "executor_feedback");
|
answer = extractStateText(stateOptional, "executor_feedback");
|
||||||
VerifierContextHolder.setExecutorFinalAnswer(answer);
|
VerifierContextHolder.setExecutorFinalAnswer(answer);
|
||||||
String verifierOutput = extractStateText(stateOptional, "verifier_output");
|
String verifierOutput = extractStateText(stateOptional, "verifier_output");
|
||||||
|
if ((verifierOutput == null || verifierOutput.isBlank()) && answer != null && !answer.isBlank()) {
|
||||||
|
verifierOutput = invokeVerifierFallback(verifier, question, round, config);
|
||||||
|
}
|
||||||
finalDecision = parseVerifierDecision(verifierOutput, round);
|
finalDecision = parseVerifierDecision(verifierOutput, round);
|
||||||
logger.debug("Supervisor round {} finished: plannerPlanLength={}, answerLength={}, verifierOutputLength={}",
|
logger.debug("Supervisor round {} finished: plannerPlanLength={}, answerLength={}, verifierOutputLength={}",
|
||||||
round,
|
round,
|
||||||
@@ -588,6 +591,17 @@ public class ChatService {
|
|||||||
return input.toString();
|
return input.toString();
|
||||||
}
|
}
|
||||||
|
|
||||||
|
private String invokeVerifierFallback(ReactAgent verifier, String question, int round, RunnableConfig config) {
|
||||||
|
try {
|
||||||
|
logger.warn("Supervisor round {} finished without verifier_output, invoking chat_verifier fallback", round);
|
||||||
|
return verifier.call("请基于 executor_final_answer 和 tool_trace_summary 输出 verifier JSON。原始问题:" + question, config)
|
||||||
|
.getText();
|
||||||
|
} catch (Exception e) {
|
||||||
|
logger.error("chat_verifier fallback 执行失败", e);
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
private VerifierDecision parseVerifierDecision(String verifierOutput, int round) {
|
private VerifierDecision parseVerifierDecision(String verifierOutput, int round) {
|
||||||
if (verifierOutput == null || verifierOutput.isBlank()) {
|
if (verifierOutput == null || verifierOutput.isBlank()) {
|
||||||
return null;
|
return null;
|
||||||
|
|||||||
@@ -0,0 +1,136 @@
|
|||||||
|
package com.superbiz.agent.service;
|
||||||
|
|
||||||
|
import com.fasterxml.jackson.core.type.TypeReference;
|
||||||
|
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||||
|
import com.superbiz.agent.domain.entity.AgentStep;
|
||||||
|
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||||
|
import com.superbiz.agent.domain.entity.ToolInvocation;
|
||||||
|
import com.superbiz.agent.dto.DiagnosisTraceResponse;
|
||||||
|
import com.superbiz.agent.exception.SessionNotFoundException;
|
||||||
|
import com.superbiz.agent.repository.AgentStepRepository;
|
||||||
|
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
||||||
|
import com.superbiz.agent.repository.ToolInvocationRepository;
|
||||||
|
import lombok.RequiredArgsConstructor;
|
||||||
|
import org.springframework.stereotype.Service;
|
||||||
|
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.Map;
|
||||||
|
|
||||||
|
@Service
|
||||||
|
@RequiredArgsConstructor
|
||||||
|
public class DiagnosisTraceService {
|
||||||
|
|
||||||
|
private static final TypeReference<Map<String, Object>> JSON_MAP_TYPE = new TypeReference<>() {
|
||||||
|
};
|
||||||
|
|
||||||
|
private final DiagnosisSessionRepository diagnosisSessionRepository;
|
||||||
|
private final AgentStepRepository agentStepRepository;
|
||||||
|
private final ToolInvocationRepository toolInvocationRepository;
|
||||||
|
private final ObjectMapper objectMapper;
|
||||||
|
|
||||||
|
public DiagnosisTraceResponse getTrace(String sessionId) {
|
||||||
|
DiagnosisSession session = diagnosisSessionRepository.findBySessionId(sessionId)
|
||||||
|
.orElseThrow(() -> new SessionNotFoundException(sessionId));
|
||||||
|
List<AgentStep> steps = agentStepRepository.findBySessionIdOrderByStepIndex(sessionId);
|
||||||
|
List<ToolInvocation> toolInvocations = toolInvocationRepository.findBySessionIdOrderByIdAsc(sessionId);
|
||||||
|
|
||||||
|
return DiagnosisTraceResponse.builder()
|
||||||
|
.session(toSessionTrace(session))
|
||||||
|
.steps(steps.stream().map(this::toAgentStepTrace).toList())
|
||||||
|
.toolInvocations(toolInvocations.stream().map(this::toToolInvocationTrace).toList())
|
||||||
|
.summary(toSummary(session, steps, toolInvocations))
|
||||||
|
.build();
|
||||||
|
}
|
||||||
|
|
||||||
|
private DiagnosisTraceResponse.SessionTrace toSessionTrace(DiagnosisSession session) {
|
||||||
|
return DiagnosisTraceResponse.SessionTrace.builder()
|
||||||
|
.id(session.getId())
|
||||||
|
.sessionId(session.getSessionId())
|
||||||
|
.query(session.getQuery())
|
||||||
|
.status(session.getStatus())
|
||||||
|
.agentFlow(session.getAgentFlow())
|
||||||
|
.totalDurationMs(session.getTotalDurationMs())
|
||||||
|
.totalTokenCount(session.getTotalTokenCount())
|
||||||
|
.stepCount(session.getStepCount())
|
||||||
|
.toolCallCount(session.getToolCallCount())
|
||||||
|
.answer(session.getAnswer())
|
||||||
|
.selfEvaluationRaw(session.getSelfEvaluation())
|
||||||
|
.selfEvaluation(parseJsonObject(session.getSelfEvaluation()))
|
||||||
|
.feedback(session.getFeedback())
|
||||||
|
.createdAt(session.getCreatedAt())
|
||||||
|
.updatedAt(session.getUpdatedAt())
|
||||||
|
.build();
|
||||||
|
}
|
||||||
|
|
||||||
|
private DiagnosisTraceResponse.AgentStepTrace toAgentStepTrace(AgentStep step) {
|
||||||
|
return DiagnosisTraceResponse.AgentStepTrace.builder()
|
||||||
|
.id(step.getId())
|
||||||
|
.sessionId(step.getSessionId())
|
||||||
|
.stepIndex(step.getStepIndex())
|
||||||
|
.agentName(step.getAgentName())
|
||||||
|
.modelInput(step.getModelInput())
|
||||||
|
.modelOutput(step.getModelOutput())
|
||||||
|
.thought(step.getThought())
|
||||||
|
.hasToolCall(step.getHasToolCall())
|
||||||
|
.durationMs(step.getDurationMs())
|
||||||
|
.tokenCount(step.getTokenCount())
|
||||||
|
.createdAt(step.getCreatedAt())
|
||||||
|
.build();
|
||||||
|
}
|
||||||
|
|
||||||
|
private DiagnosisTraceResponse.ToolInvocationTrace toToolInvocationTrace(ToolInvocation invocation) {
|
||||||
|
return DiagnosisTraceResponse.ToolInvocationTrace.builder()
|
||||||
|
.id(invocation.getId())
|
||||||
|
.sessionId(invocation.getSessionId())
|
||||||
|
.stepId(invocation.getStepId())
|
||||||
|
.toolName(invocation.getToolName())
|
||||||
|
.inputParamsRaw(invocation.getInputParams())
|
||||||
|
.inputParams(parseJsonObject(invocation.getInputParams()))
|
||||||
|
.outputPreview(invocation.getOutputPreview())
|
||||||
|
.outputLength(invocation.getOutputLength())
|
||||||
|
.retrievalLayer(invocation.getRetrievalLayer())
|
||||||
|
.l0MatchCount(invocation.getL0MatchCount())
|
||||||
|
.l1MatchCount(invocation.getL1MatchCount())
|
||||||
|
.truncated(invocation.getIsTruncated())
|
||||||
|
.relevanceLevel(invocation.getRelevanceLevel())
|
||||||
|
.dedupReason(invocation.getDedupReason())
|
||||||
|
.retrievalDetailsRaw(invocation.getRetrievalDetails())
|
||||||
|
.retrievalDetails(parseJsonObject(invocation.getRetrievalDetails()))
|
||||||
|
.durationMs(invocation.getDurationMs())
|
||||||
|
.success(invocation.getSuccess())
|
||||||
|
.errorMessage(invocation.getErrorMessage())
|
||||||
|
.createdAt(invocation.getCreatedAt())
|
||||||
|
.build();
|
||||||
|
}
|
||||||
|
|
||||||
|
private DiagnosisTraceResponse.TraceSummary toSummary(
|
||||||
|
DiagnosisSession session,
|
||||||
|
List<AgentStep> steps,
|
||||||
|
List<ToolInvocation> toolInvocations
|
||||||
|
) {
|
||||||
|
Map<String, Object> selfEvaluation = parseJsonObject(session.getSelfEvaluation());
|
||||||
|
return DiagnosisTraceResponse.TraceSummary.builder()
|
||||||
|
.persistedStepCount(defaultInt(session.getStepCount()))
|
||||||
|
.returnedStepCount(steps.size())
|
||||||
|
.persistedToolCallCount(defaultInt(session.getToolCallCount()))
|
||||||
|
.returnedToolCallCount(toolInvocations.size())
|
||||||
|
.hasVerifierEvaluation(selfEvaluation != null && selfEvaluation.containsKey("verifier_evaluation"))
|
||||||
|
.hasFeedback(session.getFeedback() != null && !session.getFeedback().isBlank())
|
||||||
|
.build();
|
||||||
|
}
|
||||||
|
|
||||||
|
private Map<String, Object> parseJsonObject(String json) {
|
||||||
|
if (json == null || json.isBlank()) {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
try {
|
||||||
|
return objectMapper.readValue(json, JSON_MAP_TYPE);
|
||||||
|
} catch (Exception ignored) {
|
||||||
|
return null;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private int defaultInt(Integer value) {
|
||||||
|
return value == null ? 0 : value;
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,25 @@
|
|||||||
|
spring:
|
||||||
|
config:
|
||||||
|
activate:
|
||||||
|
on-profile: mvp-demo
|
||||||
|
|
||||||
|
server:
|
||||||
|
port: 9900
|
||||||
|
|
||||||
|
prometheus:
|
||||||
|
mock-enabled: true
|
||||||
|
timeout: 5
|
||||||
|
|
||||||
|
cls:
|
||||||
|
mock-enabled: true
|
||||||
|
|
||||||
|
logging:
|
||||||
|
level:
|
||||||
|
root: INFO
|
||||||
|
com.superbiz.agent: DEBUG
|
||||||
|
com.alibaba.cloud: INFO
|
||||||
|
|
||||||
|
mvp:
|
||||||
|
demo:
|
||||||
|
name: payment-timeout-trace
|
||||||
|
description: Repeatable MVP flow for chat diagnosis, tool evidence, verifier evaluation, trace query, and feedback.
|
||||||
@@ -85,6 +85,43 @@ class ChatServiceSupervisorAgentTest {
|
|||||||
assertEquals(List.of("chat_planner", "chat_executor", "chat_verifier", "FINISH"), chatModel.decisions);
|
assertEquals(List.of("chat_planner", "chat_executor", "chat_verifier", "FINISH"), chatModel.decisions);
|
||||||
}
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void executeChatComplexInvokesVerifierFallbackWhenSupervisorSkipsVerifier() throws Exception {
|
||||||
|
ChatService chatService = createChatService();
|
||||||
|
ScriptedChatModel chatModel = new ScriptedChatModel(
|
||||||
|
List.of("chat_planner", "chat_executor", "FINISH"),
|
||||||
|
"""
|
||||||
|
{
|
||||||
|
"verdict": "PASS",
|
||||||
|
"groundedness_score": 1.0,
|
||||||
|
"critical_fact_count": 1,
|
||||||
|
"facts_checked": [
|
||||||
|
{
|
||||||
|
"fact": "executor answer generated",
|
||||||
|
"is_critical": true,
|
||||||
|
"verification": "direct_evidence",
|
||||||
|
"detail": "covered by fallback verifier",
|
||||||
|
"evidence_refs": []
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"rationale": "fallback verifier pass"
|
||||||
|
}
|
||||||
|
"""
|
||||||
|
);
|
||||||
|
|
||||||
|
ChatService.ChatResult result = chatService.executeChatComplex(
|
||||||
|
chatModel,
|
||||||
|
new ToolCallback[0],
|
||||||
|
"请分析订单支付超时的原因,并给出修复建议",
|
||||||
|
List.of(),
|
||||||
|
"supervisor-verifier-fallback-session"
|
||||||
|
);
|
||||||
|
|
||||||
|
assertEquals("EXECUTOR_FINAL_ANSWER", result.answer());
|
||||||
|
assertEquals(List.of("chat_planner", "chat_executor", "FINISH"), chatModel.decisions);
|
||||||
|
assertTrue(chatModel.sawVerifierPrompt);
|
||||||
|
}
|
||||||
|
|
||||||
@Test
|
@Test
|
||||||
void buildMethodToolsArrayIncludesLogsAndMetricsWhenAvailable() {
|
void buildMethodToolsArrayIncludesLogsAndMetricsWhenAvailable() {
|
||||||
ChatService chatService = new ChatService();
|
ChatService chatService = new ChatService();
|
||||||
@@ -153,7 +190,7 @@ class ChatServiceSupervisorAgentTest {
|
|||||||
}
|
}
|
||||||
|
|
||||||
private static final class ScriptedChatModel implements ChatModel {
|
private static final class ScriptedChatModel implements ChatModel {
|
||||||
private final List<String> decisionScript = List.of("chat_planner", "chat_executor", "chat_verifier", "FINISH");
|
private final List<String> decisionScript;
|
||||||
private final java.util.ArrayList<String> decisions = new java.util.ArrayList<>();
|
private final java.util.ArrayList<String> decisions = new java.util.ArrayList<>();
|
||||||
private int decisionIndex;
|
private int decisionIndex;
|
||||||
private String promptText = "";
|
private String promptText = "";
|
||||||
@@ -161,7 +198,7 @@ class ChatServiceSupervisorAgentTest {
|
|||||||
private final String verifierOutput;
|
private final String verifierOutput;
|
||||||
|
|
||||||
private ScriptedChatModel() {
|
private ScriptedChatModel() {
|
||||||
this("""
|
this(List.of("chat_planner", "chat_executor", "chat_verifier", "FINISH"), """
|
||||||
{
|
{
|
||||||
"verdict": "PASS",
|
"verdict": "PASS",
|
||||||
"groundedness_score": 1.0,
|
"groundedness_score": 1.0,
|
||||||
@@ -181,6 +218,11 @@ class ChatServiceSupervisorAgentTest {
|
|||||||
}
|
}
|
||||||
|
|
||||||
private ScriptedChatModel(String verifierOutput) {
|
private ScriptedChatModel(String verifierOutput) {
|
||||||
|
this(List.of("chat_planner", "chat_executor", "chat_verifier", "FINISH"), verifierOutput);
|
||||||
|
}
|
||||||
|
|
||||||
|
private ScriptedChatModel(List<String> decisionScript, String verifierOutput) {
|
||||||
|
this.decisionScript = decisionScript;
|
||||||
this.verifierOutput = verifierOutput;
|
this.verifierOutput = verifierOutput;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,119 @@
|
|||||||
|
package com.superbiz.agent.service;
|
||||||
|
|
||||||
|
import com.fasterxml.jackson.databind.ObjectMapper;
|
||||||
|
import com.superbiz.agent.domain.entity.AgentStep;
|
||||||
|
import com.superbiz.agent.domain.entity.DiagnosisSession;
|
||||||
|
import com.superbiz.agent.domain.entity.ToolInvocation;
|
||||||
|
import com.superbiz.agent.dto.DiagnosisTraceResponse;
|
||||||
|
import com.superbiz.agent.exception.SessionNotFoundException;
|
||||||
|
import com.superbiz.agent.repository.AgentStepRepository;
|
||||||
|
import com.superbiz.agent.repository.DiagnosisSessionRepository;
|
||||||
|
import com.superbiz.agent.repository.ToolInvocationRepository;
|
||||||
|
import org.junit.jupiter.api.Test;
|
||||||
|
|
||||||
|
import java.time.LocalDateTime;
|
||||||
|
import java.util.List;
|
||||||
|
import java.util.Optional;
|
||||||
|
|
||||||
|
import static org.junit.jupiter.api.Assertions.*;
|
||||||
|
import static org.mockito.Mockito.*;
|
||||||
|
|
||||||
|
class DiagnosisTraceServiceTest {
|
||||||
|
|
||||||
|
private final DiagnosisSessionRepository diagnosisSessionRepository = mock(DiagnosisSessionRepository.class);
|
||||||
|
private final AgentStepRepository agentStepRepository = mock(AgentStepRepository.class);
|
||||||
|
private final ToolInvocationRepository toolInvocationRepository = mock(ToolInvocationRepository.class);
|
||||||
|
private final DiagnosisTraceService service = new DiagnosisTraceService(
|
||||||
|
diagnosisSessionRepository,
|
||||||
|
agentStepRepository,
|
||||||
|
toolInvocationRepository,
|
||||||
|
new ObjectMapper()
|
||||||
|
);
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void getTraceAggregatesSessionStepsAndTools() {
|
||||||
|
String sessionId = "trace-session-001";
|
||||||
|
LocalDateTime now = LocalDateTime.of(2026, 7, 3, 14, 30);
|
||||||
|
DiagnosisSession session = DiagnosisSession.builder()
|
||||||
|
.id(1L)
|
||||||
|
.sessionId(sessionId)
|
||||||
|
.query("payment timeout")
|
||||||
|
.status("SUCCESS")
|
||||||
|
.agentFlow("COMPLEX")
|
||||||
|
.totalDurationMs(1200)
|
||||||
|
.totalTokenCount(300)
|
||||||
|
.stepCount(2)
|
||||||
|
.toolCallCount(1)
|
||||||
|
.answer("restart payment gateway pool")
|
||||||
|
.selfEvaluation("{\"verifier_evaluation\":{\"verdict\":\"PASS\"}}")
|
||||||
|
.feedback("useful")
|
||||||
|
.createdAt(now)
|
||||||
|
.updatedAt(now)
|
||||||
|
.build();
|
||||||
|
AgentStep step = AgentStep.builder()
|
||||||
|
.id(10L)
|
||||||
|
.sessionId(sessionId)
|
||||||
|
.stepIndex(1)
|
||||||
|
.agentName("chat_executor")
|
||||||
|
.modelInput("input")
|
||||||
|
.modelOutput("output")
|
||||||
|
.thought("executor finished")
|
||||||
|
.hasToolCall(true)
|
||||||
|
.durationMs(500)
|
||||||
|
.tokenCount(100)
|
||||||
|
.createdAt(now)
|
||||||
|
.build();
|
||||||
|
ToolInvocation invocation = ToolInvocation.builder()
|
||||||
|
.id(20L)
|
||||||
|
.sessionId(sessionId)
|
||||||
|
.stepId(10L)
|
||||||
|
.toolName("lookup_knowledge")
|
||||||
|
.inputParams("{\"query\":\"ERR_TIMEOUT\"}")
|
||||||
|
.outputPreview("payment timeout doc")
|
||||||
|
.outputLength(19)
|
||||||
|
.retrievalLayer("L0")
|
||||||
|
.l0MatchCount(1)
|
||||||
|
.l1MatchCount(0)
|
||||||
|
.isTruncated(false)
|
||||||
|
.relevanceLevel("HIGHLY_RELEVANT")
|
||||||
|
.dedupReason("FIRST_HIT")
|
||||||
|
.retrievalDetails("{\"documents\":[\"payment-errors.md\"]}")
|
||||||
|
.durationMs(80)
|
||||||
|
.success(true)
|
||||||
|
.createdAt(now)
|
||||||
|
.build();
|
||||||
|
|
||||||
|
when(diagnosisSessionRepository.findBySessionId(sessionId)).thenReturn(Optional.of(session));
|
||||||
|
when(agentStepRepository.findBySessionIdOrderByStepIndex(sessionId)).thenReturn(List.of(step));
|
||||||
|
when(toolInvocationRepository.findBySessionIdOrderByIdAsc(sessionId)).thenReturn(List.of(invocation));
|
||||||
|
|
||||||
|
DiagnosisTraceResponse response = service.getTrace(sessionId);
|
||||||
|
|
||||||
|
assertEquals(sessionId, response.getSession().getSessionId());
|
||||||
|
assertEquals("payment timeout", response.getSession().getQuery());
|
||||||
|
assertEquals("PASS", ((java.util.Map<?, ?>) response.getSession()
|
||||||
|
.getSelfEvaluation()
|
||||||
|
.get("verifier_evaluation")).get("verdict"));
|
||||||
|
assertEquals(1, response.getSteps().size());
|
||||||
|
assertEquals("chat_executor", response.getSteps().get(0).getAgentName());
|
||||||
|
assertEquals(1, response.getToolInvocations().size());
|
||||||
|
assertEquals("ERR_TIMEOUT", response.getToolInvocations().get(0).getInputParams().get("query"));
|
||||||
|
assertEquals(2, response.getSummary().getPersistedStepCount());
|
||||||
|
assertEquals(1, response.getSummary().getReturnedStepCount());
|
||||||
|
assertEquals(1, response.getSummary().getPersistedToolCallCount());
|
||||||
|
assertEquals(1, response.getSummary().getReturnedToolCallCount());
|
||||||
|
assertTrue(response.getSummary().isHasVerifierEvaluation());
|
||||||
|
assertTrue(response.getSummary().isHasFeedback());
|
||||||
|
}
|
||||||
|
|
||||||
|
@Test
|
||||||
|
void getTraceThrowsWhenSessionMissing() {
|
||||||
|
String sessionId = "missing-session";
|
||||||
|
when(diagnosisSessionRepository.findBySessionId(sessionId)).thenReturn(Optional.empty());
|
||||||
|
|
||||||
|
assertThrows(SessionNotFoundException.class, () -> service.getTrace(sessionId));
|
||||||
|
|
||||||
|
verify(diagnosisSessionRepository).findBySessionId(sessionId);
|
||||||
|
verifyNoInteractions(agentStepRepository, toolInvocationRepository);
|
||||||
|
}
|
||||||
|
}
|
||||||
Reference in New Issue
Block a user