feat(trace): add diagnosis trace workbench

This commit is contained in:
aruo
2026-07-07 01:24:56 +08:00
parent b3315ead52
commit aa035b828c
9 changed files with 1505 additions and 2 deletions
@@ -0,0 +1,70 @@
# Trace UI workbench design
## Subject and audience
Subject: a single-session diagnosis trace audit workbench.
Audience: backend and AIOps engineers reviewing an MVP diagnosis run after a
demo, incident drill, or regression check.
Single job: turn one trace response into an inspectable ledger that exposes
agent flow, evidence, verifier judgement, and skill boundaries without reading
raw JSON first.
## Frontend approach
- Static `trace.html`, `trace.css`, and `trace.js`.
- No dependency on a package manager, bundler, external fonts, or remote icons.
- Fetches `/api/diagnosis/{sessionId}/trace` and renders client-side.
- Accepts `?sessionId=...` and keeps the loaded session id in the URL.
## Visual direction
Palette:
- `#111827` ink rail for the trace frame.
- `#f7f3ea` warm ledger surface.
- `#1f7a6b` evidence green.
- `#b5472f` rejection red.
- `#c58a19` warning amber.
- `#5b6472` operational gray.
Type:
- UI/body: `Inter, ui-sans-serif, system-ui`.
- Data and labels: `ui-monospace, SFMono-Regular, Consolas`.
Layout:
```text
+--------------------------------------------------------------------+
| Session id input | Load | status / verdict / counts strip |
+------------------+----------------------+--------------------------+
| Agent trace rail | Evidence tool ledger | Inspector |
| ordered steps | filters + calls | verifier + RAG + skill |
+------------------+----------------------+--------------------------+
```
Signature element: a trace rail that treats each agent step as a ledger entry
with ordered markers, duration, token count, and expandable raw model excerpts.
## Data mapping
- `data.session`: summary strip, final answer, self-evaluation source.
- `data.steps`: ordered trace rail.
- `data.toolInvocations`: evidence ledger and selected tool inspector.
- `data.session.selfEvaluation.verifier_evaluation`: verdict, groundedness,
facts table, and verifier `tool_trace_summary`.
- `data.toolInvocations[*].retrievalDetails`: RAG query transform, retrieval
trace, context pack, rerank trace, and evidence blocks.
- `steps[*].modelOutput`: best-effort extraction of `selected_skill`.
- `steps[*].modelInput/modelOutput` and tool names: best-effort skill boundary
checks for `read_skill`.
## Accessibility and states
- Keyboard-focusable controls.
- Loading, missing session id, API error, empty steps, empty tools, and missing
verifier states.
- Responsive three-pane desktop layout that stacks on narrow screens.
- Respects `prefers-reduced-motion`.
@@ -0,0 +1,37 @@
# Trace UI workbench
## Why
The MVP already persists diagnosis trace data and exposes it through
`GET /api/diagnosis/{sessionId}/trace`, but reviewers still need to inspect raw
JSON to answer basic audit questions:
- Which agents ran, in what order, and with what persisted counts?
- Which evidence tools executed, how many times, and did they succeed?
- Which facts did the Verifier check, and which evidence refs support them?
- Did Planner only select skill metadata while Executor handled skill loading?
- What RAG retrieval details were used by `lookup_knowledge`?
This slows down demo review and makes trace quality issues harder to spot.
## What changes
- Add a static Trace workbench page served by Spring Boot static resources.
- Load a diagnosis trace by session id through the existing read-only Trace API.
- Render session summary, agent timeline, evidence tool ledger, verifier facts,
skill boundary checks, and RAG retrieval details.
- Add a navigation entry from the existing chat page to the Trace workbench.
## Non-goals
- No new backend endpoint.
- No mutation of diagnosis sessions, agent steps, tool invocations, or feedback.
- No new frontend framework or build pipeline.
- No change to the persisted trace schema.
## Impact
- Frontend-only runtime surface under `src/main/resources/static`.
- Uses the existing `Result<T>` API response contract.
- Works with existing trace records, including sessions that lack verifier or RAG
detail fields.
@@ -0,0 +1,36 @@
## ADDED Requirements
### Requirement: MVP demo SHALL provide a browser trace workbench
The MVP demo SHALL provide a browser-accessible static page for inspecting one
diagnosis trace by session id using the existing read-only Trace API.
#### Scenario: Existing trace renders in the workbench
- **WHEN** a reviewer opens the Trace workbench with a session id that exists
- **THEN** the page SHALL request `GET /api/diagnosis/{sessionId}/trace`
- **AND** it SHALL render session summary, ordered agent steps, ordered tool
invocations, verifier evaluation, and final answer when present
#### Scenario: Trace workbench handles missing or failed traces
- **WHEN** the Trace API returns an error or the session id is empty
- **THEN** the page SHALL show a clear error or empty state without mutating any
diagnosis data
### Requirement: Trace workbench SHALL expose evidence and skill boundaries
The Trace workbench SHALL make tool evidence and skill-loading boundaries visible
without requiring raw JSON inspection first.
#### Scenario: Tool and verifier evidence are inspectable
- **WHEN** a trace includes tool invocations and verifier facts
- **THEN** the page SHALL show tool invocation counts, success status, tool
filters, verifier facts, and evidence references
#### Scenario: RAG details are inspectable for lookup knowledge calls
- **WHEN** a `lookup_knowledge` invocation includes retrieval details
- **THEN** the page SHALL show query transform, retrieval trace, context pack,
rerank trace, and evidence block data where available
#### Scenario: Skill boundary checks are visible
- **WHEN** a trace includes planner, executor, or verifier steps
- **THEN** the page SHALL show best-effort indicators for selected skill,
planner `read_skill` text mentions, executor `read_skill` text mentions, and
verifier `read_skill` text mentions
@@ -0,0 +1,9 @@
# Tasks
- [x] 1. Add OpenSpec delta for the Trace UI workbench.
- [x] 2. Add static Trace workbench HTML/CSS/JS.
- [x] 3. Add a chat-page navigation entry to the Trace workbench.
- [x] 4. Verify Java compilation and static page syntax.
- [x] 5. Validate the page against a real trace endpoint when a local service is available.
- [x] 6. Archive the OpenSpec change after validation.
- [x] 7. Commit the completed change.
@@ -1,9 +1,7 @@
## Purpose
Provide a repeatable MVP demo flow that can run a chat diagnosis, expose its persisted execution trace, and submit feedback for the same session id.
## Requirements
### Requirement: Diagnosis trace can be queried by session id
The system SHALL expose a read-only HTTP endpoint `GET /api/diagnosis/{sessionId}/trace` that returns the persisted diagnosis trace for the requested session id.
@@ -64,3 +62,39 @@ The MVP demo SHALL document which trace fields to inspect for evidence, verifier
#### Scenario: Checklist maps fields to interview claims
- **WHEN** a developer reviews a trace response
- **THEN** the checklist SHALL map concrete JSON paths to the claims made in the interview walkthrough
### Requirement: MVP demo SHALL provide a browser trace workbench
The MVP demo SHALL provide a browser-accessible static page for inspecting one
diagnosis trace by session id using the existing read-only Trace API.
#### Scenario: Existing trace renders in the workbench
- **WHEN** a reviewer opens the Trace workbench with a session id that exists
- **THEN** the page SHALL request `GET /api/diagnosis/{sessionId}/trace`
- **AND** it SHALL render session summary, ordered agent steps, ordered tool
invocations, verifier evaluation, and final answer when present
#### Scenario: Trace workbench handles missing or failed traces
- **WHEN** the Trace API returns an error or the session id is empty
- **THEN** the page SHALL show a clear error or empty state without mutating any
diagnosis data
### Requirement: Trace workbench SHALL expose evidence and skill boundaries
The Trace workbench SHALL make tool evidence and skill-loading boundaries visible
without requiring raw JSON inspection first.
#### Scenario: Tool and verifier evidence are inspectable
- **WHEN** a trace includes tool invocations and verifier facts
- **THEN** the page SHALL show tool invocation counts, success status, tool
filters, verifier facts, and evidence references
#### Scenario: RAG details are inspectable for lookup knowledge calls
- **WHEN** a `lookup_knowledge` invocation includes retrieval details
- **THEN** the page SHALL show query transform, retrieval trace, context pack,
rerank trace, and evidence block data where available
#### Scenario: Skill boundary checks are visible
- **WHEN** a trace includes planner, executor, or verifier steps
- **THEN** the page SHALL show best-effort indicators for selected skill,
planner `read_skill` text mentions, executor `read_skill` text mentions, and
verifier `read_skill` text mentions