feat(trace): add diagnosis trace workbench
This commit is contained in:
@@ -0,0 +1,70 @@
|
||||
# Trace UI workbench design
|
||||
|
||||
## Subject and audience
|
||||
|
||||
Subject: a single-session diagnosis trace audit workbench.
|
||||
|
||||
Audience: backend and AIOps engineers reviewing an MVP diagnosis run after a
|
||||
demo, incident drill, or regression check.
|
||||
|
||||
Single job: turn one trace response into an inspectable ledger that exposes
|
||||
agent flow, evidence, verifier judgement, and skill boundaries without reading
|
||||
raw JSON first.
|
||||
|
||||
## Frontend approach
|
||||
|
||||
- Static `trace.html`, `trace.css`, and `trace.js`.
|
||||
- No dependency on a package manager, bundler, external fonts, or remote icons.
|
||||
- Fetches `/api/diagnosis/{sessionId}/trace` and renders client-side.
|
||||
- Accepts `?sessionId=...` and keeps the loaded session id in the URL.
|
||||
|
||||
## Visual direction
|
||||
|
||||
Palette:
|
||||
|
||||
- `#111827` ink rail for the trace frame.
|
||||
- `#f7f3ea` warm ledger surface.
|
||||
- `#1f7a6b` evidence green.
|
||||
- `#b5472f` rejection red.
|
||||
- `#c58a19` warning amber.
|
||||
- `#5b6472` operational gray.
|
||||
|
||||
Type:
|
||||
|
||||
- UI/body: `Inter, ui-sans-serif, system-ui`.
|
||||
- Data and labels: `ui-monospace, SFMono-Regular, Consolas`.
|
||||
|
||||
Layout:
|
||||
|
||||
```text
|
||||
+--------------------------------------------------------------------+
|
||||
| Session id input | Load | status / verdict / counts strip |
|
||||
+------------------+----------------------+--------------------------+
|
||||
| Agent trace rail | Evidence tool ledger | Inspector |
|
||||
| ordered steps | filters + calls | verifier + RAG + skill |
|
||||
+------------------+----------------------+--------------------------+
|
||||
```
|
||||
|
||||
Signature element: a trace rail that treats each agent step as a ledger entry
|
||||
with ordered markers, duration, token count, and expandable raw model excerpts.
|
||||
|
||||
## Data mapping
|
||||
|
||||
- `data.session`: summary strip, final answer, self-evaluation source.
|
||||
- `data.steps`: ordered trace rail.
|
||||
- `data.toolInvocations`: evidence ledger and selected tool inspector.
|
||||
- `data.session.selfEvaluation.verifier_evaluation`: verdict, groundedness,
|
||||
facts table, and verifier `tool_trace_summary`.
|
||||
- `data.toolInvocations[*].retrievalDetails`: RAG query transform, retrieval
|
||||
trace, context pack, rerank trace, and evidence blocks.
|
||||
- `steps[*].modelOutput`: best-effort extraction of `selected_skill`.
|
||||
- `steps[*].modelInput/modelOutput` and tool names: best-effort skill boundary
|
||||
checks for `read_skill`.
|
||||
|
||||
## Accessibility and states
|
||||
|
||||
- Keyboard-focusable controls.
|
||||
- Loading, missing session id, API error, empty steps, empty tools, and missing
|
||||
verifier states.
|
||||
- Responsive three-pane desktop layout that stacks on narrow screens.
|
||||
- Respects `prefers-reduced-motion`.
|
||||
@@ -0,0 +1,37 @@
|
||||
# Trace UI workbench
|
||||
|
||||
## Why
|
||||
|
||||
The MVP already persists diagnosis trace data and exposes it through
|
||||
`GET /api/diagnosis/{sessionId}/trace`, but reviewers still need to inspect raw
|
||||
JSON to answer basic audit questions:
|
||||
|
||||
- Which agents ran, in what order, and with what persisted counts?
|
||||
- Which evidence tools executed, how many times, and did they succeed?
|
||||
- Which facts did the Verifier check, and which evidence refs support them?
|
||||
- Did Planner only select skill metadata while Executor handled skill loading?
|
||||
- What RAG retrieval details were used by `lookup_knowledge`?
|
||||
|
||||
This slows down demo review and makes trace quality issues harder to spot.
|
||||
|
||||
## What changes
|
||||
|
||||
- Add a static Trace workbench page served by Spring Boot static resources.
|
||||
- Load a diagnosis trace by session id through the existing read-only Trace API.
|
||||
- Render session summary, agent timeline, evidence tool ledger, verifier facts,
|
||||
skill boundary checks, and RAG retrieval details.
|
||||
- Add a navigation entry from the existing chat page to the Trace workbench.
|
||||
|
||||
## Non-goals
|
||||
|
||||
- No new backend endpoint.
|
||||
- No mutation of diagnosis sessions, agent steps, tool invocations, or feedback.
|
||||
- No new frontend framework or build pipeline.
|
||||
- No change to the persisted trace schema.
|
||||
|
||||
## Impact
|
||||
|
||||
- Frontend-only runtime surface under `src/main/resources/static`.
|
||||
- Uses the existing `Result<T>` API response contract.
|
||||
- Works with existing trace records, including sessions that lack verifier or RAG
|
||||
detail fields.
|
||||
+36
@@ -0,0 +1,36 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: MVP demo SHALL provide a browser trace workbench
|
||||
The MVP demo SHALL provide a browser-accessible static page for inspecting one
|
||||
diagnosis trace by session id using the existing read-only Trace API.
|
||||
|
||||
#### Scenario: Existing trace renders in the workbench
|
||||
- **WHEN** a reviewer opens the Trace workbench with a session id that exists
|
||||
- **THEN** the page SHALL request `GET /api/diagnosis/{sessionId}/trace`
|
||||
- **AND** it SHALL render session summary, ordered agent steps, ordered tool
|
||||
invocations, verifier evaluation, and final answer when present
|
||||
|
||||
#### Scenario: Trace workbench handles missing or failed traces
|
||||
- **WHEN** the Trace API returns an error or the session id is empty
|
||||
- **THEN** the page SHALL show a clear error or empty state without mutating any
|
||||
diagnosis data
|
||||
|
||||
### Requirement: Trace workbench SHALL expose evidence and skill boundaries
|
||||
The Trace workbench SHALL make tool evidence and skill-loading boundaries visible
|
||||
without requiring raw JSON inspection first.
|
||||
|
||||
#### Scenario: Tool and verifier evidence are inspectable
|
||||
- **WHEN** a trace includes tool invocations and verifier facts
|
||||
- **THEN** the page SHALL show tool invocation counts, success status, tool
|
||||
filters, verifier facts, and evidence references
|
||||
|
||||
#### Scenario: RAG details are inspectable for lookup knowledge calls
|
||||
- **WHEN** a `lookup_knowledge` invocation includes retrieval details
|
||||
- **THEN** the page SHALL show query transform, retrieval trace, context pack,
|
||||
rerank trace, and evidence block data where available
|
||||
|
||||
#### Scenario: Skill boundary checks are visible
|
||||
- **WHEN** a trace includes planner, executor, or verifier steps
|
||||
- **THEN** the page SHALL show best-effort indicators for selected skill,
|
||||
planner `read_skill` text mentions, executor `read_skill` text mentions, and
|
||||
verifier `read_skill` text mentions
|
||||
@@ -0,0 +1,9 @@
|
||||
# Tasks
|
||||
|
||||
- [x] 1. Add OpenSpec delta for the Trace UI workbench.
|
||||
- [x] 2. Add static Trace workbench HTML/CSS/JS.
|
||||
- [x] 3. Add a chat-page navigation entry to the Trace workbench.
|
||||
- [x] 4. Verify Java compilation and static page syntax.
|
||||
- [x] 5. Validate the page against a real trace endpoint when a local service is available.
|
||||
- [x] 6. Archive the OpenSpec change after validation.
|
||||
- [x] 7. Commit the completed change.
|
||||
@@ -1,9 +1,7 @@
|
||||
## Purpose
|
||||
|
||||
Provide a repeatable MVP demo flow that can run a chat diagnosis, expose its persisted execution trace, and submit feedback for the same session id.
|
||||
|
||||
## Requirements
|
||||
|
||||
### Requirement: Diagnosis trace can be queried by session id
|
||||
The system SHALL expose a read-only HTTP endpoint `GET /api/diagnosis/{sessionId}/trace` that returns the persisted diagnosis trace for the requested session id.
|
||||
|
||||
@@ -64,3 +62,39 @@ The MVP demo SHALL document which trace fields to inspect for evidence, verifier
|
||||
#### Scenario: Checklist maps fields to interview claims
|
||||
- **WHEN** a developer reviews a trace response
|
||||
- **THEN** the checklist SHALL map concrete JSON paths to the claims made in the interview walkthrough
|
||||
|
||||
### Requirement: MVP demo SHALL provide a browser trace workbench
|
||||
The MVP demo SHALL provide a browser-accessible static page for inspecting one
|
||||
diagnosis trace by session id using the existing read-only Trace API.
|
||||
|
||||
#### Scenario: Existing trace renders in the workbench
|
||||
- **WHEN** a reviewer opens the Trace workbench with a session id that exists
|
||||
- **THEN** the page SHALL request `GET /api/diagnosis/{sessionId}/trace`
|
||||
- **AND** it SHALL render session summary, ordered agent steps, ordered tool
|
||||
invocations, verifier evaluation, and final answer when present
|
||||
|
||||
#### Scenario: Trace workbench handles missing or failed traces
|
||||
- **WHEN** the Trace API returns an error or the session id is empty
|
||||
- **THEN** the page SHALL show a clear error or empty state without mutating any
|
||||
diagnosis data
|
||||
|
||||
### Requirement: Trace workbench SHALL expose evidence and skill boundaries
|
||||
The Trace workbench SHALL make tool evidence and skill-loading boundaries visible
|
||||
without requiring raw JSON inspection first.
|
||||
|
||||
#### Scenario: Tool and verifier evidence are inspectable
|
||||
- **WHEN** a trace includes tool invocations and verifier facts
|
||||
- **THEN** the page SHALL show tool invocation counts, success status, tool
|
||||
filters, verifier facts, and evidence references
|
||||
|
||||
#### Scenario: RAG details are inspectable for lookup knowledge calls
|
||||
- **WHEN** a `lookup_knowledge` invocation includes retrieval details
|
||||
- **THEN** the page SHALL show query transform, retrieval trace, context pack,
|
||||
rerank trace, and evidence block data where available
|
||||
|
||||
#### Scenario: Skill boundary checks are visible
|
||||
- **WHEN** a trace includes planner, executor, or verifier steps
|
||||
- **THEN** the page SHALL show best-effort indicators for selected skill,
|
||||
planner `read_skill` text mentions, executor `read_skill` text mentions, and
|
||||
verifier `read_skill` text mentions
|
||||
|
||||
|
||||
Reference in New Issue
Block a user