feat(graph): cut over chat diagnosis stategraph

This commit is contained in:
zhuyongxin
2026-07-17 18:30:08 +08:00
parent 1460dd1e99
commit 99e490f227
36 changed files with 2640 additions and 1623 deletions
@@ -3,9 +3,7 @@
## Purpose
Separate multi-turn conversation metadata from per-execution diagnosis state. `sessionId` identifies the conversation context, while `runId` identifies one replayable diagnosis execution and scopes Trace, Feedback, Evaluation, AIOps, and case-library provenance.
## Requirements
### Requirement: Conversation metadata SHALL be separated from diagnosis runs
The system SHALL persist multi-turn conversation metadata in `chat_session` and one execution's auditable state in `diagnosis_run`.
@@ -179,3 +177,58 @@ The change SHALL verify both runtime behavior and evaluation baseline impact.
- **WHEN** verification is complete
- **THEN** the project SHALL run or explicitly evaluate the relevant baseline diff command
- **AND** any drift caused by run isolation SHALL be documented as expected or investigated as a regression
### Requirement: StateGraph Chat runs SHALL persist a compact orchestration trace
Each successful new StateGraph complex Chat run SHALL persist a non-empty compact orchestration summary derived from its bounded Graph events in `diagnosis_run.orchestration_trace`.
#### Scenario: Graph reaches Composer
- **WHEN** a complex Chat Graph terminates through Composer with a safe answer
- **THEN** the current DiagnosisRun SHALL store version, transitions, final node, termination reason, degraded flag, and evidence retry count
- **AND** the summary SHALL be derived from the current Run's actual orchestration events
#### Scenario: Graph reaches handled Fallback
- **WHEN** a complex Chat Graph terminates through deterministic Fallback with a safe answer
- **THEN** the current DiagnosisRun SHALL store a non-empty orchestration trace with `degraded=true`
- **AND** the Run status SHALL be SUCCESS
#### Scenario: Unhandled execution fails after events exist
- **WHEN** an unhandled failure occurs after one or more real Graph events are available
- **THEN** the service SHALL best-effort persist a partial orchestration summary for the current failed run
- **AND** it SHALL NOT add a node or transition that did not occur
#### Scenario: Orchestration trace content is inspected
- **WHEN** orchestration trace JSON is serialized
- **THEN** it SHALL NOT include Prompt text, model reasoning, raw tool output, raw Executor output, or Graph State snapshots
- **AND** it SHALL NOT contain data owned by another run
### Requirement: Trace API SHALL expose orchestration trace only on the run object
The Trace API SHALL parse the current DiagnosisRun orchestration JSON and expose it only as `run.orchestrationTrace`.
#### Scenario: Exact StateGraph run trace is queried
- **WHEN** a caller queries a successful new StateGraph Chat run
- **THEN** `run.orchestrationTrace` SHALL be a non-empty parsed JSON object
- **AND** the response top level and compatibility `session` projection SHALL NOT duplicate the field
- **AND** no raw orchestration trace field SHALL be added
#### Scenario: Historical or non-StateGraph run is queried
- **WHEN** the selected DiagnosisRun has null orchestration trace
- **THEN** `run.orchestrationTrace` MAY be null
- **AND** the service SHALL NOT synthesize historical events or read another run's trace
### Requirement: Orchestration trace migration SHALL be additive and nullable
The database migration SHALL add only one nullable JSON column named `orchestration_trace` to `diagnosis_run` for this change.
#### Scenario: Migration is applied
- **WHEN** Flyway applies the stage 3 migration
- **THEN** existing DiagnosisRun rows SHALL remain valid without backfill
- **AND** no other table or column SHALL be changed by the stage 3 schema migration