feat(graph): cut over chat diagnosis stategraph

This commit is contained in:
zhuyongxin
2026-07-17 18:30:08 +08:00
parent 1460dd1e99
commit 99e490f227
36 changed files with 2640 additions and 1623 deletions
@@ -0,0 +1,60 @@
## MODIFIED Requirements
### Requirement: Composer SHALL generate final user-facing Chat answers
The system SHALL invoke the Composer Graph Node after Verifier routing to generate the final user-facing Chat answer from Verifier-allowed material.
#### Scenario: Composer receives only filtered material
- **WHEN** the Composer Graph Node invokes its configured Agent
- **THEN** the Composer input SHALL contain `original_query`, `verdict`, `allowed_claims`, `allowed_hypotheses`, `missing_info`, `recommended_actions`, and `rationale`
- **AND** the Composer input SHALL NOT contain raw tool output
- **AND** the Composer input SHALL NOT contain the full unscreened Executor output
- **AND** the Composer input SHALL NOT contain Executor `user_facing_answer`
#### Scenario: Composer outputs strict JSON
- **WHEN** Composer completes
- **THEN** it SHALL output exactly one JSON object
- **AND** the JSON object SHALL include `answer_summary`, `recommended_actions`, and `user_facing_answer`
- **AND** it SHALL NOT output Markdown, code fences, or explanatory text outside the JSON object
#### Scenario: Composer does not introduce new facts
- **WHEN** Composer produces `answer_summary`, `recommended_actions`, or `user_facing_answer`
- **THEN** every service name, entity, timestamp, error code, metric value, root cause, and recommendation reason SHALL be derived from the Composer input
- **AND** Composer SHALL NOT add facts from model knowledge, raw tool history, or Executor raw text
### Requirement: Composer input SHALL honor Verifier claim checks
The Composer Graph Node SHALL construct Composer input by filtering verified Executor structured output through Verifier `claim_checks`.
#### Scenario: Passing claims become allowed claims
- **WHEN** a claim check verification is `direct_observation`
- **THEN** the Composer input builder SHALL include the matching verified Executor claim in `allowed_claims`
#### Scenario: Reasonable inferences remain bounded
- **WHEN** a claim check verification is `reasonable_inference`
- **THEN** the Composer input builder MAY include the matching verified Executor claim in `allowed_claims`
- **AND** the final answer SHALL NOT describe it as the sole confirmed root cause unless the allowed claim itself is a root-cause claim and the final verdict is `PASS`
#### Scenario: Overstated claims are not confirmed findings
- **WHEN** a claim check verification is `overstated`
- **THEN** the Composer input builder SHALL NOT include the matching claim as a confirmed item in `allowed_claims`
- **AND** it MAY include it as `allowed_hypotheses` or represent it in `missing_info`
#### Scenario: Unsupported or external claims are withheld
- **WHEN** a claim check verification is `unsupported`, `external_unknown`, or `contradicted`
- **THEN** the Composer input builder SHALL NOT include the matching claim in `allowed_claims`
- **AND** the final user-facing answer SHALL NOT present that claim as confirmed
### Requirement: Composer failures SHALL degrade safely
The system SHALL tolerate malformed or failed Composer execution without leaking raw JSON or unverified Executor material.
#### Scenario: malformed Composer output falls back safely
- **WHEN** Composer exhausts its fixed-input technical retry or returns a non-retryable failure
- **THEN** the deterministic Fallback Node SHALL produce a final answer using only filtered material
- **AND** the final answer SHALL NOT expose raw Composer output
- **AND** the final answer SHALL NOT expose raw Executor output
- **AND** the final answer SHALL NOT use Executor `user_facing_answer`
#### Scenario: Composer audit is persisted
- **WHEN** Graph result mapping persists verifier evaluation
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation.composer_output` SHALL record parsed Composer audit when available
- **AND** handled Composer fallback SHALL be observable through orchestration trace and available status/reason fields
- **AND** the audit SHALL remain compact and SHALL NOT store full raw tool output
@@ -0,0 +1,107 @@
## ADDED Requirements
### Requirement: Complex Chat SHALL use the real Diagnosis StateGraph as its only production orchestrator
The system SHALL execute each complex Chat request through one real Diagnosis CompiledGraph and SHALL NOT create, invoke, or fall back to a SequentialAgent workflow.
#### Scenario: Complex Chat starts
- **WHEN** `executeChatWithStrategy` classifies a valid question as complex
- **THEN** ChatService SHALL create the current Diagnosis Run and invoke one Diagnosis CompiledGraph
- **AND** the production path SHALL NOT maintain an outer score-based retry loop
#### Scenario: Graph dependencies are assembled
- **WHEN** the complex Chat Graph is constructed
- **THEN** Planner, Executor, Verifier, and Composer SHALL use the existing project Prompt, tool, skill, and AgentLoggingHook assembly rules
- **AND** the Graph Verifier SHALL NOT register VerifierInputHook
- **AND** Gatekeeper SHALL run only as the explicit Graph Node
### Requirement: Graph invocation SHALL preserve current Run ownership
The system SHALL use the current `runId` as Graph `threadId` and SHALL pass both current `sessionId` and `runId` in RunnableConfig metadata.
#### Scenario: Agent Node runs
- **WHEN** any Agent Node is invoked for a complex Chat run
- **THEN** its RunnableConfig threadId SHALL equal the current runId
- **AND** AgentStep and ToolInvocation writes SHALL retain the current sessionId and runId
#### Scenario: Gatekeeper validates Executor output
- **WHEN** the Gatekeeper Node runs
- **THEN** it SHALL validate only tool invocations belonging to the runId in RunnableConfig metadata
- **AND** data from another run in the same session SHALL NOT be considered
### Requirement: Initial Graph State SHALL contain only bounded diagnosis control data
ChatService SHALL initialize Diagnosis Graph State with the current query context, NORMAL Planner mode, zero independent retry counters, and an empty orchestration event list.
#### Scenario: Initial state is projected
- **WHEN** a complex Chat run enters Planner for the first time
- **THEN** `diagnosis_context` SHALL contain the current query/original query
- **AND** complete conversation history SHALL NOT be stored in parent Graph State
- **AND** history MAY remain in the Planner and Executor system Prompt assembled for this request
#### Scenario: Counters are initialized
- **WHEN** the Graph starts
- **THEN** planner, verifier, composer, and evidence retry counters SHALL each be zero
- **AND** Planner mode SHALL be NORMAL
### Requirement: Graph final state SHALL map to the existing Chat result and Run lifecycle
The system SHALL use non-empty `final_answer` from a handled Graph terminal state as the existing ChatResult answer. Run status SHALL express execution lifecycle rather than diagnosis quality.
#### Scenario: Composer completes
- **WHEN** Graph reaches Composer and produces a safe non-empty final answer
- **THEN** ChatResult SHALL preserve the current answer/sessionId/runId protocol
- **AND** DiagnosisRun SHALL be saved as SUCCESS with answer, duration, token count, step count, and tool count
- **AND** EvaluationService SHALL evaluate the current runId
#### Scenario: Handled Fallback completes
- **WHEN** Graph reaches deterministic Fallback and produces a safe non-empty answer
- **THEN** DiagnosisRun SHALL be SUCCESS
- **AND** diagnosis quality SHALL be expressed by verifier fields when available or `orchestrationTrace.degraded=true`
- **AND** no DEGRADED Run status SHALL be introduced
#### Scenario: Graph cannot produce a safe response
- **WHEN** Graph has an unhandled failure, final state is unavailable, final answer is blank, or required successful-result persistence fails
- **THEN** DiagnosisRun SHALL be marked FAILED when it can still be saved
- **AND** the system SHALL NOT report a successful Graph result
### Requirement: Graph state SHALL be the only source for verifier evaluation persistence
The system SHALL build `diagnosis_run.self_evaluation.verifier_evaluation` from explicit Graph State and SHALL NOT read VerifierContextHolder on the complex Chat path.
#### Scenario: Verifier completed
- **WHEN** final Graph State contains a completed Verifier result
- **THEN** verifier evaluation SHALL include execution status, model verdict, effective verdict, groundedness, claim/fact checks, rationale, round, Gatekeeper audit, verified Executor output/evidence, Prompt audit, and Composer audit when available
- **AND** the compatibility `verdict` field SHALL equal effective verdict
#### Scenario: Pre-verification Fallback completed
- **WHEN** Graph reaches Fallback before Verifier completes
- **THEN** verifier evaluation SHALL preserve available execution statuses, Gatekeeper audit, Prompt audit, and fallback context
- **AND** it SHALL NOT fabricate a model or effective verdict
#### Scenario: Evaluation payload is inspected
- **WHEN** verifier evaluation is persisted
- **THEN** it SHALL NOT contain raw Executor text or complete tool trace summary
- **AND** `executor_structured_output` SHALL contain at most the verified projection retained for compatibility
### Requirement: Stage 3 verification SHALL not run the final live E2E
The change SHALL use focused automated tests for production cutover and contracts while reserving Maven live startup, log inspection, and database querying for stage 5.
#### Scenario: Stage 3 is accepted
- **WHEN** stage 3 verification completes
- **THEN** Graph cutover, Run/Trace, Prompt, Controller/Eval regressions, Maven test compilation, and OpenSpec strict validation SHALL have passed
- **AND** live E2E, `logs/`, and `scripts/query_mysql.py` SHALL be recorded as intentionally deferred to stage 5
@@ -0,0 +1,7 @@
## REMOVED Requirements
### Requirement: Real Nodes SHALL remain isolated from the production Chat path in stage 2
**Reason**: Stage 2 production isolation has completed its migration purpose; stage 3 intentionally makes the real Diagnosis Graph the only complex Chat production orchestrator.
**Migration**: Replace `ChatService.executeChatComplex` SequentialAgent orchestration with real Graph assembly/invocation. Keep `/api/chat` compatible and use a full Git revert of stage 3 for runtime rollback rather than retaining dual paths.
@@ -0,0 +1,263 @@
## MODIFIED Requirements
### Requirement: Verifier SHALL fact-check Executor answers
The system SHALL have a Verifier Agent that reads only Gatekeeper-projected structured Executor claims and verified claim-local evidence, then produces a structured verdict based on claim derivability.
#### Scenario: PASS verdict when all claims have evidence
- **WHEN** all critical claims in `verified_executor_output.claims` have direct observation or reasonable inference support in `verified_evidence`
- **AND** at least one critical claim has direct observation
- **AND** no critical claim is contradicted, unsupported, external unknown, or overstated
- **AND** verdict ceiling is PASS
- **THEN** the Verifier MAY output model verdict="PASS" with groundedness_score ≥ 0.5
#### Scenario: LOW_CONFID verdict with partial evidence
- **WHEN** no critical claim contradicts verified evidence
- **AND** some critical claims are `unsupported`, `external_unknown`, or `overstated`
- **THEN** the Verifier SHALL output model verdict="LOW_CONFID"
#### Scenario: LOW_CONFID verdict with only inference support
- **WHEN** no critical claim contradicts verified evidence
- **AND** all critical claims are only `reasonable_inference`
- **THEN** the Verifier SHALL output model verdict="LOW_CONFID"
#### Scenario: REJECT verdict when claims contradict evidence
- **WHEN** any critical claim in `verified_executor_output.claims` contradicts verified evidence
- **OR** the claim fabricates a key entity, error code, or conclusion that does not exist in verified evidence
- **THEN** the Verifier SHALL output model verdict="REJECT"
#### Scenario: Verified structured claims are the only verification target
- **WHEN** `verified_executor_output.claims` is present
- **THEN** Verifier SHALL verify each structured claim against matching `verified_evidence` through `claim_checks`
- **AND** each claim's evidence references SHALL match existing claim/invocation/tool/path identifiers when available
- **AND** a claim without matching verified evidence SHALL NOT be classified as `direct_observation`
- **AND** Verifier SHALL NOT receive or add confirmed facts from raw Executor text
#### Scenario: Executor output is invalid
- **WHEN** Executor does not return a legal structured contract
- **THEN** the Graph SHALL route directly to pre-verification Fallback
- **AND** Verifier SHALL NOT execute or fabricate a diagnostic verdict
### Requirement: facts_checked SHALL use a fixed classification set
The system SHALL continue to expose compatibility `facts_checked` using its fixed verification classification set.
#### Scenario: claim checks are mapped to legacy facts
- **WHEN** Verifier output contains `claim_checks`
- **THEN** the shared Verifier protocol parser SHALL derive compatibility `facts_checked` when the model did not provide them
- **AND** `direct_observation` SHALL map to `direct_evidence`
- **AND** `reasonable_inference` and `overstated` SHALL map to `indirect_support`
- **AND** `unsupported` and `external_unknown` SHALL map to `no_evidence`
- **AND** `contradicted` SHALL map to `contradicted`
### Requirement: ChatService SHALL route based on Verifier verdict
The system SHALL use Diagnosis StateGraph conditional edges, rather than a ChatService outer loop, to route explicit Verifier execution status and effective verdict.
#### Scenario: PASS routes to Composer
- **WHEN** Verifier completes with effective verdict="PASS"
- **THEN** the Graph SHALL invoke Composer with filtered Verifier-allowed material
- **AND** the final user-facing answer SHALL NOT pass through raw Executor output
- **AND** the final user-facing answer SHALL NOT read Executor `user_facing_answer`
#### Scenario: LOW_CONFID does not qualify for evidence retry
- **WHEN** Verifier completes LOW_CONFID but ceiling is LOW_CONFID, no valid critical evidence gap exists, or evidence retry count is already one
- **THEN** the Graph SHALL route to Composer without another Planner cycle
- **AND** the final answer SHALL distinguish confirmed information, possible directions, and evidence gaps
#### Scenario: LOW_CONFID qualifies for evidence retry
- **WHEN** Verifier completes LOW_CONFID with ceiling PASS, at least one critical valid evidence gap, and evidence retry count zero
- **THEN** the Graph SHALL invoke one EVIDENCE_GAP_ONLY Planner cycle
- **AND** it SHALL NOT use groundedness threshold or a ChatService feature flag to decide the retry
#### Scenario: REJECT does not enter retry round
- **WHEN** Verifier completes with effective verdict="REJECT"
- **THEN** the Graph SHALL NOT start an evidence supplementation round
- **AND** it SHALL route to Composer-safe output
#### Scenario: REJECT produces bounded output
- **WHEN** effective verdict is REJECT
- **THEN** the system SHALL output a degraded result indicating current evidence cannot support a reliable conclusion
- **AND** it SHALL NOT pass through raw Executor answer
- **AND** it SHALL NOT include an unsupported root-cause conclusion
#### Scenario: Verifier execution fails
- **WHEN** Verifier exhausts technical retry or returns a non-retryable failure
- **THEN** the Graph SHALL route to pre-verification Fallback
- **AND** no execution status string SHALL be used as model or effective verdict
### Requirement: Verifier SHALL be observable
The Verifier execution, effective verdict, and downstream final-answer composition SHALL be persisted in the current Diagnosis Run self-evaluation container.
#### Scenario: claim checks written to self_evaluation
- **WHEN** a completed Verifier evaluation is persisted
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include `claim_checks`
- **AND** it SHALL continue to include compatibility `facts_checked`
- **AND** it SHALL include `verifier_status`, `model_verdict`, `effective_verdict`, `verdict`, `groundedness_score`, `rationale`, verified output/evidence, and Gatekeeper audit
#### Scenario: composer output written to self_evaluation
- **WHEN** final answer composition completes
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include compact `composer_output` when available
- **AND** handled Composer fallback SHALL remain observable through orchestration trace and status/reason fields
- **AND** existing claim/fact and Gatekeeper fields SHALL be preserved
#### Scenario: verdict written to self_evaluation
- **WHEN** Verifier completes
- **THEN** Graph result mapping SHALL write effective verdict under `diagnosis_run.self_evaluation.verifier_evaluation.verdict`
- **AND** existing `rule_evaluation` and `aiops_rule_evaluation` channels SHALL be preserved
#### Scenario: pre-verification fallback is persisted
- **WHEN** Graph reaches Fallback before Verifier completes
- **THEN** verifier evaluation SHALL include available status, Gatekeeper audit, failure reason, and Prompt audit
- **AND** it SHALL NOT fabricate `model_verdict` or `effective_verdict`
#### Scenario: gatekeeper result written to self_evaluation
- **WHEN** Graph result mapping persists available Gatekeeper state
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result`
- **AND** the result SHALL retain status, severity, checked bindings, rules, failed rules, warnings, and errors when provided by Gatekeeper
#### Scenario: prompt audit written to verifier evaluation
- **WHEN** a complex Chat Graph result is persisted
- **THEN** the system SHALL include a `prompt_audit` object under `diagnosis_run.self_evaluation.verifier_evaluation`
- **AND** `prompt_audit.version` SHALL identify the Chat Prompt audit catalog version
- **AND** `prompt_audit.prompts` SHALL include Planner, Executor, Verifier, and Composer Prompt names and versions
- **AND** full Prompt text SHALL NOT be persisted
#### Scenario: prompt audit available on fallback paths
- **WHEN** Planner, Executor, Gatekeeper, Verifier, or Composer reaches a handled Fallback
- **THEN** the persisted verifier evaluation SHALL still include `prompt_audit`
#### Scenario: evaluation payload is inspected
- **WHEN** Graph verifier evaluation is persisted
- **THEN** it SHALL NOT contain raw Executor text or complete `tool_trace_summary`
- **AND** compatibility `executor_structured_output` SHALL contain at most the verified projection
### Requirement: Verifier SHALL consume explicit verification inputs
The Verifier SHALL receive a Graph-built verified-only payload rather than inferring business inputs from conversation history, ThreadLocal state, raw Executor text, or complete tool history.
#### Scenario: explicit input blocks available to Verifier
- **WHEN** the Verifier Graph Node starts
- **THEN** the payload SHALL provide `diagnosis_context`, `verified_executor_output`, `verified_evidence`, `gatekeeper_audit`, and `verdict_ceiling`
- **AND** permitted structured `retry_context` SHALL be provided only after evidence retry preparation
#### Scenario: Verifier remains isolated from intermediate and raw material
- **WHEN** the Verifier input is serialized
- **THEN** it SHALL exclude Planner reasoning, Executor intermediate reasoning, raw Executor text, complete tool trace summary, Prompt text, and unrelated parent Graph State
#### Scenario: only passed bindings are available
- **WHEN** Gatekeeper returns mixed passed and failed checked bindings
- **THEN** `verified_executor_output` and `verified_evidence` SHALL contain only claims/material matching passed bindings
- **AND** the Verifier SHALL NOT receive failed or unreferenced tool material
#### Scenario: verified evidence preserves precise references
- **WHEN** the system prepares Verifier input
- **THEN** each verified evidence item SHALL preserve claim id, source invocation id, tool name, raw path, and matched text
- **AND** the item SHALL be traceable to current-run Gatekeeper validation
#### Scenario: gatekeeper audit and ceiling are available
- **WHEN** the system prepares Verifier input
- **THEN** the payload SHALL include raw Gatekeeper audit separately from normalized verdict ceiling
- **AND** a LOW_CONFID ceiling SHALL prevent effective PASS
#### Scenario: technical retry occurs
- **WHEN** the first Verifier attempt returns invalid output or a retryable invocation failure
- **THEN** the second attempt SHALL receive byte-identical serialized input
- **AND** Executor, Gatekeeper, and tools SHALL NOT rerun
### Requirement: Verifier facts SHALL be auditable
Verifier claims and facts SHALL be linkable to the verified binding projection used during verification.
#### Scenario: claim checks contain evidence refs
- **WHEN** the Verifier emits `claim_checks`
- **THEN** each check SHALL include an `evidence_refs` array
- **AND** any non-empty evidence ref SHALL correspond to existing verified evidence by claim id, source invocation id, tool name, or raw path
- **AND** it SHALL NOT reference a failed or unverified binding
#### Scenario: verifier evaluation persists traceability snapshot
- **WHEN** Graph result mapping persists verifier evaluation
- **THEN** it SHALL include `traceability_version`
- **AND** it SHALL include the bounded `verified_evidence` snapshot used by the Verifier
- **AND** it SHALL NOT persist a complete tool trace summary as Verifier input
### Requirement: Structured Executor output SHALL degrade safely
The StateGraph runtime SHALL tolerate malformed or absent structured Executor output without crashing the Chat flow or invoking Verifier with untrusted material.
#### Scenario: Malformed Executor JSON is classified
- **WHEN** Executor returns malformed JSON or text outside the expected contract
- **THEN** Executor Node SHALL set INVALID_OUTPUT
- **AND** the Graph SHALL route directly to deterministic pre-verification Fallback
- **AND** Gatekeeper, Verifier, and model Composer SHALL NOT execute
#### Scenario: Structured parse failure remains observable
- **WHEN** Executor output parsing fails
- **THEN** orchestration events and verifier evaluation status/failure fields SHALL make the parse failure visible
- **AND** the failure SHALL NOT be treated as a successful evidence-attribution contract or diagnostic verdict
### Requirement: Executor Gatekeeper SHALL validate deterministic structured-output failures
The system SHALL run deterministic Gatekeeper checks as an explicit Graph Node after legal Executor output parsing and before Verifier model execution.
#### Scenario: schema rule rejects removed fields
- **WHEN** Executor structured output contains `diagnosis_summary` or `user_facing_answer`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
#### Scenario: schema rule rejects missing evidence bindings
- **WHEN** a confirmed claim has no `evidence_bindings`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
#### Scenario: invocation rule rejects fabricated invocation ids
- **WHEN** a claim evidence binding references an invocation id absent from current-run `tool_invocation` rows
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
#### Scenario: invocation rule rejects tool name mismatch
- **WHEN** a claim evidence binding references an existing current-run invocation id
- **AND** binding `tool_name` does not match persisted invocation `tool_name`
- **THEN** `gatekeeper_result.status` SHALL be `fail`
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
#### Scenario: valid structured output passes initial gatekeeper rules
- **WHEN** Executor emits legal `executor_evidence_v2`
- **AND** each claim has evidence bindings pointing to current-run invocations with matching tool names and paths
- **THEN** `gatekeeper_result.status` SHALL be `pass`
- **AND** `gatekeeper_result.failed_rules` SHALL be empty
#### Scenario: gatekeeper reject bypasses Verifier
- **WHEN** normalized Gatekeeper status is REJECT
- **THEN** the Graph SHALL route directly to pre-verification Fallback
- **AND** Verifier SHALL NOT execute
#### Scenario: gatekeeper low confidence is bounded
- **WHEN** normalized Gatekeeper status is LOW_CONFID with at least one passed binding
- **THEN** verified input SHALL contain only passed bindings
- **AND** effective verdict SHALL NOT exceed LOW_CONFID
### Requirement: Verifier SHALL use verified claim-local evidence for derivability
Verifier SHALL judge structured claims only against Gatekeeper-verified claim-local evidence excerpts and their precise current-run references.
#### Scenario: Verified excerpt supports direct observation
- **WHEN** verdict ceiling is PASS
- **AND** a claim's verified evidence matched text directly contains the claim's concrete facts
- **THEN** Verifier MAY classify that claim as `direct_observation`
#### Scenario: Verified evidence is complete Verifier context
- **WHEN** verified claims and evidence are available
- **THEN** Verifier SHALL use them as its evidence context
- **AND** it SHALL NOT require or request a complete tool trace summary
- **AND** it SHALL NOT read raw Executor or unreferenced tool material
### Requirement: Gatekeeper severity SHALL constrain effective verdict
Runtime effective verdict calculation SHALL treat normalized Gatekeeper ceiling as a hard upper bound independent from Verifier model output.
#### Scenario: Reject severity bypasses Verifier
- **WHEN** `gatekeeper_result.severity=reject`
- **THEN** the Graph SHALL route to pre-verification Fallback without invoking Verifier
- **AND** it SHALL NOT fabricate an effective diagnostic verdict
#### Scenario: Low confidence severity prevents PASS
- **WHEN** `gatekeeper_result.severity=low_confid`
- **AND** the Verifier model returns `verdict=PASS`
- **THEN** deterministic effective-verdict calculation SHALL downgrade the result
- **AND** effective verdict SHALL be `LOW_CONFID`
#### Scenario: Gatekeeper audit includes severity
- **WHEN** Graph verifier evaluation is persisted
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation.gatekeeper_result` SHALL include available `status`, `severity`, `checked_bindings`, `failed_rules`, `warnings`, and `errors`
@@ -0,0 +1,56 @@
## ADDED Requirements
### Requirement: StateGraph Chat runs SHALL persist a compact orchestration trace
Each successful new StateGraph complex Chat run SHALL persist a non-empty compact orchestration summary derived from its bounded Graph events in `diagnosis_run.orchestration_trace`.
#### Scenario: Graph reaches Composer
- **WHEN** a complex Chat Graph terminates through Composer with a safe answer
- **THEN** the current DiagnosisRun SHALL store version, transitions, final node, termination reason, degraded flag, and evidence retry count
- **AND** the summary SHALL be derived from the current Run's actual orchestration events
#### Scenario: Graph reaches handled Fallback
- **WHEN** a complex Chat Graph terminates through deterministic Fallback with a safe answer
- **THEN** the current DiagnosisRun SHALL store a non-empty orchestration trace with `degraded=true`
- **AND** the Run status SHALL be SUCCESS
#### Scenario: Unhandled execution fails after events exist
- **WHEN** an unhandled failure occurs after one or more real Graph events are available
- **THEN** the service SHALL best-effort persist a partial orchestration summary for the current failed run
- **AND** it SHALL NOT add a node or transition that did not occur
#### Scenario: Orchestration trace content is inspected
- **WHEN** orchestration trace JSON is serialized
- **THEN** it SHALL NOT include Prompt text, model reasoning, raw tool output, raw Executor output, or Graph State snapshots
- **AND** it SHALL NOT contain data owned by another run
### Requirement: Trace API SHALL expose orchestration trace only on the run object
The Trace API SHALL parse the current DiagnosisRun orchestration JSON and expose it only as `run.orchestrationTrace`.
#### Scenario: Exact StateGraph run trace is queried
- **WHEN** a caller queries a successful new StateGraph Chat run
- **THEN** `run.orchestrationTrace` SHALL be a non-empty parsed JSON object
- **AND** the response top level and compatibility `session` projection SHALL NOT duplicate the field
- **AND** no raw orchestration trace field SHALL be added
#### Scenario: Historical or non-StateGraph run is queried
- **WHEN** the selected DiagnosisRun has null orchestration trace
- **THEN** `run.orchestrationTrace` MAY be null
- **AND** the service SHALL NOT synthesize historical events or read another run's trace
### Requirement: Orchestration trace migration SHALL be additive and nullable
The database migration SHALL add only one nullable JSON column named `orchestration_trace` to `diagnosis_run` for this change.
#### Scenario: Migration is applied
- **WHEN** Flyway applies the stage 3 migration
- **THEN** existing DiagnosisRun rows SHALL remain valid without backfill
- **AND** no other table or column SHALL be changed by the stage 3 schema migration