feat(graph): cut over chat diagnosis stategraph
This commit is contained in:
+60
@@ -0,0 +1,60 @@
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Composer SHALL generate final user-facing Chat answers
|
||||
The system SHALL invoke the Composer Graph Node after Verifier routing to generate the final user-facing Chat answer from Verifier-allowed material.
|
||||
|
||||
#### Scenario: Composer receives only filtered material
|
||||
- **WHEN** the Composer Graph Node invokes its configured Agent
|
||||
- **THEN** the Composer input SHALL contain `original_query`, `verdict`, `allowed_claims`, `allowed_hypotheses`, `missing_info`, `recommended_actions`, and `rationale`
|
||||
- **AND** the Composer input SHALL NOT contain raw tool output
|
||||
- **AND** the Composer input SHALL NOT contain the full unscreened Executor output
|
||||
- **AND** the Composer input SHALL NOT contain Executor `user_facing_answer`
|
||||
|
||||
#### Scenario: Composer outputs strict JSON
|
||||
- **WHEN** Composer completes
|
||||
- **THEN** it SHALL output exactly one JSON object
|
||||
- **AND** the JSON object SHALL include `answer_summary`, `recommended_actions`, and `user_facing_answer`
|
||||
- **AND** it SHALL NOT output Markdown, code fences, or explanatory text outside the JSON object
|
||||
|
||||
#### Scenario: Composer does not introduce new facts
|
||||
- **WHEN** Composer produces `answer_summary`, `recommended_actions`, or `user_facing_answer`
|
||||
- **THEN** every service name, entity, timestamp, error code, metric value, root cause, and recommendation reason SHALL be derived from the Composer input
|
||||
- **AND** Composer SHALL NOT add facts from model knowledge, raw tool history, or Executor raw text
|
||||
|
||||
### Requirement: Composer input SHALL honor Verifier claim checks
|
||||
The Composer Graph Node SHALL construct Composer input by filtering verified Executor structured output through Verifier `claim_checks`.
|
||||
|
||||
#### Scenario: Passing claims become allowed claims
|
||||
- **WHEN** a claim check verification is `direct_observation`
|
||||
- **THEN** the Composer input builder SHALL include the matching verified Executor claim in `allowed_claims`
|
||||
|
||||
#### Scenario: Reasonable inferences remain bounded
|
||||
- **WHEN** a claim check verification is `reasonable_inference`
|
||||
- **THEN** the Composer input builder MAY include the matching verified Executor claim in `allowed_claims`
|
||||
- **AND** the final answer SHALL NOT describe it as the sole confirmed root cause unless the allowed claim itself is a root-cause claim and the final verdict is `PASS`
|
||||
|
||||
#### Scenario: Overstated claims are not confirmed findings
|
||||
- **WHEN** a claim check verification is `overstated`
|
||||
- **THEN** the Composer input builder SHALL NOT include the matching claim as a confirmed item in `allowed_claims`
|
||||
- **AND** it MAY include it as `allowed_hypotheses` or represent it in `missing_info`
|
||||
|
||||
#### Scenario: Unsupported or external claims are withheld
|
||||
- **WHEN** a claim check verification is `unsupported`, `external_unknown`, or `contradicted`
|
||||
- **THEN** the Composer input builder SHALL NOT include the matching claim in `allowed_claims`
|
||||
- **AND** the final user-facing answer SHALL NOT present that claim as confirmed
|
||||
|
||||
### Requirement: Composer failures SHALL degrade safely
|
||||
The system SHALL tolerate malformed or failed Composer execution without leaking raw JSON or unverified Executor material.
|
||||
|
||||
#### Scenario: malformed Composer output falls back safely
|
||||
- **WHEN** Composer exhausts its fixed-input technical retry or returns a non-retryable failure
|
||||
- **THEN** the deterministic Fallback Node SHALL produce a final answer using only filtered material
|
||||
- **AND** the final answer SHALL NOT expose raw Composer output
|
||||
- **AND** the final answer SHALL NOT expose raw Executor output
|
||||
- **AND** the final answer SHALL NOT use Executor `user_facing_answer`
|
||||
|
||||
#### Scenario: Composer audit is persisted
|
||||
- **WHEN** Graph result mapping persists verifier evaluation
|
||||
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation.composer_output` SHALL record parsed Composer audit when available
|
||||
- **AND** handled Composer fallback SHALL be observable through orchestration trace and available status/reason fields
|
||||
- **AND** the audit SHALL remain compact and SHALL NOT store full raw tool output
|
||||
+107
@@ -0,0 +1,107 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: Complex Chat SHALL use the real Diagnosis StateGraph as its only production orchestrator
|
||||
|
||||
The system SHALL execute each complex Chat request through one real Diagnosis CompiledGraph and SHALL NOT create, invoke, or fall back to a SequentialAgent workflow.
|
||||
|
||||
#### Scenario: Complex Chat starts
|
||||
|
||||
- **WHEN** `executeChatWithStrategy` classifies a valid question as complex
|
||||
- **THEN** ChatService SHALL create the current Diagnosis Run and invoke one Diagnosis CompiledGraph
|
||||
- **AND** the production path SHALL NOT maintain an outer score-based retry loop
|
||||
|
||||
#### Scenario: Graph dependencies are assembled
|
||||
|
||||
- **WHEN** the complex Chat Graph is constructed
|
||||
- **THEN** Planner, Executor, Verifier, and Composer SHALL use the existing project Prompt, tool, skill, and AgentLoggingHook assembly rules
|
||||
- **AND** the Graph Verifier SHALL NOT register VerifierInputHook
|
||||
- **AND** Gatekeeper SHALL run only as the explicit Graph Node
|
||||
|
||||
### Requirement: Graph invocation SHALL preserve current Run ownership
|
||||
|
||||
The system SHALL use the current `runId` as Graph `threadId` and SHALL pass both current `sessionId` and `runId` in RunnableConfig metadata.
|
||||
|
||||
#### Scenario: Agent Node runs
|
||||
|
||||
- **WHEN** any Agent Node is invoked for a complex Chat run
|
||||
- **THEN** its RunnableConfig threadId SHALL equal the current runId
|
||||
- **AND** AgentStep and ToolInvocation writes SHALL retain the current sessionId and runId
|
||||
|
||||
#### Scenario: Gatekeeper validates Executor output
|
||||
|
||||
- **WHEN** the Gatekeeper Node runs
|
||||
- **THEN** it SHALL validate only tool invocations belonging to the runId in RunnableConfig metadata
|
||||
- **AND** data from another run in the same session SHALL NOT be considered
|
||||
|
||||
### Requirement: Initial Graph State SHALL contain only bounded diagnosis control data
|
||||
|
||||
ChatService SHALL initialize Diagnosis Graph State with the current query context, NORMAL Planner mode, zero independent retry counters, and an empty orchestration event list.
|
||||
|
||||
#### Scenario: Initial state is projected
|
||||
|
||||
- **WHEN** a complex Chat run enters Planner for the first time
|
||||
- **THEN** `diagnosis_context` SHALL contain the current query/original query
|
||||
- **AND** complete conversation history SHALL NOT be stored in parent Graph State
|
||||
- **AND** history MAY remain in the Planner and Executor system Prompt assembled for this request
|
||||
|
||||
#### Scenario: Counters are initialized
|
||||
|
||||
- **WHEN** the Graph starts
|
||||
- **THEN** planner, verifier, composer, and evidence retry counters SHALL each be zero
|
||||
- **AND** Planner mode SHALL be NORMAL
|
||||
|
||||
### Requirement: Graph final state SHALL map to the existing Chat result and Run lifecycle
|
||||
|
||||
The system SHALL use non-empty `final_answer` from a handled Graph terminal state as the existing ChatResult answer. Run status SHALL express execution lifecycle rather than diagnosis quality.
|
||||
|
||||
#### Scenario: Composer completes
|
||||
|
||||
- **WHEN** Graph reaches Composer and produces a safe non-empty final answer
|
||||
- **THEN** ChatResult SHALL preserve the current answer/sessionId/runId protocol
|
||||
- **AND** DiagnosisRun SHALL be saved as SUCCESS with answer, duration, token count, step count, and tool count
|
||||
- **AND** EvaluationService SHALL evaluate the current runId
|
||||
|
||||
#### Scenario: Handled Fallback completes
|
||||
|
||||
- **WHEN** Graph reaches deterministic Fallback and produces a safe non-empty answer
|
||||
- **THEN** DiagnosisRun SHALL be SUCCESS
|
||||
- **AND** diagnosis quality SHALL be expressed by verifier fields when available or `orchestrationTrace.degraded=true`
|
||||
- **AND** no DEGRADED Run status SHALL be introduced
|
||||
|
||||
#### Scenario: Graph cannot produce a safe response
|
||||
|
||||
- **WHEN** Graph has an unhandled failure, final state is unavailable, final answer is blank, or required successful-result persistence fails
|
||||
- **THEN** DiagnosisRun SHALL be marked FAILED when it can still be saved
|
||||
- **AND** the system SHALL NOT report a successful Graph result
|
||||
|
||||
### Requirement: Graph state SHALL be the only source for verifier evaluation persistence
|
||||
|
||||
The system SHALL build `diagnosis_run.self_evaluation.verifier_evaluation` from explicit Graph State and SHALL NOT read VerifierContextHolder on the complex Chat path.
|
||||
|
||||
#### Scenario: Verifier completed
|
||||
|
||||
- **WHEN** final Graph State contains a completed Verifier result
|
||||
- **THEN** verifier evaluation SHALL include execution status, model verdict, effective verdict, groundedness, claim/fact checks, rationale, round, Gatekeeper audit, verified Executor output/evidence, Prompt audit, and Composer audit when available
|
||||
- **AND** the compatibility `verdict` field SHALL equal effective verdict
|
||||
|
||||
#### Scenario: Pre-verification Fallback completed
|
||||
|
||||
- **WHEN** Graph reaches Fallback before Verifier completes
|
||||
- **THEN** verifier evaluation SHALL preserve available execution statuses, Gatekeeper audit, Prompt audit, and fallback context
|
||||
- **AND** it SHALL NOT fabricate a model or effective verdict
|
||||
|
||||
#### Scenario: Evaluation payload is inspected
|
||||
|
||||
- **WHEN** verifier evaluation is persisted
|
||||
- **THEN** it SHALL NOT contain raw Executor text or complete tool trace summary
|
||||
- **AND** `executor_structured_output` SHALL contain at most the verified projection retained for compatibility
|
||||
|
||||
### Requirement: Stage 3 verification SHALL not run the final live E2E
|
||||
|
||||
The change SHALL use focused automated tests for production cutover and contracts while reserving Maven live startup, log inspection, and database querying for stage 5.
|
||||
|
||||
#### Scenario: Stage 3 is accepted
|
||||
|
||||
- **WHEN** stage 3 verification completes
|
||||
- **THEN** Graph cutover, Run/Trace, Prompt, Controller/Eval regressions, Maven test compilation, and OpenSpec strict validation SHALL have passed
|
||||
- **AND** live E2E, `logs/`, and `scripts/query_mysql.py` SHALL be recorded as intentionally deferred to stage 5
|
||||
+7
@@ -0,0 +1,7 @@
|
||||
## REMOVED Requirements
|
||||
|
||||
### Requirement: Real Nodes SHALL remain isolated from the production Chat path in stage 2
|
||||
|
||||
**Reason**: Stage 2 production isolation has completed its migration purpose; stage 3 intentionally makes the real Diagnosis Graph the only complex Chat production orchestrator.
|
||||
|
||||
**Migration**: Replace `ChatService.executeChatComplex` SequentialAgent orchestration with real Graph assembly/invocation. Keep `/api/chat` compatible and use a full Git revert of stage 3 for runtime rollback rather than retaining dual paths.
|
||||
+263
@@ -0,0 +1,263 @@
|
||||
## MODIFIED Requirements
|
||||
|
||||
### Requirement: Verifier SHALL fact-check Executor answers
|
||||
The system SHALL have a Verifier Agent that reads only Gatekeeper-projected structured Executor claims and verified claim-local evidence, then produces a structured verdict based on claim derivability.
|
||||
|
||||
#### Scenario: PASS verdict when all claims have evidence
|
||||
- **WHEN** all critical claims in `verified_executor_output.claims` have direct observation or reasonable inference support in `verified_evidence`
|
||||
- **AND** at least one critical claim has direct observation
|
||||
- **AND** no critical claim is contradicted, unsupported, external unknown, or overstated
|
||||
- **AND** verdict ceiling is PASS
|
||||
- **THEN** the Verifier MAY output model verdict="PASS" with groundedness_score ≥ 0.5
|
||||
|
||||
#### Scenario: LOW_CONFID verdict with partial evidence
|
||||
- **WHEN** no critical claim contradicts verified evidence
|
||||
- **AND** some critical claims are `unsupported`, `external_unknown`, or `overstated`
|
||||
- **THEN** the Verifier SHALL output model verdict="LOW_CONFID"
|
||||
|
||||
#### Scenario: LOW_CONFID verdict with only inference support
|
||||
- **WHEN** no critical claim contradicts verified evidence
|
||||
- **AND** all critical claims are only `reasonable_inference`
|
||||
- **THEN** the Verifier SHALL output model verdict="LOW_CONFID"
|
||||
|
||||
#### Scenario: REJECT verdict when claims contradict evidence
|
||||
- **WHEN** any critical claim in `verified_executor_output.claims` contradicts verified evidence
|
||||
- **OR** the claim fabricates a key entity, error code, or conclusion that does not exist in verified evidence
|
||||
- **THEN** the Verifier SHALL output model verdict="REJECT"
|
||||
|
||||
#### Scenario: Verified structured claims are the only verification target
|
||||
- **WHEN** `verified_executor_output.claims` is present
|
||||
- **THEN** Verifier SHALL verify each structured claim against matching `verified_evidence` through `claim_checks`
|
||||
- **AND** each claim's evidence references SHALL match existing claim/invocation/tool/path identifiers when available
|
||||
- **AND** a claim without matching verified evidence SHALL NOT be classified as `direct_observation`
|
||||
- **AND** Verifier SHALL NOT receive or add confirmed facts from raw Executor text
|
||||
|
||||
#### Scenario: Executor output is invalid
|
||||
- **WHEN** Executor does not return a legal structured contract
|
||||
- **THEN** the Graph SHALL route directly to pre-verification Fallback
|
||||
- **AND** Verifier SHALL NOT execute or fabricate a diagnostic verdict
|
||||
|
||||
### Requirement: facts_checked SHALL use a fixed classification set
|
||||
The system SHALL continue to expose compatibility `facts_checked` using its fixed verification classification set.
|
||||
|
||||
#### Scenario: claim checks are mapped to legacy facts
|
||||
- **WHEN** Verifier output contains `claim_checks`
|
||||
- **THEN** the shared Verifier protocol parser SHALL derive compatibility `facts_checked` when the model did not provide them
|
||||
- **AND** `direct_observation` SHALL map to `direct_evidence`
|
||||
- **AND** `reasonable_inference` and `overstated` SHALL map to `indirect_support`
|
||||
- **AND** `unsupported` and `external_unknown` SHALL map to `no_evidence`
|
||||
- **AND** `contradicted` SHALL map to `contradicted`
|
||||
|
||||
### Requirement: ChatService SHALL route based on Verifier verdict
|
||||
The system SHALL use Diagnosis StateGraph conditional edges, rather than a ChatService outer loop, to route explicit Verifier execution status and effective verdict.
|
||||
|
||||
#### Scenario: PASS routes to Composer
|
||||
- **WHEN** Verifier completes with effective verdict="PASS"
|
||||
- **THEN** the Graph SHALL invoke Composer with filtered Verifier-allowed material
|
||||
- **AND** the final user-facing answer SHALL NOT pass through raw Executor output
|
||||
- **AND** the final user-facing answer SHALL NOT read Executor `user_facing_answer`
|
||||
|
||||
#### Scenario: LOW_CONFID does not qualify for evidence retry
|
||||
- **WHEN** Verifier completes LOW_CONFID but ceiling is LOW_CONFID, no valid critical evidence gap exists, or evidence retry count is already one
|
||||
- **THEN** the Graph SHALL route to Composer without another Planner cycle
|
||||
- **AND** the final answer SHALL distinguish confirmed information, possible directions, and evidence gaps
|
||||
|
||||
#### Scenario: LOW_CONFID qualifies for evidence retry
|
||||
- **WHEN** Verifier completes LOW_CONFID with ceiling PASS, at least one critical valid evidence gap, and evidence retry count zero
|
||||
- **THEN** the Graph SHALL invoke one EVIDENCE_GAP_ONLY Planner cycle
|
||||
- **AND** it SHALL NOT use groundedness threshold or a ChatService feature flag to decide the retry
|
||||
|
||||
#### Scenario: REJECT does not enter retry round
|
||||
- **WHEN** Verifier completes with effective verdict="REJECT"
|
||||
- **THEN** the Graph SHALL NOT start an evidence supplementation round
|
||||
- **AND** it SHALL route to Composer-safe output
|
||||
|
||||
#### Scenario: REJECT produces bounded output
|
||||
- **WHEN** effective verdict is REJECT
|
||||
- **THEN** the system SHALL output a degraded result indicating current evidence cannot support a reliable conclusion
|
||||
- **AND** it SHALL NOT pass through raw Executor answer
|
||||
- **AND** it SHALL NOT include an unsupported root-cause conclusion
|
||||
|
||||
#### Scenario: Verifier execution fails
|
||||
- **WHEN** Verifier exhausts technical retry or returns a non-retryable failure
|
||||
- **THEN** the Graph SHALL route to pre-verification Fallback
|
||||
- **AND** no execution status string SHALL be used as model or effective verdict
|
||||
|
||||
### Requirement: Verifier SHALL be observable
|
||||
The Verifier execution, effective verdict, and downstream final-answer composition SHALL be persisted in the current Diagnosis Run self-evaluation container.
|
||||
|
||||
#### Scenario: claim checks written to self_evaluation
|
||||
- **WHEN** a completed Verifier evaluation is persisted
|
||||
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include `claim_checks`
|
||||
- **AND** it SHALL continue to include compatibility `facts_checked`
|
||||
- **AND** it SHALL include `verifier_status`, `model_verdict`, `effective_verdict`, `verdict`, `groundedness_score`, `rationale`, verified output/evidence, and Gatekeeper audit
|
||||
|
||||
#### Scenario: composer output written to self_evaluation
|
||||
- **WHEN** final answer composition completes
|
||||
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include compact `composer_output` when available
|
||||
- **AND** handled Composer fallback SHALL remain observable through orchestration trace and status/reason fields
|
||||
- **AND** existing claim/fact and Gatekeeper fields SHALL be preserved
|
||||
|
||||
#### Scenario: verdict written to self_evaluation
|
||||
- **WHEN** Verifier completes
|
||||
- **THEN** Graph result mapping SHALL write effective verdict under `diagnosis_run.self_evaluation.verifier_evaluation.verdict`
|
||||
- **AND** existing `rule_evaluation` and `aiops_rule_evaluation` channels SHALL be preserved
|
||||
|
||||
#### Scenario: pre-verification fallback is persisted
|
||||
- **WHEN** Graph reaches Fallback before Verifier completes
|
||||
- **THEN** verifier evaluation SHALL include available status, Gatekeeper audit, failure reason, and Prompt audit
|
||||
- **AND** it SHALL NOT fabricate `model_verdict` or `effective_verdict`
|
||||
|
||||
#### Scenario: gatekeeper result written to self_evaluation
|
||||
- **WHEN** Graph result mapping persists available Gatekeeper state
|
||||
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result`
|
||||
- **AND** the result SHALL retain status, severity, checked bindings, rules, failed rules, warnings, and errors when provided by Gatekeeper
|
||||
|
||||
#### Scenario: prompt audit written to verifier evaluation
|
||||
- **WHEN** a complex Chat Graph result is persisted
|
||||
- **THEN** the system SHALL include a `prompt_audit` object under `diagnosis_run.self_evaluation.verifier_evaluation`
|
||||
- **AND** `prompt_audit.version` SHALL identify the Chat Prompt audit catalog version
|
||||
- **AND** `prompt_audit.prompts` SHALL include Planner, Executor, Verifier, and Composer Prompt names and versions
|
||||
- **AND** full Prompt text SHALL NOT be persisted
|
||||
|
||||
#### Scenario: prompt audit available on fallback paths
|
||||
- **WHEN** Planner, Executor, Gatekeeper, Verifier, or Composer reaches a handled Fallback
|
||||
- **THEN** the persisted verifier evaluation SHALL still include `prompt_audit`
|
||||
|
||||
#### Scenario: evaluation payload is inspected
|
||||
- **WHEN** Graph verifier evaluation is persisted
|
||||
- **THEN** it SHALL NOT contain raw Executor text or complete `tool_trace_summary`
|
||||
- **AND** compatibility `executor_structured_output` SHALL contain at most the verified projection
|
||||
|
||||
### Requirement: Verifier SHALL consume explicit verification inputs
|
||||
The Verifier SHALL receive a Graph-built verified-only payload rather than inferring business inputs from conversation history, ThreadLocal state, raw Executor text, or complete tool history.
|
||||
|
||||
#### Scenario: explicit input blocks available to Verifier
|
||||
- **WHEN** the Verifier Graph Node starts
|
||||
- **THEN** the payload SHALL provide `diagnosis_context`, `verified_executor_output`, `verified_evidence`, `gatekeeper_audit`, and `verdict_ceiling`
|
||||
- **AND** permitted structured `retry_context` SHALL be provided only after evidence retry preparation
|
||||
|
||||
#### Scenario: Verifier remains isolated from intermediate and raw material
|
||||
- **WHEN** the Verifier input is serialized
|
||||
- **THEN** it SHALL exclude Planner reasoning, Executor intermediate reasoning, raw Executor text, complete tool trace summary, Prompt text, and unrelated parent Graph State
|
||||
|
||||
#### Scenario: only passed bindings are available
|
||||
- **WHEN** Gatekeeper returns mixed passed and failed checked bindings
|
||||
- **THEN** `verified_executor_output` and `verified_evidence` SHALL contain only claims/material matching passed bindings
|
||||
- **AND** the Verifier SHALL NOT receive failed or unreferenced tool material
|
||||
|
||||
#### Scenario: verified evidence preserves precise references
|
||||
- **WHEN** the system prepares Verifier input
|
||||
- **THEN** each verified evidence item SHALL preserve claim id, source invocation id, tool name, raw path, and matched text
|
||||
- **AND** the item SHALL be traceable to current-run Gatekeeper validation
|
||||
|
||||
#### Scenario: gatekeeper audit and ceiling are available
|
||||
- **WHEN** the system prepares Verifier input
|
||||
- **THEN** the payload SHALL include raw Gatekeeper audit separately from normalized verdict ceiling
|
||||
- **AND** a LOW_CONFID ceiling SHALL prevent effective PASS
|
||||
|
||||
#### Scenario: technical retry occurs
|
||||
- **WHEN** the first Verifier attempt returns invalid output or a retryable invocation failure
|
||||
- **THEN** the second attempt SHALL receive byte-identical serialized input
|
||||
- **AND** Executor, Gatekeeper, and tools SHALL NOT rerun
|
||||
|
||||
### Requirement: Verifier facts SHALL be auditable
|
||||
Verifier claims and facts SHALL be linkable to the verified binding projection used during verification.
|
||||
|
||||
#### Scenario: claim checks contain evidence refs
|
||||
- **WHEN** the Verifier emits `claim_checks`
|
||||
- **THEN** each check SHALL include an `evidence_refs` array
|
||||
- **AND** any non-empty evidence ref SHALL correspond to existing verified evidence by claim id, source invocation id, tool name, or raw path
|
||||
- **AND** it SHALL NOT reference a failed or unverified binding
|
||||
|
||||
#### Scenario: verifier evaluation persists traceability snapshot
|
||||
- **WHEN** Graph result mapping persists verifier evaluation
|
||||
- **THEN** it SHALL include `traceability_version`
|
||||
- **AND** it SHALL include the bounded `verified_evidence` snapshot used by the Verifier
|
||||
- **AND** it SHALL NOT persist a complete tool trace summary as Verifier input
|
||||
|
||||
### Requirement: Structured Executor output SHALL degrade safely
|
||||
The StateGraph runtime SHALL tolerate malformed or absent structured Executor output without crashing the Chat flow or invoking Verifier with untrusted material.
|
||||
|
||||
#### Scenario: Malformed Executor JSON is classified
|
||||
- **WHEN** Executor returns malformed JSON or text outside the expected contract
|
||||
- **THEN** Executor Node SHALL set INVALID_OUTPUT
|
||||
- **AND** the Graph SHALL route directly to deterministic pre-verification Fallback
|
||||
- **AND** Gatekeeper, Verifier, and model Composer SHALL NOT execute
|
||||
|
||||
#### Scenario: Structured parse failure remains observable
|
||||
- **WHEN** Executor output parsing fails
|
||||
- **THEN** orchestration events and verifier evaluation status/failure fields SHALL make the parse failure visible
|
||||
- **AND** the failure SHALL NOT be treated as a successful evidence-attribution contract or diagnostic verdict
|
||||
|
||||
### Requirement: Executor Gatekeeper SHALL validate deterministic structured-output failures
|
||||
The system SHALL run deterministic Gatekeeper checks as an explicit Graph Node after legal Executor output parsing and before Verifier model execution.
|
||||
|
||||
#### Scenario: schema rule rejects removed fields
|
||||
- **WHEN** Executor structured output contains `diagnosis_summary` or `user_facing_answer`
|
||||
- **THEN** `gatekeeper_result.status` SHALL be `fail`
|
||||
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
|
||||
|
||||
#### Scenario: schema rule rejects missing evidence bindings
|
||||
- **WHEN** a confirmed claim has no `evidence_bindings`
|
||||
- **THEN** `gatekeeper_result.status` SHALL be `fail`
|
||||
- **AND** `gatekeeper_result.failed_rules` SHALL contain `schema.executor_v2`
|
||||
|
||||
#### Scenario: invocation rule rejects fabricated invocation ids
|
||||
- **WHEN** a claim evidence binding references an invocation id absent from current-run `tool_invocation` rows
|
||||
- **THEN** `gatekeeper_result.status` SHALL be `fail`
|
||||
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
|
||||
|
||||
#### Scenario: invocation rule rejects tool name mismatch
|
||||
- **WHEN** a claim evidence binding references an existing current-run invocation id
|
||||
- **AND** binding `tool_name` does not match persisted invocation `tool_name`
|
||||
- **THEN** `gatekeeper_result.status` SHALL be `fail`
|
||||
- **AND** `gatekeeper_result.failed_rules` SHALL contain `evidence.invocation_ref`
|
||||
|
||||
#### Scenario: valid structured output passes initial gatekeeper rules
|
||||
- **WHEN** Executor emits legal `executor_evidence_v2`
|
||||
- **AND** each claim has evidence bindings pointing to current-run invocations with matching tool names and paths
|
||||
- **THEN** `gatekeeper_result.status` SHALL be `pass`
|
||||
- **AND** `gatekeeper_result.failed_rules` SHALL be empty
|
||||
|
||||
#### Scenario: gatekeeper reject bypasses Verifier
|
||||
- **WHEN** normalized Gatekeeper status is REJECT
|
||||
- **THEN** the Graph SHALL route directly to pre-verification Fallback
|
||||
- **AND** Verifier SHALL NOT execute
|
||||
|
||||
#### Scenario: gatekeeper low confidence is bounded
|
||||
- **WHEN** normalized Gatekeeper status is LOW_CONFID with at least one passed binding
|
||||
- **THEN** verified input SHALL contain only passed bindings
|
||||
- **AND** effective verdict SHALL NOT exceed LOW_CONFID
|
||||
|
||||
### Requirement: Verifier SHALL use verified claim-local evidence for derivability
|
||||
Verifier SHALL judge structured claims only against Gatekeeper-verified claim-local evidence excerpts and their precise current-run references.
|
||||
|
||||
#### Scenario: Verified excerpt supports direct observation
|
||||
- **WHEN** verdict ceiling is PASS
|
||||
- **AND** a claim's verified evidence matched text directly contains the claim's concrete facts
|
||||
- **THEN** Verifier MAY classify that claim as `direct_observation`
|
||||
|
||||
#### Scenario: Verified evidence is complete Verifier context
|
||||
- **WHEN** verified claims and evidence are available
|
||||
- **THEN** Verifier SHALL use them as its evidence context
|
||||
- **AND** it SHALL NOT require or request a complete tool trace summary
|
||||
- **AND** it SHALL NOT read raw Executor or unreferenced tool material
|
||||
|
||||
### Requirement: Gatekeeper severity SHALL constrain effective verdict
|
||||
Runtime effective verdict calculation SHALL treat normalized Gatekeeper ceiling as a hard upper bound independent from Verifier model output.
|
||||
|
||||
#### Scenario: Reject severity bypasses Verifier
|
||||
- **WHEN** `gatekeeper_result.severity=reject`
|
||||
- **THEN** the Graph SHALL route to pre-verification Fallback without invoking Verifier
|
||||
- **AND** it SHALL NOT fabricate an effective diagnostic verdict
|
||||
|
||||
#### Scenario: Low confidence severity prevents PASS
|
||||
- **WHEN** `gatekeeper_result.severity=low_confid`
|
||||
- **AND** the Verifier model returns `verdict=PASS`
|
||||
- **THEN** deterministic effective-verdict calculation SHALL downgrade the result
|
||||
- **AND** effective verdict SHALL be `LOW_CONFID`
|
||||
|
||||
#### Scenario: Gatekeeper audit includes severity
|
||||
- **WHEN** Graph verifier evaluation is persisted
|
||||
- **THEN** `diagnosis_run.self_evaluation.verifier_evaluation.gatekeeper_result` SHALL include available `status`, `severity`, `checked_bindings`, `failed_rules`, `warnings`, and `errors`
|
||||
+56
@@ -0,0 +1,56 @@
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: StateGraph Chat runs SHALL persist a compact orchestration trace
|
||||
|
||||
Each successful new StateGraph complex Chat run SHALL persist a non-empty compact orchestration summary derived from its bounded Graph events in `diagnosis_run.orchestration_trace`.
|
||||
|
||||
#### Scenario: Graph reaches Composer
|
||||
|
||||
- **WHEN** a complex Chat Graph terminates through Composer with a safe answer
|
||||
- **THEN** the current DiagnosisRun SHALL store version, transitions, final node, termination reason, degraded flag, and evidence retry count
|
||||
- **AND** the summary SHALL be derived from the current Run's actual orchestration events
|
||||
|
||||
#### Scenario: Graph reaches handled Fallback
|
||||
|
||||
- **WHEN** a complex Chat Graph terminates through deterministic Fallback with a safe answer
|
||||
- **THEN** the current DiagnosisRun SHALL store a non-empty orchestration trace with `degraded=true`
|
||||
- **AND** the Run status SHALL be SUCCESS
|
||||
|
||||
#### Scenario: Unhandled execution fails after events exist
|
||||
|
||||
- **WHEN** an unhandled failure occurs after one or more real Graph events are available
|
||||
- **THEN** the service SHALL best-effort persist a partial orchestration summary for the current failed run
|
||||
- **AND** it SHALL NOT add a node or transition that did not occur
|
||||
|
||||
#### Scenario: Orchestration trace content is inspected
|
||||
|
||||
- **WHEN** orchestration trace JSON is serialized
|
||||
- **THEN** it SHALL NOT include Prompt text, model reasoning, raw tool output, raw Executor output, or Graph State snapshots
|
||||
- **AND** it SHALL NOT contain data owned by another run
|
||||
|
||||
### Requirement: Trace API SHALL expose orchestration trace only on the run object
|
||||
|
||||
The Trace API SHALL parse the current DiagnosisRun orchestration JSON and expose it only as `run.orchestrationTrace`.
|
||||
|
||||
#### Scenario: Exact StateGraph run trace is queried
|
||||
|
||||
- **WHEN** a caller queries a successful new StateGraph Chat run
|
||||
- **THEN** `run.orchestrationTrace` SHALL be a non-empty parsed JSON object
|
||||
- **AND** the response top level and compatibility `session` projection SHALL NOT duplicate the field
|
||||
- **AND** no raw orchestration trace field SHALL be added
|
||||
|
||||
#### Scenario: Historical or non-StateGraph run is queried
|
||||
|
||||
- **WHEN** the selected DiagnosisRun has null orchestration trace
|
||||
- **THEN** `run.orchestrationTrace` MAY be null
|
||||
- **AND** the service SHALL NOT synthesize historical events or read another run's trace
|
||||
|
||||
### Requirement: Orchestration trace migration SHALL be additive and nullable
|
||||
|
||||
The database migration SHALL add only one nullable JSON column named `orchestration_trace` to `diagnosis_run` for this change.
|
||||
|
||||
#### Scenario: Migration is applied
|
||||
|
||||
- **WHEN** Flyway applies the stage 3 migration
|
||||
- **THEN** existing DiagnosisRun rows SHALL remain valid without backfill
|
||||
- **AND** no other table or column SHALL be changed by the stage 3 schema migration
|
||||
Reference in New Issue
Block a user