5.1 KiB
5.1 KiB
MODIFIED Requirements
Requirement: ChatService SHALL route based on Verifier verdict
The system SHALL use ChatService for explicit single-round Planner -> Executor -> Verifier -> Composer orchestration and SHALL use ChatService to control whether an additional round is allowed.
Scenario: PASS -> Composer output
- WHEN Verifier outputs verdict="PASS"
- THEN ChatService SHALL filter Verifier-allowed material and invoke Composer or a safe fixed template
- AND the final user-facing answer SHALL NOT pass through raw Executor output
- AND the final user-facing answer SHALL NOT read Executor
user_facing_answer
Scenario: LOW_CONFID score>=0.5 -> Composer output with uncertainty
- WHEN Verifier outputs verdict="LOW_CONFID" with groundedness_score >= 0.5
- THEN ChatService SHALL filter Verifier-allowed material and invoke Composer or a safe fixed template
- AND the final user-facing answer SHALL distinguish confirmed information, possible directions, and evidence gaps
- AND unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
Scenario: LOW_CONFID score<0.5 -> trigger one additional round
- WHEN Verifier outputs verdict="LOW_CONFID" with groundedness_score < 0.5 and this is the first callback
- THEN the ChatService SHALL invoke one additional
Planner -> Executor -> Verifierround to supplement evidence - AND after the second Verifier run, verdict="LOW_CONFID" SHALL be routed to Composer or a safe fixed template
- AND after the second Verifier run, verdict="REJECT" SHALL still produce a degraded Composer-safe output
Scenario: REJECT does not enter retry round
- WHEN Verifier outputs verdict="REJECT"
- THEN the system SHALL NOT start a retry round for evidence supplementation
- AND it SHALL produce a degraded output directly through Composer-safe rendering
Scenario: REJECT -> degraded output
- WHEN Verifier outputs verdict="REJECT"
- THEN the system SHALL output a degraded result indicating the answer cannot be reliably generated
- AND it SHALL NOT pass through the raw Executor answer
- AND it SHALL NOT include a root-cause conclusion
Requirement: User-facing verifier outputs SHALL follow Composer-safe protocols
The system SHALL use Composer or fixed safe templates for PASS, LOW_CONFID, and REJECT user-facing responses.
Scenario: PASS uses Composer-safe output
- WHEN the final verdict is
PASS - THEN the user-facing response SHALL be generated from Verifier-allowed material through Composer or a safe fixed template
- AND it SHALL NOT use Executor
user_facing_answer - AND it SHALL NOT expose raw Executor JSON
Scenario: LOW_CONFID uses Composer-safe uncertainty output
- WHEN the final verdict is
LOW_CONFID - THEN the user-facing response SHALL include only confirmed facts, possible directions, evidence gaps, and next-step suggestions derived from Verifier-allowed material
- AND optional evidence gaps, if present, SHALL come only from verifier-identified gaps
- AND unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
Scenario: REJECT uses degraded template
- WHEN the final verdict is
REJECT - THEN the user-facing response SHALL use a degraded template or Composer-safe degraded output
- AND it SHALL include only confirmed facts, evidence gaps, and next-step suggestions
- AND it SHALL NOT include unverified raw answer content
- AND it SHALL NOT include a root-cause conclusion
Requirement: Verifier SHALL be observable
The Verifier's verdict and downstream final-answer composition SHALL be persisted for observability.
Scenario: claim checks written to self_evaluation
- WHEN the Verifier evaluation is persisted
- THEN
diagnosis_session.self_evaluation.verifier_evaluationSHALL includeclaim_checks - AND it SHALL continue to include compatibility
facts_checked - AND existing fields such as
verdict,groundedness_score,rationale,executor_output_parse_status,tool_trace_summary, andgatekeeper_resultSHALL be preserved
Scenario: composer output written to self_evaluation
- WHEN final answer composition completes
- THEN
diagnosis_session.self_evaluation.verifier_evaluationSHALL includecomposer_output - AND
composer_outputSHALL indicate whether parsed Composer output or fallback rendering was used - AND existing verifier fields such as
claim_checks,facts_checked,gatekeeper_result, andtool_trace_summarySHALL be preserved
Scenario: verdict written to self_evaluation
- WHEN the Verifier produces a verdict
- THEN the ChatService SHALL write the verdict data under
diagnosis_session.self_evaluation.verifier_evaluation - AND existing
rule_evaluationdata SHALL be preserved
Scenario: gatekeeper result written to self_evaluation
- WHEN the Verifier evaluation is persisted
- THEN
diagnosis_session.self_evaluation.verifier_evaluationSHALL includegatekeeper_result - AND existing verifier fields such as
verdict,facts_checked,executor_output_parse_status, andtool_trace_summarySHALL be preserved