feat(agent): add composer final answer
This commit is contained in:
@@ -89,48 +89,38 @@ The groundedness score SHALL be computed from critical fact classifications inst
|
||||
- **AND** the result SHALL be clamped into `[0.0, 1.0]`
|
||||
|
||||
### Requirement: ChatService SHALL route based on Verifier verdict
|
||||
The system SHALL use ChatService for explicit single-round `Planner → Executor → Verifier` orchestration and SHALL use ChatService to control whether an additional round is allowed.
|
||||
The system SHALL use ChatService for explicit single-round `Planner -> Executor -> Verifier -> Composer` orchestration and SHALL use ChatService to control whether an additional round is allowed.
|
||||
|
||||
#### Scenario: PASS → direct output
|
||||
#### Scenario: PASS -> Composer output
|
||||
- **WHEN** Verifier outputs verdict="PASS"
|
||||
- **THEN** the system SHALL output the Executor's answer directly
|
||||
- **THEN** ChatService SHALL filter Verifier-allowed material and invoke Composer or a safe fixed template
|
||||
- **AND** the final user-facing answer SHALL NOT pass through raw Executor output
|
||||
- **AND** the final user-facing answer SHALL NOT read Executor `user_facing_answer`
|
||||
|
||||
#### Scenario: LOW_CONFID score≥0.5 → output with disclaimer
|
||||
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score ≥ 0.5
|
||||
- **THEN** the system SHALL output the Executor's answer prefixed with a fixed confidence disclaimer
|
||||
#### Scenario: LOW_CONFID score>=0.5 -> Composer output with uncertainty
|
||||
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score >= 0.5
|
||||
- **THEN** ChatService SHALL filter Verifier-allowed material and invoke Composer or a safe fixed template
|
||||
- **AND** the final user-facing answer SHALL distinguish confirmed information, possible directions, and evidence gaps
|
||||
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
|
||||
|
||||
#### Scenario: LOW_CONFID score<0.5 → trigger one additional round
|
||||
#### Scenario: LOW_CONFID score<0.5 -> trigger one additional round
|
||||
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score < 0.5 and this is the first callback
|
||||
- **THEN** the ChatService SHALL invoke one additional `Planner → Executor → Verifier` round to supplement evidence
|
||||
- **AND** after the second Verifier run, verdict="LOW_CONFID" SHALL be output with a confidence disclaimer
|
||||
- **AND** after the second Verifier run, verdict="REJECT" SHALL still produce a degraded output
|
||||
- **THEN** the ChatService SHALL invoke one additional `Planner -> Executor -> Verifier` round to supplement evidence
|
||||
- **AND** after the second Verifier run, verdict="LOW_CONFID" SHALL be routed to Composer or a safe fixed template
|
||||
- **AND** after the second Verifier run, verdict="REJECT" SHALL still produce a degraded Composer-safe output
|
||||
|
||||
#### Scenario: REJECT does not enter retry round
|
||||
- **WHEN** Verifier outputs verdict="REJECT"
|
||||
- **THEN** the system SHALL NOT start a retry round for evidence补充
|
||||
- **AND** it SHALL produce a degraded output directly
|
||||
- **THEN** the system SHALL NOT start a retry round for evidence supplementation
|
||||
- **AND** it SHALL produce a degraded output directly through Composer-safe rendering
|
||||
|
||||
#### Scenario: REJECT → degraded output
|
||||
#### Scenario: REJECT -> degraded output
|
||||
- **WHEN** Verifier outputs verdict="REJECT"
|
||||
- **THEN** the system SHALL output a degraded result indicating the answer cannot be reliably generated
|
||||
- **AND** it SHALL NOT pass through the raw Executor answer
|
||||
|
||||
### Requirement: User-facing verifier outputs SHALL follow fixed templates
|
||||
The system SHALL use fixed output protocols for LOW_CONFID and REJECT user-facing responses.
|
||||
|
||||
#### Scenario: LOW_CONFID uses disclaimer template
|
||||
- **WHEN** the final verdict is `LOW_CONFID`
|
||||
- **THEN** the user-facing response SHALL prepend a fixed disclaimer before the Executor answer
|
||||
- **AND** optional evidence gaps, if present, SHALL come only from verifier-identified critical gaps
|
||||
|
||||
#### Scenario: REJECT uses degraded template
|
||||
- **WHEN** the final verdict is `REJECT`
|
||||
- **THEN** the user-facing response SHALL use a degraded template
|
||||
- **AND** it SHALL include only confirmed facts, evidence gaps, and next-step suggestions
|
||||
- **AND** it SHALL NOT include unverified raw answer content
|
||||
|
||||
- **AND** it SHALL NOT include a root-cause conclusion
|
||||
### Requirement: Verifier SHALL be observable
|
||||
The Verifier's verdict SHALL be persisted for observability.
|
||||
The Verifier's verdict and downstream final-answer composition SHALL be persisted for observability.
|
||||
|
||||
#### Scenario: claim checks written to self_evaluation
|
||||
- **WHEN** the Verifier evaluation is persisted
|
||||
@@ -138,6 +128,12 @@ The Verifier's verdict SHALL be persisted for observability.
|
||||
- **AND** it SHALL continue to include compatibility `facts_checked`
|
||||
- **AND** existing fields such as `verdict`, `groundedness_score`, `rationale`, `executor_output_parse_status`, `tool_trace_summary`, and `gatekeeper_result` SHALL be preserved
|
||||
|
||||
#### Scenario: composer output written to self_evaluation
|
||||
- **WHEN** final answer composition completes
|
||||
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `composer_output`
|
||||
- **AND** `composer_output` SHALL indicate whether parsed Composer output or fallback rendering was used
|
||||
- **AND** existing verifier fields such as `claim_checks`, `facts_checked`, `gatekeeper_result`, and `tool_trace_summary` SHALL be preserved
|
||||
|
||||
#### Scenario: verdict written to self_evaluation
|
||||
- **WHEN** the Verifier produces a verdict
|
||||
- **THEN** the ChatService SHALL write the verdict data under `diagnosis_session.self_evaluation.verifier_evaluation`
|
||||
@@ -403,3 +399,25 @@ The Verifier SHALL classify each structured claim using a fixed derivability cla
|
||||
- **WHEN** Verifier emits `claim_checks`
|
||||
- **THEN** each claim check SHALL include `claim_id`, `verification`, `detail`, and `evidence_refs`
|
||||
- **AND** every evidence ref SHALL preserve available `trace_ref`, `tool_name`, and `source_invocation_ids`
|
||||
|
||||
### Requirement: User-facing verifier outputs SHALL follow Composer-safe protocols
|
||||
The system SHALL use Composer or fixed safe templates for PASS, LOW_CONFID, and REJECT user-facing responses.
|
||||
|
||||
#### Scenario: PASS uses Composer-safe output
|
||||
- **WHEN** the final verdict is `PASS`
|
||||
- **THEN** the user-facing response SHALL be generated from Verifier-allowed material through Composer or a safe fixed template
|
||||
- **AND** it SHALL NOT use Executor `user_facing_answer`
|
||||
- **AND** it SHALL NOT expose raw Executor JSON
|
||||
|
||||
#### Scenario: LOW_CONFID uses Composer-safe uncertainty output
|
||||
- **WHEN** the final verdict is `LOW_CONFID`
|
||||
- **THEN** the user-facing response SHALL include only confirmed facts, possible directions, evidence gaps, and next-step suggestions derived from Verifier-allowed material
|
||||
- **AND** optional evidence gaps, if present, SHALL come only from verifier-identified gaps
|
||||
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
|
||||
|
||||
#### Scenario: REJECT uses degraded template
|
||||
- **WHEN** the final verdict is `REJECT`
|
||||
- **THEN** the user-facing response SHALL use a degraded template or Composer-safe degraded output
|
||||
- **AND** it SHALL include only confirmed facts, evidence gaps, and next-step suggestions
|
||||
- **AND** it SHALL NOT include unverified raw answer content
|
||||
- **AND** it SHALL NOT include a root-cause conclusion
|
||||
|
||||
Reference in New Issue
Block a user