feat(agent): add composer final answer
This commit is contained in:
@@ -0,0 +1 @@
|
||||
ready
|
||||
+31
@@ -41,3 +41,34 @@
|
||||
- No unresolved user-interview question identified.
|
||||
- No database migration required.
|
||||
- No OpenSpec/devflow conflict found.
|
||||
|
||||
## Apply Notes
|
||||
|
||||
- Implemented `chat_composer` as a post-Verifier expression Agent in `ChatService`.
|
||||
- `ChatService` now builds Composer input from `VerifierDecision` plus parsed Executor structured output, not from raw Executor answer text.
|
||||
- PASS, LOW_CONFID, and REJECT final-answer paths now use Composer output or a deterministic safe fallback.
|
||||
- Verifier-missing or Verifier-malformed paths use fixed fallback directly because there is no trustworthy Verifier decision for Composer.
|
||||
- Composer audit is persisted under `verifier_evaluation.composer_output` without adding a database table.
|
||||
|
||||
## Test Drift And Fixes
|
||||
|
||||
- Initial test compilation failed because `ChatServiceSequentialAgentTest.java` had a UTF-8 BOM at the file start. Removed the BOM.
|
||||
- Existing sequential-flow tests still expected the old three-Agent call sequence and temporary V2 renderer behavior. Updated them to include `chat_composer` when a valid Verifier decision exists.
|
||||
- Tests that exercise fixed fallback now force malformed Composer output so they verify no raw JSON or Executor final answer leakage.
|
||||
- LOW_CONFID tests were updated to allow indirect support to appear as a possible direction rather than requiring it to disappear from all final-answer text.
|
||||
|
||||
## Verification
|
||||
|
||||
Passed:
|
||||
|
||||
```powershell
|
||||
mvn "-Dtest=ChatServiceSequentialAgentTest" test
|
||||
mvn "-Dtest=ExecutorGatekeeperServiceTest,VerifierInputHookTest,ChatServiceSequentialAgentTest" test
|
||||
cmd /c openspec validate executor-composer-final-answer
|
||||
cmd /c openspec validate --specs
|
||||
```
|
||||
|
||||
Known existing warnings:
|
||||
|
||||
- Maven reports duplicate `spring-boot-starter-test` dependency in `pom.xml`.
|
||||
- Existing Lombok `@Builder` default warnings remain.
|
||||
+31
-22
@@ -32,28 +32,6 @@ The system SHALL use ChatService for explicit single-round `Planner -> Executor
|
||||
- **AND** it SHALL NOT pass through the raw Executor answer
|
||||
- **AND** it SHALL NOT include a root-cause conclusion
|
||||
|
||||
### Requirement: User-facing verifier outputs SHALL follow Composer-safe protocols
|
||||
The system SHALL use Composer or fixed safe templates for PASS, LOW_CONFID, and REJECT user-facing responses.
|
||||
|
||||
#### Scenario: PASS uses Composer-safe output
|
||||
- **WHEN** the final verdict is `PASS`
|
||||
- **THEN** the user-facing response SHALL be generated from Verifier-allowed material through Composer or a safe fixed template
|
||||
- **AND** it SHALL NOT use Executor `user_facing_answer`
|
||||
- **AND** it SHALL NOT expose raw Executor JSON
|
||||
|
||||
#### Scenario: LOW_CONFID uses Composer-safe uncertainty output
|
||||
- **WHEN** the final verdict is `LOW_CONFID`
|
||||
- **THEN** the user-facing response SHALL include only confirmed facts, possible directions, evidence gaps, and next-step suggestions derived from Verifier-allowed material
|
||||
- **AND** optional evidence gaps, if present, SHALL come only from verifier-identified gaps
|
||||
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
|
||||
|
||||
#### Scenario: REJECT uses degraded template
|
||||
- **WHEN** the final verdict is `REJECT`
|
||||
- **THEN** the user-facing response SHALL use a degraded template or Composer-safe degraded output
|
||||
- **AND** it SHALL include only confirmed facts, evidence gaps, and next-step suggestions
|
||||
- **AND** it SHALL NOT include unverified raw answer content
|
||||
- **AND** it SHALL NOT include a root-cause conclusion
|
||||
|
||||
### Requirement: Verifier SHALL be observable
|
||||
The Verifier's verdict and downstream final-answer composition SHALL be persisted for observability.
|
||||
|
||||
@@ -78,3 +56,34 @@ The Verifier's verdict and downstream final-answer composition SHALL be persiste
|
||||
- **WHEN** the Verifier evaluation is persisted
|
||||
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `gatekeeper_result`
|
||||
- **AND** existing verifier fields such as `verdict`, `facts_checked`, `executor_output_parse_status`, and `tool_trace_summary` SHALL be preserved
|
||||
|
||||
## ADDED Requirements
|
||||
|
||||
### Requirement: User-facing verifier outputs SHALL follow Composer-safe protocols
|
||||
The system SHALL use Composer or fixed safe templates for PASS, LOW_CONFID, and REJECT user-facing responses.
|
||||
|
||||
#### Scenario: PASS uses Composer-safe output
|
||||
- **WHEN** the final verdict is `PASS`
|
||||
- **THEN** the user-facing response SHALL be generated from Verifier-allowed material through Composer or a safe fixed template
|
||||
- **AND** it SHALL NOT use Executor `user_facing_answer`
|
||||
- **AND** it SHALL NOT expose raw Executor JSON
|
||||
|
||||
#### Scenario: LOW_CONFID uses Composer-safe uncertainty output
|
||||
- **WHEN** the final verdict is `LOW_CONFID`
|
||||
- **THEN** the user-facing response SHALL include only confirmed facts, possible directions, evidence gaps, and next-step suggestions derived from Verifier-allowed material
|
||||
- **AND** optional evidence gaps, if present, SHALL come only from verifier-identified gaps
|
||||
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
|
||||
|
||||
#### Scenario: REJECT uses degraded template
|
||||
- **WHEN** the final verdict is `REJECT`
|
||||
- **THEN** the user-facing response SHALL use a degraded template or Composer-safe degraded output
|
||||
- **AND** it SHALL include only confirmed facts, evidence gaps, and next-step suggestions
|
||||
- **AND** it SHALL NOT include unverified raw answer content
|
||||
- **AND** it SHALL NOT include a root-cause conclusion
|
||||
|
||||
## REMOVED Requirements
|
||||
|
||||
### Requirement: User-facing verifier outputs SHALL follow fixed templates
|
||||
**Reason**: final user-facing answers are no longer owned by legacy Verifier templates. They must be generated from Verifier-allowed material through Composer or deterministic safe fallback.
|
||||
|
||||
**Migration**: use "User-facing verifier outputs SHALL follow Composer-safe protocols" and the new `chat-composer-agent` capability.
|
||||
@@ -0,0 +1,44 @@
|
||||
## 1. Composer Prompt Contract
|
||||
|
||||
- [x] 1.1 Add `src/main/resources/prompts/chat-composer-prompt.md`.
|
||||
- [x] 1.2 Define Composer input fields: `original_query`, `verdict`, `allowed_claims`, `allowed_hypotheses`, `missing_info`, `recommended_actions`, and `rationale`.
|
||||
- [x] 1.3 Define strict JSON output fields: `answer_summary`, `recommended_actions`, and `user_facing_answer`.
|
||||
- [x] 1.4 State verdict-specific wording rules for PASS, LOW_CONFID, and REJECT.
|
||||
- [x] 1.5 State that Composer must not add facts, call tools, output Markdown, or use raw Executor/tool output.
|
||||
|
||||
## 2. Composer Invocation And Input Filtering
|
||||
|
||||
- [x] 2.1 Load the Composer prompt in `ChatService`.
|
||||
- [x] 2.2 Add a `chat_composer` Agent or equivalent Composer model call after final Verifier decision.
|
||||
- [x] 2.3 Build Composer input from Verifier decision and Executor structured output.
|
||||
- [x] 2.4 Filter `allowed_claims` from `claim_checks` using `direct_observation` and bounded `reasonable_inference`.
|
||||
- [x] 2.5 Exclude `unsupported`, `external_unknown`, and `contradicted` claims from confirmed output.
|
||||
- [x] 2.6 Downgrade `overstated` claims to `allowed_hypotheses` or `missing_info`.
|
||||
- [x] 2.7 Ensure REJECT Composer input has `allowed_hypotheses=[]`.
|
||||
- [x] 2.8 Ensure Composer input contains no raw tool output, full unscreened Executor output, or Executor `user_facing_answer`.
|
||||
|
||||
## 3. Composer Output Parsing, Fallback, And Audit
|
||||
|
||||
- [x] 3.1 Parse Composer strict JSON output.
|
||||
- [x] 3.2 Add safe fallback rendering for malformed Composer output.
|
||||
- [x] 3.3 Ensure fallback rendering never exposes raw Composer JSON, raw Executor JSON, or Executor `user_facing_answer`.
|
||||
- [x] 3.4 Persist `composer_output` under `diagnosis_session.self_evaluation.verifier_evaluation`.
|
||||
- [x] 3.5 Preserve existing verifier audit fields when writing Composer output.
|
||||
|
||||
## 4. Final Answer Routing
|
||||
|
||||
- [x] 4.1 Replace PASS temporary V2 renderer usage with Composer or safe template rendering.
|
||||
- [x] 4.2 Ensure LOW_CONFID final answer uses Composer-safe filtered material.
|
||||
- [x] 4.3 Ensure REJECT final answer does not include root-cause conclusions and does not pass through Executor raw answer.
|
||||
- [x] 4.4 Ensure Executor `user_facing_answer` is not read by any final answer path.
|
||||
|
||||
## 5. Tests And Verification
|
||||
|
||||
- [x] 5.1 Add or update tests proving Composer input excludes raw tool output and full Executor output.
|
||||
- [x] 5.2 Add or update tests proving unsupported claims do not appear as confirmed final-answer content.
|
||||
- [x] 5.3 Add or update tests for PASS root-cause wording only when an allowed root-cause claim exists.
|
||||
- [x] 5.4 Add or update tests for LOW_CONFID separation of confirmed information, possible directions, and evidence gaps.
|
||||
- [x] 5.5 Add or update tests for REJECT with `allowed_hypotheses=[]` and no root-cause conclusion.
|
||||
- [x] 5.6 Add or update tests for malformed Composer fallback with no raw JSON leakage.
|
||||
- [x] 5.7 Run targeted Maven tests.
|
||||
- [x] 5.8 Validate this OpenSpec change and all specs.
|
||||
@@ -1,44 +0,0 @@
|
||||
## 1. Composer Prompt Contract
|
||||
|
||||
- [ ] 1.1 Add `src/main/resources/prompts/chat-composer-prompt.md`.
|
||||
- [ ] 1.2 Define Composer input fields: `original_query`, `verdict`, `allowed_claims`, `allowed_hypotheses`, `missing_info`, `recommended_actions`, and `rationale`.
|
||||
- [ ] 1.3 Define strict JSON output fields: `answer_summary`, `recommended_actions`, and `user_facing_answer`.
|
||||
- [ ] 1.4 State verdict-specific wording rules for PASS, LOW_CONFID, and REJECT.
|
||||
- [ ] 1.5 State that Composer must not add facts, call tools, output Markdown, or use raw Executor/tool output.
|
||||
|
||||
## 2. Composer Invocation And Input Filtering
|
||||
|
||||
- [ ] 2.1 Load the Composer prompt in `ChatService`.
|
||||
- [ ] 2.2 Add a `chat_composer` Agent or equivalent Composer model call after final Verifier decision.
|
||||
- [ ] 2.3 Build Composer input from Verifier decision and Executor structured output.
|
||||
- [ ] 2.4 Filter `allowed_claims` from `claim_checks` using `direct_observation` and bounded `reasonable_inference`.
|
||||
- [ ] 2.5 Exclude `unsupported`, `external_unknown`, and `contradicted` claims from confirmed output.
|
||||
- [ ] 2.6 Downgrade `overstated` claims to `allowed_hypotheses` or `missing_info`.
|
||||
- [ ] 2.7 Ensure REJECT Composer input has `allowed_hypotheses=[]`.
|
||||
- [ ] 2.8 Ensure Composer input contains no raw tool output, full unscreened Executor output, or Executor `user_facing_answer`.
|
||||
|
||||
## 3. Composer Output Parsing, Fallback, And Audit
|
||||
|
||||
- [ ] 3.1 Parse Composer strict JSON output.
|
||||
- [ ] 3.2 Add safe fallback rendering for malformed Composer output.
|
||||
- [ ] 3.3 Ensure fallback rendering never exposes raw Composer JSON, raw Executor JSON, or Executor `user_facing_answer`.
|
||||
- [ ] 3.4 Persist `composer_output` under `diagnosis_session.self_evaluation.verifier_evaluation`.
|
||||
- [ ] 3.5 Preserve existing verifier audit fields when writing Composer output.
|
||||
|
||||
## 4. Final Answer Routing
|
||||
|
||||
- [ ] 4.1 Replace PASS temporary V2 renderer usage with Composer or safe template rendering.
|
||||
- [ ] 4.2 Ensure LOW_CONFID final answer uses Composer-safe filtered material.
|
||||
- [ ] 4.3 Ensure REJECT final answer does not include root-cause conclusions and does not pass through Executor raw answer.
|
||||
- [ ] 4.4 Ensure Executor `user_facing_answer` is not read by any final answer path.
|
||||
|
||||
## 5. Tests And Verification
|
||||
|
||||
- [ ] 5.1 Add or update tests proving Composer input excludes raw tool output and full Executor output.
|
||||
- [ ] 5.2 Add or update tests proving unsupported claims do not appear as confirmed final-answer content.
|
||||
- [ ] 5.3 Add or update tests for PASS root-cause wording only when an allowed root-cause claim exists.
|
||||
- [ ] 5.4 Add or update tests for LOW_CONFID separation of confirmed information, possible directions, and evidence gaps.
|
||||
- [ ] 5.5 Add or update tests for REJECT with `allowed_hypotheses=[]` and no root-cause conclusion.
|
||||
- [ ] 5.6 Add or update tests for malformed Composer fallback with no raw JSON leakage.
|
||||
- [ ] 5.7 Run targeted Maven tests.
|
||||
- [ ] 5.8 Validate this OpenSpec change and all specs.
|
||||
@@ -0,0 +1,83 @@
|
||||
# chat-composer-agent Specification
|
||||
|
||||
## Purpose
|
||||
TBD - created by archiving change executor-composer-final-answer. Update Purpose after archive.
|
||||
## Requirements
|
||||
### Requirement: Composer SHALL generate final user-facing Chat answers
|
||||
The system SHALL invoke a Composer expression layer after Verifier to generate the final user-facing Chat answer from Verifier-allowed material.
|
||||
|
||||
#### Scenario: Composer receives only filtered material
|
||||
- **WHEN** ChatService invokes Composer
|
||||
- **THEN** the Composer input SHALL contain `original_query`, `verdict`, `allowed_claims`, `allowed_hypotheses`, `missing_info`, `recommended_actions`, and `rationale`
|
||||
- **AND** the Composer input SHALL NOT contain raw tool output
|
||||
- **AND** the Composer input SHALL NOT contain the full unscreened Executor output
|
||||
- **AND** the Composer input SHALL NOT contain Executor `user_facing_answer`
|
||||
|
||||
#### Scenario: Composer outputs strict JSON
|
||||
- **WHEN** Composer completes
|
||||
- **THEN** it SHALL output exactly one JSON object
|
||||
- **AND** the JSON object SHALL include `answer_summary`, `recommended_actions`, and `user_facing_answer`
|
||||
- **AND** it SHALL NOT output Markdown, code fences, or explanatory text outside the JSON object
|
||||
|
||||
#### Scenario: Composer does not introduce new facts
|
||||
- **WHEN** Composer produces `answer_summary`, `recommended_actions`, or `user_facing_answer`
|
||||
- **THEN** every service name, entity, timestamp, error code, metric value, root cause, and recommendation reason SHALL be derived from the Composer input
|
||||
- **AND** Composer SHALL NOT add facts from model knowledge, raw tool history, or Executor raw text
|
||||
|
||||
### Requirement: Composer input SHALL honor Verifier claim checks
|
||||
ChatService SHALL construct Composer input by filtering Executor structured output through Verifier `claim_checks`.
|
||||
|
||||
#### Scenario: Passing claims become allowed claims
|
||||
- **WHEN** a claim check verification is `direct_observation`
|
||||
- **THEN** ChatService SHALL include the matching Executor claim in `allowed_claims`
|
||||
|
||||
#### Scenario: Reasonable inferences remain bounded
|
||||
- **WHEN** a claim check verification is `reasonable_inference`
|
||||
- **THEN** ChatService MAY include the matching Executor claim in `allowed_claims`
|
||||
- **AND** the final answer SHALL NOT describe it as the sole confirmed root cause unless the allowed claim itself is a root-cause claim and the final verdict is `PASS`
|
||||
|
||||
#### Scenario: Overstated claims are not confirmed findings
|
||||
- **WHEN** a claim check verification is `overstated`
|
||||
- **THEN** ChatService SHALL NOT include the matching Executor claim as a confirmed item in `allowed_claims`
|
||||
- **AND** ChatService MAY include it as `allowed_hypotheses` or represent it in `missing_info`
|
||||
|
||||
#### Scenario: Unsupported or external claims are withheld
|
||||
- **WHEN** a claim check verification is `unsupported`, `external_unknown`, or `contradicted`
|
||||
- **THEN** ChatService SHALL NOT include the matching Executor claim in `allowed_claims`
|
||||
- **AND** the final user-facing answer SHALL NOT present that claim as confirmed
|
||||
|
||||
### Requirement: Composer SHALL respect verdict-specific wording
|
||||
Composer SHALL phrase final answers according to the effective Verifier verdict.
|
||||
|
||||
#### Scenario: PASS answer uses confirmed material
|
||||
- **WHEN** the effective verdict is `PASS`
|
||||
- **THEN** the final answer MAY state confirmed findings from `allowed_claims`
|
||||
- **AND** it SHALL only state root cause confirmed when an allowed root-cause claim is present
|
||||
|
||||
#### Scenario: LOW_CONFID answer separates findings and gaps
|
||||
- **WHEN** the effective verdict is `LOW_CONFID`
|
||||
- **THEN** the final answer SHALL distinguish confirmed information from possible directions
|
||||
- **AND** it SHALL mention evidence gaps from `missing_info`
|
||||
- **AND** it SHALL NOT turn `allowed_hypotheses` into confirmed findings
|
||||
|
||||
#### Scenario: REJECT answer avoids root-cause conclusions
|
||||
- **WHEN** the effective verdict is `REJECT`
|
||||
- **THEN** Composer input SHALL have `allowed_hypotheses=[]`
|
||||
- **AND** the final answer SHALL state that current evidence cannot support a reliable conclusion
|
||||
- **AND** the final answer SHALL NOT include a root-cause conclusion
|
||||
|
||||
### Requirement: Composer failures SHALL degrade safely
|
||||
The system SHALL tolerate malformed Composer output without leaking raw JSON or unverified Executor material.
|
||||
|
||||
#### Scenario: malformed Composer output falls back safely
|
||||
- **WHEN** Composer returns malformed JSON or omits required fields
|
||||
- **THEN** ChatService SHALL produce a final answer using a fixed safe fallback template based only on filtered material
|
||||
- **AND** the final answer SHALL NOT expose raw Composer output
|
||||
- **AND** the final answer SHALL NOT expose raw Executor output
|
||||
- **AND** the final answer SHALL NOT use Executor `user_facing_answer`
|
||||
|
||||
#### Scenario: Composer audit is persisted
|
||||
- **WHEN** ChatService persists verifier evaluation
|
||||
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation.composer_output` SHALL record whether Composer output was valid or fallback was used
|
||||
- **AND** the audit SHALL include the parsed Composer fields when valid
|
||||
- **AND** the audit SHALL remain compact and SHALL NOT store full raw tool output
|
||||
@@ -89,48 +89,38 @@ The groundedness score SHALL be computed from critical fact classifications inst
|
||||
- **AND** the result SHALL be clamped into `[0.0, 1.0]`
|
||||
|
||||
### Requirement: ChatService SHALL route based on Verifier verdict
|
||||
The system SHALL use ChatService for explicit single-round `Planner → Executor → Verifier` orchestration and SHALL use ChatService to control whether an additional round is allowed.
|
||||
The system SHALL use ChatService for explicit single-round `Planner -> Executor -> Verifier -> Composer` orchestration and SHALL use ChatService to control whether an additional round is allowed.
|
||||
|
||||
#### Scenario: PASS → direct output
|
||||
#### Scenario: PASS -> Composer output
|
||||
- **WHEN** Verifier outputs verdict="PASS"
|
||||
- **THEN** the system SHALL output the Executor's answer directly
|
||||
- **THEN** ChatService SHALL filter Verifier-allowed material and invoke Composer or a safe fixed template
|
||||
- **AND** the final user-facing answer SHALL NOT pass through raw Executor output
|
||||
- **AND** the final user-facing answer SHALL NOT read Executor `user_facing_answer`
|
||||
|
||||
#### Scenario: LOW_CONFID score≥0.5 → output with disclaimer
|
||||
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score ≥ 0.5
|
||||
- **THEN** the system SHALL output the Executor's answer prefixed with a fixed confidence disclaimer
|
||||
#### Scenario: LOW_CONFID score>=0.5 -> Composer output with uncertainty
|
||||
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score >= 0.5
|
||||
- **THEN** ChatService SHALL filter Verifier-allowed material and invoke Composer or a safe fixed template
|
||||
- **AND** the final user-facing answer SHALL distinguish confirmed information, possible directions, and evidence gaps
|
||||
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
|
||||
|
||||
#### Scenario: LOW_CONFID score<0.5 → trigger one additional round
|
||||
#### Scenario: LOW_CONFID score<0.5 -> trigger one additional round
|
||||
- **WHEN** Verifier outputs verdict="LOW_CONFID" with groundedness_score < 0.5 and this is the first callback
|
||||
- **THEN** the ChatService SHALL invoke one additional `Planner → Executor → Verifier` round to supplement evidence
|
||||
- **AND** after the second Verifier run, verdict="LOW_CONFID" SHALL be output with a confidence disclaimer
|
||||
- **AND** after the second Verifier run, verdict="REJECT" SHALL still produce a degraded output
|
||||
- **THEN** the ChatService SHALL invoke one additional `Planner -> Executor -> Verifier` round to supplement evidence
|
||||
- **AND** after the second Verifier run, verdict="LOW_CONFID" SHALL be routed to Composer or a safe fixed template
|
||||
- **AND** after the second Verifier run, verdict="REJECT" SHALL still produce a degraded Composer-safe output
|
||||
|
||||
#### Scenario: REJECT does not enter retry round
|
||||
- **WHEN** Verifier outputs verdict="REJECT"
|
||||
- **THEN** the system SHALL NOT start a retry round for evidence补充
|
||||
- **AND** it SHALL produce a degraded output directly
|
||||
- **THEN** the system SHALL NOT start a retry round for evidence supplementation
|
||||
- **AND** it SHALL produce a degraded output directly through Composer-safe rendering
|
||||
|
||||
#### Scenario: REJECT → degraded output
|
||||
#### Scenario: REJECT -> degraded output
|
||||
- **WHEN** Verifier outputs verdict="REJECT"
|
||||
- **THEN** the system SHALL output a degraded result indicating the answer cannot be reliably generated
|
||||
- **AND** it SHALL NOT pass through the raw Executor answer
|
||||
|
||||
### Requirement: User-facing verifier outputs SHALL follow fixed templates
|
||||
The system SHALL use fixed output protocols for LOW_CONFID and REJECT user-facing responses.
|
||||
|
||||
#### Scenario: LOW_CONFID uses disclaimer template
|
||||
- **WHEN** the final verdict is `LOW_CONFID`
|
||||
- **THEN** the user-facing response SHALL prepend a fixed disclaimer before the Executor answer
|
||||
- **AND** optional evidence gaps, if present, SHALL come only from verifier-identified critical gaps
|
||||
|
||||
#### Scenario: REJECT uses degraded template
|
||||
- **WHEN** the final verdict is `REJECT`
|
||||
- **THEN** the user-facing response SHALL use a degraded template
|
||||
- **AND** it SHALL include only confirmed facts, evidence gaps, and next-step suggestions
|
||||
- **AND** it SHALL NOT include unverified raw answer content
|
||||
|
||||
- **AND** it SHALL NOT include a root-cause conclusion
|
||||
### Requirement: Verifier SHALL be observable
|
||||
The Verifier's verdict SHALL be persisted for observability.
|
||||
The Verifier's verdict and downstream final-answer composition SHALL be persisted for observability.
|
||||
|
||||
#### Scenario: claim checks written to self_evaluation
|
||||
- **WHEN** the Verifier evaluation is persisted
|
||||
@@ -138,6 +128,12 @@ The Verifier's verdict SHALL be persisted for observability.
|
||||
- **AND** it SHALL continue to include compatibility `facts_checked`
|
||||
- **AND** existing fields such as `verdict`, `groundedness_score`, `rationale`, `executor_output_parse_status`, `tool_trace_summary`, and `gatekeeper_result` SHALL be preserved
|
||||
|
||||
#### Scenario: composer output written to self_evaluation
|
||||
- **WHEN** final answer composition completes
|
||||
- **THEN** `diagnosis_session.self_evaluation.verifier_evaluation` SHALL include `composer_output`
|
||||
- **AND** `composer_output` SHALL indicate whether parsed Composer output or fallback rendering was used
|
||||
- **AND** existing verifier fields such as `claim_checks`, `facts_checked`, `gatekeeper_result`, and `tool_trace_summary` SHALL be preserved
|
||||
|
||||
#### Scenario: verdict written to self_evaluation
|
||||
- **WHEN** the Verifier produces a verdict
|
||||
- **THEN** the ChatService SHALL write the verdict data under `diagnosis_session.self_evaluation.verifier_evaluation`
|
||||
@@ -403,3 +399,25 @@ The Verifier SHALL classify each structured claim using a fixed derivability cla
|
||||
- **WHEN** Verifier emits `claim_checks`
|
||||
- **THEN** each claim check SHALL include `claim_id`, `verification`, `detail`, and `evidence_refs`
|
||||
- **AND** every evidence ref SHALL preserve available `trace_ref`, `tool_name`, and `source_invocation_ids`
|
||||
|
||||
### Requirement: User-facing verifier outputs SHALL follow Composer-safe protocols
|
||||
The system SHALL use Composer or fixed safe templates for PASS, LOW_CONFID, and REJECT user-facing responses.
|
||||
|
||||
#### Scenario: PASS uses Composer-safe output
|
||||
- **WHEN** the final verdict is `PASS`
|
||||
- **THEN** the user-facing response SHALL be generated from Verifier-allowed material through Composer or a safe fixed template
|
||||
- **AND** it SHALL NOT use Executor `user_facing_answer`
|
||||
- **AND** it SHALL NOT expose raw Executor JSON
|
||||
|
||||
#### Scenario: LOW_CONFID uses Composer-safe uncertainty output
|
||||
- **WHEN** the final verdict is `LOW_CONFID`
|
||||
- **THEN** the user-facing response SHALL include only confirmed facts, possible directions, evidence gaps, and next-step suggestions derived from Verifier-allowed material
|
||||
- **AND** optional evidence gaps, if present, SHALL come only from verifier-identified gaps
|
||||
- **AND** unsupported raw Executor claims SHALL NOT be presented as confirmed conclusions
|
||||
|
||||
#### Scenario: REJECT uses degraded template
|
||||
- **WHEN** the final verdict is `REJECT`
|
||||
- **THEN** the user-facing response SHALL use a degraded template or Composer-safe degraded output
|
||||
- **AND** it SHALL include only confirmed facts, evidence gaps, and next-step suggestions
|
||||
- **AND** it SHALL NOT include unverified raw answer content
|
||||
- **AND** it SHALL NOT include a root-cause conclusion
|
||||
|
||||
Reference in New Issue
Block a user