Merge branch 'emdash/shy-items-fry-f4zze' into refactor/mvp1.0
# Conflicts: # mvp/issues/README.md
This commit is contained in:
@@ -4,6 +4,11 @@
|
||||
|
||||
| 日期 | slug | 领域 | 关键词 | 状态 |
|
||||
|---|---|---|---|---|
|
||||
| 2026-07-05 | mvp-demo-interview-runbook | MVP Demo/Interview | Plan C, payment timeout, runbook, trace checklist, demo script | openspec/changes/archive/2026-07-05-mvp-demo-interview-runbook | archived |
|
||||
| 2026-07-05 | diagnosis-eval-baseline-diff | Agent 评测/回归 Diff | baseline diff, regression detection, evidence coverage, cost signal, markdown report | openspec/changes/archive/2026-07-05-diagnosis-eval-baseline-diff | archived |
|
||||
| 2026-07-04 | expand-diagnosis-eval-fixtures | Agent 评测/回归 Baseline | fixture coverage, baseline report, redis timeout, slow response, jvm memory risk | openspec/changes/archive/2026-07-05-expand-diagnosis-eval-fixtures | archived |
|
||||
| 2026-07-04 | diagnosis-eval-harness | Agent 评测/回归 Harness | fixed cases, trace validation, evidence coverage, verdict distribution, markdown report | openspec/changes/archive/2026-07-04-diagnosis-eval-harness | archived |
|
||||
| 2026-07-04 | evidence-trace-hardening | 证据链/降级契约/离线验证 | ToolInvocationRecorder, ToolTraceSummaryService, lookup_knowledge, query_logs, query_metrics, LOW_CONFID, REJECT | openspec/changes/archive/2026-07-04-evidence-trace-hardening | archived |
|
||||
| 2026-07-03 | mvp-demo-trace-acceptance | MVP Demo/trace/acceptance | mvp-demo, trace API, diagnosis_session, agent_step, tool_invocation, feedback | openspec/changes/archive/2026-07-03-mvp-demo-trace-acceptance | archived |
|
||||
| 2026-07-04 | aiops-traceable-diagnosis-entry | AIOps/trace/alert diagnosis | ai_ops, SSE, alert input, sessionId, diagnosis_session, trace API | openspec/changes/archive/2026-07-04-aiops-traceable-diagnosis-entry | archived |
|
||||
| 2026-07-04 | aiops-alert-scope-control | AIOps/scope/prompt control | payload mode, auto-discovery mode, queryPrometheusAlerts, HighCPUUsage | openspec/changes/archive/2026-07-04-aiops-alert-scope-control | archived |
|
||||
|
||||
@@ -0,0 +1,36 @@
|
||||
# Acceptance: diagnosis-eval-harness
|
||||
|
||||
## Classification
|
||||
|
||||
standard-light
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Issue and OpenSpec setup | Done | `ISS-006` and initial OpenSpec artifacts were created. |
|
||||
| Implementation | Done | Added fixed cases, fixture-mode trace evaluation, aggregate metrics, and JSON / Markdown report writer. |
|
||||
| Verification | Done | Targeted evaluator tests, compile verification, and OpenSpec validation passed. |
|
||||
|
||||
## Current State
|
||||
|
||||
- First implementation uses fixture-mode evaluation.
|
||||
- Live trace API polling remains a follow-up option.
|
||||
|
||||
## Verification
|
||||
|
||||
### Script Verification
|
||||
|
||||
- Command: `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
|
||||
- Result: passed
|
||||
- Notes: Covers fixed case loading, fixture evaluation, missing fixture reporting, reject degraded-output validation, and report writing.
|
||||
|
||||
### Static Verification
|
||||
|
||||
- Command: `mvn -q -DskipTests compile`
|
||||
- Result: passed
|
||||
|
||||
### OpenSpec Verification
|
||||
|
||||
- Command: `openspec validate diagnosis-eval-harness --strict`
|
||||
- Result: passed
|
||||
@@ -0,0 +1,31 @@
|
||||
# Brief: diagnosis-eval-harness
|
||||
|
||||
## Background
|
||||
|
||||
The MVP has a runnable demo and hardened evidence trace semantics, but it still lacks a fixed regression baseline for Agent diagnosis quality. P1-B creates a small evaluation harness that can validate diagnosis traces against fixed cases and produce repeatable reports.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Define fixed diagnosis cases for the MVP demo domain.
|
||||
2. Validate trace evidence, verifier verdicts, answer keywords, and degraded-output behavior.
|
||||
3. Produce JSON and Markdown reports for interview and regression use.
|
||||
4. Keep the first version offline by supporting trace fixtures.
|
||||
|
||||
## Scope
|
||||
|
||||
- Evaluation case definitions
|
||||
- Trace fixture shape
|
||||
- Rule-based evaluator
|
||||
- JSON / Markdown report output
|
||||
- Focused offline tests and docs
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No LLM-as-judge
|
||||
- No live end-to-end runtime requirement
|
||||
- No production API
|
||||
- No chat or verifier runtime change
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/diagnosis-eval-harness/`
|
||||
@@ -0,0 +1,28 @@
|
||||
# Diagnosis Eval Harness Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: build P1-B fixed case evaluation after evidence trace hardening.
|
||||
- Slug: `diagnosis-eval-harness`
|
||||
- Devflow scale: standard-light
|
||||
|
||||
## Context
|
||||
|
||||
- P1-A `evidence-trace-hardening` created stable evidence semantics for supported, no-evidence, deduped, and failed tool calls.
|
||||
- The MVP demo trace API already provides an aggregate trace shape suitable for evaluation.
|
||||
- The first evaluator should avoid depending on external infrastructure so it can run in regular development.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Start with rule-based trace validation instead of LLM-as-judge.
|
||||
- Reason: The first regression signal should be deterministic and tied to trace contracts.
|
||||
|
||||
- Decision: Support offline fixture traces first.
|
||||
- Reason: This makes the harness usable without MySQL, Redis, Milvus, or a real LLM.
|
||||
|
||||
- Decision: Output both JSON and Markdown.
|
||||
- Reason: JSON supports automation; Markdown is easier to discuss in interviews.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Whether live trace API polling belongs in this change or a follow-up after fixture mode lands.
|
||||
@@ -0,0 +1,10 @@
|
||||
# Diagnosis Eval Harness Evidence
|
||||
|
||||
## Evidence
|
||||
|
||||
| Source | Evidence | Conclusion | Reported |
|
||||
|---|---|---|---|
|
||||
| `openspec/specs/evidence-trace-hardening/spec.md` | Defines stable evidence states and summary behavior | Evaluation can rely on trace semantics rather than ad hoc log parsing | Yes |
|
||||
| `mvp/demo/README.md` | Documents an end-to-end demo flow with chat, trace, and feedback | Existing demo flow provides the runtime story, but not a reusable evaluation baseline | Yes |
|
||||
| `DiagnosisTraceService` | Aggregates session, steps, tools, and self-evaluation | Trace response shape can be reused as evaluation input | Yes |
|
||||
| `ToolTraceSummaryService` | Builds verifier-facing evidence summaries from persisted tool rows | Evaluator can check evidence coverage through persisted trace artifacts | Yes |
|
||||
@@ -0,0 +1,32 @@
|
||||
# Acceptance: evidence-trace-hardening
|
||||
|
||||
## Classification
|
||||
|
||||
standard-light
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Issue and OpenSpec setup | Done | `ISS-005` and the initial OpenSpec artifacts were created. |
|
||||
| Implementation | Done | Recorder contract, lookup persistence path, evidence summary semantics, and degraded-path tests were implemented. |
|
||||
| Verification | Done | Targeted offline tests and compile verification passed. |
|
||||
|
||||
## Verification
|
||||
|
||||
### Script Verification
|
||||
|
||||
- Command: `mvn -q "-Dtest=ToolInvocationRecorderTest,ToolTraceSummaryServiceTest,ChatServiceSequentialAgentTest,LookupKnowledgeToolTest" test`
|
||||
- Result: passed
|
||||
- Notes: Covers recorder contract, summary semantics for success/failure/no-evidence, and `ChatService` fallback / degraded paths.
|
||||
|
||||
### Static Verification
|
||||
|
||||
- Command: `mvn -q -DskipTests compile`
|
||||
- Result: passed
|
||||
|
||||
## Open Questions
|
||||
|
||||
| Question | Current position |
|
||||
| --- | --- |
|
||||
| Should deduped retrievals be counted separately from generic no-hit events in future evaluation metrics? | Deferred to P1-B; this change preserves enough structure to decide later. |
|
||||
@@ -0,0 +1,32 @@
|
||||
# Brief: evidence-trace-hardening
|
||||
|
||||
## Background
|
||||
|
||||
The MVP already has persisted tool traces and a verifier, but the evidence contract is still only partially standardized. For interview-focused hardening, the project now needs a tighter contract for evidence persistence, no-evidence / failure semantics, and degraded-output behavior.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Standardize the persisted evidence-tool contract across `lookup_knowledge`, `query_logs`, and `query_metrics`.
|
||||
2. Make verifier-facing summaries distinguish failed calls, no-hit calls, deduped retrievals, and actual supporting evidence.
|
||||
3. Add offline tests for verifier fallback and degraded-output paths.
|
||||
|
||||
## Scope
|
||||
|
||||
- `ToolInvocationRecorder`
|
||||
- `LookupKnowledgeTool`
|
||||
- `QueryLogsTools`
|
||||
- `QueryMetricsTools`
|
||||
- `ToolTraceSummaryService`
|
||||
- `ChatService`
|
||||
- Focused offline tests
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No new API or schema
|
||||
- No evaluation harness yet
|
||||
- No trace UI
|
||||
- No security/config cleanup
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/evidence-trace-hardening/`
|
||||
@@ -0,0 +1,24 @@
|
||||
# Evidence Trace Hardening Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: harden the MVP evidence contract before building the P1-B evaluation harness.
|
||||
- Slug: `evidence-trace-hardening`
|
||||
- Devflow scale: standard-light
|
||||
|
||||
## Context
|
||||
|
||||
- `ISS-003` raised verifier traceability and failure-path concerns.
|
||||
- Current code inspection shows `QueryLogsTools` and `QueryMetricsTools` already use `ToolInvocationRecorder`, while `LookupKnowledgeTool` still persists rows through a local helper.
|
||||
- `ChatService` already contains fallback behavior for missing/invalid `verifier_output`, but coverage is narrow.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Treat this as a contract-hardening change, not a new feature change.
|
||||
- Reason: The project already has the necessary runtime pieces; the gap is semantic consistency and testability.
|
||||
|
||||
- Decision: Keep the scope before P1-B.
|
||||
- Reason: The evaluation harness will rely on stable evidence semantics, so this contract slice should land first.
|
||||
|
||||
- Decision: Preserve schema and API stability.
|
||||
- Reason: The interview value here is engineering rigor, not more surface area.
|
||||
@@ -0,0 +1,11 @@
|
||||
# Evidence Trace Hardening Evidence
|
||||
|
||||
## Evidence
|
||||
|
||||
| Source | Evidence | Conclusion | Reported |
|
||||
|---|---|---|---|
|
||||
| `ToolInvocationRecorder` | Provides a common persistence seam for evidence tools | Contract hardening should build on the existing recorder instead of introducing a new store path | Yes |
|
||||
| `LookupKnowledgeTool` | Still constructs `ToolInvocation` rows through a local helper | Retrieval-aware evidence persistence is not yet unified with the recorder contract | Yes |
|
||||
| `QueryLogsTools` / `QueryMetricsTools` | Already record evidence invocations through `recordEvidenceTool(...)` | Current gap is semantic alignment, not missing persistence | Yes |
|
||||
| `ToolTraceSummaryService` | Merges rows by tool and topic domain and infers evidence level heuristically | Summary rules need explicit handling for failure, no-hit, and dedup cases | Yes |
|
||||
| `ChatService` | Falls back to `LOW_CONFID` when verifier output is missing or invalid | These degraded paths exist and should now be covered by focused offline tests | Yes |
|
||||
@@ -0,0 +1,37 @@
|
||||
# Acceptance: expand-diagnosis-eval-fixtures
|
||||
|
||||
## Classification
|
||||
|
||||
standard-light
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Issue and OpenSpec setup | Done | Created slug-based issue and initial OpenSpec artifacts. |
|
||||
| Implementation | Done | Added remaining fixtures, full baseline reports, and documentation updates. |
|
||||
| Verification | Done | Evaluator tests, compile verification, and OpenSpec validation passed. |
|
||||
|
||||
## Current State
|
||||
|
||||
- Fixture coverage is complete for the five fixed diagnosis cases.
|
||||
- Baseline reports are saved under `mvp/eval/reports`.
|
||||
- No production runtime behavior has been changed.
|
||||
|
||||
## Verification
|
||||
|
||||
### Script Verification
|
||||
|
||||
- Command: `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`
|
||||
- Result: passed
|
||||
- Notes: Covers full fixture coverage, baseline report matching, reject degraded-output validation, and report writing.
|
||||
|
||||
### Static Verification
|
||||
|
||||
- Command: `mvn -q -DskipTests compile`
|
||||
- Result: passed
|
||||
|
||||
### OpenSpec Verification
|
||||
|
||||
- Command: `openspec validate expand-diagnosis-eval-fixtures --strict`
|
||||
- Result: passed
|
||||
@@ -0,0 +1,31 @@
|
||||
# Brief: expand-diagnosis-eval-fixtures
|
||||
|
||||
## Background
|
||||
|
||||
The diagnosis eval harness is implemented and archived, but the fixed baseline is incomplete because three of the five diagnosis cases still reference missing fixtures.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Add representative trace fixtures for all remaining fixed diagnosis cases.
|
||||
2. Save a reproducible baseline report in JSON and Markdown.
|
||||
3. Document how to regenerate and interpret the baseline.
|
||||
4. Keep evaluation offline and deterministic.
|
||||
|
||||
## Scope
|
||||
|
||||
- Redis timeout fixture
|
||||
- Slow response fixture
|
||||
- JVM memory risk fixture
|
||||
- Baseline reports under `mvp/eval/reports`
|
||||
- Focused tests for full fixture coverage and report generation
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No new diagnosis cases
|
||||
- No production Agent runtime changes
|
||||
- No LLM-as-judge
|
||||
- No live infrastructure requirement
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/expand-diagnosis-eval-fixtures/`
|
||||
@@ -0,0 +1,28 @@
|
||||
# Expand Diagnosis Eval Fixtures Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: complete the fixed diagnosis eval baseline after the harness is in place.
|
||||
- Slug: `expand-diagnosis-eval-fixtures`
|
||||
- Devflow scale: standard-light
|
||||
|
||||
## Context
|
||||
|
||||
- `diagnosis-eval-harness` created the evaluator, case file, fixture mode, and report writer.
|
||||
- The first baseline still has missing fixtures by design.
|
||||
- This follow-up turns that partial baseline into a full fixed-case baseline.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Keep this change data-focused.
|
||||
- Reason: the evaluator rules already landed; this change should not blur fixture expansion with harness behavior changes.
|
||||
|
||||
- Decision: Save baseline reports in the repository.
|
||||
- Reason: interview review and future diffs are easier when the expected baseline is visible.
|
||||
|
||||
- Decision: Use deterministic fixture traces instead of live trace generation.
|
||||
- Reason: this baseline should run without infrastructure or external model calls.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Whether a future change should add a CLI or Maven goal for report regeneration.
|
||||
@@ -0,0 +1,11 @@
|
||||
# Evidence: expand-diagnosis-eval-fixtures
|
||||
|
||||
## Evidence Log
|
||||
|
||||
- 2026-07-04: Created slug-based issue `expand-diagnosis-eval-fixtures.md`.
|
||||
- 2026-07-04: Created OpenSpec change `expand-diagnosis-eval-fixtures`.
|
||||
- 2026-07-04: Added Redis timeout, slow response, and JVM memory risk fixtures.
|
||||
- 2026-07-04: Added baseline JSON and Markdown reports under `mvp/eval/reports`.
|
||||
- 2026-07-04: Verification passed with `mvn -q "-Dtest=DiagnosisTraceEvaluatorTest" test`.
|
||||
- 2026-07-04: Verification passed with `mvn -q -DskipTests compile`.
|
||||
- 2026-07-04: Verification passed with `openspec validate expand-diagnosis-eval-fixtures --strict`.
|
||||
@@ -0,0 +1,37 @@
|
||||
# Acceptance: diagnosis-eval-baseline-diff
|
||||
|
||||
## Classification
|
||||
|
||||
standard-light
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Issue and OpenSpec setup | Done | Created slug-based issue and initial OpenSpec artifacts. |
|
||||
| Implementation | Done | Added diff model, comparator, writer, docs, sample outputs, and focused tests. |
|
||||
| Verification | Done | Diff/evaluator tests, compile verification, and OpenSpec validation passed. |
|
||||
|
||||
## Current State
|
||||
|
||||
- Baseline diff is implemented for aggregate metrics, verdict distribution, case-level state, keyword coverage, evidence coverage, missing cases, and new cases.
|
||||
- JSON and Markdown diff output are available.
|
||||
- No production runtime behavior has been changed.
|
||||
|
||||
## Verification
|
||||
|
||||
### Script Verification
|
||||
|
||||
- Command: `mvn -q "-Dtest=DiagnosisEvalBaselineDiffTest" test`
|
||||
- Result: passed
|
||||
- Notes: Also verified with `DiagnosisTraceEvaluatorTest`.
|
||||
|
||||
### Static Verification
|
||||
|
||||
- Command: `mvn -q -DskipTests compile`
|
||||
- Result: passed
|
||||
|
||||
### OpenSpec Verification
|
||||
|
||||
- Command: `openspec validate diagnosis-eval-baseline-diff --strict`
|
||||
- Result: passed
|
||||
@@ -0,0 +1,30 @@
|
||||
# Brief: diagnosis-eval-baseline-diff
|
||||
|
||||
## Background
|
||||
|
||||
The eval harness now has a complete saved baseline. This change adds the comparison layer that turns the baseline into an actionable regression signal.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Compare baseline and current `DiagnosisEvalReport` objects.
|
||||
2. Detect aggregate and per-case regressions.
|
||||
3. Output JSON and Markdown diff reports.
|
||||
4. Document how to read the diff in interview and engineering terms.
|
||||
|
||||
## Scope
|
||||
|
||||
- Diff data structures
|
||||
- Deterministic report comparison
|
||||
- JSON / Markdown diff output
|
||||
- Focused tests and eval docs
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No live Agent execution
|
||||
- No LLM-as-judge
|
||||
- No evaluator scoring rule changes
|
||||
- No production API changes
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/diagnosis-eval-baseline-diff/`
|
||||
@@ -0,0 +1,28 @@
|
||||
# Diagnosis Eval Baseline Diff Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: add report diffing on top of the completed diagnosis eval baseline.
|
||||
- Slug: `diagnosis-eval-baseline-diff`
|
||||
- Devflow scale: standard-light
|
||||
|
||||
## Context
|
||||
|
||||
- `diagnosis-eval-harness` created deterministic fixture evaluation.
|
||||
- `expand-diagnosis-eval-fixtures` created a complete saved baseline.
|
||||
- This change compares new reports against that baseline.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Diff report DTOs instead of raw traces.
|
||||
- Reason: the report is the stable contract for regression review.
|
||||
|
||||
- Decision: Use deterministic code rules instead of LLM-as-judge.
|
||||
- Reason: baseline regression checks should be repeatable and explainable.
|
||||
|
||||
- Decision: Output both JSON and Markdown.
|
||||
- Reason: JSON supports automation; Markdown is useful in reviews and interviews.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Whether a future change should expose this through a CLI or Maven goal.
|
||||
@@ -0,0 +1,11 @@
|
||||
# Evidence: diagnosis-eval-baseline-diff
|
||||
|
||||
## Evidence Log
|
||||
|
||||
- 2026-07-05: Created slug-based issue `diagnosis-eval-baseline-diff.md`.
|
||||
- 2026-07-05: Created OpenSpec change `diagnosis-eval-baseline-diff`.
|
||||
- 2026-07-05: Added baseline diff DTOs, deterministic comparer, and JSON / Markdown writer.
|
||||
- 2026-07-05: Added sample baseline diff JSON and Markdown reports.
|
||||
- 2026-07-05: Verification passed with `mvn -q "-Dtest=DiagnosisEvalBaselineDiffTest,DiagnosisTraceEvaluatorTest" test`.
|
||||
- 2026-07-05: Verification passed with `mvn -q -DskipTests compile`.
|
||||
- 2026-07-05: Verification passed with `openspec validate diagnosis-eval-baseline-diff --strict`.
|
||||
@@ -0,0 +1,25 @@
|
||||
# Acceptance: mvp-demo-interview-runbook
|
||||
|
||||
## Classification
|
||||
|
||||
standard-light
|
||||
|
||||
## Task Status
|
||||
|
||||
| Task | Status | Notes |
|
||||
| --- | --- | --- |
|
||||
| Issue and OpenSpec setup | Done | Created slug-based issue and OpenSpec artifacts. |
|
||||
| Implementation | Done | Added request payload, runnable script, output directory docs, interview walkthrough, and trace checklist. |
|
||||
| Verification | Done | OpenSpec validation passed. |
|
||||
|
||||
## Current State
|
||||
|
||||
- No backend runtime behavior has been changed.
|
||||
- Demo is packaged under `mvp/demo` for interview use.
|
||||
|
||||
## Verification
|
||||
|
||||
### OpenSpec Verification
|
||||
|
||||
- Command: `openspec validate mvp-demo-interview-runbook --strict`
|
||||
- Result: passed
|
||||
@@ -0,0 +1,29 @@
|
||||
# Brief: mvp-demo-interview-runbook
|
||||
|
||||
## Background
|
||||
|
||||
Plan C is the interview-facing demo package. The project has the engineering pieces, but needs a single place to run and explain the MVP flow.
|
||||
|
||||
## Goals
|
||||
|
||||
1. Provide a fixed payment-timeout request payload.
|
||||
2. Provide a PowerShell script that runs chat, trace, and feedback.
|
||||
3. Save demo responses under `mvp/demo/output`.
|
||||
4. Add interview walkthrough and trace checklist.
|
||||
|
||||
## Scope
|
||||
|
||||
- Demo docs and scripts only
|
||||
- Existing local APIs only
|
||||
- Existing `mvp-demo` profile only
|
||||
|
||||
## Non-Goals
|
||||
|
||||
- No backend code changes
|
||||
- No eval extension
|
||||
- No secret cleanup
|
||||
- No full offline runtime
|
||||
|
||||
## Related OpenSpec
|
||||
|
||||
`openspec/changes/mvp-demo-interview-runbook/`
|
||||
@@ -0,0 +1,27 @@
|
||||
# MVP Demo Interview Runbook Decisions
|
||||
|
||||
## Clarify
|
||||
|
||||
- Entry summary: package existing MVP capabilities into a repeatable interview demo.
|
||||
- Slug: `mvp-demo-interview-runbook`
|
||||
- Devflow scale: standard-light
|
||||
|
||||
## Context
|
||||
|
||||
- Evidence trace and eval baseline work are already done.
|
||||
- The next useful step is not more eval tooling, but a runnable demo path.
|
||||
|
||||
## Key Decisions
|
||||
|
||||
- Decision: Keep this change documentation/script-only.
|
||||
- Reason: Plan C is about demo packaging, not new runtime capability.
|
||||
|
||||
- Decision: Use a stable session id.
|
||||
- Reason: it makes trace lookup and saved output predictable.
|
||||
|
||||
- Decision: Save outputs to `mvp/demo/output`.
|
||||
- Reason: generated artifacts should be easy to review without mixing into source fixtures.
|
||||
|
||||
## Open Questions
|
||||
|
||||
- Whether a later change should add a truly offline stubbed demo mode.
|
||||
@@ -0,0 +1,9 @@
|
||||
# Evidence: mvp-demo-interview-runbook
|
||||
|
||||
## Evidence Log
|
||||
|
||||
- 2026-07-05: Created Plan C demo packaging issue and OpenSpec change.
|
||||
- 2026-07-05: Added fixed payment-timeout request payload.
|
||||
- 2026-07-05: Added PowerShell demo script for chat, trace, and feedback.
|
||||
- 2026-07-05: Added interview walkthrough and trace inspection checklist.
|
||||
- 2026-07-05: Verification passed with `openspec validate mvp-demo-interview-runbook --strict`.
|
||||
Reference in New Issue
Block a user