Root cause:
- RunStore.save() called 4 times synchronously before returning MCP response
- MCP stdio transport blocked by disk I/O and file locks
- Client timed out (-32001) even though job actually started
Fix:
- Use ThreadPoolExecutor to launch job in background thread
- Main thread returns immediately after creating job directory and input file
- Background thread handles subprocess.Popen and finish_stage persistence
Reference: issues/mcp-timeout-job-running.md