1. Codex SSE wrapper terminates on exact `data: [DONE]` only.
The old substring check `if "[DONE]" in line` would flip
sent_done True when a normal model response carried the literal
text "[DONE]" in delta.content (for example an explanation of
the OpenAI stream sentinel). The real terminator was then
suppressed, leaving OpenAI-compatible clients that finalise on
the explicit sentinel hung on stream close. Now compares the
stripped line to the exact `data: [DONE]` form.
2. Legacy `thread.run_streaming` path no longer returns an empty
reply on completion-only streams.
If the SDK exposes `thread.run_streaming` but the stream emits
ONLY item.completed / agentMessage events with no message
deltas, the loop previously exited with emitted_any False and
never reached the agent-message fallback. The request returned
200 with an empty assistant reply even though Codex produced a
final answer. Mirror the canonical-path behavior: collect
`_completed_agent_message_text` strings in a sidecar list and
emit the last one when no deltas arrived. Match the canonical
payload-extraction (`getattr(event, "payload", event)`) so the
event-vs-payload SDK shape difference is handled the same way
in both branches.
3. Parallel-calls fan-out propagates CodexUnavailableError so the
route layer can return 503.
When the SDK is not importable or the safety enums are missing
without the dev opt-in, every worker raised the same
CodexUnavailableError. The previous catch-all converted the
error into a per-tab codex_tab_error event, the outer stream
never raised, and clients saw a 200 with only tool events and
an empty synthesis -- OpenAI-compatible consumers that ignore
_toolEvent saw a successful empty reply. Now CodexUnavailableError
re-raises out of the worker (no spurious per-tab error event),
_await_workers re-raises it when EVERY worker hit the same
setup failure, and the finally-block drain await propagates the
exception out of the parallel function so the route's existing
CodexUnavailableError handler can emit the right 503 SSE error
frame. Per-tab runtime failures (model rejected, timeout, mid-
stream SDK crash) still get swallowed into codex_tab_error
events so a single bad model in the fan-out does not kill the
others.
Test counts: 63/63 passing (60 round 6-8 plus 3 new round 9 regression
tests). Each new test was first run against a `git stash`-restored
pre-fix tree to confirm it catches the bug, then run against the
patched tree.