Why would this bug be hard to catch during normal testing?
Because it's a pure timing race condition — it only triggers when a background Subagent happens to finish at precisely the moment the main result arrives. In ordinary testing, if the background task finishes much faster or much slower than the main result, you never land in that window; the connection-close timing and the background task's completion timing miss each other, and everything works fine. That's also why this class of bug tends to get described as "intermittent" or "not reliably reproducible" — it's not that testing wasn't thorough enough; the Trigger Condition itself is a narrow timing window, and a production environment running many background tasks in parallel is actually more likely to hit that window than a development environment is.
What happens if CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS is set too long or too short?
Set it too short (a few seconds, say) and you've essentially eliminated the benefit of the longer wait — if a background task genuinely needs a few minutes to finish, too short a ceiling forces the SDK to move forward before the task has actually completed, effectively reintroducing the pre-fix "closed too early" risk, just with the trigger shifted from "race-condition coincidence" to "an inevitable consequence of too short a ceiling." Set it too long (stretched to several hours), and in a scenario where a background task is genuinely stuck and will never report idle, the connection hangs for that much longer, meaning the user or caller waits longer before noticing anything is wrong.
The more practical approach is to first observe roughly how long your actual background tasks take to complete under normal conditions, then pick a buffer value slightly above that, rather than simply sticking with the 10-minute default or setting an arbitrarily large number.
If my deployment's CLI version is older and doesn't support session_state_changed, is upgrading the SDK still worth doing?
It's still worth doing, but the benefit is reduced. The SDK side automatically detects whether the CLI emits session_state_changed events, and falls back to the older "close on first result" logic if it doesn't — meaning upgrading the SDK by itself won't trigger any error or compatibility issue, but this particular bug fix won't actually take effect, and you could still run into the original race condition.
To actually get the benefit of this fix, upgrading the SDK and upgrading the CLI to a version that supports session_state_changed are both things you need to do — doing only one of the two won't fully resolve the issue.
Does this fix only matter for applications using hooks, can_use_tool, or an SDK MCP Server? Should apps not using those features care about this bug?
The changelog specifically identifies this bug as occurring in the context of "using query() together with hooks, can_use_tool, or an SDK MCP server," which indicates the Trigger Condition is tied to background-Subagent lifecycle management — and those features all involve additional processing that can run in the background. If your application doesn't use any background-subagent mechanism at all (a purely synchronous question-and-answer flow with no parallel tool calls or background validation process), you likely won't hit this particular race condition — but upgrading the SDK itself has no downside, so doing it as routine maintenance is still reasonable.
When using the Claude Agent SDK's query() together with hooks, a can_use_tool callback, or an SDK-provided MCP Server, there was a specific scenario that could reliably trigger a connection break: if a background Subagent happened to finish right around the moment the main conversation turn's result was about to arrive, the SDK's older logic would close stdin too early. The result was that the following turn would fail outright with a "Stream closed" error, and on the model's side, this would get misread as "this tool call was refused" — when in fact nothing was wrong with the tool call itself; the problem was mistimed lifecycle management of the underlying connection.
Before this fix, the SDK's logic for deciding "this turn can close stdin now" was keyed on "received the first result message" — which works fine in most cases, but when a background subagent happens to finish right at the moment the main result arrives, a race condition appears: the SDK thinks the turn is over and closes the connection early, while the CLI side is actually still wrapping up the background subagent's work and needs that same connection to keep communicating. The new fix instead listens for the CLI's own session_state_changed messages and only closes stdin once the CLI explicitly reports its state as idle (genuinely idle, with no background work running) — rather than inferring readiness from the indirect signal of "first result received."
Simply switching to "wait until the CLI reports idle" theoretically solves the early-close problem, but introduces a new risk: if some background task gets stuck for whatever reason and never reports idle, the connection would hang indefinitely, producing a different kind of malfunction. The new fix adds an environment variable, CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS, defaulting to 10 minutes, as a ceiling on how long to wait for the idle state — past that point, even if the background task hasn't reported completion, the SDK forces things forward rather than hanging indefinitely. For scenarios where background tasks genuinely need to run longer than 10 minutes (a long-running data-processing job, say), this environment variable lets you extend the ceiling yourself.
This fix depends on the CLI proactively emitting session_state_changed messages, but not every CLI version implements this mechanism. For older CLIs without this state event, the SDK automatically falls back to the old behavior pattern — using arrival of the first result as the close signal. This means the actual benefit of this bug fix depends on the CLI version in your deployment: if the CLI is too old, even with an updated SDK, you're still running the older logic with its race-condition risk underneath.
If your application makes heavy use of background subagents — parallel processing of multiple sub-queries, or having a subagent run a time-consuming validation process in the background — this class of "Stream closed" error paired with a falsely-perceived tool-call refusal may have been treated by your team as a random, hard-to-reproduce instability issue in the past. Now that the root cause is known, upgrade the SDK to a version containing this fix, and separately confirm your deployment's CLI version supports session_state_changed — only then do you actually get the benefit of this fix. If your background tasks routinely need to run longer than 10 minutes, remember to set CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS accordingly, or the default 10-minute ceiling may force the connection forward before the task has actually finished.