If a process started by prewarm() is never claimed, what happens to it? Could this waste resources?
The official changelog doesn't detail the lifecycle management of an unclaimed spare process (for instance, whether there's a timeout-based automatic reclaim mechanism), which is worth flagging given how sparse the current documentation is. It's reasonable to infer from the design logic that, since this is an API set for managing process resources, there's likely some form of timeout or cap to prevent spare processes from accumulating indefinitely — but until the official documentation is more complete, the more cautious approach is to add your own monitoring at the application layer, confirming that the number and lifetime of prewarmed processes stay within a reasonable range, rather than assuming the underlying layer will always clean up for you.
Since prewarm/SpareProcess is conceptually similar to the existing startup(), can the two be used together, or are they mutually exclusive?
The official changelog lists these as separate feature items, and their descriptions don't fully overlap in framing — startup() is an existing performance-optimization mechanism dating back to v0.2.89, while prewarm/SpareProcess is a new, alpha-labeled mechanism that further splits "starting" and "claiming" into two steps. The public documentation doesn't clearly state whether the two are designed to be mutually exclusive or can be layered together, which is worth flagging — if your application already uses startup(), the more cautious approach before adopting prewarm is to test on a small scale first whether having both present causes any conflict or duplicate process startup, rather than assuming they can be stacked without issue.
For an API labeled alpha, how long should you generally wait before considering it for production use?
There's no universal timeframe — it depends on how this kind of API has evolved within the same SDK historically. A more practical way to judge is to watch whether this API set shows up with "breaking change" or "behavior adjustment" entries in the changelog over several subsequent versions. If the interface goes several consecutive versions without further changes, with only small bug fixes continuing, that generally signals it's moving toward stability. Conversely, if the interface is still being adjusted frequently in the near term, that signals the risk of adopting it in production right now is still elevated, and every SDK upgrade might require re-checking whether this part of your code is affected.
If my application is a long-running server-side service (not restarting a process every time), does prewarm() still matter for me?
It matters much less than in a "cold-start every time" scenario, but not zero. If your server already maintains a process pool, runs persistently, and handles multiple user sessions repeatedly, the startup cost of a single process is already amortized across the service's lifetime and isn't paid again on every request — in that architecture, prewarm's marginal benefit is limited. But if your architecture spins up a brand-new process per session or per request (certain serverless or containerized deployment patterns, where every call is a fresh execution environment), prewarm is still valuable — you can get the process ready during the gap before the request arrives and before parameters are known.
When building applications with the Claude Agent SDK, every call to query() actually requires starting a Claude Code subprocess first, and that startup process itself takes time — for scenarios where first-query latency matters a lot (a user opening the app, sending their first message, and having to wait several seconds before a response even starts streaming), this process-startup cost shows up directly as perceived user wait time. The new prewarm() and SpareProcess.claim() (currently marked alpha, meaning experimental and the interface may still change) offer an approach: decoupling "starting a process" from "actually beginning a session" so they're handled separately.
The traditional flow is: you call query(), and only then does the SDK start a Claude Code process, bind it to the folder path and various session-level options this session needs, and begin actually processing your message. prewarm() breaks that ordering — it lets you start a process at a point when you don't yet know which session it will ultimately serve, leaving it in a "spare" state. Then, once you're actually ready to begin a session and know which folder and which options to bind, you use SpareProcess.claim() to "claim" that already-started spare process, attaching it directly to this session's configuration and skipping the time it would otherwise take to start a process from scratch.
This mechanism is best suited to situations where "you roughly know a session will be needed soon, but the exact session parameters aren't settled yet" — for instance, a user opens a chat interface but hasn't typed anything yet; you can use that gap to call prewarm(), and once the user actually sends their first message and you know which folder and which tool permissions to use, you call claim() to attach the spare process. Because the process-startup latency has effectively been amortized during the time the user spent typing, the wait the user actually experiences between "sending a message" and "a response starting" shrinks noticeably. This is conceptually similar to startup() (an existing feature added in SDK v0.2.89) — both move startup cost off the critical path — but prewarm/SpareProcess goes further by splitting "starting" and "claiming" into two independent actions, giving you more flexibility over when each step happens.
Marking this API set as alpha means the interface design itself may still be adjusted in later versions, and there may even be unpublicized edge cases that aren't fully handled yet. If you're considering adopting it in production, the more cautious approach is to try it first on non-critical paths, or in scenarios tolerant of behavioral changes, while closely watching the SDK's subsequent changelogs — and rely on it fully only once the API has settled down. The official changelog's own description of this API set is currently fairly sparse with no detailed usage examples; more complete documentation needs to be found in the official TypeScript API reference.
If your application is particularly sensitive to first-query latency — a customer-support chat interface, a real-time assistant product, where the wait between a user typing and the first response appearing directly shapes the felt experience — prewarm() is worth putting on your performance-optimization backlog: evaluate in a non-production environment how much actual latency improvement it delivers first, then weigh whether it's worth accepting alpha-stage API-change risk to bring into production. If your use case isn't particularly sensitive to first-query latency (Batch Processing, background jobs), this feature offers limited benefit, and there's no need to take on the maintenance cost of an alpha API just for it.