Quick answer
Serverless inference deploys an AI model as an endpoint that starts when a request arrives, serves it, and shuts down when traffic stops, billing only for the seconds of compute used. RunPod, Modal, Baseten, and Replicate all offer it. The advantage is near-zero cost for low or bursty traffic; the disadvantage is the cold start, the seconds to a minute it takes to load a model onto a fresh GPU for the first request.
A dedicated GPU costs the same whether it serves a thousand requests an hour or one. For an app with unpredictable traffic that is money spent on idle silicon. Serverless inference is the answer, and the cold start is its price.
How it works
You package a model with its dependencies, usually as a container. The platform keeps the package ready but runs nothing. When a request arrives it allocates a GPU, loads the model into memory, runs inference, and returns the result. If more requests follow it keeps the instance warm; if they stop it releases the GPU. You pay for the seconds the instance existed.
The cold-start problem
Loading a multi-gigabyte model onto a GPU takes time, from a few seconds for small models to a minute or more for large ones. The first request after an idle period waits for it. Platforms attack this with warm pools (keeping some instances ready, which costs money), memory snapshots that restore a loaded model faster, and minimum-instance settings that turn "serverless" into "mostly dedicated". Each mitigation trades some of the cost advantage for speed.
When to use it
- Side projects and internal tools with sporadic use: serverless, no question
- Batch jobs and scheduled workloads: serverless, since latency does not matter
- User-facing apps with steady traffic: dedicated capacity is usually cheaper and always more predictable
- Latency-critical apps with bursty traffic: serverless with a warm minimum, and measure the bill
Bottom line
Serverless inference makes running your own model affordable when traffic is light or lumpy. Understand the cold start, decide how much of it your users can tolerate, and pay for warmth only where they cannot.



