Replicate vs. Fly GPU cold-start latency
venki.dev
venki.dev
Here's what we're doing:
- Fine-tuned models now boot fast: https://replicate.com/blog/fine-tune-cold-boots
- You can keep models switched on to avoid cold boots: https://replicate.com/docs/deployments
- We've optimized how weights are loaded into GPU memory for some of the models we maintain, and we're going to open this up to all custom models soon.
- We're going to be distributing images as individual files rather than as image layers, which makes pulling images much more efficient.
Although our cold boots do suck, the comparison in this blog post is comparing apples to oranges because Fly machines are much lower level than Replicate models. It is more like a warm boot.
It seems to be using a stopped Fly machine, which has already pulled the Docker image onto a node. When it starts, all it's doing is starting the Docker container. Creating the Fly machine or scaling it up would take much longer.
On Replicate, the models auto-scale on a cluster. The model could be running anywhere in our cluster so we have to pull the image to that node when it starts.
Something funny seems to be going on with the latency too. Our round-trip latency is about 200ms for a similar model. Would be curious to see the methodology, or maybe something was broken on our end.
But we do acknowledge the problem. It's going to get better soon.
I feel like it would be worth a deep dive with your team on what's happening and maybe writing a blog post on what you found?
Also, I'll gently point out that Fly not having to pull Docker images on "cold" boot isn't something your customers think much about, since a stopped Fly machine doesn't accrue additional cost (other than a few cents a month for rootfs storage). If it's roughly the same price, and roughly the same level of effort, and ends up performing the same function for the customer (inference), whether or not it's doing Docker image pulls behind the scenes doesn't matter so much to most customers. Maybe it's worth adding a pricing tier to Replicate that charges a small amount for storage even for unused models, and results in much better cold boot time for those models since you can skip the Docker image pull — or in the future, model file download — and just attach a storage device?
(I know you're also selling the infinitely autoscaling cluster, but I think for a lot of people the tradeoff between finite-autoscaling vs extremely long cold boot times is not going to be in favor of the long cold boots — so paying a small fee for a block storage tier that can be attached quickly for autoscaling up to N instances would probably make a lot of sense, even if scaling to N+1 instances is slow again and/or requires clicking a button or running a CLI command.)
(There's a lot I can say about why I think a benchmark like this is showing us unusually well! I'm not trying to argue that people should take this benchmark too seriously.)
This is one of the things we (at https://fal.ai) working very hard to solve. Because of ML workloads and their multiple GB environments (torch, all those cuda/cudnn libraries, and anything else they pull) it is a real challange just to get the container to start in a reasonable time frame. We had to write our own shared Python virtual environment runtime using SquashFS distributed thru a peer-to-peer caching system to bring it down sub-second mark.
After the container boots, there is the aspect of storing model weights, which IMHO less challenging since it is just big blobs of data (compared to Python environments where there are thousands of smaller files where each might be sequentially read and incur a really major latency penalty). Distributing them once we had the system above was super easy since just like squashfs'd virtual environments, they are immutable data blobs.
We are also starting to play with GPUDirect on some of our bare metal clusters and hopefully planning to expose it to our customers, which is especially important if your models is 40GB or higher. At that point, you are technically operating at the PCIE/SXM speeds which is ~2-3 seconds for a model of that size.
Are you using GDS or networking or both?
The main costs were:
- gpu time for training
- gpu time for inference
- storage costs for the users' models
- egress fees to download model
I ended up using banana.dev and runpod.io for the serverless gpus. Both were great, easy to hook into, and highly customizable.
I spent a bunch of time trying to optimize download speed, egress fees, gpu spot pricing, gpu location, etc.
R2 is cheaper than s3 - free egress! But the download speeds were MUCH worse than s3 - enough that it ended up not even being competitive.
It was frequently cheaper to use more expensive GPUs w/ better location and network speeds. That factored more into the pricing than how long the actual inference took on each instance.
Likewise, if your most important metric is time from boot to starting inference then network access might be the limiting factor.
While we loved the dev experience we just couldn’t make it work with frequently switching models / LORA weights.
We switched to beam (https://www.beam.cloud) and it’s so much better. Their cold start times are consistently small and they provide caching layer for model files i.e volumes which make switching between models a breeze.
Beam also has much better pricing policy. For custom models on replicate you pay for boot times (which are very long!) so you are paying a lot of $ for a single request.
With beam you only pay for inference and idle time.
Incentives are aligned for us to make it better. :)
“[…] Unlike public models, you’ll pay for boot and idle time in addition to the time it spends processing your requests.”
Apart from boot times, we actually find replicate to be an amazing platform, congrats
https://open.substack.com/pub/jonolson/p/replicatecom-review...
I can literally boot the server for the 10-20s it takes to run a bunch of generations, and have it shut down automatically afterwards. It feels like magic.
Sure, creating the image after a new deployment takes up to two minutes, but once it’s there it’s incredibly fast.
I ask this because Fly has immutable Docker containers which wouldn't store any data unless you use Fly Volumes. So it could be that Fly is downloading the 100MB model each time it cold-boots.
If that's the case, a multi-stage Dockerfile could help in bundling the model in, and perhaps reducing cold-boot time even further.
So the problem of size is exactly the same, and RUN with curl will have identical size as COPY layer from COPY --from=stage
Am I missing something?
I can see only benefit for build cache reuse, so download is independed from building code so you won't redownload 14G when you change code, is that what you had in mind?
I don't know Docker that well. I literally figured it out as I went along to deploy on Fly...
I didn't wanted to sound snarky, I thought that I wasn't aware about some cool docker optimization hack :)
I tested it initially because it’s the naivest implementation. The right implementation would bundle it in.
But I ended up primarily reporting timings that stop counting up as soon as control is handed over to user generated code - since that’s the number you care about the most.
Perhaps a good future idea would be to benchmark between bundling it in the Docker image, vs. using Fly Volumes as Simon suggested in a sibling comment.
Cold starting containers quickly is a fascinating problems. We've gotten a long way but there's still a lot more to do. For GPU-based inference, starting containers isn't enough – you also need to initialize the model GPU quickly. We are working on a long list of things that will bring down cold start latency even further.
Replicate created the cog spec and is a fantastic resource for browsing and playing with new models. They are a social destination too.
But flyIO is nice and simple for docker side projects, I hope their cog deployments are as smooth as there are a few extra pieces involved.
They could call it and just pay for the time spent, not a persistent server.
The issue is cold start time for custom models. It takes time to pull in dependencies, the model, and load the model into memory.
It’s difficult because someone has to pay for the cold start time.