24 karma · joined August 2, 2021
Where it starts to get harder is when you have multiple base stacks (different CUDA versions, frameworks, etc.) or when you need to update them frequently. You end up with lots of slightly different multi-GB bases.
Chunked images keep the benefit you mentioned (we still cache heavily on the nodes) but the caching happens at a finer granularity. That makes it much more tolerant to small differences between images and to frequent updates, since unchanged chunks can still be reused.
In practice these systems typically fetch data over a local, highly available network and aggressively cache anything that gets read. If that network path becomes unavailable, it usually indicates a much larger infrastructure issue since many other parts of the system rely on the same storage or registry endpoints.
So while it does introduce a different failure mode, in most production environments it ends up being a low practical risk compared to the startup latency improvements.
For us and our customers, the trade off is worth it.
even the smaller nvidia images (like nvidia/cuda:13.1.1-cudnn-runtime-ubuntu24.04) are about 2Gb before adding any python deps and that is a problem.
if you split the image into chunks and pull on-demand, your container will start much faster.
- clearer messaging - more tutorials - one-click deploys - clear & upfront costing
We have plans to add other runtimes (like Typescript) in the future but Python is our focus for now.
Cerebrium abstracts some functionality - like streaming and batching endpoints. I think you would need to build that yourself on paperspace.