We're running 50 LLMs on 2 GPUs – no cold starts, no overprovisioning
We built InferX, a model runtime that snapshots the full GPU execution state, weights, memory layout, KV cache and resumes any model in ~2 seconds. No reinitialization, no weight reloading, no containers.
With this, we’re running 50+ LLMs on just 2 A1000 GPUs, with cold starts eliminated and memory orchestrated like threads. Traditionally, this would take 70+ GPUs if you pinned each model.
We’re not doing speculative batching or model merging. this is native orchestration at the runtime layer.
This was built to support:
•agent stacks (each agent using its own model)
•tenant-specific fine-tunes
•long-tail workloads where models don’t get constant traffic
Would love to hear how others are solving this or what you’re seeing in the multi-model inference space. Happy to go into technical detail on snapshotting, memory management, or orchestration strategy.
Ask me anything.