HNHacker News
TopNewBestAskShowJobs

pveldandi

4 karma · joined July 17, 2023

Building InferX, a GPU-native runtime that snapshots the full execution state of LLMs so you can hot-swap models like threads.

Obsessed with inference efficiency, cold-start elimination, and agentic infra.

Previously: enterprise software, now deep in AI infra

Say hi on LinkedIn: https://www.linkedin.com/in/prashanth-v-98629b115/

submissionscomments
pveldandi··on Show HN: How We Run 60 Hugging Face Models on 2 GPUs
Happy to engage with you if Hulu have any additional questions. Thanks for all the feedback.that was helpful
pveldandi··on Show HN: How We Run 60 Hugging Face Models on 2 GPUs
Fair point. I’ll repost as a regular submission instead of Show HN. The goal was to demonstrate the runtime behavior behind multi-model serving rather than a polished end-user app. Appreciate the clarification.
pveldandi··on Show HN: How We Run 60 Hugging Face Models on 2 GPUs
It’s an open-core model. The control plane is already open source and can be deployed fairly easily. We’re not trying to replace in-house systems or hyperscalers. This can run on Kubernetes and integrate into existing infrastructure. The runtime layer is where we’re focusing the differentiation.
pveldandi··on Show HN: How We Run 60 Hugging Face Models on 2 GPUs
It’s an open-core model. The control plane is already open source and can be deployed fairly easily. We’re not trying to replace in-house systems or hyperscalers. This can run on Kubernetes and integrate into existing infrastructure. The runtime layer is where we’re focusing the differentiation.
pveldandi··on Show HN: How We Run 60 Hugging Face Models on 2 GPUs
The demo is live. It’s meant to show how snapshot restore works inside a multi-tenant runtime, not just a prompt playground. You can interact with the deployed models and observe how state is restored and managed across them. The focus is on the runtime behavior rather than a chat UI.
pveldandi··on Show HN: How We Run 60 Hugging Face Models on 2 GPUs
Happy to dig deeper and show how exactly it works under the hood. For context, here’s the main site where the architecture and deployment options are explained: https://inferx.net/
pveldandi··on Show HN: How We Run 60 Hugging Face Models on 2 GPUs
Also, For what it’s worth, this can be deployed both on-prem and in the cloud. Different teams have different constraints, so we’re trying to stay flexible on that.
pveldandi··on Show HN: How We Run 60 Hugging Face Models on 2 GPUs
Basic html? The core of what we built is at the runtime layer. We’re capturing CUDA graphs and restoring model state directly at the GPU execution level rather than just snapshotting containers. That’s what enables fast restores and higher utilization across multiple models.

If that’s not a problem space you care about, that’s totally fair. But for teams juggling many models with uneven traffic, that’s where the economics start to matter.

pveldandi··on Show HN: How We Run 60 Hugging Face Models on 2 GPUs
Vertex is great control plane. We’re not replacing them.

What we focus on is the runtime layer underneath. You can run us behind Cloud Run or inside your existing GCP setup. The difference is at the GPU utilization level when you’re serving many models with uneven demand.

If your workload is steady and high volume on a small set of models, the standard cloud stack works well. If you’re juggling dozens of models with spiky traffic, the economics start to look very different.

As an example, we’re currently being tested inside GCP environments. Some teams are experimenting with running us behind their existing Google Cloud setup rather than replacing it. The idea isn’t to swap out Cloud Run or Vertex, but to improve the runtime efficiency underneath when serving multiple models with uneven demand.

pveldandi··on Show HN: How We Run 60 Hugging Face Models on 2 GPUs
Good question.

This isn’t for single-model apps running steady traffic at high utilization. If you’re saturating GPUs 24/7, you’ll architect very differently.

This is for teams that…

• Serve many models with uneven traffic • Run per-customer fine-tunes • Offer model marketplaces • Do evaluation / experimentation at scale • Have spiky workloads • Don’t want idle GPU burn between requests

A lot of SaaS AI products fall into that category. They aren’t OpenAI-scale. They’re running dozens of models with unpredictable demand.

Lambda exists because not every workload is steady state. Same idea here.

pveldandi··on Show HN: How We Run 60 Hugging Face Models on 2 GPUs
You can try it here: https://inferx.net:8443/demo/
pveldandi··on Show HN: How We Run 60 Hugging Face Models on 2 GPUs
Ollama is great for local workflows. What we’re focused on is multi-tenant, high-throughput serving where dozens of models share the same GPUs and scale to zero without keeping them resident.

You’re right that HN expects something runnable. We’re spinning up a public endpoint so people can test with their own models directly instead of requesting access. I’ll share it shortly. Thank you for the suggestion.

pveldandi··on [dead]
Lambda made compute cheap by treating code as ephemeral and execution state as disposable.

LLM inference breaks that model. The execution state (weights shards, CUDA context, KV cache) is expensive and GPU-resident, so “scale to zero” usually means full re-initialization and cold starts.

We’ve been working on a runtime that treats model execution state as the unit of scheduling, not the container or service. Instead of tearing everything down, we snapshot per-GPU execution state and restore it on demand, so models can scale down without paying the full cold-start cost again.

This makes “model = function” closer to practical: fast scale-up, bounded latency, and GPUs shared across many models without keeping them always on.

Let us know What you think about serverless inference, especially where the abstraction breaks today (KV cache, memory pressure, multi-GPU models, scheduling).

pveldandi··on Show HN: Checkpoint K8s pods transparently (plain CPU or GPU accelerated) [video]
Really interesting work. we’ve been building a container-native snapshotting system too, but focused on cold start reduction and multi-model orchestration for LLM inference.

Different use case (sub-2s loading for large models), but very similar challenges around memory, device state, and restore reliability.