HNHacker News
TopNewBestAskShowJobs

pveldandi

4 karma · joined July 17, 2023

Building InferX, a GPU-native runtime that snapshots the full execution state of LLMs so you can hot-swap models like threads.

Obsessed with inference efficiency, cold-start elimination, and agentic infra.

Previously: enterprise software, now deep in AI infra

Say hi on LinkedIn: https://www.linkedin.com/in/prashanth-v-98629b115/

submissionscomments

Show HN: How We Run 60 Hugging Face Models on 2 GPUs

4 pts·pveldandi·
20

Benchmark: A100 vs. H100 NVMe Random Read throughput during multi-GPU loading

1 pts·pveldandi·
0

Show HN: 50+ LLMs on 2 GPUs with 2-Second Swapping? We built AI-Native Runtime

github.com·5 pts·pveldandi·
0

Show HN: InferX - AI Lambda-Like Inference Function as a Service

2 pts·pveldandi·
0

We're running 50 LLMs on 2 GPUs – no cold starts, no overprovisioning

4 pts·pveldandi·
1

Show HN: InferX – an AI-native OS for running 50 LLMs per GPU with hot swapping

3 pts·pveldandi·
2