88 karma · joined February 21, 2018
My benchmarks showed vLLM delivering up to 3.2x the requests-per-second of Ollama on identical hardware, with noticeably lower latency at high concurrency.
If you're not looking for the ultimate performance on the latest GPU hardware, then Ollama is still hard to beat. It installs in minutes, runs on laptops, supports CPU fallback, and provides a curated model hub plus on-the-fly model switching. If your typical load is a handful of concurrent users, batch jobs that can wait an extra second, or local exploration during development, Ollama’s “good-enough” performance is exactly that, good enough.
Ollama is the reliable daily driver that gets almost everyone where they need to go; vLLM is the tuned engine you unleash when the freeway opens up and you really need to fly.
An example knowledge graph it created is located here: https://robert-mcdermott.github.io/ai-knowledge-graph/
But had to stop using it last year when all the binaries it generated where being detected by CrowdStrike as a Trojan. The binaries would dissapear on execution and I'd be contacted by the security office. I uploaded a program to totalvirus and a few other AV systems also detected it as a trojan. At that point I stopped using it and switched to Go. Has that situation been fixed yet?