Why would I use those models on your cloud instead of using Google's or Anthropic's models? I'm glad there are open models available and that they get better and better, but if I'm paying money to use a cloud API I might as well use the best commercial models, I think they will remain much better than the open alternatives for quite some time.
Fast forward to now, open models are quickly catching up, and at a significantly lower price point for most and can be customized for specific tasks instead of being general purpose. For general purpose models, absolutely the closed models are currently dominating.
Or spend 20 a month for models even a 5090 couldn't run. And not have to spend your own electricity, hardware, maintenance, updates etc.
This is why everyone needs to get every flavour and speedrun building all the tools they need when the infinite money faucets are turned off.
At some point companies will start raising prices or moving towards per-token pricing (Which is sustainable, but expensive).
And with that in mind, i definetly dont use more than a couple of bucks a month in API refils. (not that i really am a power user or anything)
So if you consider the 20 bucks to be balanced between poer and non power users, and with the existing rate limits, its probably not that far off being profitable, at least on the pure inference side.
This is almost exactly how duckdb/motherduck functions and I think theyre doing an excellent job.
EDIT: grammar and readability
I tried it a while back, I was very surprised to find that simply running `uvx ramalama run deepseek-r1:1.5b` just worked. I'm on Fedora Silverblue with nothing layered on the ostree. Before RamaLama, getting llama.cpp working with my GPU was a major PITA.
* Work with somebody like System76 or Framework to create great hardware systems come with their ecosystem preinstalled.
* Build out a PaaS, perhaps in partnership with an existing provider, that makes it easy for anybody to do what Ollama search does. I'm more than half certain I could convince our cash strapped organization to ditch elastic search for that.
* Partner with Home Assistant, get into home automation and wipe the floor with Echo and its ilk (yeah basically resurrect Mycroft but add whole-house automation to it).
Each of those are half-baked, but it also took me 7 minutes to come up with them, and they seem more in line with what Ollama tries to represent than a pure cloud play using low-power models.
This is the play. Its only a matter of time till they do it. Investors will want their returns
And are they VC funded? Are they funded by Y-combinator or anything else..
I just thought it was a project by someone to write something similar to docker but for LLM's and that was its pitch for a really really long time I think
Gotta pay those VC juicy returns somehow.
Ollama is beloved by people who know how to write 5 lines of python and bash to do API calls, but can't possibly improve the actual app.
Qwen3 235b
Deepseek 3.1 671b (thinking and non thinking)
Llama 3.1 405b
GPT OSS 120b
Those are hardly "small inferior models".
What is really cool is that you can set Codex up to use Ollama's API and then have it run tools on different models.
I was thinking about trying ChatGPT Pro, but I seem to have completely missed that they bumped the price from $100 to $200. It was $100 just a while ago, right? Before GPT-5, I assume.
Like I had Codex + gpt-5-codex (20€ tier) build me a network connectivity monitor for my very specific use case.
It worked, but had some really weird choices. Gave it to Claude Code (20€ tier again) and it immediately found a few issues and simplifications.
Here's a good example. For summarization of a page of content. Content is maybe pulled down by an agentic crawler, so using a local model to summarize is great. It's fast, doesn't cost anything (or much) and I can run it without guardrails as it doesn't represent a cost risk if it ran out of control.
1. Access to specific large open models (Qwen3 235b, Deepseek 3.1 671b, Llama 3.1 405b, GPT OSS 120b)
2. Having them available via the Ollama API LOCALLY
3. The ability to set up Codex to use Ollama's API for running tools on different models
I mean, really, nothing else is even close at this point and I would rather eat a bug than use Microsoft's cloud.
At some level it's also more of a principle that I could run something locally that matters rather than actually doing it. I don't want to become dependent on technology that someone could take away from me.