We'd originally set this up to be able to locally run larger models, in the 300-400 billion parameter range, but the rate of improvement of open-weight models has been so fast that, coupled with extensive custom skill creation, we're now getting similar results out of Qwen3.8 27B to what we were getting out of Qwen 3.5 297B when we started out with the project, which frees enough memory to allow 20-25 users to have 256K context concurrently. Both the hardware, the software, and the models are improving at an accelerating rate.
Investing in data centers to support SaaS LLM providers today feels a bit like investing in mainframes and minicomputers in the late '70s, with a massive paradigm shift lurking right around the corner.
Actually, it's probably already closer to the early '80s, given that purpose-built local AI workstations are already available at price points lower than the inflation-adjusted initial price of the original IBM PC.
While I cautiously circle the AI programming concept, one thing close to my mind has definitely been whether the ultimately cloud-tied nature of it is what makes me much more stand offish to it.
I also use pi.dev as my agent harness, and ollama as my inference engine. Though I've been considering switching out the latter for something else, since the main advantage of ollama is the ease of switching models, which I don't really do often anymore.
Actually I think the AI revolution is going to have a worse outcome. It will kill off the ladder to a better life by eliminating entry level white collar work. This won’t be valuable enough to generate a UBI, and there won’t be some whale company to get it from.
Current day mid-level and up white collar workers will just adopt AI into their work, and likely benefit. In 10-15 years there won’t be replacements but maybe then the AI can run these companies anyways. In the meantime companies will adjust their goods and services to target the increasingly wealthy but shrinking upper class or the growing lower class.
And even if you do inference at home you are not training the models. Moreover, most people certainly are not doing inference locally.
What? (Edit: I think you misread "in support of their primary concern, being AI" as "in support of AI"? People who complain about pollution are usually mostly concerned about AI, and they use arguments of pollution to strengthen their argument against AI related things.)
> And even if you do inference at home you are not training the models.
No, and training does use a large amount of energy on a large amount of hardware. But:
- That sort of workload doesn't really require many distributed data centers, only a few powerful ones. I believe training is also getting cheaper for the achieving higher levels of capabilities, but I don't think the efficiency advancement has been as dramatic as it has been for inference.
- I believe we are really hitting a wall of diminishing returns, especially at the high end of large models. There is lower demand for training, because models are good enough to have a reasonably long shelf life at this point. Year+ old models that were SOTA in their time are still useful today. The demand for training is going down.
> Moreover, most people certainly are not doing inference locally.
Maybe not, but I think most people who go out of their way to use LLMs because they find them useful for their work actually are. Though most inference is probably from people who do it accidentally (as a part of a search result or something) or students who don't have the resources to do it themselves. But basically every software engineer that I personally know that uses AI for programming is either running their inference on their own hardware, or is talking about building a rig to do it.
See: they were anti-AI the whole time, and pollution was just an attempt at supporting their anti-AI argument.
Electric cars and ICE cars both use more watts than AI. If you're concern is energy consumption, cars are worse. If your concern is pollution, cars are worse. Do you care about pollution and energy consumption, or do you honestly just hate AI? It's okay if you just hate AI, but be honest about it.
Because as you obviously can tell, it's totally ridiculous to suggest people give up their primary mode of transport so people can have data centers instead.
The people who grow your food aren't going to be doing that on a bus. The people who build and run data centers need cars to do it.
I'm not sure what you think farmers need personal cars for? Tractors and trains seem like they'd do literally anything you could imagine a farmer "needing" a car for?
Tractors are crazy slow, and it's obviously stupidly inefficient to have trains running into every rural area just for the off chance a farmer needs to go pickup some hay, which by the way would usually go onto the back of a lorry.