Read the paper the author talks about yourself instead (https://arxiv.org/abs/2511.07885), and also, contrary to what the author says in the article, do not do investments based on single papers made from academic studies, regardless of how much money this guy tells you you can make.
I run models on my desktop and access them from a phone app on the go. Wireguard tunnel.
Responds fast and lets me kick off tasks or workflows via text or voice.
Maybe I could use codex or whatever to set that all up nowadays, but it was such a pain doing the initial setup and getting it all running and finding sources of media yadda yadda like a decade ago when I last tried lol
Very much a personality and/or lifestyle thing, but I've also been building infra for decades. I mean it was two container startups and a phone app download (10min). Not exactly difficult. If I didn't like it I certainly wouldn't be in tech!
> pay someone
I'd do that with something like 3 story roof work or foundation work with a contractor, but tech? If I want the outcome I will be happy with, often have to do it myself.
I was into home automation and self hosting media and some other stuff for a while. It’s the kind of thing I’d only do again if i was equally or more interested in the process than the outcome.
We're looking for different outcomes.
I am also looking to actively cut my SaaS / services cost to $0 or near $0. I don't want another company touching my fucking data or prompts.
There will be some “rising star” businessperson who somehow got an amazing deal on land and energy and just likes to spend 1/2 their time in Shenzhen or something, but the datacenter is in Vietnam or Singapore or Thailand or wherever. They might even have a US branch, to make everyone feel better!
It will collapse the market, and everyone will realize that data centers are actually worth LESS than Toyota - they’ll look more like Tulips, suddenly.
In my opinion, of course.
I'm absolutely going to run this localy if I can. The only reason I'm not doing it is the fact I'm literally priced out.
The compute capacity, however, is here to stay. There will be a seemingly endless line of people queued up to buy it at pennies on the dollar if this whole thing blows up financially. And they will use it!
Just out of curiosity: Use it for what?
The demand is there and the market will find pricing that works.
Or they'd be part of the customer base for leasing compute from someone else who bought the rack.
Although there is such a thing as hobbyists with a ton of money who want to experiment with large models on enterprise grade hardware. Probably not a lot of those, though, of course.
I'd disagree here. I see two avenues for an efficiency multiple, albeit a single-digit multiple:
* Client aggregation allows a hyperscaler to average out demand spikes from uncorrelated clients, reducing the peak:average demand ratio and allowing better budgeting of compute.
* Dynamic batching allows typical requests to run in batches of more-than-1 and/or overlap, offering better internal compute utilization ratios (e.g. interleaving output and input streams). The small limit of on-device LLMs will run with batch sizes of one with strong memory bandwidth bottlenecks.
For an example of these factors in action, see the API cost differential between batch, standard, and 'fast' processing. OpenAI prices these tiers at a 1:2:4 ratio.
Go get a job a hyperscaler, they want to smoke what you are smoking.
I am not talking about diurnal cloud workloads, I am talking about the native efficiencies of hyperscalers vs on-prem. They have no magic and they all think they are going to make up their business overheads in exorbitant saas pricing.
- buy in bulk, for lower prices and access to better hardware through big contracts
- build in bulk (i.e. spread out software improvements over a lot of data centres and customers)
- offer additional services such as edge caching and multi-region data redundancy
That doesn't mean that hyperscalers don't also do things wrong, but it seems very odd to pretend there's nothing to them.
There may be a great deal of advantage to be able to run a small language model on something I have in my hand, disconnected.
Although mainframes have their use, the pendulum of centralized to distributed has gone back and forth and there are benefits to be gleaned from either model, sometimes at the same time.
The efficiency doesn't matter as long as it's good enough, because you already have the computer.
The privacy cost of sending everything to a third party is huge. Running locally fully resolves that, so it is the obvious choice if only cost can be managed.
Maybe if you combine an SLM with a database (as a tool) then it could work, but someone should first prove that.
1. Training a large model with lots of information, then stripping the "useless" information from that model to obtain a small model => nobody has shown this.
2. Training a small model, letting it use a database tool so it scores the same as a large model without database => nobody has shown this.
To go very small (thousands of rules) so that the reasoner can be understood by humans and proven sound - might be computationally quite difficult.
But seemingly models good at programming for example, would get worse at programming if you removed everything not-programming. Train a model solely on syntax, and it'll be worse than a general purpose LLM on syntax, in general at least.
That might be because emergence of capabilities to reason about programs requires abstractions (such as fuzzy and modal logic) that are rarely present in software sources. That doesn't mean the reasoning model itself has to be large; neither does it have to emerge from the ML training on large language corpus, we might construct it by different means.