The memory situation currently is a mess. But ... is it bad enough to get us to _change_ how we build?
The memory situation currently is a mess. But ... is it bad enough to get us to _change_ how we build?
Wrt. the memory efficiency - considering that AI is going to be the main workload by far and thus will vacuum in majority of the available optimization resources - be it humans or coding AI - i think it will be even less chances that any meaningful resources will be spent on optimizing on the fringe, ie. the users' OS/browser/etc.
But in the real world, haven't we got to see whether all this investment will pay off? What happens if it turns out people actually don't want to pay for this?
The companies go bankrupt their stuff is sold for pennies & investors loose their money.
The now super cheap stuff might get reused for something else, that is actually profitable. Or will just get dumped, like many unprofitable or never finished lines from the railway mania days.
Once everyone has to pay enough for the usage to just break even, it's going to be a shitshow. And what's worse, they'll have to make up for all the debt they're accumulating now, so even pricier.
Basically, individuals will only stay at the current massive discount. Corporations that are paying token costs right now are dialing back as they realize it's not worth the expense.
We are going to have an AI crash, and the only thing left are going to be the efficient on-device or local models, which are already starting to bridge the gap.
It makes zero sense to rebase your mental model of the tradeoff between RAM versus developer time because of a temporary shortage.
In a few years time you’ll look silly for having wasted developer time on lowering BF memory consumption, while your built new features and waited for supply and demand to rebalance.
2. Say it's infeasible to have enough compute locally to do basic tasks
3. Force usage of centrally controlled, subscription-based cloud software.
4. Capitalism win???!
It claims fp8 and fp16 has had higher throughput in the cloud since 2023.
fp32 and general computing still has more compute than personal computers.
And that memory is at the crossover point, but at this point only 10% of memory is going to personal computers, so the next few years will radically change the pc to cloud memory ratios.
https://claude.ai/public/artifacts/8733239a-328b-4aaa-8d76-0...
Anyway, all of this made me a (tiny bit) sad. It is the gradual end of general purpose computing you own, and the start of everything just being cloud compute that someone else owns and you rent.
A few thoughts...
(a) There've never been more general purpose computers than there are right now. Spec progress may stall for consumer products due to the price increases, but the demand for workstations is not going to go down that much; they're critical items to get work done. Many are hoping this means a return to memory efficiency; the 8GB Macbook Neo is a good sign that this is something Apple believes to some extent also.
(b) We've arguably been doing this since the 90s. A website (especially in the days of the heavy backend and more passive client) is cloud compute that someone else owns. You rent it by paying them in the form of looking at ads...
(c) There are a lot of companies and people working on making sure you have every chance to run your AI locally. It's surprisingly good now and keeps getting better. At some point we'll see the same oscillation that we always do from thin client to fat client as the latency advantage of working at the edge starts to win out. How long until I have a laptop with a Talaas chip in it running a basic local model at 16000 tok/sec that is just good enough to use A2A to reach a more powerful agent running elsewhere if I need it?
Most compute loads are bursty. Most PCs and phones sit idle almost all the time. It's an inefficient use of capital. In the cloud you can be idle and then suddenly command 1000X more power than any affordable PC could ever have... for 10 seconds... then be idle again. Those kinds of work loads aren't possible to realize on a PC without massively overbuying hardware.
I'd like to see more research on secure private cloud computing. Homomorphic encryption has gotten better but is still very slow, but there's other approaches that can help quite a bit but don't require much of a performance hit. Look into secure multiparty computation, reversible anonymization, etc.
I don't think personal computing will die, but I think it'll go back to being a hobbyist and pro thing like it was from its birth in the 1970s up until the middle 1990s. In the 1980s the average person did not own a computer, or if they did it was just to play some games. The average person was not loading BASIC programs off bootleg tapes and calling BBSes.
BTW I've been thinking a lot lately about how AI could save personal computing by allowing us to create an off-ramp from the present rats nest of complexity to something saner and simpler. Just thoughts right now but it's interesting.
At the same time you are at the mercy of whoever runs the cloud, they can spy on your data or even manipulate your workloads, they can renegotiate the terms of use at any time - including price hikes or booting your from their cloud at any time.
It's sort of like what I say about self driving cars. Are they better than good human drivers? Not yet. Are they better than bad, drunk, or texting drivers? Yes.
LLMs are already better than bad coders and honestly most coders are bad coders.
Being able to automatically generate large quantities of code that is worse than a good human programmer can produce is not helpful in this world you're envisioning.
It's the price you pay for having an application delivery platform that also needs to be hugely sandboxed running one of the most difficult languages to run efficiently in such large amounts as so the plaintext minified versions are a dozen megabytes in size.
I don't envy browsers even slightly, peoples expectations of them are ludicrous.
The major problem is that nobody is genuinely looking at using more efficient languages, lessening the amount of required inefficient language use or gasp producing native applications. (which, is definitely an order of magnitude harder to produce, I'll grant).
Compared to what though? It's not nearly as horrible as it used to be, and you can do it fairly neatly today, and end up with something much better. Requires you to actually have skill and knowledge about there being other alternatives to Electron though, which seems like knowledge hard to come by these days.