The rest is capital intensive, and the price will approach the cost of production over time.
Thinking this is a profitable endeavor is equivalent to claiming coal plants have good margins because boilers are expensive.
The rest is capital intensive, and the price will approach the cost of production over time.
Thinking this is a profitable endeavor is equivalent to claiming coal plants have good margins because boilers are expensive.
What moat? You answered yourself: "capital intensive"
But, history says the supercomputer of today will fit in your pocket in a few years.
They've bought up all the RAM and GPUs, which pushes the capital requirements upward for everyone else. But, they can't corner the market forever, there are too many competing interests. AMD and Intel keep making new GPUs and APUs. The memory makers can't just sell to only AI companies forever, if they do Chinese manufacturers will move in and eventually eat them from below (as has happened many times before).
They have a moat today, and it's just that it's really expensive to train and host frontier models, especially at commercial scale. It used to be there was also some secret sauce to making it fast and efficient. But, secret sauce is being published daily by all sorts of researchers, folks are figuring out how to do more with less and it often finds its way into llama.cpp or vLLM or SGLang within days or weeks.
I don't think this will be true in the same time span anymore. Each miniaturization is costing more and more money.
Perhaps they'll come up with exotic fundamental improvements, but I don't think the rate of improvement of compute/watt will match the previous decades.
That said, I recently replaced my five year old self-built PC (with a top-of-the-line desktop CPU, chipset, memory, and GPU of the time) with a new everything-the-best build, and while it's clear we're not keeping up with Moore's Law anymore, it's still 4-5 times faster for compute-intensive stuff, especially parallelizable tasks. We're still getting faster/cheaper. So, the time scale is maybe ten years rather than five.
As that transition happens, hardware evolves from general purpose (because nobody knows what's needed and hardware design is slow) to fixed function high performance (once requirements are better defined).
GPUs (and TPUs) are a weird middle-ground here, as they're already fairly specialized, but I wouldn't bet against next gen AI inference-optimized hardware architectures dominating that use case in ~10 years if the pace of AI arch tweaking slows.
The efficiency/power/cost gains from fixed function optimization are always too great, and the only thing that holds that approach back is rapidly mutating requirements.
Drop the power requirements 1000 fold, and yea you will be able to make your own SOTA model on the cheap. The problem is the person that has a few exaflops of power will still leave you in the dust in the intelligence explosion that would happen after an event like this.
If training models gets way cheaper, I would expect the diminishing returns to get steeper too.
intelligence may be different. If we look at biological brains - do we get diminishing returns or completely opposite scaling law when we compare our brain against say gorilla's ?
Architecture / biological structure matters more.
I’d expect weight and wattage to be proportional for animals, at least.
A related argument is speed of intelligence vs capability at that speed. You can think of a three way trade off between latency, cost, and capability that is unlikely to be linear in any dimension and that changes in steps as technology or biology evolves.
Ultimately relating to the properties of the computing substrate and almost certainly bounded by some kind of thermodynamic limits that present systems do not approach.
I’d give that over 50% odds of happening in the next few years.
Unless we invest heavily in research and find new way to do chips. But I think there's not enough motivation and money to do that.
Apple is talking about 17.5 FP16 TFlop/s on the iphone 17 neural engine. So 20 years later we are still nowhere near, not even at reduced precision.
You can get an SoC that does 126 TOPs (strix halo) in tablet form factor, which is a factor of two. (I’ll count them as equivalent ops, since software couldn’t low precision floating point back then). So, not quite “pocket”, but probably “purse” and certainly backpack.
Depends on your world view, they might or might not come up with something better. but I guess we can agree nothing with stop them from _trying_?
China will certainly compete though.
“We’ve failed to deliver on 5 years of promises after wasting billions of dollars… sorry” is a death knell. However, “We’ve decided to not deliver on 5 years of promises after wasting billions of dollars… for safety… but keep those investments rolling in” is like crack to the true believers.
Is there an endgame where even this is considered overly complex? Instead of starving the competition by buying up all the compute, why not just buy up… all the money!? Hoover up as much investment capital as possible so that your competitors can’t get funding.
every major tech company literally have deal,ownership,alliance etc
they literally not gonna gobble up entirely to trigger anti-trust case
Anthropic / OpenAI / SpaceX going public makes it easier for capital to both flow to and away from them.
I guess that’s one way to try to make capital finite.
That is such a crazy way to start a response to someone trying to argue with you. I should try this. That's amazing. I know you didn't mean it as a trick, at least I'm pretty sure you meant it sincerely, but I'm just struck by the power of it to defuse and redirect the conversation. And this was a very low-grade example, but I could imagine this being useful in much more heated contexts.
It's a component of a few psych frameworks around improving interpersonal conflict. Ref: https://hartsteinpsychological.com/the-power-of-active-liste...
Short template form is "What I think I heard you say is (repeat their words as exactly as possible)? Did I get that right?"
I was nitpicking the use of the word “moat”. For it to be a moat, it’d need to be more expensive to traverse than to build.
Instead, the big AI firms are trying to create a monopoly on capital in an area where real costs are dropping 90% year over year.
hmm nooo ??, physic says otherwise
That was Moore's law saying that. And it seems Moore's law slowed down quite a bit for now.
(less facetiously, I think they mean "5 to 50")
To build a working prototype, sure. To operate at production scale, definitely not. The same rule would apply to WhatsApp and many other world-scale products. Turns out that, the moment you need to monetize these machines, your O(10) stops working.
I do not know why every Chinese model fan thinks that people that aren't impressed by them simply don't use them.
It's quite obvious that when you dont try to do something particularly complex there will be literally no difference between GPT, Claude, Gemini and Deepseek.
Fot many things I'm doing in gamedev Gemini 2.5 Pro was already good enough even though it released more than year ago.
Once you pass certain threshold it's just enough.
Openweight models turned a corner around kimi 2.6, deepseek v4 pro/flash, hy3 and mimo 2.5 pro. Similar to how closed LLMs turned a corner around gpt 5.2 and opus 4.5.
While they remain a step behind closed frontier models, for real world tasks ranging across functional reactive programming, distributed systems, mathematical modeling, to-the-millisecond highly optimized spatial data-structures, complex compute shaders and shader effects and non-trivial systems involving parser combinators and algebraic effect systems, I can say that open models have very recently gone from useless to productive. For my work, mimo v2.5 pro is hands down better than sonnet 4.6.
I'm not working on the frontier problems, I don't need god-in-a-box for $600 per month.
and almost nobody is working on frontier problems. they just want frontier intelligence to solve their given problems in a superior manner.
you're minimizing and exaggerating all of the wrong things. cope more i guess - more compute for us!
Anthropic can stretch the moat all they want, but in the department of trust, they put a final nail in their coffin today. Anthropic is pure evil at this point.
I don't know. If my ISP started MITMing my traffic so that they could silently rewrite packets, and/or deleting files on my computer because they thought me sharing wireless AP with my SO was me trying to compete with them, I'd call them evil.
I believe they tried something similar to the first one a few years ago in the US, and I remember people called that evil to the point where tech giants shut down their websites in protest.
> gee i wonder what would happen if they ever actually achieved SOTA? They would clamp down on that so fast Dadio's dradel would spin
Cool. Let them "achieve SOTA" and close down the models. Let the pendulum swing the other way.
You seem to not understand what China's goal is here. They want the AI bubble to burst and take your 401ks with it. And OAI/ANTs decisions are driving you towards that cliff.
Open source in quotes because they are not open source and not even close to open source.
This is just another incremental improvement, rushed out to boost the ipo, AI has the capacity to aid an engineer but this minor bump in performance will have essentially zero impact on the productivity of an engineer working on real world solutions when compared with any other major model.
We are trending towards asymtotic and it can't happen fast enough, that's when the true cost of this will become evident.