We're in the brute force phase of AI – once it ends, demand for GPUs will too
theregister.com
theregister.com
> In economics, the Jevons paradox occurs when technological progress increases the efficiency with which a resource is used (reducing the amount necessary for any one use), but the falling cost of use induces increases in demand enough that resource use is increased, rather than reduced.
> "I think there is a world market for maybe five computers." Thomas Watson, president of IBM, 1943
In 1943 there were probably only a few governments capable of paying the extreme cost, and having a valid use case, for a "computer" in 1943.
If you asked him what he thought the market was for a device that cost a month's wages, and could connect to anywhere on earth, do infinite math, remember anything perfectly, and entertain the whole household was, he probably would have had a different answer.
It was already clear then that 64K on its own was not enough, and the segment register twiddling to touch other 64K windows was seriously limiting.
The model was patchy from day 1, and would need serious effort to be future proof. Even if computers with 640K installed were rare, the design limitation of 640K possible RAM was clearly not enough.
yes companies used to hire people to just add columns of numbers.
Given all computers in that era are literal super computers of their times, that quote seems to be mostly right even till date.
only include what you need, and it minifies to zero byes!!
Writing good software is as hard as it has ever been. IDEs don’t help you with anything that makes proper software difficult. The only thing that has changed is that users have been conditioned to accept shit.
Software development hasn't even gotten a magnitude easier, let alone infinitely. Every improvement has addressed accidental aspects not the essence.
Unfortunately, market pays mostly for CRUDs with various styles of APIs on top.
People have no respect for experience and skills anymore, it's all about the profit and any bootcamp monkey can make money.
I don't agree here. It was way simpler in the 90s. The programmer experience probably peaked around the transition from TUI:s to win32 where you could do either. Different screen resolutions is probably what made programming gui:s suck. And all the churn of Microsoft and Oracle frameworks didn't help.
Nowadays the overhead of making an app that passes procurement is insurmountable. And consumers seem to not buy apps at full price anymore.
Techs advanced enough that you could make a facade of doom by yourself over the weekend (game design aside). Game devs can be so incredibly productive. But instead, demand sweeled to insatiable heights, as well as dev teams. Doom 2016 probably had hundreds of staff involved and 4-5 years of dev time.
That's basically what happens when tech aims to be bigger and better instead of how to optimize each dev themselves to be individually more productive and keep the project lean (company structures aren't helping either).
Comparing Doom 1993 vs 2016 makes no sense in that context. There's no scenario where you make gigantic scale Doom-style game worlds circa 2016 with even 20-30 people, much less 5-10. Art asset creation alone for 2016 requires far more staff than the original Doom. Optimizing each dev wouldn't begin to scratch the surface in terms of what you need to get to Doom 2016 if we're talking a dozen people or less. You'd need extraordinarily advanced AI agents creating for you, and the year would need to be more like ~2040-2050. The tech underlying a game like Doom 2016 is a modest part of the labor scale problem.
What do you think comes to mind when you hear "optimizing each dev"? My suggestion was for each dev to work wider, not deeper. There'd inevitably be a hit in raw fidelity, but that's part of the point I wanted to make.
Of course, no amount of automation even with AI will make up for millions of hand crafted man hours. But my big discrepancy with modern game dev is: for how much business want so care about costs and skimping on labor, we go far, far past middling returns in order to deliver these AAA games. I'd definitely wager that you could preserve 80% the quality of Doom 2016 with 20% of the staff (and pay that staff better, not just pocket the 3x cost reduction) and it would still look top of the line. Even then I question the middling returns.
There are good and insidious reasons why it's so rare, but I was really hoping these better tooling over time would produce more indie studios able to work at around a AA level of production. 5-10 people making games that really aren't that vast a gap from AAA presentation to the common consumer. Instead that sector is seemingly shrinking. It's a problem I at least want to try and chip away at in my career.
Specifically, our target environments distance us too greatly from the problem(s) being solved on the main/"happy" path.
Is there a WYSIWYG GUI builder for WWW/Android/iOS as good as Delphi was for win32 or Interface Builder for NeXTSTEP?
If not, then I have to waste a ton of time working on UI boilerplate code instead of the important & fun stuff.
A.k.a. The problems have become harder (stricter requirements, more ambitious objectives), which is entirely different than the tooling having become worse.
The tools themselves are better in so many ways. They just haven't caught up to what we're trying to do with them. Myself, and others, remember fondly when they had.
What you're citing is an example of the Jevons paradox, not a counterexample. Development got easier, so demand for (more depth and variety of) development rose.
We are now getting to a point where IDEs are as good as the ones we had in the 90s.
> performance constraints
Evened out by higher fidelity and less efficient programming languages and paradigms.
> deployment
More robust, perhaps, but also much more complex.
> version control
An improvement in some respects, a regression in others.
> more productive
Hard constraints make people productive. Being productive is about what not to do, impossibilities make for easy decisions.
https://bangaloremirror.indiatimes.com/opinion/others/easyno...
The idea is more lanes on the highway means more traffic means slower (even though you added a lane!). What goes off of this is what's called "Palin's Corrolary", which is to make traffic faster it's best to have fewer lanes. Politicians apply various techniques for this such as perpetual construction or allocating vast swaths of asphalt for bicycles, to make the traffic flow faster.
So it does make sense that in fact slower chips will make AI faster, and punch cards will make software development faster, as the inverse of these proposed trends.
This is worth it for the mental image of heaping them into a boiler fire by the shovel load alone.
But it’s not really a theory so much as an established fact that the only way to reduce traffic is to have viable alternatives to driving.
Des Moines needs Huston levels of freeway to meet current demand. Which is why nobody is willing to pay for it.
But, of course, the people saying "better streets would only make people drive more" are stupid. It's just that the people insisting one can do everything by car are also stupid.
of course that is even more expensive.
that's right, better streets and more parking area will deter car use.
what people want is to do things. They want better streets and more parking because they think it will help them - induced demand proves they are right. If you don't like it then work on a real answer, great transit for example (not the transit for people with 5 DWIs that we are trying to punish - which is what all most people see, no wonder they don't want it.)
> better not focus on butting down the us.
You wouldn't consider casually popping into a Paris cafe if the wait was always 4 hours or you had to have a reservation months in advance, which would be the case if travel time was a non-factor for everyone
Very large R&D expenditures for the next iterations of the models at the leading edge (the "fabs" of the world), everything downstream getting much cheaper and better with demand increasing as a result.
Like a world where Claude Opus 3.5 is incredibly expensive to train and run, but also results in a Claude Haiku that's on net better than the Opus of the prior generation, occurring every cycle.
My colleague introduced me to this idea. He had been studying ways to increase computing efficiency out of concern for the environment. Making programs more efficient would reduce energy consumption, right?
His advisor introduced him to Jevons paradox and he realized such efforts could have the exact opposite effect. So he dropped that research entirely. If you're worried about energy consumption, you need to make energy production more green, not machines more efficient.
Making data centers more efficient will probably cause us to build more data centers and use more power overall, not less.
I think it was literally lack of imagination. We were like "well, I automated most of the paper pushing we used to do in the office, guess my job is done!" and this occupied 0.001% of a computer's time. We invented all sorts of ways for people to only pay for that tiny slice of active time (serverless, async web frameworks, etc).
Now we're in an era where we can actually use the computers we've built. I don't think we're going back
I have a stack of 1080ti and Titan V GPUs that are testament to this. :-) (which, admittedly, I should sell)
Then take a Titan V at 14.9 TFLOPs (32 bit) at about 250W, for 59,000 MFLOP/Watt.
There's almost no conceivable world in which it's worth running the P5. It literally consumes ten thousand times as much power per unit compute as a not-quite modern GPU.
For compute to "always" be more valuable than the cost of the power it consumes, the value of that compute would have to be infinite. We have no such application. And I suspect we're unlikely to. :-)
Once the compute heavy pathways have established, I'd wager the next round of automation can utilise these established paths to lock-in on an answer rather than throwing more cycles at the problem.
Power consumption and heat dissipation, mostly. Amusingly, heat even turns into a performance thing itself, since you can get better performance in bursts than sustained.
For sure, once the LLM hype diminishes to more practical scale we'll figure something else to burn cycles on. And I'm not predicting demand for GPUs to die out all of a sudden, but my money isn't on the nVidias of this world.
It haven't even started. There are so many places it can be used, expect sci-fi in the next few years. There is no way back. The only thing that may change is the AI technology under the hood. I mean LLM isn't the goal, it's only a tool. Something else may replace or extend it. Many people are actively working on this. GPUs aren't the goal either, there must be more efficient way. But still they are very good at numbers crunching and that is needed for video processing.
i want inefficient so that if I feel like a large calculation I have one ready at hand.
It may not be obvious in the same ways as engines, but computation consumes a resources like power and attention and outputs waste like heat and fried circuits. These resources and wastes interact with other systems than just computers and data centers and so efficiency and necessity needs to be considered in a bigger picture than just "let's use all the transistors all the time and see what happens!"
There is so many resources we keep available and do not fully utilize. I don't see why computing devices should be different...
- [X] Text
- [X] Images
- [X] Audio
- [ ] Videos (in progress)
- [ ] 3D Meshes and Textures (in progress)
- [ ] Genetics (in progress)
- [ ] Physics Simulation (in progress)
- [ ] Mathematics
- [ ] Logic and Algorithms aka Planning and Optimization
- [ ] Reasoning
- [ ] Emotion
- [ ] Consciousness
We still have a lot of data to crunch but it's not nearly enough so we're also going to have to collect and generate a lot more of it. Some of these items require data that we don't even know how to collect yet. Barring some kind of disastrous event, draconian regulation, or politically/culturally motivated demonization of ML I don't see GPU demand dropping any time soon.Shedding some hindsight on earlier extrapolations — The billions pored into the metaverse or self-driving didn't yield the results we expected in the period we expected.
Whether we will really have cracked the physical world connections, of physics, genetics, etc that we can use it to make physical products, changes etc I am less sure. Many usecases like medicine require not just correctness, but also a degree of verifiability. It is being worked on a lot, with many promising results. But the just-scale-the-training data strategy seems less viable here, both because relevant data is less prevalent and may not give the level of correctness.
Two Reddit threads really highlight this.
- ~10 years ago: https://www.reddit.com/r/StableDiffusion/comments/y9zxj1/you...
- Today: https://www.reddit.com/r/StableDiffusion/comments/1f0b45f/fl...
The upgrade in throughput from GPT-4 to GPT-4o and GPT-4o Mini actually unlocked use cases for the startup I'm at.
People that think demand for GPU compute capacity is going to decrease are probably wrong in the same way that people who thought the demand for faster processors and more RAM would wane were wrong. We are just barely at the start of finding the use cases and how to eat those GPU cycles.
> The need for specialist hardware, he observed, is a sign of the "brute force" phase of AI, in which programming techniques are yet to be refined and powerful hardware is needed. "If you cannot find the elegant way of programming … it [the AI application] dies," he added.
The thing is that even if there is an elegant and efficient programmatic/algorithmic solution, having more and faster hardware only makes it better and pushes the limits even more.What makes you say that? I don’t really see a trend of AI generated content getting better, just more players in the space.
I think we’re at the peak of AI gen, I doubt we’ll see much improvement in quality (it’s already pretty good and it seems like all the low hanging fruit is gone), just more specialized models. Maybe some better tooling to give artists more control
having seen it grow more and more since 2016 when GANs started making fairly realistic human faces, this seems like the end goal already.
> I don’t really see a trend of AI generated content getting better
You see that second link as the endpoint? That there's nowhere to go from there? How about you can have a holodeck type experience with Apple Vision Pro? Literally generate any scenario you want? Download generated scenarios and customize it however you want in real time?Entire animation workflows changed from animating models to using voice and text to describe scenes and actions.
Lowering the barrier of digital film making to the same level and ease of use as photo editing apps today -- even easier.
You really think that the second link is the peak of gen AI? You really think that nothing else and no more major industry shifts are going to happen when gen AI gets cheaper, faster, algorithms get better, and hardware gets more powerful?
And this is coming from someone thoroughly bullish on AI!
The applications are obvious: film making, content creation, teaching, etc. This is in contrast to crypto which was/is quite abstract (as is money in the first place) and the metaverse which required investing hundreds of dollars in specialized hardware.
In contrast, our world is surrounded by visual content so the applications and utility of gen AI seems far more obvious for the layperson.
The only real accomplishments of LLMs were how good the proposed use-cases sound on paper under competent implementation, and a theoretical solution to unstructured data parsing that's still too heavy to be worth a tiny bump in performance.
Do you want to live in a future where all human thought has been replaced by its surface level reproductions, made by big tech stuffing copyrighted works into a GPU farm with near-zero human labor? We both know it won't benefit you and me, our role is merely transitory in bootstrapping their self-improvement under the guise of a paid product, nor had the relationship between us and these tools been in any shape collaborative in the first place.
This fantasy targets the owner class, which can finally dream of labor decoupled from the laborer, the work simply costing no more than the price of electricity, all without the demands for livable compensation or following best practice. Even if the LLMs gained above-human performance in all domains of knowledge shortly followed by institution of a universal basic income, their invention will still have only been a force of stagnation, learned intellectual helpless, and overconsumption.
> terrible; bland art...bad code
Ironically, all of this means that we're not at the apex and there's still a long ways to go both in terms of algorithms and the hardware to run them. > Do you want to live in a future where all human thought has been replaced by its surface level reproductions, made by big tech stuffing copyrighted works into a GPU farm with near-zero human labor?
Whether we want to or not, it's the apparent path that will unfold; there's no putting AI back into the box. The race is already on.Sure, in the way that technically this is a computable problem, but maybe not a simple one. Any exponential in the real world is a sigmoid and given all major AI labs, having spent years and incomprehensible sums of cache, have arrived at about GPT-4 performance, including OpenAI's latest release being a smaller model, should tell us something. Be the limiting factor corpus size, model parameters, or an architectural defect, we're clearly loosing momentum, at least until it's to be diagnosed and solved. Deferring to hypothetical futures without meeting the burden of evidence seems ill-advised, especially when it comes to incompetent use today.
> there's no putting AI back into the box
There's also no complete undoing of an oil spill, nor the practical possibility of unilateral nuclear disarmament. The weights are public, the corps is as well, as much as it had broken the open web to gather it, the architecture is known, ergo today's open-weight models are the baseline of capability for all future models. Still, seeing that LLMs form a natural monopoly given the required compute power to enter the field, we can enforce policies like mandatory statistical fingerprinting on all outputs of proprietary models beyond a certain size. LLM detection is also getting quite good; my hope is that adversarial fine-tuning will work out like Bayesian poisoning for email spammers, only giving the discriminator new stable patterns to look for. We as consumers do have power, however small and unconcentrated, to vote with our wallets, which a study posted here shown many are beginning to do [0]. The copyright question also gives quite a nice kill switch if our society decides this whole industry isn't that beneficial, removing the commercial incentive for training new models while providing some for detection.
There's great value from transformer networks, such as state-of-the-art speech recognition, that will make its way into consumer products and will be here to stay. As for the FADs like useless chatbots in every product, that may come under question.
Thank you.
Recording time use to cost a huge amount of money and then DAWs put that in the hands of basically everyone with a PC.
We got an infinite amount more of half finished demos that no one listened to. I would have a hard time saying with a straight face that music as a whole has got better. The sheer volume crowds out a lot the fringe from being worth the effort of creating.
AI Art is really a better example though. I just resubscribed to midjourney this weekend on the web. I could see the thousands of images I made on discord. I think some are cool but what use are they? Millions of synthetic photo realistic selfies that no one but the creator bothers to look at.
The arts ultimately need a network of people appreciating the art form or you have nothing. Just an infinite amount of board classroom notebook doodles that no one ever sees. It doesn't matter if it is Picasso doing bored doodling if there is no audience. You can't have all artists with no audience. Really good chance with no audience Picasso just does something else too.
I mean has the art world really been disrupted by midjourney? It is an absurd idea. What is most curious is how little disruption any of this seems to be having.
Now, anyone can go into TJ Maxx or Home Goods and pick up cheap printed artwork from China. Want a portrait? Snap a picture with your phone and print it out. No skill required.
It was once the case that if you wanted a bespoke sculpture, you either needed to have the skill or the wealth to pay an artist to create one. Now anyone can 3D print one or use a CNC machine or injection molding to create one.
Is there something inherently wrong with that? Does that mean that award winning photographers and acclaimed modern painters are degraded to the same level as anyone with a phone? Are world renowned sculptors and artists no longer a thing because of 3D printing and cheap access to injection molding?
> Really good chance with no audience Picasso just does something else too.
The people that want to be really good at their craft will continue to do so. The people that appreciate the effort, artistry, and skill will continue to do so. The presence of cheap mass produced wine doesn't degrade expensive wine; the presence of cheap mass produced whisky doesn't obviate the market for expensive small batch whisky. It's the opposite; in fact, it elevates it onto a pedestal.But more than that, the power of generative AI is to create an experience that otherwise doesn't exist because no game or visual experience can be tailored exactly to my tastes, preferences, and style. What gen AI promises is that every person can get exactly the experience that they are seeking by simply tweaking the input.
It’s because the people impressed by AI are impressed because we weren’t able to do it 5 years ago. It’s novel. It makes unskilled artists feel like they have skill. It makes for easy, specific, good looking images. They can get quick images that are more specific than ever before very very quickly.
But that isn’t what making art or doing graphic design is. Making the image is the easy part most of the time.
Same with code. Rarely is writing the code the hard part. Solving problems within the constraints of a system are.
Some models we still have to correct it for outputting Markdown code fences for JSON (we just brute force and string replace it).
We have to give it minute details or fine-tune it on a large enough sample set to get the results we want.
Even just LLMs right now still have a ways to go in terms of how good they are at working without needing precise instructions and micro-corrections and throughput.
You seem to be talking about a technical solution, but naming it after the problem.
Imagine YouTube except you can fully generate shorts. Realistic, 2D animated, 3D, whatever your imagination desires. Imagine how that changes storytelling and content creation.
More GPUs, please.
The ability to have a conversation with a character with a back story is definitely going to be an interesting addition in the bery near future.
With ML, its basically just matrix math. You can easily build ASIC if you know your Tensor sizes for inference, because the memory locations become static, which means your circuitry gets way simpler.
"Generative" is typically referred to as gen AI.
This is pretty standard nomenclature.
[0] https://en.wikipedia.org/wiki/Artificial_general_intelligenc...
To be clear, I agree that LLMs are not anywhere close to AGI and I don't think they ever will be (just a component). But that doesn't mean they aren't useful enough to chew up a lot of compute for the foreseeable future.
Sure, but the question is how much compute? What if we're reaching the asymptotic limit of scaling (but still with some post-training and dataset curation gains to be had), and existing datacenters go from being used for training to inference instead. How long before another data center needs to be built or updated with latest NVIDIA GPUs? Is this (LLM-based AI) just a GPU-upgrade market, or still one growing explosively ?
It's not clear what level of datacenter/GPU investment is needed for inference vs training - all the talk of massive GPU clusters seems to be about training needs, not inference.
As far as efficiency, presumably we'll eventually switch to smarter (brain-like) dataflow type designs with incremental learning, but for time being we're stuck synchronously pumping 100K contexts through hundreds of transformer layers, and trillion token pre-training runs. Brute force indeed!
Demand can't "not go anywhere". Demand has to continue to go through the roof for sales to stay where they're at.
For sales to continue increasing at the rate they were, demand would have to be insane.
That’s a big assumption.
> There has been criticism. Journalist Joel Hruska writing in ExtremeTech in 2020 said "there is no such thing as Huang's Law", calling it an "illusion" that rests on the gains made possible by Moore's law; and that it is too soon to determine a law exists.[9] The research nonprofit Epoch has found that, between 2006 and 2021, GPU price performance (in terms of FLOPS/$) has tended to double approximately every 2.5 years, much slower than predicted by Huang's law
Regardless, there is a ceiling to how fast hardware can be. Can’t just double forever.
There is a decent chance we ramp to 10+ trillion $$ of demand annually for inference here in the next decade or so. That will drive a lot of GPU sales.
Because consumption "demand" didn't "not go anywhere". It went through the roof.
My point is: you can't say that just because current usage is "here to stay" that sales at this level are "here to stay". Usage has to increase dramatically to keep sales flat.
So if you told my company they could run 2x the number of experiments for the same price they would do it because it doubles the number of hyperparameter settings we can try simultaneously. If that means fine tuning BERT in a day vs 2, that’s a huge increase in iteration velocity. It means we can work on more backlog projects simultaneously. It means engineers aren’t elbowing each other for GPU time (ok that is optimistic).
Machine learning is a trade off between model size (training cost), model run time (inference cost), and quality.
When some task is solved (e.g., hot word detection or speech to text), it becomes a commodity and some harder task becomes the priority.
I know that as soon as I can output 100req/s on the cheap on a llama-level model I will put it EVERYWHERE. And my clients too.
DMCA handling? Content flagging alerts? Fuzzy categorization? Natural UI for end user complex queries?
All LLM baby.
And much, much more.
- if a company, say, AMD, found a way to produce GPUs at a fraction (say, x=10%) of the price of NVDA, would that increase aggregate demand, or keep it about the same (substitution for nVidia)? Would the price difference be enough to incentivize creation of a CUDA- alternative ecosystem? If not, what does the fraction x need to be?
- Very reductively, GPUs seem to be universally needed because they are better at manipulating matrices than alternatives (eg, inverting a matrix, finding dot product, cosine similarity, etc). Are there alternative approaches that could come to market in the next 2-3 years that could be better, or better-per-$, than the current approach of just building bigger GPUs?
And even LLM's and photo generation open up a bajillion usecases that would have taken years of research and development before. But nobody focuses on these - instead, they focus on S&P 500 companies and how they haven't earned much with GenAI.
Because these companies are usually known as the peak of human creativity, imagination and are ready to jump on a new technology without much red tape in it.
Honestly, it's been like two-three years tops. Even talking to tech startup CEO's I don't get the feeling they remotely understand the technology or application, as 90% of things I've heard them say is "oooh we could make a chatbot!" or "let's replace developers with it - oh it can't one shot generate my whole codebase? pft that sucks".
If these folks don't know how to use it, surely Jim VP of Engineering #62 at ACME & CO that hasn't used any tech except ERP's for the last 10 years will have an idea how to.
It might sound condescending, but a lot, and I mean a loooooot of product managers, product heads, product owners, VP's, CEO's and such don't have a clue what they're doing or experience in the real world. One might be "oh but why are they CEO then", but hell, corporate incompetence is a real thing.
I've met dozens, if not hundreds of PMs/POs who were hired based on "oh they have organisational skills and aren't an autist when it comes to talking".
>"Or who exactly are you envisioning as possessing that supposed pinnacle of human creativity"
Creative people. It's okay to say some people aren't creative. Most ticket-dragging meeting-slacking Jira people aren't it.
I've met devs who had brilliant ideas on AI incorporation, only to be dismissed by the higher ups because they didn't understand it. I've talked to designers who's ideas could save hundreds of man-hours if implemented.
But do you really think an average VP knows that they can use their existing component library to train an MMM to output pseudo-code from screenshot and then translate that into their real, existing components? Or that an LLM could be hooked up to auto-correct the mistakes in the input of tens of thousands workers that they actually have people check, wasting human souls on what is basically input formatting? Or shit, that it could even translate John from Warehouse's data directly into monthly reports without him having to go around asking stuff and wasting everyone's time?
No, these people usually don't have a clue about the real-world process of actual work that goes on so, so of course they have a problem identifying leverage spots in it.
Rather, it's just that after multiple hackathons and several greenlit internal projects, there isn't a single one I wouldn't find an utter dogshit waste of money, time and effort. So while maybe we're just an unfortunate bunch who somehow all belong to that non-creative group, I struggle to imagine an actually good application of these technologies that are somehow all just being dismissed prematurely by those evil bean counters up there.
Google, Meta, et-all are working on their own AI chips but those chips will have to beat Nvidia's at Performance and TCO and Nvidia shows no signs of slowing down to let competitors catch up.
Etched, for example claims they have a chip reaching 500k tok/s in the works. Which is still far from the theoretical max with the current techology.
A similar scenario went with Bitcoin's GPU/FPGA/ASIC - the current ASICs are millions of times faster than GPUs.
TCO, yes. Raw performance, not necessarily. TCO will attack NVDA's margins. When Meta last wrote about their cluster it was presented as power equivalent to X NVDA chips. They are already bringing their own chips into the mix.
With AI, we’re constantly training different models, which can’t be trained using asics. If we ever get to the point where we no longer need to train new models, then yeah, it will go the way of bitcoin.
Wait what!? Did the Bitcoin hashing algorithm ever change?
And, as it happens, that's exactly what NVidia's done with the H100: https://developer.nvidia.com/blog/nvidia-hopper-architecture...
It still needs to be programmable though. Can't get away from that.
If you look at Apple and Google, they already have their own hardware for inference in their smartphones. They don't need NVidia for that.
https://www.cnbc.com/2024/07/29/apple-says-its-ai-models-wer...
The problem with this comparison is Bitcoin has basically just been SHA256 for 15 years and likely will continue to be for some time.
Transformers have been mostly dominant for at least several years but there are still other archs (CNN, RNN, etc) in various use-cases and we're already seeing nearly-fundamental changes in Transformers and "emerging" approaches like Mamba, RWKV, hybrids, etc. Transformers have shown remarkable versatility and adaptability (that's their whole thing) but it's already creaking and showing its age.
Startups building Transformer-specific silicon are playing a very risky game that is already somewhat problematic now and almost certainly won't end well.
AI is much newer, much more vast, and moving much more quickly. The ASIC design, tape out, manufacture, software ecosystem, actually getting to market, etc cycle is fundamentally too long and I suspect even the Transformer-specific silicon we see now will be viewed as a major blunder in the relatively near future:
"Oh yeah, remember those graveyard companies that did transformer silicon back in the first AI hype round?"
I cannot see how anything other than GPGPU, TPU, NPU, etc (or similar "generic" approaches) will have legs.
I don't know why people hire them. But I know for sure it's not because of the content.
Millions or billions of iterations to tune some algorithms to output some data which satisfies a success metric.
Really don’t see why he thinks brute force is going away. I mean human brains have billions of neurons too after all.
SOTA will remain at the edge of what compute can produce for a long time to come. SOTA is a moving frontier, and there will be demand for incrementally smarter models, because they save you time by making fewer mistakes.
No matter how much more efficient algorithmic innovations make ML model training, compute will make those algorithms smarter. It's a coefficient.