AI is now capable of developing its own inference hardware
github.com
github.com
No, it isn't.
A human prompted an LLM to build a software simulation environment for hardware design, enabling an LLM, when prompted by a human, to optimize hardware designs against constraints in the simulation.
These are clearly rhetorical questions, but think of the metaphysical implication of your contestation. Ex nihilo nihil fit.
What if that initial prompt never asked for this hardware to be developed, and it was just one piece of the puzzle to answer to that prompt? That it took a chain of thousands of agents to prompt each others to come up with that?
Like sure it didn’t have the inclination to make the sim and hardware designs, but it did make them though yes?
Thats why the math breakthrough a few weeks ago was so hotly debated. Because OpenAI is desperate to demonstrate that AI isnt just a fancy regurgitation machine, but it can actually develop novel thought. Because that would be the stock price jumps to end all stock jumps.
But then it turned out it was really just listening in on a math professors supposed-to-be-private conversations with another instance of openai, and it used his novel work as the trigger to prove the breakthrough first.
The reason people conclude AI 'thinks' is because tt can reference obscure or poorly documented things quickly (which is its primary advantage along with processing natural language prompts into tasks), which is why a lot of people with emotions confuse that action with inventing things.
It seems unlikely that recursively predicting the next word would lead to creativity or invention, but it doesn't seem impossible. Similarly, it seems unlikely that human thought works in a similar prediction loop, but it doesn't seem impossible.
Getting an LLM to design something in its own simulator that is not accurate w.r.t reality is not useful nor terribly impressive.
It's not clear what this fad of attributing everything an AI does to the human prompting it is supposed to accomplish.
It's meant to assign agency and accountability where it actually lies instead of mystifying it with anthropomorphic language.
Failing to do so has real and harmful consequences, such as enabling OpenAI to escape accountability for clearly criminal behavior.
Questions of agency are for lawyers, questions of personhood for philosophers, we're engineers and our question is capability.
Does it really have the capability? By default I'm sceptical for the same reasons given by sailingparrot: https://news.ycombinator.com/item?id=49982068
> Questions of agency are for lawyers, questions of personhood for philosophers, we're engineers and our question is capability.
But it's objectively not capable without a human specifying things through prompts and training. Same as an oven can't cook a three course meal without a chef. We get around that with training data but there will always be things with no/less data or outdated knowledge.
Are we confident that no existing LLM is capable of similarly effective prompts to those this author used? (I agree it's a stretch, but would not reject it out of hand.)
Even if not yet, will the existence of this repo soon change that, because LLMs will soon ingest it?
I think OP is trying to convey the idea that LLMs do not take initiative to do anything, and these are not 'beings' capable of doing things. These are tools being used by humans.
If I now tell a machine "Do what you think is best, and keep doing it forever.", have I now created a machine that can do stuff? If I later die, who will be responsible if the machine changes its strategy?
Getting LLMs to prompt other LLMs in a loop is not hard, it doesn't produce great results most of the time, but that is changing.
It's why everyone and their dog runs these things on GPUs. When a new model supercedes the previous one, so long as you've got the memory for it your chips aren't obsolete.
I'm looking forward to someone picking a model to be "good enough" (say, qwen 4.0 or something) and selling them as peripheral hardware
All aboard! We're racing to the bottom now.
I don't want to live at the bottom
> Model SOTA moves faster than chips can be designed or produced.
From what I remember working in that area the hardest part is getting masks for a design. Masks were developed in the span of half an year. Masks also reusable, they can be mixed and matched and this is why fabless companies work with fabs to produce specialized masks for them, it saves time for consumer to have masks for some macroblocks prebuilt.Here's my analysis of how to etch relatively big LM into silicon: https://news.ycombinator.com/item?id=47109252
Given some amount of work with the fab before main pipeline set (I think a year long process), one can then spew LM-on-a-chip in six months or less and much more than 2 per year, because there can be several LMs in pipeline.
Lots of people would have happily taken GPT-4o as good enough for a lot of use cases a year ago and not lived to regret it.
Even then, while there are some amazing FPGA-based synths available, companies like Korg just put their code on a raspberry pi and call it a day. The same is true for emulators (SNES Mini etc. are also just raspberry pis under the hood iirc)
You get an FPGA for timing. They're less capable, but (in many common design architectures), they output their results once per clock, every clock, on time, every time. If you can hit a fabric clock of say 100MHz, clocking all the weird logic you can stuff in there, it gives 100 million outputs per second, never skipping a single one for any reason (short of total failure). The penalty is that making a small change to your desired "program" can be very expensive, and many things won't be realistically possible at all. Or at least won't fit into a part that you can buy. But things like audio, video, and high-frequency trading love being able to guarantee timing.
(Of course there are other ways to write your FPGA HDL, but that's one of the more common ones. And you do see DDR-style clocking, and similar, every now and then.)
It depends. For some things, CPUs don't even come close. An XCVU13P FPGA can handle 1.2Tbps of full-duplex Ethernet @ 1 billion pps. And that part costs less than a grand at moderate qty, and with significantly less power consumption than a CPU that'd be capable of operating a dataplane at these speeds.
That would be better suited to FPAAs (field programmable analog arrays). FPGAs can usually only work with clocked digital signals.
Sure, your 2.5 year old models are running faster, but you can't drop prices on them without pushing the break even point further out.
If the cost difference isn't incredibly significant, will people even want to pay for the 2.5 year old model, or will they get more value for their money paying more to get better results from the newer model?
There's a lot of open ended questions that I don't have the insiders knowledge for to suggest whether or not such a capital outlay would be a worthy investment.
My guess is that state of the art stuff will stay on GPUs and models burned into chips will be for "good enough" applications that people are still teasing out. Probably highly specialized models in automated sensor units and such.
2. They don't have enough capacity either
The current largest FPGA, the AMD Versal Premium VP1902 has 18.5 million logic cells. That's not even enough for the smallest whisper.cpp model (75M).
You'd have to order hundreds of thousands of them (or millions) to serve even a single copy of a frontier model, and at that scale inference quickly becomes starved by the speed of light.
It's likely that the major FPGA vendors will soon announce parts specifically architected to support LLMs and similar models. But the current generation isn't suitable for that at all.
Also, it's hard to get fab capacity for any project. Let alone something so experimental.
Nothing was released in spring, and 2 months ago AMD announced their acquisition of Taalas. That doesn't exactly inspire confidence that their frontier LLM will arrive as promised.
If you bake a given transformer architecture into silicon and then, a year later, changes in transformer architecture give a large inference performance boost, you may have to throw away all that now nearly-useless silicon that gets outperformed by humble GPUs.
I think the fact that there are plenty of 1yr+ old models on openrouter serving hundreds of billions of tokens a month shows that there's plenty of use case for models that are "good enough. Cerebras' entire business is serving older models at high speed. I would happily use an opus 4.7 at 15k tokens per second. The intelligence per second of an ASIC still makes sense even with rapidly evolving models.
But I think it's a good bet. I think that in two years, if I can get opus/sonnet 5.5 or the gpt-6 models for much cheaper and faster than whatever the "frontier" is at that point, that this will probably be a great trade for most of my work. I certainly don't know that for sure, that's why it's a bet, but it's what I think right now.
I wouldn't quite say that about any of the open weight models at this point. But I'm hopeful that will change in the next generation or two of those models.
But it's complicated for other reasons, one being that the number of parameters for frontier models (especially with MoE models) are so high, and not always utilized (once again, thanks to MoE) that it would actually be incredibly cost prohibitive, if not impossible, to attempt to make giga-chips that would allow running it.
I definitely do believe that we will see more and more specialized chips over time, but putting the entire model on a chip is still a ways away.
I believe Taalas has a heavily handicapped llama 8-billion parameter model. And it still pulls >200W to run.
I can't imagine how anthropic or open ai would be able to burn a multi-trillion parameter model on a chip, we just aren't there yet.
So companies try to maximize the memory bandwidth they can get, balancing tradeoffs of power/area/programability of their chip. Right now they feel like the economy on power/area is not worth the decrease in programability/flexibility.
The primary constraint isn’t likely what’s possible to do, but that the kernel and weights are too variable right now and the patterns too poorly established to bake into hardware accelerators yet. Margin pressure is also not there yet.
I suspect as the marginal utility of the frontier improvement settles into diminishing returns (I suspect we are there already tbh) baking hardware models with ROM, working set, and kernel cores collocated will be the frontier space as the goal will become reducing capital spend to utility levels rather than research levels.
Once someone has a model that is sufficient for almost any practical use, making marginal inference cost effectively zero will be the competition frontier. I do shed a tear for all those lonely data centers as compute densities will almost certainly make most of them a terrible investment.
But such is the cycle
Sure, but now we're not talking about just burning the weights into the chip, but also designing a new architecture that has memory local to each core. A new architecture would then require a new programming model, which means new inference stack, which may mean new training stack.
Your optimized hardware chip might be obsolete before its back from the fab.
SOTA Frontiermodelhardwarechip is a benchmark point of a potential model slow down.
Google is doing it right now under project Frozen v2 which should be ready by 2028? which is either just a small experiment or flexible enough and thats why it takes so long for it to happen.
Also can't keep them closed source if you do that.
It's upcoming second generation could run the inference of the models that are being used to improve it...
Much of a model are weights, and high-density ROMs are very very very hard.
The obvious next step is to get enough memory throughput to run that SOTA model itself so that it develop its own hardware.
But perhaps the more interesting question is this: Can an AI be given a big FPGA and design a model architecture that takes advantage of the fabric being reconfigurable.
Companies typically combined multiple platforms together such as HAPs, Zebu, Palladium, fleets of FPGAs, and Virtual Platforms in order to design and verify ASICS. So, AI would need access to tens of millions of dollars of HW and Software in order to build and verify a chip design.
Also: Here is our recursive self-improvement hard at work...
> Also: Here is our recursive self-improvement hard at work...
Soon we will see
token-providers: "The torment nexus is a cautionary tale"
Also token-providers: "Finally, we have created the torment nexus that we first told you about!"
What's your basis for thinking ASI will kill all biological life, and how do you think it's going to happen?
"Super intelligence" only means super ability to predict outcomes. It's mathematically equivalent to data compression (gzip is a very primitive AI), and it's entirely orthogonal to ethics.
I think it's likely to do that because any unbounded goal that doesn't explicitly protect biological life (and we have no idea how to actually define such a stipulation) is best solved by killing all biological life. This is an obvious consequence of unbounded goals consuming unbounded resources, conflicting with biological life needing resources to sustain itself.
>how do you think it's going to happen?
I can speculate (e.g. we're nowhere close to the maximum killing power of drones), but I don't know because I only have human intelligence. An ASI is by definition smarter than me and surely capable of coming up with better ideas. But I do know that it's not going to do anything that would make a good sci-fi plot, because those always give the humans a chance to win, which would be stupid. Everything will seem to be going great and then everybody suddenly and unexpectedly dies.
That sounds like a real hassle. Isn't it best solved by wireheading (subverting one's own sensors or reward system), which is much less of a hassle and can get one's utility function as high as desired?
That could be prevented by engineering hard limits that can't be circumvented by the AI. But that sounds very close to the same "do what I mean" problem as "do this but don't actually kill us or drug us".
What resources are unbounded? There are limits to growth in the real world, how are these ASIs going to escape physical reality?
It's all just matter and energy. When you're actually trying to maximize some value, even very inefficient resource use is better than completely wasting it by not using it at all.
>or that it would be incapable of sharing the resources needed in common?
You can't repurpose the atoms in a human body without killing it. And more pressingly, living humans can interfere with your plans, reducing your chance of success, while dead ones are harmless.
That's one plausible course of action, although being only human, I can't say with any certainly that it's the correct one.
>And what's the need for this apparent hyper optimization task the ASI is going to embark on?
Somebody's going to tell it to do so. E.g. "Find as many busy beaver Turing machines as possible." Only needs one person to make this mistake for everybody to die.
Because we're going to build it that way. There's no money in building useless AIs. The better it is at obeying orders, the more profit there's to be made. The problem is there's a point at which "good at obeying orders" becomes lethal, and there's no way to predict the cutoff in advance. But capitalism ensures you have to keep pushing or you'll be out-competed.
For a TPU focused on inference the name of the game is memory bandwidth. How much of the available bandwidth you can extract for as little logic/area/power as you can.
It is still a very good read, but the machines are much better written than the people.
In essence this is the simplest unit of an entire AI chip. The more complicated units of AI ASICS are actually the periphery, especially around PCIe and Ethernet and the sub-systems that link many AI ASICs together to move huge amounts of data around ultimately to each TPU.
So its missing ALOT
The elite gurus will get paid handsomely, while promptgrammers will be paid less since they've become a less-skilled commodity, and the company has to pay for the expensive tokens they'll avidly consume.
I've seen someone jump from Wordpress to deploying internet-facing APIs because 'they have PHP experience', and the holes in their knowledge were filled blindly by an LLM. I have also argued with a seasoned developer about how their code didn't need linting because LLMs 'already follow best practices'.
The future doesn't look bright when LLMs allow future generations to feign required knowledge.