It's why everyone and their dog runs these things on GPUs. When a new model supercedes the previous one, so long as you've got the memory for it your chips aren't obsolete.
I'm looking forward to someone picking a model to be "good enough" (say, qwen 4.0 or something) and selling them as peripheral hardware
2. They don't have enough capacity either
The current largest FPGA, the AMD Versal Premium VP1902 has 18.5 million logic cells. That's not even enough for the smallest whisper.cpp model (75M).
You'd have to order hundreds of thousands of them (or millions) to serve even a single copy of a frontier model, and at that scale inference quickly becomes starved by the speed of light.
It's likely that the major FPGA vendors will soon announce parts specifically architected to support LLMs and similar models. But the current generation isn't suitable for that at all.
Never tried anything like that, though.
Additionally, it's not clear how well large mask roms scale. For instance the Nintendo switch cartridges were expected to be mask roms, but instead are Macronix's XtraRom technology, which is essentially a flash cell array made denser by removing the erase functionality. So basically the die gets manufactured with all bits at the same state, a late manufacturing step either empties or fills the floating gate of the bits you want different, and then it's treated as pretty close to a mask rom. It's not even clear if the bits can be changed without a bare die, a floating probe array, and specialized hardware. Though, like flash the electrons in the floating gates will eventually tunnel and cause the data to bitrot.
So from that it appears that even at the tens of millions of chips volumes that would make sense for essentially whatever size of maskrom, the memory manufacturers tap out at 128megabit for a mask rom chip, and push you towards something flash esque.
And at the end of the day, flash without the erase functionality is pretty damn close to a mask ROM, and lets you write it near the end of manufacturing rather than at the something close to the metal 1 layer.
Sure, your 2.5 year old models are running faster, but you can't drop prices on them without pushing the break even point further out.
If the cost difference isn't incredibly significant, will people even want to pay for the 2.5 year old model, or will they get more value for their money paying more to get better results from the newer model?
There's a lot of open ended questions that I don't have the insiders knowledge for to suggest whether or not such a capital outlay would be a worthy investment.
My guess is that state of the art stuff will stay on GPUs and models burned into chips will be for "good enough" applications that people are still teasing out. Probably highly specialized models in automated sensor units and such.
Even then, while there are some amazing FPGA-based synths available, companies like Korg just put their code on a raspberry pi and call it a day. The same is true for emulators (SNES Mini etc. are also just raspberry pis under the hood iirc)
You get an FPGA for timing. They're less capable, but (in many common design architectures), they output their results once per clock, every clock, on time, every time. If you can hit a fabric clock of say 100MHz, clocking all the weird logic you can stuff in there, it gives 100 million outputs per second, never skipping a single one for any reason (short of total failure). The penalty is that making a small change to your desired "program" can be very expensive, and many things won't be realistically possible at all. Or at least won't fit into a part that you can buy. But things like audio, video, and high-frequency trading love being able to guarantee timing.
(Of course there are other ways to write your FPGA HDL, but that's one of the more common ones. And you do see DDR-style clocking, and similar, every now and then.)
It depends. For some things, CPUs don't even come close. An XCVU13P FPGA can handle 1.2Tbps of full-duplex Ethernet @ 1 billion pps. And that part costs less than a grand at moderate qty, and with significantly less power consumption than a CPU that'd be capable of operating a dataplane at these speeds.
The point I was trying to make is that the CPU is a general-purpose creature and doesn't really care what you want it to do. If you had a CPU that could handle 1.2Tbps of Ethernet packets at 1Gpps, it could do a whole lot of other things involving 1.2Tbps of data flow too, very easily, if someone wrote the software. And more. (But you're probably not getting 2.4Tbps out of it, no matter what you do.)
An FPGA can not. There's plenty of things that those XCVU13Ps just can't do, or would do worse than a $1 microcontroller. (Setting aside for a moment implementing a CPU inside the FPGA... which does actually happen in just about every large-enough FPGA design, which is its own discussion....)
That would be better suited to FPAAs (field programmable analog arrays). FPGAs can usually only work with clocked digital signals.
All aboard! We're racing to the bottom now.
I don't want to live at the bottom
Companies being profit seeking entities that are actively hostile to social wellbeing would gladly put all of our money in a machine money making loop and leave all but a few humans out.
This may even be stable to enact and keep, as pensioners are famously more likely to vote than working people.
AI doesn't have workers rights, doesn't burn out, doesn't get sick, can be instantly onboarded, etc. The second we can be fully replaced with AI, we will be. Plan accordingly. I'm pursuing FIRE and considering moving into a trade.
But leaving them out of this conversation like it doesn’t even exist is bonkers.
And yes i am skeptical as a default on this subject. But still - photos don’t actually tell the whole story. What was the cost? How was the project schedule? Etc.
If there is really nothing for humans to do, whoever has control over (not necessarily ownership of) enough robots and AI to directly maintain and grow their collection of robots and AI, has something functionally equivalent to a breeding population of the stuff. (If they can't make more of themselves and humans made the initial batch, then that's a job humans can do so the initial condition has not been reached).
None of this helps people figure out how to look after their own interests in the meantime. The rest of society may carry on as today without AI, akin to the Amish if we're lucky (rejecting further developments unless good for us) or like Pol Pot if we're unlucky (reject anything that nerds like because nerds liked AI and everything breaks down).
On the other hand, the institutions may simply fail to handle reality, like in the Great Depression.
Doesn't matter what plans you make, the impact will be across everyone!
Moving into a trade won't help, because the supply is doubling while the demand is lowering. Fewer people with money to spend; they'll fix their own damn toilets if the decision comes down to "buy food" or "hire plumber".
https://www.mrmoneymustache.com/2012/05/29/how-much-do-i-nee...
In some kind of economic AI apocalypse, though, who knows if the assumptions behind FIRE will hold.
AI could and should benefit us all.
Lots of people would have happily taken GPT-4o as good enough for a lot of use cases a year ago and not lived to regret it.
> Model SOTA moves faster than chips can be designed or produced.
From what I remember working in that area the hardest part is getting masks for a design. Masks were developed in the span of half an year. Masks also reusable, they can be mixed and matched and this is why fabless companies work with fabs to produce specialized masks for them, it saves time for consumer to have masks for some macroblocks prebuilt.Here's my analysis of how to etch relatively big LM into silicon: https://news.ycombinator.com/item?id=47109252
Given some amount of work with the fab before main pipeline set (I think a year long process), one can then spew LM-on-a-chip in six months or less and much more than 2 per year, because there can be several LMs in pipeline.
Speed (clocks speed) , is dependent on your design layout, and for any resonably advanced layout, it takes lots of knowledge to push the clockspeed beyond 200Mhz on 'consumer'/prosumer models. (Compare to a few GHz for GPU/CPU).
FPGA speed shine where they can pipeline massively parallel calculations through pipelines with minimal lookups.
Yes SRAM is limited but that's more of a ram-process limitation than a limitation of FPGA
Wrong. This is one area where FPGAs have an insanely unfair advantage compared to CPUs and GPUs. Yes the SRAM is limited but you have so many individual blocks and all of them come with dual ports and getting the maximum frequency out of block RAM is much easier than getting the maximum frequency out of programmable logic.
If you wanted the highest possible memory bandwidth while being free to look up hundreds or thousands of independent memory addresses at the same time you're better off with an FPGA.
E.g. with an Efinix Titanium Ti180 you could hypothetically have 2560 simultaneous memory requests per cycle all pointing at a different address and process those requests at 1 Ghz.
looking up static values is quick and easy. When you need to lookup results from previous stages of pipeline rather than just feeding them forward, thats where I run into trouble. But I am a relatively fresh FPGA designer, so I am sure it can be done. And I probably need to level up my boards a bit too.
Also, it's hard to get fab capacity for any project. Let alone something so experimental.
Nothing was released in spring, and 2 months ago AMD announced their acquisition of Taalas. That doesn't exactly inspire confidence that their frontier LLM will arrive as promised.
Why not? AMDs chip design experience and production capacity are magnitudes larger than a small startup. Assuming that they acquired Taalas for their technology, I don't see a reason why they couldn't.
Because you can't use the former for anything else? You will have to buy a new chip every time a new model is released.
Not to mention that it will be much more expensive.
If you bake a given transformer architecture into silicon and then, a year later, changes in transformer architecture give a large inference performance boost, you may have to throw away all that now nearly-useless silicon that gets outperformed by humble GPUs.
I think the fact that there are plenty of 1yr+ old models on openrouter serving hundreds of billions of tokens a month shows that there's plenty of use case for models that are "good enough. Cerebras' entire business is serving older models at high speed. I would happily use an opus 4.7 at 15k tokens per second. The intelligence per second of an ASIC still makes sense even with rapidly evolving models.
But I think it's a good bet. I think that in two years, if I can get opus/sonnet 5.5 or the gpt-6 models for much cheaper and faster than whatever the "frontier" is at that point, that this will probably be a great trade for most of my work. I certainly don't know that for sure, that's why it's a bet, but it's what I think right now.
I wouldn't quite say that about any of the open weight models at this point. But I'm hopeful that will change in the next generation or two of those models.
But you're right in sense: moderate intelligence at superhuman rates (and presuming moderate energy usage) is very compelling compared to an intelligence that takes 1000 years to return "42"
So imagine taking a year-old SOTA model and running it at 100 tokens per second on an edge device. That's enough to feed a screen's worth of content through it and power decent multilingual message suggestions on IM.
Imagine running it at 1000 tps. That's enough to reparse that screen mid-keystroke, and give you semantic autocomplete in text. Or fully general "the phone has a good idea of what you're attempting to do" context at all times.
There's many, many new classes of features that will open up if decent enough models can be run on edge devices at 100+ "intelligence per second".
I'm poking (lightly) at the locallama game, have a fancy MBP5 with gobs of ram (so I can demo/trial locally) and have been semi-waffling between whether to chase a mini or studio for local "always on" type stuff.
My outcome was "CapEx v. OpEx", and dropping another $5k for an aluminum cube buys a lot of OpEx (eg: just trickle-drip HF/OpenAI credits to a raspberry pi or VPS orchestrator rather than trying to do the inference locally), ie: $5/mo inference for 1000 months.
HOWEVER, there's definitely a role for that 1-10 TPS type "ambient inference" that I wouldn't mind sustaining on any sort of always-on / local / private compute. My main email address is still on ...@yahoo.com and their spam filtering has gone to absolute shit.
Being able to have the always-on mini (local, trusted, no private data leaves my control) poke at the IMAP/email and thresh it into SPAM/HAM/Personal/Political ... random spot check, I'm getting ~5 emails per hour, and that's completely tractable for staged low-med-high processing. (Subject + rules.py? Subject + Body + LocalSlowTPS? Subject + Body + LocalDeepTPS? Subject + Body + RemoteLLM?)
Even if you did 10000tps of "jev" that's an incredible value... not quite "Literal AI Packet Router", but as you're dancing aroud saying... "Speed is a Weapon"
Look into "OODA Loop" => """The OODA loop is a four-step decision-making model—Observe, Orient, Decide, Act—created by U.S. Air Force Colonel John Boyd to help leaders make fast and accurate choices in chaotic situations. // The main goal is to cycle through the loop faster than an opponent or changing environment. By operating inside another person's loop, you create confusion and outpace their ability to respond. While initially designed for aerial combat, it is now widely used in business, sports, and crisis management."""
For example, Opus 4.6 Max was somewhere between Opus 4.7 Medium and High in some benchmarks, but it was slower. If there was a way to run it 10x faster, the economics would be different.
But it's complicated for other reasons, one being that the number of parameters for frontier models (especially with MoE models) are so high, and not always utilized (once again, thanks to MoE) that it would actually be incredibly cost prohibitive, if not impossible, to attempt to make giga-chips that would allow running it.
I definitely do believe that we will see more and more specialized chips over time, but putting the entire model on a chip is still a ways away.
I believe Taalas has a heavily handicapped llama 8-billion parameter model. And it still pulls >200W to run.
I can't imagine how anthropic or open ai would be able to burn a multi-trillion parameter model on a chip, we just aren't there yet.
The closest anyone has gotten is Cerebras with their Wafer Scale Engine. It uses SRAM embedded with the compute. A single chip is an entire wafer, but the headline spec, how much ram, only 44GB, which is tiny for the silicon area used.
I understand traditional IC production workflows are ludicrously expensive, and glacially slow, but surely at some point the economics are going to tip in favour of mask rom.
Say you setup your foundry/packaging/ai chip facility. You come up with a new set of model weights. Run your CI/CD pipeline to produce a new mask output. The only thing you've changed are the assignments of the bits, this is extremely low risk change. The new masks should be completely interchangeable with the current process.
You produce the new masks, swap them into your foundry process, and all of a sudden your new chips have the latest version of the model. This would probably manifest itself in the form of inference providers having yearly / bi-yearly "updates" to their models as new hardware is brought online.
Tiered subscription levels would gate keep access to the latest and greatest model, cheaper subscriptions will be limited to older versions of the model, and so on, until running the hardware is no longer economically viable (no demand/running costs exceeding what the market is willing to pay).
> We basically have an architecture where we are embedding the models, and we are hard coding the models and the weights into our what we call the mask ROM recall fabric, which is paired with an SRAM recall fabric. Together, they are able to store both the model as well as do all the computations of KV cache. We have adapters and customizations – we support all of that. This design allows us to be super-dense in terms of compute and in terms of storage, and we can do compute on that storage incredibly fast, which is what drives density up and cost down
Source: https://www.nextplatform.com/compute/2026/02/19/taalas-etche...
what happened to improve existing hardware in the ways not obvious to humans? Some novel transinstor placement, some new approach to make logic gates faster, matrix multiplications etc - the filed is huge
> Why aren't the labs burning their frontier models into chips already?
So companies try to maximize the memory bandwidth they can get, balancing tradeoffs of power/area/programability of their chip. Right now they feel like the economy on power/area is not worth the decrease in programability/flexibility.
The primary constraint isn’t likely what’s possible to do, but that the kernel and weights are too variable right now and the patterns too poorly established to bake into hardware accelerators yet. Margin pressure is also not there yet.
I suspect as the marginal utility of the frontier improvement settles into diminishing returns (I suspect we are there already tbh) baking hardware models with ROM, working set, and kernel cores collocated will be the frontier space as the goal will become reducing capital spend to utility levels rather than research levels.
Once someone has a model that is sufficient for almost any practical use, making marginal inference cost effectively zero will be the competition frontier. I do shed a tear for all those lonely data centers as compute densities will almost certainly make most of them a terrible investment.
But such is the cycle
Sure, but now we're not talking about just burning the weights into the chip, but also designing a new architecture that has memory local to each core. A new architecture would then require a new programming model, which means new inference stack, which may mean new training stack.
Think of something novel.
Your optimized hardware chip might be obsolete before its back from the fab.
SOTA Frontiermodelhardwarechip is a benchmark point of a potential model slow down.
Google is doing it right now under project Frozen v2 which should be ready by 2028? which is either just a small experiment or flexible enough and thats why it takes so long for it to happen.
Also can't keep them closed source if you do that.
It's upcoming second generation could run the inference of the models that are being used to improve it...
Much of a model are weights, and high-density ROMs are very very very hard.
Qwen 3.6-27B is 56 GB of weights. Let's round to 550Gbit. NOBODY makes huge ROMs in modern SoCs, but if we assume it is as dense as the densest TSMC N3 SRAM, this is 33.5 Mbit/mm^2
Then your weights are 16,400 mm^2 !!! The maximum reticle-size N3 die is ~800 mm^2 so you cannot make it.
But, you might say, SRAM is 6 transistors. ROM might be just one. So 1/6 of that... 2730 mm^2... which is still too big to manufacture.
Their stock price, be it public or estimated, is heavily pricing the notion that they are first and foremost Growth companies. Therefore their focus must remain on ever better and greater things. If they lose focus and get distracted by lesser endeavors, their valuations crumble, their ability to raise capital vanishes, and their runways collapse before they ever have a chance to reach their end goal, whatever that may be.
That means the boring job of productizing AI models into reliable systems that won't vanish in six months is left for a smaller company willing to pick up the crumbs. Unless they get acquired by Big AI before getting it done.
I picked TinyStories-1M, synthesized via yosys against SKY130 PDK. It's GPT-Neo with 8 layers. At int8 it's ~30mm^2 per layer (including kv caches, multipliers and all the attention stuff). So it's 250 mm^2 for the full toy model.
I'm kinda surprised, I spent no time, so if you want to etch a toy model, you can do that (you can probably reduce the area significantly). If you want a mask for that is probably going to cost 100k though.