It’s an ASIC with the model wired into it so it’s very low power and fast.
I’d buy these. Say $100 for a frontier class model. Maybe more.
It’s an ASIC with the model wired into it so it’s very low power and fast.
I’d buy these. Say $100 for a frontier class model. Maybe more.
There was a paper a while back that showed top-K selection like that with tiny models was able to reliably solve some 1M-step Tower of Hanoi when no frontier model could. Very big level up in capability just from horizontally scaling compute.
If I can make my small fast AI model just a tiny, tiny bit more capable, and still run 100 of them or 1000 and run an evaluation model on top of that, the overall system capability will scale quickly with tiny increases in base model intelligence.
So i guess maybe they currently try to solve a very hard problem with a small focused group before scaling or they are dysfunctional.
Also Llama 3.1 8B is a dense model AFAIK and they are fast by nature. As there are not a lot of dense models these days i could imagine that they try to optimise for MOE models.
We could have the photonic AI model ASICs for real!
So basically a static model version of consciousness uploading.
But the flip side is possibly 1000+ tok/s on a SOTA model, which would be game changing.
Could make sense for datacentres or enterprise, but I don't think we'll be getting SOTA Game Boy carts this decade.
Imagine what’s possible if you had GLM-5.2 turned into a hardware chip like this.
It's not as simple as a weight swap between identical architectures.
The speed gains are also from not having to route the weights through wiring like with ROM cartridges.
Sure you would. Running frontier class models on current hardware costs in the order of tens of thousands of dollars. It is more likely that these custom ASICs will be priced competitively with that, and not with Super Mario Bros.
Oh, and energy consumption will be in the same order.
Doesn’t need hbm or lots of memory, because the hardware can just forward the data straight to the next layer and you don’t need to round trip through memory.
They claim to be working on an approach to make the underlying hardware a bit more reusable between models.
Most big model weights will not fit a single reticle sized chip - so you’d have prob 30 different chips to split the model .
And you’d need super fast chip to chip comms for the all-reduce and similar.
So scaling to 1T models is hard - and a long lead time - but can be very power efficient.
1) the hardest, custom silicon + MCU to manage the USB interface
2) not as hard, shared memory, NPU + MCU to manage inference and USB interface
Theoretically you could do 2 with the right MCU, NPU, and memory combo. You'd stream/DMA the weights from memory into the NPU and then read the results with the MCU. From a user's perspective, it might take the form of an openAI API compatible endpoint that enumerates when they plug the USB device in. There would likely be some host-side software to ease the pain of trying to use a USB device as an HTTP API.
> “In the current generation, our density is 8 billion parameters on the hard wired part of the chip., plus the SRAM to allow us to do KV caches, adaptations like fine tuning, and etc. In our next generation, we would have the ability to go up to 20 billion parameters in a chip. Even with trillions of parameters, we’re talking about few tens of chips, which is a very, very small compared to anything else out there on the market today.”
https://www.nextplatform.com/compute/2026/02/19/taalas-etche...
Edit: i do not know how reliable this page is... it has a lot of typos.. for me it does not look like it was written by a LLM
We’ll see what the market chooses
That sounds good and practical to happen!