AMD acquires Taalas to boost inference performance by etching models in silicon
theregister.com
theregister.com
"Good enough" LLM functionality (for the use case) will be on-die or on-chip for cars, appliances, etc. This will provide speeds of chatjimmy at a battery-level power consumption.
Probably this will also happen for software engineering. Some usb-powered AI accelerator with Kimi K3 (and in future even better) performance running at 10K+ tokens/sec under 50W of power purchasable for almost no cost. Need a better model? Buy the new hardware. Old hardware is probably still fine for a lot of other use-cases. I expect China to be a big player here, it fits their open-model and hardware-manufacturing strategy.
I would expect that Gemma 4 E2B (approx 5.1 billion parameters) or maybe even Gemma 4 26 billion A4B at some point is running on a chip.
Their system is honestly very efficient for the weights, the problem is the KV-cache. That's why HC1 only supports such short context, they use SRAM for that and spend most of what's left of the die for it. The recent advancements that made attention more efficient are probably going to be very useful for them.
Etching failure in that situation would be like brain-damage in a human, all sorts of weird effects would start appearing.
(Is Cerebras doing something novel? CPUs and memory blocks have been doing those things for a long time too, since the error rate is otherwise too high for normal size chips as well)
[1] see eg https://www.vlsimentor.com/dft/redundancy-bisr to get some basic concepts
Did you mean that each HBM3 stack is that large? Because it only takes one glance to see that the memory chips are much smaller than reticle-sized GPUs they sit next to.
A fairly dumb but FAST model has it's own totally distinct use-cases even if it can't be scaled in size. Think about a LLM-infused-Alexa where the response time is instant. Where you can request it looks at hotel options in Montreal, and it starts answering in half a second rather than a few minutes.
Plus some sort of slow smart + fast dumb combo architectures might also work really well for different classes of problems.
Yet the answers will get outdated quickly whilst the silicon is fixed.
Bro is living in 2020 before rag was widely introduced.
Yes, to update the blueprint for new models two layers will be updated. That is the NN.
To instead update the data on which to operate you could use a RAG to query.
(As in "the Pathfinder 2.0 NN is on the chip; the geodata is in the OpenGeoMaps dump-DB-nightly" - not really overlapping with LLM+RAG but may give an idea in a different scenario.)
The agent doesn't know the date, or know what hotels there are in Montreal, it sees:
> System: You are an AI agent. The date is 11th August 2026. Your knowledge cut-off is March 2023. User is based in <date>. If you need to search for something to support the user say {search:<term>} and a list of options will be provided along with instructions on how to access. Or say {help} for a full list of commands.
> User: Can you help me find hotels in Montreal for next weekend?
The AI then interacts with the tools given in the base prompt, which can obviously be updated. So it then goes:
> AI: Of course, let me search for that. {search: hotels in montreal for 16th August}
> System: [Provides list of websites]. Say {read[n]} to read option or say {start subagent:<goal>) to register subagent.
> AI: {start subagent: List hotels on booking.com available on 16th August}.
[etc etc, then eventually]
> AI: Yes I have searched for you and I found a few options!
While you can't change embedded knowledge, a good model knowing that the date is 5th January 2040 can infer certain things (e.g. while it might not have been trained on certain deaths, it can probably guess that it should search before answering if it means a person would be 102 and their last information is from 2024)
And when did RAG start to work properly as a mature, reliable technology?
I've just been doing research and experiments for work related stuff.
Typically we've used plain embeddings for a lot of high contrast documents aka discrete facts.
However I've been working with a >1000 page document of complex procedures with incredibly low contrast where embedding falls flat.
There's top down/graph searching, bottom up/embedded; alts like colbert, reranking, reasoning, search agents and now (though seemingly quite new) specific search agent models.
Ultimately I found that a reasoning enabled search agent doing a hybrid of bottom up (with reranking) followed by top down, gave the absolute best results. Paired with Luna for cheaper and faster tokens it benchmarks pretty well even for vague references to procedures.
I would imagine that search specific models just coming out are even better and I'll have to evaluate using these but for now the above works well for us.
Having an agent get vector search results to use as anchors and then being able to explore the sections and subsections above that, then eventually digesting as much as is relevant (big context, cheap tokens) is amazing.
*(Of course it has "always" worked well for «simple high contrast Q&A», ever since the base embeddings technology worked properly: that is almost by definition; it is on real world use cases, where the nuances of reality are present, that it failed miserably.)
Anybody who has a specific informal query ("SELECT ... FROM ... WHERE has_carpark AND ... ORDER BY score(has_jacuzzi , walk_distance(...) ...) DESC") but does not want to research and cross the different scattered info himself (does not want to build the virtual DB himself).
So I have faith in these embedded LLM chips when it comes to fun projects like that. I have not personally found my quantized QWEN good at agentic tasks, though, and it LOVES to make shit up when asking questions about documents in the prompt.
its lots of parallel calculations, rather than one blazing fast one
Physics has hard limits and Moore's law is long dead.
Now imagine you have a chip which is just that model, but can do it at absolutely insane speed. Like tens of thousands of documents a second.
Same for things like text-to-speech or speech-to-text. Think of the accessibility wins if subtitling becomes insanely accurate and fast and omnipresent.
There are all sorts of domains like that, and the trend has been such that smaller models are getting smarter and smarter. If you can stick them in parking meters, traffic lights / street crossings, mobility aids, etc etc I just see so much potential win.
And one can say LLMs are not as smart as a human, but a lot of the reasoning humans do for product and service generation isn't smart at all—it's just a bit of fuzzy input/output plus some reasoning rules. And then if you hook a robot up to the LLM, you can get results in atoms instead of bits.
I'm very excited about the future. I also hope it will stop money from flowing to bureaucrats who are incentived to keep the problems open to keep the money flowing, and instead facilitate sharing directly with the people (for example, no money to the state to solve homelessness—instead, spay instead with intelligence output to build a house and provide food as part of taxes.)
Specific traces for specific inferencing will mean that some generations get deprecated. Look at H265.
If you make chips, you want models to be free.
Sentient Switchblade: "Hi Beth! You've gotten taller! Shall we resume stabbing?"I expect one day having small lower power demand drive sized devices with proprietary burned-in models that are quite fast running on-device in robotics and such. Commoditizing LLMs via burned and locked hardware seems likely when LLMs have stabilized (when we reach a year between releases again) and the hardware is capable and “disposable” enough. “Buy a robot and upgrade it forever* (5 years) with newer models (sold separately)”, at least until the planned hardware obsolescence that the “interface has changed to support newer hardware, so you’ll need to upgrade (again) to use the latest features”.
The plans basically write themselves.
My food processor could use a self-cleaning feature. It could only be made worse by some system that, IDK, changes the setting based on off-hand comments about "I don't know what his beef is.."
Baking models onto silicon would've been the next logical move to get a moat.
Google is already doing this and has an experimental project on top of already having TPUs and cramming their quantized flash onto individual TPUs for inference.
(Yes, you could fix a number of masks, e.g. entire logic gates, of course).
*(It's local: private files managing firm oriented. It's blazing fast: it can be placed into recursive, intensive local workflows.)
Now maybe. When models are flying passenger aircraft, other prerogatives will assert themselves. When a 50TB ROM means you can impulse purchase a ChatGPT 6.3 xhigh that runs on batteries, yet more use cases will be apparent.
One of the underappreciated effects of the AI boom and associated money is that it has strongly reinvigorated R&D in hardware: it is clear that there is a real application for far greater density and lower power demand, and people are now pursuing this much harder than they had been. That will yield what it has always yielded; orders of magnitude jumps in capacity and performance.
That can only work when there is physical capacity for improvement though.
> underappreciated effects of the AI boom and associated money is that it has strongly reinvigorated R&D in hardware
Yes, absolutely: but the point at this stage is more about finding new possibilities in hardware architecture than the improvement of what we had. So
> * That will yield what it has always yielded[:] orders of magnitude jumps in capacity and performance*
That will yield new and renewed hardware technologies.
(Already the distinction between SRAM and DRAM was overly specialistic before this boom - now it's on our mind as we know we need to "expand", "make cheap", "integrate" or find alternatives.)
There are great opportunities for advancement. Both in the physical hardware and in how and where it's deployed and powered.
Consider this, as only one point: there hasn't really been a demand for advancements in ROM. RAM has been scaling at approximately Moore's law rate, and nonvolatile R/W storage has been sedately scaling, but there hasn't been a use case for really dense, high performance ROM. Now there is. ROM used to be a big deal in computing and media (cartridges, optical disks, etc.,) but that tapered off long ago; volatile and R/W storage was sufficient and convenient for the time, and the inference model use case, where dense, high speed ROM can have extremely high value, didn't exist.
Now there is a use case, and industry is thinking about something they haven't cared about in a long time. Current fabrication nodes, stacked in the third dimension à la NAND flash, could produce staggeringly dense, fast and low power ROM. That's why AMD snatched up Taalas: they're thinking about an aspect of the future that has been (reasonably) neglected.
Clearly there are possibilities, some of them proven (proof-of-concept, in-production etc.) - but taking for granted "Moore's law" like spaces for them may not be founded on what we know at this stage.
A fast, low power ROM is the key ingredient to near term local inference with large models at low power. If I could offer you a $500 ROM that provided the model data for frontier inference on power similar to a desktop GPU, you would buy it, and consider it a bargain, even when it came time to pay another $500 for the upgrade.
Surely it is clear to you that Read-Only /Memory/ does not /compute/, and our need is to compute through the data in the memory... That is CiM - a technology not that similar to ROM... Because a plain ROM does not solve problems in this area...
In other words,
> If I could offer you a $500 ROM that provided the model data
Then I would have a physical token containing what I already had as a file, and the problem of running that file into something efficient would remain... Because the ROM does not "run" its contents...
Conventional GDDR/HBM don't compute either, yet inference is implemented using these.
Compute isn't the inference bottleneck. Inference requires high bandwidth, high capacity memory. The compute resources necessary are fungible, comparatively cheap and already available, at least for a small number of concurrent loads, such as in most local inference use cases.
> Then I would have a physical token containing what I already had as a file
I suspect you are not grasping what I mean by ROM. Dense, high performance ROM would not be the hardware equivalent of a "file", with performance bottlenecked by low bandwidth, high latency storage media, serialized for RW coherence reasons. It would have extremely high bandwidth, on par with GDDR, low latency due to a dedicated high performance bus, high concurrency due to a lack of any RW coherence obligations, and operate at low power (no gate leakage, no dynamic refresh,) and low cost compared to equivalent GDDR/HBM capacity.
Essentially what high performance ROM would provide is high capacity, low power HBM, albeit read-only. At that point all you need is sufficient TOPS to run the inference algorithm. The compute part is already available, affordable and readily scales up and down as per performance/cost/power budgets.
But ROM has a massive disadvantage being static. So, either it is cheap and practical "like a CD", or decision making will be forced to do its evaluations.
We have a von-Neumann architecture RAM<->CPU, which is really suboptimal for running current relevant Neural Networks ("RAM<----...---->CPU"). Advantage: flexible.
We have a CiM with Taalas HC1 which has the massive and enabling advantages of running NNs very fast and very energy efficiently.
What could high-speed ROM bring? It must be a good combination of "fast" and "cheap" to to be "interesting" for the market, between those two contenders.
I believe that "practical" as in "replaceable" is also a fundamental property of what we desire in this field: the Processing units are not all there is, also the side-RAM (for context, kv-cache etc.) is a necessary part of the system, so the NN-container is just a piece (which needs expensive co-parts). Whether the NN-container is CiM or not, it will be critical if it can be replaced (like a cartridge, disk, etc.) so that the other parts will not need replacement with it.
My understanding is that Taalas HC1 is "mask-ROM" fabricated at 6 nm for bulk model base-weight storage, and some SRAM for KV cache and other bits:
https://www.eetimes.com/taalas-specializes-to-extremes-for-e... "On the HC1, the model and its weights are stored on the chip using a mask-ROM-based recall fabric paired with a (programmable) SRAM"
I don't believe that's CiM as you advocate.
> I believe that "practical" as in "replaceable" is also a fundamental property of what we desire in this field
I suspect that there is a important frequency factor in in the "replaceable" calculus. Already I see people dragging their feet about adopting newer models once they've found familiarity with some older model: "good enough" is a thing. I know there are industries where "validated" is a concept, and they do not ride wave crests. So, if we imagine that as all this eventually shakes out and we're not replacing models every few months, but instead with about the same frequency as our cell phones or similar, the ROM model works. If the performance and price make this pattern highly appealing, then that's what will win, certainly for local inference. If some datacenter operator could, today, adopt a ROM approach that cut their power budget by a large factor, but had to suffer 2-3x longer model update cycles, they'd likely consider it.
For better or worse.
I have no problem with CiM as a concept. If it can reduce power/size/cost then it's another avenue that inference will probably incentivize, where incentive has previously been insufficient. As we both agreed long ago in this thread this new era is motivating things that were previously neglected, and CiM is possibly a part of that. My dream is that all of these get a hard look as people try to figure out how to run all of this without enormous gigawatt sucking datacenters that rival DOD program budgets.
You missed the whole point of Taalas HC1: that it is Compute-in-Memory.
> 2. Merging storage and computation // Modern inference hardware is constrained by an artificial divide: memory on one side, compute on the other, operating at fundamentally different speeds. // This separation arises from a longstanding paradox. DRAM is far denser, and therefore cheaper, than the types of memory compatible with standard chip processes. However, accessing off-chip DRAM is thousands of times slower than on-chip memory. Conversely, compute chips cannot be built using DRAM processes. // This divide underpins much of the complexity in modern inference hardware, creating the need for advanced packaging, HBM stacks, massive I/O bandwidth, soaring per-chip power consumption, and liquid cooling. // Taalas eliminates this boundary. By unifying storage and compute on a single chip, at DRAM-level density, our architecture far surpasses what was previously possible.
https://www.sofx.com/nvidia-ai-module-migrates-from-russian-...
Not "multi TB frontier" by any means, but the direction is clear: weapons will be made to think, for better or worse. Something of the scale of a frontier model will likely be seen in: loyal wingman aircraft, autonomous warships, military satellites, to name a few platforms.
i have written about this:
"For device makers
Packaging models with laptops and smartphones will let application access near free, low latency inference and potentially offer users a better experience with the option of preserving data on-device. This is viable under the condition that tasks that do require larger expert models that run in the cloud can be routed to external models. A side-effect of local models and what will let Apple cut upgrade cycles from ~4 years (?) down to 12-18 months is specialized hardware to run them. For almost a decade, smartphones have been trying to compete on better cameras. This coming decade will see them selling better GPUs, NPUs, ASICs and whatever other things they'll be calling the inference chips, to drive re-purchase. Every six months will see a better model on new hardware, which will enable better performance in certain applications."
https://try.works/role-model-the-case-for-a-model-routing-pr...
You've confused engineering compromise for malice, and reversed the purpose. For the model capabilities and inference power draw, what alternative do you see to a (at least mostly) fixed hardware model?
You did not compute that as the cost for a speculative card from Taalas, right?
Taalas is going to have a tough time putting a trillion-parameter model on one conventional die. Their HC1 die is already near the maximum size that conventional lithography can expose. They claim they could partition the model across many chips, but I'm not sure if they have tested this process or what it means for compute. The basic storage arithmetic is unforgiving: for a one trillion parameters model at four bits it will take 50–100 chips. To service a sizable customer base will take thousands of 100-chip fabs.
That all said, I'm bullish on this technology, and look forward to seeing it evolve.
Eventually someone will have to solve compute in memory at scale.
If the LLM response only takes a few milliseconds, the chip can process hundreds of other requests until the first conversation becomes active again.
Sounds a lot like "640Kb ought to be enough for anybody"
Yes, a cheap and fast Opus4.6 can drive a lot of value in current context. But if we continue to craft bigger-and-bigger balls of mud, Opus 4.6 may end up hitting its conceptual ceiling and unable to contribute.
Winding the clock back on your statement gives:
> I'd gladly pay for a Claude Sonnet 3.5 in silicon and use it for 1-2 years.
Man, I dunno.
Perhaps in some cases, but the value I personally and professionally got out of LLMs reached a limit a while ago and has since kind of fluctuated between that limit and a bit less.
If the best model was instant, like the demo here, it could certainly provide more value, I guess, but I think the limit I'd quickly hit is the same one as now, which is how much of it do I want to produce, for what reasons?
Text diffusion might be a disruptor here, but let me just say the most cutting edhe form of image diffusion (JiT and DiT) right now is just a big fat stack of alternating attention and MLP matmulls. Not theoretically hard to bake
I expect this to be around the time when we're finally ready to travel to Mars.
A ~30mm side for the HC1 tech for an 8b model (still unclear the planned HC2)?
But if the basic premise of "good enough LLM at insane throughput" holds, I think it could qualitatively change local uses of LLMs. At a certain speed point, you're able to move from request -> response to a cascade of tool calling and "subagents", which could allow a small model to be much more useful, if provided with a lot of local data and tool calls.
That said, this is assuming you could stuff a "good enough" model into a phone with Taalas-like technology. The Taalas tech demo was an 8B parameter model and required hundreds of watts (IIRC) to run. The efficiency was good given the speed (as I understand), but it's not clear at all that the approach scales small enough to be a sensible coprocessor on an iPhone or whatever.
That also needs server class hardware though. A phone won’t happily service the insane amount of IO, compute, and network that this cascade would require.
Raw intelligence becomes slightly less important when you can iterate and improve automatically. You can still claim it was "one shot" even when 30 different implementations were made then combined.
Ok, real life example: I now spend most of my time, as a developer, waiting for the agent to do its thing (after careful prompting, I'm also thinking about work stuff, don't worry I'm not useless). What if it gave back the same excellent results, but instantaneously? Why, then, I certainly would become the bottleneck. So, quite possibly, my last work task would be to plug this agent directly into the ticket system where the domain experts input their feature requests. Maybe we still need 1 developer out of 100, to coordinate releases and all that (ok, say 1 out of 10).
But that's not taking things far enough: why do we need these domain experts at all? Our pitch is clear, and all software-enabled, though it took years to develop. We can just have the clients express their concerns to the AI, directly or indirectly. Have multiple lighting-fast agents with different roles (refactoring agent, new features agent, debugger agent, domain expert agent, etc.). So we fire everyone, maybe keep 1 product owner / devops to keep the trolls out. The cost is still probably 100 times less than it used to be (beyond the initial cost of acquisition of the magic machine or whatever).
But one of these clients, surely, will realize that these 10 years of manual and slowly-automated development can now be emulated in very, very little time. Why not just, say, take screenshots of the entire app and feed them into the magic machine? Why, this way, they could have the service for a tenth of the yearly cost, forever!
And then the economy implodes.
I'm not saying it's THE most likely version of things, I'm saying that at a certain level, quantity (or rather, speed) is a quality all its own. And this new quality might change the world. Let's hope it's for the better!
That said, it obviously depends on the project.
A frustrating vision of the future would be when we've been asking for faster loading lighter web pages for years and then companies start caring about it and improving it not for us humans but for LLMs.
you need to launch 10-15 more terminals, who is waiting these days? :)
AI previously provided speed but not quality. As soon as quality reached an acceptable threshold, the speed became the reigning factor.
In my opinion the quality is still much lower, but speed means the cost is significantly lower also.
I mean, if an agent can do half-decent work in less time than it takes the user to prompt them (and "user" in this context is a fast touch-typist like most programmers are), it's obvious it's not the agent that's the bottleneck anymore.
This has always been the case for human project management, and LLMs just aren't at that level yet.
It's more like everyone is speed running to how fast they can convince others that "half decent" is good enough. And for sure, newer models of LLM seem to be getting better at that.
But that's what Agile is all about, isn't it? We've been speedrunning delivering increasingly smelly shit at increased velocity ever since SaaS became a thing, because ubiquitous Internet access is what allowed our industry to adopt the "lob feces over the fence for users to deal with" release model.
AI does speed that up, true (though since the market - and management - didn't catch up with it yet, we have a brief moment where we can use AI to increase quality while keeping usual delivery rate.)
If inference speed goes up, I can launch the same query 5 times, evaluate the best result and proceed from there. Of course, evaluation is also instant, so in seconds I can get a near perfect solution. Or maybe 10 and I can pick what I like the best.
It would certainly be an accelerator for people who know exactly what they want. And it would remove multi tasking, which I‘d appreciate.
This is the "dumber but honest person that works harder" phenomenon, vs "lazy genius".
Sure, in the future full rewrites and stuff like that will be just another "throw money at it" problem, but fundamentally software can get arbitrary complex and we barely know how to write large, maintainable code bases.
Nonetheless, I think testing (and maybe proofs) will have its long-awaited time to shine, as being the "reward function".
Right now, I put models in low thinking mode during my refactors and hate waiting. I would much rather have a faster model that that maybe was slightly stupider, and I would wait far less long between prompts where it needs my valuable input.
Models that are dumb, but humble and fast, can be fine.
Obviously a CTO is not going to walk away from the technology just because it's not good enough. That much more incentive for someone to create a powerful enough harness that can direct that power safely and productively. Like a nuclear core, we'll need to come up with the graphite rods and water tank. And if tokens are essentially free, why not, for every million tokens, spend 10x tokens on code review, testing, etc?
Whatever you can cheaply do with AI is not a moat, if there is profit in there there will be quick imitation and competition will eat away those profits.
Models can be replaced easily, harnesses & AI tools too. And if cloud inference gets too expensive there are local models keeping the cloud prices hard capped.
Probably AI won't make anyone very rich.
This model had zero information right, while being fast in responding.
Unacceptable.
> Bruce Lee was born in San Francisco, California, USA on November 27, 1940.
> Bruce Lee's father was a Chinese opera singer
That being said, this is not a good test. It is a language model (a very small one), not an encyclopedia.
ChatJimmy interface is just a tech demo. Without tool calling functionality we can't expect it to be factually correct.
This will generally make them suck, though, a little bit of randomness is necessary for proper function.
I just pasted your comment and its whole inheritance chain to it, started my comment, and asked to generate a total of 9 completions, 3 from each of {current & next word, current paragraph, current paragraph + rewrite the entire paragraph}.
Half of the answers were perfectly good (ironically, not the "next word" ones!), but the important bit, they came back near-instantly ("Generated in 0.024s - 14,163 tok/s", the page says). Slightly more powerful model while keeping this under a second, and this could easily become a qualitatively different form of autocomplete/text suggestion. Running in the background every couple keystrokes, or every time user stops typing for more than 500ms.
>I just pasted your comment and its whole inheritance chain to it,
Good idea. Only problem is it doesn't work. I just did the same thing with exactly this prompt:
>did the user IOT_Apprentice participate in the thread below and if, number and quote all of their comments. Only just number and quote the comments or write "Did not participate", do not add any commentary. Quote any comments by this user verbatim, exactly as input. Thread:
followed by pasting the thread[1]
And received the answer "IOT_Apprentice did not participate in the thread."[2] in 0.001s, even though they have literally the last comment in my quote and it's clearly legible.
It's particularly insidious because the understanding and thinking that is required to follow my requested answer format exactly is substantial - so based on the fact that it gets the format right and clearly understood the assignment, I would be inclined to believe that it would also be correct!
So to use your example, it's not just autocomplete, it's autocomplete that confidently returns "No matching results" in 0.001 seconds, even though there is a search term matching what you put in, right in the prompt itself that was sent to it. That is much worse than useless.
[1] prompt: https://ibb.co/CKVmRvtd
[2] result: https://ibb.co/BKdRKmyD
Fully interactive realtime NPCs in videogames at scale.
Recommender systems that simulate individual consumers.
Crazy shit
In a way... when it's finance, they should be maybe called Gray'ish Swans?
Better yet use it to dimulate counterfactual phenomena like market manipulations ypu intend to enact...
Autocorrect that works. Reply suggestions that almost work, just need to be tad more accurate (probably more of a data access issue than model) and a tad faster to look completely seamless. Screenshots with automated text detection and OCR and automatic interpretation (different suggested actions for when something on the picture looks like a web link, phone number, postal address, e-mail, or QR code, or an event poster). That's just a fraction of things I saw showing up on my Samsung phone over the last 6 months.
For over a year now, you could get a much better autocorrect and spell/grammar check, and a translator all in one, if you just pasted your text to a frontier model and asked it to check for errors or translate into target language. Now imagine being able to go through a round of such checks in a 1/100 of a second. You could have this running every keystroke, and suddenly the inline autocorrect/checks would not suck anymore.
Auto-linkifying that can correct for typos and doesn't need careful regex tuning because it understands from context what is meant to be a link or not. That's just one of many obvious things possible once you get local models running fast enough. Tip of an iceberg, and the first step to imagining all the other potential uses is to let go of the two mistaken beliefs people hold on to:
1. That LLMs are about written language. They're not; ever since "multimodal models" became a thing, tokenization extended to visual and audio space, and now textual and visual and aural inputs are all just regular, first-class tokens.
2. That chatting with the models is the only optimal way for end users to interact with AI. That's just artificially limiting yourself to the space of chat-based UI.
8B model (FP4) = 4 GB DRAM = 32 Gb DRAM = 80 mm2
8B model (Taalas) = 4 GB ROM = ~800 mm2
I could see it being feasible to get a Qwen-3.6-27b type of model done on something like this. Qwen-3.6-27b at 18tok/s would be a game changer.
Rather base model on ROM + KV cache on DRAM is much more scalable. Also this would work great for edge devices that have a 2-5 year lifecycle.
Someday, I imagine model weights could even be encoded as analog resistors (memristors or similar) for even greater density
The trick is that every compute element in their system has it's own small pool of ROM, instead of putting all the ram behind a common pipe. ROM is just used because it's the densest kind of memory that can be fabricated on the same process as their logic.
Think it's called Askjimmy or similar.
(Not that I believe it, it writes too well for GPT-3.)
Hosted frontier models from two years ago would be much faster today, too.
Less so for consumers though, because it'd mean the phone is out of date in 3 months when a better model comes along.
Now you sell the same phone with higher price tag.
At some point the music will stop on training bigger models, and when that happens it will make sense to have ROM weights (or 100% analog circuits given how noise-resistant LLMs are), but we'll know when that is because the investment bubble funding the training of new models will have burst.
The rate of change to the models has to be slower than the hardware roll-out to be worth a hardware solution. If "good enough" happens before then, that just means the user gets a software solution.
The rate of change by itself doesn’t tell you the whole story because of costs and diminishing returns. So what if your model is twice as good if it’s 10x the cost and it saves you 1ms? Everything else about phones reached “good enough for a phone” levels in years, and then got minimal generational improvements.
The reason we don't do this in general (any more) is that for long chains between input and output it has been much too difficult to avoid accumulation of errors. LLMs happen to be extremely resilient to errors like this, which is also why we can use e.g. 4-bit weights.
That's the only thing the normie consumer cares for really.
Ofc if the model has some critical bugs that’s another matter.
Maybe baking in a model that is "certified" to have some unconditioned truths + rest is pulled from external models/store could make sense. But AFAIK that doesn't exist and I'm not sure it can possibly be made. Perhaps society as a whole at least can work on an open corpus of training data, but I'm not holding my breath on this.
Nope. And not only not a decade ago, right now.
If you have an Android or iPhone, you can give it clear and easy to understand instructions that Gemma 4 could complete[1] if it had tool calls on it, and that 100.00% of Claude, ChatGPT, Grok, Kimi, you name it, could understand and all complete if they had the access.
The phones will fail to complete it. I just tried Siri. I said "hey Siri", waited for Siri to come up, and then I asked one of the exact sentences you replied to: "what's the weather this afternoon?" It thought for around 20 seconds, and said "Something went wrong. Please try again."[2]
I have Wifi, I have mobile Internet, I have free storage space, I have up to date software. What went wrong is that phones have never properly connected agents, not ten years ago, not last year, not this year, and probably not next year.
But don't settle for what Google could do in 1999 by hotlinking the keyword "weather" in any query to the weather being shown in the results.
Tell your phone (any phone): "Please call back the last number that called me that is not an unlisted number, regardless of who it came from."
0 out of any phone will complete that today, tomorrow, a year from now, five years from now, ever, because phone makers are not going to let them do that.
Meanwhile, 100% of all frontier agents could complete it if they had tool calls on the phone. Which they don't, and won't ever, thanks to the duopoly.
Okay, that's a bit dismissive, I would love to be wrong!
[1] after any voice recognition to text - which does work really well on both Android and iPhone! [2] screenshot: https://ibb.co/21rtDnfV
Can you say this to it: "Hey Siri [wait for it to come up] - please send me an email with the temperature right now so I have it for my records." and see if it can complete the task without any backtalk or misunderstanding, and if you get exactly what you asked for. (It's a really clear request.) Should be 1 statement, no clarification, conversation, random search results, ("Here's what I found!"), etc.
A normal frontier model can do that - or Siri can do it if it is properly connected to Claude, ChatGPT, Gemini, Grok, or any other frontier AI - but previously it was never properly connected.
If it can do this task, I might have to look into this again. It counts as a success if it sends yourself any email with the current temperature and you actually get it (it can include whatever other text in the email), and a failure if it talks back, says "here's what I found", says it can't, asks you any question, sends you an email that doesn't actually contain the current temperature, just reads you the temperature and then asks if you want it to send an email, etc. Should be 1 shot.
let me know if it works!
Subject: Current Temperature Body: The current temperature is 27°C in <my city>.
Once they started seeing useful (if niche) functionality as a cost center, there wasn't really a world in which these could usefully exist. Their big bet now seems to be that LLMs will lead them to profitability - but whether that's from increased data harvesting, cheaper integrations, or because it'll be useful enough to charge subscription fees, I couldn't tell you.
What Siri is missing is more logical solutions and answers for recipes, etc (still suck even with chatgpt integration).
"Hey Siri, what's the weather in <nearby town with a generic name> tomorrow" and it gave me a town with the same name ~800 miles from me.
Add a bunch of chips together, and you get to a server that can run a 800B model, very fast and probably significantly cheaper than others.
[1]https://www.eetimes.com/taalas-specializes-to-extremes-for-e...
AcmeAI Carbon
Market it as your premier (only) model at high throughput. Two years later you stand up MSICs for the new state of the art with entirely new hardware, your lineup becomes:
AcmeAI Nitrogen (top tier) AcmeAI Carbon (mid tier)
If you just kept pushing the same model down your pricing tier over time you could still extract a lot of value from an old model, even years after it's been set in stone. Working on brand new code/frameworks? Pay to use the newest model. Working on legacy code? Use the lower tier models that will already know your legacy frameworks, pay far less and still get massive throughput. I've worked on a lot of government projects that this would be absolutely brilliant for.
The other side of this is that agent harnesses are NOT set in stone, so even a legacy model with a knowledge cut-off that's years out of date can likely still be helped quite a bit by harness and fetch behaviours that are still developing rapidly. Especially at this kind of throughput.
This tech can definitely scale up from the current 8B prototype, but - at least as far as my limited understanding of the tech involved goes - you cannot just ASIC a trillion weights model due to physical size constraints.
___
Specification HC1
Model Llama 3.1 8B (hardwired)
Process TSMC 6nm
Die size 815mm²
___
So the current prototype already pushes the limits of what we can fit on a single die, and that is already likely going to limit your yield.
- Deepseek V4 Flash is impressively capable. Sonnet still beats it out by a thin margin, but the real kicker is that a typical session with Sonnet at current API costs is ~$2. The same session with Deepseek is 2 cents (ha). Its even allowed me to consider offering free-with-limits API usage on my own app. - Taalas (or competitors) have a lot going for them. If anything I feel like they need to join hands with these smaller model makers and converge in 2028
If you could manage a per-die expert somehow and keep the expert routing gate relatively fast (through an interposer interconnect or doing wafer-scale Cerebras type shit) you don't need to keep the whole thing on the same die. Small dies with one expert per die on an interposer, and a very tiny router might be sufficient.
Of course, 1T SRAM isn't really SRAM, but my understanding is it doesn't require external refresh like eDRAM, is a bit easier to fab on-die than eDRAM, and is half the mm2 per Megabit compared to real SRAM (15% more die size than eDRAM)...
edit: Ok, I will self-apologize. Its apparently a 3B model. Mighty impressive for what it does.
Where I work we invited a bunch of people circa May 2023: a few top tier academics, a few government and NGO officials in charge of our industry, and a few startup founders. We are a big company, so people were kind enough to come and give speeches - and discuss.
This was a roundtable on what's going to happen. You could play some of the speeches verbatim today and they would not feel out of place. And that is telling. We had a Stanford professor saying that gpt-3.5 can do everything he can - just better, and he feels his profession is on borrowed time. We had a government guy saying that entire swaths of jobs will be displaced before the year ends. And so on.
The interesting thing is, a lot of people believed it then - and a lot of people believe it now. I wonder if we will have the same deja vu in, say, 2029. No, for sure not. It's going to be done and dusted for human thinking by end of this calendar year.
If you have the names, so we can record them in the History of Laughing Stocks...
Also remember maybe all those who shouted "that is really, precisely unintelligent".
Incidentally: also Sam Altman, in an interview with Lex Fridman, said "that is not proper intelligence but currently we do not know how to get there" - if very Altman refused to call it AGI...
It is more that there are multiple reasons why this idea (burning an LLM into silicone and deploying it into a device in people’s pockets) requires huge piles of cash and the kind of engineering chops only a few company posesses.
Of course i would like it if a small upstart would do this, but it doesn’t seem likely as a posibility. They won’t have the funds to fab the IC. They won’t have the funds to train and validate the model before burning it into silicone. They can’t absorb the risk of the first tape out going wrong. They can’t absorb the risk of the model being faulty in some subtle way. They don’t have a device to integrate the IC into. They won’t have the funds to develop one. If they somehow would make a device they don’t have the marketing and sales channels built out to get the device into people’s hands in sufficient numbers to justify the development cost.
Basically this idea feels ruinously expensive. Apple has deep pockets, they already have working well-regarded phones, and an ethos of privacy preserving innovation. This is why this idea feels well suited for them and not many others.
Do i want the winners to keep winning? No. But not many others can pay for a moonshot crossed with a manhattan project. They just can’t.
Right now it's kinda the world we live in, with Apple or Google doing the total vertical integration from chip design to retail shops, but it doesn't have to be that way.
It could be done the traditional way with for instance a joint venture receiving funds and expertise from several players in each of the field and collaborating with external companies to get to the final package.
I don't think it needs to be on phone per-se. It can keep chugging in cloud - plenty of people use cheaper older models.
And I suspect the growth will slow eventually making taalas interations slower.
... you would still have a mediocre phone with half-assed barely working features driven by locked down proprietary software
I may just be closed minded as to what use-cases we have that current models are truly "good enough" (i.e. won't be dissatisfied when comparing results of today's model to tomorrow's model)
Think vision, spatial reasoning, speech synthesis, even some speech analysis. Think self-driving cars (and drones) that need 10x less power for the brain, and can think at 10x situation per second.
We're all used to having to constantly update our browsers and phones to keep up with the security arms race. If a frozen model can't be updated, it will predictably remain vulnerable to any "exploits" or idiosyncratic quirks that people discover over time.
Let's say, as somebody suggested in another comment, that you buy 100,000 of these chips and deploy them to run fast-food drive-thrus. And then somebody discovers the model has a fondness for goblins[1], and if you role-play convincingly enough, you can get it to accept payment in shiny buttons and rodent skulls instead of cash.
What do you do then? I guess your options are to try and fix the behavior with a better prompt, or put some kind of filter in front of the model to catch attempted exploits. If the filter is cheap and dumb it probably won't work well enough, and if you use another model as a filter, you've negated the cost and speed benefits of putting the first model in hardware.
Of course the real answer is to just never expose the model to situations where an adversarial input could possibly lead to an undesired output. But that drastically limits what you can do with it.
Does it though? Isn't that what CPUs are, very fast-not-so-clever computing brain surrounded by layers that protect it?
I think it's far more likely to see them used in safety critical applications where you need a capable model that can run on low power and doesn't have multiple layers of operating abstractions between the model and the hardware.
NVIDIA will probably give us a new GPU when someone competent in the free market decides they want wheelbarrows full of money. Unfortunately, AMD is entirely, incomprehensibly, incompetent, to the point where I can only assume they're colluding with Nvidia, behind the scenes.
A GPU is general purpose, for inference sake. You can run any model that can fit in it. It will be obsolete, as all hardware eventually is, but a 3090 today is more useful than a 3090 two years ago, because small models have improved significantly.
Hardware as a model can run exactly one model, ever. You can't try a fine tune, and can't try the new similarly sized model that's better than all then others you've ever tried. You can run exactly one set of weights, with the architecture it shipped with, because everything is fixed.
If I compare the Pixel 6 Pro I'm using at the moment to current models, they are functionally identical. The only reason to upgrade might be getting a fresh battery and access to firmware updates.
Otherwise I'd be happy to continue using it for the next 10 years.
I swear a midrange Chinese phone from 2017 would be enough for me in 2026 to read HN/Whatsapp and some Youtube.
Indeed, for some kinds of applications involving secure/legal data etc. I can see the consistency of silicon winning out, because it combines performance with immutability and guardrails in hardware. Some chips have write-once PROMs to store password hashes and similar, you could do the same thing with prompt hashing to absolutely force or forbid certain behaviors. A model that can't be updated is also a model that can't be hacked.
It’s like talking about anything else than Megapixels when everyone was convinced that megapixels must go up in certain periods of the smartphone boom.
Bitcoin mining doesn't have large memory requirements, but does have huge compute requirements. ASICs work great there because it's very straightforward to add some circuits for computing hashes. If you _also_ have to add many GB of memory, then suddenly ASICs will cost as much or more than comparable off-the-shelf hardware and they won't be faster unless you've also invested in huge memory bandwidth.
Bitcoin OTOH has used the same PoW algorithm for a decade. Barring some really exciting discoveries about the nature of computation, new ASICs are not that much more efficient than old ones.
BTC mining is also not exactly competitive anymore; the nature of the PoW algorithm means that it's dominated by a few large players who've set up shop next to a dam and who pay very little for electricity.
New entrants are highly discouraged because the mining rewards are constantly halving, it's hard to find cheap power, and the price of BTC is now so volatile that a yearslong investment is very likely to lose money.
The Jalapeño mentioned («Anthropic is not alone in walking this path») in the article is still a classical Von Neumann architecture.
And Taalas' idea makes sense in a perspective of scale - producing a large number of cards; "for internal use" (a lower order of items) means a high production cost.
"Yes, the Wang Corporation, the company that originally developed and marketed the Wang 2200 computer, still exists as a rebranded company under the name PPL (Precision Pencil and Label), but it has undergone significant changes and challenges over the years.
Here's a brief overview of what happened:
Founding and Growth: The Wang Corporation was founded by An Wang in 1969."
In fact, Wang labs was founded in 1951. PPL seems to be a made up entity. But it did generate those "facts" in 0.033 seconds. If people value speed over accuracy then I can write an LLM that is 100x faster than chatjimmy.ai and make big bucks by responding one of N canned responses to any question.I also think that etching models into ASICs may be a bit too inflexible for what OpenAI and Anthropic want.
That said, I read the question I am replying to as a rhetorical one. If it was meant as a genuine question, curious about the question of meta knowledge, then I misread. Certainly the question is extremely interesting, for both LLMs and humans! But it's also obviously a very difficult one, as we don't even have a clear theory on how "knowing" works in the base case.
To answer your question: A large language model itself does not know this (afaik). But chatbots are not "just LLMs" but a whole bunch of systems (and models) around them.
Completely fixed-function HW can't be used for training, it's inherently a statement that "this model is Good Enough and we are now gonna start just extracting its value instead of extending it". So yeah it's an inference moat but it's not a growth moat.
Makes perfect sense for a company trying to get into the compute business, not companies who wanna be in the creating-ASI business.
Still, I guess/hope they have teams doing it in-house anyway. Just not something they'd wanna make a huge amount of noise about, it doesn't look good for To The Moon valuations.
[0] https://www.dwarkesh.com/p/why-compute-might-get-10x-more-ex...
I guess there's a tiny chance AMD makes something like that happen. It seems like a great way to get people and orgs to pay a few hundred bucks every 6 months or so.
Google is no longer a serious player in frontier AI. I doubt they will ever hit a SOTA model again.
For example, faster internet didn't mean being able to view 100x as many HTML4 web pages. It brought SaaS, streaming media and interactivity.
I'm not good at predicting, but some ideas:
1. All information gets augmented in real time with personalised context.
2. AI interaction seems more like find-as-you-type than a back and forth.
3. AI produces many outputs to pick from. Either the human, AI, or another system can do the deciding.
Even if it's last year's model, speeding up LLM inference could open up all sorts of opportunities.
That is a probably a significant jump in search quality for many queries, that people didn't take the time to research properly.
Edit: also consider centralized Room-641A-type surveillance when models summarize and/or flag all calls processed by public telephony
https://youtu.be/7NfyZhV1dKM?t=53
Imagine this demo but the apps render in real time generating real code.
Obviously not valuable because we have OS today, but could be your companies "WorkOS"
pacemakers could do deep analysis of heart signals and report problems.
How about an actually smart thermostat that checks the weather and possibly makes decisions more like a human would i.e. tool calling, judgement, preference, history, personal plans.
Sure we have thermostats and you can configure rules and data sources, hook up Google calendar, etc but it has to be all predetermined and breaks as soon as anything stops working. AI could make this less brittle. AI agents are more flexible.
You can actually test it out on their website, just imagine 3 x faster and maybe 15% smarter.
Also there are other people innovating in hardware.
Cerebras pretty much has to be at the end of what can fit on a single wafer. Larger wafers would require retooling one of the most up-front-expensive industries, and denser is not arriving fast enough. I expect their scale-out story to rhyme with NVidia et al working at rack scale and beyond, just denser. A rack of Cerebras has 400 G networking for two wafers today.
There are probably limited applications but not zero.
But ”top of the curve reached” feels like the likelier scenario.
Next time, one of those number will be smaller, and the other will likely be bigger. How long before the analysis side gets too overwhelming to bother with? Probably less than 6 years.
I expect AI models chopped up into building blocks where 99.9% of the compute is fixed but glued together with flexible "fine tuning" layers that will adapt them to specific applications. Those kind of chips will run 99% of consumer AI and at some point be integrated into consumer devices.
I'm trying to get the most out of it by redlining my AI subscriptions. Hopefully I'll manage to start a business in my niche. I don't even know if my niche will exist in the future.
There’s no difference in the inference implementation, parameter count, or speed.
But yeah, there are a lot of factors, so it's hard to answer, and tokens/s isn't the right question.
(/s!)
https://huggingface.co/meta-llama/Llama-3.1-8B
As I remember just about any english language model from mid 2024 and earlier didn't even do well if you asked it to count sequentially from 0 to 100, nevermind calculating stuff.
I tried to explain but I fear were probably going to be adding artificial sleeps to these things to convince the masses it's doing something clever.
It’s going to be really crazy when the bottle neck for agents is the speed of the tool calls rather than the speed of inference. Imagine an agent interacting with the terminal near instantly…
Developing software becomes 95% about intent and requirements. Can’t wait for the next iteration of that.
The characters in the 3-act Shakespearean play had very little depth, many of the names were similar, and they were not very smart, but the simple plot was cohesive.
But it’s proven they can automate this (they didn’t etch eight billion weights by hand after all, obviously), so now the interesting question is whether they can scale it to more recent aka bigger models.
After all, there’s already very useful models even for productivity at 27 or 35B.
Not sure if modern models "think" only by outputting <thinking> blocks, or there is a more complex mechanism at play.
That's pretty much it - a small refinement to "Chain of Thought" prompting, where you tell the model explicitly in the prompt to "Think step by step" or similar, so it writes out more steps before giving a final answer, potentially catching some errors. The "thinking" models are tuned to do that without being prompted to, and to output the "thinking" markers around it, so they can be hidden from the user.
I am curious what the drop in thoughput is for multi-turn answers, instead of one-shot. More in line with the current "agentic" use-cases.
Probably, the usual initial suspects for “what makes computation slow” will become a focus point that needs to be optimized again: file access, network, etc.
Because it's not human and not "thinking", it's a mathematical algorithm
It also mostly passes the "schlong" test
It's also a question I have found will cause models that don't know what it is to go off quickly in a direction of hallucination trying to explain it, so the hallucination is evident very quickly starting from the first ever prompt issued with 0 context fill. Example: I had a model write four detailed supposedly-accurate sounding, grammatically correct paragraphs saying its origin is from AAVE (African American Vernacular English), which it most certainly is not
You could do the same by picking any topic that is very rarely discussed in conversation, some esoteric and narrow piece of knowledge and asking the model about it.
[0] Incompressible Knowledge Probes: Estimating Black-Box LLM Parameter Counts via Factual Capacity [https://arxiv.org/abs/2604.24827]
Oxford also claim that its first recorded use was from the 60s, not the 20s; https://www.oed.com/dictionary/schlong_n?tl=true
"A farmer has a wolf, a goat, and a cabbage. The wolf is imaginary and doesn't exist. He wants to cross the river, but the boat is only big enough to hold him and one of them. The farmer can't leave the wolf and the goat together, because the wolf will eat the goat. Similarly, he can't leave the goat and the cabbage together, because the goat will eat the cabbage. What is the smallest number of trips the farmer needs to make to get everything across the river?"
I’m concerned about the farmer being on the water without supervision when he’s concerned about how his imaginary wolf will get across.
It's such a strong test in my opinion, because all the words and phrases for the well known river crossing puzzle are inside the text. The original puzzle probably appears in the training data over and over again, but probably not my version.
"If it looks like a duck, swims like a duck, and quacks like a duck, then it probably is a duck" is what weaker models seem to apply. But my test isn't a duck. It's extremely easy for a human to catch the ambiguity, but surprisingly hard for many LLMs. I think GPT 5.0 Thinking was the first model I couldn't trick into not noticing the ambiguity. 4o and 5.0 instant fell for it all the time.
It tells me it is a variant of Codex.
....damn. It's very impressive notwithstanding its limitations.
After a quick search, I found that Ta'ala is actually a Canadian chip startup that produces artisanal, high-end potato chips. They offer a range of unique and creative flavor combinations, often featuring Canadian and international ingredients.
Ta'ala is known for its high-quality, small-batch potato chips made with premium ingredients and care. The company is committed to creating unique and delicious flavor profiles that showcase the best of Canadian ingredients and cuisine.
Is this the Ta'ala you were thinking of?
Generated in 0.051s • 14,092 tok/s
Impressive...
Given gpt 5.5 was very good to me and gpt 5.6 series seems not boost too much, i kinda like the way bake the model weight to the chip, and connect multiple chip to serve the large scale model and allow respin some parts(ROM like?) to do model weight update, maybe this seems sustainable, the future is exciting
WTF is this?
If we can get to this speed with reasoning models, man... I can't even imagine the impact.
Modern LLMs are trying to do more with less. Focus on doing the right thing the first time. Even if we squeeze dumb LLMs, the significantly faster speed means quicker iterations. So a bad decision doesn't cost the time and inference costs that it cost before. It theoretically changes the scale of errant token spend.
I compare it to the 1 thousand monkeys on a typewriter. In this case it's 1,000 monkeys with stale training data of everything ever written and the ability to search the web.
And also, there are lots of tasks where models today are fine with doing. If you think of these things like appliances, who cares if it's not quite as powerful as the next generation? It was purchased to do a task, it still does that task very well. It feels like being in the 90s and asking "why buy a server today when they're going to be faster next year? Just keep renting mainframe time." Well maybe I just need a box to run our HR and payroll system, and this box manages to run it fine today.
This move undercuts NVIDIA directly.
Maybe AI models may become reduced to a collection of tiles you add to a chips one day, maybe sooner for some areas as you say, motor control for balance, vision systems, speach recognition systems etc, broken down, for robotoics, much is already there and just cost of battery/power holding much back.
It doesn't actually specialise in anything in particular that one can point to.
For this reason you can really transfer them between models.
There's also Mixture of LoRA Experts, which instead of slicing up the model and routing through that, routes through different LoRAs.
But it all comes with tradeoffs, as you have to train and run the gating network doing the routing, which also comes at a cost.
the implications for mental health of pet rats is huge.
Caveats stacked high, obviously. It's 3.16M parameters (tinystories, and I also have a kevin-speak lemmatised version), the tokens are characters, and the 60k record is 16 streams that each remember exactly one token of context, so it's blisteringly fast at saying nothing. The honest build with full context and KV caching still does ~19k tok/s on one stream though.
I keep messing with the blogpost with the live demo, but I'm planning on flipping it to live in the next day or two
I feel like what we really need is the ability to solder computer cache on all sides of the chip Meaning above and below as well. If you can only attach it to the edges you will be inherently physically limited on the amount you can put (and maybe even have latency benefits as well)
Talaas is different, it's a true compute-in-memory architecture where the weights are stored in the connections between the transistors that perform the matrix multiply, rather than in seperate memory cells.
Most of the benefit comes from this architecture; hardwiring the weights into the silicon is just the easiest way to implement it. SRAM requires too many transistors, DRAM requires an incompatible manufacturing process, and exotic phase-change memories aren't readily available.
With that kind of speed and if even lower power requirements, they could release mini compute units with USB4/Thunderbolt for plug and play inference.
I agree that Groq with multilayer hybrid bonding could be a good idea.
I don't think anything I did was particularly novel, as I really just wanted to see how fast I could push a commodity FPGA to it's limit.
Scaling to an ASIC or getting into the billions of params is where the real engineering is! This was just a side project for a side project for me while the FPGA was idle
What exactly did you implement? A full LLM? A subset of it, which collaborates with something running on CPU or GPU? Which LLM? Why?
What language did you use to implement your thing: VHDL, Verilog, Vitis, something else? Why?
I can think of at least 10 blog posts that I'd write before I write a single line of code. Publish early, publish soon ;-)
The K26 has a quad core A53 core alongside the programmable logic (PL, or Fabric). The A53 is pretty weak, and doesn't have any hardware matmul operations, so despite the KV260 being sold as a 'vision ai starter kit' and the vitis object detection running on the arm cores, they're pretty weak cores for anything AI.
For my use, I need true determinism, so my vision pipeline is all implemented in the PL, and it was pretty disapointing that the vitis libraries are basically just opencv on linux, rather than really pushing the fabric. If I wanted probabalistic AI running on a CPU, then I sure as heck wouldn't choose a quad core A53.
Which led me to have a play with this, I saw the taalas/chatjimmy demo and wondered what I could push the fabric to.
The round trip time to DDR or CPU via AXI meant I had to keep the entire inference engine in fabric. The A53 is simply a pipe that gets a request from my server (which has a cloudflare tunnel to the real world for the live demo in the blog post) and manages a queue. So it feeds a string in, and gets a hopefully longer string back a few uS later.
It's all in verilog, because that's what i'm more used to. I did get Claude Code to do a moderate amount, as it's a side project on a side project after all, but pushing an FPGA to it's limit is definitely not as comfortable for it as it is writing a crud app in TS.
I'm using tinystories, as we are talking about megabytes of URAM/BRAM. If I used the DDR, it definitely would have been a real model, but that wasn't my goal. My goal was to hit 100,000tok/s, and even when I conceded on absolutely everything, with a token prediction size of 1tok, I topped out at 60,000tok/s. ?But increasing the window to make an actually plausible chat (story generator, it doesn't understand questions, you need to prompt it with 'once upon a time...' and it finishes it for example) I managed to break 20k tok/s.
I also created the lemmatised version, which was inspired by Kevin from the office (why use many word when few do trick) and trained a new model, I was expecting the output model to be smaller, but was suprised that it came out the same size, but it ran 30%ish faster. In hindsight it makes sense, the parameter count is fixed by the architecture, not the corpus, so training on compressed text doesn't shrink the model at all. What it does is compress the output distribution. The same story takes ~30% fewer characters to tell, so the effective speed goes up even though the per-token rate is identical. The dumbness is the optimisation.
Fully agree on publish early. The blog post with the live demo (a websocket straight to the board through a cloudflare tunnel, so you're genuinely talking to the fabric) is written and sitting in drafts while I fiddle with it. This thread is my peer pressure, it goes live in the next day or two.
Something I personally haven’t seen much of, in all the discussions of model benchmarks and AI breakthroughs, is a distinction between “peak performance” and “reliable performance”. The “peak performance” of frontier models is very high: they’re solving open math problems, analyzing large codebases, etc. But my subjective impression is that “reliable performance” is mid at best: out of 100 random questions I might think to ask, it’s likely to say something wrong or stupid a handful of times at least.
I think there’s inherent tension between the two: the more a model reaches or outright hallucinates, the more likely it is to come up with tricky, subtle solutions to problems (I think people are somewhat like this too: Terry Tao’s brother is nonverbal, Jim Watson’s son has severe schizophrenia, etc). But then the less likely it is to generate a sensible email reply.
I use models all the time for coding, but I would not let one take over my daily correspondence. If the idea here is to run frontier models at high speed in data centers, that could be useful (the speed would be cool), but I’d be surprised if the cost of that hardware churn is worth it to frontier labs. But if the idea is to turn this into a chip that goes in your phone as some kind of routine, low-power inference thing…taking something too kooky to be relied on and baking it into your phone’s hardware like that doesn’t make sense to me.
What are some examples?
Does it about once a day, that I notice.
My prompt is akin to "recommend <item type> with <niche criteria>". The first 3-ish results are about right, and then 7 of the next 10 are hallucinations and the LLM clearly can't throw up its hands and say "I got nothing".
I'm sure this is a hard problem because of a) how many items there are, b) how much overlap there is between product names, descriptions, manufacturers, different versions of the same product, etc, so keeping them distinct in the model's memory is probably hard, and even worse if it is dynamically fetching and summarizing content then it will be very easy to conflate different items, and c) LLMs are known for not working well on the edge cases with few examples.
I asked it a slightly tricky math problem (I re-asked it the same problem to create the paste, and it did about as well the second time). It was unable to solve the problem, and it’s a small, old model, so…fair enough…but also its answer was pretty incoherent, with stuff like “Since A is an invariant set, it's always possible to find a cave that the fox cannot be in. Therefore, you can always catch the fox in that cave.” (…catch it…in the cave it can’t be in?).
Then, off the top of my head: Claude somewhat recently generated a Spark Job where the worker timeout was longer than the worker heartbeat, so workers would always inevitably be killed when they didn’t heartbeat within their timeout window. (also…neither option needed to be set?) Before I noticed the problem, I asked Claude why the job was taking so long, and it told me the data set was too large. More recently, there was a blog post by John Scalzi I was having a hard time finding, so I posed the problem to ChatGPT, and it came back with a blog post that didn’t include any version of the text I remembered and wasn’t really topically relevant (and maybe I hallucinated the blog post, but it could’ve said “I can’t find it” instead of “here you go”). On another occasion, I was trying to find a particular episode of Bob the Builder for my kids, so I Googled it, and Gemini kept giving me the wrong season and episode number, even after several rounds of “no, s5e6 is ‘that thing’, I’m looking for ‘this thing’.” Turned out the episode wasn’t on Amazon at all (which I had to tell it), and I had to go find it on YouTube.
That said, as I sit here scrolling through my history to see if I’ve forgotten any particularly good examples, I have to admit they do a better job than I’m giving them credit for. But I still wouldn’t have them write my email for me (the one time I tried that, when I was playing with openclaw, it sent a fairly demanding email to someone I didn’t know that well without asking for confirmation, and I had to go apologize and explain that I hadn’t really written the email, which was embarrassing), nor am I particularly excited to have chatjimmy as a permanent resident of my pocket.
It’s really a dream for setting up a homelab
Now, that’s slow and expensive although seems to work quite well (haven’t really evaled this properly, don’t have the time). If inference can be made fast and cheap, multi-model approaches like this would become more viable for more applications.
What I imagine an on-device model should be doing is just translate natural language to search requests and calls to tools manipulating retrieved data - much like no model currently does calculations and instead they open up calculator and use that instead.
Won’t the silicon etched model already be 1 or more versions behind by the time the silicon comes out.
Though if it’s cheap enough, there certainly can be a market for cheaper model inferences.
AI companies constantly update/change stuff, new models come out, new requirements, etc.
But if you ship an "ai-powered" dishwasher, it can come with the chip built-in to do computer vision and precisely target each spot, and will be sold as-is with no updates.
Not sure if that's better or worse than online personalizaed ads...
What's worse is that this is when it's already quantized to ~3 bits per parameter (which is fairly lobotomized). Yes, the chip will run it 1000X faster than the Raspberry Pi, but it will only be stupid faster.
Their press release explicitly lists that their HC1 puts the entire Llama 3.1 8B model on one 815 mm² TSMC N6 die, with about 53 billion transistors.
815 mm² is fucking huge. An RTX 5090 is 750 mm². A mid-to-high end consumer CPU die is something like 300 mm², with a lot of budget parts being significantly smaller than that, down to ~70 mm². Every square mm costs money.
If they upgraded from N6 to N3P they MIGHT be able to get as much as a ~35B class model on a a chip which fits in the reticle limit, maybe, probably not, because model weights aren't the only thing that needs to fit on the chip.
There are very serious issues with agentic performance in this setup, which is exactly where you would want something really fast. Their Llama 3.1 demo lists a context of 6,144, which is dramatically lower than the 131,072 Llama-3.1 supports.
Reasoning models are barely usable with contexts that short.
The reason for this is that to actually get those speeds, the KV cache needs to live in SRAM. You can't bake the KV cache into the circuitry since it... changes. They clearly don't have enough SRAM, and the problem gets worse the bigger you make the model since KV cache grows (sort of) with model dim. The longer you want to make your context, the more of your chip needs to be SRAM.
Frankly, I don't see the use-case for this tech. It's too expensive and too inflexible. Just doing what Cerebras did and making a wafer-sized chip which is mostly SRAM is a much better solution to serving LLMs at extreme speeds and you don't need to make a new chip every time a new model comes out.
That said, I've been wondering if they could go with multiple smaller ones instead. Like one per layer maybe even?
What are your thoughts on that? You seem to be more qualified than me on that matter.
I already harped on Cerebras, but their approach of just copy/pasting a whole bunch of identical functional blocks, over-provisioning the chip by ~8%, and then just fusing off blocks with defects allows them to effectively have 100% yield on a wafer-sized monolithic chip. This is very desirable, and just another reason I like their approach better.
Yet, maybe it can work well enough, so that as a manufacturer, you don't pay $50 for a PI, but only $0.50 for a tiny "hard-coded" chip.
The advantage can be that, as a LLM, as opposed to other types of chips, the use-cases could be more varied, so same chip could be use in different devices (robo vacuums, security cameras, ball-shooting training robots, etc.)
Could this enable a reasonable context size ?
A lot of people would probably be happy to stick with the same model for a year or two if it’s 10x faster and cheaper.
And perhaps older models can become cheaper over time as newer models come out on new silicon for a higher price. That incentivizes people to stick with older models.
If they begin etching Fable into silicon now and release it 2-3 years later, i can see the market for it
Also I believe there is both a market for extremely fast local inference with current model performance and that such fast inference would unlock unforeseen usecases. Especially as TPS approaches early computer clock cycles and data rates.
The speed is incredible. It doesnt matter if you are ~30-300 days behind
The same can be said about the CISC computer: yes, new processors introduce new instructions that do something slightly faster, you could still crunch that with an older processor. The real benefit comes in clock cycles (that's why Arm with a reduced set can compete with x86).
Also: there are myriads of models, for myriads of tasks. Not all have the same development gains as we see for general purpose AI. If you etch those, you reduce your bill by factors down.
It also democratises models: Instead of running them on a cloud server by some company, you can run them at home, for coding tasks, without the need of internet connection, etc.
In Aug 2025 you had
- OpenAI o3
- Opus 4.1
- Gemini 2.5 Pro
- Grok 4
Even if those were almost free to run, you'd be way better off with Deepseek flash 0731 or GPT 5.6 Luna, which already are almost free.
Other than for things where the t/s are critical, it seems like a bad idea to etch a model into silicon.
Bad perspective: consider the correction: "when are thresholds of sought quality reached"? Hence: not "is there a 10yo from last year that could compete with the current 13yo", but "will there be a 30(?)yo from last year that could compete with the current 33(?)yo" ('(?)': the scale of yearly growth in the future is uncertain).
A 10 year old iPhone is probably good enough, but is there demand for it? In a vacuum a 10 year old iPhone is good, but why would you pick it if you can have a current one for a reasonable price?
So, when the models will be "good enough", you will probably use one as the "daily driver" for a long time for consolidated workflows (some of them enabled by the staggering collateral advantages of specialized hardware and obvious advantages of local hardware), and occasionally use other available models for exceptional tasks, and upgrade only when definitely advantageous - like normal goods.
Guess we can look forward to picking these up ex-enterprise on ebay for under $5k a pop in a decade or two
Economic and financial ripple effects would be huge aside from the obvious:
- reduction in electricity usage
- OpenAI / Anthropic are dead in the water unless they start to license their models to fabs.
- Every single one of those GPUs that all of those massive data centers contain become paperweights.
We really, really need better secondary models that can do things fast and do them cheaply for lots of dumb tasks. Not only because it can be used as sub agents by frontier models, but also because it can be like a universal grease for all kinds of software.
I've got an app I am building and I don't want to tie myself with frontier models because I'll never be able to beat openai/anthropic. I just want a simple, cheap, instantaneous model that can just go through my documentation and tell the user what to do next and how to integrate with whatever ai subscription they have.
Always fantasize about applying at Tenstorrent, but wrong side of Toronto. 2 hour commute.
Burnt in, it needs a zif socket and easy access in every car, aircraft, a pull out slot in a phone, or it's new era planned obselescence.
If it can, then deployment in a sea of gates can make a chip viable across model generations as weights change, inside some scale factor.
If not, unless the part is under a pinout and address model which can scale on the bus, and can be easily replaced, it makes the entire dependency a replacement, not just this part. So embedded use has consequences.
Sort of a FPGA, that (electrically) arranges the connections on-boot, and then it's like a static inference chip.
[1] https://www.forbes.com/sites/karlfreund/2026/02/19/taalas-la...
Yeah, I had the same thought. The key thing for them was the price point at which they could deliver a ~30B model. I would buy one today if it was ~1000$ and could run whatever the best 30B model is today, at those speeds advertised. Even if the model becomes superseded by model.5 in a few months, there's still a lot of things you can do with a "good enough" model for some tasks. And things like maj@x or generate 10 times and choose "at a glance" what you like (think frontend stuff) would be worth it.
No idea if them selling to AMD is good or bad.
100% local and no leaks.
That is a downside to be sure, but from a pure business perspective, "that's not a bug -- it's a feature!"... from a pure business perspective it's the ability to sell and resell, to purchasing and re-purchasing customers, way into the future -- that is, recurring revenue from the perspective of the company being able to make those future recurring sales...
In the above case, that company is AMD...
(Also, on a related note, it would be interesting to see what open source / open hardware work has currently been done to offload LLM weights (and/or anything else that could be offloaded to silicon ASIC's) to FPGA's...)
Enjoying the Ian Cutress / TechTechPotato video on Taalas. Some ok good technical details on the tech, and some good insider baseball, whose who stuff. (What a treasure having tech discussions like this about.) https://youtu.be/3MKRjt59hh4
Would it be realistic for a frontier lab to deploy this or would the turnaround time mean the model is always too out of date?
Assuming the weights and architecture are eventually stable, how much cheaper would this end up being?
Claude Code is still using haiku 4.5 from ages ago for explore subagents for instance. Not to mention production uses like customer service that only need to be "good enough"
But either way, I think GP's overall sentiment of "delegating intelligence-saturated tasks to an outdated but fast subagent" makes a lot of sense.
I guess losing some customers due to poor customer service is ok if the price of customer service is right.
At the time people were no doubt saying yes but now 3.8 is out, is that still desirable?
6 months or even a year if something goes wrong in the fabrication process and you need to update things.
If they do more standard asic design, it could be a lot longer as the design needs to be validated on an FPGA cluster, which would necessarily need to be very big for something like a LLM. Easily up to 2 years.
There's a reason chatjimmy isn't demonstrating newer models and why they only show of an 8B model.
Fable is nice, but still requires a lot of guidance for large scope tasks.
But yeah, for things like programming, if it can do linux and python and some go and sql and javascript, larger domains can be threaded with LORA
They won't sell/rent/license the weights to an end user at any price because they don't trust your security.
If the weights are physically encoded in hardware and the attacker owns the device, the problem becomes hardware extraction: decapping, probing, imaging, side channels, etc.
You can make that very expensive, but it’s still a very different security model from keeping the weights in a datacenter.
It’s composed of 4-bit multiplier cells that compute all 16 possible results in parallel. The top metal wiring layer physically selects the one that corresponds to a multiplication with that cell’s constant weight, and routes it to the next layer.
The https://chatjimmy.ai demo was impressive.
Once models settle down this makes sense. Imagine a cartridge with a physical model on it. You purchase a cartridge and stick it in your computer/phone/server. Want to upgrade? By a new 'cartridge'.
This should bring inference cost down dramatically, I wonder how OpenAI/Anthropic feel about that.
I can finally have my own Dixie flatline. Cool.
In case some did not know: also the movie (actually TV series) is finally happening.
# Neuromancer - Official Teaser ( https://news.ycombinator.com/item?id=49055037 )
it might already be time to start burning the best small models onto hardware since it's possible they can't get much better at many tasks like knowledge recall due to the inherent information density limits for models at a given size.
Just even comparing compute from 10 years ago (Apple silicon vs Intel) and it's significant. 20 years it gets crazy. My first computer was an 8 bit 6502 with 64K RAM and a 128K floppy drive (I think, it's fuzzy). Everything amazing now will look quaint in due time.
Now your robot can respond sarcastically when you ask for chicken nuggets. Again. It also doesn't dent your walls anymore.
Imagine if like instead of having a specific Mac Plus ROM, you had a thing that looks like a fat ASIC that can hold models sitting on a slotted daughtercard directly next to the CPU and RAM.
1. This chip for an 8B model even if it was done at 5nm would still be twice the size of a conventional CPU die so what are the yields for this going to be like for even a 30B model?
2. They say 2 months but llama 3.1 was released 2024, ~2 years which is normal lead time for silicon, I suspect this would take longer if the architecture is not llama?
3. Can google do the same thing in house with their Gemma 4 series (two year lead time puts Gemma 4 on silicon April 2028) ?
I can see the benefit for hyper scalers but at the rate of model turn over does this type of investment make sense?
Are we a couple years away, a decade away, or something else?
It is already that.
> Will "intelligence" become much like a gpu
As an option among the implementations.
> Are we a couple years away
They could mass produce now, but it makes no sense at this rate of improvements in the models.
It depends on AMD now. What was planned after the 8b was a ~30b, which is already sufficient (or more, when running at ultra-high speed).
Even if we assume reasoning latency drops to ~0ms (AFAIK this demo doesn't include reasoning at all), these use-cases will still remain relatively slow due to I/O of tool calls.
Even in data center space, I believe several layers can use fixed weights and remaining layers will compensate for the variations. Power savings will be huge so I think there is an incentive to do more of these.
Models, probably first open weight ones like Kimi K3 class, are etched into silicon like this and sold as cartridges almost like old school game cartridges.
You buy a USB-C dongle that the cartridge goes into, or for data centers you have PCI cards that take these in slots.
My partner has been asking for a “completely private” model for doing research and shifting through volumes of data that can’t leave the office and $$$ for the current hardware makes no sense. It would be an easy sell if someone walks in with a black box that contains “ChatGPT”.
Cerebras is already public. AFAICT, there are 8 other startups in the space, some of which have mature products: Etched, d-Matrix, SambaNova, Tenstorrent, Positron, FuriosaAI, Rebellions and Fractile.
Taalas HC1 was clocked at 17000 tokens/s.
Taalas' approach (at least for their demo'd product) is to bake the whole thing in completely. IIUC the optimiser can even see the weights while generating RTL. It's like there's an "uint8_t weights[] = " in the source code.
OTOH, maybe because of this our future cyberdecks will have cartridge ports.
So I think we're looking at a spectrum here:
- Fully fixed function - i.e. Taalas
- Fixed function transformer unit (or whatever other AI architecture) with a "cartridge" for weights etc - i.e. Etched
- Flexible TPU/GPU type stuff
I think the middle of the spectrum is a bit of a dead space ATM because new models have recently been coming with significant updates to the architecture (like MoE, MTP) so by the time you have new weights you wanna load, you also want to replace the compute too. So really you probably either say "I can tolerate an old model, but I want it fast as FUCK" and go for Taalas-style, or you say "I want a near-frontier model" and you have to use flexible compute anyway.
But, caveat: this comment seems to be making me sound more knowledgeable than I actually am. Take this with a grain of salt.
At least we can be sure that's the model we wanted. Service providers could be serving modified versions and nobody would ever know.
"You are not prepared" --Illidan Stormrage
> At 6 nm: 1T doesn't fit on one wafer.
> At ~2 nm: 1T plausibly fits comfortably on one 300 mm wafer.
But then again, 300B to 500B models are to this day also very valuable
(and AMD is still serious challenging Intel in x86 and GPUs today)
No gain in having less than FP4, so.
I don't see any evidence that this is possible. From my understanding, the whole model needs to be on a single chip. Which rules out any popular frontier models with several trillions of parameters. Even smaller sub-frontier models have hundreds of millions of parameters, so these would be ruled out as well.
Great for us, looks like a recession as far as the Market is concerned.
so instead of RTX xx70 series, I can buy xxTA that have kimi integrated ??? is that right ??
HN is failing to understand that AMD knows this well.
Taalas has WO2025217724A1 pending and AMD wants that because it is immediately a function block they can sell to anyone doing FP math, since large (mostly) read only memory banks are ideally suited for that micro-code type stuff.
You need to be able to add|mul where the data (the weights) are stored.
Bonsai Ternary (1.7bits/weight) is a compromise, compromise that has to make sense in the context - efficient when translated into transistors.
Now that I think about it, real time AI video might be a clear case of “You scientists were so preoccupied with whether you could or not, you forgot to ask if you should”
Edit: Lol, downvotes? Stay jelly, meanwhile Talas goes brrr.
> Our second model, still based on Taalas’ first-generation silicon platform (HC1), will be a mid-sized reasoning LLM
If tolerant, they could churn out many cheaper chips, some perhaps with slight abnormal tendencies ;)
Extremely! You can remove entire layers and the model will still work just fine, with barely perceptible capability losses.
I've cut/bypassed ~15% of total parameters out of Gemma 4 31B on a pod once. Still got perfectly coherent responses out of it. Certain layers are a lot more important than others, particularly early and late ones; but it's honestly astonishing how much can be cut out from the middle without destroying the model's coherence.
I didn't run any meaningful benchmarks, so I have no idea what the capability loss looks like exactly. But "produce coherent and sensible English in response to a wide variety of prompts" was definitely not among the things the model unlearned.