Mythic's analog compute-in-memory architecture
mythic.ai
mythic.ai
A 2026 EE Times article [1] refers to "compensation" and "calibration" techniques.
[1] https://www.eetimes.com/mythic-rises-from-the-ashes-with-125...
But like NAND, there is room for, I don’t know what to call it, “quantization”? You get a few values out instead of just binary.
My guess is, like other issues of precision, errors can get out of hand if you are not careful.
But being an EE in another life, I can tell you the power waste and slowness of ALUs is kinda wild.
I definitely believe that Gaming on CPUs is like ML on GPUs. Slow, and waiting for something more appropriate to come along.
Don’t know if these guys are the ones to do it (and they are, ahem, not alone). But someone will deliver 100x to 1000x boost either in speed, power efficiency, or both.
Within one generation, evidence of significant differences seems weaker. Sure, manufacturing had loose tolerances, you could end up with a "faulty" chip, but most SIDs sounded the same. You could also end up with a 6510 that crashed after 30 minutes.
I wish they would have done what Taalas did with chatjimmy.ai and just directly host a model for us to view, rather than just claiming it’s 50x faster than Nvidia/groq. Their claim is specifically for a 1 trillion param model. So they could have just grabbed GLM 5.2, or similar, and hosted it.
Btw, can guarantee that they are not ready to demonstrate that yet. They’re using 2D FLASH with 30M weights per die [1], so to get to 1T they will need… 33,333 dies. Interesting scaling problem to say the least
A confusing thing is that the goal is tackled through a number of proposals... Why Vanguard if they have Mead? If Mead, how to get the memory integration that are explicit on Vanguard?
And maybe one can use NAND or NOR flash with a different sort of controller to do analog computation, and maybe one could convince a flash memory fab to build it for you.
But there is no mention on the site of how they expect to manufacture the thing.
They are using GlobalFoundries’ 28nm node for the floating gate transistors, afaiu, which they then bond onto a TSCM 5nm digital I/O wafer.
Analog computation’s principles are sound. It’s mostly doing matrix vector multiplications though. The rest is digital.
Which means connecting ~350 chiplets to run a Qwen 3.8 27b and over 30000 chiplets to run Qwen3.8-2.4T-A95B. Cost? Space? Feasibility?
Edit: wrong values, lost a zero...
Edit: seemingly, the M1 is only part of the whole need. With the M1, you would run a feedforward pass of the NN but use the rest of the Von Neumann architecture to manage the data. The pass in the M1 will be lightning fast, the rest still a bottleneck. The M1 is almost explicitly not for LLMs.
3D NAND flash can indeed routinely store hundreds of GB per die, so that’s proven. The question is about all the peripheral circuitry needed. Each attention block would need its own KV cache (i.e. SRAM or DRAM somewhere), plus DAC/ADC inputs and outputs, unless they figure out a way to keep it analog all the way (really cool but unlikely).
I think this field is very interesting, at least from a technology point of view. Whether it works out or not will sadly be a matter of economics more than physics I fear.
Very unlikely for the connection to the cache RAM, seemingly impossible if we want to get an articulate output :)
> they do claim to have a “Mead” design [1] designed to run GPT-3 [ - i.e. to hold a 175b NN - ] in a single chip
Careful: that is the /intention/, but the chip is just the storage (and CiM) for the NN and other parts are missing - explained in the Vanguard, not explained in the Mead.
> Whether it works out or not will sadly be a matter of economics more than physics I fear
Yes, but:
-- what is not enabled today may often be tomorrow through advances, esp. in the economy of production;
-- a very great point about this product seems to be that the chips can be rewritten, the NN is not etched and static (cpr. Taalas): that makes the practical implications extremely relevant, the demand would be "screaming mob" like;
-- we have to go in that direction (of NNs in CiM) anyway, so it's just a matter of time, effort after effort we will get there.
Btw, I've seen some roadmaps in which foundries were heading towards integrating ferroelectrics and floating gates to accelerate switching. The two technologies are not mutually exclusive really, so I think we might see something in the relatively near future. The crazy memory market may help!
Why? I have not seen anything outlandish for a NN implementation (vs a NN simulation).
> can pull off the error correction
There lie the issues that have not been mentioned, the solutions not explained. Analog computing means: * costly digital-to-analog at the input and analog-to-digital at the output; * sensitivity to environmental conditions such as temperature; * signal dispersion hence the need to boost it in the path.
Maybe checking the patents they registered?
Flash NAND can routinely be bought with 4 bits per cell, perhaps even 5 soon (QLC and PLC drives). Since it has been proven to store 4 digital bits at production scale, I’m willing to bet an analog architecture running an LLM should be able to yield the performance analog of a 4-8 bit quantized model. Where in that 4-8 range is pretty crucial, but it depends on the specific design.
[1] https://www.nature.com/articles/s41928-023-01010-1 [2] https://www.nature.com/articles/s41586-022-04992-8
The technology that could run LLMs should be the "Vanguard", but as the homepage says, "the M1 (scope: Edge/Cameras/Drones) is there, the Vanguard should be a reality in 2027".
Bottom of page linked from HN (currently https://www.mythic.ai/) indicates they're hoping to demonstrate something that could that in 2028 or later, and both Nvidia and Cerebra are looking at 10x'ing models to 10T+ plus in 2027.
So they may never catch up on LLMs.
They're a good fit for the companies they're working with and have taken investment from, ex. Toyota, that aren't doing LLMs.
https://www.mythic.ai/vanguard
Seems big, IF true
You generally don’t need error correction. Instead, you do a retraining of sorts where you basically fine tune the model on the target chip architecture. Because these systems are so complex, errors tend to localize without affecting overall output.
My understanding is the hardest part is actually noise in the analog system.