Why the fuck was this downvoted.
Very occasionally I get the feeling HN is entering the /. phase
Very occasionally I get the feeling HN is entering the /. phase
So if you had a DIMM with 16 of these chips, you would already be on the same bandwidth as HBM. 96 DIMMs and you get 40TB/s memory bandwidth.
Just ten years back, squeezing ML models onto microcontrollers sounded completely insane, given their tight memory and power constraints. We've seen NN compilers developed, game-changing techniques like quantization, pruning, and graph-level optimizations pruning. This allowed deployment of ML models in microcontrollers with a newly developed framework like TFLite Micro.
Also, speculative basic research isn't comparable to adding more RAM to a graphics card. You can't substitute one for the other.