Climate is pretty much my #1 concern about the world, but LLM use of energy is really really far down on the list of important actions for climate.
First and foremost are removing roadblocks for deploying existing technologies for clean energy, and speeding up the necessary supporting infrastructure such as transmission and market policies for choosing cheapest possible solutions (over the objections of dinosaur execs that choose last century's solutions). Then the big hard to decarbonize parts of industry like cement and steel, as well as deploying electrolyzers to get ammonia fertilizer production switched over to carbon neutral production rather than from fossil-generated hydrogen.
Reducing energy consumption is important for advancing AI in general, but ultimately all its energy consumption will be from clean energy sources anyway, and the switch that needs to happen is that switch in energy sources. Reducing energy use by 2x or 10x is not good enough, we must change the sources fundamentally.
Very occasionally I get the feeling HN is entering the /. phase
So if you had a DIMM with 16 of these chips, you would already be on the same bandwidth as HBM. 96 DIMMs and you get 40TB/s memory bandwidth.
Just ten years back, squeezing ML models onto microcontrollers sounded completely insane, given their tight memory and power constraints. We've seen NN compilers developed, game-changing techniques like quantization, pruning, and graph-level optimizations pruning. This allowed deployment of ML models in microcontrollers with a newly developed framework like TFLite Micro.
Also, speculative basic research isn't comparable to adding more RAM to a graphics card. You can't substitute one for the other.