And the question is what do programs that max out Ironwood look like vs TPU programs written 5 years ago?
Just because it's still called CUDA doesn't mean it's portable over a not-that-long of a timeframe.
You are aware that Gemini was trained on TPU, and that most research at Deepmind is done on TPU?
The tweet gives their justification; CUDA isn't ASIC. Nvidia GPUs were popular for crypto mining, protein folding, and now AI inference too. TPUs are tensor ASICs.
FWIW I'm inclined to agree with Nvidia here. Scaling up a systolic array is impressive but nothing new.
a generation is 6 months
* Turing: September 2018
* Ampere: May 2020
* Hopper: March 2022
* Lovelace (designed to work with Hopper): October 2022
* Blackwell: November 2024
* Next: December 2025 or later
With a single exception for Lovelace (arguably not a generation), there are multiple years between generations.