They also suffer from the global optimisation problem for layout of calculations so compile time is going to be insane.
Their WSE technology is also already obsolete - Tesla's chip does it in a much more logical and cost effective way.
More than that arguably. CUDA cores are more like SIMD lanes than CPU cores like cerebras's usage of 'core'. Since they have 4 wide tensor ops on cerebras, there's arguably 3.6M CUDA equivalent cores.
Your number is off by 64x.
It can do 125 petaflops at FP16
https://www.tomshardware.com/tech-industry/artificial-intell...
And, 9 trillion flops per core in 4.4 million transistors per core. That sounds a bit too good to be true.
You would need 64 of these to get 8 exaflops.
https://www.tomshardware.com/tech-industry/artificial-intell...