Can anyone comment on the performance of this hardware? How does it compare to state of the art, human-designed hardware? Is this actually an improvement? (To get to recursive self-improvement, you first have to improve at all.)
In essence this is the simplest unit of an entire AI chip. The more complicated units of AI ASICS are actually the periphery, especially around PCIe and Ethernet and the sub-systems that link many AI ASICs together to move huge amounts of data around ultimately to each TPU.
So its missing ALOT
You should only upgrade when the bottleneck becomes memory, right now the bottleneck is still on compute until you are at 100% max perf