Amazon has their own chips for inference and training: Trainium1/2.
Memory hierarchy management across HBM/DDR/Flash is much more difficult but necessary to achieve practical inference economics.
Everyone else went the CoWoS direction, which enables heterogeneous integration and much more cost effective inference.
I think while being fast, cerebra’s probably not very economical in fleets at scale.