but now I think they probably have bitten more than they can chew.
Apple already proved with their unified memory - that as long you have the capacity you can run capable models locally - thereby goes demand for inference if everyone is running some model locally.
For training - Chinese models have proved that you don't need the latest & greatest in Nvidia hardware. Same as TPUs.
only time will tell.