Surely the two leading candidates must be (a) the model is just not that good, or (b) it is misaligned in a pretty obvious way that can't be swept under the rug.
The fact that OpenAI, Anthropic, SpaceXAI and 3 different Chinese companies were all able to train big models without these issues, yet Google could not, seems shocking.