Seems like a solvable problem though by generating synthetic data that’s guaranteed to be accurate through linters, compilers and tests.
If it’s ~9 figures to train a frontier model, it seems like training on a new language could be a rounding error if it was a priority.
I could see the appeal of a new agentic-friendly language that’s focused on minimizing tokens. The standard library could be massive with no concern of making the language easily readable or learnable.