Synthetic codebases are certainly an option, but if top models remain closed I don’t see them building datasets for every new language.
But compared to the immense amount of effort that goes into convincing a critical mass of humans to learn and write about your new language, and using _that_ material to train an LLM, I think it's fair to say things have gotten easier, not harder.