https://www.bangkokpost.com/tech/2556324/nectec-agencies-rol...
Whether that focuses in this manner is unclear, but we should expect forward looking governments to train LLMs to their own taste.
There's something like 6 million Danish-speakers for example.
(I think this is how Google Translate went about things in the past, making translations into some languages come out very formal, as most of the training corpora for that languages came from internal and international official documents.)
Countries that have been on-line for a while may also have discussion boards and comment-bearing sites that are entirely unknown to people outside those countries, too.
Maybe multi-step approach would be in order - try to get half-decent a translation system working (an OG LLM, or an LLM trained to fix grammar in translations outsourced to GPT-4), and then synthesize training data for your main LLM by having GPT-4 (or its successor) generate tons of English text of all kind, and feeding it to the translator system.
(There's a limit to synthesizing training data, beyond which it'll only amplify existing patterns, impacting model performance in bad ways - but I don't know how easy it is to reach it.)