We let models localize into 16 languages. How we made it read native.
reelang.com
reelang.com
I was hoping this was about using an LLM to produce text in the house style. That's also a task they should be good at, but that clearly doesn't happen by default.
Here's a two-page public announcement from Ironhide Games, who make Kingdom Rush. They're apologizing over a direction they chose for the upcoming 6th game in the series:
https://cdn.discordapp.com/attachments/524968371494846484/15...
https://cdn.discordapp.com/attachments/524968371494846484/15...
To me the main message is twofold:
- "People complained, and we're going in a different direction."
- "We had this translated from Spanish by an LLM."
The 11 paragraphs of text are superfluous, because... (1) anything you could learn from them you could also learn from the two-bullet-point summary on the second page; and (2) the text style clearly doesn't reflect anything that the spokesperson for Ironhide Games wrote. This is a document written specifically to prevent any information from coming through other than the two bullet points in the summary.
And that's a style that companies often invoke on purpose. PR statements are vetted for this kind of thing, which is why everyone hates them. But it appears to be the only style that LLMs want to produce, whether they're making a PR statement or not, and I feel sure that almost none of the people using them for translations actually want this.
I also feel sure that the "LLM style" is specifically intended by the LLM providers, and that it wouldn't exist if they didn't spend enormous effort causing it to.
In my experience, there's two key incredients to get to a good results (with humans OR LLMs doing the work), I've worked on plenty of sites that had human l10n before LLMs came about.
If you're interested in the actual prompt we're using for our voice that includes instructions on how to avoid sounding like bad LLM copy, I've popped that here in a Github Gist: https://gist.github.com/tobyurff/2b461c259c34dfc6758a3932982...
Hope it helps, I might turn that into a blog post one day as well, seems like this could be useful to others!
So it isn’t randomly translating “upgrade" as "upgrade to." The context tells you that "upgrade” means "move to a different plan," rather than "update the software."
You're right about the random capitalization though - it wasn't intentional and we'll edit that bit to avoid confusion.
I would leave it at that, but then I consider what you're actually using it for...
> two-person founding team running a language-learning app in dozens of languages, so the copy is read by people who are unusually sensitive to it.
They aren't "unusually sensitive to it". Language learners don't fucking know what native speech in the target language sounds like, they're learning it! Nor does your two person team. How could you possibly claim to objectively evaluate how good translations are in dozens of languages you don't speak? 58, according to your site! This is beyond scummy even for the usual AI boosters, and will deceive people trying to learn, who don't know better, with low-grade slop. Disgusting.
I don't think this is true. LLM translation will generally satisfy the core goals of (1) communicating the information that's present in the source document; and (2) not communicating information that isn't present in the source document.
There are some style issues -- I have a comment sidethread complaining about them -- but those are very much second-order effects compared to "can I get a foreigner to understand what my document says?"
Compared to who? Compared to a professional with a language degree, with a feel and love for both languages and an understanding of the subject area, LLM translation is shit, no question. But such professionals don't get hired by startups to localize their websites, too expensive. When comparing to paid- by-the-word contractors who do website localization, I am not so sure that LLM is worse - because in practice too many biologically-human-localized websites and apps are translated horrifically.
And I think using machine translation, and being honest about it, is completely understandable for some tiny startup that just wants their website localized. I think it is completely and totally unacceptable for a language learning application, though.
Most CAT tools do vet for ML/LLM translations these days, so "cheap" human translation is mostly ML-assisted and human-edited, but my entire point is that they're not to blame, the majority of translation fails happens because of the lack of context, not because the human didn't try. It's about setting the site and l10n infrastructure out in a way that gives them the freedom to translate in the way that makes most sense for their language without being forced into "English in a costume".
It is unfortunately a fact that one can be both things: a startup and a language learning app. But my entire point wasn't about LLMs vs. humans, it was about laying out the architecture in a way that enables great l10n, whether we do that with LLMs today or humans tomorrow once we can afford it.
Your literal headline claim is that LLMs are better than humans. With a caveat attached, sure, but you're immediately setting the tone with that.
> It is unfortunately a fact that one can be both things: a startup and a language learning app
You know what you do in that case? Teach the languages you know! It's arrogant and greedy beyond belief to think you have any ability or right to "teach" people 58 languages. If you were genuinely interested in helping people learn and not making a quick AI startup cash grab, you'd start with the language pair you're most fluent in, then if that has any success, slowly expand your offerings once your concept is proven, hiring more people as you can. Of course, such a rational business as that probably doesn't sound as juicy as "we used AI to SOLVE TRANSLATION" to the investors you're hoping to attract...
But the reality for us (and many other startups) is that the options are no l10n and locking people who don't speak English (well) out of our product or having one that we - at least - do our best to make feel pretty native. We also built a version in Simple English. (I wrote an article sharing that process, too: https://reelang.com/open-startup/blog/how-to-build-a-simple-...)
There is also a reality that coding agents do have the ability to build more context than the average human translator with the average tooling like Smartcat or Weglot would.
The localized sites are exactly what they can judge as native speakers. They’re judging copy in their own language, not the language they’re learning.
We noticed that a good share of our learners were already using Google Translate built into their browsers to access the site in their native languages. The quality is atrocious, so what we’re offering is a step change from that.
We offer 58 languages to learners. Only 16 have been localized. That means the chrome of the website and apps -- the learner’s native language -- not the video content they watch in the target language they’re learning.
Between my co-founder and me plus people who have kindly offered to review the site we can judge seven of those localizations well enough to see that they’re orders of magnitude better than what Google Translate would concoct for them.
Between the two of us, we do speak 7 languages well enough to judge a localized site, and we have people around us for review on some more.
We run our workflow for a new localized version of say, German, then judge the quality and add rules to our prompt, guidance and voice guide for issues we spot, never hand-editing the actual strings.
We then re-run it until it's in a place where we feel this feels fairly native (feels transcreated, not calqued) and serves its purpose well).
Once we had that done in a good amount of locales, we started running it in languages we can't judge ourselves. Here we do depend on people on the app/website to spot and help us if we got it wrong somewhere, so it's a living process that evolves.
Here's an example of how that guidance and prompt evolved over time, in case you're interested: https://gist.github.com/tobyurff/2b461c259c34dfc6758a3932982...