HNHacker News
TopNewBestAskShowJobs

stuess

2 karma · joined December 7, 2020

submissionscomments
stuess··on Show HN: OpenThat.Link – Trigger a webhook, open a link in your browser
Open source side project to serve a real need I had:

I use n8n a lot for automations, and sometimes it's helpful to have a human (myself) in the loop. New customers signs up? Open their LinkedIn profile automatically in my browser. New article in an RSS feed I care about, open it in a browser tab near real time.

Hope it's helpful for others! Can be self-hosted, and is built in a data-sparse way (minimal data retention/logging) so nothing is stored beyond what's required for the functionality.

stuess··on We let models localize into 16 languages. How we made it read native.
So our approach here to vet these:

Between the two of us, we do speak 7 languages well enough to judge a localized site, and we have people around us for review on some more.

We run our workflow for a new localized version of say, German, then judge the quality and add rules to our prompt, guidance and voice guide for issues we spot, never hand-editing the actual strings.

We then re-run it until it's in a place where we feel this feels fairly native (feels transcreated, not calqued) and serves its purpose well).

Once we had that done in a good amount of locales, we started running it in languages we can't judge ourselves. Here we do depend on people on the app/website to spot and help us if we got it wrong somewhere, so it's a living process that evolves.

Here's an example of how that guidance and prompt evolved over time, in case you're interested: https://gist.github.com/tobyurff/2b461c259c34dfc6758a3932982...

stuess··on We let models localize into 16 languages. How we made it read native.
The point I'm making isn't (in life or in this blog post) that LLMs do a better job than humans. My point is that LLMs can get pretty close with the right context, and in some circumstances they can build context better than the typical human translator would.

Most CAT tools do vet for ML/LLM translations these days, so "cheap" human translation is mostly ML-assisted and human-edited, but my entire point is that they're not to blame, the majority of translation fails happens because of the lack of context, not because the human didn't try. It's about setting the site and l10n infrastructure out in a way that gives them the freedom to translate in the way that makes most sense for their language without being forced into "English in a costume".

It is unfortunately a fact that one can be both things: a startup and a language learning app. But my entire point wasn't about LLMs vs. humans, it was about laying out the architecture in a way that enables great l10n, whether we do that with LLMs today or humans tomorrow once we can afford it.

stuess··on We let models localize into 16 languages. How we made it read native.
This is interesting! Exactly our experience, you do have to guide LLMs on author the source text in the way we would write, with few-shot examples in the prompt. And then prompt them to transcreate so it doesn't calque (read as a literal translation rather than something a native speaker would have written for other native speakers). Both of this can be solved for reasonably well with the right prompt guidance IMHO.

In my experience, there's two key incredients to get to a good results (with humans OR LLMs doing the work), I've worked on plenty of sites that had human l10n before LLMs came about.

If you're interested in the actual prompt we're using for our voice that includes instructions on how to avoid sounding like bad LLM copy, I've popped that here in a Github Gist: https://gist.github.com/tobyurff/2b461c259c34dfc6758a3932982...

Hope it helps, I might turn that into a blog post one day as well, seems like this could be useful to others!

stuess··on We let models localize into 16 languages. How we made it read native.
My point here is exactly that, as a startup, there's l10n that we can afford and l10n that's perfect. Could a human that has plenty of time to build context, browse the website, think about and learn our product do a better job at transcreating? They sure would.

But the reality for us (and many other startups) is that the options are no l10n and locking people who don't speak English (well) out of our product or having one that we - at least - do our best to make feel pretty native. We also built a version in Simple English. (I wrote an article sharing that process, too: https://reelang.com/open-startup/blog/how-to-build-a-simple-...)

There is also a reality that coding agents do have the ability to build more context than the average human translator with the average tooling like Smartcat or Weglot would.

stuess··on We let models localize into 16 languages. How we made it read native.
Sharing this in hopes it helps other founders who are scared l10n is expensive/bad quality when using Claude Code and other coding agents, it doesn't have to be.