485 karma · joined March 15, 2024
Email at b64 decode YWxleEBudWVua2kuYXBw
https://gchq.github.io/CyberChef/#recipe=From_Base64('A-Za-z0-9%2B/%3D',true,false)&input=WVd4bGVFQnVkV1Z1YTJrdVlYQnc
I've done a lot of research into LLM translation for my product[0], and I'm currently working on a deep translation service that provides reliably human-level translations.
I don't know what model you're using, but GPT-4.1 is probably the best for your use case - it's in the top few % for nearly every language, and has a low standard deviation, while also being relatively low latency and low cost.
The API itself is done. Right now I'm remaking the landing page so that it doesn't quote "look like it's from 2015". We'll see how it goes.
A friend of mine had made a spreadsheet of fusion 360 shortcuts, so I made a little webapp for use at school: https://fusion.alexcj.co.uk/
I made it use a custom stylesheet for printing, as they describe in the article, so that it produced a nice worksheet - though in retrospect, I probably ought to have made it denser when printed!
More realistically, though - the areas with the best lobbying and strongest unions.
I've had some breakthroughs with LLM translation, and I can now translate (slowly, unfortunately) at a far far higher quality than Opus, and well above DeepL. So I'm considering offering that as an API, though I don't know how much people actually care about translation quality.
DeepL's customers clearly don't care - their website is all about enterprise features, and they appear to get plenty of business despite their core product being mediocre.
Would people here be interested in that?
P.S. You might have more success with replies on Reddit. HN is very all-or-nothing.
At the moment I'm focused on translation quality, but I intend to add that.
"Translate Isolated Words" allows it to translate "sentences" of only one word, but it doesn't disable full sentences.
And yeah, atm it word splits by spaces for the dictionary. I hadn't thought to do it with LLMs, though that's a good idea. There's a somewhat related problem when doing Furigana, where it has a hashmap of strings-to-pronunciations, and it starts with a 4-character sliding window looking for matches, then a 3 character, etc.
- There's a global blacklist of sites, as well as phrases in the title/URL (e.g. "bank")
- You can blacklist sites yourself
- Each sentence is run against filters checking for medical/legal/etc info, as well as checks for addresses, card/social security numbers, etc. All the checks are done client side
- There are also some special implementations, e.g. it looks at the source code of websites to work out if they're an instance of an American health portal that I've forgotten the name of - each doctor's surgery self-hosts it.
- Websites can add `nuenki-ignore=true` on their end, if they'd like to disable it.
And of course it doesn't log anything, though there is an anonymous cache in order to make it economical.
I haven't added Anki integration, though. A few people have asked for it, but it's a big time investment for something relatively niche.
It's a browser extension that finds English sentences in webpages, and translates the ones at your difficulty level into the language you're learning.
https://nuenki.app/blog/the_more_llms_think_the_worse_they_t...
DeepL is a step up, and modern LLMs are even better. There's some data here[0], if you're curious - DeepL is beaten by 24B models, and dramatically beaten by Sonnet / Opus / https://nuenki.app/translator .
[0] https://nuenki.app/blog/claude_4_is_good_at_translation_but_... - my own blog
I also have a few heuristics (e.g. "I can't translate" in many different languages) to detect if it deviates from that.
It works pretty well.
I built something kinda similar, and made it open source. It picks the top x models based on my research, translates with them, then has a final judge model critique, compare, and synthesise a combined best translation. You can try it at https://nuenki.app/translator if you're interested, and my data is at https://nuenki.app/blog
I could improve the definitions, but that'd cost huge amounts of money in order to use proper dictionaries rather than Wiktionary, or I could add support for multiple languages at once, which two people have asked for but which would require a lot of dev time, and after that there isn't really much left to change.
And I've tried asking where I can find people, and they generally suggest "subreddits" (most ban self-promotion, and I've already posted on the ones that don't), and "discords" (all of them are very hostile to self-promotion).
I think the problem might be due to my landing page? People are far more likely to convert if I've already explained it to them. But I'm not sure how to convert a text explanation into the landing page form.
The hybrid translator is kinda a side thing based on my research into LLM translation quality.
It's a bit more complicated than using the best LLM, because it combines the results of the best ones, but yeah, that's broadly how it works. I made that part open source, anyway - I'm not trying to sell the hybrid translator, just use it as a marketing tool.
But I guess this comes back to the fact that I think my landing page might explain it poorly. I'm just not sure how to explain "It finds English sentences in webpages, filters out so that it's only the ones at your difficulty level, then translates them into your target language, so you're immersed while you browse" in the form of a ultra-low-attention-span landing page.
There's an Obsidian extension that lets you encrypt notes. I use that for my diary.
Firefox account containers are also quite nice.
Now, just a year later, DeepL is beaten by open models served by https://groq.com for most languages, and Claude 4 / GPT-4.1 / my hybrid LLM translator (https://nuenki.app/translator) produce practically perfect translations.
LLMs are also better at critiquing translations than producing them, but pre-thinking doesn't help at all, which is just fascinating. Anyway, it's a really cool topic that I'll happily talk at length about! They've made so much possible. There's a blog on the website, if anyone's curious.
It's co-published with an American group, though, and they also use 50.
It'd be nice if they were more explicit.
There are other learning tips for other areas (e.g. immersion for languages), but I think that one works quite well for STEM fields.
I also like to do organised notetaking via Obsidian, not because I read the notes super often but because the act of organising (and building flashcards while I'm there) is very effective at building understanding.
Your best bet is probably:
- Produce the sentence with Qwen 3. I didn't test down to 8B, but its 32B variant does reasonably well (see https://nuenki.app/blog/claude_4_is_good_at_translation_but_...) and Chinese models are better at Chinese in general
- Then prompt Qwen 3 again, this time telling it to critique the translation and improve it
LLMs tend to be better at post-critique than generation, though I can't say I've tested with models that small. You may find https://nuenki.app/translator interesting.
You might also be able to use some hideous distil of Deepseek V3?