485 karma · joined March 15, 2024
Email at b64 decode YWxleEBudWVua2kuYXBw
https://gchq.github.io/CyberChef/#recipe=From_Base64('A-Za-z0-9%2B/%3D',true,false)&input=WVd4bGVFQnVkV1Z1YTJrdVlYQnc
So those numbers are from an older version of the benchmark.
Coherence is done by:
- Translating English, to the target language, to English
- repeating three times
- Having 3 LLMs score how close the original English is to the new English
I like it because it's robust against LLM bias, but it obviously isn't exact, and I found that after a certain point it's actually negatively correlated with quality, because it incentivises literal, word by word translations.
Accuracy and Idiomaticity are based on asking the judge LLMs to rate by how accurate / idiomatic the translations are. I mostly focused on idiomaticity, as it was the differentiator at the upper end.
The new benchmark has gone through a few iterations, and I'm still not super happy with it. Now it's just based on LLM scoring (this time 0-100), but with better stats, prompting, etc. I've still done some small scale tests on coherence, and I did some more today that I haven't published yet, and again they have DeepL and Lingvanex doing well because they tend towards quite rigid translations over idiomatic ones. Claude 4 is also interestingly doing quite well on those metrics.
I need to sleep, but I can discuss it more tomorrow, if you'd like.
I'm curious, why electric motors vs a solid rocket motor? Volatility? Control over thrust? Making it safe to throw without worrying about backblast?
It is in the LLM comparison blog posts, at least the newer ones, though it tends to be on the first line.
I didn't end up sending many, but I've noticed that it's really difficult to get AI to write in a decent style. I've tried giving it a list of AI-isms to avoid, and it just doesn't work.
I has the most success with deepseek V3, giving a list of AI-isms, then ending with "You have been randomly assigned the following writing style/personality: [codeblock]" then a stereotype. Eg "Write in the style of a to-the-point, concise HN commenter" works alright, while "Write naturally and without AI-isms" is hopeless.
(Don't worry, I'm not using it for HN botting or whatever, it just tends to write in a nice style when you give it that)
I've documented a lot of my research into LLM translation at https://nuenki.app/blog, and I made an open source hybrid translator that beats any individual LLM at https://nuenki.app/translator
It uses the fact that
- LLMs are better at critiquing translations than producing them (even when thinking, which doesn't actually help!)
- When they make mistakes, the mistakes tend to be different to each other.
So it translates with the top 4-5 models based on my research, then has another model critique, compare, and combine.
It's more expensive than any one model, but it isn't super expensive. The main issue is that it's quite slow. Anyway, hopefully it's useful, and hopefully the data is useful too. Feel free to email/reply if you have any questions/ideas for tests etc.
I've been doing a lot of experiments evaluating LLM translation performance, and I used what I learnt (that LLMs make mistakes, but different LLMs make different mistakes, and they're better at critiquing translations than producing them) to make a hybrid translator (https://nuenki.app/translator) that beats everything else.
And I was invited to do a talk about that to a company, which was really cool! I'm 19, doing this in my gap year before uni.
I used it for some geography stuff, and that was fine I suppose, I like geography, but I stopped after a while.
I find it easy to remember things if I care about them. Ergo, if I care enough to put it in Anki, Anki is useless.
Particularly with LLMs being a thing, you don't even need to concretely know what you're looking for - just give chatgpt a vague description and let it list out suggestions until it jogs your memory.
When it comes to physics etc, you'll end up memorising everything relevant to the areas you actually use. In areas like "What is the star type of a star with a temperature of 6000K" - something that I memorised using Anki, and which took me a while - if I actually worked in that area of astrophysics I'd obviously learn that quite quickly.
I suppose it could be useful for maintaining knowledge. Had I not been burnt out I would have kept up my flashcards through my gap year, which would've presumably been quite useful - I'm currently going through the slog of relearning how to integrate so I'm not totally embarrassed at uni!
Edit: It's worth noting I had a nasty head injury that was slowing me down. Optimising my learning was a necessity, and the injury meant I spent more time studying than my peers, in more optimised and less enjoyable ways, to get the same result.
People will genuinely download top x wordlists for a language and try to learn from them. Hideous, but they do it.
I have tried different approaches, including using other people's flashcards (not as good - objectively they were high quality, but you gain a lot from writing your own + tailoring to your own way of looking at things) and learning from them (for my driving theory - terrible idea!). That hybrid approach is the best I've found, and the one I intend to use for my degree.
Many things. I think HN is a bit of a bubble here, but you'll find a lot of people prefer something enjoyable but slower to something efficient and faster, even if they won't admit it.
See the popularity of Duolingo vs Anki as an example! Or Quizlet vs Anki. Or the scores of students who revise by half-watching dopamine-ified youtube videos rather than doing past papers and flashcards. If you ask people, they'll often say they care for efficiency, but their revealed preferences say otherwise.
Doing large amounts (hours) of Anki day in day out is truly miserable, particularly when the alternatives can be quite enjoyable. And if you burn out before you achieve your goal, is the "efficiency" really worth it vs going slower but eventually getting there?
Plus, a lot of people want to learn e.g. a language because they enjoy the process as well as the end result. Making the process miserable in order to get to the end result faster isn't always a good tradeoff.
Which is what it's about. It's a tradeoff. I'm a big proponent of flashcards, but I think it's important to recognise that you're trading enjoyment for speed in most cases.
Keep that up every day and you'll burn out much faster with option 1 than option 2. Now, maybe you have enough motivation for that not to matter, or the self-discipline to keep going - as I did in my A levels - but don't be surprised if it kills your interest in the subject.
I personally used Anki the whole time, so if you're currently doing your exams some of my advice might not be super useful. I did maths, physics, and computer science. I didn't use flashcards much for maths - just for the irritating stats equations - but used it extensively for physics, and a little for compsci (I barely studied for compsci).
During my GCSEs I extensively used them for history, which is probably the closest analogue to the wordy questions you'll get in geography. I used it for facts, order of events, etc. I found that the process of organising history into a well-organised Obsidian database, then distilling it into flashcards, was as useful as the flashcard reviews itself. I recommend separating your rough working during class (e.g. the short essays you write at the end of a lesson) and your organised notes, which I split into separate, interlinked concepts with flashcards at the bottom of the file synced with Anki via an extension.
I suppose the advice I'd have is
- Cloze cards are excellent, and you should use them
- You can't flashcard your way to mental models. Absolutely don't rely on them alone, and you need to do practice questions for every separate question type you'll get until you're confident with the mental model itself.
- That said, it's easy to get into the trap of remembering the answers to flashcards as words. While this lets you "learn" quicker, and speed up reviews, I found that I had much better real-world results when I tried to actually "load in" the mental model into my head. So for example, if I had a flashcard about refraction behaviour, I'd not just answer the question, I'd also visualise a laser going from air into water and how the behaviour of the light changed as the angle changed.
For history, it's been a while, but if I had a question about one factor in a broader crisis (e.g. the Berlin Airlift) I'd try to think about the broader context of the question - not in my internal monologue, but just vaguely considering the various factors involved, the period of history, personally I instinctively visualise a map, etc, for a second or two before clicking for the answer.
Edit: Oh, and the heatmap extension is great. It gives you streaks and a heat map that you really don't want to break!
Now, those cards weren't alone - they were reinforcing content that I'd learnt in lessons. But if they were doing it for 10 minutes a day a few times a day, it seems quite plausible to me.
It about halved the amount of reviews I needed to do, and they didn't come up in bursts, so they were a lot more pleasant. I didn't quite believe it at first, and worried that it would be less effective, but it worked just as well if not better.
I really recommend giving it another try!
I credit Anki to my success at GCSEs and A Levels despite having a head injury, and I also credit it to me burning out so hard I took a gap year!
And I'm enjoying the gap year, but Anki made it a near necessity.
We spend hours a day browsing the web, so I made a browser extension[0] that translates sentences at your knowledge level into the language you're learning, so that you're always learning a little through immersion.
I also used the same "10 minutes a day on Anki" strategy with my A levels, and it made the revision process so so much nicer because stuff I'd learnt two years ago was as fresh as if I'd learnt it a couple of months ago, rather than years.
[Model name]-[Major version]-date-[minor variant/feature flags]
E.g. [Claude Sonnet] [3.7] [2025-03-02]:thinking
with a few differences based on the specific company.
Aside from some silliness (e.g. OpenAI's naming schemes; Anthropic deciding to go with "New Sonnet 3.5" rather than "Sonnet 3.6"), I think it works reasonably well.
Toucan did take off; it's fairly well known, and if their website is anything to go by they have hundreds of thousands of people using it.
It translates on a per-word basis, which means that the translations are often simply wrong due to a lack of context. Nuenki doesn't translate single word "sentences" by default, despite me spending a few days trying to improve the quality, because there just isn't enough context to go off of.
They try to mitigate that by only translating certain words, mostly nouns, which limits how much you can learn from it.
They also only have three difficulty levels, while Nuenki assigns a numerical score to each sentence's difficulty and has much finer grained control.
The flip side of that is that it's free, while Nuenki needs to pay for the cost of translation.
I've got ~20 paid subscribers. People on HN seem to love it, and most of my users are from here, while capital-L language learners are hard to market to. There's a lot of AI slop in language learning, and I'm really not good at marketing!
And yeah, I used the same data source, just heavily processed. It's a great project!
I like your approach. I wish English used neologisms more; I use them occasionally, and it's just quite fun to create a new, lexically valid word to describe something novel.
The Nuenki dictionary database is a bit over a gigabyte, albeit with ~30 languages in it, so yeah that's definitely a bonus! JSONL probably compresses quite well, too. I added compression to Nuenki's (serialised-struct) dictionary entries a while back and it reduced the size by about 30%, iirc.
But it's more the "It's not x; it's y" and "that thing saying z? doesn't mean much" format.
That said, I just double checked with some of the online LLM checkers and none agreed with me. Maybe I had a false positive.