DeepSeek-R1-Distill-Qwen-1.5B Surpasses GPT-4o in certain benchmarks
huggingface.co
huggingface.co
It still seems to me that these models are 'dumb' and often don't understand what I'm asking, where claude's intuition is much stronger.
I feel r1 14b even feels weaker than qwen 2.5 14b
Primary use-case is web technology / coding. Maybe I'm prompting it incorrectly?
O1 or even O3 might be able to crack academic level math problems, but I still wouldn't trust it to correctly fill out a McDonalds application using a PDF of my resume and a calendar of my availability.
It’s a bit like you get instruct tuned models and you get chat tuned ones. It’s not really one worse than the other just aimed at different uses
Vibes are important in this case...
So I would not put too much weight on how the models are doing on benchmarks.
Last I saw FrontierMath said they had a holdback set of problems specifically to ensure investors with access couldn’t cheat[1]. Or did that turn out to be a lie?
1. https://www.reddit.com/r/singularity/comments/1i4n0r5/this_i...
epoch is also very vague about how much access to the dataset openai had, but it's clear that they had more access than the general public: https://www.lesswrong.com/posts/8ZgLYwBmB3vLavjKE/some-lesso...
so yes, openai were caught gaming it.
OpenAI didn't train on the data or even use it to select a release candidate.
openai tried to hide their financial relationship with frontier. why do that if you're honest?
`ollama run hf.co/bartowski/DeepSeek-R1-Distill-Qwen-1.5B-GGUF:F16` and you're off to the races.