I'm trying to square the excitement over DeepSeek with its good -but not dominant- performance in evals.
I'm trying to square the excitement over DeepSeek with its good -but not dominant- performance in evals.
The excitement is probably a bit much but it's not just about the eval results themselves but the baggaged attached with them.
I was planning on spending the $200 for a month but had been thinking of prompts to try it out.
DeepSeek already answered them all for free so I am not just going to light $200 on fire for fun.
Also R1 (and its distilled models) expose their CoT & web interface has a websearch option too.
With the 14b distilled models, I found multiple math-related prompts where it gives the right answers almost immediately but then wastes 10 minutes making self-verification mistakes (e.g. "Write Python3 code that computes the modular inverse of a mod 2^32")