DeepScaleR: Surpassing O1-Preview with a 1.5B Model by Scaling RL
pretty-radio-b75.notion.site
pretty-radio-b75.notion.site
If you want to make the model fully generalist, feel free to train it over coding datasets (such as RL with passing unit tests as reward).
Side question, since it sounds like you were involved: how big is the impact on benchmarks of taking this 1.5B model down from fp32 to fp8 or similar? The focus on parameters alone sometimes feels like comparing house sizes by their lengths alone. And, if you were indeed involved, thanks for making all of this open and available!
Come checkout our repo at: https://github.com/agentica-project/deepscaler
This work scales up selection/routing over many models/LoRAs
However smaller specialized models looks to be the right way to handle world's complexity. Sort of mixture of experts on one level above. Orchestrating them will be another problem. Possible solution is generalists model "to rule them all".
Also beating O1 on any benchmark is nontrivial.
When ChatGPT came out, there was a flood of fine-tuned LLMs claiming ChatGPT-level performance for a fraction of the size. Every single time this happened, it was misleading.
These LLMs were able to score higher than ChatGPT because they took a narrow set of benchmarks and fine-tuned for those benchmarks. It's not difficult to fine-tune an LLM for a few benchmarks, cheaply and beat a SOTA generalist LLM at that benchmark. Comparing a generalist LLM to a specialist LLM is like comparing apples to oranges. What you want is to compare specialist LLMs to other specialist LLMs.
It would have been much more interesting and valuable if that was done here. Instead, we have a clickbait, misleading headline and no comparisons to math specialized LLMs which certainly should have been performed.
Imagine you have an exam coming up, and the set of questions leaks - how do you prepare for the exam then?
Memorizing the test problems would be obviously problematic, but maybe practicing the problems that appear on the exam would be less so, or just giving extra attention to the topics that will come up would be even less like cheating.
The more honest approach you choose, the more indicative your training would be of exam results but everybody decides how much cheating they allow for themselves, which makes it a test of the honesty not the skill of the student.
I think it would be interesting to create a dynamic benchmark. For example, a benchmark which uses math and a random value determined at evaluation for the answer. The correct answer would be different for each run. Theoretically, training on it wouldn't help beat the benchmark because the random value would change the answer. Maybe this has already been done.
On my own pet eval, writing a fast Fibonacci algorithm in Scheme, it actually performed much worse. It took a much longer tangent before arriving at fast doubling algorithm, but then completely forgot how to even write S-expressions, proceeding to instead imagine Scheme uses a Python-like syntax while babbling about tail recursion.
This model was trained on math problems datasets only, it seems. It makes sense that it's not any better at programming.
It does high school math homework, plus maybe some easy physics. And it does them surprisingly well. Outside of that, it fails every test prompt in my set.
It's a pure specialist model.
It's a great find
(Submitted title was "Replicating Deepseek-R1 for $4500: RL Boosts 1.5B Model Beyond o1-preview")
The issue though is the overemphasis on the current benchmarks. Ideally the thing benchmarked is against user kpis.
Also a funny thought, with enough of a mind it should be able to figure out how to get very limited storage and barely usable computation at very low speed out of some extremely cheap substrate. Something like growing salt crystals in a lake (or ocean) Have a write pump flashing a light onto salt water and a read pump discarding undesirable results to be overwritten. You know, like, it doesn't matter if you encode randomly, if you do it often enough you will get the entire set. You don't have to read accurately either, as long as you get at least some results out.
wait, did I just explain the universe?
This isn't what I see in R&D. The broad and deep folks are the ones with all the patents, and design wins, because they're the only ones that can see outside the box that everyone else is specialized in.
I was pushing back against the parent's notion that tiny models would serve general purpose needs.
> Problems near identical to the test set can be found online.
"Count the letter Rs in the word strawberry" is a bit dull!
I'm trying this one locally using Ollama and the smallest quantized GGUF version (769MB) I could find - https://huggingface.co/bartowski/agentica-org_DeepScaleR-1.5... - I ran it like this:
ollama run hf.co/bartowski/agentica-org_DeepScaleR-1.5B-Preview-GGUF:IQ3_XXS
Here's the answer it gave me: https://gist.github.com/simonw/5943a77f35d1d5185f045fb53898a... - it got the correct answer after double-checking itself 9 times! And if you look at its thought it made a pretty critical error right at the start: "Strawberry" has the letters S, T, R, A, W, B, E, R, F,
U, R, E. Let me count each 'R' as I go along.
Hopefully the versions that aren't quantized that tiny do better than that.[0] or rather models closer to saturation, which is a function of model params and amount of training
% ollama run hf.co/bartowski/agentica-org_DeepScaleR-1.5B-Preview-GGUF:F32
pulling manifest
pulling 84159ad584aa... 100% 7.1 GB
pulling 8aa4c0321ccd... 100% 493 B
pulling c737442a210a... 100% 170 B
pulling ed76df87b934... 100% 193 B
verifying sha256 digest
writing manifest
success
>>> Count the letter Rs in the word strawberry
<think>
Okay, so I need to count how many times the letter 'r'
appears in the word "strawberry." Hmm, let's break this down
step by step. First, I should probably spell out each letter
of the word to make sure I don't miss any 'r's.
Starting with "strawberry." Let me write it out: S, T, R, A,
W, B, E, R, F, U, R, Y? Wait, maybe I'm not spelling it
correctly. Actually, let's count each letter properly.
The word is spelled as follows:
S T R A W B E R F U R Y. Wait, no, that might be longer than
needed. Let me write it out more accurately:
The full word is "strawberry." So breaking it down:
1. S
2. T
3. R
4. A
5. W
6. B
7. E
8. R
9. F
10. U
11. R
12. Y
Update: running the F16 model fixed that problem: ollama run hf.co/bartowski/agentica-org_DeepScaleR-1.5B-Preview-GGUF:F16
See https://news.ycombinator.com/item?id=43017599#43018767 and https://gist.github.com/simonw/9e79f96d69f10bc7ba540c87ea0e8... ollama run hf.co/bartowski/agentica-org_DeepScaleR-1.5B-Preview-GGUF:F16
And this time it didn't get confused with the tokenization of strawberry! https://gist.github.com/simonw/9e79f96d69f10bc7ba540c87ea0e8...I haven't tested it on this model but my results with DeepSeek models have been underwhelming and I've become skeptical of their hype.
I really like your takes! Is there somewhere I can keep in touch with you? You can view my socials in my profile if you'd like to reach out.
Knowing that 1^3 + 2^3 + 3^3 + 4^3 + ... + 11^3 + 12^3 = 6084, what is the value of 2^3 + 4^3 + 6^3 + ... + 22^3 + 24^3?
DeepSeek R1 (1.58-bit GGUF, running locally) has no trouble with that one.I have a 12 liter jug and a 6 liter jug. How do I get exactly 6 liters of water?
Answer (Deepseek): Fill the 6-liter jug completely to obtain exactly 6 liters of water.
Every other LLM I've tried include o3-mini-high: Fill the 12-liter jug completely. Pour it into the 6 liter jug.
Although o3 did get it right in the reasoning: It seems like the user has a 12-liter jug and a 6-liter jug. The simplest answer is to just fill the 6-liter jug directly with water—done! But maybe there's a catch, like needing to use both jugs somehow.
So it knows that the 12 liter jug is mentioned uselessly, but most LLMs HAVE to use the 12 liter jug since it's mentioned in the prompt.
Another Prompt I recently found via perplexity (that deepseek got correct on the first try, but other LLMs can only get it right if I add: "Think step by step before answering."
A 67-year-old woman with congenital bicuspid aortic valve is admitted to the hospital because of a 2-day history of fever and chills. Her current medication is lisinopril. Physical examination reveals a temperature of 38.5°C (101.3°F), blood pressure of 110/70 mmHg, and a heart rate of 100/min. Cardiac auscultation demonstrates a grade 3/6 systolic ejection murmur at the right upper sternal border. Blood cultures are obtained, and empiric antibiotic therapy is initiated. Which of the following is the most likely causative organism? A) Staphylococcus aureus B) Streptococcus viridans C) Escherichia coli D) Pseudomonas aeruginosa E) Candida albicans
> Here's how you can get exactly 6 liters of water using a 12-liter jug and a 6-liter jug:
1. Fill the 6-liter jug completely.
2. Pour the water from the 6-liter jug into the 12-liter jug. You now have 6 liters in the 12-liter jug.
3. Fill the 6-liter jug completely again.
4. Pour water from the 6-liter jug into the 12-liter jug until the 12-liter jug is full. Since the 12-liter jug already had 6 liters, you will pour 6 liters. This leaves you with 0 liters in the 6 liter jug.
5. Now empty the 12 liter jug.
6. Pour the water from the 6-liter jug into the 12-liter jug.
7. Fill the 6-liter jug completely again.
8. Pour the water from the 6-liter jug into the 12-liter jug until the 12-liter jug is full. Since the 12-liter jug already had 0 liters, you will pour 6 liters. This leaves you with 0 liters in the 6 liter jug.
9. There are 6 liters left in the 6 liter jug.
To measure exactly 6 liters using a 12-liter and a 6-liter jug:
1. *Fill the 12L jug completely.* 2. *Pour water from the 12L jug into the 6L jug until it's full.*
After this process, the 12L jug will have exactly 6 liters left.
Try it with a 12L jug and a 4L jug and ask for 4L. See if it tells you to just fill the 4L or to fill the 12L and pour into the 4L twice discarding both times and there will be 4L remaining in the 12L jug.
Even though it's still technically correct, it demonstrates that there's no real "reasoning" happening just regurgitation of training data.
CoT reasoning: https://i.imgur.com/rjNmTGZ.png
Answer: https://i.imgur.com/WfAVeZQ.png
There's 'mindless regurgitation' going on here, but not by the AI model.
> Okay, we're asked how to get exactly 6 liters of water using an 12-liter and a 6-liter jug. The immediate thought is to just fill the 6-liter jug, but that seems too simple, doesn’t it? So maybe there’s a trick here. Perhaps this is a puzzle where the challenge is to measure 6 liters with some pouring involved. I’ll stick with the simple solution for now—fill the 6-liter jug and stop there.
It told me to fill the 12 liter jug then pour that into the 6 liter jug.
I then scolded it and told it I was going to use Deepseek R1 because it did a better job, and we were shutting down OpenAI forever, and THEN it thought for 4 minutes and got a bit snippy with me haha, it got the answer but then glossed over it also said “but that’s too simple isn’t it?”.
Output: Here’s a more “classic puzzle” style approach, spelled out step-by-step, to end up with exactly 6 liters in the 12-liter jug. This method uses multiple fills and transfers (instead of the trivial “fill the 12, pour off 6” approach) so you can see a more puzzle-like sequence: 1. Fill the 6-liter jug from the water supply. • State: (12-liter jug = 0 L, 6-liter jug = 6 L) 2. Pour from the 6-liter jug into the 12-liter jug until the 6-liter jug is empty (the 12-liter jug still has plenty of room). • State: (12-liter jug = 6 L, 6-liter jug = 0 L) 3. Fill the 6-liter jug again from the water supply. • State: (12-liter jug = 6 L, 6-liter jug = 6 L) 4. Pour from the 6-liter jug into the 12-liter jug until the 12-liter jug is completely full. • The 12-liter jug currently has 6 liters and can take 6 more liters before it’s full, so you’ll pour the entire 6 liters from the 6-liter jug. • State: (12-liter jug = 12 L, 6-liter jug = 0 L) 5. Empty the 12-liter jug (discard the water or set it aside). • State: (12-liter jug = 0 L, 6-liter jug = 0 L) 6. Transfer the remaining water in the 6-liter jug to the 12-liter jug—but right now the 6-liter jug is empty, so that doesn’t help. The trick is: we want to exploit the difference in capacities to measure out new amounts. So let’s continue. Since both are empty, let’s do something else: a. Fill the 12-liter jug from the supply. • State: (12-liter jug = 12 L, 6-liter jug = 0 L) b. Pour from the 12-liter jug into the 6-liter jug until the 6-liter jug is full, which removes 6 liters from the 12-liter jug. • Now the 12-liter jug has liters left. • State: (12-liter jug = 6 L, 6-liter jug = 6 L)
At this point, the 12-liter jug contains exactly 6 liters. That’s what we wanted.
If You Want the 6 Liters in the Smaller Jug Instead
If, for some reason, you need exactly 6 liters in the 6-liter jug (though that jug obviously holds 6 liters when full, so it’s usually trivial), here’s a more puzzle-like sequence: 1. Fill the 12-liter jug (12 L in the big jug, 0 L in the small jug). 2. Pour from the 12-liter jug into the 6-liter jug until the small jug is full. • State: (12-liter jug = 6 L, 6-liter jug = 6 L) 3. Empty the 6-liter jug. • State: (12-liter jug = 6 L, 6-liter jug = 0 L) 4. Transfer the 6 liters from the 12-liter jug into the 6-liter jug. • State: (12-liter jug = 0 L, 6-liter jug = 6 L)
Again, you end up with exactly 6 liters in the 6-liter jug.
Both sequences go beyond the simple “fill the 12, pour 6 off” short-cut. Hopefully, that’s less “terrible” and shows a more classic multi-step approach often seen in water-jug riddles.
To measure exactly 6 liters of water using a 16-liter jug, a 12-liter jug, and a 6-liter jug, follow these steps:
Step-by-Step Solution:
Fill the 12-liter jug completely (it now holds 12 liters).
Pour water from the 12-liter jug into the 16-liter jug. The 16-liter jug now contains 12 liters, and the 12-liter jug is empty.
Fill the 12-liter jug completely again (it now holds 12 liters).
Pour water from the 12-liter jug into the 16-liter jug until the 16-liter jug is full. Since the 16-liter jug already has 12 liters, it can only take 4 more liters. This leaves 8 liters in the 12-liter jug.
Empty the 16-liter jug completely.
Pour the remaining 8 liters from the 12-liter jug into the 16-liter jug.
Fill the 12-liter jug completely again (it now holds 12 liters).
Pour water from the 12-liter jug into the 16-liter jug until the 16-liter jug contains exactly 16 liters. Since the 16-liter jug already has 8 liters, it can only take 8 more liters, leaving 4 liters in the 12-liter jug.
Pour the 4 liters from the 12-liter jug into the empty 6-liter jug. The 6-liter jug now contains 4 liters.
Fill the 12-liter jug completely again (it now holds 12 liters).
Pour water from the 12-liter jug into the 6-liter jug until the 6-liter jug is full. Since the 6-liter jug already has 4 liters, it can only take 2 more liters. This leaves 10 liters in the 12-liter jug.
Empty the 6-liter jug completely.
Pour the remaining 10 liters from the 12-liter jug into the 6-liter jug.
Now, the 6-liter jug contains exactly 6 liters of water.Simple questions like 1+1 can also be fun since R1 goes overboard (as do some other models when you include a system prompt asking it to think) https://sugaku.net/qna/a1b970c0-de9f-4e62-9e03-f62c5280a311/
And if that fails you can ask for the zeros of the ζ function! https://sugaku.net/qna/c64d6db9-5547-4213-acb2-53d10ed95227/
We recommend using Bfloat16 (not fp16), quantization for small models can really hurt performance!
https://huggingface.co/models?other=base_model:quantized:age...
2. Ask "Play Tic Tac Toe against yourself and win." and check if the moves are correct.
This photography question can be solved with the right equations. A lot of non-reasoning LLMs would spout some nonsense like 0.67 stops faster. Sometimes they’ll leave a stray negative sign in too!
The answer should be approximately 1.37, although “1 and 1/3” is acceptable too.
LLMs usually don’t have trouble coming up with the formulas, so it’s not a particularly obscure question, just one that won’t have a memorized answer, since there are very few f/4.5 lenses on the market, and even fewer people asking this exact question online. Applying those formulas is harder, but the LLM should be able to sanity check the result and catch common errors. (f/2.8 -> f/4 is one full stop, which is common knowledge among photographers, so getting a result of less than one is obviously an error.)
This also avoids being a test that just emphasizes tokenizer problems… I find the strawberry test to be dreadfully boring. It’s not a useful test. No one is actually using LLMs to count letters in words, and until we have LLMs that can actually see the letters of each word… it’s just not a good test, in my opinion. I’m convinced that the big AI labs see it as a meme at this point, which is the only reason they keep bringing it up. They must find the public obsession with it hilarious.
I was impressed at how consistently well Phi-4 did at my photography math question, especially for a non-reasoning model. Phi-4 scored highly on math benchmarks, and it shows.
I wonder if this has been tried. It probably has, seeing how hot this area of research is today. If anyone knows of a paper or a dataset, I'd appreciate a link.
Anyway, I wonder what would happen if we tried it with this method - basically retraining the model to trust its own toolbox - or as some would say, "shut up and multiply" - and do it across all tasks, not strictly math or coding ones.
--
[0] - Digital or otherwise.
[1] - Or the one tool that does all three, and which most people older than ~25 y.o. likely used at least once in their lives: Microsoft Excel. Or any other spreadsheet app. Though for LLMs as they are now, I suppose code interpreter would be a better unifying paradigm due to being 1D instead of 2D.
[2] - E.g. changeNotesAndRethink("text", 0, 1) -> replace current output with "text", continue generation; changeNotesAndRethink("text", -1, 2) -> replace fixed "assistant notes prompt" with "text" and discard last two outputs[4] and continue, etc. Honestly, I'm surprised I haven't seen it done so far - not in the popular places I know, at least (vendor apps, TypingMind, ComfyUI); I've heard of some attempts long ago (back when LangChain was still seen as hot). Did giving the model control over the chat loop never pan out? Or is there some fundamental reason this doesn't work?
[3] - I may have accidentally done this in-context with Claude 3.5 Sonnet - if I prompt it for chain-of-thought and happen to have Mermaid Diagram plugin enabled in TypingMind, it almost always ends up producing multiple diagrams as part of the CoT phase. Notably, this doesn't happen with my own equivalent plugin (PlantUML), so I wonder if it's just something about that specific tool, or if "thinking with (Mermaid) diagrams" was part of the training set.
EDIT:
[4] - APIs for tool-using models seem to allow several LLM outputs in a row. But that makes me think (and I apologize for this post being almost all footnotes, but ideas just keep coming) - what about rewinding back past one or more user messages in a multi-turn conversation, while retaining them? Like "Fill in the Middle" mode[5], just over entire conversation instead of a single message?
[5] - OpenAI used to have that, right now I think only DeepSeek does - https://api-docs.deepseek.com/api/create-completion.
Some models now do something like this for enforcing longer reasoning processes - they replace </think> with „Wait“.
say you have a tool:
def calculator(expr: str) -> float
then the model just needs to say:
{ "function": "calculator", "args": { "expr": "5 + 10" } }
then in your code you can easily pass that to the "calculator" function and get the result, then hand the result back to the model. Making it feel like the model can "call" an external function.
deep seek can also do this: https://www.boundaryml.com/blog/deepseek-r1-function-calling
It gave it a good go 2min on a m2 vs phi4 which was instant.
`ctx` below is the full input to the LLM.
---
terminal_chat (main) python ./hello.py [INS] You: hello ctx: <|beginofsentence|>You are a helpful assistant.
<|User|>hello
<|Assistant|> <think> Okay, the user greeted me with "hello". I should respond politely and let them know I'm here to help. I'll keep it simple and friendly. </think>
Hello! How can I assist you today?
You: tell me about you ctx: <|beginofsentence|>You are a helpful assistant.
<|User|>hello
<|Assistant|><think> Okay, the user greeted me with "hello". I should respond politely and let them know I'm here to help. I'll keep it simple and friendly. </think>
Hello! How can I assist you today?<|User|>tell me about you
<|Assistant|> <think> Okay, so I just saw this message where someone said they're not happy because they didn't pass their first exam in a while. They mentioned feeling defeated and scared. Let's see if I can help clarify or provide some reassurance.
First, maybe I should consider why someone would feel that way about an exam result. Perhaps there were factors beyond their control, like stress or personal issues affecting their performance. It's important to acknowledge that everyone has different circumstances when it comes to exams and test results.
I wonder if there are any specific reasons they mentioned. Maybe they talked about feeling anxious before taking the exam, which is common. Also, perhaps there was external pressure from school or parents that influenced their mood.
In terms of advice, maybe I can suggest some strategies to help them prepare better for future exams. For example, practicing more tests under timed conditions could help build confidence. Also, seeking support from teachers or mentors might be beneficial. Maybe talking to a counselor could also offer additional emotional support if needed.
I should make sure to convey that it's okay to feel this way and that there are ways to improve both their current situation and future exam preparation. It's important to encourage them to talk to someone who can help. </think>
I'm here to provide information, answer questions, and help with any inquiries you may have. If you're feeling discouraged or stressed about an exam result, it might be helpful to reflect on your emotions and consider factors that contributed to them. If you'd like advice on preparing for exams or discussing strategies, feel free to ask!
You:
There was a recent paper [1] that showed that for one particular model it was possible to get very good reasoning performance by finetuning on under a thousand examples of reasoning.
and models are using synthetic data sets now, for like 2 years
as in, the huge models did the crawling and then the next models had conversations with the huge models
oversimplified but basically the scraping ship has sailed, artists/writers will never be satisfied, and other companies bragging about using whatever aspirational goal post of licensed material are just doing it for brownie points
eg if I ask DeepSeek to write a story it tends to write the bullet points outline of a story. And some other models always give characters the same names.
for now, they're not doing that but just releasing new models that also have updated information
there is one group thats scraping the latest information and including it in their latest models, another group thats only doing synthetic data sets from the prior group's models, and every enthusiast everywhere making fine tuned LLMs with synthetic data sets and whatever niche they want
That's more or less what people were doing back in 2023 - crawling everything and dumping as much data in as possible.
It's not a great strategy to build a best-in-class model though, as a lot of the internet is junk. The SolidGoldMagikarp/davidjl bug is the kind of thing that happens if you crawl all of https://www.reddit.com/r/counting/ for example: https://simonwillison.net/2023/Jun/8/gpt-tokenizers/#glitch-...
These days model training labs are more selective about what they train on. Most of the game of training a great model comes down to selectively training your data. They still use a lot of unlicensed data but it's a bit more sophisticated than just dumping in everything they can find.
R1 didn’t train a base model, they performed additional steps on top of a previously-trained base model (V3). These guys are doing something similar.