596 karma · joined January 12, 2020
If you have experience with tiny ~1B param models, its still head and shoulders above anything that has come before. IMO there have not been any other quantized/distilled/etc models as good at this size. It would not exist without the original R1 model work.
But, its free and open and the quant models are insane. My anecdotal test is running models on a 2012 mac book pro using CPU inference and a tiny amount of RAM.
The 1.5B model is still snappy, and answered the strawberry question on the first try with some minor prompt engineering (telling it to count out each letter).
This would have been unthinkable last year. Truly a watershed moment.
Years later, "man we tried, we had that meeting and everything, we just couldn't compete"
Picks up goalpost, looks for stadium exit
"Well, yeah, but its kind of expensive" -- this guy
The only thing Sam did wrong was play too fast and loose with the implication of "Her" given that he had been talking to ScarJo. Lawyers should have told him to just STFU about it.
The human brain is a marvel of mechanical and electrical engineering, not a mystical, otherworldly device. Maybe there is something going on at a quantum level to explain the incredible performance our brains achieve, but that is still nothing that we can't build.
(assuming companies won't easily share all their models for this kind of effort)
We want AI to hoop jump something we don't even understand ourselves. The only empirical evidence we have is as we increase compute, the results get better.
These are non-scientific concepts. You are basically saying "humans are doing something more, but we can't really explain it".
That assumption is getting weaker by the day. Our entire existence is a single, linear, time sequence data set. Am I "extrapolating from my experience" when I decide to scratch my head? No, I got a sequential data point of an "itch" and my reward programming has learned to output "scratch".
That is not to say this is fine, but more that we tend to get hung up on what these models do wrong rather than all the amazing stuff they do correctly.
I guess it might be surprising to some that code from HF can have exploits just like GH.