850 karma · joined November 24, 2016
It debuted as ~same score as Sol on Artificial Analysis. People couldn't accept it so they had to change the formula.
The model is a big step forward only in desktop use and 3D. That's impressive, but for software engineering, Fable is still in a league of its own.
I'm genuinely so confused when people say this with a straight face. Are you talking about coding? Desktop use? Prose? Or something else?
Sol is a much smaller models and it shows. It often misses the forest for the trees.
Still below Fable 5, let alone Fable 5.1.
EDIT: This is suspiciously low. Calls the relevance of existing benchmarks into question.
That's not clear. Need to see independent benchmarks first.
https://totalrealreturns.com/n/MRNA,SPY
MRNA has vastly outperformed the market since IPO.
Do you have a formal proof of that?
At this point, Anthropic only needs to release models to the public when the competition forces them to.
OpenAI also has a better model (Astra) that they haven't released yet.
This may be a narrative violation on HN, but the quality of Google's management is exceptional.
So it's not 6 months but it's also not a few weeks.