Needless to say, for practical purposes am back to 5.6/Opus 4.6-4.8. But hey, maybe I am not smart enough to use LLMs?
Needless to say, for practical purposes am back to 5.6/Opus 4.6-4.8. But hey, maybe I am not smart enough to use LLMs?
If we look at the math problems they're solving their just now reaching the human frontier... they weren't doing that before.
And your comparison point is model released 2.5 months ago... saying for some use case you didn't see noticeable improvement in 2.5 months (even while other people and benchmarks disagree) isn't a great argument that they aren't improving.
Success in such problems does not automatically extrapolate to other contexts.
Jane Street is apparently one of Anthropic's biggest customers. Probably engineering, finance, and some math.
That's the thing I'm most worried about - LLMs that are super clever at coding and maths, making an actually very very dangerous model that is far more efficient, and clever in a more innate (less brute force) way.
Kind of a tangent, but one thing I am curious about is to what degree the Navier-Stokes result announced today was primarily a brute-forced result based on the 'program' previously established by researchers to find counterexamples (blowups), or whether the model actually added significant/novel intellectual value beyond its ability to run at arbitrary parallelism. With 10K agents and a staggering $15M in compute (IIRC), I am feeling like a lot of the former may have been involved, but I don't really understand either the problem or the approach (or, indeed, the solution).
Obviously the potential for parallelism and coordination between so many agents is quite scary by itself, but I think brute force by 10K mediocre AI mathematicians is much less scary than ~one AI mathematician reasoning its way through the problem where all human attempts have failed. It seems fairly obvious that massive parallelism lends itself to brute-force counterexample-finding, and I suspect it isn't a coincidence that most of the touted AI math results have been counterexamples.
It's all still quite scary, but coming full circle: I really don't know what to think.