13,143 karma · joined February 17, 2022
But I took Dark Forest to mean the Dark Forest Theory which the post I was replying to mentioned, and that's just a common theory under a new name, which has been a commonly known solution to the Fermi paradox for so long that no one is credited with coming up with it. It's given the "deadly probe" name in a SETI review of the Fermi paradox literature in 1984.
Yes this has held true on Opus 5.5. I checked. It’s a massively better model, peer to Fable but with different strengths and weaknesses. But it still has this issue. Which to be fair, people do too. Planning is a learned skill.
I think what they’re saying is that the harness no longer uses a text search on “think” to engage reasoning modes. Fair, that’s good to know. That doesn’t mean asking the model to think a certain way doesn’t have the intended effect.
Why? Because 4.6 actually talked like a human being. It actually organized its thoughts well, and got the main information across without the wall of text that makes your eyes glaze over. So from the perspective of human-computer interaction and maximizing the productivity of a developer+agent team, 4.7 and 4.8 were regressions. Despite much better benchmark performance.
Even if we consider autonomous agents, that benchmark is not indicative of how well they will interpret *your* requests. Or how well they will interact with other agents in a flock/swarm situation. The benchmark just doesn't cover this. (And the difference can be nontrivial! Sakana AI's published results show two generations of uplifting potential from better harnesses.)
What benchmarks are usually good at is showing to what degree new models are better than old models. What they are not good at, by construction, is showing that harnesses are well adapted to how people use them.
Nobody does what you are saying, because physics, which leads me to believe that you are confused.